Runs 3 and 6 were killed by me restarting the runner mid-job, not by any
defect in the workflow. Config verified on disk first this time.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The migration step failed with 'could not translate host name postgres'. The
Postgres service container was healthy; the job simply could not name it,
because service DNS aliases only exist on a per-job network and I had pinned
container.network to the flat 'bridge'.
That choice was wrong in both directions: it broke service containers AND it
was weaker isolation, since every concurrent job shared one bridge and could
see its neighbours. A per-job network is stricter and correct — and still has
no route to the forge, because these networks live inside the dind daemon,
which has no forge attachment at all.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Trading one wedge for another: running the job in python:3.12-bookworm fixed
the resolver spiral but broke checkout, because actions/checkout is a
JavaScript action and the official Python images carry no node — 'executable
file not found in $PATH'.
setup-python on the act image has both. Test now accepts either a pinned
container image or an explicit setup-python version, and still checks it
against pyproject's requires-python.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first real CI run wedged for 14 minutes. Not a network problem, not the
isolation work — catthehacker/ubuntu:act-22.04 ships Python 3.10.12 while this
project declares requires-python >=3.12, and pip answered that by backtracking
through the entire release history of every dependency looking for something
3.10-compatible. At 100% CPU, with -q hiding every line of it, and it would
have churned until the runner's 30m timeout.
A version mismatch presenting as a hang rather than an error is worth a test,
so there is one: the workflow's python image must satisfy pyproject's
requires-python, checked by parsing both rather than by eyeballing them.
Also: every job now has timeout-minutes. A wedged step should be a red check in
minutes, not an occupied runner for half an hour.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first CI run failed with 'Could not resolve host: gitea' — job containers
live on dind's private network and cannot see the forge network. Two ways out,
and they are not equivalent:
(a) put job containers on the forge network. Easy, one line, and it leaves
untrusted workflow code one DNS name from the forge's Postgres. It quietly
repeals I-5.
(b) send jobs to the PUBLIC forge surface over the tunnel, exactly like any
stranger on the internet.
Took (b). The runner no longer needs the forge network at all, so there is now
NO private route from any CI container to anything — a better posture than this
file started with. Cost is a hairpin through Cloudflare plus its ~100s ceiling
per fetch, which for 0.63 GB of objects across 61 repos and depth=1 checkouts is
nowhere near binding.
Test upgraded to assert the stronger property: no CI container joins the forge
network.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
I-5 says runners execute untrusted code and must be isolated by machine
boundary. act_runner needs a Docker daemon to start job containers, and the
tempting move — what every published example does — is to mount the host's
/var/run/docker.sock. That hands every workflow, including whatever a
transitive dependency's postinstall script feels like doing, the ability to
start a privileged container mounting / — root on Grant's workstation.
Instead the runner talks to its OWN dind daemon:
- runner (TRUSTED, the act_runner daemon) sits on the forge network only to
collect jobs from gitea:3000
- dind and every job container it spawns are UNTRUSTED, on a private network
with no route to the forge, its Postgres, or its .env
- jobs cannot bind-mount from the daemon host (valid_volumes: []) and are not
handed the runner's own socket (docker_host: -)
- separate compose project, cpu/memory bounded — Veron 1 is Grant's
workstation, not a dedicated build box
The gate itself now runs as a workflow, including the migration round-trip that
already caught two bugs review did not, and the I-12 check that a COMMIT_SHA
env override cannot change what /version reports.
Labels are explicit and pinned. A workflow naming a label nobody provides
queues forever and presents as a hung CI system rather than a typo — which is
what ubuntu-latest does on every windy-registry run today.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Client 'windy-git' registered on account-server; Gitea auth source 'windy'
added against the discovery document. Auto-registration on, so a Windy account
IS the account — nobody is asked to invent a second identity for the same
person and no local password ever exists.
Proven end to end:
- app.windygit.com/user/login offers 'Sign in with windy'
- /user/oauth2/windy -> 307 to account.windyword.ai/oauth/authorize
- authenticated authorize -> 200, issues a code to the registered callback
- a bogus evil.example.com redirect_uri -> 'redirect_uri not registered for
this client'. The anti-phishing check works, which is the whole reason
redirect URIs are registered rather than accepted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Not a debt, a decision: sandbox phase, months from launch, and minting a tenth
Cloudflare token to sit in the inventory costs more than it buys. A
platform-specific scoped token is a launch-hardening item.
Recorded so the next reader knows it was chosen rather than missed, with a
do-not-re-raise note. Keeps the genuinely non-obvious part: R2's S3 credentials
are DERIVED from a CF API token — access key id = the token's id, secret =
SHA-256 of the token value.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Gitea reports the epoch for 'not yet synced', which arithmetic turns into a
56-year lag and a confident 'degraded'. Collapsing those two states is how a
backup that was never made gets read as a backup that is merely stale — which
is the more dangerous direction, because 'behind' sounds survivable.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
I-4 said 'never a one-way door' and had no implementation. Now it does.
- ensure the GitHub counterpart exists (idempotent), then ask Gitea to keep
it in step with sync_on_commit=True. An hourly timer means an hour of work
can be the thing you lose, and that window is invisible until it costs you.
- mirror status reports what is TRUE including 'we do not know'. An
unconfigured mirror reports unconfigured, NEVER healthy — same posture as
me-fleet.ts refusing to say 'online' when it only knows 'registered'.
- lag past the threshold is a P2, not a shrug. A mirror nobody checks is a
belief, not a backup, and this ecosystem already lost 37 days to a canary
everyone assumed was fine.
Gitea owns the replication rather than a hand-rolled loop, because a background
job that fails silently is exactly how the registry's integrity refresh spent
its entire life calling a 404 and incrementing a counter instead of raising.
Also fixes a real bug I had written myself: list_versions derived the Gitea
namespace from the CALLER, which is correct only while the caller is the owner
and addresses the wrong namespace the moment a collaborator asks — surfacing as
'not found', which is the hardest kind of bug to see. Now derived from the repo,
with a test that keeps it that way.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
'There are 1 saved versions' is exactly the sloppiness the vocabulary law
exists to catch. Copy is design material, not decoration.
G5.9 records something the live test surfaced: a repo created through
X-Service-Token lands in a 'u-system' namespace because the service caller has
no identity of its own. Correct for /internal plumbing, WRONG for anything a
person owns — the portal must pass an acting user and this cell must refuse to
create a user-owned project without one. Until then service-created repos are
ops artifacts, not customer data.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Gitea rejects a null password with a bare 400. These accounts are never
password-authenticated — humans arrive via OIDC, agents via passport-bound
scoped tokens, local password sign-in is disabled server-wide — so we generate
a credential that is never stored, returned or recoverable. An unusable
password is safer than a blank one or a shared default.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Alembic runs synchronously, so the container needs a sync driver. Without it
migrations fail on a fresh deploy while passing on any developer machine that
happens to have psycopg2 — exactly the class of gap that only appears in
production.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The plane Windy Cloud does not have. Verified 2026-08-11: routes/storage.py and
its models contain ZERO occurrences of share/permission/acl/collaborat/seat/
version/snapshot/history/revision. This fills a hole rather than bolting onto
something that already had one.
- repos: create/list/get, repo_type required (I-7), reserved slugs, Gitea
reached ONLY through the membrane client (I-1)
- grants: human identity OR agent passport, exactly one enforced by a database
CHECK constraint; agent grants expire in 90 days by default
- versions: history in words a person recognises — no 'commit', no 'branch',
no 'repository' in any user-facing string (D-9/I-9), with a test that greps
the speak strings and fails on developer vocabulary
- private repos 404 rather than 403, so a stranger cannot learn one exists
Auth: three first-class caller classes (human OIDC / agent EPT / internal
service token), NO fourth, and no bypass env var — copied deliberately from the
desktop control server, the ecosystem's best Principle-#5 artifact.
G3.6 status-code law implemented: 400 and 404 REFUSE, 429/5xx retry then REFUSE.
A sibling maps 400/429 to 'unreachable' and soft-ALLOWS, which is inducible —
an attacker who wants the check skipped only has to make it rate-limit itself.
A test asserts resolve_passport has exactly one return path.
And I-8 applied to ourselves: G3.2's JWKS verifier does not exist yet, so the
human token path REFUSES in production rather than accepting an unverified JWT.
An unverified JWT is an authentication bypass, not a shortcut.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
VERIFIED END TO END from outside the network:
- create repo via API -> clone -> commit -> push -> read back over HTTPS
- 3 MB LFS object pushed through the tunnel, landed in R2 at lfs/34/2d/...
- NO local lfs/ directory on the host: I-3 confirmed by measurement
- /health/full: db, gitea and r2 all green
Adds strand G4A recording four traps that each cost a crash loop, with tests:
1. [lfs] STORAGE_TYPE creates a separate storage section that does not
inherit [storage] — crash loop, error names the symptom not the cause
2. storage backend != LFS enabled; LFS_START_SERVER is separate, and its
absence reads as a permissions error
3. Gitea env-to-ini SETS but never UNSETS — removing a compose var leaves the
line in the persisted app.ini, so repo config and prod config silently
disagree. Exactly the drift this cell exists to end.
4. R2 rejects the default S3 checksum algorithm
And G4A.5, which is architecture rather than a bug: an 8 MB non-LFS push died
with HTTP 524 at Cloudflare's ~100s limit. This makes the LFS threshold
load-bearing and REQUIRES G10 to serve model weights via presigned R2 URLs
rather than proxying blobs through the tunnel — client straight to R2, which is
what Hugging Face does and which takes Grant's home upstream out of the path.
G2.4: Gitea's MIT text and a NOTICE now travel with the repo.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Naming a storage type inside [lfs] creates a separate storage section that does
NOT inherit endpoint or credentials from [storage], so Gitea crash-looped on
'Endpoint: does not follow ip address or domain name standards' — an error that
names the symptom and not the cause. LFS inherits the [storage] defaults on its
own; avatars had already proved that by initialising correctly.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three buckets created. S3 round-trip against R2 proven end to end
(PUT/GET/DELETE) before any of it was wired in.
Includes the R2 checksum trap: R2 rejects the algorithm S3 clients send by
default, and the resulting error reads like a credential problem and is not one.
Records a NAMED DEBT in SUBSTRATE.md: the R2 credential is currently the
account-wide god token, because no available token can mint a scoped one. Gated
— it must be replaced before G7 puts CI runners on this host, since I-5 exists
precisely to keep untrusted job code away from broadly-scoped credentials.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Gitea was configured to connect as a 'gitea' DB user that never existed, so it
sat unconfigured behind a working tunnel. It now gets its OWN role and database
via a first-init script — I-1 says we never write Gitea's tables, and that is
better as a permission than as a promise.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
cloudflared binds 127.0.0.1:2000 on the HOST. This process runs in a container
whose only route to the host is the bridge gateway (172.17.0.1), where nothing
is listening — so the check was permanently red regardless of what the tunnel
was actually doing.
Binding the metrics endpoint wider would have fixed the probe and made a
metrics bind failure capable of taking down ingress. That is a worse trade than
losing one row on a dashboard.
The check is not silently dropped: /health/full now carries a 'not_checked_here'
map naming the tunnel and where its health actually lives (systemd
windygit-tunnel). An observer should never have to wonder whether a missing
check means healthy or means forgotten.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
app.windygit.com / api.windygit.com / models.windygit.com are serving over
HTTPS with ZERO inbound ports open on Grant's network.
- tunnel 4e856c5d, 4 registered edge connections, systemd-managed and bounded
- three proxied single-level CNAMEs (Free Universal SSL covers them; a
two-level name would need ACM and would die in the TLS handshake)
- services bound to 127.0.0.1 with configurable host ports — Veron 1 is
Grant's workstation and 3000/3300 belong to other projects
- docs/RUNBOOK-VERON.md
I-12 PROVEN IN PRODUCTION: /version reports source=baked with a sha equal to
the deployed HEAD.
Also fixed: the tunnel health probe targeted localhost from inside a container,
so it was permanently red. A check that is always red is as useless as one that
is always green — it is how a fleet canary goes 37 days dead unnoticed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Veron 1 is Grant's workstation and already runs other projects on 3000 (node
dev server) and 3300 (nginx). A deploy must never fight a resident process for
a port, and nothing here needs to be reachable from the LAN — the Cloudflare
Tunnel runs on the host and reaches us over 127.0.0.1 (G1.6).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Strand G0 complete and VERIFIED against real Postgres, not asserted.
- FastAPI plane, fail-closed provider seams, repair-pointer error taxonomy
- migration 001: all 10 tables incl. repo_type NOT NULL and model_cards (I-7)
- 17 invariant tests, ruff clean, vocabulary audit clean
Two bugs found by RUNNING it that review would not have caught:
1. SQLAlchemy Enum persists .name, not .value — so RepoState.deleted_soft
and CreatedVia.imported would have written labels migration 001 never
declared, failing at runtime rather than at review. Pinned via
values_callable.
2. op.create_table asks each Enum to emit its own CREATE TYPE with no
checkfirst, so the second reference raised DuplicateObject and the
migration died halfway. Types are now created once, referenced with
create_type=False.
Proven live, with the hostile env var set:
- I-12: COMMIT_SHA=deadbeef... in the environment, /version reports real HEAD.
That env pin is the documented root cause of nine sibling services
misreporting their commit; here it is structurally ignored.
- I-8: three unconfigured providers -> status degraded, HTTP 503, each saying
'refusing to report healthy'. No mock, no false green.
- G0.4: upgrade -> downgrade -> upgrade round-trip clean (10 -> 0 -> 10).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Stake in the ground for Windy Git — the version, permission and provenance
plane over Windy Cloud, and an agent-native code and model host.
Captures the 2026-08-11 vision session as nine locked decisions (D-1..D-9),
thirteen invariants (I-1..I-13), and thirteen strands (G0..G12) with atomic
codons and per-codon acceptance criteria.
Grounded in measurement, not estimate:
- 61 repos = 11 GB working trees, 0.63 GB git objects
- Veron 1: 24 cores / 251 GB / 978 GB free / load 1.52 / $0
- Kit 0: 4 vCPU / load 7.62 — disqualified on four independent grounds
- Windy Cloud storage confirmed to have ZERO sharing/permission/versioning
Every warning in the plan is a lesson a sibling cell already paid for in
production, cited to its source in the 2026-08-09 audits.
Apex windygit.com purchased 2026-08-11, zone 9d8637dc, active.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>