Commit Graph

106 Commits

Author SHA1 Message Date
Grant Whitmer
830ca48705 docs: mark revocation finding resolved — it was critical, not low
All checks were successful
check / gate (push) Successful in 19s
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:20:46 -04:00
Grant Whitmer
deb3a8fafc test: prove the revocation wiring end-to-end (resolve_passport -> decide_trust)
Some checks failed
check / gate (push) Has been cancelled
Monkeypatched httpx so resolve_passport sees a real revoked trust body
(status=revoked, band=unproven, allowed=[]) and must raise
PassportNotInGoodStanding — proving the WIRING, not just the decision. This is
the path that stops a validly-signed EPT that outlived its passport's
revocation (~365-day tokens).

Live-confirmed alongside: Eternitas refuses to mint EPTs for revoked bots
("credentials are not issued for non-active bots"), so the only exposure was a
pre-existing token — exactly what this now catches. The active agent's EPT still
returns 200. A fully-live revoked test would require revoking a real fleet
passport (destructive), so the wiring is proven deterministically instead.

79 tests green.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:20:44 -04:00
Grant Whitmer
587fb05265 SECURITY: enforce passport status — revocation now takes effect live
All checks were successful
check / gate (push) Successful in 21s
A revoked passport returns HTTP 200, status=revoked, band=unproven,
allowed_actions=[] (verified live 2026-08-13). resolve_passport keyed refusal
only on HTTP 4xx and band=="untrusted", so it returned band 'unproven' and the
agent was seated. Revocation was NOT enforced on the live auth path at all — and
now that agent auth actually works, a revoked agent could authenticate and act.

Extracts decide_trust(body) -> (band, actions) | raise. Only status=="active"
is allowed; revoked/suspended/frozen/unknown all refuse, fail-closed on the
field that carries the most consequential fact about an identity. The agent call
site turns that into a clean 403 passport_revoked.

This is the REAL revocation gate — the token cannot be un-issued, but its
standing is re-checked on every request, so revocation takes effect on the next
call with no webhook required. The Eternitas webhook remains useful for
invalidating locally-issued credentials/grants (G6.3, not built yet), but it was
never the primary gate and its being unwired is no longer a live exposure.

Behavioral tests: revoked body refused, active accepted, unknown/missing status
fails closed.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:18:16 -04:00
Grant Whitmer
c96d1d3102 canary: guard both forgery shapes now that real verification is live
All checks were successful
check / gate (push) Successful in 20s
The EPT-shaped forgery is the one that matters after G3.2 — it is what
signature verification actually guards. The JWT-shaped one still exercises the
human gate. Both must_refuse; a 2xx on either pages.

Verified live: forged EPT -> 401 ept_invalid, forged JWT -> 503,
genuine EPT -> 200.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:13:44 -04:00
Grant Whitmer
79dcea4d3e G3.2/G3.4: real EPT verification + wire the throttle, reopening agent auth
All checks were successful
check / gate (push) Successful in 19s
canary / probe (push) Successful in 10s
REOPENS the agent path — but only because possession is now actually proven.

EPT verification (api/app/ept.py): ES256 against Eternitas's published key set
at /.well-known/eternitas-keys. algorithms=["ES256"] makes alg:none and
algorithm confusion unrepresentable rather than merely unlikely; issuer and exp
are enforced by the library; an unknown kid is refused.

Order is deliberate: signature FIRST, trust lookup second. These EPTs live ~365
days and carry rev/tru baked in at issuance, so a year-old "rev: false" proves
nothing — revocation and band still come from a live lookup on every request.

Found while building it: real EPTs put the passport in the "sub" claim. The old
code read "passport"/"sub_passport", which no genuine EPT carries — so real
agents were never recognised and ONLY forged tokens ever authenticated. The
bypass was not just a hole, it was the only thing that worked.

Throttle (api/app/throttle.py): BAND_MULTIPLIER and rate_*_per_day were defined
and read by nothing. Now enforced on repo.create and grant.create, counted
against agent_actions (one source of truth, not a private counter that drifts
from the audit log). Fails CLOSED — a limiter that fails open protects you until
the moment something is wrong. Untrusted band is 403 read-only, not 429, because
"slow down" would be a lie.

Tests: 14 behavioral, signing real ES256 tokens with a locally-generated key so
they exercise the crypto path with no network dependency — genuine tokens
accepted, and alg:none / foreign key / tampered payload / expired / wrong issuer
/ unknown kid / missing claims all refused. 74 green.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:09:47 -04:00
Grant Whitmer
c257cc55c6 docs: second-auditor review (Fable) — 1 critical fixed live, 4 open
All checks were successful
check / gate (push) Successful in 20s
canary / probe (push) Successful in 11s
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 22:54:15 -04:00
Grant Whitmer
85fa65a52b canary: continuously verify the forged-token bypass stays closed
All checks were successful
check / gate (push) Successful in 18s
Adds must_refuse checks: a 2xx is the alarm, a 401/403/503 is health. The
forged alg:none agent token is probed every 10 min; if it ever returns 2xx the
canary goes red and pages. Proven both ways — 503 reads ok, a 200 endpoint
alarms 'ACCEPTED — this MUST be refused'. The security property is now enforced
by a running check, not assumed.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 22:52:56 -04:00
Grant Whitmer
d8deffe4db SECURITY: close the agent-auth bypass — trust is not authentication
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 11s
Verified live 2026-08-13: a forged 'alg:none' token naming a passport lifted
from the logs returned HTTP 200 as that agent. The agent path read the passport
without verifying the EPT signature, asked Eternitas 'is this passport
reputable?', and seated the caller on a yes. That answers reputation, not
possession — anyone who knows a passport number could impersonate that agent on
the public API.

The human path already failed closed for exactly this reason
(require_verified_jwt). The gate was on the wrong path: it sat AFTER the agent
branch returned. The agent path now fails closed too, BEFORE the trust lookup,
so a forged token never even reaches Eternitas. Reopens automatically when the
ES256/JWKS verifier (G3.2/G9.1) exists and this gate consults it.

Adds BEHAVIORAL tests (not string-grep): a forged alg:none token exercised
through the real get_caller must raise, not authenticate. This is the test that
would have caught the bypass; the suite had 86 source-string assertions and
zero that ran the auth decision.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 22:49:45 -04:00
Grant Whitmer
83047def94 docs: whole account on Windy Git — 143 repos, 1.58 GB, two tiers
131 read-only mirrors (cannot run Actions, zero deploy risk) + 12 writable with
CI and deploys disabled. Total size matches the measured GitHub archive exactly,
which is the confirmation the copy is complete.

Splits the two concerns cleanly: having a copy is safe and should cover
everything now; running code needs judgement and happens per repo.

windy-pro IS included as a mirror — the G11.5 caution is about making it
writable while six checkouts disagree on HEAD, not about holding a read-only
copy. The DR copy is now complete.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:49:48 -04:00
Grant Whitmer
711c47d8c9 fix: the --all-as-mirrors flag itself was never added
The function landed but the argparse anchor did not match the real formatting,
so the command existed and could not be invoked. Caught by running it rather
than assuming the patch applied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:43:22 -04:00
Grant Whitmer
ae6ca8b9d2 G11.3: bulk DR copy — every repo as a read-only mirror
Grant's plan: clone the whole account, let it circulate, reverse direction later
when things are clean. Right plan, with one change that matters.

Sampling 40 repos found 18 carrying deploy/release/publish workflows that
trigger on push: — roughly 63 across the account. Importing those writable with
Actions enabled would arm sixty-odd production deploy triggers on Veron 1, each
needing disarming by hand.

So the bulk goes in as READ-ONLY pull mirrors. A mirror cannot run Actions at
all, so this carries zero deploy risk, and Gitea syncs them itself with no
script and no timer. What you get is a complete, current second copy of the
whole account — the disaster-recovery half — with none of the execution risk.

Converting one to writable + CI stays a deliberate per-repo act: re-import,
review its workflows, disable the deploying ones. That judgement belongs at the
moment you want CI on that repo, not in bulk sixty times by accident.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:40:52 -04:00
Grant Whitmer
89723c6ebd safety: disable deploy workflows on Windy Git before they can fire
Six workflows deploy to production on push:. Windy Git now has a working
runner, so the next synced commit to main would have attempted a production
deploy FROM VERON 1. Their secrets are unset here so they would have failed —
but loudly, on every push, with any pre-SSH step still running.

All six now disabled_manually. Tests, lints and migration checks stay active:
they need no secrets, which is exactly why Phase 1 delivers CI value with
nothing to configure.

Same class of mistake as the push-mirror direction, caught before firing this
time rather than after.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:30:56 -04:00
Grant Whitmer
b2f00821d7 docs: replace the cutover plan with the phased one that matches reality
Phase 1 requires nothing from anyone: agents keep pushing to GitHub, a timer
syncs GitHub -> Windy Git every 15 minutes, CI runs on Veron against current
code. Phase 2 flips one repo at a time, only when that repo is idle.

Records the direction mistake honestly: the source of truth is wherever people
are actually typing, not wherever the plan says it should be.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:14:21 -04:00
Grant Whitmer
51acf9f86e URGENT FIX: reverse the sync — GitHub is the source of truth, not Windy Git
I migrated nine repos writable with push-mirrors pointed AT GitHub. That was
wrong for the actual situation: a dozen agent sessions on the Mac mini are
pushing to GitHub continuously, so GitHub is where the live work is.

A push-mirror force-updates refs. On its 8-hour timer it would have pushed
Windy Git's stale copy over live work — silently, no conflict, nothing to
notice. Removed all nine before the first timer fired; verified no GitHub repo
had been touched (latest push predated the mirrors).

Replaced with the correct Phase 1 direction:

  agents --push--> GitHub --sync--> Windy Git --> CI on Veron

It requires NOTHING from anyone. No remote changes, no coordination, no
'everybody stop pushing'. Agents keep working exactly as they are and CI starts
running on 24 cores.

Windy Git is force-updated on purpose: in Phase 1 it holds nothing anyone
depends on, so GitHub always wins and there is no merge to reconcile.

Phase 2 is per-repo and only when that repo is idle. Never a big-bang cutover
across a dozen live sessions.

Fetches +refs/heads/* and tags explicitly rather than --mirror, which would drag
GitHub's refs/pull/* that Gitea rejects and bury the real errors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:11:34 -04:00
Grant Whitmer
7711adcb97 docs: cutover — Windy Git is the daily driver
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 11s
9 repos migrated writable with push-mirrors to GitHub (sync_on_commit).
Full loop proven end to end: pushed a commit to Windy Git, its existing
workflow ran on Veron 1, and GitHub received the commit within 20s.

Documents the one rule the cutover creates: do NOT push directly to GitHub for
a migrated repo. The mirror makes GitHub match Windy Git, so a direct commit is
overwritten on the next sync, silently, with no conflict. One writer is the
point — two writers with no reconciliation is how you lose work you thought was
saved.

windy-pro deliberately excluded until its six-checkout / forked-counter
ambiguity is resolved by reading (G11.5).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 21:54:47 -04:00
Grant Whitmer
2c292c501c G0.9: nightly timer + the first rehearsed restore in the ecosystem
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 23s
Timer installed with Persistent=true — Grant's workstation is not always on at
04:17, and without that a missed window is silently skipped and the backup
simply never runs.

RESTORE DRILL PASSED, which is the part that matters: pulled a bundle back out
of R2, cloned from it, and the restored HEAD (ffce529) matches live origin/main
exactly — 40 commits, all branches, 50 files.

The August audits found no rehearsed restore anywhere in the ecosystem, for
anything. This is one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 21:49:11 -04:00
Grant Whitmer
ffce529344 G0.9: nightly backup — the prerequisite for becoming the daily driver
All checks were successful
check / gate (push) Successful in 23s
Today GitHub is authoritative, so losing Veron 1 costs nothing. The moment
people push HERE first that inverts: Veron 1 holds the only current copy of the
company's source between mirror syncs, and it is Grant's workstation — no SLA,
no snapshots, residential line, and he reboots it.

git bundle over tar, deliberately: a bundle is one file that git clone reads
directly, so a restore needs no knowledge of Gitea's on-disk layout, and
bundling asks git for a consistent view instead of racing a live push.

  - --all, so every branch and tag is captured. A single-branch bundle loses
    the rest silently and you find out during the restore.
  - git bundle verify before upload. An unverified bundle is a belief.
  - empty repos are reported as skipped, not counted as failures
  - non-zero exit on ANY failure so the unit goes red. A backup script that
    swallows errors manufactures confidence.

Whole archive measured 1.58 GB across 141 repos — about two cents a month.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 21:47:46 -04:00
Grant Whitmer
db31e786e0 G7.4: default to writable repos — mirrors cannot run CI (measured)
All checks were successful
check / gate (push) Successful in 37s
canary / probe (push) Successful in 23s
windy-calendar imported as a pull mirror sat at 0 workflow runs. Gitea does not
fire Actions on mirror sync and a mirror is not a push target, so a mirrored
repo gives you the code and none of the point.

Re-imported it writable, pushed a commit, and its EXISTING .github/workflows/
ci.yml ran on Veron 1 and reported success — with zero workflow edits. That
workflow's own header says it was written for 'OUR self-hosted runners (Kit 0 +
Veron One)' because the account is billing-locked. The fleet already wanted
this; the runners just died when the repos went private.

Default is now writable + push-mirror to GitHub (I-4 steady state). --mirror
stays available for repos to copy but not move.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 19:58:15 -04:00
Grant Whitmer
14fb43b959 G7.4: advertise the fleet's runner labels, and correct my own doctrine
Some checks are pending
check / gate (push) Waiting to run
Surveyed ten repos: 36 of 36 ACTIVE workflows already say
runs-on: [self-hosted, linux, x64]. They were written for the self-hosted
runners that died when the repos went private — so advertising those three
labels makes every one of them runnable AS-IS. No workflow edits, no rewrites.

Correcting an earlier note in this plan: 'ban ubuntu-latest' was wrong as
stated. All 11 occurrences are tagged '# runner-lint-allow — CD/hosted-only;
disabled, manual until CD mission'. They are deliberately hosted-only and
deliberately off. They are correct as written and should not be 'fixed'.

I generalised that doctrine from one repo without surveying. The survey
disagreed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 19:53:44 -04:00
Grant Whitmer
e3677dddd9 docs: measure WHY login takes 18s — it is not 468 call sites
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 24s
The SOTU scopes the child-process DB bridge as 468 call sites and multi-week.
Measurement says otherwise.

  bare node -e '0'  on Kit 0 (4 vCPU, load 20): 1.70-1.99s
  bare node -e '0'  on Veron 1 (24 cores, idle): 0.01s

170x, and node startup is essentially the whole per-query cost — adding pg
connect and a real query to a bare node boot adds only ~0.2-1.2s on top of
~1.8s of interpreter start.

Login makes 9 such forks. 9 x 1.8s = 16s. Observed: 17-25s.

Two independent multipliers compound: the adapter forks per query (9x) and the
box is saturated so each fork costs 1.8s instead of 0.01s (170x).

The consequence: you do not need to touch 468 call sites. querySyncViaChild is
ONE function and its interface does not change — replace the per-query
execFileSync with a persistent worker holding a pg.Pool. All 468 call sites get
faster without being edited, including the mail-lookup that caused today's
outage.

Also measured: 54 containers, 301% CPU of 400%, load 20. dev/demo is 55% of
that — real relief, honestly not a fix. 42 non-dev containers on 4 cores is the
actual condition.

Not implemented here on purpose: sync-over-async with Atomics in the most
critical file in the ecosystem, at the end of a long session, on a box that
already had an outage today, deserves a fresh session and a load test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 19:05:25 -04:00
Grant Whitmer
ef3450d87b G7.6: never let bookkeeping kill the monitor
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 26s
An unwritable state path raised and took the whole canary down. That is the
worst possible trade for a monitoring tool: it reports nothing at all, and
reports it silently. State is an optimisation for transition detection; the
probing is the point.

Found by fat-fingering an env var, which is exactly how it would happen in
production.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 15:42:58 -04:00
Grant Whitmer
a094960da3 G3.5: accept platform.test_ping unverified — unverifiable by construction
All checks were successful
check / gate (push) Successful in 19s
canary / probe (push) Successful in 26s
Read the sender rather than guessing: Eternitas generates the webhook secret at
registration time and pings the URL to prove reachability BEFORE returning that
secret. The ping IS signed — with a secret the receiver cannot possibly hold
yet. Unverifiable by construction, not by oversight.

Accepting it is safe because the event is definitionally a no-op: nothing read,
nothing written, acted:false. Every event that changes anything still requires a
valid HMAC. The alternative, skip_validation:true, would permanently disable
reachability checking for this platform to solve a one-time ordering problem.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 15:36:03 -04:00
Grant Whitmer
5dd914b8a6 G3.5: answer the reachability probe honestly instead of skipping validation
All checks were successful
check / gate (push) Successful in 18s
Eternitas verifies a webhook URL answers BEFORE issuing the secret that signs
deliveries, so the very first request can never carry a signature — refusing it
makes registration impossible. Real chicken-and-egg, not a reason to disable
validation.

A probe is a request claiming no event and carrying no signature. Answering it
200 is honest: the endpoint exists and is ready. It changes nothing (acted:
false), and anything claiming to BE an event still goes through full HMAC
verification. Registering with skip_validation:true would have permanently
disabled a safety check to solve a one-time ordering problem.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 15:33:56 -04:00
Grant Whitmer
339ec70853 G3.5: Eternitas revocation receiver — fail-closed, both signature traps avoided
All checks were successful
check / gate (push) Successful in 18s
When a passport is revoked, every credential it holds here dies in one
transaction: tokens revoked, grants revoked. A revocation that takes effect
'eventually' is not a revocation.

Avoids two traps that each cost a sibling service a subscription that looked
wired and never once delivered:
  1. strip the 'sha256=' prefix before comparing — comparing the decorated
     header against a bare digest returns 401 forever
  2. HMAC the RAW REQUEST BYTES, never a re-serialised body — JSON.stringify of
     a parsed body reorders keys and changes whitespace, so the digest never
     matches what the sender signed

Both fail silently from the sender's side: Eternitas records a delivery, the
receiver records a rejection, nobody notices for weeks.

Unset secret REFUSES rather than accepts — accepting unverified instructions
about identity is worse than missing them. And it never acknowledges a
revocation it could not apply; a 200 there is a security hole reporting success.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 15:32:27 -04:00
Grant Whitmer
fc1937560c G7.6: fix the alert path — urllib UA was rejected 403 by Resend
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 22s
Caught by TESTING the alert path instead of assuming it. Without an explicit
User-Agent, urllib sends 'Python-urllib/3.x' and Resend rejects it 403, while
the identical request via curl succeeds.

The failure mode this avoids is the worst one a canary has: it would have
detected every outage correctly and told nobody. Same bot-filtering trap as the
Gitea migrate call earlier today — worth recognising on sight.

Also prints the HTTP body on failure. '403 Forbidden' alone sends you hunting
for a bad key; the body names the real cause.

Verified: alert sent (200).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:44:00 -04:00
Grant Whitmer
9a7030351b G7.6: fleet canary — probes what a user does, not what is cheap
All checks were successful
check / gate (push) Successful in 19s
Today's outage is the whole design brief: /health returned 200 for the entire
hour that login was dead. A canary watching /health would have stayed green
while nobody in the ecosystem could sign in. So the login probe is here and it
is the one that matters.

Three rules it obeys:
  - never green for something it did not prove (I-8)
  - alert on TRANSITIONS, not every run — a canary people filter is a dead
    canary, which is how the last one sat 37 days dead unnoticed
  - run where the watched thing cannot take it down: Veron 1, never Kit 0

Two independent signals, so losing one still leaves the other: an email via
Resend on state change, and a non-zero exit that turns the CI run red in the
forge itself.

Alerts say what broke in human terms — 'a human can actually sign in' — rather
than only naming an endpoint.

Verified against production: 7/7 green including login at 17.1s; a forced 404
reports DOWN; a 1s threshold reports SLOW at 23.3s.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:41:37 -04:00
Grant Whitmer
211a48187f docs: incident RESOLVED — roster backoff deployed, agent chat back
All checks were successful
check / gate (push) Successful in 18s
The retry fix already existed: a parallel session landed c79f196 (#172) with a
test. Kit 0 was ONE COMMIT BEHIND and did not have it — which is the entire
reason the loop ran.

Deployed by ff-merging /root/windy-chat ac61db6 -> c79f196 and rebuilding only
agent-roster (--no-deps). Never reset --hard; that checkout has a documented
history of local edits a hard reset would eat.

  account-server CPU   168-210%  ->  0.00%
  roster failures      74/min    ->  0/min
  agent chat           stopped   ->  up, healthy
  login                timeout   ->  HTTP 200 ~18s

The lesson: a merged fix that has not reached production is not a fix, it is a
belief. That is exactly what both August audits named — nothing checks whether a
decision reached production — arriving as an outage instead of a report finding.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:19:58 -04:00
Grant Whitmer
3926726309 docs: incident record — account-server login outage 2026-08-12
All checks were successful
check / gate (push) Successful in 18s
Login was timing out for ~50 minutes on the identity service every Windy
product authenticates against.

Two faults multiplying: windy-agent-roster retried a failing mail-lookup at 74
failures/minute with no backoff, and every call cost 0.6-2.1s because the
postgres adapter forks a node process per query and blocks the event loop.
Together they formed a feedback loop — the container's listening socket showed
Recv-Q 510, connections the kernel accepted that node was too blocked to pick
up. Login sat in that queue.

Resolved by stopping windy-agent-roster: login went from 45s+ timeout to HTTP
200 in ~18s. Restored, not healthy — 18s is the fork-per-query adapter on a
54-container 4-vCPU box, and it is what remains after the loop was removed.

Nothing caught this. The container was (unhealthy) with a failing healthcheck
streak of 74 and no alert exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:00:54 -04:00
Grant Whitmer
b6a7045907 G7.4/G11.3: import tool — GitHub upstream, Windy Git downstream, CI here
All checks were successful
check / gate (push) Successful in 17s
Direction is deliberate. For the migration quarter GitHub stays upstream and
Windy Git is a PULL mirror, so it cannot diverge — worst case it is stale, not
wrong. Making Windy Git authoritative before Grant flips G11.6 would create a
two-writer problem nobody asked for. I-4's push-mirror is the steady state for
repos that originate here.

Refuses windy-pro by name: six checkouts, a build counter forked three ways, and
two sessions recording different HEADs hours apart. That gets resolved by
reading, not by importing (G11.5).

Records the Cloudflare trap: the public endpoint answers 403 error 1010 because
CF blocks urllib's user-agent as a bot signature. It reads like Gitea rejecting
the token and is not — the identical call against localhost:3080 on the host
succeeds. Bulk import belongs on the host anyway.

Verified: windy-calendar imported, 731 KB, mirroring.

Survey: 135 private repos, and GitHub Actions cannot run on ANY of them. Sampled
windy-pro, windy-mind, eternitas, windy-registry, windy-agent, WindyCloud — every
recent run is startup_failure, as recently as today.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 12:12:02 -04:00
Grant Whitmer
44756091b0 G7 COMPLETE: CI is green on our own hardware
All checks were successful
check / gate (push) Successful in 17s
Run 7 passed end to end: Python 3.12.13, ruff clean, vocabulary audit clean,
42 tests, migration upgrade->downgrade->upgrade against real Postgres, and
'I-12 holds: env override ignored'.

This is the finding both August audits converged on — 'nothing anywhere checks
whether a decision reached production' — closed for this repo. GitHub Actions
is billing-locked and cannot run on private repos at all, even self-hosted.
This can, on hardware Grant already owns, for zero dollars.

I-5 proven by inspection, not assertion: CI containers sit on
windy-git-runner_jobs, the forge database on windy-git_default. Disjoint. And
neither CI container holds the host docker socket.

Adds strand G7A: seven runs to first green, and each failure presented as a
different problem than it was. The worst was a Python version mismatch that
appeared as a 14-minute hang rather than an error, because pip answered
3.10-vs-3.12 by backtracking through every dependency's release history with
-q hiding it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:38:31 -04:00
Grant Whitmer
358c9bdaa7 G7: clean CI run on a stable runner
All checks were successful
check / gate (push) Successful in 44s
Runs 3 and 6 were killed by me restarting the runner mid-job, not by any
defect in the workflow. Config verified on disk first this time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:35:14 -04:00
Grant Whitmer
6b96087554 G7: per-job networks — fixes service DNS and tightens isolation
Some checks failed
check / gate (push) Failing after 13m33s
The migration step failed with 'could not translate host name postgres'. The
Postgres service container was healthy; the job simply could not name it,
because service DNS aliases only exist on a per-job network and I had pinned
container.network to the flat 'bridge'.

That choice was wrong in both directions: it broke service containers AND it
was weaker isolation, since every concurrent job shared one bridge and could
see its neighbours. A per-job network is stricter and correct — and still has
no route to the forge, because these networks live inside the dind daemon,
which has no forge attachment at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:18:51 -04:00
Grant Whitmer
8c4bdd2bc2 G7.3: setup-python on the act image, not a python image
Some checks failed
check / gate (push) Failing after 43s
Trading one wedge for another: running the job in python:3.12-bookworm fixed
the resolver spiral but broke checkout, because actions/checkout is a
JavaScript action and the official Python images carry no node — 'executable
file not found in $PATH'.

setup-python on the act image has both. Test now accepts either a pinned
container image or an explicit setup-python version, and still checks it
against pyproject's requires-python.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:16:00 -04:00
Grant Whitmer
980fc2dddd G7.3: pin Python 3.12 in CI and cap job runtime
Some checks failed
check / gate (push) Failing after 7s
The first real CI run wedged for 14 minutes. Not a network problem, not the
isolation work — catthehacker/ubuntu:act-22.04 ships Python 3.10.12 while this
project declares requires-python >=3.12, and pip answered that by backtracking
through the entire release history of every dependency looking for something
3.10-compatible. At 100% CPU, with -q hiding every line of it, and it would
have churned until the runner's 30m timeout.

A version mismatch presenting as a hang rather than an error is worth a test,
so there is one: the workflow's python image must satisfy pyproject's
requires-python, checked by parsing both rather than by eyeballing them.

Also: every job now has timeout-minutes. A wedged step should be a red check in
minutes, not an occupied runner for half an hour.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:13:39 -04:00
Grant Whitmer
ed71102e49 G7: trigger a CI run after the routing fix
Some checks failed
check / gate (push) Has been cancelled
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 10:53:57 -04:00
Grant Whitmer
53dea84ab1 G7: route CI jobs via the public surface, not the forge network
Some checks failed
check / gate (push) Failing after 13s
The first CI run failed with 'Could not resolve host: gitea' — job containers
live on dind's private network and cannot see the forge network. Two ways out,
and they are not equivalent:

  (a) put job containers on the forge network. Easy, one line, and it leaves
      untrusted workflow code one DNS name from the forge's Postgres. It quietly
      repeals I-5.
  (b) send jobs to the PUBLIC forge surface over the tunnel, exactly like any
      stranger on the internet.

Took (b). The runner no longer needs the forge network at all, so there is now
NO private route from any CI container to anything — a better posture than this
file started with. Cost is a hairpin through Cloudflare plus its ~100s ceiling
per fetch, which for 0.63 GB of objects across 61 repos and depth=1 checkouts is
nowhere near binding.

Test upgraded to assert the stronger property: no CI container joins the forge
network.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 10:51:00 -04:00
Grant Whitmer
5f717ef74b G7: CI runners with real isolation, and the gate as a workflow
Some checks failed
check / gate (push) Failing after 50s
I-5 says runners execute untrusted code and must be isolated by machine
boundary. act_runner needs a Docker daemon to start job containers, and the
tempting move — what every published example does — is to mount the host's
/var/run/docker.sock. That hands every workflow, including whatever a
transitive dependency's postinstall script feels like doing, the ability to
start a privileged container mounting / — root on Grant's workstation.

Instead the runner talks to its OWN dind daemon:
  - runner (TRUSTED, the act_runner daemon) sits on the forge network only to
    collect jobs from gitea:3000
  - dind and every job container it spawns are UNTRUSTED, on a private network
    with no route to the forge, its Postgres, or its .env
  - jobs cannot bind-mount from the daemon host (valid_volumes: []) and are not
    handed the runner's own socket (docker_host: -)
  - separate compose project, cpu/memory bounded — Veron 1 is Grant's
    workstation, not a dedicated build box

The gate itself now runs as a workflow, including the migration round-trip that
already caught two bugs review did not, and the I-12 check that a COMMIT_SHA
env override cannot change what /version reports.

Labels are explicit and pinned. A workflow naming a label nobody provides
queues forever and presents as a hung CI system rather than a typo — which is
what ubuntu-latest does on every windy-registry run today.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 10:39:05 -04:00
Grant Whitmer
b274392e96 G3.1: OIDC live — sign in to Windy Git with a Windy account
Client 'windy-git' registered on account-server; Gitea auth source 'windy'
added against the discovery document. Auto-registration on, so a Windy account
IS the account — nobody is asked to invent a second identity for the same
person and no local password ever exists.

Proven end to end:
  - app.windygit.com/user/login offers 'Sign in with windy'
  - /user/oauth2/windy -> 307 to account.windyword.ai/oauth/authorize
  - authenticated authorize -> 200, issues a code to the registered callback
  - a bogus evil.example.com redirect_uri -> 'redirect_uri not registered for
    this client'. The anti-phishing check works, which is the whole reason
    redirect URIs are registered rather than accepted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 10:27:14 -04:00
Grant Whitmer
d684cfc2e1 G4.2: record Grant's ruling — use an existing fleet CF token
Not a debt, a decision: sandbox phase, months from launch, and minting a tenth
Cloudflare token to sit in the inventory costs more than it buys. A
platform-specific scoped token is a launch-hardening item.

Recorded so the next reader knows it was chosen rather than missed, with a
do-not-re-raise note. Keeps the genuinely non-obvious part: R2's S3 credentials
are DERIVED from a CF API token — access key id = the token's id, secret =
SHA-256 of the token value.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:34:10 -04:00
Grant Whitmer
a2430de94d I-4: distinguish a mirror that never ran from one that is behind
Gitea reports the epoch for 'not yet synced', which arithmetic turns into a
56-year lag and a confident 'degraded'. Collapsing those two states is how a
backup that was never made gets read as a backup that is merely stale — which
is the more dangerous direction, because 'behind' sounds survivable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:16:05 -04:00
Grant Whitmer
da257652c2 G11 / I-4: continuous off-site mirror, and a namespace bug fixed
I-4 said 'never a one-way door' and had no implementation. Now it does.

  - ensure the GitHub counterpart exists (idempotent), then ask Gitea to keep
    it in step with sync_on_commit=True. An hourly timer means an hour of work
    can be the thing you lose, and that window is invisible until it costs you.
  - mirror status reports what is TRUE including 'we do not know'. An
    unconfigured mirror reports unconfigured, NEVER healthy — same posture as
    me-fleet.ts refusing to say 'online' when it only knows 'registered'.
  - lag past the threshold is a P2, not a shrug. A mirror nobody checks is a
    belief, not a backup, and this ecosystem already lost 37 days to a canary
    everyone assumed was fine.

Gitea owns the replication rather than a hand-rolled loop, because a background
job that fails silently is exactly how the registry's integrity refresh spent
its entire life calling a 404 and incrementing a counter instead of raising.

Also fixes a real bug I had written myself: list_versions derived the Gitea
namespace from the CALLER, which is correct only while the caller is the owner
and addresses the wrong namespace the moment a collaborator asks — surfacing as
'not found', which is the hardest kind of bug to see. Now derived from the repo,
with a test that keeps it that way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:11:45 -04:00
Grant Whitmer
066de34489 G5: fix version-count copy; record the u-system namespace finding
'There are 1 saved versions' is exactly the sloppiness the vocabulary law
exists to catch. Copy is design material, not decoration.

G5.9 records something the live test surfaced: a repo created through
X-Service-Token lands in a 'u-system' namespace because the service caller has
no identity of its own. Correct for /internal plumbing, WRONG for anything a
person owns — the portal must pass an acting user and this cell must refuse to
create a user-owned project without one. Until then service-created repos are
ops artifacts, not customer data.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 15:59:41 -04:00
Grant Whitmer
3a9259a0da G5: generate an unusable password on Gitea user create
Gitea rejects a null password with a bare 400. These accounts are never
password-authenticated — humans arrive via OIDC, agents via passport-bound
scoped tokens, local password sign-in is disabled server-wide — so we generate
a credential that is never stored, returned or recoverable. An unusable
password is safer than a blank one or a shared default.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 15:55:11 -04:00
Grant Whitmer
f621e49770 G0.4: add psycopg2-binary so migrations run inside the image
Alembic runs synchronously, so the container needs a sync driver. Without it
migrations fail on a fresh deploy while passing on any developer machine that
happens to have psycopg2 — exactly the class of gap that only appears in
production.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 15:47:14 -04:00
Grant Whitmer
099e4be9b1 G5: the shelter — repos, grants and version history
The plane Windy Cloud does not have. Verified 2026-08-11: routes/storage.py and
its models contain ZERO occurrences of share/permission/acl/collaborat/seat/
version/snapshot/history/revision. This fills a hole rather than bolting onto
something that already had one.

  - repos: create/list/get, repo_type required (I-7), reserved slugs, Gitea
    reached ONLY through the membrane client (I-1)
  - grants: human identity OR agent passport, exactly one enforced by a database
    CHECK constraint; agent grants expire in 90 days by default
  - versions: history in words a person recognises — no 'commit', no 'branch',
    no 'repository' in any user-facing string (D-9/I-9), with a test that greps
    the speak strings and fails on developer vocabulary
  - private repos 404 rather than 403, so a stranger cannot learn one exists

Auth: three first-class caller classes (human OIDC / agent EPT / internal
service token), NO fourth, and no bypass env var — copied deliberately from the
desktop control server, the ecosystem's best Principle-#5 artifact.

G3.6 status-code law implemented: 400 and 404 REFUSE, 429/5xx retry then REFUSE.
A sibling maps 400/429 to 'unreachable' and soft-ALLOWS, which is inducible —
an attacker who wants the check skipped only has to make it rate-limit itself.
A test asserts resolve_passport has exactly one return path.

And I-8 applied to ourselves: G3.2's JWKS verifier does not exist yet, so the
human token path REFUSES in production rather than accepting an unverified JWT.
An unverified JWT is an authentication bypass, not a shortcut.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 15:45:22 -04:00
Grant Whitmer
82044933ed G2/G4 complete: Gitea live, LFS landing in R2, four traps pinned
VERIFIED END TO END from outside the network:
  - create repo via API -> clone -> commit -> push -> read back over HTTPS
  - 3 MB LFS object pushed through the tunnel, landed in R2 at lfs/34/2d/...
  - NO local lfs/ directory on the host: I-3 confirmed by measurement
  - /health/full: db, gitea and r2 all green

Adds strand G4A recording four traps that each cost a crash loop, with tests:
  1. [lfs] STORAGE_TYPE creates a separate storage section that does not
     inherit [storage] — crash loop, error names the symptom not the cause
  2. storage backend != LFS enabled; LFS_START_SERVER is separate, and its
     absence reads as a permissions error
  3. Gitea env-to-ini SETS but never UNSETS — removing a compose var leaves the
     line in the persisted app.ini, so repo config and prod config silently
     disagree. Exactly the drift this cell exists to end.
  4. R2 rejects the default S3 checksum algorithm

And G4A.5, which is architecture rather than a bug: an 8 MB non-LFS push died
with HTTP 524 at Cloudflare's ~100s limit. This makes the LFS threshold
load-bearing and REQUIRES G10 to serve model weights via presigned R2 URLs
rather than proxying blobs through the tunnel — client straight to R2, which is
what Hugging Face does and which takes Grant's home upstream out of the path.

G2.4: Gitea's MIT text and a NOTICE now travel with the repo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 15:13:58 -04:00
Grant Whitmer
647d414ec1 G4.3: fix Gitea crash loop — [lfs] STORAGE_TYPE breaks inheritance
Naming a storage type inside [lfs] creates a separate storage section that does
NOT inherit endpoint or credentials from [storage], so Gitea crash-looped on
'Endpoint: does not follow ip address or domain name standards' — an error that
names the symptom and not the cause. LFS inherits the [storage] defaults on its
own; avatars had already proved that by initialising correctly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 15:00:11 -04:00
Grant Whitmer
301aa47b52 G4.3: enable the LFS server (storage type alone does not)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 14:58:05 -04:00
Grant Whitmer
ef75ee5a9c G4: R2 buckets live, Gitea storage wired, checksum trap pinned
Three buckets created. S3 round-trip against R2 proven end to end
(PUT/GET/DELETE) before any of it was wired in.

Includes the R2 checksum trap: R2 rejects the algorithm S3 clients send by
default, and the resulting error reads like a credential problem and is not one.

Records a NAMED DEBT in SUBSTRATE.md: the R2 credential is currently the
account-wide god token, because no available token can mint a scoped one. Gated
— it must be replaced before G7 puts CI runners on this host, since I-5 exists
precisely to keep untrusted job code away from broadly-scoped credentials.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 14:49:26 -04:00
Grant Whitmer
eec087ac50 G2: Gitea auto-install, own DB role, SSH deferred, Actions on
Gitea was configured to connect as a 'gitea' DB user that never existed, so it
sat unconfigured behind a working tunnel. It now gets its OWN role and database
via a first-init script — I-1 says we never write Gitea's tables, and that is
better as a permission than as a promise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 14:44:13 -04:00