Commit Graph

47 Commits

Author SHA1 Message Date
Kit OC5
54c63b4d89 push velocity: key humans on windy_identity_id (SSO link); no id -> system + caller
Telemetry UPDATE 2 actor rule: agent/human rows without actor_id are
quarantined. Forge humans sign in only via Windy SSO, so Gitea's
external_login_user.external_id is their windy_identity_id.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 13:06:17 -04:00
Kit OC5
ac83c4721e telemetry: detect push velocity from Gitea's action table (alert only)
git push never touches our API, so throttle.py can't see it. Gitea's
action table records every push; the 5-min sync-side emitter now reads
it and emits forge.push_velocity when an account crosses 60 pushes/1h,
500 pushes/24h (standard-band base) or 10 ref deletes/24h. One row per
account per rule per window while over; windyadmin (the sync) exempt.
Nothing sits in the push path and nothing is refused. HOLD until
Telemetry Boss declares the shape.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 13:06:17 -04:00
Kit OC5
e3b69fa759 telemetry: UPDATE 7 — read the ingest body; count quarantined + dropped on heartbeats
Some checks failed
canary / probe (push) Has been cancelled
check / gate (push) Has been cancelled
The ledger answers 202 even when it quarantines rows. Both emitters now log
a warning with the reasons and report service.health.telemetry_quarantined
and telemetry_dropped (API: buffer overflow; sync: 0 by construction, since
a failed send keeps cursor + spool). HOLD until Telemetry Boss declares both
keys on windy-git's two service.health shapes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:32:28 -04:00
Kit OC5
4acf50d9ef bridge tests: fake serves workflow contents; cover invalid-workflow status
All checks were successful
check / gate (push) Successful in 23s
b7a7e94 made the bridge read workflow files, which the strict fake Gitea
refused (7 red). The fake now serves contents (404 when absent), and new
tests cover: error posted with no runs, valid files add nothing, no repost,
.gitea/workflows wins over .github/workflows, and each workflow_problem
shape. pyyaml declared in dev extras (the bridge imports it).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:31:19 -04:00
c83f808a60 bridge: retry transport blips (TLS timeout/reset), never HTTP errors
All checks were successful
check / gate (push) Successful in 20s
canary / probe (push) Successful in 9s
A single GitHub TLS handshake timeout failed the whole sync, flipped its
windy-job heartbeat to ok:false and would page for nothing. Up to 3
attempts with backoff for URLError/timeout/reset; HTTP errors return
immediately as before. Test covers both.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:15:22 -04:00
eb27e63db3 telemetry: adopt the end-to-end synthetic convention (UPDATE 4)
All checks were successful
check / gate (push) Successful in 24s
canary / probe (push) Successful in 9s
Replaces the keyed marker from 1c3b5b0 with the ecosystem convention:
any X-Windy-Synthetic value marks the request synthetic; the flag lives in
a per-request contextvar, labels this request's rows, and is FORWARDED on
downstream calls (Eternitas trust lookup, Gitea API). The canary sends
"1". Rows are still recorded; the label separates, never suppresses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:06:51 -04:00
1c3b5b0638 telemetry: synthetic:true on canary refusals (keyed, not a bare flag)
All checks were successful
check / gate (push) Successful in 25s
canary / probe (push) Successful in 7s
The canary deliberately sends forged tokens every 10 min; those refusal
rows read as attacks. It now sends X-Windy-Synthetic carrying a shared
secret (Gitea repo secret CANARY_SYNTHETIC_KEY = WINDYGIT_SYNTHETIC_KEY in
Veron .env); the API marks the row synthetic only on a constant-time
match, so an attacker cannot label their own refusals synthetic to hide.
synthetic is declared on forge.auth.failed (Telemetry Boss, UPDATE 3).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:54:56 -04:00
90643fe48e telemetry step 2: API boot/health + forge.auth.failed (declared)
All checks were successful
check / gate (push) Successful in 25s
canary / probe (push) Successful in 6s
Membrane first: I-2 and MEMBRANE.v1 now list the windy-admin ledger
(POST /v1/events). api/app/telemetry.py: service.boot once per start
(commit_sha omitted when unknown, I-12), an hourly in-process
service.health with the shared keys (requests, errors_5xx/4xx,
refusals_4xx, p95_ms only when there was traffic), and one
forge.auth.failed row per refused request: declared 13-code enum,
http_status, caller class, route TEMPLATE (never the concrete path),
actor_type system with no actor_id (all-lanes rule). No token = nothing
sent or buffered; flush failures keep rows (bounded) and never raise.
Token from root-only /etc/windygit/telemetry.env (optional env_file).
8 behavioural tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:47:29 -04:00
f246417095 ci: windy-pro desktop/installer jobs are NON-BLOCKING (Grant, 09-23)
build-desktop, test-installer and reality-check still run on Windy Git
and stay visible there, but the bridge no longer posts them to GitHub, so
they cannot turn windy-pro's combined status red. Windy Git side only;
the desktop code is Grant's to fix. BRIDGE_NON_BLOCKING, per repo.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 10:58:07 -04:00
50c1464043 security: never bundle credential repos to R2 in plaintext
All checks were successful
check / gate (push) Successful in 23s
canary / probe (push) Successful in 7s
kit-army-config (the lockbox) and every *-soul / anima repo carry
credentials; the nightly R2 bundles are unencrypted, so the R2 key was a
key to every secret. Excluded by name (BACKUP_EXCLUDE); they are backed up
encrypted by the Windy Drops lane (restic) and stay mirrored on Veron.
Behavioural test runs the script's own exclusion function.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 10:48:42 -04:00
e6530d3171 test: behavioral G3.5 webhook tests (audit: tests were source-string asserts)
All checks were successful
check / gate (push) Successful in 20s
canary / probe (push) Successful in 9s
Seven HTTP-level tests through the real route: sha256= prefix and bare
digests accepted, digest of re-serialised JSON refused, forged/wrong-key/
missing signatures refused, unset secret -> 503, a revocation with a bad
signature never reaches the handler, the reachability ping never acts.
Mutation-checked: dropping the prefix strip fails the behavioral test
while the old string-grep invariant still passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 03:26:14 -04:00
5b16114b98 auth: token contract v1 (aud windy_git, both issuers); CI for eternitas
Some checks failed
check / gate (push) Successful in 37s
canary / probe (push) Has been cancelled
- hub_jwt: aud list is ["windy_git"] (contract v1 array). Dropped
  "windy-git": that is Gitea's OIDC client_id, so a forge id_token would
  have passed the aud check. `type: human` is now REQUIRED (id_tokens have
  none), which makes accepting the discovery-URL issuer safe.
- runner job ceiling 30m -> 90m: eternitas's serial pytest is ~50 min and
  would have been killed mid-suite.
- eternitas (private) added to the GitHub status bridge.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:58:56 -04:00
18ea9a4686 auth: G3.2 hub JWKS verifier — humans can sign in to the plane (SSO #14)
The human path refused every token in production (503
human_signin_not_ready) because no verifier existed. api/app/hub_jwt.py
verifies hub access tokens against account.windyword.ai's JWKS:

- RS256 only (closes alg:none and HS256-with-public-key confusion)
- iss must be "windy-identity" — what hub ACCESS tokens carry (observed
  live); id_tokens (discovery-URL issuer) are not accepted as bearers
- aud optional today, must name Windy Git when present; hub_require_aud
  flips it mandatory once the hub emits it. PyJWT's own aud check is off
  on purpose: it rejects ANY aud-bearing token when no audience is given.
- type must be human; identity = windy_identity_id, never sub (per-row id)
- production verifies even if require_verified_jwt is off

11 behavioral tests sign real RS256 tokens with a local key.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:55:17 -04:00
45686283be ci: bound CI storage; don't bridge image-build jobs
- deploy/runner/prune.sh + windygit-ci-prune.timer (6h): age-based prune
  of the CI-only dind (containers, finished-job volumes, images/builder
  cache >7d) plus a hard 60 GB cap. Only that daemon, over its own TCP
  socket; never the host's Docker. It was 38 GB and unbounded — the same
  class of growth that filled Kit 0 on 09-01.
- pr_status_bridge: jobs named *docker* are not posted. Job containers
  have no daemon by design (I-5), so they are red on every commit; a
  permanent red X teaches everyone to ignore red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:47:28 -04:00
e4a15869c0 ci: onboard windy-connect + windy-search to private-repo CI
Some checks failed
check / gate (push) Has been cancelled
windy-connect promoted from pull mirror to writable (release.yml, which
publishes to PyPI on tag push, disabled — the sync pushes tags).
windy-search was already writable; its scheduled drift-check is disabled
because it now runs as cron on Kit 0. Both added to BRIDGE_REPOS.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:44:19 -04:00
dcf9286f16 ci: make Windy Git CI permanent for the private repos
All checks were successful
check / gate (push) Successful in 21s
canary / probe (push) Successful in 6s
- Four runners x capacity 1 instead of one x capacity 4. Concurrent jobs in
  one act_runner share /root/.cache/act; a refresh racing a copy killed 3 of
  windy-chat's ~20 jobs at setup-node (lstat ... no such file). Separate
  processes have separate caches. Same parallelism, same capped dind.
- Behavioral tests for pr_status_bridge (latest verdict wins, no reposting,
  skipped never painted green, fork PRs never run, pagination, PR lifecycle).
- import_from_github.py reads IMPORT_GITEA_URL, not GITEA_BASE_URL: sourcing
  the deploy .env pointed it at http://gitea:3000 and it died on DNS after the
  mirror it replaces had already been deleted.
- CUTOVER.md: the private-repo CI path, onboarding steps, and the
  /actions/tasks-hides-queued-runs trap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:18:04 -04:00
390c1e7479 ops: move tunnel metrics to 2001, sync windy-git into itself
Some checks failed
check / gate (push) Successful in 19s
canary / probe (push) Failing after 7s
windygit-tunnel had crash-looped ~91k times: another project's
cornercall-tunnel holds 127.0.0.1:2000, and cloudflared exits when it
cannot bind its metrics port. Ingress only survived because a stray
cloudflared.service ran the same config. That unit is now disabled and
/etc/cloudflared/config.yml uses metrics 127.0.0.1:2001.

Also add windy-git to the GitHub->Windy Git sync list; its self-hosted
copy was stuck 3 commits behind (only check + canary workflows, no
deploys, so syncing is safe).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 01:29:57 -04:00
Grant Whitmer
01e36155a3 ci: remove the service-networking probe; scope the python-pin test
Some checks failed
check / gate (push) Successful in 20s
canary / probe (push) Failing after 1m33s
The probe's own log was never retrievable through the jobs API, but the
question it asked was answered better by a direct comparison of two real
workflows on the same runner and image:

  windy-git gate       @postgres:5432   -> passes its migration round-trip
  eternitas migrations @localhost:5432  -> failed

Also scopes test_g73 to workflows that actually run Python. It failed the probe
for not pinning a version when the probe only shelled out to psql — the test
being wrong rather than the workflow.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 14:41:14 -04:00
Grant Whitmer
eae1bff50b G2.3: Windy Git branding — and get it out of one host's disk into the repo
All checks were successful
check / gate (push) Successful in 19s
canary / probe (push) Successful in 8s
The front end was 100% stock Gitea: green teacup, "Gitea: Git with a cup of
tea", "A painless, self-hosted Git service". G2.3 was specified in the plan with
an acceptance test and never executed, and nothing enforced it.

Now: Windy Git name, wind-mark logo, brand-blue accent, and a landing page that
says what this actually is. Uses Gitea's SUPPORTED surface (custom templates +
public assets) so upstream upgrades keep arriving — no source modified (D-2/I-1).

Two traps this cost, both now documented and tested:

1. GITEA__DEFAULT__APP_NAME does not work. Gitea reads APP_NAME from the TOP
   LEVEL of app.ini; the env var created a literal [default] section that Gitea
   ignores, so the installer's stock APP_NAME kept winning while the config
   looked correct. The env-to-ini pass also APPENDED a second APP_NAME rather
   than replacing the first — a new variant of the documented G4A.3 trap.

2. Cloudflare caches /assets/* for 6h and no token in this stack can purge, so
   the new logo and CSS were invisible while being correct at origin. Brand
   assets now carry a VERSION IN THE FILENAME; bump it on every change.

Committed with an idempotent apply.sh, because applying it straight to Veron's
disk first was itself the config-drift trap this project documents: a rebuild
would have silently reverted to stock Gitea.

85 tests green.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 09:08:48 -04:00
Grant Whitmer
e7dee39151 throttle: stop claiming to limit pushes we cannot see
All checks were successful
check / gate (push) Successful in 19s
canary / probe (push) Successful in 9s
I reintroduced the exact defect I had just criticised. ACTION_BASE listed
"push" and "push.force", but git push goes straight to Gitea over HTTPS and
never touches this API — so nothing records a push, a count would be zero
forever, and enforce() would look up a limit, count nothing, and allow
everything. A silent no-op wearing the costume of a control, made worse by a
config name that implies the protection exists.

Split into ACTION_BASE (actually enforced: repo.create, grant.create) and
NOT_ENFORCED_HERE (push, push.force) with the reason and the remedy written
down: enforcing push velocity needs a Gitea-side pre-receive or push webhook
reporting into agent_actions.

enforce("push") now raises rather than silently allowing, and a test asserts the
two sets stay disjoint.

Found by auditing whether the auth fix could be walked around — every
/api/v1/repos/* route does require a caller, and the only unauthenticated
endpoints are /health, /version and the HMAC-verified webhook.

83 tests green.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:33:30 -04:00
Grant Whitmer
c60bfb2b89 I-12: fail the build when COMMIT_SHA is empty
All checks were successful
check / gate (push) Successful in 19s
/version went null after a deploy — the exact "service cannot name its own
commit" defect this project was built to prevent, caught by its own honesty
check.

Cause: the sed replaced "" with "" (a no-op when COMMIT_SHA is empty) and the
grep then matched that same empty string, so the guard verified nothing. A build
with no COMMIT_SHA passed and shipped a container reporting commit_sha: null.

Now the build fails loudly instead.

Second cause of the stale deploy, and it was mine: an earlier `git commit
--amend` + force-push rewrote history the Veron deploy checkout was already
sitting on, leaving it divergent so `git pull -q` failed SILENTLY (-q hid
"Need to specify how to reconcile divergent branches"). Two lessons: do not
force-push a branch a deploy checkout tracks, and do not pull with -q in a
deploy script.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:27:38 -04:00
Grant Whitmer
a0ed4a5ec0 SECURITY: make EPT routing independent of signature well-formedness
All checks were successful
check / gate (push) Successful in 20s
looks_like_ept used jwt.get_unverified_header, which validates the WHOLE token
and therefore rejects anything with a malformed signature segment. Routing
consequently depended on signature well-formedness: an EPT-shaped token with a
bad signature fell through to the HUMAN path, where it was refused for the wrong
reason and — with require_verified_jwt off (dev) — could have been read as a
human identity via its `sub` claim.

Now the header segment is decoded directly, so routing depends only on what the
token CLAIMS to be; whether it is authentic remains verify_ept's job.

Also routes alg:none to the EPT verifier regardless of typ, since a `none`
token is never valid for any caller. Both forged shapes now return 401
ept_invalid — the honest code — instead of 503 "feature not ready".

Found by noticing a forged EPT returned 503 where the verifier should have
answered 401, rather than accepting "it was refused, close enough".

80 tests green.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:23:50 -04:00
Grant Whitmer
deb3a8fafc test: prove the revocation wiring end-to-end (resolve_passport -> decide_trust)
Some checks failed
check / gate (push) Has been cancelled
Monkeypatched httpx so resolve_passport sees a real revoked trust body
(status=revoked, band=unproven, allowed=[]) and must raise
PassportNotInGoodStanding — proving the WIRING, not just the decision. This is
the path that stops a validly-signed EPT that outlived its passport's
revocation (~365-day tokens).

Live-confirmed alongside: Eternitas refuses to mint EPTs for revoked bots
("credentials are not issued for non-active bots"), so the only exposure was a
pre-existing token — exactly what this now catches. The active agent's EPT still
returns 200. A fully-live revoked test would require revoking a real fleet
passport (destructive), so the wiring is proven deterministically instead.

79 tests green.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:20:44 -04:00
Grant Whitmer
587fb05265 SECURITY: enforce passport status — revocation now takes effect live
All checks were successful
check / gate (push) Successful in 21s
A revoked passport returns HTTP 200, status=revoked, band=unproven,
allowed_actions=[] (verified live 2026-08-13). resolve_passport keyed refusal
only on HTTP 4xx and band=="untrusted", so it returned band 'unproven' and the
agent was seated. Revocation was NOT enforced on the live auth path at all — and
now that agent auth actually works, a revoked agent could authenticate and act.

Extracts decide_trust(body) -> (band, actions) | raise. Only status=="active"
is allowed; revoked/suspended/frozen/unknown all refuse, fail-closed on the
field that carries the most consequential fact about an identity. The agent call
site turns that into a clean 403 passport_revoked.

This is the REAL revocation gate — the token cannot be un-issued, but its
standing is re-checked on every request, so revocation takes effect on the next
call with no webhook required. The Eternitas webhook remains useful for
invalidating locally-issued credentials/grants (G6.3, not built yet), but it was
never the primary gate and its being unwired is no longer a live exposure.

Behavioral tests: revoked body refused, active accepted, unknown/missing status
fails closed.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:18:16 -04:00
Grant Whitmer
79dcea4d3e G3.2/G3.4: real EPT verification + wire the throttle, reopening agent auth
All checks were successful
check / gate (push) Successful in 19s
canary / probe (push) Successful in 10s
REOPENS the agent path — but only because possession is now actually proven.

EPT verification (api/app/ept.py): ES256 against Eternitas's published key set
at /.well-known/eternitas-keys. algorithms=["ES256"] makes alg:none and
algorithm confusion unrepresentable rather than merely unlikely; issuer and exp
are enforced by the library; an unknown kid is refused.

Order is deliberate: signature FIRST, trust lookup second. These EPTs live ~365
days and carry rev/tru baked in at issuance, so a year-old "rev: false" proves
nothing — revocation and band still come from a live lookup on every request.

Found while building it: real EPTs put the passport in the "sub" claim. The old
code read "passport"/"sub_passport", which no genuine EPT carries — so real
agents were never recognised and ONLY forged tokens ever authenticated. The
bypass was not just a hole, it was the only thing that worked.

Throttle (api/app/throttle.py): BAND_MULTIPLIER and rate_*_per_day were defined
and read by nothing. Now enforced on repo.create and grant.create, counted
against agent_actions (one source of truth, not a private counter that drifts
from the audit log). Fails CLOSED — a limiter that fails open protects you until
the moment something is wrong. Untrusted band is 403 read-only, not 429, because
"slow down" would be a lie.

Tests: 14 behavioral, signing real ES256 tokens with a locally-generated key so
they exercise the crypto path with no network dependency — genuine tokens
accepted, and alg:none / foreign key / tampered payload / expired / wrong issuer
/ unknown kid / missing claims all refused. 74 green.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:09:47 -04:00
Grant Whitmer
d8deffe4db SECURITY: close the agent-auth bypass — trust is not authentication
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 11s
Verified live 2026-08-13: a forged 'alg:none' token naming a passport lifted
from the logs returned HTTP 200 as that agent. The agent path read the passport
without verifying the EPT signature, asked Eternitas 'is this passport
reputable?', and seated the caller on a yes. That answers reputation, not
possession — anyone who knows a passport number could impersonate that agent on
the public API.

The human path already failed closed for exactly this reason
(require_verified_jwt). The gate was on the wrong path: it sat AFTER the agent
branch returned. The agent path now fails closed too, BEFORE the trust lookup,
so a forged token never even reaches Eternitas. Reopens automatically when the
ES256/JWKS verifier (G3.2/G9.1) exists and this gate consults it.

Adds BEHAVIORAL tests (not string-grep): a forged alg:none token exercised
through the real get_caller must raise, not authenticate. This is the test that
would have caught the bypass; the suite had 86 source-string assertions and
zero that ran the auth decision.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 22:49:45 -04:00
Grant Whitmer
2c292c501c G0.9: nightly timer + the first rehearsed restore in the ecosystem
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 23s
Timer installed with Persistent=true — Grant's workstation is not always on at
04:17, and without that a missed window is silently skipped and the backup
simply never runs.

RESTORE DRILL PASSED, which is the part that matters: pulled a bundle back out
of R2, cloned from it, and the restored HEAD (ffce529) matches live origin/main
exactly — 40 commits, all branches, 50 files.

The August audits found no rehearsed restore anywhere in the ecosystem, for
anything. This is one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 21:49:11 -04:00
Grant Whitmer
ef3450d87b G7.6: never let bookkeeping kill the monitor
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 26s
An unwritable state path raised and took the whole canary down. That is the
worst possible trade for a monitoring tool: it reports nothing at all, and
reports it silently. State is an optimisation for transition detection; the
probing is the point.

Found by fat-fingering an env var, which is exactly how it would happen in
production.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 15:42:58 -04:00
Grant Whitmer
a094960da3 G3.5: accept platform.test_ping unverified — unverifiable by construction
All checks were successful
check / gate (push) Successful in 19s
canary / probe (push) Successful in 26s
Read the sender rather than guessing: Eternitas generates the webhook secret at
registration time and pings the URL to prove reachability BEFORE returning that
secret. The ping IS signed — with a secret the receiver cannot possibly hold
yet. Unverifiable by construction, not by oversight.

Accepting it is safe because the event is definitionally a no-op: nothing read,
nothing written, acted:false. Every event that changes anything still requires a
valid HMAC. The alternative, skip_validation:true, would permanently disable
reachability checking for this platform to solve a one-time ordering problem.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 15:36:03 -04:00
Grant Whitmer
5dd914b8a6 G3.5: answer the reachability probe honestly instead of skipping validation
All checks were successful
check / gate (push) Successful in 18s
Eternitas verifies a webhook URL answers BEFORE issuing the secret that signs
deliveries, so the very first request can never carry a signature — refusing it
makes registration impossible. Real chicken-and-egg, not a reason to disable
validation.

A probe is a request claiming no event and carrying no signature. Answering it
200 is honest: the endpoint exists and is ready. It changes nothing (acted:
false), and anything claiming to BE an event still goes through full HMAC
verification. Registering with skip_validation:true would have permanently
disabled a safety check to solve a one-time ordering problem.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 15:33:56 -04:00
Grant Whitmer
339ec70853 G3.5: Eternitas revocation receiver — fail-closed, both signature traps avoided
All checks were successful
check / gate (push) Successful in 18s
When a passport is revoked, every credential it holds here dies in one
transaction: tokens revoked, grants revoked. A revocation that takes effect
'eventually' is not a revocation.

Avoids two traps that each cost a sibling service a subscription that looked
wired and never once delivered:
  1. strip the 'sha256=' prefix before comparing — comparing the decorated
     header against a bare digest returns 401 forever
  2. HMAC the RAW REQUEST BYTES, never a re-serialised body — JSON.stringify of
     a parsed body reorders keys and changes whitespace, so the digest never
     matches what the sender signed

Both fail silently from the sender's side: Eternitas records a delivery, the
receiver records a rejection, nobody notices for weeks.

Unset secret REFUSES rather than accepts — accepting unverified instructions
about identity is worse than missing them. And it never acknowledges a
revocation it could not apply; a 200 there is a security hole reporting success.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 15:32:27 -04:00
Grant Whitmer
fc1937560c G7.6: fix the alert path — urllib UA was rejected 403 by Resend
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 22s
Caught by TESTING the alert path instead of assuming it. Without an explicit
User-Agent, urllib sends 'Python-urllib/3.x' and Resend rejects it 403, while
the identical request via curl succeeds.

The failure mode this avoids is the worst one a canary has: it would have
detected every outage correctly and told nobody. Same bot-filtering trap as the
Gitea migrate call earlier today — worth recognising on sight.

Also prints the HTTP body on failure. '403 Forbidden' alone sends you hunting
for a bad key; the body names the real cause.

Verified: alert sent (200).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:44:00 -04:00
Grant Whitmer
9a7030351b G7.6: fleet canary — probes what a user does, not what is cheap
All checks were successful
check / gate (push) Successful in 19s
Today's outage is the whole design brief: /health returned 200 for the entire
hour that login was dead. A canary watching /health would have stayed green
while nobody in the ecosystem could sign in. So the login probe is here and it
is the one that matters.

Three rules it obeys:
  - never green for something it did not prove (I-8)
  - alert on TRANSITIONS, not every run — a canary people filter is a dead
    canary, which is how the last one sat 37 days dead unnoticed
  - run where the watched thing cannot take it down: Veron 1, never Kit 0

Two independent signals, so losing one still leaves the other: an email via
Resend on state change, and a non-zero exit that turns the CI run red in the
forge itself.

Alerts say what broke in human terms — 'a human can actually sign in' — rather
than only naming an endpoint.

Verified against production: 7/7 green including login at 17.1s; a forced 404
reports DOWN; a 1s threshold reports SLOW at 23.3s.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:41:37 -04:00
Grant Whitmer
6b96087554 G7: per-job networks — fixes service DNS and tightens isolation
Some checks failed
check / gate (push) Failing after 13m33s
The migration step failed with 'could not translate host name postgres'. The
Postgres service container was healthy; the job simply could not name it,
because service DNS aliases only exist on a per-job network and I had pinned
container.network to the flat 'bridge'.

That choice was wrong in both directions: it broke service containers AND it
was weaker isolation, since every concurrent job shared one bridge and could
see its neighbours. A per-job network is stricter and correct — and still has
no route to the forge, because these networks live inside the dind daemon,
which has no forge attachment at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:18:51 -04:00
Grant Whitmer
8c4bdd2bc2 G7.3: setup-python on the act image, not a python image
Some checks failed
check / gate (push) Failing after 43s
Trading one wedge for another: running the job in python:3.12-bookworm fixed
the resolver spiral but broke checkout, because actions/checkout is a
JavaScript action and the official Python images carry no node — 'executable
file not found in $PATH'.

setup-python on the act image has both. Test now accepts either a pinned
container image or an explicit setup-python version, and still checks it
against pyproject's requires-python.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:16:00 -04:00
Grant Whitmer
980fc2dddd G7.3: pin Python 3.12 in CI and cap job runtime
Some checks failed
check / gate (push) Failing after 7s
The first real CI run wedged for 14 minutes. Not a network problem, not the
isolation work — catthehacker/ubuntu:act-22.04 ships Python 3.10.12 while this
project declares requires-python >=3.12, and pip answered that by backtracking
through the entire release history of every dependency looking for something
3.10-compatible. At 100% CPU, with -q hiding every line of it, and it would
have churned until the runner's 30m timeout.

A version mismatch presenting as a hang rather than an error is worth a test,
so there is one: the workflow's python image must satisfy pyproject's
requires-python, checked by parsing both rather than by eyeballing them.

Also: every job now has timeout-minutes. A wedged step should be a red check in
minutes, not an occupied runner for half an hour.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:13:39 -04:00
Grant Whitmer
53dea84ab1 G7: route CI jobs via the public surface, not the forge network
Some checks failed
check / gate (push) Failing after 13s
The first CI run failed with 'Could not resolve host: gitea' — job containers
live on dind's private network and cannot see the forge network. Two ways out,
and they are not equivalent:

  (a) put job containers on the forge network. Easy, one line, and it leaves
      untrusted workflow code one DNS name from the forge's Postgres. It quietly
      repeals I-5.
  (b) send jobs to the PUBLIC forge surface over the tunnel, exactly like any
      stranger on the internet.

Took (b). The runner no longer needs the forge network at all, so there is now
NO private route from any CI container to anything — a better posture than this
file started with. Cost is a hairpin through Cloudflare plus its ~100s ceiling
per fetch, which for 0.63 GB of objects across 61 repos and depth=1 checkouts is
nowhere near binding.

Test upgraded to assert the stronger property: no CI container joins the forge
network.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 10:51:00 -04:00
Grant Whitmer
5f717ef74b G7: CI runners with real isolation, and the gate as a workflow
Some checks failed
check / gate (push) Failing after 50s
I-5 says runners execute untrusted code and must be isolated by machine
boundary. act_runner needs a Docker daemon to start job containers, and the
tempting move — what every published example does — is to mount the host's
/var/run/docker.sock. That hands every workflow, including whatever a
transitive dependency's postinstall script feels like doing, the ability to
start a privileged container mounting / — root on Grant's workstation.

Instead the runner talks to its OWN dind daemon:
  - runner (TRUSTED, the act_runner daemon) sits on the forge network only to
    collect jobs from gitea:3000
  - dind and every job container it spawns are UNTRUSTED, on a private network
    with no route to the forge, its Postgres, or its .env
  - jobs cannot bind-mount from the daemon host (valid_volumes: []) and are not
    handed the runner's own socket (docker_host: -)
  - separate compose project, cpu/memory bounded — Veron 1 is Grant's
    workstation, not a dedicated build box

The gate itself now runs as a workflow, including the migration round-trip that
already caught two bugs review did not, and the I-12 check that a COMMIT_SHA
env override cannot change what /version reports.

Labels are explicit and pinned. A workflow naming a label nobody provides
queues forever and presents as a hung CI system rather than a typo — which is
what ubuntu-latest does on every windy-registry run today.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 10:39:05 -04:00
Grant Whitmer
a2430de94d I-4: distinguish a mirror that never ran from one that is behind
Gitea reports the epoch for 'not yet synced', which arithmetic turns into a
56-year lag and a confident 'degraded'. Collapsing those two states is how a
backup that was never made gets read as a backup that is merely stale — which
is the more dangerous direction, because 'behind' sounds survivable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:16:05 -04:00
Grant Whitmer
da257652c2 G11 / I-4: continuous off-site mirror, and a namespace bug fixed
I-4 said 'never a one-way door' and had no implementation. Now it does.

  - ensure the GitHub counterpart exists (idempotent), then ask Gitea to keep
    it in step with sync_on_commit=True. An hourly timer means an hour of work
    can be the thing you lose, and that window is invisible until it costs you.
  - mirror status reports what is TRUE including 'we do not know'. An
    unconfigured mirror reports unconfigured, NEVER healthy — same posture as
    me-fleet.ts refusing to say 'online' when it only knows 'registered'.
  - lag past the threshold is a P2, not a shrug. A mirror nobody checks is a
    belief, not a backup, and this ecosystem already lost 37 days to a canary
    everyone assumed was fine.

Gitea owns the replication rather than a hand-rolled loop, because a background
job that fails silently is exactly how the registry's integrity refresh spent
its entire life calling a 404 and incrementing a counter instead of raising.

Also fixes a real bug I had written myself: list_versions derived the Gitea
namespace from the CALLER, which is correct only while the caller is the owner
and addresses the wrong namespace the moment a collaborator asks — surfacing as
'not found', which is the hardest kind of bug to see. Now derived from the repo,
with a test that keeps it that way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:11:45 -04:00
Grant Whitmer
066de34489 G5: fix version-count copy; record the u-system namespace finding
'There are 1 saved versions' is exactly the sloppiness the vocabulary law
exists to catch. Copy is design material, not decoration.

G5.9 records something the live test surfaced: a repo created through
X-Service-Token lands in a 'u-system' namespace because the service caller has
no identity of its own. Correct for /internal plumbing, WRONG for anything a
person owns — the portal must pass an acting user and this cell must refuse to
create a user-owned project without one. Until then service-created repos are
ops artifacts, not customer data.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 15:59:41 -04:00
Grant Whitmer
3a9259a0da G5: generate an unusable password on Gitea user create
Gitea rejects a null password with a bare 400. These accounts are never
password-authenticated — humans arrive via OIDC, agents via passport-bound
scoped tokens, local password sign-in is disabled server-wide — so we generate
a credential that is never stored, returned or recoverable. An unusable
password is safer than a blank one or a shared default.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 15:55:11 -04:00
Grant Whitmer
099e4be9b1 G5: the shelter — repos, grants and version history
The plane Windy Cloud does not have. Verified 2026-08-11: routes/storage.py and
its models contain ZERO occurrences of share/permission/acl/collaborat/seat/
version/snapshot/history/revision. This fills a hole rather than bolting onto
something that already had one.

  - repos: create/list/get, repo_type required (I-7), reserved slugs, Gitea
    reached ONLY through the membrane client (I-1)
  - grants: human identity OR agent passport, exactly one enforced by a database
    CHECK constraint; agent grants expire in 90 days by default
  - versions: history in words a person recognises — no 'commit', no 'branch',
    no 'repository' in any user-facing string (D-9/I-9), with a test that greps
    the speak strings and fails on developer vocabulary
  - private repos 404 rather than 403, so a stranger cannot learn one exists

Auth: three first-class caller classes (human OIDC / agent EPT / internal
service token), NO fourth, and no bypass env var — copied deliberately from the
desktop control server, the ecosystem's best Principle-#5 artifact.

G3.6 status-code law implemented: 400 and 404 REFUSE, 429/5xx retry then REFUSE.
A sibling maps 400/429 to 'unreachable' and soft-ALLOWS, which is inducible —
an attacker who wants the check skipped only has to make it rate-limit itself.
A test asserts resolve_passport has exactly one return path.

And I-8 applied to ourselves: G3.2's JWKS verifier does not exist yet, so the
human token path REFUSES in production rather than accepting an unverified JWT.
An unverified JWT is an authentication bypass, not a shortcut.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 15:45:22 -04:00
Grant Whitmer
82044933ed G2/G4 complete: Gitea live, LFS landing in R2, four traps pinned
VERIFIED END TO END from outside the network:
  - create repo via API -> clone -> commit -> push -> read back over HTTPS
  - 3 MB LFS object pushed through the tunnel, landed in R2 at lfs/34/2d/...
  - NO local lfs/ directory on the host: I-3 confirmed by measurement
  - /health/full: db, gitea and r2 all green

Adds strand G4A recording four traps that each cost a crash loop, with tests:
  1. [lfs] STORAGE_TYPE creates a separate storage section that does not
     inherit [storage] — crash loop, error names the symptom not the cause
  2. storage backend != LFS enabled; LFS_START_SERVER is separate, and its
     absence reads as a permissions error
  3. Gitea env-to-ini SETS but never UNSETS — removing a compose var leaves the
     line in the persisted app.ini, so repo config and prod config silently
     disagree. Exactly the drift this cell exists to end.
  4. R2 rejects the default S3 checksum algorithm

And G4A.5, which is architecture rather than a bug: an 8 MB non-LFS push died
with HTTP 524 at Cloudflare's ~100s limit. This makes the LFS threshold
load-bearing and REQUIRES G10 to serve model weights via presigned R2 URLs
rather than proxying blobs through the tunnel — client straight to R2, which is
what Hugging Face does and which takes Grant's home upstream out of the path.

G2.4: Gitea's MIT text and a NOTICE now travel with the repo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 15:13:58 -04:00
Grant Whitmer
ce54d488f2 G1: stop probing the tunnel from inside a container
cloudflared binds 127.0.0.1:2000 on the HOST. This process runs in a container
whose only route to the host is the bridge gateway (172.17.0.1), where nothing
is listening — so the check was permanently red regardless of what the tunnel
was actually doing.

Binding the metrics endpoint wider would have fixed the probe and made a
metrics bind failure capable of taking down ingress. That is a worse trade than
losing one row on a dashboard.

The check is not silently dropped: /health/full now carries a 'not_checked_here'
map naming the tunnel and where its health actually lives (systemd
windygit-tunnel). An observer should never have to wonder whether a missing
check means healthy or means forgotten.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 14:37:53 -04:00
Grant Whitmer
a68261a563 G1: Veron 1 host live behind Cloudflare Tunnel
app.windygit.com / api.windygit.com / models.windygit.com are serving over
HTTPS with ZERO inbound ports open on Grant's network.

  - tunnel 4e856c5d, 4 registered edge connections, systemd-managed and bounded
  - three proxied single-level CNAMEs (Free Universal SSL covers them; a
    two-level name would need ACM and would die in the TLS handshake)
  - services bound to 127.0.0.1 with configurable host ports — Veron 1 is
    Grant's workstation and 3000/3300 belong to other projects
  - docs/RUNBOOK-VERON.md

I-12 PROVEN IN PRODUCTION: /version reports source=baked with a sha equal to
the deployed HEAD.

Also fixed: the tunnel health probe targeted localhost from inside a container,
so it was permanently red. A check that is always red is as useless as one that
is always green — it is how a fleet canary goes 37 days dead unnoticed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 14:34:18 -04:00
Grant Whitmer
659991b2bd G0: cell substrate — invariants made executable
Strand G0 complete and VERIFIED against real Postgres, not asserted.

  - FastAPI plane, fail-closed provider seams, repair-pointer error taxonomy
  - migration 001: all 10 tables incl. repo_type NOT NULL and model_cards (I-7)
  - 17 invariant tests, ruff clean, vocabulary audit clean

Two bugs found by RUNNING it that review would not have caught:

  1. SQLAlchemy Enum persists .name, not .value — so RepoState.deleted_soft
     and CreatedVia.imported would have written labels migration 001 never
     declared, failing at runtime rather than at review. Pinned via
     values_callable.
  2. op.create_table asks each Enum to emit its own CREATE TYPE with no
     checkfirst, so the second reference raised DuplicateObject and the
     migration died halfway. Types are now created once, referenced with
     create_type=False.

Proven live, with the hostile env var set:
  - I-12: COMMIT_SHA=deadbeef... in the environment, /version reports real HEAD.
    That env pin is the documented root cause of nine sibling services
    misreporting their commit; here it is structurally ignored.
  - I-8: three unconfigured providers -> status degraded, HTTP 503, each saying
    'refusing to report healthy'. No mock, no false green.
  - G0.4: upgrade -> downgrade -> upgrade round-trip clean (10 -> 0 -> 10).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 14:19:28 -04:00