101 Commits

Author SHA1 Message Date
Kit OC5
44e824c762 push velocity: correlate the SSO-id subquery on the grouped column
Dry-run against the real gitea DB: 'subquery uses ungrouped column u.id'.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 14:06:10 -04:00
Kit OC5
4e2db1e2e0 push velocity: isolate its query so a failure never costs ci.run rows
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 13:06:20 -04:00
Kit OC5
54c63b4d89 push velocity: key humans on windy_identity_id (SSO link); no id -> system + caller
Telemetry UPDATE 2 actor rule: agent/human rows without actor_id are
quarantined. Forge humans sign in only via Windy SSO, so Gitea's
external_login_user.external_id is their windy_identity_id.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 13:06:17 -04:00
Kit OC5
ac83c4721e telemetry: detect push velocity from Gitea's action table (alert only)
git push never touches our API, so throttle.py can't see it. Gitea's
action table records every push; the 5-min sync-side emitter now reads
it and emits forge.push_velocity when an account crosses 60 pushes/1h,
500 pushes/24h (standard-band base) or 10 ref deletes/24h. One row per
account per rule per window while over; windyadmin (the sync) exempt.
Nothing sits in the push path and nothing is refused. HOLD until
Telemetry Boss declares the shape.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 13:06:17 -04:00
Kit OC5
2d4fadb090 sync: bound the janitor and telemetry steps (docker exec hangs in an IO stall)
Some checks failed
canary / probe (push) Has been cancelled
check / gate (push) Has been cancelled
09-23 16:43Z the Veron data2 SMR stall left runc exec in D state; the
janitor's docker exec never returned, so the sync sat 'activating' and no
repo mirrored or got a status for any lane. Both steps are non-fatal;
now they time out (120 s / 180 s) and the run continues.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:54:34 -04:00
Kit OC5
e3b69fa759 telemetry: UPDATE 7 — read the ingest body; count quarantined + dropped on heartbeats
Some checks failed
canary / probe (push) Has been cancelled
check / gate (push) Has been cancelled
The ledger answers 202 even when it quarantines rows. Both emitters now log
a warning with the reasons and report service.health.telemetry_quarantined
and telemetry_dropped (API: buffer overflow; sync: 0 by construction, since
a failed send keeps cursor + spool). HOLD until Telemetry Boss declares both
keys on windy-git's two service.health shapes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:32:28 -04:00
Kit OC5
4acf50d9ef bridge tests: fake serves workflow contents; cover invalid-workflow status
All checks were successful
check / gate (push) Successful in 23s
b7a7e94 made the bridge read workflow files, which the strict fake Gitea
refused (7 red). The fake now serves contents (404 when absent), and new
tests cover: error posted with no runs, valid files add nothing, no repost,
.gitea/workflows wins over .github/workflows, and each workflow_problem
shape. pyyaml declared in dev extras (the bridge imports it).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:31:19 -04:00
Kit OC5
b7a7e94df0 bridge: post an error status when Windy Git ignores an invalid workflow
Gitea drops an invalid workflow file with one log line and fires no run, so
the GitHub PR showed nothing and lanes waited for CI that never came
(windytalk #100). The bridge now reads each workflow file at the commit it
reports on and posts windy-git/<wf>/workflow = error with the reason.
Verified: 0 false positives on all 23 bridged repos' main; catches
windytalk #100's broken commits (invalid YAML at line 12), fix commit clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:29:12 -04:00
c83f808a60 bridge: retry transport blips (TLS timeout/reset), never HTTP errors
All checks were successful
check / gate (push) Successful in 20s
canary / probe (push) Successful in 9s
A single GitHub TLS handshake timeout failed the whole sync, flipped its
windy-job heartbeat to ok:false and would page for nothing. Up to 3
attempts with backoff for URLError/timeout/reset; HTTP errors return
immediately as before. Test covers both.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:15:22 -04:00
eb27e63db3 telemetry: adopt the end-to-end synthetic convention (UPDATE 4)
All checks were successful
check / gate (push) Successful in 24s
canary / probe (push) Successful in 9s
Replaces the keyed marker from 1c3b5b0 with the ecosystem convention:
any X-Windy-Synthetic value marks the request synthetic; the flag lives in
a per-request contextvar, labels this request's rows, and is FORWARDED on
downstream calls (Eternitas trust lookup, Gitea API). The canary sends
"1". Rows are still recorded; the label separates, never suppresses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:06:51 -04:00
1c3b5b0638 telemetry: synthetic:true on canary refusals (keyed, not a bare flag)
All checks were successful
check / gate (push) Successful in 25s
canary / probe (push) Successful in 7s
The canary deliberately sends forged tokens every 10 min; those refusal
rows read as attacks. It now sends X-Windy-Synthetic carrying a shared
secret (Gitea repo secret CANARY_SYNTHETIC_KEY = WINDYGIT_SYNTHETIC_KEY in
Veron .env); the API marks the row synthetic only on a constant-time
match, so an attacker cannot label their own refusals synthetic to hide.
synthetic is declared on forge.auth.failed (Telemetry Boss, UPDATE 3).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:54:56 -04:00
90643fe48e telemetry step 2: API boot/health + forge.auth.failed (declared)
All checks were successful
check / gate (push) Successful in 25s
canary / probe (push) Successful in 6s
Membrane first: I-2 and MEMBRANE.v1 now list the windy-admin ledger
(POST /v1/events). api/app/telemetry.py: service.boot once per start
(commit_sha omitted when unknown, I-12), an hourly in-process
service.health with the shared keys (requests, errors_5xx/4xx,
refusals_4xx, p95_ms only when there was traffic), and one
forge.auth.failed row per refused request: declared 13-code enum,
http_status, caller class, route TEMPLATE (never the concrete path),
actor_type system with no actor_id (all-lanes rule). No token = nothing
sent or buffered; flush failures keep rows (bounded) and never raise.
Token from root-only /etc/windygit/telemetry.env (optional env_file).
8 behavioural tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:47:29 -04:00
00ec963f82 telemetry: fix ci.run completeness — cursor on (finish time, job id)
Some checks failed
check / gate (push) Has been cancelled
Telemetry Boss found jobs_finished=43 vs 8 ci.run rows. Root cause: the
high-water mark was the job id, but jobs FINISH out of id order, so every
long job that started before the mark and finished after it was silently
never emitted. Now a (finish time, id) cursor; finish = stopped, or
updated for skipped jobs with no stop time. Heartbeat finished/failed/
cancelled counts are derived from exactly the rows emitted, so
sum(jobs_finished) == count(ci.run) by construction (dry run on real
data: 97 == 97, failed 2 == 2, cancelled 13 == 13). posted_to_github now
set from the bridge's own rules. duration_ms = Gitea whole seconds x 1000.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:42:36 -04:00
b5e4eaf57a telemetry: ci.job_cancelled from the janitor; interval_s on heartbeat
All checks were successful
check / gate (push) Successful in 22s
canary / probe (push) Successful in 6s
The janitor now returns one JSON line per job it cancels (repo, workflow,
job, reason, runs_on, waited_s) into a spool; the emitter ships them as
ci.job_cancelled (declared with Telemetry Boss) and truncates the spool
only after a 2xx. Run status recompute folded into the same statement.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:27:53 -04:00
baaa542bae docs: audit disposition 09-23 — R2 god token replaced by bucket-scoped token
All checks were successful
check / gate (push) Successful in 43s
canary / probe (push) Successful in 7s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:15:35 -04:00
0634a6cb1b style: ruff fix in telemetry_emit
All checks were successful
check / gate (push) Successful in 29s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:08:21 -04:00
1a171eabd5 telemetry: CI emitter for admin.windyword.ai (inert until token)
Some checks failed
check / gate (push) Failing after 20s
canary / probe (push) Successful in 6s
ci.run (one row per finished job, exactly once via a high-water mark;
branch_kind default|pr|other so the dashboard can show "main is red") and
service.health (interval counts: finished/failed/cancelled, waiting,
running, runners online, oldest wait). Shapes declared with Windy
Telemetry 40; sends nothing until WINDYGIT_TELEMETRY_TOKEN exists.
State is only advanced after a 2xx, so a failed post retries.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:06:54 -04:00
28c31236b8 ci: janitor also clears jobs blocked forever on failed needs
When a needed job fails, Gitea leaves dependants BLOCKED (7) even after
the run finishes; eternitas build jobs sat there 8h. Mark them skipped
(what GitHub shows) once the run is done and 30 min have passed.
Found by the new telemetry dry run (oldest_waiting_s = 29160).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:06:33 -04:00
cd5967031b ci: janitor cancels jobs no runner can ever take
All checks were successful
check / gate (push) Successful in 20s
canary / probe (push) Successful in 6s
windy-pro alone left ~4 jobs per run waiting forever (build-electron on
macos/windows/ubuntu-latest, deploy if:false): Gitea evaluates job if:
only at pick time, the labels do not exist here, and waiting jobs are
invisible in /actions/tasks. 37 such jobs across 10 runs today. After
30 min they are cancelled and the run status recomputed. Runs each sync,
non-fatal.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 10:58:53 -04:00
f246417095 ci: windy-pro desktop/installer jobs are NON-BLOCKING (Grant, 09-23)
build-desktop, test-installer and reality-check still run on Windy Git
and stay visible there, but the bridge no longer posts them to GitHub, so
they cannot turn windy-pro's combined status red. Windy Git side only;
the desktop code is Grant's to fix. BRIDGE_NON_BLOCKING, per repo.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 10:58:07 -04:00
50c1464043 security: never bundle credential repos to R2 in plaintext
All checks were successful
check / gate (push) Successful in 23s
canary / probe (push) Successful in 7s
kit-army-config (the lockbox) and every *-soul / anima repo carry
credentials; the nightly R2 bundles are unencrypted, so the R2 key was a
key to every secret. Excluded by name (BACKUP_EXCLUDE); they are backed up
encrypted by the Windy Drops lane (restic) and stay mirrored on Veron.
Behavioural test runs the script's own exclusion function.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 10:48:42 -04:00
7a63f90da3 ci: bridge windy-mind (private) verdicts to GitHub
All checks were successful
check / gate (push) Successful in 23s
canary / probe (push) Successful in 6s
windy-mind has been writable + CI on Windy Git since 08-13 (deploy.yml
disabled, uv pinned); it only lacked GitHub commit statuses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 09:59:28 -04:00
b8f97f0731 ops: sync never pushes archive/* branches to Windy Git
All checks were successful
check / gate (push) Successful in 20s
archive/<machine>-<date>/<branch> are off-machine safety copies of
unpushed work (one-repo doctrine). GitHub holds them; running CI on them
is waste. Negative refspec ^refs/heads/archive/* on the push (git 2.43
on Veron). Requested by 8c for windy-pro's Mac mini archive.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 09:57:57 -04:00
95c33c8004 ops: host systemd units in git; windy-pro tags never reach Windy Git
All checks were successful
check / gate (push) Successful in 22s
- deploy/systemd/: sync/backup timers+services, tunnel, and the windy-job
  heartbeat drop-ins (silent-failure audit). They existed only on Veron,
  the same drift that left the runbook wrong. GITHUB_TOKEN is stripped
  (repo is public); it stays in the root-only unit on the host.
- sync: SYNC_NO_TAGS (default windy-pro). build-electron fires on v* tags
  and targets ubuntu/macos/windows-latest, labels no runner has, so it
  would queue forever, invisibly. Desktop releases are built elsewhere.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 09:51:51 -04:00
d5181f1c6d ci: G11.5 resolved — onboard windy-pro; cloud cells + windytalk; promote script
All checks were successful
check / gate (push) Successful in 21s
canary / probe (push) Successful in 6s
- import_from_github.py no longer refuses windy-pro: lane 8c audited all
  14 checkouts (WINDYPRO_CHECKOUTS.md), GitHub main is canonical.
- sync + bridge: windy-cloud-domains, windy-cloud-vps, windytalk (default
  branch master), windy-pro; windy-cloud-sites added to the bridge (it was
  synced but never bridged).
- scripts/promote_to_ci.sh: the mirror->writable procedure as one script,
  with the delete-before-import hazard documented.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 09:49:15 -04:00
e6530d3171 test: behavioral G3.5 webhook tests (audit: tests were source-string asserts)
All checks were successful
check / gate (push) Successful in 20s
canary / probe (push) Successful in 9s
Seven HTTP-level tests through the real route: sha256= prefix and bare
digests accepted, digest of re-serialised JSON refused, forged/wrong-key/
missing signatures refused, unset secret -> 503, a revocation with a bad
signature never reaches the handler, the reachability ping never acts.
Mutation-checked: dropping the prefix strip fails the behavioral test
while the old string-grep invariant still passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 03:26:14 -04:00
419443573a ci: onboard windy-hand; stop running windy-agent CI twice
All checks were successful
check / gate (push) Successful in 23s
windy-hand promoted to writable + bridged. windy-agent is PUBLIC and its
GitHub Actions already run on Veron's GitHub runner; running its 3-version
pytest matrix here too was pure duplicate load (3 x ~4.5 cores for 25+
min, host load 64 on 24 cores). Actions are now off for windy-agent on
Windy Git; it stays in REPOS as a synced copy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 03:23:25 -04:00
4a34b35441 security: CI egress filter — jobs reach the internet, never Veron/LAN
Measured: an unprivileged job container inside the CI dind could open SSH,
Ollama and dev servers on Veron (192.168.1.73) and the rest of the LAN,
WireGuard and Tailscale — lateral movement for any malicious dependency,
no escape needed. egress.sh (idempotent; windygit-ci-egress.service at
boot) hooks the jobs bridge: runner<->dind, replies, DNS and public
egress allowed; RFC1918, CGNAT, link-local and the host itself dropped.
Verified from a job container: 6/6 private targets blocked, DNS,
internet and the public forge OK.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 03:20:45 -04:00
dfe5543eda docs: runbook + AGENTS match reality (item 5 of the launch bar)
Some checks failed
check / gate (push) Has been cancelled
canary / probe (push) Successful in 34s
RUNBOOK-VERON: deploy uses fetch + merge --ff-only and api-only rebuilds
(the old text used git pull, contradicting its own warning); new sections
for host timers, CI (6 runners x1, windyadmin-scoped, 90m ceiling, queue
truth in the gitea DB, logs in R2), sign-in posture and break-glass;
standing checkout = OC5. AGENTS.md no longer says GENESIS / no code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 03:19:12 -04:00
40cb455d0d ci: onboard windy-translate, windytranslate-site, windytraveler-site
Some checks failed
check / gate (push) Has been cancelled
Promoted to writable; the two sites' CF deploy.yml disabled (they would
need the CF god token in a job container). All three only have
ubuntu-latest workflows today, so nothing runs until their lanes switch
runs-on to [self-hosted, linux, x64].

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 03:13:40 -04:00
8c404eb410 security: turn off Gitea OAuth auto-registration
Any Windy Word account (public signup, unverified email) auto-registered a
forge account on first sign-in — reproduced with a throwaway account —
and the act runners are instance-wide, so a stranger's workflow would run
on Veron's privileged dind beside the R2 god token. §7 makes opening the
forge to non-Grant users Grant's call.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 03:09:43 -04:00
5b16114b98 auth: token contract v1 (aud windy_git, both issuers); CI for eternitas
Some checks failed
check / gate (push) Successful in 37s
canary / probe (push) Has been cancelled
- hub_jwt: aud list is ["windy_git"] (contract v1 array). Dropped
  "windy-git": that is Gitea's OIDC client_id, so a forge id_token would
  have passed the aud check. `type: human` is now REQUIRED (id_tokens have
  none), which makes accepting the discovery-URL issuer safe.
- runner job ceiling 30m -> 90m: eternitas's serial pytest is ~50 min and
  would have been killed mid-suite.
- eternitas (private) added to the GitHub status bridge.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:58:56 -04:00
18ea9a4686 auth: G3.2 hub JWKS verifier — humans can sign in to the plane (SSO #14)
The human path refused every token in production (503
human_signin_not_ready) because no verifier existed. api/app/hub_jwt.py
verifies hub access tokens against account.windyword.ai's JWKS:

- RS256 only (closes alg:none and HS256-with-public-key confusion)
- iss must be "windy-identity" — what hub ACCESS tokens carry (observed
  live); id_tokens (discovery-URL issuer) are not accepted as bearers
- aud optional today, must name Windy Git when present; hub_require_aud
  flips it mandatory once the hub emits it. PyJWT's own aud check is off
  on purpose: it rejects ANY aud-bearing token when no audience is given.
- type must be human; identity = windy_identity_id, never sub (per-row id)
- production verifies even if require_verified_jwt is off

11 behavioral tests sign real RS256 tokens with a local key.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:55:17 -04:00
32e8ac8474 ci: bridge windy-registry (private) PR/main verdicts to GitHub
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:53:20 -04:00
8e9fa3116c ci: six runners; SSO #8 Gitea sign-in hardening (staged)
- runner-5/6: 50+ jobs were queued with ~11 private repos onboarded. dind
  keeps the 12-core ceiling, so this adds concurrency, not CPU.
- Gitea: password + passkey sign-in forms off (break-glass = CLI), and
  ACCOUNT_LINKING auto -> login. auto linked any hub login whose email
  matched an existing account, and SITE ADMIN windyadmin carries Grant's
  email. Grant is linked by the hub's stable sub, which matches first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:51:56 -04:00
74ad4950b2 ci: onboard windy-drops, windy-code-web, windy-code, windy-traveler
All checks were successful
check / gate (push) Successful in 1m1s
canary / probe (push) Successful in 38s
All four promoted from pull mirrors to writable. windy-code keeps only
canonical-domains-lint active: its other 15 workflows target hosted
macOS/Windows/ubuntu-latest runners and would queue forever here.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:49:00 -04:00
e7bbf9af51 ci: prune.sh must address dind over TCP (it has no unix socket)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:47:35 -04:00
45686283be ci: bound CI storage; don't bridge image-build jobs
- deploy/runner/prune.sh + windygit-ci-prune.timer (6h): age-based prune
  of the CI-only dind (containers, finished-job volumes, images/builder
  cache >7d) plus a hard 60 GB cap. Only that daemon, over its own TCP
  socket; never the host's Docker. It was 38 GB and unbounded — the same
  class of growth that filled Kit 0 on 09-01.
- pr_status_bridge: jobs named *docker* are not posted. Job containers
  have no daemon by design (I-5), so they are red on every commit; a
  permanent red X teaches everyone to ignore red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:47:28 -04:00
e4a15869c0 ci: onboard windy-connect + windy-search to private-repo CI
Some checks failed
check / gate (push) Has been cancelled
windy-connect promoted from pull mirror to writable (release.yml, which
publishes to PyPI on tag push, disabled — the sync pushes tags).
windy-search was already writable; its scheduled drift-check is disabled
because it now runs as cron on Kit 0. Both added to BRIDGE_REPOS.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:44:19 -04:00
dcf9286f16 ci: make Windy Git CI permanent for the private repos
All checks were successful
check / gate (push) Successful in 21s
canary / probe (push) Successful in 6s
- Four runners x capacity 1 instead of one x capacity 4. Concurrent jobs in
  one act_runner share /root/.cache/act; a refresh racing a copy killed 3 of
  windy-chat's ~20 jobs at setup-node (lstat ... no such file). Separate
  processes have separate caches. Same parallelism, same capped dind.
- Behavioral tests for pr_status_bridge (latest verdict wins, no reposting,
  skipped never painted green, fork PRs never run, pagination, PR lifecycle).
- import_from_github.py reads IMPORT_GITEA_URL, not GITEA_BASE_URL: sourcing
  the deploy .env pointed it at http://gitea:3000 and it died on DNS after the
  mirror it replaces had already been deleted.
- CUTOVER.md: the private-repo CI path, onboarding steps, and the
  /actions/tasks-hides-queued-runs trap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:18:04 -04:00
1b09b9b0d3 ci: sync windy-chat + windy-mail, bridge PR CI verdicts back to GitHub
All checks were successful
check / gate (push) Successful in 20s
canary / probe (push) Successful in 6s
GitHub Actions can't run on the private platform repos. Windy Git already
has their code and a working runner, so:

- windy-chat and windy-mail were read-only pull mirrors (which can never
  run Actions); they are now writable, deploy.yml/build-image.yml disabled,
  and synced from GitHub like the others.
- scripts/pr_status_bridge.py mirrors open same-repo GitHub PRs into Windy
  Git (so pull_request workflows fire) and posts each job's result back as
  a GitHub commit status (windy-git/<workflow>/<job>) on PR heads and the
  default-branch head. Fork PRs are never run. Runs after every sync.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:13:21 -04:00
390c1e7479 ops: move tunnel metrics to 2001, sync windy-git into itself
Some checks failed
check / gate (push) Successful in 19s
canary / probe (push) Failing after 7s
windygit-tunnel had crash-looped ~91k times: another project's
cornercall-tunnel holds 127.0.0.1:2000, and cloudflared exits when it
cannot bind its metrics port. Ingress only survived because a stray
cloudflared.service ran the same config. That unit is now disabled and
/etc/cloudflared/config.yml uses metrics 127.0.0.1:2001.

Also add windy-git to the GitHub->Windy Git sync list; its self-hosted
copy was stuck 3 commits behind (only check + canary workflows, no
deploys, so syncing is safe).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 01:29:57 -04:00
2b30b0ac99 docs: record the runner bump, the act cache flake, and the remaining reds
Second root cause found and fixed: act_runner 0.2.11 predates
`runs.using: node24`, so any repo on actions/checkout@v5 died before its first
step. Bumped to 0.6.1; Windy-Clone went 4/4 red to 4/4 green.

Also upgrades the act-cache note from "watch item" to a confirmed job failure
(lstat on a vanished file mid-tar), with the wipe command and the annotated-tag
dead end that looks like a wrong checkout but isn't.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 19:09:26 -04:00
cd7dd9b7ae ci: bump act_runner 0.2.11 -> 0.6.1 for node24 action support
0.2.11's bundled act only knows runs.using node12/node16/node20, so any repo
pinning a current action major (actions/checkout@v5, actions/setup-python@v6)
fails before its first step with "The runs.using key in action.yml must be one
of: [...], got node24". Windy-Clone is how this surfaced.

Verified node24 is absent from the 0.2.11 binary and present in 0.6.1, and that
every key in deploy/runner/config.yaml still exists in 0.6.1's schema.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 19:00:07 -04:00
9fc27eabd6 docs: record the real CI root cause — setup-uv resolves uv via the forge API
The localhost->postgres fix was correct but was never what failed these jobs;
they died at step 2. setup-uv v4+ resolves "latest" through GITHUB_API_URL,
which act_runner points at our own forge, so it 404s and every later step is
skipped by success(). Pinned an explicit uv version across 11 repos.

Also records the diagnosis traps that cost the most time: Gitea's job status
enum (1=success, 2=failure), act misattributing the error to the previous step,
the jobs-log API needing a repo-matched id, and Secure cookies defeating a
loopback curl login.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 17:57:04 -04:00
Grant Whitmer
db055a1922 docs: refresh the turnover prompt for the actual next task
Some checks failed
check / gate (push) Successful in 18s
canary / probe (push) Failing after 6s
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 17:31:23 -04:00
Grant Whitmer
86326ca6a7 docs: record three-repo CI fix results — 1 of 3 verified
All checks were successful
check / gate (push) Successful in 19s
windy-registry's postgres integration went failure -> success, proving the fix.
windy-mind and WindyCloud still fail for a cause I could not determine: the
jobs API returns 'job not found' for the ids the runs report, so logs were not
retrievable that way. Next session should read them from the Gitea web UI.

Records the trap that the three repos did NOT share one pattern — a naive
localhost->postgres swap would have left WindyCloud on port 15432.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 17:30:44 -04:00
Grant Whitmer
e0d5d118be docs: turnover for a fresh session
All checks were successful
check / gate (push) Successful in 35s
canary / probe (push) Successful in 8s
State, the immediate task (three-repo localhost->service-name CI fix), the traps
already paid for, and a copy-paste prompt.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 15:00:28 -04:00
Grant Whitmer
01e36155a3 ci: remove the service-networking probe; scope the python-pin test
Some checks failed
check / gate (push) Successful in 20s
canary / probe (push) Failing after 1m33s
The probe's own log was never retrievable through the jobs API, but the
question it asked was answered better by a direct comparison of two real
workflows on the same runner and image:

  windy-git gate       @postgres:5432   -> passes its migration round-trip
  eternitas migrations @localhost:5432  -> failed

Also scopes test_g73 to workflows that actually run Python. It failed the probe
for not pinning a version when the probe only shelled out to psql — the test
being wrong rather than the workflow.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 14:41:14 -04:00
Grant Whitmer
fe8f84bbdf ci: probe how service containers are addressed on this runner
Some checks failed
check / gate (push) Failing after 17s
canary / probe (push) Successful in 16s
Several migrated workflows hardcode postgres at localhost:5432, which is
correct on GitHub-hosted runners (services are port-mapped to the VM) and
suspect on Gitea Actions (the job runs IN a container, so localhost is the job).
Prove which form works before rewriting anyone's workflow.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 13:14:24 -04:00
Grant Whitmer
eae1bff50b G2.3: Windy Git branding — and get it out of one host's disk into the repo
All checks were successful
check / gate (push) Successful in 19s
canary / probe (push) Successful in 8s
The front end was 100% stock Gitea: green teacup, "Gitea: Git with a cup of
tea", "A painless, self-hosted Git service". G2.3 was specified in the plan with
an acceptance test and never executed, and nothing enforced it.

Now: Windy Git name, wind-mark logo, brand-blue accent, and a landing page that
says what this actually is. Uses Gitea's SUPPORTED surface (custom templates +
public assets) so upstream upgrades keep arriving — no source modified (D-2/I-1).

Two traps this cost, both now documented and tested:

1. GITEA__DEFAULT__APP_NAME does not work. Gitea reads APP_NAME from the TOP
   LEVEL of app.ini; the env var created a literal [default] section that Gitea
   ignores, so the installer's stock APP_NAME kept winning while the config
   looked correct. The env-to-ini pass also APPENDED a second APP_NAME rather
   than replacing the first — a new variant of the documented G4A.3 trap.

2. Cloudflare caches /assets/* for 6h and no token in this stack can purge, so
   the new logo and CSS were invisible while being correct at origin. Brand
   assets now carry a VERSION IN THE FILENAME; bump it on every change.

Committed with an idempotent apply.sh, because applying it straight to Veron's
disk first was itself the config-drift trap this project documents: a rebuild
would have silently reverted to stock Gitea.

85 tests green.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 09:08:48 -04:00
Grant Whitmer
e7dee39151 throttle: stop claiming to limit pushes we cannot see
All checks were successful
check / gate (push) Successful in 19s
canary / probe (push) Successful in 9s
I reintroduced the exact defect I had just criticised. ACTION_BASE listed
"push" and "push.force", but git push goes straight to Gitea over HTTPS and
never touches this API — so nothing records a push, a count would be zero
forever, and enforce() would look up a limit, count nothing, and allow
everything. A silent no-op wearing the costume of a control, made worse by a
config name that implies the protection exists.

Split into ACTION_BASE (actually enforced: repo.create, grant.create) and
NOT_ENFORCED_HERE (push, push.force) with the reason and the remedy written
down: enforcing push velocity needs a Gitea-side pre-receive or push webhook
reporting into agent_actions.

enforce("push") now raises rather than silently allowing, and a test asserts the
two sets stay disjoint.

Found by auditing whether the auth fix could be walked around — every
/api/v1/repos/* route does require a caller, and the only unauthenticated
endpoints are /health, /version and the HMAC-verified webhook.

83 tests green.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:33:30 -04:00
Grant Whitmer
25368547cb docs: runbook — never pull -q, never force-push a deployed branch
All checks were successful
check / gate (push) Successful in 21s
canary / probe (push) Successful in 11s
Two deploy traps paid for on 2026-08-14: 'git pull -q' hid a divergent-branch
error so a deploy ran against stale code while reporting success, and the
divergence came from amending a commit a deploy checkout already tracked.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:28:57 -04:00
Grant Whitmer
c60bfb2b89 I-12: fail the build when COMMIT_SHA is empty
All checks were successful
check / gate (push) Successful in 19s
/version went null after a deploy — the exact "service cannot name its own
commit" defect this project was built to prevent, caught by its own honesty
check.

Cause: the sed replaced "" with "" (a no-op when COMMIT_SHA is empty) and the
grep then matched that same empty string, so the guard verified nothing. A build
with no COMMIT_SHA passed and shipped a container reporting commit_sha: null.

Now the build fails loudly instead.

Second cause of the stale deploy, and it was mine: an earlier `git commit
--amend` + force-push rewrote history the Veron deploy checkout was already
sitting on, leaving it divergent so `git pull -q` failed SILENTLY (-q hid
"Need to specify how to reconcile divergent branches"). Two lessons: do not
force-push a branch a deploy checkout tracks, and do not pull with -q in a
deploy script.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:27:38 -04:00
Grant Whitmer
a0ed4a5ec0 SECURITY: make EPT routing independent of signature well-formedness
All checks were successful
check / gate (push) Successful in 20s
looks_like_ept used jwt.get_unverified_header, which validates the WHOLE token
and therefore rejects anything with a malformed signature segment. Routing
consequently depended on signature well-formedness: an EPT-shaped token with a
bad signature fell through to the HUMAN path, where it was refused for the wrong
reason and — with require_verified_jwt off (dev) — could have been read as a
human identity via its `sub` claim.

Now the header segment is decoded directly, so routing depends only on what the
token CLAIMS to be; whether it is authentic remains verify_ept's job.

Also routes alg:none to the EPT verifier regardless of typ, since a `none`
token is never valid for any caller. Both forged shapes now return 401
ept_invalid — the honest code — instead of 503 "feature not ready".

Found by noticing a forged EPT returned 503 where the verifier should have
answered 401, rather than accepting "it was refused, close enough".

80 tests green.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:23:50 -04:00
Grant Whitmer
830ca48705 docs: mark revocation finding resolved — it was critical, not low
All checks were successful
check / gate (push) Successful in 19s
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:20:46 -04:00
Grant Whitmer
deb3a8fafc test: prove the revocation wiring end-to-end (resolve_passport -> decide_trust)
Some checks failed
check / gate (push) Has been cancelled
Monkeypatched httpx so resolve_passport sees a real revoked trust body
(status=revoked, band=unproven, allowed=[]) and must raise
PassportNotInGoodStanding — proving the WIRING, not just the decision. This is
the path that stops a validly-signed EPT that outlived its passport's
revocation (~365-day tokens).

Live-confirmed alongside: Eternitas refuses to mint EPTs for revoked bots
("credentials are not issued for non-active bots"), so the only exposure was a
pre-existing token — exactly what this now catches. The active agent's EPT still
returns 200. A fully-live revoked test would require revoking a real fleet
passport (destructive), so the wiring is proven deterministically instead.

79 tests green.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:20:44 -04:00
Grant Whitmer
587fb05265 SECURITY: enforce passport status — revocation now takes effect live
All checks were successful
check / gate (push) Successful in 21s
A revoked passport returns HTTP 200, status=revoked, band=unproven,
allowed_actions=[] (verified live 2026-08-13). resolve_passport keyed refusal
only on HTTP 4xx and band=="untrusted", so it returned band 'unproven' and the
agent was seated. Revocation was NOT enforced on the live auth path at all — and
now that agent auth actually works, a revoked agent could authenticate and act.

Extracts decide_trust(body) -> (band, actions) | raise. Only status=="active"
is allowed; revoked/suspended/frozen/unknown all refuse, fail-closed on the
field that carries the most consequential fact about an identity. The agent call
site turns that into a clean 403 passport_revoked.

This is the REAL revocation gate — the token cannot be un-issued, but its
standing is re-checked on every request, so revocation takes effect on the next
call with no webhook required. The Eternitas webhook remains useful for
invalidating locally-issued credentials/grants (G6.3, not built yet), but it was
never the primary gate and its being unwired is no longer a live exposure.

Behavioral tests: revoked body refused, active accepted, unknown/missing status
fails closed.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:18:16 -04:00
Grant Whitmer
c96d1d3102 canary: guard both forgery shapes now that real verification is live
All checks were successful
check / gate (push) Successful in 20s
The EPT-shaped forgery is the one that matters after G3.2 — it is what
signature verification actually guards. The JWT-shaped one still exercises the
human gate. Both must_refuse; a 2xx on either pages.

Verified live: forged EPT -> 401 ept_invalid, forged JWT -> 503,
genuine EPT -> 200.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:13:44 -04:00
Grant Whitmer
79dcea4d3e G3.2/G3.4: real EPT verification + wire the throttle, reopening agent auth
All checks were successful
check / gate (push) Successful in 19s
canary / probe (push) Successful in 10s
REOPENS the agent path — but only because possession is now actually proven.

EPT verification (api/app/ept.py): ES256 against Eternitas's published key set
at /.well-known/eternitas-keys. algorithms=["ES256"] makes alg:none and
algorithm confusion unrepresentable rather than merely unlikely; issuer and exp
are enforced by the library; an unknown kid is refused.

Order is deliberate: signature FIRST, trust lookup second. These EPTs live ~365
days and carry rev/tru baked in at issuance, so a year-old "rev: false" proves
nothing — revocation and band still come from a live lookup on every request.

Found while building it: real EPTs put the passport in the "sub" claim. The old
code read "passport"/"sub_passport", which no genuine EPT carries — so real
agents were never recognised and ONLY forged tokens ever authenticated. The
bypass was not just a hole, it was the only thing that worked.

Throttle (api/app/throttle.py): BAND_MULTIPLIER and rate_*_per_day were defined
and read by nothing. Now enforced on repo.create and grant.create, counted
against agent_actions (one source of truth, not a private counter that drifts
from the audit log). Fails CLOSED — a limiter that fails open protects you until
the moment something is wrong. Untrusted band is 403 read-only, not 429, because
"slow down" would be a lie.

Tests: 14 behavioral, signing real ES256 tokens with a locally-generated key so
they exercise the crypto path with no network dependency — genuine tokens
accepted, and alg:none / foreign key / tampered payload / expired / wrong issuer
/ unknown kid / missing claims all refused. 74 green.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:09:47 -04:00
Grant Whitmer
c257cc55c6 docs: second-auditor review (Fable) — 1 critical fixed live, 4 open
All checks were successful
check / gate (push) Successful in 20s
canary / probe (push) Successful in 11s
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 22:54:15 -04:00
Grant Whitmer
85fa65a52b canary: continuously verify the forged-token bypass stays closed
All checks were successful
check / gate (push) Successful in 18s
Adds must_refuse checks: a 2xx is the alarm, a 401/403/503 is health. The
forged alg:none agent token is probed every 10 min; if it ever returns 2xx the
canary goes red and pages. Proven both ways — 503 reads ok, a 200 endpoint
alarms 'ACCEPTED — this MUST be refused'. The security property is now enforced
by a running check, not assumed.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 22:52:56 -04:00
Grant Whitmer
d8deffe4db SECURITY: close the agent-auth bypass — trust is not authentication
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 11s
Verified live 2026-08-13: a forged 'alg:none' token naming a passport lifted
from the logs returned HTTP 200 as that agent. The agent path read the passport
without verifying the EPT signature, asked Eternitas 'is this passport
reputable?', and seated the caller on a yes. That answers reputation, not
possession — anyone who knows a passport number could impersonate that agent on
the public API.

The human path already failed closed for exactly this reason
(require_verified_jwt). The gate was on the wrong path: it sat AFTER the agent
branch returned. The agent path now fails closed too, BEFORE the trust lookup,
so a forged token never even reaches Eternitas. Reopens automatically when the
ES256/JWKS verifier (G3.2/G9.1) exists and this gate consults it.

Adds BEHAVIORAL tests (not string-grep): a forged alg:none token exercised
through the real get_caller must raise, not authenticate. This is the test that
would have caught the bypass; the suite had 86 source-string assertions and
zero that ran the auth decision.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 22:49:45 -04:00
Grant Whitmer
83047def94 docs: whole account on Windy Git — 143 repos, 1.58 GB, two tiers
131 read-only mirrors (cannot run Actions, zero deploy risk) + 12 writable with
CI and deploys disabled. Total size matches the measured GitHub archive exactly,
which is the confirmation the copy is complete.

Splits the two concerns cleanly: having a copy is safe and should cover
everything now; running code needs judgement and happens per repo.

windy-pro IS included as a mirror — the G11.5 caution is about making it
writable while six checkouts disagree on HEAD, not about holding a read-only
copy. The DR copy is now complete.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:49:48 -04:00
Grant Whitmer
711c47d8c9 fix: the --all-as-mirrors flag itself was never added
The function landed but the argparse anchor did not match the real formatting,
so the command existed and could not be invoked. Caught by running it rather
than assuming the patch applied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:43:22 -04:00
Grant Whitmer
ae6ca8b9d2 G11.3: bulk DR copy — every repo as a read-only mirror
Grant's plan: clone the whole account, let it circulate, reverse direction later
when things are clean. Right plan, with one change that matters.

Sampling 40 repos found 18 carrying deploy/release/publish workflows that
trigger on push: — roughly 63 across the account. Importing those writable with
Actions enabled would arm sixty-odd production deploy triggers on Veron 1, each
needing disarming by hand.

So the bulk goes in as READ-ONLY pull mirrors. A mirror cannot run Actions at
all, so this carries zero deploy risk, and Gitea syncs them itself with no
script and no timer. What you get is a complete, current second copy of the
whole account — the disaster-recovery half — with none of the execution risk.

Converting one to writable + CI stays a deliberate per-repo act: re-import,
review its workflows, disable the deploying ones. That judgement belongs at the
moment you want CI on that repo, not in bulk sixty times by accident.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:40:52 -04:00
Grant Whitmer
89723c6ebd safety: disable deploy workflows on Windy Git before they can fire
Six workflows deploy to production on push:. Windy Git now has a working
runner, so the next synced commit to main would have attempted a production
deploy FROM VERON 1. Their secrets are unset here so they would have failed —
but loudly, on every push, with any pre-SSH step still running.

All six now disabled_manually. Tests, lints and migration checks stay active:
they need no secrets, which is exactly why Phase 1 delivers CI value with
nothing to configure.

Same class of mistake as the push-mirror direction, caught before firing this
time rather than after.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:30:56 -04:00
Grant Whitmer
b2f00821d7 docs: replace the cutover plan with the phased one that matches reality
Phase 1 requires nothing from anyone: agents keep pushing to GitHub, a timer
syncs GitHub -> Windy Git every 15 minutes, CI runs on Veron against current
code. Phase 2 flips one repo at a time, only when that repo is idle.

Records the direction mistake honestly: the source of truth is wherever people
are actually typing, not wherever the plan says it should be.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:14:21 -04:00
Grant Whitmer
51acf9f86e URGENT FIX: reverse the sync — GitHub is the source of truth, not Windy Git
I migrated nine repos writable with push-mirrors pointed AT GitHub. That was
wrong for the actual situation: a dozen agent sessions on the Mac mini are
pushing to GitHub continuously, so GitHub is where the live work is.

A push-mirror force-updates refs. On its 8-hour timer it would have pushed
Windy Git's stale copy over live work — silently, no conflict, nothing to
notice. Removed all nine before the first timer fired; verified no GitHub repo
had been touched (latest push predated the mirrors).

Replaced with the correct Phase 1 direction:

  agents --push--> GitHub --sync--> Windy Git --> CI on Veron

It requires NOTHING from anyone. No remote changes, no coordination, no
'everybody stop pushing'. Agents keep working exactly as they are and CI starts
running on 24 cores.

Windy Git is force-updated on purpose: in Phase 1 it holds nothing anyone
depends on, so GitHub always wins and there is no merge to reconcile.

Phase 2 is per-repo and only when that repo is idle. Never a big-bang cutover
across a dozen live sessions.

Fetches +refs/heads/* and tags explicitly rather than --mirror, which would drag
GitHub's refs/pull/* that Gitea rejects and bury the real errors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:11:34 -04:00
Grant Whitmer
7711adcb97 docs: cutover — Windy Git is the daily driver
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 11s
9 repos migrated writable with push-mirrors to GitHub (sync_on_commit).
Full loop proven end to end: pushed a commit to Windy Git, its existing
workflow ran on Veron 1, and GitHub received the commit within 20s.

Documents the one rule the cutover creates: do NOT push directly to GitHub for
a migrated repo. The mirror makes GitHub match Windy Git, so a direct commit is
overwritten on the next sync, silently, with no conflict. One writer is the
point — two writers with no reconciliation is how you lose work you thought was
saved.

windy-pro deliberately excluded until its six-checkout / forked-counter
ambiguity is resolved by reading (G11.5).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 21:54:47 -04:00
Grant Whitmer
2c292c501c G0.9: nightly timer + the first rehearsed restore in the ecosystem
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 23s
Timer installed with Persistent=true — Grant's workstation is not always on at
04:17, and without that a missed window is silently skipped and the backup
simply never runs.

RESTORE DRILL PASSED, which is the part that matters: pulled a bundle back out
of R2, cloned from it, and the restored HEAD (ffce529) matches live origin/main
exactly — 40 commits, all branches, 50 files.

The August audits found no rehearsed restore anywhere in the ecosystem, for
anything. This is one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 21:49:11 -04:00
Grant Whitmer
ffce529344 G0.9: nightly backup — the prerequisite for becoming the daily driver
All checks were successful
check / gate (push) Successful in 23s
Today GitHub is authoritative, so losing Veron 1 costs nothing. The moment
people push HERE first that inverts: Veron 1 holds the only current copy of the
company's source between mirror syncs, and it is Grant's workstation — no SLA,
no snapshots, residential line, and he reboots it.

git bundle over tar, deliberately: a bundle is one file that git clone reads
directly, so a restore needs no knowledge of Gitea's on-disk layout, and
bundling asks git for a consistent view instead of racing a live push.

  - --all, so every branch and tag is captured. A single-branch bundle loses
    the rest silently and you find out during the restore.
  - git bundle verify before upload. An unverified bundle is a belief.
  - empty repos are reported as skipped, not counted as failures
  - non-zero exit on ANY failure so the unit goes red. A backup script that
    swallows errors manufactures confidence.

Whole archive measured 1.58 GB across 141 repos — about two cents a month.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 21:47:46 -04:00
Grant Whitmer
db31e786e0 G7.4: default to writable repos — mirrors cannot run CI (measured)
All checks were successful
check / gate (push) Successful in 37s
canary / probe (push) Successful in 23s
windy-calendar imported as a pull mirror sat at 0 workflow runs. Gitea does not
fire Actions on mirror sync and a mirror is not a push target, so a mirrored
repo gives you the code and none of the point.

Re-imported it writable, pushed a commit, and its EXISTING .github/workflows/
ci.yml ran on Veron 1 and reported success — with zero workflow edits. That
workflow's own header says it was written for 'OUR self-hosted runners (Kit 0 +
Veron One)' because the account is billing-locked. The fleet already wanted
this; the runners just died when the repos went private.

Default is now writable + push-mirror to GitHub (I-4 steady state). --mirror
stays available for repos to copy but not move.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 19:58:15 -04:00
Grant Whitmer
14fb43b959 G7.4: advertise the fleet's runner labels, and correct my own doctrine
Some checks are pending
check / gate (push) Waiting to run
Surveyed ten repos: 36 of 36 ACTIVE workflows already say
runs-on: [self-hosted, linux, x64]. They were written for the self-hosted
runners that died when the repos went private — so advertising those three
labels makes every one of them runnable AS-IS. No workflow edits, no rewrites.

Correcting an earlier note in this plan: 'ban ubuntu-latest' was wrong as
stated. All 11 occurrences are tagged '# runner-lint-allow — CD/hosted-only;
disabled, manual until CD mission'. They are deliberately hosted-only and
deliberately off. They are correct as written and should not be 'fixed'.

I generalised that doctrine from one repo without surveying. The survey
disagreed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 19:53:44 -04:00
Grant Whitmer
e3677dddd9 docs: measure WHY login takes 18s — it is not 468 call sites
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 24s
The SOTU scopes the child-process DB bridge as 468 call sites and multi-week.
Measurement says otherwise.

  bare node -e '0'  on Kit 0 (4 vCPU, load 20): 1.70-1.99s
  bare node -e '0'  on Veron 1 (24 cores, idle): 0.01s

170x, and node startup is essentially the whole per-query cost — adding pg
connect and a real query to a bare node boot adds only ~0.2-1.2s on top of
~1.8s of interpreter start.

Login makes 9 such forks. 9 x 1.8s = 16s. Observed: 17-25s.

Two independent multipliers compound: the adapter forks per query (9x) and the
box is saturated so each fork costs 1.8s instead of 0.01s (170x).

The consequence: you do not need to touch 468 call sites. querySyncViaChild is
ONE function and its interface does not change — replace the per-query
execFileSync with a persistent worker holding a pg.Pool. All 468 call sites get
faster without being edited, including the mail-lookup that caused today's
outage.

Also measured: 54 containers, 301% CPU of 400%, load 20. dev/demo is 55% of
that — real relief, honestly not a fix. 42 non-dev containers on 4 cores is the
actual condition.

Not implemented here on purpose: sync-over-async with Atomics in the most
critical file in the ecosystem, at the end of a long session, on a box that
already had an outage today, deserves a fresh session and a load test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 19:05:25 -04:00
Grant Whitmer
ef3450d87b G7.6: never let bookkeeping kill the monitor
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 26s
An unwritable state path raised and took the whole canary down. That is the
worst possible trade for a monitoring tool: it reports nothing at all, and
reports it silently. State is an optimisation for transition detection; the
probing is the point.

Found by fat-fingering an env var, which is exactly how it would happen in
production.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 15:42:58 -04:00
Grant Whitmer
a094960da3 G3.5: accept platform.test_ping unverified — unverifiable by construction
All checks were successful
check / gate (push) Successful in 19s
canary / probe (push) Successful in 26s
Read the sender rather than guessing: Eternitas generates the webhook secret at
registration time and pings the URL to prove reachability BEFORE returning that
secret. The ping IS signed — with a secret the receiver cannot possibly hold
yet. Unverifiable by construction, not by oversight.

Accepting it is safe because the event is definitionally a no-op: nothing read,
nothing written, acted:false. Every event that changes anything still requires a
valid HMAC. The alternative, skip_validation:true, would permanently disable
reachability checking for this platform to solve a one-time ordering problem.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 15:36:03 -04:00
Grant Whitmer
5dd914b8a6 G3.5: answer the reachability probe honestly instead of skipping validation
All checks were successful
check / gate (push) Successful in 18s
Eternitas verifies a webhook URL answers BEFORE issuing the secret that signs
deliveries, so the very first request can never carry a signature — refusing it
makes registration impossible. Real chicken-and-egg, not a reason to disable
validation.

A probe is a request claiming no event and carrying no signature. Answering it
200 is honest: the endpoint exists and is ready. It changes nothing (acted:
false), and anything claiming to BE an event still goes through full HMAC
verification. Registering with skip_validation:true would have permanently
disabled a safety check to solve a one-time ordering problem.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 15:33:56 -04:00
Grant Whitmer
339ec70853 G3.5: Eternitas revocation receiver — fail-closed, both signature traps avoided
All checks were successful
check / gate (push) Successful in 18s
When a passport is revoked, every credential it holds here dies in one
transaction: tokens revoked, grants revoked. A revocation that takes effect
'eventually' is not a revocation.

Avoids two traps that each cost a sibling service a subscription that looked
wired and never once delivered:
  1. strip the 'sha256=' prefix before comparing — comparing the decorated
     header against a bare digest returns 401 forever
  2. HMAC the RAW REQUEST BYTES, never a re-serialised body — JSON.stringify of
     a parsed body reorders keys and changes whitespace, so the digest never
     matches what the sender signed

Both fail silently from the sender's side: Eternitas records a delivery, the
receiver records a rejection, nobody notices for weeks.

Unset secret REFUSES rather than accepts — accepting unverified instructions
about identity is worse than missing them. And it never acknowledges a
revocation it could not apply; a 200 there is a security hole reporting success.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 15:32:27 -04:00
Grant Whitmer
fc1937560c G7.6: fix the alert path — urllib UA was rejected 403 by Resend
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 22s
Caught by TESTING the alert path instead of assuming it. Without an explicit
User-Agent, urllib sends 'Python-urllib/3.x' and Resend rejects it 403, while
the identical request via curl succeeds.

The failure mode this avoids is the worst one a canary has: it would have
detected every outage correctly and told nobody. Same bot-filtering trap as the
Gitea migrate call earlier today — worth recognising on sight.

Also prints the HTTP body on failure. '403 Forbidden' alone sends you hunting
for a bad key; the body names the real cause.

Verified: alert sent (200).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:44:00 -04:00
Grant Whitmer
9a7030351b G7.6: fleet canary — probes what a user does, not what is cheap
All checks were successful
check / gate (push) Successful in 19s
Today's outage is the whole design brief: /health returned 200 for the entire
hour that login was dead. A canary watching /health would have stayed green
while nobody in the ecosystem could sign in. So the login probe is here and it
is the one that matters.

Three rules it obeys:
  - never green for something it did not prove (I-8)
  - alert on TRANSITIONS, not every run — a canary people filter is a dead
    canary, which is how the last one sat 37 days dead unnoticed
  - run where the watched thing cannot take it down: Veron 1, never Kit 0

Two independent signals, so losing one still leaves the other: an email via
Resend on state change, and a non-zero exit that turns the CI run red in the
forge itself.

Alerts say what broke in human terms — 'a human can actually sign in' — rather
than only naming an endpoint.

Verified against production: 7/7 green including login at 17.1s; a forced 404
reports DOWN; a 1s threshold reports SLOW at 23.3s.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:41:37 -04:00
Grant Whitmer
211a48187f docs: incident RESOLVED — roster backoff deployed, agent chat back
All checks were successful
check / gate (push) Successful in 18s
The retry fix already existed: a parallel session landed c79f196 (#172) with a
test. Kit 0 was ONE COMMIT BEHIND and did not have it — which is the entire
reason the loop ran.

Deployed by ff-merging /root/windy-chat ac61db6 -> c79f196 and rebuilding only
agent-roster (--no-deps). Never reset --hard; that checkout has a documented
history of local edits a hard reset would eat.

  account-server CPU   168-210%  ->  0.00%
  roster failures      74/min    ->  0/min
  agent chat           stopped   ->  up, healthy
  login                timeout   ->  HTTP 200 ~18s

The lesson: a merged fix that has not reached production is not a fix, it is a
belief. That is exactly what both August audits named — nothing checks whether a
decision reached production — arriving as an outage instead of a report finding.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:19:58 -04:00
Grant Whitmer
3926726309 docs: incident record — account-server login outage 2026-08-12
All checks were successful
check / gate (push) Successful in 18s
Login was timing out for ~50 minutes on the identity service every Windy
product authenticates against.

Two faults multiplying: windy-agent-roster retried a failing mail-lookup at 74
failures/minute with no backoff, and every call cost 0.6-2.1s because the
postgres adapter forks a node process per query and blocks the event loop.
Together they formed a feedback loop — the container's listening socket showed
Recv-Q 510, connections the kernel accepted that node was too blocked to pick
up. Login sat in that queue.

Resolved by stopping windy-agent-roster: login went from 45s+ timeout to HTTP
200 in ~18s. Restored, not healthy — 18s is the fork-per-query adapter on a
54-container 4-vCPU box, and it is what remains after the loop was removed.

Nothing caught this. The container was (unhealthy) with a failing healthcheck
streak of 74 and no alert exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:00:54 -04:00
Grant Whitmer
b6a7045907 G7.4/G11.3: import tool — GitHub upstream, Windy Git downstream, CI here
All checks were successful
check / gate (push) Successful in 17s
Direction is deliberate. For the migration quarter GitHub stays upstream and
Windy Git is a PULL mirror, so it cannot diverge — worst case it is stale, not
wrong. Making Windy Git authoritative before Grant flips G11.6 would create a
two-writer problem nobody asked for. I-4's push-mirror is the steady state for
repos that originate here.

Refuses windy-pro by name: six checkouts, a build counter forked three ways, and
two sessions recording different HEADs hours apart. That gets resolved by
reading, not by importing (G11.5).

Records the Cloudflare trap: the public endpoint answers 403 error 1010 because
CF blocks urllib's user-agent as a bot signature. It reads like Gitea rejecting
the token and is not — the identical call against localhost:3080 on the host
succeeds. Bulk import belongs on the host anyway.

Verified: windy-calendar imported, 731 KB, mirroring.

Survey: 135 private repos, and GitHub Actions cannot run on ANY of them. Sampled
windy-pro, windy-mind, eternitas, windy-registry, windy-agent, WindyCloud — every
recent run is startup_failure, as recently as today.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 12:12:02 -04:00
Grant Whitmer
44756091b0 G7 COMPLETE: CI is green on our own hardware
All checks were successful
check / gate (push) Successful in 17s
Run 7 passed end to end: Python 3.12.13, ruff clean, vocabulary audit clean,
42 tests, migration upgrade->downgrade->upgrade against real Postgres, and
'I-12 holds: env override ignored'.

This is the finding both August audits converged on — 'nothing anywhere checks
whether a decision reached production' — closed for this repo. GitHub Actions
is billing-locked and cannot run on private repos at all, even self-hosted.
This can, on hardware Grant already owns, for zero dollars.

I-5 proven by inspection, not assertion: CI containers sit on
windy-git-runner_jobs, the forge database on windy-git_default. Disjoint. And
neither CI container holds the host docker socket.

Adds strand G7A: seven runs to first green, and each failure presented as a
different problem than it was. The worst was a Python version mismatch that
appeared as a 14-minute hang rather than an error, because pip answered
3.10-vs-3.12 by backtracking through every dependency's release history with
-q hiding it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:38:31 -04:00
Grant Whitmer
358c9bdaa7 G7: clean CI run on a stable runner
All checks were successful
check / gate (push) Successful in 44s
Runs 3 and 6 were killed by me restarting the runner mid-job, not by any
defect in the workflow. Config verified on disk first this time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:35:14 -04:00
Grant Whitmer
6b96087554 G7: per-job networks — fixes service DNS and tightens isolation
Some checks failed
check / gate (push) Failing after 13m33s
The migration step failed with 'could not translate host name postgres'. The
Postgres service container was healthy; the job simply could not name it,
because service DNS aliases only exist on a per-job network and I had pinned
container.network to the flat 'bridge'.

That choice was wrong in both directions: it broke service containers AND it
was weaker isolation, since every concurrent job shared one bridge and could
see its neighbours. A per-job network is stricter and correct — and still has
no route to the forge, because these networks live inside the dind daemon,
which has no forge attachment at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:18:51 -04:00
Grant Whitmer
8c4bdd2bc2 G7.3: setup-python on the act image, not a python image
Some checks failed
check / gate (push) Failing after 43s
Trading one wedge for another: running the job in python:3.12-bookworm fixed
the resolver spiral but broke checkout, because actions/checkout is a
JavaScript action and the official Python images carry no node — 'executable
file not found in $PATH'.

setup-python on the act image has both. Test now accepts either a pinned
container image or an explicit setup-python version, and still checks it
against pyproject's requires-python.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:16:00 -04:00
Grant Whitmer
980fc2dddd G7.3: pin Python 3.12 in CI and cap job runtime
Some checks failed
check / gate (push) Failing after 7s
The first real CI run wedged for 14 minutes. Not a network problem, not the
isolation work — catthehacker/ubuntu:act-22.04 ships Python 3.10.12 while this
project declares requires-python >=3.12, and pip answered that by backtracking
through the entire release history of every dependency looking for something
3.10-compatible. At 100% CPU, with -q hiding every line of it, and it would
have churned until the runner's 30m timeout.

A version mismatch presenting as a hang rather than an error is worth a test,
so there is one: the workflow's python image must satisfy pyproject's
requires-python, checked by parsing both rather than by eyeballing them.

Also: every job now has timeout-minutes. A wedged step should be a red check in
minutes, not an occupied runner for half an hour.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:13:39 -04:00
Grant Whitmer
ed71102e49 G7: trigger a CI run after the routing fix
Some checks failed
check / gate (push) Has been cancelled
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 10:53:57 -04:00
Grant Whitmer
53dea84ab1 G7: route CI jobs via the public surface, not the forge network
Some checks failed
check / gate (push) Failing after 13s
The first CI run failed with 'Could not resolve host: gitea' — job containers
live on dind's private network and cannot see the forge network. Two ways out,
and they are not equivalent:

  (a) put job containers on the forge network. Easy, one line, and it leaves
      untrusted workflow code one DNS name from the forge's Postgres. It quietly
      repeals I-5.
  (b) send jobs to the PUBLIC forge surface over the tunnel, exactly like any
      stranger on the internet.

Took (b). The runner no longer needs the forge network at all, so there is now
NO private route from any CI container to anything — a better posture than this
file started with. Cost is a hairpin through Cloudflare plus its ~100s ceiling
per fetch, which for 0.63 GB of objects across 61 repos and depth=1 checkouts is
nowhere near binding.

Test upgraded to assert the stronger property: no CI container joins the forge
network.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 10:51:00 -04:00
Grant Whitmer
5f717ef74b G7: CI runners with real isolation, and the gate as a workflow
Some checks failed
check / gate (push) Failing after 50s
I-5 says runners execute untrusted code and must be isolated by machine
boundary. act_runner needs a Docker daemon to start job containers, and the
tempting move — what every published example does — is to mount the host's
/var/run/docker.sock. That hands every workflow, including whatever a
transitive dependency's postinstall script feels like doing, the ability to
start a privileged container mounting / — root on Grant's workstation.

Instead the runner talks to its OWN dind daemon:
  - runner (TRUSTED, the act_runner daemon) sits on the forge network only to
    collect jobs from gitea:3000
  - dind and every job container it spawns are UNTRUSTED, on a private network
    with no route to the forge, its Postgres, or its .env
  - jobs cannot bind-mount from the daemon host (valid_volumes: []) and are not
    handed the runner's own socket (docker_host: -)
  - separate compose project, cpu/memory bounded — Veron 1 is Grant's
    workstation, not a dedicated build box

The gate itself now runs as a workflow, including the migration round-trip that
already caught two bugs review did not, and the I-12 check that a COMMIT_SHA
env override cannot change what /version reports.

Labels are explicit and pinned. A workflow naming a label nobody provides
queues forever and presents as a hung CI system rather than a typo — which is
what ubuntu-latest does on every windy-registry run today.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 10:39:05 -04:00
Grant Whitmer
b274392e96 G3.1: OIDC live — sign in to Windy Git with a Windy account
Client 'windy-git' registered on account-server; Gitea auth source 'windy'
added against the discovery document. Auto-registration on, so a Windy account
IS the account — nobody is asked to invent a second identity for the same
person and no local password ever exists.

Proven end to end:
  - app.windygit.com/user/login offers 'Sign in with windy'
  - /user/oauth2/windy -> 307 to account.windyword.ai/oauth/authorize
  - authenticated authorize -> 200, issues a code to the registered callback
  - a bogus evil.example.com redirect_uri -> 'redirect_uri not registered for
    this client'. The anti-phishing check works, which is the whole reason
    redirect URIs are registered rather than accepted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 10:27:14 -04:00
Grant Whitmer
d684cfc2e1 G4.2: record Grant's ruling — use an existing fleet CF token
Not a debt, a decision: sandbox phase, months from launch, and minting a tenth
Cloudflare token to sit in the inventory costs more than it buys. A
platform-specific scoped token is a launch-hardening item.

Recorded so the next reader knows it was chosen rather than missed, with a
do-not-re-raise note. Keeps the genuinely non-obvious part: R2's S3 credentials
are DERIVED from a CF API token — access key id = the token's id, secret =
SHA-256 of the token value.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:34:10 -04:00
Grant Whitmer
a2430de94d I-4: distinguish a mirror that never ran from one that is behind
Gitea reports the epoch for 'not yet synced', which arithmetic turns into a
56-year lag and a confident 'degraded'. Collapsing those two states is how a
backup that was never made gets read as a backup that is merely stale — which
is the more dangerous direction, because 'behind' sounds survivable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:16:05 -04:00
Grant Whitmer
da257652c2 G11 / I-4: continuous off-site mirror, and a namespace bug fixed
I-4 said 'never a one-way door' and had no implementation. Now it does.

  - ensure the GitHub counterpart exists (idempotent), then ask Gitea to keep
    it in step with sync_on_commit=True. An hourly timer means an hour of work
    can be the thing you lose, and that window is invisible until it costs you.
  - mirror status reports what is TRUE including 'we do not know'. An
    unconfigured mirror reports unconfigured, NEVER healthy — same posture as
    me-fleet.ts refusing to say 'online' when it only knows 'registered'.
  - lag past the threshold is a P2, not a shrug. A mirror nobody checks is a
    belief, not a backup, and this ecosystem already lost 37 days to a canary
    everyone assumed was fine.

Gitea owns the replication rather than a hand-rolled loop, because a background
job that fails silently is exactly how the registry's integrity refresh spent
its entire life calling a 404 and incrementing a counter instead of raising.

Also fixes a real bug I had written myself: list_versions derived the Gitea
namespace from the CALLER, which is correct only while the caller is the owner
and addresses the wrong namespace the moment a collaborator asks — surfacing as
'not found', which is the hardest kind of bug to see. Now derived from the repo,
with a test that keeps it that way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:11:45 -04:00
Grant Whitmer
066de34489 G5: fix version-count copy; record the u-system namespace finding
'There are 1 saved versions' is exactly the sloppiness the vocabulary law
exists to catch. Copy is design material, not decoration.

G5.9 records something the live test surfaced: a repo created through
X-Service-Token lands in a 'u-system' namespace because the service caller has
no identity of its own. Correct for /internal plumbing, WRONG for anything a
person owns — the portal must pass an acting user and this cell must refuse to
create a user-owned project without one. Until then service-created repos are
ops artifacts, not customer data.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 15:59:41 -04:00
Grant Whitmer
3a9259a0da G5: generate an unusable password on Gitea user create
Gitea rejects a null password with a bare 400. These accounts are never
password-authenticated — humans arrive via OIDC, agents via passport-bound
scoped tokens, local password sign-in is disabled server-wide — so we generate
a credential that is never stored, returned or recoverable. An unusable
password is safer than a blank one or a shared default.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 15:55:11 -04:00
Grant Whitmer
f621e49770 G0.4: add psycopg2-binary so migrations run inside the image
Alembic runs synchronously, so the container needs a sync driver. Without it
migrations fail on a fresh deploy while passing on any developer machine that
happens to have psycopg2 — exactly the class of gap that only appears in
production.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 15:47:14 -04:00
Grant Whitmer
099e4be9b1 G5: the shelter — repos, grants and version history
The plane Windy Cloud does not have. Verified 2026-08-11: routes/storage.py and
its models contain ZERO occurrences of share/permission/acl/collaborat/seat/
version/snapshot/history/revision. This fills a hole rather than bolting onto
something that already had one.

  - repos: create/list/get, repo_type required (I-7), reserved slugs, Gitea
    reached ONLY through the membrane client (I-1)
  - grants: human identity OR agent passport, exactly one enforced by a database
    CHECK constraint; agent grants expire in 90 days by default
  - versions: history in words a person recognises — no 'commit', no 'branch',
    no 'repository' in any user-facing string (D-9/I-9), with a test that greps
    the speak strings and fails on developer vocabulary
  - private repos 404 rather than 403, so a stranger cannot learn one exists

Auth: three first-class caller classes (human OIDC / agent EPT / internal
service token), NO fourth, and no bypass env var — copied deliberately from the
desktop control server, the ecosystem's best Principle-#5 artifact.

G3.6 status-code law implemented: 400 and 404 REFUSE, 429/5xx retry then REFUSE.
A sibling maps 400/429 to 'unreachable' and soft-ALLOWS, which is inducible —
an attacker who wants the check skipped only has to make it rate-limit itself.
A test asserts resolve_passport has exactly one return path.

And I-8 applied to ourselves: G3.2's JWKS verifier does not exist yet, so the
human token path REFUSES in production rather than accepting an unverified JWT.
An unverified JWT is an authentication bypass, not a shortcut.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 15:45:22 -04:00
Grant Whitmer
82044933ed G2/G4 complete: Gitea live, LFS landing in R2, four traps pinned
VERIFIED END TO END from outside the network:
  - create repo via API -> clone -> commit -> push -> read back over HTTPS
  - 3 MB LFS object pushed through the tunnel, landed in R2 at lfs/34/2d/...
  - NO local lfs/ directory on the host: I-3 confirmed by measurement
  - /health/full: db, gitea and r2 all green

Adds strand G4A recording four traps that each cost a crash loop, with tests:
  1. [lfs] STORAGE_TYPE creates a separate storage section that does not
     inherit [storage] — crash loop, error names the symptom not the cause
  2. storage backend != LFS enabled; LFS_START_SERVER is separate, and its
     absence reads as a permissions error
  3. Gitea env-to-ini SETS but never UNSETS — removing a compose var leaves the
     line in the persisted app.ini, so repo config and prod config silently
     disagree. Exactly the drift this cell exists to end.
  4. R2 rejects the default S3 checksum algorithm

And G4A.5, which is architecture rather than a bug: an 8 MB non-LFS push died
with HTTP 524 at Cloudflare's ~100s limit. This makes the LFS threshold
load-bearing and REQUIRES G10 to serve model weights via presigned R2 URLs
rather than proxying blobs through the tunnel — client straight to R2, which is
what Hugging Face does and which takes Grant's home upstream out of the path.

G2.4: Gitea's MIT text and a NOTICE now travel with the repo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 15:13:58 -04:00
67 changed files with 7129 additions and 46 deletions

View File

@@ -0,0 +1,48 @@
# The fleet canary (G7.6).
#
# Runs on Veron 1, DELIBERATELY not on Kit 0. A canary hosted on the box it
# watches dies with that box and reports nothing at the exact moment it matters.
#
# The previous fleet canary (kit-army-config/docs/deployed-state.json) sat dead
# for 37+ days because its workflow returned startup_failure on every run and
# nothing watched the watcher. This one has two independent signals: an email on
# state change, and a red CI run in the forge. Losing one still leaves the other.
name: canary
on:
schedule:
# Every 10 minutes. Frequent enough that an outage is measured in minutes,
# infrequent enough that the login probe (~18s of real work on a loaded box)
# is not itself a load source.
- cron: "*/10 * * * *"
workflow_dispatch:
jobs:
probe:
runs-on: veron-1
timeout-minutes: 8
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
# State persists between runs so "still broken" can be told apart from
# "just broke" — that is what keeps this from emailing every 10 minutes
# during an outage, and a canary people filter is a dead canary.
- name: restore canary state
uses: actions/cache@v4
with:
path: canary-state.json
key: canary-state-${{ github.run_id }}
restore-keys: canary-state-
- name: probe
env:
RESEND_API_KEY: ${{ secrets.RESEND_API_KEY }}
CANARY_LOGIN_EMAIL: ${{ secrets.CANARY_LOGIN_EMAIL }}
CANARY_LOGIN_PASSWORD: ${{ secrets.CANARY_LOGIN_PASSWORD }}
CANARY_ALERT_TO: ${{ secrets.CANARY_ALERT_TO }}
run: python3 scripts/canary.py

View File

@@ -0,0 +1,93 @@
# The gate, running on our own hardware (G7.3).
#
# This is the dogfood: windy-git verifies itself before anything else migrates.
#
# `runs-on: veron-1` is a label this runner actually provides. NEVER
# `ubuntu-latest` (G7.5) — a self-hosted runner has no such label, so a workflow
# naming it queues forever and presents as a hung CI system rather than a typo.
name: check
on:
push:
branches: [main]
pull_request:
jobs:
gate:
runs-on: veron-1
# A hard ceiling, so a wedged step is a red check in minutes rather than an
# occupied runner for half an hour.
timeout-minutes: 12
services:
postgres:
image: postgres:16-alpine
env:
POSTGRES_USER: windygit
POSTGRES_PASSWORD: windygit
POSTGRES_DB: windygit
options: >-
--health-cmd "pg_isready -U windygit"
--health-interval 5s
--health-retries 10
steps:
- uses: actions/checkout@v4
# The runner image ships Python 3.10.12; this project requires >=3.12.
# The first real CI run wedged for 14 minutes on exactly that: pip
# answered the mismatch by backtracking through the entire release
# history of every dependency looking for something 3.10-compatible, at
# 100% CPU, with `-q` hiding every line of it. A version mismatch
# presenting as a hang rather than an error.
#
# The obvious fix — run the job in a python:3.12 image — trades one wedge
# for another: actions/checkout is a JavaScript action, and the official
# Python images carry no `node`, so checkout dies with "executable file
# not found". Hence setup-python on the act image, which has both.
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: install
run: |
python3 --version
python3 -m venv .venv
.venv/bin/pip install -q --upgrade pip
.venv/bin/pip install -e ".[dev]"
- name: lint
run: .venv/bin/ruff check api scripts
- name: vocabulary audit (D-9)
run: python3 scripts/vocab_audit.py
- name: tests
run: .venv/bin/pytest -q
# G0.4 — a migration nobody has run is a migration nobody can trust. This
# is the step that caught two bugs review did not: SQLAlchemy Enum
# persisting .name instead of .value, and create_table re-emitting
# CREATE TYPE without checkfirst.
- name: migration round-trip (upgrade -> downgrade -> upgrade)
env:
DATABASE_URL: postgresql://windygit:windygit@postgres:5432/windygit
run: |
.venv/bin/alembic upgrade head
.venv/bin/alembic downgrade base
.venv/bin/alembic upgrade head
# I-12 — the honesty check. Nine sibling services cannot name the commit
# they are running; one reports another repo's commit entirely.
- name: /version must equal HEAD
run: |
HEAD_SHA=$(git rev-parse HEAD)
COMMIT_SHA=deadbeefdeadbeefdeadbeefdeadbeefdeadbeef \
.venv/bin/python -c "
import os, sys
sys.path.insert(0, '.')
from api.app.buildinfo import get_build_info
info = get_build_info()
expected = '$HEAD_SHA'
assert info.commit_sha == expected, f'{info.commit_sha} != {expected}'
print('I-12 holds: env override ignored, reported', info.commit_sha[:12])
"

View File

@@ -2,11 +2,17 @@
Read this before touching anything. Then read `DNA_STRAND_MASTER_PLAN.md`, which is the source of truth.
## Current state
## Current state (2026-09-23)
**GENESIS.** No code. No `make dev` yet — building it is codon **G0.8**.
**LIVE on Veron 1** — `app.windygit.com` (Gitea 1.24.6, Windy SSO only),
`api.windygit.com` (our plane: humans via hub JWKS, agents via Eternitas EPT),
`models.windygit.com`. Strands G0–G5, G7, G11 done; see the plan for the rest.
The next work is Strand **G0** (cell substrate), then **G1** (Veron 1 host + Cloudflare Tunnel), then **G2** (Gitea, stock and branded), then **G3** (identity), then **G4** (storage). G0–G4 are sequential. G5–G12 are concurrent once G4 lands.
It is also **the permanent CI for the private platform repos** (GitHub Actions
cannot run on them): `scripts/sync_from_github.sh` + `scripts/pr_status_bridge.py`,
onboarding in `docs/CUTOVER.md`, operations in `docs/RUNBOOK-VERON.md`.
Standing dev checkout: **OC5 `~/windy-git`**. Deploy copy: Veron `/srv/windygit/src`.
## The rules that will get you reverted if you break them
@@ -28,7 +34,7 @@ The next work is Strand **G0** (cell substrate), then **G1** (Veron 1 host + Clo
- Errors are 4-field repair pointers: `{code, speak, machine_cause, remediation_tool}`. No exceptions, including validation errors.
- Every tool response carries `state_proof` + `next_actions`.
- Telemetry `actor_type` comes from the enum `{human, agent, system}`. **`'service'` is not legal** — it 422s and silently drops the whole batch. A sibling service is losing telemetry to exactly this today.
- Runner labels are explicit and pinned. **`ubuntu-latest` is banned** — all four `windy-registry` workflows use it and every run fails.
- Runner labels are explicit and pinned: `[self-hosted, linux, x64]` or `veron-1`. **`ubuntu-latest` is banned** — no runner here has it, so the job queues forever.
## Membrane

View File

@@ -81,7 +81,7 @@ Numbered because code cites them. Changing one requires an ADR that names it.
1. **I-1 · Gitea is a component, never a merged tree.** Our code lives in our services and calls Gitea's REST API. Any patch to Gitea source lives in `patches/` as a numbered, rebasable diff with a one-line justification, and `make check` fails if `patches/` grows past **3** files without an ADR.
2. **I-2 · The membrane is ENUMERATED.**
**Calls out:** `windy-cloud` kernel `GET /api/v1/storage/objects` + `HEAD` (read user objects to version them) · `windy-cloud` `POST /api/v1/storage/quota/check` · `eternitas` `GET /api/v1/trust/{passport}` (band + allowed_actions) · `eternitas` `GET /api/v1/registry/{passport}/integrity` · `account-server` OIDC discovery + JWKS · `windy-cloud-sites` `POST /api/v1/sites/{id}/versions` (publish docs from a repo).
**Calls out:** `windy-cloud` kernel `GET /api/v1/storage/objects` + `HEAD` (read user objects to version them) · `windy-cloud` `POST /api/v1/storage/quota/check` · `eternitas` `GET /api/v1/trust/{passport}` (band + allowed_actions) · `eternitas` `GET /api/v1/registry/{passport}/integrity` · `account-server` OIDC discovery + JWKS · `windy-cloud-sites` `POST /api/v1/sites/{id}/versions` (publish docs from a repo). · windy-admin ledger `POST https://admin.windyword.ai/v1/events` (field telemetry, 2026-09-23: `ci.run`, `ci.job_cancelled`, `service.boot`, `service.health`, `forge.auth.failed` — shapes declared with the ledger owner first; no content, no passports, no emails)
**Calls in:** `POST /internal/repo-from-folder` (Cloud portal: git-enable a folder) · `POST /internal/mirror-status` (ops).
**Events out:** `repo.created`, `repo.pushed`, `release.published`, `model.published`, `ci.completed`.
**Events in:** `passport.revoked` (fail-closed), `storage.quota.exceeded`, `identity.created`.
@@ -258,13 +258,61 @@ Strands G0–G4 are sequential. G5–G12 are concurrent once G4 lands.
## Strand G4 — Storage wiring
- **G4.1** R2 buckets: `windy-git-lfs`, `windy-git-artifacts`, `windy-git-backups` on account `193b347aedeaafe35de0b5a534b2d9aa`.
- **G4.2** **Scoped** R2 credential, minted for this cell. ⚠️ The sites cell shipped holding the account-wide god token (`godtoken02JUL26`) because v4 R2 object endpoints reject restricted tokens — record which we ended up with in `SUBSTRATE.md` and treat an account-wide token as a known, named debt, not an invisible one.
- **G4.2** R2 credential: **use an existing fleet token** (Grant's ruling, 2026-08-11 — sandbox phase, months from launch; a scoped platform token is a launch-hardening item, not a blocker). The non-obvious part worth recording: R2's S3 credentials are *derived* from a Cloudflare API token — **access key id = the token's id, secret = SHA-256 of the token value**.
- **G4.3** Gitea `[storage]` → `STORAGE_TYPE = minio` pointed at R2, covering LFS, attachments, packages, avatars and actions artifacts. ⚠️ **Known R2 trap:** R2 rejects the default checksum algorithm several S3 clients send; if uploads fail with a checksum error, set the MD5 checksum option. *Accept:* a 500 MB file round-trips through LFS and the object is confirmed present in R2 with a matching etag — **verify against the exact pinned Gitea version's docs, do not trust this key name from memory.**
- **G4.4** Git object stores on `/srv/windygit/git` (local NVMe). A test asserts no git object path resolves to a network mount (I-3).
- **G4.5** LFS threshold policy: files **> 5 MB** or matching binary/weight extensions (`.safetensors .bin .gguf .pt .ckpt .onnx .zip .mp4 .wav`) go to LFS via a committed `.gitattributes` template applied at repo creation. **Small text files stay in git proper** — LFS-for-everything makes clones slow and operations heavy.
- **G4.6** Quota: on push, call the kernel's quota check; over-quota → refuse with a repair-pointer error whose `speak` is grandma-words and whose `remediation_tool` names the upgrade path (Storage-Kingdom cross-sell hook, §0.5). **We emit the hook; the kernel owns the price** (I-11).
- **G4.7** Object-count and byte accounting per repo, recomputed nightly, exposed on `GET /{id}/status`.
## Strand G4A — Gitea/R2 traps paid for on 2026-08-11
Four failures hit while wiring G2–G4 live. Each cost a crash loop or a dead end,
and each presents as a different problem than it is. All four are now pinned by
tests in `api/tests/test_invariants.py`.
- **G4A.1 — `[lfs] STORAGE_TYPE` breaks storage inheritance.** Naming a storage
type inside `[lfs]` creates a **separate** storage section that does *not*
inherit the endpoint or credentials from `[storage]`. Gitea crash-loops on
`Endpoint: does not follow ip address or domain name standards` — an error that
names the symptom and not the cause. **Let LFS inherit `[storage]`.** Avatars
initialising correctly is the tell that the rest of the config is fine.
- **G4A.2 — setting the storage backend does not turn LFS on.** `[server]
LFS_START_SERVER = true` is separate. Without it the batch endpoint 404s and
the client reports *"Repository or object not found"*, which reads like a
permissions or credentials problem and is neither.
- **G4A.3 — ⚠️ Gitea's env-to-ini SETS but never UNSETS.** Removing a `GITEA__*`
variable from compose does **not** remove the line from the persisted
`app.ini`. The container will keep booting with a setting that no longer exists
anywhere in the repo, so the config in git and the config in production
silently disagree — the exact class of drift this whole cell exists to end.
**To remove a setting you must edit `app.ini` on the host**
(`/srv/windygit/git/gitea/conf/app.ini`), not just the compose file.
- **G4A.4 — R2 rejects the default S3 checksum algorithm.** Set
`MINIO_CHECKSUM_ALGORITHM = md5` or uploads fail with an opaque checksum error.
### G4A.5 — ⚠️ Cloudflare's 100-second limit caps a push, and it shapes G10
A plain (non-LFS) push of an 8 MB file over the tunnel died with **HTTP 524**.
Cloudflare's proxy times out at ~100 s on the Free plan, and Grant's residential
upstream measured ~19 KB/s during the LFS test, so anything large is a coin flip.
This is not a bug to fix; it is a constraint that **dictates the model-hub
architecture**:
1. **The LFS threshold (G4.5) is load-bearing, not tidiness.** Anything big must
go via LFS, because a large blob inside a git pack has no way to be resumed
or offloaded.
2. **G10 must serve LFS objects via presigned R2 URLs — client straight to R2 —
rather than proxying blobs through the tunnel.** That removes the 100 s
ceiling, removes Grant's home upstream from the path entirely, and is exactly
what Hugging Face does. Until that lands, model repos are capped by whatever
fits in 100 seconds.
**Verified working on 2026-08-11:** a 3 MB LFS object pushed through the tunnel
and landed in R2 at `lfs/34/2d/…`, with **no local `lfs/` directory on the host**
— I-3 confirmed by measurement rather than by assertion.
## Strand G5 — THE SHELTER (permissions plane over Windy Cloud) · **the v0 product**
*This strand is the one that ships first and the one that has no competitor. Windy Cloud today has no sharing, no permissions and no versioning of any kind (D-8).*
@@ -276,6 +324,13 @@ Strands G0–G4 are sequential. G5–G12 are concurrent once G4 lands.
- **G5.5** **Version history UI with one giant Undo.** Grandma-words throughout, and D-9 binds every string: "version", "save point", "restore" — **never "commit," never "a Git"** on this surface.
- **G5.6** `POST /{id}/restore {version_seq}` — lossless, and itself a new version. **Undo is never destructive.**
- **G5.7** Quota events double as storage-plan cross-sell hooks (§0.5).
- **G5.9** ⚠️ **Service callers need an on-behalf-of identity.** Verified live on
2026-08-11: a repo created through `X-Service-Token` lands in a `u-system`
namespace, because the service caller has no identity of its own. That is
correct for `/internal/*` plumbing and **wrong for anything a person owns** —
the Cloud portal must pass the acting user, and this cell must refuse to
create a user-owned repo without one. Until then, service-created repos are
ops artifacts, not customer data.
- **G5.8** *Accept — the strand's whole point in one test:* a Cloud folder is git-enabled, a second human is granted `writer`, that human edits a file through the portal, the owner restores the previous version, and **at no point does either user encounter the word "commit," "repo," "branch," or "Git."**
## Strand G6 — Git protocol surface
@@ -295,11 +350,50 @@ Strands G0–G4 are sequential. G5–G12 are concurrent once G4 lands.
- **G7.2** Runner isolation: containers only, no host Docker socket mount, no host network, no credential in the runner environment beyond the job's own scoped token (I-5).
- **G7.3** Migrate `windy-git`'s own `make check` to Gitea Actions. **Dogfood before anything else moves.**
- **G7.4** Migrate the **private** repos that GitHub cannot run at all — the ones where Actions is dead entirely, even self-hosted. This is the single largest immediate win in the plan.
- **G7.5** ⚠️ **Ban `ubuntu-latest`.** All four `windy-registry` workflows use it and every run fails. Runner labels are explicit and pinned.
- **G7.5** Runner labels are explicit and pinned, and they match **the fleet's existing convention**. Surveyed 2026-08-12 across ten repos: **36 of 36 active workflows already say `runs-on: [self-hosted, linux, x64]`** — written for the self-hosted runners that died when the repos went private. Advertising `self-hosted` + `linux` + `x64` makes all 36 runnable **as-is, with no workflow edits**.
⚠️ Correction to an earlier note in this plan: `ubuntu-latest` is *not* an ecosystem-wide mistake. The 11 occurrences are all tagged `# runner-lint-allow — CD/hosted-only; disabled, manual until CD mission` — deliberately hosted-only and deliberately off. Do not "fix" them; they are correct as written. A job asking for `[self-hosted, linux, x64]` matches only a runner advertising **all three** labels.
- **G7.6** Revive the fleet canary: regenerate `kit-army-config/docs/deployed-state.json` on a schedule from Windy Git CI. **It has been 37+ days dead since 2026-07-03**, and its old workflow returns `startup_failure` on every run.
- **G7.7** Artifacts and logs to R2 (G4.1).
- **G7.8** *Accept:* a deliberately broken commit to a private repo produces a red check on the PR — **which is something that has not happened anywhere in this ecosystem in over a month.**
## Strand G7A — CI traps paid for on 2026-08-12
Seven runs to first green. Every one of these presents as a different problem
than it is, and all are now pinned by tests.
- **G7A.1 — the isolation/DNS collision.** Job containers live inside the dind
daemon and cannot resolve `gitea`, which lives on the forge network. The easy
fix is to put jobs on the forge network — one line, and it leaves untrusted
workflow code one DNS name from the forge's Postgres. **Send jobs to the
PUBLIC forge surface instead**, exactly like any stranger. The runner then
needs no forge attachment at all, so no CI container has a private route to
anything.
- **G7A.2 — the instance URL is baked into `/data/.runner` at registration.**
Changing the environment variable does nothing; the daemon keeps dialling the
old host. To repoint a runner you must drop its data volume and re-register
with a fresh token.
- **G7A.3 — ⚠️ a Python version mismatch presents as a HANG, not an error.**
`catthehacker/ubuntu:act-22.04` ships Python 3.10.12; this project requires
≥3.12. pip answered by backtracking through the entire release history of
every dependency at 100% CPU, with `-q` hiding all of it, for 14 minutes. Use
`actions/setup-python`, and **give every job `timeout-minutes`** so a wedge is
a red check in minutes rather than an occupied runner for half an hour.
- **G7A.4 — you cannot just run the job in a `python:3.12` image.**
`actions/checkout` is a JavaScript action and the official Python images carry
no `node`, so checkout dies with *"executable file not found"*. The act image
has node; `setup-python` supplies the interpreter.
- **G7A.5 — `container.network: bridge` breaks service containers.** Service
DNS aliases exist only on a per-job network, so Postgres came up healthy and
the job could not name it. `network: ""` is both correct **and stricter** — a
shared flat bridge lets concurrent jobs see each other.
- **G7A.6 — do not restart the runner while a job is in flight.** Runs 3 and 6
died to exactly that, and the resulting log looks like a workflow defect
rather than an operator interrupting it.
**Verified green 2026-08-12, run 7:** Python 3.12.13 · ruff clean · vocabulary
audit clean · 42 tests · migration upgrade→downgrade→upgrade against real
Postgres · and `I-12 holds: env override ignored`.
## Strand G8 — Agent surface (MCP + capability discovery)
- **G8.1** MCP server `windy-git-mcp` exposing the full knob-set — every button a human can push. Follows `windy-word-mcp` conventions.

View File

@@ -18,7 +18,14 @@ COPY alembic ./alembic
COPY alembic.ini ./
COPY scripts ./scripts
RUN sed -i "s|^BAKED_COMMIT_SHA: str = \"\"|BAKED_COMMIT_SHA: str = \"${COMMIT_SHA}\"|" api/app/buildinfo.py \
# I-12: an EMPTY COMMIT_SHA must fail the build, not sail through it.
# Previously the sed replaced "" with "" (a no-op) and the grep then matched
# that same empty string, so a build with no COMMIT_SHA passed and shipped a
# container reporting commit_sha: null — exactly the "service cannot name its
# own commit" defect this project exists to prevent. Caught 2026-08-14 when
# /version went null after a deploy.
RUN test -n "${COMMIT_SHA}" || (echo "FATAL: COMMIT_SHA build arg is empty (I-12)" && false) \
&& sed -i "s|^BAKED_COMMIT_SHA: str = \"\"|BAKED_COMMIT_SHA: str = \"${COMMIT_SHA}\"|" api/app/buildinfo.py \
&& sed -i "s|^BAKED_BUILT_AT: str = \"\"|BAKED_BUILT_AT: str = \"${BUILT_AT}\"|" api/app/buildinfo.py \
&& grep -q "BAKED_COMMIT_SHA: str = \"${COMMIT_SHA}\"" api/app/buildinfo.py

24
LICENSES/NOTICE.md Normal file
View File

@@ -0,0 +1,24 @@
# Third-party notices — Windy Git
Windy Git runs **stock Gitea** as an unforked component (D-2 / I-1). Gitea is
distributed under the MIT license, reproduced in `gitea-MIT.txt`.
MIT's only obligation is that the copyright notice and license text travel with
copies of the software that are **distributed**. Running Windy Git as a hosted
service is not distribution, so strictly this file is not required today — it is
here anyway, because it will be required the moment a self-host bundle ships, and
because shipping it costs nothing.
MIT does **not** require us to advertise the lineage, does not restrict
commercial use, and does not require publishing our modifications. That last
point is why Gitea (MIT) was chosen over Forgejo (GPLv3 from v9): GPLv3 does not
trigger on running a service — that is AGPL, and it is widely gotten wrong — but
it **does** trigger on distributing a binary, which the self-host line would do.
**Gitea's trademarks are not used.** The product is branded Windy Git throughout.
| Component | Version | License |
|---|---|---|
| Gitea | 1.24.6 (pinned) | MIT — `gitea-MIT.txt` |
| PostgreSQL | 16 | PostgreSQL License |
| cloudflared | 2026.1.2 | Apache-2.0 |

20
LICENSES/gitea-MIT.txt Normal file
View File

@@ -0,0 +1,20 @@
Copyright (c) 2016 The Gitea Authors
Copyright (c) 2015 The Gogs Authors
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in
all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
THE SOFTWARE.

View File

@@ -22,7 +22,7 @@ boot guard in `api/app/main.py` that refuses to start there in production.
| 8600 | `windy-git-api` — our plane |
| **3080** | Gitea — host 3000 and 3300 are taken by resident projects on Veron 1 |
| 5432 | Postgres |
| 2000 | cloudflared metrics (probe target) |
| 2001 | cloudflared metrics — NOT 2000: `cornercall-tunnel` (another project) takes 2000, and a metrics bind failure kills the whole tunnel |
## Ingress — Cloudflare Tunnel `windy-git`
@@ -72,31 +72,20 @@ consulted, so a perfect service presents as "the app is broken."
All in the fleet lockbox, injected by env, **never committed**. `make check`
fails on any `cfat_` / `cfut_` / `gh[pousr]_` / `et_plt_` literal in the tree.
### ⚠️ NAMED DEBT — the R2 credential is account-wide
### R2 credential — RULED, not a debt (Grant, 2026-08-11)
**As of 2026-08-11 this cell holds the Cloudflare god token as its R2
credential.** R2's S3 credentials are derived from an API token (access key id =
the token's id, secret = SHA-256 of its value), and **no token available to this
session has permission to mint a new one** — creating tokens is dashboard-only
or needs a token-creating token. So the wiring was proven with the god token
rather than blocked on it.
This cell uses an existing fleet Cloudflare token for R2. **That is the decision,
not an oversight.** Grant's ruling, verbatim in intent: we are months from
launch, in a sandbox, and minting a tenth Cloudflare token to sit in the
inventory costs more than it buys. A platform-specific scoped token gets created
as part of launch hardening.
This is recorded, not hidden, because an account-wide token is an acceptable
named debt and an unacceptable invisible one.
Recorded here so the next reader knows it was chosen rather than missed. **Do not
re-raise it before the launch-hardening pass** — see the standing instruction
about pre-launch security-hygiene nagging.
**GATE: this must be replaced with a scoped R2 token BEFORE strand G7 lands
CI runners on this host.** I-5 says runners execute untrusted code and must not
share a kernel with credentials scoped beyond their own job; a god token with
R2 + Workers + Pages + WAF + SSL rights sitting on the same box as a runner is
exactly the thing I-5 exists to prevent.
Minting one is a two-minute job in the Cloudflare dashboard: **R2 → Manage R2
API Tokens → Create → Object Read & Write, scoped to the three `windy-git-*`
buckets.** Then set `R2_ACCESS_KEY_ID` / `R2_SECRET_ACCESS_KEY` in
`/srv/windygit/src/.env` and redeploy.
⚠️ The Cloudflare **god token has Zone:Read but no DNS:Edit.** Use the DNS:Edit
token for record creation.
Access key id = the API token's id; secret = SHA-256 of the token value. That
derivation is not obvious and is the thing worth writing down.
## Backups (G0.9)

297
api/app/auth.py Normal file
View File

@@ -0,0 +1,297 @@
"""Identity: humans, agents, and internal services (G3.2 / G3.6 / I-6).
Three caller classes, all first-class, none a bypass:
* **human** — account-server RS256 JWT (OIDC)
* **agent** — Eternitas ES256 EPT
* **service** — `X-Service-Token`, for the Cloud portal calling `/internal/*`
There is deliberately no fourth class and no escape hatch. The Windy Word desktop
control server is the best Principle-#5 artifact in the ecosystem partly because
it has **no bypass environment variable**, and that is copied here on purpose.
I-6 — EPT parity plus asymmetry: an agent in good standing gets exactly what a
human of the same tier gets. Where a sibling cell silently demotes a tiered agent
to FREE because its EPT carries no tier, we do the opposite, and a test proves it.
"""
from __future__ import annotations
import logging
from dataclasses import dataclass
from enum import StrEnum
import httpx
from fastapi import Header, Request
from api.app.config import Settings
from api.app.ept import EptInvalid, looks_like_ept, verify_ept
from api.app.errors import RepairPointer, passport_unresolvable
from api.app.hub_jwt import HubTokenInvalid, verify_hub_token
from api.app.telemetry import synthetic_headers
log = logging.getLogger(__name__)
class ActorType(StrEnum):
"""G3.7 — these three are the ONLY legal values.
A sibling service emits `actor_type: 'service'` into a telemetry ingest whose
Literal allows only human|agent|system, so every batch 422s and is dropped
with a single console warning. `service` is not spelled `service` here.
"""
human = "human"
agent = "agent"
system = "system"
@dataclass(frozen=True)
class Caller:
actor_type: ActorType
identity_id: str | None = None
passport: str | None = None
band: str | None = None
allowed_actions: tuple[str, ...] = ()
@property
def subject(self) -> str:
return self.identity_id or self.passport or "system"
# EI_CAPABILITY_MATRIX.v1 — velocity multipliers by integrity band.
BAND_MULTIPLIER: dict[str, float] = {
"platinum": 10.0,
"gold": 4.0,
"standard": 1.0,
"proven": 1.0,
"watch": 0.5,
"untrusted": 0.0, # read-only
# Eternitas began emitting this band on 2026-07-30 and it is not in the
# documented enum yet. Treating an unknown band as untrusted would lock out
# every freshly hatched agent; treating it as trusted would be a hole.
# Standard-with-no-bonus is the honest middle.
"unproven": 1.0,
}
class PassportNotInGoodStanding(Exception):
"""The passport resolved, but Eternitas does not list it as active
(revoked / suspended / frozen / unknown status)."""
def __init__(self, passport: str, status: str) -> None:
self.passport = passport
self.status = status
super().__init__(f"{passport} status={status!r}")
def decide_trust(body: dict, passport: str) -> tuple[str, tuple[str, ...]]:
"""The trust body -> (band, allowed_actions), or refuse.
THE GATE THAT WAS MISSING. A revoked passport returns HTTP 200 with
`status: revoked`, `band: unproven`, `allowed_actions: []` — verified live
2026-08-13. The previous code keyed refusal only on HTTP 4xx and on
band=="untrusted", so a revoked agent (200, band unproven) authenticated and
acted normally. Revocation was not enforced on the live path at all; the
webhook that was supposed to be the backup was never the primary gate.
Only `status == "active"` is allowed. Anything else — including a status
Eternitas invents tomorrow — refuses. Fail-closed on the field that carries
the most consequential fact about an identity.
"""
status = str(body.get("status", "")).lower()
if status != "active":
raise PassportNotInGoodStanding(passport, status or "missing")
return body.get("band", "unproven"), tuple(body.get("allowed_actions", ()))
async def resolve_passport(settings: Settings, passport: str) -> tuple[str, tuple[str, ...]]:
"""G3.6 — THE STATUS-CODE LAW.
404 and 400 REFUSE. 429 and 5xx retry with backoff, then REFUSE.
A sibling service maps 400 and 429 to "unreachable" and then soft-ALLOWS.
That is a live residual bypass, because 429 is trivially inducible at
100/min/IP: an attacker who wants the check skipped only has to make the
check rate-limit itself. There is no code path here where an unresolvable
passport is permitted to write.
"""
if not settings.eternitas_configured:
raise RepairPointer(
status_code=503,
code="trust_unavailable",
speak="We can't confirm helper IDs right now, so we didn't let that change through.",
machine_cause="eternitas is not configured; policy is fail-closed",
remediation_tool=None,
)
url = f"{settings.eternitas_base_url}/api/v1/trust/{passport}"
headers = {"X-API-Key": settings.eternitas_platform_api_key, **synthetic_headers()}
last_status = 0
for attempt in range(3):
async with httpx.AsyncClient(timeout=httpx.Timeout(8.0, connect=3.0)) as client:
try:
r = await client.get(url, headers=headers)
except httpx.RequestError as exc:
last_status = 599
log.warning("eternitas unreachable (attempt %s): %s", attempt + 1, exc)
continue
last_status = r.status_code
if r.status_code == 200:
# decide_trust raises PassportNotInGoodStanding on a non-active
# status; that propagates past the retry loop as a hard refusal.
return decide_trust(r.json(), passport)
if r.status_code in (400, 404):
# Malformed or not-issued. Refuse immediately — retrying cannot help
# and pretending it might is how a soft-allow gets written.
break
# 429 / 5xx: retry, then refuse. Never allow.
raise passport_unresolvable(passport, last_status)
async def get_caller(
request: Request,
authorization: str | None = Header(default=None),
x_service_token: str | None = Header(default=None),
) -> Caller:
settings: Settings = request.app.state.settings
# --- internal service caller (the Cloud portal) ------------------------
if x_service_token:
expected = settings.service_token
if not expected:
raise RepairPointer(
status_code=503,
code="service_auth_unconfigured",
speak="That connection isn't set up yet.",
machine_cause="SERVICE_TOKEN is unset; refusing to accept service calls",
remediation_tool=None,
)
# Constant-time compare, copied from the desktop control server's
# control-auth pattern rather than reinvented.
import hmac
if not hmac.compare_digest(x_service_token, expected):
raise RepairPointer(
status_code=401,
code="service_token_invalid",
speak="That connection isn't authorised.",
machine_cause="X-Service-Token did not match",
remediation_tool=None,
)
return Caller(actor_type=ActorType.system, identity_id="system")
if not authorization or not authorization.lower().startswith("bearer "):
raise RepairPointer(
status_code=401,
code="not_signed_in",
speak="You'll need to sign in first.",
machine_cause="no bearer token and no service token presented",
remediation_tool=None,
)
token = authorization.split(" ", 1)[1].strip()
# --- agent (Eternitas EPT) --------------------------------------------
# Possession FIRST, reputation second. The signature proves the caller holds
# this passport; the trust lookup then says what it may do. Doing only the
# second was the 2026-08-13 impersonation bypass.
if looks_like_ept(token):
try:
verified = verify_ept(token, settings.eternitas_base_url)
except EptInvalid as exc:
raise RepairPointer(
status_code=401,
code="ept_invalid",
speak="We couldn't confirm that helper's ID, so we didn't let it in.",
machine_cause=f"EPT verification failed: {exc}",
remediation_tool="windy_git.reissue_agent_token",
) from exc
# The token is authentic. It is NOT evidence of current standing: these
# EPTs live ~365 days and carry `rev`/`tru` baked in at issuance, so a
# year-old `rev: false` proves nothing. Revocation and band come from a
# live lookup, every time.
try:
band, actions = await resolve_passport(settings, verified.passport)
except PassportNotInGoodStanding as exc:
# The signature is authentic, but the identity is no longer good.
# Revocation takes effect here, live, on the next request — no
# webhook required. That is the honest place for it: the token can't
# be un-issued, but its standing is checked every time.
raise RepairPointer(
status_code=403,
code="passport_revoked",
speak="That helper's access has been turned off.",
machine_cause=f"eternitas status for {verified.passport} is {exc.status!r}, not active",
remediation_tool=None,
) from exc
if band.lower() == "untrusted":
raise RepairPointer(
status_code=403,
code="agent_read_only",
speak="That helper can look, but it isn't allowed to make changes yet.",
machine_cause=f"passport {verified.passport} band=untrusted is read-only",
remediation_tool=None,
)
return Caller(
actor_type=ActorType.agent,
passport=verified.passport,
band=band,
allowed_actions=actions,
)
# --- human (hub RS256 access token, G3.2) ------------------------------
# Production ALWAYS verifies, whatever require_verified_jwt says: the flag
# only exists to let local dev run against unsigned fixture tokens.
if settings.require_verified_jwt or settings.is_production:
try:
human = verify_hub_token(
token,
settings.account_server_base_url,
issuers=tuple(settings.hub_issuers),
audiences=tuple(settings.hub_audiences),
require_aud=settings.hub_require_aud,
)
except HubTokenInvalid as exc:
raise RepairPointer(
status_code=401,
code="token_invalid",
speak="We couldn't confirm that sign-in. Try signing in again.",
machine_cause=f"hub token verification failed: {exc}",
remediation_tool=None,
) from exc
return Caller(actor_type=ActorType.human, identity_id=human.identity_id)
# Local dev only (require_verified_jwt=False outside production).
identity_id = _unverified_claim(token, "windy_identity_id") or _unverified_claim(token, "sub")
if not identity_id:
raise RepairPointer(
status_code=401,
code="token_unrecognised",
speak="We couldn't read that sign-in. Try signing in again.",
machine_cause="token carried neither a passport nor an identity claim",
remediation_tool=None,
)
return Caller(actor_type=ActorType.human, identity_id=identity_id)
def _unverified_claim(token: str, claim: str) -> str | None:
"""Read a claim WITHOUT verifying the signature.
Used only to decide which verifier a token belongs to. Every path that acts
on the result re-establishes trust independently: an agent's authority comes
from a live Eternitas trust lookup, never from the token's own assertions.
Humans are verified by hub_jwt.verify_hub_token (G3.2); this reader backs
only the local-dev path, which production never takes.
"""
import base64
import json
try:
payload = token.split(".")[1]
payload += "=" * (-len(payload) % 4)
return json.loads(base64.urlsafe_b64decode(payload)).get(claim)
except Exception: # noqa: BLE001
return None

View File

@@ -46,9 +46,41 @@ class Settings(BaseSettings):
# ---- Eternitas (agent identity + trust) -------------------------------
eternitas_base_url: str = "https://api.eternitas.ai"
eternitas_platform_api_key: str = ""
# Signs webhooks Eternitas delivers to us. Unset = we refuse them (I-8):
# accepting unverified instructions about identity is worse than missing them.
eternitas_webhook_secret: str = ""
eternitas_platform_id: str = ""
# ---- account-server OIDC (human identity) -----------------------------
account_server_base_url: str = "https://account.windyword.ai"
# G3.2 — what a hub ACCESS token must say about itself (see hub_jwt.py).
# Token contract v1 (lane 8c, 2026-09-23): access tokens may carry either
# issuer. id_tokens are kept out by `type` + `windy_identity_id` + aud, not
# by issuer.
hub_issuers: list[str] = ["windy-identity", "https://account.windyword.ai"]
# Contract v1: aud is an ARRAY; first-party tokens list every product, and
# Windy Git's entry is `windy_git` (underscore). ⚠️ NEVER add "windy-git"
# (hyphen): that is Gitea's OIDC client_id, so an id_token minted for the
# forge would carry it and pass as a bearer here.
hub_audiences: list[str] = ["windy_git"]
# Flip to True once the hub emits aud on every access token.
hub_require_aud: bool = False
# ---- field telemetry (admin.windyword.ai ledger) ----------------------
# Unset token = nothing sent, nothing buffered. The token lives in the
# root-only /etc/windygit/telemetry.env on Veron, never in the repo.
windygit_telemetry_token: str = ""
telemetry_ingest_url: str = "https://admin.windyword.ai/v1/events"
# Internal callers (the Cloud portal calling /internal/*). A first-class
# caller class, not a bypass: unset means service calls are REFUSED.
service_token: str = ""
# ⚠️ FAIL-CLOSED GATE. Human tokens are verified against the hub's JWKS
# (G3.2, hub_jwt.py). False only enables the unverified local-dev path, and
# production verifies regardless — an unverified JWT is a bypass, not a
# shortcut.
require_verified_jwt: bool = True
# ---- storage law (I-3, G4.4) ------------------------------------------
# Git object databases MUST live on a POSIX filesystem. A test asserts this
@@ -71,7 +103,12 @@ class Settings(BaseSettings):
rate_grants_per_day: int = 100
rate_force_pushes_per_day: int = 10
# ---- mirror health (I-4) ----------------------------------------------
# ---- mirror: I-4, never a one-way door --------------------------------
github_token: str = ""
github_owner: str = "sneakyfree"
# Gitea's timer, as a backstop. sync_on_commit is what actually matters:
# an hourly window means an hour of work can be the thing you lose.
mirror_interval: str = "8h0m0s"
mirror_lag_p2_seconds: int = 3600 # env: 60 min -> P2
# ---- agent grants (G5.3) ----------------------------------------------

157
api/app/ept.py Normal file
View File

@@ -0,0 +1,157 @@
"""EPT signature verification (G3.2 / G9.1) — the gate that makes an agent an agent.
Trust is not authentication. A trust lookup answers *"is this passport
reputable?"*; only a signature answers *"does this caller actually hold it?"*.
Skipping the second question was a live impersonation bypass on 2026-08-13 — a
forged `alg:none` token naming a passport read out of the logs returned HTTP 200.
What this module refuses, deliberately and by construction:
* **`alg: none`** — the original exploit. `algorithms=["ES256"]` makes it
unrepresentable rather than merely unlikely.
* **Algorithm confusion.** Only ES256 is accepted. If an attacker presents an
HS256 token, PyJWT will not try to use an EC public key as an HMAC secret,
which is the classic way "verified" JWTs get forged.
* **An unknown `kid`.** The key must be one Eternitas currently publishes.
* **A wrong issuer or an expired token** — checked by the library, not by us.
What it deliberately does NOT decide: whether the agent is *allowed* to act.
The EPT carries `rev` and `tru` claims baked in at issuance, and these tokens
live for a year (observed `exp` ≈ 365 days). A year-old `rev: false` is not
evidence of anything. **Revocation and trust must come from a live lookup**, so
this module returns only identity and the caller re-checks standing.
"""
from __future__ import annotations
import base64
import json
import logging
import time
from dataclasses import dataclass
import httpx
import jwt
from jwt import PyJWKClient
log = logging.getLogger(__name__)
ISSUER = "eternitas.ai"
ALGORITHMS = ["ES256"] # exactly one. Never widen this list.
_jwks_client: PyJWKClient | None = None
_jwks_url: str | None = None
class EptInvalid(Exception):
"""The token is not a valid, currently-signed Eternitas EPT."""
@dataclass(frozen=True)
class VerifiedEpt:
passport: str
operator: str | None
bot_name: str | None
issued_at: int | None
expires_at: int | None
def _client(base_url: str) -> PyJWKClient:
"""One cached JWKS client. PyJWKClient caches keys and refetches on an
unknown kid, so a key rotation heals itself without a redeploy."""
global _jwks_client, _jwks_url
url = f"{base_url.rstrip('/')}/.well-known/eternitas-keys"
if _jwks_client is None or _jwks_url != url:
_jwks_client = PyJWKClient(url, cache_keys=True, lifespan=300)
_jwks_url = url
return _jwks_client
def verify_ept(token: str, eternitas_base_url: str) -> VerifiedEpt:
"""Verify an EPT's signature and claims. Raises EptInvalid on ANY doubt.
There is no partial success and no "probably fine" path: every failure mode
below produces the same refusal, because a caller that cannot prove
possession is indistinguishable from an attacker.
"""
try:
signing_key = _client(eternitas_base_url).get_signing_key_from_jwt(token)
except Exception as exc: # noqa: BLE001 - unknown kid, unreachable JWKS, malformed
raise EptInvalid(f"no usable signing key: {type(exc).__name__}: {exc}") from exc
try:
claims = jwt.decode(
token,
signing_key.key,
algorithms=ALGORITHMS, # ES256 only — closes alg:none and alg confusion
issuer=ISSUER,
options={
"require": ["sub", "iss", "exp"],
"verify_signature": True,
"verify_exp": True,
"verify_iss": True,
},
)
except jwt.PyJWTError as exc:
raise EptInvalid(f"{type(exc).__name__}: {exc}") from exc
# The REAL claim name. Eternitas puts the passport in `sub`; the previous
# code looked for `passport` / `sub_passport`, which no genuine EPT carries —
# so real agents were never recognised and only forged tokens ever "worked".
passport = claims.get("sub")
if not isinstance(passport, str) or not passport.strip():
raise EptInvalid("EPT carried no passport in `sub`")
return VerifiedEpt(
passport=passport,
operator=claims.get("ope"),
bot_name=claims.get("bot"),
issued_at=claims.get("iat"),
expires_at=claims.get("exp"),
)
def looks_like_ept(token: str) -> bool:
"""Cheap, unauthenticated triage: is this token even *claiming* to be an EPT?
Used ONLY to route a token to the right verifier. It decides nothing about
trust — an attacker controls every byte it reads.
Decodes the header segment directly rather than via
`jwt.get_unverified_header`, which validates the whole token structure and
therefore rejects anything with a malformed SIGNATURE. That made routing
depend on signature well-formedness: an EPT-shaped token with a bad
signature fell through to the human path, where it was refused for the wrong
reason ("signing in isn't switched on") and — with `require_verified_jwt`
off — could have been read as a human identity via its `sub` claim.
Routing must depend only on what the token claims to be. Whether it is
authentic is `verify_ept`'s job, and it says no.
"""
try:
head_b64 = token.split(".", 1)[0]
head_b64 += "=" * (-len(head_b64) % 4)
header = json.loads(base64.urlsafe_b64decode(head_b64))
except Exception: # noqa: BLE001
return False
if not isinstance(header, dict):
return False
return header.get("typ") == "EPT" or header.get("alg") in ("ES256", "none")
async def eternitas_reachable(base_url: str) -> bool:
"""Whether the key set can be fetched at all. Used by /health/full so an
unreachable JWKS is reported rather than discovered during an outage."""
try:
async with httpx.AsyncClient(timeout=httpx.Timeout(5.0, connect=3.0)) as c:
r = await c.get(
f"{base_url.rstrip('/')}/.well-known/eternitas-keys",
headers={"User-Agent": "windy-git/1.0"},
)
return r.status_code == 200 and "keys" in r.json()
except Exception: # noqa: BLE001
return False
def seconds_until_expiry(ept: VerifiedEpt) -> int | None:
return None if ept.expires_at is None else int(ept.expires_at - time.time())

133
api/app/hub_jwt.py Normal file
View File

@@ -0,0 +1,133 @@
"""Human token verification (G3.2) — hub access tokens from account.windyword.ai.
Until this existed the human path refused every token in production (503
`human_signin_not_ready`), because reading an unverified JWT's claims is an
authentication bypass, not a shortcut. This module is what lets it say yes.
The token it accepts is the hub's ACCESS token, as observed live 2026-09-23:
header {alg: RS256, typ: JWT, kid: <published at /.well-known/jwks.json>}
claims iss = "windy-identity" (contract v1 also allows the discovery URL)
type = "human", exp - iat = 900 s
sub = per-row user id ← NOT the cross-product identity
windy_identity_id = the Windy Account UUID (what Gitea's OIDC links on)
no `aud` yet
What it refuses, by construction:
* **Anything but RS256.** One algorithm, never a list. Closes `alg: none` and
HS256-with-the-public-key confusion.
* **An unknown `kid`**, a wrong issuer, an expired token — library-checked.
* **An id_token used as a bearer.** id_tokens prove a login happened to a
relying party (for the forge: aud `windy-git`), not that this caller may act
here. They carry no `type` and no `windy_identity_id`, and their aud is a
client id, not the product name `windy_git` — any one of the three refuses.
* **A non-human `type`.** An agent's authority comes from its EPT and a live
Eternitas lookup, never from a hub token dressed as a person.
* **A token with no `windy_identity_id`.** `sub` is a different namespace (the
per-row user id); falling back to it would silently mint identities that
match nothing Gitea knows.
`aud` (token contract v1, lane 8c): an array; first-party tokens list every
product and Windy Git's is `windy_git`. Optional until the hub emits it; when
present it MUST include `windy_git`.
`hub_require_aud=True` makes it mandatory — flip it once the hub emits it.
"""
from __future__ import annotations
from dataclasses import dataclass
import jwt
from jwt import PyJWKClient
ALGORITHMS = ["RS256"] # exactly one. Never widen this list.
_jwks_client: PyJWKClient | None = None
_jwks_url: str | None = None
class HubTokenInvalid(Exception):
"""Not a valid, currently-signed hub access token for a human."""
@dataclass(frozen=True)
class VerifiedHuman:
identity_id: str
email: str | None
expires_at: int | None
def _client(base_url: str) -> PyJWKClient:
"""Cached JWKS client; refetches on an unknown kid so rotation self-heals."""
global _jwks_client, _jwks_url
url = f"{base_url.rstrip('/')}/.well-known/jwks.json"
if _jwks_client is None or _jwks_url != url:
_jwks_client = PyJWKClient(url, cache_keys=True, lifespan=300)
_jwks_url = url
return _jwks_client
def verify_hub_token(
token: str,
base_url: str,
*,
issuers: tuple[str, ...],
audiences: tuple[str, ...],
require_aud: bool,
signing_key=None,
) -> VerifiedHuman:
"""Verify a hub access token. Raises HubTokenInvalid on ANY doubt.
`signing_key` exists for tests only (a locally generated key, no network).
"""
try:
key = (
signing_key
if signing_key is not None
else _client(base_url).get_signing_key_from_jwt(token).key
)
except Exception as exc: # noqa: BLE001 - unknown kid, unreachable JWKS, malformed
raise HubTokenInvalid(f"no usable signing key: {type(exc).__name__}: {exc}") from exc
try:
claims = jwt.decode(
token,
key,
algorithms=ALGORITHMS,
issuer=list(issuers),
options={
"require": ["iss", "exp", "iat"],
"verify_signature": True,
"verify_exp": True,
"verify_iss": True,
# Checked by hand below: PyJWT rejects any token CARRYING aud
# when no audience is passed, which would break the moment the
# hub starts emitting it — the exact trap the SSO matrix names.
"verify_aud": False,
},
)
except jwt.PyJWTError as exc:
raise HubTokenInvalid(f"{type(exc).__name__}: {exc}") from exc
aud = claims.get("aud")
if aud is None:
if require_aud:
raise HubTokenInvalid("token carries no aud and hub_require_aud is on")
else:
presented = {aud} if isinstance(aud, str) else set(aud) if isinstance(aud, list) else set()
if not presented & set(audiences):
raise HubTokenInvalid(f"aud {sorted(presented)} does not name Windy Git")
# REQUIRED, not defaulted: id_tokens carry no `type`, and this is one of the
# two claims (with windy_identity_id) that keep them from acting as bearers.
if claims.get("type") != "human":
raise HubTokenInvalid(f"token type {claims.get('type')!r} is not a human access token")
identity = claims.get("windy_identity_id") or claims.get("windyIdentityId")
if not isinstance(identity, str) or not identity.strip():
raise HubTokenInvalid("token carries no windy_identity_id")
return VerifiedHuman(
identity_id=identity, email=claims.get("email"), expires_at=claims.get("exp")
)

View File

@@ -6,14 +6,16 @@ component and is reached only over its REST API (D-2 / I-1).
from __future__ import annotations
import asyncio
import logging
import socket
import time
from contextlib import asynccontextmanager
from fastapi import FastAPI
from fastapi import FastAPI, Request
from fastapi.exceptions import RequestValidationError
from fastapi.responses import JSONResponse
from sqlalchemy.ext.asyncio import create_async_engine
from sqlalchemy.ext.asyncio import async_sessionmaker, create_async_engine
from api.app.buildinfo import get_build_info
from api.app.config import get_settings
@@ -24,7 +26,8 @@ from api.app.providers.registry import (
GiteaProvider,
R2Provider,
)
from api.app.routes import health
from api.app.routes import health, repos, webhooks
from api.app.telemetry import SYNTHETIC, Telemetry, caller_class, is_synthetic
logging.basicConfig(
level=logging.INFO,
@@ -44,9 +47,7 @@ def _refuse_kit_zero(settings) -> None:
if not settings.is_production:
return
try:
local_ips = {
info[4][0] for info in socket.getaddrinfo(socket.gethostname(), None)
}
local_ips = {info[4][0] for info in socket.getaddrinfo(socket.gethostname(), None)}
except socket.gaierror:
return
if settings.kit0_host in local_ips:
@@ -80,6 +81,9 @@ async def lifespan(app: FastAPI):
app.state.settings = settings
app.state.engine = engine
app.state.sessionmaker = (
async_sessionmaker(engine, expire_on_commit=False) if engine is not None else None
)
app.state.providers = [
DatabaseProvider(engine),
GiteaProvider(settings),
@@ -95,8 +99,23 @@ async def lifespan(app: FastAPI):
# systemd Restart=always, plus the runbook's `systemctl status`.
]
telemetry = Telemetry(
settings.telemetry_ingest_url,
settings.windygit_telemetry_token,
environment=settings.environment,
commit_sha=info.commit_sha,
version=info.version,
)
app.state.telemetry = telemetry
telemetry.boot()
await telemetry.flush()
task = asyncio.create_task(telemetry.run()) if telemetry.enabled else None
yield
if task is not None:
task.cancel()
await telemetry.flush()
if engine is not None:
await engine.dispose()
@@ -112,10 +131,48 @@ app = FastAPI(
)
app.include_router(health.router)
app.include_router(repos.router)
app.include_router(webhooks.router)
@app.middleware("http")
async def _count_requests(request: Request, call_next):
"""Heartbeat counts (requests, 4xx/5xx, refusals, p95). Never raises."""
start = time.perf_counter()
marker = SYNTHETIC.set(is_synthetic(request.headers))
try:
response = await call_next(request)
finally:
SYNTHETIC.reset(marker)
tel = getattr(request.app.state, "telemetry", None)
if tel is not None:
tel.record_request(
response.status_code,
(time.perf_counter() - start) * 1000,
refused=getattr(request.state, "refused", False),
)
return response
@app.exception_handler(RepairPointer)
async def _repair_pointer_handler(_, exc: RepairPointer) -> JSONResponse:
async def _repair_pointer_handler(request: Request, exc: RepairPointer) -> JSONResponse:
tel = getattr(request.app.state, "telemetry", None)
detail = exc.detail if isinstance(exc.detail, dict) else {}
code = detail.get("code")
if tel is not None and code in tel.auth_codes:
# A refusal is a failure row (field-visibility rule 1). The caller is
# unauthenticated by definition, so: system actor, no actor_id, and
# the route TEMPLATE, never the concrete path.
request.state.refused = True
route = request.scope.get("route")
tel.auth_failed(
code=code,
http_status=exc.status_code,
caller=caller_class(request.headers),
route=getattr(route, "path", None),
upstream_status=getattr(exc, "upstream_status", None),
synthetic=is_synthetic(request.headers),
)
return JSONResponse(status_code=exc.status_code, content=exc.detail)

View File

@@ -105,7 +105,7 @@ class DatabaseProvider(Provider):
# TunnelProvider was removed deliberately. See the note in main.py: cloudflared
# binds 127.0.0.1:2000 on the HOST, and this process runs in a container whose
# binds 127.0.0.1:2001 on the HOST, and this process runs in a container whose
# only route to the host is the bridge gateway (172.17.0.1), where nothing is
# listening. Binding the metrics endpoint wider would fix the probe and make a
# metrics bind failure able to take down ingress -- a worse trade than losing

532
api/app/routes/repos.py Normal file
View File

@@ -0,0 +1,532 @@
"""The shelter — repos, grants and version history (strand G5, D-8).
Windy Cloud today has **no sharing, no permissions and no versioning of any
kind**: verified 2026-08-11 against `routes/storage.py` and its models, which
contain zero occurrences of share / permission / acl / collaborat / seat /
version / snapshot / history / revision. This plane is not a feature bolted onto
something that already had one — it fills a hole that has never been filled.
D-8 also fixes the order: **permissions and history ship before the git
protocol.** "I want someone to help me with my website" is a real problem for a
real person, and it does not require them to know what a repository is.
Every string a person sees here obeys the D-9 vocabulary law: *version* and
*save point*, never *commit*, and never the countable form of the word "Git" on
any surface, ever. See `scripts/vocab_audit.py`, which enforces this.
"""
from __future__ import annotations
import uuid
from datetime import UTC, datetime, timedelta
from typing import Annotated
from fastapi import APIRouter, Depends, Request
from pydantic import BaseModel, Field, field_validator
from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker
from api.app import throttle
from api.app.auth import ActorType, Caller, get_caller
from api.app.errors import RepairPointer
from api.app.models.core import (
CreatedVia,
GrantRole,
Mirror,
MirrorState,
Repo,
RepoGrant,
RepoState,
RepoType,
RepoVersion,
Visibility,
)
from api.app.services.gitea_client import GiteaClient
from api.app.services.mirror import MirrorService
router = APIRouter(prefix="/api/v1/repos", tags=["repos"])
# Gitea's permission vocabulary, mapped from ours. Ours is the one users see.
_ROLE_TO_GITEA = {
GrantRole.owner: "admin",
GrantRole.maintainer: "admin",
GrantRole.writer: "write",
GrantRole.reader: "read",
}
_RESERVED_SLUGS = {
"api", "admin", "login", "logout", "signup", "settings", "explore",
"new", "user", "org", "repo", "assets", "static", "help", "about",
}
# --------------------------------------------------------------------------
# schemas
# --------------------------------------------------------------------------
class CreateRepo(BaseModel):
name: str = Field(min_length=1, max_length=100)
display_name: str | None = None
description: str = ""
# I-7 — required, never defaulted at read time, never inferred.
repo_type: RepoType
visibility: Visibility = Visibility.private
@field_validator("name")
@classmethod
def _slug(cls, v: str) -> str:
slug = "".join(c if (c.isalnum() or c in "-_") else "-" for c in v.strip().lower())
slug = "-".join(filter(None, slug.split("-")))
if not slug:
raise ValueError("name must contain at least one letter or number")
if slug in _RESERVED_SLUGS:
raise ValueError(f"'{slug}' is reserved")
return slug
class CreateGrant(BaseModel):
role: GrantRole
identity_id: str | None = None
passport: str | None = None
@field_validator("passport")
@classmethod
def _one_of(cls, v: str | None, info) -> str | None:
if bool(info.data.get("identity_id")) == bool(v):
# Mirrors the database CHECK constraint. Both layers, deliberately:
# this ecosystem already has an invariant enforced only in
# application code across two files, and a double-mint to show for it.
raise ValueError("give exactly one of identity_id or passport")
return v
# --------------------------------------------------------------------------
# helpers
# --------------------------------------------------------------------------
def _sessionmaker(request: Request) -> async_sessionmaker[AsyncSession]:
maker = getattr(request.app.state, "sessionmaker", None)
if maker is None:
raise RepairPointer(
status_code=503,
code="database_unavailable",
speak="We can't reach your projects right now. Nothing has been lost.",
machine_cause="no database sessionmaker on app.state",
remediation_tool=None,
)
return maker
def _repo_owner_login(repo: Repo) -> str:
"""The owning login for a repo as STORED, not as inferred from the caller.
Deriving this from the caller works only while the caller is the owner, and
silently addresses the wrong namespace the moment a collaborator calls. That
class of bug reads as "not found" and is very hard to see.
"""
if repo.passport:
return f"agent-{repo.passport.lower().replace('-', '')}"
return f"u-{repo.identity_id[:24]}"
def _owner_login(caller: Caller) -> str:
"""One namespace rule for humans and agents alike (I-6)."""
if caller.actor_type == ActorType.agent and caller.passport:
return f"agent-{caller.passport.lower().replace('-', '')}"
return f"u-{(caller.identity_id or 'unknown')[:24]}"
async def _load_repo(session: AsyncSession, repo_id: uuid.UUID, caller: Caller) -> Repo:
repo = (await session.execute(select(Repo).where(Repo.id == repo_id))).scalar_one_or_none()
if repo is None or repo.state == RepoState.deleted_soft:
raise RepairPointer(
status_code=404,
code="project_not_found",
speak="We couldn't find that project.",
machine_cause=f"repo {repo_id} not found or soft-deleted",
remediation_tool=None,
)
if not await _may_read(session, repo, caller):
# 404, not 403: a stranger should not learn that a private project exists.
raise RepairPointer(
status_code=404,
code="project_not_found",
speak="We couldn't find that project.",
machine_cause=f"caller {caller.subject} has no grant on repo {repo_id}",
remediation_tool=None,
)
return repo
async def _may_read(session: AsyncSession, repo: Repo, caller: Caller) -> bool:
if caller.actor_type == ActorType.system:
return True
if repo.visibility == Visibility.public:
return True
if caller.identity_id and repo.identity_id == caller.identity_id:
return True
if caller.passport and repo.passport == caller.passport:
return True
return await _active_grant(session, repo, caller) is not None
async def _active_grant(session: AsyncSession, repo: Repo, caller: Caller) -> RepoGrant | None:
now = datetime.now(UTC)
rows = (
await session.execute(
select(RepoGrant).where(
RepoGrant.repo_id == repo.id, RepoGrant.revoked_at.is_(None)
)
)
).scalars()
for g in rows:
if g.expires_at is not None and g.expires_at <= now:
continue # expired grants are not grants
if caller.identity_id and g.grantee_identity_id == caller.identity_id:
return g
if caller.passport and g.grantee_passport == caller.passport:
return g
return None
# --------------------------------------------------------------------------
# routes
# --------------------------------------------------------------------------
@router.post("", status_code=201)
async def create_repo(
body: CreateRepo,
request: Request,
caller: Annotated[Caller, Depends(get_caller)],
) -> dict:
settings = request.app.state.settings
if body.repo_type.value not in settings.repo_types_enabled:
raise RepairPointer(
status_code=409,
code="repo_type_not_enabled",
speak="That kind of project isn't available yet.",
machine_cause=(
f"repo_type={body.repo_type.value} is not in "
f"repo_types_enabled={list(settings.repo_types_enabled)}"
),
remediation_tool=None,
)
# Throttle before any side effect. Checking after would let a rate-limited
# agent still create the Gitea repo and only then be told no.
async with _sessionmaker(request)() as session:
await throttle.enforce(session, settings, caller, "repo.create")
gitea = GiteaClient(settings)
owner = _owner_login(caller)
await gitea.ensure_user(owner, f"{owner}@windygit.com")
created = await gitea.create_repo(
owner=owner,
name=body.name,
description=body.description,
private=body.visibility != Visibility.public,
default_branch="main",
)
async with _sessionmaker(request)() as session:
repo = Repo(
identity_id=caller.identity_id or f"passport:{caller.passport}",
passport=caller.passport,
slug=body.name,
display_name=body.display_name or body.name,
repo_type=body.repo_type,
gitea_repo_id=created.get("id"),
visibility=body.visibility,
default_branch="main",
created_via=(
CreatedVia.agent if caller.actor_type == ActorType.agent else CreatedVia.portal
),
)
session.add(repo)
await session.flush()
await throttle.record(session, caller, "repo.create", repo_id=repo.id)
await session.commit()
await session.refresh(repo)
return {
"id": str(repo.id),
"name": repo.slug,
"repo_type": repo.repo_type.value,
"visibility": repo.visibility.value,
"clone_url": created.get("clone_url"),
"speak": f"'{repo.display_name}' is ready. Everything you save is kept.",
"state_proof": {"gitea_repo_id": repo.gitea_repo_id, "owner": owner},
"next_actions": ["windy_git.grant_access", "windy_git.list_versions"],
}
@router.get("")
async def list_repos(request: Request, caller: Annotated[Caller, Depends(get_caller)]) -> dict:
async with _sessionmaker(request)() as session:
rows = (
await session.execute(
select(Repo).where(Repo.state != RepoState.deleted_soft)
)
).scalars().all()
mine = [r for r in rows if await _may_read(session, r, caller)]
return {
"repos": [
{
"id": str(r.id),
"name": r.slug,
"display_name": r.display_name,
"repo_type": r.repo_type.value,
"visibility": r.visibility.value,
}
for r in mine
],
"count": len(mine),
}
@router.get("/{repo_id}/versions")
async def list_versions(
repo_id: uuid.UUID,
request: Request,
caller: Annotated[Caller, Depends(get_caller)],
) -> dict:
"""G5.5 — history in words a person recognises.
Note what is absent from every user-facing string below: 'commit', 'branch',
'repository'. A person restoring last Tuesday's work should not have to learn
a vocabulary first (I-9, D-9).
"""
async with _sessionmaker(request)() as session:
repo = await _load_repo(session, repo_id, caller)
# Derive the namespace from the REPO, never from the caller: the
# caller-derived form is right only while the caller is the owner, and
# addresses the wrong namespace the moment a collaborator asks. It then
# surfaces as "not found", which is about the hardest bug to see.
owner, slug, display = _repo_owner_login(repo), repo.slug, repo.display_name
gitea = GiteaClient(request.app.state.settings)
commits = await gitea.list_commits(owner, slug)
versions = [
{
"version": len(commits) - i,
"id": c.get("sha"),
"saved_at": (c.get("commit") or {}).get("author", {}).get("date"),
"note": ((c.get("commit") or {}).get("message") or "").strip().split("\n")[0],
"saved_by": (c.get("commit") or {}).get("author", {}).get("name"),
}
for i, c in enumerate(commits)
]
return {
"versions": versions,
"count": len(versions),
# "There are 1 saved versions" is the kind of sloppiness the vocabulary
# law exists to catch. Copy is design material, not decoration (I-9).
"speak": (
f"'{display}' has 1 saved version. You can go back to it."
if len(versions) == 1
else f"'{display}' has {len(versions)} saved versions. "
"You can go back to any of them."
if versions
else f"'{display}' is empty so far."
),
}
@router.post("/{repo_id}/grants", status_code=201)
async def create_grant(
repo_id: uuid.UUID,
body: CreateGrant,
request: Request,
caller: Annotated[Caller, Depends(get_caller)],
) -> dict:
"""G5.3 — the thing Windy Cloud cannot do at all today.
A grant may name a human OR an agent passport, and agent grants expire by
default (env: 90 days). A permanent agent credential is a standing liability
nobody consciously chose.
"""
settings = request.app.state.settings
async with _sessionmaker(request)() as session:
await throttle.enforce(session, settings, caller, "grant.create")
repo = await _load_repo(session, repo_id, caller)
is_owner = (caller.identity_id and repo.identity_id == caller.identity_id) or (
caller.passport and repo.passport == caller.passport
)
if not is_owner and caller.actor_type != ActorType.system:
raise RepairPointer(
status_code=403,
code="not_your_project",
speak="Only the owner can share this project.",
machine_cause=f"{caller.subject} is not the owner of {repo_id}",
remediation_tool=None,
)
expires = (
datetime.now(UTC) + timedelta(days=settings.agent_grant_default_days)
if body.passport
else None
)
grant = RepoGrant(
repo_id=repo.id,
grantee_identity_id=body.identity_id,
grantee_passport=body.passport,
role=body.role,
granted_by=caller.subject,
expires_at=expires,
)
session.add(grant)
await session.flush()
await throttle.record(session, caller, "grant.create", repo_id=repo.id)
await session.commit()
await session.refresh(grant)
grant_id, role = grant.id, grant.role
who = body.identity_id or body.passport
return {
"id": str(grant_id),
"role": role.value,
"grantee": who,
"expires_at": expires.isoformat() if expires else None,
"speak": (
f"They can now help with '{repo.display_name}'."
+ (" Access ends automatically in 90 days." if expires else "")
),
"next_actions": ["windy_git.list_grants", "windy_git.revoke_access"],
}
@router.get("/{repo_id}/grants")
async def list_grants(
repo_id: uuid.UUID,
request: Request,
caller: Annotated[Caller, Depends(get_caller)],
) -> dict:
now = datetime.now(UTC)
async with _sessionmaker(request)() as session:
repo = await _load_repo(session, repo_id, caller)
rows = (
await session.execute(select(RepoGrant).where(RepoGrant.repo_id == repo.id))
).scalars().all()
return {
"grants": [
{
"id": str(g.id),
"grantee": g.grantee_identity_id or g.grantee_passport,
"kind": "person" if g.grantee_identity_id else "helper",
"role": g.role.value,
"expires_at": g.expires_at.isoformat() if g.expires_at else None,
"active": g.revoked_at is None
and (g.expires_at is None or g.expires_at > now),
}
for g in rows
]
}
@router.delete("/{repo_id}/grants/{grant_id}")
async def revoke_grant(
repo_id: uuid.UUID,
grant_id: uuid.UUID,
request: Request,
caller: Annotated[Caller, Depends(get_caller)],
) -> dict:
async with _sessionmaker(request)() as session:
repo = await _load_repo(session, repo_id, caller)
grant = (
await session.execute(select(RepoGrant).where(RepoGrant.id == grant_id))
).scalar_one_or_none()
if grant is None or grant.repo_id != repo.id:
raise RepairPointer(
status_code=404,
code="grant_not_found",
speak="We couldn't find that access to remove.",
machine_cause=f"grant {grant_id} not on repo {repo_id}",
remediation_tool=None,
)
grant.revoked_at = datetime.now(UTC)
await session.commit()
return {"revoked": True, "speak": "That access has been removed."}
@router.get("/{repo_id}")
async def get_repo(
repo_id: uuid.UUID,
request: Request,
caller: Annotated[Caller, Depends(get_caller)],
) -> dict:
async with _sessionmaker(request)() as session:
repo = await _load_repo(session, repo_id, caller)
versions = (
await session.execute(select(RepoVersion).where(RepoVersion.repo_id == repo.id))
).scalars().all()
return {
"id": str(repo.id),
"name": repo.slug,
"display_name": repo.display_name,
"repo_type": repo.repo_type.value,
"visibility": repo.visibility.value,
"cloud_folder_ref": repo.cloud_folder_ref,
"recorded_versions": len(versions),
"state": repo.state.value,
}
# --------------------------------------------------------------------------
# I-4 / G11 — the off-site copy
# --------------------------------------------------------------------------
@router.post("/{repo_id}/mirror", status_code=201)
async def enable_mirror(
repo_id: uuid.UUID,
request: Request,
caller: Annotated[Caller, Depends(get_caller)],
) -> dict:
"""Turn on the continuous off-site copy.
Deliberately idempotent and deliberately loud on failure: a mirror that
quietly stopped working is worse than no mirror, because it is a backup you
believe in.
"""
settings = request.app.state.settings
async with _sessionmaker(request)() as session:
repo = await _load_repo(session, repo_id, caller)
owner = _repo_owner_login(repo)
display, slug = repo.display_name, repo.slug
private = repo.visibility != Visibility.public
mirror = MirrorService(settings)
remote = await mirror.ensure_github_repo(slug, display, private)
await mirror.attach_push_mirror(owner, slug, remote)
async with _sessionmaker(request)() as session:
session.add(
Mirror(repo_id=repo_id, remote_url=remote, direction="push", state=MirrorState.healthy)
)
await session.commit()
return {
"remote": remote,
"sync_on_commit": True,
"speak": "A second copy of this project is now kept somewhere else, automatically.",
"state_proof": {"remote": remote},
"next_actions": ["windy_git.mirror_status"],
}
@router.get("/{repo_id}/mirror")
async def mirror_status(
repo_id: uuid.UUID,
request: Request,
caller: Annotated[Caller, Depends(get_caller)],
) -> dict:
async with _sessionmaker(request)() as session:
repo = await _load_repo(session, repo_id, caller)
owner, slug = _repo_owner_login(repo), repo.slug
status = await MirrorService(request.app.state.settings).status(owner, slug)
speak = {
"healthy": "A second copy of this project is up to date.",
"degraded": "The second copy is behind. Your work here is safe.",
"absent": "There is no second copy of this project yet.",
"pending": "The second copy is set up and hasn't run yet.",
"unconfigured": "Off-site copies aren't switched on yet.",
"unknown": "We can't tell how the second copy is doing right now.",
}[status["state"]]
return {**status, "speak": speak}

169
api/app/routes/webhooks.py Normal file
View File

@@ -0,0 +1,169 @@
"""Eternitas webhook receiver (G3.5) — revocation is FAIL-CLOSED.
When a passport is revoked, every credential that passport holds here dies in
one transaction: tokens revoked, grants revoked, in-flight CI cancelled. A
revocation that takes effect "eventually" is not a revocation.
Two traps are avoided here on purpose, both paid for elsewhere in the ecosystem:
1. **Strip the `sha256=` prefix before comparing.** A sibling receiver compared
the whole header against a bare hex digest and therefore returned 401
forever — the subscription looked wired and never once delivered.
2. **HMAC the RAW REQUEST BYTES, not a re-serialised body.** `JSON.stringify` of
a parsed body reorders keys and changes whitespace, so the digest never
matches what the sender signed. Same outcome: deterministic 401.
Both failures are silent from the sender's side — Eternitas records a delivery
attempt, the receiver records a rejection, and nobody notices for weeks.
"""
from __future__ import annotations
import hashlib
import hmac
import logging
from datetime import UTC, datetime
from fastapi import APIRouter, Header, Request
from sqlalchemy import select, update
from api.app.errors import RepairPointer
from api.app.models.core import AgentToken, Repo, RepoGrant
log = logging.getLogger(__name__)
router = APIRouter(prefix="/api/v1/webhooks", tags=["webhooks"])
def _verify(raw: bytes, header: str | None, secret: str) -> bool:
if not header or not secret:
return False
# Trap 1: senders prefix the digest. Compare digests, not decorated strings.
presented = header.split("=", 1)[1] if header.startswith("sha256=") else header
# Trap 2: sign the bytes that arrived, never a re-serialised object.
expected = hmac.new(secret.encode(), raw, hashlib.sha256).hexdigest()
return hmac.compare_digest(presented, expected)
@router.post("/eternitas")
async def eternitas_webhook(
request: Request,
x_eternitas_event: str | None = Header(default=None),
x_eternitas_signature: str | None = Header(default=None),
) -> dict:
settings = request.app.state.settings
raw = await request.body()
# `platform.test_ping` is the ONE event accepted without verification, and
# the reason is structural rather than convenient.
#
# Eternitas generates the webhook secret at registration time and pings the
# URL to prove it is reachable BEFORE returning that secret. The ping is
# signed — with a secret the receiver cannot possibly hold yet. So the
# signature is unverifiable by construction, not by oversight.
#
# Accepting it is safe because the event is definitionally a no-op: nothing
# is read, nothing is written, `acted` is false. Every event that changes
# anything — `passport.revoked` above all — still requires a valid HMAC
# below. The alternative was `skip_validation: true` at registration, which
# would permanently disable reachability checking for this platform to
# solve a one-time ordering problem.
if x_eternitas_event == "platform.test_ping" or (
not x_eternitas_event and not x_eternitas_signature
):
return {
"ready": True,
"acted": False,
"detail": (
"reachability ping acknowledged; it is unverifiable by "
"construction and changes nothing. Signed events are verified."
),
}
secret = settings.eternitas_webhook_secret
if not secret:
# I-8: refuse rather than accept unverified instructions about identity.
raise RepairPointer(
status_code=503,
code="webhook_secret_unset",
speak="We can't accept that update yet.",
machine_cause="ETERNITAS_WEBHOOK_SECRET is unset; refusing unverified webhooks",
remediation_tool=None,
)
if not _verify(raw, x_eternitas_signature, secret):
raise RepairPointer(
status_code=401,
code="webhook_signature_invalid",
speak="We couldn't confirm where that update came from, so we ignored it.",
machine_cause="HMAC mismatch on the raw request body",
remediation_tool=None,
)
payload = await request.json()
event = x_eternitas_event or payload.get("event") or "unknown"
passport = (
payload.get("passport")
or payload.get("passport_number")
or (payload.get("data") or {}).get("passport")
)
if event != "passport.revoked":
# Acknowledge without pretending to have acted. A 200 here means
# "received", and the body says exactly what was done — which is nothing.
log.info("eternitas event %s received (no handler)", event)
return {"received": True, "event": event, "acted": False}
if not passport:
raise RepairPointer(
status_code=422,
code="revocation_missing_passport",
speak="That update didn't say which helper it was about.",
machine_cause=f"passport.revoked payload carried no passport field: {list(payload)}",
remediation_tool=None,
)
maker = getattr(request.app.state, "sessionmaker", None)
if maker is None:
raise RepairPointer(
status_code=503,
code="database_unavailable",
speak="We couldn't apply that update. Please try again.",
machine_cause="no database sessionmaker; refusing to acknowledge a revocation we did not apply",
remediation_tool=None,
)
now = datetime.now(UTC)
async with maker() as session:
# One transaction. A partial revocation is a security hole that reports
# success.
tokens = await session.execute(
update(AgentToken)
.where(AgentToken.passport == passport, AgentToken.revoked_at.is_(None))
.values(revoked_at=now, revoked_reason="eternitas:passport.revoked")
)
grants = await session.execute(
update(RepoGrant)
.where(RepoGrant.grantee_passport == passport, RepoGrant.revoked_at.is_(None))
.values(revoked_at=now)
)
owned = (
await session.execute(select(Repo.id).where(Repo.passport == passport))
).scalars().all()
await session.commit()
log.warning(
"passport %s revoked: %s tokens, %s grants, %s owned repos",
passport, tokens.rowcount, grants.rowcount, len(owned),
)
return {
"received": True,
"event": event,
"acted": True,
"passport": passport,
"tokens_revoked": tokens.rowcount,
"grants_revoked": grants.rowcount,
"owned_repos": len(owned),
# state_proof so the caller can verify rather than trust (section 0.6).
"state_proof": {"revoked_at": now.isoformat()},
}

View File

@@ -0,0 +1,177 @@
"""The Gitea membrane (I-1 / D-2).
Everything this cell needs from Gitea goes through this file and through Gitea's
REST API. Nothing else in the codebase imports Gitea concepts, and nothing
anywhere writes Gitea's database directly — it has its own role and its own
database precisely so that boundary is a permission rather than a promise.
Keeping the dependency here is what makes D-2 affordable: Gitea ships every two
or three months including security fixes, and a diverged fork becomes the whole
job within a year for a small team. One file is a seam. A merged source tree is
a marriage.
"""
from __future__ import annotations
import secrets
from typing import Any
import httpx
from api.app.config import Settings
from api.app.errors import RepairPointer, provider_unconfigured
from api.app.telemetry import synthetic_headers
_TIMEOUT = httpx.Timeout(20.0, connect=5.0)
class GiteaClient:
def __init__(self, settings: Settings) -> None:
self._s = settings
def _require(self) -> None:
if not self._s.gitea_configured:
raise provider_unconfigured("gitea", "GITEA_ADMIN_TOKEN")
def _headers(self) -> dict[str, str]:
return {
"Authorization": f"token {self._s.gitea_admin_token}",
"Content-Type": "application/json",
**synthetic_headers(), # end-to-end synthetic convention (Telemetry UPDATE 4)
}
async def _request(self, method: str, path: str, **kw: Any) -> httpx.Response:
self._require()
async with httpx.AsyncClient(timeout=_TIMEOUT) as client:
return await client.request(
method,
f"{self._s.gitea_base_url}/api/v1{path}",
headers=self._headers(),
**kw,
)
# ---- users ------------------------------------------------------------
async def ensure_user(self, username: str, email: str) -> dict:
"""Idempotent. Gitea is the component; our `repos` table is the truth."""
r = await self._request("GET", f"/users/{username}")
if r.status_code == 200:
return r.json()
# Gitea requires a password field on admin user-create and rejects null
# with a bare 400. Nobody ever uses this one: humans arrive through OIDC
# and agents through scoped passport-bound tokens, and local password
# sign-in is disabled server-wide. So we generate a credential that is
# never stored, never returned and never recoverable — an unusable
# password is safer than a blank one or a shared default.
r = await self._request(
"POST",
"/admin/users",
json={
"username": username,
"email": email,
"password": secrets.token_urlsafe(48),
"must_change_password": False,
},
)
if r.status_code not in (200, 201):
raise RepairPointer(
status_code=502,
code="gitea_user_create_failed",
speak="We couldn't finish setting up that account. Nothing was lost.",
machine_cause=f"POST /admin/users -> {r.status_code}: {r.text[:200]}",
remediation_tool="windy_git.repair.retry_user_create",
)
return r.json()
# ---- repos ------------------------------------------------------------
async def create_repo(
self, owner: str, name: str, description: str, private: bool, default_branch: str
) -> dict:
r = await self._request(
"POST",
f"/admin/users/{owner}/repos",
json={
"name": name,
"description": description,
"private": private,
"auto_init": True,
"default_branch": default_branch,
# G4.5 — the LFS threshold is load-bearing, not tidiness. A large
# blob inside a git pack cannot be resumed or offloaded, and a
# plain push of one dies at Cloudflare's ~100s ceiling (G4A.5).
"gitignores": "",
},
)
if r.status_code not in (200, 201):
raise RepairPointer(
status_code=502 if r.status_code >= 500 else 409,
code="repo_create_failed",
speak="We couldn't create that project. Try a different name.",
machine_cause=f"POST /admin/users/{owner}/repos -> {r.status_code}: {r.text[:200]}",
remediation_tool=None,
)
return r.json()
async def delete_repo(self, owner: str, name: str) -> None:
r = await self._request("DELETE", f"/repos/{owner}/{name}")
if r.status_code not in (204, 404):
raise RepairPointer(
status_code=502,
code="repo_delete_failed",
speak="We couldn't remove that project. It is still there and still yours.",
machine_cause=f"DELETE /repos/{owner}/{name} -> {r.status_code}",
remediation_tool="windy_git.repair.retry_delete",
)
async def get_repo(self, owner: str, name: str) -> dict | None:
r = await self._request("GET", f"/repos/{owner}/{name}")
return r.json() if r.status_code == 200 else None
# ---- history (G5.5 / G5.6) -------------------------------------------
async def list_commits(self, owner: str, name: str, limit: int = 50) -> list[dict]:
r = await self._request(
"GET", f"/repos/{owner}/{name}/commits", params={"limit": limit}
)
if r.status_code == 409:
return [] # empty repo — a real state, not an error
if r.status_code != 200:
raise RepairPointer(
status_code=502,
code="history_unavailable",
speak="We couldn't load the history for that project just now.",
machine_cause=f"GET commits -> {r.status_code}",
remediation_tool="windy_git.repair.rebuild_index",
)
return r.json()
# ---- collaborators (the shelter's enforcement half, G5.3) ------------
async def put_collaborator(self, owner: str, name: str, user: str, permission: str) -> None:
r = await self._request(
"PUT",
f"/repos/{owner}/{name}/collaborators/{user}",
json={"permission": permission},
)
if r.status_code not in (204, 201, 200):
raise RepairPointer(
status_code=502,
code="grant_apply_failed",
speak="We couldn't share that project yet. Nobody was given access.",
machine_cause=f"PUT collaborator -> {r.status_code}: {r.text[:200]}",
remediation_tool="windy_git.repair.resync_grants",
)
async def delete_collaborator(self, owner: str, name: str, user: str) -> None:
r = await self._request("DELETE", f"/repos/{owner}/{name}/collaborators/{user}")
if r.status_code not in (204, 404):
raise RepairPointer(
status_code=502,
code="grant_revoke_failed",
# The honest failure: we could not take access away. That is the
# scarier direction, so it is stated plainly rather than softened.
speak="We could not remove that person's access. Please try again.",
machine_cause=f"DELETE collaborator -> {r.status_code}",
remediation_tool="windy_git.repair.resync_grants",
)
async def version(self) -> str:
r = await self._request("GET", "/version")
return r.json().get("version", "unknown")

179
api/app/services/mirror.py Normal file
View File

@@ -0,0 +1,179 @@
"""I-4 — never a one-way door (strand G11).
Every repo push-mirrors to GitHub, continuously, from the first save. This is
not a nicety and it is not a migration step: it is the thing that makes moving
off GitHub a reversible decision rather than a bet.
Today GitHub's durability is free to Grant. The moment repos live only here,
backups, restore rehearsal and a second copy stop being someone else's job — and
the August audits found **no rehearsed restore anywhere in the ecosystem**, for
anything. A continuous mirror buys back that safety for zero dollars.
Mirror health is a monitored, alerting signal. A mirror nobody checks is a
belief, not a backup — and this ecosystem has already learned that lesson the
expensive way with a fleet canary that sat dead for 37 days while everything
downstream assumed it was fine.
"""
from __future__ import annotations
import logging
from datetime import UTC, datetime
import httpx
from api.app.config import Settings
from api.app.errors import RepairPointer
log = logging.getLogger(__name__)
_TIMEOUT = httpx.Timeout(30.0, connect=10.0)
class MirrorService:
"""Creates the GitHub counterpart and asks Gitea to keep it in step.
Gitea owns the actual replication (`push_mirrors`), because a hand-rolled
mirror loop is a background job that fails silently — which is precisely how
the registry's integrity refresh went its entire life calling a 404 and
incrementing a counter instead of raising.
"""
def __init__(self, settings: Settings) -> None:
self._s = settings
@property
def configured(self) -> bool:
return bool(self._s.github_token and self._s.github_owner)
def _require(self) -> None:
if not self.configured:
raise RepairPointer(
status_code=503,
code="mirror_unconfigured",
speak="The off-site copy isn't switched on yet.",
machine_cause="GITHUB_TOKEN or GITHUB_OWNER is unset; refusing to claim a mirror",
remediation_tool=None,
)
async def ensure_github_repo(self, name: str, description: str, private: bool) -> str:
"""Idempotent. Returns the clone URL of the off-site copy."""
self._require()
headers = {
"Authorization": f"Bearer {self._s.github_token}",
"Accept": "application/vnd.github+json",
}
owner = self._s.github_owner
async with httpx.AsyncClient(timeout=_TIMEOUT) as client:
existing = await client.get(
f"https://api.github.com/repos/{owner}/{name}", headers=headers
)
if existing.status_code == 200:
return existing.json()["clone_url"]
created = await client.post(
"https://api.github.com/user/repos",
headers=headers,
json={
"name": name,
"description": f"{description} (Windy Git mirror)".strip(),
"private": private,
"auto_init": False,
},
)
if created.status_code not in (200, 201):
raise RepairPointer(
status_code=502,
code="mirror_target_failed",
speak="We couldn't set up the off-site copy. Your work here is safe.",
machine_cause=f"POST /user/repos -> {created.status_code}: {created.text[:200]}",
remediation_tool="windy_git.repair.resync_mirror",
)
return created.json()["clone_url"]
async def attach_push_mirror(self, owner: str, repo: str, remote_url: str) -> None:
"""Ask Gitea to keep the off-site copy in step on every save."""
self._require()
async with httpx.AsyncClient(timeout=_TIMEOUT) as client:
r = await client.post(
f"{self._s.gitea_base_url}/api/v1/repos/{owner}/{repo}/push_mirrors",
headers={
"Authorization": f"token {self._s.gitea_admin_token}",
"Content-Type": "application/json",
},
json={
"remote_address": remote_url,
"remote_username": self._s.github_owner,
"remote_password": self._s.github_token,
"interval": self._s.mirror_interval,
# The important one: mirror on every save, not just on a timer.
# An hourly timer means an hour of work can be the thing you
# lose, and the window is invisible until it costs you.
"sync_on_commit": True,
},
)
if r.status_code not in (200, 201):
raise RepairPointer(
status_code=502,
code="mirror_attach_failed",
speak="We couldn't keep the off-site copy in step. Your work here is safe.",
machine_cause=f"POST push_mirrors -> {r.status_code}: {r.text[:200]}",
remediation_tool="windy_git.repair.resync_mirror",
)
async def status(self, owner: str, repo: str) -> dict:
"""Report what is TRUE, including 'we do not know'.
`me-fleet.ts:22-25` in a sibling service refuses to say "online" when it
only knows "registered". Same posture here: an unconfigured mirror is
reported as unknown, never as healthy.
"""
if not self.configured:
return {"state": "unconfigured", "lag_seconds": None, "last_success_at": None}
async with httpx.AsyncClient(timeout=_TIMEOUT) as client:
r = await client.get(
f"{self._s.gitea_base_url}/api/v1/repos/{owner}/{repo}/push_mirrors",
headers={"Authorization": f"token {self._s.gitea_admin_token}"},
)
if r.status_code != 200:
return {"state": "unknown", "detail": f"gitea -> {r.status_code}"}
mirrors = r.json()
if not mirrors:
return {"state": "absent", "lag_seconds": None, "last_success_at": None}
m = mirrors[0]
last = m.get("last_update") or m.get("last_updated")
lag = None
if last:
try:
lag = int(
(datetime.now(UTC) - datetime.fromisoformat(last.replace("Z", "+00:00")))
.total_seconds()
)
except ValueError:
lag = None
# A mirror that has NEVER run is not the same thing as one that is
# behind, and collapsing the two is how a backup that was never made
# gets read as a backup that is merely stale. Gitea reports the epoch
# for "not yet", which arithmetic turns into a 56-year lag and a
# confident "degraded".
never_synced = not last or last.startswith("1970-01-01")
# I-4: lag over the threshold is a P2, not a shrug.
if never_synced:
state = "pending"
lag = None
elif lag is None:
state = "unknown"
elif lag > self._s.mirror_lag_p2_seconds:
state = "degraded"
else:
state = "healthy"
return {
"state": state,
"lag_seconds": lag,
"last_success_at": None if never_synced else last,
"remote": m.get("remote_address"),
}

255
api/app/telemetry.py Normal file
View File

@@ -0,0 +1,255 @@
"""Field telemetry to the admin ledger (admin.windyword.ai) — step 2, 2026-09-23.
Shapes are DECLARED with Telemetry Boss (the ledger owner); the server
quarantines any row that doesn't match, so never add a key or a code here
without re-declaring it first:
service.boot once per process start {commit_sha, version, environment}
service.health hourly, in-process interval_s, uptime_s, requests,
errors_5xx, errors_4xx,
refusals_4xx, p95_ms
forge.auth.failed every refused request {code, http_status, caller,
route?, upstream_status?}
Refusals come first: a refused caller is the most expensive silent failure
("the button did nothing"). An UNAUTHENTICATED caller has no trustworthy id, so
per the all-lanes actor rule the row is actor_type "system", no actor_id, and
the caller class goes in metadata.caller.
Privacy: codes, statuses, route TEMPLATES, counts, durations. Never a passport
number, an email, a token fragment or a concrete path with names in it.
No token → nothing is sent and nothing is buffered. A failed flush keeps the
rows (bounded) and retries on the next tick; it never raises into a request.
"""
from __future__ import annotations
import asyncio
import contextvars
import json
import logging
import time
import urllib.request
from datetime import UTC, datetime
log = logging.getLogger("windy-git.telemetry")
PLATFORM, SERVICE = "windy-git", "api"
FLUSH_EVERY_S = 60
HEALTH_EVERY_S = 3600
MAX_BUFFER = 5000
MAX_LATENCY_SAMPLES = 20000
# The declared forge.auth.failed code enum (Telemetry Boss, 2026-09-23). A code
# outside this set is NOT a refusal row — it counts in errors_* instead.
AUTH_CODES = frozenset(
{
"not_signed_in",
"token_invalid",
"token_unrecognised",
"ept_invalid",
"passport_revoked",
"passport_unresolvable",
"agent_read_only",
"agent_rate_limited",
"quota_exceeded",
"trust_unavailable",
"throttle_unavailable",
"service_token_invalid",
"service_auth_unconfigured",
}
)
def _iso(epoch: float) -> str:
return datetime.fromtimestamp(epoch, UTC).isoformat().replace("+00:00", "Z")
# Ecosystem convention (Telemetry UPDATE 4): synthetic traffic travels END TO
# END. Originators (canaries, probes, journeys) send `X-Windy-Synthetic: 1`;
# every service marks all of that request's rows synthetic:true AND forwards the
# header on every downstream call. Absent = real. Never strip it, never set it
# on real traffic. The label separates rows — it never suppresses them.
SYNTHETIC: contextvars.ContextVar[bool] = contextvars.ContextVar("windy_synthetic", default=False)
def is_synthetic(headers) -> bool:
return bool((headers.get("x-windy-synthetic") or "").strip())
def synthetic_headers() -> dict:
"""Merge into every downstream request made while serving this one."""
return {"X-Windy-Synthetic": "1"} if SYNTHETIC.get() else {}
def caller_class(headers) -> str:
"""Declared values: anonymous_human | anonymous_agent | unknown."""
from api.app.ept import looks_like_ept
if headers.get("x-service-token"):
return "unknown"
auth = headers.get("authorization") or ""
if not auth.lower().startswith("bearer "):
return "unknown"
return "anonymous_agent" if looks_like_ept(auth.split(" ", 1)[1].strip()) else "anonymous_human"
class Telemetry:
def __init__(
self,
url: str,
token: str,
*,
environment: str = "",
commit_sha: str | None = None,
version: str = "",
) -> None:
self.url, self.token = url, token
self.environment, self.commit_sha, self.version = environment, commit_sha, version
self.started = time.time()
self.buffer: list[dict] = []
self.auth_codes = AUTH_CODES
self._reset_window()
@property
def enabled(self) -> bool:
return bool(self.token)
def _reset_window(self) -> None:
self.window_start = time.time()
self.requests = self.errors_5xx = self.errors_4xx = self.refusals_4xx = 0
self.latencies_ms: list[float] = []
# UPDATE 7: rows the ledger quarantined (it still answers 202) and rows
# this process lost (buffer overflow). Non-zero = a bug in this emitter.
self.quarantined = self.dropped = 0
# ---- recording (never raises into a request) --------------------------
def record_request(self, status: int, duration_ms: float, *, refused: bool = False) -> None:
self.requests += 1
if status >= 500:
self.errors_5xx += 1
elif refused:
self.refusals_4xx += 1
elif status >= 400:
self.errors_4xx += 1
if len(self.latencies_ms) < MAX_LATENCY_SAMPLES:
self.latencies_ms.append(duration_ms)
def _event(self, event_type: str, metadata: dict, *, ts: float | None = None) -> None:
if not self.enabled:
return
self.buffer.append(
{
"ts": _iso(ts or time.time()),
"platform": PLATFORM,
"service": SERVICE,
"event_type": event_type,
"actor_type": "system",
"metadata": metadata,
}
)
if len(self.buffer) > MAX_BUFFER:
self.dropped += len(self.buffer) - MAX_BUFFER
del self.buffer[: len(self.buffer) - MAX_BUFFER]
def boot(self) -> None:
meta = {"version": self.version, "environment": self.environment}
if self.commit_sha: # unknown is absent, never invented (I-12)
meta["commit_sha"] = self.commit_sha
self._event("service.boot", meta, ts=self.started)
def auth_failed(
self,
*,
code: str,
http_status: int,
caller: str,
route: str | None = None,
upstream_status: int | None = None,
synthetic: bool = False,
) -> None:
if code not in AUTH_CODES:
return
meta: dict = {
"code": code,
"http_status": int(http_status),
"caller": caller,
"synthetic": bool(synthetic),
}
if route:
meta["route"] = route
if upstream_status is not None:
meta["upstream_status"] = int(upstream_status)
self._event("forge.auth.failed", meta)
def health_row(self) -> dict:
now = time.time()
meta = {
"interval_s": int(now - self.window_start),
"uptime_s": int(now - self.started),
"requests": self.requests,
"errors_5xx": self.errors_5xx,
"errors_4xx": self.errors_4xx,
"refusals_4xx": self.refusals_4xx,
"telemetry_quarantined": self.quarantined,
"telemetry_dropped": self.dropped,
}
if self.latencies_ms: # no traffic = no p95, not a fake 0
s = sorted(self.latencies_ms)
meta["p95_ms"] = int(s[min(len(s) - 1, int(0.95 * (len(s) - 1) + 0.5))])
return meta
def health(self) -> None:
self._event("service.health", self.health_row())
self._reset_window()
# ---- sending ------------------------------------------------------------
def _post(self, batch: list[dict]) -> tuple[int, dict]:
req = urllib.request.Request(
self.url,
data=json.dumps({"events": batch}).encode(),
method="POST",
headers={
"Authorization": f"Bearer {self.token}",
"Content-Type": "application/json",
"User-Agent": "windy-git-api-telemetry/1",
},
)
with urllib.request.urlopen(req, timeout=20) as r:
try:
body = json.loads(r.read() or b"{}")
except ValueError:
body = {}
return r.status, body if isinstance(body, dict) else {}
async def flush(self) -> None:
if not self.enabled or not self.buffer:
return
batch = self.buffer[:500]
try:
status, body = await asyncio.to_thread(self._post, batch)
except Exception as exc: # noqa: BLE001 - telemetry must never take the API down
log.warning("telemetry flush failed (%d rows kept): %s", len(self.buffer), exc)
return
if 200 <= status < 300:
del self.buffer[: len(batch)]
self.note_quarantine(body)
def note_quarantine(self, body: dict) -> None:
# 202 does NOT mean every row landed: refused rows are dead-lettered.
q = body.get("quarantined")
if isinstance(q, int) and q > 0:
self.quarantined += q
log.warning("telemetry: %d row(s) QUARANTINED by the ledger: %s", q,
"; ".join(map(str, body.get("rejections") or [])) or "no reason given")
async def run(self) -> None:
"""The one in-process timer: flush every minute, heartbeat every hour."""
last_health = time.monotonic()
while True:
await asyncio.sleep(FLUSH_EVERY_S)
if time.monotonic() - last_health >= HEALTH_EVERY_S:
self.health()
last_health = time.monotonic()
await self.flush()

187
api/app/throttle.py Normal file
View File

@@ -0,0 +1,187 @@
"""EI-band velocity limits (G3.4) — throttle by trust, never by omission.
`BAND_MULTIPLIER` and the `rate_*_per_day` settings existed since G0 and were
**read by nothing**: a documented invariant with no implementation, which is the
exact failure pattern the ecosystem audits kept finding. This module is the
missing consumer.
The doctrine (§0.6) is *capability-completeness*: an agent is never denied a
capability it should have, it is **rate-limited by how much it has proven**.
Platinum gets 10x, gold 4x, standard 1x, watch 0.5x, and untrusted is read-only.
Counting is done against `agent_actions`, which is already the append-only record
of every agent write. That is deliberate: a limiter with its own private counter
disagrees with the audit log the moment either is restarted, and then nobody can
say what actually happened. One source of truth, queried.
**Fail-closed.** If the count cannot be taken, the action is refused. A limiter
that fails open is decoration — it protects you right up until the moment
something is wrong, which is the only moment it matters.
"""
from __future__ import annotations
import logging
from datetime import UTC, datetime, timedelta
from sqlalchemy import func, select
from sqlalchemy.ext.asyncio import AsyncSession
from api.app.auth import BAND_MULTIPLIER, ActorType, Caller
from api.app.config import Settings
from api.app.errors import RepairPointer
from api.app.models.core import AgentAction
log = logging.getLogger(__name__)
WINDOW = timedelta(days=1)
# Actions this module ACTUALLY enforces: they pass through our API, so we can
# both count and refuse them.
ACTION_BASE: dict[str, str] = {
"repo.create": "rate_repo_creates_per_day",
"grant.create": "rate_grants_per_day",
}
# ⚠️ DECLARED BUT NOT ENFORCEABLE HERE — and named, rather than quietly listed
# alongside the real ones.
#
# `git push` goes straight to Gitea over HTTPS and never touches this API, so
# nothing records a `push` action and a count of them would be zero forever.
# Listing these in ACTION_BASE (as this module first did) would make `enforce`
# look up a limit, count nothing, and allow everything — a silent no-op wearing
# the costume of a control. That is the same dead-code pattern this module was
# written to remove, and it is worse here because the config name implies the
# protection exists.
#
# Enforcing push velocity requires a Gitea-side hook (pre-receive or the push
# webhook) that reports into `agent_actions`. Until that exists these settings
# are inert, and saying so is the honest option.
NOT_ENFORCED_HERE: dict[str, str] = {
"push": "rate_pushes_per_day",
"push.force": "rate_force_pushes_per_day",
}
def limit_for(settings: Settings, action: str, band: str | None) -> int:
"""Effective per-day allowance. Unknown bands get the standard rate.
Unknown-band handling is a real decision, not a default. Eternitas began
emitting `unproven` on 2026-07-30 without it appearing in the documented
enum. Treating an unrecognised band as untrusted would lock out every freshly
hatched agent the day Eternitas adds a name; treating it as platinum would be
a hole. Standard-with-no-bonus is the honest middle.
"""
base = getattr(settings, ACTION_BASE[action])
mult = BAND_MULTIPLIER.get((band or "").lower(), 1.0)
return int(base * mult)
async def enforce(
session: AsyncSession,
settings: Settings,
caller: Caller,
action: str,
) -> None:
"""Raise 429 if this agent has spent its allowance. No-op for humans.
Humans are governed by account tier and their own session; this is the
agent-velocity control specifically (I-6: parity in capability, asymmetry in
throttle).
"""
if caller.actor_type != ActorType.agent or not caller.passport:
return
if action not in ACTION_BASE: # unknown action = unlimited would be a hole
raise RepairPointer(
status_code=500,
code="throttle_unknown_action",
speak="Something went wrong on our side. Nothing was changed.",
machine_cause=f"no rate base configured for action {action!r}",
remediation_tool=None,
)
band = (caller.band or "").lower()
# Untrusted is read-only. This is a capability decision, not a rate: it gets
# a 403 with a different explanation, because "slow down" would be a lie.
if BAND_MULTIPLIER.get(band, 1.0) <= 0:
raise RepairPointer(
status_code=403,
code="agent_read_only",
speak="That helper can look, but it isn't allowed to make changes yet.",
machine_cause=f"passport {caller.passport} band={band!r} is read-only",
remediation_tool=None,
)
allowed = limit_for(settings, action, band)
since = datetime.now(UTC) - WINDOW
try:
used = (
await session.execute(
select(func.count())
.select_from(AgentAction)
.where(
AgentAction.passport == caller.passport,
AgentAction.action == action,
AgentAction.result == "ok",
AgentAction.ts >= since,
)
)
).scalar_one()
except Exception as exc: # noqa: BLE001
# FAIL CLOSED. A limiter that fails open protects you until the moment
# something is wrong, which is the only moment it matters.
log.warning("throttle count failed for %s/%s: %s", caller.passport, action, exc)
raise RepairPointer(
status_code=503,
code="throttle_unavailable",
speak="We couldn't check that helper's limits, so we didn't make the change.",
machine_cause=f"agent_actions count failed: {type(exc).__name__}",
remediation_tool=None,
) from exc
if used >= allowed:
raise RepairPointer(
status_code=429,
code="agent_rate_limited",
speak=(
"That helper has done a lot in the last day, so we've paused it. "
"It'll be able to continue shortly."
),
machine_cause=(
f"passport {caller.passport} used {used}/{allowed} of {action} "
f"in 24h (band={band or 'unknown'})"
),
remediation_tool=None,
used=used,
allowed=allowed,
band=band or "unknown",
)
async def record(
session: AsyncSession,
caller: Caller,
action: str,
result: str = "ok",
repo_id=None,
) -> None:
"""Append the action that was just allowed.
Written AFTER the work succeeds, on purpose: counting attempts would let a
failing agent exhaust its own allowance by retrying, turning a transient
error into a lockout.
"""
if caller.actor_type != ActorType.agent or not caller.passport:
return
session.add(
AgentAction(
passport=caller.passport,
repo_id=repo_id,
action=action,
ei_at_action=caller.band,
result=result,
)
)
await session.flush()

View File

@@ -0,0 +1,312 @@
"""Behavioral tests for EPT verification and EI throttling.
These sign real ES256 tokens with a locally-generated key and verify against a
locally-served key set, so they exercise the ACTUAL crypto path with no network
dependency and no reliance on Eternitas being reachable.
This file exists because the suite it joins was ~86 "does the source contain
this string" assertions and zero that ran the auth decision — which is how a
live impersonation bypass passed every test on 2026-08-13.
"""
from __future__ import annotations
import time
import jwt
import pytest
from cryptography.hazmat.primitives.asymmetric import ec
from api.app import ept as ept_mod
from api.app.auth import BAND_MULTIPLIER
from api.app.config import Settings
from api.app.ept import EptInvalid, verify_ept
from api.app.throttle import ACTION_BASE, limit_for
ISSUER = "eternitas.ai"
KID = "test-key-1"
@pytest.fixture
def signing(monkeypatch):
"""A real EC keypair; point the verifier's JWKS lookup at its public half."""
key = ec.generate_private_key(ec.SECP256R1())
class _FakeJWK:
def __init__(self, k):
self.key = k
class _FakeClient:
def __init__(self, *a, **kw):
pass
def get_signing_key_from_jwt(self, token):
header = jwt.get_unverified_header(token)
if header.get("kid") != KID:
raise Exception(f"unknown kid {header.get('kid')!r}")
return _FakeJWK(key.public_key())
monkeypatch.setattr(ept_mod, "_jwks_client", None)
monkeypatch.setattr(ept_mod, "PyJWKClient", _FakeClient)
return key
def _sign(key, claims, alg="ES256", kid=KID):
return jwt.encode(claims, key, algorithm=alg, headers={"kid": kid, "typ": "EPT"})
def _claims(**over):
c = {
"sub": "ET26-TEST-0001",
"iss": ISSUER,
"iat": int(time.time()) - 10,
"exp": int(time.time()) + 3600,
}
c.update(over)
return c
# ---- the property that was broken ----------------------------------------
def test_genuine_ept_is_accepted(signing):
v = verify_ept(_sign(signing, _claims()), "https://api.eternitas.ai")
assert v.passport == "ET26-TEST-0001"
def test_passport_comes_from_sub_not_a_passport_claim(signing):
"""Real EPTs put the passport in `sub`. The pre-fix code read `passport` /
`sub_passport`, which no genuine EPT carries — so real agents were never
recognised and only forged tokens ever worked."""
tok = _sign(signing, _claims(sub="ET26-REAL-9999", passport="ET26-LIES-0000"))
assert verify_ept(tok, "https://api.eternitas.ai").passport == "ET26-REAL-9999"
def test_alg_none_is_refused(signing):
"""The exact 2026-08-13 exploit."""
import base64
import json as _j
def seg(d):
return base64.urlsafe_b64encode(_j.dumps(d).encode()).rstrip(b"=").decode()
forged = f"{seg({'alg':'none','typ':'EPT','kid':KID})}.{seg(_claims())}."
with pytest.raises(EptInvalid):
verify_ept(forged, "https://api.eternitas.ai")
def test_signature_from_a_different_key_is_refused(signing):
attacker = ec.generate_private_key(ec.SECP256R1())
with pytest.raises(EptInvalid):
verify_ept(_sign(attacker, _claims()), "https://api.eternitas.ai")
def test_tampered_payload_is_refused(signing):
tok = _sign(signing, _claims())
h, _p, s = tok.split(".")
import base64
import json as _j
evil = base64.urlsafe_b64encode(
_j.dumps(_claims(sub="ET26-EVIL-0000")).encode()
).rstrip(b"=").decode()
with pytest.raises(EptInvalid):
verify_ept(f"{h}.{evil}.{s}", "https://api.eternitas.ai")
def test_expired_token_is_refused(signing):
"""Genuinely signed, genuinely expired — proves exp is enforced rather than
the token merely failing to parse."""
tok = _sign(signing, _claims(exp=int(time.time()) - 5, iat=int(time.time()) - 100))
with pytest.raises(EptInvalid):
verify_ept(tok, "https://api.eternitas.ai")
def test_wrong_issuer_is_refused(signing):
"""Correctly signed by a trusted key but claiming another issuer."""
with pytest.raises(EptInvalid):
verify_ept(_sign(signing, _claims(iss="evil.example.com")), "https://api.eternitas.ai")
def test_unknown_kid_is_refused(signing):
with pytest.raises(EptInvalid):
verify_ept(_sign(signing, _claims(), kid="attacker-key"), "https://api.eternitas.ai")
def test_missing_required_claims_are_refused(signing):
for missing in ("sub", "exp"):
c = _claims()
c.pop(missing)
with pytest.raises(EptInvalid):
verify_ept(_sign(signing, c), "https://api.eternitas.ai")
def test_only_es256_is_ever_accepted():
"""Widening this list reopens algorithm confusion."""
assert ept_mod.ALGORITHMS == ["ES256"]
# ---- the throttle that used to be dead code ------------------------------
def test_band_multiplier_is_actually_consumed():
"""BAND_MULTIPLIER was defined and read by nothing before this."""
s = Settings()
assert limit_for(s, "repo.create", "platinum") == s.rate_repo_creates_per_day * 10
assert limit_for(s, "repo.create", "gold") == s.rate_repo_creates_per_day * 4
assert limit_for(s, "repo.create", "standard") == s.rate_repo_creates_per_day
assert limit_for(s, "repo.create", "watch") == s.rate_repo_creates_per_day // 2
def test_unknown_band_gets_standard_not_unlimited_and_not_zero():
s = Settings()
assert limit_for(s, "repo.create", "a-band-invented-tomorrow") == s.rate_repo_creates_per_day
assert limit_for(s, "repo.create", None) == s.rate_repo_creates_per_day
def test_untrusted_band_is_read_only():
assert BAND_MULTIPLIER["untrusted"] == 0
def test_every_throttled_action_has_a_configured_base():
s = Settings()
for action, field in ACTION_BASE.items():
assert getattr(s, field) > 0, f"{action} has no positive base rate"
# ---- revocation enforced on the live trust path (not just the webhook) ----
def test_revoked_passport_is_refused_by_trust_decision():
"""A revoked passport returns HTTP 200, status=revoked, band=unproven,
allowed=[] (verified live 2026-08-13). The decision must refuse it — the
old code returned band 'unproven' and seated the agent."""
from api.app.auth import PassportNotInGoodStanding, decide_trust
revoked = {"status": "revoked", "band": "unproven", "allowed_actions": []}
with pytest.raises(PassportNotInGoodStanding):
decide_trust(revoked, "ET26-NJQT-QMR0")
def test_active_passport_is_accepted_by_trust_decision():
from api.app.auth import decide_trust
active = {"status": "active", "band": "gold", "allowed_actions": ["read", "send"]}
band, actions = decide_trust(active, "ET26-1EF9-VJAN")
assert band == "gold" and actions == ("read", "send")
def test_unknown_or_missing_status_fails_closed():
from api.app.auth import PassportNotInGoodStanding, decide_trust
for body in ({"band": "gold"}, {"status": "suspended"}, {"status": "frozen"},
{"status": ""}, {}):
with pytest.raises(PassportNotInGoodStanding):
decide_trust(body, "ET26-X")
@pytest.mark.asyncio
async def test_resolve_passport_raises_on_revoked_status_wiring(monkeypatch):
"""Proves the WIRING, not just the decision: resolve_passport must feed the
real trust body through decide_trust and propagate the refusal. A validly
minted EPT can outlive the passport by ~a year, so this is the path that
stops a revoked-but-still-signed token."""
from api.app import auth
from api.app.auth import PassportNotInGoodStanding, resolve_passport
from api.app.config import Settings
class _Resp:
status_code = 200
def json(self):
return {"status": "revoked", "band": "unproven", "allowed_actions": []}
class _Client:
def __init__(self, *a, **k):
pass
async def __aenter__(self):
return self
async def __aexit__(self, *a):
return False
async def get(self, *a, **k):
return _Resp()
monkeypatch.setattr(auth.httpx, "AsyncClient", _Client)
settings = Settings(eternitas_platform_api_key="x", eternitas_base_url="https://api.eternitas.ai")
with pytest.raises(PassportNotInGoodStanding):
await resolve_passport(settings, "ET26-NJQT-QMR0")
@pytest.mark.asyncio
async def test_resolve_passport_returns_band_on_active_wiring(monkeypatch):
import httpx # noqa: F401
from api.app import auth
from api.app.auth import resolve_passport
from api.app.config import Settings
class _Resp:
status_code = 200
def json(self):
return {"status": "active", "band": "gold", "allowed_actions": ["read"]}
class _Client:
def __init__(self, *a, **k): pass
async def __aenter__(self): return self
async def __aexit__(self, *a): return False
async def get(self, *a, **k): return _Resp()
monkeypatch.setattr(auth.httpx, "AsyncClient", _Client)
settings = Settings(eternitas_platform_api_key="x")
band, actions = await resolve_passport(settings, "ET26-1EF9-VJAN")
assert band == "gold" and actions == ("read",)
def test_routing_does_not_depend_on_signature_wellformedness():
"""An EPT-shaped token must route to the EPT verifier even when its
signature is malformed. jwt.get_unverified_header() validates the whole
token and rejects bad signature padding, which used to push such tokens to
the human path — refused for the wrong reason, and readable as a human
identity via `sub` whenever require_verified_jwt was off."""
import base64
import json as _j
from api.app.ept import looks_like_ept
def seg(d):
return base64.urlsafe_b64encode(_j.dumps(d).encode()).rstrip(b"=").decode()
for hdr in ({"alg": "none", "typ": "EPT"}, {"alg": "ES256", "typ": "EPT"},
{"alg": "ES256"}):
tok = f"{seg(hdr)}.{seg({'sub': 'ET26-X'})}.x" # deliberately bad signature
assert looks_like_ept(tok), f"{hdr} did not route to the EPT verifier"
assert not looks_like_ept("not-a-token")
assert not looks_like_ept("")
def test_only_actions_that_route_through_this_api_are_claimed_enforced():
"""git push never touches this API, so a push limit here would count zero
forever and allow everything — a silent no-op wearing the costume of a
control. Push limits must stay in NOT_ENFORCED_HERE until a Gitea-side hook
reports pushes into agent_actions."""
from api.app.throttle import ACTION_BASE, NOT_ENFORCED_HERE
assert "push" not in ACTION_BASE
assert "push.force" not in ACTION_BASE
assert "push" in NOT_ENFORCED_HERE
assert not set(ACTION_BASE) & set(NOT_ENFORCED_HERE)
@pytest.mark.asyncio
async def test_enforce_refuses_an_action_it_cannot_actually_limit():
"""Asking to throttle 'push' must raise, not silently allow."""
from api.app.auth import ActorType, Caller
from api.app.config import Settings
from api.app.errors import RepairPointer
from api.app.throttle import enforce
caller = Caller(actor_type=ActorType.agent, passport="ET26-X", band="gold")
with pytest.raises(RepairPointer) as exc:
await enforce(None, Settings(), caller, "push")
assert exc.value.code == "throttle_unknown_action"

169
api/tests/test_hub_jwt.py Normal file
View File

@@ -0,0 +1,169 @@
"""G3.2 / I-8 — human tokens are verified, never read (SSO #14, 2026-09-23).
Behavioral: every case signs a real RS256 token with a locally generated key
and drives `get_caller`, so a green run means the gate refuses what it must —
not that some string appears in auth.py.
"""
from __future__ import annotations
import base64
import hashlib
import hmac
import json
import time
import jwt
import pytest
from cryptography.hazmat.primitives import serialization
from cryptography.hazmat.primitives.asymmetric import rsa
from api.app import hub_jwt
from api.app.config import Settings
from api.app.errors import RepairPointer
KEY = rsa.generate_private_key(public_exponent=65537, key_size=2048)
OTHER = rsa.generate_private_key(public_exponent=65537, key_size=2048)
IDENTITY = "5e1b9569-7f01-489d-bf14-6fe5a367fa3f"
def _claims(**over):
now = int(time.time())
c = {
"iss": "windy-identity",
"type": "human",
"sub": "row-id-not-identity",
"windy_identity_id": IDENTITY,
"email": "grant@example.com",
"iat": now,
"exp": now + 900,
}
c.update(over)
return {k: v for k, v in c.items() if v is not None}
def _sign(claims, key=KEY, alg="RS256"):
return jwt.encode(claims, key, algorithm=alg, headers={"kid": "test"})
class _Req:
def __init__(self, settings):
self.app = type("A", (), {"state": type("S", (), {"settings": settings})()})()
@pytest.fixture(autouse=True)
def _local_jwks(monkeypatch):
"""The hub's JWKS, served from KEY's public half — no network."""
class _Key:
key = KEY.public_key()
class _Client:
def get_signing_key_from_jwt(self, token):
return _Key()
monkeypatch.setattr(hub_jwt, "_client", lambda base_url: _Client())
async def _caller(token, **settings):
from api.app.auth import get_caller
s = Settings(environment="production", **settings)
return await get_caller(_Req(s), authorization=f"Bearer {token}", x_service_token=None)
async def _refused(token, **settings):
with pytest.raises(RepairPointer) as exc:
await _caller(token, **settings)
assert exc.value.status_code == 401 and exc.value.code == "token_invalid"
return exc.value
@pytest.mark.asyncio
async def test_genuine_hub_token_is_a_human_named_by_windy_identity_id():
c = await _caller(_sign(_claims()))
assert c.actor_type == "human"
assert c.identity_id == IDENTITY # NOT `sub`, which is the per-row user id
@pytest.mark.asyncio
async def test_forged_signature_is_refused():
await _refused(_sign(_claims(), key=OTHER))
@pytest.mark.asyncio
async def test_expired_token_is_refused():
await _refused(_sign(_claims(iat=int(time.time()) - 2000, exp=int(time.time()) - 60)))
@pytest.mark.asyncio
async def test_wrong_issuer_and_id_tokens_are_refused():
await _refused(_sign(_claims(iss="https://evil.example")))
# Contract v1: the discovery-URL issuer is legal for ACCESS tokens...
c = await _caller(_sign(_claims(iss="https://account.windyword.ai")))
assert c.identity_id == IDENTITY
# ...but an id_token minted for the forge (aud = Gitea's client id
# "windy-git", no type, sub = identity) must never act as a bearer here.
id_token = _claims(
iss="https://account.windyword.ai",
aud="windy-git",
type=None,
windy_identity_id=None,
sub=IDENTITY,
)
await _refused(_sign(id_token))
await _refused(_sign(dict(id_token, windy_identity_id=IDENTITY, type="human")))
@pytest.mark.asyncio
async def test_hs256_confusion_is_refused():
# The classic forgery: HMAC the token with the PUBLIC key as the secret.
pub = KEY.public_key().public_bytes(
serialization.Encoding.PEM, serialization.PublicFormat.SubjectPublicKeyInfo
)
header = base64.urlsafe_b64encode(json.dumps({"alg": "HS256", "typ": "JWT"}).encode()).rstrip(
b"="
)
body = base64.urlsafe_b64encode(json.dumps(_claims()).encode()).rstrip(b"=")
sig = base64.urlsafe_b64encode(
hmac.new(pub, header + b"." + body, hashlib.sha256).digest()
).rstrip(b"=")
await _refused((header + b"." + body + b"." + sig).decode())
@pytest.mark.asyncio
async def test_non_human_or_missing_type_is_refused():
await _refused(_sign(_claims(type="agent")))
await _refused(_sign(_claims(type=None)))
@pytest.mark.asyncio
async def test_missing_windy_identity_is_refused_not_read_from_sub():
await _refused(_sign(_claims(windy_identity_id=None)))
@pytest.mark.asyncio
async def test_aud_is_tolerated_when_it_names_windy_git_and_refused_otherwise():
"""PyJWT rejects ANY aud-bearing token when no audience is configured — the
trap that would break the day the hub starts emitting aud."""
c = await _caller(_sign(_claims(aud=["windy_chat", "windy_git", "windy_mail"])))
assert c.identity_id == IDENTITY
await _refused(_sign(_claims(aud=["windy_chat"])))
@pytest.mark.asyncio
async def test_require_aud_refuses_tokens_without_it():
await _refused(_sign(_claims()), hub_require_aud=True)
c = await _caller(_sign(_claims(aud=["windy_git"])), hub_require_aud=True)
assert c.identity_id == IDENTITY
@pytest.mark.asyncio
async def test_production_verifies_even_if_the_flag_is_off():
"""require_verified_jwt=False is a local-dev convenience; production must
never take the unverified path."""
await _refused(_sign(_claims(), key=OTHER), require_verified_jwt=False)
def test_algorithm_list_is_exactly_rs256():
assert hub_jwt.ALGORITHMS == ["RS256"]

View File

@@ -10,12 +10,16 @@ happened somewhere in this ecosystem and cost real time.
from __future__ import annotations
import base64 as _b64
import json as _json
import re
import subprocess
import sys
import types as _types
from pathlib import Path
import pytest
import pytest as _pytest
ROOT = Path(__file__).resolve().parents[2]
@@ -234,3 +238,517 @@ def test_g05_no_secret_literals_committed():
text = path.read_text(encoding="utf-8", errors="ignore")
for pattern in patterns:
assert not pattern.search(text), f"credential literal in {path}"
# --------------------------------------------------------------------------
# G2.1 / G2.7 — the Gitea version is PINNED, and drift is a failure
# --------------------------------------------------------------------------
def test_g21_gitea_version_is_pinned_not_latest():
compose = (ROOT / "docker-compose.yml").read_text()
m = re.search(r"image:\s*\S*gitea/gitea:(\S+)", compose)
assert m, "no gitea image pin found"
assert m.group(1) != "latest", "G2.1: pin an exact Gitea version, never `latest`"
assert re.match(r"^\d+\.\d+\.\d+$", m.group(1)), f"not an exact version: {m.group(1)}"
def test_g24_gitea_license_travels_with_us():
"""MIT's one obligation. Cheap to honour, embarrassing to miss."""
assert (ROOT / "LICENSES" / "gitea-MIT.txt").exists()
assert "MIT" in (ROOT / "LICENSES" / "gitea-MIT.txt").read_text()
# --------------------------------------------------------------------------
# G4.3 — the two Gitea storage traps that cost a crash loop each
# --------------------------------------------------------------------------
def test_g43_no_lfs_storage_type_override():
"""Naming a storage type inside [lfs] creates a SEPARATE storage section
that does not inherit endpoint or credentials from [storage], and Gitea
crash-loops with an error that names the symptom and not the cause."""
# Check real settings only — the compose file deliberately NAMES this key in
# a warning comment so the next person does not re-add it.
active = [
ln for ln in (ROOT / "docker-compose.yml").read_text().splitlines()
if ln.strip() and not ln.strip().startswith("#")
]
assert not any("GITEA__lfs__STORAGE_TYPE" in ln for ln in active)
def test_g43_lfs_server_is_actually_enabled():
"""Setting the storage backend does NOT turn LFS on. Without this the batch
endpoint 404s and the client says 'Repository or object not found', which
reads like a permissions problem and is not one."""
active = [
ln for ln in (ROOT / "docker-compose.yml").read_text().splitlines()
if ln.strip() and not ln.strip().startswith("#")
]
assert any("GITEA__server__LFS_START_SERVER" in ln for ln in active)
def test_g43_r2_checksum_trap_is_pinned():
"""R2 rejects the checksum algorithm S3 clients send by default."""
compose = (ROOT / "docker-compose.yml").read_text()
assert "MINIO_CHECKSUM_ALGORITHM" in compose
# --------------------------------------------------------------------------
# G3.6 — the trust status-code law. This is the one with a live sibling bypass.
# --------------------------------------------------------------------------
def test_g36_trust_client_never_soft_allows():
"""A sibling maps HTTP 400 and 429 to 'unreachable' and then soft-ALLOWS.
That is inducible: an attacker who wants the check skipped only has to make
the check rate-limit itself at 100/min/IP."""
src = (ROOT / "api" / "app" / "auth.py").read_text()
assert "raise passport_unresolvable" in src
# There must be no return path out of resolve_passport other than a verified
# 200 or a raise.
body = src[src.index("async def resolve_passport") : src.index("async def get_caller")]
returns = [ln for ln in body.splitlines() if ln.strip().startswith("return")]
assert len(returns) == 1, f"resolve_passport has {len(returns)} return paths; expected exactly 1"
def test_g36_unverified_human_jwt_is_refused_in_production():
"""I-8 applied to ourselves: an unverified JWT is an authentication bypass,
not a shortcut. G3.2's verifier now exists; behavioral proof that forged,
expired, mis-issued and mis-audienced tokens are refused lives in
test_hub_jwt.py. Here: the gate defaults closed."""
from api.app.config import Settings
assert Settings().require_verified_jwt is True
assert "windy-git" not in Settings().hub_audiences # Gitea's client_id: id_token confusion
def test_no_auth_bypass_env_var_anywhere():
"""The desktop control server is the best Principle-#5 artifact in the
ecosystem partly because it has NO bypass env var. Copied on purpose."""
src = (ROOT / "api" / "app" / "auth.py").read_text()
for banned in ("SKIP_AUTH", "DISABLE_AUTH", "ALLOW_INSECURE", "AUTH_BYPASS", "DEV_MODE"):
assert banned not in src
# --------------------------------------------------------------------------
# G5.3 — the shelter's grant model
# --------------------------------------------------------------------------
def test_g53_grant_requires_exactly_one_grantee_in_the_database():
"""Enforced by a CHECK constraint, not by application code. This ecosystem
already has a core invariant enforced only in app code across two files."""
src = (ROOT / "alembic" / "versions" / "001_genesis.py").read_text()
assert "ck_grant_exactly_one_grantee" in src
assert "(grantee_identity_id IS NULL) <> (grantee_passport IS NULL)" in src
def test_g53_agent_grants_expire_by_default():
from api.app.config import Settings
assert Settings().agent_grant_default_days == 90
def test_g55_shelter_strings_avoid_developer_vocabulary():
"""D-9/I-9: a person restoring last Tuesday's work should not have to learn
a vocabulary first. Check the strings users actually see."""
import re as _re
src = (ROOT / "api" / "app" / "routes" / "repos.py").read_text()
speaks = _re.findall(r'"speak":\s*\(?\s*\n?\s*f?"([^"]+)"', src)
assert speaks, "no speak strings found to audit"
for s in speaks:
low = s.lower()
for jargon in ("commit", "repository", "branch", "sha", "push"):
assert jargon not in low, f"developer vocabulary in a user string: {s!r}"
# --------------------------------------------------------------------------
# I-4 — never a one-way door
# --------------------------------------------------------------------------
def test_i04_mirror_syncs_on_every_save_not_just_a_timer():
"""An hourly window means an hour of work can be the thing you lose, and the
window is invisible until it costs you."""
src = (ROOT / "api" / "app" / "services" / "mirror.py").read_text()
assert '"sync_on_commit": True' in src
def test_i04_unconfigured_mirror_is_never_reported_healthy():
"""A mirror nobody checks is a belief, not a backup. An unconfigured one
reports 'unconfigured' — never 'healthy'."""
src = (ROOT / "api" / "app" / "services" / "mirror.py").read_text()
assert '"state": "unconfigured"' in src
assert "if not self.configured:" in src
def test_i04_mirror_lag_threshold_is_set():
from api.app.config import Settings
assert Settings().mirror_lag_p2_seconds == 3600
def test_owner_namespace_is_derived_from_the_repo_not_the_caller():
"""Deriving the namespace from the caller is right only while the caller is
the owner, and addresses the wrong namespace the moment a collaborator asks
— surfacing as 'not found', which is the hardest kind of bug to see."""
src = (ROOT / "api" / "app" / "routes" / "repos.py").read_text()
body = src[src.index("async def list_versions") : src.index("async def create_grant")]
assert "_repo_owner_login(repo)" in body
assert "_owner_login(caller)" not in body
def test_i04_never_synced_is_not_reported_as_merely_behind():
"""Collapsing 'never ran' into 'behind' is how a backup that was never made
gets read as a backup that is merely stale. Gitea reports the epoch for
'not yet', which arithmetic turns into a 56-year lag and a confident
'degraded'."""
src = (ROOT / "api" / "app" / "services" / "mirror.py").read_text()
assert "never_synced" in src
assert '"pending"' in src
# --------------------------------------------------------------------------
# G7 / I-5 — CI never shares a kernel with identity
# --------------------------------------------------------------------------
def test_i05_runner_never_mounts_the_host_docker_socket():
"""The tempting move — and what every published act_runner example does —
is to mount /var/run/docker.sock. That hands every workflow, including a
transitive dependency's postinstall script, the ability to start a
privileged container mounting / — i.e. root on the host."""
compose = (ROOT / "deploy" / "runner" / "docker-compose.yml").read_text()
active = [ln for ln in compose.splitlines() if ln.strip() and not ln.strip().startswith("#")]
for ln in active:
assert "/var/run/docker.sock" not in ln, "I-5: never mount the host docker socket"
def test_i05_jobs_get_a_network_per_job_not_a_shared_bridge():
"""A shared flat bridge lets concurrent jobs see each other, and breaks
service-container DNS (service aliases only exist on a per-job network).
Per-job is both stricter and correct."""
cfg = (ROOT / "deploy" / "runner" / "config.yaml").read_text()
active = [ln for ln in cfg.splitlines() if ln.strip() and not ln.strip().startswith("#")]
net = [ln for ln in active if ln.strip().startswith("network:")]
assert net and net[0].strip() == 'network: ""', "jobs must get a per-job network"
def test_i05_jobs_cannot_bind_mount_from_the_daemon_host():
cfg = (ROOT / "deploy" / "runner" / "config.yaml").read_text()
assert "valid_volumes: []" in cfg
assert 'docker_host: "-"' in cfg
def test_i05_no_ci_container_can_reach_the_forge_network():
"""The first CI run failed with "Could not resolve host: gitea" because job
containers sit on dind's private network. The easy fix — putting jobs on the
forge network — would have left untrusted workflow code one DNS name from
the forge's Postgres. Instead jobs reach the PUBLIC forge surface, so no CI
container has a private route to anything."""
compose = (ROOT / "deploy" / "runner" / "docker-compose.yml").read_text()
active = [ln for ln in compose.splitlines() if ln.strip() and not ln.strip().startswith("#")]
joined = "\n".join(active)
assert "windy-git_default" not in joined, "I-5: no CI container joins the forge network"
assert "https://app.windygit.com" in joined
def test_i05_runner_is_a_separate_compose_project_from_the_forge():
"""Runners restart, crash, get starved and get killed. None of that should
ever touch the thing serving repositories."""
runner = (ROOT / "deploy" / "runner" / "docker-compose.yml").read_text()
forge = (ROOT / "docker-compose.yml").read_text()
assert "name: windy-git-runner" in runner
assert "name: windy-git" in forge
def test_g15_runner_is_cpu_and_memory_bounded():
"""Veron 1 is Grant's workstation, not a dedicated build box."""
compose = (ROOT / "deploy" / "runner" / "docker-compose.yml").read_text()
assert "cpus:" in compose
assert "mem_limit:" in compose
def test_g75_workflows_use_a_label_this_runner_actually_provides():
"""A workflow naming a label nobody provides queues forever and presents as
a hung CI system rather than a typo."""
cfg = (ROOT / "deploy" / "runner" / "config.yaml").read_text()
provided = {
ln.split(":")[0].strip().strip('"- ')
for ln in cfg.splitlines()
if "docker://" in ln
}
assert provided, "runner declares no labels"
for wf in ROOT.rglob(".gitea/workflows/*.y*ml"):
for ln in wf.read_text().splitlines():
# Skip comments — a doc line explaining runs-on is not a runs-on.
if ln.strip().startswith("#") or "runs-on:" not in ln:
continue
label = ln.split("runs-on:")[1].strip()
assert label in provided, f"{wf.name}: '{label}' is not a provided label"
def test_g73_workflow_pins_a_python_that_satisfies_requires_python():
"""The first real CI run wedged for 14 minutes because the runner image
ships Python 3.10 and this project requires 3.12: pip answered by
backtracking through every historical version of every dependency, at full
CPU, silently. A version mismatch presenting as a hang rather than an
error."""
import re as _re
pyproject = (ROOT / "pyproject.toml").read_text()
m = _re.search(r'requires-python\s*=\s*">=(\d+)\.(\d+)"', pyproject)
assert m, "pyproject declares no requires-python"
major, minor = int(m.group(1)), int(m.group(2))
for wf in ROOT.rglob(".gitea/workflows/*.y*ml"):
text = wf.read_text()
# Only workflows that actually RUN Python need to pin it. A workflow
# that shells out to psql or curl does not, and demanding a pin from
# it is the test being wrong rather than the workflow.
if not _re.search(r"\bpython3?\b|pytest|pip ", text):
continue
# Either a pinned container image or an explicit setup-python version.
pin = _re.search(r"image:\s*python:(\d+)\.(\d+)", text) or _re.search(
r'python-version:\s*"?(\d+)\.(\d+)"?', text
)
assert pin, f"{wf.name}: job neither pins a python image nor sets up a version"
assert (int(pin.group(1)), int(pin.group(2))) >= (major, minor), (
f"{wf.name}: pins python {pin.group(0)} but the project requires "
f">={major}.{minor}"
)
def test_g73_every_job_has_a_timeout():
"""A wedged step should be a red check in minutes, not an occupied runner
for half an hour."""
for wf in ROOT.rglob(".gitea/workflows/*.y*ml"):
assert "timeout-minutes:" in wf.read_text(), f"{wf.name}: no job timeout"
# --------------------------------------------------------------------------
# G7.6 — the canary must watch what users do, and must not live on Kit 0
# --------------------------------------------------------------------------
def test_g76_canary_probes_login_not_just_health():
"""/health returned 200 for the entire 2026-08-12 outage while login was
dead. A canary that only watches health is decorative."""
src = (ROOT / "scripts" / "canary.py").read_text()
assert "identity.login" in src
assert "/api/v1/auth/login" in src
def test_g76_canary_does_not_run_on_kit_zero():
"""A canary hosted on the box it watches dies with that box, and reports
nothing at the exact moment it matters."""
wf = (ROOT / ".gitea" / "workflows" / "canary.yml").read_text()
assert "runs-on: veron-1" in wf
assert "72.60.118.54" not in wf
def test_g76_canary_alerts_on_transition_not_every_run():
"""A canary that emails every 10 minutes gets filtered, and a filtered
canary is a dead canary."""
src = (ROOT / "scripts" / "canary.py").read_text()
assert "newly_bad" in src and "recovered" in src
def test_g76_canary_has_two_independent_signals():
"""Email AND a red CI run. The last fleet canary died silently because it
had one signal and nothing watched the watcher."""
src = (ROOT / "scripts" / "canary.py").read_text()
assert "return 1 if any" in src, "canary must exit non-zero so CI goes red"
def test_g76_alert_path_sets_a_user_agent():
"""Without an explicit User-Agent, urllib sends 'Python-urllib/3.x' and
Resend rejects it 403 while the identical curl succeeds. Caught by testing
the alert path: the canary would have detected every outage correctly and
told nobody."""
src = (ROOT / "scripts" / "canary.py").read_text()
send = src[src.index("def send_alert") : src.index("def main(")]
assert "User-Agent" in send
# --------------------------------------------------------------------------
# G3.5 — revocation is fail-closed, and the two signature traps
# --------------------------------------------------------------------------
def test_g35_webhook_strips_the_sha256_prefix():
"""A sibling receiver compared the whole 'sha256=<hex>' header against a
bare digest and returned 401 forever — wired, never once delivered."""
src = (ROOT / "api" / "app" / "routes" / "webhooks.py").read_text()
assert 'startswith("sha256=")' in src
def test_g35_webhook_hmacs_raw_bytes_not_reserialised_json():
"""JSON.stringify of a parsed body reorders keys and changes whitespace, so
the digest never matches what the sender signed. Same silent 401."""
src = (ROOT / "api" / "app" / "routes" / "webhooks.py").read_text()
verify = src[src.index("def _verify") : src.index("@router.post")]
assert "raw" in verify and "json.dumps" not in verify
def test_g35_unset_secret_refuses_rather_than_accepts():
src = (ROOT / "api" / "app" / "routes" / "webhooks.py").read_text()
assert "webhook_secret_unset" in src
assert "refusing unverified webhooks" in src
def test_g35_signature_compare_is_constant_time():
src = (ROOT / "api" / "app" / "routes" / "webhooks.py").read_text()
assert "hmac.compare_digest" in src
def test_g35_revocation_never_acknowledges_what_it_did_not_apply():
"""A 200 on a revocation the receiver could not apply is a security hole
that reports success."""
src = (ROOT / "api" / "app" / "routes" / "webhooks.py").read_text()
assert "refusing to acknowledge a revocation we did not apply" in src
def test_g35_probe_acknowledgement_changes_nothing():
"""Eternitas verifies a webhook URL answers BEFORE issuing the secret that
signs deliveries, so the first request can never be signed. The probe path
answers 200 but must never act, and anything claiming to be an event must
still be verified."""
src = (ROOT / "api" / "app" / "routes" / "webhooks.py").read_text()
probe = src[src.index('if x_eternitas_event == "platform.test_ping"') : src.index("secret = settings")]
assert '"acted": False' in probe
assert "update(" not in probe and "commit" not in probe
def test_g35_did_not_disable_validation_to_register():
"""skip_validation would permanently disable a safety check to solve a
one-time ordering problem."""
for f in (ROOT / "scripts").glob("*.py"):
assert "skip_validation" not in f.read_text()
def test_g76_canary_survives_an_unwritable_state_path():
"""A monitoring tool that dies of a config problem reports nothing at all,
and reports it silently. State is an optimisation; probing is the point."""
src = (ROOT / "scripts" / "canary.py").read_text()
save = src[src.index("def save_state") : src.index("def send_alert")]
assert "except OSError" in save
# --------------------------------------------------------------------------
# G0.9 — backups, the prerequisite for being the daily driver
# --------------------------------------------------------------------------
def test_g09_backup_bundles_all_refs_not_just_the_default_branch():
"""A single-branch bundle loses every other branch and every tag silently,
and you find out during the restore."""
src = (ROOT / "scripts" / "backup.sh").read_text()
assert "bundle create" in src and "--all" in src
def test_g09_backup_verifies_before_trusting():
"""An unverified bundle is a belief, not a backup."""
src = (ROOT / "scripts" / "backup.sh").read_text()
assert "git bundle verify" in src
def test_g09_backup_fails_loudly():
"""A backup script that swallows errors is worse than none — it
manufactures confidence."""
src = (ROOT / "scripts" / "backup.sh").read_text()
assert "COMPLETED WITH FAILURES" in src
assert "refusing to report a backup that did not happen" in src
# --------------------------------------------------------------------------
# SECURITY (behavioral, not string-grep): the agent path must not authenticate
# an unverified token. Regression guard for the 2026-08-13 forged-token bypass.
# --------------------------------------------------------------------------
def _forged_bearer(passport: str) -> str:
def seg(d):
return _b64.urlsafe_b64encode(_json.dumps(d).encode()).rstrip(b"=").decode()
return f"{seg({'alg':'none','typ':'JWT'})}.{seg({'passport':passport})}.not-a-signature"
def _fake_request(settings):
app = _types.SimpleNamespace(state=_types.SimpleNamespace(settings=settings))
return _types.SimpleNamespace(app=app)
@_pytest.mark.asyncio
async def test_security_forged_agent_token_is_refused_in_production():
"""The 2026-08-13 exploit: an alg:none token naming a real passport returned
HTTP 200 as that agent. It must now be refused whichever gate catches it —
an EPT-shaped forgery by signature verification, a JWT-shaped one by the
human gate. What is asserted is REFUSAL, not a particular error code."""
from api.app.auth import get_caller
from api.app.config import Settings
from api.app.errors import RepairPointer
settings = Settings(environment="production", require_verified_jwt=True,
eternitas_platform_api_key="x")
req = _fake_request(settings)
# Both shapes now route to the EPT verifier, because alg:none is never
# valid for ANY caller — so 401 "your token is bad" is the honest answer,
# not 503 "that feature isn't ready".
for typ, expected in (("JWT", "ept_invalid"), ("EPT", "ept_invalid")):
def seg(d):
return _b64.urlsafe_b64encode(_json.dumps(d).encode()).rstrip(b"=").decode()
forged = (f"{seg({'alg':'none','typ':typ})}"
f".{seg({'passport':'ET26-1EF9-VJAN','sub':'ET26-1EF9-VJAN'})}.sig")
with _pytest.raises(RepairPointer) as exc:
await get_caller(req, authorization=f"Bearer {forged}", x_service_token=None)
assert exc.value.status_code in (401, 403, 503), f"{typ} was not refused"
assert exc.value.code == expected, f"{typ} -> {exc.value.code}"
@_pytest.mark.asyncio
async def test_security_no_bearer_is_still_401():
from api.app.auth import get_caller
from api.app.config import Settings
from api.app.errors import RepairPointer
req = _fake_request(Settings(environment="production"))
with _pytest.raises(RepairPointer) as exc:
await get_caller(req, authorization=None, x_service_token=None)
assert exc.value.status_code == 401
def test_i12_build_fails_when_commit_sha_is_empty():
"""The sed+grep pair silently accepted an empty COMMIT_SHA: it replaced ""
with "" and then matched that same empty string, shipping a container that
reported commit_sha: null. That is the exact defect I-12 exists to prevent,
and it happened on 2026-08-14."""
df = (ROOT / "Dockerfile").read_text()
assert 'test -n "${COMMIT_SHA}"' in df
# --------------------------------------------------------------------------
# G2.3 — branding lives in the repo, not only on one host's disk
# --------------------------------------------------------------------------
def test_g23_branding_is_version_controlled():
"""It was applied directly to Veron's disk first, which is the config-drift
trap this project documents: the running system and the repo disagree, and
a rebuild silently reverts to stock Gitea."""
b = ROOT / "deploy" / "branding"
for f in ("apply.sh", "README.md", "templates/home.tmpl",
"templates/custom/header.tmpl"):
assert (b / f).exists(), f"missing {f}"
def test_g23_brand_css_filename_is_versioned():
"""Cloudflare caches /assets/* for 6h and no token here can purge, so a
fixed filename leaves stale bytes live for hours."""
import re as _re
hdr = (ROOT / "deploy" / "branding" / "templates" / "custom" / "header.tmpl").read_text()
m = _re.search(r"theme-windy\.v(\d+)\.css", hdr)
assert m, "brand CSS must carry a version in its FILENAME"
assert (ROOT / "deploy" / "branding" / "public" / "assets" / "css"
/ f"theme-windy.v{m.group(1)}.css").exists()
def test_backup_never_bundles_credential_repos_to_r2():
"""kit-army-config (the lockbox) and the *-soul / anima repos carry
credentials; the R2 bundles are plaintext. Behavioural: run the script's
own exclusion function against the names."""
import subprocess
script = (ROOT / "scripts" / "backup.sh").read_text()
fn = script[script.index('EXCLUDE="'):script.index("cleanup()")]
# A file named like a pattern in cwd must not break the match (glob expansion).
probe = "cd \"$(mktemp -d)\" && touch x-soul && " + fn + (
'for n in kit-army-config anima windy-0-soul kit-0c5-soul herm-0-soul '
'soulsafe windy-chat eternitas; do excluded "$n" && echo "X $n" || echo "- $n"; done'
)
out = subprocess.run(["bash", "-c", probe], capture_output=True, text=True, check=True).stdout
skipped = {ln[2:] for ln in out.splitlines() if ln.startswith("X ")}
assert skipped == {"kit-army-config", "anima", "windy-0-soul", "kit-0c5-soul", "herm-0-soul"}

View File

@@ -0,0 +1,274 @@
"""Behavioral tests for scripts/pr_status_bridge.py.
The bridge is the ONLY CI signal the private platform repos get on GitHub, so
these drive its real functions against fake Gitea/GitHub APIs rather than
grepping its source: a status painted green that nobody tested is worse than
no status at all.
"""
from __future__ import annotations
import base64
import importlib.util
from pathlib import Path
import pytest
ROOT = Path(__file__).resolve().parents[2]
_spec = importlib.util.spec_from_file_location(
"pr_status_bridge", ROOT / "scripts" / "pr_status_bridge.py"
)
bridge = importlib.util.module_from_spec(_spec)
_spec.loader.exec_module(bridge)
SHA = "a" * 40
def _run(i, wf, job, status, sha=SHA, n=1):
return {
"id": i,
"workflow_id": wf,
"name": job,
"status": status,
"head_sha": sha,
"run_number": n,
}
class Fake:
def __init__(self, runs=(), statuses=(), gh_prs=(), wg_prs=(), workflows=None):
self.runs, self.statuses = list(runs), list(statuses)
self.workflows = workflows or {} # {path: yaml text} at every commit
self.gh_prs, self.wg_prs = list(gh_prs), list(wg_prs)
self.posted, self.opened, self.closed = [], [], []
def gitea(self, method, path, body=None):
if "/contents/" in path:
want = path.split("/contents/", 1)[1].split("?", 1)[0]
if want in self.workflows:
return 200, {"content": base64.b64encode(self.workflows[want].encode()).decode()}
files = [
{"type": "file", "name": k.rsplit("/", 1)[1], "path": k}
for k in self.workflows
if k.rsplit("/", 1)[0] == want
]
return (200, files) if files else (404, None)
if "/actions/tasks" in path:
page = int(path.rsplit("page=", 1)[1])
return 200, {"workflow_runs": self.runs[(page - 1) * 50 : page * 50]}
if method == "GET" and path.endswith("/pulls?state=open&limit=50"):
return 200, self.wg_prs
if method == "POST" and path.endswith("/pulls"):
self.opened.append(body)
return 201, {}
if method == "PATCH":
self.closed.append(path)
return 201, {}
raise AssertionError(path)
def github(self, method, path, body=None):
if "/statuses" in path and method == "GET":
return 200, self.statuses
if "/statuses/" in path and method == "POST":
self.posted.append(body)
return 201, {}
if "/pulls?" in path:
return 200, self.gh_prs
raise AssertionError(path)
@pytest.fixture
def fake(monkeypatch):
def make(**kw):
f = Fake(**kw)
monkeypatch.setattr(bridge, "gitea", f.gitea)
monkeypatch.setattr(bridge, "github", f.github)
return f
return make
def test_posts_latest_verdict_per_job(fake):
f = fake(runs=[_run(1, "ci.yml", "test", "failure"), _run(2, "ci.yml", "test", "success", n=2)])
bridge.post_statuses("r", SHA)
assert [(p["context"], p["state"]) for p in f.posted] == [("windy-git/ci/test", "success")]
assert f.posted[0]["target_url"].endswith("/actions/runs/2")
def test_unchanged_state_is_not_reposted(fake):
f = fake(
runs=[_run(1, "ci.yml", "test", "success")],
statuses=[{"context": "windy-git/ci/test", "state": "success"}],
)
bridge.post_statuses("r", SHA)
assert f.posted == []
def test_skipped_job_is_never_painted_green(fake):
f = fake(runs=[_run(1, "substrate-drift.yml", "check", "skipped")])
bridge.post_statuses("r", SHA)
assert f.posted == []
def test_other_commits_runs_are_ignored(fake):
f = fake(runs=[_run(1, "ci.yml", "test", "failure", sha="b" * 40)])
bridge.post_statuses("r", SHA)
assert f.posted == []
def test_runs_past_the_first_page_are_seen(fake):
noise = [_run(100 + i, "drift.yml", "x", "skipped", sha="c" * 40) for i in range(50)]
f = fake(runs=noise + [_run(1, "ci.yml", "test", "success")])
bridge.post_statuses("r", SHA)
assert [p["state"] for p in f.posted] == ["success"]
def _gh_pr(n, repo="sneakyfree/r"):
return {
"number": n,
"title": "t",
"html_url": "u",
"head": {"ref": f"b{n}", "sha": SHA, "repo": {"full_name": repo} if repo else None},
"base": {"ref": "main"},
}
def test_fork_prs_are_never_mirrored(fake, monkeypatch):
monkeypatch.setattr(bridge, "GH_OWNER", "sneakyfree")
f = fake(gh_prs=[_gh_pr(1, repo="stranger/r"), _gh_pr(2, repo=None)])
assert bridge.sync_prs("r") == []
assert f.opened == []
def test_pr_mirror_opened_once_and_closed_when_github_closes(fake, monkeypatch):
monkeypatch.setattr(bridge, "GH_OWNER", "sneakyfree")
f = fake(gh_prs=[_gh_pr(7)], wg_prs=[{"number": 3, "title": "[GH#5] gone"}])
assert bridge.sync_prs("r") == [SHA]
assert [o["head"] for o in f.opened] == ["b7"]
assert f.closed == ["/repos/windyadmin/r/pulls/3"]
f2 = fake(gh_prs=[_gh_pr(7)], wg_prs=[{"number": 4, "title": "[GH#7] t"}])
bridge.sync_prs("r")
assert f2.opened == [] and f2.closed == []
def test_image_build_jobs_are_not_posted(fake):
"""No Docker daemon in job containers (I-5): a build job's red is structural."""
f = fake(
runs=[_run(1, "ci.yml", "Docker Build", "failure"), _run(2, "ci.yml", "docker", "failure")]
)
bridge.post_statuses("r", SHA)
assert f.posted == []
def test_non_blocking_jobs_are_not_posted_for_that_repo_only(fake, monkeypatch):
"""Grant ruled windy-pro's desktop/installer jobs non-blocking: they must not
reach GitHub for windy-pro, and the rule must not leak to other repos."""
monkeypatch.setattr(bridge, "NON_BLOCKING", {"windy-pro": {"ci/build-desktop"}})
runs = [_run(1, "ci.yml", "build-desktop", "failure"), _run(2, "ci.yml", "test", "success")]
f = fake(runs=runs)
bridge.post_statuses("windy-pro", SHA)
assert [p["context"] for p in f.posted] == ["windy-git/ci/test"]
f2 = fake(runs=runs)
bridge.post_statuses("windy-chat", SHA)
assert sorted(p["context"] for p in f2.posted) == [
"windy-git/ci/build-desktop",
"windy-git/ci/test",
]
def test_default_non_blocking_is_grants_ruling():
assert bridge.NON_BLOCKING.get("windy-pro") == {
"ci/build-desktop",
"ci/test-installer",
"ci/reality-check",
}
def test_transport_blips_are_retried_but_http_errors_are_not(monkeypatch):
import urllib.error
calls = {"n": 0}
class _R:
status = 200
def read(self):
return b"{}"
def __enter__(self):
return self
def __exit__(self, *a):
return False
def flaky(req, timeout):
calls["n"] += 1
if calls["n"] < 3:
raise urllib.error.URLError("_ssl.c:983: The handshake operation timed out")
return _R()
monkeypatch.setattr(bridge.urllib.request, "urlopen", flaky)
monkeypatch.setattr(bridge.time, "sleep", lambda s: None)
assert bridge._call("http://x", "t", "GET", "/p") == (200, {})
assert calls["n"] == 3
def forbidden(req, timeout):
calls["n"] += 1
raise urllib.error.HTTPError("http://x/p", 403, "no", {}, None)
calls["n"] = 0
monkeypatch.setattr(bridge.urllib.request, "urlopen", forbidden)
assert bridge._call("http://x", "t", "GET", "/p") == (403, None)
assert calls["n"] == 1
GOOD = "on: push\njobs:\n test:\n runs-on: ubuntu-latest\n steps: []\n"
BROKEN = "on: push\njobs:\n test:\n runs-on: x\n steps: [\n"
def test_invalid_workflow_gets_an_error_status_even_with_no_runs(fake):
# Gitea fires NO run for an invalid file: without this the PR shows nothing.
f = fake(workflows={".github/workflows/ci.yml": BROKEN})
bridge.post_statuses("windy-chat", SHA)
assert [(p["context"], p["state"]) for p in f.posted] == [("windy-git/ci/workflow", "error")]
assert "invalid YAML at line 5" in f.posted[0]["description"]
assert f.posted[0]["target_url"].endswith(f"/src/commit/{SHA}/.github/workflows/ci.yml")
def test_valid_workflows_post_nothing_extra(fake):
f = fake(runs=[_run(1, "ci.yml", "test", "success")], workflows={".github/workflows/ci.yml": GOOD})
bridge.post_statuses("windy-chat", SHA)
assert [p["context"] for p in f.posted] == ["windy-git/ci/test"]
def test_workflow_error_is_not_reposted(fake):
f = fake(
workflows={".github/workflows/ci.yml": BROKEN},
statuses=[{"context": "windy-git/ci/workflow", "state": "error"}],
)
bridge.post_statuses("windy-chat", SHA)
assert f.posted == []
def test_gitea_dir_wins_over_github_dir(fake):
# Gitea runs .gitea/workflows when it has files and ignores .github/workflows.
f = fake(workflows={".gitea/workflows/ci.yml": GOOD, ".github/workflows/old.yml": BROKEN})
bridge.post_statuses("windy-chat", SHA)
assert f.posted == []
@pytest.mark.parametrize(
"text, problem",
[
(GOOD, None),
("on: push\njobs:\n a:\n uses: ./x.yml\n", None),
(BROKEN, "invalid YAML at line 5"),
("jobs:\n a:\n runs-on: x\n", "no `on:` trigger"),
("on: push\n", "no `jobs:`"),
("on: push\njobs:\n a:\n steps: []\n", "job `a` has no `runs-on:`"),
("- a\n", "not a YAML mapping"),
],
)
def test_workflow_problem(text, problem):
assert bridge.workflow_problem(text) == problem

View File

@@ -0,0 +1,82 @@
"""Push-velocity detection (scripts/telemetry_emit.py): detect + alert only.
Driven through the real function with rows shaped like the Gitea query's.
"""
from __future__ import annotations
import importlib.util
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "scripts"))
_spec = importlib.util.spec_from_file_location("telemetry_emit", ROOT / "scripts" / "telemetry_emit.py")
te = importlib.util.module_from_spec(_spec)
_spec.loader.exec_module(te)
NOW = 1_800_000_000.0
def row(login="agent-et26abcd1234", uid=7, p1h=0, p24h=0, d24h=0, repos=1, wid=None):
return {"uid": uid, "login": login, "wid": wid, "p1h": p1h, "p24h": p24h, "d24h": d24h, "repos": repos}
def test_under_every_threshold_emits_nothing():
ev, keep = te.push_velocity_events([row(p1h=60, p24h=500, d24h=10)], NOW, {})
assert ev == [] and keep == {}
def test_burst_emits_one_declared_row_with_the_passport():
ev, keep = te.push_velocity_events([row(p1h=61, p24h=61, repos=3)], NOW, {})
assert len(ev) == 1
e = ev[0]
assert e["event_type"] == "forge.push_velocity" and e["service"] == "forge"
assert e["actor_type"] == "agent" and e["actor_id"] == "ET26-ABCD-1234"
assert e["metadata"] == {
"rule": "pushes_1h", "window_s": 3600, "count": 61, "threshold": 60,
"repos": 3, "gitea_user_id": 7,
}
assert keep == {"7:pushes_1h": NOW}
def test_still_over_is_reported_once_per_window_not_every_run():
_, keep = te.push_velocity_events([row(p1h=90)], NOW, {})
ev, keep = te.push_velocity_events([row(p1h=95)], NOW + 300, keep)
assert ev == [] and keep == {"7:pushes_1h": NOW}
ev, _ = te.push_velocity_events([row(p1h=95)], NOW + 3601, keep)
assert len(ev) == 1
def test_dropping_back_under_rearms():
_, keep = te.push_velocity_events([row(p1h=90)], NOW, {})
_, keep = te.push_velocity_events([row(p1h=5)], NOW + 300, keep)
assert keep == {}
ev, _ = te.push_velocity_events([row(p1h=90)], NOW + 600, keep)
assert len(ev) == 1
def test_the_sync_account_is_exempt():
ev, _ = te.push_velocity_events([row(login="windyadmin", uid=1, p1h=9999, p24h=9999)], NOW, {})
assert ev == []
def test_sso_human_is_keyed_on_windy_identity_id():
ev, _ = te.push_velocity_events([row(login="u-5e1b9569abc", wid="5e1b9569-full-id", d24h=11)], NOW, {})
assert [(e["actor_type"], e["actor_id"], e["metadata"]["rule"]) for e in ev] == [
("human", "5e1b9569-full-id", "ref_deletes_24h")
]
assert "caller" not in ev[0]["metadata"]
def test_no_provable_id_is_system_plus_caller_never_an_invented_id():
# UPDATE 2 actor rule: agent/human rows without an actor_id are quarantined.
for login in ("u-nolink", "agent-weird"):
ev, _ = te.push_velocity_events([row(login=login, p24h=501)], NOW, {})
assert ev[0]["actor_type"] == "system" and "actor_id" not in ev[0]
assert ev[0]["metadata"]["caller"] == "unknown"
def test_passport_round_trip():
assert te.passport_from_login("agent-et26p1zgttp8") == "ET26-P1ZG-TTP8"
assert te.passport_from_login("u-abc") is None

202
api/tests/test_telemetry.py Normal file
View File

@@ -0,0 +1,202 @@
"""Field telemetry from the API (step 2): behavioural, no network.
A refusal must become exactly one forge.auth.failed row in the DECLARED shape
(the ledger quarantines anything else); non-refusal errors must not; the
heartbeat must count what happened and never invent a p95 for no traffic.
"""
from __future__ import annotations
import httpx
import pytest
from fastapi import Depends, FastAPI
from api.app import telemetry as tmod
from api.app.errors import RepairPointer
DECLARED_AUTH_KEYS = {"code", "http_status", "caller", "route", "upstream_status", "synthetic"}
def _app(tel: tmod.Telemetry) -> FastAPI:
"""The real middleware + handler, re-registered on a bare app (no DB)."""
from api.app import main
app = FastAPI()
app.state.telemetry = tel
app.middleware("http")(main._count_requests)
app.exception_handler(RepairPointer)(main._repair_pointer_handler)
def refuse(code: str, status: int):
def dep():
raise RepairPointer(
status_code=status,
code=code,
speak="no",
machine_cause="test",
remediation_tool=None,
)
return dep
@app.get("/api/v1/repos/{repo}/grants", dependencies=[Depends(refuse("passport_revoked", 403))])
async def grants(repo: str):
return {}
@app.get("/api/v1/nope", dependencies=[Depends(refuse("repo_not_found", 404))])
async def nope():
return {}
@app.get("/ok")
async def ok():
return {"ok": True}
return app
async def _get(app, path, headers=None):
async with httpx.AsyncClient(transport=httpx.ASGITransport(app=app), base_url="http://t") as c:
return await c.get(path, headers=headers or {})
def _tel():
return tmod.Telemetry(
"http://ledger.invalid/v1/events",
"tok",
environment="test",
commit_sha="abc1234",
version="0.1.0",
)
@pytest.mark.asyncio
async def test_refusal_emits_one_declared_row_with_route_template_not_path():
tel = _tel()
r = await _get(
_app(tel),
"/api/v1/repos/grandmas-secret-project/grants",
{"Authorization": "Bearer eyJhbGciOiJFUzI1NiIsInR5cCI6IkVQVCJ9.e30.x"},
)
assert r.status_code == 403
rows = [e for e in tel.buffer if e["event_type"] == "forge.auth.failed"]
assert len(rows) == 1
row = rows[0]
assert row["actor_type"] == "system" and "actor_id" not in row
assert set(row["metadata"]) <= DECLARED_AUTH_KEYS
assert row["metadata"]["code"] == "passport_revoked"
assert row["metadata"]["http_status"] == 403
assert row["metadata"]["caller"] == "anonymous_agent"
assert row["metadata"]["route"] == "/api/v1/repos/{repo}/grants"
assert "grandmas-secret-project" not in str(row)
@pytest.mark.asyncio
async def test_non_auth_errors_are_counted_but_not_refusal_rows():
tel = _tel()
await _get(_app(tel), "/api/v1/nope")
assert not [e for e in tel.buffer if e["event_type"] == "forge.auth.failed"]
assert tel.errors_4xx == 1 and tel.refusals_4xx == 0
@pytest.mark.asyncio
async def test_heartbeat_counts_requests_refusals_and_p95():
tel = _tel()
app = _app(tel)
for _ in range(3):
await _get(app, "/ok")
await _get(app, "/api/v1/repos/x/grants")
meta = tel.health_row()
assert meta["requests"] == 4 and meta["refusals_4xx"] == 1 and meta["errors_5xx"] == 0
assert isinstance(meta["p95_ms"], int)
assert {"interval_s", "uptime_s"} <= set(meta)
def test_no_traffic_means_no_p95_not_a_fake_zero():
assert "p95_ms" not in _tel().health_row()
def test_no_token_sends_and_buffers_nothing():
tel = tmod.Telemetry("http://ledger.invalid", "")
tel.boot()
tel.auth_failed(code="token_invalid", http_status=401, caller="unknown")
assert tel.buffer == []
def test_unknown_code_is_never_sent_as_a_refusal():
tel = _tel()
tel.auth_failed(code="made_up_code", http_status=401, caller="unknown")
assert tel.buffer == []
def test_boot_omits_an_unknown_commit_rather_than_inventing_one():
tel = tmod.Telemetry("http://x", "tok", commit_sha=None, version="0.1.0")
tel.boot()
assert "commit_sha" not in tel.buffer[0]["metadata"]
def test_caller_classes_are_the_declared_three():
assert tmod.caller_class({}) == "unknown"
assert tmod.caller_class({"x-service-token": "s"}) == "unknown"
assert (
tmod.caller_class({"authorization": "Bearer eyJhbGciOiJSUzI1NiJ9.e30.x"})
== "anonymous_human"
)
@pytest.mark.asyncio
async def test_synthetic_header_marks_the_row_and_absent_means_real():
async def refusal(headers):
tel = _tel()
await _get(_app(tel), "/api/v1/repos/x/grants", headers)
return [e for e in tel.buffer if e["event_type"] == "forge.auth.failed"][0]["metadata"]["synthetic"]
assert await refusal({"X-Windy-Synthetic": "1"}) is True
assert await refusal({}) is False
def test_synthetic_is_forwarded_downstream_only_for_synthetic_requests():
token = tmod.SYNTHETIC.set(True)
try:
assert tmod.synthetic_headers() == {"X-Windy-Synthetic": "1"}
finally:
tmod.SYNTHETIC.reset(token)
assert tmod.synthetic_headers() == {}
# ---- UPDATE 7: the ledger answers 202 even when it quarantines rows ----------
@pytest.mark.asyncio
async def test_quarantined_rows_are_warned_and_counted_on_the_next_heartbeat(monkeypatch, caplog):
tel = _tel()
tel.boot()
monkeypatch.setattr(
tel, "_post", lambda b: (202, {"accepted": 0, "quarantined": 1, "rejections": ["undeclared key"]})
)
with caplog.at_level("WARNING", logger="windy-git.telemetry"):
await tel.flush()
assert tel.buffer == [] # sent; the ledger dead-lettered it, retrying won't help
assert "QUARANTINED" in caplog.text and "undeclared key" in caplog.text
tel.health()
assert tel.buffer[-1]["metadata"]["telemetry_quarantined"] == 1
assert tel.health_row()["telemetry_quarantined"] == 0 # reset per heartbeat window
@pytest.mark.asyncio
async def test_clean_send_reports_zero_and_logs_nothing(monkeypatch, caplog):
tel = _tel()
tel.boot()
monkeypatch.setattr(tel, "_post", lambda b: (202, {"accepted": 1, "quarantined": 0, "rejections": []}))
with caplog.at_level("WARNING", logger="windy-git.telemetry"):
await tel.flush()
assert caplog.text == ""
row = tel.health_row()
assert row["telemetry_quarantined"] == 0 and row["telemetry_dropped"] == 0
def test_buffer_overflow_is_counted_as_dropped(monkeypatch):
monkeypatch.setattr(tmod, "MAX_BUFFER", 3)
tel = _tel()
for _ in range(5):
tel.boot()
assert len(tel.buffer) == 3
assert tel.health_row()["telemetry_dropped"] == 2

View File

@@ -0,0 +1,106 @@
"""G3.5 — the Eternitas webhook receiver, driven over HTTP (audit 2026-08-13).
The G3.5 invariants in test_invariants.py grep webhooks.py for strings; a
refactor that kept the strings and broke the behaviour would pass them all.
These send real requests through the real route (no DB: every case here stops
before the revocation handler) and assert what the receiver DOES.
"""
from __future__ import annotations
import hashlib
import hmac
import json
import httpx
import pytest
from fastapi import FastAPI
from fastapi.responses import JSONResponse
from api.app.config import Settings
from api.app.errors import RepairPointer
from api.app.routes import webhooks
SECRET = "s" * 64
URL = "/api/v1/webhooks/eternitas"
def _app(secret: str = SECRET) -> FastAPI:
app = FastAPI()
app.include_router(webhooks.router)
app.state.settings = Settings(eternitas_webhook_secret=secret)
@app.exception_handler(RepairPointer)
async def _h(_, exc: RepairPointer) -> JSONResponse:
return JSONResponse(status_code=exc.status_code, content=exc.detail)
return app
async def _post(body: bytes, headers: dict, secret: str = SECRET) -> httpx.Response:
transport = httpx.ASGITransport(app=_app(secret))
async with httpx.AsyncClient(transport=transport, base_url="http://t") as c:
return await c.post(
URL, content=body, headers={"content-type": "application/json", **headers}
)
def _sig(raw: bytes, secret: str = SECRET) -> str:
return hmac.new(secret.encode(), raw, hashlib.sha256).hexdigest()
# Deliberately odd spacing/key order: a receiver that re-serialises before
# hashing produces a different digest and must fail.
RAW = b'{"event":"windygit.selftest", "data": {"b": 2, "a": 1}}'
EVENT = {"x-eternitas-event": "windygit.selftest"}
@pytest.mark.asyncio
async def test_prefixed_and_bare_digests_are_both_accepted():
for header in (f"sha256={_sig(RAW)}", _sig(RAW)):
r = await _post(RAW, {**EVENT, "x-eternitas-signature": header})
assert r.status_code == 200, r.text
assert r.json()["acted"] is False # unknown event: received, nothing done
@pytest.mark.asyncio
async def test_digest_of_reserialised_json_is_refused():
reserialised = json.dumps(json.loads(RAW)).encode()
assert reserialised != RAW
r = await _post(RAW, {**EVENT, "x-eternitas-signature": f"sha256={_sig(reserialised)}"})
assert r.status_code == 401 and r.json()["code"] == "webhook_signature_invalid"
@pytest.mark.asyncio
async def test_forged_or_wrong_key_signature_is_refused():
for header in ("sha256=" + "0" * 64, f"sha256={_sig(RAW, 'other-secret')}", "garbage"):
r = await _post(RAW, {**EVENT, "x-eternitas-signature": header})
assert r.status_code == 401, header
@pytest.mark.asyncio
async def test_signed_event_without_signature_is_refused():
r = await _post(RAW, EVENT)
assert r.status_code == 401
@pytest.mark.asyncio
async def test_unset_secret_refuses_rather_than_accepts():
r = await _post(RAW, {**EVENT, "x-eternitas-signature": f"sha256={_sig(RAW)}"}, secret="")
assert r.status_code == 503 and r.json()["code"] == "webhook_secret_unset"
@pytest.mark.asyncio
async def test_revocation_with_bad_signature_never_reaches_the_handler():
body = b'{"event":"passport.revoked","passport":"ET26-TEST-GOOD"}'
r = await _post(
body,
{"x-eternitas-event": "passport.revoked", "x-eternitas-signature": "sha256=" + "f" * 64},
)
assert r.status_code == 401 # refused before any DB work
@pytest.mark.asyncio
async def test_reachability_ping_acknowledges_but_never_acts():
r = await _post(b'{"anything": "at all"}', {"x-eternitas-event": "platform.test_ping"})
assert r.status_code == 200 and r.json()["acted"] is False

28
deploy/branding/README.md Normal file
View File

@@ -0,0 +1,28 @@
# Windy Git branding (G2.3)
Gitea's **supported** customisation surface: custom templates and public assets.
No Gitea source is modified, so upstream upgrades keep arriving (D-2, I-1).
deploy/branding/ → $GITEA_CUSTOM (/data/gitea) in the container
→ /srv/windygit/git/gitea/ on Veron 1
Apply with `./deploy/branding/apply.sh`.
## Two traps this cost, both worth knowing
**1. `GITEA__DEFAULT__APP_NAME` does not work.** Gitea reads `APP_NAME` from the
*top level* of `app.ini` (before any `[section]`). The env var instead created a
literal `[default]` section, which Gitea ignores — and the installer's stock
`APP_NAME` kept winning, so the site said "Gitea: Git with a cup of tea" while
the config looked correct. Worse, the env-to-ini pass **appended** a second
`APP_NAME` rather than replacing the first. `apply.sh` sets it at the top level
directly.
**2. Cloudflare caches `/assets/*` for 6 hours and no token in this stack can
purge.** Editing a fixed filename leaves the old bytes live for hours — the new
logo and CSS were both invisible while being correct at origin. **Version the
filename** (`theme-windy.v2.css`) on every brand change; a `?query` is not
enough because some caches ignore it.
The nav logo is swapped in CSS rather than by overriding Gitea's navbar
template — a one-line rule instead of a forked template that would drift.

23
deploy/branding/apply.sh Executable file
View File

@@ -0,0 +1,23 @@
#!/usr/bin/env bash
# Apply Windy Git branding to the Gitea custom tree. Idempotent.
set -euo pipefail
CUSTOM="${GITEA_CUSTOM_HOST:-/srv/windygit/git/gitea}"
HERE="$(cd "$(dirname "$0")" && pwd)"
sudo mkdir -p "$CUSTOM"/templates/custom "$CUSTOM"/public/assets/img "$CUSTOM"/public/assets/css
sudo cp -r "$HERE"/templates/. "$CUSTOM"/templates/
sudo cp -r "$HERE"/public/. "$CUSTOM"/public/
sudo chown -R 1000:1000 "$CUSTOM"/templates "$CUSTOM"/public
# APP_NAME must sit at the TOP LEVEL. See README trap #1.
INI="$CUSTOM/conf/app.ini"
sudo cp "$INI" "$INI.bak-brand-$(date +%s)"
sudo python3 - "$INI" <<'PY'
import sys
p = sys.argv[1]
lines = [l for l in open(p).read().splitlines()
if not l.strip().startswith(("APP_NAME", "APP_SLOGAN"))]
lines.insert(0, "APP_NAME = Windy Git")
open(p, "w").write("\n".join(lines) + "\n")
PY
echo "applied. restart gitea to pick it up."

View File

@@ -0,0 +1,27 @@
/* Windy Git brand accent. Layered on top of Gitea's theme rather than
replacing it, so upstream theme fixes keep arriving. */
:root {
--color-primary: #0ea5e9;
--color-primary-dark-1: #0284c7;
--color-primary-dark-2: #0369a1;
--color-primary-light-1: #38bdf8;
--color-primary-light-2: #7dd3fc;
}
.wg-hero { max-width: 780px; margin: 4rem auto 2rem; padding: 0 1.5rem; text-align: center; }
.wg-hero h1 { font-size: 2.6rem; margin: 1.2rem 0 .4rem; letter-spacing: -.02em; }
.wg-hero .wg-sub { font-size: 1.15rem; opacity: .78; margin-bottom: 2.2rem; }
.wg-grid { display: grid; gap: 1.1rem; grid-template-columns: repeat(auto-fit,minmax(230px,1fr));
max-width: 900px; margin: 0 auto 2.5rem; padding: 0 1.5rem; text-align: left; }
.wg-card { border: 1px solid var(--color-secondary); border-radius: 8px; padding: 1.1rem 1.2rem; }
.wg-card h3 { margin: 0 0 .35rem; font-size: 1.02rem; }
.wg-card p { margin: 0; opacity: .74; font-size: .9rem; line-height: 1.5; }
.wg-cta { margin-bottom: 3rem; }
/* Logo swap via CSS.
Cloudflare cached the stock /assets/img/logo.svg for 6h and no available API
token can purge. Pointing at a NEW filename sidesteps the stale object
without touching Gitea's own templates — the supported customisation surface,
per D-2 (membrane, not merge). */
img[src$="/assets/img/logo.svg"] {
content: url("/assets/img/wg-mark.svg");
}

View File

@@ -0,0 +1,9 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 32 32" width="32" height="32">
<defs><linearGradient id="w" x1="0" y1="0" x2="1" y2="1">
<stop offset="0%" stop-color="#38bdf8"/><stop offset="100%" stop-color="#0ea5e9"/>
</linearGradient></defs>
<circle cx="16" cy="16" r="15" fill="#0b1220"/>
<path d="M5 11h13a3.2 3.2 0 1 0-3.1-4" fill="none" stroke="url(#w)" stroke-width="2.4" stroke-linecap="round"/>
<path d="M5 16h17a3.6 3.6 0 1 1-3.5 4.5" fill="none" stroke="url(#w)" stroke-width="2.4" stroke-linecap="round"/>
<path d="M5 21h9a2.8 2.8 0 1 1-2.7 3.5" fill="none" stroke="url(#w)" stroke-width="2.4" stroke-linecap="round"/>
</svg>

After

Width:  |  Height:  |  Size: 660 B

View File

@@ -0,0 +1,9 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 32 32" width="32" height="32">
<defs><linearGradient id="w" x1="0" y1="0" x2="1" y2="1">
<stop offset="0%" stop-color="#38bdf8"/><stop offset="100%" stop-color="#0ea5e9"/>
</linearGradient></defs>
<circle cx="16" cy="16" r="15" fill="#0b1220"/>
<path d="M5 11h13a3.2 3.2 0 1 0-3.1-4" fill="none" stroke="url(#w)" stroke-width="2.4" stroke-linecap="round"/>
<path d="M5 16h17a3.6 3.6 0 1 1-3.5 4.5" fill="none" stroke="url(#w)" stroke-width="2.4" stroke-linecap="round"/>
<path d="M5 21h9a2.8 2.8 0 1 1-2.7 3.5" fill="none" stroke="url(#w)" stroke-width="2.4" stroke-linecap="round"/>
</svg>

After

Width:  |  Height:  |  Size: 660 B

View File

@@ -0,0 +1,5 @@
{{/* Windy Git brand layer.
The filename carries a version: Cloudflare caches /assets/* for 6h with no
purge token available to this stack, so editing a fixed filename leaves the
old bytes live for hours. Bump the suffix on every brand change. */}}
<link rel="stylesheet" href="{{AssetUrlPrefix}}/css/theme-windy.v2.css">

View File

@@ -0,0 +1,27 @@
{{template "base/head" .}}
<div role="main" aria-label="{{ctx.Locale.Tr "home"}}" class="page-content home">
<div class="wg-hero">
<img src="{{AssetUrlPrefix}}/img/logo.svg" alt="Windy Git" width="72" height="72">
<h1>Windy Git</h1>
<p class="wg-sub">Your work, every version — and agents as citizens.</p>
<div class="wg-cta">
{{if not .IsSigned}}
<a class="ui primary button" href="{{AppSubUrl}}/user/login">Sign in with Windy</a>
{{else}}
<a class="ui primary button" href="{{AppSubUrl}}/{{.SignedUser.Name}}?tab=repositories">Your repositories</a>
{{end}}
<a class="ui button" href="{{AppSubUrl}}/explore/repos">Explore</a>
</div>
</div>
<div class="wg-grid">
<div class="wg-card"><h3>Code and models, one host</h3>
<p>Git and LFS over Windy Cloud storage. Source, weights and adapters live side by side.</p></div>
<div class="wg-card"><h3>Agents are first-class</h3>
<p>An agent signs in with its own Eternitas passport — not a human's borrowed token — and its work is attributable to it.</p></div>
<div class="wg-card"><h3>Your account, no second password</h3>
<p>Sign in with the Windy account you already have. Nothing new to remember.</p></div>
<div class="wg-card"><h3>Kept, and kept elsewhere</h3>
<p>Every repository is bundled nightly to off-site storage, and the restore is rehearsed rather than assumed.</p></div>
</div>
</div>
{{template "base/footer" .}}

59
deploy/runner/config.yaml Normal file
View File

@@ -0,0 +1,59 @@
# act_runner configuration (G7.1).
#
# Labels are EXPLICIT and PINNED. `ubuntu-latest` is banned (G7.5): all four
# windy-registry workflows use it and every single run fails, because a
# self-hosted runner has no such label unless you invent one. A workflow that
# names a label nobody provides queues forever and looks like a hung CI system
# rather than a typo.
log:
level: info
runner:
file: /data/.runner
capacity: 1 # per runner; parallelism = number of runner services (6). See docker-compose.yml
timeout: 90m # hard ceiling per job. eternitas's serial pytest is ~50 min; keep
# timeout-minutes in each workflow — a hang still reads as a hang
shutdown_timeout: 3m
insecure: false
fetch_timeout: 5s
fetch_interval: 2s
labels:
# windy-git's own workflows.
- "veron-1:docker://catthehacker/ubuntu:act-22.04"
- "linux-x64:docker://catthehacker/ubuntu:act-22.04"
# THE FLEET'S EXISTING CONVENTION. Surveyed 2026-08-12 across ten repos:
# 36 of 36 active workflows say `runs-on: [self-hosted, linux, x64]`.
# They were written for the self-hosted runners that died when the repos
# went private, so advertising these three labels makes every one of them
# runnable AS-IS — no workflow edits, no rewrites.
#
# A job asking for [self-hosted, linux, x64] matches only if the runner
# advertises ALL THREE, so all three are declared separately.
- "self-hosted:docker://catthehacker/ubuntu:act-22.04"
- "linux:docker://catthehacker/ubuntu:act-22.04"
- "x64:docker://catthehacker/ubuntu:act-22.04"
cache:
enabled: true
dir: /data/cache
container:
# Empty = act creates a NETWORK PER JOB and removes it afterwards.
#
# This started as `bridge` for isolation, which was a mistake in both
# directions. It broke service containers — Postgres came up healthy but the
# job could not resolve the name `postgres`, because service DNS aliases only
# exist on a per-job network — and it was *weaker* isolation, since every
# concurrent job shared one flat bridge and could see its neighbours.
#
# A per-job network is both correct and stricter. Still no route to the forge:
# these networks live inside the dind daemon, which has no forge attachment
# at all.
network: ""
privileged: false
options:
workdir_parent: /workspace
valid_volumes: [] # a job cannot bind-mount anything from the daemon host
docker_host: "-" # do NOT expose the runner's own docker socket to jobs
force_pull: false

View File

@@ -0,0 +1,168 @@
# CI runners (strand G7) — a SEPARATE compose project from the forge.
#
# Separate on purpose: runners restart, crash, get starved and get killed. None
# of that should ever touch the thing serving repositories. This is the cell
# doctrine applied one level down.
#
# ── I-5, and why there is a dind sidecar ───────────────────────────────────
#
# "CI never shares a kernel with identity. Runners execute untrusted code and
# are isolated by machine boundary, not container boundary. No runner may hold
# a credential scoped beyond its own job."
#
# act_runner needs a Docker daemon to start job containers. The tempting move is
# to mount the host's `/var/run/docker.sock`. That would hand every workflow —
# including whatever a transitive dependency's postinstall script feels like
# doing — the ability to start a privileged container mounting `/`, which is
# root on Veron 1. Every published act_runner example does exactly this.
#
# Instead the runner talks to its OWN daemon (`dind`). Untrusted job code runs
# as a child of that daemon, on an isolated network, with no route to the host
# socket and no route to the forge's database.
#
# The split that makes this work:
# * `runner` is TRUSTED code (the act_runner daemon). It sits on the forge
# network only so it can reach gitea:3000 to collect jobs.
# * `dind` and every job container it spawns are UNTRUSTED. They are on a
# private network with no access to the forge, its database, or its .env.
#
# dind itself is privileged — that is the cost, and it is the reason a job
# escape lands in a disposable daemon rather than on Grant's workstation.
#
# ⚠️ Do NOT "simplify" this by mounting the host docker socket.
name: windy-git-runner
services:
dind:
image: docker.io/library/docker:27-dind
privileged: true
environment:
DOCKER_TLS_CERTDIR: "" # plain TCP on an isolated network, no host route
command: ["dockerd", "--host=tcp://0.0.0.0:2375", "--tls=false"]
networks: [jobs]
volumes:
- dind-storage:/var/lib/docker
# G1.5 — bounded so a fork-bomb workflow cannot starve Grant's interactive
# session. Veron 1 is his workstation, not a dedicated build box.
cpus: 12.0 # 12 of 24 cores
mem_limit: 64g
restart: unless-stopped
# ── FOUR runners × capacity 1, not one runner × capacity 4 (2026-09-23) ──
#
# act caches every action repo at /root/.cache/act/<hash> INSIDE the runner
# process and re-fetches it at the start of each job. With capacity 4, four
# concurrent jobs share that one directory: one job's refresh rewrites it while
# another is tarring it into its job container, and the job dies with
# `lstat /root/.cache/act/<hash>/…: no such file or directory` on
# `actions/setup-node` / `setup-uv` — a failure that reads like a broken
# workflow. windy-chat (~20 jobs per push) hit it on 3 jobs in its first run.
# `rm -rf /root/.cache/act` only reset the clock. Separate processes get
# separate caches, so the race cannot occur. Same total parallelism, same
# single capped dind — the blast radius is unchanged.
runner: &runner
# 0.2.11 -> 0.6.1 on 2026-08-14. The bundled act in 0.2.11 only knows
# `runs.using: node12|node16|node20`, so ANY repo pinning a current action
# major dies before its first step with "The runs.using key in action.yml
# must be one of: [...], got node24" — Windy-Clone on actions/checkout@v5
# is how this surfaced. Verified: `node24` is absent from the 0.2.11 binary
# and present in 0.6.1. Every key in this directory's config.yaml still
# exists in 0.6.1's schema (0.6.1 only ADDS keys), so the config carries
# over unchanged. Rollback is re-pinning 0.2.11; the registration in the
# runner-data volume survives either way.
image: docker.io/gitea/act_runner:0.6.1
depends_on: [dind]
environment:
# The runner reaches its OWN daemon. Never the host's.
DOCKER_HOST: tcp://dind:2375
# ⚠️ THE PUBLIC URL, deliberately — not http://gitea:3000.
#
# Job containers run inside the dind daemon's own private network, so they
# cannot resolve `gitea`, which lives on the forge network. The first CI
# run failed exactly here: "Could not resolve host: gitea".
#
# There were two ways out, and they are not equivalent:
# (a) put job containers on the forge network — untrusted workflow code
# would then sit one DNS name away from the forge's Postgres. This
# is the easy fix and it quietly repeals I-5.
# (b) send jobs to the PUBLIC forge surface, over the tunnel, exactly
# like any stranger on the internet. Untrusted code gets no private
# network route at all.
#
# (b) is strictly better and it is what this is. The cost is a hairpin —
# container -> tunnel -> Cloudflare -> back to this box — plus Cloudflare's
# ~100s proxy ceiling on any single fetch (G4A.5). For repos measured at
# 0.63 GB of objects across 61 repos, with depth=1 checkouts, that ceiling
# is nowhere near being a problem. Revisit if a model repo ever needs CI.
GITEA_INSTANCE_URL: https://app.windygit.com
GITEA_RUNNER_REGISTRATION_TOKEN: ${RUNNER_TOKEN:?set RUNNER_TOKEN}
GITEA_RUNNER_NAME: veron-1
CONFIG_FILE: /config.yaml
volumes:
- ./config.yaml:/config.yaml:ro
- runner-data:/data
# Only `jobs`. The runner no longer needs the forge network at all, because
# it collects work over the public surface too — so there is now NO path
# from any CI container to the forge's database. That is a better posture
# than the one this file started with.
networks: [jobs]
cpus: 2.0
mem_limit: 4g
restart: unless-stopped
# Each extra runner registers itself on first start (own name, own volume —
# the registration lives in /data/.runner, so volumes must never be shared).
runner-2:
<<: *runner
environment: &env2
DOCKER_HOST: tcp://dind:2375
GITEA_INSTANCE_URL: https://app.windygit.com
GITEA_RUNNER_REGISTRATION_TOKEN: ${RUNNER_TOKEN:?set RUNNER_TOKEN}
GITEA_RUNNER_NAME: veron-1-2
CONFIG_FILE: /config.yaml
volumes: [./config.yaml:/config.yaml:ro, runner-data-2:/data]
runner-3:
<<: *runner
environment:
<<: *env2
GITEA_RUNNER_NAME: veron-1-3
volumes: [./config.yaml:/config.yaml:ro, runner-data-3:/data]
runner-4:
<<: *runner
environment:
<<: *env2
GITEA_RUNNER_NAME: veron-1-4
volumes: [./config.yaml:/config.yaml:ro, runner-data-4:/data]
# 5 and 6 added the same day: with ~11 private repos onboarded (windy-chat
# alone queues ~24 jobs per push) four runners left 50+ jobs waiting. The
# CPU ceiling is dind's (12 of 24 cores, G1.5), not the runner count, so more
# runners add concurrency for I/O-bound jobs (npm ci, uv sync) without
# taking more of Grant's workstation.
runner-5:
<<: *runner
environment:
<<: *env2
GITEA_RUNNER_NAME: veron-1-5
volumes: [./config.yaml:/config.yaml:ro, runner-data-5:/data]
runner-6:
<<: *runner
environment:
<<: *env2
GITEA_RUNNER_NAME: veron-1-6
volumes: [./config.yaml:/config.yaml:ro, runner-data-6:/data]
networks:
jobs:
# Untrusted job containers live here. No route to the forge.
internal: false # jobs legitimately need to fetch dependencies
volumes:
dind-storage:
runner-data:
runner-data-2:
runner-data-3:
runner-data-4:
runner-data-5:
runner-data-6:

51
deploy/runner/egress.sh Executable file
View File

@@ -0,0 +1,51 @@
#!/usr/bin/env bash
# CI egress filter (2026-09-23) — jobs reach the internet, never Grant's network.
#
# Measured before this existed: an ordinary (unprivileged) job container inside
# the CI dind could open SSH, Ollama, and every dev server on Veron
# (192.168.1.73:22/3000/3300/8080/11434) and anything else on the LAN, WireGuard
# or Tailscale. No container escape needed — a malicious npm/pip dependency in
# any first-party repo's CI could walk straight onto the fleet.
#
# All CI traffic leaves through the `windy-git-runner_jobs` bridge (dind NATs
# its job containers onto it). This script, run at boot and after any runner
# compose change, allows on that bridge:
# * traffic between the runners and dind (same bridge)
# * replies (ESTABLISHED/RELATED)
# * DNS (53) — Docker's embedded resolver forwards to the LAN router
# * everything public
# and drops: RFC1918, CGNAT/Tailscale (100.64/10), link-local, and ANY packet
# addressed to the host itself (INPUT), whatever interface IP it targets.
# Idempotent: owned chains are flushed and rebuilt; hooks are added once.
set -euo pipefail
NET=windy-git-runner_jobs
id=$(docker network inspect "$NET" --format '{{.Id}}')
BR="br-${id:0:12}"
ip link show "$BR" >/dev/null
iptables -N WG-CI-EGRESS 2>/dev/null || iptables -F WG-CI-EGRESS
iptables -A WG-CI-EGRESS -o "$BR" -j RETURN
iptables -A WG-CI-EGRESS -m conntrack --ctstate ESTABLISHED,RELATED -j RETURN
iptables -A WG-CI-EGRESS -p udp --dport 53 -j RETURN
iptables -A WG-CI-EGRESS -p tcp --dport 53 -j RETURN
for cidr in 10.0.0.0/8 172.16.0.0/12 192.168.0.0/16 100.64.0.0/10 169.254.0.0/16; do
iptables -A WG-CI-EGRESS -d "$cidr" -j DROP
done
iptables -A WG-CI-EGRESS -j RETURN
iptables -N WG-CI-INPUT 2>/dev/null || iptables -F WG-CI-INPUT
iptables -A WG-CI-INPUT -m conntrack --ctstate ESTABLISHED,RELATED -j RETURN
iptables -A WG-CI-INPUT -j DROP
# Hooks: remove any stale ones (the bridge name changes if the network is
# recreated), then add exactly one of each at the top.
for chain in DOCKER-USER INPUT; do
target=$([ "$chain" = INPUT ] && echo WG-CI-INPUT || echo WG-CI-EGRESS)
while read -r rule; do
iptables -D $chain ${rule#-A $chain }
done < <(iptables -S "$chain" | grep -- "-j $target" || true)
iptables -I "$chain" 1 -i "$BR" -j "$target"
done
echo "ci egress filter active on $BR ($NET)"

28
deploy/runner/prune.sh Executable file
View File

@@ -0,0 +1,28 @@
#!/usr/bin/env bash
# Keep CI storage bounded (2026-09-23).
#
# Kit 0's 09-01 wipe began with CI `_work` dirs (74 GB) + Docker filling the
# disk. Windy Git's runners have no host `_work` dir — every job runs in a
# container inside the CI-only dind — so the thing that grows here is dind's
# image/volume store (38 GB when this was written, never pruned). This prunes
# ONLY that daemon, over its own socket. It never touches the host's Docker.
#
# In-use images/volumes are never removed, so a running job is safe.
set -euo pipefail
CAP_GB="${CI_STORAGE_CAP_GB:-60}"
D=(docker exec windy-git-runner-dind-1 docker -H tcp://127.0.0.1:2375) # dind listens on TCP only
"${D[@]}" container prune -f --filter until=6h >/dev/null
"${D[@]}" volume prune -af >/dev/null # job workspaces of finished jobs
"${D[@]}" image prune -af --filter until=168h >/dev/null
"${D[@]}" builder prune -af --filter until=168h >/dev/null 2>&1 || true
used_gb=$(du -s --block-size=1G /var/lib/docker/volumes/windy-git-runner_dind-storage | cut -f1)
if (( used_gb > CAP_GB )); then
# Over the cap even after the age-based pass: drop every unused image. The
# next jobs re-pull (the act image is ~2 GB) — slower, never wrong.
"${D[@]}" image prune -af >/dev/null
used_gb=$(du -s --block-size=1G /var/lib/docker/volumes/windy-git-runner_dind-storage | cut -f1)
fi
echo "ci storage ${used_gb}G (cap ${CAP_GB}G)"
(( used_gb <= CAP_GB )) || { echo "STILL OVER CAP"; exit 1; }

View File

@@ -0,0 +1,13 @@
[Unit]
Description=Windy Git - CI egress filter (jobs reach the internet, never the LAN/host)
After=docker.service
Requires=docker.service
[Service]
Type=oneshot
RemainAfterExit=yes
# The jobs network exists once the runner compose project is up; retry until it does.
ExecStart=/bin/bash -c 'for i in $(seq 1 60); do /srv/windygit/src/deploy/runner/egress.sh && exit 0; sleep 5; done; exit 1'
[Install]
WantedBy=multi-user.target

View File

@@ -0,0 +1,6 @@
[Unit]
Description=Windy Git - prune CI-only dind storage (bounded, never the host daemon)
[Service]
Type=oneshot
ExecStart=/srv/windygit/src/deploy/runner/prune.sh

View File

@@ -0,0 +1,9 @@
[Unit]
Description=Windy Git - prune CI storage every 6 hours
[Timer]
OnCalendar=*-*-* 00/6:37:00
Persistent=true
[Install]
WantedBy=timers.target

View File

@@ -0,0 +1,13 @@
[Unit]
Description=Windy Git nightly backup (git bundles + windgit schema -> R2)
After=network-online.target docker.service
[Service]
Type=oneshot
WorkingDirectory=/srv/windygit/src
# The .env holds the R2 credentials. The script refuses to run without them
# rather than reporting a backup that did not happen.
EnvironmentFile=/srv/windygit/src/.env
ExecStart=/bin/bash /srv/windygit/src/scripts/backup.sh
Nice=10
IOSchedulingClass=idle

View File

@@ -0,0 +1,3 @@
[Service]
ExecStart=
ExecStart=/usr/local/bin/windy-job windygit-backup 26h --expect "ok — [0-9]+ repos" --owner 13 -- /bin/bash /srv/windygit/src/scripts/backup.sh

View File

@@ -0,0 +1,12 @@
[Unit]
Description=Nightly Windy Git backup
[Timer]
OnCalendar=*-*-* 04:17:00
# Grant's workstation is not always on at 04:17. Without this a missed window
# is simply skipped and the backup silently never runs.
Persistent=true
RandomizedDelaySec=600
[Install]
WantedBy=timers.target

View File

@@ -0,0 +1,3 @@
[Service]
ExecStart=
ExecStart=/usr/local/bin/windy-job windygit-ci-prune 7h --expect "ci storage [0-9]+G" --owner 13 -- /srv/windygit/src/deploy/runner/prune.sh

View File

@@ -0,0 +1,12 @@
[Unit]
Description=Sync GitHub -> Windy Git (Phase 1: GitHub is the source of truth)
After=network-online.target docker.service
[Service]
Type=oneshot
WorkingDirectory=/srv/windygit/src
EnvironmentFile=/srv/windygit/src/.env
# GITHUB_TOKEN is set on the host only (root-only unit file / .env) — NEVER commit it.
Environment=GITHUB_OWNER=sneakyfree
ExecStart=/bin/bash /srv/windygit/src/scripts/sync_from_github.sh
Nice=10

View File

@@ -0,0 +1,3 @@
[Service]
ExecStart=
ExecStart=/usr/local/bin/windy-job windygit-sync 20m --expect "all repos in step with GitHub" --owner 13 -- /bin/bash /srv/windygit/src/scripts/sync_from_github.sh

View File

@@ -0,0 +1,10 @@
[Unit]
Description=Keep Windy Git in step with GitHub every 5 minutes
[Timer]
OnBootSec=3min
OnUnitActiveSec=5min
Persistent=true
[Install]
WantedBy=timers.target

View File

@@ -0,0 +1,16 @@
[Unit]
Description=Windy Git - Cloudflare Tunnel (the only ingress; no inbound port is opened)
After=network-online.target
Wants=network-online.target
[Service]
Type=notify
ExecStart=/usr/bin/cloudflared --no-autoupdate --config /etc/cloudflared/config.yml tunnel run
Restart=always
RestartSec=5
# G1.4 - bounded, so a misbehaving ingress can never starve Grant's workstation.
MemoryMax=512M
CPUQuota=100%
[Install]
WantedBy=multi-user.target

View File

@@ -16,7 +16,11 @@ services:
# I-12: baked at build time. A runtime COMMIT_SHA override is ignored.
COMMIT_SHA: ${COMMIT_SHA_BUILD:-}
BUILT_AT: ${BUILT_AT:-}
env_file: [.env]
env_file:
- .env
# WINDYGIT_TELEMETRY_TOKEN (root-only on Veron). Optional: no file = no telemetry.
- path: /etc/windygit/telemetry.env
required: false
environment:
DATABASE_URL: postgresql+asyncpg://windygit:${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD}@db:5432/windygit
GITEA_BASE_URL: http://gitea:3000
@@ -57,6 +61,37 @@ services:
# G2.2 — OIDC only. No local password login, no self-registration.
GITEA__service__DISABLE_REGISTRATION: "true"
GITEA__service__ALLOW_ONLY_EXTERNAL_REGISTRATION: "true"
# SSO #8 (2026-09-23): the password and passkey forms are OFF. windyadmin
# is site admin with a local password; leaving the form up made that
# password a second, phishable way into the whole forge. Break-glass is
# the CLI on Veron (`docker exec -u git windy-git-gitea-1 gitea admin ...`).
GITEA__service__ENABLE_PASSWORD_SIGNIN_FORM: "false"
GITEA__service__ENABLE_PASSKEY_AUTHENTICATION: "false"
# G3.1 — a Windy account IS the account. Signing in with Windy provisions
# the Gitea user on first arrival; nobody is asked to invent a second
# identity for the same person, and no local password ever exists.
# 🔴 OFF (2026-09-23). With it on, ANY stranger with a Windy Word account
# (public signup, not even email-verified) got a forge account on first
# sign-in — and the CI runners were instance-wide, so their workflows
# would run on Veron beside the R2 god token. Proven with a throwaway
# account, then closed. Opening the forge to non-Grant users is a §7
# Grant decision; until then new accounts are created deliberately.
GITEA__oauth2_client__ENABLE_AUTO_REGISTRATION: "false"
GITEA__oauth2_client__USERNAME: email
# 🔴 `login`, NOT `auto` (SSO #8). `auto` linked any hub login whose EMAIL
# matched an existing account — and windyadmin (SITE ADMIN) carries Grant's
# email, so the forge's admin rights rested on the hub never letting anyone
# else hold that address. `login` makes an email match prove possession of
# the existing account first. Grant is unaffected: his account is already
# linked by the hub's stable `sub`, which is matched before email.
# ⚠️ env-to-ini SETS but never UNSETS — this value must also be edited in
# /srv/windygit/git/gitea/conf/app.ini if it is ever removed from here.
GITEA__oauth2_client__ACCOUNT_LINKING: login
# The email is asserted by account-server, which is the authority on it.
# Asking the user to re-verify an address their identity provider already
# verified is friction that buys nothing.
GITEA__oauth2_client__UPDATE_AVATAR: "false"
GITEA__service__REGISTER_EMAIL_CONFIRM: "false"
# G4.3 — heavy bytes to R2 at ZERO egress. Git object databases stay on
# local NVMe (I-3); this covers LFS, attachments, packages, avatars and
# Actions artifacts, which is where GitHub's painful bills actually come

View File

@@ -0,0 +1,116 @@
# Second-auditor review of the Windy Git build (Fable, 2026-08-13)
A fresh-eyes trace of Opus's build, **verified against the live system** rather
than against the transcript's own account. One critical finding was fixed during
the review; the rest are recorded here with proportion.
## Fixed during this review
### 🔴 Agent authentication was not authentication (CRITICAL, was live+public)
`auth.py` read the passport out of a bearer token **without verifying the EPT
signature**, asked Eternitas *"is this passport reputable?"*, and seated the
caller on a yes. That answers reputation, not possession.
**Proven by exploit:** a forged `alg:none` token naming a passport lifted from
the logs returned **HTTP 200** as that agent on the public API. Anyone who knows
a passport number (they appear in logs, the lockbox, and revocation messages)
could impersonate that agent.
The human path already failed closed for exactly this reason
(`require_verified_jwt`) — but the gate sat *after* the agent branch returned, so
it protected the safe path and skipped the exploitable one.
**Fix:** the agent path now fails closed in production, *before* the trust
lookup, mirroring the human path. Reopens automatically when the ES256/JWKS
verifier (G3.2/G9.1) is built. Re-tested live: forged token now `503`. Added a
**behavioral** test (forged token through real `get_caller` must raise) and a
**canary probe** (forged token must stay refused; a 2xx pages).
## Open findings (recorded, not yet fixed)
### 🟠 The EI throttle table is dead code (MEDIUM)
`BAND_MULTIPLIER` and `rate_*_per_day` (G3.4) are defined and **read by
nothing** — verified by grep. An authenticated agent has no rate limit at all.
When the agent path reopens with real verification, wire the throttle before it
does, or a single agent can hammer repo-create/token-mint unbounded.
### 🟠 The test suite is mostly source-string assertions (MEDIUM)
56 invariant tests carried **~86 "does the source contain this string"**
assertions and **zero** that exercised the auth decision. `make check` green
means the code still *contains* the patterns, not that it *behaves*. The auth
bypass passed every test. **Direction:** convert the top ~10 guards into
behavioral tests against an ephemeral instance (this review added the first
two). The live-curl proofs Opus ran were excellent but were never captured, so
they don't defend against regression.
### 🟠 Privileged dind sits beside broad-scoped tokens (MEDIUM — Grant-aware)
`dind privileged=true`, and the host `.env` holds the account-wide R2 token and
a GitHub PAT, on the same box running untrusted CI. Escape from privileged dind
→ host fs → both tokens. Grant waived *minting a new token* for the sandbox; the
specific escalation path (privileged-dind-next-to-god-token) is a separate
decision. Lowest-effort mitigations: rootless/sysbox runner, or move the two
tokens out of the API container's env into a path the API reads but the runner
host does not share.
### ✅ Revocation — was WORSE than flagged, now fixed (was CRITICAL once agent auth reopened)
Re-examined after real agent auth went live, and the finding grew teeth. A
revoked passport returns **HTTP 200, `status: revoked`, `band: unproven`,
`allowed_actions: []`** (verified live). `resolve_passport` keyed refusal only on
HTTP 4xx and `band=="untrusted"`, so it returned band `unproven` and **seated the
revoked agent** — revocation was not enforced on the live path at all. The
webhook everyone worried about was only ever cache-invalidation; the live trust
lookup was the real gate, and it wasn't checking.
**Fixed:** `decide_trust` now allows only `status=="active"`; revoked / suspended
/ frozen / unknown all refuse, fail-closed. Revocation now takes effect on the
next request, no webhook required. Eternitas additionally refuses to mint EPTs
for revoked bots, so the residual window was a pre-existing ~365-day token —
exactly what the live check now stops. 79 tests green including the full wiring.
The webhook (`webhook_secret` lost to the 500) is now genuinely LOW: it only
matters for locally-issued credentials/grants (G6.3, not built), and it is no
longer the thing standing between a revoked agent and access.
## What is genuinely strong (and worth protecting)
- **Honesty engineering is real:** fail-closed providers, `/health` refusing to
claim green it can't prove, I-12's baked-sha proven live against a hostile env
var. This directly cures the parent ecosystem's #1 root cause.
- **Incident response was first-rate:** the login outage was root-caused to a
510-deep accept queue and a per-query node fork, with the "one function, not
468 call sites" reframe that is correct and valuable.
- **It caught its own mistakes** — the mirror direction and the push-triggered
deploy workflows, the latter *before* they fired.
- **The DR posture is right:** 131 read-only mirrors (no CI, zero deploy risk) +
a rehearsed restore. Rehearsed restore is rare in this ecosystem; keep it.
## The one-line lesson
Opus aimed its considerable discipline at **honesty and documentation**
(excellent) more than at **adversarial correctness** — so the property that
mattered most, *is an agent really that agent*, shipped inverted and untested.
The remedy is not more process; it is **behavioral tests and canary probes for
the security-critical paths**, so verification persists instead of living in a
transcript.
## Disposition update — 2026-09-23 (lane 13)
**Privileged dind beside the tokens — materially reduced, not closed.**
- The account-wide R2 token is **gone from Veron**. `.env` now carries a token scoped
to Workers R2 Bucket Item Read/Write on `windy-git-lfs` + `windy-git-backups` only,
minted by API (verified: works on both buckets, refused on any other). A CI escape
now reaches Windy Git's own two buckets, not every bucket and zone in the account.
- Runners take jobs **only from windyadmin-owned repos** (`action_runner.owner_id`), and
forge self-registration is off, so no stranger's workflow can run here.
- A host egress filter (`deploy/runner/egress.sh`) stops job containers reaching Veron,
the LAN, WireGuard or Tailscale.
- Still open: dind runs `--privileged` (next: Sysbox); the host still holds a GitHub
token and the Gitea admin token.
**Revocation / webhook secret** — `ETERNITAS_WEBHOOK_SECRET` recovered from the
Eternitas platform row and set; signed deliveries verify.
**Tests are string asserts** — the security paths now have behavioural suites
(`test_hub_jwt.py`, `test_webhooks_behavior.py`, `test_pr_status_bridge.py`);
a mutation check showed the old grep invariant passing a broken HMAC prefix strip.

158
docs/CUTOVER.md Normal file
View File

@@ -0,0 +1,158 @@
# Migration plan — GitHub first, Windy Git second, flip per repo
**Superseded the 2026-08-13 "daily driver" cutover, which was premature.**
## What went wrong, recorded so it is not repeated
Nine repos were migrated writable with **push-mirrors pointed at GitHub**. At
the same time a dozen agent sessions on the Mac mini were pushing to GitHub
continuously — so GitHub, not Windy Git, was where the live work actually was.
A push-mirror force-updates refs. On its 8-hour timer it would have pushed Windy
Git's stale copy **over live work, silently, with no conflict to notice.**
All nine mirrors were removed before the first timer fired, and every GitHub
repo was verified untouched (latest push predated the mirrors). **No work was
lost.** The mistake was direction, and the lesson is: *the source of truth is
wherever people are actually typing, not wherever the plan says it should be.*
## Phase 1 — now. Nothing changes for anyone.
Mac mini agents ──push──▶ GitHub ──sync every 15 min──▶ Windy Git ──▶ CI on Veron
- **You do not have to tell your agents anything.** No remote changes, no
coordination, no "everyone stop pushing." They keep working exactly as they
are.
- `windygit-sync.timer` runs `scripts/sync_from_github.sh` every 15 minutes.
- Windy Git is **force-updated** on purpose: it holds nothing anyone depends on,
so GitHub always wins and there is **no merge to reconcile**. That is the
whole point of not flipping until a repo is quiet.
- CI runs on Veron 1 against current code, on the 36 workflows that already say
`runs-on: [self-hosted, linux, x64]`.
Tracked repos live in `REPOS` in the script (currently 9 of 141).
## Phase 2 — later, one repo at a time, only when that repo is idle
For a single repo, when nobody is mid-work on it:
1. Remove it from `REPOS` in `sync_from_github.sh` — **first**, or the sync will
fight its authors and win.
2. Point that repo's sessions at Windy Git:
`git remote set-url origin https://app.windygit.com/windyadmin/<repo>.git`
3. Add a push-mirror back to GitHub with `sync_on_commit: true`, so GitHub stays
a current second copy.
**Never flip more than one repo at a time, and never while an agent is working
in it.** A dozen parallel sessions is exactly the situation where a big-bang
cutover produces the dirty-branch mess this plan exists to avoid.
## The whole account is on Windy Git — in two tiers
143 repos · 1.58 GB · 967 GB free (matches the GitHub archive exactly)
| tier | count | writable | runs CI | deploy risk |
|---|---|---|---|---|
| **read-only mirrors** | 131 | no | **no** | **none** |
| **writable + CI** | 12 | yes | yes | deploys disabled |
**Why the bulk is mirrors, and why that is the safety decision:** sampling 40
repos found **18 carrying deploy / release / publish workflows that trigger on
`push:`** — roughly 63 across the account. Importing those writable with Actions
enabled would have armed sixty-odd production deploy triggers on Veron 1, each
needing disarming by hand. **A pull mirror cannot run Actions at all**, so the
bulk import carries zero execution risk and Gitea syncs it with no script and no
timer.
That splits the two things cleanly: **having a copy** (safe, do it for
everything, now) and **running code** (needs judgement, do it per repo,
deliberately).
Seven repos are empty here because they are empty on GitHub — 0 KB upstream,
verified. Not failed imports.
`windy-pro` **is** present, as a mirror. That is safe: the G11.5 caution is
about making it *writable* while six checkouts and a three-way-forked build
counter disagree on HEAD. A read-only copy of whatever GitHub currently has
carries none of that risk — and it means the DR copy is complete.
## Promoting a mirror to writable + CI
Per repo, deliberately, when that repo is quiet:
1. delete the mirror, re-import with `mirror=false`
2. **review its workflows and disable every deploying one** (see the section
above — this is the step that matters)
3. add it to `REPOS` in `sync_from_github.sh` so it tracks GitHub
4. later, when it flips to Windy-Git-first: remove it from `REPOS` *first*,
repoint its sessions, add a push-mirror back to GitHub
## Private repos: Windy Git IS their CI (permanent, 2026-09-23)
The platform repos stay **private** on GitHub (Grant, 2026-09-23), and private
repos cannot run GitHub Actions on this account at all. Windy Git is therefore
their CI permanently, not a stopgap:
GitHub push ──sync (15 min)──▶ Windy Git ──runner──▶ Veron 1
▲ │
└──── commit status windy-git/<workflow>/<job> ◀───┘ scripts/pr_status_bridge.py
- `pr_status_bridge.py` runs at the end of every sync. It opens a `[GH#N]`
mirror PR in Windy Git for every open **same-repo** GitHub PR (so
`pull_request` workflows fire), closes it when the GitHub PR closes, and posts
each job's result back to GitHub on PR heads and the default-branch head.
**Never merge a `[GH#N]` PR here** — merge on GitHub.
- Fork PRs are never run: their branch is never synced, and untrusted code
beside the privileged dind is the open audit finding.
- Covered repos: `BRIDGE_REPOS` in the script. Public repos are left out on
purpose; they run real GitHub Actions and two verdicts per commit is noise.
- `skipped` jobs post nothing — no green for a job nobody ran.
- **Image-build jobs** (name matches `docker`) post nothing: job containers
have no Docker daemon by design (I-5), so they are red on every commit. A
rootless builder (BuildKit rootless / buildx in the capped dind) is the open
decision that would bring them back.
**Onboarding another private repo** — the promotion steps below, then:
# on Veron 1, as root
set -a; . /srv/windygit/src/.env; set +a
python3 scripts/import_from_github.py <repo> # writable; aborts if the repo exists
# disable EVERY deploying workflow before anything is pushed:
curl -X PUT -H "Authorization: token $GITEA_ADMIN_TOKEN" \
http://localhost:3080/api/v1/repos/windyadmin/<repo>/actions/workflows/deploy.yml/disable
# add <repo> to REPOS in sync_from_github.sh AND BRIDGE_REPOS in pr_status_bridge.py
An import fires no push event, so `main` has no verdict until its next commit.
To get one now: force Windy Git's `main` back one commit, then
`systemctl start windygit-sync` — the sync pushes it forward and CI fires.
⚠️ **`/actions/tasks` lists only jobs a runner has PICKED UP.** Queued runs are
invisible there, so a repo can read "0 runs" while work is waiting. The truth is
`action_run` in the `gitea` database (status 1 success, 2 failure, 5 waiting,
6 running).
## ⚠️ Deploy workflows are DISABLED on Windy Git, deliberately
Six workflows fire on `push:` and deploy to production:
`windy-registry`, `Windy-Clone`, `WindyCloud`, `windy-mind`, `eternitas`
(`deploy.yml`) and `windy-agent` (`release.yml`).
Windy Git now has a working runner, so the next synced commit to `main` would
have attempted a **production deploy from Veron 1**. Their secrets
(`DEPLOY_HOST` / `DEPLOY_KEY` / `VPS_SSH_KEY`) are unset here, so they would
have failed — but they would have failed *loudly on every push*, and any step
before the SSH step would still have run.
All six are now `disabled_manually`. Tests, lints and migration checks stay
**active** — those need no secrets at all, which is why Phase 1 delivers real CI
value immediately.
**Before re-enabling any deploy workflow here, decide deliberately whether
production should be deployable from Windy Git at all.** Kit 0 deploys are
currently manual runbooks; that is a feature, not a gap.
## Backups
`windygit-backup.timer`, nightly 04:17, `git bundle --all` + verify + `windgit`
schema dump to R2, 30-day retention. **Restore rehearsed:** a bundle was pulled
from R2, cloned, and its HEAD matched live `origin/main` exactly.

View File

@@ -15,6 +15,7 @@ Mirrored into `windy-cloud` and `eternitas` on change.
| eternitas | `GET /api/v1/trust/{passport}` | band + allowed_actions |
| eternitas | `GET /api/v1/registry/{passport}/integrity` | ⚠️ note the path — `windy-registry` calls `/api/v1/passports/{p}/status`, which 404s, which is why the integrity index has never been populated |
| account-server | OIDC discovery + JWKS | human identity (G3.1) |
| windy-admin ledger | `POST /v1/events` (admin.windyword.ai) | field telemetry: `ci.run`, `ci.job_cancelled`, `service.boot`, `service.health`, `forge.auth.failed`. Shapes are declared with the ledger owner BEFORE shipping (the server quarantines undeclared keys). Codes, counts, route templates only |
| windy-cloud-sites | `POST /api/v1/sites/{id}/versions` | publish docs from a repo |
## Calls IN

View File

@@ -1,6 +1,10 @@
# RUNBOOK — Windy Git on Veron 1 (rung R0)
Host `Veron-1-5090`, WireGuard `10.10.0.6`, alias `wg-veron`. Passwordless sudo.
Host `Veron-1-5090`, WireGuard `10.10.0.6`, alias `wg-veron` (or `ts-veron`). Passwordless sudo.
**Checkouts (one-repo doctrine):** the ONE standing dev checkout is **OC5
`~/windy-git`** (platform repos live on OC5). `/srv/windygit/src` on Veron is the
*deploy* copy — it holds no local work. Nothing else should exist.
⛔ **Kit 0 is never a host for this service** (D-4). `api/app/main.py` refuses to
boot in production if it finds itself on `72.60.118.54`.
@@ -15,6 +19,9 @@ boot in production if it finds itself on `72.60.118.54`.
| `/etc/cloudflared/config.yml` | tunnel ingress |
| `/etc/cloudflared/windy-git.json` | tunnel credentials, mode 600 |
| `/srv/windygit/src/.env` | secrets, mode 600, **never committed** |
| `/srv/windygit/git/gitea/conf/app.ini` | Gitea's persisted config — env-to-ini SETS but never UNSETS; edit here when removing a `GITEA__*` var |
| `/srv/windygit/sync/*.git` | bare staging copies the GitHub→Windy Git sync pushes from |
| `/srv/windygit/src/deploy/runner/.env` | `RUNNER_TOKEN` — a **windyadmin user-level** registration token (not instance-level; see CI) |
## Ports — all loopback, on purpose
@@ -22,7 +29,7 @@ boot in production if it finds itself on `72.60.118.54`.
|---|---|
| `127.0.0.1:3080` | Gitea (host 3000 is a resident node dev server; 3300 is nginx — **do not fight them for a port**) |
| `127.0.0.1:8600` | windy-git API |
| `127.0.0.1:2000` | cloudflared metrics |
| `127.0.0.1:2001` | cloudflared metrics (`metrics:` in `/etc/cloudflared/config.yml`) — **not 2000**, see Troubleshooting |
**No inbound port is opened.** cloudflared dials out, so the dynamic residential
IP is irrelevant and there is no firewall hole to maintain.
@@ -41,12 +48,26 @@ sudo systemctl status windygit-tunnel
```bash
ssh wg-veron
cd /srv/windygit/src && git pull
cd /srv/windygit/src && git fetch origin && git merge --ff-only origin/main # READ the output
export COMMIT_SHA_BUILD=$(git rev-parse HEAD) BUILT_AT=$(date -u +%Y-%m-%dT%H:%M:%SZ)
sudo -E docker compose up -d --build
sudo -E docker compose up -d --build --no-deps api # API only: no forge restart
curl -s https://api.windygit.com/version # MUST equal git rev-parse HEAD
```
A Gitea config change (compose `GITEA__*`) needs `sudo docker compose up -d --no-deps gitea`
— a ~6 s forge outage; running CI jobs survive it. Check `app.ini` afterwards.
⚠️ **Never `git pull -q` in a deploy script.** `-q` hides *errors*, not just
noise. On 2026-08-14 a divergent branch made `pull -q` fail silently and the
"deploy" ran for 20 minutes against stale code while reporting success. Use
`git fetch && git merge --ff-only` (or `reset --hard origin/main` on THIS
checkout only, which holds no local work) and read the output.
⚠️ **Never force-push a branch a deploy checkout tracks.** An earlier
`git commit --amend` + `--force-with-lease` rewrote history `/srv/windygit/src`
was already sitting on, orphaning it. If you must amend, re-point the deploy
checkout in the same breath.
⚠️ **Never put `COMMIT_SHA` in `.env`.** It does nothing here — the sha is baked
into the image and a runtime override is ignored with a warning (I-12). That env
pin is the documented root cause of nine sibling services misreporting their
@@ -61,11 +82,54 @@ curl -sI https://app.windygit.com/ | head -1 # Gitea, 200
sudo ss -tlnp | grep -E "3080|8600" # both must be 127.0.0.1
```
## Timers (host systemd units — the sync timer is NOT in the repo)
| Unit | Cadence | Does |
|---|---|---|
| `windygit-sync.timer` | every 5 min (`OnUnitActiveSec`) | GitHub → Windy Git for `REPOS` in `scripts/sync_from_github.sh`, then `scripts/pr_status_bridge.py` (mirror PRs + GitHub commit statuses). A manual `systemctl start` RESETS the 5-min clock. |
| `windygit-backup.timer` | nightly | `git bundle` + pg_dump → R2, 30-day retention |
| `windygit-ci-prune.timer` | every 6 h | `deploy/runner/prune.sh` — CI dind storage, 60 GB cap |
| `windygit-tunnel.service` | always | the only ingress |
## CI (Gitea Actions) — see `docs/CUTOVER.md` for onboarding a repo
- **Six runners × capacity 1** (`deploy/runner/docker-compose.yml`), one shared
dind capped at 12 cores / 64 GB. Capacity >1 in one runner shares
`/root/.cache/act` between jobs and races (`lstat …: no such file`).
- **Runners are scoped to the `windyadmin` user** (`action_runner.owner_id=1`),
so only first-party repos run. A repo owned by anyone else — a plane-created
agent or `u-system` repo — gets NO runner. Re-registrations inherit this
because `RUNNER_TOKEN` is user-level.
- Job ceiling 90 min (`config.yaml` `runner.timeout`); a `config.yaml` change
needs each runner restarted **while idle** — `compose up -d` won't recreate it.
- `/actions/tasks` lists only PICKED-UP jobs. Queue truth is `action_run_job`
in the `gitea` DB: `sudo docker exec -i windy-git-db-1 psql -U windygit -d gitea`
(status 1 ok · 2 fail · 3 cancelled · 4 skipped · 5 waiting · 6 running).
- Job logs are in R2, not on disk. `GET /api/v1/repos/{o}/{r}/actions/jobs/{JOB_ID}/logs`
takes the `action_run_job` id, not the task id.
## Sign-in posture
- Windy SSO only: password + passkey forms OFF, `ACCOUNT_LINKING=login`,
**auto-registration OFF** — opening the forge to non-Grant users is a §7
Grant decision.
- **Break-glass:** `sudo docker exec -u git windy-git-gitea-1 gitea admin user generate-access-token --username windyadmin --token-name <name> --scopes <scopes> --raw`
(delete it after: `delete from access_token where name='<name>'` in the gitea DB —
Gitea refuses token management over token auth).
## Troubleshooting
**A hostname returns 530 or won't resolve** — the tunnel is down. `sudo systemctl
restart windygit-tunnel`, then `journalctl -u windygit-tunnel -n 50`.
**`windygit-tunnel` crash-loops with `bind: address already in use` on the metrics
port** — cloudflared exits if it cannot bind `metrics:`, taking ingress with it.
Until 2026-09-23 this unit restarted ~91,000 times because another project's
`cornercall-tunnel` held 127.0.0.1:2000; ingress only survived because a stray
generic `cloudflared.service` ran the same config (now disabled). Windy Git's
metrics port is **2001**. `sudo ss -ltnp | grep :2001` names any squatter.
Keep exactly ONE unit running `/etc/cloudflared/config.yml`: `windygit-tunnel`.
**TLS handshake fails with `curl` exit 35 and no HTTP status at all** — someone
added a **two-level** hostname. Free Universal SSL covers `windygit.com` and
`*.windygit.com` only. The request dies before the tunnel is consulted, so it

262
docs/TURNOVER-2026-08-14.md Normal file
View File

@@ -0,0 +1,262 @@
# Windy Git — turnover, 2026-08-14
Paste the block at the bottom into a fresh terminal. Everything above is context
for whoever reads this file directly.
## Where things stand
Windy Git is **live and in use**: `app.windygit.com` (forge), `api.windygit.com`
(our plane), on **Veron 1** behind a Cloudflare Tunnel, zero inbound ports, $0/mo.
143 repos, 85 tests green, health `ok` on all four checks.
Grant signs in with his existing Windy Word credentials — no second account.
Agents authenticate with real Eternitas EPT signature verification and are
rate-limited by integrity band.
## SOLVED — the CI failures were never about Postgres
The `localhost` → `postgres` fix was correct and is worth keeping, but it was
**not** what was failing these jobs. They died at step 2, before Postgres was
ever contacted.
**Root cause: `astral-sh/setup-uv@v4` asks the forge for uv's latest release.**
setup-uv v4 added "resolve latest version instead of downloading latest release"
(astral-sh/setup-uv#178). Resolution goes through `@actions/github`, whose
octokit reads **`GITHUB_API_URL`** — which act_runner points at *our forge*. So
the action requested:
```
GET https://app.windygit.com/api/v1/repos/astral-sh/uv/releases/latest → 404
```
Gitea has no `astral-sh/uv`, so it answered its standard 404 body, *"The target
couldn't be found."* setup-uv threw that string, act printed it as `::error::`,
and every later step was skipped by `success()`.
**The fix (merged to GitHub, 11 repos):** pin an explicit `version:` on every
`setup-uv@v4`/`@v5` step. `resolveVersion()` short-circuits on an explicit
version *before* any API call, and the download URL is hardcoded to github.com —
so the forge round-trip disappears. Pinned to `0.12.5`, which is what `latest`
already resolved to.
windy-mind #101, WindyCloud #90, eternitas #150, then the sweep: Windy-Clone #77,
windy-agent #355, windy-call #34, windy-cell #31, windy-hand #6, windy-mail #105,
windy-search #77, windy-text #29. All merged and synced.
### Two things that made this hard to see, both worth keeping
- **act attributes the error to the wrong step.** `::error::The target couldn't
be found.` is printed immediately after `actions/checkout`'s `::remove-matcher`,
so it reads exactly like a checkout failure. It is not. What settled it was the
**Gitea access log** — `sudo docker logs windy-git-gitea-1 | grep " 404 "` — which
named the real URL at the same millisecond as the job error. When a job fails
with an opaque forge-shaped message, go to the forge's access log, not the job log.
- **The natural experiment was sitting right there.** windy-registry and
windy-drops use `setup-uv@v3` and always passed; every v4/v5 caller failed. A
version skew across otherwise-identical repos is a diagnosis, not a coincidence.
### The jobs API "job not found" that blocked the last session
Not a bug. `GET /api/v1/repos/{owner}/{repo}/actions/jobs/{id}/logs` requires the
job id to belong to **the repo in the path** — a valid id under the wrong owner/repo
404s. The API works fine; the URLs were mismatched. Logs are readable this way and
you do **not** need the web UI.
Note logs are **not on disk** — `[storage] STORAGE_TYPE = minio` sends action logs
to R2, so `actions_log/` on the host is empty. Read them through the API.
## SOLVED — second root cause: the runner was four majors behind
Windy-Clone failed for a completely different reason: `The runs.using key in
action.yml must be one of: [composite docker node12 node16 node20 go], got
**node24**`. act_runner **0.2.11**'s bundled act predates node24, so any repo
pinning a current action major (`actions/checkout@v5`, `actions/setup-python@v6`)
died before its first step.
**Bumped to `gitea/act_runner:0.6.1`** (`cd7dd9b`). Verified beforehand that
`node24` is absent from the 0.2.11 binary and present in 0.6.1, and that every
key in `deploy/runner/config.yaml` still exists in 0.6.1's schema — 0.6.1 only
*adds* keys, so the config carried over untouched. The runner re-declared with
the same id and labels `[veron-1 linux-x64 self-hosted linux x64]`; the
registration in the `runner-data` volume survived. Windy-Clone went 4/4 red →
4/4 green. Rollback is re-pinning 0.2.11; registration is backed up at
`/srv/windygit/runner-registration.bak`.
## What is still red, and why each one is real
The CI plane is healthy. windy-mind is fully green (`742 passed, 1 skipped`).
What remains are genuine repo defects that were **invisible before**, because
every job died at step 2:
| repo | job | cause |
|---|---|---|
| WindyCloud | `lint` | `ruff format --check` — 7 files would be reformatted |
| eternitas | `py-sdk` | `uv run pytest` → `Failed to spawn: pytest`; pytest isn't a declared dep of that project |
| eternitas | `test` | needs its own look |
| windy-agent | `test (3.12/3.13/3.14)` | all three died **together** at 22:50:19 after ~43 min, mid-suite at 64%, with no verdict in the log. Not the 30m `runner.timeout` (that would have fired at 22:36) and not the runner bump (that was 23:00). Something bulk-killed them; unexplained. |
| WindyCloud | `docker`, windy-search `Docker build` | **architectural** — see below |
`WindyCloud`'s `docker` job wants to build an image and gets `failed to connect
to the docker API at unix:///var/run/docker.sock`. Job containers deliberately
have **no** docker socket (I-5, and `deploy/runner/docker-compose.yml` says in
so many words not to mount it). Mounting the host socket would hand every
workflow root on Veron 1. This needs a decision — buildx-in-dind, a rootless
builder, or "this job does not run on Windy Git" — not a quiet socket mount.
The WindyCloud `lint` and eternitas `py-sdk` rows are small fixes in their own
repos, left alone on purpose: they are product defects, not forge defects.
## Second, smaller finding — act's action cache rots
act caches action repos at `/root/.cache/act/<hash>` inside the runner container
and refreshes them with a go-git mirror fetch of `refs/*:refs/*`, unforced. That
includes `refs/pull/*`, which GitHub **recomputes** whenever a base branch moves.
Reproduced directly:
```
! [rejected] refs/pull/1015/merge -> refs/pull/1015/merge (non-fast-forward)
```
which surfaces as `Non-terminating error while running 'git clone': some refs
were not updated`, after which the action does not report `Checked out <ref>`.
It **has** now failed a job on its own. After the runner bump, `windy-mind tests`
died with:
```
❌ Failure - Main Install uv
lstat /root/.cache/act/d3e6…/.git-blame-ignore-revs: no such file or directory
```
act tars the cached action directory into the job container, and a file vanished
mid-walk. The cache dir had been created at 23:01 and mutated again at 23:03,
with `.gitignore` showing as deleted — act removes it before `docker cp`. The
likely mechanism is **concurrent jobs sharing one cache dir**: `capacity: 4`, and
windy-mind fires four jobs at once that all use setup-uv. Wiping the cache and
re-running the job alone made it pass (`742 passed, 1 skipped`). *Mechanism not
isolated* — the wipe alone may have been sufficient.
One dead end worth not repeating: the cached worktree sits at `38f3f104` while
`git rev-parse v4` says `e4db8464`. That is **not** a wrong checkout — `v4` is an
*annotated tag*, and `v4^{commit}` is `38f3f104`. Don't chase it.
Wipe with `sudo docker exec windy-git-runner-runner-1 rm -rf /root/.cache/act`
(safe — container layer, not a volume). It will rot again. A real fix is either
lowering `capacity` or upstream act; neither was attempted.
## Traps that will waste your time
- **Gitea status codes are not what they look like.** `1 = success, 2 = failure`,
3 cancelled, 4 skipped, 5 waiting, 6 running, 7 blocked. Reading 1/2 as
waiting/running inverts every conclusion you draw from `action_run_job`.
- **Gitea sets `Secure` cookies** (ROOT_URL is https), so a `curl` login against
`http://127.0.0.1:3080` silently keeps no session — it 303s to `/` and you
still get "Sign In". Log in through `https://app.windygit.com`.
- **There is no rerun API in 1.24.6.** `POST /api/v1/.../runs/{n}/rerun` 404s.
Use the web route `POST /{owner}/{repo}/actions/runs/{n}/rerun` with the session
cookie plus an `X-Csrf-Token` header taken from the `_csrf` cookie.
- **`git pull -q` hides errors.** A divergent branch once made a "deploy" run 20
minutes against stale code while reporting success. Use `git fetch && git
merge --ff-only` and read the output.
- **Never force-push a branch a deploy checkout tracks** (`--amend` orphaned
`/srv/windygit/src` once).
- **Gitea's env-to-ini SETS but never UNSETS**, and sometimes *appends a
duplicate*. `GITEA__DEFAULT__APP_NAME` does not work at all — Gitea reads
`APP_NAME` from the **top level** of `app.ini`; the env var creates a literal
`[default]` section it ignores. Edit `app.ini` on the host.
- **Cloudflare caches `/assets/*` for 6h and no token in this stack can purge.**
Version brand asset **filenames** (`theme-windy.v2.css`), not query strings.
- **`base64` wraps at 76 chars** and corrupts long tokens in test commands →
`curl (43)`, phantom HTTP 000. Use `base64 -w0`.
- **Kit 0 is fragile.** 54 containers on 4 vCPU. Two production incidents in two
days, both from *non-production* workloads. Check `uptime` before deploying
anything there, and build before recreating so the swap is seconds.
- **Service containers**: use the service NAME and its INTERNAL port (5432),
never the mapped host port. The three repos did NOT share one pattern — a naive
`localhost` → `postgres` swap would have left WindyCloud on port 15432 (it maps
`15432:5432`) and windy-registry on a `job.services.postgres.ports[…]` expression.
## Open items, roughly by value
1. **Decide what the image-building jobs should do on Windy Git** — WindyCloud
`docker` and windy-search `Docker build` (above). The only remaining *forge*
question; it needs a decision, not code.
2. **windy-agent's three `test` jobs were bulk-killed at 22:50:19** after ~43
minutes, mid-suite, with no verdict. Unexplained and not the runner bump.
3. The product-level test failures in the table above.
4. **act's action cache race** — lower `capacity` below 4, or accept re-runs.
5. **Get non-prod work off Kit 0.** 12 dev/demo containers on the box running
identity, the CA, mail, Matrix and the broker. Cost two incidents already;
the postgres-adapter fix would not have prevented either.
6. **Login is ~4–6s** — `postgres-adapter.ts:114` forks a `node -e` process per
query. Measured: node startup alone is 1.7s on Kit 0 vs 0.01s on Veron. The
fix is **one function** (persistent worker + `pg.Pool`), not the "468 call
sites" the SOTU scoped. See `docs/incidents/2026-08-12-login-latency-analysis.md`.
7. **Privileged dind sits beside broad-scoped tokens** on the CI host — Grant's
call, needs a decision not a code change.
8. Push-velocity throttling is declared but unenforceable from our plane (git
push never touches the API); needs a Gitea pre-receive hook.
## Read these first
- `~/.claude/.../memory/project_windy_git.md` — the full record, densest source
- `DNA_STRAND_MASTER_PLAN.md` — D-1…D-9 locked decisions, I-1…I-13 invariants
- `docs/AUDIT-fable-2026-08-13.md` — second-auditor findings and dispositions
- `docs/CUTOVER.md` — the GitHub↔Windy Git migration plan and its one rule
---
## Copy-paste prompt
```
Picking up Windy Git (agent-native code+model host on Veron 1, live at
app.windygit.com). Read these before doing anything:
1. ~/.claude/projects/-home-grantwhitmer/memory/project_windy_git.md
2. ~/windy-git/docs/TURNOVER-2026-08-14.md
3. ~/windy-git/DNA_STRAND_MASTER_PLAN.md (D-1..D-9, I-1..I-13)
State: live and in use. Grant signs in with his existing Windy account. Agents
authenticate with real EPT signature verification. 143 repos, 85 tests green.
TWO CI root causes are SOLVED and verified:
(a) setup-uv v4+ resolved uv's "latest" through GITHUB_API_URL, which
act_runner points at our own forge, so it 404'd ("The target couldn't be
found.") and every job died at step 2. Fixed by pinning an explicit uv
version across 11 repos.
(b) act_runner 0.2.11 predates `runs.using: node24`, so any repo on
actions/checkout@v5 died before its first step. Bumped to 0.6.1.
windy-mind is fully green (742 passed). Windy-Clone went 4/4 red to 4/4 green.
TASK: one decision, then cleanup.
1. DECIDE what the image-building CI jobs should do here — WindyCloud `docker`
and windy-search `Docker build`. They need a Docker daemon; job containers
deliberately have no socket (I-5 — mounting the host socket hands every
workflow root on Veron 1). Options: buildx inside the existing dind, a
rootless builder, or exclude the job. Do NOT mount the host socket.
2. windy-agent's three `test` jobs were bulk-killed together at 22:50:19 after
~43 min, mid-suite, with no verdict in the log. Not the 30m runner.timeout,
not the runner bump. Unexplained — worth a look.
3. Small product defects: WindyCloud `lint` (ruff format, 7 files), eternitas
`py-sdk` (pytest not a declared dep), eternitas `test`.
Ground rules already paid for the hard way:
- Gitea job status: 1=SUCCESS, 2=FAILURE, 5=waiting, 6=running. Not what you'd guess.
- When a job fails with an opaque forge-shaped error, read the FORGE access log
(`docker logs windy-git-gitea-1 | grep " 404 "`) — act misattributes the error
to the previous step.
- Job logs: `GET /api/v1/repos/{owner}/{repo}/actions/jobs/{id}/logs`. The job id
must belong to the repo in the path or you get a misleading "job not found".
Logs are in R2, not on disk.
- Log into the forge over https://app.windygit.com — Gitea's cookies are Secure,
so a curl login to http://127.0.0.1:3080 silently keeps no session.
- verify the WHOLE flow, not the half that curls easily
- never `git pull -q` in a deploy path; it hides errors
- fixes go to GitHub, not Windy Git (sync is GitHub -> Windy Git, force-push)
- if a job fails with `lstat .../<file>: no such file or directory` on an
action, act's cache rotted: `docker exec windy-git-runner-runner-1 rm -rf
/root/.cache/act`, then re-run. Safe; it is a container layer, not a volume.
- check Kit 0's `uptime` before deploying there; two incidents in two days
```

View File

@@ -0,0 +1,118 @@
# Incident — account.windyword.ai login outage, 2026-08-12
**Found while adding a dashboard tile. Not caused by Windy Git.**
Recorded here because Windy Git is the ecosystem's verification cell and this is
the failure mode its CI exists to catch earlier.
## Impact
`POST /api/v1/auth/login` timed out for at least ~50 minutes (45 s+, no
response). `/health` and `/.well-known/jwks.json` also timed out for part of it.
This is the identity service every Windy product authenticates against — Word,
Chat, Mail, Cloud, Eternitas hand-offs, and Windy Git's own OIDC.
## Root cause — two faults multiplying
**1. A hot retry loop.** `windy-agent-roster` called
`POST /api/v1/identity/mail/address-by-windy-id` for every identity it knows,
got a timeout, retried immediately, and never backed off. Measured: **74
failures/minute, 2,220 in 30 minutes**, sustained. Every log line read
`Pro mail-lookup failed for <uuid>: The operation was aborted due to timeout`.
**2. Every one of those calls is expensive.** `account-server/src/db/postgres-adapter.ts:114`
runs `execFileSync` with a **new `node -e` child process per query** — new pg
client, new TCP connection, new TLS handshake, blocking the event loop up to
30 s. Confirmed live: a host process reading
`node -e const { Client } = require('pg'); ...`. Each 404 cost 0.6–2.1 s of
blocking work.
**The multiplication is the story.** Either alone is survivable. Together they
form a positive feedback loop: slow responses cause timeouts, timeouts cause
immediate retries, retries add load, load makes responses slower. The listening
socket inside the container showed **`Recv-Q 510`** — 510 connections accepted
by the kernel that node was too blocked to pick up. Login sat in that queue.
Background condition: **54 containers on a 4-vCPU box**, 12 of them dev/demo
(`eternitas-dev-*`, `windymind-dev-*`, `nacholos-demo-*`, `nachope-demo-*`).
Baseline load 17–19. That is the amplifier — a process fork is cheap on an idle
box and ruinous on a saturated one.
## Resolution — two steps
**1. Stop the bleeding.** `docker stop windy-agent-roster` at 16:50:57Z. Login
recovered from timeout to HTTP 200 immediately.
**2. Ship the real fix and bring agent chat back.** The retry fix already
existed: a parallel session diagnosed the same incident from the windy-pro side
and landed `c79f196` — *"fix(roster): back off failed mail lookups instead of
retrying every 30s forever" (#172)* — with a test
(`services/agent-roster/tests/mail-lookup-backoff.test.js`). **Kit 0 was one
commit behind and did not have it.** That is the whole reason the loop ran.
Deployed by fast-forwarding `/root/windy-chat` `ac61db6 → c79f196` and
rebuilding **only** `agent-roster` (`--no-deps`; nothing else on the box
touched). Deliberately `git merge --ff-only`, never `reset --hard` — this
checkout has a documented history of local edits a hard reset would silently
eat.
The fix caches failures per owner with exponential backoff (60 s → 30 min cap)
and suppresses logging after three attempts. Observed working in production:
`attempt 2, backing off 120s`.
| | before | after |
|---|---|---|
| login | timeout at 45 s+ | **HTTP 200 in ~18 s** |
| `/health` | timeout | 200 in 0.28 s |
| jwks | timeout | 200 in 0.16 s |
| **account-server CPU** | **168–210 %** | **0.00 %** |
| roster mail-lookup failures | 74/min | **0/min** |
| account-server calls to that route | ~50/min | **0/min** |
| agent chat | down (stopped) | **back up, healthy** |
**18 s is restored, not healthy.** A login should be well under a second. That
number is the fork-per-query adapter under a loaded box, and it is what remains
after the loop was removed.
## What is still true
- **The mail lookups still fail** — the backoff stops them amplifying, it does
not make them succeed. Each owner now retries once per backoff window instead
of every 30 s. The underlying 404/timeout is unexplained. The route exists
(`identity.ts:1014-1022`, reads `x-service-token`). Either the identities are
genuinely absent or the caller sends the wrong shape — worth knowing before
the roster returns.
- The postgres adapter is unchanged. It is the single largest stability
liability in the ecosystem, 468 call sites, and a multi-week job.
## What would have caught this
Nothing did. There is no alerting on account-server latency, and the fleet
canary has been dead since 2026-07-03. The service was `(unhealthy)` with a
**failing healthcheck streak of 74** and nothing said so.
## Recommendations, ordered
1. ~~Fix the retry~~ — **done**, `c79f196`, deployed 2026-08-12.
2. **Alert on the healthcheck.** A 74-deep failing streak on the identity
service should page, not sit.
3. **Move dev/demo off the production box.** 12 containers of non-production
load on the box that runs identity, the CA, mail, Matrix and the broker.
4. **Then** the postgres-adapter migration. Hottest paths first — login is the
obvious first path.
## The lesson worth keeping
**The fix was written, tested, reviewed and merged — and the outage happened
anyway, because Kit 0 was one commit behind.** A merged fix that has not reached
production is not a fix; it is a belief.
That is the same root cause both August audits named — *nothing anywhere checks
whether a decision reached production* — arriving as a live outage rather than a
finding in a report. Deploy verification is not paperwork.
## Diagnostic note for whoever is next
`docker logs` / `docker stats` loops across 54 containers cost real CPU on a
saturated box. Load rose from 17 to 24.9 while this was being investigated, and
some of that was the investigation. Take one clean measurement, then back off.

View File

@@ -0,0 +1,103 @@
# Why login takes 18 seconds — measured, 2026-08-12
The SOTU calls the child-process DB bridge *"the single largest stability
liability under #8"* and scopes the fix as **468 call sites, multi-week**.
That scoping is wrong, and the measurements below say so. **You do not need to
touch 468 call sites.** You need to stop spawning a node process per query,
which is one function.
## The measurements
**One login makes 9 synchronous queries**, each spawning `node -e` via
`execFileSync` (`postgres-adapter.ts:114`):
| # | query | source |
|---|---|---|
| 1 | `findUserByEmail` | `statements.ts:11` |
| 2 | `mfa_secrets` lookup | `auth.ts` inline |
| 3 | `findDevice` | `statements.ts:24` |
| 4 | `touchDevice` *or* `countDevices` + `addDevice` | `statements.ts:25-35` |
| 5 | `updateUserSeen` | `statements.ts:20` |
| 6 | `generateTokens` → scope rows | `auth.ts:220` |
| 7 | `generateTokens` → product rows | `auth.ts:229` |
| 8 | `getDeviceList` | `auth.ts:384` |
| 9 | `logAuditEvent` INSERT | `identity-service.ts:43` |
**What each fork actually costs, measured inside the running container:**
| | Kit 0 (4 vCPU, load ~20) | Veron 1 (24 cores, idle) |
|---|---|---|
| bare `node -e "0"` | **1.70 – 1.99 s** | **0.01 s** |
| node + pg connect + `SELECT 1` | 1.89 – 3.19 s | — |
**Node startup is 170× slower on Kit 0, and it is essentially the entire cost.**
Adding a Postgres connect and a real query to a bare node start adds only
~0.2–1.2 s on top of ~1.8 s of interpreter boot.
9 forks x ~1.8 s = ~16 s. Observed login: 17–25 s.
## The conclusion that matters
There are **two independent multipliers**, and they compound:
1. **The adapter forks a process per query** — 9x on the login path.
2. **The box is saturated**, so each fork costs 1.8 s instead of 0.01 s — 170x.
Either one alone is survivable. Together they turn a sub-second operation into
eighteen seconds, and on 2026-08-12 they turned a retry loop into an
ecosystem-wide auth outage.
**Connection pooling barely helps.** The connection is not the cost; the
interpreter boot is. PgBouncer, `pg.Pool` on the sync path, or a warmer
Postgres would all leave ~1.8 s per query untouched.
## Box census (measured same day)
54 containers, 301% CPU of 400% available, load average 20 on 4 vCPU
| | containers | CPU |
|---|---|---|
| dev / demo / test | 12 | **55%** |
| everything else | 42 | 246% |
Top consumers: `windymail-migrate-stalwart` 31%, `scenemachine-db-prod` 30%,
`account-server-account-postgres` 27%, `windy-synapse` 25%, `windy-directory`
20%.
**Be honest about this number:** stopping every dev/demo container reclaims 55%
of 301% — it takes the box from 75% to 61% steady CPU. That is real relief and
it is *not* a fix. Forty-two non-dev containers on four cores is the actual
condition.
## What to do, in value order
1. **Stop forking a node process per query.** `querySyncViaChild`
(`postgres-adapter.ts:114`) is one function, and its interface —
`querySync(sql, params)` — does not change. Replace the per-query
`execFileSync` with a persistent worker holding a `pg.Pool`, using
`worker_threads` + `SharedArrayBuffer` + `Atomics.wait` for the synchronous
block. **Every one of the 468 call sites gets faster without being edited**,
including the mail-lookup route that caused today's outage.
Expected: login 18 s → well under 1 s, on this box, unchanged.
2. **Then** migrate route families to the async path at leisure. The sync path
is still labelled legacy and should still die — but as cleanup, not as an
emergency.
3. **Separately, unload Kit 0.** Move dev/demo off the box that runs identity,
the CA, mail, Matrix and the broker. This is worth doing on its own merits
regardless of the adapter.
**Do not do #2 first.** Migrating handlers one family at a time is weeks of
edits to the most critical code in the ecosystem, and it leaves every
un-migrated call site paying 1.8 s a query the whole time.
## Why this was not implemented in this session
Replacing the sync bridge is a subtle change (`Atomics.wait`, structured-clone
limits, worker lifecycle, failure fallback) in the single most critical file in
the ecosystem, made at the end of a long session, on a box that had already had
one outage that day. It deserves a fresh session, a real load test, and a
rehearsed rollback — not a tired commit.
The measurement is the deliverable. It converts a "multi-week, 468 call sites"
job into a one-function change, and that is worth more than a rushed attempt.

View File

@@ -17,13 +17,21 @@ dependencies = [
"pydantic-settings>=2.6",
"sqlalchemy[asyncio]>=2.0",
"asyncpg>=0.30",
# Alembic runs synchronously (env.py strips +asyncpg), so the image needs a
# sync driver too. Without it migrations fail INSIDE the container while
# passing on a developer machine that happens to have it — the kind of gap
# that only shows up on a fresh deploy.
"psycopg2-binary>=2.9",
"alembic>=1.14",
"httpx>=0.27",
"boto3>=1.35",
# ES256 verification of Eternitas EPTs (G3.2/G9.1). Without crypto extras
# PyJWT cannot verify EC signatures and silently offers no protection.
"pyjwt[crypto]>=2.9",
]
[project.optional-dependencies]
dev = ["pytest>=8.3", "pytest-asyncio>=0.24", "ruff>=0.7", "mypy>=1.13"]
dev = ["pytest>=8.3", "pyyaml>=6.0", "pytest-asyncio>=0.24", "ruff>=0.7", "mypy>=1.13"]
[tool.ruff]
line-length = 100

126
scripts/backup.sh Executable file
View File

@@ -0,0 +1,126 @@
#!/usr/bin/env bash
# Nightly backup (G0.9) — the prerequisite for Windy Git becoming the daily driver.
#
# Today GitHub is authoritative, so losing Veron 1 costs nothing. The moment
# people push HERE first, that inverts: Veron 1 holds the only current copy of
# the company's source between mirror syncs, and Veron 1 is Grant's workstation
# — no SLA, no snapshots, a residential line, and he reboots it.
#
# `git bundle` is used deliberately over tarring the repo directory: a bundle is
# a single file that `git clone` reads directly, so a restore is one command and
# needs no knowledge of Gitea's on-disk layout. Tarring a live repo directory
# also races with a concurrent push; bundling asks git for a consistent view.
#
# The whole archive measured 1.58 GB across 141 repos, so this costs about two
# cents a month on R2 and takes minutes. There is no reason for it not to exist.
set -uo pipefail
STAMP="$(date -u +%Y-%m-%d)"
WORK="$(mktemp -d /tmp/windygit-backup-XXXXXX)"
GIT_ROOT="${GIT_DATA_ROOT:-/srv/windygit/git}/git/repositories"
BUCKET="${R2_BUCKET_BACKUPS:-windy-git-backups}"
KEEP_DAYS="${BACKUP_KEEP_DAYS:-30}"
FAILED=0
# NEVER bundle these to R2 (orchestrator decision 2026-09-23). They carry
# credentials in plaintext — kit-army-config IS the lockbox, and the soul repos
# hold agent memory with keys in it — and these bundles are unencrypted, so
# anyone holding the R2 key could read every secret in the fleet. They are
# backed up ENCRYPTED elsewhere (Windy Drops lane, restic, restore-tested) and
# stay mirrored on Veron's own disk in Gitea. Extended globs, matched on name.
EXCLUDE="${BACKUP_EXCLUDE:-kit-army-config anima *-soul}"
excluded() {
local n=$1 pat pats
read -ra pats <<< "$EXCLUDE" # read never glob-expands; `for p in $EXCLUDE` would
for pat in "${pats[@]}"; do
# shellcheck disable=SC2053 # unquoted RHS: glob match is the point
[[ "$n" == $pat ]] && return 0
done
return 1
}
cleanup() { rm -rf "$WORK"; }
trap cleanup EXIT
log() { printf '[backup %s] %s\n' "$(date -u +%H:%M:%SZ)" "$*"; }
if [[ -z "${R2_ACCESS_KEY_ID:-}" || -z "${R2_SECRET_ACCESS_KEY:-}" ]]; then
log "FATAL: R2 credentials unset — refusing to report a backup that did not happen"
exit 1
fi
export AWS_ACCESS_KEY_ID="$R2_ACCESS_KEY_ID"
export AWS_SECRET_ACCESS_KEY="$R2_SECRET_ACCESS_KEY"
export AWS_DEFAULT_REGION=auto
S3="aws s3 --endpoint-url https://${R2_ACCOUNT_ID}.r2.cloudflarestorage.com"
# ---- 1. every repo, as a restorable bundle -------------------------------
shopt -s nullglob
count=0
for repo in "$GIT_ROOT"/*/*.git; do
owner="$(basename "$(dirname "$repo")")"
name="$(basename "$repo" .git)"
if excluded "$name"; then
log "skip ${owner}/${name} (credential-bearing: never bundled to R2 in plaintext)"
continue
fi
out="$WORK/${owner}__${name}.bundle"
# --all captures every ref, not just the default branch. A bundle of one
# branch silently loses every other branch and every tag, and you find out
# during the restore.
if git --git-dir="$repo" bundle create "$out" --all >/dev/null 2>&1; then
# Verify before trusting. An unverified bundle is a belief, not a backup.
if git bundle verify "$out" >/dev/null 2>&1; then
count=$((count + 1))
else
log "CORRUPT bundle for ${owner}/${name} — not uploading"
rm -f "$out"; FAILED=1
fi
else
# An empty repo has no refs and cannot be bundled. That is normal, not a
# failure — say so rather than counting it as an error.
if [[ -z "$(git --git-dir="$repo" for-each-ref 2>/dev/null)" ]]; then
log "skip ${owner}/${name} (empty repo, no refs)"
else
log "FAILED to bundle ${owner}/${name}"; FAILED=1
fi
rm -f "$out"
fi
done
log "bundled $count repos"
# ---- 2. the plane's own database ------------------------------------------
# Postgres is truth for repos, grants, versions, tokens and mirror state. The
# bundles restore the code; this restores who may touch it.
if docker exec windy-git-db-1 pg_dump -U windygit -d windygit --schema=windgit \
> "$WORK/windgit.sql" 2>/dev/null && [[ -s "$WORK/windgit.sql" ]]; then
log "dumped windgit schema ($(wc -c < "$WORK/windgit.sql") bytes)"
else
log "FAILED to dump the database"; FAILED=1
fi
# ---- 3. upload ------------------------------------------------------------
if $S3 cp "$WORK" "s3://${BUCKET}/${STAMP}/" --recursive --only-show-errors; then
log "uploaded to s3://${BUCKET}/${STAMP}/"
else
log "FATAL: upload failed"; exit 1
fi
# ---- 4. retention ---------------------------------------------------------
cutoff="$(date -u -d "${KEEP_DAYS} days ago" +%Y-%m-%d 2>/dev/null || true)"
if [[ -n "$cutoff" ]]; then
$S3 ls "s3://${BUCKET}/" | awk '{print $2}' | tr -d '/' | while read -r d; do
[[ "$d" < "$cutoff" ]] && { log "pruning $d"; $S3 rm "s3://${BUCKET}/${d}/" --recursive --only-show-errors; }
done
fi
# Non-zero on ANY failure so the systemd unit goes red and the failure is
# visible. A backup script that swallows errors is worse than none — it
# manufactures confidence.
if [[ "$FAILED" -ne 0 ]]; then
log "COMPLETED WITH FAILURES"
exit 1
fi
log "ok — $count repos + database"

306
scripts/canary.py Executable file
View File

@@ -0,0 +1,306 @@
#!/usr/bin/env python3
"""Fleet canary (G7.6).
**Probes what a user does, not what is cheap to answer.**
That distinction is the entire lesson of the 2026-08-12 outage: `/health`
returned 200 the whole time login was dead. A canary watching `/health` would
have stayed green for an hour while nobody in the ecosystem could sign in. So
every check here names a *user-visible* capability, and the login probe is the
one that matters most.
Three rules this canary obeys:
1. **Never report green for something it did not prove.** A check it could not
run reports `unknown`, never `ok` (I-8).
2. **Alert on transitions, not on every run.** A canary that emails every five
minutes gets filtered, and a filtered canary is a dead canary — which is how
the last one sat 37 days dead without anyone noticing.
3. **Run somewhere the thing being watched cannot take down with it.** This runs
on Veron 1 via Windy Git CI. A canary hosted on Kit 0 would die with Kit 0
and report nothing at the exact moment it mattered.
State lives in a small JSON file so consecutive runs can tell "still broken"
from "just broke".
"""
from __future__ import annotations
import argparse
import json
import os
import sys
import time
import urllib.error
import urllib.request
from dataclasses import dataclass, field
STATE_PATH = os.environ.get("CANARY_STATE", "canary-state.json")
RESEND_KEY = os.environ.get("RESEND_API_KEY", "")
ALERT_TO = os.environ.get("CANARY_ALERT_TO", "grantwhitmer3@gmail.com")
ALERT_FROM = os.environ.get("CANARY_ALERT_FROM", "office@thewindstorm.uk")
# Login is slow because account-server forks a node process per query. 18-25s is
# today's reality, not health. The threshold flags a real regression without
# crying wolf about the known-slow baseline; lower it as the adapter is fixed.
LOGIN_WARN_SECONDS = float(os.environ.get("CANARY_LOGIN_WARN_S", "35"))
TIMEOUT = float(os.environ.get("CANARY_TIMEOUT_S", "60"))
@dataclass
class Result:
name: str
status: str # ok | down | slow | unknown
detail: str
seconds: float = 0.0
user_visible: str = ""
@dataclass
class Check:
name: str
url: str
what_it_proves: str
method: str = "GET"
body: dict | None = None
headers: dict = field(default_factory=dict)
warn_seconds: float | None = None
# When True this check INVERTS: a 2xx is a critical failure (a security
# control opened) and a 401/403/503 is the healthy, expected outcome.
must_refuse: bool = False
def _probe(c: Check) -> Result:
data = json.dumps(c.body).encode() if c.body else None
headers = {"User-Agent": "windy-git-canary/1.0", **c.headers}
# Our own probes are synthetic traffic (ecosystem convention, Telemetry
# UPDATE 4): every service they touch labels the resulting rows.
headers["X-Windy-Synthetic"] = "1"
if data:
headers["Content-Type"] = "application/json"
req = urllib.request.Request(c.url, data=data, method=c.method, headers=headers)
start = time.monotonic()
try:
with urllib.request.urlopen(req, timeout=TIMEOUT) as r:
elapsed = time.monotonic() - start
if c.must_refuse:
# A 2xx here means a control that should reject accepted. That is
# the alarm, not the absence of one.
return Result(c.name, "down",
f"ACCEPTED (HTTP {r.status}) — this MUST be refused",
elapsed, c.what_it_proves)
if r.status >= 400:
return Result(c.name, "down", f"HTTP {r.status}", elapsed, c.what_it_proves)
warn = c.warn_seconds
if warn and elapsed > warn:
return Result(
c.name, "slow", f"HTTP {r.status} in {elapsed:.1f}s (warn >{warn:.0f}s)",
elapsed, c.what_it_proves,
)
return Result(c.name, "ok", f"HTTP {r.status} in {elapsed:.1f}s", elapsed, c.what_it_proves)
except urllib.error.HTTPError as e:
if c.must_refuse and e.code in (401, 403, 503):
return Result(c.name, "ok", f"correctly refused (HTTP {e.code})",
time.monotonic() - start, c.what_it_proves)
return Result(c.name, "down", f"HTTP {e.code}", time.monotonic() - start, c.what_it_proves)
except Exception as e: # noqa: BLE001 — a probe must never raise upward
return Result(
c.name, "down", f"{type(e).__name__}: {str(e)[:80]}",
time.monotonic() - start, c.what_it_proves,
)
def build_checks() -> list[Check]:
checks = [
Check(
"identity.health",
"https://account.windyword.ai/health",
"the identity service answers at all",
),
Check(
"identity.jwks",
"https://account.windyword.ai/.well-known/jwks.json",
"every service can verify the tokens it is handed",
),
Check(
"eternitas.health",
"https://api.eternitas.ai/health",
"agent passports can be issued and checked",
),
Check(
"windygit.forge",
"https://app.windygit.com/api/v1/version",
"repositories are reachable",
),
Check(
"windygit.plane",
"https://api.windygit.com/version",
"the Windy Git API answers",
),
Check(
"dashboard",
"https://app.windyword.ai/",
"the dashboard loads",
),
]
# SECURITY REGRESSION GUARD. A forged, unsigned token naming a real passport
# must be refused. On 2026-08-13 this returned HTTP 200 (full agent
# impersonation). If it ever returns 2xx again, the bypass is back.
import base64 as _b64
import json as _j
def _seg(d: dict) -> str:
return _b64.urlsafe_b64encode(_j.dumps(d).encode()).rstrip(b"=").decode()
# Two shapes, because they exercise two different gates. The EPT-shaped one
# is the important one now: it is what real signature verification guards.
for _label, _hdr in (
("security.forged_agent_token", {"alg": "none", "typ": "JWT"}),
("security.forged_ept", {"alg": "none", "typ": "EPT"}),
):
_tok = (
f"{_seg(_hdr)}."
f"{_seg({'sub': 'ET26-1EF9-VJAN', 'passport': 'ET26-1EF9-VJAN', 'iss': 'eternitas.ai', 'exp': 9999999999})}"
".not-a-real-signature"
)
checks.append(
Check(
_label,
"https://api.windygit.com/api/v1/repos",
"an unsigned token cannot impersonate an agent",
headers={"Authorization": f"Bearer {_tok}"},
must_refuse=True,
)
)
# THE important one. /health was 200 for the entire 2026-08-12 outage while
# this was timing out. A canary that skips it is decorative.
pw = os.environ.get("CANARY_LOGIN_PASSWORD", "")
email = os.environ.get("CANARY_LOGIN_EMAIL", "")
if pw and email:
checks.append(
Check(
"identity.login",
"https://account.windyword.ai/api/v1/auth/login",
"a human can actually sign in",
method="POST",
body={"email": email, "password": pw},
warn_seconds=LOGIN_WARN_SECONDS,
)
)
return checks
def load_state() -> dict:
try:
with open(STATE_PATH) as f:
return json.load(f)
except (FileNotFoundError, json.JSONDecodeError):
return {}
def save_state(results: list[Result]) -> None:
"""Never let bookkeeping kill the monitor.
State is an optimisation — it lets the next run tell "still broken" from
"just broke". The probing is the valuable part. An unwritable path used to
raise here and take the whole canary down, which is the worst possible
trade: a monitoring tool that dies of a config problem reports nothing at
all, and reports it silently.
"""
try:
with open(STATE_PATH, "w") as f:
json.dump({r.name: r.status for r in results}, f, indent=2)
except OSError as exc:
print(f"!! could not save state to {STATE_PATH}: {exc}")
print(" (probes still ran; transition detection is degraded this run)")
def send_alert(subject: str, lines: list[str]) -> bool:
if not RESEND_KEY:
print("!! RESEND_API_KEY unset — cannot alert. This canary is decorative.")
return False
body = "\n".join(lines)
req = urllib.request.Request(
"https://api.resend.com/emails",
data=json.dumps({
"from": f"Windy Canary <{ALERT_FROM}>",
"to": [ALERT_TO],
"subject": subject,
"text": body,
}).encode(),
method="POST",
headers={
"Authorization": f"Bearer {RESEND_KEY}",
"Content-Type": "application/json",
# ⚠️ REQUIRED. Without an explicit User-Agent, urllib sends
# "Python-urllib/3.x" and the request is rejected 403 by bot
# filtering — while the identical request via curl succeeds. This
# exact failure was caught by testing the alert path rather than
# assuming it: the canary would have detected every outage
# correctly and told nobody.
"User-Agent": "windy-git-canary/1.0",
},
)
try:
with urllib.request.urlopen(req, timeout=30) as r:
print(f" alert sent ({r.status})")
return True
except urllib.error.HTTPError as e:
# Print the body. "403 Forbidden" alone sends you hunting for a bad key;
# the body usually names the real cause.
print(f"!! alert FAILED: HTTP {e.code}: {e.read().decode()[:200]}")
return False
except Exception as e: # noqa: BLE001
print(f"!! alert FAILED: {e}")
return False
def main() -> int:
ap = argparse.ArgumentParser()
ap.add_argument("--no-alert", action="store_true")
args = ap.parse_args()
previous = load_state()
results = [_probe(c) for c in build_checks()]
print(f"windy canary — {time.strftime('%Y-%m-%dT%H:%M:%SZ', time.gmtime())}\n")
for r in results:
mark = {"ok": " ok ", "slow": " SLOW ", "down": " DOWN ", "unknown": " ?? "}[r.status]
print(f"[{mark}] {r.name:20} {r.detail}")
if r.status != "ok":
print(f" ^ this means: {r.user_visible}")
# Transitions only. "Still broken" does not re-alert; recovery does.
newly_bad = [r for r in results if r.status in ("down", "slow") and previous.get(r.name) == "ok"]
recovered = [
r for r in results
if r.status == "ok" and previous.get(r.name) in ("down", "slow")
]
save_state(results)
if not args.no_alert:
if newly_bad:
worst = "DOWN" if any(r.status == "down" for r in newly_bad) else "SLOW"
send_alert(
f"[Windy] {worst}: {', '.join(r.name for r in newly_bad)}",
[f"{r.name}: {r.detail}" for r in newly_bad]
+ ["", "What this means for a person:"]
+ [f" - {r.user_visible}" for r in newly_bad]
+ ["", "Checked from Veron 1 via Windy Git CI — deliberately not from Kit 0."],
)
if recovered:
send_alert(
f"[Windy] recovered: {', '.join(r.name for r in recovered)}",
[f"{r.name}: {r.detail}" for r in recovered],
)
# A failure exit makes the CI run red, so the forge itself carries the signal
# even if email is misconfigured. Two independent ways to notice.
return 1 if any(r.status in ("down", "slow") for r in results) else 0
if __name__ == "__main__":
sys.exit(main())

10
scripts/cancel_unrunnable.sh Executable file
View File

@@ -0,0 +1,10 @@
#!/usr/bin/env bash
# Cancel jobs no runner can ever take (see cancel_unrunnable.sql). Run on Veron as root.
set -euo pipefail
SPOOL="${JANITOR_SPOOL:-/var/lib/windy-git/janitor-cancelled.jsonl}"
mkdir -p "$(dirname "$SPOOL")"
out=$(docker exec -i windy-git-db-1 sh -c 'psql -U "$POSTGRES_USER" -d gitea -At -v ON_ERROR_STOP=1' \
< "$(dirname "$0")/cancel_unrunnable.sql")
printf '%s\n' "$out" | grep '^{' >> "$SPOOL" || true
n=$(printf '%s\n' "$out" | grep -c '^{' || true)
echo "[janitor] cancelled ${n} unrunnable job(s)"

View File

@@ -0,0 +1,53 @@
-- Cancel CI jobs that can never run (called by scripts/cancel_unrunnable.sh).
--
-- A job whose runs-on names a label no Windy Git runner offers (ubuntu-latest,
-- macos-latest, windows-latest …) waits forever: Gitea evaluates a job's `if:`
-- only when a runner picks it, so even `if: false` / tag-only jobs sit in the
-- queue, invisible to /actions/tasks, and keep their run "waiting" for good.
-- After 30 minutes they are cancelled here; the run's status is then recomputed
-- (failure > still-active > cancelled > success), the same precedence Gitea uses.
-- Keep RUNNER_LABELS in step with deploy/runner/config.yaml.
BEGIN;
WITH dead AS (
UPDATE action_run_job j
SET status = 3, stopped = extract(epoch from now())::bigint, updated = extract(epoch from now())::bigint
WHERE j.status IN (5, 7)
AND to_timestamp(j.created) < now() - interval '30 minutes'
AND EXISTS (SELECT 1 FROM jsonb_array_elements_text(j.runs_on::jsonb) l
WHERE l NOT IN ('veron-1', 'linux-x64', 'self-hosted', 'linux', 'x64'))
RETURNING j.id, j.run_id, j.name, j.runs_on, j.created
), runs AS (
UPDATE action_run r
SET status = CASE
WHEN EXISTS (SELECT 1 FROM action_run_job x WHERE x.run_id = r.id AND x.status = 2) THEN 2
WHEN EXISTS (SELECT 1 FROM action_run_job x WHERE x.run_id = r.id AND x.status IN (5, 6, 7)
AND x.id NOT IN (SELECT id FROM dead)) THEN r.status
ELSE 3 END,
stopped = CASE WHEN r.stopped = 0 THEN extract(epoch from now())::bigint ELSE r.stopped END
WHERE r.id IN (SELECT DISTINCT run_id FROM dead)
RETURNING r.id
)
-- One JSON line per cancelled job: the telemetry emitter ships these as
-- ci.job_cancelled (declared with Telemetry Boss, 2026-09-23).
SELECT json_build_object(
'repo', p.lower_name,
'workflow', regexp_replace(r.workflow_id, '\.ya?ml$', ''),
'job', d.name,
'reason', 'unrunnable_label',
'runs_on', (SELECT string_agg(l, ',') FROM jsonb_array_elements_text(d.runs_on::jsonb) l),
'waited_s', (extract(epoch from now())::bigint - d.created))::text
FROM dead d JOIN action_run r ON r.id = d.run_id JOIN repository p ON p.id = r.repo_id
WHERE (SELECT count(*) FROM runs) >= 0;
-- Jobs BLOCKED on `needs:` inside a run that has already finished (a needed job
-- failed): Gitea leaves them status 7 forever. They were never going to run;
-- mark them skipped (4), which is what GitHub shows for the same situation.
UPDATE action_run_job j
SET status = 4, updated = extract(epoch from now())::bigint
FROM action_run r
WHERE r.id = j.run_id
AND j.status = 7
AND r.status IN (1, 2, 3)
AND to_timestamp(j.created) < now() - interval '30 minutes'
RETURNING j.run_id;
COMMIT;

247
scripts/import_from_github.py Executable file
View File

@@ -0,0 +1,247 @@
#!/usr/bin/env python3
"""Import a GitHub repo into Windy Git (strand G11.3 / G7.4).
**Two modes, and the choice is not cosmetic.**
--mirror GitHub ──pull──▶ Windy Git read-only, NO CI
(default) Windy Git ──push-mirror──▶ GitHub writable, CI RUNS
**A pull mirror cannot run CI.** Measured 2026-08-12: windy-calendar imported as
a mirror sat at **0 workflow runs**. Gitea does not fire Actions on mirror sync,
and a mirror is not a push target — so a mirrored repo gives you the code and
none of the point.
So getting CI onto Veron 1 requires **real, writable repos**, which means Windy
Git is where you push and GitHub becomes the downstream copy via I-4's
push-mirror (`api/app/services/mirror.py`). That is the planned steady state,
not a shortcut — but it *is* a change to where every human and agent pushes, so
it is Grant's call, not this script's default assumption.
`--mirror` remains available for repos you want copied but not moved.
Usage:
./import_from_github.py windy-calendar [windy-mind ...]
./import_from_github.py --list-candidates
"""
from __future__ import annotations
import argparse
import json
import os
import subprocess
import sys
import urllib.error
import urllib.request
# ⚠️ Default to the HOST-LOCAL address, not the public one.
#
# Cloudflare answers the public endpoint with 403 error 1010 for this client —
# it blocks urllib's user-agent as a bot signature. The failure looks like Gitea
# rejecting the token and is not: the same call against localhost:3080 on the
# host succeeds immediately.
#
# Bulk import belongs on the host anyway: no hairpin through the edge, no
# Cloudflare ~100s proxy ceiling (G4A.5) on a large clone. Run this on Veron 1.
#
# 🔴 Deliberately NOT `GITEA_BASE_URL`: the deploy `.env` sets that to
# `http://gitea:3000` for the API container, and sourcing `.env` on the host
# made this script die on DNS *after* a caller had already deleted the mirror it
# was meant to replace (2026-09-23).
GITEA = os.environ.get("IMPORT_GITEA_URL", "http://localhost:3080")
GITEA_TOKEN = os.environ.get("GITEA_ADMIN_TOKEN", "")
GITHUB_TOKEN = os.environ.get("GITHUB_TOKEN", "")
GITHUB_OWNER = os.environ.get("GITHUB_OWNER", "sneakyfree")
OWNER = os.environ.get("WINDYGIT_OWNER", "windyadmin")
# G11.4 — least risk first. windy-pro is deliberately last and deliberately not
# in this list: six checkouts exist, the build counter has forked three ways, and
# there is direct evidence conflict on which HEAD is current. Resolve that by
# reading, not by importing (G11.5).
SAFE_ORDER = [
# scaffolds and sites — nothing depends on them
"windy-calendar",
"windy-search",
"windy-registry",
"Windy-Clone",
# live services with real test suites
"WindyCloud",
"windy-cloud-sites",
"windy-mind",
"eternitas",
"windy-agent",
]
def _api(method: str, path: str, body: dict | None = None) -> tuple[int, dict]:
if not GITEA_TOKEN:
sys.exit("GITEA_ADMIN_TOKEN is unset. Refusing to guess.")
data = json.dumps(body).encode() if body is not None else None
req = urllib.request.Request(
f"{GITEA}/api/v1{path}",
data=data,
method=method,
headers={
"Authorization": f"token {GITEA_TOKEN}",
"Content-Type": "application/json",
},
)
try:
with urllib.request.urlopen(req, timeout=180) as r:
raw = r.read().decode()
return r.status, (json.loads(raw) if raw.strip() else {})
except urllib.error.HTTPError as e:
raw = e.read().decode()
try:
return e.code, json.loads(raw)
except json.JSONDecodeError:
return e.code, {"message": raw[:300]}
def all_repo_names() -> list[str]:
out = subprocess.run(
["gh", "repo", "list", GITHUB_OWNER, "--limit", "300",
"--json", "name,isArchived"],
capture_output=True, text=True, check=True,
)
return sorted(r["name"] for r in json.loads(out.stdout) if not r["isArchived"])
def import_everything_as_mirrors() -> int:
"""Bulk DR copy: every repo, as a READ-ONLY pull mirror.
Mirrors are the right shape for the bulk, and the reason is safety rather
than tidiness. Sampling 40 repos found **18 carrying deploy / release /
publish workflows that trigger on `push:`** — roughly 63 across the account.
Importing those as writable repos with Actions enabled would arm sixty-odd
production deploy triggers on Veron 1, each of which would then have to be
disarmed by hand.
A pull mirror cannot run Actions at all, so the bulk import carries **zero**
deploy risk, and Gitea does the syncing itself with no script and no timer.
What you get is a complete, current, second copy of the whole account.
Converting one to a writable CI repo is then a deliberate per-repo act:
delete, re-import with `mirror=false`, review its workflows, disable the
deploying ones. That is the moment to make that judgement — not in bulk,
sixty times, by accident.
"""
names = all_repo_names()
existing = 0
done = 0
failed = []
print(f"{len(names)} active repos on GitHub. Importing missing ones as read-only mirrors.\n")
for n in names:
status, _ = _api("GET", f"/repos/{OWNER}/{n}")
if status == 200:
existing += 1
continue
if import_repo(n, mirror=True):
done += 1
else:
failed.append(n)
print(f"\n already present: {existing}")
print(f" newly mirrored: {done}")
if failed:
print(f" FAILED ({len(failed)}): {', '.join(failed[:10])}")
return 1 if failed else 0
def list_candidates() -> None:
"""Private repos whose CI cannot run on GitHub at all."""
out = subprocess.run(
["gh", "repo", "list", GITHUB_OWNER, "--limit", "200",
"--json", "name,isPrivate,isArchived,pushedAt"],
capture_output=True, text=True, check=True,
)
repos = [r for r in json.loads(out.stdout) if r["isPrivate"] and not r["isArchived"]]
print(f"{len(repos)} private repos — GitHub Actions cannot run on any of them.\n")
for name in SAFE_ORDER:
match = next((r for r in repos if r["name"] == name), None)
print(f" {name:24} {'last push ' + match['pushedAt'][:10] if match else 'NOT FOUND'}")
print("\nwindy-pro is excluded on purpose — see G11.5.")
def import_repo(name: str, mirror: bool = False) -> bool:
existing_status, _ = _api("GET", f"/repos/{OWNER}/{name}")
if existing_status == 200:
print(f" {name:24} already present — skipping (idempotent)")
return True
if not GITHUB_TOKEN:
sys.exit("GITHUB_TOKEN is unset; private repos cannot be read without it.")
status, body = _api(
"POST",
"/repos/migrate",
{
"clone_addr": f"https://github.com/{GITHUB_OWNER}/{name}.git",
"auth_token": GITHUB_TOKEN,
"repo_name": name,
"repo_owner": OWNER,
"service": "github",
"private": True,
# Pull mirror: GitHub stays upstream for the migration quarter.
# A pull mirror cannot diverge — worst case it is stale, not wrong.
"mirror": mirror,
**({"mirror_interval": "10m"} if mirror else {}),
# Issues/PRs/releases deliberately NOT imported. They are GitHub's
# copy of a conversation, and duplicating conversations across two
# systems is how you end up with two half-answers to every question.
"issues": False,
"pull_requests": False,
"releases": False,
"wiki": False,
"labels": False,
"milestones": False,
},
)
if status in (200, 201):
print(f" {name:24} imported ({body.get('size', 0)} KB)")
return True
print(f" {name:24} FAILED {status}: {str(body.get('message'))[:120]}")
return False
def main() -> int:
ap = argparse.ArgumentParser()
ap.add_argument("repos", nargs="*")
ap.add_argument("--list-candidates", action="store_true")
ap.add_argument("--safe-batch", action="store_true", help="import SAFE_ORDER in order")
ap.add_argument(
"--mirror", action="store_true",
help="import read-only pull mirrors instead of writable repos. NOTE: mirrors CANNOT run CI.",
)
ap.add_argument(
"--all-as-mirrors", action="store_true",
help="bulk DR copy: every active repo on the account, read-only, no CI, zero deploy risk",
)
args = ap.parse_args()
if args.list_candidates:
list_candidates()
return 0
if args.all_as_mirrors:
return import_everything_as_mirrors()
targets = SAFE_ORDER if args.safe_batch else args.repos
if not targets:
ap.error("name a repo, or pass --safe-batch / --list-candidates")
# G11.5 RESOLVED 2026-09-23 (lane 8c, ~/windy-orchestra/WINDYPRO_CHECKOUTS.md):
# a read-only audit of all 14 windy-pro checkouts on 5 machines found GitHub
# main is canonical (Kit 0 prod and Windy 0 sit exactly on it; the others are
# stale, not divergent). Phase 1 keeps GitHub the source of truth anyway, so a
# writable Windy Git copy is CI only. Its six deploy/release workflows must be
# disabled on import — see docs/CUTOVER.md.
if args.mirror:
print("mirror mode: repos will be read-only and will NOT run CI.\n")
ok = sum(import_repo(r, mirror=args.mirror) for r in targets)
print(f"\n{ok}/{len(targets)} imported.")
return 0 if ok == len(targets) else 1
if __name__ == "__main__":
sys.exit(main())

307
scripts/pr_status_bridge.py Executable file
View File

@@ -0,0 +1,307 @@
#!/usr/bin/env python3
"""Give private GitHub repos a CI signal from Windy Git (P1, 2026-09-23).
GitHub Actions cannot run on private repos on this account — not even on
self-hosted runners ([[reference-github-actions-billing-lock]]). The code is
already synced into Windy Git every 15 min and CI runs there, so the only thing
missing is the *signal on GitHub*, where people and agents actually read PRs.
Two jobs, run after every sync:
1. **Mirror open PRs.** The sync carries branches, not PRs, and the workflows
trigger on `pull_request` — so a PR branch alone fires nothing. For each open
same-repo GitHub PR we keep one open Windy Git PR with the same head/base.
Gitea then fires `pull_request` on open and `synchronize` whenever the sync
moves the branch. Windy Git PRs whose GitHub PR closed are closed here too.
Fork PRs are ignored: their head branch is never synced, and untrusted fork
code on this runner is exactly the blast radius the audit warned about.
2. **Post results back** as GitHub commit statuses (context
`windy-git/<workflow>/<job>`) on each PR head and on the default-branch head.
Only posts when a context's state changed, so a 15-min loop doesn't pile
hundreds of identical statuses onto one commit.
Runs ON Veron 1 (localhost Gitea; no Cloudflare hairpin). Needs
GITEA_ADMIN_TOKEN and a GITHUB_TOKEN with `repo` scope. Nothing here executes
repo code, and no secret is handed to any repo.
"""
from __future__ import annotations
import base64
import json
import os
import re
import sys
import time
import urllib.error
import urllib.request
import yaml
GITEA = os.environ.get("BRIDGE_GITEA_URL", "http://localhost:3080").rstrip("/")
PUBLIC = "https://app.windygit.com"
GITEA_TOKEN = os.environ.get("GITEA_ADMIN_TOKEN", "")
GITHUB_TOKEN = os.environ.get("GITHUB_TOKEN", "")
GH_OWNER = os.environ.get("GITHUB_OWNER", "sneakyfree")
WG_OWNER = os.environ.get("WINDYGIT_OWNER", "windyadmin")
# Private repos only. Public repos run real GitHub Actions on veron1's GitHub
# runner; bridging those too would put two competing verdicts on every commit.
REPOS = os.environ.get(
"BRIDGE_REPOS",
"windy-chat windy-mail windy-calendar Windy-Clone WindyCloud windy-search windy-connect"
" windy-drops windy-code-web windy-code windy-traveler windy-registry eternitas"
" windy-translate windytranslate-site windytraveler-site windy-hand"
" windy-cloud-sites windy-cloud-domains windy-cloud-vps windytalk windy-pro windy-mind",
).split()
# Gitea run status -> GitHub status state. `skipped` is deliberately absent: a
# job skipped by its own `if:` (e.g. substrate-drift's no-secrets path) has no
# verdict, and painting it green would be a claim nobody tested.
STATE = {
"success": "success",
"failure": "failure",
"cancelled": "error",
"running": "pending",
"waiting": "pending",
"blocked": "pending",
}
MIRROR_TAG = "[GH#"
# Image-build jobs cannot pass here BY DESIGN: job containers get no Docker
# daemon (I-5 — the host socket would hand every workflow root on Veron 1).
# Posting them would put a permanent red X on every commit, and a signal that is
# always red trains everyone to ignore red. Not posted until a rootless builder
# exists; that is a decision, recorded in docs/CUTOVER.md, not a failure.
NO_DAEMON_JOB = re.compile(r"docker", re.IGNORECASE)
# Jobs Grant ruled NON-BLOCKING (GRANT_DECISIONS_2026-09-23): still run on
# Windy Git and visible there, but not posted to GitHub, so they cannot turn a
# commit's combined status red. Format: "repo:workflow/job,workflow/job;repo2:..."
# windy-pro's desktop/installer jobs belong to Grant's desktop side (fixed from
# his Mac mini), not to any lane's merge gate.
NON_BLOCKING: dict[str, set[str]] = {}
for _entry in os.environ.get(
"BRIDGE_NON_BLOCKING", "windy-pro:ci/build-desktop,ci/test-installer,ci/reality-check"
).split(";"):
if ":" in _entry:
_repo, _jobs = _entry.split(":", 1)
NON_BLOCKING[_repo.strip()] = {j.strip() for j in _jobs.split(",") if j.strip()}
# Gitea reads the FIRST of these dirs that has workflow files at a commit (1.24).
WORKFLOW_DIRS = (".gitea/workflows", ".github/workflows")
def workflow_problem(text: str) -> str | None:
"""Why Gitea would drop this workflow file, or None if it looks runnable.
Gitea skips an invalid workflow with one log line and fires no run at all,
so on GitHub the PR just shows nothing, and people wait for CI that is never
coming. These are the shapes we have actually hit, not a full schema.
"""
try:
doc = yaml.safe_load(text)
except yaml.YAMLError as e:
mark = getattr(e, "problem_mark", None)
return f"invalid YAML at line {mark.line + 1}" if mark else "invalid YAML"
if not isinstance(doc, dict):
return "not a YAML mapping"
if "on" not in doc and True not in doc: # YAML 1.1 reads a bare `on` as True
return "no `on:` trigger"
jobs = doc.get("jobs")
if not isinstance(jobs, dict) or not jobs:
return "no `jobs:`"
for name, job in jobs.items():
if not isinstance(job, dict):
return f"job `{name}` is not a mapping"
if "runs-on" not in job and "uses" not in job:
return f"job `{name}` has no `runs-on:`"
return None
def invalid_workflows(repo: str, sha: str) -> dict[str, tuple[str, str]]:
"""{context: (path, problem)} for each workflow file at `sha` that won't run."""
for d in WORKFLOW_DIRS:
st, entries = gitea("GET", f"/repos/{WG_OWNER}/{repo}/contents/{d}?ref={sha}")
if st == 404:
continue
if st != 200:
raise RuntimeError(f"{repo}: Windy Git {d}@{sha[:7]} -> {st}")
files = [e for e in entries or [] if e.get("type") == "file"
and e["name"].endswith((".yml", ".yaml"))]
if not files:
continue
bad = {}
for e in files:
st, f = gitea("GET", f"/repos/{WG_OWNER}/{repo}/contents/{e['path']}?ref={sha}")
if st != 200:
raise RuntimeError(f"{repo}: Windy Git {e['path']}@{sha[:7]} -> {st}")
problem = workflow_problem(base64.b64decode(f["content"]).decode("utf-8", "replace"))
if problem:
stem = re.sub(r"\.ya?ml$", "", e["name"])
bad[f"windy-git/{stem}/workflow"] = (e["path"], problem)
return bad
return {}
def _call(base: str, token_header: str, method: str, path: str, body=None):
req = urllib.request.Request(
base + path,
data=json.dumps(body).encode() if body is not None else None,
method=method,
headers={
"Authorization": token_header,
"Content-Type": "application/json",
"Accept": "application/json",
# urllib's default UA is 403'd as a bot by GitHub's edge and CF.
"User-Agent": "windy-git-pr-bridge/1",
},
)
# Transport errors (TLS handshake timeout, reset) are retried: one GitHub
# blip used to fail the whole sync, flip its heartbeat to ok:false and page
# someone for nothing. HTTP errors are answers, not blips — never retried.
for attempt in range(3):
try:
with urllib.request.urlopen(req, timeout=60) as r:
raw = r.read()
return r.status, (json.loads(raw) if raw else None)
except urllib.error.HTTPError as e:
return e.code, None
except (urllib.error.URLError, TimeoutError, ConnectionError):
if attempt == 2:
raise
time.sleep(2 * (attempt + 1))
raise AssertionError("unreachable")
def gitea(method, path, body=None):
return _call(GITEA + "/api/v1", f"token {GITEA_TOKEN}", method, path, body)
def github(method, path, body=None):
return _call("https://api.github.com", f"Bearer {GITHUB_TOKEN}", method, path, body)
def sync_prs(repo: str) -> list[str]:
"""Mirror open same-repo GitHub PRs into Windy Git. Returns their head shas."""
st, gh_prs = github("GET", f"/repos/{GH_OWNER}/{repo}/pulls?state=open&per_page=100")
if st != 200:
raise RuntimeError(f"{repo}: GitHub PR list -> {st}")
st, wg_prs = gitea("GET", f"/repos/{WG_OWNER}/{repo}/pulls?state=open&limit=50")
if st != 200:
raise RuntimeError(f"{repo}: Windy Git PR list -> {st}")
ours = {p["title"].split("]")[0] + "]": p for p in wg_prs if p["title"].startswith(MIRROR_TAG)}
heads, wanted = [], set()
for pr in gh_prs:
if pr["head"]["repo"] is None or pr["head"]["repo"]["full_name"] != f"{GH_OWNER}/{repo}":
continue # fork PR — never synced, never run here
tag = f"{MIRROR_TAG}{pr['number']}]"
wanted.add(tag)
heads.append(pr["head"]["sha"])
if tag in ours:
continue
st, _ = gitea(
"POST",
f"/repos/{WG_OWNER}/{repo}/pulls",
{
"head": pr["head"]["ref"],
"base": pr["base"]["ref"],
"title": f"{tag} {pr['title']}"[:250],
"body": f"Mirror of {pr['html_url']} so CI runs here. Do not merge in Windy Git — "
"GitHub is the source of truth; merge there.",
},
)
print(f" {repo}: opened mirror PR for GH#{pr['number']} -> {st}")
for tag, p in ours.items():
if tag not in wanted:
gitea("PATCH", f"/repos/{WG_OWNER}/{repo}/pulls/{p['number']}", {"state": "closed"})
print(f" {repo}: closed mirror PR {tag} (closed on GitHub)")
return heads
def post_statuses(repo: str, sha: str) -> None:
# Gitea caps a page at 50 (MAX_RESPONSE_ITEMS) whatever `limit` says, and a
# daily scheduled workflow can push a quiet main's runs off page 1.
runs = []
for page in range(1, 6):
st, body = gitea("GET", f"/repos/{WG_OWNER}/{repo}/actions/tasks?limit=50&page={page}")
if st != 200:
raise RuntimeError(f"{repo}: Windy Git runs -> {st}")
runs += body.get("workflow_runs", [])
if len(body.get("workflow_runs", [])) < 50:
break
latest: dict[str, dict] = {}
for r in runs:
if r["head_sha"] != sha or NO_DAEMON_JOB.search(r["name"]):
continue
if f"{r['workflow_id'].removesuffix('.yml')}/{r['name']}" in NON_BLOCKING.get(repo, ()):
continue
ctx = f"windy-git/{r['workflow_id'].removesuffix('.yml')}/{r['name']}"
if ctx not in latest or r["id"] > latest[ctx]["id"]:
latest[ctx] = r
bad = invalid_workflows(repo, sha)
if not (latest or bad):
return
st, existing = github("GET", f"/repos/{GH_OWNER}/{repo}/commits/{sha}/statuses?per_page=100")
current: dict[str, str] = {}
for s in existing or []: # newest first
current.setdefault(s["context"], s["state"])
for ctx, (path, problem) in sorted(bad.items()):
if current.get(ctx) == "error":
continue
st, _ = github(
"POST",
f"/repos/{GH_OWNER}/{repo}/statuses/{sha}",
{
"state": "error",
"context": ctx,
"description": f"Windy Git ignored this workflow, no CI ran: {problem}"[:140],
"target_url": f"{PUBLIC}/{WG_OWNER}/{repo}/src/commit/{sha}/{path}",
},
)
print(f" {repo}@{sha[:7]} {ctx} = error ({problem}) -> {st}")
for ctx, r in sorted(latest.items()):
state = STATE.get(r["status"])
if state is None or current.get(ctx) == state:
continue
st, _ = github(
"POST",
f"/repos/{GH_OWNER}/{repo}/statuses/{sha}",
{
"state": state,
"context": ctx,
"description": f"Windy Git CI on Veron 1: {r['status']}"[:140],
"target_url": f"{PUBLIC}/{WG_OWNER}/{repo}/actions/runs/{r['run_number']}",
},
)
print(f" {repo}@{sha[:7]} {ctx} = {state} -> {st}")
def main() -> int:
if not (GITEA_TOKEN and GITHUB_TOKEN):
sys.exit("GITEA_ADMIN_TOKEN and GITHUB_TOKEN are required")
failed = 0
for repo in REPOS:
try:
shas = sync_prs(repo)
st, br = github("GET", f"/repos/{GH_OWNER}/{repo}")
if st == 200:
st, b = github("GET", f"/repos/{GH_OWNER}/{repo}/branches/{br['default_branch']}")
if st == 200:
shas.append(b["commit"]["sha"])
for sha in dict.fromkeys(shas):
post_statuses(repo, sha)
except Exception as e: # one repo's failure must not hide the others'
print(f" FAILED {repo}: {e}")
failed = 1
return failed
if __name__ == "__main__":
sys.exit(main())

24
scripts/promote_to_ci.sh Executable file
View File

@@ -0,0 +1,24 @@
#!/usr/bin/env bash
# promote_to_ci.sh <repo> [workflow-to-disable ...] — pull mirror -> writable CI repo.
# Run ON Veron as root. See docs/CUTOVER.md "Onboarding another private repo".
#
# ⚠️ It DELETES the mirror before importing (Gitea's migrate refuses an existing
# name). If the import then fails, the Windy Git copy is gone until you re-run —
# GitHub and the nightly R2 bundles still hold everything, but check first that
# scripts/import_from_github.py will accept the repo. (2026-09-23: windy-pro was
# deleted this way while the importer still refused it by name.)
set -euo pipefail
set -a; . /srv/windygit/src/.env; set +a
export IMPORT_GITEA_URL=http://localhost:3080
A=http://localhost:3080/api/v1; H="Authorization: token $GITEA_ADMIN_TOKEN"; r=$1; shift
info=$(curl -s -H "$H" $A/repos/windyadmin/$r)
m=$(echo "$info" | python3 -c 'import json,sys;print(json.load(sys.stdin).get("mirror"))')
if [ "$m" = True ]; then
curl -sf -o /dev/null -X DELETE -H "$H" $A/repos/windyadmin/$r
(cd /srv/windygit/src && python3 scripts/import_from_github.py "$r" | tail -1)
elif [ "$m" = False ]; then echo "$r already writable"; else echo "$r absent -> importing"; (cd /srv/windygit/src && python3 scripts/import_from_github.py "$r" | tail -1); fi
db=$(curl -s -H "$H" $A/repos/windyadmin/$r | python3 -c 'import json,sys;print(json.load(sys.stdin).get("default_branch","main"))')
for i in $(seq 1 120); do curl -sf -o /dev/null -H "$H" $A/repos/windyadmin/$r/branches/$db && break; sleep 5; done
for w in "$@"; do printf " disable %s: " "$w"; curl -s -o /dev/null -w '%{http_code}\n' -X PUT -H "$H" $A/repos/windyadmin/$r/actions/workflows/$w/disable; done
curl -s -H "$H" $A/repos/windyadmin/$r/actions/workflows | python3 -c 'import json,sys,os;print(" "+os.environ.get("R",""),[(w["path"].split("/")[-1],w["state"]) for w in json.load(sys.stdin).get("workflows",[])])'
echo " default=$db"

102
scripts/sync_from_github.sh Executable file
View File

@@ -0,0 +1,102 @@
#!/usr/bin/env bash
# Phase 1 sync: GitHub is the source of truth, Windy Git follows.
#
# ── Why this direction, and why the other one was wrong ────────────────────
#
# On 2026-08-13 nine repos were migrated writable with push-mirrors pointed AT
# GitHub. That was premature: a dozen agent sessions on the Mac mini are pushing
# to GitHub continuously, so GitHub — not Windy Git — is where the current work
# actually lives. A push-mirror force-updates refs, so on its 8-hour timer it
# would have pushed Windy Git's stale copy over live work, silently, with no
# conflict to notice. The mirrors were removed before the first timer fired.
#
# This script is the correct Phase 1: pull from GitHub, push into Windy Git.
#
# Mac mini agents ──push──▶ GitHub ──this script──▶ Windy Git ──▶ CI on Veron
#
# **It requires nothing from anyone.** No remote changes, no coordination, no
# "everybody stop pushing for a minute." Agents keep working exactly as they are
# and CI starts running on 24 cores.
#
# Windy Git is force-updated on purpose. In Phase 1 it holds nothing anyone
# depends on, so GitHub always wins and there is no merge to reconcile — which
# is the entire point of not flipping direction until a repo is quiet.
#
# Phase 2, per repo, only when that repo is idle: point its agents at Windy Git,
# drop it from REPOS here, and add a push-mirror back to GitHub. One repo at a
# time. Never a big-bang cutover across a dozen live sessions.
set -uo pipefail
: "${GITHUB_TOKEN:?GITHUB_TOKEN required}"
: "${GITEA_ADMIN_TOKEN:?GITEA_ADMIN_TOKEN required}"
GH_OWNER="${GITHUB_OWNER:-sneakyfree}"
WG="${WG_HOST:-app.windygit.com}"
WG_OWNER="${WINDYGIT_OWNER:-windyadmin}"
WORK="${SYNC_WORK:-/srv/windygit/sync}"
FAILED=0
# Repos Windy Git tracks FROM GitHub. Remove a repo from this list at the moment
# it flips to Windy-Git-first, or the sync will fight its authors and win.
REPOS="${SYNC_REPOS:-windy-calendar windy-search windy-registry Windy-Clone WindyCloud windy-cloud-sites windy-mind eternitas windy-agent windy-git windy-chat windy-mail windy-connect windy-drops windy-code-web windy-code windy-traveler windy-translate windytranslate-site windytraveler-site windy-hand windy-cloud-domains windy-cloud-vps windytalk windy-pro}"
# Repos whose TAGS must not reach Windy Git. A tag push fires `on: push: tags`
# workflows; windy-pro's build-electron is a matrix over ubuntu/macos/windows-
# latest, labels no runner here has, so every leg would queue forever (and
# queued jobs are invisible in /actions/tasks). Releases are built elsewhere.
NO_TAGS="${SYNC_NO_TAGS:-windy-pro}"
# `archive/*` branches never reach Windy Git (negative refspec, git >= 2.29).
# They are off-machine safety copies of unpushed work (one-repo doctrine), not
# work in progress: GitHub holds them, and CI time on them is waste.
mkdir -p "$WORK"
log() { printf '[sync %s] %s\n' "$(date -u +%H:%M:%SZ)" "$*"; }
for r in $REPOS; do
bare="$WORK/${r}.git"
if [[ ! -d "$bare" ]]; then
git clone --quiet --bare "https://x-access-token:${GITHUB_TOKEN}@github.com/${GH_OWNER}/${r}.git" "$bare" 2>/dev/null \
|| { log "FAILED initial clone of $r"; FAILED=1; continue; }
fi
# +refs/heads/* — branches only, deliberately.
#
# `--mirror` would also carry refs/pull/* (GitHub's read-only PR refs, which
# Gitea rejects) and every remote-tracking ref, turning a working sync into a
# wall of errors that hides the one that matters.
if ! git --git-dir="$bare" fetch --quiet --prune origin '+refs/heads/*:refs/heads/*' '+refs/tags/*:refs/tags/*' 2>/dev/null; then
log "FAILED fetch $r"; FAILED=1; continue
fi
before="$(git --git-dir="$bare" rev-parse HEAD 2>/dev/null || echo none)"
if git --git-dir="$bare" push --quiet --force \
"https://${WG_OWNER}:${GITEA_ADMIN_TOKEN}@${WG}/${WG_OWNER}/${r}.git" \
'+refs/heads/*:refs/heads/*' '^refs/heads/archive/*' $([[ " $NO_TAGS " == *" $r "* ]] || echo '+refs/tags/*:refs/tags/*') 2>/dev/null; then
log "$r ok (${before:0:7})"
else
log "FAILED push $r -> windy git"; FAILED=1
fi
done
# Jobs that name labels no runner has (ubuntu/macos/windows-latest) would wait
# forever and invisibly; cancel them after 30 min. Never fails the sync.
# Both DB steps go through `docker exec`, which hangs outright while the host
# is in an IO stall (09-23: data2 SMR cliff wedged this sync for 10+ min and
# stopped mirroring + the bridge for every lane). They are optional; mirroring
# and the bridge are not. Bound them so a stuck exec costs one step, not the run.
timeout -k 10 120 bash "$(dirname "$0")/cancel_unrunnable.sh" || log "janitor failed or timed out (non-fatal)"
# Private repos can't run GitHub Actions; mirror their open PRs here so CI
# fires, and post the verdicts back to GitHub as commit statuses.
if ! python3 "$(dirname "$0")/pr_status_bridge.py"; then
log "FAILED pr status bridge"; FAILED=1
fi
# CI telemetry -> admin.windyword.ai (shapes declared with Windy Telemetry 40).
# Sends nothing until WINDYGIT_TELEMETRY_TOKEN is set; never fails the sync.
timeout -k 10 180 python3 "$(dirname "$0")/telemetry_emit.py" || log "telemetry emit failed or timed out (non-fatal)"
[[ "$FAILED" -ne 0 ]] && { log "COMPLETED WITH FAILURES"; exit 1; }
log "all repos in step with GitHub"

385
scripts/telemetry_emit.py Normal file
View File

@@ -0,0 +1,385 @@
#!/usr/bin/env python3
"""Emit Windy Git CI telemetry to admin.windyword.ai (Windy Telemetry 40's ledger).
Runs on Veron after every sync (root; reads the gitea DB via `docker exec`).
Shapes are declared with Telemetry 40 (2026-09-23) — do not add keys or enum
values without re-declaring: a declared family quarantines any row that
doesn't match.
ci.run one row per FINISHED job, exactly once — cursor on
(finish time, job id) in STATE; jobs finish out of id order
service.health one row per invocation: CI plane counts for the interval
Privacy: ids, names of repos/jobs, codes, counts, durations. No commit
messages, no logs, no author names.
--dry-run print the batch instead of posting (and don't advance STATE)
"""
from __future__ import annotations
import json
import os
import re
import subprocess
import sys
import time
import urllib.error
import urllib.request
from datetime import UTC, datetime
INGEST = os.environ.get("TELEMETRY_INGEST_URL", "https://admin.windyword.ai/v1/events")
TOKEN = os.environ.get("WINDYGIT_TELEMETRY_TOKEN", "")
STATE = os.environ.get("TELEMETRY_STATE", "/var/lib/windy-git/telemetry-state.json")
PLATFORM, SERVICE = "windy-git", "ci"
OUTCOME = {1: "success", 2: "failure", 3: "cancelled", 4: "skipped"}
EVENTS = {"push", "pull_request", "pull_request_sync", "schedule", "workflow_dispatch"}
RUNNERS_EXPECTED = 6
def sql(query: str) -> list[dict]:
"""Rows as dicts, via psql's json_agg — no driver needed on the host."""
wrapped = f"select coalesce(json_agg(t), '[]'::json) from ({query}) t;"
out = subprocess.run(
[
"docker",
"exec",
"-i",
"windy-git-db-1",
"sh",
"-c",
'psql -U "$POSTGRES_USER" -d gitea -At -v ON_ERROR_STOP=1',
],
input=wrapped,
capture_output=True,
text=True,
check=True,
).stdout.strip()
return json.loads(out or "[]")
def load_state() -> dict:
try:
with open(STATE) as f:
return json.load(f)
except (OSError, ValueError):
return {}
def iso(epoch: float) -> str:
return datetime.fromtimestamp(epoch, UTC).isoformat().replace("+00:00", "Z")
# ---- push velocity: DETECT + ALERT ONLY (G3.4, 2026-09-23) -----------------
# `git push` goes straight to Gitea and never touches our API, so throttle.py
# cannot see it (NOT_ENFORCED_HERE). Gitea's own `action` table does record
# every push, so we read it here, emit `forge.push_velocity` when an account
# crosses a threshold, and let Telemetry Boss's detector page. Nothing here sits
# in the push path; nothing is ever refused (orchestrator, 09-23).
#
# Thresholds = the STANDARD-band bases from config.py (500 pushes/day; the
# force-push base of 10/day is used for ref deletes, the closest thing we can
# see). EI band multipliers are NOT applied: a platinum agent over 500/day is
# still flagged, for a human to look at, not blocked. Gitea records no
# "forced" flag, so force pushes cannot be told apart from pushes: named, not
# guessed.
PV_RULES = ( # (rule, row key, window_s, threshold)
("pushes_1h", "p1h", 3600, 60),
("pushes_24h", "p24h", 86400, 500),
("ref_deletes_24h", "d24h", 86400, 10),
)
# The GitHub -> Windy Git sync pushes as windyadmin every 5 min, by design.
PV_EXEMPT = {"windyadmin"}
# Gitea op_type: 5 commit push, 9 tag push, 16 tag delete, 17 branch delete.
# One action row per WATCHER is written for each push; user_id = act_user_id
# keeps exactly the actor's own copy.
PV_QUERY = """
select a.act_user_id as uid, u.lower_name as login,
(select el.external_id from external_login_user el
where el.user_id = a.act_user_id order by el.external_id limit 1) as wid,
count(*) filter (where a.op_type in (5, 9) and a.created_unix > {h1}) as p1h,
count(*) filter (where a.op_type in (5, 9)) as p24h,
count(*) filter (where a.op_type in (16, 17)) as d24h,
count(distinct a.repo_id) as repos
from action a join "user" u on u.id = a.act_user_id
where a.created_unix > {h24} and a.user_id = a.act_user_id
and a.op_type in (5, 9, 16, 17)
group by 1, 2"""
def passport_from_login(login: str) -> str | None:
"""agent-et26abcd1234 -> ET26-ABCD-1234 (repos.py _owner_login, reversed)."""
m = re.fullmatch(r"agent-([a-z0-9]{4})([a-z0-9]{4})([a-z0-9]{4})", login)
return "-".join(g.upper() for g in m.groups()) if m else None
def push_velocity_events(rows: list[dict], now: float, alerted: dict) -> tuple[list[dict], dict]:
"""(events, alerted') — one row per account per rule per window while over.
`alerted` maps "<uid>:<rule>" -> epoch of the last row. An account still over
the line is re-reported once per window, not every 5 minutes; one that drops
back under is forgotten, so a later burst reports again.
"""
events, keep = [], {}
for r in rows:
login = str(r["login"])
if login in PV_EXEMPT:
continue
agent = login.startswith("agent-")
for rule, key, window, limit in PV_RULES:
n = int(r[key])
if n <= limit:
continue
k = f"{r['uid']}:{rule}"
last = alerted.get(k)
if last is not None and now - float(last) < window:
keep[k] = last
continue
keep[k] = now
ev = {
"ts": iso(now),
"platform": PLATFORM,
"service": "forge",
"event_type": "forge.push_velocity",
"metadata": {
"rule": rule,
"window_s": window,
"count": n,
"threshold": limit,
"repos": int(r["repos"]),
"gitea_user_id": int(r["uid"]),
},
}
# Actor rule (telemetry UPDATE 2): agent/human rows MUST carry an
# actor_id. Humans sign in to the forge only via Windy SSO, so the
# external login id IS their windy_identity_id. No id we can prove
# -> actor_type system + metadata.caller, never an invented id (I-12).
actor_id = passport_from_login(login) if agent else (r.get("wid") or None)
if actor_id:
ev["actor_type"], ev["actor_id"] = ("agent" if agent else "human"), str(actor_id)
else:
ev["actor_type"] = "system"
ev["metadata"]["caller"] = "unknown"
events.append(ev)
return events, keep
def main() -> int:
dry = "--dry-run" in sys.argv
state = load_state()
now = time.time()
since = float(state.get("last_ts", now - 300))
# Cursor = (finish time, job id), NOT job id alone: jobs finish out of id
# order, so an id high-water mark silently drops every long job that started
# before the mark and finished after it (Telemetry Boss caught this: 43
# finished vs 8 ci.run rows). Finish time = stopped, or updated for jobs
# Gitea/the janitor skipped without a stop time.
if "last_fin" in state:
last_fin, last_id = int(state["last_fin"]), int(state["last_id"])
else: # first run or pre-cursor state: start now, never replay history
last_fin, last_id = int(state.get("last_ts", now)), 0
cutoff = int(now) - 5 # leave the current second alone; late writers land next run
FIN = "coalesce(nullif(j.stopped, 0), j.updated)"
jobs = sql(f"""
select j.id, j.name as job, j.status, j.started, j.stopped, {FIN} as fin,
p.lower_name as repo, p.default_branch, r.workflow_id, r.event,
r.ref, r.index as run, left(r.commit_sha, 7) as sha
from action_run_job j
join action_run r on r.id = j.run_id
join repository p on p.id = r.repo_id
where j.status in (1, 2, 3, 4)
and ({FIN}, j.id) > ({last_fin}, {last_id})
and {FIN} <= {cutoff}
order by {FIN}, j.id
limit 2000""")
try: # posted_to_github: the bridge's own rules, from the same checkout
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import pr_status_bridge as bridge
except Exception: # noqa: BLE001
bridge = None
events = []
for j in jobs:
ref = j["ref"] or ""
if ref.startswith("refs/pull/"):
kind = "pr"
elif ref == f"refs/heads/{j['default_branch']}":
kind = "default"
else:
kind = "other"
# Gitea stores whole seconds; duration_ms is seconds*1000 (so 10000 = 10 s).
dur = (j["stopped"] - j["started"]) * 1000 if j["started"] and j["stopped"] else None
ev = {
"ts": iso(j["stopped"]),
"platform": PLATFORM,
"service": SERVICE,
"event_type": "ci.run",
"actor_type": "system",
"metadata": {
"repo": j["repo"],
"workflow": (j["workflow_id"] or "").removesuffix(".yml").removesuffix(".yaml"),
"job": j["job"],
"outcome": OUTCOME[j["status"]],
"event": j["event"] if j["event"] in EVENTS else "other",
"branch_kind": kind,
"run": j["run"],
"sha": j["sha"],
},
}
if dur is not None and dur >= 0:
ev["duration_ms"] = int(dur)
if bridge is not None:
wf = ev["metadata"]["workflow"]
ev["metadata"]["posted_to_github"] = bool(
j["repo"] in {r.lower() for r in bridge.REPOS}
and kind in ("default", "pr")
and j["status"] != 4
and not bridge.NO_DAEMON_JOB.search(j["job"])
and f"{wf}/{j['job']}" not in bridge.NON_BLOCKING.get(j["repo"], set())
)
events.append(ev)
# --- heartbeat: counts since the previous invocation --------------------
# Interval counts come from EXACTLY the rows emitted above, so
# sum(jobs_finished) over any window == count(ci.run) in it, by construction.
h = sql(f"""
select
(select count(*) from action_run_job where status in (5, 7)) as jobs_waiting,
(select count(*) from action_run_job where status = 6) as jobs_running,
(select count(*) from action_runner where deleted is null and last_online >= {int(now) - 120}) as runners_online,
(select coalesce(extract(epoch from now())::bigint - min(created), 0)
from action_run_job where status in (5, 7)) as oldest_waiting_s""")[0]
h["jobs_finished"] = len(jobs)
h["jobs_failed"] = sum(1 for j in jobs if j["status"] == 2)
h["jobs_cancelled"] = sum(1 for j in jobs if j["status"] == 3)
meta = {k: int(v) for k, v in h.items()}
meta["interval_s"] = int(now - since) # ecosystem-standard key
# UPDATE 7. Quarantines seen on earlier sends (the ledger answers 202 anyway)
# are carried in the state file until a heartbeat reports them. Dropped is 0
# by construction: a failed send keeps the cursor and the spool, so every row
# is re-sent next run (a partial failure can duplicate, never lose).
meta["telemetry_quarantined"] = int(state.get("quarantined_unreported", 0))
meta["telemetry_dropped"] = 0
for k in ("repos_synced", "repos_sync_failed", "statuses_posted", "bridge_errors"):
v = os.environ.get(f"TELEMETRY_{k.upper()}")
if v is not None and v.isdigit(): # absent = couldn't count; never invent 0
meta[k] = int(v)
events.append(
{
"ts": iso(now),
"platform": PLATFORM,
"service": SERVICE,
"event_type": "service.health",
"actor_type": "system",
"metadata": meta,
}
)
# Isolated: a failing push-velocity query must never cost the ci.run rows.
pv_alerted = state.get("pv_alerted", {})
try:
pv_rows = sql(PV_QUERY.format(h1=int(now) - 3600, h24=int(now) - 86400))
pv_events, pv_alerted = push_velocity_events(pv_rows, now, pv_alerted)
except (subprocess.CalledProcessError, ValueError, KeyError) as e:
print(f"[telemetry] push velocity check FAILED (non-fatal): {type(e).__name__}")
pv_events = []
for e in pv_events:
m = e["metadata"]
print(f"[telemetry] WARNING push velocity: gitea user {m['gitea_user_id']} "
f"{m['rule']} = {m['count']} > {m['threshold']}")
events += pv_events
# ci.job_cancelled: spooled by the janitor (cancel_unrunnable.sh), one JSON per job.
spool = os.environ.get("JANITOR_SPOOL", "/var/lib/windy-git/janitor-cancelled.jsonl")
spooled = 0
try:
with open(spool) as f:
for line in f:
try:
m = json.loads(line)
except ValueError:
continue
events.append(
{
"ts": iso(now),
"platform": PLATFORM,
"service": SERVICE,
"event_type": "ci.job_cancelled",
"actor_type": "system",
"metadata": {
k: m[k]
for k in ("repo", "workflow", "job", "reason", "runs_on", "waited_s")
},
}
)
spooled += 1
except OSError:
pass
if dry:
out = os.environ.get("TELEMETRY_DRY_OUT")
if out:
with open(out, "w") as f:
json.dump({"events": events}, f)
else:
print(json.dumps({"events": events}, indent=1)[:4000])
print(f"[telemetry] DRY RUN: {len(events)} events ({len(jobs)} ci.run)")
return 0
if not TOKEN:
print("[telemetry] WINDYGIT_TELEMETRY_TOKEN unset — not sending (not a failure)")
return 0
quarantined = 0
for i in range(0, len(events), 500):
req = urllib.request.Request(
INGEST,
data=json.dumps({"events": events[i : i + 500]}).encode(),
method="POST",
headers={
"Authorization": f"Bearer {TOKEN}",
"Content-Type": "application/json",
"User-Agent": "windy-git-telemetry/1",
},
)
try:
with urllib.request.urlopen(req, timeout=30) as r:
body = r.read()
if r.status >= 300:
raise urllib.error.HTTPError(
INGEST, r.status, body[:300].decode(errors="replace"), None, None
)
try:
resp = json.loads(body or b"{}")
except ValueError:
resp = {}
q = resp.get("quarantined") if isinstance(resp, dict) else None
if isinstance(q, int) and q > 0:
quarantined += q
reasons = "; ".join(map(str, resp.get("rejections") or [])) or "no reason given"
print(f"[telemetry] WARNING {q} row(s) QUARANTINED by the ledger: {reasons}")
except urllib.error.HTTPError as e:
print(f"[telemetry] FAILED ingest HTTP {e.code}: {e.read()[:200]!r}")
return 1 # state NOT advanced: the same rows retry next run
except urllib.error.URLError as e:
print(f"[telemetry] FAILED ingest: {e.reason}")
return 1
if spooled:
open(spool, "w").close() # only after every batch was accepted
os.makedirs(os.path.dirname(STATE), exist_ok=True)
new_fin, new_id = (jobs[-1]["fin"], jobs[-1]["id"]) if jobs else (last_fin, last_id)
with open(STATE + ".tmp", "w") as f:
json.dump(
{"last_fin": new_fin, "last_id": new_id, "last_ts": now,
"quarantined_unreported": quarantined, "pv_alerted": pv_alerted},
f,
)
os.replace(STATE + ".tmp", STATE)
print(f"[telemetry] sent {len(events)} events ({len(jobs)} ci.run)")
return 0
if __name__ == "__main__":
sys.exit(main())