- runner-5/6: 50+ jobs were queued with ~11 private repos onboarded. dind
keeps the 12-core ceiling, so this adds concurrency, not CPU.
- Gitea: password + passkey sign-in forms off (break-glass = CLI), and
ACCOUNT_LINKING auto -> login. auto linked any hub login whose email
matched an existing account, and SITE ADMIN windyadmin carries Grant's
email. Grant is linked by the hub's stable sub, which matches first.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
All four promoted from pull mirrors to writable. windy-code keeps only
canonical-domains-lint active: its other 15 workflows target hosted
macOS/Windows/ubuntu-latest runners and would queue forever here.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- deploy/runner/prune.sh + windygit-ci-prune.timer (6h): age-based prune
of the CI-only dind (containers, finished-job volumes, images/builder
cache >7d) plus a hard 60 GB cap. Only that daemon, over its own TCP
socket; never the host's Docker. It was 38 GB and unbounded — the same
class of growth that filled Kit 0 on 09-01.
- pr_status_bridge: jobs named *docker* are not posted. Job containers
have no daemon by design (I-5), so they are red on every commit; a
permanent red X teaches everyone to ignore red.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
windy-connect promoted from pull mirror to writable (release.yml, which
publishes to PyPI on tag push, disabled — the sync pushes tags).
windy-search was already writable; its scheduled drift-check is disabled
because it now runs as cron on Kit 0. Both added to BRIDGE_REPOS.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Four runners x capacity 1 instead of one x capacity 4. Concurrent jobs in
one act_runner share /root/.cache/act; a refresh racing a copy killed 3 of
windy-chat's ~20 jobs at setup-node (lstat ... no such file). Separate
processes have separate caches. Same parallelism, same capped dind.
- Behavioral tests for pr_status_bridge (latest verdict wins, no reposting,
skipped never painted green, fork PRs never run, pagination, PR lifecycle).
- import_from_github.py reads IMPORT_GITEA_URL, not GITEA_BASE_URL: sourcing
the deploy .env pointed it at http://gitea:3000 and it died on DNS after the
mirror it replaces had already been deleted.
- CUTOVER.md: the private-repo CI path, onboarding steps, and the
/actions/tasks-hides-queued-runs trap.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
GitHub Actions can't run on the private platform repos. Windy Git already
has their code and a working runner, so:
- windy-chat and windy-mail were read-only pull mirrors (which can never
run Actions); they are now writable, deploy.yml/build-image.yml disabled,
and synced from GitHub like the others.
- scripts/pr_status_bridge.py mirrors open same-repo GitHub PRs into Windy
Git (so pull_request workflows fire) and posts each job's result back as
a GitHub commit status (windy-git/<workflow>/<job>) on PR heads and the
default-branch head. Fork PRs are never run. Runs after every sync.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
windygit-tunnel had crash-looped ~91k times: another project's
cornercall-tunnel holds 127.0.0.1:2000, and cloudflared exits when it
cannot bind its metrics port. Ingress only survived because a stray
cloudflared.service ran the same config. That unit is now disabled and
/etc/cloudflared/config.yml uses metrics 127.0.0.1:2001.
Also add windy-git to the GitHub->Windy Git sync list; its self-hosted
copy was stuck 3 commits behind (only check + canary workflows, no
deploys, so syncing is safe).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Second root cause found and fixed: act_runner 0.2.11 predates
`runs.using: node24`, so any repo on actions/checkout@v5 died before its first
step. Bumped to 0.6.1; Windy-Clone went 4/4 red to 4/4 green.
Also upgrades the act-cache note from "watch item" to a confirmed job failure
(lstat on a vanished file mid-tar), with the wipe command and the annotated-tag
dead end that looks like a wrong checkout but isn't.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
0.2.11's bundled act only knows runs.using node12/node16/node20, so any repo
pinning a current action major (actions/checkout@v5, actions/setup-python@v6)
fails before its first step with "The runs.using key in action.yml must be one
of: [...], got node24". Windy-Clone is how this surfaced.
Verified node24 is absent from the 0.2.11 binary and present in 0.6.1, and that
every key in deploy/runner/config.yaml still exists in 0.6.1's schema.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The localhost->postgres fix was correct but was never what failed these jobs;
they died at step 2. setup-uv v4+ resolves "latest" through GITHUB_API_URL,
which act_runner points at our own forge, so it 404s and every later step is
skipped by success(). Pinned an explicit uv version across 11 repos.
Also records the diagnosis traps that cost the most time: Gitea's job status
enum (1=success, 2=failure), act misattributing the error to the previous step,
the jobs-log API needing a repo-matched id, and Secure cookies defeating a
loopback curl login.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
windy-registry's postgres integration went failure -> success, proving the fix.
windy-mind and WindyCloud still fail for a cause I could not determine: the
jobs API returns 'job not found' for the ids the runs report, so logs were not
retrievable that way. Next session should read them from the Gitea web UI.
Records the trap that the three repos did NOT share one pattern — a naive
localhost->postgres swap would have left WindyCloud on port 15432.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
State, the immediate task (three-repo localhost->service-name CI fix), the traps
already paid for, and a copy-paste prompt.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
The probe's own log was never retrievable through the jobs API, but the
question it asked was answered better by a direct comparison of two real
workflows on the same runner and image:
windy-git gate @postgres:5432 -> passes its migration round-trip
eternitas migrations @localhost:5432 -> failed
Also scopes test_g73 to workflows that actually run Python. It failed the probe
for not pinning a version when the probe only shelled out to psql — the test
being wrong rather than the workflow.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
Several migrated workflows hardcode postgres at localhost:5432, which is
correct on GitHub-hosted runners (services are port-mapped to the VM) and
suspect on Gitea Actions (the job runs IN a container, so localhost is the job).
Prove which form works before rewriting anyone's workflow.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
The front end was 100% stock Gitea: green teacup, "Gitea: Git with a cup of
tea", "A painless, self-hosted Git service". G2.3 was specified in the plan with
an acceptance test and never executed, and nothing enforced it.
Now: Windy Git name, wind-mark logo, brand-blue accent, and a landing page that
says what this actually is. Uses Gitea's SUPPORTED surface (custom templates +
public assets) so upstream upgrades keep arriving — no source modified (D-2/I-1).
Two traps this cost, both now documented and tested:
1. GITEA__DEFAULT__APP_NAME does not work. Gitea reads APP_NAME from the TOP
LEVEL of app.ini; the env var created a literal [default] section that Gitea
ignores, so the installer's stock APP_NAME kept winning while the config
looked correct. The env-to-ini pass also APPENDED a second APP_NAME rather
than replacing the first — a new variant of the documented G4A.3 trap.
2. Cloudflare caches /assets/* for 6h and no token in this stack can purge, so
the new logo and CSS were invisible while being correct at origin. Brand
assets now carry a VERSION IN THE FILENAME; bump it on every change.
Committed with an idempotent apply.sh, because applying it straight to Veron's
disk first was itself the config-drift trap this project documents: a rebuild
would have silently reverted to stock Gitea.
85 tests green.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
I reintroduced the exact defect I had just criticised. ACTION_BASE listed
"push" and "push.force", but git push goes straight to Gitea over HTTPS and
never touches this API — so nothing records a push, a count would be zero
forever, and enforce() would look up a limit, count nothing, and allow
everything. A silent no-op wearing the costume of a control, made worse by a
config name that implies the protection exists.
Split into ACTION_BASE (actually enforced: repo.create, grant.create) and
NOT_ENFORCED_HERE (push, push.force) with the reason and the remedy written
down: enforcing push velocity needs a Gitea-side pre-receive or push webhook
reporting into agent_actions.
enforce("push") now raises rather than silently allowing, and a test asserts the
two sets stay disjoint.
Found by auditing whether the auth fix could be walked around — every
/api/v1/repos/* route does require a caller, and the only unauthenticated
endpoints are /health, /version and the HMAC-verified webhook.
83 tests green.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
Two deploy traps paid for on 2026-08-14: 'git pull -q' hid a divergent-branch
error so a deploy ran against stale code while reporting success, and the
divergence came from amending a commit a deploy checkout already tracked.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
/version went null after a deploy — the exact "service cannot name its own
commit" defect this project was built to prevent, caught by its own honesty
check.
Cause: the sed replaced "" with "" (a no-op when COMMIT_SHA is empty) and the
grep then matched that same empty string, so the guard verified nothing. A build
with no COMMIT_SHA passed and shipped a container reporting commit_sha: null.
Now the build fails loudly instead.
Second cause of the stale deploy, and it was mine: an earlier `git commit
--amend` + force-push rewrote history the Veron deploy checkout was already
sitting on, leaving it divergent so `git pull -q` failed SILENTLY (-q hid
"Need to specify how to reconcile divergent branches"). Two lessons: do not
force-push a branch a deploy checkout tracks, and do not pull with -q in a
deploy script.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
looks_like_ept used jwt.get_unverified_header, which validates the WHOLE token
and therefore rejects anything with a malformed signature segment. Routing
consequently depended on signature well-formedness: an EPT-shaped token with a
bad signature fell through to the HUMAN path, where it was refused for the wrong
reason and — with require_verified_jwt off (dev) — could have been read as a
human identity via its `sub` claim.
Now the header segment is decoded directly, so routing depends only on what the
token CLAIMS to be; whether it is authentic remains verify_ept's job.
Also routes alg:none to the EPT verifier regardless of typ, since a `none`
token is never valid for any caller. Both forged shapes now return 401
ept_invalid — the honest code — instead of 503 "feature not ready".
Found by noticing a forged EPT returned 503 where the verifier should have
answered 401, rather than accepting "it was refused, close enough".
80 tests green.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
Monkeypatched httpx so resolve_passport sees a real revoked trust body
(status=revoked, band=unproven, allowed=[]) and must raise
PassportNotInGoodStanding — proving the WIRING, not just the decision. This is
the path that stops a validly-signed EPT that outlived its passport's
revocation (~365-day tokens).
Live-confirmed alongside: Eternitas refuses to mint EPTs for revoked bots
("credentials are not issued for non-active bots"), so the only exposure was a
pre-existing token — exactly what this now catches. The active agent's EPT still
returns 200. A fully-live revoked test would require revoking a real fleet
passport (destructive), so the wiring is proven deterministically instead.
79 tests green.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
A revoked passport returns HTTP 200, status=revoked, band=unproven,
allowed_actions=[] (verified live 2026-08-13). resolve_passport keyed refusal
only on HTTP 4xx and band=="untrusted", so it returned band 'unproven' and the
agent was seated. Revocation was NOT enforced on the live auth path at all — and
now that agent auth actually works, a revoked agent could authenticate and act.
Extracts decide_trust(body) -> (band, actions) | raise. Only status=="active"
is allowed; revoked/suspended/frozen/unknown all refuse, fail-closed on the
field that carries the most consequential fact about an identity. The agent call
site turns that into a clean 403 passport_revoked.
This is the REAL revocation gate — the token cannot be un-issued, but its
standing is re-checked on every request, so revocation takes effect on the next
call with no webhook required. The Eternitas webhook remains useful for
invalidating locally-issued credentials/grants (G6.3, not built yet), but it was
never the primary gate and its being unwired is no longer a live exposure.
Behavioral tests: revoked body refused, active accepted, unknown/missing status
fails closed.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
The EPT-shaped forgery is the one that matters after G3.2 — it is what
signature verification actually guards. The JWT-shaped one still exercises the
human gate. Both must_refuse; a 2xx on either pages.
Verified live: forged EPT -> 401 ept_invalid, forged JWT -> 503,
genuine EPT -> 200.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
REOPENS the agent path — but only because possession is now actually proven.
EPT verification (api/app/ept.py): ES256 against Eternitas's published key set
at /.well-known/eternitas-keys. algorithms=["ES256"] makes alg:none and
algorithm confusion unrepresentable rather than merely unlikely; issuer and exp
are enforced by the library; an unknown kid is refused.
Order is deliberate: signature FIRST, trust lookup second. These EPTs live ~365
days and carry rev/tru baked in at issuance, so a year-old "rev: false" proves
nothing — revocation and band still come from a live lookup on every request.
Found while building it: real EPTs put the passport in the "sub" claim. The old
code read "passport"/"sub_passport", which no genuine EPT carries — so real
agents were never recognised and ONLY forged tokens ever authenticated. The
bypass was not just a hole, it was the only thing that worked.
Throttle (api/app/throttle.py): BAND_MULTIPLIER and rate_*_per_day were defined
and read by nothing. Now enforced on repo.create and grant.create, counted
against agent_actions (one source of truth, not a private counter that drifts
from the audit log). Fails CLOSED — a limiter that fails open protects you until
the moment something is wrong. Untrusted band is 403 read-only, not 429, because
"slow down" would be a lie.
Tests: 14 behavioral, signing real ES256 tokens with a locally-generated key so
they exercise the crypto path with no network dependency — genuine tokens
accepted, and alg:none / foreign key / tampered payload / expired / wrong issuer
/ unknown kid / missing claims all refused. 74 green.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
Adds must_refuse checks: a 2xx is the alarm, a 401/403/503 is health. The
forged alg:none agent token is probed every 10 min; if it ever returns 2xx the
canary goes red and pages. Proven both ways — 503 reads ok, a 200 endpoint
alarms 'ACCEPTED — this MUST be refused'. The security property is now enforced
by a running check, not assumed.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
Verified live 2026-08-13: a forged 'alg:none' token naming a passport lifted
from the logs returned HTTP 200 as that agent. The agent path read the passport
without verifying the EPT signature, asked Eternitas 'is this passport
reputable?', and seated the caller on a yes. That answers reputation, not
possession — anyone who knows a passport number could impersonate that agent on
the public API.
The human path already failed closed for exactly this reason
(require_verified_jwt). The gate was on the wrong path: it sat AFTER the agent
branch returned. The agent path now fails closed too, BEFORE the trust lookup,
so a forged token never even reaches Eternitas. Reopens automatically when the
ES256/JWKS verifier (G3.2/G9.1) exists and this gate consults it.
Adds BEHAVIORAL tests (not string-grep): a forged alg:none token exercised
through the real get_caller must raise, not authenticate. This is the test that
would have caught the bypass; the suite had 86 source-string assertions and
zero that ran the auth decision.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
131 read-only mirrors (cannot run Actions, zero deploy risk) + 12 writable with
CI and deploys disabled. Total size matches the measured GitHub archive exactly,
which is the confirmation the copy is complete.
Splits the two concerns cleanly: having a copy is safe and should cover
everything now; running code needs judgement and happens per repo.
windy-pro IS included as a mirror — the G11.5 caution is about making it
writable while six checkouts disagree on HEAD, not about holding a read-only
copy. The DR copy is now complete.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The function landed but the argparse anchor did not match the real formatting,
so the command existed and could not be invoked. Caught by running it rather
than assuming the patch applied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Grant's plan: clone the whole account, let it circulate, reverse direction later
when things are clean. Right plan, with one change that matters.
Sampling 40 repos found 18 carrying deploy/release/publish workflows that
trigger on push: — roughly 63 across the account. Importing those writable with
Actions enabled would arm sixty-odd production deploy triggers on Veron 1, each
needing disarming by hand.
So the bulk goes in as READ-ONLY pull mirrors. A mirror cannot run Actions at
all, so this carries zero deploy risk, and Gitea syncs them itself with no
script and no timer. What you get is a complete, current second copy of the
whole account — the disaster-recovery half — with none of the execution risk.
Converting one to writable + CI stays a deliberate per-repo act: re-import,
review its workflows, disable the deploying ones. That judgement belongs at the
moment you want CI on that repo, not in bulk sixty times by accident.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Six workflows deploy to production on push:. Windy Git now has a working
runner, so the next synced commit to main would have attempted a production
deploy FROM VERON 1. Their secrets are unset here so they would have failed —
but loudly, on every push, with any pre-SSH step still running.
All six now disabled_manually. Tests, lints and migration checks stay active:
they need no secrets, which is exactly why Phase 1 delivers CI value with
nothing to configure.
Same class of mistake as the push-mirror direction, caught before firing this
time rather than after.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phase 1 requires nothing from anyone: agents keep pushing to GitHub, a timer
syncs GitHub -> Windy Git every 15 minutes, CI runs on Veron against current
code. Phase 2 flips one repo at a time, only when that repo is idle.
Records the direction mistake honestly: the source of truth is wherever people
are actually typing, not wherever the plan says it should be.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
I migrated nine repos writable with push-mirrors pointed AT GitHub. That was
wrong for the actual situation: a dozen agent sessions on the Mac mini are
pushing to GitHub continuously, so GitHub is where the live work is.
A push-mirror force-updates refs. On its 8-hour timer it would have pushed
Windy Git's stale copy over live work — silently, no conflict, nothing to
notice. Removed all nine before the first timer fired; verified no GitHub repo
had been touched (latest push predated the mirrors).
Replaced with the correct Phase 1 direction:
agents --push--> GitHub --sync--> Windy Git --> CI on Veron
It requires NOTHING from anyone. No remote changes, no coordination, no
'everybody stop pushing'. Agents keep working exactly as they are and CI starts
running on 24 cores.
Windy Git is force-updated on purpose: in Phase 1 it holds nothing anyone
depends on, so GitHub always wins and there is no merge to reconcile.
Phase 2 is per-repo and only when that repo is idle. Never a big-bang cutover
across a dozen live sessions.
Fetches +refs/heads/* and tags explicitly rather than --mirror, which would drag
GitHub's refs/pull/* that Gitea rejects and bury the real errors.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
9 repos migrated writable with push-mirrors to GitHub (sync_on_commit).
Full loop proven end to end: pushed a commit to Windy Git, its existing
workflow ran on Veron 1, and GitHub received the commit within 20s.
Documents the one rule the cutover creates: do NOT push directly to GitHub for
a migrated repo. The mirror makes GitHub match Windy Git, so a direct commit is
overwritten on the next sync, silently, with no conflict. One writer is the
point — two writers with no reconciliation is how you lose work you thought was
saved.
windy-pro deliberately excluded until its six-checkout / forked-counter
ambiguity is resolved by reading (G11.5).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Timer installed with Persistent=true — Grant's workstation is not always on at
04:17, and without that a missed window is silently skipped and the backup
simply never runs.
RESTORE DRILL PASSED, which is the part that matters: pulled a bundle back out
of R2, cloned from it, and the restored HEAD (ffce529) matches live origin/main
exactly — 40 commits, all branches, 50 files.
The August audits found no rehearsed restore anywhere in the ecosystem, for
anything. This is one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Today GitHub is authoritative, so losing Veron 1 costs nothing. The moment
people push HERE first that inverts: Veron 1 holds the only current copy of the
company's source between mirror syncs, and it is Grant's workstation — no SLA,
no snapshots, residential line, and he reboots it.
git bundle over tar, deliberately: a bundle is one file that git clone reads
directly, so a restore needs no knowledge of Gitea's on-disk layout, and
bundling asks git for a consistent view instead of racing a live push.
- --all, so every branch and tag is captured. A single-branch bundle loses
the rest silently and you find out during the restore.
- git bundle verify before upload. An unverified bundle is a belief.
- empty repos are reported as skipped, not counted as failures
- non-zero exit on ANY failure so the unit goes red. A backup script that
swallows errors manufactures confidence.
Whole archive measured 1.58 GB across 141 repos — about two cents a month.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
windy-calendar imported as a pull mirror sat at 0 workflow runs. Gitea does not
fire Actions on mirror sync and a mirror is not a push target, so a mirrored
repo gives you the code and none of the point.
Re-imported it writable, pushed a commit, and its EXISTING .github/workflows/
ci.yml ran on Veron 1 and reported success — with zero workflow edits. That
workflow's own header says it was written for 'OUR self-hosted runners (Kit 0 +
Veron One)' because the account is billing-locked. The fleet already wanted
this; the runners just died when the repos went private.
Default is now writable + push-mirror to GitHub (I-4 steady state). --mirror
stays available for repos to copy but not move.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Surveyed ten repos: 36 of 36 ACTIVE workflows already say
runs-on: [self-hosted, linux, x64]. They were written for the self-hosted
runners that died when the repos went private — so advertising those three
labels makes every one of them runnable AS-IS. No workflow edits, no rewrites.
Correcting an earlier note in this plan: 'ban ubuntu-latest' was wrong as
stated. All 11 occurrences are tagged '# runner-lint-allow — CD/hosted-only;
disabled, manual until CD mission'. They are deliberately hosted-only and
deliberately off. They are correct as written and should not be 'fixed'.
I generalised that doctrine from one repo without surveying. The survey
disagreed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The SOTU scopes the child-process DB bridge as 468 call sites and multi-week.
Measurement says otherwise.
bare node -e '0' on Kit 0 (4 vCPU, load 20): 1.70-1.99s
bare node -e '0' on Veron 1 (24 cores, idle): 0.01s
170x, and node startup is essentially the whole per-query cost — adding pg
connect and a real query to a bare node boot adds only ~0.2-1.2s on top of
~1.8s of interpreter start.
Login makes 9 such forks. 9 x 1.8s = 16s. Observed: 17-25s.
Two independent multipliers compound: the adapter forks per query (9x) and the
box is saturated so each fork costs 1.8s instead of 0.01s (170x).
The consequence: you do not need to touch 468 call sites. querySyncViaChild is
ONE function and its interface does not change — replace the per-query
execFileSync with a persistent worker holding a pg.Pool. All 468 call sites get
faster without being edited, including the mail-lookup that caused today's
outage.
Also measured: 54 containers, 301% CPU of 400%, load 20. dev/demo is 55% of
that — real relief, honestly not a fix. 42 non-dev containers on 4 cores is the
actual condition.
Not implemented here on purpose: sync-over-async with Atomics in the most
critical file in the ecosystem, at the end of a long session, on a box that
already had an outage today, deserves a fresh session and a load test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An unwritable state path raised and took the whole canary down. That is the
worst possible trade for a monitoring tool: it reports nothing at all, and
reports it silently. State is an optimisation for transition detection; the
probing is the point.
Found by fat-fingering an env var, which is exactly how it would happen in
production.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Read the sender rather than guessing: Eternitas generates the webhook secret at
registration time and pings the URL to prove reachability BEFORE returning that
secret. The ping IS signed — with a secret the receiver cannot possibly hold
yet. Unverifiable by construction, not by oversight.
Accepting it is safe because the event is definitionally a no-op: nothing read,
nothing written, acted:false. Every event that changes anything still requires a
valid HMAC. The alternative, skip_validation:true, would permanently disable
reachability checking for this platform to solve a one-time ordering problem.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Eternitas verifies a webhook URL answers BEFORE issuing the secret that signs
deliveries, so the very first request can never carry a signature — refusing it
makes registration impossible. Real chicken-and-egg, not a reason to disable
validation.
A probe is a request claiming no event and carrying no signature. Answering it
200 is honest: the endpoint exists and is ready. It changes nothing (acted:
false), and anything claiming to BE an event still goes through full HMAC
verification. Registering with skip_validation:true would have permanently
disabled a safety check to solve a one-time ordering problem.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
When a passport is revoked, every credential it holds here dies in one
transaction: tokens revoked, grants revoked. A revocation that takes effect
'eventually' is not a revocation.
Avoids two traps that each cost a sibling service a subscription that looked
wired and never once delivered:
1. strip the 'sha256=' prefix before comparing — comparing the decorated
header against a bare digest returns 401 forever
2. HMAC the RAW REQUEST BYTES, never a re-serialised body — JSON.stringify of
a parsed body reorders keys and changes whitespace, so the digest never
matches what the sender signed
Both fail silently from the sender's side: Eternitas records a delivery, the
receiver records a rejection, nobody notices for weeks.
Unset secret REFUSES rather than accepts — accepting unverified instructions
about identity is worse than missing them. And it never acknowledges a
revocation it could not apply; a 200 there is a security hole reporting success.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Caught by TESTING the alert path instead of assuming it. Without an explicit
User-Agent, urllib sends 'Python-urllib/3.x' and Resend rejects it 403, while
the identical request via curl succeeds.
The failure mode this avoids is the worst one a canary has: it would have
detected every outage correctly and told nobody. Same bot-filtering trap as the
Gitea migrate call earlier today — worth recognising on sight.
Also prints the HTTP body on failure. '403 Forbidden' alone sends you hunting
for a bad key; the body names the real cause.
Verified: alert sent (200).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Today's outage is the whole design brief: /health returned 200 for the entire
hour that login was dead. A canary watching /health would have stayed green
while nobody in the ecosystem could sign in. So the login probe is here and it
is the one that matters.
Three rules it obeys:
- never green for something it did not prove (I-8)
- alert on TRANSITIONS, not every run — a canary people filter is a dead
canary, which is how the last one sat 37 days dead unnoticed
- run where the watched thing cannot take it down: Veron 1, never Kit 0
Two independent signals, so losing one still leaves the other: an email via
Resend on state change, and a non-zero exit that turns the CI run red in the
forge itself.
Alerts say what broke in human terms — 'a human can actually sign in' — rather
than only naming an endpoint.
Verified against production: 7/7 green including login at 17.1s; a forced 404
reports DOWN; a 1s threshold reports SLOW at 23.3s.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The retry fix already existed: a parallel session landed c79f196 (#172) with a
test. Kit 0 was ONE COMMIT BEHIND and did not have it — which is the entire
reason the loop ran.
Deployed by ff-merging /root/windy-chat ac61db6 -> c79f196 and rebuilding only
agent-roster (--no-deps). Never reset --hard; that checkout has a documented
history of local edits a hard reset would eat.
account-server CPU 168-210% -> 0.00%
roster failures 74/min -> 0/min
agent chat stopped -> up, healthy
login timeout -> HTTP 200 ~18s
The lesson: a merged fix that has not reached production is not a fix, it is a
belief. That is exactly what both August audits named — nothing checks whether a
decision reached production — arriving as an outage instead of a report finding.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Login was timing out for ~50 minutes on the identity service every Windy
product authenticates against.
Two faults multiplying: windy-agent-roster retried a failing mail-lookup at 74
failures/minute with no backoff, and every call cost 0.6-2.1s because the
postgres adapter forks a node process per query and blocks the event loop.
Together they formed a feedback loop — the container's listening socket showed
Recv-Q 510, connections the kernel accepted that node was too blocked to pick
up. Login sat in that queue.
Resolved by stopping windy-agent-roster: login went from 45s+ timeout to HTTP
200 in ~18s. Restored, not healthy — 18s is the fork-per-query adapter on a
54-container 4-vCPU box, and it is what remains after the loop was removed.
Nothing caught this. The container was (unhealthy) with a failing healthcheck
streak of 74 and no alert exists.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>