Private-key matches are only the BEGIN line (same hash everywhere), so they
are allowed by path+kind; everything else by hash. Real revoked tokens
(1354fc9b, d49dc2ba) are pinned by a test to never be allowed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
New guard windy-git/secret-guard (warn-only) over EVERY text file: Telegram,
GitHub, AWS, Slack, Anthropic, OpenAI, Stripe live, Google API keys and
private-key blocks (scripts/secret_shapes.py, shared with the weekly public
scan). A finding carries "<kind> #<sha256[:8]>", never the value (house rule
10). Known fakes allowed BY HASH (ci/secret-guard-allow.yml). GUARDS_STATUS
gets a secrets column. Leak hunt 09-24: @Windy_0_bot token in a public fixture.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
windy-search #96 "Boot smoke (no Docker)" passed but was hidden by the
image-build name filter. Negative lookbehind for no/no-/without.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
status_for(lane, whole_tree, grant=...): only lane-owned findings fail in
MODE=block; Grant-owned (ci/grant-owned.yml) post WARN. The bridge splits via
guards_report.split_grant; if the split cannot run it WARNs (never blocks).
Orchestrator 09-23: block compute-guard for lane-owned paths only.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
deploy/release workflows run on the target host (real daemon) and are
disabled here (repo_unit DisabledWorkflows); one bounded query per process,
flag everything if it fails.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- bridge: BRIDGE_NO_DAEMON names image-build jobs whose name lacks docker
(default eternitas:ci/build); never posted, like the docker-named ones.
- ci-hygiene: flag docker build/buildx/run/compose, docker-compose and
docker/build-push-action in workflow steps ("needs docker") with the fix:
job services: + a no-Docker smoke test; the image builds at deploy.
- test_guards_report: owner column (14ed23a broke it).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Non-secret inputs git-ignored in windy-pro (models, linux-x64 portable
bundle, enter-monitor build) that build-desktop needs. Mounted :ro into
dind; valid_volumes allows only /ci-inputs/windy-pro; refresh-ci-inputs.sh
copies them from the frozen release clone (read-only on the source).
Invariant I-5 narrowed, not dropped: exactly that one path, read-only in
dind, no other service mounts it, still no docker socket (proven to fail
on :rw). Orchestrator-approved (option a). Applied in an idle window.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The mirror kept its creation-time base forever. eternitas #167 was stacked
on fix/one-hallway, retargeted to main after #166 merged; ci.yml
(pull_request: branches [main]) then silently never ran for it, while
unfiltered workflows did. An edited event triggers nothing, so close the
stale mirror and open a fresh one on the new base (runs CI at once).
An unknown base is left alone, never guessed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Allow entries may carry matches: (regexes); then only matching lines are
allowed, so an allowed file can't smuggle in a new call. windy-pro #609
MindKeychain.jsx: openrouter.ai/auth? and /api/v1/auth/keys (BYOK key
acquisition via OAuth PKCE, no inference; successor of the MindPanel
allow, ADR-064). An inference call in the same file still flags (tested).
Orchestrator-approved.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scans every bridged default branch with both guards; lane-owned vs
Grant-owned (ci/grant-owned.yml: windy-pro desktop paths + its desktop CI
jobs, attributed per job) so Grant's code never holds up a block.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The bridge reads PR/default heads from GitHub after the sync's fetch; a
push in between isn't in the clone until the next cycle. Both guards logged
a CalledProcessError for it (windy-pro main 40 s after the fetch). Now
None = nothing posted this cycle; the next one scans it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
House rule 6 (09-23). The bridge now also posts windy-git/ci-hygiene on
every PR head (added lines) and default branch (whole files), scanning CI
workflows and Dockerfiles for: floating pip / uv pip installs (not -r,
not --no-deps, not exact pins), uv sync without --locked/--frozen,
npm install instead of npm ci (unless every package is exact-pinned),
yarn/pnpm without a frozen lockfile, :latest images and COPY lock* globs
(Windy Mail #147), and CI services publishing a HOST port (every job
shares one dind: Windy Mind runs 147/176 died on 5432). Warn-only;
CI_HYGIENE_MODE=block later. Allow-list ci/ci-hygiene-allow.yml (empty).
compute_guard's walker is now parameterised (line_fn / path_ok /
prefilter) so both guards share one scanner, cache and allow loader; the
bridge posts both through one _post_guard. Today: 95 issues in 21 repos;
windy-git, calendar, traveler, traveler-site clean.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Grant's rule (09-23): every model call goes through Windy Mind. The bridge
now posts windy-git/compute-guard on every PR head (lines the PR ADDS vs its
merge-base) and default-branch head (whole tree): provider hosts, provider
SDK imports/deps and raw provider key names. Warn-only: success + "⚠ WARN"
and a link to the first hit; COMPUTE_GUARD_MODE=block turns it red later.
Exceptions live in ci/compute-guard-allow.yml, each with a reason (Mind
itself, user-BYOK windy-agent / windy-code extension / windy-pro desktop +
MindPanel, windy-connect config writers). Tests, docs, comments, lockfiles,
vendored code and CI config are never scanned. Reads the sync's bare clones
(no docker exec); cached per (repo, sha, rules). Non-fatal; never a fake OK.
First cases = COMPUTE_BYPASS_AUDIT.md. Today on default branches: 38
findings in 3 repos (windy-chat audit #2, windy-pro account-server #3/#4,
windytalk reference/), 0 elsewhere.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Gitea 1.24 lists only picked-up jobs, so a queued PR showed NOTHING on
GitHub and lanes asked whether their push was lost (Windy Mind #131,
Windy Cloud today). The bridge now reads waiting jobs from the gitea DB
and posts pending where nothing newer was picked up; a queued re-run
supersedes the stale failure it replaces.
Only status 5 jobs whose runs-on labels a live runner has: blocked jobs
often end skipped and label-unrunnable jobs are cancelled unpicked, and
neither ever reaches /actions/tasks, so their pending would never resolve.
Lookup is bounded (30 s) and non-fatal: the IO-stall lesson.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Telemetry UPDATE 2 actor rule: agent/human rows without actor_id are
quarantined. Forge humans sign in only via Windy SSO, so Gitea's
external_login_user.external_id is their windy_identity_id.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
git push never touches our API, so throttle.py can't see it. Gitea's
action table records every push; the 5-min sync-side emitter now reads
it and emits forge.push_velocity when an account crosses 60 pushes/1h,
500 pushes/24h (standard-band base) or 10 ref deletes/24h. One row per
account per rule per window while over; windyadmin (the sync) exempt.
Nothing sits in the push path and nothing is refused. HOLD until
Telemetry Boss declares the shape.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
identity.login created a live hub session every 10 min and never ended it.
It now logs out with the token it got: retried on 5xx / no response
(8 x 15 s), 401/404/410 = already over, any other 4xx fails fast, and a
cleanup it can't finish is reported as identity.logout DOWN "CLEANUP
FAILED" (alerts + red run). The hub's /auth/logout revokes every refresh
token of the account (verified live), so the next run's logout heals a
leftover; no ledger needed. Proven end to end: login 200, logout 200,
10/10 checks.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The ledger answers 202 even when it quarantines rows. Both emitters now log
a warning with the reasons and report service.health.telemetry_quarantined
and telemetry_dropped (API: buffer overflow; sync: 0 by construction, since
a failed send keeps cursor + spool). HOLD until Telemetry Boss declares both
keys on windy-git's two service.health shapes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
b7a7e94 made the bridge read workflow files, which the strict fake Gitea
refused (7 red). The fake now serves contents (404 when absent), and new
tests cover: error posted with no runs, valid files add nothing, no repost,
.gitea/workflows wins over .github/workflows, and each workflow_problem
shape. pyyaml declared in dev extras (the bridge imports it).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A single GitHub TLS handshake timeout failed the whole sync, flipped its
windy-job heartbeat to ok:false and would page for nothing. Up to 3
attempts with backoff for URLError/timeout/reset; HTTP errors return
immediately as before. Test covers both.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replaces the keyed marker from 1c3b5b0 with the ecosystem convention:
any X-Windy-Synthetic value marks the request synthetic; the flag lives in
a per-request contextvar, labels this request's rows, and is FORWARDED on
downstream calls (Eternitas trust lookup, Gitea API). The canary sends
"1". Rows are still recorded; the label separates, never suppresses.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The canary deliberately sends forged tokens every 10 min; those refusal
rows read as attacks. It now sends X-Windy-Synthetic carrying a shared
secret (Gitea repo secret CANARY_SYNTHETIC_KEY = WINDYGIT_SYNTHETIC_KEY in
Veron .env); the API marks the row synthetic only on a constant-time
match, so an attacker cannot label their own refusals synthetic to hide.
synthetic is declared on forge.auth.failed (Telemetry Boss, UPDATE 3).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Membrane first: I-2 and MEMBRANE.v1 now list the windy-admin ledger
(POST /v1/events). api/app/telemetry.py: service.boot once per start
(commit_sha omitted when unknown, I-12), an hourly in-process
service.health with the shared keys (requests, errors_5xx/4xx,
refusals_4xx, p95_ms only when there was traffic), and one
forge.auth.failed row per refused request: declared 13-code enum,
http_status, caller class, route TEMPLATE (never the concrete path),
actor_type system with no actor_id (all-lanes rule). No token = nothing
sent or buffered; flush failures keep rows (bounded) and never raise.
Token from root-only /etc/windygit/telemetry.env (optional env_file).
8 behavioural tests.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
build-desktop, test-installer and reality-check still run on Windy Git
and stay visible there, but the bridge no longer posts them to GitHub, so
they cannot turn windy-pro's combined status red. Windy Git side only;
the desktop code is Grant's to fix. BRIDGE_NON_BLOCKING, per repo.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
kit-army-config (the lockbox) and every *-soul / anima repo carry
credentials; the nightly R2 bundles are unencrypted, so the R2 key was a
key to every secret. Excluded by name (BACKUP_EXCLUDE); they are backed up
encrypted by the Windy Drops lane (restic) and stay mirrored on Veron.
Behavioural test runs the script's own exclusion function.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Seven HTTP-level tests through the real route: sha256= prefix and bare
digests accepted, digest of re-serialised JSON refused, forged/wrong-key/
missing signatures refused, unset secret -> 503, a revocation with a bad
signature never reaches the handler, the reachability ping never acts.
Mutation-checked: dropping the prefix strip fails the behavioral test
while the old string-grep invariant still passes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- hub_jwt: aud list is ["windy_git"] (contract v1 array). Dropped
"windy-git": that is Gitea's OIDC client_id, so a forge id_token would
have passed the aud check. `type: human` is now REQUIRED (id_tokens have
none), which makes accepting the discovery-URL issuer safe.
- runner job ceiling 30m -> 90m: eternitas's serial pytest is ~50 min and
would have been killed mid-suite.
- eternitas (private) added to the GitHub status bridge.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The human path refused every token in production (503
human_signin_not_ready) because no verifier existed. api/app/hub_jwt.py
verifies hub access tokens against account.windyword.ai's JWKS:
- RS256 only (closes alg:none and HS256-with-public-key confusion)
- iss must be "windy-identity" — what hub ACCESS tokens carry (observed
live); id_tokens (discovery-URL issuer) are not accepted as bearers
- aud optional today, must name Windy Git when present; hub_require_aud
flips it mandatory once the hub emits it. PyJWT's own aud check is off
on purpose: it rejects ANY aud-bearing token when no audience is given.
- type must be human; identity = windy_identity_id, never sub (per-row id)
- production verifies even if require_verified_jwt is off
11 behavioral tests sign real RS256 tokens with a local key.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- deploy/runner/prune.sh + windygit-ci-prune.timer (6h): age-based prune
of the CI-only dind (containers, finished-job volumes, images/builder
cache >7d) plus a hard 60 GB cap. Only that daemon, over its own TCP
socket; never the host's Docker. It was 38 GB and unbounded — the same
class of growth that filled Kit 0 on 09-01.
- pr_status_bridge: jobs named *docker* are not posted. Job containers
have no daemon by design (I-5), so they are red on every commit; a
permanent red X teaches everyone to ignore red.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
windy-connect promoted from pull mirror to writable (release.yml, which
publishes to PyPI on tag push, disabled — the sync pushes tags).
windy-search was already writable; its scheduled drift-check is disabled
because it now runs as cron on Kit 0. Both added to BRIDGE_REPOS.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Four runners x capacity 1 instead of one x capacity 4. Concurrent jobs in
one act_runner share /root/.cache/act; a refresh racing a copy killed 3 of
windy-chat's ~20 jobs at setup-node (lstat ... no such file). Separate
processes have separate caches. Same parallelism, same capped dind.
- Behavioral tests for pr_status_bridge (latest verdict wins, no reposting,
skipped never painted green, fork PRs never run, pagination, PR lifecycle).
- import_from_github.py reads IMPORT_GITEA_URL, not GITEA_BASE_URL: sourcing
the deploy .env pointed it at http://gitea:3000 and it died on DNS after the
mirror it replaces had already been deleted.
- CUTOVER.md: the private-repo CI path, onboarding steps, and the
/actions/tasks-hides-queued-runs trap.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
windygit-tunnel had crash-looped ~91k times: another project's
cornercall-tunnel holds 127.0.0.1:2000, and cloudflared exits when it
cannot bind its metrics port. Ingress only survived because a stray
cloudflared.service ran the same config. That unit is now disabled and
/etc/cloudflared/config.yml uses metrics 127.0.0.1:2001.
Also add windy-git to the GitHub->Windy Git sync list; its self-hosted
copy was stuck 3 commits behind (only check + canary workflows, no
deploys, so syncing is safe).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The probe's own log was never retrievable through the jobs API, but the
question it asked was answered better by a direct comparison of two real
workflows on the same runner and image:
windy-git gate @postgres:5432 -> passes its migration round-trip
eternitas migrations @localhost:5432 -> failed
Also scopes test_g73 to workflows that actually run Python. It failed the probe
for not pinning a version when the probe only shelled out to psql — the test
being wrong rather than the workflow.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
The front end was 100% stock Gitea: green teacup, "Gitea: Git with a cup of
tea", "A painless, self-hosted Git service". G2.3 was specified in the plan with
an acceptance test and never executed, and nothing enforced it.
Now: Windy Git name, wind-mark logo, brand-blue accent, and a landing page that
says what this actually is. Uses Gitea's SUPPORTED surface (custom templates +
public assets) so upstream upgrades keep arriving — no source modified (D-2/I-1).
Two traps this cost, both now documented and tested:
1. GITEA__DEFAULT__APP_NAME does not work. Gitea reads APP_NAME from the TOP
LEVEL of app.ini; the env var created a literal [default] section that Gitea
ignores, so the installer's stock APP_NAME kept winning while the config
looked correct. The env-to-ini pass also APPENDED a second APP_NAME rather
than replacing the first — a new variant of the documented G4A.3 trap.
2. Cloudflare caches /assets/* for 6h and no token in this stack can purge, so
the new logo and CSS were invisible while being correct at origin. Brand
assets now carry a VERSION IN THE FILENAME; bump it on every change.
Committed with an idempotent apply.sh, because applying it straight to Veron's
disk first was itself the config-drift trap this project documents: a rebuild
would have silently reverted to stock Gitea.
85 tests green.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
I reintroduced the exact defect I had just criticised. ACTION_BASE listed
"push" and "push.force", but git push goes straight to Gitea over HTTPS and
never touches this API — so nothing records a push, a count would be zero
forever, and enforce() would look up a limit, count nothing, and allow
everything. A silent no-op wearing the costume of a control, made worse by a
config name that implies the protection exists.
Split into ACTION_BASE (actually enforced: repo.create, grant.create) and
NOT_ENFORCED_HERE (push, push.force) with the reason and the remedy written
down: enforcing push velocity needs a Gitea-side pre-receive or push webhook
reporting into agent_actions.
enforce("push") now raises rather than silently allowing, and a test asserts the
two sets stay disjoint.
Found by auditing whether the auth fix could be walked around — every
/api/v1/repos/* route does require a caller, and the only unauthenticated
endpoints are /health, /version and the HMAC-verified webhook.
83 tests green.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
/version went null after a deploy — the exact "service cannot name its own
commit" defect this project was built to prevent, caught by its own honesty
check.
Cause: the sed replaced "" with "" (a no-op when COMMIT_SHA is empty) and the
grep then matched that same empty string, so the guard verified nothing. A build
with no COMMIT_SHA passed and shipped a container reporting commit_sha: null.
Now the build fails loudly instead.
Second cause of the stale deploy, and it was mine: an earlier `git commit
--amend` + force-push rewrote history the Veron deploy checkout was already
sitting on, leaving it divergent so `git pull -q` failed SILENTLY (-q hid
"Need to specify how to reconcile divergent branches"). Two lessons: do not
force-push a branch a deploy checkout tracks, and do not pull with -q in a
deploy script.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
looks_like_ept used jwt.get_unverified_header, which validates the WHOLE token
and therefore rejects anything with a malformed signature segment. Routing
consequently depended on signature well-formedness: an EPT-shaped token with a
bad signature fell through to the HUMAN path, where it was refused for the wrong
reason and — with require_verified_jwt off (dev) — could have been read as a
human identity via its `sub` claim.
Now the header segment is decoded directly, so routing depends only on what the
token CLAIMS to be; whether it is authentic remains verify_ept's job.
Also routes alg:none to the EPT verifier regardless of typ, since a `none`
token is never valid for any caller. Both forged shapes now return 401
ept_invalid — the honest code — instead of 503 "feature not ready".
Found by noticing a forged EPT returned 503 where the verifier should have
answered 401, rather than accepting "it was refused, close enough".
80 tests green.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
Monkeypatched httpx so resolve_passport sees a real revoked trust body
(status=revoked, band=unproven, allowed=[]) and must raise
PassportNotInGoodStanding — proving the WIRING, not just the decision. This is
the path that stops a validly-signed EPT that outlived its passport's
revocation (~365-day tokens).
Live-confirmed alongside: Eternitas refuses to mint EPTs for revoked bots
("credentials are not issued for non-active bots"), so the only exposure was a
pre-existing token — exactly what this now catches. The active agent's EPT still
returns 200. A fully-live revoked test would require revoking a real fleet
passport (destructive), so the wiring is proven deterministically instead.
79 tests green.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
A revoked passport returns HTTP 200, status=revoked, band=unproven,
allowed_actions=[] (verified live 2026-08-13). resolve_passport keyed refusal
only on HTTP 4xx and band=="untrusted", so it returned band 'unproven' and the
agent was seated. Revocation was NOT enforced on the live auth path at all — and
now that agent auth actually works, a revoked agent could authenticate and act.
Extracts decide_trust(body) -> (band, actions) | raise. Only status=="active"
is allowed; revoked/suspended/frozen/unknown all refuse, fail-closed on the
field that carries the most consequential fact about an identity. The agent call
site turns that into a clean 403 passport_revoked.
This is the REAL revocation gate — the token cannot be un-issued, but its
standing is re-checked on every request, so revocation takes effect on the next
call with no webhook required. The Eternitas webhook remains useful for
invalidating locally-issued credentials/grants (G6.3, not built yet), but it was
never the primary gate and its being unwired is no longer a live exposure.
Behavioral tests: revoked body refused, active accepted, unknown/missing status
fails closed.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
REOPENS the agent path — but only because possession is now actually proven.
EPT verification (api/app/ept.py): ES256 against Eternitas's published key set
at /.well-known/eternitas-keys. algorithms=["ES256"] makes alg:none and
algorithm confusion unrepresentable rather than merely unlikely; issuer and exp
are enforced by the library; an unknown kid is refused.
Order is deliberate: signature FIRST, trust lookup second. These EPTs live ~365
days and carry rev/tru baked in at issuance, so a year-old "rev: false" proves
nothing — revocation and band still come from a live lookup on every request.
Found while building it: real EPTs put the passport in the "sub" claim. The old
code read "passport"/"sub_passport", which no genuine EPT carries — so real
agents were never recognised and ONLY forged tokens ever authenticated. The
bypass was not just a hole, it was the only thing that worked.
Throttle (api/app/throttle.py): BAND_MULTIPLIER and rate_*_per_day were defined
and read by nothing. Now enforced on repo.create and grant.create, counted
against agent_actions (one source of truth, not a private counter that drifts
from the audit log). Fails CLOSED — a limiter that fails open protects you until
the moment something is wrong. Untrusted band is 403 read-only, not 429, because
"slow down" would be a lie.
Tests: 14 behavioral, signing real ES256 tokens with a locally-generated key so
they exercise the crypto path with no network dependency — genuine tokens
accepted, and alg:none / foreign key / tampered payload / expired / wrong issuer
/ unknown kid / missing claims all refused. 74 green.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
Verified live 2026-08-13: a forged 'alg:none' token naming a passport lifted
from the logs returned HTTP 200 as that agent. The agent path read the passport
without verifying the EPT signature, asked Eternitas 'is this passport
reputable?', and seated the caller on a yes. That answers reputation, not
possession — anyone who knows a passport number could impersonate that agent on
the public API.
The human path already failed closed for exactly this reason
(require_verified_jwt). The gate was on the wrong path: it sat AFTER the agent
branch returned. The agent path now fails closed too, BEFORE the trust lookup,
so a forged token never even reaches Eternitas. Reopens automatically when the
ES256/JWKS verifier (G3.2/G9.1) exists and this gate consults it.
Adds BEHAVIORAL tests (not string-grep): a forged alg:none token exercised
through the real get_caller must raise, not authenticate. This is the test that
would have caught the bypass; the suite had 86 source-string assertions and
zero that ran the auth decision.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Timer installed with Persistent=true — Grant's workstation is not always on at
04:17, and without that a missed window is silently skipped and the backup
simply never runs.
RESTORE DRILL PASSED, which is the part that matters: pulled a bundle back out
of R2, cloned from it, and the restored HEAD (ffce529) matches live origin/main
exactly — 40 commits, all branches, 50 files.
The August audits found no rehearsed restore anywhere in the ecosystem, for
anything. This is one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An unwritable state path raised and took the whole canary down. That is the
worst possible trade for a monitoring tool: it reports nothing at all, and
reports it silently. State is an optimisation for transition detection; the
probing is the point.
Found by fat-fingering an env var, which is exactly how it would happen in
production.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Read the sender rather than guessing: Eternitas generates the webhook secret at
registration time and pings the URL to prove reachability BEFORE returning that
secret. The ping IS signed — with a secret the receiver cannot possibly hold
yet. Unverifiable by construction, not by oversight.
Accepting it is safe because the event is definitionally a no-op: nothing read,
nothing written, acted:false. Every event that changes anything still requires a
valid HMAC. The alternative, skip_validation:true, would permanently disable
reachability checking for this platform to solve a one-time ordering problem.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Eternitas verifies a webhook URL answers BEFORE issuing the secret that signs
deliveries, so the very first request can never carry a signature — refusing it
makes registration impossible. Real chicken-and-egg, not a reason to disable
validation.
A probe is a request claiming no event and carrying no signature. Answering it
200 is honest: the endpoint exists and is ready. It changes nothing (acted:
false), and anything claiming to BE an event still goes through full HMAC
verification. Registering with skip_validation:true would have permanently
disabled a safety check to solve a one-time ordering problem.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
When a passport is revoked, every credential it holds here dies in one
transaction: tokens revoked, grants revoked. A revocation that takes effect
'eventually' is not a revocation.
Avoids two traps that each cost a sibling service a subscription that looked
wired and never once delivered:
1. strip the 'sha256=' prefix before comparing — comparing the decorated
header against a bare digest returns 401 forever
2. HMAC the RAW REQUEST BYTES, never a re-serialised body — JSON.stringify of
a parsed body reorders keys and changes whitespace, so the digest never
matches what the sender signed
Both fail silently from the sender's side: Eternitas records a delivery, the
receiver records a rejection, nobody notices for weeks.
Unset secret REFUSES rather than accepts — accepting unverified instructions
about identity is worse than missing them. And it never acknowledges a
revocation it could not apply; a 200 there is a security hole reporting success.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Caught by TESTING the alert path instead of assuming it. Without an explicit
User-Agent, urllib sends 'Python-urllib/3.x' and Resend rejects it 403, while
the identical request via curl succeeds.
The failure mode this avoids is the worst one a canary has: it would have
detected every outage correctly and told nobody. Same bot-filtering trap as the
Gitea migrate call earlier today — worth recognising on sight.
Also prints the HTTP body on failure. '403 Forbidden' alone sends you hunting
for a bad key; the body names the real cause.
Verified: alert sent (200).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>