Commit Graph

65 Commits

Author SHA1 Message Date
dcf9286f16 ci: make Windy Git CI permanent for the private repos
All checks were successful
check / gate (push) Successful in 21s
canary / probe (push) Successful in 6s
- Four runners x capacity 1 instead of one x capacity 4. Concurrent jobs in
  one act_runner share /root/.cache/act; a refresh racing a copy killed 3 of
  windy-chat's ~20 jobs at setup-node (lstat ... no such file). Separate
  processes have separate caches. Same parallelism, same capped dind.
- Behavioral tests for pr_status_bridge (latest verdict wins, no reposting,
  skipped never painted green, fork PRs never run, pagination, PR lifecycle).
- import_from_github.py reads IMPORT_GITEA_URL, not GITEA_BASE_URL: sourcing
  the deploy .env pointed it at http://gitea:3000 and it died on DNS after the
  mirror it replaces had already been deleted.
- CUTOVER.md: the private-repo CI path, onboarding steps, and the
  /actions/tasks-hides-queued-runs trap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:18:04 -04:00
1b09b9b0d3 ci: sync windy-chat + windy-mail, bridge PR CI verdicts back to GitHub
All checks were successful
check / gate (push) Successful in 20s
canary / probe (push) Successful in 6s
GitHub Actions can't run on the private platform repos. Windy Git already
has their code and a working runner, so:

- windy-chat and windy-mail were read-only pull mirrors (which can never
  run Actions); they are now writable, deploy.yml/build-image.yml disabled,
  and synced from GitHub like the others.
- scripts/pr_status_bridge.py mirrors open same-repo GitHub PRs into Windy
  Git (so pull_request workflows fire) and posts each job's result back as
  a GitHub commit status (windy-git/<workflow>/<job>) on PR heads and the
  default-branch head. Fork PRs are never run. Runs after every sync.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:13:21 -04:00
390c1e7479 ops: move tunnel metrics to 2001, sync windy-git into itself
Some checks failed
check / gate (push) Successful in 19s
canary / probe (push) Failing after 7s
windygit-tunnel had crash-looped ~91k times: another project's
cornercall-tunnel holds 127.0.0.1:2000, and cloudflared exits when it
cannot bind its metrics port. Ingress only survived because a stray
cloudflared.service ran the same config. That unit is now disabled and
/etc/cloudflared/config.yml uses metrics 127.0.0.1:2001.

Also add windy-git to the GitHub->Windy Git sync list; its self-hosted
copy was stuck 3 commits behind (only check + canary workflows, no
deploys, so syncing is safe).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 01:29:57 -04:00
Grant Whitmer
c96d1d3102 canary: guard both forgery shapes now that real verification is live
All checks were successful
check / gate (push) Successful in 20s
The EPT-shaped forgery is the one that matters after G3.2 — it is what
signature verification actually guards. The JWT-shaped one still exercises the
human gate. Both must_refuse; a 2xx on either pages.

Verified live: forged EPT -> 401 ept_invalid, forged JWT -> 503,
genuine EPT -> 200.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:13:44 -04:00
Grant Whitmer
85fa65a52b canary: continuously verify the forged-token bypass stays closed
All checks were successful
check / gate (push) Successful in 18s
Adds must_refuse checks: a 2xx is the alarm, a 401/403/503 is health. The
forged alg:none agent token is probed every 10 min; if it ever returns 2xx the
canary goes red and pages. Proven both ways — 503 reads ok, a 200 endpoint
alarms 'ACCEPTED — this MUST be refused'. The security property is now enforced
by a running check, not assumed.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 22:52:56 -04:00
Grant Whitmer
711c47d8c9 fix: the --all-as-mirrors flag itself was never added
The function landed but the argparse anchor did not match the real formatting,
so the command existed and could not be invoked. Caught by running it rather
than assuming the patch applied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:43:22 -04:00
Grant Whitmer
ae6ca8b9d2 G11.3: bulk DR copy — every repo as a read-only mirror
Grant's plan: clone the whole account, let it circulate, reverse direction later
when things are clean. Right plan, with one change that matters.

Sampling 40 repos found 18 carrying deploy/release/publish workflows that
trigger on push: — roughly 63 across the account. Importing those writable with
Actions enabled would arm sixty-odd production deploy triggers on Veron 1, each
needing disarming by hand.

So the bulk goes in as READ-ONLY pull mirrors. A mirror cannot run Actions at
all, so this carries zero deploy risk, and Gitea syncs them itself with no
script and no timer. What you get is a complete, current second copy of the
whole account — the disaster-recovery half — with none of the execution risk.

Converting one to writable + CI stays a deliberate per-repo act: re-import,
review its workflows, disable the deploying ones. That judgement belongs at the
moment you want CI on that repo, not in bulk sixty times by accident.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:40:52 -04:00
Grant Whitmer
51acf9f86e URGENT FIX: reverse the sync — GitHub is the source of truth, not Windy Git
I migrated nine repos writable with push-mirrors pointed AT GitHub. That was
wrong for the actual situation: a dozen agent sessions on the Mac mini are
pushing to GitHub continuously, so GitHub is where the live work is.

A push-mirror force-updates refs. On its 8-hour timer it would have pushed
Windy Git's stale copy over live work — silently, no conflict, nothing to
notice. Removed all nine before the first timer fired; verified no GitHub repo
had been touched (latest push predated the mirrors).

Replaced with the correct Phase 1 direction:

  agents --push--> GitHub --sync--> Windy Git --> CI on Veron

It requires NOTHING from anyone. No remote changes, no coordination, no
'everybody stop pushing'. Agents keep working exactly as they are and CI starts
running on 24 cores.

Windy Git is force-updated on purpose: in Phase 1 it holds nothing anyone
depends on, so GitHub always wins and there is no merge to reconcile.

Phase 2 is per-repo and only when that repo is idle. Never a big-bang cutover
across a dozen live sessions.

Fetches +refs/heads/* and tags explicitly rather than --mirror, which would drag
GitHub's refs/pull/* that Gitea rejects and bury the real errors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:11:34 -04:00
Grant Whitmer
ffce529344 G0.9: nightly backup — the prerequisite for becoming the daily driver
All checks were successful
check / gate (push) Successful in 23s
Today GitHub is authoritative, so losing Veron 1 costs nothing. The moment
people push HERE first that inverts: Veron 1 holds the only current copy of the
company's source between mirror syncs, and it is Grant's workstation — no SLA,
no snapshots, residential line, and he reboots it.

git bundle over tar, deliberately: a bundle is one file that git clone reads
directly, so a restore needs no knowledge of Gitea's on-disk layout, and
bundling asks git for a consistent view instead of racing a live push.

  - --all, so every branch and tag is captured. A single-branch bundle loses
    the rest silently and you find out during the restore.
  - git bundle verify before upload. An unverified bundle is a belief.
  - empty repos are reported as skipped, not counted as failures
  - non-zero exit on ANY failure so the unit goes red. A backup script that
    swallows errors manufactures confidence.

Whole archive measured 1.58 GB across 141 repos — about two cents a month.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 21:47:46 -04:00
Grant Whitmer
db31e786e0 G7.4: default to writable repos — mirrors cannot run CI (measured)
All checks were successful
check / gate (push) Successful in 37s
canary / probe (push) Successful in 23s
windy-calendar imported as a pull mirror sat at 0 workflow runs. Gitea does not
fire Actions on mirror sync and a mirror is not a push target, so a mirrored
repo gives you the code and none of the point.

Re-imported it writable, pushed a commit, and its EXISTING .github/workflows/
ci.yml ran on Veron 1 and reported success — with zero workflow edits. That
workflow's own header says it was written for 'OUR self-hosted runners (Kit 0 +
Veron One)' because the account is billing-locked. The fleet already wanted
this; the runners just died when the repos went private.

Default is now writable + push-mirror to GitHub (I-4 steady state). --mirror
stays available for repos to copy but not move.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 19:58:15 -04:00
Grant Whitmer
ef3450d87b G7.6: never let bookkeeping kill the monitor
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 26s
An unwritable state path raised and took the whole canary down. That is the
worst possible trade for a monitoring tool: it reports nothing at all, and
reports it silently. State is an optimisation for transition detection; the
probing is the point.

Found by fat-fingering an env var, which is exactly how it would happen in
production.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 15:42:58 -04:00
Grant Whitmer
fc1937560c G7.6: fix the alert path — urllib UA was rejected 403 by Resend
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 22s
Caught by TESTING the alert path instead of assuming it. Without an explicit
User-Agent, urllib sends 'Python-urllib/3.x' and Resend rejects it 403, while
the identical request via curl succeeds.

The failure mode this avoids is the worst one a canary has: it would have
detected every outage correctly and told nobody. Same bot-filtering trap as the
Gitea migrate call earlier today — worth recognising on sight.

Also prints the HTTP body on failure. '403 Forbidden' alone sends you hunting
for a bad key; the body names the real cause.

Verified: alert sent (200).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:44:00 -04:00
Grant Whitmer
9a7030351b G7.6: fleet canary — probes what a user does, not what is cheap
All checks were successful
check / gate (push) Successful in 19s
Today's outage is the whole design brief: /health returned 200 for the entire
hour that login was dead. A canary watching /health would have stayed green
while nobody in the ecosystem could sign in. So the login probe is here and it
is the one that matters.

Three rules it obeys:
  - never green for something it did not prove (I-8)
  - alert on TRANSITIONS, not every run — a canary people filter is a dead
    canary, which is how the last one sat 37 days dead unnoticed
  - run where the watched thing cannot take it down: Veron 1, never Kit 0

Two independent signals, so losing one still leaves the other: an email via
Resend on state change, and a non-zero exit that turns the CI run red in the
forge itself.

Alerts say what broke in human terms — 'a human can actually sign in' — rather
than only naming an endpoint.

Verified against production: 7/7 green including login at 17.1s; a forced 404
reports DOWN; a 1s threshold reports SLOW at 23.3s.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:41:37 -04:00
Grant Whitmer
b6a7045907 G7.4/G11.3: import tool — GitHub upstream, Windy Git downstream, CI here
All checks were successful
check / gate (push) Successful in 17s
Direction is deliberate. For the migration quarter GitHub stays upstream and
Windy Git is a PULL mirror, so it cannot diverge — worst case it is stale, not
wrong. Making Windy Git authoritative before Grant flips G11.6 would create a
two-writer problem nobody asked for. I-4's push-mirror is the steady state for
repos that originate here.

Refuses windy-pro by name: six checkouts, a build counter forked three ways, and
two sessions recording different HEADs hours apart. That gets resolved by
reading, not by importing (G11.5).

Records the Cloudflare trap: the public endpoint answers 403 error 1010 because
CF blocks urllib's user-agent as a bot signature. It reads like Gitea rejecting
the token and is not — the identical call against localhost:3080 on the host
succeeds. Bulk import belongs on the host anyway.

Verified: windy-calendar imported, 731 KB, mirroring.

Survey: 135 private repos, and GitHub Actions cannot run on ANY of them. Sampled
windy-pro, windy-mind, eternitas, windy-registry, windy-agent, WindyCloud — every
recent run is startup_failure, as recently as today.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 12:12:02 -04:00
Grant Whitmer
659991b2bd G0: cell substrate — invariants made executable
Strand G0 complete and VERIFIED against real Postgres, not asserted.

  - FastAPI plane, fail-closed provider seams, repair-pointer error taxonomy
  - migration 001: all 10 tables incl. repo_type NOT NULL and model_cards (I-7)
  - 17 invariant tests, ruff clean, vocabulary audit clean

Two bugs found by RUNNING it that review would not have caught:

  1. SQLAlchemy Enum persists .name, not .value — so RepoState.deleted_soft
     and CreatedVia.imported would have written labels migration 001 never
     declared, failing at runtime rather than at review. Pinned via
     values_callable.
  2. op.create_table asks each Enum to emit its own CREATE TYPE with no
     checkfirst, so the second reference raised DuplicateObject and the
     migration died halfway. Types are now created once, referenced with
     create_type=False.

Proven live, with the hostile env var set:
  - I-12: COMMIT_SHA=deadbeef... in the environment, /version reports real HEAD.
    That env pin is the documented root cause of nine sibling services
    misreporting their commit; here it is structurally ignored.
  - I-8: three unconfigured providers -> status degraded, HTTP 503, each saying
    'refusing to report healthy'. No mock, no false green.
  - G0.4: upgrade -> downgrade -> upgrade round-trip clean (10 -> 0 -> 10).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 14:19:28 -04:00