Commit Graph

28 Commits

Author SHA1 Message Date
Kit OC5
cbcb4c206b docs: CUTOVER-PRIMARY.md checklist (draft, nothing flipped)
All checks were successful
check / gate (push) Successful in 45s
canary / probe (push) Successful in 7s
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-01 04:03:09 -04:00
Kit OC5
1b552c0882 docs: RESTORE-DRILL.md (state backup + restore procedure)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-01 02:46:53 -04:00
Kit OC5
790d794900 bridge: drop eternitas ci/build from BRIDGE_NO_DAEMON (converted to ci/smoke, #179)
All checks were successful
check / gate (push) Successful in 10s
canary / probe (push) Successful in 6s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 19:51:26 -04:00
Kit OC5
0a57d96f2e bridge + ci-hygiene: Docker-needing CI jobs (option A)
- bridge: BRIDGE_NO_DAEMON names image-build jobs whose name lacks docker
  (default eternitas:ci/build); never posted, like the docker-named ones.
- ci-hygiene: flag docker build/buildx/run/compose, docker-compose and
  docker/build-push-action in workflow steps ("needs docker") with the fix:
  job services: + a no-Docker smoke test; the image builds at deploy.
- test_guards_report: owner column (14ed23a broke it).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 19:09:09 -04:00
Kit OC5
6b0b60eb4a scripts: rerun_ci.sh — re-fire a PR's CI without the web button
Gitea 1.24 has no rerun API and the web button needs Grant's SSO identity.
Guarded branch rewind that the next sync undoes; restores the branch itself
on timeout. Used today for eternitas #166 and windy-mind #131 after the
Veron IO stall.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 14:17:10 -04:00
90643fe48e telemetry step 2: API boot/health + forge.auth.failed (declared)
All checks were successful
check / gate (push) Successful in 25s
canary / probe (push) Successful in 6s
Membrane first: I-2 and MEMBRANE.v1 now list the windy-admin ledger
(POST /v1/events). api/app/telemetry.py: service.boot once per start
(commit_sha omitted when unknown, I-12), an hourly in-process
service.health with the shared keys (requests, errors_5xx/4xx,
refusals_4xx, p95_ms only when there was traffic), and one
forge.auth.failed row per refused request: declared 13-code enum,
http_status, caller class, route TEMPLATE (never the concrete path),
actor_type system with no actor_id (all-lanes rule). No token = nothing
sent or buffered; flush failures keep rows (bounded) and never raise.
Token from root-only /etc/windygit/telemetry.env (optional env_file).
8 behavioural tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:47:29 -04:00
baaa542bae docs: audit disposition 09-23 — R2 god token replaced by bucket-scoped token
All checks were successful
check / gate (push) Successful in 43s
canary / probe (push) Successful in 7s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:15:35 -04:00
dfe5543eda docs: runbook + AGENTS match reality (item 5 of the launch bar)
Some checks failed
check / gate (push) Has been cancelled
canary / probe (push) Successful in 34s
RUNBOOK-VERON: deploy uses fetch + merge --ff-only and api-only rebuilds
(the old text used git pull, contradicting its own warning); new sections
for host timers, CI (6 runners x1, windyadmin-scoped, 90m ceiling, queue
truth in the gitea DB, logs in R2), sign-in posture and break-glass;
standing checkout = OC5. AGENTS.md no longer says GENESIS / no code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 03:19:12 -04:00
45686283be ci: bound CI storage; don't bridge image-build jobs
- deploy/runner/prune.sh + windygit-ci-prune.timer (6h): age-based prune
  of the CI-only dind (containers, finished-job volumes, images/builder
  cache >7d) plus a hard 60 GB cap. Only that daemon, over its own TCP
  socket; never the host's Docker. It was 38 GB and unbounded — the same
  class of growth that filled Kit 0 on 09-01.
- pr_status_bridge: jobs named *docker* are not posted. Job containers
  have no daemon by design (I-5), so they are red on every commit; a
  permanent red X teaches everyone to ignore red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:47:28 -04:00
dcf9286f16 ci: make Windy Git CI permanent for the private repos
All checks were successful
check / gate (push) Successful in 21s
canary / probe (push) Successful in 6s
- Four runners x capacity 1 instead of one x capacity 4. Concurrent jobs in
  one act_runner share /root/.cache/act; a refresh racing a copy killed 3 of
  windy-chat's ~20 jobs at setup-node (lstat ... no such file). Separate
  processes have separate caches. Same parallelism, same capped dind.
- Behavioral tests for pr_status_bridge (latest verdict wins, no reposting,
  skipped never painted green, fork PRs never run, pagination, PR lifecycle).
- import_from_github.py reads IMPORT_GITEA_URL, not GITEA_BASE_URL: sourcing
  the deploy .env pointed it at http://gitea:3000 and it died on DNS after the
  mirror it replaces had already been deleted.
- CUTOVER.md: the private-repo CI path, onboarding steps, and the
  /actions/tasks-hides-queued-runs trap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:18:04 -04:00
390c1e7479 ops: move tunnel metrics to 2001, sync windy-git into itself
Some checks failed
check / gate (push) Successful in 19s
canary / probe (push) Failing after 7s
windygit-tunnel had crash-looped ~91k times: another project's
cornercall-tunnel holds 127.0.0.1:2000, and cloudflared exits when it
cannot bind its metrics port. Ingress only survived because a stray
cloudflared.service ran the same config. That unit is now disabled and
/etc/cloudflared/config.yml uses metrics 127.0.0.1:2001.

Also add windy-git to the GitHub->Windy Git sync list; its self-hosted
copy was stuck 3 commits behind (only check + canary workflows, no
deploys, so syncing is safe).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 01:29:57 -04:00
2b30b0ac99 docs: record the runner bump, the act cache flake, and the remaining reds
Second root cause found and fixed: act_runner 0.2.11 predates
`runs.using: node24`, so any repo on actions/checkout@v5 died before its first
step. Bumped to 0.6.1; Windy-Clone went 4/4 red to 4/4 green.

Also upgrades the act-cache note from "watch item" to a confirmed job failure
(lstat on a vanished file mid-tar), with the wipe command and the annotated-tag
dead end that looks like a wrong checkout but isn't.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 19:09:26 -04:00
9fc27eabd6 docs: record the real CI root cause — setup-uv resolves uv via the forge API
The localhost->postgres fix was correct but was never what failed these jobs;
they died at step 2. setup-uv v4+ resolves "latest" through GITHUB_API_URL,
which act_runner points at our own forge, so it 404s and every later step is
skipped by success(). Pinned an explicit uv version across 11 repos.

Also records the diagnosis traps that cost the most time: Gitea's job status
enum (1=success, 2=failure), act misattributing the error to the previous step,
the jobs-log API needing a repo-matched id, and Secure cookies defeating a
loopback curl login.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 17:57:04 -04:00
Grant Whitmer
db055a1922 docs: refresh the turnover prompt for the actual next task
Some checks failed
check / gate (push) Successful in 18s
canary / probe (push) Failing after 6s
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 17:31:23 -04:00
Grant Whitmer
86326ca6a7 docs: record three-repo CI fix results — 1 of 3 verified
All checks were successful
check / gate (push) Successful in 19s
windy-registry's postgres integration went failure -> success, proving the fix.
windy-mind and WindyCloud still fail for a cause I could not determine: the
jobs API returns 'job not found' for the ids the runs report, so logs were not
retrievable that way. Next session should read them from the Gitea web UI.

Records the trap that the three repos did NOT share one pattern — a naive
localhost->postgres swap would have left WindyCloud on port 15432.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 17:30:44 -04:00
Grant Whitmer
e0d5d118be docs: turnover for a fresh session
All checks were successful
check / gate (push) Successful in 35s
canary / probe (push) Successful in 8s
State, the immediate task (three-repo localhost->service-name CI fix), the traps
already paid for, and a copy-paste prompt.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 15:00:28 -04:00
Grant Whitmer
25368547cb docs: runbook — never pull -q, never force-push a deployed branch
All checks were successful
check / gate (push) Successful in 21s
canary / probe (push) Successful in 11s
Two deploy traps paid for on 2026-08-14: 'git pull -q' hid a divergent-branch
error so a deploy ran against stale code while reporting success, and the
divergence came from amending a commit a deploy checkout already tracked.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:28:57 -04:00
Grant Whitmer
830ca48705 docs: mark revocation finding resolved — it was critical, not low
All checks were successful
check / gate (push) Successful in 19s
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:20:46 -04:00
Grant Whitmer
c257cc55c6 docs: second-auditor review (Fable) — 1 critical fixed live, 4 open
All checks were successful
check / gate (push) Successful in 20s
canary / probe (push) Successful in 11s
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 22:54:15 -04:00
Grant Whitmer
83047def94 docs: whole account on Windy Git — 143 repos, 1.58 GB, two tiers
131 read-only mirrors (cannot run Actions, zero deploy risk) + 12 writable with
CI and deploys disabled. Total size matches the measured GitHub archive exactly,
which is the confirmation the copy is complete.

Splits the two concerns cleanly: having a copy is safe and should cover
everything now; running code needs judgement and happens per repo.

windy-pro IS included as a mirror — the G11.5 caution is about making it
writable while six checkouts disagree on HEAD, not about holding a read-only
copy. The DR copy is now complete.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:49:48 -04:00
Grant Whitmer
89723c6ebd safety: disable deploy workflows on Windy Git before they can fire
Six workflows deploy to production on push:. Windy Git now has a working
runner, so the next synced commit to main would have attempted a production
deploy FROM VERON 1. Their secrets are unset here so they would have failed —
but loudly, on every push, with any pre-SSH step still running.

All six now disabled_manually. Tests, lints and migration checks stay active:
they need no secrets, which is exactly why Phase 1 delivers CI value with
nothing to configure.

Same class of mistake as the push-mirror direction, caught before firing this
time rather than after.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:30:56 -04:00
Grant Whitmer
b2f00821d7 docs: replace the cutover plan with the phased one that matches reality
Phase 1 requires nothing from anyone: agents keep pushing to GitHub, a timer
syncs GitHub -> Windy Git every 15 minutes, CI runs on Veron against current
code. Phase 2 flips one repo at a time, only when that repo is idle.

Records the direction mistake honestly: the source of truth is wherever people
are actually typing, not wherever the plan says it should be.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:14:21 -04:00
Grant Whitmer
7711adcb97 docs: cutover — Windy Git is the daily driver
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 11s
9 repos migrated writable with push-mirrors to GitHub (sync_on_commit).
Full loop proven end to end: pushed a commit to Windy Git, its existing
workflow ran on Veron 1, and GitHub received the commit within 20s.

Documents the one rule the cutover creates: do NOT push directly to GitHub for
a migrated repo. The mirror makes GitHub match Windy Git, so a direct commit is
overwritten on the next sync, silently, with no conflict. One writer is the
point — two writers with no reconciliation is how you lose work you thought was
saved.

windy-pro deliberately excluded until its six-checkout / forked-counter
ambiguity is resolved by reading (G11.5).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 21:54:47 -04:00
Grant Whitmer
e3677dddd9 docs: measure WHY login takes 18s — it is not 468 call sites
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 24s
The SOTU scopes the child-process DB bridge as 468 call sites and multi-week.
Measurement says otherwise.

  bare node -e '0'  on Kit 0 (4 vCPU, load 20): 1.70-1.99s
  bare node -e '0'  on Veron 1 (24 cores, idle): 0.01s

170x, and node startup is essentially the whole per-query cost — adding pg
connect and a real query to a bare node boot adds only ~0.2-1.2s on top of
~1.8s of interpreter start.

Login makes 9 such forks. 9 x 1.8s = 16s. Observed: 17-25s.

Two independent multipliers compound: the adapter forks per query (9x) and the
box is saturated so each fork costs 1.8s instead of 0.01s (170x).

The consequence: you do not need to touch 468 call sites. querySyncViaChild is
ONE function and its interface does not change — replace the per-query
execFileSync with a persistent worker holding a pg.Pool. All 468 call sites get
faster without being edited, including the mail-lookup that caused today's
outage.

Also measured: 54 containers, 301% CPU of 400%, load 20. dev/demo is 55% of
that — real relief, honestly not a fix. 42 non-dev containers on 4 cores is the
actual condition.

Not implemented here on purpose: sync-over-async with Atomics in the most
critical file in the ecosystem, at the end of a long session, on a box that
already had an outage today, deserves a fresh session and a load test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 19:05:25 -04:00
Grant Whitmer
211a48187f docs: incident RESOLVED — roster backoff deployed, agent chat back
All checks were successful
check / gate (push) Successful in 18s
The retry fix already existed: a parallel session landed c79f196 (#172) with a
test. Kit 0 was ONE COMMIT BEHIND and did not have it — which is the entire
reason the loop ran.

Deployed by ff-merging /root/windy-chat ac61db6 -> c79f196 and rebuilding only
agent-roster (--no-deps). Never reset --hard; that checkout has a documented
history of local edits a hard reset would eat.

  account-server CPU   168-210%  ->  0.00%
  roster failures      74/min    ->  0/min
  agent chat           stopped   ->  up, healthy
  login                timeout   ->  HTTP 200 ~18s

The lesson: a merged fix that has not reached production is not a fix, it is a
belief. That is exactly what both August audits named — nothing checks whether a
decision reached production — arriving as an outage instead of a report finding.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:19:58 -04:00
Grant Whitmer
3926726309 docs: incident record — account-server login outage 2026-08-12
All checks were successful
check / gate (push) Successful in 18s
Login was timing out for ~50 minutes on the identity service every Windy
product authenticates against.

Two faults multiplying: windy-agent-roster retried a failing mail-lookup at 74
failures/minute with no backoff, and every call cost 0.6-2.1s because the
postgres adapter forks a node process per query and blocks the event loop.
Together they formed a feedback loop — the container's listening socket showed
Recv-Q 510, connections the kernel accepted that node was too blocked to pick
up. Login sat in that queue.

Resolved by stopping windy-agent-roster: login went from 45s+ timeout to HTTP
200 in ~18s. Restored, not healthy — 18s is the fork-per-query adapter on a
54-container 4-vCPU box, and it is what remains after the loop was removed.

Nothing caught this. The container was (unhealthy) with a failing healthcheck
streak of 74 and no alert exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:00:54 -04:00
Grant Whitmer
a68261a563 G1: Veron 1 host live behind Cloudflare Tunnel
app.windygit.com / api.windygit.com / models.windygit.com are serving over
HTTPS with ZERO inbound ports open on Grant's network.

  - tunnel 4e856c5d, 4 registered edge connections, systemd-managed and bounded
  - three proxied single-level CNAMEs (Free Universal SSL covers them; a
    two-level name would need ACM and would die in the TLS handshake)
  - services bound to 127.0.0.1 with configurable host ports — Veron 1 is
    Grant's workstation and 3000/3300 belong to other projects
  - docs/RUNBOOK-VERON.md

I-12 PROVEN IN PRODUCTION: /version reports source=baked with a sha equal to
the deployed HEAD.

Also fixed: the tunnel health probe targeted localhost from inside a container,
so it was permanently red. A check that is always red is as useless as one that
is always green — it is how a fleet canary goes 37 days dead unnoticed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 14:34:18 -04:00
Grant Whitmer
659991b2bd G0: cell substrate — invariants made executable
Strand G0 complete and VERIFIED against real Postgres, not asserted.

  - FastAPI plane, fail-closed provider seams, repair-pointer error taxonomy
  - migration 001: all 10 tables incl. repo_type NOT NULL and model_cards (I-7)
  - 17 invariant tests, ruff clean, vocabulary audit clean

Two bugs found by RUNNING it that review would not have caught:

  1. SQLAlchemy Enum persists .name, not .value — so RepoState.deleted_soft
     and CreatedVia.imported would have written labels migration 001 never
     declared, failing at runtime rather than at review. Pinned via
     values_callable.
  2. op.create_table asks each Enum to emit its own CREATE TYPE with no
     checkfirst, so the second reference raised DuplicateObject and the
     migration died halfway. Types are now created once, referenced with
     create_type=False.

Proven live, with the hostile env var set:
  - I-12: COMMIT_SHA=deadbeef... in the environment, /version reports real HEAD.
    That env pin is the documented root cause of nine sibling services
    misreporting their commit; here it is structurally ignored.
  - I-8: three unconfigured providers -> status degraded, HTTP 503, each saying
    'refusing to report healthy'. No mock, no false green.
  - G0.4: upgrade -> downgrade -> upgrade round-trip clean (10 -> 0 -> 10).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 14:19:28 -04:00