Commit Graph

4 Commits

Author SHA1 Message Date
Grant Whitmer
211a48187f docs: incident RESOLVED — roster backoff deployed, agent chat back
All checks were successful
check / gate (push) Successful in 18s
The retry fix already existed: a parallel session landed c79f196 (#172) with a
test. Kit 0 was ONE COMMIT BEHIND and did not have it — which is the entire
reason the loop ran.

Deployed by ff-merging /root/windy-chat ac61db6 -> c79f196 and rebuilding only
agent-roster (--no-deps). Never reset --hard; that checkout has a documented
history of local edits a hard reset would eat.

  account-server CPU   168-210%  ->  0.00%
  roster failures      74/min    ->  0/min
  agent chat           stopped   ->  up, healthy
  login                timeout   ->  HTTP 200 ~18s

The lesson: a merged fix that has not reached production is not a fix, it is a
belief. That is exactly what both August audits named — nothing checks whether a
decision reached production — arriving as an outage instead of a report finding.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:19:58 -04:00
Grant Whitmer
3926726309 docs: incident record — account-server login outage 2026-08-12
All checks were successful
check / gate (push) Successful in 18s
Login was timing out for ~50 minutes on the identity service every Windy
product authenticates against.

Two faults multiplying: windy-agent-roster retried a failing mail-lookup at 74
failures/minute with no backoff, and every call cost 0.6-2.1s because the
postgres adapter forks a node process per query and blocks the event loop.
Together they formed a feedback loop — the container's listening socket showed
Recv-Q 510, connections the kernel accepted that node was too blocked to pick
up. Login sat in that queue.

Resolved by stopping windy-agent-roster: login went from 45s+ timeout to HTTP
200 in ~18s. Restored, not healthy — 18s is the fork-per-query adapter on a
54-container 4-vCPU box, and it is what remains after the loop was removed.

Nothing caught this. The container was (unhealthy) with a failing healthcheck
streak of 74 and no alert exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:00:54 -04:00
Grant Whitmer
a68261a563 G1: Veron 1 host live behind Cloudflare Tunnel
app.windygit.com / api.windygit.com / models.windygit.com are serving over
HTTPS with ZERO inbound ports open on Grant's network.

  - tunnel 4e856c5d, 4 registered edge connections, systemd-managed and bounded
  - three proxied single-level CNAMEs (Free Universal SSL covers them; a
    two-level name would need ACM and would die in the TLS handshake)
  - services bound to 127.0.0.1 with configurable host ports — Veron 1 is
    Grant's workstation and 3000/3300 belong to other projects
  - docs/RUNBOOK-VERON.md

I-12 PROVEN IN PRODUCTION: /version reports source=baked with a sha equal to
the deployed HEAD.

Also fixed: the tunnel health probe targeted localhost from inside a container,
so it was permanently red. A check that is always red is as useless as one that
is always green — it is how a fleet canary goes 37 days dead unnoticed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 14:34:18 -04:00
Grant Whitmer
659991b2bd G0: cell substrate — invariants made executable
Strand G0 complete and VERIFIED against real Postgres, not asserted.

  - FastAPI plane, fail-closed provider seams, repair-pointer error taxonomy
  - migration 001: all 10 tables incl. repo_type NOT NULL and model_cards (I-7)
  - 17 invariant tests, ruff clean, vocabulary audit clean

Two bugs found by RUNNING it that review would not have caught:

  1. SQLAlchemy Enum persists .name, not .value — so RepoState.deleted_soft
     and CreatedVia.imported would have written labels migration 001 never
     declared, failing at runtime rather than at review. Pinned via
     values_callable.
  2. op.create_table asks each Enum to emit its own CREATE TYPE with no
     checkfirst, so the second reference raised DuplicateObject and the
     migration died halfway. Types are now created once, referenced with
     create_type=False.

Proven live, with the hostile env var set:
  - I-12: COMMIT_SHA=deadbeef... in the environment, /version reports real HEAD.
    That env pin is the documented root cause of nine sibling services
    misreporting their commit; here it is structurally ignored.
  - I-8: three unconfigured providers -> status degraded, HTTP 503, each saying
    'refusing to report healthy'. No mock, no false green.
  - G0.4: upgrade -> downgrade -> upgrade round-trip clean (10 -> 0 -> 10).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 14:19:28 -04:00