Commit Graph

3 Commits

Author SHA1 Message Date
Grant Whitmer
e3677dddd9 docs: measure WHY login takes 18s — it is not 468 call sites
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 24s
The SOTU scopes the child-process DB bridge as 468 call sites and multi-week.
Measurement says otherwise.

  bare node -e '0'  on Kit 0 (4 vCPU, load 20): 1.70-1.99s
  bare node -e '0'  on Veron 1 (24 cores, idle): 0.01s

170x, and node startup is essentially the whole per-query cost — adding pg
connect and a real query to a bare node boot adds only ~0.2-1.2s on top of
~1.8s of interpreter start.

Login makes 9 such forks. 9 x 1.8s = 16s. Observed: 17-25s.

Two independent multipliers compound: the adapter forks per query (9x) and the
box is saturated so each fork costs 1.8s instead of 0.01s (170x).

The consequence: you do not need to touch 468 call sites. querySyncViaChild is
ONE function and its interface does not change — replace the per-query
execFileSync with a persistent worker holding a pg.Pool. All 468 call sites get
faster without being edited, including the mail-lookup that caused today's
outage.

Also measured: 54 containers, 301% CPU of 400%, load 20. dev/demo is 55% of
that — real relief, honestly not a fix. 42 non-dev containers on 4 cores is the
actual condition.

Not implemented here on purpose: sync-over-async with Atomics in the most
critical file in the ecosystem, at the end of a long session, on a box that
already had an outage today, deserves a fresh session and a load test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 19:05:25 -04:00
Grant Whitmer
211a48187f docs: incident RESOLVED — roster backoff deployed, agent chat back
All checks were successful
check / gate (push) Successful in 18s
The retry fix already existed: a parallel session landed c79f196 (#172) with a
test. Kit 0 was ONE COMMIT BEHIND and did not have it — which is the entire
reason the loop ran.

Deployed by ff-merging /root/windy-chat ac61db6 -> c79f196 and rebuilding only
agent-roster (--no-deps). Never reset --hard; that checkout has a documented
history of local edits a hard reset would eat.

  account-server CPU   168-210%  ->  0.00%
  roster failures      74/min    ->  0/min
  agent chat           stopped   ->  up, healthy
  login                timeout   ->  HTTP 200 ~18s

The lesson: a merged fix that has not reached production is not a fix, it is a
belief. That is exactly what both August audits named — nothing checks whether a
decision reached production — arriving as an outage instead of a report finding.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:19:58 -04:00
Grant Whitmer
3926726309 docs: incident record — account-server login outage 2026-08-12
All checks were successful
check / gate (push) Successful in 18s
Login was timing out for ~50 minutes on the identity service every Windy
product authenticates against.

Two faults multiplying: windy-agent-roster retried a failing mail-lookup at 74
failures/minute with no backoff, and every call cost 0.6-2.1s because the
postgres adapter forks a node process per query and blocks the event loop.
Together they formed a feedback loop — the container's listening socket showed
Recv-Q 510, connections the kernel accepted that node was too blocked to pick
up. Login sat in that queue.

Resolved by stopping windy-agent-roster: login went from 45s+ timeout to HTTP
200 in ~18s. Restored, not healthy — 18s is the fork-per-query adapter on a
54-container 4-vCPU box, and it is what remains after the loop was removed.

Nothing caught this. The container was (unhealthy) with a failing healthcheck
streak of 74 and no alert exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 13:00:54 -04:00