The SOTU scopes the child-process DB bridge as 468 call sites and multi-week.
Measurement says otherwise.
bare node -e '0' on Kit 0 (4 vCPU, load 20): 1.70-1.99s
bare node -e '0' on Veron 1 (24 cores, idle): 0.01s
170x, and node startup is essentially the whole per-query cost — adding pg
connect and a real query to a bare node boot adds only ~0.2-1.2s on top of
~1.8s of interpreter start.
Login makes 9 such forks. 9 x 1.8s = 16s. Observed: 17-25s.
Two independent multipliers compound: the adapter forks per query (9x) and the
box is saturated so each fork costs 1.8s instead of 0.01s (170x).
The consequence: you do not need to touch 468 call sites. querySyncViaChild is
ONE function and its interface does not change — replace the per-query
execFileSync with a persistent worker holding a pg.Pool. All 468 call sites get
faster without being edited, including the mail-lookup that caused today's
outage.
Also measured: 54 containers, 301% CPU of 400%, load 20. dev/demo is 55% of
that — real relief, honestly not a fix. 42 non-dev containers on 4 cores is the
actual condition.
Not implemented here on purpose: sync-over-async with Atomics in the most
critical file in the ecosystem, at the end of a long session, on a box that
already had an outage today, deserves a fresh session and a load test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The retry fix already existed: a parallel session landed c79f196 (#172) with a
test. Kit 0 was ONE COMMIT BEHIND and did not have it — which is the entire
reason the loop ran.
Deployed by ff-merging /root/windy-chat ac61db6 -> c79f196 and rebuilding only
agent-roster (--no-deps). Never reset --hard; that checkout has a documented
history of local edits a hard reset would eat.
account-server CPU 168-210% -> 0.00%
roster failures 74/min -> 0/min
agent chat stopped -> up, healthy
login timeout -> HTTP 200 ~18s
The lesson: a merged fix that has not reached production is not a fix, it is a
belief. That is exactly what both August audits named — nothing checks whether a
decision reached production — arriving as an outage instead of a report finding.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Login was timing out for ~50 minutes on the identity service every Windy
product authenticates against.
Two faults multiplying: windy-agent-roster retried a failing mail-lookup at 74
failures/minute with no backoff, and every call cost 0.6-2.1s because the
postgres adapter forks a node process per query and blocks the event loop.
Together they formed a feedback loop — the container's listening socket showed
Recv-Q 510, connections the kernel accepted that node was too blocked to pick
up. Login sat in that queue.
Resolved by stopping windy-agent-roster: login went from 45s+ timeout to HTTP
200 in ~18s. Restored, not healthy — 18s is the fork-per-query adapter on a
54-container 4-vCPU box, and it is what remains after the loop was removed.
Nothing caught this. The container was (unhealthy) with a failing healthcheck
streak of 74 and no alert exists.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>