The retry fix already existed: a parallel session landed c79f196 (#172) with a
test. Kit 0 was ONE COMMIT BEHIND and did not have it — which is the entire
reason the loop ran.
Deployed by ff-merging /root/windy-chat ac61db6 -> c79f196 and rebuilding only
agent-roster (--no-deps). Never reset --hard; that checkout has a documented
history of local edits a hard reset would eat.
account-server CPU 168-210% -> 0.00%
roster failures 74/min -> 0/min
agent chat stopped -> up, healthy
login timeout -> HTTP 200 ~18s
The lesson: a merged fix that has not reached production is not a fix, it is a
belief. That is exactly what both August audits named — nothing checks whether a
decision reached production — arriving as an outage instead of a report finding.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Login was timing out for ~50 minutes on the identity service every Windy
product authenticates against.
Two faults multiplying: windy-agent-roster retried a failing mail-lookup at 74
failures/minute with no backoff, and every call cost 0.6-2.1s because the
postgres adapter forks a node process per query and blocks the event loop.
Together they formed a feedback loop — the container's listening socket showed
Recv-Q 510, connections the kernel accepted that node was too blocked to pick
up. Login sat in that queue.
Resolved by stopping windy-agent-roster: login went from 45s+ timeout to HTTP
200 in ~18s. Restored, not healthy — 18s is the fork-per-query adapter on a
54-container 4-vCPU box, and it is what remains after the loop was removed.
Nothing caught this. The container was (unhealthy) with a failing healthcheck
streak of 74 and no alert exists.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>