docs: turnover for a fresh session
All checks were successful
check / gate (push) Successful in 35s
canary / probe (push) Successful in 8s

State, the immediate task (three-repo localhost->service-name CI fix), the traps
already paid for, and a copy-paste prompt.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
This commit is contained in:
Grant Whitmer
2026-08-14 15:00:28 -04:00
parent 01e36155a3
commit e0d5d118be

113
docs/TURNOVER-2026-08-14.md Normal file
View File

@@ -0,0 +1,113 @@
# Windy Git — turnover, 2026-08-14
Paste the block at the bottom into a fresh terminal. Everything above is context
for whoever reads this file directly.
## Where things stand
Windy Git is **live and in use**: `app.windygit.com` (forge), `api.windygit.com`
(our plane), on **Veron 1** behind a Cloudflare Tunnel, zero inbound ports, $0/mo.
143 repos, 85 tests green, health `ok` on all four checks.
Grant signs in with his existing Windy Word credentials — no second account.
Agents authenticate with real Eternitas EPT signature verification and are
rate-limited by integrity band.
## The immediate next task
**Three repos need a one-line CI fix.** Their workflows reach a Postgres service
at `@localhost:5432`, which works on GitHub-hosted runners (services are
port-mapped to the VM) and fails on ours (the job runs *inside* a container, so
`localhost` is the job itself). The service is reachable as **`postgres`**.
| repo | workflow |
|---|---|
| `windy-mind` | `migrations.yml` |
| `windy-registry` | `ci.yml` |
| `WindyCloud` | `ci.yml` |
`eternitas` was already fixed this way — see **eternitas PR #149** for the exact
shape, including the comment explaining why. Fix must go to **GitHub**, not Windy
Git: the sync runs GitHub → Windy Git and force-pushes over local edits.
Proven by direct comparison, same runner and same `postgres:16-alpine` image:
windy-git's own gate uses `@postgres:5432` and passes its migration round-trip;
eternitas' used `@localhost:5432` and failed.
## Traps that will waste your time
- **`git pull -q` hides errors.** A divergent branch once made a "deploy" run 20
minutes against stale code while reporting success. Use `git fetch && git
merge --ff-only` and read the output.
- **Never force-push a branch a deploy checkout tracks** (`--amend` orphaned
`/srv/windygit/src` once).
- **Gitea's env-to-ini SETS but never UNSETS**, and sometimes *appends a
duplicate*. `GITEA__DEFAULT__APP_NAME` does not work at all — Gitea reads
`APP_NAME` from the **top level** of `app.ini`; the env var creates a literal
`[default]` section it ignores. Edit `app.ini` on the host.
- **Cloudflare caches `/assets/*` for 6h and no token in this stack can purge.**
Version brand asset **filenames** (`theme-windy.v2.css`), not query strings.
- **`base64` wraps at 76 chars** and corrupts long tokens in test commands →
`curl (43)`, phantom HTTP 000. Use `base64 -w0`.
- **Kit 0 is fragile.** 54 containers on 4 vCPU. Two production incidents in two
days, both from *non-production* workloads. Check `uptime` before deploying
anything there, and build before recreating so the swap is seconds.
## Open items, roughly by value
1. The three-repo `localhost` fix above.
2. **Get non-prod work off Kit 0.** 12 dev/demo containers on the box running
identity, the CA, mail, Matrix and the broker. Cost two incidents already;
the postgres-adapter fix would not have prevented either.
3. **Login is ~4–6s** — `postgres-adapter.ts:114` forks a `node -e` process per
query. Measured: node startup alone is 1.7s on Kit 0 vs 0.01s on Veron. The
fix is **one function** (persistent worker + `pg.Pool`), not the "468 call
sites" the SOTU scoped. See `docs/incidents/2026-08-12-login-latency-analysis.md`.
4. **Privileged dind sits beside broad-scoped tokens** on the CI host — Grant's
call, needs a decision not a code change.
5. Push-velocity throttling is declared but unenforceable from our plane (git
push never touches the API); needs a Gitea pre-receive hook.
## Read these first
- `~/.claude/.../memory/project_windy_git.md` — the full record, densest source
- `DNA_STRAND_MASTER_PLAN.md` — D-1…D-9 locked decisions, I-1…I-13 invariants
- `docs/AUDIT-fable-2026-08-13.md` — second-auditor findings and dispositions
- `docs/CUTOVER.md` — the GitHub↔Windy Git migration plan and its one rule
---
## Copy-paste prompt
```
Picking up Windy Git (agent-native code+model host on Veron 1, live at
app.windygit.com). Read these before doing anything:
1. ~/.claude/projects/-home-grantwhitmer/memory/project_windy_git.md
2. ~/windy-git/docs/TURNOVER-2026-08-14.md
3. ~/windy-git/DNA_STRAND_MASTER_PLAN.md (D-1..D-9, I-1..I-13)
Current state: live and in use. Grant signs in with his existing Windy account
(SSO fixed 08-14 across windy-pro PRs #346/#347). Agents authenticate with real
EPT signature verification. 143 repos, 85 tests green, health ok.
TASK: three repos have workflows that reach their Postgres service at
@localhost:5432, which fails on our runner because the job runs inside a
container — the service is reachable as `postgres`. Fix:
windy-mind .github/workflows/migrations.yml
windy-registry .github/workflows/ci.yml
WindyCloud .github/workflows/ci.yml
Copy the exact shape from eternitas PR #149 (already merged), comment included.
The fix must go to GitHub, not Windy Git — the sync is GitHub -> Windy Git and
force-pushes over local edits. Open one PR per repo, then sync and confirm the
migration job actually goes green rather than assuming it.
Ground rules that have already been paid for the hard way:
- verify the WHOLE flow, not the half that curls easily
- never `git pull -q` in a deploy path; it hides errors
- check Kit 0's `uptime` before deploying there; it has had two incidents
in two days from non-production workloads
- build before recreating a container so the swap is seconds
```