Files
windy-git/docs/TURNOVER-2026-08-14.md
Grant Whitmer db055a1922
Some checks failed
check / gate (push) Successful in 18s
canary / probe (push) Failing after 6s
docs: refresh the turnover prompt for the actual next task
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 17:31:23 -04:00

146 lines
7.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Windy Git — turnover, 2026-08-14
Paste the block at the bottom into a fresh terminal. Everything above is context
for whoever reads this file directly.
## Where things stand
Windy Git is **live and in use**: `app.windygit.com` (forge), `api.windygit.com`
(our plane), on **Veron 1** behind a Cloudflare Tunnel, zero inbound ports, $0/mo.
143 repos, 85 tests green, health `ok` on all four checks.
Grant signs in with his existing Windy Word credentials — no second account.
Agents authenticate with real Eternitas EPT signature verification and are
rate-limited by integrity band.
## DONE since this was written — the three-repo CI fix
All three PRs are **merged and synced**: windy-mind #100, WindyCloud #89,
windy-registry #31 (eternitas #149 earlier).
**Result: 1 of 3 verified fixed, 2 still failing for an undetermined reason.**
- ✅ **windy-registry** — `postgres integration` went **failure → success**. The
fix is proven correct.
- ❌ **windy-mind**, **WindyCloud** — `migrations` still fails. The DATABASE_URL
is definitely right now; the cause is something else and was **not
determined** — the jobs API returns "job not found" for the ids the runs
report, so logs could not be retrieved that way.
**Next session: read those job logs from the Gitea web UI** (`app.windygit.com`
→ repo → Actions → the failing run), not the jobs API. Suspicion worth checking
first: both use `astral-sh/setup-uv`, and eternitas' equivalent job failed with
`error: Failed to spawn: pytest` even after the action resolved — so the uv
toolchain may not be landing on PATH in these containers. That would be a
different, shared root cause.
**A trap worth keeping:** the three repos did NOT share one pattern. A naive
`localhost` → `postgres` swap would have left **WindyCloud on port 15432** (it
maps `15432:5432`) and windy-registry on a `job.services.postgres.ports[…]`
expression. Service-name networking always uses the container's **internal**
port — 5432 — never the mapped host port.
## The original task description (superseded above)
**Three repos need a one-line CI fix.** Their workflows reach a Postgres service
at `@localhost:5432`, which works on GitHub-hosted runners (services are
port-mapped to the VM) and fails on ours (the job runs *inside* a container, so
`localhost` is the job itself). The service is reachable as **`postgres`**.
| repo | workflow |
|---|---|
| `windy-mind` | `migrations.yml` |
| `windy-registry` | `ci.yml` |
| `WindyCloud` | `ci.yml` |
`eternitas` was already fixed this way — see **eternitas PR #149** for the exact
shape, including the comment explaining why. Fix must go to **GitHub**, not Windy
Git: the sync runs GitHub → Windy Git and force-pushes over local edits.
Proven by direct comparison, same runner and same `postgres:16-alpine` image:
windy-git's own gate uses `@postgres:5432` and passes its migration round-trip;
eternitas' used `@localhost:5432` and failed.
## Traps that will waste your time
- **`git pull -q` hides errors.** A divergent branch once made a "deploy" run 20
minutes against stale code while reporting success. Use `git fetch && git
merge --ff-only` and read the output.
- **Never force-push a branch a deploy checkout tracks** (`--amend` orphaned
`/srv/windygit/src` once).
- **Gitea's env-to-ini SETS but never UNSETS**, and sometimes *appends a
duplicate*. `GITEA__DEFAULT__APP_NAME` does not work at all — Gitea reads
`APP_NAME` from the **top level** of `app.ini`; the env var creates a literal
`[default]` section it ignores. Edit `app.ini` on the host.
- **Cloudflare caches `/assets/*` for 6h and no token in this stack can purge.**
Version brand asset **filenames** (`theme-windy.v2.css`), not query strings.
- **`base64` wraps at 76 chars** and corrupts long tokens in test commands →
`curl (43)`, phantom HTTP 000. Use `base64 -w0`.
- **Kit 0 is fragile.** 54 containers on 4 vCPU. Two production incidents in two
days, both from *non-production* workloads. Check `uptime` before deploying
anything there, and build before recreating so the swap is seconds.
## Open items, roughly by value
1. The three-repo `localhost` fix above.
2. **Get non-prod work off Kit 0.** 12 dev/demo containers on the box running
identity, the CA, mail, Matrix and the broker. Cost two incidents already;
the postgres-adapter fix would not have prevented either.
3. **Login is ~4–6s** — `postgres-adapter.ts:114` forks a `node -e` process per
query. Measured: node startup alone is 1.7s on Kit 0 vs 0.01s on Veron. The
fix is **one function** (persistent worker + `pg.Pool`), not the "468 call
sites" the SOTU scoped. See `docs/incidents/2026-08-12-login-latency-analysis.md`.
4. **Privileged dind sits beside broad-scoped tokens** on the CI host — Grant's
call, needs a decision not a code change.
5. Push-velocity throttling is declared but unenforceable from our plane (git
push never touches the API); needs a Gitea pre-receive hook.
## Read these first
- `~/.claude/.../memory/project_windy_git.md` — the full record, densest source
- `DNA_STRAND_MASTER_PLAN.md` — D-1…D-9 locked decisions, I-1…I-13 invariants
- `docs/AUDIT-fable-2026-08-13.md` — second-auditor findings and dispositions
- `docs/CUTOVER.md` — the GitHub↔Windy Git migration plan and its one rule
---
## Copy-paste prompt
```
Picking up Windy Git (agent-native code+model host on Veron 1, live at
app.windygit.com). Read these before doing anything:
1. ~/.claude/projects/-home-grantwhitmer/memory/project_windy_git.md
2. ~/windy-git/docs/TURNOVER-2026-08-14.md
3. ~/windy-git/DNA_STRAND_MASTER_PLAN.md (D-1..D-9, I-1..I-13)
State: live and in use. Grant signs in with his existing Windy account (SSO
fixed across windy-pro #346/#347). Agents authenticate with real EPT signature
verification. 143 repos, 85 tests green, health ok.
TASK: finish the CI fix. Four repos had workflows reaching Postgres through a
host port; all four are patched and merged (eternitas #149, windy-mind #100,
WindyCloud #89, windy-registry #31). windy-registry's `postgres integration`
went failure -> success, proving the approach. But windy-mind and WindyCloud
`migrations` still FAIL and I could not determine why.
Start by reading those job logs from the GITEA WEB UI (app.windygit.com -> repo
-> Actions -> failing run). Do NOT use the jobs API — it returns "job not found"
for the ids the runs report, which is what blocked the last session.
First hypothesis to test: windy-mind, WindyCloud and eternitas all use
`astral-sh/setup-uv`, and eternitas' job failed with `error: Failed to spawn:
pytest` even after the action resolved correctly. The uv toolchain may not be
landing on PATH inside these job containers — one shared root cause rather than
three.
Ground rules already paid for the hard way:
- verify the WHOLE flow, not the half that curls easily
- never `git pull -q` in a deploy path; it hides errors
- fixes go to GitHub, not Windy Git (sync is GitHub -> Windy Git, force-push)
- service containers: use the service NAME and its INTERNAL port (5432),
never the mapped host port
- check Kit 0's `uptime` before deploying there; two incidents in two days
from non-production workloads
```