diff --git a/docs/TURNOVER-2026-08-14.md b/docs/TURNOVER-2026-08-14.md index 3bd495a..9a2e68f 100644 --- a/docs/TURNOVER-2026-08-14.md +++ b/docs/TURNOVER-2026-08-14.md @@ -13,56 +13,113 @@ Grant signs in with his existing Windy Word credentials — no second account. Agents authenticate with real Eternitas EPT signature verification and are rate-limited by integrity band. -## DONE since this was written — the three-repo CI fix +## SOLVED — the CI failures were never about Postgres -All three PRs are **merged and synced**: windy-mind #100, WindyCloud #89, -windy-registry #31 (eternitas #149 earlier). +The `localhost` → `postgres` fix was correct and is worth keeping, but it was +**not** what was failing these jobs. They died at step 2, before Postgres was +ever contacted. -**Result: 1 of 3 verified fixed, 2 still failing for an undetermined reason.** +**Root cause: `astral-sh/setup-uv@v4` asks the forge for uv's latest release.** -- ✅ **windy-registry** — `postgres integration` went **failure → success**. The - fix is proven correct. -- ❌ **windy-mind**, **WindyCloud** — `migrations` still fails. The DATABASE_URL - is definitely right now; the cause is something else and was **not - determined** — the jobs API returns "job not found" for the ids the runs - report, so logs could not be retrieved that way. +setup-uv v4 added "resolve latest version instead of downloading latest release" +(astral-sh/setup-uv#178). Resolution goes through `@actions/github`, whose +octokit reads **`GITHUB_API_URL`** — which act_runner points at *our forge*. So +the action requested: -**Next session: read those job logs from the Gitea web UI** (`app.windygit.com` -→ repo → Actions → the failing run), not the jobs API. Suspicion worth checking -first: both use `astral-sh/setup-uv`, and eternitas' equivalent job failed with -`error: Failed to spawn: pytest` even after the action resolved — so the uv -toolchain may not be landing on PATH in these containers. That would be a -different, shared root cause. +``` +GET https://app.windygit.com/api/v1/repos/astral-sh/uv/releases/latest → 404 +``` -**A trap worth keeping:** the three repos did NOT share one pattern. A naive -`localhost` → `postgres` swap would have left **WindyCloud on port 15432** (it -maps `15432:5432`) and windy-registry on a `job.services.postgres.ports[…]` -expression. Service-name networking always uses the container's **internal** -port — 5432 — never the mapped host port. +Gitea has no `astral-sh/uv`, so it answered its standard 404 body, *"The target +couldn't be found."* setup-uv threw that string, act printed it as `::error::`, +and every later step was skipped by `success()`. -## The original task description (superseded above) +**The fix (merged to GitHub, 11 repos):** pin an explicit `version:` on every +`setup-uv@v4`/`@v5` step. `resolveVersion()` short-circuits on an explicit +version *before* any API call, and the download URL is hardcoded to github.com — +so the forge round-trip disappears. Pinned to `0.12.5`, which is what `latest` +already resolved to. -**Three repos need a one-line CI fix.** Their workflows reach a Postgres service -at `@localhost:5432`, which works on GitHub-hosted runners (services are -port-mapped to the VM) and fails on ours (the job runs *inside* a container, so -`localhost` is the job itself). The service is reachable as **`postgres`**. +windy-mind #101, WindyCloud #90, eternitas #150, then the sweep: Windy-Clone #77, +windy-agent #355, windy-call #34, windy-cell #31, windy-hand #6, windy-mail #105, +windy-search #77, windy-text #29. All merged and synced. -| repo | workflow | -|---|---| -| `windy-mind` | `migrations.yml` | -| `windy-registry` | `ci.yml` | -| `WindyCloud` | `ci.yml` | +### Two things that made this hard to see, both worth keeping -`eternitas` was already fixed this way — see **eternitas PR #149** for the exact -shape, including the comment explaining why. Fix must go to **GitHub**, not Windy -Git: the sync runs GitHub → Windy Git and force-pushes over local edits. +- **act attributes the error to the wrong step.** `::error::The target couldn't + be found.` is printed immediately after `actions/checkout`'s `::remove-matcher`, + so it reads exactly like a checkout failure. It is not. What settled it was the + **Gitea access log** — `sudo docker logs windy-git-gitea-1 | grep " 404 "` — which + named the real URL at the same millisecond as the job error. When a job fails + with an opaque forge-shaped message, go to the forge's access log, not the job log. -Proven by direct comparison, same runner and same `postgres:16-alpine` image: -windy-git's own gate uses `@postgres:5432` and passes its migration round-trip; -eternitas' used `@localhost:5432` and failed. +- **The natural experiment was sitting right there.** windy-registry and + windy-drops use `setup-uv@v3` and always passed; every v4/v5 caller failed. A + version skew across otherwise-identical repos is a diagnosis, not a coincidence. + +### The jobs API "job not found" that blocked the last session + +Not a bug. `GET /api/v1/repos/{owner}/{repo}/actions/jobs/{id}/logs` requires the +job id to belong to **the repo in the path** — a valid id under the wrong owner/repo +404s. The API works fine; the URLs were mismatched. Logs are readable this way and +you do **not** need the web UI. + +Note logs are **not on disk** — `[storage] STORAGE_TYPE = minio` sends action logs +to R2, so `actions_log/` on the host is empty. Read them through the API. + +## What is still red, and why each one is real + +The CI plane is healthy. These are genuine repo defects that were **invisible +before**, because every job died at step 2: + +| repo | job | cause | +|---|---|---| +| windy-mind | `tests` | `ruff check` — 6 real errors, 5 auto-fixable (`ruff check --fix`) | +| WindyCloud | `lint` | `ruff format --check` — 7 files would be reformatted | +| eternitas | `py-sdk` | `uv run pytest` → `Failed to spawn: pytest`; pytest isn't a declared dep of that project | +| WindyCloud | `docker` | **architectural** — see below | + +`WindyCloud`'s `docker` job wants to build an image and gets `failed to connect +to the docker API at unix:///var/run/docker.sock`. Job containers deliberately +have **no** docker socket (I-5, and `deploy/runner/docker-compose.yml` says in +so many words not to mount it). Mounting the host socket would hand every +workflow root on Veron 1. This needs a decision — buildx-in-dind, a rootless +builder, or "this job does not run on Windy Git" — not a quiet socket mount. + +The first three are one-line code fixes in their own repos and were left alone +on purpose: they are product defects, not forge defects. + +## Second, smaller finding — act's action cache rots + +act caches action repos at `/root/.cache/act/` inside the runner container +and refreshes them with a go-git mirror fetch of `refs/*:refs/*`, unforced. That +includes `refs/pull/*`, which GitHub **recomputes** whenever a base branch moves. +Reproduced directly: + +``` +! [rejected] refs/pull/1015/merge -> refs/pull/1015/merge (non-fast-forward) +``` + +which surfaces as `Non-terminating error while running 'git clone': some refs +were not updated`, after which the action does not report `Checked out `. + +The cache was wiped this session (`rm -rf /root/.cache/act`, safe — it is in the +container layer, not a volume) and the actions resolved cleanly afterwards. +**This was never proven to fail a job on its own** — the setup-uv 404 masked it. +It is a watch item, not a closed issue. If actions start failing to resolve, wipe +that directory first. It will rot again. ## Traps that will waste your time +- **Gitea status codes are not what they look like.** `1 = success, 2 = failure`, + 3 cancelled, 4 skipped, 5 waiting, 6 running, 7 blocked. Reading 1/2 as + waiting/running inverts every conclusion you draw from `action_run_job`. +- **Gitea sets `Secure` cookies** (ROOT_URL is https), so a `curl` login against + `http://127.0.0.1:3080` silently keeps no session — it 303s to `/` and you + still get "Sign In". Log in through `https://app.windygit.com`. +- **There is no rerun API in 1.24.6.** `POST /api/v1/.../runs/{n}/rerun` 404s. + Use the web route `POST /{owner}/{repo}/actions/runs/{n}/rerun` with the session + cookie plus an `X-Csrf-Token` header taken from the `_csrf` cookie. - **`git pull -q` hides errors.** A divergent branch once made a "deploy" run 20 minutes against stale code while reporting success. Use `git fetch && git merge --ff-only` and read the output. @@ -79,20 +136,26 @@ eternitas' used `@localhost:5432` and failed. - **Kit 0 is fragile.** 54 containers on 4 vCPU. Two production incidents in two days, both from *non-production* workloads. Check `uptime` before deploying anything there, and build before recreating so the swap is seconds. +- **Service containers**: use the service NAME and its INTERNAL port (5432), + never the mapped host port. The three repos did NOT share one pattern — a naive + `localhost` → `postgres` swap would have left WindyCloud on port 15432 (it maps + `15432:5432`) and windy-registry on a `job.services.postgres.ports[…]` expression. ## Open items, roughly by value -1. The three-repo `localhost` fix above. -2. **Get non-prod work off Kit 0.** 12 dev/demo containers on the box running +1. **Decide what `WindyCloud`'s `docker` job should do on Windy Git** (above). + This is the only remaining *forge* question; it needs a decision, not code. +2. The three product-level test failures in the table above. +3. **Get non-prod work off Kit 0.** 12 dev/demo containers on the box running identity, the CA, mail, Matrix and the broker. Cost two incidents already; the postgres-adapter fix would not have prevented either. -3. **Login is ~4–6s** — `postgres-adapter.ts:114` forks a `node -e` process per +4. **Login is ~4–6s** — `postgres-adapter.ts:114` forks a `node -e` process per query. Measured: node startup alone is 1.7s on Kit 0 vs 0.01s on Veron. The fix is **one function** (persistent worker + `pg.Pool`), not the "468 call sites" the SOTU scoped. See `docs/incidents/2026-08-12-login-latency-analysis.md`. -4. **Privileged dind sits beside broad-scoped tokens** on the CI host — Grant's +5. **Privileged dind sits beside broad-scoped tokens** on the CI host — Grant's call, needs a decision not a code change. -5. Push-velocity throttling is declared but unenforceable from our plane (git +6. Push-velocity throttling is declared but unenforceable from our plane (git push never touches the API); needs a Gitea pre-receive hook. ## Read these first @@ -114,32 +177,37 @@ app.windygit.com). Read these before doing anything: 2. ~/windy-git/docs/TURNOVER-2026-08-14.md 3. ~/windy-git/DNA_STRAND_MASTER_PLAN.md (D-1..D-9, I-1..I-13) -State: live and in use. Grant signs in with his existing Windy account (SSO -fixed across windy-pro #346/#347). Agents authenticate with real EPT signature -verification. 143 repos, 85 tests green, health ok. +State: live and in use. Grant signs in with his existing Windy account. Agents +authenticate with real EPT signature verification. 143 repos, 85 tests green. -TASK: finish the CI fix. Four repos had workflows reaching Postgres through a -host port; all four are patched and merged (eternitas #149, windy-mind #100, -WindyCloud #89, windy-registry #31). windy-registry's `postgres integration` -went failure -> success, proving the approach. But windy-mind and WindyCloud -`migrations` still FAIL and I could not determine why. +The CI breakage is SOLVED and verified: setup-uv v4+ resolved uv's "latest" +through GITHUB_API_URL, which act_runner points at our own forge, so it 404'd +("The target couldn't be found.") and every job died at step 2. Fixed by pinning +an explicit uv version across 11 repos; all merged and synced. The migrations +jobs in windy-mind, WindyCloud and eternitas are green. -Start by reading those job logs from the GITEA WEB UI (app.windygit.com -> repo --> Actions -> failing run). Do NOT use the jobs API — it returns "job not found" -for the ids the runs report, which is what blocked the last session. - -First hypothesis to test: windy-mind, WindyCloud and eternitas all use -`astral-sh/setup-uv`, and eternitas' job failed with `error: Failed to spawn: -pytest` even after the action resolved correctly. The uv toolchain may not be -landing on PATH inside these job containers — one shared root cause rather than -three. +TASK: one decision, then cleanup. + 1. DECIDE what WindyCloud's `docker` CI job should do here. It needs a Docker + daemon; job containers deliberately have no socket (I-5 — mounting the host + socket hands every workflow root on Veron 1). Options: buildx inside the + existing dind, a rootless builder, or exclude the job. Do NOT mount the + host socket. + 2. Three genuine product defects, newly visible now that jobs get past step 2: + windy-mind `tests` (ruff check, 6 errors), WindyCloud `lint` (ruff format, + 7 files), eternitas `py-sdk` (pytest not a declared dep). Ground rules already paid for the hard way: + - Gitea job status: 1=SUCCESS, 2=FAILURE, 5=waiting, 6=running. Not what you'd guess. + - When a job fails with an opaque forge-shaped error, read the FORGE access log + (`docker logs windy-git-gitea-1 | grep " 404 "`) — act misattributes the error + to the previous step. + - Job logs: `GET /api/v1/repos/{owner}/{repo}/actions/jobs/{id}/logs`. The job id + must belong to the repo in the path or you get a misleading "job not found". + Logs are in R2, not on disk. + - Log into the forge over https://app.windygit.com — Gitea's cookies are Secure, + so a curl login to http://127.0.0.1:3080 silently keeps no session. - verify the WHOLE flow, not the half that curls easily - never `git pull -q` in a deploy path; it hides errors - fixes go to GitHub, not Windy Git (sync is GitHub -> Windy Git, force-push) - - service containers: use the service NAME and its INTERNAL port (5432), - never the mapped host port - check Kit 0's `uptime` before deploying there; two incidents in two days - from non-production workloads ```