From 2b30b0ac9912ec52812dc517ca44b32acb3bb786 Mon Sep 17 00:00:00 2001 From: Grant Whitmer Date: Fri, 14 Aug 2026 19:09:26 -0400 Subject: [PATCH] docs: record the runner bump, the act cache flake, and the remaining reds Second root cause found and fixed: act_runner 0.2.11 predates `runs.using: node24`, so any repo on actions/checkout@v5 died before its first step. Bumped to 0.6.1; Windy-Clone went 4/4 red to 4/4 green. Also upgrades the act-cache note from "watch item" to a confirmed job failure (lstat on a vanished file mid-tar), with the wipe command and the annotated-tag dead end that looks like a wrong checkout but isn't. Co-Authored-By: Claude Opus 5 --- docs/TURNOVER-2026-08-14.md | 111 ++++++++++++++++++++++++++---------- 1 file changed, 80 insertions(+), 31 deletions(-) diff --git a/docs/TURNOVER-2026-08-14.md b/docs/TURNOVER-2026-08-14.md index 9a2e68f..c352b8a 100644 --- a/docs/TURNOVER-2026-08-14.md +++ b/docs/TURNOVER-2026-08-14.md @@ -67,17 +67,36 @@ you do **not** need the web UI. Note logs are **not on disk** — `[storage] STORAGE_TYPE = minio` sends action logs to R2, so `actions_log/` on the host is empty. Read them through the API. +## SOLVED — second root cause: the runner was four majors behind + +Windy-Clone failed for a completely different reason: `The runs.using key in +action.yml must be one of: [composite docker node12 node16 node20 go], got +**node24**`. act_runner **0.2.11**'s bundled act predates node24, so any repo +pinning a current action major (`actions/checkout@v5`, `actions/setup-python@v6`) +died before its first step. + +**Bumped to `gitea/act_runner:0.6.1`** (`cd7dd9b`). Verified beforehand that +`node24` is absent from the 0.2.11 binary and present in 0.6.1, and that every +key in `deploy/runner/config.yaml` still exists in 0.6.1's schema — 0.6.1 only +*adds* keys, so the config carried over untouched. The runner re-declared with +the same id and labels `[veron-1 linux-x64 self-hosted linux x64]`; the +registration in the `runner-data` volume survived. Windy-Clone went 4/4 red → +4/4 green. Rollback is re-pinning 0.2.11; registration is backed up at +`/srv/windygit/runner-registration.bak`. + ## What is still red, and why each one is real -The CI plane is healthy. These are genuine repo defects that were **invisible -before**, because every job died at step 2: +The CI plane is healthy. windy-mind is fully green (`742 passed, 1 skipped`). +What remains are genuine repo defects that were **invisible before**, because +every job died at step 2: | repo | job | cause | |---|---|---| -| windy-mind | `tests` | `ruff check` — 6 real errors, 5 auto-fixable (`ruff check --fix`) | | WindyCloud | `lint` | `ruff format --check` — 7 files would be reformatted | | eternitas | `py-sdk` | `uv run pytest` → `Failed to spawn: pytest`; pytest isn't a declared dep of that project | -| WindyCloud | `docker` | **architectural** — see below | +| eternitas | `test` | needs its own look | +| windy-agent | `test (3.12/3.13/3.14)` | all three died **together** at 22:50:19 after ~43 min, mid-suite at 64%, with no verdict in the log. Not the 30m `runner.timeout` (that would have fired at 22:36) and not the runner bump (that was 23:00). Something bulk-killed them; unexplained. | +| WindyCloud | `docker`, windy-search `Docker build` | **architectural** — see below | `WindyCloud`'s `docker` job wants to build an image and gets `failed to connect to the docker API at unix:///var/run/docker.sock`. Job containers deliberately @@ -86,8 +105,8 @@ so many words not to mount it). Mounting the host socket would hand every workflow root on Veron 1. This needs a decision — buildx-in-dind, a rootless builder, or "this job does not run on Windy Git" — not a quiet socket mount. -The first three are one-line code fixes in their own repos and were left alone -on purpose: they are product defects, not forge defects. +The WindyCloud `lint` and eternitas `py-sdk` rows are small fixes in their own +repos, left alone on purpose: they are product defects, not forge defects. ## Second, smaller finding — act's action cache rots @@ -103,11 +122,29 @@ Reproduced directly: which surfaces as `Non-terminating error while running 'git clone': some refs were not updated`, after which the action does not report `Checked out `. -The cache was wiped this session (`rm -rf /root/.cache/act`, safe — it is in the -container layer, not a volume) and the actions resolved cleanly afterwards. -**This was never proven to fail a job on its own** — the setup-uv 404 masked it. -It is a watch item, not a closed issue. If actions start failing to resolve, wipe -that directory first. It will rot again. +It **has** now failed a job on its own. After the runner bump, `windy-mind tests` +died with: + +``` +❌ Failure - Main Install uv +lstat /root/.cache/act/d3e6…/.git-blame-ignore-revs: no such file or directory +``` + +act tars the cached action directory into the job container, and a file vanished +mid-walk. The cache dir had been created at 23:01 and mutated again at 23:03, +with `.gitignore` showing as deleted — act removes it before `docker cp`. The +likely mechanism is **concurrent jobs sharing one cache dir**: `capacity: 4`, and +windy-mind fires four jobs at once that all use setup-uv. Wiping the cache and +re-running the job alone made it pass (`742 passed, 1 skipped`). *Mechanism not +isolated* — the wipe alone may have been sufficient. + +One dead end worth not repeating: the cached worktree sits at `38f3f104` while +`git rev-parse v4` says `e4db8464`. That is **not** a wrong checkout — `v4` is an +*annotated tag*, and `v4^{commit}` is `38f3f104`. Don't chase it. + +Wipe with `sudo docker exec windy-git-runner-runner-1 rm -rf /root/.cache/act` +(safe — container layer, not a volume). It will rot again. A real fix is either +lowering `capacity` or upstream act; neither was attempted. ## Traps that will waste your time @@ -143,19 +180,23 @@ that directory first. It will rot again. ## Open items, roughly by value -1. **Decide what `WindyCloud`'s `docker` job should do on Windy Git** (above). - This is the only remaining *forge* question; it needs a decision, not code. -2. The three product-level test failures in the table above. -3. **Get non-prod work off Kit 0.** 12 dev/demo containers on the box running +1. **Decide what the image-building jobs should do on Windy Git** — WindyCloud + `docker` and windy-search `Docker build` (above). The only remaining *forge* + question; it needs a decision, not code. +2. **windy-agent's three `test` jobs were bulk-killed at 22:50:19** after ~43 + minutes, mid-suite, with no verdict. Unexplained and not the runner bump. +3. The product-level test failures in the table above. +4. **act's action cache race** — lower `capacity` below 4, or accept re-runs. +5. **Get non-prod work off Kit 0.** 12 dev/demo containers on the box running identity, the CA, mail, Matrix and the broker. Cost two incidents already; the postgres-adapter fix would not have prevented either. -4. **Login is ~4–6s** — `postgres-adapter.ts:114` forks a `node -e` process per +6. **Login is ~4–6s** — `postgres-adapter.ts:114` forks a `node -e` process per query. Measured: node startup alone is 1.7s on Kit 0 vs 0.01s on Veron. The fix is **one function** (persistent worker + `pg.Pool`), not the "468 call sites" the SOTU scoped. See `docs/incidents/2026-08-12-login-latency-analysis.md`. -5. **Privileged dind sits beside broad-scoped tokens** on the CI host — Grant's +7. **Privileged dind sits beside broad-scoped tokens** on the CI host — Grant's call, needs a decision not a code change. -6. Push-velocity throttling is declared but unenforceable from our plane (git +8. Push-velocity throttling is declared but unenforceable from our plane (git push never touches the API); needs a Gitea pre-receive hook. ## Read these first @@ -180,21 +221,26 @@ app.windygit.com). Read these before doing anything: State: live and in use. Grant signs in with his existing Windy account. Agents authenticate with real EPT signature verification. 143 repos, 85 tests green. -The CI breakage is SOLVED and verified: setup-uv v4+ resolved uv's "latest" -through GITHUB_API_URL, which act_runner points at our own forge, so it 404'd -("The target couldn't be found.") and every job died at step 2. Fixed by pinning -an explicit uv version across 11 repos; all merged and synced. The migrations -jobs in windy-mind, WindyCloud and eternitas are green. +TWO CI root causes are SOLVED and verified: + (a) setup-uv v4+ resolved uv's "latest" through GITHUB_API_URL, which + act_runner points at our own forge, so it 404'd ("The target couldn't be + found.") and every job died at step 2. Fixed by pinning an explicit uv + version across 11 repos. + (b) act_runner 0.2.11 predates `runs.using: node24`, so any repo on + actions/checkout@v5 died before its first step. Bumped to 0.6.1. +windy-mind is fully green (742 passed). Windy-Clone went 4/4 red to 4/4 green. TASK: one decision, then cleanup. - 1. DECIDE what WindyCloud's `docker` CI job should do here. It needs a Docker - daemon; job containers deliberately have no socket (I-5 — mounting the host - socket hands every workflow root on Veron 1). Options: buildx inside the - existing dind, a rootless builder, or exclude the job. Do NOT mount the - host socket. - 2. Three genuine product defects, newly visible now that jobs get past step 2: - windy-mind `tests` (ruff check, 6 errors), WindyCloud `lint` (ruff format, - 7 files), eternitas `py-sdk` (pytest not a declared dep). + 1. DECIDE what the image-building CI jobs should do here — WindyCloud `docker` + and windy-search `Docker build`. They need a Docker daemon; job containers + deliberately have no socket (I-5 — mounting the host socket hands every + workflow root on Veron 1). Options: buildx inside the existing dind, a + rootless builder, or exclude the job. Do NOT mount the host socket. + 2. windy-agent's three `test` jobs were bulk-killed together at 22:50:19 after + ~43 min, mid-suite, with no verdict in the log. Not the 30m runner.timeout, + not the runner bump. Unexplained — worth a look. + 3. Small product defects: WindyCloud `lint` (ruff format, 7 files), eternitas + `py-sdk` (pytest not a declared dep), eternitas `test`. Ground rules already paid for the hard way: - Gitea job status: 1=SUCCESS, 2=FAILURE, 5=waiting, 6=running. Not what you'd guess. @@ -209,5 +255,8 @@ Ground rules already paid for the hard way: - verify the WHOLE flow, not the half that curls easily - never `git pull -q` in a deploy path; it hides errors - fixes go to GitHub, not Windy Git (sync is GitHub -> Windy Git, force-push) + - if a job fails with `lstat .../: no such file or directory` on an + action, act's cache rotted: `docker exec windy-git-runner-runner-1 rm -rf + /root/.cache/act`, then re-run. Safe; it is a container layer, not a volume. - check Kit 0's `uptime` before deploying there; two incidents in two days ```