docs: record the runner bump, the act cache flake, and the remaining reds

Second root cause found and fixed: act_runner 0.2.11 predates
`runs.using: node24`, so any repo on actions/checkout@v5 died before its first
step. Bumped to 0.6.1; Windy-Clone went 4/4 red to 4/4 green.

Also upgrades the act-cache note from "watch item" to a confirmed job failure
(lstat on a vanished file mid-tar), with the wipe command and the annotated-tag
dead end that looks like a wrong checkout but isn't.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-14 19:09:26 -04:00
parent cd7dd9b7ae
commit 2b30b0ac99

View File

@@ -67,17 +67,36 @@ you do **not** need the web UI.
Note logs are **not on disk** — `[storage] STORAGE_TYPE = minio` sends action logs Note logs are **not on disk** — `[storage] STORAGE_TYPE = minio` sends action logs
to R2, so `actions_log/` on the host is empty. Read them through the API. to R2, so `actions_log/` on the host is empty. Read them through the API.
## SOLVED — second root cause: the runner was four majors behind
Windy-Clone failed for a completely different reason: `The runs.using key in
action.yml must be one of: [composite docker node12 node16 node20 go], got
**node24**`. act_runner **0.2.11**'s bundled act predates node24, so any repo
pinning a current action major (`actions/checkout@v5`, `actions/setup-python@v6`)
died before its first step.
**Bumped to `gitea/act_runner:0.6.1`** (`cd7dd9b`). Verified beforehand that
`node24` is absent from the 0.2.11 binary and present in 0.6.1, and that every
key in `deploy/runner/config.yaml` still exists in 0.6.1's schema — 0.6.1 only
*adds* keys, so the config carried over untouched. The runner re-declared with
the same id and labels `[veron-1 linux-x64 self-hosted linux x64]`; the
registration in the `runner-data` volume survived. Windy-Clone went 4/4 red →
4/4 green. Rollback is re-pinning 0.2.11; registration is backed up at
`/srv/windygit/runner-registration.bak`.
## What is still red, and why each one is real ## What is still red, and why each one is real
The CI plane is healthy. These are genuine repo defects that were **invisible The CI plane is healthy. windy-mind is fully green (`742 passed, 1 skipped`).
before**, because every job died at step 2: What remains are genuine repo defects that were **invisible before**, because
every job died at step 2:
| repo | job | cause | | repo | job | cause |
|---|---|---| |---|---|---|
| windy-mind | `tests` | `ruff check` — 6 real errors, 5 auto-fixable (`ruff check --fix`) |
| WindyCloud | `lint` | `ruff format --check` — 7 files would be reformatted | | WindyCloud | `lint` | `ruff format --check` — 7 files would be reformatted |
| eternitas | `py-sdk` | `uv run pytest` → `Failed to spawn: pytest`; pytest isn't a declared dep of that project | | eternitas | `py-sdk` | `uv run pytest` → `Failed to spawn: pytest`; pytest isn't a declared dep of that project |
| WindyCloud | `docker` | **architectural** — see below | | eternitas | `test` | needs its own look |
| windy-agent | `test (3.12/3.13/3.14)` | all three died **together** at 22:50:19 after ~43 min, mid-suite at 64%, with no verdict in the log. Not the 30m `runner.timeout` (that would have fired at 22:36) and not the runner bump (that was 23:00). Something bulk-killed them; unexplained. |
| WindyCloud | `docker`, windy-search `Docker build` | **architectural** — see below |
`WindyCloud`'s `docker` job wants to build an image and gets `failed to connect `WindyCloud`'s `docker` job wants to build an image and gets `failed to connect
to the docker API at unix:///var/run/docker.sock`. Job containers deliberately to the docker API at unix:///var/run/docker.sock`. Job containers deliberately
@@ -86,8 +105,8 @@ so many words not to mount it). Mounting the host socket would hand every
workflow root on Veron 1. This needs a decision — buildx-in-dind, a rootless workflow root on Veron 1. This needs a decision — buildx-in-dind, a rootless
builder, or "this job does not run on Windy Git" — not a quiet socket mount. builder, or "this job does not run on Windy Git" — not a quiet socket mount.
The first three are one-line code fixes in their own repos and were left alone The WindyCloud `lint` and eternitas `py-sdk` rows are small fixes in their own
on purpose: they are product defects, not forge defects. repos, left alone on purpose: they are product defects, not forge defects.
## Second, smaller finding — act's action cache rots ## Second, smaller finding — act's action cache rots
@@ -103,11 +122,29 @@ Reproduced directly:
which surfaces as `Non-terminating error while running 'git clone': some refs which surfaces as `Non-terminating error while running 'git clone': some refs
were not updated`, after which the action does not report `Checked out <ref>`. were not updated`, after which the action does not report `Checked out <ref>`.
The cache was wiped this session (`rm -rf /root/.cache/act`, safe — it is in the It **has** now failed a job on its own. After the runner bump, `windy-mind tests`
container layer, not a volume) and the actions resolved cleanly afterwards. died with:
**This was never proven to fail a job on its own** — the setup-uv 404 masked it.
It is a watch item, not a closed issue. If actions start failing to resolve, wipe ```
that directory first. It will rot again. ❌ Failure - Main Install uv
lstat /root/.cache/act/d3e6…/.git-blame-ignore-revs: no such file or directory
```
act tars the cached action directory into the job container, and a file vanished
mid-walk. The cache dir had been created at 23:01 and mutated again at 23:03,
with `.gitignore` showing as deleted — act removes it before `docker cp`. The
likely mechanism is **concurrent jobs sharing one cache dir**: `capacity: 4`, and
windy-mind fires four jobs at once that all use setup-uv. Wiping the cache and
re-running the job alone made it pass (`742 passed, 1 skipped`). *Mechanism not
isolated* — the wipe alone may have been sufficient.
One dead end worth not repeating: the cached worktree sits at `38f3f104` while
`git rev-parse v4` says `e4db8464`. That is **not** a wrong checkout — `v4` is an
*annotated tag*, and `v4^{commit}` is `38f3f104`. Don't chase it.
Wipe with `sudo docker exec windy-git-runner-runner-1 rm -rf /root/.cache/act`
(safe — container layer, not a volume). It will rot again. A real fix is either
lowering `capacity` or upstream act; neither was attempted.
## Traps that will waste your time ## Traps that will waste your time
@@ -143,19 +180,23 @@ that directory first. It will rot again.
## Open items, roughly by value ## Open items, roughly by value
1. **Decide what `WindyCloud`'s `docker` job should do on Windy Git** (above). 1. **Decide what the image-building jobs should do on Windy Git** — WindyCloud
This is the only remaining *forge* question; it needs a decision, not code. `docker` and windy-search `Docker build` (above). The only remaining *forge*
2. The three product-level test failures in the table above. question; it needs a decision, not code.
3. **Get non-prod work off Kit 0.** 12 dev/demo containers on the box running 2. **windy-agent's three `test` jobs were bulk-killed at 22:50:19** after ~43
minutes, mid-suite, with no verdict. Unexplained and not the runner bump.
3. The product-level test failures in the table above.
4. **act's action cache race** — lower `capacity` below 4, or accept re-runs.
5. **Get non-prod work off Kit 0.** 12 dev/demo containers on the box running
identity, the CA, mail, Matrix and the broker. Cost two incidents already; identity, the CA, mail, Matrix and the broker. Cost two incidents already;
the postgres-adapter fix would not have prevented either. the postgres-adapter fix would not have prevented either.
4. **Login is ~4–6s** — `postgres-adapter.ts:114` forks a `node -e` process per 6. **Login is ~4–6s** — `postgres-adapter.ts:114` forks a `node -e` process per
query. Measured: node startup alone is 1.7s on Kit 0 vs 0.01s on Veron. The query. Measured: node startup alone is 1.7s on Kit 0 vs 0.01s on Veron. The
fix is **one function** (persistent worker + `pg.Pool`), not the "468 call fix is **one function** (persistent worker + `pg.Pool`), not the "468 call
sites" the SOTU scoped. See `docs/incidents/2026-08-12-login-latency-analysis.md`. sites" the SOTU scoped. See `docs/incidents/2026-08-12-login-latency-analysis.md`.
5. **Privileged dind sits beside broad-scoped tokens** on the CI host — Grant's 7. **Privileged dind sits beside broad-scoped tokens** on the CI host — Grant's
call, needs a decision not a code change. call, needs a decision not a code change.
6. Push-velocity throttling is declared but unenforceable from our plane (git 8. Push-velocity throttling is declared but unenforceable from our plane (git
push never touches the API); needs a Gitea pre-receive hook. push never touches the API); needs a Gitea pre-receive hook.
## Read these first ## Read these first
@@ -180,21 +221,26 @@ app.windygit.com). Read these before doing anything:
State: live and in use. Grant signs in with his existing Windy account. Agents State: live and in use. Grant signs in with his existing Windy account. Agents
authenticate with real EPT signature verification. 143 repos, 85 tests green. authenticate with real EPT signature verification. 143 repos, 85 tests green.
The CI breakage is SOLVED and verified: setup-uv v4+ resolved uv's "latest" TWO CI root causes are SOLVED and verified:
through GITHUB_API_URL, which act_runner points at our own forge, so it 404'd (a) setup-uv v4+ resolved uv's "latest" through GITHUB_API_URL, which
("The target couldn't be found.") and every job died at step 2. Fixed by pinning act_runner points at our own forge, so it 404'd ("The target couldn't be
an explicit uv version across 11 repos; all merged and synced. The migrations found.") and every job died at step 2. Fixed by pinning an explicit uv
jobs in windy-mind, WindyCloud and eternitas are green. version across 11 repos.
(b) act_runner 0.2.11 predates `runs.using: node24`, so any repo on
actions/checkout@v5 died before its first step. Bumped to 0.6.1.
windy-mind is fully green (742 passed). Windy-Clone went 4/4 red to 4/4 green.
TASK: one decision, then cleanup. TASK: one decision, then cleanup.
1. DECIDE what WindyCloud's `docker` CI job should do here. It needs a Docker 1. DECIDE what the image-building CI jobs should do here — WindyCloud `docker`
daemon; job containers deliberately have no socket (I-5 — mounting the host and windy-search `Docker build`. They need a Docker daemon; job containers
socket hands every workflow root on Veron 1). Options: buildx inside the deliberately have no socket (I-5 — mounting the host socket hands every
existing dind, a rootless builder, or exclude the job. Do NOT mount the workflow root on Veron 1). Options: buildx inside the existing dind, a
host socket. rootless builder, or exclude the job. Do NOT mount the host socket.
2. Three genuine product defects, newly visible now that jobs get past step 2: 2. windy-agent's three `test` jobs were bulk-killed together at 22:50:19 after
windy-mind `tests` (ruff check, 6 errors), WindyCloud `lint` (ruff format, ~43 min, mid-suite, with no verdict in the log. Not the 30m runner.timeout,
7 files), eternitas `py-sdk` (pytest not a declared dep). not the runner bump. Unexplained — worth a look.
3. Small product defects: WindyCloud `lint` (ruff format, 7 files), eternitas
`py-sdk` (pytest not a declared dep), eternitas `test`.
Ground rules already paid for the hard way: Ground rules already paid for the hard way:
- Gitea job status: 1=SUCCESS, 2=FAILURE, 5=waiting, 6=running. Not what you'd guess. - Gitea job status: 1=SUCCESS, 2=FAILURE, 5=waiting, 6=running. Not what you'd guess.
@@ -209,5 +255,8 @@ Ground rules already paid for the hard way:
- verify the WHOLE flow, not the half that curls easily - verify the WHOLE flow, not the half that curls easily
- never `git pull -q` in a deploy path; it hides errors - never `git pull -q` in a deploy path; it hides errors
- fixes go to GitHub, not Windy Git (sync is GitHub -> Windy Git, force-push) - fixes go to GitHub, not Windy Git (sync is GitHub -> Windy Git, force-push)
- if a job fails with `lstat .../<file>: no such file or directory` on an
action, act's cache rotted: `docker exec windy-git-runner-runner-1 rm -rf
/root/.cache/act`, then re-run. Safe; it is a container layer, not a volume.
- check Kit 0's `uptime` before deploying there; two incidents in two days - check Kit 0's `uptime` before deploying there; two incidents in two days
``` ```