Second root cause found and fixed: act_runner 0.2.11 predates `runs.using: node24`, so any repo on actions/checkout@v5 died before its first step. Bumped to 0.6.1; Windy-Clone went 4/4 red to 4/4 green. Also upgrades the act-cache note from "watch item" to a confirmed job failure (lstat on a vanished file mid-tar), with the wipe command and the annotated-tag dead end that looks like a wrong checkout but isn't. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
263 lines
14 KiB
Markdown
263 lines
14 KiB
Markdown
# Windy Git — turnover, 2026-08-14
|
||
|
||
Paste the block at the bottom into a fresh terminal. Everything above is context
|
||
for whoever reads this file directly.
|
||
|
||
## Where things stand
|
||
|
||
Windy Git is **live and in use**: `app.windygit.com` (forge), `api.windygit.com`
|
||
(our plane), on **Veron 1** behind a Cloudflare Tunnel, zero inbound ports, $0/mo.
|
||
143 repos, 85 tests green, health `ok` on all four checks.
|
||
|
||
Grant signs in with his existing Windy Word credentials — no second account.
|
||
Agents authenticate with real Eternitas EPT signature verification and are
|
||
rate-limited by integrity band.
|
||
|
||
## SOLVED — the CI failures were never about Postgres
|
||
|
||
The `localhost` → `postgres` fix was correct and is worth keeping, but it was
|
||
**not** what was failing these jobs. They died at step 2, before Postgres was
|
||
ever contacted.
|
||
|
||
**Root cause: `astral-sh/setup-uv@v4` asks the forge for uv's latest release.**
|
||
|
||
setup-uv v4 added "resolve latest version instead of downloading latest release"
|
||
(astral-sh/setup-uv#178). Resolution goes through `@actions/github`, whose
|
||
octokit reads **`GITHUB_API_URL`** — which act_runner points at *our forge*. So
|
||
the action requested:
|
||
|
||
```
|
||
GET https://app.windygit.com/api/v1/repos/astral-sh/uv/releases/latest → 404
|
||
```
|
||
|
||
Gitea has no `astral-sh/uv`, so it answered its standard 404 body, *"The target
|
||
couldn't be found."* setup-uv threw that string, act printed it as `::error::`,
|
||
and every later step was skipped by `success()`.
|
||
|
||
**The fix (merged to GitHub, 11 repos):** pin an explicit `version:` on every
|
||
`setup-uv@v4`/`@v5` step. `resolveVersion()` short-circuits on an explicit
|
||
version *before* any API call, and the download URL is hardcoded to github.com —
|
||
so the forge round-trip disappears. Pinned to `0.12.5`, which is what `latest`
|
||
already resolved to.
|
||
|
||
windy-mind #101, WindyCloud #90, eternitas #150, then the sweep: Windy-Clone #77,
|
||
windy-agent #355, windy-call #34, windy-cell #31, windy-hand #6, windy-mail #105,
|
||
windy-search #77, windy-text #29. All merged and synced.
|
||
|
||
### Two things that made this hard to see, both worth keeping
|
||
|
||
- **act attributes the error to the wrong step.** `::error::The target couldn't
|
||
be found.` is printed immediately after `actions/checkout`'s `::remove-matcher`,
|
||
so it reads exactly like a checkout failure. It is not. What settled it was the
|
||
**Gitea access log** — `sudo docker logs windy-git-gitea-1 | grep " 404 "` — which
|
||
named the real URL at the same millisecond as the job error. When a job fails
|
||
with an opaque forge-shaped message, go to the forge's access log, not the job log.
|
||
|
||
- **The natural experiment was sitting right there.** windy-registry and
|
||
windy-drops use `setup-uv@v3` and always passed; every v4/v5 caller failed. A
|
||
version skew across otherwise-identical repos is a diagnosis, not a coincidence.
|
||
|
||
### The jobs API "job not found" that blocked the last session
|
||
|
||
Not a bug. `GET /api/v1/repos/{owner}/{repo}/actions/jobs/{id}/logs` requires the
|
||
job id to belong to **the repo in the path** — a valid id under the wrong owner/repo
|
||
404s. The API works fine; the URLs were mismatched. Logs are readable this way and
|
||
you do **not** need the web UI.
|
||
|
||
Note logs are **not on disk** — `[storage] STORAGE_TYPE = minio` sends action logs
|
||
to R2, so `actions_log/` on the host is empty. Read them through the API.
|
||
|
||
## SOLVED — second root cause: the runner was four majors behind
|
||
|
||
Windy-Clone failed for a completely different reason: `The runs.using key in
|
||
action.yml must be one of: [composite docker node12 node16 node20 go], got
|
||
**node24**`. act_runner **0.2.11**'s bundled act predates node24, so any repo
|
||
pinning a current action major (`actions/checkout@v5`, `actions/setup-python@v6`)
|
||
died before its first step.
|
||
|
||
**Bumped to `gitea/act_runner:0.6.1`** (`cd7dd9b`). Verified beforehand that
|
||
`node24` is absent from the 0.2.11 binary and present in 0.6.1, and that every
|
||
key in `deploy/runner/config.yaml` still exists in 0.6.1's schema — 0.6.1 only
|
||
*adds* keys, so the config carried over untouched. The runner re-declared with
|
||
the same id and labels `[veron-1 linux-x64 self-hosted linux x64]`; the
|
||
registration in the `runner-data` volume survived. Windy-Clone went 4/4 red →
|
||
4/4 green. Rollback is re-pinning 0.2.11; registration is backed up at
|
||
`/srv/windygit/runner-registration.bak`.
|
||
|
||
## What is still red, and why each one is real
|
||
|
||
The CI plane is healthy. windy-mind is fully green (`742 passed, 1 skipped`).
|
||
What remains are genuine repo defects that were **invisible before**, because
|
||
every job died at step 2:
|
||
|
||
| repo | job | cause |
|
||
|---|---|---|
|
||
| WindyCloud | `lint` | `ruff format --check` — 7 files would be reformatted |
|
||
| eternitas | `py-sdk` | `uv run pytest` → `Failed to spawn: pytest`; pytest isn't a declared dep of that project |
|
||
| eternitas | `test` | needs its own look |
|
||
| windy-agent | `test (3.12/3.13/3.14)` | all three died **together** at 22:50:19 after ~43 min, mid-suite at 64%, with no verdict in the log. Not the 30m `runner.timeout` (that would have fired at 22:36) and not the runner bump (that was 23:00). Something bulk-killed them; unexplained. |
|
||
| WindyCloud | `docker`, windy-search `Docker build` | **architectural** — see below |
|
||
|
||
`WindyCloud`'s `docker` job wants to build an image and gets `failed to connect
|
||
to the docker API at unix:///var/run/docker.sock`. Job containers deliberately
|
||
have **no** docker socket (I-5, and `deploy/runner/docker-compose.yml` says in
|
||
so many words not to mount it). Mounting the host socket would hand every
|
||
workflow root on Veron 1. This needs a decision — buildx-in-dind, a rootless
|
||
builder, or "this job does not run on Windy Git" — not a quiet socket mount.
|
||
|
||
The WindyCloud `lint` and eternitas `py-sdk` rows are small fixes in their own
|
||
repos, left alone on purpose: they are product defects, not forge defects.
|
||
|
||
## Second, smaller finding — act's action cache rots
|
||
|
||
act caches action repos at `/root/.cache/act/<hash>` inside the runner container
|
||
and refreshes them with a go-git mirror fetch of `refs/*:refs/*`, unforced. That
|
||
includes `refs/pull/*`, which GitHub **recomputes** whenever a base branch moves.
|
||
Reproduced directly:
|
||
|
||
```
|
||
! [rejected] refs/pull/1015/merge -> refs/pull/1015/merge (non-fast-forward)
|
||
```
|
||
|
||
which surfaces as `Non-terminating error while running 'git clone': some refs
|
||
were not updated`, after which the action does not report `Checked out <ref>`.
|
||
|
||
It **has** now failed a job on its own. After the runner bump, `windy-mind tests`
|
||
died with:
|
||
|
||
```
|
||
❌ Failure - Main Install uv
|
||
lstat /root/.cache/act/d3e6…/.git-blame-ignore-revs: no such file or directory
|
||
```
|
||
|
||
act tars the cached action directory into the job container, and a file vanished
|
||
mid-walk. The cache dir had been created at 23:01 and mutated again at 23:03,
|
||
with `.gitignore` showing as deleted — act removes it before `docker cp`. The
|
||
likely mechanism is **concurrent jobs sharing one cache dir**: `capacity: 4`, and
|
||
windy-mind fires four jobs at once that all use setup-uv. Wiping the cache and
|
||
re-running the job alone made it pass (`742 passed, 1 skipped`). *Mechanism not
|
||
isolated* — the wipe alone may have been sufficient.
|
||
|
||
One dead end worth not repeating: the cached worktree sits at `38f3f104` while
|
||
`git rev-parse v4` says `e4db8464`. That is **not** a wrong checkout — `v4` is an
|
||
*annotated tag*, and `v4^{commit}` is `38f3f104`. Don't chase it.
|
||
|
||
Wipe with `sudo docker exec windy-git-runner-runner-1 rm -rf /root/.cache/act`
|
||
(safe — container layer, not a volume). It will rot again. A real fix is either
|
||
lowering `capacity` or upstream act; neither was attempted.
|
||
|
||
## Traps that will waste your time
|
||
|
||
- **Gitea status codes are not what they look like.** `1 = success, 2 = failure`,
|
||
3 cancelled, 4 skipped, 5 waiting, 6 running, 7 blocked. Reading 1/2 as
|
||
waiting/running inverts every conclusion you draw from `action_run_job`.
|
||
- **Gitea sets `Secure` cookies** (ROOT_URL is https), so a `curl` login against
|
||
`http://127.0.0.1:3080` silently keeps no session — it 303s to `/` and you
|
||
still get "Sign In". Log in through `https://app.windygit.com`.
|
||
- **There is no rerun API in 1.24.6.** `POST /api/v1/.../runs/{n}/rerun` 404s.
|
||
Use the web route `POST /{owner}/{repo}/actions/runs/{n}/rerun` with the session
|
||
cookie plus an `X-Csrf-Token` header taken from the `_csrf` cookie.
|
||
- **`git pull -q` hides errors.** A divergent branch once made a "deploy" run 20
|
||
minutes against stale code while reporting success. Use `git fetch && git
|
||
merge --ff-only` and read the output.
|
||
- **Never force-push a branch a deploy checkout tracks** (`--amend` orphaned
|
||
`/srv/windygit/src` once).
|
||
- **Gitea's env-to-ini SETS but never UNSETS**, and sometimes *appends a
|
||
duplicate*. `GITEA__DEFAULT__APP_NAME` does not work at all — Gitea reads
|
||
`APP_NAME` from the **top level** of `app.ini`; the env var creates a literal
|
||
`[default]` section it ignores. Edit `app.ini` on the host.
|
||
- **Cloudflare caches `/assets/*` for 6h and no token in this stack can purge.**
|
||
Version brand asset **filenames** (`theme-windy.v2.css`), not query strings.
|
||
- **`base64` wraps at 76 chars** and corrupts long tokens in test commands →
|
||
`curl (43)`, phantom HTTP 000. Use `base64 -w0`.
|
||
- **Kit 0 is fragile.** 54 containers on 4 vCPU. Two production incidents in two
|
||
days, both from *non-production* workloads. Check `uptime` before deploying
|
||
anything there, and build before recreating so the swap is seconds.
|
||
- **Service containers**: use the service NAME and its INTERNAL port (5432),
|
||
never the mapped host port. The three repos did NOT share one pattern — a naive
|
||
`localhost` → `postgres` swap would have left WindyCloud on port 15432 (it maps
|
||
`15432:5432`) and windy-registry on a `job.services.postgres.ports[…]` expression.
|
||
|
||
## Open items, roughly by value
|
||
|
||
1. **Decide what the image-building jobs should do on Windy Git** — WindyCloud
|
||
`docker` and windy-search `Docker build` (above). The only remaining *forge*
|
||
question; it needs a decision, not code.
|
||
2. **windy-agent's three `test` jobs were bulk-killed at 22:50:19** after ~43
|
||
minutes, mid-suite, with no verdict. Unexplained and not the runner bump.
|
||
3. The product-level test failures in the table above.
|
||
4. **act's action cache race** — lower `capacity` below 4, or accept re-runs.
|
||
5. **Get non-prod work off Kit 0.** 12 dev/demo containers on the box running
|
||
identity, the CA, mail, Matrix and the broker. Cost two incidents already;
|
||
the postgres-adapter fix would not have prevented either.
|
||
6. **Login is ~4–6s** — `postgres-adapter.ts:114` forks a `node -e` process per
|
||
query. Measured: node startup alone is 1.7s on Kit 0 vs 0.01s on Veron. The
|
||
fix is **one function** (persistent worker + `pg.Pool`), not the "468 call
|
||
sites" the SOTU scoped. See `docs/incidents/2026-08-12-login-latency-analysis.md`.
|
||
7. **Privileged dind sits beside broad-scoped tokens** on the CI host — Grant's
|
||
call, needs a decision not a code change.
|
||
8. Push-velocity throttling is declared but unenforceable from our plane (git
|
||
push never touches the API); needs a Gitea pre-receive hook.
|
||
|
||
## Read these first
|
||
|
||
- `~/.claude/.../memory/project_windy_git.md` — the full record, densest source
|
||
- `DNA_STRAND_MASTER_PLAN.md` — D-1…D-9 locked decisions, I-1…I-13 invariants
|
||
- `docs/AUDIT-fable-2026-08-13.md` — second-auditor findings and dispositions
|
||
- `docs/CUTOVER.md` — the GitHub↔Windy Git migration plan and its one rule
|
||
|
||
---
|
||
|
||
## Copy-paste prompt
|
||
|
||
```
|
||
Picking up Windy Git (agent-native code+model host on Veron 1, live at
|
||
app.windygit.com). Read these before doing anything:
|
||
|
||
1. ~/.claude/projects/-home-grantwhitmer/memory/project_windy_git.md
|
||
2. ~/windy-git/docs/TURNOVER-2026-08-14.md
|
||
3. ~/windy-git/DNA_STRAND_MASTER_PLAN.md (D-1..D-9, I-1..I-13)
|
||
|
||
State: live and in use. Grant signs in with his existing Windy account. Agents
|
||
authenticate with real EPT signature verification. 143 repos, 85 tests green.
|
||
|
||
TWO CI root causes are SOLVED and verified:
|
||
(a) setup-uv v4+ resolved uv's "latest" through GITHUB_API_URL, which
|
||
act_runner points at our own forge, so it 404'd ("The target couldn't be
|
||
found.") and every job died at step 2. Fixed by pinning an explicit uv
|
||
version across 11 repos.
|
||
(b) act_runner 0.2.11 predates `runs.using: node24`, so any repo on
|
||
actions/checkout@v5 died before its first step. Bumped to 0.6.1.
|
||
windy-mind is fully green (742 passed). Windy-Clone went 4/4 red to 4/4 green.
|
||
|
||
TASK: one decision, then cleanup.
|
||
1. DECIDE what the image-building CI jobs should do here — WindyCloud `docker`
|
||
and windy-search `Docker build`. They need a Docker daemon; job containers
|
||
deliberately have no socket (I-5 — mounting the host socket hands every
|
||
workflow root on Veron 1). Options: buildx inside the existing dind, a
|
||
rootless builder, or exclude the job. Do NOT mount the host socket.
|
||
2. windy-agent's three `test` jobs were bulk-killed together at 22:50:19 after
|
||
~43 min, mid-suite, with no verdict in the log. Not the 30m runner.timeout,
|
||
not the runner bump. Unexplained — worth a look.
|
||
3. Small product defects: WindyCloud `lint` (ruff format, 7 files), eternitas
|
||
`py-sdk` (pytest not a declared dep), eternitas `test`.
|
||
|
||
Ground rules already paid for the hard way:
|
||
- Gitea job status: 1=SUCCESS, 2=FAILURE, 5=waiting, 6=running. Not what you'd guess.
|
||
- When a job fails with an opaque forge-shaped error, read the FORGE access log
|
||
(`docker logs windy-git-gitea-1 | grep " 404 "`) — act misattributes the error
|
||
to the previous step.
|
||
- Job logs: `GET /api/v1/repos/{owner}/{repo}/actions/jobs/{id}/logs`. The job id
|
||
must belong to the repo in the path or you get a misleading "job not found".
|
||
Logs are in R2, not on disk.
|
||
- Log into the forge over https://app.windygit.com — Gitea's cookies are Secure,
|
||
so a curl login to http://127.0.0.1:3080 silently keeps no session.
|
||
- verify the WHOLE flow, not the half that curls easily
|
||
- never `git pull -q` in a deploy path; it hides errors
|
||
- fixes go to GitHub, not Windy Git (sync is GitHub -> Windy Git, force-push)
|
||
- if a job fails with `lstat .../<file>: no such file or directory` on an
|
||
action, act's cache rotted: `docker exec windy-git-runner-runner-1 rm -rf
|
||
/root/.cache/act`, then re-run. Safe; it is a container layer, not a volume.
|
||
- check Kit 0's `uptime` before deploying there; two incidents in two days
|
||
```
|