Files
windy-git/docs/TURNOVER-2026-08-14.md
Grant Whitmer 9fc27eabd6 docs: record the real CI root cause — setup-uv resolves uv via the forge API
The localhost->postgres fix was correct but was never what failed these jobs;
they died at step 2. setup-uv v4+ resolves "latest" through GITHUB_API_URL,
which act_runner points at our own forge, so it 404s and every later step is
skipped by success(). Pinned an explicit uv version across 11 repos.

Also records the diagnosis traps that cost the most time: Gitea's job status
enum (1=success, 2=failure), act misattributing the error to the previous step,
the jobs-log API needing a repo-matched id, and Secure cookies defeating a
loopback curl login.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 17:57:04 -04:00

11 KiB
Raw Blame History

Windy Git — turnover, 2026-08-14

Paste the block at the bottom into a fresh terminal. Everything above is context for whoever reads this file directly.

Where things stand

Windy Git is live and in use: app.windygit.com (forge), api.windygit.com (our plane), on Veron 1 behind a Cloudflare Tunnel, zero inbound ports, $0/mo. 143 repos, 85 tests green, health ok on all four checks.

Grant signs in with his existing Windy Word credentials — no second account. Agents authenticate with real Eternitas EPT signature verification and are rate-limited by integrity band.

SOLVED — the CI failures were never about Postgres

The localhost → postgres fix was correct and is worth keeping, but it was not what was failing these jobs. They died at step 2, before Postgres was ever contacted.

Root cause: astral-sh/setup-uv@v4 asks the forge for uv's latest release.

setup-uv v4 added "resolve latest version instead of downloading latest release" (astral-sh/setup-uv#178). Resolution goes through @actions/github, whose octokit reads GITHUB_API_URL — which act_runner points at our forge. So the action requested:

GET https://app.windygit.com/api/v1/repos/astral-sh/uv/releases/latest  →  404

Gitea has no astral-sh/uv, so it answered its standard 404 body, "The target couldn't be found." setup-uv threw that string, act printed it as ::error::, and every later step was skipped by success().

The fix (merged to GitHub, 11 repos): pin an explicit version: on every setup-uv@v4/@v5 step. resolveVersion() short-circuits on an explicit version before any API call, and the download URL is hardcoded to github.com — so the forge round-trip disappears. Pinned to 0.12.5, which is what latest already resolved to.

windy-mind #101, WindyCloud #90, eternitas #150, then the sweep: Windy-Clone #77, windy-agent #355, windy-call #34, windy-cell #31, windy-hand #6, windy-mail #105, windy-search #77, windy-text #29. All merged and synced.

Two things that made this hard to see, both worth keeping

  • act attributes the error to the wrong step. ::error::The target couldn't be found. is printed immediately after actions/checkout's ::remove-matcher, so it reads exactly like a checkout failure. It is not. What settled it was the Gitea access log — sudo docker logs windy-git-gitea-1 | grep " 404 " — which named the real URL at the same millisecond as the job error. When a job fails with an opaque forge-shaped message, go to the forge's access log, not the job log.

  • The natural experiment was sitting right there. windy-registry and windy-drops use setup-uv@v3 and always passed; every v4/v5 caller failed. A version skew across otherwise-identical repos is a diagnosis, not a coincidence.

The jobs API "job not found" that blocked the last session

Not a bug. GET /api/v1/repos/{owner}/{repo}/actions/jobs/{id}/logs requires the job id to belong to the repo in the path — a valid id under the wrong owner/repo 404s. The API works fine; the URLs were mismatched. Logs are readable this way and you do not need the web UI.

Note logs are not on disk — [storage] STORAGE_TYPE = minio sends action logs to R2, so actions_log/ on the host is empty. Read them through the API.

What is still red, and why each one is real

The CI plane is healthy. These are genuine repo defects that were invisible before, because every job died at step 2:

repo job cause
windy-mind tests ruff check — 6 real errors, 5 auto-fixable (ruff check --fix)
WindyCloud lint ruff format --check — 7 files would be reformatted
eternitas py-sdk uv run pytest → Failed to spawn: pytest; pytest isn't a declared dep of that project
WindyCloud docker architectural — see below

WindyCloud's docker job wants to build an image and gets failed to connect to the docker API at unix:///var/run/docker.sock. Job containers deliberately have no docker socket (I-5, and deploy/runner/docker-compose.yml says in so many words not to mount it). Mounting the host socket would hand every workflow root on Veron 1. This needs a decision — buildx-in-dind, a rootless builder, or "this job does not run on Windy Git" — not a quiet socket mount.

The first three are one-line code fixes in their own repos and were left alone on purpose: they are product defects, not forge defects.

Second, smaller finding — act's action cache rots

act caches action repos at /root/.cache/act/<hash> inside the runner container and refreshes them with a go-git mirror fetch of refs/*:refs/*, unforced. That includes refs/pull/*, which GitHub recomputes whenever a base branch moves. Reproduced directly:

! [rejected]  refs/pull/1015/merge -> refs/pull/1015/merge  (non-fast-forward)

which surfaces as Non-terminating error while running 'git clone': some refs were not updated, after which the action does not report Checked out <ref>.

The cache was wiped this session (rm -rf /root/.cache/act, safe — it is in the container layer, not a volume) and the actions resolved cleanly afterwards. This was never proven to fail a job on its own — the setup-uv 404 masked it. It is a watch item, not a closed issue. If actions start failing to resolve, wipe that directory first. It will rot again.

Traps that will waste your time

  • Gitea status codes are not what they look like. 1 = success, 2 = failure, 3 cancelled, 4 skipped, 5 waiting, 6 running, 7 blocked. Reading 1/2 as waiting/running inverts every conclusion you draw from action_run_job.
  • Gitea sets Secure cookies (ROOT_URL is https), so a curl login against http://127.0.0.1:3080 silently keeps no session — it 303s to / and you still get "Sign In". Log in through https://app.windygit.com.
  • There is no rerun API in 1.24.6. POST /api/v1/.../runs/{n}/rerun 404s. Use the web route POST /{owner}/{repo}/actions/runs/{n}/rerun with the session cookie plus an X-Csrf-Token header taken from the _csrf cookie.
  • git pull -q hides errors. A divergent branch once made a "deploy" run 20 minutes against stale code while reporting success. Use git fetch && git merge --ff-only and read the output.
  • Never force-push a branch a deploy checkout tracks (--amend orphaned /srv/windygit/src once).
  • Gitea's env-to-ini SETS but never UNSETS, and sometimes appends a duplicate. GITEA__DEFAULT__APP_NAME does not work at all — Gitea reads APP_NAME from the top level of app.ini; the env var creates a literal [default] section it ignores. Edit app.ini on the host.
  • Cloudflare caches /assets/* for 6h and no token in this stack can purge. Version brand asset filenames (theme-windy.v2.css), not query strings.
  • base64 wraps at 76 chars and corrupts long tokens in test commands → curl (43), phantom HTTP 000. Use base64 -w0.
  • Kit 0 is fragile. 54 containers on 4 vCPU. Two production incidents in two days, both from non-production workloads. Check uptime before deploying anything there, and build before recreating so the swap is seconds.
  • Service containers: use the service NAME and its INTERNAL port (5432), never the mapped host port. The three repos did NOT share one pattern — a naive localhost → postgres swap would have left WindyCloud on port 15432 (it maps 15432:5432) and windy-registry on a job.services.postgres.ports[…] expression.

Open items, roughly by value

  1. Decide what WindyCloud's docker job should do on Windy Git (above). This is the only remaining forge question; it needs a decision, not code.
  2. The three product-level test failures in the table above.
  3. Get non-prod work off Kit 0. 12 dev/demo containers on the box running identity, the CA, mail, Matrix and the broker. Cost two incidents already; the postgres-adapter fix would not have prevented either.
  4. Login is ~4–6s — postgres-adapter.ts:114 forks a node -e process per query. Measured: node startup alone is 1.7s on Kit 0 vs 0.01s on Veron. The fix is one function (persistent worker + pg.Pool), not the "468 call sites" the SOTU scoped. See docs/incidents/2026-08-12-login-latency-analysis.md.
  5. Privileged dind sits beside broad-scoped tokens on the CI host — Grant's call, needs a decision not a code change.
  6. Push-velocity throttling is declared but unenforceable from our plane (git push never touches the API); needs a Gitea pre-receive hook.

Read these first

  • ~/.claude/.../memory/project_windy_git.md — the full record, densest source
  • DNA_STRAND_MASTER_PLAN.md — D-1…D-9 locked decisions, I-1…I-13 invariants
  • docs/AUDIT-fable-2026-08-13.md — second-auditor findings and dispositions
  • docs/CUTOVER.md — the GitHub↔Windy Git migration plan and its one rule

Copy-paste prompt

Picking up Windy Git (agent-native code+model host on Veron 1, live at
app.windygit.com). Read these before doing anything:

  1. ~/.claude/projects/-home-grantwhitmer/memory/project_windy_git.md
  2. ~/windy-git/docs/TURNOVER-2026-08-14.md
  3. ~/windy-git/DNA_STRAND_MASTER_PLAN.md  (D-1..D-9, I-1..I-13)

State: live and in use. Grant signs in with his existing Windy account. Agents
authenticate with real EPT signature verification. 143 repos, 85 tests green.

The CI breakage is SOLVED and verified: setup-uv v4+ resolved uv's "latest"
through GITHUB_API_URL, which act_runner points at our own forge, so it 404'd
("The target couldn't be found.") and every job died at step 2. Fixed by pinning
an explicit uv version across 11 repos; all merged and synced. The migrations
jobs in windy-mind, WindyCloud and eternitas are green.

TASK: one decision, then cleanup.
  1. DECIDE what WindyCloud's `docker` CI job should do here. It needs a Docker
     daemon; job containers deliberately have no socket (I-5 — mounting the host
     socket hands every workflow root on Veron 1). Options: buildx inside the
     existing dind, a rootless builder, or exclude the job. Do NOT mount the
     host socket.
  2. Three genuine product defects, newly visible now that jobs get past step 2:
     windy-mind `tests` (ruff check, 6 errors), WindyCloud `lint` (ruff format,
     7 files), eternitas `py-sdk` (pytest not a declared dep).

Ground rules already paid for the hard way:
  - Gitea job status: 1=SUCCESS, 2=FAILURE, 5=waiting, 6=running. Not what you'd guess.
  - When a job fails with an opaque forge-shaped error, read the FORGE access log
    (`docker logs windy-git-gitea-1 | grep " 404 "`) — act misattributes the error
    to the previous step.
  - Job logs: `GET /api/v1/repos/{owner}/{repo}/actions/jobs/{id}/logs`. The job id
    must belong to the repo in the path or you get a misleading "job not found".
    Logs are in R2, not on disk.
  - Log into the forge over https://app.windygit.com — Gitea's cookies are Secure,
    so a curl login to http://127.0.0.1:3080 silently keeps no session.
  - verify the WHOLE flow, not the half that curls easily
  - never `git pull -q` in a deploy path; it hides errors
  - fixes go to GitHub, not Windy Git (sync is GitHub -> Windy Git, force-push)
  - check Kit 0's `uptime` before deploying there; two incidents in two days