docs: record the real CI root cause — setup-uv resolves uv via the forge API
The localhost->postgres fix was correct but was never what failed these jobs; they died at step 2. setup-uv v4+ resolves "latest" through GITHUB_API_URL, which act_runner points at our own forge, so it 404s and every later step is skipped by success(). Pinned an explicit uv version across 11 repos. Also records the diagnosis traps that cost the most time: Gitea's job status enum (1=success, 2=failure), act misattributing the error to the previous step, the jobs-log API needing a repo-matched id, and Secure cookies defeating a loopback curl login. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -13,56 +13,113 @@ Grant signs in with his existing Windy Word credentials — no second account.
|
||||
Agents authenticate with real Eternitas EPT signature verification and are
|
||||
rate-limited by integrity band.
|
||||
|
||||
## DONE since this was written — the three-repo CI fix
|
||||
## SOLVED — the CI failures were never about Postgres
|
||||
|
||||
All three PRs are **merged and synced**: windy-mind #100, WindyCloud #89,
|
||||
windy-registry #31 (eternitas #149 earlier).
|
||||
The `localhost` → `postgres` fix was correct and is worth keeping, but it was
|
||||
**not** what was failing these jobs. They died at step 2, before Postgres was
|
||||
ever contacted.
|
||||
|
||||
**Result: 1 of 3 verified fixed, 2 still failing for an undetermined reason.**
|
||||
**Root cause: `astral-sh/setup-uv@v4` asks the forge for uv's latest release.**
|
||||
|
||||
- ✅ **windy-registry** — `postgres integration` went **failure → success**. The
|
||||
fix is proven correct.
|
||||
- ❌ **windy-mind**, **WindyCloud** — `migrations` still fails. The DATABASE_URL
|
||||
is definitely right now; the cause is something else and was **not
|
||||
determined** — the jobs API returns "job not found" for the ids the runs
|
||||
report, so logs could not be retrieved that way.
|
||||
setup-uv v4 added "resolve latest version instead of downloading latest release"
|
||||
(astral-sh/setup-uv#178). Resolution goes through `@actions/github`, whose
|
||||
octokit reads **`GITHUB_API_URL`** — which act_runner points at *our forge*. So
|
||||
the action requested:
|
||||
|
||||
**Next session: read those job logs from the Gitea web UI** (`app.windygit.com`
|
||||
→ repo → Actions → the failing run), not the jobs API. Suspicion worth checking
|
||||
first: both use `astral-sh/setup-uv`, and eternitas' equivalent job failed with
|
||||
`error: Failed to spawn: pytest` even after the action resolved — so the uv
|
||||
toolchain may not be landing on PATH in these containers. That would be a
|
||||
different, shared root cause.
|
||||
```
|
||||
GET https://app.windygit.com/api/v1/repos/astral-sh/uv/releases/latest → 404
|
||||
```
|
||||
|
||||
**A trap worth keeping:** the three repos did NOT share one pattern. A naive
|
||||
`localhost` → `postgres` swap would have left **WindyCloud on port 15432** (it
|
||||
maps `15432:5432`) and windy-registry on a `job.services.postgres.ports[…]`
|
||||
expression. Service-name networking always uses the container's **internal**
|
||||
port — 5432 — never the mapped host port.
|
||||
Gitea has no `astral-sh/uv`, so it answered its standard 404 body, *"The target
|
||||
couldn't be found."* setup-uv threw that string, act printed it as `::error::`,
|
||||
and every later step was skipped by `success()`.
|
||||
|
||||
## The original task description (superseded above)
|
||||
**The fix (merged to GitHub, 11 repos):** pin an explicit `version:` on every
|
||||
`setup-uv@v4`/`@v5` step. `resolveVersion()` short-circuits on an explicit
|
||||
version *before* any API call, and the download URL is hardcoded to github.com —
|
||||
so the forge round-trip disappears. Pinned to `0.12.5`, which is what `latest`
|
||||
already resolved to.
|
||||
|
||||
**Three repos need a one-line CI fix.** Their workflows reach a Postgres service
|
||||
at `@localhost:5432`, which works on GitHub-hosted runners (services are
|
||||
port-mapped to the VM) and fails on ours (the job runs *inside* a container, so
|
||||
`localhost` is the job itself). The service is reachable as **`postgres`**.
|
||||
windy-mind #101, WindyCloud #90, eternitas #150, then the sweep: Windy-Clone #77,
|
||||
windy-agent #355, windy-call #34, windy-cell #31, windy-hand #6, windy-mail #105,
|
||||
windy-search #77, windy-text #29. All merged and synced.
|
||||
|
||||
| repo | workflow |
|
||||
|---|---|
|
||||
| `windy-mind` | `migrations.yml` |
|
||||
| `windy-registry` | `ci.yml` |
|
||||
| `WindyCloud` | `ci.yml` |
|
||||
### Two things that made this hard to see, both worth keeping
|
||||
|
||||
`eternitas` was already fixed this way — see **eternitas PR #149** for the exact
|
||||
shape, including the comment explaining why. Fix must go to **GitHub**, not Windy
|
||||
Git: the sync runs GitHub → Windy Git and force-pushes over local edits.
|
||||
- **act attributes the error to the wrong step.** `::error::The target couldn't
|
||||
be found.` is printed immediately after `actions/checkout`'s `::remove-matcher`,
|
||||
so it reads exactly like a checkout failure. It is not. What settled it was the
|
||||
**Gitea access log** — `sudo docker logs windy-git-gitea-1 | grep " 404 "` — which
|
||||
named the real URL at the same millisecond as the job error. When a job fails
|
||||
with an opaque forge-shaped message, go to the forge's access log, not the job log.
|
||||
|
||||
Proven by direct comparison, same runner and same `postgres:16-alpine` image:
|
||||
windy-git's own gate uses `@postgres:5432` and passes its migration round-trip;
|
||||
eternitas' used `@localhost:5432` and failed.
|
||||
- **The natural experiment was sitting right there.** windy-registry and
|
||||
windy-drops use `setup-uv@v3` and always passed; every v4/v5 caller failed. A
|
||||
version skew across otherwise-identical repos is a diagnosis, not a coincidence.
|
||||
|
||||
### The jobs API "job not found" that blocked the last session
|
||||
|
||||
Not a bug. `GET /api/v1/repos/{owner}/{repo}/actions/jobs/{id}/logs` requires the
|
||||
job id to belong to **the repo in the path** — a valid id under the wrong owner/repo
|
||||
404s. The API works fine; the URLs were mismatched. Logs are readable this way and
|
||||
you do **not** need the web UI.
|
||||
|
||||
Note logs are **not on disk** — `[storage] STORAGE_TYPE = minio` sends action logs
|
||||
to R2, so `actions_log/` on the host is empty. Read them through the API.
|
||||
|
||||
## What is still red, and why each one is real
|
||||
|
||||
The CI plane is healthy. These are genuine repo defects that were **invisible
|
||||
before**, because every job died at step 2:
|
||||
|
||||
| repo | job | cause |
|
||||
|---|---|---|
|
||||
| windy-mind | `tests` | `ruff check` — 6 real errors, 5 auto-fixable (`ruff check --fix`) |
|
||||
| WindyCloud | `lint` | `ruff format --check` — 7 files would be reformatted |
|
||||
| eternitas | `py-sdk` | `uv run pytest` → `Failed to spawn: pytest`; pytest isn't a declared dep of that project |
|
||||
| WindyCloud | `docker` | **architectural** — see below |
|
||||
|
||||
`WindyCloud`'s `docker` job wants to build an image and gets `failed to connect
|
||||
to the docker API at unix:///var/run/docker.sock`. Job containers deliberately
|
||||
have **no** docker socket (I-5, and `deploy/runner/docker-compose.yml` says in
|
||||
so many words not to mount it). Mounting the host socket would hand every
|
||||
workflow root on Veron 1. This needs a decision — buildx-in-dind, a rootless
|
||||
builder, or "this job does not run on Windy Git" — not a quiet socket mount.
|
||||
|
||||
The first three are one-line code fixes in their own repos and were left alone
|
||||
on purpose: they are product defects, not forge defects.
|
||||
|
||||
## Second, smaller finding — act's action cache rots
|
||||
|
||||
act caches action repos at `/root/.cache/act/<hash>` inside the runner container
|
||||
and refreshes them with a go-git mirror fetch of `refs/*:refs/*`, unforced. That
|
||||
includes `refs/pull/*`, which GitHub **recomputes** whenever a base branch moves.
|
||||
Reproduced directly:
|
||||
|
||||
```
|
||||
! [rejected] refs/pull/1015/merge -> refs/pull/1015/merge (non-fast-forward)
|
||||
```
|
||||
|
||||
which surfaces as `Non-terminating error while running 'git clone': some refs
|
||||
were not updated`, after which the action does not report `Checked out <ref>`.
|
||||
|
||||
The cache was wiped this session (`rm -rf /root/.cache/act`, safe — it is in the
|
||||
container layer, not a volume) and the actions resolved cleanly afterwards.
|
||||
**This was never proven to fail a job on its own** — the setup-uv 404 masked it.
|
||||
It is a watch item, not a closed issue. If actions start failing to resolve, wipe
|
||||
that directory first. It will rot again.
|
||||
|
||||
## Traps that will waste your time
|
||||
|
||||
- **Gitea status codes are not what they look like.** `1 = success, 2 = failure`,
|
||||
3 cancelled, 4 skipped, 5 waiting, 6 running, 7 blocked. Reading 1/2 as
|
||||
waiting/running inverts every conclusion you draw from `action_run_job`.
|
||||
- **Gitea sets `Secure` cookies** (ROOT_URL is https), so a `curl` login against
|
||||
`http://127.0.0.1:3080` silently keeps no session — it 303s to `/` and you
|
||||
still get "Sign In". Log in through `https://app.windygit.com`.
|
||||
- **There is no rerun API in 1.24.6.** `POST /api/v1/.../runs/{n}/rerun` 404s.
|
||||
Use the web route `POST /{owner}/{repo}/actions/runs/{n}/rerun` with the session
|
||||
cookie plus an `X-Csrf-Token` header taken from the `_csrf` cookie.
|
||||
- **`git pull -q` hides errors.** A divergent branch once made a "deploy" run 20
|
||||
minutes against stale code while reporting success. Use `git fetch && git
|
||||
merge --ff-only` and read the output.
|
||||
@@ -79,20 +136,26 @@ eternitas' used `@localhost:5432` and failed.
|
||||
- **Kit 0 is fragile.** 54 containers on 4 vCPU. Two production incidents in two
|
||||
days, both from *non-production* workloads. Check `uptime` before deploying
|
||||
anything there, and build before recreating so the swap is seconds.
|
||||
- **Service containers**: use the service NAME and its INTERNAL port (5432),
|
||||
never the mapped host port. The three repos did NOT share one pattern — a naive
|
||||
`localhost` → `postgres` swap would have left WindyCloud on port 15432 (it maps
|
||||
`15432:5432`) and windy-registry on a `job.services.postgres.ports[…]` expression.
|
||||
|
||||
## Open items, roughly by value
|
||||
|
||||
1. The three-repo `localhost` fix above.
|
||||
2. **Get non-prod work off Kit 0.** 12 dev/demo containers on the box running
|
||||
1. **Decide what `WindyCloud`'s `docker` job should do on Windy Git** (above).
|
||||
This is the only remaining *forge* question; it needs a decision, not code.
|
||||
2. The three product-level test failures in the table above.
|
||||
3. **Get non-prod work off Kit 0.** 12 dev/demo containers on the box running
|
||||
identity, the CA, mail, Matrix and the broker. Cost two incidents already;
|
||||
the postgres-adapter fix would not have prevented either.
|
||||
3. **Login is ~4–6s** — `postgres-adapter.ts:114` forks a `node -e` process per
|
||||
4. **Login is ~4–6s** — `postgres-adapter.ts:114` forks a `node -e` process per
|
||||
query. Measured: node startup alone is 1.7s on Kit 0 vs 0.01s on Veron. The
|
||||
fix is **one function** (persistent worker + `pg.Pool`), not the "468 call
|
||||
sites" the SOTU scoped. See `docs/incidents/2026-08-12-login-latency-analysis.md`.
|
||||
4. **Privileged dind sits beside broad-scoped tokens** on the CI host — Grant's
|
||||
5. **Privileged dind sits beside broad-scoped tokens** on the CI host — Grant's
|
||||
call, needs a decision not a code change.
|
||||
5. Push-velocity throttling is declared but unenforceable from our plane (git
|
||||
6. Push-velocity throttling is declared but unenforceable from our plane (git
|
||||
push never touches the API); needs a Gitea pre-receive hook.
|
||||
|
||||
## Read these first
|
||||
@@ -114,32 +177,37 @@ app.windygit.com). Read these before doing anything:
|
||||
2. ~/windy-git/docs/TURNOVER-2026-08-14.md
|
||||
3. ~/windy-git/DNA_STRAND_MASTER_PLAN.md (D-1..D-9, I-1..I-13)
|
||||
|
||||
State: live and in use. Grant signs in with his existing Windy account (SSO
|
||||
fixed across windy-pro #346/#347). Agents authenticate with real EPT signature
|
||||
verification. 143 repos, 85 tests green, health ok.
|
||||
State: live and in use. Grant signs in with his existing Windy account. Agents
|
||||
authenticate with real EPT signature verification. 143 repos, 85 tests green.
|
||||
|
||||
TASK: finish the CI fix. Four repos had workflows reaching Postgres through a
|
||||
host port; all four are patched and merged (eternitas #149, windy-mind #100,
|
||||
WindyCloud #89, windy-registry #31). windy-registry's `postgres integration`
|
||||
went failure -> success, proving the approach. But windy-mind and WindyCloud
|
||||
`migrations` still FAIL and I could not determine why.
|
||||
The CI breakage is SOLVED and verified: setup-uv v4+ resolved uv's "latest"
|
||||
through GITHUB_API_URL, which act_runner points at our own forge, so it 404'd
|
||||
("The target couldn't be found.") and every job died at step 2. Fixed by pinning
|
||||
an explicit uv version across 11 repos; all merged and synced. The migrations
|
||||
jobs in windy-mind, WindyCloud and eternitas are green.
|
||||
|
||||
Start by reading those job logs from the GITEA WEB UI (app.windygit.com -> repo
|
||||
-> Actions -> failing run). Do NOT use the jobs API — it returns "job not found"
|
||||
for the ids the runs report, which is what blocked the last session.
|
||||
|
||||
First hypothesis to test: windy-mind, WindyCloud and eternitas all use
|
||||
`astral-sh/setup-uv`, and eternitas' job failed with `error: Failed to spawn:
|
||||
pytest` even after the action resolved correctly. The uv toolchain may not be
|
||||
landing on PATH inside these job containers — one shared root cause rather than
|
||||
three.
|
||||
TASK: one decision, then cleanup.
|
||||
1. DECIDE what WindyCloud's `docker` CI job should do here. It needs a Docker
|
||||
daemon; job containers deliberately have no socket (I-5 — mounting the host
|
||||
socket hands every workflow root on Veron 1). Options: buildx inside the
|
||||
existing dind, a rootless builder, or exclude the job. Do NOT mount the
|
||||
host socket.
|
||||
2. Three genuine product defects, newly visible now that jobs get past step 2:
|
||||
windy-mind `tests` (ruff check, 6 errors), WindyCloud `lint` (ruff format,
|
||||
7 files), eternitas `py-sdk` (pytest not a declared dep).
|
||||
|
||||
Ground rules already paid for the hard way:
|
||||
- Gitea job status: 1=SUCCESS, 2=FAILURE, 5=waiting, 6=running. Not what you'd guess.
|
||||
- When a job fails with an opaque forge-shaped error, read the FORGE access log
|
||||
(`docker logs windy-git-gitea-1 | grep " 404 "`) — act misattributes the error
|
||||
to the previous step.
|
||||
- Job logs: `GET /api/v1/repos/{owner}/{repo}/actions/jobs/{id}/logs`. The job id
|
||||
must belong to the repo in the path or you get a misleading "job not found".
|
||||
Logs are in R2, not on disk.
|
||||
- Log into the forge over https://app.windygit.com — Gitea's cookies are Secure,
|
||||
so a curl login to http://127.0.0.1:3080 silently keeps no session.
|
||||
- verify the WHOLE flow, not the half that curls easily
|
||||
- never `git pull -q` in a deploy path; it hides errors
|
||||
- fixes go to GitHub, not Windy Git (sync is GitHub -> Windy Git, force-push)
|
||||
- service containers: use the service NAME and its INTERNAL port (5432),
|
||||
never the mapped host port
|
||||
- check Kit 0's `uptime` before deploying there; two incidents in two days
|
||||
from non-production workloads
|
||||
```
|
||||
|
||||
Reference in New Issue
Block a user