4 Commits

Author SHA1 Message Date
390c1e7479 ops: move tunnel metrics to 2001, sync windy-git into itself
Some checks failed
check / gate (push) Successful in 19s
canary / probe (push) Failing after 7s
windygit-tunnel had crash-looped ~91k times: another project's
cornercall-tunnel holds 127.0.0.1:2000, and cloudflared exits when it
cannot bind its metrics port. Ingress only survived because a stray
cloudflared.service ran the same config. That unit is now disabled and
/etc/cloudflared/config.yml uses metrics 127.0.0.1:2001.

Also add windy-git to the GitHub->Windy Git sync list; its self-hosted
copy was stuck 3 commits behind (only check + canary workflows, no
deploys, so syncing is safe).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 01:29:57 -04:00
2b30b0ac99 docs: record the runner bump, the act cache flake, and the remaining reds
Second root cause found and fixed: act_runner 0.2.11 predates
`runs.using: node24`, so any repo on actions/checkout@v5 died before its first
step. Bumped to 0.6.1; Windy-Clone went 4/4 red to 4/4 green.

Also upgrades the act-cache note from "watch item" to a confirmed job failure
(lstat on a vanished file mid-tar), with the wipe command and the annotated-tag
dead end that looks like a wrong checkout but isn't.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 19:09:26 -04:00
cd7dd9b7ae ci: bump act_runner 0.2.11 -> 0.6.1 for node24 action support
0.2.11's bundled act only knows runs.using node12/node16/node20, so any repo
pinning a current action major (actions/checkout@v5, actions/setup-python@v6)
fails before its first step with "The runs.using key in action.yml must be one
of: [...], got node24". Windy-Clone is how this surfaced.

Verified node24 is absent from the 0.2.11 binary and present in 0.6.1, and that
every key in deploy/runner/config.yaml still exists in 0.6.1's schema.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 19:00:07 -04:00
9fc27eabd6 docs: record the real CI root cause — setup-uv resolves uv via the forge API
The localhost->postgres fix was correct but was never what failed these jobs;
they died at step 2. setup-uv v4+ resolves "latest" through GITHUB_API_URL,
which act_runner points at our own forge, so it 404s and every later step is
skipped by success(). Pinned an explicit uv version across 11 repos.

Also records the diagnosis traps that cost the most time: Gitea's job status
enum (1=success, 2=failure), act misattributing the error to the previous step,
the jobs-log API needing a repo-matched id, and Secure cookies defeating a
loopback curl login.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 17:57:04 -04:00
6 changed files with 201 additions and 67 deletions

View File

@@ -22,7 +22,7 @@ boot guard in `api/app/main.py` that refuses to start there in production.
| 8600 | `windy-git-api` — our plane |
| **3080** | Gitea — host 3000 and 3300 are taken by resident projects on Veron 1 |
| 5432 | Postgres |
| 2000 | cloudflared metrics (probe target) |
| 2001 | cloudflared metrics — NOT 2000: `cornercall-tunnel` (another project) takes 2000, and a metrics bind failure kills the whole tunnel |
## Ingress — Cloudflare Tunnel `windy-git`

View File

@@ -105,7 +105,7 @@ class DatabaseProvider(Provider):
# TunnelProvider was removed deliberately. See the note in main.py: cloudflared
# binds 127.0.0.1:2000 on the HOST, and this process runs in a container whose
# binds 127.0.0.1:2001 on the HOST, and this process runs in a container whose
# only route to the host is the bridge gateway (172.17.0.1), where nothing is
# listening. Binding the metrics endpoint wider would fix the probe and make a
# metrics bind failure able to take down ingress -- a worse trade than losing

View File

@@ -50,7 +50,16 @@ services:
restart: unless-stopped
runner:
image: docker.io/gitea/act_runner:0.2.11
# 0.2.11 -> 0.6.1 on 2026-08-14. The bundled act in 0.2.11 only knows
# `runs.using: node12|node16|node20`, so ANY repo pinning a current action
# major dies before its first step with "The runs.using key in action.yml
# must be one of: [...], got node24" — Windy-Clone on actions/checkout@v5
# is how this surfaced. Verified: `node24` is absent from the 0.2.11 binary
# and present in 0.6.1. Every key in this directory's config.yaml still
# exists in 0.6.1's schema (0.6.1 only ADDS keys), so the config carries
# over unchanged. Rollback is re-pinning 0.2.11; the registration in the
# runner-data volume survives either way.
image: docker.io/gitea/act_runner:0.6.1
depends_on: [dind]
environment:
# The runner reaches its OWN daemon. Never the host's.

View File

@@ -22,7 +22,7 @@ boot in production if it finds itself on `72.60.118.54`.
|---|---|
| `127.0.0.1:3080` | Gitea (host 3000 is a resident node dev server; 3300 is nginx — **do not fight them for a port**) |
| `127.0.0.1:8600` | windy-git API |
| `127.0.0.1:2000` | cloudflared metrics |
| `127.0.0.1:2001` | cloudflared metrics (`metrics:` in `/etc/cloudflared/config.yml`) — **not 2000**, see Troubleshooting |
**No inbound port is opened.** cloudflared dials out, so the dynamic residential
IP is irrelevant and there is no firewall hole to maintain.
@@ -77,6 +77,14 @@ sudo ss -tlnp | grep -E "3080|8600" # both must be 127.0.0.1
**A hostname returns 530 or won't resolve** — the tunnel is down. `sudo systemctl
restart windygit-tunnel`, then `journalctl -u windygit-tunnel -n 50`.
**`windygit-tunnel` crash-loops with `bind: address already in use` on the metrics
port** — cloudflared exits if it cannot bind `metrics:`, taking ingress with it.
Until 2026-09-23 this unit restarted ~91,000 times because another project's
`cornercall-tunnel` held 127.0.0.1:2000; ingress only survived because a stray
generic `cloudflared.service` ran the same config (now disabled). Windy Git's
metrics port is **2001**. `sudo ss -ltnp | grep :2001` names any squatter.
Keep exactly ONE unit running `/etc/cloudflared/config.yml`: `windygit-tunnel`.
**TLS handshake fails with `curl` exit 35 and no HTTP status at all** — someone
added a **two-level** hostname. Free Universal SSL covers `windygit.com` and
`*.windygit.com` only. The request dies before the tunnel is consulted, so it

View File

@@ -13,56 +13,150 @@ Grant signs in with his existing Windy Word credentials — no second account.
Agents authenticate with real Eternitas EPT signature verification and are
rate-limited by integrity band.
## DONE since this was written — the three-repo CI fix
## SOLVED — the CI failures were never about Postgres
All three PRs are **merged and synced**: windy-mind #100, WindyCloud #89,
windy-registry #31 (eternitas #149 earlier).
The `localhost` → `postgres` fix was correct and is worth keeping, but it was
**not** what was failing these jobs. They died at step 2, before Postgres was
ever contacted.
**Result: 1 of 3 verified fixed, 2 still failing for an undetermined reason.**
**Root cause: `astral-sh/setup-uv@v4` asks the forge for uv's latest release.**
- ✅ **windy-registry** — `postgres integration` went **failure → success**. The
fix is proven correct.
- ❌ **windy-mind**, **WindyCloud** — `migrations` still fails. The DATABASE_URL
is definitely right now; the cause is something else and was **not
determined** — the jobs API returns "job not found" for the ids the runs
report, so logs could not be retrieved that way.
setup-uv v4 added "resolve latest version instead of downloading latest release"
(astral-sh/setup-uv#178). Resolution goes through `@actions/github`, whose
octokit reads **`GITHUB_API_URL`** — which act_runner points at *our forge*. So
the action requested:
**Next session: read those job logs from the Gitea web UI** (`app.windygit.com`
→ repo → Actions → the failing run), not the jobs API. Suspicion worth checking
first: both use `astral-sh/setup-uv`, and eternitas' equivalent job failed with
`error: Failed to spawn: pytest` even after the action resolved — so the uv
toolchain may not be landing on PATH in these containers. That would be a
different, shared root cause.
```
GET https://app.windygit.com/api/v1/repos/astral-sh/uv/releases/latest → 404
```
**A trap worth keeping:** the three repos did NOT share one pattern. A naive
`localhost` → `postgres` swap would have left **WindyCloud on port 15432** (it
maps `15432:5432`) and windy-registry on a `job.services.postgres.ports[…]`
expression. Service-name networking always uses the container's **internal**
port — 5432 — never the mapped host port.
Gitea has no `astral-sh/uv`, so it answered its standard 404 body, *"The target
couldn't be found."* setup-uv threw that string, act printed it as `::error::`,
and every later step was skipped by `success()`.
## The original task description (superseded above)
**The fix (merged to GitHub, 11 repos):** pin an explicit `version:` on every
`setup-uv@v4`/`@v5` step. `resolveVersion()` short-circuits on an explicit
version *before* any API call, and the download URL is hardcoded to github.com —
so the forge round-trip disappears. Pinned to `0.12.5`, which is what `latest`
already resolved to.
**Three repos need a one-line CI fix.** Their workflows reach a Postgres service
at `@localhost:5432`, which works on GitHub-hosted runners (services are
port-mapped to the VM) and fails on ours (the job runs *inside* a container, so
`localhost` is the job itself). The service is reachable as **`postgres`**.
windy-mind #101, WindyCloud #90, eternitas #150, then the sweep: Windy-Clone #77,
windy-agent #355, windy-call #34, windy-cell #31, windy-hand #6, windy-mail #105,
windy-search #77, windy-text #29. All merged and synced.
| repo | workflow |
|---|---|
| `windy-mind` | `migrations.yml` |
| `windy-registry` | `ci.yml` |
| `WindyCloud` | `ci.yml` |
### Two things that made this hard to see, both worth keeping
`eternitas` was already fixed this way — see **eternitas PR #149** for the exact
shape, including the comment explaining why. Fix must go to **GitHub**, not Windy
Git: the sync runs GitHub → Windy Git and force-pushes over local edits.
- **act attributes the error to the wrong step.** `::error::The target couldn't
be found.` is printed immediately after `actions/checkout`'s `::remove-matcher`,
so it reads exactly like a checkout failure. It is not. What settled it was the
**Gitea access log** — `sudo docker logs windy-git-gitea-1 | grep " 404 "` — which
named the real URL at the same millisecond as the job error. When a job fails
with an opaque forge-shaped message, go to the forge's access log, not the job log.
Proven by direct comparison, same runner and same `postgres:16-alpine` image:
windy-git's own gate uses `@postgres:5432` and passes its migration round-trip;
eternitas' used `@localhost:5432` and failed.
- **The natural experiment was sitting right there.** windy-registry and
windy-drops use `setup-uv@v3` and always passed; every v4/v5 caller failed. A
version skew across otherwise-identical repos is a diagnosis, not a coincidence.
### The jobs API "job not found" that blocked the last session
Not a bug. `GET /api/v1/repos/{owner}/{repo}/actions/jobs/{id}/logs` requires the
job id to belong to **the repo in the path** — a valid id under the wrong owner/repo
404s. The API works fine; the URLs were mismatched. Logs are readable this way and
you do **not** need the web UI.
Note logs are **not on disk** — `[storage] STORAGE_TYPE = minio` sends action logs
to R2, so `actions_log/` on the host is empty. Read them through the API.
## SOLVED — second root cause: the runner was four majors behind
Windy-Clone failed for a completely different reason: `The runs.using key in
action.yml must be one of: [composite docker node12 node16 node20 go], got
**node24**`. act_runner **0.2.11**'s bundled act predates node24, so any repo
pinning a current action major (`actions/checkout@v5`, `actions/setup-python@v6`)
died before its first step.
**Bumped to `gitea/act_runner:0.6.1`** (`cd7dd9b`). Verified beforehand that
`node24` is absent from the 0.2.11 binary and present in 0.6.1, and that every
key in `deploy/runner/config.yaml` still exists in 0.6.1's schema — 0.6.1 only
*adds* keys, so the config carried over untouched. The runner re-declared with
the same id and labels `[veron-1 linux-x64 self-hosted linux x64]`; the
registration in the `runner-data` volume survived. Windy-Clone went 4/4 red →
4/4 green. Rollback is re-pinning 0.2.11; registration is backed up at
`/srv/windygit/runner-registration.bak`.
## What is still red, and why each one is real
The CI plane is healthy. windy-mind is fully green (`742 passed, 1 skipped`).
What remains are genuine repo defects that were **invisible before**, because
every job died at step 2:
| repo | job | cause |
|---|---|---|
| WindyCloud | `lint` | `ruff format --check` — 7 files would be reformatted |
| eternitas | `py-sdk` | `uv run pytest` → `Failed to spawn: pytest`; pytest isn't a declared dep of that project |
| eternitas | `test` | needs its own look |
| windy-agent | `test (3.12/3.13/3.14)` | all three died **together** at 22:50:19 after ~43 min, mid-suite at 64%, with no verdict in the log. Not the 30m `runner.timeout` (that would have fired at 22:36) and not the runner bump (that was 23:00). Something bulk-killed them; unexplained. |
| WindyCloud | `docker`, windy-search `Docker build` | **architectural** — see below |
`WindyCloud`'s `docker` job wants to build an image and gets `failed to connect
to the docker API at unix:///var/run/docker.sock`. Job containers deliberately
have **no** docker socket (I-5, and `deploy/runner/docker-compose.yml` says in
so many words not to mount it). Mounting the host socket would hand every
workflow root on Veron 1. This needs a decision — buildx-in-dind, a rootless
builder, or "this job does not run on Windy Git" — not a quiet socket mount.
The WindyCloud `lint` and eternitas `py-sdk` rows are small fixes in their own
repos, left alone on purpose: they are product defects, not forge defects.
## Second, smaller finding — act's action cache rots
act caches action repos at `/root/.cache/act/<hash>` inside the runner container
and refreshes them with a go-git mirror fetch of `refs/*:refs/*`, unforced. That
includes `refs/pull/*`, which GitHub **recomputes** whenever a base branch moves.
Reproduced directly:
```
! [rejected] refs/pull/1015/merge -> refs/pull/1015/merge (non-fast-forward)
```
which surfaces as `Non-terminating error while running 'git clone': some refs
were not updated`, after which the action does not report `Checked out <ref>`.
It **has** now failed a job on its own. After the runner bump, `windy-mind tests`
died with:
```
❌ Failure - Main Install uv
lstat /root/.cache/act/d3e6…/.git-blame-ignore-revs: no such file or directory
```
act tars the cached action directory into the job container, and a file vanished
mid-walk. The cache dir had been created at 23:01 and mutated again at 23:03,
with `.gitignore` showing as deleted — act removes it before `docker cp`. The
likely mechanism is **concurrent jobs sharing one cache dir**: `capacity: 4`, and
windy-mind fires four jobs at once that all use setup-uv. Wiping the cache and
re-running the job alone made it pass (`742 passed, 1 skipped`). *Mechanism not
isolated* — the wipe alone may have been sufficient.
One dead end worth not repeating: the cached worktree sits at `38f3f104` while
`git rev-parse v4` says `e4db8464`. That is **not** a wrong checkout — `v4` is an
*annotated tag*, and `v4^{commit}` is `38f3f104`. Don't chase it.
Wipe with `sudo docker exec windy-git-runner-runner-1 rm -rf /root/.cache/act`
(safe — container layer, not a volume). It will rot again. A real fix is either
lowering `capacity` or upstream act; neither was attempted.
## Traps that will waste your time
- **Gitea status codes are not what they look like.** `1 = success, 2 = failure`,
3 cancelled, 4 skipped, 5 waiting, 6 running, 7 blocked. Reading 1/2 as
waiting/running inverts every conclusion you draw from `action_run_job`.
- **Gitea sets `Secure` cookies** (ROOT_URL is https), so a `curl` login against
`http://127.0.0.1:3080` silently keeps no session — it 303s to `/` and you
still get "Sign In". Log in through `https://app.windygit.com`.
- **There is no rerun API in 1.24.6.** `POST /api/v1/.../runs/{n}/rerun` 404s.
Use the web route `POST /{owner}/{repo}/actions/runs/{n}/rerun` with the session
cookie plus an `X-Csrf-Token` header taken from the `_csrf` cookie.
- **`git pull -q` hides errors.** A divergent branch once made a "deploy" run 20
minutes against stale code while reporting success. Use `git fetch && git
merge --ff-only` and read the output.
@@ -79,20 +173,30 @@ eternitas' used `@localhost:5432` and failed.
- **Kit 0 is fragile.** 54 containers on 4 vCPU. Two production incidents in two
days, both from *non-production* workloads. Check `uptime` before deploying
anything there, and build before recreating so the swap is seconds.
- **Service containers**: use the service NAME and its INTERNAL port (5432),
never the mapped host port. The three repos did NOT share one pattern — a naive
`localhost` → `postgres` swap would have left WindyCloud on port 15432 (it maps
`15432:5432`) and windy-registry on a `job.services.postgres.ports[…]` expression.
## Open items, roughly by value
1. The three-repo `localhost` fix above.
2. **Get non-prod work off Kit 0.** 12 dev/demo containers on the box running
1. **Decide what the image-building jobs should do on Windy Git** — WindyCloud
`docker` and windy-search `Docker build` (above). The only remaining *forge*
question; it needs a decision, not code.
2. **windy-agent's three `test` jobs were bulk-killed at 22:50:19** after ~43
minutes, mid-suite, with no verdict. Unexplained and not the runner bump.
3. The product-level test failures in the table above.
4. **act's action cache race** — lower `capacity` below 4, or accept re-runs.
5. **Get non-prod work off Kit 0.** 12 dev/demo containers on the box running
identity, the CA, mail, Matrix and the broker. Cost two incidents already;
the postgres-adapter fix would not have prevented either.
3. **Login is ~4–6s** — `postgres-adapter.ts:114` forks a `node -e` process per
6. **Login is ~4–6s** — `postgres-adapter.ts:114` forks a `node -e` process per
query. Measured: node startup alone is 1.7s on Kit 0 vs 0.01s on Veron. The
fix is **one function** (persistent worker + `pg.Pool`), not the "468 call
sites" the SOTU scoped. See `docs/incidents/2026-08-12-login-latency-analysis.md`.
4. **Privileged dind sits beside broad-scoped tokens** on the CI host — Grant's
7. **Privileged dind sits beside broad-scoped tokens** on the CI host — Grant's
call, needs a decision not a code change.
5. Push-velocity throttling is declared but unenforceable from our plane (git
8. Push-velocity throttling is declared but unenforceable from our plane (git
push never touches the API); needs a Gitea pre-receive hook.
## Read these first
@@ -114,32 +218,45 @@ app.windygit.com). Read these before doing anything:
2. ~/windy-git/docs/TURNOVER-2026-08-14.md
3. ~/windy-git/DNA_STRAND_MASTER_PLAN.md (D-1..D-9, I-1..I-13)
State: live and in use. Grant signs in with his existing Windy account (SSO
fixed across windy-pro #346/#347). Agents authenticate with real EPT signature
verification. 143 repos, 85 tests green, health ok.
State: live and in use. Grant signs in with his existing Windy account. Agents
authenticate with real EPT signature verification. 143 repos, 85 tests green.
TASK: finish the CI fix. Four repos had workflows reaching Postgres through a
host port; all four are patched and merged (eternitas #149, windy-mind #100,
WindyCloud #89, windy-registry #31). windy-registry's `postgres integration`
went failure -> success, proving the approach. But windy-mind and WindyCloud
`migrations` still FAIL and I could not determine why.
TWO CI root causes are SOLVED and verified:
(a) setup-uv v4+ resolved uv's "latest" through GITHUB_API_URL, which
act_runner points at our own forge, so it 404'd ("The target couldn't be
found.") and every job died at step 2. Fixed by pinning an explicit uv
version across 11 repos.
(b) act_runner 0.2.11 predates `runs.using: node24`, so any repo on
actions/checkout@v5 died before its first step. Bumped to 0.6.1.
windy-mind is fully green (742 passed). Windy-Clone went 4/4 red to 4/4 green.
Start by reading those job logs from the GITEA WEB UI (app.windygit.com -> repo
-> Actions -> failing run). Do NOT use the jobs API — it returns "job not found"
for the ids the runs report, which is what blocked the last session.
First hypothesis to test: windy-mind, WindyCloud and eternitas all use
`astral-sh/setup-uv`, and eternitas' job failed with `error: Failed to spawn:
pytest` even after the action resolved correctly. The uv toolchain may not be
landing on PATH inside these job containers — one shared root cause rather than
three.
TASK: one decision, then cleanup.
1. DECIDE what the image-building CI jobs should do here — WindyCloud `docker`
and windy-search `Docker build`. They need a Docker daemon; job containers
deliberately have no socket (I-5 — mounting the host socket hands every
workflow root on Veron 1). Options: buildx inside the existing dind, a
rootless builder, or exclude the job. Do NOT mount the host socket.
2. windy-agent's three `test` jobs were bulk-killed together at 22:50:19 after
~43 min, mid-suite, with no verdict in the log. Not the 30m runner.timeout,
not the runner bump. Unexplained — worth a look.
3. Small product defects: WindyCloud `lint` (ruff format, 7 files), eternitas
`py-sdk` (pytest not a declared dep), eternitas `test`.
Ground rules already paid for the hard way:
- Gitea job status: 1=SUCCESS, 2=FAILURE, 5=waiting, 6=running. Not what you'd guess.
- When a job fails with an opaque forge-shaped error, read the FORGE access log
(`docker logs windy-git-gitea-1 | grep " 404 "`) — act misattributes the error
to the previous step.
- Job logs: `GET /api/v1/repos/{owner}/{repo}/actions/jobs/{id}/logs`. The job id
must belong to the repo in the path or you get a misleading "job not found".
Logs are in R2, not on disk.
- Log into the forge over https://app.windygit.com — Gitea's cookies are Secure,
so a curl login to http://127.0.0.1:3080 silently keeps no session.
- verify the WHOLE flow, not the half that curls easily
- never `git pull -q` in a deploy path; it hides errors
- fixes go to GitHub, not Windy Git (sync is GitHub -> Windy Git, force-push)
- service containers: use the service NAME and its INTERNAL port (5432),
never the mapped host port
- if a job fails with `lstat .../<file>: no such file or directory` on an
action, act's cache rotted: `docker exec windy-git-runner-runner-1 rm -rf
/root/.cache/act`, then re-run. Safe; it is a container layer, not a volume.
- check Kit 0's `uptime` before deploying there; two incidents in two days
from non-production workloads
```

View File

@@ -38,7 +38,7 @@ FAILED=0
# Repos Windy Git tracks FROM GitHub. Remove a repo from this list at the moment
# it flips to Windy-Git-first, or the sync will fight its authors and win.
REPOS="${SYNC_REPOS:-windy-calendar windy-search windy-registry Windy-Clone WindyCloud windy-cloud-sites windy-mind eternitas windy-agent}"
REPOS="${SYNC_REPOS:-windy-calendar windy-search windy-registry Windy-Clone WindyCloud windy-cloud-sites windy-mind eternitas windy-agent windy-git}"
mkdir -p "$WORK"
log() { printf '[sync %s] %s\n' "$(date -u +%H:%M:%SZ)" "$*"; }