The probe's own log was never retrievable through the jobs API, but the
question it asked was answered better by a direct comparison of two real
workflows on the same runner and image:
windy-git gate @postgres:5432 -> passes its migration round-trip
eternitas migrations @localhost:5432 -> failed
Also scopes test_g73 to workflows that actually run Python. It failed the probe
for not pinning a version when the probe only shelled out to psql — the test
being wrong rather than the workflow.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
Several migrated workflows hardcode postgres at localhost:5432, which is
correct on GitHub-hosted runners (services are port-mapped to the VM) and
suspect on Gitea Actions (the job runs IN a container, so localhost is the job).
Prove which form works before rewriting anyone's workflow.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
Today's outage is the whole design brief: /health returned 200 for the entire
hour that login was dead. A canary watching /health would have stayed green
while nobody in the ecosystem could sign in. So the login probe is here and it
is the one that matters.
Three rules it obeys:
- never green for something it did not prove (I-8)
- alert on TRANSITIONS, not every run — a canary people filter is a dead
canary, which is how the last one sat 37 days dead unnoticed
- run where the watched thing cannot take it down: Veron 1, never Kit 0
Two independent signals, so losing one still leaves the other: an email via
Resend on state change, and a non-zero exit that turns the CI run red in the
forge itself.
Alerts say what broke in human terms — 'a human can actually sign in' — rather
than only naming an endpoint.
Verified against production: 7/7 green including login at 17.1s; a forced 404
reports DOWN; a 1s threshold reports SLOW at 23.3s.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Run 7 passed end to end: Python 3.12.13, ruff clean, vocabulary audit clean,
42 tests, migration upgrade->downgrade->upgrade against real Postgres, and
'I-12 holds: env override ignored'.
This is the finding both August audits converged on — 'nothing anywhere checks
whether a decision reached production' — closed for this repo. GitHub Actions
is billing-locked and cannot run on private repos at all, even self-hosted.
This can, on hardware Grant already owns, for zero dollars.
I-5 proven by inspection, not assertion: CI containers sit on
windy-git-runner_jobs, the forge database on windy-git_default. Disjoint. And
neither CI container holds the host docker socket.
Adds strand G7A: seven runs to first green, and each failure presented as a
different problem than it was. The worst was a Python version mismatch that
appeared as a 14-minute hang rather than an error, because pip answered
3.10-vs-3.12 by backtracking through every dependency's release history with
-q hiding it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Runs 3 and 6 were killed by me restarting the runner mid-job, not by any
defect in the workflow. Config verified on disk first this time.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Trading one wedge for another: running the job in python:3.12-bookworm fixed
the resolver spiral but broke checkout, because actions/checkout is a
JavaScript action and the official Python images carry no node — 'executable
file not found in $PATH'.
setup-python on the act image has both. Test now accepts either a pinned
container image or an explicit setup-python version, and still checks it
against pyproject's requires-python.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first real CI run wedged for 14 minutes. Not a network problem, not the
isolation work — catthehacker/ubuntu:act-22.04 ships Python 3.10.12 while this
project declares requires-python >=3.12, and pip answered that by backtracking
through the entire release history of every dependency looking for something
3.10-compatible. At 100% CPU, with -q hiding every line of it, and it would
have churned until the runner's 30m timeout.
A version mismatch presenting as a hang rather than an error is worth a test,
so there is one: the workflow's python image must satisfy pyproject's
requires-python, checked by parsing both rather than by eyeballing them.
Also: every job now has timeout-minutes. A wedged step should be a red check in
minutes, not an occupied runner for half an hour.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
I-5 says runners execute untrusted code and must be isolated by machine
boundary. act_runner needs a Docker daemon to start job containers, and the
tempting move — what every published example does — is to mount the host's
/var/run/docker.sock. That hands every workflow, including whatever a
transitive dependency's postinstall script feels like doing, the ability to
start a privileged container mounting / — root on Grant's workstation.
Instead the runner talks to its OWN dind daemon:
- runner (TRUSTED, the act_runner daemon) sits on the forge network only to
collect jobs from gitea:3000
- dind and every job container it spawns are UNTRUSTED, on a private network
with no route to the forge, its Postgres, or its .env
- jobs cannot bind-mount from the daemon host (valid_volumes: []) and are not
handed the runner's own socket (docker_host: -)
- separate compose project, cpu/memory bounded — Veron 1 is Grant's
workstation, not a dedicated build box
The gate itself now runs as a workflow, including the migration round-trip that
already caught two bugs review did not, and the I-12 check that a COMMIT_SHA
env override cannot change what /version reports.
Labels are explicit and pinned. A workflow naming a label nobody provides
queues forever and presents as a hung CI system rather than a typo — which is
what ubuntu-latest does on every windy-registry run today.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>