windygit-tunnel had crash-looped ~91k times: another project's
cornercall-tunnel holds 127.0.0.1:2000, and cloudflared exits when it
cannot bind its metrics port. Ingress only survived because a stray
cloudflared.service ran the same config. That unit is now disabled and
/etc/cloudflared/config.yml uses metrics 127.0.0.1:2001.
Also add windy-git to the GitHub->Windy Git sync list; its self-hosted
copy was stuck 3 commits behind (only check + canary workflows, no
deploys, so syncing is safe).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cloudflared binds 127.0.0.1:2000 on the HOST. This process runs in a container
whose only route to the host is the bridge gateway (172.17.0.1), where nothing
is listening — so the check was permanently red regardless of what the tunnel
was actually doing.
Binding the metrics endpoint wider would have fixed the probe and made a
metrics bind failure capable of taking down ingress. That is a worse trade than
losing one row on a dashboard.
The check is not silently dropped: /health/full now carries a 'not_checked_here'
map naming the tunnel and where its health actually lives (systemd
windygit-tunnel). An observer should never have to wonder whether a missing
check means healthy or means forgotten.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
app.windygit.com / api.windygit.com / models.windygit.com are serving over
HTTPS with ZERO inbound ports open on Grant's network.
- tunnel 4e856c5d, 4 registered edge connections, systemd-managed and bounded
- three proxied single-level CNAMEs (Free Universal SSL covers them; a
two-level name would need ACM and would die in the TLS handshake)
- services bound to 127.0.0.1 with configurable host ports — Veron 1 is
Grant's workstation and 3000/3300 belong to other projects
- docs/RUNBOOK-VERON.md
I-12 PROVEN IN PRODUCTION: /version reports source=baked with a sha equal to
the deployed HEAD.
Also fixed: the tunnel health probe targeted localhost from inside a container,
so it was permanently red. A check that is always red is as useless as one that
is always green — it is how a fleet canary goes 37 days dead unnoticed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Strand G0 complete and VERIFIED against real Postgres, not asserted.
- FastAPI plane, fail-closed provider seams, repair-pointer error taxonomy
- migration 001: all 10 tables incl. repo_type NOT NULL and model_cards (I-7)
- 17 invariant tests, ruff clean, vocabulary audit clean
Two bugs found by RUNNING it that review would not have caught:
1. SQLAlchemy Enum persists .name, not .value — so RepoState.deleted_soft
and CreatedVia.imported would have written labels migration 001 never
declared, failing at runtime rather than at review. Pinned via
values_callable.
2. op.create_table asks each Enum to emit its own CREATE TYPE with no
checkfirst, so the second reference raised DuplicateObject and the
migration died halfway. Types are now created once, referenced with
create_type=False.
Proven live, with the hostile env var set:
- I-12: COMMIT_SHA=deadbeef... in the environment, /version reports real HEAD.
That env pin is the documented root cause of nine sibling services
misreporting their commit; here it is structurally ignored.
- I-8: three unconfigured providers -> status degraded, HTTP 503, each saying
'refusing to report healthy'. No mock, no false green.
- G0.4: upgrade -> downgrade -> upgrade round-trip clean (10 -> 0 -> 10).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>