G1: stop probing the tunnel from inside a container

cloudflared binds 127.0.0.1:2000 on the HOST. This process runs in a container
whose only route to the host is the bridge gateway (172.17.0.1), where nothing
is listening — so the check was permanently red regardless of what the tunnel
was actually doing.

Binding the metrics endpoint wider would have fixed the probe and made a
metrics bind failure capable of taking down ingress. That is a worse trade than
losing one row on a dashboard.

The check is not silently dropped: /health/full now carries a 'not_checked_here'
map naming the tunnel and where its health actually lives (systemd
windygit-tunnel). An observer should never have to wonder whether a missing
check means healthy or means forgotten.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Grant Whitmer
2026-08-11 14:37:53 -04:00
parent a68261a563
commit ce54d488f2
5 changed files with 22 additions and 33 deletions

View File

@@ -23,7 +23,6 @@ from api.app.providers.registry import (
EternitasProvider,
GiteaProvider,
R2Provider,
TunnelProvider,
)
from api.app.routes import health
@@ -86,7 +85,14 @@ async def lifespan(app: FastAPI):
GiteaProvider(settings),
R2Provider(settings),
EternitasProvider(settings),
TunnelProvider(settings),
# NOTE: the tunnel is deliberately NOT probed from here. cloudflared
# binds its metrics on the host's loopback, so a container can never
# reach it -- the check would be permanently red no matter what the
# tunnel is doing. A check that structurally cannot succeed is worse
# than no check: it trains people to ignore the dashboard, which is
# exactly how a fleet canary goes 37 days dead without anyone noticing.
# Tunnel health is a host concern and lives where it can be observed:
# systemd Restart=always, plus the runbook's `systemctl status`.
]
yield