G1: stop probing the tunnel from inside a container
cloudflared binds 127.0.0.1:2000 on the HOST. This process runs in a container whose only route to the host is the bridge gateway (172.17.0.1), where nothing is listening — so the check was permanently red regardless of what the tunnel was actually doing. Binding the metrics endpoint wider would have fixed the probe and made a metrics bind failure capable of taking down ingress. That is a worse trade than losing one row on a dashboard. The check is not silently dropped: /health/full now carries a 'not_checked_here' map naming the tunnel and where its health actually lives (systemd windygit-tunnel). An observer should never have to wonder whether a missing check means healthy or means forgotten. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -23,7 +23,6 @@ from api.app.providers.registry import (
|
||||
EternitasProvider,
|
||||
GiteaProvider,
|
||||
R2Provider,
|
||||
TunnelProvider,
|
||||
)
|
||||
from api.app.routes import health
|
||||
|
||||
@@ -86,7 +85,14 @@ async def lifespan(app: FastAPI):
|
||||
GiteaProvider(settings),
|
||||
R2Provider(settings),
|
||||
EternitasProvider(settings),
|
||||
TunnelProvider(settings),
|
||||
# NOTE: the tunnel is deliberately NOT probed from here. cloudflared
|
||||
# binds its metrics on the host's loopback, so a container can never
|
||||
# reach it -- the check would be permanently red no matter what the
|
||||
# tunnel is doing. A check that structurally cannot succeed is worse
|
||||
# than no check: it trains people to ignore the dashboard, which is
|
||||
# exactly how a fleet canary goes 37 days dead without anyone noticing.
|
||||
# Tunnel health is a host concern and lives where it can be observed:
|
||||
# systemd Restart=always, plus the runbook's `systemctl status`.
|
||||
]
|
||||
|
||||
yield
|
||||
|
||||
Reference in New Issue
Block a user