G7: per-job networks — fixes service DNS and tightens isolation
Some checks failed
check / gate (push) Failing after 13m33s

The migration step failed with 'could not translate host name postgres'. The
Postgres service container was healthy; the job simply could not name it,
because service DNS aliases only exist on a per-job network and I had pinned
container.network to the flat 'bridge'.

That choice was wrong in both directions: it broke service containers AND it
was weaker isolation, since every concurrent job shared one bridge and could
see its neighbours. A per-job network is stricter and correct — and still has
no route to the forge, because these networks live inside the dind daemon,
which has no forge attachment at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Grant Whitmer
2026-08-12 11:18:51 -04:00
parent 8c4bdd2bc2
commit 6b96087554
2 changed files with 22 additions and 4 deletions

View File

@@ -409,6 +409,16 @@ def test_i05_runner_never_mounts_the_host_docker_socket():
assert "/var/run/docker.sock" not in ln, "I-5: never mount the host docker socket" assert "/var/run/docker.sock" not in ln, "I-5: never mount the host docker socket"
def test_i05_jobs_get_a_network_per_job_not_a_shared_bridge():
"""A shared flat bridge lets concurrent jobs see each other, and breaks
service-container DNS (service aliases only exist on a per-job network).
Per-job is both stricter and correct."""
cfg = (ROOT / "deploy" / "runner" / "config.yaml").read_text()
active = [ln for ln in cfg.splitlines() if ln.strip() and not ln.strip().startswith("#")]
net = [ln for ln in active if ln.strip().startswith("network:")]
assert net and net[0].strip() == 'network: ""', "jobs must get a per-job network"
def test_i05_jobs_cannot_bind_mount_from_the_daemon_host(): def test_i05_jobs_cannot_bind_mount_from_the_daemon_host():
cfg = (ROOT / "deploy" / "runner" / "config.yaml").read_text() cfg = (ROOT / "deploy" / "runner" / "config.yaml").read_text()
assert "valid_volumes: []" in cfg assert "valid_volumes: []" in cfg

View File

@@ -26,10 +26,18 @@ cache:
dir: /data/cache dir: /data/cache
container: container:
# Job containers join the dind daemon's own bridge. NOT the forge network: # Empty = act creates a NETWORK PER JOB and removes it afterwards.
# untrusted code must never be able to reach the forge's Postgres or its #
# environment (I-5). # This started as `bridge` for isolation, which was a mistake in both
network: bridge # directions. It broke service containers — Postgres came up healthy but the
# job could not resolve the name `postgres`, because service DNS aliases only
# exist on a per-job network — and it was *weaker* isolation, since every
# concurrent job shared one flat bridge and could see its neighbours.
#
# A per-job network is both correct and stricter. Still no route to the forge:
# these networks live inside the dind daemon, which has no forge attachment
# at all.
network: ""
privileged: false privileged: false
options: options:
workdir_parent: /workspace workdir_parent: /workspace