Some checks failed
check / gate (push) Failing after 13s
The first CI run failed with 'Could not resolve host: gitea' — job containers
live on dind's private network and cannot see the forge network. Two ways out,
and they are not equivalent:
(a) put job containers on the forge network. Easy, one line, and it leaves
untrusted workflow code one DNS name from the forge's Postgres. It quietly
repeals I-5.
(b) send jobs to the PUBLIC forge surface over the tunnel, exactly like any
stranger on the internet.
Took (b). The runner no longer needs the forge network at all, so there is now
NO private route from any CI container to anything — a better posture than this
file started with. Cost is a hairpin through Cloudflare plus its ~100s ceiling
per fetch, which for 0.63 GB of objects across 61 repos and depth=1 checkouts is
nowhere near binding.
Test upgraded to assert the stronger property: no CI container joins the forge
network.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
102 lines
4.5 KiB
YAML
102 lines
4.5 KiB
YAML
# CI runners (strand G7) — a SEPARATE compose project from the forge.
|
|
#
|
|
# Separate on purpose: runners restart, crash, get starved and get killed. None
|
|
# of that should ever touch the thing serving repositories. This is the cell
|
|
# doctrine applied one level down.
|
|
#
|
|
# ── I-5, and why there is a dind sidecar ───────────────────────────────────
|
|
#
|
|
# "CI never shares a kernel with identity. Runners execute untrusted code and
|
|
# are isolated by machine boundary, not container boundary. No runner may hold
|
|
# a credential scoped beyond its own job."
|
|
#
|
|
# act_runner needs a Docker daemon to start job containers. The tempting move is
|
|
# to mount the host's `/var/run/docker.sock`. That would hand every workflow —
|
|
# including whatever a transitive dependency's postinstall script feels like
|
|
# doing — the ability to start a privileged container mounting `/`, which is
|
|
# root on Veron 1. Every published act_runner example does exactly this.
|
|
#
|
|
# Instead the runner talks to its OWN daemon (`dind`). Untrusted job code runs
|
|
# as a child of that daemon, on an isolated network, with no route to the host
|
|
# socket and no route to the forge's database.
|
|
#
|
|
# The split that makes this work:
|
|
# * `runner` is TRUSTED code (the act_runner daemon). It sits on the forge
|
|
# network only so it can reach gitea:3000 to collect jobs.
|
|
# * `dind` and every job container it spawns are UNTRUSTED. They are on a
|
|
# private network with no access to the forge, its database, or its .env.
|
|
#
|
|
# dind itself is privileged — that is the cost, and it is the reason a job
|
|
# escape lands in a disposable daemon rather than on Grant's workstation.
|
|
#
|
|
# ⚠️ Do NOT "simplify" this by mounting the host docker socket.
|
|
|
|
name: windy-git-runner
|
|
|
|
services:
|
|
dind:
|
|
image: docker.io/library/docker:27-dind
|
|
privileged: true
|
|
environment:
|
|
DOCKER_TLS_CERTDIR: "" # plain TCP on an isolated network, no host route
|
|
command: ["dockerd", "--host=tcp://0.0.0.0:2375", "--tls=false"]
|
|
networks: [jobs]
|
|
volumes:
|
|
- dind-storage:/var/lib/docker
|
|
# G1.5 — bounded so a fork-bomb workflow cannot starve Grant's interactive
|
|
# session. Veron 1 is his workstation, not a dedicated build box.
|
|
cpus: 12.0 # 12 of 24 cores
|
|
mem_limit: 64g
|
|
restart: unless-stopped
|
|
|
|
runner:
|
|
image: docker.io/gitea/act_runner:0.2.11
|
|
depends_on: [dind]
|
|
environment:
|
|
# The runner reaches its OWN daemon. Never the host's.
|
|
DOCKER_HOST: tcp://dind:2375
|
|
# ⚠️ THE PUBLIC URL, deliberately — not http://gitea:3000.
|
|
#
|
|
# Job containers run inside the dind daemon's own private network, so they
|
|
# cannot resolve `gitea`, which lives on the forge network. The first CI
|
|
# run failed exactly here: "Could not resolve host: gitea".
|
|
#
|
|
# There were two ways out, and they are not equivalent:
|
|
# (a) put job containers on the forge network — untrusted workflow code
|
|
# would then sit one DNS name away from the forge's Postgres. This
|
|
# is the easy fix and it quietly repeals I-5.
|
|
# (b) send jobs to the PUBLIC forge surface, over the tunnel, exactly
|
|
# like any stranger on the internet. Untrusted code gets no private
|
|
# network route at all.
|
|
#
|
|
# (b) is strictly better and it is what this is. The cost is a hairpin —
|
|
# container -> tunnel -> Cloudflare -> back to this box — plus Cloudflare's
|
|
# ~100s proxy ceiling on any single fetch (G4A.5). For repos measured at
|
|
# 0.63 GB of objects across 61 repos, with depth=1 checkouts, that ceiling
|
|
# is nowhere near being a problem. Revisit if a model repo ever needs CI.
|
|
GITEA_INSTANCE_URL: https://app.windygit.com
|
|
GITEA_RUNNER_REGISTRATION_TOKEN: ${RUNNER_TOKEN:?set RUNNER_TOKEN}
|
|
GITEA_RUNNER_NAME: veron-1
|
|
CONFIG_FILE: /config.yaml
|
|
volumes:
|
|
- ./config.yaml:/config.yaml:ro
|
|
- runner-data:/data
|
|
# Only `jobs`. The runner no longer needs the forge network at all, because
|
|
# it collects work over the public surface too — so there is now NO path
|
|
# from any CI container to the forge's database. That is a better posture
|
|
# than the one this file started with.
|
|
networks: [jobs]
|
|
cpus: 2.0
|
|
mem_limit: 4g
|
|
restart: unless-stopped
|
|
|
|
networks:
|
|
jobs:
|
|
# Untrusted job containers live here. No route to the forge.
|
|
internal: false # jobs legitimately need to fetch dependencies
|
|
|
|
volumes:
|
|
dind-storage:
|
|
runner-data:
|
|
|