Files
windy-git/deploy/runner/docker-compose.yml
Kit OC5 b438b6a053 ci: run dind under Sysbox, not privileged (rollback override kept)
dind was privileged: true, so a job that escaped into dind was root on
Veron 1, which is Grant's workstation. Under sysbox-runc (sysbox-ce 0.7.1,
installed 09-23 with no docker restart) dind root is an unprivileged host
uid. Smoke-tested standalone: nested containers, internet, a services-style
postgres on a private network and a python image all pass unprivileged.
Fresh volume dind-storage-sysbox; the old dind-storage stays for
docker-compose.privileged.yml, the one-command rollback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:41:42 -04:00

176 lines
8.0 KiB
YAML
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# CI runners (strand G7) — a SEPARATE compose project from the forge.
#
# Separate on purpose: runners restart, crash, get starved and get killed. None
# of that should ever touch the thing serving repositories. This is the cell
# doctrine applied one level down.
#
# ── I-5, and why there is a dind sidecar ───────────────────────────────────
#
# "CI never shares a kernel with identity. Runners execute untrusted code and
# are isolated by machine boundary, not container boundary. No runner may hold
# a credential scoped beyond its own job."
#
# act_runner needs a Docker daemon to start job containers. The tempting move is
# to mount the host's `/var/run/docker.sock`. That would hand every workflow —
# including whatever a transitive dependency's postinstall script feels like
# doing — the ability to start a privileged container mounting `/`, which is
# root on Veron 1. Every published act_runner example does exactly this.
#
# Instead the runner talks to its OWN daemon (`dind`). Untrusted job code runs
# as a child of that daemon, on an isolated network, with no route to the host
# socket and no route to the forge's database.
#
# The split that makes this work:
# * `runner` is TRUSTED code (the act_runner daemon). It sits on the forge
# network only so it can reach gitea:3000 to collect jobs.
# * `dind` and every job container it spawns are UNTRUSTED. They are on a
# private network with no access to the forge, its database, or its .env.
#
# dind is NOT privileged (2026-09-23): it runs under the Sysbox runtime
# (sysbox-ce on Veron, `runtime: sysbox-runc`), a user-namespaced system
# container whose root is an unprivileged host uid. A job that escapes its own
# container lands in dind as a nobody on the host, not as root on Grant's
# workstation. Before Sysbox, dind was `privileged: true`; that config is kept
# as docker-compose.privileged.yml (ROLLBACK ONLY, one command, see that file).
#
# ⚠️ Do NOT "simplify" this by mounting the host docker socket.
name: windy-git-runner
services:
dind:
image: docker.io/library/docker:27-dind
runtime: sysbox-runc # NOT privileged: see the I-5 note above
environment:
DOCKER_TLS_CERTDIR: "" # plain TCP on an isolated network, no host route
command: ["dockerd", "--host=tcp://0.0.0.0:2375", "--tls=false"]
networks: [jobs]
volumes:
# A fresh volume: Sysbox shifts ownership to its own uid range. The old
# `dind-storage` is kept untouched for the privileged rollback.
- dind-storage-sysbox:/var/lib/docker
# G1.5 — bounded so a fork-bomb workflow cannot starve Grant's interactive
# session. Veron 1 is his workstation, not a dedicated build box.
cpus: 12.0 # 12 of 24 cores
mem_limit: 64g
restart: unless-stopped
# ── FOUR runners × capacity 1, not one runner × capacity 4 (2026-09-23) ──
#
# act caches every action repo at /root/.cache/act/<hash> INSIDE the runner
# process and re-fetches it at the start of each job. With capacity 4, four
# concurrent jobs share that one directory: one job's refresh rewrites it while
# another is tarring it into its job container, and the job dies with
# `lstat /root/.cache/act/<hash>/…: no such file or directory` on
# `actions/setup-node` / `setup-uv` — a failure that reads like a broken
# workflow. windy-chat (~20 jobs per push) hit it on 3 jobs in its first run.
# `rm -rf /root/.cache/act` only reset the clock. Separate processes get
# separate caches, so the race cannot occur. Same total parallelism, same
# single capped dind — the blast radius is unchanged.
runner: &runner
# 0.2.11 -> 0.6.1 on 2026-08-14. The bundled act in 0.2.11 only knows
# `runs.using: node12|node16|node20`, so ANY repo pinning a current action
# major dies before its first step with "The runs.using key in action.yml
# must be one of: [...], got node24" — Windy-Clone on actions/checkout@v5
# is how this surfaced. Verified: `node24` is absent from the 0.2.11 binary
# and present in 0.6.1. Every key in this directory's config.yaml still
# exists in 0.6.1's schema (0.6.1 only ADDS keys), so the config carries
# over unchanged. Rollback is re-pinning 0.2.11; the registration in the
# runner-data volume survives either way.
image: docker.io/gitea/act_runner:0.6.1
depends_on: [dind]
environment:
# The runner reaches its OWN daemon. Never the host's.
DOCKER_HOST: tcp://dind:2375
# ⚠️ THE PUBLIC URL, deliberately — not http://gitea:3000.
#
# Job containers run inside the dind daemon's own private network, so they
# cannot resolve `gitea`, which lives on the forge network. The first CI
# run failed exactly here: "Could not resolve host: gitea".
#
# There were two ways out, and they are not equivalent:
# (a) put job containers on the forge network — untrusted workflow code
# would then sit one DNS name away from the forge's Postgres. This
# is the easy fix and it quietly repeals I-5.
# (b) send jobs to the PUBLIC forge surface, over the tunnel, exactly
# like any stranger on the internet. Untrusted code gets no private
# network route at all.
#
# (b) is strictly better and it is what this is. The cost is a hairpin —
# container -> tunnel -> Cloudflare -> back to this box — plus Cloudflare's
# ~100s proxy ceiling on any single fetch (G4A.5). For repos measured at
# 0.63 GB of objects across 61 repos, with depth=1 checkouts, that ceiling
# is nowhere near being a problem. Revisit if a model repo ever needs CI.
GITEA_INSTANCE_URL: https://app.windygit.com
GITEA_RUNNER_REGISTRATION_TOKEN: ${RUNNER_TOKEN:?set RUNNER_TOKEN}
GITEA_RUNNER_NAME: veron-1
CONFIG_FILE: /config.yaml
volumes:
- ./config.yaml:/config.yaml:ro
- runner-data:/data
# Only `jobs`. The runner no longer needs the forge network at all, because
# it collects work over the public surface too — so there is now NO path
# from any CI container to the forge's database. That is a better posture
# than the one this file started with.
networks: [jobs]
cpus: 2.0
mem_limit: 4g
restart: unless-stopped
# Each extra runner registers itself on first start (own name, own volume —
# the registration lives in /data/.runner, so volumes must never be shared).
runner-2:
<<: *runner
environment: &env2
DOCKER_HOST: tcp://dind:2375
GITEA_INSTANCE_URL: https://app.windygit.com
GITEA_RUNNER_REGISTRATION_TOKEN: ${RUNNER_TOKEN:?set RUNNER_TOKEN}
GITEA_RUNNER_NAME: veron-1-2
CONFIG_FILE: /config.yaml
volumes: [./config.yaml:/config.yaml:ro, runner-data-2:/data]
runner-3:
<<: *runner
environment:
<<: *env2
GITEA_RUNNER_NAME: veron-1-3
volumes: [./config.yaml:/config.yaml:ro, runner-data-3:/data]
runner-4:
<<: *runner
environment:
<<: *env2
GITEA_RUNNER_NAME: veron-1-4
volumes: [./config.yaml:/config.yaml:ro, runner-data-4:/data]
# 5 and 6 added the same day: with ~11 private repos onboarded (windy-chat
# alone queues ~24 jobs per push) four runners left 50+ jobs waiting. The
# CPU ceiling is dind's (12 of 24 cores, G1.5), not the runner count, so more
# runners add concurrency for I/O-bound jobs (npm ci, uv sync) without
# taking more of Grant's workstation.
runner-5:
<<: *runner
environment:
<<: *env2
GITEA_RUNNER_NAME: veron-1-5
volumes: [./config.yaml:/config.yaml:ro, runner-data-5:/data]
runner-6:
<<: *runner
environment:
<<: *env2
GITEA_RUNNER_NAME: veron-1-6
volumes: [./config.yaml:/config.yaml:ro, runner-data-6:/data]
networks:
jobs:
# Untrusted job containers live here. No route to the forge.
internal: false # jobs legitimately need to fetch dependencies
volumes:
dind-storage:
dind-storage-sysbox:
runner-data:
runner-data-2:
runner-data-3:
runner-data-4:
runner-data-5:
runner-data-6: