Files
windy-git/deploy/runner/docker-compose.yml
Grant Whitmer cd7dd9b7ae ci: bump act_runner 0.2.11 -> 0.6.1 for node24 action support
0.2.11's bundled act only knows runs.using node12/node16/node20, so any repo
pinning a current action major (actions/checkout@v5, actions/setup-python@v6)
fails before its first step with "The runs.using key in action.yml must be one
of: [...], got node24". Windy-Clone is how this surfaced.

Verified node24 is absent from the 0.2.11 binary and present in 0.6.1, and that
every key in deploy/runner/config.yaml still exists in 0.6.1's schema.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 19:00:07 -04:00

111 lines
5.1 KiB
YAML

# CI runners (strand G7) — a SEPARATE compose project from the forge.
#
# Separate on purpose: runners restart, crash, get starved and get killed. None
# of that should ever touch the thing serving repositories. This is the cell
# doctrine applied one level down.
#
# ── I-5, and why there is a dind sidecar ───────────────────────────────────
#
# "CI never shares a kernel with identity. Runners execute untrusted code and
# are isolated by machine boundary, not container boundary. No runner may hold
# a credential scoped beyond its own job."
#
# act_runner needs a Docker daemon to start job containers. The tempting move is
# to mount the host's `/var/run/docker.sock`. That would hand every workflow —
# including whatever a transitive dependency's postinstall script feels like
# doing — the ability to start a privileged container mounting `/`, which is
# root on Veron 1. Every published act_runner example does exactly this.
#
# Instead the runner talks to its OWN daemon (`dind`). Untrusted job code runs
# as a child of that daemon, on an isolated network, with no route to the host
# socket and no route to the forge's database.
#
# The split that makes this work:
# * `runner` is TRUSTED code (the act_runner daemon). It sits on the forge
# network only so it can reach gitea:3000 to collect jobs.
# * `dind` and every job container it spawns are UNTRUSTED. They are on a
# private network with no access to the forge, its database, or its .env.
#
# dind itself is privileged — that is the cost, and it is the reason a job
# escape lands in a disposable daemon rather than on Grant's workstation.
#
# ⚠️ Do NOT "simplify" this by mounting the host docker socket.
name: windy-git-runner
services:
dind:
image: docker.io/library/docker:27-dind
privileged: true
environment:
DOCKER_TLS_CERTDIR: "" # plain TCP on an isolated network, no host route
command: ["dockerd", "--host=tcp://0.0.0.0:2375", "--tls=false"]
networks: [jobs]
volumes:
- dind-storage:/var/lib/docker
# G1.5 — bounded so a fork-bomb workflow cannot starve Grant's interactive
# session. Veron 1 is his workstation, not a dedicated build box.
cpus: 12.0 # 12 of 24 cores
mem_limit: 64g
restart: unless-stopped
runner:
# 0.2.11 -> 0.6.1 on 2026-08-14. The bundled act in 0.2.11 only knows
# `runs.using: node12|node16|node20`, so ANY repo pinning a current action
# major dies before its first step with "The runs.using key in action.yml
# must be one of: [...], got node24" — Windy-Clone on actions/checkout@v5
# is how this surfaced. Verified: `node24` is absent from the 0.2.11 binary
# and present in 0.6.1. Every key in this directory's config.yaml still
# exists in 0.6.1's schema (0.6.1 only ADDS keys), so the config carries
# over unchanged. Rollback is re-pinning 0.2.11; the registration in the
# runner-data volume survives either way.
image: docker.io/gitea/act_runner:0.6.1
depends_on: [dind]
environment:
# The runner reaches its OWN daemon. Never the host's.
DOCKER_HOST: tcp://dind:2375
# ⚠️ THE PUBLIC URL, deliberately — not http://gitea:3000.
#
# Job containers run inside the dind daemon's own private network, so they
# cannot resolve `gitea`, which lives on the forge network. The first CI
# run failed exactly here: "Could not resolve host: gitea".
#
# There were two ways out, and they are not equivalent:
# (a) put job containers on the forge network — untrusted workflow code
# would then sit one DNS name away from the forge's Postgres. This
# is the easy fix and it quietly repeals I-5.
# (b) send jobs to the PUBLIC forge surface, over the tunnel, exactly
# like any stranger on the internet. Untrusted code gets no private
# network route at all.
#
# (b) is strictly better and it is what this is. The cost is a hairpin —
# container -> tunnel -> Cloudflare -> back to this box — plus Cloudflare's
# ~100s proxy ceiling on any single fetch (G4A.5). For repos measured at
# 0.63 GB of objects across 61 repos, with depth=1 checkouts, that ceiling
# is nowhere near being a problem. Revisit if a model repo ever needs CI.
GITEA_INSTANCE_URL: https://app.windygit.com
GITEA_RUNNER_REGISTRATION_TOKEN: ${RUNNER_TOKEN:?set RUNNER_TOKEN}
GITEA_RUNNER_NAME: veron-1
CONFIG_FILE: /config.yaml
volumes:
- ./config.yaml:/config.yaml:ro
- runner-data:/data
# Only `jobs`. The runner no longer needs the forge network at all, because
# it collects work over the public surface too — so there is now NO path
# from any CI container to the forge's database. That is a better posture
# than the one this file started with.
networks: [jobs]
cpus: 2.0
mem_limit: 4g
restart: unless-stopped
networks:
jobs:
# Untrusted job containers live here. No route to the forge.
internal: false # jobs legitimately need to fetch dependencies
volumes:
dind-storage:
runner-data: