0.2.11's bundled act only knows runs.using node12/node16/node20, so any repo pinning a current action major (actions/checkout@v5, actions/setup-python@v6) fails before its first step with "The runs.using key in action.yml must be one of: [...], got node24". Windy-Clone is how this surfaced. Verified node24 is absent from the 0.2.11 binary and present in 0.6.1, and that every key in deploy/runner/config.yaml still exists in 0.6.1's schema. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
111 lines
5.1 KiB
YAML
111 lines
5.1 KiB
YAML
# CI runners (strand G7) — a SEPARATE compose project from the forge.
|
|
#
|
|
# Separate on purpose: runners restart, crash, get starved and get killed. None
|
|
# of that should ever touch the thing serving repositories. This is the cell
|
|
# doctrine applied one level down.
|
|
#
|
|
# ── I-5, and why there is a dind sidecar ───────────────────────────────────
|
|
#
|
|
# "CI never shares a kernel with identity. Runners execute untrusted code and
|
|
# are isolated by machine boundary, not container boundary. No runner may hold
|
|
# a credential scoped beyond its own job."
|
|
#
|
|
# act_runner needs a Docker daemon to start job containers. The tempting move is
|
|
# to mount the host's `/var/run/docker.sock`. That would hand every workflow —
|
|
# including whatever a transitive dependency's postinstall script feels like
|
|
# doing — the ability to start a privileged container mounting `/`, which is
|
|
# root on Veron 1. Every published act_runner example does exactly this.
|
|
#
|
|
# Instead the runner talks to its OWN daemon (`dind`). Untrusted job code runs
|
|
# as a child of that daemon, on an isolated network, with no route to the host
|
|
# socket and no route to the forge's database.
|
|
#
|
|
# The split that makes this work:
|
|
# * `runner` is TRUSTED code (the act_runner daemon). It sits on the forge
|
|
# network only so it can reach gitea:3000 to collect jobs.
|
|
# * `dind` and every job container it spawns are UNTRUSTED. They are on a
|
|
# private network with no access to the forge, its database, or its .env.
|
|
#
|
|
# dind itself is privileged — that is the cost, and it is the reason a job
|
|
# escape lands in a disposable daemon rather than on Grant's workstation.
|
|
#
|
|
# ⚠️ Do NOT "simplify" this by mounting the host docker socket.
|
|
|
|
name: windy-git-runner
|
|
|
|
services:
|
|
dind:
|
|
image: docker.io/library/docker:27-dind
|
|
privileged: true
|
|
environment:
|
|
DOCKER_TLS_CERTDIR: "" # plain TCP on an isolated network, no host route
|
|
command: ["dockerd", "--host=tcp://0.0.0.0:2375", "--tls=false"]
|
|
networks: [jobs]
|
|
volumes:
|
|
- dind-storage:/var/lib/docker
|
|
# G1.5 — bounded so a fork-bomb workflow cannot starve Grant's interactive
|
|
# session. Veron 1 is his workstation, not a dedicated build box.
|
|
cpus: 12.0 # 12 of 24 cores
|
|
mem_limit: 64g
|
|
restart: unless-stopped
|
|
|
|
runner:
|
|
# 0.2.11 -> 0.6.1 on 2026-08-14. The bundled act in 0.2.11 only knows
|
|
# `runs.using: node12|node16|node20`, so ANY repo pinning a current action
|
|
# major dies before its first step with "The runs.using key in action.yml
|
|
# must be one of: [...], got node24" — Windy-Clone on actions/checkout@v5
|
|
# is how this surfaced. Verified: `node24` is absent from the 0.2.11 binary
|
|
# and present in 0.6.1. Every key in this directory's config.yaml still
|
|
# exists in 0.6.1's schema (0.6.1 only ADDS keys), so the config carries
|
|
# over unchanged. Rollback is re-pinning 0.2.11; the registration in the
|
|
# runner-data volume survives either way.
|
|
image: docker.io/gitea/act_runner:0.6.1
|
|
depends_on: [dind]
|
|
environment:
|
|
# The runner reaches its OWN daemon. Never the host's.
|
|
DOCKER_HOST: tcp://dind:2375
|
|
# ⚠️ THE PUBLIC URL, deliberately — not http://gitea:3000.
|
|
#
|
|
# Job containers run inside the dind daemon's own private network, so they
|
|
# cannot resolve `gitea`, which lives on the forge network. The first CI
|
|
# run failed exactly here: "Could not resolve host: gitea".
|
|
#
|
|
# There were two ways out, and they are not equivalent:
|
|
# (a) put job containers on the forge network — untrusted workflow code
|
|
# would then sit one DNS name away from the forge's Postgres. This
|
|
# is the easy fix and it quietly repeals I-5.
|
|
# (b) send jobs to the PUBLIC forge surface, over the tunnel, exactly
|
|
# like any stranger on the internet. Untrusted code gets no private
|
|
# network route at all.
|
|
#
|
|
# (b) is strictly better and it is what this is. The cost is a hairpin —
|
|
# container -> tunnel -> Cloudflare -> back to this box — plus Cloudflare's
|
|
# ~100s proxy ceiling on any single fetch (G4A.5). For repos measured at
|
|
# 0.63 GB of objects across 61 repos, with depth=1 checkouts, that ceiling
|
|
# is nowhere near being a problem. Revisit if a model repo ever needs CI.
|
|
GITEA_INSTANCE_URL: https://app.windygit.com
|
|
GITEA_RUNNER_REGISTRATION_TOKEN: ${RUNNER_TOKEN:?set RUNNER_TOKEN}
|
|
GITEA_RUNNER_NAME: veron-1
|
|
CONFIG_FILE: /config.yaml
|
|
volumes:
|
|
- ./config.yaml:/config.yaml:ro
|
|
- runner-data:/data
|
|
# Only `jobs`. The runner no longer needs the forge network at all, because
|
|
# it collects work over the public surface too — so there is now NO path
|
|
# from any CI container to the forge's database. That is a better posture
|
|
# than the one this file started with.
|
|
networks: [jobs]
|
|
cpus: 2.0
|
|
mem_limit: 4g
|
|
restart: unless-stopped
|
|
|
|
networks:
|
|
jobs:
|
|
# Untrusted job containers live here. No route to the forge.
|
|
internal: false # jobs legitimately need to fetch dependencies
|
|
|
|
volumes:
|
|
dind-storage:
|
|
runner-data:
|
|
|