- Four runners x capacity 1 instead of one x capacity 4. Concurrent jobs in one act_runner share /root/.cache/act; a refresh racing a copy killed 3 of windy-chat's ~20 jobs at setup-node (lstat ... no such file). Separate processes have separate caches. Same parallelism, same capped dind. - Behavioral tests for pr_status_bridge (latest verdict wins, no reposting, skipped never painted green, fork PRs never run, pagination, PR lifecycle). - import_from_github.py reads IMPORT_GITEA_URL, not GITEA_BASE_URL: sourcing the deploy .env pointed it at http://gitea:3000 and it died on DNS after the mirror it replaces had already been deleted. - CUTOVER.md: the private-repo CI path, onboarding steps, and the /actions/tasks-hides-queued-runs trap. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
150 lines
6.8 KiB
YAML
150 lines
6.8 KiB
YAML
# CI runners (strand G7) — a SEPARATE compose project from the forge.
|
||
#
|
||
# Separate on purpose: runners restart, crash, get starved and get killed. None
|
||
# of that should ever touch the thing serving repositories. This is the cell
|
||
# doctrine applied one level down.
|
||
#
|
||
# ── I-5, and why there is a dind sidecar ───────────────────────────────────
|
||
#
|
||
# "CI never shares a kernel with identity. Runners execute untrusted code and
|
||
# are isolated by machine boundary, not container boundary. No runner may hold
|
||
# a credential scoped beyond its own job."
|
||
#
|
||
# act_runner needs a Docker daemon to start job containers. The tempting move is
|
||
# to mount the host's `/var/run/docker.sock`. That would hand every workflow —
|
||
# including whatever a transitive dependency's postinstall script feels like
|
||
# doing — the ability to start a privileged container mounting `/`, which is
|
||
# root on Veron 1. Every published act_runner example does exactly this.
|
||
#
|
||
# Instead the runner talks to its OWN daemon (`dind`). Untrusted job code runs
|
||
# as a child of that daemon, on an isolated network, with no route to the host
|
||
# socket and no route to the forge's database.
|
||
#
|
||
# The split that makes this work:
|
||
# * `runner` is TRUSTED code (the act_runner daemon). It sits on the forge
|
||
# network only so it can reach gitea:3000 to collect jobs.
|
||
# * `dind` and every job container it spawns are UNTRUSTED. They are on a
|
||
# private network with no access to the forge, its database, or its .env.
|
||
#
|
||
# dind itself is privileged — that is the cost, and it is the reason a job
|
||
# escape lands in a disposable daemon rather than on Grant's workstation.
|
||
#
|
||
# ⚠️ Do NOT "simplify" this by mounting the host docker socket.
|
||
|
||
name: windy-git-runner
|
||
|
||
services:
|
||
dind:
|
||
image: docker.io/library/docker:27-dind
|
||
privileged: true
|
||
environment:
|
||
DOCKER_TLS_CERTDIR: "" # plain TCP on an isolated network, no host route
|
||
command: ["dockerd", "--host=tcp://0.0.0.0:2375", "--tls=false"]
|
||
networks: [jobs]
|
||
volumes:
|
||
- dind-storage:/var/lib/docker
|
||
# G1.5 — bounded so a fork-bomb workflow cannot starve Grant's interactive
|
||
# session. Veron 1 is his workstation, not a dedicated build box.
|
||
cpus: 12.0 # 12 of 24 cores
|
||
mem_limit: 64g
|
||
restart: unless-stopped
|
||
|
||
# ── FOUR runners × capacity 1, not one runner × capacity 4 (2026-09-23) ──
|
||
#
|
||
# act caches every action repo at /root/.cache/act/<hash> INSIDE the runner
|
||
# process and re-fetches it at the start of each job. With capacity 4, four
|
||
# concurrent jobs share that one directory: one job's refresh rewrites it while
|
||
# another is tarring it into its job container, and the job dies with
|
||
# `lstat /root/.cache/act/<hash>/…: no such file or directory` on
|
||
# `actions/setup-node` / `setup-uv` — a failure that reads like a broken
|
||
# workflow. windy-chat (~20 jobs per push) hit it on 3 jobs in its first run.
|
||
# `rm -rf /root/.cache/act` only reset the clock. Separate processes get
|
||
# separate caches, so the race cannot occur. Same total parallelism, same
|
||
# single capped dind — the blast radius is unchanged.
|
||
runner: &runner
|
||
# 0.2.11 -> 0.6.1 on 2026-08-14. The bundled act in 0.2.11 only knows
|
||
# `runs.using: node12|node16|node20`, so ANY repo pinning a current action
|
||
# major dies before its first step with "The runs.using key in action.yml
|
||
# must be one of: [...], got node24" — Windy-Clone on actions/checkout@v5
|
||
# is how this surfaced. Verified: `node24` is absent from the 0.2.11 binary
|
||
# and present in 0.6.1. Every key in this directory's config.yaml still
|
||
# exists in 0.6.1's schema (0.6.1 only ADDS keys), so the config carries
|
||
# over unchanged. Rollback is re-pinning 0.2.11; the registration in the
|
||
# runner-data volume survives either way.
|
||
image: docker.io/gitea/act_runner:0.6.1
|
||
depends_on: [dind]
|
||
environment:
|
||
# The runner reaches its OWN daemon. Never the host's.
|
||
DOCKER_HOST: tcp://dind:2375
|
||
# ⚠️ THE PUBLIC URL, deliberately — not http://gitea:3000.
|
||
#
|
||
# Job containers run inside the dind daemon's own private network, so they
|
||
# cannot resolve `gitea`, which lives on the forge network. The first CI
|
||
# run failed exactly here: "Could not resolve host: gitea".
|
||
#
|
||
# There were two ways out, and they are not equivalent:
|
||
# (a) put job containers on the forge network — untrusted workflow code
|
||
# would then sit one DNS name away from the forge's Postgres. This
|
||
# is the easy fix and it quietly repeals I-5.
|
||
# (b) send jobs to the PUBLIC forge surface, over the tunnel, exactly
|
||
# like any stranger on the internet. Untrusted code gets no private
|
||
# network route at all.
|
||
#
|
||
# (b) is strictly better and it is what this is. The cost is a hairpin —
|
||
# container -> tunnel -> Cloudflare -> back to this box — plus Cloudflare's
|
||
# ~100s proxy ceiling on any single fetch (G4A.5). For repos measured at
|
||
# 0.63 GB of objects across 61 repos, with depth=1 checkouts, that ceiling
|
||
# is nowhere near being a problem. Revisit if a model repo ever needs CI.
|
||
GITEA_INSTANCE_URL: https://app.windygit.com
|
||
GITEA_RUNNER_REGISTRATION_TOKEN: ${RUNNER_TOKEN:?set RUNNER_TOKEN}
|
||
GITEA_RUNNER_NAME: veron-1
|
||
CONFIG_FILE: /config.yaml
|
||
volumes:
|
||
- ./config.yaml:/config.yaml:ro
|
||
- runner-data:/data
|
||
# Only `jobs`. The runner no longer needs the forge network at all, because
|
||
# it collects work over the public surface too — so there is now NO path
|
||
# from any CI container to the forge's database. That is a better posture
|
||
# than the one this file started with.
|
||
networks: [jobs]
|
||
cpus: 2.0
|
||
mem_limit: 4g
|
||
restart: unless-stopped
|
||
|
||
# Each extra runner registers itself on first start (own name, own volume —
|
||
# the registration lives in /data/.runner, so volumes must never be shared).
|
||
runner-2:
|
||
<<: *runner
|
||
environment: &env2
|
||
DOCKER_HOST: tcp://dind:2375
|
||
GITEA_INSTANCE_URL: https://app.windygit.com
|
||
GITEA_RUNNER_REGISTRATION_TOKEN: ${RUNNER_TOKEN:?set RUNNER_TOKEN}
|
||
GITEA_RUNNER_NAME: veron-1-2
|
||
CONFIG_FILE: /config.yaml
|
||
volumes: [./config.yaml:/config.yaml:ro, runner-data-2:/data]
|
||
runner-3:
|
||
<<: *runner
|
||
environment:
|
||
<<: *env2
|
||
GITEA_RUNNER_NAME: veron-1-3
|
||
volumes: [./config.yaml:/config.yaml:ro, runner-data-3:/data]
|
||
runner-4:
|
||
<<: *runner
|
||
environment:
|
||
<<: *env2
|
||
GITEA_RUNNER_NAME: veron-1-4
|
||
volumes: [./config.yaml:/config.yaml:ro, runner-data-4:/data]
|
||
|
||
networks:
|
||
jobs:
|
||
# Untrusted job containers live here. No route to the forge.
|
||
internal: false # jobs legitimately need to fetch dependencies
|
||
|
||
volumes:
|
||
dind-storage:
|
||
runner-data:
|
||
runner-data-2:
|
||
runner-data-3:
|
||
runner-data-4:
|
||
|