Files
windy-git/deploy/runner/docker-compose.yml
Grant Whitmer dcf9286f16
All checks were successful
check / gate (push) Successful in 21s
canary / probe (push) Successful in 6s
ci: make Windy Git CI permanent for the private repos
- Four runners x capacity 1 instead of one x capacity 4. Concurrent jobs in
  one act_runner share /root/.cache/act; a refresh racing a copy killed 3 of
  windy-chat's ~20 jobs at setup-node (lstat ... no such file). Separate
  processes have separate caches. Same parallelism, same capped dind.
- Behavioral tests for pr_status_bridge (latest verdict wins, no reposting,
  skipped never painted green, fork PRs never run, pagination, PR lifecycle).
- import_from_github.py reads IMPORT_GITEA_URL, not GITEA_BASE_URL: sourcing
  the deploy .env pointed it at http://gitea:3000 and it died on DNS after the
  mirror it replaces had already been deleted.
- CUTOVER.md: the private-repo CI path, onboarding steps, and the
  /actions/tasks-hides-queued-runs trap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:18:04 -04:00

150 lines
6.8 KiB
YAML
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# CI runners (strand G7) — a SEPARATE compose project from the forge.
#
# Separate on purpose: runners restart, crash, get starved and get killed. None
# of that should ever touch the thing serving repositories. This is the cell
# doctrine applied one level down.
#
# ── I-5, and why there is a dind sidecar ───────────────────────────────────
#
# "CI never shares a kernel with identity. Runners execute untrusted code and
# are isolated by machine boundary, not container boundary. No runner may hold
# a credential scoped beyond its own job."
#
# act_runner needs a Docker daemon to start job containers. The tempting move is
# to mount the host's `/var/run/docker.sock`. That would hand every workflow —
# including whatever a transitive dependency's postinstall script feels like
# doing — the ability to start a privileged container mounting `/`, which is
# root on Veron 1. Every published act_runner example does exactly this.
#
# Instead the runner talks to its OWN daemon (`dind`). Untrusted job code runs
# as a child of that daemon, on an isolated network, with no route to the host
# socket and no route to the forge's database.
#
# The split that makes this work:
# * `runner` is TRUSTED code (the act_runner daemon). It sits on the forge
# network only so it can reach gitea:3000 to collect jobs.
# * `dind` and every job container it spawns are UNTRUSTED. They are on a
# private network with no access to the forge, its database, or its .env.
#
# dind itself is privileged — that is the cost, and it is the reason a job
# escape lands in a disposable daemon rather than on Grant's workstation.
#
# ⚠️ Do NOT "simplify" this by mounting the host docker socket.
name: windy-git-runner
services:
dind:
image: docker.io/library/docker:27-dind
privileged: true
environment:
DOCKER_TLS_CERTDIR: "" # plain TCP on an isolated network, no host route
command: ["dockerd", "--host=tcp://0.0.0.0:2375", "--tls=false"]
networks: [jobs]
volumes:
- dind-storage:/var/lib/docker
# G1.5 — bounded so a fork-bomb workflow cannot starve Grant's interactive
# session. Veron 1 is his workstation, not a dedicated build box.
cpus: 12.0 # 12 of 24 cores
mem_limit: 64g
restart: unless-stopped
# ── FOUR runners × capacity 1, not one runner × capacity 4 (2026-09-23) ──
#
# act caches every action repo at /root/.cache/act/<hash> INSIDE the runner
# process and re-fetches it at the start of each job. With capacity 4, four
# concurrent jobs share that one directory: one job's refresh rewrites it while
# another is tarring it into its job container, and the job dies with
# `lstat /root/.cache/act/<hash>/…: no such file or directory` on
# `actions/setup-node` / `setup-uv` — a failure that reads like a broken
# workflow. windy-chat (~20 jobs per push) hit it on 3 jobs in its first run.
# `rm -rf /root/.cache/act` only reset the clock. Separate processes get
# separate caches, so the race cannot occur. Same total parallelism, same
# single capped dind — the blast radius is unchanged.
runner: &runner
# 0.2.11 -> 0.6.1 on 2026-08-14. The bundled act in 0.2.11 only knows
# `runs.using: node12|node16|node20`, so ANY repo pinning a current action
# major dies before its first step with "The runs.using key in action.yml
# must be one of: [...], got node24" — Windy-Clone on actions/checkout@v5
# is how this surfaced. Verified: `node24` is absent from the 0.2.11 binary
# and present in 0.6.1. Every key in this directory's config.yaml still
# exists in 0.6.1's schema (0.6.1 only ADDS keys), so the config carries
# over unchanged. Rollback is re-pinning 0.2.11; the registration in the
# runner-data volume survives either way.
image: docker.io/gitea/act_runner:0.6.1
depends_on: [dind]
environment:
# The runner reaches its OWN daemon. Never the host's.
DOCKER_HOST: tcp://dind:2375
# ⚠️ THE PUBLIC URL, deliberately — not http://gitea:3000.
#
# Job containers run inside the dind daemon's own private network, so they
# cannot resolve `gitea`, which lives on the forge network. The first CI
# run failed exactly here: "Could not resolve host: gitea".
#
# There were two ways out, and they are not equivalent:
# (a) put job containers on the forge network — untrusted workflow code
# would then sit one DNS name away from the forge's Postgres. This
# is the easy fix and it quietly repeals I-5.
# (b) send jobs to the PUBLIC forge surface, over the tunnel, exactly
# like any stranger on the internet. Untrusted code gets no private
# network route at all.
#
# (b) is strictly better and it is what this is. The cost is a hairpin —
# container -> tunnel -> Cloudflare -> back to this box — plus Cloudflare's
# ~100s proxy ceiling on any single fetch (G4A.5). For repos measured at
# 0.63 GB of objects across 61 repos, with depth=1 checkouts, that ceiling
# is nowhere near being a problem. Revisit if a model repo ever needs CI.
GITEA_INSTANCE_URL: https://app.windygit.com
GITEA_RUNNER_REGISTRATION_TOKEN: ${RUNNER_TOKEN:?set RUNNER_TOKEN}
GITEA_RUNNER_NAME: veron-1
CONFIG_FILE: /config.yaml
volumes:
- ./config.yaml:/config.yaml:ro
- runner-data:/data
# Only `jobs`. The runner no longer needs the forge network at all, because
# it collects work over the public surface too — so there is now NO path
# from any CI container to the forge's database. That is a better posture
# than the one this file started with.
networks: [jobs]
cpus: 2.0
mem_limit: 4g
restart: unless-stopped
# Each extra runner registers itself on first start (own name, own volume —
# the registration lives in /data/.runner, so volumes must never be shared).
runner-2:
<<: *runner
environment: &env2
DOCKER_HOST: tcp://dind:2375
GITEA_INSTANCE_URL: https://app.windygit.com
GITEA_RUNNER_REGISTRATION_TOKEN: ${RUNNER_TOKEN:?set RUNNER_TOKEN}
GITEA_RUNNER_NAME: veron-1-2
CONFIG_FILE: /config.yaml
volumes: [./config.yaml:/config.yaml:ro, runner-data-2:/data]
runner-3:
<<: *runner
environment:
<<: *env2
GITEA_RUNNER_NAME: veron-1-3
volumes: [./config.yaml:/config.yaml:ro, runner-data-3:/data]
runner-4:
<<: *runner
environment:
<<: *env2
GITEA_RUNNER_NAME: veron-1-4
volumes: [./config.yaml:/config.yaml:ro, runner-data-4:/data]
networks:
jobs:
# Untrusted job containers live here. No route to the forge.
internal: false # jobs legitimately need to fetch dependencies
volumes:
dind-storage:
runner-data:
runner-data-2:
runner-data-3:
runner-data-4: