Commit Graph

12 Commits

Author SHA1 Message Date
Kit OC5
b438b6a053 ci: run dind under Sysbox, not privileged (rollback override kept)
dind was privileged: true, so a job that escaped into dind was root on
Veron 1, which is Grant's workstation. Under sysbox-runc (sysbox-ce 0.7.1,
installed 09-23 with no docker restart) dind root is an unprivileged host
uid. Smoke-tested standalone: nested containers, internet, a services-style
postgres on a private network and a python image all pass unprivileged.
Fresh volume dind-storage-sysbox; the old dind-storage stays for
docker-compose.privileged.yml, the one-command rollback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:41:42 -04:00
4a34b35441 security: CI egress filter — jobs reach the internet, never Veron/LAN
Measured: an unprivileged job container inside the CI dind could open SSH,
Ollama and dev servers on Veron (192.168.1.73) and the rest of the LAN,
WireGuard and Tailscale — lateral movement for any malicious dependency,
no escape needed. egress.sh (idempotent; windygit-ci-egress.service at
boot) hooks the jobs bridge: runner<->dind, replies, DNS and public
egress allowed; RFC1918, CGNAT, link-local and the host itself dropped.
Verified from a job container: 6/6 private targets blocked, DNS,
internet and the public forge OK.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 03:20:45 -04:00
5b16114b98 auth: token contract v1 (aud windy_git, both issuers); CI for eternitas
Some checks failed
check / gate (push) Successful in 37s
canary / probe (push) Has been cancelled
- hub_jwt: aud list is ["windy_git"] (contract v1 array). Dropped
  "windy-git": that is Gitea's OIDC client_id, so a forge id_token would
  have passed the aud check. `type: human` is now REQUIRED (id_tokens have
  none), which makes accepting the discovery-URL issuer safe.
- runner job ceiling 30m -> 90m: eternitas's serial pytest is ~50 min and
  would have been killed mid-suite.
- eternitas (private) added to the GitHub status bridge.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:58:56 -04:00
8e9fa3116c ci: six runners; SSO #8 Gitea sign-in hardening (staged)
- runner-5/6: 50+ jobs were queued with ~11 private repos onboarded. dind
  keeps the 12-core ceiling, so this adds concurrency, not CPU.
- Gitea: password + passkey sign-in forms off (break-glass = CLI), and
  ACCOUNT_LINKING auto -> login. auto linked any hub login whose email
  matched an existing account, and SITE ADMIN windyadmin carries Grant's
  email. Grant is linked by the hub's stable sub, which matches first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:51:56 -04:00
e7bbf9af51 ci: prune.sh must address dind over TCP (it has no unix socket)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:47:35 -04:00
45686283be ci: bound CI storage; don't bridge image-build jobs
- deploy/runner/prune.sh + windygit-ci-prune.timer (6h): age-based prune
  of the CI-only dind (containers, finished-job volumes, images/builder
  cache >7d) plus a hard 60 GB cap. Only that daemon, over its own TCP
  socket; never the host's Docker. It was 38 GB and unbounded — the same
  class of growth that filled Kit 0 on 09-01.
- pr_status_bridge: jobs named *docker* are not posted. Job containers
  have no daemon by design (I-5), so they are red on every commit; a
  permanent red X teaches everyone to ignore red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:47:28 -04:00
dcf9286f16 ci: make Windy Git CI permanent for the private repos
All checks were successful
check / gate (push) Successful in 21s
canary / probe (push) Successful in 6s
- Four runners x capacity 1 instead of one x capacity 4. Concurrent jobs in
  one act_runner share /root/.cache/act; a refresh racing a copy killed 3 of
  windy-chat's ~20 jobs at setup-node (lstat ... no such file). Separate
  processes have separate caches. Same parallelism, same capped dind.
- Behavioral tests for pr_status_bridge (latest verdict wins, no reposting,
  skipped never painted green, fork PRs never run, pagination, PR lifecycle).
- import_from_github.py reads IMPORT_GITEA_URL, not GITEA_BASE_URL: sourcing
  the deploy .env pointed it at http://gitea:3000 and it died on DNS after the
  mirror it replaces had already been deleted.
- CUTOVER.md: the private-repo CI path, onboarding steps, and the
  /actions/tasks-hides-queued-runs trap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:18:04 -04:00
cd7dd9b7ae ci: bump act_runner 0.2.11 -> 0.6.1 for node24 action support
0.2.11's bundled act only knows runs.using node12/node16/node20, so any repo
pinning a current action major (actions/checkout@v5, actions/setup-python@v6)
fails before its first step with "The runs.using key in action.yml must be one
of: [...], got node24". Windy-Clone is how this surfaced.

Verified node24 is absent from the 0.2.11 binary and present in 0.6.1, and that
every key in deploy/runner/config.yaml still exists in 0.6.1's schema.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 19:00:07 -04:00
Grant Whitmer
14fb43b959 G7.4: advertise the fleet's runner labels, and correct my own doctrine
Some checks are pending
check / gate (push) Waiting to run
Surveyed ten repos: 36 of 36 ACTIVE workflows already say
runs-on: [self-hosted, linux, x64]. They were written for the self-hosted
runners that died when the repos went private — so advertising those three
labels makes every one of them runnable AS-IS. No workflow edits, no rewrites.

Correcting an earlier note in this plan: 'ban ubuntu-latest' was wrong as
stated. All 11 occurrences are tagged '# runner-lint-allow — CD/hosted-only;
disabled, manual until CD mission'. They are deliberately hosted-only and
deliberately off. They are correct as written and should not be 'fixed'.

I generalised that doctrine from one repo without surveying. The survey
disagreed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 19:53:44 -04:00
Grant Whitmer
6b96087554 G7: per-job networks — fixes service DNS and tightens isolation
Some checks failed
check / gate (push) Failing after 13m33s
The migration step failed with 'could not translate host name postgres'. The
Postgres service container was healthy; the job simply could not name it,
because service DNS aliases only exist on a per-job network and I had pinned
container.network to the flat 'bridge'.

That choice was wrong in both directions: it broke service containers AND it
was weaker isolation, since every concurrent job shared one bridge and could
see its neighbours. A per-job network is stricter and correct — and still has
no route to the forge, because these networks live inside the dind daemon,
which has no forge attachment at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:18:51 -04:00
Grant Whitmer
53dea84ab1 G7: route CI jobs via the public surface, not the forge network
Some checks failed
check / gate (push) Failing after 13s
The first CI run failed with 'Could not resolve host: gitea' — job containers
live on dind's private network and cannot see the forge network. Two ways out,
and they are not equivalent:

  (a) put job containers on the forge network. Easy, one line, and it leaves
      untrusted workflow code one DNS name from the forge's Postgres. It quietly
      repeals I-5.
  (b) send jobs to the PUBLIC forge surface over the tunnel, exactly like any
      stranger on the internet.

Took (b). The runner no longer needs the forge network at all, so there is now
NO private route from any CI container to anything — a better posture than this
file started with. Cost is a hairpin through Cloudflare plus its ~100s ceiling
per fetch, which for 0.63 GB of objects across 61 repos and depth=1 checkouts is
nowhere near binding.

Test upgraded to assert the stronger property: no CI container joins the forge
network.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 10:51:00 -04:00
Grant Whitmer
5f717ef74b G7: CI runners with real isolation, and the gate as a workflow
Some checks failed
check / gate (push) Failing after 50s
I-5 says runners execute untrusted code and must be isolated by machine
boundary. act_runner needs a Docker daemon to start job containers, and the
tempting move — what every published example does — is to mount the host's
/var/run/docker.sock. That hands every workflow, including whatever a
transitive dependency's postinstall script feels like doing, the ability to
start a privileged container mounting / — root on Grant's workstation.

Instead the runner talks to its OWN dind daemon:
  - runner (TRUSTED, the act_runner daemon) sits on the forge network only to
    collect jobs from gitea:3000
  - dind and every job container it spawns are UNTRUSTED, on a private network
    with no route to the forge, its Postgres, or its .env
  - jobs cannot bind-mount from the daemon host (valid_volumes: []) and are not
    handed the runner's own socket (docker_host: -)
  - separate compose project, cpu/memory bounded — Veron 1 is Grant's
    workstation, not a dedicated build box

The gate itself now runs as a workflow, including the migration round-trip that
already caught two bugs review did not, and the I-12 check that a COMMIT_SHA
env override cannot change what /version reports.

Labels are explicit and pinned. A workflow naming a label nobody provides
queues forever and presents as a hung CI system rather than a typo — which is
what ubuntu-latest does on every windy-registry run today.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 10:39:05 -04:00