Commit Graph

16 Commits

Author SHA1 Message Date
Kit OC5
a2daca61ef ci: mount windy-pro's read-only build inputs into dind; allow exactly that path for jobs
Non-secret inputs git-ignored in windy-pro (models, linux-x64 portable
bundle, enter-monitor build) that build-desktop needs. Mounted :ro into
dind; valid_volumes allows only /ci-inputs/windy-pro; refresh-ci-inputs.sh
copies them from the frozen release clone (read-only on the source).
Invariant I-5 narrowed, not dropped: exactly that one path, read-only in
dind, no other service mounts it, still no docker socket (proven to fail
on :rw). Orchestrator-approved (option a). Applied in an idle window.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 14:56:48 -04:00
Kit OC5
fdb0f5989e ci: run dind under Sysbox, not privileged (rollback override kept)
Some checks failed
canary / probe (push) Has been cancelled
check / gate (push) Has been cancelled
dind was privileged: true, so a job that escaped into dind was root on
Veron 1, which is Grant's workstation. Under sysbox-runc (sysbox-ce 0.7.1,
installed 09-23 with no docker restart) dind root is an unprivileged host
uid. Smoke-tested standalone: nested containers, internet, a services-style
postgres on a private network and a python image all pass unprivileged.
Fresh volume dind-storage-sysbox; the old dind-storage stays for
docker-compose.privileged.yml, the one-command rollback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 14:23:55 -04:00
95c33c8004 ops: host systemd units in git; windy-pro tags never reach Windy Git
All checks were successful
check / gate (push) Successful in 22s
- deploy/systemd/: sync/backup timers+services, tunnel, and the windy-job
  heartbeat drop-ins (silent-failure audit). They existed only on Veron,
  the same drift that left the runbook wrong. GITHUB_TOKEN is stripped
  (repo is public); it stays in the root-only unit on the host.
- sync: SYNC_NO_TAGS (default windy-pro). build-electron fires on v* tags
  and targets ubuntu/macos/windows-latest, labels no runner has, so it
  would queue forever, invisibly. Desktop releases are built elsewhere.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 09:51:51 -04:00
4a34b35441 security: CI egress filter — jobs reach the internet, never Veron/LAN
Measured: an unprivileged job container inside the CI dind could open SSH,
Ollama and dev servers on Veron (192.168.1.73) and the rest of the LAN,
WireGuard and Tailscale — lateral movement for any malicious dependency,
no escape needed. egress.sh (idempotent; windygit-ci-egress.service at
boot) hooks the jobs bridge: runner<->dind, replies, DNS and public
egress allowed; RFC1918, CGNAT, link-local and the host itself dropped.
Verified from a job container: 6/6 private targets blocked, DNS,
internet and the public forge OK.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 03:20:45 -04:00
5b16114b98 auth: token contract v1 (aud windy_git, both issuers); CI for eternitas
Some checks failed
check / gate (push) Successful in 37s
canary / probe (push) Has been cancelled
- hub_jwt: aud list is ["windy_git"] (contract v1 array). Dropped
  "windy-git": that is Gitea's OIDC client_id, so a forge id_token would
  have passed the aud check. `type: human` is now REQUIRED (id_tokens have
  none), which makes accepting the discovery-URL issuer safe.
- runner job ceiling 30m -> 90m: eternitas's serial pytest is ~50 min and
  would have been killed mid-suite.
- eternitas (private) added to the GitHub status bridge.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:58:56 -04:00
8e9fa3116c ci: six runners; SSO #8 Gitea sign-in hardening (staged)
- runner-5/6: 50+ jobs were queued with ~11 private repos onboarded. dind
  keeps the 12-core ceiling, so this adds concurrency, not CPU.
- Gitea: password + passkey sign-in forms off (break-glass = CLI), and
  ACCOUNT_LINKING auto -> login. auto linked any hub login whose email
  matched an existing account, and SITE ADMIN windyadmin carries Grant's
  email. Grant is linked by the hub's stable sub, which matches first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:51:56 -04:00
e7bbf9af51 ci: prune.sh must address dind over TCP (it has no unix socket)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:47:35 -04:00
45686283be ci: bound CI storage; don't bridge image-build jobs
- deploy/runner/prune.sh + windygit-ci-prune.timer (6h): age-based prune
  of the CI-only dind (containers, finished-job volumes, images/builder
  cache >7d) plus a hard 60 GB cap. Only that daemon, over its own TCP
  socket; never the host's Docker. It was 38 GB and unbounded — the same
  class of growth that filled Kit 0 on 09-01.
- pr_status_bridge: jobs named *docker* are not posted. Job containers
  have no daemon by design (I-5), so they are red on every commit; a
  permanent red X teaches everyone to ignore red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:47:28 -04:00
dcf9286f16 ci: make Windy Git CI permanent for the private repos
All checks were successful
check / gate (push) Successful in 21s
canary / probe (push) Successful in 6s
- Four runners x capacity 1 instead of one x capacity 4. Concurrent jobs in
  one act_runner share /root/.cache/act; a refresh racing a copy killed 3 of
  windy-chat's ~20 jobs at setup-node (lstat ... no such file). Separate
  processes have separate caches. Same parallelism, same capped dind.
- Behavioral tests for pr_status_bridge (latest verdict wins, no reposting,
  skipped never painted green, fork PRs never run, pagination, PR lifecycle).
- import_from_github.py reads IMPORT_GITEA_URL, not GITEA_BASE_URL: sourcing
  the deploy .env pointed it at http://gitea:3000 and it died on DNS after the
  mirror it replaces had already been deleted.
- CUTOVER.md: the private-repo CI path, onboarding steps, and the
  /actions/tasks-hides-queued-runs trap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:18:04 -04:00
cd7dd9b7ae ci: bump act_runner 0.2.11 -> 0.6.1 for node24 action support
0.2.11's bundled act only knows runs.using node12/node16/node20, so any repo
pinning a current action major (actions/checkout@v5, actions/setup-python@v6)
fails before its first step with "The runs.using key in action.yml must be one
of: [...], got node24". Windy-Clone is how this surfaced.

Verified node24 is absent from the 0.2.11 binary and present in 0.6.1, and that
every key in deploy/runner/config.yaml still exists in 0.6.1's schema.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 19:00:07 -04:00
Grant Whitmer
eae1bff50b G2.3: Windy Git branding — and get it out of one host's disk into the repo
All checks were successful
check / gate (push) Successful in 19s
canary / probe (push) Successful in 8s
The front end was 100% stock Gitea: green teacup, "Gitea: Git with a cup of
tea", "A painless, self-hosted Git service". G2.3 was specified in the plan with
an acceptance test and never executed, and nothing enforced it.

Now: Windy Git name, wind-mark logo, brand-blue accent, and a landing page that
says what this actually is. Uses Gitea's SUPPORTED surface (custom templates +
public assets) so upstream upgrades keep arriving — no source modified (D-2/I-1).

Two traps this cost, both now documented and tested:

1. GITEA__DEFAULT__APP_NAME does not work. Gitea reads APP_NAME from the TOP
   LEVEL of app.ini; the env var created a literal [default] section that Gitea
   ignores, so the installer's stock APP_NAME kept winning while the config
   looked correct. The env-to-ini pass also APPENDED a second APP_NAME rather
   than replacing the first — a new variant of the documented G4A.3 trap.

2. Cloudflare caches /assets/* for 6h and no token in this stack can purge, so
   the new logo and CSS were invisible while being correct at origin. Brand
   assets now carry a VERSION IN THE FILENAME; bump it on every change.

Committed with an idempotent apply.sh, because applying it straight to Veron's
disk first was itself the config-drift trap this project documents: a rebuild
would have silently reverted to stock Gitea.

85 tests green.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 09:08:48 -04:00
Grant Whitmer
14fb43b959 G7.4: advertise the fleet's runner labels, and correct my own doctrine
Some checks are pending
check / gate (push) Waiting to run
Surveyed ten repos: 36 of 36 ACTIVE workflows already say
runs-on: [self-hosted, linux, x64]. They were written for the self-hosted
runners that died when the repos went private — so advertising those three
labels makes every one of them runnable AS-IS. No workflow edits, no rewrites.

Correcting an earlier note in this plan: 'ban ubuntu-latest' was wrong as
stated. All 11 occurrences are tagged '# runner-lint-allow — CD/hosted-only;
disabled, manual until CD mission'. They are deliberately hosted-only and
deliberately off. They are correct as written and should not be 'fixed'.

I generalised that doctrine from one repo without surveying. The survey
disagreed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 19:53:44 -04:00
Grant Whitmer
6b96087554 G7: per-job networks — fixes service DNS and tightens isolation
Some checks failed
check / gate (push) Failing after 13m33s
The migration step failed with 'could not translate host name postgres'. The
Postgres service container was healthy; the job simply could not name it,
because service DNS aliases only exist on a per-job network and I had pinned
container.network to the flat 'bridge'.

That choice was wrong in both directions: it broke service containers AND it
was weaker isolation, since every concurrent job shared one bridge and could
see its neighbours. A per-job network is stricter and correct — and still has
no route to the forge, because these networks live inside the dind daemon,
which has no forge attachment at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:18:51 -04:00
Grant Whitmer
53dea84ab1 G7: route CI jobs via the public surface, not the forge network
Some checks failed
check / gate (push) Failing after 13s
The first CI run failed with 'Could not resolve host: gitea' — job containers
live on dind's private network and cannot see the forge network. Two ways out,
and they are not equivalent:

  (a) put job containers on the forge network. Easy, one line, and it leaves
      untrusted workflow code one DNS name from the forge's Postgres. It quietly
      repeals I-5.
  (b) send jobs to the PUBLIC forge surface over the tunnel, exactly like any
      stranger on the internet.

Took (b). The runner no longer needs the forge network at all, so there is now
NO private route from any CI container to anything — a better posture than this
file started with. Cost is a hairpin through Cloudflare plus its ~100s ceiling
per fetch, which for 0.63 GB of objects across 61 repos and depth=1 checkouts is
nowhere near binding.

Test upgraded to assert the stronger property: no CI container joins the forge
network.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 10:51:00 -04:00
Grant Whitmer
5f717ef74b G7: CI runners with real isolation, and the gate as a workflow
Some checks failed
check / gate (push) Failing after 50s
I-5 says runners execute untrusted code and must be isolated by machine
boundary. act_runner needs a Docker daemon to start job containers, and the
tempting move — what every published example does — is to mount the host's
/var/run/docker.sock. That hands every workflow, including whatever a
transitive dependency's postinstall script feels like doing, the ability to
start a privileged container mounting / — root on Grant's workstation.

Instead the runner talks to its OWN dind daemon:
  - runner (TRUSTED, the act_runner daemon) sits on the forge network only to
    collect jobs from gitea:3000
  - dind and every job container it spawns are UNTRUSTED, on a private network
    with no route to the forge, its Postgres, or its .env
  - jobs cannot bind-mount from the daemon host (valid_volumes: []) and are not
    handed the runner's own socket (docker_host: -)
  - separate compose project, cpu/memory bounded — Veron 1 is Grant's
    workstation, not a dedicated build box

The gate itself now runs as a workflow, including the migration round-trip that
already caught two bugs review did not, and the I-12 check that a COMMIT_SHA
env override cannot change what /version reports.

Labels are explicit and pinned. A workflow naming a label nobody provides
queues forever and presents as a hung CI system rather than a typo — which is
what ubuntu-latest does on every windy-registry run today.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 10:39:05 -04:00
Grant Whitmer
eec087ac50 G2: Gitea auto-install, own DB role, SSH deferred, Actions on
Gitea was configured to connect as a 'gitea' DB user that never existed, so it
sat unconfigured behind a working tunnel. It now gets its OWN role and database
via a first-init script — I-1 says we never write Gitea's tables, and that is
better as a permission than as a promise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 14:44:13 -04:00