- bridge: BRIDGE_NO_DAEMON names image-build jobs whose name lacks docker
(default eternitas:ci/build); never posted, like the docker-named ones.
- ci-hygiene: flag docker build/buildx/run/compose, docker-compose and
docker/build-push-action in workflow steps ("needs docker") with the fix:
job services: + a no-Docker smoke test; the image builds at deploy.
- test_guards_report: owner column (14ed23a broke it).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The mirror kept its creation-time base forever. eternitas #167 was stacked
on fix/one-hallway, retargeted to main after #166 merged; ci.yml
(pull_request: branches [main]) then silently never ran for it, while
unfiltered workflows did. An edited event triggers nothing, so close the
stale mirror and open a fresh one on the new base (runs CI at once).
An unknown base is left alone, never guessed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Allow entries may carry matches: (regexes); then only matching lines are
allowed, so an allowed file can't smuggle in a new call. windy-pro #609
MindKeychain.jsx: openrouter.ai/auth? and /api/v1/auth/keys (BYOK key
acquisition via OAuth PKCE, no inference; successor of the MindPanel
allow, ADR-064). An inference call in the same file still flags (tested).
Orchestrator-approved.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scans every bridged default branch with both guards; lane-owned vs
Grant-owned (ci/grant-owned.yml: windy-pro desktop paths + its desktop CI
jobs, attributed per job) so Grant's code never holds up a block.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The bridge reads PR/default heads from GitHub after the sync's fetch; a
push in between isn't in the clone until the next cycle. Both guards logged
a CalledProcessError for it (windy-pro main 40 s after the fetch). Now
None = nothing posted this cycle; the next one scans it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
House rule 6 (09-23). The bridge now also posts windy-git/ci-hygiene on
every PR head (added lines) and default branch (whole files), scanning CI
workflows and Dockerfiles for: floating pip / uv pip installs (not -r,
not --no-deps, not exact pins), uv sync without --locked/--frozen,
npm install instead of npm ci (unless every package is exact-pinned),
yarn/pnpm without a frozen lockfile, :latest images and COPY lock* globs
(Windy Mail #147), and CI services publishing a HOST port (every job
shares one dind: Windy Mind runs 147/176 died on 5432). Warn-only;
CI_HYGIENE_MODE=block later. Allow-list ci/ci-hygiene-allow.yml (empty).
compute_guard's walker is now parameterised (line_fn / path_ok /
prefilter) so both guards share one scanner, cache and allow loader; the
bridge posts both through one _post_guard. Today: 95 issues in 21 repos;
windy-git, calendar, traveler, traveler-site clean.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Grant's rule (09-23): every model call goes through Windy Mind. The bridge
now posts windy-git/compute-guard on every PR head (lines the PR ADDS vs its
merge-base) and default-branch head (whole tree): provider hosts, provider
SDK imports/deps and raw provider key names. Warn-only: success + "⚠ WARN"
and a link to the first hit; COMPUTE_GUARD_MODE=block turns it red later.
Exceptions live in ci/compute-guard-allow.yml, each with a reason (Mind
itself, user-BYOK windy-agent / windy-code extension / windy-pro desktop +
MindPanel, windy-connect config writers). Tests, docs, comments, lockfiles,
vendored code and CI config are never scanned. Reads the sync's bare clones
(no docker exec); cached per (repo, sha, rules). Non-fatal; never a fake OK.
First cases = COMPUTE_BYPASS_AUDIT.md. Today on default branches: 38
findings in 3 repos (windy-chat audit #2, windy-pro account-server #3/#4,
windytalk reference/), 0 elsewhere.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Gitea 1.24 lists only picked-up jobs, so a queued PR showed NOTHING on
GitHub and lanes asked whether their push was lost (Windy Mind #131,
Windy Cloud today). The bridge now reads waiting jobs from the gitea DB
and posts pending where nothing newer was picked up; a queued re-run
supersedes the stale failure it replaces.
Only status 5 jobs whose runs-on labels a live runner has: blocked jobs
often end skipped and label-unrunnable jobs are cancelled unpicked, and
neither ever reaches /actions/tasks, so their pending would never resolve.
Lookup is bounded (30 s) and non-fatal: the IO-stall lesson.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Gitea 1.24 has no rerun API and the web button needs Grant's SSO identity.
Guarded branch rewind that the next sync undoes; restores the branch itself
on timeout. Used today for eternitas #166 and windy-mind #131 after the
Veron IO stall.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Telemetry UPDATE 2 actor rule: agent/human rows without actor_id are
quarantined. Forge humans sign in only via Windy SSO, so Gitea's
external_login_user.external_id is their windy_identity_id.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
git push never touches our API, so throttle.py can't see it. Gitea's
action table records every push; the 5-min sync-side emitter now reads
it and emits forge.push_velocity when an account crosses 60 pushes/1h,
500 pushes/24h (standard-band base) or 10 ref deletes/24h. One row per
account per rule per window while over; windyadmin (the sync) exempt.
Nothing sits in the push path and nothing is refused. HOLD until
Telemetry Boss declares the shape.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
identity.login created a live hub session every 10 min and never ended it.
It now logs out with the token it got: retried on 5xx / no response
(8 x 15 s), 401/404/410 = already over, any other 4xx fails fast, and a
cleanup it can't finish is reported as identity.logout DOWN "CLEANUP
FAILED" (alerts + red run). The hub's /auth/logout revokes every refresh
token of the account (verified live), so the next run's logout heals a
leftover; no ledger needed. Proven end to end: login 200, logout 200,
10/10 checks.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
09-23 16:43Z the Veron data2 SMR stall left runc exec in D state; the
janitor's docker exec never returned, so the sync sat 'activating' and no
repo mirrored or got a status for any lane. Both steps are non-fatal;
now they time out (120 s / 180 s) and the run continues.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The ledger answers 202 even when it quarantines rows. Both emitters now log
a warning with the reasons and report service.health.telemetry_quarantined
and telemetry_dropped (API: buffer overflow; sync: 0 by construction, since
a failed send keeps cursor + spool). HOLD until Telemetry Boss declares both
keys on windy-git's two service.health shapes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Gitea drops an invalid workflow file with one log line and fires no run, so
the GitHub PR showed nothing and lanes waited for CI that never came
(windytalk #100). The bridge now reads each workflow file at the commit it
reports on and posts windy-git/<wf>/workflow = error with the reason.
Verified: 0 false positives on all 23 bridged repos' main; catches
windytalk #100's broken commits (invalid YAML at line 12), fix commit clean.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A single GitHub TLS handshake timeout failed the whole sync, flipped its
windy-job heartbeat to ok:false and would page for nothing. Up to 3
attempts with backoff for URLError/timeout/reset; HTTP errors return
immediately as before. Test covers both.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replaces the keyed marker from 1c3b5b0 with the ecosystem convention:
any X-Windy-Synthetic value marks the request synthetic; the flag lives in
a per-request contextvar, labels this request's rows, and is FORWARDED on
downstream calls (Eternitas trust lookup, Gitea API). The canary sends
"1". Rows are still recorded; the label separates, never suppresses.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The canary deliberately sends forged tokens every 10 min; those refusal
rows read as attacks. It now sends X-Windy-Synthetic carrying a shared
secret (Gitea repo secret CANARY_SYNTHETIC_KEY = WINDYGIT_SYNTHETIC_KEY in
Veron .env); the API marks the row synthetic only on a constant-time
match, so an attacker cannot label their own refusals synthetic to hide.
synthetic is declared on forge.auth.failed (Telemetry Boss, UPDATE 3).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Telemetry Boss found jobs_finished=43 vs 8 ci.run rows. Root cause: the
high-water mark was the job id, but jobs FINISH out of id order, so every
long job that started before the mark and finished after it was silently
never emitted. Now a (finish time, id) cursor; finish = stopped, or
updated for skipped jobs with no stop time. Heartbeat finished/failed/
cancelled counts are derived from exactly the rows emitted, so
sum(jobs_finished) == count(ci.run) by construction (dry run on real
data: 97 == 97, failed 2 == 2, cancelled 13 == 13). posted_to_github now
set from the bridge's own rules. duration_ms = Gitea whole seconds x 1000.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The janitor now returns one JSON line per job it cancels (repo, workflow,
job, reason, runs_on, waited_s) into a spool; the emitter ships them as
ci.job_cancelled (declared with Telemetry Boss) and truncates the spool
only after a 2xx. Run status recompute folded into the same statement.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ci.run (one row per finished job, exactly once via a high-water mark;
branch_kind default|pr|other so the dashboard can show "main is red") and
service.health (interval counts: finished/failed/cancelled, waiting,
running, runners online, oldest wait). Shapes declared with Windy
Telemetry 40; sends nothing until WINDYGIT_TELEMETRY_TOKEN exists.
State is only advanced after a 2xx, so a failed post retries.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
When a needed job fails, Gitea leaves dependants BLOCKED (7) even after
the run finishes; eternitas build jobs sat there 8h. Mark them skipped
(what GitHub shows) once the run is done and 30 min have passed.
Found by the new telemetry dry run (oldest_waiting_s = 29160).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
windy-pro alone left ~4 jobs per run waiting forever (build-electron on
macos/windows/ubuntu-latest, deploy if:false): Gitea evaluates job if:
only at pick time, the labels do not exist here, and waiting jobs are
invisible in /actions/tasks. 37 such jobs across 10 runs today. After
30 min they are cancelled and the run status recomputed. Runs each sync,
non-fatal.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
build-desktop, test-installer and reality-check still run on Windy Git
and stay visible there, but the bridge no longer posts them to GitHub, so
they cannot turn windy-pro's combined status red. Windy Git side only;
the desktop code is Grant's to fix. BRIDGE_NON_BLOCKING, per repo.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
kit-army-config (the lockbox) and every *-soul / anima repo carry
credentials; the nightly R2 bundles are unencrypted, so the R2 key was a
key to every secret. Excluded by name (BACKUP_EXCLUDE); they are backed up
encrypted by the Windy Drops lane (restic) and stay mirrored on Veron.
Behavioural test runs the script's own exclusion function.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
windy-mind has been writable + CI on Windy Git since 08-13 (deploy.yml
disabled, uv pinned); it only lacked GitHub commit statuses.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
archive/<machine>-<date>/<branch> are off-machine safety copies of
unpushed work (one-repo doctrine). GitHub holds them; running CI on them
is waste. Negative refspec ^refs/heads/archive/* on the push (git 2.43
on Veron). Requested by 8c for windy-pro's Mac mini archive.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- deploy/systemd/: sync/backup timers+services, tunnel, and the windy-job
heartbeat drop-ins (silent-failure audit). They existed only on Veron,
the same drift that left the runbook wrong. GITHUB_TOKEN is stripped
(repo is public); it stays in the root-only unit on the host.
- sync: SYNC_NO_TAGS (default windy-pro). build-electron fires on v* tags
and targets ubuntu/macos/windows-latest, labels no runner has, so it
would queue forever, invisibly. Desktop releases are built elsewhere.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- import_from_github.py no longer refuses windy-pro: lane 8c audited all
14 checkouts (WINDYPRO_CHECKOUTS.md), GitHub main is canonical.
- sync + bridge: windy-cloud-domains, windy-cloud-vps, windytalk (default
branch master), windy-pro; windy-cloud-sites added to the bridge (it was
synced but never bridged).
- scripts/promote_to_ci.sh: the mirror->writable procedure as one script,
with the delete-before-import hazard documented.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
windy-hand promoted to writable + bridged. windy-agent is PUBLIC and its
GitHub Actions already run on Veron's GitHub runner; running its 3-version
pytest matrix here too was pure duplicate load (3 x ~4.5 cores for 25+
min, host load 64 on 24 cores). Actions are now off for windy-agent on
Windy Git; it stays in REPOS as a synced copy.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Promoted to writable; the two sites' CF deploy.yml disabled (they would
need the CF god token in a job container). All three only have
ubuntu-latest workflows today, so nothing runs until their lanes switch
runs-on to [self-hosted, linux, x64].
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- hub_jwt: aud list is ["windy_git"] (contract v1 array). Dropped
"windy-git": that is Gitea's OIDC client_id, so a forge id_token would
have passed the aud check. `type: human` is now REQUIRED (id_tokens have
none), which makes accepting the discovery-URL issuer safe.
- runner job ceiling 30m -> 90m: eternitas's serial pytest is ~50 min and
would have been killed mid-suite.
- eternitas (private) added to the GitHub status bridge.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
All four promoted from pull mirrors to writable. windy-code keeps only
canonical-domains-lint active: its other 15 workflows target hosted
macOS/Windows/ubuntu-latest runners and would queue forever here.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- deploy/runner/prune.sh + windygit-ci-prune.timer (6h): age-based prune
of the CI-only dind (containers, finished-job volumes, images/builder
cache >7d) plus a hard 60 GB cap. Only that daemon, over its own TCP
socket; never the host's Docker. It was 38 GB and unbounded — the same
class of growth that filled Kit 0 on 09-01.
- pr_status_bridge: jobs named *docker* are not posted. Job containers
have no daemon by design (I-5), so they are red on every commit; a
permanent red X teaches everyone to ignore red.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
windy-connect promoted from pull mirror to writable (release.yml, which
publishes to PyPI on tag push, disabled — the sync pushes tags).
windy-search was already writable; its scheduled drift-check is disabled
because it now runs as cron on Kit 0. Both added to BRIDGE_REPOS.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Four runners x capacity 1 instead of one x capacity 4. Concurrent jobs in
one act_runner share /root/.cache/act; a refresh racing a copy killed 3 of
windy-chat's ~20 jobs at setup-node (lstat ... no such file). Separate
processes have separate caches. Same parallelism, same capped dind.
- Behavioral tests for pr_status_bridge (latest verdict wins, no reposting,
skipped never painted green, fork PRs never run, pagination, PR lifecycle).
- import_from_github.py reads IMPORT_GITEA_URL, not GITEA_BASE_URL: sourcing
the deploy .env pointed it at http://gitea:3000 and it died on DNS after the
mirror it replaces had already been deleted.
- CUTOVER.md: the private-repo CI path, onboarding steps, and the
/actions/tasks-hides-queued-runs trap.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
GitHub Actions can't run on the private platform repos. Windy Git already
has their code and a working runner, so:
- windy-chat and windy-mail were read-only pull mirrors (which can never
run Actions); they are now writable, deploy.yml/build-image.yml disabled,
and synced from GitHub like the others.
- scripts/pr_status_bridge.py mirrors open same-repo GitHub PRs into Windy
Git (so pull_request workflows fire) and posts each job's result back as
a GitHub commit status (windy-git/<workflow>/<job>) on PR heads and the
default-branch head. Fork PRs are never run. Runs after every sync.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
windygit-tunnel had crash-looped ~91k times: another project's
cornercall-tunnel holds 127.0.0.1:2000, and cloudflared exits when it
cannot bind its metrics port. Ingress only survived because a stray
cloudflared.service ran the same config. That unit is now disabled and
/etc/cloudflared/config.yml uses metrics 127.0.0.1:2001.
Also add windy-git to the GitHub->Windy Git sync list; its self-hosted
copy was stuck 3 commits behind (only check + canary workflows, no
deploys, so syncing is safe).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The EPT-shaped forgery is the one that matters after G3.2 — it is what
signature verification actually guards. The JWT-shaped one still exercises the
human gate. Both must_refuse; a 2xx on either pages.
Verified live: forged EPT -> 401 ept_invalid, forged JWT -> 503,
genuine EPT -> 200.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
Adds must_refuse checks: a 2xx is the alarm, a 401/403/503 is health. The
forged alg:none agent token is probed every 10 min; if it ever returns 2xx the
canary goes red and pages. Proven both ways — 503 reads ok, a 200 endpoint
alarms 'ACCEPTED — this MUST be refused'. The security property is now enforced
by a running check, not assumed.
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
The function landed but the argparse anchor did not match the real formatting,
so the command existed and could not be invoked. Caught by running it rather
than assuming the patch applied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Grant's plan: clone the whole account, let it circulate, reverse direction later
when things are clean. Right plan, with one change that matters.
Sampling 40 repos found 18 carrying deploy/release/publish workflows that
trigger on push: — roughly 63 across the account. Importing those writable with
Actions enabled would arm sixty-odd production deploy triggers on Veron 1, each
needing disarming by hand.
So the bulk goes in as READ-ONLY pull mirrors. A mirror cannot run Actions at
all, so this carries zero deploy risk, and Gitea syncs them itself with no
script and no timer. What you get is a complete, current second copy of the
whole account — the disaster-recovery half — with none of the execution risk.
Converting one to writable + CI stays a deliberate per-repo act: re-import,
review its workflows, disable the deploying ones. That judgement belongs at the
moment you want CI on that repo, not in bulk sixty times by accident.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
I migrated nine repos writable with push-mirrors pointed AT GitHub. That was
wrong for the actual situation: a dozen agent sessions on the Mac mini are
pushing to GitHub continuously, so GitHub is where the live work is.
A push-mirror force-updates refs. On its 8-hour timer it would have pushed
Windy Git's stale copy over live work — silently, no conflict, nothing to
notice. Removed all nine before the first timer fired; verified no GitHub repo
had been touched (latest push predated the mirrors).
Replaced with the correct Phase 1 direction:
agents --push--> GitHub --sync--> Windy Git --> CI on Veron
It requires NOTHING from anyone. No remote changes, no coordination, no
'everybody stop pushing'. Agents keep working exactly as they are and CI starts
running on 24 cores.
Windy Git is force-updated on purpose: in Phase 1 it holds nothing anyone
depends on, so GitHub always wins and there is no merge to reconcile.
Phase 2 is per-repo and only when that repo is idle. Never a big-bang cutover
across a dozen live sessions.
Fetches +refs/heads/* and tags explicitly rather than --mirror, which would drag
GitHub's refs/pull/* that Gitea rejects and bury the real errors.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>