Commit Graph

156 Commits

Author SHA1 Message Date
Kit OC5
e3b69fa759 telemetry: UPDATE 7 — read the ingest body; count quarantined + dropped on heartbeats
Some checks failed
canary / probe (push) Has been cancelled
check / gate (push) Has been cancelled
The ledger answers 202 even when it quarantines rows. Both emitters now log
a warning with the reasons and report service.health.telemetry_quarantined
and telemetry_dropped (API: buffer overflow; sync: 0 by construction, since
a failed send keeps cursor + spool). HOLD until Telemetry Boss declares both
keys on windy-git's two service.health shapes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:32:28 -04:00
Kit OC5
4acf50d9ef bridge tests: fake serves workflow contents; cover invalid-workflow status
All checks were successful
check / gate (push) Successful in 23s
b7a7e94 made the bridge read workflow files, which the strict fake Gitea
refused (7 red). The fake now serves contents (404 when absent), and new
tests cover: error posted with no runs, valid files add nothing, no repost,
.gitea/workflows wins over .github/workflows, and each workflow_problem
shape. pyyaml declared in dev extras (the bridge imports it).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:31:19 -04:00
Kit OC5
b7a7e94df0 bridge: post an error status when Windy Git ignores an invalid workflow
Gitea drops an invalid workflow file with one log line and fires no run, so
the GitHub PR showed nothing and lanes waited for CI that never came
(windytalk #100). The bridge now reads each workflow file at the commit it
reports on and posts windy-git/<wf>/workflow = error with the reason.
Verified: 0 false positives on all 23 bridged repos' main; catches
windytalk #100's broken commits (invalid YAML at line 12), fix commit clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:29:12 -04:00
c83f808a60 bridge: retry transport blips (TLS timeout/reset), never HTTP errors
All checks were successful
check / gate (push) Successful in 20s
canary / probe (push) Successful in 9s
A single GitHub TLS handshake timeout failed the whole sync, flipped its
windy-job heartbeat to ok:false and would page for nothing. Up to 3
attempts with backoff for URLError/timeout/reset; HTTP errors return
immediately as before. Test covers both.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:15:22 -04:00
eb27e63db3 telemetry: adopt the end-to-end synthetic convention (UPDATE 4)
All checks were successful
check / gate (push) Successful in 24s
canary / probe (push) Successful in 9s
Replaces the keyed marker from 1c3b5b0 with the ecosystem convention:
any X-Windy-Synthetic value marks the request synthetic; the flag lives in
a per-request contextvar, labels this request's rows, and is FORWARDED on
downstream calls (Eternitas trust lookup, Gitea API). The canary sends
"1". Rows are still recorded; the label separates, never suppresses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:06:51 -04:00
1c3b5b0638 telemetry: synthetic:true on canary refusals (keyed, not a bare flag)
All checks were successful
check / gate (push) Successful in 25s
canary / probe (push) Successful in 7s
The canary deliberately sends forged tokens every 10 min; those refusal
rows read as attacks. It now sends X-Windy-Synthetic carrying a shared
secret (Gitea repo secret CANARY_SYNTHETIC_KEY = WINDYGIT_SYNTHETIC_KEY in
Veron .env); the API marks the row synthetic only on a constant-time
match, so an attacker cannot label their own refusals synthetic to hide.
synthetic is declared on forge.auth.failed (Telemetry Boss, UPDATE 3).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:54:56 -04:00
90643fe48e telemetry step 2: API boot/health + forge.auth.failed (declared)
All checks were successful
check / gate (push) Successful in 25s
canary / probe (push) Successful in 6s
Membrane first: I-2 and MEMBRANE.v1 now list the windy-admin ledger
(POST /v1/events). api/app/telemetry.py: service.boot once per start
(commit_sha omitted when unknown, I-12), an hourly in-process
service.health with the shared keys (requests, errors_5xx/4xx,
refusals_4xx, p95_ms only when there was traffic), and one
forge.auth.failed row per refused request: declared 13-code enum,
http_status, caller class, route TEMPLATE (never the concrete path),
actor_type system with no actor_id (all-lanes rule). No token = nothing
sent or buffered; flush failures keep rows (bounded) and never raise.
Token from root-only /etc/windygit/telemetry.env (optional env_file).
8 behavioural tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:47:29 -04:00
00ec963f82 telemetry: fix ci.run completeness — cursor on (finish time, job id)
Some checks failed
check / gate (push) Has been cancelled
Telemetry Boss found jobs_finished=43 vs 8 ci.run rows. Root cause: the
high-water mark was the job id, but jobs FINISH out of id order, so every
long job that started before the mark and finished after it was silently
never emitted. Now a (finish time, id) cursor; finish = stopped, or
updated for skipped jobs with no stop time. Heartbeat finished/failed/
cancelled counts are derived from exactly the rows emitted, so
sum(jobs_finished) == count(ci.run) by construction (dry run on real
data: 97 == 97, failed 2 == 2, cancelled 13 == 13). posted_to_github now
set from the bridge's own rules. duration_ms = Gitea whole seconds x 1000.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:42:36 -04:00
b5e4eaf57a telemetry: ci.job_cancelled from the janitor; interval_s on heartbeat
All checks were successful
check / gate (push) Successful in 22s
canary / probe (push) Successful in 6s
The janitor now returns one JSON line per job it cancels (repo, workflow,
job, reason, runs_on, waited_s) into a spool; the emitter ships them as
ci.job_cancelled (declared with Telemetry Boss) and truncates the spool
only after a 2xx. Run status recompute folded into the same statement.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:27:53 -04:00
baaa542bae docs: audit disposition 09-23 — R2 god token replaced by bucket-scoped token
All checks were successful
check / gate (push) Successful in 43s
canary / probe (push) Successful in 7s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:15:35 -04:00
0634a6cb1b style: ruff fix in telemetry_emit
All checks were successful
check / gate (push) Successful in 29s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:08:21 -04:00
1a171eabd5 telemetry: CI emitter for admin.windyword.ai (inert until token)
Some checks failed
check / gate (push) Failing after 20s
canary / probe (push) Successful in 6s
ci.run (one row per finished job, exactly once via a high-water mark;
branch_kind default|pr|other so the dashboard can show "main is red") and
service.health (interval counts: finished/failed/cancelled, waiting,
running, runners online, oldest wait). Shapes declared with Windy
Telemetry 40; sends nothing until WINDYGIT_TELEMETRY_TOKEN exists.
State is only advanced after a 2xx, so a failed post retries.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:06:54 -04:00
28c31236b8 ci: janitor also clears jobs blocked forever on failed needs
When a needed job fails, Gitea leaves dependants BLOCKED (7) even after
the run finishes; eternitas build jobs sat there 8h. Mark them skipped
(what GitHub shows) once the run is done and 30 min have passed.
Found by the new telemetry dry run (oldest_waiting_s = 29160).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:06:33 -04:00
cd5967031b ci: janitor cancels jobs no runner can ever take
All checks were successful
check / gate (push) Successful in 20s
canary / probe (push) Successful in 6s
windy-pro alone left ~4 jobs per run waiting forever (build-electron on
macos/windows/ubuntu-latest, deploy if:false): Gitea evaluates job if:
only at pick time, the labels do not exist here, and waiting jobs are
invisible in /actions/tasks. 37 such jobs across 10 runs today. After
30 min they are cancelled and the run status recomputed. Runs each sync,
non-fatal.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 10:58:53 -04:00
f246417095 ci: windy-pro desktop/installer jobs are NON-BLOCKING (Grant, 09-23)
build-desktop, test-installer and reality-check still run on Windy Git
and stay visible there, but the bridge no longer posts them to GitHub, so
they cannot turn windy-pro's combined status red. Windy Git side only;
the desktop code is Grant's to fix. BRIDGE_NON_BLOCKING, per repo.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 10:58:07 -04:00
50c1464043 security: never bundle credential repos to R2 in plaintext
All checks were successful
check / gate (push) Successful in 23s
canary / probe (push) Successful in 7s
kit-army-config (the lockbox) and every *-soul / anima repo carry
credentials; the nightly R2 bundles are unencrypted, so the R2 key was a
key to every secret. Excluded by name (BACKUP_EXCLUDE); they are backed up
encrypted by the Windy Drops lane (restic) and stay mirrored on Veron.
Behavioural test runs the script's own exclusion function.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 10:48:42 -04:00
7a63f90da3 ci: bridge windy-mind (private) verdicts to GitHub
All checks were successful
check / gate (push) Successful in 23s
canary / probe (push) Successful in 6s
windy-mind has been writable + CI on Windy Git since 08-13 (deploy.yml
disabled, uv pinned); it only lacked GitHub commit statuses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 09:59:28 -04:00
b8f97f0731 ops: sync never pushes archive/* branches to Windy Git
All checks were successful
check / gate (push) Successful in 20s
archive/<machine>-<date>/<branch> are off-machine safety copies of
unpushed work (one-repo doctrine). GitHub holds them; running CI on them
is waste. Negative refspec ^refs/heads/archive/* on the push (git 2.43
on Veron). Requested by 8c for windy-pro's Mac mini archive.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 09:57:57 -04:00
95c33c8004 ops: host systemd units in git; windy-pro tags never reach Windy Git
All checks were successful
check / gate (push) Successful in 22s
- deploy/systemd/: sync/backup timers+services, tunnel, and the windy-job
  heartbeat drop-ins (silent-failure audit). They existed only on Veron,
  the same drift that left the runbook wrong. GITHUB_TOKEN is stripped
  (repo is public); it stays in the root-only unit on the host.
- sync: SYNC_NO_TAGS (default windy-pro). build-electron fires on v* tags
  and targets ubuntu/macos/windows-latest, labels no runner has, so it
  would queue forever, invisibly. Desktop releases are built elsewhere.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 09:51:51 -04:00
d5181f1c6d ci: G11.5 resolved — onboard windy-pro; cloud cells + windytalk; promote script
All checks were successful
check / gate (push) Successful in 21s
canary / probe (push) Successful in 6s
- import_from_github.py no longer refuses windy-pro: lane 8c audited all
  14 checkouts (WINDYPRO_CHECKOUTS.md), GitHub main is canonical.
- sync + bridge: windy-cloud-domains, windy-cloud-vps, windytalk (default
  branch master), windy-pro; windy-cloud-sites added to the bridge (it was
  synced but never bridged).
- scripts/promote_to_ci.sh: the mirror->writable procedure as one script,
  with the delete-before-import hazard documented.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 09:49:15 -04:00
e6530d3171 test: behavioral G3.5 webhook tests (audit: tests were source-string asserts)
All checks were successful
check / gate (push) Successful in 20s
canary / probe (push) Successful in 9s
Seven HTTP-level tests through the real route: sha256= prefix and bare
digests accepted, digest of re-serialised JSON refused, forged/wrong-key/
missing signatures refused, unset secret -> 503, a revocation with a bad
signature never reaches the handler, the reachability ping never acts.
Mutation-checked: dropping the prefix strip fails the behavioral test
while the old string-grep invariant still passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 03:26:14 -04:00
419443573a ci: onboard windy-hand; stop running windy-agent CI twice
All checks were successful
check / gate (push) Successful in 23s
windy-hand promoted to writable + bridged. windy-agent is PUBLIC and its
GitHub Actions already run on Veron's GitHub runner; running its 3-version
pytest matrix here too was pure duplicate load (3 x ~4.5 cores for 25+
min, host load 64 on 24 cores). Actions are now off for windy-agent on
Windy Git; it stays in REPOS as a synced copy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 03:23:25 -04:00
4a34b35441 security: CI egress filter — jobs reach the internet, never Veron/LAN
Measured: an unprivileged job container inside the CI dind could open SSH,
Ollama and dev servers on Veron (192.168.1.73) and the rest of the LAN,
WireGuard and Tailscale — lateral movement for any malicious dependency,
no escape needed. egress.sh (idempotent; windygit-ci-egress.service at
boot) hooks the jobs bridge: runner<->dind, replies, DNS and public
egress allowed; RFC1918, CGNAT, link-local and the host itself dropped.
Verified from a job container: 6/6 private targets blocked, DNS,
internet and the public forge OK.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 03:20:45 -04:00
dfe5543eda docs: runbook + AGENTS match reality (item 5 of the launch bar)
Some checks failed
check / gate (push) Has been cancelled
canary / probe (push) Successful in 34s
RUNBOOK-VERON: deploy uses fetch + merge --ff-only and api-only rebuilds
(the old text used git pull, contradicting its own warning); new sections
for host timers, CI (6 runners x1, windyadmin-scoped, 90m ceiling, queue
truth in the gitea DB, logs in R2), sign-in posture and break-glass;
standing checkout = OC5. AGENTS.md no longer says GENESIS / no code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 03:19:12 -04:00
40cb455d0d ci: onboard windy-translate, windytranslate-site, windytraveler-site
Some checks failed
check / gate (push) Has been cancelled
Promoted to writable; the two sites' CF deploy.yml disabled (they would
need the CF god token in a job container). All three only have
ubuntu-latest workflows today, so nothing runs until their lanes switch
runs-on to [self-hosted, linux, x64].

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 03:13:40 -04:00
8c404eb410 security: turn off Gitea OAuth auto-registration
Any Windy Word account (public signup, unverified email) auto-registered a
forge account on first sign-in — reproduced with a throwaway account —
and the act runners are instance-wide, so a stranger's workflow would run
on Veron's privileged dind beside the R2 god token. §7 makes opening the
forge to non-Grant users Grant's call.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 03:09:43 -04:00
5b16114b98 auth: token contract v1 (aud windy_git, both issuers); CI for eternitas
Some checks failed
check / gate (push) Successful in 37s
canary / probe (push) Has been cancelled
- hub_jwt: aud list is ["windy_git"] (contract v1 array). Dropped
  "windy-git": that is Gitea's OIDC client_id, so a forge id_token would
  have passed the aud check. `type: human` is now REQUIRED (id_tokens have
  none), which makes accepting the discovery-URL issuer safe.
- runner job ceiling 30m -> 90m: eternitas's serial pytest is ~50 min and
  would have been killed mid-suite.
- eternitas (private) added to the GitHub status bridge.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:58:56 -04:00
18ea9a4686 auth: G3.2 hub JWKS verifier — humans can sign in to the plane (SSO #14)
The human path refused every token in production (503
human_signin_not_ready) because no verifier existed. api/app/hub_jwt.py
verifies hub access tokens against account.windyword.ai's JWKS:

- RS256 only (closes alg:none and HS256-with-public-key confusion)
- iss must be "windy-identity" — what hub ACCESS tokens carry (observed
  live); id_tokens (discovery-URL issuer) are not accepted as bearers
- aud optional today, must name Windy Git when present; hub_require_aud
  flips it mandatory once the hub emits it. PyJWT's own aud check is off
  on purpose: it rejects ANY aud-bearing token when no audience is given.
- type must be human; identity = windy_identity_id, never sub (per-row id)
- production verifies even if require_verified_jwt is off

11 behavioral tests sign real RS256 tokens with a local key.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:55:17 -04:00
32e8ac8474 ci: bridge windy-registry (private) PR/main verdicts to GitHub
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:53:20 -04:00
8e9fa3116c ci: six runners; SSO #8 Gitea sign-in hardening (staged)
- runner-5/6: 50+ jobs were queued with ~11 private repos onboarded. dind
  keeps the 12-core ceiling, so this adds concurrency, not CPU.
- Gitea: password + passkey sign-in forms off (break-glass = CLI), and
  ACCOUNT_LINKING auto -> login. auto linked any hub login whose email
  matched an existing account, and SITE ADMIN windyadmin carries Grant's
  email. Grant is linked by the hub's stable sub, which matches first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:51:56 -04:00
74ad4950b2 ci: onboard windy-drops, windy-code-web, windy-code, windy-traveler
All checks were successful
check / gate (push) Successful in 1m1s
canary / probe (push) Successful in 38s
All four promoted from pull mirrors to writable. windy-code keeps only
canonical-domains-lint active: its other 15 workflows target hosted
macOS/Windows/ubuntu-latest runners and would queue forever here.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:49:00 -04:00
e7bbf9af51 ci: prune.sh must address dind over TCP (it has no unix socket)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:47:35 -04:00
45686283be ci: bound CI storage; don't bridge image-build jobs
- deploy/runner/prune.sh + windygit-ci-prune.timer (6h): age-based prune
  of the CI-only dind (containers, finished-job volumes, images/builder
  cache >7d) plus a hard 60 GB cap. Only that daemon, over its own TCP
  socket; never the host's Docker. It was 38 GB and unbounded — the same
  class of growth that filled Kit 0 on 09-01.
- pr_status_bridge: jobs named *docker* are not posted. Job containers
  have no daemon by design (I-5), so they are red on every commit; a
  permanent red X teaches everyone to ignore red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:47:28 -04:00
e4a15869c0 ci: onboard windy-connect + windy-search to private-repo CI
Some checks failed
check / gate (push) Has been cancelled
windy-connect promoted from pull mirror to writable (release.yml, which
publishes to PyPI on tag push, disabled — the sync pushes tags).
windy-search was already writable; its scheduled drift-check is disabled
because it now runs as cron on Kit 0. Both added to BRIDGE_REPOS.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:44:19 -04:00
dcf9286f16 ci: make Windy Git CI permanent for the private repos
All checks were successful
check / gate (push) Successful in 21s
canary / probe (push) Successful in 6s
- Four runners x capacity 1 instead of one x capacity 4. Concurrent jobs in
  one act_runner share /root/.cache/act; a refresh racing a copy killed 3 of
  windy-chat's ~20 jobs at setup-node (lstat ... no such file). Separate
  processes have separate caches. Same parallelism, same capped dind.
- Behavioral tests for pr_status_bridge (latest verdict wins, no reposting,
  skipped never painted green, fork PRs never run, pagination, PR lifecycle).
- import_from_github.py reads IMPORT_GITEA_URL, not GITEA_BASE_URL: sourcing
  the deploy .env pointed it at http://gitea:3000 and it died on DNS after the
  mirror it replaces had already been deleted.
- CUTOVER.md: the private-repo CI path, onboarding steps, and the
  /actions/tasks-hides-queued-runs trap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:18:04 -04:00
1b09b9b0d3 ci: sync windy-chat + windy-mail, bridge PR CI verdicts back to GitHub
All checks were successful
check / gate (push) Successful in 20s
canary / probe (push) Successful in 6s
GitHub Actions can't run on the private platform repos. Windy Git already
has their code and a working runner, so:

- windy-chat and windy-mail were read-only pull mirrors (which can never
  run Actions); they are now writable, deploy.yml/build-image.yml disabled,
  and synced from GitHub like the others.
- scripts/pr_status_bridge.py mirrors open same-repo GitHub PRs into Windy
  Git (so pull_request workflows fire) and posts each job's result back as
  a GitHub commit status (windy-git/<workflow>/<job>) on PR heads and the
  default-branch head. Fork PRs are never run. Runs after every sync.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 02:13:21 -04:00
390c1e7479 ops: move tunnel metrics to 2001, sync windy-git into itself
Some checks failed
check / gate (push) Successful in 19s
canary / probe (push) Failing after 7s
windygit-tunnel had crash-looped ~91k times: another project's
cornercall-tunnel holds 127.0.0.1:2000, and cloudflared exits when it
cannot bind its metrics port. Ingress only survived because a stray
cloudflared.service ran the same config. That unit is now disabled and
/etc/cloudflared/config.yml uses metrics 127.0.0.1:2001.

Also add windy-git to the GitHub->Windy Git sync list; its self-hosted
copy was stuck 3 commits behind (only check + canary workflows, no
deploys, so syncing is safe).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 01:29:57 -04:00
2b30b0ac99 docs: record the runner bump, the act cache flake, and the remaining reds
Second root cause found and fixed: act_runner 0.2.11 predates
`runs.using: node24`, so any repo on actions/checkout@v5 died before its first
step. Bumped to 0.6.1; Windy-Clone went 4/4 red to 4/4 green.

Also upgrades the act-cache note from "watch item" to a confirmed job failure
(lstat on a vanished file mid-tar), with the wipe command and the annotated-tag
dead end that looks like a wrong checkout but isn't.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 19:09:26 -04:00
cd7dd9b7ae ci: bump act_runner 0.2.11 -> 0.6.1 for node24 action support
0.2.11's bundled act only knows runs.using node12/node16/node20, so any repo
pinning a current action major (actions/checkout@v5, actions/setup-python@v6)
fails before its first step with "The runs.using key in action.yml must be one
of: [...], got node24". Windy-Clone is how this surfaced.

Verified node24 is absent from the 0.2.11 binary and present in 0.6.1, and that
every key in deploy/runner/config.yaml still exists in 0.6.1's schema.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 19:00:07 -04:00
9fc27eabd6 docs: record the real CI root cause — setup-uv resolves uv via the forge API
The localhost->postgres fix was correct but was never what failed these jobs;
they died at step 2. setup-uv v4+ resolves "latest" through GITHUB_API_URL,
which act_runner points at our own forge, so it 404s and every later step is
skipped by success(). Pinned an explicit uv version across 11 repos.

Also records the diagnosis traps that cost the most time: Gitea's job status
enum (1=success, 2=failure), act misattributing the error to the previous step,
the jobs-log API needing a repo-matched id, and Secure cookies defeating a
loopback curl login.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 17:57:04 -04:00
Grant Whitmer
db055a1922 docs: refresh the turnover prompt for the actual next task
Some checks failed
check / gate (push) Successful in 18s
canary / probe (push) Failing after 6s
Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 17:31:23 -04:00
Grant Whitmer
86326ca6a7 docs: record three-repo CI fix results — 1 of 3 verified
All checks were successful
check / gate (push) Successful in 19s
windy-registry's postgres integration went failure -> success, proving the fix.
windy-mind and WindyCloud still fail for a cause I could not determine: the
jobs API returns 'job not found' for the ids the runs report, so logs were not
retrievable that way. Next session should read them from the Gitea web UI.

Records the trap that the three repos did NOT share one pattern — a naive
localhost->postgres swap would have left WindyCloud on port 15432.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 17:30:44 -04:00
Grant Whitmer
e0d5d118be docs: turnover for a fresh session
All checks were successful
check / gate (push) Successful in 35s
canary / probe (push) Successful in 8s
State, the immediate task (three-repo localhost->service-name CI fix), the traps
already paid for, and a copy-paste prompt.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 15:00:28 -04:00
Grant Whitmer
01e36155a3 ci: remove the service-networking probe; scope the python-pin test
Some checks failed
check / gate (push) Successful in 20s
canary / probe (push) Failing after 1m33s
The probe's own log was never retrievable through the jobs API, but the
question it asked was answered better by a direct comparison of two real
workflows on the same runner and image:

  windy-git gate       @postgres:5432   -> passes its migration round-trip
  eternitas migrations @localhost:5432  -> failed

Also scopes test_g73 to workflows that actually run Python. It failed the probe
for not pinning a version when the probe only shelled out to psql — the test
being wrong rather than the workflow.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 14:41:14 -04:00
Grant Whitmer
fe8f84bbdf ci: probe how service containers are addressed on this runner
Some checks failed
check / gate (push) Failing after 17s
canary / probe (push) Successful in 16s
Several migrated workflows hardcode postgres at localhost:5432, which is
correct on GitHub-hosted runners (services are port-mapped to the VM) and
suspect on Gitea Actions (the job runs IN a container, so localhost is the job).
Prove which form works before rewriting anyone's workflow.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 13:14:24 -04:00
Grant Whitmer
eae1bff50b G2.3: Windy Git branding — and get it out of one host's disk into the repo
All checks were successful
check / gate (push) Successful in 19s
canary / probe (push) Successful in 8s
The front end was 100% stock Gitea: green teacup, "Gitea: Git with a cup of
tea", "A painless, self-hosted Git service". G2.3 was specified in the plan with
an acceptance test and never executed, and nothing enforced it.

Now: Windy Git name, wind-mark logo, brand-blue accent, and a landing page that
says what this actually is. Uses Gitea's SUPPORTED surface (custom templates +
public assets) so upstream upgrades keep arriving — no source modified (D-2/I-1).

Two traps this cost, both now documented and tested:

1. GITEA__DEFAULT__APP_NAME does not work. Gitea reads APP_NAME from the TOP
   LEVEL of app.ini; the env var created a literal [default] section that Gitea
   ignores, so the installer's stock APP_NAME kept winning while the config
   looked correct. The env-to-ini pass also APPENDED a second APP_NAME rather
   than replacing the first — a new variant of the documented G4A.3 trap.

2. Cloudflare caches /assets/* for 6h and no token in this stack can purge, so
   the new logo and CSS were invisible while being correct at origin. Brand
   assets now carry a VERSION IN THE FILENAME; bump it on every change.

Committed with an idempotent apply.sh, because applying it straight to Veron's
disk first was itself the config-drift trap this project documents: a rebuild
would have silently reverted to stock Gitea.

85 tests green.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-14 09:08:48 -04:00
Grant Whitmer
e7dee39151 throttle: stop claiming to limit pushes we cannot see
All checks were successful
check / gate (push) Successful in 19s
canary / probe (push) Successful in 9s
I reintroduced the exact defect I had just criticised. ACTION_BASE listed
"push" and "push.force", but git push goes straight to Gitea over HTTPS and
never touches this API — so nothing records a push, a count would be zero
forever, and enforce() would look up a limit, count nothing, and allow
everything. A silent no-op wearing the costume of a control, made worse by a
config name that implies the protection exists.

Split into ACTION_BASE (actually enforced: repo.create, grant.create) and
NOT_ENFORCED_HERE (push, push.force) with the reason and the remedy written
down: enforcing push velocity needs a Gitea-side pre-receive or push webhook
reporting into agent_actions.

enforce("push") now raises rather than silently allowing, and a test asserts the
two sets stay disjoint.

Found by auditing whether the auth fix could be walked around — every
/api/v1/repos/* route does require a caller, and the only unauthenticated
endpoints are /health, /version and the HMAC-verified webhook.

83 tests green.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:33:30 -04:00
Grant Whitmer
25368547cb docs: runbook — never pull -q, never force-push a deployed branch
All checks were successful
check / gate (push) Successful in 21s
canary / probe (push) Successful in 11s
Two deploy traps paid for on 2026-08-14: 'git pull -q' hid a divergent-branch
error so a deploy ran against stale code while reporting success, and the
divergence came from amending a commit a deploy checkout already tracked.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:28:57 -04:00
Grant Whitmer
c60bfb2b89 I-12: fail the build when COMMIT_SHA is empty
All checks were successful
check / gate (push) Successful in 19s
/version went null after a deploy — the exact "service cannot name its own
commit" defect this project was built to prevent, caught by its own honesty
check.

Cause: the sed replaced "" with "" (a no-op when COMMIT_SHA is empty) and the
grep then matched that same empty string, so the guard verified nothing. A build
with no COMMIT_SHA passed and shipped a container reporting commit_sha: null.

Now the build fails loudly instead.

Second cause of the stale deploy, and it was mine: an earlier `git commit
--amend` + force-push rewrote history the Veron deploy checkout was already
sitting on, leaving it divergent so `git pull -q` failed SILENTLY (-q hid
"Need to specify how to reconcile divergent branches"). Two lessons: do not
force-push a branch a deploy checkout tracks, and do not pull with -q in a
deploy script.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:27:38 -04:00
Grant Whitmer
a0ed4a5ec0 SECURITY: make EPT routing independent of signature well-formedness
All checks were successful
check / gate (push) Successful in 20s
looks_like_ept used jwt.get_unverified_header, which validates the WHOLE token
and therefore rejects anything with a malformed signature segment. Routing
consequently depended on signature well-formedness: an EPT-shaped token with a
bad signature fell through to the HUMAN path, where it was refused for the wrong
reason and — with require_verified_jwt off (dev) — could have been read as a
human identity via its `sub` claim.

Now the header segment is decoded directly, so routing depends only on what the
token CLAIMS to be; whether it is authentic remains verify_ept's job.

Also routes alg:none to the EPT verifier regardless of typ, since a `none`
token is never valid for any caller. Both forged shapes now return 401
ept_invalid — the honest code — instead of 503 "feature not ready".

Found by noticing a forged EPT returned 503 where the verifier should have
answered 401, rather than accepting "it was refused, close enough".

80 tests green.

Co-Authored-By: Claude (Fable 5) <noreply@anthropic.com>
2026-08-13 23:23:50 -04:00