7 Commits

Author SHA1 Message Date
Grant Whitmer
d8deffe4db SECURITY: close the agent-auth bypass — trust is not authentication
All checks were successful
check / gate (push) Successful in 18s
canary / probe (push) Successful in 11s
Verified live 2026-08-13: a forged 'alg:none' token naming a passport lifted
from the logs returned HTTP 200 as that agent. The agent path read the passport
without verifying the EPT signature, asked Eternitas 'is this passport
reputable?', and seated the caller on a yes. That answers reputation, not
possession — anyone who knows a passport number could impersonate that agent on
the public API.

The human path already failed closed for exactly this reason
(require_verified_jwt). The gate was on the wrong path: it sat AFTER the agent
branch returned. The agent path now fails closed too, BEFORE the trust lookup,
so a forged token never even reaches Eternitas. Reopens automatically when the
ES256/JWKS verifier (G3.2/G9.1) exists and this gate consults it.

Adds BEHAVIORAL tests (not string-grep): a forged alg:none token exercised
through the real get_caller must raise, not authenticate. This is the test that
would have caught the bypass; the suite had 86 source-string assertions and
zero that ran the auth decision.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 22:49:45 -04:00
Grant Whitmer
83047def94 docs: whole account on Windy Git — 143 repos, 1.58 GB, two tiers
131 read-only mirrors (cannot run Actions, zero deploy risk) + 12 writable with
CI and deploys disabled. Total size matches the measured GitHub archive exactly,
which is the confirmation the copy is complete.

Splits the two concerns cleanly: having a copy is safe and should cover
everything now; running code needs judgement and happens per repo.

windy-pro IS included as a mirror — the G11.5 caution is about making it
writable while six checkouts disagree on HEAD, not about holding a read-only
copy. The DR copy is now complete.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:49:48 -04:00
Grant Whitmer
711c47d8c9 fix: the --all-as-mirrors flag itself was never added
The function landed but the argparse anchor did not match the real formatting,
so the command existed and could not be invoked. Caught by running it rather
than assuming the patch applied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:43:22 -04:00
Grant Whitmer
ae6ca8b9d2 G11.3: bulk DR copy — every repo as a read-only mirror
Grant's plan: clone the whole account, let it circulate, reverse direction later
when things are clean. Right plan, with one change that matters.

Sampling 40 repos found 18 carrying deploy/release/publish workflows that
trigger on push: — roughly 63 across the account. Importing those writable with
Actions enabled would arm sixty-odd production deploy triggers on Veron 1, each
needing disarming by hand.

So the bulk goes in as READ-ONLY pull mirrors. A mirror cannot run Actions at
all, so this carries zero deploy risk, and Gitea syncs them itself with no
script and no timer. What you get is a complete, current second copy of the
whole account — the disaster-recovery half — with none of the execution risk.

Converting one to writable + CI stays a deliberate per-repo act: re-import,
review its workflows, disable the deploying ones. That judgement belongs at the
moment you want CI on that repo, not in bulk sixty times by accident.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:40:52 -04:00
Grant Whitmer
89723c6ebd safety: disable deploy workflows on Windy Git before they can fire
Six workflows deploy to production on push:. Windy Git now has a working
runner, so the next synced commit to main would have attempted a production
deploy FROM VERON 1. Their secrets are unset here so they would have failed —
but loudly, on every push, with any pre-SSH step still running.

All six now disabled_manually. Tests, lints and migration checks stay active:
they need no secrets, which is exactly why Phase 1 delivers CI value with
nothing to configure.

Same class of mistake as the push-mirror direction, caught before firing this
time rather than after.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:30:56 -04:00
Grant Whitmer
b2f00821d7 docs: replace the cutover plan with the phased one that matches reality
Phase 1 requires nothing from anyone: agents keep pushing to GitHub, a timer
syncs GitHub -> Windy Git every 15 minutes, CI runs on Veron against current
code. Phase 2 flips one repo at a time, only when that repo is idle.

Records the direction mistake honestly: the source of truth is wherever people
are actually typing, not wherever the plan says it should be.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:14:21 -04:00
Grant Whitmer
51acf9f86e URGENT FIX: reverse the sync — GitHub is the source of truth, not Windy Git
I migrated nine repos writable with push-mirrors pointed AT GitHub. That was
wrong for the actual situation: a dozen agent sessions on the Mac mini are
pushing to GitHub continuously, so GitHub is where the live work is.

A push-mirror force-updates refs. On its 8-hour timer it would have pushed
Windy Git's stale copy over live work — silently, no conflict, nothing to
notice. Removed all nine before the first timer fired; verified no GitHub repo
had been touched (latest push predated the mirrors).

Replaced with the correct Phase 1 direction:

  agents --push--> GitHub --sync--> Windy Git --> CI on Veron

It requires NOTHING from anyone. No remote changes, no coordination, no
'everybody stop pushing'. Agents keep working exactly as they are and CI starts
running on 24 cores.

Windy Git is force-updated on purpose: in Phase 1 it holds nothing anyone
depends on, so GitHub always wins and there is no merge to reconcile.

Phase 2 is per-repo and only when that repo is idle. Never a big-bang cutover
across a dozen live sessions.

Fetches +refs/heads/* and tags explicitly rather than --mirror, which would drag
GitHub's refs/pull/* that Gitea rejects and bury the real errors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 22:11:34 -04:00
5 changed files with 310 additions and 48 deletions

View File

@@ -159,10 +159,37 @@ async def get_caller(
token = authorization.split(" ", 1)[1].strip()
# --- agent (Eternitas EPT) --------------------------------------------
# An EPT names its passport; the trust API is the authority on whether that
# passport may act. We never read a band out of the token itself.
passport = _unverified_claim(token, "passport") or _unverified_claim(token, "sub_passport")
if passport:
# ⚠️ SECURITY — trust is not authentication.
#
# A trust lookup answers "is this passport reputable?". It does NOT
# answer "does this caller actually hold this passport?". Skipping the
# second question is an authentication bypass: anyone who knows a
# passport number (they appear in logs, the lockbox and revocation
# messages) could present an UNSIGNED token naming it and be treated as
# that agent. Verified live 2026-08-13 — a forged `alg:none` token
# returned HTTP 200.
#
# ES256/JWKS verification of the EPT against Eternitas is not built yet
# (the G3.2/G9.1 verifier). Until it is, the agent path FAILS CLOSED in
# production — exactly as the human path below already does. This is not
# a downgrade of the "agents are citizens" design; it is refusing to
# seat a citizen whose ID we cannot yet check. It reopens automatically
# the moment `verify_ept_signature` exists and this gate consults it.
if settings.is_production and settings.require_verified_jwt:
raise RepairPointer(
status_code=503,
code="agent_signin_not_ready",
speak="Helper sign-in isn't switched on yet. Nothing you have is affected.",
machine_cause=(
"EPT signature verification (G3.2/G9.1) is not implemented; "
"refusing an unverified agent token in production. A trust "
"lookup proves reputation, not possession."
),
remediation_tool=None,
)
band, actions = await resolve_passport(settings, passport)
if band.lower() == "untrusted":
raise RepairPointer(

View File

@@ -10,12 +10,16 @@ happened somewhere in this ecosystem and cost real time.
from __future__ import annotations
import base64 as _b64
import json as _json
import re
import subprocess
import sys
import types as _types
from pathlib import Path
import pytest
import pytest as _pytest
ROOT = Path(__file__).resolve().parents[2]
@@ -633,3 +637,53 @@ def test_g09_backup_fails_loudly():
src = (ROOT / "scripts" / "backup.sh").read_text()
assert "COMPLETED WITH FAILURES" in src
assert "refusing to report a backup that did not happen" in src
# --------------------------------------------------------------------------
# SECURITY (behavioral, not string-grep): the agent path must not authenticate
# an unverified token. Regression guard for the 2026-08-13 forged-token bypass.
# --------------------------------------------------------------------------
def _forged_bearer(passport: str) -> str:
def seg(d):
return _b64.urlsafe_b64encode(_json.dumps(d).encode()).rstrip(b"=").decode()
return f"{seg({'alg':'none','typ':'JWT'})}.{seg({'passport':passport})}.not-a-signature"
def _fake_request(settings):
app = _types.SimpleNamespace(state=_types.SimpleNamespace(settings=settings))
return _types.SimpleNamespace(app=app)
@_pytest.mark.asyncio
async def test_security_forged_agent_token_is_refused_in_production():
"""A token with alg:none naming a real passport must NOT authenticate.
This is the exploit that returned HTTP 200 on 2026-08-13, exercised through
the real get_caller decision rather than by grepping for a string."""
from api.app.auth import get_caller
from api.app.config import Settings
from api.app.errors import RepairPointer
settings = Settings(environment="production", require_verified_jwt=True,
eternitas_platform_api_key="x", eternitas_base_url="https://api.eternitas.ai")
req = _fake_request(settings)
with _pytest.raises(RepairPointer) as exc:
await get_caller(req, authorization=f"Bearer {_forged_bearer('ET26-1EF9-VJAN')}",
x_service_token=None)
# Must be refused, and must be refused BEFORE any trust lookup could seat it.
assert exc.value.status_code in (401, 503)
assert exc.value.code == "agent_signin_not_ready"
@_pytest.mark.asyncio
async def test_security_no_bearer_is_still_401():
from api.app.auth import get_caller
from api.app.config import Settings
from api.app.errors import RepairPointer
req = _fake_request(Settings(environment="production"))
with _pytest.raises(RepairPointer) as exc:
await get_caller(req, authorization=None, x_service_token=None)
assert exc.value.status_code == 401

View File

@@ -1,63 +1,114 @@
# Windy Git is the daily driver — 2026-08-13
# Migration plan — GitHub first, Windy Git second, flip per repo
Grant's call, 2026-08-12: **push to Windy Git; GitHub is the second copy.**
**Superseded the 2026-08-13 "daily driver" cutover, which was premature.**
you ──push──▶ Windy Git (Veron 1) ──▶ CI on 24 cores
│
└──push-mirror on every commit──▶ GitHub
## What went wrong, recorded so it is not repeated
## 🔴 The one rule this creates
Nine repos were migrated writable with **push-mirrors pointed at GitHub**. At
the same time a dozen agent sessions on the Mac mini were pushing to GitHub
continuously — so GitHub, not Windy Git, was where the live work actually was.
**Do not push directly to GitHub for a migrated repo.**
A push-mirror force-updates refs. On its 8-hour timer it would have pushed Windy
Git's stale copy **over live work, silently, with no conflict to notice.**
A push mirror makes GitHub match Windy Git. Anything committed straight to
GitHub is **overwritten on the next sync**, silently, with no conflict and no
warning. That is the cost of having one writer, and one writer is the point —
two writers with no reconciliation is how you lose work you thought was saved.
All nine mirrors were removed before the first timer fired, and every GitHub
repo was verified untouched (latest push predated the mirrors). **No work was
lost.** The mistake was direction, and the lesson is: *the source of truth is
wherever people are actually typing, not wherever the plan says it should be.*
If you must hotfix on GitHub: push there, then immediately pull that commit into
Windy Git *before* anything triggers a mirror sync. Better: don't.
## Phase 1 — now. Nothing changes for anyone.
## Migrated (9)
Mac mini agents ──push──▶ GitHub ──sync every 15 min──▶ Windy Git ──▶ CI on Veron
`windy-calendar` · `windy-search` · `windy-registry` · `Windy-Clone` ·
`WindyCloud` · `windy-cloud-sites` · `windy-mind` · `eternitas` · `windy-agent`
- **You do not have to tell your agents anything.** No remote changes, no
coordination, no "everyone stop pushing." They keep working exactly as they
are.
- `windygit-sync.timer` runs `scripts/sync_from_github.sh` every 15 minutes.
- Windy Git is **force-updated** on purpose: it holds nothing anyone depends on,
so GitHub always wins and there is **no merge to reconcile**. That is the
whole point of not flipping until a repo is quiet.
- CI runs on Veron 1 against current code, on the 36 workflows that already say
`runs-on: [self-hosted, linux, x64]`.
All writable (`mirror=false`), all push-mirroring to GitHub with
`sync_on_commit: true`. Clone from `https://app.windygit.com/windyadmin/<repo>.git`.
Tracked repos live in `REPOS` in the script (currently 9 of 141).
**Not migrated on purpose:** `windy-pro`. Six checkouts exist, the build counter
has forked three ways (main 12 / overnight 34 / wave-44 56), and two sessions
recorded different HEADs hours apart. Resolve which is current and write it
down first (G11.5). The import script refuses it by name.
## Phase 2 — later, one repo at a time, only when that repo is idle
## CI
For a single repo, when nobody is mid-work on it:
The runner advertises `veron-1`, `linux-x64`, `self-hosted`, `linux`, `x64`.
**36 of 36 active workflows in the fleet already say
`runs-on: [self-hosted, linux, x64]`** — they were written for the self-hosted
runners that died when the repos went private, so they run **as-is, unedited**.
1. Remove it from `REPOS` in `sync_from_github.sh` — **first**, or the sync will
fight its authors and win.
2. Point that repo's sessions at Windy Git:
`git remote set-url origin https://app.windygit.com/windyadmin/<repo>.git`
3. Add a push-mirror back to GitHub with `sync_on_commit: true`, so GitHub stays
a current second copy.
Proven: `windy-calendar`'s existing `.github/workflows/ci.yml` ran on Veron 1
and reported success with no changes.
**Never flip more than one repo at a time, and never while an agent is working
in it.** A dozen parallel sessions is exactly the situation where a big-bang
cutover produces the dirty-branch mess this plan exists to avoid.
**Per-repo secrets are not imported.** A repo whose CI needs a database URL or
an API key will fail until those are set in its Gitea repo settings. Set them as
each repo needs them, not speculatively.
## The whole account is on Windy Git — in two tiers
143 repos · 1.58 GB · 967 GB free (matches the GitHub archive exactly)
| tier | count | writable | runs CI | deploy risk |
|---|---|---|---|---|
| **read-only mirrors** | 131 | no | **no** | **none** |
| **writable + CI** | 12 | yes | yes | deploys disabled |
**Why the bulk is mirrors, and why that is the safety decision:** sampling 40
repos found **18 carrying deploy / release / publish workflows that trigger on
`push:`** — roughly 63 across the account. Importing those writable with Actions
enabled would have armed sixty-odd production deploy triggers on Veron 1, each
needing disarming by hand. **A pull mirror cannot run Actions at all**, so the
bulk import carries zero execution risk and Gitea syncs it with no script and no
timer.
That splits the two things cleanly: **having a copy** (safe, do it for
everything, now) and **running code** (needs judgement, do it per repo,
deliberately).
Seven repos are empty here because they are empty on GitHub — 0 KB upstream,
verified. Not failed imports.
`windy-pro` **is** present, as a mirror. That is safe: the G11.5 caution is
about making it *writable* while six checkouts and a three-way-forked build
counter disagree on HEAD. A read-only copy of whatever GitHub currently has
carries none of that risk — and it means the DR copy is complete.
## Promoting a mirror to writable + CI
Per repo, deliberately, when that repo is quiet:
1. delete the mirror, re-import with `mirror=false`
2. **review its workflows and disable every deploying one** (see the section
above — this is the step that matters)
3. add it to `REPOS` in `sync_from_github.sh` so it tracks GitHub
4. later, when it flips to Windy-Git-first: remove it from `REPOS` *first*,
repoint its sessions, add a push-mirror back to GitHub
## ⚠️ Deploy workflows are DISABLED on Windy Git, deliberately
Six workflows fire on `push:` and deploy to production:
`windy-registry`, `Windy-Clone`, `WindyCloud`, `windy-mind`, `eternitas`
(`deploy.yml`) and `windy-agent` (`release.yml`).
Windy Git now has a working runner, so the next synced commit to `main` would
have attempted a **production deploy from Veron 1**. Their secrets
(`DEPLOY_HOST` / `DEPLOY_KEY` / `VPS_SSH_KEY`) are unset here, so they would
have failed — but they would have failed *loudly on every push*, and any step
before the SSH step would still have run.
All six are now `disabled_manually`. Tests, lints and migration checks stay
**active** — those need no secrets at all, which is why Phase 1 delivers real CI
value immediately.
**Before re-enabling any deploy workflow here, decide deliberately whether
production should be deployable from Windy Git at all.** Kit 0 deploys are
currently manual runbooks; that is a feature, not a gap.
## Backups
Nightly `windygit-backup.timer` at 04:17 (`Persistent=true`, so a window missed
while the workstation is off is caught up rather than skipped). Every repo is
bundled `--all`, verified, and uploaded to R2 with the `windgit` schema.
**Restore is rehearsed, not assumed:** a bundle was pulled back from R2, cloned,
and its HEAD matched live `origin/main` exactly.
## Verify the loop yourself
```bash
git clone https://app.windygit.com/windyadmin/windy-calendar.git
cd windy-calendar && git commit --allow-empty -m "probe" && git push
# CI runs on Veron 1; GitHub receives the commit within ~20s
```
`windygit-backup.timer`, nightly 04:17, `git bundle --all` + verify + `windgit`
schema dump to R2, 30-day retention. **Restore rehearsed:** a bundle was pulled
from R2, cloned, and its HEAD matched live `origin/main` exactly.

View File

@@ -93,6 +93,55 @@ def _api(method: str, path: str, body: dict | None = None) -> tuple[int, dict]:
return e.code, {"message": raw[:300]}
def all_repo_names() -> list[str]:
out = subprocess.run(
["gh", "repo", "list", GITHUB_OWNER, "--limit", "300",
"--json", "name,isArchived"],
capture_output=True, text=True, check=True,
)
return sorted(r["name"] for r in json.loads(out.stdout) if not r["isArchived"])
def import_everything_as_mirrors() -> int:
"""Bulk DR copy: every repo, as a READ-ONLY pull mirror.
Mirrors are the right shape for the bulk, and the reason is safety rather
than tidiness. Sampling 40 repos found **18 carrying deploy / release /
publish workflows that trigger on `push:`** — roughly 63 across the account.
Importing those as writable repos with Actions enabled would arm sixty-odd
production deploy triggers on Veron 1, each of which would then have to be
disarmed by hand.
A pull mirror cannot run Actions at all, so the bulk import carries **zero**
deploy risk, and Gitea does the syncing itself with no script and no timer.
What you get is a complete, current, second copy of the whole account.
Converting one to a writable CI repo is then a deliberate per-repo act:
delete, re-import with `mirror=false`, review its workflows, disable the
deploying ones. That is the moment to make that judgement — not in bulk,
sixty times, by accident.
"""
names = all_repo_names()
existing = 0
done = 0
failed = []
print(f"{len(names)} active repos on GitHub. Importing missing ones as read-only mirrors.\n")
for n in names:
status, _ = _api("GET", f"/repos/{OWNER}/{n}")
if status == 200:
existing += 1
continue
if import_repo(n, mirror=True):
done += 1
else:
failed.append(n)
print(f"\n already present: {existing}")
print(f" newly mirrored: {done}")
if failed:
print(f" FAILED ({len(failed)}): {', '.join(failed[:10])}")
return 1 if failed else 0
def list_candidates() -> None:
"""Private repos whose CI cannot run on GitHub at all."""
out = subprocess.run(
@@ -158,12 +207,19 @@ def main() -> int:
"--mirror", action="store_true",
help="import read-only pull mirrors instead of writable repos. NOTE: mirrors CANNOT run CI.",
)
ap.add_argument(
"--all-as-mirrors", action="store_true",
help="bulk DR copy: every active repo on the account, read-only, no CI, zero deploy risk",
)
args = ap.parse_args()
if args.list_candidates:
list_candidates()
return 0
if args.all_as_mirrors:
return import_everything_as_mirrors()
targets = SAFE_ORDER if args.safe_batch else args.repos
if not targets:
ap.error("name a repo, or pass --safe-batch / --list-candidates")

74
scripts/sync_from_github.sh Executable file
View File

@@ -0,0 +1,74 @@
#!/usr/bin/env bash
# Phase 1 sync: GitHub is the source of truth, Windy Git follows.
#
# ── Why this direction, and why the other one was wrong ────────────────────
#
# On 2026-08-13 nine repos were migrated writable with push-mirrors pointed AT
# GitHub. That was premature: a dozen agent sessions on the Mac mini are pushing
# to GitHub continuously, so GitHub — not Windy Git — is where the current work
# actually lives. A push-mirror force-updates refs, so on its 8-hour timer it
# would have pushed Windy Git's stale copy over live work, silently, with no
# conflict to notice. The mirrors were removed before the first timer fired.
#
# This script is the correct Phase 1: pull from GitHub, push into Windy Git.
#
# Mac mini agents ──push──▶ GitHub ──this script──▶ Windy Git ──▶ CI on Veron
#
# **It requires nothing from anyone.** No remote changes, no coordination, no
# "everybody stop pushing for a minute." Agents keep working exactly as they are
# and CI starts running on 24 cores.
#
# Windy Git is force-updated on purpose. In Phase 1 it holds nothing anyone
# depends on, so GitHub always wins and there is no merge to reconcile — which
# is the entire point of not flipping direction until a repo is quiet.
#
# Phase 2, per repo, only when that repo is idle: point its agents at Windy Git,
# drop it from REPOS here, and add a push-mirror back to GitHub. One repo at a
# time. Never a big-bang cutover across a dozen live sessions.
set -uo pipefail
: "${GITHUB_TOKEN:?GITHUB_TOKEN required}"
: "${GITEA_ADMIN_TOKEN:?GITEA_ADMIN_TOKEN required}"
GH_OWNER="${GITHUB_OWNER:-sneakyfree}"
WG="${WG_HOST:-app.windygit.com}"
WG_OWNER="${WINDYGIT_OWNER:-windyadmin}"
WORK="${SYNC_WORK:-/srv/windygit/sync}"
FAILED=0
# Repos Windy Git tracks FROM GitHub. Remove a repo from this list at the moment
# it flips to Windy-Git-first, or the sync will fight its authors and win.
REPOS="${SYNC_REPOS:-windy-calendar windy-search windy-registry Windy-Clone WindyCloud windy-cloud-sites windy-mind eternitas windy-agent}"
mkdir -p "$WORK"
log() { printf '[sync %s] %s\n' "$(date -u +%H:%M:%SZ)" "$*"; }
for r in $REPOS; do
bare="$WORK/${r}.git"
if [[ ! -d "$bare" ]]; then
git clone --quiet --bare "https://x-access-token:${GITHUB_TOKEN}@github.com/${GH_OWNER}/${r}.git" "$bare" 2>/dev/null \
|| { log "FAILED initial clone of $r"; FAILED=1; continue; }
fi
# +refs/heads/* — branches only, deliberately.
#
# `--mirror` would also carry refs/pull/* (GitHub's read-only PR refs, which
# Gitea rejects) and every remote-tracking ref, turning a working sync into a
# wall of errors that hides the one that matters.
if ! git --git-dir="$bare" fetch --quiet --prune origin '+refs/heads/*:refs/heads/*' '+refs/tags/*:refs/tags/*' 2>/dev/null; then
log "FAILED fetch $r"; FAILED=1; continue
fi
before="$(git --git-dir="$bare" rev-parse HEAD 2>/dev/null || echo none)"
if git --git-dir="$bare" push --quiet --force \
"https://${WG_OWNER}:${GITEA_ADMIN_TOKEN}@${WG}/${WG_OWNER}/${r}.git" \
'+refs/heads/*:refs/heads/*' '+refs/tags/*:refs/tags/*' 2>/dev/null; then
log "$r ok (${before:0:7})"
else
log "FAILED push $r -> windy git"; FAILED=1
fi
done
[[ "$FAILED" -ne 0 ]] && { log "COMPLETED WITH FAILURES"; exit 1; }
log "all repos in step with GitHub"