How our gauntlet loop works

A gauntlet lap takes a batch of work from "approved idea" to "proven end to end on a real running copy of the shop", without pushing anything until you say so. Agents build in parallel, reviewers check them, then the whole thing gets deployed to kali and hammered by the purchase matrix until it passes twice in a row.

Arcade version

Watch it run press Start, or click the picture

    You (サトル). Set the goal, answer the questions, say when to push.
    Claude, the controller. Plans, writes the rulings, sends the agents out, keeps the ledger, sorts every red.
    T1T1 harness. Writes the new matrix tests first.
    T2T2 login API (backend).
    T3T3 mail link (backend): the link in the confirmation mail.
    T4T4 storefront. The pages buyers see.
    ReviewReviewer. Never writes code, only checks a branch against the rulings.
    ExploreExplorer. Tries to break the app by hand in a real browser.
    Stops on a track. That agent's tests and commits. Red: a new test failing before the fix.
    Squares on macv2. Matrix cells, one real flow each (a purchase, a login, MyPage).

    Laps, rounds, waves, runs boxes inside boxes

    LapOne batch of work, from your GO to the handoff. Its own base tag, rulings, ledger and handoff. Lap 4 = the buyer login (V3-2459). Laps 1 to 3 were the LP purchase flow.
    RoundOne trip from "build" to "exit pair" inside the lap. Round 1 starts with your GO. Round 2 starts when your Q&A answers (or a big finding) ask for changes. Lap 4 had 2.
    WaveOne merge of every branch, deployed to kali. A real bug fixed by an agent means a new wave. Lap 4 had 3.
    RunOne pass of the matrix. A setup or test fix only needs another run, not a new wave. Two green runs in a row on one build = the exit pair, which ends the round.

    The Q&A part

    • Questions only you can answer are collected in the handoff (§0), never pinged one by one mid-lap.
    • Each question comes with my pick, so "keep" or "yes" is a full answer.
    • You answer them in one message, like "1 keep 2 keep 3 yes ...".
    • Each answer becomes a ruling. If any of them means new work, that is the next round.

    What still pings you right away

    • Passwords and logins (only you can type them).
    • Anything outside the safety rails, like pruning Docker on kali (as this morning).
    • Risks to money or live data.
    • Push day: it never starts without your word.

    The loop at a glance

    1-2Rule + plan rulings, base tag, tasks 3Build agents in parallel 4Review reviewer per task 5Wave merge, deploy to kali 6Matrix + checks tests, gates, explore 7Triage reds bug, harness, or setup 8Exit pair 2 green runs in a row 9-11Handoff, test-server night, push your questions, real AWS test, PRs when you say so real bug: fix round + new wave all green setup or harness fix: just rerun

    Red dashed: a product bug goes back to its implementer, then a new wave. Grey: a test-setup or harness problem only needs the matrix run again. Green: two clean runs in a row end the lap.

    Who runs what four machines, each with one job

    This Mac

    controller

    • Me plus the agents (implementers, reviewers, auditor, explorer)
    • Git branches, the ledger, the scripts in tools/lap1/
    • Frontend tsc / vitest only. Never backend tests or containers.

    kali (Pi 5)

    the lap's test server

    • Serves backend, admin, storefront at one https address on the office LAN
    • Its own Postgres, Redis, a SCORE mock
    • Test slots for backend pytest
    • Also a CI runner for the devs' PRs (shared CPU)

    macv2

    browser + test box

    • Runs the Playwright browser (pwremote) against kali
    • Backend pytest when CI is not busy there
    • Exploratory browser over an ssh tunnel
    • Also a CI runner

    AWS test server

    night only

    • Off limits during the lap itself
    • Used on a test-server night: branch images in, matrix at 01/03/05 JST, master back by morning
    • Real Stripe, SCORE, Amazon sandbox, SES

    Step by step click a step to open it

    Before any code
    1
    Rule it and check it isn't already builtYour GO, the rulings file, the scope sweep
    ›

    You say GO and pick the options. I write every decision into one binding file, lapN-constraints-and-rulings.md, as numbered rulings (R4-SESSION, R4-READONLY, ...). Every agent reads it before it starts.

    • Scope sweep: before building, search master, open PRs and other devs' branches so we don't rebuild something that exists.
    • Standing rules ride along: commit identity, never push, never rebase mid-lap, tests only on kali/macv2, never prune Docker without you.
    Lap 4: option A (our own buyer session, no Firebase), cancel by e-mail only, kali only on night one. The sweep found four approved storefront PRs touching MyPage, so they were merged into the test build too.
    2
    Freeze a base and split into tasksBase tag, one branch per task, a plan
    ›

    Tag today's master as the lap base (lap4/base). Every branch starts there and the lap never rebases, so a master change mid-lap can't break a run halfway.

    • One task = one branch = one future PR. Tasks that don't touch each other run at the same time.
    • There's always a harness task: the new matrix tests, written first, which must fail on the old code.
    Lap 4: T1 harness (test/lap4-matrix), T2 login backend, T3 confirmation-mail link, T4 storefront. All four were dispatched at once at about 21:00.
    Build and review
    3
    Build in parallelOne implementer agent per task, in its own worktree
    ›

    Each implementer works alone on its branch and has to prove its work before it reports back:

    • Red then green: every new test fails on the old code and passes on the fix, with the output in the report.
    • Gates: ruff + ruff format, tsc, vitest/jest; backend pytest through lap-pytest.sh (macv2 if free, else kali).
    • Codex cap: one codex review against the base. Fix P1, and P2 on money, security, privacy or a broken flow. Anything else is a follow-up.
    • Commit locally with your name, no trailers. Nothing is pushed.
    Why the cap: without it, review rounds never end. Small issues go to the follow-up list instead of blocking the lap.
    4
    Review each taskA fresh reviewer agent, then fix rounds
    ›

    A separate reviewer reads the branch against the rulings and either approves it or lists problems by severity.

    • Critical / High / Medium go back to the same implementer as a fix round.
    • Low goes into the handoff's follow-up list.
    • A reviewer finding that changes the design becomes a new ruling (for example R4-READONLY), sent to every task it touches.
    • Questions only you can answer are batched into the handoff, not pinged one by one.
    Lap 4: the T4 review found that the buyer token sits where shop tags can read it, so the buyer session became read-only on the server for every write route (ruling R4-READONLY).
    Prove it on kali
    5
    Run a wavewave.sh: merge everything, deploy to kali
    ›

    One script builds a throwaway integration branch (base + every lap branch merged) and puts it on kali:

    • Rebuild both integrations (backend, storefront) from the branch lists.
    • Snapshot kali's DB, then deploy backend (with alembic upgrade heads), admin and storefront.
    • Recreate the test slots, seed the lap data, convert the MIG forms, warm the pages, verify served commits.
    • A disk guard stops it if kali has under 8 GB free.
    Why throwaway: the integration is never pushed. Each real branch stays clean, so a conflict stays one task wide.
    6
    Matrix and checks, side by sideNew cells first, then the whole matrix, plus three parallel checks
    ›

    The matrix is our Playwright suite of real purchases and flows (Stripe card, SCORE, Amazon Pay, coupons, MyPage), run from macv2's browser against kali, on desktop Chrome and a Pixel 7 profile. run-matrix.sh reads kali's settings first and warns if something would make a cell lie.

    • New cells first: just this lap's tests, so a basic break shows in minutes.
    • Then the whole matrix: about 18 minutes, 84 cells now.

    Running alongside:

    • Union gates: every test the lap touched, run on the merged tree and compared with the base. Only new reds count.
    • Cross-branch audit (after wave 1): one agent reads all branches together for things no single reviewer could see.
    • Exploratory testing: an agent drives a real browser on macv2 with charters ("try to read another buyer's orders").
    Lap 4: exploring confirmed that all 32 storefront write routes refuse a buyer session, and that forged and tampered tokens are rejected.
    7
    Triage every redIs it the product, the test, or the setup?
    ›
    Kind of redWhat happensLap 4 example
    Product bugFix round to the implementer, then a new wave.Codex caught a race in the duplicate e-mail check, fixed with a lock.
    HarnessFix the test helper, rerun. No redeploy.A buy hit the 10-a-minute submit limit, so the helper now waits it out.
    Setup / dataFix kali's env or data, rerun.The new 10-codes-a-day cap would block BL-10, so kali gets 1000.
    LoadFind the cause, rerun on a quiet box.BL-1 mobile missed while a CI job ran on kali, then passed.
    Pre-existingSame red on the base too: noted, not ours.13 order-pack unit tests already fail on master.

    Every slip becomes a rule so it can't happen twice: R4-RESEED (re-seed is always followed by the MIG convert), R4-DETACH (long runs are detached with a "done" file so a timeout can't orphan a test).

    8
    Exit pairTwo whole-matrix runs, back to back, both green
    ›

    The lap ends only when the whole matrix passes twice in a row on the same served build, with a fresh re-seed before each run. A red is allowed only if it has its own ticket.

    Why two: one green run can be luck. Two in a row on fresh data means it's repeatable.
    Lap 4: round 1 exited at 03:06 with 81 passed twice. Round 2 (wave 3) exited at 08:29 with 83 passed twice. The one skip is the real Amazon cell, which only runs in your Keychain job on the test server.
    After the lap
    9
    Handoffdocs/gauntlet/lapN-handoff.md
    ›
    • §0 your questions, each with my pick, so you answer in one go (lap 4: 7 questions, answered "1 keep, 2 keep, 3 yes ...").
    • What was delivered, results, branch heads, push order, the test-server night plan, chat/Jira drafts, follow-ups.

    If your answers ask for changes, that's round 2: back to step 3 with new rulings, same loop, same exit bar.

    10
    Test-server nightThe real AWS environment, overnight, while nobody's testing
    ›

    kali proves the code; the test server proves it with the real services (Stripe, SCORE, Amazon sandbox, SES, the WAF).

    • Evening (after 20:00 JST): TOALL message, build branch images with unique tags (never :latest), migrate, switch the services, card payments to 仮売上, seed, arm the night.
    • Night: the launchd job runs the matrix at 01, 03, 05 JST using your Keychain passwords (I never see them). I triage between runs.
    • Morning (before 07:30 JST): template edit removed, 即時売上 back, DB downgraded, master services back, branch revisions deleted, TOALL done. A backup watcher does the rollback by itself at 06:45 JST if I haven't.
    11
    Push dayOnly when you say so
    ›
    • Fetch master, rebase each lap branch onto it (the no-rebase rule ends here).
    • One PR per ticket (lap 4: backend, storefront, matrix), CI, merge, the lead dev deploys.
    • Jira 完了, the QA heads-up, notes to devs whose work overlaps (lap 4: a teammate's V3-2451).

    Safety rails that are always on

    RailWhy
    Nothing pushed during the lapMaster and the devs never see half-done work.
    No rebase mid-lapA run that starts on one build finishes on it.
    Tests only on kali / macv2This Mac stays free for the controller.
    Disk guard (8 GB)A full Pi breaks the DB and CI.
    RailWhy
    Ledger line for every stepAnyone can see what ran, when, on which build.
    Detached long runs + done filesA tool timeout can't leave a test half-running.
    Passwords stay in your KeychainThe night job reads them; I never do.
    Questions batched, not pingedYou answer once, in the handoff.

    Lap 4 as it actually ran

    1. 09-30 ~21:00 PHTGO. Four tasks dispatched in parallel.
    2. 09-30 nightReviews, fix rounds, wave 1: 81 passed. Audit + explore wave 1 found no Critical/High/Medium issues.
    3. 10-01 03:06Wave 2 exit pair green (81/81). Handoff with 7 questions.
    4. 10-01 morningYour answers became round 2: guest-only mail block, duplicate e-mail block, 10 codes a day.
    5. 10-01 07:40-07:56Wave 3 deployed, new cells 13/14 (one load flake), review H1 fixed on kali.
    6. 10-01 08:29Round 2 EXIT: 83 passed, twice. Union gates clean, harness fixes merged.
    7. 10-01 19:00 PHT (20:00 JST)Test-server night.
    8. When you say soPush day.