A gauntlet lap takes a batch of work from "approved idea" to "proven end to end on a real running copy of the shop", without pushing anything until you say so. Agents build in parallel, reviewers check them, then the whole thing gets deployed to kali and hammered by the purchase matrix until it passes twice in a row.
Red dashed: a product bug goes back to its implementer, then a new wave. Grey: a test-setup or harness problem only needs the matrix run again. Green: two clean runs in a row end the lap.
controller
tools/lap1/the lap's test server
browser + test box
pwremote) against kalinight only
You say GO and pick the options. I write every decision into one binding file, lapN-constraints-and-rulings.md,
as numbered rulings (R4-SESSION, R4-READONLY, ...). Every agent reads it before it starts.
Tag today's master as the lap base (lap4/base). Every branch starts there and the lap never rebases,
so a master change mid-lap can't break a run halfway.
test/lap4-matrix), T2 login backend, T3 confirmation-mail link,
T4 storefront. All four were dispatched at once at about 21:00.Each implementer works alone on its branch and has to prove its work before it reports back:
lap-pytest.sh (macv2 if free, else kali).codex review against the base. Fix P1, and P2 on money, security, privacy or a broken flow. Anything else is a follow-up.A separate reviewer reads the branch against the rulings and either approves it or lists problems by severity.
R4-READONLY), sent to every task it touches.wave.sh: merge everything, deploy to kaliOne script builds a throwaway integration branch (base + every lap branch merged) and puts it on kali:
alembic upgrade heads), admin and storefront.The matrix is our Playwright suite of real purchases and flows (Stripe card, SCORE, Amazon Pay, coupons, MyPage),
run from macv2's browser against kali, on desktop Chrome and a Pixel 7 profile. run-matrix.sh reads kali's
settings first and warns if something would make a cell lie.
Running alongside:
| Kind of red | What happens | Lap 4 example |
|---|---|---|
| Product bug | Fix round to the implementer, then a new wave. | Codex caught a race in the duplicate e-mail check, fixed with a lock. |
| Harness | Fix the test helper, rerun. No redeploy. | A buy hit the 10-a-minute submit limit, so the helper now waits it out. |
| Setup / data | Fix kali's env or data, rerun. | The new 10-codes-a-day cap would block BL-10, so kali gets 1000. |
| Load | Find the cause, rerun on a quiet box. | BL-1 mobile missed while a CI job ran on kali, then passed. |
| Pre-existing | Same red on the base too: noted, not ours. | 13 order-pack unit tests already fail on master. |
Every slip becomes a rule so it can't happen twice: R4-RESEED (re-seed is always followed by the MIG convert),
R4-DETACH (long runs are detached with a "done" file so a timeout can't orphan a test).
The lap ends only when the whole matrix passes twice in a row on the same served build, with a fresh re-seed before each run. A red is allowed only if it has its own ticket.
docs/gauntlet/lapN-handoff.mdIf your answers ask for changes, that's round 2: back to step 3 with new rulings, same loop, same exit bar.
kali proves the code; the test server proves it with the real services (Stripe, SCORE, Amazon sandbox, SES, the WAF).
:latest), migrate, switch the services, card payments to 仮売上, seed, arm the night.| Rail | Why |
|---|---|
| Nothing pushed during the lap | Master and the devs never see half-done work. |
| No rebase mid-lap | A run that starts on one build finishes on it. |
| Tests only on kali / macv2 | This Mac stays free for the controller. |
| Disk guard (8 GB) | A full Pi breaks the DB and CI. |
| Rail | Why |
|---|---|
| Ledger line for every step | Anyone can see what ran, when, on which build. |
| Detached long runs + done files | A tool timeout can't leave a test half-running. |
| Passwords stay in your Keychain | The night job reads them; I never do. |
| Questions batched, not pinged | You answer once, in the handoff. |