The reviewer prompt was identical for every round, and the canon said
nothing about the scope of a repeat pass, so r2 re-derived the product
framing and re-checked acceptance criteria the fix never touched: the r2
pass on #150 cost a full pipeline run over one line in a test fixture.
From the second cycle on, the subject is the delta against the SHA the
previous verdict was given on: each earlier finding must be shown closed
by a line of code or text, only the criteria the delta can reach are
re-verified, and whatever is carried over is listed with the round and SHA
it came from. Cheap gates still run every round.
The scope shrinks, the strictness does not. A fix can break a criterion an
earlier round accepted — that is how regression #102 happened — so the
boundary is the findings plus everything the delta can reach, and a
non-local delta (a rebase onto a moved dev, a behaviour contract change, a
new subsystem) still gets the full pass.
Issue: #214
User-Visible: no
The reviewer prompt was identical for every round, and the canon said
nothing about the scope of a repeat pass, so r2 re-derived the product
framing and re-checked acceptance criteria the fix never touched: the r2
pass on #150 cost a full pipeline run over one line in a test fixture.
From the second cycle on, the subject is the delta against the SHA the
previous verdict was given on: each earlier finding must be shown closed
by a line of code or text, only the criteria the delta can reach are
re-verified, and whatever is carried over is listed with the round and SHA
it came from. Cheap gates still run every round.
The scope shrinks, the strictness does not. A fix can break a criterion an
earlier round accepted — that is how regression #102 happened — so the
boundary is the findings plus everything the delta can reach, and a
non-local delta (a rebase onto a moved dev, a behaviour contract change, a
new subsystem) still gets the full pass.
Issue: #214
User-Visible: no
Every push to dev paid for the full browser trio and the backend suite,
including commits that touch only documentation, workflows or process
scripts — the bundle and the harness were byte-identical, so the runs
proved nothing new. On 2026-08-19 alone that was roughly six pushes at
about seven minutes each.
The reuse key per heavy job is sourceFingerprint (src, demo fixtures,
golden scenarios, build manifests) plus that job's own harness: smoke
takes demo/smoke_*.mjs, golden takes demo/golden/** including baselines,
performance_smoke takes demo/performance/**, backend takes tests_backend
and the Python sources. A cache marker is written only by a successful run
of the same key, so a hit proves a job with identical inputs already
passed. scripts/** is deliberately outside every key: infrastructure work
edits it constantly and reuse would never fire.
This is not the path filter from the `changes` job, which stays disabled
on dev on purpose: there the scope is guessed from paths and "green" means
different things, here input equivalence is proven by a hash. And a
release candidate always bumps the version, which is part of the
fingerprint, so its keys are new by construction and the full gate set
still runs before every beta and release.
A waived job is announced with a notice and a run summary line rather than
skipped in silence, and the marker save tolerates a concurrent identical
run instead of reddening the job.
Issue: #208
User-Visible: no