Ревью шло по ветке как есть, слияние делало ребейз: проверенный SHA и
слитый SHA были разными коммитами. Текстовое расхождение ловил конфликт,
смысловое git склеивал молча — так пришёл регресс #234. Заодно конфликт
обнаруживался после сорока минут работы ревьюера, хотя виден до них.
Новый шаг для этапа code, сразу после выбора ветки: потомок dev —
ничего; отстала и ребейзится — ребейз, push с --force-with-lease, ревью
приведённого состояния и запись о ребейзе в промпт (§7.2 требует полного
разбора); конфликт — возврат в S6-in-progress без запуска ревью.
Issue: #257
User-Visible: no
(cherry picked from commit 793a6486d8)
Three code-review rounds on #220 published a verdict and then failed the
run: the document never reached the branch, so the #171 guard refused
before the label step and neither the merge nor S8-merged happened. The
cause was structural. The document lived as an untracked file inside the
very checkout the reviewer edits while proving that a test can fail, and
restoring that tree — git checkout, git clean — deletes an untracked file.
Spec rounds survived only because they never mutate anything.
The reviewer now writes to REVIEW_DOC under RUNNER_TEMP, outside the
repository, and the publish step copies it into docs/reviews before
committing. Tree cleanup can no longer destroy the artefact, and the
reviewer no longer needs to touch docs/reviews at all.
Verified against a local git fixture on five paths: document outside the
repo with a mutated tree, nothing anywhere (loud failure), document only in
the working copy, document already committed by the reviewer, and a branch
that moved during the review.
Same file as main, byte for byte.
Issue: #220
User-Visible: no
The pipeline punished what it prescribed: after a failed merge it tells the
author to rebase and restore S7-code-review, and that attempt finished the
budget. On #225 a green code review with green CI ended in review-4.
Only yellow and red verdicts spend the budget now; a green verdict returned
nothing and consumes nothing. Attempts and cycles became separate
quantities: the attempt number names the review document, the limit
compares blocking cycles. The exhaustion comment lists what it counted, and
the guard reports a recount instead of stripping review-4 on its own.
Same file as main (41325a8), byte for byte.
Issue: #227
User-Visible: no
The reviewer prompt was identical for every round, and the canon said
nothing about the scope of a repeat pass, so r2 re-derived the product
framing and re-checked acceptance criteria the fix never touched: the r2
pass on #150 cost a full pipeline run over one line in a test fixture.
From the second cycle on, the subject is the delta against the SHA the
previous verdict was given on: each earlier finding must be shown closed
by a line of code or text, only the criteria the delta can reach are
re-verified, and whatever is carried over is listed with the round and SHA
it came from. Cheap gates still run every round.
The scope shrinks, the strictness does not. A fix can break a criterion an
earlier round accepted — that is how regression #102 happened — so the
boundary is the findings plus everything the delta can reach, and a
non-local delta (a rebase onto a moved dev, a behaviour contract change, a
new subsystem) still gets the full pass.
Issue: #214
User-Visible: no
Every push to dev paid for the full browser trio and the backend suite,
including commits that touch only documentation, workflows or process
scripts — the bundle and the harness were byte-identical, so the runs
proved nothing new. On 2026-08-19 alone that was roughly six pushes at
about seven minutes each.
The reuse key per heavy job is sourceFingerprint (src, demo fixtures,
golden scenarios, build manifests) plus that job's own harness: smoke
takes demo/smoke_*.mjs, golden takes demo/golden/** including baselines,
performance_smoke takes demo/performance/**, backend takes tests_backend
and the Python sources. A cache marker is written only by a successful run
of the same key, so a hit proves a job with identical inputs already
passed. scripts/** is deliberately outside every key: infrastructure work
edits it constantly and reuse would never fire.
This is not the path filter from the `changes` job, which stays disabled
on dev on purpose: there the scope is guessed from paths and "green" means
different things, here input equivalence is proven by a hash. And a
release candidate always bumps the version, which is part of the
fingerprint, so its keys are new by construction and the full gate set
still runs before every beta and release.
A waived job is announced with a notice and a run summary line rather than
skipped in silence, and the marker save tolerates a concurrent identical
run instead of reddening the job.
Issue: #208
User-Visible: no
performance_smoke burned nearly all of its 15-minute budget before the
benchmark even started, twice in a row: validate.yml had no browser cache
at all, so every browser job paid for a full `playwright install
--with-deps` — apt work the ubuntu-latest image makes redundant, with
unbounded retries against an unreachable azure mirror on top. For a
measuring job that is worse than lost minutes: the timing window competes
with package installation on the same runner.
#175 fixed this for the review pipeline but deliberately left the flag
here, reasoning that a prerelease gate values predictability over
minutes. That reasoning was wrong — the flag is what made the gate
unpredictable.
Browsers are now cached per package-lock hash in smoke, golden,
performance_smoke and the full performance run; installation happens only
on a cache miss and no longer touches apt. performance_smoke keeps
headroom for a cold cache at 20 minutes. If the image ever drops a
required library, Chromium fails to launch with a clear missing-libraries
error; that is the moment to bring the flag back.
Issue: #206
User-Visible: no
Filing and servicing a separate issue costs far more than fixing a small
problem in place — the owner's call of 2026-08-19 (#202). A Medium finding
inside the task's scope no longer becomes its own issue: with no High
findings the verdict is yellow, the author fixes it and the fix passes
another review cycle. Only an out-of-scope Medium is still filed
separately, because foreign scope is never patched from a task branch.
Applied to the canon (PROCESS.md), the reviewer prompt in process.yml and
AGENTS.md; the verdict format now writes "Medium: N -> in-task | #NN".
Issue: #202
User-Visible: no
On a Playwright cache miss the flag pulled Chromium's system libraries
through apt, spending minutes of the 45-minute review budget on packages
the ubuntu-latest image already ships — and the runner's retries against
the unreachable azure mirror made the step look hung on a live run. If the
image ever drops a required library, Chromium fails to launch with a clear
missing-libraries error; that is the moment to bring the flag back.
validate.yml keeps the flag deliberately: it is the prerelease gate, where
predictability is worth more than minutes.
Issue: #175
User-Visible: no
On #150 both spec-review verdicts survived only as issue comments: the
publish step found nothing staged, printed a warning, and exited zero, so
the label moved and the missing artifact went unnoticed until the next
review caught it (#171). A verdict without a document in docs/reviews/ now
fails the run before the label step, preserving the invariant that an
unchanged label means a failed run.
An empty working copy alone is not a failure: the reviewer occasionally
commits the document itself through its app token, bypassing this step
(CODE-REVIEW-150-r1, committer GitHub), so the branch is checked first. A
postcondition verifies the exact expected filename reached the branch, and
the rebase-conflict path no longer exits zero either.
Issue: #171
User-Visible: no
Issue #150 reached a green verdict and then hit two pipeline defects at once.
The review document push came back 403 as github-actions[bot]: the PAT had
died, and checkout's persisted credential quietly took its place — a masked
actor instead of a loud failure. Credentials are no longer persisted, and the
token is now proven alive before the review starts, not after forty minutes of
reviewer work.
Branch selection took the first match alphabetically, and with a spec-era
branch sitting next to the implementation branch that meant the stale one.
The freshest branch by commit date is chosen instead, with a warning naming
every candidate when more than one exists.
Verified against the real #150 branches: the fix branch wins, the warning
fires.
Issue: #114
User-Visible: no
Every push to every branch ran 128 browser smokes, 50 golden scenes, a Home
Assistant install and a performance pass — including a push that added one spec
file. The pipeline made such pushes routine: every spec revision and every
review document is a push to a task branch and used to cost the full suite.
A changes job classifies the push range; frontend, smoke, golden, performance
and backend now run only when their paths moved, and hacs and hassfest only for
manifests, translations or Python. provenance and process-gate always run — they
judge commits, not code.
The exception carries the design. On dev everything runs, always, unfiltered:
the beta gate accepts "green Validate at the exact SHA", and if the volume of a
run depends on the diff, green stops meaning one thing — a release candidate
touches manifests and changelogs, would skip the browser suites under filtering,
and a run with skipped jobs still concludes success. That would be the sixth
silent success of the week. Filters save time on task branches, where Validate is
an early signal and the real acceptance is the code review running gates itself.
A new branch with a zero before-sha is classified from the merge-base with dev,
not from the root of history.
Issue: #136
User-Visible: no
Five times in this project a green test meant nothing was checked. The
continuity smoke stayed green after the entire mechanism it guards was cut out.
The golden scene created to protect doorway light was empty — 1,177 warm pixels
against 107,119, all of them icons. The shadow smoke passed while no shadow was
drawn. Each time the test had been written alongside the code, went green at
once, and nobody ever asked whether it could go red.
The gate makes that question routine. Each mutant is a few lines of patch that
reproduce a known breakage, plus the name of the test that must fail on it. A
worktree is patched, the bundle rebuilt, the guard run — and a guard that stays
green fails the gate. Six mutants cover the holes documented in #85; the anchors
are exact strings from today's source, so the registry cannot silently drift —
a unit test that runs with the ordinary suite refuses a stale anchor.
The full run rebuilds the bundle per mutant, so it lives in its own workflow,
before a stable release and on a weekly schedule, not in Validate. The rules for
new tests are written at the top of docs/TESTING.md, and the sixth of them is
the cheapest: an assertion that reads back the property the code just set is
not written at all.
Issue: #85
User-Visible: no
The owner's report: the process works but every stage takes a long time even on
simple bugs. Two causes, and neither was the one that first comes to mind.
The reviewer ran everything regardless. On #89 it installed Chromium, ran all 127
smoke files and a full golden capture — right for a task rated 10/10 for
complexity, absurd for a bug about a room divider. Full suites are the pre-beta
gate; the review now runs typecheck, unit and build always, and smokes, golden,
pytest or performance only where the diff and the AC call for them. The price of
narrowing it is honesty: the reviewer must list which gates it ran, which it did
not, and why, so a skipped gate is a visible decision rather than a silent one.
The reviewer also built its own environment out of model turns, with no npm cache
and no browser cache, paid for from the same forty-five minutes. The workflow now
installs dependencies and Chromium as ordinary cached steps, after switching to
the task branch so the lockfile is the branch's own.
Second, ceremony did not scale down. The light track makes a spec cheap; the new
trivial track does without one — S2-analysis straight to S5-ready, no spec review,
AC in the issue body. It is deliberately hard to qualify for: a bug on one surface,
no new UX contract, no migration, no i18n, no perf or touch effect, three checkable
AC at most, and expected behaviour already on record. Nothing left to decide is the
criterion that holds the whole thing up, and it cannot be met by feeling sure.
Code review is never skipped on either track. It is what stands in for testing
here, so it is the one stage speed may not buy.
Issue: #127
Issue: #128
User-Visible: no
Issues labelled before the pipeline existed keep their spec straight in dev and
have no issue/NN branch. The publish step quietly exited zero for them, so the
verdict would arrive as a comment and the analysis behind it would be thrown
away — the fifth instance today of a step reporting success by doing nothing.
The document now goes wherever the spec itself lives: the task branch when there
is one, dev otherwise. Publishing also survives dev moving on while the review
ran, which takes up to forty-five minutes, by rebasing once before it gives up.
Four issues are waiting on this — #12, #30, #44 and #52 — each with a spec in dev,
a status label applied during the bulk pass in August and a review that never ran
because nothing was there to raise the event.
Issue: #114
User-Visible: no
The two copies of this file must match byte for byte; a comment line had drifted
by one character. main is the copy the issues event actually reads, so it is the
reference. Trivial in itself, and worth closing anyway: the file's own header
warns that a divergence between these two branches is one of the ways this
pipeline fails quietly.
Issue: #114
User-Visible: no
The guard refused to review any issue the owner had not filed himself. The rule
was meant to keep malformed outside reports out of the pipeline, but it checked at
every step instead of at the entrance, and it duplicated a guarantee the platform
already gives: only someone with write access can apply a label. Applying the
first status label is the owner's explicit decision, and it is the only place the
question belongs.
So the author check is gone. While an issue carries no status label it sits
outside the process and the invariants do not apply; once labelled, the task is in
flight and who filed it stops mattering.
The old rule also cost real work. On #123 an outside bug report had been analysed
and specified before the guard turned it away in nine seconds, and the remedy on
offer was to refile the same thing as the owner's own issue.
Issue: #114
User-Visible: no
A review label promises work. When the guard declined it wrote the reason to the
run log and nothing else, so the issue sat in a status nobody was acting on and
nobody could tell. #123 showed it: an outside reporter's issue was walked up to
S4-spec-review, the guard refused in nine seconds because only the owner's issues
enter the process, and the issue itself said not a word.
Refusals that a human can act on now become a comment: wrong author, blocked,
review-4. Only when a stage was actually recognised, so an unrelated label change
stays silent.
This is the same defect as the merge conflict that left the label untouched, seen
from the other side. The pattern is worth naming: doing nothing quietly is the
most expensive thing a pipeline can do.
Issue: #114
User-Visible: no
A green code review whose merge conflicted used to leave the label where it was.
That is a dead end: the author waits for the label to change, so it polled thirty
times and reported the limit as exhausted — on a task the reviewer had already
passed. The verdict existed and nobody could act on it.
The merge step no longer fails the job. It reports whether it merged, and a green
review that did not merge sends the task back to S6-in-progress, because the work
did return to the author — a rebase rather than a code fix, and the comment says
so and says the verdict still stands.
The invariant is now stronger and worth stating plainly: after a review run the
label always changes. A pipeline whose state can stall silently is worse than one
that reports the wrong state loudly.
Issue: #114
User-Visible: no
The step that comments on the issue when a review run dies carried a literal
backslash instead of a line continuation, so gh received four arguments and
--repo ran as a command of its own. The handler for failures would itself have
failed, silently, and only when something had already gone wrong.
bash -n does not catch this: the syntax is valid, the meaning is not. Checking
run blocks now also means looking for a doubled backslash at end of line.
Issue: #114
User-Visible: no
The guard counted every verdict comment on the issue, so a spec-review verdict
consumed a cycle from the code-review budget. On #89 the first code review came
out as r2/4. With two spec cycles the second code review would have hit review-4
after a single fix — the limit would have fired on a task nobody had reviewed
twice.
The stage is now resolved first and only its own verdicts are counted, recognised
by the review document named in the comment. If the document is missing the
verdict is not counted: undercounting grants an extra cycle, overcounting would
stop the work early, and of the two mistakes the recoverable one wins.
Issue: #114
User-Visible: no
PROCESS.md 10.2 item 10 asks for this to happen because a beta shipped, not
because someone remembered. The manual cleanup was skipped twice and both times
it broke the invariant that a closed issue carries no status label — the one
thing `verify` leans on. A manual step that falls due right after a successful
release is the worst kind: the work already looks finished, which is precisely
why it gets forgotten.
The job comments the tag, removes the label, then closes. That order is
deliberate: dying between the two steps leaves an open issue without a status,
which is visible and fixable in the flow, where the reverse order would recreate
the breakage this exists to prevent. It ends by asserting that no closed issue
still carries S8-merged — aimed at the defect that actually recurs rather than at
the invariant in general.
The stock token is used on purpose. Events caused by GITHUB_TOKEN do not start
workflows, so stripping the label cannot wake the review pipeline; a PAT here
would turn bookkeeping into a cascade.
Issue: #120
User-Visible: no
PROCESS.md 10.2 describes scripts/process-gate.mjs; the script never existed.
Commits go straight to dev without PRs and GitHub blocks nothing on its side, so
until now the only thing standing between the process and rule #1 was the good
faith of whoever was committing. Hooks catch a violation on the author's machine
but --no-verify walks past them; this job is the catch-up pass that cannot be
skipped locally.
Checks 1-7 offline, 8 through gh, plus the escalation of check 3: a class A
commit with neither a spec file nor the `small` label is a failure, not a
warning. Check 8 is fail closed — an unreachable or closed issue is a refusal,
never a silent pass.
Two things surfaced while wiring it up and are recorded in the script header.
S8-merged had to join the allowed statuses: the pipeline merges into dev before
it moves the label, so Validate reads the issue already advanced and a strict set
would redden every accepted task. And the status question now applies only to
class A/B commits — asking it of a review document would fail every time, since
that document lands while the issue sits in S4-spec-review or S7-code-review.
Issue: #105
User-Visible: no
Keeps dev identical to main so the broken revision does not come back at the
next promotion. The workflow only fires from the default branch, but a stale
copy here would overwrite the working one.
Issue: #114
User-Visible: no
Pushing the hook through the GitHub contents API dropped its mode to 100644.
assertHookMode caught it on the next commit, which is the gate working as
intended — a non-executable commit-msg would simply never run.
Also brings dev in line with the turn-limit fix already on main.
Issue: #116
User-Visible: no
The first live run failed with "Could not fetch an OIDC token": the action
needs id-token: write to authenticate the GitHub App.
The reviewer also checked out dev, where the material under review does not
exist yet — specs and code are committed to issue/<NN>-slug. The job now
switches to that branch when it is pushed, and warns loudly when it is not.
Issue: #114
User-Visible: no
Adds .github/workflows/process.yml. A status label change is the trigger:
S4-spec-review runs the spec review, S7-code-review runs the code review,
and the verdict decides the next label. Only a green verdict advances;
yellow and red return the task to its author. Cycle limits (4, or 2 on the
light track) are counted from the verdicts already posted on the issue.
Labels are moved with HP_PROCESS_TOKEN, not GITHUB_TOKEN, so the change
emits an event and the chain continues.
Issue: #114
User-Visible: no