Five times in this project a green test meant nothing was checked. The
continuity smoke stayed green after the entire mechanism it guards was cut out.
The golden scene created to protect doorway light was empty — 1,177 warm pixels
against 107,119, all of them icons. The shadow smoke passed while no shadow was
drawn. Each time the test had been written alongside the code, went green at
once, and nobody ever asked whether it could go red.
The gate makes that question routine. Each mutant is a few lines of patch that
reproduce a known breakage, plus the name of the test that must fail on it. A
worktree is patched, the bundle rebuilt, the guard run — and a guard that stays
green fails the gate. Six mutants cover the holes documented in #85; the anchors
are exact strings from today's source, so the registry cannot silently drift —
a unit test that runs with the ordinary suite refuses a stale anchor.
The full run rebuilds the bundle per mutant, so it lives in its own workflow,
before a stable release and on a weekly schedule, not in Validate. The rules for
new tests are written at the top of docs/TESTING.md, and the sixth of them is
the cheapest: an assertion that reads back the property the code just set is
not written at all.
Issue: #85
User-Visible: no
Five times in this project a green test meant nothing was checked. The
continuity smoke stayed green after the entire mechanism it guards was cut out.
The golden scene created to protect doorway light was empty — 1,177 warm pixels
against 107,119, all of them icons. The shadow smoke passed while no shadow was
drawn. Each time the test had been written alongside the code, went green at
once, and nobody ever asked whether it could go red.
The gate makes that question routine. Each mutant is a few lines of patch that
reproduce a known breakage, plus the name of the test that must fail on it. A
worktree is patched, the bundle rebuilt, the guard run — and a guard that stays
green fails the gate. Six mutants cover the holes documented in #85; the anchors
are exact strings from today's source, so the registry cannot silently drift —
a unit test that runs with the ordinary suite refuses a stale anchor.
The full run rebuilds the bundle per mutant, so it lives in its own workflow,
before a stable release and on a weekly schedule, not in Validate. The rules for
new tests are written at the top of docs/TESTING.md, and the sixth of them is
the cheapest: an assertion that reads back the property the code just set is
not written at all.
Issue: #85
User-Visible: no
The owner's report: the process works but every stage takes a long time even on
simple bugs. Two causes, and neither was the one that first comes to mind.
The reviewer ran everything regardless. On #89 it installed Chromium, ran all 127
smoke files and a full golden capture — right for a task rated 10/10 for
complexity, absurd for a bug about a room divider. Full suites are the pre-beta
gate; the review now runs typecheck, unit and build always, and smokes, golden,
pytest or performance only where the diff and the AC call for them. The price of
narrowing it is honesty: the reviewer must list which gates it ran, which it did
not, and why, so a skipped gate is a visible decision rather than a silent one.
The reviewer also built its own environment out of model turns, with no npm cache
and no browser cache, paid for from the same forty-five minutes. The workflow now
installs dependencies and Chromium as ordinary cached steps, after switching to
the task branch so the lockfile is the branch's own.
Second, ceremony did not scale down. The light track makes a spec cheap; the new
trivial track does without one — S2-analysis straight to S5-ready, no spec review,
AC in the issue body. It is deliberately hard to qualify for: a bug on one surface,
no new UX contract, no migration, no i18n, no perf or touch effect, three checkable
AC at most, and expected behaviour already on record. Nothing left to decide is the
criterion that holds the whole thing up, and it cannot be met by feeling sure.
Code review is never skipped on either track. It is what stands in for testing
here, so it is the one stage speed may not buy.
Issue: #127
Issue: #128
User-Visible: no
The owner's report: the process works but every stage takes a long time even on
simple bugs. Two causes, and neither was the one that first comes to mind.
The reviewer ran everything regardless. On #89 it installed Chromium, ran all 127
smoke files and a full golden capture — right for a task rated 10/10 for
complexity, absurd for a bug about a room divider. Full suites are the pre-beta
gate; the review now runs typecheck, unit and build always, and smokes, golden,
pytest or performance only where the diff and the AC call for them. The price of
narrowing it is honesty: the reviewer must list which gates it ran, which it did
not, and why, so a skipped gate is a visible decision rather than a silent one.
The reviewer also built its own environment out of model turns, with no npm cache
and no browser cache, paid for from the same forty-five minutes. The workflow now
installs dependencies and Chromium as ordinary cached steps, after switching to
the task branch so the lockfile is the branch's own.
Second, ceremony did not scale down. The light track makes a spec cheap; the new
trivial track does without one — S2-analysis straight to S5-ready, no spec review,
AC in the issue body. It is deliberately hard to qualify for: a bug on one surface,
no new UX contract, no migration, no i18n, no perf or touch effect, three checkable
AC at most, and expected behaviour already on record. Nothing left to decide is the
criterion that holds the whole thing up, and it cannot be met by feeling sure.
Code review is never skipped on either track. It is what stands in for testing
here, so it is the one stage speed may not buy.
Issue: #127
Issue: #128
User-Visible: no
The owner's report: the process works but every stage takes a long time even on
simple bugs. Two causes, and neither was the one that first comes to mind.
The reviewer ran everything regardless. On #89 it installed Chromium, ran all 127
smoke files and a full golden capture — right for a task rated 10/10 for
complexity, absurd for a bug about a room divider. Full suites are the pre-beta
gate; the review now runs typecheck, unit and build always, and smokes, golden,
pytest or performance only where the diff and the AC call for them. The price of
narrowing it is honesty: the reviewer must list which gates it ran, which it did
not, and why, so a skipped gate is a visible decision rather than a silent one.
The reviewer also built its own environment out of model turns, with no npm cache
and no browser cache, paid for from the same forty-five minutes. The workflow now
installs dependencies and Chromium as ordinary cached steps, after switching to
the task branch so the lockfile is the branch's own.
Second, ceremony did not scale down. The light track makes a spec cheap; the new
trivial track does without one — S2-analysis straight to S5-ready, no spec review,
AC in the issue body. It is deliberately hard to qualify for: a bug on one surface,
no new UX contract, no migration, no i18n, no perf or touch effect, three checkable
AC at most, and expected behaviour already on record. Nothing left to decide is the
criterion that holds the whole thing up, and it cannot be met by feeling sure.
Code review is never skipped on either track. It is what stands in for testing
here, so it is the one stage speed may not buy.
Issue: #127
Issue: #128
User-Visible: no
The owner's report: the process works but every stage takes a long time even on
simple bugs. Two causes, and neither was the one that first comes to mind.
The reviewer ran everything regardless. On #89 it installed Chromium, ran all 127
smoke files and a full golden capture — right for a task rated 10/10 for
complexity, absurd for a bug about a room divider. Full suites are the pre-beta
gate; the review now runs typecheck, unit and build always, and smokes, golden,
pytest or performance only where the diff and the AC call for them. The price of
narrowing it is honesty: the reviewer must list which gates it ran, which it did
not, and why, so a skipped gate is a visible decision rather than a silent one.
The reviewer also built its own environment out of model turns, with no npm cache
and no browser cache, paid for from the same forty-five minutes. The workflow now
installs dependencies and Chromium as ordinary cached steps, after switching to
the task branch so the lockfile is the branch's own.
Second, ceremony did not scale down. The light track makes a spec cheap; the new
trivial track does without one — S2-analysis straight to S5-ready, no spec review,
AC in the issue body. It is deliberately hard to qualify for: a bug on one surface,
no new UX contract, no migration, no i18n, no perf or touch effect, three checkable
AC at most, and expected behaviour already on record. Nothing left to decide is the
criterion that holds the whole thing up, and it cannot be met by feeling sure.
Code review is never skipped on either track. It is what stands in for testing
here, so it is the one stage speed may not buy.
Issue: #127
Issue: #128
User-Visible: no
Issues labelled before the pipeline existed keep their spec straight in dev and
have no issue/NN branch. The publish step quietly exited zero for them, so the
verdict would arrive as a comment and the analysis behind it would be thrown
away — the fifth instance today of a step reporting success by doing nothing.
The document now goes wherever the spec itself lives: the task branch when there
is one, dev otherwise. Publishing also survives dev moving on while the review
ran, which takes up to forty-five minutes, by rebasing once before it gives up.
Four issues are waiting on this — #12, #30, #44 and #52 — each with a spec in dev,
a status label applied during the bulk pass in August and a review that never ran
because nothing was there to raise the event.
Issue: #114
User-Visible: no
Issues labelled before the pipeline existed keep their spec straight in dev and
have no issue/NN branch. The publish step quietly exited zero for them, so the
verdict would arrive as a comment and the analysis behind it would be thrown
away — the fifth instance today of a step reporting success by doing nothing.
The document now goes wherever the spec itself lives: the task branch when there
is one, dev otherwise. Publishing also survives dev moving on while the review
ran, which takes up to forty-five minutes, by rebasing once before it gives up.
Four issues are waiting on this — #12, #30, #44 and #52 — each with a spec in dev,
a status label applied during the bulk pass in August and a review that never ran
because nothing was there to raise the event.
Issue: #114
User-Visible: no
The two copies of this file must match byte for byte; a comment line had drifted
by one character. main is the copy the issues event actually reads, so it is the
reference. Trivial in itself, and worth closing anyway: the file's own header
warns that a divergence between these two branches is one of the ways this
pipeline fails quietly.
Issue: #114
User-Visible: no
The guard refused to review any issue the owner had not filed himself. The rule
was meant to keep malformed outside reports out of the pipeline, but it checked at
every step instead of at the entrance, and it duplicated a guarantee the platform
already gives: only someone with write access can apply a label. Applying the
first status label is the owner's explicit decision, and it is the only place the
question belongs.
So the author check is gone. While an issue carries no status label it sits
outside the process and the invariants do not apply; once labelled, the task is in
flight and who filed it stops mattering.
The old rule also cost real work. On #123 an outside bug report had been analysed
and specified before the guard turned it away in nine seconds, and the remedy on
offer was to refile the same thing as the owner's own issue.
Issue: #114
User-Visible: no
The guard refused to review any issue the owner had not filed himself. The rule
was meant to keep malformed outside reports out of the pipeline, but it checked at
every step instead of at the entrance, and it duplicated a guarantee the platform
already gives: only someone with write access can apply a label. Applying the
first status label is the owner's explicit decision, and it is the only place the
question belongs.
So the author check is gone. While an issue carries no status label it sits
outside the process and the invariants do not apply; once labelled, the task is in
flight and who filed it stops mattering.
The old rule also cost real work. On #123 an outside bug report had been analysed
and specified before the guard turned it away in nine seconds, and the remedy on
offer was to refile the same thing as the owner's own issue.
Issue: #114
User-Visible: no
A review label promises work. When the guard declined it wrote the reason to the
run log and nothing else, so the issue sat in a status nobody was acting on and
nobody could tell. #123 showed it: an outside reporter's issue was walked up to
S4-spec-review, the guard refused in nine seconds because only the owner's issues
enter the process, and the issue itself said not a word.
Refusals that a human can act on now become a comment: wrong author, blocked,
review-4. Only when a stage was actually recognised, so an unrelated label change
stays silent.
This is the same defect as the merge conflict that left the label untouched, seen
from the other side. The pattern is worth naming: doing nothing quietly is the
most expensive thing a pipeline can do.
Issue: #114
User-Visible: no
A review label promises work. When the guard declined it wrote the reason to the
run log and nothing else, so the issue sat in a status nobody was acting on and
nobody could tell. #123 showed it: an outside reporter's issue was walked up to
S4-spec-review, the guard refused in nine seconds because only the owner's issues
enter the process, and the issue itself said not a word.
Refusals that a human can act on now become a comment: wrong author, blocked,
review-4. Only when a stage was actually recognised, so an unrelated label change
stays silent.
This is the same defect as the merge conflict that left the label untouched, seen
from the other side. The pattern is worth naming: doing nothing quietly is the
most expensive thing a pipeline can do.
Issue: #114
User-Visible: no
The implementation loop runs typecheck, unit and build. Golden, browser smokes,
performance and the full HA harness run before a beta — after the code review has
passed and the issue already sits in S8-merged. Some defects cannot surface any
earlier, and until now the process had nothing to say about them, so the honest
reading was a second full review cycle at the most expensive possible moment.
The owner's decision: fix it, re-run what failed, and a green run carries the
release on. The gate named the defect precisely and the same gate proves the fix,
so the check is objective and depends on nobody's judgement.
The boundary is written down with it, because "the gate found something" could
otherwise absorb an arbitrary amount of new work. A fix that changes a behaviour
contract, reaches an untouched subsystem or rivals the task in size goes through
the normal flow. Editing a test so it stops failing is concealment rather than
repair — the exception is a defect proven to be in the fixture, as on #89.
The rule also records what it costs: the author judges his own work here, which
the process refuses everywhere else. That is the price of speed at the one point
where a review cycle is dearest, and the compensation is that the re-run command
and its result are written into the issue where the release manager reads them.
Issue: #114
User-Visible: no
A review run always moves the label. The rule is written down because its absence
cost a real stall: a green code review whose merge conflicted left the label alone,
the waiting author polled thirty times and reported the limit as exhausted, and a
verdict that existed reached nobody.
Both documents now say what S6-in-progress means when the verdict was green and
only the merge failed — rebase, not rework, and the verdict still stands. They also
say that a label which did not change means the run failed rather than the work, so
the answer is logs and the owner, not more polling. Cycles are counted per stage.
Issue: #114
User-Visible: no
A green code review whose merge conflicted used to leave the label where it was.
That is a dead end: the author waits for the label to change, so it polled thirty
times and reported the limit as exhausted — on a task the reviewer had already
passed. The verdict existed and nobody could act on it.
The merge step no longer fails the job. It reports whether it merged, and a green
review that did not merge sends the task back to S6-in-progress, because the work
did return to the author — a rebase rather than a code fix, and the comment says
so and says the verdict still stands.
The invariant is now stronger and worth stating plainly: after a review run the
label always changes. A pipeline whose state can stall silently is worse than one
that reports the wrong state loudly.
Issue: #114
User-Visible: no
A green code review whose merge conflicted used to leave the label where it was.
That is a dead end: the author waits for the label to change, so it polled thirty
times and reported the limit as exhausted — on a task the reviewer had already
passed. The verdict existed and nobody could act on it.
The merge step no longer fails the job. It reports whether it merged, and a green
review that did not merge sends the task back to S6-in-progress, because the work
did return to the author — a rebase rather than a code fix, and the comment says
so and says the verdict still stands.
The invariant is now stronger and worth stating plainly: after a review run the
label always changes. A pipeline whose state can stall silently is worse than one
that reports the wrong state loudly.
Issue: #114
User-Visible: no
The step that comments on the issue when a review run dies carried a literal
backslash instead of a line continuation, so gh received four arguments and
--repo ran as a command of its own. The handler for failures would itself have
failed, silently, and only when something had already gone wrong.
bash -n does not catch this: the syntax is valid, the meaning is not. Checking
run blocks now also means looking for a doubled backslash at end of line.
Issue: #114
User-Visible: no
The step that comments on the issue when a review run dies carried a literal
backslash instead of a line continuation, so gh received four arguments and
--repo ran as a command of its own. The handler for failures would itself have
failed, silently, and only when something had already gone wrong.
bash -n does not catch this: the syntax is valid, the meaning is not. Checking
run blocks now also means looking for a doubled backslash at end of line.
Issue: #114
User-Visible: no
The guard counted every verdict comment on the issue, so a spec-review verdict
consumed a cycle from the code-review budget. On #89 the first code review came
out as r2/4. With two spec cycles the second code review would have hit review-4
after a single fix — the limit would have fired on a task nobody had reviewed
twice.
The stage is now resolved first and only its own verdicts are counted, recognised
by the review document named in the comment. If the document is missing the
verdict is not counted: undercounting grants an extra cycle, overcounting would
stop the work early, and of the two mistakes the recoverable one wins.
Issue: #114
User-Visible: no
The guard counted every verdict comment on the issue, so a spec-review verdict
consumed a cycle from the code-review budget. On #89 the first code review came
out as r2/4. With two spec cycles the second code review would have hit review-4
after a single fix — the limit would have fired on a task nobody had reviewed
twice.
The stage is now resolved first and only its own verdicts are counted, recognised
by the review document named in the comment. If the document is missing the
verdict is not counted: undercounting grants an extra cycle, overcounting would
stop the work early, and of the two mistakes the recoverable one wins.
Issue: #114
User-Visible: no
docs/specs/README.md kept a "Статус ТЗ" column with its own vocabulary — draft,
in implementation, done — next to the labels that already hold the status. Two
dictionaries for one fact drift apart, and these had: the column still called
issues "in implementation" that were closed weeks ago. The table now says only
which issue a spec belongs to.
AGENTS.md was telling agents that PROCESS.md §1 does not cover package.json and
the rest of the configuration, and to report it as missing. It covers them now.
The same paragraph gained the rule that D beats A where paths overlap, which is
what keeps the built bundle under custom_components/houseplan/frontend/ from
reading as product source.
Issue: #119
User-Visible: no
Section 10.1 promised pre-push as the blocking gate that replaces pull requests.
The hook did not exist, so the document promised a check that was not there —
worse than saying nothing, because a promise like that gets relied on. Until now
process-gate ran only as the catch-up job in CI, which reports after the code is
already in dev.
The hook skips branch deletions and tags, and for a branch the remote has not
seen it measures from the merge-base with dev rather than from the root, or every
violation committed before the gate existed would make it impossible to pass. A
missing script does not block a push: old checkouts and worktrees have to stay
usable.
gh is optional on purpose. Reading issue status needs the network, and a hook
that cannot work on a train is a hook people switch off; offline it runs what it
can and CI does the strict pass.
The executable bit is the quiet part. Git skips a hook without +x and says
nothing — the gate reports success by being absent. Measured on a real push:
mode 644 produces zero lines from the gate and the push goes through, 755 stops
it. The API cannot set the bit, so install-hooks restores it on every install.
Issue: #121
User-Visible: no
PROCESS.md 10.2 item 10 asks for this to happen because a beta shipped, not
because someone remembered. The manual cleanup was skipped twice and both times
it broke the invariant that a closed issue carries no status label — the one
thing `verify` leans on. A manual step that falls due right after a successful
release is the worst kind: the work already looks finished, which is precisely
why it gets forgotten.
The job comments the tag, removes the label, then closes. That order is
deliberate: dying between the two steps leaves an open issue without a status,
which is visible and fixable in the flow, where the reverse order would recreate
the breakage this exists to prevent. It ends by asserting that no closed issue
still carries S8-merged — aimed at the defect that actually recurs rather than at
the invariant in general.
The stock token is used on purpose. Events caused by GITHUB_TOKEN do not start
workflows, so stripping the label cannot wake the review pipeline; a PAT here
would turn bookkeeping into a cascade.
Issue: #120
User-Visible: no
Check 2 compared the Issue trailers against whatever branch the working tree
happened to be on, over whatever range it was given. Those two are not the same
set. After a rebase the CI range widens — `before` points at a discarded commit,
the merge-base slides back, and commits that belong to dev arrive carrying other
issue numbers. Every one of them then looks like a violation.
Running the gate over real history from issue/89 with a dev range produced 26
false refusals out of 26 commits, which would have reddened Validate on the next
force-push of any task branch.
The rule now reads origin/dev..HEAD for its own verdict and leaves the event
range to the other checks. A commit that genuinely carries the wrong trailer for
its branch is still caught; the integration test covers both directions.
Issue: #105
User-Visible: no
The canon moved into the repository in #112 and then stood still while the
process kept moving. A document that lags is worse than no document: an agent
reading it as truth acts on rules that no longer exist. It promised a pre-push
hook that was never written, named labels in Russian that the repository has
never used, listed a status set the gate no longer applies, and said nothing at
all about the event-driven pipeline — the largest mechanism the process has.
Label names are now English throughout and S8-merged is documented. Section 1
covers the configuration files the gate kept reporting as unclassified, and
records that D beats A where paths overlap, since the built bundle lives inside
custom_components/houseplan/frontend/. Section 10.1 admits that pre-push does
not exist. Section 10.2 matches ALLOWED_STATUS in scripts/process-gate.mjs,
including the two caveats that only surfaced once the pipeline ran. Section 10.4
is new and documents the four silent-failure traps that cost a working day each.
The source-of-truth order now says that actual automation outranks its own
description — this document included.
Issue: #119
User-Visible: no
The label asserts the code is in dev. The workflow used to set it on a green
code review while the commits were still only on the task branch, so between
the verdict and the author's merge the state machine stated something untrue —
which is exactly what happened on #104.
The merge now runs inside the pipeline, before the label. A conflict leaves
the issue in S7-code-review and comments instead.
Issue: #114
User-Visible: no
PROCESS.md wants a review document in docs/reviews/; the CI reviewer could
only leave a comment, and flagged the gap itself. It may now write there.
What lands in the commit is decided by the workflow, not by the model: every
path outside docs/reviews/ is reverted before staging, and the commit carries
the usual trailers so the provenance gate accepts it.
Issue: #114
User-Visible: no
The multi-line --body started at column zero, which ends the YAML block
scalar. The parser silently truncated the run script and left an unclosed
double quote, so the whole workflow became unusable and blocked the code
review on #104.
The body now goes through a heredoc. Validating YAML alone did not catch
this; every run block is checked with bash -n from now on.
Issue: #114
User-Visible: no
The r2 spec review on #104 produced a complete green verdict and then failed
on --max-turns 40 at turn 43, so the label step never ran and the transition
had to be reconciled by hand. Forty was a guess; a review that reads SCOPE,
AGENTS, PROCESS, the issue thread and the spec exceeds it routinely, and a
code review that also runs gates needs far more.
The real guard against a runaway run is the job timeout, not the turn count.
Issue: #114
User-Visible: no
Agents were escalating technical calls. The owner answers what a person sees
or does and how much user-visible change belongs in an issue; storage, module
layout, naming, test strategy and migration mechanics are settled by the
agents, recorded as assumptions and challenged in review.
Issue: #114
User-Visible: no
Without it S8-merged would be a lie: the label asserts the code is in dev,
while the branch-only push leaves it on the task branch. What lands is what
the reviewer just accepted, and dev is allowed to carry unreviewed code
anyway, so the merge adds no risk the branch did not already carry.
Issue: #114
User-Visible: no
The hook parsed the message path with basename, an external command. Where
it is missing from PATH the substitution yields an empty string, set -e does
not trip on it, and the MERGE_MSG guard silently stops working — a generated
merge commit would then be rejected for missing trailers it cannot have.
POSIX parameter expansion needs no external command and behaves the same in
sh, dash, bash and Git Bash.
Issue: #116
User-Visible: no
The HACS submission check does not read hacs.json to find the integration: it
globs `*manifest.json` over the whole clone of the default branch and exits 1
unless there is exactly one (hacs/default, scripts/helpers/integration_path.py).
Three files matched — the two stand-only integrations added on 2026-07-31 and
the golden baseline index added on 2026-08-11 — so the Hassfest job of PR #9004
went red five weeks into the review queue, with a log that named no file.
The stand manifests ship as manifest.template.json and demo/stand/install.sh
renames them at install time; the golden index becomes baselines-index.json
(the exported constant keeps its name, so no consumer changes).
test/repo-hygiene.test.mjs fails if a second manifest ever appears, and the
existing golden-policy assertion — which compared against 'manifest.json' and
happily passed on 'baseline-manifest.json' — now checks the suffix.