Three lines of section 10.4 and the trailing newline went missing in transit.
The lost paragraph is the one that says a label which did not change means the
run failed rather than the work — the sentence that tells a waiting author to
read the logs instead of polling for another forty-five minutes. Losing exactly
that one while copying a document about silent failures is a joke the situation
made on its own.
Caught by the byte comparison that follows every publish, which is the whole
reason it follows every publish.
Issue: #139
User-Visible: no
The owner stopped using GitHub Projects. Most of this is wording, but one part
was not: release-prerelease.mjs talked to the Project in code. finishIssues
looked up the project id, listed its items and its Status=Done option, and threw
when an issue was missing from the board — so the first release that closed an
issue would have died on a step with nothing to do with publishing. Found by
reading rather than by releasing, which was luck.
Closing issues stays, and now strips the status label first. That order is not
cosmetic: the invariant that a closed issue carries no status label has broken
twice already, both times because a manual step did it the other way round. The
close-merged job already does it in this order.
The documents now say labels and only labels. The explicit "no longer used"
lines are kept on purpose, in PROCESS.md and next to the code that used to sync:
a decision that vanishes quietly gets reintroduced a month later by someone who
never knew it was made.
Issue: #139
User-Visible: no
The owner stopped using GitHub Projects. Most of this is wording, but one part
was not: release-prerelease.mjs talked to the Project in code. finishIssues
looked up the project id, listed its items and its Status=Done option, and threw
when an issue was missing from the board — so the first release that closed an
issue would have died on a step with nothing to do with publishing. Found by
reading rather than by releasing, which was luck.
Closing issues stays, and now strips the status label first. That order is not
cosmetic: the invariant that a closed issue carries no status label has broken
twice already, both times because a manual step did it the other way round. The
close-merged job already does it in this order.
The documents now say labels and only labels. The explicit "no longer used"
lines are kept on purpose, in PROCESS.md and next to the code that used to sync:
a decision that vanishes quietly gets reintroduced a month later by someone who
never knew it was made.
Issue: #139
User-Visible: no
The owner stopped using GitHub Projects. Most of this is wording, but one part
was not: release-prerelease.mjs talked to the Project in code. finishIssues
looked up the project id, listed its items and its Status=Done option, and threw
when an issue was missing from the board — so the first release that closed an
issue would have died on a step with nothing to do with publishing. Found by
reading rather than by releasing, which was luck.
Closing issues stays, and now strips the status label first. That order is not
cosmetic: the invariant that a closed issue carries no status label has broken
twice already, both times because a manual step did it the other way round. The
close-merged job already does it in this order.
The documents now say labels and only labels. The explicit "no longer used"
lines are kept on purpose, in PROCESS.md and next to the code that used to sync:
a decision that vanishes quietly gets reintroduced a month later by someone who
never knew it was made.
Issue: #139
User-Visible: no
The owner stopped using GitHub Projects. Most of this is wording, but one part
was not: release-prerelease.mjs talked to the Project in code. finishIssues
looked up the project id, listed its items and its Status=Done option, and threw
when an issue was missing from the board — so the first release that closed an
issue would have died on a step with nothing to do with publishing. Found by
reading rather than by releasing, which was luck.
Closing issues stays, and now strips the status label first. That order is not
cosmetic: the invariant that a closed issue carries no status label has broken
twice already, both times because a manual step did it the other way round. The
close-merged job already does it in this order.
The documents now say labels and only labels. The explicit "no longer used"
lines are kept on purpose, in PROCESS.md and next to the code that used to sync:
a decision that vanishes quietly gets reintroduced a month later by someone who
never knew it was made.
Issue: #139
User-Visible: no
Every push to every branch ran 128 browser smokes, 50 golden scenes, a Home
Assistant install and a performance pass — including a push that added one spec
file. The pipeline made such pushes routine: every spec revision and every
review document is a push to a task branch and used to cost the full suite.
A changes job classifies the push range; frontend, smoke, golden, performance
and backend now run only when their paths moved, and hacs and hassfest only for
manifests, translations or Python. provenance and process-gate always run — they
judge commits, not code.
The exception carries the design. On dev everything runs, always, unfiltered:
the beta gate accepts "green Validate at the exact SHA", and if the volume of a
run depends on the diff, green stops meaning one thing — a release candidate
touches manifests and changelogs, would skip the browser suites under filtering,
and a run with skipped jobs still concludes success. That would be the sixth
silent success of the week. Filters save time on task branches, where Validate is
an early signal and the real acceptance is the code review running gates itself.
A new branch with a zero before-sha is classified from the merge-base with dev,
not from the root of history.
Issue: #136
User-Visible: no
The file declared itself pure but imported the module through the package, and
the package __init__ unconditionally imports homeassistant. Without homeassistant
installed pytest did not skip the file — it stopped collecting the whole
tests_backend directory, taking the previously working pure suite down with it.
test_validation.py had already established the by-path pattern; virtual_lights.py
imports nothing beyond the standard library, so it loads cleanly.
The async tests also dropped their pytest-asyncio dependency in favour of
asyncio.run: the offline environment does not carry the plugin, and without it
the two tests failed as unsupported async defs. The offline gate has to be green,
or nobody runs it.
Verified in both environments: pytest+voluptuous only — 129 passed where
collection previously stopped dead; with pytest-asyncio as in CI — 129 passed.
Issue: #135
User-Visible: no
Five times in this project a green test meant nothing was checked. The
continuity smoke stayed green after the entire mechanism it guards was cut out.
The golden scene created to protect doorway light was empty — 1,177 warm pixels
against 107,119, all of them icons. The shadow smoke passed while no shadow was
drawn. Each time the test had been written alongside the code, went green at
once, and nobody ever asked whether it could go red.
The gate makes that question routine. Each mutant is a few lines of patch that
reproduce a known breakage, plus the name of the test that must fail on it. A
worktree is patched, the bundle rebuilt, the guard run — and a guard that stays
green fails the gate. Six mutants cover the holes documented in #85; the anchors
are exact strings from today's source, so the registry cannot silently drift —
a unit test that runs with the ordinary suite refuses a stale anchor.
The full run rebuilds the bundle per mutant, so it lives in its own workflow,
before a stable release and on a weekly schedule, not in Validate. The rules for
new tests are written at the top of docs/TESTING.md, and the sixth of them is
the cheapest: an assertion that reads back the property the code just set is
not written at all.
Issue: #85
User-Visible: no
Five times in this project a green test meant nothing was checked. The
continuity smoke stayed green after the entire mechanism it guards was cut out.
The golden scene created to protect doorway light was empty — 1,177 warm pixels
against 107,119, all of them icons. The shadow smoke passed while no shadow was
drawn. Each time the test had been written alongside the code, went green at
once, and nobody ever asked whether it could go red.
The gate makes that question routine. Each mutant is a few lines of patch that
reproduce a known breakage, plus the name of the test that must fail on it. A
worktree is patched, the bundle rebuilt, the guard run — and a guard that stays
green fails the gate. Six mutants cover the holes documented in #85; the anchors
are exact strings from today's source, so the registry cannot silently drift —
a unit test that runs with the ordinary suite refuses a stale anchor.
The full run rebuilds the bundle per mutant, so it lives in its own workflow,
before a stable release and on a weekly schedule, not in Validate. The rules for
new tests are written at the top of docs/TESTING.md, and the sixth of them is
the cheapest: an assertion that reads back the property the code just set is
not written at all.
Issue: #85
User-Visible: no
The owner's report: the process works but every stage takes a long time even on
simple bugs. Two causes, and neither was the one that first comes to mind.
The reviewer ran everything regardless. On #89 it installed Chromium, ran all 127
smoke files and a full golden capture — right for a task rated 10/10 for
complexity, absurd for a bug about a room divider. Full suites are the pre-beta
gate; the review now runs typecheck, unit and build always, and smokes, golden,
pytest or performance only where the diff and the AC call for them. The price of
narrowing it is honesty: the reviewer must list which gates it ran, which it did
not, and why, so a skipped gate is a visible decision rather than a silent one.
The reviewer also built its own environment out of model turns, with no npm cache
and no browser cache, paid for from the same forty-five minutes. The workflow now
installs dependencies and Chromium as ordinary cached steps, after switching to
the task branch so the lockfile is the branch's own.
Second, ceremony did not scale down. The light track makes a spec cheap; the new
trivial track does without one — S2-analysis straight to S5-ready, no spec review,
AC in the issue body. It is deliberately hard to qualify for: a bug on one surface,
no new UX contract, no migration, no i18n, no perf or touch effect, three checkable
AC at most, and expected behaviour already on record. Nothing left to decide is the
criterion that holds the whole thing up, and it cannot be met by feeling sure.
Code review is never skipped on either track. It is what stands in for testing
here, so it is the one stage speed may not buy.
Issue: #127
Issue: #128
User-Visible: no
The owner's report: the process works but every stage takes a long time even on
simple bugs. Two causes, and neither was the one that first comes to mind.
The reviewer ran everything regardless. On #89 it installed Chromium, ran all 127
smoke files and a full golden capture — right for a task rated 10/10 for
complexity, absurd for a bug about a room divider. Full suites are the pre-beta
gate; the review now runs typecheck, unit and build always, and smokes, golden,
pytest or performance only where the diff and the AC call for them. The price of
narrowing it is honesty: the reviewer must list which gates it ran, which it did
not, and why, so a skipped gate is a visible decision rather than a silent one.
The reviewer also built its own environment out of model turns, with no npm cache
and no browser cache, paid for from the same forty-five minutes. The workflow now
installs dependencies and Chromium as ordinary cached steps, after switching to
the task branch so the lockfile is the branch's own.
Second, ceremony did not scale down. The light track makes a spec cheap; the new
trivial track does without one — S2-analysis straight to S5-ready, no spec review,
AC in the issue body. It is deliberately hard to qualify for: a bug on one surface,
no new UX contract, no migration, no i18n, no perf or touch effect, three checkable
AC at most, and expected behaviour already on record. Nothing left to decide is the
criterion that holds the whole thing up, and it cannot be met by feeling sure.
Code review is never skipped on either track. It is what stands in for testing
here, so it is the one stage speed may not buy.
Issue: #127
Issue: #128
User-Visible: no
The owner's report: the process works but every stage takes a long time even on
simple bugs. Two causes, and neither was the one that first comes to mind.
The reviewer ran everything regardless. On #89 it installed Chromium, ran all 127
smoke files and a full golden capture — right for a task rated 10/10 for
complexity, absurd for a bug about a room divider. Full suites are the pre-beta
gate; the review now runs typecheck, unit and build always, and smokes, golden,
pytest or performance only where the diff and the AC call for them. The price of
narrowing it is honesty: the reviewer must list which gates it ran, which it did
not, and why, so a skipped gate is a visible decision rather than a silent one.
The reviewer also built its own environment out of model turns, with no npm cache
and no browser cache, paid for from the same forty-five minutes. The workflow now
installs dependencies and Chromium as ordinary cached steps, after switching to
the task branch so the lockfile is the branch's own.
Second, ceremony did not scale down. The light track makes a spec cheap; the new
trivial track does without one — S2-analysis straight to S5-ready, no spec review,
AC in the issue body. It is deliberately hard to qualify for: a bug on one surface,
no new UX contract, no migration, no i18n, no perf or touch effect, three checkable
AC at most, and expected behaviour already on record. Nothing left to decide is the
criterion that holds the whole thing up, and it cannot be met by feeling sure.
Code review is never skipped on either track. It is what stands in for testing
here, so it is the one stage speed may not buy.
Issue: #127
Issue: #128
User-Visible: no
The owner's report: the process works but every stage takes a long time even on
simple bugs. Two causes, and neither was the one that first comes to mind.
The reviewer ran everything regardless. On #89 it installed Chromium, ran all 127
smoke files and a full golden capture — right for a task rated 10/10 for
complexity, absurd for a bug about a room divider. Full suites are the pre-beta
gate; the review now runs typecheck, unit and build always, and smokes, golden,
pytest or performance only where the diff and the AC call for them. The price of
narrowing it is honesty: the reviewer must list which gates it ran, which it did
not, and why, so a skipped gate is a visible decision rather than a silent one.
The reviewer also built its own environment out of model turns, with no npm cache
and no browser cache, paid for from the same forty-five minutes. The workflow now
installs dependencies and Chromium as ordinary cached steps, after switching to
the task branch so the lockfile is the branch's own.
Second, ceremony did not scale down. The light track makes a spec cheap; the new
trivial track does without one — S2-analysis straight to S5-ready, no spec review,
AC in the issue body. It is deliberately hard to qualify for: a bug on one surface,
no new UX contract, no migration, no i18n, no perf or touch effect, three checkable
AC at most, and expected behaviour already on record. Nothing left to decide is the
criterion that holds the whole thing up, and it cannot be met by feeling sure.
Code review is never skipped on either track. It is what stands in for testing
here, so it is the one stage speed may not buy.
Issue: #127
Issue: #128
User-Visible: no
Issues labelled before the pipeline existed keep their spec straight in dev and
have no issue/NN branch. The publish step quietly exited zero for them, so the
verdict would arrive as a comment and the analysis behind it would be thrown
away — the fifth instance today of a step reporting success by doing nothing.
The document now goes wherever the spec itself lives: the task branch when there
is one, dev otherwise. Publishing also survives dev moving on while the review
ran, which takes up to forty-five minutes, by rebasing once before it gives up.
Four issues are waiting on this — #12, #30, #44 and #52 — each with a spec in dev,
a status label applied during the bulk pass in August and a review that never ran
because nothing was there to raise the event.
Issue: #114
User-Visible: no
Issues labelled before the pipeline existed keep their spec straight in dev and
have no issue/NN branch. The publish step quietly exited zero for them, so the
verdict would arrive as a comment and the analysis behind it would be thrown
away — the fifth instance today of a step reporting success by doing nothing.
The document now goes wherever the spec itself lives: the task branch when there
is one, dev otherwise. Publishing also survives dev moving on while the review
ran, which takes up to forty-five minutes, by rebasing once before it gives up.
Four issues are waiting on this — #12, #30, #44 and #52 — each with a spec in dev,
a status label applied during the bulk pass in August and a review that never ran
because nothing was there to raise the event.
Issue: #114
User-Visible: no
The two copies of this file must match byte for byte; a comment line had drifted
by one character. main is the copy the issues event actually reads, so it is the
reference. Trivial in itself, and worth closing anyway: the file's own header
warns that a divergence between these two branches is one of the ways this
pipeline fails quietly.
Issue: #114
User-Visible: no
The guard refused to review any issue the owner had not filed himself. The rule
was meant to keep malformed outside reports out of the pipeline, but it checked at
every step instead of at the entrance, and it duplicated a guarantee the platform
already gives: only someone with write access can apply a label. Applying the
first status label is the owner's explicit decision, and it is the only place the
question belongs.
So the author check is gone. While an issue carries no status label it sits
outside the process and the invariants do not apply; once labelled, the task is in
flight and who filed it stops mattering.
The old rule also cost real work. On #123 an outside bug report had been analysed
and specified before the guard turned it away in nine seconds, and the remedy on
offer was to refile the same thing as the owner's own issue.
Issue: #114
User-Visible: no
The guard refused to review any issue the owner had not filed himself. The rule
was meant to keep malformed outside reports out of the pipeline, but it checked at
every step instead of at the entrance, and it duplicated a guarantee the platform
already gives: only someone with write access can apply a label. Applying the
first status label is the owner's explicit decision, and it is the only place the
question belongs.
So the author check is gone. While an issue carries no status label it sits
outside the process and the invariants do not apply; once labelled, the task is in
flight and who filed it stops mattering.
The old rule also cost real work. On #123 an outside bug report had been analysed
and specified before the guard turned it away in nine seconds, and the remedy on
offer was to refile the same thing as the owner's own issue.
Issue: #114
User-Visible: no
A review label promises work. When the guard declined it wrote the reason to the
run log and nothing else, so the issue sat in a status nobody was acting on and
nobody could tell. #123 showed it: an outside reporter's issue was walked up to
S4-spec-review, the guard refused in nine seconds because only the owner's issues
enter the process, and the issue itself said not a word.
Refusals that a human can act on now become a comment: wrong author, blocked,
review-4. Only when a stage was actually recognised, so an unrelated label change
stays silent.
This is the same defect as the merge conflict that left the label untouched, seen
from the other side. The pattern is worth naming: doing nothing quietly is the
most expensive thing a pipeline can do.
Issue: #114
User-Visible: no
A review label promises work. When the guard declined it wrote the reason to the
run log and nothing else, so the issue sat in a status nobody was acting on and
nobody could tell. #123 showed it: an outside reporter's issue was walked up to
S4-spec-review, the guard refused in nine seconds because only the owner's issues
enter the process, and the issue itself said not a word.
Refusals that a human can act on now become a comment: wrong author, blocked,
review-4. Only when a stage was actually recognised, so an unrelated label change
stays silent.
This is the same defect as the merge conflict that left the label untouched, seen
from the other side. The pattern is worth naming: doing nothing quietly is the
most expensive thing a pipeline can do.
Issue: #114
User-Visible: no
The implementation loop runs typecheck, unit and build. Golden, browser smokes,
performance and the full HA harness run before a beta — after the code review has
passed and the issue already sits in S8-merged. Some defects cannot surface any
earlier, and until now the process had nothing to say about them, so the honest
reading was a second full review cycle at the most expensive possible moment.
The owner's decision: fix it, re-run what failed, and a green run carries the
release on. The gate named the defect precisely and the same gate proves the fix,
so the check is objective and depends on nobody's judgement.
The boundary is written down with it, because "the gate found something" could
otherwise absorb an arbitrary amount of new work. A fix that changes a behaviour
contract, reaches an untouched subsystem or rivals the task in size goes through
the normal flow. Editing a test so it stops failing is concealment rather than
repair — the exception is a defect proven to be in the fixture, as on #89.
The rule also records what it costs: the author judges his own work here, which
the process refuses everywhere else. That is the price of speed at the one point
where a review cycle is dearest, and the compensation is that the re-run command
and its result are written into the issue where the release manager reads them.
Issue: #114
User-Visible: no