Commit Graph
86 Commits
Author SHA1 Message Date
Matysh e4b4f33c6d fix: verify demo bundle freshness for smokes too
Golden runs, benchmarks and documentation captures each called
assertFreshDemoBundle; the smoke launcher never did, so all ~128 smokes could
silently test a stale demo/srv/assets bundle. On #234 that cost a round of
analysis: three assertions went red and a fourth went green, because the old
code was wrong in two places that agreed with each other, and a mixed result
reads as a logic defect rather than a stale artefact.

launch() now runs the check once for every smoke, against the repository root
rather than the serving root — demo/srv has no src/** to fingerprint.
HP_ALLOW_STALE_BUNDLE=1 skips it for debugging and warns out loud, because a
guard that says nothing when it steps aside is the silent success this project
keeps removing. A mutation entry proves the call cannot quietly disappear.

Issue: #236
User-Visible: no
2026-08-22 01:47:47 +03:00
Matysh 5fd9716906 fix: verify demo bundle freshness for smokes too
Golden runs, benchmarks and documentation captures each called
assertFreshDemoBundle; the smoke launcher never did, so all ~128 smokes could
silently test a stale demo/srv/assets bundle. On #234 that cost a round of
analysis: three assertions went red and a fourth went green, because the old
code was wrong in two places that agreed with each other, and a mixed result
reads as a logic defect rather than a stale artefact.

launch() now runs the check once for every smoke, against the repository root
rather than the serving root — demo/srv has no src/** to fingerprint.
HP_ALLOW_STALE_BUNDLE=1 skips it for debugging and warns out loud, because a
guard that says nothing when it steps aside is the silent success this project
keeps removing. A mutation entry proves the call cannot quietly disappear.

Issue: #236
User-Visible: no
2026-08-22 01:46:39 +03:00
Matysh 1fb3a5754e fix: verify demo bundle freshness for smokes too
Golden runs, benchmarks and documentation captures each called
assertFreshDemoBundle; the smoke launcher never did, so all ~128 smokes could
silently test a stale demo/srv/assets bundle. On #234 that cost a round of
analysis: three assertions went red and a fourth went green, because the old
code was wrong in two places that agreed with each other, and a mixed result
reads as a logic defect rather than a stale artefact.

launch() now runs the check once for every smoke, against the repository root
rather than the serving root — demo/srv has no src/** to fingerprint.
HP_ALLOW_STALE_BUNDLE=1 skips it for debugging and warns out loud, because a
guard that says nothing when it steps aside is the silent success this project
keeps removing. A mutation entry proves the call cannot quietly disappear.

Issue: #236
User-Visible: no
2026-08-22 01:45:56 +03:00
MatyshandSergey Matyunin b2024c6626 docs(spec): state that i18n is untouched, name the zero case
The spec review found the i18n section missing outright — the analysis
comment claimed it was untouched, the document itself said nothing, and a
DoR section cannot be inferred from a comment. It now says so explicitly,
and adds that the absence of i18n files from the diff is part of the
contract rather than an accident.

The waived Low is closed too: the old wallChainSegments treated a recorded
zero as valid while the new contract requires strictly positive values, so
zero is now named in the AC1 examples instead of being derivable from the
prose.

Issue: #234
User-Visible: no
2026-08-21 19:41:37 +03:00
MatyshandSergey Matyunin f4098fcf09 docs(spec): register the #234 spec in the index
Issue: #234
User-Visible: no
2026-08-21 19:41:37 +03:00
MatyshandSergey Matyunin 05b2c74442 docs(spec): define one thickness resolver for a wall chain
Issue: #234
User-Visible: no
2026-08-21 19:41:37 +03:00
Matysh 33470306d2 fix: keep the review document outside the tree the reviewer mutates
Three code-review rounds on #220 published a verdict and then failed the
run: the document never reached the branch, so the #171 guard refused
before the label step and neither the merge nor S8-merged happened. The
cause was structural. The document lived as an untracked file inside the
very checkout the reviewer edits while proving that a test can fail, and
restoring that tree — git checkout, git clean — deletes an untracked file.
Spec rounds survived only because they never mutate anything.

The reviewer now writes to REVIEW_DOC under RUNNER_TEMP, outside the
repository, and the publish step copies it into docs/reviews before
committing. Tree cleanup can no longer destroy the artefact, and the
reviewer no longer needs to touch docs/reviews at all.

Verified against a local git fixture on five paths: document outside the
repo with a mutated tree, nothing anywhere (loud failure), document only in
the working copy, document already committed by the reviewer, and a branch
that moved during the review.

Same file as main, byte for byte.

Issue: #220
User-Visible: no
2026-08-21 10:53:07 +03:00
Matysh 07f1b8c674 docs: record that a green verdict spends no review budget
Validate / docs (push) Failing after 27s
Validate / provenance (push) Successful in 54s
Validate / process-gate (push) Failing after 44s
Validate / changes (push) Successful in 39s
Validate / reuse (push) Successful in 45s
Validate / hacs (push) Failing after 16s
Validate / hassfest (push) Failing after 13s
Validate / frontend (push) Successful in 6m37s
Validate / backend (push) Failing after 9m6s
Validate / performance_smoke (push) Failing after 18m53s
Validate / smoke (push) Failing after 39m24s
Validate / golden (push) Failing after 11m5s
Section 4 now says what the pipeline does: a cycle is a verdict with
blocking findings followed by a return to the author, so a green verdict
consumes nothing — the case that cost #225 an arbitration after a failed
merge forced a rebase and a third attempt.

The canon also separates the two quantities the verdict line carries. The
attempt number names the review document, because two runs sharing a
number would overwrite each other's artefact; the budget counts blocking
cycles only. That is why the document threshold in the process gate sits
above the cycle limit, and why the guard reports a recount of review-4
rather than removing the label itself.

Issue: #227
User-Visible: no
2026-08-21 00:01:19 +03:00
Matysh afd5d589d5 fix: spend the review budget on blocking verdicts only
The pipeline punished what it prescribed: after a failed merge it tells the
author to rebase and restore S7-code-review, and that attempt finished the
budget. On #225 a green code review with green CI ended in review-4.

Only yellow and red verdicts spend the budget now; a green verdict returned
nothing and consumes nothing. Attempts and cycles became separate
quantities: the attempt number names the review document, the limit
compares blocking cycles. The exhaustion comment lists what it counted, and
the guard reports a recount instead of stripping review-4 on its own.

Same file as main (41325a8), byte for byte.

Issue: #227
User-Visible: no
2026-08-20 23:54:13 +03:00
Matysh 99c6cd09f3 fix: restore the exact process-gate source for #227
Issue: #227
User-Visible: no
2026-08-20 23:49:12 +03:00
Matysh 6176f02430 test: expect the raised review-document threshold
Issue: #227
User-Visible: no
2026-08-20 23:42:29 +03:00
Matysh 01d7554607 test: expect the raised review-document threshold
Issue: #227
User-Visible: no
2026-08-20 23:37:57 +03:00
Matysh 9848f4a0cb fix: spend the review budget on blocking verdicts only
The pipeline punished what it prescribed: after a failed merge it tells the
author to rebase and restore S7-code-review, and that attempt finished the
budget. On #225 (light track, limit 2) the sequence yellow, green, rebase
produced review-4 on a task whose code review was green and whose CI was
green, with no product change after the verdict — the owner had to
arbitrate work that was already accepted.

A cycle under section 4 is a verdict with blocking findings followed by a
return to the author, so only yellow and red verdicts spend the budget now.
A green verdict returned nothing and consumes nothing, which also removes
any need to mark rebase re-runs specially.

Attempts and cycles are now separate quantities. The attempt number keeps
naming the document, because two runs sharing a number would overwrite each
other's review artefact, while the limit compares blocking cycles only. The
exhaustion comment lists the verdicts it counted, and the guard no longer
strips review-4 — it reports the recount and leaves the decision with the
owner.

Rule 7 of the process gate follows: its document threshold rises above the
cycle limit, because legitimate attempts can exceed cycles and a threshold
equal to the limit would refuse the very rebase the pipeline demands.

Issue: #227
User-Visible: no
2026-08-20 23:33:08 +03:00
Matysh 2b45086794 docs: scope a repeat review round to the delta
The reviewer prompt was identical for every round, and the canon said
nothing about the scope of a repeat pass, so r2 re-derived the product
framing and re-checked acceptance criteria the fix never touched: the r2
pass on #150 cost a full pipeline run over one line in a test fixture.

From the second cycle on, the subject is the delta against the SHA the
previous verdict was given on: each earlier finding must be shown closed
by a line of code or text, only the criteria the delta can reach are
re-verified, and whatever is carried over is listed with the round and SHA
it came from. Cheap gates still run every round.

The scope shrinks, the strictness does not. A fix can break a criterion an
earlier round accepted — that is how regression #102 happened — so the
boundary is the findings plus everything the delta can reach, and a
non-local delta (a rebase onto a moved dev, a behaviour contract change, a
new subsystem) still gets the full pass.

Issue: #214
User-Visible: no
2026-08-20 11:50:16 +03:00
Matysh cbaa7b0cb9 docs: scope a repeat review round to the delta
The reviewer prompt was identical for every round, and the canon said
nothing about the scope of a repeat pass, so r2 re-derived the product
framing and re-checked acceptance criteria the fix never touched: the r2
pass on #150 cost a full pipeline run over one line in a test fixture.

From the second cycle on, the subject is the delta against the SHA the
previous verdict was given on: each earlier finding must be shown closed
by a line of code or text, only the criteria the delta can reach are
re-verified, and whatever is carried over is listed with the round and SHA
it came from. Cheap gates still run every round.

The scope shrinks, the strictness does not. A fix can break a criterion an
earlier round accepted — that is how regression #102 happened — so the
boundary is the findings plus everything the delta can reach, and a
non-local delta (a rebase onto a moved dev, a behaviour contract change, a
new subsystem) still gets the full pass.

Issue: #214
User-Visible: no
2026-08-20 11:29:45 +03:00
Matysh 72eae1059c perf: skip a heavy gate whose inputs are byte-identical to a green run
Every push to dev paid for the full browser trio and the backend suite,
including commits that touch only documentation, workflows or process
scripts — the bundle and the harness were byte-identical, so the runs
proved nothing new. On 2026-08-19 alone that was roughly six pushes at
about seven minutes each.

The reuse key per heavy job is sourceFingerprint (src, demo fixtures,
golden scenarios, build manifests) plus that job's own harness: smoke
takes demo/smoke_*.mjs, golden takes demo/golden/** including baselines,
performance_smoke takes demo/performance/**, backend takes tests_backend
and the Python sources. A cache marker is written only by a successful run
of the same key, so a hit proves a job with identical inputs already
passed. scripts/** is deliberately outside every key: infrastructure work
edits it constantly and reuse would never fire.

This is not the path filter from the `changes` job, which stays disabled
on dev on purpose: there the scope is guessed from paths and "green" means
different things, here input equivalence is proven by a hash. And a
release candidate always bumps the version, which is part of the
fingerprint, so its keys are new by construction and the full gate set
still runs before every beta and release.

A waived job is announced with a notice and a run summary line rather than
skipped in silence, and the marker save tolerates a concurrent identical
run instead of reddening the job.

Issue: #208
User-Visible: no
2026-08-19 23:33:44 +03:00
Matysh ad8e7a50cc fix: waive the issue status for a class-A-free range in the process gate
Rule 8 demanded a working S-label from every class A/B commit's issue,
while owner decision #118 sends infrastructure work outside the S1..S8
flow entirely — such an issue has no status label by construction. The
two rules contradicted each other and the machine-checked one won, so
Validate on dev went red on every infrastructure commit (#175, #191,
#202, #206) and the catch-up signal stopped meaning anything. A gate that
is always red is not a gate.

The waiver keys on the diff, not on a permission label: a range with no
class A file at all. An `infra` label could be pinned on a product task
to walk a product commit past the status check; ceasing to touch class A
without ceasing to be infrastructure work is not possible. Issue
existence, open state, `blocked` and fail-closed on an unreachable gh all
still apply, and the waiver prints a visible warning rather than passing
in silence.

Mutation-checked both ways: unwiring the waiver reddens the CLI test,
and letting class A keep the waiver reddens both new tests.

Issue: #207
User-Visible: no
2026-08-19 20:39:07 +03:00
Matysh f287bddd97 fix: cache Playwright browsers instead of reinstalling them via apt
performance_smoke burned nearly all of its 15-minute budget before the
benchmark even started, twice in a row: validate.yml had no browser cache
at all, so every browser job paid for a full `playwright install
--with-deps` — apt work the ubuntu-latest image makes redundant, with
unbounded retries against an unreachable azure mirror on top. For a
measuring job that is worse than lost minutes: the timing window competes
with package installation on the same runner.

#175 fixed this for the review pipeline but deliberately left the flag
here, reasoning that a prerelease gate values predictability over
minutes. That reasoning was wrong — the flag is what made the gate
unpredictable.

Browsers are now cached per package-lock hash in smoke, golden,
performance_smoke and the full performance run; installation happens only
on a cache miss and no longer touches apt. performance_smoke keeps
headroom for a cold cache at 20 minutes. If the image ever drops a
required library, Chromium fails to launch with a clear missing-libraries
error; that is the moment to bring the flag back.

Issue: #206
User-Visible: no
2026-08-19 20:06:38 +03:00
Matysh 6e93aa705c docs: fix in-scope Medium findings inside the current issue
Filing and servicing a separate issue costs far more than fixing a small
problem in place — the owner's call of 2026-08-19 (#202). A Medium finding
inside the task's scope no longer becomes its own issue: with no High
findings the verdict is yellow, the author fixes it and the fix passes
another review cycle. Only an out-of-scope Medium is still filed
separately, because foreign scope is never patched from a task branch.

Applied to the canon (PROCESS.md), the reviewer prompt in process.yml and
AGENTS.md; the verdict format now writes "Medium: N -> in-task | #NN".

Issue: #202
User-Visible: no
2026-08-19 13:46:25 +03:00
Matysh c1dde9a0cf docs: fix in-scope Medium findings inside the current issue
Validate / provenance (push) Successful in 42s
Validate / process-gate (push) Failing after 36s
Validate / changes (push) Successful in 44s
Validate / hacs (push) Skipped
Validate / performance_smoke (push) Skipped
Validate / hassfest (push) Skipped
Validate / frontend (push) Skipped
Validate / smoke (push) Skipped
Validate / golden (push) Skipped
Validate / backend (push) Skipped
Full Performance / performance (push) Failing after 2h16m25s
Filing and servicing a separate issue costs far more than fixing a small
problem in place — the owner's call of 2026-08-19 (#202). A Medium finding
inside the task's scope no longer becomes its own issue: with no High
findings the verdict is yellow, the author fixes it and the fix passes
another review cycle. Only an out-of-scope Medium is still filed
separately, because foreign scope is never patched from a task branch.

Applied to the canon (PROCESS.md), the reviewer prompt in process.yml and
AGENTS.md; the verdict format now writes "Medium: N -> in-task | #NN".

Issue: #202
User-Visible: no
2026-08-19 13:16:30 +03:00
Matysh a55ba3de8b test: make the junction-patch bridge test able to fail
The old fixture ran the branch into the end of p1, so the T-patch lived
beyond the host (x>100) and the probe point [92,0] sat outside it under
any code behaviour: deleting the cutPartitionBody flatMap over patches
kept all subtests green (#188, found at the #132 code review).

The branch now meets the middle of the span, putting both node patches
inside the default opening's slot. The test proves its own fixture first:
an uncut run must show the patches bridging the slot, so if the geometry
ever stops producing them the control goes red instead of silently
devaluing the real assertion. Mutation-checked: reverting the flatMap
fails exactly this test, 888/889.

Issue: #188
User-Visible: no
2026-08-19 09:57:43 +03:00
Matysh 2f0dc44f27 fix: clamp an issue-branch gate range to the branch's own commits
After the mandatory rebase of a published issue branch the pre-push hook
still passes remote_old..local_new, and once the old tip is no longer an
ancestor that range drags in the whole advanced dev history: on #117 it
meant 84 foreign commits and 20 false rule-8 rejections over already
closed issues, leaving --no-verify as the only exit.

The clamp lives in the gate rather than the hook: .githooks/pre-push
carries an executable bit that MCP publication strips (the commit-msg
precedent), so editing it needs an owner-side commit. When the target is
an issue branch and the declared base is not an ancestor of the head, the
base becomes the merge-base with origin/dev. Fast-forward pushes keep
their exact range, every own commit is still judged, and a real violation
in a post-rebase commit still blocks — covered by a scenario test that
goes red without the wiring.

Issue: #190
User-Visible: no
2026-08-19 09:48:39 +03:00
Matysh 3e335b8808 docs: list the actual Validate gate jobs in AGENTS.md
The canonical job list predated the `docs` job (added 2026-08-16) and
omitted `process-gate`. `docs` is a real blocking gate — its fingerprint
check went red right after the #113 merge and cost an extra review cycle
of confusion. The list now matches validate.yml and names `changes` as a
service path-filter rather than a gate.

Issue: #191
User-Visible: no
2026-08-19 09:16:54 +03:00
Matysh e727023bf4 fix: install Chromium without --with-deps in the review pipeline
Validate / smoke (push) Failing after 31s
Validate / golden (push) Failing after 12s
Validate / backend (push) Failing after 19m25s
Validate / performance_smoke (push) Failing after 14m0s
Validate / docs (push) Failing after 27s
Validate / changes (push) Successful in 44s
Validate / provenance (push) Successful in 48s
Validate / process-gate (push) Failing after 46s
Validate / hacs (push) Failing after 12s
Validate / hassfest (push) Failing after 24s
Validate / frontend (push) Successful in 6m14s
On a Playwright cache miss the flag pulled Chromium's system libraries
through apt, spending minutes of the 45-minute review budget on packages
the ubuntu-latest image already ships — and the runner's retries against
the unreachable azure mirror made the step look hung on a live run. If the
image ever drops a required library, Chromium fails to launch with a clear
missing-libraries error; that is the moment to bring the flag back.

validate.yml keeps the flag deliberately: it is the prerelease gate, where
predictability is worth more than minutes.

Issue: #175
User-Visible: no
2026-08-18 20:01:23 +03:00
Matysh 91f2c23539 fix: install Chromium without --with-deps in the review pipeline
Full Performance / performance (push) Failing after 1h30m16s
Validate / provenance (push) Successful in 48s
Validate / changes (push) Successful in 38s
Validate / performance_smoke (push) Skipped
Validate / backend (push) Skipped
Validate / process-gate (push) Failing after 45s
Validate / hacs (push) Skipped
Validate / hassfest (push) Skipped
Validate / frontend (push) Skipped
Validate / smoke (push) Skipped
Validate / golden (push) Skipped
On a Playwright cache miss the flag pulled Chromium's system libraries
through apt, spending minutes of the 45-minute review budget on packages
the ubuntu-latest image already ships — and the runner's retries against
the unreachable azure mirror made the step look hung on a live run. If the
image ever drops a required library, Chromium fails to launch with a clear
missing-libraries error; that is the moment to bring the flag back.

validate.yml keeps the flag deliberately: it is the prerelease gate, where
predictability is worth more than minutes.

Issue: #175
User-Visible: no
2026-08-18 19:53:15 +03:00
Matysh ae7fb621e7 fix: fail loudly when a review verdict has no document
On #150 both spec-review verdicts survived only as issue comments: the
publish step found nothing staged, printed a warning, and exited zero, so
the label moved and the missing artifact went unnoticed until the next
review caught it (#171). A verdict without a document in docs/reviews/ now
fails the run before the label step, preserving the invariant that an
unchanged label means a failed run.

An empty working copy alone is not a failure: the reviewer occasionally
commits the document itself through its app token, bypassing this step
(CODE-REVIEW-150-r1, committer GitHub), so the branch is checked first. A
postcondition verifies the exact expected filename reached the branch, and
the rebase-conflict path no longer exits zero either.

Issue: #171
User-Visible: no
2026-08-18 18:02:44 +03:00
Matysh ba32234b52 fix: fail loudly when a review verdict has no document
On #150 both spec-review verdicts survived only as issue comments: the
publish step found nothing staged, printed a warning, and exited zero, so
the label moved and the missing artifact went unnoticed until the next
review caught it (#171). A verdict without a document in docs/reviews/ now
fails the run before the label step, preserving the invariant that an
unchanged label means a failed run.

An empty working copy alone is not a failure: the reviewer occasionally
commits the document itself through its app token, bypassing this step
(CODE-REVIEW-150-r1, committer GitHub), so the branch is checked first. A
postcondition verifies the exact expected filename reached the branch, and
the rebase-conflict path no longer exits zero either.

Issue: #171
User-Visible: no
2026-08-18 17:54:51 +03:00
Matysh c27185cfa4 fix: verify the PAT before reviewing, pick the freshest task branch
Issue #150 reached a green verdict and then hit two pipeline defects at once.
The review document push came back 403 as github-actions[bot]: the PAT had
died, and checkout's persisted credential quietly took its place — a masked
actor instead of a loud failure. Credentials are no longer persisted, and the
token is now proven alive before the review starts, not after forty minutes of
reviewer work.

Branch selection took the first match alphabetically, and with a spec-era
branch sitting next to the implementation branch that meant the stale one.
The freshest branch by commit date is chosen instead, with a warning naming
every candidate when more than one exists.

Verified against the real #150 branches: the fix branch wins, the warning
fires.

Issue: #114
User-Visible: no
2026-08-18 17:21:06 +03:00
Matysh 382afd2766 fix: verify the PAT before reviewing, pick the freshest task branch
Issue #150 reached a green verdict and then hit two pipeline defects at once.
The review document push came back 403 as github-actions[bot]: the PAT had
died, and checkout's persisted credential quietly took its place — a masked
actor instead of a loud failure. Credentials are no longer persisted, and the
token is now proven alive before the review starts, not after forty minutes of
reviewer work.

Branch selection took the first match alphabetically, and with a spec-era
branch sitting next to the implementation branch that meant the stale one.
The freshest branch by commit date is chosen instead, with a warning naming
every candidate when more than one exists.

Verified against the real #150 branches: the fix branch wins, the warning
fires.

Issue: #114
User-Visible: no
2026-08-18 17:13:35 +03:00
Matysh 7642c484d2 feat: analysis proceeds on its own, questions are product-only and spec-stage
The analyst used to ask the owner to confirm every estimate and waited for an
answer on each point. Most issues are unambiguous, and most of that waiting
changed nothing — the owner's own measure is that seven issues in ten should
travel from S1-new to a finished spec without a single question.

Section 2.2 flips the default. Estimates, type, priority and track go on as
labels immediately; the analysis comment is a notification, not a request —
the owner's silence is consent, his disagreement is a label edit, and neither
stops the work. The analyst moves the issue to S3-spec himself. The only
questions that ever reach the owner are product questions, asked at the spec
stage in one batch with defaults and blocked, and only when the spec cannot be
written without the answer; anything that can wait for the spec waits, anything
that does not block it becomes a recorded assumption instead. The one full stop
left in analysis is a genuine SCOPE conflict, where the analyst proposes
rejection and the owner decides.

Issue: #114
User-Visible: no
2026-08-14 20:31:08 +03:00
Matysh ec7408f3d5 docs: the pre-1.62 "no tests, no commits" workflow line is dead
The author agent read STATUS.md, saw the owner's 2026-08-07 rule that ordinary
fixes are made locally without tests or commits, correctly ranked it below
AGENTS.md and PROCESS.md, and followed the canon instead. That is the trust
order doing its job — and the canon's second half says a divergence is not
ignored but fixed.

The line now says what replaced it: since release 1.62 every product change goes
through the process — an issue in S5-ready or later, a task branch, trailers on
every commit, the review pipeline. The release mechanics in the same cell were
still accurate and stay.

Issue: #114
User-Visible: no
2026-08-14 20:17:39 +03:00
Matysh 0f8d35f516 docs: run the AC-named smokes locally before S7-code-review
The owner's machine now carries Playwright with Chromium on Windows and a full
WSL environment — verified by execution: 34/34 smoke assertions, and 242 backend
tests passed where native Windows silently skips every test_ha_* file. A red
smoke that reaches the review costs a cycle of forty-five minutes plus the
return trip; run locally it costs a minute, and #89 already paid that price
once.

WSL runs of the full harness and golden verify are advisory. The canon does not
move: the beta gate is CI at the exact SHA, and baselines are accepted only via
golden:accept --reviewed on a complete Linux CI artefact.

Issue: #151
User-Visible: no
2026-08-14 19:50:47 +03:00
Matysh 737e7b62aa fix: the golden mutant guard runs capture, because verify forbids one scene
First real run of the gate failed before reaching a single mutant: the clean
run of the golden guard was red on untouched code. demo/golden/policy.mjs
refuses `verify --scenario=...` on purpose — a partial verify is the "make CI
green" loophole the policy exists to close. The gate built to catch dishonest
tests had reached for a dishonest shortcut, and the policy caught it.

capture keeps the whole check: a failed semantic assertion becomes status
error, and goldenRunFailed treats an error as failure in either mode. The scene
carries warmPixelRegion with minPixels 2500 over the receiving half, so a lamp
moved out of reach still fails it — which is exactly what this mutant asserts.

Issue: #85
User-Visible: no
2026-08-14 15:18:46 +03:00
Matysh 0cf10613f2 ci: mutation-gate must live on the default branch to be dispatchable
Validate / hacs (push) Failing after 15s
Validate / hassfest (push) Failing after 13s
Validate / provenance (push) Successful in 43s
Validate / process-gate (push) Failing after 44s
Validate / frontend (push) Successful in 6m54s
Validate / backend (push) Failing after 10m31s
Validate / golden (push) Failing after 9m38s
Validate / performance_smoke (push) Failing after 10m29s
Validate / smoke (push) Failing after 29m50s
Full Performance / performance (push) Failing after 1h43m37s
gh workflow run answered 404: workflow_dispatch and schedule both resolve the
workflow file against the default branch, and the file sat only in dev. The
same trap as the process pipeline — even documented in that file's header — and
still stepped in a second time. The job itself checks out dev, so running from
main tests exactly the code it should.

Issue: #85
User-Visible: no
2026-08-14 15:08:31 +03:00
Matysh e894ce2986 docs: fix the working-tree layout in writing
Two agents sharing one checkout share one HEAD, and twice in an hour a commit
landed on someone else's task branch that way. The layout that ends it: the main
clone belongs to the author and its task branches, hp-dev is the owner's
permanent worktree on dev, and the reviewer and the infrastructure agent own no
local tree at all — one runs in CI on a fresh checkout, the other reads through
git show and publishes through the API, so it has no HEAD to collide with.

Also recorded: a worktree is only usable on the machine that created it, because
its .git file stores an absolute path in that machine's format. We hit this in
both directions within a day.

Issue: #115
User-Visible: no
2026-08-14 12:31:12 +03:00
Matysh ac30f8913d fix: build the gate CLI path with fileURLToPath, not URL.pathname
On Windows URL.pathname yields /C:/..., which spawnSync then reads as C:\C:\...
and the whole npm test run dies in this one test. Linux CI never caught it
because both spellings coincide there — which is exactly why the canonical gate
lives on Linux and the local run is advisory.

Issue: #133
User-Visible: no
2026-08-14 12:15:06 +03:00
Matysh de46db3343 docs: restore exact wording in the moved #89 spec review
One word was mistyped while transferring the file: "на каждый HA state
update" instead of "на каждом". The moved document must match the original
byte for byte.

Issue: #142
User-Visible: no
2026-08-14 11:50:54 +03:00
Matysh 782ff54e0f docs: remove originals after the move to docs/reviews and legacy
Issue: #142
User-Visible: no
2026-08-14 11:41:05 +03:00
Matysh b989c84b71 docs: remove originals after the move to docs/reviews and legacy
Issue: #142
User-Visible: no
2026-08-14 11:40:57 +03:00
Matysh 400ca7043e docs: remove originals after the move to docs/reviews and legacy
Issue: #142
User-Visible: no
2026-08-14 11:40:48 +03:00
Matysh 9f5d729538 docs: remove originals after the move to docs/reviews and legacy
Issue: #142
User-Visible: no
2026-08-14 11:40:36 +03:00
Matysh 8eb4bab7c6 docs: file reviews where reviews live, retire the #89 draft, honest markers
Three review documents sat in the repository root, committed before the pipeline
existed and before docs/reviews/ did. The directory exists now and the pipeline
writes into it, so they move there and the root stops being a second place to
look.

The #89 spec had a twin: the research draft next to the normative stage1
document, two files for one issue. The draft goes to legacy — it fed the
decisions and is worth keeping, but nothing should read it as current.

ROADMAP.md carried a live link to the Project v2 board that was dropped
yesterday; missed then because the sweep grepped for status-canon wording, not
for every link. And docs/README.ru.md said "verified against v1.60.0" as if
that were fresh — the line is now an explicit warning naming what to trust
instead: USER-GUIDE.ru.md and the changelogs.

Issue: #142
User-Visible: no
2026-08-14 11:40:26 +03:00
Matysh e9a148315a docs: restore the paragraph lost while publishing PROCESS.md
Three lines of section 10.4 and the trailing newline went missing in transit.
The lost paragraph is the one that says a label which did not change means the
run failed rather than the work — the sentence that tells a waiting author to
read the logs instead of polling for another forty-five minutes. Losing exactly
that one while copying a document about silent failures is a joke the situation
made on its own.

Caught by the byte comparison that follows every publish, which is the whole
reason it follows every publish.

Issue: #139
User-Visible: no
2026-08-14 11:11:41 +03:00
Matysh fb4096f67b chore: drop Project v2 from the process, the docs and the release script
The owner stopped using GitHub Projects. Most of this is wording, but one part
was not: release-prerelease.mjs talked to the Project in code. finishIssues
looked up the project id, listed its items and its Status=Done option, and threw
when an issue was missing from the board — so the first release that closed an
issue would have died on a step with nothing to do with publishing. Found by
reading rather than by releasing, which was luck.

Closing issues stays, and now strips the status label first. That order is not
cosmetic: the invariant that a closed issue carries no status label has broken
twice already, both times because a manual step did it the other way round. The
close-merged job already does it in this order.

The documents now say labels and only labels. The explicit "no longer used"
lines are kept on purpose, in PROCESS.md and next to the code that used to sync:
a decision that vanishes quietly gets reintroduced a month later by someone who
never knew it was made.

Issue: #139
User-Visible: no
2026-08-14 10:59:09 +03:00
Matysh e83da25085 chore: drop Project v2 from the process, the docs and the release script
The owner stopped using GitHub Projects. Most of this is wording, but one part
was not: release-prerelease.mjs talked to the Project in code. finishIssues
looked up the project id, listed its items and its Status=Done option, and threw
when an issue was missing from the board — so the first release that closed an
issue would have died on a step with nothing to do with publishing. Found by
reading rather than by releasing, which was luck.

Closing issues stays, and now strips the status label first. That order is not
cosmetic: the invariant that a closed issue carries no status label has broken
twice already, both times because a manual step did it the other way round. The
close-merged job already does it in this order.

The documents now say labels and only labels. The explicit "no longer used"
lines are kept on purpose, in PROCESS.md and next to the code that used to sync:
a decision that vanishes quietly gets reintroduced a month later by someone who
never knew it was made.

Issue: #139
User-Visible: no
2026-08-14 10:54:34 +03:00
Matysh ae10b2861b chore: drop Project v2 from the process, the docs and the release script
The owner stopped using GitHub Projects. Most of this is wording, but one part
was not: release-prerelease.mjs talked to the Project in code. finishIssues
looked up the project id, listed its items and its Status=Done option, and threw
when an issue was missing from the board — so the first release that closed an
issue would have died on a step with nothing to do with publishing. Found by
reading rather than by releasing, which was luck.

Closing issues stays, and now strips the status label first. That order is not
cosmetic: the invariant that a closed issue carries no status label has broken
twice already, both times because a manual step did it the other way round. The
close-merged job already does it in this order.

The documents now say labels and only labels. The explicit "no longer used"
lines are kept on purpose, in PROCESS.md and next to the code that used to sync:
a decision that vanishes quietly gets reintroduced a month later by someone who
never knew it was made.

Issue: #139
User-Visible: no
2026-08-14 10:50:56 +03:00
Matysh eef3634f23 chore: drop Project v2 from the process, the docs and the release script
The owner stopped using GitHub Projects. Most of this is wording, but one part
was not: release-prerelease.mjs talked to the Project in code. finishIssues
looked up the project id, listed its items and its Status=Done option, and threw
when an issue was missing from the board — so the first release that closed an
issue would have died on a step with nothing to do with publishing. Found by
reading rather than by releasing, which was luck.

Closing issues stays, and now strips the status label first. That order is not
cosmetic: the invariant that a closed issue carries no status label has broken
twice already, both times because a manual step did it the other way round. The
close-merged job already does it in this order.

The documents now say labels and only labels. The explicit "no longer used"
lines are kept on purpose, in PROCESS.md and next to the code that used to sync:
a decision that vanishes quietly gets reintroduced a month later by someone who
never knew it was made.

Issue: #139
User-Visible: no
2026-08-14 10:48:58 +03:00
Matysh 6ecbedfb85 ci: heavy Validate jobs run only where relevant paths changed
Validate / changes (push) Successful in 52s
Validate / process-gate (push) Failing after 1m17s
Validate / provenance (push) Successful in 1m20s
Validate / hacs (push) Failing after 16s
Validate / hassfest (push) Failing after 13s
Validate / frontend (push) Successful in 7m30s
Validate / backend (push) Failing after 8m41s
Validate / performance_smoke (push) Failing after 9m45s
Validate / golden (push) Failing after 14m25s
Validate / smoke (push) Failing after 28m53s
Every push to every branch ran 128 browser smokes, 50 golden scenes, a Home
Assistant install and a performance pass — including a push that added one spec
file. The pipeline made such pushes routine: every spec revision and every
review document is a push to a task branch and used to cost the full suite.

A changes job classifies the push range; frontend, smoke, golden, performance
and backend now run only when their paths moved, and hacs and hassfest only for
manifests, translations or Python. provenance and process-gate always run — they
judge commits, not code.

The exception carries the design. On dev everything runs, always, unfiltered:
the beta gate accepts "green Validate at the exact SHA", and if the volume of a
run depends on the diff, green stops meaning one thing — a release candidate
touches manifests and changelogs, would skip the browser suites under filtering,
and a run with skipped jobs still concludes success. That would be the sixth
silent success of the week. Filters save time on task branches, where Validate is
an early signal and the real acceptance is the code review running gates itself.

A new branch with a zero before-sha is classified from the merge-base with dev,
not from the root of history.

Issue: #136
User-Visible: no
2026-08-14 10:16:18 +03:00
Matysh e6366b6548 fix: load virtual_lights by path so offline collection survives
The file declared itself pure but imported the module through the package, and
the package __init__ unconditionally imports homeassistant. Without homeassistant
installed pytest did not skip the file — it stopped collecting the whole
tests_backend directory, taking the previously working pure suite down with it.
test_validation.py had already established the by-path pattern; virtual_lights.py
imports nothing beyond the standard library, so it loads cleanly.

The async tests also dropped their pytest-asyncio dependency in favour of
asyncio.run: the offline environment does not carry the plugin, and without it
the two tests failed as unsupported async defs. The offline gate has to be green,
or nobody runs it.

Verified in both environments: pytest+voluptuous only — 129 passed where
collection previously stopped dead; with pytest-asyncio as in CI — 129 passed.

Issue: #135
User-Visible: no
2026-08-14 10:10:34 +03:00
Matysh bc98116a31 test: a registry of known breakages that tests must catch
Five times in this project a green test meant nothing was checked. The
continuity smoke stayed green after the entire mechanism it guards was cut out.
The golden scene created to protect doorway light was empty — 1,177 warm pixels
against 107,119, all of them icons. The shadow smoke passed while no shadow was
drawn. Each time the test had been written alongside the code, went green at
once, and nobody ever asked whether it could go red.

The gate makes that question routine. Each mutant is a few lines of patch that
reproduce a known breakage, plus the name of the test that must fail on it. A
worktree is patched, the bundle rebuilt, the guard run — and a guard that stays
green fails the gate. Six mutants cover the holes documented in #85; the anchors
are exact strings from today's source, so the registry cannot silently drift —
a unit test that runs with the ordinary suite refuses a stale anchor.

The full run rebuilds the bundle per mutant, so it lives in its own workflow,
before a stable release and on a weekly schedule, not in Validate. The rules for
new tests are written at the top of docs/TESTING.md, and the sixth of them is
the cheapest: an assertion that reads back the property the code just set is
not written at all.

Issue: #85
User-Visible: no
2026-08-14 02:15:12 +03:00
Matysh 328ed7afc0 test: a registry of known breakages that tests must catch
Validate / provenance (push) Successful in 46s
Validate / hacs (push) Failing after 10s
Validate / hassfest (push) Failing after 12s
Validate / process-gate (push) Failing after 35s
Validate / frontend (push) Successful in 6m11s
Validate / backend (push) Failing after 9m24s
Validate / golden (push) Failing after 10m10s
Validate / performance_smoke (push) Failing after 12m46s
Validate / smoke (push) Failing after 30m15s
Five times in this project a green test meant nothing was checked. The
continuity smoke stayed green after the entire mechanism it guards was cut out.
The golden scene created to protect doorway light was empty — 1,177 warm pixels
against 107,119, all of them icons. The shadow smoke passed while no shadow was
drawn. Each time the test had been written alongside the code, went green at
once, and nobody ever asked whether it could go red.

The gate makes that question routine. Each mutant is a few lines of patch that
reproduce a known breakage, plus the name of the test that must fail on it. A
worktree is patched, the bundle rebuilt, the guard run — and a guard that stays
green fails the gate. Six mutants cover the holes documented in #85; the anchors
are exact strings from today's source, so the registry cannot silently drift —
a unit test that runs with the ordinary suite refuses a stale anchor.

The full run rebuilds the bundle per mutant, so it lives in its own workflow,
before a stable release and on a weekly schedule, not in Validate. The rules for
new tests are written at the top of docs/TESTING.md, and the sixth of them is
the cheapest: an assertion that reads back the property the code just set is
not written at all.

Issue: #85
User-Visible: no
2026-08-14 02:05:53 +03:00
Matysh 888e90450a perf: make review scope and ceremony fit the size of the task
The owner's report: the process works but every stage takes a long time even on
simple bugs. Two causes, and neither was the one that first comes to mind.

The reviewer ran everything regardless. On #89 it installed Chromium, ran all 127
smoke files and a full golden capture — right for a task rated 10/10 for
complexity, absurd for a bug about a room divider. Full suites are the pre-beta
gate; the review now runs typecheck, unit and build always, and smokes, golden,
pytest or performance only where the diff and the AC call for them. The price of
narrowing it is honesty: the reviewer must list which gates it ran, which it did
not, and why, so a skipped gate is a visible decision rather than a silent one.

The reviewer also built its own environment out of model turns, with no npm cache
and no browser cache, paid for from the same forty-five minutes. The workflow now
installs dependencies and Chromium as ordinary cached steps, after switching to
the task branch so the lockfile is the branch's own.

Second, ceremony did not scale down. The light track makes a spec cheap; the new
trivial track does without one — S2-analysis straight to S5-ready, no spec review,
AC in the issue body. It is deliberately hard to qualify for: a bug on one surface,
no new UX contract, no migration, no i18n, no perf or touch effect, three checkable
AC at most, and expected behaviour already on record. Nothing left to decide is the
criterion that holds the whole thing up, and it cannot be met by feeling sure.

Code review is never skipped on either track. It is what stands in for testing
here, so it is the one stage speed may not buy.

Issue: #127
Issue: #128
User-Visible: no
2026-08-13 22:07:42 +03:00
Matysh 565f518dcd perf: make review scope and ceremony fit the size of the task
The owner's report: the process works but every stage takes a long time even on
simple bugs. Two causes, and neither was the one that first comes to mind.

The reviewer ran everything regardless. On #89 it installed Chromium, ran all 127
smoke files and a full golden capture — right for a task rated 10/10 for
complexity, absurd for a bug about a room divider. Full suites are the pre-beta
gate; the review now runs typecheck, unit and build always, and smokes, golden,
pytest or performance only where the diff and the AC call for them. The price of
narrowing it is honesty: the reviewer must list which gates it ran, which it did
not, and why, so a skipped gate is a visible decision rather than a silent one.

The reviewer also built its own environment out of model turns, with no npm cache
and no browser cache, paid for from the same forty-five minutes. The workflow now
installs dependencies and Chromium as ordinary cached steps, after switching to
the task branch so the lockfile is the branch's own.

Second, ceremony did not scale down. The light track makes a spec cheap; the new
trivial track does without one — S2-analysis straight to S5-ready, no spec review,
AC in the issue body. It is deliberately hard to qualify for: a bug on one surface,
no new UX contract, no migration, no i18n, no perf or touch effect, three checkable
AC at most, and expected behaviour already on record. Nothing left to decide is the
criterion that holds the whole thing up, and it cannot be met by feeling sure.

Code review is never skipped on either track. It is what stands in for testing
here, so it is the one stage speed may not buy.

Issue: #127
Issue: #128
User-Visible: no
2026-08-13 21:59:55 +03:00
Matysh 516257e322 perf: make review scope and ceremony fit the size of the task
The owner's report: the process works but every stage takes a long time even on
simple bugs. Two causes, and neither was the one that first comes to mind.

The reviewer ran everything regardless. On #89 it installed Chromium, ran all 127
smoke files and a full golden capture — right for a task rated 10/10 for
complexity, absurd for a bug about a room divider. Full suites are the pre-beta
gate; the review now runs typecheck, unit and build always, and smokes, golden,
pytest or performance only where the diff and the AC call for them. The price of
narrowing it is honesty: the reviewer must list which gates it ran, which it did
not, and why, so a skipped gate is a visible decision rather than a silent one.

The reviewer also built its own environment out of model turns, with no npm cache
and no browser cache, paid for from the same forty-five minutes. The workflow now
installs dependencies and Chromium as ordinary cached steps, after switching to
the task branch so the lockfile is the branch's own.

Second, ceremony did not scale down. The light track makes a spec cheap; the new
trivial track does without one — S2-analysis straight to S5-ready, no spec review,
AC in the issue body. It is deliberately hard to qualify for: a bug on one surface,
no new UX contract, no migration, no i18n, no perf or touch effect, three checkable
AC at most, and expected behaviour already on record. Nothing left to decide is the
criterion that holds the whole thing up, and it cannot be met by feeling sure.

Code review is never skipped on either track. It is what stands in for testing
here, so it is the one stage speed may not buy.

Issue: #127
Issue: #128
User-Visible: no
2026-08-13 21:51:42 +03:00
Matysh 9177c9a944 perf: make review scope and ceremony fit the size of the task
The owner's report: the process works but every stage takes a long time even on
simple bugs. Two causes, and neither was the one that first comes to mind.

The reviewer ran everything regardless. On #89 it installed Chromium, ran all 127
smoke files and a full golden capture — right for a task rated 10/10 for
complexity, absurd for a bug about a room divider. Full suites are the pre-beta
gate; the review now runs typecheck, unit and build always, and smokes, golden,
pytest or performance only where the diff and the AC call for them. The price of
narrowing it is honesty: the reviewer must list which gates it ran, which it did
not, and why, so a skipped gate is a visible decision rather than a silent one.

The reviewer also built its own environment out of model turns, with no npm cache
and no browser cache, paid for from the same forty-five minutes. The workflow now
installs dependencies and Chromium as ordinary cached steps, after switching to
the task branch so the lockfile is the branch's own.

Second, ceremony did not scale down. The light track makes a spec cheap; the new
trivial track does without one — S2-analysis straight to S5-ready, no spec review,
AC in the issue body. It is deliberately hard to qualify for: a bug on one surface,
no new UX contract, no migration, no i18n, no perf or touch effect, three checkable
AC at most, and expected behaviour already on record. Nothing left to decide is the
criterion that holds the whole thing up, and it cannot be met by feeling sure.

Code review is never skipped on either track. It is what stands in for testing
here, so it is the one stage speed may not buy.

Issue: #127
Issue: #128
User-Visible: no
2026-08-13 21:33:36 +03:00
Matysh 8a3f6efa0a fix: the review document is published even without a task branch
Issues labelled before the pipeline existed keep their spec straight in dev and
have no issue/NN branch. The publish step quietly exited zero for them, so the
verdict would arrive as a comment and the analysis behind it would be thrown
away — the fifth instance today of a step reporting success by doing nothing.

The document now goes wherever the spec itself lives: the task branch when there
is one, dev otherwise. Publishing also survives dev moving on while the review
ran, which takes up to forty-five minutes, by rebasing once before it gives up.

Four issues are waiting on this — #12, #30, #44 and #52 — each with a spec in dev,
a status label applied during the bulk pass in August and a review that never ran
because nothing was there to raise the event.

Issue: #114
User-Visible: no
2026-08-13 21:11:41 +03:00
Matysh be7d6b9706 fix: the review document is published even without a task branch
Issues labelled before the pipeline existed keep their spec straight in dev and
have no issue/NN branch. The publish step quietly exited zero for them, so the
verdict would arrive as a comment and the analysis behind it would be thrown
away — the fifth instance today of a step reporting success by doing nothing.

The document now goes wherever the spec itself lives: the task branch when there
is one, dev otherwise. Publishing also survives dev moving on while the review
ran, which takes up to forty-five minutes, by rebasing once before it gives up.

Four issues are waiting on this — #12, #30, #44 and #52 — each with a spec in dev,
a status label applied during the bulk pass in August and a review that never ran
because nothing was there to raise the event.

Issue: #114
User-Visible: no
2026-08-13 21:05:17 +03:00
Matysh 7c1edbfa9b ci: realign the workflow copy in dev with main
The two copies of this file must match byte for byte; a comment line had drifted
by one character. main is the copy the issues event actually reads, so it is the
reference. Trivial in itself, and worth closing anyway: the file's own header
warns that a divergence between these two branches is one of the ways this
pipeline fails quietly.

Issue: #114
User-Visible: no
2026-08-13 20:53:35 +03:00
Matysh 024cdc0d94 feat: an outsider's issue is worked like any other once admitted
The guard refused to review any issue the owner had not filed himself. The rule
was meant to keep malformed outside reports out of the pipeline, but it checked at
every step instead of at the entrance, and it duplicated a guarantee the platform
already gives: only someone with write access can apply a label. Applying the
first status label is the owner's explicit decision, and it is the only place the
question belongs.

So the author check is gone. While an issue carries no status label it sits
outside the process and the invariants do not apply; once labelled, the task is in
flight and who filed it stops mattering.

The old rule also cost real work. On #123 an outside bug report had been analysed
and specified before the guard turned it away in nine seconds, and the remedy on
offer was to refile the same thing as the owner's own issue.

Issue: #114
User-Visible: no
2026-08-13 20:49:04 +03:00
Matysh 2fd042a7de feat: an outsider's issue is worked like any other once admitted
The guard refused to review any issue the owner had not filed himself. The rule
was meant to keep malformed outside reports out of the pipeline, but it checked at
every step instead of at the entrance, and it duplicated a guarantee the platform
already gives: only someone with write access can apply a label. Applying the
first status label is the owner's explicit decision, and it is the only place the
question belongs.

So the author check is gone. While an issue carries no status label it sits
outside the process and the invariants do not apply; once labelled, the task is in
flight and who filed it stops mattering.

The old rule also cost real work. On #123 an outside bug report had been analysed
and specified before the guard turned it away in nine seconds, and the remedy on
offer was to refile the same thing as the owner's own issue.

Issue: #114
User-Visible: no
2026-08-13 20:41:36 +03:00
Matysh 4e539b02df fix: the guard says why it refused, in the issue
A review label promises work. When the guard declined it wrote the reason to the
run log and nothing else, so the issue sat in a status nobody was acting on and
nobody could tell. #123 showed it: an outside reporter's issue was walked up to
S4-spec-review, the guard refused in nine seconds because only the owner's issues
enter the process, and the issue itself said not a word.

Refusals that a human can act on now become a comment: wrong author, blocked,
review-4. Only when a stage was actually recognised, so an unrelated label change
stays silent.

This is the same defect as the merge conflict that left the label untouched, seen
from the other side. The pattern is worth naming: doing nothing quietly is the
most expensive thing a pipeline can do.

Issue: #114
User-Visible: no
2026-08-13 20:36:25 +03:00
Matysh d7e2c4d4f0 fix: the guard says why it refused, in the issue
A review label promises work. When the guard declined it wrote the reason to the
run log and nothing else, so the issue sat in a status nobody was acting on and
nobody could tell. #123 showed it: an outside reporter's issue was walked up to
S4-spec-review, the guard refused in nine seconds because only the owner's issues
enter the process, and the issue itself said not a word.

Refusals that a human can act on now become a comment: wrong author, blocked,
review-4. Only when a stage was actually recognised, so an unrelated label change
stays silent.

This is the same defect as the merge conflict that left the label untouched, seen
from the other side. The pattern is worth naming: doing nothing quietly is the
most expensive thing a pipeline can do.

Issue: #114
User-Visible: no
2026-08-13 20:30:15 +03:00
Matysh 948f2848dd docs: a failed pre-release gate does not send the issue back to review
The implementation loop runs typecheck, unit and build. Golden, browser smokes,
performance and the full HA harness run before a beta — after the code review has
passed and the issue already sits in S8-merged. Some defects cannot surface any
earlier, and until now the process had nothing to say about them, so the honest
reading was a second full review cycle at the most expensive possible moment.

The owner's decision: fix it, re-run what failed, and a green run carries the
release on. The gate named the defect precisely and the same gate proves the fix,
so the check is objective and depends on nobody's judgement.

The boundary is written down with it, because "the gate found something" could
otherwise absorb an arbitrary amount of new work. A fix that changes a behaviour
contract, reaches an untouched subsystem or rivals the task in size goes through
the normal flow. Editing a test so it stops failing is concealment rather than
repair — the exception is a defect proven to be in the fixture, as on #89.

The rule also records what it costs: the author judges his own work here, which
the process refuses everywhere else. That is the price of speed at the one point
where a review cycle is dearest, and the compensation is that the re-run command
and its result are written into the issue where the release manager reads them.

Issue: #114
User-Visible: no
2026-08-13 19:36:39 +03:00
Matysh dbe12f1a54 docs: state the invariant the pipeline was missing
A review run always moves the label. The rule is written down because its absence
cost a real stall: a green code review whose merge conflicted left the label alone,
the waiting author polled thirty times and reported the limit as exhausted, and a
verdict that existed reached nobody.

Both documents now say what S6-in-progress means when the verdict was green and
only the merge failed — rebase, not rework, and the verdict still stands. They also
say that a label which did not change means the run failed rather than the work, so
the answer is logs and the owner, not more polling. Cycles are counted per stage.

Issue: #114
User-Visible: no
2026-08-13 17:13:54 +03:00
Matysh 1e9952db35 fix: a review run always moves the label, conflict or not
A green code review whose merge conflicted used to leave the label where it was.
That is a dead end: the author waits for the label to change, so it polled thirty
times and reported the limit as exhausted — on a task the reviewer had already
passed. The verdict existed and nobody could act on it.

The merge step no longer fails the job. It reports whether it merged, and a green
review that did not merge sends the task back to S6-in-progress, because the work
did return to the author — a rebase rather than a code fix, and the comment says
so and says the verdict still stands.

The invariant is now stronger and worth stating plainly: after a review run the
label always changes. A pipeline whose state can stall silently is worse than one
that reports the wrong state loudly.

Issue: #114
User-Visible: no
2026-08-13 17:03:43 +03:00
Matysh d1be6891b2 fix: a review run always moves the label, conflict or not
Validate / hacs (push) Failing after 55s
Validate / hassfest (push) Failing after 13s
Validate / frontend (push) Successful in 6m16s
Validate / backend (push) Failing after 8m39s
Validate / provenance (push) Successful in 37s
Validate / golden (push) Failing after 8m4s
Validate / smoke (push) Failing after 27m15s
Validate / performance_smoke (push) Failing after 14m5s
Full Performance / performance (push) Failing after 1h15m11s
A green code review whose merge conflicted used to leave the label where it was.
That is a dead end: the author waits for the label to change, so it polled thirty
times and reported the limit as exhausted — on a task the reviewer had already
passed. The verdict existed and nobody could act on it.

The merge step no longer fails the job. It reports whether it merged, and a green
review that did not merge sends the task back to S6-in-progress, because the work
did return to the author — a rebase rather than a code fix, and the comment says
so and says the verdict still stands.

The invariant is now stronger and worth stating plainly: after a review run the
label always changes. A pipeline whose state can stall silently is worse than one
that reports the wrong state loudly.

Issue: #114
User-Visible: no
2026-08-13 16:58:18 +03:00
Matysh 316ee76a29 fix: repair the line continuation in the failure handler
The step that comments on the issue when a review run dies carried a literal
backslash instead of a line continuation, so gh received four arguments and
--repo ran as a command of its own. The handler for failures would itself have
failed, silently, and only when something had already gone wrong.

bash -n does not catch this: the syntax is valid, the meaning is not. Checking
run blocks now also means looking for a doubled backslash at end of line.

Issue: #114
User-Visible: no
2026-08-13 16:21:33 +03:00
Matysh 9be81c1413 fix: repair the line continuation in the failure handler
The step that comments on the issue when a review run dies carried a literal
backslash instead of a line continuation, so gh received four arguments and
--repo ran as a command of its own. The handler for failures would itself have
failed, silently, and only when something had already gone wrong.

bash -n does not catch this: the syntax is valid, the meaning is not. Checking
run blocks now also means looking for a doubled backslash at end of line.

Issue: #114
User-Visible: no
2026-08-13 16:16:39 +03:00
Matysh 869fe169d8 fix: count review cycles per stage, not across the whole issue
The guard counted every verdict comment on the issue, so a spec-review verdict
consumed a cycle from the code-review budget. On #89 the first code review came
out as r2/4. With two spec cycles the second code review would have hit review-4
after a single fix — the limit would have fired on a task nobody had reviewed
twice.

The stage is now resolved first and only its own verdicts are counted, recognised
by the review document named in the comment. If the document is missing the
verdict is not counted: undercounting grants an extra cycle, overcounting would
stop the work early, and of the two mistakes the recoverable one wins.

Issue: #114
User-Visible: no
2026-08-13 16:13:13 +03:00
Matysh fafeca4540 fix: count review cycles per stage, not across the whole issue
The guard counted every verdict comment on the issue, so a spec-review verdict
consumed a cycle from the code-review budget. On #89 the first code review came
out as r2/4. With two spec cycles the second code review would have hit review-4
after a single fix — the limit would have fired on a task nobody had reviewed
twice.

The stage is now resolved first and only its own verdicts are counted, recognised
by the review document named in the comment. If the document is missing the
verdict is not counted: undercounting grants an extra cycle, overcounting would
stop the work early, and of the two mistakes the recoverable one wins.

Issue: #114
User-Visible: no
2026-08-13 16:08:20 +03:00
Matysh d38a5be68b docs: drop the second status dictionary and the stale class note
docs/specs/README.md kept a "Статус ТЗ" column with its own vocabulary — draft,
in implementation, done — next to the labels that already hold the status. Two
dictionaries for one fact drift apart, and these had: the column still called
issues "in implementation" that were closed weeks ago. The table now says only
which issue a spec belongs to.

AGENTS.md was telling agents that PROCESS.md §1 does not cover package.json and
the rest of the configuration, and to report it as missing. It covers them now.
The same paragraph gained the rule that D beats A where paths overlap, which is
what keeps the built bundle under custom_components/houseplan/frontend/ from
reading as product source.

Issue: #119
User-Visible: no
2026-08-13 15:02:42 +03:00
Matysh 42335bc16d ci: add the pre-push gate and stop lying about it in the canon
Section 10.1 promised pre-push as the blocking gate that replaces pull requests.
The hook did not exist, so the document promised a check that was not there —
worse than saying nothing, because a promise like that gets relied on. Until now
process-gate ran only as the catch-up job in CI, which reports after the code is
already in dev.

The hook skips branch deletions and tags, and for a branch the remote has not
seen it measures from the merge-base with dev rather than from the root, or every
violation committed before the gate existed would make it impossible to pass. A
missing script does not block a push: old checkouts and worktrees have to stay
usable.

gh is optional on purpose. Reading issue status needs the network, and a hook
that cannot work on a train is a hook people switch off; offline it runs what it
can and CI does the strict pass.

The executable bit is the quiet part. Git skips a hook without +x and says
nothing — the gate reports success by being absent. Measured on a real push:
mode 644 produces zero lines from the gate and the push goes through, 755 stops
it. The API cannot set the bit, so install-hooks restores it on every install.

Issue: #121
User-Visible: no
2026-08-13 14:58:57 +03:00
Matysh a36b3129f6 ci: close the S8-merged queue when a beta is published
PROCESS.md 10.2 item 10 asks for this to happen because a beta shipped, not
because someone remembered. The manual cleanup was skipped twice and both times
it broke the invariant that a closed issue carries no status label — the one
thing `verify` leans on. A manual step that falls due right after a successful
release is the worst kind: the work already looks finished, which is precisely
why it gets forgotten.

The job comments the tag, removes the label, then closes. That order is
deliberate: dying between the two steps leaves an open issue without a status,
which is visible and fixable in the flow, where the reverse order would recreate
the breakage this exists to prevent. It ends by asserting that no closed issue
still carries S8-merged — aimed at the defect that actually recurs rather than at
the invariant in general.

The stock token is used on purpose. Events caused by GITHUB_TOKEN do not start
workflows, so stripping the label cannot wake the review pipeline; a PAT here
would turn bookkeeping into a cascade.

Issue: #120
User-Visible: no
2026-08-13 14:43:32 +03:00
Matysh 8cecaf2c5e fix: judge the branch rule only by the branch's own commits
Check 2 compared the Issue trailers against whatever branch the working tree
happened to be on, over whatever range it was given. Those two are not the same
set. After a rebase the CI range widens — `before` points at a discarded commit,
the merge-base slides back, and commits that belong to dev arrive carrying other
issue numbers. Every one of them then looks like a violation.

Running the gate over real history from issue/89 with a dev range produced 26
false refusals out of 26 commits, which would have reddened Validate on the next
force-push of any task branch.

The rule now reads origin/dev..HEAD for its own verdict and leaves the event
range to the other checks. A commit that genuinely carries the wrong trailer for
its branch is still caught; the integration test covers both directions.

Issue: #105
User-Visible: no
2026-08-13 14:30:13 +03:00
Matysh 5aa8771dc3 docs: bring the process canon back in line with what actually runs
The canon moved into the repository in #112 and then stood still while the
process kept moving. A document that lags is worse than no document: an agent
reading it as truth acts on rules that no longer exist. It promised a pre-push
hook that was never written, named labels in Russian that the repository has
never used, listed a status set the gate no longer applies, and said nothing at
all about the event-driven pipeline — the largest mechanism the process has.

Label names are now English throughout and S8-merged is documented. Section 1
covers the configuration files the gate kept reporting as unclassified, and
records that D beats A where paths overlap, since the built bundle lives inside
custom_components/houseplan/frontend/. Section 10.1 admits that pre-push does
not exist. Section 10.2 matches ALLOWED_STATUS in scripts/process-gate.mjs,
including the two caveats that only surfaced once the pipeline ran. Section 10.4
is new and documents the four silent-failure traps that cost a working day each.

The source-of-truth order now says that actual automation outranks its own
description — this document included.

Issue: #119
User-Visible: no
2026-08-13 14:24:50 +03:00
Matysh 3ade633538 ci: merge into dev before setting S8-merged
The label asserts the code is in dev. The workflow used to set it on a green
code review while the commits were still only on the task branch, so between
the verdict and the author's merge the state machine stated something untrue —
which is exactly what happened on #104.

The merge now runs inside the pipeline, before the label. A conflict leaves
the issue in S7-code-review and comments instead.

Issue: #114
User-Visible: no
2026-08-13 13:20:26 +03:00
Matysh 68596a75a0 ci: the reviewer writes a review document to the task branch
PROCESS.md wants a review document in docs/reviews/; the CI reviewer could
only leave a comment, and flagged the gap itself. It may now write there.

What lands in the commit is decided by the workflow, not by the model: every
path outside docs/reviews/ is reverted before staging, and the commit carries
the usual trailers so the provenance gate accepts it.

Issue: #114
User-Visible: no
2026-08-13 13:04:26 +03:00
Matysh 65a86db122 fix(ci): repair the failure comment step
The multi-line --body started at column zero, which ends the YAML block
scalar. The parser silently truncated the run script and left an unclosed
double quote, so the whole workflow became unusable and blocked the code
review on #104.

The body now goes through a heredoc. Validating YAML alone did not catch
this; every run block is checked with bash -n from now on.

Issue: #114
User-Visible: no
2026-08-13 12:57:04 +03:00
Matysh a29df12e0b ci: raise the turn limit, bound the run by time instead
The r2 spec review on #104 produced a complete green verdict and then failed
on --max-turns 40 at turn 43, so the label step never ran and the transition
had to be reconciled by hand. Forty was a guess; a review that reads SCOPE,
AGENTS, PROCESS, the issue thread and the spec exceeds it routinely, and a
code review that also runs gates needs far more.

The real guard against a runaway run is the job timeout, not the turn count.

Issue: #114
User-Visible: no
2026-08-13 12:37:14 +03:00
Matysh 0509e1a008 docs: only product ambiguity goes to the owner
Agents were escalating technical calls. The owner answers what a person sees
or does and how much user-visible change belongs in an issue; storage, module
layout, naming, test strategy and migration mechanics are settled by the
agents, recorded as assumptions and challenged in review.

Issue: #114
User-Visible: no
2026-08-13 12:18:57 +03:00
Matysh e043974c44 docs: standing permission to merge a reviewed branch into dev
Without it S8-merged would be a lie: the label asserts the code is in dev,
while the branch-only push leaves it on the task branch. What lands is what
the reviewer just accepted, and dev is allowed to carry unreviewed code
anyway, so the merge adds no risk the branch did not already carry.

Issue: #114
User-Visible: no
2026-08-13 12:11:04 +03:00
Matysh 05b38e67c4 fix(hooks): drop basename from commit-msg
The hook parsed the message path with basename, an external command. Where
it is missing from PATH the substitution yields an empty string, set -e does
not trip on it, and the MERGE_MSG guard silently stops working — a generated
merge commit would then be rejected for missing trailers it cannot have.

POSIX parameter expansion needs no external command and behaves the same in
sh, dash, bash and Git Bash.

Issue: #116
User-Visible: no
2026-08-13 12:09:02 +03:00
Matysh 9419842333 fix(hacs): keep exactly one *manifest.json in the tree
Validate / hassfest (push) Failing after 14s
Validate / frontend (push) Successful in 8m18s
Validate / backend (push) Failing after 10m34s
Validate / smoke (push) Failing after 24m45s
Validate / performance_smoke (push) Failing after 9m42s
Full Performance / performance (push) Failing after 54m31s
Validate / golden (push) Failing after 6m50s
Validate / hacs (push) Failing after 11s
The HACS submission check does not read hacs.json to find the integration: it
globs `*manifest.json` over the whole clone of the default branch and exits 1
unless there is exactly one (hacs/default, scripts/helpers/integration_path.py).
Three files matched — the two stand-only integrations added on 2026-07-31 and
the golden baseline index added on 2026-08-11 — so the Hassfest job of PR #9004
went red five weeks into the review queue, with a log that named no file.

The stand manifests ship as manifest.template.json and demo/stand/install.sh
renames them at install time; the golden index becomes baselines-index.json
(the exported constant keeps its name, so no consumer changes).
test/repo-hygiene.test.mjs fails if a second manifest ever appears, and the
existing golden-policy assertion — which compared against 'manifest.json' and
happily passed on 'baseline-manifest.json' — now checks the suffix.
2026-08-11 14:03:38 +03:00
MatyshandGitHub 4c260f0ea0 Merge pull request #5 from Matysh/cursor/project-audit-6009
Validate / frontend (push) Successful in 2m3s
Validate / smoke (push) Failing after 20m45s
Validate / hacs (push) Failing after 11s
Validate / backend (push) Failing after 7m50s
Validate / hassfest (push) Failing after 11s
docs: полный аудит проекта (рынок, качество, пробелы, рекомендации)
2026-08-05 11:04:55 +03:00
MatyshandGitHub 159f4f43c0 Merge pull request #4 from Matysh/cursor/setup-dev-environment-6009
chore: document Cursor Cloud dev environment setup
2026-08-05 11:04:53 +03:00
MatyshandGitHub eea669fa12 Add files via upload 2026-07-23 21:45:45 +03:00