Windows portability (the three red tests on the owner's machine at green CI):
- scripts/spawn-portable.mjs: isMainModule via pathToFileURL (the
`file://${argv[1]}` form gives file:///C:/C:/... and the CLI stays silent);
portableCommand — a shell only for npm/npx/.cmd, node and git run directly
(spawn via shell dropped the quotes of `node -e "…"`). Applied to
classify-changes, mutation-gate-report, review-doc-guard, check-inputs,
merge-candidate, gate-small, rebase-on-dev.
- the rebase-on-dev test pins core.autocrlf=false / core.eol=lf through
GIT_CONFIG_* for its temp repository instead of touching the user's git
config; test/windows-portability.test.mjs forbids both anti-patterns.
Pins: scripts/toolchain-pins.mjs reads Node/Python from validate.yml, the HA
stack from tests_backend/requirements.txt, Playwright/Chromium from the
lockfile — no second dictionary; `npm run toolchain:check` compares the
machine; .nvmrc/.python-version are derived and tested equal;
scripts/wsl-setup.sh provisions WSL/Linux with those pins.
scripts/task-packet.mjs: one derived view of an issue (status, rights, owner
decisions, branch vs dev, Validate on the tip, previous verdict with recorded
tree, AC → evidence, unwitnessed). scripts/wait-verdict.mjs: polls labels,
pipeline comments and optionally Validate, prints only on state change, exit
0/3/4, writes nothing.
gate:small --smokes: after build the browser phase runs bundle-sync and then
the directly matched and registered smokes, two at a time; broad matches stay
with the reviewer. package.json changed, so the bundle is rebuilt here.
Issue: #496
User-Visible: no
Review pipeline (process.yml):
- concurrency moves from the workflow to the guard/review jobs and the guard
runs only for S4-spec-review / S7-code-review. Any other label used to enter
the issue's concurrency group and evict the pending review run (sample of
150 runs since 2026-09-01: 92 empty guard-only runs, 30 cancelled).
- the guard reads the issue's current labels instead of the event snapshot; a
label removed before the run starts is a withdrawn request, no comment.
- a green verdict is re-applied without calling the model when the latest
review document carries the pipeline-recorded verdict `green`/High 0 and the
tree differs from its anchor in nothing outside docs/reviews/** (#437 r4
re-reviewed an unchanged tree for 7 minutes). The verdict from
structured_output is now written into the anchor block for that purpose.
- the reviewer is pinned to the captured material SHA in the prompt; the
broken escaping in the "merge cancelled" comment (empty SHAs) is fixed.
Mutation gate: nine browser guards started with `npm run bundle:sync` although
the runner already builds the mutant bundle — a second rollup plus a
`tsc --noEmit` that fails on a non-strict mutant before the smoke even runs.
Prefix removed; `--check` refuses guards that build the bundle themselves.
Docs: SCOPE (Project v2 dropped, three editors), STATUS (#437 merged, HACS zip
automated), USER-GUIDE ru/en (static card shows live states; kiosk double tap
on free background fits all), #34 → #425 references, #367 named as closed in
bundle-budget messages, PROCESS §10.4 and AGENTS.md describe the controller.
Issue: #499
User-Visible: no
Validate ran three smoke shards, golden and performance_smoke on every push,
check-docs went red on any src/** change until screenshots were re-captured,
and a parallel bundle build made every second task branch fail to rebase.
None of these gates ever failed at review time; they fail before betas.
- `heavy` output in job `changes` (scripts/classify-changes.mjs): smoke,
smoke_done, golden, performance_smoke run only for a head commit with a
`Release:` trailer, `workflow_dispatch full=true` and pull requests.
- nightly.yml dispatches Validate on dev with full=true every night.
- check-docs `--screenshots=warn|strict`: freshness of the screenshot index
warns on a plain push, errors on the candidate; everything else still errors.
- publish-prerelease.yml and release.yml refuse a candidate without the
`Release:` trailer and (prerelease) require fresh screenshots — a green
Validate without the heavy jobs cannot pass for a release.
- scripts/rebase-on-dev.mjs: rebase on origin/dev taking dev's copy of the
committed bundle, rebuild with bundle:sync, amend; any other conflict aborts.
- npm run gate:small: mandatory PROCESS §8 part in one parallel run.
Issue: #479
User-Visible: no
Вопрос владельца: зачем агенты снимают PNG на Windows, если кадры мы всё
равно не принимаем, тем более что WSL есть на обеих машинах. Ответ по
коду: этому ничто не мешало. Ни один из шести скриптов съёмки и приёмки
не знал, на какой он ОС — ни `process.platform`, ни win32, ни WSL не
упоминались нигде. Съёмка отрабатывала штатно, а стена появлялась на
приёмке, и текст стены говорил «сцен-свидетелей 0 из 10», то есть
подсказывал неверный вывод «надо объявить больше сцен» — от которого до
`--expect-change` на всю матрицу одна команда.
Причина запрета не политика, а физика: Windows растеризует текст через
DirectWrite, с другим субпиксельным сглаживанием и DPI, поэтому
байтового совпадения с принятым эталоном не даёт никогда, свидетелей
среды быть не может, и приёмка откажет всё равно. Флаги детерминизма из
#410 убирают разброс внутри среды, а не между ОС.
Что сделано:
- `golden:capture` отказывается до запуска браузера, в тексте отказа
готовая команда для WSL;
- `golden:verify` остаётся законным в любой среде: он ничего не
принимает, а как грубая проверка полезен;
- обе приёмки (golden и документации) отказываются в чужой среде;
- к отказу свидетелей приписывается фраза про расхождение среды — та
самая, которой не хватало, чтобы отказ не читался как «объяви больше
сцен»;
- платформа уезжает в манифесты рядом с версией Chromium: у кадров
появился провенанс среды;
- осознанный обход есть и требует причину: HP_ALLOW_FOREIGN_CAPTURE.
Главный урок задачи — про цену правки файла, у которого записан хеш.
Первая редакция встроила проверку в `demo/docs/capture.mjs`, и гейт
документации сразу покраснел: его sha записан в индексе скриншотов, и
`scripts/check-docs.mjs` их сверяет. То есть проверка, которая ничего не
рисует, стоила бы пересъёмки всех картинок документации и визуальной
приёмки владельца. Поэтому для документации отказ живёт шагом раньше —
`npm run docs:capture` вызывает `scripts/assert-capture-env.mjs` — и
шагом позже, на приёмке. По той же причине гейт golden стоит в
`demo/golden/policy.mjs`, а не в `run.mjs`: последний входит в корпус
sourceFingerprint. Тест закрепляет обе границы: гейт обязан быть в
policy.mjs и в npm-скрипте и обязан отсутствовать в двух
фингерпринтуемых файлах.
Платформа в юнитах — параметр, а не `process.platform`: иначе тест был
бы зелёным на Linux и красным на машине владельца, то есть тестом про
хост, а не про правило.
AGENTS.md приведён к состоянию после #401 (принимается любая среда,
доказавшая себя байтовым совпадением непринятых кадров) и разводит
проверку и съёмку — прежний текст сливал их в «advisory» и утверждал
«accepted only on a complete Linux CI artefact».
Свидетели, все проверены отрицательным прогоном: снятый гейт съёмки,
гейт, отказывающий и на verify, обход без причины, отказ без команды,
убранная приписка про среду, отцепленные гейты обеих приёмок,
переставшая бросать обёртка, npm-скрипт без проверки и возврат гейта в
каждый из двух фингерпринтуемых файлов. Проверка подключения сначала
смотрела только на импорт модуля и молча проходила, когда отказ
заменяли на `void` — теперь она проверяет вызов бросающей обёртки.
Гейты: npm test 1918 tests, 1917 pass, 0 fail; typecheck зелёный;
check-docs зелёный (индекс скриншотов не задет); pytest без HA 378
passed, 3 skipped.
Issue: #455
User-Visible: no
Прежде полный трек был бесплатен, а выбор лёгкого требовал обоснования. Цена —
2.9 ревью-документа на задачу и до шести на одну issue (#329, #316, #290), при
том что Medium-находки всё равно чинятся в той же задаче без отдельного цикла.
Порог не изменился: критерии §5 те же и обязательны все одновременно. Изменилась
сторона доказательства — в S2-analysis называется критерий, который задача НЕ
проходит, если идёт полным треком. «Обычный трек» без названного критерия
обоснованием не является.
Правка идёт и в AGENTS.md: там трек описан как «shortcut для мелкой работы», а
это ровно та формулировка, из-за которой полный трек остаётся умолчанием на
практике. AGENTS.md стоит вторым в порядке доверия, поэтому без него правка
канона поведение не меняет.
Бюджет четырёх циклов, арбитраж владельца, обязательность ТЗ на полном треке и
правило «ревью до мержа» не тронуты.
Issue: #338
User-Visible: no
The contract named benchmark and golden tooling, which is exactly why the smoke
launcher was allowed to skip the check for so long. It now covers every browser
check, names where each one gets it, and records why a stale bundle is worse
than a plain failure: part of the assertions go red and part stay green.
Issue: #236
User-Visible: no
Filing and servicing a separate issue costs far more than fixing a small
problem in place — the owner's call of 2026-08-19 (#202). A Medium finding
inside the task's scope no longer becomes its own issue: with no High
findings the verdict is yellow, the author fixes it and the fix passes
another review cycle. Only an out-of-scope Medium is still filed
separately, because foreign scope is never patched from a task branch.
Applied to the canon (PROCESS.md), the reviewer prompt in process.yml and
AGENTS.md; the verdict format now writes "Medium: N -> in-task | #NN".
Issue: #202
User-Visible: no
The canonical job list predated the `docs` job (added 2026-08-16) and
omitted `process-gate`. `docs` is a real blocking gate — its fingerprint
check went red right after the #113 merge and cost an extra review cycle
of confusion. The list now matches validate.yml and names `changes` as a
service path-filter rather than a gate.
Issue: #191
User-Visible: no
The owner's machine now carries Playwright with Chromium on Windows and a full
WSL environment — verified by execution: 34/34 smoke assertions, and 242 backend
tests passed where native Windows silently skips every test_ha_* file. A red
smoke that reaches the review costs a cycle of forty-five minutes plus the
return trip; run locally it costs a minute, and #89 already paid that price
once.
WSL runs of the full harness and golden verify are advisory. The canon does not
move: the beta gate is CI at the exact SHA, and baselines are accepted only via
golden:accept --reviewed on a complete Linux CI artefact.
Issue: #151
User-Visible: no
Two agents sharing one checkout share one HEAD, and twice in an hour a commit
landed on someone else's task branch that way. The layout that ends it: the main
clone belongs to the author and its task branches, hp-dev is the owner's
permanent worktree on dev, and the reviewer and the infrastructure agent own no
local tree at all — one runs in CI on a fresh checkout, the other reads through
git show and publishes through the API, so it has no HEAD to collide with.
Also recorded: a worktree is only usable on the machine that created it, because
its .git file stores an absolute path in that machine's format. We hit this in
both directions within a day.
Issue: #115
User-Visible: no
The owner stopped using GitHub Projects. Most of this is wording, but one part
was not: release-prerelease.mjs talked to the Project in code. finishIssues
looked up the project id, listed its items and its Status=Done option, and threw
when an issue was missing from the board — so the first release that closed an
issue would have died on a step with nothing to do with publishing. Found by
reading rather than by releasing, which was luck.
Closing issues stays, and now strips the status label first. That order is not
cosmetic: the invariant that a closed issue carries no status label has broken
twice already, both times because a manual step did it the other way round. The
close-merged job already does it in this order.
The documents now say labels and only labels. The explicit "no longer used"
lines are kept on purpose, in PROCESS.md and next to the code that used to sync:
a decision that vanishes quietly gets reintroduced a month later by someone who
never knew it was made.
Issue: #139
User-Visible: no
The owner's report: the process works but every stage takes a long time even on
simple bugs. Two causes, and neither was the one that first comes to mind.
The reviewer ran everything regardless. On #89 it installed Chromium, ran all 127
smoke files and a full golden capture — right for a task rated 10/10 for
complexity, absurd for a bug about a room divider. Full suites are the pre-beta
gate; the review now runs typecheck, unit and build always, and smokes, golden,
pytest or performance only where the diff and the AC call for them. The price of
narrowing it is honesty: the reviewer must list which gates it ran, which it did
not, and why, so a skipped gate is a visible decision rather than a silent one.
The reviewer also built its own environment out of model turns, with no npm cache
and no browser cache, paid for from the same forty-five minutes. The workflow now
installs dependencies and Chromium as ordinary cached steps, after switching to
the task branch so the lockfile is the branch's own.
Second, ceremony did not scale down. The light track makes a spec cheap; the new
trivial track does without one — S2-analysis straight to S5-ready, no spec review,
AC in the issue body. It is deliberately hard to qualify for: a bug on one surface,
no new UX contract, no migration, no i18n, no perf or touch effect, three checkable
AC at most, and expected behaviour already on record. Nothing left to decide is the
criterion that holds the whole thing up, and it cannot be met by feeling sure.
Code review is never skipped on either track. It is what stands in for testing
here, so it is the one stage speed may not buy.
Issue: #127
Issue: #128
User-Visible: no
The guard refused to review any issue the owner had not filed himself. The rule
was meant to keep malformed outside reports out of the pipeline, but it checked at
every step instead of at the entrance, and it duplicated a guarantee the platform
already gives: only someone with write access can apply a label. Applying the
first status label is the owner's explicit decision, and it is the only place the
question belongs.
So the author check is gone. While an issue carries no status label it sits
outside the process and the invariants do not apply; once labelled, the task is in
flight and who filed it stops mattering.
The old rule also cost real work. On #123 an outside bug report had been analysed
and specified before the guard turned it away in nine seconds, and the remedy on
offer was to refile the same thing as the owner's own issue.
Issue: #114
User-Visible: no
The implementation loop runs typecheck, unit and build. Golden, browser smokes,
performance and the full HA harness run before a beta — after the code review has
passed and the issue already sits in S8-merged. Some defects cannot surface any
earlier, and until now the process had nothing to say about them, so the honest
reading was a second full review cycle at the most expensive possible moment.
The owner's decision: fix it, re-run what failed, and a green run carries the
release on. The gate named the defect precisely and the same gate proves the fix,
so the check is objective and depends on nobody's judgement.
The boundary is written down with it, because "the gate found something" could
otherwise absorb an arbitrary amount of new work. A fix that changes a behaviour
contract, reaches an untouched subsystem or rivals the task in size goes through
the normal flow. Editing a test so it stops failing is concealment rather than
repair — the exception is a defect proven to be in the fixture, as on #89.
The rule also records what it costs: the author judges his own work here, which
the process refuses everywhere else. That is the price of speed at the one point
where a review cycle is dearest, and the compensation is that the re-run command
and its result are written into the issue where the release manager reads them.
Issue: #114
User-Visible: no
A review run always moves the label. The rule is written down because its absence
cost a real stall: a green code review whose merge conflicted left the label alone,
the waiting author polled thirty times and reported the limit as exhausted, and a
verdict that existed reached nobody.
Both documents now say what S6-in-progress means when the verdict was green and
only the merge failed — rebase, not rework, and the verdict still stands. They also
say that a label which did not change means the run failed rather than the work, so
the answer is logs and the owner, not more polling. Cycles are counted per stage.
Issue: #114
User-Visible: no
docs/specs/README.md kept a "Статус ТЗ" column with its own vocabulary — draft,
in implementation, done — next to the labels that already hold the status. Two
dictionaries for one fact drift apart, and these had: the column still called
issues "in implementation" that were closed weeks ago. The table now says only
which issue a spec belongs to.
AGENTS.md was telling agents that PROCESS.md §1 does not cover package.json and
the rest of the configuration, and to report it as missing. It covers them now.
The same paragraph gained the rule that D beats A where paths overlap, which is
what keeps the built bundle under custom_components/houseplan/frontend/ from
reading as product source.
Issue: #119
User-Visible: no
Practice had already diverged from the documents: #105, #112, #114 and #116 were
all done without a spec and without review, and that was right. Nothing said it
was allowed.
The test is mechanical — not a single class A file — rather than left to the
executor's judgement, because a loose reading is exactly how product changes
would learn to skip review.
Issue: #118
User-Visible: no
Review fires from the label and runs on its own, but nothing was picking the
result up: the author reported "handed over for review" and stopped, so the
conveyor stalled until the owner said a sentence. The author now polls the label
and continues from whatever it became.
Also drops the merge-into-dev standing permission: the pipeline does the merge
before setting S8-merged, so a hand merge would race it.
Issue: #114
User-Visible: no
Agents were escalating technical calls. The owner answers what a person sees
or does and how much user-visible change belongs in an issue; storage, module
layout, naming, test strategy and migration mechanics are settled by the
agents, recorded as assumptions and challenged in review.
Issue: #114
User-Visible: no
Without it S8-merged would be a lie: the label asserts the code is in dev,
while the branch-only push leaves it on the task branch. What lands is what
the reviewer just accepted, and dev is allowed to carry unreviewed code
anyway, so the merge adds no risk the branch did not already carry.
Issue: #114
User-Visible: no
The reviewer runs in CI and can only read the remote, so an unpushed spec or
commit either stalls the review or points it at the wrong tree. Pushing
issue/<NN>-slug now needs no command; dev, main, merges, tags and releases
still do.
Issue: #114
User-Visible: no
Publishes the 538-line process canon into the repository, replacing the
51-line provenance stub that pointed at a non-existent .agents/PROTOCOL.md.
Rewrites AGENTS.md: product context first, labels as the canonical status,
rule #1 with the status check, change classes, trailers, push cadence,
Codex/Claude roles and review cycle limits.
Issue: #112
User-Visible: no