No Full Performance profile walked the backdrop (imagePlan) path: every
large-house fixture has plan_url null. So the #739 K1 double render -- a
warm 2.5D floor switch with a backdrop cleared the ready paper, inserted
the veil, probed the card background and rendered a second time -- was
invisible to CI by time and structurally; a temporary probe found it.
large-house-isometric-backdrop-v1 is the twin of large-house-isometric-v1
with the shipped f1.svg under a URL of its own on every floor
(plan_aspect 1, room geometry unchanged). The variant is derived in the
runner as plan-snap's is, so demo/fixtures and the bundle fingerprint do
not change. A first stable frame without the backdrop image fails the
sample. After the switchCycle window and its #735 guard, before forced
GC and outside every timed window, a probe makes six warm switches and
counts performUpdate passes until updateComplete resolves true: the K1
second pass starts after the first updateComplete resolves, so a count
taken right after the first await reads one on both sides.
evaluate.mjs rejects a candidate of this profile unless every switch took
one pass, or when the probe is missing; the base is reported, not judged
(v1.78.0 and dev before #739 take two). The budget is a copy of the
historical isometric budget under the new profile id. performance.yml
gains the isometric-backdrop matrix entry with exact-SHA comparison.
Witness: a tree with #739 reverted reads perSwitch 2 in every sample and
benchmark:compare against the branch report throws on the pass count;
v1.78.0 reads 2, the branch reads 1.
Issue: #743
User-Visible: no
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018qZfe7YS4rqEMKoVeS3GKd
Since #735 switchCycleMs times the warmed twelve-switch cycle, and the
absolute ceilings of 7000 ms (flat) and 8000 ms (2.5D) sat 7.6-10.2
times above the level. performance_smoke judges only these ceilings, so
between full runs the warm floor switch that #694/#725 just sped up was
guarded only against a several-fold collapse.
Series: every Full Performance run after #735, both sides (the base is
measured by the candidate runner, so it is warm too), 7 samples each -
36821241343 (#735, base 76558bf2), 36838891001 (#740), 36838952536
(#742) and 36839009721 (#739), the last three against dev 7ff2b5ae.
flat large-house-v1 / plan-snap / interaction 666.9-812.7 ms
2.5D large-house-isometric / stage3-dense 766.5-1333.9 ms
The 2.5D maximum is the dense pair of 36838952536, whose base on the
same runner read 1249.7 ms against 849.9-982.4 ms elsewhere: runner
noise the series is meant to contain. No 3-sample performance_smoke
median is in the series yet; those profiles join Validate only on a
src/** diff.
Rule (as #692, #675 falls in the same band): one number per family, the
first multiple of 50 ms at or above 1.15 x M and no higher than 1.2 x M,
M being the family's maximum median: flat 1.15 x 812.7 = 934.6 -> 950
(+16.9 %), 2.5D 1.15 x 1333.9 = 1534.0 -> 1550 (+16.2 %). One number
per family keeps the smoke = full (#473 AC4), plan-snap/interaction
"every original ceiling" and dense = historical (#160) contracts; the
price is wider headroom for the faster profiles. The base-relative
comparison of the full workflow (0.35 / 0.2, 250 ms) is unchanged and
stays the detector for smaller growth.
The new test pins both families: one ceiling in every file of a family,
every point of the series passes the smoke budget with --absolute-only,
the ceiling follows the rule and stays inside [1.15, 1.2] x M, doubling
the level fails, and the full profiles' ratio and noise allowance are
unchanged. 7000 left in any flat file or a ceiling under 1.15 x M reds
it. The README records the series and the reasoning.
Issue: #747
User-Visible: no
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018qZfe7YS4rqEMKoVeS3GKd
Every large-house sample mounts a new card that has visited only floors 1
and 2 before the twelve-switch cycle, so the cycle's second step was always
the first visit to floor 3: 20 new clean-floor entries, and in 2.5D one Iso
geometry entry plus one structural build. In large-house-interaction-v1 the
editor series also moves the config epoch that keys the clean-floor cache,
so floor 1 was cold as well. That one cold step was about half of
switchCycleMs, which the README and the cycle comment describe as warmed
navigation, and a 35% warm regression drowned in it.
The runner now visits every fixture floor once in cycle order after the
settings dialog closes, outside every timed and Long Task window, and
returns to floor 2, so the cycle still starts with 2 -> 1. A guard snapshots
the hot caches and the 2.5D structural build counter around the window and
fails the sample when anything grew. Caches an older base lacks read as 0
and its null counter is not judged, so a v1.78.0 base still passes.
Budgets, hardMaxMs, metric names, the report schema, profiles and the
workflow are unchanged. Base and candidate are both measured by the
candidate runner, so the comparison is unaffected; the absolute
switchCycleMs level steps down, which the README now explains. A unit
anchor pins the warm-up position and the guard message.
Issue: #735
User-Visible: no
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018qZfe7YS4rqEMKoVeS3GKd
The isometric-stage3-dense-v1 runner still demanded at least one bounded
#651 nudge. Since #713 every raised device tile and lock badge is lifted by
the one shared wall-top rise and carries data-hp-iso-nudged="false", so the
Full Performance profile failed its input contract before any timing.
The contract is inverted: a single nudged raised root now fails the sample,
matching the golden requireOneRise preflight. The performance README states
the current contract, and the #570 runner-contract unit pins the new failure
text.
Issue: #719
User-Visible: no
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018qZfe7YS4rqEMKoVeS3GKd
Owner decision of 2026-09-30 in #694. Switching Flat <-> 2.5D is a one-off
General settings change, and since #649 the runner measures a full config
reload for the candidate against a per-device projection flip for v1.77.0,
which alone explains most of 73.8 -> 195.7 ms. The scene build stays gated by
modelReady, firstStableRender and spaceSwitch, a UI freeze by the single
long-task ceiling; the runner still reports viewToggleMs.
Issue: #720
User-Visible: no
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018qZfe7YS4rqEMKoVeS3GKd
The full workflow's 7-sample median of firstStableRenderMs for
large-house-interaction-v1 was about 2790 ms across the 1.77 line and
about 2920 ms at v1.78.0 (7d4d75bd: 2925.0; the neighbouring run of the
same SHA read 3144.8). The 3000 ms ceiling sat 2.7 % above the level and
failed on runner noise; #689's 3-sample smoke read 3002.4.
- hardMaxMs 3000 → 3400 in the full profile and its smoke twin (one
number, #473 AC4): +16 % over the 1.78 level, +8 % over the worst run.
The base-relative ratio and noise allowance are unchanged.
- README: the series and the reasoning; the test pins the number and the
series points.
Issue: #692
User-Visible: no
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018qZfe7YS4rqEMKoVeS3GKd
spaceSwitchMs изометрии — одно холодное переключение этажа со сборкой
2.5D-геометрии; он растёт с каждой стадией 2.5D, и потолок 1800 мс из #89
stage 1 стал самим уровнем. Медиана perf-smoke на hosted-раннере:
1500 мс (09-09…09-12, 12 прогонов) → 1668 (09-19…09-25, 6) → 1798 после
#649 (9, σ 42, максимум 1867,7). Попытки одного SHA расходятся на 0,4 %,
разные прогоны одного SHA — до 5,5 %: больше образцов вердикт не меняют,
рычаг — потолок.
2200 мс: +17,8 % над наблюдённым максимумом и ниже единственной настоящей
регрессии окна (2719,6 мс, #583 до решётки, 09-16); удвоение уровня
краснеет с запасом. Одно число в трёх файлах: смок = полный профиль
(#473 AC4), плотный двойник = исторический (#160). Коэффициент и допуск
полного сравнения не тронуты. Обоснование — demo/performance/README.md,
тест закрепляет число и три точки ряда.
Issue: #675
User-Visible: no
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018qZfe7YS4rqEMKoVeS3GKd
Относительная половина «Полных бенчмарков» сравнивала кандидата с прошлой
вершиной `main`. Для стабильного релиза это давало круг, в котором гейт не
может покраснеть дважды: прогон идёт только на push в `main`, кандидат обязан
там оказаться, и следующий коммит той же линейки берёт базой первый — то есть
линейку саму. На выпуске v1.76.0 это видно построчно: прогон 35097102695 на
`c3d64789` честно показал resizePreview 603 → 981 и panZoom 91 → 205 в скрытой
изометрии, а прогон на `9683a590` был зелёным и был бы зелёным без всякой
правки бюджетов.
Теперь база выбирается по намерению коммита: head несёт трейлер `Release:` без
пре-релизного суффикса — сравниваем с предыдущим стабильным тегом. Бета,
обычный push и ручной `comparison_ref` не меняются.
Решение вынесено из shell в `scripts/performance-baseline.mjs` по тому же
доводу, что и разбор вердикта ревью (#556): отрицательные случаи — тега нет,
тег стоит на самой голове, база перестала быть предком, база старше
HP-PERF-01 — в YAML не прогнать ни одним тестом. Обращения к git инжектируются,
фикстуры описывают дерево. Отказы по-прежнему уводят в сторону БОЛЬШЕГО
сравнения: непригодная база → родитель → последний достижимый релизный тег.
Проверено исполнением на этом репозитории: стабильный кандидат v1.76.0 →
`2c6410bb` (v1.75.0); бета v1.76.0-beta.5 и обычный push → `push before`;
dispatch с `comparison_ref=v1.74.0` → `e63460f0`.
Свидетели: `test/performance-baseline.test.mjs` (10 проверок, включая AC2 —
второй коммит линейки не сравнивается сам с собой) и мутант
`stable-candidate-compares-against-itself`, который возвращает прежнее
поведение и обязан краснеть; проверено подменой руками — AC2 падает, оригинал
проходит.
AC4: других релизных гейтов, судящих о родителя, нет. `validate.yml` берёт
`github.event.before` только для ДИАПАЗОНА файлов, и там база уже заменена
доказанно зелёным предком (#387/#388), а не сырым родителем.
npm test 2731/2730/0 fail, typecheck чистый, check-docs зелёный (кроме
известного отпечатка скриншотов, #586).
Issue: #587
User-Visible: no
Полные бенчмарки кандидата v1.76.0 (прогон 35097102695, SHA c3d64789) против
v1.75.0 покраснели на двух скрытых изометрических профилях:
| метрика | v1.75.0 | v1.76.0 | предел |
|---|---|---|---|
| large-house-isometric resizePreview | 603 | 981.4 | 753 |
| large-house-isometric panZoom | 90.8 | 205.2 | 150.8 |
| stage3-dense stateUpdate | 77.5 | 169.1 | 152.5 |
| stage3-dense resizePreview | 582.5 | 765.6 | 732.5 |
| stage3-dense panZoom | 90.6 | 171.7 | 150.6 |
Шаг настоящий и объяснённый: #583 добавил скрытому 2.5D-виду геометрии, а
поиск свободного места для подписей даже после ускорения решёткой стоит вдвое
дороже, чем до #583. Абсолютные потолки не тронуты и держатся с запасом
(panZoom 205 при 600, resize 981 при 2200); все семь пользовательских профилей
зелёные — регрессия целиком внутри вида за `hp_alpha`.
Решение владельца 2026-09-16: принять. Рычаг — допуск в миллисекундах, а не
коэффициент: он покрывает разовый сдвиг уровня и продолжает ловить рост от
нового уровня, тогда как поднятый коэффициент разрешил бы удвоение навсегда.
Значения одинаковы у обоих профилей — контракт #160 требует, чтобы плотный
двойник Stage 3 делил с историческим профилем каждый общий потолок, и тест
`performance-workflow` это стережёт.
Допуски временные. #585 переписывает поиск на перебор границ препятствий
вместо скана диска 48 px; когда он приедет, значения возвращаются к 150/60/75.
Тест `performance-budget` фиксирует и числа, и то, что наблюдённый шаг проходит,
а следующий такой же — уже нет.
Честно о границе метода: сравнение идёт с ПРЕДЫДУЩЕЙ вершиной main, а она уже
несёт эту же линейку, поэтому следующий прогон был бы зелёным и без правки
бюджетов. Правка сделана не ради зелёного прогона, а чтобы принятый уровень был
записан явно и проверялся тестом. Отдельно завожу, что базой стабильного
кандидата должен быть предыдущий стабильный тег, а не вершина main.
npm test 2721/2720/0 fail.
Issue: #585
User-Visible: no
Release: v1.76.0
The gate used to require every Validate run on the tag SHA to be green:
a cancelled duplicate or a red flake that a later re-run had fixed kept
the stable release blocked (v1.73.0, 09.09 — released by hand). Now the
verdict comes from the newest run that was not cancelled: not completed →
wait, success → pass, anything else → fail, no run → wait. The same rule
is documented for the perf workflow and the release runbook.
Mutants: release-gate-counts-cancelled-runs, release-gate-oldest-run-wins.
Issue: #511
User-Visible: no
The #160 contract test keeps the dense profile's longTasks block equal to
the historical isometric one, and both profiles boot through the same lazy
iso-scene-render chunk; the accepted split applies to both. Validate
34356856702 caught the divergence.
Issue: #507
User-Visible: no
Release: v1.73.0
Full Performance of the v1.73.0 stable candidate against the v1.72.0 product
(run 34354409872) was red on one check of the isometric profile:
longTask.countP95 16 → 20 against max(16×1.2, 16+3) = 19.2, with every
timing, longTask.totalP95Ms and longTask.maxSingleMs green. The trace behind
#506 shows why: since v1.73.0-beta.1 the isometric renderer is the lazy
iso-scene-render chunk (#160 Stage 3), so the single v1.72.0 boot task is
split in two around that import — the same work, +2 tasks.
Owner decision 2026-09-09: accept the split. countNoiseAllowance 3 → 5 for
large-house-isometric-v1 only; the ratio, the hard ceiling, total and
maximum single task keep gating real growth. The downloaded CI artefact
re-evaluated with this budget passes (limit 21, actual 20, no failures).
Documented in demo/performance/README.md; the budget test pins the
allowance and the untouched profiles.
Issue: #507
User-Visible: no
Release: v1.73.0
Классификация changes вынесена в scripts/classify-changes.mjs (выходы
perf_iso/perf_interaction, fallback --all). performance_smoke добавляет
large-house-isometric-v1 при правке src/iso-* и large-house-interaction-v1
при правке живого пути, по 3 образца против hardMaxMs полных профилей
(budgets-*-smoke.json). Набор профилей входит в ключ reuse. PROCESS.md §8.
Issue: #473
User-Visible: no