mirror of
https://github.com/Matysh/houseplan-card
synced 2026-10-06 14:39:22 +00:00
104 lines
4.6 KiB
Markdown
104 lines
4.6 KiB
Markdown
# Large-house performance gate
|
|
|
|
`benchmark_large_house.mjs` exercises the deterministic `large-house-v1`
|
|
fixture. The fixture has 60 rooms, 200 devices, 100 openings, 60 partitions,
|
|
40 columns and 500 decor objects on three floors.
|
|
|
|
The runner records seven measured samples after one discarded warm-up. With
|
|
this intentionally small CI sample, the nearest-rank `p95` is the observed
|
|
maximum; reports keep the conventional field name but should be read as a
|
|
high-tail guard rather than a population estimate:
|
|
|
|
- model readiness and first stable render;
|
|
- space switch, HA state update, pan/zoom and opening the settings dialog;
|
|
- a shared-wall room-resize preview which is cancelled before persistence;
|
|
- a twelve-switch navigation cycle;
|
|
- Long Tasks for every measured window;
|
|
- heap growth after four additional navigation rounds with forced GC;
|
|
- hot-cache size and growth after the same warmed cycles.
|
|
|
|
Every report is tied to the source fingerprint embedded by Rollup. A stale
|
|
bundle is a hard failure.
|
|
|
|
## CI contract
|
|
|
|
The `performance` job checks out the candidate and its base SHA, builds both,
|
|
and runs them sequentially with the same Node.js 22 process family, pinned
|
|
Playwright Chromium and hosted runner. `compare.mjs` then applies two limits:
|
|
|
|
1. a relative regression allowance against the base-SHA report;
|
|
2. an absolute safety ceiling from `budgets.json`.
|
|
|
|
The tighter limit wins. The absolute values are catastrophic safety ceilings,
|
|
not normal-performance targets; the base-relative comparison catches smaller
|
|
regressions. Small fast operations receive an absolute noise
|
|
allowance so normal scheduler jitter does not become a false regression. Heap,
|
|
Long Tasks, warmed-cache growth and the expected rendered-device count are
|
|
gated separately. Long-Task maximum/count/total checks use the same
|
|
relative-plus-absolute policy as timings. Both raw reports and the comparison are always uploaded as
|
|
the `large-house-performance` artifact, and the table is written to the GitHub
|
|
job summary.
|
|
|
|
This base-vs-candidate design intentionally does not compare timings captured
|
|
on different machines or different Chromium builds. A runtime/profile mismatch
|
|
fails closed.
|
|
|
|
## Local diagnostics
|
|
|
|
Build and copy a fresh demo bundle first, then run:
|
|
|
|
```bash
|
|
npm run benchmark:large-house -- --samples=7 --warmups=1 --output=artifacts/performance/local.json
|
|
```
|
|
|
|
A local report is diagnostic only; it cannot replace the CI comparison.
|
|
|
|
To reproduce the comparison against another checkout using one harness and one
|
|
browser installation:
|
|
|
|
```bash
|
|
npm run benchmark:large-house -- --target-root=../base --samples=7 --output=artifacts/performance/baseline.json
|
|
npm run benchmark:large-house -- --target-root=. --samples=7 --output=artifacts/performance/candidate.json
|
|
npm run benchmark:compare
|
|
```
|
|
|
|
## Changing budgets
|
|
|
|
Budget changes require an explicit review of recent CI artifacts and a written
|
|
rationale in the change. Do not loosen a threshold merely to make a single red
|
|
run pass. A new fixture profile gets a new profile id instead of silently
|
|
changing the meaning of `large-house-v1`.
|
|
|
|
The `cleanFloor` entry ceiling is 160: the reviewed fixture currently warms
|
|
120 deterministic room/physical-body entries, and the extra 40 slots allow a
|
|
legitimate fixture extension without weakening the separate zero-growth gate.
|
|
|
|
## Glow profiles
|
|
|
|
Both Glow profiles run deterministic 1/10/30/60-pool variants at DPR 1 and
|
|
Chromium CPU throttling x4, but deliberately exercise different fixtures:
|
|
|
|
- `large-light-blend-v1` compares the isolated screen group with the previous
|
|
normal-layer implementation on the shared frontend/backend schema fixture
|
|
`test/fixtures/glow/additive-pools.json`;
|
|
- `large-house-glow-overlay-v1` measures simultaneous temperature fill and
|
|
independent Glow on the existing 60-room/200-device large-house fixture,
|
|
without changing `large-house-v1`.
|
|
|
|
```bash
|
|
npm run benchmark:glow -- --profile=large-light-blend-v1 --output=artifacts/performance/glow.json
|
|
npm run benchmark:glow -- --profile=large-house-glow-overlay-v1 --output=artifacts/performance/overlay.json
|
|
```
|
|
|
|
Reports include per-variant state-update timings, render/pool counts, Long
|
|
Tasks, screenshot time, heap and cache growth. The first CI comparison against
|
|
a base SHA that predates `glow_enabled` bootstraps only the overlay profile's
|
|
relative baseline from the candidate; its absolute ceilings still gate that
|
|
introduction. Every subsequent revision compares both profiles to the real
|
|
base SHA.
|
|
|
|
The initial absolute ceilings are intentionally conservative bootstrap limits;
|
|
they must be reviewed against the first paired Ubuntu artifacts before the
|
|
feature is promoted from beta. Same-runner relative checks remain the primary
|
|
candidate-only regression signal.
|