Files
houseplan-card/demo/performance/README.md
T

3.7 KiB
Raw Blame History

Large-house performance gate

benchmark_large_house.mjs exercises the deterministic large-house-v1 fixture. The fixture has 60 rooms, 200 devices, 100 openings, 60 partitions, 40 columns and 500 decor objects on three floors.

The runner records seven measured samples after one discarded warm-up. With this intentionally small CI sample, the nearest-rank p95 is the observed maximum; reports keep the conventional field name but should be read as a high-tail guard rather than a population estimate:

  • model readiness and first stable render;
  • space switch, HA state update, pan/zoom and opening the settings dialog;
  • a shared-wall room-resize preview which is cancelled before persistence;
  • a twelve-switch navigation cycle;
  • Long Tasks for every measured window;
  • heap growth after four additional navigation rounds with forced GC;
  • hot-cache size and growth after the same warmed cycles.

Every report is tied to the source fingerprint embedded by Rollup. A stale bundle is a hard failure.

CI contract

The performance job checks out the candidate and its base SHA, builds both, and runs them sequentially with the same Node.js 22 process family, pinned Playwright Chromium and hosted runner. compare.mjs then applies two limits:

  1. a relative regression allowance against the base-SHA report;
  2. an absolute safety ceiling from budgets.json.

The tighter limit wins. The absolute values are catastrophic safety ceilings, not normal-performance targets; the base-relative comparison catches smaller regressions. Small fast operations receive an absolute noise allowance so normal scheduler jitter does not become a false regression. Heap, Long Tasks, warmed-cache growth and the expected rendered-device count are gated separately. Long-Task maximum/count/total checks use the same relative-plus-absolute policy as timings. Both raw reports and the comparison are always uploaded as the large-house-performance artifact, and the table is written to the GitHub job summary.

This base-vs-candidate design intentionally does not compare timings captured on different machines or different Chromium builds. A runtime/profile mismatch fails closed.

Local diagnostics

Build and copy a fresh demo bundle first, then run:

npm run benchmark:large-house -- --samples=7 --warmups=1 --output=artifacts/performance/local.json

A local report is diagnostic only; it cannot replace the CI comparison.

To reproduce the comparison against another checkout using one harness and one browser installation:

npm run benchmark:large-house -- --target-root=../base --samples=7 --output=artifacts/performance/baseline.json
npm run benchmark:large-house -- --target-root=. --samples=7 --output=artifacts/performance/candidate.json
npm run benchmark:compare

Changing budgets

Budget changes require an explicit review of recent CI artifacts and a written rationale in the change. Do not loosen a threshold merely to make a single red run pass. A new fixture profile gets a new profile id instead of silently changing the meaning of large-house-v1.

The cleanFloor entry ceiling is 160: the reviewed fixture currently warms 120 deterministic room/physical-body entries, and the extra 40 slots allow a legitimate fixture extension without weakening the separate zero-growth gate.

The absolute switch-cycle/Long-Task ceilings include roughly 20–30% headroom over the paired 2026-08-09 Ubuntu run where the unchanged base and candidate both reached about 5.3 s / 2.45 s / 22 tasks / 9.9 s total under runner load. The same-runner relative checks remain tighter for an actual candidate-only regression; this prevents an overloaded but symmetric runner from turning an absolute safety ceiling into a flaky code-regression signal.