* perf(canvas): add traced navigation benchmarks - Record and replay timestamped pan and zoom gestures through DOM and CDP input paths\n- Correlate input, viewport, render, long-task, and retained-backing events in Chromium traces\n- Report frame pacing, latency, jump, anchor drift, and crisp-settlement metrics * perf(canvas): stabilize navigation comparisons - Separate low-overhead metric runs from optional CPU-profile traces\n- Warm scenarios before recording and use a consistent SwiftShader browser configuration\n- Add a canonical momentum-pan reversal gesture alongside pinch reversal * fix(canvas): require hardware GPU navigation benchmarks - Run macOS performance captures through Metal-backed ANGLE and reject accidental SwiftShader fallback\n- Record the GL renderer and reserve software GPU mode for portable correctness smoke runs * perf(canvas): cache shadow rasters for crisp backing - Rasterize local drop and inner shadows only while constructing retained scene backing\n- Bound native image memory and invalidate cached entries with node and renderer lifecycle changes\n- Quantize zoom-aware raster resolution and reuse nearby scales without lowering normal scene quality * test(canvas): verify retained shadow raster fidelity - Compare settled retained-backing shadow output with direct CanvasKit rendering\n- Keep backdrop blur on the picture fallback and exercise graph-driven cache invalidation\n- Cover updates, deletion, and reparenting through actual SceneGraph events * perf(canvas): benchmark real FIG fixtures - Serve exact local fixture bytes through an isolated Playwright route for production preview runs\n- Wait for document loading and page population before zooming to fit and recording navigation\n- Record the resolved fixture path in benchmark environment artifacts * fix(canvas): preserve nested effect subtree pictures - Keep deeply nested shadow documents on one retained subtree picture instead of exploding them into per-node image draws\n- Restrict shadow raster acceleration to effect-bearing page children\n- Cover nested shadow fallback and restore gold-preview FIG pinch performance to master levels * refactor(canvas): share recorded wheel sample type * perf(canvas): defer backing settlement across zoom reversals - Track explicit navigation phases and gesture generations instead of inferring idle from viewport timing\n- Cancel or defer retained backing construction while pan, zoom, momentum, or tentative settlement is active\n- Add a repeated short-pause pinch reversal fixture based on the user trace * perf(canvas): index bounded render chunks - Split oversized painter subtrees into self-paint and bounded descendant chunks without dropping container visuals\n- Bulk-load chunk visual bounds into RBush for selective world-space queries\n- Cover bounded updates and gold-preview build/query complexity before tile rendering consumes the index * refactor(canvas): namespace render chunk coverage * perf(canvas): model chunk paint context - Preserve ancestor transform and clip dependencies for independently renderable chunks\n- Keep opacity, blend, blur, and mask isolation subtrees atomic until command-level splitting exists\n- Report oversized atomic chunks and lock gold-preview to bounded painter units * perf(canvas): record pixel-correct render chunks - Record interruptible chunks in world coordinates with ancestor transforms, clips, and chunk-local culling bounds\n- Draw opacity, blend, blur, and mask isolation chunks directly into destination surfaces in painter order\n- Compare composited chunk output with direct CanvasKit rendering instead of relaxing visual thresholds * perf(canvas): render selective world tiles - Map world regions to fixed 256-device-pixel tile targets and quantized sharpness levels\n- Query only intersecting render chunks and preserve atomic destination compositing\n- Match multi-tile CanvasKit output against direct rendering and measure gold-preview tile cost * perf(canvas): cache chunk pictures across tiles - Reuse world-space chunk command pictures for every intersecting tile\n- Pool 256-pixel tile surfaces and expose allocation, draw, flush, and snapshot timings\n- Keep expensive atomic foreground blur visible as an over-budget scheduler constraint * perf(canvas): schedule cached tile rendering - Bound tile images with an LRU cache and reuse pooled CanvasKit surfaces\n- Plan mandatory holes, stale visible refreshes, and overscan by navigation and content generation\n- Stop jobs at a strict deadline while reporting fallbacks, stale work, overruns, and over-budget effects * perf(canvas): integrate progressive tiled rendering - Keep retained scene output as the interaction fallback while exact tiles refine only after navigation becomes idle - Centralize runtime URL flags and pass renderer selection through the typed Vue canvas API - Replace benchmark sleeps with explicit mode-aware renderer settlement and report exact tiled coverage - Preserve bounded scheduler metrics, generation cancellation, native resource cleanup, and shared visual-bounds logic * refactor(app): centralize runtime query configuration - Parse collaboration, recent-files, benchmark, presentation, and renderer flags in one typed app module - Remove ad hoc URL parsing from workspace and collaboration runtime consumers - Cover supported values and production-safe defaults without adding a repository lint rule * fix(canvas): replace fallback pixels with exact tiles - Render opaque page-background tile cells and install them with source replacement instead of double-compositing translucent scene content - Exercise the live progressive controller against direct rendering across masks, effects, blend isolation, images, fallback text, transforms, and clipping - Preserve the bounded reversal path with zero Long Tasks and exact settlement near 128 ms on gold-preview.fig * perf(canvas): invalidate tiled content selectively - Index chunk dependencies across contained nodes and transform or clipping ancestors - Re-record affected chunk pictures and invalidate tiles intersecting old or new visual bounds - Advance unaffected cached tiles to the new scene generation instead of rebuilding the full chunk index and tile cache - Keep structural graph mutations on the safe full-rebuild path and cover selective refresh end to end * perf(canvas): bound atomic blur tile refresh - Render atomic blur chunks with tile-local isolation bounds and blur halos instead of replaying full-subtree layers - Keep content refresh behind the retained fallback, cap GPU submissions to four tile jobs per frame, and adapt estimates from measured work - Preserve large-radius CPU over-budget visibility while preventing Metal-backed refresh bursts and deferred GPU overload - Add deterministic node-mutation benchmarks and summarize scheduler throughput, job duration, overruns, and exact content settlement * perf(canvas): cancel obsolete tile refresh generations - Count and report queued jobs removed by content or navigation generation changes - Add deterministic mutation-then-reversal benchmark support without sleeps - Assert exact tile work remains suspended during navigation and resumes for the final viewport - Summarize cancellation alongside scheduler throughput, overruns, and settlement metrics * test(canvas): cover live tiled blur settlement - Load gold-preview.fig through the real tiled canvas surface and wait on explicit renderer settlement - Commit the settled radius-210 large-blur browser snapshot - Replay the canonical zoom reversal during refresh and require byte-identical final canvas convergence * fix(canvas): harden renderer resource lifecycle - Release tiled surfaces, images, pictures, and queued work across surface, font, graph, page, structure, and renderer lifecycle boundaries - Restore pooled canvas, viewport, and backing state through exception-safe native recording and raster paths - Rebuild tiled chunk topology only when isolation requirements actually change, preserving selective blur mutation performance - Document deterministic active-renderer settlement and add lifecycle, graph replacement, cache failure, and surface replacement regressions * test(canvas): remove source-matching renderer claims - Delete the autopsy suite that inferred runtime correctness from source text, regexes, line placement, and symbol counts - Keep renderer ordering, cache cleanup, effect behavior, and pixel fidelity covered by executable behavioral and lifecycle tests * perf(canvas): present retained backing during tiled navigation - Profile production reversal traces and attribute tiled p95 cost to GPU command-buffer flushes from full-scene fallback replay and tile presentation - Use the retained backing as the moving fallback while tile scheduling and cached lookup remain allocation-free - Defer tile image presentation until idle and expose visible versus presented tile counts in navigation telemetry - Reduce tiled reversal render p95 from about 8ms to 0.3ms while preserving exact idle replacement and visual parity * perf(canvas): prioritize visible tile settlement - Profile per-tile allocation, draw, flush, snapshot, and chunk costs through scheduler telemetry - Defer overscan until all visible exact tiles are covered - Replace the four-job idle cap with a higher safety ceiling while the measured five-millisecond deadline controls cheap work - Reduce mutation-plus-reversal exact settlement from about 272ms to 160ms without Long Tasks, overruns, or over-budget jobs * refactor(canvas): clarify renderer lifecycle boundaries - Extract retained backing state types and navigation preview timing\n- Isolate tiled scheduler telemetry from frame orchestration\n- Document settlement and CanvasKit ownership invariants\n- Preserve hot drawing loops, budgets, cache limits, and rendering decisions * fix(canvas): preserve current label rendering Retain the merged paragraph-label cache lifecycle and substituted-font readiness while reconstructing the renderer stack on current master. * test(canvas): keep tile benchmark assertions deterministic Keep performance timing in benchmark telemetry while asserting structural tile selectivity and cache behavior in CI. * feat(canvas): expose experimental tiled rendering - Persist retained or tiled canvas mode in General settings\n- Keep retained rendering as the default and apply changes after reload\n- Preserve URL overrides for deterministic benchmarks and support reproduction * refactor(app): centralize renderer preference state Expose renderer override provenance from runtime configuration and keep the settings control's derived state separate from its explicit persistence action. * refactor(app): share settings layout anatomy Reuse slot-based section headers and bordered groups while keeping each settings control row explicit. * fix(canvas): harden tiled renderer boundaries - Bound low-zoom tile planning and handle failed tile surface allocation\n- Preserve effect raster dependencies, runtime-safe clocks, and navigation timing contracts\n- Keep benchmarks deterministic, backward compatible, and accurately localized * fix(canvas): invalidate dependent node pictures Track first-child shadow dependencies for retained node pictures so child geometry updates cannot leave stale parent shadows.
8.7 KiB
Navigation performance benchmark
The navigation benchmark measures the complete wheel/trackpad-to-render pipeline. It is intentionally separate from the synchronous microbenchmark in tests/e2e/viewport/zoom-pan.spec.ts: accepting JavaScript calls quickly does not prove smooth navigation.
Renderer lifecycle and ownership
The benchmark validates the retained and tiled renderer contracts documented in Renderer lifecycle and ownership. Structural renderer changes must preserve those generation, settlement, and CanvasKit ownership rules.
What it captures
Every run writes recording.json, metrics.json, and environment.json. When tracing is enabled, it also writes trace.json.gz:
trace.json.gz(when tracing is enabled): Chromium tracing data for Perfetto or Chrome tracing tools.recording.json: wheel samples plus OpenPencil input, viewport, render, and retained-backing events.metrics.json: frame pacing, input latency, zoom-anchor drift, viewport jumps, exact active-renderer settlement, and tiled scheduler throughput/cancellation when enabled.environment.json: browser, runtime, replay mode, and source gesture information.
OpenPencil emits User Timing marks under openpencil:*, including wheel receipt/flush, viewport mutation, render start/end, backing preview/build, crisp-backing completion, and exact tiled coverage. The recording also runs a continuous requestAnimationFrame heartbeat and observes browser Long Tasks, so display stalls remain visible even when OpenPencil does not render.
Record a physical macOS trackpad gesture
-
Start a production preview (preferred) or development server:
bun run build bun run preview -- --port 1420 -
Open:
http://localhost:1420/?test&no-chrome&no-rulers&navigation-benchmark -
In DevTools, start recording:
openPencil.test.navigation.startRecording('macbook-fast-pinch-reversal') -
Perform exactly one gesture, allow the canvas to become crisp, then stop and copy the result:
copy(JSON.stringify(openPencil.test.navigation.stopRecording(), null, 2)) -
Save the result under
tests/fixtures/navigation/gestures/. Do not edit delta values or timestamps. Remove document names if they contain private information.
Record at least slow pan, momentum pan, direction reversal, slow pinch, fast pinch in/out, pinch reversal, diagonal pan, and the personally observed failing gesture. Recordings from actual hardware must use source: "macos-trackpad"; generated fixtures use source: "synthetic".
Benchmark a real .fig fixture
Use current-document with an explicit browser-served fixture path. The runner waits for loading and page population, applies layout through the normal document-open path, zooms to fit, and waits on the active renderer's explicit settlement contract before recording:
bun run benchmark:navigation \
--url http://localhost:1420/ \
--gesture tests/fixtures/navigation/gestures/synthetic-pinch-reversal.json \
--mode dom \
--scenario current-document \
--document tests/fixtures/gold-preview.fig \
--no-trace \
--output artifacts/navigation-benchmark/gold-preview
environment.json records the resolved local document path so reports cannot silently mix generated and real-document runs. The runner serves the exact local bytes through an isolated Playwright route, so production Vite preview does not return its SPA fallback for non-public test fixtures.
Replay with Chromium tracing
bun run benchmark:navigation \
--url http://localhost:1420/ \
--gesture tests/fixtures/navigation/gestures/synthetic-pinch-reversal.json \
--mode cdp \
--scenario raster-stress \
--output artifacts/navigation-benchmark/current
--mode cdp sends browser input through the Chrome DevTools Protocol. --mode dom deterministically dispatches events inside the page and is useful for scheduler/correctness debugging. Neither mode is represented as physical WKWebView input.
Use --no-trace for the lowest-overhead metrics pass. Run a separate diagnostic pass with tracing enabled, and add --cpu-profile only when stack sampling is needed: CPU profiling can perturb the timing being measured.
Performance runs require a hardware-accelerated Metal/ANGLE context and fail if Chromium falls back to SwiftShader. Use --software-gpu only for portable correctness and CI smoke runs; never compare its timings with hardware results. Every run records the unmasked GL renderer and vendor in environment.json.
Open the trace:
open https://ui.perfetto.dev
Then load trace.json.gz and search for openpencil:.
Baseline comparison
Build and run the same gesture against v0.14.0 and the candidate in separate clean worktrees, on the same machine and power state. Alternate baseline/candidate runs rather than completing every baseline run first. Use release builds, fixed 1280×800 CSS viewport and DPR 2, and at least five measured repetitions after one warmup.
Do not gate shared CI on absolute timing. Dedicated benchmark hardware may gate on same-run baseline ratios and correctness invariants. Always retain raw traces for regressions.
Compare two completed runs:
bun tools/navigation-benchmark/src/cli.ts compare \
--baseline artifacts/navigation-benchmark/v0.14.0/metrics.json \
--candidate artifacts/navigation-benchmark/current/metrics.json \
--output artifacts/navigation-benchmark/comparison.json
Metrics
The report includes:
- Frame interval median, p95, p99, maximum, and counts over 8.33/16.67/33.33/50 ms.
- Render CPU interval and CanvasKit flush timing.
- Wheel-to-viewport and wheel-to-render-end latency.
- Maximum zoom focal-point drift in screen pixels.
- Maximum presented viewport displacement between rendered frames.
- Final input to exact active-renderer settlement: crisp retained backing for the existing renderer or exact visible tile coverage for tiled mode.
- Tiled scheduler frame count, maximum jobs per frame, maximum measured job submission, over-budget jobs, deadline overrun, and cancelled obsolete jobs.
The runner contains no warmup or settlement sleeps. waitForSettlement() requires an idle navigation lifecycle and renderer-owned exact coverage; its timeout rejects the benchmark and reports renderer state.
Benchmark content mutation and cancellation
Use the exact imported node ID plus a deterministic property mutation to measure selective refresh:
bun run benchmark:navigation \
--url 'http://localhost:1420/?renderer=tiled' \
--gesture tests/fixtures/navigation/gestures/synthetic-repeated-pinch-reversal.json \
--mode dom \
--scenario current-document \
--document tests/fixtures/gold-preview.fig \
--mutate-node '0:5' \
--mutate-opacity 0.11 \
--no-trace \
--output artifacts/navigation-benchmark/gold-preview-mutation
Add --replay-after-mutation to wait until refresh is observably pending and then replay the gesture before exact settlement. This measures generation cancellation and final-viewport prioritization without a fixed delay.
Mutation-only and combined runs record their mutation parameters in environment.json. For tiled runs, inspect scheduler.cancelledJobs, maximumJobsPerFrame, maximumJobRenderMs, overBudgetJobs, and maximumDeadlineOverrunMs.
Averages alone are not acceptance criteria. Inspect p95/p99, maximum stalls, contiguous missed frames, motion discontinuities, and crisp-settlement latency.
Required fixture matrix
Permanent benchmark scenarios should cover:
- Light: isolates input and scheduling overhead.
- Large flat: stresses culling and traversal.
- Realistic: nested frames, text, images, gradients, effects, masks, and instances.
- Raster stress: expensive retained-backing creation and coverage changes.
- Imported
.fig: exercises imported-document rendering paths.
Keep large or binary .fig fixtures in Git LFS. Synthetic fixture generation must be deterministic.
Native macOS acceptance
Chromium replay does not establish WKWebView or physical trackpad behavior. Before a release with navigation changes:
- Build a release-mode Tauri application.
- Record the same gesture on physical target hardware.
- Capture Instruments Time Profiler and Core Animation traces, including OpenPencil and WebKit processes.
- Check input delivery, main-thread/WASM work, GPU/compositor stalls, viewport continuity, and final crisp settlement.
- Store the recording, metrics, hardware/macOS metadata, and Instruments trace as release artifacts.
Native automation may add OS-level scroll injection, but it must preserve continuous-scroll and momentum phases before being called trackpad-equivalent. Synthetic composition or DOM wheel events are not native acceptance evidence.