Tools live under tools/<role>/<domain> (checks, generate, release, ci, dev), every tool is a workspace named @open-pencil/<domain>-tools, a shared tools/tsconfig.json backs the new check:tools gate that fixed 55 latent type errors, test:tools runs through bun --filter, the placement check is its own checks/test-homes package, and every tool resolves the repository through resolveWorkspaceRoot. Bun, Node, and mdast types live in a tools-root workspace so they never reach the app program.
4.8 KiB
4.8 KiB
Tests
packages/docs/development/testing.md is the canonical testing architecture: ownership table, fixture/driver/probe roles, browser and native boundaries, execution commands, server ownership, and worktree ports. This file holds the rules that bite during implementation.
Placement
- Package-local tests mirror source domains under
packages/<owner>/tests/; central app tests mirrorsrc/app/**undertests/app/;tests/integration/requires a genuinely cross-owner contract; E2E undertests/e2e/follows user workflows; native and Figma acceptance are explicit exceptions. - Existing
tests/engine/**domains migrate together with runner discovery:tools/dev/unit-tests/src/shards.tslists each owner's canonical home and its currenttests/enginedirectories, so a move is agit mvplus imports.bun run check:test-homesrejects any test added undertests/engineand any stale entry intools/dev/unit-tests/engine-baseline.txt; new tests go to the canonical home, and a moved file is removed from the baseline (--writeregenerates it). Scene Graph has migrated topackages/scene-graph/tests. - Owner-local helpers and fixtures stay local; only genuinely shared support goes under central
tests/helpers/<domain>/andtests/fixtures/. MCP transport tests:tests/engine/mcp/{server,stdio,transport}withtests/helpers/mcp. - Never commit temporary, diagnostic, or profile specs; keep them in ignored
scratch/.
Writing specs
- Test contracts and observable behavior, not source text. Specs use domain drivers and probes, not scattered Window/store traversal or unrestricted evaluator wrappers.
- Locate behavior by accessible role and name, then label, then visible text. Scope repeated controls to a named region. Use scoped
data-slotanatomy or semantic attributes (data-property,data-command,data-node-id) when needed; reservedata-test-idfor integration boundaries and never add test-hook props or compound IDs. - Prefer test-runner-owned fixtures and request/route counters over browser globals. For in-page performance instrumentation, return a scoped
JSHandlefromevaluateHandle(), restore patched methods and listeners, and dispose the handle infinally; handles do not survive navigation. Assert transient DOM state with locators before the interaction ends. - Do not create a catch-all test Window interface or ad-hoc counter properties on
window. Native-test declarations live intests/helpers/tauri/native-global.d.ts; never expand production Window declarations for fixtures. - In Bun tests prefer injected dependencies or scoped spies with explicit cleanup.
mock.restore()restores spies but does not undomock.module()overrides; do not assume module mocks are isolated by cleanup hooks. Read the installed runner's lifecycle and mocking docs before adding global or module-level instrumentation. - Make flaky tests deterministic; raising a timeout is never the fix. Keep package-manager invocations out of
bun testsuites; their cold start is not bounded.
Browser runs
- Use the canonical
playwright.config.ts; do not create task-specific config copies or server runners. Playwright owns Vite; the Vite automation plugin owns MCP startup and cleanup; browser fixtures own interactions, not server processes. - Test scripts select their server; direct Playwright commands start both servers unless
OPENPENCIL_TEST_SERVER=app|storybook|allis set. Managed runs start the intended checkout; server reuse is opt-in for local development only, never for baseline comparisons or CI. Isolate the app URL, MCP endpoint, CORS origin, socket, and discovery path together. - Pixel-affecting renderer changes need committed canvas snapshots (
packages/core/AGENTS.md, Renderer). Update only the justified affected snapshot and rerun without update mode.
Native WebView
- Native checks live under
tests/e2e/native/**and run through WebdriverIO against an explicit test-only Tauri binary:bun run test:nativebuilds and runs,bun run build:native-testonly builds. The binary uses a separate application identifier, an ephemeral WebView data store, and process-memory credentials. - Never run UI smoke tests against production Keychain entries or clear user recovery data to unblock tests. Persistence across restarts needs a dedicated test-owned persistent profile.
- Native tests answer only whether the real WebView and Tauri shell deliver an interaction; engine tests cover state contracts and Playwright covers app integration. Platform-limited checks skip rather than claim coverage. Synthetic composition does not prove IME behavior; native clipboard remains a separate acceptance gap without trusted OS clipboard events.
- Centralize native invocation in a guarded helper using vendor-derived types; do not import packages inside serialized WebView callbacks or repeat direct Tauri-global access in specs.