openpencil/tests/AGENTS.md
Danila Poyarkov aa86873dd7
test: typecheck the test suites and fix what that found (#896)
* build: typecheck the test suites

Tests were in no TypeScript program: no tsconfig included tests/** or
packages/*/tests/**, and bun strips types without checking them, so a
fixture could drop a required field and keep passing until something
read it.

@types/bun moves to the root because it was installed per package only,
and #cli-tests/* joins the paths the root config already carries.

* test: fix the type errors the test suites were hiding

Typechecking the tests turned up 1123 errors. Most were ordinary
strictness, but some were real: `NodeChange` bound to Figma's plugin
typings rather than the Kiwi codec in thirteen .fig tests,
materializeInstance was called with six arguments against five so the
blobs and source children were dropped, CanvasKit pixels were written
to a plain object that never reached WASM, and assertions were made
through accessors that do not exist, so they asserted nothing.

Fixtures that had quietly lost a required field now carry it, nullable
results are narrowed through the existing expectDefined helper rather
than assumed, and stand-ins for CanvasKit and the editor go through one
named helper instead of an unexplained cast at each site.

No test was deleted, skipped, or weakened, and no `any`, non-null
assertion, or ts-expect-error was introduced.

* docs: record what typechecking the tests established

Pins the app program's global types with an assertion rather than a
note, since an unpinned types list lets any root @types package decide
which platform src/** is judged against.

The two environment faults that look like code regressions — Vite's
dependency pre-bundle outliving a package rebuild, and heavy .fig
suites failing under load — go to the development docs, where an
explanation belongs.

* fix: align @types/bun and keep node types resolvable when extended

The root manifest declared a newer @types/bun than every package, which
check:monorepo rejects, and pinning the app program's types left them
unresolvable from a config that extends this one out of tree.

* fix: fail the test typecheck when the compiler itself fails

The gate matched diagnostics by substring, so a compiler or config
failure that named no test file printed a pass while having checked
nothing. Diagnostics are now split by whether they name a file: an
unscoped one is the run failing and stops the gate, a test file's is a
finding, and a source file's stays out by design.

Also drops the parameter planComponentConstruction never read, and
makes the inner-shadow verification script exit non-zero when it
renders no image instead of logging and succeeding.

* chore: merge master into tests-typecheck
2026-10-05 12:42:38 +00:00

5.9 KiB

Tests

packages/docs/development/testing.md is the canonical testing architecture: ownership table, fixture/driver/probe roles, browser and native boundaries, execution commands, server ownership, and worktree ports. This file holds the rules that bite during implementation.

Placement

  • Package-local tests mirror source domains under packages/<owner>/tests/; central app tests mirror src/app/** under tests/app/; tests/integration/ requires a genuinely cross-owner contract; E2E under tests/e2e/ follows user workflows; native and Figma acceptance are explicit exceptions.
  • Existing tests/engine/** domains migrate together with runner discovery: tools/dev/unit-tests/src/shards.ts lists each owner's canonical home and its current tests/engine directories, so a move is a git mv plus imports. bun run check:test-homes rejects any test added under tests/engine and any stale entry in tools/checks/test-homes/engine-baseline.txt; new tests go to the canonical home, and a moved file is removed from the baseline (--write regenerates it). Scene Graph has migrated to packages/scene-graph/tests and the editor domain to packages/core/tests/editor; put new Core tests there, mirroring packages/core/src.
  • Address a package's source by #<pkg>/* and its own helpers by #<pkg>-tests/* (#core-tests/*, #fig-tests/*); never drill with ../../. Registering a new alias means imports in the package manifest and PACKAGE_ALIASES/PACKAGE_ALIAS_OWNERS in tools/checks/architecture/src/steiger-rules/support.ts.
  • Owner-local helpers and fixtures stay local; only genuinely shared support goes under central tests/helpers/<domain>/ and tests/fixtures/. MCP transport tests: tests/engine/mcp/{server,stdio,transport} with tests/helpers/mcp.
  • Never commit temporary, diagnostic, or profile specs; keep them in ignored scratch/.

Writing specs

  • Test contracts and observable behavior, not source text. Specs use domain drivers and probes, not scattered Window/store traversal or unrestricted evaluator wrappers.
  • Locate behavior by accessible role and name, then label, then visible text. Scope repeated controls to a named region. Use scoped data-slot anatomy or semantic attributes (data-property, data-command, data-node-id) when needed; reserve data-test-id for integration boundaries and never add test-hook props or compound IDs.
  • Prefer test-runner-owned fixtures and request/route counters over browser globals. For in-page performance instrumentation, return a scoped JSHandle from evaluateHandle(), restore patched methods and listeners, and dispose the handle in finally; handles do not survive navigation. Assert transient DOM state with locators before the interaction ends.
  • Do not create a catch-all test Window interface or ad-hoc counter properties on window. Native-test declarations live in tests/helpers/tauri/native-global.d.ts; never expand production Window declarations for fixtures.
  • Tests are typechecked: bun run check:test-types covers tests/** and packages/*/tests/** through tsconfig.tests.json, and reports test files only because sources are judged by bun run typecheck against the globals they ship with.
  • Build a fixture with the owning factory (createDefaultNode, colorToFill) rather than a partial literal asserted to the type; a partial that drops a required field is exactly what the typecheck exists to catch.
  • A stand-in for an interface a test cannot build (CanvasKit, EditorStore) goes through asDouble from #tests/helpers/doubles, which is the one sanctioned widening; open-pencil/no-broad-double-cast rejects the as unknown as spelling.
  • In Bun tests prefer injected dependencies or scoped spies with explicit cleanup. mock.restore() restores spies but does not undo mock.module() overrides; do not assume module mocks are isolated by cleanup hooks. Read the installed runner's lifecycle and mocking docs before adding global or module-level instrumentation.
  • Make flaky tests deterministic; raising a timeout is never the fix. Keep package-manager invocations out of bun test suites; their cold start is not bounded.

Browser runs

  • Use the canonical playwright.config.ts; do not create task-specific config copies or server runners. Playwright owns Vite; the Vite automation plugin owns MCP startup and cleanup; browser fixtures own interactions, not server processes.
  • Test scripts select their server; direct Playwright commands start both servers unless OPENPENCIL_TEST_SERVER=app|storybook|all is set. Managed runs start the intended checkout; server reuse is opt-in for local development only, never for baseline comparisons or CI. Isolate the app URL, MCP endpoint, CORS origin, socket, and discovery path together.
  • Pixel-affecting renderer changes need committed canvas snapshots (packages/core/AGENTS.md, Renderer). Update only the justified affected snapshot and rerun without update mode.

Native WebView

  • Native checks live under tests/e2e/native/** and run through WebdriverIO against an explicit test-only Tauri binary: bun run test:native builds and runs, bun run build:native-test only builds. The binary uses a separate application identifier, an ephemeral WebView data store, and process-memory credentials.
  • Never run UI smoke tests against production Keychain entries or clear user recovery data to unblock tests. Persistence across restarts needs a dedicated test-owned persistent profile.
  • Native tests answer only whether the real WebView and Tauri shell deliver an interaction; engine tests cover state contracts and Playwright covers app integration. Platform-limited checks skip rather than claim coverage. Synthetic composition does not prove IME behavior; native clipboard remains a separate acceptance gap without trusted OS clipboard events.
  • Centralize native invocation in a guarded helper using vendor-derived types; do not import packages inside serialized WebView callbacks or repeat direct Tauri-global access in specs.