* build: typecheck the test suites
Tests were in no TypeScript program: no tsconfig included tests/** or
packages/*/tests/**, and bun strips types without checking them, so a
fixture could drop a required field and keep passing until something
read it.
@types/bun moves to the root because it was installed per package only,
and #cli-tests/* joins the paths the root config already carries.
* test: fix the type errors the test suites were hiding
Typechecking the tests turned up 1123 errors. Most were ordinary
strictness, but some were real: `NodeChange` bound to Figma's plugin
typings rather than the Kiwi codec in thirteen .fig tests,
materializeInstance was called with six arguments against five so the
blobs and source children were dropped, CanvasKit pixels were written
to a plain object that never reached WASM, and assertions were made
through accessors that do not exist, so they asserted nothing.
Fixtures that had quietly lost a required field now carry it, nullable
results are narrowed through the existing expectDefined helper rather
than assumed, and stand-ins for CanvasKit and the editor go through one
named helper instead of an unexplained cast at each site.
No test was deleted, skipped, or weakened, and no `any`, non-null
assertion, or ts-expect-error was introduced.
* docs: record what typechecking the tests established
Pins the app program's global types with an assertion rather than a
note, since an unpinned types list lets any root @types package decide
which platform src/** is judged against.
The two environment faults that look like code regressions — Vite's
dependency pre-bundle outliving a package rebuild, and heavy .fig
suites failing under load — go to the development docs, where an
explanation belongs.
* fix: align @types/bun and keep node types resolvable when extended
The root manifest declared a newer @types/bun than every package, which
check:monorepo rejects, and pinning the app program's types left them
unresolvable from a config that extends this one out of tree.
* fix: fail the test typecheck when the compiler itself fails
The gate matched diagnostics by substring, so a compiler or config
failure that named no test file printed a pass while having checked
nothing. Diagnostics are now split by whether they name a file: an
unscoped one is the run failing and stops the gate, a test file's is a
finding, and a source file's stays out by design.
Also drops the parameter planComponentConstruction never read, and
makes the inner-shadow verification script exit non-zero when it
renders no image instead of logging and succeeding.
* chore: merge master into tests-typecheck
* feat(app): record runtime errors in diagnostics with their stack
Uncaught errors and unhandled rejections only showed a toast, Vue component errors after boot only reached the console, and a failed chat kept just its error name, so a failure like WebKit's 'Attempting to define property on object that is not extensible.' left nothing to diagnose. They now record a runtime.error, and chat.failed its code, message, and stack. Messages and stacks are scrubbed of URL queries, key- and token-like strings, and home folder names and bounded; AI SDK and provider errors keep no message, since it can quote prompts or responses. Copied diagnostics start with the app version, shell, browser, and language.
* feat(app): label, filter, and page diagnostics events
Every row in Settings → Diagnostics read 'Technical event': the summary looked labels up under diagnostics-prefixed keys the messages do not have, and only a few event kinds had labels at all. Each event now has a specific label and a short detail, such as 'Tool: render · 162 ms', 'Model step · <model>', or an error's message, expands to its recorded fields and stack, and the list filters by level and category and grows a page at a time. The copy action passes the environment header, which moves out of the recorder so tooling that compiles it needs no build-time globals.
* test(app): stream a reasoning reply in WebKit without page errors
Errors such as WebKit's "not extensible" TypeError appear only in that engine, so run a streamed reasoning reply there and fail on any page or console error.
* fix(app): scrub queries on bare paths in diagnostic errors
Only URLs had their query removed, so a message like 'Failed to load /Designs/app.fig?token=…' kept the token.
* fix(app): count diagnostics recorded before Settings opens
The event count and size updated only on new events, so the panel showed 0 events beside a full list.
* feat(app): record failed AI tool calls as problems, with the stack of engine errors
A tool catches what it throws and returns only the message to the model, so diagnostics saw a failed tool as an info event without details. The adapter now passes the thrown value to the tool log. A failed call is a warning; a TypeError, ReferenceError, or RangeError, which comes from a bug in OpenPencil rather than a wrong call, is an error with its message and stack. Other tool errors keep only their name, since their messages quote layer names and arguments.
* fix(core): log tool calls that return an error as failed
Most tools report a failure by returning { error } rather than throwing, such as describe with an unknown node, so the tool log and diagnostics counted them as successful calls while the chat showed them failed.
* feat(app): scrub cloud keys, JWTs, private keys, URL credentials, and emails from diagnostics
The scrubber caught keys by shape only, so 20-character AWS access key IDs, user:pass@ in URLs, and emails reached the log, and a JWT's payload survived because its dots split it into short runs. It moves into its own module with rules grouped by what they protect. The added credential formats follow gitleaks; keys the shape rules already catch, such as GitHub, OpenAI, Anthropic, and Stripe ones, get no separate rule. No maintained browser library fits: secretlint needs Node built-ins and adds at least 23 KB gzipped, and the PII redactors miss tokens. The scrubber is 0.8 KB gzipped.
* refactor(app): name how a tool call is recorded and import diagnostics from its index
* fix(app): record demo document loads in diagnostics
The preparation event's schema listed its kinds, phases, cancel reasons, and failure codes by hand and lacked demo-load, so every demo load failed validation and was dropped. The schema now validates against the same lists the preparation types derive from.
* feat(app): label document preparation events in Settings diagnostics
Preparation events showed their raw name, editor.preparation.finished, because the summary had no label for them. They now read as their kind, such as Switch page, with the outcome and duration below. Event names are a typed union and the labels a map keyed by it, so recording a new event without a label fails type-checking; names stored by older versions still fall back to the raw name.
* fix(app): keep source paths and scrub provider stacks, auth headers, and spaced home folders
Review follow-up. A provider error's message was dropped but repeated on its stack's first line, so it is now removed there too. The long-run rule redacted source paths of 40 or more characters, losing the failing file; a run with slashes now loses only its key-like segments. The bare-path query rule cut optional chaining such as a.b?.c and now needs name= after the question mark. Authorization header values in any scheme, credential assignments such as api_key= or password:, and home folder names with spaces are now scrubbed.
* fix(app): suppress repeats of alternating runtime errors
Repeat suppression compared each error only with the previous one, so a loop alternating between two errors recorded every occurrence. Recent errors are now kept in a small bounded map.
* test(app): validate copied diagnostics and wait for the copy to finish
Master now rejects JSON.parse with a type assertion, so the copied report is read through a Valibot schema. The uncaught-error test read the clipboard before its copy finished and could see the previous test's report; it now waits for the confirmation, as the export test does.