openpencil/tests/e2e/settings/diagnostics.spec.ts
Danila Poyarkov a50444c265
feat(app): record runtime errors in diagnostics and make events readable (#874)
* feat(app): record runtime errors in diagnostics with their stack

Uncaught errors and unhandled rejections only showed a toast, Vue component errors after boot only reached the console, and a failed chat kept just its error name, so a failure like WebKit's 'Attempting to define property on object that is not extensible.' left nothing to diagnose. They now record a runtime.error, and chat.failed its code, message, and stack. Messages and stacks are scrubbed of URL queries, key- and token-like strings, and home folder names and bounded; AI SDK and provider errors keep no message, since it can quote prompts or responses. Copied diagnostics start with the app version, shell, browser, and language.

* feat(app): label, filter, and page diagnostics events

Every row in Settings → Diagnostics read 'Technical event': the summary looked labels up under diagnostics-prefixed keys the messages do not have, and only a few event kinds had labels at all. Each event now has a specific label and a short detail, such as 'Tool: render · 162 ms', 'Model step · <model>', or an error's message, expands to its recorded fields and stack, and the list filters by level and category and grows a page at a time. The copy action passes the environment header, which moves out of the recorder so tooling that compiles it needs no build-time globals.

* test(app): stream a reasoning reply in WebKit without page errors

Errors such as WebKit's "not extensible" TypeError appear only in that engine, so run a streamed reasoning reply there and fail on any page or console error.

* fix(app): scrub queries on bare paths in diagnostic errors

Only URLs had their query removed, so a message like 'Failed to load /Designs/app.fig?token=…' kept the token.

* fix(app): count diagnostics recorded before Settings opens

The event count and size updated only on new events, so the panel showed 0 events beside a full list.

* feat(app): record failed AI tool calls as problems, with the stack of engine errors

A tool catches what it throws and returns only the message to the model, so diagnostics saw a failed tool as an info event without details. The adapter now passes the thrown value to the tool log. A failed call is a warning; a TypeError, ReferenceError, or RangeError, which comes from a bug in OpenPencil rather than a wrong call, is an error with its message and stack. Other tool errors keep only their name, since their messages quote layer names and arguments.

* fix(core): log tool calls that return an error as failed

Most tools report a failure by returning { error } rather than throwing, such as describe with an unknown node, so the tool log and diagnostics counted them as successful calls while the chat showed them failed.

* feat(app): scrub cloud keys, JWTs, private keys, URL credentials, and emails from diagnostics

The scrubber caught keys by shape only, so 20-character AWS access key IDs, user:pass@ in URLs, and emails reached the log, and a JWT's payload survived because its dots split it into short runs. It moves into its own module with rules grouped by what they protect. The added credential formats follow gitleaks; keys the shape rules already catch, such as GitHub, OpenAI, Anthropic, and Stripe ones, get no separate rule. No maintained browser library fits: secretlint needs Node built-ins and adds at least 23 KB gzipped, and the PII redactors miss tokens. The scrubber is 0.8 KB gzipped.

* refactor(app): name how a tool call is recorded and import diagnostics from its index

* fix(app): record demo document loads in diagnostics

The preparation event's schema listed its kinds, phases, cancel reasons, and failure codes by hand and lacked demo-load, so every demo load failed validation and was dropped. The schema now validates against the same lists the preparation types derive from.

* feat(app): label document preparation events in Settings diagnostics

Preparation events showed their raw name, editor.preparation.finished, because the summary had no label for them. They now read as their kind, such as Switch page, with the outcome and duration below. Event names are a typed union and the labels a map keyed by it, so recording a new event without a label fails type-checking; names stored by older versions still fall back to the raw name.

* fix(app): keep source paths and scrub provider stacks, auth headers, and spaced home folders

Review follow-up. A provider error's message was dropped but repeated on its stack's first line, so it is now removed there too. The long-run rule redacted source paths of 40 or more characters, losing the failing file; a run with slashes now loses only its key-like segments. The bare-path query rule cut optional chaining such as a.b?.c and now needs name= after the question mark. Authorization header values in any scheme, credential assignments such as api_key= or password:, and home folder names with spaces are now scrubbed.

* fix(app): suppress repeats of alternating runtime errors

Repeat suppression compared each error only with the previous one, so a loop alternating between two errors recorded every occurrence. Recent errors are now kept in a small bounded map.

* test(app): validate copied diagnostics and wait for the copy to finish

Master now rejects JSON.parse with a type assertion, so the copied report is read through a Valibot schema. The uncaught-error test read the clipboard before its copy finished and could see the previous test's report; it now waits for the confirmation, as the export test does.
2026-10-05 07:59:44 +00:00

142 lines
6 KiB
TypeScript

import type { Page } from '@playwright/test'
import * as v from 'valibot'
import { CanvasHelper } from '#tests/helpers/canvas'
import { expect, test } from '#tests/helpers/chat/fixture'
/** The JSON that Copy diagnostics puts on the clipboard, as far as these tests read it. */
const ReportJSON = v.pipe(
v.string(),
v.parseJson(),
v.object({
environment: v.object({ app: v.string(), shell: v.string() }),
events: v.array(
v.object({
name: v.string(),
sessionId: v.optional(v.string()),
runId: v.optional(v.string()),
attributes: v.record(v.string(), v.unknown())
})
)
})
)
test('Settings exports correlated chat telemetry without conversation content', async ({
configuredChat: chat,
page,
context
}) => {
await context.grantPermissions(['clipboard-read', 'clipboard-write'])
await chat.submit('Private conversation content must not enter diagnostics')
await expect(chat.assistantMessage()).toBeVisible()
await expect(chat.input).toBeEnabled()
await chat.submit('A second request in the same conversation')
await expect(page.getByTestId('chat-message-assistant')).toHaveCount(2)
await page.getByTestId('app-settings-trigger').click()
await page.getByTestId('settings-section-diagnostics').click()
await page.getByRole('button', { name: 'Copy diagnostics', exact: true }).click()
await expect(page.getByText('Diagnostics copied to clipboard.', { exact: true })).toBeVisible()
const text = await page.evaluate(() => navigator.clipboard.readText())
const { environment, events } = v.parse(ReportJSON, text)
expect(environment).toMatchObject({ app: expect.any(String), shell: 'browser' })
const completed = events.filter((event) => event.name === 'chat.completed')
expect(completed).toHaveLength(2)
expect(completed[0].sessionId).toEqual(expect.any(String))
expect(completed[1].sessionId).toBe(completed[0].sessionId)
expect(completed[0].runId).toEqual(expect.any(String))
expect(completed[1].runId).toEqual(expect.any(String))
expect(completed[1].runId).not.toBe(completed[0].runId)
expect(text).not.toContain('Private conversation content')
expect(text).not.toContain('A second request')
})
async function openDiagnostics(page: Page) {
await page.keyboard.press(process.platform === 'darwin' ? 'Meta+,' : 'Control+,')
await page.getByTestId('settings-section-diagnostics').click()
}
async function reloadAndOpenDiagnostics(page: Page) {
await page.reload()
await new CanvasHelper(page).waitForInit()
await openDiagnostics(page)
}
test('diagnostics retention offers presets and accepts a bounded custom value', async ({
page
}) => {
await page.goto('/?test')
await new CanvasHelper(page).waitForInit()
await openDiagnostics(page)
const row = page
.locator('[data-slot="settings-row"]')
.filter({ hasText: 'Diagnostics retention' })
await expect(row).toContainText('Keep up to this many recent events locally.')
const retention = row.getByRole('combobox', { name: 'Diagnostics retention' })
await expect(retention).toHaveText('500')
// A custom value is validated against the supported range before it is kept.
await retention.click()
await page.getByRole('option', { name: 'Custom…' }).click()
const custom = row.getByRole('spinbutton', { name: /Diagnostics retention/ })
await custom.fill('10')
await custom.press('Enter')
await expect(custom).toHaveAttribute('aria-invalid', 'true')
await expect(custom).toHaveAccessibleDescription(/Enter a whole number from 50 to 20000/)
await custom.fill('750')
await custom.press('Enter')
await expect(custom).toHaveAttribute('aria-invalid', 'false')
await reloadAndOpenDiagnostics(page)
await expect(retention).toHaveText('Custom…')
await expect(row.getByRole('spinbutton', { name: /Diagnostics retention/ })).toHaveValue('750')
// Presets remain one click and persist without revealing the field.
await retention.click()
await page.getByRole('option', { name: '1000', exact: true }).click()
await expect(retention).toHaveText('1000')
await reloadAndOpenDiagnostics(page)
await expect(retention).toHaveText('1000')
await expect(row.getByRole('spinbutton')).toHaveCount(0)
})
test('uncaught errors are listed with their stack and can be filtered and copied', async ({
configuredChat: chat,
page,
context
}) => {
await context.grantPermissions(['clipboard-read', 'clipboard-write'])
await chat.submit('A request that records info events')
await expect(chat.assistantMessage()).toBeVisible()
await page.evaluate(() => {
setTimeout(() => {
throw new TypeError('Attempting to define property on object that is not extensible.')
})
})
await openDiagnostics(page)
const events = page.locator('[data-slot="diagnostics-events"]')
const failure = events.getByRole('button', { name: /Error: TypeError/ })
await expect(failure).toContainText(
'Attempting to define property on object that is not extensible.'
)
await failure.click()
await expect(events.locator('pre')).toContainText('TypeError')
// Events recorded before Settings opened are counted.
await expect(page.getByText(/^[1-9]\d* events · /)).toBeVisible()
// Info events such as the completed chat drop out when only problems are shown.
await expect(events.getByRole('button', { name: /AI chat completed/ })).toBeVisible()
await events.getByRole('button', { name: 'Errors and warnings' }).click()
await expect(events.getByRole('button', { name: /AI chat completed/ })).toHaveCount(0)
await expect(failure).toBeVisible()
await page.getByRole('button', { name: 'Copy diagnostics', exact: true }).click()
// The clipboard outlives the test's page, so read it only once this copy has finished.
await expect(page.getByText('Diagnostics copied to clipboard.', { exact: true })).toBeVisible()
const report = v.parse(ReportJSON, await page.evaluate(() => navigator.clipboard.readText()))
const runtime = report.events.find((event) => event.name === 'runtime.error')
expect(runtime?.attributes).toMatchObject({ source: 'window', errorName: 'TypeError' })
expect(String(runtime?.attributes.stack)).toContain('TypeError')
})