* build: typecheck the test suites
Tests were in no TypeScript program: no tsconfig included tests/** or
packages/*/tests/**, and bun strips types without checking them, so a
fixture could drop a required field and keep passing until something
read it.
@types/bun moves to the root because it was installed per package only,
and #cli-tests/* joins the paths the root config already carries.
* test: fix the type errors the test suites were hiding
Typechecking the tests turned up 1123 errors. Most were ordinary
strictness, but some were real: `NodeChange` bound to Figma's plugin
typings rather than the Kiwi codec in thirteen .fig tests,
materializeInstance was called with six arguments against five so the
blobs and source children were dropped, CanvasKit pixels were written
to a plain object that never reached WASM, and assertions were made
through accessors that do not exist, so they asserted nothing.
Fixtures that had quietly lost a required field now carry it, nullable
results are narrowed through the existing expectDefined helper rather
than assumed, and stand-ins for CanvasKit and the editor go through one
named helper instead of an unexplained cast at each site.
No test was deleted, skipped, or weakened, and no `any`, non-null
assertion, or ts-expect-error was introduced.
* docs: record what typechecking the tests established
Pins the app program's global types with an assertion rather than a
note, since an unpinned types list lets any root @types package decide
which platform src/** is judged against.
The two environment faults that look like code regressions — Vite's
dependency pre-bundle outliving a package rebuild, and heavy .fig
suites failing under load — go to the development docs, where an
explanation belongs.
* fix: align @types/bun and keep node types resolvable when extended
The root manifest declared a newer @types/bun than every package, which
check:monorepo rejects, and pinning the app program's types left them
unresolvable from a config that extends this one out of tree.
* fix: fail the test typecheck when the compiler itself fails
The gate matched diagnostics by substring, so a compiler or config
failure that named no test file printed a pass while having checked
nothing. Diagnostics are now split by whether they name a file: an
unscoped one is the run failing and stops the gate, a test file's is a
finding, and a source file's stays out by design.
Also drops the parameter planComponentConstruction never read, and
makes the inner-shadow verification script exit non-zero when it
renders no image instead of logging and succeeding.
* chore: merge master into tests-typecheck
* feat(ai): render tool calls as summarized, highlighted cards
Every tool call showed only a status and its output as a JSON string,
so render calls hid their JSX, export_image dumped base64, and long
runs filled the transcript with identical rows.
A call now shows a one-line summary read from its input and chips that
select and zoom to the layers it touched, switching to the run's page
when needed. Expanded, it shows the JSX or script it wrote and its
JSON input and output in a read-only CodeMirror view, and exported
images inline. Render calls can be expanded while their input streams,
so the JSX appears alongside the canvas preview. Consecutive calls
beyond three fold into one row that keeps the latest call visible.
CodeMirror loads with the first expanded call. The code theme gains a
monospace fallback because the editor font variable is not always
emitted.
* refactor(ai): drop the unused tool JSON slot and place the JSX summary comment
* fix(ai): keep an opened tool call in place instead of following the output
Opening reasoning already stopped the transcript from following new output; tool calls and tool groups did not, so expanding one near the bottom re-pinned the bottom on every animation frame and slid the card away as it opened. Any disclosure in the transcript now stops following.
* fix(ai): show a pointer over chat tool calls, tool groups, and reasoning
* refactor(app): share CodeMirror setup between the code editor and viewer
CodeViewer repeated CodeEditor's view lifecycle: mounting the EditorView, label, theme, and language compartments, the app-theme watcher, and teardown. useCodeMirror owns that once; each component passes its own fixed and reactive extensions.
* refactor(ai): move tool node lookup and focusing into useToolNodes
ToolNodeChips looked nodes up in the active document and ran the show-on-canvas flow, with its superseded-switch and error handling, inside the component. The composable owns both; the component renders the chips.
* refactor(ai): derive tool call state and input once
ToolCallCard and ToolCallGroup each rebuilt classifyToolState's input from the part, and the card decided inline whether a call had input to show. toolCallState and toolHasInput own those rules beside the other per-call helpers.
* fix(app): use the thin app scrollbar in code editors and viewers
CodeMirror scrolls its own .cm-scroller, which fell back to the platform scrollbar, thick and light in the dark chat. The hosts now give it the shared scrollbar-thin utility.
Reasoning effort was a free-text profile field that only reached OpenAI
and OpenRouter, so Anthropic, Google, and DeepSeek models never thought
in direct chat. AI SDK 7 standardizes a `reasoning` call option that
those providers map to their own thinking settings, so profiles now
store one typed thinking level, shared with Pi, and requests pass it
through that option. OpenRouter's provider ignores the standard option
and receives its own reasoning option instead.
The composer offers the level next to the Design profile and reads it
per request, so a change applies to the next message without
rebuilding the transport. Saved profiles migrate from the Pi level or
the old effort string. Finished reasoning shows how long the model
thought while the block streamed.
* feat: preview streamed JSX on the canvas
Project incomplete JSX into isolated scene graphs and disposable pictures without mutating the document or adding intermediate undo entries. Share placement with final rendering and cover lifecycle and placement parity with AI SDK mocks and visual tests.
* test: require partial input for unfinished coordinates
Assert the complete partial object so rejecting the entire input cannot satisfy the truncated-exponent regression test. Addresses CodeRabbit's review finding on #692.
* feat(ai): keep a chat run on its page across page switches
Page switches go through the editor's preparation flow, and the chat panel treated every preparation as a document change: it dropped its Chat and reloaded history, detaching the panel from a reply still in progress. The panel now keeps the live chat unless the tab or the conversation changes.
AI tools also followed the page on screen, so a user browsing mid-run sent the next edits elsewhere, and the agent's own switch_page affected only one call. A run now pins the page where the message started; switch_page moves the run and the user's view, and streamed previews stay attached to the run's page, which the renderer draws only while that page is on screen.
Page snapshots now restore the page they were taken of, so undoing an AI edit works while another page is visible.
* refactor(core): share picture recording and export preparation with previews
Preview recording reimplemented three pieces Core already had: world-bounds picture recording (also duplicated by render chunks and the retained backing), font and layout preparation (prepareForExport), and page subgraph extraction. Extract recordWorldPicture and withWorldViewport for all three recorders, reuse prepareForExport, and add extractPageContext and findPageChildId next to the other subgraph helpers instead of editing a cloned graph's nodes.
prepareForExport also kept the shared layout text measurer overridden across an await, so a concurrent layout could measure with the export renderer. withTextMeasurer scopes the override to the synchronous layout.
* fix(design-jsx): inline nested fragments in streamed previews
The streaming projection kept a nested fragment as an empty-type node, which rendered trees inline, so a preview of <Frame><>…</></Frame> failed with 'Unknown element: <>'.
* refactor(ai): schedule previews and gate test streams with VueUse
The preview controller hand-rolled a trailing timer and abort-listener cleanup, and the test stream gate a promise resolver and listener set. Use useDebounceFn with maxWait (a lone delta still flushes, unlike useThrottleFn with leading off), useEventListener, and until(). Share the mock token usage between chat tests.
* fix(ai): keep previews alive through document edits and slow builds
Document edits finished every preview call, and onInputStart never restarts one, so a render call committing while a second was still streaming ended the second call's preview for good. Edits now invalidate: drop the shown artifact and rebuild on the new document.
A build that finished after another delta arrived was discarded, so a steady stream that outpaced staging and recording never showed a preview. Show it, then render the newer revision.
* docs(changelog): separate the Fixed heading from its entries
Add the blank line markdownlint (MD022) expects after the heading, and drop the one that split the Fixed list in two.
* feat(settings): configure tool access and agent step limits
Built-in AI exposed only a hardcoded subset of the tool registry, and the
maximum agent steps was a constant, so users could neither enable
extended tools such as create_component nor adjust long-running tasks.
Built-in AI and the local MCP server now keep independent, locally saved
tool permissions over one shared catalog, with searchable read-only and
side-effect groups and per-target defaults. Chat settings gain a validated
maximum-steps field whose captured value drives the stop condition,
remaining-step warnings, and limit detection for each message.
Tool access, the local server, browser access, and MCP connections are
grouped under a single Automation settings page.
Closes#573Closes#584
* refactor(settings): split automation into MCP and Tool access pages
The Automation page mixed a permission matrix with server endpoints behind
a Tools/Connections switch, and the view switch was indistinguishable from
the provider switch. The nested scroll region showed three of 110 tools.
Rename the MCP-facing page to MCP and give tool permissions their own Tool
access page. The page owns a fixed toolbar for the target, count, defaults,
and search, so the list uses the full dialog body and no row is clipped.
* fix(automation): explain MCP startup failures with localized guidance
Every startup failure collapsed into "MCP server did not become healthy":
the spawn layer recorded the real error but the runtime discarded it, and
health probes could not distinguish a rejected token from a missing server.
The message also surfaced raw English text as the alert heading.
Classify failures by reason (not installed, denied command, early exit,
startup timeout, rejected token, unexpected response, unreachable) and
render translated heading and guidance from the catalog, keeping captured
stderr or HTTP status as labeled diagnostic detail.
* refactor(ui): share one collapsible disclosure primitive
Six features each wired Reka's collapsible with their own motion classes and
one settings-only theme token, so the same interaction drifted in spacing,
icon size, and reduced-motion handling.
Add AppCollapsible with a family theme and move the settings disclosure and
the model editor's advanced settings onto it. Chat and frame-preset call
sites keep their distinct visuals for a follow-up.
* fix(automation): explain MCP failures with localized details
The failure alert carried raw English error text as its heading, and the
diagnostic payload sat in a sibling block outside the alert with no
relationship to it.
Classify failures by reason, render translated heading and guidance from
the catalog, and keep the payload in a collapsible inside the alert, which
unmounts while collapsed so the live region announces only the summary.
Add a copy action for issue reports.
Find the executable where a graphical launch can: extend PATH with the
common global bin directories before the lookup and report the searched
directories as diagnostic detail.
* fix(automation): keep MCP failure details out of reasons already explained
An unreachable address and a rejected token already name their cause in the
translated guidance, so repeating it under Details added noise. Details now
carry only output the summary cannot: stderr, HTTP status, or an unknown
error message.
* test(settings): browse every MCP failure reason in Storybook
The failure copy lived inside the settings panel, so reviewing the eight
reasons meant reproducing each failure and the mapping could only be
checked through the panel's dependencies.
Extract MCPFailureAlert, which owns the reason-to-copy mapping, detail
visibility, copy action, and restart action, and add a story covering
every reason plus the collapsed-details behavior.
* fix(ui): order alert details above the recovery actions
The alert rendered its action buttons before the details slot, so the
collapsible explanation of a failure appeared under the controls it
explains. Details now render directly after the description.
* fix(automation): correct MCP failure classification and detail
Review follow-ups on the failure diagnostics.
Only 401 and 403 mean the server refused our token; any other status now
reports an unexpected response instead of telling the user to replace a
token that was never the problem.
The install hint rendered the whole diagnostic detail as its package
argument, so searched directories appeared inside the install command.
The install target is now a domain constant and the searched directories
stay as detail, which not-installed failures surface again since they are
the actionable desktop diagnostic.
Exited failures also record the process exit code and signal so copied
diagnostics stay conclusive when stderr is empty. The bundled PATH test
now covers the append branch instead of only the unchanged path.
* feat(settings): accept custom values for presets and retention
Retention was a closed set of three counts while the AI step limit was a
free number, so two bounded numeric preferences looked and behaved
differently for no product reason.
Add a shared preset-or-custom field: presets stay one click, the escape
hatch reveals a validated numeric field, and the model carries only the
resolved number. Diagnostics retention becomes a bounded number (50 to
20,000) with the presets as shortcuts, and the hardcoded revalidation in
the panel is replaced by one domain resolver.
* fix(settings): label the preset and custom fields
Replacing the labeled provider field with the shared control left the AI
step limit as a bare select with a detached hint paragraph, outside the
settings group, so nothing on screen said what the number meant. The
accessibility name came from aria-label, which is why behavior tests
passed while the panel was unreadable.
Move both controls into labeled settings rows with their descriptions, and
give the revealed field its own accessible name so the two controls in one
row differ. The specs now assert the control lives inside the row that
names it, which is the check that would have caught this.
* fix(mcp): allow the desktop app origin by default
A server started manually bound the port and answered curl but the app
webview could not use it: no CORS origin was configured, so the browser
blocked every fetch and the app reported the server as unhealthy. The
workaround required an undocumented environment variable.
Allow the desktop app origins by default, accept a comma-separated
override, and document the default in the CLI help and the security notes.
Authenticated requests still need the bearer token, and browsers set Origin
themselves, so only the app webview can present these origins.
* fix(settings): address review findings on the new controls
Copy details awaited nothing and confirmed the copy before the write
finished. VueUse never rejects and falls back to a legacy write, so the
await is what makes the confirmation honest rather than an error branch.
The preset field only left custom mode when a preset arrived; a non-preset
value assigned from the owner left the select showing a value absent from
its options with the field still hidden. The watcher now follows the model
in both directions.
The story play functions queried the revealed field by the row label, which
Testing Library matches as a whole string, so those interactions could not
find it. The Storybook smoke assertion also assumed a button or tab, which
skipped every story built from other primitives.
* feat: persist AI conversations and add chat history
Store transcripts and attachment previews in IndexedDB, separate conversation history from transport lifetime, and add document-aware switching with shared Storybook coverage. Preserve interrupted activity and guard stale writes and async switches.
* test: group chat history tests by domain
* feat: simplify chat history navigation and diagnostics
Use a compact header with searchable document-scoped history and a correctly anchored conversation menu. Move diagnostic copying into the menu with copy-result feedback, and separate Storybook fixtures from composition.
* fix: address chat history review findings
* test: cover harness shutdown and restored chat scrolling
* fix: preserve Portless proxy port in MCP routes
* docs: clarify UI animation conventions
* feat: configure reasoning display and animate disclosure
* fix: separate transcript following from reasoning disclosure
* fix: restore selected chat after document recovery
* refactor: separate chat history persistence and sessions
* style: format reasoning story imports
Merges the contributor fix with maintainer follow-up coverage. MCP results now treat omitted isError as success, scope detection to mcp__ tools, and preserve generic tool error handling.