* feat(MCP): follow agent activity in canvas
* fix(fig): preserve imported design fidelity
Keep component overrides, variable-backed icon colors, page backgrounds, and fixed text sizing intact across lazy FIG materialization.
* feat: add selection-context MCP tools and mode
* chore: scope work branch to MCP selection and canvas follow
* fix: honor MCP-only tool contracts in CI
* refactor(mcp): drop the follow and selection-context tools this branch carried
Following agents landed in #725, through the agents registry and the
chat's follow toggle, so this branch's MCP follow setting and its
follow-agent module are superseded. The see_user_selection and
get_user_selection_details tools duplicated get_selection, get_node,
describe, get_page_tree, and export_image; the selection-only workflow
they served is rebuilt on those tools in the following commits.
Co-authored-by: Victor Wads <victor@wads.dev>
* feat(mcp): make get_selection the compact entry point with a depth
get_selection returned every selected layer's whole subtree, which is
too much as the first call when the user points at a large frame. It now
returns the selection with direct children by default, counts deeper
children as childCount, and takes a depth.
Co-authored-by: Victor Wads <victor@wads.dev>
* feat(mcp): share only the selection with MCP clients
A selection scope, set with Share only the selection in the local
server settings or OPENPENCIL_MCP_SCOPE=selection, limits MCP clients
to get_selection, get_node, get_page_tree, describe, and export_image
on the selected layers and what they hold.
The server enforces the scope on everything it sends to the app: MCP
sessions and /rpc, which stdio clients also go through, carry only
those tool calls and the session-closed notice, each stamped with the
scope, so a client cannot reach other tools or the settings that would
widen it. The app's bridge rejects node IDs outside the selection,
points describe and export_image at the selection when they name no
nodes, and asks get_page_tree for a root inside it. A stdio client can
ask for the scope itself while the server shares the whole document.
Co-authored-by: Victor Wads <victor@wads.dev>
* fix(mcp): keep selection-scoped clients from writing files or listing wider tools
export_image writes its result to a file when given a path and an MCP
root is set, which reaches past reading the selection. A path is now
refused in selection scope, by the tool registration before the call
and by the app's bridge, so a client with a stale scope cannot write
either; the image itself is still returned.
A stdio client follows the narrower of its own scope and the scope the
server records, instead of letting OPENPENCIL_MCP_SCOPE=document list
tools a selection-scoped server rejects.
Co-authored-by: Victor Wads <victor@wads.dev>
* test(mcp): name the selection scope's tools instead of reading the allowlist
The server test compared the listed tools with SELECTION_SCOPE_TOOLS,
the same list that decides registration, so a tool added to it by
mistake would still pass. It now names the five tools the scope offers.
Co-authored-by: Victor Wads <victor@wads.dev>
---------
Co-authored-by: Danila Poyarkov <dev@dannote.net>
* feat(MCP): follow agent activity in canvas
* feat(collab): show MCP, ACP, and harness sessions as agents
MCP clients worked on the document unseen: only the built-in chat had a
presence, and following agent activity meant a separate setting that
moved the viewport after every MCP tool. Each MCP session now shows as
an agent with a callsign in its owner's color, like the chat: the MCP
server forwards the session and the client's name with each tool call,
the app's own ACP and Pi harness chats mark their sessions with a
header, and the agent points at the layers each call reads or changes
on their page. It rests after a quiet spell, leaves when its session
ends or the server disconnects, and collaborators see it through
awareness. Following it works like following anyone else, from the
avatars, so the default-on follow setting and its viewport fitting go.
Co-authored-by: Victor Wads <victor@wads.dev>
* feat(ai): move the chat's agent through JSX as it streams
While the built-in chat streamed a render call, its preview grew on the
canvas but the agent stood still until the tool finished. The preview
now reports, after each update, the element that appeared last and the
bounds of what the JSX builds; the agent's cursor follows that element
and its outline traces the preview, for collaborators too, until the
tool runs and the agent outlines the real layers. Peers' outlines are
validated and capped like their selections.
Co-authored-by: Victor Wads <victor@wads.dev>
* feat(collab): glide cursors and the followed view instead of jumping
People's and agents' cursors jumped to each new point, which with
throttled awareness and an agent streaming JSX made them stutter, and
following re-centered the view in one jump on every update. Cursors now
ease to each new point from wherever they are drawn, and following pans
and zooms the view there the same way, stopping in place when you take
over. With animations off or reduced motion, both move at once. Each
cursor carries an id so it keeps its glide between updates.
Co-authored-by: Victor Wads <victor@wads.dev>
* test(collab): cover MCP agents and the streaming agent in the browser
Test runs send MCP requests through the bridge's own command handler, so
a browser test can show an MCP session as an agent to the editor and to
a collaborator without a separate server. The streaming JSX test checks
that the chat's agent cursor and outline follow the newest element.
Co-authored-by: Victor Wads <victor@wads.dev>
* docs: describe agent sessions, streaming cursors, and gliding follow
Co-authored-by: Victor Wads <victor@wads.dev>
* fix(ai): keep the streaming agent's name off the text it writes
The chat's agent sat at the newest streamed element's top-left corner,
so its name label covered the text being written. It now sits at the
element's trailing corner, where content grows.
Co-authored-by: Victor Wads <victor@wads.dev>
* fix(collab): end only a closed connection's own MCP sessions
When the app's connection to the MCP server closed, every MCP session's
agent left, including sessions that never came over that connection,
such as a second editor's. The bridge now remembers the sessions each
connection carried and ends only those.
Co-authored-by: Victor Wads <victor@wads.dev>
* feat(ai): follow your agents automatically while they work
Agents spun up from the chat or an MCP client worked out of sight and
then went idle, so people had to find what changed, and the agent skill
told agents to move the user's view and selection after every edit. A
Follow agents toggle in the AI panel's header, on by default, now has
the view follow our agents from the start of each run: whichever starts
first, then the next one at work once the followed agent rests. Leaving
the page, moving the view, Escape, or Stop following leaves that agent
alone until it rests; people are never followed this way. The skill no
longer asks agents to select and zoom to their work.
Co-authored-by: Victor Wads <victor@wads.dev>
* fix(collab): keep an MCP agent when a restarted bridge carries its session
Restarting MCP disconnects the old bridge and opens a new one at once,
but the browser reports the old socket's close only after its close
handshake. A tool call that reached the new connection in between was
undone by that close, which ended the session and removed its agent.
Sessions now record every connection their calls came over, across
bridges, and end only when the last of them closes.
The docs no longer say that following an agent keeps it from editing
out of sight: following moves your view, and the stale variables and
follow bullets the changelog's union merge brought back are removed.
Co-authored-by: Victor Wads <victor@wads.dev>
* test(collab): record what the bridge sends instead of an empty fake method
Co-authored-by: Victor Wads <victor@wads.dev>
---------
Co-authored-by: Danila Poyarkov <dev@dannote.net>
* build: typecheck the test suites
Tests were in no TypeScript program: no tsconfig included tests/** or
packages/*/tests/**, and bun strips types without checking them, so a
fixture could drop a required field and keep passing until something
read it.
@types/bun moves to the root because it was installed per package only,
and #cli-tests/* joins the paths the root config already carries.
* test: fix the type errors the test suites were hiding
Typechecking the tests turned up 1123 errors. Most were ordinary
strictness, but some were real: `NodeChange` bound to Figma's plugin
typings rather than the Kiwi codec in thirteen .fig tests,
materializeInstance was called with six arguments against five so the
blobs and source children were dropped, CanvasKit pixels were written
to a plain object that never reached WASM, and assertions were made
through accessors that do not exist, so they asserted nothing.
Fixtures that had quietly lost a required field now carry it, nullable
results are narrowed through the existing expectDefined helper rather
than assumed, and stand-ins for CanvasKit and the editor go through one
named helper instead of an unexplained cast at each site.
No test was deleted, skipped, or weakened, and no `any`, non-null
assertion, or ts-expect-error was introduced.
* docs: record what typechecking the tests established
Pins the app program's global types with an assertion rather than a
note, since an unpinned types list lets any root @types package decide
which platform src/** is judged against.
The two environment faults that look like code regressions — Vite's
dependency pre-bundle outliving a package rebuild, and heavy .fig
suites failing under load — go to the development docs, where an
explanation belongs.
* fix: align @types/bun and keep node types resolvable when extended
The root manifest declared a newer @types/bun than every package, which
check:monorepo rejects, and pinning the app program's types left them
unresolvable from a config that extends this one out of tree.
* fix: fail the test typecheck when the compiler itself fails
The gate matched diagnostics by substring, so a compiler or config
failure that named no test file printed a pass while having checked
nothing. Diagnostics are now split by whether they name a file: an
unscoped one is the run failing and stops the gate, a test file's is a
finding, and a source file's stays out by design.
Also drops the parameter planComponentConstruction never read, and
makes the inner-shadow verification script exit non-zero when it
renders no image instead of logging and succeeding.
* chore: merge master into tests-typecheck
* fix: validate parsed JSON at untrusted boundaries with Valibot
Clipboard HTML, library revisions from shared storage, MCP and automation
WebSocket messages, the MCP discovery file, sidecar output and AI/MCP tool
arguments were JSON.parse'd and cast to their expected types, so a
malformed payload reached the document or crashed paste. They now go
through v.pipe(v.string(), v.parseJson(), Schema), which reports bad JSON
and a wrong shape as the same validation failure.
The path_set tool rejects an invalid VectorNetwork and shares its parser
with create_vector. The CLI library catalog validates its files and runs
revisions through the same size, identity and content-hash checks as the
app; reading image bytes as index-keyed records also stops them coming
back empty. Hand-rolled typeof readers for plugin data, document metadata,
caches and preferences become schemas with their behaviour preserved, and
readCacheJSON takes a schema for its payload.
open-pencil/no-unvalidated-json-parse rejects type assertions on
JSON.parse results other than `as unknown` in src and packages/*/src.
* refactor: validate parsed JSON in tests and tooling
Extend open-pencil/no-unvalidated-json-parse beyond source: tests, helpers and repo tooling now parse JSON through Valibot schemas instead of asserting a type. The shared fixture reader returns a validated object; its old array annotation never matched the fixtures.
* fix: validate clipboard geometry bytes, library images and model catalogs
Clipboard geometry blobs and library image bytes must be bytes at contiguous indexes, so out-of-range or gapped values are rejected instead of silently becoming different geometry or images; serialized library nodes must carry source metadata. The models.dev and OpenRouter responses are validated like their cached copies, and activate-tab rejects a CDP frame it cannot read instead of hanging.
* refactor: extend the JSON validation lint to .json() results
no-unvalidated-json-parse now also rejects type assertions on Response, Bun.file and shell .json() results, the same unchecked parse in another form. MCP server tests read /health through a validated readHealth helper and discovery files through parseDiscoveryInfo; the remaining tooling reads its JSON through schemas.
* test: validate the RPC request body in the CLI app export test
* test: validate CLI JSON output in the tool and app command tests
* test: compare the malformed models.dev fallback with the curated list
* fix(app): record MCP and CLI structural edits as undo steps
The automation bridge ran non-atomic tools, render, and eval without an
undo entry, so Edit > Undo could not revert layers an MCP client or the
CLI created, deleted, or rearranged. Snapshot the page around these
edits as the AI chat does, and skip the entry when nothing changed so
read-only scripts leave the history alone.
* feat(app): activate documents, undo, redo, and change settings over automation
Add activate_document, undo, redo, get_settings, and update_settings to
the app's automation bridge. Settings cover appearance, snapping, canvas
rendering, recovery, and chat preferences, validated with Valibot and
applied through their owning stores; credentials, models, MCP
connections, storage, and tool access stay out of reach.
* feat(mcp): expose document activation, history, and settings tools
* feat(cli): manage documents, history, settings, and tools in the running app
Turn documents into a command group (list, open, new, save, close,
activate), add undo, redo, and settings get/set, and add tool
list/describe/call so every MCP tool runs from the shell, against the
running app or headlessly on a file.
* docs: document app control from the CLI and MCP
* fix: never prompt in the app from automation closes and saves
close_file opened the app's Save changes dialog, which an agent cannot
answer: the call timed out and the dialog stayed open. It now fails on
unsaved changes unless the caller passes unsaved "save" or "discard"
(CLI --save or --discard). save_file and new_document no longer open a
Save dialog for a document that was never saved, report a failed save
as an error, and leave the document untouched when the path is refused.
* docs: describe non-interactive close and save
* fix: address review findings in app automation
Keep a document's source when a save to a new path fails, report
vector-edit undo and redo no-ops as unapplied, echo only the applied
patch from update_settings so writing cannot read settings, reject
tool call --write/--output without a file, and stop settings get from
following inherited keys.
* fix(app): record render undo on the page that receives the layers
A render into a parent on another page was snapshotted against the
target page, so undo left the new layers in place. Snapshot the page
that contains the parent instead, and document that eval edits made
after switching pages stay outside the undo step.
* feat(app): limit automation undo to its own steps and expose design check settings
The undo history is shared with the person in the editor, so an agent's
undo could revert the user's last edit. Automation undo and redo now act
only on steps made through the bridge, and only while they are newest;
otherwise they fail and leave the history alone. Vector edit mode's
session history is off limits entirely. Settings automation also covers
the design check preferences that landed on master.
* feat(settings): configure tool access and agent step limits
Built-in AI exposed only a hardcoded subset of the tool registry, and the
maximum agent steps was a constant, so users could neither enable
extended tools such as create_component nor adjust long-running tasks.
Built-in AI and the local MCP server now keep independent, locally saved
tool permissions over one shared catalog, with searchable read-only and
side-effect groups and per-target defaults. Chat settings gain a validated
maximum-steps field whose captured value drives the stop condition,
remaining-step warnings, and limit detection for each message.
Tool access, the local server, browser access, and MCP connections are
grouped under a single Automation settings page.
Closes#573Closes#584
* refactor(settings): split automation into MCP and Tool access pages
The Automation page mixed a permission matrix with server endpoints behind
a Tools/Connections switch, and the view switch was indistinguishable from
the provider switch. The nested scroll region showed three of 110 tools.
Rename the MCP-facing page to MCP and give tool permissions their own Tool
access page. The page owns a fixed toolbar for the target, count, defaults,
and search, so the list uses the full dialog body and no row is clipped.
* fix(automation): explain MCP startup failures with localized guidance
Every startup failure collapsed into "MCP server did not become healthy":
the spawn layer recorded the real error but the runtime discarded it, and
health probes could not distinguish a rejected token from a missing server.
The message also surfaced raw English text as the alert heading.
Classify failures by reason (not installed, denied command, early exit,
startup timeout, rejected token, unexpected response, unreachable) and
render translated heading and guidance from the catalog, keeping captured
stderr or HTTP status as labeled diagnostic detail.
* refactor(ui): share one collapsible disclosure primitive
Six features each wired Reka's collapsible with their own motion classes and
one settings-only theme token, so the same interaction drifted in spacing,
icon size, and reduced-motion handling.
Add AppCollapsible with a family theme and move the settings disclosure and
the model editor's advanced settings onto it. Chat and frame-preset call
sites keep their distinct visuals for a follow-up.
* fix(automation): explain MCP failures with localized details
The failure alert carried raw English error text as its heading, and the
diagnostic payload sat in a sibling block outside the alert with no
relationship to it.
Classify failures by reason, render translated heading and guidance from
the catalog, and keep the payload in a collapsible inside the alert, which
unmounts while collapsed so the live region announces only the summary.
Add a copy action for issue reports.
Find the executable where a graphical launch can: extend PATH with the
common global bin directories before the lookup and report the searched
directories as diagnostic detail.
* fix(automation): keep MCP failure details out of reasons already explained
An unreachable address and a rejected token already name their cause in the
translated guidance, so repeating it under Details added noise. Details now
carry only output the summary cannot: stderr, HTTP status, or an unknown
error message.
* test(settings): browse every MCP failure reason in Storybook
The failure copy lived inside the settings panel, so reviewing the eight
reasons meant reproducing each failure and the mapping could only be
checked through the panel's dependencies.
Extract MCPFailureAlert, which owns the reason-to-copy mapping, detail
visibility, copy action, and restart action, and add a story covering
every reason plus the collapsed-details behavior.
* fix(ui): order alert details above the recovery actions
The alert rendered its action buttons before the details slot, so the
collapsible explanation of a failure appeared under the controls it
explains. Details now render directly after the description.
* fix(automation): correct MCP failure classification and detail
Review follow-ups on the failure diagnostics.
Only 401 and 403 mean the server refused our token; any other status now
reports an unexpected response instead of telling the user to replace a
token that was never the problem.
The install hint rendered the whole diagnostic detail as its package
argument, so searched directories appeared inside the install command.
The install target is now a domain constant and the searched directories
stay as detail, which not-installed failures surface again since they are
the actionable desktop diagnostic.
Exited failures also record the process exit code and signal so copied
diagnostics stay conclusive when stderr is empty. The bundled PATH test
now covers the append branch instead of only the unchanged path.
* feat(settings): accept custom values for presets and retention
Retention was a closed set of three counts while the AI step limit was a
free number, so two bounded numeric preferences looked and behaved
differently for no product reason.
Add a shared preset-or-custom field: presets stay one click, the escape
hatch reveals a validated numeric field, and the model carries only the
resolved number. Diagnostics retention becomes a bounded number (50 to
20,000) with the presets as shortcuts, and the hardcoded revalidation in
the panel is replaced by one domain resolver.
* fix(settings): label the preset and custom fields
Replacing the labeled provider field with the shared control left the AI
step limit as a bare select with a detached hint paragraph, outside the
settings group, so nothing on screen said what the number meant. The
accessibility name came from aria-label, which is why behavior tests
passed while the panel was unreadable.
Move both controls into labeled settings rows with their descriptions, and
give the revealed field its own accessible name so the two controls in one
row differ. The specs now assert the control lives inside the row that
names it, which is the check that would have caught this.
* fix(mcp): allow the desktop app origin by default
A server started manually bound the port and answered curl but the app
webview could not use it: no CORS origin was configured, so the browser
blocked every fetch and the app reported the server as unhealthy. The
workaround required an undocumented environment variable.
Allow the desktop app origins by default, accept a comma-separated
override, and document the default in the CLI help and the security notes.
Authenticated requests still need the bearer token, and browsers set Origin
themselves, so only the app webview can present these origins.
* fix(settings): address review findings on the new controls
Copy details awaited nothing and confirmed the copy before the write
finished. VueUse never rejects and falls back to a legacy write, so the
await is what makes the confirmation honest rather than an error branch.
The preset field only left custom mode when a preset arrived; a non-preset
value assigned from the owner left the select showing a value absent from
its options with the field still hidden. The watcher now follows the model
in both directions.
The story play functions queried the revealed field by the row label, which
Testing Library matches as a whole string, so those interactions could not
find it. The Storybook smoke assertion also assumed a button or tab, which
skipped every story built from other primitives.
Share typed MCP, AI, and WebMCP exclusions across adapters. Preserve the browser tool inventory through explicit exclusions and keep execution support and user permissions independent.
Define native Valibot inputs and execution/exposure metadata on each tool. Derive effects and default capabilities, consume upstream Standard Schema conversion, and validate finite numeric inputs consistently across adapters.
Move atomic execution to Core and restore failures from Scene Graph checkpoints without relying on a property diff. Preserve topology, collections, indexes and surviving object identities during rollback.
BREAKING CHANGE: custom tools use input schemas and execution metadata instead of params, ParamDef and independently declared mutation flags. Direct tool execution validates inputs before invoking the handler.
Register inspection and atomic editing tools through document.modelContext with input validation, result bounds, captured targets, cancellation guards, and workspace cleanup. Verify native browser discovery and cross-page undo.
Share canonical tool inputs with AI and browser adapters through Valibot and Standard Schema while retaining MCP numeric coercion and existing transports.
BREAKING CHANGE: programmatic integrations use SDK v2 server/client types; paramToZod is removed in favor of the shared Core tool input contract.
- Hold Vite startup and MCP restarts until the sibling service health endpoint responds\n- Tolerate transient Portless 404 responses during service registration\n- Cover worktree-origin CORS preflight and restart health polling
- Publish explicit effective tool state while retaining disabled tools for Settings
- Classify filesystem writes as side effects and localize category labels
- Stop failed restarts and return precise development control status codes
- Remove duplicated document-access declarations from core tools
- Replace the MCP catalog with typed descriptors, capabilities, and policy
- Emit standard MCP annotations through type-safe registration
- Collect tool metadata through the existing registration wrapper
- Expose the runtime catalog to Settings without a parallel MCP-only list
- Keep disabled tools discoverable so they can be re-enabled
- Declare inspection or modification access on every canonical tool definition
- Add bulk category controls while preserving individual disabled-tool storage
- Keep runtime availability separate from document access semantics
* fix(mcp): close orphaned servers that no app ever claims
- Add ServerOptions.appAttachTimeoutMs: if no app registers within this
window after startup, the server closes itself and removes its
discovery file, instead of squatting the port indefinitely.
- Wire it through the openpencil-mcp-http CLI as
OPENPENCIL_MCP_APP_TIMEOUT_MS (opt-in, unset/0 disables it — a bare
CLI invocation for manual testing should not self-terminate).
- The desktop app opts in with a 30s timeout when it spawns the server.
Without this, a server that outlives its spawning app (renderer crash,
forced reload) keeps holding its port with a stale discovery file. The
app's liveness check only asks whether something answers /health, not
whether an app has ever registered (see /health's no_app status) — so
every later launch finds the orphan already listening and defers to
it, and MCP tool calls fail with "app is not connected" until someone
manually kills the orphaned process. Closing self-caused orphans at
the source means the next launch finds no discovery file and takes
the normal fresh-spawn path.
Fixes#488
* fix(mcp): clean up servers after app disconnects
- Re-arm the orphan watchdog when the registered app disconnects
- Cancel pending shutdown when the app reconnects within the grace period
- Reject timeout values that overflow the runtime timer range
* test(mcp): make watchdog reconnect coverage deterministic
- Wait for the disconnected health state before reconnecting
- Report distinct safe-integer and timer-range validation errors
---------
Co-authored-by: swe-sanad <sanad.arousi@export119.com>
- Rename first-party API, RPC, JSON, CORS, SVG, JSX, and related identifiers to preserve acronym casing
- Keep upstream and serialized boundary names unchanged
- Add a lint guardrail and migration notes for exported APIs
- Resolve MCP workspace subpaths during Vite development
- Remove discovery state before shutdown and close upgraded sockets
- Cover discovery cleanup with a connected WebSocket client
Co-authored-by: Joseph Cumines <joeycumines@gmail.com>
- Prefer private Unix sockets with TCP fallback for local MCP clients
- Unify HTTP and WebSocket lifecycle, authentication, and cleanup
- Discover transport details from an owner-only runtime file
- Add AST-based lint checks for broad unknown object assertions and local JsonObject aliases
- Centralize JsonObject in core and package-local MCP RPC JSON typing
- Replace baseline Record<string, unknown> assertions with named shared/domain types
- Remove stale imports and unused locals from split engine and e2e tests
- Drop the targeted no-unused-vars test allowances
- Keep check and affected test suites warning-free
- Configure oxfmt custom import groups for workspace, app, package, and test aliases
- Keep type imports grouped with their matching source category instead of one global tail group
- Expand the format script to cover formatter config, Vite files, and scripts
- Move top-level engine and e2e prefixed test files under domain folders
- Update fixture path helpers after moving render and pen tests
- Add lint coverage to prevent new top-level prefixed test files
- Refresh testing docs for the new fig and layout paths