Adds an N-tool for the floating card that drops from a "⋯ more"
button or appears on right-click. Emits the OPEN state: vertical
stack of padded icon+label rows in a white card with subtle stroke
and shadow. Positioning and show/hide are caller concerns (same
philosophy as add_modal_shell_v0 / add_toast_v0).
Destructive items (destructive=true) render in red with role
`action-menu-item-destructive` so renderers can style the hover
state separately. divider_before=true on any item (except first,
where it's ignored) inserts a 1px hairline above — useful for
"Edit / Share / Report / Delete" grouping patterns.
Wired through all 3 paths + parity/contract/design-prompt tests.
Handler test covers 7 cases: simple list, destructive red fill,
divider between groups, leading-divider ignored, label-only no
icon, width clamp, bogus parent_id rejection.
Adds an N-tool for the "no data yet" tile that sits in the exact
footprint where a real chart would go. Default 320×200 matches the
line/bar chart default footprint; dashed border + slate-50 fill +
slate icon signal "chart slot, currently empty". Caller can hint
at the widget type via icon ("line-chart" / "pie-chart" / default
"bar-chart-2").
Intentionally separate from add_empty_state_v0 — that tool is for
inbox/onboarding/no-results full-page empties (has optional CTA,
no dashed border). add_empty_chart_v0 reads as "chart widget is
live, just lacks data yet" rather than "nothing to show on this
screen at all".
Wired through all 3 paths + parity/contract/design-prompt tests.
Handler test covers 6 cases: defaults + dashed stroke + icon
override + size clamping + title/subtitle override + bogus
parent_id rejection.
Adds an N-tool for the "variable-N pill-plus-cursor" pattern: a
labeled form control that holds N removable tag pills followed by
an inline placeholder caret. Wrap layout (layoutWrap=wrap) so
chips flow onto additional rows as they accumulate — a horizontal
fit_content row would clip after ~4 chips.
Each chip: pill (cornerRadius=16, slate-100 fill, padding 10/4/6/6
L/R/T/B) + label + 14×14 lucide "x". Default caret placeholder is
"Add tag…" when chips is empty; caller overrides with `placeholder`
(e.g. "Enter emails" for recipient lists).
Wired through all 3 paths with matching handler test (7 cases) +
parity tests auto-picking-up the entry. elements.md: decision tree
§54, trigger list, minimal usage (populated + empty).
Adds an N-tool for one row in a FAQ list. Collapsed (default):
bold question + chevron-right header. Expanded (expanded=true):
chevron-down + multi-line answer paragraph beneath. Optional
show_divider draws a 1px slate-200 rectangle at the bottom for
visual separation between items (no implicit padding — caller
stacks in a vertical parent).
Wired through all 3 paths (pen-core buildFaqItem + pen-mcp handler
+ apps/web shim + Nitro SERVER_BUILDERS); both parity tests and
the stale-integration guard auto-pick-up the entry. Handler test
covers 5 cases: collapsed default, expanded with answer, expanded
without answer (guards against undefined), show_divider hairline,
bogus parent_id rejection.
Adds an N-tool for list/table footer pagination: row of page-
number pills flanked by optional prev/next chevron buttons. Active
page renders filled with the accent color, inactive pages are
ghost. Long ranges collapse with "…" Google-style (always show 1
and total, plus a ±siblings window around current).
Wired through all 3 paths: pen-core buildPagination + pen-mcp
handler + apps/web browser shim + Nitro SERVER_BUILDERS. Both
parity tests (shim-server-parity, element-tool-registry-parity)
pick up the entry automatically. Handler test covers 7 cases:
small range no ellipsis, 10-page ellipsis, start-edge current,
accent override, show_arrows=false, total=1 single pill, bogus
parent_id rejection.
Also updates packages/pen-ai-skills/skills/phases/generation/
elements.md: decision tree §52, keyword triggers (pagination /
page nav / 分页 / 分页条), minimal usage example. The stale-
integration guard (design-prompt-elements.test.ts) now passes.
Dark slate (#334155) box + centered white play icon + optional
caption. Default 320×180 for 16:9. The "future video embed"
affordance — semantically distinct from add_image_placeholder_v0:
dark bg + play icon reads as "play me later", not "picture coming".
Play affordance is a lucide `play` icon_font, NEVER a hand-drawn
path triangle (classic LLM anti-pattern for video placeholders).
Regression test locks that invariant.
Wired across all three paths + elements.md entry (44b) + keyword
map + example + parametric test CASES in 8 files. Handler test
has 6 assertions including the path-vs-icon_font anti-pattern
guard.
Tool count: 55 → 56. Test count: 3225 → 3241 (+16).
- add_metric_comparison_v0: KPI cell with trend. label above + big
value + optional arrow icon + change amount. trend enum (up/down/
flat) drives arrow icon (trending-up/-down/minus) + color (emerald/
red/slate). Distinct from add_metric_row_v0 (scroll row of label+
value cells without trend affordance). Required: label + value.
- add_notification_row_v0: leading icon + (title + optional
timestamp + optional unread red dot) header + optional body
preview. Distinct from add_list_row_v0 which has no timestamp
or unread affordance. Required: title only.
Both wired across all three paths + elements.md ("Analytics / KPIs"
and "Notifications" sections) + keyword map + examples + 2 corpus
prompts in ab-v1/ (dashboard-revenue-trend, mobile-notification-item).
Fixed in same turn: ab-v1 file that was created with the wrong
filename got renamed.
Tool count: 53 → 55. Test count: 3180 → 3225 (+45).
- add_spinner_v0: static loading spinner — full ring (track) + 270°
arc (active). Sits at size=32 default, clamped 16..128. Two
ellipses at SAME origin with DIFFERENT sweep ranges — NOT the
"stacked ellipses for ring" anti-pattern (rewriteLlmAntiPatterns
only fires when both are full-sweep duplicates).
- add_tooltip_v0: small dark pill (#111827) + white text for
hover hints. position param ("top"/"bottom"/"left"/"right")
encodes a role hint (`tooltip-top` etc.) for downstream position
logic; visual body is identical. NO arrow pointer (pen-core has
no clean triangle primitive — caller composes via batch_design
rectangle + rotate if needed).
Wired across all three paths. elements.md adds "Feedback / loading"
section (#48, #49). ab-v1 corpus +1 prompt (mobile-help-tooltip);
spinner omitted from corpus for now — the prompt wording is too
ambiguous for "obvious" difficulty.
Tool count: 51 → 53. Test count: 3138 → 3180 (+42).
Small colored dot + short label: "● Online" / "● Busy" / "● Error"
pattern. Distinguished from the more general add_badge_v0 (just a
pill label) by always having a dot.
tone enum picks dot color:
- success → emerald #10B981
- warning → amber #F59E0B
- error → red #EF4444
- info → blue #3B82F6
- neutral → slate #94A3B8 (default)
Dot uses `frame + cornerRadius=4`, NEVER `ellipse` — an 8×8 ellipse
is the classic "status dot via stacked ellipses" anti-pattern bait.
Keeping it a frame stays clean of rewriteLlmAntiPatterns. Regression
test locked in pen-mcp/add-status-badge-v0.test.ts.
Wired across all three paths + elements.md entry + keyword map +
examples + ab-v1/dashboard-server-status.yaml corpus prompt +
parametric builder test CASES in 7 files.
Tool count: 50 → 51. Test count: 3114 → 3138 (+24).
Fills 3 common UI gaps the 47-tool set didn't cover:
- add_image_placeholder_v0: gray box + centered icon + optional
caption. The "future image slot" affordance. Separate from G()
(which fetches real images). Emits frame+fill, NEVER image node
(empty image renders as broken indicator).
- add_comment_v0: avatar + (author + timestamp) header + body.
Social / UGC / feedback unit. Does NOT handle replies / likes /
action menu — compose via batch_design.
- add_modal_shell_v0: dimmed scrim + centered card (rounded,
shadowed) + title + optional subtitle. "Shell" in the name is
deliberate — this is chrome only; body content goes into the
`modal-shell-card` role via a follow-up insert.
Wired across all three paths (pen-core buildX + pen-mcp handler +
routes + schema + apps/web shim + Nitro SERVER_BUILDERS) + elements.md
decision-tree entries + keyword map + examples + 3 pen-mcp handler
tests (22 cases total) + 3 A/B v1 corpus prompts in ab-v1/ +
parametric builder test CASES auto-extended in 9 files.
Milestones:
- Tool count: 47 → 50
- Test count: 3045 → 3114 (+69)
- A/B v1 corpus: 5 → 8 prompts
Defensive gate: the browser DSL executor clones the doc, mutates
it across all ops, then applies the final doc back in ONE call.
A regression that applied per-op would thrash React + history
state for no benefit (the surrounding startBatch already wraps it
into one undo entry anyway). Spy on `applyExternalDocument` and
assert it's called exactly once per 4-op batch.
Pairs with the #44 commit (browser-safe DSL executor). Every test
stubs fetch to reject so any regression falling back to HTTP fails
loudly.
Coverage:
- Single I() at root → frame inserted, no HTTP
- Binding chain: root + nested child land in correct parent
- U() update applied: properties merged on bound node
- 6-op realistic screen (nav + cards + divider): order preserved
- Multi-op batch → exactly ONE undo entry (dispatcher's
startBatch/endBatch wrap survives the browser path)
- Malformed op in the middle: per-line errors surfaced as
status=failed (not opaque HTTP 500)
- Empty DSL: zero ops, applied + no insertions
- G() without fetcher: image node inserted with empty src (the
apps/web scanAndFillImages pipeline enriches later)
Complements the 5 static browser-safety checks in pen-mcp
(`batch-design-dsl-browser-safe.test.ts`) — that file gates the
import tree, this one gates the runtime behavior.
Total: 3033 → 3041 passing.
Extract ~600 lines of pure DSL logic from pen-mcp/tools/batch-design.ts
into a sibling batch-design-dsl.ts that does not import document-manager
(node:fs) or hooks (server-injected). `handleBatchDesign` becomes a
thin server-side wrapper that opens/saves the .op file around the pure
executor. Backward-compat: pen-mcp barrel + batch-design.ts both
re-export `runBatchDesignDsl` so existing callers keep working.
apps/web dispatcher changes:
- `applyBatchDesignDsl` now runs `runBatchDesignDsl` DIRECTLY in the
browser against useDocumentStore.getState().document (structuredClone
+ apply via applyExternalDocument).
- HTTP `/api/mcp/exec-tool` fallback fires only when the in-browser
executor throws (rare — caller error or future regression).
- Removes per-tag HTTP latency on the common batch_design path.
- ctx.defaultParentId is intentionally NOT applied here: the DSL is
the AI's verbatim instruction set and rewriting `null` parents
would change author intent. Element-tool calls still honor it.
Image search (`G()` op) is swapped from `getSyncUrl`-based absolute
URL (server) to an injectable `ImageSearchFetcher` callback. Server
wrapper keeps the old behavior; browser path omits the fetcher so
`src` stays empty for the apps/web image pipeline (scanAndFillImages)
to enrich later.
Tests:
- New browser-safety gate `batch-design-dsl-browser-safe.test.ts`:
walks the transitive import graph from batch-design-dsl.ts and
fails if any reachable file imports node:fs / node:os / node:path
/ document-manager / hooks. 5 checks including a negative control
on batch-design.ts (wrapper) to confirm the split is meaningful.
- Dispatcher tests updated to match new behavior: happy path does
NOT hit fetch; malformed DSL returns `status=failed` with per-op
error surfaced (via pure executor's `errors[]`), not an opaque
HTTP 500.
- Total: 3028 → 3033 passing.
chart_line: polyline through N data points (normalized to max),
optional dots at each vertex. Emits a `path` node with computed SVG
`d`="M x y L x y …" + N `ellipse` dots. fit_content width = values
× point_spacing.
chart_pie: N colored slices via ellipse `startAngle`/`sweepAngle`
arc support. NOT the "stacked full ellipses" anti-pattern — each
slice has a UNIQUE sweep range (sums to 360°). Supports donut cut-
out via `inner_radius_ratio`. Default 6-color palette rotates.
All-zero input throws (degenerate chart can't be drawn).
Wired across all three paths + elements.md decision-tree entries +
keyword map + examples. Both use layout=none (absolute positioning
for vertices / slices stacked at origin).
Tests:
- pen-mcp: 8 cases chart_line + 8 cases chart_pie (height math,
clamp bounds, custom colors, donut, error paths, id uniqueness)
- Parametric coverage auto-extended (+2 cases each in 9 files)
- Total delta: 2979 → 3027 passing (+48).
Closes#45 + #46; covers part of #50 (charts entries).
Dropdown/picker closed-state display. Same label-above-input shape
as add_form_field_v0 with:
- always-present trailing chevron-down icon
- when `value` is set: black value text + chevron
- when absent: placeholder text styled gray (#94A3B8) + chevron
- justifyContent=space_between pushes the chevron to the right edge
Explicitly NOT modeled: open-menu state (dropdown list). An open
dropdown needs absolute positioning + scrim + per-option states
that belong in a different builder; compose via batch_design for now.
Wired across all three paths + elements.md keyword map + examples.
Tests:
- pen-mcp: 7 new cases (value/placeholder rendering, custom
trailing_icon, required suffix, id uniqueness, parent_id rollback)
- Parametric coverage: auto-extended (+1 case each in 9 files).
Total delta: 2955 → 2979 passing.
Closes#48; covers part of #50 (select entry).
Same label-above-input shape as add_form_field_v0 but the input
grows vertically by `rows` (default 4, clamped 2..12) for notes /
bio / feedback use cases. Input frame layout is vertical with
placeholder top-aligned, matching native iOS/Material behavior.
Wired across all three paths:
- pen-core: `buildTextarea` + TextareaParams in element-builders
- pen-mcp: `handleAddTextareaV0` + tool schema in element-tool-defs-ext
- apps/web: shim + Nitro SERVER_BUILDERS entry
- elements.md: PREFER list + keyword map + 2 example lines
Tests:
- pen-mcp: 8 new cases in add-textarea-v0.test.ts (height math,
rows clamp, required suffix, placeholder wiring, id uniqueness,
parent_id rollback)
- Parametric builder tests auto-extended (+1 case each in 9 files):
layout smoke, post-process idempotency, normalize preservation,
performance, role coverage, anti-patterns clean, detectors
clean, shim-server parity. Total delta: 2907 → 2931 passing.
Closes#47; covers part of #50 (textarea entry).
Simulates realistic "full screen design" orchestrator turns where
one dispatchElementToolCalls call handles 40-60 tools at once.
Coverage:
- 40-tool batch: <500ms end-to-end, 1 undo entry, 40 children, all
ids unique.
- 60-tool batch (upper bound): <800ms, all land, no id collisions.
- 3 × 20 consecutive batches: 60 total children, 3 undo entries,
global id uniqueness preserved across rounds.
- Promise.all of 3 parallel dispatches into 3 distinct roots: each
completes independently, full document id set remains unique.
The mixed pattern (heading + body + list-row + divider + stat-grid)
exercises multi-level tree inserts, not just flat heading stacks —
which is closer to what real AI orchestration emits.
Why the latency budget: AI thinking dominates generation time
(5-30s typical). The dispatch pipeline shouldn't be the bottleneck;
500ms for 40 tools = ~12ms per tool including store round-trip,
which is reasonable. A regression past this threshold signals
quadratic behavior somewhere in the pipeline.
Guards against builders hardcoding icon names that don't resolve at
runtime — a regression would render as an empty glyph or fallback
circle on canvas, silent-but-broken.
Three layers:
1. Per-builder (17 tests): collect every icon_font iconFontName
from default output, assert lookupIconByName resolves each. Fail
message names the specific builder + unresolved slugs.
2. Aggregate (2 tests): full cross-builder icon vocabulary resolves
at runtime; icon set is non-trivial (≥10 distinct icons).
3. Invariants (2 tests): text-only builders emit zero icons;
every icon_font node has iconFontFamily='lucide' (prevents a
regression to a font-family the renderer doesn't bundle).
Note: uses lookupIconByName (not AVAILABLE_LUCIDE_ICONS directly)
because the dictionary has prefix/substring fallbacks that resolve
common names ("home", "more-vertical") even when the literal slug
isn't in the exported list. The lookup is the authoritative runtime
resolver, so matching its behavior is correct.
Bilateral drift guard across three registries + two executable paths:
Registries:
- ELEMENT_SHIMS (client shim, apps/web)
- SUPPORTED_EMBEDDED_ELEMENT_TOOLS (canonical exported list)
- ELEMENT_TOOL_NAMES (pen-mcp source of truth for all 42 tools)
Executable paths:
- A: client shim → pen-core buildX
- B: server /api/mcp/exec-tool → pen-core buildX
Since both paths delegate to the SAME pen-core buildX, the structural
parity is transitive: if shim output matches direct buildX output,
server output matches too.
Tests:
- CASES (42 fixtures) covers every ELEMENT_SHIMS key — refactor adding
a new shim without a test row fails here.
- CASES covers SUPPORTED_EMBEDDED_ELEMENT_TOOLS — same guarantee at
the exported constant.
- SUPPORTED_EMBEDDED_ELEMENT_TOOLS ⊆ ELEMENT_TOOL_NAMES — shim must
only expose tools that pen-mcp actually defines.
- No duplicate keys in ELEMENT_SHIMS.
- For each tool: stripIds(shim(args).node) === stripIds(buildX(args)).
The shim is a pure delegation layer plus id stamping.
- Meta-param extraction: parent_id / pageId / filePath are split out
BEFORE the builder sees them (no spurious field leak into the node).
If any of these diverge in future refactors, the failure points
directly at the broken registry or transformation.
Verifies the user-visible history contract the dispatcher owes:
1. One batch dispatch = exactly ONE undo entry (even for 8 tools).
No Ctrl-Z spam to reverse one AI turn.
2. Separate dispatches = separate undo entries. The batch window
closes at endBatch; subsequent dispatches don't piggyback.
3. Undo reverts the whole batch atomically (all tools snap back).
4. Undo → redo restores the whole batch atomically.
5. All-unsupported batch creates ZERO undo entries — endBatch's
"no changes" short-circuit prevents ghost undos that would jump
the UI between identical states.
6. Mixed valid+unsupported batch → 1 undo entry for the valid ones.
Complements element-tools-dispatcher.test.ts (which spies on
startBatch/endBatch) by exercising the actual history-store round-
trip and asserting stack length deltas.
Full matrix of parse → dispatch → builder preservation for
basic / standard / full tier resolutions. The wire format
(`<op_tool>{...}</op_tool>`) is tier-independent TODAY; this test
anchors that as a hard contract so a future "tier-specific argument
escape" surfaces immediately.
Coverage:
- Tier gating sanity (3 model ids → expected tier)
- ASCII content (3 tiers)
- CJK + mixed scripts + RTL (15 tests: 3 tiers × 5 fixtures)
- Emoji deliberately stripped (anchors applyNoEmojiIconHeuristic
behavior — emojis become icon_font nodes, not embedded text)
- Embedded quotes, backslashes, newlines, tabs, unicode punctuation
(18 tests: 3 tiers × 6 tricky fixtures)
- Large number arrays + nested item objects + timeline shape
(9 tests: 3 tiers × 3 tool shapes)
- Multi-tag batch: 5 tools with varied shapes all parse + match
original args (1 test)
Notable findings:
- gpt-4o-mini resolves to "standard" (matches 'gpt-4o' rule first),
so claude-haiku is the stable basic-tier fixture id
- The pipeline scrubs emojis AND collapses 2+ whitespace chars;
this is intentional (applyNoEmojiIconHeuristic), now anchored
- Unicode punctuation (em-dash, ellipsis, curly quotes) passes
through untouched — the emoji regex is conservative
44 tests (42 builders × clean assertion + 2 sanity/negative anchors).
For every builder output:
- run detectAllIssues (invisible-container + empty-path +
text-explicit-height + sibling-inconsistency detectors)
- filter to severity !== 'info' (info is detect-only, skipped
by the auto-fix pipeline, not a regression signal)
- fail with a per-issue summary if any fire
Plus 2 negative-case sanity checks so passing tests can't mask
broken detectors:
- text with explicit pixel height → height detector fires
- same-fill-as-parent container → invisible-container detector fires
The test lives in apps/web/__tests__ (not pen-ai-skills/__tests__)
because pen-ai-skills doesn't depend on pen-core — apps/web is the
first place both are available.
Walks every branch of the parent_id resolution rule:
payload.parent_id > ctx.defaultParentId > page root
Matrix:
payload.parent_id ∈ {present+valid, present+invalid, absent}
ctx.defaultParentId ∈ {set+valid, set+stale, null, undefined}
Notable cases that happy-path tests miss:
- parent_id exists + defaultParentId set → parent_id wins (default
MUST NOT contaminate when explicit id is valid)
- defaultParentId set to stale id + parent_id absent → fails fast
with "stale" in diagnostic (2026-04-21 regression anchor)
- valid parent_id + stale defaultParentId → still applies (the
dispatcher must not inspect default when explicit is valid)
Plus 3 result-shape assertions: insertedNodes, route, toolName on
applied/failed/unsupported results — orchestrator consumes these
for its progress + inserted-node accounting.
5 integration tests where one raw AI response contains 5-10 op_tool
tags forming a complete screen. Verifies:
- 8-tool login screen applies each tool, one undo batch wraps all
- dashboard: 5 tools land in emitted order (top-nav → stat-grid →
section-header → scroll-row-wrapper → bottom-tab-bar)
- settings: interleaved list-row + divider preserves order
- mixed known/unknown: 2 apply + 1 short-circuit, still one batch
- empty emission: parser returns [], dispatcher reports 'empty'
Complements ai-pipeline-e2e.test.ts (one-tag-at-a-time chain). This
is the multi-tag shape the N-tool orchestrator actually emits for a
full screen — catches ordering / batching / partial-failure
regressions that single-tag tests miss.
44 tests (42 builders × no-op assertion + 2 regression anchors for
activity-ring / progress-bar primitives).
Builders are clean-by-construction templates — they should never
trip an LLM anti-pattern detector. If any detector mutates a
builder tree, this test fails on that row with a visible diff,
pointing directly at either:
- a builder regression (e.g. drifted to stacked ellipses), or
- a false-positive in the detector on valid builder output.
Two regression anchors hard-code the ring rule from auto-memory:
activity-ring and progress-bar must use frame/rectangle, never
stacked ellipses — the same anti-pattern from the 2026-04-07 lesson.
- 128 tests (42 builders × 3 assertions + 2 vocabulary sanity checks):
1. resolveTreeRoles doesn't throw on light theme
2. resolveTreeRoles doesn't throw on dark theme (forced via 7th arg)
3. node count preserved + top-level role survives resolve pass
- Aggregate role set assertion (>= 60 distinct roles) catches mass
stripping if a refactor drops role annotations.
- Covers 85 unique role strings emitted by builders. Unknown roles
are documented pass-through per role-resolver.ts:292, so a typo
wouldn't throw; this test at least anchors the vocabulary in place.
The test is in apps/web because resolveTreeRoles + role-definitions
live there (browser-side post-generation pipeline).
- element-builders-layout.test.ts: 44 tests wrap each of the 42 builder
outputs in a 375x812 frame and run computeLayoutPositions, asserting
no NaN/Infinity coords, every child positioned, widths fit parent
bbox. Proves the real renderer path accepts every builder tree.
- element-builders-composition.test.ts: 3 screens (login / dashboard /
settings) assemble 4-8 builders into a vertical frame, stamp ids,
recurse computeLayoutPositions at every level, and assert expected
role presence. Proves multi-builder assembly survives layout end to
end.
- Drive-by: oxfmt reformat on ai-pipeline-e2e.test.ts imports.
Covers the full embedded orchestrator path as it runs in production:
raw <op_tool> response → tryParseElementToolOutput → dispatcher →
document-store. Nothing mocked between parser and store.
Cases: happy-path element tool lands as text node with correct
content; parent_id targets seeded container; stale parent_id
fails without write; batch_design fallback detected + HTTP
attempted (fetch stubbed to fail); multi-tag response batches
into one undo entry; malformed tag returns null (orchestrator
falls to legacy JSONL); uncovered tool name short-circuits
before HTTP; <think> wrapper stripped (reasoning models);
ctx.defaultParentId applied when payload is rootless.
Pairs with the isolated dispatcher / parser / shim tests — this
one's the "everything wired together" proof. Full repo: 2052
tests across 212 files.
design-parser.ts gains tryParseAllElementToolOutputs(raw) → returns
every `<op_tool>` tag in emit order (element-tool or
batch-design-dsl shape). Single-tag helper stays for orchestrator's
current path; this one's for future prompts that emit multiple
tags per response.
element-tools-dispatcher.ts gains dispatchElementToolCalls(shapes,
ctx) — wraps the whole loop in ONE startBatch/endBatch pair so
N tags collapse to one undo entry. Per-shape results preserve
emit order. Individual shape failure does NOT abort the batch
(matches pen-mcp handleBatchDesign's "collect-errors-keep-going"
philosophy; the AI's later tags may depend on earlier successful
inserts). BatchDispatchResult.status rolls up to applied /
partial / all-failed / empty.
5 new unit tests: empty list skips batch, 3-successful one undo
entry, partial status, all-failed status, result order matches
input order. Full repo: 2043 tests across 211 files.
Two assertions that fire if the embedded shim registry drifts out
of sync with the pen-mcp catalog:
1. Every add_*_v0 name pen-mcp exposes must have a shim —
catches "added a pen-mcp tool, forgot the builder / shim /
Nitro SERVER_BUILDERS update" triple-edit drift.
2. Every shim key must exist in pen-mcp — catches stale shims
for removed or renamed tools.
Failure message names exactly which tools are missing on each
side so the fix is a copy-paste, not a search. Pairs with the
existing short-circuit test — together they lock the invariant
that elements.md catalog = shim set = Nitro registry.
Final 11 builders moved to pen-core: rating_stars, carousel_dots,
link, kbd, price, quote_block, code_block, color_swatch, chart_bars,
timeline, calendar_grid. pen-mcp handlers delegate; shim + Nitro
SERVER_BUILDERS now match the full 42-tool pen-mcp catalog.
With this batch the embedded orchestrator can execute any element
tool the AI emits — no more fallback-to-batch_design routing on
elements.md names that happened to be outside the shim registry.
The "advertised vs executable" asymmetry is closed: elements.md
catalog = pen-mcp handler set = shim set = SERVER_BUILDERS set.
Test suite updated: the "unsupported tool short-circuits before
HTTP" case now uses a fictional name (add_fictional_future_v1)
since every real add_*_v0 is now wired. 1907/1907 pass, zero
pen-mcp handler behavior regressions (builders are byte-identical
to the local tree build they replaced).
Batch B — five controls (switch, checkbox, radio, tabs,
segmented_control) moved to pen-core builders; pen-mcp delegates;
embedded shim + Nitro SERVER_BUILDERS gain direct coverage. 311/311
handler tests still pass unchanged. AI generation under the flag
can now emit common form controls without batch_design fallback.
Batch A — six atomic/single-node tools moved from pen-mcp-local to
pen-core builders so the embedded shim + Nitro SERVER_BUILDERS
actually cover them: divider, badge, avatar, icon_button,
icon_label, stat_grid. pen-mcp handlers delegate; 311/311 handler
tests still pass (zero behavior change). Dispatcher short-circuit
list now names 16 tools instead of 10 — AI generation under the
flag can emit these directly without bouncing through batch_design
fallback.
fetchFn's vi.fn() return type inferred its mock.calls entries as
empty tuples `[]`, so destructuring `[url, init]` tripped TS2493
("no element at index 0/1") and the follow-up `as { body: string }`
cast of possibly-undefined `init` tripped TS2352. Cast the call
tuple through `unknown` to `[string, { body: string } | undefined]`
and guard the body-access with an optional chain — same behavioral
assertions, no untyped any escape.
ELEMENT_TOOL_OUTPUT_FORMAT tells the AI to emit
`<op_tool>{"name":"batch_design", ...}` when no element-tool fits.
Prior Nitro implementation hard-coded a 501 for any DSL payload,
so the FALLBACK branch advertised to the AI was a lie — any AI
that actually took the guidance would see its generation fail.
Fix extracts pen-mcp's `handleBatchDesign` pure executor
(`runBatchDesignDsl`) from the file-I/O wrapper and exposes it on
the package's main barrel. Nitro's `/api/mcp/exec-tool` now
accepts `{dsl}`, runs the executor against a clone of the
sync-state doc (no file I/O, no post-processing hooks — those
belong to the pen-mcp server process), and calls setSyncDocument
to broadcast the result via SSE. Response shape gains
`insertedNodeIds: string[]` so batch inserts (multiple root
bindings in one DSL) surface all their root nodes to the
orchestrator's progress accounting, not just the first.
Client dispatcher updated to prefer the array form with fallback
to the legacy single-id field. Adds a test asserting the
dispatcher actually calls fetch when taking the DSL fallback
(proves the route wires end-to-end). JSDoc in dispatcher +
endpoint updated so code and docs agree.
handleBatchDesign's external behavior is unchanged — it still
opens / post-processes / saves around the refactored executor;
311/311 pen-mcp tests pass unchanged.
elements.md's 42-tool catalog is authored for external MCP clients
that talk to pen-mcp's full handler set via stdio/HTTP. The embedded
orchestrator (this runtime) can only execute tools with BOTH a
client-side shim (element-tool-shims) AND a matching Nitro
SERVER_BUILDERS entry — currently 10 of the 42. Prior code
advertised the full catalog to the AI and promised HTTP fallback
coverage without restriction, so 32/42 tool names would silently
route through to a 404 → surfaced-error path.
Fix:
- Export SUPPORTED_EMBEDDED_ELEMENT_TOOLS from the shim module as
the canonical covered list. Shim + Nitro registries stay in sync
by convention; extending coverage requires updating both.
- Dispatcher short-circuits on tool names not in the list — no
wasted HTTP roundtrip, diagnostic carries the covered-list so the
caller can route to batch_design.
- ELEMENT_TOOL_OUTPUT_FORMAT in orchestrator-sub-agent names the
available subset inline so the AI knows which add_*_v0 it can
emit and when to fall back to batch_design.
- JSDoc in dispatcher, shim module, and exec-tool endpoint updated
to reflect actual behavior (insertStreamingNode path, embedded-
vs-external coverage asymmetry) instead of the stale "HTTP
fallback covers everything" story.
- New test locks the short-circuit: calling an uncovered tool name
(e.g. add_divider_v0) must not attempt fetch.
Real follow-up work is still to extract the remaining ~32 pen-mcp
tool tree-build functions into pen-core, shim them, and extend
SERVER_BUILDERS. Until then, the routes advertised to the AI
actually match what the runtime can execute.
Replaces orchestrator-sub-agent.ts's Phase 1 stub (which only
logged and errored on `<op_tool>` output) with a real apply path:
- element-tools-dispatcher.ts: dispatchElementToolCall(shape, ctx)
runs the full pass inside one startBatch/endBatch pair so a
generation collapses to a single undo entry. Routes element-tool
calls through the shim registry; falls back to /api/mcp/exec-tool
HTTP when shim misses. Validates parent_id existence, rejects
filePath unless it is the "live://canvas" sentinel, rejects
pageId that diverges from the active page, rejects stale
defaultParentId — all with structured failure messages instead
of silent drops with "applied" reports.
- element-tool-shims/: 10-tool registry backed by pen-core
element-builders. wrap<T>() strips parent_id/pageId/filePath
before invoking the builder and surfaces them on ElementShimResult
so the dispatcher can honor them during insert. Same pen-core
builders the server-side pen-mcp handler uses — drift impossible.
- orchestrator-sub-agent.ts: invokes dispatcher with
defaultParentId = subtask.parentFrameId ?? plan.rootFrame.id,
mirroring StreamingDesignRenderer's construction so rootless
payloads land inside the generation's target frame. DispatchResult
carries insertedNodes[] and the orchestrator uses them to update
progressEntry.nodeCount / progress.totalNodes / onApplyPartial
so a successful element-tool subtask does not look like a failure
to upstream accounting.
- Dispatcher uses insertStreamingNode (not raw addNode) so element
tool output goes through the same canonical path the streaming
renderer uses: id collision guard, parent remap, layout-aware
child normalization, phone-placeholder guards, append semantics,
and auto expandRootFrameHeight.
13 tests lock invariants: batch wrap fires exactly once per
dispatch (including back-to-back), applied/failed/unsupported
return shapes, parent_id / pageId / filePath / defaultParentId
validation branches, shim-hit success, HTTP fallback path under
fetch failure, "stale default + valid payload parent_id" payload-wins
precedence.
Live smoke test with VITE_ENABLE_ELEMENT_TOOLS=1 showed `elements`
missing from the sub-agent prompt — the skill was correctly included
by resolveSkills (hasMcpTools flag fired) but then stripped by
compactSubAgentSkills's basic-tier allow-list. Result: the feature
flag was effectively a no-op on basic-tier models, which is exactly
the tier the A/B v1 data says benefits most (MiniMax/GLM +8-21pp ΔM1).
- Add 'elements' to the basic-tier allowed set in
compactSubAgentSkills. The `hasMcpTools` gate at resolveSkills is
still the primary ON/OFF — this just stops the compact step from
silently dropping the skill downstream.
- Deliberately OMIT 'elements' from the reducedComplexity retry-
allowed set. Retries are the last-ditch fallback after a full-
skill attempt already failed; elements.md is ~17k chars and adds
to the prompt budget we're trying to shrink.
Test fix: model-profiles-element-tools.test.ts was passing in
isolation but failing under the full suite. Root cause: vitest's Node
runner `vi.stubEnv` doesn't reach `import.meta.env` across modules
(per-module import.meta instance) and dev's `.env.local` sets
VITE_ENABLE_ELEMENT_TOOLS=1 at Vite transform time. Changes:
- setFlag() now writes to both process.env AND the test file's own
import.meta.env object (belt and braces; doesn't cross modules
but removes the test file's own leakage path).
- Browser-safe (`process` broken) tests changed from
`toBe(false)` to `not.toThrow()`. The actual regression guarded by
these tests is the no-throw contract; cross-module env stubbing is
intractable in the current setup and the boolean path is already
covered by the "flag OFF" suite through process.env stubs.
Full suite: 200/200 files, 1866/1866 tests, format/tsc clean.
Browser-side helper (`isElementToolsFlagEnabled` in model-profiles.ts)
reads via `import.meta.env` as the client fallback, but Vite's default
`envPrefix` only exposes `VITE_`-prefixed variables to the browser
bundle. The previous name `ENABLE_ELEMENT_TOOLS_IN_ORCHESTRATOR`
would be inlined as `undefined` at build time for client code —
meaning flipping the flag in `.env.local` could NEVER actually
enable the feature from the embedded orchestrator, defeating the
Phase 2 rollout plan.
Rename to `VITE_ENABLE_ELEMENT_TOOLS` so client code can actually
see the toggle. Server-side `process.env` reads work with any name,
so one variable name now covers both sides of the SSR boundary.
Docstring in model-profiles.ts now explicitly calls out the VITE_
prefix requirement so future edits don't regress — the "bare name
would be inlined as undefined" point is worth preserving in-file.
Also updated the orchestrator-sub-agent.ts error message that points
users at the flag so its instructions match the real var name.
Tests: 23 → 23 (renamed FLAG constant, all cases still pass). Full
suite 1866/1866 green.
Default-off path crashed in browser bundles before reaching the
false return: Vite doesn't polyfill `process`, and
orchestrator-sub-agent.ts runs client-side, so a bare
`process.env.ENABLE_ELEMENT_TOOLS_IN_ORCHESTRATOR` read raised
`ReferenceError: process is not defined` — defeating the whole
"scaffolding off until explicitly enabled" rollout premise.
Fix:
- Extract `readFlagFromEnv(name)` with three layered safety nets:
1. `typeof process !== 'undefined'` guard around process.env read
2. try/catch around the access itself (Deno/workerd throw on
env inspection rather than returning undefined)
3. Fall through to `import.meta.env` (Vite's canonical browser
env reader) so a dev can toggle the flag via `.env.local`
and hit the same behavior on both sides of SSR
- Any failure path returns undefined → default-off survives
Regression tests (3 new, 23 total):
- simulated browser (`globalThis.process = undefined`) → returns false
- simulated sandbox (`process.env = undefined`) → returns false
- simulated Deno/workerd (getter throws on process.env) → returns false
Full suite 1866/1866 (was 1863; +3 tests).
Implements plan §3.1-§3.5 of the tier-aware embedded-orchestrator
integration behind ENABLE_ELEMENT_TOOLS_IN_ORCHESTRATOR env var. With
the flag unset (default production state) this change is a no-op —
every path added here short-circuits on !needsElementTools(profile).
§3.1 model-profiles.ts:
- needsElementTools(profile) — returns true iff env flag truthy AND
tier in {basic, standard}. Full tier stays OFF per A/B v1 Kimi K2.5
ceiling-effect finding (Δ M1 -12.5pp).
- 20 unit tests cover the 2×3 flag × tier matrix + truthy-value
allow-list parsing.
§3.2 orchestrator-sub-agent.ts:
- Pass hasMcpTools: needsElementTools(modelProfile) into
resolveSkills('generation', ...) so elements.md auto-loads for
gated models, matching the A/B v1 treatment arm.
§3.3 orchestrator-sub-agent.ts:
- When flag fires, append ELEMENT_TOOL_OUTPUT_FORMAT block to the
sub-agent system prompt. Verbatim from
scripts/ab-corpus/build-prompt.ts::T_TOOL_CALL_INSTRUCTIONS so
production reproduces the measured behavior (PRIMARY element-tool
call / FALLBACK batch_design wrapped in op_tool).
§3.4 design-parser.ts:
- tryParseElementToolOutput(raw) wraps pen-ai-skills parseModelOutput
and returns a tagged union {kind:'element-tool'|'batch-design-dsl'}
when <op_tool> is detected, or null to route back through the
legacy extractJsonFromResponse flow.
- 9 unit tests cover happy-path detection, <think> stripping,
multi-tag preference (element tool wins over scaffold batch_design),
legacy passthrough, and malformed-tag graceful fallback.
§3.5 orchestrator-sub-agent.ts:
- STUB: when streaming applied zero nodes AND the completed response
is element-tool-shape, return a clear error pointing at plan §3.5
as the Phase 2 work item. Apply-path dispatch (server-side pen-mcp
handler invocation, live://canvas merge) is deferred to avoid
shipping a path that's untested against the live-canvas sync
machinery.
Tests: 1863/1863 (was 1834; +20 profile tests + 9 parser tests).
Format and tsc clean. No behavior change with flag off.
design.md was stored in a global Zustand store + per-file-key localStorage
in apps/web, and in a module-level cache in pen-mcp. Both leaked across
files: a newly-created document could pick up the previous file's dark
palette (async clearForNewDocument raced with AI chat reads; hydrate()
could rehydrate the last file's designMd on refresh; shared .pen files
lost the spec entirely because it wasn't inside the document).
Fix:
- Add `designMd?: DesignMdSpec` to PenDocument (pen-types). It now
serializes with .pen/.op and travels across sessions/users.
- Add `setDesignMd` action to document-store.
- Rewrite design-md-store as a thin mirror over document-store so the
legacy hook API still works. On document load it migrates any legacy
localStorage entry into the opened document and deletes the localStorage
key; hydrate() wipes the orphan `openpencil-design-md-current-key`.
- MCP handleGetDesignMd / handleSetDesignMd / handleExportDesignMd read
`doc.designMd` directly and persist via saveDocument. Removed the
process-level `_mcpDesignMd` cache.
Verified via MCP live round-trip: set on file A → persists to A's .op on
disk → new file B returns hasDesignMd:false (no leak).
isBadgeOverlayNode matched role:'badge'|'pill'|'tag' and pulled those
children out of their parent's auto-layout, rendering them at (0,0) of
the parent and stacking them on top of siblings. But in this repo
badge/pill/tag are inline-component roles (see role-resolver NAME_EXACT_MAP
and strip-redundant-section-fills PROTECTED_ROLES) — they're meant to
flow in layout like any other child.
Rename to isOverlayNode and narrow to role:'overlay'. Add matching
"Layout-escape roles" guidance in role-definitions.md so generation
prompts can reach the new opt-in. Inline roles now flow correctly;
true floating decorations (notification dots, corner ribbons) still
have a dedicated marker.
snapshot_layout now emits an `overlaps` array listing sibling pairs whose
rendered bounds intersect, so text-only agents can diagnose stacking bugs
without a screenshot. When the shared parent has `layout: "none"` the
reason string points at the real cause (absolute x/y stacking) instead
of letting models hedge with height/padding tweaks. Handler prompt adds
a matching diagnosis workflow so agents fix the parent layout rather
than resizing the overlapping children.
* fix(ai): stop white section bands on dark-themed pages
- role-resolver: skip fixSectionAlternation when parent fill luminance < 0.5, so we no longer paint #FFFFFF/#F8FAFC over a dark root
- strip-redundant-section-fills: add SAFE_LIGHT_HEXES so stale whites from earlier runs (or weak-model hedges) are cleaned up on the sink side
- regression tests for both layers
* feat(ai): design.md-driven background + sidebar color pipeline
- orchestrator-sidebar-color: extract sidebar surface picker; prefer design.md palette role (sidebar/panel/surface) over catalog style-guide legacy cell
- orchestrator-planning: force rootFrame fill from design.md background when a user spec is provided, so sections don't inherit a bright catalog default
- orchestrator-prompt-optimizer: infer design.md background + neutral theme fallback for sub-agent prompts
- orchestrator-sub-agent / ai-prompts: tell sub-agents to leave section root fills unset when design.md drives the palette
- design-md-style-policy: surface-colors policy block keeps MCP and web pipeline aligned
- add planning + prompt-optimizer regression tests
* chore: ignore .omx/ directory
* Enable local OS fonts with vector rendering and proper permission handling (#110)
* docs(readme): update cover screenshot
* fix(renderer): enable local OS fonts with vector rendering and proper permission handling
* test(renderer): refactoring names and creating vi.stubGlobal for the navigator as it's not available in the test environment.
---------
Co-authored-by: Fini <fini.yang@gmail.com>
Co-authored-by: Daniel Chettiar <danielc@snapwork.com>
* feat(types): add AppendContext and SubTask.existingSectionLabels
* feat(ai): add detectAppendIntent for continue/append prompts
* feat(ai): detect append intent before generate_design dispatch
* feat(ai): add applyAppendContextToPlan helper
* feat(ai): reuse existing content-root in append mode
* feat(ai): sub-agent APPEND MODE preamble for existing siblings
* docs(ai): teach horizontal scroll card-row pattern
* chore(ai): enable incremental-add skill in generation phase
* fix(canvas): render synchronously on resize to prevent white flash
Setting canvas.width/height clears the pixel buffer to transparent.
resize() previously only marked dirty, leaving the canvas transparent
until the next RAF and showing the container bg-muted through for one
frame whenever the flex layout shifted (e.g. RightPanel mount on first
selection after idle). Rendering inline after recreateSurface fills
the new surface before the browser paints, closing that window.
* style: apply oxfmt formatting drift across web and renderer files
Non-semantic line-break and wrapping adjustments picked up by oxfmt.
No behavior changes.
* fix(mcp): run codex via shell on Windows to handle .cmd shims
Since Node 18.20/20.12 (CVE-2024-27980) execFileSync refuses to spawn
.cmd/.bat files directly and throws EINVAL. On Windows route through
execSync with shell resolution so PATHEXT picks whichever shim exists
(codex.exe / codex.cmd / codex.ps1).
* feat(editor): anchor paste to selected container or sibling
Pressing Cmd/Ctrl+V now inserts pasted nodes into the selected
container (if it can hold children) or immediately after the selected
node as a sibling, falling back to the root when nothing is selected.
Previously every paste landed at document root, which broke expected
behavior when working inside nested frames.
* docs(ai): expand horizontal scroll card-row example in overflow skill
Flesh out the inline JSON example so the generation-phase skill shows
the full clipContent + nested fit_content row pattern, instead of a
truncated snippet that left model output inconsistent.
* style(lint): clear 7 oxlint warnings from recent commits
- orchestrator-planning.test.ts: narrow fill-array type to
Array<{...}> | undefined and use ?.[0] instead of unchecked [0]
so optional chain does not throw on short-circuit
- mcp-install.ts: drop `?? {}` fallbacks when spreading
config.mcpServers; spread of undefined in an object literal
is a no-op (ES2018+)
* style(lint): clear remaining 15 oxlint warnings across repo
Removes pre-existing warnings not related to any single feature:
- no-useless-fallback-in-spread (6): drop `?? {}` when spreading
possibly-undefined records (document-store-variable-actions,
pen-mcp/tools/{variables,theme-presets}, variable-theme-manager)
- no-useless-spread (2): replace `[...iterable]` with `Array.from`
in for-of snapshots (document-events, agent-indicator), keeping
the re-entry-safe copy intent explicit
- no-control-regex (2): use `\P{ASCII}` unicode property escape
instead of `[^\x00-\x7F]` to express "non-ASCII" without
referencing U+0000 (opencode clients)
- no-new-array (1): `Array.from({ length }, () => '..')` in
document-assets
- no-unused-vars (3): drop unused catch params (agent.ts,
code-generation-pipeline) and unused globSync import
(patch-srvx-bun)
- no-useless-escape (1): `[[{]` instead of `[\[{]` in
chat-message-content regex
---------
Co-authored-by: Fini <fini.yang@gmail.com>
Co-authored-by: Daniel Chettiar <74943095+1MochaChan1@users.noreply.github.com>
Co-authored-by: Daniel Chettiar <danielc@snapwork.com>
* Stabilize synced main for AI handoff, drag nesting, and Electron dev (#104)
* docs(readme): update cover screenshot
* fix: stabilize electron dev sync and codex env passthrough
* Preserve nested frame behavior during drag reparenting
Reparenting across containers used raw local coordinates and root-only clipping assumptions, which made nodes jump visually and caused dragged frames to lose clip/corner semantics after nesting. This adapts the drag-reparent fix to the current upstream store architecture, keeps frame/shape nodes from auto-detaching on canvas drags, and promotes formerly root-only frame clipping to explicit clipContent when nested.
Constraint: Latest upstream workspace checkout is incomplete locally (missing workspaces/deps), so full upstream verification could not be rerun in this environment
Rejected: Keep using raw local x/y during parent changes | fails for auto-layout/padding-rendered positions
Rejected: Make all nested frames clip unconditionally | would change non-clipping containers
Confidence: medium
Scope-risk: moderate
Reversibility: clean
Directive: Preserve visual-position conversion through rendered coordinates when parent changes; local coordinates alone are insufficient once layout participates
Not-tested: Fresh full workspace typecheck/test/build on latest upstream checkout (blocked by missing workspace/dependency setup in this local clone)
* Keep AI codegen requests bounded while exporting asset bundles
The AI codegen pipeline needed two stability fixes: exported design images had to flow through chunk/assembly prompts as reusable asset hints, and oversized chat payloads needed a local guard before hitting provider limits. This commit wires asset extraction into the planning pipeline, threads exported asset paths into prompt assembly, and rejects obviously overlarge chat requests with an actionable client-side error.
Constraint: This branch is split out from a larger local fix stack, so only codegen/prompt/context files are included here
Constraint: Provider request limits are approximate locally, so the payload guard must be conservative rather than exact
Rejected: Inline base64 assets directly into prompts | explodes request size and repeats the same payload per chunk
Rejected: Let provider errors handle oversized payloads | too slow and opaque for users
Confidence: high
Scope-risk: moderate
Reversibility: clean
Directive: Keep asset references flowing as stable ./assets paths and enforce payload limits before fetch to avoid silent request bloat
Tested: bun x tsc -p apps/web/tsconfig.json --noEmit; cd apps/web && bun --bun vitest run src/services/ai/__tests__/context-optimizer.test.ts src/services/ai/__tests__/codegen-assets.test.ts src/services/ai/__tests__/structure-bundle.test.ts; bun run build
Not-tested: Manual end-to-end AI generation with live providers
* Explain sanitized design views instead of leaving AI to guess
The sanitized structure bundle already stabilized asset paths, but it still exposed low-level image/layout/component fields that models had to interpret on their own. This change adds explicit consumer-view enrichment for fills, layout, text, variables, themes, and component semantics, carries original image size through the Figma import path, and augments sanitized bundles with summary/highlight guidance for downstream AI consumers.
Constraint: This branch is intentionally stacked on the asset-bundle PR because it extends the sanitized/codegen asset pipeline rather than replacing it
Constraint: Figma import data is not always complete, so original image size must be preserved when present and inferred only as a fallback downstream
Rejected: Keep sanitized.json as a pure field-level dump | still leaves AI to misread transforms, layout, and component relationships
Rejected: Put all explain text directly in asset extraction helpers | mixes resource stabilization with semantic enrichment responsibilities
Confidence: high
Scope-risk: moderate
Reversibility: clean
Directive: Treat consumer-view enrichment as a distinct layer on top of stable asset extraction; future AI-facing semantics should land there instead of leaking into unrelated pipeline code
Tested: bun x tsc -p apps/web/tsconfig.json --noEmit; cd apps/web && bun --bun vitest run src/services/ai/__tests__/consumer-view-enrichment.test.ts src/services/ai/__tests__/codegen-assets.test.ts src/services/ai/__tests__/structure-bundle.test.ts ../../packages/pen-figma/src/figma-fill-mapper.test.ts; bun run build
Not-tested: Manual prompt-to-code generation quality with live provider responses
* Restore code-panel bundle exports for AI handoff flows
The code generation backend still produced asset manifests and AI structure bundles, but the code panel UI no longer exposed those export paths after later sync work. This commit reconnects the panel to bundle export actions, restores ZIP download behavior when generated code includes exported assets, and locks the affordances with focused panel tests.
Constraint: Other local fixes are still in progress in the working tree, so this commit is intentionally limited to the code-panel export surface
Rejected: Rebuild export support in a separate panel | users expect the export actions to remain where generation results are shown
Confidence: high
Scope-risk: narrow
Reversibility: clean
Directive: Keep code-panel UI aligned with codegen asset/bundle backends whenever generation result shape changes
Tested: cd apps/web && bun --bun vitest run src/components/panels/code-panel.test.tsx src/services/ai/__tests__/codegen-assets.test.ts src/services/ai/__tests__/structure-bundle.test.ts; bun run build
Not-tested: Manual click-through of AI Bundle and Download ZIP in the desktop/web UI
* Unblock electron dev startup in the incomplete local workspace
The local workspace was failing before the app could even start: the skills plugin hard-required js-yaml from a node_modules layout that was not present, Vite dev under Bun hit Nitro NodeResponse incompatibilities, and the web tsconfig was missing path mappings for local packages. This commit removes the unnecessary js-yaml dependency from the skills loader, runs Vite under Node for dev startup, hardens readiness probing with socket checks, and points TypeScript/Vite at the in-repo package sources.
Constraint: The current local clone has incomplete hoisted/workspace installation state, so dev startup must not depend on root package links being perfectly present
Constraint: Bun + Nitro dev currently mis-handle NodeResponse in this environment, so the safest startup path is Node-hosted Vite
Rejected: Keep js-yaml and require everyone to fix local hoisting first | still leaves electron:dev broken in the current environment
Rejected: Continue running Vite dev through Bun | reproduces the NodeResponse/Parse Error failure on /api and /editor requests
Confidence: high
Scope-risk: narrow
Reversibility: clean
Directive: Keep the dev launcher biased toward resilient local startup, even when the workspace install shape is imperfect
Tested: bun -e import('./packages/pen-ai-skills/vite-plugin-skills.ts').then(() => console.log('SKILL_PLUGIN_IMPORT_OK')); bun electron:dev verified Vite ready, MCP/Electron compiled, Electron launched, MCP sync log emitted
Not-tested: Long-running interactive desktop session after startup
* fix(figma): preserve cropped image fill transforms
The synced branch started exporting original image dimensions but dropped the
existing crop transform semantics from the shared image-fill type and both
Figma mappers. That broke the new regression test and stripped metadata that
AI consumer-view/bundle code already relies on.
Constraint: keep app and package Figma mappers in lockstep
Rejected: loosen the new regression test | would hide a real metadata regression
Confidence: high
Scope-risk: narrow
Reversibility: clean
Directive: when extending image fill metadata, update shared pen-types and both Figma mapper copies together
Tested: bun --bun run test (148/149 files passed; only server/__tests__/sse-keepalive.test.ts blocked by missing agent_napi.node), cd apps/web && bun --bun vitest run src/canvas/skia/drag-reparent-policy.test.ts src/components/panels/layer-dnd-utils.test.ts src/stores/document-position-utils.test.ts src/components/panels/code-panel.test.tsx ../../packages/pen-renderer/src/__tests__/document-flattener.test.ts ../../packages/pen-figma/src/figma-fill-mapper.test.ts, cd apps/web && bun --bun vitest run src/services/ai/__tests__/codegen-assets.test.ts src/services/ai/__tests__/structure-bundle.test.ts src/services/ai/__tests__/consumer-view-enrichment.test.ts, cd apps/web && bun --bun vitest run src/utils/__tests__/security.test.ts, bun test scripts/loopback-no-proxy.test.ts, npx tsc --noEmit, bun --bun run build
Not-tested: server/__tests__/sse-keepalive.test.ts without a locally built @zseven-w/agent-native addon
* docs(editor): normalize new PR comments to English
The PR had a handful of newly introduced Chinese code comments in dev, sync, and AI helper paths. This follow-up keeps the implementation unchanged while translating those comments to English so the PR stays consistent with the repository comment-language expectation.
Constraint: The request was limited to comment language cleanup after the conflict-resolution merge, so behavior had to remain unchanged
Rejected: Leave the mixed-language comments in place | conflicts with the PR requirement for English comments
Rejected: Broader repository-wide translation sweep | unnecessary scope expansion beyond the PR-introduced comments
Confidence: high
Scope-risk: narrow
Reversibility: clean
Directive: Keep code comments in English on this branch, even when local notes or working memory are in another language
Tested: bun test scripts/loopback-no-proxy.test.ts apps/desktop/__tests__/dev-utils.test.ts; cd apps/web && bun --bun vitest run server/__tests__/mcp-sync-state-active.test.ts src/canvas/skia/__tests__/skia-interaction.test.ts; npx tsc --noEmit; branch-diff comment scan for Han characters in comment lines
Not-tested: Manual runtime behavior, since this change only rewrote comments
* style(editor): apply repository formatting expected by CI
The PR was failing the CI Format check after the conflict-resolution and comment-normalization follow-ups. This commit applies the repository formatter output to the files touched by the branch so CI sees the exact formatting it expects, without changing behavior.
Constraint: The failing GitHub Actions job stopped at Format check, so the fix had to match oxfmt output rather than introduce functional changes
Rejected: Leave the branch as-is and rely on local formatting differences being acceptable | CI explicitly rejects the current formatting
Rejected: Broader code cleanup beyond formatter output | unnecessary scope while repairing the failing check
Confidence: high
Scope-risk: narrow
Reversibility: clean
Directive: After conflict resolution or comment-only edits on this repo, run bun run format:check before pushing because formatter expectations are stricter than the existing file style in some touched files
Tested: bun run format:check; bun run lint; npx tsc --noEmit
Not-tested: Full test suite after this formatting-only commit (previous run showed formatting was the first CI blocker)
* refactor(editor): remove proxy-specific dev workarounds from PR
The PR no longer needs the loopback proxy bypass layer, so this cleanup removes the proxy-specific dev entrypoint, environment bootstrap, helper module, and its tests while keeping the unrelated Electron and AI handoff changes intact.
Constraint: Removal had to be limited to proxy-related code on PR #104 without undoing the other merged fixes on the branch
Rejected: Keep the helper and stop using it | leaves proxy-specific maintenance surface and tests in the PR
Rejected: Revert the entire Electron dev file to upstream earlier than necessary | would risk dropping unrelated local conflict-resolution choices beyond the proxy scope
Confidence: high
Scope-risk: moderate
Reversibility: clean
Directive: If proxy handling is reintroduced later, keep it out of this PR unless there is a dedicated, separately justified change for it
Tested: bun run format:check; bun run lint; npx tsc --noEmit
Not-tested: Manual electron:dev behavior after removing the proxy-specific launcher path
* docs(ai): translate JSON-facing semantic descriptions to English
The PR still emitted Chinese semantic description strings inside the AI consumer-view and structure-bundle JSON outputs. This change translates those JSON-facing runtime descriptions and updates the affected tests so exported AI-facing structure data is consistently English.
Constraint: The request was limited to JSON description strings, so the change had to preserve the same semantics and structure while only translating output text
Rejected: Leave Chinese test fixtures and runtime descriptions in place | conflicts with the requirement for English JSON descriptions
Rejected: Broader i18n cleanup outside these AI JSON description paths | unnecessary scope expansion beyond the requested exported-description surface
Confidence: high
Scope-risk: moderate
Reversibility: clean
Directive: Keep AI/exported JSON explanation strings in English unless a future change explicitly adds localized output modes
Tested: cd apps/web && bun --bun vitest run src/services/ai/__tests__/consumer-view-enrichment.test.ts src/services/ai/__tests__/structure-bundle.test.ts src/services/ai/__tests__/codegen-assets.test.ts; bun run format:check; npx tsc --noEmit
Not-tested: Full app runtime flows that consume these JSON descriptions outside the covered unit tests
* refactor(ai): remove remaining network-proxy handling
The current project still carried Anthropic proxy-specific heuristics and environment handling outside the PR-specific cleanup. Since the earlier crashes and connectivity issues were unrelated to proxying, this removes the remaining network-proxy branches, model remapping, and TLS override advice while leaving unrelated request flows intact.
Constraint: The cleanup needed to remove proxy-specific logic without disturbing unrelated transport concepts such as app-internal API proxy routes or React proxy objects used in tests
Rejected: Keep the proxy heuristics as dormant fallback logic | preserves misleading operational guidance and dead maintenance surface
Rejected: Rename every remaining literal use of the word proxy in the repo | would overreach into unrelated concepts like internal API proxying and JS Proxy-based test setup
Confidence: medium
Scope-risk: moderate
Reversibility: clean
Directive: If endpoint-specific compatibility logic is needed later, add it as explicit endpoint handling rather than generic proxy heuristics
Tested: bun run format:check; bun run lint; npx tsc --noEmit; repo-wide search for network-proxy env references after cleanup
Not-tested: End-to-end Claude connection flows against custom base URLs after removing proxy-specific remapping
* fix(electron): keep Node-backed dev launch for Nitro compatibility
Comparing against upstream commit 7271a03 confirms the current Electron dev fix is not the same idea as the original Bun-based launcher. The upstream version starts Vite with Bun, while the observed failure shows Nitro now crashes in that path with "Vite environment nitro is unavailable". This keeps the non-proxy Node-backed launcher because it fixes the actual regression without restoring the removed proxy code.
Constraint: The request preferred reverting to the upstream original only if the intent matched, but the current Nitro/Electron failure proves the upstream Bun launcher is no longer equivalent in behavior
Rejected: Restore the exact 7271a03 Bun launcher | reproduces the Nitro dev-worker crash and ERR_EMPTY_RESPONSE in Electron
Rejected: Reintroduce the old proxy workaround bundle | unrelated to the reproduced failure and already removed by request
Confidence: high
Scope-risk: narrow
Reversibility: clean
Directive: Keep Electron dev on the Node-backed Vite launcher unless Nitro/Bun dev compatibility is revalidated with a real startup test
Tested: bun run electron:dev (reached Electron launch after Vite/MCP/Electron compile steps); bun test apps/desktop/__tests__/dev-utils.test.ts; bun run format:check; npx tsc --noEmit
Not-tested: Full interactive manual editor workflow after Electron launch
---------
Co-authored-by: Fini <fini.yang@gmail.com>
* fix(ai,cli): openai-compat turn-2, StepFun reasoning+451, Mac CLI discovery
Round up the v0.7.2 stability fixes for AI connectivity and local CLI
detection that surfaced during real user runs against GLM, StepFun, and
Mac users on nvm/fnm/pnpm/bun/mise/asdf/fish shells.
Provider (via @zseven-w/agent-native v0.3.0 submodule bump):
- OpenAI-compat providers can now complete multi-turn tool-calling loops:
the request builder translates Anthropic-shaped message history
(tool_use / tool_result blocks, thinking) into OpenAI's tool_calls +
role="tool" form so turn 2 no longer 400s. system_prompt is finally
injected instead of being silently dropped.
- The SSE parser accepts `delta.reasoning` (StepFun step_plan) alongside
`reasoning_content` (GLM / DeepSeek / Qwen), and also streams tool_call
fragments, which unblocks GLM / dashscope and stops the
firstTextTimeout → fetch abort → std.http panic → Bun segfault cascade.
- HTTP 451 (StepFun content-safety) surfaces as InvalidRequest with a
specific "content blocked by provider safety filter" message instead
of an opaque error_server.
Server route + client watchdog:
- /api/ai/chat forwards the provider's last_error string
(result.errors[0]) so users see "HTTP 451 content blocked" rather than
"Provider error: error_server".
- streamChat clears firstTextTimeout on thinking chunks (when
thinkingResetsTimeout=true), so models that stream long reasoning
before any text aren't falsely killed as "stuck".
Orchestrator sub-agent resilience:
- Failed sub-agents (empty response / unparseable output) now retry once
with a minimal ~3KB kernel prompt (schema + jsonl-format only). Only
the failing subtask re-runs — successful earlier sections are kept.
- Deterministic refusals (HTTP 400/401/429/451, "content blocked",
"censorship", "authentication failed") short-circuit the retry ladder
so a 4-minute StepFun safety scan isn't spent twice in a row.
Local CLI discovery (Mac users on managed shells):
- New server/utils/cli-resolver-helpers.ts exports probeViaLoginShell()
and posixUserBinDirs(). Login-shell probe asks $SHELL (or zsh/bash
fallback — fish added at /opt/homebrew/bin/fish and friends) with
`-ilc 'command -v <cli>'` so nvm/pnpm/bun/mise/asdf/volta/fnm shims
are visible even when Electron scrubs the inherited PATH.
- resolveClaudeCli / resolveGeminiCli / resolveCopilotCli and the
inline codex/opencode resolvers in connect-agent.ts all run the same
PATH → login-shell → npm-prefix → user-bin candidates ladder. Each
step logs via serverLog to ~/.openpencil/logs/server-YYYY-MM-DD.log
for remote diagnosis.
Builtin provider preset:
- Add StepFun Coding Plan (api.stepfun.com/step_plan/v1, label "StepFun
Coding Plan") alongside the existing StepFun preset.
Version bump 0.7.1 → 0.7.2 across all workspaces.
---------
Co-authored-by: RaisCui <857943+raiscui@users.noreply.github.com>
Co-authored-by: Fini <fini.yang@gmail.com>