Long cell content rode `width=auto` and the cell itself had
`height=fit_content` with no clip, so a string longer than its
allotted column would push past the cell edge into the adjacent
column at render time. Switch each cell to `width/height=
fill_container + clipContent=true`, give the text child
`width=fill_container + textGrowth=fixed-width` so the layout
engine wraps inside the cell, and clip the row itself so a runaway
cell can't push siblings off-row either. Lock the contract in the
handler test.
Greps every builder source for paddingTop/Right/Bottom/Left as a
property key and fails if any reappear. The unified padding field is
the only form resolvePadding reads; the CSS-side siblings render with
zero inset, which slipped past type checks because ElementTree is
Record<string, unknown>. JSDoc / line-comment mentions stay allowed
so builders can still document the trap inline.
origin's v0.8.0 had cherry-picks of the v0.7.5 deepseek/image-search
fixes (a727632a, a5952bc8) overlapping local 2073cf5b / 04f4fbc1, plus
new commits (model-selector ark-coding deepseek-v4-pro/flash IDs that
ARK rejects, fetch error.cause unwrap, CI agent-native build, op
export docs cleanup, main merge). Resolved the ark-coding list in
favor of HEAD's deepseek-v3.2 entry (only model ARK Coding Plan
actually supports — see openpencil-docs note).
The `op export` command was removed in 0.7.x but the README still
advertised it (#116). The pen-mcp README also documented an
`npx @zseven-w/pen-mcp` quick-start that never worked because the
package ships TypeScript source against workspace-only deps with no
`bin` entry (#117).
- Strip `op export` references from all 15 root and 15 cli READMEs
- Sync AGENTS.md, CLAUDE.md, apps/cli/CLAUDE.md to match the codegen-
pipeline reality (no standalone export command anymore)
- Rewrite pen-mcp README's quick-start: explain the package ships as
part of the OpenPencil app and external clients connect over HTTP
Closes#116Closes#117
`/models` now returns only deepseek-v4-pro and deepseek-v4-flash;
deepseek-chat / deepseek-reasoner sunset 2026-07-24 and the
deepseek-v3.2 hard-coded in the ark-coding fallback list never
existed. Both v4 models default to thinking enabled and the API
toggles via `{"thinking":{"type":"disabled"}}` — keep
`thinkingMode: 'disabled'` so the app's fast/non-thinking default
stays intact (server reasoning paths honor it; the Zig openai-compat
path doesn't emit the toggle yet, so calls through that path still
get provider-default thinking until it's wired). v4-pro promoted to
full tier; legacy aliases pinned to an exact RegExp so future
deepseek-* variants don't inherit a forced disabled mode.
Bandaid for the unwired toggle: v4-pro gets `timeoutMultiplier: 2`
because its default-on reasoning blows past the orchestrator's
planning timeout on long system prompts (observed in dev: planning
phase falls back, sub-agent then succeeds — UX degraded but
functional). Drop the multiplier once the Zig path actually sends
`thinking:{type:disabled}`.
Don't add a BUILTIN_MODEL_LISTS.deepseek entry — DeepSeek exposes
/v1/models, so let `fetchProviderModels` pull the live catalog
through `/api/ai/provider-models` instead of pinning a snapshot
(the ark-coding `deepseek-v3.2` ghost above shows what those
snapshots drift into).
paddingTop/Right/Bottom/Left are silently dropped by resolvePadding in
the layout engine, causing all four sides to render as 0. Switch every
affected builder to the canonical padding: [T,R,B,L] form.
pen-core's resolvePadding only reads the array/scalar `padding`
field; the CSS-style paddingTop/Right/Bottom/Left siblings I shipped
in 5c275ba1 were silently dropped, so the rail rendered flush. Switch
to padding=[16,12] / [8,12,24,12] / [0,12] and add a regression
assertion that catches the same trap on any future tweak.
Adds the obvious-difficulty prompt covering the 77th element tool;
bumps the corpus-loader test from 26 to 27 entries and lists the new
2026-04-27 batch in the README so future corpus drift is caught at
test time.
Top_nav_bar / bottom_nav / nav_chip_row covered mobile and inline
chrome but desktop dashboards still had to hand-roll their left
sidebar via batch_design. Adds a 240px-wide vertical rail with
icon+label rows, optional brand title, and slate-100 pill bg on the
active item — distinct surface from the existing nav tools so the
decision tree picks it cleanly for "sidebar / side nav / 侧边栏".
Sticky bottom-of-page disclosure card with title, body, accept /
decline buttons (decline-then-accept order), and an optional
"Cookie settings" link for fine-grained consent. Caller positions
the banner; the tool emits the card itself with shadow + 1px slate
border for visual lift above the page content. Plus the v1 corpus
prompt landing-cookie-banner (26 v1 prompts total).
The "Subscribe to newsletter" / "Apply discount code" / "Send chat
message" pattern. Two action variants:
- text (default): pill button with label like "Subscribe"
- icon: 44×44 square icon button (chat send arrow / search apply)
Distinct from add_form_field_v0 (label-above, no inline button)
and add_search_bar_v0 (no trailing action button). Optional
leading_icon adds an icon inside the input itself. Plus the v1
corpus prompt landing-newsletter-signup (25 v1 prompts total).
The "+1 (555) …" pattern from every modern signup / login screen.
A 44px row with leading country selector (flag + dial code +
chevron-down), a 1px slate divider, and the digits input on the
right. Country selector is a button-shape (no actual dropdown
menu); caller handles picker UX as a separate concern. `value`
toggles between placeholder (slate-400) and populated (slate-900).
Third theme-aware v1 (after add_modal_shell_v1 and add_toast_v1).
Same dashed-border "no data yet" tile shape as v0 with a `theme`
param that swaps 5 colors (bg / border / icon / title / subtitle):
- light (default): byte-parity with v0 (slate-50 / slate-300)
- dark: slate-800 bg + slate-600 border + slate-200 title
- system: $color-surface-2 / $color-border / $color-text-primary
refs — requires applySemanticPalette(doc) seeded
Lets dark-theme dashboards keep a matching empty-slot surface
instead of punching a light rectangle out of dark cards.
Visual static representation of a horizontal slider. Track splits
into left fill (accent) + 20×20 thumb (white + accent stroke) +
right remaining (slate), all aligned on a 20px track wrap. Optional
label + value readout row above; `value_suffix` renders "60%" /
"128px" / "0°" style readouts. Pixel math: fill = (width-20) * pct,
auto-collapses fill or remaining at either extreme.
Second theme-aware v1 (after add_modal_shell_v1). Same pill shape as
v0 with a `theme` param:
- light (default): byte-parity with v0 (dark #111827 pill + white fg)
- dark: INVERTED contrast — light pill (#F1F5F9) + dark fg (#0F172A)
- system: $color-text-primary bg + $color-surface fg (inverted swap)
Unlike surface-like v1s (modal-shell), toasts use inverted contrast by
design — a dark pill on light bg, a light pill on dark bg — so the
dark variant flips the pill rather than darkening it.
The "Pro $29/month" column for pricing tables. Tier name + big
price (currency + amount + period) + check-mark feature list + CTA
button. Two emphases: default (slate border + slate CTA) and
featured (accent border + accent CTA + auto "Most popular" badge
unless overridden by explicit `badge` value).
Two independent changes rolled together since they both serve the
same goal — "can real LLMs actually route to the tools we shipped
today?":
1. scripts/ab-corpus/run.ts gains a `--corpus` flag (ab-v0 |
ab-v1, default ab-v0 for back-compat). The harness was
hardcoded to ab-v0 — adding ab-v1 prompts was worthless
without a way to run them. Validated 17 prompts × 2 models ×
2 variants during a live 方舟-CP run.
2. 5 new ab-v1 prompts cover the 2026-04-24 tool batch:
- mobile-upload-dropzone.yaml → add_upload_dropzone_v0
- mobile-otp-verification.yaml → add_otp_input_v0
- mobile-file-attachment.yaml → add_attachment_row_v0
- mobile-chat-message.yaml → add_chat_bubble_v0
- dashboard-dark-modal.yaml → add_modal_shell_v1
corpus-loader.test.ts bumps its count assertion 12 → 17 and
extends the tool-coverage set. Also loosens the regex to accept
`_v\d+$` (was `_v0$`) so add_modal_shell_v1 passes. No other
test file needed changes — the existing registry-parity and
mock-llm tests already use `_v\d+$` or the registry directly.
.gitignore gains:
- .playwright-mcp/ (MCP Playwright session artifacts)
- editor-*.png (local verification screenshots)
- scripts/ab-corpus/runs/ (live-run outputs / reports)
None of those belong in version control — they're artifacts
from local verification runs.
Extends element-tools-composition.test.ts with a 7th real-screen
scenario that exercises every tool added after the 62-tool mark:
- add_top_nav_bar_v0 (existing anchor)
- add_chat_bubble_v0 × 2 (left from-others + right from-self)
- add_attachment_row_v0 (file on the self-message path)
- add_upload_dropzone_v0 (drop area for screenshots)
- add_chip_input_v0 (conversation tags)
- add_action_menu_v0 (floating menu, open state)
- add_modal_shell_v1 (theme="dark" confirm dialog)
- add_otp_input_v0 (phone verification step)
8 tool calls, chained through the full MCP handler pipeline
(ensureParentExists → builder → assignIdsRecursively →
batch_design insert with rollback-on-failure → post-insert
landing check → save → re-read from disk). Asserts:
- every call emits a nodeId (no silent no-ops)
- final doc has exactly 9 root children
- every tool's role marker survives round-trip save/load
- modal v1 theme=dark → card fill #1E293B (not v0's #FFFFFF)
- chat right-side bubble surface → #2563EB accent
- OTP focused slot stroke → #2563EB accent
This is end-to-end at the DATA layer, not the visual layer.
Still not covered by any test:
- Skia rendering (needs debug_screenshot against a live
canvas)
- Real LLM tool routing (needs ab-corpus harness with live
API keys — route to 方舟 CP works but has not run since
the routing fix)
Document-layer coverage is enough to rule out handler-pipeline
bugs (bad parent_id threading / cache-stale doc-state / silent
rollback swallowing); visual regressions require the next gate.
Codex stop-hook caught element-tool-defs-ext-2.ts at 832 lines
after add_chat_bubble_v0 landed there — 32 over the repo's 800-
line ceiling. Same trap the original single ext file hit at 1329
lines, same fix pattern: carve the second half into a new shard.
Split at `add_modal_shell_v1` (tool #16 of 24 in old ext-2):
- ext-2 keeps tools 1-15 (calendar_grid through textarea) →
482 lines
- ext-3 (new) holds tools 16-24 (modal_shell_v1 through
chat_bubble) → 371 lines
All registry shards now:
base 683
ext 671
ext-2 482
ext-3 371
(props 26)
(top 254)
element-tool-defs.ts concatenates all three ext shards into the
single ELEMENT_TOOL_DEFINITIONS — external API unchanged.
ELEMENT_TOOL_DEFINITIONS_EXT_3 is the new import; 67 tools
still resolve.
Header in ext-1 updated to reflect the three-way split + advise
"when ANY shard crosses 700, carve a ~5-tool chunk into the
shortest shard" so the next rebalance happens proactively instead
of after a stop-hook trip.
Chat / messaging / customer-support UI message unit. Two variants
via the `side` enum:
- side="left" (default): from-others bubble. Slate-100 fill,
slate-900 text, alignItems=flex-start. Optional `author` text
shown above the bubble (group-chat pattern).
- side="right": from-self bubble. Accent-color fill (customizable
via `accent_color`), white text, alignItems=flex-end. Author
intentionally suppressed on this side — a self-bubble never
carries "You:".
Optional `timestamp` below the bubble on either side.
Max-width mechanic: pen-core has no native max-width primitive, so
`max_width` becomes the bubble's fixed width (clamped 160..480).
Short messages get extra padding on one side — matches every real
chat client (iMessage / WhatsApp / Slack). Message text uses
`textGrowth: 'fixed-width'` + `width: 'fill_container'` to wrap
correctly inside the fixed-width surface.
Full wiring: pen-core builder + pen-mcp handler + schema (into
ext-2, shorter shard — 24 tools vs ext-1's 24 after this) + shim
+ SERVER_BUILDERS + parity CASES + contract allow-list +
elements.md decision tree + triggers + minimal usage for both
sides. Handler test covers 9 cases: registration, left defaults,
left+author, right with self-dropped-author, right+accent_color,
timestamp both sides, max_width clamps (low + high split into
separate tests to avoid cache-interference), textGrowth wiring,
bogus parent_id rejection.
Fills another common UI gap: the "here's an already-uploaded file"
row you see in email composers, chat attachments, and form upload
summaries. Compact horizontal layout: type-icon + filename (bold) +
optional muted size string + optional right-side × remove affordance.
Structure: horizontal frame (slate-50 bg, cornerRadius=8) with
three children:
1. attachment-icon — lucide file-* (caller picks: file / file-
text / file-image / file-video / file-audio / file-archive /
file-spreadsheet / file-code)
2. attachment-meta — vertical frame with filename + optional size
3. attachment-remove — × icon, suppressed via removable=false
Intentionally NOT embedding an upload-progress variant in v0. The
pen-core schema lacks percentage-width primitives, so a %-filled
progress bar would either need a fixed track width (brittle across
parents) or a caller-computed pixel value (awkward API). Callers
who need the uploading state compose `add_progress_bar_v0` directly
below the row — cleaner separation.
Wired through all standard points: schema into ext-1 (balanced
shards 23/23 after upload-dropzone landed there last commit) +
shim + SERVER_BUILDERS + parity CASES + contract allow-list +
elements.md decision tree + triggers + minimal usage.
Handler test covers 7 cases: registration + minimal (no size) +
size rendered + custom icon + removable=false + default icon +
bogus parent_id rejection.
Fills the auth-flow gap: 2FA / PIN / phone-verification codes.
Horizontal row of N square slots (4..8), one digit per slot.
Renders three states per caller intent:
- blank (no `digits`): all slots empty, `focused_index` marks
the currently-typing slot with an accent-color 2px outline
- partial: first M slots filled with digit text, slot M+1
focused, rest empty
- full: all N slots filled (final submittable state)
Filled slots get role=otp-slot-filled + slate-700 border + 20/600
digit text. Focused empty slot gets role=otp-slot-focused +
2px accent border. Blank unfocused slots get role=otp-slot +
1px slate-300 border.
Wired through all standard points — schema into ext-2 (shorter
shard) + shim + SERVER_BUILDERS + parity CASES + contract
allow-list + elements.md decision tree + triggers + minimal
usage for each state.
Handler test covers 8 cases: registration + defaults (6 blank
focused-first) + partial state / full state / length clamp low
(< 4 → 4) + length clamp high (> 8 → 8) + accent color override
+ bogus parent_id rejection.
Fills a real gap in the element-tool family: upload / drag-and-
drop surfaces. Dashed border + cloud icon + two-line instruction
("Drop files to upload" / "or click to browse") — the classic
pattern from every modern file-upload UI.
Deliberately structurally similar to add_empty_chart_v0 (dashed
border + icon + title/subtitle) but semantically distinct:
- empty_chart = "chart widget will render when data arrives"
(320×200, icon chart-typed)
- upload_dropzone = "users drop files here" (480×200, icon
semantic: upload-cloud / upload / file-up)
elements.md routes them by intent, and the tool descriptions
cross-reference each other to prevent the AI from picking the
wrong one on ambiguous prompts.
Wired through all the standard points per the add-new-tool
checklist: builder + handler + schema (into ext-1, the shorter
shard) + shim + SERVER_BUILDERS + parity CASES + contract
allow-list + elements.md decision tree + keyword triggers +
minimal usage. Handler test covers 5 cases: defaults, dashed
stroke, overrides, size clamping, bogus parent_id rejection.
[Codex P3] element-tool-defs-ext.ts had grown to 1329 lines —
over the repo's documented 800-line ceiling. Ironically the file
header comment claimed it existed to keep its parent under the
cap, but the shard itself had outgrown the limit.
Split into three files:
- element-tool-def-props.ts (26 lines) — shared JSON-Schema
fragments (schemaVersionProp / filePathProp / parentIdProp /
pageIdProp) that every definition file uses. Deduplicating
these unblocks the split cleanly.
- element-tool-defs-ext.ts (596 lines) — first 22 tools
(add_switch_v0 through add_segmented_control_v0 era). Imports
the shared props.
- element-tool-defs-ext-2.ts (744 lines, new) — remaining 22
tools starting at add_calendar_grid_v0. Imports the shared
props.
element-tool-defs.ts concatenates all three arrays into the single
ELEMENT_TOOL_DEFINITIONS registry — external API surface unchanged.
Header comments in both shards now document the split convention:
"pick whichever shard has fewer tools" when adding a new entry,
to keep the files balanced as the family grows toward ~100.
Incidentally the previous commit's file also carried the P2 fix
(pageId threading through the in-browser and HTTP DSL paths, so
multi-page docs land the generation on the ACTIVE page instead of
doc.pages[0]). Both touched the same file, didn't make sense to
split. Title-wise the previous commit is P1 but functionally it's
P1+P2.
All under the 800-line ceiling now:
element-tool-defs-base.ts 683
element-tool-defs-ext.ts 596
element-tool-defs-ext-2.ts 744
element-tool-defs.ts 240
element-tool-def-props.ts 26
[Codex P1] The browser-side element-tools-dispatcher imported
runBatchDesignDsl from the \`@zseven-w/pen-mcp\` package barrel.
That barrel re-exports node-only modules — document-manager,
log-utils, theme-presets — which import node:fs / node:path at
top level. Vite / esbuild resolve the barrel BEFORE tree-shaking
can drop those branches, so browser builds failed on unresolved
node built-ins.
Fix:
- packages/pen-mcp/package.json: add \`./dsl\` subpath export
pointing at tools/batch-design-dsl.ts — the pure executor
file already guarded as browser-safe by the adjacent
regression test.
- apps/web dispatcher: switch import to
\`@zseven-w/pen-mcp/dsl\`. No other changes — the re-exported
symbols (runBatchDesignDsl / OpResult / ImageSearchFetcher /
RunBatchDesignDslOptions) are identical shape.
- batch-design-dsl-browser-safe.test.ts: add an assertion that
package.json's exports field preserves the \`./dsl\` key
pointing at the expected file. Without this, silently
removing the subpath would re-introduce the browser-breaking
resolution path.
The package barrel keeps its current export of runBatchDesignDsl
too (a few internal test files still import from it). Browser
callers should migrate to \`@zseven-w/pen-mcp/dsl\` per the JSDoc
note now in the dispatcher.
Wires the buildModalShellV1 pen-core builder through the full MCP
toolchain — handler + schema + dispatch + shim + SERVER_BUILDERS
+ parity CASES + handler tests + elements.md skill — so external
MCP clients (Claude Code / Codex / Gemini CLI) can call
add_modal_shell_v1 as a first-class tool alongside the 62 v0
tools.
Three theme variants exposed via `theme` param (enum [light, dark,
system]):
- omitted / `'light'`: byte-parity with add_modal_shell_v0
(same hex, same structure, same role tree)
- `'dark'`: hardcoded dark palette (#1E293B card, #F1F5F9 title,
#94A3B8 muted). No \$refs needed.
- `'system'`: emits \$color-surface / \$color-text-primary /
\$color-text-muted refs. Caller MUST have run
applySemanticPalette(doc) first or refs resolve to undefined
(documented in schema description).
Scrim stays #000000 in ALL themes — modal backdrops are a dim
effect, not a themeable surface. Pinned by handler test.
9 handler test cases cover: registration + schema shape +
required[title] + theme variants + scrim invariant + bogus
parent_id rejection. Parity test added a CASES entry with
\`theme:'dark'\` args (exercises the theme branch in both
shim and server paths).
elements.md gained:
- §46b decision-tree entry pointing to the v1 variant
- Trigger list entry for dark-mode / theme-aware prompts
- Minimal usage showing \`theme:'dark'\` + \`theme:'system'\`
This is the reference implementation for the remaining 9 theme-
aware v1 tools in the top-10 offenders list (empty-chart,
chip-input, toast, pagination, notification-row, image-placeholder,
faq-item, comment, checkbox — per dark-theme-audit §offenders).
Five regex sites across pen-ai-skills + pen-mcp + apps/web were
anchored at `_v0$`, blocking the _v1 family from being recognized
as element tools:
- packages/pen-ai-skills/src/corpus/output-parser.ts
ELEMENT_TOOL_NAME_RE (filters tool_call outputs in A/B
scorer)
- apps/web/src/services/ai/design-parser.ts:106 (embedded
orchestrator dispatch)
- packages/pen-mcp/src/__tests__/design-prompt-elements.test.ts
(×2 — stale-integration guard for elements.md)
- packages/pen-mcp/src/__tests__/element-tool-registry-parity.test.ts
("every tool name matches convention" — renamed to _vN)
- apps/web/src/services/ai/__tests__/element-tools-dispatcher.test.ts
(drift guard for SUPPORTED_EMBEDDED_ELEMENT_TOOLS)
All now accept /^add_[a-z_]+_v\d+$/. Registry-parity test's
expectedBuilder mapping already handled both v0 (strip suffix →
buildModalShell) and v1+ (preserve → buildModalShellV1) via the
existing `.replace(/_v0$/, '')` — no change there.
Prerequisite for landing add_modal_shell_v1 as a first-class MCP
tool in the next commit.
Ships the first theme-aware element builder demonstrating the v1
contract end-to-end. `buildModalShellV1({ title, theme })` accepts
three theme variants:
- `'light'` (default): byte-parity with buildModalShell v0 —
same hex literals, same role tree, same structural shape.
Structural test asserts stripIds(v0) === stripIds(v1) for
the default-theme path.
- `'dark'`: hardcoded dark-palette hex (#1E293B card, #F1F5F9
title, #94A3B8 muted text). No \$refs — this path is for
callers who want a dark modal without the theme-switching
infrastructure.
- `'system'`: emits \$color-surface / \$color-text-primary /
\$color-text-muted refs. Renders track \`themes.Mode\` at
paint time. Requires \`applySemanticPalette(doc)\` to have
been seeded; if not, refs resolve to undefined (caller's
responsibility per the v1 contract).
End-to-end tests validate the 'system' path's round-trip through
resolveColorRef for both Light (→ #FFFFFF) and Dark (→ #1E293B)
modes. That's the full chain working:
buildModalShellV1({theme:'system'}) → tree with \$refs
→ applySemanticPalette(doc) → palette seeded
→ resolveColorRef(ref, doc.variables, {Mode:'Dark'}) → hex
One intentional design note tested: scrim stays #000000 in BOTH
light and dark themes. Modal backdrops are a "dim everything
below" effect, not a themeable surface — dimming a dark surface
with a dark color is a better visual than a themed shade.
18 tests. v0 byte-parity verified against the existing
buildModalShell for full structural equality. This is the
reference implementation for all subsequent v1 tools (top-10
offenders per the dark-theme audit).
Extends the generation-phase variables skill with a table of the
14 semantic palette tokens that \`applySemanticPalette(doc)\`
seeds. Models consuming this skill learn:
1. The exact token names + their light/dark resolved values
(can cross-reference against what the user's document
actually has via \`hasSemanticPalette\`)
2. When to PREFER \`\$color-*\` refs over hex literals (theme-
aware intent: dark-mode design, system-follow apps, user-
toggleable themes)
3. When to FALL BACK to hex (default createEmptyDocument state
where the palette isn't seeded)
4. That semantic tokens override theme — \`\$color-success\`
stays green in both light and dark modes because "green"
is the semantic signal, not a visual choice
Without this guidance, an AI asked for a dark-themed dashboard
could either: (a) emit hex literals that don't track theme
(defeats the purpose), or (b) emit \$color-* refs blindly into
a doc that lacks the palette (resolves to undefined, renders as
raw string). The table + fallback rule close both gaps.
Ships the canonical 14-token palette called for in the dark-theme
audit (openpencil-docs/superpowers/notes/2026-04-22-dark-theme-
defaults-audit.md §role clusters). Every token has paired Light +
Dark values on a single `Mode` theme axis.
API surface:
- getSemanticPalette() → {themes, variables} for merge
- getSemanticPaletteHex(mode='Light') → flat Record<name, hex>
- applySemanticPalette(doc) → non-destructive merge (user-
defined variables + theme axes WIN on collision; palette is
purely additive)
- hasSemanticPalette(doc) → runtime check for v1 tools before
emitting \$color-* refs
- getSemanticPaletteDescription(name) → human-readable string
for token-picker UI
- SEMANTIC_PALETTE_NAMES + theme-axis constants exported
The 14 tokens:
- Surfaces: color-surface, color-surface-2, color-surface-3,
color-bg-deep
- Borders: color-border, color-border-strong
- Text: color-text-primary, color-text-body, color-text-muted,
color-text-subtle
- Semantic: color-accent, color-destructive, color-success
- Other: color-scrim (with alpha for modal backdrop)
Intentionally NOT wired into createEmptyDocument(). Seeding by
default would alter every existing document on re-save and
violate the v0 byte-parity contract (the whole point of the
audit). v1 tools will call applySemanticPalette(doc) as a pre-
flight, OR the app shell offers a "enable dark theme" user
action that triggers the apply.
29 tests cover: palette shape (14 variables, 2 themed values each,
hex-formatted, light≠dark), hex getter for both modes, apply
non-mutation + user-var-wins-on-collision + additive theme-axis
merge, hasSemanticPalette (empty / full / partial), and full
round-trip through the existing resolveVariableRef / resolveColorRef
paths for every palette name.
Two tsc errors surfaced by a later tsc run:
1. browser-image-search-fetcher.test.ts — imported `beforeEach`
but never used; also `spy.mock.calls[0]` typed as empty tuple
since vi.fn()'s signature isn't inferred. Cast via `unknown` +
explicit tuple shape.
2. chart-builders-visual-smoke.test.ts — custom-dimensions case
passed `width`/`height` to buildChartLine. The real shape is
`point_spacing` + `chart_height`. Fixed the test to use the
actual param names (test itself wasn't broken, just the type).
Geometric invariants for the three chart builders (bars/line/pie)
that a rendering failure would start from. Can't run Skia
headlessly in unit tests (CanvasKit WASM is heavy + GPU-context-
dependent), so this is the cheap smoke layer that catches shape-
level regressions before they reach the renderer. The app-level
debug_screenshot MCP tool gives us real visual regression on top.
Per chart type:
- buildChartBars: 8 variants (default, all-equal, single, small
values, large values, custom dims, zeros, empty-throws).
Asserts one chart-bar per value, all dims finite.
- buildChartLine: 9 variants (smooth, monotonic, flat, two-point,
spike, fractional, custom dims, single-value, empty-throws).
Asserts chart-line geometry present, coords finite.
- buildChartPie: 9 variants (equal, skewed, two, many-thin,
custom diameter, donut at 0.5 and 0.9 ratio, single 100%,
empty-throws, all-zero-throws). Asserts one slice per value,
startAngle/sweepAngle finite, sweepAngle positive, total
sweep = 360°.
Cross-chart invariants: same input length → same geometry-child
count; all three types produce finite-coord trees on identical
input.
Empty-input behavior: all three builders throw with clear
messages. Test pins the throws as intended behavior — the
alternative (silently returning an empty tree) would let an
all-zeros dataset produce a "chart is there but invisible" UI
bug that's much harder to diagnose than the explicit throw.
Deterministic mock-LLM for local unit-level testing of the scorer
+ output-parser + apply pipeline — the same three components real
A/B runs use, minus the network.
mockLlmRaw(prompt, variant) returns the raw string a "well-
behaved" model would emit:
- Treatment on obvious prompt → <op_tool>{...}</op_tool> naming
expected_tool_if_any with empty args
- Baseline → minimal batch_design DSL with a single frame role-
stamped from must_contain_roles[0]
- Optional + no hint → falls through to the baseline path
mockLlmParsed() skips the raw-string round-trip and returns a
ParsedOutput directly for tests pinning a specific kind. Both
respect (promptId, variant) overrides so tests can simulate
garbage / wrong-tool / empty outputs inline.
Integration test loads the real ab-v1 corpus from disk and
exercises every prompt through the full pipeline:
corpus → mockLlmRaw → parseModelOutput → scoreRun → ScoreRow
Verifies all 4 routing outcomes (right-tool / wrong-tool /
fallback / garbage) classify correctly on mocked input. 17 test
cases; corpus sweep runs in ~6ms — fast enough to gate every
PR without slowing CI.
This is the prerequisite for future "real" A/B test runners: if
the harness misclassifies obvious mock inputs, no conclusion
from a real run would be trustworthy.
In-memory counters for the element-tool dispatcher, exposed via
\`getElementToolMetric(name)\` / \`getAllElementToolMetrics()\` /
\`getTopElementToolCalls(n)\` / \`resetElementToolMetrics()\` in
packages/pen-mcp/src/metrics/.
\`handleElementToolCall\` now wraps the existing switch in a
try/record — every dispatch increments \`calls\`, thrown handlers
additionally bump \`errors\` and stash the last error message.
Unknown tool names still fire a counter (useful signal: "the AI
picked a tool we don't have").
Process-local / in-memory by design:
- Test determinism: resetElementToolMetrics() in beforeEach
- Matches stdio MCP server's one-client-one-server model
- No persistence backend choice baked in — if we need
cross-restart persistence later, a thin serializer drops on
top without touching this API
Unlocks #92 Local A/B harness: feed a corpus through the MCP
server, read back getAllElementToolMetrics() to see which tools
the model actually picked vs what the corpus expected. Core
observability for non-Claude regression detection.
Pins that text-carrying element builders preserve the caller's
byte representation verbatim — no silent normal-form conversion,
no zero-width stripping, no fullwidth↔ASCII collapse.
Tests each of NFC/NFD/NFKC/NFKD forms through 6 representative
builders (heading / body-text / list-row / form-field / faq-item /
comment), plus 5 targeted fixtures:
- Zero-width joiner mid-word ("emoji" stays 6 codepoints)
- BOM at string start
- ZWJ emoji family sequence (4-person glyph)
- Vietnamese combining-marks (NFC vs NFD both preserved as-is)
- Halfwidth/fullwidth CJK distinction (NFKC would collapse
fullwidth "A" to ASCII "A"; we assert the builder does NOT)
Why this matters: macOS ships filenames in NFD, Windows/web in
NFC; copy-paste carries any form; some CJK inputs emit
precomposed, others decomposed base+combining. Exact-match
lookups in external systems (especially emoji-less fallback
keys) break silently if the builder pre-empts the downstream
validator's normalization decision. Builders must pass through
bytes unmodified.
AI orchestrators occasionally emit very large batches (one sub-
agent producing a whole section in a single batch_design call).
Existing multi-line regression only covered pretty-printed JSON
in SINGLE ops — nothing pinned behavior when N itself grows.
5 scenarios:
1. 100 sibling I() ops → all land, <5s wall-clock
2. 200 sibling ops → 2x node count, <10s (catches O(n²) regressions)
3. 30-level nested I() chain via parent_id threading
4. 250 mixed ops (50 sections × 4 children) with parent refs
5. Partial failure: 1 bogus parent_id among 100 good ops — good
ones still land (don't let one bad op poison the batch)
Observed: 250-op mixed batch completes in ~42ms on an M-series
machine. Budgets are "reasonable" (5s / 10s / 15s), not "fast" —
they're meant to catch O(n²) regressions in the DSL parser / tree
insert / save loop, not enforce a perf target.
Per-tool handler tests cover "this tool emits the correct shape"
individually. This file covers the next layer: can N element-tool
calls chain together into a realistic multi-section screen without
breaking tree invariants?
Scenarios (each spans multiple tool families to catch cross-
family regressions):
1. Mobile settings — top_nav + 2 sections × 3 list_rows + bottom_nav (10 calls)
2. Dashboard home — top_nav + stat_grid + section + 3 metric_comparisons + chart (7)
3. Login form — heading + body + 2 form_fields + button + link (6)
4. Profile + UGC — top_nav + avatar + heading + badge + 2 faq_items + action_menu (7)
5. Listing — search + card_row + divider + empty_chart + date_picker + chip_input + pagination (7)
6. parent_id threading invariant — nested insert actually lands under named parent
Each scenario asserts:
- Every call emits a nodeId (no silent no-ops)
- Final document parses as valid JSON with expected root children count
- Every call's nodeId is findable in the saved tree
- Every tool's canonical role survives post-save
- parent_id threading works (child lands under named parent, not root)
This is the integration gate that catches "tool wiring works
individually but composes wrong" — the ghost regression that can
slip past per-tool tests.