Commit graph

152 commits

Author SHA1 Message Date
Fini a2bf13bfce feat(ab-corpus): dashboard-applied-filter-tag ab-v1 prompt for add_tag_v0 2026-04-27 08:58:00 +08:00
Fini 072856e522 feat(ai): add_tag_v0 — single closable filter chip (80th tool) 2026-04-27 08:55:00 +08:00
Fini e14b9446c8 fix(ai): data-table cells clip overflow so long content stays in column
Long cell content rode `width=auto` and the cell itself had
`height=fit_content` with no clip, so a string longer than its
allotted column would push past the cell edge into the adjacent
column at render time. Switch each cell to `width/height=
fill_container + clipContent=true`, give the text child
`width=fill_container + textGrowth=fixed-width` so the layout
engine wraps inside the cell, and clip the row itself so a runaway
cell can't push siblings off-row either. Lock the contract in the
handler test.
2026-04-27 08:50:00 +08:00
Fini 68411c9e8b feat(ab-corpus): dashboard-customers-table ab-v1 prompt for add_data_table_row_v0 2026-04-27 08:45:00 +08:00
Fini d24b587b6c feat(ai): add_data_table_row_v0 — desktop tabular row (79th tool) 2026-04-27 08:40:00 +08:00
Fini 4442d8fe26 feat(ab-corpus): dashboard-team-presence ab-v1 prompt for add_avatar_group_v0 2026-04-27 08:35:00 +08:00
Fini e225104e43 feat(ai): add_avatar_group_v0 — stacked presence tile group (78th tool) 2026-04-27 08:30:00 +08:00
Fini b79ce5233d test(pen-core): static guard against CSS-side padding shorthands
Greps every builder source for paddingTop/Right/Bottom/Left as a
property key and fails if any reappear. The unified padding field is
the only form resolvePadding reads; the CSS-side siblings render with
zero inset, which slipped past type checks because ElementTree is
Record<string, unknown>. JSDoc / line-comment mentions stay allowed
so builders can still document the trap inline.
2026-04-27 08:25:00 +08:00
Fini c5655000fd chore(release): bump to v0.8.0 2026-04-27 08:20:00 +08:00
Fini 6e54054c73 chore(merge): integrate origin/v0.8.0 — main pre-release sync + CI fixes
origin's v0.8.0 had cherry-picks of the v0.7.5 deepseek/image-search
fixes (a727632a, a5952bc8) overlapping local 2073cf5b / 04f4fbc1, plus
new commits (model-selector ark-coding deepseek-v4-pro/flash IDs that
ARK rejects, fetch error.cause unwrap, CI agent-native build, op
export docs cleanup, main merge). Resolved the ark-coding list in
favor of HEAD's deepseek-v3.2 entry (only model ARK Coding Plan
actually supports — see openpencil-docs note).
2026-04-27 08:15:00 +08:00
Kayshen-X b554b4f1a6 Merge branch 'main' of github.com:ZSeven-W/openpencil into v0.8.0 2026-04-26 19:39:14 +08:00
Kayshen-X b80db5b798 docs: drop op export from CLI docs and clarify pen-mcp usage
The `op export` command was removed in 0.7.x but the README still
advertised it (#116). The pen-mcp README also documented an
`npx @zseven-w/pen-mcp` quick-start that never worked because the
package ships TypeScript source against workspace-only deps with no
`bin` entry (#117).

- Strip `op export` references from all 15 root and 15 cli READMEs
- Sync AGENTS.md, CLAUDE.md, apps/cli/CLAUDE.md to match the codegen-
  pipeline reality (no standalone export command anymore)
- Rewrite pen-mcp README's quick-start: explain the package ships as
  part of the OpenPencil app and external clients connect over HTTP

Closes #116
Closes #117
2026-04-26 19:20:14 +08:00
Fini abbccc1ba2 fix(ai): refresh DeepSeek defaults to v4 model series
`/models` now returns only deepseek-v4-pro and deepseek-v4-flash;
deepseek-chat / deepseek-reasoner sunset 2026-07-24 and the
deepseek-v3.2 hard-coded in the ark-coding fallback list never
existed. Both v4 models default to thinking enabled and the API
toggles via `{"thinking":{"type":"disabled"}}` — keep
`thinkingMode: 'disabled'` so the app's fast/non-thinking default
stays intact (server reasoning paths honor it; the Zig openai-compat
path doesn't emit the toggle yet, so calls through that path still
get provider-default thinking until it's wired). v4-pro promoted to
full tier; legacy aliases pinned to an exact RegExp so future
deepseek-* variants don't inherit a forced disabled mode.

Bandaid for the unwired toggle: v4-pro gets `timeoutMultiplier: 2`
because its default-on reasoning blows past the orchestrator's
planning timeout on long system prompts (observed in dev: planning
phase falls back, sub-agent then succeeds — UX degraded but
functional). Drop the multiplier once the Zig path actually sends
`thinking:{type:disabled}`.

Don't add a BUILTIN_MODEL_LISTS.deepseek entry — DeepSeek exposes
/v1/models, so let `fetchProviderModels` pull the live catalog
through `/api/ai/provider-models` instead of pinning a snapshot
(the ark-coding `deepseek-v3.2` ghost above shows what those
snapshots drift into).
2026-04-26 06:30:00 +08:00
Fini 2f62cb262f fix(core): replace CSS padding shorthands with canonical padding array in 14 element builders
paddingTop/Right/Bottom/Left are silently dropped by resolvePadding in
the layout engine, causing all four sides to render as 0. Switch every
affected builder to the canonical padding: [T,R,B,L] form.
2026-04-26 06:29:59 +08:00
Fini ae40046b17 fix(ai): sidebar_nav uses unified padding so insets actually render
pen-core's resolvePadding only reads the array/scalar `padding`
field; the CSS-style paddingTop/Right/Bottom/Left siblings I shipped
in 5c275ba1 were silently dropped, so the rail rendered flush. Switch
to padding=[16,12] / [8,12,24,12] / [0,12] and add a regression
assertion that catches the same trap on any future tweak.
2026-04-26 06:29:58 +08:00
Fini 0cb0e81a78 feat(ab-corpus): dashboard-team-sidebar ab-v1 prompt for add_sidebar_nav_v0
Adds the obvious-difficulty prompt covering the 77th element tool;
bumps the corpus-loader test from 26 to 27 entries and lists the new
2026-04-27 batch in the README so future corpus drift is caught at
test time.
2026-04-26 06:29:57 +08:00
Fini e9d9e37770 feat(ai): add_sidebar_nav_v0 — desktop persistent left rail (77th tool)
Top_nav_bar / bottom_nav / nav_chip_row covered mobile and inline
chrome but desktop dashboards still had to hand-roll their left
sidebar via batch_design. Adds a 240px-wide vertical rail with
icon+label rows, optional brand title, and slate-100 pill bg on the
active item — distinct surface from the existing nav tools so the
decision tree picks it cleanly for "sidebar / side nav / 侧边栏".
2026-04-26 06:29:56 +08:00
Fini 6c44d58a31 feat(ai): add_cookie_banner_v0 — GDPR/CCPA cookie banner (76th tool)
Sticky bottom-of-page disclosure card with title, body, accept /
decline buttons (decline-then-accept order), and an optional
"Cookie settings" link for fine-grained consent. Caller positions
the banner; the tool emits the card itself with shadow + 1px slate
border for visual lift above the page content. Plus the v1 corpus
prompt landing-cookie-banner (26 v1 prompts total).
2026-04-25 08:15:00 +08:00
Fini 1acd31a2e1 feat(ai): add_input_with_action_v0 — input + inline action button (75th tool)
The "Subscribe to newsletter" / "Apply discount code" / "Send chat
message" pattern. Two action variants:
  - text (default): pill button with label like "Subscribe"
  - icon: 44×44 square icon button (chat send arrow / search apply)

Distinct from add_form_field_v0 (label-above, no inline button)
and add_search_bar_v0 (no trailing action button). Optional
leading_icon adds an icon inside the input itself. Plus the v1
corpus prompt landing-newsletter-signup (25 v1 prompts total).
2026-04-25 08:00:00 +08:00
Fini 3527f6ff6a feat(ab-corpus): mobile-phone-input ab-v1 prompt for add_phone_input_v0
Adds 24th v1 prompt and rolls the 2026-04-25 batch from 6 to 7
tools (5 v0 + 2 v1). Loader test count + tool-set assertions
updated.
2026-04-25 07:45:00 +08:00
Fini 577488546a feat(ai): add_phone_input_v0 — international phone input (74th tool)
The "+1 (555) …" pattern from every modern signup / login screen.
A 44px row with leading country selector (flag + dial code +
chevron-down), a 1px slate divider, and the digits input on the
right. Country selector is a button-shape (no actual dropdown
menu); caller handles picker UX as a separate concern. `value`
toggles between placeholder (slate-400) and populated (slate-900).
2026-04-25 07:30:00 +08:00
Fini 008bf0ca46 test(ab-corpus): update v1 corpus expectations to 23 prompts
Three batches now: 2026-04-22 (12) + 2026-04-24 (5) + 2026-04-25 (6
— 4 v0 + 2 v1). The 2026-04-25 batch corresponds to the tools
shipped in this session: social_login / pricing_card / stat_card /
range_slider / toast_v1 / empty_chart_v1.
2026-04-25 01:47:23 +08:00
Fini dd5c85dcc0 feat(ab-corpus): ab-v1 prompts for 6 new tools shipped this session
Adds NL prompts exercising the new narrow tools + theme-aware v1s:
- mobile-social-login (→ add_social_login_row_v0)
- dashboard-pricing-card (→ add_pricing_card_v0)
- dashboard-revenue-stat (→ add_stat_card_v0)
- mobile-volume-slider (→ add_range_slider_v0)
- mobile-dark-toast (→ add_toast_v1, inverted contrast)
- dashboard-dark-empty-chart (→ add_empty_chart_v1)

Each prompt describes the visual intent in natural language, asserts
the expected roles, and pins the expected_tool_if_any so the harness
can attribute tool-choice correctness.
2026-04-25 01:39:40 +08:00
Fini d1022e74d5 feat(ai): add_empty_chart_v1 — theme-aware no-data placeholder
Third theme-aware v1 (after add_modal_shell_v1 and add_toast_v1).
Same dashed-border "no data yet" tile shape as v0 with a `theme`
param that swaps 5 colors (bg / border / icon / title / subtitle):
  - light (default): byte-parity with v0 (slate-50 / slate-300)
  - dark: slate-800 bg + slate-600 border + slate-200 title
  - system: $color-surface-2 / $color-border / $color-text-primary
    refs — requires applySemanticPalette(doc) seeded

Lets dark-theme dashboards keep a matching empty-slot surface
instead of punching a light rectangle out of dark cards.
2026-04-25 01:37:40 +08:00
Fini aea75fa198 feat(ai): add_range_slider_v0 — single-thumb range slider (71st tool)
Visual static representation of a horizontal slider. Track splits
into left fill (accent) + 20×20 thumb (white + accent stroke) +
right remaining (slate), all aligned on a 20px track wrap. Optional
label + value readout row above; `value_suffix` renders "60%" /
"128px" / "0°" style readouts. Pixel math: fill = (width-20) * pct,
auto-collapses fill or remaining at either extreme.
2026-04-25 01:33:46 +08:00
Fini 984bb10bd1 feat(ai): add_toast_v1 — theme-aware floating pill toast
Second theme-aware v1 (after add_modal_shell_v1). Same pill shape as
v0 with a `theme` param:
  - light (default): byte-parity with v0 (dark #111827 pill + white fg)
  - dark: INVERTED contrast — light pill (#F1F5F9) + dark fg (#0F172A)
  - system: $color-text-primary bg + $color-surface fg (inverted swap)

Unlike surface-like v1s (modal-shell), toasts use inverted contrast by
design — a dark pill on light bg, a light pill on dark bg — so the
dark variant flips the pill rather than darkening it.
2026-04-25 01:28:43 +08:00
Fini 91d303d0f0 feat(ai): add_pricing_card_v0 — SaaS pricing tier card (70th tool)
The "Pro $29/month" column for pricing tables. Tier name + big
price (currency + amount + period) + check-mark feature list + CTA
button. Two emphases: default (slate border + slate CTA) and
featured (accent border + accent CTA + auto "Most popular" badge
unless overridden by explicit `badge` value).
2026-04-25 01:24:24 +08:00
Fini e0b926c256 feat(ai): add_social_login_row_v0 — social auth button row (69th tool)
The "Continue with Google / Apple / Microsoft" row. Vertical default
(stacked full-width 48px buttons) or horizontal (compact icon-only
48×48 pills). Known provider names (google, apple, github, microsoft,
facebook, twitter, linkedin, discord, slack, gitlab, email, phone)
auto-map to lucide icons; `icon` param overrides for SSO/SAML/Okta.
2026-04-25 01:19:00 +08:00
Fini 127eccdefb feat(ai): add_stat_card_v0 — big-number KPI tile (68th tool)
Featured-metric dashboard tile: label (uppercase muted) above a
huge 32/700 primary value, with optional tone-colored delta line
and corner icon. Distinct from its neighbors:

  - add_stat_grid_v0   — multi-cell side-by-side, smaller values
  - add_metric_comparison_v0 — horizontal label+value+inline arrow
  - add_stat_card_v0 (this) — single featured metric, whole-card focus

Trend enum tones the delta line only (value stays slate-900):
  up    → #10B981 (emerald)
  down  → #EF4444 (red)
  flat  → #64748B (slate, default)

Wired through all standard points — schema into ext-3 (shortest
shard at 371 lines + new tool → 408, still comfortably under 800)
+ shim + SERVER_BUILDERS + parity CASES + contract allow-list +
elements.md decision tree + triggers + minimal usage.

Handler test covers 6 cases: minimal registration + defaults +
icon+delta+up-tone + down-tone + flat-tone default + width
clamp + bogus parent_id rejection.
2026-04-25 01:09:49 +08:00
Fini 54ee6e0ec7 feat(ab-corpus): --corpus flag + 5 new prompts for tools 63-67 + v1
Two independent changes rolled together since they both serve the
same goal — "can real LLMs actually route to the tools we shipped
today?":

  1. scripts/ab-corpus/run.ts gains a `--corpus` flag (ab-v0 |
     ab-v1, default ab-v0 for back-compat). The harness was
     hardcoded to ab-v0 — adding ab-v1 prompts was worthless
     without a way to run them. Validated 17 prompts × 2 models ×
     2 variants during a live 方舟-CP run.

  2. 5 new ab-v1 prompts cover the 2026-04-24 tool batch:
      - mobile-upload-dropzone.yaml        → add_upload_dropzone_v0
      - mobile-otp-verification.yaml       → add_otp_input_v0
      - mobile-file-attachment.yaml        → add_attachment_row_v0
      - mobile-chat-message.yaml           → add_chat_bubble_v0
      - dashboard-dark-modal.yaml          → add_modal_shell_v1

corpus-loader.test.ts bumps its count assertion 12 → 17 and
extends the tool-coverage set. Also loosens the regex to accept
`_v\d+$` (was `_v0$`) so add_modal_shell_v1 passes. No other
test file needed changes — the existing registry-parity and
mock-llm tests already use `_v\d+$` or the registry directly.

.gitignore gains:
  - .playwright-mcp/ (MCP Playwright session artifacts)
  - editor-*.png      (local verification screenshots)
  - scripts/ab-corpus/runs/  (live-run outputs / reports)

None of those belong in version control — they're artifacts
from local verification runs.
2026-04-25 00:23:11 +08:00
Fini 6d4bdb88da test(pen-mcp): support-chat composition scenario covering #63-#67 + v1
Extends element-tools-composition.test.ts with a 7th real-screen
scenario that exercises every tool added after the 62-tool mark:

  - add_top_nav_bar_v0         (existing anchor)
  - add_chat_bubble_v0         × 2  (left from-others + right from-self)
  - add_attachment_row_v0      (file on the self-message path)
  - add_upload_dropzone_v0     (drop area for screenshots)
  - add_chip_input_v0          (conversation tags)
  - add_action_menu_v0         (floating menu, open state)
  - add_modal_shell_v1         (theme="dark" confirm dialog)
  - add_otp_input_v0           (phone verification step)

8 tool calls, chained through the full MCP handler pipeline
(ensureParentExists → builder → assignIdsRecursively →
batch_design insert with rollback-on-failure → post-insert
landing check → save → re-read from disk). Asserts:

  - every call emits a nodeId (no silent no-ops)
  - final doc has exactly 9 root children
  - every tool's role marker survives round-trip save/load
  - modal v1 theme=dark → card fill #1E293B (not v0's #FFFFFF)
  - chat right-side bubble surface → #2563EB accent
  - OTP focused slot stroke → #2563EB accent

This is end-to-end at the DATA layer, not the visual layer.
Still not covered by any test:
  - Skia rendering (needs debug_screenshot against a live
    canvas)
  - Real LLM tool routing (needs ab-corpus harness with live
    API keys — route to 方舟 CP works but has not run since
    the routing fix)

Document-layer coverage is enough to rule out handler-pipeline
bugs (bad parent_id threading / cache-stale doc-state / silent
rollback swallowing); visual regressions require the next gate.
2026-04-22 22:31:49 +08:00
Fini 06362a4040 refactor(mcp): split ext-2 into ext-2 + ext-3 (keep shards under 800)
Codex stop-hook caught element-tool-defs-ext-2.ts at 832 lines
after add_chat_bubble_v0 landed there — 32 over the repo's 800-
line ceiling. Same trap the original single ext file hit at 1329
lines, same fix pattern: carve the second half into a new shard.

Split at `add_modal_shell_v1` (tool #16 of 24 in old ext-2):
  - ext-2 keeps tools 1-15 (calendar_grid through textarea) →
    482 lines
  - ext-3 (new) holds tools 16-24 (modal_shell_v1 through
    chat_bubble) → 371 lines

All registry shards now:
  base    683
  ext     671
  ext-2   482
  ext-3   371
  (props   26)
  (top    254)

element-tool-defs.ts concatenates all three ext shards into the
single ELEMENT_TOOL_DEFINITIONS — external API unchanged.
ELEMENT_TOOL_DEFINITIONS_EXT_3 is the new import; 67 tools
still resolve.

Header in ext-1 updated to reflect the three-way split + advise
"when ANY shard crosses 700, carve a ~5-tool chunk into the
shortest shard" so the next rebalance happens proactively instead
of after a stop-hook trip.
2026-04-22 22:12:26 +08:00
Fini 165fb94e44 feat(ai): add_chat_bubble_v0 — messaging bubble (67th tool)
Chat / messaging / customer-support UI message unit. Two variants
via the `side` enum:

  - side="left" (default): from-others bubble. Slate-100 fill,
    slate-900 text, alignItems=flex-start. Optional `author` text
    shown above the bubble (group-chat pattern).
  - side="right": from-self bubble. Accent-color fill (customizable
    via `accent_color`), white text, alignItems=flex-end. Author
    intentionally suppressed on this side — a self-bubble never
    carries "You:".

Optional `timestamp` below the bubble on either side.

Max-width mechanic: pen-core has no native max-width primitive, so
`max_width` becomes the bubble's fixed width (clamped 160..480).
Short messages get extra padding on one side — matches every real
chat client (iMessage / WhatsApp / Slack). Message text uses
`textGrowth: 'fixed-width'` + `width: 'fill_container'` to wrap
correctly inside the fixed-width surface.

Full wiring: pen-core builder + pen-mcp handler + schema (into
ext-2, shorter shard — 24 tools vs ext-1's 24 after this) + shim
+ SERVER_BUILDERS + parity CASES + contract allow-list +
elements.md decision tree + triggers + minimal usage for both
sides. Handler test covers 9 cases: registration, left defaults,
left+author, right with self-dropped-author, right+accent_color,
timestamp both sides, max_width clamps (low + high split into
separate tests to avoid cache-interference), textGrowth wiring,
bogus parent_id rejection.
2026-04-22 22:03:44 +08:00
Fini 49d7d1baf2 feat(ai): add_attachment_row_v0 — file attachment list unit (66th tool)
Fills another common UI gap: the "here's an already-uploaded file"
row you see in email composers, chat attachments, and form upload
summaries. Compact horizontal layout: type-icon + filename (bold) +
optional muted size string + optional right-side × remove affordance.

Structure: horizontal frame (slate-50 bg, cornerRadius=8) with
three children:
  1. attachment-icon — lucide file-* (caller picks: file / file-
     text / file-image / file-video / file-audio / file-archive /
     file-spreadsheet / file-code)
  2. attachment-meta — vertical frame with filename + optional size
  3. attachment-remove — × icon, suppressed via removable=false

Intentionally NOT embedding an upload-progress variant in v0. The
pen-core schema lacks percentage-width primitives, so a %-filled
progress bar would either need a fixed track width (brittle across
parents) or a caller-computed pixel value (awkward API). Callers
who need the uploading state compose `add_progress_bar_v0` directly
below the row — cleaner separation.

Wired through all standard points: schema into ext-1 (balanced
shards 23/23 after upload-dropzone landed there last commit) +
shim + SERVER_BUILDERS + parity CASES + contract allow-list +
elements.md decision tree + triggers + minimal usage.

Handler test covers 7 cases: registration + minimal (no size) +
size rendered + custom icon + removable=false + default icon +
bogus parent_id rejection.
2026-04-22 21:53:47 +08:00
Fini dc73885015 feat(ai): add_otp_input_v0 — verification code input (65th tool)
Fills the auth-flow gap: 2FA / PIN / phone-verification codes.
Horizontal row of N square slots (4..8), one digit per slot.
Renders three states per caller intent:

  - blank (no `digits`): all slots empty, `focused_index` marks
    the currently-typing slot with an accent-color 2px outline
  - partial: first M slots filled with digit text, slot M+1
    focused, rest empty
  - full: all N slots filled (final submittable state)

Filled slots get role=otp-slot-filled + slate-700 border + 20/600
digit text. Focused empty slot gets role=otp-slot-focused +
2px accent border. Blank unfocused slots get role=otp-slot +
1px slate-300 border.

Wired through all standard points — schema into ext-2 (shorter
shard) + shim + SERVER_BUILDERS + parity CASES + contract
allow-list + elements.md decision tree + triggers + minimal
usage for each state.

Handler test covers 8 cases: registration + defaults (6 blank
focused-first) + partial state / full state / length clamp low
(< 4 → 4) + length clamp high (> 8 → 8) + accent color override
+ bogus parent_id rejection.
2026-04-22 21:48:05 +08:00
Fini 45214efb70 feat(ai): add_upload_dropzone_v0 — file drop zone (64th tool)
Fills a real gap in the element-tool family: upload / drag-and-
drop surfaces. Dashed border + cloud icon + two-line instruction
("Drop files to upload" / "or click to browse") — the classic
pattern from every modern file-upload UI.

Deliberately structurally similar to add_empty_chart_v0 (dashed
border + icon + title/subtitle) but semantically distinct:
  - empty_chart = "chart widget will render when data arrives"
    (320×200, icon chart-typed)
  - upload_dropzone = "users drop files here" (480×200, icon
    semantic: upload-cloud / upload / file-up)

elements.md routes them by intent, and the tool descriptions
cross-reference each other to prevent the AI from picking the
wrong one on ambiguous prompts.

Wired through all the standard points per the add-new-tool
checklist: builder + handler + schema (into ext-1, the shorter
shard) + shim + SERVER_BUILDERS + parity CASES + contract
allow-list + elements.md decision tree + keyword triggers +
minimal usage. Handler test covers 5 cases: defaults, dashed
stroke, overrides, size clamping, bogus parent_id rejection.
2026-04-22 21:41:13 +08:00
Fini ed449bc896 refactor(mcp): split element-tool-defs-ext in half + extract shared props
[Codex P3] element-tool-defs-ext.ts had grown to 1329 lines —
over the repo's documented 800-line ceiling. Ironically the file
header comment claimed it existed to keep its parent under the
cap, but the shard itself had outgrown the limit.

Split into three files:

  - element-tool-def-props.ts (26 lines) — shared JSON-Schema
    fragments (schemaVersionProp / filePathProp / parentIdProp /
    pageIdProp) that every definition file uses. Deduplicating
    these unblocks the split cleanly.
  - element-tool-defs-ext.ts (596 lines) — first 22 tools
    (add_switch_v0 through add_segmented_control_v0 era). Imports
    the shared props.
  - element-tool-defs-ext-2.ts (744 lines, new) — remaining 22
    tools starting at add_calendar_grid_v0. Imports the shared
    props.

element-tool-defs.ts concatenates all three arrays into the single
ELEMENT_TOOL_DEFINITIONS registry — external API surface unchanged.

Header comments in both shards now document the split convention:
"pick whichever shard has fewer tools" when adding a new entry,
to keep the files balanced as the family grows toward ~100.

Incidentally the previous commit's file also carried the P2 fix
(pageId threading through the in-browser and HTTP DSL paths, so
multi-page docs land the generation on the ACTIVE page instead of
doc.pages[0]). Both touched the same file, didn't make sense to
split. Title-wise the previous commit is P1 but functionally it's
P1+P2.

All under the 800-line ceiling now:
  element-tool-defs-base.ts  683
  element-tool-defs-ext.ts   596
  element-tool-defs-ext-2.ts 744
  element-tool-defs.ts       240
  element-tool-def-props.ts   26
2026-04-22 11:20:00 +08:00
Fini 0ce113c733 fix(ai): dispatcher imports DSL executor via browser-safe subpath
[Codex P1] The browser-side element-tools-dispatcher imported
runBatchDesignDsl from the \`@zseven-w/pen-mcp\` package barrel.
That barrel re-exports node-only modules — document-manager,
log-utils, theme-presets — which import node:fs / node:path at
top level. Vite / esbuild resolve the barrel BEFORE tree-shaking
can drop those branches, so browser builds failed on unresolved
node built-ins.

Fix:
  - packages/pen-mcp/package.json: add \`./dsl\` subpath export
    pointing at tools/batch-design-dsl.ts — the pure executor
    file already guarded as browser-safe by the adjacent
    regression test.
  - apps/web dispatcher: switch import to
    \`@zseven-w/pen-mcp/dsl\`. No other changes — the re-exported
    symbols (runBatchDesignDsl / OpResult / ImageSearchFetcher /
    RunBatchDesignDslOptions) are identical shape.
  - batch-design-dsl-browser-safe.test.ts: add an assertion that
    package.json's exports field preserves the \`./dsl\` key
    pointing at the expected file. Without this, silently
    removing the subpath would re-introduce the browser-breaking
    resolution path.

The package barrel keeps its current export of runBatchDesignDsl
too (a few internal test files still import from it). Browser
callers should migrate to \`@zseven-w/pen-mcp/dsl\` per the JSDoc
note now in the dispatcher.
2026-04-22 11:15:00 +08:00
Fini dd5d156160 feat(ai): add_modal_shell_v1 — first theme-aware MCP tool (63rd, v1-family debut)
Wires the buildModalShellV1 pen-core builder through the full MCP
toolchain — handler + schema + dispatch + shim + SERVER_BUILDERS
+ parity CASES + handler tests + elements.md skill — so external
MCP clients (Claude Code / Codex / Gemini CLI) can call
add_modal_shell_v1 as a first-class tool alongside the 62 v0
tools.

Three theme variants exposed via `theme` param (enum [light, dark,
system]):
  - omitted / `'light'`: byte-parity with add_modal_shell_v0
    (same hex, same structure, same role tree)
  - `'dark'`: hardcoded dark palette (#1E293B card, #F1F5F9 title,
    #94A3B8 muted). No \$refs needed.
  - `'system'`: emits \$color-surface / \$color-text-primary /
    \$color-text-muted refs. Caller MUST have run
    applySemanticPalette(doc) first or refs resolve to undefined
    (documented in schema description).

Scrim stays #000000 in ALL themes — modal backdrops are a dim
effect, not a themeable surface. Pinned by handler test.

9 handler test cases cover: registration + schema shape +
required[title] + theme variants + scrim invariant + bogus
parent_id rejection. Parity test added a CASES entry with
\`theme:'dark'\` args (exercises the theme branch in both
shim and server paths).

elements.md gained:
  - §46b decision-tree entry pointing to the v1 variant
  - Trigger list entry for dark-mode / theme-aware prompts
  - Minimal usage showing \`theme:'dark'\` + \`theme:'system'\`

This is the reference implementation for the remaining 9 theme-
aware v1 tools in the top-10 offenders list (empty-chart,
chip-input, toast, pagination, notification-row, image-placeholder,
faq-item, comment, checkbox — per dark-theme-audit §offenders).
2026-04-22 11:00:00 +08:00
Fini 315ded51ad chore(ai): generalize element-tool name regex to accept _v\d+
Five regex sites across pen-ai-skills + pen-mcp + apps/web were
anchored at `_v0$`, blocking the _v1 family from being recognized
as element tools:

  - packages/pen-ai-skills/src/corpus/output-parser.ts
    ELEMENT_TOOL_NAME_RE (filters tool_call outputs in A/B
    scorer)
  - apps/web/src/services/ai/design-parser.ts:106 (embedded
    orchestrator dispatch)
  - packages/pen-mcp/src/__tests__/design-prompt-elements.test.ts
    (×2 — stale-integration guard for elements.md)
  - packages/pen-mcp/src/__tests__/element-tool-registry-parity.test.ts
    ("every tool name matches convention" — renamed to _vN)
  - apps/web/src/services/ai/__tests__/element-tools-dispatcher.test.ts
    (drift guard for SUPPORTED_EMBEDDED_ELEMENT_TOOLS)

All now accept /^add_[a-z_]+_v\d+$/. Registry-parity test's
expectedBuilder mapping already handled both v0 (strip suffix →
buildModalShell) and v1+ (preserve → buildModalShellV1) via the
existing `.replace(/_v0$/, '')` — no change there.

Prerequisite for landing add_modal_shell_v1 as a first-class MCP
tool in the next commit.
2026-04-22 10:55:00 +08:00
Fini b07b970daa feat(pen-core): buildModalShellV1 — first theme-aware v1 element (proof of chain)
Ships the first theme-aware element builder demonstrating the v1
contract end-to-end. `buildModalShellV1({ title, theme })` accepts
three theme variants:

  - `'light'` (default): byte-parity with buildModalShell v0 —
    same hex literals, same role tree, same structural shape.
    Structural test asserts stripIds(v0) === stripIds(v1) for
    the default-theme path.
  - `'dark'`: hardcoded dark-palette hex (#1E293B card, #F1F5F9
    title, #94A3B8 muted text). No \$refs — this path is for
    callers who want a dark modal without the theme-switching
    infrastructure.
  - `'system'`: emits \$color-surface / \$color-text-primary /
    \$color-text-muted refs. Renders track \`themes.Mode\` at
    paint time. Requires \`applySemanticPalette(doc)\` to have
    been seeded; if not, refs resolve to undefined (caller's
    responsibility per the v1 contract).

End-to-end tests validate the 'system' path's round-trip through
resolveColorRef for both Light (→ #FFFFFF) and Dark (→ #1E293B)
modes. That's the full chain working:

  buildModalShellV1({theme:'system'}) → tree with \$refs
  → applySemanticPalette(doc) → palette seeded
  → resolveColorRef(ref, doc.variables, {Mode:'Dark'}) → hex

One intentional design note tested: scrim stays #000000 in BOTH
light and dark themes. Modal backdrops are a "dim everything
below" effect, not a themeable surface — dimming a dark surface
with a dark color is a better visual than a themed shade.

18 tests. v0 byte-parity verified against the existing
buildModalShell for full structural equality. This is the
reference implementation for all subsequent v1 tools (top-10
offenders per the dark-theme audit).
2026-04-22 10:50:00 +08:00
Fini 19a1b13764 docs(ai-skills): teach variables.md about the 14 semantic tokens
Extends the generation-phase variables skill with a table of the
14 semantic palette tokens that \`applySemanticPalette(doc)\`
seeds. Models consuming this skill learn:

  1. The exact token names + their light/dark resolved values
     (can cross-reference against what the user's document
     actually has via \`hasSemanticPalette\`)
  2. When to PREFER \`\$color-*\` refs over hex literals (theme-
     aware intent: dark-mode design, system-follow apps, user-
     toggleable themes)
  3. When to FALL BACK to hex (default createEmptyDocument state
     where the palette isn't seeded)
  4. That semantic tokens override theme — \`\$color-success\`
     stays green in both light and dark modes because "green"
     is the semantic signal, not a visual choice

Without this guidance, an AI asked for a dark-themed dashboard
could either: (a) emit hex literals that don't track theme
(defeats the purpose), or (b) emit \$color-* refs blindly into
a doc that lacks the palette (resolves to undefined, renders as
raw string). The table + fallback rule close both gaps.
2026-04-22 10:45:00 +08:00
Fini ce32f9f572 feat(pen-core): 14-variable semantic palette (unblocks v1 theme-aware tools)
Ships the canonical 14-token palette called for in the dark-theme
audit (openpencil-docs/superpowers/notes/2026-04-22-dark-theme-
defaults-audit.md §role clusters). Every token has paired Light +
Dark values on a single `Mode` theme axis.

API surface:
  - getSemanticPalette() → {themes, variables} for merge
  - getSemanticPaletteHex(mode='Light') → flat Record<name, hex>
  - applySemanticPalette(doc) → non-destructive merge (user-
    defined variables + theme axes WIN on collision; palette is
    purely additive)
  - hasSemanticPalette(doc) → runtime check for v1 tools before
    emitting \$color-* refs
  - getSemanticPaletteDescription(name) → human-readable string
    for token-picker UI
  - SEMANTIC_PALETTE_NAMES + theme-axis constants exported

The 14 tokens:
  - Surfaces: color-surface, color-surface-2, color-surface-3,
    color-bg-deep
  - Borders: color-border, color-border-strong
  - Text: color-text-primary, color-text-body, color-text-muted,
    color-text-subtle
  - Semantic: color-accent, color-destructive, color-success
  - Other: color-scrim (with alpha for modal backdrop)

Intentionally NOT wired into createEmptyDocument(). Seeding by
default would alter every existing document on re-save and
violate the v0 byte-parity contract (the whole point of the
audit). v1 tools will call applySemanticPalette(doc) as a pre-
flight, OR the app shell offers a "enable dark theme" user
action that triggers the apply.

29 tests cover: palette shape (14 variables, 2 themed values each,
hex-formatted, light≠dark), hex getter for both modes, apply
non-mutation + user-var-wins-on-collision + additive theme-axis
merge, hasSemanticPalette (empty / full / partial), and full
round-trip through the existing resolveVariableRef / resolveColorRef
paths for every palette name.
2026-04-22 10:40:00 +08:00
Fini 42b6492404 chore(tests): fix tsc errors in earlier test files
Two tsc errors surfaced by a later tsc run:

1. browser-image-search-fetcher.test.ts — imported `beforeEach`
   but never used; also `spy.mock.calls[0]` typed as empty tuple
   since vi.fn()'s signature isn't inferred. Cast via `unknown` +
   explicit tuple shape.

2. chart-builders-visual-smoke.test.ts — custom-dimensions case
   passed `width`/`height` to buildChartLine. The real shape is
   `point_spacing` + `chart_height`. Fixed the test to use the
   actual param names (test itself wasn't broken, just the type).
2026-04-22 10:35:00 +08:00
Fini cf3aa4240a test(pen-core): chart builders visual smoke (30 cases)
Geometric invariants for the three chart builders (bars/line/pie)
that a rendering failure would start from. Can't run Skia
headlessly in unit tests (CanvasKit WASM is heavy + GPU-context-
dependent), so this is the cheap smoke layer that catches shape-
level regressions before they reach the renderer. The app-level
debug_screenshot MCP tool gives us real visual regression on top.

Per chart type:
  - buildChartBars: 8 variants (default, all-equal, single, small
    values, large values, custom dims, zeros, empty-throws).
    Asserts one chart-bar per value, all dims finite.
  - buildChartLine: 9 variants (smooth, monotonic, flat, two-point,
    spike, fractional, custom dims, single-value, empty-throws).
    Asserts chart-line geometry present, coords finite.
  - buildChartPie: 9 variants (equal, skewed, two, many-thin,
    custom diameter, donut at 0.5 and 0.9 ratio, single 100%,
    empty-throws, all-zero-throws). Asserts one slice per value,
    startAngle/sweepAngle finite, sweepAngle positive, total
    sweep = 360°.

Cross-chart invariants: same input length → same geometry-child
count; all three types produce finite-coord trees on identical
input.

Empty-input behavior: all three builders throw with clear
messages. Test pins the throws as intended behavior — the
alternative (silently returning an empty tree) would let an
all-zeros dataset produce a "chart is there but invisible" UI
bug that's much harder to diagnose than the explicit throw.
2026-04-22 10:30:00 +08:00
Fini 6a2e80022c feat(ai-skills): mock-LLM local A/B harness (no API tokens)
Deterministic mock-LLM for local unit-level testing of the scorer
+ output-parser + apply pipeline — the same three components real
A/B runs use, minus the network.

mockLlmRaw(prompt, variant) returns the raw string a "well-
behaved" model would emit:
  - Treatment on obvious prompt → <op_tool>{...}</op_tool> naming
    expected_tool_if_any with empty args
  - Baseline → minimal batch_design DSL with a single frame role-
    stamped from must_contain_roles[0]
  - Optional + no hint → falls through to the baseline path

mockLlmParsed() skips the raw-string round-trip and returns a
ParsedOutput directly for tests pinning a specific kind. Both
respect (promptId, variant) overrides so tests can simulate
garbage / wrong-tool / empty outputs inline.

Integration test loads the real ab-v1 corpus from disk and
exercises every prompt through the full pipeline:
  corpus → mockLlmRaw → parseModelOutput → scoreRun → ScoreRow

Verifies all 4 routing outcomes (right-tool / wrong-tool /
fallback / garbage) classify correctly on mocked input. 17 test
cases; corpus sweep runs in ~6ms — fast enough to gate every
PR without slowing CI.

This is the prerequisite for future "real" A/B test runners: if
the harness misclassifies obvious mock inputs, no conclusion
from a real run would be trustworthy.
2026-04-22 10:10:00 +08:00
Fini 271ea1b317 feat(mcp): dispatcher metrics — per-tool call/error counters
In-memory counters for the element-tool dispatcher, exposed via
\`getElementToolMetric(name)\` / \`getAllElementToolMetrics()\` /
\`getTopElementToolCalls(n)\` / \`resetElementToolMetrics()\` in
packages/pen-mcp/src/metrics/.

\`handleElementToolCall\` now wraps the existing switch in a
try/record — every dispatch increments \`calls\`, thrown handlers
additionally bump \`errors\` and stash the last error message.
Unknown tool names still fire a counter (useful signal: "the AI
picked a tool we don't have").

Process-local / in-memory by design:
  - Test determinism: resetElementToolMetrics() in beforeEach
  - Matches stdio MCP server's one-client-one-server model
  - No persistence backend choice baked in — if we need
    cross-restart persistence later, a thin serializer drops on
    top without touching this API

Unlocks #92 Local A/B harness: feed a corpus through the MCP
server, read back getAllElementToolMetrics() to see which tools
the model actually picked vs what the corpus expected. Core
observability for non-Claude regression detection.
2026-04-22 10:05:00 +08:00
Fini 801ff4c532 test(pen-core): unicode normalization preservation (30 cases)
Pins that text-carrying element builders preserve the caller's
byte representation verbatim — no silent normal-form conversion,
no zero-width stripping, no fullwidth↔ASCII collapse.

Tests each of NFC/NFD/NFKC/NFKD forms through 6 representative
builders (heading / body-text / list-row / form-field / faq-item /
comment), plus 5 targeted fixtures:
  - Zero-width joiner mid-word ("emo‍ji" stays 6 codepoints)
  - BOM at string start
  - ZWJ emoji family sequence (4-person glyph)
  - Vietnamese combining-marks (NFC vs NFD both preserved as-is)
  - Halfwidth/fullwidth CJK distinction (NFKC would collapse
    fullwidth "A" to ASCII "A"; we assert the builder does NOT)

Why this matters: macOS ships filenames in NFD, Windows/web in
NFC; copy-paste carries any form; some CJK inputs emit
precomposed, others decomposed base+combining. Exact-match
lookups in external systems (especially emoji-less fallback
keys) break silently if the builder pre-empts the downstream
validator's normalization decision. Builders must pass through
bytes unmodified.
2026-04-22 10:00:00 +08:00
Fini 3ee85c1bc5 test(pen-mcp): large-scale batch_design stress (100-250 ops)
AI orchestrators occasionally emit very large batches (one sub-
agent producing a whole section in a single batch_design call).
Existing multi-line regression only covered pretty-printed JSON
in SINGLE ops — nothing pinned behavior when N itself grows.

5 scenarios:
  1. 100 sibling I() ops → all land, <5s wall-clock
  2. 200 sibling ops → 2x node count, <10s (catches O(n²) regressions)
  3. 30-level nested I() chain via parent_id threading
  4. 250 mixed ops (50 sections × 4 children) with parent refs
  5. Partial failure: 1 bogus parent_id among 100 good ops — good
     ones still land (don't let one bad op poison the batch)

Observed: 250-op mixed batch completes in ~42ms on an M-series
machine. Budgets are "reasonable" (5s / 10s / 15s), not "fast" —
they're meant to catch O(n²) regressions in the DSL parser / tree
insert / save loop, not enforce a perf target.
2026-04-22 09:55:00 +08:00
Fini df9f37522f test(pen-mcp): composition — N-tool real screens (6 scenarios)
Per-tool handler tests cover "this tool emits the correct shape"
individually. This file covers the next layer: can N element-tool
calls chain together into a realistic multi-section screen without
breaking tree invariants?

Scenarios (each spans multiple tool families to catch cross-
family regressions):
  1. Mobile settings — top_nav + 2 sections × 3 list_rows + bottom_nav (10 calls)
  2. Dashboard home — top_nav + stat_grid + section + 3 metric_comparisons + chart (7)
  3. Login form — heading + body + 2 form_fields + button + link (6)
  4. Profile + UGC — top_nav + avatar + heading + badge + 2 faq_items + action_menu (7)
  5. Listing — search + card_row + divider + empty_chart + date_picker + chip_input + pagination (7)
  6. parent_id threading invariant — nested insert actually lands under named parent

Each scenario asserts:
  - Every call emits a nodeId (no silent no-ops)
  - Final document parses as valid JSON with expected root children count
  - Every call's nodeId is findable in the saved tree
  - Every tool's canonical role survives post-save
  - parent_id threading works (child lands under named parent, not root)

This is the integration gate that catches "tool wiring works
individually but composes wrong" — the ghost regression that can
slip past per-tool tests.
2026-04-22 09:50:00 +08:00