Commit graph

300 commits

Author SHA1 Message Date
Fini e0b926c256 feat(ai): add_social_login_row_v0 — social auth button row (69th tool)
The "Continue with Google / Apple / Microsoft" row. Vertical default
(stacked full-width 48px buttons) or horizontal (compact icon-only
48×48 pills). Known provider names (google, apple, github, microsoft,
facebook, twitter, linkedin, discord, slack, gitlab, email, phone)
auto-map to lucide icons; `icon` param overrides for SSO/SAML/Okta.
2026-04-25 01:19:00 +08:00
Fini 127eccdefb feat(ai): add_stat_card_v0 — big-number KPI tile (68th tool)
Featured-metric dashboard tile: label (uppercase muted) above a
huge 32/700 primary value, with optional tone-colored delta line
and corner icon. Distinct from its neighbors:

  - add_stat_grid_v0   — multi-cell side-by-side, smaller values
  - add_metric_comparison_v0 — horizontal label+value+inline arrow
  - add_stat_card_v0 (this) — single featured metric, whole-card focus

Trend enum tones the delta line only (value stays slate-900):
  up    → #10B981 (emerald)
  down  → #EF4444 (red)
  flat  → #64748B (slate, default)

Wired through all standard points — schema into ext-3 (shortest
shard at 371 lines + new tool → 408, still comfortably under 800)
+ shim + SERVER_BUILDERS + parity CASES + contract allow-list +
elements.md decision tree + triggers + minimal usage.

Handler test covers 6 cases: minimal registration + defaults +
icon+delta+up-tone + down-tone + flat-tone default + width
clamp + bogus parent_id rejection.
2026-04-25 01:09:49 +08:00
Fini 23988b9e84 fix(ab-corpus): per-call timeout + progressive scores.jsonl writes
The first 12-prompt × 2-model × 2-variant sweep (48 API calls)
ran 29 minutes before I killed it. A Kimi call on the 10th
prompt hung indefinitely — no client-side timeout — and the
harness writes scores.jsonl + report.md ONLY at the end, so
partial progress was unrecoverable. Lost 9/12 completed prompts
because the aggregate step never ran.

Two fixes:

1. openai-compat.ts: AbortController with default 120s timeout
   (overridable via AB_CORPUS_CALL_TIMEOUT_MS env). When a call
   exceeds the budget, the harness catches the abort, records it
   as __HARNESS_ERROR__ (routing=garbage, M1=false), and moves on.
   Verified by dialing the timeout to 60s — GLM-5.1's first call
   took >60s, got aborted cleanly, run continued to completion
   instead of hanging.

2. run.ts: append each ScoreRow to scores.jsonl immediately after
   scoring. Truncate at start (so re-runs overwrite). Lost-work
   window now bounded to "the currently-executing API call," not
   "everything since the run started." report.md and report.json
   still write once at the end (aggregate needs the full set) but
   scores.jsonl alone is enough for any partial-run analysis.

Post-hardening validation (live 方舟 CP runs):
  - mobile-upload-dropzone → add_upload_dropzone_v0 ✓ right-tool
  - dashboard-dark-modal   → add_modal_shell_v1 (theme=dark) ✓

Second one is the first end-to-end proof that the v1 theme-aware
tool family routes correctly with a real LLM — GLM-5.1 inferred
\`theme: "dark"\` from the natural-language prompt.
2026-04-25 00:43:48 +08:00
Fini 54ee6e0ec7 feat(ab-corpus): --corpus flag + 5 new prompts for tools 63-67 + v1
Two independent changes rolled together since they both serve the
same goal — "can real LLMs actually route to the tools we shipped
today?":

  1. scripts/ab-corpus/run.ts gains a `--corpus` flag (ab-v0 |
     ab-v1, default ab-v0 for back-compat). The harness was
     hardcoded to ab-v0 — adding ab-v1 prompts was worthless
     without a way to run them. Validated 17 prompts × 2 models ×
     2 variants during a live 方舟-CP run.

  2. 5 new ab-v1 prompts cover the 2026-04-24 tool batch:
      - mobile-upload-dropzone.yaml        → add_upload_dropzone_v0
      - mobile-otp-verification.yaml       → add_otp_input_v0
      - mobile-file-attachment.yaml        → add_attachment_row_v0
      - mobile-chat-message.yaml           → add_chat_bubble_v0
      - dashboard-dark-modal.yaml          → add_modal_shell_v1

corpus-loader.test.ts bumps its count assertion 12 → 17 and
extends the tool-coverage set. Also loosens the regex to accept
`_v\d+$` (was `_v0$`) so add_modal_shell_v1 passes. No other
test file needed changes — the existing registry-parity and
mock-llm tests already use `_v\d+$` or the registry directly.

.gitignore gains:
  - .playwright-mcp/ (MCP Playwright session artifacts)
  - editor-*.png      (local verification screenshots)
  - scripts/ab-corpus/runs/  (live-run outputs / reports)

None of those belong in version control — they're artifacts
from local verification runs.
2026-04-25 00:23:11 +08:00
Fini 6d4bdb88da test(pen-mcp): support-chat composition scenario covering #63-#67 + v1
Extends element-tools-composition.test.ts with a 7th real-screen
scenario that exercises every tool added after the 62-tool mark:

  - add_top_nav_bar_v0         (existing anchor)
  - add_chat_bubble_v0         × 2  (left from-others + right from-self)
  - add_attachment_row_v0      (file on the self-message path)
  - add_upload_dropzone_v0     (drop area for screenshots)
  - add_chip_input_v0          (conversation tags)
  - add_action_menu_v0         (floating menu, open state)
  - add_modal_shell_v1         (theme="dark" confirm dialog)
  - add_otp_input_v0           (phone verification step)

8 tool calls, chained through the full MCP handler pipeline
(ensureParentExists → builder → assignIdsRecursively →
batch_design insert with rollback-on-failure → post-insert
landing check → save → re-read from disk). Asserts:

  - every call emits a nodeId (no silent no-ops)
  - final doc has exactly 9 root children
  - every tool's role marker survives round-trip save/load
  - modal v1 theme=dark → card fill #1E293B (not v0's #FFFFFF)
  - chat right-side bubble surface → #2563EB accent
  - OTP focused slot stroke → #2563EB accent

This is end-to-end at the DATA layer, not the visual layer.
Still not covered by any test:
  - Skia rendering (needs debug_screenshot against a live
    canvas)
  - Real LLM tool routing (needs ab-corpus harness with live
    API keys — route to 方舟 CP works but has not run since
    the routing fix)

Document-layer coverage is enough to rule out handler-pipeline
bugs (bad parent_id threading / cache-stale doc-state / silent
rollback swallowing); visual regressions require the next gate.
2026-04-22 22:31:49 +08:00
Fini 06362a4040 refactor(mcp): split ext-2 into ext-2 + ext-3 (keep shards under 800)
Codex stop-hook caught element-tool-defs-ext-2.ts at 832 lines
after add_chat_bubble_v0 landed there — 32 over the repo's 800-
line ceiling. Same trap the original single ext file hit at 1329
lines, same fix pattern: carve the second half into a new shard.

Split at `add_modal_shell_v1` (tool #16 of 24 in old ext-2):
  - ext-2 keeps tools 1-15 (calendar_grid through textarea) →
    482 lines
  - ext-3 (new) holds tools 16-24 (modal_shell_v1 through
    chat_bubble) → 371 lines

All registry shards now:
  base    683
  ext     671
  ext-2   482
  ext-3   371
  (props   26)
  (top    254)

element-tool-defs.ts concatenates all three ext shards into the
single ELEMENT_TOOL_DEFINITIONS — external API unchanged.
ELEMENT_TOOL_DEFINITIONS_EXT_3 is the new import; 67 tools
still resolve.

Header in ext-1 updated to reflect the three-way split + advise
"when ANY shard crosses 700, carve a ~5-tool chunk into the
shortest shard" so the next rebalance happens proactively instead
of after a stop-hook trip.
2026-04-22 22:12:26 +08:00
Fini 165fb94e44 feat(ai): add_chat_bubble_v0 — messaging bubble (67th tool)
Chat / messaging / customer-support UI message unit. Two variants
via the `side` enum:

  - side="left" (default): from-others bubble. Slate-100 fill,
    slate-900 text, alignItems=flex-start. Optional `author` text
    shown above the bubble (group-chat pattern).
  - side="right": from-self bubble. Accent-color fill (customizable
    via `accent_color`), white text, alignItems=flex-end. Author
    intentionally suppressed on this side — a self-bubble never
    carries "You:".

Optional `timestamp` below the bubble on either side.

Max-width mechanic: pen-core has no native max-width primitive, so
`max_width` becomes the bubble's fixed width (clamped 160..480).
Short messages get extra padding on one side — matches every real
chat client (iMessage / WhatsApp / Slack). Message text uses
`textGrowth: 'fixed-width'` + `width: 'fill_container'` to wrap
correctly inside the fixed-width surface.

Full wiring: pen-core builder + pen-mcp handler + schema (into
ext-2, shorter shard — 24 tools vs ext-1's 24 after this) + shim
+ SERVER_BUILDERS + parity CASES + contract allow-list +
elements.md decision tree + triggers + minimal usage for both
sides. Handler test covers 9 cases: registration, left defaults,
left+author, right with self-dropped-author, right+accent_color,
timestamp both sides, max_width clamps (low + high split into
separate tests to avoid cache-interference), textGrowth wiring,
bogus parent_id rejection.
2026-04-22 22:03:44 +08:00
Fini 49d7d1baf2 feat(ai): add_attachment_row_v0 — file attachment list unit (66th tool)
Fills another common UI gap: the "here's an already-uploaded file"
row you see in email composers, chat attachments, and form upload
summaries. Compact horizontal layout: type-icon + filename (bold) +
optional muted size string + optional right-side × remove affordance.

Structure: horizontal frame (slate-50 bg, cornerRadius=8) with
three children:
  1. attachment-icon — lucide file-* (caller picks: file / file-
     text / file-image / file-video / file-audio / file-archive /
     file-spreadsheet / file-code)
  2. attachment-meta — vertical frame with filename + optional size
  3. attachment-remove — × icon, suppressed via removable=false

Intentionally NOT embedding an upload-progress variant in v0. The
pen-core schema lacks percentage-width primitives, so a %-filled
progress bar would either need a fixed track width (brittle across
parents) or a caller-computed pixel value (awkward API). Callers
who need the uploading state compose `add_progress_bar_v0` directly
below the row — cleaner separation.

Wired through all standard points: schema into ext-1 (balanced
shards 23/23 after upload-dropzone landed there last commit) +
shim + SERVER_BUILDERS + parity CASES + contract allow-list +
elements.md decision tree + triggers + minimal usage.

Handler test covers 7 cases: registration + minimal (no size) +
size rendered + custom icon + removable=false + default icon +
bogus parent_id rejection.
2026-04-22 21:53:47 +08:00
Fini dc73885015 feat(ai): add_otp_input_v0 — verification code input (65th tool)
Fills the auth-flow gap: 2FA / PIN / phone-verification codes.
Horizontal row of N square slots (4..8), one digit per slot.
Renders three states per caller intent:

  - blank (no `digits`): all slots empty, `focused_index` marks
    the currently-typing slot with an accent-color 2px outline
  - partial: first M slots filled with digit text, slot M+1
    focused, rest empty
  - full: all N slots filled (final submittable state)

Filled slots get role=otp-slot-filled + slate-700 border + 20/600
digit text. Focused empty slot gets role=otp-slot-focused +
2px accent border. Blank unfocused slots get role=otp-slot +
1px slate-300 border.

Wired through all standard points — schema into ext-2 (shorter
shard) + shim + SERVER_BUILDERS + parity CASES + contract
allow-list + elements.md decision tree + triggers + minimal
usage for each state.

Handler test covers 8 cases: registration + defaults (6 blank
focused-first) + partial state / full state / length clamp low
(< 4 → 4) + length clamp high (> 8 → 8) + accent color override
+ bogus parent_id rejection.
2026-04-22 21:48:05 +08:00
Fini 45214efb70 feat(ai): add_upload_dropzone_v0 — file drop zone (64th tool)
Fills a real gap in the element-tool family: upload / drag-and-
drop surfaces. Dashed border + cloud icon + two-line instruction
("Drop files to upload" / "or click to browse") — the classic
pattern from every modern file-upload UI.

Deliberately structurally similar to add_empty_chart_v0 (dashed
border + icon + title/subtitle) but semantically distinct:
  - empty_chart = "chart widget will render when data arrives"
    (320×200, icon chart-typed)
  - upload_dropzone = "users drop files here" (480×200, icon
    semantic: upload-cloud / upload / file-up)

elements.md routes them by intent, and the tool descriptions
cross-reference each other to prevent the AI from picking the
wrong one on ambiguous prompts.

Wired through all the standard points per the add-new-tool
checklist: builder + handler + schema (into ext-1, the shorter
shard) + shim + SERVER_BUILDERS + parity CASES + contract
allow-list + elements.md decision tree + keyword triggers +
minimal usage. Handler test covers 5 cases: defaults, dashed
stroke, overrides, size clamping, bogus parent_id rejection.
2026-04-22 21:41:13 +08:00
Fini ed449bc896 refactor(mcp): split element-tool-defs-ext in half + extract shared props
[Codex P3] element-tool-defs-ext.ts had grown to 1329 lines —
over the repo's documented 800-line ceiling. Ironically the file
header comment claimed it existed to keep its parent under the
cap, but the shard itself had outgrown the limit.

Split into three files:

  - element-tool-def-props.ts (26 lines) — shared JSON-Schema
    fragments (schemaVersionProp / filePathProp / parentIdProp /
    pageIdProp) that every definition file uses. Deduplicating
    these unblocks the split cleanly.
  - element-tool-defs-ext.ts (596 lines) — first 22 tools
    (add_switch_v0 through add_segmented_control_v0 era). Imports
    the shared props.
  - element-tool-defs-ext-2.ts (744 lines, new) — remaining 22
    tools starting at add_calendar_grid_v0. Imports the shared
    props.

element-tool-defs.ts concatenates all three arrays into the single
ELEMENT_TOOL_DEFINITIONS registry — external API surface unchanged.

Header comments in both shards now document the split convention:
"pick whichever shard has fewer tools" when adding a new entry,
to keep the files balanced as the family grows toward ~100.

Incidentally the previous commit's file also carried the P2 fix
(pageId threading through the in-browser and HTTP DSL paths, so
multi-page docs land the generation on the ACTIVE page instead of
doc.pages[0]). Both touched the same file, didn't make sense to
split. Title-wise the previous commit is P1 but functionally it's
P1+P2.

All under the 800-line ceiling now:
  element-tool-defs-base.ts  683
  element-tool-defs-ext.ts   596
  element-tool-defs-ext-2.ts 744
  element-tool-defs.ts       240
  element-tool-def-props.ts   26
2026-04-22 11:20:00 +08:00
Fini 0ce113c733 fix(ai): dispatcher imports DSL executor via browser-safe subpath
[Codex P1] The browser-side element-tools-dispatcher imported
runBatchDesignDsl from the \`@zseven-w/pen-mcp\` package barrel.
That barrel re-exports node-only modules — document-manager,
log-utils, theme-presets — which import node:fs / node:path at
top level. Vite / esbuild resolve the barrel BEFORE tree-shaking
can drop those branches, so browser builds failed on unresolved
node built-ins.

Fix:
  - packages/pen-mcp/package.json: add \`./dsl\` subpath export
    pointing at tools/batch-design-dsl.ts — the pure executor
    file already guarded as browser-safe by the adjacent
    regression test.
  - apps/web dispatcher: switch import to
    \`@zseven-w/pen-mcp/dsl\`. No other changes — the re-exported
    symbols (runBatchDesignDsl / OpResult / ImageSearchFetcher /
    RunBatchDesignDslOptions) are identical shape.
  - batch-design-dsl-browser-safe.test.ts: add an assertion that
    package.json's exports field preserves the \`./dsl\` key
    pointing at the expected file. Without this, silently
    removing the subpath would re-introduce the browser-breaking
    resolution path.

The package barrel keeps its current export of runBatchDesignDsl
too (a few internal test files still import from it). Browser
callers should migrate to \`@zseven-w/pen-mcp/dsl\` per the JSDoc
note now in the dispatcher.
2026-04-22 11:15:00 +08:00
Fini e10b3a37c9 fix(ab-corpus): kimi-2.6 alias fell through to Bailian instead of Ark
The Ark router regex was `/^kimi-k2\.6/i` — required the `k`
prefix. But mapKimiArkId's alias list accepted both `kimi-k2.6`
AND `kimi-2.6` (no-prefix form). Result: `kimi-2.6` failed the
Ark regex, fell through to the generic `/^kimi/i` branch, got
routed to Bailian — which doesn't host K2.6. Bailian would
return HTTP 400 "model not supported" with no hint that the id
belonged on Ark.

Fix: Ark router regex now `/^kimi-k?2\.6(-ark)?$/i` — optional
`k` prefix + optional `-ark` suffix, anchored at end to prevent
accidentally matching a hypothetical later version. mapKimiArkId
normalizes all four accepted aliases (kimi-k2.6, kimi-2.6,
kimi-k2.6-ark, kimi-2.6-ark) to the canonical on-Ark id.

Same latent bug fixed on the glm-5.1 route: regex `/^glm-5\.1/i`
would prefix-match a hypothetical `glm-5.10` and wrongly route it
to Ark. Tightened to `/^glm-5\.1(-coding|-ark)?$/i` with the same
anchored-end + suffix-allowlist pattern.

Caught by Codex stop-hook review during 2026-04-22 session.
2026-04-22 11:10:00 +08:00
Fini 5538350a2f chore(ab-corpus): route glm-5.1 + kimi-k2.6 through 方舟 CP (Ark)
Volcengine 方舟 (Ark) added GLM-5.1 and Kimi-K2.6 to its coding
plan on 2026-04-22 — single ARK_CODING_KEY covers both. Harness
now prefers this route over the previous paths:

  - glm-5.1 was routed to clients/glm.ts (GLM official CP via
    open.bigmodel.cn with GLM_OFFICIAL_CODING_KEY). Now routed to
    new clients/ark.ts. The old glm.ts file is kept on disk for
    historical comparison but not wired into the default router —
    callers who want to A/B the old GLM-official path vs. new Ark
    path can import callGlm directly.
  - kimi-k2.6 is new — added as a dedicated router branch above
    the kimi-k2.5 (bailian) branch so the version-specific match
    lands on Ark.

Old kimi-k2.5 continues to route through clients/bailian.ts
(DashScope aggregator) for continuity with earlier A/B runs.

Key management (unchanged from the harness convention):
  - ARK_CODING_KEY — Volcengine 方舟 CP UUID format key. Export
    in shell before running --live; never committed.
  - Existing MINIMAX_API_KEY / GLM_OFFICIAL_CODING_KEY /
    DASHSCOPE_BAILIAN_CODING_KEY all still honored for their
    respective routes.

Throw message updated so missing-key errors surface the correct
env var for each route.
2026-04-22 11:05:00 +08:00
Fini dd5d156160 feat(ai): add_modal_shell_v1 — first theme-aware MCP tool (63rd, v1-family debut)
Wires the buildModalShellV1 pen-core builder through the full MCP
toolchain — handler + schema + dispatch + shim + SERVER_BUILDERS
+ parity CASES + handler tests + elements.md skill — so external
MCP clients (Claude Code / Codex / Gemini CLI) can call
add_modal_shell_v1 as a first-class tool alongside the 62 v0
tools.

Three theme variants exposed via `theme` param (enum [light, dark,
system]):
  - omitted / `'light'`: byte-parity with add_modal_shell_v0
    (same hex, same structure, same role tree)
  - `'dark'`: hardcoded dark palette (#1E293B card, #F1F5F9 title,
    #94A3B8 muted). No \$refs needed.
  - `'system'`: emits \$color-surface / \$color-text-primary /
    \$color-text-muted refs. Caller MUST have run
    applySemanticPalette(doc) first or refs resolve to undefined
    (documented in schema description).

Scrim stays #000000 in ALL themes — modal backdrops are a dim
effect, not a themeable surface. Pinned by handler test.

9 handler test cases cover: registration + schema shape +
required[title] + theme variants + scrim invariant + bogus
parent_id rejection. Parity test added a CASES entry with
\`theme:'dark'\` args (exercises the theme branch in both
shim and server paths).

elements.md gained:
  - §46b decision-tree entry pointing to the v1 variant
  - Trigger list entry for dark-mode / theme-aware prompts
  - Minimal usage showing \`theme:'dark'\` + \`theme:'system'\`

This is the reference implementation for the remaining 9 theme-
aware v1 tools in the top-10 offenders list (empty-chart,
chip-input, toast, pagination, notification-row, image-placeholder,
faq-item, comment, checkbox — per dark-theme-audit §offenders).
2026-04-22 11:00:00 +08:00
Fini 315ded51ad chore(ai): generalize element-tool name regex to accept _v\d+
Five regex sites across pen-ai-skills + pen-mcp + apps/web were
anchored at `_v0$`, blocking the _v1 family from being recognized
as element tools:

  - packages/pen-ai-skills/src/corpus/output-parser.ts
    ELEMENT_TOOL_NAME_RE (filters tool_call outputs in A/B
    scorer)
  - apps/web/src/services/ai/design-parser.ts:106 (embedded
    orchestrator dispatch)
  - packages/pen-mcp/src/__tests__/design-prompt-elements.test.ts
    (×2 — stale-integration guard for elements.md)
  - packages/pen-mcp/src/__tests__/element-tool-registry-parity.test.ts
    ("every tool name matches convention" — renamed to _vN)
  - apps/web/src/services/ai/__tests__/element-tools-dispatcher.test.ts
    (drift guard for SUPPORTED_EMBEDDED_ELEMENT_TOOLS)

All now accept /^add_[a-z_]+_v\d+$/. Registry-parity test's
expectedBuilder mapping already handled both v0 (strip suffix →
buildModalShell) and v1+ (preserve → buildModalShellV1) via the
existing `.replace(/_v0$/, '')` — no change there.

Prerequisite for landing add_modal_shell_v1 as a first-class MCP
tool in the next commit.
2026-04-22 10:55:00 +08:00
Fini b07b970daa feat(pen-core): buildModalShellV1 — first theme-aware v1 element (proof of chain)
Ships the first theme-aware element builder demonstrating the v1
contract end-to-end. `buildModalShellV1({ title, theme })` accepts
three theme variants:

  - `'light'` (default): byte-parity with buildModalShell v0 —
    same hex literals, same role tree, same structural shape.
    Structural test asserts stripIds(v0) === stripIds(v1) for
    the default-theme path.
  - `'dark'`: hardcoded dark-palette hex (#1E293B card, #F1F5F9
    title, #94A3B8 muted text). No \$refs — this path is for
    callers who want a dark modal without the theme-switching
    infrastructure.
  - `'system'`: emits \$color-surface / \$color-text-primary /
    \$color-text-muted refs. Renders track \`themes.Mode\` at
    paint time. Requires \`applySemanticPalette(doc)\` to have
    been seeded; if not, refs resolve to undefined (caller's
    responsibility per the v1 contract).

End-to-end tests validate the 'system' path's round-trip through
resolveColorRef for both Light (→ #FFFFFF) and Dark (→ #1E293B)
modes. That's the full chain working:

  buildModalShellV1({theme:'system'}) → tree with \$refs
  → applySemanticPalette(doc) → palette seeded
  → resolveColorRef(ref, doc.variables, {Mode:'Dark'}) → hex

One intentional design note tested: scrim stays #000000 in BOTH
light and dark themes. Modal backdrops are a "dim everything
below" effect, not a themeable surface — dimming a dark surface
with a dark color is a better visual than a themed shade.

18 tests. v0 byte-parity verified against the existing
buildModalShell for full structural equality. This is the
reference implementation for all subsequent v1 tools (top-10
offenders per the dark-theme audit).
2026-04-22 10:50:00 +08:00
Fini 19a1b13764 docs(ai-skills): teach variables.md about the 14 semantic tokens
Extends the generation-phase variables skill with a table of the
14 semantic palette tokens that \`applySemanticPalette(doc)\`
seeds. Models consuming this skill learn:

  1. The exact token names + their light/dark resolved values
     (can cross-reference against what the user's document
     actually has via \`hasSemanticPalette\`)
  2. When to PREFER \`\$color-*\` refs over hex literals (theme-
     aware intent: dark-mode design, system-follow apps, user-
     toggleable themes)
  3. When to FALL BACK to hex (default createEmptyDocument state
     where the palette isn't seeded)
  4. That semantic tokens override theme — \`\$color-success\`
     stays green in both light and dark modes because "green"
     is the semantic signal, not a visual choice

Without this guidance, an AI asked for a dark-themed dashboard
could either: (a) emit hex literals that don't track theme
(defeats the purpose), or (b) emit \$color-* refs blindly into
a doc that lacks the palette (resolves to undefined, renders as
raw string). The table + fallback rule close both gaps.
2026-04-22 10:45:00 +08:00
Fini ce32f9f572 feat(pen-core): 14-variable semantic palette (unblocks v1 theme-aware tools)
Ships the canonical 14-token palette called for in the dark-theme
audit (openpencil-docs/superpowers/notes/2026-04-22-dark-theme-
defaults-audit.md §role clusters). Every token has paired Light +
Dark values on a single `Mode` theme axis.

API surface:
  - getSemanticPalette() → {themes, variables} for merge
  - getSemanticPaletteHex(mode='Light') → flat Record<name, hex>
  - applySemanticPalette(doc) → non-destructive merge (user-
    defined variables + theme axes WIN on collision; palette is
    purely additive)
  - hasSemanticPalette(doc) → runtime check for v1 tools before
    emitting \$color-* refs
  - getSemanticPaletteDescription(name) → human-readable string
    for token-picker UI
  - SEMANTIC_PALETTE_NAMES + theme-axis constants exported

The 14 tokens:
  - Surfaces: color-surface, color-surface-2, color-surface-3,
    color-bg-deep
  - Borders: color-border, color-border-strong
  - Text: color-text-primary, color-text-body, color-text-muted,
    color-text-subtle
  - Semantic: color-accent, color-destructive, color-success
  - Other: color-scrim (with alpha for modal backdrop)

Intentionally NOT wired into createEmptyDocument(). Seeding by
default would alter every existing document on re-save and
violate the v0 byte-parity contract (the whole point of the
audit). v1 tools will call applySemanticPalette(doc) as a pre-
flight, OR the app shell offers a "enable dark theme" user
action that triggers the apply.

29 tests cover: palette shape (14 variables, 2 themed values each,
hex-formatted, light≠dark), hex getter for both modes, apply
non-mutation + user-var-wins-on-collision + additive theme-axis
merge, hasSemanticPalette (empty / full / partial), and full
round-trip through the existing resolveVariableRef / resolveColorRef
paths for every palette name.
2026-04-22 10:40:00 +08:00
Fini 42b6492404 chore(tests): fix tsc errors in earlier test files
Two tsc errors surfaced by a later tsc run:

1. browser-image-search-fetcher.test.ts — imported `beforeEach`
   but never used; also `spy.mock.calls[0]` typed as empty tuple
   since vi.fn()'s signature isn't inferred. Cast via `unknown` +
   explicit tuple shape.

2. chart-builders-visual-smoke.test.ts — custom-dimensions case
   passed `width`/`height` to buildChartLine. The real shape is
   `point_spacing` + `chart_height`. Fixed the test to use the
   actual param names (test itself wasn't broken, just the type).
2026-04-22 10:35:00 +08:00
Fini cf3aa4240a test(pen-core): chart builders visual smoke (30 cases)
Geometric invariants for the three chart builders (bars/line/pie)
that a rendering failure would start from. Can't run Skia
headlessly in unit tests (CanvasKit WASM is heavy + GPU-context-
dependent), so this is the cheap smoke layer that catches shape-
level regressions before they reach the renderer. The app-level
debug_screenshot MCP tool gives us real visual regression on top.

Per chart type:
  - buildChartBars: 8 variants (default, all-equal, single, small
    values, large values, custom dims, zeros, empty-throws).
    Asserts one chart-bar per value, all dims finite.
  - buildChartLine: 9 variants (smooth, monotonic, flat, two-point,
    spike, fractional, custom dims, single-value, empty-throws).
    Asserts chart-line geometry present, coords finite.
  - buildChartPie: 9 variants (equal, skewed, two, many-thin,
    custom diameter, donut at 0.5 and 0.9 ratio, single 100%,
    empty-throws, all-zero-throws). Asserts one slice per value,
    startAngle/sweepAngle finite, sweepAngle positive, total
    sweep = 360°.

Cross-chart invariants: same input length → same geometry-child
count; all three types produce finite-coord trees on identical
input.

Empty-input behavior: all three builders throw with clear
messages. Test pins the throws as intended behavior — the
alternative (silently returning an empty tree) would let an
all-zeros dataset produce a "chart is there but invisible" UI
bug that's much harder to diagnose than the explicit throw.
2026-04-22 10:30:00 +08:00
Fini 90d409207c feat(ai): browser-side G() image-search fetcher (relative URL)
Adds makeBrowserImageSearchFetcher() — a browser-safe
ImageSearchFetcher that POSTs to /api/ai/image-search (relative,
same-origin) for inline G() resolution in the batch_design DSL.

The server-side fetcher uses absolute URL via getSyncUrl() — not
applicable in the browser, where fetch resolves relative paths
against window.origin. This helper mirrors the server-side shape
but drops the sync-URL dependency, so any browser caller that
wants inline image-search can opt in.

NOT wired into the default dispatch path. The existing behavior
(applyBatchDesignDsl omits the fetcher → empty src → enriched
asynchronously by scanAndFillImages) stays the default because
per-G() round-trips would blow up latency on batches with many
images. This helper is opt-in for callers that accept that
trade-off (composition smoke tests, user-opt-in preview modes).

Never throws. Returns null on: empty query, network error,
non-ok response, non-JSON body, missing/empty results, invalid
shape. Callers can drop it in without try/catch wrappers.

16 test cases covering happy path (URL/body/headers/first-of-
many) + 11 failure modes + 1 concurrency invariant.
2026-04-22 10:25:00 +08:00
Fini 59e6c09b18 test(ai): orchestrator cancel mid-batch — structural + behavioral (12 cases)
Two-layered coverage for the abort-signal pattern that's the sole
mechanism preventing the orchestrator from continuing to call the
LLM after the user hits Stop.

A. **Structural** (grep-level): pins that every sequential / per-
   screen-group loop in orchestrator-sub-agent.ts AND the "no nodes"
   throw + validation gate in orchestrator.ts check abortSignal at
   the right points. A new loop that forgets the guard silently
   wastes tokens AND mutates canvas post-stop; this grep-style
   check catches it at unit-test time.

B. **Behavioral** (pure): stubbed sequencer that mirrors the
   actual loop body, exercised with 7 scenarios:
     - no abort → all run
     - abort during iteration N → N+1 onwards never run
     - pre-aborted signal → zero execution
     - post-completion abort → no-op
     - no signal → normal behavior
     - concurrent abort during one worker → other workers stop at
       next iteration-start check
     - pre-aborted signal with concurrent workers → both skip

Real orchestrator logic lives in orchestrator-sub-agent.ts but is
wrapped in LLM-calling code that's expensive to mock. The
behavioral stub is pattern-isomorphic: if the stub works here, the
real loop does too; if the real loop changes shape, the structural
check catches the drift.
2026-04-22 10:20:00 +08:00
Fini a22dc6a639 test(ai): model-tier × elements-skill injection e2e (18 cases)
Existing model-profiles-element-tools.test.ts only covers the
upstream flag boolean (\`needsElementTools\`). This file pins the
downstream filter — \`compactSubAgentSkills\` — specifically
around the elements skill, the spot where a regression would
silently drop elements.md content from the sub-agent prompt even
though VITE_ENABLE_ELEMENT_TOOLS=1 is set.

Covers:
  - basic tier: allow-list preserves elements (mobile + non-mobile)
  - basic + reducedComplexity: elements INTENTIONALLY dropped for
    retry path (~17k char savings when fallback to batch_design)
  - standard / full: elements always passes through
  - jsonl-format vs jsonl-format-simplified conflict: simplified
    wins, elements survives both resolutions
  - Screen-type gates (mobile-app vs landing-page/copywriting/
    anti-slop) are orthogonal to elements — elements survives
    every combination
  - Determinism: same input → same output; original array not
    mutated

Also exercises edge cases: empty skill list, elements-only list,
unknown-name skill at each tier.
2026-04-22 10:15:00 +08:00
Fini 6a2e80022c feat(ai-skills): mock-LLM local A/B harness (no API tokens)
Deterministic mock-LLM for local unit-level testing of the scorer
+ output-parser + apply pipeline — the same three components real
A/B runs use, minus the network.

mockLlmRaw(prompt, variant) returns the raw string a "well-
behaved" model would emit:
  - Treatment on obvious prompt → <op_tool>{...}</op_tool> naming
    expected_tool_if_any with empty args
  - Baseline → minimal batch_design DSL with a single frame role-
    stamped from must_contain_roles[0]
  - Optional + no hint → falls through to the baseline path

mockLlmParsed() skips the raw-string round-trip and returns a
ParsedOutput directly for tests pinning a specific kind. Both
respect (promptId, variant) overrides so tests can simulate
garbage / wrong-tool / empty outputs inline.

Integration test loads the real ab-v1 corpus from disk and
exercises every prompt through the full pipeline:
  corpus → mockLlmRaw → parseModelOutput → scoreRun → ScoreRow

Verifies all 4 routing outcomes (right-tool / wrong-tool /
fallback / garbage) classify correctly on mocked input. 17 test
cases; corpus sweep runs in ~6ms — fast enough to gate every
PR without slowing CI.

This is the prerequisite for future "real" A/B test runners: if
the harness misclassifies obvious mock inputs, no conclusion
from a real run would be trustworthy.
2026-04-22 10:10:00 +08:00
Fini 271ea1b317 feat(mcp): dispatcher metrics — per-tool call/error counters
In-memory counters for the element-tool dispatcher, exposed via
\`getElementToolMetric(name)\` / \`getAllElementToolMetrics()\` /
\`getTopElementToolCalls(n)\` / \`resetElementToolMetrics()\` in
packages/pen-mcp/src/metrics/.

\`handleElementToolCall\` now wraps the existing switch in a
try/record — every dispatch increments \`calls\`, thrown handlers
additionally bump \`errors\` and stash the last error message.
Unknown tool names still fire a counter (useful signal: "the AI
picked a tool we don't have").

Process-local / in-memory by design:
  - Test determinism: resetElementToolMetrics() in beforeEach
  - Matches stdio MCP server's one-client-one-server model
  - No persistence backend choice baked in — if we need
    cross-restart persistence later, a thin serializer drops on
    top without touching this API

Unlocks #92 Local A/B harness: feed a corpus through the MCP
server, read back getAllElementToolMetrics() to see which tools
the model actually picked vs what the corpus expected. Core
observability for non-Claude regression detection.
2026-04-22 10:05:00 +08:00
Fini 801ff4c532 test(pen-core): unicode normalization preservation (30 cases)
Pins that text-carrying element builders preserve the caller's
byte representation verbatim — no silent normal-form conversion,
no zero-width stripping, no fullwidth↔ASCII collapse.

Tests each of NFC/NFD/NFKC/NFKD forms through 6 representative
builders (heading / body-text / list-row / form-field / faq-item /
comment), plus 5 targeted fixtures:
  - Zero-width joiner mid-word ("emo‍ji" stays 6 codepoints)
  - BOM at string start
  - ZWJ emoji family sequence (4-person glyph)
  - Vietnamese combining-marks (NFC vs NFD both preserved as-is)
  - Halfwidth/fullwidth CJK distinction (NFKC would collapse
    fullwidth "A" to ASCII "A"; we assert the builder does NOT)

Why this matters: macOS ships filenames in NFD, Windows/web in
NFC; copy-paste carries any form; some CJK inputs emit
precomposed, others decomposed base+combining. Exact-match
lookups in external systems (especially emoji-less fallback
keys) break silently if the builder pre-empts the downstream
validator's normalization decision. Builders must pass through
bytes unmodified.
2026-04-22 10:00:00 +08:00
Fini 3ee85c1bc5 test(pen-mcp): large-scale batch_design stress (100-250 ops)
AI orchestrators occasionally emit very large batches (one sub-
agent producing a whole section in a single batch_design call).
Existing multi-line regression only covered pretty-printed JSON
in SINGLE ops — nothing pinned behavior when N itself grows.

5 scenarios:
  1. 100 sibling I() ops → all land, <5s wall-clock
  2. 200 sibling ops → 2x node count, <10s (catches O(n²) regressions)
  3. 30-level nested I() chain via parent_id threading
  4. 250 mixed ops (50 sections × 4 children) with parent refs
  5. Partial failure: 1 bogus parent_id among 100 good ops — good
     ones still land (don't let one bad op poison the batch)

Observed: 250-op mixed batch completes in ~42ms on an M-series
machine. Budgets are "reasonable" (5s / 10s / 15s), not "fast" —
they're meant to catch O(n²) regressions in the DSL parser / tree
insert / save loop, not enforce a perf target.
2026-04-22 09:55:00 +08:00
Fini df9f37522f test(pen-mcp): composition — N-tool real screens (6 scenarios)
Per-tool handler tests cover "this tool emits the correct shape"
individually. This file covers the next layer: can N element-tool
calls chain together into a realistic multi-section screen without
breaking tree invariants?

Scenarios (each spans multiple tool families to catch cross-
family regressions):
  1. Mobile settings — top_nav + 2 sections × 3 list_rows + bottom_nav (10 calls)
  2. Dashboard home — top_nav + stat_grid + section + 3 metric_comparisons + chart (7)
  3. Login form — heading + body + 2 form_fields + button + link (6)
  4. Profile + UGC — top_nav + avatar + heading + badge + 2 faq_items + action_menu (7)
  5. Listing — search + card_row + divider + empty_chart + date_picker + chip_input + pagination (7)
  6. parent_id threading invariant — nested insert actually lands under named parent

Each scenario asserts:
  - Every call emits a nodeId (no silent no-ops)
  - Final document parses as valid JSON with expected root children count
  - Every call's nodeId is findable in the saved tree
  - Every tool's canonical role survives post-save
  - parent_id threading works (child lands under named parent, not root)

This is the integration gate that catches "tool wiring works
individually but composes wrong" — the ghost regression that can
slip past per-tool tests.
2026-04-22 09:50:00 +08:00
Fini 4a75b53c45 feat(ai): add_date_picker_v0 — date input closed state (62nd tool)
Adds an N-tool for the labeled date input + calendar-icon trigger.
Emits ONLY the CLOSED state; the open month grid lives in
add_calendar_grid_v0 and is typically shown inside a popover,
not stacked directly below. Two visual states:

- placeholder (no value): slate-400 "Select date" + calendar icon
- populated (value): slate-900 date text + calendar icon

`clearable: true` adds a small X affordance to the right of the
value (only when value is present — no-op for the placeholder
state since there's nothing to clear). `required: true` appends
" *" to the label.

Keeping closed + open as separate tools is intentional: AI specs
often ask for only the closed trigger inside a form, and a single
combined tool would either force an unwanted grid or require a
mode flag that splits the parameter surface. Separate narrow tools
compose cleanly via batch_design when a designer DOES want both.

Wired through all 3 paths + parity/contract/design-prompt tests.
Handler test covers 7 cases: placeholder state fills, populated
value fills, clearable X behaviors (both present + absent value),
custom placeholder override, required marker, bogus parent_id
rejection.
2026-04-22 09:45:00 +08:00
Fini 1b4f362d59 feat(ai): add_action_menu_v0 — context/kebab dropdown panel (61st tool)
Adds an N-tool for the floating card that drops from a "⋯ more"
button or appears on right-click. Emits the OPEN state: vertical
stack of padded icon+label rows in a white card with subtle stroke
and shadow. Positioning and show/hide are caller concerns (same
philosophy as add_modal_shell_v0 / add_toast_v0).

Destructive items (destructive=true) render in red with role
`action-menu-item-destructive` so renderers can style the hover
state separately. divider_before=true on any item (except first,
where it's ignored) inserts a 1px hairline above — useful for
"Edit / Share / Report / Delete" grouping patterns.

Wired through all 3 paths + parity/contract/design-prompt tests.
Handler test covers 7 cases: simple list, destructive red fill,
divider between groups, leading-divider ignored, label-only no
icon, width clamp, bogus parent_id rejection.
2026-04-22 09:40:00 +08:00
Fini 97343c1d74 feat(ai): add_empty_chart_v0 — chart-slot empty state (60th tool)
Adds an N-tool for the "no data yet" tile that sits in the exact
footprint where a real chart would go. Default 320×200 matches the
line/bar chart default footprint; dashed border + slate-50 fill +
slate icon signal "chart slot, currently empty". Caller can hint
at the widget type via icon ("line-chart" / "pie-chart" / default
"bar-chart-2").

Intentionally separate from add_empty_state_v0 — that tool is for
inbox/onboarding/no-results full-page empties (has optional CTA,
no dashed border). add_empty_chart_v0 reads as "chart widget is
live, just lacks data yet" rather than "nothing to show on this
screen at all".

Wired through all 3 paths + parity/contract/design-prompt tests.
Handler test covers 6 cases: defaults + dashed stroke + icon
override + size clamping + title/subtitle override + bogus
parent_id rejection.
2026-04-22 09:35:00 +08:00
Fini 6102d54cd9 feat(ai): add_chip_input_v0 — tag / multi-select input (59th tool)
Adds an N-tool for the "variable-N pill-plus-cursor" pattern: a
labeled form control that holds N removable tag pills followed by
an inline placeholder caret. Wrap layout (layoutWrap=wrap) so
chips flow onto additional rows as they accumulate — a horizontal
fit_content row would clip after ~4 chips.

Each chip: pill (cornerRadius=16, slate-100 fill, padding 10/4/6/6
L/R/T/B) + label + 14×14 lucide "x". Default caret placeholder is
"Add tag…" when chips is empty; caller overrides with `placeholder`
(e.g. "Enter emails" for recipient lists).

Wired through all 3 paths with matching handler test (7 cases) +
parity tests auto-picking-up the entry. elements.md: decision tree
§54, trigger list, minimal usage (populated + empty).
2026-04-22 09:30:00 +08:00
Fini e948a081d2 feat(ai): add_faq_item_v0 — accordion/FAQ item (58th tool)
Adds an N-tool for one row in a FAQ list. Collapsed (default):
bold question + chevron-right header. Expanded (expanded=true):
chevron-down + multi-line answer paragraph beneath. Optional
show_divider draws a 1px slate-200 rectangle at the bottom for
visual separation between items (no implicit padding — caller
stacks in a vertical parent).

Wired through all 3 paths (pen-core buildFaqItem + pen-mcp handler
+ apps/web shim + Nitro SERVER_BUILDERS); both parity tests and
the stale-integration guard auto-pick-up the entry. Handler test
covers 5 cases: collapsed default, expanded with answer, expanded
without answer (guards against undefined), show_divider hairline,
bogus parent_id rejection.
2026-04-22 09:25:00 +08:00
Fini fa69986fc8 feat(ai): add_pagination_v0 — Google-style pagination bar (57th tool)
Adds an N-tool for list/table footer pagination: row of page-
number pills flanked by optional prev/next chevron buttons. Active
page renders filled with the accent color, inactive pages are
ghost. Long ranges collapse with "…" Google-style (always show 1
and total, plus a ±siblings window around current).

Wired through all 3 paths: pen-core buildPagination + pen-mcp
handler + apps/web browser shim + Nitro SERVER_BUILDERS. Both
parity tests (shim-server-parity, element-tool-registry-parity)
pick up the entry automatically. Handler test covers 7 cases:
small range no ellipsis, 10-page ellipsis, start-edge current,
accent override, show_arrows=false, total=1 single pill, bogus
parent_id rejection.

Also updates packages/pen-ai-skills/skills/phases/generation/
elements.md: decision tree §52, keyword triggers (pagination /
page nav / 分页 / 分页条), minimal usage example. The stale-
integration guard (design-prompt-elements.test.ts) now passes.
2026-04-22 09:20:00 +08:00
Fini c9620b42df test(pen-core): lock 10k/100k-char boundary for text-carrying builders
Feeds a 10,200-char Lorem-ipsum through every text-accepting
element builder (heading, body-text, list-row subtitle, form-field
placeholder, textarea placeholder, alert, toast, quote-block,
tooltip, comment body, notification-row body, modal-shell subtitle)
and asserts: build doesn't throw, computeLayoutPositions doesn't
throw, no NaN/Infinity coords leak in, and content is preserved
verbatim (no silent truncation inside the builder — that's a
caller/renderer concern).

Also pins a 100k-char stress on buildHeading + buildBodyText with
a <100ms budget — catches accidental O(n²) in any character scan.
2026-04-22 09:15:00 +08:00
Fini 194621b474 test(ai): dedicated regression for applyNoEmojiIconHeuristic behavior
Pins the 2026-04-22 finding (element-tool-round-trip surfaced the
behavior indirectly) as a focused unit test. The heuristic runs
inside insertStreamingNode → applyGenerationHeuristics for every
text node and:
  1. strips emojis (EMOJI_REGEX)
  2. collapses 2+ whitespace to 1
  3. trims leading/trailing whitespace

Coverage:
- Emoji scrubbing: single, multiple, leading, trailing positions
- Preservation of: ASCII, CJK (zh/ja/ko), Arabic RTL, unicode
  punctuation (em-dash, ellipsis, curly quotes), arrows
- Non-text nodes (frame, icon_font) skipped entirely
- Empty / missing content no-op
- Emoji-only content → converted to `path` (fallback icon geometry),
  the behavior branch that downstream code depends on

16 tests, closes #102.
2026-04-22 09:10:00 +08:00
Fini 6e1d523514 feat(ai): add_video_placeholder_v0 — video embed placeholder (56th)
Dark slate (#334155) box + centered white play icon + optional
caption. Default 320×180 for 16:9. The "future video embed"
affordance — semantically distinct from add_image_placeholder_v0:
dark bg + play icon reads as "play me later", not "picture coming".

Play affordance is a lucide `play` icon_font, NEVER a hand-drawn
path triangle (classic LLM anti-pattern for video placeholders).
Regression test locks that invariant.

Wired across all three paths + elements.md entry (44b) + keyword
map + example + parametric test CASES in 8 files. Handler test
has 6 assertions including the path-vs-icon_font anti-pattern
guard.

Tool count: 55 → 56. Test count: 3225 → 3241 (+16).
2026-04-22 09:05:00 +08:00
Fini 053e9d0b1f feat(ai): add_metric_comparison_v0 + add_notification_row_v0 (54th, 55th)
- add_metric_comparison_v0: KPI cell with trend. label above + big
  value + optional arrow icon + change amount. trend enum (up/down/
  flat) drives arrow icon (trending-up/-down/minus) + color (emerald/
  red/slate). Distinct from add_metric_row_v0 (scroll row of label+
  value cells without trend affordance). Required: label + value.
- add_notification_row_v0: leading icon + (title + optional
  timestamp + optional unread red dot) header + optional body
  preview. Distinct from add_list_row_v0 which has no timestamp
  or unread affordance. Required: title only.

Both wired across all three paths + elements.md ("Analytics / KPIs"
and "Notifications" sections) + keyword map + examples + 2 corpus
prompts in ab-v1/ (dashboard-revenue-trend, mobile-notification-item).

Fixed in same turn: ab-v1 file that was created with the wrong
filename got renamed.

Tool count: 53 → 55. Test count: 3180 → 3225 (+45).
2026-04-22 09:00:00 +08:00
Fini fb306dcf79 feat(ai): add_spinner_v0 + add_tooltip_v0 (52nd, 53rd tools)
- add_spinner_v0: static loading spinner — full ring (track) + 270°
  arc (active). Sits at size=32 default, clamped 16..128. Two
  ellipses at SAME origin with DIFFERENT sweep ranges — NOT the
  "stacked ellipses for ring" anti-pattern (rewriteLlmAntiPatterns
  only fires when both are full-sweep duplicates).
- add_tooltip_v0: small dark pill (#111827) + white text for
  hover hints. position param ("top"/"bottom"/"left"/"right")
  encodes a role hint (`tooltip-top` etc.) for downstream position
  logic; visual body is identical. NO arrow pointer (pen-core has
  no clean triangle primitive — caller composes via batch_design
  rectangle + rotate if needed).

Wired across all three paths. elements.md adds "Feedback / loading"
section (#48, #49). ab-v1 corpus +1 prompt (mobile-help-tooltip);
spinner omitted from corpus for now — the prompt wording is too
ambiguous for "obvious" difficulty.

Tool count: 51 → 53. Test count: 3138 → 3180 (+42).
2026-04-22 08:55:00 +08:00
Fini baa414bded feat(ai): add_status_badge_v0 — semantic status indicator (51st tool)
Small colored dot + short label: "● Online" / "● Busy" / "● Error"
pattern. Distinguished from the more general add_badge_v0 (just a
pill label) by always having a dot.

tone enum picks dot color:
- success → emerald #10B981
- warning → amber   #F59E0B
- error   → red     #EF4444
- info    → blue    #3B82F6
- neutral → slate   #94A3B8 (default)

Dot uses `frame + cornerRadius=4`, NEVER `ellipse` — an 8×8 ellipse
is the classic "status dot via stacked ellipses" anti-pattern bait.
Keeping it a frame stays clean of rewriteLlmAntiPatterns. Regression
test locked in pen-mcp/add-status-badge-v0.test.ts.

Wired across all three paths + elements.md entry + keyword map +
examples + ab-v1/dashboard-server-status.yaml corpus prompt +
parametric builder test CASES in 7 files.

Tool count: 50 → 51. Test count: 3114 → 3138 (+24).
2026-04-22 08:50:00 +08:00
Fini c6e938e66a docs(pen-core): element-builders README — 42 → 50 tool count
Pure text bump: the diagram legend + file-layout note now say
"50 today (as of 2026-04-22)". 3 new tools added in afternoon
batch (image_placeholder, comment, modal_shell).
2026-04-22 08:45:00 +08:00
Fini 54131c03c0 feat(ai): 3 new element tools — image_placeholder / comment / modal_shell (48th-50th)
Fills 3 common UI gaps the 47-tool set didn't cover:

- add_image_placeholder_v0: gray box + centered icon + optional
  caption. The "future image slot" affordance. Separate from G()
  (which fetches real images). Emits frame+fill, NEVER image node
  (empty image renders as broken indicator).
- add_comment_v0: avatar + (author + timestamp) header + body.
  Social / UGC / feedback unit. Does NOT handle replies / likes /
  action menu — compose via batch_design.
- add_modal_shell_v0: dimmed scrim + centered card (rounded,
  shadowed) + title + optional subtitle. "Shell" in the name is
  deliberate — this is chrome only; body content goes into the
  `modal-shell-card` role via a follow-up insert.

Wired across all three paths (pen-core buildX + pen-mcp handler +
routes + schema + apps/web shim + Nitro SERVER_BUILDERS) + elements.md
decision-tree entries + keyword map + examples + 3 pen-mcp handler
tests (22 cases total) + 3 A/B v1 corpus prompts in ab-v1/ +
parametric builder test CASES auto-extended in 9 files.

Milestones:
- Tool count: 47 → 50
- Test count: 3045 → 3114 (+69)
- A/B v1 corpus: 5 → 8 prompts
2026-04-22 08:30:00 +08:00
Fini 37807bda39 test(ai-skills): A/B corpus v1 — 5 obvious prompts for new tools (#51)
Supplemental corpus covering the 5 element tools added after v0
froze (2026-04-20): textarea, skeleton, select, chart_line, chart_pie.
One obvious prompt per tool, so an A/B v2 run can measure routing +
legality on the new surface without re-running the full 24-prompt v0.

v0 stays frozen in `corpus/ab-v0/` so the published v1 results
remain reproducible (openpencil-docs 2026-04-20-ab-v1-results.md).

Per prompt:
  - mobile-bio-textarea    → add_textarea_v0
  - mobile-loading-skeleton → add_skeleton_v0
  - mobile-country-select  → add_select_v0
  - dashboard-revenue-line → add_chart_line_v0
  - dashboard-category-pie → add_chart_pie_v0

corpus-loader tests extended with 3 v1-specific cases:
  - exactly 5 prompts, all obvious
  - covers the 5 expected tool names
  - every prompt anchors must_contain_roles (non-empty)

Monorepo test count: 3042 → 3045. v2 corpus is load-only; running
the actual A/B experiment requires the external harness (not wired
in this commit).
2026-04-22 07:20:00 +08:00
Fini 5af04070e0 test(ai): batch_design browser exec: applyExternalDocument one-shot invariant
Defensive gate: the browser DSL executor clones the doc, mutates
it across all ops, then applies the final doc back in ONE call.
A regression that applied per-op would thrash React + history
state for no benefit (the surrounding startBatch already wraps it
into one undo entry anyway). Spy on `applyExternalDocument` and
assert it's called exactly once per 4-op batch.
2026-04-22 07:10:00 +08:00
Fini 9087345dec test(ai): batch_design DSL browser executor integration (8 cases)
Pairs with the #44 commit (browser-safe DSL executor). Every test
stubs fetch to reject so any regression falling back to HTTP fails
loudly.

Coverage:
- Single I() at root → frame inserted, no HTTP
- Binding chain: root + nested child land in correct parent
- U() update applied: properties merged on bound node
- 6-op realistic screen (nav + cards + divider): order preserved
- Multi-op batch → exactly ONE undo entry (dispatcher's
  startBatch/endBatch wrap survives the browser path)
- Malformed op in the middle: per-line errors surfaced as
  status=failed (not opaque HTTP 500)
- Empty DSL: zero ops, applied + no insertions
- G() without fetcher: image node inserted with empty src (the
  apps/web scanAndFillImages pipeline enriches later)

Complements the 5 static browser-safety checks in pen-mcp
(`batch-design-dsl-browser-safe.test.ts`) — that file gates the
import tree, this one gates the runtime behavior.

Total: 3033 → 3041 passing.
2026-04-22 07:05:00 +08:00
Fini 52f32a8dad feat(ai): browser-safe batch_design DSL executor (closes #44)
Extract ~600 lines of pure DSL logic from pen-mcp/tools/batch-design.ts
into a sibling batch-design-dsl.ts that does not import document-manager
(node:fs) or hooks (server-injected). `handleBatchDesign` becomes a
thin server-side wrapper that opens/saves the .op file around the pure
executor. Backward-compat: pen-mcp barrel + batch-design.ts both
re-export `runBatchDesignDsl` so existing callers keep working.

apps/web dispatcher changes:
- `applyBatchDesignDsl` now runs `runBatchDesignDsl` DIRECTLY in the
  browser against useDocumentStore.getState().document (structuredClone
  + apply via applyExternalDocument).
- HTTP `/api/mcp/exec-tool` fallback fires only when the in-browser
  executor throws (rare — caller error or future regression).
- Removes per-tag HTTP latency on the common batch_design path.
- ctx.defaultParentId is intentionally NOT applied here: the DSL is
  the AI's verbatim instruction set and rewriting `null` parents
  would change author intent. Element-tool calls still honor it.

Image search (`G()` op) is swapped from `getSyncUrl`-based absolute
URL (server) to an injectable `ImageSearchFetcher` callback. Server
wrapper keeps the old behavior; browser path omits the fetcher so
`src` stays empty for the apps/web image pipeline (scanAndFillImages)
to enrich later.

Tests:
- New browser-safety gate `batch-design-dsl-browser-safe.test.ts`:
  walks the transitive import graph from batch-design-dsl.ts and
  fails if any reachable file imports node:fs / node:os / node:path
  / document-manager / hooks. 5 checks including a negative control
  on batch-design.ts (wrapper) to confirm the split is meaningful.
- Dispatcher tests updated to match new behavior: happy path does
  NOT hit fetch; malformed DSL returns `status=failed` with per-op
  error surfaced (via pure executor's `errors[]`), not an opaque
  HTTP 500.
- Total: 3028 → 3033 passing.
2026-04-22 06:55:00 +08:00
Fini 2404769253 fix(ai): chart-pie donut innerRadius is a RATIO, not pixels
EllipseNode.innerRadius is a ratio 0..1 (arc-path.ts docstring;
skia-interaction clamps dist to [0, 0.99]; renderer does `rx * inner`).
Previous builder stored diameter/2 × ratio in pixels, which the
renderer would then multiply by rx AGAIN — blowing past the outer
radius and clipping every slice.

Pass the ratio directly. Test that asserted pixels is updated +
a new test proves the stored value is invariant across diameter
(which pixels-based storage would fail by construction).

Caught by Codex stop-hook review before the donut regression
shipped. No runtime regression to undo — chart_pie hadn't been
called with inner_radius_ratio > 0 yet in any wired path.
2026-04-22 06:45:00 +08:00
Fini 59f25dac82 feat(ai): add_chart_line_v0 + add_chart_pie_v0 — two chart skeletons (46th, 47th)
chart_line: polyline through N data points (normalized to max),
optional dots at each vertex. Emits a `path` node with computed SVG
`d`="M x y L x y …" + N `ellipse` dots. fit_content width = values
× point_spacing.

chart_pie: N colored slices via ellipse `startAngle`/`sweepAngle`
arc support. NOT the "stacked full ellipses" anti-pattern — each
slice has a UNIQUE sweep range (sums to 360°). Supports donut cut-
out via `inner_radius_ratio`. Default 6-color palette rotates.
All-zero input throws (degenerate chart can't be drawn).

Wired across all three paths + elements.md decision-tree entries +
keyword map + examples. Both use layout=none (absolute positioning
for vertices / slices stacked at origin).

Tests:
- pen-mcp: 8 cases chart_line + 8 cases chart_pie (height math,
  clamp bounds, custom colors, donut, error paths, id uniqueness)
- Parametric coverage auto-extended (+2 cases each in 9 files)
- Total delta: 2979 → 3027 passing (+48).

Closes #45 + #46; covers part of #50 (charts entries).
2026-04-22 06:35:00 +08:00
Fini 912ed1cd5a feat(ai): add_select_v0 — dropdown display element tool (45th)
Dropdown/picker closed-state display. Same label-above-input shape
as add_form_field_v0 with:
- always-present trailing chevron-down icon
- when `value` is set: black value text + chevron
- when absent: placeholder text styled gray (#94A3B8) + chevron
- justifyContent=space_between pushes the chevron to the right edge

Explicitly NOT modeled: open-menu state (dropdown list). An open
dropdown needs absolute positioning + scrim + per-option states
that belong in a different builder; compose via batch_design for now.

Wired across all three paths + elements.md keyword map + examples.
Tests:
- pen-mcp: 7 new cases (value/placeholder rendering, custom
  trailing_icon, required suffix, id uniqueness, parent_id rollback)
- Parametric coverage: auto-extended (+1 case each in 9 files).
  Total delta: 2955 → 2979 passing.

Closes #48; covers part of #50 (select entry).
2026-04-22 06:25:00 +08:00