Commit graph

319 commits

Author SHA1 Message Date
Fini 829772f60e fix(ai): swallow image-search network failures into the existing fallback
`fetchFromOpenverse` and `fetchFromWikimedia` were missing try/catch,
so a ConnectTimeoutError on `api.openverse.org` (frequent on networks
that can't reach Openverse) bubbled up to nitro's default handler and
turned a single image-search lookup into a HTTP 500 for the whole
design generation flow. The handler already treats `null` (Openverse)
and `[]` (Wikimedia) as the documented fallback signals — wrap the
fetches and return those on any throw, plus an explicit 8s
AbortSignal.timeout so the wait is bounded.
2026-04-26 07:30:00 +08:00
Fini abbccc1ba2 fix(ai): refresh DeepSeek defaults to v4 model series
`/models` now returns only deepseek-v4-pro and deepseek-v4-flash;
deepseek-chat / deepseek-reasoner sunset 2026-07-24 and the
deepseek-v3.2 hard-coded in the ark-coding fallback list never
existed. Both v4 models default to thinking enabled and the API
toggles via `{"thinking":{"type":"disabled"}}` — keep
`thinkingMode: 'disabled'` so the app's fast/non-thinking default
stays intact (server reasoning paths honor it; the Zig openai-compat
path doesn't emit the toggle yet, so calls through that path still
get provider-default thinking until it's wired). v4-pro promoted to
full tier; legacy aliases pinned to an exact RegExp so future
deepseek-* variants don't inherit a forced disabled mode.

Bandaid for the unwired toggle: v4-pro gets `timeoutMultiplier: 2`
because its default-on reasoning blows past the orchestrator's
planning timeout on long system prompts (observed in dev: planning
phase falls back, sub-agent then succeeds — UX degraded but
functional). Drop the multiplier once the Zig path actually sends
`thinking:{type:disabled}`.

Don't add a BUILTIN_MODEL_LISTS.deepseek entry — DeepSeek exposes
/v1/models, so let `fetchProviderModels` pull the live catalog
through `/api/ai/provider-models` instead of pinning a snapshot
(the ark-coding `deepseek-v3.2` ghost above shows what those
snapshots drift into).
2026-04-26 06:30:00 +08:00
Fini 2f62cb262f fix(core): replace CSS padding shorthands with canonical padding array in 14 element builders
paddingTop/Right/Bottom/Left are silently dropped by resolvePadding in
the layout engine, causing all four sides to render as 0. Switch every
affected builder to the canonical padding: [T,R,B,L] form.
2026-04-26 06:29:59 +08:00
Fini ae40046b17 fix(ai): sidebar_nav uses unified padding so insets actually render
pen-core's resolvePadding only reads the array/scalar `padding`
field; the CSS-style paddingTop/Right/Bottom/Left siblings I shipped
in 5c275ba1 were silently dropped, so the rail rendered flush. Switch
to padding=[16,12] / [8,12,24,12] / [0,12] and add a regression
assertion that catches the same trap on any future tweak.
2026-04-26 06:29:58 +08:00
Fini 0cb0e81a78 feat(ab-corpus): dashboard-team-sidebar ab-v1 prompt for add_sidebar_nav_v0
Adds the obvious-difficulty prompt covering the 77th element tool;
bumps the corpus-loader test from 26 to 27 entries and lists the new
2026-04-27 batch in the README so future corpus drift is caught at
test time.
2026-04-26 06:29:57 +08:00
Fini e9d9e37770 feat(ai): add_sidebar_nav_v0 — desktop persistent left rail (77th tool)
Top_nav_bar / bottom_nav / nav_chip_row covered mobile and inline
chrome but desktop dashboards still had to hand-roll their left
sidebar via batch_design. Adds a 240px-wide vertical rail with
icon+label rows, optional brand title, and slate-100 pill bg on the
active item — distinct surface from the existing nav tools so the
decision tree picks it cleanly for "sidebar / side nav / 侧边栏".
2026-04-26 06:29:56 +08:00
Fini 6d91f60f1e fix(ai): keep deepseek-v3.2 in ark-coding fallback list
The previous DeepSeek refresh swapped this entry for v4-pro / v4-flash,
but ark-coding routes to Volcengine's Coding Plan, not the DeepSeek
direct API — and ARK Coding ships its own DeepSeek catalog where only
deepseek-v3.2 is supported. Selecting v4-pro / v4-flash through ARK
returns "404 The xxxxxx model does not support the coding plan
feature". Restore v3.2 in this list and add a comment so the next
person doesn't repeat the mistake. Direct-DeepSeek preset/profile
keeps the v4 changes — those are correct for api.deepseek.com.

Source: https://developer.volcengine.com/articles/7615528054736945158
2026-04-26 06:29:55 +08:00
Fini b9daac8a22 fix(ai): swallow image-search network failures into the existing fallback
`fetchFromOpenverse` and `fetchFromWikimedia` were missing try/catch,
so a ConnectTimeoutError on `api.openverse.org` (frequent on networks
that can't reach Openverse) bubbled up to nitro's default handler and
turned a single image-search lookup into a HTTP 500 for the whole
design generation flow. The handler already treats `null` (Openverse)
and `[]` (Wikimedia) as the documented fallback signals — wrap the
fetches and return those on any throw, plus an explicit 8s
AbortSignal.timeout so the wait is bounded.
2026-04-26 06:29:54 +08:00
Fini cc4ca08da7 fix(ai): refresh DeepSeek defaults to v4 model series
`/models` now returns only deepseek-v4-pro and deepseek-v4-flash;
deepseek-chat / deepseek-reasoner sunset 2026-07-24 and the
deepseek-v3.2 hard-coded in the ark-coding fallback list never
existed. Both v4 models default to thinking enabled and the API
toggles via `{"thinking":{"type":"disabled"}}` — keep
`thinkingMode: 'disabled'` so the app's fast/non-thinking default
stays intact (server reasoning paths honor it; the Zig openai-compat
path doesn't emit the toggle yet, so calls through that path still
get provider-default thinking until it's wired). v4-pro promoted to
full tier; legacy aliases pinned to an exact RegExp so future
deepseek-* variants don't inherit a forced disabled mode.

Bandaid for the unwired toggle: v4-pro gets `timeoutMultiplier: 2`
because its default-on reasoning blows past the orchestrator's
planning timeout on long system prompts (observed in dev: planning
phase falls back, sub-agent then succeeds — UX degraded but
functional). Drop the multiplier once the Zig path actually sends
`thinking:{type:disabled}`.

Don't add a BUILTIN_MODEL_LISTS.deepseek entry — DeepSeek exposes
/v1/models, so let `fetchProviderModels` pull the live catalog
through `/api/ai/provider-models` instead of pinning a snapshot
(the ark-coding `deepseek-v3.2` ghost above shows what those
snapshots drift into).
2026-04-26 06:29:53 +08:00
Fini 6c44d58a31 feat(ai): add_cookie_banner_v0 — GDPR/CCPA cookie banner (76th tool)
Sticky bottom-of-page disclosure card with title, body, accept /
decline buttons (decline-then-accept order), and an optional
"Cookie settings" link for fine-grained consent. Caller positions
the banner; the tool emits the card itself with shadow + 1px slate
border for visual lift above the page content. Plus the v1 corpus
prompt landing-cookie-banner (26 v1 prompts total).
2026-04-25 08:15:00 +08:00
Fini 1acd31a2e1 feat(ai): add_input_with_action_v0 — input + inline action button (75th tool)
The "Subscribe to newsletter" / "Apply discount code" / "Send chat
message" pattern. Two action variants:
  - text (default): pill button with label like "Subscribe"
  - icon: 44×44 square icon button (chat send arrow / search apply)

Distinct from add_form_field_v0 (label-above, no inline button)
and add_search_bar_v0 (no trailing action button). Optional
leading_icon adds an icon inside the input itself. Plus the v1
corpus prompt landing-newsletter-signup (25 v1 prompts total).
2026-04-25 08:00:00 +08:00
Fini 3527f6ff6a feat(ab-corpus): mobile-phone-input ab-v1 prompt for add_phone_input_v0
Adds 24th v1 prompt and rolls the 2026-04-25 batch from 6 to 7
tools (5 v0 + 2 v1). Loader test count + tool-set assertions
updated.
2026-04-25 07:45:00 +08:00
Fini 577488546a feat(ai): add_phone_input_v0 — international phone input (74th tool)
The "+1 (555) …" pattern from every modern signup / login screen.
A 44px row with leading country selector (flag + dial code +
chevron-down), a 1px slate divider, and the digits input on the
right. Country selector is a button-shape (no actual dropdown
menu); caller handles picker UX as a separate concern. `value`
toggles between placeholder (slate-400) and populated (slate-900).
2026-04-25 07:30:00 +08:00
Fini 008bf0ca46 test(ab-corpus): update v1 corpus expectations to 23 prompts
Three batches now: 2026-04-22 (12) + 2026-04-24 (5) + 2026-04-25 (6
— 4 v0 + 2 v1). The 2026-04-25 batch corresponds to the tools
shipped in this session: social_login / pricing_card / stat_card /
range_slider / toast_v1 / empty_chart_v1.
2026-04-25 01:47:23 +08:00
Fini dd5c85dcc0 feat(ab-corpus): ab-v1 prompts for 6 new tools shipped this session
Adds NL prompts exercising the new narrow tools + theme-aware v1s:
- mobile-social-login (→ add_social_login_row_v0)
- dashboard-pricing-card (→ add_pricing_card_v0)
- dashboard-revenue-stat (→ add_stat_card_v0)
- mobile-volume-slider (→ add_range_slider_v0)
- mobile-dark-toast (→ add_toast_v1, inverted contrast)
- dashboard-dark-empty-chart (→ add_empty_chart_v1)

Each prompt describes the visual intent in natural language, asserts
the expected roles, and pins the expected_tool_if_any so the harness
can attribute tool-choice correctness.
2026-04-25 01:39:40 +08:00
Fini d1022e74d5 feat(ai): add_empty_chart_v1 — theme-aware no-data placeholder
Third theme-aware v1 (after add_modal_shell_v1 and add_toast_v1).
Same dashed-border "no data yet" tile shape as v0 with a `theme`
param that swaps 5 colors (bg / border / icon / title / subtitle):
  - light (default): byte-parity with v0 (slate-50 / slate-300)
  - dark: slate-800 bg + slate-600 border + slate-200 title
  - system: $color-surface-2 / $color-border / $color-text-primary
    refs — requires applySemanticPalette(doc) seeded

Lets dark-theme dashboards keep a matching empty-slot surface
instead of punching a light rectangle out of dark cards.
2026-04-25 01:37:40 +08:00
Fini aea75fa198 feat(ai): add_range_slider_v0 — single-thumb range slider (71st tool)
Visual static representation of a horizontal slider. Track splits
into left fill (accent) + 20×20 thumb (white + accent stroke) +
right remaining (slate), all aligned on a 20px track wrap. Optional
label + value readout row above; `value_suffix` renders "60%" /
"128px" / "0°" style readouts. Pixel math: fill = (width-20) * pct,
auto-collapses fill or remaining at either extreme.
2026-04-25 01:33:46 +08:00
Fini 984bb10bd1 feat(ai): add_toast_v1 — theme-aware floating pill toast
Second theme-aware v1 (after add_modal_shell_v1). Same pill shape as
v0 with a `theme` param:
  - light (default): byte-parity with v0 (dark #111827 pill + white fg)
  - dark: INVERTED contrast — light pill (#F1F5F9) + dark fg (#0F172A)
  - system: $color-text-primary bg + $color-surface fg (inverted swap)

Unlike surface-like v1s (modal-shell), toasts use inverted contrast by
design — a dark pill on light bg, a light pill on dark bg — so the
dark variant flips the pill rather than darkening it.
2026-04-25 01:28:43 +08:00
Fini 91d303d0f0 feat(ai): add_pricing_card_v0 — SaaS pricing tier card (70th tool)
The "Pro $29/month" column for pricing tables. Tier name + big
price (currency + amount + period) + check-mark feature list + CTA
button. Two emphases: default (slate border + slate CTA) and
featured (accent border + accent CTA + auto "Most popular" badge
unless overridden by explicit `badge` value).
2026-04-25 01:24:24 +08:00
Fini e0b926c256 feat(ai): add_social_login_row_v0 — social auth button row (69th tool)
The "Continue with Google / Apple / Microsoft" row. Vertical default
(stacked full-width 48px buttons) or horizontal (compact icon-only
48×48 pills). Known provider names (google, apple, github, microsoft,
facebook, twitter, linkedin, discord, slack, gitlab, email, phone)
auto-map to lucide icons; `icon` param overrides for SSO/SAML/Okta.
2026-04-25 01:19:00 +08:00
Fini 127eccdefb feat(ai): add_stat_card_v0 — big-number KPI tile (68th tool)
Featured-metric dashboard tile: label (uppercase muted) above a
huge 32/700 primary value, with optional tone-colored delta line
and corner icon. Distinct from its neighbors:

  - add_stat_grid_v0   — multi-cell side-by-side, smaller values
  - add_metric_comparison_v0 — horizontal label+value+inline arrow
  - add_stat_card_v0 (this) — single featured metric, whole-card focus

Trend enum tones the delta line only (value stays slate-900):
  up    → #10B981 (emerald)
  down  → #EF4444 (red)
  flat  → #64748B (slate, default)

Wired through all standard points — schema into ext-3 (shortest
shard at 371 lines + new tool → 408, still comfortably under 800)
+ shim + SERVER_BUILDERS + parity CASES + contract allow-list +
elements.md decision tree + triggers + minimal usage.

Handler test covers 6 cases: minimal registration + defaults +
icon+delta+up-tone + down-tone + flat-tone default + width
clamp + bogus parent_id rejection.
2026-04-25 01:09:49 +08:00
Fini 23988b9e84 fix(ab-corpus): per-call timeout + progressive scores.jsonl writes
The first 12-prompt × 2-model × 2-variant sweep (48 API calls)
ran 29 minutes before I killed it. A Kimi call on the 10th
prompt hung indefinitely — no client-side timeout — and the
harness writes scores.jsonl + report.md ONLY at the end, so
partial progress was unrecoverable. Lost 9/12 completed prompts
because the aggregate step never ran.

Two fixes:

1. openai-compat.ts: AbortController with default 120s timeout
   (overridable via AB_CORPUS_CALL_TIMEOUT_MS env). When a call
   exceeds the budget, the harness catches the abort, records it
   as __HARNESS_ERROR__ (routing=garbage, M1=false), and moves on.
   Verified by dialing the timeout to 60s — GLM-5.1's first call
   took >60s, got aborted cleanly, run continued to completion
   instead of hanging.

2. run.ts: append each ScoreRow to scores.jsonl immediately after
   scoring. Truncate at start (so re-runs overwrite). Lost-work
   window now bounded to "the currently-executing API call," not
   "everything since the run started." report.md and report.json
   still write once at the end (aggregate needs the full set) but
   scores.jsonl alone is enough for any partial-run analysis.

Post-hardening validation (live 方舟 CP runs):
  - mobile-upload-dropzone → add_upload_dropzone_v0 ✓ right-tool
  - dashboard-dark-modal   → add_modal_shell_v1 (theme=dark) ✓

Second one is the first end-to-end proof that the v1 theme-aware
tool family routes correctly with a real LLM — GLM-5.1 inferred
\`theme: "dark"\` from the natural-language prompt.
2026-04-25 00:43:48 +08:00
Fini 54ee6e0ec7 feat(ab-corpus): --corpus flag + 5 new prompts for tools 63-67 + v1
Two independent changes rolled together since they both serve the
same goal — "can real LLMs actually route to the tools we shipped
today?":

  1. scripts/ab-corpus/run.ts gains a `--corpus` flag (ab-v0 |
     ab-v1, default ab-v0 for back-compat). The harness was
     hardcoded to ab-v0 — adding ab-v1 prompts was worthless
     without a way to run them. Validated 17 prompts × 2 models ×
     2 variants during a live 方舟-CP run.

  2. 5 new ab-v1 prompts cover the 2026-04-24 tool batch:
      - mobile-upload-dropzone.yaml        → add_upload_dropzone_v0
      - mobile-otp-verification.yaml       → add_otp_input_v0
      - mobile-file-attachment.yaml        → add_attachment_row_v0
      - mobile-chat-message.yaml           → add_chat_bubble_v0
      - dashboard-dark-modal.yaml          → add_modal_shell_v1

corpus-loader.test.ts bumps its count assertion 12 → 17 and
extends the tool-coverage set. Also loosens the regex to accept
`_v\d+$` (was `_v0$`) so add_modal_shell_v1 passes. No other
test file needed changes — the existing registry-parity and
mock-llm tests already use `_v\d+$` or the registry directly.

.gitignore gains:
  - .playwright-mcp/ (MCP Playwright session artifacts)
  - editor-*.png      (local verification screenshots)
  - scripts/ab-corpus/runs/  (live-run outputs / reports)

None of those belong in version control — they're artifacts
from local verification runs.
2026-04-25 00:23:11 +08:00
Fini 6d4bdb88da test(pen-mcp): support-chat composition scenario covering #63-#67 + v1
Extends element-tools-composition.test.ts with a 7th real-screen
scenario that exercises every tool added after the 62-tool mark:

  - add_top_nav_bar_v0         (existing anchor)
  - add_chat_bubble_v0         × 2  (left from-others + right from-self)
  - add_attachment_row_v0      (file on the self-message path)
  - add_upload_dropzone_v0     (drop area for screenshots)
  - add_chip_input_v0          (conversation tags)
  - add_action_menu_v0         (floating menu, open state)
  - add_modal_shell_v1         (theme="dark" confirm dialog)
  - add_otp_input_v0           (phone verification step)

8 tool calls, chained through the full MCP handler pipeline
(ensureParentExists → builder → assignIdsRecursively →
batch_design insert with rollback-on-failure → post-insert
landing check → save → re-read from disk). Asserts:

  - every call emits a nodeId (no silent no-ops)
  - final doc has exactly 9 root children
  - every tool's role marker survives round-trip save/load
  - modal v1 theme=dark → card fill #1E293B (not v0's #FFFFFF)
  - chat right-side bubble surface → #2563EB accent
  - OTP focused slot stroke → #2563EB accent

This is end-to-end at the DATA layer, not the visual layer.
Still not covered by any test:
  - Skia rendering (needs debug_screenshot against a live
    canvas)
  - Real LLM tool routing (needs ab-corpus harness with live
    API keys — route to 方舟 CP works but has not run since
    the routing fix)

Document-layer coverage is enough to rule out handler-pipeline
bugs (bad parent_id threading / cache-stale doc-state / silent
rollback swallowing); visual regressions require the next gate.
2026-04-22 22:31:49 +08:00
Fini 06362a4040 refactor(mcp): split ext-2 into ext-2 + ext-3 (keep shards under 800)
Codex stop-hook caught element-tool-defs-ext-2.ts at 832 lines
after add_chat_bubble_v0 landed there — 32 over the repo's 800-
line ceiling. Same trap the original single ext file hit at 1329
lines, same fix pattern: carve the second half into a new shard.

Split at `add_modal_shell_v1` (tool #16 of 24 in old ext-2):
  - ext-2 keeps tools 1-15 (calendar_grid through textarea) →
    482 lines
  - ext-3 (new) holds tools 16-24 (modal_shell_v1 through
    chat_bubble) → 371 lines

All registry shards now:
  base    683
  ext     671
  ext-2   482
  ext-3   371
  (props   26)
  (top    254)

element-tool-defs.ts concatenates all three ext shards into the
single ELEMENT_TOOL_DEFINITIONS — external API unchanged.
ELEMENT_TOOL_DEFINITIONS_EXT_3 is the new import; 67 tools
still resolve.

Header in ext-1 updated to reflect the three-way split + advise
"when ANY shard crosses 700, carve a ~5-tool chunk into the
shortest shard" so the next rebalance happens proactively instead
of after a stop-hook trip.
2026-04-22 22:12:26 +08:00
Fini 165fb94e44 feat(ai): add_chat_bubble_v0 — messaging bubble (67th tool)
Chat / messaging / customer-support UI message unit. Two variants
via the `side` enum:

  - side="left" (default): from-others bubble. Slate-100 fill,
    slate-900 text, alignItems=flex-start. Optional `author` text
    shown above the bubble (group-chat pattern).
  - side="right": from-self bubble. Accent-color fill (customizable
    via `accent_color`), white text, alignItems=flex-end. Author
    intentionally suppressed on this side — a self-bubble never
    carries "You:".

Optional `timestamp` below the bubble on either side.

Max-width mechanic: pen-core has no native max-width primitive, so
`max_width` becomes the bubble's fixed width (clamped 160..480).
Short messages get extra padding on one side — matches every real
chat client (iMessage / WhatsApp / Slack). Message text uses
`textGrowth: 'fixed-width'` + `width: 'fill_container'` to wrap
correctly inside the fixed-width surface.

Full wiring: pen-core builder + pen-mcp handler + schema (into
ext-2, shorter shard — 24 tools vs ext-1's 24 after this) + shim
+ SERVER_BUILDERS + parity CASES + contract allow-list +
elements.md decision tree + triggers + minimal usage for both
sides. Handler test covers 9 cases: registration, left defaults,
left+author, right with self-dropped-author, right+accent_color,
timestamp both sides, max_width clamps (low + high split into
separate tests to avoid cache-interference), textGrowth wiring,
bogus parent_id rejection.
2026-04-22 22:03:44 +08:00
Fini 49d7d1baf2 feat(ai): add_attachment_row_v0 — file attachment list unit (66th tool)
Fills another common UI gap: the "here's an already-uploaded file"
row you see in email composers, chat attachments, and form upload
summaries. Compact horizontal layout: type-icon + filename (bold) +
optional muted size string + optional right-side × remove affordance.

Structure: horizontal frame (slate-50 bg, cornerRadius=8) with
three children:
  1. attachment-icon — lucide file-* (caller picks: file / file-
     text / file-image / file-video / file-audio / file-archive /
     file-spreadsheet / file-code)
  2. attachment-meta — vertical frame with filename + optional size
  3. attachment-remove — × icon, suppressed via removable=false

Intentionally NOT embedding an upload-progress variant in v0. The
pen-core schema lacks percentage-width primitives, so a %-filled
progress bar would either need a fixed track width (brittle across
parents) or a caller-computed pixel value (awkward API). Callers
who need the uploading state compose `add_progress_bar_v0` directly
below the row — cleaner separation.

Wired through all standard points: schema into ext-1 (balanced
shards 23/23 after upload-dropzone landed there last commit) +
shim + SERVER_BUILDERS + parity CASES + contract allow-list +
elements.md decision tree + triggers + minimal usage.

Handler test covers 7 cases: registration + minimal (no size) +
size rendered + custom icon + removable=false + default icon +
bogus parent_id rejection.
2026-04-22 21:53:47 +08:00
Fini dc73885015 feat(ai): add_otp_input_v0 — verification code input (65th tool)
Fills the auth-flow gap: 2FA / PIN / phone-verification codes.
Horizontal row of N square slots (4..8), one digit per slot.
Renders three states per caller intent:

  - blank (no `digits`): all slots empty, `focused_index` marks
    the currently-typing slot with an accent-color 2px outline
  - partial: first M slots filled with digit text, slot M+1
    focused, rest empty
  - full: all N slots filled (final submittable state)

Filled slots get role=otp-slot-filled + slate-700 border + 20/600
digit text. Focused empty slot gets role=otp-slot-focused +
2px accent border. Blank unfocused slots get role=otp-slot +
1px slate-300 border.

Wired through all standard points — schema into ext-2 (shorter
shard) + shim + SERVER_BUILDERS + parity CASES + contract
allow-list + elements.md decision tree + triggers + minimal
usage for each state.

Handler test covers 8 cases: registration + defaults (6 blank
focused-first) + partial state / full state / length clamp low
(< 4 → 4) + length clamp high (> 8 → 8) + accent color override
+ bogus parent_id rejection.
2026-04-22 21:48:05 +08:00
Fini 45214efb70 feat(ai): add_upload_dropzone_v0 — file drop zone (64th tool)
Fills a real gap in the element-tool family: upload / drag-and-
drop surfaces. Dashed border + cloud icon + two-line instruction
("Drop files to upload" / "or click to browse") — the classic
pattern from every modern file-upload UI.

Deliberately structurally similar to add_empty_chart_v0 (dashed
border + icon + title/subtitle) but semantically distinct:
  - empty_chart = "chart widget will render when data arrives"
    (320×200, icon chart-typed)
  - upload_dropzone = "users drop files here" (480×200, icon
    semantic: upload-cloud / upload / file-up)

elements.md routes them by intent, and the tool descriptions
cross-reference each other to prevent the AI from picking the
wrong one on ambiguous prompts.

Wired through all the standard points per the add-new-tool
checklist: builder + handler + schema (into ext-1, the shorter
shard) + shim + SERVER_BUILDERS + parity CASES + contract
allow-list + elements.md decision tree + keyword triggers +
minimal usage. Handler test covers 5 cases: defaults, dashed
stroke, overrides, size clamping, bogus parent_id rejection.
2026-04-22 21:41:13 +08:00
Fini ed449bc896 refactor(mcp): split element-tool-defs-ext in half + extract shared props
[Codex P3] element-tool-defs-ext.ts had grown to 1329 lines —
over the repo's documented 800-line ceiling. Ironically the file
header comment claimed it existed to keep its parent under the
cap, but the shard itself had outgrown the limit.

Split into three files:

  - element-tool-def-props.ts (26 lines) — shared JSON-Schema
    fragments (schemaVersionProp / filePathProp / parentIdProp /
    pageIdProp) that every definition file uses. Deduplicating
    these unblocks the split cleanly.
  - element-tool-defs-ext.ts (596 lines) — first 22 tools
    (add_switch_v0 through add_segmented_control_v0 era). Imports
    the shared props.
  - element-tool-defs-ext-2.ts (744 lines, new) — remaining 22
    tools starting at add_calendar_grid_v0. Imports the shared
    props.

element-tool-defs.ts concatenates all three arrays into the single
ELEMENT_TOOL_DEFINITIONS registry — external API surface unchanged.

Header comments in both shards now document the split convention:
"pick whichever shard has fewer tools" when adding a new entry,
to keep the files balanced as the family grows toward ~100.

Incidentally the previous commit's file also carried the P2 fix
(pageId threading through the in-browser and HTTP DSL paths, so
multi-page docs land the generation on the ACTIVE page instead of
doc.pages[0]). Both touched the same file, didn't make sense to
split. Title-wise the previous commit is P1 but functionally it's
P1+P2.

All under the 800-line ceiling now:
  element-tool-defs-base.ts  683
  element-tool-defs-ext.ts   596
  element-tool-defs-ext-2.ts 744
  element-tool-defs.ts       240
  element-tool-def-props.ts   26
2026-04-22 11:20:00 +08:00
Fini 0ce113c733 fix(ai): dispatcher imports DSL executor via browser-safe subpath
[Codex P1] The browser-side element-tools-dispatcher imported
runBatchDesignDsl from the \`@zseven-w/pen-mcp\` package barrel.
That barrel re-exports node-only modules — document-manager,
log-utils, theme-presets — which import node:fs / node:path at
top level. Vite / esbuild resolve the barrel BEFORE tree-shaking
can drop those branches, so browser builds failed on unresolved
node built-ins.

Fix:
  - packages/pen-mcp/package.json: add \`./dsl\` subpath export
    pointing at tools/batch-design-dsl.ts — the pure executor
    file already guarded as browser-safe by the adjacent
    regression test.
  - apps/web dispatcher: switch import to
    \`@zseven-w/pen-mcp/dsl\`. No other changes — the re-exported
    symbols (runBatchDesignDsl / OpResult / ImageSearchFetcher /
    RunBatchDesignDslOptions) are identical shape.
  - batch-design-dsl-browser-safe.test.ts: add an assertion that
    package.json's exports field preserves the \`./dsl\` key
    pointing at the expected file. Without this, silently
    removing the subpath would re-introduce the browser-breaking
    resolution path.

The package barrel keeps its current export of runBatchDesignDsl
too (a few internal test files still import from it). Browser
callers should migrate to \`@zseven-w/pen-mcp/dsl\` per the JSDoc
note now in the dispatcher.
2026-04-22 11:15:00 +08:00
Fini e10b3a37c9 fix(ab-corpus): kimi-2.6 alias fell through to Bailian instead of Ark
The Ark router regex was `/^kimi-k2\.6/i` — required the `k`
prefix. But mapKimiArkId's alias list accepted both `kimi-k2.6`
AND `kimi-2.6` (no-prefix form). Result: `kimi-2.6` failed the
Ark regex, fell through to the generic `/^kimi/i` branch, got
routed to Bailian — which doesn't host K2.6. Bailian would
return HTTP 400 "model not supported" with no hint that the id
belonged on Ark.

Fix: Ark router regex now `/^kimi-k?2\.6(-ark)?$/i` — optional
`k` prefix + optional `-ark` suffix, anchored at end to prevent
accidentally matching a hypothetical later version. mapKimiArkId
normalizes all four accepted aliases (kimi-k2.6, kimi-2.6,
kimi-k2.6-ark, kimi-2.6-ark) to the canonical on-Ark id.

Same latent bug fixed on the glm-5.1 route: regex `/^glm-5\.1/i`
would prefix-match a hypothetical `glm-5.10` and wrongly route it
to Ark. Tightened to `/^glm-5\.1(-coding|-ark)?$/i` with the same
anchored-end + suffix-allowlist pattern.

Caught by Codex stop-hook review during 2026-04-22 session.
2026-04-22 11:10:00 +08:00
Fini 5538350a2f chore(ab-corpus): route glm-5.1 + kimi-k2.6 through 方舟 CP (Ark)
Volcengine 方舟 (Ark) added GLM-5.1 and Kimi-K2.6 to its coding
plan on 2026-04-22 — single ARK_CODING_KEY covers both. Harness
now prefers this route over the previous paths:

  - glm-5.1 was routed to clients/glm.ts (GLM official CP via
    open.bigmodel.cn with GLM_OFFICIAL_CODING_KEY). Now routed to
    new clients/ark.ts. The old glm.ts file is kept on disk for
    historical comparison but not wired into the default router —
    callers who want to A/B the old GLM-official path vs. new Ark
    path can import callGlm directly.
  - kimi-k2.6 is new — added as a dedicated router branch above
    the kimi-k2.5 (bailian) branch so the version-specific match
    lands on Ark.

Old kimi-k2.5 continues to route through clients/bailian.ts
(DashScope aggregator) for continuity with earlier A/B runs.

Key management (unchanged from the harness convention):
  - ARK_CODING_KEY — Volcengine 方舟 CP UUID format key. Export
    in shell before running --live; never committed.
  - Existing MINIMAX_API_KEY / GLM_OFFICIAL_CODING_KEY /
    DASHSCOPE_BAILIAN_CODING_KEY all still honored for their
    respective routes.

Throw message updated so missing-key errors surface the correct
env var for each route.
2026-04-22 11:05:00 +08:00
Fini dd5d156160 feat(ai): add_modal_shell_v1 — first theme-aware MCP tool (63rd, v1-family debut)
Wires the buildModalShellV1 pen-core builder through the full MCP
toolchain — handler + schema + dispatch + shim + SERVER_BUILDERS
+ parity CASES + handler tests + elements.md skill — so external
MCP clients (Claude Code / Codex / Gemini CLI) can call
add_modal_shell_v1 as a first-class tool alongside the 62 v0
tools.

Three theme variants exposed via `theme` param (enum [light, dark,
system]):
  - omitted / `'light'`: byte-parity with add_modal_shell_v0
    (same hex, same structure, same role tree)
  - `'dark'`: hardcoded dark palette (#1E293B card, #F1F5F9 title,
    #94A3B8 muted). No \$refs needed.
  - `'system'`: emits \$color-surface / \$color-text-primary /
    \$color-text-muted refs. Caller MUST have run
    applySemanticPalette(doc) first or refs resolve to undefined
    (documented in schema description).

Scrim stays #000000 in ALL themes — modal backdrops are a dim
effect, not a themeable surface. Pinned by handler test.

9 handler test cases cover: registration + schema shape +
required[title] + theme variants + scrim invariant + bogus
parent_id rejection. Parity test added a CASES entry with
\`theme:'dark'\` args (exercises the theme branch in both
shim and server paths).

elements.md gained:
  - §46b decision-tree entry pointing to the v1 variant
  - Trigger list entry for dark-mode / theme-aware prompts
  - Minimal usage showing \`theme:'dark'\` + \`theme:'system'\`

This is the reference implementation for the remaining 9 theme-
aware v1 tools in the top-10 offenders list (empty-chart,
chip-input, toast, pagination, notification-row, image-placeholder,
faq-item, comment, checkbox — per dark-theme-audit §offenders).
2026-04-22 11:00:00 +08:00
Fini 315ded51ad chore(ai): generalize element-tool name regex to accept _v\d+
Five regex sites across pen-ai-skills + pen-mcp + apps/web were
anchored at `_v0$`, blocking the _v1 family from being recognized
as element tools:

  - packages/pen-ai-skills/src/corpus/output-parser.ts
    ELEMENT_TOOL_NAME_RE (filters tool_call outputs in A/B
    scorer)
  - apps/web/src/services/ai/design-parser.ts:106 (embedded
    orchestrator dispatch)
  - packages/pen-mcp/src/__tests__/design-prompt-elements.test.ts
    (×2 — stale-integration guard for elements.md)
  - packages/pen-mcp/src/__tests__/element-tool-registry-parity.test.ts
    ("every tool name matches convention" — renamed to _vN)
  - apps/web/src/services/ai/__tests__/element-tools-dispatcher.test.ts
    (drift guard for SUPPORTED_EMBEDDED_ELEMENT_TOOLS)

All now accept /^add_[a-z_]+_v\d+$/. Registry-parity test's
expectedBuilder mapping already handled both v0 (strip suffix →
buildModalShell) and v1+ (preserve → buildModalShellV1) via the
existing `.replace(/_v0$/, '')` — no change there.

Prerequisite for landing add_modal_shell_v1 as a first-class MCP
tool in the next commit.
2026-04-22 10:55:00 +08:00
Fini b07b970daa feat(pen-core): buildModalShellV1 — first theme-aware v1 element (proof of chain)
Ships the first theme-aware element builder demonstrating the v1
contract end-to-end. `buildModalShellV1({ title, theme })` accepts
three theme variants:

  - `'light'` (default): byte-parity with buildModalShell v0 —
    same hex literals, same role tree, same structural shape.
    Structural test asserts stripIds(v0) === stripIds(v1) for
    the default-theme path.
  - `'dark'`: hardcoded dark-palette hex (#1E293B card, #F1F5F9
    title, #94A3B8 muted text). No \$refs — this path is for
    callers who want a dark modal without the theme-switching
    infrastructure.
  - `'system'`: emits \$color-surface / \$color-text-primary /
    \$color-text-muted refs. Renders track \`themes.Mode\` at
    paint time. Requires \`applySemanticPalette(doc)\` to have
    been seeded; if not, refs resolve to undefined (caller's
    responsibility per the v1 contract).

End-to-end tests validate the 'system' path's round-trip through
resolveColorRef for both Light (→ #FFFFFF) and Dark (→ #1E293B)
modes. That's the full chain working:

  buildModalShellV1({theme:'system'}) → tree with \$refs
  → applySemanticPalette(doc) → palette seeded
  → resolveColorRef(ref, doc.variables, {Mode:'Dark'}) → hex

One intentional design note tested: scrim stays #000000 in BOTH
light and dark themes. Modal backdrops are a "dim everything
below" effect, not a themeable surface — dimming a dark surface
with a dark color is a better visual than a themed shade.

18 tests. v0 byte-parity verified against the existing
buildModalShell for full structural equality. This is the
reference implementation for all subsequent v1 tools (top-10
offenders per the dark-theme audit).
2026-04-22 10:50:00 +08:00
Fini 19a1b13764 docs(ai-skills): teach variables.md about the 14 semantic tokens
Extends the generation-phase variables skill with a table of the
14 semantic palette tokens that \`applySemanticPalette(doc)\`
seeds. Models consuming this skill learn:

  1. The exact token names + their light/dark resolved values
     (can cross-reference against what the user's document
     actually has via \`hasSemanticPalette\`)
  2. When to PREFER \`\$color-*\` refs over hex literals (theme-
     aware intent: dark-mode design, system-follow apps, user-
     toggleable themes)
  3. When to FALL BACK to hex (default createEmptyDocument state
     where the palette isn't seeded)
  4. That semantic tokens override theme — \`\$color-success\`
     stays green in both light and dark modes because "green"
     is the semantic signal, not a visual choice

Without this guidance, an AI asked for a dark-themed dashboard
could either: (a) emit hex literals that don't track theme
(defeats the purpose), or (b) emit \$color-* refs blindly into
a doc that lacks the palette (resolves to undefined, renders as
raw string). The table + fallback rule close both gaps.
2026-04-22 10:45:00 +08:00
Fini ce32f9f572 feat(pen-core): 14-variable semantic palette (unblocks v1 theme-aware tools)
Ships the canonical 14-token palette called for in the dark-theme
audit (openpencil-docs/superpowers/notes/2026-04-22-dark-theme-
defaults-audit.md §role clusters). Every token has paired Light +
Dark values on a single `Mode` theme axis.

API surface:
  - getSemanticPalette() → {themes, variables} for merge
  - getSemanticPaletteHex(mode='Light') → flat Record<name, hex>
  - applySemanticPalette(doc) → non-destructive merge (user-
    defined variables + theme axes WIN on collision; palette is
    purely additive)
  - hasSemanticPalette(doc) → runtime check for v1 tools before
    emitting \$color-* refs
  - getSemanticPaletteDescription(name) → human-readable string
    for token-picker UI
  - SEMANTIC_PALETTE_NAMES + theme-axis constants exported

The 14 tokens:
  - Surfaces: color-surface, color-surface-2, color-surface-3,
    color-bg-deep
  - Borders: color-border, color-border-strong
  - Text: color-text-primary, color-text-body, color-text-muted,
    color-text-subtle
  - Semantic: color-accent, color-destructive, color-success
  - Other: color-scrim (with alpha for modal backdrop)

Intentionally NOT wired into createEmptyDocument(). Seeding by
default would alter every existing document on re-save and
violate the v0 byte-parity contract (the whole point of the
audit). v1 tools will call applySemanticPalette(doc) as a pre-
flight, OR the app shell offers a "enable dark theme" user
action that triggers the apply.

29 tests cover: palette shape (14 variables, 2 themed values each,
hex-formatted, light≠dark), hex getter for both modes, apply
non-mutation + user-var-wins-on-collision + additive theme-axis
merge, hasSemanticPalette (empty / full / partial), and full
round-trip through the existing resolveVariableRef / resolveColorRef
paths for every palette name.
2026-04-22 10:40:00 +08:00
Fini 42b6492404 chore(tests): fix tsc errors in earlier test files
Two tsc errors surfaced by a later tsc run:

1. browser-image-search-fetcher.test.ts — imported `beforeEach`
   but never used; also `spy.mock.calls[0]` typed as empty tuple
   since vi.fn()'s signature isn't inferred. Cast via `unknown` +
   explicit tuple shape.

2. chart-builders-visual-smoke.test.ts — custom-dimensions case
   passed `width`/`height` to buildChartLine. The real shape is
   `point_spacing` + `chart_height`. Fixed the test to use the
   actual param names (test itself wasn't broken, just the type).
2026-04-22 10:35:00 +08:00
Fini cf3aa4240a test(pen-core): chart builders visual smoke (30 cases)
Geometric invariants for the three chart builders (bars/line/pie)
that a rendering failure would start from. Can't run Skia
headlessly in unit tests (CanvasKit WASM is heavy + GPU-context-
dependent), so this is the cheap smoke layer that catches shape-
level regressions before they reach the renderer. The app-level
debug_screenshot MCP tool gives us real visual regression on top.

Per chart type:
  - buildChartBars: 8 variants (default, all-equal, single, small
    values, large values, custom dims, zeros, empty-throws).
    Asserts one chart-bar per value, all dims finite.
  - buildChartLine: 9 variants (smooth, monotonic, flat, two-point,
    spike, fractional, custom dims, single-value, empty-throws).
    Asserts chart-line geometry present, coords finite.
  - buildChartPie: 9 variants (equal, skewed, two, many-thin,
    custom diameter, donut at 0.5 and 0.9 ratio, single 100%,
    empty-throws, all-zero-throws). Asserts one slice per value,
    startAngle/sweepAngle finite, sweepAngle positive, total
    sweep = 360°.

Cross-chart invariants: same input length → same geometry-child
count; all three types produce finite-coord trees on identical
input.

Empty-input behavior: all three builders throw with clear
messages. Test pins the throws as intended behavior — the
alternative (silently returning an empty tree) would let an
all-zeros dataset produce a "chart is there but invisible" UI
bug that's much harder to diagnose than the explicit throw.
2026-04-22 10:30:00 +08:00
Fini 90d409207c feat(ai): browser-side G() image-search fetcher (relative URL)
Adds makeBrowserImageSearchFetcher() — a browser-safe
ImageSearchFetcher that POSTs to /api/ai/image-search (relative,
same-origin) for inline G() resolution in the batch_design DSL.

The server-side fetcher uses absolute URL via getSyncUrl() — not
applicable in the browser, where fetch resolves relative paths
against window.origin. This helper mirrors the server-side shape
but drops the sync-URL dependency, so any browser caller that
wants inline image-search can opt in.

NOT wired into the default dispatch path. The existing behavior
(applyBatchDesignDsl omits the fetcher → empty src → enriched
asynchronously by scanAndFillImages) stays the default because
per-G() round-trips would blow up latency on batches with many
images. This helper is opt-in for callers that accept that
trade-off (composition smoke tests, user-opt-in preview modes).

Never throws. Returns null on: empty query, network error,
non-ok response, non-JSON body, missing/empty results, invalid
shape. Callers can drop it in without try/catch wrappers.

16 test cases covering happy path (URL/body/headers/first-of-
many) + 11 failure modes + 1 concurrency invariant.
2026-04-22 10:25:00 +08:00
Fini 59e6c09b18 test(ai): orchestrator cancel mid-batch — structural + behavioral (12 cases)
Two-layered coverage for the abort-signal pattern that's the sole
mechanism preventing the orchestrator from continuing to call the
LLM after the user hits Stop.

A. **Structural** (grep-level): pins that every sequential / per-
   screen-group loop in orchestrator-sub-agent.ts AND the "no nodes"
   throw + validation gate in orchestrator.ts check abortSignal at
   the right points. A new loop that forgets the guard silently
   wastes tokens AND mutates canvas post-stop; this grep-style
   check catches it at unit-test time.

B. **Behavioral** (pure): stubbed sequencer that mirrors the
   actual loop body, exercised with 7 scenarios:
     - no abort → all run
     - abort during iteration N → N+1 onwards never run
     - pre-aborted signal → zero execution
     - post-completion abort → no-op
     - no signal → normal behavior
     - concurrent abort during one worker → other workers stop at
       next iteration-start check
     - pre-aborted signal with concurrent workers → both skip

Real orchestrator logic lives in orchestrator-sub-agent.ts but is
wrapped in LLM-calling code that's expensive to mock. The
behavioral stub is pattern-isomorphic: if the stub works here, the
real loop does too; if the real loop changes shape, the structural
check catches the drift.
2026-04-22 10:20:00 +08:00
Fini a22dc6a639 test(ai): model-tier × elements-skill injection e2e (18 cases)
Existing model-profiles-element-tools.test.ts only covers the
upstream flag boolean (\`needsElementTools\`). This file pins the
downstream filter — \`compactSubAgentSkills\` — specifically
around the elements skill, the spot where a regression would
silently drop elements.md content from the sub-agent prompt even
though VITE_ENABLE_ELEMENT_TOOLS=1 is set.

Covers:
  - basic tier: allow-list preserves elements (mobile + non-mobile)
  - basic + reducedComplexity: elements INTENTIONALLY dropped for
    retry path (~17k char savings when fallback to batch_design)
  - standard / full: elements always passes through
  - jsonl-format vs jsonl-format-simplified conflict: simplified
    wins, elements survives both resolutions
  - Screen-type gates (mobile-app vs landing-page/copywriting/
    anti-slop) are orthogonal to elements — elements survives
    every combination
  - Determinism: same input → same output; original array not
    mutated

Also exercises edge cases: empty skill list, elements-only list,
unknown-name skill at each tier.
2026-04-22 10:15:00 +08:00
Fini 6a2e80022c feat(ai-skills): mock-LLM local A/B harness (no API tokens)
Deterministic mock-LLM for local unit-level testing of the scorer
+ output-parser + apply pipeline — the same three components real
A/B runs use, minus the network.

mockLlmRaw(prompt, variant) returns the raw string a "well-
behaved" model would emit:
  - Treatment on obvious prompt → <op_tool>{...}</op_tool> naming
    expected_tool_if_any with empty args
  - Baseline → minimal batch_design DSL with a single frame role-
    stamped from must_contain_roles[0]
  - Optional + no hint → falls through to the baseline path

mockLlmParsed() skips the raw-string round-trip and returns a
ParsedOutput directly for tests pinning a specific kind. Both
respect (promptId, variant) overrides so tests can simulate
garbage / wrong-tool / empty outputs inline.

Integration test loads the real ab-v1 corpus from disk and
exercises every prompt through the full pipeline:
  corpus → mockLlmRaw → parseModelOutput → scoreRun → ScoreRow

Verifies all 4 routing outcomes (right-tool / wrong-tool /
fallback / garbage) classify correctly on mocked input. 17 test
cases; corpus sweep runs in ~6ms — fast enough to gate every
PR without slowing CI.

This is the prerequisite for future "real" A/B test runners: if
the harness misclassifies obvious mock inputs, no conclusion
from a real run would be trustworthy.
2026-04-22 10:10:00 +08:00
Fini 271ea1b317 feat(mcp): dispatcher metrics — per-tool call/error counters
In-memory counters for the element-tool dispatcher, exposed via
\`getElementToolMetric(name)\` / \`getAllElementToolMetrics()\` /
\`getTopElementToolCalls(n)\` / \`resetElementToolMetrics()\` in
packages/pen-mcp/src/metrics/.

\`handleElementToolCall\` now wraps the existing switch in a
try/record — every dispatch increments \`calls\`, thrown handlers
additionally bump \`errors\` and stash the last error message.
Unknown tool names still fire a counter (useful signal: "the AI
picked a tool we don't have").

Process-local / in-memory by design:
  - Test determinism: resetElementToolMetrics() in beforeEach
  - Matches stdio MCP server's one-client-one-server model
  - No persistence backend choice baked in — if we need
    cross-restart persistence later, a thin serializer drops on
    top without touching this API

Unlocks #92 Local A/B harness: feed a corpus through the MCP
server, read back getAllElementToolMetrics() to see which tools
the model actually picked vs what the corpus expected. Core
observability for non-Claude regression detection.
2026-04-22 10:05:00 +08:00
Fini 801ff4c532 test(pen-core): unicode normalization preservation (30 cases)
Pins that text-carrying element builders preserve the caller's
byte representation verbatim — no silent normal-form conversion,
no zero-width stripping, no fullwidth↔ASCII collapse.

Tests each of NFC/NFD/NFKC/NFKD forms through 6 representative
builders (heading / body-text / list-row / form-field / faq-item /
comment), plus 5 targeted fixtures:
  - Zero-width joiner mid-word ("emo‍ji" stays 6 codepoints)
  - BOM at string start
  - ZWJ emoji family sequence (4-person glyph)
  - Vietnamese combining-marks (NFC vs NFD both preserved as-is)
  - Halfwidth/fullwidth CJK distinction (NFKC would collapse
    fullwidth "A" to ASCII "A"; we assert the builder does NOT)

Why this matters: macOS ships filenames in NFD, Windows/web in
NFC; copy-paste carries any form; some CJK inputs emit
precomposed, others decomposed base+combining. Exact-match
lookups in external systems (especially emoji-less fallback
keys) break silently if the builder pre-empts the downstream
validator's normalization decision. Builders must pass through
bytes unmodified.
2026-04-22 10:00:00 +08:00
Fini 3ee85c1bc5 test(pen-mcp): large-scale batch_design stress (100-250 ops)
AI orchestrators occasionally emit very large batches (one sub-
agent producing a whole section in a single batch_design call).
Existing multi-line regression only covered pretty-printed JSON
in SINGLE ops — nothing pinned behavior when N itself grows.

5 scenarios:
  1. 100 sibling I() ops → all land, <5s wall-clock
  2. 200 sibling ops → 2x node count, <10s (catches O(n²) regressions)
  3. 30-level nested I() chain via parent_id threading
  4. 250 mixed ops (50 sections × 4 children) with parent refs
  5. Partial failure: 1 bogus parent_id among 100 good ops — good
     ones still land (don't let one bad op poison the batch)

Observed: 250-op mixed batch completes in ~42ms on an M-series
machine. Budgets are "reasonable" (5s / 10s / 15s), not "fast" —
they're meant to catch O(n²) regressions in the DSL parser / tree
insert / save loop, not enforce a perf target.
2026-04-22 09:55:00 +08:00
Fini df9f37522f test(pen-mcp): composition — N-tool real screens (6 scenarios)
Per-tool handler tests cover "this tool emits the correct shape"
individually. This file covers the next layer: can N element-tool
calls chain together into a realistic multi-section screen without
breaking tree invariants?

Scenarios (each spans multiple tool families to catch cross-
family regressions):
  1. Mobile settings — top_nav + 2 sections × 3 list_rows + bottom_nav (10 calls)
  2. Dashboard home — top_nav + stat_grid + section + 3 metric_comparisons + chart (7)
  3. Login form — heading + body + 2 form_fields + button + link (6)
  4. Profile + UGC — top_nav + avatar + heading + badge + 2 faq_items + action_menu (7)
  5. Listing — search + card_row + divider + empty_chart + date_picker + chip_input + pagination (7)
  6. parent_id threading invariant — nested insert actually lands under named parent

Each scenario asserts:
  - Every call emits a nodeId (no silent no-ops)
  - Final document parses as valid JSON with expected root children count
  - Every call's nodeId is findable in the saved tree
  - Every tool's canonical role survives post-save
  - parent_id threading works (child lands under named parent, not root)

This is the integration gate that catches "tool wiring works
individually but composes wrong" — the ghost regression that can
slip past per-tool tests.
2026-04-22 09:50:00 +08:00
Fini 4a75b53c45 feat(ai): add_date_picker_v0 — date input closed state (62nd tool)
Adds an N-tool for the labeled date input + calendar-icon trigger.
Emits ONLY the CLOSED state; the open month grid lives in
add_calendar_grid_v0 and is typically shown inside a popover,
not stacked directly below. Two visual states:

- placeholder (no value): slate-400 "Select date" + calendar icon
- populated (value): slate-900 date text + calendar icon

`clearable: true` adds a small X affordance to the right of the
value (only when value is present — no-op for the placeholder
state since there's nothing to clear). `required: true` appends
" *" to the label.

Keeping closed + open as separate tools is intentional: AI specs
often ask for only the closed trigger inside a form, and a single
combined tool would either force an unwanted grid or require a
mode flag that splits the parameter surface. Separate narrow tools
compose cleanly via batch_design when a designer DOES want both.

Wired through all 3 paths + parity/contract/design-prompt tests.
Handler test covers 7 cases: placeholder state fills, populated
value fills, clearable X behaviors (both present + absent value),
custom placeholder override, required marker, bogus parent_id
rejection.
2026-04-22 09:45:00 +08:00
Fini 1b4f362d59 feat(ai): add_action_menu_v0 — context/kebab dropdown panel (61st tool)
Adds an N-tool for the floating card that drops from a "⋯ more"
button or appears on right-click. Emits the OPEN state: vertical
stack of padded icon+label rows in a white card with subtle stroke
and shadow. Positioning and show/hide are caller concerns (same
philosophy as add_modal_shell_v0 / add_toast_v0).

Destructive items (destructive=true) render in red with role
`action-menu-item-destructive` so renderers can style the hover
state separately. divider_before=true on any item (except first,
where it's ignored) inserts a 1px hairline above — useful for
"Edit / Share / Report / Delete" grouping patterns.

Wired through all 3 paths + parity/contract/design-prompt tests.
Handler test covers 7 cases: simple list, destructive red fill,
divider between groups, leading-divider ignored, label-only no
icon, width clamp, bogus parent_id rejection.
2026-04-22 09:40:00 +08:00