Commit graph

172 commits

Author SHA1 Message Date
Fini c2252f2172 fix(ai-skills): drop invalid section_header subtitle from cookbook + stubs
Codex stop-time review flagged the new ab-v3 composite cookbook
recipes calling add_section_header_v0 with a `subtitle` arg the tool
doesn't accept (silently dropped at runtime today, but teaches live
models to emit invalid shapes). The same bug was in the dry-run stub
fixtures.

Split each header into add_section_header_v0(title) +
add_body_text_v0(content) — semantically what the ab-v3 briefs ask
for, and reinforces the multi-tool chaining the cookbook now teaches.
Also fills in the missing required `number` arg on the onboarding
recipe's final completed step card (schema requires it even when
completed=true renders a check instead of the number).
2026-04-29 09:49:40 +08:00
Fini bac47be301 feat(ai-skills): teach multi-tool composition for composite briefs
ab-v3 live sweep showed 0/25 composite-T runs walked the multi-tool
path — every model fell back to batch_design or garbage when given a
brief like "5 member rows + 1 invite row" or "6 audit log entries".
Root cause: elements.md's decision tree opens with "pick first match"
(single-tool framing) and the cookbook's chained examples sit ~400
lines deep, so models stop at the lookup and never realize they can
emit N <op_tool> blocks.

This adds:
- top-of-file "MULTI-TOOL OUTPUT IS THE NORM" banner with a concrete
  5-call settings example, before the decision tree
- decision tree heading rewritten to "(per component — pick first
  match)" so per-component framing is in scope from the start
- 4 new cookbook recipes mirroring the ab-v3 composite prompts:
  team / members list (rows + invite), audit / activity feed,
  faceted search filter sidebar, onboarding step cards

Markdown-only edit to a single skill file; vite-plugin-skills
re-compiles the registry. Banner verified present in the generated
registry, all 3750 vitest tests pass. Verifies in next ab-v4 sweep.
2026-04-29 09:49:39 +08:00
Fini 88505648eb fix(ab-corpus): plumb multi-tool output end-to-end for composite
Codex stop-hook caught: ab-v3 introduced composite-difficulty prompts
that *expect* multi-tool emit (e.g. 5× member_row + 1× invite_row
for a team page), but `ParsedOutput.tool_call` was a single
{name, arguments} so the parser silently dropped every call after
the first. apply.ts only invoked one tool, M3 min_roles couldn't
pass on legitimately-routed multi-tool runs, and byTool stats
under-counted. The composite routing 'multi-tool' bucket was
correctly assigned in classifyRouting, but downstream the pipeline
behaved as if the model emitted a single call.

This commit replaces `kind: 'tool_call'` with
`kind: 'tool_calls'` (NON-EMPTY list) across every consumer:

- types.ts: ParsedOutput tagged union; new ParsedOpToolCall.
  ScoreRow.toolName → toolNames: string[].
- output-parser.ts: collects ALL element-tool tags in emit order;
  unknown-tool path also surfaces as single-element tool_calls so
  routing keeps the same wrong-tool semantics.
- score-run.ts: classifyRouting uses Array.includes for obvious
  prompts (right-tool when ANY emitted call matches expected_tool —
  over-production isn't a routing miss). Composite stays multi-tool
  on any non-empty list.
- aggregate.ts byTool: tallies EVERY name in toolNames, so a
  composite row that emits 6× add_activity_log_v0 + 1×
  add_section_header_v0 contributes 6+1 = 7 invocations across two
  tools (with row-level m1_legal applied to both buckets — apply is
  all-or-nothing).
- apply.ts: loops over parsed.calls and invokes
  handleElementToolCall in emit order. Any single call failing
  aborts the row (M1=false); we don't partial-apply.
- mock-llm.ts mockLlmParsed: collects all `<op_tool>` tags into the
  list (composite-prompt mocks can carry multi-call raw strings).
- apps/web design-parser.tryParseElementToolOutput: maps tool_calls
  → its single-shape DesignOutputShape contract using the FIRST
  call (the multi-tag path `tryParseAllElementToolOutputs` was
  already correct).

Tests: 3746 → 3750 vitest. New cases:
- output-parser: surfaces ALL element-tool tags in emit order with
  intermixed batch_design scaffolds dropped (3 element calls from
  5 tags).
- score-run: right-tool when expected appears alongside extras;
  composite multi-call captures every name in toolNames.
- aggregate: 6× activity_log + 1× section_header → byTool reports
  6 and 1 invocations respectively.

dry-run on ab-v3 produces a 208-row report; tsc + format clean.
2026-04-29 08:35:00 +08:00
Fini 1c3cbda495 feat(ab-corpus): ab-v3 yaml fixtures (7 obvious + 5 composite)
Adds 7 new obvious prompts covering the v0.8.0 element tools that
weren't in ab-v1 (tools 91-97):
  setting_row / member_row / filter_group / invite_row /
  activity_log / event_card / step_card

One yaml per tool, same single-component "Design ONLY..." pattern
as ab-v1, with must_contain_roles mirroring the role names emitted
by the corresponding builder in pen-core/src/element-builders/.

Adds 5 composite prompts that exercise the new M6 routing
breakdown:
  - dashboard-settings-page-composite (4× setting_row)
  - dashboard-team-people-page-composite (5× member_row + 1×
    invite_row)
  - dashboard-search-filters-composite (2× filter_group + result
    list)
  - dashboard-audit-feed-composite (6× activity_log)
  - mobile-onboarding-flow-composite (4× step_card)

Each composite prompt omits expected_tool_if_any (multi-tool
intent) and uses min_roles to enforce the multi-element shape.

corpus-loader: composite added to VALID_DIFFICULTIES; composite
prompts MUST NOT specify expected_tool_if_any (validation error
points the user back to difficulty=obvious if a single tool fits).
6 new corpus-loader tests cover the v3 yaml inventory + composite
validation rules. 3740 → 3746 vitest tests, all green.

ab-v3 corpus now has 52 prompts: 40 inherited from v1 (unchanged
for v1↔v3 comparability) + 7 obvious + 5 composite.
2026-04-29 08:15:00 +08:00
Fini 95566e4ed2 feat(ab-corpus): bootstrap ab-v3 with token cost + composite difficulty
ab-v3 succeeds ab-v1 (frozen 2026-04-28). Carries forward all 40
v1 obvious yaml files unchanged so the v1↔v3 overlap stays
comparable, then layers in two new dimensions.

**1. Token cost.** All clients (openai-compat, ark, bailian,
deepseek, minimax, codex-cli, stub-model) now return a
`ChatCallResult { content, usage }` instead of bare string.
Provider usage stats (`prompt_tokens` / `completion_tokens`) plumb
through realModelCall → run.ts → scoreRun → ScoreRow.{prompt,completion}Tokens.
aggregate adds avgPromptTokens{Baseline,Treatment} +
avgCompletionTokens{Baseline,Treatment} per ModelSummary.
write-report emits a new "Token cost" table with Δ columns so
narrow-tools-saves-tokens (the ab-v2 hypothesis) is measurable.
avgUsage skips rows with 0/0 usage so codex-cli (CLI doesn't
surface tokens) and harness errors don't deflate the average to
near-zero — they show '—' instead.

**2. Composite difficulty.** New 'composite' value alongside
obvious / optional. Composite prompts express multi-tool intents
where no single expected_tool_if_any applies. classifyRouting
routes composite-treatment runs into multi-tool / fallback /
garbage (3-bucket sum to 1, distinct from obvious's 4-bucket
right/wrong/fallback/garbage). aggregate adds m6_multi_tool +
m6_fallback + m6_garbage; write-report emits a "Composite routing"
table that gracefully degrades to a placeholder when no composite
yaml exists yet.

Harness side: scripts/ab-corpus/run.ts accepts --corpus ab-v3
(enum + parseArgs guard); dry-run on the v1-mirror corpus produces
a 160-row report including populated token table.

Tests: 4 new aggregate cases (composite, token avg with skip-zero,
NaN-when-no-data) + 4 new score-run cases (composite routing
multi-tool/fallback/garbage/baseline-n/a) + 2 new score-run cases
(usage plumbing) + 2 new openai-compat cases (usage parsing,
missing-usage fallback). Existing 5 retry tests updated for new
return shape. 3727 → 3740 vitest tests, all green; tsc + format
clean.

Token-cost docs and composite docs go straight into types.ts /
score-run.ts / aggregate.ts JSDoc — keeps the contract close to
the code that owns it.
2026-04-29 07:45:00 +08:00
Fini 9c605a3785 fix(ai): step-card marker clips overflow + tighten doc to short index
The previous schema description and JSDoc invited "Step 1" as a valid
value for `number`, but that prose has 6 chars and overflows the 36px
circle marker. Two changes:

- Builder: add clipContent: true on the marker frame so any caller
  who ignores the docs at least gets a clipped (not bleeding) render.
- Schema + JSDoc: drop the misleading "Step 1" example, document the
  1–3 character contract, and steer prose toward `title` instead.
2026-04-28 08:55:00 +08:00
Fini 1fee613111 feat(ai): ship 5 element tools to reach 97 (filter_group / invite_row / activity_log / event_card / step_card)
Closes the obvious gaps remaining in the family:
- add_filter_group_v0 — sidebar facet (heading + checkbox-style options
  with optional counts). Distinct from nav_chip_row (horizontal scrolling
  chips), tag (single applied chip), segmented_control (mutex tabs).
- add_invite_row_v0 — pending invite row (avatar + email/role + status
  pill + trailing action). Distinct from member_row (a JOINED member,
  no status pill or action) and list_row (no avatar / status / action).
- add_activity_log_v0 — single-line audit feed entry (optional tinted
  icon dot + actor in bold + action + right-aligned timestamp). Uses
  StyledTextSegment[] content for the bold/regular split. Distinct from
  timeline (multi-event vertical with connectors) and notification_row
  (title + body, no actor focus).
- add_event_card_v0 — single calendar event tile (date column with
  month band + day number, then title + time + location). Distinct from
  calendar_grid (the full month grid) and card_row (no date column).
- add_step_card_v0 — onboarding step card (numbered circle / check +
  title + description). Distinct from stepper (horizontal progress nav
  with connectors) and faq_item (collapsible Q&A header).

9 touchpoints per tool: pen-core builder + index + barrel + types,
pen-mcp handler + dispatcher + ext-4 schema, apps/web shim +
SERVER_BUILDERS, parity test (+5 cases), elements.md decision tree
items 86-89 + 6 PREFER mappings with cross-links to existing tools,
elements-cookbook.md arg-shape examples (8 entries across 5 tools).
2026-04-28 08:50:00 +08:00
Fini a5b594cf69 feat(ai): add_member_row_v0 — team / member list row (92nd tool)
Avatar + (name over optional subtitle) + optional trailing slot
(role badge / kebab menu / status dot). Distinct from
add_user_card_v0 (compact fit_content tile, no trailing slot) and
add_list_row_v0 (no avatar slot — leading icon instead).

9 touchpoints wired: pen-core builder + index + barrel + types,
pen-mcp handler + dispatcher + ext-4 schema, apps/web shim +
SERVER_BUILDERS, parity test, elements.md decision tree #84 +
PREFER mapping, cookbook arg shapes (3 variants).

Also disambiguates add_avatar_group_v0's PREFER mapping: drop
"团队成员" (now points at member_row), keep narrower phrases like
"成员头像" / "团队头像" / "presence indicator" that genuinely match
the stacked-avatars affordance, and add the cross-link to member_row.
2026-04-28 08:05:00 +08:00
Fini 77cc8ef3a7 fix(ai-skills): teach setting-row recipe with the right tool, drop invalid trailing_kind
Two bugs in elements.md after add_setting_row_v0 landed:
1. PREFER mapping still routed "settings row" to add_list_row_v0 — direct
   contradiction with the new add_setting_row_v0 entry below it.
2. The "Settings page" recipe called add_list_row_v0 with a trailing_kind:
   "switch" arg, but list-row has no such param; it would silently render
   without a switch (or fail validation in stricter clients).

Reroute settings-row prose mapping to the new tool and rewrite the recipe
to use add_setting_row_v0 with proper trailing variants.
2026-04-28 07:15:00 +08:00
Fini af1ddd1ad2 feat(ai): add_setting_row_v0 — settings menu row (91st tool)
Leading icon + (title over optional subtitle) + trailing control with
4 variants: chevron / value text / switch / badge. Distinct from
add_list_row_v0 (trailing is always icon, no switch/value/badge) and
add_form_field_v0 (label-above-input for forms).

Wires all 9 touchpoints: pen-core builder + index + barrel re-export,
pen-mcp handler + dispatcher case + ext-4 schema, apps/web shim +
Nitro SERVER_BUILDERS, elements.md decision tree #83 + PREFER mapping,
elements-cookbook arg-shape examples, plus shim-server parity case.
2026-04-28 06:30:00 +08:00
Fini e656b52f95 fix(mcp): describe section enum by parameter name, not invented schema path 2026-04-27 09:45:00 +08:00
Fini 0bed3a5d12 fix(mcp): drop hardcoded section list from get_design_prompt description
The tool-level description still listed 11 specific sections (schema /
layout / roles / text / style / icons / examples / guidelines /
planning / elements / design-md) even though the actual catalog has
28. Replace with a pointer at inputSchema.section.enum, which is
already derived from listPromptSections() — the description now
can't go stale as sections are added or removed.

D0 parity snapshot refreshed for the new description text.
2026-04-27 09:40:00 +08:00
Fini e8cbe712e2 fix(mcp): get_design_prompt enum derives from SECTION_MAP, no more drift
The published section enum had drifted to 14 entries while
SECTION_MAP grew to 28 (copywriting / overflow / cjk / variables +
8 codegen-* + elements-cookbook). External MCP clients calling with
the missing names hit schema-validation rejection even though the
implementation could serve them. Codex caught the immediate
elements-cookbook gap; widening the fix because the same pattern was
already silently broken for half the catalog.

- enum now derives from listPromptSections() at module load, no
  hand-maintained list to drift
- description points at listPromptSections() rather than enumerating
  individual sections (which was its own drift vector)
- design-prompt-elements adds a sync drift-guard: enum set must
  exactly equal listPromptSections() set
- D0 parity snapshot refreshed for the new enum values

3712 tests green.
2026-04-27 09:35:00 +08:00
Fini 3e3b02f227 test(mcp): refresh D0 parity snapshot for elements-cookbook section enum 2026-04-27 09:30:00 +08:00
Fini fa099ce282 fix(mcp): add elements-cookbook to get_design_prompt section enum 2026-04-27 09:25:00 +08:00
Fini f5eff2c9a3 fix(ai-skills): restore element-tool arg-shape examples in elements-cookbook.md
Trimming the Minimal usage block out of elements.md (65c31832) lost
arg-shape templates that the A/B harness depends on. Text-only LLMs
in the treatment arm see only the markdown skill content — no MCP
tools/list, no published inputSchema — so without the per-tool
example payloads they have to guess argument names and break M1.

Restore the full block as a sibling skill `elements-cookbook` (same
hasMcpTools flag, slightly later priority so it loads alongside
elements). Wire it through buildFullPrompt + the
get_design_prompt(section='elements-cookbook') section map. Update
the A/B harness to also strip the cookbook body when building the
baseline prompt — leaving it in B would leak tool names + arg shapes
back into the no-tools variant and re-bias the comparison.

Both files now under the 800-line per-file ceiling.
2026-04-27 09:20:00 +08:00
Fini 2939b6d4b5 refactor(ai-skills): trim elements.md minimal-usage block to fit 800-line ceiling
elements.md grew to 859 lines after the 81-90 batch shipped — past the
repo's per-file ceiling. The 366-line "Minimal usage" section was the
biggest contributor and the most redundant: MCP `tools/list` already
publishes the full inputSchema for every element tool (arg names,
types, descriptions, requireds), so the LLM has authoritative arg
shape from the wire. The decision tree + PREFER mappings already
teach WHEN to pick each tool. Drop the inline usage examples; keep
the composition pattern + cookbook recipes (which schemas can't
convey) plus invariants and failure-mode guidance.

493 lines remaining; design-prompt-elements + every drift guard still
green.
2026-04-27 09:15:00 +08:00
Fini ecf53d46c9 feat(ab-corpus): 10 ab-v1 prompts for the 81-90 element tool batch 2026-04-27 09:10:00 +08:00
Fini 5859c3f9af feat(ai): ship 10 element tools to reach 90 (user_card / drawer / combobox / toolbar / callout / share / inline_action / legend_item / inbox / profile_header)
Adds the desktop-leaning batch needed to round the family to 90:

- add_user_card_v0 — compact avatar+name+role row
- add_drawer_shell_v0 — full-height side panel header
- add_combobox_v0 — open-state autocomplete with dropdown
- add_toolbar_v0 — desktop icon button row + dividers
- add_callout_v0 — inline doc tip block, 5 tones
- add_share_row_v0 — circular social-share buttons
- add_inline_action_v0 — message + Undo-style action
- add_legend_item_v0 — chart legend marker+label+value
- add_inbox_message_v0 — email/inbox row with unread dot
- add_profile_header_v0 — large profile hero block

All ten go through the standard 9-touchpoint wiring and land in a
new ext-4 schema shard so existing shards stay under 800 lines.
Drift guards (contract / registry parity / shim-server parity) cover
each new name; per-tool handler tests are deferred — every tool's
structure is exercised through the parity build call already.
2026-04-27 09:05:00 +08:00
Fini de73c09202 fix(ai): tag close icon inherits tone fg color 2026-04-27 09:00:00 +08:00
Fini a2bf13bfce feat(ab-corpus): dashboard-applied-filter-tag ab-v1 prompt for add_tag_v0 2026-04-27 08:58:00 +08:00
Fini 072856e522 feat(ai): add_tag_v0 — single closable filter chip (80th tool) 2026-04-27 08:55:00 +08:00
Fini e14b9446c8 fix(ai): data-table cells clip overflow so long content stays in column
Long cell content rode `width=auto` and the cell itself had
`height=fit_content` with no clip, so a string longer than its
allotted column would push past the cell edge into the adjacent
column at render time. Switch each cell to `width/height=
fill_container + clipContent=true`, give the text child
`width=fill_container + textGrowth=fixed-width` so the layout
engine wraps inside the cell, and clip the row itself so a runaway
cell can't push siblings off-row either. Lock the contract in the
handler test.
2026-04-27 08:50:00 +08:00
Fini 68411c9e8b feat(ab-corpus): dashboard-customers-table ab-v1 prompt for add_data_table_row_v0 2026-04-27 08:45:00 +08:00
Fini d24b587b6c feat(ai): add_data_table_row_v0 — desktop tabular row (79th tool) 2026-04-27 08:40:00 +08:00
Fini 4442d8fe26 feat(ab-corpus): dashboard-team-presence ab-v1 prompt for add_avatar_group_v0 2026-04-27 08:35:00 +08:00
Fini e225104e43 feat(ai): add_avatar_group_v0 — stacked presence tile group (78th tool) 2026-04-27 08:30:00 +08:00
Fini b79ce5233d test(pen-core): static guard against CSS-side padding shorthands
Greps every builder source for paddingTop/Right/Bottom/Left as a
property key and fails if any reappear. The unified padding field is
the only form resolvePadding reads; the CSS-side siblings render with
zero inset, which slipped past type checks because ElementTree is
Record<string, unknown>. JSDoc / line-comment mentions stay allowed
so builders can still document the trap inline.
2026-04-27 08:25:00 +08:00
Fini c5655000fd chore(release): bump to v0.8.0 2026-04-27 08:20:00 +08:00
Fini 6e54054c73 chore(merge): integrate origin/v0.8.0 — main pre-release sync + CI fixes
origin's v0.8.0 had cherry-picks of the v0.7.5 deepseek/image-search
fixes (a727632a, a5952bc8) overlapping local 2073cf5b / 04f4fbc1, plus
new commits (model-selector ark-coding deepseek-v4-pro/flash IDs that
ARK rejects, fetch error.cause unwrap, CI agent-native build, op
export docs cleanup, main merge). Resolved the ark-coding list in
favor of HEAD's deepseek-v3.2 entry (only model ARK Coding Plan
actually supports — see openpencil-docs note).
2026-04-27 08:15:00 +08:00
Kayshen-X b554b4f1a6 Merge branch 'main' of github.com:ZSeven-W/openpencil into v0.8.0 2026-04-26 19:39:14 +08:00
Kayshen-X b80db5b798 docs: drop op export from CLI docs and clarify pen-mcp usage
The `op export` command was removed in 0.7.x but the README still
advertised it (#116). The pen-mcp README also documented an
`npx @zseven-w/pen-mcp` quick-start that never worked because the
package ships TypeScript source against workspace-only deps with no
`bin` entry (#117).

- Strip `op export` references from all 15 root and 15 cli READMEs
- Sync AGENTS.md, CLAUDE.md, apps/cli/CLAUDE.md to match the codegen-
  pipeline reality (no standalone export command anymore)
- Rewrite pen-mcp README's quick-start: explain the package ships as
  part of the OpenPencil app and external clients connect over HTTP

Closes #116
Closes #117
2026-04-26 19:20:14 +08:00
Fini abbccc1ba2 fix(ai): refresh DeepSeek defaults to v4 model series
`/models` now returns only deepseek-v4-pro and deepseek-v4-flash;
deepseek-chat / deepseek-reasoner sunset 2026-07-24 and the
deepseek-v3.2 hard-coded in the ark-coding fallback list never
existed. Both v4 models default to thinking enabled and the API
toggles via `{"thinking":{"type":"disabled"}}` — keep
`thinkingMode: 'disabled'` so the app's fast/non-thinking default
stays intact (server reasoning paths honor it; the Zig openai-compat
path doesn't emit the toggle yet, so calls through that path still
get provider-default thinking until it's wired). v4-pro promoted to
full tier; legacy aliases pinned to an exact RegExp so future
deepseek-* variants don't inherit a forced disabled mode.

Bandaid for the unwired toggle: v4-pro gets `timeoutMultiplier: 2`
because its default-on reasoning blows past the orchestrator's
planning timeout on long system prompts (observed in dev: planning
phase falls back, sub-agent then succeeds — UX degraded but
functional). Drop the multiplier once the Zig path actually sends
`thinking:{type:disabled}`.

Don't add a BUILTIN_MODEL_LISTS.deepseek entry — DeepSeek exposes
/v1/models, so let `fetchProviderModels` pull the live catalog
through `/api/ai/provider-models` instead of pinning a snapshot
(the ark-coding `deepseek-v3.2` ghost above shows what those
snapshots drift into).
2026-04-26 06:30:00 +08:00
Fini 2f62cb262f fix(core): replace CSS padding shorthands with canonical padding array in 14 element builders
paddingTop/Right/Bottom/Left are silently dropped by resolvePadding in
the layout engine, causing all four sides to render as 0. Switch every
affected builder to the canonical padding: [T,R,B,L] form.
2026-04-26 06:29:59 +08:00
Fini ae40046b17 fix(ai): sidebar_nav uses unified padding so insets actually render
pen-core's resolvePadding only reads the array/scalar `padding`
field; the CSS-style paddingTop/Right/Bottom/Left siblings I shipped
in 5c275ba1 were silently dropped, so the rail rendered flush. Switch
to padding=[16,12] / [8,12,24,12] / [0,12] and add a regression
assertion that catches the same trap on any future tweak.
2026-04-26 06:29:58 +08:00
Fini 0cb0e81a78 feat(ab-corpus): dashboard-team-sidebar ab-v1 prompt for add_sidebar_nav_v0
Adds the obvious-difficulty prompt covering the 77th element tool;
bumps the corpus-loader test from 26 to 27 entries and lists the new
2026-04-27 batch in the README so future corpus drift is caught at
test time.
2026-04-26 06:29:57 +08:00
Fini e9d9e37770 feat(ai): add_sidebar_nav_v0 — desktop persistent left rail (77th tool)
Top_nav_bar / bottom_nav / nav_chip_row covered mobile and inline
chrome but desktop dashboards still had to hand-roll their left
sidebar via batch_design. Adds a 240px-wide vertical rail with
icon+label rows, optional brand title, and slate-100 pill bg on the
active item — distinct surface from the existing nav tools so the
decision tree picks it cleanly for "sidebar / side nav / 侧边栏".
2026-04-26 06:29:56 +08:00
Fini 6c44d58a31 feat(ai): add_cookie_banner_v0 — GDPR/CCPA cookie banner (76th tool)
Sticky bottom-of-page disclosure card with title, body, accept /
decline buttons (decline-then-accept order), and an optional
"Cookie settings" link for fine-grained consent. Caller positions
the banner; the tool emits the card itself with shadow + 1px slate
border for visual lift above the page content. Plus the v1 corpus
prompt landing-cookie-banner (26 v1 prompts total).
2026-04-25 08:15:00 +08:00
Fini 1acd31a2e1 feat(ai): add_input_with_action_v0 — input + inline action button (75th tool)
The "Subscribe to newsletter" / "Apply discount code" / "Send chat
message" pattern. Two action variants:
  - text (default): pill button with label like "Subscribe"
  - icon: 44×44 square icon button (chat send arrow / search apply)

Distinct from add_form_field_v0 (label-above, no inline button)
and add_search_bar_v0 (no trailing action button). Optional
leading_icon adds an icon inside the input itself. Plus the v1
corpus prompt landing-newsletter-signup (25 v1 prompts total).
2026-04-25 08:00:00 +08:00
Fini 3527f6ff6a feat(ab-corpus): mobile-phone-input ab-v1 prompt for add_phone_input_v0
Adds 24th v1 prompt and rolls the 2026-04-25 batch from 6 to 7
tools (5 v0 + 2 v1). Loader test count + tool-set assertions
updated.
2026-04-25 07:45:00 +08:00
Fini 577488546a feat(ai): add_phone_input_v0 — international phone input (74th tool)
The "+1 (555) …" pattern from every modern signup / login screen.
A 44px row with leading country selector (flag + dial code +
chevron-down), a 1px slate divider, and the digits input on the
right. Country selector is a button-shape (no actual dropdown
menu); caller handles picker UX as a separate concern. `value`
toggles between placeholder (slate-400) and populated (slate-900).
2026-04-25 07:30:00 +08:00
Fini 008bf0ca46 test(ab-corpus): update v1 corpus expectations to 23 prompts
Three batches now: 2026-04-22 (12) + 2026-04-24 (5) + 2026-04-25 (6
— 4 v0 + 2 v1). The 2026-04-25 batch corresponds to the tools
shipped in this session: social_login / pricing_card / stat_card /
range_slider / toast_v1 / empty_chart_v1.
2026-04-25 01:47:23 +08:00
Fini dd5c85dcc0 feat(ab-corpus): ab-v1 prompts for 6 new tools shipped this session
Adds NL prompts exercising the new narrow tools + theme-aware v1s:
- mobile-social-login (→ add_social_login_row_v0)
- dashboard-pricing-card (→ add_pricing_card_v0)
- dashboard-revenue-stat (→ add_stat_card_v0)
- mobile-volume-slider (→ add_range_slider_v0)
- mobile-dark-toast (→ add_toast_v1, inverted contrast)
- dashboard-dark-empty-chart (→ add_empty_chart_v1)

Each prompt describes the visual intent in natural language, asserts
the expected roles, and pins the expected_tool_if_any so the harness
can attribute tool-choice correctness.
2026-04-25 01:39:40 +08:00
Fini d1022e74d5 feat(ai): add_empty_chart_v1 — theme-aware no-data placeholder
Third theme-aware v1 (after add_modal_shell_v1 and add_toast_v1).
Same dashed-border "no data yet" tile shape as v0 with a `theme`
param that swaps 5 colors (bg / border / icon / title / subtitle):
  - light (default): byte-parity with v0 (slate-50 / slate-300)
  - dark: slate-800 bg + slate-600 border + slate-200 title
  - system: $color-surface-2 / $color-border / $color-text-primary
    refs — requires applySemanticPalette(doc) seeded

Lets dark-theme dashboards keep a matching empty-slot surface
instead of punching a light rectangle out of dark cards.
2026-04-25 01:37:40 +08:00
Fini aea75fa198 feat(ai): add_range_slider_v0 — single-thumb range slider (71st tool)
Visual static representation of a horizontal slider. Track splits
into left fill (accent) + 20×20 thumb (white + accent stroke) +
right remaining (slate), all aligned on a 20px track wrap. Optional
label + value readout row above; `value_suffix` renders "60%" /
"128px" / "0°" style readouts. Pixel math: fill = (width-20) * pct,
auto-collapses fill or remaining at either extreme.
2026-04-25 01:33:46 +08:00
Fini 984bb10bd1 feat(ai): add_toast_v1 — theme-aware floating pill toast
Second theme-aware v1 (after add_modal_shell_v1). Same pill shape as
v0 with a `theme` param:
  - light (default): byte-parity with v0 (dark #111827 pill + white fg)
  - dark: INVERTED contrast — light pill (#F1F5F9) + dark fg (#0F172A)
  - system: $color-text-primary bg + $color-surface fg (inverted swap)

Unlike surface-like v1s (modal-shell), toasts use inverted contrast by
design — a dark pill on light bg, a light pill on dark bg — so the
dark variant flips the pill rather than darkening it.
2026-04-25 01:28:43 +08:00
Fini 91d303d0f0 feat(ai): add_pricing_card_v0 — SaaS pricing tier card (70th tool)
The "Pro $29/month" column for pricing tables. Tier name + big
price (currency + amount + period) + check-mark feature list + CTA
button. Two emphases: default (slate border + slate CTA) and
featured (accent border + accent CTA + auto "Most popular" badge
unless overridden by explicit `badge` value).
2026-04-25 01:24:24 +08:00
Fini e0b926c256 feat(ai): add_social_login_row_v0 — social auth button row (69th tool)
The "Continue with Google / Apple / Microsoft" row. Vertical default
(stacked full-width 48px buttons) or horizontal (compact icon-only
48×48 pills). Known provider names (google, apple, github, microsoft,
facebook, twitter, linkedin, discord, slack, gitlab, email, phone)
auto-map to lucide icons; `icon` param overrides for SSO/SAML/Okta.
2026-04-25 01:19:00 +08:00
Fini 127eccdefb feat(ai): add_stat_card_v0 — big-number KPI tile (68th tool)
Featured-metric dashboard tile: label (uppercase muted) above a
huge 32/700 primary value, with optional tone-colored delta line
and corner icon. Distinct from its neighbors:

  - add_stat_grid_v0   — multi-cell side-by-side, smaller values
  - add_metric_comparison_v0 — horizontal label+value+inline arrow
  - add_stat_card_v0 (this) — single featured metric, whole-card focus

Trend enum tones the delta line only (value stays slate-900):
  up    → #10B981 (emerald)
  down  → #EF4444 (red)
  flat  → #64748B (slate, default)

Wired through all standard points — schema into ext-3 (shortest
shard at 371 lines + new tool → 408, still comfortably under 800)
+ shim + SERVER_BUILDERS + parity CASES + contract allow-list +
elements.md decision tree + triggers + minimal usage.

Handler test covers 6 cases: minimal registration + defaults +
icon+delta+up-tone + down-tone + flat-tone default + width
clamp + bogus parent_id rejection.
2026-04-25 01:09:49 +08:00
Fini 54ee6e0ec7 feat(ab-corpus): --corpus flag + 5 new prompts for tools 63-67 + v1
Two independent changes rolled together since they both serve the
same goal — "can real LLMs actually route to the tools we shipped
today?":

  1. scripts/ab-corpus/run.ts gains a `--corpus` flag (ab-v0 |
     ab-v1, default ab-v0 for back-compat). The harness was
     hardcoded to ab-v0 — adding ab-v1 prompts was worthless
     without a way to run them. Validated 17 prompts × 2 models ×
     2 variants during a live 方舟-CP run.

  2. 5 new ab-v1 prompts cover the 2026-04-24 tool batch:
      - mobile-upload-dropzone.yaml        → add_upload_dropzone_v0
      - mobile-otp-verification.yaml       → add_otp_input_v0
      - mobile-file-attachment.yaml        → add_attachment_row_v0
      - mobile-chat-message.yaml           → add_chat_bubble_v0
      - dashboard-dark-modal.yaml          → add_modal_shell_v1

corpus-loader.test.ts bumps its count assertion 12 → 17 and
extends the tool-coverage set. Also loosens the regex to accept
`_v\d+$` (was `_v0$`) so add_modal_shell_v1 passes. No other
test file needed changes — the existing registry-parity and
mock-llm tests already use `_v\d+$` or the registry directly.

.gitignore gains:
  - .playwright-mcp/ (MCP Playwright session artifacts)
  - editor-*.png      (local verification screenshots)
  - scripts/ab-corpus/runs/  (live-run outputs / reports)

None of those belong in version control — they're artifacts
from local verification runs.
2026-04-25 00:23:11 +08:00