Commit graph

178 commits

Author SHA1 Message Date
Fini 2902bdc88e feat(ai-skills): teach parent_id rule + trim PREFER list "Different from"
Two diet/teaching changes to elements.md, both motivated by today's
ab-v4 smoke results:

(a) parent_id teaching — minimax-m2.7 invented "members-section" /
    "canvas" as parent_id values on its multi-tool composite output,
    then every one of its 14 tags failed apply with "parent_id X not
    found in document". Same pattern as yesterday. Adds an explicit
    rule near the top of elements.md (right under the multi-tool
    banner) — `parent_id` is REAL or OMITTED, never invented. The
    `<page>` / `<panel>` / `<sidebar>` placeholders in the cookbook
    recipes are documentation conventions; in actual output, OMIT
    the field. Names the failure mode by reproducing the error
    message format so models learn to avoid it.

(b) Phase 3 token diet — strip ". Different from <tool> (...)"
    disambiguation suffixes from the PREFER list (22 entries had
    them, ~80-200 chars each = ~2.5kb / ~625 tokens saved). The
    primary keyword + tool-name + capability description survives
    intact; the cross-references pointing at sibling tools get
    dropped. Risk: slight increase in wrong-tool routing on
    ambiguous prompts. Worth it for the size reduction; the
    decision tree alone still shows the tool family in context.

Verified by mechanical diff — perl in-place edit, then visual
review confirms no other content was touched. 3785 vitest pass,
format clean, tsc silent.

Net effect on T-prompt size (chars / 4 estimate):
  T + composite + mobile: was 17,930 → now 17,475 (-455 chars)
  T + composite + dashboard: was 18,437 → now 17,983 (-454 chars)
  T + obvious + dashboard: ~14,009 → ~13,556 (-453 chars)

Per-arm savings are smaller than the raw 2.5kb trim because (a)
adds ~700 chars of parent_id teaching. Net win: ~450 chars / ~110
tokens per call. Modest but compounds across 520+ ab-v4 runs.
2026-04-29 09:49:58 +08:00
Fini bafbdfaeef fix(pen-core): reject invalid enum values in 5 more element builders
Sweep follow-up to 113bd55a — same defensive pattern (reject
unknown enum strings at the entry boundary) applied to every other
builder that indexed a Record<EnumLiteral, T> with a value sourced
from raw JSON args.

Builders + enums covered:
- buildTag — TagTone (default | accent | success | warning | error)
- buildCallout — CalloutTone (info | success | warning | danger | note)
- buildActivityLog — tone (info | success | warning | danger | neutral)
- buildInviteRow — InviteStatus (pending | expired | accepted)
- buildMemberRow — trailing.tone for status_dot (online | busy | away
  | offline). role_badge / menu variants skip the check (no tone field)

Same failure mode each one fixed: when a model invents an
out-of-enum string (gpt-5.4 did this with `level: "caption"` in
ab-v4), the lookup `TONES[bad]` / `STATUS_TONE[bad]` returned
undefined, the next property access crashed mid-batch with a
cryptic `undefined is not an object`, and the surrounding dispatch
loop dropped every remaining tag (until df33e937 + 07639f6d landed
the per-shape continuation + partial-doc scoring earlier today).
With validation in place, a bad enum becomes a clean per-shape
error message + the rest of the batch still applies.

13 new edge-case tests cover throw on bad input + valid path on
every enum value + omitted-default for each builder. 3785 vitest
pass, format clean, tsc silent.

Builders not touched: heading.ts (already done in 113bd55a).
Builders that don't fit this pattern (no enum→Record lookup of a
user-controlled string): everything else surveyed via grep on
`Record<.*Tone|Status|Level|Mode|Kind`.
2026-04-29 09:49:57 +08:00
Fini b4b931211c fix(pen-ai-skills): score partial PenDocument when apply.ok=false
Codex stop-time review caught the previous fix (df33e937) handing
the scorer a partial PenDocument that the scorer immediately
ignored. score-run.ts:87 short-circuited on `!applied.ok || !applied.doc`,
so even though apply.ts now surfaces 12 of 13 successfully-applied
tags as a populated `applied.doc`, the row still scored as a total
failure (M1=false, M3=false, m3_failure_reason="apply failed before
shape checks") — exactly the noise df33e937 was meant to eliminate.

Loosens the short-circuit to `!applied.doc` only. When apply.ok=false
but apply.doc is populated, the scorer now:
  - runs the issue detector against the partial doc (issues surface)
  - keeps M1 strict (apply.ok=false → M1=false regardless of detector)
  - decouples M3 from M1: M3 = shape.ok against the partial doc
  - sets m3_failure_reason to the shape miss when shape fails;
    otherwise to "partial apply (M3 met by what landed): <error>" so
    the row reads "tag 12 of 13 broke, but role coverage still met"
    instead of silently swallowing the partial signal
  - surfaces applied.error in row.applyError so per-shape failure
    messages flow into reports

Two new tests cover the new path:
  1. partial apply + shape match → M1=false, M3=true, reason mentions
     "partial apply"
  2. partial apply + shape miss → M1=false, M3=false, shape-miss
     reason wins (structural verdict trumps the partial-apply notice)

Plumbing chain across today's session is now consistent:
  - apply.ts continues past per-shape failures (df33e937)
  - score-run.ts scores the partial doc that lands (this commit)
  - scoring no longer over-attributes to "apply failed" when the model
    actually produced most of the brief

3774 vitest pass (+2), format clean, tsc silent.
2026-04-29 09:49:56 +08:00
Fini 39aeb0c30e fix(pen-core): reject invalid level in buildHeading with a clear error
ab-v4 partial sweep (2026-05-01) caught gpt-5.4 emitting
`add_heading_v0({"content":"Pending invitations","level":"caption"})`
as the 12th tag of a 13-tag composite multi-tool response. Even
though the MCP tool def has `enum: ['display','h1','h2','h3']`, the
ab-corpus harness and the in-process production dispatcher both call
buildHeading() with raw JSON args (no jsonschema gate), so the model's
invented "caption" reached the preset lookup. LATIN_PRESETS["caption"]
is undefined, and the next line `fontSize: preset.fontSize` crashed
the WHOLE batch with `undefined is not an object (evaluating
'preset.fontSize')` — the 11 valid tags ahead of it never landed.

Adds an entry-point validation in buildHeading: if `level` is set and
not in the {display, h1, h2, h3} set, throw with a clear message.
The dispatch loops in apps/web/element-tools-dispatcher and
scripts/ab-corpus/apply both catch per-shape and keep running the
remaining tags, so a single bad level on tag 12 no longer kills tags
1-11 + 13.

3 new edge-case tests in element-builders-edge-cases.test.ts
cover the throw + the four valid levels + the omitted-default case.
3772 vitest pass (+3), format clean, tsc silent.

Other element builders likely have the same pattern (preset lookup
on a string enum without runtime validation) — separate sweep, not
shotgunning here.
2026-04-29 09:49:54 +08:00
Fini c835976479 feat(ab-corpus): per-domain cookbook filter (Phase 2 of token diet)
ab-v3 / ab-v4 showed Phase 1A (cookbook strip on obvious difficulty)
shaved ~4.4k tokens off T-obvious. Phase 2 adds a per-category gate
that strips cookbook recipes whose domain doesn't match the prompt's
category — mobile briefs don't see dashboard recipes, dashboard
briefs don't see mobile / landing recipes, etc.

Mechanism: HTML comment block markers in elements.md
(`<!-- @domain:dashboard --> ... <!-- /@domain -->`) plus a
stripNonMatchingDomains() pass in buildSystemPrompt that drops blocks
whose tag list doesn't include the active category. Untagged content
is "general" and stays in every variant — the safe default.

Tagged 7 single-domain cookbook recipes:
- dashboard: Team / members list, Audit / activity feed, Faceted
  search filter sidebar, Dashboard KPI strip
- landing: Pricing section
- mobile: Onboarding "How it works", Support chat thread

Cross-domain recipes (Login, Signup, Settings page, OTP, Empty
inbox) stay untagged so they load for every category. Decision tree
+ PREFER list also untagged today; the per-tool annotations there
would be a much larger judgment pass for marginal additional savings.

Token measurements (chars / 4 estimate):
                       full     mobile  dashboard  landing
- T + composite       19.0k    17.9k    18.4k     17.7k
                              (-1.1k)  (-0.6k)   (-1.3k)
- T + obvious         14.6k    13.5k    14.0k     13.3k
                              (-1.1k)  (-0.6k)   (-1.3k)

Modest absolute savings — Phase 2 only filters cookbook RECIPES (in
elements.md), and most cookbook content is in elements-cookbook.md
which Phase 1A already strips on obvious. To hit the 6-8k T target
we still need decision-tree compression or PREFER-list trim, but
both are lossier than this gate. Phase 3 candidates noted in the
ab-v4 results doc.

real-model.ts plumbs call.prompt.category through to buildSystemPrompt.
3766 vitest pass (+6 category filter tests including a 500-char
floor regression guard that the filter actually shaves bytes).
2026-04-29 09:49:49 +08:00
Fini cae5501021 fix(ai-skills): purge mixed-strategy teaching from elements.md cookbook
Codex stop-time review (3rd round) caught residual "T prompt
includes mixed-strategy instructions": even after the prompt
forbade mixing, elements.md still taught the mixed pattern in
several places that are also part of the T system prompt.

Cleaned out every spot that paired batch_design with add_*_v0:

- Login screen / Pricing section / Dashboard KPI strip recipes:
  dropped the leading `batch_design: foo = I("page", {...})` line
  and renamed `<foo>` / `<row>` placeholders to `<page>` so each
  recipe is now Strategy A (element tools only). Lost: explicit
  page-level layout/padding/horizontal row — acceptable, the
  recipes still teach the tool selection + chain pattern.

- Intro paragraph: dropped "override via a follow-up batch_design
  U-op if needed" (taught a per-component fallback that the parser
  drops).

- Banner: rewrote the fall-back clause from "when no element tool
  fits a specific component shape" to "when at least one component
  truly needs a custom shape no element tool covers — and then use
  a SINGLE batch_design for the WHOLE response, never mixed."

- "STILL use batch_design when" list: collapsed 3 mixed-strategy
  bullets ("larger composite via batch_design then element tool",
  "post-hoc styling via batch_design U-ops") into 3 clean
  Strategy-B-only bullets, all explicitly emit a SINGLE batch_design.

- Removed the "## Composition pattern" section entirely — its 3-step
  plan was a textbook mixed pattern (batch_design root → element
  tool inserts → batch_design U-op styling) that depended on real
  MCP multi-round semantics the corpus harness can't provide.

- Composition rules of thumb: replaced "Don't mix N-tool and
  batch_design DSL ops in a single call" + "Style overrides come
  AFTER structure" with a single "Don't mix in the same output"
  rule that names the corpus parser behavior explicitly.

3760 vitest pass, format clean, tsc silent. T-obvious 14.5k tokens
(down ~100 from the cleanup); composite still 18.9k.
2026-04-29 09:49:46 +08:00
Fini c2252f2172 fix(ai-skills): drop invalid section_header subtitle from cookbook + stubs
Codex stop-time review flagged the new ab-v3 composite cookbook
recipes calling add_section_header_v0 with a `subtitle` arg the tool
doesn't accept (silently dropped at runtime today, but teaches live
models to emit invalid shapes). The same bug was in the dry-run stub
fixtures.

Split each header into add_section_header_v0(title) +
add_body_text_v0(content) — semantically what the ab-v3 briefs ask
for, and reinforces the multi-tool chaining the cookbook now teaches.
Also fills in the missing required `number` arg on the onboarding
recipe's final completed step card (schema requires it even when
completed=true renders a check instead of the number).
2026-04-29 09:49:40 +08:00
Fini bac47be301 feat(ai-skills): teach multi-tool composition for composite briefs
ab-v3 live sweep showed 0/25 composite-T runs walked the multi-tool
path — every model fell back to batch_design or garbage when given a
brief like "5 member rows + 1 invite row" or "6 audit log entries".
Root cause: elements.md's decision tree opens with "pick first match"
(single-tool framing) and the cookbook's chained examples sit ~400
lines deep, so models stop at the lookup and never realize they can
emit N <op_tool> blocks.

This adds:
- top-of-file "MULTI-TOOL OUTPUT IS THE NORM" banner with a concrete
  5-call settings example, before the decision tree
- decision tree heading rewritten to "(per component — pick first
  match)" so per-component framing is in scope from the start
- 4 new cookbook recipes mirroring the ab-v3 composite prompts:
  team / members list (rows + invite), audit / activity feed,
  faceted search filter sidebar, onboarding step cards

Markdown-only edit to a single skill file; vite-plugin-skills
re-compiles the registry. Banner verified present in the generated
registry, all 3750 vitest tests pass. Verifies in next ab-v4 sweep.
2026-04-29 09:49:39 +08:00
Fini 88505648eb fix(ab-corpus): plumb multi-tool output end-to-end for composite
Codex stop-hook caught: ab-v3 introduced composite-difficulty prompts
that *expect* multi-tool emit (e.g. 5× member_row + 1× invite_row
for a team page), but `ParsedOutput.tool_call` was a single
{name, arguments} so the parser silently dropped every call after
the first. apply.ts only invoked one tool, M3 min_roles couldn't
pass on legitimately-routed multi-tool runs, and byTool stats
under-counted. The composite routing 'multi-tool' bucket was
correctly assigned in classifyRouting, but downstream the pipeline
behaved as if the model emitted a single call.

This commit replaces `kind: 'tool_call'` with
`kind: 'tool_calls'` (NON-EMPTY list) across every consumer:

- types.ts: ParsedOutput tagged union; new ParsedOpToolCall.
  ScoreRow.toolName → toolNames: string[].
- output-parser.ts: collects ALL element-tool tags in emit order;
  unknown-tool path also surfaces as single-element tool_calls so
  routing keeps the same wrong-tool semantics.
- score-run.ts: classifyRouting uses Array.includes for obvious
  prompts (right-tool when ANY emitted call matches expected_tool —
  over-production isn't a routing miss). Composite stays multi-tool
  on any non-empty list.
- aggregate.ts byTool: tallies EVERY name in toolNames, so a
  composite row that emits 6× add_activity_log_v0 + 1×
  add_section_header_v0 contributes 6+1 = 7 invocations across two
  tools (with row-level m1_legal applied to both buckets — apply is
  all-or-nothing).
- apply.ts: loops over parsed.calls and invokes
  handleElementToolCall in emit order. Any single call failing
  aborts the row (M1=false); we don't partial-apply.
- mock-llm.ts mockLlmParsed: collects all `<op_tool>` tags into the
  list (composite-prompt mocks can carry multi-call raw strings).
- apps/web design-parser.tryParseElementToolOutput: maps tool_calls
  → its single-shape DesignOutputShape contract using the FIRST
  call (the multi-tag path `tryParseAllElementToolOutputs` was
  already correct).

Tests: 3746 → 3750 vitest. New cases:
- output-parser: surfaces ALL element-tool tags in emit order with
  intermixed batch_design scaffolds dropped (3 element calls from
  5 tags).
- score-run: right-tool when expected appears alongside extras;
  composite multi-call captures every name in toolNames.
- aggregate: 6× activity_log + 1× section_header → byTool reports
  6 and 1 invocations respectively.

dry-run on ab-v3 produces a 208-row report; tsc + format clean.
2026-04-29 08:35:00 +08:00
Fini 1c3cbda495 feat(ab-corpus): ab-v3 yaml fixtures (7 obvious + 5 composite)
Adds 7 new obvious prompts covering the v0.8.0 element tools that
weren't in ab-v1 (tools 91-97):
  setting_row / member_row / filter_group / invite_row /
  activity_log / event_card / step_card

One yaml per tool, same single-component "Design ONLY..." pattern
as ab-v1, with must_contain_roles mirroring the role names emitted
by the corresponding builder in pen-core/src/element-builders/.

Adds 5 composite prompts that exercise the new M6 routing
breakdown:
  - dashboard-settings-page-composite (4× setting_row)
  - dashboard-team-people-page-composite (5× member_row + 1×
    invite_row)
  - dashboard-search-filters-composite (2× filter_group + result
    list)
  - dashboard-audit-feed-composite (6× activity_log)
  - mobile-onboarding-flow-composite (4× step_card)

Each composite prompt omits expected_tool_if_any (multi-tool
intent) and uses min_roles to enforce the multi-element shape.

corpus-loader: composite added to VALID_DIFFICULTIES; composite
prompts MUST NOT specify expected_tool_if_any (validation error
points the user back to difficulty=obvious if a single tool fits).
6 new corpus-loader tests cover the v3 yaml inventory + composite
validation rules. 3740 → 3746 vitest tests, all green.

ab-v3 corpus now has 52 prompts: 40 inherited from v1 (unchanged
for v1↔v3 comparability) + 7 obvious + 5 composite.
2026-04-29 08:15:00 +08:00
Fini 95566e4ed2 feat(ab-corpus): bootstrap ab-v3 with token cost + composite difficulty
ab-v3 succeeds ab-v1 (frozen 2026-04-28). Carries forward all 40
v1 obvious yaml files unchanged so the v1↔v3 overlap stays
comparable, then layers in two new dimensions.

**1. Token cost.** All clients (openai-compat, ark, bailian,
deepseek, minimax, codex-cli, stub-model) now return a
`ChatCallResult { content, usage }` instead of bare string.
Provider usage stats (`prompt_tokens` / `completion_tokens`) plumb
through realModelCall → run.ts → scoreRun → ScoreRow.{prompt,completion}Tokens.
aggregate adds avgPromptTokens{Baseline,Treatment} +
avgCompletionTokens{Baseline,Treatment} per ModelSummary.
write-report emits a new "Token cost" table with Δ columns so
narrow-tools-saves-tokens (the ab-v2 hypothesis) is measurable.
avgUsage skips rows with 0/0 usage so codex-cli (CLI doesn't
surface tokens) and harness errors don't deflate the average to
near-zero — they show '—' instead.

**2. Composite difficulty.** New 'composite' value alongside
obvious / optional. Composite prompts express multi-tool intents
where no single expected_tool_if_any applies. classifyRouting
routes composite-treatment runs into multi-tool / fallback /
garbage (3-bucket sum to 1, distinct from obvious's 4-bucket
right/wrong/fallback/garbage). aggregate adds m6_multi_tool +
m6_fallback + m6_garbage; write-report emits a "Composite routing"
table that gracefully degrades to a placeholder when no composite
yaml exists yet.

Harness side: scripts/ab-corpus/run.ts accepts --corpus ab-v3
(enum + parseArgs guard); dry-run on the v1-mirror corpus produces
a 160-row report including populated token table.

Tests: 4 new aggregate cases (composite, token avg with skip-zero,
NaN-when-no-data) + 4 new score-run cases (composite routing
multi-tool/fallback/garbage/baseline-n/a) + 2 new score-run cases
(usage plumbing) + 2 new openai-compat cases (usage parsing,
missing-usage fallback). Existing 5 retry tests updated for new
return shape. 3727 → 3740 vitest tests, all green; tsc + format
clean.

Token-cost docs and composite docs go straight into types.ts /
score-run.ts / aggregate.ts JSDoc — keeps the contract close to
the code that owns it.
2026-04-29 07:45:00 +08:00
Fini 9c605a3785 fix(ai): step-card marker clips overflow + tighten doc to short index
The previous schema description and JSDoc invited "Step 1" as a valid
value for `number`, but that prose has 6 chars and overflows the 36px
circle marker. Two changes:

- Builder: add clipContent: true on the marker frame so any caller
  who ignores the docs at least gets a clipped (not bleeding) render.
- Schema + JSDoc: drop the misleading "Step 1" example, document the
  1–3 character contract, and steer prose toward `title` instead.
2026-04-28 08:55:00 +08:00
Fini 1fee613111 feat(ai): ship 5 element tools to reach 97 (filter_group / invite_row / activity_log / event_card / step_card)
Closes the obvious gaps remaining in the family:
- add_filter_group_v0 — sidebar facet (heading + checkbox-style options
  with optional counts). Distinct from nav_chip_row (horizontal scrolling
  chips), tag (single applied chip), segmented_control (mutex tabs).
- add_invite_row_v0 — pending invite row (avatar + email/role + status
  pill + trailing action). Distinct from member_row (a JOINED member,
  no status pill or action) and list_row (no avatar / status / action).
- add_activity_log_v0 — single-line audit feed entry (optional tinted
  icon dot + actor in bold + action + right-aligned timestamp). Uses
  StyledTextSegment[] content for the bold/regular split. Distinct from
  timeline (multi-event vertical with connectors) and notification_row
  (title + body, no actor focus).
- add_event_card_v0 — single calendar event tile (date column with
  month band + day number, then title + time + location). Distinct from
  calendar_grid (the full month grid) and card_row (no date column).
- add_step_card_v0 — onboarding step card (numbered circle / check +
  title + description). Distinct from stepper (horizontal progress nav
  with connectors) and faq_item (collapsible Q&A header).

9 touchpoints per tool: pen-core builder + index + barrel + types,
pen-mcp handler + dispatcher + ext-4 schema, apps/web shim +
SERVER_BUILDERS, parity test (+5 cases), elements.md decision tree
items 86-89 + 6 PREFER mappings with cross-links to existing tools,
elements-cookbook.md arg-shape examples (8 entries across 5 tools).
2026-04-28 08:50:00 +08:00
Fini a5b594cf69 feat(ai): add_member_row_v0 — team / member list row (92nd tool)
Avatar + (name over optional subtitle) + optional trailing slot
(role badge / kebab menu / status dot). Distinct from
add_user_card_v0 (compact fit_content tile, no trailing slot) and
add_list_row_v0 (no avatar slot — leading icon instead).

9 touchpoints wired: pen-core builder + index + barrel + types,
pen-mcp handler + dispatcher + ext-4 schema, apps/web shim +
SERVER_BUILDERS, parity test, elements.md decision tree #84 +
PREFER mapping, cookbook arg shapes (3 variants).

Also disambiguates add_avatar_group_v0's PREFER mapping: drop
"团队成员" (now points at member_row), keep narrower phrases like
"成员头像" / "团队头像" / "presence indicator" that genuinely match
the stacked-avatars affordance, and add the cross-link to member_row.
2026-04-28 08:05:00 +08:00
Fini 77cc8ef3a7 fix(ai-skills): teach setting-row recipe with the right tool, drop invalid trailing_kind
Two bugs in elements.md after add_setting_row_v0 landed:
1. PREFER mapping still routed "settings row" to add_list_row_v0 — direct
   contradiction with the new add_setting_row_v0 entry below it.
2. The "Settings page" recipe called add_list_row_v0 with a trailing_kind:
   "switch" arg, but list-row has no such param; it would silently render
   without a switch (or fail validation in stricter clients).

Reroute settings-row prose mapping to the new tool and rewrite the recipe
to use add_setting_row_v0 with proper trailing variants.
2026-04-28 07:15:00 +08:00
Fini af1ddd1ad2 feat(ai): add_setting_row_v0 — settings menu row (91st tool)
Leading icon + (title over optional subtitle) + trailing control with
4 variants: chevron / value text / switch / badge. Distinct from
add_list_row_v0 (trailing is always icon, no switch/value/badge) and
add_form_field_v0 (label-above-input for forms).

Wires all 9 touchpoints: pen-core builder + index + barrel re-export,
pen-mcp handler + dispatcher case + ext-4 schema, apps/web shim +
Nitro SERVER_BUILDERS, elements.md decision tree #83 + PREFER mapping,
elements-cookbook arg-shape examples, plus shim-server parity case.
2026-04-28 06:30:00 +08:00
Fini e656b52f95 fix(mcp): describe section enum by parameter name, not invented schema path 2026-04-27 09:45:00 +08:00
Fini 0bed3a5d12 fix(mcp): drop hardcoded section list from get_design_prompt description
The tool-level description still listed 11 specific sections (schema /
layout / roles / text / style / icons / examples / guidelines /
planning / elements / design-md) even though the actual catalog has
28. Replace with a pointer at inputSchema.section.enum, which is
already derived from listPromptSections() — the description now
can't go stale as sections are added or removed.

D0 parity snapshot refreshed for the new description text.
2026-04-27 09:40:00 +08:00
Fini e8cbe712e2 fix(mcp): get_design_prompt enum derives from SECTION_MAP, no more drift
The published section enum had drifted to 14 entries while
SECTION_MAP grew to 28 (copywriting / overflow / cjk / variables +
8 codegen-* + elements-cookbook). External MCP clients calling with
the missing names hit schema-validation rejection even though the
implementation could serve them. Codex caught the immediate
elements-cookbook gap; widening the fix because the same pattern was
already silently broken for half the catalog.

- enum now derives from listPromptSections() at module load, no
  hand-maintained list to drift
- description points at listPromptSections() rather than enumerating
  individual sections (which was its own drift vector)
- design-prompt-elements adds a sync drift-guard: enum set must
  exactly equal listPromptSections() set
- D0 parity snapshot refreshed for the new enum values

3712 tests green.
2026-04-27 09:35:00 +08:00
Fini 3e3b02f227 test(mcp): refresh D0 parity snapshot for elements-cookbook section enum 2026-04-27 09:30:00 +08:00
Fini fa099ce282 fix(mcp): add elements-cookbook to get_design_prompt section enum 2026-04-27 09:25:00 +08:00
Fini f5eff2c9a3 fix(ai-skills): restore element-tool arg-shape examples in elements-cookbook.md
Trimming the Minimal usage block out of elements.md (65c31832) lost
arg-shape templates that the A/B harness depends on. Text-only LLMs
in the treatment arm see only the markdown skill content — no MCP
tools/list, no published inputSchema — so without the per-tool
example payloads they have to guess argument names and break M1.

Restore the full block as a sibling skill `elements-cookbook` (same
hasMcpTools flag, slightly later priority so it loads alongside
elements). Wire it through buildFullPrompt + the
get_design_prompt(section='elements-cookbook') section map. Update
the A/B harness to also strip the cookbook body when building the
baseline prompt — leaving it in B would leak tool names + arg shapes
back into the no-tools variant and re-bias the comparison.

Both files now under the 800-line per-file ceiling.
2026-04-27 09:20:00 +08:00
Fini 2939b6d4b5 refactor(ai-skills): trim elements.md minimal-usage block to fit 800-line ceiling
elements.md grew to 859 lines after the 81-90 batch shipped — past the
repo's per-file ceiling. The 366-line "Minimal usage" section was the
biggest contributor and the most redundant: MCP `tools/list` already
publishes the full inputSchema for every element tool (arg names,
types, descriptions, requireds), so the LLM has authoritative arg
shape from the wire. The decision tree + PREFER mappings already
teach WHEN to pick each tool. Drop the inline usage examples; keep
the composition pattern + cookbook recipes (which schemas can't
convey) plus invariants and failure-mode guidance.

493 lines remaining; design-prompt-elements + every drift guard still
green.
2026-04-27 09:15:00 +08:00
Fini ecf53d46c9 feat(ab-corpus): 10 ab-v1 prompts for the 81-90 element tool batch 2026-04-27 09:10:00 +08:00
Fini 5859c3f9af feat(ai): ship 10 element tools to reach 90 (user_card / drawer / combobox / toolbar / callout / share / inline_action / legend_item / inbox / profile_header)
Adds the desktop-leaning batch needed to round the family to 90:

- add_user_card_v0 — compact avatar+name+role row
- add_drawer_shell_v0 — full-height side panel header
- add_combobox_v0 — open-state autocomplete with dropdown
- add_toolbar_v0 — desktop icon button row + dividers
- add_callout_v0 — inline doc tip block, 5 tones
- add_share_row_v0 — circular social-share buttons
- add_inline_action_v0 — message + Undo-style action
- add_legend_item_v0 — chart legend marker+label+value
- add_inbox_message_v0 — email/inbox row with unread dot
- add_profile_header_v0 — large profile hero block

All ten go through the standard 9-touchpoint wiring and land in a
new ext-4 schema shard so existing shards stay under 800 lines.
Drift guards (contract / registry parity / shim-server parity) cover
each new name; per-tool handler tests are deferred — every tool's
structure is exercised through the parity build call already.
2026-04-27 09:05:00 +08:00
Fini de73c09202 fix(ai): tag close icon inherits tone fg color 2026-04-27 09:00:00 +08:00
Fini a2bf13bfce feat(ab-corpus): dashboard-applied-filter-tag ab-v1 prompt for add_tag_v0 2026-04-27 08:58:00 +08:00
Fini 072856e522 feat(ai): add_tag_v0 — single closable filter chip (80th tool) 2026-04-27 08:55:00 +08:00
Fini e14b9446c8 fix(ai): data-table cells clip overflow so long content stays in column
Long cell content rode `width=auto` and the cell itself had
`height=fit_content` with no clip, so a string longer than its
allotted column would push past the cell edge into the adjacent
column at render time. Switch each cell to `width/height=
fill_container + clipContent=true`, give the text child
`width=fill_container + textGrowth=fixed-width` so the layout
engine wraps inside the cell, and clip the row itself so a runaway
cell can't push siblings off-row either. Lock the contract in the
handler test.
2026-04-27 08:50:00 +08:00
Fini 68411c9e8b feat(ab-corpus): dashboard-customers-table ab-v1 prompt for add_data_table_row_v0 2026-04-27 08:45:00 +08:00
Fini d24b587b6c feat(ai): add_data_table_row_v0 — desktop tabular row (79th tool) 2026-04-27 08:40:00 +08:00
Fini 4442d8fe26 feat(ab-corpus): dashboard-team-presence ab-v1 prompt for add_avatar_group_v0 2026-04-27 08:35:00 +08:00
Fini e225104e43 feat(ai): add_avatar_group_v0 — stacked presence tile group (78th tool) 2026-04-27 08:30:00 +08:00
Fini b79ce5233d test(pen-core): static guard against CSS-side padding shorthands
Greps every builder source for paddingTop/Right/Bottom/Left as a
property key and fails if any reappear. The unified padding field is
the only form resolvePadding reads; the CSS-side siblings render with
zero inset, which slipped past type checks because ElementTree is
Record<string, unknown>. JSDoc / line-comment mentions stay allowed
so builders can still document the trap inline.
2026-04-27 08:25:00 +08:00
Fini c5655000fd chore(release): bump to v0.8.0 2026-04-27 08:20:00 +08:00
Fini 6e54054c73 chore(merge): integrate origin/v0.8.0 — main pre-release sync + CI fixes
origin's v0.8.0 had cherry-picks of the v0.7.5 deepseek/image-search
fixes (a727632a, a5952bc8) overlapping local 2073cf5b / 04f4fbc1, plus
new commits (model-selector ark-coding deepseek-v4-pro/flash IDs that
ARK rejects, fetch error.cause unwrap, CI agent-native build, op
export docs cleanup, main merge). Resolved the ark-coding list in
favor of HEAD's deepseek-v3.2 entry (only model ARK Coding Plan
actually supports — see openpencil-docs note).
2026-04-27 08:15:00 +08:00
Kayshen-X b554b4f1a6 Merge branch 'main' of github.com:ZSeven-W/openpencil into v0.8.0 2026-04-26 19:39:14 +08:00
Kayshen-X b80db5b798 docs: drop op export from CLI docs and clarify pen-mcp usage
The `op export` command was removed in 0.7.x but the README still
advertised it (#116). The pen-mcp README also documented an
`npx @zseven-w/pen-mcp` quick-start that never worked because the
package ships TypeScript source against workspace-only deps with no
`bin` entry (#117).

- Strip `op export` references from all 15 root and 15 cli READMEs
- Sync AGENTS.md, CLAUDE.md, apps/cli/CLAUDE.md to match the codegen-
  pipeline reality (no standalone export command anymore)
- Rewrite pen-mcp README's quick-start: explain the package ships as
  part of the OpenPencil app and external clients connect over HTTP

Closes #116
Closes #117
2026-04-26 19:20:14 +08:00
Fini abbccc1ba2 fix(ai): refresh DeepSeek defaults to v4 model series
`/models` now returns only deepseek-v4-pro and deepseek-v4-flash;
deepseek-chat / deepseek-reasoner sunset 2026-07-24 and the
deepseek-v3.2 hard-coded in the ark-coding fallback list never
existed. Both v4 models default to thinking enabled and the API
toggles via `{"thinking":{"type":"disabled"}}` — keep
`thinkingMode: 'disabled'` so the app's fast/non-thinking default
stays intact (server reasoning paths honor it; the Zig openai-compat
path doesn't emit the toggle yet, so calls through that path still
get provider-default thinking until it's wired). v4-pro promoted to
full tier; legacy aliases pinned to an exact RegExp so future
deepseek-* variants don't inherit a forced disabled mode.

Bandaid for the unwired toggle: v4-pro gets `timeoutMultiplier: 2`
because its default-on reasoning blows past the orchestrator's
planning timeout on long system prompts (observed in dev: planning
phase falls back, sub-agent then succeeds — UX degraded but
functional). Drop the multiplier once the Zig path actually sends
`thinking:{type:disabled}`.

Don't add a BUILTIN_MODEL_LISTS.deepseek entry — DeepSeek exposes
/v1/models, so let `fetchProviderModels` pull the live catalog
through `/api/ai/provider-models` instead of pinning a snapshot
(the ark-coding `deepseek-v3.2` ghost above shows what those
snapshots drift into).
2026-04-26 06:30:00 +08:00
Fini 2f62cb262f fix(core): replace CSS padding shorthands with canonical padding array in 14 element builders
paddingTop/Right/Bottom/Left are silently dropped by resolvePadding in
the layout engine, causing all four sides to render as 0. Switch every
affected builder to the canonical padding: [T,R,B,L] form.
2026-04-26 06:29:59 +08:00
Fini ae40046b17 fix(ai): sidebar_nav uses unified padding so insets actually render
pen-core's resolvePadding only reads the array/scalar `padding`
field; the CSS-style paddingTop/Right/Bottom/Left siblings I shipped
in 5c275ba1 were silently dropped, so the rail rendered flush. Switch
to padding=[16,12] / [8,12,24,12] / [0,12] and add a regression
assertion that catches the same trap on any future tweak.
2026-04-26 06:29:58 +08:00
Fini 0cb0e81a78 feat(ab-corpus): dashboard-team-sidebar ab-v1 prompt for add_sidebar_nav_v0
Adds the obvious-difficulty prompt covering the 77th element tool;
bumps the corpus-loader test from 26 to 27 entries and lists the new
2026-04-27 batch in the README so future corpus drift is caught at
test time.
2026-04-26 06:29:57 +08:00
Fini e9d9e37770 feat(ai): add_sidebar_nav_v0 — desktop persistent left rail (77th tool)
Top_nav_bar / bottom_nav / nav_chip_row covered mobile and inline
chrome but desktop dashboards still had to hand-roll their left
sidebar via batch_design. Adds a 240px-wide vertical rail with
icon+label rows, optional brand title, and slate-100 pill bg on the
active item — distinct surface from the existing nav tools so the
decision tree picks it cleanly for "sidebar / side nav / 侧边栏".
2026-04-26 06:29:56 +08:00
Fini 6c44d58a31 feat(ai): add_cookie_banner_v0 — GDPR/CCPA cookie banner (76th tool)
Sticky bottom-of-page disclosure card with title, body, accept /
decline buttons (decline-then-accept order), and an optional
"Cookie settings" link for fine-grained consent. Caller positions
the banner; the tool emits the card itself with shadow + 1px slate
border for visual lift above the page content. Plus the v1 corpus
prompt landing-cookie-banner (26 v1 prompts total).
2026-04-25 08:15:00 +08:00
Fini 1acd31a2e1 feat(ai): add_input_with_action_v0 — input + inline action button (75th tool)
The "Subscribe to newsletter" / "Apply discount code" / "Send chat
message" pattern. Two action variants:
  - text (default): pill button with label like "Subscribe"
  - icon: 44×44 square icon button (chat send arrow / search apply)

Distinct from add_form_field_v0 (label-above, no inline button)
and add_search_bar_v0 (no trailing action button). Optional
leading_icon adds an icon inside the input itself. Plus the v1
corpus prompt landing-newsletter-signup (25 v1 prompts total).
2026-04-25 08:00:00 +08:00
Fini 3527f6ff6a feat(ab-corpus): mobile-phone-input ab-v1 prompt for add_phone_input_v0
Adds 24th v1 prompt and rolls the 2026-04-25 batch from 6 to 7
tools (5 v0 + 2 v1). Loader test count + tool-set assertions
updated.
2026-04-25 07:45:00 +08:00
Fini 577488546a feat(ai): add_phone_input_v0 — international phone input (74th tool)
The "+1 (555) …" pattern from every modern signup / login screen.
A 44px row with leading country selector (flag + dial code +
chevron-down), a 1px slate divider, and the digits input on the
right. Country selector is a button-shape (no actual dropdown
menu); caller handles picker UX as a separate concern. `value`
toggles between placeholder (slate-400) and populated (slate-900).
2026-04-25 07:30:00 +08:00
Fini 008bf0ca46 test(ab-corpus): update v1 corpus expectations to 23 prompts
Three batches now: 2026-04-22 (12) + 2026-04-24 (5) + 2026-04-25 (6
— 4 v0 + 2 v1). The 2026-04-25 batch corresponds to the tools
shipped in this session: social_login / pricing_card / stat_card /
range_slider / toast_v1 / empty_chart_v1.
2026-04-25 01:47:23 +08:00
Fini dd5c85dcc0 feat(ab-corpus): ab-v1 prompts for 6 new tools shipped this session
Adds NL prompts exercising the new narrow tools + theme-aware v1s:
- mobile-social-login (→ add_social_login_row_v0)
- dashboard-pricing-card (→ add_pricing_card_v0)
- dashboard-revenue-stat (→ add_stat_card_v0)
- mobile-volume-slider (→ add_range_slider_v0)
- mobile-dark-toast (→ add_toast_v1, inverted contrast)
- dashboard-dark-empty-chart (→ add_empty_chart_v1)

Each prompt describes the visual intent in natural language, asserts
the expected roles, and pins the expected_tool_if_any so the harness
can attribute tool-choice correctness.
2026-04-25 01:39:40 +08:00
Fini d1022e74d5 feat(ai): add_empty_chart_v1 — theme-aware no-data placeholder
Third theme-aware v1 (after add_modal_shell_v1 and add_toast_v1).
Same dashed-border "no data yet" tile shape as v0 with a `theme`
param that swaps 5 colors (bg / border / icon / title / subtitle):
  - light (default): byte-parity with v0 (slate-50 / slate-300)
  - dark: slate-800 bg + slate-600 border + slate-200 title
  - system: $color-surface-2 / $color-border / $color-text-primary
    refs — requires applySemanticPalette(doc) seeded

Lets dark-theme dashboards keep a matching empty-slot surface
instead of punching a light rectangle out of dark cards.
2026-04-25 01:37:40 +08:00