Two diet/teaching changes to elements.md, both motivated by today's
ab-v4 smoke results:
(a) parent_id teaching — minimax-m2.7 invented "members-section" /
"canvas" as parent_id values on its multi-tool composite output,
then every one of its 14 tags failed apply with "parent_id X not
found in document". Same pattern as yesterday. Adds an explicit
rule near the top of elements.md (right under the multi-tool
banner) — `parent_id` is REAL or OMITTED, never invented. The
`<page>` / `<panel>` / `<sidebar>` placeholders in the cookbook
recipes are documentation conventions; in actual output, OMIT
the field. Names the failure mode by reproducing the error
message format so models learn to avoid it.
(b) Phase 3 token diet — strip ". Different from <tool> (...)"
disambiguation suffixes from the PREFER list (22 entries had
them, ~80-200 chars each = ~2.5kb / ~625 tokens saved). The
primary keyword + tool-name + capability description survives
intact; the cross-references pointing at sibling tools get
dropped. Risk: slight increase in wrong-tool routing on
ambiguous prompts. Worth it for the size reduction; the
decision tree alone still shows the tool family in context.
Verified by mechanical diff — perl in-place edit, then visual
review confirms no other content was touched. 3785 vitest pass,
format clean, tsc silent.
Net effect on T-prompt size (chars / 4 estimate):
T + composite + mobile: was 17,930 → now 17,475 (-455 chars)
T + composite + dashboard: was 18,437 → now 17,983 (-454 chars)
T + obvious + dashboard: ~14,009 → ~13,556 (-453 chars)
Per-arm savings are smaller than the raw 2.5kb trim because (a)
adds ~700 chars of parent_id teaching. Net win: ~450 chars / ~110
tokens per call. Modest but compounds across 520+ ab-v4 runs.
Sweep follow-up to 113bd55a — same defensive pattern (reject
unknown enum strings at the entry boundary) applied to every other
builder that indexed a Record<EnumLiteral, T> with a value sourced
from raw JSON args.
Builders + enums covered:
- buildTag — TagTone (default | accent | success | warning | error)
- buildCallout — CalloutTone (info | success | warning | danger | note)
- buildActivityLog — tone (info | success | warning | danger | neutral)
- buildInviteRow — InviteStatus (pending | expired | accepted)
- buildMemberRow — trailing.tone for status_dot (online | busy | away
| offline). role_badge / menu variants skip the check (no tone field)
Same failure mode each one fixed: when a model invents an
out-of-enum string (gpt-5.4 did this with `level: "caption"` in
ab-v4), the lookup `TONES[bad]` / `STATUS_TONE[bad]` returned
undefined, the next property access crashed mid-batch with a
cryptic `undefined is not an object`, and the surrounding dispatch
loop dropped every remaining tag (until df33e937 + 07639f6d landed
the per-shape continuation + partial-doc scoring earlier today).
With validation in place, a bad enum becomes a clean per-shape
error message + the rest of the batch still applies.
13 new edge-case tests cover throw on bad input + valid path on
every enum value + omitted-default for each builder. 3785 vitest
pass, format clean, tsc silent.
Builders not touched: heading.ts (already done in 113bd55a).
Builders that don't fit this pattern (no enum→Record lookup of a
user-controlled string): everything else surveyed via grep on
`Record<.*Tone|Status|Level|Mode|Kind`.
Codex stop-time review caught the previous fix (df33e937) handing
the scorer a partial PenDocument that the scorer immediately
ignored. score-run.ts:87 short-circuited on `!applied.ok || !applied.doc`,
so even though apply.ts now surfaces 12 of 13 successfully-applied
tags as a populated `applied.doc`, the row still scored as a total
failure (M1=false, M3=false, m3_failure_reason="apply failed before
shape checks") — exactly the noise df33e937 was meant to eliminate.
Loosens the short-circuit to `!applied.doc` only. When apply.ok=false
but apply.doc is populated, the scorer now:
- runs the issue detector against the partial doc (issues surface)
- keeps M1 strict (apply.ok=false → M1=false regardless of detector)
- decouples M3 from M1: M3 = shape.ok against the partial doc
- sets m3_failure_reason to the shape miss when shape fails;
otherwise to "partial apply (M3 met by what landed): <error>" so
the row reads "tag 12 of 13 broke, but role coverage still met"
instead of silently swallowing the partial signal
- surfaces applied.error in row.applyError so per-shape failure
messages flow into reports
Two new tests cover the new path:
1. partial apply + shape match → M1=false, M3=true, reason mentions
"partial apply"
2. partial apply + shape miss → M1=false, M3=false, shape-miss
reason wins (structural verdict trumps the partial-apply notice)
Plumbing chain across today's session is now consistent:
- apply.ts continues past per-shape failures (df33e937)
- score-run.ts scores the partial doc that lands (this commit)
- scoring no longer over-attributes to "apply failed" when the model
actually produced most of the brief
3774 vitest pass (+2), format clean, tsc silent.
ab-v4 partial sweep (2026-05-01) caught gpt-5.4 emitting
`add_heading_v0({"content":"Pending invitations","level":"caption"})`
as the 12th tag of a 13-tag composite multi-tool response. Even
though the MCP tool def has `enum: ['display','h1','h2','h3']`, the
ab-corpus harness and the in-process production dispatcher both call
buildHeading() with raw JSON args (no jsonschema gate), so the model's
invented "caption" reached the preset lookup. LATIN_PRESETS["caption"]
is undefined, and the next line `fontSize: preset.fontSize` crashed
the WHOLE batch with `undefined is not an object (evaluating
'preset.fontSize')` — the 11 valid tags ahead of it never landed.
Adds an entry-point validation in buildHeading: if `level` is set and
not in the {display, h1, h2, h3} set, throw with a clear message.
The dispatch loops in apps/web/element-tools-dispatcher and
scripts/ab-corpus/apply both catch per-shape and keep running the
remaining tags, so a single bad level on tag 12 no longer kills tags
1-11 + 13.
3 new edge-case tests in element-builders-edge-cases.test.ts
cover the throw + the four valid levels + the omitted-default case.
3772 vitest pass (+3), format clean, tsc silent.
Other element builders likely have the same pattern (preset lookup
on a string enum without runtime validation) — separate sweep, not
shotgunning here.
ab-v3 / ab-v4 showed Phase 1A (cookbook strip on obvious difficulty)
shaved ~4.4k tokens off T-obvious. Phase 2 adds a per-category gate
that strips cookbook recipes whose domain doesn't match the prompt's
category — mobile briefs don't see dashboard recipes, dashboard
briefs don't see mobile / landing recipes, etc.
Mechanism: HTML comment block markers in elements.md
(`<!-- @domain:dashboard --> ... <!-- /@domain -->`) plus a
stripNonMatchingDomains() pass in buildSystemPrompt that drops blocks
whose tag list doesn't include the active category. Untagged content
is "general" and stays in every variant — the safe default.
Tagged 7 single-domain cookbook recipes:
- dashboard: Team / members list, Audit / activity feed, Faceted
search filter sidebar, Dashboard KPI strip
- landing: Pricing section
- mobile: Onboarding "How it works", Support chat thread
Cross-domain recipes (Login, Signup, Settings page, OTP, Empty
inbox) stay untagged so they load for every category. Decision tree
+ PREFER list also untagged today; the per-tool annotations there
would be a much larger judgment pass for marginal additional savings.
Token measurements (chars / 4 estimate):
full mobile dashboard landing
- T + composite 19.0k 17.9k 18.4k 17.7k
(-1.1k) (-0.6k) (-1.3k)
- T + obvious 14.6k 13.5k 14.0k 13.3k
(-1.1k) (-0.6k) (-1.3k)
Modest absolute savings — Phase 2 only filters cookbook RECIPES (in
elements.md), and most cookbook content is in elements-cookbook.md
which Phase 1A already strips on obvious. To hit the 6-8k T target
we still need decision-tree compression or PREFER-list trim, but
both are lossier than this gate. Phase 3 candidates noted in the
ab-v4 results doc.
real-model.ts plumbs call.prompt.category through to buildSystemPrompt.
3766 vitest pass (+6 category filter tests including a 500-char
floor regression guard that the filter actually shaves bytes).
Codex stop-time review (3rd round) caught residual "T prompt
includes mixed-strategy instructions": even after the prompt
forbade mixing, elements.md still taught the mixed pattern in
several places that are also part of the T system prompt.
Cleaned out every spot that paired batch_design with add_*_v0:
- Login screen / Pricing section / Dashboard KPI strip recipes:
dropped the leading `batch_design: foo = I("page", {...})` line
and renamed `<foo>` / `<row>` placeholders to `<page>` so each
recipe is now Strategy A (element tools only). Lost: explicit
page-level layout/padding/horizontal row — acceptable, the
recipes still teach the tool selection + chain pattern.
- Intro paragraph: dropped "override via a follow-up batch_design
U-op if needed" (taught a per-component fallback that the parser
drops).
- Banner: rewrote the fall-back clause from "when no element tool
fits a specific component shape" to "when at least one component
truly needs a custom shape no element tool covers — and then use
a SINGLE batch_design for the WHOLE response, never mixed."
- "STILL use batch_design when" list: collapsed 3 mixed-strategy
bullets ("larger composite via batch_design then element tool",
"post-hoc styling via batch_design U-ops") into 3 clean
Strategy-B-only bullets, all explicitly emit a SINGLE batch_design.
- Removed the "## Composition pattern" section entirely — its 3-step
plan was a textbook mixed pattern (batch_design root → element
tool inserts → batch_design U-op styling) that depended on real
MCP multi-round semantics the corpus harness can't provide.
- Composition rules of thumb: replaced "Don't mix N-tool and
batch_design DSL ops in a single call" + "Style overrides come
AFTER structure" with a single "Don't mix in the same output"
rule that names the corpus parser behavior explicitly.
3760 vitest pass, format clean, tsc silent. T-obvious 14.5k tokens
(down ~100 from the cleanup); composite still 18.9k.
Codex stop-time review flagged the new ab-v3 composite cookbook
recipes calling add_section_header_v0 with a `subtitle` arg the tool
doesn't accept (silently dropped at runtime today, but teaches live
models to emit invalid shapes). The same bug was in the dry-run stub
fixtures.
Split each header into add_section_header_v0(title) +
add_body_text_v0(content) — semantically what the ab-v3 briefs ask
for, and reinforces the multi-tool chaining the cookbook now teaches.
Also fills in the missing required `number` arg on the onboarding
recipe's final completed step card (schema requires it even when
completed=true renders a check instead of the number).
ab-v3 live sweep showed 0/25 composite-T runs walked the multi-tool
path — every model fell back to batch_design or garbage when given a
brief like "5 member rows + 1 invite row" or "6 audit log entries".
Root cause: elements.md's decision tree opens with "pick first match"
(single-tool framing) and the cookbook's chained examples sit ~400
lines deep, so models stop at the lookup and never realize they can
emit N <op_tool> blocks.
This adds:
- top-of-file "MULTI-TOOL OUTPUT IS THE NORM" banner with a concrete
5-call settings example, before the decision tree
- decision tree heading rewritten to "(per component — pick first
match)" so per-component framing is in scope from the start
- 4 new cookbook recipes mirroring the ab-v3 composite prompts:
team / members list (rows + invite), audit / activity feed,
faceted search filter sidebar, onboarding step cards
Markdown-only edit to a single skill file; vite-plugin-skills
re-compiles the registry. Banner verified present in the generated
registry, all 3750 vitest tests pass. Verifies in next ab-v4 sweep.
Codex stop-hook caught: ab-v3 introduced composite-difficulty prompts
that *expect* multi-tool emit (e.g. 5× member_row + 1× invite_row
for a team page), but `ParsedOutput.tool_call` was a single
{name, arguments} so the parser silently dropped every call after
the first. apply.ts only invoked one tool, M3 min_roles couldn't
pass on legitimately-routed multi-tool runs, and byTool stats
under-counted. The composite routing 'multi-tool' bucket was
correctly assigned in classifyRouting, but downstream the pipeline
behaved as if the model emitted a single call.
This commit replaces `kind: 'tool_call'` with
`kind: 'tool_calls'` (NON-EMPTY list) across every consumer:
- types.ts: ParsedOutput tagged union; new ParsedOpToolCall.
ScoreRow.toolName → toolNames: string[].
- output-parser.ts: collects ALL element-tool tags in emit order;
unknown-tool path also surfaces as single-element tool_calls so
routing keeps the same wrong-tool semantics.
- score-run.ts: classifyRouting uses Array.includes for obvious
prompts (right-tool when ANY emitted call matches expected_tool —
over-production isn't a routing miss). Composite stays multi-tool
on any non-empty list.
- aggregate.ts byTool: tallies EVERY name in toolNames, so a
composite row that emits 6× add_activity_log_v0 + 1×
add_section_header_v0 contributes 6+1 = 7 invocations across two
tools (with row-level m1_legal applied to both buckets — apply is
all-or-nothing).
- apply.ts: loops over parsed.calls and invokes
handleElementToolCall in emit order. Any single call failing
aborts the row (M1=false); we don't partial-apply.
- mock-llm.ts mockLlmParsed: collects all `<op_tool>` tags into the
list (composite-prompt mocks can carry multi-call raw strings).
- apps/web design-parser.tryParseElementToolOutput: maps tool_calls
→ its single-shape DesignOutputShape contract using the FIRST
call (the multi-tag path `tryParseAllElementToolOutputs` was
already correct).
Tests: 3746 → 3750 vitest. New cases:
- output-parser: surfaces ALL element-tool tags in emit order with
intermixed batch_design scaffolds dropped (3 element calls from
5 tags).
- score-run: right-tool when expected appears alongside extras;
composite multi-call captures every name in toolNames.
- aggregate: 6× activity_log + 1× section_header → byTool reports
6 and 1 invocations respectively.
dry-run on ab-v3 produces a 208-row report; tsc + format clean.
Adds 7 new obvious prompts covering the v0.8.0 element tools that
weren't in ab-v1 (tools 91-97):
setting_row / member_row / filter_group / invite_row /
activity_log / event_card / step_card
One yaml per tool, same single-component "Design ONLY..." pattern
as ab-v1, with must_contain_roles mirroring the role names emitted
by the corresponding builder in pen-core/src/element-builders/.
Adds 5 composite prompts that exercise the new M6 routing
breakdown:
- dashboard-settings-page-composite (4× setting_row)
- dashboard-team-people-page-composite (5× member_row + 1×
invite_row)
- dashboard-search-filters-composite (2× filter_group + result
list)
- dashboard-audit-feed-composite (6× activity_log)
- mobile-onboarding-flow-composite (4× step_card)
Each composite prompt omits expected_tool_if_any (multi-tool
intent) and uses min_roles to enforce the multi-element shape.
corpus-loader: composite added to VALID_DIFFICULTIES; composite
prompts MUST NOT specify expected_tool_if_any (validation error
points the user back to difficulty=obvious if a single tool fits).
6 new corpus-loader tests cover the v3 yaml inventory + composite
validation rules. 3740 → 3746 vitest tests, all green.
ab-v3 corpus now has 52 prompts: 40 inherited from v1 (unchanged
for v1↔v3 comparability) + 7 obvious + 5 composite.
ab-v3 succeeds ab-v1 (frozen 2026-04-28). Carries forward all 40
v1 obvious yaml files unchanged so the v1↔v3 overlap stays
comparable, then layers in two new dimensions.
**1. Token cost.** All clients (openai-compat, ark, bailian,
deepseek, minimax, codex-cli, stub-model) now return a
`ChatCallResult { content, usage }` instead of bare string.
Provider usage stats (`prompt_tokens` / `completion_tokens`) plumb
through realModelCall → run.ts → scoreRun → ScoreRow.{prompt,completion}Tokens.
aggregate adds avgPromptTokens{Baseline,Treatment} +
avgCompletionTokens{Baseline,Treatment} per ModelSummary.
write-report emits a new "Token cost" table with Δ columns so
narrow-tools-saves-tokens (the ab-v2 hypothesis) is measurable.
avgUsage skips rows with 0/0 usage so codex-cli (CLI doesn't
surface tokens) and harness errors don't deflate the average to
near-zero — they show '—' instead.
**2. Composite difficulty.** New 'composite' value alongside
obvious / optional. Composite prompts express multi-tool intents
where no single expected_tool_if_any applies. classifyRouting
routes composite-treatment runs into multi-tool / fallback /
garbage (3-bucket sum to 1, distinct from obvious's 4-bucket
right/wrong/fallback/garbage). aggregate adds m6_multi_tool +
m6_fallback + m6_garbage; write-report emits a "Composite routing"
table that gracefully degrades to a placeholder when no composite
yaml exists yet.
Harness side: scripts/ab-corpus/run.ts accepts --corpus ab-v3
(enum + parseArgs guard); dry-run on the v1-mirror corpus produces
a 160-row report including populated token table.
Tests: 4 new aggregate cases (composite, token avg with skip-zero,
NaN-when-no-data) + 4 new score-run cases (composite routing
multi-tool/fallback/garbage/baseline-n/a) + 2 new score-run cases
(usage plumbing) + 2 new openai-compat cases (usage parsing,
missing-usage fallback). Existing 5 retry tests updated for new
return shape. 3727 → 3740 vitest tests, all green; tsc + format
clean.
Token-cost docs and composite docs go straight into types.ts /
score-run.ts / aggregate.ts JSDoc — keeps the contract close to
the code that owns it.
The previous schema description and JSDoc invited "Step 1" as a valid
value for `number`, but that prose has 6 chars and overflows the 36px
circle marker. Two changes:
- Builder: add clipContent: true on the marker frame so any caller
who ignores the docs at least gets a clipped (not bleeding) render.
- Schema + JSDoc: drop the misleading "Step 1" example, document the
1–3 character contract, and steer prose toward `title` instead.
Closes the obvious gaps remaining in the family:
- add_filter_group_v0 — sidebar facet (heading + checkbox-style options
with optional counts). Distinct from nav_chip_row (horizontal scrolling
chips), tag (single applied chip), segmented_control (mutex tabs).
- add_invite_row_v0 — pending invite row (avatar + email/role + status
pill + trailing action). Distinct from member_row (a JOINED member,
no status pill or action) and list_row (no avatar / status / action).
- add_activity_log_v0 — single-line audit feed entry (optional tinted
icon dot + actor in bold + action + right-aligned timestamp). Uses
StyledTextSegment[] content for the bold/regular split. Distinct from
timeline (multi-event vertical with connectors) and notification_row
(title + body, no actor focus).
- add_event_card_v0 — single calendar event tile (date column with
month band + day number, then title + time + location). Distinct from
calendar_grid (the full month grid) and card_row (no date column).
- add_step_card_v0 — onboarding step card (numbered circle / check +
title + description). Distinct from stepper (horizontal progress nav
with connectors) and faq_item (collapsible Q&A header).
9 touchpoints per tool: pen-core builder + index + barrel + types,
pen-mcp handler + dispatcher + ext-4 schema, apps/web shim +
SERVER_BUILDERS, parity test (+5 cases), elements.md decision tree
items 86-89 + 6 PREFER mappings with cross-links to existing tools,
elements-cookbook.md arg-shape examples (8 entries across 5 tools).
Two bugs in elements.md after add_setting_row_v0 landed:
1. PREFER mapping still routed "settings row" to add_list_row_v0 — direct
contradiction with the new add_setting_row_v0 entry below it.
2. The "Settings page" recipe called add_list_row_v0 with a trailing_kind:
"switch" arg, but list-row has no such param; it would silently render
without a switch (or fail validation in stricter clients).
Reroute settings-row prose mapping to the new tool and rewrite the recipe
to use add_setting_row_v0 with proper trailing variants.
Leading icon + (title over optional subtitle) + trailing control with
4 variants: chevron / value text / switch / badge. Distinct from
add_list_row_v0 (trailing is always icon, no switch/value/badge) and
add_form_field_v0 (label-above-input for forms).
Wires all 9 touchpoints: pen-core builder + index + barrel re-export,
pen-mcp handler + dispatcher case + ext-4 schema, apps/web shim +
Nitro SERVER_BUILDERS, elements.md decision tree #83 + PREFER mapping,
elements-cookbook arg-shape examples, plus shim-server parity case.
The tool-level description still listed 11 specific sections (schema /
layout / roles / text / style / icons / examples / guidelines /
planning / elements / design-md) even though the actual catalog has
28. Replace with a pointer at inputSchema.section.enum, which is
already derived from listPromptSections() — the description now
can't go stale as sections are added or removed.
D0 parity snapshot refreshed for the new description text.
The published section enum had drifted to 14 entries while
SECTION_MAP grew to 28 (copywriting / overflow / cjk / variables +
8 codegen-* + elements-cookbook). External MCP clients calling with
the missing names hit schema-validation rejection even though the
implementation could serve them. Codex caught the immediate
elements-cookbook gap; widening the fix because the same pattern was
already silently broken for half the catalog.
- enum now derives from listPromptSections() at module load, no
hand-maintained list to drift
- description points at listPromptSections() rather than enumerating
individual sections (which was its own drift vector)
- design-prompt-elements adds a sync drift-guard: enum set must
exactly equal listPromptSections() set
- D0 parity snapshot refreshed for the new enum values
3712 tests green.
Trimming the Minimal usage block out of elements.md (65c31832) lost
arg-shape templates that the A/B harness depends on. Text-only LLMs
in the treatment arm see only the markdown skill content — no MCP
tools/list, no published inputSchema — so without the per-tool
example payloads they have to guess argument names and break M1.
Restore the full block as a sibling skill `elements-cookbook` (same
hasMcpTools flag, slightly later priority so it loads alongside
elements). Wire it through buildFullPrompt + the
get_design_prompt(section='elements-cookbook') section map. Update
the A/B harness to also strip the cookbook body when building the
baseline prompt — leaving it in B would leak tool names + arg shapes
back into the no-tools variant and re-bias the comparison.
Both files now under the 800-line per-file ceiling.
elements.md grew to 859 lines after the 81-90 batch shipped — past the
repo's per-file ceiling. The 366-line "Minimal usage" section was the
biggest contributor and the most redundant: MCP `tools/list` already
publishes the full inputSchema for every element tool (arg names,
types, descriptions, requireds), so the LLM has authoritative arg
shape from the wire. The decision tree + PREFER mappings already
teach WHEN to pick each tool. Drop the inline usage examples; keep
the composition pattern + cookbook recipes (which schemas can't
convey) plus invariants and failure-mode guidance.
493 lines remaining; design-prompt-elements + every drift guard still
green.
Adds the desktop-leaning batch needed to round the family to 90:
- add_user_card_v0 — compact avatar+name+role row
- add_drawer_shell_v0 — full-height side panel header
- add_combobox_v0 — open-state autocomplete with dropdown
- add_toolbar_v0 — desktop icon button row + dividers
- add_callout_v0 — inline doc tip block, 5 tones
- add_share_row_v0 — circular social-share buttons
- add_inline_action_v0 — message + Undo-style action
- add_legend_item_v0 — chart legend marker+label+value
- add_inbox_message_v0 — email/inbox row with unread dot
- add_profile_header_v0 — large profile hero block
All ten go through the standard 9-touchpoint wiring and land in a
new ext-4 schema shard so existing shards stay under 800 lines.
Drift guards (contract / registry parity / shim-server parity) cover
each new name; per-tool handler tests are deferred — every tool's
structure is exercised through the parity build call already.
Long cell content rode `width=auto` and the cell itself had
`height=fit_content` with no clip, so a string longer than its
allotted column would push past the cell edge into the adjacent
column at render time. Switch each cell to `width/height=
fill_container + clipContent=true`, give the text child
`width=fill_container + textGrowth=fixed-width` so the layout
engine wraps inside the cell, and clip the row itself so a runaway
cell can't push siblings off-row either. Lock the contract in the
handler test.
Greps every builder source for paddingTop/Right/Bottom/Left as a
property key and fails if any reappear. The unified padding field is
the only form resolvePadding reads; the CSS-side siblings render with
zero inset, which slipped past type checks because ElementTree is
Record<string, unknown>. JSDoc / line-comment mentions stay allowed
so builders can still document the trap inline.
origin's v0.8.0 had cherry-picks of the v0.7.5 deepseek/image-search
fixes (a727632a, a5952bc8) overlapping local 2073cf5b / 04f4fbc1, plus
new commits (model-selector ark-coding deepseek-v4-pro/flash IDs that
ARK rejects, fetch error.cause unwrap, CI agent-native build, op
export docs cleanup, main merge). Resolved the ark-coding list in
favor of HEAD's deepseek-v3.2 entry (only model ARK Coding Plan
actually supports — see openpencil-docs note).
The `op export` command was removed in 0.7.x but the README still
advertised it (#116). The pen-mcp README also documented an
`npx @zseven-w/pen-mcp` quick-start that never worked because the
package ships TypeScript source against workspace-only deps with no
`bin` entry (#117).
- Strip `op export` references from all 15 root and 15 cli READMEs
- Sync AGENTS.md, CLAUDE.md, apps/cli/CLAUDE.md to match the codegen-
pipeline reality (no standalone export command anymore)
- Rewrite pen-mcp README's quick-start: explain the package ships as
part of the OpenPencil app and external clients connect over HTTP
Closes#116Closes#117
`/models` now returns only deepseek-v4-pro and deepseek-v4-flash;
deepseek-chat / deepseek-reasoner sunset 2026-07-24 and the
deepseek-v3.2 hard-coded in the ark-coding fallback list never
existed. Both v4 models default to thinking enabled and the API
toggles via `{"thinking":{"type":"disabled"}}` — keep
`thinkingMode: 'disabled'` so the app's fast/non-thinking default
stays intact (server reasoning paths honor it; the Zig openai-compat
path doesn't emit the toggle yet, so calls through that path still
get provider-default thinking until it's wired). v4-pro promoted to
full tier; legacy aliases pinned to an exact RegExp so future
deepseek-* variants don't inherit a forced disabled mode.
Bandaid for the unwired toggle: v4-pro gets `timeoutMultiplier: 2`
because its default-on reasoning blows past the orchestrator's
planning timeout on long system prompts (observed in dev: planning
phase falls back, sub-agent then succeeds — UX degraded but
functional). Drop the multiplier once the Zig path actually sends
`thinking:{type:disabled}`.
Don't add a BUILTIN_MODEL_LISTS.deepseek entry — DeepSeek exposes
/v1/models, so let `fetchProviderModels` pull the live catalog
through `/api/ai/provider-models` instead of pinning a snapshot
(the ark-coding `deepseek-v3.2` ghost above shows what those
snapshots drift into).
paddingTop/Right/Bottom/Left are silently dropped by resolvePadding in
the layout engine, causing all four sides to render as 0. Switch every
affected builder to the canonical padding: [T,R,B,L] form.
pen-core's resolvePadding only reads the array/scalar `padding`
field; the CSS-style paddingTop/Right/Bottom/Left siblings I shipped
in 5c275ba1 were silently dropped, so the rail rendered flush. Switch
to padding=[16,12] / [8,12,24,12] / [0,12] and add a regression
assertion that catches the same trap on any future tweak.
Adds the obvious-difficulty prompt covering the 77th element tool;
bumps the corpus-loader test from 26 to 27 entries and lists the new
2026-04-27 batch in the README so future corpus drift is caught at
test time.
Top_nav_bar / bottom_nav / nav_chip_row covered mobile and inline
chrome but desktop dashboards still had to hand-roll their left
sidebar via batch_design. Adds a 240px-wide vertical rail with
icon+label rows, optional brand title, and slate-100 pill bg on the
active item — distinct surface from the existing nav tools so the
decision tree picks it cleanly for "sidebar / side nav / 侧边栏".
Sticky bottom-of-page disclosure card with title, body, accept /
decline buttons (decline-then-accept order), and an optional
"Cookie settings" link for fine-grained consent. Caller positions
the banner; the tool emits the card itself with shadow + 1px slate
border for visual lift above the page content. Plus the v1 corpus
prompt landing-cookie-banner (26 v1 prompts total).
The "Subscribe to newsletter" / "Apply discount code" / "Send chat
message" pattern. Two action variants:
- text (default): pill button with label like "Subscribe"
- icon: 44×44 square icon button (chat send arrow / search apply)
Distinct from add_form_field_v0 (label-above, no inline button)
and add_search_bar_v0 (no trailing action button). Optional
leading_icon adds an icon inside the input itself. Plus the v1
corpus prompt landing-newsletter-signup (25 v1 prompts total).
The "+1 (555) …" pattern from every modern signup / login screen.
A 44px row with leading country selector (flag + dial code +
chevron-down), a 1px slate divider, and the digits input on the
right. Country selector is a button-shape (no actual dropdown
menu); caller handles picker UX as a separate concern. `value`
toggles between placeholder (slate-400) and populated (slate-900).
Third theme-aware v1 (after add_modal_shell_v1 and add_toast_v1).
Same dashed-border "no data yet" tile shape as v0 with a `theme`
param that swaps 5 colors (bg / border / icon / title / subtitle):
- light (default): byte-parity with v0 (slate-50 / slate-300)
- dark: slate-800 bg + slate-600 border + slate-200 title
- system: $color-surface-2 / $color-border / $color-text-primary
refs — requires applySemanticPalette(doc) seeded
Lets dark-theme dashboards keep a matching empty-slot surface
instead of punching a light rectangle out of dark cards.