Commit graph

305 commits

Author SHA1 Message Date
Fini b2f6ae99ff Merge remote-tracking branch 'origin/v0.8.0' into v0.8.0 2026-05-08 22:33:12 +08:00
Kayshen-X 136274a3ec chore(vendor): bump jian submodule to d5d358e (Step 1b §3.2 P0.5A)
Picks up the keyboard/IME/focus event additions + W3C wheel deltaMode
landed in jian commit d5d358e. shell-core re-exports of the new types
land in the next commit; this commit only moves the pointer + Cargo.lock.

cargo test -p openpencil-shell-core --test gesture_re_export → 6/6 PASS
against the pinned submodule.
2026-05-08 22:03:08 +08:00
Fini 6c5a2c21c6 fix(ai): rank 4 builder + multi-page vision-validation hardening
ab-corpus rerun (gpt-5.5, ab-v3, 52 prompts × 2 arms): obvious-T M3
59.6% -> 91.5% (+31.9pp); composite-T 0% -> 40% (+40pp). Lift on top
of d5d1a8cd (Rank 1 schema coerce), 9e90cffe (Rank 2 prompt fail watch),
e34d9238 (Rank 3 vision toggle).

Builder fallback minima:
- chart-pie/line/bars-v1: values [1] -> [30,25,20,15,10] / [10,15,12,20,18]
  so chart-pie-slice (>=4) and chart-line-dot (>=7) corpus minima are met
- toolbar-v1: fallback items include a divider_after entry so toolbar-divider
  role emits even when the model passes only icons
- avatar-group-v1: entry-coerce items with 5 placeholder initials so the
  builder always emits avatar-group-{item,initial,overflow,overflow-count}
- combobox/data-table-row/share-row-v1: fallback arrays grown to 3 items
  matching the corpus shape minimums

Optional-content discipline (codex stop-time round 2):
- user-card-v1: name field fuzzy-coerce (required field, real fix for the
  "Element tool insert failed: I(null,...)" handler bug); the optional role
  text stays conditional, never invented. content empty-string was tried
  but rejected (empty text nodes still consume flex gap).
- image-placeholder-v1: label stays conditional for the same reason.

Prompt:
- elements.md fail-watch table extended with 5 components (chart legend,
  skeleton, inline-action, share-row, combobox) so models routing to
  batch_design at least know the role names.

Multi-page vision validation (codex stop-time round 1):
- design-validation.ts: countNodesInActivePage + buildNodeTreeDump now
  read getActivePageChildren(activePageId) instead of DEFAULT_FRAME_ID,
  so the size-gate and the LLM's tree dump both reflect the page the user
  is actually editing rather than the default page. Was a latent bug
  surfaced when VALIDATION_ENABLED flipped to true in e34d9238.

Tests: 4223/4223 pass; format:check + tsc clean. 12 files changed.
2026-05-08 15:00:00 +08:00
Fini f7f9226599 feat(ai-skills): elements.md missing-role fail watch table (Rank 2)
ab-v8 obvious-T 40 fails matched "missing required role(s)" — model
went batch_design fallback rather than the matching add_*_v1 tool, and
forgot the role names the validator checks. Surface the top 12
fail-mode component-to-tool mappings + their explicit role names at
the top of elements.md (was previously buried 400 lines down in the
keyword section).

Components covered: modal-shell, avatar-group, metric-comparison,
image-placeholder, tag, toolbar, callout, profile-header, inbox-message,
drawer-shell, cookie-banner, user-card.

Even if the model still insists on batch_design (no v1 fits), the
explicit role list helps it emit the correct role strings on each
child node.

Tests: 84/84 pen-ai-skills pass; format:check + tsc clean; skill
budget under 2400 tokens unchanged.

Predicted KPI lift: M3 obvious-T +3-5pp on top of Rank 1's +6pp.
Recovers ~1/3 of the 40 missing-role fails on stronger models
(deepseek/gpt-5.5); weaker models (kimi/minimax) still need the
vision-feedback loop in Rank 3.
2026-05-08 08:15:00 +08:00
Fini 9513b4692d feat(ai): v1 builder fuzzy-coerce hallucinated params (Rank 1)
ab-v8 KPI showed ~14/95 obvious-T fail = v1 schema throw on hallucinated
enum values (tone='info') / missing required arrays (params.columns
undefined). Builder rejected the whole tool call instead of degrading.

Replace throw paths in 12 v1 builders + entry-coerce in 4 builders that
directly accessed params.X.map() / .forEach():

- chart-{pie,line,bars}-v1: coerceNumberArray fallback [1]
- tag-v1 / heading-v1 / callout-v1 / member-row-v1 / invite-row-v1 /
  activity-log-v1: coerceEnum fallback to schema default
- timeline-v1 / social-login-row-v1: coerceNonEmptyArray with placeholder
- kbd-v1: coerceStringArray fallback ['?']
- data-table-row-v1 / combobox-v1 / toolbar-v1 / share-row-v1: entry
  coerceNonEmptyArray (no prior throw, but params.X.map() crashed on
  undefined input)

New helper packages/pen-core/src/element-builders/coerce-params.ts with
five primitives (coerceEnum, coerceNonEmptyArray, coerceNumberArray,
coerceStringArray, coerceNonEmptyString) + process-global warning sink
for orchestrators to surface coercions to the LLM.

v0 builders unchanged (byte-parity contract still holds — verified by
existing parity tests).

3 pen-core tests + 1 pen-mcp test updated: previously asserted toThrow
on invalid input -> now assert coerce success + warning emission.

Tests: 4223/4223 pass; format:check + tsc clean.

Predicted KPI lift: M3 obvious-T 58% -> ~64% (~14/235 schema-throw fail
recovered; ~6pp). Composite-T unchanged (composite fail is routing/
parse, not schema — verified by ab-v8 raw analysis).
2026-05-08 08:00:00 +08:00
Fini c81cfe82a4 Merge branch 'v0.8.0' of github.com:ZSeven-W/openpencil into v0.8.0
# Conflicts:
#	.github/workflows/rust-multiplatform.yml
#	README.md
#	crates/openpencil-shell-core/src/lib.rs
#	crates/openpencil-shell-native/examples/basic_window.rs
#	crates/openpencil-shell-native/src/lib.rs
2026-05-05 22:51:00 +08:00
Kayshen-X c55807e432 ci: remove TS/Electron workflows (build-electron / ci / docker / publish-cli)
Rust-ification 阶段,CI 只保留 Rust 相关:
- rust-check.yml: cargo fmt + build + test (with STEP1A_REQUIRE_GPU=1 on Linux) + clippy + cargo-deny
- wasm-bundle-check.yml: wasm32 target check

删除:
- build-electron.yml: Electron desktop build (Rust 化后用 openpencil-shell-native)
- ci.yml: TS type-check + Vitest + web build (Rust 化后已废)
- docker.yml: TS Docker image (Rust 化后重做)
- publish-cli.yml: npm packages (Rust 化后改 cargo publish)
2026-05-05 22:09:00 +08:00
Fini bd464a04ea fix(pen-core): overlay converter only matches explicit layout=vertical
Codex flagged: when the convert pass runs BEFORE
normalizeTreeLayout (required to preserve child x/y offsets — see
2fa66bc1), accepting \`layout === undefined\` as a vertical signal
mis-classifies layout-less horizontal rows. A model that emits
two equal-height images side by side without an explicit \`layout\`
field intends a horizontal row; \`inferLayout\` (which normalize
later runs) often agrees. The earlier converter saw the absent
keyword as "vertical-shaped" and flipped the row to absolute,
collapsing both images to (0,0).

Tightened the gate to require explicit \`layout: 'vertical'\`. A
hero that omits the keyword is now an acceptable miss — the
convert pass leaves it for normalize to classify, after which
nothing else fires the layered-detection rule (normalize would
have stripped the children's x/y by then anyway, so even running
convert again post-normalize wouldn't help). The cost is a small
miss rate on extremely sloppy hero outputs; the benefit is no
false positives on legit horizontal rows.

New regression test: layout-less frame with two side-by-side
height-200 images stays untouched. Verified by reverting the
gate to also accept \`undefined\` — the new test correctly fails
("expected false to be true"). All 8 tests pass with the
tightened gate.
2026-05-05 22:06:00 +08:00
Fini a4225ba776 fix(ai): convert overlay-to-absolute runs BEFORE normalizeTreeLayout
Codex flagged: \`normalizeTreeLayout\` strips \`x\` / \`y\` from
non-overlay children of any vertical / horizontal layout container
as a stale-coordinate cleanup. The new
\`convertStackedOverlayToAbsolute\` post-pass was wired in AFTER
normalize, so when a sub-agent emitted an intentional content
offset on a layered hero — e.g.

    hero { layout: 'vertical', height: 200, children: [
      image { full bg },
      overlay { full bg gradient },
      content { x: 16, y: 80 } ← inset above the gradient
    ]}

normalize would delete the \`x: 16, y: 80\` first, then convert
would flip layout to 'none' on a hero whose children have no
positions to honor. The content frame ends up at (0,0) overlapping
the bg image instead of where the model placed it.

Move convert to run BEFORE normalize. After convert, the
container's layout is 'none' so normalize sees an absolute-
positioning container and leaves the children's x/y untouched.
The function is a no-op when no layered pattern matches, so
running it earlier doesn't add cost on the common path.

New test asserts: convert + normalize (in that order) preserves
content's x=16, y=80 through the chain. Verified by reversing the
order in the test — assertion correctly fails with
"expected undefined to be 16", proving the regression coverage
actually exercises the bug condition.
2026-05-05 22:03:00 +08:00
Fini f8f7d0e1ca fix(pen-core): convert stacked-overlay heroes to layout=none
M2.7 food-app run shipped a hero whose content piled into the next
section. Live doc inspection showed:

  hero-image-container { width: 'fill_container', height: 200,
                         layout: 'vertical' }
    ├─ hero-image       { width: 'fill_container', height: 200 }
    ├─ hero-overlay     { width: 'fill_container', height: 200 }  // gradient
    └─ hero-content     { width: 'fill_container', height: 'fit_content' }
        ├─ "Hungry?" title
        └─ search-bar (48 tall)

The model intended the image + overlay to LAYER on top of each
other as bg+gradient with content floating on top. With
\`layout: 'vertical'\` the layout engine instead stacked them
sequentially: 200 + 200 + ~80 = 480, far past the 200 declared
height. No clipContent on the container, so the overflow rendered
into the NEXT sibling section — the user's screenshot showed
"Hungry?" search and category icons piled over the "Near You"
restaurant cards.

\`convertStackedOverlayToAbsolute\` post-pass detects the pattern
conservatively:
  - frame, layout='vertical' (or undefined → infers vertical)
  - numeric fixed height H
  - >= 2 children of types image / rectangle / frame whose height
    is exactly H or 'fill_container'
The repair: switch \`layout\` to 'none' so the layout engine
respects each child's own x/y (defaulting to 0/0 = layered) — the
image lands at (0,0), the overlay layers on top, and the content
frame floats on top. Children with explicit positions stay
respected.

Wired into \`design-canvas-ops.ts::applyPostStreamingTreeHeuristics\`
right after \`expandOverflowingFixedHeightCards\` so both layered
and overflowing-fixed-height fixes run together.

6 tests cover: hero pattern conversion, fill_container variant,
plain content stacks left alone (only one bg-like child), no
fixed height left alone, horizontal-layout side-by-side rows
left alone, nested heroes detected.
2026-05-05 22:00:00 +08:00
Fini 0248c17209 Merge branch 'v0.8.0' of github.com:ZSeven-W/openpencil into v0.8.0 2026-05-05 21:54:00 +08:00
Kayshen-X 22003f9b0c chore(shell-native): add transient P0 probe gate (Step 1a)
Drives the three-OS CI matrix verification of the skia-safe + glutin +
glow + winit dep stack per Step 1a spec §7.

- examples/p0_probe.rs: stencil_visibility + readback chain runner (must
  own a real OS main thread because winit on macOS rejects
  EventLoop::new() from cargo test worker threads).
- tests/p0_probe.rs: subprocess-invoke wrapper, gated
  #[ignore = "P0_PROBE_GATE"] so default cargo test stays untouched.
- Cargo.toml: add transient [target.'cfg(not(target_arch = "wasm32"))'.
  dev-dependencies] block (skia-safe 0.97 + glutin 0.32.3 + glutin-winit
  0.5.0 + glow 0.17.0 + raw-window-handle 0.6.2 + scopeguard 1.2.0 +
  winit defaults). Pinned to versions resolved in /tmp/skia-glow-probe.
- .github/workflows/rust-check.yml: install Linux GL prereqs (xvfb,
  mesa, libxkbcommon, libwayland) and add a P0-probe-gate step running
  cargo test --ignored on each OS (Linux through xvfb-run; Windows
  early-returns per spec §8.2 WINDOWS_GPU_DEFERRED_NO_RUNNER).

All three artefacts are TRANSIENT — reverted in a follow-up cleanup
commit after CI is green and the loader-compat notes commit lands.
Task 1 owns the permanent integration.
2026-05-05 12:23:49 +08:00
Fini 90cd86fc11 Merge branch 'v0.8.0' of github.com:ZSeven-W/openpencil into v0.8.0
# Conflicts:
#	.github/workflows/rust-multiplatform.yml
#	README.md
#	crates/openpencil-shell-core/src/lib.rs
#	crates/openpencil-shell-native/examples/basic_window.rs
#	crates/openpencil-shell-native/src/lib.rs
2026-05-05 12:23:47 +08:00
Kayshen-X d5011547f1 ci: remove TS/Electron workflows (build-electron / ci / docker / publish-cli)
Rust-ification 阶段,CI 只保留 Rust 相关:
- rust-check.yml: cargo fmt + build + test (with STEP1A_REQUIRE_GPU=1 on Linux) + clippy + cargo-deny
- wasm-bundle-check.yml: wasm32 target check

删除:
- build-electron.yml: Electron desktop build (Rust 化后用 openpencil-shell-native)
- ci.yml: TS type-check + Vitest + web build (Rust 化后已废)
- docker.yml: TS Docker image (Rust 化后重做)
- publish-cli.yml: npm packages (Rust 化后改 cargo publish)
2026-05-05 12:23:33 +08:00
Fini c929c6a8b4 fix(pen-core): overlay converter only matches explicit layout=vertical
Codex flagged: when the convert pass runs BEFORE
normalizeTreeLayout (required to preserve child x/y offsets — see
2fa66bc1), accepting \`layout === undefined\` as a vertical signal
mis-classifies layout-less horizontal rows. A model that emits
two equal-height images side by side without an explicit \`layout\`
field intends a horizontal row; \`inferLayout\` (which normalize
later runs) often agrees. The earlier converter saw the absent
keyword as "vertical-shaped" and flipped the row to absolute,
collapsing both images to (0,0).

Tightened the gate to require explicit \`layout: 'vertical'\`. A
hero that omits the keyword is now an acceptable miss — the
convert pass leaves it for normalize to classify, after which
nothing else fires the layered-detection rule (normalize would
have stripped the children's x/y by then anyway, so even running
convert again post-normalize wouldn't help). The cost is a small
miss rate on extremely sloppy hero outputs; the benefit is no
false positives on legit horizontal rows.

New regression test: layout-less frame with two side-by-side
height-200 images stays untouched. Verified by reverting the
gate to also accept \`undefined\` — the new test correctly fails
("expected false to be true"). All 8 tests pass with the
tightened gate.
2026-05-05 12:23:32 +08:00
Fini 999cbaa9c4 fix(ai): convert overlay-to-absolute runs BEFORE normalizeTreeLayout
Codex flagged: \`normalizeTreeLayout\` strips \`x\` / \`y\` from
non-overlay children of any vertical / horizontal layout container
as a stale-coordinate cleanup. The new
\`convertStackedOverlayToAbsolute\` post-pass was wired in AFTER
normalize, so when a sub-agent emitted an intentional content
offset on a layered hero — e.g.

    hero { layout: 'vertical', height: 200, children: [
      image { full bg },
      overlay { full bg gradient },
      content { x: 16, y: 80 } ← inset above the gradient
    ]}

normalize would delete the \`x: 16, y: 80\` first, then convert
would flip layout to 'none' on a hero whose children have no
positions to honor. The content frame ends up at (0,0) overlapping
the bg image instead of where the model placed it.

Move convert to run BEFORE normalize. After convert, the
container's layout is 'none' so normalize sees an absolute-
positioning container and leaves the children's x/y untouched.
The function is a no-op when no layered pattern matches, so
running it earlier doesn't add cost on the common path.

New test asserts: convert + normalize (in that order) preserves
content's x=16, y=80 through the chain. Verified by reversing the
order in the test — assertion correctly fails with
"expected undefined to be 16", proving the regression coverage
actually exercises the bug condition.
2026-05-05 12:23:31 +08:00
Fini 5da1ff84e3 fix(pen-core): convert stacked-overlay heroes to layout=none
M2.7 food-app run shipped a hero whose content piled into the next
section. Live doc inspection showed:

  hero-image-container { width: 'fill_container', height: 200,
                         layout: 'vertical' }
    ├─ hero-image       { width: 'fill_container', height: 200 }
    ├─ hero-overlay     { width: 'fill_container', height: 200 }  // gradient
    └─ hero-content     { width: 'fill_container', height: 'fit_content' }
        ├─ "Hungry?" title
        └─ search-bar (48 tall)

The model intended the image + overlay to LAYER on top of each
other as bg+gradient with content floating on top. With
\`layout: 'vertical'\` the layout engine instead stacked them
sequentially: 200 + 200 + ~80 = 480, far past the 200 declared
height. No clipContent on the container, so the overflow rendered
into the NEXT sibling section — the user's screenshot showed
"Hungry?" search and category icons piled over the "Near You"
restaurant cards.

\`convertStackedOverlayToAbsolute\` post-pass detects the pattern
conservatively:
  - frame, layout='vertical' (or undefined → infers vertical)
  - numeric fixed height H
  - >= 2 children of types image / rectangle / frame whose height
    is exactly H or 'fill_container'
The repair: switch \`layout\` to 'none' so the layout engine
respects each child's own x/y (defaulting to 0/0 = layered) — the
image lands at (0,0), the overlay layers on top, and the content
frame floats on top. Children with explicit positions stay
respected.

Wired into \`design-canvas-ops.ts::applyPostStreamingTreeHeuristics\`
right after \`expandOverflowingFixedHeightCards\` so both layered
and overflowing-fixed-height fixes run together.

6 tests cover: hero pattern conversion, fill_container variant,
plain content stacks left alone (only one bg-like child), no
fixed height left alone, horizontal-layout side-by-side rows
left alone, nested heroes detected.
2026-05-05 12:23:30 +08:00
Fini 46c853f650 Merge branch 'v0.8.0' of github.com:ZSeven-W/openpencil into v0.8.0 2026-05-05 12:23:28 +08:00
Kayshen-X 9b7d96c60e chore(shell-native): add transient P0 probe gate (Step 1a)
Drives the three-OS CI matrix verification of the skia-safe + glutin +
glow + winit dep stack per Step 1a spec §7.

- examples/p0_probe.rs: stencil_visibility + readback chain runner (must
  own a real OS main thread because winit on macOS rejects
  EventLoop::new() from cargo test worker threads).
- tests/p0_probe.rs: subprocess-invoke wrapper, gated
  #[ignore = "P0_PROBE_GATE"] so default cargo test stays untouched.
- Cargo.toml: add transient [target.'cfg(not(target_arch = "wasm32"))'.
  dev-dependencies] block (skia-safe 0.97 + glutin 0.32.3 + glutin-winit
  0.5.0 + glow 0.17.0 + raw-window-handle 0.6.2 + scopeguard 1.2.0 +
  winit defaults). Pinned to versions resolved in /tmp/skia-glow-probe.
- .github/workflows/rust-check.yml: install Linux GL prereqs (xvfb,
  mesa, libxkbcommon, libwayland) and add a P0-probe-gate step running
  cargo test --ignored on each OS (Linux through xvfb-run; Windows
  early-returns per spec §8.2 WINDOWS_GPU_DEFERRED_NO_RUNNER).

All three artefacts are TRANSIENT — reverted in a follow-up cleanup
commit after CI is green and the loader-compat notes commit lands.
Task 1 owns the permanent integration.
2026-05-05 12:23:03 +08:00
Fini 15e51465b3 fix(pen-core): hex prefix repair drops 4-digit RGBA (renderer can't parse)
Codex flagged: my previous repair regex matched 3/4/6/8 hex digits,
but \`pen-renderer/paint-utils.ts::parseColor\` only handles
lengths 3, 6, and 8 — the length-4 branch falls through to the
gray fallback. So a raw 4-digit string like \`F00A\` got the \`#\`
prepended and looked like a valid \`#F00A\` color downstream, but
the renderer still painted gray. Net effect: traded one broken
render path (raw-string → fallback) for another (length-4 → fallback)
while masking the schema error so upstream callers couldn't see
it had a problem.

Tightened RAW_HEX_RE to only the three lengths parseColor actually
accepts. 4-digit strings now stay un-prefixed so the schema error
stays visible to tooling that flags malformed hex.

Test updated: drops the F00A → #F00A case from the "repairs N-digit
shapes" matrix and adds a dedicated negative-test case asserting
F00A survives normalization unchanged. Comment in RAW_HEX_RE also
captures the parseColor support matrix and the rationale for not
expanding 4-digit shorthand here — that would require an actual
RGBA-to-RRGGBBAA expansion (e.g. F00A → #FF0000AA), which is a
separate concern that belongs in the renderer or a dedicated
shorthand expander, not in a schema repair pass.
2026-05-05 12:23:02 +08:00
Fini 07fb3d0d3c fix(pen-core): repair hex colors missing the leading # prefix
M2.7 food-app run shipped the page root with
  fill: [{ type: 'solid', color: 'FFF8F0' }]
(no leading \`#\`). The renderer's hex parser failed → root frame
fell back to its default gray fill → the warm-food cream page bg
disappeared and the whole design read as a generic gray app
instead of the warm-light theme. Bottom nav and other surfaces
were similarly affected when sub-agents emitted raw 6-digit hex
without the prefix.

normalizer now adds the missing \`#\` in place when:
- entry is a SolidFill with a string color
- color starts with neither \`#\` nor \`$\` (so we don't touch
  variable refs)
- color matches one of the four hex shapes the renderer accepts:
  \`/^[0-9A-Fa-f]{3}([0-9A-Fa-f]([0-9A-Fa-f]{2}([0-9A-Fa-f]{2})?)?)?$/\`
  — exactly 3, 4, 6, or 8 hex digits. 5 and 7 digit strings
  intentionally don't match (those aren't repairable hex).

Same repair applies to:
- gradient stop colors (linear_gradient + radial_gradient)
- stroke.fill colors (M2.7 also drops the prefix on stroke colors)

6 new tests cover: 6-digit repair, 3/4/8-digit shapes, valid hex
unchanged, \$color-* refs unchanged, non-hex strings (named
colors / partial / 5-7 digit) untouched, stroke fill repair,
gradient stop repair. Verified by temporarily commenting out the
repair calls — 4 tests correctly fail "expected '#FFF8F0' to be
'FFF8F0'", confirming the regression coverage actually exercises
the bug condition.
2026-05-05 12:23:01 +08:00
Fini b23fba48dc test(pen-core): image-card test actually triggers the bug condition
Codex flagged: the previous image-card test had \`height: 180\` with
an image fill_container child + a moderately long caption. With
the way fitContentHeight resolves a fill_container image's height
(returns 0 when no parent height context), the natural height
landed at ~120 — well below the declared 180 — so the bug
condition \`natural > declared\` never fired and the assertion
\`changed === false\` would have passed even with image-card back
in CARD_ROLES.

Rewrite to actually exercise the regression:

- Drop declared height to 80 (a tight 1:3.75 crop).
- Use a multi-paragraph caption that wraps to ~10 lines at the
  card's 300px width — natural height lands at ~210, well past 80.
- Add a sanity assertion (\`fitContentHeight(card) > 80\`) before the
  no-change check so future edits to the test fixture can't
  silently re-introduce the vacuous-pass shape without setting off
  this guard.
- Mirror the same shape under \`role: 'card'\` and assert it DOES
  get expanded. The role-based gate is the whole point of the
  fix; asserting the contrast across two near-identical fixtures
  makes the regression's blast radius and behavior obvious.

Verified by temporarily putting \`image-card\` back into CARD_ROLES:
the test correctly fails with "expected true to be false". With
the fix in place, all 6 tests pass.
2026-05-05 12:23:00 +08:00
Fini 1c0ed4dcc4 fix(pen-core): expand pass skips image-card (fixed crop is intentional)
Codex flagged: \`image-card\` was in CARD_ROLES, so a 16:9 photo
tile or a 1:1 thumbnail could get silently switched to fit_content
when its computed natural height exceeded the declared one
(image+caption pattern: caption text wraps past the photo crop,
fitContentHeight returns more than the fixed height, my pass
auto-expanded). That breaks the intended visual proportion —
\`image-card\` exists precisely to lock in a fixed crop / aspect
ratio.

Removed \`image-card\` from CARD_ROLES with a scope note explaining
the rationale. Authors who want an image card to grow with content
should use the generic \`role: 'card'\` with an image child instead.

Other card-family roles (card, stat-card, pricing-card,
feature-card, testimonial, event-card, product-card) keep the
auto-expand because they're text-content first and overflow there
is the bug we're trying to fix.

New regression test seeds an image-card with a 16:9 crop + a long
wrapped caption that pushes natural height past the declared 180,
asserts the height stays at 180 and the pass returns false.
2026-05-05 12:22:59 +08:00
Fini 3691eb5f20 fix(pen-core): expand cards whose fixed height clips their content
Image #44 banner shipped with the "Order now" button cut in half:
\`featured-promo-card { role: 'card', height: 165, clipContent: true }\`
held a vertical content stack (badge + title + body + button) whose
natural height was ~220px on the model's wrapped column width. The
card role default sets \`clipContent: true\` to keep image children
inside rounded corners, so the overflow got rendered then clipped at
y=165, making the bottom row of content disappear.

New \`expandOverflowingFixedHeightCards\` post-pass:
- Walks the tree.
- For each frame whose \`role\` is in CARD_ROLES (card, stat-card,
  pricing-card, feature-card, image-card, testimonial, event-card,
  product-card) AND \`height\` is a positive number AND
  \`fitContentHeight(node) > height\`, switches \`height\` to
  \`'fit_content'\`.
- Returns true if any card was patched.

Why fit_content, not removing clipContent: clipContent is what makes
nested image children respect the card's rounded corners. Removing
it would un-clip the button (good) but un-clip the image edges (bad
— image bleeds past the card's corner radius). Just letting the
card grow keeps both invariants right.

Also wired in: \`design-canvas-ops.ts\` calls the new pass right
after \`injectMissingNavSurfaceFill(pageRoot)\` in the streaming /
dispatcher post-pass chain. The card-overflow fix runs ONCE per
post-pass invocation on the page root, so all card-family children
on the page get checked together.

Side fix: button role default for tab-style buttons (parent role is
bottom-tab-bar / tab-bar / tab-row AND layout='vertical') now
returns \`padding: [6, 4], gap: 4\` instead of falling through to
the text-button \`[12, 24]\` default. Only affects the case where
the model omits padding on the tab cell — sub-agents that emit
explicit padding still win (per applyDefaults' missing-only rule).

5 new tests cover: banner-style overflow gets fit_content, fitting
content stays at fixed, non-card roles never get touched, already
auto-sizing cards stay alone, and overflow detection walks into
nested sections.
2026-05-05 12:22:58 +08:00
Fini 6b2e4bf18e fix(ai): nav inject hops single-child wrappers + button icon matches text
Two visible regressions in Image #44:

1. Bottom nav reverted to no-background even though earlier runs
   worked. GPT-5.5 wrapped its bottom nav in a single-child section:
     root > frame{role:'section',id:'bottom-tabs-root'}
          > frame{role:'bottom-tab-bar'} > [tabs]
   The inject pass only walked DIRECT children of root and bailed on
   the section wrapper. Now we hop one level when the wrapper is a
   single-child section AND its sole child is a nav-role frame, so
   the nested nav gets the surface fill + position-aware shadow.
   Multi-child sections still bail (those are real content sections,
   not wrappers).

2. Banner "Order now" CTA shipped with white text + dark icon. My
   prior contrast fix used a luminance-delta threshold of 0.4, but
   #0F172A icon vs #F97316 (orange accent) actually has delta 0.48
   — the threshold said "good contrast, leave it alone" while the
   user sees an obvious mismatch with the white text label.
   Wrong axis: the user's complaint is about CONSISTENCY (icon
   should read as the same token as text), not raw contrast.

   Refactored fixButtonForegroundContrast:
     PASS 1 — find a "reference" foreground from sibling text fill
       (after refs resolve). The model's own text color is the
       authoritative signal for what the button's foreground should
       look like, regardless of what bg/fg luminance suggests.
     PASS 2 — for each icon_font sibling, override when its
       resolved hex differs from the reference fg. Icon-only
       buttons (no text sibling) fall back to a luminance-based
       check at threshold 0.5 — catches dark-on-dark / light-on-
       light pairs that motivated the original rule, without the
       false-negative on saturated mid-luminance bgs (orange).

3 new tests: wrapper-section nav reach, multi-child wrapper bail,
and the regression test for the original "dark-on-dark icon-only
button" still passing under the new luminance-fallback path.
141 tests in the affected suites all green.

Side effect: applyNavSurfaceFill now bails entirely (returns false)
when the nav already has a fill — earlier version still added a
shadow even when fill was preserved, which violated the
"preserves sub-agent intent" semantics the existing tests rely on.
2026-05-05 12:22:57 +08:00
Fini 08bc403f4e fix(pen-core): nav inject also stamps a separating shadow
The food-app run on warm-light theme shipped a bottom-tab-bar with
a valid \$color-surface (white) fill, but the page bg
(\$color-bg-deep) is cream #FFF8F0. The luminance delta between
white and cream is ~0.03 — visually indistinguishable, so the user
reads the nav as having no background even though it does. Image #40
made this concrete: the nav fill landed correctly per live-doc
inspection, but the screenshot still showed icons floating over an
unbroken cream background.

The inject pass already set the surface fill. To survive the
low-fill-contrast case we also stamp a soft shadow:

- bottom-tab-bar → upward shadow (offsetY: -4) lifts the nav off
  the content above. A downward shadow would clip off-screen.
- top-app-bar / top-nav-bar / navbar → downward shadow
  (offsetY: 4). An upward shadow would cling to the screen edge.
- nav / tab-bar / tab-row → ambiguous position, default downward.

Shadow specs (offsetY: ±4, blur: 12, spread: 0, color: #0000000F)
match conventional iOS/Android nav lift values and survive on
ANY page bg color, not just cream — even on dark themes the
extra subtle shadow is invisible (already-dark page) without
breaking the design.

Existing effects on the nav are preserved — sub-agents that
intentionally emit a drop-shadow / glow keep their declaration.

3 new tests cover: bottom-nav gets upward shadow,
top-nav variants get downward shadow, sub-agent's existing
effects survive the inject pass.
2026-05-05 12:22:52 +08:00
Fini 319d6fab60 docs(ai-skills): nudge image_search_query toward 2 keywords (3 zero-results)
The food-app live doc had two unfilled placeholders because the model
emitted 3-keyword queries that Openverse returned zero results for
("burger combo fries", "sakura sushi platter"). The 2-keyword forms
have plenty of matches (240 each).

The previous skill text said "2-3 English keywords" which the model
read as "3 is fine"; concrete examples like `image_search_query:
"burger fries combo"` reinforced the 3-word habit. Updated to:

- "Strongly prefer 2 keywords; never more than 3"
- Explanation of WHY (strict AND-search, concrete zero-result vs hit
  comparisons for "burger fries" / "sushi platter")
- Rule for the 3rd keyword: only when it's a strong common-phrase
  noun ("iced latte" yes, "iced latte coffee" no)
- All worked examples in elements.md row 44 updated to 2 keywords
  ("burger fries", "sushi platter", "chicken bowl")

The server-side fallback (commit 93e5847f) catches 3-keyword
zero-results by retrying with 2 words, so this is a quality nudge
on top of a working safety net — fewer retries means tighter
relevance and faster fill.
2026-05-05 12:22:45 +08:00
Fini f116a89bba docs(ai-skills): per-image image_search_query — never reuse one across cards
Yesterday's GPT-5.5 food-app run shipped with all 5 placeholder frames
carrying `image_search_query: "salmon sushi"`, even though only one of
the dishes was actually salmon sushi (the others were burger combo,
sushi restaurant card, chicken bowl, etc). Once the proxy fix lets
the search reach Openverse, the screen would render five identical
salmon-sushi photos instead of five different food shots.

Root cause is teaching: the previous skill text gave a single example
("burger fries") which the model copy-pasted to every placeholder on
the screen instead of mining each card's own title.

Updates:

- `elements.md` row 44: explicit "MUST receive its own query" + four
  worked examples mapping card titles to per-card queries (Burger
  House → "burger restaurant", Sakura Sushi → "sushi japanese",
  etc.).
- `elements-cookbook.md`: replaces the single example with three
  context-distinct calls + an inline comment warning against reuse.
- `jsonl-format.md` TYPES line: bolded "imageSearchQuery MUST be
  UNIQUE per image — derive it from the surrounding card/dish/section
  text" so the JSONL fallback path gets the same signal.
- `schema.md` image bullet: same uniqueness clause inline.

No code change in this commit — purely prompt-side teaching for the
JSONL + element-tool generation paths.
2026-05-05 12:22:39 +08:00
Fini e04648e651 docs(ai-skills): teach image_search_query on add_image_placeholder
elements.md row 44 now spells out that passing 2-3 English keywords
(e.g. "burger fries", "modern office") via image_search_query is what
lets the auto-search pass swap the gray box for a relevant photo —
otherwise it searches the label or falls back to a generic placeholder.
elements-cookbook adds two example calls so the model has copy-paste
templates for the common case.
2026-05-05 04:02:50 +08:00
Fini 71d3d8b5b7 feat(ai): add image_search_query param to add_image_placeholder_v0/v1
Without an explicit query, the auto-search pipeline can only fall back
to the placeholder's `label` (often unset for context-rich cards) or
finally a generic "placeholder" string — both produce off-topic stock
photos instead of, e.g., burger / sushi shots for a food-app brief.

Builders (`buildImagePlaceholder`, `buildImagePlaceholderV1`) now accept
an optional `image_search_query` param (snake_case to match the rest of
the params interface). When set, it gets stamped onto the resulting
frame as `imageSearchQuery` — the same camelCase field
`image-search-pipeline.ts::extractQueryForNode` already prefers over
`name` and the label child.

Tool definitions in `element-tool-defs-ext-2.ts` (v0) and
`element-tool-defs-ext-6.ts` (v1) expose the new property with a
description that nudges callers to pass 2-3 keywords ("burger fries",
"modern office workspace") for product / restaurant / hero contexts.

3 new tests in `add-image-placeholder-v0.test.ts`: query stamps onto
frame, omitted query leaves field undefined, empty-string query is
treated as missing.
2026-05-05 04:01:51 +08:00
Fini 49964a42cd fix(pen-core): nav fill inject validates per-type required fields
`hasAnyFill` only checked that the first entry's `type` was a string,
which let several malformed shapes bypass injection: `[{type:'solid'}]`
(missing color), `[{type:'solid',color:''}]` (empty color), and
`[{type:'invalid'}]` (unknown variant). All three render as
transparent — effectively unfilled — so the inject pass should patch
them, but the truthy `type` made the function short-circuit and the
nav stayed bare.

Per-type validation:
  - solid: color must be a non-empty string
  - linear_gradient / radial_gradient: stops must be non-empty array
  - image: src must be a non-empty string
  - any other type: treated as unfilled (renderer can't paint it)

Two new tests: malformed solids (missing/empty color, unknown type) and
empty gradient + image-with-empty-src — all properly patched. Existing
preservation tests (real solid, linear_gradient with stops, radial
with stops, image with src) still pass.
2026-05-05 03:07:05 +08:00
Fini 41f49c66e5 fix(pen-core): nav fill inject preserves gradient / image fills
Previous `hasSolidFill` only matched `type === 'solid'`. Sub-agents
legitimately put `linear_gradient` (sunrise hero, accent ribbon),
`radial_gradient` (splash entries), or `image` (branded photo banners)
on top app bars and other nav surfaces, and `hasSolidFill` would
return false for those — making the inject pass overwrite the
gradient/image with a flat `$color-surface` solid.

Renamed to `hasAnyFill`; matches any first-entry shape with a
recognized `type` field. Sub-agent intent (any non-empty fill) now
short-circuits the inject. Three new tests cover linear gradient,
radial gradient, and image fills explicitly — all preserved.
2026-05-05 02:47:22 +08:00
Fini 4212220d57 fix(pen-core): inject default surface fill on top-level nav frames
The previous "navbar in PROTECTED_ROLES" change was Codex-flagged as a
no-op: PROTECTED_ROLES only PREVENTS strip-pass deletion of an existing
fill, it doesn't ADD one. The actual food-app brief failure was that the
sub-agent emitted a bottom navigation row WITHOUT any fill at all,
relying on the parent surface for visual contrast — but the parent (the
cream root frame) doesn't supply that contrast, so the nav blends
straight into the cream background and visually disappears.

New deterministic pass: `injectMissingNavSurfaceFill`. For each direct
child of the page root whose role is one of {navbar, nav, tab-bar,
bottom-tab-bar, top-nav-bar, top-app-bar, tab-row} AND whose fill is
empty/missing, set `fill = [{type: solid, color: '$color-surface'}]`
so the renderer resolves it through the seeded palette and the nav
gets a visible white surface separation from the cream root.

Scope contract:
- Only direct children of the passed root frame (page root). Nav frames
  nested inside cards / sections / banners are left alone.
- Never overrides an existing fill — sub-agent intent (e.g. an
  intentionally dark `top-app-bar`) is preserved.
- Pure mutation; returns `true` when any nav was patched.

Wired into the same hook point as `stripRedundantSectionFills` (via
`design-canvas-ops.ts::generationCleanup`), so every generation cycle
sees both a strip pass (remove hedge fills) and an inject pass (add
the missing nav surface). Five new tests cover all nav role variants,
preservation of existing fills, scope (no recurse into cards), and
no-op on unrelated roles.
2026-05-05 02:42:52 +08:00
Fini d9f8d2d40f fix: navbar fill protection + dispatcher fires image search at subtask level
Two related issues from the GPT-5.5 food-app run:

1. Bottom navigation rendered without its surface fill, blending into
   the cream root background. The strip-redundant-section-fills pass
   didn't have any of the navigation roles (`navbar`, `nav`, `tab-bar`,
   `bottom-tab-bar`, `top-nav-bar`) in PROTECTED_ROLES, so a navbar
   carrying `fill: #FFFFFF` (or any SAFE_LIGHT tint) hit the
   "safe-light hedge" branch and got stripped. Real-world navs
   intentionally use a white surface to separate from a tinted root —
   that fill is intended, not a hedge.

   Fix: add the five navigation role names to PROTECTED_ROLES. New
   test asserts a `role: navbar` frame with `fill: #FFFFFF` on a
   `#FFF8F0` cream root keeps its fill.

2. Empty-src image placeholders inserted by the dispatcher's JSONL
   fallback only got auto-filled at the orchestrator's tail (line
   ~1219, after every subtask completes). On a long brief that's a
   visible lag; on an aborted/throwing brief the tail never runs and
   images stay placeholder forever.

   Fire-and-forget `scanAndFillImages(parentId)` from the dispatcher's
   applied path so each subtask's image set starts searching as soon
   as it lands. The orchestrator-tail scan still runs and dedups
   through `queuedNodeIds`, so this is purely a latency / robustness
   improvement (no double fetch).
2026-05-05 02:29:17 +08:00
Fini 1dd7015c80 fix(pen-core): wrapper detection skips secondary atomics inside primary atomics
The prior pass treated ANY atomic-role frame containing another atomic-
role child with a fill as a wrapper. That's still too aggressive: real
atomic components legitimately compose secondary atomics inside them
(input + trailing icon-button for clear/reveal-password, search-bar +
voice-search icon-button, etc). Stripping the parent's fill in those
cases erases the input/search-bar surface — a regression.

Refine: split atomic protected roles into PRIMARY (input, form-input,
search-bar — input-class components that constitute the "main" atom)
and SECONDARY (button, icon-button, badge, chip, tag, pill — sub-action
or decoration atomics that legitimately nest inside primary atomics).

Wrapper detection now triggers only when:
  - same-role nesting (search-bar > search-bar, input > input), OR
  - PRIMARY atomic nested inside another atomic (search-bar > input —
    the canonical sub-agent misroll).

Two new tests:
- input atom with trailing icon-button (filled clear button) → input
  fill kept
- search-bar atom with voice icon-button (filled accent) → search-bar
  fill kept

Original misroll case (search-bar wrapper > inner input) still strips —
covered by prior test.
2026-05-05 00:33:15 +08:00
Fini e867fcdc50 fix(pen-core): wrapper detection only fires for atomic protected roles
Previous nested-wrapper detection treated any PROTECTED_ROLES frame
containing another protected/structural-fill child as a wrapper. That
swept too widely and could strip fills from real container components:

- card containing a CTA `button` (button is filled, card surface is
  intentional) — card fill stripped if its surface was in SAFE_LIGHT.
- pricing-card with a `badge` ribbon and a CTA button — same issue.
- banner with a nested card — banner fill stripped.

Real component composition is normal; the problem is specifically
sub-agent role mislabels where an ATOMIC component (search-bar, button,
input, badge, chip) is reused as a section wrapper. Container roles
(card, pricing-card, feature-card, banner, etc) NEVER appear as
wrappers — their fill is always intentional.

Fix: introduce ATOMIC_PROTECTED_ROLES (subset of PROTECTED_ROLES) and
restrict wrapper detection to firing only when the OUTER role is in this
atomic set. Container roles stay fully protected.

Three new tests added:
- card with filled button child → card fill kept
- pricing-card with badge + button children → pricing-card fill kept
- banner with nested filled card → banner fill kept

The original misroll case (search-bar > input wrapper) still strips —
covered by the prior test.
2026-05-04 23:27:19 +08:00
Fini b719efd3e8 fix(pen-core): strip safe-light fills on misrolled component-wrapper sections
Real repro from MiniMax-M2.7: sub-agent emits a section wrapper with
the WRONG role applied — Search Bar(role=search-bar) > Search Input
Container(role=input,fill=$color-surface). The outer "search-bar" frame
is actually a section-level wrapper (its child carries the real atom),
but its role is `search-bar` which is in PROTECTED_ROLES, so the strip
pass treated it as the real atom and left its #F8FAFC hedge fill alone.
Result: visible double-cream nesting against the cream root background.

Detect this misroll: a frame whose role IS protected but ALSO contains
a child carrying either the same role or another protected/structural
role with its own solid fill is a wrapper, not the atom — its fill is
eligible for the same safe-light/safe-dark hedge stripping that pure
section frames get.

Counter-case kept covered: a real `search-bar` atom whose children are
just icons / placeholder text (no nested input/search-bar/card/etc with
its own fill) keeps its fill — that fill is intentional, not a hedge.

Two new tests:
- M2.7 misrolled wrapper (search-bar > input + safe-light fill) — outer
  fill stripped, inner input fill preserved.
- Real search-bar atom (no fill-bearing component children) — fill
  preserved.
2026-05-04 23:22:52 +08:00
Fini dfb055eb6a fix(ai-skills): scope JSONL output instructions to fallback-only sections
Codex flagged: even after the previous CRITICAL preamble told the model
to defer to `<op_tool>` mode when an OUTPUT FORMAT block exists later,
the rest of jsonl-format / jsonl-format-simplified still contained
specific JSONL-output directives ("Output a ```json block with ONE node
per line", "FORMAT: _parent (null=root, …)", a full ```json example).
Those specific instructions can dominate over the abstract preamble for
weak models — they read concrete rules and execute them, ignoring the
top-of-skill conditional.

Restructured both skills so JSONL-specific output mechanics are scoped
to a clearly-marked "JSONL FALLBACK MODE" section and the schema
content (TYPES / RULES / DESIGN SYSTEM TOKENS) is mode-agnostic.

- New top-of-skill comment explicitly states TYPES / RULES / TOKENS
  apply to BOTH `<op_tool>` argument shape AND JSONL — neither mode
  contradicts them.
- The "Output ```json block" directive, the "FORMAT: _parent" directive,
  and the ```json example are now wrapped under a "JSONL FALLBACK MODE"
  header that explicitly says "ignore this section if `<op_tool>` mode
  is in effect".

In `<op_tool>` mode the model now reads schema rules without reading
JSONL-specific output mechanics; in JSONL fallback the JSONL section is
unambiguously authoritative. No conflicting instructions for either
output path.
2026-05-04 22:54:37 +08:00
Fini 71f3c04dbf fix(ai): jsonl-format skills coexist with ELEMENT_TOOL_OUTPUT_FORMAT (dual-mode)
The previous fix dropped jsonl-format / jsonl-format-simplified entirely
when elementToolsEnabled was true, on the theory that their CRITICAL
"Output ONLY ```json … Do NOT use tool calls" line conflicted with the
appended `<op_tool>` instruction. But empirically dropping them made
weak-model output WORSE: MiniMax-M2.7 still emits raw JSONL most of the
time (it can't reliably emit `<op_tool>`), and without the JSONL
schema/format teaching its output degrades — role coverage dropped
from 74% to 22%, color-ref% from 84% to 49%.

The right fix is dual-mode coexistence: keep BOTH skills loaded so the
model has the JSONL fallback teaching, but rewrite each skill's CRITICAL
opener to defer to the ELEMENT_TOOL_OUTPUT_FORMAT block when present.

- jsonl-format / jsonl-format-simplified now lead with: "If a separate
  OUTPUT FORMAT — EMIT AS TOOL CALL(S) block appears later in the system
  prompt, FOLLOW THAT block. Use the JSONL form below ONLY when no
  <op_tool> instruction is present."

- Removed the orchestrator-sub-agent.ts skill-filtering branch; both
  skills load unconditionally now.

Net effect: strong models that can follow `<op_tool>` will use the
element-tool path (preserving the n-tools-per-element design intent for
weak-model stability — MiniMax/GLM/Kimi will emit `<op_tool>` when they
can). Weak models that fall back to raw JSONL still get the schema /
sizing / fill / token rules they need to produce coherent output. No
forced choice, no degraded fallback.
2026-05-04 22:47:25 +08:00
Fini 1f2d3c5e1d Merge branch 'v0.8.0' of github.com:ZSeven-W/openpencil into v0.8.0 2026-05-04 21:45:19 +08:00
Kayshen-X a4b7f62e9a Merge feat/rust-ification into v0.8.0 (Step 0 Rust workspace bootstrap)
Step 0 of OP Rust-ification (per kickoff spec v7 FROZEN):
- Cargo workspace at root (members = ["crates/*"], glob)
- 9 skeleton crates: openpencil-app, openpencil-shell-{core,web,native},
  pen-{types,core,engine,codegen,figma}
- rust-toolchain.toml pinned 1.85 (forced from 1.80 → 1.82 → 1.85
  due to crates.io ecosystem edition2024 requirements)
- deny.toml with kickoff §1.2 wasm32 ban invariant
- 2 GitHub Actions: rust-check.yml (3-platform native + cargo-deny)
  and wasm-bundle-check.yml (wasm32 forward + reverse cargo-deny bans)
- vendor/agent submodule → github.com/ZSeven-W/agent-rs
- Bun script wrappers (cargo:check / :test / :wasm-check / :deny)
- README "Rust subsystem" section + Phase boundary note

§1.2 invariants live:
- Forward wasm32 check: shell-web + 5 bucket A crates compile
- Reverse cargo-deny check bans: native + wasm32 both clean
- compile_error guard: shell-native fails wasm32 build with explicit
  message, validated by canary

Step 1+ owns real implementation; Phase 0 docs (snapshot / plan
patches / IPC inventory / parley-taffy matrix / cargo-deny validation)
in openpencil-docs.
2026-05-04 21:00:00 +08:00
Kayshen-X 4e92f0c250 Merge origin/v0.8.0 into feat/rust-ification 2026-05-03 21:00:00 +08:00
MseeP.ai 112921c9ea Add MseeP.ai badge to README.md (#124) 2026-04-29 09:50:57 +08:00
Fini 3275314f40 fix(ai-skills): remove JSONL output references from element-tool path
elements.md still carried two stale claims from the P6 spec era when I
incorrectly assumed the apps/web sub-agent always emitted JSONL:

1. Frontmatter comment (line 17): "embedded orchestrator in apps/web
   emits single-shot JSON and cannot call MCP tools — this skill would
   be 1500 tokens of dead weight there, so it stays excluded."

   Wrong now. With VITE_ENABLE_ELEMENT_TOOLS=1 the embedded orchestrator
   sets `hasMcpTools` and the sub-agent DOES emit `<op_tool>` blocks
   parsed by tryParseAllElementToolOutputs and dispatched via
   element-tools-dispatcher.ts. The skill loads in BOTH paths.

2. Theme handling section opener: "This section applies to the MCP
   tool-call path only … the web-app sub-agent JSONL path forbids tool
   calls — there, write $color-* / $type-* refs directly in JSONL".

   Wrong now. With element tools enabled both paths use `<op_tool>` and
   the same `theme: 'system'` advice applies uniformly. The pointer to
   "DESIGN SYSTEM TOKENS in jsonl-format.md" is dead — that skill was
   just dropped from the element-tool path.

Removed the stale comment, rewrote the comment positively to describe
the dual-path loading. Removed the misleading sub-section header so the
"Default to theme: 'system'" rule applies cleanly to every caller.
2026-04-29 09:50:56 +08:00
Fini 511cc5a1e4 fix(ai): style-guide ranking surfaces warm/industry mobile guides correctly
Two ranking bugs were silently sending mobile food/wellness/fintech briefs
to a desktop landing-page palette:

1. Substring tag inference. /red|red/ matched 'Featured', /health/ matched
   'Healthy' (a category in the food brief), so a food prompt picked up a
   spurious 'wellness' tag and a desktop wellness guide jumped above the
   mobile food guide via tag-overlap math. Added \b word boundaries to
   every English keyword in inferTagsFromPrompt; CJK rules unchanged
   because \b doesn't apply.

2. Industry vs style tag weighting + platform mismatch penalty. Each
   matched tag was worth +10 regardless of meaning, and a platform
   mismatch was a tiny -3 vs +0. So a desktop ecommerce-modern guide
   beating mobile warm-food on the same brief was just `clean+modern+
   rounded` overlapping more than `warm-tones+friendly+rounded` while
   the platform penalty was negligible.
   Now: industry tags (warm-tones / wellness / fintech / developer /
   monospace) score 30, generic style tags 10, platform mismatch -30.
   Empirically pushes warm-food-mobile-light to the top of the food
   brief shortlist (verified with the actual expanded prompt that the
   user's MiniMax-M2.7 run logged).

Same fix applies to every brief that was getting "wrong palette" results
because the planner snippets only contain the top-4 ranked guides — if
the right answer falls past 4, the planner literally never sees it and
the model invents its own (default-blue) palette.

Also: jsonl-format-simplified.md (basic-tier sub-agent prompt) now mirrors
jsonl-format.md's design-system-tokens teaching — basic-tier models like
MiniMax-M2.7 currently emit 0% typography refs because the simplified
prompt doesn't mention $type-* refs at all. The expanded simplified
prompt is 3981 chars, well under the bumped budget=1700 (=6800 char cap).
CRITICAL contract moved to top-of-file as the same defense-in-depth
pattern applied earlier to jsonl-format.md.
2026-04-29 09:50:46 +08:00
Fini 4aa59677be fix(renderer): RenderNode.clipRect → clipStack so each ancestor rrect is preserved
Single ClipInfo can't faithfully encode `(rrect ∩ rrect)` whenever one rect
cuts inside the other's corner. The previous fix collapsed nested clips
into one ClipInfo and dropped one side's rounded corner — which meant a
rounded modal containing rounded cards would silently lose either the
modal's rounding or the card's rounding at paint time.

Fix: replace the single `RenderNode.clipRect: ClipInfo | undefined` with
`clipStack: ClipInfo[]`. Flatten time accumulates a stack from outer-most
ancestor down to the immediate clip-introducing parent. Paint time pushes
each entry as its own canvas.save+clipRect/clipRRect — Skia's clip stack
intersects them naturally, so each level's rounded corner is enforced
independently.

Touched:
- types.ts: export ClipInfo, replace clipRect with clipStack
- document-flattener.ts: thread `clipStack: ClipInfo[]` through recursion;
  push to a copy when isRootFrame || explicitClip
- node-renderer.ts paint: loop over clipStack, push N save+clip ops, pop
  the same N at the end
- renderer.ts (root frame label loop) + skia-engine.ts (root frame label
  loop) + focus-fit.ts (auto-fit excludes clipped descendants) +
  global-export.ts (page bounds): all check clipStack.length instead of
  truthy single field
- skia-interaction.ts: drag/resize/rotate snapshots store and restore
  clipStack arrays (deep-cloned per entry)
- Tests updated + 1 new test: rounded modal containing rounded card
  preserves both rrects on the inner content's clip stack
2026-04-29 09:50:41 +08:00
Fini 220a92c501 fix(renderer): nested clipContent intersects with ancestor clip
Previously `flattenToRenderNodes` overrode the inherited `clipCtx` whenever
a frame had `clipContent: true` (e.g. a card masking its rounded image),
which let the card's children paint past any outer clip — including the
root frame's artboard clip and any horizontal scrolling row's clip.

Repro: a horizontal `clipContent: true` row containing 3 rounded cards
whose total width exceeds the row's visible width. The 3rd card's children
(thumbnail, name text, etc) painted all the way out to the card's own
right edge — past the row, past the root frame, onto the canvas
background.

Fix: introduce `intersectClip(inner, outer)` and use it whenever a nested
clipContent is enabled. `inner` is the new clip we want to introduce (the
current frame's own bounds + cornerRadius); `outer` is the inherited clip
from the ancestor chain. We intersect the rectangles and drop the rounded
corner only if the inner was actually cut on either axis (a single
ClipInfo can't faithfully encode a rrect ∩ rect when the rect cuts inside
a corner).

`outer.rx` is intentionally not propagated — the ambient canvas clip
stack at paint time already enforces the outer rounded shape, so each new
clip just needs to refine the rectangular extent.

Test added: overflowing horizontal scroll row with 3 rounded cards.
Pre-fix: 3rd card's inner text gets clipRect={x:324, w:150}, escaping
the row clip. Post-fix: clipRect={x:324, w:76, rx:0}, properly clipped.
2026-04-29 09:50:40 +08:00
Fini e0039a303e fix(ai): JSONL sub-agent path emits design-system refs not hex
Two-part fix for the web-app chat path (both built-in and CLI mode go through
sub-agent JSONL output, not MCP tool calls — `jsonl-format.md` line 50 forbids
tool calls). Without this, every fill in generated designs was a hex literal
even after the 5/3 design-system-aware work — the 188 v1 element tools and
DEFAULT_PALETTE_FALLBACK were dead weight here.

A. STYLE GUIDE injection now uses ref + (hex) double form so the model sees
   `$color-accent` paired with the resolved hex it represents:

       Before: - Background: #FFF8F0  Surface: #FFFFFF
       After:  - Background: `$color-bg-deep` (resolves to #FFF8F0)
               - Surface: `$color-surface` (#FFFFFF)

   Applied to both `buildSubAgentStyleGuideInstruction` (selectedStyleGuideContent
   path) and the inline `plan.styleGuide` injection in orchestrator-sub-agent.ts.

B. `seedDocVariablesFromStyleGuide` runs once before sub-agent execution: when
   `doc.variables` is empty AND a style guide is selected, it maps the palette
   to v1 token names (`color-bg-deep` / `color-accent` / etc) and seeds them
   into `doc.variables`. This makes refs emitted by the model resolve to the
   user's chosen palette at render time instead of falling back to the default
   #2563EB blue.

A and B are coupled — A alone would make designs render in the wrong color
(every design becomes blue regardless of style guide); B alone leaves the model
mimicking the hex from the prompt. Both must ship together.

Also:
- jsonl-format.md: DESIGN SYSTEM TOKENS section + example fills converted to
  refs + CRITICAL contract moved to top (so future budget overruns can't
  truncate it). Budget bumped 1500 → 1700 for safety margin.
- elements.md: Theme handling section now flags MCP path vs JSONL path so
  models on either path know which guidance applies.
2026-04-29 09:50:35 +08:00
Fini dc656be409 fix(pen-core/variables): resolveNodeForCanvas — fontWeight + empty-vars fallback
GAP-1: add 'fontWeight' to the text-node key list in resolveNodeForCanvas so
$type-*-weight refs resolve to a number before reaching the renderer.

GAP-2: replace the early-exit `if (!variables || Object.keys(variables).length === 0)
return node` with `if (!variables) variables = {}` so DEFAULT_PALETTE_FALLBACK
fires even when the document has an empty variables map (un-seeded v1 docs).

Adds 2 new tests to fallback-equivalence.test.ts covering both gaps.
2026-04-29 09:50:32 +08:00
Fini 1650f81fa8 feat(ai-skills): default decision tree to v1 system mode (P4)
Switch all 89 element tool recommendations in elements.md from v0 to v1
with theme: 'system' as the default. Add "Theme handling" section explaining
when to use system/light/dark/v0. v0 entries retained as "Byte-frozen escape
hatch" sub-entries for rare byte-parity requirements.
2026-04-29 09:50:30 +08:00
Fini 238ab344e2 feat(element-tools): add 11 v1 tools with theme parameter (P3 batch 9 — FINAL)
Converts tabs, tag, text_button, textarea, timeline, toolbar, tooltip,
top_nav_bar, upload_dropzone, user_card, video_placeholder to theme-aware v1.
All 9 touchpoints wired; ext-8 extended and new ext-9 shard created for overflow.
ListTools count: 177 → 188. All 4127 tests pass.

Classification:
- Pass-through (all modes identical, no surface colors): text_button, textarea,
  top_nav_bar, tabs (accent brand-invariant), tooltip (dark=inverted per §3.4),
  tag (status tones per §3.4), video_placeholder (dark bg per §3.4)
- Surface-tint (light/dark/system tokenized): timeline (inactive dot+connector+subtitle),
  toolbar (surface+border+active-bg+icon), upload_dropzone (5 tokens),
  user_card (name+role text)
2026-04-29 09:50:29 +08:00
Fini d547e6917b feat(element-tools): add 10 v1 tools with theme parameter (P3 batch 8)
Converts sidebar_nav, skeleton, social_login_row, spinner, stat_card,
stat_grid, status_badge, step_card, stepper, switch to theme-aware v1.
All 9 touchpoints wired; ext-8 shard extended for schema definitions.
ListTools count: 167 → 177. All 4116 tests pass.

Notable: spinner/stat_grid/status_badge/switch emit identical trees
across all theme modes (caller-param colors, status semantics, or iOS
HIG builder-private literals per spec §3.4) — theme param accepted
for API consistency only.
2026-04-29 09:50:28 +08:00
Fini e1a931e4ff feat(element-tools): add 10 v1 tools with theme parameter (P3 batch 7)
Converts progress_bar, quote_block, radio, range_slider, rating_stars,
search_bar, section_header, segmented_control, select, share_row to
theme-aware v1. All 9 touchpoints wired; adds ext-8 shard for schema
definitions. ListTools count: 157 → 167. All 4106 tests pass.
2026-04-29 09:50:27 +08:00
Fini ab695d9633 feat(element-tools): add 10 v1 tools with theme parameter (P3 batch 6)
Converts metric_comparison, metric_row, nav_chip_row, notification_row,
otp_input, pagination, phone_input, price, pricing_card, profile_header
to theme-aware v1. All 9 touchpoints wired; adds ext-7 shard for schema
definitions. ListTools count: 147 → 157. All 4096 tests pass.
2026-04-29 09:50:26 +08:00
Fini 028ef9048b feat(element-tools): add 10 v1 tools with theme parameter (P3 batch 5)
Adds theme-aware v1 builders for icon_button, image_placeholder,
inbox_message, inline_action, input_with_action, invite_row, kbd,
legend_item, link, and list_row. Group A (zero-color: icon_button,
link, list_row) — no hardcoded colors in v0, all three modes identical.
Group B (kbd) — key bg → surface2, stroke → border in dark/system.
Group C (remaining 6) — surface/text/border/accent/alertColors tokens
applied in dark/system modes, full byte-parity with v0 in light mode.

Extends ext-6 shard (357→647 lines, within 800-line ceiling) housing
all 20 batch-4 + batch-5 tool schema definitions. All 9 touchpoints
wired per playbook: builder, index.ts, pen-core barrel, handler,
dispatcher, ext-6 shard, client shim, server builder, elements.md entries.

Verified: format:check clean, tsc --noEmit clean, 4086/4086 tests pass.
2026-04-29 09:50:25 +08:00
Fini f4a38dc508 feat(element-tools): add 10 v1 tools with theme parameter (P3 batch 4)
Adds theme-aware v1 builders for cookie_banner, data_table_row,
date_picker, drawer_shell, empty_state, event_card, fab, faq_item,
filter_group, and form_field. Group A (zero-color: empty_state,
form_field) — no hardcoded colors in v0, all three modes identical.
Group B (fab) — accent bg is brand-invariant, maps to accent token
in dark/system; icon stays white in all modes. Group C (remaining 7)
— surface/text/border/accent tokens applied in dark/system modes, full
byte-parity with v0 in light mode.

Creates ext-6 shard (ext-5 was at 798-line ceiling) housing all 10
new tool schema definitions (357 lines). All 9 touchpoints wired per
playbook: builder, index.ts, pen-core barrel, handler, dispatcher,
ext-6 shard, client shim, server builder, elements.md entries.

Verified: format:check clean, tsc --noEmit clean, 4076/4076 tests pass.
2026-04-29 09:50:24 +08:00
Fini 957f845bf5 feat(element-tools): add 10 v1 tools with theme parameter (P3 batch 3)
Adds theme-aware v1 builders for chart_bars, chart_line, chart_pie,
chat_bubble, checkbox, chip_input, code_block, color_swatch, combobox,
and comment. Group A (chart tools) maps bar/line color to chart-1 token
and pie default palette to chart-1..6 tokens in dark/system modes.
Group B (color_swatch) is theme-invariant — swatch color is caller-
supplied and passes through unchanged. Group C (chat_bubble, checkbox,
chip_input, code_block, combobox, comment) resolves surface/text/border
via semantic palette tokens. All light modes are byte-parity with v0.
2026-04-29 09:50:23 +08:00
Fini c0717d43c4 feat(element-tools): add 10 v1 tools with theme parameter (P3 batch 2)
Adds theme-aware v1 builders for alert, bottom_nav, breadcrumb,
activity_ring, carousel_dots, action_menu, attachment_row,
calendar_grid, avatar_group, and callout. Group A (zero-color:
alert/bottom_nav/breadcrumb/activity_ring) produce identical output
across all three theme modes. Group B (carousel_dots) maps active=
text-primary, inactive=border in dark/system modes. Group C (action_menu/
attachment_row/calendar_grid/avatar_group) resolve surface/text/border via
semantic palette. Group D (callout) maps tone-keyed bg/fg to alert palette
tokens in dark/system modes. All light modes are byte-parity with v0.
2026-04-29 09:50:22 +08:00
Fini ddfb3fa81b feat(element-tools): add 5 atom v1 tools with theme parameter (P3 batch 1)
avatar-v1, badge-v1, divider-v1, body_text-v1, icon_label-v1 — each with
full 9-touchpoint coverage (pen-core builder + index + pen-mcp handler +
schema shard + dispatcher + apps/web shim + SERVER_BUILDERS + parity test
+ elements.md). Light mode is byte-equal to v0; dark/system modes produce
identical output since all 5 tools emit zero hardcoded color fills — theme
param accepted for API consistency across all v1 tools. New shard
element-tool-defs-ext-5.ts created (ext-4 was at 739 lines). All 2026
pen-core + pen-mcp tests pass; format:check + tsc clean.
2026-04-29 09:50:21 +08:00
Fini 4962d6176e feat(element-tools): add 4 representative v1 tools (Task 2.4)
card_row-v1, setting_row-v1, member_row-v1, activity_log-v1 — each with
full 9-touchpoint coverage (pen-core builder + index + pen-mcp handler +
schema shard + dispatcher + apps/web shim + SERVER_BUILDERS + parity test
+ elements.md). Light mode is byte-equal to v0; dark/system use resolveTheme()
for all color fills. activity_log-v1 maps tone×theme to alertColors tokens
(info/success/warning/danger) with neutral falling back to surface/textMuted.
All 3998 tests pass. Completes P2 representative phase.
2026-04-29 09:50:20 +08:00
Fini 6a74933945 feat(element-tools): add heading-v1 with theme parameter (9 touchpoints)
Task 2.3 — representative v1 tool walkthrough for Plan 14 byte-parity contract.
Light mode is byte-equal to add_heading_v0 (V0_LATIN_PRESETS table reused);
dark/system modes use resolveTheme() for fill color and typography token refs.
Adds theme enum [light, dark, system] to add_heading_v1 MCP schema, elements.md
decision tree, shim-server-parity CASES, and SERVER_BUILDERS. All 3937 tests pass.
2026-04-29 09:50:19 +08:00
Fini 6d055c1bbb feat(empty-chart-v1): full token coverage via resolveTheme (P2.2) 2026-04-29 09:50:18 +08:00
Fini c57bdab2d8 feat(toast-v1): full token coverage via resolveTheme (P2.2) 2026-04-29 09:50:17 +08:00
Fini 8bf59014f0 feat(modal-shell-v1): full token coverage via resolveTheme (P2.2) 2026-04-29 09:50:16 +08:00
Fini f927a200a8 feat(element-builders): resolveTheme helper for v1 builders (P2.1) 2026-04-29 09:50:15 +08:00
Fini caa35e5d8b fix(pen-core): resolve fontSize/lineHeight/letterSpacing/cornerRadius $refs in resolveNodeForCanvas
P1.5 consumer compatibility gate before P2 element-tool v1 builders.

resolveNodeForCanvas was already resolving gap/padding/opacity/color refs,
but missed text typography fields (fontSize, lineHeight, letterSpacing) and
cornerRadius. Since both skia-engine and pen-renderer consume resolver output
before layout runs, unresolved string refs would pass arithmetic as NaN.

layout/engine.ts also patched to guard typeof === 'number' at the 4 text
font-size/line-height access points — defensive fallback to 16/1.x even when
called on pre-resolution nodes (e.g. MCP server, normalizer pipeline).

Adds 21-test p1-5-consumer-compat.test.ts covering all 6 consumer paths.
All 1183 pen-core tests pass. format:check and tsc --noEmit clean.
2026-04-29 09:50:14 +08:00
Fini 8a8151d35c feat(pen-core/variables): resolver-side fallback to default palette (P1.6)
Introduce DEFAULT_PALETTE_FALLBACK (56-entry map built from semantic-palette
at module init) and wire it into resolveVariableRef so that v1 'system' mode
on an un-seeded doc resolves to canonical hex/numeric defaults rather than
undefined — fulfilling the equivalence guarantee from spec §5.3.

Update modal-shell-v1 test to reflect the new contract: un-seeded doc +
Mode:Dark now yields #1E293B (color-surface dark) instead of undefined.
2026-04-29 09:50:13 +08:00
Fini f18f94da84 fix(pen-core/variables): remove 4 extra tokens, add merge map per spec §3.1
The 4 extra single-value tokens (color-accent-dark, color-info-surface,
color-warning-text-strong, color-danger-text-strong) introduced in P1.1.6
violated spec §3.1 / §7.4 — those hex were INTENDED to merge into existing
tokens with ≤ 5% accepted color drift, not become new tokens.

Replaced with MERGE_MAP in measure-v0-hex-coverage.ts that tracks the 4
near-shade redirections (#1D4ED8→color-accent, #EFF6FF→color-info-bg,
#B45309→color-warning-text, #B91C1C→color-danger-text). Cover rate
calculation now reports direct + merge breakdown.

Final palette token count: 56 (28 color + 18 type + 2 letterSpacing +
5 spacing + 3 radius). Cover rate: 28 direct + 4 merge = 32/32 = 100.0%.
2026-04-29 09:50:12 +08:00
Fini cfabfe8a3d style(types): format P1.1-P1.5 changes with oxfmt (no logic change) 2026-04-29 09:50:11 +08:00
Fini 8ab607bc9d feat(types): add 3 radius tokens to semantic-palette (P1.5)
radius-sm=4 / radius-md=8 / radius-lg=12 px — covers chip, card, and sheet
border-radius tiers. All P1 token additions complete (60 total tokens:
32 color + 18 type + 2 letterSpacing + 5 spacing + 3 radius).
2026-04-29 09:50:10 +08:00
Fini eb14bada88 feat(types): add 5 spacing tokens to semantic-palette (P1.4)
spacing-1..5 = 4/8/12/16/24 px — a 4-point base scale covering xs through xl.
All tokens use type="number" with plain scalar values. Palette total grows to 57.
2026-04-29 09:50:09 +08:00
Fini c9aad9f3c9 feat(types): add 2 letterSpacing tokens to semantic-palette (P1.3)
type-display-letter-spacing (-0.5 px) for large display text, and
type-uppercase-label-letter-spacing (1.5 px) for uppercase / overline labels.
Palette total grows to 52.
2026-04-29 09:50:08 +08:00
Fini 67003d72f6 feat(types): add 18 typography tokens to semantic-palette (P1.2)
size + weight + line-height for 6 roles: display / h1 / h2 / h3 / body / caption.
All 18 tokens use type="number" with a plain scalar value (no theme axis).
getSemanticPaletteHex() omits numeric tokens from its return map.
Palette grows from 32 color → 50 total variables.
2026-04-29 09:50:07 +08:00
Fini 7eefa08acf fix(types): add §3.4 exclusion list to cover-rate script, reach 100% (P1.1.6)
10 builder-private hex literals excluded from denominator per Codex B-route
decision. 4 uncovered semantic hex added as new single-value tokens
(color-accent-dark, color-info-surface, color-warning-text-strong,
color-danger-text-strong) → semantic cover rate 87.5% → 100%. Hard gate passes.
2026-04-29 09:50:06 +08:00
Fini b23c55ea10 feat(types): add 14 alert + chart color tokens to semantic-palette (P1.1.5)
8 light/dark alert pairs (info/success/warning/danger bg+text) and 6
single-value chart series colors added to PALETTE. PaletteEntry union
type introduced to support LightDarkEntry | SingleColorEntry | SingleNumberEntry.
getSemanticPalette(), getSemanticPaletteHex(), applySemanticPalette()
updated to handle all three entry shapes. Palette grows from 14 → 28 tokens.
2026-04-29 09:50:05 +08:00
Fini 98921c5846 fix(ai-skills): correct setting-row schema in parent_id teaching examples
df750c98 加的 ❌/✅ 反例里 setting_row 字段名写错了:用了
\`label\` 但 SettingRowParams 实际字段是 \`title\`;trailing.switch
用了 \`value: true\` 但 schema 实际是 \`{ kind: "switch", on: boolean }\`。
Codex stop-time review 抓到 "invalid prompt example would teach a
bad tool schema" — 直接用 source builder 的 SettingRowParams /
SettingRowTrailing 类型作为 ground truth 更新两处例子。
2026-04-29 09:50:03 +08:00
Fini 01b925bd7e feat(ai-skills): teach parent_id rule with concrete WRONG/RIGHT examples
ab-v5 (2026-05-02) 显示 5acde087 教学只是把模型发明的 placeholder
名字从 "root" 改成 "entry-1" / "members-section",没根除"先发明 id
再引用"的习惯。加 ❌/✅ side-by-side 例子让模型直接看到 invented
name 跟 omit 的区别。token 增量 ~200,仍在 Phase 3 trim (-540) 净盈
余范围内。
2026-04-29 09:50:01 +08:00
Fini 2902bdc88e feat(ai-skills): teach parent_id rule + trim PREFER list "Different from"
Two diet/teaching changes to elements.md, both motivated by today's
ab-v4 smoke results:

(a) parent_id teaching — minimax-m2.7 invented "members-section" /
    "canvas" as parent_id values on its multi-tool composite output,
    then every one of its 14 tags failed apply with "parent_id X not
    found in document". Same pattern as yesterday. Adds an explicit
    rule near the top of elements.md (right under the multi-tool
    banner) — `parent_id` is REAL or OMITTED, never invented. The
    `<page>` / `<panel>` / `<sidebar>` placeholders in the cookbook
    recipes are documentation conventions; in actual output, OMIT
    the field. Names the failure mode by reproducing the error
    message format so models learn to avoid it.

(b) Phase 3 token diet — strip ". Different from <tool> (...)"
    disambiguation suffixes from the PREFER list (22 entries had
    them, ~80-200 chars each = ~2.5kb / ~625 tokens saved). The
    primary keyword + tool-name + capability description survives
    intact; the cross-references pointing at sibling tools get
    dropped. Risk: slight increase in wrong-tool routing on
    ambiguous prompts. Worth it for the size reduction; the
    decision tree alone still shows the tool family in context.

Verified by mechanical diff — perl in-place edit, then visual
review confirms no other content was touched. 3785 vitest pass,
format clean, tsc silent.

Net effect on T-prompt size (chars / 4 estimate):
  T + composite + mobile: was 17,930 → now 17,475 (-455 chars)
  T + composite + dashboard: was 18,437 → now 17,983 (-454 chars)
  T + obvious + dashboard: ~14,009 → ~13,556 (-453 chars)

Per-arm savings are smaller than the raw 2.5kb trim because (a)
adds ~700 chars of parent_id teaching. Net win: ~450 chars / ~110
tokens per call. Modest but compounds across 520+ ab-v4 runs.
2026-04-29 09:49:58 +08:00
Fini bafbdfaeef fix(pen-core): reject invalid enum values in 5 more element builders
Sweep follow-up to 113bd55a — same defensive pattern (reject
unknown enum strings at the entry boundary) applied to every other
builder that indexed a Record<EnumLiteral, T> with a value sourced
from raw JSON args.

Builders + enums covered:
- buildTag — TagTone (default | accent | success | warning | error)
- buildCallout — CalloutTone (info | success | warning | danger | note)
- buildActivityLog — tone (info | success | warning | danger | neutral)
- buildInviteRow — InviteStatus (pending | expired | accepted)
- buildMemberRow — trailing.tone for status_dot (online | busy | away
  | offline). role_badge / menu variants skip the check (no tone field)

Same failure mode each one fixed: when a model invents an
out-of-enum string (gpt-5.4 did this with `level: "caption"` in
ab-v4), the lookup `TONES[bad]` / `STATUS_TONE[bad]` returned
undefined, the next property access crashed mid-batch with a
cryptic `undefined is not an object`, and the surrounding dispatch
loop dropped every remaining tag (until df33e937 + 07639f6d landed
the per-shape continuation + partial-doc scoring earlier today).
With validation in place, a bad enum becomes a clean per-shape
error message + the rest of the batch still applies.

13 new edge-case tests cover throw on bad input + valid path on
every enum value + omitted-default for each builder. 3785 vitest
pass, format clean, tsc silent.

Builders not touched: heading.ts (already done in 113bd55a).
Builders that don't fit this pattern (no enum→Record lookup of a
user-controlled string): everything else surveyed via grep on
`Record<.*Tone|Status|Level|Mode|Kind`.
2026-04-29 09:49:57 +08:00
Fini b4b931211c fix(pen-ai-skills): score partial PenDocument when apply.ok=false
Codex stop-time review caught the previous fix (df33e937) handing
the scorer a partial PenDocument that the scorer immediately
ignored. score-run.ts:87 short-circuited on `!applied.ok || !applied.doc`,
so even though apply.ts now surfaces 12 of 13 successfully-applied
tags as a populated `applied.doc`, the row still scored as a total
failure (M1=false, M3=false, m3_failure_reason="apply failed before
shape checks") — exactly the noise df33e937 was meant to eliminate.

Loosens the short-circuit to `!applied.doc` only. When apply.ok=false
but apply.doc is populated, the scorer now:
  - runs the issue detector against the partial doc (issues surface)
  - keeps M1 strict (apply.ok=false → M1=false regardless of detector)
  - decouples M3 from M1: M3 = shape.ok against the partial doc
  - sets m3_failure_reason to the shape miss when shape fails;
    otherwise to "partial apply (M3 met by what landed): <error>" so
    the row reads "tag 12 of 13 broke, but role coverage still met"
    instead of silently swallowing the partial signal
  - surfaces applied.error in row.applyError so per-shape failure
    messages flow into reports

Two new tests cover the new path:
  1. partial apply + shape match → M1=false, M3=true, reason mentions
     "partial apply"
  2. partial apply + shape miss → M1=false, M3=false, shape-miss
     reason wins (structural verdict trumps the partial-apply notice)

Plumbing chain across today's session is now consistent:
  - apply.ts continues past per-shape failures (df33e937)
  - score-run.ts scores the partial doc that lands (this commit)
  - scoring no longer over-attributes to "apply failed" when the model
    actually produced most of the brief

3774 vitest pass (+2), format clean, tsc silent.
2026-04-29 09:49:56 +08:00
Fini 39aeb0c30e fix(pen-core): reject invalid level in buildHeading with a clear error
ab-v4 partial sweep (2026-05-01) caught gpt-5.4 emitting
`add_heading_v0({"content":"Pending invitations","level":"caption"})`
as the 12th tag of a 13-tag composite multi-tool response. Even
though the MCP tool def has `enum: ['display','h1','h2','h3']`, the
ab-corpus harness and the in-process production dispatcher both call
buildHeading() with raw JSON args (no jsonschema gate), so the model's
invented "caption" reached the preset lookup. LATIN_PRESETS["caption"]
is undefined, and the next line `fontSize: preset.fontSize` crashed
the WHOLE batch with `undefined is not an object (evaluating
'preset.fontSize')` — the 11 valid tags ahead of it never landed.

Adds an entry-point validation in buildHeading: if `level` is set and
not in the {display, h1, h2, h3} set, throw with a clear message.
The dispatch loops in apps/web/element-tools-dispatcher and
scripts/ab-corpus/apply both catch per-shape and keep running the
remaining tags, so a single bad level on tag 12 no longer kills tags
1-11 + 13.

3 new edge-case tests in element-builders-edge-cases.test.ts
cover the throw + the four valid levels + the omitted-default case.
3772 vitest pass (+3), format clean, tsc silent.

Other element builders likely have the same pattern (preset lookup
on a string enum without runtime validation) — separate sweep, not
shotgunning here.
2026-04-29 09:49:54 +08:00
Fini c835976479 feat(ab-corpus): per-domain cookbook filter (Phase 2 of token diet)
ab-v3 / ab-v4 showed Phase 1A (cookbook strip on obvious difficulty)
shaved ~4.4k tokens off T-obvious. Phase 2 adds a per-category gate
that strips cookbook recipes whose domain doesn't match the prompt's
category — mobile briefs don't see dashboard recipes, dashboard
briefs don't see mobile / landing recipes, etc.

Mechanism: HTML comment block markers in elements.md
(`<!-- @domain:dashboard --> ... <!-- /@domain -->`) plus a
stripNonMatchingDomains() pass in buildSystemPrompt that drops blocks
whose tag list doesn't include the active category. Untagged content
is "general" and stays in every variant — the safe default.

Tagged 7 single-domain cookbook recipes:
- dashboard: Team / members list, Audit / activity feed, Faceted
  search filter sidebar, Dashboard KPI strip
- landing: Pricing section
- mobile: Onboarding "How it works", Support chat thread

Cross-domain recipes (Login, Signup, Settings page, OTP, Empty
inbox) stay untagged so they load for every category. Decision tree
+ PREFER list also untagged today; the per-tool annotations there
would be a much larger judgment pass for marginal additional savings.

Token measurements (chars / 4 estimate):
                       full     mobile  dashboard  landing
- T + composite       19.0k    17.9k    18.4k     17.7k
                              (-1.1k)  (-0.6k)   (-1.3k)
- T + obvious         14.6k    13.5k    14.0k     13.3k
                              (-1.1k)  (-0.6k)   (-1.3k)

Modest absolute savings — Phase 2 only filters cookbook RECIPES (in
elements.md), and most cookbook content is in elements-cookbook.md
which Phase 1A already strips on obvious. To hit the 6-8k T target
we still need decision-tree compression or PREFER-list trim, but
both are lossier than this gate. Phase 3 candidates noted in the
ab-v4 results doc.

real-model.ts plumbs call.prompt.category through to buildSystemPrompt.
3766 vitest pass (+6 category filter tests including a 500-char
floor regression guard that the filter actually shaves bytes).
2026-04-29 09:49:49 +08:00
Fini cae5501021 fix(ai-skills): purge mixed-strategy teaching from elements.md cookbook
Codex stop-time review (3rd round) caught residual "T prompt
includes mixed-strategy instructions": even after the prompt
forbade mixing, elements.md still taught the mixed pattern in
several places that are also part of the T system prompt.

Cleaned out every spot that paired batch_design with add_*_v0:

- Login screen / Pricing section / Dashboard KPI strip recipes:
  dropped the leading `batch_design: foo = I("page", {...})` line
  and renamed `<foo>` / `<row>` placeholders to `<page>` so each
  recipe is now Strategy A (element tools only). Lost: explicit
  page-level layout/padding/horizontal row — acceptable, the
  recipes still teach the tool selection + chain pattern.

- Intro paragraph: dropped "override via a follow-up batch_design
  U-op if needed" (taught a per-component fallback that the parser
  drops).

- Banner: rewrote the fall-back clause from "when no element tool
  fits a specific component shape" to "when at least one component
  truly needs a custom shape no element tool covers — and then use
  a SINGLE batch_design for the WHOLE response, never mixed."

- "STILL use batch_design when" list: collapsed 3 mixed-strategy
  bullets ("larger composite via batch_design then element tool",
  "post-hoc styling via batch_design U-ops") into 3 clean
  Strategy-B-only bullets, all explicitly emit a SINGLE batch_design.

- Removed the "## Composition pattern" section entirely — its 3-step
  plan was a textbook mixed pattern (batch_design root → element
  tool inserts → batch_design U-op styling) that depended on real
  MCP multi-round semantics the corpus harness can't provide.

- Composition rules of thumb: replaced "Don't mix N-tool and
  batch_design DSL ops in a single call" + "Style overrides come
  AFTER structure" with a single "Don't mix in the same output"
  rule that names the corpus parser behavior explicitly.

3760 vitest pass, format clean, tsc silent. T-obvious 14.5k tokens
(down ~100 from the cleanup); composite still 18.9k.
2026-04-29 09:49:46 +08:00
Fini c2252f2172 fix(ai-skills): drop invalid section_header subtitle from cookbook + stubs
Codex stop-time review flagged the new ab-v3 composite cookbook
recipes calling add_section_header_v0 with a `subtitle` arg the tool
doesn't accept (silently dropped at runtime today, but teaches live
models to emit invalid shapes). The same bug was in the dry-run stub
fixtures.

Split each header into add_section_header_v0(title) +
add_body_text_v0(content) — semantically what the ab-v3 briefs ask
for, and reinforces the multi-tool chaining the cookbook now teaches.
Also fills in the missing required `number` arg on the onboarding
recipe's final completed step card (schema requires it even when
completed=true renders a check instead of the number).
2026-04-29 09:49:40 +08:00
Fini bac47be301 feat(ai-skills): teach multi-tool composition for composite briefs
ab-v3 live sweep showed 0/25 composite-T runs walked the multi-tool
path — every model fell back to batch_design or garbage when given a
brief like "5 member rows + 1 invite row" or "6 audit log entries".
Root cause: elements.md's decision tree opens with "pick first match"
(single-tool framing) and the cookbook's chained examples sit ~400
lines deep, so models stop at the lookup and never realize they can
emit N <op_tool> blocks.

This adds:
- top-of-file "MULTI-TOOL OUTPUT IS THE NORM" banner with a concrete
  5-call settings example, before the decision tree
- decision tree heading rewritten to "(per component — pick first
  match)" so per-component framing is in scope from the start
- 4 new cookbook recipes mirroring the ab-v3 composite prompts:
  team / members list (rows + invite), audit / activity feed,
  faceted search filter sidebar, onboarding step cards

Markdown-only edit to a single skill file; vite-plugin-skills
re-compiles the registry. Banner verified present in the generated
registry, all 3750 vitest tests pass. Verifies in next ab-v4 sweep.
2026-04-29 09:49:39 +08:00
Fini 88505648eb fix(ab-corpus): plumb multi-tool output end-to-end for composite
Codex stop-hook caught: ab-v3 introduced composite-difficulty prompts
that *expect* multi-tool emit (e.g. 5× member_row + 1× invite_row
for a team page), but `ParsedOutput.tool_call` was a single
{name, arguments} so the parser silently dropped every call after
the first. apply.ts only invoked one tool, M3 min_roles couldn't
pass on legitimately-routed multi-tool runs, and byTool stats
under-counted. The composite routing 'multi-tool' bucket was
correctly assigned in classifyRouting, but downstream the pipeline
behaved as if the model emitted a single call.

This commit replaces `kind: 'tool_call'` with
`kind: 'tool_calls'` (NON-EMPTY list) across every consumer:

- types.ts: ParsedOutput tagged union; new ParsedOpToolCall.
  ScoreRow.toolName → toolNames: string[].
- output-parser.ts: collects ALL element-tool tags in emit order;
  unknown-tool path also surfaces as single-element tool_calls so
  routing keeps the same wrong-tool semantics.
- score-run.ts: classifyRouting uses Array.includes for obvious
  prompts (right-tool when ANY emitted call matches expected_tool —
  over-production isn't a routing miss). Composite stays multi-tool
  on any non-empty list.
- aggregate.ts byTool: tallies EVERY name in toolNames, so a
  composite row that emits 6× add_activity_log_v0 + 1×
  add_section_header_v0 contributes 6+1 = 7 invocations across two
  tools (with row-level m1_legal applied to both buckets — apply is
  all-or-nothing).
- apply.ts: loops over parsed.calls and invokes
  handleElementToolCall in emit order. Any single call failing
  aborts the row (M1=false); we don't partial-apply.
- mock-llm.ts mockLlmParsed: collects all `<op_tool>` tags into the
  list (composite-prompt mocks can carry multi-call raw strings).
- apps/web design-parser.tryParseElementToolOutput: maps tool_calls
  → its single-shape DesignOutputShape contract using the FIRST
  call (the multi-tag path `tryParseAllElementToolOutputs` was
  already correct).

Tests: 3746 → 3750 vitest. New cases:
- output-parser: surfaces ALL element-tool tags in emit order with
  intermixed batch_design scaffolds dropped (3 element calls from
  5 tags).
- score-run: right-tool when expected appears alongside extras;
  composite multi-call captures every name in toolNames.
- aggregate: 6× activity_log + 1× section_header → byTool reports
  6 and 1 invocations respectively.

dry-run on ab-v3 produces a 208-row report; tsc + format clean.
2026-04-29 08:35:00 +08:00
Fini 1c3cbda495 feat(ab-corpus): ab-v3 yaml fixtures (7 obvious + 5 composite)
Adds 7 new obvious prompts covering the v0.8.0 element tools that
weren't in ab-v1 (tools 91-97):
  setting_row / member_row / filter_group / invite_row /
  activity_log / event_card / step_card

One yaml per tool, same single-component "Design ONLY..." pattern
as ab-v1, with must_contain_roles mirroring the role names emitted
by the corresponding builder in pen-core/src/element-builders/.

Adds 5 composite prompts that exercise the new M6 routing
breakdown:
  - dashboard-settings-page-composite (4× setting_row)
  - dashboard-team-people-page-composite (5× member_row + 1×
    invite_row)
  - dashboard-search-filters-composite (2× filter_group + result
    list)
  - dashboard-audit-feed-composite (6× activity_log)
  - mobile-onboarding-flow-composite (4× step_card)

Each composite prompt omits expected_tool_if_any (multi-tool
intent) and uses min_roles to enforce the multi-element shape.

corpus-loader: composite added to VALID_DIFFICULTIES; composite
prompts MUST NOT specify expected_tool_if_any (validation error
points the user back to difficulty=obvious if a single tool fits).
6 new corpus-loader tests cover the v3 yaml inventory + composite
validation rules. 3740 → 3746 vitest tests, all green.

ab-v3 corpus now has 52 prompts: 40 inherited from v1 (unchanged
for v1↔v3 comparability) + 7 obvious + 5 composite.
2026-04-29 08:15:00 +08:00
Fini 95566e4ed2 feat(ab-corpus): bootstrap ab-v3 with token cost + composite difficulty
ab-v3 succeeds ab-v1 (frozen 2026-04-28). Carries forward all 40
v1 obvious yaml files unchanged so the v1↔v3 overlap stays
comparable, then layers in two new dimensions.

**1. Token cost.** All clients (openai-compat, ark, bailian,
deepseek, minimax, codex-cli, stub-model) now return a
`ChatCallResult { content, usage }` instead of bare string.
Provider usage stats (`prompt_tokens` / `completion_tokens`) plumb
through realModelCall → run.ts → scoreRun → ScoreRow.{prompt,completion}Tokens.
aggregate adds avgPromptTokens{Baseline,Treatment} +
avgCompletionTokens{Baseline,Treatment} per ModelSummary.
write-report emits a new "Token cost" table with Δ columns so
narrow-tools-saves-tokens (the ab-v2 hypothesis) is measurable.
avgUsage skips rows with 0/0 usage so codex-cli (CLI doesn't
surface tokens) and harness errors don't deflate the average to
near-zero — they show '—' instead.

**2. Composite difficulty.** New 'composite' value alongside
obvious / optional. Composite prompts express multi-tool intents
where no single expected_tool_if_any applies. classifyRouting
routes composite-treatment runs into multi-tool / fallback /
garbage (3-bucket sum to 1, distinct from obvious's 4-bucket
right/wrong/fallback/garbage). aggregate adds m6_multi_tool +
m6_fallback + m6_garbage; write-report emits a "Composite routing"
table that gracefully degrades to a placeholder when no composite
yaml exists yet.

Harness side: scripts/ab-corpus/run.ts accepts --corpus ab-v3
(enum + parseArgs guard); dry-run on the v1-mirror corpus produces
a 160-row report including populated token table.

Tests: 4 new aggregate cases (composite, token avg with skip-zero,
NaN-when-no-data) + 4 new score-run cases (composite routing
multi-tool/fallback/garbage/baseline-n/a) + 2 new score-run cases
(usage plumbing) + 2 new openai-compat cases (usage parsing,
missing-usage fallback). Existing 5 retry tests updated for new
return shape. 3727 → 3740 vitest tests, all green; tsc + format
clean.

Token-cost docs and composite docs go straight into types.ts /
score-run.ts / aggregate.ts JSDoc — keeps the contract close to
the code that owns it.
2026-04-29 07:45:00 +08:00
Fini 9c605a3785 fix(ai): step-card marker clips overflow + tighten doc to short index
The previous schema description and JSDoc invited "Step 1" as a valid
value for `number`, but that prose has 6 chars and overflows the 36px
circle marker. Two changes:

- Builder: add clipContent: true on the marker frame so any caller
  who ignores the docs at least gets a clipped (not bleeding) render.
- Schema + JSDoc: drop the misleading "Step 1" example, document the
  1–3 character contract, and steer prose toward `title` instead.
2026-04-28 08:55:00 +08:00
Fini 1fee613111 feat(ai): ship 5 element tools to reach 97 (filter_group / invite_row / activity_log / event_card / step_card)
Closes the obvious gaps remaining in the family:
- add_filter_group_v0 — sidebar facet (heading + checkbox-style options
  with optional counts). Distinct from nav_chip_row (horizontal scrolling
  chips), tag (single applied chip), segmented_control (mutex tabs).
- add_invite_row_v0 — pending invite row (avatar + email/role + status
  pill + trailing action). Distinct from member_row (a JOINED member,
  no status pill or action) and list_row (no avatar / status / action).
- add_activity_log_v0 — single-line audit feed entry (optional tinted
  icon dot + actor in bold + action + right-aligned timestamp). Uses
  StyledTextSegment[] content for the bold/regular split. Distinct from
  timeline (multi-event vertical with connectors) and notification_row
  (title + body, no actor focus).
- add_event_card_v0 — single calendar event tile (date column with
  month band + day number, then title + time + location). Distinct from
  calendar_grid (the full month grid) and card_row (no date column).
- add_step_card_v0 — onboarding step card (numbered circle / check +
  title + description). Distinct from stepper (horizontal progress nav
  with connectors) and faq_item (collapsible Q&A header).

9 touchpoints per tool: pen-core builder + index + barrel + types,
pen-mcp handler + dispatcher + ext-4 schema, apps/web shim +
SERVER_BUILDERS, parity test (+5 cases), elements.md decision tree
items 86-89 + 6 PREFER mappings with cross-links to existing tools,
elements-cookbook.md arg-shape examples (8 entries across 5 tools).
2026-04-28 08:50:00 +08:00
Fini a5b594cf69 feat(ai): add_member_row_v0 — team / member list row (92nd tool)
Avatar + (name over optional subtitle) + optional trailing slot
(role badge / kebab menu / status dot). Distinct from
add_user_card_v0 (compact fit_content tile, no trailing slot) and
add_list_row_v0 (no avatar slot — leading icon instead).

9 touchpoints wired: pen-core builder + index + barrel + types,
pen-mcp handler + dispatcher + ext-4 schema, apps/web shim +
SERVER_BUILDERS, parity test, elements.md decision tree #84 +
PREFER mapping, cookbook arg shapes (3 variants).

Also disambiguates add_avatar_group_v0's PREFER mapping: drop
"团队成员" (now points at member_row), keep narrower phrases like
"成员头像" / "团队头像" / "presence indicator" that genuinely match
the stacked-avatars affordance, and add the cross-link to member_row.
2026-04-28 08:05:00 +08:00
Fini 77cc8ef3a7 fix(ai-skills): teach setting-row recipe with the right tool, drop invalid trailing_kind
Two bugs in elements.md after add_setting_row_v0 landed:
1. PREFER mapping still routed "settings row" to add_list_row_v0 — direct
   contradiction with the new add_setting_row_v0 entry below it.
2. The "Settings page" recipe called add_list_row_v0 with a trailing_kind:
   "switch" arg, but list-row has no such param; it would silently render
   without a switch (or fail validation in stricter clients).

Reroute settings-row prose mapping to the new tool and rewrite the recipe
to use add_setting_row_v0 with proper trailing variants.
2026-04-28 07:15:00 +08:00
Fini af1ddd1ad2 feat(ai): add_setting_row_v0 — settings menu row (91st tool)
Leading icon + (title over optional subtitle) + trailing control with
4 variants: chevron / value text / switch / badge. Distinct from
add_list_row_v0 (trailing is always icon, no switch/value/badge) and
add_form_field_v0 (label-above-input for forms).

Wires all 9 touchpoints: pen-core builder + index + barrel re-export,
pen-mcp handler + dispatcher case + ext-4 schema, apps/web shim +
Nitro SERVER_BUILDERS, elements.md decision tree #83 + PREFER mapping,
elements-cookbook arg-shape examples, plus shim-server parity case.
2026-04-28 06:30:00 +08:00
Fini e656b52f95 fix(mcp): describe section enum by parameter name, not invented schema path 2026-04-27 09:45:00 +08:00
Fini 0bed3a5d12 fix(mcp): drop hardcoded section list from get_design_prompt description
The tool-level description still listed 11 specific sections (schema /
layout / roles / text / style / icons / examples / guidelines /
planning / elements / design-md) even though the actual catalog has
28. Replace with a pointer at inputSchema.section.enum, which is
already derived from listPromptSections() — the description now
can't go stale as sections are added or removed.

D0 parity snapshot refreshed for the new description text.
2026-04-27 09:40:00 +08:00
Fini e8cbe712e2 fix(mcp): get_design_prompt enum derives from SECTION_MAP, no more drift
The published section enum had drifted to 14 entries while
SECTION_MAP grew to 28 (copywriting / overflow / cjk / variables +
8 codegen-* + elements-cookbook). External MCP clients calling with
the missing names hit schema-validation rejection even though the
implementation could serve them. Codex caught the immediate
elements-cookbook gap; widening the fix because the same pattern was
already silently broken for half the catalog.

- enum now derives from listPromptSections() at module load, no
  hand-maintained list to drift
- description points at listPromptSections() rather than enumerating
  individual sections (which was its own drift vector)
- design-prompt-elements adds a sync drift-guard: enum set must
  exactly equal listPromptSections() set
- D0 parity snapshot refreshed for the new enum values

3712 tests green.
2026-04-27 09:35:00 +08:00
Fini 3e3b02f227 test(mcp): refresh D0 parity snapshot for elements-cookbook section enum 2026-04-27 09:30:00 +08:00
Fini fa099ce282 fix(mcp): add elements-cookbook to get_design_prompt section enum 2026-04-27 09:25:00 +08:00
Fini f5eff2c9a3 fix(ai-skills): restore element-tool arg-shape examples in elements-cookbook.md
Trimming the Minimal usage block out of elements.md (65c31832) lost
arg-shape templates that the A/B harness depends on. Text-only LLMs
in the treatment arm see only the markdown skill content — no MCP
tools/list, no published inputSchema — so without the per-tool
example payloads they have to guess argument names and break M1.

Restore the full block as a sibling skill `elements-cookbook` (same
hasMcpTools flag, slightly later priority so it loads alongside
elements). Wire it through buildFullPrompt + the
get_design_prompt(section='elements-cookbook') section map. Update
the A/B harness to also strip the cookbook body when building the
baseline prompt — leaving it in B would leak tool names + arg shapes
back into the no-tools variant and re-bias the comparison.

Both files now under the 800-line per-file ceiling.
2026-04-27 09:20:00 +08:00
Fini 2939b6d4b5 refactor(ai-skills): trim elements.md minimal-usage block to fit 800-line ceiling
elements.md grew to 859 lines after the 81-90 batch shipped — past the
repo's per-file ceiling. The 366-line "Minimal usage" section was the
biggest contributor and the most redundant: MCP `tools/list` already
publishes the full inputSchema for every element tool (arg names,
types, descriptions, requireds), so the LLM has authoritative arg
shape from the wire. The decision tree + PREFER mappings already
teach WHEN to pick each tool. Drop the inline usage examples; keep
the composition pattern + cookbook recipes (which schemas can't
convey) plus invariants and failure-mode guidance.

493 lines remaining; design-prompt-elements + every drift guard still
green.
2026-04-27 09:15:00 +08:00