Why: f1923ff1 added coerceNavTabIcon to bottom-nav-v1 and noted that
sidebar-nav-v1 should adopt the same helper. Without it, a sidebar
nav with \`{ label: 'Profile', icon: 'profile' }\` would still render
a placeholder circle (resolver doesn't know "profile" is a known
wrong-glyph alias for "user") instead of the lucide:user glyph.
What: import coerceNavTabIcon and apply it in buildItemV1 before
stamping iconFontName onto the Icon child node. Same convention
single-sourced. Existing 1435 tests still pass — no test relied on
the prior pass-through behavior for known wrong-glyph names.
Why: end-to-end test of "Design a bottom nav with Home / Search /
Orders / Cart / Profile" with MiniMax-M2.7 surfaced that the model
emits \`{ title: 'Cart', icon: 'shopping-bag' }\` for the Cart tab
~half the time. Both icons exist in lucide but they are different
glyphs — bag is for carrying, cart has wheels for checkout. The
icon-catalog skill update (19ca1c66) fixed it for the planning side
but not for the builder's runtime input — direct \`add_bottom_nav_v1\`
calls still pass through whatever icon the model picks.
What: new \`coerceNavTabIcon(title, icon, builder)\` helper in
coerce-params.ts. Maintains a small Title→canonical-lucide-name map
(Cart→shopping-cart, Profile→user, Home→house, etc., with Chinese
labels) AND a per-canonical KNOWN_WRONG_ALTS list so the swap only
fires when the emitted icon is one of the known wrong-glyph choices
for that title:
Cart + shopping-bag → shopping-cart (warn)
Cart + package → shopping-cart (warn)
Cart + rocket → rocket (pass-through)
Cart + shopping-cart → shopping-cart (silent)
Custom + anything → anything (silent)
The pass-through rule keeps user / model intentional custom choices
intact. Warnings flow through the existing coerce-params sink so
orchestrators can surface them.
bottom-nav-v1.ts now calls coerceNavTabIcon before stamping
iconFontName onto the Tab frame. Sidebar-nav-v1 + similar nav
builders can adopt the same helper later without re-implementing the
map.
10 new unit + integration tests cover: positive swaps for cart /
profile / notifications / Chinese 购物车, pass-through for
custom titles, case-insensitivity, and the builder integration
that the emitted Tab tree carries the canonical iconFontName.
1098 / 1098 tests pass overall (1080 AI + 10 new + 8 elsewhere).
Why: the 6 aesthetic detectors added in 53435bf7 / 7aef1b14 / cd1e4325 /
ad025c95 catch problems AFTER the model emits them. Telling the model
upfront — in the always-loaded layout skill — prevents the same patterns
in the first place. Cheaper than running a corrective post-pass on
every generation, and the model produces cleaner output that doesn't
trip the detectors at all.
What: AESTHETIC HYGIENE block appended to layout.md (priority 10, base,
loaded for every generation). 4 rules each backed by a corresponding
detector:
- Text never gets cornerRadius / stroke / effects / rotation. Mirrors
detectTextCornerRadius / detectTextStroke / detectTextEffect.
- Rotation on UI frames is almost always wrong. Mirrors
detectUnexpectedRotation (with the same 90/180/270 + path/line/polygon
/image escape hatches).
- Same-role siblings must share cornerRadius AND padding. Mirrors
detectMixedSiblingCornerRadius / detectMixedSiblingPadding.
- Inner layout frames (sections, wrappers) inherit from page/card —
only opt into fill/stroke/shadow on the outer card/button/badge/chip.
Mirrors the existing invisible-container detector.
Phrased as a "keep these silent" pre-condition since the post-pass
also strips them. 1080/1080 AI tests + 234/234 pen-ai-skills tests
still pass.
Why: continuation of the aesthetic detector series. Mirrors
detectMixedSiblingCornerRadius (53435bf7) for the padding axis. Three
cards with padding 16 / 16 / 20 looks ragged on canvas; the existing
sibling-inconsistency detector covers cards-vs-cards but dedupes
against cornerRadius and other props so the padding outlier
sometimes drops.
What: detectMixedSiblingPadding normalises padding values to a
4-tuple [top, right, bottom, left] before comparison, so
padding: 16 → [16,16,16,16]
padding: [12, 24] → [12,24,12,24] (CSS 2-tuple shorthand)
padding: [16,16,16,16] → [16,16,16,16]
all compare equal and don't trigger false positives. Modal value
collapses back to a scalar when all four sides are equal so the
suggested fix matches the model's preferred shorthand.
Same 60% modal-majority threshold as the cornerRadius detector —
1-1-1 three-way splits are skipped because there's no canonical value
to suggest. Same divider / spacer skip and same-type-and-role grouping.
Wired through detectAllIssues + index.ts exports + the
debug_validation_report MCP categories enum.
6 new tests cover: number shorthand outlier, number-vs-array
equivalence, 2-tuple-vs-4-tuple equivalence, 1-1-1 split skip,
mixed-role groups skipped, no-padding siblings excluded from modal.
57 / 57 diagnostics tests pass (was 51; +6).
Why: continuation of the aesthetic detector series. Outlined text on a
UI label is almost always an AI mistake — Lucide / SF / Material icons
get stroked, but body / heading / label text is filled. The model
occasionally copies a generic "give it a stroke" instruction onto text
nodes; on canvas the result reads as double-rendered glyphs. The
existing sibling-inconsistency detector doesn't catch this because
text stroke is rarely a sibling-by-sibling outlier — it's emitted
across the whole tree at once.
What:
- detectTextStroke added with the same shape as the other text-only
aesthetic detectors (text node + property check + warning severity +
suggestedValue undefined).
- Skips stroke.thickness === 0 (some model JSON keeps an empty stroke
object as a placeholder; flagging that would be noise).
- Wired through detectAllIssues + index.ts public exports + the
debug_validation_report MCP tool's categories enum.
Tests: 4 new positive + negative cases (text with stroke, text without
stroke, text with thickness=0 placeholder, frame with stroke). 51 / 51
diagnostics tests pass (was 47; +4); 228 / 228 pen-ai-skills overall.
Why: continuation of the aesthetic detector family added in 53435bf7.
The model frequently sprinkles \`effects: [{type:'shadow', …}]\` onto
body / label / caption text. On canvas the type goes fuzzy and reads
"AI-designed". Real product UIs use text shadows extremely sparingly
(hero overlays on photos, a few brand elements). Detection is cheap
(walk + isArray check) and the suggested fix (remove effects array)
is safe — text shadow on UI labels is almost never intentional.
What:
- detectTextEffect added to packages/pen-ai-skills/diagnostics with the
same shape as the prior 3 (warning severity, suggestedValue undefined,
reason string for logs).
- Wired through detectAllIssues + index.ts public exports + the
debug_validation_report MCP tool's categories enum.
Tests: 5 new it() cases covering positive (shadow / blur on text),
negative (text without effects, empty effects array, frame with
effects), and tree-walk (multiple text effects in nested frames).
47 / 47 diagnostics tests pass (was 42; +5).
Why: user reports the validation pipeline lacks "aesthetic standards"
— it accepts misalignment, unwanted corner radius, and other visual
issues as "normal". Existing detectors are pure code-quality (invisible
container / empty path / text height / sibling inconsistency); they
don't catch design-system violations the user can see at a glance.
Vision validation does, but it only runs on Anthropic / Codex /
OpenCode / Gemini providers and only above 30 nodes — leaving a long
tail of small-design / builtin-provider runs with no aesthetic check
at all. Adding cheap pure-function detectors closes that gap with no
upstream provider dependency.
What: 3 new pure detectors in pen-ai-skills/diagnostics:
- detectUnexpectedRotation — flags non-axis-aligned rotation on
UI-bearing nodes (frame / text / shape). Skips path / line /
polygon / image (legitimate decorative geometry frequently
rotated), skips multiples of 90° (intentional vertical text /
grid). Catches the "tilted card" hallucination cleanly.
- detectTextCornerRadius — flags text nodes with cornerRadius > 0.
Text isn't drawn into a clipped rectangle so the prop is silently
dropped at render time, but it survives in the doc and burns
LLM context on subsequent batch_get calls. Suggested fix: remove.
- detectMixedSiblingCornerRadius — stricter than the existing
sibling-inconsistency check on cornerRadius alone. Flags outliers
when 2+ of 3 same-type-and-role siblings share a value and one
differs (e.g. three cards with cornerRadius 8 / 8 / 12 reads as
ragged on canvas). Skips 1-1-1 three-way splits (no canonical
modal) and divider / spacer nodes (visual primitives).
All three are wired through detectAllIssues + the index.ts public
exports + the debug_validation_report MCP tool's `categories` enum so
the user / agent can opt-in or filter via `op debug_validation_report
--categories unexpected-rotation`.
35 new tests cover the load-bearing positive + negative cases for each
detector. 219/219 pen-ai-skills tests pass (was 184; +35). 1080/1080
AI service tests still pass.
Codex P0 mini-gate Round 2 finding (Q5) fix: gesture_re_export.rs tests
set the new W3C fields (KeyEvent.is_composing, FocusEvent.related_node_
id_hint, WheelEvent.delta_z + WheelEvent.mode mutability) but only
asserted the structural compile-time identity, not value readback.
Strengthened to assert every W3C field reads back what was written so
cross-crate type identity AND field-level binary compat are both verified
through the OP re-export path:
- key_event_is_re_exported_from_jian_with_all_w3c_fields: 7-field assert
- focus_event_is_re_exported_from_jian_with_all_w3c_fields: 3-field assert
- wheel_event_is_re_exported_from_jian_with_w3c_fields: defaults +
mutate-and-assert mode + delta_z + delta.x/y
cargo test -p openpencil-shell-core --test gesture_re_export → 6/6 PASS.
Why: MiniMax-M2.7 food-app run rendered Header with white fill on cream
page bg, and used iconFontName=shopping-bag for Cart tab. Two skill-side
issues: icon-catalog.md was self-contradictory ("use path nodes" vs
"use icon_font"), and layout.md had no rule for inner-section bg.
What:
- icon-catalog.md rewritten as "ALWAYS USE icon_font, NEVER path NODES"
with role→name map (Cart→shopping-cart not shopping-bag, Pizza→pizza,
Sushi→fish via alias, etc) and food-category icon list appended.
- layout.md adds: interior section wrappers (Header, Search Section,
Categories Section) MUST have fill:[] (transparent / inherit page bg);
only opt into a fill when the section is intentionally a card with
its own surface tone.
Why: "Design a profile card" through MiniMax-M2.7 produced a 375×803 mobile
screen with auto-injected status bar, because the planner skill listed
"profiles" as a Type 2 single-task screen and the orchestrator's
isMobileScreen heuristic ran on width≤480 alone.
What: design-type.md + decomposition.md add Type 0 (single component:
card / badge / chip / modal) with width=400 height=0 1 subtask no chrome.
isMobileFullScreen helper extracted to orchestrator-plan-classify.ts and
required by both orchestrator.ts and orchestrator-sub-agent.ts so the
two paths can't drift on what "mobile" means (Codex review caught this
when only orchestrator.ts had the new check).
Verified with same MiniMax + same prompt: 400×320 component, 8 nodes,
firstChildRole=card, no status-bar.
The 12th detector landed in 55f1f5d7 fired on any frame 320–480 wide
regardless of height — so a 393×88 bottom-nav, a 393×200 tile, or a
393×500 settings panel being designed in isolation got page-level
gutters auto-applied as if it were the page root.
Tighten the page predicate to require height >= 568 (iPhone SE 1st gen,
the shortest production phone) AND aspect ratio h/w >= 1.5. Real phones
clear both bars; mobile-width components do not.
Surfaced by Codex stop-hook review on the previous commit.
Add detectEdgeSectionPadding (12th pre-validation detector). Flags a
mobile-shaped page root (width 320–480) when its horizontal padding is
0 AND a child content section also has 0 left padding AND that section
contains visible text or icon descendants — the chain that produced the
"Categories" no-padding bug. Suggested fix sets root.padding to
[top, 16, bottom, 16] preserving vertical padding.
False-positive guards skip top-nav / bottom-nav / hero / banner roles
and image-only sections that are intentionally full-bleed.
The new detector lives in its own detectors-spacing.ts since detectors.ts
is already over the 800-line file limit.
Why: 2026-05-10 user report — "你看不到任何真正的生产级设计工具的希望"
called out three persistent visible issues with the food-app design:
1. Search bar shows a "weird inner rounded border" (image 7) — the
wrapper around address+search section sets fill + stroke +
cornerRadius together. strip-redundant-section-fills only cleared
the fill; the leftover stroke + cornerRadius kept drawing the
visible inner pill, which has been there for a long time.
2. Card image has all 4 corners rounded (image 8) — Taco Fiesta /
Bella Italia cards show the food image with bottom corners
rounded too, leaving them visually "ragged" against the title
text below that's flush with the card surface.
(The third — Categories padding — is a layout/intent question; left
for later since it can't be auto-fixed safely without knowing whether
the design wants edge-to-edge content.)
What — two coordinated fixes:
Fix#1: strip-redundant-section-fills now removes stroke and
cornerRadius alongside the fill. A misroll wrapper sets all three
together to "look like a card"; the fill gets stripped on detection
but the leftover chrome kept drawing a phantom card outline. The
three travel together so they should be cleared together. New test
pins the search-bar wrapper case.
Fix#2: new clipCardImageCorners pass in pen-core. When a card-shape
parent (scalar cornerRadius > 0, 2+ children, first child is an
image or canonical image-placeholder, image has its own scalar
cornerRadius) is detected, set parent.clipContent = true and remove
the image's scalar cornerRadius. The card's own corner clip then
cleanly handles the image — top corners round with the card, bottom
corners flush against the title below. 9 unit tests pin the
conservative match policy (silent on standalone image, title-first
card, array-form cornerRadius, cornerRadius:0, no image
cornerRadius, nested cards, existing clipContent).
Wired into applyPostStreamingTreeHeuristics right after
unwrapFakePhoneMockups. 1448 / 1448 pen-core tests pass (was 1438; +10);
1110 / 1110 AI service tests still pass.
Why: 2026-05-10 user report — "Mexican" badge in the food-app screenshot
landed with a "带尖的背景阴影" (pointy / spiked background shadow). The
underlying cause is the model emitting effects with positive spread,
which "bleeds" the shadow color outward and creates a visible bloom /
halo around the badge that doesn't match real product UI shadows.
Real UI shadows are tight: blur 4-16, spread 0, near-black low-alpha.
Existing detectors only handle text effects (text-effect from 7aef1b14);
frame-level effects had no aesthetic gate.
What: detectExcessiveFrameEffects flags a frame node iff ANY of:
- blur > 40 (glow / halo signature; modal-shell scrim uses exactly
40 so the threshold is strict-greater to keep that legitimate use
untouched — verified by detectors-builder-clean.test.ts)
- any effect carries spread > 0 (bleeding outward = the "spiked
shadow" the user called out)
- 3+ stacked effects on one frame (typical UI uses 0-2)
Suggested fix is to remove the effects array; the user / agent can
re-add a proper subtle shadow afterwards if intentional.
Wired through detectAllIssues + index.ts public exports + the
debug_validation_report MCP categories enum. Skips text nodes (those
go through detectTextEffect with a stricter zero-tolerance rule).
7 new tests cover: positive on spread > 0 / blur > 40 / 3+ stacked,
negative on typical subtle shadow / no effects / blur exactly 40
(modal-shell legit) / text node (different detector). 241 / 241
pen-ai-skills tests pass (was 234; +7). 1110 / 1110 AI service tests
pass (unchanged — production builders pre-clean).
Why: 2026-05-10 user report — the food-app "Featured" block landed
with a black background on a cream page. The wrapper had role='card'
AND held 3 restaurant cards each with role='card'. The existing
strip-redundant-section-fills pass treats role='card' as PROTECTED
(cards legitimately own their fills) so the black hedge fill survived.
Visible result: a giant black band between Categories and Popular Near
You that doesn't fit the cream page bg — exactly the "莫名其妙的背景颜色"
issue the user called out.
Root cause: the existing wrapper detection (hasNestedFilledComponent)
only fires for ATOMIC roles (search-bar / button / input / badge /
chip / tag / pill). Container-role wrappers around same-role children
were never matched, so a card-of-cards misroll kept its fill.
What: new hasMultipleSameRoleChildren predicate. Treats a frame as a
section-level wrapper (eligible for safe-dark / safe-light fill
stripping) when ALL of:
- frame role is in CONTAINER_PROTECTED_ROLES (card / banner /
pricing-card / feature-card / image-card / testimonial /
metric-card / gallery-item / phone-mockup)
- frame has ≥ 2 children with the SAME role
- existing safe-hex / root-match check still applies
Net effect: the Featured wrapper's #000000 fill is now stripped, the
section inherits the cream page bg as intended, and the 3 inner
restaurant cards keep their own white surface fills (each was
role='card' but only 1 child of the same role per card, so the
predicate is silent on them).
3 new it() cases pin: Featured-block misroll (3 cards inside) gets
stripped, single-same-role-child stays untouched, banner-wrapping-
banners pattern also strips. 25 / 25 strip tests pass (was 22; +3).
1438 / 1438 pen-core tests pass overall.
Why: f1923ff1 added coerceNavTabIcon to bottom-nav-v1 and noted that
sidebar-nav-v1 should adopt the same helper. Without it, a sidebar
nav with \`{ label: 'Profile', icon: 'profile' }\` would still render
a placeholder circle (resolver doesn't know "profile" is a known
wrong-glyph alias for "user") instead of the lucide:user glyph.
What: import coerceNavTabIcon and apply it in buildItemV1 before
stamping iconFontName onto the Icon child node. Same convention
single-sourced. Existing 1435 tests still pass — no test relied on
the prior pass-through behavior for known wrong-glyph names.
Why: end-to-end test of "Design a bottom nav with Home / Search /
Orders / Cart / Profile" with MiniMax-M2.7 surfaced that the model
emits \`{ title: 'Cart', icon: 'shopping-bag' }\` for the Cart tab
~half the time. Both icons exist in lucide but they are different
glyphs — bag is for carrying, cart has wheels for checkout. The
icon-catalog skill update (19ca1c66) fixed it for the planning side
but not for the builder's runtime input — direct \`add_bottom_nav_v1\`
calls still pass through whatever icon the model picks.
What: new \`coerceNavTabIcon(title, icon, builder)\` helper in
coerce-params.ts. Maintains a small Title→canonical-lucide-name map
(Cart→shopping-cart, Profile→user, Home→house, etc., with Chinese
labels) AND a per-canonical KNOWN_WRONG_ALTS list so the swap only
fires when the emitted icon is one of the known wrong-glyph choices
for that title:
Cart + shopping-bag → shopping-cart (warn)
Cart + package → shopping-cart (warn)
Cart + rocket → rocket (pass-through)
Cart + shopping-cart → shopping-cart (silent)
Custom + anything → anything (silent)
The pass-through rule keeps user / model intentional custom choices
intact. Warnings flow through the existing coerce-params sink so
orchestrators can surface them.
bottom-nav-v1.ts now calls coerceNavTabIcon before stamping
iconFontName onto the Tab frame. Sidebar-nav-v1 + similar nav
builders can adopt the same helper later without re-implementing the
map.
10 new unit + integration tests cover: positive swaps for cart /
profile / notifications / Chinese 购物车, pass-through for
custom titles, case-insensitivity, and the builder integration
that the emitted Tab tree carries the canonical iconFontName.
1098 / 1098 tests pass overall (1080 AI + 10 new + 8 elsewhere).
Why: the 6 aesthetic detectors added in 53435bf7 / 7aef1b14 / cd1e4325 /
ad025c95 catch problems AFTER the model emits them. Telling the model
upfront — in the always-loaded layout skill — prevents the same patterns
in the first place. Cheaper than running a corrective post-pass on
every generation, and the model produces cleaner output that doesn't
trip the detectors at all.
What: AESTHETIC HYGIENE block appended to layout.md (priority 10, base,
loaded for every generation). 4 rules each backed by a corresponding
detector:
- Text never gets cornerRadius / stroke / effects / rotation. Mirrors
detectTextCornerRadius / detectTextStroke / detectTextEffect.
- Rotation on UI frames is almost always wrong. Mirrors
detectUnexpectedRotation (with the same 90/180/270 + path/line/polygon
/image escape hatches).
- Same-role siblings must share cornerRadius AND padding. Mirrors
detectMixedSiblingCornerRadius / detectMixedSiblingPadding.
- Inner layout frames (sections, wrappers) inherit from page/card —
only opt into fill/stroke/shadow on the outer card/button/badge/chip.
Mirrors the existing invisible-container detector.
Phrased as a "keep these silent" pre-condition since the post-pass
also strips them. 1080/1080 AI tests + 234/234 pen-ai-skills tests
still pass.
Why: continuation of the aesthetic detector series. Mirrors
detectMixedSiblingCornerRadius (53435bf7) for the padding axis. Three
cards with padding 16 / 16 / 20 looks ragged on canvas; the existing
sibling-inconsistency detector covers cards-vs-cards but dedupes
against cornerRadius and other props so the padding outlier
sometimes drops.
What: detectMixedSiblingPadding normalises padding values to a
4-tuple [top, right, bottom, left] before comparison, so
padding: 16 → [16,16,16,16]
padding: [12, 24] → [12,24,12,24] (CSS 2-tuple shorthand)
padding: [16,16,16,16] → [16,16,16,16]
all compare equal and don't trigger false positives. Modal value
collapses back to a scalar when all four sides are equal so the
suggested fix matches the model's preferred shorthand.
Same 60% modal-majority threshold as the cornerRadius detector —
1-1-1 three-way splits are skipped because there's no canonical value
to suggest. Same divider / spacer skip and same-type-and-role grouping.
Wired through detectAllIssues + index.ts exports + the
debug_validation_report MCP categories enum.
6 new tests cover: number shorthand outlier, number-vs-array
equivalence, 2-tuple-vs-4-tuple equivalence, 1-1-1 split skip,
mixed-role groups skipped, no-padding siblings excluded from modal.
57 / 57 diagnostics tests pass (was 51; +6).
Why: continuation of the aesthetic detector series. Outlined text on a
UI label is almost always an AI mistake — Lucide / SF / Material icons
get stroked, but body / heading / label text is filled. The model
occasionally copies a generic "give it a stroke" instruction onto text
nodes; on canvas the result reads as double-rendered glyphs. The
existing sibling-inconsistency detector doesn't catch this because
text stroke is rarely a sibling-by-sibling outlier — it's emitted
across the whole tree at once.
What:
- detectTextStroke added with the same shape as the other text-only
aesthetic detectors (text node + property check + warning severity +
suggestedValue undefined).
- Skips stroke.thickness === 0 (some model JSON keeps an empty stroke
object as a placeholder; flagging that would be noise).
- Wired through detectAllIssues + index.ts public exports + the
debug_validation_report MCP tool's categories enum.
Tests: 4 new positive + negative cases (text with stroke, text without
stroke, text with thickness=0 placeholder, frame with stroke). 51 / 51
diagnostics tests pass (was 47; +4); 228 / 228 pen-ai-skills overall.
Why: continuation of the aesthetic detector family added in 53435bf7.
The model frequently sprinkles \`effects: [{type:'shadow', …}]\` onto
body / label / caption text. On canvas the type goes fuzzy and reads
"AI-designed". Real product UIs use text shadows extremely sparingly
(hero overlays on photos, a few brand elements). Detection is cheap
(walk + isArray check) and the suggested fix (remove effects array)
is safe — text shadow on UI labels is almost never intentional.
What:
- detectTextEffect added to packages/pen-ai-skills/diagnostics with the
same shape as the prior 3 (warning severity, suggestedValue undefined,
reason string for logs).
- Wired through detectAllIssues + index.ts public exports + the
debug_validation_report MCP tool's categories enum.
Tests: 5 new it() cases covering positive (shadow / blur on text),
negative (text without effects, empty effects array, frame with
effects), and tree-walk (multiple text effects in nested frames).
47 / 47 diagnostics tests pass (was 42; +5).
Why: user reports the validation pipeline lacks "aesthetic standards"
— it accepts misalignment, unwanted corner radius, and other visual
issues as "normal". Existing detectors are pure code-quality (invisible
container / empty path / text height / sibling inconsistency); they
don't catch design-system violations the user can see at a glance.
Vision validation does, but it only runs on Anthropic / Codex /
OpenCode / Gemini providers and only above 30 nodes — leaving a long
tail of small-design / builtin-provider runs with no aesthetic check
at all. Adding cheap pure-function detectors closes that gap with no
upstream provider dependency.
What: 3 new pure detectors in pen-ai-skills/diagnostics:
- detectUnexpectedRotation — flags non-axis-aligned rotation on
UI-bearing nodes (frame / text / shape). Skips path / line /
polygon / image (legitimate decorative geometry frequently
rotated), skips multiples of 90° (intentional vertical text /
grid). Catches the "tilted card" hallucination cleanly.
- detectTextCornerRadius — flags text nodes with cornerRadius > 0.
Text isn't drawn into a clipped rectangle so the prop is silently
dropped at render time, but it survives in the doc and burns
LLM context on subsequent batch_get calls. Suggested fix: remove.
- detectMixedSiblingCornerRadius — stricter than the existing
sibling-inconsistency check on cornerRadius alone. Flags outliers
when 2+ of 3 same-type-and-role siblings share a value and one
differs (e.g. three cards with cornerRadius 8 / 8 / 12 reads as
ragged on canvas). Skips 1-1-1 three-way splits (no canonical
modal) and divider / spacer nodes (visual primitives).
All three are wired through detectAllIssues + the index.ts public
exports + the debug_validation_report MCP tool's `categories` enum so
the user / agent can opt-in or filter via `op debug_validation_report
--categories unexpected-rotation`.
35 new tests cover the load-bearing positive + negative cases for each
detector. 219/219 pen-ai-skills tests pass (was 184; +35). 1080/1080
AI service tests still pass.
Codex P0 mini-gate Round 2 finding (Q5) fix: gesture_re_export.rs tests
set the new W3C fields (KeyEvent.is_composing, FocusEvent.related_node_
id_hint, WheelEvent.delta_z + WheelEvent.mode mutability) but only
asserted the structural compile-time identity, not value readback.
Strengthened to assert every W3C field reads back what was written so
cross-crate type identity AND field-level binary compat are both verified
through the OP re-export path:
- key_event_is_re_exported_from_jian_with_all_w3c_fields: 7-field assert
- focus_event_is_re_exported_from_jian_with_all_w3c_fields: 3-field assert
- wheel_event_is_re_exported_from_jian_with_w3c_fields: defaults +
mutate-and-assert mode + delta_z + delta.x/y
cargo test -p openpencil-shell-core --test gesture_re_export → 6/6 PASS.
Why: MiniMax-M2.7 food-app run rendered Header with white fill on cream
page bg, and used iconFontName=shopping-bag for Cart tab. Two skill-side
issues: icon-catalog.md was self-contradictory ("use path nodes" vs
"use icon_font"), and layout.md had no rule for inner-section bg.
What:
- icon-catalog.md rewritten as "ALWAYS USE icon_font, NEVER path NODES"
with role→name map (Cart→shopping-cart not shopping-bag, Pizza→pizza,
Sushi→fish via alias, etc) and food-category icon list appended.
- layout.md adds: interior section wrappers (Header, Search Section,
Categories Section) MUST have fill:[] (transparent / inherit page bg);
only opt into a fill when the section is intentionally a card with
its own surface tone.
Why: "Design a profile card" through MiniMax-M2.7 produced a 375×803 mobile
screen with auto-injected status bar, because the planner skill listed
"profiles" as a Type 2 single-task screen and the orchestrator's
isMobileScreen heuristic ran on width≤480 alone.
What: design-type.md + decomposition.md add Type 0 (single component:
card / badge / chip / modal) with width=400 height=0 1 subtask no chrome.
isMobileFullScreen helper extracted to orchestrator-plan-classify.ts and
required by both orchestrator.ts and orchestrator-sub-agent.ts so the
two paths can't drift on what "mobile" means (Codex review caught this
when only orchestrator.ts had the new check).
Verified with same MiniMax + same prompt: 400×320 component, 8 nodes,
firstChildRole=card, no status-bar.
Picks up the keyboard/IME/focus event additions + W3C wheel deltaMode
landed in jian commit d5d358e. shell-core re-exports of the new types
land in the next commit; this commit only moves the pointer + Cargo.lock.
cargo test -p openpencil-shell-core --test gesture_re_export → 6/6 PASS
against the pinned submodule.
ab-corpus rerun (gpt-5.5, ab-v3, 52 prompts × 2 arms): obvious-T M3
59.6% -> 91.5% (+31.9pp); composite-T 0% -> 40% (+40pp). Lift on top
of d5d1a8cd (Rank 1 schema coerce), 9e90cffe (Rank 2 prompt fail watch),
e34d9238 (Rank 3 vision toggle).
Builder fallback minima:
- chart-pie/line/bars-v1: values [1] -> [30,25,20,15,10] / [10,15,12,20,18]
so chart-pie-slice (>=4) and chart-line-dot (>=7) corpus minima are met
- toolbar-v1: fallback items include a divider_after entry so toolbar-divider
role emits even when the model passes only icons
- avatar-group-v1: entry-coerce items with 5 placeholder initials so the
builder always emits avatar-group-{item,initial,overflow,overflow-count}
- combobox/data-table-row/share-row-v1: fallback arrays grown to 3 items
matching the corpus shape minimums
Optional-content discipline (codex stop-time round 2):
- user-card-v1: name field fuzzy-coerce (required field, real fix for the
"Element tool insert failed: I(null,...)" handler bug); the optional role
text stays conditional, never invented. content empty-string was tried
but rejected (empty text nodes still consume flex gap).
- image-placeholder-v1: label stays conditional for the same reason.
Prompt:
- elements.md fail-watch table extended with 5 components (chart legend,
skeleton, inline-action, share-row, combobox) so models routing to
batch_design at least know the role names.
Multi-page vision validation (codex stop-time round 1):
- design-validation.ts: countNodesInActivePage + buildNodeTreeDump now
read getActivePageChildren(activePageId) instead of DEFAULT_FRAME_ID,
so the size-gate and the LLM's tree dump both reflect the page the user
is actually editing rather than the default page. Was a latent bug
surfaced when VALIDATION_ENABLED flipped to true in e34d9238.
Tests: 4223/4223 pass; format:check + tsc clean. 12 files changed.
ab-v8 obvious-T 40 fails matched "missing required role(s)" — model
went batch_design fallback rather than the matching add_*_v1 tool, and
forgot the role names the validator checks. Surface the top 12
fail-mode component-to-tool mappings + their explicit role names at
the top of elements.md (was previously buried 400 lines down in the
keyword section).
Components covered: modal-shell, avatar-group, metric-comparison,
image-placeholder, tag, toolbar, callout, profile-header, inbox-message,
drawer-shell, cookie-banner, user-card.
Even if the model still insists on batch_design (no v1 fits), the
explicit role list helps it emit the correct role strings on each
child node.
Tests: 84/84 pen-ai-skills pass; format:check + tsc clean; skill
budget under 2400 tokens unchanged.
Predicted KPI lift: M3 obvious-T +3-5pp on top of Rank 1's +6pp.
Recovers ~1/3 of the 40 missing-role fails on stronger models
(deepseek/gpt-5.5); weaker models (kimi/minimax) still need the
vision-feedback loop in Rank 3.
Codex flagged: when the convert pass runs BEFORE
normalizeTreeLayout (required to preserve child x/y offsets — see
2fa66bc1), accepting \`layout === undefined\` as a vertical signal
mis-classifies layout-less horizontal rows. A model that emits
two equal-height images side by side without an explicit \`layout\`
field intends a horizontal row; \`inferLayout\` (which normalize
later runs) often agrees. The earlier converter saw the absent
keyword as "vertical-shaped" and flipped the row to absolute,
collapsing both images to (0,0).
Tightened the gate to require explicit \`layout: 'vertical'\`. A
hero that omits the keyword is now an acceptable miss — the
convert pass leaves it for normalize to classify, after which
nothing else fires the layered-detection rule (normalize would
have stripped the children's x/y by then anyway, so even running
convert again post-normalize wouldn't help). The cost is a small
miss rate on extremely sloppy hero outputs; the benefit is no
false positives on legit horizontal rows.
New regression test: layout-less frame with two side-by-side
height-200 images stays untouched. Verified by reverting the
gate to also accept \`undefined\` — the new test correctly fails
("expected false to be true"). All 8 tests pass with the
tightened gate.
Codex flagged: \`normalizeTreeLayout\` strips \`x\` / \`y\` from
non-overlay children of any vertical / horizontal layout container
as a stale-coordinate cleanup. The new
\`convertStackedOverlayToAbsolute\` post-pass was wired in AFTER
normalize, so when a sub-agent emitted an intentional content
offset on a layered hero — e.g.
hero { layout: 'vertical', height: 200, children: [
image { full bg },
overlay { full bg gradient },
content { x: 16, y: 80 } ← inset above the gradient
]}
normalize would delete the \`x: 16, y: 80\` first, then convert
would flip layout to 'none' on a hero whose children have no
positions to honor. The content frame ends up at (0,0) overlapping
the bg image instead of where the model placed it.
Move convert to run BEFORE normalize. After convert, the
container's layout is 'none' so normalize sees an absolute-
positioning container and leaves the children's x/y untouched.
The function is a no-op when no layered pattern matches, so
running it earlier doesn't add cost on the common path.
New test asserts: convert + normalize (in that order) preserves
content's x=16, y=80 through the chain. Verified by reversing the
order in the test — assertion correctly fails with
"expected undefined to be 16", proving the regression coverage
actually exercises the bug condition.
M2.7 food-app run shipped a hero whose content piled into the next
section. Live doc inspection showed:
hero-image-container { width: 'fill_container', height: 200,
layout: 'vertical' }
├─ hero-image { width: 'fill_container', height: 200 }
├─ hero-overlay { width: 'fill_container', height: 200 } // gradient
└─ hero-content { width: 'fill_container', height: 'fit_content' }
├─ "Hungry?" title
└─ search-bar (48 tall)
The model intended the image + overlay to LAYER on top of each
other as bg+gradient with content floating on top. With
\`layout: 'vertical'\` the layout engine instead stacked them
sequentially: 200 + 200 + ~80 = 480, far past the 200 declared
height. No clipContent on the container, so the overflow rendered
into the NEXT sibling section — the user's screenshot showed
"Hungry?" search and category icons piled over the "Near You"
restaurant cards.
\`convertStackedOverlayToAbsolute\` post-pass detects the pattern
conservatively:
- frame, layout='vertical' (or undefined → infers vertical)
- numeric fixed height H
- >= 2 children of types image / rectangle / frame whose height
is exactly H or 'fill_container'
The repair: switch \`layout\` to 'none' so the layout engine
respects each child's own x/y (defaulting to 0/0 = layered) — the
image lands at (0,0), the overlay layers on top, and the content
frame floats on top. Children with explicit positions stay
respected.
Wired into \`design-canvas-ops.ts::applyPostStreamingTreeHeuristics\`
right after \`expandOverflowingFixedHeightCards\` so both layered
and overflowing-fixed-height fixes run together.
6 tests cover: hero pattern conversion, fill_container variant,
plain content stacks left alone (only one bg-like child), no
fixed height left alone, horizontal-layout side-by-side rows
left alone, nested heroes detected.
Drives the three-OS CI matrix verification of the skia-safe + glutin +
glow + winit dep stack per Step 1a spec §7.
- examples/p0_probe.rs: stencil_visibility + readback chain runner (must
own a real OS main thread because winit on macOS rejects
EventLoop::new() from cargo test worker threads).
- tests/p0_probe.rs: subprocess-invoke wrapper, gated
#[ignore = "P0_PROBE_GATE"] so default cargo test stays untouched.
- Cargo.toml: add transient [target.'cfg(not(target_arch = "wasm32"))'.
dev-dependencies] block (skia-safe 0.97 + glutin 0.32.3 + glutin-winit
0.5.0 + glow 0.17.0 + raw-window-handle 0.6.2 + scopeguard 1.2.0 +
winit defaults). Pinned to versions resolved in /tmp/skia-glow-probe.
- .github/workflows/rust-check.yml: install Linux GL prereqs (xvfb,
mesa, libxkbcommon, libwayland) and add a P0-probe-gate step running
cargo test --ignored on each OS (Linux through xvfb-run; Windows
early-returns per spec §8.2 WINDOWS_GPU_DEFERRED_NO_RUNNER).
All three artefacts are TRANSIENT — reverted in a follow-up cleanup
commit after CI is green and the loader-compat notes commit lands.
Task 1 owns the permanent integration.
Codex flagged: when the convert pass runs BEFORE
normalizeTreeLayout (required to preserve child x/y offsets — see
2fa66bc1), accepting \`layout === undefined\` as a vertical signal
mis-classifies layout-less horizontal rows. A model that emits
two equal-height images side by side without an explicit \`layout\`
field intends a horizontal row; \`inferLayout\` (which normalize
later runs) often agrees. The earlier converter saw the absent
keyword as "vertical-shaped" and flipped the row to absolute,
collapsing both images to (0,0).
Tightened the gate to require explicit \`layout: 'vertical'\`. A
hero that omits the keyword is now an acceptable miss — the
convert pass leaves it for normalize to classify, after which
nothing else fires the layered-detection rule (normalize would
have stripped the children's x/y by then anyway, so even running
convert again post-normalize wouldn't help). The cost is a small
miss rate on extremely sloppy hero outputs; the benefit is no
false positives on legit horizontal rows.
New regression test: layout-less frame with two side-by-side
height-200 images stays untouched. Verified by reverting the
gate to also accept \`undefined\` — the new test correctly fails
("expected false to be true"). All 8 tests pass with the
tightened gate.
Codex flagged: \`normalizeTreeLayout\` strips \`x\` / \`y\` from
non-overlay children of any vertical / horizontal layout container
as a stale-coordinate cleanup. The new
\`convertStackedOverlayToAbsolute\` post-pass was wired in AFTER
normalize, so when a sub-agent emitted an intentional content
offset on a layered hero — e.g.
hero { layout: 'vertical', height: 200, children: [
image { full bg },
overlay { full bg gradient },
content { x: 16, y: 80 } ← inset above the gradient
]}
normalize would delete the \`x: 16, y: 80\` first, then convert
would flip layout to 'none' on a hero whose children have no
positions to honor. The content frame ends up at (0,0) overlapping
the bg image instead of where the model placed it.
Move convert to run BEFORE normalize. After convert, the
container's layout is 'none' so normalize sees an absolute-
positioning container and leaves the children's x/y untouched.
The function is a no-op when no layered pattern matches, so
running it earlier doesn't add cost on the common path.
New test asserts: convert + normalize (in that order) preserves
content's x=16, y=80 through the chain. Verified by reversing the
order in the test — assertion correctly fails with
"expected undefined to be 16", proving the regression coverage
actually exercises the bug condition.
M2.7 food-app run shipped a hero whose content piled into the next
section. Live doc inspection showed:
hero-image-container { width: 'fill_container', height: 200,
layout: 'vertical' }
├─ hero-image { width: 'fill_container', height: 200 }
├─ hero-overlay { width: 'fill_container', height: 200 } // gradient
└─ hero-content { width: 'fill_container', height: 'fit_content' }
├─ "Hungry?" title
└─ search-bar (48 tall)
The model intended the image + overlay to LAYER on top of each
other as bg+gradient with content floating on top. With
\`layout: 'vertical'\` the layout engine instead stacked them
sequentially: 200 + 200 + ~80 = 480, far past the 200 declared
height. No clipContent on the container, so the overflow rendered
into the NEXT sibling section — the user's screenshot showed
"Hungry?" search and category icons piled over the "Near You"
restaurant cards.
\`convertStackedOverlayToAbsolute\` post-pass detects the pattern
conservatively:
- frame, layout='vertical' (or undefined → infers vertical)
- numeric fixed height H
- >= 2 children of types image / rectangle / frame whose height
is exactly H or 'fill_container'
The repair: switch \`layout\` to 'none' so the layout engine
respects each child's own x/y (defaulting to 0/0 = layered) — the
image lands at (0,0), the overlay layers on top, and the content
frame floats on top. Children with explicit positions stay
respected.
Wired into \`design-canvas-ops.ts::applyPostStreamingTreeHeuristics\`
right after \`expandOverflowingFixedHeightCards\` so both layered
and overflowing-fixed-height fixes run together.
6 tests cover: hero pattern conversion, fill_container variant,
plain content stacks left alone (only one bg-like child), no
fixed height left alone, horizontal-layout side-by-side rows
left alone, nested heroes detected.
Drives the three-OS CI matrix verification of the skia-safe + glutin +
glow + winit dep stack per Step 1a spec §7.
- examples/p0_probe.rs: stencil_visibility + readback chain runner (must
own a real OS main thread because winit on macOS rejects
EventLoop::new() from cargo test worker threads).
- tests/p0_probe.rs: subprocess-invoke wrapper, gated
#[ignore = "P0_PROBE_GATE"] so default cargo test stays untouched.
- Cargo.toml: add transient [target.'cfg(not(target_arch = "wasm32"))'.
dev-dependencies] block (skia-safe 0.97 + glutin 0.32.3 + glutin-winit
0.5.0 + glow 0.17.0 + raw-window-handle 0.6.2 + scopeguard 1.2.0 +
winit defaults). Pinned to versions resolved in /tmp/skia-glow-probe.
- .github/workflows/rust-check.yml: install Linux GL prereqs (xvfb,
mesa, libxkbcommon, libwayland) and add a P0-probe-gate step running
cargo test --ignored on each OS (Linux through xvfb-run; Windows
early-returns per spec §8.2 WINDOWS_GPU_DEFERRED_NO_RUNNER).
All three artefacts are TRANSIENT — reverted in a follow-up cleanup
commit after CI is green and the loader-compat notes commit lands.
Task 1 owns the permanent integration.
Codex flagged: my previous repair regex matched 3/4/6/8 hex digits,
but \`pen-renderer/paint-utils.ts::parseColor\` only handles
lengths 3, 6, and 8 — the length-4 branch falls through to the
gray fallback. So a raw 4-digit string like \`F00A\` got the \`#\`
prepended and looked like a valid \`#F00A\` color downstream, but
the renderer still painted gray. Net effect: traded one broken
render path (raw-string → fallback) for another (length-4 → fallback)
while masking the schema error so upstream callers couldn't see
it had a problem.
Tightened RAW_HEX_RE to only the three lengths parseColor actually
accepts. 4-digit strings now stay un-prefixed so the schema error
stays visible to tooling that flags malformed hex.
Test updated: drops the F00A → #F00A case from the "repairs N-digit
shapes" matrix and adds a dedicated negative-test case asserting
F00A survives normalization unchanged. Comment in RAW_HEX_RE also
captures the parseColor support matrix and the rationale for not
expanding 4-digit shorthand here — that would require an actual
RGBA-to-RRGGBBAA expansion (e.g. F00A → #FF0000AA), which is a
separate concern that belongs in the renderer or a dedicated
shorthand expander, not in a schema repair pass.
M2.7 food-app run shipped the page root with
fill: [{ type: 'solid', color: 'FFF8F0' }]
(no leading \`#\`). The renderer's hex parser failed → root frame
fell back to its default gray fill → the warm-food cream page bg
disappeared and the whole design read as a generic gray app
instead of the warm-light theme. Bottom nav and other surfaces
were similarly affected when sub-agents emitted raw 6-digit hex
without the prefix.
normalizer now adds the missing \`#\` in place when:
- entry is a SolidFill with a string color
- color starts with neither \`#\` nor \`$\` (so we don't touch
variable refs)
- color matches one of the four hex shapes the renderer accepts:
\`/^[0-9A-Fa-f]{3}([0-9A-Fa-f]([0-9A-Fa-f]{2}([0-9A-Fa-f]{2})?)?)?$/\`
— exactly 3, 4, 6, or 8 hex digits. 5 and 7 digit strings
intentionally don't match (those aren't repairable hex).
Same repair applies to:
- gradient stop colors (linear_gradient + radial_gradient)
- stroke.fill colors (M2.7 also drops the prefix on stroke colors)
6 new tests cover: 6-digit repair, 3/4/8-digit shapes, valid hex
unchanged, \$color-* refs unchanged, non-hex strings (named
colors / partial / 5-7 digit) untouched, stroke fill repair,
gradient stop repair. Verified by temporarily commenting out the
repair calls — 4 tests correctly fail "expected '#FFF8F0' to be
'FFF8F0'", confirming the regression coverage actually exercises
the bug condition.
Codex flagged: the previous image-card test had \`height: 180\` with
an image fill_container child + a moderately long caption. With
the way fitContentHeight resolves a fill_container image's height
(returns 0 when no parent height context), the natural height
landed at ~120 — well below the declared 180 — so the bug
condition \`natural > declared\` never fired and the assertion
\`changed === false\` would have passed even with image-card back
in CARD_ROLES.
Rewrite to actually exercise the regression:
- Drop declared height to 80 (a tight 1:3.75 crop).
- Use a multi-paragraph caption that wraps to ~10 lines at the
card's 300px width — natural height lands at ~210, well past 80.
- Add a sanity assertion (\`fitContentHeight(card) > 80\`) before the
no-change check so future edits to the test fixture can't
silently re-introduce the vacuous-pass shape without setting off
this guard.
- Mirror the same shape under \`role: 'card'\` and assert it DOES
get expanded. The role-based gate is the whole point of the
fix; asserting the contrast across two near-identical fixtures
makes the regression's blast radius and behavior obvious.
Verified by temporarily putting \`image-card\` back into CARD_ROLES:
the test correctly fails with "expected true to be false". With
the fix in place, all 6 tests pass.
Codex flagged: \`image-card\` was in CARD_ROLES, so a 16:9 photo
tile or a 1:1 thumbnail could get silently switched to fit_content
when its computed natural height exceeded the declared one
(image+caption pattern: caption text wraps past the photo crop,
fitContentHeight returns more than the fixed height, my pass
auto-expanded). That breaks the intended visual proportion —
\`image-card\` exists precisely to lock in a fixed crop / aspect
ratio.
Removed \`image-card\` from CARD_ROLES with a scope note explaining
the rationale. Authors who want an image card to grow with content
should use the generic \`role: 'card'\` with an image child instead.
Other card-family roles (card, stat-card, pricing-card,
feature-card, testimonial, event-card, product-card) keep the
auto-expand because they're text-content first and overflow there
is the bug we're trying to fix.
New regression test seeds an image-card with a 16:9 crop + a long
wrapped caption that pushes natural height past the declared 180,
asserts the height stays at 180 and the pass returns false.