Visible regression in the GPT-5.5 food-app run: the bottom nav
shipped with no surface fill (floating icons on the cream root
background), even though `injectMissingNavSurfaceFill` was wired in
and verified to add a fill on top-level nav-role frames. Live doc
inspection showed `bottom-tab-bar` carrying `fill: undefined`
post-generation — the inject pass simply never ran.
Root cause: `orchestrator-sub-agent.ts` runs the dispatcher branch
(Strategy A `<op_tool>` element-tools AND Strategy B JSONL-in-
batch_design) and **early-returns before** reaching
`applyPostStreamingTreeHeuristics(rootId)` further down in the
function. That post-pass is what runs:
- normalizeStrokeFillSchema
- unwrapFakePhoneMockups
- resolveTreeRoles + resolveTreePostPass
- normalizeTreeLayout
- stripRedundantSectionFills
- injectMissingNavSurfaceFill
- publish (forcePageResync)
Skipping it on the dispatcher path means EVERY sub-agent that emits
via element tools or JSONL fallback bypasses role resolution, layout
normalization, redundant-fill stripping, AND nav-surface injection.
The streaming path was the only branch that fired the cleanup.
Fix: call `applyPostStreamingTreeHeuristics(subtask.parentFrameId ??
plan.rootFrame.id)` right before the dispatcher branch returns, when
at least one node was inserted. The post-pass walks up to the page
root via `getParentOf()` for the inject step, so passing the section
root that the dispatcher inserted into is correct.
This also un-blocks several heuristics that depend on the full
subtree being in the store: button width / frame height equalization,
clipContent on cards-with-image-children, and theme detection on the
sub-agent's root (which feeds icon/text color defaults).
The previous commit ran `bun install` while ~/.npmrc set
`registry=https://registry.npmmirror.com/`, so every tgz URL in
bun.lock got rewritten from the empty-default-registry form (`""`)
to an explicit `https://registry.npmmirror.com/...` URL. That pinned
the entire workspace install to a regional mirror that other
contributors and CI don't have access to.
Re-ran the install with ~/.npmrc temporarily moved aside so undici's
new entry lands with the workspace's normal "" URL convention while
every other dep's URL is reset to "" too. apps/web's workspace block
now has `"undici": "^7.22.0"` while the package metadata block keeps
the same empty-URL shape as the rest of the lockfile.
No content / version changes — only URL field resets.
Previous version did `require('undici')` inside a try/catch on the
theory that would let it run on both CJS and ESM. In Vite/Nitro's dev
path the helper loads as an ESM module, where `require` is undefined —
the call threw `ReferenceError: require is not defined`, the catch
block silenced it, `configured=true` still got flipped, and every
subsequent call short-circuited. Net effect: the proxy was never
installed in the very dev environment the fix was meant to repair, so
image-search kept ECONNREFUSED-ing on Openverse + Wikimedia and
landing zero filled placeholders.
Switch to a static `import { setGlobalDispatcher, EnvHttpProxyAgent }
from 'undici'`. undici is a transitive dep of h3 in this workspace
(verified resolved in node_modules), and it's also the package Node 18+
uses internally for fetch — pinning it as a direct dep on apps/web
makes the resolution intentional rather than reliant on the h3 chain.
Also swapped the hand-rolled ProxyAgent for undici's built-in
`EnvHttpProxyAgent`: it reads HTTPS_PROXY / HTTP_PROXY / NO_PROXY
itself (case-insensitive) and applies the no-proxy bypass list, which
saves us from re-implementing those rules.
Verified with both `bun -e` (workspace deps) AND a direct ESM Node
context: with HTTPS_PROXY set, `fetch(api.openverse.org/...)` now
returns 240 results for "salmon sushi" instead of the earlier
ECONNREFUSED. The "configured" guard still makes calls idempotent so
multiple endpoints can opt in without coordinating.
Yesterday's GPT-5.5 food-app run shipped with all 5 placeholder frames
carrying `image_search_query: "salmon sushi"`, even though only one of
the dishes was actually salmon sushi (the others were burger combo,
sushi restaurant card, chicken bowl, etc). Once the proxy fix lets
the search reach Openverse, the screen would render five identical
salmon-sushi photos instead of five different food shots.
Root cause is teaching: the previous skill text gave a single example
("burger fries") which the model copy-pasted to every placeholder on
the screen instead of mining each card's own title.
Updates:
- `elements.md` row 44: explicit "MUST receive its own query" + four
worked examples mapping card titles to per-card queries (Burger
House → "burger restaurant", Sakura Sushi → "sushi japanese",
etc.).
- `elements-cookbook.md`: replaces the single example with three
context-distinct calls + an inline comment warning against reuse.
- `jsonl-format.md` TYPES line: bolded "imageSearchQuery MUST be
UNIQUE per image — derive it from the surrounding card/dish/section
text" so the JSONL fallback path gets the same signal.
- `schema.md` image bullet: same uniqueness clause inline.
No code change in this commit — purely prompt-side teaching for the
JSONL + element-tool generation paths.
Root cause for all-blank-placeholders on the food-app brief: Node's
native fetch (used by the Nitro dev server's image-search endpoint)
ignores the system proxy by default. On machines that route outbound
HTTPS through a local proxy (clash / mihomo / corporate gateway —
mine sits at 127.0.0.1:7897), every Openverse + Wikimedia call from
the server silently ECONNREFUSEDs. The endpoint's catch block returns
`null` for Openverse → falls back to Wikimedia → that ECONNREFUSEDs
too → returns `[]`. Browser shows zero filled images.
Direct curl from the same machine uses HTTPS_PROXY automatically, which
is why a manual API check (e.g. `curl https://api.openverse.org/...`)
returned 240 results for "salmon sushi" while
`/api/ai/image-search?query=salmon%20sushi` returned `{results:[]}`.
`apps/web/server/utils/proxy-dispatcher.ts::configureProxyDispatcher`:
- Reads HTTPS_PROXY / https_proxy / HTTP_PROXY / http_proxy.
- If set, installs `undici.ProxyAgent` as the global fetch dispatcher
via `setGlobalDispatcher`. From that point on every server-side
`fetch()` routes through the proxy.
- Idempotent — multiple endpoints can call it without re-installing.
- No-op when no proxy env var is present (production / CI).
- Dynamic `require('undici')` so a build target that strips undici
doesn't crash at import time.
Wired into `image-search.ts` at module top so the dispatcher is
configured before the first request lands. Other endpoints making
external fetches can opt in with the same single-line call.
Verified standalone via Bun: with the helper in place,
`fetch('https://api.openverse.org/v1/images/?q=salmon+sushi')` returns
240 results. The dev server itself needs a restart to pick up the
server-side change (Vite server-code HMR doesn't re-evaluate Nitro
modules).
Previous fix left `role: 'image-placeholder'` on the frame even after
its fill was swapped to `[{type:'image', url, mode:'crop'}]` — the role
is what makes "this slot is meant to hold a photo" semantics survive
into history / codegen / downstream tooling, so stripping it would
trade one regression for another.
But that meant any follow-up generation (which calls
`resetImageSearchQueue` to clear `queuedNodeIds`) would re-walk the
tree, re-enqueue the same placeholder via role match, and overwrite
the already-good photo with whatever the next search returned.
`isUnfilledImagePlaceholderFrame` now gates every read: role match AND
fill is not already `type: 'image'`. Used in three places:
- `collectImageSearchTargets` only collects unfilled placeholders.
- `enqueueImageForSearch` early-returns if the caller passes an
already-filled placeholder (defense-in-depth for direct callers).
- `processQueue`'s re-check uses it instead of a plain
`isImagePlaceholderFrame`, so even a stale queue entry from before
someone else filled the frame gets dropped.
4 new tests in image-search-pipeline.test.ts cover the predicate
(default solid fill = unfilled; missing/empty fill = unfilled; image
fill = filled; non-placeholder role = always false) and a regression
test in `collectImageSearchTargets` that keeps the already-filled
placeholder out of the result while still picking up its sibling.
elements.md row 44 now spells out that passing 2-3 English keywords
(e.g. "burger fries", "modern office") via image_search_query is what
lets the auto-search pass swap the gray box for a relevant photo —
otherwise it searches the label or falls back to a generic placeholder.
elements-cookbook adds two example calls so the model has copy-paste
templates for the common case.
Without an explicit query, the auto-search pipeline can only fall back
to the placeholder's `label` (often unset for context-rich cards) or
finally a generic "placeholder" string — both produce off-topic stock
photos instead of, e.g., burger / sushi shots for a food-app brief.
Builders (`buildImagePlaceholder`, `buildImagePlaceholderV1`) now accept
an optional `image_search_query` param (snake_case to match the rest of
the params interface). When set, it gets stamped onto the resulting
frame as `imageSearchQuery` — the same camelCase field
`image-search-pipeline.ts::extractQueryForNode` already prefers over
`name` and the label child.
Tool definitions in `element-tool-defs-ext-2.ts` (v0) and
`element-tool-defs-ext-6.ts` (v1) expose the new property with a
description that nudges callers to pass 2-3 keywords ("burger fries",
"modern office workspace") for product / restaurant / hero contexts.
3 new tests in `add-image-placeholder-v0.test.ts`: query stamps onto
frame, omitted query leaves field undefined, empty-string query is
treated as missing.
The `add_image_placeholder_v0` / `_v1` element tools and JSONL payloads
that mimic them emit a `frame` carrying `role: 'image-placeholder'` (a
gray slate-100 box + centered icon_font child + optional label) — NOT
an `image` node. The auto-search pipeline only filtered on
`type === 'image'`, so every placeholder produced via element tools
silently bypassed the search hook. Latest GPT-5.5 food-app run shipped
8 placeholder frames; zero got auto-filled and the design landed with
all dashed-border icons instead of real photos.
Pipeline now:
- `isImagePlaceholderFrame` predicate identifies placeholder frames.
- `collectImageSearchTargets` returns mixed `{node, kind}` pairs
('image' for `type==='image'` with placeholder src, 'placeholder-frame'
for the role-keyed frames). Skips descending into placeholder
children (icon_font + label get wiped on fill anyway).
- `enqueueImageForSearch` accepts both shapes; queue items track `kind`.
- `processQueue` re-checks the right invariant per kind, and on success
uses `updateNode(id, { fill: [{type:'image',url,mode:'crop'}], children: [] })`
for placeholder frames (vs `updateNode(id, { src })` for image nodes).
Clearing children prevents the icon/label from rendering on top of
the searched photo.
- Streaming path (insertStreamingNode line 382) intentionally still
gates on `type === 'image'` — placeholder frames stream their
children separately, so enqueueing mid-stream would race with the
late-arriving icon. Placeholder frames are only enqueued via the
post-tree `scanAndFillImages` scan (orchestrator-tail + dispatcher
per-subtask), where the full tree is already in the doc.
`extractQueryForNode` looks for `imageSearchQuery` first, falls back
to a non-default `name`, then mines the optional
`role: 'image-placeholder-label'` text child for a hint. Generic
default still works ("placeholder") if nothing useful is on the frame.
7 new tests cover `isImagePlaceholderFrame` and
`collectImageSearchTargets` (placeholder + image mix, no descent into
placeholder children, missing root id).
`hasAnyFill` only checked that the first entry's `type` was a string,
which let several malformed shapes bypass injection: `[{type:'solid'}]`
(missing color), `[{type:'solid',color:''}]` (empty color), and
`[{type:'invalid'}]` (unknown variant). All three render as
transparent — effectively unfilled — so the inject pass should patch
them, but the truthy `type` made the function short-circuit and the
nav stayed bare.
Per-type validation:
- solid: color must be a non-empty string
- linear_gradient / radial_gradient: stops must be non-empty array
- image: src must be a non-empty string
- any other type: treated as unfilled (renderer can't paint it)
Two new tests: malformed solids (missing/empty color, unknown type) and
empty gradient + image-with-empty-src — all properly patched. Existing
preservation tests (real solid, linear_gradient with stops, radial
with stops, image with src) still pass.
Previous `hasSolidFill` only matched `type === 'solid'`. Sub-agents
legitimately put `linear_gradient` (sunrise hero, accent ribbon),
`radial_gradient` (splash entries), or `image` (branded photo banners)
on top app bars and other nav surfaces, and `hasSolidFill` would
return false for those — making the inject pass overwrite the
gradient/image with a flat `$color-surface` solid.
Renamed to `hasAnyFill`; matches any first-entry shape with a
recognized `type` field. Sub-agent intent (any non-empty fill) now
short-circuits the inject. Three new tests cover linear gradient,
radial gradient, and image fills explicitly — all preserved.
The previous "navbar in PROTECTED_ROLES" change was Codex-flagged as a
no-op: PROTECTED_ROLES only PREVENTS strip-pass deletion of an existing
fill, it doesn't ADD one. The actual food-app brief failure was that the
sub-agent emitted a bottom navigation row WITHOUT any fill at all,
relying on the parent surface for visual contrast — but the parent (the
cream root frame) doesn't supply that contrast, so the nav blends
straight into the cream background and visually disappears.
New deterministic pass: `injectMissingNavSurfaceFill`. For each direct
child of the page root whose role is one of {navbar, nav, tab-bar,
bottom-tab-bar, top-nav-bar, top-app-bar, tab-row} AND whose fill is
empty/missing, set `fill = [{type: solid, color: '$color-surface'}]`
so the renderer resolves it through the seeded palette and the nav
gets a visible white surface separation from the cream root.
Scope contract:
- Only direct children of the passed root frame (page root). Nav frames
nested inside cards / sections / banners are left alone.
- Never overrides an existing fill — sub-agent intent (e.g. an
intentionally dark `top-app-bar`) is preserved.
- Pure mutation; returns `true` when any nav was patched.
Wired into the same hook point as `stripRedundantSectionFills` (via
`design-canvas-ops.ts::generationCleanup`), so every generation cycle
sees both a strip pass (remove hedge fills) and an inject pass (add
the missing nav surface). Five new tests cover all nav role variants,
preservation of existing fills, scope (no recurse into cards), and
no-op on unrelated roles.
Two related issues from the GPT-5.5 food-app run:
1. Bottom navigation rendered without its surface fill, blending into
the cream root background. The strip-redundant-section-fills pass
didn't have any of the navigation roles (`navbar`, `nav`, `tab-bar`,
`bottom-tab-bar`, `top-nav-bar`) in PROTECTED_ROLES, so a navbar
carrying `fill: #FFFFFF` (or any SAFE_LIGHT tint) hit the
"safe-light hedge" branch and got stripped. Real-world navs
intentionally use a white surface to separate from a tinted root —
that fill is intended, not a hedge.
Fix: add the five navigation role names to PROTECTED_ROLES. New
test asserts a `role: navbar` frame with `fill: #FFFFFF` on a
`#FFF8F0` cream root keeps its fill.
2. Empty-src image placeholders inserted by the dispatcher's JSONL
fallback only got auto-filled at the orchestrator's tail (line
~1219, after every subtask completes). On a long brief that's a
visible lag; on an aborted/throwing brief the tail never runs and
images stay placeholder forever.
Fire-and-forget `scanAndFillImages(parentId)` from the dispatcher's
applied path so each subtask's image set starts searching as soon
as it lands. The orchestrator-tail scan still runs and dedups
through `queuedNodeIds`, so this is purely a latency / robustness
improvement (no double fetch).
`store.addNode(null, …)` routes the insert through `_children()` →
`getActivePageChildren(doc, activePageId)` — meaning the parent list
is the ACTIVE PAGE's children, not `doc.children`. The previous
append-index calc read `doc.children?.length` directly, which only
holds the legacy single-page fallback array. On a multi-page doc the
two diverge: `doc.children` may be empty or stale while the active
page already has N siblings, so the computed append index doesn't
correspond to the actual insertion target — landing either before
existing siblings (off-by-N) or out of bounds.
Use `getActivePageChildren(document, activePageId)` to read the same
list `addNode` writes into. Sub-agent generation runs on whichever
page the user has active, so this matches dispatch behavior exactly.
The non-null parent path (`getNodeById(parentId)` then read its
children length) was already correct — only the null-parent branch
needed fixing.
`store.addNode` defaults to `index: 0` (prepend) — appropriate for new
shapes a user draws on canvas (topmost in z-order), but wrong for
sub-agent generation where each subtask emits a section that should
appear AFTER the previous subtask's output in document order.
The previous JSONL fallback called `addNode(parentId, root)` without
an explicit index, so every subtask's output prepended to the previous
ones. Result: the last subtask landed first in the root frame's children
and the first subtask was pushed to the bottom. Real-world repro:
a food-app brief with sections [status-bar, header, search, categories,
banner, popular, recommended, bottom-nav] produced [recommended,
popular, banner, header, search, status-bar, categories, "what are
you craving", bottom-nav] — same nodes, reversed order.
Compute the parent's current `children.length` for each insert and
pass it as the explicit index so roots append at the end of the
parent. Order matches subtask iteration order, matches the brief.
The default-prepend behavior of `addNode` is unchanged for other
callers (drawing tools, paste, etc) — only this dispatcher path
overrides it.
parseJsonlToTree resolves `_parent` via a `Map<id, node>` that's
overwritten on duplicate ids. A JSONL line whose `_parent` equals its
own `id` (or otherwise references a node that ends up being itself
after the map overwrite) produces `node.children = [node]` — a cyclic
graph. Without a cycle guard, `collectIds` would recurse forever
walking node → node.children[0] → node → … and never surface the
duplicate or fail the dispatch.
Two-part guard:
1. Reference-identity `WeakSet` (visitedRefs) — short-circuits the
recursion the moment we re-enter the same node object.
2. Early return after pushing a duplicate id — once we've recorded
the dup, descending into its (potentially cyclic) subtree adds no
information.
Either guard alone would prevent the stack overflow; together they
make the dup-detection robust against any malformed input shape that
parseJsonlToTree might produce.
Prior precheck only compared the JSONL payload's ids against the live
doc — it didn't detect duplicates within the payload itself. A model
emitting the same id twice (two roots sharing an id, a nested child
reusing a root id, etc) would still slip past: addNode appends both
copies, getNodeById returns the first match for both verifies, and a
later rollback removeNode would delete only one of the two duplicates,
leaving an orphan with the same id in the doc.
Detect internal duplicates during the same id-collection pass:
collectIds tracks `treeIds` as a Set and pushes any id seen twice into
`internalDuplicates`. If non-empty, return `failed` immediately with
the duplicate list — same shape as the existing live-doc collision
branch, no doc mutation, no side effects, retry path takes over.
`insertNodeInTree` does not dedupe — it appends. So if the model emits
a JSONL root whose id collides with a pre-existing live node, the doc
ends up with two nodes sharing that id. `getNodeById(root.id)` returns
the FIRST match (the pre-existing one), making the post-insert verify
look successful even though the new node was appended elsewhere. A
later rollback `removeNode(id)` then deletes the PRE-EXISTING node
instead of the duplicate, corrupting the doc.
Precheck: collect every id the JSONL tree introduces (roots and
descendants). If ANY of them already exists in the live doc, refuse to
insert and return `failed` immediately — doc state is preserved, no
rollback needed, the orchestrator's retry path takes over.
Once we know all ids are fresh, the existing post-insert verify and
rollback paths are safe: every id we touch was provably absent before
the dispatch, so `removeNode(id)` targets only what we just added.
Previous "failed with non-empty insertedNodes" combination still bypassed
retry. orchestrator-sub-agent.ts gates retry on `result.nodes.length === 0`
— the partial inserts surfaced through DispatchResult.insertedNodes
flowed through to the subtask's `nodes` field, made it look non-empty,
and skipped the retry / minimal-skills / batch_design fallback chain.
Hard-rollback partial inserts on JSONL fallback failure: call
`store.removeNode(id)` for every root that did land, then return
`failed` with `insertedNodes: []`. The dispatcher's surrounding
history-batch wrapper absorbs both the addNode and removeNode calls so
the user-visible undo entry is a net no-op, and the retry condition
upstream now sees a genuinely empty result and re-runs the subtask
cleanly.
Three outcomes after this:
- All roots land → `applied` with full insertedNodes.
- Partial / total failure → `failed` with `insertedNodes: []` (any
partial successes rolled back) so retry fires and the doc returns to
its pre-dispatch state.
Previous version reported `status: 'applied'` whenever at least one root
landed, with a partial-failure note in `message`. But the orchestrator's
retry / minimal-skills / batch_design-fallback chain checks
`status === 'applied'` to decide whether to bypass retry — a partial
insert (e.g. 1/5 roots landed because `defaultParentId` was stale)
would short-circuit retry and leave the user with a degraded design
that the system never tried to fix.
Now any failed root flips the dispatch to `status: 'failed'` so the
orchestrator's retry path can take over. The successful partial inserts
are still surfaced in `insertedNodes` so the surrounding history-batch
wrapper can roll them back / clean up — `failed` with non-empty
`insertedNodes` is a legitimate combination meaning "side effects
happened but the dispatch did not complete its contract".
Three outcomes now:
- All N roots land → `applied` with full count.
- 1..N-1 land → `failed` with partial-success `insertedNodes` and a
message naming the parent id + failed root ids.
- 0 land → `failed` with empty `insertedNodes` and the same diagnostic.
Previous JSONL fallback called `store.addNode(defaultParentId, root)`
and reported `status: 'applied'` regardless of outcome. But `addNode`
returns void and silently no-ops via `insertNodeInTree` when the parent
id can't be resolved (stale `defaultParentId`, empty doc, etc). The
caller would then count the dispatch as a successful insert even
though the doc was unchanged.
Verify each root via `getNodeById(root.id)` immediately after addNode.
Outcomes now:
- All roots land → `applied`, message lists count.
- Some land, some don't → `applied` with partial-failure note in
message; only the live roots are returned in `insertedNodes`.
- No roots land → `failed` with diagnostic naming the parent id and the
first few failed root ids — caller surfaces this to the orchestrator
retry path instead of silently absorbing the loss.
Mid-tier models (observed: GPT-5.5 standard tier in web-app CLI mode)
correctly emit `<op_tool>{name:"batch_design",arguments:{operations:...}}`
when the brief doesn't fit any embedded element tool — Strategy B in
ELEMENT_TOOL_OUTPUT_FORMAT. But the prompt only declares the operations
value as `<DSL_STRING>` without showing the DSL syntax, so models stuff
flat JSONL (`{"_parent":null,"id":"…","type":"frame",…}`) into the
`operations` field instead of `foo=I("parent",{…})\nbar=U(foo,…)`.
The browser DSL executor then rejects every line ("Cannot parse
operation: …"), all retries fail, and the user sees a degenerate result
(303B / 1 node) despite the model having streamed a full design.
Detect at dispatch time: if `operations` looks like JSONL (starts with
`{` AND contains a `_parent` key or a typed PenNode shape near the top),
route through `parseJsonlToTree` + `store.addNode(defaultParentId, root)`
loop instead of the DSL parser. Same dispatch invariants (single
history batch, dispatch result accounting) apply.
This unblocks the most common Strategy B failure: model emits JSONL
inside a `batch_design` envelope. Strategy A (per-component element
tools) and DSL-shaped Strategy B both still go through their existing
paths unchanged.
The prior pass treated ANY atomic-role frame containing another atomic-
role child with a fill as a wrapper. That's still too aggressive: real
atomic components legitimately compose secondary atomics inside them
(input + trailing icon-button for clear/reveal-password, search-bar +
voice-search icon-button, etc). Stripping the parent's fill in those
cases erases the input/search-bar surface — a regression.
Refine: split atomic protected roles into PRIMARY (input, form-input,
search-bar — input-class components that constitute the "main" atom)
and SECONDARY (button, icon-button, badge, chip, tag, pill — sub-action
or decoration atomics that legitimately nest inside primary atomics).
Wrapper detection now triggers only when:
- same-role nesting (search-bar > search-bar, input > input), OR
- PRIMARY atomic nested inside another atomic (search-bar > input —
the canonical sub-agent misroll).
Two new tests:
- input atom with trailing icon-button (filled clear button) → input
fill kept
- search-bar atom with voice icon-button (filled accent) → search-bar
fill kept
Original misroll case (search-bar wrapper > inner input) still strips —
covered by prior test.
Previous nested-wrapper detection treated any PROTECTED_ROLES frame
containing another protected/structural-fill child as a wrapper. That
swept too widely and could strip fills from real container components:
- card containing a CTA `button` (button is filled, card surface is
intentional) — card fill stripped if its surface was in SAFE_LIGHT.
- pricing-card with a `badge` ribbon and a CTA button — same issue.
- banner with a nested card — banner fill stripped.
Real component composition is normal; the problem is specifically
sub-agent role mislabels where an ATOMIC component (search-bar, button,
input, badge, chip) is reused as a section wrapper. Container roles
(card, pricing-card, feature-card, banner, etc) NEVER appear as
wrappers — their fill is always intentional.
Fix: introduce ATOMIC_PROTECTED_ROLES (subset of PROTECTED_ROLES) and
restrict wrapper detection to firing only when the OUTER role is in this
atomic set. Container roles stay fully protected.
Three new tests added:
- card with filled button child → card fill kept
- pricing-card with badge + button children → pricing-card fill kept
- banner with nested filled card → banner fill kept
The original misroll case (search-bar > input wrapper) still strips —
covered by the prior test.
Real repro from MiniMax-M2.7: sub-agent emits a section wrapper with
the WRONG role applied — Search Bar(role=search-bar) > Search Input
Container(role=input,fill=$color-surface). The outer "search-bar" frame
is actually a section-level wrapper (its child carries the real atom),
but its role is `search-bar` which is in PROTECTED_ROLES, so the strip
pass treated it as the real atom and left its #F8FAFC hedge fill alone.
Result: visible double-cream nesting against the cream root background.
Detect this misroll: a frame whose role IS protected but ALSO contains
a child carrying either the same role or another protected/structural
role with its own solid fill is a wrapper, not the atom — its fill is
eligible for the same safe-light/safe-dark hedge stripping that pure
section frames get.
Counter-case kept covered: a real `search-bar` atom whose children are
just icons / placeholder text (no nested input/search-bar/card/etc with
its own fill) keeps its fill — that fill is intentional, not a hedge.
Two new tests:
- M2.7 misrolled wrapper (search-bar > input + safe-light fill) — outer
fill stripped, inner input fill preserved.
- Real search-bar atom (no fill-bearing component children) — fill
preserved.
Codex flagged: even after the previous CRITICAL preamble told the model
to defer to `<op_tool>` mode when an OUTPUT FORMAT block exists later,
the rest of jsonl-format / jsonl-format-simplified still contained
specific JSONL-output directives ("Output a ```json block with ONE node
per line", "FORMAT: _parent (null=root, …)", a full ```json example).
Those specific instructions can dominate over the abstract preamble for
weak models — they read concrete rules and execute them, ignoring the
top-of-skill conditional.
Restructured both skills so JSONL-specific output mechanics are scoped
to a clearly-marked "JSONL FALLBACK MODE" section and the schema
content (TYPES / RULES / DESIGN SYSTEM TOKENS) is mode-agnostic.
- New top-of-skill comment explicitly states TYPES / RULES / TOKENS
apply to BOTH `<op_tool>` argument shape AND JSONL — neither mode
contradicts them.
- The "Output ```json block" directive, the "FORMAT: _parent" directive,
and the ```json example are now wrapped under a "JSONL FALLBACK MODE"
header that explicitly says "ignore this section if `<op_tool>` mode
is in effect".
In `<op_tool>` mode the model now reads schema rules without reading
JSONL-specific output mechanics; in JSONL fallback the JSONL section is
unambiguously authoritative. No conflicting instructions for either
output path.
The previous fix dropped jsonl-format / jsonl-format-simplified entirely
when elementToolsEnabled was true, on the theory that their CRITICAL
"Output ONLY ```json … Do NOT use tool calls" line conflicted with the
appended `<op_tool>` instruction. But empirically dropping them made
weak-model output WORSE: MiniMax-M2.7 still emits raw JSONL most of the
time (it can't reliably emit `<op_tool>`), and without the JSONL
schema/format teaching its output degrades — role coverage dropped
from 74% to 22%, color-ref% from 84% to 49%.
The right fix is dual-mode coexistence: keep BOTH skills loaded so the
model has the JSONL fallback teaching, but rewrite each skill's CRITICAL
opener to defer to the ELEMENT_TOOL_OUTPUT_FORMAT block when present.
- jsonl-format / jsonl-format-simplified now lead with: "If a separate
OUTPUT FORMAT — EMIT AS TOOL CALL(S) block appears later in the system
prompt, FOLLOW THAT block. Use the JSONL form below ONLY when no
<op_tool> instruction is present."
- Removed the orchestrator-sub-agent.ts skill-filtering branch; both
skills load unconditionally now.
Net effect: strong models that can follow `<op_tool>` will use the
element-tool path (preserving the n-tools-per-element design intent for
weak-model stability — MiniMax/GLM/Kimi will emit `<op_tool>` when they
can). Weak models that fall back to raw JSONL still get the schema /
sizing / fill / token rules they need to produce coherent output. No
forced choice, no degraded fallback.
- deny.toml: add [graph].targets to limit metadata to native+wasm32
(avoid Android/iOS edition-2024 deps that fail rustc 1.82 cargo metadata)
- deny.toml: [bans] allow-wildcard-paths = true for workspace path deps
- crates/*/Cargo.toml: add explicit version="0.1.0" alongside path = "..."
(cargo-deny rejects wildcard-path deps for publishable crates)
cargo-deny 0.16.4 hits a CVSS 4.0 parse error AND lacks edition-2024 cargo
metadata support; bumped to 0.18.9 (installed via stable toolchain). Run
cargo-deny with RUSTUP_TOOLCHAIN=stable so it uses cargo 1.95 for metadata
parsing while project itself still builds on 1.82.
Verified: advisories ok, bans ok, licenses ok, sources ok (exit 0)
on both native and wasm32-unknown-unknown targets.
Phase 1 batch 3 implementer found 1.80 incompatible with current
crates.io ecosystem: parley → fontique → litemap 0.7.5 needs 1.81;
accesskit chain → indexmap 2.14 → hashbrown 0.17 needs edition2024
(1.85); skia-safe 0.75+ → home 0.5.12 needs 1.88. 1.82 is the sweet
spot that fixes litemap (and matches what Task 0.4 actually probed
with — 1.95).
shell-native dep set deviation (winit only, skia-safe + accesskit
deferred to Step 1 kill-spike when actually used) is documented in
the Phase 1 review trail. compile_error guard for wasm32 still fires
correctly — the load-bearing §1.2 invariant is satisfied.
Phase 1 skeleton: declare crate, add compile_error! wasm32 guard so accidental
inclusion in the web bundle fails at compile time (kickoff spec §1.2 invariant).
Native deps intentionally minimal (just winit, no default features). skia-safe /
accesskit / accesskit_winit deferred to Stage F when RenderBackend is actually
implemented. Reason: current top-tier versions of these crates pull transitive
deps (home 0.5.12, litemap 0.7.5, hashbrown 0.17) that require Rust 1.81+ /
edition2024, but our pinned toolchain is 1.80. Pinning to spec versions
(skia-safe=0.74) also fails since 0.74 was never published. Will revisit when
either the toolchain bumps or upstream stabilizes around an MSRV-1.80 line.
Verified:
- cargo build -p openpencil-shell-native PASS
- cargo test -p openpencil-shell-native PASS (1 test)
- cargo check --target wasm32-unknown-unknown -p openpencil-shell-native
fails with the compile_error! guard text (NOT a winit/skia build error).