Retire the flat-JSONL retry rung. Every subagent rung now emits a JS
program (script-gen); reduced_complexity and minimal_skills only
narrow the loaded skill set — they no longer switch the output format
to positional _parent JSONL, whose omittable parent field collapsed a
whole tree into flat siblings when a model skipped it. script-gen's
I(parent, node) makes parenting a positional argument that cannot be
dropped, and the reasoning-harvest fix made it robust across models.
parse_nodes stays for the modify/chat paths that still consume flat
node JSON; the jsonl-format generation skills are removed.
A rectangle is a container in the canonical schema — it carries
clipContent like Frame/Group and models nest content inside one (an
image-area rectangle wrapping a photo, a card body, a badge holder).
The painter's NodeKind::Rect branch drew the rectangle's own fill and
returned without recursing, so every child of a rectangle vanished
behind that fill. Measured: an AI-generated travel page whose seven
destination photos each sat inside an image-area rectangle rendered
as blank cards despite the photos being fetched and embedded. Recurse
into the children (honouring clipContent) after the rectangle's own
paint, mirroring the Frame branch.
A component-swapped instance (overriddenSymbolID present) carries
derivedSymbolData for the swapped-in component, not its base symbolID.
Pooling or geometry-seeding it under the base component's cache pinned
wrong pk→node mappings that poisoned genuine base-component instances
reusing that cache. Skip swapped instances in both seeding passes.
A nested instance swapped via overriddenSymbolID keeps the pre-swap
component's derivedSymbolData alongside the swapped-in component's. When
the two frames are the same size the fingerprint can't tell them apart
and the stale (earlier-listed) cluster hijacks the mapping, sizing the
swapped frame wrong and clipping its icon. Cluster the derived entries
by localID and keep only the one that geometrically fits the swapped
subtree. Adds a near-exact geometric-match bonus so a near-perfect size
match outweighs the walk-order prior a stale sibling would otherwise
win on. Splits instance.rs (swap_filter.rs) and fingerprint_tests.rs
(foreign_tests.rs) to honor the 800-line cap.
Renders the Test.fig Sales-card pie icon and the other swapped card
icons at their correct size instead of clipped.
Replace the walk-order virtual-GUID guessing in Strategy 2 with an
evidence-based fingerprint (size / transform / text-class / fill hints)
plus cross-instance pooled + geometry pre-seeding, so nested instance
overrides land on the right nodes. Adds foreign-session subtree
anchoring with a uniform-family fallback, nested-derived field merge,
strong fill-hint routing (image / rare-solid) with conflict-vs-
inapplicable rejection handling, and a single-axis transform-drift
score tier. Splits instance.rs walk/apply helpers into apply.rs and the
test module into fingerprint_tests.rs to honor the 800-line cap.
Renders Test.fig order-row thumbnails, breadcrumb, sidebar logo,
summary-card filters, and status chips to match the Figma source.
A section-heavy new-design prompt ('Design a … page. Include a search
section …') trips is_section_add_request, so requests_new_whole_screen
returns false; the selection-modify bias then routed the whole prompt
into run_modify_turn, where M3's flat-JSONL output was renest-dropped
to nothing ('Could not parse design nodes'). Add a section-add-blind
creation-signal veto to both routing gates so a new-design request
reaches the design pipeline regardless of an active selection.
The script runner only stripped a code fence anchored at position
zero, and never stripped <think> reasoning at all. A reasoning model
that keeps its thinking (MiniMax-M3 rides Adaptive) prefixes the
program with a <think> block full of draft JS plus a prose lead-in;
that went to QuickJS verbatim as source, threw a syntax error, and
dropped the model onto the fragile flat-JSONL retry rung where
omitted _parent fields collapse the whole tree into flat siblings
under the root (measured: a full travel page, 44 nodes, all piled at
origin). Strip reasoning first, then extract the fenced block from
anywhere in the response. Models with thinking disabled (GLM) were
unaffected, which is why this read as GLM-only handling.
Reasoning models burn the whole output budget inside think blocks and
emit zero design nodes on a modify turn (measured on MiniMax-M3: the
turn died in analysis prose). Same policy the design subtasks already
use; the HTTP layer maps it per provider.
The A/B measured the loop ahead of the single-shot orchestrator on
audit cleanliness (11 vs 33 issues over 10 prompt pairs) at
comparable wall time, with the artboard seed guard closing its one
failure mode. OPENPENCIL_DESIGN_AGENT_LOOP=0|false|off opts back into
the orchestrator; CLI providers keep their existing path.
A delayed write from a dead editing session re-added an assertion
encoding the pre-guard behavior (dashboard 'web app' treated as a
mobile ask); the negative case is speced by the sibling test.
The guard landed in the previous commit's tests but its function body
got clobbered by a stale editing session between validation and
commit; a dashboard 'web app' seeded 390x844 again.
Weak models sometimes ignore the seed-the-artboard-first instruction
and leave the root frame sizeless, collapsing the whole screen into a
thin strip (measured on a travel-app A/B cell). The executor now
seeds missing root axes deterministically on the first applied batch
— 390x844 for mobile asks, 1440x900 otherwise (a dashboard 'web app'
is not a mobile ask) — tells the model about the contract in the tool
result, and leaves authored numeric axes alone.
The 0.56em/char estimate under-measures wider platform font stacks;
on Linux the centered label wrapped inside its authored width and
sank 6px below the ring center. Give the forced width single-line
headroom and let the sentinel tolerate single-line metric variance.
Same two shifts as the native mirror: the palette slot removal moved
the right cluster 28px, and the input-width wrap basis fix lifts the
band when the probe text wraps.
The palette slot removal shifted the right cluster 28px right, and
the input-height width basis fix (measuring wrap against the real
inner width) made the old probe text wrap and lift the footer band;
probe with a single-line text at the new offsets.
An active canvas selection now biases intent routing to modify (the
selection IS the target — only an explicit new-whole-screen request
or a plain chat question escapes it), restoring select-then-ask
editing. The panel paints a selected-count chip above the input with
a clear affordance, so the armed state is visible. Chat-question CJK
keywords keep questions about a selection conversational, and the
footer hover math reserves the chip row like paint does.
A sidebar subtask can come back as a landing-page archetype — a
horizontal nav row (brand + links + actions) plus a display-size hero
headline — inside a 260px rail, where everything overlaps and clips.
When geometry proves the squeeze (fixed rail <= 320, link row wider
than its slot), restack the row vertically, clamp display text to
rail scale, relax oversized rail padding, and emit a diagnostic for
the agent loop. Planning and dashboard skills state the rule
outright: a sidebar is a vertical rail, never a navbar or hero.
56x28 left the centered label ~3px from each edge; 76x30 matches the
reference chrome. Long install-guidance lines may now ellipsize at
the tail — the actionable prefix still paints.
The + button pushed a defaulted ChatState whose empty discovered and
available model lists painted 'No models connected' until the next
provider probe. Model discovery is app-level state — carry it (and
the selected index) into new tabs and the last-tab in-place reset.
The section-granularity dwell comparison was reversed: the page root
reveals first at scaffold time, so 'ancestor earlier than child' held
for every descendant and all waypoints were suppressed — the cursor
sat on the root center for the whole run. Suppress a child only when
it pops in within the dwell window AFTER its nearest revealed
ancestor, and propagate that nearest ancestor instead of the
min-start (which pinned every window to the root).
The loop reveal gate still matched the removed emit_elements constant
after rebasing onto its retirement; batch_design is the only
canvas-writing loop tool left.
Every revealed node used to become a cursor waypoint, so dense batch
reveals sent the pointer darting between leaves. Children of a
just-revealed ancestor no longer schedule waypoints (the cursor parks
on the section while its contents pop in), tiny standalone leaves are
not chased, and stops closer than a dwell window collapse to the last
placement so each move stays readable.
Document-content commands allocate a revision from a monotonic
counter (never revision+1: after save, undo, then a new edit that
would collide with the saved value and report a clean file over
divergent content); undo back to the save point turns clean again.
Selection, viewport and panel state do not count. The top bar paints
a muted edited label next to the file name; load/save/save-as reset
the baseline.
CLI subprocesses fell back to the macOS system proxy, whose
stream-reconnect path some CLIs reject mid-turn; pass through the
HTTP(S)/NO_PROXY vars users actually validate in their shell and keep
socks-scheme all_proxy out.
Loop-applied batches now register node reveals against the live run
epoch (no more whole-sections popping in silently), the checklist
surfaces tool-call progress per turn, per-batch layout feedback keeps
flowing, and the design loop gets a bounded turn budget sized for
section-by-section batching plus repair turns.
A re-bound binding now updates in place instead of duplicating its
result row, and bindless container lines get monotonic auto bindings
so later lines can still parent onto them.
Dashboard/mobile constants measured from strong references (row
rhythm, KPI card anatomy, chip specs, badge content rules); the agent
phase seeds the artboard full-size in its first batch, keeps batches
small, fixes reported layout issues before continuing, and moves
misplaced sections instead of rebuilding them; planning appends the
bottom-nav subtask for app home screens.
A transport-config rejection repeats identically on every attempt;
burning the retry ladder and salvage pass on it costs minutes and
ends the same way.
finalize's ReplaceSubtree allocates ids the reveal overlay never saw,
so a mid-animation section snapped in at once and the agent cursor
lost its target. Wait (worker thread, abort-aware, capped) for the
scheduled sweep to finish first.
A container filled with a $color-text-* token is a slot-category
error (a search pill filled text-primary rendered as a white capsule
on a dark theme); remap to the surface token family.
New finalize passes measured on real generations: radial stacks
(donut/ring segments authored as flow siblings) recentre under a
layout:none parent; empty decorated stub frames are removed unless
positioned (a pinned overlay is intent, not an abandoned shell);
corner badges and notification dots adopt into their image/icon
anchors; pill-chip rows clip instead of flexifying; card overflow
gains deep-descendant clip; row-gap, table-scale and overfull-row
repairs get sturdier floors so they cannot compound a design into
mush. Diagnostics for starved fills, over-packed rows and vertical
overlap feed the agent feedback loop.
An authored layout:none whose children carry real x/y is deliberate
overlay placement; the horizontal-layout inference pass flattened it
(a pinned badge dot snapped into flow). Legacy frames with absent
layout and stale 0,0 children still infer.
Bump jian for the shared font resolver, path/line flex sizing, and
input-chrome measurement; route the native backend's draw_text and
measure_text_weighted through the resolver so fit_content boxes match
what actually paints (synthetic bold and multi-line heights included).
MergeAppState is additive by design — doc-owned keys win and lower
plan_idx wins, so "nothing to add" is the designed outcome, not a
failure. Returning false made Batch[merge, insert] reject the entire
generated insert whenever the target document already carried every
declared key (regenerating a section over an opened file), and made
batch_program report a line failed after its insert had already
landed. The return now signals "command processed"; regression tests
cover the batch-survival and pre-seeded-doc cases.
The fixture is a live file on the developer Desktop (v2.8, loaded
best-effort) and currently trips a REAL main-axis overflow (tab section
bottom 2456 > root bottom 2418) in the jian growth path being reworked
around jian 57068a6 — permanently red on this machine through no fault
of the loader. Keep the net, run it on demand with --ignored.