Corpus-run refinements, each anchored to a measured case:
- a JAM participant must be a text-bearing CONTAINER cell — two bare
text siblings set tight on purpose (price + unit) are typography;
- the row-gap repair also covers a TWO-column jam (date column against
a details stack at 0px) when the row top-packs;
- text overflow now triggers on the right EDGE, catching the combined
overflow a width-only check misses (avatar + name pair);
- one sibling centered inside another is an intentional overlay (a
number set on a ring), not an overlap accident;
- drop an unreachable duplicate match arm the merge left in the lint
contrast detector.
Corpus re-audit: 51/52 clean as generated, 52/52 after finalize.
The desktop LLM adapter capped design turns at 8192 output tokens; a
rich plan truncated mid-JSON, the parse failed, and the heuristic
fallback shipped a skeleton (hero + three cards) with no visible error.
Raise the budget to the headless-harness value that ran a 52-prompt
corpus with zero truncations (16384; 24576 for M3 which spends budget
on reasoning), and give planning a second attempt before the fallback —
a truncated stream or transient blip usually parses fine on retry.
Self-loop run v1 findings, fixed at the root:
- geometry row-gap repair: a horizontal row whose >=3 text-bearing frame
cells ALL resolve jammed gets a gap injected — name-blind, so rows
buried under any depth of unnamed wrappers are covered (the name-gated
table pass also now recurses through wrapper chains as a fast path).
- rigid fit_content children overflowing a narrow flex parent are
retargeted to fill_container (fit never shrinks; an icon+text pair
painted over its siblings inside an 80px card); the text inside then
wraps via the text-overflow fixer on the next loop round.
- prompt: one brand name reused verbatim across logo/footer/sample data.
Offline replay of both v1 artifacts: 12 audit issues -> 0.
Conflicts were the two mesh/shader implementations meeting: kept the
remote's newer complete version (typed shader uniforms, shader color
uniform binding, mesh vertex editing defaults, status-bar shell stroke
handling); deduped two identically-replayed RenderBackend methods.
The backend + dispatch wiring landed, but the real editor scene
builder (editor_state_to_layout_scene → pen_document_to_payload)
still collapsed PenFill::MeshGradient / PenFill::Shader at the
extraction layer — fill_type mapped to "solid", the gradient payload
only knew linear/radial, and SceneNode.shader was hardwired None —
so the new painters were unreachable from an actual document. Extend
the payload chain end to end: GradientPayload grows a Mesh variant
(row-major lattice, sparse vertices stay transparent) and NodePayload
grows a resolved ShaderPayload (uniforms pre-expanded, colour hex →
premultiplied vec4, fallback = first colour uniform else mid-gray),
both mirroring jian-core's canonical try_mesh_gradient / try_shader
rules; fill_type maps "mesh"/"shader" into the scene enum; node
opacity folds into the shader alpha like the gradient path. The flat
`fill` now matches the shader fallback so painters without SkSL (web)
show the same colour. Reachability is pinned by scene-threading tests
driving the real `.op` → EditorState → scene chain.
The mesh/shader migration only papered over the compile errors: the
canvas dispatched Mesh bodies to a flat first-vertex fill and never
dispatched SceneNode.shader at all, even though the jian Painter trait
already carries fill_round_rect_mesh_gradient / fill_round_rect_shader
with real Skia implementations available. Port jian-skia's Gouraud
vertices + cached-RuntimeEffect paint onto NativeBackend (sibling
mesh_shader.rs, ShaderCache shared with jian-skia so sources compile
once), delegate through NativeFrameBackend, and route the canvas fill
dispatch through the trait methods — shader wins over gradient over
solid. CanvasKit/web inherits the trait's first-vertex / fallback-solid
defaults (documented parity gap, same as jian's own non-native
painters). Pixel-level raster tests prove interpolation actually
happens and a compiled shader beats its fallback.
Two web-parity gaps from the audit:
- AI/MCP writes were not undoable in the browser: the live-sync pull
replaced the document through replace_document, which resets History.
EditorState grows replace_document_with_undo (same stale-draft
clearing, pre-replace snapshot pushed as one undo step); the glue
uses it for every apply after the mount-time first pull, gated on
WebSyncClient::initialized().
- Agent-team indicators never showed on web: design runs execute in
the serve-web daemon, whose process-global registry the browser wasm
can't see. The registry gains a JSON relay (relay_json /
parse_relay_json / apply_remote — reveal timestamps insert-only so
locally rebased clocks survive re-polls), the daemon serves it at
GET /api/mcp/indicators, and the web host polls on the live-sync
cadence, mirrors into its local registry, and drives a
self-terminating rAF pump so reveal/breathing animations play
between DOM events.
The image Search/Generate buttons, the six MCP CLI-integration
toggles, and the auto-update card all painted on web with no executor
behind them — clicking did nothing (silent dead UI, worse than a
missing feature). Gate each behind a cfg const following the existing
GIT_BUTTON_AVAILABLE pattern, with paint, hit-test, and layout heights
kept in sync; the experimental-features card moves up into the hidden
auto-update slot.
Four Windows runtime defects from the platform audit:
- The binary stayed in the console subsystem, parking a console window
behind the GUI when launched from Explorer. Release builds now set
windows_subsystem = "windows"; debug keeps stderr tracing visible.
- Background CLI probes (model discovery, provider version checks) and
the vendored Claude SDK's per-turn spawns lacked CREATE_NO_WINDOW,
flashing console windows once the GUI detaches from the console.
- MCP stdio servers naming .cmd/.bat shims (npx and most npm-installed
servers) could not spawn: CreateProcess cannot execute shims and Rust
1.77+ refuses them as program names. vendor/agent now resolves the
command PATHEXT-style against the PATH the server will actually see
(per-server env override wins) and routes only genuine shims through
cmd /c — real executables keep direct spawn semantics.
- cmd /C start truncated URLs at `&` (every OAuth authorize URL). The
URL now travels double-quoted via raw_arg so cmd keeps it literal.
Two Linux startup blockers:
1. Skia's Interface::new_native() dlopens libGL/GLX, which fails on
the EGL contexts this project always creates on Linux — GL init
errored and the app exited (the tracked LINUX_GPU_SKIA_LOADER_TBD
gap). GlContextProvider now exposes gl_proc_address; the glutin
provider resolves through its display (eglGetProcAddress on EGL)
and SharedSkiaContext builds the interface via new_load_with,
falling back to new_native for providers without a loader. The
Linux GPU smoke tests are un-ignored (they soft-skip without a
working EGL stack unless STEP1A_REQUIRE_GPU=1).
2. Dropping casement's default features also dropped wayland-dlopen,
making libwayland-client a DT_NEEDED hard dependency — the binary
could not even load on X11-only hosts. Restore the feature so
libwayland loads at runtime when present.
vendor/jian ff39a41 added PenFill::MeshGradient / PenFill::Shader plus
SceneGradient::Mesh and SceneNode.shader; every exhaustive match over
those types stopped compiling. Cover them with faithful conservative
semantics: opacity getters/setters treat the new bodies like their
gradient siblings, variable binding walks mesh vertex stops (shader
uniforms are untyped — left alone), canvas paint degrades mesh to its
first vertex colour until the RenderBackend grows a mesh method, the
property panel reports the neutral Solid type (no lossy conversion
offered), and lint contrast checks skip both like image fills.
finish_if_epoch previously routed the empty-queue case through
drain_finished_run, which set the process-global needs_final_frame flag
even though a run that never queued a reveal never put a cursor on
screen. That stray flag made next_reveal_deadline_ms return a redraw
deadline out of an idle registry, and — because the registry is shared
across the whole test binary — perturbed an unrelated exact
animation-deadline assertion whenever a design-session test dropped an
empty session in parallel.
Clear an empty finish inline without arming the flag; the paint-path
drain still arms it after real reveals prune, where a cursor genuinely
was on screen. Adds a regression test.
Verification note: the op-editor-core suite could not be run for this
commit because the shared workspace is transiently non-compiling under a
concurrent mesh-gradient/SkSL-shader jian bump (new PenFill variants not
yet handled in fills.rs — unrelated files). Change is trace-verified and
touches only agent_indicators; re-run pending the tree compiling again.
47 commits from the align branch merged onto the force-updated remote
base (which had replayed an earlier snapshot of the same work plus new
overlay/pointer features and CI fixes). Conflict resolution: kept the
newer align side for the generation pipeline (orchestrator, mcp, skills,
design tools), kept the base side for the chat-panel test semantics and
graceful overlay teardown, fused both in sub_agent_session (design-turn
thinking policy + graceful epoch finish), and dropped the files each
side had deleted (legacy concurrent/dashboard paths, retired TS skills).
Deduped two identical replayed hunks (export.rs, chat_session_tests.rs).
Known issue carried over: provider_probe_host::landed_connected_outcome_
without_models_is_failure fails on a host with a live provider config
(env-sensitive test, both sides byte-identical there; green on CI).
- OPENPENCIL_SMOKE_AUDIT=<file.op>: load the doc, run the real-layout
geometry diagnostics, print a JSON report, exit non-zero on issues —
the machine-checkable leg of the generate→render→audit loop.
- extend the harness thinking-disable gate to glm models (reasoning
burned the whole token budget and returned empty content; an
orchestrator sidebar subtask failed 3x and shipped missing).
- scripts/self-loop.sh: prompts file → generate → render (the real
canvas pipeline) → audit → scorecard.json, fully unattended.
- design-agent: layoutIssues in a tool result are measured facts — fix
them with a follow-up batch before building the next section.
- avatar image searches must include face/headshot (bare 'portrait'
returns torsos and cropped bodies).
- dashboard domain: hard completeness bar — data tables carry >=6 varied
rows and a dashboard ships its full section set.
paint_text_field used to paint an unconditional white fill, a default
grey border and a 6px radius floor, with near-black value text — a model
embedding an input in its own styled wrapper zeroes all of that out
(fill:[], stroke thickness 0, cornerRadius 0) and got a glaring white
pill with a duplicate icon on a dark themed search bar. No authored fill
now paints no box, no authored stroke draws no border, radius is taken
as-is, and the value text color adapts to the authored fill's luminance.
A model groups [toolbar, rows-wrapper] inside the table frame; the
gap-less rows all lived behind that unnamed vertical wrapper, so the
name-gated column-gap pass never saw them and columns rendered touching
('Oct 24, 202442Marcus Thorne'). The table name gate stays on the outer
node; an unnamed structural wrapper now inherits it, for both the gap
injection and the container row-count gate.
A plan's rootFrame.height is often 0 ('compute from content'); the
desktop path let that literal 0 through, so every fill_container
descendant resolved to 0px mid-pipeline and the geometry pass demoted a
healthy fill-height sidebar before adjust_root_height ever assigned the
real number (footer floated mid-page on three consecutive runs). Map
non-positive scaffold root heights to fit_content; the final adjust pass
still writes the definitive numeric height. Golden end-to-end test uses
a height-0 plan.
Also add an end-of-run salvage pass: transient provider failures (network
blips, empty responses) used to burn all 3 back-to-back attempts and the
section shipped missing with no visible signal; every failed subtask now
gets one late retry after the others complete.
Every mutating batch_design / emit_elements result now carries
layoutIssues (real-layout geometry diagnostics) plus an actionable hint,
so the model sees each batch's geometric consequences immediately and
repairs them in-process instead of piling defects up for the loop-end
finalize. Verified end-to-end: the model received a table-overflow
report, recomputed its column widths and fixed them in the next batch.
- geometry_diagnostics: report-mode counterpart of the fix loop (collapse,
table overflow, text/frame overflow, sibling jam with row-cell gating),
powering per-batch layout feedback and the audit gate.
- new fixer: a numeric-width child resolved wider than its flex parent is
retargeted to fill_container (an 800px avatar bar inside a ~550px row
spilled across the design).
- retire the tree-shape circular-height demoter: the layout engine now
resolves a fill-height child of a hugging parent to its content size on
both axes, so the collapse it guessed at no longer exists while its
demotions broke healthy shells; real collapses stay covered by the
geometry loop. Engine-contract sentinel test added.
- theme variable polarity repair: a dark design leaving stock light
values in surface-family slots (chips resolving #F1F5F9 on #0A0A0A)
adopts the correct-polarity value from the variable's other theme slot,
written to the exact slot the resolver reads.
- backdrop sibling merge: an 'active background' authored as a flex
sibling (empty fill×fill rectangle) moves its fill onto the container
instead of eating half the row.
- text-fill injection is now tri-state: gradient/image/unresolvable
backgrounds get NO guessed color (was: mirror-image dark-on-dark).
- env-gated per-pass debug probes + pre-geometry tree dump + repair/
finalize/rect-probe harnesses for offline diagnosis.
A generated program that parents its first line onto a binding that was
never created used to ride into InsertAuthoredSubtree unvalidated; the
host's existence check then rejected the WHOLE otherwise-valid program
and the section was lost to retries. Insert at the document root instead
and surface a warning in the envelope.
Also: drop the ambiguous bare-word 'auto' sizing (schema default wins
instead of inverting intent), and recover word-based textGrowth
spellings (fixed_width_and_height and friends) instead of silently
removing them.
Generation density was inconsistent — some runs shipped a full dashboard,
others stopped at a header + four stat cards + a two-row table. A vision-less
model cannot use the get_screenshot completeness check (it never sees the
render), so it declared done on a coherent-but-sparse skeleton. Add a text-based
completeness bar it CAN follow: tables need >=6 realistic rows, a dashboard needs
its full section set (stats + primary table + a secondary section), and every
list/card carries varied data.
An already-row shell whose sidebar hugs its content (height=fit_content) is
only as tall as its nav, so a space_between / fill_container footer inside it has
no room to sink and floats mid-page (measured: a 260x532 sidebar in a 260x1234
row stranded the profile card 700px above the bottom). ensure_split_shell_is_row
now also promotes the sidebar of an ALREADY-horizontal shell to fill_container
height, not just the freshly-flipped case.
A fill_container block keeps min-size 0 so a fixed sibling gets its space, but
that lets it shrink below its fit_content text — the text then overflows into
the next column (measured: a 260px sidebar schedule row painted the client name
over the appointment time). Add a geometry-driven pass: when a text resolves
wider than its parent block, constrain it (width fill_container + textGrowth
fixed-width) so it wraps inside instead of overflowing.
The design-loop prompt teaches the G(...) image-fill op but never required
using it, so a model built avatar/photo frames and left them as empty colored
squares. Add a hard rule: every avatar, photo, thumbnail, hero, or logo slot
must get an image fill via G(id, "search", "<subject>") in the same batch
that creates the frame.
Add a dashboard domain skill and teach models to emit an imagePrompt alongside
imageSearchQuery so a configured generation model produces a rich image while
stock search stays the fallback.
When a generation profile is configured, generate images from the node-bound
prompt; otherwise search stock placeholders. Both paths keep the prompt/query on
the node so it can be re-generated or re-searched later.
The OpenAI-compat tool-loop body never sent the reasoning-off field, so a GLM/
MiniMax reasoning model spent its whole per-turn budget on hidden reasoning and
truncated the design mid-op. Send thinking:{type:disabled} for those models in
the loop body (matching the single-shot path) and raise the per-turn output
budget for headroom.
Resolve real layout geometry to correct table-column overflow and collapsed
fill-container heights, iterating until stable. Add whole-root passes: sidebar
app-shell reshape + content eviction + split-shell row layout, table column
gap/regroup, footer sink, surface discipline, scaffold padding. The design-loop
finalize backstop runs these once at loop end and injects background-contrasting
fills for fill-less text. run.rs: continue past a failed subtask and compute
zero-content before cleanup.
Add a parent-by-reference program DSL and an executable-script protocol with
best-effort per-line parsing, so a truncated or malformed line drops only that
op instead of the whole design. Route protocol selection in prompt/subagent and
normalize flex keywords in parse. Registers the new modules in lib.
Map Figma/flex auto-layout field names (layoutMode/direction/itemSpacing/
strokeWeight and nested {type,gap,padding} / {Horizontal:{…}} layout objects)
onto the flat schema, recover image src from url/source aliases, normalize
textGrowth/sizing keywords, and quote bareword DSL values — so a single
unfamiliar spelling no longer drops the whole node or op.
Reads a program file, runs it through op_mcp::batch_design_snapshot, applies, saves
the .op (postProcess off = raw structure) — the headless harness for benchmarking
the program-DSL path vs flat JSONL across weak models.
OPENPENCIL_PROGRAM_GEN makes the sub-agent emit a batch_design program
(binding=I(parent,{...})) instead of flat _parent JSONL; nesting by captured
binding makes weak-model row decomposition / header-only tables near-inexpressible.
Runs the existing op_mcp executor to a section forest. Gated to the full first
attempt; drops the conflicting jsonl-format skill so PROGRAM_FORMAT governs alone.
split_operations now groups by operation-start grammar + bracket depth (not a
quote/bracket state machine), so a stray unbalanced quote can no longer swallow
following operations. parse_json_arg gains lenient repairs for fused close-quote+
comma, missing opening quote on a string value, and unclosed trailing brackets;
normalize_node_shape drops an empty stroke (0-length PenStroke) instead of failing
the whole node. +5 regression tests (369 pass).
Weak models emit a table as a header row followed by FLAT row-indexed sibling
cells (R1 Client Cell, R1 Visit, …, R2 Client Cell, …) that render stacked
vertically with full-width status bars. table_repair::regroup_flat_table_rows
groups them into Table→Row→Cell with header-aligned column widths. Detection is
narrow and never guesses (adversarial review caught heuristic-header /
chunk-of-N over-firing on toolbars/feeds): it requires an explicit table header
AND every cell to carry an R{n} row index, aborting on any ragged/ambiguous run.
ReplaceSubtree allocates a fresh root id, so the structural restructures
(app_shell + table_repair) ran via a new apply_root_transform helper that
returns the current root id; run_cleanup_passes threads it into the subsequent
per-root passes instead of the stale id (which they'd otherwise no-op on).
Weak models emit a desktop dashboard's sidebar either as a full-width band on a
vertical root or crammed into a horizontal row with every section — both render
broken (the removed bespoke scaffold used to pre-build the two-column root).
app_shell::reshape_sidebar_to_app_shell detects a sidebar dashboard and rewrites
it to a horizontal [sidebar(260) | content-column(fill)] shell, run in the shared
cleanup::run_cleanup_passes finalize point (covers orchestrator + agentic loop).
A layout==vertical|none guard on fix_horizontal_overflow stops it re-widening the
narrowed sidebar. Detection is hardened (strong sidebar token + structural
dashboard-content gate) against restaurant/landing/top-nav/mobile/multi-screen
false-positives; 14 unit + 1 integration test, verified end-to-end via op-smoke.
classify_intent matched only creation verbs, so a noun-phrase request like
"Luxury webapp for managing barbershop clients" fell through to the weak chat
loop instead of the orchestrator. Add product/app nouns (webapp/web app/website/
app for/mobile app/admin panel/saas + 网站/网页/小程序/后台) to DESIGN_KEYWORDS.
Reasoning models (glm-5.x / minimax) burn their whole token budget on hidden
<think> and emit an empty design when thinking is left on (glm-5.2: thinking≈30k,
text=0). design_turn_thinking_mode forces ThinkingMode::Disabled for any model
whose profile is thinking_disabled, applied across all design-capable paths —
the design-agent loop, the builtin tool-executing chat loop, and sub-agent
spawns. Claude (thinking productive) keeps the chat default.
A clipped frame with an explicit numeric height now honours that height and
clips overflow (layout_repair no longer lets the root grow to content), matching
Pencil's fixed-height artboards. Status-bar shell strokes are suppressed in the
scene path (width-independent) so they don't paint a stray black frame, and
legacy .op path nodes remap geometry→d so authored paths aren't dropped.
Bundle 11 OFL design-font TTFs and register them with jian-skia at startup so
--render-shots reproduces Pencil's glyph metrics without system fonts.
OPENPENCIL_RENDER_MARGIN tightens the export crop to match Pencil export_nodes
(no frame), and OPENPENCIL_DUMP_LAYOUT prints every node's computed rect as JSONL
to diff our layout against Pencil snapshot_layout element-for-element. Bumps
vendor/jian to the bundled-font + Auto=no-wrap render-parity commits.
Data verdict (copy-pencil-verdict.md): weak-model M3 quality comes entirely from the
SEQUENTIAL deterministic core (manifest element-builders + role-resolver + post-passes
+ finalize), exercised at concurrency=1. The 3 scaffold strategies / mode-rotation /
per-subtask retry ladder are removable complexity not paid for by M3. Collapsed to a
single sequential path; concurrent (perf-only) folds to sequential; dashboard folds into
the generic path (cleanup_desktop_dashboard already handles sparse rows). Kept the
deterministic core 100% intact + BufferDocSink/clamp_concurrency (spawn_agents) + the
dashboard normalizer trio (feeds plan_normalize). M3 gate held within noise (50%→ the one
flip re-ran 2/3 PASS; 7/8 prompts identical; no new failures from removed code).