Second increment of the role-resolver port. New `role_defaults` module ports
all 43 `registerRole` rules from role-definitions/index.ts + `applyDefaults` +
`detectThemeFromNode`. `resolve_tree_roles` now, after inferring a node's role
(I1), fills in that role's default visual/layout properties — theme-aware
(card/input/navbar/divider fills), parent-aware (cards stretch in a horizontal
row), and canvas-width-aware (mobile vs desktop padding).
Defaults are applied via a JSON round-trip merge that sets ONLY keys the model
left unset (AI-explicit always wins), mirroring the TS `record[key] ===
undefined` check; jian's camelCase + flattened container/text props let the
merge map straight onto the schema with no per-field hand-mapping. On any
(de)serialize failure the node is left untouched, so defaults can never drop a
node. Theme + canvas width are derived from the plan's root frame in
run_subtask and threaded down the walk.
The tiny-card guard now applies to EXPLICIT roles too (not just inferred): an
LLM-emitted role:"card"/"stat-card" on a node too small to be a card is stripped
before the defaults pass, so it can't inflate a 6×6 dot into a padded, shadowed
card (Codex review).
29 tests (theme detection, set-if-absent, AI-explicit-wins, theme/parent/
canvas-width variants, CJK typography, divider orientation, explicit+inferred
tiny-card strip, unknown role). Verified on a MiniMax-M3 run: 106 nodes,
search-bars/cards received role defaults, an explicit cornerRadius preserved.
First increment of the role-resolver port (spec
openpencil-docs/.../2026-06-06-role-resolver-rust-port.md). Port of
inferRoleFromName + the name-inference path of resolveNodeRole + the
resolveTreeRoles walk from role-resolver.ts.
New `role_infer` module: NAME_EXACT_MAP + NAME_PATTERN_MAP (regex) + the
container-suffix and role-part-word guards, plus the page-chrome-in-card guard
(a card's inner "Header" must not become a navbar) and the tiny-card guard
("Status Dot" must not become stat-card). Runs on each sub-agent subtree in
run_subtask BEFORE the fallback sizing normalize (semantic-before-fallback).
No role DEFAULTS injection yet (that is increment I2) — but setting `role`
already has effect: the existing cleanup passes (nav-surface repair, section
logic) key off node.role, so an inferred navbar/footer is now repaired the
same way an explicitly-tagged one is.
Adds the `regex` dep (already locked at 1.12.3). 12 tests cover the exact /
pattern maps, all guards, explicit-role-wins, and the tree walk.
P1 ported compactSubAgentSkills' design-system drop. Getting it right needs to
balance three cases, since Rust has NOT ported TS's buildSubAgentStyleGuideInstruction
block and design-system.md carries a conflicting "output ONLY a JSON token object"
header on top of its useful styling RULES:
- design.md present → the design-md skill replaces design-system → DROP.
- no style guide → noStyleGuideMatch → the style-defaults skill covers
styling → DROP (keeping both injected design-system's conflicting header).
- style guide NAMED, no design.md → neither covers it and Rust injects no
style-guide block → design-system is the only styling guidance → KEEP.
Implement by passing `has_design_md || no_style_guide_match` as the
`has_design_system_substitute` signal to the base filter, and keeping
design-system in the Basic allow-set + reduced-retry kernel so the gap case
survives there too. Three Codex review rounds (2026-06-06): suppress-without-
replacement (style guide), same in reduced retry, and conflict-when-covered.
604 tests pass; adds unit guards for each case + a Full-tier integration test
asserting design-system is dropped with style-defaults present and kept in the
gap case. Porting buildSubAgentStyleGuideInstruction would let Rust drop
design-system for the style-guide case too (tracked follow-up).
execute_visual_ref_orchestration + the visual_ref module ported a TS pipeline
upstream had already deleted (commit 0f12b6e9). Nothing ever dispatched on
visual_ref_enabled and no symbol was used outside the module (verified). Delete
visual_ref.rs + visual_ref_tests.rs (1552 L) and the lib.rs re-exports.
VisualRefProvider trait, SkippedVisualRefProvider stub, and the visual_ref_enabled
field stay (they live in types.rs/stub_providers.rs; removing the field would
churn ~40 construction sites) — the field is now documented as vestigial.
resolve_generation_skills called resolve_skills with an empty ResolveOptions,
so every Flags-gated generation skill was silently dropped — most critically
jsonl-format-simplified (isBasicTier) for the weak models it targets, plus
design-md / variables / style-defaults / elements. Compute the five flags
(isBasicTier / hasDesignMd / hasVariables / noStyleGuideMatch / hasMcpTools),
the {{designMdContent}} dynamic value, and the tier-scaled budget override,
then pass them through.
Also port compactSubAgentSkills' content-aware base filter (mobile /
style-guide drops + jsonl-format↔simplified dedup + Basic-tier allow-set)
into apply_skill_filter; previously only the minimal/reduced retry modes
were ported, so Basic-tier models kept skills TS would have dropped.
Faithful port of orchestrator-sub-agent.ts:396-453 and
orchestrator-sub-agent-compact.ts:4-76. Verified end-to-end: MiniMax-M3
(Basic tier) now loads the simplified format skill and produces a 93-node
nested design.
Two token-resolution fixes for canonical PenNode fields:
- Scope $type/$spacing/$radius resolution to numeric-typed fields only.
A text `content`/`name` literally "$spacing-3" (a design-system
showcase) was previously corrupted into a number.
- Resolve $type-*-weight tokens to their numeric value. FontWeight is an
untagged {Number, Keyword} enum, so an unresolved weight token
deserialized as a keyword and the weight resolver then fell back to the
default — a 700 heading silently rendered at 400.
op-smoke's QueryEngine path can't send MiniMax's `thinking:{type:
"disabled"}` field, so reasoning models (M3 etc.) thought themselves out
of the token budget and couldn't be benchmarked end-to-end. Add a
non-streaming DirectClient (OPENPENCIL_SMOKE_DIRECT=1) that posts
openai-compat directly and disables thinking at the wire for MiniMax —
or any endpoint via OPENPENCIL_SMOKE_DISABLE_THINKING=1 (Volcengine
honors the same param) — so the production thinking-disable path can be
exercised headless without a GUI.
MiniMax reasoning models (M3 etc.) inject `<think>` into content, which
burns the output-token budget (truncating the JSON) and forces the node
parser to scrub reasoning. Send `thinking:{type:"disabled"}` for MiniMax
models, gated on the caller's ThinkingMode being Disabled: the
orchestrator requests that for design subtasks, while normal chat
(Adaptive) keeps the model's reasoning intact.
The subtask prompt contradicted itself: the jsonl-format skill taught
flat `_parent` while NODE_FORMAT (appended last, so dominant) demanded
a JSON array with nested `children`. Weaker models followed the last
instruction and emitted flat, unnested nodes, collapsing every layout
to a vertical stack. Align NODE_FORMAT and the basic-tier simplified
skill (both Rust + TS copies of the corpus) to flat `_parent`, and make
the user-prompt format line format-agnostic so it conflicts with neither
tier's skill.
Weak models emit dirty node JSON that previously failed whole subtasks.
Re-nest flat `_parent` output into a tree (port of TS parseJsonlToTree;
horizontal containers were collapsing to vertical), strip reasoning
`<think>` blocks (including truncated ones) so drafts aren't parsed as
nodes, resolve documented numeric design tokens (`$type`/`$spacing`/
`$radius`) and wrap bare `fill` color strings that canonical PenNode
rejects, and restore per-node tolerance — re-nesting had collapsed the
flat array so one bad deep node failed the entire root.
Split the test module into sibling parse_tests.rs to keep parse.rs
under the 800-line cap; ark_metrics.jsonl is the captured regression.
The bracket scan could return a nested children array or a stray field array instead of the real node tree. Earlier heuristics each had a hole: combined-depth missed valid JSON after an unmatched prose bracket, and colon-skip dropped legitimately-labelled outputs (a wrapper object whose node array sits under a key).
balanced_spans now just collects candidate spans robustly: each bracket tried independently (an unmatched prose bracket cannot hide a later array), skip-past so nested sub-spans are not separate candidates, string-aware. parse_nodes tries both the array-form candidates and the bare-object/JSONL collection, then keeps the result with the MOST total nodes including descendants — a full tree (root + children) always beats a stray inner array, with no colon/label heuristic.
Adds regression tests for nested-children, JSONL, unmatched-prose, and labelled-wrapper inputs. 40 parse tests pass. Hardened across Codex stop-time reviews.
The headless benchmark paths reported success even when their output/render step failed: op-smoke still returned SUCCESS when the OPENPENCIL_SMOKE_OUT write failed, and --render-shots returned the GUI-skip 'handled' signal (exit 0) on an unreadable/unparsable .op, an empty page, or a node that failed to render.
op-smoke now exits 4 when a requested save fails; --render-shots exits 1/2 on any read/parse/empty/partial-render failure and only exits 0 when every top-level node rendered. Verified: good .op -> 0, missing file -> 1. Found by Codex stop-time review.
main.rs reached 927 lines (over the project 800-line cap, mostly pre-existing) after wiring --render-shots. Move the 177-line handle_menu_action into a sibling menu_action.rs (same pattern as the widget_host split), dropping main.rs to 748. Behaviour unchanged; the method becomes pub(crate). Also drops the now-unused ActiveEventLoop/Fullscreen imports from main.rs. Addresses Codex stop-time review (touched file over 800-cap).
DeepSeek emits node JSON as JSONL: bare top-level objects, one per line inside a json code fence, with no enclosing array brackets. The bracket scan only saw the inner fill sub-arrays (not node arrays), so the parser returned 'no JSON array found' and DeepSeek produced nothing.
parse_nodes now falls back to collecting top-level brace objects and deserializing each as a PenNode when the array path yields no nodes. balanced_arrays/balanced_end were generalized to a shared balanced_spans(open, close) used by both the bracket and brace scans. Adds a JSONL test; 37 parse tests pass.
Multi-candidate selection picked the candidate with the most nodes, but balanced_arrays also collected nested children arrays. For nested-format output (a single root frame whose children array holds 3 nodes), the inner array could outrank the outer (1 node), returning the children and dropping the real root. Flat _parent-style output (what M2.7 emitted in testing) hid it. balanced_arrays now collects only top-level arrays, skipping past each captured span. Adds a regression test. Found by Codex stop-time review.
Adds a fully-headless pipeline to benchmark built-in-Agent design generation across models, with no GUI or live canvas.
op-smoke: OPENPENCIL_SMOKE_OUT saves the generated PenDocument to a .op; OPENPENCIL_SMOKE_DUMP dumps the raw LLM response for diagnosing weak-model JSON malformations. openpencil-desktop --render-shots FILE OUTDIR SCALE loads a .op, runs the jian layout pass, and exports one cropped PNG per top-level node via export_node_raster (skia CPU raster) for node-only benchmark screenshots.
M2.7 and other weak models emit malformed PenNode JSON the strict parser rejected, losing whole subtasks (Food Delivery restaurants; Dashboard charts-row dropped to 0 nodes and aborted downstream subtasks).
Multi-candidate extraction scans every balanced bracket span and keeps the result with the most nodes instead of grabbing the first open bracket, so reasoning-prose brackets (e.g. step markers) no longer trigger 'expected ident'. Per-element tolerant deserialize keeps valid PenNodes and skips+logs bad elements instead of failing the whole array on one stray node (e.g. a fill of type 'solid' written where a node was expected). Both are strict generalizations; fully-valid JSON is unchanged. Adds tests for both classes; 35 parse tests pass.
The Maximize icon (rightmost of the TopBar right cluster) was painted
but had no hit-test mapping, so clicking it did nothing.
- Add TopBarHit::ToggleFullscreen + a hit-test rect matching its paint
position; add a regression test.
- Native: the host raises a pending_fullscreen_toggle intent (it can't
reach the winit window); app_handler consumes it after the press and
routes through the existing handle_menu_action(ToggleFullscreen) ->
window.set_fullscreen(), the same path the macOS menu item used.
- Web: the WASM host toggles the browser Fullscreen API directly
(document request_fullscreen / exit_fullscreen) — no runner round-trip
and no unconsumed intent.
Right-press menu-open + Delete/Duplicate dispatch kept the multi-selection instead of collapsing to the right-clicked row (web + native). Adds a multi-delete core test.
Extend the op-orchestrator phased design runner (single-screen scaffold, sequential sub-agent execution, dashboard columns, validation-fix apply) and adapt the headless smoke runner to the new InsertSubtree page_id field.
Expand the Rust MCP server toward pen-mcp parity: layered batch_design helpers with TS-shaped structured args, semantic element-alias builders for high-frequency elements, insert/update node-data parsers, and desktop mcp_serve/runtime wiring (incl. a file-path tool) with tests.
Add EditorCommand support powering the MCP layered-design workflow: authored-id subtree insertion for design skeletons, deterministic RefineDesign cleanup, plus batch-page and replace-node apply paths and an InsertSubtree page_id target. Includes command_apply/command_node wiring and tests.
Bring the Rust git panel to TS parity across four surfaces:
- Commit rows expand inline into a semantic diff card (ported
diffDocuments / indexNodesById / engineDiff over serde_json::Value)
instead of navigating to a separate diff page; the card shows the
"~modify N nodes" summary plus a per-node patch list. blob_at_commit
reads the {rev} and {rev}^ blobs to compute base-vs-next.
- Overflow menu (...) with the 5 TS actions — switch tracked file, clear
commit author, remote settings, SSH keys, close repository — backed by
GitOverflowView subviews (tracked-file picker + SSH-keys list) and
remote-settings rows (origin URL + ahead/behind + fetch; no HTTPS token
field, matching TS).
- Commit signature form: empty commits are refused, and committing with no
configured author opens an inline name/email form (Enter saves, Escape
cancels) that writes the repo-local identity before retrying the commit.
- Polish: taller commit-detail card, removed panel shadow, lengthened the
ready panel, and unified the commit-textarea caret blink to the app's
500 ms cadence.
New overflow/SSH/author/remote labels resolve via Document::t with an
English fallback for keys not yet in the locale tables.
The TS top-bar shows the active provider glyph inside a rounded bg-muted
box; the Rust chip painted the bare icon, so a spinning/active provider
read as a floating glyph with no chip affordance. Paint a theme.muted
rounded backdrop (pad 4, radius 5) behind the brand icon(s) to match.