Plain-data types for the Code panel (Framework / CodegenPhase /
ChunkStatus / CodeGenProgress / ChunkProgress / AssetMeta /
CodegenState) in a new op-editor-core::codegen module. Keeps
op-editor-core wasm-clean; pipeline logic stays in op-codegen::ai
which already depends on this crate (acyclic).
The `Install Bun` step (oven-sh/setup-bun@v2) intermittently fails the Bun
download and reds the whole Rust Check job even when every Rust step (fmt /
build / test / clippy -D warnings) passed and windows/macos jobs are green.
Replace it with the official install script in a real 3-attempt retry loop
(no new third-party action), exporting the bin dir to GITHUB_PATH for the
JS-deps + planner-prompt drift-guard steps.
`set -o pipefail` is REQUIRED so `curl | bash` surfaces a curl download failure
(otherwise the pipeline returns bash's exit and the flake is never retried),
and an explicit ok-flag + `exit 1` after 3 attempts keeps a persistent real
failure from silently passing the step. (Codex review.)
Port of buildSubAgentStyleGuideInstruction (orchestrator-sub-agent-compact.ts):
look up the planner-selected guide by name (select_style_guide + the registry),
extract its palette/fonts/radius (extract_style_guide_values), and append a
tier-aware "follow these specs" block to the sub-agent system prompt so the
model uses the selected aesthetic instead of inventing a conflicting palette.
RUST ADAPTATION: TS emits $color-* refs (which it seeds into doc.variables);
Rust does not seed style-guide vars, so refs wouldn't resolve — we emit the
guide's CONCRETE HEX values instead. Same effect, correct for the Rust path.
This also CLOSES the design-system gap the earlier fix left open: when the
style-guide block is injected it REPLACES design-system, so design_system_covered
now also includes style_guide_instruction.is_some() — a named guide no longer
falls back to the generic (and conflicting) design-system skill.
Tests: a named style guide injects "VISUAL STYLE GUIDE" + drops design-system;
the no-style-guide path still drops design-system via style-defaults.
equalizeHorizontalSiblings (the dashboard column-equalization pass) is the same
logic as the equalizeCardRow ported in I4, except it also excludes badge/pill/tag
roles from card candidates. Add those exclusions so equalize_card_row covers both
— closing the P3 "dashboard 等宽" item.
The other two P3 items are architectural non-gaps, not ports:
- 变量 seeding: plans carry no palette (variables.rs is a deliberate no-op); the
design-system path (design_system_to_seed_commands) already seeds variables.
- intent-modify: the Rust Design path is AGENTIC — an agent::Provider whose
modify-level tools (update_node / batch_design / ...) handle create-vs-modify
organically, so the 2-way Design/Chat gate suffices. TS's 3-way classify
existed for its non-agentic routing and would be redundant here.
Fourth increment of the role-resolver port — the SAFE, property-only subset of
resolveTreePostPass's layout fixes, added to the same JSON post-pass walk before
the contrast cluster (TS order 1/3/4/8):
- equalizeCardRow: a row of unequal fixed-width card frames (width ratio < 0.6,
similar heights) is promoted to fill_container so taffy stretches them evenly.
- normalizeFormInputWidths: fixed-width inputs become fill_container when a
fill_container frame sibling is present.
- normalizeInputTrailingIconAlignment: text before a trailing icon becomes
fill_container + fixed-width so the icon is pushed to the right edge.
- clipContent: a rounded frame containing an image gets clipContent=true.
These only SET properties (taffy then lays out) — no pixel recomputation. The
layout-COMPUTING fixes are deliberately deferred as they overlap Kayshen's
jian/taffy work or need text measurement: fixHorizontalOverflow (recomputes
child/parent pixel widths), fixTextHeights + frame-height expansion (font
metrics), repairPlaceholderIcons (icon catalog). Tracked in the spec.
5 tests (card-row equalize + similar-width skip, form-input promote, trailing-
icon text-fill, clipContent for rounded image frame).
hex_luminance (role_post_pass / I3) and parse_hex_rgb (role_defaults / I2)
byte-sliced a color string after only a byte-LENGTH check, so a malformed
multi-byte color like "#héllo!" passed the length check and `h[0..2]` landed
mid-codepoint, panicking. Reject non-ASCII before slicing in both. (Codex review.)
Third increment of the role-resolver port. New `role_post_pass` module ports the
CONTRAST fixes from resolveTreePostPass: fixButtonForegroundContrast (button
text/icon get a readable fg vs the button bg; a sibling text's color is the
reference so text+icon read as one unit), fixSectionAlternation (alternating
white/slate section backgrounds on light pages), fixOrphanContainerContrast (an
unfilled card-family node with cornerRadius sitting on an unfilled parent gets a
white fill + shadow so it's visible), fixInputSiblingConsistency (sibling inputs
unified to the first's fill/stroke), plus the color helpers (hexLuminance,
hasFill, hasVisibleFill, getFirstSolidColor, needsLuminanceContrastOverride,
resolveColorMaybeRef).
Runs on a JSON Value tree after role resolution (keys off I1/I2 roles) and
before the fallback normalize; each root round-trips through PenNode so a fix
can never drop a node. resolveColorMaybeRef returns None for $color-* refs (no
doc-variable context at sub-agent time), so ref-dependent fixes skip — matching
TS. The LAYOUT fixes (equalizeCardRow / fixHorizontalOverflow / fixTextHeights)
are I4 and overlap Kayshen's taffy work — not included.
13 tests (button dark/light/icon-copy/transparent/unresolved-ref, orphan
fill+shadow / parent-filled / root / structural-role / no-radius skips, section
alternation + dark-page skip, input sibling unify, forest round-trip).
Second increment of the role-resolver port. New `role_defaults` module ports
all 43 `registerRole` rules from role-definitions/index.ts + `applyDefaults` +
`detectThemeFromNode`. `resolve_tree_roles` now, after inferring a node's role
(I1), fills in that role's default visual/layout properties — theme-aware
(card/input/navbar/divider fills), parent-aware (cards stretch in a horizontal
row), and canvas-width-aware (mobile vs desktop padding).
Defaults are applied via a JSON round-trip merge that sets ONLY keys the model
left unset (AI-explicit always wins), mirroring the TS `record[key] ===
undefined` check; jian's camelCase + flattened container/text props let the
merge map straight onto the schema with no per-field hand-mapping. On any
(de)serialize failure the node is left untouched, so defaults can never drop a
node. Theme + canvas width are derived from the plan's root frame in
run_subtask and threaded down the walk.
The tiny-card guard now applies to EXPLICIT roles too (not just inferred): an
LLM-emitted role:"card"/"stat-card" on a node too small to be a card is stripped
before the defaults pass, so it can't inflate a 6×6 dot into a padded, shadowed
card (Codex review).
29 tests (theme detection, set-if-absent, AI-explicit-wins, theme/parent/
canvas-width variants, CJK typography, divider orientation, explicit+inferred
tiny-card strip, unknown role). Verified on a MiniMax-M3 run: 106 nodes,
search-bars/cards received role defaults, an explicit cornerRadius preserved.
First increment of the role-resolver port (spec
openpencil-docs/.../2026-06-06-role-resolver-rust-port.md). Port of
inferRoleFromName + the name-inference path of resolveNodeRole + the
resolveTreeRoles walk from role-resolver.ts.
New `role_infer` module: NAME_EXACT_MAP + NAME_PATTERN_MAP (regex) + the
container-suffix and role-part-word guards, plus the page-chrome-in-card guard
(a card's inner "Header" must not become a navbar) and the tiny-card guard
("Status Dot" must not become stat-card). Runs on each sub-agent subtree in
run_subtask BEFORE the fallback sizing normalize (semantic-before-fallback).
No role DEFAULTS injection yet (that is increment I2) — but setting `role`
already has effect: the existing cleanup passes (nav-surface repair, section
logic) key off node.role, so an inferred navbar/footer is now repaired the
same way an explicitly-tagged one is.
Adds the `regex` dep (already locked at 1.12.3). 12 tests cover the exact /
pattern maps, all guards, explicit-role-wins, and the tree walk.
P1 ported compactSubAgentSkills' design-system drop. Getting it right needs to
balance three cases, since Rust has NOT ported TS's buildSubAgentStyleGuideInstruction
block and design-system.md carries a conflicting "output ONLY a JSON token object"
header on top of its useful styling RULES:
- design.md present → the design-md skill replaces design-system → DROP.
- no style guide → noStyleGuideMatch → the style-defaults skill covers
styling → DROP (keeping both injected design-system's conflicting header).
- style guide NAMED, no design.md → neither covers it and Rust injects no
style-guide block → design-system is the only styling guidance → KEEP.
Implement by passing `has_design_md || no_style_guide_match` as the
`has_design_system_substitute` signal to the base filter, and keeping
design-system in the Basic allow-set + reduced-retry kernel so the gap case
survives there too. Three Codex review rounds (2026-06-06): suppress-without-
replacement (style guide), same in reduced retry, and conflict-when-covered.
604 tests pass; adds unit guards for each case + a Full-tier integration test
asserting design-system is dropped with style-defaults present and kept in the
gap case. Porting buildSubAgentStyleGuideInstruction would let Rust drop
design-system for the style-guide case too (tracked follow-up).
execute_visual_ref_orchestration + the visual_ref module ported a TS pipeline
upstream had already deleted (commit 0f12b6e9). Nothing ever dispatched on
visual_ref_enabled and no symbol was used outside the module (verified). Delete
visual_ref.rs + visual_ref_tests.rs (1552 L) and the lib.rs re-exports.
VisualRefProvider trait, SkippedVisualRefProvider stub, and the visual_ref_enabled
field stay (they live in types.rs/stub_providers.rs; removing the field would
churn ~40 construction sites) — the field is now documented as vestigial.
resolve_generation_skills called resolve_skills with an empty ResolveOptions,
so every Flags-gated generation skill was silently dropped — most critically
jsonl-format-simplified (isBasicTier) for the weak models it targets, plus
design-md / variables / style-defaults / elements. Compute the five flags
(isBasicTier / hasDesignMd / hasVariables / noStyleGuideMatch / hasMcpTools),
the {{designMdContent}} dynamic value, and the tier-scaled budget override,
then pass them through.
Also port compactSubAgentSkills' content-aware base filter (mobile /
style-guide drops + jsonl-format↔simplified dedup + Basic-tier allow-set)
into apply_skill_filter; previously only the minimal/reduced retry modes
were ported, so Basic-tier models kept skills TS would have dropped.
Faithful port of orchestrator-sub-agent.ts:396-453 and
orchestrator-sub-agent-compact.ts:4-76. Verified end-to-end: MiniMax-M3
(Basic tier) now loads the simplified format skill and produces a 93-node
nested design.
Two token-resolution fixes for canonical PenNode fields:
- Scope $type/$spacing/$radius resolution to numeric-typed fields only.
A text `content`/`name` literally "$spacing-3" (a design-system
showcase) was previously corrupted into a number.
- Resolve $type-*-weight tokens to their numeric value. FontWeight is an
untagged {Number, Keyword} enum, so an unresolved weight token
deserialized as a keyword and the weight resolver then fell back to the
default — a 700 heading silently rendered at 400.
op-smoke's QueryEngine path can't send MiniMax's `thinking:{type:
"disabled"}` field, so reasoning models (M3 etc.) thought themselves out
of the token budget and couldn't be benchmarked end-to-end. Add a
non-streaming DirectClient (OPENPENCIL_SMOKE_DIRECT=1) that posts
openai-compat directly and disables thinking at the wire for MiniMax —
or any endpoint via OPENPENCIL_SMOKE_DISABLE_THINKING=1 (Volcengine
honors the same param) — so the production thinking-disable path can be
exercised headless without a GUI.
MiniMax reasoning models (M3 etc.) inject `<think>` into content, which
burns the output-token budget (truncating the JSON) and forces the node
parser to scrub reasoning. Send `thinking:{type:"disabled"}` for MiniMax
models, gated on the caller's ThinkingMode being Disabled: the
orchestrator requests that for design subtasks, while normal chat
(Adaptive) keeps the model's reasoning intact.
The subtask prompt contradicted itself: the jsonl-format skill taught
flat `_parent` while NODE_FORMAT (appended last, so dominant) demanded
a JSON array with nested `children`. Weaker models followed the last
instruction and emitted flat, unnested nodes, collapsing every layout
to a vertical stack. Align NODE_FORMAT and the basic-tier simplified
skill (both Rust + TS copies of the corpus) to flat `_parent`, and make
the user-prompt format line format-agnostic so it conflicts with neither
tier's skill.
Weak models emit dirty node JSON that previously failed whole subtasks.
Re-nest flat `_parent` output into a tree (port of TS parseJsonlToTree;
horizontal containers were collapsing to vertical), strip reasoning
`<think>` blocks (including truncated ones) so drafts aren't parsed as
nodes, resolve documented numeric design tokens (`$type`/`$spacing`/
`$radius`) and wrap bare `fill` color strings that canonical PenNode
rejects, and restore per-node tolerance — re-nesting had collapsed the
flat array so one bad deep node failed the entire root.
Split the test module into sibling parse_tests.rs to keep parse.rs
under the 800-line cap; ark_metrics.jsonl is the captured regression.
The bracket scan could return a nested children array or a stray field array instead of the real node tree. Earlier heuristics each had a hole: combined-depth missed valid JSON after an unmatched prose bracket, and colon-skip dropped legitimately-labelled outputs (a wrapper object whose node array sits under a key).
balanced_spans now just collects candidate spans robustly: each bracket tried independently (an unmatched prose bracket cannot hide a later array), skip-past so nested sub-spans are not separate candidates, string-aware. parse_nodes tries both the array-form candidates and the bare-object/JSONL collection, then keeps the result with the MOST total nodes including descendants — a full tree (root + children) always beats a stray inner array, with no colon/label heuristic.
Adds regression tests for nested-children, JSONL, unmatched-prose, and labelled-wrapper inputs. 40 parse tests pass. Hardened across Codex stop-time reviews.
The headless benchmark paths reported success even when their output/render step failed: op-smoke still returned SUCCESS when the OPENPENCIL_SMOKE_OUT write failed, and --render-shots returned the GUI-skip 'handled' signal (exit 0) on an unreadable/unparsable .op, an empty page, or a node that failed to render.
op-smoke now exits 4 when a requested save fails; --render-shots exits 1/2 on any read/parse/empty/partial-render failure and only exits 0 when every top-level node rendered. Verified: good .op -> 0, missing file -> 1. Found by Codex stop-time review.
main.rs reached 927 lines (over the project 800-line cap, mostly pre-existing) after wiring --render-shots. Move the 177-line handle_menu_action into a sibling menu_action.rs (same pattern as the widget_host split), dropping main.rs to 748. Behaviour unchanged; the method becomes pub(crate). Also drops the now-unused ActiveEventLoop/Fullscreen imports from main.rs. Addresses Codex stop-time review (touched file over 800-cap).
DeepSeek emits node JSON as JSONL: bare top-level objects, one per line inside a json code fence, with no enclosing array brackets. The bracket scan only saw the inner fill sub-arrays (not node arrays), so the parser returned 'no JSON array found' and DeepSeek produced nothing.
parse_nodes now falls back to collecting top-level brace objects and deserializing each as a PenNode when the array path yields no nodes. balanced_arrays/balanced_end were generalized to a shared balanced_spans(open, close) used by both the bracket and brace scans. Adds a JSONL test; 37 parse tests pass.
Multi-candidate selection picked the candidate with the most nodes, but balanced_arrays also collected nested children arrays. For nested-format output (a single root frame whose children array holds 3 nodes), the inner array could outrank the outer (1 node), returning the children and dropping the real root. Flat _parent-style output (what M2.7 emitted in testing) hid it. balanced_arrays now collects only top-level arrays, skipping past each captured span. Adds a regression test. Found by Codex stop-time review.
Adds a fully-headless pipeline to benchmark built-in-Agent design generation across models, with no GUI or live canvas.
op-smoke: OPENPENCIL_SMOKE_OUT saves the generated PenDocument to a .op; OPENPENCIL_SMOKE_DUMP dumps the raw LLM response for diagnosing weak-model JSON malformations. openpencil-desktop --render-shots FILE OUTDIR SCALE loads a .op, runs the jian layout pass, and exports one cropped PNG per top-level node via export_node_raster (skia CPU raster) for node-only benchmark screenshots.
M2.7 and other weak models emit malformed PenNode JSON the strict parser rejected, losing whole subtasks (Food Delivery restaurants; Dashboard charts-row dropped to 0 nodes and aborted downstream subtasks).
Multi-candidate extraction scans every balanced bracket span and keeps the result with the most nodes instead of grabbing the first open bracket, so reasoning-prose brackets (e.g. step markers) no longer trigger 'expected ident'. Per-element tolerant deserialize keeps valid PenNodes and skips+logs bad elements instead of failing the whole array on one stray node (e.g. a fill of type 'solid' written where a node was expected). Both are strict generalizations; fully-valid JSON is unchanged. Adds tests for both classes; 35 parse tests pass.
The Maximize icon (rightmost of the TopBar right cluster) was painted
but had no hit-test mapping, so clicking it did nothing.
- Add TopBarHit::ToggleFullscreen + a hit-test rect matching its paint
position; add a regression test.
- Native: the host raises a pending_fullscreen_toggle intent (it can't
reach the winit window); app_handler consumes it after the press and
routes through the existing handle_menu_action(ToggleFullscreen) ->
window.set_fullscreen(), the same path the macOS menu item used.
- Web: the WASM host toggles the browser Fullscreen API directly
(document request_fullscreen / exit_fullscreen) — no runner round-trip
and no unconsumed intent.
Right-press menu-open + Delete/Duplicate dispatch kept the multi-selection instead of collapsing to the right-clicked row (web + native). Adds a multi-delete core test.
Extend the op-orchestrator phased design runner (single-screen scaffold, sequential sub-agent execution, dashboard columns, validation-fix apply) and adapt the headless smoke runner to the new InsertSubtree page_id field.
Expand the Rust MCP server toward pen-mcp parity: layered batch_design helpers with TS-shaped structured args, semantic element-alias builders for high-frequency elements, insert/update node-data parsers, and desktop mcp_serve/runtime wiring (incl. a file-path tool) with tests.
Add EditorCommand support powering the MCP layered-design workflow: authored-id subtree insertion for design skeletons, deterministic RefineDesign cleanup, plus batch-page and replace-node apply paths and an InsertSubtree page_id target. Includes command_apply/command_node wiring and tests.