Flip the manifest repair-layer default for v1 element kinds from
`light` to `system`, so weak-model designs adapt to the document's
light/dark axis instead of baking light-mode hex. Both prerequisites
are now in place: the orchestrator seeds the 56-token semantic palette
at run start (`-*` refs resolve at paint via scene_vars) and
every v1 builder emits numeric gap/padding in system mode. Explicit
`theme=light|dark` still wins.
ab-v9.3 validation (52 prompts × 4 arms, theme=system default):
- Rendering: 0 dangling $color-* refs across all 208 .op; 207/208
carry the seeded palette + Mode:[Light,Dark] axis. The D2 grey-ref
risk does not materialize at corpus scale.
- Structural inertness: theme is injected post-model by repair_args
and only selects fills; role inference + the contrast post-pass are
already $ref-aware (role_post_pass resolves a ref to None and skips).
M3/M5 cannot move on the flip — glm's 15 M3 misses are 100%
missing-role (model output), theme-independent.
- No collapse: M3 ark 81% / glm 71% / DS 85% / m3-think 94%, 0-1
garbage per arm. Per-arm deltas vs ab-v9.2 are cross-codebase
(Kayshen's d31734e2 edited the manifest skill + builders between
runs) + sampling, NOT the theme flip.
Two failures surfaced once the casement build blocker cleared and the
`cargo test --workspace` step finally ran on v0.8.0-new:
- wheel_over_panel_scrolls_rows_and_clamps: `over_topmost_panel` ran
first in apply_wheel and lists the variables panel, so a wheel over
the panel was swallowed (returned true) WITHOUT ever reaching
try_scroll_variables_panel — rows never advanced. Move the variables
scroll handler ahead of the topmost-panel guard; it already swallows
the wheel over its own rect, so the canvas-zoom guard still holds.
- variables_panel_preset_button_toggles_menu: the second press used a
+58 y-offset that lands in the open dropdown's SaveCurrent row, which
intentionally focuses the name input and keeps the popover open (TS
parity, see anchor_press_does_not_double_toggle). Re-press the button
itself (+22) so the test exercises the toggle it is named for.
op-host-native lib: 296 passed / 0 failed.
theme=system audit (D2 prerequisite): the schema accepts $spacing-*
expression refs in gap/padding, but both Rust layout paths zero them
out — jian-core container_to_style and op-pen-loader gap_value fall
through Expression to 0 — so a ref collapses the layout instead of
resolving. Exactly three emission sites existed: modal_shell's card
padding ($spacing-5) and gap ($spacing-3), and toast's pill gap
($spacing-2). Emit the palette-equal numbers (24/12/8) in every mode,
same pattern as the numeric font-size toast fix (011f4104). A new
catalog-wide invariant test builds all kinds with theme=system and
rejects any expression gap/padding, so future ports can't reintroduce
the class. Color $refs are unaffected — scene_vars resolves fill and
stroke refs at paint.
Replace the global OPENPENCIL_MANIFEST env gate with per-model routing:
manifest defaults ON for the families that cleared the ab-v9.2 KPI gate
(minimax 92% / ark-code 92% / glm 98% / deepseek 83%, all >=70%) and
stays OFF for models without benchmark data — the catalog is a floor
for weak models, not a ceiling for strong ones. The env var remains a
both-ways override (1/true/on forces on, 0/false/off forces off) so
op-smoke benchmarks and rollback keep their one-knob workflow.
Also plumb the selected built-in agent's model id into DesignRequest:
the production chat path always sent model:None, which routed every
model to Full-tier prompts and dead-armed both this gate and the M3
thinking policy outside op-smoke. Live-verified: glm-5.1 with no env
set engages the manifest protocol purely via routing.
Port the TS applySemanticPalette 56-token palette (22 themed light/dark
colors on a Mode axis, 6 chart singles, 28 numerics) and wake the
dormant variables.rs seeding path. Merge semantics differ from TS:
editor-core set_variables_bulk extends (incoming wins) while TS keeps
existing values, so the snapshot computes the missing set first and
seeds only absent tokens via one MergeThemePreset; rollback removes
exactly the created names/axes when no content survives. Live-smoke
verified: generated .op carries the 56 vars + Mode axis and $color-*
refs resolve at paint.
A childless Frame with explicit stroke or fill is pixel-equivalent to a
childless Rectangle, which the blank check already admits. otp_input's
empty stroked digit slots made perfect manifests read as blank
scaffolding -> retry ladder -> raw fallback degradation (ark emitted
flawless otp manifests twice and got rejected both times). Also
serialize the OPENPENCIL_MANIFEST env tests behind a shared lock --
parallel set/remove races only surfaced in full-suite runs.
ab-v9.1 misses were attention failures, not vocabulary failures: kinds
named literally in the brief (toolbar 5/10) still got hand-composed,
while anti-pattern-named kinds rarely missed. Token-match the subtask
label+elements text (plus the whole brief for single-section plans)
against the catalog, expand via a data-grounded synonym table, and
inject up to 6 nominations as an ELEMENT HINTS block with composite-
first and no-hand-compose rules. ab-v9.2 matrix: 51 FAIL->PASS vs 4
noise-shaped regressions; all four arms >=83% M3-pass, M5 adoption 96%.
Full-tier prompts run ~8k tokens and deepseek kept hand-rolling past a
top-of-prompt element catalog (ab-v9 58% vs Basic-tier 85%). Pull the
element-manifest skill out of priority order and append it last, right
above the output contract, so recency keeps the catalog in attention.
The skill's COVERAGE section now also teaches composite-first checking
instead of only naming examples.
M3 with thinking disabled drops to M2.7-level element adoption; route
design sub-agent calls through Adaptive thinking with a 16384 budget
for the M3 family while other built-in models keep Disabled+8192.
Validated by the ab-v9.2 matrix (m3-think arm 92% M3-pass).
check-widget-boundary.sh F4 forbids op_editor_ui::widgets references
outside widget_host.rs + siblings; the remote-icon fetcher now calls the
new icon_ingest wrapper instead of icon_catalog directly.
Real chat streaming (echo stub retired; production bundle now builds with
codegen), live_sync glue for browser<->daemon<->MCP, codegen panel actions,
structure-bundle zip (stored-zip encoder), iconify web search, component
browser dispatch, accessibility DOM mirror v1, IME + clipboard paths.
scripts/ab-v9/run_matrix.py runs the ab-v3 corpus (52 prompts) through
the full Rust orchestrator per provider (op-smoke headless,
OPENPENCIL_MANIFEST=1) and scores M3 expected-shape (required roles in
the saved .op tree) + M5 element selection, appending rows per cell for
crash-safe resume. Keys come from env only (MM_KEY/ARK_KEY/DS_KEY).
op-smoke grows OPENPENCIL_SMOKE_KEEP_THINKING=1 to keep MiniMax
reasoning ON: ab-v9 showed M3-nothink emits lazy minimal manifests
(17%, ~10s answers) while M3-with-thinking lands 60% with composite
tied-best at ~110s — the MiniMax production routing target.
Weak models stop tool-calling 188 element schemas and instead emit a
JSONL element manifest — one {"el":"stat_card",...} declaration per
line. Nesting uses manifest-local line numbers ("in"), node ids are
system-assigned (the id-hallucination class disappears at the protocol
level), and a repairing argument layer replaces strict rejection:
enum synonyms, schema-driven placeholders for missing required params,
plain-string-to-JSON shape coercion, and a bounded second-chance retry
when a builder names its missing arg. Raw PenNode lines stay available
as an escape hatch and raw frames can serve as "in" containers.
- op-mcp: element_manifest.rs — build_element facade over
semantic_alias_node, kind normalization (prefers theme-aware _v1,
light default until palette seeding lands), schema-generated prompt
catalog; every known kind builds from empty args (the never-reject
guarantee, covered by a 94-kind fixture test)
- op-orchestrator: manifest.rs parser (line-tolerant, id injection for
raw subtrees, depth-clamped sections) wired into run_subtask behind
OPENPENCIL_MANIFEST=1; retry ladder falls back to the raw-JSONL path
- op-ai-skills: element-manifest skill (priority 0, replaces the
jsonl-format skills when active) with anti-hand-roll teaching;
build.rs adds rerun-if-changed so include_dir! sees new skill files
- subagent.rs: childless rectangles count as content — skeleton-screen
designs no longer reject as blank
ab-v9 (52-prompt corpus, full orchestration): ark-code-latest and
glm-5.1 hit 85% expected-shape success vs the 58% tool-calling
baseline, at 5-8k prompt tokens per call vs ~19.9k.
add_toast_v1's system branch put the $type-body-size token ref into
fontSize, a strict f64 schema slot — every system-themed toast failed
PenNode deserialization. Emit 14 in every mode (the ref resolves to the
same default at render time; see the fidelity caveat on
build_modal_shell). gap keeps its $spacing-2 ref — that slot accepts
expressions.
ai_chat_panel/tests.rs had grown to 960 lines. Move the paint-assertion
tests + the recording PanelPaintBackend into a tests_paint.rs sibling
(369 lines), leaving layout/hit-test in tests.rs (599 lines); shared
helpers stay in tests.rs as pub(super).
- run the divider above the input across the full panel width
- paint the built-in model key with the real lucide Key glyph instead
of a hand-rolled ring+shaft approximation
- concurrency chip rests as a faint ghost at 1x and gets a
primary-tinted chip when the agent team is staffed, keeping the
footer hover + cycle wiring
- attach/send are ghost icons; send fades to muted-foreground/30 when
there is nothing to send
Re-lands 94b9aa1b (dropped during a re-batch) merged with the
agent-team footer layout.
Stopping a turn or starting a new chat drops the DesignSession, but the
canvas indicators only cleared when the worker thread next touched its
channel and unwound — which lags seconds behind when the worker is
mid-LLM-call. The breathing borders lingered on a canvas the user had
already moved on from.
Give the host (which owns the turn lifecycle) the indicator epoch:
DesignSession::start mints it (clearing the prior turn at once) and the
session's Drop ends it immediately via end_if_epoch. The worker registers
under the same epoch via Orchestrator::with_indicator_epoch; headless and
test callers pass none and the concurrent path mints its own, so run()'s
signature and the DesignRequest construction sites are untouched.
end_if_epoch clears AND retires the epoch (vs clear_if_epoch which only
clears), so a worker still in its add_frame loop when the turn is stopped
can't re-populate the set under the now-stale epoch. It's epoch-scoped, so
a newer turn that already began is left untouched.
Snapshots the full working tree: web select-all picker wiring
(op-host-web) plus in-progress chrome work (chat input rework,
agent settings, codegen panel, variables modal, toolbar actions,
model discovery).