Flip the manifest repair-layer default for v1 element kinds from
`light` to `system`, so weak-model designs adapt to the document's
light/dark axis instead of baking light-mode hex. Both prerequisites
are now in place: the orchestrator seeds the 56-token semantic palette
at run start (`-*` refs resolve at paint via scene_vars) and
every v1 builder emits numeric gap/padding in system mode. Explicit
`theme=light|dark` still wins.
ab-v9.3 validation (52 prompts × 4 arms, theme=system default):
- Rendering: 0 dangling $color-* refs across all 208 .op; 207/208
carry the seeded palette + Mode:[Light,Dark] axis. The D2 grey-ref
risk does not materialize at corpus scale.
- Structural inertness: theme is injected post-model by repair_args
and only selects fills; role inference + the contrast post-pass are
already $ref-aware (role_post_pass resolves a ref to None and skips).
M3/M5 cannot move on the flip — glm's 15 M3 misses are 100%
missing-role (model output), theme-independent.
- No collapse: M3 ark 81% / glm 71% / DS 85% / m3-think 94%, 0-1
garbage per arm. Per-arm deltas vs ab-v9.2 are cross-codebase
(Kayshen's d31734e2 edited the manifest skill + builders between
runs) + sampling, NOT the theme flip.
Two failures surfaced once the casement build blocker cleared and the
`cargo test --workspace` step finally ran on v0.8.0-new:
- wheel_over_panel_scrolls_rows_and_clamps: `over_topmost_panel` ran
first in apply_wheel and lists the variables panel, so a wheel over
the panel was swallowed (returned true) WITHOUT ever reaching
try_scroll_variables_panel — rows never advanced. Move the variables
scroll handler ahead of the topmost-panel guard; it already swallows
the wheel over its own rect, so the canvas-zoom guard still holds.
- variables_panel_preset_button_toggles_menu: the second press used a
+58 y-offset that lands in the open dropdown's SaveCurrent row, which
intentionally focuses the name input and keeps the popover open (TS
parity, see anchor_press_does_not_double_toggle). Re-press the button
itself (+22) so the test exercises the toggle it is named for.
op-host-native lib: 296 passed / 0 failed.
theme=system audit (D2 prerequisite): the schema accepts $spacing-*
expression refs in gap/padding, but both Rust layout paths zero them
out — jian-core container_to_style and op-pen-loader gap_value fall
through Expression to 0 — so a ref collapses the layout instead of
resolving. Exactly three emission sites existed: modal_shell's card
padding ($spacing-5) and gap ($spacing-3), and toast's pill gap
($spacing-2). Emit the palette-equal numbers (24/12/8) in every mode,
same pattern as the numeric font-size toast fix (011f4104). A new
catalog-wide invariant test builds all kinds with theme=system and
rejects any expression gap/padding, so future ports can't reintroduce
the class. Color $refs are unaffected — scene_vars resolves fill and
stroke refs at paint.
Replace the global OPENPENCIL_MANIFEST env gate with per-model routing:
manifest defaults ON for the families that cleared the ab-v9.2 KPI gate
(minimax 92% / ark-code 92% / glm 98% / deepseek 83%, all >=70%) and
stays OFF for models without benchmark data — the catalog is a floor
for weak models, not a ceiling for strong ones. The env var remains a
both-ways override (1/true/on forces on, 0/false/off forces off) so
op-smoke benchmarks and rollback keep their one-knob workflow.
Also plumb the selected built-in agent's model id into DesignRequest:
the production chat path always sent model:None, which routed every
model to Full-tier prompts and dead-armed both this gate and the M3
thinking policy outside op-smoke. Live-verified: glm-5.1 with no env
set engages the manifest protocol purely via routing.
Port the TS applySemanticPalette 56-token palette (22 themed light/dark
colors on a Mode axis, 6 chart singles, 28 numerics) and wake the
dormant variables.rs seeding path. Merge semantics differ from TS:
editor-core set_variables_bulk extends (incoming wins) while TS keeps
existing values, so the snapshot computes the missing set first and
seeds only absent tokens via one MergeThemePreset; rollback removes
exactly the created names/axes when no content survives. Live-smoke
verified: generated .op carries the 56 vars + Mode axis and $color-*
refs resolve at paint.
A childless Frame with explicit stroke or fill is pixel-equivalent to a
childless Rectangle, which the blank check already admits. otp_input's
empty stroked digit slots made perfect manifests read as blank
scaffolding -> retry ladder -> raw fallback degradation (ark emitted
flawless otp manifests twice and got rejected both times). Also
serialize the OPENPENCIL_MANIFEST env tests behind a shared lock --
parallel set/remove races only surfaced in full-suite runs.
ab-v9.1 misses were attention failures, not vocabulary failures: kinds
named literally in the brief (toolbar 5/10) still got hand-composed,
while anti-pattern-named kinds rarely missed. Token-match the subtask
label+elements text (plus the whole brief for single-section plans)
against the catalog, expand via a data-grounded synonym table, and
inject up to 6 nominations as an ELEMENT HINTS block with composite-
first and no-hand-compose rules. ab-v9.2 matrix: 51 FAIL->PASS vs 4
noise-shaped regressions; all four arms >=83% M3-pass, M5 adoption 96%.
Full-tier prompts run ~8k tokens and deepseek kept hand-rolling past a
top-of-prompt element catalog (ab-v9 58% vs Basic-tier 85%). Pull the
element-manifest skill out of priority order and append it last, right
above the output contract, so recency keeps the catalog in attention.
The skill's COVERAGE section now also teaches composite-first checking
instead of only naming examples.
M3 with thinking disabled drops to M2.7-level element adoption; route
design sub-agent calls through Adaptive thinking with a 16384 budget
for the M3 family while other built-in models keep Disabled+8192.
Validated by the ab-v9.2 matrix (m3-think arm 92% M3-pass).
check-widget-boundary.sh F4 forbids op_editor_ui::widgets references
outside widget_host.rs + siblings; the remote-icon fetcher now calls the
new icon_ingest wrapper instead of icon_catalog directly.
Real chat streaming (echo stub retired; production bundle now builds with
codegen), live_sync glue for browser<->daemon<->MCP, codegen panel actions,
structure-bundle zip (stored-zip encoder), iconify web search, component
browser dispatch, accessibility DOM mirror v1, IME + clipboard paths.
scripts/ab-v9/run_matrix.py runs the ab-v3 corpus (52 prompts) through
the full Rust orchestrator per provider (op-smoke headless,
OPENPENCIL_MANIFEST=1) and scores M3 expected-shape (required roles in
the saved .op tree) + M5 element selection, appending rows per cell for
crash-safe resume. Keys come from env only (MM_KEY/ARK_KEY/DS_KEY).
op-smoke grows OPENPENCIL_SMOKE_KEEP_THINKING=1 to keep MiniMax
reasoning ON: ab-v9 showed M3-nothink emits lazy minimal manifests
(17%, ~10s answers) while M3-with-thinking lands 60% with composite
tied-best at ~110s — the MiniMax production routing target.
Weak models stop tool-calling 188 element schemas and instead emit a
JSONL element manifest — one {"el":"stat_card",...} declaration per
line. Nesting uses manifest-local line numbers ("in"), node ids are
system-assigned (the id-hallucination class disappears at the protocol
level), and a repairing argument layer replaces strict rejection:
enum synonyms, schema-driven placeholders for missing required params,
plain-string-to-JSON shape coercion, and a bounded second-chance retry
when a builder names its missing arg. Raw PenNode lines stay available
as an escape hatch and raw frames can serve as "in" containers.
- op-mcp: element_manifest.rs — build_element facade over
semantic_alias_node, kind normalization (prefers theme-aware _v1,
light default until palette seeding lands), schema-generated prompt
catalog; every known kind builds from empty args (the never-reject
guarantee, covered by a 94-kind fixture test)
- op-orchestrator: manifest.rs parser (line-tolerant, id injection for
raw subtrees, depth-clamped sections) wired into run_subtask behind
OPENPENCIL_MANIFEST=1; retry ladder falls back to the raw-JSONL path
- op-ai-skills: element-manifest skill (priority 0, replaces the
jsonl-format skills when active) with anti-hand-roll teaching;
build.rs adds rerun-if-changed so include_dir! sees new skill files
- subagent.rs: childless rectangles count as content — skeleton-screen
designs no longer reject as blank
ab-v9 (52-prompt corpus, full orchestration): ark-code-latest and
glm-5.1 hit 85% expected-shape success vs the 58% tool-calling
baseline, at 5-8k prompt tokens per call vs ~19.9k.
add_toast_v1's system branch put the $type-body-size token ref into
fontSize, a strict f64 schema slot — every system-themed toast failed
PenNode deserialization. Emit 14 in every mode (the ref resolves to the
same default at render time; see the fidelity caveat on
build_modal_shell). gap keeps its $spacing-2 ref — that slot accepts
expressions.
ai_chat_panel/tests.rs had grown to 960 lines. Move the paint-assertion
tests + the recording PanelPaintBackend into a tests_paint.rs sibling
(369 lines), leaving layout/hit-test in tests.rs (599 lines); shared
helpers stay in tests.rs as pub(super).
- run the divider above the input across the full panel width
- paint the built-in model key with the real lucide Key glyph instead
of a hand-rolled ring+shaft approximation
- concurrency chip rests as a faint ghost at 1x and gets a
primary-tinted chip when the agent team is staffed, keeping the
footer hover + cycle wiring
- attach/send are ghost icons; send fades to muted-foreground/30 when
there is nothing to send
Re-lands 94b9aa1b (dropped during a re-batch) merged with the
agent-team footer layout.
Stopping a turn or starting a new chat drops the DesignSession, but the
canvas indicators only cleared when the worker thread next touched its
channel and unwound — which lags seconds behind when the worker is
mid-LLM-call. The breathing borders lingered on a canvas the user had
already moved on from.
Give the host (which owns the turn lifecycle) the indicator epoch:
DesignSession::start mints it (clearing the prior turn at once) and the
session's Drop ends it immediately via end_if_epoch. The worker registers
under the same epoch via Orchestrator::with_indicator_epoch; headless and
test callers pass none and the concurrent path mints its own, so run()'s
signature and the DesignRequest construction sites are untouched.
end_if_epoch clears AND retires the epoch (vs clear_if_epoch which only
clears), so a worker still in its add_frame loop when the turn is stopped
can't re-populate the set under the now-stale epoch. It's epoch-scoped, so
a newer turn that already began is left untouched.
Snapshots the full working tree: web select-all picker wiring
(op-host-web) plus in-progress chrome work (chat input rework,
agent settings, codegen panel, variables modal, toolbar actions,
model discovery).
dev-desktop wraps cargo-watch to rebuild + relaunch op-host-desktop on
save; test-ui runs the windowless op-editor-ui geometry tests for a fast
inner loop. Both are opt-in conveniences (cargo-watch is a local
install) and don't affect the build.
The agent-team canvas indicators are a process-global registry shared by
the design worker (which registers them) and the paint pass (which reads
them). Two ways a stale/cancelled run could corrupt the run that
replaced it:
1. Teardown — the host cleared the registry unconditionally, so an old
run finishing late wiped a newer run's live indicators.
2. Registration — add_frame did no epoch check, so an old worker still
in its registration loop after a newer begin() folded its frames into
the new run's registry, leaving stale indicators that only cleared
when the new run ended.
Stamp each run with an epoch (begin()), kept inside the registry mutex.
Every registration (add_node/add_frame/mark_preview/clear_preview) and
the RAII teardown guard compare the epoch and mutate under the same
lock, so anything from a superseded run is dropped on the floor and a
begin() racing in between can't slip through the gap. The redundant
host-side clears are removed — the worker owns the lifecycle. At most
the latest run ever has live indicators.
The badge filled at 0.92 alpha, so in light theme it composited over the
pale canvas and lifted the real background luminance past the contrast
crossover — white-on-purple stopped being legible. Fill opaque so the
foreground pick matches what is actually drawn.
The perceptual-luminance threshold mis-classified coral (#FF6B6B) as
dark and gave it unreadable white text. Switch to gamma-corrected WCAG
relative luminance with the standard 0.179 crossover: every palette
colour but the dark purple now takes dark glyphs. Test pins all six.