scripts/ab-v9/run_matrix.py runs the ab-v3 corpus (52 prompts) through
the full Rust orchestrator per provider (op-smoke headless,
OPENPENCIL_MANIFEST=1) and scores M3 expected-shape (required roles in
the saved .op tree) + M5 element selection, appending rows per cell for
crash-safe resume. Keys come from env only (MM_KEY/ARK_KEY/DS_KEY).
op-smoke grows OPENPENCIL_SMOKE_KEEP_THINKING=1 to keep MiniMax
reasoning ON: ab-v9 showed M3-nothink emits lazy minimal manifests
(17%, ~10s answers) while M3-with-thinking lands 60% with composite
tied-best at ~110s — the MiniMax production routing target.
Weak models stop tool-calling 188 element schemas and instead emit a
JSONL element manifest — one {"el":"stat_card",...} declaration per
line. Nesting uses manifest-local line numbers ("in"), node ids are
system-assigned (the id-hallucination class disappears at the protocol
level), and a repairing argument layer replaces strict rejection:
enum synonyms, schema-driven placeholders for missing required params,
plain-string-to-JSON shape coercion, and a bounded second-chance retry
when a builder names its missing arg. Raw PenNode lines stay available
as an escape hatch and raw frames can serve as "in" containers.
- op-mcp: element_manifest.rs — build_element facade over
semantic_alias_node, kind normalization (prefers theme-aware _v1,
light default until palette seeding lands), schema-generated prompt
catalog; every known kind builds from empty args (the never-reject
guarantee, covered by a 94-kind fixture test)
- op-orchestrator: manifest.rs parser (line-tolerant, id injection for
raw subtrees, depth-clamped sections) wired into run_subtask behind
OPENPENCIL_MANIFEST=1; retry ladder falls back to the raw-JSONL path
- op-ai-skills: element-manifest skill (priority 0, replaces the
jsonl-format skills when active) with anti-hand-roll teaching;
build.rs adds rerun-if-changed so include_dir! sees new skill files
- subagent.rs: childless rectangles count as content — skeleton-screen
designs no longer reject as blank
ab-v9 (52-prompt corpus, full orchestration): ark-code-latest and
glm-5.1 hit 85% expected-shape success vs the 58% tool-calling
baseline, at 5-8k prompt tokens per call vs ~19.9k.
add_toast_v1's system branch put the $type-body-size token ref into
fontSize, a strict f64 schema slot — every system-themed toast failed
PenNode deserialization. Emit 14 in every mode (the ref resolves to the
same default at render time; see the fidelity caveat on
build_modal_shell). gap keeps its $spacing-2 ref — that slot accepts
expressions.
ai_chat_panel/tests.rs had grown to 960 lines. Move the paint-assertion
tests + the recording PanelPaintBackend into a tests_paint.rs sibling
(369 lines), leaving layout/hit-test in tests.rs (599 lines); shared
helpers stay in tests.rs as pub(super).
- run the divider above the input across the full panel width
- paint the built-in model key with the real lucide Key glyph instead
of a hand-rolled ring+shaft approximation
- concurrency chip rests as a faint ghost at 1x and gets a
primary-tinted chip when the agent team is staffed, keeping the
footer hover + cycle wiring
- attach/send are ghost icons; send fades to muted-foreground/30 when
there is nothing to send
Re-lands 94b9aa1b (dropped during a re-batch) merged with the
agent-team footer layout.
Stopping a turn or starting a new chat drops the DesignSession, but the
canvas indicators only cleared when the worker thread next touched its
channel and unwound — which lags seconds behind when the worker is
mid-LLM-call. The breathing borders lingered on a canvas the user had
already moved on from.
Give the host (which owns the turn lifecycle) the indicator epoch:
DesignSession::start mints it (clearing the prior turn at once) and the
session's Drop ends it immediately via end_if_epoch. The worker registers
under the same epoch via Orchestrator::with_indicator_epoch; headless and
test callers pass none and the concurrent path mints its own, so run()'s
signature and the DesignRequest construction sites are untouched.
end_if_epoch clears AND retires the epoch (vs clear_if_epoch which only
clears), so a worker still in its add_frame loop when the turn is stopped
can't re-populate the set under the now-stale epoch. It's epoch-scoped, so
a newer turn that already began is left untouched.
Snapshots the full working tree: web select-all picker wiring
(op-host-web) plus in-progress chrome work (chat input rework,
agent settings, codegen panel, variables modal, toolbar actions,
model discovery).
dev-desktop wraps cargo-watch to rebuild + relaunch op-host-desktop on
save; test-ui runs the windowless op-editor-ui geometry tests for a fast
inner loop. Both are opt-in conveniences (cargo-watch is a local
install) and don't affect the build.
The agent-team canvas indicators are a process-global registry shared by
the design worker (which registers them) and the paint pass (which reads
them). Two ways a stale/cancelled run could corrupt the run that
replaced it:
1. Teardown — the host cleared the registry unconditionally, so an old
run finishing late wiped a newer run's live indicators.
2. Registration — add_frame did no epoch check, so an old worker still
in its registration loop after a newer begin() folded its frames into
the new run's registry, leaving stale indicators that only cleared
when the new run ended.
Stamp each run with an epoch (begin()), kept inside the registry mutex.
Every registration (add_node/add_frame/mark_preview/clear_preview) and
the RAII teardown guard compare the epoch and mutate under the same
lock, so anything from a superseded run is dropped on the floor and a
begin() racing in between can't slip through the gap. The redundant
host-side clears are removed — the worker owns the lifecycle. At most
the latest run ever has live indicators.
The badge filled at 0.92 alpha, so in light theme it composited over the
pale canvas and lifted the real background luminance past the contrast
crossover — white-on-purple stopped being legible. Fill opaque so the
foreground pick matches what is actually drawn.
The perceptual-luminance threshold mis-classified coral (#FF6B6B) as
dark and gave it unreadable white text. Switch to gamma-corrected WCAG
relative luminance with the standard 0.179 crossover: every palette
colour but the dark purple now takes dark glyphs. Test pins all six.
The badge drew white text on the agent colour, which vanished on the
light palette entries (yellow, mint). Pick the glyph + dot colour from
the badge luminance so the name reads on every agent colour.
Above each agent-owned frame, paint a colour-matched pill with the
agent's name and a pulsing status dot, so a running team is legible at
a glance — not just distinguishable by border colour.
A cancelled turn (Stop / New Chat) drops the DesignSession without
hitting the finished path, which left agent indicators registered — a
stale breathing border plus a permanent 30fps redraw loop. Clear them
in Drop so every termination path resets them.
During a multi-screen concurrent generation, every screen group's root
frame is tagged with a distinct agent colour; the canvas now strokes a
pulsing glow + ring in that colour around each one so you can see which
agent is building what. Indicators are registered as the concurrent
scaffold lands and cleared once the turn finishes.
A process-global registry of active per-agent indicators (node borders,
root-frame glow/badge, preview pulses), mirroring the design worker /
paint-pass split. The canvas will read snapshot()/is_active() each frame;
the generation path tags nodes via add_node/add_frame as they stream in.
Each parallel design sub-agent gets a distinct colour (6-colour palette,
cycled by index) and name, so the canvas can later draw per-agent
breathing borders and badges. assign_agent_identities(count) is the
entry point; colours wrap past the palette size.
Drop the bare "panel" trigger keyword so ordinary two-column rows
whose second child merely contains "panel" are no longer rewritten to
full desktop width; require right/insight to signal a real right rail.
Fade the faked drop shadow out (faint far layer) instead of a hard
full-width grey band, and vertically center the example-card emoji
against a two-line title instead of pinning it to the first line.
Replace the flat bg-muted icon row with one rounded bg-foreground/10
chip per connected provider, overlapped 6px with a card-coloured ring
(TS top-bar.tsx: -space-x-1.5 + ring-1 ring-card).
canvas_surface light #fafafa->#e5e5e5, dark #181818->#1a1a1a to equal
the TS CANVAS_BACKGROUND_LIGHT/DARK constants, so the white frame reads
against the canvas instead of vanishing into a near-white surface.
Soften the floating chat panel shadow (shadow-2xl -> shadow-lg), add
top padding before the first message bubble, and give the input text
more room below the divider. Mirrors the native editor panel fix.
Soften the floating panel's drop shadow (two tight low-alpha layers
instead of a heavy offset pair so it no longer stacks into a dark band),
add top breathing room before the first transcript bubble, and drop the
input text below the separator line.
Splits the hit-test half (hit_test / resize_edge_at /
design_block_hover_at / rect_contains) out into a sibling
ai_chat_panel_hit module so the edited file stays under the 800-line
cap; the painting half keeps the spacing fixes.
Add a desktop-dashboard cleanup pass that widens a narrow main-content
row to the desktop width when a no-sidebar generation emits a 1200px
root with a sparse two-column row, preserving the authored intent.
Paint NodeKind::Path nodes that carry an svg_path via a dedicated
svg_path module (fill + stroke, EvenOdd for multi-subpath) instead of
the polyline fallback, and give them proper export corner bounds.