Both of today's image repairs trusted an authored fixed size as the
design's intent. A weak model routinely drops a 400x300 plate into a
390px phone card and CROPS it - which looks right - so the rail repair
widened every card to 400 (one card ate the screen, the rest clipped
away) and the band repair grew every photo band to 300 (a wall of empty
space under each deal). Both now demand that the number be plausible for
the slot it is in: a card is starved only when it is genuinely unusable
and its demand still leaves the rail scrollable, and a band only takes a
photo's height when the result is not taller than the band is wide.
Oversized plates go back to the fixer that shrinks them.
Suppressing the queue left holes in the middle of the skeleton while the
model was working the top - the skeleton-first effect is wanted, only one
shell may look ACTIVE. The on-deck shell (first empty in fill order)
keeps the radar sweep; every queued shell now paints a still wireframe -
outline plus a whisper of wash, no band - so the whole page reads as
planned-but-unwritten instead of half missing, and nothing downstream
pretends to be under the cursor.
A chip sat 6px under the paragraph it followed but 14px above the next
one, so it read as belonging to the story below it rather than the one
it came from; both sides now take the same breathing room. The expander
also becomes the reference's stacked up/down caret pair instead of a
navigational chevron-right.
Five destination cards rendered as five identical blue bands. The photos
were fine and all different - the cards were the problem. Each card's
image slot carried TWO images (an absolutely-positioned image-filled
plate AND the real image node) inside a band authored 56px tall against
a 130px photo, so all that survived the clip was each photo's sky.
The duplicate plate is dropped (the image node wins - it is what the
image pipeline fills and re-searches), the band grows to the photo's own
declared height when it would hide more than half of it, and the starved
rail-card repair now reads fixed-width DESCENDANTS and assigns the
demand as a definite width: an out-of-flow photo plate contributes
nothing to a hug, so fit_content re-starved the very cards it was meant
to rescue (126px against a 200px plate). Replaying the measured document
through the repair loop: 4 geometry diagnostics to 0, cards 79px to
200px.
Three defects from one live run. The tool chips were 48px slabs with a
right-flushed status glyph; the reference reads as a 28px row with the
check inline after the verb. The cursor halo was three concentric
STROKES, whose joins banded into a dirty grey ring on the pencil's
corners - it is now a feathered stack of filled expansions (no joins,
no banding) that doubles as the contact shadow. And the work-order gate
was per-container, so a root's trailing bottom-nav shell lit up and took
the cursor while the model was still filling the header nested inside
the content wrapper; fill order is document order, so the deck is now
global to the generating root.
Design-loop turns used to render as one collapsed 'N tool calls' row
above a monolithic prose wall. The transcript now splits the narration
at each call's stamped content offset and lays prose paragraphs and
headerless verb-chip panels in chronological order, so the story reads
top-to-bottom the way it happened; chips stay visible (nothing folded
away) and each card toggles its own expand override via the original
call index. Plain chat turns keep the grouped panel.
A rail whose fill-width cards resolve skinnier than the fixed-width
content they carry paints clipped image slivers and truncated labels
(measured: a 5-card destination rail sharing one screen width at ~58px
per card around 160px images). The authored fixed content proves the
intended card width, so cards hug it and the rail becomes a clipped
scroller - the clip also keeps the overfull-row flexifier from undoing
the repair next round. Forensic harness gains OP_FORENSIC_FIX=1 to
replay the repair loop on a saved document.
New designs are taught to apply a built-in design system first and build
on $--tokens (own variables reuse the shadcn vocabulary); the preset
tool is covered by a runtime test proving the full table lands, the Mode
axis registers and themed tokens flip between Light and Dark.
The loop's own request path posted raw - a provider rate limit killed
the design run instantly with a raw JSON error, bypassing the adaptive
throttle, the retry ladder and the humanized message that the plain
chat path already had.
A single wide 28% band read as a dirty grey ring - the glow is now three
concentric strokes with falling alpha, approximating a soft blur with
the primitives available.
Preflight headless document loading so malformed or binary archives fail with a clear, actionable error before the MCP server starts.
Keep the CLI parser dependency lightweight by disabling op-pen-loader default features and include the updated lockfile plus regression coverage.
User-picked O4 treatment: a wide translucent slate halo OUTSIDE the
white rim keeps the cursor silhouette readable on light designs while
the white edge stays crisp on dark ones (contact shadow relaxed to
compensate). The default rounded silhouette is additionally fattened
1.28x across the pencil axis with the tip hotspot fixed.
apply_design_system applies one of the bundled presets (full variable
table + Mode theme axis) as a single undoable batch on both the MCP and
design-agent surfaces. The transcript additionally normalizes streamed
narration markdown - emphasis/backtick markers stripped, glued bold
headings re-broken onto their own lines, dash bullets rendered as
bullets - so the design-loop commentary reads as prose instead of
asterisk soup.
Complete shadcn-vocabulary token tables (background/card/primary/muted/
sidebar families plus font and radius tokens) with Light+Dark theme
values, parsed once from a bundled asset and exposed with a
set_variables-ready payload per preset. Groundwork for the
apply-design-system tool and shadcn-aligned variable naming.
The nested-shell pass misfired on a search bar whose children are bare
icon/text leaves (the bar IS the input): its authored background was
stripped, the search glyph painted white and the filter glyph re-inked
with a dangling $color-accent that rendered the fallback blue. The pass
now requires container children and, for the legit nested case, moves
the button's own demoted accent hex onto the glyph instead of a symbolic
ref.
Settings > System gains a pencil-cursor picker - five silhouettes
(classic, rounded, chubby with a pink eraser, crayon, round-nib marker)
drawn as live swatches, rounded as the default, choice persisted in
~/.openpencil/ui.json and restored on launch. Design-loop narration
additionally streams as visible prose instead of folding into the
collapsed thinking area, and every tool call stamps the content offset
where it landed so the transcript can interleave prose with per-call
verb chips.
User-picked variant B: round shoulders at the eraser butt, a curved
collar seam and a soft tip wedge - the quadratic arcs are baked as
dense polygon samples so the existing polygon painters render them
smoothly with no new backend primitive.
Waiting skeleton shells painted their author-given dark fills as odd
slabs below the section being worked. They now keep their layout slot
but paint nothing until they become the on-deck shell (Pencil shows
plain canvas where work has not reached). The agent colour pool also
swaps golden yellow and pale mint for cobalt and emerald - the white
pill label was unreadable on both.
The eye/lock hover glyphs drew straight over long layer names - they
now sit on a locally rebuilt row surface (panel bg + the row's own
wash). The locale globe drops its chevron and paints like its icon
siblings.
A DeepSeek turn was labelled 'Codex CLI' - built-in models no longer
wear a CLI's name (they show their provider group label), and a design
run restamps the streaming bubble with its assigned persona name and
colour, matching the canvas cursor pill and badge.
The size-based exclusion kept the dashboard's main column dark even
while its own subtask was running - content then popped in with no
skeleton at all. Work order alone expresses both measured expectations:
a later region stays plain while an earlier shell is unfilled, and the
big column washes the moment it becomes the first empty shell.
The pinned checklist duplicated what the transcript already renders
inline (action steps, subtask cards, verb chips, agent narration - the
Pencil-style reading flow), while eating a third of the panel's height.
The transcript is now the single progress surface; the checklist
widget, its hit/scroll wiring on both hosts and its ChatState fields
are removed.
The two-column scaffold's Main Content was fit_content and collapsed to
its own padding (a 940x64 strip beside a full-height sidebar) until the
first subtask landed. With the root now presetting a full artboard
height the empty column can safely fill it; the finalize height-adjust
still sizes the root off real content.
DeepSeek V4 sometimes echoes a whole script-gen source a second time
glued mid-line, and the re-declared consts threw 'invalid redefinition'
- a valid section then burned its retries. A new repair rung detects the
first declaration recurring and runs the first copy alone. The
placeholder scan additionally follows work order: within each container
only the first empty shell (the one on deck) washes, so a queued right
column never glows while the sidebar is still being output.
Headless audits now report chrome completeness, node-kind vocabulary and
density alongside the issue counts; scripts/ab-g3 runs the same prompts
through both generation paths and tabulates the rubric for routing
decisions.
The missing/failed-image fallback derives its dash + glyph grey from the
slot's own fill luminance (one fixed grey was invisible on dark
designs); the search-failure sentinel src routes here so a failed
enrichment reads as an intentional placeholder.
CLI provider connections persist (connected flags only) and replay as
silent probes on launch; GUI launches graft the login-shell PATH onto
the process so CLI agents and API keys resolve without a terminal;
provider rate limits ride an adaptive throttle with longer retries and
a human-readable terminal error; quota errors fold into the transcript
as a sentence instead of raw JSON. Design turns keep the artboard in
view (fit on first sized root, refit only when growth pushes it out),
run the structural finalize backstop even when a turn dies early, and
image enrichment fills small and anonymous media slots with junk-free,
session-deduplicated, artifact-word-stripped searches - failures land a
theme-adaptive placeholder instead of a bare grey box.