Commit graph

472 commits

Author SHA1 Message Date
Fini 5ffbd1a86a fix(ai): partial JSONL failures hard-rollback so retry actually fires
Previous "failed with non-empty insertedNodes" combination still bypassed
retry. orchestrator-sub-agent.ts gates retry on `result.nodes.length === 0`
— the partial inserts surfaced through DispatchResult.insertedNodes
flowed through to the subtask's `nodes` field, made it look non-empty,
and skipped the retry / minimal-skills / batch_design fallback chain.

Hard-rollback partial inserts on JSONL fallback failure: call
`store.removeNode(id)` for every root that did land, then return
`failed` with `insertedNodes: []`. The dispatcher's surrounding
history-batch wrapper absorbs both the addNode and removeNode calls so
the user-visible undo entry is a net no-op, and the retry condition
upstream now sees a genuinely empty result and re-runs the subtask
cleanly.

Three outcomes after this:
- All roots land → `applied` with full insertedNodes.
- Partial / total failure → `failed` with `insertedNodes: []` (any
  partial successes rolled back) so retry fires and the doc returns to
  its pre-dispatch state.
2026-05-05 01:09:23 +08:00
Fini 4043e9915f fix(ai): JSONL fallback returns 'failed' on partial-insert (was 'applied')
Previous version reported `status: 'applied'` whenever at least one root
landed, with a partial-failure note in `message`. But the orchestrator's
retry / minimal-skills / batch_design-fallback chain checks
`status === 'applied'` to decide whether to bypass retry — a partial
insert (e.g. 1/5 roots landed because `defaultParentId` was stale)
would short-circuit retry and leave the user with a degraded design
that the system never tried to fix.

Now any failed root flips the dispatch to `status: 'failed'` so the
orchestrator's retry path can take over. The successful partial inserts
are still surfaced in `insertedNodes` so the surrounding history-batch
wrapper can roll them back / clean up — `failed` with non-empty
`insertedNodes` is a legitimate combination meaning "side effects
happened but the dispatch did not complete its contract".

Three outcomes now:
- All N roots land → `applied` with full count.
- 1..N-1 land → `failed` with partial-success `insertedNodes` and a
  message naming the parent id + failed root ids.
- 0 land → `failed` with empty `insertedNodes` and the same diagnostic.
2026-05-05 01:04:47 +08:00
Fini f6bc47b347 fix(ai): batch_design JSONL fallback verifies actual insertion
Previous JSONL fallback called `store.addNode(defaultParentId, root)`
and reported `status: 'applied'` regardless of outcome. But `addNode`
returns void and silently no-ops via `insertNodeInTree` when the parent
id can't be resolved (stale `defaultParentId`, empty doc, etc). The
caller would then count the dispatch as a successful insert even
though the doc was unchanged.

Verify each root via `getNodeById(root.id)` immediately after addNode.

Outcomes now:
- All roots land → `applied`, message lists count.
- Some land, some don't → `applied` with partial-failure note in
  message; only the live roots are returned in `insertedNodes`.
- No roots land → `failed` with diagnostic naming the parent id and the
  first few failed root ids — caller surfaces this to the orchestrator
  retry path instead of silently absorbing the loss.
2026-05-05 00:59:41 +08:00
Fini 9731a3c76f fix(ai): batch_design dispatcher accepts JSONL operations as fallback
Mid-tier models (observed: GPT-5.5 standard tier in web-app CLI mode)
correctly emit `<op_tool>{name:"batch_design",arguments:{operations:...}}`
when the brief doesn't fit any embedded element tool — Strategy B in
ELEMENT_TOOL_OUTPUT_FORMAT. But the prompt only declares the operations
value as `<DSL_STRING>` without showing the DSL syntax, so models stuff
flat JSONL (`{"_parent":null,"id":"…","type":"frame",…}`) into the
`operations` field instead of `foo=I("parent",{…})\nbar=U(foo,…)`.

The browser DSL executor then rejects every line ("Cannot parse
operation: …"), all retries fail, and the user sees a degenerate result
(303B / 1 node) despite the model having streamed a full design.

Detect at dispatch time: if `operations` looks like JSONL (starts with
`{` AND contains a `_parent` key or a typed PenNode shape near the top),
route through `parseJsonlToTree` + `store.addNode(defaultParentId, root)`
loop instead of the DSL parser. Same dispatch invariants (single
history batch, dispatch result accounting) apply.

This unblocks the most common Strategy B failure: model emits JSONL
inside a `batch_design` envelope. Strategy A (per-component element
tools) and DSL-shaped Strategy B both still go through their existing
paths unchanged.
2026-05-05 00:54:23 +08:00
Fini 1dd7015c80 fix(pen-core): wrapper detection skips secondary atomics inside primary atomics
The prior pass treated ANY atomic-role frame containing another atomic-
role child with a fill as a wrapper. That's still too aggressive: real
atomic components legitimately compose secondary atomics inside them
(input + trailing icon-button for clear/reveal-password, search-bar +
voice-search icon-button, etc). Stripping the parent's fill in those
cases erases the input/search-bar surface — a regression.

Refine: split atomic protected roles into PRIMARY (input, form-input,
search-bar — input-class components that constitute the "main" atom)
and SECONDARY (button, icon-button, badge, chip, tag, pill — sub-action
or decoration atomics that legitimately nest inside primary atomics).

Wrapper detection now triggers only when:
  - same-role nesting (search-bar > search-bar, input > input), OR
  - PRIMARY atomic nested inside another atomic (search-bar > input —
    the canonical sub-agent misroll).

Two new tests:
- input atom with trailing icon-button (filled clear button) → input
  fill kept
- search-bar atom with voice icon-button (filled accent) → search-bar
  fill kept

Original misroll case (search-bar wrapper > inner input) still strips —
covered by prior test.
2026-05-05 00:33:15 +08:00
Fini e867fcdc50 fix(pen-core): wrapper detection only fires for atomic protected roles
Previous nested-wrapper detection treated any PROTECTED_ROLES frame
containing another protected/structural-fill child as a wrapper. That
swept too widely and could strip fills from real container components:

- card containing a CTA `button` (button is filled, card surface is
  intentional) — card fill stripped if its surface was in SAFE_LIGHT.
- pricing-card with a `badge` ribbon and a CTA button — same issue.
- banner with a nested card — banner fill stripped.

Real component composition is normal; the problem is specifically
sub-agent role mislabels where an ATOMIC component (search-bar, button,
input, badge, chip) is reused as a section wrapper. Container roles
(card, pricing-card, feature-card, banner, etc) NEVER appear as
wrappers — their fill is always intentional.

Fix: introduce ATOMIC_PROTECTED_ROLES (subset of PROTECTED_ROLES) and
restrict wrapper detection to firing only when the OUTER role is in this
atomic set. Container roles stay fully protected.

Three new tests added:
- card with filled button child → card fill kept
- pricing-card with badge + button children → pricing-card fill kept
- banner with nested filled card → banner fill kept

The original misroll case (search-bar > input wrapper) still strips —
covered by the prior test.
2026-05-04 23:27:19 +08:00
Fini b719efd3e8 fix(pen-core): strip safe-light fills on misrolled component-wrapper sections
Real repro from MiniMax-M2.7: sub-agent emits a section wrapper with
the WRONG role applied — Search Bar(role=search-bar) > Search Input
Container(role=input,fill=$color-surface). The outer "search-bar" frame
is actually a section-level wrapper (its child carries the real atom),
but its role is `search-bar` which is in PROTECTED_ROLES, so the strip
pass treated it as the real atom and left its #F8FAFC hedge fill alone.
Result: visible double-cream nesting against the cream root background.

Detect this misroll: a frame whose role IS protected but ALSO contains
a child carrying either the same role or another protected/structural
role with its own solid fill is a wrapper, not the atom — its fill is
eligible for the same safe-light/safe-dark hedge stripping that pure
section frames get.

Counter-case kept covered: a real `search-bar` atom whose children are
just icons / placeholder text (no nested input/search-bar/card/etc with
its own fill) keeps its fill — that fill is intentional, not a hedge.

Two new tests:
- M2.7 misrolled wrapper (search-bar > input + safe-light fill) — outer
  fill stripped, inner input fill preserved.
- Real search-bar atom (no fill-bearing component children) — fill
  preserved.
2026-05-04 23:22:52 +08:00
Fini dfb055eb6a fix(ai-skills): scope JSONL output instructions to fallback-only sections
Codex flagged: even after the previous CRITICAL preamble told the model
to defer to `<op_tool>` mode when an OUTPUT FORMAT block exists later,
the rest of jsonl-format / jsonl-format-simplified still contained
specific JSONL-output directives ("Output a ```json block with ONE node
per line", "FORMAT: _parent (null=root, …)", a full ```json example).
Those specific instructions can dominate over the abstract preamble for
weak models — they read concrete rules and execute them, ignoring the
top-of-skill conditional.

Restructured both skills so JSONL-specific output mechanics are scoped
to a clearly-marked "JSONL FALLBACK MODE" section and the schema
content (TYPES / RULES / DESIGN SYSTEM TOKENS) is mode-agnostic.

- New top-of-skill comment explicitly states TYPES / RULES / TOKENS
  apply to BOTH `<op_tool>` argument shape AND JSONL — neither mode
  contradicts them.
- The "Output ```json block" directive, the "FORMAT: _parent" directive,
  and the ```json example are now wrapped under a "JSONL FALLBACK MODE"
  header that explicitly says "ignore this section if `<op_tool>` mode
  is in effect".

In `<op_tool>` mode the model now reads schema rules without reading
JSONL-specific output mechanics; in JSONL fallback the JSONL section is
unambiguously authoritative. No conflicting instructions for either
output path.
2026-05-04 22:54:37 +08:00
Fini 71f3c04dbf fix(ai): jsonl-format skills coexist with ELEMENT_TOOL_OUTPUT_FORMAT (dual-mode)
The previous fix dropped jsonl-format / jsonl-format-simplified entirely
when elementToolsEnabled was true, on the theory that their CRITICAL
"Output ONLY ```json … Do NOT use tool calls" line conflicted with the
appended `<op_tool>` instruction. But empirically dropping them made
weak-model output WORSE: MiniMax-M2.7 still emits raw JSONL most of the
time (it can't reliably emit `<op_tool>`), and without the JSONL
schema/format teaching its output degrades — role coverage dropped
from 74% to 22%, color-ref% from 84% to 49%.

The right fix is dual-mode coexistence: keep BOTH skills loaded so the
model has the JSONL fallback teaching, but rewrite each skill's CRITICAL
opener to defer to the ELEMENT_TOOL_OUTPUT_FORMAT block when present.

- jsonl-format / jsonl-format-simplified now lead with: "If a separate
  OUTPUT FORMAT — EMIT AS TOOL CALL(S) block appears later in the system
  prompt, FOLLOW THAT block. Use the JSONL form below ONLY when no
  <op_tool> instruction is present."

- Removed the orchestrator-sub-agent.ts skill-filtering branch; both
  skills load unconditionally now.

Net effect: strong models that can follow `<op_tool>` will use the
element-tool path (preserving the n-tools-per-element design intent for
weak-model stability — MiniMax/GLM/Kimi will emit `<op_tool>` when they
can). Weak models that fall back to raw JSONL still get the schema /
sizing / fill / token rules they need to produce coherent output. No
forced choice, no degraded fallback.
2026-05-04 22:47:25 +08:00
Fini 1f2d3c5e1d Merge branch 'v0.8.0' of github.com:ZSeven-W/openpencil into v0.8.0 2026-05-04 21:45:19 +08:00
Kayshen-X a4b7f62e9a Merge feat/rust-ification into v0.8.0 (Step 0 Rust workspace bootstrap)
Step 0 of OP Rust-ification (per kickoff spec v7 FROZEN):
- Cargo workspace at root (members = ["crates/*"], glob)
- 9 skeleton crates: openpencil-app, openpencil-shell-{core,web,native},
  pen-{types,core,engine,codegen,figma}
- rust-toolchain.toml pinned 1.85 (forced from 1.80 → 1.82 → 1.85
  due to crates.io ecosystem edition2024 requirements)
- deny.toml with kickoff §1.2 wasm32 ban invariant
- 2 GitHub Actions: rust-check.yml (3-platform native + cargo-deny)
  and wasm-bundle-check.yml (wasm32 forward + reverse cargo-deny bans)
- vendor/agent submodule → github.com/ZSeven-W/agent-rs
- Bun script wrappers (cargo:check / :test / :wasm-check / :deny)
- README "Rust subsystem" section + Phase boundary note

§1.2 invariants live:
- Forward wasm32 check: shell-web + 5 bucket A crates compile
- Reverse cargo-deny check bans: native + wasm32 both clean
- compile_error guard: shell-native fails wasm32 build with explicit
  message, validated by canary

Step 1+ owns real implementation; Phase 0 docs (snapshot / plan
patches / IPC inventory / parley-taffy matrix / cargo-deny validation)
in openpencil-docs.
2026-05-04 21:00:00 +08:00
Kayshen-X 536ab91d0c chore(workspace): bump CI yaml + README toolchain refs 1.82 → 1.85 2026-05-03 23:50:00 +08:00
Kayshen-X d0e8ca6516 chore(workspace): drop nightly-only rustfmt features (warnings under stable) 2026-05-03 23:45:00 +08:00
Kayshen-X 535a405dab chore(workspace): bump rust-toolchain 1.82 → 1.85 (cargo-deny edition2024 fix)
Phase 2 Gate codex round 1 BLOCK: cargo-deny check fails on
1.82 because wit-bindgen v0.57.1 requires edition2024 manifest
parsing (introduced in Rust 1.85). Bumping to 1.85 unblocks
both `cargo deny check` and `cargo deny --target wasm32 check
bans` — both now exit 0 (advisories ok, bans ok, licenses ok,
sources ok / bans ok).

Also drops `imports_granularity` + `group_imports` from
rustfmt.toml (nightly-only; were emitting warnings under
stable toolchain). Comment preserved to remind future
nightly-pinning to re-enable.

Cargo.lock regenerated under 1.85 (drops the litemap precise
pin from Phase 1 Task 1.4 — no longer needed).

Verification on 1.85:
- cargo build --workspace: PASS
- cargo test --workspace: PASS (skeleton tests)
- cargo clippy --workspace --all-targets -- -D warnings: PASS
- cargo fmt --all -- --check: PASS (with two nightly warnings now removed)
- cargo check --target wasm32 -p {shell-web --no-default --features web | pen-types/-core/-engine/-codegen/-figma}: PASS
- cargo check --target wasm32 -p openpencil-shell-native: FAIL with compile_error guard text (correct)
- cargo deny check: PASS
- cargo deny --target wasm32 check bans: PASS
2026-05-03 23:40:00 +08:00
Kayshen-X 4448f9110b docs(readme): rust subsystem getting-started section 2026-05-03 23:35:00 +08:00
Kayshen-X e690b35716 chore(workspace): bun scripts wrap cargo commands 2026-05-03 23:30:00 +08:00
Kayshen-X 6510a56822 ci(workspace): wasm32 bundle invariant (forward + reverse cargo-deny) 2026-05-03 23:25:00 +08:00
Kayshen-X 5e12e39c99 ci(workspace): native rust-check (fmt + build + test + clippy + deny) 2026-05-03 23:20:00 +08:00
Kayshen-X cebe6614cc chore(vendor): add agent-rs submodule at vendor/agent 2026-05-03 23:15:00 +08:00
Kayshen-X 1bb13d4508 chore(workspace): commit Cargo.lock after skeleton bootstrap 2026-05-03 23:10:00 +08:00
Kayshen-X c54a5facee chore(workspace): cargo-deny 0.18 activation (Phase 1 Task 1.8 Step 6)
- deny.toml: add [graph].targets to limit metadata to native+wasm32
  (avoid Android/iOS edition-2024 deps that fail rustc 1.82 cargo metadata)
- deny.toml: [bans] allow-wildcard-paths = true for workspace path deps
- crates/*/Cargo.toml: add explicit version="0.1.0" alongside path = "..."
  (cargo-deny rejects wildcard-path deps for publishable crates)

cargo-deny 0.16.4 hits a CVSS 4.0 parse error AND lacks edition-2024 cargo
metadata support; bumped to 0.18.9 (installed via stable toolchain). Run
cargo-deny with RUSTUP_TOOLCHAIN=stable so it uses cargo 1.95 for metadata
parsing while project itself still builds on 1.82.

Verified: advisories ok, bans ok, licenses ok, sources ok (exit 0)
on both native and wasm32-unknown-unknown targets.
2026-05-03 23:05:00 +08:00
Kayshen-X 4764be8dc5 style: rustfmt placeholder format! macros (Phase 1 Task 1.8 Step 3) 2026-05-03 23:00:00 +08:00
Kayshen-X 701c7670e2 feat(pen-figma): skeleton crate (bucket A) 2026-05-03 22:55:00 +08:00
Kayshen-X d2554eaa4e feat(pen-codegen): skeleton crate (bucket A) 2026-05-03 22:50:00 +08:00
Kayshen-X aabd681444 feat(pen-engine): skeleton crate (bucket A) 2026-05-03 22:45:00 +08:00
Kayshen-X fbeb66324c feat(pen-core): skeleton crate (bucket A) 2026-05-03 22:40:00 +08:00
Kayshen-X 05da632559 feat(pen-types): skeleton crate (bucket A) 2026-05-03 22:35:00 +08:00
Kayshen-X 2ea9b23b66 chore(workspace): bump rust-toolchain 1.80 → 1.82
Phase 1 batch 3 implementer found 1.80 incompatible with current
crates.io ecosystem: parley → fontique → litemap 0.7.5 needs 1.81;
accesskit chain → indexmap 2.14 → hashbrown 0.17 needs edition2024
(1.85); skia-safe 0.75+ → home 0.5.12 needs 1.88. 1.82 is the sweet
spot that fixes litemap (and matches what Task 0.4 actually probed
with — 1.95).

shell-native dep set deviation (winit only, skia-safe + accesskit
deferred to Step 1 kill-spike when actually used) is documented in
the Phase 1 review trail. compile_error guard for wasm32 still fires
correctly — the load-bearing §1.2 invariant is satisfied.
2026-05-03 22:30:00 +08:00
Kayshen-X 059a7f3d73 feat(openpencil-shell-native): skeleton crate (kickoff §1.2 native-only)
Phase 1 skeleton: declare crate, add compile_error! wasm32 guard so accidental
inclusion in the web bundle fails at compile time (kickoff spec §1.2 invariant).

Native deps intentionally minimal (just winit, no default features). skia-safe /
accesskit / accesskit_winit deferred to Stage F when RenderBackend is actually
implemented. Reason: current top-tier versions of these crates pull transitive
deps (home 0.5.12, litemap 0.7.5, hashbrown 0.17) that require Rust 1.81+ /
edition2024, but our pinned toolchain is 1.80. Pinning to spec versions
(skia-safe=0.74) also fails since 0.74 was never published. Will revisit when
either the toolchain bumps or upstream stabilizes around an MSRV-1.80 line.

Verified:
- cargo build -p openpencil-shell-native        PASS
- cargo test  -p openpencil-shell-native        PASS (1 test)
- cargo check --target wasm32-unknown-unknown -p openpencil-shell-native
  fails with the compile_error! guard text (NOT a winit/skia build error).
2026-05-03 22:25:00 +08:00
Kayshen-X 79b2a766af feat(openpencil-shell-web): skeleton crate (kickoff §1.2 wasm bundle entry) 2026-05-03 22:20:00 +08:00
Kayshen-X 8fc97c0a48 chore(workspace): switch members to crates/* glob (avoid masking when adding crates incrementally) 2026-05-03 22:15:00 +08:00
Kayshen-X 1a6b28698d feat(openpencil-shell-core): skeleton crate (kickoff §1.2 three-crate split) 2026-05-03 22:10:00 +08:00
Kayshen-X fcaf791b18 feat(openpencil-app): skeleton crate (Stage F entry placeholder) 2026-05-03 22:05:00 +08:00
Kayshen-X 5685c7d745 chore(workspace): cargo-deny config (kickoff §1.2 wasm32 bans) 2026-05-03 22:00:00 +08:00
Kayshen-X 025d1763a5 chore(workspace): bootstrap Cargo workspace + toolchain pin 2026-05-03 21:55:00 +08:00
Kayshen-X 4e92f0c250 Merge origin/v0.8.0 into feat/rust-ification 2026-05-03 21:00:00 +08:00
Kayshen-X 5022a2f9e3 docs: wrap MseeP badge in Assessments section above License 2026-04-29 22:05:00 +08:00
Kayshen-X b2a5f4616c docs: move MseeP badge to bottom of all README versions 2026-04-29 22:00:00 +08:00
MseeP.ai 112921c9ea Add MseeP.ai badge to README.md (#124) 2026-04-29 09:50:57 +08:00
Fini 3275314f40 fix(ai-skills): remove JSONL output references from element-tool path
elements.md still carried two stale claims from the P6 spec era when I
incorrectly assumed the apps/web sub-agent always emitted JSONL:

1. Frontmatter comment (line 17): "embedded orchestrator in apps/web
   emits single-shot JSON and cannot call MCP tools — this skill would
   be 1500 tokens of dead weight there, so it stays excluded."

   Wrong now. With VITE_ENABLE_ELEMENT_TOOLS=1 the embedded orchestrator
   sets `hasMcpTools` and the sub-agent DOES emit `<op_tool>` blocks
   parsed by tryParseAllElementToolOutputs and dispatched via
   element-tools-dispatcher.ts. The skill loads in BOTH paths.

2. Theme handling section opener: "This section applies to the MCP
   tool-call path only … the web-app sub-agent JSONL path forbids tool
   calls — there, write $color-* / $type-* refs directly in JSONL".

   Wrong now. With element tools enabled both paths use `<op_tool>` and
   the same `theme: 'system'` advice applies uniformly. The pointer to
   "DESIGN SYSTEM TOKENS in jsonl-format.md" is dead — that skill was
   just dropped from the element-tool path.

Removed the stale comment, rewrote the comment positively to describe
the dual-path loading. Removed the misleading sub-section header so the
"Default to theme: 'system'" rule applies cleanly to every caller.
2026-04-29 09:50:56 +08:00
Fini 49e61eba06 fix(ai): n-tools actually fire on the sub-agent path (drop jsonl-format conflict)
The whole point of the n-tools-per-element design is stability for weak
models in the BUILT-IN AGENT path (MiniMax / GLM / Kimi). But empirically
no element tool was firing on that path — the model emitted raw JSONL
and bypassed every `<op_tool>` strategy.

Root cause: when `elementToolsEnabled` is true, the prompt mixed two
incompatible output-format instructions:

  - jsonl-format / jsonl-format-simplified — early in the prompt, leads
    with `CRITICAL: Output ONLY ```json. Do NOT use [TOOL_CALL] or
    {tool => ...} syntax.`
  - ELEMENT_TOOL_OUTPUT_FORMAT — appended at the end, says `Respond
    with one or more <op_tool> tags, nothing else.`

Weak models anchor on the early CRITICAL ("Do NOT use tool calls"),
read `<op_tool>` as a forbidden tool-call form, and silently fall back
to raw JSONL. Result: every brief on basic tier with element tools
enabled bypassed the whole element-tool surface — which is the opposite
of the design intent.

Fix: when `elementToolsEnabled` is true, drop both `jsonl-format` and
`jsonl-format-simplified` from `resolvedSkills`. ELEMENT_TOOL_OUTPUT_FORMAT
becomes the sole output-format instruction. Content rules (schema /
layout / text-rules / overflow / icon-catalog / elements) stay loaded.

This was P5/ab-v8's blind spot: the test harness either ran on standard
tier (no jsonl-format-simplified swap) or didn't observe `<op_tool>`
emit rate directly, so the conflict masked real-world failure on basic
tier in the built-in agent path.
2026-04-29 09:50:55 +08:00
Fini 9e18f6ebb6 fix(ai): catalog style-guide palette beats AI-invented palette in seed
The planner output frequently contains BOTH `styleGuideName` (catalog
pick) and `styleGuide.palette` (AI's hallucinated palette) — and the
two often disagree. Empirically MiniMax / GLM gravitate to indigo
`#6366F1` for the accent regardless of what catalog snippet they were
just shown: the model picks 'warm-food-mobile-light' (orange catalog),
copies the cream background `#FFF8F0` correctly, then invents
`accent: #6366F1` for `plan.styleGuide.palette`.

The previous seedDocVariablesFromStyleGuide preferred
`plan.styleGuide.palette` first and only fell back to
`plan.selectedStyleGuideContent` when the AI palette was missing — so
the catalog accent was always overridden by the AI's invented one.
Result: every brief seeded indigo, no matter how good the ranking and
catalog match upstream were.

Swap the priority: catalog content (designed by humans, high
confidence) wins; AI-generated palette is the fallback when no catalog
content was attached. The planner's catalog choice is preserved
(`plan.styleGuideName`) so visible UX is unchanged for that signal —
just the COLORS now come from the catalog rather than the model's bias.
2026-04-29 09:50:54 +08:00
Fini db4434a221 fix(ai): cover plural/card variants in Apple Wallet exclusion list
The previous wallet-app exclusion list only had singular forms ('gift
card' not 'gift cards', 'coupon' not 'coupons') and missed common
membership/loyalty card variants. So briefs like 'wallet app for gift
cards' / 'wallet app for coupons and discounts' / 'wallet app for
membership cards' still routed to a fintech style guide despite being
generic Apple-Wallet contexts.

Extracted the exclusion list into APPLE_WALLET_CONTEXT and added:
- gift card → gift card(s) (singular OR plural)
- coupon → coupon | coupons
- membership / membership card(s)
- punch card(s) — restaurant loyalty cards
- stamp card(s) — coffee shop loyalty cards
- vaccination card(s) — pandemic Apple Wallet pass type

Verified with 12 representative briefs:
- All 7 plural/card-variant briefs now fall back to neutrals
- Singular forms (gift card, coupon) keep their existing fallback
- Real fintech briefs (generic wallet app, send money, crypto wallet) keep
  triggering fintech
2026-04-29 09:50:53 +08:00
Fini 7757dfe47e fix(ai): 'wallet app' triggers fintech unless Apple-Wallet pass context
The previous fix removed 'wallet app' from the fintech phrase list to
stop Apple-Wallet-pass briefs from being routed to a fintech style guide.
But that swung too far: bare 'design a wallet app' or 'wallet app to
send money' are common fintech briefs that don't carry a 'crypto'/
'digital'/'payment' modifier and now fell back to generic neutrals.

Hybrid rule: 'wallet app' triggers fintech UNLESS the brief also mentions
an Apple-Wallet-style context word (pass / passes / boarding / ticket /
tickets / ticketing / gift card / coupon / loyalty). Real fintech briefs
that center on a wallet app rarely use any of those words; Apple Wallet
briefs almost always do.

Verified:
- 'design a wallet app' / 'wallet app to send money' / 'wallet app with
  QR code support' → fintech ✓
- 'Apple Wallet app for boarding passes' / 'wallet app pass viewer' /
  'wallet app to store concert tickets' / 'wallet app for loyalty
  cards' → neutral fallback ✓
- crypto/digital/payment wallet, wallet payment(s), wallet connect,
  budget tracker — unchanged ✓
2026-04-29 09:50:52 +08:00
Fini 291e39e653 fix(ai): drop 'wallet app' from fintech triggers (Apple Wallet pass UI)
The previous wallet-pass fix only removed 'wallet pass' but left
'wallet app' in the fintech phrase list. That still routes Apple-Wallet
contexts like "Apple Wallet app for boarding passes", "wallet app pass
viewer", or a bare "wallet app" brief to a fintech style guide — none
of which want banking aesthetics.

Restrict wallet right-side triggers to phrases that are unambiguously
fintech: 'wallet payment(s)' and 'wallet connect'. Real fintech briefs
that center on a wallet almost always qualify it ('crypto wallet app',
'payment wallet flow', 'digital wallet onboarding') and those still
trigger via the left-side modifier list.

Verified:
- Apple Wallet app passes / wallet app pass viewer / generic wallet app
  all fall back to neutrals (no fintech force)
- crypto wallet / crypto wallet app / digital wallet / payment wallet /
  wallet payment(s) / wallet connect all still trigger fintech
- Other fintech (budget tracker, crypto trading) unchanged
2026-04-29 09:50:51 +08:00
Fini 88766e5c61 fix(ai): drop 'wallet pass' from fintech triggers (generic iOS feature)
The previous commit re-added 'wallet pass' to the fintech phrase list
along with 'wallet app' / 'wallet payment' / 'wallet connect'. But
'wallet pass' specifically is the generic Apple Wallet feature for
boarding passes, event tickets, gift cards, and vaccination cards —
none of those are fintech UI briefs and forcing a fintech guide makes
the design come out banking-styled when the user wanted a clean ticket
or boarding-pass layout.

Removed 'pass' from the wallet right-side phrase list. The other three
('wallet app', 'wallet payment', 'wallet connect') are still
unambiguously fintech briefs.

Verified:
- 'Apple Wallet pass for an event ticket' falls back to neutral guides
- 'wallet pass for a boarding pass' falls back to neutrals
- 'crypto wallet', 'wallet app', 'wallet payment' still trigger fintech
2026-04-29 09:50:50 +08:00
Fini 0eb9537891 fix(ai): add contextual phrase matches for finance/dev domain briefs
Removing 'wallet' / 'budget' / 'expense' / 'api' / 'dev' wholesale to
fix generic-UI over-trigger swung the regex too far the other way: real
fintech and developer briefs that legitimately use these words as their
primary signal lost their domain guide.

Apply the same contextual two-word pattern that 'code' uses to bring
them back without re-introducing the over-trigger:

Finance phrases:
- (crypto|digital|payment|hot|cold|hardware|web3) wallet
- wallet (app|pass|payment|connect)
- (budget|expense) (tracker|app|report|management|manager|tracking)

Developer phrases:
- code (editor|review|repo|repository|completion|snippet|base) [kept]
- api (console|platform|portal|docs|documentation|reference|sdk|gateway|playground|key|keys)
- dev (tool|tools|portal|experience|environment|console|platform)
- (developer is already in the unconditional standalone list)

Verified with 15 representative briefs:
- All 8 Codex-flagged regressions (crypto wallet / digital wallet /
  budget tracker / expense tracker / API console / API docs / dev tools
  / developer portal) now hit a fintech or developer guide in top-4.
- 4 generic UI checks (settings menu / expense form / API integration in
  fintech / Apple Wallet pass) still fall back to neutrals or the
  contextually-correct guide instead of forcing a wrong one.
- Food / wellness / modernist briefs unchanged.
2026-04-29 09:50:49 +08:00
Fini 6247255195 fix(ai): drop generic UI/tech words from domain keyword lists
The previous over-correction-recovery commit kept synonyms a bit too
generously and re-introduced over-trigger problems Codex flagged:

- 'menu' would force a food guide on every \"settings menu\" / \"side
  menu\" / \"dropdown menu\" brief.
- 'api' would force a developer guide on every brief that mentions API
  integration (fintech, ecommerce, etc).
- 'dev' would force a developer guide on any tech context.
- 'wallet' would force a fintech guide on Apple Wallet passes / generic
  iOS wallet UI features.
- 'budget' / 'expense' would force a fintech guide on every form that
  tracks costs (project mgmt, travel apps, design feedback).
- 'mint' / 'brass' / 'sage' would force color tags on common English
  phrases (\"mint condition\", \"brass instrument\", \"sage advice\").

Fix: remove all of those from the unconditional domain keyword lists.

'code' is the special case worth preserving — it IS the most-defining
single word for a developer brief — but it has too many non-dev uses
(QR code, promo code, area code, country code) to match unconditionally.
Replaced with a contextual two-word match: 'code' followed immediately
by editor / review / repo / repository / completion / snippet / base
triggers the dev tag. \"QR code\" / \"promo code\" do not.

Verified:
- 5 generic UI/tech briefs no longer force a domain guide (top 4 falls
  back to alphabetical neutrals).
- 'code editor' and 'code review' still match developer-terminal-dark.
- Food / wellness briefs unchanged from prior fix.
2026-04-29 09:50:48 +08:00
Fini e34d491cb4 fix(ai): restore exact-keyword matches lost in word-boundary regex tightening
The previous ranking fix added \b boundaries to fight substring traps
(\"Featured\" → red, \"Healthy\" → wellness in a food category list),
but in the process I dropped several exact domain keywords that were
NOT substring traps and that legitimate briefs use:

- \"code\" (developer brief: \"code editor\", \"VS Code app\") — was
  silently removed; now restored as `\\bcode\\b` so it matches the
  standalone word but still won't trip on \"decoder\" / \"encode\".
- \"health\" / \"healthy\" (wellness brief: \"design a healthy
  lifestyle app\") — was lost; restored as `\\bhealth\\b` /
  `\\bhealthy\\b`. The food-category-list \"Healthy\" still matches
  too, but that's a smaller harm than missing genuine wellness briefs
  — and the rest of the ranking fix (industry tag weight 30, platform
  mismatch -30) keeps mobile food guides above desktop wellness guides
  even when both are tagged.

Also restored derivative forms that earlier substring matches caught
by accident (modern → modernist/contemporary, luxury → luxurious,
brutal → brutalist/brutalism, minimal → minimalist) and broadened
each domain block with common synonyms so we don't regress brief
coverage on real prompts:

- food: + menu, diner, kitchen, dining, eatery, cafe/café
- finance: + trading, wallet, crypto, budget, expense
- developer: + api, engineering, dev (alongside restored code)
- wellness: + wellbeing, spa, gym, exercise, workout (alongside
  restored health/healthy)
- accents: each color block expanded with common synonyms
  (orange→peach/amber/tangerine, blue→navy/sapphire/cobalt,
  green→emerald/sage/mint, gold→golden/brass, red→ruby).

Verified end-to-end with 5 representative briefs:
- Food brief still puts warm-food-mobile-light in top-4
- \"Healthy lifestyle\" wellness brief now picks wellness-green-mobile
- \"code editor\" developer brief picks developer-terminal-dark
- Modernist brand picks ecommerce-modern-light
2026-04-29 09:50:47 +08:00
Fini 511cc5a1e4 fix(ai): style-guide ranking surfaces warm/industry mobile guides correctly
Two ranking bugs were silently sending mobile food/wellness/fintech briefs
to a desktop landing-page palette:

1. Substring tag inference. /red|red/ matched 'Featured', /health/ matched
   'Healthy' (a category in the food brief), so a food prompt picked up a
   spurious 'wellness' tag and a desktop wellness guide jumped above the
   mobile food guide via tag-overlap math. Added \b word boundaries to
   every English keyword in inferTagsFromPrompt; CJK rules unchanged
   because \b doesn't apply.

2. Industry vs style tag weighting + platform mismatch penalty. Each
   matched tag was worth +10 regardless of meaning, and a platform
   mismatch was a tiny -3 vs +0. So a desktop ecommerce-modern guide
   beating mobile warm-food on the same brief was just `clean+modern+
   rounded` overlapping more than `warm-tones+friendly+rounded` while
   the platform penalty was negligible.
   Now: industry tags (warm-tones / wellness / fintech / developer /
   monospace) score 30, generic style tags 10, platform mismatch -30.
   Empirically pushes warm-food-mobile-light to the top of the food
   brief shortlist (verified with the actual expanded prompt that the
   user's MiniMax-M2.7 run logged).

Same fix applies to every brief that was getting "wrong palette" results
because the planner snippets only contain the top-4 ranked guides — if
the right answer falls past 4, the planner literally never sees it and
the model invents its own (default-blue) palette.

Also: jsonl-format-simplified.md (basic-tier sub-agent prompt) now mirrors
jsonl-format.md's design-system-tokens teaching — basic-tier models like
MiniMax-M2.7 currently emit 0% typography refs because the simplified
prompt doesn't mention $type-* refs at all. The expanded simplified
prompt is 3981 chars, well under the bumped budget=1700 (=6800 char cap).
CRITICAL contract moved to top-of-file as the same defense-in-depth
pattern applied earlier to jsonl-format.md.
2026-04-29 09:50:46 +08:00