Adds a self-contained evaluation subsystem in packages/pen-ai-skills:
corpus/ab-v0/*.yaml (24 prompts)
- 8 per category (mobile / dashboard / landing), 4 obvious + 4 optional
- Obvious entries are single-component-scope (one element tool fits 1:1)
and carry expected_tool_if_any; optional entries test composite
prompts where batch_design is the expected path.
src/corpus/
- types.ts — CorpusPrompt / ScoreRow / Report / ApplyFn
- corpus-loader.ts — yaml → CorpusPrompt[] with schema validation
- output-parser.ts — model output → tool_call / batch_design / garbage
tagged union; handles <op_tool> wrapper, multi-tag preference
(element tools win over batch_design scaffolding), malformed-tag
recovery, reasoning-model <think> stripping, and strict cleanDsl
validation mirroring handleBatchDesign's splitter so mixed payloads
don't falsely pass as batch_design
- score-run.ts — M1 (apply + no detector errors) + M3 (shape checks
on roles) + 4-way routing classification (right-tool / wrong-tool
/ fallback / garbage; denominator includes garbage so right-tool
rate stays interpretable)
- aggregate.ts — per-model, per-category, per-tool summaries
- index.ts — public exports
- js-yaml.d.ts — minimal ambient decl (avoids @types/js-yaml churn)
Decoupled from pen-mcp via ApplyFn injection — the scorer never imports
handleBatchDesign; scripts/ab-corpus provides the concrete impl.
61 unit tests cover the corpus pipeline end-to-end; no live API keys
required. Full test suite: 1834/1834.
|
||
|---|---|---|
| .. | ||
| agent-native@e1f90cab96 | ||
| pen-acp | ||
| pen-ai-skills | ||
| pen-core | ||
| pen-engine | ||
| pen-figma | ||
| pen-mcp | ||
| pen-react | ||
| pen-renderer | ||
| pen-sdk | ||
| pen-types | ||
| CLAUDE.md | ||