Deterministic mock-LLM for local unit-level testing of the scorer
+ output-parser + apply pipeline — the same three components real
A/B runs use, minus the network.
mockLlmRaw(prompt, variant) returns the raw string a "well-
behaved" model would emit:
- Treatment on obvious prompt → <op_tool>{...}</op_tool> naming
expected_tool_if_any with empty args
- Baseline → minimal batch_design DSL with a single frame role-
stamped from must_contain_roles[0]
- Optional + no hint → falls through to the baseline path
mockLlmParsed() skips the raw-string round-trip and returns a
ParsedOutput directly for tests pinning a specific kind. Both
respect (promptId, variant) overrides so tests can simulate
garbage / wrong-tool / empty outputs inline.
Integration test loads the real ab-v1 corpus from disk and
exercises every prompt through the full pipeline:
corpus → mockLlmRaw → parseModelOutput → scoreRun → ScoreRow
Verifies all 4 routing outcomes (right-tool / wrong-tool /
fallback / garbage) classify correctly on mocked input. 17 test
cases; corpus sweep runs in ~6ms — fast enough to gate every
PR without slowing CI.
This is the prerequisite for future "real" A/B test runners: if
the harness misclassifies obvious mock inputs, no conclusion
from a real run would be trustworthy.
|
||
|---|---|---|
| .. | ||
| agent-native@e1f90cab96 | ||
| pen-acp | ||
| pen-ai-skills | ||
| pen-core | ||
| pen-engine | ||
| pen-figma | ||
| pen-mcp | ||
| pen-react | ||
| pen-renderer | ||
| pen-sdk | ||
| pen-types | ||
| CLAUDE.md | ||