ab-v4 raw output capture on dashboard-search-filters-composite shows minimax-m2.7 emitting <think>...</think> + 4 op_tool tags that fit inside the 4096 default — its measured completion-token average for this run was 697, well under the cap. So thinking-budget truncation is NOT the actual root cause of minimax's lower multi-tool hit rate (25% vs gpt+deepseek 50%); the real issues are instruction-following (mixed Strategy A + B despite the explicit forbidance, invented "canvas" parent_id placeholder). Still doubling the cap defensively: composite multi-tool outputs can chain 12-13 op_tool tags + thinking, and "fit easy" today doesn't mean "fits headroom-free on a longer brief tomorrow." The bump is free on the happy path (provider stops generating when done, doesn't bill unused headroom) and only ever helps when the model would otherwise hit a real ceiling. Real follow-up for minimax: instruction compliance — the no-mix rule needs to land harder than a single trailing sentence. Probably wants the rule moved to top-of-prompt + a few-shot bad-example contrast. Out of scope here; tracked under Phase 2 prompt design. |
||
|---|---|---|
| .. | ||
| ab-corpus | ||
| bundle-skill.ts | ||
| ensure-agent-native.cjs | ||
| patch-srvx-bun.ts | ||
| publish-beta.sh | ||
| unpublish.sh | ||