Per-tool handler tests cover "this tool emits the correct shape" individually. This file covers the next layer: can N element-tool calls chain together into a realistic multi-section screen without breaking tree invariants? Scenarios (each spans multiple tool families to catch cross- family regressions): 1. Mobile settings — top_nav + 2 sections × 3 list_rows + bottom_nav (10 calls) 2. Dashboard home — top_nav + stat_grid + section + 3 metric_comparisons + chart (7) 3. Login form — heading + body + 2 form_fields + button + link (6) 4. Profile + UGC — top_nav + avatar + heading + badge + 2 faq_items + action_menu (7) 5. Listing — search + card_row + divider + empty_chart + date_picker + chip_input + pagination (7) 6. parent_id threading invariant — nested insert actually lands under named parent Each scenario asserts: - Every call emits a nodeId (no silent no-ops) - Final document parses as valid JSON with expected root children count - Every call's nodeId is findable in the saved tree - Every tool's canonical role survives post-save - parent_id threading works (child lands under named parent, not root) This is the integration gate that catches "tool wiring works individually but composes wrong" — the ghost regression that can slip past per-tool tests. |
||
|---|---|---|
| .. | ||
| agent-native@e1f90cab96 | ||
| pen-acp | ||
| pen-ai-skills | ||
| pen-core | ||
| pen-engine | ||
| pen-figma | ||
| pen-mcp | ||
| pen-react | ||
| pen-renderer | ||
| pen-sdk | ||
| pen-types | ||
| CLAUDE.md | ||