diff --git a/docs/superpowers/plans/2026-06-30-cross-model-behavior-matrix.md b/docs/superpowers/plans/2026-06-30-cross-model-behavior-matrix.md new file mode 100644 index 00000000..ffb772e4 --- /dev/null +++ b/docs/superpowers/plans/2026-06-30-cross-model-behavior-matrix.md @@ -0,0 +1,77 @@ +# Plan — Workstream G: cross-model behavior matrix + +Scoping doc (not yet executed). The UI/UX redesign workstreams D, E, B, C, F, A are done and on +`main`. G is the final, separate piece: prove the model-facing workflow is robust across **any** +configured model, not just GLM 5.2 — and close the two live checks deferred into G: + +- **F's live check** — does the model actually emit a single-pick `reviewer_select` carrying a + `decision` payload (auto-confirm), with **no** redundant follow-up gate, and the decision landing + in `review_decisions.jsonl`? +- **A's resume robustness** — does each model chain into the resume bootstrap tool calls in-turn + (no narrate-and-stop), as GLM 5.2 now does on pi 0.79.4? + +## Why this is its own workstream + +A2 already showed model behavior is the variable that matters (the resume stall was a model/pi +artifact, not a code bug). The gate contract (kickoffs, `reviewer_*` widgets, SKILL discipline) is +model-agnostic by design, but only GLM 5.2 has been exercised live. Weaker/non-thinking models are +the realistic failure surface: turn-dropping, narrate-and-stop, stringified-array tool args, +ignoring the auto-confirm decision payload, choosing the wrong widget. + +## Models in play (from backend `/models`, 2026-06-30) + +| Plan name | id(s) | tier / notes | +|--------------|-----------------------------------------|--------------| +| GLM 5.2 | `zai/glm-5.2` | baseline (verified live); thinking | +| Deepseek V4 | `deepseek/deepseek-v4-pro`, `…-flash` | pro + a cheaper/faster flash | +| Qwen3.6 | `aritmolab/qwen3.6-35b-a3b` | local AritmoLab endpoint; MoE | +| (others) | `zai/glm-4.7`, `glm-5.1`, `gemma4-26b` | breadth / regression coverage | + +Prioritise the three named primaries first (GLM 5.2 baseline, Deepseek V4 pro, Qwen3.6), then add +one non-thinking / smaller model (gemma4 or glm-4.5-air) to probe the weak end. + +## Method — two tiers (cheap signal first, full run sparingly) + +**Tier 1 — clean-room first-turn harness (fast, ~15-30s/run, cheap).** Generalise the A2 +`repro-driver.mjs` (the proven clean-room driver: spawn `pi --mode rpc` with the backend env, send +`set_model {provider,modelId}` → `set_thinking_level {level}` → `prompt`, capture raw JSONL, +classify the first turn). Parameterise over `(model, thinking, scenario)` and record: +`CHAINED_INTO_TOOLCALL` vs `STALLED_TURN_ENDED_IDLE`, time-to-first-tool, the first tool name, and +the assistant-text tail. Scenarios that only need the FIRST turn: +- **new-question kickoff** — first tool should be `read` (SKILL.md). (kickoff robustness) +- **resume kickoff** — first tools should be `tht session show` + `read SKILL.md` in-turn. (A) + +**Tier 2 — full live run via Playwright (slow, ~3-4 min/F1, run sparingly).** Only for the +end-to-end widget/persistence checks that need a real gate round-trip. Drive ONE canonical session +per model partway through F1 and assert: +- **single-select (F)** — the model emits `reviewer_select` whose chosen option carries a + `decision`; answering it persists ONE decision (`review_decisions.jsonl`) and shows **no** + follow-up confirmation widget; an "Altro"/"back" answer is still handled (non-persisting). +- **multiselect** — a genuinely multi-answer ambiguity uses `reviewer_decide` (checkboxes). +- **thinking vs non-thinking pacing** — note latency / whether non-thinking models skip steps. + +## What to record + +A matrix `(model × scenario) → {verdict, time-to-first-tool, widget/decision correctness, notes}`. +Capture in PROJECT_STATE + a new `thothii-cross-model-matrix` memory. For every model that +misbehaves, harden the **model-facing prompt** (kickoffs in `tht-gate.js`, discipline in +`SKILL.md`) and re-run — never special-case per model in code; keep the contract uniform. + +## Deliverables + +1. A reusable matrix harness (promote the A2 driver out of scratchpad into e.g. + `harness/scripts/model-matrix.mjs`, committed). +2. The filled results matrix (PROJECT_STATE + memory). +3. Prompt hardening PRs where a model diverges, each re-verified. +4. A short "supported models" note (which models drive the workflow reliably; which to avoid). + +## Risks / notes + +- **psd is the active workspace (real client data).** Use throwaway sessions + delete after (the A2 + pattern); never drive turns on the user's real sessions. +- **Cost/time:** Tier 1 is cheap; cap Tier 2 to one full F1 per model. Local AritmoLab models + (Qwen3.6, gemma4) may need the AritmoLab endpoint reachable (VPN) and may differ on tool-call + formatting — `prepareReviewerArguments` already parses stringified-array args, a known quirk. +- **Non-thinking models** may need `thinking:"low"`/none; sweep the thinking level as a variable. +- Strictly separate from the shipped workstreams: G changes prompts/docs only, never the gate's + control flow.