Files
ThothII/docs/superpowers/plans/2026-06-30-cross-model-behavior-matrix.md
T
marcopanandClaude Opus 4.8 e8cdd0037c docs(plan): scope workstream G (cross-model behavior matrix)
Actionable scoping for the final workstream: a two-tier method (clean-room
first-turn harness generalised from the A2 repro-driver for cheap kickoff/resume
signals; sparing full Playwright F1 runs for the gate round-trip) over the
configured models (GLM 5.2 baseline, Deepseek V4, Qwen3.6, + breadth). Folds in
F's deferred single-select live check and A's resume-on-weaker-models check.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 18:20:46 +02:00

78 lines
4.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Plan — Workstream G: cross-model behavior matrix
Scoping doc (not yet executed). The UI/UX redesign workstreams D, E, B, C, F, A are done and on
`main`. G is the final, separate piece: prove the model-facing workflow is robust across **any**
configured model, not just GLM 5.2 — and close the two live checks deferred into G:
- **F's live check** — does the model actually emit a single-pick `reviewer_select` carrying a
`decision` payload (auto-confirm), with **no** redundant follow-up gate, and the decision landing
in `review_decisions.jsonl`?
- **A's resume robustness** — does each model chain into the resume bootstrap tool calls in-turn
(no narrate-and-stop), as GLM 5.2 now does on pi 0.79.4?
## Why this is its own workstream
A2 already showed model behavior is the variable that matters (the resume stall was a model/pi
artifact, not a code bug). The gate contract (kickoffs, `reviewer_*` widgets, SKILL discipline) is
model-agnostic by design, but only GLM 5.2 has been exercised live. Weaker/non-thinking models are
the realistic failure surface: turn-dropping, narrate-and-stop, stringified-array tool args,
ignoring the auto-confirm decision payload, choosing the wrong widget.
## Models in play (from backend `/models`, 2026-06-30)
| Plan name | id(s) | tier / notes |
|--------------|-----------------------------------------|--------------|
| GLM 5.2 | `zai/glm-5.2` | baseline (verified live); thinking |
| Deepseek V4 | `deepseek/deepseek-v4-pro`, `…-flash` | pro + a cheaper/faster flash |
| Qwen3.6 | `aritmolab/qwen3.6-35b-a3b` | local AritmoLab endpoint; MoE |
| (others) | `zai/glm-4.7`, `glm-5.1`, `gemma4-26b` | breadth / regression coverage |
Prioritise the three named primaries first (GLM 5.2 baseline, Deepseek V4 pro, Qwen3.6), then add
one non-thinking / smaller model (gemma4 or glm-4.5-air) to probe the weak end.
## Method — two tiers (cheap signal first, full run sparingly)
**Tier 1 — clean-room first-turn harness (fast, ~15-30s/run, cheap).** Generalise the A2
`repro-driver.mjs` (the proven clean-room driver: spawn `pi --mode rpc` with the backend env, send
`set_model {provider,modelId}` → `set_thinking_level {level}` → `prompt`, capture raw JSONL,
classify the first turn). Parameterise over `(model, thinking, scenario)` and record:
`CHAINED_INTO_TOOLCALL` vs `STALLED_TURN_ENDED_IDLE`, time-to-first-tool, the first tool name, and
the assistant-text tail. Scenarios that only need the FIRST turn:
- **new-question kickoff** — first tool should be `read` (SKILL.md). (kickoff robustness)
- **resume kickoff** — first tools should be `tht session show` + `read SKILL.md` in-turn. (A)
**Tier 2 — full live run via Playwright (slow, ~3-4 min/F1, run sparingly).** Only for the
end-to-end widget/persistence checks that need a real gate round-trip. Drive ONE canonical session
per model partway through F1 and assert:
- **single-select (F)** — the model emits `reviewer_select` whose chosen option carries a
`decision`; answering it persists ONE decision (`review_decisions.jsonl`) and shows **no**
follow-up confirmation widget; an "Altro"/"back" answer is still handled (non-persisting).
- **multiselect** — a genuinely multi-answer ambiguity uses `reviewer_decide` (checkboxes).
- **thinking vs non-thinking pacing** — note latency / whether non-thinking models skip steps.
## What to record
A matrix `(model × scenario) → {verdict, time-to-first-tool, widget/decision correctness, notes}`.
Capture in PROJECT_STATE + a new `thothii-cross-model-matrix` memory. For every model that
misbehaves, harden the **model-facing prompt** (kickoffs in `tht-gate.js`, discipline in
`SKILL.md`) and re-run — never special-case per model in code; keep the contract uniform.
## Deliverables
1. A reusable matrix harness (promote the A2 driver out of scratchpad into e.g.
`harness/scripts/model-matrix.mjs`, committed).
2. The filled results matrix (PROJECT_STATE + memory).
3. Prompt hardening PRs where a model diverges, each re-verified.
4. A short "supported models" note (which models drive the workflow reliably; which to avoid).
## Risks / notes
- **psd is the active workspace (real client data).** Use throwaway sessions + delete after (the A2
pattern); never drive turns on the user's real sessions.
- **Cost/time:** Tier 1 is cheap; cap Tier 2 to one full F1 per model. Local AritmoLab models
(Qwen3.6, gemma4) may need the AritmoLab endpoint reachable (VPN) and may differ on tool-call
formatting — `prepareReviewerArguments` already parses stringified-array args, a known quirk.
- **Non-thinking models** may need `thinking:"low"`/none; sweep the thinking level as a variable.
- Strictly separate from the shipped workstreams: G changes prompts/docs only, never the gate's
control flow.