Files
ThothII/docs/superpowers/plans/2026-06-30-cross-model-behavior-matrix.md
T
marcopanandClaude Opus 4.8 e8cdd0037c docs(plan): scope workstream G (cross-model behavior matrix)
Actionable scoping for the final workstream: a two-tier method (clean-room
first-turn harness generalised from the A2 repro-driver for cheap kickoff/resume
signals; sparing full Playwright F1 runs for the gate round-trip) over the
configured models (GLM 5.2 baseline, Deepseek V4, Qwen3.6, + breadth). Folds in
F's deferred single-select live check and A's resume-on-weaker-models check.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 18:20:46 +02:00

4.8 KiB
Raw Blame History

Plan — Workstream G: cross-model behavior matrix

Scoping doc (not yet executed). The UI/UX redesign workstreams D, E, B, C, F, A are done and on main. G is the final, separate piece: prove the model-facing workflow is robust across any configured model, not just GLM 5.2 — and close the two live checks deferred into G:

  • F's live check — does the model actually emit a single-pick reviewer_select carrying a decision payload (auto-confirm), with no redundant follow-up gate, and the decision landing in review_decisions.jsonl?
  • A's resume robustness — does each model chain into the resume bootstrap tool calls in-turn (no narrate-and-stop), as GLM 5.2 now does on pi 0.79.4?

Why this is its own workstream

A2 already showed model behavior is the variable that matters (the resume stall was a model/pi artifact, not a code bug). The gate contract (kickoffs, reviewer_* widgets, SKILL discipline) is model-agnostic by design, but only GLM 5.2 has been exercised live. Weaker/non-thinking models are the realistic failure surface: turn-dropping, narrate-and-stop, stringified-array tool args, ignoring the auto-confirm decision payload, choosing the wrong widget.

Models in play (from backend /models, 2026-06-30)

Plan name id(s) tier / notes
GLM 5.2 zai/glm-5.2 baseline (verified live); thinking
Deepseek V4 deepseek/deepseek-v4-pro, …-flash pro + a cheaper/faster flash
Qwen3.6 aritmolab/qwen3.6-35b-a3b local AritmoLab endpoint; MoE
(others) zai/glm-4.7, glm-5.1, gemma4-26b breadth / regression coverage

Prioritise the three named primaries first (GLM 5.2 baseline, Deepseek V4 pro, Qwen3.6), then add one non-thinking / smaller model (gemma4 or glm-4.5-air) to probe the weak end.

Method — two tiers (cheap signal first, full run sparingly)

Tier 1 — clean-room first-turn harness (fast, ~15-30s/run, cheap). Generalise the A2 repro-driver.mjs (the proven clean-room driver: spawn pi --mode rpc with the backend env, send set_model {provider,modelId} → set_thinking_level {level} → prompt, capture raw JSONL, classify the first turn). Parameterise over (model, thinking, scenario) and record: CHAINED_INTO_TOOLCALL vs STALLED_TURN_ENDED_IDLE, time-to-first-tool, the first tool name, and the assistant-text tail. Scenarios that only need the FIRST turn:

  • new-question kickoff — first tool should be read (SKILL.md). (kickoff robustness)
  • resume kickoff — first tools should be tht session show + read SKILL.md in-turn. (A)

Tier 2 — full live run via Playwright (slow, ~3-4 min/F1, run sparingly). Only for the end-to-end widget/persistence checks that need a real gate round-trip. Drive ONE canonical session per model partway through F1 and assert:

  • single-select (F) — the model emits reviewer_select whose chosen option carries a decision; answering it persists ONE decision (review_decisions.jsonl) and shows no follow-up confirmation widget; an "Altro"/"back" answer is still handled (non-persisting).
  • multiselect — a genuinely multi-answer ambiguity uses reviewer_decide (checkboxes).
  • thinking vs non-thinking pacing — note latency / whether non-thinking models skip steps.

What to record

A matrix (model × scenario) → {verdict, time-to-first-tool, widget/decision correctness, notes}. Capture in PROJECT_STATE + a new thothii-cross-model-matrix memory. For every model that misbehaves, harden the model-facing prompt (kickoffs in tht-gate.js, discipline in SKILL.md) and re-run — never special-case per model in code; keep the contract uniform.

Deliverables

  1. A reusable matrix harness (promote the A2 driver out of scratchpad into e.g. harness/scripts/model-matrix.mjs, committed).
  2. The filled results matrix (PROJECT_STATE + memory).
  3. Prompt hardening PRs where a model diverges, each re-verified.
  4. A short "supported models" note (which models drive the workflow reliably; which to avoid).

Risks / notes

  • psd is the active workspace (real client data). Use throwaway sessions + delete after (the A2 pattern); never drive turns on the user's real sessions.
  • Cost/time: Tier 1 is cheap; cap Tier 2 to one full F1 per model. Local AritmoLab models (Qwen3.6, gemma4) may need the AritmoLab endpoint reachable (VPN) and may differ on tool-call formatting — prepareReviewerArguments already parses stringified-array args, a known quirk.
  • Non-thinking models may need thinking:"low"/none; sweep the thinking level as a variable.
  • Strictly separate from the shipped workstreams: G changes prompts/docs only, never the gate's control flow.