Files
ThothII/docs/superpowers/plans/2026-06-30-cross-model-behavior-matrix.md
T
marcopanandClaude Opus 4.8 cbb8e184f1 feat(harness): cross-model behavior matrix (G) — harness + results
Tier 1 (clean-room first-turn harness, harness/scripts/model-matrix.mjs):
kickoff + resume chain in-turn on ALL available models — zai/glm-5.2,
deepseek/deepseek-v4-{pro,flash}, aritmolab/qwen3.6-35b-a3b, zai/glm-4.5-air.
The resume cold-start stall recurs on none (closes A's cross-model robustness).
aritmolab/gemma4-26b-a4b is a 404 at the endpoint (listed but not served) — an
availability gap classified as MODEL_ERROR, not a workflow issue.

Tier 2 (live, baseline zai/glm-5.2): F single-select auto-confirm verified
end-to-end — answering the first reviewer_select persisted a concept_clarified
decision (review_decisions.jsonl 0->1) with no follow-up confirmation gate.
Closes F's deferred live check.

No prompt hardening needed. Results in the G plan doc + memory. Throwaway psd
sessions used and deleted; real sessions untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 18:46:55 +02:00

107 lines
6.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Plan — Workstream G: cross-model behavior matrix
Scoping doc (not yet executed). The UI/UX redesign workstreams D, E, B, C, F, A are done and on
`main`. G is the final, separate piece: prove the model-facing workflow is robust across **any**
configured model, not just GLM 5.2 — and close the two live checks deferred into G:
- **F's live check** — does the model actually emit a single-pick `reviewer_select` carrying a
`decision` payload (auto-confirm), with **no** redundant follow-up gate, and the decision landing
in `review_decisions.jsonl`?
- **A's resume robustness** — does each model chain into the resume bootstrap tool calls in-turn
(no narrate-and-stop), as GLM 5.2 now does on pi 0.79.4?
## Why this is its own workstream
A2 already showed model behavior is the variable that matters (the resume stall was a model/pi
artifact, not a code bug). The gate contract (kickoffs, `reviewer_*` widgets, SKILL discipline) is
model-agnostic by design, but only GLM 5.2 has been exercised live. Weaker/non-thinking models are
the realistic failure surface: turn-dropping, narrate-and-stop, stringified-array tool args,
ignoring the auto-confirm decision payload, choosing the wrong widget.
## Models in play (from backend `/models`, 2026-06-30)
| Plan name | id(s) | tier / notes |
|--------------|-----------------------------------------|--------------|
| GLM 5.2 | `zai/glm-5.2` | baseline (verified live); thinking |
| Deepseek V4 | `deepseek/deepseek-v4-pro`, `…-flash` | pro + a cheaper/faster flash |
| Qwen3.6 | `aritmolab/qwen3.6-35b-a3b` | local AritmoLab endpoint; MoE |
| (others) | `zai/glm-4.7`, `glm-5.1`, `gemma4-26b` | breadth / regression coverage |
Prioritise the three named primaries first (GLM 5.2 baseline, Deepseek V4 pro, Qwen3.6), then add
one non-thinking / smaller model (gemma4 or glm-4.5-air) to probe the weak end.
## Method — two tiers (cheap signal first, full run sparingly)
**Tier 1 — clean-room first-turn harness (fast, ~15-30s/run, cheap).** Generalise the A2
`repro-driver.mjs` (the proven clean-room driver: spawn `pi --mode rpc` with the backend env, send
`set_model {provider,modelId}` → `set_thinking_level {level}` → `prompt`, capture raw JSONL,
classify the first turn). Parameterise over `(model, thinking, scenario)` and record:
`CHAINED_INTO_TOOLCALL` vs `STALLED_TURN_ENDED_IDLE`, time-to-first-tool, the first tool name, and
the assistant-text tail. Scenarios that only need the FIRST turn:
- **new-question kickoff** — first tool should be `read` (SKILL.md). (kickoff robustness)
- **resume kickoff** — first tools should be `tht session show` + `read SKILL.md` in-turn. (A)
**Tier 2 — full live run via Playwright (slow, ~3-4 min/F1, run sparingly).** Only for the
end-to-end widget/persistence checks that need a real gate round-trip. Drive ONE canonical session
per model partway through F1 and assert:
- **single-select (F)** — the model emits `reviewer_select` whose chosen option carries a
`decision`; answering it persists ONE decision (`review_decisions.jsonl`) and shows **no**
follow-up confirmation widget; an "Altro"/"back" answer is still handled (non-persisting).
- **multiselect** — a genuinely multi-answer ambiguity uses `reviewer_decide` (checkboxes).
- **thinking vs non-thinking pacing** — note latency / whether non-thinking models skip steps.
## What to record
A matrix `(model × scenario) → {verdict, time-to-first-tool, widget/decision correctness, notes}`.
Capture in PROJECT_STATE + a new `thothii-cross-model-matrix` memory. For every model that
misbehaves, harden the **model-facing prompt** (kickoffs in `tht-gate.js`, discipline in
`SKILL.md`) and re-run — never special-case per model in code; keep the contract uniform.
## Deliverables
1. A reusable matrix harness (promote the A2 driver out of scratchpad into e.g.
`harness/scripts/model-matrix.mjs`, committed).
2. The filled results matrix (PROJECT_STATE + memory).
3. Prompt hardening PRs where a model diverges, each re-verified.
4. A short "supported models" note (which models drive the workflow reliably; which to avoid).
## Results (executed 2026-06-30)
Harness committed at `harness/scripts/model-matrix.mjs`. Throwaway sessions used + deleted; the two
real psd sessions were never driven.
**Tier 1 — kickoff + resume first-turn (medium thinking), all *available* models CHAINED in-turn:**
| model | new (→`read` SKILL) | resume (→`tht session show`+`read`) |
|---|---|---|
| `zai/glm-5.2` | ✅ 10.9s | ✅ 14.5s |
| `deepseek/deepseek-v4-pro` | ✅ 8.9s | ✅ 7.8s |
| `deepseek/deepseek-v4-flash` | ✅ 5.9s | ✅ 6.4s |
| `aritmolab/qwen3.6-35b-a3b` | ✅ 6.5s | ✅ 7.0s |
| `zai/glm-4.5-air` | ✅ 18.3s | ✅ 11.3s |
| `aritmolab/gemma4-26b-a4b` | ⚠️ `MODEL_ERROR` — 404 "model does not exist" at the endpoint (in the registry but not served); not a workflow issue |
→ The kickoff contract is model-agnostic across the available fleet; the **resume cold-start stall
does not recur on any model** (closes A's cross-model robustness). No prompt hardening needed.
**Tier 2 — F single-select auto-confirm, live on baseline `zai/glm-5.2`:** F1 reached the first
`reviewer_select` ("cosa significa 'fibrillazione atriale'?") at ~321s; answering a concrete option
drove `review_decisions.jsonl` **0 → 1** (a full `concept_clarified` decision persisted directly)
with **no follow-up confirmation gate**. → **F's auto-confirm contract verified end-to-end.**
**Supported-models note:** GLM 5.2 (baseline), Deepseek V4 (pro + flash), Qwen3.6 35B, and the
lighter GLM-4.5-air all drive the workflow reliably. `gemma4-26b-a4b` is listed but **not served**
by the AritmoLab endpoint (404) — exclude until the endpoint provides it. Tier-2 per-model F/
multiselect behavior beyond the baseline remains a cheap future add (re-run with each model).
## Risks / notes
- **psd is the active workspace (real client data).** Use throwaway sessions + delete after (the A2
pattern); never drive turns on the user's real sessions.
- **Cost/time:** Tier 1 is cheap; cap Tier 2 to one full F1 per model. Local AritmoLab models
(Qwen3.6, gemma4) may need the AritmoLab endpoint reachable (VPN) and may differ on tool-call
formatting — `prepareReviewerArguments` already parses stringified-array args, a known quirk.
- **Non-thinking models** may need `thinking:"low"`/none; sweep the thinking level as a variable.
- Strictly separate from the shipped workstreams: G changes prompts/docs only, never the gate's
control flow.