Tier 1 (clean-room first-turn harness, harness/scripts/model-matrix.mjs):
kickoff + resume chain in-turn on ALL available models — zai/glm-5.2,
deepseek/deepseek-v4-{pro,flash}, aritmolab/qwen3.6-35b-a3b, zai/glm-4.5-air.
The resume cold-start stall recurs on none (closes A's cross-model robustness).
aritmolab/gemma4-26b-a4b is a 404 at the endpoint (listed but not served) — an
availability gap classified as MODEL_ERROR, not a workflow issue.
Tier 2 (live, baseline zai/glm-5.2): F single-select auto-confirm verified
end-to-end — answering the first reviewer_select persisted a concept_clarified
decision (review_decisions.jsonl 0->1) with no follow-up confirmation gate.
Closes F's deferred live check.
No prompt hardening needed. Results in the G plan doc + memory. Throwaway psd
sessions used and deleted; real sessions untouched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
6.5 KiB
Plan — Workstream G: cross-model behavior matrix
Scoping doc (not yet executed). The UI/UX redesign workstreams D, E, B, C, F, A are done and on
main. G is the final, separate piece: prove the model-facing workflow is robust across any
configured model, not just GLM 5.2 — and close the two live checks deferred into G:
- F's live check — does the model actually emit a single-pick
reviewer_selectcarrying adecisionpayload (auto-confirm), with no redundant follow-up gate, and the decision landing inreview_decisions.jsonl? - A's resume robustness — does each model chain into the resume bootstrap tool calls in-turn (no narrate-and-stop), as GLM 5.2 now does on pi 0.79.4?
Why this is its own workstream
A2 already showed model behavior is the variable that matters (the resume stall was a model/pi
artifact, not a code bug). The gate contract (kickoffs, reviewer_* widgets, SKILL discipline) is
model-agnostic by design, but only GLM 5.2 has been exercised live. Weaker/non-thinking models are
the realistic failure surface: turn-dropping, narrate-and-stop, stringified-array tool args,
ignoring the auto-confirm decision payload, choosing the wrong widget.
Models in play (from backend /models, 2026-06-30)
| Plan name | id(s) | tier / notes |
|---|---|---|
| GLM 5.2 | zai/glm-5.2 |
baseline (verified live); thinking |
| Deepseek V4 | deepseek/deepseek-v4-pro, …-flash |
pro + a cheaper/faster flash |
| Qwen3.6 | aritmolab/qwen3.6-35b-a3b |
local AritmoLab endpoint; MoE |
| (others) | zai/glm-4.7, glm-5.1, gemma4-26b |
breadth / regression coverage |
Prioritise the three named primaries first (GLM 5.2 baseline, Deepseek V4 pro, Qwen3.6), then add one non-thinking / smaller model (gemma4 or glm-4.5-air) to probe the weak end.
Method — two tiers (cheap signal first, full run sparingly)
Tier 1 — clean-room first-turn harness (fast, ~15-30s/run, cheap). Generalise the A2
repro-driver.mjs (the proven clean-room driver: spawn pi --mode rpc with the backend env, send
set_model {provider,modelId} → set_thinking_level {level} → prompt, capture raw JSONL,
classify the first turn). Parameterise over (model, thinking, scenario) and record:
CHAINED_INTO_TOOLCALL vs STALLED_TURN_ENDED_IDLE, time-to-first-tool, the first tool name, and
the assistant-text tail. Scenarios that only need the FIRST turn:
- new-question kickoff — first tool should be
read(SKILL.md). (kickoff robustness) - resume kickoff — first tools should be
tht session show+read SKILL.mdin-turn. (A)
Tier 2 — full live run via Playwright (slow, ~3-4 min/F1, run sparingly). Only for the end-to-end widget/persistence checks that need a real gate round-trip. Drive ONE canonical session per model partway through F1 and assert:
- single-select (F) — the model emits
reviewer_selectwhose chosen option carries adecision; answering it persists ONE decision (review_decisions.jsonl) and shows no follow-up confirmation widget; an "Altro"/"back" answer is still handled (non-persisting). - multiselect — a genuinely multi-answer ambiguity uses
reviewer_decide(checkboxes). - thinking vs non-thinking pacing — note latency / whether non-thinking models skip steps.
What to record
A matrix (model × scenario) → {verdict, time-to-first-tool, widget/decision correctness, notes}.
Capture in PROJECT_STATE + a new thothii-cross-model-matrix memory. For every model that
misbehaves, harden the model-facing prompt (kickoffs in tht-gate.js, discipline in
SKILL.md) and re-run — never special-case per model in code; keep the contract uniform.
Deliverables
- A reusable matrix harness (promote the A2 driver out of scratchpad into e.g.
harness/scripts/model-matrix.mjs, committed). - The filled results matrix (PROJECT_STATE + memory).
- Prompt hardening PRs where a model diverges, each re-verified.
- A short "supported models" note (which models drive the workflow reliably; which to avoid).
Results (executed 2026-06-30)
Harness committed at harness/scripts/model-matrix.mjs. Throwaway sessions used + deleted; the two
real psd sessions were never driven.
Tier 1 — kickoff + resume first-turn (medium thinking), all available models CHAINED in-turn:
| model | new (→read SKILL) |
resume (→tht session show+read) |
|---|---|---|
zai/glm-5.2 |
✅ 10.9s | ✅ 14.5s |
deepseek/deepseek-v4-pro |
✅ 8.9s | ✅ 7.8s |
deepseek/deepseek-v4-flash |
✅ 5.9s | ✅ 6.4s |
aritmolab/qwen3.6-35b-a3b |
✅ 6.5s | ✅ 7.0s |
zai/glm-4.5-air |
✅ 18.3s | ✅ 11.3s |
aritmolab/gemma4-26b-a4b |
⚠️ MODEL_ERROR — 404 "model does not exist" at the endpoint (in the registry but not served); not a workflow issue |
→ The kickoff contract is model-agnostic across the available fleet; the resume cold-start stall does not recur on any model (closes A's cross-model robustness). No prompt hardening needed.
Tier 2 — F single-select auto-confirm, live on baseline zai/glm-5.2: F1 reached the first
reviewer_select ("cosa significa 'fibrillazione atriale'?") at ~321s; answering a concrete option
drove review_decisions.jsonl 0 → 1 (a full concept_clarified decision persisted directly)
with no follow-up confirmation gate. → F's auto-confirm contract verified end-to-end.
Supported-models note: GLM 5.2 (baseline), Deepseek V4 (pro + flash), Qwen3.6 35B, and the
lighter GLM-4.5-air all drive the workflow reliably. gemma4-26b-a4b is listed but not served
by the AritmoLab endpoint (404) — exclude until the endpoint provides it. Tier-2 per-model F/
multiselect behavior beyond the baseline remains a cheap future add (re-run with each model).
Risks / notes
- psd is the active workspace (real client data). Use throwaway sessions + delete after (the A2 pattern); never drive turns on the user's real sessions.
- Cost/time: Tier 1 is cheap; cap Tier 2 to one full F1 per model. Local AritmoLab models
(Qwen3.6, gemma4) may need the AritmoLab endpoint reachable (VPN) and may differ on tool-call
formatting —
prepareReviewerArgumentsalready parses stringified-array args, a known quirk. - Non-thinking models may need
thinking:"low"/none; sweep the thinking level as a variable. - Strictly separate from the shipped workstreams: G changes prompts/docs only, never the gate's control flow.