Tier 1 (clean-room first-turn harness, harness/scripts/model-matrix.mjs):
kickoff + resume chain in-turn on ALL available models — zai/glm-5.2,
deepseek/deepseek-v4-{pro,flash}, aritmolab/qwen3.6-35b-a3b, zai/glm-4.5-air.
The resume cold-start stall recurs on none (closes A's cross-model robustness).
aritmolab/gemma4-26b-a4b is a 404 at the endpoint (listed but not served) — an
availability gap classified as MODEL_ERROR, not a workflow issue.
Tier 2 (live, baseline zai/glm-5.2): F single-select auto-confirm verified
end-to-end — answering the first reviewer_select persisted a concept_clarified
decision (review_decisions.jsonl 0->1) with no follow-up confirmation gate.
Closes F's deferred live check.
No prompt hardening needed. Results in the G plan doc + memory. Throwaway psd
sessions used and deleted; real sessions untouched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
107 lines
6.5 KiB
Markdown
107 lines
6.5 KiB
Markdown
# Plan — Workstream G: cross-model behavior matrix
|
||
|
||
Scoping doc (not yet executed). The UI/UX redesign workstreams D, E, B, C, F, A are done and on
|
||
`main`. G is the final, separate piece: prove the model-facing workflow is robust across **any**
|
||
configured model, not just GLM 5.2 — and close the two live checks deferred into G:
|
||
|
||
- **F's live check** — does the model actually emit a single-pick `reviewer_select` carrying a
|
||
`decision` payload (auto-confirm), with **no** redundant follow-up gate, and the decision landing
|
||
in `review_decisions.jsonl`?
|
||
- **A's resume robustness** — does each model chain into the resume bootstrap tool calls in-turn
|
||
(no narrate-and-stop), as GLM 5.2 now does on pi 0.79.4?
|
||
|
||
## Why this is its own workstream
|
||
|
||
A2 already showed model behavior is the variable that matters (the resume stall was a model/pi
|
||
artifact, not a code bug). The gate contract (kickoffs, `reviewer_*` widgets, SKILL discipline) is
|
||
model-agnostic by design, but only GLM 5.2 has been exercised live. Weaker/non-thinking models are
|
||
the realistic failure surface: turn-dropping, narrate-and-stop, stringified-array tool args,
|
||
ignoring the auto-confirm decision payload, choosing the wrong widget.
|
||
|
||
## Models in play (from backend `/models`, 2026-06-30)
|
||
|
||
| Plan name | id(s) | tier / notes |
|
||
|--------------|-----------------------------------------|--------------|
|
||
| GLM 5.2 | `zai/glm-5.2` | baseline (verified live); thinking |
|
||
| Deepseek V4 | `deepseek/deepseek-v4-pro`, `…-flash` | pro + a cheaper/faster flash |
|
||
| Qwen3.6 | `aritmolab/qwen3.6-35b-a3b` | local AritmoLab endpoint; MoE |
|
||
| (others) | `zai/glm-4.7`, `glm-5.1`, `gemma4-26b` | breadth / regression coverage |
|
||
|
||
Prioritise the three named primaries first (GLM 5.2 baseline, Deepseek V4 pro, Qwen3.6), then add
|
||
one non-thinking / smaller model (gemma4 or glm-4.5-air) to probe the weak end.
|
||
|
||
## Method — two tiers (cheap signal first, full run sparingly)
|
||
|
||
**Tier 1 — clean-room first-turn harness (fast, ~15-30s/run, cheap).** Generalise the A2
|
||
`repro-driver.mjs` (the proven clean-room driver: spawn `pi --mode rpc` with the backend env, send
|
||
`set_model {provider,modelId}` → `set_thinking_level {level}` → `prompt`, capture raw JSONL,
|
||
classify the first turn). Parameterise over `(model, thinking, scenario)` and record:
|
||
`CHAINED_INTO_TOOLCALL` vs `STALLED_TURN_ENDED_IDLE`, time-to-first-tool, the first tool name, and
|
||
the assistant-text tail. Scenarios that only need the FIRST turn:
|
||
- **new-question kickoff** — first tool should be `read` (SKILL.md). (kickoff robustness)
|
||
- **resume kickoff** — first tools should be `tht session show` + `read SKILL.md` in-turn. (A)
|
||
|
||
**Tier 2 — full live run via Playwright (slow, ~3-4 min/F1, run sparingly).** Only for the
|
||
end-to-end widget/persistence checks that need a real gate round-trip. Drive ONE canonical session
|
||
per model partway through F1 and assert:
|
||
- **single-select (F)** — the model emits `reviewer_select` whose chosen option carries a
|
||
`decision`; answering it persists ONE decision (`review_decisions.jsonl`) and shows **no**
|
||
follow-up confirmation widget; an "Altro"/"back" answer is still handled (non-persisting).
|
||
- **multiselect** — a genuinely multi-answer ambiguity uses `reviewer_decide` (checkboxes).
|
||
- **thinking vs non-thinking pacing** — note latency / whether non-thinking models skip steps.
|
||
|
||
## What to record
|
||
|
||
A matrix `(model × scenario) → {verdict, time-to-first-tool, widget/decision correctness, notes}`.
|
||
Capture in PROJECT_STATE + a new `thothii-cross-model-matrix` memory. For every model that
|
||
misbehaves, harden the **model-facing prompt** (kickoffs in `tht-gate.js`, discipline in
|
||
`SKILL.md`) and re-run — never special-case per model in code; keep the contract uniform.
|
||
|
||
## Deliverables
|
||
|
||
1. A reusable matrix harness (promote the A2 driver out of scratchpad into e.g.
|
||
`harness/scripts/model-matrix.mjs`, committed).
|
||
2. The filled results matrix (PROJECT_STATE + memory).
|
||
3. Prompt hardening PRs where a model diverges, each re-verified.
|
||
4. A short "supported models" note (which models drive the workflow reliably; which to avoid).
|
||
|
||
## Results (executed 2026-06-30)
|
||
|
||
Harness committed at `harness/scripts/model-matrix.mjs`. Throwaway sessions used + deleted; the two
|
||
real psd sessions were never driven.
|
||
|
||
**Tier 1 — kickoff + resume first-turn (medium thinking), all *available* models CHAINED in-turn:**
|
||
|
||
| model | new (→`read` SKILL) | resume (→`tht session show`+`read`) |
|
||
|---|---|---|
|
||
| `zai/glm-5.2` | ✅ 10.9s | ✅ 14.5s |
|
||
| `deepseek/deepseek-v4-pro` | ✅ 8.9s | ✅ 7.8s |
|
||
| `deepseek/deepseek-v4-flash` | ✅ 5.9s | ✅ 6.4s |
|
||
| `aritmolab/qwen3.6-35b-a3b` | ✅ 6.5s | ✅ 7.0s |
|
||
| `zai/glm-4.5-air` | ✅ 18.3s | ✅ 11.3s |
|
||
| `aritmolab/gemma4-26b-a4b` | ⚠️ `MODEL_ERROR` — 404 "model does not exist" at the endpoint (in the registry but not served); not a workflow issue |
|
||
|
||
→ The kickoff contract is model-agnostic across the available fleet; the **resume cold-start stall
|
||
does not recur on any model** (closes A's cross-model robustness). No prompt hardening needed.
|
||
|
||
**Tier 2 — F single-select auto-confirm, live on baseline `zai/glm-5.2`:** F1 reached the first
|
||
`reviewer_select` ("cosa significa 'fibrillazione atriale'?") at ~321s; answering a concrete option
|
||
drove `review_decisions.jsonl` **0 → 1** (a full `concept_clarified` decision persisted directly)
|
||
with **no follow-up confirmation gate**. → **F's auto-confirm contract verified end-to-end.**
|
||
|
||
**Supported-models note:** GLM 5.2 (baseline), Deepseek V4 (pro + flash), Qwen3.6 35B, and the
|
||
lighter GLM-4.5-air all drive the workflow reliably. `gemma4-26b-a4b` is listed but **not served**
|
||
by the AritmoLab endpoint (404) — exclude until the endpoint provides it. Tier-2 per-model F/
|
||
multiselect behavior beyond the baseline remains a cheap future add (re-run with each model).
|
||
|
||
## Risks / notes
|
||
|
||
- **psd is the active workspace (real client data).** Use throwaway sessions + delete after (the A2
|
||
pattern); never drive turns on the user's real sessions.
|
||
- **Cost/time:** Tier 1 is cheap; cap Tier 2 to one full F1 per model. Local AritmoLab models
|
||
(Qwen3.6, gemma4) may need the AritmoLab endpoint reachable (VPN) and may differ on tool-call
|
||
formatting — `prepareReviewerArguments` already parses stringified-array args, a known quirk.
|
||
- **Non-thinking models** may need `thinking:"low"`/none; sweep the thinking level as a variable.
|
||
- Strictly separate from the shipped workstreams: G changes prompts/docs only, never the gate's
|
||
control flow.
|