feat(harness): cross-model behavior matrix (G) — harness + results

Tier 1 (clean-room first-turn harness, harness/scripts/model-matrix.mjs):
kickoff + resume chain in-turn on ALL available models — zai/glm-5.2,
deepseek/deepseek-v4-{pro,flash}, aritmolab/qwen3.6-35b-a3b, zai/glm-4.5-air.
The resume cold-start stall recurs on none (closes A's cross-model robustness).
aritmolab/gemma4-26b-a4b is a 404 at the endpoint (listed but not served) — an
availability gap classified as MODEL_ERROR, not a workflow issue.

Tier 2 (live, baseline zai/glm-5.2): F single-select auto-confirm verified
end-to-end — answering the first reviewer_select persisted a concept_clarified
decision (review_decisions.jsonl 0->1) with no follow-up confirmation gate.
Closes F's deferred live check.

No prompt hardening needed. Results in the G plan doc + memory. Throwaway psd
sessions used and deleted; real sessions untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
2026-06-30 18:46:55 +02:00
co-authored by Claude Opus 4.8
parent e8cdd0037c
commit cbb8e184f1
3 changed files with 201 additions and 5 deletions
@@ -65,6 +65,35 @@ misbehaves, harden the **model-facing prompt** (kickoffs in `tht-gate.js`, disci
3. Prompt hardening PRs where a model diverges, each re-verified.
4. A short "supported models" note (which models drive the workflow reliably; which to avoid).
## Results (executed 2026-06-30)
Harness committed at `harness/scripts/model-matrix.mjs`. Throwaway sessions used + deleted; the two
real psd sessions were never driven.
**Tier 1 — kickoff + resume first-turn (medium thinking), all *available* models CHAINED in-turn:**
| model | new (→`read` SKILL) | resume (→`tht session show`+`read`) |
|---|---|---|
| `zai/glm-5.2` | ✅ 10.9s | ✅ 14.5s |
| `deepseek/deepseek-v4-pro` | ✅ 8.9s | ✅ 7.8s |
| `deepseek/deepseek-v4-flash` | ✅ 5.9s | ✅ 6.4s |
| `aritmolab/qwen3.6-35b-a3b` | ✅ 6.5s | ✅ 7.0s |
| `zai/glm-4.5-air` | ✅ 18.3s | ✅ 11.3s |
| `aritmolab/gemma4-26b-a4b` | ⚠️ `MODEL_ERROR` — 404 "model does not exist" at the endpoint (in the registry but not served); not a workflow issue |
→ The kickoff contract is model-agnostic across the available fleet; the **resume cold-start stall
does not recur on any model** (closes A's cross-model robustness). No prompt hardening needed.
**Tier 2 — F single-select auto-confirm, live on baseline `zai/glm-5.2`:** F1 reached the first
`reviewer_select` ("cosa significa 'fibrillazione atriale'?") at ~321s; answering a concrete option
drove `review_decisions.jsonl` **0 → 1** (a full `concept_clarified` decision persisted directly)
with **no follow-up confirmation gate**. → **F's auto-confirm contract verified end-to-end.**
**Supported-models note:** GLM 5.2 (baseline), Deepseek V4 (pro + flash), Qwen3.6 35B, and the
lighter GLM-4.5-air all drive the workflow reliably. `gemma4-26b-a4b` is listed but **not served**
by the AritmoLab endpoint (404) — exclude until the endpoint provides it. Tier-2 per-model F/
multiselect behavior beyond the baseline remains a cheap future add (re-run with each model).
## Risks / notes
- **psd is the active workspace (real client data).** Use throwaway sessions + delete after (the A2