feat(harness): cross-model behavior matrix (G) — harness + results
Tier 1 (clean-room first-turn harness, harness/scripts/model-matrix.mjs):
kickoff + resume chain in-turn on ALL available models — zai/glm-5.2,
deepseek/deepseek-v4-{pro,flash}, aritmolab/qwen3.6-35b-a3b, zai/glm-4.5-air.
The resume cold-start stall recurs on none (closes A's cross-model robustness).
aritmolab/gemma4-26b-a4b is a 404 at the endpoint (listed but not served) — an
availability gap classified as MODEL_ERROR, not a workflow issue.
Tier 2 (live, baseline zai/glm-5.2): F single-select auto-confirm verified
end-to-end — answering the first reviewer_select persisted a concept_clarified
decision (review_decisions.jsonl 0->1) with no follow-up confirmation gate.
Closes F's deferred live check.
No prompt hardening needed. Results in the G plan doc + memory. Throwaway psd
sessions used and deleted; real sessions untouched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
+12
-5
@@ -122,12 +122,19 @@ all pushed to origin. A (Resume menu + stall diagnosis/hardening) DONE, uncommit
|
|||||||
`tht session show`+`read SKILL.md` in-turn). The earlier narrate-and-stop predates the pi upgrade.
|
`tht session show`+`read SKILL.md` in-turn). The earlier narrate-and-stop predates the pi upgrade.
|
||||||
Defense-in-depth applied: `RIPRENDI_KICKOFF` hardened to force the in-turn tool call (gate test +
|
Defense-in-depth applied: `RIPRENDI_KICKOFF` hardened to force the in-turn tool call (gate test +
|
||||||
live regression 2/2). The cross-model angle (weaker/older models) lives in **G**.
|
live regression 2/2). The cross-model angle (weaker/older models) lives in **G**.
|
||||||
- **G (last, separate):** cross-model behavior matrix (Qwen3.6 / GLM 5.2 / Deepseek V4 / others) —
|
- **G — DONE** (uncommitted): cross-model behavior matrix via a committed clean-room harness
|
||||||
also the home for **F's live check** and **A's cross-model resume robustness**.
|
(`harness/scripts/model-matrix.mjs`). **Tier 1** — kickoff + resume first-turn: all *available*
|
||||||
|
models chain in-turn (`zai/glm-5.2`, `deepseek/deepseek-v4-{pro,flash}`, `aritmolab/qwen3.6-35b-a3b`,
|
||||||
|
`zai/glm-4.5-air`); the resume stall recurs on none (closes A's cross-model robustness).
|
||||||
|
`aritmolab/gemma4-26b-a4b` = **404 unavailable** at the endpoint (listed but not served) — infra
|
||||||
|
gap, not a workflow issue. **Tier 2** — F single-select auto-confirm verified live on `glm-5.2`:
|
||||||
|
answering the first `reviewer_select` persisted a `concept_clarified` decision **0→1** with **no
|
||||||
|
follow-up gate** (closes F's deferred live check). Full results: the G plan doc + memory
|
||||||
|
`thothii-cross-model-matrix`. No prompt hardening needed.
|
||||||
|
|
||||||
**Status:** A-G core work done (D, E, B, C, F, A). **Remaining:** **G** (cross-model matrix,
|
**Status:** **All UI-redesign + resume workstreams done — D, E, B, C, F, A, G.** D/E/B/C/F/A pushed;
|
||||||
incl. F's live single-select check + A's resume-on-weaker-models), plus a one-off manual Playwright
|
**G uncommitted** (harness + docs + PROJECT_STATE). Remaining nice-to-haves: Tier-2 F/multiselect for
|
||||||
kebab→resume pass through the live UI (open item 2).
|
the non-baseline models (cheap re-run), and a one-off manual Playwright kebab→resume pass in the UI.
|
||||||
|
|
||||||
## Live verification + reviewer_select fix (2026-06-30, afternoon)
|
## Live verification + reviewer_select fix (2026-06-30, afternoon)
|
||||||
|
|
||||||
|
|||||||
@@ -65,6 +65,35 @@ misbehaves, harden the **model-facing prompt** (kickoffs in `tht-gate.js`, disci
|
|||||||
3. Prompt hardening PRs where a model diverges, each re-verified.
|
3. Prompt hardening PRs where a model diverges, each re-verified.
|
||||||
4. A short "supported models" note (which models drive the workflow reliably; which to avoid).
|
4. A short "supported models" note (which models drive the workflow reliably; which to avoid).
|
||||||
|
|
||||||
|
## Results (executed 2026-06-30)
|
||||||
|
|
||||||
|
Harness committed at `harness/scripts/model-matrix.mjs`. Throwaway sessions used + deleted; the two
|
||||||
|
real psd sessions were never driven.
|
||||||
|
|
||||||
|
**Tier 1 — kickoff + resume first-turn (medium thinking), all *available* models CHAINED in-turn:**
|
||||||
|
|
||||||
|
| model | new (→`read` SKILL) | resume (→`tht session show`+`read`) |
|
||||||
|
|---|---|---|
|
||||||
|
| `zai/glm-5.2` | ✅ 10.9s | ✅ 14.5s |
|
||||||
|
| `deepseek/deepseek-v4-pro` | ✅ 8.9s | ✅ 7.8s |
|
||||||
|
| `deepseek/deepseek-v4-flash` | ✅ 5.9s | ✅ 6.4s |
|
||||||
|
| `aritmolab/qwen3.6-35b-a3b` | ✅ 6.5s | ✅ 7.0s |
|
||||||
|
| `zai/glm-4.5-air` | ✅ 18.3s | ✅ 11.3s |
|
||||||
|
| `aritmolab/gemma4-26b-a4b` | ⚠️ `MODEL_ERROR` — 404 "model does not exist" at the endpoint (in the registry but not served); not a workflow issue |
|
||||||
|
|
||||||
|
→ The kickoff contract is model-agnostic across the available fleet; the **resume cold-start stall
|
||||||
|
does not recur on any model** (closes A's cross-model robustness). No prompt hardening needed.
|
||||||
|
|
||||||
|
**Tier 2 — F single-select auto-confirm, live on baseline `zai/glm-5.2`:** F1 reached the first
|
||||||
|
`reviewer_select` ("cosa significa 'fibrillazione atriale'?") at ~321s; answering a concrete option
|
||||||
|
drove `review_decisions.jsonl` **0 → 1** (a full `concept_clarified` decision persisted directly)
|
||||||
|
with **no follow-up confirmation gate**. → **F's auto-confirm contract verified end-to-end.**
|
||||||
|
|
||||||
|
**Supported-models note:** GLM 5.2 (baseline), Deepseek V4 (pro + flash), Qwen3.6 35B, and the
|
||||||
|
lighter GLM-4.5-air all drive the workflow reliably. `gemma4-26b-a4b` is listed but **not served**
|
||||||
|
by the AritmoLab endpoint (404) — exclude until the endpoint provides it. Tier-2 per-model F/
|
||||||
|
multiselect behavior beyond the baseline remains a cheap future add (re-run with each model).
|
||||||
|
|
||||||
## Risks / notes
|
## Risks / notes
|
||||||
|
|
||||||
- **psd is the active workspace (real client data).** Use throwaway sessions + delete after (the A2
|
- **psd is the active workspace (real client data).** Use throwaway sessions + delete after (the A2
|
||||||
|
|||||||
@@ -0,0 +1,160 @@
|
|||||||
|
// model-matrix.mjs — workstream G Tier 1: clean-room cross-model behavior matrix.
|
||||||
|
//
|
||||||
|
// Drives `pi --mode rpc` exactly as the backend does (set_model -> set_thinking_level ->
|
||||||
|
// prompt) for each (model x scenario) and classifies the FIRST turn. Decisive signal: a
|
||||||
|
// `toolcall_start`/`tool_execution_start` means the model chained into work in-turn; its
|
||||||
|
// absence + quiescence = the narrate-and-stop stall.
|
||||||
|
//
|
||||||
|
// Scenarios (first-turn only, cheap ~15-30s each):
|
||||||
|
// new -> prompt `/nuova-domanda "kickoff"` (THT_SESSION set => PROVIDED kickoff; first
|
||||||
|
// tool should be `read` of SKILL.md)
|
||||||
|
// resume -> prompt `/riprendi-sessione <id>` (first tools should be `tht session show` +
|
||||||
|
// `read SKILL.md` in-turn) — workstream A robustness across models.
|
||||||
|
//
|
||||||
|
// Usage:
|
||||||
|
// node harness/scripts/model-matrix.mjs <sessionId> [provider/model,...] [thinking] [outDir]
|
||||||
|
// Defaults: the named primaries + one weak/local model; thinking=medium.
|
||||||
|
//
|
||||||
|
// NOTE: spawns real models against the configured workspace. Use a THROWAWAY session and
|
||||||
|
// delete it afterwards (psd is real client data). Kills each child after the first decisive
|
||||||
|
// signal, so mutation is minimal (reads only; no gate is answered).
|
||||||
|
|
||||||
|
import { spawn } from "node:child_process";
|
||||||
|
import fs from "node:fs";
|
||||||
|
import path from "node:path";
|
||||||
|
import { fileURLToPath } from "node:url";
|
||||||
|
|
||||||
|
const harnessDir = path.resolve(path.dirname(fileURLToPath(import.meta.url)), "..");
|
||||||
|
|
||||||
|
const sessionId = process.argv[2];
|
||||||
|
if (!sessionId) {
|
||||||
|
console.error("usage: node model-matrix.mjs <sessionId> [provider/model,...] [thinking] [outDir]");
|
||||||
|
process.exit(2);
|
||||||
|
}
|
||||||
|
const models = (process.argv[3] ??
|
||||||
|
"zai/glm-5.2,deepseek/deepseek-v4-pro,aritmolab/qwen3.6-35b-a3b,aritmolab/gemma4-26b-a4b")
|
||||||
|
.split(",").map((s) => s.trim()).filter(Boolean);
|
||||||
|
const thinking = process.argv[4] ?? "medium";
|
||||||
|
const outDir = process.argv[5] ?? "/tmp/model-matrix";
|
||||||
|
fs.mkdirSync(outDir, { recursive: true });
|
||||||
|
|
||||||
|
const scenarios = ["new", "resume"];
|
||||||
|
const DEADLINE = 90000; // hard cap per cell
|
||||||
|
const QUIESCENCE = 14000; // turn ended idle if this long with no new events and no toolcall
|
||||||
|
const sleep = (n) => new Promise((r) => setTimeout(r, n));
|
||||||
|
|
||||||
|
function runCell(provider, modelId, scenario, rawOut) {
|
||||||
|
return new Promise((resolve) => {
|
||||||
|
const env = {
|
||||||
|
...process.env,
|
||||||
|
THT_SESSION: sessionId,
|
||||||
|
THT_AUTHOR: "matrix@local",
|
||||||
|
PATH: `${harnessDir}/.venv/bin:${process.env.PATH ?? ""}`,
|
||||||
|
};
|
||||||
|
const child = spawn("pi", ["--mode", "rpc"], { cwd: harnessDir, env });
|
||||||
|
const raw = fs.createWriteStream(rawOut);
|
||||||
|
child.stderr.resume();
|
||||||
|
|
||||||
|
const t0 = Date.now();
|
||||||
|
const ms = () => Date.now() - t0;
|
||||||
|
const counts = {};
|
||||||
|
let firstTool = null;
|
||||||
|
let assistantText = "";
|
||||||
|
let modelError = null; // a 404/endpoint error is "model unavailable", not a behavioral stall
|
||||||
|
let lastEventAtMs = 0;
|
||||||
|
let promptAtMs = Infinity;
|
||||||
|
let buf = "";
|
||||||
|
|
||||||
|
child.stdout.on("data", (d) => {
|
||||||
|
raw.write(d);
|
||||||
|
buf += d.toString();
|
||||||
|
let i;
|
||||||
|
while ((i = buf.indexOf("\n")) >= 0) {
|
||||||
|
const line = buf.slice(0, i);
|
||||||
|
buf = buf.slice(i + 1);
|
||||||
|
if (!line.trim()) continue;
|
||||||
|
lastEventAtMs = ms();
|
||||||
|
let e;
|
||||||
|
try { e = JSON.parse(line); } catch { continue; }
|
||||||
|
const type = e.type ?? "?";
|
||||||
|
counts[type] = (counts[type] || 0) + 1;
|
||||||
|
if ((type === "toolcall_start" || type === "tool_execution_start") && !firstTool) {
|
||||||
|
firstTool = {
|
||||||
|
atMs: ms(),
|
||||||
|
name: e.toolName ?? e.name ?? e.tool?.name ?? e.input?.command ?? "?",
|
||||||
|
};
|
||||||
|
}
|
||||||
|
if (type === "message_update") {
|
||||||
|
const delta = e.assistantMessageEvent?.delta;
|
||||||
|
if (typeof delta === "string") assistantText += delta;
|
||||||
|
}
|
||||||
|
if (!modelError) {
|
||||||
|
const m = e.message;
|
||||||
|
if (m && (m.errorMessage || m.stopReason === "error")) {
|
||||||
|
modelError = m.errorMessage || "stopReason=error";
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
function send(obj) { try { child.stdin.write(JSON.stringify(obj) + "\n"); } catch {} }
|
||||||
|
|
||||||
|
function finish(verdict) {
|
||||||
|
clearInterval(timer);
|
||||||
|
try { child.kill("SIGKILL"); } catch {}
|
||||||
|
resolve({
|
||||||
|
provider, modelId, scenario,
|
||||||
|
verdict: modelError ? "MODEL_ERROR" : verdict,
|
||||||
|
firstToolAtMs: firstTool?.atMs ?? null,
|
||||||
|
firstToolName: firstTool?.name ?? null,
|
||||||
|
error: modelError,
|
||||||
|
events: counts,
|
||||||
|
textTail: assistantText.slice(-240).replace(/\s+/g, " ").trim(),
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
(async () => {
|
||||||
|
await sleep(700);
|
||||||
|
send({ type: "set_model", provider, modelId });
|
||||||
|
await sleep(1800);
|
||||||
|
send({ type: "set_thinking_level", level: thinking });
|
||||||
|
await sleep(1000);
|
||||||
|
const message = scenario === "resume"
|
||||||
|
? `/riprendi-sessione ${sessionId}`
|
||||||
|
: `/nuova-domanda "kickoff"`;
|
||||||
|
send({ type: "prompt", message });
|
||||||
|
promptAtMs = ms();
|
||||||
|
})();
|
||||||
|
|
||||||
|
const timer = setInterval(() => {
|
||||||
|
if (modelError) return finish("MODEL_ERROR");
|
||||||
|
if (firstTool) return finish("CHAINED");
|
||||||
|
const sawTurn = lastEventAtMs > promptAtMs && (counts.agent_start || counts.turn_start || counts.message_start);
|
||||||
|
const quiet = ms() - lastEventAtMs > QUIESCENCE;
|
||||||
|
if (sawTurn && quiet) return finish("STALLED");
|
||||||
|
if (ms() > DEADLINE) return finish(sawTurn ? "STALLED" : "NO_TURN");
|
||||||
|
}, 1000);
|
||||||
|
|
||||||
|
child.on("exit", () => { if (!firstTool) finish("CHILD_EXIT"); });
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
const results = [];
|
||||||
|
for (const m of models) {
|
||||||
|
const [provider, modelId] = m.split("/");
|
||||||
|
for (const scenario of scenarios) {
|
||||||
|
const tag = `${provider}_${modelId}_${scenario}`.replace(/[^a-z0-9_.-]/gi, "-");
|
||||||
|
const r = await runCell(provider, modelId, scenario, path.join(outDir, `${tag}.jsonl`));
|
||||||
|
results.push(r);
|
||||||
|
console.error(`done ${m} ${scenario}: ${r.verdict} (${r.firstToolName ?? "-"} @ ${r.firstToolAtMs ?? "-"}ms)`);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Markdown summary table + raw JSON for the record.
|
||||||
|
const rows = results.map((r) =>
|
||||||
|
`| ${r.provider}/${r.modelId} | ${r.scenario} | ${r.verdict} | ${r.firstToolName ?? "-"} | ${r.firstToolAtMs ?? "-"} | ${r.error ?? ""} |`,
|
||||||
|
);
|
||||||
|
console.log("\n| model | scenario | verdict | first tool | t(ms) | error |");
|
||||||
|
console.log("|---|---|---|---|---|---|");
|
||||||
|
console.log(rows.join("\n"));
|
||||||
|
console.log("\nJSON " + JSON.stringify(results));
|
||||||
Reference in New Issue
Block a user