diff --git a/PROJECT_STATE.md b/PROJECT_STATE.md index 51981af8..cd2c048e 100644 --- a/PROJECT_STATE.md +++ b/PROJECT_STATE.md @@ -122,12 +122,19 @@ all pushed to origin. A (Resume menu + stall diagnosis/hardening) DONE, uncommit `tht session show`+`read SKILL.md` in-turn). The earlier narrate-and-stop predates the pi upgrade. Defense-in-depth applied: `RIPRENDI_KICKOFF` hardened to force the in-turn tool call (gate test + live regression 2/2). The cross-model angle (weaker/older models) lives in **G**. -- **G (last, separate):** cross-model behavior matrix (Qwen3.6 / GLM 5.2 / Deepseek V4 / others) — - also the home for **F's live check** and **A's cross-model resume robustness**. +- **G — DONE** (uncommitted): cross-model behavior matrix via a committed clean-room harness + (`harness/scripts/model-matrix.mjs`). **Tier 1** — kickoff + resume first-turn: all *available* + models chain in-turn (`zai/glm-5.2`, `deepseek/deepseek-v4-{pro,flash}`, `aritmolab/qwen3.6-35b-a3b`, + `zai/glm-4.5-air`); the resume stall recurs on none (closes A's cross-model robustness). + `aritmolab/gemma4-26b-a4b` = **404 unavailable** at the endpoint (listed but not served) — infra + gap, not a workflow issue. **Tier 2** — F single-select auto-confirm verified live on `glm-5.2`: + answering the first `reviewer_select` persisted a `concept_clarified` decision **0→1** with **no + follow-up gate** (closes F's deferred live check). Full results: the G plan doc + memory + `thothii-cross-model-matrix`. No prompt hardening needed. -**Status:** A-G core work done (D, E, B, C, F, A). **Remaining:** **G** (cross-model matrix, -incl. F's live single-select check + A's resume-on-weaker-models), plus a one-off manual Playwright -kebab→resume pass through the live UI (open item 2). +**Status:** **All UI-redesign + resume workstreams done — D, E, B, C, F, A, G.** D/E/B/C/F/A pushed; +**G uncommitted** (harness + docs + PROJECT_STATE). Remaining nice-to-haves: Tier-2 F/multiselect for +the non-baseline models (cheap re-run), and a one-off manual Playwright kebab→resume pass in the UI. ## Live verification + reviewer_select fix (2026-06-30, afternoon) diff --git a/docs/superpowers/plans/2026-06-30-cross-model-behavior-matrix.md b/docs/superpowers/plans/2026-06-30-cross-model-behavior-matrix.md index ffb772e4..e312bf7f 100644 --- a/docs/superpowers/plans/2026-06-30-cross-model-behavior-matrix.md +++ b/docs/superpowers/plans/2026-06-30-cross-model-behavior-matrix.md @@ -65,6 +65,35 @@ misbehaves, harden the **model-facing prompt** (kickoffs in `tht-gate.js`, disci 3. Prompt hardening PRs where a model diverges, each re-verified. 4. A short "supported models" note (which models drive the workflow reliably; which to avoid). +## Results (executed 2026-06-30) + +Harness committed at `harness/scripts/model-matrix.mjs`. Throwaway sessions used + deleted; the two +real psd sessions were never driven. + +**Tier 1 — kickoff + resume first-turn (medium thinking), all *available* models CHAINED in-turn:** + +| model | new (→`read` SKILL) | resume (→`tht session show`+`read`) | +|---|---|---| +| `zai/glm-5.2` | ✅ 10.9s | ✅ 14.5s | +| `deepseek/deepseek-v4-pro` | ✅ 8.9s | ✅ 7.8s | +| `deepseek/deepseek-v4-flash` | ✅ 5.9s | ✅ 6.4s | +| `aritmolab/qwen3.6-35b-a3b` | ✅ 6.5s | ✅ 7.0s | +| `zai/glm-4.5-air` | ✅ 18.3s | ✅ 11.3s | +| `aritmolab/gemma4-26b-a4b` | ⚠️ `MODEL_ERROR` — 404 "model does not exist" at the endpoint (in the registry but not served); not a workflow issue | + +→ The kickoff contract is model-agnostic across the available fleet; the **resume cold-start stall +does not recur on any model** (closes A's cross-model robustness). No prompt hardening needed. + +**Tier 2 — F single-select auto-confirm, live on baseline `zai/glm-5.2`:** F1 reached the first +`reviewer_select` ("cosa significa 'fibrillazione atriale'?") at ~321s; answering a concrete option +drove `review_decisions.jsonl` **0 → 1** (a full `concept_clarified` decision persisted directly) +with **no follow-up confirmation gate**. → **F's auto-confirm contract verified end-to-end.** + +**Supported-models note:** GLM 5.2 (baseline), Deepseek V4 (pro + flash), Qwen3.6 35B, and the +lighter GLM-4.5-air all drive the workflow reliably. `gemma4-26b-a4b` is listed but **not served** +by the AritmoLab endpoint (404) — exclude until the endpoint provides it. Tier-2 per-model F/ +multiselect behavior beyond the baseline remains a cheap future add (re-run with each model). + ## Risks / notes - **psd is the active workspace (real client data).** Use throwaway sessions + delete after (the A2 diff --git a/harness/scripts/model-matrix.mjs b/harness/scripts/model-matrix.mjs new file mode 100644 index 00000000..f0f7c0f1 --- /dev/null +++ b/harness/scripts/model-matrix.mjs @@ -0,0 +1,160 @@ +// model-matrix.mjs — workstream G Tier 1: clean-room cross-model behavior matrix. +// +// Drives `pi --mode rpc` exactly as the backend does (set_model -> set_thinking_level -> +// prompt) for each (model x scenario) and classifies the FIRST turn. Decisive signal: a +// `toolcall_start`/`tool_execution_start` means the model chained into work in-turn; its +// absence + quiescence = the narrate-and-stop stall. +// +// Scenarios (first-turn only, cheap ~15-30s each): +// new -> prompt `/nuova-domanda "kickoff"` (THT_SESSION set => PROVIDED kickoff; first +// tool should be `read` of SKILL.md) +// resume -> prompt `/riprendi-sessione ` (first tools should be `tht session show` + +// `read SKILL.md` in-turn) — workstream A robustness across models. +// +// Usage: +// node harness/scripts/model-matrix.mjs [provider/model,...] [thinking] [outDir] +// Defaults: the named primaries + one weak/local model; thinking=medium. +// +// NOTE: spawns real models against the configured workspace. Use a THROWAWAY session and +// delete it afterwards (psd is real client data). Kills each child after the first decisive +// signal, so mutation is minimal (reads only; no gate is answered). + +import { spawn } from "node:child_process"; +import fs from "node:fs"; +import path from "node:path"; +import { fileURLToPath } from "node:url"; + +const harnessDir = path.resolve(path.dirname(fileURLToPath(import.meta.url)), ".."); + +const sessionId = process.argv[2]; +if (!sessionId) { + console.error("usage: node model-matrix.mjs [provider/model,...] [thinking] [outDir]"); + process.exit(2); +} +const models = (process.argv[3] ?? + "zai/glm-5.2,deepseek/deepseek-v4-pro,aritmolab/qwen3.6-35b-a3b,aritmolab/gemma4-26b-a4b") + .split(",").map((s) => s.trim()).filter(Boolean); +const thinking = process.argv[4] ?? "medium"; +const outDir = process.argv[5] ?? "/tmp/model-matrix"; +fs.mkdirSync(outDir, { recursive: true }); + +const scenarios = ["new", "resume"]; +const DEADLINE = 90000; // hard cap per cell +const QUIESCENCE = 14000; // turn ended idle if this long with no new events and no toolcall +const sleep = (n) => new Promise((r) => setTimeout(r, n)); + +function runCell(provider, modelId, scenario, rawOut) { + return new Promise((resolve) => { + const env = { + ...process.env, + THT_SESSION: sessionId, + THT_AUTHOR: "matrix@local", + PATH: `${harnessDir}/.venv/bin:${process.env.PATH ?? ""}`, + }; + const child = spawn("pi", ["--mode", "rpc"], { cwd: harnessDir, env }); + const raw = fs.createWriteStream(rawOut); + child.stderr.resume(); + + const t0 = Date.now(); + const ms = () => Date.now() - t0; + const counts = {}; + let firstTool = null; + let assistantText = ""; + let modelError = null; // a 404/endpoint error is "model unavailable", not a behavioral stall + let lastEventAtMs = 0; + let promptAtMs = Infinity; + let buf = ""; + + child.stdout.on("data", (d) => { + raw.write(d); + buf += d.toString(); + let i; + while ((i = buf.indexOf("\n")) >= 0) { + const line = buf.slice(0, i); + buf = buf.slice(i + 1); + if (!line.trim()) continue; + lastEventAtMs = ms(); + let e; + try { e = JSON.parse(line); } catch { continue; } + const type = e.type ?? "?"; + counts[type] = (counts[type] || 0) + 1; + if ((type === "toolcall_start" || type === "tool_execution_start") && !firstTool) { + firstTool = { + atMs: ms(), + name: e.toolName ?? e.name ?? e.tool?.name ?? e.input?.command ?? "?", + }; + } + if (type === "message_update") { + const delta = e.assistantMessageEvent?.delta; + if (typeof delta === "string") assistantText += delta; + } + if (!modelError) { + const m = e.message; + if (m && (m.errorMessage || m.stopReason === "error")) { + modelError = m.errorMessage || "stopReason=error"; + } + } + } + }); + + function send(obj) { try { child.stdin.write(JSON.stringify(obj) + "\n"); } catch {} } + + function finish(verdict) { + clearInterval(timer); + try { child.kill("SIGKILL"); } catch {} + resolve({ + provider, modelId, scenario, + verdict: modelError ? "MODEL_ERROR" : verdict, + firstToolAtMs: firstTool?.atMs ?? null, + firstToolName: firstTool?.name ?? null, + error: modelError, + events: counts, + textTail: assistantText.slice(-240).replace(/\s+/g, " ").trim(), + }); + } + + (async () => { + await sleep(700); + send({ type: "set_model", provider, modelId }); + await sleep(1800); + send({ type: "set_thinking_level", level: thinking }); + await sleep(1000); + const message = scenario === "resume" + ? `/riprendi-sessione ${sessionId}` + : `/nuova-domanda "kickoff"`; + send({ type: "prompt", message }); + promptAtMs = ms(); + })(); + + const timer = setInterval(() => { + if (modelError) return finish("MODEL_ERROR"); + if (firstTool) return finish("CHAINED"); + const sawTurn = lastEventAtMs > promptAtMs && (counts.agent_start || counts.turn_start || counts.message_start); + const quiet = ms() - lastEventAtMs > QUIESCENCE; + if (sawTurn && quiet) return finish("STALLED"); + if (ms() > DEADLINE) return finish(sawTurn ? "STALLED" : "NO_TURN"); + }, 1000); + + child.on("exit", () => { if (!firstTool) finish("CHILD_EXIT"); }); + }); +} + +const results = []; +for (const m of models) { + const [provider, modelId] = m.split("/"); + for (const scenario of scenarios) { + const tag = `${provider}_${modelId}_${scenario}`.replace(/[^a-z0-9_.-]/gi, "-"); + const r = await runCell(provider, modelId, scenario, path.join(outDir, `${tag}.jsonl`)); + results.push(r); + console.error(`done ${m} ${scenario}: ${r.verdict} (${r.firstToolName ?? "-"} @ ${r.firstToolAtMs ?? "-"}ms)`); + } +} + +// Markdown summary table + raw JSON for the record. +const rows = results.map((r) => + `| ${r.provider}/${r.modelId} | ${r.scenario} | ${r.verdict} | ${r.firstToolName ?? "-"} | ${r.firstToolAtMs ?? "-"} | ${r.error ?? ""} |`, +); +console.log("\n| model | scenario | verdict | first tool | t(ms) | error |"); +console.log("|---|---|---|---|---|---|"); +console.log(rows.join("\n")); +console.log("\nJSON " + JSON.stringify(results));