2 Commits
Author SHA1 Message Date
marcopanandClaude Fable 5 4479683cf4 chore(pi): scope AritmoLab provider to project, drop Gemma model
Move the AritmoLab provider (Qwen) into harness/.pi/extensions/ so it is
auto-discovered only when Pi runs with cwd=harness, keeping it project-local.
GLM (~/.pi/agent/models.json) and DeepSeek (built-in) stay user-wide. File is
.js (not .mjs) because Pi extension auto-discovery matches only /\.(ts|js)$/.

Removes the gemma4-26b-a4b model (404 at the endpoint) from both the provider
and the model-matrix default list.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-02 11:49:11 +02:00
marcopanandClaude Opus 4.8 cbb8e184f1 feat(harness): cross-model behavior matrix (G) — harness + results
Tier 1 (clean-room first-turn harness, harness/scripts/model-matrix.mjs):
kickoff + resume chain in-turn on ALL available models — zai/glm-5.2,
deepseek/deepseek-v4-{pro,flash}, aritmolab/qwen3.6-35b-a3b, zai/glm-4.5-air.
The resume cold-start stall recurs on none (closes A's cross-model robustness).
aritmolab/gemma4-26b-a4b is a 404 at the endpoint (listed but not served) — an
availability gap classified as MODEL_ERROR, not a workflow issue.

Tier 2 (live, baseline zai/glm-5.2): F single-select auto-confirm verified
end-to-end — answering the first reviewer_select persisted a concept_clarified
decision (review_decisions.jsonl 0->1) with no follow-up confirmation gate.
Closes F's deferred live check.

No prompt hardening needed. Results in the G plan doc + memory. Throwaway psd
sessions used and deleted; real sessions untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 18:46:55 +02:00