Bite-sized TDD tasks: schema-linking types + columns modal, gate widget with
staged per-table column selection, registry/summary wiring, and a replay
fixture generated from the real physical.yaml catalog. Offline-verifiable in
the replay server; harness wiring is Plan 2.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Tier 1 (clean-room first-turn harness, harness/scripts/model-matrix.mjs):
kickoff + resume chain in-turn on ALL available models — zai/glm-5.2,
deepseek/deepseek-v4-{pro,flash}, aritmolab/qwen3.6-35b-a3b, zai/glm-4.5-air.
The resume cold-start stall recurs on none (closes A's cross-model robustness).
aritmolab/gemma4-26b-a4b is a 404 at the endpoint (listed but not served) — an
availability gap classified as MODEL_ERROR, not a workflow issue.
Tier 2 (live, baseline zai/glm-5.2): F single-select auto-confirm verified
end-to-end — answering the first reviewer_select persisted a concept_clarified
decision (review_decisions.jsonl 0->1) with no follow-up confirmation gate.
Closes F's deferred live check.
No prompt hardening needed. Results in the G plan doc + memory. Throwaway psd
sessions used and deleted; real sessions untouched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Actionable scoping for the final workstream: a two-tier method (clean-room
first-turn harness generalised from the A2 repro-driver for cheap kickoff/resume
signals; sparing full Playwright F1 runs for the gate round-trip) over the
configured models (GLM 5.2 baseline, Deepseek V4, Qwen3.6, + breadth). Folds in
F's deferred single-select live check and A's resume-on-weaker-models check.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Lo spike Task 1 ha provato che ctx.sendRaw non esiste e pi.on(extension_ui_response)
non e' dispatchato. Aggiornati: Piano1 Task2 (mock ctx.ui), Task4 (rewrite gate a
ctx.ui.input con descriptor in title), Task10 (fake-pi-rpc shape nativa); Piano2
Task3/Task4 (SessionBridge decodifica title<->value). Contratto FE invariato.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Due piani separati (decomposizione concordata): il backend dipende da lavoro
harness non testato (gate RPC-ready, id injection, prereq CLI --json/--offset,
fake-Pi). Piano 1 rende l'harness pilotabile via RPC; Piano 2 costruisce il
backend Node/Fastify testato contro il fake-pi-rpc.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Task 3.2: nuovo Step 1b — drop funzioni registry da memory_cmd + tht/memory.py
(load/save/update/delete/promote/reusable_promotions); promotion (F5) riscritta
come upsert batch al vectordb; memory list/delete su vectordb (no registry).
- Skill F2/F5: drop riferimenti a 'memory promote + memory index' con registry.
- Onda 0b riscritta: repo workspace per-cliente separato (no copia in harness/),
LSH scarica-tutti-i-valori-distinti (no 'campiona'), indice nel repo workspace cliente.
Piano da spec 2026-06-27-cli-port-completo-skill-riscritta-design.md.
Struttura in onde: -1 (renaming tht isolato), 0 (backend), 0b (evidence+LSH setup),
1 (radici CLI + phase drift), 2 (vector), 3 (foglia + metadata memory), 4 (SQL/CTE),
S (skill riscritta), L2 (sessione manuale).
TDD per logica nuova (schema_linking_phase, require_phase_or_exit, arricchimento
metadata memory); port+smoke+commit per i port verbatim (logica gia' validata in
ChironeWp3). pytest verde (109 passed) a ogni task come gate di regressione.
Self-review: copertura spec completa (11 decisioni), nessun placeholder, type
consistency verificata. Gap residui onesti: circularita' session_cmd<->sql_cmd
(risolto con import lazy), RPC server mancanti (fuori piano codice), L2 manuale.