F3/F4 close with reviewer_confirm kind:'phase' (they do NOT auto-advance);
advance:true only auto-advances F2-empty/F6-skip; rewrite_question belongs
to F3 not F1; F4 uses the new write_schema_linking tool with the documented
SchemaLinking shape. Adds a per-phase artifact/close cheat-sheet.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
tht() gains an optional stdin arg; the tool pipes the schema-linking
object to 'tht session set-schema-linking --file -', which validates
against SchemaLinking and returns the exact error on failure — so the
model stops hand-writing the artifact and validating with ad-hoc python.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
reviewer_confirm kind:'cte_result' registered cte_approved --subject
phase:6, which decision_cmd rejects (exit 5) and next_cte never
recognizes — dead-ending F6. Derive the CTE name from 'tht cte next'
(plan-order single source of truth) and approve by name. sql path
(sql_approved:phase:N) unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Tier 1 (clean-room first-turn harness, harness/scripts/model-matrix.mjs):
kickoff + resume chain in-turn on ALL available models — zai/glm-5.2,
deepseek/deepseek-v4-{pro,flash}, aritmolab/qwen3.6-35b-a3b, zai/glm-4.5-air.
The resume cold-start stall recurs on none (closes A's cross-model robustness).
aritmolab/gemma4-26b-a4b is a 404 at the endpoint (listed but not served) — an
availability gap classified as MODEL_ERROR, not a workflow issue.
Tier 2 (live, baseline zai/glm-5.2): F single-select auto-confirm verified
end-to-end — answering the first reviewer_select persisted a concept_clarified
decision (review_decisions.jsonl 0->1) with no follow-up confirmation gate.
Closes F's deferred live check.
No prompt hardening needed. Results in the G plan doc + memory. Throwaway psd
sessions used and deleted; real sessions untouched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A1 — SessionMenu gains a Resume item, gated to status!=="finalized" && !archived
(matching the backend's 409 read-only guard), wired in AppShell to doResume ->
POST /sessions/:id/resume. SessionMenu.test.tsx (3 tests); frontend 96/96, tsc clean.
A2 — diagnosis-first clean-room repro driving `pi --mode rpc` with the backend's
exact resume handshake shows the cold-start stall NO LONGER reproduces on pi
0.79.4 (8/8 chained into `tht session show` + `read SKILL.md` in-turn, fresh and
partway sessions). The earlier narrate-and-stop predates the pi upgrade.
Defense-in-depth anyway: RIPRENDI_KICKOFF hardened to force the in-turn tool call
(gate_resume_kickoff.test.js + live regression 2/2). Gate JS 34/34.
PROJECT_STATE open-item #1 (resume stall) flipped to RESOLVED; cross-model resume
robustness folded into workstream G.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
F — reviewer_select options may now carry a `decision` payload {type, subject,
detail?, rationale?} plus an optional `advance`. Picking such an option IS the
confirmation: the gate persists it directly (tht decision add) and optionally
advances, with no redundant reviewer_decide/reviewer_confirm follow-up gate.
Options without a payload stay ask-only; back/exit/Other never persist.
Pure logic extracted + exported for unit tests: resolveSelectOutcome (classifies
the response) and decisionAddArgs (shared with reviewer_decide, DRY). Gate JS
suite 33/33 (gate_select_decision.test.js, +5); harness pytest 269 unchanged.
Contract docs updated together: reviewer_select tool description, SKILL.md
(widget summary, disciplines 2-3, Phase-1 single-pick), and the CLAUDE.md gate
note. Live verification (model truly emits reviewer_select+decision, decision in
review_decisions.jsonl, no follow-up gate) deferred to workstream G — it is
model-behavior-dependent.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
New sessions get a concise Italian-keyword `name` instead of the truncated
question. `tht session new` (when no --name is given) derives it via a new
`_extract_name` helper using YAKE (pure-Python, unsupervised, Italian, no LLM),
dropping generic query verbs and keeping the top keywords in reading order;
falls back to `_summarize` if YAKE is unavailable. `create_session` core keeps
its `name=None` default — the policy lives at the CLI layer.
TDD: tests/test_session_name.py (unit + CliRunner integration). Full harness
suite 269 passed; ruff clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The frontend's reviewer widgets uniformly send the picked option in a `choices`
array (SelectWidget/ArtifactGateWidget: `choices: [optionId]`), but the gate's
reviewer_select and reviewer_confirm(reject) handlers read `resp.choice`
(singular). Result: every single-select gate saw an undefined choice, answered
"Nessuna scelta ricevuta", and re-presented forever — the workflow could never
pass F1. (reviewer_decide/multiselect already read `resp.choices`, so it worked.)
Add a shared selectedChoice(resp) helper reading choices[0] (falling back to the
legacy singular choice); both handlers use it.
TDD: gate/__tests__/gate_choice.test.js RED->GREEN; full gate suite 28/28.
Verified LIVE (Playwright -> real Pi -> GLM 5.2): a single-select F1 answer is
now accepted and the workflow advances (clarification 2/4 -> 3/4). The same run
also live-verified the F1 hang fix (418187a).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
ctx.ui.input in `pi --mode rpc` correlates extension_ui_response on its own
top-level RPC id (crypto.randomUUID), not the descriptor id the gate carries
in `title`. SessionBridge replied with the descriptor id, so Pi silently
dropped the response and the model never resumed — every reviewer widget hung
after the human answered.
SessionBridge now stores Pi's top-level m.id (pendingPiId) and replies
extension_ui_response{ id: pendingPiId, value: <uiResponse> }; value still
carries the descriptor id so the gate's internal resp.id === descriptor.id
check still holds.
The fake-pi double had masked the bug by forcing m.id == descriptor.id; it now
mirrors real Pi (distinct randomUUID, correlate on it, drop unknown ids), with
a negative regression test. SKILL.md Phase 1 also now steers multi-answer
disambiguation to reviewer_decide (multiselect).
Tests: backend 67/67, tsc clean, fake-pi contract 2/2.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- loadEnvFromDotenv: read ctx.cwd/.env into the process environment
- prepareReviewerArguments: normalize/parse reviewer tool inputs
- tidy reserved-option filtering and tht command argument assembly
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
workspace/provider/model/thinking now come from getSettings() injected into
sessionRoutes; the request body supplies only question+name. Also teaches
fake_pi_rpc to respond to set_model and set_thinking_level RPC commands so
tests that pass real model settings don't hang.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Harness:
- preview_cmd FILE positional arg made optional; when omitted with --session,
path is derived via _session_sql_file (mirrors export_cmd) — fixes the
deferred Task-5 bug where the backend passed sessions/<id>/sql_final.sql
relative to harnessDir, which broke for workspace-dependent paths.
- New pytest: test_preview_session_no_file_resolves_sql_final
Backend:
- ThtRunner.sqlPreview: drop positional file arg; use --session only
- New routes/sql.ts: POST /sessions/:id/sql/preview + /export
- New routes/meta.ts: GET /workspaces (yaml scan) + GET /models (injectable
seam + graceful fallback to {models:[]})
- app.ts: register sqlRoutes + metaRoutes; add listModels to BuildAppDeps
- tht-runner.test.ts: add sqlPreview argv assertion (no file path)
- test/routes-sql-meta.test.ts: 9 tests (sql preview/export + meta routes)
Tests: harness 233 passed; backend 29 passed; build clean.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Set quietStartup:true in .pi/settings.json to suppress Pi banner on RPC stdout.
Add test pinning both quietStartup and theme values. Document project-local trust
requirement (--approve) so no trust prompt blocks RPC loop iteration.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- test_list_sessions_required_keys: assert full spec key set
(adds summary/updated_at/author) so dropping any goes caught.
- test_cli_show_json_valid: assert "schema" in / "db_schema" not in
data to pin by_alias=True on the alias-sensitive field.
Production code unchanged. 8 passed; full suite 230 passed.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When offset>0 the wrapper added an outer LIMIT N, so run_controlled's
_inject_limit bailed (a LIMIT IS present) and truncated was always False —
AGGrid could never detect more rows. Fix: for offset>0 probe with LIMIT (N+1)
OFFSET M, then compute truncated = len(rows) > N in do_run and slice back to N.
offset==0 path unchanged (delegates to extracted _run_transport helper). JSON
still reports the user's requested limit N and correct truncated. Adds 3 tests
exercising the real do_run offset>0 path (N+1 -> truncated True, N -> False,
offset==0 verbatim). 222/222 passing.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add inject_limit_offset (tht/execute/limit.py) — pure subquery wrapper that
applies LIMIT/OFFSET non-destructively without clobbering user-supplied LIMITs.
Wire offset param into do_run (pre-processing when offset>0) and add --offset /
--json flags to preview_cmd; JSON mode emits pristine stdout with columns, rows,
execution_ms, truncated, limit, offset. 5 new tests (4 unit + 1 JSON-purity),
219/219 total passing (no regressions).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When THT_SESSION env var is set, /nuova-domanda injects a kickoff that
tells the model to use the pre-created session id instead of running
`tht session new`. /riprendi-sessione and no-env-var paths unchanged.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Removed event.source === "interactive" guard from entry detection so
/nuova-domanda and /riprendi-sessione activate lockActive+pendingKickoff
regardless of source (TUI or RPC prompt).
- Removed event.source !== "interactive" from free-input filter; lock now
blocks/steers all user input when active, not only interactive keystrokes.
- Added typebox@1.1.38 devDep + fake_pi_runtime.registerCommand (gap from Task 2).
- New test: gate_entry.test.js (2 tests: lock activates on RPC; !-steer passes).
- Full suite: 19/19 pass, zero regressions.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Implementazione del piano di remediation progressiva sui difetti emersi
dall'analisi dell'harness. Tutto verificato: 214 test Python (incl. L0 su
Postgres reale), 14 test JS del gate, ruff pulito.
Blocco 1 (CRITICA, integrazione gate↔CLI):
- phase advance: gate usa --auto + exit 6; reviewer_confirm kind:phase fa
advance esplicito che applica i prerequisiti (prima non avanzava per le
fasi a conferma umana).
- cte plan riceve i --name dal gate (param names); set-question con id
posizionale; skill `tht search find`; nuovo comando `tht memory save-one`
con dedup hash client-side in save_one_memory.
Blocco 2 (D15, stato post-rollback):
- campo `phase` su DecisionRecord + effective_decisions phase-aware per i
subject "a nome" (cte_approved ecc.); _compute_promotions e finalize sulla
vista effective; finalize confronta col piano CTE effettivo, non glob;
`decision add --retracts` + comando `decision retract`.
Blocco 3 (D7 read-only + D6 manifest):
- assert_read_only su tutti e quattro i codepath (direct + REST);
- manifest author/summary/updated_at/updated_by/schema_version popolati +
helper touch_manifest sulle mutazioni.
Blocco 4-5 (D14a/D14b):
- decision_min_phase data-driven via `emits:` in workflow.yaml;
- formula evidence: status auto, search_formulas, gruppo CLI `tht formula`,
`search find --kind formula`, load_evidence_dir salta i .sql.md.
Blocco 6 (robustezza):
- taskdoc slice promoted_tables + bound enforced; report escaping/bound +
rsplit note; filtro kind reader REST/direct; conteggio upserted robusto;
guard REST run_query non-list; LSH disallineato -> LshIndexError.
Blocco 7 (pulizia):
- dead code gate e KIND_TO_TABLE morto rimossi; doc Postgres-only
(README + connection.py).
Blocco 0 (parziale): test di compatibilità firma gate↔CLI
(tests/integration). Rinviati: fake-Pi runtime completo, artifact-gate da
disco (#23), parità eligibility REST/direct (#28), unificazione
reserved-labels (#30), memory_rejected da deselezione (#33).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Prima sessione L2 end-to-end dopo il porting. Il loop skill->LLM->gate funziona nel
dominio (ricerche, evidence, quadro corretto su fact_cardioversione/fact_see_ablazione)
ma si blocca a F1 sul bug fatale #4 (ctx.sendRaw non esiste nel runtime Pi).
Bug emersi (4):
#1 tht non nel PATH di Pi (basso, workaround wrapper)
#2 phase show non passava config (medio, FIXATO: _cfg() risolve env+default)
#3 session check signature inconsistente (basso, da verificare)
#4 ctx.sendRaw is not a function (FATALE, mismatch architetturale: il porting ha
sostituito i dialog nativi ctx.ui.* con ctx.sendRaw, API non esposta in questo Pi)
Analisi #4 (verificata sul runtime installato):
- ctx.sendRaw non esiste; extension_ui_request e' emesso solo dal runtime
(modes/rpc) come traduzione di ctx.ui.select/confirm/input, non come API extension
- canali RPC per decisioni: enum chiuso select/confirm/input/editor (no custom)
- ctx.ui.custom (multiselect TUI di ChironeWp3) e' no-op in RPC mode
- conseguenza: multiselect F4 non ha canale in RPC -> decisione di design del gate
aperta (opzioni A/B/C nel report, C=scartato), merita brainstorming dedicato
Fix#2 (dal modello in sessione, validato e ripulito): phase_cmd._cfg() ora risolve
THT_WORKSPACE/THT_CONFIG env poi fallback config/tht.yaml (stessa convenzione CONFIG_OPT).
Testato phase show OK, suite 165 passed.
Report: docs/l2-run-report-2026-06-27.md.
Bug di porting emerso in L2 prep: tht session new falliva con
'AttributeError: SessionManifest has no to_yaml'. Lo stub locale _YamlModel in
session/models.py (placeholder pre-porting mschema) definiva solo
populate_by_name, senza i metodi to_yaml/from_yaml che store.py e session_cmd.py
usano. Aggiunti gli stessi metodi della controparte mschema (mantenendo
populate_by_name, necessario per db_schema alias='schema').
config/tht.yaml: symlink locale al workspace cliente attivo (psd.yaml). Il gate
chiama tht senza -c (default config/tht.yaml), quindi serve questo ponte per il
deployment per-cliente. Gitignored (per-cliente). .gitignore: + config/tht.yaml,
+ artifacts/.
Verifica: pytest 165 passed; tht session new crea sessione nel repo cliente;
pi vede i 3 tool reviewer_* (gate caricato); GLM 5.2 e' il model default di pi.
Due cambiamenti interconnessi da user review:
1. language come parametro workspace (spec decisione 9):
- Config.language (default 'en') + workspaces PSD con 'language: it'
- Generalizza Thoth oltre l'italiano: descrizioni tabelle/colonne ed evidence
sono nel workspace language; le istruzioni della skill restano in inglese
(piu' affidabili per modelli piccoli, meno ambigue)
2. Skill riscritta in INGLESE preservando la semantica COMPLETA dell'originale
(autocritica: la mia riscrittura precedente aveva perso ~10 vincoli precisi):
- 'promuovere' ambiguo (3 accezioni: phase advance / recommend / memory promote)
-> 'never advance a phase or record a decision without confirmation'
- recuperati vincoli persi: choice-is-confirmation (no reviewer_confirm dopo
reviewer_decide), reviewer_select SOLO per iterazione no-decision, messaggi
auto-contenuti obbligatori, artefatto = superficie di decisione (gate rilegge
da disco per CTE/SQL), candidati con provenienza+score non verita', opzione
'leave ambiguity open', F1 passa lista completa non solo ultima
- language contract esplicito (istruzioni EN, output nel workspace language)
Sottomoduli cte/memoria/rewriting/sql-generation in inglese, semantica tecnica
intatta (regole AV-SQL, dim_time trick, max 5 memorie solo 3 tipi riusabili).
Verifica: 0 residui nsp/chirone, tutti i tht <cmd> citati registrati, 165 passed.
Skill ex-novo che riflette Thoth (non copia di ChironeWp3):
- vocabolario widget-descriptor (reviewer_select/decide/confirm) invece di 'dialog native'
- D11 save-one in F2 (upsert mirato vs full resync)
- D14a value_grounded (LSH multi-colonna non collassa) + D14b concept_formula in F4
- D13 free-text e D15 rollback nelle discipline trasversali
- memory vive SOLO nel vectordb (drop registry, spec 5): save-one/promote senza registry
- F8 datamart onesto (stub NotImplementedError)
Sottomoduli cte/memoria/rewriting/sql-generation portati adattando nsp->tht, con
i vincoli precisi trasferiti fedelmente (max 5 memorie, solo 3 tipi riusabili,
CTE solo WITH senza SELECT, dim_time join non aritmetica, sql_final pulito).
Verifica: zero residui nsp/chirone/psd nella skill; ogni 'tht <cmd>' citato e'
registrato (correzione: 'tht formula retrieve' era inesistente -> riformulato in
'ricerca nelle evidence'). Suite: 165 passed.
Ultima onda CLI. 4 cmd portati con rename + grep-per-file (3 residui nsp nei messaggi
fixati). Nessun drift costanti phase in questi cmd.
La CLI tht e' ora COMPLETA: 14 gruppi di comandi (phase config schema session vector
memory search evidence db decision sql cte datamart lsh). tht --help li list tutti.
Suite: 165 passed.
Il loop skill->LLM->gate ora ha tutti i comandi che la skill chiamera'. Resta:
skill riscritta (S), setup pre-sessione (0b), sessione L2 manuale.
Correzione del gap ereditato (resosi NECESSARIO dal drop del registry, spec 5): il
metadata del VectorRecord memory ora porta subject/detail/rationale oltre a
type/session_id/tables/concepts. pack_metadata li serializza nel jsonb via
**record.metadata. search_similar proietta metadata completo -> la F2 ricostruisce
la decisione direttamente dall'hit, senza lookup registro.
L1: 4 test (subject/detail/rationale presenti, campi esistenti preservati,
no cross-contamination multi-record, save_one_memory propaga il metadata alla riga).
Suite: 165 passed.
Stesso bug del precedente: VENDORED.md e' arrivato con Onda 0 (dopo l'Onda -1 che
aveva pulito i riferimenti cliente). Neutralizzato a 'il datawarehouse' (coerente
con le altre neutralizzazioni). Sweep completo tht/ ora vuoto per chirone/psdwp3/
policlinico/sandonato.