The model sometimes stops after receiving 'Fase approvata' without calling
`tht session finalize`, leaving the session open. Now the gate itself calls
finalize after advancing the max phase (F8), making session closure
deterministic regardless of model behavior.
- reviewer_confirm kind:phase: after phase advance at max_phase, gate calls
`tht session finalize <session>` (best-effort with recovery message)
- SKILL.md updated: model no longer needs to call finalize itself
- Tests: 2 new JS tests (auto-finalize at max phase; no-finalize at non-max)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Lever 1: Join-graph via FK logics in annotations + suggest-fks command
- TableAnnotation.foreign_keys field stores curated logical FKs (DWH has no FK constraints)
- tht schema suggest-fks: mine from approved SQL, heuristics (time_key → dim_time),
same-name discovery + explicit --assume flag for multi-owner PKs
- mschema renders 【Foreign keys】 section populated; validation in merge.py
- SKILL.md F4 now reads FKs from mschema-text, no custom data_time_key logic
Lever 2: Context-pack consolidation at kickoff (tht search pack)
- Single embedding of question, reused for schema + evidence + solved searches
- One command: tht search pack <question> --session <id> → retrieval_pack.md
- Graceful degradation when Ollama/vector store unreachable (exit 0, empty sections)
- SKILL.md F1 prescribes as first call; reduces model thinking turns via pre-retrieval
Lever 3: Phase-summary recap v2 auto-construction from session ledger
- tht session show --json includes full decisions ledger
- tht phase meta --json exports 'emits' (substantive decision types per phase)
- Gate appends deterministic 【Decisioni registrate in questa fase】 section (appendLedgerSection)
- Model authors only summary + checks; recap table comes from persisted state (exact by construction)
- SKILL.md Disciplina 6: brief model output, gate fills the rest
Tests: 358 Python (including 10 FK + 3 pack + 1 session-ledger tests) + 111 JS gate tests, all pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Transcript analysis (session 2026-07-06-175012, GLM 5.2) showed F1 at 567s:
182s wasted on a useless `tht schema introspect` (re-introspecting the remote
DWH although physical.yaml was already materialized) plus ~220s of model
thinking inflated by ~7 exploratory turns (--help/find/cat). The actual
searches cost ~15s; reviewer gates (~145s, untouched) are the quality contract.
- schema_cmd.py: introspect now exits 0 with "OK (cache)" in ~1s when
physical.yaml exists; --refresh forces the real re-introspection.
Deterministic cross-model guarantee, verified live on psd (163 tables, 1.2s).
- SKILL.md: F1 toolbox (only `tht search find` + `tht schema render`; no
introspect/--help/filesystem browsing; batch all searches in one turn);
F4 step 1 is render-only with a one-shot introspect fallback.
- tht-gate.js: `tht schema introspect ... --refresh` added to FORBIDDEN
(maintenance stays shell-only, never in-session).
- tests: 4 new pytest cases (cache hit placement proven with fake credentials,
refresh bypass, corrupt-catalog fall-through, render fallback message) and
2 gate anti-bypass JS cases.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The tht-gate.js reviewer_confirm now builds structured v2 artifacts before the
widget so the reviewer approves gate-derived data, not raw model text:
- new pure modules gate/artifact-contracts.js (soft validators, {ok,errors},
legacy-passthrough) and gate/enrich.js (index/description enrichment,
buildCteResultV2 fusing thin model data with `tht cte info`, phase enrichment)
- cte_plan v2: validate + enrich + persist via `tht cte plan --name … --doc -`
(names derived from data.ctes[]); legacy `names` param kept as fallback
- cte_result v2: rebuild from `tht cte next`/`tht cte info` (sql + preview from
the persisted test record); null/error last_test -> actionable textResult
- phase v2: soft-validate + fill phase from meta + catalog descriptions
- prepareReviewerArguments coerces artifact.data too (GLM double-stringify);
legacy markdown strings pass through unchanged
- SKILL.md: Phase 6 cte_plan payload A + thin cte_result guidance; Discipline 6
payload C example; Discipline 7 reworded for gate-rebuilt cte_result
Legacy (non-v2) paths unchanged. TypeBox stays Type.Any() for artifact.data;
validation is soft (textResult) so models self-correct instead of looping.
All 102 gate JS tests green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The cheat-sheet, Discipline 2, and Phase 7 said F7 closes with
reviewer_confirm kind:"sql" — but the gate's kind:"sql" only records
sql_approved:phase:7; F7 is not in _AUTO_ADVANCE_PHASES, so advancing to
F8 still needs reviewer_confirm kind:"phase" (mirrors F6). Same
record-vs-advance trap this branch removes elsewhere.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
F3/F4 close with reviewer_confirm kind:'phase' (they do NOT auto-advance);
advance:true only auto-advances F2-empty/F6-skip; rewrite_question belongs
to F3 not F1; F4 uses the new write_schema_linking tool with the documented
SchemaLinking shape. Adds a per-phase artifact/close cheat-sheet.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
F — reviewer_select options may now carry a `decision` payload {type, subject,
detail?, rationale?} plus an optional `advance`. Picking such an option IS the
confirmation: the gate persists it directly (tht decision add) and optionally
advances, with no redundant reviewer_decide/reviewer_confirm follow-up gate.
Options without a payload stay ask-only; back/exit/Other never persist.
Pure logic extracted + exported for unit tests: resolveSelectOutcome (classifies
the response) and decisionAddArgs (shared with reviewer_decide, DRY). Gate JS
suite 33/33 (gate_select_decision.test.js, +5); harness pytest 269 unchanged.
Contract docs updated together: reviewer_select tool description, SKILL.md
(widget summary, disciplines 2-3, Phase-1 single-pick), and the CLAUDE.md gate
note. Live verification (model truly emits reviewer_select+decision, decision in
review_decisions.jsonl, no follow-up gate) deferred to workstream G — it is
model-behavior-dependent.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
ctx.ui.input in `pi --mode rpc` correlates extension_ui_response on its own
top-level RPC id (crypto.randomUUID), not the descriptor id the gate carries
in `title`. SessionBridge replied with the descriptor id, so Pi silently
dropped the response and the model never resumed — every reviewer widget hung
after the human answered.
SessionBridge now stores Pi's top-level m.id (pendingPiId) and replies
extension_ui_response{ id: pendingPiId, value: <uiResponse> }; value still
carries the descriptor id so the gate's internal resp.id === descriptor.id
check still holds.
The fake-pi double had masked the bug by forcing m.id == descriptor.id; it now
mirrors real Pi (distinct randomUUID, correlate on it, drop unknown ids), with
a negative regression test. SKILL.md Phase 1 also now steers multi-answer
disambiguation to reviewer_decide (multiselect).
Tests: backend 67/67, tsc clean, fake-pi contract 2/2.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Implementazione del piano di remediation progressiva sui difetti emersi
dall'analisi dell'harness. Tutto verificato: 214 test Python (incl. L0 su
Postgres reale), 14 test JS del gate, ruff pulito.
Blocco 1 (CRITICA, integrazione gate↔CLI):
- phase advance: gate usa --auto + exit 6; reviewer_confirm kind:phase fa
advance esplicito che applica i prerequisiti (prima non avanzava per le
fasi a conferma umana).
- cte plan riceve i --name dal gate (param names); set-question con id
posizionale; skill `tht search find`; nuovo comando `tht memory save-one`
con dedup hash client-side in save_one_memory.
Blocco 2 (D15, stato post-rollback):
- campo `phase` su DecisionRecord + effective_decisions phase-aware per i
subject "a nome" (cte_approved ecc.); _compute_promotions e finalize sulla
vista effective; finalize confronta col piano CTE effettivo, non glob;
`decision add --retracts` + comando `decision retract`.
Blocco 3 (D7 read-only + D6 manifest):
- assert_read_only su tutti e quattro i codepath (direct + REST);
- manifest author/summary/updated_at/updated_by/schema_version popolati +
helper touch_manifest sulle mutazioni.
Blocco 4-5 (D14a/D14b):
- decision_min_phase data-driven via `emits:` in workflow.yaml;
- formula evidence: status auto, search_formulas, gruppo CLI `tht formula`,
`search find --kind formula`, load_evidence_dir salta i .sql.md.
Blocco 6 (robustezza):
- taskdoc slice promoted_tables + bound enforced; report escaping/bound +
rsplit note; filtro kind reader REST/direct; conteggio upserted robusto;
guard REST run_query non-list; LSH disallineato -> LshIndexError.
Blocco 7 (pulizia):
- dead code gate e KIND_TO_TABLE morto rimossi; doc Postgres-only
(README + connection.py).
Blocco 0 (parziale): test di compatibilità firma gate↔CLI
(tests/integration). Rinviati: fake-Pi runtime completo, artifact-gate da
disco (#23), parità eligibility REST/direct (#28), unificazione
reserved-labels (#30), memory_rejected da deselezione (#33).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Due cambiamenti interconnessi da user review:
1. language come parametro workspace (spec decisione 9):
- Config.language (default 'en') + workspaces PSD con 'language: it'
- Generalizza Thoth oltre l'italiano: descrizioni tabelle/colonne ed evidence
sono nel workspace language; le istruzioni della skill restano in inglese
(piu' affidabili per modelli piccoli, meno ambigue)
2. Skill riscritta in INGLESE preservando la semantica COMPLETA dell'originale
(autocritica: la mia riscrittura precedente aveva perso ~10 vincoli precisi):
- 'promuovere' ambiguo (3 accezioni: phase advance / recommend / memory promote)
-> 'never advance a phase or record a decision without confirmation'
- recuperati vincoli persi: choice-is-confirmation (no reviewer_confirm dopo
reviewer_decide), reviewer_select SOLO per iterazione no-decision, messaggi
auto-contenuti obbligatori, artefatto = superficie di decisione (gate rilegge
da disco per CTE/SQL), candidati con provenienza+score non verita', opzione
'leave ambiguity open', F1 passa lista completa non solo ultima
- language contract esplicito (istruzioni EN, output nel workspace language)
Sottomoduli cte/memoria/rewriting/sql-generation in inglese, semantica tecnica
intatta (regole AV-SQL, dim_time trick, max 5 memorie solo 3 tipi riusabili).
Verifica: 0 residui nsp/chirone, tutti i tht <cmd> citati registrati, 165 passed.
Skill ex-novo che riflette Thoth (non copia di ChironeWp3):
- vocabolario widget-descriptor (reviewer_select/decide/confirm) invece di 'dialog native'
- D11 save-one in F2 (upsert mirato vs full resync)
- D14a value_grounded (LSH multi-colonna non collassa) + D14b concept_formula in F4
- D13 free-text e D15 rollback nelle discipline trasversali
- memory vive SOLO nel vectordb (drop registry, spec 5): save-one/promote senza registry
- F8 datamart onesto (stub NotImplementedError)
Sottomoduli cte/memoria/rewriting/sql-generation portati adattando nsp->tht, con
i vincoli precisi trasferiti fedelmente (max 5 memorie, solo 3 tipi riusabili,
CTE solo WITH senza SELECT, dim_time join non aritmetica, sql_final pulito).
Verifica: zero residui nsp/chirone/psd nella skill; ogni 'tht <cmd>' citato e'
registrato (correzione: 'tht formula retrieve' era inesistente -> riformulato in
'ricerca nelle evidence'). Suite: 165 passed.