Embeddings timeout was 120s, causing multi-minute hangs when Ollama was
down during F4/F6/F7 solved-search. Now: connect_timeout=5s across all
HTTP clients (REST + Ollama), read_timeout reduced to 30s for embeddings,
and OllamaEmbeddings auto-restarts the server on ConnectionError before
degrading gracefully.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Lever 1: Join-graph via FK logics in annotations + suggest-fks command
- TableAnnotation.foreign_keys field stores curated logical FKs (DWH has no FK constraints)
- tht schema suggest-fks: mine from approved SQL, heuristics (time_key → dim_time),
same-name discovery + explicit --assume flag for multi-owner PKs
- mschema renders 【Foreign keys】 section populated; validation in merge.py
- SKILL.md F4 now reads FKs from mschema-text, no custom data_time_key logic
Lever 2: Context-pack consolidation at kickoff (tht search pack)
- Single embedding of question, reused for schema + evidence + solved searches
- One command: tht search pack <question> --session <id> → retrieval_pack.md
- Graceful degradation when Ollama/vector store unreachable (exit 0, empty sections)
- SKILL.md F1 prescribes as first call; reduces model thinking turns via pre-retrieval
Lever 3: Phase-summary recap v2 auto-construction from session ledger
- tht session show --json includes full decisions ledger
- tht phase meta --json exports 'emits' (substantive decision types per phase)
- Gate appends deterministic 【Decisioni registrate in questa fase】 section (appendLedgerSection)
- Model authors only summary + checks; recap table comes from persisted state (exact by construction)
- SKILL.md Disciplina 6: brief model output, gate fills the rest
Tests: 358 Python (including 10 FK + 3 pack + 1 session-ledger tests) + 111 JS gate tests, all pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Transcript analysis (session 2026-07-06-175012, GLM 5.2) showed F1 at 567s:
182s wasted on a useless `tht schema introspect` (re-introspecting the remote
DWH although physical.yaml was already materialized) plus ~220s of model
thinking inflated by ~7 exploratory turns (--help/find/cat). The actual
searches cost ~15s; reviewer gates (~145s, untouched) are the quality contract.
- schema_cmd.py: introspect now exits 0 with "OK (cache)" in ~1s when
physical.yaml exists; --refresh forces the real re-introspection.
Deterministic cross-model guarantee, verified live on psd (163 tables, 1.2s).
- SKILL.md: F1 toolbox (only `tht search find` + `tht schema render`; no
introspect/--help/filesystem browsing; batch all searches in one turn);
F4 step 1 is render-only with a one-shot introspect fallback.
- tht-gate.js: `tht schema introspect ... --refresh` added to FORBIDDEN
(maintenance stays shell-only, never in-session).
- tests: 4 new pytest cases (cache hit placement proven with fake credentials,
refresh bypass, corrupt-catalog fall-through, render fallback message) and
2 gate anti-bypass JS cases.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
WS1 of review-gates-v2: gives the JS gate (WS2) deterministic data to build the
cte_result v2 payload.
- CteTestRecord gains optional preview_rows (JSON-coerced, truncated cells);
test_cmd populates it from the bounded result rows.
- New read-only `tht cte info <name> --session <id> [--json]`: plan
index/total, persisted .sql, cte_plan_doc.json entry (if any), last
CteTestRecord. Exits 1 with a clean stderr message on missing
session/plan/name/sql.
- `tht cte plan --doc -` validates a chain-doc JSON (ctes[].name must match
--name, same order) and writes it to cte_plan_doc.json; cte_plan.json stays
a plain list[str] (load-bearing for tht.phase.next_cte). --doc is optional.
New sessions get a concise Italian-keyword `name` instead of the truncated
question. `tht session new` (when no --name is given) derives it via a new
`_extract_name` helper using YAKE (pure-Python, unsupervised, Italian, no LLM),
dropping generic query verbs and keeping the top keywords in reading order;
falls back to `_summarize` if YAKE is unavailable. `create_session` core keeps
its `name=None` default — the policy lives at the CLI layer.
TDD: tests/test_session_name.py (unit + CliRunner integration). Full harness
suite 269 passed; ruff clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Harness:
- preview_cmd FILE positional arg made optional; when omitted with --session,
path is derived via _session_sql_file (mirrors export_cmd) — fixes the
deferred Task-5 bug where the backend passed sessions/<id>/sql_final.sql
relative to harnessDir, which broke for workspace-dependent paths.
- New pytest: test_preview_session_no_file_resolves_sql_final
Backend:
- ThtRunner.sqlPreview: drop positional file arg; use --session only
- New routes/sql.ts: POST /sessions/:id/sql/preview + /export
- New routes/meta.ts: GET /workspaces (yaml scan) + GET /models (injectable
seam + graceful fallback to {models:[]})
- app.ts: register sqlRoutes + metaRoutes; add listModels to BuildAppDeps
- tht-runner.test.ts: add sqlPreview argv assertion (no file path)
- test/routes-sql-meta.test.ts: 9 tests (sql preview/export + meta routes)
Tests: harness 233 passed; backend 29 passed; build clean.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When offset>0 the wrapper added an outer LIMIT N, so run_controlled's
_inject_limit bailed (a LIMIT IS present) and truncated was always False —
AGGrid could never detect more rows. Fix: for offset>0 probe with LIMIT (N+1)
OFFSET M, then compute truncated = len(rows) > N in do_run and slice back to N.
offset==0 path unchanged (delegates to extracted _run_transport helper). JSON
still reports the user's requested limit N and correct truncated. Adds 3 tests
exercising the real do_run offset>0 path (N+1 -> truncated True, N -> False,
offset==0 verbatim). 222/222 passing.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add inject_limit_offset (tht/execute/limit.py) — pure subquery wrapper that
applies LIMIT/OFFSET non-destructively without clobbering user-supplied LIMITs.
Wire offset param into do_run (pre-processing when offset>0) and add --offset /
--json flags to preview_cmd; JSON mode emits pristine stdout with columns, rows,
execution_ms, truncated, limit, offset. 5 new tests (4 unit + 1 JSON-purity),
219/219 total passing (no regressions).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Implementazione del piano di remediation progressiva sui difetti emersi
dall'analisi dell'harness. Tutto verificato: 214 test Python (incl. L0 su
Postgres reale), 14 test JS del gate, ruff pulito.
Blocco 1 (CRITICA, integrazione gate↔CLI):
- phase advance: gate usa --auto + exit 6; reviewer_confirm kind:phase fa
advance esplicito che applica i prerequisiti (prima non avanzava per le
fasi a conferma umana).
- cte plan riceve i --name dal gate (param names); set-question con id
posizionale; skill `tht search find`; nuovo comando `tht memory save-one`
con dedup hash client-side in save_one_memory.
Blocco 2 (D15, stato post-rollback):
- campo `phase` su DecisionRecord + effective_decisions phase-aware per i
subject "a nome" (cte_approved ecc.); _compute_promotions e finalize sulla
vista effective; finalize confronta col piano CTE effettivo, non glob;
`decision add --retracts` + comando `decision retract`.
Blocco 3 (D7 read-only + D6 manifest):
- assert_read_only su tutti e quattro i codepath (direct + REST);
- manifest author/summary/updated_at/updated_by/schema_version popolati +
helper touch_manifest sulle mutazioni.
Blocco 4-5 (D14a/D14b):
- decision_min_phase data-driven via `emits:` in workflow.yaml;
- formula evidence: status auto, search_formulas, gruppo CLI `tht formula`,
`search find --kind formula`, load_evidence_dir salta i .sql.md.
Blocco 6 (robustezza):
- taskdoc slice promoted_tables + bound enforced; report escaping/bound +
rsplit note; filtro kind reader REST/direct; conteggio upserted robusto;
guard REST run_query non-list; LSH disallineato -> LshIndexError.
Blocco 7 (pulizia):
- dead code gate e KIND_TO_TABLE morto rimossi; doc Postgres-only
(README + connection.py).
Blocco 0 (parziale): test di compatibilità firma gate↔CLI
(tests/integration). Rinviati: fake-Pi runtime completo, artifact-gate da
disco (#23), parità eligibility REST/direct (#28), unificazione
reserved-labels (#30), memory_rejected da deselezione (#33).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Prima sessione L2 end-to-end dopo il porting. Il loop skill->LLM->gate funziona nel
dominio (ricerche, evidence, quadro corretto su fact_cardioversione/fact_see_ablazione)
ma si blocca a F1 sul bug fatale #4 (ctx.sendRaw non esiste nel runtime Pi).
Bug emersi (4):
#1 tht non nel PATH di Pi (basso, workaround wrapper)
#2 phase show non passava config (medio, FIXATO: _cfg() risolve env+default)
#3 session check signature inconsistente (basso, da verificare)
#4 ctx.sendRaw is not a function (FATALE, mismatch architetturale: il porting ha
sostituito i dialog nativi ctx.ui.* con ctx.sendRaw, API non esposta in questo Pi)
Analisi #4 (verificata sul runtime installato):
- ctx.sendRaw non esiste; extension_ui_request e' emesso solo dal runtime
(modes/rpc) come traduzione di ctx.ui.select/confirm/input, non come API extension
- canali RPC per decisioni: enum chiuso select/confirm/input/editor (no custom)
- ctx.ui.custom (multiselect TUI di ChironeWp3) e' no-op in RPC mode
- conseguenza: multiselect F4 non ha canale in RPC -> decisione di design del gate
aperta (opzioni A/B/C nel report, C=scartato), merita brainstorming dedicato
Fix#2 (dal modello in sessione, validato e ripulito): phase_cmd._cfg() ora risolve
THT_WORKSPACE/THT_CONFIG env poi fallback config/tht.yaml (stessa convenzione CONFIG_OPT).
Testato phase show OK, suite 165 passed.
Report: docs/l2-run-report-2026-06-27.md.
Bug di porting emerso in L2 prep: tht session new falliva con
'AttributeError: SessionManifest has no to_yaml'. Lo stub locale _YamlModel in
session/models.py (placeholder pre-porting mschema) definiva solo
populate_by_name, senza i metodi to_yaml/from_yaml che store.py e session_cmd.py
usano. Aggiunti gli stessi metodi della controparte mschema (mantenendo
populate_by_name, necessario per db_schema alias='schema').
config/tht.yaml: symlink locale al workspace cliente attivo (psd.yaml). Il gate
chiama tht senza -c (default config/tht.yaml), quindi serve questo ponte per il
deployment per-cliente. Gitignored (per-cliente). .gitignore: + config/tht.yaml,
+ artifacts/.
Verifica: pytest 165 passed; tht session new crea sessione nel repo cliente;
pi vede i 3 tool reviewer_* (gate caricato); GLM 5.2 e' il model default di pi.
Due cambiamenti interconnessi da user review:
1. language come parametro workspace (spec decisione 9):
- Config.language (default 'en') + workspaces PSD con 'language: it'
- Generalizza Thoth oltre l'italiano: descrizioni tabelle/colonne ed evidence
sono nel workspace language; le istruzioni della skill restano in inglese
(piu' affidabili per modelli piccoli, meno ambigue)
2. Skill riscritta in INGLESE preservando la semantica COMPLETA dell'originale
(autocritica: la mia riscrittura precedente aveva perso ~10 vincoli precisi):
- 'promuovere' ambiguo (3 accezioni: phase advance / recommend / memory promote)
-> 'never advance a phase or record a decision without confirmation'
- recuperati vincoli persi: choice-is-confirmation (no reviewer_confirm dopo
reviewer_decide), reviewer_select SOLO per iterazione no-decision, messaggi
auto-contenuti obbligatori, artefatto = superficie di decisione (gate rilegge
da disco per CTE/SQL), candidati con provenienza+score non verita', opzione
'leave ambiguity open', F1 passa lista completa non solo ultima
- language contract esplicito (istruzioni EN, output nel workspace language)
Sottomoduli cte/memoria/rewriting/sql-generation in inglese, semantica tecnica
intatta (regole AV-SQL, dim_time trick, max 5 memorie solo 3 tipi riusabili).
Verifica: 0 residui nsp/chirone, tutti i tht <cmd> citati registrati, 165 passed.
Ultima onda CLI. 4 cmd portati con rename + grep-per-file (3 residui nsp nei messaggi
fixati). Nessun drift costanti phase in questi cmd.
La CLI tht e' ora COMPLETA: 14 gruppi di comandi (phase config schema session vector
memory search evidence db decision sql cte datamart lsh). tht --help li list tutti.
Suite: 165 passed.
Il loop skill->LLM->gate ora ha tutti i comandi che la skill chiamera'. Resta:
skill riscritta (S), setup pre-sessione (0b), sessione L2 manuale.
Correzione del gap ereditato (resosi NECESSARIO dal drop del registry, spec 5): il
metadata del VectorRecord memory ora porta subject/detail/rationale oltre a
type/session_id/tables/concepts. pack_metadata li serializza nel jsonb via
**record.metadata. search_similar proietta metadata completo -> la F2 ricostruisce
la decisione direttamente dall'hit, senza lookup registro.
L1: 4 test (subject/detail/rationale presenti, campi esistenti preservati,
no cross-contamination multi-record, save_one_memory propaga il metadata alla riga).
Suite: 165 passed.
Stesso bug del precedente: VENDORED.md e' arrivato con Onda 0 (dopo l'Onda -1 che
aveva pulito i riferimenti cliente). Neutralizzato a 'il datawarehouse' (coerente
con le altre neutralizzazioni). Sweep completo tht/ ora vuoto per chirone/psdwp3/
policlinico/sandonato.
User review ha trovato 2 residui 'nsp' sfuggiti al renaming: erano nei moduli
portati in Onda 0 (DOPO l'Onda -1 che aveva pulito), in messaggi utente/docstring
non in import. L'import-smoke di Onda 0 non li catturava (verifica solo import, non
stringhe). Corretti a tht: 'tht lsh build' (lshindex:61), 'tht schema introspect'
(VENDORED.md:25).
Lesson: dopo ogni port di file sorgente, grep di nsp su quel file, non solo import-smoke.
Suite: 161 passed, zero residui nsp nel codice.
require_phase_or_exit: guard riscritto vs Workflow (load_workflow().phase_name invece
della costante PHASE_NAMES drift). Exit 1 se la sessione e' sotto soglia. Usato da
cte/decision/datamart cmd.
Comandi phase (portati + adattati al modello ThothII, non copia cieca):
- advance: persiste phase_approved; --auto exit 6 se la fase non e' completa
(contratto col gate)
- reopen: persiste phase_reopened + teardown_to_phase degli artefatti oltre il target
- show: stato sessione (fase corrente, ultime decisioni)
session_dir helper tenuto qui (mirror di session_cmd) per evitare circular import.
_cfg() fa fallback a THT_WORKSPACE env finche' _load_config_or_exit (Onda 1.4) non
sara' portato.
L1: 4 test require_phase_or_exit (allow at/above, exit below, message con nome fase
dal workflow). Suite: 161 passed.
Aggiunge il metodo che i cmd CLI useranno al posto della vecchia costante
SCHEMA_LINKING_PHASE (drift fix Onda 1). Ritorna il num della fase il cui
artifacts_out contiene schema_linking.json, default 5 se nessuna la dichiara.
L1: 4 test (fase reale F4, posizione arbitraria, default 5, artefatti multipli).
Suite: 157 passed.
44 test L1 sui 3 moduli backend con logica non banale (opzione 2 della user review):
- sqlcheck.validate_sql (16 test): parse/single-statement, read-only enforcement
(INSERT/UPDATE/DELETE/CREATE/DROP/ALTER/TRUNCATE/GRANT rifiutati, WITH/UNION ok),
forbidden functions (dblink default blacklist, custom set, allowed not flagged),
object-existence (tabella inesistente, CTE non flaggata, perimetro promoted warning,
colonna inesistente con alias). Documenta una limitazione reale: le funzioni
aggregate specializzate (count/sum/coalesce) NON sono catturate dal name-matcher
perche' sqlglot modella .name come argomento, non come nome funzione.
- ctetest (14 test): has_trailing_select (semantica controintuitiva: True = violazione),
last_cte_name, build_test_sql, ledger I/O (load/append roundtrip, JSON-array e
JSONL tolleranti, corrupt-ledger raise).
- execute._inject_limit (6 test): LIMIT iniettato quando assente (limit+1 per
troncamento), rispettato quando presente, non iniettato su non-query, UNION/WITH ok.
Suite: 153 passed (109 + 44). Bonus: __psd_probe__ -> __tht_probe__ (riferimento
cliente neutralizzato in ctetest).