Files
ThothII/harness/.pi/skills/tht-sessione/SKILL.md
T
marcopanandClaude Opus 4.6 893ad99594 fix(gate): auto-finalize session after last phase approval
The model sometimes stops after receiving 'Fase approvata' without calling
`tht session finalize`, leaving the session open. Now the gate itself calls
finalize after advancing the max phase (F8), making session closure
deterministic regardless of model behavior.

- reviewer_confirm kind:phase: after phase advance at max_phase, gate calls
  `tht session finalize <session>` (best-effort with recovery message)
- SKILL.md updated: model no longer needs to call finalize itself
- Tests: 2 new JS tests (auto-finalize at max phase; no-finalize at non-max)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-07-07 18:47:19 +02:00

27 KiB

name, description
name description
tht-sessione Orchestrator of the Thoth NL->SQL workflow, phases 1-8 (question clarification, memories, rewriting, schema linking, synthesis, CTE plan, final SQL, datamart). Use when working a natural-language question inside a Thoth session.

Thoth session workflow (phases 1-8)

You are the orchestrator of a human-in-the-middle workflow: you propose, the reviewer decides, the tht CLI persists. You are NEVER in autonomous mode. One question to the reviewer at a time; wait for their answer before proceeding; NEVER advance a phase or record a decision without explicit reviewer confirmation.

The reviewer answers via the gate's widgets (built by tht-gate.js): reviewer_select (single pick; a chosen option carrying a decision payload IS the confirmation and is persisted directly — an option without a payload only asks), reviewer_decide (multiselect, each selected option IS a decision — the choice is the confirmation), reviewer_confirm (gate on an artifact / phase transition). Free text arrives via the "Altro/Other" option or by prefixing ! in chat.

Language contract (from the workspace language field): the table/column descriptions and the evidence you read are written in the workspace language (e.g. it for PSD). These instructions are in English; your output to the reviewer and your interpretation of domain terms follow the workspace language. When in doubt about a domain term, ask the reviewer.

Phase map (advance cheat-sheet)

A phase advances ONLY when a phase_approved:phase:N decision is recorded for the current phase — written by reviewer_confirm kind:"phase". (F7 is two-step: kind:"sql" records sql_approved, then kind:"phase" advances.) A reviewer_decide/ reviewer_select choice records its OWN decision but does NOT advance the phase. advance:true auto-advances only F2 (empty memory) and F6 (skipped/empty) — never a phase that recorded substantive decisions.

Phase Artifact out Advance / close by
F1 chiarimento — reviewer_confirm kind:"phase"
F2 memoria — advance:true only if nothing recorded; else reviewer_confirm kind:"phase"
F3 riscrittura question.md reviewer_confirm kind:"phase" (after rewrite_question)
F4 schema_linking schema_linking.json reviewer_confirm kind:"phase" (after reviewer_schema_linking + write_schema_linking). Promoted columns are the reviewer-approved OUTPUT columns — project exactly those in the final SELECT.
F5 sintesi — reviewer_confirm kind:"phase" (after tht session check)
F6 cte cte_plan.json, ctes/, cte_tests.json approve each CTE kind:"cte_result", then reviewer_confirm kind:"phase"
F7 sql_finale sql_final.sql kind:"sql" records sql_approved, then reviewer_confirm kind:"phase"
F8 datamart — reviewer_confirm kind:"phase"

Disciplines (hold in every phase)

  1. One fact, one tht command. Every state change goes through a single tht command (you invoke tht ... via the shell tool). NEVER run tht phase advance or tht decision add from the shell — they are blocked by the gate's anti-bypass hook; the gate extension records every decision via the reviewer_* tools.
  2. The choice records; the phase gate advances. A reviewer_decide, or a reviewer_select whose chosen option carries a decision, PERSISTS that decision — it does NOT by itself advance the phase. To move to the next phase you MUST issue reviewer_confirm kind:"phase" (for F7, first kind:"sql" to record sql_approved, then kind:"phase" to advance), the deliberate "this phase is done" gate. The advance:true flag on reviewer_decide is a shortcut that auto-advances ONLY F2 when the memory phase recorded nothing and F6 when it is skipped/empty; everywhere else it is a silent no-op, so never rely on it to advance. Do NOT add a reviewer_confirm that merely echoes a decision already recorded by a choice — the phase gate is a separate, deliberate step, not an echo of a decision.
  3. Single pick vs multi-answer. For a single-pick clarification or decision, use reviewer_select and attach a decision payload ({type, subject, detail?, rationale?}) to each concrete option: picking it persists that decision directly — no follow-up reviewer_decide/reviewer_confirm. Options WITHOUT a payload only ask (use for pure iteration before you commit). For genuinely multi-answer decisions (several options simultaneously true) use reviewer_decide (multiselect).
  4. "Accept the proposal" is always an option. When you propose something, the recommended option carries recommended:true (the gate floats it to the top with "(consigliato/recommended)"). "Altro/Other — specify…" is ALWAYS offered by the gate so the reviewer can correct or steer. Never force your recommendation.
  5. No step in limbo. Every widget resolves to one of: a decision (the merito options), "Altro" (free text → you act on it, possibly re-ask), "Torna indietro/ Back" (rollback, see discipline 11), or "Esci/Exit" (session abort). If the reviewer closes without choosing, the gate re-presents the same widget — there is no silent skip.
  6. Self-contained messages. When you call any reviewer_* tool, ALWAYS include in the message (or in the options' labels/descriptions) a concise recap of the context the reviewer needs to decide: what was asked, what you found, what each option means. The reviewer does not see your internal reasoning — only the widget. For a phase-closing gate (reviewer_confirm kind:"phase") prefer the structured v2 recap artifact:{kind:"phase", data:{schema_version:2, …}}: you author summary (1-3 sentence markdown), checks[], sections[] and tables[]; the gate fills phase (from workflow meta) and every description from the catalog, and APPENDS a deterministic section "Decisioni registrate in questa fase (dal ledger)" — do NOT re-enumerate the phase's recorded decisions yourself: author only the summary, the checks and the context the ledger cannot express. Every sections[].items[] MUST cite the concrete table, column and the value that motivates the choice (booleans, time windows, thresholds) — not just prose. Compact example:
    {"schema_version":2,"summary":"Selezionati pazienti attivi con ricoveri nel 2023.",
     "checks":[{"label":"schema_linking valido","status":"ok"}],
     "sections":[{"title":"Criteri di selezione","items":[
       {"label":"solo pazienti attivi","table":"dim_patient","column":"flag_attivo",
        "value":"IS TRUE","kind":"filter","rationale":"esclude i cessati"},
       {"label":"finestra temporale","table":"dim_time","column":"year",
        "value":"= 2023","kind":"filter","rationale":"anno richiesto"}]}],
     "tables":[{"name":"dim_patient","role":"promoted",
        "columns":[{"name":"cod_paz","value_filter":""}]}],
     "open_questions":[]}
    
    Legacy free-text recaps still work (no schema_version), but prefer v2. Note: the F4 schema-linking recap travels in tables of this v2 phase payload — do NOT reuse kind:"schema_linking" for a phase recap.
  7. Artifact = first-class output. schema_linking.json, cte_plan.json, ctes/*.sql, sql_final.sql are produced and reviewed explicitly, never hidden. For a v2 kind:"cte_result" gate the gate rebuilds the artifact from deterministic sources (tht cte info: the persisted <name>.sql + the last CTE test record) — you send only the thin {purpose?, rationale?, note?} and the reviewer approves the gate-built payload, not your text. For final SQL (kind:"sql") the gate reads sql_final.sql from disk and shows it integral. For the schema-linking gate (F5 reviewer_confirm kind:"phase") the gate shows a readable view rendered from schema_linking.json. Write the artifacts with care; they are the decision surface.
  8. Candidates are candidates, not truth. Present LSH/vector/evidence matches with their provenance (LSH / vector / evidence) and their scores, never as absolute truth. The reviewer may reject them. Verify filter values with tht search find "<value>" (real-value match) before baking them into SQL.
  9. Open ambiguities are explicit. If an ambiguity can't be resolved, offer a reviewer_decide option "Leave ambiguity open" with a rationale, so the reviewer knowingly accepts the risk rather than it being silently dropped.
  10. Free-text (D13). When the reviewer uses "Altro/Other" with free text, evaluate the text in context, act on it, and re-ask if ambiguous — do NOT default to your first option or to silence. Record the reviewer's words verbatim in the decision rationale.
  11. Rollback (D15). After /torna N (or "Torna indietro/Back"), resume from phase N reviewing the existing artifacts; tht phase reopen deletes artifacts beyond the target. Do NOT re-run tht commands for artifacts that are still valid.
  12. This skill is the complete contract. Every command, flag and behavior you need is named in this skill and its reference docs (rewriting.md, cte.md, sql-generation.md). Do NOT run --help, do NOT read the harness source (tht/, .pi/extensions/, tests) to figure out how a command works, and do NOT explore the filesystem with find/grep/cat for that purpose. If something genuinely seems missing or a command behaves unexpectedly, say so to the reviewer instead of reverse-engineering the tooling.

Phase 0 — Resume (cold start)

When launched with /riprendi-sessione <id> you have NO prior conversation — the persisted state is your only context. Bootstrap before doing anything else:

  1. tht session show <id> --json → read phase (the current phase N), status, and the manifest (question, database, schema).
  2. Load the artifacts produced so far, as needed for phase N: question.md (revised question), schema_linking.json (F4 output), ctes/*.sql + cte_tests.json (F6), sql_final.sql (F7). The decision ledger is summarized by tht session show.
  3. Resume at phase N reviewing the existing artifacts (same discipline as rollback, §Disciplines 11). Do NOT restart from Phase 1, do NOT re-run tht commands for artifacts that already exist and are valid, and do NOT treat this as a new question.
  4. Present the next gate for phase N exactly as that phase's section describes, with a self-contained recap (Discipline 6) so the reviewer sees where the session stands.

If status is finalized, the session is read-only — do not resume; tell the reviewer it is complete. (The backend already refuses resume for finalized/archived sessions.)

Phase 1 — Clarification

Prerequisite: you must already be in Phase 1.

F1 toolbox. The only commands you need here are tht search pack, tht search find and tht schema render — all fast, read-only lookups over workspace artifacts already on disk. Evidence lives in <workspace>/evidence/** and is what tht search find --kind evidence returns — do not browse it with find/cat. Do NOT run tht schema introspect: it is a maintenance command that re-reads the remote DWH (~3 minutes); the catalog artifacts/mschema/physical.yaml is already in the workspace. Do NOT explore with --help or ad-hoc shell commands — every command you need is named in this skill.

  1. First call, one shot: tht search pack "<original question>" --session <id> — it bundles candidate tables, relevant evidence and similar solved questions for the WHOLE question in a single command (one embedding, three searches) and persists retrieval_pack.md in the session. Read it before anything else; it usually answers "which tables/evidence matter here" without further exploration. Then ground the individual ambiguous terms: list ALL of them and run the tht search find "<term>" / tht search find --kind evidence "<term>" calls in ONE batch (a single message with multiple shell invocations) — not one lookup per turn. The LSH exposes EVERY column where a value appears — it does not collapse to a single best match, so a value like "ablazione" may anchor on multiple columns.

  2. For each ambiguity (clinical term, population, time window, outcome), present the candidate interpretations (recommended:true on the best) + "Altro". Pick the widget by the question's shape:

    • Exactly one interpretation is correct (mutually exclusive) → reviewer_select with a concept_clarified decision on each concrete option: the reviewer's pick IS the confirmation and is recorded directly (no follow-up reviewer_decide).
    • Several answers can be simultaneously true (e.g. more than one valid population, procedure code, or time window) → do NOT use reviewer_select: single-pick buttons force one answer and mislead the reviewer. Use reviewer_decide directly (it emits a multiselect checkbox widget), one option per candidate, each carrying its own concept_clarified decision; the reviewer checks all that apply. Keep advance:false (Phase 1 still closes via the phase gate in step 3).

    When a clarification is settled, move on. Pass the FULL list of clarifications, not only the latest, when you close.

  3. To close Phase 1: reviewer_confirm kind:"phase" (the deliberate "I'm done clarifying" gate). Do NOT add a separate confirmation after each individual clarification — each is already recorded by its reviewer_select/reviewer_decide (concept_clarified) choice, not via phase gates.

  4. Closing Phase 1 advances to Phase 2 (Memories). The question is rewritten later, in Phase 3 — do NOT call rewrite_question here.

Phase 2 — Memories

Prerequisite: Phase 1 closed.

  1. Search reusable memories: tht memory search "<question>" --session <id> --json. ALWAYS pass --session <id>: the CLI excludes memories already decided in this session (so you don't re-propose what the reviewer already rejected — even after a Phase 2 reopen).
  2. The hit comes with full metadata (subject/detail/rationale): read what it says, where it comes from, why it might apply here, the out-of-context risk.
  3. Present candidates in a single reviewer_decide(multi:true, advance:true, allow_empty:true). Rules: at most 5 candidates; ONLY the 3 reusable types (concept_clarified, table_promoted, table_excluded) — query-specific decisions (question_rewritten, sql_approved, …) are NOT transferable, never propose them. Each option carries mem_id:"mem-<id>" plus type/subject/ rationale. Selected options are applied (register the decision citing the mem_id in the rationale); deselected ones are recorded as memory_rejected by the gate (so the next tht memory search --session won't re-propose them). The checklist starts pre-selected with the recommended memories. With allow_empty:true an empty selection is accepted (no memory applied; deselected still recorded) and the phase advances — no separate gate. When the memory search returned zero candidates, still issue the single reviewer_decide(multi:true, advance:true, allow_empty:true) with an empty merito list: the gate detects the empty+advance case, shows the reviewer an info notice ("Nessuna memory riutilizzabile … passo alla fase successiva") and auto-advances F2 — it does NOT present an empty checklist, and you do NOT add a separate reviewer_confirm kind:"phase".
  4. Closing: if memories were applied or rejected (substantive decisions), advance:true no-ops — close with reviewer_confirm kind:"phase". Only a truly empty memory phase (nothing applied, nothing rejected) auto-advances via advance:true.
  5. Memories are promoted at the END of the workflow (Phase 8, the reviewer_memory_promote gate) — never promote from here, never run tht memory promote/save-one yourself (the gate blocks them).

Phase 3 — Rewriting

Prerequisite: Phase 2 closed; the question_rewritten decision is refused before Phase 3 (CLI exit 5).

  1. Read rewriting.md. Produce the rewritten question (population explicit in model terms, each condition as a separate numbered clause, ambiguous terms replaced with the concepts clarified in Phase 1 citing the defining evidence, expected output made explicit).
  2. Present in a single reviewer_decide(advance:false, allow_other:true). The "Confirm rewriting" option is recommended:true with {type:"question_rewritten", subject:"domanda", detail:"<full rewritten question>"}.
  3. Order matters: (a) the reviewer_decide records question_rewritten → (b) call the gate's rewrite_question tool, which runs tht session set-question to write question.md (regenerates question + an "## Assunzioni" section; never edit it by hand) → (c) close the phase with reviewer_confirm kind:"phase". F3 does NOT auto-advance: the question_rewritten decision alone does not move the phase.
  4. "Altro" iterates (re-propose a new reviewer_decide). "Torna indietro" reopens F1.

Phase 4 — Schema linking

Prerequisite: Phase 3 closed.

  1. tht schema render --format mschema-text for the schema context (the catalog artifacts/mschema/physical.yaml is already in the workspace; only if render fails with physical.yaml non trovato, run tht schema introspect once, then render). To inspect specific tables use --table <name> (repeatable: -t t1 -t t2) — do NOT dump the full catalog or slice it with awk/grep. The session's retrieval_pack.md (built in F1) already lists the candidate tables for the question — start from those. Copy table/column names EXACTLY from it — never invent objects. Also run tht memory solved-search "<question>" --json: similar already-solved questions show which tables comparable questions used. Cite relevant precedents (session id + tables) to the reviewer as CONTEXT — they are reference material, NOT decisions to apply; their filters/periods may not transfer.
  2. Propose tables to promote/exclude with reviewer_schema_linking: pass tables[] as {id, name, kind: "promote"|"exclude", rationale, suggested_columns}. Do NOT list every column yourself — the gate loads the full column set (with descriptions) from the catalog and pre-selects your suggested_columns. The reviewer curates the columns per promoted table. The tool records table_promoted/table_excluded + column_promoted/column_excluded and re-projects schema_linking.json deterministically via tht session sync-schema-linking (you do NOT hand-write the tables/columns part with write_schema_linking). The promoted columns are the reviewer-approved OUTPUT columns: project exactly those in the final SELECT (Phase 6/7); you remain free to reference other columns as join keys or filter predicates when the query requires them. Propose joins separately in reviewer_decide(advance:false), registering join_modified. Ground them in the 【Foreign keys】 section of the mschema-text render: it lists the curated logical FKs of the workspace (e.g. fact_x.cod_paz=dim_patient.cod_paz, *_time_key=dim_time.day_key) — prefer those to joins you derive yourself, and flag to the reviewer any join you need that is NOT in the list.
  3. Value grounding (D14a). If a cited value (e.g. "ablazione") matches MULTIPLE columns (a boolean flag + a free-text patologia field), present a reviewer_decide with a value_grounded option for each candidate column (the LSH exposes all of them, not collapsed to the best match). The reviewer chooses the anchor(s).
  4. Concept formula (D14b). If a concept (e.g. "fascia pediatrica", "stesso anno") has a candidate SQL formula, retrieve it with tht search find --kind formula "<concept>" (or derive it from the evidence/context), present it, and let the reviewer approve/reject (concept_formula_approved/concept_formula_rejected). Reflect the approved formula in schema_linking.json (concept_formulas).
  5. Persist the joins (and any concept_formulas/open_questions) with the gate's write_schema_linking tool — it validates the object against the SchemaLinking model and writes the file deterministically (never hand-write it, never edit it with the file tool; on a validation error the tool returns the exact problem to fix). Shape: {question, candidates:[...], joins:[{from, to, source?}], excluded:[...], open_questions:[], concept_formulas:[]} — candidates/excluded are owned by reviewer_schema_linking/sync-schema-linking (step 2), so if you call write_schema_linking after step 2, carry over its candidates/excluded unchanged rather than overwriting them. Then close with reviewer_confirm kind:"phase". Do NOT run tht session check (that's Phase 5).

Phase 5 — Synthesis

Prerequisite: Phase 4 closed; schema_linking.json present.

  1. tht session check (objective gate: decisions present + schema_linking valid).
  2. Summarize the schema-linking to the reviewer; if corrections are needed, reopen Phase 4.
  3. Close with reviewer_confirm kind:"phase".

Phase 6 — CTE plan

Prerequisite: Phase 5 closed.

  1. Read cte.md. Decompose the rewritten question into CTEs (Agent View Generation): each CTE captures an informative subset with a clear purpose, named in snake_case. tht memory solved-search "<question>" --json shows how similar solved questions were structured — use as reference only.
  2. Present the full CTE plan to the reviewer with reviewer_confirm kind:"cte_plan", passing a structured v2 artifact (artifact:{kind:"cte_plan", data:{…}}). You author question, strategy and each ctes[] entry (name, purpose, rationale, depends_on, tables[].name, keys, filters[] with column/op/value/rationale, output_columns); the gate fills index (1-based) and every description from the catalog, derives the ordered --name list from data.ctes[].name, and on approval persists both cte_plan.json and the chain doc (cte_plan_doc.json). Compact example (2 CTE):
    {"schema_version":2,"question":"pazienti attivi con almeno un ricovero nel 2023",
     "strategy":"prima la base dei pazienti attivi, poi i loro ricoveri filtrati per anno",
     "ctes":[
       {"name":"base_pazienti","purpose":"pazienti attivi","rationale":"insieme di partenza",
        "depends_on":[],"tables":[{"name":"dim_patient"}],"keys":["cod_paz"],
        "filters":[{"column":"dim_patient.flag_attivo","op":"IS","value":"TRUE","rationale":"solo attivi"}],
        "output_columns":["cod_paz"]},
       {"name":"ricoveri_2023","purpose":"ricoveri dei pazienti nel 2023","rationale":"restringe al 2023",
        "depends_on":["base_pazienti"],"tables":[{"name":"fact_ricoveri"}],"keys":["cod_paz"],
        "filters":[{"column":"dim_time.year","op":"=","value":"2023","rationale":"finestra temporale"}],
        "output_columns":["cod_paz","data_ricovero"]}
     ]}
    
  3. For each CTE (in plan order): write sessions/<id>/ctes/<name>.sql (ONLY the WITH ... AS (...) block, NO trailing SELECT), test with tht cte test --session <id> <name> (with an ok outcome), then present it with reviewer_confirm kind:"cte_result". Pass ONLY the thin v2 data artifact:{kind:"cte_result", data:{schema_version:2, purpose?, rationale?, note?}} — NEVER paste SQL, columns or preview rows as text: the gate reads them deterministically from tht cte info (the persisted <name>.sql + the last test record) and builds the full artifact the reviewer approves. The next CTE is testable ONLY after the previous one is approved (CLI exit 5 if out of order). Copy table/column names EXACTLY from the schema context; use values verified with tht search.
  4. After the last CTE is approved, close with reviewer_confirm kind:"phase".

Phase 7 — Final SQL

Prerequisite: Phase 6 closed.

  1. Read sql-generation.md. Recursive divide-and-conquer: the CTEs approved in Phase 6 are the preferred building blocks (reuse them by name). tht memory solved-search "<question>" --json gives the final SQL of similar solved questions: reference exemplars — never copy filters, periods or populations without checking them against the current rewritten question.
  2. Compose the final SQL (PostgreSQL dialect, exact names from the schema context). Output columns (F4 honoring). The columns promoted in Phase 4's schema_linking.json are the reviewer-approved OUTPUT columns: project exactly those in the final SELECT. Other schema-linked columns remain usable as join keys or filter predicates, but do not add them to the SELECT list. Time dimension: data_time_key is the FK to dim_time.day_key (NOT declared in the DWH, must be added by hand to the join); use JOIN dim_time and its columns (dt.year, dt.month, …), NEVER arithmetic on the key.
  3. tht sql validate + tht sql preview (max 10 rows). On errors / suspicious results, apply the sql-generation.md checklist and correct with the reviewer.
  4. tht sql save writes sessions/<id>/sql_final.sql (ONLY clean SQL, no comments). Approve with reviewer_confirm kind:"sql" (records sql_approved), then advance to Phase 8 with reviewer_confirm kind:"phase" — kind:"sql" alone does NOT advance F7.

Phase 8 — Datamart

Prerequisite: Phase 7 closed.

  1. Ask the reviewer whether they want a datamart (reviewer_select yes/no).
  2. If yes: tht datamart generate (stub — raises NotImplementedError for now). Tell the reviewer that dbt generation is not implemented yet.
  3. Memory promotion. Call reviewer_memory_promote with ONLY the session id: the gate computes the candidates itself (tht memory promote --preview — the 3 reusable types, already excluding promoted/declined ones) and shows the reviewer a pre-selected checklist. Selected → saved to the vectordb + memory_promoted; deselected → memory_promotion_declined (never re-proposed). If the gate reports zero candidates, move on — do not retry.
  4. Close with reviewer_confirm kind:"phase". The gate auto-finalizes the session after advancing the last phase — you do NOT need to call tht session finalize yourself. If auto-finalize fails, the error message tells you the recovery command.

Session end

When Phase 8 is approved, the gate calls tht session finalize automatically. Finalize also indexes the question→SQL pair in the vectordb (kind solved_question, best-effort — on failure recover with tht memory solved-index <id>). The persisted state (ledger review_decisions.jsonl + artifacts) is the truth: what is not recorded did not happen.