feat(opt): three efficiency levers for NL→SQL workflow

Lever 1: Join-graph via FK logics in annotations + suggest-fks command
  - TableAnnotation.foreign_keys field stores curated logical FKs (DWH has no FK constraints)
  - tht schema suggest-fks: mine from approved SQL, heuristics (time_key → dim_time),
    same-name discovery + explicit --assume flag for multi-owner PKs
  - mschema renders 【Foreign keys】 section populated; validation in merge.py
  - SKILL.md F4 now reads FKs from mschema-text, no custom data_time_key logic

Lever 2: Context-pack consolidation at kickoff (tht search pack)
  - Single embedding of question, reused for schema + evidence + solved searches
  - One command: tht search pack <question> --session <id> → retrieval_pack.md
  - Graceful degradation when Ollama/vector store unreachable (exit 0, empty sections)
  - SKILL.md F1 prescribes as first call; reduces model thinking turns via pre-retrieval

Lever 3: Phase-summary recap v2 auto-construction from session ledger
  - tht session show --json includes full decisions ledger
  - tht phase meta --json exports 'emits' (substantive decision types per phase)
  - Gate appends deterministic 【Decisioni registrate in questa fase】 section (appendLedgerSection)
  - Model authors only summary + checks; recap table comes from persisted state (exact by construction)
  - SKILL.md Disciplina 6: brief model output, gate fills the rest

Tests: 358 Python (including 10 FK + 3 pack + 1 session-ledger tests) + 111 JS gate tests, all pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-07 17:43:08 +02:00
co-authored by Claude Fable 5
parent 87e875bc81
commit e24b41b156
19 changed files with 936 additions and 29 deletions
+20 -5
View File
@@ -1,11 +1,25 @@
from typing import Any
from tht.mschema.eligibility import effective_eligibility
from tht.mschema.models import Annotations, ColumnAnnotation, PhysicalSchema
from tht.mschema.models import Annotations, ColumnAnnotation, ForeignKey, PhysicalSchema
MAX_EXAMPLES_IN_PROMPT = 5
def table_foreign_keys(
physical: PhysicalSchema, annotations: Annotations, table: str
) -> list[ForeignKey]:
"""FK fisiche + FK logiche dalle annotations (dedup su columns/ref)."""
fks = list(physical.tables[table].foreign_keys)
ann = annotations.tables.get(table)
if ann:
seen = {(tuple(f.columns), f.ref_table, tuple(f.ref_columns)) for f in fks}
for fk in ann.foreign_keys:
if (tuple(fk.columns), fk.ref_table, tuple(fk.ref_columns)) not in seen:
fks.append(fk)
return fks
def _ann_col(annotations: Annotations, table: str, column: str) -> ColumnAnnotation | None:
ann = annotations.tables.get(table)
if ann is None:
@@ -59,7 +73,7 @@ def to_mschema_text(
shown = ", ".join(column.examples[:MAX_EXAMPLES_IN_PROMPT])
lines.append(f" -- Examples: {shown}")
lines.append(");")
for fk in table.foreign_keys:
for fk in table_foreign_keys(physical, annotations, table_name):
for src, dst in zip(fk.columns, fk.ref_columns):
fk_lines.append(f"{table_name}.{src}={fk.ref_table}.{dst}")
lines.extend(["", "【Foreign keys】", *fk_lines])
@@ -89,7 +103,7 @@ def to_schema_dict(
"primary_keys": [c for c in cols if table.columns[c].pk],
"foreign_keys": [
{"columns": fk.columns, "ref_table": fk.ref_table, "ref_columns": fk.ref_columns}
for fk in table.foreign_keys
for fk in table_foreign_keys(physical, annotations, table_name)
],
}
return out
@@ -132,9 +146,10 @@ def to_markdown(physical: PhysicalSchema, annotations: Annotations | None = None
f"| {column_name} | {column.type} | {'sì' if column.nullable else 'no'} "
f"| {'sì' if column.pk else ''} | {cdesc} | {examples} |"
)
if table.foreign_keys:
fks = table_foreign_keys(physical, annotations, table_name)
if fks:
lines += ["", "Foreign keys:"]
for fk in table.foreign_keys:
for fk in fks:
lines.append(
f"- ({', '.join(fk.columns)}) → {fk.ref_table} ({', '.join(fk.ref_columns)})"
)