feat(opt): three efficiency levers for NL→SQL workflow
Lever 1: Join-graph via FK logics in annotations + suggest-fks command
- TableAnnotation.foreign_keys field stores curated logical FKs (DWH has no FK constraints)
- tht schema suggest-fks: mine from approved SQL, heuristics (time_key → dim_time),
same-name discovery + explicit --assume flag for multi-owner PKs
- mschema renders 【Foreign keys】 section populated; validation in merge.py
- SKILL.md F4 now reads FKs from mschema-text, no custom data_time_key logic
Lever 2: Context-pack consolidation at kickoff (tht search pack)
- Single embedding of question, reused for schema + evidence + solved searches
- One command: tht search pack <question> --session <id> → retrieval_pack.md
- Graceful degradation when Ollama/vector store unreachable (exit 0, empty sections)
- SKILL.md F1 prescribes as first call; reduces model thinking turns via pre-retrieval
Lever 3: Phase-summary recap v2 auto-construction from session ledger
- tht session show --json includes full decisions ledger
- tht phase meta --json exports 'emits' (substantive decision types per phase)
- Gate appends deterministic 【Decisioni registrate in questa fase】 section (appendLedgerSection)
- Model authors only summary + checks; recap table comes from persisted state (exact by construction)
- SKILL.md Disciplina 6: brief model output, gate fills the rest
Tests: 358 Python (including 10 FK + 3 pack + 1 session-ledger tests) + 111 JS gate tests, all pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -97,8 +97,12 @@ def combined_search(
|
||||
top: int,
|
||||
rrf_k: int,
|
||||
kinds: list[str] | None,
|
||||
query_vec: list[float] | None = None,
|
||||
) -> list[SearchResult]:
|
||||
"""Fonde LSH (valori di campo) e pgvector con Reciprocal Rank Fusion."""
|
||||
"""Fonde LSH (valori di campo) e pgvector con Reciprocal Rank Fusion.
|
||||
|
||||
`query_vec` permette di riusare un embedding gia' calcolato della stessa
|
||||
keyword (es. `tht search pack`, che fa piu' ricerche sulla stessa domanda)."""
|
||||
rankings: dict[str, list[tuple[str, float]]] = {}
|
||||
lsh_values: dict[str, str] = {}
|
||||
if lsh_hits:
|
||||
@@ -106,7 +110,9 @@ def combined_search(
|
||||
rankings["lsh"] = [(key, score) for key, score, _ in aggregated]
|
||||
lsh_values = {key: value for key, _, value in aggregated}
|
||||
|
||||
vector_hits = store.search(embedder.embed_query(keyword), top_n=top * 2, kinds=kinds)
|
||||
if query_vec is None:
|
||||
query_vec = embedder.embed_query(keyword)
|
||||
vector_hits = store.search(query_vec, top_n=top * 2, kinds=kinds)
|
||||
rankings["vector"] = [(_vector_key(h), h.similarity) for h in vector_hits]
|
||||
by_key = {_vector_key(h): h for h in vector_hits}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user