feat(opt): three efficiency levers for NL→SQL workflow
Lever 1: Join-graph via FK logics in annotations + suggest-fks command
- TableAnnotation.foreign_keys field stores curated logical FKs (DWH has no FK constraints)
- tht schema suggest-fks: mine from approved SQL, heuristics (time_key → dim_time),
same-name discovery + explicit --assume flag for multi-owner PKs
- mschema renders 【Foreign keys】 section populated; validation in merge.py
- SKILL.md F4 now reads FKs from mschema-text, no custom data_time_key logic
Lever 2: Context-pack consolidation at kickoff (tht search pack)
- Single embedding of question, reused for schema + evidence + solved searches
- One command: tht search pack <question> --session <id> → retrieval_pack.md
- Graceful degradation when Ollama/vector store unreachable (exit 0, empty sections)
- SKILL.md F1 prescribes as first call; reduces model thinking turns via pre-retrieval
Lever 3: Phase-summary recap v2 auto-construction from session ledger
- tht session show --json includes full decisions ledger
- tht phase meta --json exports 'emits' (substantive decision types per phase)
- Gate appends deterministic 【Decisioni registrate in questa fase】 section (appendLedgerSection)
- Model authors only summary + checks; recap table comes from persisted state (exact by construction)
- SKILL.md Disciplina 6: brief model output, gate fills the rest
Tests: 358 Python (including 10 FK + 3 pack + 1 session-ledger tests) + 111 JS gate tests, all pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -13,8 +13,8 @@ Rules (from the AV-SQL discipline, hold verbatim):
|
||||
3. Better one column too many than one too few: if unsure, include it.
|
||||
4. One CTE = one informative subset with a clear purpose (e.g. "ricoveri with
|
||||
ablazione in 2025"), named in a speaking snake_case.
|
||||
5. CTEs can chain-reference each other; the last one in the file is the one that
|
||||
`tht cte test` will query.
|
||||
5. CTEs can chain-reference each other **within the same file**; the last one in the
|
||||
file is the one that `tht cte test` will query (see the execution contract below).
|
||||
6. Each file in `sessions/<id>/ctes/<name>.sql` contains ONLY the `WITH ... AS (...)`
|
||||
block (multi-CTE allowed), WITHOUT a trailing SELECT. A `SELECT ...` line after
|
||||
the WITH block causes an error in `tht cte test`: never add it. CTEs are tested
|
||||
@@ -24,6 +24,24 @@ Rules (from the AV-SQL discipline, hold verbatim):
|
||||
7. Filters: use field values verified with `tht search` (LSH match on real values),
|
||||
not imagined values.
|
||||
|
||||
## How `tht cte test` executes (complete contract — do not read the harness source)
|
||||
|
||||
- Each CTE file is **standalone**: the test reads ONLY `ctes/<name>.sql`, appends
|
||||
`SELECT * FROM <last CTE defined in that file>` and runs it read-only against the
|
||||
DWH with an injected LIMIT and statement timeout. Cross-file references are NOT
|
||||
resolved: to build on an earlier CTE, repeat its definition in the same `WITH`
|
||||
chain (that is why multi-CTE files are allowed). A test normally completes in
|
||||
well under a second.
|
||||
- Before execution the SQL is validated statically: parsable, a single statement,
|
||||
read-only by structure, no blacklisted functions, and every referenced table must
|
||||
exist in the catalog (names defined in the `WITH` chain are exempt). Tables outside
|
||||
the promoted perimeter produce warnings. Unqualified columns are not statically
|
||||
checked — the DWH will catch them at run time.
|
||||
- Test order is enforced by the CLI from the approved plan (exit 5 names the CTE
|
||||
whose turn it is). Every outcome (ok or error) is appended to `cte_tests.json`;
|
||||
the gate reads the persisted file + last test via `tht cte info`, so never paste
|
||||
SQL, columns or preview rows into the gate text.
|
||||
|
||||
Presentation to the reviewer, for each CTE in the plan:
|
||||
|
||||
> **<name>** — purpose: <one line>
|
||||
|
||||
Reference in New Issue
Block a user