Files
ThothII/harness/.pi/skills/tht-sessione/sql-generation.md
T
marcopanandClaude Fable 5 e24b41b156 feat(opt): three efficiency levers for NL→SQL workflow
Lever 1: Join-graph via FK logics in annotations + suggest-fks command
  - TableAnnotation.foreign_keys field stores curated logical FKs (DWH has no FK constraints)
  - tht schema suggest-fks: mine from approved SQL, heuristics (time_key → dim_time),
    same-name discovery + explicit --assume flag for multi-owner PKs
  - mschema renders 【Foreign keys】 section populated; validation in merge.py
  - SKILL.md F4 now reads FKs from mschema-text, no custom data_time_key logic

Lever 2: Context-pack consolidation at kickoff (tht search pack)
  - Single embedding of question, reused for schema + evidence + solved searches
  - One command: tht search pack <question> --session <id> → retrieval_pack.md
  - Graceful degradation when Ollama/vector store unreachable (exit 0, empty sections)
  - SKILL.md F1 prescribes as first call; reduces model thinking turns via pre-retrieval

Lever 3: Phase-summary recap v2 auto-construction from session ledger
  - tht session show --json includes full decisions ledger
  - tht phase meta --json exports 'emits' (substantive decision types per phase)
  - Gate appends deterministic 【Decisioni registrate in questa fase】 section (appendLedgerSection)
  - Model authors only summary + checks; recap table comes from persisted state (exact by construction)
  - SKILL.md Disciplina 6: brief model output, gate fills the rest

Tests: 358 Python (including 10 FK + 3 pack + 1 session-ledger tests) + 111 JS gate tests, all pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 17:43:08 +02:00

53 lines
2.8 KiB
Markdown

# Final SQL generation technique
Adapted from the "query generation" step of AV-SQL (recursive divide-and-conquer)
and its review checklist.
## Generation (divide-and-conquer)
1. **Divide**: decompose the rewritten question into sub-questions, each aimed at a
piece of information or logic (a population, a filter, an aggregate).
2. **Conquer**: for each sub-question formulate a pseudo-SQL, with placeholders for
sub-questions not yet resolved. The CTEs already tested in Phase 6 are the
preferred building blocks: reuse them by name, with their known outcome.
3. **Recombine**: replace placeholders bottom-up until the full SQL. The final SQL
may include the CTEs in its own WITH.
4. Dialect: PostgreSQL. Copy table and column names EXACTLY from the schema context;
never invent objects.
The file `sessions/<id>/sql_final.sql` must contain ONLY the SQL, clean and
copy-pasteable: no rationale comments (that lives in the audit artifacts).
## Time dimension (analysis by year/month/quarter)
Fact tables have `data_time_key` (`integer`, format `YYYYMMDD`): it is the FK to
`dim_time.day_key`. The DWH does not declare it, but the workspace annotations do:
it appears in the `【Foreign keys】` section of the mschema-text render (every
`*_time_key` column maps to `dim_time.day_key`) — take it from there for the join.
- To extract year, month, quarter, semester etc. do
`JOIN dim_time dt ON dt.day_key = <fact>.data_time_key` and use the dimension's
columns: `dt.year`, `dt.month`, `dt.quarter`, `dt.semester`, `dt.full_date`,
`dt.month_name_it`, `dt.year_month`.
- **Do NOT** do arithmetic on the key (e.g. `data_time_key / 10000` for the year):
it works by accident but is fragile and breaks as soon as you need to format a date
or do a cast. Always use `dim_time`.
- "Last N years from the most recent year":
`dt.year >= (SELECT MAX(year) FROM dim_time WHERE day_key IN (SELECT data_time_key FROM <fact>)) - (N-1)`,
or compute the max year on the rows actually present in the fact.
## Review checklist (on errors or suspicious results)
- Do the returned columns answer the question exactly?
- Do the filters (WHERE/HAVING) reflect ALL the conditions of the rewritten question?
- Are aggregations, groupings and orderings the required ones?
- Empty or zero result: almost always indicates a problem in conditions or joins.
Verify the filter values with `tht search find "<value>"` (match on real values).
- Do the joins follow those promoted in `schema_linking.json`? (exception: the FK
`data_time_key → dim_time.day_key` is not declared, see the time section.)
- Time analyses: are you using `JOIN dim_time` and not key arithmetic?
Every substantive revision is recorded with:
`reviewer_decide(options:[{label:"Register revision", type:"sql_revised",
subject:"sql_final", detail:"<what changed>", rationale:"<why>"}], allow_other:false)`.