Implementazione del piano di remediation progressiva sui difetti emersi dall'analisi dell'harness. Tutto verificato: 214 test Python (incl. L0 su Postgres reale), 14 test JS del gate, ruff pulito. Blocco 1 (CRITICA, integrazione gate↔CLI): - phase advance: gate usa --auto + exit 6; reviewer_confirm kind:phase fa advance esplicito che applica i prerequisiti (prima non avanzava per le fasi a conferma umana). - cte plan riceve i --name dal gate (param names); set-question con id posizionale; skill `tht search find`; nuovo comando `tht memory save-one` con dedup hash client-side in save_one_memory. Blocco 2 (D15, stato post-rollback): - campo `phase` su DecisionRecord + effective_decisions phase-aware per i subject "a nome" (cte_approved ecc.); _compute_promotions e finalize sulla vista effective; finalize confronta col piano CTE effettivo, non glob; `decision add --retracts` + comando `decision retract`. Blocco 3 (D7 read-only + D6 manifest): - assert_read_only su tutti e quattro i codepath (direct + REST); - manifest author/summary/updated_at/updated_by/schema_version popolati + helper touch_manifest sulle mutazioni. Blocco 4-5 (D14a/D14b): - decision_min_phase data-driven via `emits:` in workflow.yaml; - formula evidence: status auto, search_formulas, gruppo CLI `tht formula`, `search find --kind formula`, load_evidence_dir salta i .sql.md. Blocco 6 (robustezza): - taskdoc slice promoted_tables + bound enforced; report escaping/bound + rsplit note; filtro kind reader REST/direct; conteggio upserted robusto; guard REST run_query non-list; LSH disallineato -> LshIndexError. Blocco 7 (pulizia): - dead code gate e KIND_TO_TABLE morto rimossi; doc Postgres-only (README + connection.py). Blocco 0 (parziale): test di compatibilità firma gate↔CLI (tests/integration). Rinviati: fake-Pi runtime completo, artifact-gate da disco (#23), parità eligibility REST/direct (#28), unificazione reserved-labels (#30), memory_rejected da deselezione (#33). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
53 lines
2.8 KiB
Markdown
53 lines
2.8 KiB
Markdown
# Final SQL generation technique
|
|
|
|
Adapted from the "query generation" step of AV-SQL (recursive divide-and-conquer)
|
|
and its review checklist.
|
|
|
|
## Generation (divide-and-conquer)
|
|
|
|
1. **Divide**: decompose the rewritten question into sub-questions, each aimed at a
|
|
piece of information or logic (a population, a filter, an aggregate).
|
|
2. **Conquer**: for each sub-question formulate a pseudo-SQL, with placeholders for
|
|
sub-questions not yet resolved. The CTEs already tested in Phase 6 are the
|
|
preferred building blocks: reuse them by name, with their known outcome.
|
|
3. **Recombine**: replace placeholders bottom-up until the full SQL. The final SQL
|
|
may include the CTEs in its own WITH.
|
|
4. Dialect: PostgreSQL. Copy table and column names EXACTLY from the schema context;
|
|
never invent objects.
|
|
|
|
The file `sessions/<id>/sql_final.sql` must contain ONLY the SQL, clean and
|
|
copy-pasteable: no rationale comments (that lives in the audit artifacts).
|
|
|
|
## Time dimension (analysis by year/month/quarter)
|
|
|
|
Fact tables have `data_time_key` (`integer`, format `YYYYMMDD`): it is the FK to
|
|
`dim_time.day_key`. **This FK is NOT declared** in the DWH (facts have
|
|
`foreign_keys: []`), so it will NOT appear in `schema_linking.json`: you must add it
|
|
by hand to the join.
|
|
|
|
- To extract year, month, quarter, semester etc. do
|
|
`JOIN dim_time dt ON dt.day_key = <fact>.data_time_key` and use the dimension's
|
|
columns: `dt.year`, `dt.month`, `dt.quarter`, `dt.semester`, `dt.full_date`,
|
|
`dt.month_name_it`, `dt.year_month`.
|
|
- **Do NOT** do arithmetic on the key (e.g. `data_time_key / 10000` for the year):
|
|
it works by accident but is fragile and breaks as soon as you need to format a date
|
|
or do a cast. Always use `dim_time`.
|
|
- "Last N years from the most recent year":
|
|
`dt.year >= (SELECT MAX(year) FROM dim_time WHERE day_key IN (SELECT data_time_key FROM <fact>)) - (N-1)`,
|
|
or compute the max year on the rows actually present in the fact.
|
|
|
|
## Review checklist (on errors or suspicious results)
|
|
|
|
- Do the returned columns answer the question exactly?
|
|
- Do the filters (WHERE/HAVING) reflect ALL the conditions of the rewritten question?
|
|
- Are aggregations, groupings and orderings the required ones?
|
|
- Empty or zero result: almost always indicates a problem in conditions or joins.
|
|
Verify the filter values with `tht search find "<value>"` (match on real values).
|
|
- Do the joins follow those promoted in `schema_linking.json`? (exception: the FK
|
|
`data_time_key → dim_time.day_key` is not declared, see the time section.)
|
|
- Time analyses: are you using `JOIN dim_time` and not key arithmetic?
|
|
|
|
Every substantive revision is recorded with:
|
|
`reviewer_decide(options:[{label:"Register revision", type:"sql_revised",
|
|
subject:"sql_final", detail:"<what changed>", rationale:"<why>"}], allow_other:false)`.
|