Files
ThothII/harness/.pi/skills/tht-sessione/sql-generation.md
T
marcopan 292048f777 feat(harness): language workspace param + skill riscritta in inglese (semantica completa)
Due cambiamenti interconnessi da user review:

1. language come parametro workspace (spec decisione 9):
   - Config.language (default 'en') + workspaces PSD con 'language: it'
   - Generalizza Thoth oltre l'italiano: descrizioni tabelle/colonne ed evidence
     sono nel workspace language; le istruzioni della skill restano in inglese
     (piu' affidabili per modelli piccoli, meno ambigue)

2. Skill riscritta in INGLESE preservando la semantica COMPLETA dell'originale
   (autocritica: la mia riscrittura precedente aveva perso ~10 vincoli precisi):
   - 'promuovere' ambiguo (3 accezioni: phase advance / recommend / memory promote)
     -> 'never advance a phase or record a decision without confirmation'
   - recuperati vincoli persi: choice-is-confirmation (no reviewer_confirm dopo
     reviewer_decide), reviewer_select SOLO per iterazione no-decision, messaggi
     auto-contenuti obbligatori, artefatto = superficie di decisione (gate rilegge
     da disco per CTE/SQL), candidati con provenienza+score non verita', opzione
     'leave ambiguity open', F1 passa lista completa non solo ultima
   - language contract esplicito (istruzioni EN, output nel workspace language)

Sottomoduli cte/memoria/rewriting/sql-generation in inglese, semantica tecnica
intatta (regole AV-SQL, dim_time trick, max 5 memorie solo 3 tipi riusabili).

Verifica: 0 residui nsp/chirone, tutti i tht <cmd> citati registrati, 165 passed.
2026-06-27 14:50:30 +02:00

2.7 KiB

Final SQL generation technique

Adapted from the "query generation" step of AV-SQL (recursive divide-and-conquer) and its review checklist.

Generation (divide-and-conquer)

  1. Divide: decompose the rewritten question into sub-questions, each aimed at a piece of information or logic (a population, a filter, an aggregate).
  2. Conquer: for each sub-question formulate a pseudo-SQL, with placeholders for sub-questions not yet resolved. The CTEs already tested in Phase 6 are the preferred building blocks: reuse them by name, with their known outcome.
  3. Recombine: replace placeholders bottom-up until the full SQL. The final SQL may include the CTEs in its own WITH.
  4. Dialect: PostgreSQL. Copy table and column names EXACTLY from the schema context; never invent objects.

The file sessions/<id>/sql_final.sql must contain ONLY the SQL, clean and copy-pasteable: no rationale comments (that lives in the audit artifacts).

Time dimension (analysis by year/month/quarter)

Fact tables have data_time_key (integer, format YYYYMMDD): it is the FK to dim_time.day_key. This FK is NOT declared in the DWH (facts have foreign_keys: []), so it will NOT appear in schema_linking.json: you must add it by hand to the join.

  • To extract year, month, quarter, semester etc. do JOIN dim_time dt ON dt.day_key = <fact>.data_time_key and use the dimension's columns: dt.year, dt.month, dt.quarter, dt.semester, dt.full_date, dt.month_name_it, dt.year_month.
  • Do NOT do arithmetic on the key (e.g. data_time_key / 10000 for the year): it works by accident but is fragile and breaks as soon as you need to format a date or do a cast. Always use dim_time.
  • "Last N years from the most recent year": dt.year >= (SELECT MAX(year) FROM dim_time WHERE day_key IN (SELECT data_time_key FROM <fact>)) - (N-1), or compute the max year on the rows actually present in the fact.

Review checklist (on errors or suspicious results)

  • Do the returned columns answer the question exactly?
  • Do the filters (WHERE/HAVING) reflect ALL the conditions of the rewritten question?
  • Are aggregations, groupings and orderings the required ones?
  • Empty or zero result: almost always indicates a problem in conditions or joins. Verify the filter values with tht search "<value>" (match on real values).
  • Do the joins follow those promoted in schema_linking.json? (exception: the FK data_time_key → dim_time.day_key is not declared, see the time section.)
  • Time analyses: are you using JOIN dim_time and not key arithmetic?

Every substantive revision is recorded with: reviewer_decide(options:[{label:"Register revision", type:"sql_revised", subject:"sql_final", detail:"<what changed>", rationale:"<why>"}], allow_other:false).