feat(harness): language workspace param + skill riscritta in inglese (semantica completa)

Due cambiamenti interconnessi da user review:

1. language come parametro workspace (spec decisione 9):
   - Config.language (default 'en') + workspaces PSD con 'language: it'
   - Generalizza Thoth oltre l'italiano: descrizioni tabelle/colonne ed evidence
     sono nel workspace language; le istruzioni della skill restano in inglese
     (piu' affidabili per modelli piccoli, meno ambigue)

2. Skill riscritta in INGLESE preservando la semantica COMPLETA dell'originale
   (autocritica: la mia riscrittura precedente aveva perso ~10 vincoli precisi):
   - 'promuovere' ambiguo (3 accezioni: phase advance / recommend / memory promote)
     -> 'never advance a phase or record a decision without confirmation'
   - recuperati vincoli persi: choice-is-confirmation (no reviewer_confirm dopo
     reviewer_decide), reviewer_select SOLO per iterazione no-decision, messaggi
     auto-contenuti obbligatori, artefatto = superficie di decisione (gate rilegge
     da disco per CTE/SQL), candidati con provenienza+score non verita', opzione
     'leave ambiguity open', F1 passa lista completa non solo ultima
   - language contract esplicito (istruzioni EN, output nel workspace language)

Sottomoduli cte/memoria/rewriting/sql-generation in inglese, semantica tecnica
intatta (regole AV-SQL, dim_time trick, max 5 memorie solo 3 tipi riusabili).

Verifica: 0 residui nsp/chirone, tutti i tht <cmd> citati registrati, 165 passed.
This commit is contained in:
2026-06-27 14:50:30 +02:00
parent de61034a8d
commit 292048f777
9 changed files with 317 additions and 262 deletions
@@ -1,53 +1,52 @@
# Tecnica di generazione del SQL finale
# Final SQL generation technique
Adattata dallo step "query generation" di AV-SQL (divide-and-conquer ricorsivo)
e dalla sua checklist di revisione.
Adapted from the "query generation" step of AV-SQL (recursive divide-and-conquer)
and its review checklist.
## Generazione (divide-and-conquer)
## Generation (divide-and-conquer)
1. **Dividi**: scomponi la domanda riscritta in sotto-domande, ognuna mirata a
un pezzo di informazione o logica (una popolazione, un filtro, un aggregato).
2. **Conquista**: per ogni sotto-domanda formula uno pseudo-SQL, con segnaposto
per le sotto-domande non ancora risolte. I CTE gia' testati in Fase 6 sono i
mattoni preferenziali: riusali per nome, col loro esito noto.
3. **Ricombina**: sostituisci i segnaposto dal basso verso l'alto fino al SQL
completo. Il SQL finale puo' includere i CTE nel proprio WITH.
4. Dialetto: PostgreSQL. Copia ESATTAMENTE i nomi di tabelle e colonne dal
contesto schema; mai inventare oggetti.
1. **Divide**: decompose the rewritten question into sub-questions, each aimed at a
piece of information or logic (a population, a filter, an aggregate).
2. **Conquer**: for each sub-question formulate a pseudo-SQL, with placeholders for
sub-questions not yet resolved. The CTEs already tested in Phase 6 are the
preferred building blocks: reuse them by name, with their known outcome.
3. **Recombine**: replace placeholders bottom-up until the full SQL. The final SQL
may include the CTEs in its own WITH.
4. Dialect: PostgreSQL. Copy table and column names EXACTLY from the schema context;
never invent objects.
Il file `sessions/<id>/sql_final.sql` deve contenere SOLO il SQL, pulito e
copiabile: niente commenti di razionale (quello vive negli artefatti di audit).
The file `sessions/<id>/sql_final.sql` must contain ONLY the SQL, clean and
copy-pasteable: no rationale comments (that lives in the audit artifacts).
## Dimensione tempo (analisi per anno/mese/trimestre)
## Time dimension (analysis by year/month/quarter)
Le fact table hanno `data_time_key` (`integer`, formato `YYYYMMDD`): è la FK
verso `dim_time.day_key`. **Questa FK non è dichiarata** nel DWH (le fact hanno
`foreign_keys: []`), quindi non comparirà in `schema_linking.json`: vai aggiunta
a mano nel join.
Fact tables have `data_time_key` (`integer`, format `YYYYMMDD`): it is the FK to
`dim_time.day_key`. **This FK is NOT declared** in the DWH (facts have
`foreign_keys: []`), so it will NOT appear in `schema_linking.json`: you must add it
by hand to the join.
- Per estrarre anno, mese, trimestre, semestre ecc. fai
`JOIN dim_time dt ON dt.day_key = <fact>.data_time_key` e usa le colonne della
dimensione: `dt.year`, `dt.month`, `dt.quarter`, `dt.semester`, `dt.full_date`,
- To extract year, month, quarter, semester etc. do
`JOIN dim_time dt ON dt.day_key = <fact>.data_time_key` and use the dimension's
columns: `dt.year`, `dt.month`, `dt.quarter`, `dt.semester`, `dt.full_date`,
`dt.month_name_it`, `dt.year_month`.
- **Non** fare aritmetica sulla chiave (es. `data_time_key / 10000` per l'anno):
funziona per caso ma è fragile e si rompe appena serve formattare una data o
fare cast. Usa sempre `dim_time`.
- "Ultimi N anni dall'anno più recente": `dt.year >= (SELECT MAX(year) FROM
dim_time WHERE day_key IN (SELECT data_time_key FROM <fact>)) - (N-1)`, oppure
calcola il max anno sulle righe effettivamente presenti nella fact.
- **Do NOT** do arithmetic on the key (e.g. `data_time_key / 10000` for the year):
it works by accident but is fragile and breaks as soon as you need to format a date
or do a cast. Always use `dim_time`.
- "Last N years from the most recent year":
`dt.year >= (SELECT MAX(year) FROM dim_time WHERE day_key IN (SELECT data_time_key FROM <fact>)) - (N-1)`,
or compute the max year on the rows actually present in the fact.
## Checklist di revisione (su errori o risultati sospetti)
## Review checklist (on errors or suspicious results)
- Le colonne restituite rispondono esattamente alla domanda?
- I filtri (WHERE/HAVING) riflettono tutte le condizioni della domanda riscritta?
- Aggregazioni, raggruppamenti e ordinamenti sono quelli richiesti?
- Risultato vuoto o zero: quasi sempre indica un problema in condizioni o join.
Verifica i valori dei filtri con `tht search "<valore>"` (match sui valori reali).
- I join seguono quelli promossi in schema_linking.json? (eccezione: la FK
`data_time_key → dim_time.day_key` non è dichiarata, vedi sezione tempo.)
- Analisi temporali: stai usando `JOIN dim_time` e non aritmetica sulla chiave?
- Do the returned columns answer the question exactly?
- Do the filters (WHERE/HAVING) reflect ALL the conditions of the rewritten question?
- Are aggregations, groupings and orderings the required ones?
- Empty or zero result: almost always indicates a problem in conditions or joins.
Verify the filter values with `tht search "<value>"` (match on real values).
- Do the joins follow those promoted in `schema_linking.json`? (exception: the FK
`data_time_key → dim_time.day_key` is not declared, see the time section.)
- Time analyses: are you using `JOIN dim_time` and not key arithmetic?
Ogni revisione sostanziale va registrata con:
`reviewer_decide(options:[{label:"Registra revisione", type:"sql_revised",
subject:"sql_final", detail:"<cosa e' cambiato>", rationale:"<perche'>"}],
allow_other:false)`.
Every substantive revision is recorded with:
`reviewer_decide(options:[{label:"Register revision", type:"sql_revised",
subject:"sql_final", detail:"<what changed>", rationale:"<why>"}], allow_other:false)`.