feat(harness): language workspace param + skill riscritta in inglese (semantica completa)
Due cambiamenti interconnessi da user review:
1. language come parametro workspace (spec decisione 9):
- Config.language (default 'en') + workspaces PSD con 'language: it'
- Generalizza Thoth oltre l'italiano: descrizioni tabelle/colonne ed evidence
sono nel workspace language; le istruzioni della skill restano in inglese
(piu' affidabili per modelli piccoli, meno ambigue)
2. Skill riscritta in INGLESE preservando la semantica COMPLETA dell'originale
(autocritica: la mia riscrittura precedente aveva perso ~10 vincoli precisi):
- 'promuovere' ambiguo (3 accezioni: phase advance / recommend / memory promote)
-> 'never advance a phase or record a decision without confirmation'
- recuperati vincoli persi: choice-is-confirmation (no reviewer_confirm dopo
reviewer_decide), reviewer_select SOLO per iterazione no-decision, messaggi
auto-contenuti obbligatori, artefatto = superficie di decisione (gate rilegge
da disco per CTE/SQL), candidati con provenienza+score non verita', opzione
'leave ambiguity open', F1 passa lista completa non solo ultima
- language contract esplicito (istruzioni EN, output nel workspace language)
Sottomoduli cte/memoria/rewriting/sql-generation in inglese, semantica tecnica
intatta (regole AV-SQL, dim_time trick, max 5 memorie solo 3 tipi riusabili).
Verifica: 0 residui nsp/chirone, tutti i tht <cmd> citati registrati, 165 passed.
This commit is contained in:
@@ -1,53 +1,52 @@
|
||||
# Tecnica di generazione del SQL finale
|
||||
# Final SQL generation technique
|
||||
|
||||
Adattata dallo step "query generation" di AV-SQL (divide-and-conquer ricorsivo)
|
||||
e dalla sua checklist di revisione.
|
||||
Adapted from the "query generation" step of AV-SQL (recursive divide-and-conquer)
|
||||
and its review checklist.
|
||||
|
||||
## Generazione (divide-and-conquer)
|
||||
## Generation (divide-and-conquer)
|
||||
|
||||
1. **Dividi**: scomponi la domanda riscritta in sotto-domande, ognuna mirata a
|
||||
un pezzo di informazione o logica (una popolazione, un filtro, un aggregato).
|
||||
2. **Conquista**: per ogni sotto-domanda formula uno pseudo-SQL, con segnaposto
|
||||
per le sotto-domande non ancora risolte. I CTE gia' testati in Fase 6 sono i
|
||||
mattoni preferenziali: riusali per nome, col loro esito noto.
|
||||
3. **Ricombina**: sostituisci i segnaposto dal basso verso l'alto fino al SQL
|
||||
completo. Il SQL finale puo' includere i CTE nel proprio WITH.
|
||||
4. Dialetto: PostgreSQL. Copia ESATTAMENTE i nomi di tabelle e colonne dal
|
||||
contesto schema; mai inventare oggetti.
|
||||
1. **Divide**: decompose the rewritten question into sub-questions, each aimed at a
|
||||
piece of information or logic (a population, a filter, an aggregate).
|
||||
2. **Conquer**: for each sub-question formulate a pseudo-SQL, with placeholders for
|
||||
sub-questions not yet resolved. The CTEs already tested in Phase 6 are the
|
||||
preferred building blocks: reuse them by name, with their known outcome.
|
||||
3. **Recombine**: replace placeholders bottom-up until the full SQL. The final SQL
|
||||
may include the CTEs in its own WITH.
|
||||
4. Dialect: PostgreSQL. Copy table and column names EXACTLY from the schema context;
|
||||
never invent objects.
|
||||
|
||||
Il file `sessions/<id>/sql_final.sql` deve contenere SOLO il SQL, pulito e
|
||||
copiabile: niente commenti di razionale (quello vive negli artefatti di audit).
|
||||
The file `sessions/<id>/sql_final.sql` must contain ONLY the SQL, clean and
|
||||
copy-pasteable: no rationale comments (that lives in the audit artifacts).
|
||||
|
||||
## Dimensione tempo (analisi per anno/mese/trimestre)
|
||||
## Time dimension (analysis by year/month/quarter)
|
||||
|
||||
Le fact table hanno `data_time_key` (`integer`, formato `YYYYMMDD`): è la FK
|
||||
verso `dim_time.day_key`. **Questa FK non è dichiarata** nel DWH (le fact hanno
|
||||
`foreign_keys: []`), quindi non comparirà in `schema_linking.json`: vai aggiunta
|
||||
a mano nel join.
|
||||
Fact tables have `data_time_key` (`integer`, format `YYYYMMDD`): it is the FK to
|
||||
`dim_time.day_key`. **This FK is NOT declared** in the DWH (facts have
|
||||
`foreign_keys: []`), so it will NOT appear in `schema_linking.json`: you must add it
|
||||
by hand to the join.
|
||||
|
||||
- Per estrarre anno, mese, trimestre, semestre ecc. fai
|
||||
`JOIN dim_time dt ON dt.day_key = <fact>.data_time_key` e usa le colonne della
|
||||
dimensione: `dt.year`, `dt.month`, `dt.quarter`, `dt.semester`, `dt.full_date`,
|
||||
- To extract year, month, quarter, semester etc. do
|
||||
`JOIN dim_time dt ON dt.day_key = <fact>.data_time_key` and use the dimension's
|
||||
columns: `dt.year`, `dt.month`, `dt.quarter`, `dt.semester`, `dt.full_date`,
|
||||
`dt.month_name_it`, `dt.year_month`.
|
||||
- **Non** fare aritmetica sulla chiave (es. `data_time_key / 10000` per l'anno):
|
||||
funziona per caso ma è fragile e si rompe appena serve formattare una data o
|
||||
fare cast. Usa sempre `dim_time`.
|
||||
- "Ultimi N anni dall'anno più recente": `dt.year >= (SELECT MAX(year) FROM
|
||||
dim_time WHERE day_key IN (SELECT data_time_key FROM <fact>)) - (N-1)`, oppure
|
||||
calcola il max anno sulle righe effettivamente presenti nella fact.
|
||||
- **Do NOT** do arithmetic on the key (e.g. `data_time_key / 10000` for the year):
|
||||
it works by accident but is fragile and breaks as soon as you need to format a date
|
||||
or do a cast. Always use `dim_time`.
|
||||
- "Last N years from the most recent year":
|
||||
`dt.year >= (SELECT MAX(year) FROM dim_time WHERE day_key IN (SELECT data_time_key FROM <fact>)) - (N-1)`,
|
||||
or compute the max year on the rows actually present in the fact.
|
||||
|
||||
## Checklist di revisione (su errori o risultati sospetti)
|
||||
## Review checklist (on errors or suspicious results)
|
||||
|
||||
- Le colonne restituite rispondono esattamente alla domanda?
|
||||
- I filtri (WHERE/HAVING) riflettono tutte le condizioni della domanda riscritta?
|
||||
- Aggregazioni, raggruppamenti e ordinamenti sono quelli richiesti?
|
||||
- Risultato vuoto o zero: quasi sempre indica un problema in condizioni o join.
|
||||
Verifica i valori dei filtri con `tht search "<valore>"` (match sui valori reali).
|
||||
- I join seguono quelli promossi in schema_linking.json? (eccezione: la FK
|
||||
`data_time_key → dim_time.day_key` non è dichiarata, vedi sezione tempo.)
|
||||
- Analisi temporali: stai usando `JOIN dim_time` e non aritmetica sulla chiave?
|
||||
- Do the returned columns answer the question exactly?
|
||||
- Do the filters (WHERE/HAVING) reflect ALL the conditions of the rewritten question?
|
||||
- Are aggregations, groupings and orderings the required ones?
|
||||
- Empty or zero result: almost always indicates a problem in conditions or joins.
|
||||
Verify the filter values with `tht search "<value>"` (match on real values).
|
||||
- Do the joins follow those promoted in `schema_linking.json`? (exception: the FK
|
||||
`data_time_key → dim_time.day_key` is not declared, see the time section.)
|
||||
- Time analyses: are you using `JOIN dim_time` and not key arithmetic?
|
||||
|
||||
Ogni revisione sostanziale va registrata con:
|
||||
`reviewer_decide(options:[{label:"Registra revisione", type:"sql_revised",
|
||||
subject:"sql_final", detail:"<cosa e' cambiato>", rationale:"<perche'>"}],
|
||||
allow_other:false)`.
|
||||
Every substantive revision is recorded with:
|
||||
`reviewer_decide(options:[{label:"Register revision", type:"sql_revised",
|
||||
subject:"sql_final", detail:"<what changed>", rationale:"<why>"}], allow_other:false)`.
|
||||
|
||||
Reference in New Issue
Block a user