docs(skill): Phase 4 uses reviewer_schema_linking; honor curated output columns

This commit is contained in:
2026-07-06 23:25:00 +02:00
committed by Marco Pancotti
parent ad788278dd
commit bd63e3194f
+29 -10
View File
@@ -37,7 +37,7 @@ substantive decisions.
| F1 chiarimento | — | `reviewer_confirm kind:"phase"` | | F1 chiarimento | — | `reviewer_confirm kind:"phase"` |
| F2 memoria | — | `advance:true` only if nothing recorded; else `reviewer_confirm kind:"phase"` | | F2 memoria | — | `advance:true` only if nothing recorded; else `reviewer_confirm kind:"phase"` |
| F3 riscrittura | `question.md` | `reviewer_confirm kind:"phase"` (after `rewrite_question`) | | F3 riscrittura | `question.md` | `reviewer_confirm kind:"phase"` (after `rewrite_question`) |
| F4 schema_linking | `schema_linking.json` | `reviewer_confirm kind:"phase"` (after `write_schema_linking`) | | F4 schema_linking | `schema_linking.json` | `reviewer_confirm kind:"phase"` (after `reviewer_schema_linking` + `write_schema_linking`). Promoted columns are the reviewer-approved OUTPUT columns — project exactly those in the final SELECT. |
| F5 sintesi | — | `reviewer_confirm kind:"phase"` (after `tht session check`) | | F5 sintesi | — | `reviewer_confirm kind:"phase"` (after `tht session check`) |
| F6 cte | `cte_plan.json`, `ctes/`, `cte_tests.json` | approve each CTE `kind:"cte_result"`, then `reviewer_confirm kind:"phase"` | | F6 cte | `cte_plan.json`, `ctes/`, `cte_tests.json` | approve each CTE `kind:"cte_result"`, then `reviewer_confirm kind:"phase"` |
| F7 sql_finale | `sql_final.sql` | `kind:"sql"` records `sql_approved`, then `reviewer_confirm kind:"phase"` | | F7 sql_finale | `sql_final.sql` | `kind:"sql"` records `sql_approved`, then `reviewer_confirm kind:"phase"` |
@@ -208,8 +208,20 @@ Prerequisite: Phase 3 closed.
1. `tht schema introspect` + `tht schema render --format mschema-text` for the schema 1. `tht schema introspect` + `tht schema render --format mschema-text` for the schema
context. Copy table/column names EXACTLY from it — never invent objects. context. Copy table/column names EXACTLY from it — never invent objects.
2. Propose tables/columns/joins in `reviewer_decide(advance:false)`. Register 2. Propose tables to promote/exclude with **`reviewer_schema_linking`**: pass
`table_promoted`/`table_excluded`/`column_corrected`/`join_modified`. `tables[]` as `{id, name, kind: "promote"|"exclude", rationale, suggested_columns}`.
Do NOT list every column yourself — the gate loads the full column set (with
descriptions) from the catalog and pre-selects your `suggested_columns`. The
reviewer curates the columns per promoted table. The tool records
`table_promoted`/`table_excluded` + `column_promoted`/`column_excluded` and
re-projects `schema_linking.json` deterministically via `tht session
sync-schema-linking` (you do NOT hand-write the tables/columns part with
`write_schema_linking`). The promoted columns are the reviewer-approved OUTPUT
columns: project exactly those in the final SELECT (Phase 6/7); you remain free
to reference other columns as join keys or filter predicates when the query
requires them.
Propose joins separately in `reviewer_decide(advance:false)`, registering
`join_modified`.
3. **Value grounding (D14a).** If a cited value (e.g. "ablazione") matches MULTIPLE 3. **Value grounding (D14a).** If a cited value (e.g. "ablazione") matches MULTIPLE
columns (a boolean flag + a free-text patologia field), present a `reviewer_decide` columns (a boolean flag + a free-text patologia field), present a `reviewer_decide`
with a `value_grounded` option for each candidate column (the LSH exposes all of with a `value_grounded` option for each candidate column (the LSH exposes all of
@@ -219,13 +231,16 @@ Prerequisite: Phase 3 closed.
"<concept>"` (or derive it from the evidence/context), present it, and let the "<concept>"` (or derive it from the evidence/context), present it, and let the
reviewer approve/reject (`concept_formula_approved`/`concept_formula_rejected`). reviewer approve/reject (`concept_formula_approved`/`concept_formula_rejected`).
Reflect the approved formula in `schema_linking.json` (`concept_formulas`). Reflect the approved formula in `schema_linking.json` (`concept_formulas`).
5. Persist `schema_linking.json` with the gate's `write_schema_linking` tool — it validates 5. Persist the **joins** (and any `concept_formulas`/`open_questions`) with the gate's
the object against the `SchemaLinking` model and writes the file deterministically (never `write_schema_linking` tool — it validates the object against the `SchemaLinking`
hand-write it, never edit it with the file tool; on a validation error the tool returns the model and writes the file deterministically (never hand-write it, never edit it
exact problem to fix). Shape: `{question, candidates:[{kind:"table"|"column", name, with the file tool; on a validation error the tool returns the exact problem to
evidence?, decision?:"promoted"|"excluded"|"pending"}], joins:[{from, to, source?}], fix). Shape: `{question, candidates:[...], joins:[{from, to, source?}],
excluded:[{kind, name}], open_questions:[], concept_formulas:[]}`. Then close with excluded:[...], open_questions:[], concept_formulas:[]}` — `candidates`/`excluded`
`reviewer_confirm kind:"phase"`. Do NOT run `tht session check` (that's Phase 5). are owned by `reviewer_schema_linking`/`sync-schema-linking` (step 2), so if you
call `write_schema_linking` after step 2, carry over its `candidates`/`excluded`
unchanged rather than overwriting them. Then close with `reviewer_confirm
kind:"phase"`. Do NOT run `tht session check` (that's Phase 5).
## Phase 5 — Synthesis ## Phase 5 — Synthesis
@@ -261,6 +276,10 @@ Prerequisite: Phase 6 closed.
1. Read `sql-generation.md`. Recursive divide-and-conquer: the CTEs approved in 1. Read `sql-generation.md`. Recursive divide-and-conquer: the CTEs approved in
Phase 6 are the preferred building blocks (reuse them by name). Phase 6 are the preferred building blocks (reuse them by name).
2. Compose the final SQL (PostgreSQL dialect, exact names from the schema context). 2. Compose the final SQL (PostgreSQL dialect, exact names from the schema context).
**Output columns (F4 honoring).** The columns promoted in Phase 4's
`schema_linking.json` are the reviewer-approved OUTPUT columns: project exactly
those in the final SELECT. Other schema-linked columns remain usable as join
keys or filter predicates, but do not add them to the SELECT list.
**Time dimension:** `data_time_key` is the FK to `dim_time.day_key` (NOT declared **Time dimension:** `data_time_key` is the FK to `dim_time.day_key` (NOT declared
in the DWH, must be added by hand to the join); use `JOIN dim_time` and its in the DWH, must be added by hand to the join); use `JOIN dim_time` and its
columns (`dt.year`, `dt.month`, …), NEVER arithmetic on the key. columns (`dt.year`, `dt.month`, …), NEVER arithmetic on the key.