diff --git a/CONTEXT.md b/CONTEXT.md index 344fde1e..a53c99e3 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -95,28 +95,70 @@ terminale della sessione. **Evidence Module** — Il modulo autonomo che possiede la preparazione delle Evidence e la loro consultazione durante il workflow. La preparazione avviene fuori dalle singole -sessioni; il workflow usa soltanto contenuti già pubblicati. +sessioni; il workflow usa soltanto contenuti già pubblicati. A runtime contribuisce agli +stage semantici esistenti, senza diventare uno stage visibile e senza modificare ledger, +artifact o stato del workflow. **Source Evidence** — Un documento originale del workspace, conservato senza modifiche come riferimento umano e origine della successiva ristrutturazione. **Evidence Unit** — La più piccola unità semantica coerente, revisionabile e ricercabile -derivata da una sola Source Evidence. Fonti diverse non vengono fuse automaticamente. +derivata da una sola Source Evidence. Possiede un identificatore stabile indipendente +dal kind, assegnato una volta nella forma `evidence:`; fonti diverse non vengono +fuse automaticamente. **Evidence kind** — La categoria semantica di una Evidence Unit, che ne determina i -campi specifici e ne orienta l'uso. I tipi iniziali sono `glossary`, `domain`, `enum`, -`example`, `mapping`, `normalization`, `formula` e `reference`. +campi specifici e ne orienta l'uso. Ogni unità ha un solo kind primario; i tipi iniziali +sono `glossary`, `domain`, `enum`, `example`, `mapping`, `normalization`, `formula` e +`reference`. + +**Glossary Evidence** — Una Evidence Unit che definisce il significato linguistico, i +sinonimi o le varianti di un termine. + +**Domain Evidence** — Una Evidence Unit che esprime una regola o un vincolo del dominio +non rappresentato da un kind più specifico. + +**Enum Evidence** — Una Evidence Unit che collega un insieme finito di valori +memorizzati ai relativi significati. + +**Example Evidence** — Una Evidence Unit che associa un input o una domanda alla sua +interpretazione o al risultato atteso. + +**Mapping Evidence** — Una Evidence Unit che collega un concetto logico agli elementi +del relativo schema fisico. + +**Normalization Evidence** — Una Evidence Unit che descrive la trasformazione di una +rappresentazione in una forma canonica. + +**Formula Evidence** — Una Evidence Unit che contiene una singola espressione PostgreSQL +componibile e ne dichiara gli input. Una query SQL completa non è una Formula Evidence. + +**Reference Evidence** — Una Evidence Unit che rappresenta un collegamento esterno da +restituire come contenuto autonomo, anziché come semplice provenienza. **Evidence purpose** — La destinazione dichiarata di una Evidence Unit nel workflow: -disambiguation, rewriting, schema linking, SQL generation o memory. È distinta -dall'Evidence kind: il tipo descrive cosa contiene, il purpose quando può essere utile. +disambiguation, rewriting, schema linking o SQL generation. È distinta dall'Evidence +kind: il tipo descrive cosa contiene, il purpose quando può essere utile; durante la +ricerca il purpose richiesto è un filtro obbligatorio. Il recupero di esperienze e +soluzioni precedenti appartiene al Memory Module e non è un Evidence purpose. + +**Evidence Search Outcome** — Il risultato tipizzato di una consultazione del modulo +Evidence. Distingue una ricerca disponibile, che può legittimamente non trovare +corrispondenze, da un'indisponibilità tecnica che impedisce allo stage chiamante di +avanzare fino a un retry riuscito. + +**Evidence receipt** — La traccia minima di una consultazione disponibile conservata +nella sessione: stage semantico, purpose, generazione interrogata e identificatori delle +Evidence restituite. Non duplica il contenuto delle Evidence. **Curated Evidence** — Una o più Evidence Unit ristrutturate a partire da una Source -Evidence e conservate nel repository del workspace per la revisione umana. Non sono -ancora contenuto autorevole del runtime. +Evidence e conservate nel repository del workspace come proposte per la revisione +umana. Git conserva la versione precedente e rende visibile ogni modifica; una Curated +Evidence non è ancora contenuto autorevole del runtime. -**Published Evidence** — Le Curated Evidence appartenenti a una revisione Git approvata -e attivata del workspace. Sono le sole Evidence utilizzabili dalle sessioni ThothII. +**Published Evidence** — Le Curated Evidence valide appartenenti alla revisione attiva +del workspace e alla generazione Evidence pubblicata. L'approvazione umana precede +l'attivazione, ma non viene duplicata come stato nel manifest. **Evidence Index** — La proiezione ricercabile e ricostruibile delle Published Evidence. Accelera il recupero delle informazioni, ma non è una fonte di verità. @@ -125,30 +167,74 @@ Accelera il recupero delle informazioni, ma non è una fonte di verità. Curated Evidence mediante estrazione e normalizzazione deterministiche, una singola ristrutturazione assistita dal modello e una validazione finale deterministica. Nella prima versione accetta Markdown o testo UTF-8 e non acquisisce automaticamente il -contenuto di URL o documenti esterni. +contenuto di URL o documenti esterni. Prepara l'intero insieme delle modifiche in +un'area temporanea e lo applica atomicamente soltanto se tutti gli output sono validi; +non ritenta automaticamente una chiamata al modello fallita. -**Review item** — Un'ambiguità o un'informazione incompleta segnalata durante l'Evidence -preparation. Finché un Review item non viene risolto, oppure trasformato dal revisore in -una limitazione esplicita del contenuto, l'Evidence Unit non può essere indicizzata. +**Supporting excerpt** — Un breve estratto presente nel Source Evidence che sostiene +una Evidence Unit. Il sistema ne verifica deterministicamente la presenza dopo la +normalizzazione meccanica; il curatore resta responsabile di verificarne la sufficienza +semantica. + +**Evidence resolution** — L'operazione esplicita con cui un curatore ritira una +Evidence Unit oppure la ricollega a un Source Evidence esistente. Aggiorna documento e +manifest insieme, lascia un diff Git revisionabile e non pubblica né crea commit. + +**Review item** — Un blocco di revisione descritto da codice stabile, messaggio umano e +campo opzionale. Finché viene mantenuto nell'Evidence Unit, ne impedisce la +pubblicazione; la sua storia è conservata da Git, non da uno stato interno all'item. + +**Retirement candidate** — Una Curated Evidence che il Source Evidence esistente non +sostiene più. Rimane visibile con un Review item e blocca la pubblicazione finché il +curatore non la elimina oppure la rende nuovamente coerente con il sorgente. **Evidence evaluation set** — Un piccolo insieme versionato di domande rappresentative -e relativi risultati attesi, usato per verificare in modo ripetibile la qualità della -ricerca senza introdurre una piattaforma di valutazione separata. +e relativi risultati attesi. La baseline è accettabile quando ogni domanda recupera +almeno un risultato atteso nei primi dieci risultati della fusione RRF; il risultato +nei primi cinque è informativo. Comprende almeno un caso lessicale, uno semantico e uno +misto e conserva, a fini diagnostici, le posizioni dense, BM25 e fused. + +**Candidate Evidence Generation** — Una generazione completa dell'Evidence Index che +può essere valutata ma non è ancora visibile alle sessioni. Diventa attiva soltanto se +supera l'Evidence evaluation set. **Evidence manifest** — Il file versionato e gestito dal sistema che collega ogni Source Evidence al suo hash e alle Evidence Unit derivate. Conserva gli identificatori stabili, permette l'elaborazione incrementale e segnala le unità rimaste orfane senza cancellarle automaticamente. +**Orphaned Evidence Unit** — Una Curated Evidence il cui Source Evidence non esiste più. +Rimane disponibile per la revisione, ma blocca la pubblicazione finché non viene +eliminata, ricollegata oppure ne viene ripristinato il sorgente. + **Evidence Fragment** — Una proiezione ricercabile di una sezione semanticamente -coerente di una Published Evidence. Qdrant indicizza i frammenti, mentre l'Evidence -Module li raggruppa e restituisce al workflow l'Evidence Unit completa. +coerente di una Published Evidence. La divisione segue intestazioni e confini di +paragrafo; formule, coppie valore/significato, mapping, regole e URL non vengono mai +tagliati. Il testo completo reso per il frammento usa il solo limite esistente +`max_chunk_chars`, pari per default a 4.000 caratteri; un elemento atomico troppo grande +produce un Review item bloccante. Qdrant indicizza i frammenti, mentre l'Evidence Module +li raggruppa per Evidence Unit. + +**Evidence Result** — La rappresentazione di una singola Evidence Unit restituita dalla +ricerca con metadati, migliori estratti, provenienza e riferimento al documento completo. **Hybrid Evidence retrieval** — La ricerca che combina in Qdrant una graduatoria semantica dense e una graduatoria lessicale BM25 sparse mediante Reciprocal Rank Fusion. I metadati tipizzati restringono o orientano i risultati senza creare una collezione separata per ogni Evidence kind. +**Evidence query text** — La rappresentazione deterministica condivisa dalla ricerca +dense e BM25: domanda originale, concetti, tabelle e colonne in ordine fisso. I campi +vuoti sono omessi; domanda e contesto ricevono soltanto normalizzazione Unicode NFC, +conversione degli a-capo e rimozione degli spazi esterni. Gli elementi contestuali sono +poi deduplicati e ordinati senza conversione delle maiuscole, mentre punteggiatura e +spazi interni della domanda non vengono riscritti. + +**Additive BM25 upgrade** — L'estensione non distruttiva della collezione semantica di +un workspace che conserva il vettore dense predefinito e aggiunge il solo vettore +sparse `bm25`. Soltanto gli Evidence Fragment ricevono valori BM25; Schema e Memory +mantengono invariati dati e ricerca dense. + **Formula proposal** — Una formula individuata durante una sessione e conservata come artefatto della sessione. Non diventa Published Evidence finché non viene importata, revisionata e approvata nel repository del workspace. diff --git a/docs/adr/0001-evidence-publication-boundary.md b/docs/adr/0001-evidence-publication-boundary.md new file mode 100644 index 00000000..2892f5d6 --- /dev/null +++ b/docs/adr/0001-evidence-publication-boundary.md @@ -0,0 +1,7 @@ +# Use workspace activation as the Evidence publication boundary + +Curated Evidence becomes Published Evidence only when it is valid, belongs to the +active workspace revision, and belongs to the atomically published Evidence generation. +Human Git review remains required before activation, but its approval is not duplicated +as mutable state in the Evidence manifest; this keeps Git and the workspace registry as +the existing sources of truth instead of introducing a second approval mechanism. diff --git a/docs/adr/0002-kind-independent-evidence-identifiers.md b/docs/adr/0002-kind-independent-evidence-identifiers.md new file mode 100644 index 00000000..4b2ea0ca --- /dev/null +++ b/docs/adr/0002-kind-independent-evidence-identifiers.md @@ -0,0 +1,8 @@ +# Keep Evidence identifiers independent from kind + +An Evidence Unit keeps the same identifier when its kind is corrected or its Source +Evidence is unambiguously renamed; kind remains a separate validated field. This avoids +breaking citations and evaluation fixtures for a classification change, while genuine +semantic splits receive new identifiers so independent units never share an identity. +Identifiers use `evidence:`, are assigned once, persisted in the manifest and are +never recomputed automatically from mutable titles, paths or content hashes. diff --git a/docs/adr/0003-evaluate-evidence-before-activation.md b/docs/adr/0003-evaluate-evidence-before-activation.md new file mode 100644 index 00000000..73bc1741 --- /dev/null +++ b/docs/adr/0003-evaluate-evidence-before-activation.md @@ -0,0 +1,7 @@ +# Evaluate candidate Evidence before activation + +Evidence preprocessing builds a candidate vector generation and runs the workspace's +versioned evaluation set against that exact generation before atomically activating it. +The workspace revision may temporarily have no matching active Evidence during this +maintenance window, which is accepted as fail-closed degradation instead of introducing +a distributed transaction across Git, the workspace registry and Qdrant. diff --git a/docs/adr/0004-formula-evidence-is-an-expression.md b/docs/adr/0004-formula-evidence-is-an-expression.md new file mode 100644 index 00000000..22b2b946 --- /dev/null +++ b/docs/adr/0004-formula-evidence-is-an-expression.md @@ -0,0 +1,6 @@ +# Treat Formula Evidence as a composable SQL expression + +Formula Evidence contains one validated PostgreSQL expression with declared input +columns, not a complete query or statement. This makes formulas safely composable by +SQL generation; complete documented queries remain Example Evidence, and incompatible +legacy formulas require human review during migration. diff --git a/docs/adr/0005-evidence-contributes-to-semantic-stages.md b/docs/adr/0005-evidence-contributes-to-semantic-stages.md new file mode 100644 index 00000000..97420631 --- /dev/null +++ b/docs/adr/0005-evidence-contributes-to-semantic-stages.md @@ -0,0 +1,13 @@ +# Evidence contributes to semantic workflow stages + +The Evidence Module is a contributor to the existing semantic stages, not a visible +workflow stage. `clarification`, `rewriting`, `schema_linking`, `cte` and `final_sql` +query it with their corresponding purpose; `memory` remains owned by the Memory Module +and `synthesis` performs no Evidence search. Each mapped stage searches independently +and the workflow records only a minimal receipt containing stage, purpose, vector +generation and returned Evidence IDs. + +A successful search may return no matches and does not block the stage. Technical +unavailability is a distinct typed outcome that blocks the calling stage until retry, +without stale-generation or purpose fallback. This favors an explicit temporary stop +over silently treating a broken Evidence dependency as absence of domain knowledge. diff --git a/docs/adr/0006-grounded-atomic-evidence-preparation.md b/docs/adr/0006-grounded-atomic-evidence-preparation.md new file mode 100644 index 00000000..e75c12a1 --- /dev/null +++ b/docs/adr/0006-grounded-atomic-evidence-preparation.md @@ -0,0 +1,12 @@ +# Ground and atomically apply Evidence preparation + +Every proposed Evidence Unit carries one to five short supporting excerpts that can be +found in its normalized Source Evidence. The model may reuse only identifiers supplied +for previous units; deterministic application code assigns all new canonical IDs. This +makes provenance and identity mechanically checkable without pretending that automated +validation can replace human semantic review. + +Preparation validates the complete changed batch in a temporary area and applies +curated files plus the manifest atomically. A model timeout or invalid response is not +retried automatically and leaves the worktree unchanged. Explicit `evidence resolve` +actions retire or relink units while preserving an ordinary, recoverable Git diff. diff --git a/docs/adr/0007-add-bm25-without-rebuilding-qdrant.md b/docs/adr/0007-add-bm25-without-rebuilding-qdrant.md new file mode 100644 index 00000000..86fe7742 --- /dev/null +++ b/docs/adr/0007-add-bm25-without-rebuilding-qdrant.md @@ -0,0 +1,13 @@ +# Add BM25 without rebuilding the shared Qdrant collection + +The workspace keeps its existing unnamed dense vector and adds only the sparse `bm25` +vector with IDF through Qdrant's additive vector-schema operation. Only Evidence points +are repopulated with both default dense and BM25 values. Schema, Memory and solved +questions retain their current dense points and are verified before and after the +upgrade. + +This replaces the planned destructive conversion to a named `dense` vector. If BM25 is +missing, Evidence preprocessing may add it and verify the resulting schema; session +runtime remains read-only. An incompatible existing BM25 definition fails without +mutation. A candidate failure leaves the additive schema in place while Evidence stays +unavailable, avoiding data loss in the other workflow modules. diff --git a/docs/adr/0008-make-hybrid-evidence-retrieval-deterministic.md b/docs/adr/0008-make-hybrid-evidence-retrieval-deterministic.md new file mode 100644 index 00000000..803ff57f --- /dev/null +++ b/docs/adr/0008-make-hybrid-evidence-retrieval-deterministic.md @@ -0,0 +1,24 @@ +# Make hybrid Evidence retrieval deterministic and diagnosable + +Dense and BM25 retrieval receive the same deterministic query text. It preserves the +original question and appends nonempty concepts, tables and columns in a fixed order; +question and context receive Unicode NFC, newline canonicalization and outer trimming. +Context values are then deduplicated exactly and sorted, without lowercasing. Case, +punctuation and internal whitespace remain intact. This avoids accidental ranking +changes caused only by metadata ordering, preserves quoted PostgreSQL identifiers and +keeps the two retrieval branches directly comparable. + +Evidence fragmentation follows semantic headings, typed fields and paragraph +boundaries. Formulas, value/meaning pairs, mappings, rules and URLs remain atomic. An +atomic element larger than the existing `max_chunk_chars` limit creates the blocking +`atomic_content_too_large` Review item rather than being split mechanically. The limit +applies to the complete rendered text, defaults to 4,000 characters and is not +duplicated by an Evidence-specific setting. + +An L0 contract test starts the exact Qdrant image referenced by `compose.yaml` and +proves Italian server-side `qdrant/bm25` ingestion and search in a temporary collection. +There is no FastEmbed or dense fallback for Evidence when that capability is absent. + +The versioned evaluation set contains lexical, semantic and mixed queries. Its report +shows dense-only, BM25-only and fused ranks for expected Evidence. Publication remains +governed only by the simple fused top-10 rule; branch ranks and hit@5 are diagnostic. diff --git a/docs/plans/2026-08-24-evidence-restructuring-design.md b/docs/plans/2026-08-24-evidence-restructuring-design.md index 05f7f064..9c90268e 100644 --- a/docs/plans/2026-08-24-evidence-restructuring-design.md +++ b/docs/plans/2026-08-24-evidence-restructuring-design.md @@ -111,7 +111,7 @@ una parte specializzata determinata da `kind`. ```yaml schema_version: 1 -id: formula:fascia-pediatrica +id: evidence:fascia-pediatrica title: Fascia pediatrica kind: formula purposes: @@ -128,6 +128,8 @@ language: it provenance: source_file: source/10-domini-clinici/paziente.md source_sha256: sha256:0123456789abcdef... + supporting_excerpts: + - Per fascia pediatrica si intendono i pazienti con età inferiore a 18 anni. review_items: [] ``` @@ -139,6 +141,23 @@ I campi hanno ruoli diversi: - `provenance` permette di risalire al testo di origine; - `review_items` rende visibili i dubbi ancora da risolvere. +Ogni unità contiene da uno a cinque `supporting_excerpts`, ciascuno lungo al massimo +1.000 caratteri. Sono citazioni brevi che il validatore deve ritrovare nel sorgente dopo +la stessa normalizzazione meccanica. Provano la tracciabilità, non la correttezza +semantica: il revisore umano deve comunque verificare che sostengano davvero il +contenuto ristrutturato. + +Ogni `review_item` contiene soltanto: + +```yaml +code: ambiguous_source_statement +message: Il sorgente non chiarisce se l'età sia calcolata alla data di ricovero. +field: formula.sql # opzionale +``` + +Non possiede stato, autore o timestamp. Tutti i review item bloccano la pubblicazione; +il curatore corregge il documento e rimuove l'item, mentre Git conserva la storia. + ### 4.2 Tipi iniziali | `kind` | Contenuto | Esempio d'uso | @@ -155,6 +174,28 @@ I campi hanno ruoli diversi: Un URL che documenta un'altra Evidence appartiene alla sua `provenance`. Un URL che deve essere recuperato come risposta autonoma è invece una Evidence `reference`. +Ogni Evidence Unit possiede un solo `kind`, scelto in base ai campi strutturati che ne +definiscono il contenuto principale. `purposes` e `applies_to` possono invece avere più +valori. Quando parti dello stesso sorgente hanno identità e regole di validazione +indipendenti, vengono prodotte unità distinte; non si duplica un'unità soltanto perché è +utile in più fasi del workflow. + +La classificazione procede dai contenuti più strutturati a quelli più generali: +`formula`, `enum`, `mapping`, `normalization`, `glossary`, `domain`, `example` e +`reference`. `domain` è il tipo di ripiego per una regola del dominio che non soddisfa +uno schema più specifico; `reference` si applica soltanto quando il collegamento deve +essere restituito come contenuto autonomo. + +L'identificatore non incorpora il `kind`: una riclassificazione conserva l'ID, mentre +una vera divisione semantica assegna nuovi ID alle nuove unità. Un rinominamento +univocamente riconoscibile del Source Evidence tramite hash aggiorna la provenienza e +conserva gli ID esistenti. + +Un nuovo identificatore usa la forma leggibile `evidence:`, viene assegnato una +sola volta e non viene ricalcolato da titolo, percorso o hash. Le collisioni ricevono un +suffisso deterministico. Dopo la prima pubblicazione cambiare ID equivale a ritirare +l'unità esistente e crearne una nuova. + ### 4.3 Dati specifici per tipo La parte specializzata è una unione discriminata: ogni `kind` ammette e richiede campi @@ -189,7 +230,9 @@ enum: I tipi restano quindi sfruttabili sia in validazione sia in ricerca. Una formula non è un semplice testo etichettato: possiede obbligatoriamente un concetto, le colonne di -input e SQL valido come contenuto strutturato. +input e una singola espressione PostgreSQL componibile. `SELECT`, `WITH`, DDL e DML come +statement completi non sono Formula Evidence; una query completa documentata appartiene +a `example`. Le formule legacy incompatibili diventano review item durante la migrazione. ## 5. Pre-processing dei testi sorgente @@ -213,6 +256,11 @@ Il sistema: Un sorgente invariato non viene nuovamente elaborato. +Se la `pipeline_version` del manifest non è compatibile con quella installata, il +normale `prepare` termina senza scrivere. Il curatore può scegliere esplicitamente +`prepare --upgrade`, esclusivamente su un repository pulito, per rielaborare tutti i +sorgenti e revisionare il diff completo. + ### 5.2 Normalizzazione deterministica Prima del modello vengono normalizzati soltanto aspetti meccanici: @@ -235,6 +283,11 @@ Per ogni sorgente cambiato il modello riceve: - le precedenti Evidence curate derivate da quel sorgente; - gli identificatori già assegnati. +Per un'unità già esistente il modello può restituire soltanto uno degli identificatori +ricevuti. Per una nuova unità non propone l'ID: il preparatore assegna una sola volta +`evidence:` e l'eventuale suffisso deterministico. Un identificatore sconosciuto +prodotto dal modello rende la risposta non valida. + Può: - assegnare titoli; @@ -242,6 +295,7 @@ Può: - separare un sorgente in più Evidence Unit; - riordinare e riscrivere per chiarezza; - compilare campi strutturati con fatti presenti nel sorgente. +- citare da uno a cinque brevi estratti del sorgente che sostengono ciascuna unità. Non può: @@ -250,8 +304,19 @@ Non può: - risolvere silenziosamente un'ambiguità; - cancellare un'unità precedentemente revisionata. +Se il sorgente esiste ancora ma non sostiene più un'unità precedente, il modello la +restituisce come retirement candidate con il `review_item` +`source_no_longer_supports_unit`. Il curatore decide se eliminarla o riscriverla; fino a +quel momento la pubblicazione resta bloccata. + +Quando una precedente unità viene realmente divisa in più unità autonome, le nuove +unità ricevono nuovi ID e la precedente rimane una retirement candidate finché il +curatore non la ritira esplicitamente. + Pi viene usato in modalità non interattiva e senza strumenti di scrittura. È un dettaglio interno del comando, non una nuova tipologia di sessione ThothII. +Timeout, uscita non valida o JSON malformato interrompono il comando con un errore +attribuito al sorgente. La prima versione non esegue retry automatici. ### 5.4 Validazione deterministica @@ -261,6 +326,7 @@ L'output del modello non viene scritto direttamente. Viene prima controllato: - unicità e stabilità degli identificatori; - appartenenza alle enumerazioni ammesse; - esistenza e hash del sorgente; +- presenza nel sorgente normalizzato di ogni `supporting_excerpt`; - correttezza sintattica di URL, tabelle, colonne e SQL dove applicabile; - assenza di credenziali; - coerenza tra directory e `kind`; @@ -277,31 +343,60 @@ Esistono tre esiti. Un dubbio reale può essere mantenuto soltanto se il revisore lo trasforma in una limitazione esplicita del contenuto e svuota `review_items`. -## 6. Aggiornamenti incrementali e protezione delle correzioni umane +### 5.5 Applicazione atomica -La precedente versione curata è un input, non un file usa-e-getta. In questo modo il -modello può proporre una modifica minima senza ricominciare da zero. +Tutti gli output dei sorgenti cambiati vengono costruiti e validati in un'area +temporanea. Soltanto quando l'intero batch è valido, il comando sostituisce insieme i +documenti interessati e `manifest.yaml`. Un singolo errore lascia il worktree invariato +e il rapporto limitato elenca tutti i problemi rilevati. Non esiste successo parziale. + +## 6. Aggiornamenti incrementali e revisione delle correzioni umane + +La precedente versione curata è un input, non un file usa-e-getta. Il modello deve +proporre una modifica minima senza ricominciare da zero, ma questa istruzione non viene +presentata come una garanzia semantica. La garanzia è Git: la versione precedente resta +recuperabile, ogni variazione è visibile nel diff e nessuna proposta diventa Published +Evidence senza una nuova revisione umana. Il comando: - si rifiuta di operare se `evidence/curated/` o `evidence/manifest.yaml` contengono modifiche Git non salvate; - mantiene gli ID associati a contenuti che rappresentano ancora la stessa unità; +- riconosce come rinominato un sorgente nuovo che corrisponde univocamente all'hash di + un sorgente rimosso e ne aggiorna la provenienza senza cambiare gli ID; - mostra come diff le variazioni proposte; - non modifica i file derivati da sorgenti invariati; - segnala come orfana un'unità il cui sorgente è stato rimosso; -- non elimina mai automaticamente un'unità orfana. +- non elimina mai automaticamente un'unità orfana; +- blocca la pubblicazione finché ogni unità orfana non viene eliminata, ricollegata o + ricondotta a un sorgente ripristinato. Git fornisce confronto, revisione, cronologia e recupero. Non viene introdotto un database di authoring parallelo. +Orfani e retirement candidate vengono risolti senza modificare manualmente il manifest: + +```text +tht evidence resolve --retire +tht evidence resolve --source +``` + +Le due azioni sono mutuamente esclusive, richiedono un worktree pulito e aggiornano +atomicamente file curato e manifest. `--retire` rimuove l'unità dal corpus di authoring; +`--source` aggiorna provenienza e hash soltanto verso un sorgente esistente. Entrambe +lasciano un diff Git recuperabile, senza commit o pubblicazione automatica. Se l'unità +deve essere riscritta, il curatore modifica invece il documento e poi esegue +`validate`. + ## 7. Revisione e pubblicazione Il flusso di pubblicazione è: ```text prepare → revisione Git → validate → merge → attivazione workspace - → preprocess evidence → nuova generazione Qdrant attiva + → preprocess evidence candidata → evaluate candidata + → pubblicazione atomica della generazione Qdrant ``` ### 7.1 Approvazione umana @@ -321,11 +416,17 @@ La prima versione non crea automaticamente branch, commit o pull request. Una Curated Evidence diventa Published Evidence soltanto quando: -- appartiene a una revisione Git approvata e pulita; +- appartiene a una revisione Git pulita che ha superato il processo umano di revisione; - la revisione è stata attivata dal registry di ThothII; - non contiene `review_items` irrisolti; +- il manifest non contiene unità orfane; - l'intero corpus supera la validazione; -- la generazione Qdrant viene pubblicata atomicamente. +- la Candidate Evidence Generation supera il gate top-10; +- la generazione Qdrant viene quindi pubblicata atomicamente. + +Il runtime non tenta di ricostruire come sia avvenuta l'approvazione Git e il manifest +non contiene un flag `approved`. Il confine verificabile di pubblicazione è la +combinazione di revisione attiva, validazione superata e generazione Evidence attiva. Il descriptor filesystem deve indicizzare solo `curated/**/*.md`. I sorgenti e i file di supporto restano materializzati per tracciabilità, ma non entrano nell'indice. @@ -337,15 +438,27 @@ L'indicizzazione continua a usare il meccanismo già implementato dal modulo Evi 1. legge la radice materializzata della revisione Git attiva; 2. valida nuovamente tutte le Evidence; 3. costruisce Evidence Fragment secondo sezioni semantiche; -4. genera le rappresentazioni dense; -5. chiede a Qdrant di generare la rappresentazione lessicale BM25; -6. carica i punti con la nuova `vector_generation`; -7. verifica manifest, conteggi e leggibilità; -8. rende attiva la nuova generazione; -9. conserva le generazioni precedenti previste dalla policy. +4. verifica il vettore dense predefinito già usato da Schema e Memory; +5. aggiunge in modo non distruttivo il vettore sparse `bm25` se manca; +6. genera le rappresentazioni dense degli Evidence Fragment; +7. chiede a Qdrant di generare per gli stessi frammenti la rappresentazione lessicale + BM25; +8. carica i punti con la nuova `vector_generation`; +9. verifica manifest, conteggi e leggibilità; +10. rende attiva la nuova generazione; +11. conserva le generazioni precedenti previste dalla policy. Se uno dei passaggi fallisce, la generazione precedente rimane attiva. I punti caricati parzialmente vengono compensati secondo il meccanismo transazionale già esistente. +L'eventuale configurazione `bm25` già aggiunta rimane: è compatibile con i punti dense +esistenti e non richiede rollback. Se `bm25` esiste con una configurazione diversa da +`modifier: idf`, la procedura fallisce senza modificarla. + +L'upgrade non ricrea la collezione. I record `schema_table`, `schema_column`, `memory` e +`solved_question` restano invariati e continuano a usare il vettore dense predefinito. +Soltanto gli Evidence Fragment ricevono anche `bm25`. Durante la finestra fra aggiunta +del vettore e pubblicazione della prima candidata ibrida, Evidence è `unavailable`, ma +Schema e Memory continuano a funzionare. ## 9. Qdrant spiegato senza presupporre conoscenze vettoriali @@ -381,6 +494,12 @@ Qdrant 1.18.2 può generare questa rappresentazione direttamente sul server usan `qdrant/bm25`; per il corpus italiano si passa `language: italian` sia durante il caricamento sia durante la ricerca. Non serve aggiungere FastEmbed o un nuovo servizio. +Un test L0 avvia esattamente l'immagine Qdrant dichiarata da `compose.yaml`, crea una +collezione temporanea, indicizza due testi italiani mediante `qdrant/bm25`, verifica una +ricerca lessicale e infine elimina la collezione. Questo rende controllabile la capacità +locale richiesta prima di qualunque migrazione reale. Se la prova fallisce non esiste +un fallback silenzioso a un motore diverso. + Il nome “sparse” significa soltanto che, tra moltissime parole possibili, ogni testo ne usa poche. Qdrant mantiene anche l'IDF: una parola rara pesa più di una parola presente quasi ovunque. @@ -433,14 +552,32 @@ fragment_ordinal I metadati hanno due usi: -- workspace, revisione e generazione sono filtri obbligatori di sicurezza; -- tipo, scopo e ambito orientano la ricerca oppure diventano filtri quando il chiamante - formula una richiesta esplicita. +- workspace, revisione, generazione e purpose sono filtri obbligatori; +- tipo, concetti, tabelle e colonne diventano filtri soltanto quando il chiamante li + dichiara vincolanti; altrimenti contribuiscono al testo della query e alla spiegazione + del risultato. -Durante la generazione SQL, per esempio, `formula` e `mapping` ricevono priorità, ma una -regola `domain` molto pertinente può ancora apparire. Se il workflow chiede -esplicitamente soltanto formule, `evidence_kind=formula` diventa invece un filtro -vincolante. +La prima versione non aggiunge bonus automatici per `kind` o `applies_to`. Durante la +generazione SQL una Formula Evidence viene imposta soltanto quando il workflow richiede +esplicitamente `kind=formula`; negli altri casi dense, BM25 e RRF determinano l'ordine. + +Quando concetti, tabelle o colonne non sono vincoli, il modulo costruisce un solo testo +deterministico, identico per dense e BM25: + +```text +Domanda: +Concetti: +Tabelle: +Colonne: +``` + +Le righe vuote sono omesse. La domanda conserva formulazione e ordine originali; il +renderer applica soltanto Unicode NFC, converte CRLF e CR in `\n`, rimuove gli spazi +esterni e rifiuta una domanda vuota. Non cambia maiuscole, punteggiatura o spazi interni. +Ai valori contestuali applica NFC e `strip`, elimina stringhe vuote e duplicati esatti e +li ordina per valore Unicode senza `lower()` o `casefold()`: gli identificatori +PostgreSQL quotati possono essere sensibili alle maiuscole. In questo modo due richieste +equivalenti non cambiano per effetto dell'ordine occasionale dei metadati. ### 9.7 Perché non creare una collezione per tipo @@ -448,9 +585,10 @@ Una domanda spesso attraversa più tipi: una formula può dipendere da un mappin enum e da una regola di dominio. Collezioni separate richiederebbero più interrogazioni, fusione applicativa e più operazioni di manutenzione. -La soluzione usa la collezione semantica già posseduta dal workspace e aggiunge vettori -denominati `dense` e `bm25`. I payload indicizzati distinguono i tipi. È più semplice e -permette a Qdrant di eseguire ricerca ibrida e filtri nella stessa Query API. +La soluzione usa la collezione semantica già posseduta dal workspace, conserva il suo +vettore dense predefinito e aggiunge soltanto il vettore sparse denominato `bm25`. I +payload indicizzati distinguono i tipi. È più semplice, evita di ricostruire Schema e +Memory e permette a Qdrant di eseguire ricerca ibrida e filtri nella stessa Query API. ### 9.8 Perché indicizzare frammenti ma restituire unità @@ -458,9 +596,18 @@ Un documento lungo può contenere sezioni diverse. Un unico vettore ne diluirebb significato; frammenti arbitrari di lunghezza fissa spezzerebbero invece formule o regole. -La divisione segue intestazioni e campi tipizzati. Qdrant trova i frammenti, poi -l'Evidence Module li raggruppa per `evidence_id` e restituisce l'unità completa con -provenienza e citazione. +La divisione segue intestazioni, campi tipizzati e confini di paragrafo. Non divide mai +una formula, una coppia valore/significato, un mapping, una regola o un URL. Se uno di +questi elementi atomici supera da solo `max_chunk_chars`, la preparazione aggiunge il +Review item stabile `atomic_content_too_large` e blocca la pubblicazione: non usa un +taglio a dimensione fissa che ne altererebbe il significato. Il limite è quello già +presente nella configurazione degli embeddings, pari per default a 4.000 caratteri, e +si applica all'intero testo reso che sarà inviato all'embedder, incluse etichette e +metadati testuali. Non viene introdotta una seconda impostazione. Qdrant trova i +frammenti, poi l'Evidence Module li raggruppa per `evidence_id` e restituisce un solo +Evidence Result con i migliori estratti, provenienza, citazione e riferimento al +documento completo. Il contenuto completo viene risolto soltanto quando il workflow ne +ha bisogno. ### 9.9 Cosa non introduciamo nella prima versione @@ -478,6 +625,7 @@ Riferimenti tecnici ufficiali: - [Qdrant: Text Search](https://qdrant.tech/documentation/search/text-search/) - [Qdrant: server-side BM25](https://qdrant.tech/documentation/inference/inference-bm25/) - [Qdrant: Hybrid Queries e RRF](https://qdrant.tech/documentation/search/hybrid-queries/) +- [Qdrant: aggiornamento dello schema dei vettori](https://qdrant.tech/documentation/manage-data/collections/#update-vector-schema) - [Qdrant: payload indexing](https://qdrant.tech/documentation/manage-data/indexing/) - [Qdrant: multitenancy](https://qdrant.tech/documentation/manage-data/multitenancy/) @@ -490,19 +638,32 @@ search( query: str, purpose: EvidencePurpose, context: EvidenceSearchContext, -) -> list[EvidenceResult] +) -> EvidenceSearchOutcome ``` -`EvidenceSearchContext` può specificare tabelle, colonne, concetti e, solo quando -necessario, tipi obbligatori. +`EvidenceSearchContext` può specificare tabelle, colonne e concetti da aggiungere alla +query, oltre a vincoli espliciti su tipo, tabelle, colonne o concetti. + +Il modulo rende domanda e contesto una sola volta nel formato `Domanda`, `Concetti`, +`Tabelle`, `Colonne` definito sopra e passa esattamente quel testo sia all'embedder dense +sia a `qdrant/bm25`. + +`EvidenceSearchOutcome` distingue due stati: + +- `available`, con la generazione interrogata e zero o più `EvidenceResult`; +- `unavailable`, senza risultati e con un codice di errore stabile e un messaggio + limitato. + +Una lista vuota nello stato `available` significa che la ricerca ha funzionato ma non +ha trovato corrispondenze. Non equivale a un errore tecnico. Il modulo Evidence possiede interamente: - generazione della query dense; - query BM25 con lingua coerente; -- filtri su revisione e generazione; +- filtri su workspace, revisione, generazione e purpose; - RRF; -- preferenze per `kind`, `purpose` e `applies_to`; +- vincoli espliciti su `kind` e `applies_to`; - raggruppamento dei frammenti; - risoluzione di provenienza e citazioni; - controllo della revisione attiva. @@ -527,7 +688,33 @@ Questa superficie è usata dal curatore e non dalle sessioni. search → resolve citation → project into session ``` -F1, F3 e F4 passano `purpose` e contesto al modulo. Non conoscono collezioni, nomi di +Evidence è un contributore degli stage esistenti, non uno stage aggiuntivo. Non emette +decisioni, non scrive gli artifact canonici e non modifica il ledger o lo stato del +workflow. Lo stage chiamante decide come usare i candidati restituiti. + +L'integrazione usa l'identità semantica dello stage, non il display code: + +| Stage semantico | Display code attuale | Evidence purpose | +|---|---:|---| +| `clarification` | F1 | `disambiguation` | +| `rewriting` | F3 | `rewriting` | +| `schema_linking` | F4 | `schema_linking` | +| `cte` | F6 | `sql_generation` | +| `final_sql` | F7 | `sql_generation` | + +Lo stage `memory` (F2) usa il Memory Module. Lo stage `synthesis` (F5) verifica e +riassume lo schema linking già approvato e non avvia una nuova ricerca Evidence. + +Ogni stage elencato esegue una ricerca indipendente con gli input disponibili in quel +momento. In particolare `cte` usa domanda riscritta e schema approvato, mentre +`final_sql` aggiunge il piano CTE approvato. La prima versione non introduce una cache +condivisa fra stage. + +Il chiamante conserva nella sessione una Evidence receipt con stage, purpose, +generazione e ID restituiti. Il testo non viene copiato: rimane nel repository del +workspace e viene risolto attraverso la provenienza della Published Evidence. + +Gli stage passano `purpose` e contesto al modulo, ma non conoscono collezioni, nomi di vettori, generazioni o sintassi Qdrant. Le istruzioni Pi relative alla consultazione delle Evidence vengono spostate in @@ -559,48 +746,73 @@ globale nel workspace. - un file non UTF-8, troppo grande o strutturalmente invalido produce un errore chiaro; - un dubbio semantico produce un `review_item`; -- un albero Git sporco impedisce la sovrascrittura delle modifiche umane; -- un sorgente rimosso produce un'unità orfana, non una cancellazione. +- un albero Git sporco impedisce la scrittura di nuove proposte; +- un sorgente rimosso produce un'unità orfana, non una cancellazione, e blocca la + pubblicazione finché il curatore non la risolve. ### 13.2 Durante l'indicizzazione - la nuova generazione viene preparata senza toccare quella attiva; +- la generazione candidata viene interrogata esplicitamente per la valutazione senza + renderla visibile alle sessioni; +- una valutazione fallita lascia inattiva la candidata; - un caricamento o una verifica falliti non cambiano il puntatore attivo; - i dati parziali vengono rimossi quando possibile e comunque non sono leggibili dal runtime perché manca l'attivazione. +Poiché la revisione del workspace viene attivata prima di costruire la candidata, il +runtime può attraversare una finestra di manutenzione in cui la vecchia generazione non +corrisponde alla revisione. In questa finestra la ricerca Evidence è `unavailable` e +blocca lo stage chiamante. La prima versione accetta questa degradazione fail-closed +invece di introdurre una transazione distribuita fra Git, registry e Qdrant. Se il gate +fallisce, l'operatore corregge il corpus oppure ripristina esplicitamente la revisione +precedente. + ### 13.3 Durante una sessione +Una ricerca `available` senza corrispondenze produce una Evidence receipt vuota, viene +mostrata come tale e non impedisce allo stage di continuare. + Se Qdrant, il corpus attivo o la revisione attesa non sono disponibili: -- il risultato Evidence è vuoto; -- viene emesso un avviso esplicito; +- l'outcome è `unavailable`, non una lista vuota valida; +- viene restituito un codice stabile con un messaggio limitato; +- lo stage chiamante resta bloccato e può essere ritentato; - non vengono usate revisioni precedenti; -- la sessione può continuare con gli altri meccanismi e con i normali gate umani. +- non viene ripetuta la ricerca con un purpose diverso. -Questo è il comportamento fail-closed già presente e viene preservato. +Questo comportamento è fail-closed: un'assenza reale di corrispondenze non ferma il +workflow, mentre un guasto non viene mascherato come assenza di conoscenza. ## 14. Valutazione minima -`tht evidence evaluate` esegue le domande in `evaluation.yaml` contro l'indice attivo e -riporta almeno: +`tht evidence evaluate` esegue le domande in `evaluation.yaml` contro una generazione +indicata esplicitamente oppure, per il monitoraggio ordinario, contro l'indice attivo. +Riporta almeno: - quante domande hanno trovato una Evidence attesa nei primi 5 e nei primi 10 risultati; - quali tipi attesi sono mancati; - quali query non hanno prodotto risultati; +- per ogni Evidence attesa, la posizione nella graduatoria dense, BM25 e fused; - revisione Git, generazione e configurazione di ricerca usate. -La prima baseline deve essere salvata prima di regolare pesi o introdurre altri modelli. -Il comando non modifica l'indice. +Il file classifica ogni domanda come `lexical`, `semantic` o `mixed` e contiene almeno +un caso per profilo. L'evaluator esegue i due rami anche separatamente per renderli +diagnosticabili, oltre alla ricerca ibrida usata dal runtime. La prima baseline deve +essere salvata prima di regolare pesi o introdurre altri modelli. Il comando non +modifica l'indice. La valutazione supera il gate minimo soltanto quando ogni query trova +almeno una delle Evidence attese nei primi dieci risultati fused. Le posizioni dei +singoli rami e `hit@5` restano informative e non bloccano la pubblicazione. ## 15. Comandi e responsabilità | Comando | Dove opera | Scrive | | --- | --- | --- | -| `tht evidence prepare ` | clone Git di authoring | `curated/`, `manifest.yaml` | +| `tht evidence prepare [--upgrade]` | clone Git di authoring | `curated/`, `manifest.yaml` | | `tht evidence validate ` | clone Git o CI | nulla | -| `tht evidence evaluate ...` | indice attivo | solo rapporto su stdout/JSON | -| `tht ... workspace preprocess evidence` | installazione/runtime | nuova generazione corpus e Qdrant | +| `tht evidence resolve (--retire | --source )` | clone Git di authoring | unità interessata, `manifest.yaml` | +| `tht evidence evaluate ... [--generation ]` | generazione candidata o attiva | solo rapporto su stdout/JSON | +| `tht ... workspace preprocess evidence` | installazione/runtime | generazione candidata, poi attiva soltanto dopo il gate | `prepare` non crea commit. `preprocess evidence` non modifica il repository Git. @@ -615,10 +827,15 @@ enum, esempi NLQ, mapping e normalizzazione. La migrazione avviene così: 4. importare eventuali formule approvate come `kind: formula`; 5. compilare circa venti query in `evaluation.yaml`; 6. configurare il descriptor con `patterns: ["curated/**/*.md"]`; -7. validare, fare merge e attivare la revisione; -8. ricostruire in modo controllato la collezione per il nuovo contratto dense+BM25; -9. eseguire `preprocess evidence`; -10. salvare la baseline di valutazione e svolgere una verifica umana F1/F3/F4. +7. validare, fare merge e attivare la revisione in una finestra di manutenzione; +8. registrare conteggi e ID campione di Schema e Memory; +9. eseguire `preprocess evidence`, che aggiunge `bm25` senza ricreare la collezione e + costruisce la generazione candidata; +10. verificare che conteggi, ID campione e ricerche dense di Schema e Memory siano + invariati; +11. valutare la candidata e pubblicarla soltanto se supera il gate top-10; +12. salvare la baseline e svolgere una verifica umana degli stage `clarification`, + `rewriting`, `schema_linking`, `cte` e `final_sql`. Non serve mantenere v1 e v2 attivi contemporaneamente nel runtime: Git conserva la vecchia revisione e il meccanismo delle generazioni conserva il rollback dell'indice. @@ -630,19 +847,52 @@ La prima versione è completa quando: 1. un sorgente poco strutturato produce una o più unità tipizzate senza perdere la provenienza; 2. sorgenti invariati sono un no-op; -3. modifiche umane non vengono sovrascritte; -4. un `review_item` impedisce l'indicizzazione; -5. tutte le otto varianti hanno validazione specifica; -6. le formule approvate sono ricercate tramite lo stesso modulo delle altre Evidence; -7. soltanto `curated/**/*.md` entra nel corpus runtime; -8. la collezione Qdrant espone `dense` e `bm25` e gli indici payload richiesti; -9. la Query API esegue i due prefetch e la fusione RRF; -10. risultati di frammenti della stessa unità vengono raggruppati; -11. una revisione o generazione non corrispondente restituisce zero Evidence e un - avviso; -12. il set di valutazione produce un rapporto ripetibile; -13. una sessione completa continua a funzionare anche con Evidence non disponibili; -14. documentazione e comandi descrivono lo stesso contratto. +3. ogni modifica proposta a contenuti revisionati è recuperabile e visibile nel diff + Git prima della pubblicazione; +4. ogni unità cita brevi estratti verificabili del proprio sorgente; +5. gli ID nuovi sono assegnati dal codice e il modello può soltanto riutilizzare ID + precedenti esplicitamente forniti; +6. il batch di preparazione è tutto-o-niente e non esegue retry automatici; +7. un `review_item` impedisce l'indicizzazione; +8. un'unità orfana blocca la pubblicazione finché non viene risolta; +9. un'unità non più sostenuta dal proprio sorgente diventa una retirement candidate e + blocca la pubblicazione; +10. il comando `resolve` ritira o ricollega un'unità con un diff Git recuperabile; +11. tutte le otto varianti hanno validazione specifica e un solo `kind` primario; +12. una riclassificazione conserva l'ID e un rinominamento univoco del sorgente conserva + gli ID delle unità collegate; +13. ogni ID usa `evidence:`, non viene ricalcolato automaticamente e non contiene + il kind; +14. ogni Formula Evidence contiene una sola espressione PostgreSQL componibile; +15. le formule approvate sono ricercate tramite lo stesso modulo delle altre Evidence; +16. soltanto `curated/**/*.md` entra nel corpus runtime; +17. la collezione Qdrant conserva il dense predefinito e aggiunge `bm25` con IDF senza + ricostruzione distruttiva; +18. la Query API esegue i due prefetch e la fusione RRF; +19. risultati di frammenti della stessa unità diventano un solo Evidence Result con + estratti e riferimento al documento completo; +20. una ricerca disponibile può restituire zero Evidence, mentre una revisione o + generazione non corrispondente produce `unavailable` e blocca lo stage; +21. ogni ricerca applica il purpose come filtro obbligatorio e non usa bonus impliciti + per kind o ambito; +22. la generazione candidata diventa attiva soltanto quando ogni query recupera almeno + una Evidence attesa nei primi dieci risultati; il rapporto include anche `hit@5`; +23. i cinque stage mappati interrogano Evidence indipendentemente, mentre `memory` e + `synthesis` non lo invocano; +24. ogni ricerca disponibile conserva una ricevuta minima senza duplicare il testo; +25. una sessione completa riprende dopo la finestra fail-closed mediante retry o + rollback esplicito; +26. conteggi, ID campione e ricerche dense dimostrano che Schema e Memory non cambiano + durante l'upgrade; +27. una prova L0 dimostra `qdrant/bm25` sull'immagine locale effettivamente dichiarata; +28. dense e BM25 ricevono lo stesso testo di query deterministico; +29. nessun elemento atomico viene spezzato per rispettare la dimensione dei frammenti; +30. la valutazione copre casi lessicali, semantici e misti e mostra separatamente i tre + ranking; +31. il testo reso di ogni frammento rispetta il solo `max_chunk_chars` esistente; +32. la query conserva maiuscole, punteggiatura e spazi interni e non altera gli + identificatori PostgreSQL sensibili alle maiuscole; +33. documentazione e comandi descrivono lo stesso contratto. ## 18. Decisioni rinviate diff --git a/docs/plans/2026-08-24-evidence-restructuring.md b/docs/plans/2026-08-24-evidence-restructuring.md index 5a83e9cf..082e2a9a 100644 --- a/docs/plans/2026-08-24-evidence-restructuring.md +++ b/docs/plans/2026-08-24-evidence-restructuring.md @@ -1,6 +1,9 @@ # Evidence Restructuring Implementation Plan -> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task. +> **Execution:** GitHub issue #35 is the approved parent specification. Issues #36–#47 +> are the executable tracer-bullet tickets; implement one unblocked ticket at a time +> with `/implement`. This document remains the detailed technical reference and must +> not be executed as a second, parallel work queue. **Goal:** Build a Git-reviewed, typed Evidence authoring pipeline and publish its approved output to the existing revision-scoped Qdrant lifecycle with dense+BM25 hybrid retrieval. @@ -37,13 +40,16 @@ Cover: - all eight `kind` values; -- all five `purpose` values; +- all four `purpose` values; - immutable provenance; - strict unknown-field rejection; -- namespaced stable IDs; +- kind-independent stable IDs; - SHA-256 syntax; - directory/`kind` agreement; - parsing and dumping Markdown with YAML frontmatter. +- review items containing a stable `code`, human `message` and optional `field`, with no + status or history fields; +- IDs matching `evidence:` and remaining independent from `kind`. Start with: @@ -91,7 +97,7 @@ EvidenceKind = Literal[ "normalization", "formula", "reference", ] EvidencePurpose = Literal[ - "disambiguation", "rewriting", "schema_linking", "sql_generation", "memory", + "disambiguation", "rewriting", "schema_linking", "sql_generation", ] class EvidenceScope(StrictModel): @@ -102,12 +108,18 @@ class EvidenceScope(StrictModel): class EvidenceProvenance(StrictModel): source_file: str source_sha256: str + supporting_excerpts: tuple[str, ...] class FormulaPayload(StrictModel): concept: str columns: tuple[str, ...] sql: str +class ReviewItem(StrictModel): + code: str + message: str + field: str | None = None + class ReferencePayload(StrictModel): url: AnyHttpUrl label: str @@ -126,8 +138,10 @@ def dump_curated_markdown(value: CuratedEvidence) -> str: ... def load_curated_tree(root: Path) -> list[CuratedEvidence]: ... ``` -Validate formula SQL with `sqlglot` and validate `schema.table` / +Validate formula SQL with `sqlglot` as one PostgreSQL expression, rejecting complete +`SELECT`/`WITH` statements, DDL, DML and multiple statements. Validate `schema.table` / `schema.table.column` identifiers without querying the DWH. +Require one to five supporting excerpts, each nonempty and at most 1,000 characters. **Step 4: Export only the stable API** @@ -167,7 +181,10 @@ git commit -m "feat(evidence): add typed curated evidence model" Test that: - `review_items != []` is valid as a draft but blocks publication; +- `orphans != []` is preserved but blocks publication; - a missing or mismatched source hash blocks publication; +- every supporting excerpt is found after applying the same mechanical normalization + to the excerpt and its source; - two units cannot share an ID; - one unit cannot claim two source files; - `curated/formula/x.md` must contain `kind: formula`; @@ -202,7 +219,7 @@ orphans: [] ``` Test deterministic key ordering, stable round-trip, unknown fields, duplicate unit IDs, -and orphan preservation. +orphan preservation and incompatible `pipeline_version` refusal. **Step 3: Run and confirm RED** @@ -226,7 +243,7 @@ def validate_workspace_evidence(workspace_root: Path) -> ValidationReport: ... ``` `ValidationReport.publishable` is true only when there are no errors and no unresolved -review items. Warnings alone do not block publication. +review items or orphaned units. Warnings alone do not block publication. Do not use status fields such as `draft/reviewed` as an approval mechanism. The approved Git revision is the publication boundary. @@ -264,16 +281,24 @@ Add: ```python class EvidenceRestructurer(Protocol): - def restructure(self, request: RestructureRequest) -> tuple[CuratedEvidence, ...]: ... + def restructure(self, request: RestructureRequest) -> tuple[RestructureCandidate, ...]: ... class RestructureRequest(StrictModel): source_file: str source_sha256: str normalized_text: str previous_units: tuple[CuratedEvidence, ...] = () + +class RestructureCandidate(StrictModel): + existing_id: str | None = None + # The same title, kind, purposes, scope, provenance excerpts and typed payload + # needed to construct CuratedEvidence, but no model-assigned canonical ID. ``` -The response is validated through the models from Task 1 before any write. +`existing_id`, when present, must belong to `previous_units`. The preparer rejects any +unknown ID and deterministically allocates `evidence:` plus a collision suffix for +every candidate without one. The response is converted to and validated through the +models from Task 1 before any write. **Step 2: Write RED tests for the incremental rules** @@ -283,11 +308,23 @@ Test: - changed source: exactly one model call; - new source: new stable IDs; - removed source: old units become orphans and remain on disk; +- uniquely renamed source with the same hash: provenance changes and unit IDs remain; +- existing source no longer supporting a prior unit: retain it with + `source_no_longer_supports_unit` and block publication; +- kind-only reclassification: preserve the unit ID; +- semantic split: allocate new IDs for the new independent units; +- a semantic split retains the previous unit as a blocking retirement candidate until + the curator explicitly retires it; - one source may produce several units; +- every returned unit contains one to five exact supporting excerpts found in its + normalized source; - no returned unit may cite another source; - prior curated units are included in the request; - a dirty `evidence/curated/` or `evidence/manifest.yaml` fails before the model call; +- committed human edits remain recoverable in Git and every proposed change is exposed + in the working-tree diff; - all output writes are staged and atomically replaced only after complete validation. +- one invalid source leaves the complete batch and manifest unchanged; Inject the Git status runner and filesystem writer in tests; do not require a real Git repository for every unit test. @@ -323,14 +360,23 @@ state: - use only facts present in `normalized_text`; - preserve prior reviewed wording where it remains supported; +- retain unsupported prior units with a `source_no_longer_supports_unit` review item; - never merge sources; - emit `review_items` for uncertainty; +- copy short exact supporting excerpts from `normalized_text`; +- use only a supplied `existing_id` or omit it for deterministic allocation; - emit exactly the schema-versioned JSON object and no Markdown fence. +Do not retry automatically after a timeout, nonzero exit, malformed JSON or invalid +response. Return a stable error code and the affected source so the curator can rerun +the command explicitly. + **Step 5: Implement `prepare_workspace_evidence`** Return a bounded report with `changed`, `unchanged`, `created`, `orphaned`, `findings`, -and `model_calls`. Keep the model implementation behind `EvidenceRestructurer`. +and `model_calls`. Keep the model implementation behind `EvidenceRestructurer`. Build +the complete batch in a private staging directory and replace curated files plus the +manifest only after every candidate validates; never expose partial success. **Step 6: Run tests** @@ -364,13 +410,19 @@ git commit -m "feat(evidence): prepare curated evidence incrementally" Cover: ```text -tht evidence prepare [--json] +tht evidence prepare [--upgrade] [--json] tht evidence validate [--json] +tht evidence resolve --retire [--json] +tht evidence resolve --source [--json] ``` Require an existing canonical Git worktree root. Reject unknown flags, symlinks, a -workspace path outside the Git root, duplicate options and dirty curated state. Ensure -`--json` writes pristine JSON to stdout. +workspace path outside the Git root, duplicate options and dirty curated state. A +pipeline-version mismatch must fail without writes unless `--upgrade` is explicit; +`--upgrade` reprocesses every source. Ensure `--json` writes pristine JSON to stdout. +For `resolve`, require exactly one of `--retire` and `--source`, an existing Evidence ID, +an existing in-root source path for relinking and a clean worktree. Test atomic updates +to the curated file and manifest, with no commit or publication side effect. **Step 2: Run and confirm RED** @@ -393,6 +445,15 @@ def prepare_cmd(workspace_root: Path, json_output: bool = False) -> None: ... @evidence_app.command("validate") def validate_cmd(workspace_root: Path, json_output: bool = False) -> None: ... + +@evidence_app.command("resolve") +def resolve_cmd( + workspace_root: Path, + evidence_id: str, + retire: bool = False, + source: Path | None = None, + json_output: bool = False, +) -> None: ... ``` Register it in `harness/tht/cli/__init__.py`. Keep this separate from the existing @@ -444,7 +505,12 @@ Test that: - a short formula remains one fragment; - a domain document splits only at semantic section boundaries; - enum entries are not split in the middle of a value/meaning pair; -- fixed-size fallback is used only for a single oversized section; +- formulas, value/meaning pairs, mappings, rules and URLs are never split; +- the complete rendered fragment, including labels and textual metadata, uses the + existing `max_chunk_chars` setting whose default is 4,000 characters; +- an atomic element over `max_chunk_chars` produces the blocking + `atomic_content_too_large` review item instead of fixed-size chunks, with no second + size setting; - all fragments carry `evidence_id`, `evidence_kind`, `purposes`, scope, language and provenance; - fragment IDs and ordinals are deterministic; @@ -492,12 +558,14 @@ git add harness/tht/evidence/corpus harness/tht/evidence/preprocessing.py \ git commit -m "feat(evidence): build semantic fragments from typed units" ``` -### Task 6: Upgrade the Qdrant collection contract to named dense and BM25 vectors +### Task 6: Add BM25 to the existing Qdrant collection without rebuilding it **Files:** - Modify: `backend/src/workspaces/qdrant-collection.ts` +- Modify: `backend/src/workspaces/evidence/preprocessing.ts` - Modify: `backend/test/qdrant-collection.test.ts` +- Modify: `backend/test/workspaces/evidence/preprocessing.test.ts` - Modify: `harness/tht/adapters/vector/qdrant.py` - Modify: `harness/tests/test_qdrant_vector_store.py` - Modify: `docs/contracts/workspace-preprocessing-cli.md` @@ -508,18 +576,24 @@ The required vector contract is: ```json { - "vectors": { - "dense": {"size": 1024, "distance": "Cosine"} - }, + "vectors": {"size": 1024, "distance": "Cosine"}, "sparse_vectors": { "bm25": {"modifier": "idf"} } } ``` -Test that unnamed dense-only, wrong named dimension/distance, missing BM25, or wrong BM25 -modifier are incompatible. Self-heal may add missing payload indexes, but must not -silently convert an incompatible vector configuration. +Keep the current unnamed dense vector. Test three distinct states: + +- compatible: unnamed dense is correct and `bm25` has `modifier: idf`; +- Evidence-upgradeable: unnamed dense is correct and `bm25` is absent; +- incompatible: dense dimension/distance is wrong, dense is unexpectedly named, or + `bm25` exists with a different configuration. + +Only the Evidence maintenance path may turn the upgradeable state into compatible by +adding `bm25`. Session admission and searches remain read-only. Missing payload indexes +may still use the existing additive reconciliation; no path may delete or rename a +vector. Add keyword indexes only for fields used by filters: @@ -536,12 +610,19 @@ cd backend npx vitest run test/qdrant-collection.test.ts ``` -Expected: current unnamed-vector expectations fail. +Expected: the new additive BM25 expectations fail. -**Step 3: Implement strict named-vector reconciliation** +**Step 3: Implement additive BM25 reconciliation** -Update `createCollection`, `vectorCompatibility` and payload-index reconciliation. -Retain `require_existing` behavior: validation never performs an incompatible migration. +Keep base `vectorCompatibility` concerned with the unnamed dense contract used by +Schema and Memory. Add an Evidence-specific compatibility result that also classifies +`bm25` as compatible, upgradeable or incompatible. Update `createCollection` and the +Evidence preprocessing preflight: a new collection is created with unnamed dense plus +`bm25`; on an existing upgradeable collection, `workspace preprocess evidence` calls +Qdrant's additive vector-schema endpoint to create only `bm25` with IDF, then verifies +the result before upload. Retain `require_existing` behavior for ordinary runtime +validation: it never mutates. An incompatible state fails without changes. Missing +BM25 does not make Schema or Memory unavailable; only Evidence reports `unavailable`. **Step 4: Update the Python adapter's collection validation** @@ -553,7 +634,7 @@ not share implementation code. ```bash cd backend -npx vitest run test/qdrant-collection.test.ts +npx vitest run test/qdrant-collection.test.ts test/workspaces/evidence/preprocessing.test.ts npx tsc --noEmit -p . cd ../harness .venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q @@ -561,25 +642,22 @@ cd ../harness Expected: PASS. -**Step 6: Document the required guarded rebuild** +**Step 6: Document the additive maintenance behavior** -Update the CLI contract to say that the vector-shape change is incompatible and must be -applied with the existing exact-name guarded command: - -```text -tht --installation /thothii-installation.yaml workspace vector rebuild - --workspace --collection --confirm --destroy -``` - -No automatic deletion is allowed. +Update the CLI contract to state that `workspace preprocess evidence` may add the +missing `bm25` definition but may not delete or rename vectors. The existing destructive +`workspace vector rebuild` command remains available for unrelated operator recovery +and is not used by this migration. **Step 7: Commit** ```bash git add backend/src/workspaces/qdrant-collection.ts \ - backend/test/qdrant-collection.test.ts harness/tht/adapters/vector/qdrant.py \ + backend/src/workspaces/evidence/preprocessing.ts \ + backend/test/qdrant-collection.test.ts backend/test/workspaces/evidence/preprocessing.test.ts \ + harness/tht/adapters/vector/qdrant.py \ harness/tests/test_qdrant_vector_store.py docs/contracts/workspace-preprocessing-cli.md -git commit -m "feat(evidence): require dense and bm25 qdrant vectors" +git commit -m "feat(evidence): add qdrant bm25 vector in place" ``` ### Task 7: Add server-side BM25 ingestion and hybrid Query API retrieval @@ -593,6 +671,7 @@ git commit -m "feat(evidence): require dense and bm25 qdrant vectors" - Modify: `harness/tests/test_vector_port_contract.py` - Modify: `harness/tests/test_qdrant_vector_store.py` - Modify: `harness/tests/test_corpus_pipeline.py` +- Create: `harness/tests/l0/test_qdrant_bm25_inference.py` **Step 1: Write RED port tests** @@ -617,7 +696,7 @@ For Evidence upsert, assert: ```json "vector": { - "dense": [0.1, 0.2], + "": [0.1, 0.2], "bm25": { "text": "...", "model": "qdrant/bm25", @@ -631,7 +710,7 @@ For hybrid search, assert two filtered prefetches and default RRF: ```json { "prefetch": [ - {"query": [0.1, 0.2], "using": "dense", "limit": 20, "filter": {}}, + {"query": [0.1, 0.2], "limit": 20, "filter": {}}, { "query": { "text": "fascia pediatrica", @@ -652,42 +731,54 @@ For hybrid search, assert two filtered prefetches and default RRF: Both prefetch filters must include workspace, revision, active generation and record kind. Do not add hand-tuned weights. -**Step 3: Implement named dense writes for all semantic records** +**Step 3: Prove BM25 on the actual local Qdrant image** -All existing schema/memory/Evidence records use the `dense` vector name after the -collection rebuild. Only records with `sparse_text` receive `bm25`. +Add an `l0` test that reads the Qdrant image reference from the root `compose.yaml`, +starts that exact image with testcontainers, creates a uniquely named temporary +collection, adds `bm25` with IDF, indexes two Italian texts through server-side +`qdrant/bm25`, retrieves the expected text and deletes the collection. The test must +not use FastEmbed or accept a dense fallback. It fails if the Compose reference and the +tested image diverge. -**Step 4: Implement BM25 Evidence ingestion** +**Step 4: Preserve dense writes for all existing semantic records** + +Schema, Memory and solved-question records keep their current unnamed dense writes. +Evidence records use the empty default-vector name plus `bm25` in the mixed-vector +upsert shape. Only records with `sparse_text` receive `bm25`; no full reindex of Schema +or Memory is performed. + +**Step 5: Implement BM25 Evidence ingestion** Populate `sparse_text` and map workspace language `it` to Qdrant's `italian`. Reject an unsupported language before uploading the generation. Use the same options at ingest and query time. -**Step 5: Implement hybrid search with dense fallback only for non-Evidence callers** +**Step 6: Implement hybrid search with dense fallback only for non-Evidence callers** An Evidence hybrid request must fail as unavailable if the configured Qdrant version or collection contract does not support BM25. It must not silently claim to have run hybrid search. Existing non-Evidence dense requests continue to work. -**Step 6: Run focused tests** +**Step 7: Run focused tests** ```bash cd harness .venv/bin/pytest tests/test_vector_port_contract.py tests/test_qdrant_vector_store.py \ tests/test_corpus_pipeline.py tests/test_search_pack.py -q +.venv/bin/pytest tests/l0/test_qdrant_bm25_inference.py -m l0 -q .venv/bin/ruff check tht/ports/vector.py tht/adapters/vector/qdrant.py \ tht/evidence/corpus/pipeline.py ``` Expected: PASS. -**Step 7: Commit** +**Step 8: Commit** ```bash git add harness/tht/ports/vector.py harness/tht/adapters/vector/qdrant.py \ harness/tht/vectorstore/records.py harness/tht/evidence/corpus/pipeline.py \ harness/tests/test_vector_port_contract.py harness/tests/test_qdrant_vector_store.py \ - harness/tests/test_corpus_pipeline.py + harness/tests/test_corpus_pipeline.py harness/tests/l0/test_qdrant_bm25_inference.py git commit -m "feat(evidence): add qdrant bm25 hybrid retrieval" ``` @@ -697,8 +788,12 @@ git commit -m "feat(evidence): add qdrant bm25 hybrid retrieval" - Modify: `harness/tht/evidence/search.py` - Modify: `harness/tht/evidence/__init__.py` +- Modify: `harness/tht/evidence/session.py` - Modify: `harness/tht/cli/search_cmd.py` +- Modify: `harness/tht/session/filesystem_repository.py` +- Modify: `harness/tht/session/postgres_repository.py` - Modify: `harness/tests/test_evidence_facade_contract.py` +- Create: `harness/tests/test_evidence_session_receipts.py` - Modify: `harness/tests/test_search_pack.py` - Create: `harness/.pi/skills/tht-sessione/modules/evidence/runtime-search.md` - Modify: `harness/.pi/skills/tht-sessione/projection.md.tmpl` @@ -716,6 +811,9 @@ class EvidenceSearchContext(BaseModel): tables: tuple[str, ...] = () columns: tuple[str, ...] = () required_kinds: tuple[EvidenceKind, ...] = () + required_concepts: tuple[str, ...] = () + required_tables: tuple[str, ...] = () + required_columns: tuple[str, ...] = () def search_evidence( query: str, @@ -725,50 +823,105 @@ def search_evidence( searcher: ActiveEvidenceSearcher, embedder: EvidenceQueryEmbedder, top_n: int = 10, -) -> list[EvidenceResult]: ... +) -> EvidenceSearchOutcome: ... ``` -Test hard filters for workspace/revision/generation and explicit `required_kinds`. Test -that purpose/kind/scope preferences are deterministic tie-breakers when they were not -requested as hard filters. +Test hard filters for workspace, revision, generation, purpose and every explicit +`required_*` constraint. Concepts, tables and columns without the `required_` prefix +enrich the dense and BM25 query text but do not create payload filters. Do not add +implicit kind or scope bonuses in v1. + +Render that enrichment once, identically for both retrieval branches, in this fixed +order: + +```text +Domanda: +Concetti: +Tabelle: +Colonne: +``` + +Omit empty lines and do not rewrite the original question. Test that permutations and +duplicates in context produce the same rendered text and that the dense embedder and +BM25 request receive exactly that same value. + +The renderer normalizes the question to Unicode NFC, converts CRLF and CR to `\n`, +strips only leading and trailing whitespace and rejects an empty result. It preserves +case, punctuation and internal whitespace. Apply NFC and `strip` to context values, +remove empty strings and exact duplicates, then sort by Unicode value. Do not call +`lower()` or `casefold()` because quoted PostgreSQL identifiers can be case-sensitive. + +`EvidenceSearchOutcome` has an `available` state with a generation and zero or more +results, and an `unavailable` state with a stable error code, a bounded message and no +results. `memory` is not an `EvidencePurpose`: the `memory` stage remains owned by the +Memory Module. **Step 2: Test fragment grouping** Two returned fragments with the same `evidence_id` must become one `EvidenceResult`, -with the best score, ordered matching excerpts, canonical citation and no duplicate -unit. +with the best score, ordered matching excerpts, canonical citation, a reference for +resolving the full document and no duplicate unit. Do not insert the full document into +the search pack automatically. -**Step 3: Preserve fail-closed graceful degradation** +**Step 3: Distinguish an empty search from technical unavailability** -Test absent ACTIVE corpus, revision mismatch, unavailable Qdrant and malformed payload. -All return no Evidence plus a bounded warning through the existing search-pack contract; -none uses stale rows. +Test a successful search with zero matches separately from absent ACTIVE corpus, +revision mismatch, unavailable Qdrant and malformed payload. The first returns +`available` with an empty result list and may continue. Every technical case returns +`unavailable`, blocks the calling stage until retry, and never uses stale rows or a +fallback purpose. **Step 4: Implement the facade and CLI mapping** Keep Qdrant syntax inside `tht.evidence`. `search_cmd.py` translates command inputs to the facade and renders results; it must not duplicate ranking logic. -**Step 5: Extract Evidence instructions into a module fragment** +**Step 5: Integrate the contributor with semantic stages** -Move the common F1/F3/F4 Evidence rules from the projection template into +Move the shared Evidence rules from the projection template into `modules/evidence/runtime-search.md`. The fragment must state: - candidates are not truth; - pass the phase-appropriate purpose; - show provenance; - formulas are `kind=formula`, not a separate search store; -- absence of Evidence is visible but does not stop the whole session. +- an available empty result is visible and does not block the stage; +- an unavailable outcome blocks the stage and is retriable; +- Evidence never writes decisions, canonical artifacts or workflow state. + +Map semantic stages, independently of their display codes: + +```text +clarification -> disambiguation +rewriting -> rewriting +schema_linking -> schema_linking +cte -> sql_generation +final_sql -> sql_generation +``` + +Do not invoke Evidence from `memory` or `synthesis`. Run an independent search in every +mapped stage; the `final_sql` query includes the approved CTE plan. + +**Step 6: Persist the minimal Evidence receipt** + +Use the existing session artifact repositories to maintain one +`evidence_receipts.json` artifact. For every available search, append or replace the +receipt identified by semantic stage with exactly `stage`, `purpose`, +`vector_generation` and ordered `evidence_ids`. Do not copy excerpts or complete +Evidence text. The workflow integration owns the write; `tht.evidence` only constructs +and returns the typed receipt. Test filesystem and PostgreSQL repositories, resume and +retry replacement. Register the fragment in the static `FRAGMENT_ORDER` and regenerate. -**Step 6: Run tests** +**Step 7: Run tests** ```bash cd harness python -m tht.pi_skill_projection --write python -m tht.pi_skill_projection --check .venv/bin/pytest tests/test_evidence_facade_contract.py tests/test_search_pack.py \ + tests/test_evidence_session_receipts.py \ tests/test_pi_skill_projection.py -q ``` @@ -878,28 +1031,39 @@ schema_version: 1 queries: - id: pediatric-formula query: Come distinguo i pazienti pediatrici? + profile: semantic purpose: sql_generation expected: - - formula:fascia-pediatrica + - evidence:fascia-pediatrica ``` -Require unique query IDs, nonempty expected IDs and only public purpose values. +Require unique query IDs, nonempty expected IDs, one of `lexical`, `semantic` or +`mixed` for every profile, only public purpose values and at least one query of every +profile in the complete file. **Step 2: Write RED metric tests** Compute `hit_at_5`, `hit_at_10`, missing expected IDs, empty-result queries and counts by -expected `kind`. Do not add nDCG, relevance grading or an evaluation database in v1. +expected `kind`. For each expected ID, also report its nullable dense-only, BM25-only +and fused rank. The evaluator runs both branches separately for diagnosis and the same +hybrid request used by runtime. The report passes only when every query finds at least +one expected ID in its first ten fused results; branch ranks and `hit_at_5` are +informative. Do not add nDCG, relevance grading or an evaluation database in v1. **Step 3: Implement the evaluator and CLI** Expose: ```text -tht evidence evaluate -c [--json] +tht evidence evaluate -c + [--generation ] [--json] ``` -The report includes workspace revision, active vector generation and the fixed default -RRF configuration. It is read-only. +The report includes workspace revision, evaluated vector generation and the fixed +default RRF configuration. It is read-only. With no `--generation`, it evaluates the +active generation for monitoring. Before publication, the preprocessing pipeline calls +the same evaluator against its candidate generation and atomically activates it only +when the report passes. A failed candidate remains inactive. **Step 4: Run tests** @@ -957,7 +1121,7 @@ Document: - Git publication boundary; - no runtime writes; - validation before indexing; -- dense+BM25 collection contract and guarded rebuild; +- unnamed-dense plus BM25 contract and additive Evidence upgrade; - exact public operation names and JSON status additions, if any. **Step 4: Run backend and harness contract gates** @@ -1017,10 +1181,24 @@ and a SQL formula candidate. The acceptance runner must prove: - ambiguity becomes `review_items`; - no cross-source merge; - unchanged rerun is a no-op; -- human edit is preserved; +- committed human content remains recoverable and proposed changes are visible in Git; - dirty-tree refusal; -- validation blocks unresolved review; -- validated corpus indexes and searches hybrid; +- validation blocks unresolved review and orphaned units; +- pipeline-version mismatch refusal and explicit full-corpus `--upgrade`; +- validated corpus builds an inactive candidate and searches it hybrid; +- every evaluation query retrieves at least one expected ID in the first ten results + before the candidate becomes active; +- the evaluation set contains lexical, semantic and mixed cases and reports dense, + BM25 and fused ranks separately; +- dense and BM25 receive the same deterministic query text; +- no atomic formula, enum pair, mapping, rule or URL is split into fixed-size chunks; +- the complete rendered fragment stays within the existing 4,000-character + `max_chunk_chars` default and no parallel size option exists; +- the L0 probe passes against the Qdrant image referenced by `compose.yaml` without + FastEmbed or fallback; +- query normalization is limited to NFC, newline canonicalization and outer trimming, + preserving internal whitespace, punctuation and case-sensitive identifiers; +- Formula Evidence accepts a PostgreSQL expression and rejects a complete query; - Qdrant failure remains fail-closed. Use a fake restructurer for hermetic CI. The real Pi call is a separate manual check. @@ -1054,7 +1232,7 @@ Provide: - proposed PSD branch name; - exact list of 36 source files to move; - rollback command based on the pre-migration PSD commit; -- notice that Qdrant rebuild is destructive but scoped by exact collection-name guards. +- before/after Schema and Memory counts plus sample IDs used to prove non-regression. Do not infer approval from prior design acceptance. @@ -1064,7 +1242,7 @@ Use `tht evidence prepare`, review the Git diff, resolve all review items manual `tht evidence validate`, and create the approximately twenty evaluation queries. Do not auto-merge or auto-push unless separately requested. -**Step 5: Activate and rebuild with the existing guarded operator path** +**Step 5: Activate and perform the additive BM25 upgrade** After the PSD merge/pull and activation, inspect first: @@ -1073,20 +1251,24 @@ tht --installation /thothii-installation.yaml workspace vector inspect --workspace psd-clinical --json ``` -Then use the exact descriptor-owned name in the guarded rebuild command. Run -`workspace preprocess evidence`, `tht evidence evaluate`, and the manual F1/F3/F4 -walkthrough. +Before preprocessing, record counts and representative IDs for `schema_table`, +`schema_column`, `memory` and `solved_question`. Run `workspace preprocess evidence`: +it adds `bm25` when absent, builds the candidate, evaluates that exact generation and +publishes it only if the report passes. Repeat the counts, verify the representative IDs +and run dense smoke searches for Schema and Memory before completing the manual +walkthrough of `clarification`, `rewriting`, `schema_linking`, `cte` and `final_sql`. **Step 6: Record acceptance and commit ThothII documentation** `docs/testing/evidence-restructuring-manual.md` must record separate outcomes for: - authoring and Git review; -- collection rebuild; +- additive BM25 schema upgrade; +- Schema and Memory before/after non-regression evidence; - preprocessing generation publication; - hybrid retrieval evaluation; - formula retrieval; -- graceful degradation; +- distinzione fra risultato vuoto e indisponibilità bloccante; - complete session behavior. Update `PROJECT_STATE.md` only with observed results and immutable commit/run IDs. @@ -1105,17 +1287,38 @@ Before claiming completion, verify: - [ ] `CONTEXT.md` and both Evidence plan documents use the same terminology. - [ ] All eight Evidence kinds have type-specific positive and negative tests. - [ ] Pi is called once per changed source, with no tools and no saved session. -- [ ] `prepare` refuses dirty curated state and never deletes orphans. -- [ ] `validate` blocks unresolved review items. +- [ ] `prepare` refuses dirty curated state, exposes every proposal in Git and never + deletes orphans. +- [ ] `validate` blocks unresolved review items and orphaned units. +- [ ] Source renames and kind-only reclassifications preserve unit IDs; semantic splits + receive new IDs. +- [ ] IDs use `evidence:`, remain independent from kind and are not recomputed + after their initial assignment. +- [ ] Formula Evidence accepts one PostgreSQL expression and rejects full queries. +- [ ] A pipeline-version mismatch performs no writes without explicit `--upgrade`. - [ ] Runtime reads only `curated/**/*.md` from the pinned Git revision. - [ ] The TypeScript and Python Qdrant compatibility checks agree. -- [ ] Qdrant collection uses named `dense` plus `bm25` with IDF. +- [ ] Qdrant retains the unnamed dense vector and adds only `bm25` with IDF. +- [ ] Evidence preprocessing adds a missing `bm25` definition but never deletes, + renames or destructively rebuilds collection vectors. +- [ ] Schema and Memory counts, sample IDs and dense searches remain unchanged across + the additive upgrade. - [ ] Italian BM25 options are identical during ingest and query. - [ ] Hybrid search uses two prefetches and default RRF. -- [ ] Hard filters always include workspace, revision and active generation. +- [ ] Hard filters always include workspace, revision, active generation and purpose; + only explicit `required_*` context values add further filters. - [ ] Formula runtime lookup uses typed Evidence; session proposals remain non-published. -- [ ] Fragment hits are grouped into complete Evidence Units. -- [ ] Evaluation reports hit@5 and hit@10 against a versioned fixture. +- [ ] Fragment hits are grouped into one Evidence Result with excerpts and a reference; + full documents are loaded only on demand. +- [ ] Evidence is queried independently from the five mapped semantic stages and is not + invoked from `memory` or `synthesis`. +- [ ] An available empty result can continue; technical unavailability blocks the + current stage without stale-generation or purpose fallback. +- [ ] Each available stage search persists only its minimal Evidence receipt. +- [ ] Evaluation reports hit@5 and hit@10 against a versioned fixture and passes only + when every query has at least one expected result in the first ten. +- [ ] Evaluation runs against the candidate generation before atomic activation; a + failed candidate remains invisible to sessions. - [ ] Existing corpus rollback, compensation, retention and resume tests still pass. - [ ] Backend Vitest and TypeScript gates pass. - [ ] Harness pytest and Ruff gates pass.