diff --git a/CLAUDE.md b/CLAUDE.md index 5c2a112c..9c37d7e3 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -9,6 +9,24 @@ built, pending manual gates, workspace/secret layout, and design-doc locations. holds the stable commands + architecture mental model; PROJECT_STATE.md holds the evolving detail. Design history lives in `docs/superpowers/specs/` and `docs/superpowers/plans/`. +## Agent skills + +### Issue tracker + +Issues and specifications are tracked in GitHub Issues for `mptyl/ThothII`. +See `docs/agents/issue-tracker.md`. + +### Triage labels + +Use the standard Matt Pocock triage roles and their corresponding GitHub labels. +See `docs/agents/triage-labels.md`. + +### Domain documentation + +This repository uses a single-context domain layout: `CONTEXT.md` at the +repository root, with repository-wide ADRs stored under `docs/adr/`. +See `docs/agents/domain.md`. + ## Commands The repo has three independently-built layers. Run the **full stack** (real Pi + DWH, needs diff --git a/CONTEXT.md b/CONTEXT.md index e244e670..344fde1e 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -91,6 +91,73 @@ correzione successiva crea una nuova sessione derivata, collegata a quella prece dopo la finalizzazione. Non può modificare il ledger, gli artifact canonici o lo stato terminale della sessione. +## Evidence + +**Evidence Module** — Il modulo autonomo che possiede la preparazione delle Evidence e +la loro consultazione durante il workflow. La preparazione avviene fuori dalle singole +sessioni; il workflow usa soltanto contenuti già pubblicati. + +**Source Evidence** — Un documento originale del workspace, conservato senza modifiche +come riferimento umano e origine della successiva ristrutturazione. + +**Evidence Unit** — La più piccola unità semantica coerente, revisionabile e ricercabile +derivata da una sola Source Evidence. Fonti diverse non vengono fuse automaticamente. + +**Evidence kind** — La categoria semantica di una Evidence Unit, che ne determina i +campi specifici e ne orienta l'uso. I tipi iniziali sono `glossary`, `domain`, `enum`, +`example`, `mapping`, `normalization`, `formula` e `reference`. + +**Evidence purpose** — La destinazione dichiarata di una Evidence Unit nel workflow: +disambiguation, rewriting, schema linking, SQL generation o memory. È distinta +dall'Evidence kind: il tipo descrive cosa contiene, il purpose quando può essere utile. + +**Curated Evidence** — Una o più Evidence Unit ristrutturate a partire da una Source +Evidence e conservate nel repository del workspace per la revisione umana. Non sono +ancora contenuto autorevole del runtime. + +**Published Evidence** — Le Curated Evidence appartenenti a una revisione Git approvata +e attivata del workspace. Sono le sole Evidence utilizzabili dalle sessioni ThothII. + +**Evidence Index** — La proiezione ricercabile e ricostruibile delle Published Evidence. +Accelera il recupero delle informazioni, ma non è una fonte di verità. + +**Evidence preparation** — Il processo di authoring che trasforma Source Evidence in +Curated Evidence mediante estrazione e normalizzazione deterministiche, una singola +ristrutturazione assistita dal modello e una validazione finale deterministica. Nella +prima versione accetta Markdown o testo UTF-8 e non acquisisce automaticamente il +contenuto di URL o documenti esterni. + +**Review item** — Un'ambiguità o un'informazione incompleta segnalata durante l'Evidence +preparation. Finché un Review item non viene risolto, oppure trasformato dal revisore in +una limitazione esplicita del contenuto, l'Evidence Unit non può essere indicizzata. + +**Evidence evaluation set** — Un piccolo insieme versionato di domande rappresentative +e relativi risultati attesi, usato per verificare in modo ripetibile la qualità della +ricerca senza introdurre una piattaforma di valutazione separata. + +**Evidence manifest** — Il file versionato e gestito dal sistema che collega ogni +Source Evidence al suo hash e alle Evidence Unit derivate. Conserva gli identificatori +stabili, permette l'elaborazione incrementale e segnala le unità rimaste orfane senza +cancellarle automaticamente. + +**Evidence Fragment** — Una proiezione ricercabile di una sezione semanticamente +coerente di una Published Evidence. Qdrant indicizza i frammenti, mentre l'Evidence +Module li raggruppa e restituisce al workflow l'Evidence Unit completa. + +**Hybrid Evidence retrieval** — La ricerca che combina in Qdrant una graduatoria +semantica dense e una graduatoria lessicale BM25 sparse mediante Reciprocal Rank +Fusion. I metadati tipizzati restringono o orientano i risultati senza creare una +collezione separata per ogni Evidence kind. + +**Formula proposal** — Una formula individuata durante una sessione e conservata come +artefatto della sessione. Non diventa Published Evidence finché non viene importata, +revisionata e approvata nel repository del workspace. + +**Fail-closed Evidence retrieval** — Il comportamento per cui un indice assente, +incompatibile o non aggiornato produce nessuna Evidence e un avviso esplicito. Il +workflow può continuare, ma non usa mai silenziosamente contenuti di una revisione +precedente o di un altro workspace. + ## Catalogo dei metadati **Workspace Database** — Il database associato a un workspace, considerato nella sua diff --git a/docs/agents/domain.md b/docs/agents/domain.md new file mode 100644 index 00000000..771882e1 --- /dev/null +++ b/docs/agents/domain.md @@ -0,0 +1,76 @@ +# Domain documentation + +This repository uses a single-context domain-documentation layout. + +## Sources + +Before changing behavior or terminology, read: + +1. `CONTEXT.md` at the repository root; +2. any relevant architectural decision records under `docs/adr/`; +3. the implementation and tests for the affected module. + +`CONTEXT.md` contains the shared domain vocabulary and the system's main +concepts. Use its terminology consistently in code, documentation, issues, and +user-facing explanations. + +ADRs explain important architectural decisions and their rationale. They are +created only when a durable decision needs to be recorded; the absence of +`docs/adr/` is not an error. + +If one of these optional sources does not exist, continue without reporting an +error. + +## Layout + +```text +/ +├── CONTEXT.md +└── docs/ + └── adr/ + └── .md +``` + +Do not introduce `CONTEXT-MAP.md` unless the repository later becomes a +genuine multi-context system whose domains require separate context documents. + +## Working with domain concepts + +When implementing or reviewing work: + +- identify the domain concepts involved; +- reuse the names defined in `CONTEXT.md`; +- distinguish domain rules from infrastructure details; +- avoid creating synonyms for established terms; +- update `CONTEXT.md` when a new durable concept is introduced or an existing + definition materially changes. + +For ThothII, the Evidence module and its concepts belong to this shared domain +context even though Evidence is implemented as an autonomous workflow module. + +## Architectural decisions + +Create an ADR when a decision: + +- affects multiple parts of the system; +- establishes a durable constraint; +- selects between meaningful alternatives; +- would otherwise be difficult to reconstruct later. + +Do not create an ADR for routine implementation details. + +If current code or a proposed change conflicts with an ADR, flag the conflict +explicitly. Do not silently override the recorded decision. + +## Keeping documentation aligned + +When a change affects the domain model: + +1. update the implementation; +2. update the relevant tests; +3. update `CONTEXT.md`; +4. add or update an ADR when the decision is architectural; +5. update linked plans and GitHub issues. + +The persisted repository documentation, not the chat transcript, is the +long-term source of truth. diff --git a/docs/agents/issue-tracker.md b/docs/agents/issue-tracker.md new file mode 100644 index 00000000..6dc19b7b --- /dev/null +++ b/docs/agents/issue-tracker.md @@ -0,0 +1,165 @@ +# Issue tracker: GitHub + +Issues and specifications for this repository live in GitHub Issues under +`mptyl/ThothII`. + +Use the GitHub CLI (`gh`) for issue operations. Infer the repository from the +current Git remote when possible. + +## Conventions + +Create an issue: + +```bash +gh issue create --title "" --body-file <file> +``` + +Read an issue: + +```bash +gh issue view <number> +``` + +List issues: + +```bash +gh issue list +``` + +Add a comment: + +```bash +gh issue comment <number> --body-file <file> +``` + +Apply or remove labels: + +```bash +gh issue edit <number> --add-label "<label>" +gh issue edit <number> --remove-label "<label>" +``` + +Close an issue: + +```bash +gh issue close <number> +``` + +## Pull requests as a triage surface + +Pull requests are not used as the primary request or triage surface. + +A pull request may implement or resolve an issue, but the issue remains the +canonical location for: + +- the request; +- its scope and acceptance criteria; +- triage status; +- dependencies and sub-issues; +- implementation progress; +- the final resolution summary. + +## Publishing work + +When a workflow or skill says to publish a plan, specification, finding, or +request, create or update a GitHub issue. + +Do not leave the only authoritative copy in a chat transcript. + +Long implementation documents may also be committed to the repository. In that +case, the corresponding issue should link to the committed document and track +its execution status. + +## Fetching work + +When a workflow or skill refers to an issue number, retrieve the current issue +and its comments before acting: + +```bash +gh issue view <number> --comments +``` + +Treat the live issue state as authoritative for assignment, labels, closure, +and subsequent decisions. + +## Wayfinding operations + +A wayfinding map is represented by a parent GitHub issue and, when useful, +smaller child issues. + +### Map + +Create or update one parent issue describing: + +- the intended outcome; +- relevant context; +- known constraints; +- the proposed decomposition; +- dependencies between tasks; +- completion criteria. + +Label it according to `docs/agents/triage-labels.md`. + +### Child issues + +Create a separate issue for each independently actionable unit of work. + +Keep the parent issue readable: summarize the decomposition there and link the +child issues instead of copying every implementation detail. + +When GitHub sub-issues are available, register the relationship through the +GitHub API. Otherwise, maintain a checklist of linked child issues in the +parent issue. + +### Dependencies + +Represent blocking relationships with GitHub's native issue-dependency API +when available. + +First obtain the database ID of the blocking issue: + +```bash +gh api repos/mptyl/ThothII/issues/<blocking-number> --jq '.id' +``` + +Then register it as a blocker: + +```bash +gh api \ + --method POST \ + repos/mptyl/ThothII/issues/<blocked-number>/dependencies/blocked_by \ + -F issue_id=<blocking-issue-database-id> +``` + +If native dependencies are unavailable, record the relationship explicitly in +both issues. + +### Frontier + +The frontier is the set of open child issues that: + +- have no unresolved blockers; +- are sufficiently specified; +- can be worked on independently; +- are not already being worked on. + +Use labels and current issue relationships to identify the frontier. + +### Claim + +Before starting an issue: + +1. confirm that it is still open and unblocked; +2. assign it to the current operator when appropriate; +3. apply the label `ready-for-agent` only if it is genuinely executable; +4. add a short comment stating that work has started. + +### Resolve + +When the work is complete: + +1. verify the issue's acceptance criteria; +2. add a concise resolution comment with relevant files, tests, or decisions; +3. update the parent issue or dependent issues; +4. close the issue; +5. reconsider the frontier, because resolving a blocker may unlock more work. diff --git a/docs/agents/triage-labels.md b/docs/agents/triage-labels.md new file mode 100644 index 00000000..9510a8e4 --- /dev/null +++ b/docs/agents/triage-labels.md @@ -0,0 +1,19 @@ +# Triage labels + +These labels represent workflow roles rather than subject areas. + +| Label | Meaning | +| --- | --- | +| `needs-triage` | The request has not yet been classified or evaluated. | +| `needs-info` | More information or a human decision is required before work can proceed. | +| `ready-for-agent` | The work is sufficiently specified, unblocked, and suitable for an agent. | +| `ready-for-human` | The work requires human review, approval, or an action only a human can perform. | +| `wontfix` | The request has been deliberately declined or will not be implemented. | + +Use only the labels that describe the issue's current workflow state. + +Remove obsolete workflow labels when the state changes. For example, remove +`needs-info` when the missing information has been supplied. + +Subject-area labels may be added separately, but they must not replace these +workflow roles. diff --git a/docs/plans/2026-08-18-evidence-canonica-design.md b/docs/plans/2026-08-18-evidence-canonica-design.md index 012f89c4..4e272754 100644 --- a/docs/plans/2026-08-18-evidence-canonica-design.md +++ b/docs/plans/2026-08-18-evidence-canonica-design.md @@ -1,5 +1,11 @@ # Evidence canonica — struttura tipizzata per disambiguazione, schema linking e SQL +> **Superseded (2026-08-24).** Questo documento conserva la storia della prima +> proposta. Il disegno approvato è +> [`2026-08-24-evidence-restructuring-design.md`](2026-08-24-evidence-restructuring-design.md) +> e il relativo piano esecutivo è +> [`2026-08-24-evidence-restructuring.md`](2026-08-24-evidence-restructuring.md). + ## Contesto e decisioni prese ThothII ha già due livelli separati che non si parlano: diff --git a/docs/plans/2026-08-24-evidence-restructuring-design.md b/docs/plans/2026-08-24-evidence-restructuring-design.md new file mode 100644 index 00000000..05f7f064 --- /dev/null +++ b/docs/plans/2026-08-24-evidence-restructuring-design.md @@ -0,0 +1,659 @@ +# Ristrutturazione delle Evidence — disegno approvato + +**Stato:** approvato il 24 agosto 2026 +**Sostituisce:** `docs/plans/2026-08-18-evidence-canonica-design.md` +**Ambito:** authoring, revisione, pubblicazione, indicizzazione e uso runtime delle Evidence + +## 1. Obiettivo + +Questo disegno introduce un processo semplice e verificabile per trasformare documenti +di partenza non necessariamente ben organizzati in Evidence strutturate, revisionabili +da una persona e ricercabili in modo efficace da ThothII. + +La soluzione deve: + +1. partire dai testi oggi presenti nel repository del workspace; +2. riorganizzarli senza inventare informazioni; +3. conservare sorgenti e risultato nello stesso repository Git; +4. affidare a Git la revisione e l'approvazione umana; +5. indicizzare soltanto le versioni approvate; +6. sfruttare Qdrant senza moltiplicare collezioni e componenti; +7. inserirsi nel workflow modulare attuale, nel quale Evidence è un modulo autonomo. + +La fonte di verità rimane sempre il repository Git. Qdrant è un indice derivato che può +essere ricostruito. + +## 2. Principio guida + +Il processo è diviso in due percorsi distinti. + +- Il **percorso di authoring** prepara e revisiona le Evidence fuori dalle sessioni + domanda→SQL. +- Il **percorso runtime** è in sola lettura e consulta esclusivamente Evidence già + pubblicate. + +```mermaid +flowchart LR + S["Testi sorgente"] --> P["Pre-processing"] + P --> C["Evidence curate"] + C --> R["Revisione Git umana"] + R --> M["Merge e attivazione revisione"] + M --> I["Indicizzazione atomica"] + I --> Q["Qdrant: indice attivo"] + Q --> E["Evidence Module"] + E --> W["Workflow F1-F8"] +``` + +Una sessione può proporre una nuova formula o segnalare una lacuna, ma non modifica il +repository e non pubblica autonomamente conoscenza. + +## 3. Struttura nel repository del workspace + +Ogni workspace adotta questa struttura sotto la propria directory `evidence/`: + +```text +evidence/ +├── README.md +├── source/ +│ └── ... documenti originali ... +├── curated/ +│ ├── glossary/ +│ ├── domain/ +│ ├── enum/ +│ ├── example/ +│ ├── mapping/ +│ ├── normalization/ +│ ├── formula/ +│ └── reference/ +├── manifest.yaml +└── evaluation.yaml +``` + +### 3.1 `source/` + +Contiene i documenti originali. La prima versione accetta file Markdown, testo UTF-8 e +file `.sql.md`. Un URL può essere descritto in un documento, ma non viene scaricato né +interpretato automaticamente. + +I sorgenti vengono preservati: il pre-processing non li riscrive. + +### 3.2 `curated/` + +Contiene una Evidence Unit per file. Le sottodirectory rendono immediatamente visibile +il tipo anche a un lettore umano. Il campo `kind` nel documento resta comunque +obbligatorio: la directory aiuta la navigazione, il campo è il contratto macchina. + +### 3.3 `manifest.yaml` + +È gestito dal comando di preparazione e registra: + +- hash di ciascun sorgente; +- Evidence Unit derivate da quel sorgente; +- identificatori stabili; +- versione del processo di preparazione; +- unità orfane da controllare. + +Il manifest permette di elaborare soltanto ciò che è cambiato. Non sostituisce Git e +non contiene lo stato di approvazione. + +### 3.4 `evaluation.yaml` + +Contiene inizialmente circa venti domande rappresentative e gli identificatori delle +Evidence che ci aspettiamo di recuperare. È il controllo minimo per evitare di +considerare “migliore” una ricerca soltanto perché sembra sofisticata. + +## 4. Una struttura comune, otto tipi distinti + +La separazione tra tipi non viene eliminata. Ogni documento ha un involucro comune e +una parte specializzata determinata da `kind`. + +### 4.1 Campi comuni + +```yaml +schema_version: 1 +id: formula:fascia-pediatrica +title: Fascia pediatrica +kind: formula +purposes: + - sql_generation + - schema_linking +applies_to: + concepts: + - fascia pediatrica + tables: + - clinical.patient + columns: + - clinical.patient.birth_date +language: it +provenance: + source_file: source/10-domini-clinici/paziente.md + source_sha256: sha256:0123456789abcdef... +review_items: [] +``` + +I campi hanno ruoli diversi: + +- `kind` dice **che cosa contiene** il documento; +- `purposes` dice **in quali attività può essere utile**; +- `applies_to` dice **a quali concetti o elementi del database si riferisce**; +- `provenance` permette di risalire al testo di origine; +- `review_items` rende visibili i dubbi ancora da risolvere. + +### 4.2 Tipi iniziali + +| `kind` | Contenuto | Esempio d'uso | +| --- | --- | --- | +| `glossary` | Definizione, sinonimi e varianti linguistiche | Capire che “ricovero” e “degenza” possono indicare lo stesso concetto | +| `domain` | Regole e vincoli del dominio | Interpretare correttamente un episodio clinico | +| `enum` | Valori ammessi e loro significato | Tradurre “dimesso” nel codice memorizzato nel DWH | +| `example` | Domanda esemplificativa e interpretazione attesa | Riconoscere una formulazione già documentata | +| `mapping` | Collegamento fra concetto e schema fisico | Individuare tabella e colonne pertinenti | +| `normalization` | Regole di normalizzazione | Uniformare codici, date o varianti testuali | +| `formula` | Espressione SQL riutilizzabile e relativi input | Calcolare la fascia pediatrica dalla data di nascita | +| `reference` | Un riferimento esterno che è esso stesso contenuto recuperabile | Proporre all'utente il link a una specifica linea guida | + +Un URL che documenta un'altra Evidence appartiene alla sua `provenance`. Un URL che +deve essere recuperato come risposta autonoma è invece una Evidence `reference`. + +### 4.3 Dati specifici per tipo + +La parte specializzata è una unione discriminata: ogni `kind` ammette e richiede campi +diversi. Alcuni esempi: + +```yaml +# formula +formula: + concept: fascia pediatrica + columns: + - clinical.patient.birth_date + sql: | + CASE WHEN age < 18 THEN 'pediatrica' ELSE 'adulta' END +``` + +```yaml +# reference +reference: + url: https://example.org/linea-guida + label: Linea guida clinica + description: Criteri usati per classificare gli episodi. +``` + +```yaml +# enum +enum: + column: clinical.episode.discharge_status + values: + D: dimesso + T: trasferito +``` + +I tipi restano quindi sfruttabili sia in validazione sia in ricerca. Una formula non è +un semplice testo etichettato: possiede obbligatoriamente un concetto, le colonne di +input e SQL valido come contenuto strutturato. + +## 5. Pre-processing dei testi sorgente + +Il comando concettuale è: + +```text +tht evidence prepare <workspace-root> +``` + +Per l'utente è una sola operazione. Internamente esegue quattro passaggi. + +### 5.1 Estrazione deterministica + +Il sistema: + +- individua i file ammessi in `evidence/source/`; +- verifica dimensione, codifica UTF-8 e percorso sicuro; +- calcola l'hash del contenuto; +- confronta il risultato con `manifest.yaml`; +- carica, quando esiste, la precedente versione curata collegata al sorgente. + +Un sorgente invariato non viene nuovamente elaborato. + +### 5.2 Normalizzazione deterministica + +Prima del modello vengono normalizzati soltanto aspetti meccanici: + +- terminatori di riga e Unicode; +- spaziatura e intestazioni palesemente riconoscibili; +- elenchi, tabelle, blocchi SQL e URL; +- metadati già esplicitamente presenti; +- riferimenti a tabelle e colonne riconoscibili. + +Questa fase non interpreta il significato e non inventa strutture semantiche. + +### 5.3 Una sola ristrutturazione assistita dal modello + +Per ogni sorgente cambiato il modello riceve: + +- il testo normalizzato; +- gli otto schemi ammessi; +- le regole “non inventare” e “segnala il dubbio”; +- le precedenti Evidence curate derivate da quel sorgente; +- gli identificatori già assegnati. + +Può: + +- assegnare titoli; +- classificare il tipo; +- separare un sorgente in più Evidence Unit; +- riordinare e riscrivere per chiarezza; +- compilare campi strutturati con fatti presenti nel sorgente. + +Non può: + +- fondere automaticamente sorgenti diversi; +- aggiungere fatti non documentati; +- risolvere silenziosamente un'ambiguità; +- cancellare un'unità precedentemente revisionata. + +Pi viene usato in modalità non interattiva e senza strumenti di scrittura. È un +dettaglio interno del comando, non una nuova tipologia di sessione ThothII. + +### 5.4 Validazione deterministica + +L'output del modello non viene scritto direttamente. Viene prima controllato: + +- schema comune e schema specifico del `kind`; +- unicità e stabilità degli identificatori; +- appartenenza alle enumerazioni ammesse; +- esistenza e hash del sorgente; +- correttezza sintattica di URL, tabelle, colonne e SQL dove applicabile; +- assenza di credenziali; +- coerenza tra directory e `kind`; +- assenza di collegamenti a sorgenti diversi nella stessa unità. + +Esistono tre esiti. + +| Esito | Comportamento | +| --- | --- | +| Valido | Il documento è pronto per la revisione Git | +| Valido con dubbi | Il documento viene scritto con `review_items`; non è indicizzabile | +| Non valido | Il documento non è pubblicabile e il rapporto spiega l'errore | + +Un dubbio reale può essere mantenuto soltanto se il revisore lo trasforma in una +limitazione esplicita del contenuto e svuota `review_items`. + +## 6. Aggiornamenti incrementali e protezione delle correzioni umane + +La precedente versione curata è un input, non un file usa-e-getta. In questo modo il +modello può proporre una modifica minima senza ricominciare da zero. + +Il comando: + +- si rifiuta di operare se `evidence/curated/` o `evidence/manifest.yaml` contengono + modifiche Git non salvate; +- mantiene gli ID associati a contenuti che rappresentano ancora la stessa unità; +- mostra come diff le variazioni proposte; +- non modifica i file derivati da sorgenti invariati; +- segnala come orfana un'unità il cui sorgente è stato rimosso; +- non elimina mai automaticamente un'unità orfana. + +Git fornisce confronto, revisione, cronologia e recupero. Non viene introdotto un +database di authoring parallelo. + +## 7. Revisione e pubblicazione + +Il flusso di pubblicazione è: + +```text +prepare → revisione Git → validate → merge → attivazione workspace + → preprocess evidence → nuova generazione Qdrant attiva +``` + +### 7.1 Approvazione umana + +L'approvazione coincide con il normale processo Git del repository del workspace: + +1. il curatore esegue `prepare` in un clone di authoring; +2. legge i documenti e il diff; +3. corregge i contenuti; +4. esegue `tht evidence validate`; +5. apre o approva la pull request; +6. esegue il merge. + +La prima versione non crea automaticamente branch, commit o pull request. + +### 7.2 Quando un documento diventa Published Evidence + +Una Curated Evidence diventa Published Evidence soltanto quando: + +- appartiene a una revisione Git approvata e pulita; +- la revisione è stata attivata dal registry di ThothII; +- non contiene `review_items` irrisolti; +- l'intero corpus supera la validazione; +- la generazione Qdrant viene pubblicata atomicamente. + +Il descriptor filesystem deve indicizzare solo `curated/**/*.md`. I sorgenti e i file +di supporto restano materializzati per tracciabilità, ma non entrano nell'indice. + +## 8. Indicizzazione e generazioni + +L'indicizzazione continua a usare il meccanismo già implementato dal modulo Evidence: + +1. legge la radice materializzata della revisione Git attiva; +2. valida nuovamente tutte le Evidence; +3. costruisce Evidence Fragment secondo sezioni semantiche; +4. genera le rappresentazioni dense; +5. chiede a Qdrant di generare la rappresentazione lessicale BM25; +6. carica i punti con la nuova `vector_generation`; +7. verifica manifest, conteggi e leggibilità; +8. rende attiva la nuova generazione; +9. conserva le generazioni precedenti previste dalla policy. + +Se uno dei passaggi fallisce, la generazione precedente rimane attiva. I punti caricati +parzialmente vengono compensati secondo il meccanismo transazionale già esistente. + +## 9. Qdrant spiegato senza presupporre conoscenze vettoriali + +### 9.1 L'analogia della biblioteca + +Si può immaginare Qdrant come il catalogo di una biblioteca. + +- Le **Evidence Unit** sono i documenti completi conservati negli scaffali Git. +- Gli **Evidence Fragment** sono le schede del catalogo relative alle singole sezioni. +- I **vettori** sono rappresentazioni numeriche usate per confrontare una domanda con + quelle schede. +- Il **payload** è l'insieme delle etichette leggibili: tipo, scopo, tabelle, colonne, + revisione e documento di origine. + +Qdrant non decide se una Evidence è vera e non sostituisce il documento. Aiuta soltanto +a trovare rapidamente le schede più promettenti. + +### 9.2 Ricerca per significato: vettore dense + +La rappresentazione dense descrive il significato generale di una frase. Permette, per +esempio, di avvicinare “pazienti minorenni” a “fascia pediatrica” anche quando le parole +non coincidono. + +È utile per il linguaggio naturale, ma può essere meno precisa con codici, acronimi, +nomi di colonne e formule. + +### 9.3 Ricerca per parole e identificatori: BM25 sparse + +La rappresentazione sparse conserva il peso delle parole presenti. È adatta a termini +come `ICD-10`, `discharge_status`, `ADT`, un valore enum o un nome esatto di colonna. + +Qdrant 1.18.2 può generare questa rappresentazione direttamente sul server usando +`qdrant/bm25`; per il corpus italiano si passa `language: italian` sia durante il +caricamento sia durante la ricerca. Non serve aggiungere FastEmbed o un nuovo servizio. + +Il nome “sparse” significa soltanto che, tra moltissime parole possibili, ogni testo ne +usa poche. Qdrant mantiene anche l'IDF: una parola rara pesa più di una parola presente +quasi ovunque. + +### 9.4 Perché combinarle + +Una domanda può richiedere contemporaneamente comprensione e precisione lessicale: + +> “Qual è la formula per distinguere la fascia pediatrica usando +> `patient.birth_date`?” + +La ricerca dense riconosce il concetto; BM25 riconosce con forza “formula” e il nome +della colonna. Qdrant esegue entrambe e produce due graduatorie. + +### 9.5 Reciprocal Rank Fusion + +Reciprocal Rank Fusion, o RRF, combina le due graduatorie usando la posizione dei +risultati invece di confrontare direttamente punteggi di natura diversa. + +In termini pratici: + +- un documento alto in entrambe le liste sale; +- un documento molto forte in una sola lista può comunque emergere; +- non occorre inventare una conversione fragile fra “similarità semantica” e “punteggio + delle parole”. + +Si parte con i pesi predefiniti. Pesi diversi saranno introdotti soltanto se +`evaluation.yaml` dimostrerà un miglioramento. + +### 9.6 Il ruolo dei metadati + +Ogni punto Qdrant conserva almeno: + +```text +workspace_id +workspace_revision +vector_generation +record_kind +evidence_id +evidence_kind +purposes +concepts +tables +columns +language +source_file +source_sha256 +fragment_ordinal +``` + +I metadati hanno due usi: + +- workspace, revisione e generazione sono filtri obbligatori di sicurezza; +- tipo, scopo e ambito orientano la ricerca oppure diventano filtri quando il chiamante + formula una richiesta esplicita. + +Durante la generazione SQL, per esempio, `formula` e `mapping` ricevono priorità, ma una +regola `domain` molto pertinente può ancora apparire. Se il workflow chiede +esplicitamente soltanto formule, `evidence_kind=formula` diventa invece un filtro +vincolante. + +### 9.7 Perché non creare una collezione per tipo + +Una domanda spesso attraversa più tipi: una formula può dipendere da un mapping, da un +enum e da una regola di dominio. Collezioni separate richiederebbero più interrogazioni, +fusione applicativa e più operazioni di manutenzione. + +La soluzione usa la collezione semantica già posseduta dal workspace e aggiunge vettori +denominati `dense` e `bm25`. I payload indicizzati distinguono i tipi. È più semplice e +permette a Qdrant di eseguire ricerca ibrida e filtri nella stessa Query API. + +### 9.8 Perché indicizzare frammenti ma restituire unità + +Un documento lungo può contenere sezioni diverse. Un unico vettore ne diluirebbe il +significato; frammenti arbitrari di lunghezza fissa spezzerebbero invece formule o +regole. + +La divisione segue intestazioni e campi tipizzati. Qdrant trova i frammenti, poi +l'Evidence Module li raggruppa per `evidence_id` e restituisce l'unità completa con +provenienza e citazione. + +### 9.9 Cosa non introduciamo nella prima versione + +- una collezione per ogni tipo; +- ColBERT o multivettori late-interaction; +- un reranker basato su un altro modello; +- pesi RRF regolati a mano senza misurazioni; +- un servizio separato per BM25; +- ricerca automatica sul web. + +Queste possibilità rimangono future ottimizzazioni, non prerequisiti. + +Riferimenti tecnici ufficiali: + +- [Qdrant: Text Search](https://qdrant.tech/documentation/search/text-search/) +- [Qdrant: server-side BM25](https://qdrant.tech/documentation/inference/inference-bm25/) +- [Qdrant: Hybrid Queries e RRF](https://qdrant.tech/documentation/search/hybrid-queries/) +- [Qdrant: payload indexing](https://qdrant.tech/documentation/manage-data/indexing/) +- [Qdrant: multitenancy](https://qdrant.tech/documentation/manage-data/multitenancy/) + +## 10. Contratto di ricerca del modulo Evidence + +Il workflow non costruisce query Qdrant. Usa una sola interfaccia concettuale: + +```python +search( + query: str, + purpose: EvidencePurpose, + context: EvidenceSearchContext, +) -> list[EvidenceResult] +``` + +`EvidenceSearchContext` può specificare tabelle, colonne, concetti e, solo quando +necessario, tipi obbligatori. + +Il modulo Evidence possiede interamente: + +- generazione della query dense; +- query BM25 con lingua coerente; +- filtri su revisione e generazione; +- RRF; +- preferenze per `kind`, `purpose` e `applies_to`; +- raggruppamento dei frammenti; +- risoluzione di provenienza e citazioni; +- controllo della revisione attiva. + +Il workflow riceve candidati spiegabili, mai verità automatiche. + +## 11. Inserimento nel workflow modulare ThothII + +Evidence rimane un modulo autonomo con due responsabilità pubbliche. + +### 11.1 Authoring + +```text +prepare → validate → evaluate +``` + +Questa superficie è usata dal curatore e non dalle sessioni. + +### 11.2 Runtime + +```text +search → resolve citation → project into session +``` + +F1, F3 e F4 passano `purpose` e contesto al modulo. Non conoscono collezioni, nomi di +vettori, generazioni o sintassi Qdrant. + +Le istruzioni Pi relative alla consultazione delle Evidence vengono spostate in +frammenti del modulo Evidence e poi proiettate nel `SKILL.md` generato, seguendo il +meccanismo modulare già usato da Disambiguation e Memory. + +## 12. Formule + +Le formule approvate oggi presenti nello store `formulas/*.sql.md` vengono convertite in +Evidence `kind: formula`. Dopo la migrazione non esistono due archivi runtime. + +Una nuova formula scoperta in F4 segue invece questo percorso: + +```text +sessione → Formula proposal nell'artefatto di sessione + → importazione di manutenzione + → Curated Evidence formula + → revisione Git + → Published Evidence +``` + +Le decisioni `concept_formula_approved` e `concept_formula_rejected` continuano a +descrivere la scelta fatta nella singola sessione. Non equivalgono alla pubblicazione +globale nel workspace. + +## 13. Comportamento in caso di errore + +### 13.1 Durante l'authoring + +- un file non UTF-8, troppo grande o strutturalmente invalido produce un errore chiaro; +- un dubbio semantico produce un `review_item`; +- un albero Git sporco impedisce la sovrascrittura delle modifiche umane; +- un sorgente rimosso produce un'unità orfana, non una cancellazione. + +### 13.2 Durante l'indicizzazione + +- la nuova generazione viene preparata senza toccare quella attiva; +- un caricamento o una verifica falliti non cambiano il puntatore attivo; +- i dati parziali vengono rimossi quando possibile e comunque non sono leggibili dal + runtime perché manca l'attivazione. + +### 13.3 Durante una sessione + +Se Qdrant, il corpus attivo o la revisione attesa non sono disponibili: + +- il risultato Evidence è vuoto; +- viene emesso un avviso esplicito; +- non vengono usate revisioni precedenti; +- la sessione può continuare con gli altri meccanismi e con i normali gate umani. + +Questo è il comportamento fail-closed già presente e viene preservato. + +## 14. Valutazione minima + +`tht evidence evaluate` esegue le domande in `evaluation.yaml` contro l'indice attivo e +riporta almeno: + +- quante domande hanno trovato una Evidence attesa nei primi 5 e nei primi 10 risultati; +- quali tipi attesi sono mancati; +- quali query non hanno prodotto risultati; +- revisione Git, generazione e configurazione di ricerca usate. + +La prima baseline deve essere salvata prima di regolare pesi o introdurre altri modelli. +Il comando non modifica l'indice. + +## 15. Comandi e responsabilità + +| Comando | Dove opera | Scrive | +| --- | --- | --- | +| `tht evidence prepare <workspace-root>` | clone Git di authoring | `curated/`, `manifest.yaml` | +| `tht evidence validate <workspace-root>` | clone Git o CI | nulla | +| `tht evidence evaluate ...` | indice attivo | solo rapporto su stdout/JSON | +| `tht ... workspace preprocess evidence` | installazione/runtime | nuova generazione corpus e Qdrant | + +`prepare` non crea commit. `preprocess evidence` non modifica il repository Git. + +## 16. Migrazione iniziale del workspace PSD + +Il corpus attuale comprende 36 file Markdown organizzati in glossario, domini clinici, +enum, esempi NLQ, mapping e normalizzazione. La migrazione avviene così: + +1. spostare gli originali sotto `evidence/source/`, conservandone la gerarchia; +2. eseguire `prepare` e generare `curated/`; +3. revisionare tutte le unità e risolvere i `review_items`; +4. importare eventuali formule approvate come `kind: formula`; +5. compilare circa venti query in `evaluation.yaml`; +6. configurare il descriptor con `patterns: ["curated/**/*.md"]`; +7. validare, fare merge e attivare la revisione; +8. ricostruire in modo controllato la collezione per il nuovo contratto dense+BM25; +9. eseguire `preprocess evidence`; +10. salvare la baseline di valutazione e svolgere una verifica umana F1/F3/F4. + +Non serve mantenere v1 e v2 attivi contemporaneamente nel runtime: Git conserva la +vecchia revisione e il meccanismo delle generazioni conserva il rollback dell'indice. + +## 17. Criteri di accettazione + +La prima versione è completa quando: + +1. un sorgente poco strutturato produce una o più unità tipizzate senza perdere la + provenienza; +2. sorgenti invariati sono un no-op; +3. modifiche umane non vengono sovrascritte; +4. un `review_item` impedisce l'indicizzazione; +5. tutte le otto varianti hanno validazione specifica; +6. le formule approvate sono ricercate tramite lo stesso modulo delle altre Evidence; +7. soltanto `curated/**/*.md` entra nel corpus runtime; +8. la collezione Qdrant espone `dense` e `bm25` e gli indici payload richiesti; +9. la Query API esegue i due prefetch e la fusione RRF; +10. risultati di frammenti della stessa unità vengono raggruppati; +11. una revisione o generazione non corrispondente restituisce zero Evidence e un + avviso; +12. il set di valutazione produce un rapporto ripetibile; +13. una sessione completa continua a funzionare anche con Evidence non disponibili; +14. documentazione e comandi descrivono lo stesso contratto. + +## 18. Decisioni rinviate + +Saranno considerate soltanto dopo la baseline: + +- pesi RRF diversi da quelli predefiniti; +- reranking; +- ColBERT o multivettori; +- acquisizione automatica di PDF, Word, HTML o pagine web; +- creazione automatica di branch e pull request; +- fusione assistita di Evidence provenienti da sorgenti diversi. + +Queste esclusioni mantengono la prima implementazione comprensibile, realizzabile, +manutenibile e documentabile. diff --git a/docs/plans/2026-08-24-evidence-restructuring.md b/docs/plans/2026-08-24-evidence-restructuring.md new file mode 100644 index 00000000..5a83e9cf --- /dev/null +++ b/docs/plans/2026-08-24-evidence-restructuring.md @@ -0,0 +1,1123 @@ +# Evidence Restructuring Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task. + +**Goal:** Build a Git-reviewed, typed Evidence authoring pipeline and publish its approved output to the existing revision-scoped Qdrant lifecycle with dense+BM25 hybrid retrieval. + +**Architecture:** Keep authoring outside NL→SQL sessions: deterministic preparation wraps one read-only Pi restructuring call, writes reviewable Markdown into the workspace repository, and blocks publication on unresolved review items. Reuse the current Evidence module, corpus generations, rollback, active-revision checks, and workspace-owned semantic collection; extend them instead of introducing a parallel store. + +**Tech Stack:** Python 3.12, Pydantic 2, Typer, PyYAML, sqlglot, pytest, Pi CLI, TypeScript, Fastify workspace maintenance, Vitest, Qdrant 1.18.2 Query API, server-side `qdrant/bm25`, Git. + +--- + +## Preconditions + +- Work in this dedicated worktree. +- Read `PROJECT_STATE.md`, `CONTEXT.md`, + `docs/plans/2026-08-24-evidence-restructuring-design.md`, + `docs/contracts/workspace-evidence-v3.md`, and + `docs/contracts/workspace-preprocessing-cli.md`. +- Preserve the public facade in `harness/tht/evidence/__init__.py`. +- Do not edit `harness/.pi/skills/tht-sessione/SKILL.md` directly; regenerate it with + `python -m tht.pi_skill_projection --write`. +- Keep runtime workspace access read-only. Only the authoring CLI may write + `evidence/curated/` and `evidence/manifest.yaml` in a curator clone. +- Do not migrate the external PSD repository until all code and contract gates pass. + +### Task 1: Add the typed Curated Evidence model + +**Files:** + +- Create: `harness/tht/evidence/canonical.py` +- Modify: `harness/tht/evidence/__init__.py` +- Create: `harness/tests/test_evidence_canonical.py` + +**Step 1: Write failing tests for the common envelope** + +Cover: + +- all eight `kind` values; +- all five `purpose` values; +- immutable provenance; +- strict unknown-field rejection; +- namespaced stable IDs; +- SHA-256 syntax; +- directory/`kind` agreement; +- parsing and dumping Markdown with YAML frontmatter. + +Start with: + +```python +def test_formula_requires_formula_payload(): + with pytest.raises(ValidationError): + CuratedEvidence.model_validate({ + **COMMON, + "kind": "formula", + "payload": {"concept": "fascia pediatrica"}, + }) + + +def test_reference_rejects_formula_payload(): + with pytest.raises(ValidationError): + CuratedEvidence.model_validate({ + **COMMON, + "kind": "reference", + "payload": { + "concept": "x", + "columns": ["clinical.patient.birth_date"], + "sql": "CASE WHEN true THEN 1 END", + }, + }) +``` + +**Step 2: Run the focused tests and confirm RED** + +Run: + +```bash +cd harness +.venv/bin/pytest tests/test_evidence_canonical.py -q +``` + +Expected: import failure for `tht.evidence.canonical`. + +**Step 3: Implement the discriminated model** + +Use a strict Pydantic model with these public types: + +```python +EvidenceKind = Literal[ + "glossary", "domain", "enum", "example", "mapping", + "normalization", "formula", "reference", +] +EvidencePurpose = Literal[ + "disambiguation", "rewriting", "schema_linking", "sql_generation", "memory", +] + +class EvidenceScope(StrictModel): + concepts: tuple[str, ...] = () + tables: tuple[str, ...] = () + columns: tuple[str, ...] = () + +class EvidenceProvenance(StrictModel): + source_file: str + source_sha256: str + +class FormulaPayload(StrictModel): + concept: str + columns: tuple[str, ...] + sql: str + +class ReferencePayload(StrictModel): + url: AnyHttpUrl + label: str + description: str +``` + +Define equally strict payloads for the other six kinds and expose a single +`CuratedEvidence` API. The implementation may use an internal Pydantic discriminated +union, but callers must not switch between eight unrelated loaders. + +Add: + +```python +def parse_curated_markdown(text: str, *, path: Path | None = None) -> CuratedEvidence: ... +def dump_curated_markdown(value: CuratedEvidence) -> str: ... +def load_curated_tree(root: Path) -> list[CuratedEvidence]: ... +``` + +Validate formula SQL with `sqlglot` and validate `schema.table` / +`schema.table.column` identifiers without querying the DWH. + +**Step 4: Export only the stable API** + +Add the model and loader functions to `harness/tht/evidence/__init__.py`. Do not export +internal union member helpers unless another module needs them. + +**Step 5: Run tests and lint** + +Run: + +```bash +cd harness +.venv/bin/pytest tests/test_evidence_canonical.py tests/test_evidence_facade_contract.py -q +.venv/bin/ruff check tht/evidence/canonical.py tests/test_evidence_canonical.py +``` + +Expected: PASS. + +**Step 6: Commit** + +```bash +git add harness/tht/evidence/canonical.py harness/tht/evidence/__init__.py \ + harness/tests/test_evidence_canonical.py +git commit -m "feat(evidence): add typed curated evidence model" +``` + +### Task 2: Add deterministic validation and the Evidence manifest + +**Files:** + +- Create: `harness/tht/evidence/authoring.py` +- Create: `harness/tests/test_evidence_authoring.py` +- Modify: `harness/tht/evidence/__init__.py` + +**Step 1: Write failing tests for publication validation** + +Test that: + +- `review_items != []` is valid as a draft but blocks publication; +- a missing or mismatched source hash blocks publication; +- two units cannot share an ID; +- one unit cannot claim two source files; +- `curated/formula/x.md` must contain `kind: formula`; +- credentials in URLs, YAML, or body are rejected; +- only Markdown, `.txt`, and `.sql.md` sources are accepted; +- UTF-8 and per-file limits are enforced. + +Expose errors as bounded structured values: + +```python +@dataclass(frozen=True) +class ValidationFinding: + severity: Literal["error", "warning"] + code: str + path: str + message: str +``` + +**Step 2: Write failing tests for the versioned manifest** + +Use this minimum shape: + +```yaml +schema_version: 1 +pipeline_version: evidence-authoring-v1 +sources: + source/domain/patient.md: + sha256: sha256:... + units: + - domain:patient +orphans: [] +``` + +Test deterministic key ordering, stable round-trip, unknown fields, duplicate unit IDs, +and orphan preservation. + +**Step 3: Run and confirm RED** + +Run: + +```bash +cd harness +.venv/bin/pytest tests/test_evidence_authoring.py -q +``` + +Expected: missing authoring API. + +**Step 4: Implement `EvidenceManifest` and `validate_workspace_evidence`** + +Provide: + +```python +def load_manifest(path: Path) -> EvidenceManifest: ... +def dump_manifest(manifest: EvidenceManifest) -> str: ... +def validate_workspace_evidence(workspace_root: Path) -> ValidationReport: ... +``` + +`ValidationReport.publishable` is true only when there are no errors and no unresolved +review items. Warnings alone do not block publication. + +Do not use status fields such as `draft/reviewed` as an approval mechanism. The approved +Git revision is the publication boundary. + +**Step 5: Run tests and lint** + +```bash +cd harness +.venv/bin/pytest tests/test_evidence_authoring.py tests/test_evidence_canonical.py -q +.venv/bin/ruff check tht/evidence/authoring.py tests/test_evidence_authoring.py +``` + +Expected: PASS. + +**Step 6: Commit** + +```bash +git add harness/tht/evidence/authoring.py harness/tht/evidence/__init__.py \ + harness/tests/test_evidence_authoring.py +git commit -m "feat(evidence): validate curated corpus and manifest" +``` + +### Task 3: Implement incremental preparation with one Pi restructuring call + +**Files:** + +- Modify: `harness/tht/evidence/authoring.py` +- Create: `harness/.pi/skills/tht-evidence-authoring/SKILL.md` +- Modify: `harness/tests/test_evidence_authoring.py` +- Create: `harness/tests/test_evidence_pi_restructurer.py` + +**Step 1: Define the restructuring port and request/response models** + +Add: + +```python +class EvidenceRestructurer(Protocol): + def restructure(self, request: RestructureRequest) -> tuple[CuratedEvidence, ...]: ... + +class RestructureRequest(StrictModel): + source_file: str + source_sha256: str + normalized_text: str + previous_units: tuple[CuratedEvidence, ...] = () +``` + +The response is validated through the models from Task 1 before any write. + +**Step 2: Write RED tests for the incremental rules** + +Test: + +- unchanged source: no model call and no file write; +- changed source: exactly one model call; +- new source: new stable IDs; +- removed source: old units become orphans and remain on disk; +- one source may produce several units; +- no returned unit may cite another source; +- prior curated units are included in the request; +- a dirty `evidence/curated/` or `evidence/manifest.yaml` fails before the model call; +- all output writes are staged and atomically replaced only after complete validation. + +Inject the Git status runner and filesystem writer in tests; do not require a real Git +repository for every unit test. + +**Step 3: Implement deterministic source normalization** + +Normalize UTF-8 text with NFC, LF newlines and terminal newline. Preserve meaningful +Markdown, tables, fenced SQL, URLs and list structure. Do not rewrite vocabulary or +infer domain facts in this step. + +**Step 4: Implement `PiEvidenceRestructurer`** + +Invoke Pi as an ephemeral, no-tools process using an argument list, never a shell: + +```python +argv = [ + pi_executable, + "--mode", "text", + "--print", + "--no-session", + "--no-tools", + "--no-extensions", + "--no-context-files", + "--skill", str(skill_path), + f"@{request_path}", + "Return only the JSON object required by the Evidence authoring skill.", +] +``` + +Use a private `TemporaryDirectory`, a bounded timeout, bounded stdout/stderr, and strict +JSON parsing. Do not forward the model's raw output to public JSON errors. The skill must +state: + +- use only facts present in `normalized_text`; +- preserve prior reviewed wording where it remains supported; +- never merge sources; +- emit `review_items` for uncertainty; +- emit exactly the schema-versioned JSON object and no Markdown fence. + +**Step 5: Implement `prepare_workspace_evidence`** + +Return a bounded report with `changed`, `unchanged`, `created`, `orphaned`, `findings`, +and `model_calls`. Keep the model implementation behind `EvidenceRestructurer`. + +**Step 6: Run tests** + +```bash +cd harness +.venv/bin/pytest tests/test_evidence_authoring.py tests/test_evidence_pi_restructurer.py -q +.venv/bin/ruff check tht/evidence/authoring.py tests/test_evidence_pi_restructurer.py +``` + +Expected: PASS with no live model call. + +**Step 7: Commit** + +```bash +git add harness/tht/evidence/authoring.py \ + harness/.pi/skills/tht-evidence-authoring/SKILL.md \ + harness/tests/test_evidence_authoring.py harness/tests/test_evidence_pi_restructurer.py +git commit -m "feat(evidence): prepare curated evidence incrementally" +``` + +### Task 4: Add the authoring CLI + +**Files:** + +- Create: `harness/tht/cli/evidence_cmd.py` +- Modify: `harness/tht/cli/__init__.py` +- Create: `harness/tests/test_evidence_cli.py` + +**Step 1: Write CLI grammar tests** + +Cover: + +```text +tht evidence prepare <workspace-root> [--json] +tht evidence validate <workspace-root> [--json] +``` + +Require an existing canonical Git worktree root. Reject unknown flags, symlinks, a +workspace path outside the Git root, duplicate options and dirty curated state. Ensure +`--json` writes pristine JSON to stdout. + +**Step 2: Run and confirm RED** + +```bash +cd harness +.venv/bin/pytest tests/test_evidence_cli.py -q +``` + +Expected: `evidence` command group is unknown. + +**Step 3: Implement the Typer group** + +Use one top-level authoring group: + +```python +evidence_app = typer.Typer(help="Prepare and validate workspace Evidence") + +@evidence_app.command("prepare") +def prepare_cmd(workspace_root: Path, json_output: bool = False) -> None: ... + +@evidence_app.command("validate") +def validate_cmd(workspace_root: Path, json_output: bool = False) -> None: ... +``` + +Register it in `harness/tht/cli/__init__.py`. Keep this separate from the existing +runtime command `tht preprocess evidence`. + +Exit codes: + +- `0`: prepared/unchanged or valid; +- `2`: unsafe path or CLI misuse; +- `3`: valid drafts but review required; +- `1`: operational/model/structural failure. + +**Step 4: Run tests and CLI help** + +```bash +cd harness +.venv/bin/pytest tests/test_evidence_cli.py tests/test_preprocess_cli.py -q +.venv/bin/tht evidence --help +.venv/bin/tht preprocess evidence --help +``` + +Expected: both command families are present and unambiguous. + +**Step 5: Commit** + +```bash +git add harness/tht/cli/evidence_cmd.py harness/tht/cli/__init__.py \ + harness/tests/test_evidence_cli.py +git commit -m "feat(evidence): expose prepare and validate commands" +``` + +### Task 5: Project typed units into semantic Evidence Fragments + +**Files:** + +- Modify: `harness/tht/evidence/corpus/models.py` +- Modify: `harness/tht/evidence/corpus/chunk.py` +- Modify: `harness/tht/evidence/corpus/normalize.py` +- Modify: `harness/tht/evidence/corpus/pipeline.py` +- Modify: `harness/tht/evidence/preprocessing.py` +- Modify: `harness/tests/test_corpus_models.py` +- Modify: `harness/tests/test_corpus_chunk.py` +- Modify: `harness/tests/test_corpus_pipeline.py` + +**Step 1: Write failing projection tests** + +Test that: + +- a short formula remains one fragment; +- a domain document splits only at semantic section boundaries; +- enum entries are not split in the middle of a value/meaning pair; +- fixed-size fallback is used only for a single oversized section; +- all fragments carry `evidence_id`, `evidence_kind`, `purposes`, scope, language and + provenance; +- fragment IDs and ordinals are deterministic; +- only `curated/**/*.md` is accepted as canonical filesystem content. + +**Step 2: Extend the immutable corpus models** + +Keep `CanonicalDocument` and `CanonicalChunk` as transport-neutral storage models. Put +typed Evidence metadata in their already-safe `metadata` field, with exact allowlisted +keys. Do not make corpus storage depend on Pydantic subtype classes at read time. + +**Step 3: Implement kind-aware fragment rendering** + +Construct embedding text from title, purpose, scope and type-specific data. Example for +a formula: + +```text +Formula: Fascia pediatrica +Concept: fascia pediatrica +Columns: clinical.patient.birth_date +SQL: CASE WHEN ... END +Limitations: ... +``` + +The rendered text is derived; provenance and canonical content remain in the corpus +manifest. + +**Step 4: Run focused pipeline tests** + +```bash +cd harness +.venv/bin/pytest tests/test_corpus_models.py tests/test_corpus_chunk.py \ + tests/test_corpus_normalize.py tests/test_corpus_pipeline.py \ + tests/test_corpus_publish.py -q +``` + +Expected: PASS, including existing rollback, compensation, retention and resume tests. + +**Step 5: Commit** + +```bash +git add harness/tht/evidence/corpus harness/tht/evidence/preprocessing.py \ + harness/tests/test_corpus_models.py harness/tests/test_corpus_chunk.py \ + harness/tests/test_corpus_normalize.py harness/tests/test_corpus_pipeline.py +git commit -m "feat(evidence): build semantic fragments from typed units" +``` + +### Task 6: Upgrade the Qdrant collection contract to named dense and BM25 vectors + +**Files:** + +- Modify: `backend/src/workspaces/qdrant-collection.ts` +- Modify: `backend/test/qdrant-collection.test.ts` +- Modify: `harness/tht/adapters/vector/qdrant.py` +- Modify: `harness/tests/test_qdrant_vector_store.py` +- Modify: `docs/contracts/workspace-preprocessing-cli.md` + +**Step 1: Write RED TypeScript collection-contract tests** + +The required vector contract is: + +```json +{ + "vectors": { + "dense": {"size": 1024, "distance": "Cosine"} + }, + "sparse_vectors": { + "bm25": {"modifier": "idf"} + } +} +``` + +Test that unnamed dense-only, wrong named dimension/distance, missing BM25, or wrong BM25 +modifier are incompatible. Self-heal may add missing payload indexes, but must not +silently convert an incompatible vector configuration. + +Add keyword indexes only for fields used by filters: + +```text +content_hash, document_id, kind, record_key, record_kind, +vector_generation, workspace_id, workspace_revision, +evidence_id, evidence_kind, purposes, concepts, tables, columns, language +``` + +**Step 2: Run the TypeScript test and confirm RED** + +```bash +cd backend +npx vitest run test/qdrant-collection.test.ts +``` + +Expected: current unnamed-vector expectations fail. + +**Step 3: Implement strict named-vector reconciliation** + +Update `createCollection`, `vectorCompatibility` and payload-index reconciliation. +Retain `require_existing` behavior: validation never performs an incompatible migration. + +**Step 4: Update the Python adapter's collection validation** + +`QdrantVectorStore.health()` and `_ensure_collection()` must recognize exactly the same +contract as TypeScript. Add a shared test fixture shape even though the two languages do +not share implementation code. + +**Step 5: Run backend and harness tests** + +```bash +cd backend +npx vitest run test/qdrant-collection.test.ts +npx tsc --noEmit -p . +cd ../harness +.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q +``` + +Expected: PASS. + +**Step 6: Document the required guarded rebuild** + +Update the CLI contract to say that the vector-shape change is incompatible and must be +applied with the existing exact-name guarded command: + +```text +tht --installation <absolute>/thothii-installation.yaml workspace vector rebuild + --workspace <id> --collection <name> --confirm <name> --destroy +``` + +No automatic deletion is allowed. + +**Step 7: Commit** + +```bash +git add backend/src/workspaces/qdrant-collection.ts \ + backend/test/qdrant-collection.test.ts harness/tht/adapters/vector/qdrant.py \ + harness/tests/test_qdrant_vector_store.py docs/contracts/workspace-preprocessing-cli.md +git commit -m "feat(evidence): require dense and bm25 qdrant vectors" +``` + +### Task 7: Add server-side BM25 ingestion and hybrid Query API retrieval + +**Files:** + +- Modify: `harness/tht/ports/vector.py` +- Modify: `harness/tht/adapters/vector/qdrant.py` +- Modify: `harness/tht/vectorstore/records.py` +- Modify: `harness/tht/evidence/corpus/pipeline.py` +- Modify: `harness/tests/test_vector_port_contract.py` +- Modify: `harness/tests/test_qdrant_vector_store.py` +- Modify: `harness/tests/test_corpus_pipeline.py` + +**Step 1: Write RED port tests** + +Extend, do not replace, the current record: + +```python +@dataclass(frozen=True) +class VectorWriteRecord: + record: VectorRecord + embedding: list[float] + content_hash: str + sparse_text: str | None = None + sparse_language: str | None = None +``` + +Extend `VectorStore.search` with keyword-only `query_text` and `query_language`. Existing +callers that omit them remain dense-only. + +**Step 2: Write exact Qdrant request tests** + +For Evidence upsert, assert: + +```json +"vector": { + "dense": [0.1, 0.2], + "bm25": { + "text": "...", + "model": "qdrant/bm25", + "options": {"language": "italian"} + } +} +``` + +For hybrid search, assert two filtered prefetches and default RRF: + +```json +{ + "prefetch": [ + {"query": [0.1, 0.2], "using": "dense", "limit": 20, "filter": {}}, + { + "query": { + "text": "fascia pediatrica", + "model": "qdrant/bm25", + "options": {"language": "italian"} + }, + "using": "bm25", + "limit": 20, + "filter": {} + } + ], + "query": {"rrf": {}}, + "limit": 10, + "with_payload": true +} +``` + +Both prefetch filters must include workspace, revision, active generation and record +kind. Do not add hand-tuned weights. + +**Step 3: Implement named dense writes for all semantic records** + +All existing schema/memory/Evidence records use the `dense` vector name after the +collection rebuild. Only records with `sparse_text` receive `bm25`. + +**Step 4: Implement BM25 Evidence ingestion** + +Populate `sparse_text` and map workspace language `it` to Qdrant's `italian`. Reject an +unsupported language before uploading the generation. Use the same options at ingest +and query time. + +**Step 5: Implement hybrid search with dense fallback only for non-Evidence callers** + +An Evidence hybrid request must fail as unavailable if the configured Qdrant version or +collection contract does not support BM25. It must not silently claim to have run hybrid +search. Existing non-Evidence dense requests continue to work. + +**Step 6: Run focused tests** + +```bash +cd harness +.venv/bin/pytest tests/test_vector_port_contract.py tests/test_qdrant_vector_store.py \ + tests/test_corpus_pipeline.py tests/test_search_pack.py -q +.venv/bin/ruff check tht/ports/vector.py tht/adapters/vector/qdrant.py \ + tht/evidence/corpus/pipeline.py +``` + +Expected: PASS. + +**Step 7: Commit** + +```bash +git add harness/tht/ports/vector.py harness/tht/adapters/vector/qdrant.py \ + harness/tht/vectorstore/records.py harness/tht/evidence/corpus/pipeline.py \ + harness/tests/test_vector_port_contract.py harness/tests/test_qdrant_vector_store.py \ + harness/tests/test_corpus_pipeline.py +git commit -m "feat(evidence): add qdrant bm25 hybrid retrieval" +``` + +### Task 8: Add the typed Evidence search contract and workflow-owned purposes + +**Files:** + +- Modify: `harness/tht/evidence/search.py` +- Modify: `harness/tht/evidence/__init__.py` +- Modify: `harness/tht/cli/search_cmd.py` +- Modify: `harness/tests/test_evidence_facade_contract.py` +- Modify: `harness/tests/test_search_pack.py` +- Create: `harness/.pi/skills/tht-sessione/modules/evidence/runtime-search.md` +- Modify: `harness/.pi/skills/tht-sessione/projection.md.tmpl` +- Modify: `harness/tht/pi_skill_projection.py` +- Modify: `harness/tests/test_pi_skill_projection.py` +- Regenerate: `harness/.pi/skills/tht-sessione/SKILL.md` + +**Step 1: Write RED search-facade tests** + +Introduce: + +```python +class EvidenceSearchContext(BaseModel): + concepts: tuple[str, ...] = () + tables: tuple[str, ...] = () + columns: tuple[str, ...] = () + required_kinds: tuple[EvidenceKind, ...] = () + +def search_evidence( + query: str, + purpose: EvidencePurpose, + context: EvidenceSearchContext, + *, + searcher: ActiveEvidenceSearcher, + embedder: EvidenceQueryEmbedder, + top_n: int = 10, +) -> list[EvidenceResult]: ... +``` + +Test hard filters for workspace/revision/generation and explicit `required_kinds`. Test +that purpose/kind/scope preferences are deterministic tie-breakers when they were not +requested as hard filters. + +**Step 2: Test fragment grouping** + +Two returned fragments with the same `evidence_id` must become one `EvidenceResult`, +with the best score, ordered matching excerpts, canonical citation and no duplicate +unit. + +**Step 3: Preserve fail-closed graceful degradation** + +Test absent ACTIVE corpus, revision mismatch, unavailable Qdrant and malformed payload. +All return no Evidence plus a bounded warning through the existing search-pack contract; +none uses stale rows. + +**Step 4: Implement the facade and CLI mapping** + +Keep Qdrant syntax inside `tht.evidence`. `search_cmd.py` translates command inputs to +the facade and renders results; it must not duplicate ranking logic. + +**Step 5: Extract Evidence instructions into a module fragment** + +Move the common F1/F3/F4 Evidence rules from the projection template into +`modules/evidence/runtime-search.md`. The fragment must state: + +- candidates are not truth; +- pass the phase-appropriate purpose; +- show provenance; +- formulas are `kind=formula`, not a separate search store; +- absence of Evidence is visible but does not stop the whole session. + +Register the fragment in the static `FRAGMENT_ORDER` and regenerate. + +**Step 6: Run tests** + +```bash +cd harness +python -m tht.pi_skill_projection --write +python -m tht.pi_skill_projection --check +.venv/bin/pytest tests/test_evidence_facade_contract.py tests/test_search_pack.py \ + tests/test_pi_skill_projection.py -q +``` + +Expected: PASS and no direct edit drift in generated `SKILL.md`. + +**Step 7: Commit** + +```bash +git add harness/tht/evidence/search.py harness/tht/evidence/__init__.py \ + harness/tht/cli/search_cmd.py harness/tests/test_evidence_facade_contract.py \ + harness/tests/test_search_pack.py \ + harness/.pi/skills/tht-sessione/modules/evidence/runtime-search.md \ + harness/.pi/skills/tht-sessione/projection.md.tmpl \ + harness/.pi/skills/tht-sessione/SKILL.md harness/tht/pi_skill_projection.py \ + harness/tests/test_pi_skill_projection.py +git commit -m "refactor(evidence): own typed runtime retrieval" +``` + +### Task 9: Migrate formulas into Curated Evidence + +**Files:** + +- Modify: `harness/tht/evidence/formula_store.py` +- Modify: `harness/tht/evidence/session.py` +- Modify: `harness/tht/cli/search_cmd.py` +- Create: `harness/.pi/skills/tht-sessione/modules/evidence/formula-proposals.md` +- Modify: `harness/.pi/skills/tht-sessione/projection.md.tmpl` +- Modify: `harness/tht/pi_skill_projection.py` +- Modify: `harness/tests/test_formula.py` +- Modify: `harness/tests/test_formula_wiring.py` +- Create: `harness/tests/test_evidence_formula_migration.py` +- Modify: `harness/tests/test_pi_skill_projection.py` +- Regenerate: `harness/.pi/skills/tht-sessione/SKILL.md` + +**Step 1: Write RED migration tests** + +Map an approved legacy `ConceptFormula` to a Curated Evidence formula while preserving: + +- concept; +- SQL; +- columns; +- sources as provenance notes; +- stable deterministic ID; +- reviewed content wording. + +Reject `auto` and unresolved `draft` formulas from direct publication; they become +Formula proposals. + +**Step 2: Define the session proposal contract** + +Add a schema-versioned formula-proposal projection in the session artifact. Keep +`concept_formula_approved` / `concept_formula_rejected` decisions unchanged because they +record a session-local choice, not repository publication. + +**Step 3: Remove the separate runtime formula lookup** + +Change `tht search find --kind formula` to call typed Evidence search with +`required_kinds=("formula",)`. Keep the legacy store readable only for the migration +command/window, with a deprecation warning in human output and no warning leakage into +pristine JSON. + +**Step 4: Add and project formula instructions** + +The Evidence fragment must say that a newly synthesized formula is a session proposal +and cannot be treated as Published Evidence. + +**Step 5: Run tests** + +```bash +cd harness +python -m tht.pi_skill_projection --write +.venv/bin/pytest tests/test_formula.py tests/test_formula_wiring.py \ + tests/test_evidence_formula_migration.py tests/test_pi_skill_projection.py \ + tests/test_decision_min_phase.py tests/test_workflow_observable_contract.py -q +``` + +Expected: PASS; decision phase ownership remains F4. + +**Step 6: Commit** + +```bash +git add harness/tht/evidence/formula_store.py harness/tht/evidence/session.py \ + harness/tht/cli/search_cmd.py \ + harness/.pi/skills/tht-sessione/modules/evidence/formula-proposals.md \ + harness/.pi/skills/tht-sessione/projection.md.tmpl \ + harness/.pi/skills/tht-sessione/SKILL.md harness/tht/pi_skill_projection.py \ + harness/tests/test_formula.py harness/tests/test_formula_wiring.py \ + harness/tests/test_evidence_formula_migration.py harness/tests/test_pi_skill_projection.py +git commit -m "refactor(evidence): unify formulas with typed evidence" +``` + +### Task 10: Add the small retrieval evaluation command + +**Files:** + +- Create: `harness/tht/evidence/evaluation.py` +- Modify: `harness/tht/cli/evidence_cmd.py` +- Create: `harness/tests/test_evidence_evaluation.py` +- Modify: `harness/tests/test_evidence_cli.py` + +**Step 1: Write RED schema tests** + +Use a deliberately small format: + +```yaml +schema_version: 1 +queries: + - id: pediatric-formula + query: Come distinguo i pazienti pediatrici? + purpose: sql_generation + expected: + - formula:fascia-pediatrica +``` + +Require unique query IDs, nonempty expected IDs and only public purpose values. + +**Step 2: Write RED metric tests** + +Compute `hit_at_5`, `hit_at_10`, missing expected IDs, empty-result queries and counts by +expected `kind`. Do not add nDCG, relevance grading or an evaluation database in v1. + +**Step 3: Implement the evaluator and CLI** + +Expose: + +```text +tht evidence evaluate <workspace-root> -c <runtime-config> [--json] +``` + +The report includes workspace revision, active vector generation and the fixed default +RRF configuration. It is read-only. + +**Step 4: Run tests** + +```bash +cd harness +.venv/bin/pytest tests/test_evidence_evaluation.py tests/test_evidence_cli.py -q +.venv/bin/ruff check tht/evidence/evaluation.py tests/test_evidence_evaluation.py +``` + +Expected: PASS. + +**Step 5: Commit** + +```bash +git add harness/tht/evidence/evaluation.py harness/tht/cli/evidence_cmd.py \ + harness/tests/test_evidence_evaluation.py harness/tests/test_evidence_cli.py +git commit -m "feat(evidence): evaluate retrieval with a small fixture" +``` + +### Task 11: Enforce curated-only runtime ingestion and update contracts + +**Files:** + +- Modify: `backend/src/workspaces/schema.ts` +- Modify: `backend/src/workspaces/runtime-renderer.ts` +- Modify: `backend/test/workspaces/schema.test.ts` +- Modify: `backend/test/workspaces/runtime-renderer.test.ts` +- Modify: `docs/contracts/workspace-evidence-v3.md` +- Modify: `docs/contracts/workspace-preprocessing-cli.md` +- Modify: `docs/architecture/overview.md` +- Modify: `PROJECT_STATE.md` + +**Step 1: Write RED descriptor/rendering tests** + +For filesystem Evidence, make `patterns: ["curated/**/*.md"]` the documented and +generated default for the new authoring layout. Continue accepting an explicitly +configured safe pattern for non-Git HTTP/S3 compatibility, but reject a filesystem +descriptor that includes both `source/**` and `curated/**` once it declares the new +layout version. + +If a schema-version field is required to preserve compatibility, add it to the Evidence +subcontract, not to the whole workspace descriptor. + +**Step 2: Implement the narrowest compatible descriptor change** + +The materializer continues to copy the entire `evidence/` tree at the pinned commit. +Only the rendered runtime acquisition patterns restrict preprocessing to `curated/`. +Do not duplicate or move P6 materialization logic. + +**Step 3: Update documentation contracts** + +Document: + +- source/curated layout; +- Git publication boundary; +- no runtime writes; +- validation before indexing; +- dense+BM25 collection contract and guarded rebuild; +- exact public operation names and JSON status additions, if any. + +**Step 4: Run backend and harness contract gates** + +```bash +cd backend +npx vitest run test/workspaces/schema.test.ts test/workspaces/runtime-renderer.test.ts \ + test/workspaces/evidence/materialization.test.ts \ + test/workspaces/evidence/preprocessing.test.ts +npx tsc --noEmit -p . +cd ../harness +.venv/bin/pytest tests/test_registry_evidence_config.py \ + tests/test_filesystem_evidence_source.py tests/test_preprocess_cli.py -q +``` + +Expected: PASS; P6 materialization safety remains unchanged. + +**Step 5: Commit** + +```bash +git add backend/src/workspaces/schema.ts backend/src/workspaces/runtime-renderer.ts \ + backend/test/workspaces/schema.test.ts \ + backend/test/workspaces/runtime-renderer.test.ts \ + docs/contracts/workspace-evidence-v3.md \ + docs/contracts/workspace-preprocessing-cli.md docs/architecture/overview.md PROJECT_STATE.md +git commit -m "docs(evidence): publish curated-only workspace contract" +``` + +### Task 12: Migrate PSD and perform acceptance + +**Files in ThothII:** + +- Create: `docs/testing/evidence-restructuring-manual.md` +- Create: `scripts/evidence-restructuring-acceptance.sh` +- Create: `harness/tests/fixtures/evidence_authoring/poorly_structured.md` +- Modify: `PROJECT_STATE.md` + +**Files in the external authoring repository:** + +- Move: `/Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/<current-folders>` + to `/Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/source/` +- Create: `/Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/curated/<kind>/` +- Create: `/Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/manifest.yaml` +- Create: `/Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/evaluation.yaml` +- Modify: `/Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/README.md` +- Modify: `/Users/mp/projects/tht-workspace-psd/psd-clinical/workspace.yaml` + +Do not modify the external repository until the owner confirms the migration window and +the exact target branch. Treat that as the only manual authorization gate in this task. + +**Step 1: Add a hermetic badly-structured fixture** + +The fixture must contain prose, a rough list, an enum, an URL, an ambiguous statement +and a SQL formula candidate. The acceptance runner must prove: + +- split into multiple typed units; +- ambiguity becomes `review_items`; +- no cross-source merge; +- unchanged rerun is a no-op; +- human edit is preserved; +- dirty-tree refusal; +- validation blocks unresolved review; +- validated corpus indexes and searches hybrid; +- Qdrant failure remains fail-closed. + +Use a fake restructurer for hermetic CI. The real Pi call is a separate manual check. + +**Step 2: Run the complete automated gates before external writes** + +```bash +cd harness +.venv/bin/pytest -q +.venv/bin/ruff check . +cd ../backend +npx vitest run +npx tsc --noEmit -p . +cd ../tools/tht +go test ./... +go build ./cmd/tht +cd ../.. +bash scripts/evidence-restructuring-acceptance.sh +``` + +Expected: all suites and acceptance checks PASS. If pre-existing unrelated failures +remain, record exact names and prove they reproduce at the baseline commit before +continuing. + +**Step 3: Stop for the owner migration gate** + +Provide: + +- clean ThothII commit; +- test and acceptance summary; +- proposed PSD branch name; +- exact list of 36 source files to move; +- rollback command based on the pre-migration PSD commit; +- notice that Qdrant rebuild is destructive but scoped by exact collection-name guards. + +Do not infer approval from prior design acceptance. + +**Step 4: Migrate the PSD repository after approval** + +Use `tht evidence prepare`, review the Git diff, resolve all review items manually, run +`tht evidence validate`, and create the approximately twenty evaluation queries. Do not +auto-merge or auto-push unless separately requested. + +**Step 5: Activate and rebuild with the existing guarded operator path** + +After the PSD merge/pull and activation, inspect first: + +```text +tht --installation <absolute>/thothii-installation.yaml workspace vector inspect + --workspace psd-clinical --json +``` + +Then use the exact descriptor-owned name in the guarded rebuild command. Run +`workspace preprocess evidence`, `tht evidence evaluate`, and the manual F1/F3/F4 +walkthrough. + +**Step 6: Record acceptance and commit ThothII documentation** + +`docs/testing/evidence-restructuring-manual.md` must record separate outcomes for: + +- authoring and Git review; +- collection rebuild; +- preprocessing generation publication; +- hybrid retrieval evaluation; +- formula retrieval; +- graceful degradation; +- complete session behavior. + +Update `PROJECT_STATE.md` only with observed results and immutable commit/run IDs. + +```bash +git add docs/testing/evidence-restructuring-manual.md \ + scripts/evidence-restructuring-acceptance.sh \ + harness/tests/fixtures/evidence_authoring/poorly_structured.md PROJECT_STATE.md +git commit -m "test(evidence): record restructuring acceptance" +``` + +## Final verification checklist + +Before claiming completion, verify: + +- [ ] `CONTEXT.md` and both Evidence plan documents use the same terminology. +- [ ] All eight Evidence kinds have type-specific positive and negative tests. +- [ ] Pi is called once per changed source, with no tools and no saved session. +- [ ] `prepare` refuses dirty curated state and never deletes orphans. +- [ ] `validate` blocks unresolved review items. +- [ ] Runtime reads only `curated/**/*.md` from the pinned Git revision. +- [ ] The TypeScript and Python Qdrant compatibility checks agree. +- [ ] Qdrant collection uses named `dense` plus `bm25` with IDF. +- [ ] Italian BM25 options are identical during ingest and query. +- [ ] Hybrid search uses two prefetches and default RRF. +- [ ] Hard filters always include workspace, revision and active generation. +- [ ] Formula runtime lookup uses typed Evidence; session proposals remain non-published. +- [ ] Fragment hits are grouped into complete Evidence Units. +- [ ] Evaluation reports hit@5 and hit@10 against a versioned fixture. +- [ ] Existing corpus rollback, compensation, retention and resume tests still pass. +- [ ] Backend Vitest and TypeScript gates pass. +- [ ] Harness pytest and Ruff gates pass. +- [ ] Native `tht` Go tests/build pass. +- [ ] External PSD writes occurred only after explicit migration authorization.