docs: consolidate historical records and verify public manual publication
Publish documentation / publish (push) Successful in 29s

This commit is contained in:
Codex
2026-09-15 14:37:29 +02:00
parent 5f3a7f5975
commit 4ff91e8d6e
63 changed files with 633 additions and 6598 deletions
@@ -1,695 +0,0 @@
# Metadata Catalog di ThothII: ricognizione ThothAI e percorso incrementale
Data: 2026-08-26; aggiornato 2026-08-27
Stato: ricognizione e progettazione completate; navigazione, CRUD Workspace Database, Catalog
Table, Catalog Column, Catalog Relationship e sincronizzazione durevole dello schema implementati
il 2026-08-27. Generazione AI e integrazione con il workflow core restano negli step successivi.
## Obiettivo
ThothII deve introdurre un contesto amministrativo separato, il **Metadata Catalog**, per gestire
il database associato a ciascun workspace, la sua struttura fisica introspezionata e i metadati
semantici oggi rappresentati da `schema/annotations.yaml`.
Il programma procede per step indipendenti. Il primo step ha aggiunto l'accesso dalla sidebar; il
secondo ha sostituito la superficie vuota con il CRUD di configurazione, il PostgreSQL interno e i
test di connessione; gli step successivi hanno aggiunto navigazione gerarchica, colonne, relazioni
fisiche e sincronizzazione durevole dell'intero schema. Non introduce ancora generazione AI o
integrazione con il workflow core.
Questa analisi usa come riferimento il working tree legacy osservato in
`Thoth/ThothAI`. Non è stato verificato che quel contenuto corrisponda a una release o a un tag
canonico; i percorsi e i comportamenti descrivono il sorgente disponibile il 2026-08-26.
## Decisioni già confermate
1. Ogni Workspace Database appartiene a un solo workspace tramite un `workspace_id` obbligatorio e
univoco; un workspace può avere al massimo un Workspace Database. Poiché i workspace non sono
righe del catalogo PostgreSQL, l'associazione è un riferimento logico validato contro
`thoth-workspaces.yaml`, non una foreign key SQL.
2. Il CRUD non crea né rinomina workspace. Identità e lista ordinata dei workspace restano
autorevoli in `thoth-workspaces.yaml`; il catalogo conserva il loro identificatore stabile.
3. La struttura fisica viene acquisita interrogando il database esterno tramite i dati di
connessione registrati per il Workspace Database.
4. I contenuti semantici equivalenti a `annotations.yaml` vengono generati con l'AI e conservati nel
PostgreSQL interno.
5. Per PSD è prevista l'importazione delle annotations esistenti. Gli altri database partiranno
dalla struttura introspezionata e genereranno i metadati semantici da zero.
6. `annotations.yaml` sarà sostituito anche come input del core in uno step futuro. Il repository è
in fase di test e non è richiesta la conservazione delle sessioni esistenti durante il cutover.
7. La gestione catalogo resta una superficie separata dal processo NL→SQL. La futura integrazione
deve essere esplicita e non deve modificare fasi, gate o semantica del workflow.
8. Il link iniziale è visibile agli utenti con `workspace.manage`, usa stato React locale e non
introduce un router.
9. La pagina iniziale è vuota, segue il tema, nasconde l'intera colonna core e non interrompe una
sessione live. Le azioni di apertura, resume o creazione sessione riportano al core.
10. La compatibilità con il modello ThothAI è semantica, non una copia letterale: configurazione e
contenuti semantici sono campi relazionali mutabili, mentre identità e appartenenza della
struttura fisica derivano dall'introspezione; i segreti restano nel secret store e lo stato dei
job non viene mescolato ai dati amministrativi.
11. Il CRUD amministra il Metadata Catalog e non esegue DDL sul database esterno, che resta
read-only.
12. La prima versione supporta PostgreSQL; il confine di introspezione dovrà permettere di
aggiungere altri dialetti senza cambiare il modello del catalogo.
13. I segreti dei Workspace Database riusano il secret store cifrato di ThothII. Il catalogo
conserva riferimenti ai segreti e nessuna API, esportazione o log ne restituisce i valori.
14. La UI usa AG Grid Community per la lista master e un pannello React separato per il dettaglio;
non dipende dalle funzionalità master-detail di AG Grid Enterprise.
15. Un Workspace Database il cui `workspace_id` scompare dal catalogo YAML non viene cancellato
automaticamente: diventa orphaned e può soltanto essere recuperato, riassegnato o eliminato
esplicitamente da un amministratore.
16. La prima vertical slice gestisce configurazione del Workspace Database, riferimenti ai segreti,
test di connessione e stato. La seconda gestisce le Catalog Table: la collezione e i nomi sono
controllati dall'introspezione, mentre la descrizione curata è modificabile. Le slice successive
hanno aggiunto Catalog Column, Catalog Relationship e sincronizzazione durevole dello schema.
17. Il modello non conserva il `name` libero di ThothAI: nome e ID visualizzati appartengono al
workspace YAML, mentre `database_name` identifica il database PostgreSQL esterno.
18. Database management supporta i tre trasporti già riconosciuti da ThothII: `postgres_direct`,
`rest_api` e `ssh_tunnel`. PSD rimane un solo Workspace Database: usa la connessione diretta sul
server e l'endpoint REST in locale tramite una Database Binding specifica dell'installazione.
Questo supporto non abilita automaticamente `ssh_tunnel` nel runtime NL→SQL.
19. Una configurazione può essere salvata prima di una connessione riuscita. Il test separato
produce uno stato `untested`, `reachable` o `failed`; attivazione e introspezione richiedono uno
stato raggiungibile.
20. Il CRUD e il test di connessione richiedono `database.manage`; inserimento e sostituzione dei
segreti continuano a richiedere `workspace.secrets.manage`.
21. Il Workspace Database e il modo di raggiungerlo sono entità distinte. Ogni catalogo di
installazione conserva una sola Database Binding attiva per workspace: PSD usa `rest_api` in
locale e `postgres_direct` sul server senza duplicare il Workspace Database.
22. Nel modello finale il Metadata Catalog è autorevole per engine, `database_name`, schema,
capacità e binding. Lo YAML resta autorevole per identità e contenuti del workspace; i campi
DWH correnti saranno importati, confrontati e rimossi soltanto durante un cutover esplicito.
23. La lista master è l'unione fra workspace YAML e record del catalogo: mostra workspace
`unconfigured`, database configurati e record `orphaned`.
24. Ogni introspezione registra le capability disponibili. Una capability `unavailable` non viene
rappresentata come una collezione osservata ma vuota; REST può completare con successo anche
quando indici o enum non sono supportati.
25. Il Metadata Catalog non introduce snapshot, draft o pubblicazioni. Configurazione e contenuti
semantici, inclusi quelli futuri generati dall'AI, sono normali campi modificabili; la struttura
osservata cambia soltanto con una sincronizzazione esplicita.
26. Il normale Delete elimina realmente il Workspace Database, la Database Binding e i relativi
record catalogo e segreti. Non modifica il DWH esterno né il repository YAML; il workspace torna
visibile nella lista master come `unconfigured`.
27. La prima versione gestisce un solo schema obbligatorio per Workspace Database, identificato
dalla coppia `database_name + schema`; per PSD la coppia è `postgres + datawarehouse`.
28. I record mantengono soltanto `created_at`, `updated_at` e un contatore `version` per optimistic
concurrency. Non esistono storico delle revisioni, rollback o audit applicativo delle modifiche.
29. `workspace_databases` conserva soltanto UUID, `workspace_id` unique, engine, `database_name`,
schema, timestamp e version. Il nome visualizzato appartiene al workspace YAML.
30. Ogni Workspace Database ha al massimo una riga `database_bindings`. Una singola tabella usa
check constraint dipendenti da `transport` per i campi direct, REST e SSH; non esiste un flag
`active`, perché ciascuna installazione conserva una sola binding.
31. `rest_api` configura il Thoth REST Connector tipizzato: base URL, autenticazione e TLS sono dati
della binding, mentre path RPC e shape delle risposte appartengono al contratto applicativo e non
sono liberamente configurabili.
32. Il test connessione usa soltanto una configurazione già salvata ed è associato alla sua
`version`. Ogni modifica della binding o dei segreti invalida il risultato precedente e riporta
lo stato a `untested`.
33. Password, API key e chiavi sono write-only: l'API espone soltanto `configured`, un campo vuoto
conserva il valore esistente e la sostituzione è un'azione esplicita. Delete rimuove anche i
segreti associati.
34. La pagina usa AG Grid come master e un form React come detail, con sezioni Database, Connection
e TLS/SSH condizionali. Non esiste un'azione globale `Add database`: ogni riga `unconfigured`
offre `Configure catalog`, apre il form già vincolato a quello specifico workspace YAML e crea il
record soltanto al Save; `workspace_id` non è selezionabile né modificabile.
35. La grid mostra separatamente revisione/Evidence del workspace, binding runtime NL→SQL e
configurazione del Metadata Catalog, oltre a database, schema, endpoint e ultimo aggiornamento.
Su schermi piccoli il dettaglio occupa il pannello completo. Il cambio riga con modifiche non
salvate e Delete richiedono conferma, senza conferma testuale tipizzata.
36. La Database Binding conserva `connection_status`, `tested_version`, `last_tested_at`, un codice
errore e un messaggio breve sanificato. Non conserva stack trace, DSN, credenziali o output grezzo
del driver.
37. Le API vivono sotto `/api/catalog`: list/create di `/databases`, get/patch/delete di
`/databases/:id`, sostituzione dei segreti sotto `/databases/:id/secrets`, test connessione sotto
`/databases/:id/test` e list/patch/sync delle tabelle sotto `/databases/:id/tables`.
38. `GET /api/catalog/databases` restituisce l'intera master list unificata; AG Grid Community applica
client-side ricerca, filtri e ordinamento. La prima versione non introduce paginazione server o
funzionalità AG Grid Enterprise.
39. Il Metadata Catalog vive nello stesso processo Fastify come modulo isolato con repository,
service, route, diagnostica e readiness proprie. L'indisponibilità del catalogo non modifica
sessioni, SSE o health del core e non giustifica ancora un microservizio separato.
40. Il backend mantiene `pg@8.22.0` e aggiunge `kysely@0.29.5` per query e transazioni tipizzate. Le
migrazioni Kysely sono timestampate, compilate con il backend ed eseguite da un comando
`catalog:migrate` separato; l'applicazione non migra automaticamente il database all'avvio.
41. Lo stack aggiunge un servizio interno `catalog-db` con volume persistente, ruolo runtime DML,
ruolo migrator DDL e job one-shot `catalog-migrate`. Un catalogo indisponibile produce 503 sulle
sole route catalogo.
42. La prima vertical slice è amministrativa: scrive il catalogo ma non cambia ancora il runtime di
sessioni e workflow, che continua a usare YAML e binding correnti fino al cutover esplicito.
43. `Configure` precompila senza salvare engine, database e schema dal descriptor e i dati non
sensibili dalla binding effettiva. L'amministratore verifica, inserisce i segreti e salva; non
esiste importazione silenziosa.
44. Unit e route test usano un repository fake; una suite PostgreSQL Testcontainers separata verifica
migrazioni, constraint, transazioni, optimistic concurrency e cascade. SQLite ed emulatori non
sono sostituti ammessi per questi test.
45. La navigazione delle entità catalogo è gerarchica e senza scorciatoie globali: `Databases →
Database → Overview | Tables → Table`. Non esistono una voce globale Tables, un filtro globale
Database o una preselezione implicita; Columns continuerà sotto Table e Relationships sotto
Database.
46. Una Catalog Table conserva nome fisico, `source_comment`, descrizione curata nullable,
`generated_description` nullable per lo step AI futuro, version e timestamp. La UI mostra come
tre campi indipendenti senza fallback visivo: source comment read-only, generated description
modificabile e description modificabile. I valori null restano celle e controlli vuoti.
47. Le Catalog Table non possono essere aggiunte o rinominate manualmente. Un amministratore può
però ripulire esplicitamente le proiezioni nel Metadata Catalog senza modificare il database
esterno; `Sync tables` legge le tabelle PostgreSQL ordinarie e partizionate dello schema scelto,
mentre viste e materialized view sono escluse.
48. La sincronizzazione è esplicita. La scansione avviene fuori dalla transazione del catalogo; il
diff viene applicato atomicamente soltanto se la version del Workspace Database è ancora quella
sottoposta a scansione. Una scansione fallita non modifica il catalogo.
49. Tabelle nuove vengono create, i commenti sorgente vengono aggiornati e quelle non più osservate
vengono eliminate definitivamente. La rimozione di tabelle, colonne o relazioni richiede la
conferma dell'esatto piano distruttivo; se il secondo scan produce una fotografia differente,
l'applicazione richiede una nuova conferma.
50. Un rename fisico è intenzionalmente delete più create e perde i metadati curati. Le colonne e
relazioni dipendenti vengono eliminate in cascade insieme alla Catalog Table.
51. L'introspezione vive nel modulo catalogo Fastify dietro un adapter. PostgreSQL diretto e tunnel
SSH usano il catalogo `pg_catalog`; REST preferisce il contratto tipizzato
`POST /rpc/schema_snapshot` e, quando quell'RPC non è esposto, usa come fallback compatibile una
singola query read-only tramite `POST /rpc/run_query`. Entrambi i percorsi devono produrre la
stessa fotografia v1 stretta descritta in `docs/contracts/catalog-schema-snapshot.md`.
52. Test connessione e sincronizzazione sono serializzati per Workspace Database, hanno timeout e
richiedono che la binding nella version corrente abbia un test `reachable` prima di qualsiasi
Catalog Sync Run. La scansione asincrona ha un timeout separato, di default dieci minuti.
53. Il tunnel SSH usa OpenSSH in modalità stdio `-W`, chiave privata e passphrase opzionale dal
secret store, `known_hosts` obbligatorio, `StrictHostKeyChecking=yes`, agent e configurazione
globale disabilitati. Non è ammesso TOFU. TLS PostgreSQL con CA e server name resta verificato
anche attraverso il tunnel.
54. In questo slice `ssh_tunnel` è una binding supportata da Database management per Test connection
e Schema Sync. Il renderer e il runtime delle sessioni NL→SQL restano fuori scope e continuano a
rifiutarla finché non verrà deciso il relativo cutover.
55. I menu di azione a livello Workspace Database espongono separatamente `Synchronize tables`,
`Synchronize relationships` e `Synchronize all`. Su una selezione di
database lo scope scelto viene avviato per ogni database idoneo; non viene sostituito
implicitamente con una sincronizzazione completa.
56. Lo scope Columns è disponibile dalla grid Tables e limita la riconciliazione alle tabelle
selezionate; la pagina Columns non espone azioni di sincronizzazione. La grid Tables espone
`Synchronize columns` sulle tabelle selezionate.
## Correzione del modello mentale corrente
`schema/annotations.yaml` non contiene l'intero schema del database.
- `physical.yaml` è un artefatto derivato dall'introspezione. Contiene database, schema, timestamp,
tabelle, colonne, tipi, nullability, default, primary key, commenti sorgente, esempi, foreign key
fisiche e indici.
- `annotations.yaml` contiene metadati curati: descrizioni e concetti delle tabelle; descrizioni,
sinonimi, concetti, evidence, note e override `eligible` delle colonne; foreign key logiche.
- Il rendering M-Schema fonde questi due input. Le annotations prevalgono sui commenti sorgente e
le relazioni logiche vengono unite alle foreign key fisiche.
La sostituzione del solo file annotations non elimina automaticamente l'introspezione fisica. Il
nuovo catalogo dovrà conservare una distinzione esplicita fra fatti osservati nel database e
contenuto semantico modificabile.
## Architettura ThothII rilevante
### Autorità e revisionamento attuali
Il repository dei workspace contiene:
```text
thoth-workspaces.yaml
<workspace-id>/workspace.yaml
<workspace-id>/schema/annotations.yaml
<workspace-id>/evidence/**
```
Il backend legge descriptor e annotations allo stesso commit Git. Durante l'attivazione valida il
blob, lo copia atomicamente nello snapshot immutabile della revisione e registra commit, blob ID e
digest. Le nuove sessioni vengono legate a quella revisione; resume e SQL salvato riaprono lo stesso
snapshot.
Punti principali:
- `backend/src/workspaces/schema.ts`: descriptor v3 e singolo `dwh.database`/`dwh.schema`;
- `backend/src/workspaces/git-repository.ts`: lettura sicura del blob annotations al commit;
- `backend/src/workspaces/registry.ts`: validazione e attivazione atomica;
- `backend/src/workspaces/annotations-sync.ts`: materializzazione revision-qualified;
- `backend/src/workspaces/runtime-config-lease.ts`: binding dello snapshot al runtime;
- `harness/tht/mschema/models.py`: contratti `PhysicalSchema` e `Annotations`;
- `harness/tht/mschema/render.py`: fusione fisico/semantico;
- `harness/tht/cli/vector_cmd.py`: indicizzazione schema in Qdrant.
### Consumatori da preservare al cutover futuro
Le annotations incidono oggi su:
- override `eligible` prima del campionamento LSH;
- suggerimento, controllo e accettazione delle foreign key logiche;
- descrizioni, concetti e sinonimi dei record schema in Qdrant;
- retrieval delle tabelle e colonne candidate;
- rendering M-Schema usato dal gate F4 e dalla generazione SQL;
- digest della revisione accettata durante il preprocessing.
Il futuro cutover non potrà limitarsi a rimuovere il file: dovrà fornire al core lo stesso contenuto
effettivo, con un'identità coerente e test di equivalenza. Poiché non occorre preservare le sessioni
di test esistenti, non serve progettare compatibilità con i vecchi manifest, ma resta necessario
evitare letture parziali o semanticamente incoerenti.
## Inventario ThothAI
### Modelli legacy
I modelli sono definiti in `Thoth/ThothAI/backend/thoth_core/models.py`.
#### `SqlDb`
Campi di connessione osservati:
- `name`;
- `db_host`, `db_port`;
- `db_type`;
- `db_name`, `schema`;
- `user_name`, `password`;
- `db_mode`;
- configurazione SSH e Informix opzionale.
Il modello contiene anche scope, JSON dello scope, ERD, direttive, campi GDPR, collegamento a
`VectorDb` e numerosi campi di stato/task/log per lavori AI asincroni.
I tipi legacy dichiarati sono Informix, MariaDB, MySQL, Oracle, PostgreSQL, SQL Server e SQLite.
Questo elenco non costituisce automaticamente un requisito per ThothII: il core corrente supporta
PostgreSQL e l'estensione ad altri dialetti dovrà essere decisa separatamente.
#### `SqlTable`
- `name`;
- `description`;
- `generated_comment`;
- foreign key obbligatoria a `SqlDb`, con cancellazione cascade.
#### `SqlColumn`
- `original_column_name` e alias `column_name`;
- `data_format` normalizzato;
- `column_description`;
- `generated_comment`;
- `value_description`;
- stringhe denormalizzate `pk_field` e `fk_field`;
- foreign key obbligatoria a `SqlTable`, con cancellazione cascade.
#### `Relationship`
Contiene quattro foreign key obbligatorie:
- `source_table` e `source_column`;
- `target_table` e `target_column`.
Il form admin verifica che le tabelle appartengano allo stesso database e che ogni colonna
appartenga alla tabella selezionata. Il database non impone però gli stessi check.
#### `Workspace`
ThothAI usa `Workspace.sql_db` come foreign key nullable verso `SqlDb`: un workspace seleziona un
solo DB, mentre lo stesso DB può essere riusato da più workspace. ThothII adotterà invece una
relazione uno-a-uno: `workspace_id` deve essere unico nel catalogo.
### Lacune dei constraint legacy
Non risultano constraint database-level per:
- unicità del nome database nel workspace;
- unicità `(database, table name)`;
- unicità `(table, column name)`;
- unicità degli estremi di una relationship;
- appartenenza degli estremi della relationship allo stesso database;
- corrispondenza fra colonna e tabella dichiarata.
ThothII deve applicare queste invarianti sia nel database interno sia nel servizio applicativo. La
sola validazione del form non è sufficiente perché API, import e job la possono aggirare.
### Django Admin e UX da replicare concettualmente
ThothAI espone il CRUD tramite il Django Admin standard, registrato da
`backend/thoth_core/admin.py` e pubblicato su `/admin/`.
Capacità utili:
- lista database con ricerca per nome, host, tipo, database e schema;
- fieldset separati per identità, connessione, autenticazione, SSH e stato;
- lista tabelle filtrabile per database;
- lista colonne filtrabile in cascata per database e tabella;
- lista relazioni con estremi leggibili e filtri per database e tabelle;
- form relazione con dropdown dipendenti database → tabella → colonna;
- validazione degli estremi prima del salvataggio;
- azioni separate per test connessione, introspezione, import/export e generazione AI;
- azioni bulk sulle righe selezionate.
ThothII deve replicare i contratti di interazione e validazione, non il rendering server-side o i
template Django.
### Introspezione legacy
`Thoth/ThothAI/backend/thoth_core/dbmanagement.py` usa `thoth-dbmanager` per:
1. costruire l'adapter del dialetto;
2. acquisire tabelle;
3. acquisire e normalizzare colonne e tipi;
4. acquisire relazioni;
5. creare le eventuali colonne mancanti necessarie alle relazioni;
6. aggiornare i campi PK/FK denormalizzati.
Il comportamento è principalmente additivo: usa `get_or_create` o controlli `exists`, aggiorna
alcuni commenti, ma non riconcilia in modo completo rename, rimozioni o drift. Non va copiato così
com'è. Il processo ThothII implementato distingue scansione, differenze osservate e applicazione
della nuova snapshot.
### Generazione AI legacy
ThothAI dispone di azioni e workflow per:
- commenti delle tabelle;
- commenti delle colonne;
- scope del database;
- ERD Mermaid;
- documentazione del database;
- analisi GDPR.
Per il requisito attuale sono direttamente rilevanti descrizioni di tabelle e colonne, scope e
metadati semantici. ERD, documentazione aggregata e GDPR sono estensioni future, non prerequisiti
del CRUD iniziale.
La separazione `description`/`generated_comment` del legacy non offre versioning o approvazione
robusti. Nei passi successivi andrà deciso se l'output AI è una proposta revisionabile o diventa
immediatamente il valore editabile corrente.
### Import ed export legacy
ThothAI offre:
- CSV di database, tabelle, colonne e relazioni;
- export di struttura per workspace;
- import mediante `import_db_structure`;
- script SQL dei commenti per più dialetti;
- aggiornamento delle descrizioni colonna da CSV.
Il futuro import PSD dovrà leggere il contratto YAML corrente e convertirlo su chiavi naturali,
non riutilizzare gli ID numerici Django. Deve essere idempotente e produrre un report di elementi
creati, aggiornati, ignorati o non risolti.
## Comandi osservati in ThothAI
### Backend locale
Eseguiti da `Thoth/ThothAI/backend`:
```sh
uv sync
uv run python manage.py migrate
uv run python manage.py createsuperuser
uv run python manage.py runserver 8200
uv run pytest
```
Import catalogo legacy:
```sh
uv run python manage.py import_db_structure --source local
uv run python manage.py load_defaults --only-level 4 --source local
```
Test mirati rilevanti:
```sh
uv run pytest tests/test_relational_database_operations.py -v
uv run pytest tests/test_ssh_tunnel_configuration.py -v
```
### Stack Docker legacy
ThothAI dichiara `postgres:16-alpine` nel profilo `internal-db`, con volume persistente e
healthcheck `pg_isready`.
```sh
docker compose --profile internal-db up --build
```
Il wrapper legacy abilita lo stesso profilo quando `POSTGRES_INTERNAL=true`:
```sh
POSTGRES_INTERNAL=true ./docker-up.sh
```
Questi comandi documentano il riferimento osservato; non sono comandi di installazione per
ThothII.
## Cosa copiare in ThothII
### Parità necessaria
- gerarchia Workspace Database → Table → Column;
- relazione strutturale fra colonne sorgente e destinazione;
- navigazione e filtri dipendenti workspace/database/tabella;
- test di connessione separato dal salvataggio;
- introspezione esplicita e ripetibile;
- descrizioni generate dall'AI ma modificabili dall'utente;
- validazione cross-entity delle relazioni;
- azioni di import/export senza segreti;
- stato leggibile dei job lunghi;
- PostgreSQL interno persistente con migrazioni esplicite;
- test di CRUD, cardinalità, cascade/restrict, isolamento per workspace e idempotenza.
### Parità semantica con `annotations.yaml`
Il modello futuro deve poter rappresentare almeno:
- descrizione, concetti e note per tabella;
- descrizione, sinonimi, concetti, evidence, note ed `eligible` per colonna;
- foreign key logiche;
- distinzione fra commento fisico osservato e descrizione curata;
- provenienza del contenuto importato o generato.
L'eventuale esclusione di uno di questi campi deve essere una decisione esplicita perché cambia
rendering, retrieval, LSH o SQL generation.
### Vincoli minimi da progettare
- `workspace_id` obbligatorio e unico sul Workspace Database, con esistenza validata contro il
catalogo YAML dal servizio applicativo;
- nome tabella unico nel database e schema appropriato;
- nome colonna unico nella tabella;
- relationship unica secondo il modello, anche per chiavi composite;
- estremi della relationship nello stesso Workspace Database;
- appartenenza certa della colonna alla tabella;
- mutazioni aggregate transazionali;
- gestione esplicita di concorrenza fra CRUD e introspezione.
## Cosa non copiare
- Django, Django Admin, Django ORM, DRF, template admin e frontend Next;
- modello Workspace legacy e condivisione dello stesso DB fra più workspace;
- password o passphrase come normali campi testuali;
- password incluse in CSV o export completi;
- token SSO inseriti nella query string;
- migrazioni generate automaticamente all'avvio;
- validazioni presenti soltanto nel form;
- `pk_field` e `fk_field` testuali come fonte di verità;
- duplicazione di tabella e colonna negli estremi senza constraint coerenti;
- introspezione additiva che non segnala rename, delete o drift;
- azioni admin che possono mostrare successo dopo output AI non valido;
- dipendenza del workflow core dalla disponibilità della UI o del PostgreSQL amministrativo.
## Aspetti di sicurezza da non ereditare
L'export legacy della struttura include username e password in chiaro. Il modello conserva inoltre
password, passphrase SSH e altri segreti in `CharField`; non è stata trovata cifratura applicativa,
nonostante un testo admin affermi il contrario.
ThothII distingue i metadati di connessione dai riferimenti al secret store cifrato. In ogni caso:
- nessun endpoint o export deve restituire segreti;
- log ed errori devono sanificare DSN e credenziali;
- le credenziali di migrazione non devono essere disponibili al runtime CRUD;
- il catalogo non deve riusare credenziali del DWH, delle sessioni o di Qdrant;
- test connessione e introspezione devono usare timeout e privilegi read-only.
La binding REST corrente richiede una verifica prima del cutover: il renderer emette
`ssl_ca_file`, mentre il modello Python espone `ssl_ca`; il percorso della CA privata potrebbe quindi
non essere consumato. PSD richiede TLS con CA privata in locale, perciò questo disallineamento deve
essere corretto e coperto da un test end-to-end prima di affidare il profilo REST al catalogo.
## Percorso incrementale
### Step 1: accesso alla superficie vuota
Implementato in questo worktree:
- pulsante `Database management` nella sidebar destra;
- visibilità legata a `workspace.manage`;
- superficie centrale React separata e vuota;
- nessun router, endpoint, fetch o stato catalogo;
- sessione e SSE conservati in background;
- ritorno al core tramite creazione, apertura o resume di una sessione;
- test frontend dedicati.
Comandi di verifica:
```sh
cd frontend
npx vitest run src/shell/AppShell.database-management.test.tsx
npx vitest run src/shell/AppShell.new-session.test.tsx \
src/shell/AppShell.session-target.test.tsx \
src/shell/AppShell.session-mgmt.test.tsx
npx tsc -b
```
### Step 2: contratto di dominio e schema relazionale
Progettazione della vertical slice completata: Workspace Database, Database Binding, singolo schema,
riferimenti al secret store, optimistic concurrency e capability per trasporto hanno contratti
espliciti. Configurazione e contenuti semantici restano mutabili; la struttura fisica osservata è
sincronizzata e non modificabile manualmente.
### Step 3: PostgreSQL interno e migrazioni
PostgreSQL interno con volume e ruoli runtime/migrator separati. Il modulo catalogo usa Kysely sopra
il driver `pg`; le migrazioni compilate vengono applicate soltanto dal comando `catalog:migrate` e
mai allo startup Fastify. Health, readiness e diagnostica restano dedicate; l'indisponibilità del
catalogo non cambia `core /health` e non interrompe una sessione.
### Step 4: API CRUD
Contratti HTTP, autorizzazione, paginazione, filtri, errori, optimistic concurrency e transazioni.
Gli endpoint dovranno vivere sotto un namespace catalogo e non riutilizzare le route sessione.
### Step 5: UI CRUD
Workspace Database, Catalog Table, Catalog Column e Catalog Relationship sono implementati con
React/Vite e il design system ThothII.
La navigazione è gerarchica e locale al database (`Overview | Tables`), senza menu o filtri globali
per tipo di entità. La grid delle tabelle non offre Add o cancellazione della singola configurazione;
le selezioni espongono invece la pulizia esplicita dei metadati. Il dettaglio full-width mantiene
immutabili i fatti fisici e consente di modificare separatamente Description e Generated
Description. Colonne e relazioni seguono la stessa gerarchia: Columns appartiene al dettaglio
della tabella, Relationships al database. I valori descrittivi null sono mostrati come celle e
campi vuoti, senza fallback visivi o placeholder `Not set` che nascondano quale sorgente è
effettivamente valorizzata.
Le griglie che dispongono di azioni massive usano checkbox e una toolbar contestuale con conteggio,
menu `Actions` e cancellazione della selezione. La selezione identifica ID espliciti, può essere
accumulata attraverso i filtri e viene azzerata dopo successo, nuova sincronizzazione o uscita
dalla pagina; un'azione è all-or-nothing se un elemento non è idoneo. I menu a livello database
espongono gli scope fisici come azioni distinte: `Synchronize tables`, `Synchronize relationships`
e `Synchronize all`. La grid Tables espone invece `Synchronize columns` per le tabelle selezionate;
la pagina Columns non espone sincronizzazione. Le selezioni database aggiungono `Delete all tables` e
`Delete all relationships`; le selezioni tabelle aggiungono `Delete all columns` e `Delete all
relationships`. Queste operazioni sono atomiche, richiedono conferma e non modificano database
esterno, binding, configurazione o segreti. Test connection resta un'azione distinta; griglie senza
azioni non mostrano controlli di selezione inerti.
### Step 6: introspezione
Catalog Table, Catalog Column e Catalog Relationship sono implementate per PostgreSQL diretto,
Thoth REST Connector e tunnel SSH. La scansione read-only è separata dalla transazione; una
riconciliazione atomica crea, aggiorna i commenti sorgente ed elimina, dopo conferma, i fatti fisici
assenti senza rendere modificabile manualmente la struttura osservata. Gli scope autorevoli sono
Tables per database e Physical Relationships per database. Per Columns, `tableIds` vuoto include
tutte le Catalog Table correnti, mentre una lista di ID limita lo scope al sottoinsieme esplicito;
`Synchronize all` osserva tutti e tre gli scope in un unico snapshot e li riconcilia insieme. Tutti
gli scope sono eseguiti come Catalog Sync Run durevoli in background, non attraverso implementazioni
sincrone e asincrone separate. Un run che prevede cancellazioni conserva il diff, attende una
conferma esplicita e verifica nuovamente lo snapshot prima dell'applicazione; se la sorgente è
cambiata, invalida la conferma. Ogni applicazione è atomica e fail-closed: errori, timeout o
capability non disponibili non producono aggiornamenti parziali.
PK e FK devono essere visibili sulle Catalog Column senza duplicare le stringhe denormalizzate di
ThothAI. La posizione nella primary key è un fatto osservato della colonna; membership e conteggio
FK sono proiezioni derivate dalle Catalog Relationship e dalle loro coppie ordinate, aggiornate
nella stessa transazione di riconciliazione.
Ogni scope registra la versione della Database Binding osservata e l'istante dell'ultima
sincronizzazione. Una modifica della binding conserva il catalogo precedente ma lo marca stale;
solo un `Synchronize all` riuscito rende nuovamente corrente l'intero schema.
### Step 7: generazione AI dei metadati
Generated Description è una proposta distinta e modificabile: un revisore può correggerla prima
di consolidarla esplicitamente come Description. Lo slice AI dovrà decidere e implementare anche
alias semantici, descrizioni dei valori, sinonimi e concetti per tabelle e colonne, oltre alla
gestione esplicita di errori e output non validi. La generazione AI e l'azione di consolidamento non
appartengono allo slice di introspezione dello schema.
### Step 8: migrazione PSD
Import idempotente delle annotations PSD, riconciliazione contro la struttura introspezionata,
report degli orfani e confronto semantico con il rendering corrente. Gli altri workspace non
ricevono import legacy.
### Step 9: sostituzione dell'input core
Rimuovere la dipendenza da `annotations.yaml` soltanto dopo avere un contratto equivalente,
test di rendering/search/Qdrant e una policy di disponibilità. Le sessioni di test esistenti
possono essere eliminate, ma le nuove sessioni non devono osservare aggiornamenti parziali.
Questo cutover è esplicitamente rinviato fino al completamento del database dei metadati. Il primo
gate successivo obbligatorio sarà valutare l'integrazione del Catalog Schema Snapshot con il
workflow core e lo schema-linking corrente; il rinvio non autorizza a dimenticare o assorbire
implicitamente il lavoro in altri slice.
### Step 10: operazioni e accettazione
Backup/restore reale, diagnostica, metriche, permessi definitivi, hardening degli export e
test di failure isolation fra catalogo e workflow. I Catalog Sync Run hanno un solo job attivo per
Workspace Database, sono concorrenti fra database diversi e usano un lock persistente. Un pannello
operativo non modale rimane visibile durante la navigazione del database, mostra fasi, contatori,
tempo trascorso e log sanitizzato via SSE con polling di fallback, e offre Confirm, Cancel e Retry
quando consentiti. Un restart marca `interrupted` i run rimasti attivi; il retry crea un nuovo run.
Le modifiche ai metadati restano consentite durante la scansione e sono preservate dall'applicazione.
Il worker gira inizialmente nello stesso servizio Fastify ma dietro un'interfaccia estraibile, con
coda, lease e heartbeat persistiti nel catalog-db. I riepiloghi dei run non scadono; gli eventi
dettagliati sono conservati per 30 giorni, mentre snapshot e diff completi vengono eliminati dopo
la conclusione lasciando conteggi, decisioni e una sintesi sanitizzata dell'esito.
## Verifiche del core da conservare per il cutover
Comandi attuali rilevanti:
```sh
tht --installation <absolute>/thothii-installation.yaml workspace preprocess dwh \
--workspace <id> --json
tht --installation <absolute>/thothii-installation.yaml workspace schema suggest-fks \
--workspace <id> --json
tht --installation <absolute>/thothii-installation.yaml workspace schema check \
--workspace <id> --json
tht --installation <absolute>/thothii-installation.yaml workspace schema accept \
--workspace <id> --run <run-id> --yes --json
tht --installation <absolute>/thothii-installation.yaml workspace index-schema \
--workspace <id> --json
```
Suite che documentano il comportamento da preservare:
```sh
cd backend
npx vitest run test/workspaces-git-annotations.test.ts \
test/registry-annotations.test.ts \
test/annotations-sync.test.ts \
test/workspace-runtime-config-lease.test.ts \
test/workspace-preprocessing-service.test.ts
npx tsc --noEmit -p .
cd ../harness
.venv/bin/pytest -q \
tests/test_annotations_root.py \
tests/test_schema_fk_annotations.py \
tests/test_mschema_render.py \
tests/test_qdrant_cli_commands.py
```
Questi test non implicano che la futura implementazione debba continuare a usare file YAML.
Definiscono gli effetti semantici e le guardie da mantenere o sostituire consapevolmente.
## Decisioni rinviate
Le seguenti scelte non appartengono allo step 1:
- lifecycle dei riferimenti ai segreti durante sostituzione e cancellazione;
- criteri per aggiungere dialetti successivi a PostgreSQL;
- criteri per un'eventuale estensione futura a più schemi per database;
- lifecycle e gestione amministrativa delle future Logical Relationship;
- alias semantici, descrizioni dei valori, sinonimi e concetti prodotti o assistiti dall'AI;
- formato e momento del cutover dal file al database interno;
- permission definitiva separata da `workspace.manage`.
Ognuna sarà affrontata nel relativo step, senza anticipare scelte tecnologiche nel presente
documento.
@@ -1,251 +0,0 @@
# AI-generated descriptions for Catalog Tables and Catalog Columns
## Problem Statement
ThothII already stores a Generated Description separately from the curated Description for Catalog
Tables and Catalog Columns, but administrators cannot populate it with AI. ThothAI provides the
useful core workflow—generate table and column comments from schema context and small real-data
samples—but its execution, configuration, and interaction model cannot be copied directly into
ThothII.
Administrators need an asynchronous workflow integrated into Database Management. They must be
able to choose an installation-approved model, generate descriptions for selected or missing
targets, observe understandable progress, stop or recover a stuck operation, review generated
text, and explicitly consolidate it. The solution must retain ThothAI's practical simplicity and
must not introduce a general job platform, model gateway, distributed scheduler, or competing
user-facing CLI.
## Solution
Add Description Generation to Database Management as one installation-wide, sequential background
run owned by the Fastify backend. The browser starts a run and remains responsive while the backend
processes bounded requests one at a time. Each completion is delegated to a short-lived internal
Python helper using LiteLLM. Models, their default, and any API-key secret references are declared in
application setup YAML and are independent of both workspaces and Pi configuration.
Each valid result is written immediately to the target's Generated Description. A minimal run row
and ordered text events provide status, counters, history, and a live log. A stopped or crashed run
is not resumed automatically; completed results remain in place and Generate Missing supplies the
simple recovery path. An Unlock action marks a stale recorded run interrupted only when no helper
or backend generation loop is alive.
Prompts use catalog context and, when available, no more than five real rows and five representative
non-null examples. Samples are transient and never logged or persisted. A valid inability to infer
a description produces a standard application-localized value such as `Non generabile`; provider,
timeout, and response-validation failures remain technical errors.
Generated text remains separate from Description until an administrator uses the existing
checkbox selection and Actions control to consolidate it. Consolidation retains Generated
Description and never writes comments to the external Workspace Database.
## User Stories
1. As an installation operator, I want to declare the models allowed for metadata generation in setup YAML, so that model availability is controlled centrally.
2. As an installation operator, I want to declare one default metadata-generation model, so that administrators begin with a safe operational choice.
3. As an installation operator, I want each model to reference its own API-key secret, so that credentials are not stored in workspaces or browser-visible settings.
4. As an installation operator, I want metadata-generation models to remain independent of Pi models, so that changing this workflow cannot disrupt the core NL-to-SQL experience.
5. As an installation operator, I want invalid model setup to fail validation clearly, so that the application does not start with ambiguous provider behavior.
6. As a Catalog Administrator, I want generation controls to explain when no model is configured, so that I know why the action is unavailable.
7. As a Catalog Administrator, I want to select an approved model from a selector initialized to the setup default, so that I control which model performs the work.
8. As a Catalog Administrator, I want to generate descriptions for selected Catalog Tables, so that I can work on a focused part of the catalog.
9. As a Catalog Administrator, I want to generate descriptions for selected Catalog Columns, so that I can work on individual fields without regenerating a whole table.
10. As a Catalog Administrator, I want to generate all eligible descriptions for a Workspace Database, so that I can initialize a catalog in one operation.
11. As a Catalog Administrator, I want to generate only missing descriptions, so that I can continue interrupted work without replacing completed proposals.
12. As a Catalog Administrator, I want a full run to process Catalog Columns before their Catalog Tables, so that table descriptions can benefit from column descriptions.
13. As a Catalog Administrator, I want generation to run asynchronously after I start it, so that the browser remains usable and progress is not tied to one HTTP request.
14. As a Catalog Administrator, I want only one Description Generation Run active in the installation, so that provider traffic and operational behavior remain predictable.
15. As a Catalog Administrator, I want a second start attempt to return a clear conflict, so that I cannot accidentally overlap generation runs.
16. As a Catalog Administrator, I want synchronization, cleanup, consolidation, and edits for the target Workspace Database blocked during generation, so that the simple sequential run sees stable catalog state.
17. As a Catalog Administrator, I want to see the run's model, scope, status, counters, and timestamps, so that I understand what is happening.
18. As a Catalog Administrator, I want a chronological text log, so that I can follow completed targets and diagnose errors.
19. As a Catalog Administrator, I want live log updates with a polling fallback, so that temporary SSE problems do not hide run progress.
20. As a Catalog Administrator, I want completed and interrupted runs to remain inspectable, so that I can understand prior activity.
21. As a Catalog Administrator, I want to stop an active run, so that I can halt an incorrect or unexpectedly costly operation.
22. As a Catalog Administrator, I want stopping a run to terminate its current model helper and prevent later targets from starting, so that stop has prompt operational effect.
23. As a Catalog Administrator, I want valid results completed before a stop or failure to remain saved, so that useful work is not discarded.
24. As a Catalog Administrator, I want a run left active by a backend restart to become interrupted, so that the UI does not claim nonexistent work is still running.
25. As a Catalog Administrator, I want to unlock a stale active run when no generation process is alive, so that an erroneous recorded lock cannot block future work.
26. As a Catalog Administrator, I want Unlock rejected while a live generation process exists, so that recovery cannot create an overlapping run.
27. As a Catalog Administrator, I want Generate Missing to continue after interruption, so that recovery does not require a special resume mechanism.
28. As a Catalog Administrator, I want one retry for a transient model failure, so that a brief provider fault does not immediately lose a batch.
29. As a Catalog Administrator, I want the run to fail after three consecutive technical failures, so that a broken provider does not generate an unbounded stream of attempts.
30. As a Catalog Administrator, I want a successful request to reset the consecutive-failure count, so that isolated errors do not prematurely stop a useful run.
31. As a Catalog Administrator, I want a completed-with-errors result when isolated batches fail but the run reaches its end, so that partial problems remain visible.
32. As a Catalog Administrator, I want no automatic fallback to a different model, so that the selected model remains truthful and predictable.
33. As a Catalog Administrator, I want malformed or ambiguous model output rejected without writing it, so that descriptions cannot be assigned to the wrong target.
34. As a Catalog Administrator, I want an inability to infer a description represented by standard localized text, so that every valid outcome is understandable in the workspace language.
35. As a Catalog Administrator, I want technical failures kept distinct from non-generatable outcomes, so that provider problems are not mistaken for catalog knowledge.
36. As a Catalog Administrator, I want generated prose written in the workspace language, so that it matches the catalog's intended audience.
37. As a Catalog Administrator, I want generated text stored separately from curated Description, so that AI output remains a reviewable proposal.
38. As a Catalog Administrator, I want to edit a Generated Description manually, so that I can improve a proposal before consolidation.
39. As a Catalog Administrator, I want to select one or more tables or columns and run “Move generated description to Description” from the existing Actions control, so that review remains integrated into the current grids.
40. As a Catalog Administrator, I want consolidation to retain the Generated Description, so that I can still see the proposal from which the curated text was copied.
41. As a Catalog Administrator, I want selected records without a Generated Description skipped and reported, so that the bulk action does not erase curated text.
42. As a Catalog Administrator, I want consolidation and generation to modify only the Metadata Catalog, so that no external database comment is changed.
43. As a Catalog Administrator, I want prompts to use schema facts and existing catalog text, so that generated descriptions are grounded in available metadata.
44. As a Catalog Administrator, I want prompts to use at most five real source rows and five representative values when available, so that the model has useful examples without unbounded disclosure.
45. As a Catalog Administrator, I want to be warned that real source samples are sent to the selected provider, so that I can make an informed disclosure decision.
46. As a Catalog Administrator, I want sampled rows and values excluded from persistence and logs, so that operational history does not become a secondary data store.
47. As a security operator, I want API keys, prompts, samples, and complete provider payloads redacted from logs, so that diagnostics do not leak secrets or source data.
48. As a support operator, I want concise per-target and per-batch event messages, so that failures can be diagnosed without provider-specific internals.
49. As an authorized administrator, I want all generation, cancellation, unlock, and consolidation actions protected by database-management permission, so that ordinary users cannot mutate catalog metadata.
50. As an unauthorized user, I want generation controls hidden or disabled and API calls rejected, so that frontend visibility is not treated as authorization.
51. As an operator, I want setup changes to take effect after an application restart, so that configuration lifecycle remains simple and explicit.
52. As a product owner, I want the first release to avoid queues, parallel calls, distributed locks, and automatic resume, so that effort remains focused on generating and reviewing useful descriptions.
## Implementation Decisions
- The Fastify backend owns one installation-wide Description Generation Run and its sequential
processing loop. It does not delegate lifecycle ownership to Pi or Python.
- A Description Generation Run has one of `queued`, `running`, `completed`,
`completed_with_errors`, `cancelled`, `failed`, or `interrupted`. It stores the Workspace
Database, requested scope, selected model identifier, workspace language, progress counters,
timestamps, and an optional final error summary.
- Ordered Description Generation Events store timestamp, severity, and safe human-readable text.
When an event identifies a target, it uses the object type and qualified physical name, such as
`Column "patients.birth_date"` or `Table "patients"`; catalog UUIDs remain internal identifiers.
No durable per-target jobs, model invocation rows, prompt snapshots, sample snapshots, leases,
heartbeats, registry revisions, or provenance chains are introduced.
- Starting a run schedules an in-process background loop and returns the run immediately. The API
exposes start, run/history lookup, event listing and streaming, cancellation, stale-run unlock,
and the safe list of configured model choices. There are no retry-item or resume endpoints.
- One in-memory generation manager enforces the installation-wide active-run rule. The existing
Catalog Operation Coordinator reserves the target Workspace Database for the duration of the
run, without being generalized into a new operation framework.
- On backend startup, persisted `queued` or `running` Description Generation Runs become
`interrupted`. The application performs no automatic replay or resume.
- Unlock succeeds only when no live generation loop or helper child exists. It marks the stale run
interrupted and releases the local reservation; it is not a distributed lock recovery protocol.
- Each model completion uses a short-lived Python helper backed by LiteLLM. Structured input is
supplied over stdin, structured output alone is emitted on stdout, diagnostics use stderr, and
the helper can be terminated by cancellation.
- A completion request contains no more than ten targets. Requests run one at a time. The helper
performs at most one retry for a transient technical failure.
- Three consecutive model-request failures fail the run. A successful request resets that count.
Isolated exhausted failures may be logged and skipped, producing `completed_with_errors` if the
run later reaches its end.
- Every valid generated or non-generatable result is applied immediately to Generated Description.
Earlier writes are retained after cancellation, interruption, or later failure.
- A response must identify requested targets unambiguously and classify each returned result as
generated or non-generatable. Duplicate, unknown, missing, or malformed mappings cause a
technical request failure and no result from that ambiguous response is applied.
- The parser also tolerates one JSON object enclosed by one complete `json` code fence, because
some supported models add that formatting despite the prompt. Any prose outside the fence,
multiple payloads, or malformed/ambiguous mappings remain invalid.
- The application supplies localized standard non-generatable text. Provider wording is not used
as the standard value, and technical errors never write that value.
- A full-database run generates eligible Catalog Columns before Catalog Tables. Generate Missing
excludes targets whose Generated Description is already non-empty; all-generation may replace
existing generated proposals only after the initiating action makes that scope explicit.
- Model choices are declared under a metadata-generation section in installation setup YAML. Each
choice has a stable identifier, display label, LiteLLM provider/model settings, optional endpoint
settings, and an optional environment-secret reference for its API key. The reference may be
omitted only when an explicit endpoint is configured for unauthenticated access. One identifier
is the default.
- An explicit endpoint may set `disableThinking: true`; the helper translates it only to the
Qwen-compatible chat-template switch needed to keep the response within the strict JSON contract.
- Metadata-generation setup is separate from application settings for Pi and from workspace
`llm_policy`. Raw keys never enter setup YAML, the catalog database, API responses, process
arguments, or event text. Configuration reload is restart-only.
- If setup defines no usable model, the safe model-list response is empty and the UI disables
generation with an explanation. The backend still rejects direct generation attempts.
- Prompt construction treats schema names, comments, descriptions, and values as untrusted data.
It requests output in the workspace language and separates instructions from catalog content.
- A request may contain up to five real source rows and up to five representative distinct,
non-null values for relevant columns. Inputs are bounded before prompt construction and are not
persisted or logged.
- The UI discloses that real data can be sent to the selected provider. A future Sensitive Data
Policy will classify values and exclude or anonymize protected data; that policy is not silently
approximated in this slice.
- The generation UI reuses Database Management's table and column selections, model selector,
Actions control, run drawer conventions, SSE delivery, and polling fallback where practical.
Visual parity with Catalog Sync Run logs is not required.
- The consolidation action copies each selected, non-empty Generated Description into Description
in a catalog transaction, retains Generated Description, skips empty proposals, and reports
copied and skipped counts. It never writes to the external Workspace Database.
- Generation, cancellation, unlock, and consolidation require the existing database-management
permission and are validated by the backend independently of UI state.
- No user-facing generation CLI is added. The Python process is an internal completion adapter,
not an operator surface or a long-lived service.
## Testing Decisions
- Tests assert externally observable behavior rather than private loop structure, process timing,
or LiteLLM implementation details.
- The primary and highest test seam is the Fastify catalog API with a test PostgreSQL catalog and
an injected fake Model Completer. It verifies complete paths through authorization, run
persistence, sequential processing, event delivery, Generated Description updates, and final
status without contacting a real provider.
- API tests cover each generation scope, column-before-table order, the ten-target request bound,
model validation, one-active-run conflict, target-database exclusion, cancellation, startup
interruption, Unlock safeguards, Generate Missing, partial success, consecutive failure
handling, non-generatable localization, malformed responses, redacted events, and permissions.
- Catalog repository integration tests verify the migration, run and event ordering, active-run
constraint, immediate description writes, history queries, startup interruption, and bulk
consolidation behavior against PostgreSQL.
- The Python helper has a small black-box contract suite using a simulated LiteLLM adapter. It
verifies stdin/stdout framing, pristine stdout, stderr diagnostics, normalized success and
failure output, one transient retry, secret redaction, and termination behavior.
- Setup-validation tests cover duplicate model identifiers, missing or unknown defaults, malformed
provider settings, missing secret references, safe public model projection, and strict separation
from Pi and workspace model settings.
- Database Management tests use the existing browser-level component seam with MSW. They verify
model selection and default, selected/all/missing actions, disabled state without models, running
progress and logs, polling recovery, cancellation, Unlock visibility, terminal summaries,
generated-text refresh, and selected consolidation with copied/skipped counts.
- Existing Catalog Sync Run route, repository, SSE, and drawer tests are prior art for asynchronous
status and event behavior. Existing catalog table/column editing and Database Management tests
are prior art for optimistic catalog updates, permissions, selection, and action controls.
- One required manual acceptance gate, outside deterministic CI, uses the installation's configured
default model and a disposable PostgreSQL database containing only invented data. Its application
credentials are read-only. It generates Italian text for one Catalog Column and one Catalog
Table, verifies their Generated Description, inspects the safe activity log, confirms that no key
or sample value is exposed, and consolidates one selected result. If the configured secret is not
available, acceptance stops without exposing or requesting the key in conversation.
- Delivery includes a strict MkDocs build executed through repository-managed, reproducible
documentation dependencies rather than globally installed Python packages. A readable direct
dependency file is retained, a complete transitive lock is generated with `uv`, and one canonical
repository command performs the strict build from that lock.
- Successful real-provider acceptance is recorded in a short sanitized report under
`docs/testing/`. It identifies the model and checks performed but contains no credentials,
prompts, source samples, complete provider payloads, or generated database values.
- No tests are added for worker queues, parallel generation, distributed locking, multi-replica
recovery, automatic resume, cost accounting, or model fallback because those behaviors are out
of scope.
## Out of Scope
- Reusing Pi to execute Description Generation or changing Pi's model configuration.
- A shared Installation Model Registry, model gateway, long-lived Python sidecar, or provider
management platform.
- A user-facing generation CLI.
- Parallel model calls, worker queues, adaptive rate limiting, distributed locks, leases,
heartbeats, automatic resume, or multi-replica execution.
- Durable target jobs, invocation history, prompts, samples, token usage, cost accounting,
provenance chains, target snapshots, or advanced retention controls.
- Automatic retry or resume of individual targets beyond one technical helper retry and a new
Generate Missing run.
- Automatic fallback to a different model.
- Writing generated text into comments of the external Workspace Database.
- Generating logical relationships or other catalog metadata beyond Catalog Table and Catalog
Column descriptions.
- Implementing the Sensitive Data Policy. Its definition and exclusion/anonymization behavior are
a required follow-up improvement.
- Generalizing the log viewer across unrelated metadata operations. That broader concern remains
related to Gitea issue #2.
## Further Notes
- The design deliberately follows ThothAI's proven simple workflow while adapting it to ThothII's
asynchronous browser interaction, setup ownership, and existing Generated Description model.
- The source-sampling disclosure is a release requirement, not merely documentation for operators.
- `completed_with_errors` is reserved for a run that reaches the end after isolated technical
failures. Three consecutive failures end the run as `failed`.
- Successful values are their own recovery record: after interruption, Generate Missing naturally
skips them without needing replay state.
- The implementation is available without a feature flag once the catalog migration and valid
setup are present. With no configured model, the feature remains visibly unavailable rather than
partially initialized.
- The final manual gate is intentionally narrow: one real-provider run covers one Catalog Column
and one Catalog Table, generated Italian text, safe events, and one consolidation. Automated
tests remain the evidence for All, Missing, Stop, restart interruption, Unlock, and failure paths.
@@ -1,173 +0,0 @@
# AI catalog description generation
Status: simplified design, API, persistence, test seams, and delivery tickets accepted.
## Objective
Bring ThothAI's useful AI comment-generation workflow into the ThothII Metadata Catalog without
turning it into a general job platform. Administrators can generate editable descriptions for
catalog tables and columns, inspect progress, stop a run, recover a stale run, and explicitly copy
approved generated text into the curated Description field.
The implementation is UI/API only. There is no user-facing generation command.
## ThothAI behavior retained
- Generate descriptions for selected tables, selected columns, missing descriptions, or all
eligible targets.
- Generate columns before their containing table when running the full workflow, so table prompts
can benefit from the resulting column descriptions.
- Process bounded batches of at most ten targets, one model request at a time.
- Include schema context, existing catalog text, up to five real source rows, and up to five
representative non-null values when available.
- Keep generated text separate from the curated Description until an administrator consolidates
it.
- Use the existing table and column checkboxes plus the Actions selector to copy Generated
Description into Description for one or more selected records. The generated value is retained.
- Store a localized standard value such as `Non generabile` when a valid model response says that
a description cannot be inferred.
Unlike ThothAI, every generation action is asynchronous from the browser's perspective and exposes
a persistent, readable activity log.
## Minimal architecture
The Fastify backend owns the run lifecycle and sequential loop. It starts one short-lived Python
helper for each model completion. The helper uses LiteLLM, accepts structured input on stdin,
returns structured output on stdout, and writes diagnostics only to stderr.
This is preferred over reusing Pi. Pi remains the interactive NL-to-SQL orchestration surface,
whereas description generation is a bounded batch transformation with no conversational state or
human gate. A LiteLLM helper avoids inventing a Pi session protocol for a task that needs one
request and one structured response.
There is no Python daemon, model gateway, queue service, worker pool, or generation CLI. Python is
already a core implementation language in ThothII's harness and core image; this helper does not
introduce a new runtime family.
## Run lifecycle and exclusion
- At most one Description Generation Run may be queued or running in the installation.
- Start returns immediately after creating the run and scheduling the in-process backend loop.
- Requests are sequential; there is no parallel provider traffic.
- The target Workspace Database is reserved through the existing in-memory catalog-operation
coordinator. Synchronization, cleanup, consolidation, and direct catalog edits for that database
are rejected while generation is active.
- A second generation start is rejected with a conflict response.
- Stop terminates the current helper process, stops further targets, and marks the run cancelled.
- Backend startup marks any queued or running generation row interrupted. It does not resume work.
- Generate Missing is the normal manual continuation mechanism because successful values were
already saved.
- Unlock is available only when the backend has no live generation process; it marks a stale
recorded run interrupted and clears the local reservation.
This is intentionally a single-process policy. Multi-replica coordination is out of scope.
## Persistence
Persist only:
- a Description Generation Run with database, scope, selected model, language, status, counters,
timestamps, and an optional final error summary;
- ordered Description Generation Events containing timestamp, level, and human-readable text;
- each successful or non-generatable result directly in the target's Generated Description.
Do not add per-target job rows, invocation history, prompt or sample snapshots, provider cost
accounting, leases, heartbeats, registry revisions, or generated-description provenance. The event
log is operational evidence, not a replay mechanism.
## Model setup
Selectable models and their default belong to application setup YAML, not to a workspace. Each
entry supplies a stable display identifier, LiteLLM provider/model information, optional endpoint
settings, and—unless that explicit endpoint is unauthenticated—a reference to an installation
secret containing the API key. Keyless entries without an explicit endpoint are invalid. Raw keys must not be
stored in the YAML, database, frontend, events, or process arguments.
An explicit endpoint may opt into `disableThinking: true` when its Qwen-compatible chat template
would otherwise place reasoning text around the required JSON result.
This metadata-generation configuration is independent of the existing Pi provider/model settings
and workspace `llm_policy`. A setup change takes effect after application restart. If no model is
configured, generation controls are disabled with an explanatory message.
The browser receives only the selectable identifiers and labels. The selected value defaults to
the setup default and is validated again by the backend when a run starts.
## Prompt inputs and outputs
Targets are grouped in model requests of at most ten. Prompts distinguish instructions from
untrusted schema names, comments, descriptions, and sampled values. A response must map every
returned result to a requested target and classify it as generated or non-generatable. Missing,
duplicate, unknown, or malformed target results make that request a technical failure rather than
silently writing ambiguous text.
One complete `json` code fence around the object is tolerated for model compatibility; prose
outside it, multiple payloads, and ambiguous mappings are still rejected.
For a complete database run, eligible columns are processed before tables. A table request can use
the current Generated Description or Description of its columns. The output language is the
workspace language; the standard non-generatable text is localized by the application rather than
trusted to arbitrary model wording.
Up to five source rows and five representative examples may be sent to the provider and are never
persisted. Delivery must call out this disclosure. A follow-up Sensitive Data Policy will define
which values are excluded or anonymized.
## Errors, retry, and logs
The helper performs at most one retry for a transient technical provider failure. A final failed
request produces an error event and increments the consecutive-error count. The run stops as
failed after three consecutive technical failures; any successful request resets the count. There
is no automatic fallback to another model.
Valid non-generatable outcomes are results, not technical errors. Successful results from earlier
requests remain stored when a later request fails or the run is stopped.
The UI shows status, counters, selected model, start/end times, and a chronological text log. Live
delivery may reuse the existing SSE infrastructure with polling as fallback; exact visual parity
with synchronization logs is not required. Logs must not contain API keys, prompts, source sample
values, or full provider payloads.
## Explicitly deferred complexity
- shared model registry or cutover of Pi configuration;
- long-lived Python sidecar or internal HTTP model gateway;
- generic catalog-operation kernel;
- durable target items, invocation records, target snapshots, or provenance chains;
- distributed locks, leases, heartbeats, worker queues, automatic resume, or multi-replica support;
- parallel calls, adaptive rate limiting, cost estimation, advanced metrics, or model fallback;
- user-facing generation CLI;
- automatic writeback to comments in the external database;
- Sensitive Data Policy implementation, which remains a required improvement after this slice.
## Delivery tracking
The accepted specification is Gitea issue #4 and the implementation is split into issues #5–#11.
Each ticket is a bounded vertical slice with explicit Gitea dependencies. Implementation proceeds
from the unblocked frontier, using a fresh subagent context for each ticket; integration and final
verification remain centralized so later slices cannot silently reopen the deferred platform
features above.
Issue #4 remains open until delivery completes four final gates: the stale Compose service-set
contract is corrected in its own commit; documentation dependencies are repository-managed and a
strict MkDocs build passes; one narrow real-provider acceptance run succeeds against non-sensitive
test data; and a separate, non-blocking Sensitive Data Policy design ticket is linked as required
follow-up work.
The documentation toolchain retains a readable direct-dependency input, adds a complete lock
generated with `uv`, and exposes one canonical strict-build command. Real-provider acceptance uses
the installation's configured default model and a disposable PostgreSQL database seeded only with
invented values and accessed read-only by the application. A missing protected model secret stops
the gate without disclosing it. The successful gate is captured in a sanitized report under
`docs/testing/` without prompts, samples, full generated values, payloads, or credentials.
Delivery is organized as four reviewable commits: the stale Compose contract correction, the
reproducible documentation toolchain, the AI-description feature, and—only after acceptance—the
sanitized acceptance report. A failed real-provider gate does not invalidate already verified
commits, but issue #4 remains open and no acceptance report claims success. Application defects are
fixed and reverified; missing configuration or provider unavailability is recorded and retried.
After every gate passes, the existing `codex/db-management` branch is pushed to its configured
origin without introducing a new pull-request workflow, then issue #4 is closed with links to the
delivery evidence. The separate Sensitive Data Policy issue is created as non-blocking follow-up,
linked to #4, and labeled `enhancement` plus `ready-for-human` because its design requires a future
`grill-with-docs` before agent implementation.
@@ -1,248 +0,0 @@
# Installation Model Catalog
Status: implemented on 2026-09-02.
## Outcome
`thothii-installation.yaml` is the only operator-authored source for models used by interactive
sessions, metadata generation, and embedding. Runtime-specific files are deterministic projections,
not additional configuration sources. Workspace descriptors contain database and Evidence concerns
and no model, provider, allowlist, default, embedding, or vector-store configuration.
This design does not merge execution lifecycles. Pi continues to run interactive sessions, the
short-lived LiteLLM helper continues to perform metadata generation, and the internal Ollama service
continues to provide embeddings. They share model declaration, not execution machinery.
## Canonical installation shape
The following example covers all currently required cases: a Pi built-in model, an authenticated
custom endpoint, a keyless internal endpoint, metadata generation, and the single embedding model.
```yaml
schemaVersion: 2
profile: server
projectDirectory: /srv/thothii
envFile: /srv/thothii/operator.env
workspaceRepository:
remote: git@git.example.com:organization/workspaces.git
branch: main
access: ssh
modelCatalog:
defaults:
session: zai/glm-5.3
metadataGeneration: local-qwen/qwen3.6-35b-a3b
embedding:
id: ollama/qwen3-embedding:0.6b
dimensions: 1024
providers:
deepseek:
authentication:
mode: pi_auth
session:
mode: pi_builtin
models:
deepseek-v4-pro:
session: {}
deepseek-v4-flash:
session: {}
zai:
endpoint:
baseUrl: https://api.z.ai/api/coding/paas/v4
authentication:
mode: secret_env
apiKeyEnv: ZAI_API_KEY
session:
mode: openai_compatible
metadataGeneration:
litellmProvider: openai
models:
glm-5.3:
label: GLM-5.3
session:
reasoning: true
contextWindow: 200000
maxTokens: 131072
metadataGeneration: {}
local-qwen:
endpoint:
baseUrl: https://ml-aritmolab.policlinicosandonato.it/v1
authentication:
mode: none
session:
mode: openai_compatible
metadataGeneration:
litellmProvider: openai
models:
qwen3.6-35b-a3b:
label: Qwen3.6 35B A3B
session:
reasoning: false
contextWindow: 131072
maxTokens: 16384
compatibility:
supportsDeveloperRole: false
supportsReasoningEffort: false
supportsStore: false
maxTokensField: max_tokens
metadataGeneration:
disableThinking: true
authentication:
configDirectory: /srv/thothii/auth-canonical
runtimeProjection:
directory: /srv/thothii/auth-runtime
uid: 10001
gid: 10001
```
The catalog uses maps instead of repeated IDs. The canonical identity of a model is always derived
as `<provider-key>/<model-key>`. `upstreamModel` may be added to a model only when the endpoint uses
a different identifier. `label` is optional and falls back to the canonical identity.
Model eligibility is not repeated in an `usages` array. A `session` block makes the model eligible
for sessions; a `metadataGeneration` block makes it eligible for metadata generation. The embedding
is a single required installation value rather than a list plus default.
## Provider and authentication rules
A provider owns one endpoint, one authentication mode, and zero or one adapter for each runtime.
Model entries cannot override provider endpoint or credentials. If the same upstream service needs
different endpoints or credentials, the installation declares two provider identities.
Supported session modes are intentionally closed:
- `pi_builtin`: Pi already owns the model's technical descriptor; the model's `session` block is
empty and ThothII does not copy context-window or compatibility facts.
- `openai_compatible`: ThothII generates a Pi custom-provider descriptor; each session model supplies
the technical values required by Pi.
Metadata generation uses the provider-level `litellmProvider`. A model-level
`metadataGeneration.disableThinking: true` is permitted only for an explicit compatible endpoint.
There is no generic adapter or plugin abstraction in schema version 2.
Exactly one provider authentication mode is allowed:
- `secret_env` requires an approved API-key environment reference present in the protected secret
bundle. Secret values never enter YAML, generated files, logs, arguments, or API responses.
- `pi_auth` is valid only for session-only `pi_builtin` providers and resolves through Pi's protected
authentication projection.
- `none` is valid only for an explicit endpoint. Runtime projections may supply a fixed non-secret
compatibility placeholder when a client library requires a non-empty key.
## Defaults and selections
`defaults.session` and `embedding` are required. `defaults.metadataGeneration` is required exactly
when at least one model has a `metadataGeneration` block; metadata generation may otherwise be
absent and its UI controls are disabled.
`modelCatalog.defaults.session` is the only configured session-model default. `PI_PROVIDER`,
`PI_MODEL`, and provider/model fields in installation-default settings are removed. A user choice is
a Model Selection containing only the canonical model identity and runtime controls such as thinking
level. A session manifest pins the selected canonical identity.
Removing the currently selected model causes new-session selection to fall back to the catalog
default with an explicit administrative warning. An existing session is never silently moved to a
different model; resume fails with `model_unavailable` when its pinned identity can no longer be
resolved.
## Generated runtime projections
Before Compose starts, `tht` strictly validates schema version 2 and generates installation-local
artifacts below `deploy/<installation-id>/generated/`:
- a normalized catalog JSON consumed defensively by the backend;
- Pi `models.json` for custom providers;
- Pi `settings.json`, combining fixed product settings with the session-eligible canonical IDs;
- a Compose override that mounts the projections and supplies embedding identity and dimensions to
core, preprocessing, and `embedding-model-init`.
Generation is deterministic and published only after every candidate artifact validates. A failed
generation aborts start before Compose is invoked. `tht doctor` recomputes expected bytes and reports
differences; no digest manifest or separate apply command exists. When projection bytes change,
`tht start` recreates the affected services so they cannot continue with an older bind mount.
Pi-only restart, update, and rollback operations reject projection drift and direct the operator to
`tht start`, because applying only the core-facing files could leave embedding services stale.
Generated projections are not backed up. Restore validates the canonical installation descriptor,
regenerates every projection, and only then starts services. Base Compose files and `operator.env`
must contain no model identities, defaults, endpoints, or dimensions.
## Workspace schema v4
Workspace schema v4 removes both top-level `llm_policy` and `semantic_index`. The entire latter
block is redundant today: its engine and distance are product constants, its collection duplicates
the workspace ID, and its model and dimensions are installation facts.
The runtime derives:
- Qdrant collection identity from the workspace ID;
- engine and distance from the supported product contract;
- embedding identity and dimensions from the Installation Model Catalog.
The published index generation records the canonical embedding identity and dimensions that created
it. A mismatch makes the index explicitly incompatible and requires operator-triggered
preprocessing. No existing index is deleted or rebuilt automatically.
The v3-to-v4 workspace migration is deterministic: set `workspace.schema_version` to `4`, remove
`llm_policy`, and remove `semantic_index`. It does not alter database, Evidence, diagnostics, or
binding data.
## Installation migration
Legacy installation migration must inspect all three former sources:
1. `metadataGeneration` in `thothii-installation.yaml`;
2. `deploy/pi/models.json`;
3. `deploy/pi/settings.json`.
The migrator emits a version-2 candidate only when it can reconcile identities, endpoints,
credentials, and runtime-specific facts without guessing. Ambiguous aliases such as `glm-53`,
`zai/glm-5.3`, and `openai/glm-5.3` are not silently equated. A conflict produces a field-level
report and leaves every input unchanged for operator resolution.
After migration, the strict loader rejects `metadataGeneration`, workspace `llm_policy`, workspace
`semantic_index`, legacy Pi source files, unknown fields, duplicate YAML keys, invalid defaults, and
incompatible authentication/adapter combinations with an actionable `migration_required` or
validation error.
## Final simplicity audit
The accepted design removes every configuration duplication that can be removed without inference:
- one authored installation file instead of an installation block plus two Pi files;
- one canonical `provider/model` identity instead of display IDs and runtime IDs;
- per-use blocks instead of a duplicated usages list;
- one embedding entry instead of a selectable embedding catalog;
- one catalog session default instead of environment and settings defaults;
- no model or vector-store fields in workspace descriptors;
- provider-level credentials instead of per-model credentials;
- no generic runtime-plugin abstraction;
- no persisted digest, apply command, or backup of generated projections.
The remaining generated files are necessary boundary adapters, not configuration concepts. Making
the backend parse the authoring YAML independently would remove one file but restore two semantic
validators. Hard-coding embedding values in Compose would remove one projection but restore a model
source outside the catalog. Inferring authentication from missing fields would save one YAML key but
turn a safe explicit choice into ambiguity. These apparent simplifications are therefore rejected.
No further reduction was found that preserves one authority, strict validation, explicit security,
session determinism, and model-free workspaces.
## Implementation surface
Implementation must update the host `tht` installation loader, setup and lifecycle projection,
doctor, backup/restore, Compose mounts and embedding inputs, backend catalog/settings/session model
resolution, workspace schema and migration, runtime rendering and diagnostics, frontend workspace
drafts and model filtering, examples, fixtures, and documentation. Existing session manifests remain
readable and keep their pinned provider/model identity; only resume resolution changes to the new
catalog.
Implementation completed after explicit approval. The installation schema, deterministic runtime
projections, migration path, model-free workspace schema v4, backend consumers, operator UI,
fixtures, and documentation now enforce this contract.
+3 -3
View File
@@ -5,9 +5,9 @@ manuale con commit/push dell'operatore accettati per la release 0;
restanti semplificazioni confermate, con chiarimenti su impatto core e controllo Git;
E1, E2, E3 e il collegamento ai gate X1 implementati il 2026-09-09.
Il contratto X1 è in [Session corrections](../contracts/archive-repair.md). Risultati e limiti in
[E1 — validazione](2026-09-09-evidence-e1-validation.md) e
[E2 — validazione](2026-09-09-evidence-e2-validation.md) e
[E3 — validazione](2026-09-09-evidence-e3-validation.md).
[E1 — validazione](../reports/knowledge-archives-release.md) e
[E2 — validazione](../reports/knowledge-archives-release.md) e
[E3 — validazione](../reports/knowledge-archives-release.md).
La [revisione di semplicità](2026-09-08-memory-evidence-simplification-review.md)
registra il riesame di Q1–Q15. Il flusso principale è draft scritta dallo specialista
@@ -2,7 +2,7 @@
Data del piano: 2026-09-08. Aggiornamento 2026-09-09: M1–M3, E1–E3 e X1
implementati. Risultati e limiti della verifica finale sono raccolti nel
[rapporto X1](2026-09-09-archive-repair-x1-validation.md).
[rapporto X1](../reports/knowledge-archives-release.md).
La [revisione di semplicità](2026-09-08-memory-evidence-simplification-review.md)
riesamina tutte le decisioni Q1–Q15 alla luce degli ultimi chiarimenti del
+1 -1
View File
@@ -6,7 +6,7 @@ confini di test confermati dal proprietario il 2026-09-08.
Primo incremento del progetto Memory management. Attua le decisioni già approvate
nel piano del 2026-09-08 e nell'ADR 0018. Le scelte tecniche di dettaglio qui
proposte derivano dalla ricognizione del runtime. Gli esiti dell'implementazione
sono riportati nel [rapporto di verifica](2026-09-08-memory-m1-validation.md).
sono riportati nel [rapporto di verifica](../reports/knowledge-archives-release.md).
## Problem Statement
@@ -1,103 +0,0 @@
# M1 — Implementazione e verifica
Data: 2026-09-08. Implementazione locale della
[specifica approvata](2026-09-08-memory-m1-spec.md), associata all'
[issue 27](https://git.tylconsulting.it/mptyl/ThothII/issues/27).
## Risultato
La pagina **Memory management** è disponibile nell'Administration dopo Database
management. Gestisce le quattro famiglie di card, elenco completo, ricerca e filtri,
ordinamento, dettaglio, creazione, modifica, cancellazione, collegamenti e dipendenze.
Richiede un amministratore autenticato e una selezione esplicita del workspace;
non richiede una sessione o un database DWH configurato.
Il harness possiede l'archivio PostgreSQL `thoth_memory`. Card, collegamenti,
dipendenze e lavoro di propagazione sono salvati nella stessa transazione.
La pagina distingue salvataggio fallito e contenuto salvato con indice incompleto,
offrendo retry anche per le cancellazioni. Recall Memory ed exemplar verificano
esistenza, workspace e proiezione corrente nell'archivio prima di restituire contenuto.
Promozione, salvataggio singolo e finalizzazione corrente passano dal servizio
autorevole. Le ricevute della sorgente impediscono duplicati e ricreazione di card
cancellate. Reindicizzazione e preprocessing non importano vecchi payload o sessioni.
L'errore Memory non annulla una sessione già finalizzata; il gate segnala anche
una promozione salvata con indicizzazione incompleta.
Le migrazioni sono versionate, controllate tramite checksum e incluse nel wheel
e nell'immagine core. Il servizio di preparazione `catalog-migrate` le esegue dopo
quelle del Catalog. Il runtime assume il ruolo limitato `thoth_memory_runtime`,
con isolamento del workspace tramite RLS e senza privilegi DDL.
## Verifiche eseguite
| Confine | Esito |
| --- | --- |
| Harness, test senza L0/L2 | 1.134 passati; i 9 test dei percorsi portabili sono stati eseguiti separatamente e sono passati. |
| Servizio Memory, PostgreSQL e Qdrant reali | 17 passati, inclusi CLI, migrazioni, ruolo runtime, isolamento, transazioni, outage, retry, cancellazioni, cambio famiglia e rebuild. |
| Gate Pi | 190 passati, inclusi identità UUID e avviso dopo salvataggio con indice incompleto. |
| Backend | 1.345 passati nella suite completa, 40 esclusi dalle condizioni previste dai test; un test di autenticazione ha superato il timeout sotto carico. Il relativo file è stato rieseguito isolato: tutti i 17 test passati. |
| Frontend | 632 passati, inclusi ingresso dall'AppShell, form, filtri, collegamenti, dipendenze e retry delle cancellazioni. |
| Browser integrato | Passato: autenticazione amministratore, creazione, modifica, riavvio del backend, rilettura, cancellazione e assenza nel recall. |
| Build e tipi | Build backend e frontend, typecheck TypeScript e build documentale strict superati. |
| Lint e diff | Ruff sui file Python modificati e `git diff --check` superati. Il lint globale segnala tre rilievi in file non modificati, elencati sotto. |
Il browser utilizza autenticamente frontend, login locale, Fastify, ThtRunner,
CLI Python, PostgreSQL e Qdrant. Gli embedding sono deterministici e le attività
Pi/sessione estranee al percorso Memory usano le fixture esistenti. Non sono state
intercettate le API Memory. Sono stati usati container temporanei PostgreSQL 16 e
Qdrant 1.18.2, senza accesso a un DWH remoto o a un modello generativo.
Il test browser ha consentito di correggere etichette accessibili instabili nei
campi compilati e la sovrapposizione del pannello di recupero ai comandi del dettaglio.
La selezione del workspace e l'uscita dalla pagina sono bloccate durante le operazioni.
Il lint globale preesistente riguarda soltanto:
- ordinamento import in `harness/tests/test_effective_relationships.py`;
- ordinamento import in `harness/tests/test_p3_dwh_binding.py`;
- uso di `datetime.UTC` in `harness/tht/mschema/catalog_snapshot.py`.
## Riproduzione
Usare Node 24 e le dipendenze installate dei tre layer. Per eseguire il harness
in un ambiente con home non scrivibile si può impostare `THT_HOME` su una directory
di prova. I test dei percorsi portabili devono essere eseguiti senza questo override,
perché verificano deliberatamente la risoluzione dell'home e di `THT_DATA_ROOT`.
```sh
cd harness
THT_HOME=/private/tmp/thothii-m1-test-home .venv/bin/pytest -m 'not l0 and not l2' --ignore=tests/test_portable_paths.py -q
.venv/bin/pytest tests/test_portable_paths.py -q
.venv/bin/pytest tests/memory/test_administration.py -q
npm test
```
```sh
cd backend
npx vitest run
npx tsc --noEmit -p .
npm run build
```
```sh
cd frontend
npx vitest run
npx tsc -b
npm run build
THT_MEMORY_BROWSER_E2E=1 npx playwright test e2e/memory-real.spec.ts
```
Il percorso browser richiede Docker, Python del harness, Go per il bridge di
autenticazione e Chromium di Playwright. Avvia risorse isolate e le rimuove alla
fine. Su macOS il browser deve poter avviare i processi Chromium fuori dalle
restrizioni della sandbox. La build documentale si esegue dalla radice con
`./scripts/build-docs.sh`.
## Stato della consegna
Le modifiche sono nel worktree locale. Nessuno stack già attivo è stato aggiornato
e nessun dato esistente è stato migrato o eliminato. Prima di usare M1 su
un'installazione occorrono il nuovo core e la preparazione `catalog-migrate`.
M2 (retrieval ibrido ed espansione dei collegamenti), M3 (integrazione estesa nel
workflow) ed Evidence management restano incrementi successivi.
+1 -1
View File
@@ -2,7 +2,7 @@
Data del piano: 2026-09-08. Aggiornamento 2026-09-09: M1–M3 implementati,
con integrazione X1 per le correzioni persistenti dei conflitti. Risultati e limiti
sono raccolti nel [rapporto X1](2026-09-09-archive-repair-x1-validation.md).
sono raccolti nel [rapporto X1](../reports/knowledge-archives-release.md).
La [revisione di semplicità](2026-09-08-memory-evidence-simplification-review.md)
mantiene il perimetro Memory e precisa un salvataggio sequenziale e una pulizia
@@ -1,9 +1,31 @@
# PRD — Security hardening per Docker personale e server multiutente
## Ripresa del lavoro e limiti di autorizzazione
Questo PRD resta una bozza, non una specifica di implementazione approvata. Il precedente
prompt di ripresa è stato consolidato qui il 15 settembre 2026. Riconfermare i rilievi
SEC-01–SEC-12 contro codice e dipendenze correnti, separando fatti, ipotesi, rischi del
profilo locale/server e problemi già risolti. Le vecchie survey non provano lo stato del server.
Usare inizialmente controlli in sola lettura e dati sintetici; accesso a server, IdP, DWH
o provider e relative mutazioni richiedono target e operazioni concordati.
Prima di implementare, il proprietario deve confermare profili di fiducia, isolamento
del runtime Pi, trattamento dei valori sensibili, revoca, retention, limiti di risorse,
priorità e criteri misurabili. I modelli visibili sul server sono un problema separato.
Registrare le decisioni in questo PRD; pubblicare spec e ticket Gitea soltanto dopo la
conferma del perimetro e della granularità. La bonifica documentale non autorizza il
codice di sicurezza, né il rollout di tutti i punti SEC.
L'implementazione futura richiede un worktree dedicato, test positivi/negativi ai confini
concordati, typecheck e revisione contro standard e specifica. Conservare gate manuali,
evidenze e rollback; un ticket completato non chiude automaticamente il PRD. Non copiare
segreti o dati operativi nei worktree. Usare le skill effettivamente disponibili per
chiarimento, diagnosi e revisione, senza assumere che i vecchi nomi dei comandi esistano.
**Stato:** bozza da validare con `grill-with-docs`; implementazione rinviata.
**Data:** 8 settembre 2026.
**Owner delle decisioni:** il maintainer di ThothII.
**Ripresa:** [prompt per la prossima sessione](2026-09-08-security-hardening-resume-prompt.md).
**Ripresa:** seguire la sezione «Ripresa del lavoro e limiti di autorizzazione» sopra.
Questo documento conserva la survey di sicurezza discussa con il maintainer e propone requisiti,
priorità e criteri di accettazione. Non è una spec approvata, un penetration test, una certificazione
@@ -1,100 +0,0 @@
# Riprendere il PRD di sicurezza con le skill di Pocock
**Stato:** prompt conservato per uso futuro; nessuna esecuzione programmata.
**PRD:** [Security hardening per Docker personale e server multiutente](2026-09-08-security-hardening-prd.md).
Apri una sessione nella codebase ThothII e incolla il blocco seguente. Il nome corretto della
skill è `grill-with-docs`, che combina `grilling` e `domain-modeling`. Il prompt apre la fase di
chiarimento; il passaggio a issue, worktree e implementazione resta soggetto alle conferme indicate.
```text
Riprendiamo il lavoro di sicurezza rinviato l'8 settembre 2026.
Leggi docs/plans/2026-09-08-security-hardening-prd.md. È una bozza di PRD ricavata da una
survey storica, non una spec approvata né una prova della configurazione del server remoto.
Voglio preparare interventi proporzionati per Docker su Mac/PC personale e per un server
multiutente con autenticazione built-in oppure OIDC. Il problema separato dei modelli
visibili sul server è fuori perimetro.
Usa realmente le skill di Matt Pocock: leggi le istruzioni installate, dichiarando quali
applichi. Parti da ask-matt per verificare il percorso e da grill-with-docs per il lavoro
di design; quest'ultima richiede grilling e domain-modeling. Se una skill non è disponibile,
segnalalo e concorda il fallback, senza installarla o fingere di averla eseguita.
FASE 1 — Riconferma delle evidenze, senza modificare il runtime
1. Leggi AGENTS.md, PROJECT_STATE.md, CONTEXT.md, le istruzioni docs/agents/ su dominio,
issue tracker e label, gli ADR pertinenti e il PRD. Controlla HEAD, stato del worktree
e differenze dalla baseline della survey. Preserva tutte le modifiche preesistenti.
2. Riconferma i rilievi SEC-01…SEC-12 nel codice corrente. Separa fatti verificati,
ipotesi, rischi condizionati al profilo e problemi già risolti. Non trattare i vecchi
conteggi delle dipendenze come una scansione aggiornata.
3. Usa controlli locali read-only e dati sintetici. Per un difetto da riprodurre, usa
diagnosing-bugs con un segnale ripetibile sul comportamento effettivo; una diagnosi
non autorizza ancora il fix. Confronta fatti di librerie e advisory con fonti primarie
correnti quando necessario, senza inviare codice privato o segreti ai servizi di ricerca.
4. Non accedere o intervenire su server, IdP, DWH o provider reali senza aver concordato
target e operazioni. Non mostrare API key, cookie, password o campioni di dati reali.
Esito della fase: una matrice aggiornata che conserva gli ID dei rilievi, con evidenza,
profilo interessato e stato. Un fatto ancora non verificabile resta esplicitamente aperto.
FASE 2 — grill-with-docs, con me presente
5. Costruisci l'albero delle decisioni. Parti da profili di deploy, fiducia fra utenti e
condivisione dei dati; poi affronta isolamento di Pi, policy dei valori sensibili,
revoca, retention e limiti seguendo le dipendenze effettive.
6. A ogni round presenta soltanto le domande attualmente sbloccate, numerate, con la tua
raccomandazione e i trade-off. Attendi le mie risposte prima di assumere le decisioni
successive. Cerca autonomamente i fatti ricavabili dal repository; usa agenti di
ricerca mirati quando previsto dalla skill, senza delegare a loro le mie decisioni.
7. Aggiorna il PRD distinguendo proposte e decisioni confermate. Aggiorna CONTEXT.md solo
per termini realmente risolti. Proponi ADR soltanto per scelte difficili da invertire,
sorprendenti senza contesto e fondate su alternative reali: basta il formato minimo.
8. Concorda requisiti, priorità, rischi accettati e criteri misurabili, inclusi i tempi
di revoca e i limiti di risorse. Conferma con me i confini pubblici dei test prima
di scriverli. Mantieni espliciti i gate manuali già presenti in PROJECT_STATE.md.
Esito della fase: nessuna decisione bloccante lasciata implicitamente all'agente;
riepilogo e mia conferma della comprensione condivisa. Fino a quella conferma rimani
su analisi e documentazione: nessun cambiamento applicativo o di deployment.
FASE 3 — Spec e ticket, soltanto dopo mia conferma
9. Usa to-spec per sintetizzare le decisioni già prese, senza riaprire arbitrariamente
l'intervista. Chiedimi conferma della pubblicazione prima di creare la spec nel
tracker canonico Gitea indicato in docs/agents/issue-tracker.md, non nel mirror GitHub.
Collega la spec canonica dal PRD e rendi chiaro quale documento è la fonte aggiornata.
10. Usa to-tickets per proporre fette verticali verificabili autonomamente, dimensionate
per un contesto fresco. Collega ogni ticket ai requisiti e ai rilievi pertinenti,
indica i veri blocker e includi criteri positivi e negativi. Fai approvare granularità
e dipendenze prima di pubblicare. Solo i ticket approvati e completi ricevono
ready-for-agent; non rimetterli in triage e non chiudere automaticamente la spec padre.
11. Se una decisione richiede una prova eseguibile, proponi un prototype limitato a quella
domanda prima di fissare la spec. Usa wayfinder solo se il lavoro risulta realmente
troppo ampio e incerto per essere chiarito con grill-with-docs.
Esito della fase: spec approvata e ticket autosufficienti con dipendenze risolte o esplicite.
Chiedimi se autorizzo il primo ticket: l'approvazione del design non avvia da sola il codice.
FASE 4 — Implementazione futura autorizzata
12. Prima di modificare codice, concorda e crea un worktree dedicato, verificando percorso,
branch e commit base. Non riusare una directory occupata, non alterare il worktree
originario e non copiare automaticamente segreti o dati operativi. Assicurati che
PRD e prompt siano disponibili nel worktree attraverso un passaggio esplicito.
13. Esegui implement su un ticket sbloccato per volta, in un contesto fresco. Segui tdd
ai confini concordati: un test rosso sul comportamento, implementazione minima,
test verde. Esegui typecheck e test mirati durante il lavoro e le suite pertinenti
al termine; usa fixture locali per IdP, DWH e provider.
14. Esegui code-review sui due assi Standards e Spec, usando i due agenti previsti dalla
skill e una base Git fissata. Assicurati che il diff esaminato includa tutto il lavoro
del ticket, anche se ancora non committato; un diff vuoto non è una review superata.
Risolvi i rilievi e verifica di nuovo. Commit soltanto del lavoro pertinente nel
worktree autorizzato; push, merge e deploy richiedono un'autorizzazione distinta.
15. Consegna evidenze dei test, istruzioni di adozione e rollback, gate manuali pendenti
e rischi residui. Non dichiarare chiuso il PRD intero se è concluso soltanto un ticket
o se resta un'accettazione dell'owner.
Inizia dalla Fase 1, poi proponimi il primo round di grill-with-docs.
```
@@ -1,112 +0,0 @@
# X1 — validation of session archive corrections
Date: 2026-09-09. The joint Memory/Evidence repair increment is implemented.
The authoritative contract is [Session archive corrections](../contracts/archive-repair.md).
## Delivered behavior
The session gate shows complete before/after content for specific alternatives targeting
Memory or Evidence. The reviewer chooses one correction or rejects all proposals as
inadequate and requests reformulation. The resulting receipt survives interruption;
saved content and index activation are reported separately. Pending activation offers
retry of the same chosen operation. A subsequent curator change blocks replay.
Application requires an administrator in the harness and the responding browser
principal's archive-management permission. Cross-principal runtime responses cannot
misattribute the correction. A non-administrator can decline or continue the current
question without modifying shared archives. The gate does not advance a workflow phase.
Memory and Evidence remain separate domains. The integration coordinator reuses their
canonical persistence and activation operations. A session Evidence correction requires
a consolidated archive, preserves source lineage, and cannot publish unrelated external
edits. Existing administration, source import, dependency cleanup and final Memory review
remain available. No automatic Git commit or push was added.
## Verification
- Harness regression: **1,264 passed**, one skipped, five deselected. All **nine**
portable-path checks passed separately without `THT_HOME`. Python lint passed on changed modules.
- Backend: **1,366 passed**, 40 skipped. Tests include actual session-response routes
for both target archives, unauthorized response, malformed choice and runtime ownership.
- Frontend: complete suite **645 passed**; the final display adjustment passed all four
focused widget tests. Backend and frontend TypeScript checks passed.
- Pi extension: **199 passed**, including closed human choices, rejection, failure/retry,
forged selections and the updated public tool schema. The modular skill projection is
byte-identical to its updated approved template.
- PostgreSQL/Qdrant integration traverses the actual Python CLI for preparation,
application and recovery inspection. Corrected Memory and Evidence are retrieved from
real indexes, and Evidence activation preserves the Memory card. Session loading is a
controlled fixture and embeddings are deterministic; this is not an LLM quality test.
- Failure tests cover both targets, index outage, replay, later edits, workspace/session
isolation, changed session context, rejection, non-admin writes, and interruption after
the Evidence file write but before its saved receipt.
- Playwright desktop/mobile: **one passed**. The real widget renders both alternatives,
accepts an Evidence choice, displays pending activation and allows retry to active.
No page errors or mobile horizontal overflow. Screenshots are
`/private/tmp/thothii-x1-repair-desktop.png` and `/private/tmp/thothii-x1-repair-mobile.png`.
This browser fixture controls operation outcomes; persistent behavior is tested above.
- Strict MkDocs build and `git diff --check` passed.
## Local installation and reviewer acceptance
Core and frontend images were rebuilt from this worktree using the existing local
preview launcher. Migration `004_archive_repairs.sql` was applied to the existing
installation catalog. The new gate is available to session workflows; it is not an
always-visible administration panel. Existing PSD archive content was not changed by
the synthetic validation cases.
All five local services are healthy at `http://127.0.0.1:8080/`.
The technical increments and their planned checks are complete. The end-user acceptance
check remains a real session containing a meaningful domain conflict, with the reviewer
evaluating the proposed correction. Automated browser validation uses temporary accounts
and data, not the user's authenticated PSD session. Source import retains its E3
validation boundaries; no broader model-quality benchmark was added.
## Follow-up acceptance: configured model
The opt-in `test_real_model_proposes_a_reviewable_persistent_archive_correction`
passed with the installation's **zai/glm-5.3** model. Synthetic Memory asserted an
order-ID-only join; synthetic Evidence required the financial year too. The model
returned two schema-valid, specific alternatives with complete content and the exact
target revisions. The test reviewer selected Memory, persisted the correction through
the real coordinator and PostgreSQL, and retrieved the updated rule. Evidence stayed
unchanged. This test uses deterministic vectors and the configured completion helper;
it does not claim a full autonomous Pi session or human acceptance of PSD semantics.
The run log is `/private/tmp/x1-acceptance-model.log`. Reproduce with
`THT_MEMORY_L2_INSTALLATION=<installation.yaml>` and `THT_MEMORY_L2_CORE=<core-container>`
using `pytest -q -s -m l2 tests/memory/test_administration.py -k real_model_proposes`.
Credentials are resolved inside core and are not returned to the test runner.
## Follow-up acceptance: both administration pages
The opt-in `frontend/e2e/memory-real.spec.ts` passed through real authentication,
Fastify, ThtRunner, Python, isolated PostgreSQL and Qdrant. It verifies:
- Database management, Memory management and Evidence management appear as peers in
that order, with no active core session or DWH binding required.
- Memory creation, editing, persistence across backend restart, deletion and absence
from subsequent recall.
- Canonical Evidence remains intact after the Memory deletion. Its full rule is read
through the real Evidence administration worker; content filtering finds it and an
unmatched filter produces the empty state.
- Requests for an unregistered workspace return 404 for both archives.
- Desktop and mobile Evidence views render without horizontal document overflow.
On phones, both archive pages have at least 380px of usable width at a 390px viewport.
Navigation opens in the shared accessible dialog, closes with Escape or archive selection,
and returns focus to the trigger after Escape.
The temporary PostgreSQL readiness probe now waits for TCP, avoiding the image's
socket-only initialization server. The browser waits for Memory refresh to finish
before leaving its page, matching the existing navigation guard. Visual inspection
also exposed a real mobile layout issue: the fixed sidebar left only 134px for the
Evidence page. `ArchiveNavigation` now moves that sidebar into the shared dialog below
768px on Memory/Evidence pages. Desktop behavior is unchanged. The frontend image
was rebuilt for the local preview.
Run log: `/private/tmp/x1-acceptance-browser7.log` (**one passed**).
Screenshots: `/private/tmp/thothii-acceptance-evidence-desktop.png` and
`/private/tmp/thothii-acceptance-evidence-mobile.png`. Reproduce with
`THT_MEMORY_BROWSER_E2E=1 npx playwright test e2e/memory-real.spec.ts` from `frontend/`.
The fixture removes its temporary containers, accounts and checkout on completion.
The TypeScript check, Python lint, strict documentation build and diff check also pass.
@@ -1,67 +0,0 @@
# Evidence E1 — validation
Date: 2026-09-09. Scope: editable Curated Evidence v4 and the persistent local archive.
## Implemented behavior
- Parser, renderer, authoring output and normalization share the existing typed payloads.
Visible Markdown edits determine content for all eight kinds. Legacy v1–v3 conversion
is explicit and lossless, with errors for content that cannot be represented exactly.
- Manual declarations record the curator. A correction preserves the original document
as lineage, separately from the current declaration. No source hash is needed to
create a manual file.
- The local archive records baselines, immutable candidates, active revisions and
deletion/source suppression metadata. Unresolved review items and invalid edits block
consolidation. Missing archive directories are availability failures, not deletions.
- Activation failures preserve the previous active revision. Interrupted normalization
replays only unchanged input bytes; later operator edits survive recovery.
- Revision-checked correction methods reject stale workflow updates. Legacy preparation
and resolution cannot overwrite an initialized local archive; explicit import/refresh
integration is deferred to E3.
## Verification
The final harness suite excluding opt-in L0/L2 and portable-layout cases passed with
**1,180 tests** (58 deselected). All **9 portable-layout tests** passed separately with
`THT_HOME` unset. The dedicated real-Qdrant integration test passed, including the
optional 35-unit PSD probe. Ruff passed on the changed Evidence implementation and
tests, and the strict documentation build succeeded. No frontend or backend TypeScript
changes are part of E1.
The integration test uses an isolated Qdrant 1.18.2 container, the actual corpus
pipeline, semantic chunking, vector adapter and active Evidence searcher. Deterministic
three-dimensional embeddings isolate file/content correctness from model behavior.
It verifies that raw edits do not change recall, consolidation updates recalled content
and curator identity, a blocked candidate preserves prior recall, and deletions remove
recall. Existing schema and Memory records survive each operation.
All **35 PSD units** were copied from the owner's workspace into
`/private/tmp/thothii-e1-psd.bsW4cp`. Deterministic conversion preserved every ID, payload,
scope, provenance and review item. There were no unresolved review items. The optional
integration probe then indexed all 35 converted units and compared their complete ID
set to the original. It uses PSD's actual `max_chunk_chars: 5000`; a preliminary probe
at 4000 correctly blocked an oversized atomic unit.
Reproduce the isolated real-corpus probe after creating a converted workspace copy:
```sh
cd harness
THT_E1_PSD_COPY=/absolute/path/to/converted-copy \
.venv/bin/pytest -q -s tests/test_evidence_editable_integration.py
```
The environment variable is optional. Ordinary CI uses only synthetic Evidence. No
source refresh, external document download, DWH call or model request is involved.
## Delivery boundary
E1 is a core/library increment. E2 must add the installed manual consolidation command,
connect runtime source selection to the active local snapshot, and build administrative
list/filter/detail with real persistent host paths and manual Git instructions. E3
adds source acquisition and explicit refresh/conflict handling. X1 later wires deliberate
joint Memory/Evidence corrections into review gates.
The actual PSD Evidence checkout was not converted. The live Docker preview at
`http://127.0.0.1:8080` remains the previously deployed M3 stack, with no new Evidence
administration page. The corpus conversion and reindexing described here used copies
and disposable test resources.
@@ -1,86 +0,0 @@
# Evidence E2 — validation
Date: 2026-09-09. E2 is implemented locally and installed on the existing Docker preview.
E3 source imports/refresh and X1 deliberate Memory/Evidence workflow corrections remain open.
## Delivered behavior
- Independent **Administration → Evidence management**, after Memory, protected by
`evidence.manage`: complete typed content, provenance and original excerpts, review items,
pagination, search, kind/purpose/status and scope/source filters, sort, and refresh.
- Working-file states distinguish active, modified, new, removed, invalid, legacy and review
required. Detail shows the actual configured host path with copy controls. Instructions
cover external editing, all eight Markdown templates, consolidation and manual Git.
- Installed `tht workspace evidence consolidate --workspace <id> [--json]` uses a closed
maintenance envelope. First use converts legacy units. Validation, immutable candidates,
activation and retry run through the existing corpus pipeline without a DWH scan or Git.
- Runtime and ordinary preprocessing consume the active local snapshot. Unconsolidated
edits remain excluded. Catalog/Schema readiness is not advanced by this operation.
Clear preserves curated files, archive metadata and Memory; full preprocessing must
recreate the missing Reference/Schema derivations afterward.
- Immutable runtime lease filenames now identify rendered bytes as well as logical input
identity. This fixes upgrades colliding with old runtime files without changing Catalog
fingerprints or removing the checks against tampered files.
## Automated checks
The complete backend suite passed: **1,355 tests**, 40 skipped. The complete frontend
suite passed: **639 tests**. Both TypeScript checks passed. Native Go workspace operation
and CLI tests passed, including rejection of arbitrary consolidation flags. Ruff passed
for changed Python implementation and test files.
The harness run passed **1,248 tests**, with one skipped and five deselected. Its three
portable-path tests failed because that run deliberately set `THT_HOME` to the test
runtime; rerunning the portable tests with `THT_HOME` unset passed. The final focused
administration/path suite passed all 18 tests, including actionable migration errors and invalid
consolidation combinations rejected before cleanup or indexing.
The real-Qdrant integration test exercised the actual harness consolidation CLI with
deterministic embeddings: all 35 PSD units converted and indexed with stable identities;
active-only source selection; separate Schema and Memory canaries; Clear and rebuild
from the retained snapshot. Unit tests cover validation, saved-but-unindexed failure,
retry, browsing/filtering, no automatic Git, and no Catalog mutation from consolidation.
A temporary Git repository and bare local remote exercise the documented manual sequence:
edit, add and remove files, consolidate, inspect, stage the complete Evidence tree,
commit, push and clone. The clone retains changed content, additions, deletions, managed
metadata and an accessible active snapshot. No remote user repository was pushed.
## Installed preview
The existing Compose project is `thothii-18998cca7b0a`, at `http://127.0.0.1:8080`.
The persistent editable checkout is:
```text
/Users/mp/projects/ThothII/deploy/psd/evidence-registry/repo/psd-clinical/evidence
```
The original registry checkout was copied from its retained Docker volume. The original
author repository was not changed. Core and maintenance share a nested host bind for
`repo`; registry state/snapshots and all other existing data volumes were retained.
The installation descriptor includes the existing workspace bindings and the new
`evidence-host.yaml` override. The previous descriptor and native binary are backed up
at `/private/tmp/thothii-installation-before-e2.yaml` and `/private/tmp/tht-before-e2`.
The real installed command succeeded with **35 documents, 35 chunks, 35 changed, zero
removed**, using the configured embedding service and Qdrant. A second run succeeded
with **35 unchanged, zero changed**. Reading the actual archive from core returned
35 active units and no file errors. All five long-running services are healthy.
Only the Evidence stage ran. The strict documentation build and `git diff --check`
also passed. The stack launcher is `bash /private/tmp/thothii-memory-preview.sh`; keep its
worktree image-build override until this branch is integrated into the main checkout.
Browser verification reached the local login page. The saved administrator password
does not match the current account hash, so the authenticated visual check remains
manual. No account or password was modified. React interaction tests cover navigation,
detail, host paths, templates, filtering, pending activation and invalid files.
## Boundaries
There is no web content editor, watcher, automatic commit/push, or implicit source refresh.
The API exposes administration reads and consolidation; the archive's revision-checked
save/remove operations remain available for the later explicit workflow corrections.
These gates are not claimed as implemented by E2. Initialized local archives retain
structural/review checks but bypass the legacy fixed retrieval-evaluation fixture so
its old expected IDs cannot veto deliberate deletions. A general retrieval benchmark
is outside the agreed scope.
@@ -1,100 +0,0 @@
# Evidence E3 — validation
Date: 2026-09-09. Explicit source import/refresh and decisions are implemented. X1,
the integration of deliberate Memory/Evidence corrections into workflow gates, remains next.
## Delivered behavior
The independent Evidence page now offers **Sources and imports**. An operator copies
a specialist's draft into `evidence/incoming/`, then explicitly imports/refreshes.
Original local Markdown and configured HTTP/S3 sources use existing read-only adapters.
Acquisition retains raw bytes, source identity and versioned normalized documents.
The existing Pi authoring refiner prepares typed, editable v4 proposals.
Unchanged hashes skip refinement. All source acquisitions/refinements must succeed
before saving a new set of comparisons. Missing sources are recorded as unavailable,
never interpreted as permission to delete. No runtime lookup, ordinary consolidation
or preprocessing triggers remote refresh once the local archive is initialized.
The administrator sees current and proposed units, scope, content, excerpts, review
items and explicit retirement IDs. **Keep local Evidence** records the retained wording
as a manual declaration with original lineage. **Use proposed Evidence** adopts the
proposal and its source version. Both save and activate through the existing archive
and corpus pipeline; review items block adoption. Comparisons use optimistic checks
on affected file bytes. Interrupted decisions have a durable replay journal and retry
without reacquisition, while intervening external edits are preserved and reported.
Deleted IDs remain reserved. New model-generated identities from sources with curated
deletions are also conservatively suppressed; surviving IDs can still receive reviewed
updates. Deliberate new knowledge can be authored as a manual file. This mechanical
protection does not depend on the model detecting semantic duplication or contradictions.
Installed commands are `workspace evidence refresh` and `workspace evidence decide`,
alongside E2 consolidation. Decision envelopes carry a source identity, comparison
revision and keep/replace choice. Extra URLs, arbitrary paths, forged actors and unknown
fields are rejected at the public API/CLI boundary. HTTP requests bind the authenticated
curator. Source operations do not mutate Catalog readiness or run DWH/schema stages.
The Python source worker is internal; the workflow CLI's visible surface is preserved.
## Checks
- Complete backend suite: **1,359 passed**, 40 skipped. Complete frontend suite:
**641 passed**. Both TypeScript checks passed; native Go CLI/workspace tests passed.
- Harness regression run: **1,256 passed**, one skipped and five deselected, with
portable-path tests run separately without `THT_HOME`. All **24 focused import,
CLI-surface and portable-path checks** passed. These
cover import, unchanged refresh, access failure, missing source, manual correction,
keep/replace, deletion suppression, stale comparisons, failure/retry and interrupted
journal writes. Ruff passed on the changed Python implementation and tests.
- The real-Qdrant test traverses the actual harness source CLI with deterministic
refinement/embedding boundaries: import is absent from recall before a decision,
accepted content becomes searchable, refreshed proposals preserve active manual
corrections, replacement removes the former text, deletion remains absent after
another refresh, and unrelated Schema/Memory canaries survive.
- Source contract fixtures cover controlled HTTP and S3 identities, exact acquired
bytes, and acquisition call counts. Existing adapter tests retain transport/egress
coverage. The test does not claim to exercise a live S3 account.
- React interaction tests verify explicit refresh, comparison content, exact decisions,
saved-decision retry and failure feedback. Route tests cover admin authorization,
workspace isolation, strict inputs and principal attribution. Service tests verify
the trusted config file descriptor and absence of Catalog mutation.
## Local preview
Core/frontend were rebuilt for the existing `thothii-18998cca7b0a` stack. Its persistent
archive and data volumes are retained. The native `/usr/local/bin/tht` was updated;
the previous executable is at `/private/tmp/tht-before-e3`.
The installed refresh command ran against `psd-clinical` successfully: **35 unchanged
sources, zero changed, zero pending comparisons**. All 35 source hashes matched their
existing units, so this probe required no refinement and changed no active Evidence.
Source registry metadata was saved locally; no Git commit or push was performed.
A separate synthetic draft was passed to the configured Pi/model inside core. It
produced one domain proposal with one review item, which was not activated. The probe
exposed a deployment issue: Python wheel modules and Pi skills live in different
directories. The refiner now resolves resources through `THT_HARNESS_DIR`, with the
source-tree location as its development fallback; a regression test covers this layout.
The corrected installed worker was then exercised end to end in a temporary workspace
inside core, using the real configured Pi/model and a synthetic `incoming/orders.md`.
It returned success, one changed source, one comparison and one proposal with a review
item. The active snapshot remained absent. The temporary directory was removed on exit;
the probe did not open the DWH or activate an index. Strict docs build and
`git diff --check` also passed.
The in-app browser still showed the login page with the prior credential error.
Authenticated visual verification remains manual; no credentials were reset or retried.
## Operational instructions and boundaries
See [Import drafts and refresh sources](../contracts/curated-evidence-v4.md#import-drafts-and-refresh-sources)
for commands, local paths, review, retry and backup/Git requirements. Preserve the
complete Evidence tree, including source comparisons, journals and acquired versions.
Local activation and transfer to the remote Git repository remain separate operator steps.
No web content editor, automatic Git, background watcher, new job queue, general
retrieval benchmark or automatic source merge was introduced. Review/refinement is
sequential and bounded by per-source limits plus 200 documents/100 MiB per refresh.
New changed-source decisions and failure scenarios use isolated test data; live PSD
curated content was kept unchanged. X1 is not included in this increment.
@@ -1,96 +0,0 @@
# M2 — Ricerca ibrida e collegamenti
Data: 2026-09-09. Implementazione locale dell'incremento M2 del
[piano Memory approvato](2026-09-08-memory-management.md#piano-esecutivo).
## Risultato
La ricerca Memory ed exemplar usa embedding dense e BM25 in Qdrant. I filtri sono
applicati in entrambi i rami prima della selezione dei candidati. Il core espande
i collegamenti uscenti, deduplica, limita cicli e visite, riordina insieme i risultati
e restituisce il contenuto corrente verificato in PostgreSQL. Non usa un database
a grafo né una chiamata LLM per il riordinamento.
Il contesto fisico distingue database, schema, tabella e colonna e richiede che
corrispondano alla stessa dipendenza strutturata. Le card senza dipendenze valgono
per il workspace; un riferimento a un antenato fisico si applica ai suoi discendenti.
Ambito descrittivo e concetti possono restringere ulteriormente la ricerca tramite
`--filters`. La CLI impedisce di sostituire database/schema del runtime. Il dettaglio
dei limiti e della formula di ranking è nel [contratto operativo](../gestione-memory.md#hybrid-recall-and-links).
L'incremento conserva i confini dei gate correnti: F2 riceve chiarimenti di dominio,
gli exemplar restano consultativi. Un collegamento non autorizza a consumare una
famiglia diversa o una card fuori ambito. Il riepilogo finale, i nuovi gate e la
pulizia delle dipendenze dopo sincronizzazione fisica appartengono a M3.
## Transizione e recupero
La migrazione versionata `002_hybrid_projection.sql` aggiunge il formato delle
proiezioni. Le card M1 restano autorevoli e consultabili; le loro proiezioni dense
risultano pendenti e non possono alimentare il recall. Un retry esplicito o
`tht memory index -c <runtime.yaml>` costruisce dense e BM25 dal contenuto corrente.
Il formato della proiezione e la revisione della card sono verificati prima dell'uso.
Il salvataggio può aggiungere il vettore sparse mancante nella sola collezione
Memory. Non sostituisce configurazioni incompatibili e non modifica Reference.
Il rebuild esplicito ricrea anche una collezione Memory assente; il test elimina
la collezione, ricostruisce da PostgreSQL e verifica il ritorno della sola card conservata.
Fallimenti lasciano il lavoro di propagazione persistito e recuperabile. Nessuna
importazione da JSONL, sessioni storiche o vecchi payload Qdrant.
## Verifiche
| Controllo | Esito |
| --- | --- |
| Suite harness senza L0/L2, escluso il file dei percorsi portabili | 1.152 test passati nell'esecuzione finale. |
| Suite mirata Memory, adapter e CLI, con embedding reale | 78 test passati. |
| Verifica aggiuntiva del rebuild con collezione assente e adapter | 54 passati; il solo test del modello reale era escluso in questa riesecuzione. |
| API Fastify Memory | 13 test passati, compresa propagazione della lingua del workspace. |
| Browser amministrativo integrato | Passato: creazione, modifica, riavvio, rilettura, cancellazione e verifica del recall. |
| Wheel e casi CLI | 10 test passati; il wheel include entrambe le migrazioni Memory. |
| Build e controlli statici | Build/typecheck backend, Ruff sui file Python interessati e build documentale strict passati. |
- Test deterministici: collegamenti necessari, contenuto corrente, duplicati,
cicli, profondità e limiti, card mancanti o pendenti, rifiuti già registrati,
famiglie e ambiti esclusi, dipendenze omonime e filtri CLI vincolati al runtime.
- Adapter: stesso filtro nei prefetch dense/BM25 per Memory ed exemplar; restano
coperti i contratti Evidence esistenti.
- PostgreSQL/Qdrant: salvataggio, modifica, cambio famiglia, cancellazione,
ricostruzione, riferimento separato, isolamento, RLS, errori, retry e transizione
delle proiezioni M1 al formato ibrido.
- Recupero effettivo: client Ollama di produzione con il modello configurato
`qwen3-embedding:0.6b`, dimensione 1024, e Qdrant dell'immagine fissata in Compose.
La domanda sulla chiave commessa fra esercizi recupera la regola attesa e la
granularità collegata, escludendo un altro database, un altro ambito e dipendenze
che corrispondono soltanto combinando riferimenti distinti. Verifica separatamente
dense, BM25 e fusione, poi recall, cancellazione e rebuild.
Il test effettivo avvia un processo Ollama separato, montando il volume del modello
installato in sola lettura. PostgreSQL e Qdrant sono container temporanei dedicati,
eliminati a fine test. Non usa ID di risultati predisposti. Il browser amministrativo
usa invece embedding deterministici: verifica il collegamento fra UI, autenticazione,
Fastify, ThtRunner e persistenza, senza essere una misura di qualità del recupero.
Non è una valutazione generale della qualità semantica su un corpus di produzione;
verifica i casi di recupero richiesti da M2, con il percorso reale configurato.
## Riproduzione
Dalla directory `harness`, con Docker disponibile:
```sh
THT_HOME=/private/tmp/thothii-m2-test-home \
THT_MEMORY_TEST_OLLAMA_VOLUME=<volume-modelli-installazione> \
THT_MEMORY_TEST_MODEL=qwen3-embedding:0.6b \
THT_MEMORY_TEST_DIMENSIONS=1024 \
.venv/bin/pytest -q tests/memory/test_administration.py \
tests/memory/test_retrieval.py tests/memory/test_recall.py \
tests/test_qdrant_vector_store.py tests/test_solved_search_cli.py
```
Senza il volume esplicito, il solo test con modello reale viene escluso; gli altri
test restano eseguibili. Il volume deve contenere il modello indicato. Le immagini
Qdrant e Ollama del test sono lette da `compose.yaml`.
Consegna locale: non sono stati eseguiti deploy, migrazioni delle installazioni
attive o aggiornamenti remoti dell'issue tracker.
@@ -1,78 +0,0 @@
# Memory M3 — validation
Implemented on 2026-09-09 in the rapid-harbor worktree.
## Delivered behavior
- F8 presents an editable summary of proposed additions, explicit updates and links.
Only selected content is saved, including the optional solved-question exemplar.
- Proposals reference effective approved decisions. Exact existing content is reused.
Concurrent edits invalidate an update; manual identities and origins are preserved.
- Selected cards and links commit atomically. Durable receipts recover repeat delivery
and the gap before the session review marker. Finalization does not add Memory.
- F4/F6/F7 retrieve SQL rules and explained errors for the existing approval gates.
Retrieval is consultative and does not write an approval decision.
- Successful Catalog physical sync deletes cards with matching removed dependencies.
The same Catalog transaction marks pending Memory cleanup. Retry keeps the original
removals and does not rescan; deletion receipts and projection tombstones survive restarts.
- Migration 003 adds minimal review and physical-cleanup receipts.
## Executed checks
| Check | Result |
| --- | --- |
| Harness deterministic suite, excluding portable-path environment cases | 1,152 passed |
| Portable-path suite without the temporary THT_HOME override | 9 passed |
| Memory service/retrieval with isolated PostgreSQL and Qdrant | 41 passed, 1 optional real-embedding case skipped, 1 L2 case excluded |
| Pi gate suite | 195 passed |
| Backend full suite | 1,346 passed; one auth timing test exceeded 5 seconds under concurrent load |
| Isolated auth and Catalog route rerun | All 40 passed, including the timed-out case |
| Catalog PostgreSQL integration after adding atomic cleanup-marker coverage | All 5 passed |
| Frontend full suite | 635 passed |
| Chromium summary review, desktop and 390px mobile | Passed; no page errors or horizontal overflow |
| Configured real GLM 5.3 generation | Passed on synthetic PostgreSQL data |
| Backend/frontend production builds, modified Python lint, strict docs build | Passed |
The existing local Docker preview was rebuilt from this worktree, migration 003
was applied, and core/frontend were recreated with the existing persistent volumes.
The preview remains at `http://127.0.0.1:8080`.
The browser check uses the production widget in an isolated Vite fixture. It edits
the rule, declines the exemplar, submits only the selected card and checks responsive
layout. Gate tests separately verify request ordering through the production Pi
composition root; service and Catalog tests use real PostgreSQL. This is not a claim
of an automated complete live Pi conversation.
The L2 case retrieves an approved SQL rule, excludes a card bound to another database,
and asks the configured GLM 5.3 model to generate a query. Order IDs repeat between
financial years; the correct composite join returns 120 on the synthetic fixture.
The generated SQL is validated and executed in a read-only PostgreSQL transaction.
No real DWH rows are sent. Embeddings in this case are deterministic; the real
embedding/hybrid retrieval evidence remains documented in M2.
## Reproduction
From the harness:
```sh
THT_HOME=/private/tmp/thothii-m1-harness-home .venv/bin/pytest -q -m 'not l0 and not l2' --ignore=tests/test_portable_paths.py
.venv/bin/pytest -q tests/test_portable_paths.py
THT_HOME=/private/tmp/thothii-m1-harness-home .venv/bin/pytest -q tests/memory/test_administration.py tests/memory/test_retrieval.py -m 'not l2'
npm test
```
The optional generation case requires an installation YAML path and its running
core container. It resolves the model credential inside core without printing it:
```sh
THT_MEMORY_L2_INSTALLATION=<installation.yaml> THT_MEMORY_L2_CORE=<core-container> \
.venv/bin/pytest -q -s -m l2 tests/memory/test_administration.py -k real_model
```
From frontend: `npx playwright test e2e/memory-review.spec.ts`.
Screenshots are written to `/private/tmp/thothii-m3-summary-desktop.png` and
`/private/tmp/thothii-m3-summary-mobile.png`.
Evidence authoring and the joint X1 persistent Memory/Evidence conflict repair remain
outside M3. This increment does not infer knowledge from unexplained failures or
promise general improvements in SQL-generation accuracy.