diff --git a/docs/research/2026-08-23-legacy-thothai-metadata-capabilities.md b/docs/research/2026-08-23-legacy-thothai-metadata-capabilities.md index 26d2825b..36eb403a 100644 --- a/docs/research/2026-08-23-legacy-thothai-metadata-capabilities.md +++ b/docs/research/2026-08-23-legacy-thothai-metadata-capabilities.md @@ -1,21 +1,80 @@ # Inventario delle funzionalità legacy di gestione metadati di ThothAI -Data: 2026-08-23 +Data dell'inventario: 2026-08-23 -Issue: `mptyl/ThothII#6` +Ultima verifica rispetto a ThothII: 2026-08-31 + +Stato: **inventario storico verificato; non è una specifica dello stato corrente**. + +Issue originaria: [mptyl/ThothII#6](https://git.tylconsulting.it/mptyl/ThothII/issues/6) Fonte primaria: repository legacy annidato `Thoth/ThothAI`, commit `55855de0f18e5cb4bc72f0a2ab0a7317995186dd`. -## Sintesi +Le citazioni che iniziano con `/Thoth/ThothAI/` sono relative alla radice di quella copia legacy +fissata al commit indicato. Alla verifica del 2026-08-31 tutti i 68 riferimenti univoci puntavano a +file esistenti e a intervalli di righe validi. -La parità funzionale richiesta a ThothII comprende un catalogo dei database, l'inventario +> Questo documento conserva l'inventario e le evidenze degli anti-pattern di ThothAI. Le frasi di +> requisito nelle sezioni successive descrivono la baseline proposta il 2026-08-23; non prevalgono +> su ADR, contratti, codice o `PROJECT_STATE.md` correnti. + +## Stato rispetto al codice corrente + +### Fonti autorevoli correnti + +Per capire cosa esiste oggi, usare nell'ordine: + +- `PROJECT_STATE.md`, per lo snapshot operativo aggiornato; +- `CONTEXT.md`, per il modello di dominio corrente; +- il [piano accettato del Metadata Catalog](../plans/2026-08-26-metadata-catalog-from-thothai.md) + e il [contratto dello schema snapshot](../contracts/catalog-schema-snapshot.md); +- gli ADR del catalogo, in particolare + [0001](../adr/0001-postgres-metadata-catalog.md), + [0004](../adr/0004-fastify-kysely-metadata-catalog.md), + [0006](../adr/0006-separate-physical-and-logical-relationships.md), + [0007](../adr/0007-durable-authoritative-schema-synchronization.md), + [0008](../adr/0008-allow-manual-catalog-metadata-cleanup.md), + [0009](../adr/0009-use-one-sequential-description-generation-run.md), + [0010](../adr/0010-allow-bounded-real-source-samples-for-description-generation.md) e + [0011](../adr/0011-gate-source-samples-with-a-sensitive-data-flag.md); +- l'implementazione, soprattutto `backend/src/catalog/types.ts`, + `backend/src/routes/catalog-databases.ts`, `backend/src/routes/catalog-schema.ts`, + `backend/src/routes/catalog-description-generation.ts`, + `backend/src/routes/catalog-description-consolidation.ts` e + `frontend/src/shell/DatabaseManagementPage.tsx`. + +### Disposizione della baseline storica + +| Capacità inventariata | Disposizione al 2026-08-31 | Evidenza corrente o nota | +|---|---|---| +| Configurazione per workspace, secret separati e test connessione | **Adottata e implementata** | PostgreSQL interno, binding `postgres_direct`, `rest_api` o `ssh_tunnel`, secret write-only; ADR [0001](../adr/0001-postgres-metadata-catalog.md)–[0004](../adr/0004-fastify-kysely-metadata-catalog.md) e `backend/src/routes/catalog-databases.ts`. | +| Inventario fisico di tabelle, colonne, PK e FK | **Adottato e implementato** | La struttura osservata è distinta dai contenuti curati; `backend/src/catalog/types.ts`, `backend/src/routes/catalog-schema.ts` e ADR [0006](../adr/0006-separate-physical-and-logical-relationships.md). | +| Riconciliazione completa e osservabile dello schema | **Implementata; semantica storica parzialmente superata** | I durable Catalog Sync Runs sono autoritativi, fail-closed, atomici e richiedono conferma per diff distruttivi. Non esiste il lifecycle `drift/removed` proposto qui: la sincronizzazione riconcilia la membership e ADR [0008](../adr/0008-allow-manual-catalog-metadata-cleanup.md) consente anche cleanup manuale esplicito, superando ADR-0005. | +| Descrizione curata e Generated Description separate | **Adottata e implementata** | Tabelle e colonne espongono entrambi i campi in `backend/src/catalog/types.ts`; la copia selettiva AI → curato è in `backend/src/routes/catalog-description-consolidation.ts`. | +| Generazione AI selettiva con run, stato e log | **Adottata e implementata** | Un run asincrono installazione-wide, sequenziale, con eventi persistiti e senza resume automatico; ADR [0009](../adr/0009-use-one-sequential-description-generation-run.md) e `backend/src/routes/catalog-description-generation.ts`. I thread daemon legacy sono **esclusi**. | +| Campioni reali per la generazione e protezione dei campi sensibili | **Implementata come estensione correttiva** | ADR [0010](../adr/0010-allow-bounded-real-source-samples-for-description-generation.md) e [0011](../adr/0011-gate-source-samples-with-a-sensitive-data-flag.md): campioni bounded per colonne non sensibili e valori sintetici deterministici per quelle sensibili. | +| Proposte persistenti, diff e versioni dei testi AI | **Escluse** | Resta un solo Generated Description modificabile. La cronologia dei run è operativa: non conserva prompt, output grezzo, proposta per colonna o audit della decisione umana. | +| Relazioni fisiche | **Adottate e implementate** | Sono constraint immutabili con coppie ordinate di colonne; ADR [0006](../adr/0006-separate-physical-and-logical-relationships.md). | +| Relazioni logiche curate o inferite | **Differite/aperto** | ADR-0006 riserva un modello e un lifecycle separati; non sono ancora parte del catalogo corrente. | +| Alias semantici, descrizioni dei valori, sinonimi e concetti | **Differiti/aperti** | `PROJECT_STATE.md` li assegna a slice dedicate; non vanno dedotti dai campi fisici già implementati. | +| Scope AI, ERD Mermaid e documentazione aggregata | **Differiti/aperti** | Restano capacità legacy inventariate, non una feature corrente del Metadata Catalog. Un eventuale lavoro dovrà avere contratto e gate propri. | +| Export CSV e script SQL dei commenti | **Differiti/aperti; varianti insicure escluse** | Non risultano nella API corrente. Export di segreti, CSV incoerenti e SQL ricostruito da tipi incompleti restano vietati dagli anti-pattern sotto. | +| Inferenza euristica e validazione di relazioni candidate | **Differita/aperta** | Non va confusa con la sincronizzazione delle FK fisiche; dipende dal futuro lifecycle delle relazioni logiche. | +| Pubblicazione del catalogo al core/schema-linking/Qdrant | **Differita e richiesta come design gate successivo** | Il runtime NL→SQL continua a usare configurazione e annotations del workspace; il cutover è esplicitamente rinviato in `PROJECT_STATE.md`. | +| Sette motori database del legacy | **Baseline superata** | La prima versione corrente supporta PostgreSQL; l'aggiunta di altri dialetti è una decisione futura, non parità automatica. | +| GDPR e import da installazioni ThothAI | **Fuori dalla baseline iniziale; import differito** | GDPR resta escluso. Un eventuale import richiede una iniziativa idempotente e un cutover separati, non il riuso degli ID Django. | +| Django Admin, modifica manuale della struttura sorgente, password nel catalogo, duplicazione opaca dei database | **Esclusi** | La UI e le API correnti amministrano il catalogo e non eseguono DDL sul DWH esterno; binding e segreti hanno ownership separata. | + +## Sintesi storica + +La baseline di parità proposta il 2026-08-23 comprende un catalogo dei database, l'inventario completo di tabelle, colonne e relazioni fisiche, metadati descrittivi modificabili, relazioni logiche, introspezione dello schema, generazione AI di descrizioni, scope, ERD Mermaid, documentazione ed esportazioni operative. La UI Django Admin è soltanto l'interfaccia legacy: non è un requisito architetturale da riprodurre. -Il flusso AI per le descrizioni è volutamente semplice e va mantenuto tale: +Il flusso AI per le descrizioni individuato come baseline è volutamente semplice: 1. l'AI scrive nel campo `generated_comment` della tabella o colonna; 2. l'utente seleziona gli elementi desiderati; @@ -193,7 +252,7 @@ scrive soltanto `column.generated_comment` Provider e modello AI provengono dalla configurazione globale del backend; la lingua è invece per database (`/Thoth/ThothAI/backend/thoth_core/thoth_ai/thoth_workflow/comment_generation_utils.py:156-219`). -### Comportamento di parità da implementare +### Comportamento di parità proposto il 2026-08-23 Il contratto funzionale minimo è esattamente questo: @@ -314,7 +373,7 @@ Il modulo contiene anche un helper verso `mermaid.ink`, ma non risultano call si legacy; il percorso attivo è il servizio locale. Non va quindi considerata una dipendenza funzionale da conservare. -## 9. Baseline di parità per il Metadata Catalog di ThothII +## 9. Baseline di parità proposta il 2026-08-23 ### Necessario diff --git a/docs/research/2026-08-23-postgresql-catalog-deployment-constraints.md b/docs/research/2026-08-23-postgresql-catalog-deployment-constraints.md index 87d5d137..698bfb45 100644 --- a/docs/research/2026-08-23-postgresql-catalog-deployment-constraints.md +++ b/docs/research/2026-08-23-postgresql-catalog-deployment-constraints.md @@ -1,219 +1,206 @@ # PostgreSQL deployment and failure-isolation constraints for the Metadata Catalog -Date: 2026-08-23 -Issue: [#8 — Assess PostgreSQL deployment and failure-isolation constraints](https://github.com/mptyl/ThothII/issues/8) +Original research: 2026-08-23 + +Last verified against the repository: 2026-08-31 + +Issue: [#8 — Assess PostgreSQL deployment and failure-isolation constraints](https://git.tylconsulting.it/mptyl/ThothII/issues/8) + +> **Status: partially superseded by ADR-0004.** This is historical research, not the current +> architecture contract. The recommendation to run a separate `catalog-api` process was rejected +> by [ADR-0004](../adr/0004-fastify-kysely-metadata-catalog.md). Operational constraints that do +> not depend on that process boundary remain useful, but the status table below is authoritative +> for what the 2026-08-31 code actually adopts, rejects, or leaves pending. ## Question How can a Metadata Catalog PostgreSQL service be deployed, backed up, diagnosed, and made -non-blocking for the core across local and server installations? +non-blocking for the NL→SQL workflow across local and server installations? -## Recommendation +## Current decision and implementation -Use PostgreSQL as the operational store, but place it behind a **separate Metadata Catalog API -process**. The SQL workflow `core` must not receive the catalog DSN, catalog credentials, a client -pool, or a Compose dependency on either the catalog API or PostgreSQL. +[ADR-0001](../adr/0001-postgres-metadata-catalog.md) selects PostgreSQL as the catalog authority. +[ADR-0003](../adr/0003-installation-local-database-bindings.md) keeps database bindings +installation-local. [ADR-0004](../adr/0004-fastify-kysely-metadata-catalog.md) then places the +catalog in the existing Fastify backend as an isolated Kysely module rather than a microservice. -The resulting dependency graph is: +The implemented dependency graph is therefore: ```mermaid flowchart LR UI[Frontend] - Core[Core workflow API] + Core[Core: Fastify workflow and catalog modules] + Pi[Pi and tht workflow] + PG[(catalog-db PostgreSQL)] Semantic[Qdrant and embedding] - Catalog[Metadata Catalog API] - PG[(Catalog PostgreSQL)] - Publish[Publication worker] - UI -->|workflow routes| Core + UI --> Core + Core --> Pi + Core --> PG Core --> Semantic - UI -->|catalog routes| Catalog - Catalog --> PG - PG --> Publish - Publish -->|immutable approved Publication| Semantic ``` -The core continues to use the last successfully published, revision-qualified semantic snapshot. -Unpublished catalog edits are never a runtime dependency. This matches the domain boundary: the -Metadata Catalog does not select workflow SQL elements, while a Publication is an immutable value -made available to consumers ([`CONTEXT.md`, lines 96–110](../../CONTEXT.md#L96-L110)). It also -preserves the existing rule that Qdrant is a derived index rather than a canonical source -([`PROJECT_STATE.md`, lines 405–410](../../PROJECT_STATE.md#L405-L410)). +The catalog module is isolated behind a repository interface, but it is not process-isolated. +`core` owns the catalog pool and Compose waits for `catalog-db` health before starting `core` +(`compose.yaml`, `backend/src/app.ts`, `backend/src/catalog/repository.ts`). Database-management records still do +not feed the NL→SQL handoff, so a running workflow does not read catalog rows; that cutover remains +future work (`PROJECT_STATE.md`). -## Evidence from the current architecture +## Verification result -1. The mandatory workflow stack is currently `frontend`, `core`, `qdrant`, `embedding`, and the - model-init job. `core` waits only for Qdrant and model initialization - ([`compose.yaml`, lines 50–60](../../compose.yaml#L50-L60)); the frontend waits only for `core` - ([`compose.yaml`, lines 123–141](../../compose.yaml#L123-L141)). Adding a catalog dependency to - either chain would enlarge the workflow failure domain. -2. `tht start` runs Compose and then checks an explicit allowlist containing only those five - mandatory services. Other Compose services are ignored by the readiness fold - ([`service.go`, lines 40–57](../../tools/tht/internal/service/service.go#L40-L57), - [`service.go`, lines 166–190](../../tools/tht/internal/service/service.go#L166-L190)). Separate - catalog services can therefore start with the installation without redefining core readiness. -3. `/health` is deliberately a process-liveness endpoint and does not probe external services - ([`app.ts`, lines 271–284](../../backend/src/app.ts#L271-L284)). Tests preserve a 200 response - even when PostgreSQL session storage is configured but not contacted - ([`health.test.ts`, lines 9–32](../../backend/test/health.test.ts#L9-L32)). Catalog readiness must - not be folded into this endpoint. -4. ThothII already has a sound server-PostgreSQL precedent: runtime and migrator credentials are - distinct; the one-shot migrator is profile-gated and is explicitly not a dependency of `core` - ([`compose.session-server.yaml.example`, lines 24–49](../../deploy/compose.session-server.yaml.example#L24-L49)). - Its documented failure behavior is route-scoped 503 while process liveness remains healthy - ([`PROJECT_STATE.md`, lines 621–636](../../PROJECT_STATE.md#L621-L636)). -5. Existing backup archives contain selected Compose volumes only. Local archives include seven - named volumes and server archives include only Qdrant and embedding volumes; no PostgreSQL - logical dump exists today ([`create.go`, lines 29–52](../../tools/tht/internal/backup/create.go#L29-L52)). - A catalog database therefore requires an explicit backup contract rather than an assumption - that current installation backup already covers it. - -## Deployment contract - -### Shared rules - -- Run `catalog-api` as a separate, non-privileged service. The frontend reverse proxy may route - `/api/catalog/*` to it, but frontend startup must not depend on catalog readiness. -- Give `catalog-api` a bounded pool and bounded operations. As an initial ceiling for this small - administrative workload: pool size 5, overflow 0–2, `connect_timeout=3`, `lock_timeout=2s`, and - a normal request `statement_timeout=10s`. Long AI calls must not hold database transactions. - PostgreSQL documents that an omitted or zero `connect_timeout` waits indefinitely, so an - explicit value is required for failure isolation - ([libpq connection parameters](https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-CONNECT-TIMEOUT)). -- Use three roles: `catalog_runtime` (CRUD and job claims only), `catalog_migrator` (DDL, used only - by a one-shot job), and `catalog_backup` (minimum privileges needed by the approved dump policy). - No role is shared with the DWH or `thoth_sessions`, and none is a PostgreSQL superuser. -- Store passwords and the CA as separate secret files. For non-local connections default to - `sslmode=verify-full`; this follows the current server-session configuration - ([`compose.session-server.yaml.example`, lines 5–20](../../deploy/compose.session-server.yaml.example#L5-L20)) - and PostgreSQL's hostname-verifying TLS guidance - ([libpq SSL parameters](https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-PARAMKEYWORDS)). -- Apply ordered, checksummed migrations under one transaction and an advisory lock. Refuse pending, - unknown, or checksum-drifted migrations in catalog readiness. The session repository already - implements this pattern - ([`postgres_repository.py`, lines 118–187](../../harness/tht/session/postgres_repository.py#L118-L187)); - the catalog should have its own migration table, lock key, schema, and roles. - -### Local installation - -Add an installation-owned `catalog-postgres` container on the private `thothii` network, with no -published host port and a dedicated `catalog-postgres-data` named volume. Pin the image by version -and digest, add a `pg_isready` healthcheck, and let only `catalog-api` use -`depends_on: condition: service_healthy`. `core` and `frontend` keep their existing dependency -graph. Docker confirms that `service_healthy` delays only the declaring dependent service, while -plain container start does not imply application readiness -([Compose startup order](https://docs.docker.com/compose/how-tos/startup-order/)). - -The local catalog services can be included in the normal installation Compose set because the -host CLI's mandatory-service allowlist excludes them. A separate `catalog-migrate` one-shot profile -retains explicit operator control; Docker profiles are intended for selectively activated and -one-off services ([Compose profiles](https://docs.docker.com/compose/how-tos/profiles/)). - -### Server installation - -Use an operator-provided PostgreSQL endpoint configured through a reviewed, untracked overlay. -Prefer a separate database and dedicated roles. A separate PostgreSQL cluster gives the strongest -resource and outage isolation; sharing the existing server cluster is acceptable only when the -operator accepts the residual cluster-wide blast radius and enforces connection limits, timeouts, -and separate ownership. Merely using another schema does not isolate connection exhaustion, -maintenance outages, WAL pressure, or storage failure. - -The server overlay adds secrets and catalog configuration only to `catalog-api` and the one-shot -migrator. It must not mount those secrets into `core`. This mirrors the current rule that `core` -never receives the session migrator credential -([`PROJECT_STATE.md`, lines 623–630](../../PROJECT_STATE.md#L623-L630)). - -## Failure behavior - -| Failure | Catalog behavior | Workflow behavior | +| Area | Status on 2026-08-31 | Evidence and consequence | | --- | --- | --- | -| PostgreSQL unavailable or connection pool exhausted | Catalog routes return sanitized `503 catalog_unavailable` with `Retry-After`; UI remains read-only/unavailable | Existing sessions and new SQL workflow requests continue against the last published snapshot | -| Catalog API process unavailable | `/api/catalog/*` fails; main frontend shell and workflow route remain usable | No effect | -| Pending or drifted migration | Catalog readiness fails and writes are refused | No effect | -| AI generation fails | Job records a retryable/terminal failure; no database transaction remains open | No effect | -| Publication indexing fails | Publication remains non-active/failed and can be retried idempotently | Previous active Publication remains in Qdrant | -| Qdrant unavailable | Publication is queued/failed without losing canonical catalog state | Existing core readiness rules already govern this independent failure | +| PostgreSQL as canonical catalog store | **Adopted** | ADR-0001 is implemented by the PostgreSQL-backed Kysely repository and the internal `catalog-db` service. | +| One Workspace Database per workspace and installation-local bindings | **Adopted** | ADR-0003 and the catalog migrations enforce the model; workspace identity remains in the workspace registry. | +| Separate `catalog-api` process | **Superseded/rejected** | ADR-0004 explicitly chooses an isolated module inside the existing Fastify process. There is no catalog microservice or separate catalog liveness endpoint. | +| No catalog pool, credential, or Compose dependency in `core` | **Superseded/rejected** | `core` owns the runtime pool, receives the runtime password secret, and declares `depends_on: catalog-db: service_healthy`. The old process-level isolation acceptance criterion is not current architecture. | +| Private PostgreSQL service and durable volume | **Adopted** | `catalog-db` uses a version-and-digest-pinned PostgreSQL 17.6 image, the private `thothii` network, no published host port, a `pg_isready`-based healthcheck, and `catalog-data`. Compose contract tests assert this topology (`scripts/test-default-compose.sh`). | +| Different local and server database topology | **Superseded/rejected** | Both current profiles inherit the same internal `catalog-db` and named volume. `deploy/compose.server.yaml` does not replace it with an operator-provided endpoint. | +| Runtime and migrator role separation | **Adopted** | `core` receives only `thothii_catalog_runtime`; the profile-gated `catalog-migrate` job receives only `thothii_catalog_migrate`. Bootstrap grants runtime DML/sequence privileges without DDL (`docker/catalog-db-init.sql`). | +| Dedicated backup role | **Pending** | There is no `catalog_backup` role or backup secret. | +| Explicit one-shot migrations | **Adopted** | `backend/src/catalog/migrate.ts` registers ordered Kysely migrations and uses a pool of one. `scripts/run-stack.sh` starts PostgreSQL and runs `catalog-migrate` before local startup; migrations are not hidden in backend startup. | +| Checksums, drift/pending refusal, and migration readiness | **Pending** | No repository-owned checksum policy, drift report, or readiness gate exists for catalog migrations. The backend can start without checking the Kysely migration head when launched outside the local helper. | +| Bounded runtime connection pool | **Adopted** | `backend/src/catalog/repository.ts` sets `max: 5` and `connectionTimeoutMillis: 3000`, and closes Kysely with the Fastify lifecycle. | +| `lock_timeout`, `statement_timeout`, and catalog TLS policy | **Pending** | The runtime pool does not set query or lock timeouts. Catalog connection configuration has no explicit CA/hostname-verification contract; the server profile still uses the private Compose network. | +| Process liveness independent of catalog queries | **Adopted** | `GET /health` returns `{status: "ok"}` without probing PostgreSQL, and `backend/test/health.test.ts` preserves that behavior. `/catalog/status` performs the catalog-specific availability check. | +| Stack startup independent of catalog availability | **Superseded/rejected** | Compose blocks `core` on healthy `catalog-db`. The host service health fold does not list `catalog-db`, but it cannot make `core` start while its Compose dependency is unhealthy. | +| Uniform catalog-outage response (`503` plus `Retry-After`) | **Pending** | Routes map the domain `CatalogUnavailableError` to a sanitized `503`, and an omitted catalog configuration uses an unavailable repository. PostgreSQL driver failures are not uniformly translated to that domain error, and no `Retry-After` contract is implemented. | +| Catalog-specific readiness and `tht doctor` checks | **Pending** | There is no `/health/ready` for the catalog and no doctor section for connection, migration head, pool saturation, backup age, or publication lag (`tools/tht/internal/doctor/report.go`). | +| Logical catalog backup and restore | **Pending** | Current backup archives omit `catalog-data` and do not run `pg_dump`; server archives include only Qdrant and embedding volumes. A backup can therefore succeed without preserving the Metadata Catalog (`tools/tht/internal/backup/create.go`). | +| Last-good Publication and catalog-to-Qdrant cutover | **Pending** | The current database-management slice does not change the NL→SQL runtime or publish catalog metadata to Qdrant. | -Do not silently fall back to JSON files, dual-write, or read unpublished PostgreSQL rows from the -core. Those paths would create split-brain state. The only degraded-mode contract is use of the -last successfully activated Publication. +## Current operational contract -## Backup and restore +### Deployment and credentials + +- Local and server Compose profiles currently use the same installation-owned `catalog-db` + container and `catalog-data` named volume. PostgreSQL is private to the Compose network. +- The database bootstrap login is the migrator. The init script creates the separate runtime login + from a Docker secret and grants only runtime DML and sequence access. +- Runtime and migrator passwords are separate protected host files exposed as separate Docker + secrets. Neither value belongs in tracked environment files + (`deploy/env/local.env.example`, `deploy/env/server.env.example`). +- The Fastify catalog repository uses a five-connection pool with a three-second connection + timeout. Fastify closes the repository pool during shutdown. + +### Migrations + +The compiled `catalog-migrate` entry point owns schema changes and uses the migrator credential. +The current ordered series is under +`backend/src/catalog/migrations/`. The local launcher runs +it before the normal stack, while production rollout must invoke the profile-gated service +explicitly. + +This is weaker than the original recommendation in two ways: there is no catalog readiness check +for pending or unknown migrations, and the repository does not maintain content checksums for +migration drift. Until those checks exist, “the migrator completed” is the available deployment +gate; the application itself does not prove migration compatibility. + +### Failure behavior actually provided + +| Failure | Current catalog behavior | Current workflow behavior | +| --- | --- | --- | +| Catalog configuration omitted when launching the backend directly | Fastify uses `UnavailableCatalogRepository`; `/catalog/status` reports unavailable and domain-mapped catalog operations return sanitized `503` | `/health`, sessions, and SSE remain available | +| `catalog-db` unhealthy before Compose startup | `core` is not started because its dependency is not healthy | Workflow startup is blocked | +| PostgreSQL becomes unavailable after startup | `/health` remains process-only and `/catalog/status` reports unavailable; route errors are sanitized, but a uniform `503`/`Retry-After` mapping is not guaranteed | Existing session code does not read catalog rows, but both surfaces still share one Fastify process | +| Migrations are pending or incompatible | No dedicated readiness refusal exists; affected catalog operations fail | No catalog-to-workflow handoff exists yet, but local deployment correctness depends on running `catalog-migrate` first | +| Catalog backup is requested through current `tht backup` | No PostgreSQL dump is added and `catalog-data` is omitted | The archive may succeed while being unable to restore catalog state | + +The shared process means resource exhaustion, fatal process errors, and startup hooks remain a +common failure domain even though repository calls are separated. Conversely, placing the module +inside Fastify does not require workflow code to consume catalog rows: preserving that data-flow +boundary is still the useful part of the original isolation recommendation. + +## Historical recommendations that remain valid backlog + +The following constraints survive ADR-0004 because they can be implemented inside the current +Fastify deployment: + +1. **Bound every database operation.** Keep the adopted pool and connection timeout; add explicit + request query and lock timeouts. Long model calls and source sampling must not hold catalog + transactions. +2. **Make failures component-specific.** Translate connection, timeout, and pool-exhaustion errors + into one sanitized catalog-unavailable response, add bounded retry guidance, and keep `/health` + process-only. +3. **Make migration compatibility observable.** Report applied, pending, unknown, and drifted + migrations through a catalog readiness check and `tht doctor`; do not put migrator credentials + in `core`. +4. **Define a server TLS topology before externalizing PostgreSQL.** If the server profile moves to + an operator-provided endpoint, use a dedicated database and roles, protected CA material, and + hostname verification. Sharing a cluster leaves connection, maintenance, WAL, and storage blast + radius even when schemas are separate. PostgreSQL documents the relevant connection and TLS + parameters in its + [connection parameter reference](https://www.postgresql.org/docs/current/libpq-connect.html). +5. **Keep derived semantic data non-canonical.** When catalog publication is implemented, activate + complete immutable revisions and retain the last successfully activated revision rather than + exposing mutable catalog rows to the workflow. + +## Backup and restore target + +The original logical-backup recommendation is still valid and is now a confirmed implementation +gap. ### Backup -Extend `tht backup create` or add a catalog-specific command invoked by it. Create a logical -`pg_dump --format=custom` archive through a one-shot helper, checksum it, and place it in the -installation archive manifest. Do **not** tar a live PostgreSQL data volume: the existing raw-volume -backup mechanism was designed for the current named volumes, not database crash consistency. +Extend `tht backup create` with a catalog step that runs `pg_dump --format=custom` through a +one-shot helper and adds the dump plus checksum to the installation manifest. Do not treat a tar of +a live data volume as a PostgreSQL consistency contract. PostgreSQL documents that `pg_dump` +creates a consistent export while the database remains in use and that custom format supports +selective restore ([`pg_dump`](https://www.postgresql.org/docs/current/app-pgdump.html)). -`pg_dump` produces a consistent export while the database remains in use, and custom format is -compressed and supports selective/reordered restore -([PostgreSQL `pg_dump`](https://www.postgresql.org/docs/current/app-pgdump.html)). Because metadata -volume is small, favor correctness over parallelism. Run a backup before every migration and on an -operator-defined schedule; record the server version, migration head, dump checksum, creation time, -and catalog Publication identifiers in the manifest. Never write a dump into the workspace Git -repository. - -For an external server database, the operator must supply the dedicated backup credential through -the same secret-file boundary. A successful filesystem/Qdrant archive with a failed catalog dump -is an incomplete installation backup and must not be reported as success. +Use a dedicated least-privilege backup role, never write dumps into a workspace repository, and +record at least the server version, migration head, checksum, creation time, and catalog revision +identifiers. A filesystem/Qdrant archive with a failed or absent catalog dump must be reported as +incomplete. ### Restore -Restore only during an explicit catalog maintenance window into an empty target database (or a -deliberately cleaned target), using `pg_restore --exit-on-error --single-transaction`. Validate the -archive checksum first, then require: +Restore only through an explicit maintenance operation into an empty or deliberately cleaned +target. Validate the archive checksum, run `pg_restore --exit-on-error --single-transaction`, then +verify migration compatibility, referential integrity, bounded entity counts, and an authenticated +catalog smoke test. The controls are documented by PostgreSQL +([`pg_restore`](https://www.postgresql.org/docs/current/app-pgrestore.html)). Keep the old database +until validation passes; rollback should switch the endpoint or retained volume, not dual-write. -1. expected migration history with no checksum drift; -2. referential-integrity checks and bounded entity counts; -3. Publication content hashes matching the restored records; -4. an authenticated catalog smoke test. +## Diagnostics target -The relevant restore controls and their behavior are documented by -[PostgreSQL `pg_restore`](https://www.postgresql.org/docs/current/app-pgrestore.html). After restore, -rebuild catalog-derived Qdrant content from the restored active Publication rather than treating -the vector index as the source of truth. Keep the old database untouched until this validation -passes; rollback is endpoint reversal, not dual-write. +- Keep the shared `GET /health` endpoint process-only. +- Treat `/catalog/status` as the current minimal availability surface; add a bounded catalog + readiness check covering connection and migration compatibility before using it as a rollout + gate. +- Add read-only `tht doctor` checks for secret-file presence and permissions, PostgreSQL + reachability, authentication, migration state, pool saturation, and last successful catalog + backup. Sanitize all DSNs and errors. +- Use `pg_isready` only for server transport readiness. Its result does not prove schema or runtime + authorization correctness + ([`pg_isready`](https://www.postgresql.org/docs/current/app-pg-isready.html)). +- Add structured metrics for connection acquisition failures, pool use, query/lock timeout, + migration head, background-job backlog, and backup age. Do not log connection strings, source + samples, prompts, or generated descriptions at info level. -## Diagnostics and observability +## Recommended closure criteria -- Keep `core /health` unchanged. Give `catalog-api` separate `/health/live` (process only) and - `/health/ready` (bounded connection, `SELECT 1`, and migration compatibility) endpoints. -- Add a named `metadata-catalog` section to `tht doctor`: configuration completeness; secret-file - presence and permissions; DNS/TCP/TLS/authentication; `SELECT 1`; migration applied/pending/drifted - state; pool saturation; last successful backup age; and oldest pending Publication. Checks are - read-only and all errors/DSNs are sanitized. The current doctor already separates Compose service - status from container-local workflow checks - ([`report.go`, lines 275–310](../../tools/tht/internal/doctor/report.go#L275-L310)). -- Use `pg_isready` only for transport readiness; its exit codes distinguish accepting, rejecting, - no-response, and invalid-parameter states, but it does not prove schema or authorization - correctness ([PostgreSQL `pg_isready`](https://www.postgresql.org/docs/current/app-pg-isready.html)). -- Expose structured metrics/logs for connection acquisition failures, pool utilization, query - timeout/lock timeout, migration head, AI-job backlog, Publication lag, and backup age. Never log - connection strings, prompts containing schema metadata, or generated descriptions at info level. -- `tht status` may show the catalog as `healthy`, `degraded`, or `disabled`, but catalog state must - not change the existing core health result. `tht doctor` may return a failing catalog check for - operator visibility without stopping or restarting the workflow stack. - -## Acceptance constraints for later design and implementation - -1. Killing `catalog-postgres` and `catalog-api` must not change `core /health`, stop Pi, interrupt an - existing SQL session, or prevent a new session that uses the already active Publication. -2. No catalog database secret, client dependency, migration, or network call exists in the harness - workflow path. -3. Catalog failure is explicit on catalog routes; no JSON fallback or dual-write is introduced. -4. A failed Publication leaves the previous active semantic content addressable and retryable. -5. Local and server backup tests perform a real dump/restore round trip and prove that a backup is - incomplete when the PostgreSQL dump fails. -6. Migration credentials are absent from runtime containers; drift and pending migrations are - visible in readiness and `tht doctor`. +1. Decide whether ADR-0004's statement that catalog unavailability does not make sessions or SSE + unavailable must also hold at Compose startup. If yes, remove or soften the hard `core` → + `catalog-db` startup dependency without reintroducing a microservice. +2. Prove a configured PostgreSQL outage produces the same sanitized catalog response across every + catalog route while `/health`, session creation, and SSE continue to work. +3. Add catalog migration compatibility to readiness and `tht doctor`, including pending, unknown, + and drifted states. +4. Add a real `pg_dump`/`pg_restore` round-trip to local and server backup tests, and fail backup + publication when the catalog dump is absent or fails. +5. Before enabling a remote server catalog, define TLS verification, credential files, connection + limits, and the accepted cluster-level blast radius. +6. Before the NL→SQL cutover, define and test an immutable last-good Publication boundary; do not + dual-read mutable PostgreSQL rows and legacy metadata as competing authorities. ## Decision summary -The Metadata Catalog should be a separately deployed bounded context backed by PostgreSQL. Local -installations own a private container and volume; server installations use a TLS-verified external -database, preferably on a separate cluster when strict blast-radius isolation is required. Logical -dump/restore, one-shot migrations, role separation, bounded connections, component-specific -diagnostics, and last-good-Publication semantics make PostgreSQL operationally safe without making -it a dependency of the SQL workflow. +PostgreSQL, installation-local bindings, private deployment, role separation, one-shot migrations, +bounded pooling, and process-only liveness are implemented. The separate `catalog-api` process and +the claim that `core` has no catalog dependency are not current design: ADR-0004 chose the existing +Fastify process, and Compose currently blocks `core` startup on `catalog-db` health. Logical +backup/restore, catalog migration readiness and drift detection, consistent outage mapping, +catalog-specific doctor checks, server TLS/external topology, and last-good publication remain +pending. Those open constraints should be treated as backlog, not as capabilities already provided +by the repository. diff --git a/docs/research/2026-08-23-thothii-metadata-publication-qdrant-seams.md b/docs/research/2026-08-23-thothii-metadata-publication-qdrant-seams.md index 51cc13ef..5097b522 100644 --- a/docs/research/2026-08-23-thothii-metadata-publication-qdrant-seams.md +++ b/docs/research/2026-08-23-thothii-metadata-publication-qdrant-seams.md @@ -4,21 +4,51 @@ session-pinning, and runtime read-only contracts constrain metadata publication without changing the NL→SQL workflow? +**Last verified:** 2026-08-31 + +**Validity:** Active architectural research; the catalog-to-core integration described below is +still **deferred**, not implemented. + +**Current authorities:** `PROJECT_STATE.md`, +[`workspace-evidence-v3.md`](../contracts/workspace-evidence-v3.md), +[`workspace-preprocessing-cli.md`](../contracts/workspace-preprocessing-cli.md), +[`tht-dwh.md`](../contracts/tht-dwh.md), +[`ADR-0001`](../adr/0001-postgres-metadata-catalog.md), and +[`ADR-0004`](../adr/0004-fastify-kysely-metadata-catalog.md). These sources and current code +override this research note if they diverge. + +## Revalidation against the current implementation + +| Finding | Status on 2026-08-31 | Current evidence and consequence | +| --- | --- | --- | +| PostgreSQL metadata authority in the existing Fastify backend | **Adopted/current** | ADR-0001 and ADR-0004 are implemented; the catalog is the management-plane authority. | +| Git workspace revision, immutable registry snapshot, and revision-pinned runtime | **Adopted/current** | Registry activation and the Workspace Evidence v3 contract still provide the publication boundary for runtime-owned files. | +| Qdrant read isolation for schema and Evidence by `workspace_id` + `workspace_revision` | **Adopted/current** | `QdrantVectorStore.search()` applies `_revision_filter()` to schema/Evidence; Memory and solved questions intentionally remain workspace-wide (`harness/tht/adapters/vector/qdrant.py`). | +| Physical schema as a file in the Git publication | **Superseded clarification** | `physical.yaml` is owned by the immutable `.tht-dwh` generation selected by `ACTIVE`, not by the Git workspace revision ([DWH contract](../contracts/tht-dwh.md)). Git currently supplies the revision-pinned curated `schema/annotations.yaml`. | +| Catalog-to-core publisher | **Open/deferred** | No current route or service projects catalog records into core artifacts. `PROJECT_STATE.md` explicitly defers the schema-linking integration (`PROJECT_STATE.md:160-171`). | +| Explicit Core Schema Selection | **Open/deferred** | The workspace descriptor selects one database and one physical schema, but has no table/column allowlist (`backend/src/workspaces/schema.ts`); no catalog selection is handed to `tht`. | +| Cross-revision hash lookup used by schema synchronization | **Open defect** | Reads are revision-filtered, but `existing_hashes()` is not. A same-key/same-content point from an older revision can suppress the required upsert into the new revision (`harness/tht/adapters/vector/qdrant.py`, `harness/tht/cli/vector_cmd.py`). | +| Deletion/GC | **Current for Evidence; open for schema** | Evidence has generation inventory, retention, compensation, and exact-generation deletion. `sync_canonical_records()` never deletes schema records absent from the new canonical set, and no revision-retention GC exists for schema points. | +| Annotation consumption | **Current but incomplete** | M-Schema rendering and schema embeddings consume `Annotations`; `tht schema columns` still returns only physical comments, so F4 does not display annotation descriptions (`harness/tht/cli/schema_cmd.py`, `harness/.pi/extensions/tht-gate.js`). | + ## Conclusion -The lowest-impact publication seam is **not** a new writer inside the NL→SQL workflow and it is -not a direct CRUD-to-Qdrant path. The existing core already consumes two canonical schema inputs: -an introspected `PhysicalSchema` and a curated `Annotations` document. An explicit publication -operation should project the approved, workspace-selected subset of the new metadata catalog into -those inputs, bind the projection to a new Git `workspace_revision`, activate that immutable +The lowest-impact **proposed** publication seam is not a new writer inside the NL→SQL workflow and +is not a direct CRUD-to-Qdrant path. The existing core consumes an introspected `PhysicalSchema` +from the active immutable DWH generation and a curated, Git-revision-pinned `Annotations` +document. A future explicit publication operation should project the approved catalog subset into +a new Core Schema Selection contract and the curated annotations, activate a new immutable Git revision, and then invoke the existing `workspace index-schema` preprocessing operation. Qdrant -remains a derived, rebuildable projection. +remains a derived, rebuildable projection. None of this catalog-to-core handoff exists yet; +`PROJECT_STATE.md:160-164` deliberately defers it to the next design +gate. -This preserves the existing runtime path: +The proposed flow would preserve the existing runtime path: ```text approved catalog data - -> workspace Git revision (physical schema selection + annotations) + -> explicit Core Schema Selection + curated annotations projection (future) + -> workspace Git revision (annotations and selection contract; physical.yaml stays DWH-owned) -> immutable registry snapshot -> revision-bound runtime configuration -> existing tht vector index-schema @@ -26,23 +56,23 @@ approved catalog data -> existing retrieval_pack / schema render / F4 review ``` -The complete database inventory may remain in the Metadata Catalog. Only the explicit Core Schema -Selection should be projected into the workspace artifacts used by the SQL workflow. ThothII does -not currently model such a table-level selection in its descriptor, so that projection contract is -new work. +The complete database inventory remains in the Metadata Catalog. Only an explicit Core Schema +Selection and approved semantic fields should be projected into the artifacts used by the SQL +workflow. ThothII does not currently model or publish that table/column-level selection, so both +the projection contract and its operational publisher are new work. ## 1. Current authority and publication boundary The workspace descriptor identifies one PostgreSQL database and one physical schema, plus one workspace-owned Qdrant collection; it has no table or column allowlist -([`backend/src/workspaces/schema.ts:160-178`](../../backend/src/workspaces/schema.ts#L160-L178)). +(`backend/src/workspaces/schema.ts`). Consequently, a full-database metadata catalog and the subset eligible for the SQL core cannot be represented as the same current descriptor object. The existing public contract makes the Git workspace repository curator-owned. Changes occur in a separate authoring clone followed by installation pull, and the API does not write workspace, schema, or Evidence paths -([`docs/contracts/workspace-evidence-v3.md:148-156`](../contracts/workspace-evidence-v3.md#L148-L156)). +([Workspace Evidence v3 contract](../contracts/workspace-evidence-v3.md)). Therefore a browser CRUD service cannot silently make its PostgreSQL state authoritative for the core without either: @@ -54,10 +84,12 @@ The first option preserves current architecture and session reproducibility. Registry activation validates every descriptor at one exact commit, validates and synchronizes the commit's `schema/annotations.yaml`, and rejects duplicate ownership of a Qdrant collection -([`backend/src/workspaces/registry.ts:531-576`](../../backend/src/workspaces/registry.ts#L531-L576)). +(`backend/src/workspaces/registry.ts`, +`backend/src/workspaces/git-repository.ts`, +`backend/src/workspaces/annotations-sync.ts`). It writes the candidate snapshot into a staging directory, records file digests in `snapshot.json`, atomically renames the directory, and only then moves active state -([`backend/src/workspaces/registry.ts:582-629`](../../backend/src/workspaces/registry.ts#L582-L629)). +(`backend/src/workspaces/registry.ts`). That is the existing atomic publication boundary to reuse. ## 2. Canonical schema inputs already consumed by the core @@ -67,15 +99,14 @@ The harness separates source facts from curated semantics: - `PhysicalSchema` contains database/schema identity and tables; table facts include comments, columns, physical foreign keys, and indexes; columns contain type, nullability, primary-key, default, source comment, examples, and eligibility - ([`harness/tht/mschema/models.py:23-64`](../../harness/tht/mschema/models.py#L23-L64)). + (`harness/tht/mschema/models.py`). - `Annotations` contains curated table descriptions, concepts and notes, column descriptions, synonyms, concepts, evidence, notes and eligibility overrides, plus logical foreign keys - ([`harness/tht/mschema/models.py:67-88`](../../harness/tht/mschema/models.py#L67-L88)). + (`harness/tht/mschema/models.py`). Rendering already gives annotations precedence over source comments and merges physical and logical foreign keys. It also applies column eligibility before producing M-Schema context -([`harness/tht/mschema/render.py:9-43`](../../harness/tht/mschema/render.py#L9-L43), -[`harness/tht/mschema/render.py:46-80`](../../harness/tht/mschema/render.py#L46-L80)). This makes +(`harness/tht/mschema/render.py`). This makes `Annotations` the natural narrow projection target for approved descriptions, synonyms, logical relationships, and eligibility from the new catalog. @@ -83,14 +114,14 @@ There are two current compatibility gaps: - `schema_records()` embeds table descriptions/concepts and column descriptions/synonyms/examples, but does not include physical or logical foreign keys in vector record content - ([`harness/tht/vectorstore/records.py:101-132`](../../harness/tht/vectorstore/records.py#L101-L132)). + (`harness/tht/vectorstore/records.py`). Relationships still reach the model through deterministic schema rendering, not through schema candidate embeddings. - The F4 widget loads columns with `tht schema columns` - ([`harness/.pi/extensions/tht-gate.js:1220-1254`](../../harness/.pi/extensions/tht-gate.js#L1220-L1254)), - but that command currently returns only `PhysicalSchema.comment` values and does not merge + (`harness/.pi/extensions/tht-gate.js`), but that + command still returns only `PhysicalSchema.comment` values and does not merge `Annotations` - ([`harness/tht/cli/schema_cmd.py:592-621`](../../harness/tht/cli/schema_cmd.py#L592-L621)). + (`harness/tht/cli/schema_cmd.py`). A publisher that writes only annotations would improve vector search and rendered M-Schema but not the table/column descriptions displayed by this existing reviewer widget. Fixing the command to use the existing merged description helpers would preserve the workflow shape while closing @@ -101,32 +132,33 @@ There are two current compatibility gaps: `workspace index-schema` is already the supported operator seam. It creates a revision-bound runtime, checks collection compatibility, and runs the harness command `vector index-schema --json` -([`backend/src/workspaces/preprocessing-service.ts:312-330`](../../backend/src/workspaces/preprocessing-service.ts#L312-L330)). +(`backend/src/workspaces/preprocessing-service.ts`, +[`workspace-preprocessing-cli.md`](../contracts/workspace-preprocessing-cli.md)). The full preprocessing operation performs DWH preparation, FK suggestion/review, schema indexing, and optional Evidence preprocessing as separate resumable stages -([`backend/src/workspaces/preprocessing-service.ts:367-450`](../../backend/src/workspaces/preprocessing-service.ts#L367-L450)). +(`backend/src/workspaces/preprocessing-service.ts`). Metadata-only publication should normally use the narrow `index-schema` operation after its workspace artifacts are valid, rather than coupling catalog CRUD to the full pipeline. The harness indexer reads the active immutable physical schema and revision-specific annotations, constructs schema records, embeds only changed content, and writes them through the configured vector adapter -([`harness/tht/cli/vector_cmd.py:100-140`](../../harness/tht/cli/vector_cmd.py#L100-L140)). Its +(`harness/tht/cli/vector_cmd.py`). Its machine result carries the physical/annotation artifact digests, workspace revision, collection, and counts -([`harness/tht/cli/vector_cmd.py:157-179`](../../harness/tht/cli/vector_cmd.py#L157-L179)). This is +(`harness/tht/cli/vector_cmd.py`). This is the right place to retain publication evidence and audit linkage. Preprocessing state is already revision- and binding-aware. A resumed job must match operation, workspace revision, descriptor/catalog blobs, runtime config and binding digests or it fails with `preprocessing_resume_mismatch` -([`backend/src/workspaces/preprocessing-state.ts:262-313`](../../backend/src/workspaces/preprocessing-state.ts#L262-L313)). +(`backend/src/workspaces/preprocessing-state.ts`). There is an important operational gate: every preprocessing operation calls `assertSessionInventoryCompatible`; a non-finalized, non-archived session pinned to another revision blocks preprocessing -([`backend/src/workspaces/preprocessing-service.ts:459-472`](../../backend/src/workspaces/preprocessing-service.ts#L459-L472), -[`backend/src/workspaces/preprocessing-state.ts:363-378`](../../backend/src/workspaces/preprocessing-state.ts#L363-L378)). +(`backend/src/workspaces/preprocessing-service.ts`, +`backend/src/workspaces/preprocessing-state.ts`). A catalog publication UX must expose this as a pending/blocking condition rather than report a generic indexing failure. @@ -135,23 +167,22 @@ generic indexing failure. The collection contract is fixed at the descriptor's dimensions/distance and eight keyword payload indexes: `content_hash`, `document_id`, `kind`, `record_key`, `record_kind`, `vector_generation`, `workspace_id`, and `workspace_revision` -([`backend/src/workspaces/qdrant-collection.ts:1-20`](../../backend/src/workspaces/qdrant-collection.ts#L1-L20)). +(`backend/src/workspaces/qdrant-collection.ts`). `self_heal` may create a missing compatible collection or indexes; `require_existing` only validates and refuses an incompatible collection -([`backend/src/workspaces/qdrant-collection.ts:71-104`](../../backend/src/workspaces/qdrant-collection.ts#L71-L104)). +(`backend/src/workspaces/qdrant-collection.ts`). The publication path must use this shared manager instead of inventing collection setup. Schema payloads include both the workspace and workspace revision, record identity, semantic kind, content and content hash -([`harness/tht/vectorstore/records.py:30-49`](../../harness/tht/vectorstore/records.py#L30-L49)). +(`harness/tht/vectorstore/records.py`). Qdrant point IDs for schema and Evidence also include the revision, and upserts use the same revision in the payload -([`harness/tht/adapters/vector/qdrant.py:42-46`](../../harness/tht/adapters/vector/qdrant.py#L42-L46), -[`harness/tht/adapters/vector/qdrant.py:196-220`](../../harness/tht/adapters/vector/qdrant.py#L196-L220)). +(`harness/tht/adapters/vector/qdrant.py`). Reads always filter by `workspace_id`; schema and Evidence reads additionally filter by the bound `workspace_revision`, while memory and solved-question records intentionally remain workspace-wide -([`harness/tht/adapters/vector/qdrant.py:291-308`](../../harness/tht/adapters/vector/qdrant.py#L291-L308)). +(`harness/tht/adapters/vector/qdrant.py:125-145`). This means an approved semantic change needs a new workspace revision if old sessions must retain their previous view. Directly overwriting points under the same Git revision would mutate the @@ -160,43 +191,52 @@ would require changing the current runtime filter contract. ### Qdrant synchronization defect to resolve before catalog publication -The current incremental synchronizer compares canonical records by content hash and upserts +The current incremental schema synchronizer compares canonical records by content hash and upserts changed records, but it has no deletion step -([`harness/tht/cli/vector_cmd.py:67-89`](../../harness/tht/cli/vector_cmd.py#L67-L89)). More +(`harness/tht/cli/vector_cmd.py:77-99`). More importantly, `QdrantVectorStore.existing_hashes()` filters by workspace and record kind but does not apply `_revision_filter()` -([`harness/tht/adapters/vector/qdrant.py:174-194`](../../harness/tht/adapters/vector/qdrant.py#L174-L194)), +(`harness/tht/adapters/vector/qdrant.py:266-286`), even though point IDs and reads are revision-scoped. Therefore a same-key/same-content record from an older revision may be classified as unchanged and never written under the new revision. This is an implementation defect/risk inferred directly from the two code paths, and it should be fixed and regression-tested before the catalog relies on `index-schema` for multi-revision publication. +Deletion/GC must be stated per record family. Evidence cleanup is implemented: the corpus pipeline +retains the configured number of published generations, protects active/job-referenced +generations, and calls exact-generation deletion for evicted or compensated generations +(`harness/tht/evidence/corpus/pipeline.py`, +`harness/tht/adapters/vector/qdrant.py:350-390`). Schema cleanup is +not implemented: `delete_kinds()` exists as a workspace-scoped adapter primitive, but the schema +synchronizer never calls it, it is not revision-scoped, and there is no retention policy for old +schema revisions. Publication therefore still needs exact current-revision deletion semantics and +separate safe GC for unleased historical schema revisions. + ## 5. Session pinning and why the SQL workflow can remain unchanged Normal session creation acquires an immutable registry revision, passes its snapshot path, workspace ID and commit to `tht session new`, and only releases the retention lease after the manifest has been persisted -([`backend/src/routes/sessions.ts:341-367`](../../backend/src/routes/sessions.ts#L341-L367), -[`backend/src/routes/sessions.ts:423-440`](../../backend/src/routes/sessions.ts#L423-L440)). The +(`backend/src/routes/sessions.ts`). The manifest stores `workspace_id` and `workspace_revision` next to database/schema identity -([`harness/tht/session/store.py:101-132`](../../harness/tht/session/store.py#L101-L132)). Resume and +(`harness/tht/session/store.py`). Resume and saved-SQL paths reopen that exact retained snapshot rather than the current installation default -([`backend/src/routes/sessions.ts:184-198`](../../backend/src/routes/sessions.ts#L184-L198), -[`backend/src/routes/sql.ts:24-43`](../../backend/src/routes/sql.ts#L24-L43)). +(`backend/src/routes/sessions.ts`, +`backend/src/routes/sql.ts`). The runtime renderer places the same workspace revision in `runtime_identity`, points the harness at the descriptor-owned Qdrant collection, and supplies the internal embedding service -([`backend/src/workspaces/runtime-renderer.ts:248-277`](../../backend/src/workspaces/runtime-renderer.ts#L248-L277)). +(`backend/src/workspaces/runtime-renderer.ts`). The vector adapter is constructed directly from that configuration, including the revision -([`harness/tht/adapters/factory.py:27-42`](../../harness/tht/adapters/factory.py#L27-L42)). +(`harness/tht/adapters/factory.py`). At session bootstrap the backend invokes the existing `search pack` command -([`backend/src/routes/sessions.ts:465-470`](../../backend/src/routes/sessions.ts#L465-L470)). That +(`backend/src/routes/sessions.ts`). That command queries schema records with the existing schema kinds, ranks tables and persists the candidate list -([`harness/tht/cli/search_cmd.py:291-310`](../../harness/tht/cli/search_cmd.py#L291-L310)). The Pi +(`harness/tht/cli/search_cmd.py`). The Pi extension reads the persisted retrieval pack through the CLI -([`harness/.pi/extensions/tht-gate.js:61-75`](../../harness/.pi/extensions/tht-gate.js#L61-L75)), +(`harness/.pi/extensions/tht-gate.js`), and F4 starts from those candidates while loading full table/column context through existing schema commands. Thus catalog publication can improve the inputs without changing phases, gate semantics, or persisted session artifacts. @@ -204,17 +244,18 @@ or persisted session artifacts. ## 6. DWH reuse across metadata-only revisions DWH preparation uses immutable generations selected by an `ACTIVE` pointer -([`docs/contracts/tht-dwh.md:18-24`](../contracts/tht-dwh.md#L18-L24)). The effective-configuration +([DWH contract](../contracts/tht-dwh.md)). The effective-configuration identity deliberately excludes `runtime_identity`, so a content-only Git revision does not force a database re-introspection -([`docs/contracts/tht-dwh.md:74-90`](../contracts/tht-dwh.md#L74-L90)). This is the key enabling -property for metadata publication: a new revision can carry updated curated annotations/Core Schema -Selection, reuse the compatible physical DWH generation, and rebuild only the revision-scoped -schema projection. +([DWH contract](../contracts/tht-dwh.md)). This is the key enabling property for future metadata +publication: a new Git revision can carry updated curated annotations and a future Core Schema +Selection contract, reuse the compatible DWH-owned `physical.yaml`, and rebuild only the +revision-scoped schema projection. Today only the curated annotations part of that statement exists. ## 7. Constraints for the Metadata Catalog design -The following should be treated as requirements for the architecture map: +The following remain proposed requirements for the deferred catalog-to-core design gate; they are +not claims about current implementation: 1. **Separate full inventory from core projection.** PostgreSQL may hold the complete database catalog, drafts and AI-generated text. Only an explicit Core Schema Selection and approved @@ -225,10 +266,10 @@ The following should be treated as requirements for the architecture map: 3. **Use a new Git revision as the publication identity.** This preserves current snapshot, retention, session resume and Qdrant filtering semantics. A separate mutable catalog revision cannot be safely introduced without changing runtime reads. -4. **Reuse canonical formats.** Project physical facts/Core Schema Selection into a validated - `PhysicalSchema` view and semantic edits into `Annotations`; keep Mermaid and long-form database - documentation outside Qdrant unless a separate, explicit record kind and retrieval policy is - designed. +4. **Respect artifact ownership.** Keep introspected physical facts in the immutable DWH generation; + project the future Core Schema Selection through an explicit contract and approved semantic + edits into `Annotations`. Keep Mermaid and long-form database documentation outside Qdrant + unless a separate record kind and retrieval policy is designed. 5. **Reuse the operator boundary.** Trigger `workspace index-schema`, observe its schema-versioned result and persist its artifact identities. Do not call Qdrant from browser CRUD handlers. 6. **Keep runtime read-only.** The session/Pi process continues to read the pinned snapshot, @@ -236,14 +277,15 @@ The following should be treated as requirements for the architecture map: management control plane. 7. **Surface publication gates.** UI status must distinguish Git activation, incompatible/missing collection, resumable-session revision conflict, embedding failure and completed publication. -8. **Repair and test cross-revision synchronization first.** Scope `existing_hashes()` to the bound - revision and define deletion/garbage-collection semantics before depending on repeated catalog - publication. +8. **Repair and test cross-revision schema synchronization first.** Scope `existing_hashes()` to + the bound revision, delete records removed from the current revision's canonical schema, and + define safe schema-revision GC. Reuse rather than duplicate the already implemented Evidence + generation GC. 9. **Close the annotation display gap without changing the workflow.** Make `schema columns` read the same merged descriptions used by M-Schema/vector rendering, so the current F4 widget sees the approved catalog text. -## Decision summary for the Wayfinder map +## Proposed decision summary for the Wayfinder map - Keep the Metadata Catalog as a separate management subsystem and source of editable metadata. - Keep Qdrant derived and revision-scoped; it is not the catalog database or source of truth. @@ -253,3 +295,7 @@ The following should be treated as requirements for the architecture map: the current descriptor. - Treat direct same-revision Qdrant writes, implicit CRUD publication, and bypassing the Git snapshot boundary as rejected integration paths. + +This summary remains design input. The authoritative current state is that catalog-to-core +integration and Sensitive Data Policy enforcement in schema-linking are deferred in +`PROJECT_STATE.md:160-171`.