docs: revalidate metadata catalog research

This commit is contained in:
Codex
2026-08-31 14:54:52 +02:00
parent 16b7477861
commit 866aee4249
3 changed files with 345 additions and 253 deletions
@@ -1,21 +1,80 @@
# Inventario delle funzionalità legacy di gestione metadati di ThothAI # Inventario delle funzionalità legacy di gestione metadati di ThothAI
Data: 2026-08-23 Data dell'inventario: 2026-08-23
Issue: `mptyl/ThothII#6` Ultima verifica rispetto a ThothII: 2026-08-31
Stato: **inventario storico verificato; non è una specifica dello stato corrente**.
Issue originaria: [mptyl/ThothII#6](https://git.tylconsulting.it/mptyl/ThothII/issues/6)
Fonte primaria: repository legacy annidato `Thoth/ThothAI`, commit Fonte primaria: repository legacy annidato `Thoth/ThothAI`, commit
`55855de0f18e5cb4bc72f0a2ab0a7317995186dd`. `55855de0f18e5cb4bc72f0a2ab0a7317995186dd`.
## Sintesi Le citazioni che iniziano con `/Thoth/ThothAI/` sono relative alla radice di quella copia legacy
fissata al commit indicato. Alla verifica del 2026-08-31 tutti i 68 riferimenti univoci puntavano a
file esistenti e a intervalli di righe validi.
La parità funzionale richiesta a ThothII comprende un catalogo dei database, l'inventario > Questo documento conserva l'inventario e le evidenze degli anti-pattern di ThothAI. Le frasi di
> requisito nelle sezioni successive descrivono la baseline proposta il 2026-08-23; non prevalgono
> su ADR, contratti, codice o `PROJECT_STATE.md` correnti.
## Stato rispetto al codice corrente
### Fonti autorevoli correnti
Per capire cosa esiste oggi, usare nell'ordine:
- `PROJECT_STATE.md`, per lo snapshot operativo aggiornato;
- `CONTEXT.md`, per il modello di dominio corrente;
- il [piano accettato del Metadata Catalog](../plans/2026-08-26-metadata-catalog-from-thothai.md)
e il [contratto dello schema snapshot](../contracts/catalog-schema-snapshot.md);
- gli ADR del catalogo, in particolare
[0001](../adr/0001-postgres-metadata-catalog.md),
[0004](../adr/0004-fastify-kysely-metadata-catalog.md),
[0006](../adr/0006-separate-physical-and-logical-relationships.md),
[0007](../adr/0007-durable-authoritative-schema-synchronization.md),
[0008](../adr/0008-allow-manual-catalog-metadata-cleanup.md),
[0009](../adr/0009-use-one-sequential-description-generation-run.md),
[0010](../adr/0010-allow-bounded-real-source-samples-for-description-generation.md) e
[0011](../adr/0011-gate-source-samples-with-a-sensitive-data-flag.md);
- l'implementazione, soprattutto `backend/src/catalog/types.ts`,
`backend/src/routes/catalog-databases.ts`, `backend/src/routes/catalog-schema.ts`,
`backend/src/routes/catalog-description-generation.ts`,
`backend/src/routes/catalog-description-consolidation.ts` e
`frontend/src/shell/DatabaseManagementPage.tsx`.
### Disposizione della baseline storica
| Capacità inventariata | Disposizione al 2026-08-31 | Evidenza corrente o nota |
|---|---|---|
| Configurazione per workspace, secret separati e test connessione | **Adottata e implementata** | PostgreSQL interno, binding `postgres_direct`, `rest_api` o `ssh_tunnel`, secret write-only; ADR [0001](../adr/0001-postgres-metadata-catalog.md)–[0004](../adr/0004-fastify-kysely-metadata-catalog.md) e `backend/src/routes/catalog-databases.ts`. |
| Inventario fisico di tabelle, colonne, PK e FK | **Adottato e implementato** | La struttura osservata è distinta dai contenuti curati; `backend/src/catalog/types.ts`, `backend/src/routes/catalog-schema.ts` e ADR [0006](../adr/0006-separate-physical-and-logical-relationships.md). |
| Riconciliazione completa e osservabile dello schema | **Implementata; semantica storica parzialmente superata** | I durable Catalog Sync Runs sono autoritativi, fail-closed, atomici e richiedono conferma per diff distruttivi. Non esiste il lifecycle `drift/removed` proposto qui: la sincronizzazione riconcilia la membership e ADR [0008](../adr/0008-allow-manual-catalog-metadata-cleanup.md) consente anche cleanup manuale esplicito, superando ADR-0005. |
| Descrizione curata e Generated Description separate | **Adottata e implementata** | Tabelle e colonne espongono entrambi i campi in `backend/src/catalog/types.ts`; la copia selettiva AI → curato è in `backend/src/routes/catalog-description-consolidation.ts`. |
| Generazione AI selettiva con run, stato e log | **Adottata e implementata** | Un run asincrono installazione-wide, sequenziale, con eventi persistiti e senza resume automatico; ADR [0009](../adr/0009-use-one-sequential-description-generation-run.md) e `backend/src/routes/catalog-description-generation.ts`. I thread daemon legacy sono **esclusi**. |
| Campioni reali per la generazione e protezione dei campi sensibili | **Implementata come estensione correttiva** | ADR [0010](../adr/0010-allow-bounded-real-source-samples-for-description-generation.md) e [0011](../adr/0011-gate-source-samples-with-a-sensitive-data-flag.md): campioni bounded per colonne non sensibili e valori sintetici deterministici per quelle sensibili. |
| Proposte persistenti, diff e versioni dei testi AI | **Escluse** | Resta un solo Generated Description modificabile. La cronologia dei run è operativa: non conserva prompt, output grezzo, proposta per colonna o audit della decisione umana. |
| Relazioni fisiche | **Adottate e implementate** | Sono constraint immutabili con coppie ordinate di colonne; ADR [0006](../adr/0006-separate-physical-and-logical-relationships.md). |
| Relazioni logiche curate o inferite | **Differite/aperto** | ADR-0006 riserva un modello e un lifecycle separati; non sono ancora parte del catalogo corrente. |
| Alias semantici, descrizioni dei valori, sinonimi e concetti | **Differiti/aperti** | `PROJECT_STATE.md` li assegna a slice dedicate; non vanno dedotti dai campi fisici già implementati. |
| Scope AI, ERD Mermaid e documentazione aggregata | **Differiti/aperti** | Restano capacità legacy inventariate, non una feature corrente del Metadata Catalog. Un eventuale lavoro dovrà avere contratto e gate propri. |
| Export CSV e script SQL dei commenti | **Differiti/aperti; varianti insicure escluse** | Non risultano nella API corrente. Export di segreti, CSV incoerenti e SQL ricostruito da tipi incompleti restano vietati dagli anti-pattern sotto. |
| Inferenza euristica e validazione di relazioni candidate | **Differita/aperta** | Non va confusa con la sincronizzazione delle FK fisiche; dipende dal futuro lifecycle delle relazioni logiche. |
| Pubblicazione del catalogo al core/schema-linking/Qdrant | **Differita e richiesta come design gate successivo** | Il runtime NL→SQL continua a usare configurazione e annotations del workspace; il cutover è esplicitamente rinviato in `PROJECT_STATE.md`. |
| Sette motori database del legacy | **Baseline superata** | La prima versione corrente supporta PostgreSQL; l'aggiunta di altri dialetti è una decisione futura, non parità automatica. |
| GDPR e import da installazioni ThothAI | **Fuori dalla baseline iniziale; import differito** | GDPR resta escluso. Un eventuale import richiede una iniziativa idempotente e un cutover separati, non il riuso degli ID Django. |
| Django Admin, modifica manuale della struttura sorgente, password nel catalogo, duplicazione opaca dei database | **Esclusi** | La UI e le API correnti amministrano il catalogo e non eseguono DDL sul DWH esterno; binding e segreti hanno ownership separata. |
## Sintesi storica
La baseline di parità proposta il 2026-08-23 comprende un catalogo dei database, l'inventario
completo di tabelle, colonne e relazioni fisiche, metadati descrittivi modificabili, relazioni completo di tabelle, colonne e relazioni fisiche, metadati descrittivi modificabili, relazioni
logiche, introspezione dello schema, generazione AI di descrizioni, scope, ERD Mermaid, logiche, introspezione dello schema, generazione AI di descrizioni, scope, ERD Mermaid,
documentazione ed esportazioni operative. La UI Django Admin è soltanto l'interfaccia legacy: documentazione ed esportazioni operative. La UI Django Admin è soltanto l'interfaccia legacy:
non è un requisito architetturale da riprodurre. non è un requisito architetturale da riprodurre.
Il flusso AI per le descrizioni è volutamente semplice e va mantenuto tale: Il flusso AI per le descrizioni individuato come baseline è volutamente semplice:
1. l'AI scrive nel campo `generated_comment` della tabella o colonna; 1. l'AI scrive nel campo `generated_comment` della tabella o colonna;
2. l'utente seleziona gli elementi desiderati; 2. l'utente seleziona gli elementi desiderati;
@@ -193,7 +252,7 @@ scrive soltanto `column.generated_comment`
Provider e modello AI provengono dalla configurazione globale del backend; la lingua è invece Provider e modello AI provengono dalla configurazione globale del backend; la lingua è invece
per database (`/Thoth/ThothAI/backend/thoth_core/thoth_ai/thoth_workflow/comment_generation_utils.py:156-219`). per database (`/Thoth/ThothAI/backend/thoth_core/thoth_ai/thoth_workflow/comment_generation_utils.py:156-219`).
### Comportamento di parità da implementare ### Comportamento di parità proposto il 2026-08-23
Il contratto funzionale minimo è esattamente questo: Il contratto funzionale minimo è esattamente questo:
@@ -314,7 +373,7 @@ Il modulo contiene anche un helper verso `mermaid.ink`, ma non risultano call si
legacy; il percorso attivo è il servizio locale. Non va quindi considerata una dipendenza legacy; il percorso attivo è il servizio locale. Non va quindi considerata una dipendenza
funzionale da conservare. funzionale da conservare.
## 9. Baseline di parità per il Metadata Catalog di ThothII ## 9. Baseline di parità proposta il 2026-08-23
### Necessario ### Necessario
@@ -1,219 +1,206 @@
# PostgreSQL deployment and failure-isolation constraints for the Metadata Catalog # PostgreSQL deployment and failure-isolation constraints for the Metadata Catalog
Date: 2026-08-23 Original research: 2026-08-23
Issue: [#8 — Assess PostgreSQL deployment and failure-isolation constraints](https://github.com/mptyl/ThothII/issues/8)
Last verified against the repository: 2026-08-31
Issue: [#8 — Assess PostgreSQL deployment and failure-isolation constraints](https://git.tylconsulting.it/mptyl/ThothII/issues/8)
> **Status: partially superseded by ADR-0004.** This is historical research, not the current
> architecture contract. The recommendation to run a separate `catalog-api` process was rejected
> by [ADR-0004](../adr/0004-fastify-kysely-metadata-catalog.md). Operational constraints that do
> not depend on that process boundary remain useful, but the status table below is authoritative
> for what the 2026-08-31 code actually adopts, rejects, or leaves pending.
## Question ## Question
How can a Metadata Catalog PostgreSQL service be deployed, backed up, diagnosed, and made How can a Metadata Catalog PostgreSQL service be deployed, backed up, diagnosed, and made
non-blocking for the core across local and server installations? non-blocking for the NL→SQL workflow across local and server installations?
## Recommendation ## Current decision and implementation
Use PostgreSQL as the operational store, but place it behind a **separate Metadata Catalog API [ADR-0001](../adr/0001-postgres-metadata-catalog.md) selects PostgreSQL as the catalog authority.
process**. The SQL workflow `core` must not receive the catalog DSN, catalog credentials, a client [ADR-0003](../adr/0003-installation-local-database-bindings.md) keeps database bindings
pool, or a Compose dependency on either the catalog API or PostgreSQL. installation-local. [ADR-0004](../adr/0004-fastify-kysely-metadata-catalog.md) then places the
catalog in the existing Fastify backend as an isolated Kysely module rather than a microservice.
The resulting dependency graph is: The implemented dependency graph is therefore:
```mermaid ```mermaid
flowchart LR flowchart LR
UI[Frontend] UI[Frontend]
Core[Core workflow API] Core[Core: Fastify workflow and catalog modules]
Pi[Pi and tht workflow]
PG[(catalog-db PostgreSQL)]
Semantic[Qdrant and embedding] Semantic[Qdrant and embedding]
Catalog[Metadata Catalog API]
PG[(Catalog PostgreSQL)]
Publish[Publication worker]
UI -->|workflow routes| Core UI --> Core
Core --> Pi
Core --> PG
Core --> Semantic Core --> Semantic
UI -->|catalog routes| Catalog
Catalog --> PG
PG --> Publish
Publish -->|immutable approved Publication| Semantic
``` ```
The core continues to use the last successfully published, revision-qualified semantic snapshot. The catalog module is isolated behind a repository interface, but it is not process-isolated.
Unpublished catalog edits are never a runtime dependency. This matches the domain boundary: the `core` owns the catalog pool and Compose waits for `catalog-db` health before starting `core`
Metadata Catalog does not select workflow SQL elements, while a Publication is an immutable value (`compose.yaml`, `backend/src/app.ts`, `backend/src/catalog/repository.ts`). Database-management records still do
made available to consumers ([`CONTEXT.md`, lines 96–110](../../CONTEXT.md#L96-L110)). It also not feed the NL→SQL handoff, so a running workflow does not read catalog rows; that cutover remains
preserves the existing rule that Qdrant is a derived index rather than a canonical source future work (`PROJECT_STATE.md`).
([`PROJECT_STATE.md`, lines 405–410](../../PROJECT_STATE.md#L405-L410)).
## Evidence from the current architecture ## Verification result
1. The mandatory workflow stack is currently `frontend`, `core`, `qdrant`, `embedding`, and the | Area | Status on 2026-08-31 | Evidence and consequence |
model-init job. `core` waits only for Qdrant and model initialization
([`compose.yaml`, lines 50–60](../../compose.yaml#L50-L60)); the frontend waits only for `core`
([`compose.yaml`, lines 123–141](../../compose.yaml#L123-L141)). Adding a catalog dependency to
either chain would enlarge the workflow failure domain.
2. `tht start` runs Compose and then checks an explicit allowlist containing only those five
mandatory services. Other Compose services are ignored by the readiness fold
([`service.go`, lines 40–57](../../tools/tht/internal/service/service.go#L40-L57),
[`service.go`, lines 166–190](../../tools/tht/internal/service/service.go#L166-L190)). Separate
catalog services can therefore start with the installation without redefining core readiness.
3. `/health` is deliberately a process-liveness endpoint and does not probe external services
([`app.ts`, lines 271–284](../../backend/src/app.ts#L271-L284)). Tests preserve a 200 response
even when PostgreSQL session storage is configured but not contacted
([`health.test.ts`, lines 9–32](../../backend/test/health.test.ts#L9-L32)). Catalog readiness must
not be folded into this endpoint.
4. ThothII already has a sound server-PostgreSQL precedent: runtime and migrator credentials are
distinct; the one-shot migrator is profile-gated and is explicitly not a dependency of `core`
([`compose.session-server.yaml.example`, lines 24–49](../../deploy/compose.session-server.yaml.example#L24-L49)).
Its documented failure behavior is route-scoped 503 while process liveness remains healthy
([`PROJECT_STATE.md`, lines 621–636](../../PROJECT_STATE.md#L621-L636)).
5. Existing backup archives contain selected Compose volumes only. Local archives include seven
named volumes and server archives include only Qdrant and embedding volumes; no PostgreSQL
logical dump exists today ([`create.go`, lines 29–52](../../tools/tht/internal/backup/create.go#L29-L52)).
A catalog database therefore requires an explicit backup contract rather than an assumption
that current installation backup already covers it.
## Deployment contract
### Shared rules
- Run `catalog-api` as a separate, non-privileged service. The frontend reverse proxy may route
`/api/catalog/*` to it, but frontend startup must not depend on catalog readiness.
- Give `catalog-api` a bounded pool and bounded operations. As an initial ceiling for this small
administrative workload: pool size 5, overflow 0–2, `connect_timeout=3`, `lock_timeout=2s`, and
a normal request `statement_timeout=10s`. Long AI calls must not hold database transactions.
PostgreSQL documents that an omitted or zero `connect_timeout` waits indefinitely, so an
explicit value is required for failure isolation
([libpq connection parameters](https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-CONNECT-TIMEOUT)).
- Use three roles: `catalog_runtime` (CRUD and job claims only), `catalog_migrator` (DDL, used only
by a one-shot job), and `catalog_backup` (minimum privileges needed by the approved dump policy).
No role is shared with the DWH or `thoth_sessions`, and none is a PostgreSQL superuser.
- Store passwords and the CA as separate secret files. For non-local connections default to
`sslmode=verify-full`; this follows the current server-session configuration
([`compose.session-server.yaml.example`, lines 5–20](../../deploy/compose.session-server.yaml.example#L5-L20))
and PostgreSQL's hostname-verifying TLS guidance
([libpq SSL parameters](https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-PARAMKEYWORDS)).
- Apply ordered, checksummed migrations under one transaction and an advisory lock. Refuse pending,
unknown, or checksum-drifted migrations in catalog readiness. The session repository already
implements this pattern
([`postgres_repository.py`, lines 118–187](../../harness/tht/session/postgres_repository.py#L118-L187));
the catalog should have its own migration table, lock key, schema, and roles.
### Local installation
Add an installation-owned `catalog-postgres` container on the private `thothii` network, with no
published host port and a dedicated `catalog-postgres-data` named volume. Pin the image by version
and digest, add a `pg_isready` healthcheck, and let only `catalog-api` use
`depends_on: condition: service_healthy`. `core` and `frontend` keep their existing dependency
graph. Docker confirms that `service_healthy` delays only the declaring dependent service, while
plain container start does not imply application readiness
([Compose startup order](https://docs.docker.com/compose/how-tos/startup-order/)).
The local catalog services can be included in the normal installation Compose set because the
host CLI's mandatory-service allowlist excludes them. A separate `catalog-migrate` one-shot profile
retains explicit operator control; Docker profiles are intended for selectively activated and
one-off services ([Compose profiles](https://docs.docker.com/compose/how-tos/profiles/)).
### Server installation
Use an operator-provided PostgreSQL endpoint configured through a reviewed, untracked overlay.
Prefer a separate database and dedicated roles. A separate PostgreSQL cluster gives the strongest
resource and outage isolation; sharing the existing server cluster is acceptable only when the
operator accepts the residual cluster-wide blast radius and enforces connection limits, timeouts,
and separate ownership. Merely using another schema does not isolate connection exhaustion,
maintenance outages, WAL pressure, or storage failure.
The server overlay adds secrets and catalog configuration only to `catalog-api` and the one-shot
migrator. It must not mount those secrets into `core`. This mirrors the current rule that `core`
never receives the session migrator credential
([`PROJECT_STATE.md`, lines 623–630](../../PROJECT_STATE.md#L623-L630)).
## Failure behavior
| Failure | Catalog behavior | Workflow behavior |
| --- | --- | --- | | --- | --- | --- |
| PostgreSQL unavailable or connection pool exhausted | Catalog routes return sanitized `503 catalog_unavailable` with `Retry-After`; UI remains read-only/unavailable | Existing sessions and new SQL workflow requests continue against the last published snapshot | | PostgreSQL as canonical catalog store | **Adopted** | ADR-0001 is implemented by the PostgreSQL-backed Kysely repository and the internal `catalog-db` service. |
| Catalog API process unavailable | `/api/catalog/*` fails; main frontend shell and workflow route remain usable | No effect | | One Workspace Database per workspace and installation-local bindings | **Adopted** | ADR-0003 and the catalog migrations enforce the model; workspace identity remains in the workspace registry. |
| Pending or drifted migration | Catalog readiness fails and writes are refused | No effect | | Separate `catalog-api` process | **Superseded/rejected** | ADR-0004 explicitly chooses an isolated module inside the existing Fastify process. There is no catalog microservice or separate catalog liveness endpoint. |
| AI generation fails | Job records a retryable/terminal failure; no database transaction remains open | No effect | | No catalog pool, credential, or Compose dependency in `core` | **Superseded/rejected** | `core` owns the runtime pool, receives the runtime password secret, and declares `depends_on: catalog-db: service_healthy`. The old process-level isolation acceptance criterion is not current architecture. |
| Publication indexing fails | Publication remains non-active/failed and can be retried idempotently | Previous active Publication remains in Qdrant | | Private PostgreSQL service and durable volume | **Adopted** | `catalog-db` uses a version-and-digest-pinned PostgreSQL 17.6 image, the private `thothii` network, no published host port, a `pg_isready`-based healthcheck, and `catalog-data`. Compose contract tests assert this topology (`scripts/test-default-compose.sh`). |
| Qdrant unavailable | Publication is queued/failed without losing canonical catalog state | Existing core readiness rules already govern this independent failure | | Different local and server database topology | **Superseded/rejected** | Both current profiles inherit the same internal `catalog-db` and named volume. `deploy/compose.server.yaml` does not replace it with an operator-provided endpoint. |
| Runtime and migrator role separation | **Adopted** | `core` receives only `thothii_catalog_runtime`; the profile-gated `catalog-migrate` job receives only `thothii_catalog_migrate`. Bootstrap grants runtime DML/sequence privileges without DDL (`docker/catalog-db-init.sql`). |
| Dedicated backup role | **Pending** | There is no `catalog_backup` role or backup secret. |
| Explicit one-shot migrations | **Adopted** | `backend/src/catalog/migrate.ts` registers ordered Kysely migrations and uses a pool of one. `scripts/run-stack.sh` starts PostgreSQL and runs `catalog-migrate` before local startup; migrations are not hidden in backend startup. |
| Checksums, drift/pending refusal, and migration readiness | **Pending** | No repository-owned checksum policy, drift report, or readiness gate exists for catalog migrations. The backend can start without checking the Kysely migration head when launched outside the local helper. |
| Bounded runtime connection pool | **Adopted** | `backend/src/catalog/repository.ts` sets `max: 5` and `connectionTimeoutMillis: 3000`, and closes Kysely with the Fastify lifecycle. |
| `lock_timeout`, `statement_timeout`, and catalog TLS policy | **Pending** | The runtime pool does not set query or lock timeouts. Catalog connection configuration has no explicit CA/hostname-verification contract; the server profile still uses the private Compose network. |
| Process liveness independent of catalog queries | **Adopted** | `GET /health` returns `{status: "ok"}` without probing PostgreSQL, and `backend/test/health.test.ts` preserves that behavior. `/catalog/status` performs the catalog-specific availability check. |
| Stack startup independent of catalog availability | **Superseded/rejected** | Compose blocks `core` on healthy `catalog-db`. The host service health fold does not list `catalog-db`, but it cannot make `core` start while its Compose dependency is unhealthy. |
| Uniform catalog-outage response (`503` plus `Retry-After`) | **Pending** | Routes map the domain `CatalogUnavailableError` to a sanitized `503`, and an omitted catalog configuration uses an unavailable repository. PostgreSQL driver failures are not uniformly translated to that domain error, and no `Retry-After` contract is implemented. |
| Catalog-specific readiness and `tht doctor` checks | **Pending** | There is no `/health/ready` for the catalog and no doctor section for connection, migration head, pool saturation, backup age, or publication lag (`tools/tht/internal/doctor/report.go`). |
| Logical catalog backup and restore | **Pending** | Current backup archives omit `catalog-data` and do not run `pg_dump`; server archives include only Qdrant and embedding volumes. A backup can therefore succeed without preserving the Metadata Catalog (`tools/tht/internal/backup/create.go`). |
| Last-good Publication and catalog-to-Qdrant cutover | **Pending** | The current database-management slice does not change the NL→SQL runtime or publish catalog metadata to Qdrant. |
Do not silently fall back to JSON files, dual-write, or read unpublished PostgreSQL rows from the ## Current operational contract
core. Those paths would create split-brain state. The only degraded-mode contract is use of the
last successfully activated Publication.
## Backup and restore ### Deployment and credentials
- Local and server Compose profiles currently use the same installation-owned `catalog-db`
container and `catalog-data` named volume. PostgreSQL is private to the Compose network.
- The database bootstrap login is the migrator. The init script creates the separate runtime login
from a Docker secret and grants only runtime DML and sequence access.
- Runtime and migrator passwords are separate protected host files exposed as separate Docker
secrets. Neither value belongs in tracked environment files
(`deploy/env/local.env.example`, `deploy/env/server.env.example`).
- The Fastify catalog repository uses a five-connection pool with a three-second connection
timeout. Fastify closes the repository pool during shutdown.
### Migrations
The compiled `catalog-migrate` entry point owns schema changes and uses the migrator credential.
The current ordered series is under
`backend/src/catalog/migrations/`. The local launcher runs
it before the normal stack, while production rollout must invoke the profile-gated service
explicitly.
This is weaker than the original recommendation in two ways: there is no catalog readiness check
for pending or unknown migrations, and the repository does not maintain content checksums for
migration drift. Until those checks exist, “the migrator completed” is the available deployment
gate; the application itself does not prove migration compatibility.
### Failure behavior actually provided
| Failure | Current catalog behavior | Current workflow behavior |
| --- | --- | --- |
| Catalog configuration omitted when launching the backend directly | Fastify uses `UnavailableCatalogRepository`; `/catalog/status` reports unavailable and domain-mapped catalog operations return sanitized `503` | `/health`, sessions, and SSE remain available |
| `catalog-db` unhealthy before Compose startup | `core` is not started because its dependency is not healthy | Workflow startup is blocked |
| PostgreSQL becomes unavailable after startup | `/health` remains process-only and `/catalog/status` reports unavailable; route errors are sanitized, but a uniform `503`/`Retry-After` mapping is not guaranteed | Existing session code does not read catalog rows, but both surfaces still share one Fastify process |
| Migrations are pending or incompatible | No dedicated readiness refusal exists; affected catalog operations fail | No catalog-to-workflow handoff exists yet, but local deployment correctness depends on running `catalog-migrate` first |
| Catalog backup is requested through current `tht backup` | No PostgreSQL dump is added and `catalog-data` is omitted | The archive may succeed while being unable to restore catalog state |
The shared process means resource exhaustion, fatal process errors, and startup hooks remain a
common failure domain even though repository calls are separated. Conversely, placing the module
inside Fastify does not require workflow code to consume catalog rows: preserving that data-flow
boundary is still the useful part of the original isolation recommendation.
## Historical recommendations that remain valid backlog
The following constraints survive ADR-0004 because they can be implemented inside the current
Fastify deployment:
1. **Bound every database operation.** Keep the adopted pool and connection timeout; add explicit
request query and lock timeouts. Long model calls and source sampling must not hold catalog
transactions.
2. **Make failures component-specific.** Translate connection, timeout, and pool-exhaustion errors
into one sanitized catalog-unavailable response, add bounded retry guidance, and keep `/health`
process-only.
3. **Make migration compatibility observable.** Report applied, pending, unknown, and drifted
migrations through a catalog readiness check and `tht doctor`; do not put migrator credentials
in `core`.
4. **Define a server TLS topology before externalizing PostgreSQL.** If the server profile moves to
an operator-provided endpoint, use a dedicated database and roles, protected CA material, and
hostname verification. Sharing a cluster leaves connection, maintenance, WAL, and storage blast
radius even when schemas are separate. PostgreSQL documents the relevant connection and TLS
parameters in its
[connection parameter reference](https://www.postgresql.org/docs/current/libpq-connect.html).
5. **Keep derived semantic data non-canonical.** When catalog publication is implemented, activate
complete immutable revisions and retain the last successfully activated revision rather than
exposing mutable catalog rows to the workflow.
## Backup and restore target
The original logical-backup recommendation is still valid and is now a confirmed implementation
gap.
### Backup ### Backup
Extend `tht backup create` or add a catalog-specific command invoked by it. Create a logical Extend `tht backup create` with a catalog step that runs `pg_dump --format=custom` through a
`pg_dump --format=custom` archive through a one-shot helper, checksum it, and place it in the one-shot helper and adds the dump plus checksum to the installation manifest. Do not treat a tar of
installation archive manifest. Do **not** tar a live PostgreSQL data volume: the existing raw-volume a live data volume as a PostgreSQL consistency contract. PostgreSQL documents that `pg_dump`
backup mechanism was designed for the current named volumes, not database crash consistency. creates a consistent export while the database remains in use and that custom format supports
selective restore ([`pg_dump`](https://www.postgresql.org/docs/current/app-pgdump.html)).
`pg_dump` produces a consistent export while the database remains in use, and custom format is Use a dedicated least-privilege backup role, never write dumps into a workspace repository, and
compressed and supports selective/reordered restore record at least the server version, migration head, checksum, creation time, and catalog revision
([PostgreSQL `pg_dump`](https://www.postgresql.org/docs/current/app-pgdump.html)). Because metadata identifiers. A filesystem/Qdrant archive with a failed or absent catalog dump must be reported as
volume is small, favor correctness over parallelism. Run a backup before every migration and on an incomplete.
operator-defined schedule; record the server version, migration head, dump checksum, creation time,
and catalog Publication identifiers in the manifest. Never write a dump into the workspace Git
repository.
For an external server database, the operator must supply the dedicated backup credential through
the same secret-file boundary. A successful filesystem/Qdrant archive with a failed catalog dump
is an incomplete installation backup and must not be reported as success.
### Restore ### Restore
Restore only during an explicit catalog maintenance window into an empty target database (or a Restore only through an explicit maintenance operation into an empty or deliberately cleaned
deliberately cleaned target), using `pg_restore --exit-on-error --single-transaction`. Validate the target. Validate the archive checksum, run `pg_restore --exit-on-error --single-transaction`, then
archive checksum first, then require: verify migration compatibility, referential integrity, bounded entity counts, and an authenticated
catalog smoke test. The controls are documented by PostgreSQL
([`pg_restore`](https://www.postgresql.org/docs/current/app-pgrestore.html)). Keep the old database
until validation passes; rollback should switch the endpoint or retained volume, not dual-write.
1. expected migration history with no checksum drift; ## Diagnostics target
2. referential-integrity checks and bounded entity counts;
3. Publication content hashes matching the restored records;
4. an authenticated catalog smoke test.
The relevant restore controls and their behavior are documented by - Keep the shared `GET /health` endpoint process-only.
[PostgreSQL `pg_restore`](https://www.postgresql.org/docs/current/app-pgrestore.html). After restore, - Treat `/catalog/status` as the current minimal availability surface; add a bounded catalog
rebuild catalog-derived Qdrant content from the restored active Publication rather than treating readiness check covering connection and migration compatibility before using it as a rollout
the vector index as the source of truth. Keep the old database untouched until this validation gate.
passes; rollback is endpoint reversal, not dual-write. - Add read-only `tht doctor` checks for secret-file presence and permissions, PostgreSQL
reachability, authentication, migration state, pool saturation, and last successful catalog
backup. Sanitize all DSNs and errors.
- Use `pg_isready` only for server transport readiness. Its result does not prove schema or runtime
authorization correctness
([`pg_isready`](https://www.postgresql.org/docs/current/app-pg-isready.html)).
- Add structured metrics for connection acquisition failures, pool use, query/lock timeout,
migration head, background-job backlog, and backup age. Do not log connection strings, source
samples, prompts, or generated descriptions at info level.
## Diagnostics and observability ## Recommended closure criteria
- Keep `core /health` unchanged. Give `catalog-api` separate `/health/live` (process only) and 1. Decide whether ADR-0004's statement that catalog unavailability does not make sessions or SSE
`/health/ready` (bounded connection, `SELECT 1`, and migration compatibility) endpoints. unavailable must also hold at Compose startup. If yes, remove or soften the hard `core` →
- Add a named `metadata-catalog` section to `tht doctor`: configuration completeness; secret-file `catalog-db` startup dependency without reintroducing a microservice.
presence and permissions; DNS/TCP/TLS/authentication; `SELECT 1`; migration applied/pending/drifted 2. Prove a configured PostgreSQL outage produces the same sanitized catalog response across every
state; pool saturation; last successful backup age; and oldest pending Publication. Checks are catalog route while `/health`, session creation, and SSE continue to work.
read-only and all errors/DSNs are sanitized. The current doctor already separates Compose service 3. Add catalog migration compatibility to readiness and `tht doctor`, including pending, unknown,
status from container-local workflow checks and drifted states.
([`report.go`, lines 275–310](../../tools/tht/internal/doctor/report.go#L275-L310)). 4. Add a real `pg_dump`/`pg_restore` round-trip to local and server backup tests, and fail backup
- Use `pg_isready` only for transport readiness; its exit codes distinguish accepting, rejecting, publication when the catalog dump is absent or fails.
no-response, and invalid-parameter states, but it does not prove schema or authorization 5. Before enabling a remote server catalog, define TLS verification, credential files, connection
correctness ([PostgreSQL `pg_isready`](https://www.postgresql.org/docs/current/app-pg-isready.html)). limits, and the accepted cluster-level blast radius.
- Expose structured metrics/logs for connection acquisition failures, pool utilization, query 6. Before the NL→SQL cutover, define and test an immutable last-good Publication boundary; do not
timeout/lock timeout, migration head, AI-job backlog, Publication lag, and backup age. Never log dual-read mutable PostgreSQL rows and legacy metadata as competing authorities.
connection strings, prompts containing schema metadata, or generated descriptions at info level.
- `tht status` may show the catalog as `healthy`, `degraded`, or `disabled`, but catalog state must
not change the existing core health result. `tht doctor` may return a failing catalog check for
operator visibility without stopping or restarting the workflow stack.
## Acceptance constraints for later design and implementation
1. Killing `catalog-postgres` and `catalog-api` must not change `core /health`, stop Pi, interrupt an
existing SQL session, or prevent a new session that uses the already active Publication.
2. No catalog database secret, client dependency, migration, or network call exists in the harness
workflow path.
3. Catalog failure is explicit on catalog routes; no JSON fallback or dual-write is introduced.
4. A failed Publication leaves the previous active semantic content addressable and retryable.
5. Local and server backup tests perform a real dump/restore round trip and prove that a backup is
incomplete when the PostgreSQL dump fails.
6. Migration credentials are absent from runtime containers; drift and pending migrations are
visible in readiness and `tht doctor`.
## Decision summary ## Decision summary
The Metadata Catalog should be a separately deployed bounded context backed by PostgreSQL. Local PostgreSQL, installation-local bindings, private deployment, role separation, one-shot migrations,
installations own a private container and volume; server installations use a TLS-verified external bounded pooling, and process-only liveness are implemented. The separate `catalog-api` process and
database, preferably on a separate cluster when strict blast-radius isolation is required. Logical the claim that `core` has no catalog dependency are not current design: ADR-0004 chose the existing
dump/restore, one-shot migrations, role separation, bounded connections, component-specific Fastify process, and Compose currently blocks `core` startup on `catalog-db` health. Logical
diagnostics, and last-good-Publication semantics make PostgreSQL operationally safe without making backup/restore, catalog migration readiness and drift detection, consistent outage mapping,
it a dependency of the SQL workflow. catalog-specific doctor checks, server TLS/external topology, and last-good publication remain
pending. Those open constraints should be treated as backlog, not as capabilities already provided
by the repository.
@@ -4,21 +4,51 @@
session-pinning, and runtime read-only contracts constrain metadata publication without changing session-pinning, and runtime read-only contracts constrain metadata publication without changing
the NL→SQL workflow? the NL→SQL workflow?
**Last verified:** 2026-08-31
**Validity:** Active architectural research; the catalog-to-core integration described below is
still **deferred**, not implemented.
**Current authorities:** `PROJECT_STATE.md`,
[`workspace-evidence-v3.md`](../contracts/workspace-evidence-v3.md),
[`workspace-preprocessing-cli.md`](../contracts/workspace-preprocessing-cli.md),
[`tht-dwh.md`](../contracts/tht-dwh.md),
[`ADR-0001`](../adr/0001-postgres-metadata-catalog.md), and
[`ADR-0004`](../adr/0004-fastify-kysely-metadata-catalog.md). These sources and current code
override this research note if they diverge.
## Revalidation against the current implementation
| Finding | Status on 2026-08-31 | Current evidence and consequence |
| --- | --- | --- |
| PostgreSQL metadata authority in the existing Fastify backend | **Adopted/current** | ADR-0001 and ADR-0004 are implemented; the catalog is the management-plane authority. |
| Git workspace revision, immutable registry snapshot, and revision-pinned runtime | **Adopted/current** | Registry activation and the Workspace Evidence v3 contract still provide the publication boundary for runtime-owned files. |
| Qdrant read isolation for schema and Evidence by `workspace_id` + `workspace_revision` | **Adopted/current** | `QdrantVectorStore.search()` applies `_revision_filter()` to schema/Evidence; Memory and solved questions intentionally remain workspace-wide (`harness/tht/adapters/vector/qdrant.py`). |
| Physical schema as a file in the Git publication | **Superseded clarification** | `physical.yaml` is owned by the immutable `.tht-dwh` generation selected by `ACTIVE`, not by the Git workspace revision ([DWH contract](../contracts/tht-dwh.md)). Git currently supplies the revision-pinned curated `schema/annotations.yaml`. |
| Catalog-to-core publisher | **Open/deferred** | No current route or service projects catalog records into core artifacts. `PROJECT_STATE.md` explicitly defers the schema-linking integration (`PROJECT_STATE.md:160-171`). |
| Explicit Core Schema Selection | **Open/deferred** | The workspace descriptor selects one database and one physical schema, but has no table/column allowlist (`backend/src/workspaces/schema.ts`); no catalog selection is handed to `tht`. |
| Cross-revision hash lookup used by schema synchronization | **Open defect** | Reads are revision-filtered, but `existing_hashes()` is not. A same-key/same-content point from an older revision can suppress the required upsert into the new revision (`harness/tht/adapters/vector/qdrant.py`, `harness/tht/cli/vector_cmd.py`). |
| Deletion/GC | **Current for Evidence; open for schema** | Evidence has generation inventory, retention, compensation, and exact-generation deletion. `sync_canonical_records()` never deletes schema records absent from the new canonical set, and no revision-retention GC exists for schema points. |
| Annotation consumption | **Current but incomplete** | M-Schema rendering and schema embeddings consume `Annotations`; `tht schema columns` still returns only physical comments, so F4 does not display annotation descriptions (`harness/tht/cli/schema_cmd.py`, `harness/.pi/extensions/tht-gate.js`). |
## Conclusion ## Conclusion
The lowest-impact publication seam is **not** a new writer inside the NL→SQL workflow and it is The lowest-impact **proposed** publication seam is not a new writer inside the NL→SQL workflow and
not a direct CRUD-to-Qdrant path. The existing core already consumes two canonical schema inputs: is not a direct CRUD-to-Qdrant path. The existing core consumes an introspected `PhysicalSchema`
an introspected `PhysicalSchema` and a curated `Annotations` document. An explicit publication from the active immutable DWH generation and a curated, Git-revision-pinned `Annotations`
operation should project the approved, workspace-selected subset of the new metadata catalog into document. A future explicit publication operation should project the approved catalog subset into
those inputs, bind the projection to a new Git `workspace_revision`, activate that immutable a new Core Schema Selection contract and the curated annotations, activate a new immutable Git
revision, and then invoke the existing `workspace index-schema` preprocessing operation. Qdrant revision, and then invoke the existing `workspace index-schema` preprocessing operation. Qdrant
remains a derived, rebuildable projection. remains a derived, rebuildable projection. None of this catalog-to-core handoff exists yet;
`PROJECT_STATE.md:160-164` deliberately defers it to the next design
gate.
This preserves the existing runtime path: The proposed flow would preserve the existing runtime path:
```text ```text
approved catalog data approved catalog data
-> workspace Git revision (physical schema selection + annotations) -> explicit Core Schema Selection + curated annotations projection (future)
-> workspace Git revision (annotations and selection contract; physical.yaml stays DWH-owned)
-> immutable registry snapshot -> immutable registry snapshot
-> revision-bound runtime configuration -> revision-bound runtime configuration
-> existing tht vector index-schema -> existing tht vector index-schema
@@ -26,23 +56,23 @@ approved catalog data
-> existing retrieval_pack / schema render / F4 review -> existing retrieval_pack / schema render / F4 review
``` ```
The complete database inventory may remain in the Metadata Catalog. Only the explicit Core Schema The complete database inventory remains in the Metadata Catalog. Only an explicit Core Schema
Selection should be projected into the workspace artifacts used by the SQL workflow. ThothII does Selection and approved semantic fields should be projected into the artifacts used by the SQL
not currently model such a table-level selection in its descriptor, so that projection contract is workflow. ThothII does not currently model or publish that table/column-level selection, so both
new work. the projection contract and its operational publisher are new work.
## 1. Current authority and publication boundary ## 1. Current authority and publication boundary
The workspace descriptor identifies one PostgreSQL database and one physical schema, plus one The workspace descriptor identifies one PostgreSQL database and one physical schema, plus one
workspace-owned Qdrant collection; it has no table or column allowlist workspace-owned Qdrant collection; it has no table or column allowlist
([`backend/src/workspaces/schema.ts:160-178`](../../backend/src/workspaces/schema.ts#L160-L178)). (`backend/src/workspaces/schema.ts`).
Consequently, a full-database metadata catalog and the subset eligible for the SQL core cannot be Consequently, a full-database metadata catalog and the subset eligible for the SQL core cannot be
represented as the same current descriptor object. represented as the same current descriptor object.
The existing public contract makes the Git workspace repository curator-owned. Changes occur in a The existing public contract makes the Git workspace repository curator-owned. Changes occur in a
separate authoring clone followed by installation pull, and the API does not write workspace, separate authoring clone followed by installation pull, and the API does not write workspace,
schema, or Evidence paths schema, or Evidence paths
([`docs/contracts/workspace-evidence-v3.md:148-156`](../contracts/workspace-evidence-v3.md#L148-L156)). ([Workspace Evidence v3 contract](../contracts/workspace-evidence-v3.md)).
Therefore a browser CRUD service cannot silently make its PostgreSQL state authoritative for the Therefore a browser CRUD service cannot silently make its PostgreSQL state authoritative for the
core without either: core without either:
@@ -54,10 +84,12 @@ The first option preserves current architecture and session reproducibility.
Registry activation validates every descriptor at one exact commit, validates and synchronizes Registry activation validates every descriptor at one exact commit, validates and synchronizes
the commit's `schema/annotations.yaml`, and rejects duplicate ownership of a Qdrant collection the commit's `schema/annotations.yaml`, and rejects duplicate ownership of a Qdrant collection
([`backend/src/workspaces/registry.ts:531-576`](../../backend/src/workspaces/registry.ts#L531-L576)). (`backend/src/workspaces/registry.ts`,
`backend/src/workspaces/git-repository.ts`,
`backend/src/workspaces/annotations-sync.ts`).
It writes the candidate snapshot into a staging directory, records file digests in It writes the candidate snapshot into a staging directory, records file digests in
`snapshot.json`, atomically renames the directory, and only then moves active state `snapshot.json`, atomically renames the directory, and only then moves active state
([`backend/src/workspaces/registry.ts:582-629`](../../backend/src/workspaces/registry.ts#L582-L629)). (`backend/src/workspaces/registry.ts`).
That is the existing atomic publication boundary to reuse. That is the existing atomic publication boundary to reuse.
## 2. Canonical schema inputs already consumed by the core ## 2. Canonical schema inputs already consumed by the core
@@ -67,15 +99,14 @@ The harness separates source facts from curated semantics:
- `PhysicalSchema` contains database/schema identity and tables; table facts include comments, - `PhysicalSchema` contains database/schema identity and tables; table facts include comments,
columns, physical foreign keys, and indexes; columns contain type, nullability, primary-key, columns, physical foreign keys, and indexes; columns contain type, nullability, primary-key,
default, source comment, examples, and eligibility default, source comment, examples, and eligibility
([`harness/tht/mschema/models.py:23-64`](../../harness/tht/mschema/models.py#L23-L64)). (`harness/tht/mschema/models.py`).
- `Annotations` contains curated table descriptions, concepts and notes, column descriptions, - `Annotations` contains curated table descriptions, concepts and notes, column descriptions,
synonyms, concepts, evidence, notes and eligibility overrides, plus logical foreign keys synonyms, concepts, evidence, notes and eligibility overrides, plus logical foreign keys
([`harness/tht/mschema/models.py:67-88`](../../harness/tht/mschema/models.py#L67-L88)). (`harness/tht/mschema/models.py`).
Rendering already gives annotations precedence over source comments and merges physical and Rendering already gives annotations precedence over source comments and merges physical and
logical foreign keys. It also applies column eligibility before producing M-Schema context logical foreign keys. It also applies column eligibility before producing M-Schema context
([`harness/tht/mschema/render.py:9-43`](../../harness/tht/mschema/render.py#L9-L43), (`harness/tht/mschema/render.py`). This makes
[`harness/tht/mschema/render.py:46-80`](../../harness/tht/mschema/render.py#L46-L80)). This makes
`Annotations` the natural narrow projection target for approved descriptions, synonyms, logical `Annotations` the natural narrow projection target for approved descriptions, synonyms, logical
relationships, and eligibility from the new catalog. relationships, and eligibility from the new catalog.
@@ -83,14 +114,14 @@ There are two current compatibility gaps:
- `schema_records()` embeds table descriptions/concepts and column descriptions/synonyms/examples, - `schema_records()` embeds table descriptions/concepts and column descriptions/synonyms/examples,
but does not include physical or logical foreign keys in vector record content but does not include physical or logical foreign keys in vector record content
([`harness/tht/vectorstore/records.py:101-132`](../../harness/tht/vectorstore/records.py#L101-L132)). (`harness/tht/vectorstore/records.py`).
Relationships still reach the model through deterministic schema rendering, not through schema Relationships still reach the model through deterministic schema rendering, not through schema
candidate embeddings. candidate embeddings.
- The F4 widget loads columns with `tht schema columns` - The F4 widget loads columns with `tht schema columns`
([`harness/.pi/extensions/tht-gate.js:1220-1254`](../../harness/.pi/extensions/tht-gate.js#L1220-L1254)), (`harness/.pi/extensions/tht-gate.js`), but that
but that command currently returns only `PhysicalSchema.comment` values and does not merge command still returns only `PhysicalSchema.comment` values and does not merge
`Annotations` `Annotations`
([`harness/tht/cli/schema_cmd.py:592-621`](../../harness/tht/cli/schema_cmd.py#L592-L621)). (`harness/tht/cli/schema_cmd.py`).
A publisher that writes only annotations would improve vector search and rendered M-Schema but A publisher that writes only annotations would improve vector search and rendered M-Schema but
not the table/column descriptions displayed by this existing reviewer widget. Fixing the command not the table/column descriptions displayed by this existing reviewer widget. Fixing the command
to use the existing merged description helpers would preserve the workflow shape while closing to use the existing merged description helpers would preserve the workflow shape while closing
@@ -101,32 +132,33 @@ There are two current compatibility gaps:
`workspace index-schema` is already the supported operator seam. It creates a revision-bound `workspace index-schema` is already the supported operator seam. It creates a revision-bound
runtime, checks collection compatibility, and runs the harness command runtime, checks collection compatibility, and runs the harness command
`vector index-schema --json` `vector index-schema --json`
([`backend/src/workspaces/preprocessing-service.ts:312-330`](../../backend/src/workspaces/preprocessing-service.ts#L312-L330)). (`backend/src/workspaces/preprocessing-service.ts`,
[`workspace-preprocessing-cli.md`](../contracts/workspace-preprocessing-cli.md)).
The full preprocessing operation performs DWH preparation, FK suggestion/review, schema indexing, The full preprocessing operation performs DWH preparation, FK suggestion/review, schema indexing,
and optional Evidence preprocessing as separate resumable stages and optional Evidence preprocessing as separate resumable stages
([`backend/src/workspaces/preprocessing-service.ts:367-450`](../../backend/src/workspaces/preprocessing-service.ts#L367-L450)). (`backend/src/workspaces/preprocessing-service.ts`).
Metadata-only publication should normally use the narrow `index-schema` operation after its Metadata-only publication should normally use the narrow `index-schema` operation after its
workspace artifacts are valid, rather than coupling catalog CRUD to the full pipeline. workspace artifacts are valid, rather than coupling catalog CRUD to the full pipeline.
The harness indexer reads the active immutable physical schema and revision-specific annotations, The harness indexer reads the active immutable physical schema and revision-specific annotations,
constructs schema records, embeds only changed content, and writes them through the configured constructs schema records, embeds only changed content, and writes them through the configured
vector adapter vector adapter
([`harness/tht/cli/vector_cmd.py:100-140`](../../harness/tht/cli/vector_cmd.py#L100-L140)). Its (`harness/tht/cli/vector_cmd.py`). Its
machine result carries the physical/annotation artifact digests, workspace revision, collection, machine result carries the physical/annotation artifact digests, workspace revision, collection,
and counts and counts
([`harness/tht/cli/vector_cmd.py:157-179`](../../harness/tht/cli/vector_cmd.py#L157-L179)). This is (`harness/tht/cli/vector_cmd.py`). This is
the right place to retain publication evidence and audit linkage. the right place to retain publication evidence and audit linkage.
Preprocessing state is already revision- and binding-aware. A resumed job must match operation, Preprocessing state is already revision- and binding-aware. A resumed job must match operation,
workspace revision, descriptor/catalog blobs, runtime config and binding digests or it fails with workspace revision, descriptor/catalog blobs, runtime config and binding digests or it fails with
`preprocessing_resume_mismatch` `preprocessing_resume_mismatch`
([`backend/src/workspaces/preprocessing-state.ts:262-313`](../../backend/src/workspaces/preprocessing-state.ts#L262-L313)). (`backend/src/workspaces/preprocessing-state.ts`).
There is an important operational gate: every preprocessing operation calls There is an important operational gate: every preprocessing operation calls
`assertSessionInventoryCompatible`; a non-finalized, non-archived session pinned to another `assertSessionInventoryCompatible`; a non-finalized, non-archived session pinned to another
revision blocks preprocessing revision blocks preprocessing
([`backend/src/workspaces/preprocessing-service.ts:459-472`](../../backend/src/workspaces/preprocessing-service.ts#L459-L472), (`backend/src/workspaces/preprocessing-service.ts`,
[`backend/src/workspaces/preprocessing-state.ts:363-378`](../../backend/src/workspaces/preprocessing-state.ts#L363-L378)). `backend/src/workspaces/preprocessing-state.ts`).
A catalog publication UX must expose this as a pending/blocking condition rather than report a A catalog publication UX must expose this as a pending/blocking condition rather than report a
generic indexing failure. generic indexing failure.
@@ -135,23 +167,22 @@ generic indexing failure.
The collection contract is fixed at the descriptor's dimensions/distance and eight keyword payload The collection contract is fixed at the descriptor's dimensions/distance and eight keyword payload
indexes: `content_hash`, `document_id`, `kind`, `record_key`, `record_kind`, indexes: `content_hash`, `document_id`, `kind`, `record_key`, `record_kind`,
`vector_generation`, `workspace_id`, and `workspace_revision` `vector_generation`, `workspace_id`, and `workspace_revision`
([`backend/src/workspaces/qdrant-collection.ts:1-20`](../../backend/src/workspaces/qdrant-collection.ts#L1-L20)). (`backend/src/workspaces/qdrant-collection.ts`).
`self_heal` may create a missing compatible collection or indexes; `require_existing` only validates `self_heal` may create a missing compatible collection or indexes; `require_existing` only validates
and refuses an incompatible collection and refuses an incompatible collection
([`backend/src/workspaces/qdrant-collection.ts:71-104`](../../backend/src/workspaces/qdrant-collection.ts#L71-L104)). (`backend/src/workspaces/qdrant-collection.ts`).
The publication path must use this shared manager instead of inventing collection setup. The publication path must use this shared manager instead of inventing collection setup.
Schema payloads include both the workspace and workspace revision, record identity, semantic kind, Schema payloads include both the workspace and workspace revision, record identity, semantic kind,
content and content hash content and content hash
([`harness/tht/vectorstore/records.py:30-49`](../../harness/tht/vectorstore/records.py#L30-L49)). (`harness/tht/vectorstore/records.py`).
Qdrant point IDs for schema and Evidence also include the revision, and upserts use the same Qdrant point IDs for schema and Evidence also include the revision, and upserts use the same
revision in the payload revision in the payload
([`harness/tht/adapters/vector/qdrant.py:42-46`](../../harness/tht/adapters/vector/qdrant.py#L42-L46), (`harness/tht/adapters/vector/qdrant.py`).
[`harness/tht/adapters/vector/qdrant.py:196-220`](../../harness/tht/adapters/vector/qdrant.py#L196-L220)).
Reads always filter by `workspace_id`; schema and Evidence reads additionally filter by the bound Reads always filter by `workspace_id`; schema and Evidence reads additionally filter by the bound
`workspace_revision`, while memory and solved-question records intentionally remain `workspace_revision`, while memory and solved-question records intentionally remain
workspace-wide workspace-wide
([`harness/tht/adapters/vector/qdrant.py:291-308`](../../harness/tht/adapters/vector/qdrant.py#L291-L308)). (`harness/tht/adapters/vector/qdrant.py:125-145`).
This means an approved semantic change needs a new workspace revision if old sessions must retain This means an approved semantic change needs a new workspace revision if old sessions must retain
their previous view. Directly overwriting points under the same Git revision would mutate the their previous view. Directly overwriting points under the same Git revision would mutate the
@@ -160,43 +191,52 @@ would require changing the current runtime filter contract.
### Qdrant synchronization defect to resolve before catalog publication ### Qdrant synchronization defect to resolve before catalog publication
The current incremental synchronizer compares canonical records by content hash and upserts The current incremental schema synchronizer compares canonical records by content hash and upserts
changed records, but it has no deletion step changed records, but it has no deletion step
([`harness/tht/cli/vector_cmd.py:67-89`](../../harness/tht/cli/vector_cmd.py#L67-L89)). More (`harness/tht/cli/vector_cmd.py:77-99`). More
importantly, `QdrantVectorStore.existing_hashes()` filters by workspace and record kind but does not importantly, `QdrantVectorStore.existing_hashes()` filters by workspace and record kind but does not
apply `_revision_filter()` apply `_revision_filter()`
([`harness/tht/adapters/vector/qdrant.py:174-194`](../../harness/tht/adapters/vector/qdrant.py#L174-L194)), (`harness/tht/adapters/vector/qdrant.py:266-286`),
even though point IDs and reads are revision-scoped. Therefore a same-key/same-content record from even though point IDs and reads are revision-scoped. Therefore a same-key/same-content record from
an older revision may be classified as unchanged and never written under the new revision. This is an older revision may be classified as unchanged and never written under the new revision. This is
an implementation defect/risk inferred directly from the two code paths, and it should be fixed an implementation defect/risk inferred directly from the two code paths, and it should be fixed
and regression-tested before the catalog relies on `index-schema` for multi-revision publication. and regression-tested before the catalog relies on `index-schema` for multi-revision publication.
Deletion/GC must be stated per record family. Evidence cleanup is implemented: the corpus pipeline
retains the configured number of published generations, protects active/job-referenced
generations, and calls exact-generation deletion for evicted or compensated generations
(`harness/tht/evidence/corpus/pipeline.py`,
`harness/tht/adapters/vector/qdrant.py:350-390`). Schema cleanup is
not implemented: `delete_kinds()` exists as a workspace-scoped adapter primitive, but the schema
synchronizer never calls it, it is not revision-scoped, and there is no retention policy for old
schema revisions. Publication therefore still needs exact current-revision deletion semantics and
separate safe GC for unleased historical schema revisions.
## 5. Session pinning and why the SQL workflow can remain unchanged ## 5. Session pinning and why the SQL workflow can remain unchanged
Normal session creation acquires an immutable registry revision, passes its snapshot path, Normal session creation acquires an immutable registry revision, passes its snapshot path,
workspace ID and commit to `tht session new`, and only releases the retention lease after the workspace ID and commit to `tht session new`, and only releases the retention lease after the
manifest has been persisted manifest has been persisted
([`backend/src/routes/sessions.ts:341-367`](../../backend/src/routes/sessions.ts#L341-L367), (`backend/src/routes/sessions.ts`). The
[`backend/src/routes/sessions.ts:423-440`](../../backend/src/routes/sessions.ts#L423-L440)). The
manifest stores `workspace_id` and `workspace_revision` next to database/schema identity manifest stores `workspace_id` and `workspace_revision` next to database/schema identity
([`harness/tht/session/store.py:101-132`](../../harness/tht/session/store.py#L101-L132)). Resume and (`harness/tht/session/store.py`). Resume and
saved-SQL paths reopen that exact retained snapshot rather than the current installation default saved-SQL paths reopen that exact retained snapshot rather than the current installation default
([`backend/src/routes/sessions.ts:184-198`](../../backend/src/routes/sessions.ts#L184-L198), (`backend/src/routes/sessions.ts`,
[`backend/src/routes/sql.ts:24-43`](../../backend/src/routes/sql.ts#L24-L43)). `backend/src/routes/sql.ts`).
The runtime renderer places the same workspace revision in `runtime_identity`, points the harness The runtime renderer places the same workspace revision in `runtime_identity`, points the harness
at the descriptor-owned Qdrant collection, and supplies the internal embedding service at the descriptor-owned Qdrant collection, and supplies the internal embedding service
([`backend/src/workspaces/runtime-renderer.ts:248-277`](../../backend/src/workspaces/runtime-renderer.ts#L248-L277)). (`backend/src/workspaces/runtime-renderer.ts`).
The vector adapter is constructed directly from that configuration, including the revision The vector adapter is constructed directly from that configuration, including the revision
([`harness/tht/adapters/factory.py:27-42`](../../harness/tht/adapters/factory.py#L27-L42)). (`harness/tht/adapters/factory.py`).
At session bootstrap the backend invokes the existing `search pack` command At session bootstrap the backend invokes the existing `search pack` command
([`backend/src/routes/sessions.ts:465-470`](../../backend/src/routes/sessions.ts#L465-L470)). That (`backend/src/routes/sessions.ts`). That
command queries schema records with the existing schema kinds, ranks tables and persists the command queries schema records with the existing schema kinds, ranks tables and persists the
candidate list candidate list
([`harness/tht/cli/search_cmd.py:291-310`](../../harness/tht/cli/search_cmd.py#L291-L310)). The Pi (`harness/tht/cli/search_cmd.py`). The Pi
extension reads the persisted retrieval pack through the CLI extension reads the persisted retrieval pack through the CLI
([`harness/.pi/extensions/tht-gate.js:61-75`](../../harness/.pi/extensions/tht-gate.js#L61-L75)), (`harness/.pi/extensions/tht-gate.js`),
and F4 starts from those candidates while loading full table/column context through existing schema and F4 starts from those candidates while loading full table/column context through existing schema
commands. Thus catalog publication can improve the inputs without changing phases, gate semantics, commands. Thus catalog publication can improve the inputs without changing phases, gate semantics,
or persisted session artifacts. or persisted session artifacts.
@@ -204,17 +244,18 @@ or persisted session artifacts.
## 6. DWH reuse across metadata-only revisions ## 6. DWH reuse across metadata-only revisions
DWH preparation uses immutable generations selected by an `ACTIVE` pointer DWH preparation uses immutable generations selected by an `ACTIVE` pointer
([`docs/contracts/tht-dwh.md:18-24`](../contracts/tht-dwh.md#L18-L24)). The effective-configuration ([DWH contract](../contracts/tht-dwh.md)). The effective-configuration
identity deliberately excludes `runtime_identity`, so a content-only Git revision does not force a identity deliberately excludes `runtime_identity`, so a content-only Git revision does not force a
database re-introspection database re-introspection
([`docs/contracts/tht-dwh.md:74-90`](../contracts/tht-dwh.md#L74-L90)). This is the key enabling ([DWH contract](../contracts/tht-dwh.md)). This is the key enabling property for future metadata
property for metadata publication: a new revision can carry updated curated annotations/Core Schema publication: a new Git revision can carry updated curated annotations and a future Core Schema
Selection, reuse the compatible physical DWH generation, and rebuild only the revision-scoped Selection contract, reuse the compatible DWH-owned `physical.yaml`, and rebuild only the
schema projection. revision-scoped schema projection. Today only the curated annotations part of that statement exists.
## 7. Constraints for the Metadata Catalog design ## 7. Constraints for the Metadata Catalog design
The following should be treated as requirements for the architecture map: The following remain proposed requirements for the deferred catalog-to-core design gate; they are
not claims about current implementation:
1. **Separate full inventory from core projection.** PostgreSQL may hold the complete database 1. **Separate full inventory from core projection.** PostgreSQL may hold the complete database
catalog, drafts and AI-generated text. Only an explicit Core Schema Selection and approved catalog, drafts and AI-generated text. Only an explicit Core Schema Selection and approved
@@ -225,10 +266,10 @@ The following should be treated as requirements for the architecture map:
3. **Use a new Git revision as the publication identity.** This preserves current snapshot, 3. **Use a new Git revision as the publication identity.** This preserves current snapshot,
retention, session resume and Qdrant filtering semantics. A separate mutable catalog revision retention, session resume and Qdrant filtering semantics. A separate mutable catalog revision
cannot be safely introduced without changing runtime reads. cannot be safely introduced without changing runtime reads.
4. **Reuse canonical formats.** Project physical facts/Core Schema Selection into a validated 4. **Respect artifact ownership.** Keep introspected physical facts in the immutable DWH generation;
`PhysicalSchema` view and semantic edits into `Annotations`; keep Mermaid and long-form database project the future Core Schema Selection through an explicit contract and approved semantic
documentation outside Qdrant unless a separate, explicit record kind and retrieval policy is edits into `Annotations`. Keep Mermaid and long-form database documentation outside Qdrant
designed. unless a separate record kind and retrieval policy is designed.
5. **Reuse the operator boundary.** Trigger `workspace index-schema`, observe its schema-versioned 5. **Reuse the operator boundary.** Trigger `workspace index-schema`, observe its schema-versioned
result and persist its artifact identities. Do not call Qdrant from browser CRUD handlers. result and persist its artifact identities. Do not call Qdrant from browser CRUD handlers.
6. **Keep runtime read-only.** The session/Pi process continues to read the pinned snapshot, 6. **Keep runtime read-only.** The session/Pi process continues to read the pinned snapshot,
@@ -236,14 +277,15 @@ The following should be treated as requirements for the architecture map:
management control plane. management control plane.
7. **Surface publication gates.** UI status must distinguish Git activation, incompatible/missing 7. **Surface publication gates.** UI status must distinguish Git activation, incompatible/missing
collection, resumable-session revision conflict, embedding failure and completed publication. collection, resumable-session revision conflict, embedding failure and completed publication.
8. **Repair and test cross-revision synchronization first.** Scope `existing_hashes()` to the bound 8. **Repair and test cross-revision schema synchronization first.** Scope `existing_hashes()` to
revision and define deletion/garbage-collection semantics before depending on repeated catalog the bound revision, delete records removed from the current revision's canonical schema, and
publication. define safe schema-revision GC. Reuse rather than duplicate the already implemented Evidence
generation GC.
9. **Close the annotation display gap without changing the workflow.** Make `schema columns` read 9. **Close the annotation display gap without changing the workflow.** Make `schema columns` read
the same merged descriptions used by M-Schema/vector rendering, so the current F4 widget sees the same merged descriptions used by M-Schema/vector rendering, so the current F4 widget sees
the approved catalog text. the approved catalog text.
## Decision summary for the Wayfinder map ## Proposed decision summary for the Wayfinder map
- Keep the Metadata Catalog as a separate management subsystem and source of editable metadata. - Keep the Metadata Catalog as a separate management subsystem and source of editable metadata.
- Keep Qdrant derived and revision-scoped; it is not the catalog database or source of truth. - Keep Qdrant derived and revision-scoped; it is not the catalog database or source of truth.
@@ -253,3 +295,7 @@ The following should be treated as requirements for the architecture map:
the current descriptor. the current descriptor.
- Treat direct same-revision Qdrant writes, implicit CRUD publication, and bypassing the Git - Treat direct same-revision Qdrant writes, implicit CRUD publication, and bypassing the Git
snapshot boundary as rejected integration paths. snapshot boundary as rejected integration paths.
This summary remains design input. The authoritative current state is that catalog-to-core
integration and Sensitive Data Policy enforcement in schema-linking are deferred in
`PROJECT_STATE.md:160-171`.