docs: revalidate metadata catalog research

This commit is contained in:
Codex
2026-08-31 14:54:52 +02:00
parent 16b7477861
commit 866aee4249
3 changed files with 345 additions and 253 deletions
@@ -1,21 +1,80 @@
# Inventario delle funzionalità legacy di gestione metadati di ThothAI
Data: 2026-08-23
Data dell'inventario: 2026-08-23
Issue: `mptyl/ThothII#6`
Ultima verifica rispetto a ThothII: 2026-08-31
Stato: **inventario storico verificato; non è una specifica dello stato corrente**.
Issue originaria: [mptyl/ThothII#6](https://git.tylconsulting.it/mptyl/ThothII/issues/6)
Fonte primaria: repository legacy annidato `Thoth/ThothAI`, commit
`55855de0f18e5cb4bc72f0a2ab0a7317995186dd`.
## Sintesi
Le citazioni che iniziano con `/Thoth/ThothAI/` sono relative alla radice di quella copia legacy
fissata al commit indicato. Alla verifica del 2026-08-31 tutti i 68 riferimenti univoci puntavano a
file esistenti e a intervalli di righe validi.
La parità funzionale richiesta a ThothII comprende un catalogo dei database, l'inventario
> Questo documento conserva l'inventario e le evidenze degli anti-pattern di ThothAI. Le frasi di
> requisito nelle sezioni successive descrivono la baseline proposta il 2026-08-23; non prevalgono
> su ADR, contratti, codice o `PROJECT_STATE.md` correnti.
## Stato rispetto al codice corrente
### Fonti autorevoli correnti
Per capire cosa esiste oggi, usare nell'ordine:
- `PROJECT_STATE.md`, per lo snapshot operativo aggiornato;
- `CONTEXT.md`, per il modello di dominio corrente;
- il [piano accettato del Metadata Catalog](../plans/2026-08-26-metadata-catalog-from-thothai.md)
e il [contratto dello schema snapshot](../contracts/catalog-schema-snapshot.md);
- gli ADR del catalogo, in particolare
[0001](../adr/0001-postgres-metadata-catalog.md),
[0004](../adr/0004-fastify-kysely-metadata-catalog.md),
[0006](../adr/0006-separate-physical-and-logical-relationships.md),
[0007](../adr/0007-durable-authoritative-schema-synchronization.md),
[0008](../adr/0008-allow-manual-catalog-metadata-cleanup.md),
[0009](../adr/0009-use-one-sequential-description-generation-run.md),
[0010](../adr/0010-allow-bounded-real-source-samples-for-description-generation.md) e
[0011](../adr/0011-gate-source-samples-with-a-sensitive-data-flag.md);
- l'implementazione, soprattutto `backend/src/catalog/types.ts`,
`backend/src/routes/catalog-databases.ts`, `backend/src/routes/catalog-schema.ts`,
`backend/src/routes/catalog-description-generation.ts`,
`backend/src/routes/catalog-description-consolidation.ts` e
`frontend/src/shell/DatabaseManagementPage.tsx`.
### Disposizione della baseline storica
| Capacità inventariata | Disposizione al 2026-08-31 | Evidenza corrente o nota |
|---|---|---|
| Configurazione per workspace, secret separati e test connessione | **Adottata e implementata** | PostgreSQL interno, binding `postgres_direct`, `rest_api` o `ssh_tunnel`, secret write-only; ADR [0001](../adr/0001-postgres-metadata-catalog.md)–[0004](../adr/0004-fastify-kysely-metadata-catalog.md) e `backend/src/routes/catalog-databases.ts`. |
| Inventario fisico di tabelle, colonne, PK e FK | **Adottato e implementato** | La struttura osservata è distinta dai contenuti curati; `backend/src/catalog/types.ts`, `backend/src/routes/catalog-schema.ts` e ADR [0006](../adr/0006-separate-physical-and-logical-relationships.md). |
| Riconciliazione completa e osservabile dello schema | **Implementata; semantica storica parzialmente superata** | I durable Catalog Sync Runs sono autoritativi, fail-closed, atomici e richiedono conferma per diff distruttivi. Non esiste il lifecycle `drift/removed` proposto qui: la sincronizzazione riconcilia la membership e ADR [0008](../adr/0008-allow-manual-catalog-metadata-cleanup.md) consente anche cleanup manuale esplicito, superando ADR-0005. |
| Descrizione curata e Generated Description separate | **Adottata e implementata** | Tabelle e colonne espongono entrambi i campi in `backend/src/catalog/types.ts`; la copia selettiva AI → curato è in `backend/src/routes/catalog-description-consolidation.ts`. |
| Generazione AI selettiva con run, stato e log | **Adottata e implementata** | Un run asincrono installazione-wide, sequenziale, con eventi persistiti e senza resume automatico; ADR [0009](../adr/0009-use-one-sequential-description-generation-run.md) e `backend/src/routes/catalog-description-generation.ts`. I thread daemon legacy sono **esclusi**. |
| Campioni reali per la generazione e protezione dei campi sensibili | **Implementata come estensione correttiva** | ADR [0010](../adr/0010-allow-bounded-real-source-samples-for-description-generation.md) e [0011](../adr/0011-gate-source-samples-with-a-sensitive-data-flag.md): campioni bounded per colonne non sensibili e valori sintetici deterministici per quelle sensibili. |
| Proposte persistenti, diff e versioni dei testi AI | **Escluse** | Resta un solo Generated Description modificabile. La cronologia dei run è operativa: non conserva prompt, output grezzo, proposta per colonna o audit della decisione umana. |
| Relazioni fisiche | **Adottate e implementate** | Sono constraint immutabili con coppie ordinate di colonne; ADR [0006](../adr/0006-separate-physical-and-logical-relationships.md). |
| Relazioni logiche curate o inferite | **Differite/aperto** | ADR-0006 riserva un modello e un lifecycle separati; non sono ancora parte del catalogo corrente. |
| Alias semantici, descrizioni dei valori, sinonimi e concetti | **Differiti/aperti** | `PROJECT_STATE.md` li assegna a slice dedicate; non vanno dedotti dai campi fisici già implementati. |
| Scope AI, ERD Mermaid e documentazione aggregata | **Differiti/aperti** | Restano capacità legacy inventariate, non una feature corrente del Metadata Catalog. Un eventuale lavoro dovrà avere contratto e gate propri. |
| Export CSV e script SQL dei commenti | **Differiti/aperti; varianti insicure escluse** | Non risultano nella API corrente. Export di segreti, CSV incoerenti e SQL ricostruito da tipi incompleti restano vietati dagli anti-pattern sotto. |
| Inferenza euristica e validazione di relazioni candidate | **Differita/aperta** | Non va confusa con la sincronizzazione delle FK fisiche; dipende dal futuro lifecycle delle relazioni logiche. |
| Pubblicazione del catalogo al core/schema-linking/Qdrant | **Differita e richiesta come design gate successivo** | Il runtime NL→SQL continua a usare configurazione e annotations del workspace; il cutover è esplicitamente rinviato in `PROJECT_STATE.md`. |
| Sette motori database del legacy | **Baseline superata** | La prima versione corrente supporta PostgreSQL; l'aggiunta di altri dialetti è una decisione futura, non parità automatica. |
| GDPR e import da installazioni ThothAI | **Fuori dalla baseline iniziale; import differito** | GDPR resta escluso. Un eventuale import richiede una iniziativa idempotente e un cutover separati, non il riuso degli ID Django. |
| Django Admin, modifica manuale della struttura sorgente, password nel catalogo, duplicazione opaca dei database | **Esclusi** | La UI e le API correnti amministrano il catalogo e non eseguono DDL sul DWH esterno; binding e segreti hanno ownership separata. |
## Sintesi storica
La baseline di parità proposta il 2026-08-23 comprende un catalogo dei database, l'inventario
completo di tabelle, colonne e relazioni fisiche, metadati descrittivi modificabili, relazioni
logiche, introspezione dello schema, generazione AI di descrizioni, scope, ERD Mermaid,
documentazione ed esportazioni operative. La UI Django Admin è soltanto l'interfaccia legacy:
non è un requisito architetturale da riprodurre.
Il flusso AI per le descrizioni è volutamente semplice e va mantenuto tale:
Il flusso AI per le descrizioni individuato come baseline è volutamente semplice:
1. l'AI scrive nel campo `generated_comment` della tabella o colonna;
2. l'utente seleziona gli elementi desiderati;
@@ -193,7 +252,7 @@ scrive soltanto `column.generated_comment`
Provider e modello AI provengono dalla configurazione globale del backend; la lingua è invece
per database (`/Thoth/ThothAI/backend/thoth_core/thoth_ai/thoth_workflow/comment_generation_utils.py:156-219`).
### Comportamento di parità da implementare
### Comportamento di parità proposto il 2026-08-23
Il contratto funzionale minimo è esattamente questo:
@@ -314,7 +373,7 @@ Il modulo contiene anche un helper verso `mermaid.ink`, ma non risultano call si
legacy; il percorso attivo è il servizio locale. Non va quindi considerata una dipendenza
funzionale da conservare.
## 9. Baseline di parità per il Metadata Catalog di ThothII
## 9. Baseline di parità proposta il 2026-08-23
### Necessario
@@ -1,219 +1,206 @@
# PostgreSQL deployment and failure-isolation constraints for the Metadata Catalog
Date: 2026-08-23
Issue: [#8 — Assess PostgreSQL deployment and failure-isolation constraints](https://github.com/mptyl/ThothII/issues/8)
Original research: 2026-08-23
Last verified against the repository: 2026-08-31
Issue: [#8 — Assess PostgreSQL deployment and failure-isolation constraints](https://git.tylconsulting.it/mptyl/ThothII/issues/8)
> **Status: partially superseded by ADR-0004.** This is historical research, not the current
> architecture contract. The recommendation to run a separate `catalog-api` process was rejected
> by [ADR-0004](../adr/0004-fastify-kysely-metadata-catalog.md). Operational constraints that do
> not depend on that process boundary remain useful, but the status table below is authoritative
> for what the 2026-08-31 code actually adopts, rejects, or leaves pending.
## Question
How can a Metadata Catalog PostgreSQL service be deployed, backed up, diagnosed, and made
non-blocking for the core across local and server installations?
non-blocking for the NL→SQL workflow across local and server installations?
## Recommendation
## Current decision and implementation
Use PostgreSQL as the operational store, but place it behind a **separate Metadata Catalog API
process**. The SQL workflow `core` must not receive the catalog DSN, catalog credentials, a client
pool, or a Compose dependency on either the catalog API or PostgreSQL.
[ADR-0001](../adr/0001-postgres-metadata-catalog.md) selects PostgreSQL as the catalog authority.
[ADR-0003](../adr/0003-installation-local-database-bindings.md) keeps database bindings
installation-local. [ADR-0004](../adr/0004-fastify-kysely-metadata-catalog.md) then places the
catalog in the existing Fastify backend as an isolated Kysely module rather than a microservice.
The resulting dependency graph is:
The implemented dependency graph is therefore:
```mermaid
flowchart LR
UI[Frontend]
Core[Core workflow API]
Core[Core: Fastify workflow and catalog modules]
Pi[Pi and tht workflow]
PG[(catalog-db PostgreSQL)]
Semantic[Qdrant and embedding]
Catalog[Metadata Catalog API]
PG[(Catalog PostgreSQL)]
Publish[Publication worker]
UI -->|workflow routes| Core
UI --> Core
Core --> Pi
Core --> PG
Core --> Semantic
UI -->|catalog routes| Catalog
Catalog --> PG
PG --> Publish
Publish -->|immutable approved Publication| Semantic
```
The core continues to use the last successfully published, revision-qualified semantic snapshot.
Unpublished catalog edits are never a runtime dependency. This matches the domain boundary: the
Metadata Catalog does not select workflow SQL elements, while a Publication is an immutable value
made available to consumers ([`CONTEXT.md`, lines 96–110](../../CONTEXT.md#L96-L110)). It also
preserves the existing rule that Qdrant is a derived index rather than a canonical source
([`PROJECT_STATE.md`, lines 405–410](../../PROJECT_STATE.md#L405-L410)).
The catalog module is isolated behind a repository interface, but it is not process-isolated.
`core` owns the catalog pool and Compose waits for `catalog-db` health before starting `core`
(`compose.yaml`, `backend/src/app.ts`, `backend/src/catalog/repository.ts`). Database-management records still do
not feed the NL→SQL handoff, so a running workflow does not read catalog rows; that cutover remains
future work (`PROJECT_STATE.md`).
## Evidence from the current architecture
## Verification result
1. The mandatory workflow stack is currently `frontend`, `core`, `qdrant`, `embedding`, and the
model-init job. `core` waits only for Qdrant and model initialization
([`compose.yaml`, lines 50–60](../../compose.yaml#L50-L60)); the frontend waits only for `core`
([`compose.yaml`, lines 123–141](../../compose.yaml#L123-L141)). Adding a catalog dependency to
either chain would enlarge the workflow failure domain.
2. `tht start` runs Compose and then checks an explicit allowlist containing only those five
mandatory services. Other Compose services are ignored by the readiness fold
([`service.go`, lines 40–57](../../tools/tht/internal/service/service.go#L40-L57),
[`service.go`, lines 166–190](../../tools/tht/internal/service/service.go#L166-L190)). Separate
catalog services can therefore start with the installation without redefining core readiness.
3. `/health` is deliberately a process-liveness endpoint and does not probe external services
([`app.ts`, lines 271–284](../../backend/src/app.ts#L271-L284)). Tests preserve a 200 response
even when PostgreSQL session storage is configured but not contacted
([`health.test.ts`, lines 9–32](../../backend/test/health.test.ts#L9-L32)). Catalog readiness must
not be folded into this endpoint.
4. ThothII already has a sound server-PostgreSQL precedent: runtime and migrator credentials are
distinct; the one-shot migrator is profile-gated and is explicitly not a dependency of `core`
([`compose.session-server.yaml.example`, lines 24–49](../../deploy/compose.session-server.yaml.example#L24-L49)).
Its documented failure behavior is route-scoped 503 while process liveness remains healthy
([`PROJECT_STATE.md`, lines 621–636](../../PROJECT_STATE.md#L621-L636)).
5. Existing backup archives contain selected Compose volumes only. Local archives include seven
named volumes and server archives include only Qdrant and embedding volumes; no PostgreSQL
logical dump exists today ([`create.go`, lines 29–52](../../tools/tht/internal/backup/create.go#L29-L52)).
A catalog database therefore requires an explicit backup contract rather than an assumption
that current installation backup already covers it.
## Deployment contract
### Shared rules
- Run `catalog-api` as a separate, non-privileged service. The frontend reverse proxy may route
`/api/catalog/*` to it, but frontend startup must not depend on catalog readiness.
- Give `catalog-api` a bounded pool and bounded operations. As an initial ceiling for this small
administrative workload: pool size 5, overflow 0–2, `connect_timeout=3`, `lock_timeout=2s`, and
a normal request `statement_timeout=10s`. Long AI calls must not hold database transactions.
PostgreSQL documents that an omitted or zero `connect_timeout` waits indefinitely, so an
explicit value is required for failure isolation
([libpq connection parameters](https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-CONNECT-TIMEOUT)).
- Use three roles: `catalog_runtime` (CRUD and job claims only), `catalog_migrator` (DDL, used only
by a one-shot job), and `catalog_backup` (minimum privileges needed by the approved dump policy).
No role is shared with the DWH or `thoth_sessions`, and none is a PostgreSQL superuser.
- Store passwords and the CA as separate secret files. For non-local connections default to
`sslmode=verify-full`; this follows the current server-session configuration
([`compose.session-server.yaml.example`, lines 5–20](../../deploy/compose.session-server.yaml.example#L5-L20))
and PostgreSQL's hostname-verifying TLS guidance
([libpq SSL parameters](https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-PARAMKEYWORDS)).
- Apply ordered, checksummed migrations under one transaction and an advisory lock. Refuse pending,
unknown, or checksum-drifted migrations in catalog readiness. The session repository already
implements this pattern
([`postgres_repository.py`, lines 118–187](../../harness/tht/session/postgres_repository.py#L118-L187));
the catalog should have its own migration table, lock key, schema, and roles.
### Local installation
Add an installation-owned `catalog-postgres` container on the private `thothii` network, with no
published host port and a dedicated `catalog-postgres-data` named volume. Pin the image by version
and digest, add a `pg_isready` healthcheck, and let only `catalog-api` use
`depends_on: condition: service_healthy`. `core` and `frontend` keep their existing dependency
graph. Docker confirms that `service_healthy` delays only the declaring dependent service, while
plain container start does not imply application readiness
([Compose startup order](https://docs.docker.com/compose/how-tos/startup-order/)).
The local catalog services can be included in the normal installation Compose set because the
host CLI's mandatory-service allowlist excludes them. A separate `catalog-migrate` one-shot profile
retains explicit operator control; Docker profiles are intended for selectively activated and
one-off services ([Compose profiles](https://docs.docker.com/compose/how-tos/profiles/)).
### Server installation
Use an operator-provided PostgreSQL endpoint configured through a reviewed, untracked overlay.
Prefer a separate database and dedicated roles. A separate PostgreSQL cluster gives the strongest
resource and outage isolation; sharing the existing server cluster is acceptable only when the
operator accepts the residual cluster-wide blast radius and enforces connection limits, timeouts,
and separate ownership. Merely using another schema does not isolate connection exhaustion,
maintenance outages, WAL pressure, or storage failure.
The server overlay adds secrets and catalog configuration only to `catalog-api` and the one-shot
migrator. It must not mount those secrets into `core`. This mirrors the current rule that `core`
never receives the session migrator credential
([`PROJECT_STATE.md`, lines 623–630](../../PROJECT_STATE.md#L623-L630)).
## Failure behavior
| Failure | Catalog behavior | Workflow behavior |
| Area | Status on 2026-08-31 | Evidence and consequence |
| --- | --- | --- |
| PostgreSQL unavailable or connection pool exhausted | Catalog routes return sanitized `503 catalog_unavailable` with `Retry-After`; UI remains read-only/unavailable | Existing sessions and new SQL workflow requests continue against the last published snapshot |
| Catalog API process unavailable | `/api/catalog/*` fails; main frontend shell and workflow route remain usable | No effect |
| Pending or drifted migration | Catalog readiness fails and writes are refused | No effect |
| AI generation fails | Job records a retryable/terminal failure; no database transaction remains open | No effect |
| Publication indexing fails | Publication remains non-active/failed and can be retried idempotently | Previous active Publication remains in Qdrant |
| Qdrant unavailable | Publication is queued/failed without losing canonical catalog state | Existing core readiness rules already govern this independent failure |
| PostgreSQL as canonical catalog store | **Adopted** | ADR-0001 is implemented by the PostgreSQL-backed Kysely repository and the internal `catalog-db` service. |
| One Workspace Database per workspace and installation-local bindings | **Adopted** | ADR-0003 and the catalog migrations enforce the model; workspace identity remains in the workspace registry. |
| Separate `catalog-api` process | **Superseded/rejected** | ADR-0004 explicitly chooses an isolated module inside the existing Fastify process. There is no catalog microservice or separate catalog liveness endpoint. |
| No catalog pool, credential, or Compose dependency in `core` | **Superseded/rejected** | `core` owns the runtime pool, receives the runtime password secret, and declares `depends_on: catalog-db: service_healthy`. The old process-level isolation acceptance criterion is not current architecture. |
| Private PostgreSQL service and durable volume | **Adopted** | `catalog-db` uses a version-and-digest-pinned PostgreSQL 17.6 image, the private `thothii` network, no published host port, a `pg_isready`-based healthcheck, and `catalog-data`. Compose contract tests assert this topology (`scripts/test-default-compose.sh`). |
| Different local and server database topology | **Superseded/rejected** | Both current profiles inherit the same internal `catalog-db` and named volume. `deploy/compose.server.yaml` does not replace it with an operator-provided endpoint. |
| Runtime and migrator role separation | **Adopted** | `core` receives only `thothii_catalog_runtime`; the profile-gated `catalog-migrate` job receives only `thothii_catalog_migrate`. Bootstrap grants runtime DML/sequence privileges without DDL (`docker/catalog-db-init.sql`). |
| Dedicated backup role | **Pending** | There is no `catalog_backup` role or backup secret. |
| Explicit one-shot migrations | **Adopted** | `backend/src/catalog/migrate.ts` registers ordered Kysely migrations and uses a pool of one. `scripts/run-stack.sh` starts PostgreSQL and runs `catalog-migrate` before local startup; migrations are not hidden in backend startup. |
| Checksums, drift/pending refusal, and migration readiness | **Pending** | No repository-owned checksum policy, drift report, or readiness gate exists for catalog migrations. The backend can start without checking the Kysely migration head when launched outside the local helper. |
| Bounded runtime connection pool | **Adopted** | `backend/src/catalog/repository.ts` sets `max: 5` and `connectionTimeoutMillis: 3000`, and closes Kysely with the Fastify lifecycle. |
| `lock_timeout`, `statement_timeout`, and catalog TLS policy | **Pending** | The runtime pool does not set query or lock timeouts. Catalog connection configuration has no explicit CA/hostname-verification contract; the server profile still uses the private Compose network. |
| Process liveness independent of catalog queries | **Adopted** | `GET /health` returns `{status: "ok"}` without probing PostgreSQL, and `backend/test/health.test.ts` preserves that behavior. `/catalog/status` performs the catalog-specific availability check. |
| Stack startup independent of catalog availability | **Superseded/rejected** | Compose blocks `core` on healthy `catalog-db`. The host service health fold does not list `catalog-db`, but it cannot make `core` start while its Compose dependency is unhealthy. |
| Uniform catalog-outage response (`503` plus `Retry-After`) | **Pending** | Routes map the domain `CatalogUnavailableError` to a sanitized `503`, and an omitted catalog configuration uses an unavailable repository. PostgreSQL driver failures are not uniformly translated to that domain error, and no `Retry-After` contract is implemented. |
| Catalog-specific readiness and `tht doctor` checks | **Pending** | There is no `/health/ready` for the catalog and no doctor section for connection, migration head, pool saturation, backup age, or publication lag (`tools/tht/internal/doctor/report.go`). |
| Logical catalog backup and restore | **Pending** | Current backup archives omit `catalog-data` and do not run `pg_dump`; server archives include only Qdrant and embedding volumes. A backup can therefore succeed without preserving the Metadata Catalog (`tools/tht/internal/backup/create.go`). |
| Last-good Publication and catalog-to-Qdrant cutover | **Pending** | The current database-management slice does not change the NL→SQL runtime or publish catalog metadata to Qdrant. |
Do not silently fall back to JSON files, dual-write, or read unpublished PostgreSQL rows from the
core. Those paths would create split-brain state. The only degraded-mode contract is use of the
last successfully activated Publication.
## Current operational contract
## Backup and restore
### Deployment and credentials
- Local and server Compose profiles currently use the same installation-owned `catalog-db`
container and `catalog-data` named volume. PostgreSQL is private to the Compose network.
- The database bootstrap login is the migrator. The init script creates the separate runtime login
from a Docker secret and grants only runtime DML and sequence access.
- Runtime and migrator passwords are separate protected host files exposed as separate Docker
secrets. Neither value belongs in tracked environment files
(`deploy/env/local.env.example`, `deploy/env/server.env.example`).
- The Fastify catalog repository uses a five-connection pool with a three-second connection
timeout. Fastify closes the repository pool during shutdown.
### Migrations
The compiled `catalog-migrate` entry point owns schema changes and uses the migrator credential.
The current ordered series is under
`backend/src/catalog/migrations/`. The local launcher runs
it before the normal stack, while production rollout must invoke the profile-gated service
explicitly.
This is weaker than the original recommendation in two ways: there is no catalog readiness check
for pending or unknown migrations, and the repository does not maintain content checksums for
migration drift. Until those checks exist, “the migrator completed” is the available deployment
gate; the application itself does not prove migration compatibility.
### Failure behavior actually provided
| Failure | Current catalog behavior | Current workflow behavior |
| --- | --- | --- |
| Catalog configuration omitted when launching the backend directly | Fastify uses `UnavailableCatalogRepository`; `/catalog/status` reports unavailable and domain-mapped catalog operations return sanitized `503` | `/health`, sessions, and SSE remain available |
| `catalog-db` unhealthy before Compose startup | `core` is not started because its dependency is not healthy | Workflow startup is blocked |
| PostgreSQL becomes unavailable after startup | `/health` remains process-only and `/catalog/status` reports unavailable; route errors are sanitized, but a uniform `503`/`Retry-After` mapping is not guaranteed | Existing session code does not read catalog rows, but both surfaces still share one Fastify process |
| Migrations are pending or incompatible | No dedicated readiness refusal exists; affected catalog operations fail | No catalog-to-workflow handoff exists yet, but local deployment correctness depends on running `catalog-migrate` first |
| Catalog backup is requested through current `tht backup` | No PostgreSQL dump is added and `catalog-data` is omitted | The archive may succeed while being unable to restore catalog state |
The shared process means resource exhaustion, fatal process errors, and startup hooks remain a
common failure domain even though repository calls are separated. Conversely, placing the module
inside Fastify does not require workflow code to consume catalog rows: preserving that data-flow
boundary is still the useful part of the original isolation recommendation.
## Historical recommendations that remain valid backlog
The following constraints survive ADR-0004 because they can be implemented inside the current
Fastify deployment:
1. **Bound every database operation.** Keep the adopted pool and connection timeout; add explicit
request query and lock timeouts. Long model calls and source sampling must not hold catalog
transactions.
2. **Make failures component-specific.** Translate connection, timeout, and pool-exhaustion errors
into one sanitized catalog-unavailable response, add bounded retry guidance, and keep `/health`
process-only.
3. **Make migration compatibility observable.** Report applied, pending, unknown, and drifted
migrations through a catalog readiness check and `tht doctor`; do not put migrator credentials
in `core`.
4. **Define a server TLS topology before externalizing PostgreSQL.** If the server profile moves to
an operator-provided endpoint, use a dedicated database and roles, protected CA material, and
hostname verification. Sharing a cluster leaves connection, maintenance, WAL, and storage blast
radius even when schemas are separate. PostgreSQL documents the relevant connection and TLS
parameters in its
[connection parameter reference](https://www.postgresql.org/docs/current/libpq-connect.html).
5. **Keep derived semantic data non-canonical.** When catalog publication is implemented, activate
complete immutable revisions and retain the last successfully activated revision rather than
exposing mutable catalog rows to the workflow.
## Backup and restore target
The original logical-backup recommendation is still valid and is now a confirmed implementation
gap.
### Backup
Extend `tht backup create` or add a catalog-specific command invoked by it. Create a logical
`pg_dump --format=custom` archive through a one-shot helper, checksum it, and place it in the
installation archive manifest. Do **not** tar a live PostgreSQL data volume: the existing raw-volume
backup mechanism was designed for the current named volumes, not database crash consistency.
Extend `tht backup create` with a catalog step that runs `pg_dump --format=custom` through a
one-shot helper and adds the dump plus checksum to the installation manifest. Do not treat a tar of
a live data volume as a PostgreSQL consistency contract. PostgreSQL documents that `pg_dump`
creates a consistent export while the database remains in use and that custom format supports
selective restore ([`pg_dump`](https://www.postgresql.org/docs/current/app-pgdump.html)).
`pg_dump` produces a consistent export while the database remains in use, and custom format is
compressed and supports selective/reordered restore
([PostgreSQL `pg_dump`](https://www.postgresql.org/docs/current/app-pgdump.html)). Because metadata
volume is small, favor correctness over parallelism. Run a backup before every migration and on an
operator-defined schedule; record the server version, migration head, dump checksum, creation time,
and catalog Publication identifiers in the manifest. Never write a dump into the workspace Git
repository.
For an external server database, the operator must supply the dedicated backup credential through
the same secret-file boundary. A successful filesystem/Qdrant archive with a failed catalog dump
is an incomplete installation backup and must not be reported as success.
Use a dedicated least-privilege backup role, never write dumps into a workspace repository, and
record at least the server version, migration head, checksum, creation time, and catalog revision
identifiers. A filesystem/Qdrant archive with a failed or absent catalog dump must be reported as
incomplete.
### Restore
Restore only during an explicit catalog maintenance window into an empty target database (or a
deliberately cleaned target), using `pg_restore --exit-on-error --single-transaction`. Validate the
archive checksum first, then require:
Restore only through an explicit maintenance operation into an empty or deliberately cleaned
target. Validate the archive checksum, run `pg_restore --exit-on-error --single-transaction`, then
verify migration compatibility, referential integrity, bounded entity counts, and an authenticated
catalog smoke test. The controls are documented by PostgreSQL
([`pg_restore`](https://www.postgresql.org/docs/current/app-pgrestore.html)). Keep the old database
until validation passes; rollback should switch the endpoint or retained volume, not dual-write.
1. expected migration history with no checksum drift;
2. referential-integrity checks and bounded entity counts;
3. Publication content hashes matching the restored records;
4. an authenticated catalog smoke test.
## Diagnostics target
The relevant restore controls and their behavior are documented by
[PostgreSQL `pg_restore`](https://www.postgresql.org/docs/current/app-pgrestore.html). After restore,
rebuild catalog-derived Qdrant content from the restored active Publication rather than treating
the vector index as the source of truth. Keep the old database untouched until this validation
passes; rollback is endpoint reversal, not dual-write.
- Keep the shared `GET /health` endpoint process-only.
- Treat `/catalog/status` as the current minimal availability surface; add a bounded catalog
readiness check covering connection and migration compatibility before using it as a rollout
gate.
- Add read-only `tht doctor` checks for secret-file presence and permissions, PostgreSQL
reachability, authentication, migration state, pool saturation, and last successful catalog
backup. Sanitize all DSNs and errors.
- Use `pg_isready` only for server transport readiness. Its result does not prove schema or runtime
authorization correctness
([`pg_isready`](https://www.postgresql.org/docs/current/app-pg-isready.html)).
- Add structured metrics for connection acquisition failures, pool use, query/lock timeout,
migration head, background-job backlog, and backup age. Do not log connection strings, source
samples, prompts, or generated descriptions at info level.
## Diagnostics and observability
## Recommended closure criteria
- Keep `core /health` unchanged. Give `catalog-api` separate `/health/live` (process only) and
`/health/ready` (bounded connection, `SELECT 1`, and migration compatibility) endpoints.
- Add a named `metadata-catalog` section to `tht doctor`: configuration completeness; secret-file
presence and permissions; DNS/TCP/TLS/authentication; `SELECT 1`; migration applied/pending/drifted
state; pool saturation; last successful backup age; and oldest pending Publication. Checks are
read-only and all errors/DSNs are sanitized. The current doctor already separates Compose service
status from container-local workflow checks
([`report.go`, lines 275–310](../../tools/tht/internal/doctor/report.go#L275-L310)).
- Use `pg_isready` only for transport readiness; its exit codes distinguish accepting, rejecting,
no-response, and invalid-parameter states, but it does not prove schema or authorization
correctness ([PostgreSQL `pg_isready`](https://www.postgresql.org/docs/current/app-pg-isready.html)).
- Expose structured metrics/logs for connection acquisition failures, pool utilization, query
timeout/lock timeout, migration head, AI-job backlog, Publication lag, and backup age. Never log
connection strings, prompts containing schema metadata, or generated descriptions at info level.
- `tht status` may show the catalog as `healthy`, `degraded`, or `disabled`, but catalog state must
not change the existing core health result. `tht doctor` may return a failing catalog check for
operator visibility without stopping or restarting the workflow stack.
## Acceptance constraints for later design and implementation
1. Killing `catalog-postgres` and `catalog-api` must not change `core /health`, stop Pi, interrupt an
existing SQL session, or prevent a new session that uses the already active Publication.
2. No catalog database secret, client dependency, migration, or network call exists in the harness
workflow path.
3. Catalog failure is explicit on catalog routes; no JSON fallback or dual-write is introduced.
4. A failed Publication leaves the previous active semantic content addressable and retryable.
5. Local and server backup tests perform a real dump/restore round trip and prove that a backup is
incomplete when the PostgreSQL dump fails.
6. Migration credentials are absent from runtime containers; drift and pending migrations are
visible in readiness and `tht doctor`.
1. Decide whether ADR-0004's statement that catalog unavailability does not make sessions or SSE
unavailable must also hold at Compose startup. If yes, remove or soften the hard `core` →
`catalog-db` startup dependency without reintroducing a microservice.
2. Prove a configured PostgreSQL outage produces the same sanitized catalog response across every
catalog route while `/health`, session creation, and SSE continue to work.
3. Add catalog migration compatibility to readiness and `tht doctor`, including pending, unknown,
and drifted states.
4. Add a real `pg_dump`/`pg_restore` round-trip to local and server backup tests, and fail backup
publication when the catalog dump is absent or fails.
5. Before enabling a remote server catalog, define TLS verification, credential files, connection
limits, and the accepted cluster-level blast radius.
6. Before the NL→SQL cutover, define and test an immutable last-good Publication boundary; do not
dual-read mutable PostgreSQL rows and legacy metadata as competing authorities.
## Decision summary
The Metadata Catalog should be a separately deployed bounded context backed by PostgreSQL. Local
installations own a private container and volume; server installations use a TLS-verified external
database, preferably on a separate cluster when strict blast-radius isolation is required. Logical
dump/restore, one-shot migrations, role separation, bounded connections, component-specific
diagnostics, and last-good-Publication semantics make PostgreSQL operationally safe without making
it a dependency of the SQL workflow.
PostgreSQL, installation-local bindings, private deployment, role separation, one-shot migrations,
bounded pooling, and process-only liveness are implemented. The separate `catalog-api` process and
the claim that `core` has no catalog dependency are not current design: ADR-0004 chose the existing
Fastify process, and Compose currently blocks `core` startup on `catalog-db` health. Logical
backup/restore, catalog migration readiness and drift detection, consistent outage mapping,
catalog-specific doctor checks, server TLS/external topology, and last-good publication remain
pending. Those open constraints should be treated as backlog, not as capabilities already provided
by the repository.
@@ -4,21 +4,51 @@
session-pinning, and runtime read-only contracts constrain metadata publication without changing
the NL→SQL workflow?
**Last verified:** 2026-08-31
**Validity:** Active architectural research; the catalog-to-core integration described below is
still **deferred**, not implemented.
**Current authorities:** `PROJECT_STATE.md`,
[`workspace-evidence-v3.md`](../contracts/workspace-evidence-v3.md),
[`workspace-preprocessing-cli.md`](../contracts/workspace-preprocessing-cli.md),
[`tht-dwh.md`](../contracts/tht-dwh.md),
[`ADR-0001`](../adr/0001-postgres-metadata-catalog.md), and
[`ADR-0004`](../adr/0004-fastify-kysely-metadata-catalog.md). These sources and current code
override this research note if they diverge.
## Revalidation against the current implementation
| Finding | Status on 2026-08-31 | Current evidence and consequence |
| --- | --- | --- |
| PostgreSQL metadata authority in the existing Fastify backend | **Adopted/current** | ADR-0001 and ADR-0004 are implemented; the catalog is the management-plane authority. |
| Git workspace revision, immutable registry snapshot, and revision-pinned runtime | **Adopted/current** | Registry activation and the Workspace Evidence v3 contract still provide the publication boundary for runtime-owned files. |
| Qdrant read isolation for schema and Evidence by `workspace_id` + `workspace_revision` | **Adopted/current** | `QdrantVectorStore.search()` applies `_revision_filter()` to schema/Evidence; Memory and solved questions intentionally remain workspace-wide (`harness/tht/adapters/vector/qdrant.py`). |
| Physical schema as a file in the Git publication | **Superseded clarification** | `physical.yaml` is owned by the immutable `.tht-dwh` generation selected by `ACTIVE`, not by the Git workspace revision ([DWH contract](../contracts/tht-dwh.md)). Git currently supplies the revision-pinned curated `schema/annotations.yaml`. |
| Catalog-to-core publisher | **Open/deferred** | No current route or service projects catalog records into core artifacts. `PROJECT_STATE.md` explicitly defers the schema-linking integration (`PROJECT_STATE.md:160-171`). |
| Explicit Core Schema Selection | **Open/deferred** | The workspace descriptor selects one database and one physical schema, but has no table/column allowlist (`backend/src/workspaces/schema.ts`); no catalog selection is handed to `tht`. |
| Cross-revision hash lookup used by schema synchronization | **Open defect** | Reads are revision-filtered, but `existing_hashes()` is not. A same-key/same-content point from an older revision can suppress the required upsert into the new revision (`harness/tht/adapters/vector/qdrant.py`, `harness/tht/cli/vector_cmd.py`). |
| Deletion/GC | **Current for Evidence; open for schema** | Evidence has generation inventory, retention, compensation, and exact-generation deletion. `sync_canonical_records()` never deletes schema records absent from the new canonical set, and no revision-retention GC exists for schema points. |
| Annotation consumption | **Current but incomplete** | M-Schema rendering and schema embeddings consume `Annotations`; `tht schema columns` still returns only physical comments, so F4 does not display annotation descriptions (`harness/tht/cli/schema_cmd.py`, `harness/.pi/extensions/tht-gate.js`). |
## Conclusion
The lowest-impact publication seam is **not** a new writer inside the NL→SQL workflow and it is
not a direct CRUD-to-Qdrant path. The existing core already consumes two canonical schema inputs:
an introspected `PhysicalSchema` and a curated `Annotations` document. An explicit publication
operation should project the approved, workspace-selected subset of the new metadata catalog into
those inputs, bind the projection to a new Git `workspace_revision`, activate that immutable
The lowest-impact **proposed** publication seam is not a new writer inside the NL→SQL workflow and
is not a direct CRUD-to-Qdrant path. The existing core consumes an introspected `PhysicalSchema`
from the active immutable DWH generation and a curated, Git-revision-pinned `Annotations`
document. A future explicit publication operation should project the approved catalog subset into
a new Core Schema Selection contract and the curated annotations, activate a new immutable Git
revision, and then invoke the existing `workspace index-schema` preprocessing operation. Qdrant
remains a derived, rebuildable projection.
remains a derived, rebuildable projection. None of this catalog-to-core handoff exists yet;
`PROJECT_STATE.md:160-164` deliberately defers it to the next design
gate.
This preserves the existing runtime path:
The proposed flow would preserve the existing runtime path:
```text
approved catalog data
-> workspace Git revision (physical schema selection + annotations)
-> explicit Core Schema Selection + curated annotations projection (future)
-> workspace Git revision (annotations and selection contract; physical.yaml stays DWH-owned)
-> immutable registry snapshot
-> revision-bound runtime configuration
-> existing tht vector index-schema
@@ -26,23 +56,23 @@ approved catalog data
-> existing retrieval_pack / schema render / F4 review
```
The complete database inventory may remain in the Metadata Catalog. Only the explicit Core Schema
Selection should be projected into the workspace artifacts used by the SQL workflow. ThothII does
not currently model such a table-level selection in its descriptor, so that projection contract is
new work.
The complete database inventory remains in the Metadata Catalog. Only an explicit Core Schema
Selection and approved semantic fields should be projected into the artifacts used by the SQL
workflow. ThothII does not currently model or publish that table/column-level selection, so both
the projection contract and its operational publisher are new work.
## 1. Current authority and publication boundary
The workspace descriptor identifies one PostgreSQL database and one physical schema, plus one
workspace-owned Qdrant collection; it has no table or column allowlist
([`backend/src/workspaces/schema.ts:160-178`](../../backend/src/workspaces/schema.ts#L160-L178)).
(`backend/src/workspaces/schema.ts`).
Consequently, a full-database metadata catalog and the subset eligible for the SQL core cannot be
represented as the same current descriptor object.
The existing public contract makes the Git workspace repository curator-owned. Changes occur in a
separate authoring clone followed by installation pull, and the API does not write workspace,
schema, or Evidence paths
([`docs/contracts/workspace-evidence-v3.md:148-156`](../contracts/workspace-evidence-v3.md#L148-L156)).
([Workspace Evidence v3 contract](../contracts/workspace-evidence-v3.md)).
Therefore a browser CRUD service cannot silently make its PostgreSQL state authoritative for the
core without either:
@@ -54,10 +84,12 @@ The first option preserves current architecture and session reproducibility.
Registry activation validates every descriptor at one exact commit, validates and synchronizes
the commit's `schema/annotations.yaml`, and rejects duplicate ownership of a Qdrant collection
([`backend/src/workspaces/registry.ts:531-576`](../../backend/src/workspaces/registry.ts#L531-L576)).
(`backend/src/workspaces/registry.ts`,
`backend/src/workspaces/git-repository.ts`,
`backend/src/workspaces/annotations-sync.ts`).
It writes the candidate snapshot into a staging directory, records file digests in
`snapshot.json`, atomically renames the directory, and only then moves active state
([`backend/src/workspaces/registry.ts:582-629`](../../backend/src/workspaces/registry.ts#L582-L629)).
(`backend/src/workspaces/registry.ts`).
That is the existing atomic publication boundary to reuse.
## 2. Canonical schema inputs already consumed by the core
@@ -67,15 +99,14 @@ The harness separates source facts from curated semantics:
- `PhysicalSchema` contains database/schema identity and tables; table facts include comments,
columns, physical foreign keys, and indexes; columns contain type, nullability, primary-key,
default, source comment, examples, and eligibility
([`harness/tht/mschema/models.py:23-64`](../../harness/tht/mschema/models.py#L23-L64)).
(`harness/tht/mschema/models.py`).
- `Annotations` contains curated table descriptions, concepts and notes, column descriptions,
synonyms, concepts, evidence, notes and eligibility overrides, plus logical foreign keys
([`harness/tht/mschema/models.py:67-88`](../../harness/tht/mschema/models.py#L67-L88)).
(`harness/tht/mschema/models.py`).
Rendering already gives annotations precedence over source comments and merges physical and
logical foreign keys. It also applies column eligibility before producing M-Schema context
([`harness/tht/mschema/render.py:9-43`](../../harness/tht/mschema/render.py#L9-L43),
[`harness/tht/mschema/render.py:46-80`](../../harness/tht/mschema/render.py#L46-L80)). This makes
(`harness/tht/mschema/render.py`). This makes
`Annotations` the natural narrow projection target for approved descriptions, synonyms, logical
relationships, and eligibility from the new catalog.
@@ -83,14 +114,14 @@ There are two current compatibility gaps:
- `schema_records()` embeds table descriptions/concepts and column descriptions/synonyms/examples,
but does not include physical or logical foreign keys in vector record content
([`harness/tht/vectorstore/records.py:101-132`](../../harness/tht/vectorstore/records.py#L101-L132)).
(`harness/tht/vectorstore/records.py`).
Relationships still reach the model through deterministic schema rendering, not through schema
candidate embeddings.
- The F4 widget loads columns with `tht schema columns`
([`harness/.pi/extensions/tht-gate.js:1220-1254`](../../harness/.pi/extensions/tht-gate.js#L1220-L1254)),
but that command currently returns only `PhysicalSchema.comment` values and does not merge
(`harness/.pi/extensions/tht-gate.js`), but that
command still returns only `PhysicalSchema.comment` values and does not merge
`Annotations`
([`harness/tht/cli/schema_cmd.py:592-621`](../../harness/tht/cli/schema_cmd.py#L592-L621)).
(`harness/tht/cli/schema_cmd.py`).
A publisher that writes only annotations would improve vector search and rendered M-Schema but
not the table/column descriptions displayed by this existing reviewer widget. Fixing the command
to use the existing merged description helpers would preserve the workflow shape while closing
@@ -101,32 +132,33 @@ There are two current compatibility gaps:
`workspace index-schema` is already the supported operator seam. It creates a revision-bound
runtime, checks collection compatibility, and runs the harness command
`vector index-schema --json`
([`backend/src/workspaces/preprocessing-service.ts:312-330`](../../backend/src/workspaces/preprocessing-service.ts#L312-L330)).
(`backend/src/workspaces/preprocessing-service.ts`,
[`workspace-preprocessing-cli.md`](../contracts/workspace-preprocessing-cli.md)).
The full preprocessing operation performs DWH preparation, FK suggestion/review, schema indexing,
and optional Evidence preprocessing as separate resumable stages
([`backend/src/workspaces/preprocessing-service.ts:367-450`](../../backend/src/workspaces/preprocessing-service.ts#L367-L450)).
(`backend/src/workspaces/preprocessing-service.ts`).
Metadata-only publication should normally use the narrow `index-schema` operation after its
workspace artifacts are valid, rather than coupling catalog CRUD to the full pipeline.
The harness indexer reads the active immutable physical schema and revision-specific annotations,
constructs schema records, embeds only changed content, and writes them through the configured
vector adapter
([`harness/tht/cli/vector_cmd.py:100-140`](../../harness/tht/cli/vector_cmd.py#L100-L140)). Its
(`harness/tht/cli/vector_cmd.py`). Its
machine result carries the physical/annotation artifact digests, workspace revision, collection,
and counts
([`harness/tht/cli/vector_cmd.py:157-179`](../../harness/tht/cli/vector_cmd.py#L157-L179)). This is
(`harness/tht/cli/vector_cmd.py`). This is
the right place to retain publication evidence and audit linkage.
Preprocessing state is already revision- and binding-aware. A resumed job must match operation,
workspace revision, descriptor/catalog blobs, runtime config and binding digests or it fails with
`preprocessing_resume_mismatch`
([`backend/src/workspaces/preprocessing-state.ts:262-313`](../../backend/src/workspaces/preprocessing-state.ts#L262-L313)).
(`backend/src/workspaces/preprocessing-state.ts`).
There is an important operational gate: every preprocessing operation calls
`assertSessionInventoryCompatible`; a non-finalized, non-archived session pinned to another
revision blocks preprocessing
([`backend/src/workspaces/preprocessing-service.ts:459-472`](../../backend/src/workspaces/preprocessing-service.ts#L459-L472),
[`backend/src/workspaces/preprocessing-state.ts:363-378`](../../backend/src/workspaces/preprocessing-state.ts#L363-L378)).
(`backend/src/workspaces/preprocessing-service.ts`,
`backend/src/workspaces/preprocessing-state.ts`).
A catalog publication UX must expose this as a pending/blocking condition rather than report a
generic indexing failure.
@@ -135,23 +167,22 @@ generic indexing failure.
The collection contract is fixed at the descriptor's dimensions/distance and eight keyword payload
indexes: `content_hash`, `document_id`, `kind`, `record_key`, `record_kind`,
`vector_generation`, `workspace_id`, and `workspace_revision`
([`backend/src/workspaces/qdrant-collection.ts:1-20`](../../backend/src/workspaces/qdrant-collection.ts#L1-L20)).
(`backend/src/workspaces/qdrant-collection.ts`).
`self_heal` may create a missing compatible collection or indexes; `require_existing` only validates
and refuses an incompatible collection
([`backend/src/workspaces/qdrant-collection.ts:71-104`](../../backend/src/workspaces/qdrant-collection.ts#L71-L104)).
(`backend/src/workspaces/qdrant-collection.ts`).
The publication path must use this shared manager instead of inventing collection setup.
Schema payloads include both the workspace and workspace revision, record identity, semantic kind,
content and content hash
([`harness/tht/vectorstore/records.py:30-49`](../../harness/tht/vectorstore/records.py#L30-L49)).
(`harness/tht/vectorstore/records.py`).
Qdrant point IDs for schema and Evidence also include the revision, and upserts use the same
revision in the payload
([`harness/tht/adapters/vector/qdrant.py:42-46`](../../harness/tht/adapters/vector/qdrant.py#L42-L46),
[`harness/tht/adapters/vector/qdrant.py:196-220`](../../harness/tht/adapters/vector/qdrant.py#L196-L220)).
(`harness/tht/adapters/vector/qdrant.py`).
Reads always filter by `workspace_id`; schema and Evidence reads additionally filter by the bound
`workspace_revision`, while memory and solved-question records intentionally remain
workspace-wide
([`harness/tht/adapters/vector/qdrant.py:291-308`](../../harness/tht/adapters/vector/qdrant.py#L291-L308)).
(`harness/tht/adapters/vector/qdrant.py:125-145`).
This means an approved semantic change needs a new workspace revision if old sessions must retain
their previous view. Directly overwriting points under the same Git revision would mutate the
@@ -160,43 +191,52 @@ would require changing the current runtime filter contract.
### Qdrant synchronization defect to resolve before catalog publication
The current incremental synchronizer compares canonical records by content hash and upserts
The current incremental schema synchronizer compares canonical records by content hash and upserts
changed records, but it has no deletion step
([`harness/tht/cli/vector_cmd.py:67-89`](../../harness/tht/cli/vector_cmd.py#L67-L89)). More
(`harness/tht/cli/vector_cmd.py:77-99`). More
importantly, `QdrantVectorStore.existing_hashes()` filters by workspace and record kind but does not
apply `_revision_filter()`
([`harness/tht/adapters/vector/qdrant.py:174-194`](../../harness/tht/adapters/vector/qdrant.py#L174-L194)),
(`harness/tht/adapters/vector/qdrant.py:266-286`),
even though point IDs and reads are revision-scoped. Therefore a same-key/same-content record from
an older revision may be classified as unchanged and never written under the new revision. This is
an implementation defect/risk inferred directly from the two code paths, and it should be fixed
and regression-tested before the catalog relies on `index-schema` for multi-revision publication.
Deletion/GC must be stated per record family. Evidence cleanup is implemented: the corpus pipeline
retains the configured number of published generations, protects active/job-referenced
generations, and calls exact-generation deletion for evicted or compensated generations
(`harness/tht/evidence/corpus/pipeline.py`,
`harness/tht/adapters/vector/qdrant.py:350-390`). Schema cleanup is
not implemented: `delete_kinds()` exists as a workspace-scoped adapter primitive, but the schema
synchronizer never calls it, it is not revision-scoped, and there is no retention policy for old
schema revisions. Publication therefore still needs exact current-revision deletion semantics and
separate safe GC for unleased historical schema revisions.
## 5. Session pinning and why the SQL workflow can remain unchanged
Normal session creation acquires an immutable registry revision, passes its snapshot path,
workspace ID and commit to `tht session new`, and only releases the retention lease after the
manifest has been persisted
([`backend/src/routes/sessions.ts:341-367`](../../backend/src/routes/sessions.ts#L341-L367),
[`backend/src/routes/sessions.ts:423-440`](../../backend/src/routes/sessions.ts#L423-L440)). The
(`backend/src/routes/sessions.ts`). The
manifest stores `workspace_id` and `workspace_revision` next to database/schema identity
([`harness/tht/session/store.py:101-132`](../../harness/tht/session/store.py#L101-L132)). Resume and
(`harness/tht/session/store.py`). Resume and
saved-SQL paths reopen that exact retained snapshot rather than the current installation default
([`backend/src/routes/sessions.ts:184-198`](../../backend/src/routes/sessions.ts#L184-L198),
[`backend/src/routes/sql.ts:24-43`](../../backend/src/routes/sql.ts#L24-L43)).
(`backend/src/routes/sessions.ts`,
`backend/src/routes/sql.ts`).
The runtime renderer places the same workspace revision in `runtime_identity`, points the harness
at the descriptor-owned Qdrant collection, and supplies the internal embedding service
([`backend/src/workspaces/runtime-renderer.ts:248-277`](../../backend/src/workspaces/runtime-renderer.ts#L248-L277)).
(`backend/src/workspaces/runtime-renderer.ts`).
The vector adapter is constructed directly from that configuration, including the revision
([`harness/tht/adapters/factory.py:27-42`](../../harness/tht/adapters/factory.py#L27-L42)).
(`harness/tht/adapters/factory.py`).
At session bootstrap the backend invokes the existing `search pack` command
([`backend/src/routes/sessions.ts:465-470`](../../backend/src/routes/sessions.ts#L465-L470)). That
(`backend/src/routes/sessions.ts`). That
command queries schema records with the existing schema kinds, ranks tables and persists the
candidate list
([`harness/tht/cli/search_cmd.py:291-310`](../../harness/tht/cli/search_cmd.py#L291-L310)). The Pi
(`harness/tht/cli/search_cmd.py`). The Pi
extension reads the persisted retrieval pack through the CLI
([`harness/.pi/extensions/tht-gate.js:61-75`](../../harness/.pi/extensions/tht-gate.js#L61-L75)),
(`harness/.pi/extensions/tht-gate.js`),
and F4 starts from those candidates while loading full table/column context through existing schema
commands. Thus catalog publication can improve the inputs without changing phases, gate semantics,
or persisted session artifacts.
@@ -204,17 +244,18 @@ or persisted session artifacts.
## 6. DWH reuse across metadata-only revisions
DWH preparation uses immutable generations selected by an `ACTIVE` pointer
([`docs/contracts/tht-dwh.md:18-24`](../contracts/tht-dwh.md#L18-L24)). The effective-configuration
([DWH contract](../contracts/tht-dwh.md)). The effective-configuration
identity deliberately excludes `runtime_identity`, so a content-only Git revision does not force a
database re-introspection
([`docs/contracts/tht-dwh.md:74-90`](../contracts/tht-dwh.md#L74-L90)). This is the key enabling
property for metadata publication: a new revision can carry updated curated annotations/Core Schema
Selection, reuse the compatible physical DWH generation, and rebuild only the revision-scoped
schema projection.
([DWH contract](../contracts/tht-dwh.md)). This is the key enabling property for future metadata
publication: a new Git revision can carry updated curated annotations and a future Core Schema
Selection contract, reuse the compatible DWH-owned `physical.yaml`, and rebuild only the
revision-scoped schema projection. Today only the curated annotations part of that statement exists.
## 7. Constraints for the Metadata Catalog design
The following should be treated as requirements for the architecture map:
The following remain proposed requirements for the deferred catalog-to-core design gate; they are
not claims about current implementation:
1. **Separate full inventory from core projection.** PostgreSQL may hold the complete database
catalog, drafts and AI-generated text. Only an explicit Core Schema Selection and approved
@@ -225,10 +266,10 @@ The following should be treated as requirements for the architecture map:
3. **Use a new Git revision as the publication identity.** This preserves current snapshot,
retention, session resume and Qdrant filtering semantics. A separate mutable catalog revision
cannot be safely introduced without changing runtime reads.
4. **Reuse canonical formats.** Project physical facts/Core Schema Selection into a validated
`PhysicalSchema` view and semantic edits into `Annotations`; keep Mermaid and long-form database
documentation outside Qdrant unless a separate, explicit record kind and retrieval policy is
designed.
4. **Respect artifact ownership.** Keep introspected physical facts in the immutable DWH generation;
project the future Core Schema Selection through an explicit contract and approved semantic
edits into `Annotations`. Keep Mermaid and long-form database documentation outside Qdrant
unless a separate record kind and retrieval policy is designed.
5. **Reuse the operator boundary.** Trigger `workspace index-schema`, observe its schema-versioned
result and persist its artifact identities. Do not call Qdrant from browser CRUD handlers.
6. **Keep runtime read-only.** The session/Pi process continues to read the pinned snapshot,
@@ -236,14 +277,15 @@ The following should be treated as requirements for the architecture map:
management control plane.
7. **Surface publication gates.** UI status must distinguish Git activation, incompatible/missing
collection, resumable-session revision conflict, embedding failure and completed publication.
8. **Repair and test cross-revision synchronization first.** Scope `existing_hashes()` to the bound
revision and define deletion/garbage-collection semantics before depending on repeated catalog
publication.
8. **Repair and test cross-revision schema synchronization first.** Scope `existing_hashes()` to
the bound revision, delete records removed from the current revision's canonical schema, and
define safe schema-revision GC. Reuse rather than duplicate the already implemented Evidence
generation GC.
9. **Close the annotation display gap without changing the workflow.** Make `schema columns` read
the same merged descriptions used by M-Schema/vector rendering, so the current F4 widget sees
the approved catalog text.
## Decision summary for the Wayfinder map
## Proposed decision summary for the Wayfinder map
- Keep the Metadata Catalog as a separate management subsystem and source of editable metadata.
- Keep Qdrant derived and revision-scoped; it is not the catalog database or source of truth.
@@ -253,3 +295,7 @@ The following should be treated as requirements for the architecture map:
the current descriptor.
- Treat direct same-revision Qdrant writes, implicit CRUD publication, and bypassing the Git
snapshot boundary as rejected integration paths.
This summary remains design input. The authoritative current state is that catalog-to-core
integration and Sensitive Data Policy enforcement in schema-linking are deferred in
`PROJECT_STATE.md:160-171`.