# PostgreSQL deployment and failure-isolation constraints for the Metadata Catalog Original research: 2026-08-23 Last verified against the repository: 2026-08-31 Issue: [#8 — Assess PostgreSQL deployment and failure-isolation constraints](https://git.tylconsulting.it/mptyl/ThothII/issues/8) > **Status: partially superseded by ADR-0004.** This is historical research, not the current > architecture contract. The recommendation to run a separate `catalog-api` process was rejected > by [ADR-0004](../adr/0004-fastify-kysely-metadata-catalog.md). Operational constraints that do > not depend on that process boundary remain useful, but the status table below is authoritative > for what the 2026-08-31 code actually adopts, rejects, or leaves pending. ## Question How can a Metadata Catalog PostgreSQL service be deployed, backed up, diagnosed, and made non-blocking for the NL→SQL workflow across local and server installations? ## Current decision and implementation [ADR-0001](../adr/0001-postgres-metadata-catalog.md) selects PostgreSQL as the catalog authority. [ADR-0003](../adr/0003-installation-local-database-bindings.md) keeps database bindings installation-local. [ADR-0004](../adr/0004-fastify-kysely-metadata-catalog.md) then places the catalog in the existing Fastify backend as an isolated Kysely module rather than a microservice. The implemented dependency graph is therefore: ```mermaid flowchart LR UI[Frontend] Core[Core: Fastify workflow and catalog modules] Pi[Pi and tht workflow] PG[(catalog-db PostgreSQL)] Semantic[Qdrant and embedding] UI --> Core Core --> Pi Core --> PG Core --> Semantic ``` The catalog module is isolated behind a repository interface, but it is not process-isolated. `core` owns the catalog pool and Compose waits for `catalog-db` health before starting `core` (`compose.yaml`, `backend/src/app.ts`, `backend/src/catalog/repository.ts`). Database-management records still do not feed the NL→SQL handoff, so a running workflow does not read catalog rows; that cutover remains future work (`PROJECT_STATE.md`). ## Verification result | Area | Status on 2026-08-31 | Evidence and consequence | | --- | --- | --- | | PostgreSQL as canonical catalog store | **Adopted** | ADR-0001 is implemented by the PostgreSQL-backed Kysely repository and the internal `catalog-db` service. | | One Workspace Database per workspace and installation-local bindings | **Adopted** | ADR-0003 and the catalog migrations enforce the model; workspace identity remains in the workspace registry. | | Separate `catalog-api` process | **Superseded/rejected** | ADR-0004 explicitly chooses an isolated module inside the existing Fastify process. There is no catalog microservice or separate catalog liveness endpoint. | | No catalog pool, credential, or Compose dependency in `core` | **Superseded/rejected** | `core` owns the runtime pool, receives the runtime password secret, and declares `depends_on: catalog-db: service_healthy`. The old process-level isolation acceptance criterion is not current architecture. | | Private PostgreSQL service and durable volume | **Adopted** | `catalog-db` uses a version-and-digest-pinned PostgreSQL 17.6 image, the private `thothii` network, no published host port, a `pg_isready`-based healthcheck, and `catalog-data`. Compose contract tests assert this topology (`scripts/test-default-compose.sh`). | | Different local and server database topology | **Superseded/rejected** | Both current profiles inherit the same internal `catalog-db` and named volume. `deploy/compose.server.yaml` does not replace it with an operator-provided endpoint. | | Runtime and migrator role separation | **Adopted** | `core` receives only `thothii_catalog_runtime`; the profile-gated `catalog-migrate` job receives only `thothii_catalog_migrate`. Bootstrap grants runtime DML/sequence privileges without DDL (`docker/catalog-db-init.sql`). | | Dedicated backup role | **Pending** | There is no `catalog_backup` role or backup secret. | | Explicit one-shot migrations | **Adopted** | `backend/src/catalog/migrate.ts` registers ordered Kysely migrations and uses a pool of one. `scripts/run-stack.sh` starts PostgreSQL and runs `catalog-migrate` before local startup; migrations are not hidden in backend startup. | | Checksums, drift/pending refusal, and migration readiness | **Pending** | No repository-owned checksum policy, drift report, or readiness gate exists for catalog migrations. The backend can start without checking the Kysely migration head when launched outside the local helper. | | Bounded runtime connection pool | **Adopted** | `backend/src/catalog/repository.ts` sets `max: 5` and `connectionTimeoutMillis: 3000`, and closes Kysely with the Fastify lifecycle. | | `lock_timeout`, `statement_timeout`, and catalog TLS policy | **Pending** | The runtime pool does not set query or lock timeouts. Catalog connection configuration has no explicit CA/hostname-verification contract; the server profile still uses the private Compose network. | | Process liveness independent of catalog queries | **Adopted** | `GET /health` returns `{status: "ok"}` without probing PostgreSQL, and `backend/test/health.test.ts` preserves that behavior. `/catalog/status` performs the catalog-specific availability check. | | Stack startup independent of catalog availability | **Superseded/rejected** | Compose blocks `core` on healthy `catalog-db`. The host service health fold does not list `catalog-db`, but it cannot make `core` start while its Compose dependency is unhealthy. | | Uniform catalog-outage response (`503` plus `Retry-After`) | **Pending** | Routes map the domain `CatalogUnavailableError` to a sanitized `503`, and an omitted catalog configuration uses an unavailable repository. PostgreSQL driver failures are not uniformly translated to that domain error, and no `Retry-After` contract is implemented. | | Catalog-specific readiness and `tht doctor` checks | **Pending** | There is no `/health/ready` for the catalog and no doctor section for connection, migration head, pool saturation, backup age, or publication lag (`tools/tht/internal/doctor/report.go`). | | Logical catalog backup and restore | **Pending** | Current backup archives omit `catalog-data` and do not run `pg_dump`; server archives include only Qdrant and embedding volumes. A backup can therefore succeed without preserving the Metadata Catalog (`tools/tht/internal/backup/create.go`). | | Last-good Publication and catalog-to-Qdrant cutover | **Pending** | The current database-management slice does not change the NL→SQL runtime or publish catalog metadata to Qdrant. | ## Current operational contract ### Deployment and credentials - Local and server Compose profiles currently use the same installation-owned `catalog-db` container and `catalog-data` named volume. PostgreSQL is private to the Compose network. - The database bootstrap login is the migrator. The init script creates the separate runtime login from a Docker secret and grants only runtime DML and sequence access. - Runtime and migrator passwords are separate protected host files exposed as separate Docker secrets. Neither value belongs in tracked environment files (`deploy/env/local.env.example`, `deploy/env/server.env.example`). - The Fastify catalog repository uses a five-connection pool with a three-second connection timeout. Fastify closes the repository pool during shutdown. ### Migrations The compiled `catalog-migrate` entry point owns schema changes and uses the migrator credential. The current ordered series is under `backend/src/catalog/migrations/`. The local launcher runs it before the normal stack, while production rollout must invoke the profile-gated service explicitly. This is weaker than the original recommendation in two ways: there is no catalog readiness check for pending or unknown migrations, and the repository does not maintain content checksums for migration drift. Until those checks exist, “the migrator completed” is the available deployment gate; the application itself does not prove migration compatibility. ### Failure behavior actually provided | Failure | Current catalog behavior | Current workflow behavior | | --- | --- | --- | | Catalog configuration omitted when launching the backend directly | Fastify uses `UnavailableCatalogRepository`; `/catalog/status` reports unavailable and domain-mapped catalog operations return sanitized `503` | `/health`, sessions, and SSE remain available | | `catalog-db` unhealthy before Compose startup | `core` is not started because its dependency is not healthy | Workflow startup is blocked | | PostgreSQL becomes unavailable after startup | `/health` remains process-only and `/catalog/status` reports unavailable; route errors are sanitized, but a uniform `503`/`Retry-After` mapping is not guaranteed | Existing session code does not read catalog rows, but both surfaces still share one Fastify process | | Migrations are pending or incompatible | No dedicated readiness refusal exists; affected catalog operations fail | No catalog-to-workflow handoff exists yet, but local deployment correctness depends on running `catalog-migrate` first | | Catalog backup is requested through current `tht backup` | No PostgreSQL dump is added and `catalog-data` is omitted | The archive may succeed while being unable to restore catalog state | The shared process means resource exhaustion, fatal process errors, and startup hooks remain a common failure domain even though repository calls are separated. Conversely, placing the module inside Fastify does not require workflow code to consume catalog rows: preserving that data-flow boundary is still the useful part of the original isolation recommendation. ## Historical recommendations that remain valid backlog The following constraints survive ADR-0004 because they can be implemented inside the current Fastify deployment: 1. **Bound every database operation.** Keep the adopted pool and connection timeout; add explicit request query and lock timeouts. Long model calls and source sampling must not hold catalog transactions. 2. **Make failures component-specific.** Translate connection, timeout, and pool-exhaustion errors into one sanitized catalog-unavailable response, add bounded retry guidance, and keep `/health` process-only. 3. **Make migration compatibility observable.** Report applied, pending, unknown, and drifted migrations through a catalog readiness check and `tht doctor`; do not put migrator credentials in `core`. 4. **Define a server TLS topology before externalizing PostgreSQL.** If the server profile moves to an operator-provided endpoint, use a dedicated database and roles, protected CA material, and hostname verification. Sharing a cluster leaves connection, maintenance, WAL, and storage blast radius even when schemas are separate. PostgreSQL documents the relevant connection and TLS parameters in its [connection parameter reference](https://www.postgresql.org/docs/current/libpq-connect.html). 5. **Keep derived semantic data non-canonical.** When catalog publication is implemented, activate complete immutable revisions and retain the last successfully activated revision rather than exposing mutable catalog rows to the workflow. ## Backup and restore target The original logical-backup recommendation is still valid and is now a confirmed implementation gap. ### Backup Extend `tht backup create` with a catalog step that runs `pg_dump --format=custom` through a one-shot helper and adds the dump plus checksum to the installation manifest. Do not treat a tar of a live data volume as a PostgreSQL consistency contract. PostgreSQL documents that `pg_dump` creates a consistent export while the database remains in use and that custom format supports selective restore ([`pg_dump`](https://www.postgresql.org/docs/current/app-pgdump.html)). Use a dedicated least-privilege backup role, never write dumps into a workspace repository, and record at least the server version, migration head, checksum, creation time, and catalog revision identifiers. A filesystem/Qdrant archive with a failed or absent catalog dump must be reported as incomplete. ### Restore Restore only through an explicit maintenance operation into an empty or deliberately cleaned target. Validate the archive checksum, run `pg_restore --exit-on-error --single-transaction`, then verify migration compatibility, referential integrity, bounded entity counts, and an authenticated catalog smoke test. The controls are documented by PostgreSQL ([`pg_restore`](https://www.postgresql.org/docs/current/app-pgrestore.html)). Keep the old database until validation passes; rollback should switch the endpoint or retained volume, not dual-write. ## Diagnostics target - Keep the shared `GET /health` endpoint process-only. - Treat `/catalog/status` as the current minimal availability surface; add a bounded catalog readiness check covering connection and migration compatibility before using it as a rollout gate. - Add read-only `tht doctor` checks for secret-file presence and permissions, PostgreSQL reachability, authentication, migration state, pool saturation, and last successful catalog backup. Sanitize all DSNs and errors. - Use `pg_isready` only for server transport readiness. Its result does not prove schema or runtime authorization correctness ([`pg_isready`](https://www.postgresql.org/docs/current/app-pg-isready.html)). - Add structured metrics for connection acquisition failures, pool use, query/lock timeout, migration head, background-job backlog, and backup age. Do not log connection strings, source samples, prompts, or generated descriptions at info level. ## Recommended closure criteria 1. Decide whether ADR-0004's statement that catalog unavailability does not make sessions or SSE unavailable must also hold at Compose startup. If yes, remove or soften the hard `core` → `catalog-db` startup dependency without reintroducing a microservice. 2. Prove a configured PostgreSQL outage produces the same sanitized catalog response across every catalog route while `/health`, session creation, and SSE continue to work. 3. Add catalog migration compatibility to readiness and `tht doctor`, including pending, unknown, and drifted states. 4. Add a real `pg_dump`/`pg_restore` round-trip to local and server backup tests, and fail backup publication when the catalog dump is absent or fails. 5. Before enabling a remote server catalog, define TLS verification, credential files, connection limits, and the accepted cluster-level blast radius. 6. Before the NL→SQL cutover, define and test an immutable last-good Publication boundary; do not dual-read mutable PostgreSQL rows and legacy metadata as competing authorities. ## Decision summary PostgreSQL, installation-local bindings, private deployment, role separation, one-shot migrations, bounded pooling, and process-only liveness are implemented. The separate `catalog-api` process and the claim that `core` has no catalog dependency are not current design: ADR-0004 chose the existing Fastify process, and Compose currently blocks `core` startup on `catalog-db` health. Logical backup/restore, catalog migration readiness and drift detection, consistent outage mapping, catalog-specific doctor checks, server TLS/external topology, and last-good publication remain pending. Those open constraints should be treated as backlog, not as capabilities already provided by the repository.