Files
ThothII/docs/research/2026-08-23-postgresql-catalog-deployment-constraints.md
T

15 KiB

PostgreSQL deployment and failure-isolation constraints for the Metadata Catalog

Original research: 2026-08-23

Last verified against the repository: 2026-08-31

Issue: #8 — Assess PostgreSQL deployment and failure-isolation constraints

Status: partially superseded by ADR-0004. This is historical research, not the current architecture contract. The recommendation to run a separate catalog-api process was rejected by ADR-0004. Operational constraints that do not depend on that process boundary remain useful, but the status table below is authoritative for what the 2026-08-31 code actually adopts, rejects, or leaves pending.

Question

How can a Metadata Catalog PostgreSQL service be deployed, backed up, diagnosed, and made non-blocking for the NL→SQL workflow across local and server installations?

Current decision and implementation

ADR-0001 selects PostgreSQL as the catalog authority. ADR-0003 keeps database bindings installation-local. ADR-0004 then places the catalog in the existing Fastify backend as an isolated Kysely module rather than a microservice.

The implemented dependency graph is therefore:

flowchart LR
    UI[Frontend]
    Core[Core: Fastify workflow and catalog modules]
    Pi[Pi and tht workflow]
    PG[(catalog-db PostgreSQL)]
    Semantic[Qdrant and embedding]

    UI --> Core
    Core --> Pi
    Core --> PG
    Core --> Semantic

The catalog module is isolated behind a repository interface, but it is not process-isolated. core owns the catalog pool and Compose waits for catalog-db health before starting core (compose.yaml, backend/src/app.ts, backend/src/catalog/repository.ts). Database-management records still do not feed the NL→SQL handoff, so a running workflow does not read catalog rows; that cutover remains future work (PROJECT_STATE.md).

Verification result

Area Status on 2026-08-31 Evidence and consequence
PostgreSQL as canonical catalog store Adopted ADR-0001 is implemented by the PostgreSQL-backed Kysely repository and the internal catalog-db service.
One Workspace Database per workspace and installation-local bindings Adopted ADR-0003 and the catalog migrations enforce the model; workspace identity remains in the workspace registry.
Separate catalog-api process Superseded/rejected ADR-0004 explicitly chooses an isolated module inside the existing Fastify process. There is no catalog microservice or separate catalog liveness endpoint.
No catalog pool, credential, or Compose dependency in core Superseded/rejected core owns the runtime pool, receives the runtime password secret, and declares depends_on: catalog-db: service_healthy. The old process-level isolation acceptance criterion is not current architecture.
Private PostgreSQL service and durable volume Adopted catalog-db uses a version-and-digest-pinned PostgreSQL 17.6 image, the private thothii network, no published host port, a pg_isready-based healthcheck, and catalog-data. Compose contract tests assert this topology (scripts/test-default-compose.sh).
Different local and server database topology Superseded/rejected Both current profiles inherit the same internal catalog-db and named volume. deploy/compose.server.yaml does not replace it with an operator-provided endpoint.
Runtime and migrator role separation Adopted core receives only thothii_catalog_runtime; the profile-gated catalog-migrate job receives only thothii_catalog_migrate. Bootstrap grants runtime DML/sequence privileges without DDL (docker/catalog-db-init.sql).
Dedicated backup role Pending There is no catalog_backup role or backup secret.
Explicit one-shot migrations Adopted backend/src/catalog/migrate.ts registers ordered Kysely migrations and uses a pool of one. scripts/run-stack.sh starts PostgreSQL and runs catalog-migrate before local startup; migrations are not hidden in backend startup.
Checksums, drift/pending refusal, and migration readiness Pending No repository-owned checksum policy, drift report, or readiness gate exists for catalog migrations. The backend can start without checking the Kysely migration head when launched outside the local helper.
Bounded runtime connection pool Adopted backend/src/catalog/repository.ts sets max: 5 and connectionTimeoutMillis: 3000, and closes Kysely with the Fastify lifecycle.
lock_timeout, statement_timeout, and catalog TLS policy Pending The runtime pool does not set query or lock timeouts. Catalog connection configuration has no explicit CA/hostname-verification contract; the server profile still uses the private Compose network.
Process liveness independent of catalog queries Adopted GET /health returns {status: "ok"} without probing PostgreSQL, and backend/test/health.test.ts preserves that behavior. /catalog/status performs the catalog-specific availability check.
Stack startup independent of catalog availability Superseded/rejected Compose blocks core on healthy catalog-db. The host service health fold does not list catalog-db, but it cannot make core start while its Compose dependency is unhealthy.
Uniform catalog-outage response (503 plus Retry-After) Pending Routes map the domain CatalogUnavailableError to a sanitized 503, and an omitted catalog configuration uses an unavailable repository. PostgreSQL driver failures are not uniformly translated to that domain error, and no Retry-After contract is implemented.
Catalog-specific readiness and tht doctor checks Pending There is no /health/ready for the catalog and no doctor section for connection, migration head, pool saturation, backup age, or publication lag (tools/tht/internal/doctor/report.go).
Logical catalog backup and restore Pending Current backup archives omit catalog-data and do not run pg_dump; server archives include only Qdrant and embedding volumes. A backup can therefore succeed without preserving the Metadata Catalog (tools/tht/internal/backup/create.go).
Last-good Publication and catalog-to-Qdrant cutover Pending The current database-management slice does not change the NL→SQL runtime or publish catalog metadata to Qdrant.

Current operational contract

Deployment and credentials

  • Local and server Compose profiles currently use the same installation-owned catalog-db container and catalog-data named volume. PostgreSQL is private to the Compose network.
  • The database bootstrap login is the migrator. The init script creates the separate runtime login from a Docker secret and grants only runtime DML and sequence access.
  • Runtime and migrator passwords are separate protected host files exposed as separate Docker secrets. Neither value belongs in tracked environment files (deploy/env/local.env.example, deploy/env/server.env.example).
  • The Fastify catalog repository uses a five-connection pool with a three-second connection timeout. Fastify closes the repository pool during shutdown.

Migrations

The compiled catalog-migrate entry point owns schema changes and uses the migrator credential. The current ordered series is under backend/src/catalog/migrations/. The local launcher runs it before the normal stack, while production rollout must invoke the profile-gated service explicitly.

This is weaker than the original recommendation in two ways: there is no catalog readiness check for pending or unknown migrations, and the repository does not maintain content checksums for migration drift. Until those checks exist, “the migrator completed” is the available deployment gate; the application itself does not prove migration compatibility.

Failure behavior actually provided

Failure Current catalog behavior Current workflow behavior
Catalog configuration omitted when launching the backend directly Fastify uses UnavailableCatalogRepository; /catalog/status reports unavailable and domain-mapped catalog operations return sanitized 503 /health, sessions, and SSE remain available
catalog-db unhealthy before Compose startup core is not started because its dependency is not healthy Workflow startup is blocked
PostgreSQL becomes unavailable after startup /health remains process-only and /catalog/status reports unavailable; route errors are sanitized, but a uniform 503/Retry-After mapping is not guaranteed Existing session code does not read catalog rows, but both surfaces still share one Fastify process
Migrations are pending or incompatible No dedicated readiness refusal exists; affected catalog operations fail No catalog-to-workflow handoff exists yet, but local deployment correctness depends on running catalog-migrate first
Catalog backup is requested through current tht backup No PostgreSQL dump is added and catalog-data is omitted The archive may succeed while being unable to restore catalog state

The shared process means resource exhaustion, fatal process errors, and startup hooks remain a common failure domain even though repository calls are separated. Conversely, placing the module inside Fastify does not require workflow code to consume catalog rows: preserving that data-flow boundary is still the useful part of the original isolation recommendation.

Historical recommendations that remain valid backlog

The following constraints survive ADR-0004 because they can be implemented inside the current Fastify deployment:

  1. Bound every database operation. Keep the adopted pool and connection timeout; add explicit request query and lock timeouts. Long model calls and source sampling must not hold catalog transactions.
  2. Make failures component-specific. Translate connection, timeout, and pool-exhaustion errors into one sanitized catalog-unavailable response, add bounded retry guidance, and keep /health process-only.
  3. Make migration compatibility observable. Report applied, pending, unknown, and drifted migrations through a catalog readiness check and tht doctor; do not put migrator credentials in core.
  4. Define a server TLS topology before externalizing PostgreSQL. If the server profile moves to an operator-provided endpoint, use a dedicated database and roles, protected CA material, and hostname verification. Sharing a cluster leaves connection, maintenance, WAL, and storage blast radius even when schemas are separate. PostgreSQL documents the relevant connection and TLS parameters in its connection parameter reference.
  5. Keep derived semantic data non-canonical. When catalog publication is implemented, activate complete immutable revisions and retain the last successfully activated revision rather than exposing mutable catalog rows to the workflow.

Backup and restore target

The original logical-backup recommendation is still valid and is now a confirmed implementation gap.

Backup

Extend tht backup create with a catalog step that runs pg_dump --format=custom through a one-shot helper and adds the dump plus checksum to the installation manifest. Do not treat a tar of a live data volume as a PostgreSQL consistency contract. PostgreSQL documents that pg_dump creates a consistent export while the database remains in use and that custom format supports selective restore (pg_dump).

Use a dedicated least-privilege backup role, never write dumps into a workspace repository, and record at least the server version, migration head, checksum, creation time, and catalog revision identifiers. A filesystem/Qdrant archive with a failed or absent catalog dump must be reported as incomplete.

Restore

Restore only through an explicit maintenance operation into an empty or deliberately cleaned target. Validate the archive checksum, run pg_restore --exit-on-error --single-transaction, then verify migration compatibility, referential integrity, bounded entity counts, and an authenticated catalog smoke test. The controls are documented by PostgreSQL (pg_restore). Keep the old database until validation passes; rollback should switch the endpoint or retained volume, not dual-write.

Diagnostics target

  • Keep the shared GET /health endpoint process-only.
  • Treat /catalog/status as the current minimal availability surface; add a bounded catalog readiness check covering connection and migration compatibility before using it as a rollout gate.
  • Add read-only tht doctor checks for secret-file presence and permissions, PostgreSQL reachability, authentication, migration state, pool saturation, and last successful catalog backup. Sanitize all DSNs and errors.
  • Use pg_isready only for server transport readiness. Its result does not prove schema or runtime authorization correctness (pg_isready).
  • Add structured metrics for connection acquisition failures, pool use, query/lock timeout, migration head, background-job backlog, and backup age. Do not log connection strings, source samples, prompts, or generated descriptions at info level.
  1. Decide whether ADR-0004's statement that catalog unavailability does not make sessions or SSE unavailable must also hold at Compose startup. If yes, remove or soften the hard core → catalog-db startup dependency without reintroducing a microservice.
  2. Prove a configured PostgreSQL outage produces the same sanitized catalog response across every catalog route while /health, session creation, and SSE continue to work.
  3. Add catalog migration compatibility to readiness and tht doctor, including pending, unknown, and drifted states.
  4. Add a real pg_dump/pg_restore round-trip to local and server backup tests, and fail backup publication when the catalog dump is absent or fails.
  5. Before enabling a remote server catalog, define TLS verification, credential files, connection limits, and the accepted cluster-level blast radius.
  6. Before the NL→SQL cutover, define and test an immutable last-good Publication boundary; do not dual-read mutable PostgreSQL rows and legacy metadata as competing authorities.

Decision summary

PostgreSQL, installation-local bindings, private deployment, role separation, one-shot migrations, bounded pooling, and process-only liveness are implemented. The separate catalog-api process and the claim that core has no catalog dependency are not current design: ADR-0004 chose the existing Fastify process, and Compose currently blocks core startup on catalog-db health. Logical backup/restore, catalog migration readiness and drift detection, consistent outage mapping, catalog-specific doctor checks, server TLS/external topology, and last-good publication remain pending. Those open constraints should be treated as backlog, not as capabilities already provided by the repository.