15 KiB
PostgreSQL deployment and failure-isolation constraints for the Metadata Catalog
Original research: 2026-08-23
Last verified against the repository: 2026-08-31
Issue: #8 — Assess PostgreSQL deployment and failure-isolation constraints
Status: partially superseded by ADR-0004. This is historical research, not the current architecture contract. The recommendation to run a separate
catalog-apiprocess was rejected by ADR-0004. Operational constraints that do not depend on that process boundary remain useful, but the status table below is authoritative for what the 2026-08-31 code actually adopts, rejects, or leaves pending.
Question
How can a Metadata Catalog PostgreSQL service be deployed, backed up, diagnosed, and made non-blocking for the NL→SQL workflow across local and server installations?
Current decision and implementation
ADR-0001 selects PostgreSQL as the catalog authority. ADR-0003 keeps database bindings installation-local. ADR-0004 then places the catalog in the existing Fastify backend as an isolated Kysely module rather than a microservice.
The implemented dependency graph is therefore:
flowchart LR
UI[Frontend]
Core[Core: Fastify workflow and catalog modules]
Pi[Pi and tht workflow]
PG[(catalog-db PostgreSQL)]
Semantic[Qdrant and embedding]
UI --> Core
Core --> Pi
Core --> PG
Core --> Semantic
The catalog module is isolated behind a repository interface, but it is not process-isolated.
core owns the catalog pool and Compose waits for catalog-db health before starting core
(compose.yaml, backend/src/app.ts, backend/src/catalog/repository.ts). Database-management records still do
not feed the NL→SQL handoff, so a running workflow does not read catalog rows; that cutover remains
future work (PROJECT_STATE.md).
Verification result
| Area | Status on 2026-08-31 | Evidence and consequence |
|---|---|---|
| PostgreSQL as canonical catalog store | Adopted | ADR-0001 is implemented by the PostgreSQL-backed Kysely repository and the internal catalog-db service. |
| One Workspace Database per workspace and installation-local bindings | Adopted | ADR-0003 and the catalog migrations enforce the model; workspace identity remains in the workspace registry. |
Separate catalog-api process |
Superseded/rejected | ADR-0004 explicitly chooses an isolated module inside the existing Fastify process. There is no catalog microservice or separate catalog liveness endpoint. |
No catalog pool, credential, or Compose dependency in core |
Superseded/rejected | core owns the runtime pool, receives the runtime password secret, and declares depends_on: catalog-db: service_healthy. The old process-level isolation acceptance criterion is not current architecture. |
| Private PostgreSQL service and durable volume | Adopted | catalog-db uses a version-and-digest-pinned PostgreSQL 17.6 image, the private thothii network, no published host port, a pg_isready-based healthcheck, and catalog-data. Compose contract tests assert this topology (scripts/test-default-compose.sh). |
| Different local and server database topology | Superseded/rejected | Both current profiles inherit the same internal catalog-db and named volume. deploy/compose.server.yaml does not replace it with an operator-provided endpoint. |
| Runtime and migrator role separation | Adopted | core receives only thothii_catalog_runtime; the profile-gated catalog-migrate job receives only thothii_catalog_migrate. Bootstrap grants runtime DML/sequence privileges without DDL (docker/catalog-db-init.sql). |
| Dedicated backup role | Pending | There is no catalog_backup role or backup secret. |
| Explicit one-shot migrations | Adopted | backend/src/catalog/migrate.ts registers ordered Kysely migrations and uses a pool of one. scripts/run-stack.sh starts PostgreSQL and runs catalog-migrate before local startup; migrations are not hidden in backend startup. |
| Checksums, drift/pending refusal, and migration readiness | Pending | No repository-owned checksum policy, drift report, or readiness gate exists for catalog migrations. The backend can start without checking the Kysely migration head when launched outside the local helper. |
| Bounded runtime connection pool | Adopted | backend/src/catalog/repository.ts sets max: 5 and connectionTimeoutMillis: 3000, and closes Kysely with the Fastify lifecycle. |
lock_timeout, statement_timeout, and catalog TLS policy |
Pending | The runtime pool does not set query or lock timeouts. Catalog connection configuration has no explicit CA/hostname-verification contract; the server profile still uses the private Compose network. |
| Process liveness independent of catalog queries | Adopted | GET /health returns {status: "ok"} without probing PostgreSQL, and backend/test/health.test.ts preserves that behavior. /catalog/status performs the catalog-specific availability check. |
| Stack startup independent of catalog availability | Superseded/rejected | Compose blocks core on healthy catalog-db. The host service health fold does not list catalog-db, but it cannot make core start while its Compose dependency is unhealthy. |
Uniform catalog-outage response (503 plus Retry-After) |
Pending | Routes map the domain CatalogUnavailableError to a sanitized 503, and an omitted catalog configuration uses an unavailable repository. PostgreSQL driver failures are not uniformly translated to that domain error, and no Retry-After contract is implemented. |
Catalog-specific readiness and tht doctor checks |
Pending | There is no /health/ready for the catalog and no doctor section for connection, migration head, pool saturation, backup age, or publication lag (tools/tht/internal/doctor/report.go). |
| Logical catalog backup and restore | Pending | Current backup archives omit catalog-data and do not run pg_dump; server archives include only Qdrant and embedding volumes. A backup can therefore succeed without preserving the Metadata Catalog (tools/tht/internal/backup/create.go). |
| Last-good Publication and catalog-to-Qdrant cutover | Pending | The current database-management slice does not change the NL→SQL runtime or publish catalog metadata to Qdrant. |
Current operational contract
Deployment and credentials
- Local and server Compose profiles currently use the same installation-owned
catalog-dbcontainer andcatalog-datanamed volume. PostgreSQL is private to the Compose network. - The database bootstrap login is the migrator. The init script creates the separate runtime login from a Docker secret and grants only runtime DML and sequence access.
- Runtime and migrator passwords are separate protected host files exposed as separate Docker
secrets. Neither value belongs in tracked environment files
(
deploy/env/local.env.example,deploy/env/server.env.example). - The Fastify catalog repository uses a five-connection pool with a three-second connection timeout. Fastify closes the repository pool during shutdown.
Migrations
The compiled catalog-migrate entry point owns schema changes and uses the migrator credential.
The current ordered series is under
backend/src/catalog/migrations/. The local launcher runs
it before the normal stack, while production rollout must invoke the profile-gated service
explicitly.
This is weaker than the original recommendation in two ways: there is no catalog readiness check for pending or unknown migrations, and the repository does not maintain content checksums for migration drift. Until those checks exist, “the migrator completed” is the available deployment gate; the application itself does not prove migration compatibility.
Failure behavior actually provided
| Failure | Current catalog behavior | Current workflow behavior |
|---|---|---|
| Catalog configuration omitted when launching the backend directly | Fastify uses UnavailableCatalogRepository; /catalog/status reports unavailable and domain-mapped catalog operations return sanitized 503 |
/health, sessions, and SSE remain available |
catalog-db unhealthy before Compose startup |
core is not started because its dependency is not healthy |
Workflow startup is blocked |
| PostgreSQL becomes unavailable after startup | /health remains process-only and /catalog/status reports unavailable; route errors are sanitized, but a uniform 503/Retry-After mapping is not guaranteed |
Existing session code does not read catalog rows, but both surfaces still share one Fastify process |
| Migrations are pending or incompatible | No dedicated readiness refusal exists; affected catalog operations fail | No catalog-to-workflow handoff exists yet, but local deployment correctness depends on running catalog-migrate first |
Catalog backup is requested through current tht backup |
No PostgreSQL dump is added and catalog-data is omitted |
The archive may succeed while being unable to restore catalog state |
The shared process means resource exhaustion, fatal process errors, and startup hooks remain a common failure domain even though repository calls are separated. Conversely, placing the module inside Fastify does not require workflow code to consume catalog rows: preserving that data-flow boundary is still the useful part of the original isolation recommendation.
Historical recommendations that remain valid backlog
The following constraints survive ADR-0004 because they can be implemented inside the current Fastify deployment:
- Bound every database operation. Keep the adopted pool and connection timeout; add explicit request query and lock timeouts. Long model calls and source sampling must not hold catalog transactions.
- Make failures component-specific. Translate connection, timeout, and pool-exhaustion errors
into one sanitized catalog-unavailable response, add bounded retry guidance, and keep
/healthprocess-only. - Make migration compatibility observable. Report applied, pending, unknown, and drifted
migrations through a catalog readiness check and
tht doctor; do not put migrator credentials incore. - Define a server TLS topology before externalizing PostgreSQL. If the server profile moves to an operator-provided endpoint, use a dedicated database and roles, protected CA material, and hostname verification. Sharing a cluster leaves connection, maintenance, WAL, and storage blast radius even when schemas are separate. PostgreSQL documents the relevant connection and TLS parameters in its connection parameter reference.
- Keep derived semantic data non-canonical. When catalog publication is implemented, activate complete immutable revisions and retain the last successfully activated revision rather than exposing mutable catalog rows to the workflow.
Backup and restore target
The original logical-backup recommendation is still valid and is now a confirmed implementation gap.
Backup
Extend tht backup create with a catalog step that runs pg_dump --format=custom through a
one-shot helper and adds the dump plus checksum to the installation manifest. Do not treat a tar of
a live data volume as a PostgreSQL consistency contract. PostgreSQL documents that pg_dump
creates a consistent export while the database remains in use and that custom format supports
selective restore (pg_dump).
Use a dedicated least-privilege backup role, never write dumps into a workspace repository, and record at least the server version, migration head, checksum, creation time, and catalog revision identifiers. A filesystem/Qdrant archive with a failed or absent catalog dump must be reported as incomplete.
Restore
Restore only through an explicit maintenance operation into an empty or deliberately cleaned
target. Validate the archive checksum, run pg_restore --exit-on-error --single-transaction, then
verify migration compatibility, referential integrity, bounded entity counts, and an authenticated
catalog smoke test. The controls are documented by PostgreSQL
(pg_restore). Keep the old database
until validation passes; rollback should switch the endpoint or retained volume, not dual-write.
Diagnostics target
- Keep the shared
GET /healthendpoint process-only. - Treat
/catalog/statusas the current minimal availability surface; add a bounded catalog readiness check covering connection and migration compatibility before using it as a rollout gate. - Add read-only
tht doctorchecks for secret-file presence and permissions, PostgreSQL reachability, authentication, migration state, pool saturation, and last successful catalog backup. Sanitize all DSNs and errors. - Use
pg_isreadyonly for server transport readiness. Its result does not prove schema or runtime authorization correctness (pg_isready). - Add structured metrics for connection acquisition failures, pool use, query/lock timeout, migration head, background-job backlog, and backup age. Do not log connection strings, source samples, prompts, or generated descriptions at info level.
Recommended closure criteria
- Decide whether ADR-0004's statement that catalog unavailability does not make sessions or SSE
unavailable must also hold at Compose startup. If yes, remove or soften the hard
core→catalog-dbstartup dependency without reintroducing a microservice. - Prove a configured PostgreSQL outage produces the same sanitized catalog response across every
catalog route while
/health, session creation, and SSE continue to work. - Add catalog migration compatibility to readiness and
tht doctor, including pending, unknown, and drifted states. - Add a real
pg_dump/pg_restoreround-trip to local and server backup tests, and fail backup publication when the catalog dump is absent or fails. - Before enabling a remote server catalog, define TLS verification, credential files, connection limits, and the accepted cluster-level blast radius.
- Before the NL→SQL cutover, define and test an immutable last-good Publication boundary; do not dual-read mutable PostgreSQL rows and legacy metadata as competing authorities.
Decision summary
PostgreSQL, installation-local bindings, private deployment, role separation, one-shot migrations,
bounded pooling, and process-only liveness are implemented. The separate catalog-api process and
the claim that core has no catalog dependency are not current design: ADR-0004 chose the existing
Fastify process, and Compose currently blocks core startup on catalog-db health. Logical
backup/restore, catalog migration readiness and drift detection, consistent outage mapping,
catalog-specific doctor checks, server TLS/external topology, and last-good publication remain
pending. Those open constraints should be treated as backlog, not as capabilities already provided
by the repository.