Files
ThothII/docs/research/2026-08-23-postgresql-catalog-deployment-constraints.md
T

14 KiB
Raw Blame History

PostgreSQL deployment and failure-isolation constraints for the Metadata Catalog

Date: 2026-08-23
Issue: #8 — Assess PostgreSQL deployment and failure-isolation constraints

Question

How can a Metadata Catalog PostgreSQL service be deployed, backed up, diagnosed, and made non-blocking for the core across local and server installations?

Recommendation

Use PostgreSQL as the operational store, but place it behind a separate Metadata Catalog API process. The SQL workflow core must not receive the catalog DSN, catalog credentials, a client pool, or a Compose dependency on either the catalog API or PostgreSQL.

The resulting dependency graph is:

flowchart LR
    UI[Frontend]
    Core[Core workflow API]
    Semantic[Qdrant and embedding]
    Catalog[Metadata Catalog API]
    PG[(Catalog PostgreSQL)]
    Publish[Publication worker]

    UI -->|workflow routes| Core
    Core --> Semantic
    UI -->|catalog routes| Catalog
    Catalog --> PG
    PG --> Publish
    Publish -->|immutable approved Publication| Semantic

The core continues to use the last successfully published, revision-qualified semantic snapshot. Unpublished catalog edits are never a runtime dependency. This matches the domain boundary: the Metadata Catalog does not select workflow SQL elements, while a Publication is an immutable value made available to consumers (CONTEXT.md, lines 96–110). It also preserves the existing rule that Qdrant is a derived index rather than a canonical source (PROJECT_STATE.md, lines 405–410).

Evidence from the current architecture

  1. The mandatory workflow stack is currently frontend, core, qdrant, embedding, and the model-init job. core waits only for Qdrant and model initialization (compose.yaml, lines 50–60); the frontend waits only for core (compose.yaml, lines 123–141). Adding a catalog dependency to either chain would enlarge the workflow failure domain.
  2. tht start runs Compose and then checks an explicit allowlist containing only those five mandatory services. Other Compose services are ignored by the readiness fold (service.go, lines 40–57, service.go, lines 166–190). Separate catalog services can therefore start with the installation without redefining core readiness.
  3. /health is deliberately a process-liveness endpoint and does not probe external services (app.ts, lines 271–284). Tests preserve a 200 response even when PostgreSQL session storage is configured but not contacted (health.test.ts, lines 9–32). Catalog readiness must not be folded into this endpoint.
  4. ThothII already has a sound server-PostgreSQL precedent: runtime and migrator credentials are distinct; the one-shot migrator is profile-gated and is explicitly not a dependency of core (compose.session-server.yaml.example, lines 24–49). Its documented failure behavior is route-scoped 503 while process liveness remains healthy (PROJECT_STATE.md, lines 621–636).
  5. Existing backup archives contain selected Compose volumes only. Local archives include seven named volumes and server archives include only Qdrant and embedding volumes; no PostgreSQL logical dump exists today (create.go, lines 29–52). A catalog database therefore requires an explicit backup contract rather than an assumption that current installation backup already covers it.

Deployment contract

Shared rules

  • Run catalog-api as a separate, non-privileged service. The frontend reverse proxy may route /api/catalog/* to it, but frontend startup must not depend on catalog readiness.
  • Give catalog-api a bounded pool and bounded operations. As an initial ceiling for this small administrative workload: pool size 5, overflow 0–2, connect_timeout=3, lock_timeout=2s, and a normal request statement_timeout=10s. Long AI calls must not hold database transactions. PostgreSQL documents that an omitted or zero connect_timeout waits indefinitely, so an explicit value is required for failure isolation (libpq connection parameters).
  • Use three roles: catalog_runtime (CRUD and job claims only), catalog_migrator (DDL, used only by a one-shot job), and catalog_backup (minimum privileges needed by the approved dump policy). No role is shared with the DWH or thoth_sessions, and none is a PostgreSQL superuser.
  • Store passwords and the CA as separate secret files. For non-local connections default to sslmode=verify-full; this follows the current server-session configuration (compose.session-server.yaml.example, lines 5–20) and PostgreSQL's hostname-verifying TLS guidance (libpq SSL parameters).
  • Apply ordered, checksummed migrations under one transaction and an advisory lock. Refuse pending, unknown, or checksum-drifted migrations in catalog readiness. The session repository already implements this pattern (postgres_repository.py, lines 118–187); the catalog should have its own migration table, lock key, schema, and roles.

Local installation

Add an installation-owned catalog-postgres container on the private thothii network, with no published host port and a dedicated catalog-postgres-data named volume. Pin the image by version and digest, add a pg_isready healthcheck, and let only catalog-api use depends_on: condition: service_healthy. core and frontend keep their existing dependency graph. Docker confirms that service_healthy delays only the declaring dependent service, while plain container start does not imply application readiness (Compose startup order).

The local catalog services can be included in the normal installation Compose set because the host CLI's mandatory-service allowlist excludes them. A separate catalog-migrate one-shot profile retains explicit operator control; Docker profiles are intended for selectively activated and one-off services (Compose profiles).

Server installation

Use an operator-provided PostgreSQL endpoint configured through a reviewed, untracked overlay. Prefer a separate database and dedicated roles. A separate PostgreSQL cluster gives the strongest resource and outage isolation; sharing the existing server cluster is acceptable only when the operator accepts the residual cluster-wide blast radius and enforces connection limits, timeouts, and separate ownership. Merely using another schema does not isolate connection exhaustion, maintenance outages, WAL pressure, or storage failure.

The server overlay adds secrets and catalog configuration only to catalog-api and the one-shot migrator. It must not mount those secrets into core. This mirrors the current rule that core never receives the session migrator credential (PROJECT_STATE.md, lines 623–630).

Failure behavior

Failure Catalog behavior Workflow behavior
PostgreSQL unavailable or connection pool exhausted Catalog routes return sanitized 503 catalog_unavailable with Retry-After; UI remains read-only/unavailable Existing sessions and new SQL workflow requests continue against the last published snapshot
Catalog API process unavailable /api/catalog/* fails; main frontend shell and workflow route remain usable No effect
Pending or drifted migration Catalog readiness fails and writes are refused No effect
AI generation fails Job records a retryable/terminal failure; no database transaction remains open No effect
Publication indexing fails Publication remains non-active/failed and can be retried idempotently Previous active Publication remains in Qdrant
Qdrant unavailable Publication is queued/failed without losing canonical catalog state Existing core readiness rules already govern this independent failure

Do not silently fall back to JSON files, dual-write, or read unpublished PostgreSQL rows from the core. Those paths would create split-brain state. The only degraded-mode contract is use of the last successfully activated Publication.

Backup and restore

Backup

Extend tht backup create or add a catalog-specific command invoked by it. Create a logical pg_dump --format=custom archive through a one-shot helper, checksum it, and place it in the installation archive manifest. Do not tar a live PostgreSQL data volume: the existing raw-volume backup mechanism was designed for the current named volumes, not database crash consistency.

pg_dump produces a consistent export while the database remains in use, and custom format is compressed and supports selective/reordered restore (PostgreSQL pg_dump). Because metadata volume is small, favor correctness over parallelism. Run a backup before every migration and on an operator-defined schedule; record the server version, migration head, dump checksum, creation time, and catalog Publication identifiers in the manifest. Never write a dump into the workspace Git repository.

For an external server database, the operator must supply the dedicated backup credential through the same secret-file boundary. A successful filesystem/Qdrant archive with a failed catalog dump is an incomplete installation backup and must not be reported as success.

Restore

Restore only during an explicit catalog maintenance window into an empty target database (or a deliberately cleaned target), using pg_restore --exit-on-error --single-transaction. Validate the archive checksum first, then require:

  1. expected migration history with no checksum drift;
  2. referential-integrity checks and bounded entity counts;
  3. Publication content hashes matching the restored records;
  4. an authenticated catalog smoke test.

The relevant restore controls and their behavior are documented by PostgreSQL pg_restore. After restore, rebuild catalog-derived Qdrant content from the restored active Publication rather than treating the vector index as the source of truth. Keep the old database untouched until this validation passes; rollback is endpoint reversal, not dual-write.

Diagnostics and observability

  • Keep core /health unchanged. Give catalog-api separate /health/live (process only) and /health/ready (bounded connection, SELECT 1, and migration compatibility) endpoints.
  • Add a named metadata-catalog section to tht doctor: configuration completeness; secret-file presence and permissions; DNS/TCP/TLS/authentication; SELECT 1; migration applied/pending/drifted state; pool saturation; last successful backup age; and oldest pending Publication. Checks are read-only and all errors/DSNs are sanitized. The current doctor already separates Compose service status from container-local workflow checks (report.go, lines 275–310).
  • Use pg_isready only for transport readiness; its exit codes distinguish accepting, rejecting, no-response, and invalid-parameter states, but it does not prove schema or authorization correctness (PostgreSQL pg_isready).
  • Expose structured metrics/logs for connection acquisition failures, pool utilization, query timeout/lock timeout, migration head, AI-job backlog, Publication lag, and backup age. Never log connection strings, prompts containing schema metadata, or generated descriptions at info level.
  • tht status may show the catalog as healthy, degraded, or disabled, but catalog state must not change the existing core health result. tht doctor may return a failing catalog check for operator visibility without stopping or restarting the workflow stack.

Acceptance constraints for later design and implementation

  1. Killing catalog-postgres and catalog-api must not change core /health, stop Pi, interrupt an existing SQL session, or prevent a new session that uses the already active Publication.
  2. No catalog database secret, client dependency, migration, or network call exists in the harness workflow path.
  3. Catalog failure is explicit on catalog routes; no JSON fallback or dual-write is introduced.
  4. A failed Publication leaves the previous active semantic content addressable and retryable.
  5. Local and server backup tests perform a real dump/restore round trip and prove that a backup is incomplete when the PostgreSQL dump fails.
  6. Migration credentials are absent from runtime containers; drift and pending migrations are visible in readiness and tht doctor.

Decision summary

The Metadata Catalog should be a separately deployed bounded context backed by PostgreSQL. Local installations own a private container and volume; server installations use a TLS-verified external database, preferably on a separate cluster when strict blast-radius isolation is required. Logical dump/restore, one-shot migrations, role separation, bounded connections, component-specific diagnostics, and last-good-Publication semantics make PostgreSQL operationally safe without making it a dependency of the SQL workflow.