14 KiB
PostgreSQL deployment and failure-isolation constraints for the Metadata Catalog
Date: 2026-08-23
Issue: #8 — Assess PostgreSQL deployment and failure-isolation constraints
Question
How can a Metadata Catalog PostgreSQL service be deployed, backed up, diagnosed, and made non-blocking for the core across local and server installations?
Recommendation
Use PostgreSQL as the operational store, but place it behind a separate Metadata Catalog API
process. The SQL workflow core must not receive the catalog DSN, catalog credentials, a client
pool, or a Compose dependency on either the catalog API or PostgreSQL.
The resulting dependency graph is:
flowchart LR
UI[Frontend]
Core[Core workflow API]
Semantic[Qdrant and embedding]
Catalog[Metadata Catalog API]
PG[(Catalog PostgreSQL)]
Publish[Publication worker]
UI -->|workflow routes| Core
Core --> Semantic
UI -->|catalog routes| Catalog
Catalog --> PG
PG --> Publish
Publish -->|immutable approved Publication| Semantic
The core continues to use the last successfully published, revision-qualified semantic snapshot.
Unpublished catalog edits are never a runtime dependency. This matches the domain boundary: the
Metadata Catalog does not select workflow SQL elements, while a Publication is an immutable value
made available to consumers (CONTEXT.md, lines 96–110). It also
preserves the existing rule that Qdrant is a derived index rather than a canonical source
(PROJECT_STATE.md, lines 405–410).
Evidence from the current architecture
- The mandatory workflow stack is currently
frontend,core,qdrant,embedding, and the model-init job.corewaits only for Qdrant and model initialization (compose.yaml, lines 50–60); the frontend waits only forcore(compose.yaml, lines 123–141). Adding a catalog dependency to either chain would enlarge the workflow failure domain. tht startruns Compose and then checks an explicit allowlist containing only those five mandatory services. Other Compose services are ignored by the readiness fold (service.go, lines 40–57,service.go, lines 166–190). Separate catalog services can therefore start with the installation without redefining core readiness./healthis deliberately a process-liveness endpoint and does not probe external services (app.ts, lines 271–284). Tests preserve a 200 response even when PostgreSQL session storage is configured but not contacted (health.test.ts, lines 9–32). Catalog readiness must not be folded into this endpoint.- ThothII already has a sound server-PostgreSQL precedent: runtime and migrator credentials are
distinct; the one-shot migrator is profile-gated and is explicitly not a dependency of
core(compose.session-server.yaml.example, lines 24–49). Its documented failure behavior is route-scoped 503 while process liveness remains healthy (PROJECT_STATE.md, lines 621–636). - Existing backup archives contain selected Compose volumes only. Local archives include seven
named volumes and server archives include only Qdrant and embedding volumes; no PostgreSQL
logical dump exists today (
create.go, lines 29–52). A catalog database therefore requires an explicit backup contract rather than an assumption that current installation backup already covers it.
Deployment contract
Shared rules
- Run
catalog-apias a separate, non-privileged service. The frontend reverse proxy may route/api/catalog/*to it, but frontend startup must not depend on catalog readiness. - Give
catalog-apia bounded pool and bounded operations. As an initial ceiling for this small administrative workload: pool size 5, overflow 0–2,connect_timeout=3,lock_timeout=2s, and a normal requeststatement_timeout=10s. Long AI calls must not hold database transactions. PostgreSQL documents that an omitted or zeroconnect_timeoutwaits indefinitely, so an explicit value is required for failure isolation (libpq connection parameters). - Use three roles:
catalog_runtime(CRUD and job claims only),catalog_migrator(DDL, used only by a one-shot job), andcatalog_backup(minimum privileges needed by the approved dump policy). No role is shared with the DWH orthoth_sessions, and none is a PostgreSQL superuser. - Store passwords and the CA as separate secret files. For non-local connections default to
sslmode=verify-full; this follows the current server-session configuration (compose.session-server.yaml.example, lines 5–20) and PostgreSQL's hostname-verifying TLS guidance (libpq SSL parameters). - Apply ordered, checksummed migrations under one transaction and an advisory lock. Refuse pending,
unknown, or checksum-drifted migrations in catalog readiness. The session repository already
implements this pattern
(
postgres_repository.py, lines 118–187); the catalog should have its own migration table, lock key, schema, and roles.
Local installation
Add an installation-owned catalog-postgres container on the private thothii network, with no
published host port and a dedicated catalog-postgres-data named volume. Pin the image by version
and digest, add a pg_isready healthcheck, and let only catalog-api use
depends_on: condition: service_healthy. core and frontend keep their existing dependency
graph. Docker confirms that service_healthy delays only the declaring dependent service, while
plain container start does not imply application readiness
(Compose startup order).
The local catalog services can be included in the normal installation Compose set because the
host CLI's mandatory-service allowlist excludes them. A separate catalog-migrate one-shot profile
retains explicit operator control; Docker profiles are intended for selectively activated and
one-off services (Compose profiles).
Server installation
Use an operator-provided PostgreSQL endpoint configured through a reviewed, untracked overlay. Prefer a separate database and dedicated roles. A separate PostgreSQL cluster gives the strongest resource and outage isolation; sharing the existing server cluster is acceptable only when the operator accepts the residual cluster-wide blast radius and enforces connection limits, timeouts, and separate ownership. Merely using another schema does not isolate connection exhaustion, maintenance outages, WAL pressure, or storage failure.
The server overlay adds secrets and catalog configuration only to catalog-api and the one-shot
migrator. It must not mount those secrets into core. This mirrors the current rule that core
never receives the session migrator credential
(PROJECT_STATE.md, lines 623–630).
Failure behavior
| Failure | Catalog behavior | Workflow behavior |
|---|---|---|
| PostgreSQL unavailable or connection pool exhausted | Catalog routes return sanitized 503 catalog_unavailable with Retry-After; UI remains read-only/unavailable |
Existing sessions and new SQL workflow requests continue against the last published snapshot |
| Catalog API process unavailable | /api/catalog/* fails; main frontend shell and workflow route remain usable |
No effect |
| Pending or drifted migration | Catalog readiness fails and writes are refused | No effect |
| AI generation fails | Job records a retryable/terminal failure; no database transaction remains open | No effect |
| Publication indexing fails | Publication remains non-active/failed and can be retried idempotently | Previous active Publication remains in Qdrant |
| Qdrant unavailable | Publication is queued/failed without losing canonical catalog state | Existing core readiness rules already govern this independent failure |
Do not silently fall back to JSON files, dual-write, or read unpublished PostgreSQL rows from the core. Those paths would create split-brain state. The only degraded-mode contract is use of the last successfully activated Publication.
Backup and restore
Backup
Extend tht backup create or add a catalog-specific command invoked by it. Create a logical
pg_dump --format=custom archive through a one-shot helper, checksum it, and place it in the
installation archive manifest. Do not tar a live PostgreSQL data volume: the existing raw-volume
backup mechanism was designed for the current named volumes, not database crash consistency.
pg_dump produces a consistent export while the database remains in use, and custom format is
compressed and supports selective/reordered restore
(PostgreSQL pg_dump). Because metadata
volume is small, favor correctness over parallelism. Run a backup before every migration and on an
operator-defined schedule; record the server version, migration head, dump checksum, creation time,
and catalog Publication identifiers in the manifest. Never write a dump into the workspace Git
repository.
For an external server database, the operator must supply the dedicated backup credential through the same secret-file boundary. A successful filesystem/Qdrant archive with a failed catalog dump is an incomplete installation backup and must not be reported as success.
Restore
Restore only during an explicit catalog maintenance window into an empty target database (or a
deliberately cleaned target), using pg_restore --exit-on-error --single-transaction. Validate the
archive checksum first, then require:
- expected migration history with no checksum drift;
- referential-integrity checks and bounded entity counts;
- Publication content hashes matching the restored records;
- an authenticated catalog smoke test.
The relevant restore controls and their behavior are documented by
PostgreSQL pg_restore. After restore,
rebuild catalog-derived Qdrant content from the restored active Publication rather than treating
the vector index as the source of truth. Keep the old database untouched until this validation
passes; rollback is endpoint reversal, not dual-write.
Diagnostics and observability
- Keep
core /healthunchanged. Givecatalog-apiseparate/health/live(process only) and/health/ready(bounded connection,SELECT 1, and migration compatibility) endpoints. - Add a named
metadata-catalogsection totht doctor: configuration completeness; secret-file presence and permissions; DNS/TCP/TLS/authentication;SELECT 1; migration applied/pending/drifted state; pool saturation; last successful backup age; and oldest pending Publication. Checks are read-only and all errors/DSNs are sanitized. The current doctor already separates Compose service status from container-local workflow checks (report.go, lines 275–310). - Use
pg_isreadyonly for transport readiness; its exit codes distinguish accepting, rejecting, no-response, and invalid-parameter states, but it does not prove schema or authorization correctness (PostgreSQLpg_isready). - Expose structured metrics/logs for connection acquisition failures, pool utilization, query timeout/lock timeout, migration head, AI-job backlog, Publication lag, and backup age. Never log connection strings, prompts containing schema metadata, or generated descriptions at info level.
tht statusmay show the catalog ashealthy,degraded, ordisabled, but catalog state must not change the existing core health result.tht doctormay return a failing catalog check for operator visibility without stopping or restarting the workflow stack.
Acceptance constraints for later design and implementation
- Killing
catalog-postgresandcatalog-apimust not changecore /health, stop Pi, interrupt an existing SQL session, or prevent a new session that uses the already active Publication. - No catalog database secret, client dependency, migration, or network call exists in the harness workflow path.
- Catalog failure is explicit on catalog routes; no JSON fallback or dual-write is introduced.
- A failed Publication leaves the previous active semantic content addressable and retryable.
- Local and server backup tests perform a real dump/restore round trip and prove that a backup is incomplete when the PostgreSQL dump fails.
- Migration credentials are absent from runtime containers; drift and pending migrations are
visible in readiness and
tht doctor.
Decision summary
The Metadata Catalog should be a separately deployed bounded context backed by PostgreSQL. Local installations own a private container and volume; server installations use a TLS-verified external database, preferably on a separate cluster when strict blast-radius isolation is required. Logical dump/restore, one-shot migrations, role separation, bounded connections, component-specific diagnostics, and last-good-Publication semantics make PostgreSQL operationally safe without making it a dependency of the SQL workflow.