207 lines
15 KiB
Markdown
207 lines
15 KiB
Markdown
# PostgreSQL deployment and failure-isolation constraints for the Metadata Catalog
|
|
|
|
Original research: 2026-08-23
|
|
|
|
Last verified against the repository: 2026-08-31
|
|
|
|
Issue: [#8 — Assess PostgreSQL deployment and failure-isolation constraints](https://git.tylconsulting.it/mptyl/ThothII/issues/8)
|
|
|
|
> **Status: partially superseded by ADR-0004.** This is historical research, not the current
|
|
> architecture contract. The recommendation to run a separate `catalog-api` process was rejected
|
|
> by [ADR-0004](../adr/0004-fastify-kysely-metadata-catalog.md). Operational constraints that do
|
|
> not depend on that process boundary remain useful, but the status table below is authoritative
|
|
> for what the 2026-08-31 code actually adopts, rejects, or leaves pending.
|
|
|
|
## Question
|
|
|
|
How can a Metadata Catalog PostgreSQL service be deployed, backed up, diagnosed, and made
|
|
non-blocking for the NL→SQL workflow across local and server installations?
|
|
|
|
## Current decision and implementation
|
|
|
|
[ADR-0001](../adr/0001-postgres-metadata-catalog.md) selects PostgreSQL as the catalog authority.
|
|
[ADR-0003](../adr/0003-installation-local-database-bindings.md) keeps database bindings
|
|
installation-local. [ADR-0004](../adr/0004-fastify-kysely-metadata-catalog.md) then places the
|
|
catalog in the existing Fastify backend as an isolated Kysely module rather than a microservice.
|
|
|
|
The implemented dependency graph is therefore:
|
|
|
|
```mermaid
|
|
flowchart LR
|
|
UI[Frontend]
|
|
Core[Core: Fastify workflow and catalog modules]
|
|
Pi[Pi and tht workflow]
|
|
PG[(catalog-db PostgreSQL)]
|
|
Semantic[Qdrant and embedding]
|
|
|
|
UI --> Core
|
|
Core --> Pi
|
|
Core --> PG
|
|
Core --> Semantic
|
|
```
|
|
|
|
The catalog module is isolated behind a repository interface, but it is not process-isolated.
|
|
`core` owns the catalog pool and Compose waits for `catalog-db` health before starting `core`
|
|
(`compose.yaml`, `backend/src/app.ts`, `backend/src/catalog/repository.ts`). Database-management records still do
|
|
not feed the NL→SQL handoff, so a running workflow does not read catalog rows; that cutover remains
|
|
future work (`PROJECT_STATE.md`).
|
|
|
|
## Verification result
|
|
|
|
| Area | Status on 2026-08-31 | Evidence and consequence |
|
|
| --- | --- | --- |
|
|
| PostgreSQL as canonical catalog store | **Adopted** | ADR-0001 is implemented by the PostgreSQL-backed Kysely repository and the internal `catalog-db` service. |
|
|
| One Workspace Database per workspace and installation-local bindings | **Adopted** | ADR-0003 and the catalog migrations enforce the model; workspace identity remains in the workspace registry. |
|
|
| Separate `catalog-api` process | **Superseded/rejected** | ADR-0004 explicitly chooses an isolated module inside the existing Fastify process. There is no catalog microservice or separate catalog liveness endpoint. |
|
|
| No catalog pool, credential, or Compose dependency in `core` | **Superseded/rejected** | `core` owns the runtime pool, receives the runtime password secret, and declares `depends_on: catalog-db: service_healthy`. The old process-level isolation acceptance criterion is not current architecture. |
|
|
| Private PostgreSQL service and durable volume | **Adopted** | `catalog-db` uses a version-and-digest-pinned PostgreSQL 17.6 image, the private `thothii` network, no published host port, a `pg_isready`-based healthcheck, and `catalog-data`. Compose contract tests assert this topology (`scripts/test-default-compose.sh`). |
|
|
| Different local and server database topology | **Superseded/rejected** | Both current profiles inherit the same internal `catalog-db` and named volume. `deploy/compose.server.yaml` does not replace it with an operator-provided endpoint. |
|
|
| Runtime and migrator role separation | **Adopted** | `core` receives only `thothii_catalog_runtime`; the profile-gated `catalog-migrate` job receives only `thothii_catalog_migrate`. Bootstrap grants runtime DML/sequence privileges without DDL (`docker/catalog-db-init.sql`). |
|
|
| Dedicated backup role | **Pending** | There is no `catalog_backup` role or backup secret. |
|
|
| Explicit one-shot migrations | **Adopted** | `backend/src/catalog/migrate.ts` registers ordered Kysely migrations and uses a pool of one. `scripts/run-stack.sh` starts PostgreSQL and runs `catalog-migrate` before local startup; migrations are not hidden in backend startup. |
|
|
| Checksums, drift/pending refusal, and migration readiness | **Pending** | No repository-owned checksum policy, drift report, or readiness gate exists for catalog migrations. The backend can start without checking the Kysely migration head when launched outside the local helper. |
|
|
| Bounded runtime connection pool | **Adopted** | `backend/src/catalog/repository.ts` sets `max: 5` and `connectionTimeoutMillis: 3000`, and closes Kysely with the Fastify lifecycle. |
|
|
| `lock_timeout`, `statement_timeout`, and catalog TLS policy | **Pending** | The runtime pool does not set query or lock timeouts. Catalog connection configuration has no explicit CA/hostname-verification contract; the server profile still uses the private Compose network. |
|
|
| Process liveness independent of catalog queries | **Adopted** | `GET /health` returns `{status: "ok"}` without probing PostgreSQL, and `backend/test/health.test.ts` preserves that behavior. `/catalog/status` performs the catalog-specific availability check. |
|
|
| Stack startup independent of catalog availability | **Superseded/rejected** | Compose blocks `core` on healthy `catalog-db`. The host service health fold does not list `catalog-db`, but it cannot make `core` start while its Compose dependency is unhealthy. |
|
|
| Uniform catalog-outage response (`503` plus `Retry-After`) | **Pending** | Routes map the domain `CatalogUnavailableError` to a sanitized `503`, and an omitted catalog configuration uses an unavailable repository. PostgreSQL driver failures are not uniformly translated to that domain error, and no `Retry-After` contract is implemented. |
|
|
| Catalog-specific readiness and `tht doctor` checks | **Pending** | There is no `/health/ready` for the catalog and no doctor section for connection, migration head, pool saturation, backup age, or publication lag (`tools/tht/internal/doctor/report.go`). |
|
|
| Logical catalog backup and restore | **Pending** | Current backup archives omit `catalog-data` and do not run `pg_dump`; server archives include only Qdrant and embedding volumes. A backup can therefore succeed without preserving the Metadata Catalog (`tools/tht/internal/backup/create.go`). |
|
|
| Last-good Publication and catalog-to-Qdrant cutover | **Pending** | The current database-management slice does not change the NL→SQL runtime or publish catalog metadata to Qdrant. |
|
|
|
|
## Current operational contract
|
|
|
|
### Deployment and credentials
|
|
|
|
- Local and server Compose profiles currently use the same installation-owned `catalog-db`
|
|
container and `catalog-data` named volume. PostgreSQL is private to the Compose network.
|
|
- The database bootstrap login is the migrator. The init script creates the separate runtime login
|
|
from a Docker secret and grants only runtime DML and sequence access.
|
|
- Runtime and migrator passwords are separate protected host files exposed as separate Docker
|
|
secrets. Neither value belongs in tracked environment files
|
|
(`deploy/env/local.env.example`, `deploy/env/server.env.example`).
|
|
- The Fastify catalog repository uses a five-connection pool with a three-second connection
|
|
timeout. Fastify closes the repository pool during shutdown.
|
|
|
|
### Migrations
|
|
|
|
The compiled `catalog-migrate` entry point owns schema changes and uses the migrator credential.
|
|
The current ordered series is under
|
|
`backend/src/catalog/migrations/`. The local launcher runs
|
|
it before the normal stack, while production rollout must invoke the profile-gated service
|
|
explicitly.
|
|
|
|
This is weaker than the original recommendation in two ways: there is no catalog readiness check
|
|
for pending or unknown migrations, and the repository does not maintain content checksums for
|
|
migration drift. Until those checks exist, “the migrator completed” is the available deployment
|
|
gate; the application itself does not prove migration compatibility.
|
|
|
|
### Failure behavior actually provided
|
|
|
|
| Failure | Current catalog behavior | Current workflow behavior |
|
|
| --- | --- | --- |
|
|
| Catalog configuration omitted when launching the backend directly | Fastify uses `UnavailableCatalogRepository`; `/catalog/status` reports unavailable and domain-mapped catalog operations return sanitized `503` | `/health`, sessions, and SSE remain available |
|
|
| `catalog-db` unhealthy before Compose startup | `core` is not started because its dependency is not healthy | Workflow startup is blocked |
|
|
| PostgreSQL becomes unavailable after startup | `/health` remains process-only and `/catalog/status` reports unavailable; route errors are sanitized, but a uniform `503`/`Retry-After` mapping is not guaranteed | Existing session code does not read catalog rows, but both surfaces still share one Fastify process |
|
|
| Migrations are pending or incompatible | No dedicated readiness refusal exists; affected catalog operations fail | No catalog-to-workflow handoff exists yet, but local deployment correctness depends on running `catalog-migrate` first |
|
|
| Catalog backup is requested through current `tht backup` | No PostgreSQL dump is added and `catalog-data` is omitted | The archive may succeed while being unable to restore catalog state |
|
|
|
|
The shared process means resource exhaustion, fatal process errors, and startup hooks remain a
|
|
common failure domain even though repository calls are separated. Conversely, placing the module
|
|
inside Fastify does not require workflow code to consume catalog rows: preserving that data-flow
|
|
boundary is still the useful part of the original isolation recommendation.
|
|
|
|
## Historical recommendations that remain valid backlog
|
|
|
|
The following constraints survive ADR-0004 because they can be implemented inside the current
|
|
Fastify deployment:
|
|
|
|
1. **Bound every database operation.** Keep the adopted pool and connection timeout; add explicit
|
|
request query and lock timeouts. Long model calls and source sampling must not hold catalog
|
|
transactions.
|
|
2. **Make failures component-specific.** Translate connection, timeout, and pool-exhaustion errors
|
|
into one sanitized catalog-unavailable response, add bounded retry guidance, and keep `/health`
|
|
process-only.
|
|
3. **Make migration compatibility observable.** Report applied, pending, unknown, and drifted
|
|
migrations through a catalog readiness check and `tht doctor`; do not put migrator credentials
|
|
in `core`.
|
|
4. **Define a server TLS topology before externalizing PostgreSQL.** If the server profile moves to
|
|
an operator-provided endpoint, use a dedicated database and roles, protected CA material, and
|
|
hostname verification. Sharing a cluster leaves connection, maintenance, WAL, and storage blast
|
|
radius even when schemas are separate. PostgreSQL documents the relevant connection and TLS
|
|
parameters in its
|
|
[connection parameter reference](https://www.postgresql.org/docs/current/libpq-connect.html).
|
|
5. **Keep derived semantic data non-canonical.** When catalog publication is implemented, activate
|
|
complete immutable revisions and retain the last successfully activated revision rather than
|
|
exposing mutable catalog rows to the workflow.
|
|
|
|
## Backup and restore target
|
|
|
|
The original logical-backup recommendation is still valid and is now a confirmed implementation
|
|
gap.
|
|
|
|
### Backup
|
|
|
|
Extend `tht backup create` with a catalog step that runs `pg_dump --format=custom` through a
|
|
one-shot helper and adds the dump plus checksum to the installation manifest. Do not treat a tar of
|
|
a live data volume as a PostgreSQL consistency contract. PostgreSQL documents that `pg_dump`
|
|
creates a consistent export while the database remains in use and that custom format supports
|
|
selective restore ([`pg_dump`](https://www.postgresql.org/docs/current/app-pgdump.html)).
|
|
|
|
Use a dedicated least-privilege backup role, never write dumps into a workspace repository, and
|
|
record at least the server version, migration head, checksum, creation time, and catalog revision
|
|
identifiers. A filesystem/Qdrant archive with a failed or absent catalog dump must be reported as
|
|
incomplete.
|
|
|
|
### Restore
|
|
|
|
Restore only through an explicit maintenance operation into an empty or deliberately cleaned
|
|
target. Validate the archive checksum, run `pg_restore --exit-on-error --single-transaction`, then
|
|
verify migration compatibility, referential integrity, bounded entity counts, and an authenticated
|
|
catalog smoke test. The controls are documented by PostgreSQL
|
|
([`pg_restore`](https://www.postgresql.org/docs/current/app-pgrestore.html)). Keep the old database
|
|
until validation passes; rollback should switch the endpoint or retained volume, not dual-write.
|
|
|
|
## Diagnostics target
|
|
|
|
- Keep the shared `GET /health` endpoint process-only.
|
|
- Treat `/catalog/status` as the current minimal availability surface; add a bounded catalog
|
|
readiness check covering connection and migration compatibility before using it as a rollout
|
|
gate.
|
|
- Add read-only `tht doctor` checks for secret-file presence and permissions, PostgreSQL
|
|
reachability, authentication, migration state, pool saturation, and last successful catalog
|
|
backup. Sanitize all DSNs and errors.
|
|
- Use `pg_isready` only for server transport readiness. Its result does not prove schema or runtime
|
|
authorization correctness
|
|
([`pg_isready`](https://www.postgresql.org/docs/current/app-pg-isready.html)).
|
|
- Add structured metrics for connection acquisition failures, pool use, query/lock timeout,
|
|
migration head, background-job backlog, and backup age. Do not log connection strings, source
|
|
samples, prompts, or generated descriptions at info level.
|
|
|
|
## Recommended closure criteria
|
|
|
|
1. Decide whether ADR-0004's statement that catalog unavailability does not make sessions or SSE
|
|
unavailable must also hold at Compose startup. If yes, remove or soften the hard `core` →
|
|
`catalog-db` startup dependency without reintroducing a microservice.
|
|
2. Prove a configured PostgreSQL outage produces the same sanitized catalog response across every
|
|
catalog route while `/health`, session creation, and SSE continue to work.
|
|
3. Add catalog migration compatibility to readiness and `tht doctor`, including pending, unknown,
|
|
and drifted states.
|
|
4. Add a real `pg_dump`/`pg_restore` round-trip to local and server backup tests, and fail backup
|
|
publication when the catalog dump is absent or fails.
|
|
5. Before enabling a remote server catalog, define TLS verification, credential files, connection
|
|
limits, and the accepted cluster-level blast radius.
|
|
6. Before the NL→SQL cutover, define and test an immutable last-good Publication boundary; do not
|
|
dual-read mutable PostgreSQL rows and legacy metadata as competing authorities.
|
|
|
|
## Decision summary
|
|
|
|
PostgreSQL, installation-local bindings, private deployment, role separation, one-shot migrations,
|
|
bounded pooling, and process-only liveness are implemented. The separate `catalog-api` process and
|
|
the claim that `core` has no catalog dependency are not current design: ADR-0004 chose the existing
|
|
Fastify process, and Compose currently blocks `core` startup on `catalog-db` health. Logical
|
|
backup/restore, catalog migration readiness and drift detection, consistent outage mapping,
|
|
catalog-specific doctor checks, server TLS/external topology, and last-good publication remain
|
|
pending. Those open constraints should be treated as backlog, not as capabilities already provided
|
|
by the repository.
|