docs: revalidate metadata catalog research

This commit is contained in:
Codex
2026-08-31 14:54:52 +02:00
parent 16b7477861
commit 866aee4249
3 changed files with 345 additions and 253 deletions
@@ -1,219 +1,206 @@
# PostgreSQL deployment and failure-isolation constraints for the Metadata Catalog
Date: 2026-08-23
Issue: [#8 — Assess PostgreSQL deployment and failure-isolation constraints](https://github.com/mptyl/ThothII/issues/8)
Original research: 2026-08-23
Last verified against the repository: 2026-08-31
Issue: [#8 — Assess PostgreSQL deployment and failure-isolation constraints](https://git.tylconsulting.it/mptyl/ThothII/issues/8)
> **Status: partially superseded by ADR-0004.** This is historical research, not the current
> architecture contract. The recommendation to run a separate `catalog-api` process was rejected
> by [ADR-0004](../adr/0004-fastify-kysely-metadata-catalog.md). Operational constraints that do
> not depend on that process boundary remain useful, but the status table below is authoritative
> for what the 2026-08-31 code actually adopts, rejects, or leaves pending.
## Question
How can a Metadata Catalog PostgreSQL service be deployed, backed up, diagnosed, and made
non-blocking for the core across local and server installations?
non-blocking for the NL→SQL workflow across local and server installations?
## Recommendation
## Current decision and implementation
Use PostgreSQL as the operational store, but place it behind a **separate Metadata Catalog API
process**. The SQL workflow `core` must not receive the catalog DSN, catalog credentials, a client
pool, or a Compose dependency on either the catalog API or PostgreSQL.
[ADR-0001](../adr/0001-postgres-metadata-catalog.md) selects PostgreSQL as the catalog authority.
[ADR-0003](../adr/0003-installation-local-database-bindings.md) keeps database bindings
installation-local. [ADR-0004](../adr/0004-fastify-kysely-metadata-catalog.md) then places the
catalog in the existing Fastify backend as an isolated Kysely module rather than a microservice.
The resulting dependency graph is:
The implemented dependency graph is therefore:
```mermaid
flowchart LR
UI[Frontend]
Core[Core workflow API]
Core[Core: Fastify workflow and catalog modules]
Pi[Pi and tht workflow]
PG[(catalog-db PostgreSQL)]
Semantic[Qdrant and embedding]
Catalog[Metadata Catalog API]
PG[(Catalog PostgreSQL)]
Publish[Publication worker]
UI -->|workflow routes| Core
UI --> Core
Core --> Pi
Core --> PG
Core --> Semantic
UI -->|catalog routes| Catalog
Catalog --> PG
PG --> Publish
Publish -->|immutable approved Publication| Semantic
```
The core continues to use the last successfully published, revision-qualified semantic snapshot.
Unpublished catalog edits are never a runtime dependency. This matches the domain boundary: the
Metadata Catalog does not select workflow SQL elements, while a Publication is an immutable value
made available to consumers ([`CONTEXT.md`, lines 96–110](../../CONTEXT.md#L96-L110)). It also
preserves the existing rule that Qdrant is a derived index rather than a canonical source
([`PROJECT_STATE.md`, lines 405–410](../../PROJECT_STATE.md#L405-L410)).
The catalog module is isolated behind a repository interface, but it is not process-isolated.
`core` owns the catalog pool and Compose waits for `catalog-db` health before starting `core`
(`compose.yaml`, `backend/src/app.ts`, `backend/src/catalog/repository.ts`). Database-management records still do
not feed the NL→SQL handoff, so a running workflow does not read catalog rows; that cutover remains
future work (`PROJECT_STATE.md`).
## Evidence from the current architecture
## Verification result
1. The mandatory workflow stack is currently `frontend`, `core`, `qdrant`, `embedding`, and the
model-init job. `core` waits only for Qdrant and model initialization
([`compose.yaml`, lines 50–60](../../compose.yaml#L50-L60)); the frontend waits only for `core`
([`compose.yaml`, lines 123–141](../../compose.yaml#L123-L141)). Adding a catalog dependency to
either chain would enlarge the workflow failure domain.
2. `tht start` runs Compose and then checks an explicit allowlist containing only those five
mandatory services. Other Compose services are ignored by the readiness fold
([`service.go`, lines 40–57](../../tools/tht/internal/service/service.go#L40-L57),
[`service.go`, lines 166–190](../../tools/tht/internal/service/service.go#L166-L190)). Separate
catalog services can therefore start with the installation without redefining core readiness.
3. `/health` is deliberately a process-liveness endpoint and does not probe external services
([`app.ts`, lines 271–284](../../backend/src/app.ts#L271-L284)). Tests preserve a 200 response
even when PostgreSQL session storage is configured but not contacted
([`health.test.ts`, lines 9–32](../../backend/test/health.test.ts#L9-L32)). Catalog readiness must
not be folded into this endpoint.
4. ThothII already has a sound server-PostgreSQL precedent: runtime and migrator credentials are
distinct; the one-shot migrator is profile-gated and is explicitly not a dependency of `core`
([`compose.session-server.yaml.example`, lines 24–49](../../deploy/compose.session-server.yaml.example#L24-L49)).
Its documented failure behavior is route-scoped 503 while process liveness remains healthy
([`PROJECT_STATE.md`, lines 621–636](../../PROJECT_STATE.md#L621-L636)).
5. Existing backup archives contain selected Compose volumes only. Local archives include seven
named volumes and server archives include only Qdrant and embedding volumes; no PostgreSQL
logical dump exists today ([`create.go`, lines 29–52](../../tools/tht/internal/backup/create.go#L29-L52)).
A catalog database therefore requires an explicit backup contract rather than an assumption
that current installation backup already covers it.
## Deployment contract
### Shared rules
- Run `catalog-api` as a separate, non-privileged service. The frontend reverse proxy may route
`/api/catalog/*` to it, but frontend startup must not depend on catalog readiness.
- Give `catalog-api` a bounded pool and bounded operations. As an initial ceiling for this small
administrative workload: pool size 5, overflow 0–2, `connect_timeout=3`, `lock_timeout=2s`, and
a normal request `statement_timeout=10s`. Long AI calls must not hold database transactions.
PostgreSQL documents that an omitted or zero `connect_timeout` waits indefinitely, so an
explicit value is required for failure isolation
([libpq connection parameters](https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-CONNECT-TIMEOUT)).
- Use three roles: `catalog_runtime` (CRUD and job claims only), `catalog_migrator` (DDL, used only
by a one-shot job), and `catalog_backup` (minimum privileges needed by the approved dump policy).
No role is shared with the DWH or `thoth_sessions`, and none is a PostgreSQL superuser.
- Store passwords and the CA as separate secret files. For non-local connections default to
`sslmode=verify-full`; this follows the current server-session configuration
([`compose.session-server.yaml.example`, lines 5–20](../../deploy/compose.session-server.yaml.example#L5-L20))
and PostgreSQL's hostname-verifying TLS guidance
([libpq SSL parameters](https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-PARAMKEYWORDS)).
- Apply ordered, checksummed migrations under one transaction and an advisory lock. Refuse pending,
unknown, or checksum-drifted migrations in catalog readiness. The session repository already
implements this pattern
([`postgres_repository.py`, lines 118–187](../../harness/tht/session/postgres_repository.py#L118-L187));
the catalog should have its own migration table, lock key, schema, and roles.
### Local installation
Add an installation-owned `catalog-postgres` container on the private `thothii` network, with no
published host port and a dedicated `catalog-postgres-data` named volume. Pin the image by version
and digest, add a `pg_isready` healthcheck, and let only `catalog-api` use
`depends_on: condition: service_healthy`. `core` and `frontend` keep their existing dependency
graph. Docker confirms that `service_healthy` delays only the declaring dependent service, while
plain container start does not imply application readiness
([Compose startup order](https://docs.docker.com/compose/how-tos/startup-order/)).
The local catalog services can be included in the normal installation Compose set because the
host CLI's mandatory-service allowlist excludes them. A separate `catalog-migrate` one-shot profile
retains explicit operator control; Docker profiles are intended for selectively activated and
one-off services ([Compose profiles](https://docs.docker.com/compose/how-tos/profiles/)).
### Server installation
Use an operator-provided PostgreSQL endpoint configured through a reviewed, untracked overlay.
Prefer a separate database and dedicated roles. A separate PostgreSQL cluster gives the strongest
resource and outage isolation; sharing the existing server cluster is acceptable only when the
operator accepts the residual cluster-wide blast radius and enforces connection limits, timeouts,
and separate ownership. Merely using another schema does not isolate connection exhaustion,
maintenance outages, WAL pressure, or storage failure.
The server overlay adds secrets and catalog configuration only to `catalog-api` and the one-shot
migrator. It must not mount those secrets into `core`. This mirrors the current rule that `core`
never receives the session migrator credential
([`PROJECT_STATE.md`, lines 623–630](../../PROJECT_STATE.md#L623-L630)).
## Failure behavior
| Failure | Catalog behavior | Workflow behavior |
| Area | Status on 2026-08-31 | Evidence and consequence |
| --- | --- | --- |
| PostgreSQL unavailable or connection pool exhausted | Catalog routes return sanitized `503 catalog_unavailable` with `Retry-After`; UI remains read-only/unavailable | Existing sessions and new SQL workflow requests continue against the last published snapshot |
| Catalog API process unavailable | `/api/catalog/*` fails; main frontend shell and workflow route remain usable | No effect |
| Pending or drifted migration | Catalog readiness fails and writes are refused | No effect |
| AI generation fails | Job records a retryable/terminal failure; no database transaction remains open | No effect |
| Publication indexing fails | Publication remains non-active/failed and can be retried idempotently | Previous active Publication remains in Qdrant |
| Qdrant unavailable | Publication is queued/failed without losing canonical catalog state | Existing core readiness rules already govern this independent failure |
| PostgreSQL as canonical catalog store | **Adopted** | ADR-0001 is implemented by the PostgreSQL-backed Kysely repository and the internal `catalog-db` service. |
| One Workspace Database per workspace and installation-local bindings | **Adopted** | ADR-0003 and the catalog migrations enforce the model; workspace identity remains in the workspace registry. |
| Separate `catalog-api` process | **Superseded/rejected** | ADR-0004 explicitly chooses an isolated module inside the existing Fastify process. There is no catalog microservice or separate catalog liveness endpoint. |
| No catalog pool, credential, or Compose dependency in `core` | **Superseded/rejected** | `core` owns the runtime pool, receives the runtime password secret, and declares `depends_on: catalog-db: service_healthy`. The old process-level isolation acceptance criterion is not current architecture. |
| Private PostgreSQL service and durable volume | **Adopted** | `catalog-db` uses a version-and-digest-pinned PostgreSQL 17.6 image, the private `thothii` network, no published host port, a `pg_isready`-based healthcheck, and `catalog-data`. Compose contract tests assert this topology (`scripts/test-default-compose.sh`). |
| Different local and server database topology | **Superseded/rejected** | Both current profiles inherit the same internal `catalog-db` and named volume. `deploy/compose.server.yaml` does not replace it with an operator-provided endpoint. |
| Runtime and migrator role separation | **Adopted** | `core` receives only `thothii_catalog_runtime`; the profile-gated `catalog-migrate` job receives only `thothii_catalog_migrate`. Bootstrap grants runtime DML/sequence privileges without DDL (`docker/catalog-db-init.sql`). |
| Dedicated backup role | **Pending** | There is no `catalog_backup` role or backup secret. |
| Explicit one-shot migrations | **Adopted** | `backend/src/catalog/migrate.ts` registers ordered Kysely migrations and uses a pool of one. `scripts/run-stack.sh` starts PostgreSQL and runs `catalog-migrate` before local startup; migrations are not hidden in backend startup. |
| Checksums, drift/pending refusal, and migration readiness | **Pending** | No repository-owned checksum policy, drift report, or readiness gate exists for catalog migrations. The backend can start without checking the Kysely migration head when launched outside the local helper. |
| Bounded runtime connection pool | **Adopted** | `backend/src/catalog/repository.ts` sets `max: 5` and `connectionTimeoutMillis: 3000`, and closes Kysely with the Fastify lifecycle. |
| `lock_timeout`, `statement_timeout`, and catalog TLS policy | **Pending** | The runtime pool does not set query or lock timeouts. Catalog connection configuration has no explicit CA/hostname-verification contract; the server profile still uses the private Compose network. |
| Process liveness independent of catalog queries | **Adopted** | `GET /health` returns `{status: "ok"}` without probing PostgreSQL, and `backend/test/health.test.ts` preserves that behavior. `/catalog/status` performs the catalog-specific availability check. |
| Stack startup independent of catalog availability | **Superseded/rejected** | Compose blocks `core` on healthy `catalog-db`. The host service health fold does not list `catalog-db`, but it cannot make `core` start while its Compose dependency is unhealthy. |
| Uniform catalog-outage response (`503` plus `Retry-After`) | **Pending** | Routes map the domain `CatalogUnavailableError` to a sanitized `503`, and an omitted catalog configuration uses an unavailable repository. PostgreSQL driver failures are not uniformly translated to that domain error, and no `Retry-After` contract is implemented. |
| Catalog-specific readiness and `tht doctor` checks | **Pending** | There is no `/health/ready` for the catalog and no doctor section for connection, migration head, pool saturation, backup age, or publication lag (`tools/tht/internal/doctor/report.go`). |
| Logical catalog backup and restore | **Pending** | Current backup archives omit `catalog-data` and do not run `pg_dump`; server archives include only Qdrant and embedding volumes. A backup can therefore succeed without preserving the Metadata Catalog (`tools/tht/internal/backup/create.go`). |
| Last-good Publication and catalog-to-Qdrant cutover | **Pending** | The current database-management slice does not change the NL→SQL runtime or publish catalog metadata to Qdrant. |
Do not silently fall back to JSON files, dual-write, or read unpublished PostgreSQL rows from the
core. Those paths would create split-brain state. The only degraded-mode contract is use of the
last successfully activated Publication.
## Current operational contract
## Backup and restore
### Deployment and credentials
- Local and server Compose profiles currently use the same installation-owned `catalog-db`
container and `catalog-data` named volume. PostgreSQL is private to the Compose network.
- The database bootstrap login is the migrator. The init script creates the separate runtime login
from a Docker secret and grants only runtime DML and sequence access.
- Runtime and migrator passwords are separate protected host files exposed as separate Docker
secrets. Neither value belongs in tracked environment files
(`deploy/env/local.env.example`, `deploy/env/server.env.example`).
- The Fastify catalog repository uses a five-connection pool with a three-second connection
timeout. Fastify closes the repository pool during shutdown.
### Migrations
The compiled `catalog-migrate` entry point owns schema changes and uses the migrator credential.
The current ordered series is under
`backend/src/catalog/migrations/`. The local launcher runs
it before the normal stack, while production rollout must invoke the profile-gated service
explicitly.
This is weaker than the original recommendation in two ways: there is no catalog readiness check
for pending or unknown migrations, and the repository does not maintain content checksums for
migration drift. Until those checks exist, “the migrator completed” is the available deployment
gate; the application itself does not prove migration compatibility.
### Failure behavior actually provided
| Failure | Current catalog behavior | Current workflow behavior |
| --- | --- | --- |
| Catalog configuration omitted when launching the backend directly | Fastify uses `UnavailableCatalogRepository`; `/catalog/status` reports unavailable and domain-mapped catalog operations return sanitized `503` | `/health`, sessions, and SSE remain available |
| `catalog-db` unhealthy before Compose startup | `core` is not started because its dependency is not healthy | Workflow startup is blocked |
| PostgreSQL becomes unavailable after startup | `/health` remains process-only and `/catalog/status` reports unavailable; route errors are sanitized, but a uniform `503`/`Retry-After` mapping is not guaranteed | Existing session code does not read catalog rows, but both surfaces still share one Fastify process |
| Migrations are pending or incompatible | No dedicated readiness refusal exists; affected catalog operations fail | No catalog-to-workflow handoff exists yet, but local deployment correctness depends on running `catalog-migrate` first |
| Catalog backup is requested through current `tht backup` | No PostgreSQL dump is added and `catalog-data` is omitted | The archive may succeed while being unable to restore catalog state |
The shared process means resource exhaustion, fatal process errors, and startup hooks remain a
common failure domain even though repository calls are separated. Conversely, placing the module
inside Fastify does not require workflow code to consume catalog rows: preserving that data-flow
boundary is still the useful part of the original isolation recommendation.
## Historical recommendations that remain valid backlog
The following constraints survive ADR-0004 because they can be implemented inside the current
Fastify deployment:
1. **Bound every database operation.** Keep the adopted pool and connection timeout; add explicit
request query and lock timeouts. Long model calls and source sampling must not hold catalog
transactions.
2. **Make failures component-specific.** Translate connection, timeout, and pool-exhaustion errors
into one sanitized catalog-unavailable response, add bounded retry guidance, and keep `/health`
process-only.
3. **Make migration compatibility observable.** Report applied, pending, unknown, and drifted
migrations through a catalog readiness check and `tht doctor`; do not put migrator credentials
in `core`.
4. **Define a server TLS topology before externalizing PostgreSQL.** If the server profile moves to
an operator-provided endpoint, use a dedicated database and roles, protected CA material, and
hostname verification. Sharing a cluster leaves connection, maintenance, WAL, and storage blast
radius even when schemas are separate. PostgreSQL documents the relevant connection and TLS
parameters in its
[connection parameter reference](https://www.postgresql.org/docs/current/libpq-connect.html).
5. **Keep derived semantic data non-canonical.** When catalog publication is implemented, activate
complete immutable revisions and retain the last successfully activated revision rather than
exposing mutable catalog rows to the workflow.
## Backup and restore target
The original logical-backup recommendation is still valid and is now a confirmed implementation
gap.
### Backup
Extend `tht backup create` or add a catalog-specific command invoked by it. Create a logical
`pg_dump --format=custom` archive through a one-shot helper, checksum it, and place it in the
installation archive manifest. Do **not** tar a live PostgreSQL data volume: the existing raw-volume
backup mechanism was designed for the current named volumes, not database crash consistency.
Extend `tht backup create` with a catalog step that runs `pg_dump --format=custom` through a
one-shot helper and adds the dump plus checksum to the installation manifest. Do not treat a tar of
a live data volume as a PostgreSQL consistency contract. PostgreSQL documents that `pg_dump`
creates a consistent export while the database remains in use and that custom format supports
selective restore ([`pg_dump`](https://www.postgresql.org/docs/current/app-pgdump.html)).
`pg_dump` produces a consistent export while the database remains in use, and custom format is
compressed and supports selective/reordered restore
([PostgreSQL `pg_dump`](https://www.postgresql.org/docs/current/app-pgdump.html)). Because metadata
volume is small, favor correctness over parallelism. Run a backup before every migration and on an
operator-defined schedule; record the server version, migration head, dump checksum, creation time,
and catalog Publication identifiers in the manifest. Never write a dump into the workspace Git
repository.
For an external server database, the operator must supply the dedicated backup credential through
the same secret-file boundary. A successful filesystem/Qdrant archive with a failed catalog dump
is an incomplete installation backup and must not be reported as success.
Use a dedicated least-privilege backup role, never write dumps into a workspace repository, and
record at least the server version, migration head, checksum, creation time, and catalog revision
identifiers. A filesystem/Qdrant archive with a failed or absent catalog dump must be reported as
incomplete.
### Restore
Restore only during an explicit catalog maintenance window into an empty target database (or a
deliberately cleaned target), using `pg_restore --exit-on-error --single-transaction`. Validate the
archive checksum first, then require:
Restore only through an explicit maintenance operation into an empty or deliberately cleaned
target. Validate the archive checksum, run `pg_restore --exit-on-error --single-transaction`, then
verify migration compatibility, referential integrity, bounded entity counts, and an authenticated
catalog smoke test. The controls are documented by PostgreSQL
([`pg_restore`](https://www.postgresql.org/docs/current/app-pgrestore.html)). Keep the old database
until validation passes; rollback should switch the endpoint or retained volume, not dual-write.
1. expected migration history with no checksum drift;
2. referential-integrity checks and bounded entity counts;
3. Publication content hashes matching the restored records;
4. an authenticated catalog smoke test.
## Diagnostics target
The relevant restore controls and their behavior are documented by
[PostgreSQL `pg_restore`](https://www.postgresql.org/docs/current/app-pgrestore.html). After restore,
rebuild catalog-derived Qdrant content from the restored active Publication rather than treating
the vector index as the source of truth. Keep the old database untouched until this validation
passes; rollback is endpoint reversal, not dual-write.
- Keep the shared `GET /health` endpoint process-only.
- Treat `/catalog/status` as the current minimal availability surface; add a bounded catalog
readiness check covering connection and migration compatibility before using it as a rollout
gate.
- Add read-only `tht doctor` checks for secret-file presence and permissions, PostgreSQL
reachability, authentication, migration state, pool saturation, and last successful catalog
backup. Sanitize all DSNs and errors.
- Use `pg_isready` only for server transport readiness. Its result does not prove schema or runtime
authorization correctness
([`pg_isready`](https://www.postgresql.org/docs/current/app-pg-isready.html)).
- Add structured metrics for connection acquisition failures, pool use, query/lock timeout,
migration head, background-job backlog, and backup age. Do not log connection strings, source
samples, prompts, or generated descriptions at info level.
## Diagnostics and observability
## Recommended closure criteria
- Keep `core /health` unchanged. Give `catalog-api` separate `/health/live` (process only) and
`/health/ready` (bounded connection, `SELECT 1`, and migration compatibility) endpoints.
- Add a named `metadata-catalog` section to `tht doctor`: configuration completeness; secret-file
presence and permissions; DNS/TCP/TLS/authentication; `SELECT 1`; migration applied/pending/drifted
state; pool saturation; last successful backup age; and oldest pending Publication. Checks are
read-only and all errors/DSNs are sanitized. The current doctor already separates Compose service
status from container-local workflow checks
([`report.go`, lines 275–310](../../tools/tht/internal/doctor/report.go#L275-L310)).
- Use `pg_isready` only for transport readiness; its exit codes distinguish accepting, rejecting,
no-response, and invalid-parameter states, but it does not prove schema or authorization
correctness ([PostgreSQL `pg_isready`](https://www.postgresql.org/docs/current/app-pg-isready.html)).
- Expose structured metrics/logs for connection acquisition failures, pool utilization, query
timeout/lock timeout, migration head, AI-job backlog, Publication lag, and backup age. Never log
connection strings, prompts containing schema metadata, or generated descriptions at info level.
- `tht status` may show the catalog as `healthy`, `degraded`, or `disabled`, but catalog state must
not change the existing core health result. `tht doctor` may return a failing catalog check for
operator visibility without stopping or restarting the workflow stack.
## Acceptance constraints for later design and implementation
1. Killing `catalog-postgres` and `catalog-api` must not change `core /health`, stop Pi, interrupt an
existing SQL session, or prevent a new session that uses the already active Publication.
2. No catalog database secret, client dependency, migration, or network call exists in the harness
workflow path.
3. Catalog failure is explicit on catalog routes; no JSON fallback or dual-write is introduced.
4. A failed Publication leaves the previous active semantic content addressable and retryable.
5. Local and server backup tests perform a real dump/restore round trip and prove that a backup is
incomplete when the PostgreSQL dump fails.
6. Migration credentials are absent from runtime containers; drift and pending migrations are
visible in readiness and `tht doctor`.
1. Decide whether ADR-0004's statement that catalog unavailability does not make sessions or SSE
unavailable must also hold at Compose startup. If yes, remove or soften the hard `core` →
`catalog-db` startup dependency without reintroducing a microservice.
2. Prove a configured PostgreSQL outage produces the same sanitized catalog response across every
catalog route while `/health`, session creation, and SSE continue to work.
3. Add catalog migration compatibility to readiness and `tht doctor`, including pending, unknown,
and drifted states.
4. Add a real `pg_dump`/`pg_restore` round-trip to local and server backup tests, and fail backup
publication when the catalog dump is absent or fails.
5. Before enabling a remote server catalog, define TLS verification, credential files, connection
limits, and the accepted cluster-level blast radius.
6. Before the NL→SQL cutover, define and test an immutable last-good Publication boundary; do not
dual-read mutable PostgreSQL rows and legacy metadata as competing authorities.
## Decision summary
The Metadata Catalog should be a separately deployed bounded context backed by PostgreSQL. Local
installations own a private container and volume; server installations use a TLS-verified external
database, preferably on a separate cluster when strict blast-radius isolation is required. Logical
dump/restore, one-shot migrations, role separation, bounded connections, component-specific
diagnostics, and last-good-Publication semantics make PostgreSQL operationally safe without making
it a dependency of the SQL workflow.
PostgreSQL, installation-local bindings, private deployment, role separation, one-shot migrations,
bounded pooling, and process-only liveness are implemented. The separate `catalog-api` process and
the claim that `core` has no catalog dependency are not current design: ADR-0004 chose the existing
Fastify process, and Compose currently blocks `core` startup on `catalog-db` health. Logical
backup/restore, catalog migration readiness and drift detection, consistent outage mapping,
catalog-specific doctor checks, server TLS/external topology, and last-good publication remain
pending. Those open constraints should be treated as backlog, not as capabilities already provided
by the repository.