Files
ThothII/docs/adr/0016-use-postgres-metadata-for-all-core-consumers.md
T
Codex cffa60772e
Publish documentation / publish (push) Successful in 2m12s
feat: complete catalog-driven preprocessing
2026-09-06 17:49:35 +02:00

64 lines
4.3 KiB
Markdown

# Use PostgreSQL metadata for all core consumers
The Metadata Catalog in PostgreSQL is the sole authority for every fact about a Workspace Database
used by the core. This completes the cutover anticipated by ADR-0001 and ADR-0003 and supersedes the
ADR-0012 statements that kept workspace annotations authoritative for descriptions and materialized
a relationship-only snapshot per session. The Workspace Descriptor retains workspace identity and
Evidence scope, but contains no database identity, binding, structure, description, or relationship
data.
## Decision
All downstream consumers receive one deterministic, versioned **Catalog Metadata Snapshot** produced
from PostgreSQL. It contains the complete cataloged table and column structure, sensitivity metadata,
publishable descriptions, and the active Effective Relationship Map. Description precedence is
`Description`, then `Generated Description`, then the observed source comment; generated text is
therefore publishable without prior consolidation. Qdrant stores first-class table, column, and
relationship records, and a relationship hit contributes to both endpoint tables.
The native host CLI exposes one mutating preprocessing operation. The Node
`workspace-maintenance` process acquires a PostgreSQL lock shared with catalog synchronization,
description generation, and administrative catalog writes; it refuses to run while one of those
operations is active. It reads the catalog once, writes the snapshot to a temporary immutable JSON
file, and passes only that file to the Python harness. Catalog credentials are excluded from the
child environment. The harness never reads the Metadata Catalog directly.
Preprocessing is deliberately offline: no session uses the core while it is running, nobody starts
the core until it completes, and concurrent catalog modification is forbidden. PostgreSQL owns a
minimal `running | succeeded | failed` Preprocessing State, the input fingerprint, the Metadata
Content Revision, and the last successfully processed revision. Every relevant catalog mutation
increments the revision and makes the prior success stale in the same transaction. Core admission
requires a current `succeeded` state, but the system does not implement session draining, per-session
schema pinning, coexisting preprocessing generations, retention, or rollback.
Each run may overwrite the previous derived state. It replaces the Qdrant `schema` and `evidence`
slices while preserving `memory` and `solved_question`, rebuilds the current LSH and other required
artifacts, verifies the result, and marks success only at the end. A failure leaves the workspace
unready; rerunning starts from a clean derived state. Existing internal DWH or Evidence staging may
remain an implementation detail with minimum retention, but it is not an application-level
generation contract. A workspace without Evidence is valid and publishes an explicit absent state.
`physical.yaml`, `annotations.yaml`, the YAML FK review flow, and public partial preprocessing
commands are removed from the runtime contract. No legacy metadata is imported. Concepts, synonyms,
notes, annotation evidence, and the manual `eligible` override are retired; eligibility is computed
deterministically. The DWH is queried during preprocessing only for derived samples and LSH values,
not as an alternative metadata authority. The existing **Catalog Schema Snapshot** name remains
reserved for the DWH-to-catalog synchronization RPC.
## Considered Options
- Direct PostgreSQL access from every consumer would duplicate catalog queries and expose storage
details and credentials to the harness.
- Immutable application-level generations and an atomic composite release would permit concurrent
core use and rollback, but those capabilities are unnecessary under the accepted offline operating
constraints.
- Retaining or importing workspace annotations would preserve legacy semantic fields but would keep
an obsolete second authority despite there being no production data to migrate.
## Consequences
Every runtime metadata consumer uses the Catalog snapshot, not only Qdrant indexing. Read-only
inspection may remain, but only the complete preprocessing command can mutate derived state. The
workspace-preprocessing contract and tests describe this cutover directly; superseded partial
command designs remain available only in Git history.