Files
ThothII/docs/adr/0016-use-postgres-metadata-for-all-core-consumers.md
Codex cffa60772e
Publish documentation / publish (push) Successful in 2m12s
feat: complete catalog-driven preprocessing
2026-09-06 17:49:35 +02:00

4.3 KiB

Use PostgreSQL metadata for all core consumers

The Metadata Catalog in PostgreSQL is the sole authority for every fact about a Workspace Database used by the core. This completes the cutover anticipated by ADR-0001 and ADR-0003 and supersedes the ADR-0012 statements that kept workspace annotations authoritative for descriptions and materialized a relationship-only snapshot per session. The Workspace Descriptor retains workspace identity and Evidence scope, but contains no database identity, binding, structure, description, or relationship data.

Decision

All downstream consumers receive one deterministic, versioned Catalog Metadata Snapshot produced from PostgreSQL. It contains the complete cataloged table and column structure, sensitivity metadata, publishable descriptions, and the active Effective Relationship Map. Description precedence is Description, then Generated Description, then the observed source comment; generated text is therefore publishable without prior consolidation. Qdrant stores first-class table, column, and relationship records, and a relationship hit contributes to both endpoint tables.

The native host CLI exposes one mutating preprocessing operation. The Node workspace-maintenance process acquires a PostgreSQL lock shared with catalog synchronization, description generation, and administrative catalog writes; it refuses to run while one of those operations is active. It reads the catalog once, writes the snapshot to a temporary immutable JSON file, and passes only that file to the Python harness. Catalog credentials are excluded from the child environment. The harness never reads the Metadata Catalog directly.

Preprocessing is deliberately offline: no session uses the core while it is running, nobody starts the core until it completes, and concurrent catalog modification is forbidden. PostgreSQL owns a minimal running | succeeded | failed Preprocessing State, the input fingerprint, the Metadata Content Revision, and the last successfully processed revision. Every relevant catalog mutation increments the revision and makes the prior success stale in the same transaction. Core admission requires a current succeeded state, but the system does not implement session draining, per-session schema pinning, coexisting preprocessing generations, retention, or rollback.

Each run may overwrite the previous derived state. It replaces the Qdrant schema and evidence slices while preserving memory and solved_question, rebuilds the current LSH and other required artifacts, verifies the result, and marks success only at the end. A failure leaves the workspace unready; rerunning starts from a clean derived state. Existing internal DWH or Evidence staging may remain an implementation detail with minimum retention, but it is not an application-level generation contract. A workspace without Evidence is valid and publishes an explicit absent state.

physical.yaml, annotations.yaml, the YAML FK review flow, and public partial preprocessing commands are removed from the runtime contract. No legacy metadata is imported. Concepts, synonyms, notes, annotation evidence, and the manual eligible override are retired; eligibility is computed deterministically. The DWH is queried during preprocessing only for derived samples and LSH values, not as an alternative metadata authority. The existing Catalog Schema Snapshot name remains reserved for the DWH-to-catalog synchronization RPC.

Considered Options

  • Direct PostgreSQL access from every consumer would duplicate catalog queries and expose storage details and credentials to the harness.
  • Immutable application-level generations and an atomic composite release would permit concurrent core use and rollback, but those capabilities are unnecessary under the accepted offline operating constraints.
  • Retaining or importing workspace annotations would preserve legacy semantic fields but would keep an obsolete second authority despite there being no production data to migrate.

Consequences

Every runtime metadata consumer uses the Catalog snapshot, not only Qdrant indexing. Read-only inspection may remain, but only the complete preprocessing command can mutate derived state. The workspace-preprocessing contract and tests describe this cutover directly; superseded partial command designs remain available only in Git history.