feat: complete catalog-driven preprocessing
Publish documentation / publish (push) Successful in 2m12s

This commit is contained in:
Codex
2026-09-06 17:49:35 +02:00
parent 8707ae1d46
commit cffa60772e
141 changed files with 5898 additions and 3015 deletions
@@ -0,0 +1,63 @@
# Use PostgreSQL metadata for all core consumers
The Metadata Catalog in PostgreSQL is the sole authority for every fact about a Workspace Database
used by the core. This completes the cutover anticipated by ADR-0001 and ADR-0003 and supersedes the
ADR-0012 statements that kept workspace annotations authoritative for descriptions and materialized
a relationship-only snapshot per session. The Workspace Descriptor retains workspace identity and
Evidence scope, but contains no database identity, binding, structure, description, or relationship
data.
## Decision
All downstream consumers receive one deterministic, versioned **Catalog Metadata Snapshot** produced
from PostgreSQL. It contains the complete cataloged table and column structure, sensitivity metadata,
publishable descriptions, and the active Effective Relationship Map. Description precedence is
`Description`, then `Generated Description`, then the observed source comment; generated text is
therefore publishable without prior consolidation. Qdrant stores first-class table, column, and
relationship records, and a relationship hit contributes to both endpoint tables.
The native host CLI exposes one mutating preprocessing operation. The Node
`workspace-maintenance` process acquires a PostgreSQL lock shared with catalog synchronization,
description generation, and administrative catalog writes; it refuses to run while one of those
operations is active. It reads the catalog once, writes the snapshot to a temporary immutable JSON
file, and passes only that file to the Python harness. Catalog credentials are excluded from the
child environment. The harness never reads the Metadata Catalog directly.
Preprocessing is deliberately offline: no session uses the core while it is running, nobody starts
the core until it completes, and concurrent catalog modification is forbidden. PostgreSQL owns a
minimal `running | succeeded | failed` Preprocessing State, the input fingerprint, the Metadata
Content Revision, and the last successfully processed revision. Every relevant catalog mutation
increments the revision and makes the prior success stale in the same transaction. Core admission
requires a current `succeeded` state, but the system does not implement session draining, per-session
schema pinning, coexisting preprocessing generations, retention, or rollback.
Each run may overwrite the previous derived state. It replaces the Qdrant `schema` and `evidence`
slices while preserving `memory` and `solved_question`, rebuilds the current LSH and other required
artifacts, verifies the result, and marks success only at the end. A failure leaves the workspace
unready; rerunning starts from a clean derived state. Existing internal DWH or Evidence staging may
remain an implementation detail with minimum retention, but it is not an application-level
generation contract. A workspace without Evidence is valid and publishes an explicit absent state.
`physical.yaml`, `annotations.yaml`, the YAML FK review flow, and public partial preprocessing
commands are removed from the runtime contract. No legacy metadata is imported. Concepts, synonyms,
notes, annotation evidence, and the manual `eligible` override are retired; eligibility is computed
deterministically. The DWH is queried during preprocessing only for derived samples and LSH values,
not as an alternative metadata authority. The existing **Catalog Schema Snapshot** name remains
reserved for the DWH-to-catalog synchronization RPC.
## Considered Options
- Direct PostgreSQL access from every consumer would duplicate catalog queries and expose storage
details and credentials to the harness.
- Immutable application-level generations and an atomic composite release would permit concurrent
core use and rollback, but those capabilities are unnecessary under the accepted offline operating
constraints.
- Retaining or importing workspace annotations would preserve legacy semantic fields but would keep
an obsolete second authority despite there being no production data to migrate.
## Consequences
Every runtime metadata consumer uses the Catalog snapshot, not only Qdrant indexing. Read-only
inspection may remain, but only the complete preprocessing command can mutate derived state. The
workspace-preprocessing contract and tests describe this cutover directly; superseded partial
command designs remain available only in Git history.
@@ -0,0 +1,48 @@
# Separate reference vectors from runtime memory
Schema, relationships, and Evidence are deterministic projections of PostgreSQL Catalog metadata
and the pinned workspace Evidence revision. Memory and solved questions are produced by core use and
have a different lifecycle. Keeping both classes of records in one Qdrant collection makes a full
cleanup of derived preprocessing data unsafe because it can also erase durable runtime knowledge.
## Decision
Each workspace has two physical Qdrant collections behind the existing logical vector-store port:
- `<workspace>-reference` contains `schema_table`, `schema_column`, `schema_relationship`, and
`evidence`;
- `<workspace>-memory` contains `memory` and `solved_question`.
Callers continue to use logical record groups. The Qdrant adapter owns physical routing and merges
results when a logical search spans both collections. BM25 configuration belongs only to the
reference collection.
The public preprocessing clear operation deletes the reference collection, LSH, Evidence corpus,
private Catalog snapshot, and derived checkpoints. It never deletes the memory collection, session
artifacts, Catalog metadata, or source data. PostgreSQL records `derived_data_cleared`, which projects
to a required preprocessing state and prevents core admission until a complete run succeeds.
On the first preprocessing write or clear after this split, the adapter copies `memory` and
`solved_question` points from an existing legacy workspace-named collection into the new memory
collection before retiring the legacy collection. There is no general metadata migration, history,
generation retention, or rollback mechanism.
LSH ownership includes workspace ID, Catalog database ID, Metadata Content Revision, effective
configuration fingerprint, and input fingerprint. Reuse fails closed when any identity differs.
## Considered Options
- Deleting only Schema/Evidence points from one shared collection would keep memory but retain a
coupled physical lifecycle and make collection-level repair or recreation unsafe.
- Rebuilding memory after clear is impossible because preprocessing does not own its source of
truth.
- Keeping multiple preprocessing generations would add retention and rollback behavior that the
accepted offline operating model does not require.
## Consequences
Preprocessing can now clear and recreate all of its vector output without touching runtime memory.
Diagnostics and effective configuration must validate two collections. Operators get one clear
command and one inline-confirmed sidebar action; after either succeeds, preprocessing is required.
The split remains encapsulated in the vector adapter, so workflow and retrieval callers do not need
to know physical Qdrant collection names.