feat: complete catalog-driven preprocessing
Publish documentation / publish (push) Successful in 2m12s
Publish documentation / publish (push) Successful in 2m12s
This commit is contained in:
@@ -0,0 +1,63 @@
|
||||
# Use PostgreSQL metadata for all core consumers
|
||||
|
||||
The Metadata Catalog in PostgreSQL is the sole authority for every fact about a Workspace Database
|
||||
used by the core. This completes the cutover anticipated by ADR-0001 and ADR-0003 and supersedes the
|
||||
ADR-0012 statements that kept workspace annotations authoritative for descriptions and materialized
|
||||
a relationship-only snapshot per session. The Workspace Descriptor retains workspace identity and
|
||||
Evidence scope, but contains no database identity, binding, structure, description, or relationship
|
||||
data.
|
||||
|
||||
## Decision
|
||||
|
||||
All downstream consumers receive one deterministic, versioned **Catalog Metadata Snapshot** produced
|
||||
from PostgreSQL. It contains the complete cataloged table and column structure, sensitivity metadata,
|
||||
publishable descriptions, and the active Effective Relationship Map. Description precedence is
|
||||
`Description`, then `Generated Description`, then the observed source comment; generated text is
|
||||
therefore publishable without prior consolidation. Qdrant stores first-class table, column, and
|
||||
relationship records, and a relationship hit contributes to both endpoint tables.
|
||||
|
||||
The native host CLI exposes one mutating preprocessing operation. The Node
|
||||
`workspace-maintenance` process acquires a PostgreSQL lock shared with catalog synchronization,
|
||||
description generation, and administrative catalog writes; it refuses to run while one of those
|
||||
operations is active. It reads the catalog once, writes the snapshot to a temporary immutable JSON
|
||||
file, and passes only that file to the Python harness. Catalog credentials are excluded from the
|
||||
child environment. The harness never reads the Metadata Catalog directly.
|
||||
|
||||
Preprocessing is deliberately offline: no session uses the core while it is running, nobody starts
|
||||
the core until it completes, and concurrent catalog modification is forbidden. PostgreSQL owns a
|
||||
minimal `running | succeeded | failed` Preprocessing State, the input fingerprint, the Metadata
|
||||
Content Revision, and the last successfully processed revision. Every relevant catalog mutation
|
||||
increments the revision and makes the prior success stale in the same transaction. Core admission
|
||||
requires a current `succeeded` state, but the system does not implement session draining, per-session
|
||||
schema pinning, coexisting preprocessing generations, retention, or rollback.
|
||||
|
||||
Each run may overwrite the previous derived state. It replaces the Qdrant `schema` and `evidence`
|
||||
slices while preserving `memory` and `solved_question`, rebuilds the current LSH and other required
|
||||
artifacts, verifies the result, and marks success only at the end. A failure leaves the workspace
|
||||
unready; rerunning starts from a clean derived state. Existing internal DWH or Evidence staging may
|
||||
remain an implementation detail with minimum retention, but it is not an application-level
|
||||
generation contract. A workspace without Evidence is valid and publishes an explicit absent state.
|
||||
|
||||
`physical.yaml`, `annotations.yaml`, the YAML FK review flow, and public partial preprocessing
|
||||
commands are removed from the runtime contract. No legacy metadata is imported. Concepts, synonyms,
|
||||
notes, annotation evidence, and the manual `eligible` override are retired; eligibility is computed
|
||||
deterministically. The DWH is queried during preprocessing only for derived samples and LSH values,
|
||||
not as an alternative metadata authority. The existing **Catalog Schema Snapshot** name remains
|
||||
reserved for the DWH-to-catalog synchronization RPC.
|
||||
|
||||
## Considered Options
|
||||
|
||||
- Direct PostgreSQL access from every consumer would duplicate catalog queries and expose storage
|
||||
details and credentials to the harness.
|
||||
- Immutable application-level generations and an atomic composite release would permit concurrent
|
||||
core use and rollback, but those capabilities are unnecessary under the accepted offline operating
|
||||
constraints.
|
||||
- Retaining or importing workspace annotations would preserve legacy semantic fields but would keep
|
||||
an obsolete second authority despite there being no production data to migrate.
|
||||
|
||||
## Consequences
|
||||
|
||||
Every runtime metadata consumer uses the Catalog snapshot, not only Qdrant indexing. Read-only
|
||||
inspection may remain, but only the complete preprocessing command can mutate derived state. The
|
||||
workspace-preprocessing contract and tests describe this cutover directly; superseded partial
|
||||
command designs remain available only in Git history.
|
||||
@@ -0,0 +1,48 @@
|
||||
# Separate reference vectors from runtime memory
|
||||
|
||||
Schema, relationships, and Evidence are deterministic projections of PostgreSQL Catalog metadata
|
||||
and the pinned workspace Evidence revision. Memory and solved questions are produced by core use and
|
||||
have a different lifecycle. Keeping both classes of records in one Qdrant collection makes a full
|
||||
cleanup of derived preprocessing data unsafe because it can also erase durable runtime knowledge.
|
||||
|
||||
## Decision
|
||||
|
||||
Each workspace has two physical Qdrant collections behind the existing logical vector-store port:
|
||||
|
||||
- `<workspace>-reference` contains `schema_table`, `schema_column`, `schema_relationship`, and
|
||||
`evidence`;
|
||||
- `<workspace>-memory` contains `memory` and `solved_question`.
|
||||
|
||||
Callers continue to use logical record groups. The Qdrant adapter owns physical routing and merges
|
||||
results when a logical search spans both collections. BM25 configuration belongs only to the
|
||||
reference collection.
|
||||
|
||||
The public preprocessing clear operation deletes the reference collection, LSH, Evidence corpus,
|
||||
private Catalog snapshot, and derived checkpoints. It never deletes the memory collection, session
|
||||
artifacts, Catalog metadata, or source data. PostgreSQL records `derived_data_cleared`, which projects
|
||||
to a required preprocessing state and prevents core admission until a complete run succeeds.
|
||||
|
||||
On the first preprocessing write or clear after this split, the adapter copies `memory` and
|
||||
`solved_question` points from an existing legacy workspace-named collection into the new memory
|
||||
collection before retiring the legacy collection. There is no general metadata migration, history,
|
||||
generation retention, or rollback mechanism.
|
||||
|
||||
LSH ownership includes workspace ID, Catalog database ID, Metadata Content Revision, effective
|
||||
configuration fingerprint, and input fingerprint. Reuse fails closed when any identity differs.
|
||||
|
||||
## Considered Options
|
||||
|
||||
- Deleting only Schema/Evidence points from one shared collection would keep memory but retain a
|
||||
coupled physical lifecycle and make collection-level repair or recreation unsafe.
|
||||
- Rebuilding memory after clear is impossible because preprocessing does not own its source of
|
||||
truth.
|
||||
- Keeping multiple preprocessing generations would add retention and rollback behavior that the
|
||||
accepted offline operating model does not require.
|
||||
|
||||
## Consequences
|
||||
|
||||
Preprocessing can now clear and recreate all of its vector output without touching runtime memory.
|
||||
Diagnostics and effective configuration must validate two collections. Operators get one clear
|
||||
command and one inline-confirmed sidebar action; after either succeeds, preprocessing is required.
|
||||
The split remains encapsulated in the vector adapter, so workflow and retrieval callers do not need
|
||||
to know physical Qdrant collection names.
|
||||
Reference in New Issue
Block a user