feat: complete catalog-driven preprocessing
Publish documentation / publish (push) Successful in 2m12s
Publish documentation / publish (push) Successful in 2m12s
This commit is contained in:
@@ -0,0 +1,48 @@
|
||||
# Separate reference vectors from runtime memory
|
||||
|
||||
Schema, relationships, and Evidence are deterministic projections of PostgreSQL Catalog metadata
|
||||
and the pinned workspace Evidence revision. Memory and solved questions are produced by core use and
|
||||
have a different lifecycle. Keeping both classes of records in one Qdrant collection makes a full
|
||||
cleanup of derived preprocessing data unsafe because it can also erase durable runtime knowledge.
|
||||
|
||||
## Decision
|
||||
|
||||
Each workspace has two physical Qdrant collections behind the existing logical vector-store port:
|
||||
|
||||
- `<workspace>-reference` contains `schema_table`, `schema_column`, `schema_relationship`, and
|
||||
`evidence`;
|
||||
- `<workspace>-memory` contains `memory` and `solved_question`.
|
||||
|
||||
Callers continue to use logical record groups. The Qdrant adapter owns physical routing and merges
|
||||
results when a logical search spans both collections. BM25 configuration belongs only to the
|
||||
reference collection.
|
||||
|
||||
The public preprocessing clear operation deletes the reference collection, LSH, Evidence corpus,
|
||||
private Catalog snapshot, and derived checkpoints. It never deletes the memory collection, session
|
||||
artifacts, Catalog metadata, or source data. PostgreSQL records `derived_data_cleared`, which projects
|
||||
to a required preprocessing state and prevents core admission until a complete run succeeds.
|
||||
|
||||
On the first preprocessing write or clear after this split, the adapter copies `memory` and
|
||||
`solved_question` points from an existing legacy workspace-named collection into the new memory
|
||||
collection before retiring the legacy collection. There is no general metadata migration, history,
|
||||
generation retention, or rollback mechanism.
|
||||
|
||||
LSH ownership includes workspace ID, Catalog database ID, Metadata Content Revision, effective
|
||||
configuration fingerprint, and input fingerprint. Reuse fails closed when any identity differs.
|
||||
|
||||
## Considered Options
|
||||
|
||||
- Deleting only Schema/Evidence points from one shared collection would keep memory but retain a
|
||||
coupled physical lifecycle and make collection-level repair or recreation unsafe.
|
||||
- Rebuilding memory after clear is impossible because preprocessing does not own its source of
|
||||
truth.
|
||||
- Keeping multiple preprocessing generations would add retention and rollback behavior that the
|
||||
accepted offline operating model does not require.
|
||||
|
||||
## Consequences
|
||||
|
||||
Preprocessing can now clear and recreate all of its vector output without touching runtime memory.
|
||||
Diagnostics and effective configuration must validate two collections. Operators get one clear
|
||||
command and one inline-confirmed sidebar action; after either succeeds, preprocessing is required.
|
||||
The split remains encapsulated in the vector adapter, so workflow and retrieval callers do not need
|
||||
to know physical Qdrant collection names.
|
||||
Reference in New Issue
Block a user