2.7 KiB
Separate reference vectors from runtime memory
Schema, relationships, and Evidence are deterministic projections of PostgreSQL Catalog metadata and the pinned workspace Evidence revision. Memory and solved questions are produced by core use and have a different lifecycle. Keeping both classes of records in one Qdrant collection makes a full cleanup of derived preprocessing data unsafe because it can also erase durable runtime knowledge.
Decision
Each workspace has two physical Qdrant collections behind the existing logical vector-store port:
<workspace>-referencecontainsschema_table,schema_column,schema_relationship, andevidence;<workspace>-memorycontainsmemoryandsolved_question.
Callers continue to use logical record groups. The Qdrant adapter owns physical routing and merges results when a logical search spans both collections. BM25 configuration belongs only to the reference collection.
The public preprocessing clear operation deletes the reference collection, LSH, Evidence corpus,
private Catalog snapshot, and derived checkpoints. It never deletes the memory collection, session
artifacts, Catalog metadata, or source data. PostgreSQL records derived_data_cleared, which projects
to a required preprocessing state and prevents core admission until a complete run succeeds.
On the first preprocessing write or clear after this split, the adapter copies memory and
solved_question points from an existing legacy workspace-named collection into the new memory
collection before retiring the legacy collection. There is no general metadata migration, history,
generation retention, or rollback mechanism.
LSH ownership includes workspace ID, Catalog database ID, Metadata Content Revision, effective configuration fingerprint, and input fingerprint. Reuse fails closed when any identity differs.
Considered Options
- Deleting only Schema/Evidence points from one shared collection would keep memory but retain a coupled physical lifecycle and make collection-level repair or recreation unsafe.
- Rebuilding memory after clear is impossible because preprocessing does not own its source of truth.
- Keeping multiple preprocessing generations would add retention and rollback behavior that the accepted offline operating model does not require.
Consequences
Preprocessing can now clear and recreate all of its vector output without touching runtime memory. Diagnostics and effective configuration must validate two collections. Operators get one clear command and one inline-confirmed sidebar action; after either succeeds, preprocessing is required. The split remains encapsulated in the vector adapter, so workflow and retrieval callers do not need to know physical Qdrant collection names.