Files
ThothII/docs/adr/0017-separate-reference-vectors-from-runtime-memory.md
Codex cffa60772e
Publish documentation / publish (push) Successful in 2m12s
feat: complete catalog-driven preprocessing
2026-09-06 17:49:35 +02:00

49 lines
2.7 KiB
Markdown

# Separate reference vectors from runtime memory
Schema, relationships, and Evidence are deterministic projections of PostgreSQL Catalog metadata
and the pinned workspace Evidence revision. Memory and solved questions are produced by core use and
have a different lifecycle. Keeping both classes of records in one Qdrant collection makes a full
cleanup of derived preprocessing data unsafe because it can also erase durable runtime knowledge.
## Decision
Each workspace has two physical Qdrant collections behind the existing logical vector-store port:
- `<workspace>-reference` contains `schema_table`, `schema_column`, `schema_relationship`, and
`evidence`;
- `<workspace>-memory` contains `memory` and `solved_question`.
Callers continue to use logical record groups. The Qdrant adapter owns physical routing and merges
results when a logical search spans both collections. BM25 configuration belongs only to the
reference collection.
The public preprocessing clear operation deletes the reference collection, LSH, Evidence corpus,
private Catalog snapshot, and derived checkpoints. It never deletes the memory collection, session
artifacts, Catalog metadata, or source data. PostgreSQL records `derived_data_cleared`, which projects
to a required preprocessing state and prevents core admission until a complete run succeeds.
On the first preprocessing write or clear after this split, the adapter copies `memory` and
`solved_question` points from an existing legacy workspace-named collection into the new memory
collection before retiring the legacy collection. There is no general metadata migration, history,
generation retention, or rollback mechanism.
LSH ownership includes workspace ID, Catalog database ID, Metadata Content Revision, effective
configuration fingerprint, and input fingerprint. Reuse fails closed when any identity differs.
## Considered Options
- Deleting only Schema/Evidence points from one shared collection would keep memory but retain a
coupled physical lifecycle and make collection-level repair or recreation unsafe.
- Rebuilding memory after clear is impossible because preprocessing does not own its source of
truth.
- Keeping multiple preprocessing generations would add retention and rollback behavior that the
accepted offline operating model does not require.
## Consequences
Preprocessing can now clear and recreate all of its vector output without touching runtime memory.
Diagnostics and effective configuration must validate two collections. Operators get one clear
command and one inline-confirmed sidebar action; after either succeeds, preprocessing is required.
The split remains encapsulated in the vector adapter, so workflow and retrieval callers do not need
to know physical Qdrant collection names.