Files
ThothII/docs/adr/0008-make-hybrid-evidence-retrieval-deterministic.md
T

25 lines
1.6 KiB
Markdown

# Make hybrid Evidence retrieval deterministic and diagnosable
Dense and BM25 retrieval receive the same deterministic query text. It preserves the
original question and appends nonempty concepts, tables and columns in a fixed order;
question and context receive Unicode NFC, newline canonicalization and outer trimming.
Context values are then deduplicated exactly and sorted, without lowercasing. Case,
punctuation and internal whitespace remain intact. This avoids accidental ranking
changes caused only by metadata ordering, preserves quoted PostgreSQL identifiers and
keeps the two retrieval branches directly comparable.
Evidence fragmentation follows semantic headings, typed fields and paragraph
boundaries. Formulas, value/meaning pairs, mappings, rules and URLs remain atomic. An
atomic element larger than the existing `max_chunk_chars` limit creates the blocking
`atomic_content_too_large` Review item rather than being split mechanically. The limit
applies to the complete rendered text, defaults to 4,000 characters and is not
duplicated by an Evidence-specific setting.
An L0 contract test starts the exact Qdrant image referenced by `compose.yaml` and
proves Italian server-side `qdrant/bm25` ingestion and search in a temporary collection.
There is no FastEmbed or dense fallback for Evidence when that capability is absent.
The versioned evaluation set contains lexical, semantic and mixed queries. Its report
shows dense-only, BM25-only and fused ranks for expected Evidence. Publication remains
governed only by the simple fused top-10 rule; branch ranks and hit@5 are diagnostic.