Files
ThothII/docs/adr/0008-make-hybrid-evidence-retrieval-deterministic.md
T

1.6 KiB

Make hybrid Evidence retrieval deterministic and diagnosable

Dense and BM25 retrieval receive the same deterministic query text. It preserves the original question and appends nonempty concepts, tables and columns in a fixed order; question and context receive Unicode NFC, newline canonicalization and outer trimming. Context values are then deduplicated exactly and sorted, without lowercasing. Case, punctuation and internal whitespace remain intact. This avoids accidental ranking changes caused only by metadata ordering, preserves quoted PostgreSQL identifiers and keeps the two retrieval branches directly comparable.

Evidence fragmentation follows semantic headings, typed fields and paragraph boundaries. Formulas, value/meaning pairs, mappings, rules and URLs remain atomic. An atomic element larger than the existing max_chunk_chars limit creates the blocking atomic_content_too_large Review item rather than being split mechanically. The limit applies to the complete rendered text, defaults to 4,000 characters and is not duplicated by an Evidence-specific setting.

An L0 contract test starts the exact Qdrant image referenced by compose.yaml and proves Italian server-side qdrant/bm25 ingestion and search in a temporary collection. There is no FastEmbed or dense fallback for Evidence when that capability is absent.

The versioned evaluation set contains lexical, semantic and mixed queries. Its report shows dense-only, BM25-only and fused ranks for expected Evidence. Publication remains governed only by the simple fused top-10 rule; branch ranks and hit@5 are diagnostic.