25 lines
1.6 KiB
Markdown
25 lines
1.6 KiB
Markdown
# Make hybrid Evidence retrieval deterministic and diagnosable
|
|
|
|
Dense and BM25 retrieval receive the same deterministic query text. It preserves the
|
|
original question and appends nonempty concepts, tables and columns in a fixed order;
|
|
question and context receive Unicode NFC, newline canonicalization and outer trimming.
|
|
Context values are then deduplicated exactly and sorted, without lowercasing. Case,
|
|
punctuation and internal whitespace remain intact. This avoids accidental ranking
|
|
changes caused only by metadata ordering, preserves quoted PostgreSQL identifiers and
|
|
keeps the two retrieval branches directly comparable.
|
|
|
|
Evidence fragmentation follows semantic headings, typed fields and paragraph
|
|
boundaries. Formulas, value/meaning pairs, mappings, rules and URLs remain atomic. An
|
|
atomic element larger than the existing `max_chunk_chars` limit creates the blocking
|
|
`atomic_content_too_large` Review item rather than being split mechanically. The limit
|
|
applies to the complete rendered text, defaults to 4,000 characters and is not
|
|
duplicated by an Evidence-specific setting.
|
|
|
|
An L0 contract test starts the exact Qdrant image referenced by `compose.yaml` and
|
|
proves Italian server-side `qdrant/bm25` ingestion and search in a temporary collection.
|
|
There is no FastEmbed or dense fallback for Evidence when that capability is absent.
|
|
|
|
The versioned evaluation set contains lexical, semantic and mixed queries. Its report
|
|
shows dense-only, BM25-only and fused ranks for expected Evidence. Publication remains
|
|
governed only by the simple fused top-10 rule; branch ranks and hit@5 are diagnostic.
|