docs(evidence): finalize ticketed restructuring specification
This commit is contained in:
@@ -0,0 +1,24 @@
|
||||
# Make hybrid Evidence retrieval deterministic and diagnosable
|
||||
|
||||
Dense and BM25 retrieval receive the same deterministic query text. It preserves the
|
||||
original question and appends nonempty concepts, tables and columns in a fixed order;
|
||||
question and context receive Unicode NFC, newline canonicalization and outer trimming.
|
||||
Context values are then deduplicated exactly and sorted, without lowercasing. Case,
|
||||
punctuation and internal whitespace remain intact. This avoids accidental ranking
|
||||
changes caused only by metadata ordering, preserves quoted PostgreSQL identifiers and
|
||||
keeps the two retrieval branches directly comparable.
|
||||
|
||||
Evidence fragmentation follows semantic headings, typed fields and paragraph
|
||||
boundaries. Formulas, value/meaning pairs, mappings, rules and URLs remain atomic. An
|
||||
atomic element larger than the existing `max_chunk_chars` limit creates the blocking
|
||||
`atomic_content_too_large` Review item rather than being split mechanically. The limit
|
||||
applies to the complete rendered text, defaults to 4,000 characters and is not
|
||||
duplicated by an Evidence-specific setting.
|
||||
|
||||
An L0 contract test starts the exact Qdrant image referenced by `compose.yaml` and
|
||||
proves Italian server-side `qdrant/bm25` ingestion and search in a temporary collection.
|
||||
There is no FastEmbed or dense fallback for Evidence when that capability is absent.
|
||||
|
||||
The versioned evaluation set contains lexical, semantic and mixed queries. Its report
|
||||
shows dense-only, BM25-only and fused ranks for expected Evidence. Publication remains
|
||||
governed only by the simple fused top-10 rule; branch ranks and hit@5 are diagnostic.
|
||||
Reference in New Issue
Block a user