docs(evidence): finalize ticketed restructuring specification

This commit is contained in:
2026-08-24 17:11:37 +02:00
parent d970e10264
commit 5c6228f8c2
11 changed files with 796 additions and 167 deletions
@@ -0,0 +1,7 @@
# Use workspace activation as the Evidence publication boundary
Curated Evidence becomes Published Evidence only when it is valid, belongs to the
active workspace revision, and belongs to the atomically published Evidence generation.
Human Git review remains required before activation, but its approval is not duplicated
as mutable state in the Evidence manifest; this keeps Git and the workspace registry as
the existing sources of truth instead of introducing a second approval mechanism.
@@ -0,0 +1,8 @@
# Keep Evidence identifiers independent from kind
An Evidence Unit keeps the same identifier when its kind is corrected or its Source
Evidence is unambiguously renamed; kind remains a separate validated field. This avoids
breaking citations and evaluation fixtures for a classification change, while genuine
semantic splits receive new identifiers so independent units never share an identity.
Identifiers use `evidence:<slug>`, are assigned once, persisted in the manifest and are
never recomputed automatically from mutable titles, paths or content hashes.
@@ -0,0 +1,7 @@
# Evaluate candidate Evidence before activation
Evidence preprocessing builds a candidate vector generation and runs the workspace's
versioned evaluation set against that exact generation before atomically activating it.
The workspace revision may temporarily have no matching active Evidence during this
maintenance window, which is accepted as fail-closed degradation instead of introducing
a distributed transaction across Git, the workspace registry and Qdrant.
@@ -0,0 +1,6 @@
# Treat Formula Evidence as a composable SQL expression
Formula Evidence contains one validated PostgreSQL expression with declared input
columns, not a complete query or statement. This makes formulas safely composable by
SQL generation; complete documented queries remain Example Evidence, and incompatible
legacy formulas require human review during migration.
@@ -0,0 +1,13 @@
# Evidence contributes to semantic workflow stages
The Evidence Module is a contributor to the existing semantic stages, not a visible
workflow stage. `clarification`, `rewriting`, `schema_linking`, `cte` and `final_sql`
query it with their corresponding purpose; `memory` remains owned by the Memory Module
and `synthesis` performs no Evidence search. Each mapped stage searches independently
and the workflow records only a minimal receipt containing stage, purpose, vector
generation and returned Evidence IDs.
A successful search may return no matches and does not block the stage. Technical
unavailability is a distinct typed outcome that blocks the calling stage until retry,
without stale-generation or purpose fallback. This favors an explicit temporary stop
over silently treating a broken Evidence dependency as absence of domain knowledge.
@@ -0,0 +1,12 @@
# Ground and atomically apply Evidence preparation
Every proposed Evidence Unit carries one to five short supporting excerpts that can be
found in its normalized Source Evidence. The model may reuse only identifiers supplied
for previous units; deterministic application code assigns all new canonical IDs. This
makes provenance and identity mechanically checkable without pretending that automated
validation can replace human semantic review.
Preparation validates the complete changed batch in a temporary area and applies
curated files plus the manifest atomically. A model timeout or invalid response is not
retried automatically and leaves the worktree unchanged. Explicit `evidence resolve`
actions retire or relink units while preserving an ordinary, recoverable Git diff.
@@ -0,0 +1,13 @@
# Add BM25 without rebuilding the shared Qdrant collection
The workspace keeps its existing unnamed dense vector and adds only the sparse `bm25`
vector with IDF through Qdrant's additive vector-schema operation. Only Evidence points
are repopulated with both default dense and BM25 values. Schema, Memory and solved
questions retain their current dense points and are verified before and after the
upgrade.
This replaces the planned destructive conversion to a named `dense` vector. If BM25 is
missing, Evidence preprocessing may add it and verify the resulting schema; session
runtime remains read-only. An incompatible existing BM25 definition fails without
mutation. A candidate failure leaves the additive schema in place while Evidence stays
unavailable, avoiding data loss in the other workflow modules.
@@ -0,0 +1,24 @@
# Make hybrid Evidence retrieval deterministic and diagnosable
Dense and BM25 retrieval receive the same deterministic query text. It preserves the
original question and appends nonempty concepts, tables and columns in a fixed order;
question and context receive Unicode NFC, newline canonicalization and outer trimming.
Context values are then deduplicated exactly and sorted, without lowercasing. Case,
punctuation and internal whitespace remain intact. This avoids accidental ranking
changes caused only by metadata ordering, preserves quoted PostgreSQL identifiers and
keeps the two retrieval branches directly comparable.
Evidence fragmentation follows semantic headings, typed fields and paragraph
boundaries. Formulas, value/meaning pairs, mappings, rules and URLs remain atomic. An
atomic element larger than the existing `max_chunk_chars` limit creates the blocking
`atomic_content_too_large` Review item rather than being split mechanically. The limit
applies to the complete rendered text, defaults to 4,000 characters and is not
duplicated by an Evidence-specific setting.
An L0 contract test starts the exact Qdrant image referenced by `compose.yaml` and
proves Italian server-side `qdrant/bm25` ingestion and search in a temporary collection.
There is no FastEmbed or dense fallback for Evidence when that capability is absent.
The versioned evaluation set contains lexical, semantic and mixed queries. Its report
shows dense-only, BM25-only and fused ranks for expected Evidence. Publication remains
governed only by the simple fused top-10 rule; branch ranks and hit@5 are diagnostic.