docs(evidence): finalize ticketed restructuring specification

This commit is contained in:
2026-08-24 17:11:37 +02:00
parent d970e10264
commit 5c6228f8c2
11 changed files with 796 additions and 167 deletions
+288 -85
View File
@@ -1,6 +1,9 @@
# Evidence Restructuring Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
> **Execution:** GitHub issue #35 is the approved parent specification. Issues #36–#47
> are the executable tracer-bullet tickets; implement one unblocked ticket at a time
> with `/implement`. This document remains the detailed technical reference and must
> not be executed as a second, parallel work queue.
**Goal:** Build a Git-reviewed, typed Evidence authoring pipeline and publish its approved output to the existing revision-scoped Qdrant lifecycle with dense+BM25 hybrid retrieval.
@@ -37,13 +40,16 @@
Cover:
- all eight `kind` values;
- all five `purpose` values;
- all four `purpose` values;
- immutable provenance;
- strict unknown-field rejection;
- namespaced stable IDs;
- kind-independent stable IDs;
- SHA-256 syntax;
- directory/`kind` agreement;
- parsing and dumping Markdown with YAML frontmatter.
- review items containing a stable `code`, human `message` and optional `field`, with no
status or history fields;
- IDs matching `evidence:<slug>` and remaining independent from `kind`.
Start with:
@@ -91,7 +97,7 @@ EvidenceKind = Literal[
"normalization", "formula", "reference",
]
EvidencePurpose = Literal[
"disambiguation", "rewriting", "schema_linking", "sql_generation", "memory",
"disambiguation", "rewriting", "schema_linking", "sql_generation",
]
class EvidenceScope(StrictModel):
@@ -102,12 +108,18 @@ class EvidenceScope(StrictModel):
class EvidenceProvenance(StrictModel):
source_file: str
source_sha256: str
supporting_excerpts: tuple[str, ...]
class FormulaPayload(StrictModel):
concept: str
columns: tuple[str, ...]
sql: str
class ReviewItem(StrictModel):
code: str
message: str
field: str | None = None
class ReferencePayload(StrictModel):
url: AnyHttpUrl
label: str
@@ -126,8 +138,10 @@ def dump_curated_markdown(value: CuratedEvidence) -> str: ...
def load_curated_tree(root: Path) -> list[CuratedEvidence]: ...
```
Validate formula SQL with `sqlglot` and validate `schema.table` /
Validate formula SQL with `sqlglot` as one PostgreSQL expression, rejecting complete
`SELECT`/`WITH` statements, DDL, DML and multiple statements. Validate `schema.table` /
`schema.table.column` identifiers without querying the DWH.
Require one to five supporting excerpts, each nonempty and at most 1,000 characters.
**Step 4: Export only the stable API**
@@ -167,7 +181,10 @@ git commit -m "feat(evidence): add typed curated evidence model"
Test that:
- `review_items != []` is valid as a draft but blocks publication;
- `orphans != []` is preserved but blocks publication;
- a missing or mismatched source hash blocks publication;
- every supporting excerpt is found after applying the same mechanical normalization
to the excerpt and its source;
- two units cannot share an ID;
- one unit cannot claim two source files;
- `curated/formula/x.md` must contain `kind: formula`;
@@ -202,7 +219,7 @@ orphans: []
```
Test deterministic key ordering, stable round-trip, unknown fields, duplicate unit IDs,
and orphan preservation.
orphan preservation and incompatible `pipeline_version` refusal.
**Step 3: Run and confirm RED**
@@ -226,7 +243,7 @@ def validate_workspace_evidence(workspace_root: Path) -> ValidationReport: ...
```
`ValidationReport.publishable` is true only when there are no errors and no unresolved
review items. Warnings alone do not block publication.
review items or orphaned units. Warnings alone do not block publication.
Do not use status fields such as `draft/reviewed` as an approval mechanism. The approved
Git revision is the publication boundary.
@@ -264,16 +281,24 @@ Add:
```python
class EvidenceRestructurer(Protocol):
def restructure(self, request: RestructureRequest) -> tuple[CuratedEvidence, ...]: ...
def restructure(self, request: RestructureRequest) -> tuple[RestructureCandidate, ...]: ...
class RestructureRequest(StrictModel):
source_file: str
source_sha256: str
normalized_text: str
previous_units: tuple[CuratedEvidence, ...] = ()
class RestructureCandidate(StrictModel):
existing_id: str | None = None
# The same title, kind, purposes, scope, provenance excerpts and typed payload
# needed to construct CuratedEvidence, but no model-assigned canonical ID.
```
The response is validated through the models from Task 1 before any write.
`existing_id`, when present, must belong to `previous_units`. The preparer rejects any
unknown ID and deterministically allocates `evidence:<slug>` plus a collision suffix for
every candidate without one. The response is converted to and validated through the
models from Task 1 before any write.
**Step 2: Write RED tests for the incremental rules**
@@ -283,11 +308,23 @@ Test:
- changed source: exactly one model call;
- new source: new stable IDs;
- removed source: old units become orphans and remain on disk;
- uniquely renamed source with the same hash: provenance changes and unit IDs remain;
- existing source no longer supporting a prior unit: retain it with
`source_no_longer_supports_unit` and block publication;
- kind-only reclassification: preserve the unit ID;
- semantic split: allocate new IDs for the new independent units;
- a semantic split retains the previous unit as a blocking retirement candidate until
the curator explicitly retires it;
- one source may produce several units;
- every returned unit contains one to five exact supporting excerpts found in its
normalized source;
- no returned unit may cite another source;
- prior curated units are included in the request;
- a dirty `evidence/curated/` or `evidence/manifest.yaml` fails before the model call;
- committed human edits remain recoverable in Git and every proposed change is exposed
in the working-tree diff;
- all output writes are staged and atomically replaced only after complete validation.
- one invalid source leaves the complete batch and manifest unchanged;
Inject the Git status runner and filesystem writer in tests; do not require a real Git
repository for every unit test.
@@ -323,14 +360,23 @@ state:
- use only facts present in `normalized_text`;
- preserve prior reviewed wording where it remains supported;
- retain unsupported prior units with a `source_no_longer_supports_unit` review item;
- never merge sources;
- emit `review_items` for uncertainty;
- copy short exact supporting excerpts from `normalized_text`;
- use only a supplied `existing_id` or omit it for deterministic allocation;
- emit exactly the schema-versioned JSON object and no Markdown fence.
Do not retry automatically after a timeout, nonzero exit, malformed JSON or invalid
response. Return a stable error code and the affected source so the curator can rerun
the command explicitly.
**Step 5: Implement `prepare_workspace_evidence`**
Return a bounded report with `changed`, `unchanged`, `created`, `orphaned`, `findings`,
and `model_calls`. Keep the model implementation behind `EvidenceRestructurer`.
and `model_calls`. Keep the model implementation behind `EvidenceRestructurer`. Build
the complete batch in a private staging directory and replace curated files plus the
manifest only after every candidate validates; never expose partial success.
**Step 6: Run tests**
@@ -364,13 +410,19 @@ git commit -m "feat(evidence): prepare curated evidence incrementally"
Cover:
```text
tht evidence prepare <workspace-root> [--json]
tht evidence prepare <workspace-root> [--upgrade] [--json]
tht evidence validate <workspace-root> [--json]
tht evidence resolve <workspace-root> <evidence-id> --retire [--json]
tht evidence resolve <workspace-root> <evidence-id> --source <source-path> [--json]
```
Require an existing canonical Git worktree root. Reject unknown flags, symlinks, a
workspace path outside the Git root, duplicate options and dirty curated state. Ensure
`--json` writes pristine JSON to stdout.
workspace path outside the Git root, duplicate options and dirty curated state. A
pipeline-version mismatch must fail without writes unless `--upgrade` is explicit;
`--upgrade` reprocesses every source. Ensure `--json` writes pristine JSON to stdout.
For `resolve`, require exactly one of `--retire` and `--source`, an existing Evidence ID,
an existing in-root source path for relinking and a clean worktree. Test atomic updates
to the curated file and manifest, with no commit or publication side effect.
**Step 2: Run and confirm RED**
@@ -393,6 +445,15 @@ def prepare_cmd(workspace_root: Path, json_output: bool = False) -> None: ...
@evidence_app.command("validate")
def validate_cmd(workspace_root: Path, json_output: bool = False) -> None: ...
@evidence_app.command("resolve")
def resolve_cmd(
workspace_root: Path,
evidence_id: str,
retire: bool = False,
source: Path | None = None,
json_output: bool = False,
) -> None: ...
```
Register it in `harness/tht/cli/__init__.py`. Keep this separate from the existing
@@ -444,7 +505,12 @@ Test that:
- a short formula remains one fragment;
- a domain document splits only at semantic section boundaries;
- enum entries are not split in the middle of a value/meaning pair;
- fixed-size fallback is used only for a single oversized section;
- formulas, value/meaning pairs, mappings, rules and URLs are never split;
- the complete rendered fragment, including labels and textual metadata, uses the
existing `max_chunk_chars` setting whose default is 4,000 characters;
- an atomic element over `max_chunk_chars` produces the blocking
`atomic_content_too_large` review item instead of fixed-size chunks, with no second
size setting;
- all fragments carry `evidence_id`, `evidence_kind`, `purposes`, scope, language and
provenance;
- fragment IDs and ordinals are deterministic;
@@ -492,12 +558,14 @@ git add harness/tht/evidence/corpus harness/tht/evidence/preprocessing.py \
git commit -m "feat(evidence): build semantic fragments from typed units"
```
### Task 6: Upgrade the Qdrant collection contract to named dense and BM25 vectors
### Task 6: Add BM25 to the existing Qdrant collection without rebuilding it
**Files:**
- Modify: `backend/src/workspaces/qdrant-collection.ts`
- Modify: `backend/src/workspaces/evidence/preprocessing.ts`
- Modify: `backend/test/qdrant-collection.test.ts`
- Modify: `backend/test/workspaces/evidence/preprocessing.test.ts`
- Modify: `harness/tht/adapters/vector/qdrant.py`
- Modify: `harness/tests/test_qdrant_vector_store.py`
- Modify: `docs/contracts/workspace-preprocessing-cli.md`
@@ -508,18 +576,24 @@ The required vector contract is:
```json
{
"vectors": {
"dense": {"size": 1024, "distance": "Cosine"}
},
"vectors": {"size": 1024, "distance": "Cosine"},
"sparse_vectors": {
"bm25": {"modifier": "idf"}
}
}
```
Test that unnamed dense-only, wrong named dimension/distance, missing BM25, or wrong BM25
modifier are incompatible. Self-heal may add missing payload indexes, but must not
silently convert an incompatible vector configuration.
Keep the current unnamed dense vector. Test three distinct states:
- compatible: unnamed dense is correct and `bm25` has `modifier: idf`;
- Evidence-upgradeable: unnamed dense is correct and `bm25` is absent;
- incompatible: dense dimension/distance is wrong, dense is unexpectedly named, or
`bm25` exists with a different configuration.
Only the Evidence maintenance path may turn the upgradeable state into compatible by
adding `bm25`. Session admission and searches remain read-only. Missing payload indexes
may still use the existing additive reconciliation; no path may delete or rename a
vector.
Add keyword indexes only for fields used by filters:
@@ -536,12 +610,19 @@ cd backend
npx vitest run test/qdrant-collection.test.ts
```
Expected: current unnamed-vector expectations fail.
Expected: the new additive BM25 expectations fail.
**Step 3: Implement strict named-vector reconciliation**
**Step 3: Implement additive BM25 reconciliation**
Update `createCollection`, `vectorCompatibility` and payload-index reconciliation.
Retain `require_existing` behavior: validation never performs an incompatible migration.
Keep base `vectorCompatibility` concerned with the unnamed dense contract used by
Schema and Memory. Add an Evidence-specific compatibility result that also classifies
`bm25` as compatible, upgradeable or incompatible. Update `createCollection` and the
Evidence preprocessing preflight: a new collection is created with unnamed dense plus
`bm25`; on an existing upgradeable collection, `workspace preprocess evidence` calls
Qdrant's additive vector-schema endpoint to create only `bm25` with IDF, then verifies
the result before upload. Retain `require_existing` behavior for ordinary runtime
validation: it never mutates. An incompatible state fails without changes. Missing
BM25 does not make Schema or Memory unavailable; only Evidence reports `unavailable`.
**Step 4: Update the Python adapter's collection validation**
@@ -553,7 +634,7 @@ not share implementation code.
```bash
cd backend
npx vitest run test/qdrant-collection.test.ts
npx vitest run test/qdrant-collection.test.ts test/workspaces/evidence/preprocessing.test.ts
npx tsc --noEmit -p .
cd ../harness
.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
@@ -561,25 +642,22 @@ cd ../harness
Expected: PASS.
**Step 6: Document the required guarded rebuild**
**Step 6: Document the additive maintenance behavior**
Update the CLI contract to say that the vector-shape change is incompatible and must be
applied with the existing exact-name guarded command:
```text
tht --installation <absolute>/thothii-installation.yaml workspace vector rebuild
--workspace <id> --collection <name> --confirm <name> --destroy
```
No automatic deletion is allowed.
Update the CLI contract to state that `workspace preprocess evidence` may add the
missing `bm25` definition but may not delete or rename vectors. The existing destructive
`workspace vector rebuild` command remains available for unrelated operator recovery
and is not used by this migration.
**Step 7: Commit**
```bash
git add backend/src/workspaces/qdrant-collection.ts \
backend/test/qdrant-collection.test.ts harness/tht/adapters/vector/qdrant.py \
backend/src/workspaces/evidence/preprocessing.ts \
backend/test/qdrant-collection.test.ts backend/test/workspaces/evidence/preprocessing.test.ts \
harness/tht/adapters/vector/qdrant.py \
harness/tests/test_qdrant_vector_store.py docs/contracts/workspace-preprocessing-cli.md
git commit -m "feat(evidence): require dense and bm25 qdrant vectors"
git commit -m "feat(evidence): add qdrant bm25 vector in place"
```
### Task 7: Add server-side BM25 ingestion and hybrid Query API retrieval
@@ -593,6 +671,7 @@ git commit -m "feat(evidence): require dense and bm25 qdrant vectors"
- Modify: `harness/tests/test_vector_port_contract.py`
- Modify: `harness/tests/test_qdrant_vector_store.py`
- Modify: `harness/tests/test_corpus_pipeline.py`
- Create: `harness/tests/l0/test_qdrant_bm25_inference.py`
**Step 1: Write RED port tests**
@@ -617,7 +696,7 @@ For Evidence upsert, assert:
```json
"vector": {
"dense": [0.1, 0.2],
"": [0.1, 0.2],
"bm25": {
"text": "...",
"model": "qdrant/bm25",
@@ -631,7 +710,7 @@ For hybrid search, assert two filtered prefetches and default RRF:
```json
{
"prefetch": [
{"query": [0.1, 0.2], "using": "dense", "limit": 20, "filter": {}},
{"query": [0.1, 0.2], "limit": 20, "filter": {}},
{
"query": {
"text": "fascia pediatrica",
@@ -652,42 +731,54 @@ For hybrid search, assert two filtered prefetches and default RRF:
Both prefetch filters must include workspace, revision, active generation and record
kind. Do not add hand-tuned weights.
**Step 3: Implement named dense writes for all semantic records**
**Step 3: Prove BM25 on the actual local Qdrant image**
All existing schema/memory/Evidence records use the `dense` vector name after the
collection rebuild. Only records with `sparse_text` receive `bm25`.
Add an `l0` test that reads the Qdrant image reference from the root `compose.yaml`,
starts that exact image with testcontainers, creates a uniquely named temporary
collection, adds `bm25` with IDF, indexes two Italian texts through server-side
`qdrant/bm25`, retrieves the expected text and deletes the collection. The test must
not use FastEmbed or accept a dense fallback. It fails if the Compose reference and the
tested image diverge.
**Step 4: Implement BM25 Evidence ingestion**
**Step 4: Preserve dense writes for all existing semantic records**
Schema, Memory and solved-question records keep their current unnamed dense writes.
Evidence records use the empty default-vector name plus `bm25` in the mixed-vector
upsert shape. Only records with `sparse_text` receive `bm25`; no full reindex of Schema
or Memory is performed.
**Step 5: Implement BM25 Evidence ingestion**
Populate `sparse_text` and map workspace language `it` to Qdrant's `italian`. Reject an
unsupported language before uploading the generation. Use the same options at ingest
and query time.
**Step 5: Implement hybrid search with dense fallback only for non-Evidence callers**
**Step 6: Implement hybrid search with dense fallback only for non-Evidence callers**
An Evidence hybrid request must fail as unavailable if the configured Qdrant version or
collection contract does not support BM25. It must not silently claim to have run hybrid
search. Existing non-Evidence dense requests continue to work.
**Step 6: Run focused tests**
**Step 7: Run focused tests**
```bash
cd harness
.venv/bin/pytest tests/test_vector_port_contract.py tests/test_qdrant_vector_store.py \
tests/test_corpus_pipeline.py tests/test_search_pack.py -q
.venv/bin/pytest tests/l0/test_qdrant_bm25_inference.py -m l0 -q
.venv/bin/ruff check tht/ports/vector.py tht/adapters/vector/qdrant.py \
tht/evidence/corpus/pipeline.py
```
Expected: PASS.
**Step 7: Commit**
**Step 8: Commit**
```bash
git add harness/tht/ports/vector.py harness/tht/adapters/vector/qdrant.py \
harness/tht/vectorstore/records.py harness/tht/evidence/corpus/pipeline.py \
harness/tests/test_vector_port_contract.py harness/tests/test_qdrant_vector_store.py \
harness/tests/test_corpus_pipeline.py
harness/tests/test_corpus_pipeline.py harness/tests/l0/test_qdrant_bm25_inference.py
git commit -m "feat(evidence): add qdrant bm25 hybrid retrieval"
```
@@ -697,8 +788,12 @@ git commit -m "feat(evidence): add qdrant bm25 hybrid retrieval"
- Modify: `harness/tht/evidence/search.py`
- Modify: `harness/tht/evidence/__init__.py`
- Modify: `harness/tht/evidence/session.py`
- Modify: `harness/tht/cli/search_cmd.py`
- Modify: `harness/tht/session/filesystem_repository.py`
- Modify: `harness/tht/session/postgres_repository.py`
- Modify: `harness/tests/test_evidence_facade_contract.py`
- Create: `harness/tests/test_evidence_session_receipts.py`
- Modify: `harness/tests/test_search_pack.py`
- Create: `harness/.pi/skills/tht-sessione/modules/evidence/runtime-search.md`
- Modify: `harness/.pi/skills/tht-sessione/projection.md.tmpl`
@@ -716,6 +811,9 @@ class EvidenceSearchContext(BaseModel):
tables: tuple[str, ...] = ()
columns: tuple[str, ...] = ()
required_kinds: tuple[EvidenceKind, ...] = ()
required_concepts: tuple[str, ...] = ()
required_tables: tuple[str, ...] = ()
required_columns: tuple[str, ...] = ()
def search_evidence(
query: str,
@@ -725,50 +823,105 @@ def search_evidence(
searcher: ActiveEvidenceSearcher,
embedder: EvidenceQueryEmbedder,
top_n: int = 10,
) -> list[EvidenceResult]: ...
) -> EvidenceSearchOutcome: ...
```
Test hard filters for workspace/revision/generation and explicit `required_kinds`. Test
that purpose/kind/scope preferences are deterministic tie-breakers when they were not
requested as hard filters.
Test hard filters for workspace, revision, generation, purpose and every explicit
`required_*` constraint. Concepts, tables and columns without the `required_` prefix
enrich the dense and BM25 query text but do not create payload filters. Do not add
implicit kind or scope bonuses in v1.
Render that enrichment once, identically for both retrieval branches, in this fixed
order:
```text
Domanda: <original query>
Concetti: <normalized, deduplicated, sorted values>
Tabelle: <normalized, deduplicated, sorted values>
Colonne: <normalized, deduplicated, sorted values>
```
Omit empty lines and do not rewrite the original question. Test that permutations and
duplicates in context produce the same rendered text and that the dense embedder and
BM25 request receive exactly that same value.
The renderer normalizes the question to Unicode NFC, converts CRLF and CR to `\n`,
strips only leading and trailing whitespace and rejects an empty result. It preserves
case, punctuation and internal whitespace. Apply NFC and `strip` to context values,
remove empty strings and exact duplicates, then sort by Unicode value. Do not call
`lower()` or `casefold()` because quoted PostgreSQL identifiers can be case-sensitive.
`EvidenceSearchOutcome` has an `available` state with a generation and zero or more
results, and an `unavailable` state with a stable error code, a bounded message and no
results. `memory` is not an `EvidencePurpose`: the `memory` stage remains owned by the
Memory Module.
**Step 2: Test fragment grouping**
Two returned fragments with the same `evidence_id` must become one `EvidenceResult`,
with the best score, ordered matching excerpts, canonical citation and no duplicate
unit.
with the best score, ordered matching excerpts, canonical citation, a reference for
resolving the full document and no duplicate unit. Do not insert the full document into
the search pack automatically.
**Step 3: Preserve fail-closed graceful degradation**
**Step 3: Distinguish an empty search from technical unavailability**
Test absent ACTIVE corpus, revision mismatch, unavailable Qdrant and malformed payload.
All return no Evidence plus a bounded warning through the existing search-pack contract;
none uses stale rows.
Test a successful search with zero matches separately from absent ACTIVE corpus,
revision mismatch, unavailable Qdrant and malformed payload. The first returns
`available` with an empty result list and may continue. Every technical case returns
`unavailable`, blocks the calling stage until retry, and never uses stale rows or a
fallback purpose.
**Step 4: Implement the facade and CLI mapping**
Keep Qdrant syntax inside `tht.evidence`. `search_cmd.py` translates command inputs to
the facade and renders results; it must not duplicate ranking logic.
**Step 5: Extract Evidence instructions into a module fragment**
**Step 5: Integrate the contributor with semantic stages**
Move the common F1/F3/F4 Evidence rules from the projection template into
Move the shared Evidence rules from the projection template into
`modules/evidence/runtime-search.md`. The fragment must state:
- candidates are not truth;
- pass the phase-appropriate purpose;
- show provenance;
- formulas are `kind=formula`, not a separate search store;
- absence of Evidence is visible but does not stop the whole session.
- an available empty result is visible and does not block the stage;
- an unavailable outcome blocks the stage and is retriable;
- Evidence never writes decisions, canonical artifacts or workflow state.
Map semantic stages, independently of their display codes:
```text
clarification -> disambiguation
rewriting -> rewriting
schema_linking -> schema_linking
cte -> sql_generation
final_sql -> sql_generation
```
Do not invoke Evidence from `memory` or `synthesis`. Run an independent search in every
mapped stage; the `final_sql` query includes the approved CTE plan.
**Step 6: Persist the minimal Evidence receipt**
Use the existing session artifact repositories to maintain one
`evidence_receipts.json` artifact. For every available search, append or replace the
receipt identified by semantic stage with exactly `stage`, `purpose`,
`vector_generation` and ordered `evidence_ids`. Do not copy excerpts or complete
Evidence text. The workflow integration owns the write; `tht.evidence` only constructs
and returns the typed receipt. Test filesystem and PostgreSQL repositories, resume and
retry replacement.
Register the fragment in the static `FRAGMENT_ORDER` and regenerate.
**Step 6: Run tests**
**Step 7: Run tests**
```bash
cd harness
python -m tht.pi_skill_projection --write
python -m tht.pi_skill_projection --check
.venv/bin/pytest tests/test_evidence_facade_contract.py tests/test_search_pack.py \
tests/test_evidence_session_receipts.py \
tests/test_pi_skill_projection.py -q
```
@@ -878,28 +1031,39 @@ schema_version: 1
queries:
- id: pediatric-formula
query: Come distinguo i pazienti pediatrici?
profile: semantic
purpose: sql_generation
expected:
- formula:fascia-pediatrica
- evidence:fascia-pediatrica
```
Require unique query IDs, nonempty expected IDs and only public purpose values.
Require unique query IDs, nonempty expected IDs, one of `lexical`, `semantic` or
`mixed` for every profile, only public purpose values and at least one query of every
profile in the complete file.
**Step 2: Write RED metric tests**
Compute `hit_at_5`, `hit_at_10`, missing expected IDs, empty-result queries and counts by
expected `kind`. Do not add nDCG, relevance grading or an evaluation database in v1.
expected `kind`. For each expected ID, also report its nullable dense-only, BM25-only
and fused rank. The evaluator runs both branches separately for diagnosis and the same
hybrid request used by runtime. The report passes only when every query finds at least
one expected ID in its first ten fused results; branch ranks and `hit_at_5` are
informative. Do not add nDCG, relevance grading or an evaluation database in v1.
**Step 3: Implement the evaluator and CLI**
Expose:
```text
tht evidence evaluate <workspace-root> -c <runtime-config> [--json]
tht evidence evaluate <workspace-root> -c <runtime-config>
[--generation <id>] [--json]
```
The report includes workspace revision, active vector generation and the fixed default
RRF configuration. It is read-only.
The report includes workspace revision, evaluated vector generation and the fixed
default RRF configuration. It is read-only. With no `--generation`, it evaluates the
active generation for monitoring. Before publication, the preprocessing pipeline calls
the same evaluator against its candidate generation and atomically activates it only
when the report passes. A failed candidate remains inactive.
**Step 4: Run tests**
@@ -957,7 +1121,7 @@ Document:
- Git publication boundary;
- no runtime writes;
- validation before indexing;
- dense+BM25 collection contract and guarded rebuild;
- unnamed-dense plus BM25 contract and additive Evidence upgrade;
- exact public operation names and JSON status additions, if any.
**Step 4: Run backend and harness contract gates**
@@ -1017,10 +1181,24 @@ and a SQL formula candidate. The acceptance runner must prove:
- ambiguity becomes `review_items`;
- no cross-source merge;
- unchanged rerun is a no-op;
- human edit is preserved;
- committed human content remains recoverable and proposed changes are visible in Git;
- dirty-tree refusal;
- validation blocks unresolved review;
- validated corpus indexes and searches hybrid;
- validation blocks unresolved review and orphaned units;
- pipeline-version mismatch refusal and explicit full-corpus `--upgrade`;
- validated corpus builds an inactive candidate and searches it hybrid;
- every evaluation query retrieves at least one expected ID in the first ten results
before the candidate becomes active;
- the evaluation set contains lexical, semantic and mixed cases and reports dense,
BM25 and fused ranks separately;
- dense and BM25 receive the same deterministic query text;
- no atomic formula, enum pair, mapping, rule or URL is split into fixed-size chunks;
- the complete rendered fragment stays within the existing 4,000-character
`max_chunk_chars` default and no parallel size option exists;
- the L0 probe passes against the Qdrant image referenced by `compose.yaml` without
FastEmbed or fallback;
- query normalization is limited to NFC, newline canonicalization and outer trimming,
preserving internal whitespace, punctuation and case-sensitive identifiers;
- Formula Evidence accepts a PostgreSQL expression and rejects a complete query;
- Qdrant failure remains fail-closed.
Use a fake restructurer for hermetic CI. The real Pi call is a separate manual check.
@@ -1054,7 +1232,7 @@ Provide:
- proposed PSD branch name;
- exact list of 36 source files to move;
- rollback command based on the pre-migration PSD commit;
- notice that Qdrant rebuild is destructive but scoped by exact collection-name guards.
- before/after Schema and Memory counts plus sample IDs used to prove non-regression.
Do not infer approval from prior design acceptance.
@@ -1064,7 +1242,7 @@ Use `tht evidence prepare`, review the Git diff, resolve all review items manual
`tht evidence validate`, and create the approximately twenty evaluation queries. Do not
auto-merge or auto-push unless separately requested.
**Step 5: Activate and rebuild with the existing guarded operator path**
**Step 5: Activate and perform the additive BM25 upgrade**
After the PSD merge/pull and activation, inspect first:
@@ -1073,20 +1251,24 @@ tht --installation <absolute>/thothii-installation.yaml workspace vector inspect
--workspace psd-clinical --json
```
Then use the exact descriptor-owned name in the guarded rebuild command. Run
`workspace preprocess evidence`, `tht evidence evaluate`, and the manual F1/F3/F4
walkthrough.
Before preprocessing, record counts and representative IDs for `schema_table`,
`schema_column`, `memory` and `solved_question`. Run `workspace preprocess evidence`:
it adds `bm25` when absent, builds the candidate, evaluates that exact generation and
publishes it only if the report passes. Repeat the counts, verify the representative IDs
and run dense smoke searches for Schema and Memory before completing the manual
walkthrough of `clarification`, `rewriting`, `schema_linking`, `cte` and `final_sql`.
**Step 6: Record acceptance and commit ThothII documentation**
`docs/testing/evidence-restructuring-manual.md` must record separate outcomes for:
- authoring and Git review;
- collection rebuild;
- additive BM25 schema upgrade;
- Schema and Memory before/after non-regression evidence;
- preprocessing generation publication;
- hybrid retrieval evaluation;
- formula retrieval;
- graceful degradation;
- distinzione fra risultato vuoto e indisponibilità bloccante;
- complete session behavior.
Update `PROJECT_STATE.md` only with observed results and immutable commit/run IDs.
@@ -1105,17 +1287,38 @@ Before claiming completion, verify:
- [ ] `CONTEXT.md` and both Evidence plan documents use the same terminology.
- [ ] All eight Evidence kinds have type-specific positive and negative tests.
- [ ] Pi is called once per changed source, with no tools and no saved session.
- [ ] `prepare` refuses dirty curated state and never deletes orphans.
- [ ] `validate` blocks unresolved review items.
- [ ] `prepare` refuses dirty curated state, exposes every proposal in Git and never
deletes orphans.
- [ ] `validate` blocks unresolved review items and orphaned units.
- [ ] Source renames and kind-only reclassifications preserve unit IDs; semantic splits
receive new IDs.
- [ ] IDs use `evidence:<slug>`, remain independent from kind and are not recomputed
after their initial assignment.
- [ ] Formula Evidence accepts one PostgreSQL expression and rejects full queries.
- [ ] A pipeline-version mismatch performs no writes without explicit `--upgrade`.
- [ ] Runtime reads only `curated/**/*.md` from the pinned Git revision.
- [ ] The TypeScript and Python Qdrant compatibility checks agree.
- [ ] Qdrant collection uses named `dense` plus `bm25` with IDF.
- [ ] Qdrant retains the unnamed dense vector and adds only `bm25` with IDF.
- [ ] Evidence preprocessing adds a missing `bm25` definition but never deletes,
renames or destructively rebuilds collection vectors.
- [ ] Schema and Memory counts, sample IDs and dense searches remain unchanged across
the additive upgrade.
- [ ] Italian BM25 options are identical during ingest and query.
- [ ] Hybrid search uses two prefetches and default RRF.
- [ ] Hard filters always include workspace, revision and active generation.
- [ ] Hard filters always include workspace, revision, active generation and purpose;
only explicit `required_*` context values add further filters.
- [ ] Formula runtime lookup uses typed Evidence; session proposals remain non-published.
- [ ] Fragment hits are grouped into complete Evidence Units.
- [ ] Evaluation reports hit@5 and hit@10 against a versioned fixture.
- [ ] Fragment hits are grouped into one Evidence Result with excerpts and a reference;
full documents are loaded only on demand.
- [ ] Evidence is queried independently from the five mapped semantic stages and is not
invoked from `memory` or `synthesis`.
- [ ] An available empty result can continue; technical unavailability blocks the
current stage without stale-generation or purpose fallback.
- [ ] Each available stage search persists only its minimal Evidence receipt.
- [ ] Evaluation reports hit@5 and hit@10 against a versioned fixture and passes only
when every query has at least one expected result in the first ten.
- [ ] Evaluation runs against the candidate generation before atomic activation; a
failed candidate remains invisible to sessions.
- [ ] Existing corpus rollback, compensation, retention and resume tests still pass.
- [ ] Backend Vitest and TypeScript gates pass.
- [ ] Harness pytest and Ruff gates pass.