docs(evidence): finalize ticketed restructuring specification
This commit is contained in:
@@ -1,6 +1,9 @@
|
||||
# Evidence Restructuring Implementation Plan
|
||||
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
|
||||
> **Execution:** GitHub issue #35 is the approved parent specification. Issues #36–#47
|
||||
> are the executable tracer-bullet tickets; implement one unblocked ticket at a time
|
||||
> with `/implement`. This document remains the detailed technical reference and must
|
||||
> not be executed as a second, parallel work queue.
|
||||
|
||||
**Goal:** Build a Git-reviewed, typed Evidence authoring pipeline and publish its approved output to the existing revision-scoped Qdrant lifecycle with dense+BM25 hybrid retrieval.
|
||||
|
||||
@@ -37,13 +40,16 @@
|
||||
Cover:
|
||||
|
||||
- all eight `kind` values;
|
||||
- all five `purpose` values;
|
||||
- all four `purpose` values;
|
||||
- immutable provenance;
|
||||
- strict unknown-field rejection;
|
||||
- namespaced stable IDs;
|
||||
- kind-independent stable IDs;
|
||||
- SHA-256 syntax;
|
||||
- directory/`kind` agreement;
|
||||
- parsing and dumping Markdown with YAML frontmatter.
|
||||
- review items containing a stable `code`, human `message` and optional `field`, with no
|
||||
status or history fields;
|
||||
- IDs matching `evidence:<slug>` and remaining independent from `kind`.
|
||||
|
||||
Start with:
|
||||
|
||||
@@ -91,7 +97,7 @@ EvidenceKind = Literal[
|
||||
"normalization", "formula", "reference",
|
||||
]
|
||||
EvidencePurpose = Literal[
|
||||
"disambiguation", "rewriting", "schema_linking", "sql_generation", "memory",
|
||||
"disambiguation", "rewriting", "schema_linking", "sql_generation",
|
||||
]
|
||||
|
||||
class EvidenceScope(StrictModel):
|
||||
@@ -102,12 +108,18 @@ class EvidenceScope(StrictModel):
|
||||
class EvidenceProvenance(StrictModel):
|
||||
source_file: str
|
||||
source_sha256: str
|
||||
supporting_excerpts: tuple[str, ...]
|
||||
|
||||
class FormulaPayload(StrictModel):
|
||||
concept: str
|
||||
columns: tuple[str, ...]
|
||||
sql: str
|
||||
|
||||
class ReviewItem(StrictModel):
|
||||
code: str
|
||||
message: str
|
||||
field: str | None = None
|
||||
|
||||
class ReferencePayload(StrictModel):
|
||||
url: AnyHttpUrl
|
||||
label: str
|
||||
@@ -126,8 +138,10 @@ def dump_curated_markdown(value: CuratedEvidence) -> str: ...
|
||||
def load_curated_tree(root: Path) -> list[CuratedEvidence]: ...
|
||||
```
|
||||
|
||||
Validate formula SQL with `sqlglot` and validate `schema.table` /
|
||||
Validate formula SQL with `sqlglot` as one PostgreSQL expression, rejecting complete
|
||||
`SELECT`/`WITH` statements, DDL, DML and multiple statements. Validate `schema.table` /
|
||||
`schema.table.column` identifiers without querying the DWH.
|
||||
Require one to five supporting excerpts, each nonempty and at most 1,000 characters.
|
||||
|
||||
**Step 4: Export only the stable API**
|
||||
|
||||
@@ -167,7 +181,10 @@ git commit -m "feat(evidence): add typed curated evidence model"
|
||||
Test that:
|
||||
|
||||
- `review_items != []` is valid as a draft but blocks publication;
|
||||
- `orphans != []` is preserved but blocks publication;
|
||||
- a missing or mismatched source hash blocks publication;
|
||||
- every supporting excerpt is found after applying the same mechanical normalization
|
||||
to the excerpt and its source;
|
||||
- two units cannot share an ID;
|
||||
- one unit cannot claim two source files;
|
||||
- `curated/formula/x.md` must contain `kind: formula`;
|
||||
@@ -202,7 +219,7 @@ orphans: []
|
||||
```
|
||||
|
||||
Test deterministic key ordering, stable round-trip, unknown fields, duplicate unit IDs,
|
||||
and orphan preservation.
|
||||
orphan preservation and incompatible `pipeline_version` refusal.
|
||||
|
||||
**Step 3: Run and confirm RED**
|
||||
|
||||
@@ -226,7 +243,7 @@ def validate_workspace_evidence(workspace_root: Path) -> ValidationReport: ...
|
||||
```
|
||||
|
||||
`ValidationReport.publishable` is true only when there are no errors and no unresolved
|
||||
review items. Warnings alone do not block publication.
|
||||
review items or orphaned units. Warnings alone do not block publication.
|
||||
|
||||
Do not use status fields such as `draft/reviewed` as an approval mechanism. The approved
|
||||
Git revision is the publication boundary.
|
||||
@@ -264,16 +281,24 @@ Add:
|
||||
|
||||
```python
|
||||
class EvidenceRestructurer(Protocol):
|
||||
def restructure(self, request: RestructureRequest) -> tuple[CuratedEvidence, ...]: ...
|
||||
def restructure(self, request: RestructureRequest) -> tuple[RestructureCandidate, ...]: ...
|
||||
|
||||
class RestructureRequest(StrictModel):
|
||||
source_file: str
|
||||
source_sha256: str
|
||||
normalized_text: str
|
||||
previous_units: tuple[CuratedEvidence, ...] = ()
|
||||
|
||||
class RestructureCandidate(StrictModel):
|
||||
existing_id: str | None = None
|
||||
# The same title, kind, purposes, scope, provenance excerpts and typed payload
|
||||
# needed to construct CuratedEvidence, but no model-assigned canonical ID.
|
||||
```
|
||||
|
||||
The response is validated through the models from Task 1 before any write.
|
||||
`existing_id`, when present, must belong to `previous_units`. The preparer rejects any
|
||||
unknown ID and deterministically allocates `evidence:<slug>` plus a collision suffix for
|
||||
every candidate without one. The response is converted to and validated through the
|
||||
models from Task 1 before any write.
|
||||
|
||||
**Step 2: Write RED tests for the incremental rules**
|
||||
|
||||
@@ -283,11 +308,23 @@ Test:
|
||||
- changed source: exactly one model call;
|
||||
- new source: new stable IDs;
|
||||
- removed source: old units become orphans and remain on disk;
|
||||
- uniquely renamed source with the same hash: provenance changes and unit IDs remain;
|
||||
- existing source no longer supporting a prior unit: retain it with
|
||||
`source_no_longer_supports_unit` and block publication;
|
||||
- kind-only reclassification: preserve the unit ID;
|
||||
- semantic split: allocate new IDs for the new independent units;
|
||||
- a semantic split retains the previous unit as a blocking retirement candidate until
|
||||
the curator explicitly retires it;
|
||||
- one source may produce several units;
|
||||
- every returned unit contains one to five exact supporting excerpts found in its
|
||||
normalized source;
|
||||
- no returned unit may cite another source;
|
||||
- prior curated units are included in the request;
|
||||
- a dirty `evidence/curated/` or `evidence/manifest.yaml` fails before the model call;
|
||||
- committed human edits remain recoverable in Git and every proposed change is exposed
|
||||
in the working-tree diff;
|
||||
- all output writes are staged and atomically replaced only after complete validation.
|
||||
- one invalid source leaves the complete batch and manifest unchanged;
|
||||
|
||||
Inject the Git status runner and filesystem writer in tests; do not require a real Git
|
||||
repository for every unit test.
|
||||
@@ -323,14 +360,23 @@ state:
|
||||
|
||||
- use only facts present in `normalized_text`;
|
||||
- preserve prior reviewed wording where it remains supported;
|
||||
- retain unsupported prior units with a `source_no_longer_supports_unit` review item;
|
||||
- never merge sources;
|
||||
- emit `review_items` for uncertainty;
|
||||
- copy short exact supporting excerpts from `normalized_text`;
|
||||
- use only a supplied `existing_id` or omit it for deterministic allocation;
|
||||
- emit exactly the schema-versioned JSON object and no Markdown fence.
|
||||
|
||||
Do not retry automatically after a timeout, nonzero exit, malformed JSON or invalid
|
||||
response. Return a stable error code and the affected source so the curator can rerun
|
||||
the command explicitly.
|
||||
|
||||
**Step 5: Implement `prepare_workspace_evidence`**
|
||||
|
||||
Return a bounded report with `changed`, `unchanged`, `created`, `orphaned`, `findings`,
|
||||
and `model_calls`. Keep the model implementation behind `EvidenceRestructurer`.
|
||||
and `model_calls`. Keep the model implementation behind `EvidenceRestructurer`. Build
|
||||
the complete batch in a private staging directory and replace curated files plus the
|
||||
manifest only after every candidate validates; never expose partial success.
|
||||
|
||||
**Step 6: Run tests**
|
||||
|
||||
@@ -364,13 +410,19 @@ git commit -m "feat(evidence): prepare curated evidence incrementally"
|
||||
Cover:
|
||||
|
||||
```text
|
||||
tht evidence prepare <workspace-root> [--json]
|
||||
tht evidence prepare <workspace-root> [--upgrade] [--json]
|
||||
tht evidence validate <workspace-root> [--json]
|
||||
tht evidence resolve <workspace-root> <evidence-id> --retire [--json]
|
||||
tht evidence resolve <workspace-root> <evidence-id> --source <source-path> [--json]
|
||||
```
|
||||
|
||||
Require an existing canonical Git worktree root. Reject unknown flags, symlinks, a
|
||||
workspace path outside the Git root, duplicate options and dirty curated state. Ensure
|
||||
`--json` writes pristine JSON to stdout.
|
||||
workspace path outside the Git root, duplicate options and dirty curated state. A
|
||||
pipeline-version mismatch must fail without writes unless `--upgrade` is explicit;
|
||||
`--upgrade` reprocesses every source. Ensure `--json` writes pristine JSON to stdout.
|
||||
For `resolve`, require exactly one of `--retire` and `--source`, an existing Evidence ID,
|
||||
an existing in-root source path for relinking and a clean worktree. Test atomic updates
|
||||
to the curated file and manifest, with no commit or publication side effect.
|
||||
|
||||
**Step 2: Run and confirm RED**
|
||||
|
||||
@@ -393,6 +445,15 @@ def prepare_cmd(workspace_root: Path, json_output: bool = False) -> None: ...
|
||||
|
||||
@evidence_app.command("validate")
|
||||
def validate_cmd(workspace_root: Path, json_output: bool = False) -> None: ...
|
||||
|
||||
@evidence_app.command("resolve")
|
||||
def resolve_cmd(
|
||||
workspace_root: Path,
|
||||
evidence_id: str,
|
||||
retire: bool = False,
|
||||
source: Path | None = None,
|
||||
json_output: bool = False,
|
||||
) -> None: ...
|
||||
```
|
||||
|
||||
Register it in `harness/tht/cli/__init__.py`. Keep this separate from the existing
|
||||
@@ -444,7 +505,12 @@ Test that:
|
||||
- a short formula remains one fragment;
|
||||
- a domain document splits only at semantic section boundaries;
|
||||
- enum entries are not split in the middle of a value/meaning pair;
|
||||
- fixed-size fallback is used only for a single oversized section;
|
||||
- formulas, value/meaning pairs, mappings, rules and URLs are never split;
|
||||
- the complete rendered fragment, including labels and textual metadata, uses the
|
||||
existing `max_chunk_chars` setting whose default is 4,000 characters;
|
||||
- an atomic element over `max_chunk_chars` produces the blocking
|
||||
`atomic_content_too_large` review item instead of fixed-size chunks, with no second
|
||||
size setting;
|
||||
- all fragments carry `evidence_id`, `evidence_kind`, `purposes`, scope, language and
|
||||
provenance;
|
||||
- fragment IDs and ordinals are deterministic;
|
||||
@@ -492,12 +558,14 @@ git add harness/tht/evidence/corpus harness/tht/evidence/preprocessing.py \
|
||||
git commit -m "feat(evidence): build semantic fragments from typed units"
|
||||
```
|
||||
|
||||
### Task 6: Upgrade the Qdrant collection contract to named dense and BM25 vectors
|
||||
### Task 6: Add BM25 to the existing Qdrant collection without rebuilding it
|
||||
|
||||
**Files:**
|
||||
|
||||
- Modify: `backend/src/workspaces/qdrant-collection.ts`
|
||||
- Modify: `backend/src/workspaces/evidence/preprocessing.ts`
|
||||
- Modify: `backend/test/qdrant-collection.test.ts`
|
||||
- Modify: `backend/test/workspaces/evidence/preprocessing.test.ts`
|
||||
- Modify: `harness/tht/adapters/vector/qdrant.py`
|
||||
- Modify: `harness/tests/test_qdrant_vector_store.py`
|
||||
- Modify: `docs/contracts/workspace-preprocessing-cli.md`
|
||||
@@ -508,18 +576,24 @@ The required vector contract is:
|
||||
|
||||
```json
|
||||
{
|
||||
"vectors": {
|
||||
"dense": {"size": 1024, "distance": "Cosine"}
|
||||
},
|
||||
"vectors": {"size": 1024, "distance": "Cosine"},
|
||||
"sparse_vectors": {
|
||||
"bm25": {"modifier": "idf"}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Test that unnamed dense-only, wrong named dimension/distance, missing BM25, or wrong BM25
|
||||
modifier are incompatible. Self-heal may add missing payload indexes, but must not
|
||||
silently convert an incompatible vector configuration.
|
||||
Keep the current unnamed dense vector. Test three distinct states:
|
||||
|
||||
- compatible: unnamed dense is correct and `bm25` has `modifier: idf`;
|
||||
- Evidence-upgradeable: unnamed dense is correct and `bm25` is absent;
|
||||
- incompatible: dense dimension/distance is wrong, dense is unexpectedly named, or
|
||||
`bm25` exists with a different configuration.
|
||||
|
||||
Only the Evidence maintenance path may turn the upgradeable state into compatible by
|
||||
adding `bm25`. Session admission and searches remain read-only. Missing payload indexes
|
||||
may still use the existing additive reconciliation; no path may delete or rename a
|
||||
vector.
|
||||
|
||||
Add keyword indexes only for fields used by filters:
|
||||
|
||||
@@ -536,12 +610,19 @@ cd backend
|
||||
npx vitest run test/qdrant-collection.test.ts
|
||||
```
|
||||
|
||||
Expected: current unnamed-vector expectations fail.
|
||||
Expected: the new additive BM25 expectations fail.
|
||||
|
||||
**Step 3: Implement strict named-vector reconciliation**
|
||||
**Step 3: Implement additive BM25 reconciliation**
|
||||
|
||||
Update `createCollection`, `vectorCompatibility` and payload-index reconciliation.
|
||||
Retain `require_existing` behavior: validation never performs an incompatible migration.
|
||||
Keep base `vectorCompatibility` concerned with the unnamed dense contract used by
|
||||
Schema and Memory. Add an Evidence-specific compatibility result that also classifies
|
||||
`bm25` as compatible, upgradeable or incompatible. Update `createCollection` and the
|
||||
Evidence preprocessing preflight: a new collection is created with unnamed dense plus
|
||||
`bm25`; on an existing upgradeable collection, `workspace preprocess evidence` calls
|
||||
Qdrant's additive vector-schema endpoint to create only `bm25` with IDF, then verifies
|
||||
the result before upload. Retain `require_existing` behavior for ordinary runtime
|
||||
validation: it never mutates. An incompatible state fails without changes. Missing
|
||||
BM25 does not make Schema or Memory unavailable; only Evidence reports `unavailable`.
|
||||
|
||||
**Step 4: Update the Python adapter's collection validation**
|
||||
|
||||
@@ -553,7 +634,7 @@ not share implementation code.
|
||||
|
||||
```bash
|
||||
cd backend
|
||||
npx vitest run test/qdrant-collection.test.ts
|
||||
npx vitest run test/qdrant-collection.test.ts test/workspaces/evidence/preprocessing.test.ts
|
||||
npx tsc --noEmit -p .
|
||||
cd ../harness
|
||||
.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
|
||||
@@ -561,25 +642,22 @@ cd ../harness
|
||||
|
||||
Expected: PASS.
|
||||
|
||||
**Step 6: Document the required guarded rebuild**
|
||||
**Step 6: Document the additive maintenance behavior**
|
||||
|
||||
Update the CLI contract to say that the vector-shape change is incompatible and must be
|
||||
applied with the existing exact-name guarded command:
|
||||
|
||||
```text
|
||||
tht --installation <absolute>/thothii-installation.yaml workspace vector rebuild
|
||||
--workspace <id> --collection <name> --confirm <name> --destroy
|
||||
```
|
||||
|
||||
No automatic deletion is allowed.
|
||||
Update the CLI contract to state that `workspace preprocess evidence` may add the
|
||||
missing `bm25` definition but may not delete or rename vectors. The existing destructive
|
||||
`workspace vector rebuild` command remains available for unrelated operator recovery
|
||||
and is not used by this migration.
|
||||
|
||||
**Step 7: Commit**
|
||||
|
||||
```bash
|
||||
git add backend/src/workspaces/qdrant-collection.ts \
|
||||
backend/test/qdrant-collection.test.ts harness/tht/adapters/vector/qdrant.py \
|
||||
backend/src/workspaces/evidence/preprocessing.ts \
|
||||
backend/test/qdrant-collection.test.ts backend/test/workspaces/evidence/preprocessing.test.ts \
|
||||
harness/tht/adapters/vector/qdrant.py \
|
||||
harness/tests/test_qdrant_vector_store.py docs/contracts/workspace-preprocessing-cli.md
|
||||
git commit -m "feat(evidence): require dense and bm25 qdrant vectors"
|
||||
git commit -m "feat(evidence): add qdrant bm25 vector in place"
|
||||
```
|
||||
|
||||
### Task 7: Add server-side BM25 ingestion and hybrid Query API retrieval
|
||||
@@ -593,6 +671,7 @@ git commit -m "feat(evidence): require dense and bm25 qdrant vectors"
|
||||
- Modify: `harness/tests/test_vector_port_contract.py`
|
||||
- Modify: `harness/tests/test_qdrant_vector_store.py`
|
||||
- Modify: `harness/tests/test_corpus_pipeline.py`
|
||||
- Create: `harness/tests/l0/test_qdrant_bm25_inference.py`
|
||||
|
||||
**Step 1: Write RED port tests**
|
||||
|
||||
@@ -617,7 +696,7 @@ For Evidence upsert, assert:
|
||||
|
||||
```json
|
||||
"vector": {
|
||||
"dense": [0.1, 0.2],
|
||||
"": [0.1, 0.2],
|
||||
"bm25": {
|
||||
"text": "...",
|
||||
"model": "qdrant/bm25",
|
||||
@@ -631,7 +710,7 @@ For hybrid search, assert two filtered prefetches and default RRF:
|
||||
```json
|
||||
{
|
||||
"prefetch": [
|
||||
{"query": [0.1, 0.2], "using": "dense", "limit": 20, "filter": {}},
|
||||
{"query": [0.1, 0.2], "limit": 20, "filter": {}},
|
||||
{
|
||||
"query": {
|
||||
"text": "fascia pediatrica",
|
||||
@@ -652,42 +731,54 @@ For hybrid search, assert two filtered prefetches and default RRF:
|
||||
Both prefetch filters must include workspace, revision, active generation and record
|
||||
kind. Do not add hand-tuned weights.
|
||||
|
||||
**Step 3: Implement named dense writes for all semantic records**
|
||||
**Step 3: Prove BM25 on the actual local Qdrant image**
|
||||
|
||||
All existing schema/memory/Evidence records use the `dense` vector name after the
|
||||
collection rebuild. Only records with `sparse_text` receive `bm25`.
|
||||
Add an `l0` test that reads the Qdrant image reference from the root `compose.yaml`,
|
||||
starts that exact image with testcontainers, creates a uniquely named temporary
|
||||
collection, adds `bm25` with IDF, indexes two Italian texts through server-side
|
||||
`qdrant/bm25`, retrieves the expected text and deletes the collection. The test must
|
||||
not use FastEmbed or accept a dense fallback. It fails if the Compose reference and the
|
||||
tested image diverge.
|
||||
|
||||
**Step 4: Implement BM25 Evidence ingestion**
|
||||
**Step 4: Preserve dense writes for all existing semantic records**
|
||||
|
||||
Schema, Memory and solved-question records keep their current unnamed dense writes.
|
||||
Evidence records use the empty default-vector name plus `bm25` in the mixed-vector
|
||||
upsert shape. Only records with `sparse_text` receive `bm25`; no full reindex of Schema
|
||||
or Memory is performed.
|
||||
|
||||
**Step 5: Implement BM25 Evidence ingestion**
|
||||
|
||||
Populate `sparse_text` and map workspace language `it` to Qdrant's `italian`. Reject an
|
||||
unsupported language before uploading the generation. Use the same options at ingest
|
||||
and query time.
|
||||
|
||||
**Step 5: Implement hybrid search with dense fallback only for non-Evidence callers**
|
||||
**Step 6: Implement hybrid search with dense fallback only for non-Evidence callers**
|
||||
|
||||
An Evidence hybrid request must fail as unavailable if the configured Qdrant version or
|
||||
collection contract does not support BM25. It must not silently claim to have run hybrid
|
||||
search. Existing non-Evidence dense requests continue to work.
|
||||
|
||||
**Step 6: Run focused tests**
|
||||
**Step 7: Run focused tests**
|
||||
|
||||
```bash
|
||||
cd harness
|
||||
.venv/bin/pytest tests/test_vector_port_contract.py tests/test_qdrant_vector_store.py \
|
||||
tests/test_corpus_pipeline.py tests/test_search_pack.py -q
|
||||
.venv/bin/pytest tests/l0/test_qdrant_bm25_inference.py -m l0 -q
|
||||
.venv/bin/ruff check tht/ports/vector.py tht/adapters/vector/qdrant.py \
|
||||
tht/evidence/corpus/pipeline.py
|
||||
```
|
||||
|
||||
Expected: PASS.
|
||||
|
||||
**Step 7: Commit**
|
||||
**Step 8: Commit**
|
||||
|
||||
```bash
|
||||
git add harness/tht/ports/vector.py harness/tht/adapters/vector/qdrant.py \
|
||||
harness/tht/vectorstore/records.py harness/tht/evidence/corpus/pipeline.py \
|
||||
harness/tests/test_vector_port_contract.py harness/tests/test_qdrant_vector_store.py \
|
||||
harness/tests/test_corpus_pipeline.py
|
||||
harness/tests/test_corpus_pipeline.py harness/tests/l0/test_qdrant_bm25_inference.py
|
||||
git commit -m "feat(evidence): add qdrant bm25 hybrid retrieval"
|
||||
```
|
||||
|
||||
@@ -697,8 +788,12 @@ git commit -m "feat(evidence): add qdrant bm25 hybrid retrieval"
|
||||
|
||||
- Modify: `harness/tht/evidence/search.py`
|
||||
- Modify: `harness/tht/evidence/__init__.py`
|
||||
- Modify: `harness/tht/evidence/session.py`
|
||||
- Modify: `harness/tht/cli/search_cmd.py`
|
||||
- Modify: `harness/tht/session/filesystem_repository.py`
|
||||
- Modify: `harness/tht/session/postgres_repository.py`
|
||||
- Modify: `harness/tests/test_evidence_facade_contract.py`
|
||||
- Create: `harness/tests/test_evidence_session_receipts.py`
|
||||
- Modify: `harness/tests/test_search_pack.py`
|
||||
- Create: `harness/.pi/skills/tht-sessione/modules/evidence/runtime-search.md`
|
||||
- Modify: `harness/.pi/skills/tht-sessione/projection.md.tmpl`
|
||||
@@ -716,6 +811,9 @@ class EvidenceSearchContext(BaseModel):
|
||||
tables: tuple[str, ...] = ()
|
||||
columns: tuple[str, ...] = ()
|
||||
required_kinds: tuple[EvidenceKind, ...] = ()
|
||||
required_concepts: tuple[str, ...] = ()
|
||||
required_tables: tuple[str, ...] = ()
|
||||
required_columns: tuple[str, ...] = ()
|
||||
|
||||
def search_evidence(
|
||||
query: str,
|
||||
@@ -725,50 +823,105 @@ def search_evidence(
|
||||
searcher: ActiveEvidenceSearcher,
|
||||
embedder: EvidenceQueryEmbedder,
|
||||
top_n: int = 10,
|
||||
) -> list[EvidenceResult]: ...
|
||||
) -> EvidenceSearchOutcome: ...
|
||||
```
|
||||
|
||||
Test hard filters for workspace/revision/generation and explicit `required_kinds`. Test
|
||||
that purpose/kind/scope preferences are deterministic tie-breakers when they were not
|
||||
requested as hard filters.
|
||||
Test hard filters for workspace, revision, generation, purpose and every explicit
|
||||
`required_*` constraint. Concepts, tables and columns without the `required_` prefix
|
||||
enrich the dense and BM25 query text but do not create payload filters. Do not add
|
||||
implicit kind or scope bonuses in v1.
|
||||
|
||||
Render that enrichment once, identically for both retrieval branches, in this fixed
|
||||
order:
|
||||
|
||||
```text
|
||||
Domanda: <original query>
|
||||
Concetti: <normalized, deduplicated, sorted values>
|
||||
Tabelle: <normalized, deduplicated, sorted values>
|
||||
Colonne: <normalized, deduplicated, sorted values>
|
||||
```
|
||||
|
||||
Omit empty lines and do not rewrite the original question. Test that permutations and
|
||||
duplicates in context produce the same rendered text and that the dense embedder and
|
||||
BM25 request receive exactly that same value.
|
||||
|
||||
The renderer normalizes the question to Unicode NFC, converts CRLF and CR to `\n`,
|
||||
strips only leading and trailing whitespace and rejects an empty result. It preserves
|
||||
case, punctuation and internal whitespace. Apply NFC and `strip` to context values,
|
||||
remove empty strings and exact duplicates, then sort by Unicode value. Do not call
|
||||
`lower()` or `casefold()` because quoted PostgreSQL identifiers can be case-sensitive.
|
||||
|
||||
`EvidenceSearchOutcome` has an `available` state with a generation and zero or more
|
||||
results, and an `unavailable` state with a stable error code, a bounded message and no
|
||||
results. `memory` is not an `EvidencePurpose`: the `memory` stage remains owned by the
|
||||
Memory Module.
|
||||
|
||||
**Step 2: Test fragment grouping**
|
||||
|
||||
Two returned fragments with the same `evidence_id` must become one `EvidenceResult`,
|
||||
with the best score, ordered matching excerpts, canonical citation and no duplicate
|
||||
unit.
|
||||
with the best score, ordered matching excerpts, canonical citation, a reference for
|
||||
resolving the full document and no duplicate unit. Do not insert the full document into
|
||||
the search pack automatically.
|
||||
|
||||
**Step 3: Preserve fail-closed graceful degradation**
|
||||
**Step 3: Distinguish an empty search from technical unavailability**
|
||||
|
||||
Test absent ACTIVE corpus, revision mismatch, unavailable Qdrant and malformed payload.
|
||||
All return no Evidence plus a bounded warning through the existing search-pack contract;
|
||||
none uses stale rows.
|
||||
Test a successful search with zero matches separately from absent ACTIVE corpus,
|
||||
revision mismatch, unavailable Qdrant and malformed payload. The first returns
|
||||
`available` with an empty result list and may continue. Every technical case returns
|
||||
`unavailable`, blocks the calling stage until retry, and never uses stale rows or a
|
||||
fallback purpose.
|
||||
|
||||
**Step 4: Implement the facade and CLI mapping**
|
||||
|
||||
Keep Qdrant syntax inside `tht.evidence`. `search_cmd.py` translates command inputs to
|
||||
the facade and renders results; it must not duplicate ranking logic.
|
||||
|
||||
**Step 5: Extract Evidence instructions into a module fragment**
|
||||
**Step 5: Integrate the contributor with semantic stages**
|
||||
|
||||
Move the common F1/F3/F4 Evidence rules from the projection template into
|
||||
Move the shared Evidence rules from the projection template into
|
||||
`modules/evidence/runtime-search.md`. The fragment must state:
|
||||
|
||||
- candidates are not truth;
|
||||
- pass the phase-appropriate purpose;
|
||||
- show provenance;
|
||||
- formulas are `kind=formula`, not a separate search store;
|
||||
- absence of Evidence is visible but does not stop the whole session.
|
||||
- an available empty result is visible and does not block the stage;
|
||||
- an unavailable outcome blocks the stage and is retriable;
|
||||
- Evidence never writes decisions, canonical artifacts or workflow state.
|
||||
|
||||
Map semantic stages, independently of their display codes:
|
||||
|
||||
```text
|
||||
clarification -> disambiguation
|
||||
rewriting -> rewriting
|
||||
schema_linking -> schema_linking
|
||||
cte -> sql_generation
|
||||
final_sql -> sql_generation
|
||||
```
|
||||
|
||||
Do not invoke Evidence from `memory` or `synthesis`. Run an independent search in every
|
||||
mapped stage; the `final_sql` query includes the approved CTE plan.
|
||||
|
||||
**Step 6: Persist the minimal Evidence receipt**
|
||||
|
||||
Use the existing session artifact repositories to maintain one
|
||||
`evidence_receipts.json` artifact. For every available search, append or replace the
|
||||
receipt identified by semantic stage with exactly `stage`, `purpose`,
|
||||
`vector_generation` and ordered `evidence_ids`. Do not copy excerpts or complete
|
||||
Evidence text. The workflow integration owns the write; `tht.evidence` only constructs
|
||||
and returns the typed receipt. Test filesystem and PostgreSQL repositories, resume and
|
||||
retry replacement.
|
||||
|
||||
Register the fragment in the static `FRAGMENT_ORDER` and regenerate.
|
||||
|
||||
**Step 6: Run tests**
|
||||
**Step 7: Run tests**
|
||||
|
||||
```bash
|
||||
cd harness
|
||||
python -m tht.pi_skill_projection --write
|
||||
python -m tht.pi_skill_projection --check
|
||||
.venv/bin/pytest tests/test_evidence_facade_contract.py tests/test_search_pack.py \
|
||||
tests/test_evidence_session_receipts.py \
|
||||
tests/test_pi_skill_projection.py -q
|
||||
```
|
||||
|
||||
@@ -878,28 +1031,39 @@ schema_version: 1
|
||||
queries:
|
||||
- id: pediatric-formula
|
||||
query: Come distinguo i pazienti pediatrici?
|
||||
profile: semantic
|
||||
purpose: sql_generation
|
||||
expected:
|
||||
- formula:fascia-pediatrica
|
||||
- evidence:fascia-pediatrica
|
||||
```
|
||||
|
||||
Require unique query IDs, nonempty expected IDs and only public purpose values.
|
||||
Require unique query IDs, nonempty expected IDs, one of `lexical`, `semantic` or
|
||||
`mixed` for every profile, only public purpose values and at least one query of every
|
||||
profile in the complete file.
|
||||
|
||||
**Step 2: Write RED metric tests**
|
||||
|
||||
Compute `hit_at_5`, `hit_at_10`, missing expected IDs, empty-result queries and counts by
|
||||
expected `kind`. Do not add nDCG, relevance grading or an evaluation database in v1.
|
||||
expected `kind`. For each expected ID, also report its nullable dense-only, BM25-only
|
||||
and fused rank. The evaluator runs both branches separately for diagnosis and the same
|
||||
hybrid request used by runtime. The report passes only when every query finds at least
|
||||
one expected ID in its first ten fused results; branch ranks and `hit_at_5` are
|
||||
informative. Do not add nDCG, relevance grading or an evaluation database in v1.
|
||||
|
||||
**Step 3: Implement the evaluator and CLI**
|
||||
|
||||
Expose:
|
||||
|
||||
```text
|
||||
tht evidence evaluate <workspace-root> -c <runtime-config> [--json]
|
||||
tht evidence evaluate <workspace-root> -c <runtime-config>
|
||||
[--generation <id>] [--json]
|
||||
```
|
||||
|
||||
The report includes workspace revision, active vector generation and the fixed default
|
||||
RRF configuration. It is read-only.
|
||||
The report includes workspace revision, evaluated vector generation and the fixed
|
||||
default RRF configuration. It is read-only. With no `--generation`, it evaluates the
|
||||
active generation for monitoring. Before publication, the preprocessing pipeline calls
|
||||
the same evaluator against its candidate generation and atomically activates it only
|
||||
when the report passes. A failed candidate remains inactive.
|
||||
|
||||
**Step 4: Run tests**
|
||||
|
||||
@@ -957,7 +1121,7 @@ Document:
|
||||
- Git publication boundary;
|
||||
- no runtime writes;
|
||||
- validation before indexing;
|
||||
- dense+BM25 collection contract and guarded rebuild;
|
||||
- unnamed-dense plus BM25 contract and additive Evidence upgrade;
|
||||
- exact public operation names and JSON status additions, if any.
|
||||
|
||||
**Step 4: Run backend and harness contract gates**
|
||||
@@ -1017,10 +1181,24 @@ and a SQL formula candidate. The acceptance runner must prove:
|
||||
- ambiguity becomes `review_items`;
|
||||
- no cross-source merge;
|
||||
- unchanged rerun is a no-op;
|
||||
- human edit is preserved;
|
||||
- committed human content remains recoverable and proposed changes are visible in Git;
|
||||
- dirty-tree refusal;
|
||||
- validation blocks unresolved review;
|
||||
- validated corpus indexes and searches hybrid;
|
||||
- validation blocks unresolved review and orphaned units;
|
||||
- pipeline-version mismatch refusal and explicit full-corpus `--upgrade`;
|
||||
- validated corpus builds an inactive candidate and searches it hybrid;
|
||||
- every evaluation query retrieves at least one expected ID in the first ten results
|
||||
before the candidate becomes active;
|
||||
- the evaluation set contains lexical, semantic and mixed cases and reports dense,
|
||||
BM25 and fused ranks separately;
|
||||
- dense and BM25 receive the same deterministic query text;
|
||||
- no atomic formula, enum pair, mapping, rule or URL is split into fixed-size chunks;
|
||||
- the complete rendered fragment stays within the existing 4,000-character
|
||||
`max_chunk_chars` default and no parallel size option exists;
|
||||
- the L0 probe passes against the Qdrant image referenced by `compose.yaml` without
|
||||
FastEmbed or fallback;
|
||||
- query normalization is limited to NFC, newline canonicalization and outer trimming,
|
||||
preserving internal whitespace, punctuation and case-sensitive identifiers;
|
||||
- Formula Evidence accepts a PostgreSQL expression and rejects a complete query;
|
||||
- Qdrant failure remains fail-closed.
|
||||
|
||||
Use a fake restructurer for hermetic CI. The real Pi call is a separate manual check.
|
||||
@@ -1054,7 +1232,7 @@ Provide:
|
||||
- proposed PSD branch name;
|
||||
- exact list of 36 source files to move;
|
||||
- rollback command based on the pre-migration PSD commit;
|
||||
- notice that Qdrant rebuild is destructive but scoped by exact collection-name guards.
|
||||
- before/after Schema and Memory counts plus sample IDs used to prove non-regression.
|
||||
|
||||
Do not infer approval from prior design acceptance.
|
||||
|
||||
@@ -1064,7 +1242,7 @@ Use `tht evidence prepare`, review the Git diff, resolve all review items manual
|
||||
`tht evidence validate`, and create the approximately twenty evaluation queries. Do not
|
||||
auto-merge or auto-push unless separately requested.
|
||||
|
||||
**Step 5: Activate and rebuild with the existing guarded operator path**
|
||||
**Step 5: Activate and perform the additive BM25 upgrade**
|
||||
|
||||
After the PSD merge/pull and activation, inspect first:
|
||||
|
||||
@@ -1073,20 +1251,24 @@ tht --installation <absolute>/thothii-installation.yaml workspace vector inspect
|
||||
--workspace psd-clinical --json
|
||||
```
|
||||
|
||||
Then use the exact descriptor-owned name in the guarded rebuild command. Run
|
||||
`workspace preprocess evidence`, `tht evidence evaluate`, and the manual F1/F3/F4
|
||||
walkthrough.
|
||||
Before preprocessing, record counts and representative IDs for `schema_table`,
|
||||
`schema_column`, `memory` and `solved_question`. Run `workspace preprocess evidence`:
|
||||
it adds `bm25` when absent, builds the candidate, evaluates that exact generation and
|
||||
publishes it only if the report passes. Repeat the counts, verify the representative IDs
|
||||
and run dense smoke searches for Schema and Memory before completing the manual
|
||||
walkthrough of `clarification`, `rewriting`, `schema_linking`, `cte` and `final_sql`.
|
||||
|
||||
**Step 6: Record acceptance and commit ThothII documentation**
|
||||
|
||||
`docs/testing/evidence-restructuring-manual.md` must record separate outcomes for:
|
||||
|
||||
- authoring and Git review;
|
||||
- collection rebuild;
|
||||
- additive BM25 schema upgrade;
|
||||
- Schema and Memory before/after non-regression evidence;
|
||||
- preprocessing generation publication;
|
||||
- hybrid retrieval evaluation;
|
||||
- formula retrieval;
|
||||
- graceful degradation;
|
||||
- distinzione fra risultato vuoto e indisponibilità bloccante;
|
||||
- complete session behavior.
|
||||
|
||||
Update `PROJECT_STATE.md` only with observed results and immutable commit/run IDs.
|
||||
@@ -1105,17 +1287,38 @@ Before claiming completion, verify:
|
||||
- [ ] `CONTEXT.md` and both Evidence plan documents use the same terminology.
|
||||
- [ ] All eight Evidence kinds have type-specific positive and negative tests.
|
||||
- [ ] Pi is called once per changed source, with no tools and no saved session.
|
||||
- [ ] `prepare` refuses dirty curated state and never deletes orphans.
|
||||
- [ ] `validate` blocks unresolved review items.
|
||||
- [ ] `prepare` refuses dirty curated state, exposes every proposal in Git and never
|
||||
deletes orphans.
|
||||
- [ ] `validate` blocks unresolved review items and orphaned units.
|
||||
- [ ] Source renames and kind-only reclassifications preserve unit IDs; semantic splits
|
||||
receive new IDs.
|
||||
- [ ] IDs use `evidence:<slug>`, remain independent from kind and are not recomputed
|
||||
after their initial assignment.
|
||||
- [ ] Formula Evidence accepts one PostgreSQL expression and rejects full queries.
|
||||
- [ ] A pipeline-version mismatch performs no writes without explicit `--upgrade`.
|
||||
- [ ] Runtime reads only `curated/**/*.md` from the pinned Git revision.
|
||||
- [ ] The TypeScript and Python Qdrant compatibility checks agree.
|
||||
- [ ] Qdrant collection uses named `dense` plus `bm25` with IDF.
|
||||
- [ ] Qdrant retains the unnamed dense vector and adds only `bm25` with IDF.
|
||||
- [ ] Evidence preprocessing adds a missing `bm25` definition but never deletes,
|
||||
renames or destructively rebuilds collection vectors.
|
||||
- [ ] Schema and Memory counts, sample IDs and dense searches remain unchanged across
|
||||
the additive upgrade.
|
||||
- [ ] Italian BM25 options are identical during ingest and query.
|
||||
- [ ] Hybrid search uses two prefetches and default RRF.
|
||||
- [ ] Hard filters always include workspace, revision and active generation.
|
||||
- [ ] Hard filters always include workspace, revision, active generation and purpose;
|
||||
only explicit `required_*` context values add further filters.
|
||||
- [ ] Formula runtime lookup uses typed Evidence; session proposals remain non-published.
|
||||
- [ ] Fragment hits are grouped into complete Evidence Units.
|
||||
- [ ] Evaluation reports hit@5 and hit@10 against a versioned fixture.
|
||||
- [ ] Fragment hits are grouped into one Evidence Result with excerpts and a reference;
|
||||
full documents are loaded only on demand.
|
||||
- [ ] Evidence is queried independently from the five mapped semantic stages and is not
|
||||
invoked from `memory` or `synthesis`.
|
||||
- [ ] An available empty result can continue; technical unavailability blocks the
|
||||
current stage without stale-generation or purpose fallback.
|
||||
- [ ] Each available stage search persists only its minimal Evidence receipt.
|
||||
- [ ] Evaluation reports hit@5 and hit@10 against a versioned fixture and passes only
|
||||
when every query has at least one expected result in the first ten.
|
||||
- [ ] Evaluation runs against the candidate generation before atomic activation; a
|
||||
failed candidate remains invisible to sessions.
|
||||
- [ ] Existing corpus rollback, compensation, retention and resume tests still pass.
|
||||
- [ ] Backend Vitest and TypeScript gates pass.
|
||||
- [ ] Harness pytest and Ruff gates pass.
|
||||
|
||||
Reference in New Issue
Block a user