Files
ThothII/docs/plans/2026-08-24-evidence-restructuring.md
T

48 KiB
Raw Blame History

Evidence Restructuring Implementation Plan

Execution: GitHub issue #35 is the approved parent specification. Issues #36–#47 are the executable tracer-bullet tickets; implement one unblocked ticket at a time with /implement. This document remains the detailed technical reference and must not be executed as a second, parallel work queue.

Goal: Build a Git-reviewed, typed Evidence authoring pipeline and publish its approved output to the existing revision-scoped Qdrant lifecycle with dense+BM25 hybrid retrieval.

Architecture: Keep authoring outside NL→SQL sessions: deterministic preparation wraps one read-only Pi restructuring call, writes reviewable Markdown into the workspace repository, and blocks publication on unresolved review items. Reuse the current Evidence module, corpus generations, rollback, active-revision checks, and workspace-owned semantic collection; extend them instead of introducing a parallel store.

Tech Stack: Python 3.12, Pydantic 2, Typer, PyYAML, sqlglot, pytest, Pi CLI, TypeScript, Fastify workspace maintenance, Vitest, Qdrant 1.18.2 Query API, server-side qdrant/bm25, Git.


Preconditions

  • Work in this dedicated worktree.
  • Read PROJECT_STATE.md, CONTEXT.md, docs/plans/2026-08-24-evidence-restructuring-design.md, docs/contracts/workspace-evidence-v3.md, and docs/contracts/workspace-preprocessing-cli.md.
  • Preserve the public facade in harness/tht/evidence/__init__.py.
  • Do not edit harness/.pi/skills/tht-sessione/SKILL.md directly; regenerate it with python -m tht.pi_skill_projection --write.
  • Keep runtime workspace access read-only. Only the authoring CLI may write evidence/curated/ and evidence/manifest.yaml in a curator clone.
  • Do not migrate the external PSD repository until all code and contract gates pass.

Task 1: Add the typed Curated Evidence model

Files:

  • Create: harness/tht/evidence/canonical.py
  • Modify: harness/tht/evidence/__init__.py
  • Create: harness/tests/test_evidence_canonical.py

Step 1: Write failing tests for the common envelope

Cover:

  • all eight kind values;
  • all four purpose values;
  • immutable provenance;
  • strict unknown-field rejection;
  • kind-independent stable IDs;
  • SHA-256 syntax;
  • directory/kind agreement;
  • parsing and dumping Markdown with YAML frontmatter.
  • review items containing a stable code, human message and optional field, with no status or history fields;
  • IDs matching evidence:<slug> and remaining independent from kind.

Start with:

def test_formula_requires_formula_payload():
    with pytest.raises(ValidationError):
        CuratedEvidence.model_validate({
            **COMMON,
            "kind": "formula",
            "payload": {"concept": "fascia pediatrica"},
        })


def test_reference_rejects_formula_payload():
    with pytest.raises(ValidationError):
        CuratedEvidence.model_validate({
            **COMMON,
            "kind": "reference",
            "payload": {
                "concept": "x",
                "columns": ["clinical.patient.birth_date"],
                "sql": "CASE WHEN true THEN 1 END",
            },
        })

Step 2: Run the focused tests and confirm RED

Run:

cd harness
.venv/bin/pytest tests/test_evidence_canonical.py -q

Expected: import failure for tht.evidence.canonical.

Step 3: Implement the discriminated model

Use a strict Pydantic model with these public types:

EvidenceKind = Literal[
    "glossary", "domain", "enum", "example", "mapping",
    "normalization", "formula", "reference",
]
EvidencePurpose = Literal[
    "disambiguation", "rewriting", "schema_linking", "sql_generation",
]

class EvidenceScope(StrictModel):
    concepts: tuple[str, ...] = ()
    tables: tuple[str, ...] = ()
    columns: tuple[str, ...] = ()

class EvidenceProvenance(StrictModel):
    source_file: str
    source_sha256: str
    supporting_excerpts: tuple[str, ...]

class FormulaPayload(StrictModel):
    concept: str
    columns: tuple[str, ...]
    sql: str

class ReviewItem(StrictModel):
    code: str
    message: str
    field: str | None = None

class ReferencePayload(StrictModel):
    url: AnyHttpUrl
    label: str
    description: str

Define equally strict payloads for the other six kinds and expose a single CuratedEvidence API. The implementation may use an internal Pydantic discriminated union, but callers must not switch between eight unrelated loaders.

Add:

def parse_curated_markdown(text: str, *, path: Path | None = None) -> CuratedEvidence: ...
def dump_curated_markdown(value: CuratedEvidence) -> str: ...
def load_curated_tree(root: Path) -> list[CuratedEvidence]: ...

Validate formula SQL with sqlglot as one PostgreSQL expression, rejecting complete SELECT/WITH statements, DDL, DML and multiple statements. Validate schema.table / schema.table.column identifiers without querying the DWH. Require one to five supporting excerpts, each nonempty and at most 1,000 characters.

Step 4: Export only the stable API

Add the model and loader functions to harness/tht/evidence/__init__.py. Do not export internal union member helpers unless another module needs them.

Step 5: Run tests and lint

Run:

cd harness
.venv/bin/pytest tests/test_evidence_canonical.py tests/test_evidence_facade_contract.py -q
.venv/bin/ruff check tht/evidence/canonical.py tests/test_evidence_canonical.py

Expected: PASS.

Step 6: Commit

git add harness/tht/evidence/canonical.py harness/tht/evidence/__init__.py \
  harness/tests/test_evidence_canonical.py
git commit -m "feat(evidence): add typed curated evidence model"

Task 2: Add deterministic validation and the Evidence manifest

Files:

  • Create: harness/tht/evidence/authoring.py
  • Create: harness/tests/test_evidence_authoring.py
  • Modify: harness/tht/evidence/__init__.py

Step 1: Write failing tests for publication validation

Test that:

  • review_items != [] is valid as a draft but blocks publication;
  • orphans != [] is preserved but blocks publication;
  • a missing or mismatched source hash blocks publication;
  • every supporting excerpt is found after applying the same mechanical normalization to the excerpt and its source;
  • two units cannot share an ID;
  • one unit cannot claim two source files;
  • curated/formula/x.md must contain kind: formula;
  • credentials in URLs, YAML, or body are rejected;
  • only Markdown, .txt, and .sql.md sources are accepted;
  • UTF-8 and per-file limits are enforced.

Expose errors as bounded structured values:

@dataclass(frozen=True)
class ValidationFinding:
    severity: Literal["error", "warning"]
    code: str
    path: str
    message: str

Step 2: Write failing tests for the versioned manifest

Use this minimum shape:

schema_version: 1
pipeline_version: evidence-authoring-v1
sources:
  source/domain/patient.md:
    sha256: sha256:...
    units:
      - domain:patient
orphans: []

Test deterministic key ordering, stable round-trip, unknown fields, duplicate unit IDs, orphan preservation and incompatible pipeline_version refusal.

Step 3: Run and confirm RED

Run:

cd harness
.venv/bin/pytest tests/test_evidence_authoring.py -q

Expected: missing authoring API.

Step 4: Implement EvidenceManifest and validate_workspace_evidence

Provide:

def load_manifest(path: Path) -> EvidenceManifest: ...
def dump_manifest(manifest: EvidenceManifest) -> str: ...
def validate_workspace_evidence(workspace_root: Path) -> ValidationReport: ...

ValidationReport.publishable is true only when there are no errors and no unresolved review items or orphaned units. Warnings alone do not block publication.

Do not use status fields such as draft/reviewed as an approval mechanism. The approved Git revision is the publication boundary.

Step 5: Run tests and lint

cd harness
.venv/bin/pytest tests/test_evidence_authoring.py tests/test_evidence_canonical.py -q
.venv/bin/ruff check tht/evidence/authoring.py tests/test_evidence_authoring.py

Expected: PASS.

Step 6: Commit

git add harness/tht/evidence/authoring.py harness/tht/evidence/__init__.py \
  harness/tests/test_evidence_authoring.py
git commit -m "feat(evidence): validate curated corpus and manifest"

Task 3: Implement incremental preparation with one Pi restructuring call

Files:

  • Modify: harness/tht/evidence/authoring.py
  • Create: harness/.pi/skills/tht-evidence-authoring/SKILL.md
  • Modify: harness/tests/test_evidence_authoring.py
  • Create: harness/tests/test_evidence_pi_restructurer.py

Step 1: Define the restructuring port and request/response models

Add:

class EvidenceRestructurer(Protocol):
    def restructure(self, request: RestructureRequest) -> tuple[RestructureCandidate, ...]: ...

class RestructureRequest(StrictModel):
    source_file: str
    source_sha256: str
    normalized_text: str
    previous_units: tuple[CuratedEvidence, ...] = ()

class RestructureCandidate(StrictModel):
    existing_id: str | None = None
    # The same title, kind, purposes, scope, provenance excerpts and typed payload
    # needed to construct CuratedEvidence, but no model-assigned canonical ID.

existing_id, when present, must belong to previous_units. The preparer rejects any unknown ID and deterministically allocates evidence:<slug> plus a collision suffix for every candidate without one. The response is converted to and validated through the models from Task 1 before any write.

Step 2: Write RED tests for the incremental rules

Test:

  • unchanged source: no model call and no file write;
  • changed source: exactly one model call;
  • new source: new stable IDs;
  • removed source: old units become orphans and remain on disk;
  • uniquely renamed source with the same hash: provenance changes and unit IDs remain;
  • existing source no longer supporting a prior unit: retain it with source_no_longer_supports_unit and block publication;
  • kind-only reclassification: preserve the unit ID;
  • semantic split: allocate new IDs for the new independent units;
  • a semantic split retains the previous unit as a blocking retirement candidate until the curator explicitly retires it;
  • one source may produce several units;
  • every returned unit contains one to five exact supporting excerpts found in its normalized source;
  • no returned unit may cite another source;
  • prior curated units are included in the request;
  • a dirty evidence/curated/ or evidence/manifest.yaml fails before the model call;
  • committed human edits remain recoverable in Git and every proposed change is exposed in the working-tree diff;
  • all output writes are staged and atomically replaced only after complete validation.
  • one invalid source leaves the complete batch and manifest unchanged;

Inject the Git status runner and filesystem writer in tests; do not require a real Git repository for every unit test.

Step 3: Implement deterministic source normalization

Normalize UTF-8 text with NFC, LF newlines and terminal newline. Preserve meaningful Markdown, tables, fenced SQL, URLs and list structure. Do not rewrite vocabulary or infer domain facts in this step.

Step 4: Implement PiEvidenceRestructurer

Invoke Pi as an ephemeral, no-tools process using an argument list, never a shell:

argv = [
    pi_executable,
    "--mode", "text",
    "--print",
    "--no-session",
    "--no-tools",
    "--no-extensions",
    "--no-context-files",
    "--skill", str(skill_path),
    f"@{request_path}",
    "Return only the JSON object required by the Evidence authoring skill.",
]

Use a private TemporaryDirectory, a bounded timeout, bounded stdout/stderr, and strict JSON parsing. Do not forward the model's raw output to public JSON errors. The skill must state:

  • use only facts present in normalized_text;
  • preserve prior reviewed wording where it remains supported;
  • retain unsupported prior units with a source_no_longer_supports_unit review item;
  • never merge sources;
  • emit review_items for uncertainty;
  • copy short exact supporting excerpts from normalized_text;
  • use only a supplied existing_id or omit it for deterministic allocation;
  • emit exactly the schema-versioned JSON object and no Markdown fence.

Do not retry automatically after a timeout, nonzero exit, malformed JSON or invalid response. Return a stable error code and the affected source so the curator can rerun the command explicitly.

Step 5: Implement prepare_workspace_evidence

Return a bounded report with changed, unchanged, created, orphaned, findings, and model_calls. Keep the model implementation behind EvidenceRestructurer. Build the complete batch in a private staging directory and replace curated files plus the manifest only after every candidate validates; never expose partial success.

Step 6: Run tests

cd harness
.venv/bin/pytest tests/test_evidence_authoring.py tests/test_evidence_pi_restructurer.py -q
.venv/bin/ruff check tht/evidence/authoring.py tests/test_evidence_pi_restructurer.py

Expected: PASS with no live model call.

Step 7: Commit

git add harness/tht/evidence/authoring.py \
  harness/.pi/skills/tht-evidence-authoring/SKILL.md \
  harness/tests/test_evidence_authoring.py harness/tests/test_evidence_pi_restructurer.py
git commit -m "feat(evidence): prepare curated evidence incrementally"

Task 4: Add the authoring CLI

Files:

  • Create: harness/tht/cli/evidence_cmd.py
  • Modify: harness/tht/cli/__init__.py
  • Create: harness/tests/test_evidence_cli.py

Step 1: Write CLI grammar tests

Cover:

tht evidence prepare <workspace-root> [--upgrade] [--json]
tht evidence validate <workspace-root> [--json]
tht evidence resolve <workspace-root> <evidence-id> --retire [--json]
tht evidence resolve <workspace-root> <evidence-id> --source <source-path> [--json]

Require an existing canonical Git worktree root. Reject unknown flags, symlinks, a workspace path outside the Git root, duplicate options and dirty curated state. A pipeline-version mismatch must fail without writes unless --upgrade is explicit; --upgrade reprocesses every source. Ensure --json writes pristine JSON to stdout. For resolve, require exactly one of --retire and --source, an existing Evidence ID, an existing in-root source path for relinking and a clean worktree. Test atomic updates to the curated file and manifest, with no commit or publication side effect.

Step 2: Run and confirm RED

cd harness
.venv/bin/pytest tests/test_evidence_cli.py -q

Expected: evidence command group is unknown.

Step 3: Implement the Typer group

Use one top-level authoring group:

evidence_app = typer.Typer(help="Prepare and validate workspace Evidence")

@evidence_app.command("prepare")
def prepare_cmd(workspace_root: Path, json_output: bool = False) -> None: ...

@evidence_app.command("validate")
def validate_cmd(workspace_root: Path, json_output: bool = False) -> None: ...

@evidence_app.command("resolve")
def resolve_cmd(
    workspace_root: Path,
    evidence_id: str,
    retire: bool = False,
    source: Path | None = None,
    json_output: bool = False,
) -> None: ...

Register it in harness/tht/cli/__init__.py. Keep this separate from the existing runtime command tht preprocess evidence.

Exit codes:

  • 0: prepared/unchanged or valid;
  • 2: unsafe path or CLI misuse;
  • 3: valid drafts but review required;
  • 1: operational/model/structural failure.

Step 4: Run tests and CLI help

cd harness
.venv/bin/pytest tests/test_evidence_cli.py tests/test_preprocess_cli.py -q
.venv/bin/tht evidence --help
.venv/bin/tht preprocess evidence --help

Expected: both command families are present and unambiguous.

Step 5: Commit

git add harness/tht/cli/evidence_cmd.py harness/tht/cli/__init__.py \
  harness/tests/test_evidence_cli.py
git commit -m "feat(evidence): expose prepare and validate commands"

Task 5: Project typed units into semantic Evidence Fragments

Files:

  • Modify: harness/tht/evidence/corpus/models.py
  • Modify: harness/tht/evidence/corpus/chunk.py
  • Modify: harness/tht/evidence/corpus/normalize.py
  • Modify: harness/tht/evidence/corpus/pipeline.py
  • Modify: harness/tht/evidence/preprocessing.py
  • Modify: harness/tests/test_corpus_models.py
  • Modify: harness/tests/test_corpus_chunk.py
  • Modify: harness/tests/test_corpus_pipeline.py

Step 1: Write failing projection tests

Test that:

  • a short formula remains one fragment;
  • a domain document splits only at semantic section boundaries;
  • enum entries are not split in the middle of a value/meaning pair;
  • formulas, value/meaning pairs, mappings, rules and URLs are never split;
  • the complete rendered fragment, including labels and textual metadata, uses the existing max_chunk_chars setting whose default is 4,000 characters;
  • an atomic element over max_chunk_chars produces the blocking atomic_content_too_large review item instead of fixed-size chunks, with no second size setting;
  • all fragments carry evidence_id, evidence_kind, purposes, scope, language and provenance;
  • fragment IDs and ordinals are deterministic;
  • only curated/**/*.md is accepted as canonical filesystem content.

Step 2: Extend the immutable corpus models

Keep CanonicalDocument and CanonicalChunk as transport-neutral storage models. Put typed Evidence metadata in their already-safe metadata field, with exact allowlisted keys. Do not make corpus storage depend on Pydantic subtype classes at read time.

Step 3: Implement kind-aware fragment rendering

Construct embedding text from title, purpose, scope and type-specific data. Example for a formula:

Formula: Fascia pediatrica
Concept: fascia pediatrica
Columns: clinical.patient.birth_date
SQL: CASE WHEN ... END
Limitations: ...

The rendered text is derived; provenance and canonical content remain in the corpus manifest.

Step 4: Run focused pipeline tests

cd harness
.venv/bin/pytest tests/test_corpus_models.py tests/test_corpus_chunk.py \
  tests/test_corpus_normalize.py tests/test_corpus_pipeline.py \
  tests/test_corpus_publish.py -q

Expected: PASS, including existing rollback, compensation, retention and resume tests.

Step 5: Commit

git add harness/tht/evidence/corpus harness/tht/evidence/preprocessing.py \
  harness/tests/test_corpus_models.py harness/tests/test_corpus_chunk.py \
  harness/tests/test_corpus_normalize.py harness/tests/test_corpus_pipeline.py
git commit -m "feat(evidence): build semantic fragments from typed units"

Task 6: Add BM25 to the existing Qdrant collection without rebuilding it

Files:

  • Modify: backend/src/workspaces/qdrant-collection.ts
  • Modify: backend/src/workspaces/evidence/preprocessing.ts
  • Modify: backend/test/qdrant-collection.test.ts
  • Modify: backend/test/workspaces/evidence/preprocessing.test.ts
  • Modify: harness/tht/adapters/vector/qdrant.py
  • Modify: harness/tests/test_qdrant_vector_store.py
  • Modify: docs/contracts/workspace-preprocessing-cli.md

Step 1: Write RED TypeScript collection-contract tests

The required vector contract is:

{
  "vectors": {"size": 1024, "distance": "Cosine"},
  "sparse_vectors": {
    "bm25": {"modifier": "idf"}
  }
}

Keep the current unnamed dense vector. Test three distinct states:

  • compatible: unnamed dense is correct and bm25 has modifier: idf;
  • Evidence-upgradeable: unnamed dense is correct and bm25 is absent;
  • incompatible: dense dimension/distance is wrong, dense is unexpectedly named, or bm25 exists with a different configuration.

Only the Evidence maintenance path may turn the upgradeable state into compatible by adding bm25. Session admission and searches remain read-only. Missing payload indexes may still use the existing additive reconciliation; no path may delete or rename a vector.

Add keyword indexes only for fields used by filters:

content_hash, document_id, kind, record_key, record_kind,
vector_generation, workspace_id, workspace_revision,
evidence_id, evidence_kind, purposes, concepts, tables, columns, language

Step 2: Run the TypeScript test and confirm RED

cd backend
npx vitest run test/qdrant-collection.test.ts

Expected: the new additive BM25 expectations fail.

Step 3: Implement additive BM25 reconciliation

Keep base vectorCompatibility concerned with the unnamed dense contract used by Schema and Memory. Add an Evidence-specific compatibility result that also classifies bm25 as compatible, upgradeable or incompatible. Update createCollection and the Evidence preprocessing preflight: a new collection is created with unnamed dense plus bm25; on an existing upgradeable collection, workspace preprocess evidence calls Qdrant's additive vector-schema endpoint to create only bm25 with IDF, then verifies the result before upload. Retain require_existing behavior for ordinary runtime validation: it never mutates. An incompatible state fails without changes. Missing BM25 does not make Schema or Memory unavailable; only Evidence reports unavailable.

Step 4: Update the Python adapter's collection validation

QdrantVectorStore.health() and _ensure_collection() must recognize exactly the same contract as TypeScript. Add a shared test fixture shape even though the two languages do not share implementation code.

Step 5: Run backend and harness tests

cd backend
npx vitest run test/qdrant-collection.test.ts test/workspaces/evidence/preprocessing.test.ts
npx tsc --noEmit -p .
cd ../harness
.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q

Expected: PASS.

Step 6: Document the additive maintenance behavior

Update the CLI contract to state that workspace preprocess evidence may add the missing bm25 definition but may not delete or rename vectors. The existing destructive workspace vector rebuild command remains available for unrelated operator recovery and is not used by this migration.

Step 7: Commit

git add backend/src/workspaces/qdrant-collection.ts \
  backend/src/workspaces/evidence/preprocessing.ts \
  backend/test/qdrant-collection.test.ts backend/test/workspaces/evidence/preprocessing.test.ts \
  harness/tht/adapters/vector/qdrant.py \
  harness/tests/test_qdrant_vector_store.py docs/contracts/workspace-preprocessing-cli.md
git commit -m "feat(evidence): add qdrant bm25 vector in place"

Task 7: Add server-side BM25 ingestion and hybrid Query API retrieval

Files:

  • Modify: harness/tht/ports/vector.py
  • Modify: harness/tht/adapters/vector/qdrant.py
  • Modify: harness/tht/vectorstore/records.py
  • Modify: harness/tht/evidence/corpus/pipeline.py
  • Modify: harness/tests/test_vector_port_contract.py
  • Modify: harness/tests/test_qdrant_vector_store.py
  • Modify: harness/tests/test_corpus_pipeline.py
  • Create: harness/tests/l0/test_qdrant_bm25_inference.py

Step 1: Write RED port tests

Extend, do not replace, the current record:

@dataclass(frozen=True)
class VectorWriteRecord:
    record: VectorRecord
    embedding: list[float]
    content_hash: str
    sparse_text: str | None = None
    sparse_language: str | None = None

Extend VectorStore.search with keyword-only query_text and query_language. Existing callers that omit them remain dense-only.

Step 2: Write exact Qdrant request tests

For Evidence upsert, assert:

"vector": {
  "": [0.1, 0.2],
  "bm25": {
    "text": "...",
    "model": "qdrant/bm25",
    "options": {"language": "italian"}
  }
}

For hybrid search, assert two filtered prefetches and default RRF:

{
  "prefetch": [
    {"query": [0.1, 0.2], "limit": 20, "filter": {}},
    {
      "query": {
        "text": "fascia pediatrica",
        "model": "qdrant/bm25",
        "options": {"language": "italian"}
      },
      "using": "bm25",
      "limit": 20,
      "filter": {}
    }
  ],
  "query": {"rrf": {}},
  "limit": 10,
  "with_payload": true
}

Both prefetch filters must include workspace, revision, active generation and record kind. Do not add hand-tuned weights.

Step 3: Prove BM25 on the actual local Qdrant image

Add an l0 test that reads the Qdrant image reference from the root compose.yaml, starts that exact image with testcontainers, creates a uniquely named temporary collection, adds bm25 with IDF, indexes two Italian texts through server-side qdrant/bm25, retrieves the expected text and deletes the collection. The test must not use FastEmbed or accept a dense fallback. It fails if the Compose reference and the tested image diverge.

Step 4: Preserve dense writes for all existing semantic records

Schema, Memory and solved-question records keep their current unnamed dense writes. Evidence records use the empty default-vector name plus bm25 in the mixed-vector upsert shape. Only records with sparse_text receive bm25; no full reindex of Schema or Memory is performed.

Step 5: Implement BM25 Evidence ingestion

Populate sparse_text and map workspace language it to Qdrant's italian. Reject an unsupported language before uploading the generation. Use the same options at ingest and query time.

Step 6: Implement hybrid search with dense fallback only for non-Evidence callers

An Evidence hybrid request must fail as unavailable if the configured Qdrant version or collection contract does not support BM25. It must not silently claim to have run hybrid search. Existing non-Evidence dense requests continue to work.

Step 7: Run focused tests

cd harness
.venv/bin/pytest tests/test_vector_port_contract.py tests/test_qdrant_vector_store.py \
  tests/test_corpus_pipeline.py tests/test_search_pack.py -q
.venv/bin/pytest tests/l0/test_qdrant_bm25_inference.py -m l0 -q
.venv/bin/ruff check tht/ports/vector.py tht/adapters/vector/qdrant.py \
  tht/evidence/corpus/pipeline.py

Expected: PASS.

Step 8: Commit

git add harness/tht/ports/vector.py harness/tht/adapters/vector/qdrant.py \
  harness/tht/vectorstore/records.py harness/tht/evidence/corpus/pipeline.py \
  harness/tests/test_vector_port_contract.py harness/tests/test_qdrant_vector_store.py \
  harness/tests/test_corpus_pipeline.py harness/tests/l0/test_qdrant_bm25_inference.py
git commit -m "feat(evidence): add qdrant bm25 hybrid retrieval"

Task 8: Add the typed Evidence search contract and workflow-owned purposes

Files:

  • Modify: harness/tht/evidence/search.py
  • Modify: harness/tht/evidence/__init__.py
  • Modify: harness/tht/evidence/session.py
  • Modify: harness/tht/cli/search_cmd.py
  • Modify: harness/tht/session/filesystem_repository.py
  • Modify: harness/tht/session/postgres_repository.py
  • Modify: harness/tests/test_evidence_facade_contract.py
  • Create: harness/tests/test_evidence_session_receipts.py
  • Modify: harness/tests/test_search_pack.py
  • Create: harness/.pi/skills/tht-sessione/modules/evidence/runtime-search.md
  • Modify: harness/.pi/skills/tht-sessione/projection.md.tmpl
  • Modify: harness/tht/pi_skill_projection.py
  • Modify: harness/tests/test_pi_skill_projection.py
  • Regenerate: harness/.pi/skills/tht-sessione/SKILL.md

Step 1: Write RED search-facade tests

Introduce:

class EvidenceSearchContext(BaseModel):
    concepts: tuple[str, ...] = ()
    tables: tuple[str, ...] = ()
    columns: tuple[str, ...] = ()
    required_kinds: tuple[EvidenceKind, ...] = ()
    required_concepts: tuple[str, ...] = ()
    required_tables: tuple[str, ...] = ()
    required_columns: tuple[str, ...] = ()

def search_evidence(
    query: str,
    purpose: EvidencePurpose,
    context: EvidenceSearchContext,
    *,
    searcher: ActiveEvidenceSearcher,
    embedder: EvidenceQueryEmbedder,
    top_n: int = 10,
) -> EvidenceSearchOutcome: ...

Test hard filters for workspace, revision, generation, purpose and every explicit required_* constraint. Concepts, tables and columns without the required_ prefix enrich the dense and BM25 query text but do not create payload filters. Do not add implicit kind or scope bonuses in v1.

Render that enrichment once, identically for both retrieval branches, in this fixed order:

Domanda: <original query>
Concetti: <normalized, deduplicated, sorted values>
Tabelle: <normalized, deduplicated, sorted values>
Colonne: <normalized, deduplicated, sorted values>

Omit empty lines and do not rewrite the original question. Test that permutations and duplicates in context produce the same rendered text and that the dense embedder and BM25 request receive exactly that same value.

The renderer normalizes the question to Unicode NFC, converts CRLF and CR to \n, strips only leading and trailing whitespace and rejects an empty result. It preserves case, punctuation and internal whitespace. Apply NFC and strip to context values, remove empty strings and exact duplicates, then sort by Unicode value. Do not call lower() or casefold() because quoted PostgreSQL identifiers can be case-sensitive.

EvidenceSearchOutcome has an available state with a generation and zero or more results, and an unavailable state with a stable error code, a bounded message and no results. memory is not an EvidencePurpose: the memory stage remains owned by the Memory Module.

Step 2: Test fragment grouping

Two returned fragments with the same evidence_id must become one EvidenceResult, with the best score, ordered matching excerpts, canonical citation, a reference for resolving the full document and no duplicate unit. Do not insert the full document into the search pack automatically.

Step 3: Distinguish an empty search from technical unavailability

Test a successful search with zero matches separately from absent ACTIVE corpus, revision mismatch, unavailable Qdrant and malformed payload. The first returns available with an empty result list and may continue. Every technical case returns unavailable, blocks the calling stage until retry, and never uses stale rows or a fallback purpose.

Step 4: Implement the facade and CLI mapping

Keep Qdrant syntax inside tht.evidence. search_cmd.py translates command inputs to the facade and renders results; it must not duplicate ranking logic.

Step 5: Integrate the contributor with semantic stages

Move the shared Evidence rules from the projection template into modules/evidence/runtime-search.md. The fragment must state:

  • candidates are not truth;
  • pass the phase-appropriate purpose;
  • show provenance;
  • formulas are kind=formula, not a separate search store;
  • an available empty result is visible and does not block the stage;
  • an unavailable outcome blocks the stage and is retriable;
  • Evidence never writes decisions, canonical artifacts or workflow state.

Map semantic stages, independently of their display codes:

clarification  -> disambiguation
rewriting      -> rewriting
schema_linking -> schema_linking
cte            -> sql_generation
final_sql      -> sql_generation

Do not invoke Evidence from memory or synthesis. Run an independent search in every mapped stage; the final_sql query includes the approved CTE plan.

Step 6: Persist the minimal Evidence receipt

Use the existing session artifact repositories to maintain one evidence_receipts.json artifact. For every available search, append or replace the receipt identified by semantic stage with exactly stage, purpose, vector_generation and ordered evidence_ids. Do not copy excerpts or complete Evidence text. The workflow integration owns the write; tht.evidence only constructs and returns the typed receipt. Test filesystem and PostgreSQL repositories, resume and retry replacement.

Register the fragment in the static FRAGMENT_ORDER and regenerate.

Step 7: Run tests

cd harness
python -m tht.pi_skill_projection --write
python -m tht.pi_skill_projection --check
.venv/bin/pytest tests/test_evidence_facade_contract.py tests/test_search_pack.py \
  tests/test_evidence_session_receipts.py \
  tests/test_pi_skill_projection.py -q

Expected: PASS and no direct edit drift in generated SKILL.md.

Step 7: Commit

git add harness/tht/evidence/search.py harness/tht/evidence/__init__.py \
  harness/tht/cli/search_cmd.py harness/tests/test_evidence_facade_contract.py \
  harness/tests/test_search_pack.py \
  harness/.pi/skills/tht-sessione/modules/evidence/runtime-search.md \
  harness/.pi/skills/tht-sessione/projection.md.tmpl \
  harness/.pi/skills/tht-sessione/SKILL.md harness/tht/pi_skill_projection.py \
  harness/tests/test_pi_skill_projection.py
git commit -m "refactor(evidence): own typed runtime retrieval"

Task 9: Migrate formulas into Curated Evidence

Files:

  • Modify: harness/tht/evidence/formula_store.py
  • Modify: harness/tht/evidence/session.py
  • Modify: harness/tht/cli/search_cmd.py
  • Create: harness/.pi/skills/tht-sessione/modules/evidence/formula-proposals.md
  • Modify: harness/.pi/skills/tht-sessione/projection.md.tmpl
  • Modify: harness/tht/pi_skill_projection.py
  • Modify: harness/tests/test_formula.py
  • Modify: harness/tests/test_formula_wiring.py
  • Create: harness/tests/test_evidence_formula_migration.py
  • Modify: harness/tests/test_pi_skill_projection.py
  • Regenerate: harness/.pi/skills/tht-sessione/SKILL.md

Step 1: Write RED migration tests

Map an approved legacy ConceptFormula to a Curated Evidence formula while preserving:

  • concept;
  • SQL;
  • columns;
  • sources as provenance notes;
  • stable deterministic ID;
  • reviewed content wording.

Reject auto and unresolved draft formulas from direct publication; they become Formula proposals.

Step 2: Define the session proposal contract

Add a schema-versioned formula-proposal projection in the session artifact. Keep concept_formula_approved / concept_formula_rejected decisions unchanged because they record a session-local choice, not repository publication.

Step 3: Remove the separate runtime formula lookup

Change tht search find --kind formula to call typed Evidence search with required_kinds=("formula",). Keep the legacy store readable only for the migration command/window, with a deprecation warning in human output and no warning leakage into pristine JSON.

Step 4: Add and project formula instructions

The Evidence fragment must say that a newly synthesized formula is a session proposal and cannot be treated as Published Evidence.

Step 5: Run tests

cd harness
python -m tht.pi_skill_projection --write
.venv/bin/pytest tests/test_formula.py tests/test_formula_wiring.py \
  tests/test_evidence_formula_migration.py tests/test_pi_skill_projection.py \
  tests/test_decision_min_phase.py tests/test_workflow_observable_contract.py -q

Expected: PASS; decision phase ownership remains F4.

Step 6: Commit

git add harness/tht/evidence/formula_store.py harness/tht/evidence/session.py \
  harness/tht/cli/search_cmd.py \
  harness/.pi/skills/tht-sessione/modules/evidence/formula-proposals.md \
  harness/.pi/skills/tht-sessione/projection.md.tmpl \
  harness/.pi/skills/tht-sessione/SKILL.md harness/tht/pi_skill_projection.py \
  harness/tests/test_formula.py harness/tests/test_formula_wiring.py \
  harness/tests/test_evidence_formula_migration.py harness/tests/test_pi_skill_projection.py
git commit -m "refactor(evidence): unify formulas with typed evidence"

Task 10: Add the small retrieval evaluation command

Files:

  • Create: harness/tht/evidence/evaluation.py
  • Modify: harness/tht/cli/evidence_cmd.py
  • Create: harness/tests/test_evidence_evaluation.py
  • Modify: harness/tests/test_evidence_cli.py

Step 1: Write RED schema tests

Use a deliberately small format:

schema_version: 1
queries:
  - id: pediatric-formula
    query: Come distinguo i pazienti pediatrici?
    profile: semantic
    purpose: sql_generation
    expected:
      - evidence:fascia-pediatrica

Require unique query IDs, nonempty expected IDs, one of lexical, semantic or mixed for every profile, only public purpose values and at least one query of every profile in the complete file.

Step 2: Write RED metric tests

Compute hit_at_5, hit_at_10, missing expected IDs, empty-result queries and counts by expected kind. For each expected ID, also report its nullable dense-only, BM25-only and fused rank. The evaluator runs both branches separately for diagnosis and the same hybrid request used by runtime. The report passes only when every query finds at least one expected ID in its first ten fused results; branch ranks and hit_at_5 are informative. Do not add nDCG, relevance grading or an evaluation database in v1.

Step 3: Implement the evaluator and CLI

Expose:

tht evidence evaluate <workspace-root> -c <runtime-config>
  [--generation <id>] [--json]

The report includes workspace revision, evaluated vector generation and the fixed default RRF configuration. It is read-only. With no --generation, it evaluates the active generation for monitoring. Before publication, the preprocessing pipeline calls the same evaluator against its candidate generation and atomically activates it only when the report passes. A failed candidate remains inactive.

Step 4: Run tests

cd harness
.venv/bin/pytest tests/test_evidence_evaluation.py tests/test_evidence_cli.py -q
.venv/bin/ruff check tht/evidence/evaluation.py tests/test_evidence_evaluation.py

Expected: PASS.

Step 5: Commit

git add harness/tht/evidence/evaluation.py harness/tht/cli/evidence_cmd.py \
  harness/tests/test_evidence_evaluation.py harness/tests/test_evidence_cli.py
git commit -m "feat(evidence): evaluate retrieval with a small fixture"

Task 11: Enforce curated-only runtime ingestion and update contracts

Files:

  • Modify: backend/src/workspaces/schema.ts
  • Modify: backend/src/workspaces/runtime-renderer.ts
  • Modify: backend/test/workspaces/schema.test.ts
  • Modify: backend/test/workspaces/runtime-renderer.test.ts
  • Modify: docs/contracts/workspace-evidence-v3.md
  • Modify: docs/contracts/workspace-preprocessing-cli.md
  • Modify: docs/architecture/overview.md
  • Modify: PROJECT_STATE.md

Step 1: Write RED descriptor/rendering tests

For filesystem Evidence, make patterns: ["curated/**/*.md"] the documented and generated default for the new authoring layout. Continue accepting an explicitly configured safe pattern for non-Git HTTP/S3 compatibility, but reject a filesystem descriptor that includes both source/** and curated/** once it declares the new layout version.

If a schema-version field is required to preserve compatibility, add it to the Evidence subcontract, not to the whole workspace descriptor.

Step 2: Implement the narrowest compatible descriptor change

The materializer continues to copy the entire evidence/ tree at the pinned commit. Only the rendered runtime acquisition patterns restrict preprocessing to curated/. Do not duplicate or move P6 materialization logic.

Step 3: Update documentation contracts

Document:

  • source/curated layout;
  • Git publication boundary;
  • no runtime writes;
  • validation before indexing;
  • unnamed-dense plus BM25 contract and additive Evidence upgrade;
  • exact public operation names and JSON status additions, if any.

Step 4: Run backend and harness contract gates

cd backend
npx vitest run test/workspaces/schema.test.ts test/workspaces/runtime-renderer.test.ts \
  test/workspaces/evidence/materialization.test.ts \
  test/workspaces/evidence/preprocessing.test.ts
npx tsc --noEmit -p .
cd ../harness
.venv/bin/pytest tests/test_registry_evidence_config.py \
  tests/test_filesystem_evidence_source.py tests/test_preprocess_cli.py -q

Expected: PASS; P6 materialization safety remains unchanged.

Step 5: Commit

git add backend/src/workspaces/schema.ts backend/src/workspaces/runtime-renderer.ts \
  backend/test/workspaces/schema.test.ts \
  backend/test/workspaces/runtime-renderer.test.ts \
  docs/contracts/workspace-evidence-v3.md \
  docs/contracts/workspace-preprocessing-cli.md docs/architecture/overview.md PROJECT_STATE.md
git commit -m "docs(evidence): publish curated-only workspace contract"

Task 12: Migrate PSD and perform acceptance

Files in ThothII:

  • Create: docs/testing/evidence-restructuring-manual.md
  • Create: scripts/evidence-restructuring-acceptance.sh
  • Create: harness/tests/fixtures/evidence_authoring/poorly_structured.md
  • Modify: PROJECT_STATE.md

Files in the external authoring repository:

  • Move: /Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/<current-folders> to /Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/source/
  • Create: /Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/curated/<kind>/
  • Create: /Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/manifest.yaml
  • Create: /Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/evaluation.yaml
  • Modify: /Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/README.md
  • Modify: /Users/mp/projects/tht-workspace-psd/psd-clinical/workspace.yaml

Do not modify the external repository until the owner confirms the migration window and the exact target branch. Treat that as the only manual authorization gate in this task.

Step 1: Add a hermetic badly-structured fixture

The fixture must contain prose, a rough list, an enum, an URL, an ambiguous statement and a SQL formula candidate. The acceptance runner must prove:

  • split into multiple typed units;
  • ambiguity becomes review_items;
  • no cross-source merge;
  • unchanged rerun is a no-op;
  • committed human content remains recoverable and proposed changes are visible in Git;
  • dirty-tree refusal;
  • validation blocks unresolved review and orphaned units;
  • pipeline-version mismatch refusal and explicit full-corpus --upgrade;
  • validated corpus builds an inactive candidate and searches it hybrid;
  • every evaluation query retrieves at least one expected ID in the first ten results before the candidate becomes active;
  • the evaluation set contains lexical, semantic and mixed cases and reports dense, BM25 and fused ranks separately;
  • dense and BM25 receive the same deterministic query text;
  • no atomic formula, enum pair, mapping, rule or URL is split into fixed-size chunks;
  • the complete rendered fragment stays within the existing 4,000-character max_chunk_chars default and no parallel size option exists;
  • the L0 probe passes against the Qdrant image referenced by compose.yaml without FastEmbed or fallback;
  • query normalization is limited to NFC, newline canonicalization and outer trimming, preserving internal whitespace, punctuation and case-sensitive identifiers;
  • Formula Evidence accepts a PostgreSQL expression and rejects a complete query;
  • Qdrant failure remains fail-closed.

Use a fake restructurer for hermetic CI. The real Pi call is a separate manual check.

Step 2: Run the complete automated gates before external writes

cd harness
.venv/bin/pytest -q
.venv/bin/ruff check .
cd ../backend
npx vitest run
npx tsc --noEmit -p .
cd ../tools/tht
go test ./...
go build ./cmd/tht
cd ../..
bash scripts/evidence-restructuring-acceptance.sh

Expected: all suites and acceptance checks PASS. If pre-existing unrelated failures remain, record exact names and prove they reproduce at the baseline commit before continuing.

Step 3: Stop for the owner migration gate

Provide:

  • clean ThothII commit;
  • test and acceptance summary;
  • proposed PSD branch name;
  • exact list of 36 source files to move;
  • rollback command based on the pre-migration PSD commit;
  • before/after Schema and Memory counts plus sample IDs used to prove non-regression.

Do not infer approval from prior design acceptance.

Step 4: Migrate the PSD repository after approval

Use tht evidence prepare, review the Git diff, resolve all review items manually, run tht evidence validate, and create the approximately twenty evaluation queries. Do not auto-merge or auto-push unless separately requested.

Step 5: Activate and perform the additive BM25 upgrade

After the PSD merge/pull and activation, inspect first:

tht --installation <absolute>/thothii-installation.yaml workspace vector inspect
  --workspace psd-clinical --json

Before preprocessing, record counts and representative IDs for schema_table, schema_column, memory and solved_question. Run workspace preprocess evidence: it adds bm25 when absent, builds the candidate, evaluates that exact generation and publishes it only if the report passes. Repeat the counts, verify the representative IDs and run dense smoke searches for Schema and Memory before completing the manual walkthrough of clarification, rewriting, schema_linking, cte and final_sql.

Step 6: Record acceptance and commit ThothII documentation

docs/testing/evidence-restructuring-manual.md must record separate outcomes for:

  • authoring and Git review;
  • additive BM25 schema upgrade;
  • Schema and Memory before/after non-regression evidence;
  • preprocessing generation publication;
  • hybrid retrieval evaluation;
  • formula retrieval;
  • distinzione fra risultato vuoto e indisponibilità bloccante;
  • complete session behavior.

Update PROJECT_STATE.md only with observed results and immutable commit/run IDs.

git add docs/testing/evidence-restructuring-manual.md \
  scripts/evidence-restructuring-acceptance.sh \
  harness/tests/fixtures/evidence_authoring/poorly_structured.md PROJECT_STATE.md
git commit -m "test(evidence): record restructuring acceptance"

Final verification checklist

Before claiming completion, verify:

  • CONTEXT.md and both Evidence plan documents use the same terminology.
  • All eight Evidence kinds have type-specific positive and negative tests.
  • Pi is called once per changed source, with no tools and no saved session.
  • prepare refuses dirty curated state, exposes every proposal in Git and never deletes orphans.
  • validate blocks unresolved review items and orphaned units.
  • Source renames and kind-only reclassifications preserve unit IDs; semantic splits receive new IDs.
  • IDs use evidence:<slug>, remain independent from kind and are not recomputed after their initial assignment.
  • Formula Evidence accepts one PostgreSQL expression and rejects full queries.
  • A pipeline-version mismatch performs no writes without explicit --upgrade.
  • Runtime reads only curated/**/*.md from the pinned Git revision.
  • The TypeScript and Python Qdrant compatibility checks agree.
  • Qdrant retains the unnamed dense vector and adds only bm25 with IDF.
  • Evidence preprocessing adds a missing bm25 definition but never deletes, renames or destructively rebuilds collection vectors.
  • Schema and Memory counts, sample IDs and dense searches remain unchanged across the additive upgrade.
  • Italian BM25 options are identical during ingest and query.
  • Hybrid search uses two prefetches and default RRF.
  • Hard filters always include workspace, revision, active generation and purpose; only explicit required_* context values add further filters.
  • Formula runtime lookup uses typed Evidence; session proposals remain non-published.
  • Fragment hits are grouped into one Evidence Result with excerpts and a reference; full documents are loaded only on demand.
  • Evidence is queried independently from the five mapped semantic stages and is not invoked from memory or synthesis.
  • An available empty result can continue; technical unavailability blocks the current stage without stale-generation or purpose fallback.
  • Each available stage search persists only its minimal Evidence receipt.
  • Evaluation reports hit@5 and hit@10 against a versioned fixture and passes only when every query has at least one expected result in the first ten.
  • Evaluation runs against the candidate generation before atomic activation; a failed candidate remains invisible to sessions.
  • Existing corpus rollback, compensation, retention and resume tests still pass.
  • Backend Vitest and TypeScript gates pass.
  • Harness pytest and Ruff gates pass.
  • Native tht Go tests/build pass.
  • External PSD writes occurred only after explicit migration authorization.