Files
ThothII/docs/plans/2026-08-24-evidence-restructuring.md
T

1124 lines
36 KiB
Markdown

# Evidence Restructuring Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
**Goal:** Build a Git-reviewed, typed Evidence authoring pipeline and publish its approved output to the existing revision-scoped Qdrant lifecycle with dense+BM25 hybrid retrieval.
**Architecture:** Keep authoring outside NL→SQL sessions: deterministic preparation wraps one read-only Pi restructuring call, writes reviewable Markdown into the workspace repository, and blocks publication on unresolved review items. Reuse the current Evidence module, corpus generations, rollback, active-revision checks, and workspace-owned semantic collection; extend them instead of introducing a parallel store.
**Tech Stack:** Python 3.12, Pydantic 2, Typer, PyYAML, sqlglot, pytest, Pi CLI, TypeScript, Fastify workspace maintenance, Vitest, Qdrant 1.18.2 Query API, server-side `qdrant/bm25`, Git.
---
## Preconditions
- Work in this dedicated worktree.
- Read `PROJECT_STATE.md`, `CONTEXT.md`,
`docs/plans/2026-08-24-evidence-restructuring-design.md`,
`docs/contracts/workspace-evidence-v3.md`, and
`docs/contracts/workspace-preprocessing-cli.md`.
- Preserve the public facade in `harness/tht/evidence/__init__.py`.
- Do not edit `harness/.pi/skills/tht-sessione/SKILL.md` directly; regenerate it with
`python -m tht.pi_skill_projection --write`.
- Keep runtime workspace access read-only. Only the authoring CLI may write
`evidence/curated/` and `evidence/manifest.yaml` in a curator clone.
- Do not migrate the external PSD repository until all code and contract gates pass.
### Task 1: Add the typed Curated Evidence model
**Files:**
- Create: `harness/tht/evidence/canonical.py`
- Modify: `harness/tht/evidence/__init__.py`
- Create: `harness/tests/test_evidence_canonical.py`
**Step 1: Write failing tests for the common envelope**
Cover:
- all eight `kind` values;
- all five `purpose` values;
- immutable provenance;
- strict unknown-field rejection;
- namespaced stable IDs;
- SHA-256 syntax;
- directory/`kind` agreement;
- parsing and dumping Markdown with YAML frontmatter.
Start with:
```python
def test_formula_requires_formula_payload():
with pytest.raises(ValidationError):
CuratedEvidence.model_validate({
**COMMON,
"kind": "formula",
"payload": {"concept": "fascia pediatrica"},
})
def test_reference_rejects_formula_payload():
with pytest.raises(ValidationError):
CuratedEvidence.model_validate({
**COMMON,
"kind": "reference",
"payload": {
"concept": "x",
"columns": ["clinical.patient.birth_date"],
"sql": "CASE WHEN true THEN 1 END",
},
})
```
**Step 2: Run the focused tests and confirm RED**
Run:
```bash
cd harness
.venv/bin/pytest tests/test_evidence_canonical.py -q
```
Expected: import failure for `tht.evidence.canonical`.
**Step 3: Implement the discriminated model**
Use a strict Pydantic model with these public types:
```python
EvidenceKind = Literal[
"glossary", "domain", "enum", "example", "mapping",
"normalization", "formula", "reference",
]
EvidencePurpose = Literal[
"disambiguation", "rewriting", "schema_linking", "sql_generation", "memory",
]
class EvidenceScope(StrictModel):
concepts: tuple[str, ...] = ()
tables: tuple[str, ...] = ()
columns: tuple[str, ...] = ()
class EvidenceProvenance(StrictModel):
source_file: str
source_sha256: str
class FormulaPayload(StrictModel):
concept: str
columns: tuple[str, ...]
sql: str
class ReferencePayload(StrictModel):
url: AnyHttpUrl
label: str
description: str
```
Define equally strict payloads for the other six kinds and expose a single
`CuratedEvidence` API. The implementation may use an internal Pydantic discriminated
union, but callers must not switch between eight unrelated loaders.
Add:
```python
def parse_curated_markdown(text: str, *, path: Path | None = None) -> CuratedEvidence: ...
def dump_curated_markdown(value: CuratedEvidence) -> str: ...
def load_curated_tree(root: Path) -> list[CuratedEvidence]: ...
```
Validate formula SQL with `sqlglot` and validate `schema.table` /
`schema.table.column` identifiers without querying the DWH.
**Step 4: Export only the stable API**
Add the model and loader functions to `harness/tht/evidence/__init__.py`. Do not export
internal union member helpers unless another module needs them.
**Step 5: Run tests and lint**
Run:
```bash
cd harness
.venv/bin/pytest tests/test_evidence_canonical.py tests/test_evidence_facade_contract.py -q
.venv/bin/ruff check tht/evidence/canonical.py tests/test_evidence_canonical.py
```
Expected: PASS.
**Step 6: Commit**
```bash
git add harness/tht/evidence/canonical.py harness/tht/evidence/__init__.py \
harness/tests/test_evidence_canonical.py
git commit -m "feat(evidence): add typed curated evidence model"
```
### Task 2: Add deterministic validation and the Evidence manifest
**Files:**
- Create: `harness/tht/evidence/authoring.py`
- Create: `harness/tests/test_evidence_authoring.py`
- Modify: `harness/tht/evidence/__init__.py`
**Step 1: Write failing tests for publication validation**
Test that:
- `review_items != []` is valid as a draft but blocks publication;
- a missing or mismatched source hash blocks publication;
- two units cannot share an ID;
- one unit cannot claim two source files;
- `curated/formula/x.md` must contain `kind: formula`;
- credentials in URLs, YAML, or body are rejected;
- only Markdown, `.txt`, and `.sql.md` sources are accepted;
- UTF-8 and per-file limits are enforced.
Expose errors as bounded structured values:
```python
@dataclass(frozen=True)
class ValidationFinding:
severity: Literal["error", "warning"]
code: str
path: str
message: str
```
**Step 2: Write failing tests for the versioned manifest**
Use this minimum shape:
```yaml
schema_version: 1
pipeline_version: evidence-authoring-v1
sources:
source/domain/patient.md:
sha256: sha256:...
units:
- domain:patient
orphans: []
```
Test deterministic key ordering, stable round-trip, unknown fields, duplicate unit IDs,
and orphan preservation.
**Step 3: Run and confirm RED**
Run:
```bash
cd harness
.venv/bin/pytest tests/test_evidence_authoring.py -q
```
Expected: missing authoring API.
**Step 4: Implement `EvidenceManifest` and `validate_workspace_evidence`**
Provide:
```python
def load_manifest(path: Path) -> EvidenceManifest: ...
def dump_manifest(manifest: EvidenceManifest) -> str: ...
def validate_workspace_evidence(workspace_root: Path) -> ValidationReport: ...
```
`ValidationReport.publishable` is true only when there are no errors and no unresolved
review items. Warnings alone do not block publication.
Do not use status fields such as `draft/reviewed` as an approval mechanism. The approved
Git revision is the publication boundary.
**Step 5: Run tests and lint**
```bash
cd harness
.venv/bin/pytest tests/test_evidence_authoring.py tests/test_evidence_canonical.py -q
.venv/bin/ruff check tht/evidence/authoring.py tests/test_evidence_authoring.py
```
Expected: PASS.
**Step 6: Commit**
```bash
git add harness/tht/evidence/authoring.py harness/tht/evidence/__init__.py \
harness/tests/test_evidence_authoring.py
git commit -m "feat(evidence): validate curated corpus and manifest"
```
### Task 3: Implement incremental preparation with one Pi restructuring call
**Files:**
- Modify: `harness/tht/evidence/authoring.py`
- Create: `harness/.pi/skills/tht-evidence-authoring/SKILL.md`
- Modify: `harness/tests/test_evidence_authoring.py`
- Create: `harness/tests/test_evidence_pi_restructurer.py`
**Step 1: Define the restructuring port and request/response models**
Add:
```python
class EvidenceRestructurer(Protocol):
def restructure(self, request: RestructureRequest) -> tuple[CuratedEvidence, ...]: ...
class RestructureRequest(StrictModel):
source_file: str
source_sha256: str
normalized_text: str
previous_units: tuple[CuratedEvidence, ...] = ()
```
The response is validated through the models from Task 1 before any write.
**Step 2: Write RED tests for the incremental rules**
Test:
- unchanged source: no model call and no file write;
- changed source: exactly one model call;
- new source: new stable IDs;
- removed source: old units become orphans and remain on disk;
- one source may produce several units;
- no returned unit may cite another source;
- prior curated units are included in the request;
- a dirty `evidence/curated/` or `evidence/manifest.yaml` fails before the model call;
- all output writes are staged and atomically replaced only after complete validation.
Inject the Git status runner and filesystem writer in tests; do not require a real Git
repository for every unit test.
**Step 3: Implement deterministic source normalization**
Normalize UTF-8 text with NFC, LF newlines and terminal newline. Preserve meaningful
Markdown, tables, fenced SQL, URLs and list structure. Do not rewrite vocabulary or
infer domain facts in this step.
**Step 4: Implement `PiEvidenceRestructurer`**
Invoke Pi as an ephemeral, no-tools process using an argument list, never a shell:
```python
argv = [
pi_executable,
"--mode", "text",
"--print",
"--no-session",
"--no-tools",
"--no-extensions",
"--no-context-files",
"--skill", str(skill_path),
f"@{request_path}",
"Return only the JSON object required by the Evidence authoring skill.",
]
```
Use a private `TemporaryDirectory`, a bounded timeout, bounded stdout/stderr, and strict
JSON parsing. Do not forward the model's raw output to public JSON errors. The skill must
state:
- use only facts present in `normalized_text`;
- preserve prior reviewed wording where it remains supported;
- never merge sources;
- emit `review_items` for uncertainty;
- emit exactly the schema-versioned JSON object and no Markdown fence.
**Step 5: Implement `prepare_workspace_evidence`**
Return a bounded report with `changed`, `unchanged`, `created`, `orphaned`, `findings`,
and `model_calls`. Keep the model implementation behind `EvidenceRestructurer`.
**Step 6: Run tests**
```bash
cd harness
.venv/bin/pytest tests/test_evidence_authoring.py tests/test_evidence_pi_restructurer.py -q
.venv/bin/ruff check tht/evidence/authoring.py tests/test_evidence_pi_restructurer.py
```
Expected: PASS with no live model call.
**Step 7: Commit**
```bash
git add harness/tht/evidence/authoring.py \
harness/.pi/skills/tht-evidence-authoring/SKILL.md \
harness/tests/test_evidence_authoring.py harness/tests/test_evidence_pi_restructurer.py
git commit -m "feat(evidence): prepare curated evidence incrementally"
```
### Task 4: Add the authoring CLI
**Files:**
- Create: `harness/tht/cli/evidence_cmd.py`
- Modify: `harness/tht/cli/__init__.py`
- Create: `harness/tests/test_evidence_cli.py`
**Step 1: Write CLI grammar tests**
Cover:
```text
tht evidence prepare <workspace-root> [--json]
tht evidence validate <workspace-root> [--json]
```
Require an existing canonical Git worktree root. Reject unknown flags, symlinks, a
workspace path outside the Git root, duplicate options and dirty curated state. Ensure
`--json` writes pristine JSON to stdout.
**Step 2: Run and confirm RED**
```bash
cd harness
.venv/bin/pytest tests/test_evidence_cli.py -q
```
Expected: `evidence` command group is unknown.
**Step 3: Implement the Typer group**
Use one top-level authoring group:
```python
evidence_app = typer.Typer(help="Prepare and validate workspace Evidence")
@evidence_app.command("prepare")
def prepare_cmd(workspace_root: Path, json_output: bool = False) -> None: ...
@evidence_app.command("validate")
def validate_cmd(workspace_root: Path, json_output: bool = False) -> None: ...
```
Register it in `harness/tht/cli/__init__.py`. Keep this separate from the existing
runtime command `tht preprocess evidence`.
Exit codes:
- `0`: prepared/unchanged or valid;
- `2`: unsafe path or CLI misuse;
- `3`: valid drafts but review required;
- `1`: operational/model/structural failure.
**Step 4: Run tests and CLI help**
```bash
cd harness
.venv/bin/pytest tests/test_evidence_cli.py tests/test_preprocess_cli.py -q
.venv/bin/tht evidence --help
.venv/bin/tht preprocess evidence --help
```
Expected: both command families are present and unambiguous.
**Step 5: Commit**
```bash
git add harness/tht/cli/evidence_cmd.py harness/tht/cli/__init__.py \
harness/tests/test_evidence_cli.py
git commit -m "feat(evidence): expose prepare and validate commands"
```
### Task 5: Project typed units into semantic Evidence Fragments
**Files:**
- Modify: `harness/tht/evidence/corpus/models.py`
- Modify: `harness/tht/evidence/corpus/chunk.py`
- Modify: `harness/tht/evidence/corpus/normalize.py`
- Modify: `harness/tht/evidence/corpus/pipeline.py`
- Modify: `harness/tht/evidence/preprocessing.py`
- Modify: `harness/tests/test_corpus_models.py`
- Modify: `harness/tests/test_corpus_chunk.py`
- Modify: `harness/tests/test_corpus_pipeline.py`
**Step 1: Write failing projection tests**
Test that:
- a short formula remains one fragment;
- a domain document splits only at semantic section boundaries;
- enum entries are not split in the middle of a value/meaning pair;
- fixed-size fallback is used only for a single oversized section;
- all fragments carry `evidence_id`, `evidence_kind`, `purposes`, scope, language and
provenance;
- fragment IDs and ordinals are deterministic;
- only `curated/**/*.md` is accepted as canonical filesystem content.
**Step 2: Extend the immutable corpus models**
Keep `CanonicalDocument` and `CanonicalChunk` as transport-neutral storage models. Put
typed Evidence metadata in their already-safe `metadata` field, with exact allowlisted
keys. Do not make corpus storage depend on Pydantic subtype classes at read time.
**Step 3: Implement kind-aware fragment rendering**
Construct embedding text from title, purpose, scope and type-specific data. Example for
a formula:
```text
Formula: Fascia pediatrica
Concept: fascia pediatrica
Columns: clinical.patient.birth_date
SQL: CASE WHEN ... END
Limitations: ...
```
The rendered text is derived; provenance and canonical content remain in the corpus
manifest.
**Step 4: Run focused pipeline tests**
```bash
cd harness
.venv/bin/pytest tests/test_corpus_models.py tests/test_corpus_chunk.py \
tests/test_corpus_normalize.py tests/test_corpus_pipeline.py \
tests/test_corpus_publish.py -q
```
Expected: PASS, including existing rollback, compensation, retention and resume tests.
**Step 5: Commit**
```bash
git add harness/tht/evidence/corpus harness/tht/evidence/preprocessing.py \
harness/tests/test_corpus_models.py harness/tests/test_corpus_chunk.py \
harness/tests/test_corpus_normalize.py harness/tests/test_corpus_pipeline.py
git commit -m "feat(evidence): build semantic fragments from typed units"
```
### Task 6: Upgrade the Qdrant collection contract to named dense and BM25 vectors
**Files:**
- Modify: `backend/src/workspaces/qdrant-collection.ts`
- Modify: `backend/test/qdrant-collection.test.ts`
- Modify: `harness/tht/adapters/vector/qdrant.py`
- Modify: `harness/tests/test_qdrant_vector_store.py`
- Modify: `docs/contracts/workspace-preprocessing-cli.md`
**Step 1: Write RED TypeScript collection-contract tests**
The required vector contract is:
```json
{
"vectors": {
"dense": {"size": 1024, "distance": "Cosine"}
},
"sparse_vectors": {
"bm25": {"modifier": "idf"}
}
}
```
Test that unnamed dense-only, wrong named dimension/distance, missing BM25, or wrong BM25
modifier are incompatible. Self-heal may add missing payload indexes, but must not
silently convert an incompatible vector configuration.
Add keyword indexes only for fields used by filters:
```text
content_hash, document_id, kind, record_key, record_kind,
vector_generation, workspace_id, workspace_revision,
evidence_id, evidence_kind, purposes, concepts, tables, columns, language
```
**Step 2: Run the TypeScript test and confirm RED**
```bash
cd backend
npx vitest run test/qdrant-collection.test.ts
```
Expected: current unnamed-vector expectations fail.
**Step 3: Implement strict named-vector reconciliation**
Update `createCollection`, `vectorCompatibility` and payload-index reconciliation.
Retain `require_existing` behavior: validation never performs an incompatible migration.
**Step 4: Update the Python adapter's collection validation**
`QdrantVectorStore.health()` and `_ensure_collection()` must recognize exactly the same
contract as TypeScript. Add a shared test fixture shape even though the two languages do
not share implementation code.
**Step 5: Run backend and harness tests**
```bash
cd backend
npx vitest run test/qdrant-collection.test.ts
npx tsc --noEmit -p .
cd ../harness
.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
```
Expected: PASS.
**Step 6: Document the required guarded rebuild**
Update the CLI contract to say that the vector-shape change is incompatible and must be
applied with the existing exact-name guarded command:
```text
tht --installation <absolute>/thothii-installation.yaml workspace vector rebuild
--workspace <id> --collection <name> --confirm <name> --destroy
```
No automatic deletion is allowed.
**Step 7: Commit**
```bash
git add backend/src/workspaces/qdrant-collection.ts \
backend/test/qdrant-collection.test.ts harness/tht/adapters/vector/qdrant.py \
harness/tests/test_qdrant_vector_store.py docs/contracts/workspace-preprocessing-cli.md
git commit -m "feat(evidence): require dense and bm25 qdrant vectors"
```
### Task 7: Add server-side BM25 ingestion and hybrid Query API retrieval
**Files:**
- Modify: `harness/tht/ports/vector.py`
- Modify: `harness/tht/adapters/vector/qdrant.py`
- Modify: `harness/tht/vectorstore/records.py`
- Modify: `harness/tht/evidence/corpus/pipeline.py`
- Modify: `harness/tests/test_vector_port_contract.py`
- Modify: `harness/tests/test_qdrant_vector_store.py`
- Modify: `harness/tests/test_corpus_pipeline.py`
**Step 1: Write RED port tests**
Extend, do not replace, the current record:
```python
@dataclass(frozen=True)
class VectorWriteRecord:
record: VectorRecord
embedding: list[float]
content_hash: str
sparse_text: str | None = None
sparse_language: str | None = None
```
Extend `VectorStore.search` with keyword-only `query_text` and `query_language`. Existing
callers that omit them remain dense-only.
**Step 2: Write exact Qdrant request tests**
For Evidence upsert, assert:
```json
"vector": {
"dense": [0.1, 0.2],
"bm25": {
"text": "...",
"model": "qdrant/bm25",
"options": {"language": "italian"}
}
}
```
For hybrid search, assert two filtered prefetches and default RRF:
```json
{
"prefetch": [
{"query": [0.1, 0.2], "using": "dense", "limit": 20, "filter": {}},
{
"query": {
"text": "fascia pediatrica",
"model": "qdrant/bm25",
"options": {"language": "italian"}
},
"using": "bm25",
"limit": 20,
"filter": {}
}
],
"query": {"rrf": {}},
"limit": 10,
"with_payload": true
}
```
Both prefetch filters must include workspace, revision, active generation and record
kind. Do not add hand-tuned weights.
**Step 3: Implement named dense writes for all semantic records**
All existing schema/memory/Evidence records use the `dense` vector name after the
collection rebuild. Only records with `sparse_text` receive `bm25`.
**Step 4: Implement BM25 Evidence ingestion**
Populate `sparse_text` and map workspace language `it` to Qdrant's `italian`. Reject an
unsupported language before uploading the generation. Use the same options at ingest
and query time.
**Step 5: Implement hybrid search with dense fallback only for non-Evidence callers**
An Evidence hybrid request must fail as unavailable if the configured Qdrant version or
collection contract does not support BM25. It must not silently claim to have run hybrid
search. Existing non-Evidence dense requests continue to work.
**Step 6: Run focused tests**
```bash
cd harness
.venv/bin/pytest tests/test_vector_port_contract.py tests/test_qdrant_vector_store.py \
tests/test_corpus_pipeline.py tests/test_search_pack.py -q
.venv/bin/ruff check tht/ports/vector.py tht/adapters/vector/qdrant.py \
tht/evidence/corpus/pipeline.py
```
Expected: PASS.
**Step 7: Commit**
```bash
git add harness/tht/ports/vector.py harness/tht/adapters/vector/qdrant.py \
harness/tht/vectorstore/records.py harness/tht/evidence/corpus/pipeline.py \
harness/tests/test_vector_port_contract.py harness/tests/test_qdrant_vector_store.py \
harness/tests/test_corpus_pipeline.py
git commit -m "feat(evidence): add qdrant bm25 hybrid retrieval"
```
### Task 8: Add the typed Evidence search contract and workflow-owned purposes
**Files:**
- Modify: `harness/tht/evidence/search.py`
- Modify: `harness/tht/evidence/__init__.py`
- Modify: `harness/tht/cli/search_cmd.py`
- Modify: `harness/tests/test_evidence_facade_contract.py`
- Modify: `harness/tests/test_search_pack.py`
- Create: `harness/.pi/skills/tht-sessione/modules/evidence/runtime-search.md`
- Modify: `harness/.pi/skills/tht-sessione/projection.md.tmpl`
- Modify: `harness/tht/pi_skill_projection.py`
- Modify: `harness/tests/test_pi_skill_projection.py`
- Regenerate: `harness/.pi/skills/tht-sessione/SKILL.md`
**Step 1: Write RED search-facade tests**
Introduce:
```python
class EvidenceSearchContext(BaseModel):
concepts: tuple[str, ...] = ()
tables: tuple[str, ...] = ()
columns: tuple[str, ...] = ()
required_kinds: tuple[EvidenceKind, ...] = ()
def search_evidence(
query: str,
purpose: EvidencePurpose,
context: EvidenceSearchContext,
*,
searcher: ActiveEvidenceSearcher,
embedder: EvidenceQueryEmbedder,
top_n: int = 10,
) -> list[EvidenceResult]: ...
```
Test hard filters for workspace/revision/generation and explicit `required_kinds`. Test
that purpose/kind/scope preferences are deterministic tie-breakers when they were not
requested as hard filters.
**Step 2: Test fragment grouping**
Two returned fragments with the same `evidence_id` must become one `EvidenceResult`,
with the best score, ordered matching excerpts, canonical citation and no duplicate
unit.
**Step 3: Preserve fail-closed graceful degradation**
Test absent ACTIVE corpus, revision mismatch, unavailable Qdrant and malformed payload.
All return no Evidence plus a bounded warning through the existing search-pack contract;
none uses stale rows.
**Step 4: Implement the facade and CLI mapping**
Keep Qdrant syntax inside `tht.evidence`. `search_cmd.py` translates command inputs to
the facade and renders results; it must not duplicate ranking logic.
**Step 5: Extract Evidence instructions into a module fragment**
Move the common F1/F3/F4 Evidence rules from the projection template into
`modules/evidence/runtime-search.md`. The fragment must state:
- candidates are not truth;
- pass the phase-appropriate purpose;
- show provenance;
- formulas are `kind=formula`, not a separate search store;
- absence of Evidence is visible but does not stop the whole session.
Register the fragment in the static `FRAGMENT_ORDER` and regenerate.
**Step 6: Run tests**
```bash
cd harness
python -m tht.pi_skill_projection --write
python -m tht.pi_skill_projection --check
.venv/bin/pytest tests/test_evidence_facade_contract.py tests/test_search_pack.py \
tests/test_pi_skill_projection.py -q
```
Expected: PASS and no direct edit drift in generated `SKILL.md`.
**Step 7: Commit**
```bash
git add harness/tht/evidence/search.py harness/tht/evidence/__init__.py \
harness/tht/cli/search_cmd.py harness/tests/test_evidence_facade_contract.py \
harness/tests/test_search_pack.py \
harness/.pi/skills/tht-sessione/modules/evidence/runtime-search.md \
harness/.pi/skills/tht-sessione/projection.md.tmpl \
harness/.pi/skills/tht-sessione/SKILL.md harness/tht/pi_skill_projection.py \
harness/tests/test_pi_skill_projection.py
git commit -m "refactor(evidence): own typed runtime retrieval"
```
### Task 9: Migrate formulas into Curated Evidence
**Files:**
- Modify: `harness/tht/evidence/formula_store.py`
- Modify: `harness/tht/evidence/session.py`
- Modify: `harness/tht/cli/search_cmd.py`
- Create: `harness/.pi/skills/tht-sessione/modules/evidence/formula-proposals.md`
- Modify: `harness/.pi/skills/tht-sessione/projection.md.tmpl`
- Modify: `harness/tht/pi_skill_projection.py`
- Modify: `harness/tests/test_formula.py`
- Modify: `harness/tests/test_formula_wiring.py`
- Create: `harness/tests/test_evidence_formula_migration.py`
- Modify: `harness/tests/test_pi_skill_projection.py`
- Regenerate: `harness/.pi/skills/tht-sessione/SKILL.md`
**Step 1: Write RED migration tests**
Map an approved legacy `ConceptFormula` to a Curated Evidence formula while preserving:
- concept;
- SQL;
- columns;
- sources as provenance notes;
- stable deterministic ID;
- reviewed content wording.
Reject `auto` and unresolved `draft` formulas from direct publication; they become
Formula proposals.
**Step 2: Define the session proposal contract**
Add a schema-versioned formula-proposal projection in the session artifact. Keep
`concept_formula_approved` / `concept_formula_rejected` decisions unchanged because they
record a session-local choice, not repository publication.
**Step 3: Remove the separate runtime formula lookup**
Change `tht search find --kind formula` to call typed Evidence search with
`required_kinds=("formula",)`. Keep the legacy store readable only for the migration
command/window, with a deprecation warning in human output and no warning leakage into
pristine JSON.
**Step 4: Add and project formula instructions**
The Evidence fragment must say that a newly synthesized formula is a session proposal
and cannot be treated as Published Evidence.
**Step 5: Run tests**
```bash
cd harness
python -m tht.pi_skill_projection --write
.venv/bin/pytest tests/test_formula.py tests/test_formula_wiring.py \
tests/test_evidence_formula_migration.py tests/test_pi_skill_projection.py \
tests/test_decision_min_phase.py tests/test_workflow_observable_contract.py -q
```
Expected: PASS; decision phase ownership remains F4.
**Step 6: Commit**
```bash
git add harness/tht/evidence/formula_store.py harness/tht/evidence/session.py \
harness/tht/cli/search_cmd.py \
harness/.pi/skills/tht-sessione/modules/evidence/formula-proposals.md \
harness/.pi/skills/tht-sessione/projection.md.tmpl \
harness/.pi/skills/tht-sessione/SKILL.md harness/tht/pi_skill_projection.py \
harness/tests/test_formula.py harness/tests/test_formula_wiring.py \
harness/tests/test_evidence_formula_migration.py harness/tests/test_pi_skill_projection.py
git commit -m "refactor(evidence): unify formulas with typed evidence"
```
### Task 10: Add the small retrieval evaluation command
**Files:**
- Create: `harness/tht/evidence/evaluation.py`
- Modify: `harness/tht/cli/evidence_cmd.py`
- Create: `harness/tests/test_evidence_evaluation.py`
- Modify: `harness/tests/test_evidence_cli.py`
**Step 1: Write RED schema tests**
Use a deliberately small format:
```yaml
schema_version: 1
queries:
- id: pediatric-formula
query: Come distinguo i pazienti pediatrici?
purpose: sql_generation
expected:
- formula:fascia-pediatrica
```
Require unique query IDs, nonempty expected IDs and only public purpose values.
**Step 2: Write RED metric tests**
Compute `hit_at_5`, `hit_at_10`, missing expected IDs, empty-result queries and counts by
expected `kind`. Do not add nDCG, relevance grading or an evaluation database in v1.
**Step 3: Implement the evaluator and CLI**
Expose:
```text
tht evidence evaluate <workspace-root> -c <runtime-config> [--json]
```
The report includes workspace revision, active vector generation and the fixed default
RRF configuration. It is read-only.
**Step 4: Run tests**
```bash
cd harness
.venv/bin/pytest tests/test_evidence_evaluation.py tests/test_evidence_cli.py -q
.venv/bin/ruff check tht/evidence/evaluation.py tests/test_evidence_evaluation.py
```
Expected: PASS.
**Step 5: Commit**
```bash
git add harness/tht/evidence/evaluation.py harness/tht/cli/evidence_cmd.py \
harness/tests/test_evidence_evaluation.py harness/tests/test_evidence_cli.py
git commit -m "feat(evidence): evaluate retrieval with a small fixture"
```
### Task 11: Enforce curated-only runtime ingestion and update contracts
**Files:**
- Modify: `backend/src/workspaces/schema.ts`
- Modify: `backend/src/workspaces/runtime-renderer.ts`
- Modify: `backend/test/workspaces/schema.test.ts`
- Modify: `backend/test/workspaces/runtime-renderer.test.ts`
- Modify: `docs/contracts/workspace-evidence-v3.md`
- Modify: `docs/contracts/workspace-preprocessing-cli.md`
- Modify: `docs/architecture/overview.md`
- Modify: `PROJECT_STATE.md`
**Step 1: Write RED descriptor/rendering tests**
For filesystem Evidence, make `patterns: ["curated/**/*.md"]` the documented and
generated default for the new authoring layout. Continue accepting an explicitly
configured safe pattern for non-Git HTTP/S3 compatibility, but reject a filesystem
descriptor that includes both `source/**` and `curated/**` once it declares the new
layout version.
If a schema-version field is required to preserve compatibility, add it to the Evidence
subcontract, not to the whole workspace descriptor.
**Step 2: Implement the narrowest compatible descriptor change**
The materializer continues to copy the entire `evidence/` tree at the pinned commit.
Only the rendered runtime acquisition patterns restrict preprocessing to `curated/`.
Do not duplicate or move P6 materialization logic.
**Step 3: Update documentation contracts**
Document:
- source/curated layout;
- Git publication boundary;
- no runtime writes;
- validation before indexing;
- dense+BM25 collection contract and guarded rebuild;
- exact public operation names and JSON status additions, if any.
**Step 4: Run backend and harness contract gates**
```bash
cd backend
npx vitest run test/workspaces/schema.test.ts test/workspaces/runtime-renderer.test.ts \
test/workspaces/evidence/materialization.test.ts \
test/workspaces/evidence/preprocessing.test.ts
npx tsc --noEmit -p .
cd ../harness
.venv/bin/pytest tests/test_registry_evidence_config.py \
tests/test_filesystem_evidence_source.py tests/test_preprocess_cli.py -q
```
Expected: PASS; P6 materialization safety remains unchanged.
**Step 5: Commit**
```bash
git add backend/src/workspaces/schema.ts backend/src/workspaces/runtime-renderer.ts \
backend/test/workspaces/schema.test.ts \
backend/test/workspaces/runtime-renderer.test.ts \
docs/contracts/workspace-evidence-v3.md \
docs/contracts/workspace-preprocessing-cli.md docs/architecture/overview.md PROJECT_STATE.md
git commit -m "docs(evidence): publish curated-only workspace contract"
```
### Task 12: Migrate PSD and perform acceptance
**Files in ThothII:**
- Create: `docs/testing/evidence-restructuring-manual.md`
- Create: `scripts/evidence-restructuring-acceptance.sh`
- Create: `harness/tests/fixtures/evidence_authoring/poorly_structured.md`
- Modify: `PROJECT_STATE.md`
**Files in the external authoring repository:**
- Move: `/Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/<current-folders>`
to `/Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/source/`
- Create: `/Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/curated/<kind>/`
- Create: `/Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/manifest.yaml`
- Create: `/Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/evaluation.yaml`
- Modify: `/Users/mp/projects/tht-workspace-psd/psd-clinical/evidence/README.md`
- Modify: `/Users/mp/projects/tht-workspace-psd/psd-clinical/workspace.yaml`
Do not modify the external repository until the owner confirms the migration window and
the exact target branch. Treat that as the only manual authorization gate in this task.
**Step 1: Add a hermetic badly-structured fixture**
The fixture must contain prose, a rough list, an enum, an URL, an ambiguous statement
and a SQL formula candidate. The acceptance runner must prove:
- split into multiple typed units;
- ambiguity becomes `review_items`;
- no cross-source merge;
- unchanged rerun is a no-op;
- human edit is preserved;
- dirty-tree refusal;
- validation blocks unresolved review;
- validated corpus indexes and searches hybrid;
- Qdrant failure remains fail-closed.
Use a fake restructurer for hermetic CI. The real Pi call is a separate manual check.
**Step 2: Run the complete automated gates before external writes**
```bash
cd harness
.venv/bin/pytest -q
.venv/bin/ruff check .
cd ../backend
npx vitest run
npx tsc --noEmit -p .
cd ../tools/tht
go test ./...
go build ./cmd/tht
cd ../..
bash scripts/evidence-restructuring-acceptance.sh
```
Expected: all suites and acceptance checks PASS. If pre-existing unrelated failures
remain, record exact names and prove they reproduce at the baseline commit before
continuing.
**Step 3: Stop for the owner migration gate**
Provide:
- clean ThothII commit;
- test and acceptance summary;
- proposed PSD branch name;
- exact list of 36 source files to move;
- rollback command based on the pre-migration PSD commit;
- notice that Qdrant rebuild is destructive but scoped by exact collection-name guards.
Do not infer approval from prior design acceptance.
**Step 4: Migrate the PSD repository after approval**
Use `tht evidence prepare`, review the Git diff, resolve all review items manually, run
`tht evidence validate`, and create the approximately twenty evaluation queries. Do not
auto-merge or auto-push unless separately requested.
**Step 5: Activate and rebuild with the existing guarded operator path**
After the PSD merge/pull and activation, inspect first:
```text
tht --installation <absolute>/thothii-installation.yaml workspace vector inspect
--workspace psd-clinical --json
```
Then use the exact descriptor-owned name in the guarded rebuild command. Run
`workspace preprocess evidence`, `tht evidence evaluate`, and the manual F1/F3/F4
walkthrough.
**Step 6: Record acceptance and commit ThothII documentation**
`docs/testing/evidence-restructuring-manual.md` must record separate outcomes for:
- authoring and Git review;
- collection rebuild;
- preprocessing generation publication;
- hybrid retrieval evaluation;
- formula retrieval;
- graceful degradation;
- complete session behavior.
Update `PROJECT_STATE.md` only with observed results and immutable commit/run IDs.
```bash
git add docs/testing/evidence-restructuring-manual.md \
scripts/evidence-restructuring-acceptance.sh \
harness/tests/fixtures/evidence_authoring/poorly_structured.md PROJECT_STATE.md
git commit -m "test(evidence): record restructuring acceptance"
```
## Final verification checklist
Before claiming completion, verify:
- [ ] `CONTEXT.md` and both Evidence plan documents use the same terminology.
- [ ] All eight Evidence kinds have type-specific positive and negative tests.
- [ ] Pi is called once per changed source, with no tools and no saved session.
- [ ] `prepare` refuses dirty curated state and never deletes orphans.
- [ ] `validate` blocks unresolved review items.
- [ ] Runtime reads only `curated/**/*.md` from the pinned Git revision.
- [ ] The TypeScript and Python Qdrant compatibility checks agree.
- [ ] Qdrant collection uses named `dense` plus `bm25` with IDF.
- [ ] Italian BM25 options are identical during ingest and query.
- [ ] Hybrid search uses two prefetches and default RRF.
- [ ] Hard filters always include workspace, revision and active generation.
- [ ] Formula runtime lookup uses typed Evidence; session proposals remain non-published.
- [ ] Fragment hits are grouped into complete Evidence Units.
- [ ] Evaluation reports hit@5 and hit@10 against a versioned fixture.
- [ ] Existing corpus rollback, compensation, retention and resume tests still pass.
- [ ] Backend Vitest and TypeScript gates pass.
- [ ] Harness pytest and Ruff gates pass.
- [ ] Native `tht` Go tests/build pass.
- [ ] External PSD writes occurred only after explicit migration authorization.