feat(preprocess): publish incremental Evidence corpus

This commit is contained in:
2026-07-12 04:26:24 +02:00
parent 16a8bd9df6
commit b96d4f13b9
10 changed files with 708 additions and 0 deletions
@@ -0,0 +1,61 @@
# Evidence Task 5 report
## Outcome
Implemented an incremental Evidence corpus pipeline with immutable materialized generations,
generation-scoped vector records, and an fsynced atomic `ACTIVE` pointer. Runtime Evidence
artifact lookup reads the active canonical manifest and keeps a legacy source-tree fallback only
when no corpus has been published.
The CLI is available as `tht preprocess evidence [--dry-run] [--resume RUN_ID] [--json]`.
JSON success and failure output is pristine and failure details are sanitized.
## Safety and failure model
- A workspace writer lock serializes preprocess writers; readers never take the lock.
- Generation directories, manifests, materialized files, locks, and `ACTIVE` reject symlink/path
escape cases and use owner-only durable writes.
- Vector records use generation-specific keys and metadata. The active manifest maps each active
document to its valid vector generation, allowing unchanged documents to retain their vectors.
- Runtime retrieval admits only active document IDs and their manifest-selected generations.
Removed documents and partial writes from failed generations are therefore unreachable.
- Embedding count and dimension checks occur before vector upsert; vector write count is checked
before staging/publish. Any failure leaves `ACTIVE` unchanged.
- Dry runs perform discovery/fingerprint planning only and never acquire, embed, write vectors, or
publish. Fully unchanged runs return the active generation without creating a replacement.
- Resume can safely retry idempotent generation-scoped upserts and publish an already staged,
compatibility-checked generation after a crash between staging and pointer replacement.
## TDD evidence
Initial focused collection failed because `tht.corpus.pipeline` and `tht.corpus.store` did not
exist. The implemented suite covers incremental skips, removals, model/policy rebuilds, acquire and
partial-vector failures, dry-run isolation, dimension validation, atomic reader snapshots, pointer
validation, symlink defense, and pristine CLI JSON.
Fresh focused verification:
```text
18 passed, 3 warnings in 0.39s
```
Command:
```text
.venv/bin/pytest tests/test_corpus_pipeline.py tests/test_corpus_publish.py \
tests/test_preprocess_cli.py tests/test_search_pack.py tests/test_session_documents.py -q
```
Scoped Ruff: `All checks passed!`
Broader non-Docker/non-packaging run reached `560 passed, 5 deselected`; ten pre-existing HTTP
adapter tests could not bind localhost under the sandbox. The complete suite reached `570 passed,
5 deselected`, with the remaining failures/errors caused by denied Docker socket, localhost bind,
and offline wheel-build access. No task-focused test failed.
## Remaining operational gate
Live pgvector integration needs Docker or an authorized local pgvector endpoint. The compensation
strategy is logical isolation rather than destructive cleanup because the shared `VectorStore`
port intentionally exposes no delete/transaction API; unreachable failed generations can be
garbage-collected by a future maintenance job.