4.3 KiB
Evidence Task 5 report
Outcome
Implemented an incremental Evidence corpus pipeline with immutable materialized generations,
generation-scoped vector records, and an fsynced atomic ACTIVE pointer. Runtime Evidence
artifact lookup reads the active canonical manifest and keeps a legacy source-tree fallback only
when no corpus has been published.
The CLI is available as tht preprocess evidence [--dry-run] [--resume RUN_ID] [--json].
JSON success and failure output is pristine and failure details are sanitized.
Safety and failure model
- A workspace writer lock serializes preprocess writers; readers never take the lock.
- Generation directories, manifests, materialized files, locks, and
ACTIVEreject symlink/path escape cases and use owner-only durable writes. - Vector records use generation-specific keys and metadata. The active manifest maps each active document to its valid vector generation, allowing unchanged documents to retain their vectors.
- Runtime retrieval admits only active document IDs and their manifest-selected generations. Removed documents and partial writes from failed generations are therefore unreachable.
- Embedding count and dimension checks occur before vector upsert; vector write count is checked
before staging/publish. Any failure leaves
ACTIVEunchanged. - Dry runs perform discovery/fingerprint planning only and never acquire, embed, write vectors, or publish. Fully unchanged runs return the active generation without creating a replacement.
- Resume can safely retry idempotent generation-scoped upserts and publish an already staged, compatibility-checked generation after a crash between staging and pointer replacement.
TDD evidence
Initial focused collection failed because tht.corpus.pipeline and tht.corpus.store did not
exist. The implemented suite covers incremental skips, removals, model/policy rebuilds, acquire and
partial-vector failures, dry-run isolation, dimension validation, atomic reader snapshots, pointer
validation, symlink defense, and pristine CLI JSON.
Fresh focused verification:
18 passed, 3 warnings in 0.39s
Command:
.venv/bin/pytest tests/test_corpus_pipeline.py tests/test_corpus_publish.py \
tests/test_preprocess_cli.py tests/test_search_pack.py tests/test_session_documents.py -q
Scoped Ruff: All checks passed!
Broader non-Docker/non-packaging run reached 560 passed, 5 deselected; ten pre-existing HTTP
adapter tests could not bind localhost under the sandbox. The complete suite reached 570 passed, 5 deselected, with the remaining failures/errors caused by denied Docker socket, localhost bind,
and offline wheel-build access. No task-focused test failed.
Remaining operational gate
Live pgvector integration needs Docker or an authorized local pgvector endpoint. The compensation
strategy is logical isolation rather than destructive cleanup because the shared VectorStore
port intentionally exposes no delete/transaction API; unreachable failed generations can be
garbage-collected by a future maintenance job.
Review integration wave
Added an enforceable metadata_filter vector-port contract and capability flags. Direct pgvector
places exact Evidence generation/document predicates in SQL before LIMIT; HTTP sends the same
filter to the RPC and deliberately does not use the legacy 404 fallback. The reader RPC script now
validates and applies that filter. Normal Evidence search and search-pack use an ACTIVE-aware
searcher that groups active documents by generation, executes complete server-filtered searches,
and merges the results.
Added exact-generation Evidence cleanup to direct and HTTP writers plus the allowlisted writer RPC. Pipeline failures compensate both staged filesystem state and vector writes; cleanup failures stay sanitized and ACTIVE filtering remains the exposure boundary. Corpus-present session artifact resolution now fails closed on corrupt/missing ACTIVE rather than falling through to source files.
Focused review-wave verification: 45 passed, scoped Ruff clean. A mocked REST regression proves the exact filter payload and fail-closed legacy 404 behavior.
Still outstanding from the expanded review request: Task-4 JobRunner stage-by-stage integration,
published-generation retention/garbage collection, same-fd dirfd materialized-file reads, and
live local pgvector integration could not be completed in this wave.