# Evidence Task 5C report ## Delivered - Added `vector.retain_published_generations` (default `3`, validation minimum `1`). - Retention runs only after publication. It keeps ACTIVE, the newest configured generations, and generations referenced by running or resumable failed job checkpoints. - Cleanup deletes the exact Evidence generation from the vector store before removing its immutable filesystem directory. Vector failures retain filesystem metadata for retry and produce credential-free partial reports. - Added idempotent `tht preprocess evidence gc [--dry-run] --json` reconciliation with pristine JSON output. - Materialized document reads now open generation/documents components with directory file descriptors and `O_NOFOLLOW`, require a regular file owned by the process with one link, and hash the bytes read from the same descriptor against the canonical manifest. - HTTP generation deletion is pinned to `delete_vector_generation` with exact table/kind/generation arguments. Legacy 404 responses fail closed with an actionable, sanitized migration message. ## Evidence - Focused retention, safe-read, CLI, and HTTP contract tests: `51 passed` (Docker-backed direct parametrizations excluded from that focused invocation). - Real Docker pgvector adapter suites: `33 passed`. - Full harness suite, including Docker-backed tests: `668 passed, 5 deselected`. - Changed-file Ruff: clean. - `git diff --check`: clean. The five deselected tests are the repository's opt-in `l2` tests requiring external services; they are not local pgvector tests. Test output retains pre-existing Pydantic serialization and legacy-config deprecation warnings. ## Review fix wave - Publication is now explicit and durable (`PUBLISHED` marker). Retention candidates require a valid generation manifest and publication marker (ACTIVE remains backward-compatible), so staged and malformed directories neither consume retention slots nor become deletion targets. - The policy retains ACTIVE plus exactly `N-1` newest rollback publications, ordered by durable publication time and generation id. Running and failed-resumable JobRunner checkpoints protect every referenced plan generation. - `VectorStore` now exposes exact Evidence generation inventory. Direct pgvector uses a constrained `SELECT DISTINCT` over `kind='evidence'` and `metadata.vector_generation`; HTTP uses the allowlisted `list_evidence_generations` RPC and fails closed on legacy 404. The writer RPC SQL, revokes, and grants are packaged in `create_vector_writer_rpc.sql`. - Explicit GC reconciles the union of published filesystem generations and vector-only orphans, preserving vector-before-filesystem deletion and retry semantics. - `run_as_job` holds the same corpus writer lock across checkpoint recovery, staging, publish, and retention. Explicit GC already uses this lock, serializing candidate snapshots with publishers. - Session artifact consumers no longer receive the corpus source path after validation. They get an owned, read-only copy atomically written from the bytes read and hash-validated on the same descriptor. Fresh verification after the fix wave: full harness `672 passed, 5 deselected`; Docker pgvector, HTTP parity, and migration suites `43 passed`; exact direct inventory/delete integration `1 passed`; changed-file Ruff and `git diff --check` clean. ## Final hardening verification - Canonical generation validation is exact (`^gen:[0-9a-f]{32}$`) before HTTP/direct deletion; malformed HTTP inventory rows fail closed rather than entering the GC candidate set. - Added explicit protection coverage for running and failed-resumable JobRunner checkpoints, plus a second-GC idempotence assertion for vector-only orphan reconciliation. - Added deterministic concurrent locking coverage: a job paused after discovery retains the corpus writer lock, explicit GC blocks, then completes after publication without deleting the active run. - Added a descriptor-race regression: replacing the corpus pathname immediately after `read(2)` leaves the atomically materialized session-owned copy byte-for-byte equal to the validated ACTIVE document and its manifest hash. Final fresh evidence: Docker pgvector/HTTP/migration suites `48 passed`; full harness `680 passed, 5 external L2 deselected`; changed-file Ruff and `git diff --check` clean. ## Integrated Task 5 dependency fixes - GC now distinguishes filesystem retention from vector dependencies. ACTIVE and the newest `N-1` published manifests keep their directories; every exact generation in their `document_generations` maps remains vector-protected even after its old publication directory is evicted. Job-protected manifests receive the same dependency treatment. - The real four-publication Docker lifecycle now includes an unchanged document whose vectors come from the first generation. With retention `N=2`, only the final two publication directories remain while the first generation's vectors remain searchable from ACTIVE and survive restart/explicit GC. - Evidence lookup is always wrapped by the ACTIVE-aware searcher. With no corpus/ACTIVE, Evidence returns no rows and search packs cannot expose legacy vectors; non-Evidence kinds are unchanged. - Session artifact resolution holds the corpus writer lock, snapshots the active manifest once, and materializes bytes using that exact `manifest_id`, preventing a concurrent publish/retain-1 GC from changing or deleting the selected source generation. Focused unit tests, the updated real Docker lifecycle, changed-file Ruff, and `git diff --check` pass. The final full harness invocation completed with exit code 0, including the concurrently added DWH JobRunner tests. ## Final ACTIVE search review fixes - `ActiveEvidenceSearcher` now treats default (`kinds=None`) and mixed-kind searches as explicit split queries: non-Evidence kinds are queried separately, while Evidence is queried only with ACTIVE manifest generation/document predicates applied server-side before every limit. - Results are merged deterministically by descending similarity then stable id and truncated once to the caller's global `top_n`. Pure non-Evidence searches retain their original delegate path. - The corpus writer lock now covers manifest snapshot construction and all corresponding vector queries, preventing retain-1 publication/GC from switching or deleting generations mid-search. - Removed the public post-LIMIT `active_evidence_hits` helper; no public Evidence path performs client filtering after limit. Focused default/mixed/no-ACTIVE/search-pack tests pass, the real Docker pgvector lifecycle passes, and the final full harness plus scoped Ruff/diff invocation completed with exit code 0. ## Workspace-scoped Evidence isolation - Evidence manifests, vector metadata, and record keys now carry the stable JobRunner workspace id derived from the configured workspace identity (config stem), never credentials or absolute paths. - Every ACTIVE server-side predicate includes `workspace_id`. Legacy unscoped rows therefore fail closed and cannot appear in Evidence results. - Vector generation inventory and deletion require the workspace namespace across the port, direct pgvector adapter, HTTP client/adapter, and allowlisted RPC SQL. Legacy unscoped RPC overloads are explicitly dropped during migration; destructive SQL matches collection, kind, generation, and workspace together. - GC recovers the persisted namespace from ACTIVE for explicit/restarted cleanup and can only list or delete that workspace's generations. Real shared-pgvector coverage proves deleting a generation for workspace A preserves the same generation in workspace B. - `PipelineResult.model_dump` now serializes fields explicitly instead of `dataclasses.asdict`, avoiding deepcopy of immutable `FrozenDict` metadata while preserving pristine JSON CLI output. Final focused verification: `89 passed` across corpus/CLI JSON, direct/HTTP parity, migrations, and real Docker pgvector lifecycle; scoped Ruff and `git diff --check` clean. A contemporaneous full-suite run reached unrelated Task 6 immutable-file tamper tests; those files were deliberately not changed.