189 lines
12 KiB
Markdown
189 lines
12 KiB
Markdown
# Evidence Task 5C report
|
|
|
|
## Delivered
|
|
|
|
- Added `vector.retain_published_generations` (default `3`, validation minimum `1`).
|
|
- Retention runs only after publication. It keeps ACTIVE, the newest configured generations,
|
|
and generations referenced by running or resumable failed job checkpoints.
|
|
- Cleanup deletes the exact Evidence generation from the vector store before removing its
|
|
immutable filesystem directory. Vector failures retain filesystem metadata for retry and
|
|
produce credential-free partial reports.
|
|
- Added idempotent `tht preprocess evidence gc [--dry-run] --json` reconciliation with pristine
|
|
JSON output.
|
|
- Materialized document reads now open generation/documents components with directory file
|
|
descriptors and `O_NOFOLLOW`, require a regular file owned by the process with one link, and
|
|
hash the bytes read from the same descriptor against the canonical manifest.
|
|
- HTTP generation deletion is pinned to `delete_vector_generation` with exact
|
|
table/kind/generation arguments. Legacy 404 responses fail closed with an actionable,
|
|
sanitized migration message.
|
|
|
|
## Evidence
|
|
|
|
- Focused retention, safe-read, CLI, and HTTP contract tests: `51 passed` (Docker-backed direct
|
|
parametrizations excluded from that focused invocation).
|
|
- Real Docker pgvector adapter suites: `33 passed`.
|
|
- Full harness suite, including Docker-backed tests: `668 passed, 5 deselected`.
|
|
- Changed-file Ruff: clean.
|
|
- `git diff --check`: clean.
|
|
|
|
The five deselected tests are the repository's opt-in `l2` tests requiring external services;
|
|
they are not local pgvector tests. Test output retains pre-existing Pydantic serialization and
|
|
legacy-config deprecation warnings.
|
|
|
|
## Review fix wave
|
|
|
|
- Publication is now explicit and durable (`PUBLISHED` marker). Retention candidates require a
|
|
valid generation manifest and publication marker (ACTIVE remains backward-compatible), so
|
|
staged and malformed directories neither consume retention slots nor become deletion targets.
|
|
- The policy retains ACTIVE plus exactly `N-1` newest rollback publications, ordered by durable
|
|
publication time and generation id. Running and failed-resumable JobRunner checkpoints protect
|
|
every referenced plan generation.
|
|
- `VectorStore` now exposes exact Evidence generation inventory. Direct pgvector uses a constrained
|
|
`SELECT DISTINCT` over `kind='evidence'` and `metadata.vector_generation`; HTTP uses the
|
|
allowlisted `list_evidence_generations` RPC and fails closed on legacy 404. The writer RPC SQL,
|
|
revokes, and grants are packaged in `create_vector_writer_rpc.sql`.
|
|
- Explicit GC reconciles the union of published filesystem generations and vector-only orphans,
|
|
preserving vector-before-filesystem deletion and retry semantics.
|
|
- `run_as_job` holds the same corpus writer lock across checkpoint recovery, staging, publish, and
|
|
retention. Explicit GC already uses this lock, serializing candidate snapshots with publishers.
|
|
- Session artifact consumers no longer receive the corpus source path after validation. They get
|
|
an owned, read-only copy atomically written from the bytes read and hash-validated on the same
|
|
descriptor.
|
|
|
|
Fresh verification after the fix wave: full harness `672 passed, 5 deselected`; Docker pgvector,
|
|
HTTP parity, and migration suites `43 passed`; exact direct inventory/delete integration `1 passed`;
|
|
changed-file Ruff and `git diff --check` clean.
|
|
|
|
## Final hardening verification
|
|
|
|
- Canonical generation validation is exact (`^gen:[0-9a-f]{32}$`) before HTTP/direct deletion;
|
|
malformed HTTP inventory rows fail closed rather than entering the GC candidate set.
|
|
- Added explicit protection coverage for running and failed-resumable JobRunner checkpoints, plus
|
|
a second-GC idempotence assertion for vector-only orphan reconciliation.
|
|
- Added deterministic concurrent locking coverage: a job paused after discovery retains the corpus
|
|
writer lock, explicit GC blocks, then completes after publication without deleting the active run.
|
|
- Added a descriptor-race regression: replacing the corpus pathname immediately after `read(2)`
|
|
leaves the atomically materialized session-owned copy byte-for-byte equal to the validated ACTIVE
|
|
document and its manifest hash.
|
|
|
|
Final fresh evidence: Docker pgvector/HTTP/migration suites `48 passed`; full harness `680 passed,
|
|
5 external L2 deselected`; changed-file Ruff and `git diff --check` clean.
|
|
|
|
## Integrated Task 5 dependency fixes
|
|
|
|
- GC now distinguishes filesystem retention from vector dependencies. ACTIVE and the newest
|
|
`N-1` published manifests keep their directories; every exact generation in their
|
|
`document_generations` maps remains vector-protected even after its old publication directory is
|
|
evicted. Job-protected manifests receive the same dependency treatment.
|
|
- The real four-publication Docker lifecycle now includes an unchanged document whose vectors come
|
|
from the first generation. With retention `N=2`, only the final two publication directories remain
|
|
while the first generation's vectors remain searchable from ACTIVE and survive restart/explicit GC.
|
|
- Evidence lookup is always wrapped by the ACTIVE-aware searcher. With no corpus/ACTIVE, Evidence
|
|
returns no rows and search packs cannot expose legacy vectors; non-Evidence kinds are unchanged.
|
|
- Session artifact resolution holds the corpus writer lock, snapshots the active manifest once, and
|
|
materializes bytes using that exact `manifest_id`, preventing a concurrent publish/retain-1 GC from
|
|
changing or deleting the selected source generation.
|
|
|
|
Focused unit tests, the updated real Docker lifecycle, changed-file Ruff, and `git diff --check` pass.
|
|
The final full harness invocation completed with exit code 0, including the concurrently added DWH
|
|
JobRunner tests.
|
|
|
|
## Final ACTIVE search review fixes
|
|
|
|
- `ActiveEvidenceSearcher` now treats default (`kinds=None`) and mixed-kind searches as explicit
|
|
split queries: non-Evidence kinds are queried separately, while Evidence is queried only with
|
|
ACTIVE manifest generation/document predicates applied server-side before every limit.
|
|
- Results are merged deterministically by descending similarity then stable id and truncated once
|
|
to the caller's global `top_n`. Pure non-Evidence searches retain their original delegate path.
|
|
- The corpus writer lock now covers manifest snapshot construction and all corresponding vector
|
|
queries, preventing retain-1 publication/GC from switching or deleting generations mid-search.
|
|
- Removed the public post-LIMIT `active_evidence_hits` helper; no public Evidence path performs
|
|
client filtering after limit.
|
|
|
|
Focused default/mixed/no-ACTIVE/search-pack tests pass, the real Docker pgvector lifecycle passes,
|
|
and the final full harness plus scoped Ruff/diff invocation completed with exit code 0.
|
|
|
|
## Workspace-scoped Evidence isolation
|
|
|
|
- Evidence manifests, vector metadata, and record keys now carry the stable JobRunner workspace id
|
|
derived from the configured workspace identity (config stem), never credentials or absolute paths.
|
|
- Every ACTIVE server-side predicate includes `workspace_id`. Legacy unscoped rows therefore fail
|
|
closed and cannot appear in Evidence results.
|
|
- Vector generation inventory and deletion require the workspace namespace across the port, direct
|
|
pgvector adapter, HTTP client/adapter, and allowlisted RPC SQL. Legacy unscoped RPC overloads are
|
|
explicitly dropped during migration; destructive SQL matches collection, kind, generation, and
|
|
workspace together.
|
|
- GC recovers the persisted namespace from ACTIVE for explicit/restarted cleanup and can only list
|
|
or delete that workspace's generations. Real shared-pgvector coverage proves deleting a generation
|
|
for workspace A preserves the same generation in workspace B.
|
|
- `PipelineResult.model_dump` now serializes fields explicitly instead of `dataclasses.asdict`,
|
|
avoiding deepcopy of immutable `FrozenDict` metadata while preserving pristine JSON CLI output.
|
|
|
|
Final focused verification: `89 passed` across corpus/CLI JSON, direct/HTTP parity, migrations, and
|
|
real Docker pgvector lifecycle; scoped Ruff and `git diff --check` clean. A contemporaneous full-suite
|
|
run reached unrelated Task 6 immutable-file tamper tests; those files were deliberately not changed.
|
|
|
|
## Immutable corpus/workspace binding
|
|
|
|
- A corpus root becomes bound to the workspace id persisted in its ACTIVE manifest. Job, non-job,
|
|
explicit GC, and ACTIVE search entry points compare the configured namespace before discovery,
|
|
vector access, staging, deletion, or ACTIVE mutation.
|
|
- Reusing the same paths after renaming a workspace now fails closed with a typed/sanitized message:
|
|
use a new corpus root or perform an intentional explicit rebuild. Unscoped legacy manifests also
|
|
fail this ownership check.
|
|
- Tests prove unchanged-document reuse cannot silently mix workspace A vectors into a workspace B
|
|
manifest, and that mismatched job, GC, and search paths perform no vector/filesystem mutations.
|
|
|
|
Focused workspace-binding, search-pack, preprocess JSON, and scoped Ruff/diff tests pass.
|
|
|
|
Compatibility follow-up: direct/internal `CorpusPipeline` instances now distinguish an omitted
|
|
workspace identity from an explicit config/job identity. An unbound instance adopts the persisted
|
|
ACTIVE owner (or `default` only for a brand-new direct corpus), preserving safe resume/GC tests and
|
|
the real pgvector lifecycle. Explicit config/job identities still fail closed on any mismatch. The
|
|
two reported regressions, workspace mismatch guards, real Docker lifecycle, scoped Ruff/diff, and
|
|
the full harness suite all pass.
|
|
|
|
Final fail-closed follow-up: persisted ACTIVE ownership is now validated under the corpus lock before
|
|
every configured search delegate, including default, mixed, pack, and non-Evidence-only operations.
|
|
Malformed or missing `metadata.workspace_id` is intrinsically rejected even for unbound direct
|
|
callers; source discovery, vector operations, GC, files, and ACTIVE remain untouched. Focused tests,
|
|
real Docker lifecycle, scoped Ruff/diff, and the full harness regression run pass.
|
|
|
|
Final lock/preflight follow-up: `CorpusPipeline.gc()` now acquires the corpus writer lock itself for
|
|
ownership validation through vector/filesystem cleanup. The store lock is thread-reentrant so nested
|
|
job retention is safe without weakening cross-thread/process exclusion; the CLI wrapper no longer
|
|
double-locks. Search find/pack performs locked corpus ownership preflight immediately after config
|
|
load, before DWH leasing, vector/searcher factories, embeddings, or schema work. Focused concurrency
|
|
and fail-closed tests, real Docker lifecycle, scoped Ruff/diff, and the full harness pass.
|
|
|
|
## Compact public Evidence reports
|
|
|
|
- Public `PipelineResult.model_dump()` is now a bounded operational envelope: terminal status,
|
|
run/resume/publication/generation/manifest identifiers, capped changed/unchanged/removed source
|
|
identifiers, and aggregate document/chunk counts. Full manifests, bodies, and metadata remain
|
|
internal/on disk and are never serialized to CLI stdout.
|
|
- `tht preprocess evidence` exits `1` for any durable terminal status other than `succeeded` in
|
|
both JSON and text modes. JSON stdout remains one pristine sanitized object; text mode emits one
|
|
compact stderr error without traceback, exception identity, evidence content, or credentials.
|
|
- Tests cover a real failed acquisition job, sensitive evidence content, capped thousand-item
|
|
summaries, bounded report size, and smoke-compatible changed/unchanged fields.
|
|
|
|
Focused tests and scoped Ruff/diff pass. The contemporaneous full suite reaches an unrelated Task 6
|
|
DWH snapshot fixture missing its newly required workspace identity.
|
|
|
|
### Safe result representation and exact text totals
|
|
|
|
- `PipelineResult.manifest` is explicitly excluded from dataclass representation and the custom
|
|
representation is fixed-size operational data only. It omits manifest ids, documents, chunks,
|
|
content, metadata, and errors; `str(result)` inherits the same safe representation.
|
|
- Text-mode Evidence success output reads the uncapped aggregate totals from `payload["counts"]`
|
|
rather than the intentionally capped identifier arrays.
|
|
- Regression coverage builds a thousand-document/chunk manifest containing content and
|
|
credential-like metadata secrets, checks bounded `repr`/`str`, and verifies exact totals above
|
|
the 100-item public-array cap.
|
|
|
|
Focused Evidence verification passes (`67 passed`), and scoped Ruff is clean. The full harness run
|
|
is not green in this sandbox: Docker-backed tests cannot access the daemon, wheel packaging cannot
|
|
use the restricted build environment, and concurrent Task 6 DWH binding changes currently fail two
|
|
DWH tests. None of those failures touch the Evidence files in this follow-up.
|