3.9 KiB
Local pgvector Task 1 report
Status
Implemented the direct PgVectorStore behind the transport-neutral VectorStore port.
The adapter uses separate optional reader and writer database configurations, derives
capabilities from configured authority, validates strict positive search limits, filters kinds
in SQL before limiting, and merges multi-collection results by cosine similarity.
All collection identifiers are selected from the fixed schema_records, evidence, and
memory allowlist and composed with psycopg2.sql.Identifier. Values, vectors, kinds, hashes,
and limits remain bound parameters. Collection/kind mismatches fail with VectorStoreError.
Upserts preserve the canonical metadata shape, use record_key conflict semantics, update the
transport hash and embedding, and leave semantic metadata fields intact. Health probes reader
and writer independently and reports observed vector(N) dimensions against the configured
embedding dimension.
Configuration and factory
pgvector_direct now accepts explicit optional reader and writer DatabaseConfig entries.
The former connection entry remains supported as a deprecated read-only compatibility path.
build_vector_store(..., require_write=True) accepts writer-only direct configurations and
fails early when no explicit writer is present.
The transitional build_vector_loader bulk-sync path remains in place. It uses an explicit
direct writer when present, or the legacy connection; it deliberately does not treat a new
reader-only credential as writable. No production schema migration was added.
TDD and verification
- RED: the new tests initially failed at collection because
PgVectorStoredid not exist. - Docker L0 pgvector tests:
11 passed. - Direct + HTTP parity/factory/config focus:
51 passed. - Full harness:
461 passed, 5 deselected. - Changed-file Ruff lint: clean.
- Changed-file Ruff format check: clean.
git diff --check: clean.
The repository-wide ruff check . still reports 34 pre-existing test-file findings outside
Task 1; none are in changed files. The full pytest suite emits 17 existing legacy-config
deprecation warnings.
Scope and concerns
- Test fixtures create only the three existing vector tables needed to exercise the adapter; migration/versioning remains Task 2.
- The legacy single
connectionform stays read-only through the public port, matching its previous adapter behavior, while remaining available to the explicitly documented bulk-loader transition.
Review fix wave
The Task 1 review findings were addressed in a follow-up TDD cycle:
- Search now validates requested kinds against the global known-kind set, intersects valid kinds with each collection, and skips unrelated collections. A direct-versus-HTTP parity test covers the multi-collection case.
- Health requires all three allowlisted tables, an
embedding vector(N)column on every table, the expected dimension on every table, and the appropriate read or write table privileges for each configured side. Empty and partial schemas return deterministic, credential-free details; unexpected database failures expose only their exception class. - The Docker L0 fixture now provisions separate least-privilege reader and writer roles. Tests
prove the reader cannot insert, the writer cannot execute the cosine-search SELECT, and the
adapter still routes search to the reader and upsert/hash operations to the writer. Direct
upsert uses an atomic
INSERT ... ON CONFLICT DO NOTHINGfollowed byUPDATEfor an existing key, avoiding broad SELECT authority while retaining conflict-safe hash/upsert semantics.
Fresh verification after the fix wave:
- Docker L0 + HTTP port/search parity:
42 passed(earlier checkpoint); the final L0 file has16 passedincluding the stricter raw-role search denial. - Expanded focused adapter/config suite:
56 passed. - Full harness:
466 passed, 5 deselected. - Changed-file Ruff lint/format and
git diff --check: clean.