Implementazione del piano di remediation progressiva sui difetti emersi dall'analisi dell'harness. Tutto verificato: 214 test Python (incl. L0 su Postgres reale), 14 test JS del gate, ruff pulito. Blocco 1 (CRITICA, integrazione gate↔CLI): - phase advance: gate usa --auto + exit 6; reviewer_confirm kind:phase fa advance esplicito che applica i prerequisiti (prima non avanzava per le fasi a conferma umana). - cte plan riceve i --name dal gate (param names); set-question con id posizionale; skill `tht search find`; nuovo comando `tht memory save-one` con dedup hash client-side in save_one_memory. Blocco 2 (D15, stato post-rollback): - campo `phase` su DecisionRecord + effective_decisions phase-aware per i subject "a nome" (cte_approved ecc.); _compute_promotions e finalize sulla vista effective; finalize confronta col piano CTE effettivo, non glob; `decision add --retracts` + comando `decision retract`. Blocco 3 (D7 read-only + D6 manifest): - assert_read_only su tutti e quattro i codepath (direct + REST); - manifest author/summary/updated_at/updated_by/schema_version popolati + helper touch_manifest sulle mutazioni. Blocco 4-5 (D14a/D14b): - decision_min_phase data-driven via `emits:` in workflow.yaml; - formula evidence: status auto, search_formulas, gruppo CLI `tht formula`, `search find --kind formula`, load_evidence_dir salta i .sql.md. Blocco 6 (robustezza): - taskdoc slice promoted_tables + bound enforced; report escaping/bound + rsplit note; filtro kind reader REST/direct; conteggio upserted robusto; guard REST run_query non-list; LSH disallineato -> LshIndexError. Blocco 7 (pulizia): - dead code gate e KIND_TO_TABLE morto rimossi; doc Postgres-only (README + connection.py). Blocco 0 (parziale): test di compatibilità firma gate↔CLI (tests/integration). Rinviati: fake-Pi runtime completo, artifact-gate da disco (#23), parità eligibility REST/direct (#28), unificazione reserved-labels (#30), memory_rejected da deselezione (#33). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
74 lines
2.9 KiB
Python
74 lines
2.9 KiB
Python
"""Lettura del pgvector dietro un'unica interfaccia `.search(query_vec, top_n, kinds)`, così
|
|
`search.combined_search` resta agnostico al transport. Due implementazioni:
|
|
|
|
- `RestSearcher` → produzione: similarity search via REST (`search_similar`).
|
|
- `DirectSearcher` → dev/test: connessione diretta a Postgres/pgvector.
|
|
|
|
Entrambe mappano i `kind` sulle tabelle per-dominio dello schema `vectors`.
|
|
"""
|
|
|
|
from sqlalchemy import Engine
|
|
|
|
from tht.vectorstore.rest_client import VectorRestClient
|
|
from tht.vectorstore.store import VectorHit, VectorStore, hit_from_metadata
|
|
|
|
# kind Thoth → tabella dello schema `vectors`.
|
|
KIND_TO_TABLE = {
|
|
"schema_table": "schema_records",
|
|
"schema_column": "schema_records",
|
|
"evidence": "evidence",
|
|
"memory": "memory",
|
|
}
|
|
ALL_TABLES = ["schema_records", "evidence", "memory"]
|
|
|
|
|
|
def tables_for_kinds(kinds: list[str] | None) -> list[str]:
|
|
"""Tabelle da interrogare per i kind richiesti (tutte se kinds è vuoto/None)."""
|
|
if not kinds:
|
|
return list(ALL_TABLES)
|
|
return sorted({KIND_TO_TABLE[k] for k in kinds if k in KIND_TO_TABLE})
|
|
|
|
|
|
def _merge(hits: list[VectorHit], top_n: int) -> list[VectorHit]:
|
|
return sorted(hits, key=lambda h: h.similarity, reverse=True)[:top_n]
|
|
|
|
|
|
class RestSearcher:
|
|
"""Similarity search via REST: una chiamata `search_similar` per tabella, poi fusione."""
|
|
|
|
def __init__(self, client: VectorRestClient):
|
|
self.client = client
|
|
|
|
def search(
|
|
self, query_vec: list[float], top_n: int = 10, kinds: list[str] | None = None
|
|
) -> list[VectorHit]:
|
|
hits: list[VectorHit] = []
|
|
for table in tables_for_kinds(kinds):
|
|
for row in self.client.search_similar(table, query_vec, top_n):
|
|
hits.append(hit_from_metadata(row.get("similarity", 0.0), row.get("metadata")))
|
|
# schema_records contiene sia schema_table sia schema_column: la RPC non filtra
|
|
# per kind, quindi lo facciamo lato client per parita' col path diretto (#25).
|
|
if kinds:
|
|
allowed = set(kinds)
|
|
hits = [h for h in hits if h.kind in allowed]
|
|
return _merge(hits, top_n)
|
|
|
|
|
|
class DirectSearcher:
|
|
"""Similarity search diretta su Postgres/pgvector, interrogando le tabelle per-dominio."""
|
|
|
|
def __init__(self, engine: Engine, schema: str = "vectors", dim: int = 768):
|
|
self.engine = engine
|
|
self.schema = schema
|
|
self.dim = dim
|
|
|
|
def search(
|
|
self, query_vec: list[float], top_n: int = 10, kinds: list[str] | None = None
|
|
) -> list[VectorHit]:
|
|
hits: list[VectorHit] = []
|
|
for table in tables_for_kinds(kinds):
|
|
store = VectorStore(self.engine, schema=self.schema, table=table, dim=self.dim)
|
|
# passa kinds: dentro schema_records filtra schema_table vs schema_column (#25).
|
|
hits.extend(store.search(query_vec, top_n=top_n, kinds=kinds))
|
|
return _merge(hits, top_n)
|