fix: scope workspace registry smoke cleanup

This commit is contained in:
2026-08-08 22:32:57 +02:00
parent a3e348cf22
commit 43d8063922
14 changed files with 187 additions and 52 deletions
+8 -5
View File
@@ -22,16 +22,19 @@ The `tht` command is now on PATH. Node ≥ 20 is needed for the gate JS tests
```bash
cp .env.example .env
# fill in: THT_PROFILE, THT_DB_*, THT_DWH_API_KEY, THT_VEC_API_KEY,
# THT_VEC_WRITE_API_KEY, THT_SSL_CA, THT_OLLAMA_URL, ...
# fill in: THT_PROFILE, THT_DB_*, THT_DWH_API_KEY, THT_SSL_CA, ...
```
Keys are never logged; URLs are fine. Rotate any key that appeared in chat.
### `workspaces/<name>.yaml`
A workspace wires the relational DWH + the pgvector (dual-key) + embeddings + evidence.
See `workspaces/tht.example.yaml`. `${THT_*}}` tokens expand from `.env`.
A schema-v3 workspace wires the external relational DWH to one internal Qdrant collection
and the internal Ollama embedding model. Evidence paths remain workspace-local, while Qdrant
stores the derived semantic projection for schema, Evidence, Memory, and solved-question
records. The legacy files under `workspaces/` are retained as migration fixtures; new
operator-facing descriptors live in the workspace Git registry. `${THT_*}` tokens expand
from `.env`.
> **DB support (MVP):** the `direct` transport supports **PostgreSQL only** (psycopg2
> driver, `pg_*` catalog introspection, postgres-dialect sqlcheck/EXPLAIN). The central
@@ -130,7 +133,7 @@ tht/ Python package (CLI + workflow + phase + decisions + db/res
.pi/ Pi project (settings, prompts, themes, extensions/tht-gate.js + gate/)
workflow.yaml single source of workflow truth (F2)
workspaces/ workspace YAML definitions (D3)
scripts/ reader/writer RPC SQL for pgvector (D11)
scripts/ retained legacy SQL fixtures and workspace utilities
tests/ L0 (testcontainers), L1 (logic + builders), L2 (real model + DB)
docs/ testing guide + workflow editing
```
+8 -6
View File
@@ -59,20 +59,22 @@ the anti-bypass hooks. The glue depends on the Pi runtime (`pi.registerTool`,
## L2 — real LLM + real remote DB, manual / pre-release (NOT automated)
**What:** end-to-end sessions with GLM 5.2 + the real Chirone DWH + pgvector, reached
via REST over VPN. Plus the gate-glue validation (the part L1 cannot reach).
**What:** end-to-end sessions with GLM 5.2 + the real Chirone DWH, plus the internal
Qdrant/Ollama semantic services started by the ThothII stack. Plus the gate-glue
validation (the part L1 cannot reach).
**Dependencies (all required, skip cleanly if missing):**
- LLM: Pi configured locally with GLM 5.2.
- DB: the remote Supabase endpoints (DWH read-only + pgvector reader/writer), via VPN.
- `harness/.env` populated with the API keys + CA path.
- DB: the remote DWH endpoint, via VPN when required.
- ThothII stack: internal Qdrant and Ollama services reachable from `core`.
- `harness/.env` populated with the required DWH/model API keys + CA path.
**Coverage (honest):** validates the assumption L1 cannot — that GLM 5.2 produces tool
calls the gate accepts, that the skill's prompts lead to the expected interaction shape,
that the gate glue handles real tool-call sequences (incl. Altro/Rifiuta/rollback),
that value grounding and formula approval surface correctly on the real schema, that
`memory save-one` upserts to the real pgvector. **Closes the skill→LLM→gate loop AND
exercises the gate glue.**
`memory save-one` upserts to the configured semantic store. **Closes the
skill→LLM→gate loop AND exercises the gate glue.**
**Honest limitation:** L2 is non-deterministic (the model may behave differently across
runs) and slow/costly. It is a **pre-release safety net, not a regression gate**.
@@ -62,8 +62,8 @@ def test_memory_command_writes_through_factory_vector_store(monkeypatch):
captured = []
original_upsert = store.upsert
store.upsert = lambda table, rows: captured.extend(rows) or original_upsert(table, rows)
# Server deployments write directly to pgvector and intentionally do not
# configure the workstation-only REST writer key.
# Legacy server deployments wrote directly through the factory and intentionally
# did not configure the workstation-only REST writer key.
cfg = SimpleNamespace(profile="server", embeddings=object(), vector_write_rest=None)
manifest = SimpleNamespace(id="s1")
snapshot = SimpleNamespace(manifest=manifest, decisions=[], artifacts={})
+2 -2
View File
@@ -1,8 +1,8 @@
"""L1: tht memory save-one -- targeted upsert via the writer key (spec D11).
The D11 deviation: instead of a full vectorstore resync (tht memory index / sync),
a remote workstation with a writer key can push a SINGLE promoted decision to
pgvector as a one-row upsert. This test pins the pure core of that behavior:
the workflow can push a SINGLE promoted decision to the configured semantic store as
a one-row upsert. This test pins the pure core of that behavior:
- exactly one VectorRecord is built for the chosen decision_seq
- the writer.upsert_records is called once with a single row
- writer.sync is NEVER called (that is the full-resync path)
+2 -2
View File
@@ -1,7 +1,7 @@
"""L1: `tht memory solved-search` — degrado gentile e mapping dei risultati.
SKILL.md prescrive solved-search in F4/F6/F7 di OGNI sessione: a vectordb
irraggiungibile (VPN giu', Ollama spento) il comando non deve morire con un
SKILL.md prescrive solved-search in F4/F6/F7 di OGNI sessione: se lo store
semantico è irraggiungibile (Qdrant/Ollama non disponibili) il comando non deve morire con un
traceback grezzo ma degradare a un avviso di una riga su stderr, con stdout
puro (`[]` in modalita' --json) ed exit 0, cosi' il modello prosegue senza
exemplar. Il finalize-hook gestisce gia' lo stesso scenario in modo analogo.
+3 -3
View File
@@ -299,10 +299,10 @@ def memory_vector_record_for_decision(
def save_one_memory(
records: list[MemoryRecord], decision_seq: int, *, store, embedder
) -> int:
"""Targeted one-row upsert of a promoted decision to pgvector via the writer key
"""Targeted one-row upsert of a promoted decision to the configured semantic store
(spec D11). This is NOT a full vectorstore resync: it embeds and pushes a single
record, so a workstation with a writer key can publish one memory without
rebuilding the index. Returns the upsert count (0 if no record matched or the
record, so a memory can be published without rebuilding the index.
Returns the upsert count (0 if no record matched or the
record is already up to date).
Hash dedup client-side (spec §5.4): the SHA-256 of the content is compared with
+3 -3
View File
@@ -59,8 +59,8 @@ def aggregate_lsh_multi(hits: list[dict]) -> dict[str, list[dict]]:
grouped: dict[str, list[dict]] = {}
for (table, _), row in best.items():
grouped.setdefault(table, []).append(row)
for table in grouped:
grouped[table].sort(key=lambda r: r["score"], reverse=True)
for rows in grouped.values():
rows.sort(key=lambda r: r["score"], reverse=True)
return grouped
@@ -99,7 +99,7 @@ def combined_search(
kinds: list[str] | None,
query_vec: list[float] | None = None,
) -> list[SearchResult]:
"""Fonde LSH (valori di campo) e pgvector con Reciprocal Rank Fusion.
"""Fonde LSH (valori di campo) e ricerca semantica con Reciprocal Rank Fusion.
`query_vec` permette di riusare un embedding gia' calcolato della stessa
keyword (es. `tht search pack`, che fa piu' ricerche sulla stessa domanda)."""
+7 -5
View File
@@ -52,11 +52,13 @@ def hit_from_metadata(similarity: float, metadata: dict | None) -> VectorHit:
class VectorStore:
"""Tabella pgvector table-scoped: scrittura diretta (loading) su una tabella dello schema
`vectors`. La lettura via REST avviene su `search_similar`; questo store serve al loading e
alla lettura diretta (dev/test). Il contratto della tabella remota richiede `id` (BIGSERIAL),
`embedding vector(N)` e `metadata jsonb`; le colonne extra (`record_key`, `kind`,
`content_hash`) servono solo al loader e non sono esposte dalla REST."""
"""Legacy table-scoped vector store retained for compatibility fixtures.
New operational semantic storage is handled by the Qdrant adapter. This class preserves
the older SQL-table contract used by historical tests and migration checks: `id`
(BIGSERIAL), `embedding vector(N)`, and `metadata jsonb`; the extra columns
(`record_key`, `kind`, `content_hash`) serve only the loader and are not exposed by REST.
"""
def __init__(
self, engine: Engine, schema: str = "vectors", table: str = "records", dim: int = 768
+3 -2
View File
@@ -4,8 +4,9 @@ Reads workspaces/<name>.yaml, expands ${VAR} from env, validates via the Config
(ported from the reference implementation). Future migration to a DB store would replace only this module.
La struttura YAML rispecchia esattamente tht/config.py:
database + rest (DWH), vector_rest + vector_write_rest (pgvector, doppia key top-level),
vector_db (loading diretto, server-only), embeddings, evidence, execution.
database/rest o resources.dwh per il DWH, resources.vector/resources.embeddings
per Qdrant/Ollama interni, più evidence ed execution. I vecchi campi vector_db e
vector_rest restano solo per leggere fixture legacy durante la migrazione.
"""
from __future__ import annotations