44 test L1 sui 3 moduli backend con logica non banale (opzione 2 della user review):
- sqlcheck.validate_sql (16 test): parse/single-statement, read-only enforcement
(INSERT/UPDATE/DELETE/CREATE/DROP/ALTER/TRUNCATE/GRANT rifiutati, WITH/UNION ok),
forbidden functions (dblink default blacklist, custom set, allowed not flagged),
object-existence (tabella inesistente, CTE non flaggata, perimetro promoted warning,
colonna inesistente con alias). Documenta una limitazione reale: le funzioni
aggregate specializzate (count/sum/coalesce) NON sono catturate dal name-matcher
perche' sqlglot modella .name come argomento, non come nome funzione.
- ctetest (14 test): has_trailing_select (semantica controintuitiva: True = violazione),
last_cte_name, build_test_sql, ledger I/O (load/append roundtrip, JSON-array e
JSONL tolleranti, corrupt-ledger raise).
- execute._inject_limit (6 test): LIMIT iniettato quando assente (limit+1 per
troncamento), rispettato quando presente, non iniettato su non-query, UNION/WITH ok.
Suite: 153 passed (109 + 44). Bonus: __psd_probe__ -> __tht_probe__ (riferimento
cliente neutralizzato in ctetest).
- Task 3.2: nuovo Step 1b — drop funzioni registry da memory_cmd + tht/memory.py
(load/save/update/delete/promote/reusable_promotions); promotion (F5) riscritta
come upsert batch al vectordb; memory list/delete su vectordb (no registry).
- Skill F2/F5: drop riferimenti a 'memory promote + memory index' con registry.
- Onda 0b riscritta: repo workspace per-cliente separato (no copia in harness/),
LSH scarica-tutti-i-valori-distinti (no 'campiona'), indice nel repo workspace cliente.
Tre correzioni da user review:
5. Memory SOLO pgvector, niente registry. Il registry.jsonl di ChironeWp3 è vestigiale:
una volta che il metadata del vectordb ha subject/detail/rationale (decisione 6), il
registry non serve. Promotion (F5) = upsert diretto a vectordb. tht/memory.py non
porta le 6 funzioni registry. 'Cancellare' = metadata.status='superseded' (audit
trail; il writer e' upsert-only). Multi-workstation OK per costruzione.
7. Evidence: repo workspace separato per-cliente, NON dentro ThothII. ThothII e'
generico; un repo tht-workspace-<cliente>/ contiene evidence/ + workspace YAML +
indici LSH. Deploy = checkout ThothII + checkout workspace-cliente. Niente copia
in harness/.
8. LSH: scarica TUTTI i valori distinti (non 'campiona'), costruisce MinHash+LSH
dentro harness come preprocessing. Indice per-cliente (nel repo workspace cliente).
Sezione 3 punto 2 (F2) riscritta: save-one diretto, drop reference a promote/index
con registry. Arricchimento metadata ora 'obbligatorio' (non opzionale): senza
registry, il vectordb e' l'unica fonte. Residui nsp nei path spec corretti a tht.
Piano da spec 2026-06-27-cli-port-completo-skill-riscritta-design.md.
Struttura in onde: -1 (renaming tht isolato), 0 (backend), 0b (evidence+LSH setup),
1 (radici CLI + phase drift), 2 (vector), 3 (foglia + metadata memory), 4 (SQL/CTE),
S (skill riscritta), L2 (sessione manuale).
TDD per logica nuova (schema_linking_phase, require_phase_or_exit, arricchimento
metadata memory); port+smoke+commit per i port verbatim (logica gia' validata in
ChironeWp3). pytest verde (109 passed) a ogni task come gate di regressione.
Self-review: copertura spec completa (11 decisioni), nessun placeholder, type
consistency verificata. Gap residui onesti: circularita' session_cmd<->sql_cmd
(risolto con import lazy), RPC server mancanti (fuori piano codice), L2 manuale.
Renaming richiesto in user review: Thoth (tht) e' il prodotto, PSD e' il cliente.
Nessun riferimento al contesto clinico nel codice.
Decisioni 8-10:
8. Rinomine: nsp->tht (comando+package+46 import), nsp-sessione->tht-sessione,
nsp-gate.js->tht-gate.js, chirone.*->tht.{example,test}.yaml (generici; il deploy
cliente crea il suo psd.yaml non-committato), THOTH_*->THT_* env.
9. Neutralizzazione riferimenti chirone/psd/policlinico/sandonato nei commenti/
docstring (resi generici o rimossi). Il contesto cliente vive SOLO nei file di
config reali (.env gitignored, workspace cliente non-committato).
10. Onda -1 isolata PRIMA del porting CLI: pytest resta 109 passed (rename verificato
da solo), poi il porting avviene col nome nuovo (niente doppio lavoro).
Ordine esecuzione aggiornato a 7 step (Onda -1 prima di tutto). Self-review:
corretti i residui incoerenti di nsp/chirone nello spec (righe che usavano ancora
i nomi vecchi dove dovevano essere tht). Residui rimasti sono legittimi (descrivono
il renaming o il path sorgente one-shot della copia evidence).
Tre decisioni su dipendenze implicite (domanda user review):
4. Indipendenza da ChironeWp3 (proprieta' architetturale): quando ThothII e' pronto,
il server non deve avere ChironeWp3 — solo Supabase (RPC SECURITY DEFINER nel DB,
indipendenti dal codice app) + cartella evidence (dentro ThothII). Verificato:
codice ThothII non ha riferimenti ChironeWp3/psdwp3, .env non punta a path chirone.
5. Registro memory: locale per-workstation (registry.jsonl in harness/artifacts/memory/
su ciascuna). La F2 legge dal vectordb condiviso e, grazie all'arricchimento metadata
(decisione 6), ricostruisce la decisione senza lookup registro. Il registro serve
solo per la F5 (promozione: locale + indicizza condiviso). Multi-workstation OK.
7. Evidence: dentro ThothII. La cartella (229 markdown statici curati, 11M, nessun ETL)
si sposta in harness/evidence/. ThothII self-contained. Nota: revisionare per PII
prima di committare.
Onda 0b aggiornata: cp evidence in harness/evidence/ invece di puntare path esterno.
Gap trovato in user review: la tabella vectors.memory aveva metadata
{type,session_id,tables,concepts} — mancavano subject/detail/rationale strutturati,
quindi l'hit vettoriale non bastava per applicare la memoria. ChironeWp3 faceva
lookup nel registro canonico via mem_id; stesso difetto ereditato in ThothII.
Decisione: arricchire il metadata del VectorRecord memory con subject/detail/rationale
(in memory_vector_records, nsp/memory.py). pack_metadata (rest_writer.py:26) li
serializza gia' nel jsonb via **record.metadata. Nessuna modifica al writer RPC,
nessuna modifica allo schema DB. search_similar proietta gia' metadata completo ->
la F2 ricostruisce la decisione direttamente dall'hit, senza lookup registro.
Momento ideale: tabella memory vuota, niente re-indicizzazione. Aggiunto come task
esplicito in Onda 3 (dove si porta memory_cmd). Registro globale resta source-of-truth
per la promozione (F5), ma la F2 legge solo dal vectordb.
Buco trovato prima della user review: lo spec claims D14 value-grounding ed
evidence-based F4, ma non setup né evidence né l'indice LSH. Senza, la sessione
L2 girerebbe degradata (solo segnali vettoriali) e i claim sarebbero falsi.
Onda 0b (dopo Onda 0 + 4, prima della sessione L2):
- Evidence: cablare THOTH_DOCS_ROOT=/Users/mp/Chirone/chirone/etl/docs nel .env +
blocco evidence nel chirone-test.yaml. La cartella esiste già.
- LSH: nsp lsh build sul workspace chirone-test (one-shot, richiede VPN + Ollama).
Verifica: nsp search ritorna match multi-colonna + test_value_grounding_real
smette di skip-piare.
Ordine esecuzione aggiornato (6 step), D14a value-grounding marcato 'sì (se Onda 0b)',
nsp lsh build tolto dal fuori-scope (ora dentro).
Bug trovato provando la connessione reale col .env: il write endpoint vive su un
PATH DEDICATO /vector/write/v1/ (non /vector/v1/), e il modello Config ha write_rest/
vector_write_rest a TOP-LEVEL (non nidificati in vector_db).
- .env.example: aggiunge THOTH_VEC_WRITE_REST_URL (path dedicato del writer, con
avviso che le due chiavi valgono su path separati).
- workspaces/chirone-test.yaml: riscritto allineato a chirone.example.yaml + config.py
(vector_rest/vector_write_rest top-level; write_rest punta a THOTH_VEC_WRITE_REST_URL).
- tests/l2/*: corretti gli accessi strutturali (ws.vector_write_rest invece di
ws.vector_db.write_rest; ws.vector_rest invece di ws.vector_db.rest).
test_value_grounding_real skip-when-import-fails su nsp.lshindex (modulo deferred da B3).
Verificato end-to-end: save_one_memory (embeddings -> writer REST /vector/write/v1/
-> upsert pgvector -> read-back reader) PASSED. Suite L0+L1: 109 passed. Suite L2:
4 passed, 1 skipped (lshindex deferred).
Nota operativa: THOTH_SSL_CA va lasciato VUOTO sulla workstation (cert GoDaddy
pubblico in certifi). I campi direct-transport (THOTH_DB_*, THOTH_VEC_PASSWORD)
sono obbligatori per il modello ma inutilizzati in transport=rest: riempiti con
dummy nel .env locale (come faceva ChironeWp3).
Two fixes found while unblocking the L2 setup:
1. conftest: THOTH_SSL_CA is NOT an L2 prerequisite. The DWH endpoint presents a
public cert (*.policlinicosandonato.it, signed by GoDaddy), already in the
certifi bundle, so the REST clients validate TLS with verify=True -- no CA file
needed. The ssl_ca line was commented out in ChironeWp3's nsp.yaml too.
2. test_workspace: the profile-default assertion collided with the operator's real
harness/.env once load_dotenv (D3) started injecting THOTH_PROFILE into the
process env. The test now dels THOTH_PROFILE to assert the actual *default*
(server), regardless of what the operator set in .env.
Suite: 109 passed.
README: install, configure (.env + workspaces/), the workflow, run inside Pi, the
three-level test commands, layout, references.
docs/workflow-editing.md: how to edit workflow.yaml (add/reorder/merge/skip phases,
advance kinds, prerequisite predicates, decision_min_phase derivation, artifacts_out
+ teardown) -- referencing spec §5.3. Emphasizes no mirrored constants (the F2 point).
docs/testing.md: the honest L0/L1/L2 split in plain language -- what each covers and
does NOT. States the headline plainly: the skill->LLM->gate loop has NO automated
regression coverage (L2 only, pre-release). Documents the fake-Pi follow-up as the
gap-closer. Security note on keys (.env gitignored, never logged, rotate leaked keys).
Pre-release, non-deterministic tests (marker l2, skipped without .env + VPN). They
close the gaps L1 leaves open: real value grounding on the live schema, real memory
save-one upsert to pgvector, and the full GLM 5.2 -> gate conversation on the
'ablazione' question (which exercises D14 value grounding + formula on a multi-
column case + the gate glue L1 cannot reach).
workspaces/chirone-test.yaml points at the remote endpoints (DWH read-only +
pgvector dual-key, TLS self-signed); secrets via ${THOTH_*}.
- test_session_ablazione: precondition checks (workspace loads, env present, pi on
PATH) + the documented manual run protocol (human-in-the-loop; scripted-answers
variant is a follow-up). Default run skips cleanly.
- test_value_grounding_real: 'ablazione' grounds to multiple columns on the real
schema (D14a non-collapsing), needs a built LSH index.
- test_memory_save_one_real: save_one_memory upserts one row via the writer key
(D11) and search_similar retrieves it via the reader key.
Operator runs before release (pytest -m l2). Default run: 109 passed, 5 skipped.
conftest loads harness/.env once (session, autouse) via python-dotenv, and exposes
an l2_env fixture that SKIPS (not fails) when any L2 prerequisite var is missing/
empty: THOTH_DWH_API_KEY, THOTH_VEC_API_KEY, THOTH_VEC_WRITE_API_KEY, THOTH_SSL_CA.
So the default run (pytest = L0+L1, addopts '-m not l2') stays green without .env;
only pytest -m l2 (pre-release, with .env + VPN) exercises them. l0/l2 markers were
registered in A9. tests/l2/ package created for the L2 tests (D4, D5).
settings.json (theme: thothii-mono), the two slash-command prompts
(/nuova-domanda, /riprendi-sessione), and the theme JSON. Renamed PsdWp3 -> ThothII
in prompt prose and the theme name. No tests (config files). pi --mode rpc launched
with cwd=harness/ finds the .pi/ directory + the extensions (nsp-gate.js, gate/).
Pure-logic smoke (no LLM, no DB) that builds a synthetic ledger by hand and asserts
the Phase-A substrate stays coherent: full F1->F8 walk reaches terminal phase
(max+1); rollback truncates the effective view (stale phase-7 decision excluded
after reopen to F4) and resets current_phase; teardown deletes artifacts beyond the
target while preserving the target phase's; re-approve after rollback advances
correctly; taskdoc stays under byte budget across all phases; decision_retraction
excludes the retracted seq + the marker itself from effective_decisions.
This is the CI-runnable coherence net for the L2 session test (which exercises the
LLM->gate loop that L1 cannot).
Rewrite of ChironeWp3's gate extension. The pure widget-descriptor CONSTRUCTION
is in ./gate/builders.js (L1-tested, C1); this file is the GLUE -- it depends on
the Pi runtime (pi.on, pi.registerTool, ctx.sendRaw) and is verified end-to-end at
L2 (Task D4), NOT unit-tested here. A fake-Pi runtime mock (cross-cutting
follow-up) would let it run in CI.
PRESERVED VERBATIM (load-bearing runtime glue, spec D4):
- anti-bypass tool_call hook: FORBIDDEN (nsp phase advance|reopen, decision add,
cte plan) + PROTECTED_FILES (review_decisions.jsonl, session_manifest.yaml,
cte_plan.json)
- input lock + the input hook: /nuova-domanda|/riprendi-sessione entry detection,
free-input block, the `!`-prefixed steer channel
- before_agent_start kickoff injection + the two kickoff payloads (model prose)
- agent_end prose safety net (nudges the model back to reviewer_* tools)
- session_start state reset
- exit-code contracts with the CLI (5 = gate refusal, 6 = needs human,
7 = not-ready silent no-op)
- textResult / nsp() / relayIfNspFails / advanceIfReady helpers
TWO CORRECTIVE CHANGES vs source:
1. F2 single source: workflow facts (max_phase, phase names, schema-linking phase)
come from `nsp phase meta --json`, NOT from JS-mirrored constants. The source's
PHASE_NAMES array (truncated to 7) is gone; F8/datamart can no longer drift.
2. D2/D4 widget-descriptor: reviewer interaction is emitted as a widget-descriptor
(built by ./gate/builders.js) and awaited by id via emitAndWait + the
extension_ui_response dispatcher. This replaces the source's blocking native TUI
primitives (ctx.ui.select/custom) and introduces the correlation-by-id layer
ChironeWp3 never had.
Four tools wired: reviewer_select, reviewer_decide (persists via nsp decision add),
reviewer_confirm (gate; privileged action on approve), rewrite_question. Plus the
/torna slash command for rollback. No-limbo invariant preserved: cancel/undefined
re-presents the widget; real escapes are always in the descriptor's reserved field.
The gate's widget-descriptor CONSTRUCTION, extracted into pure testable functions.
Each builder turns plain params into a ui_request descriptor (spec §4.1 taxonomy):
buildSelectRequest, buildMultiselectRequest, buildArtifactGate, buildInfoRequest,
buildFreetextRequest, withChildLinkage. No Pi context, no I/O -- the part of the
gate fully testable in L1 (in JS, in-language, no Python mirror).
Validation in the builders (not just happy-path): select requires title + array
options; multiselect allow_empty:false with zero options throws (a broken widget);
artifact-gate requires an artifact with a kind + a valid action.kind
(confirm/approve_reject/view_only); info level must be info/warning/error. The
Altro escape hatch with freetext linkage is always injected on blocking pick
widgets (no-limbo invariant).
L1: 14 node:test cases -- 3 golden files (select_F1, multiselect_F4,
artifact_gate_F5) pin the exact descriptor shape; fuzzy tests assert bad params
throw clearly rather than silently producing a broken widget.
package.json wires 'npm test' -> node --test (runs alongside pytest). The gate
GLUE (emission, anti-bypass, no-limbo loop) is C2, verified at L2.
D13 instructs the model to evaluate Altro/Rifiuta/steering free text in context,
act on it, re-ask if ambiguous, and record the user's words in the decision
rationale. The actual interpretation is model behavior enforced by the skill prose
+ gate, validated at L2; this test pins the RECORDING contract the gate relies on:
free text from 'Altro' round-trips into the decision rationale and survives
persistence, never silently discarded.
L1: test_freetext_interpretation (4 tests) -- Altro text preserved, persistence
roundtrip (exact), multiline steering intact, empty rationale allowed.
The skill prose ('Interpretazione del testo libero') ports with the .pi/ skill
in Phase D.
New evidence/formula_store.py: ConceptFormula (concept, columns, sql, status,
sources) as a frontmatter-YAML + SQL-body unit, stored one-file-per-formula under
<formulas>/<slug>-<n>.sql.md. save_formula is append-only (competing drafts and
reviewed versions coexist); retrieve_formula(concept) returns all of them so the
gate can surface candidates and let the reviewer choose.
concept_formula_approved / concept_formula_rejected added to DecisionType
(records the reviewer's choice; approved formulas travel with schema-linking).
L1: test_formula (7 tests) -- retrieval by concept, save/reload roundtrip (SQL
body preserved, frontmatter well-formed), multiple formulas per concept, empty
on no-match / missing dir, decision-type existence, default draft status.
Deferred: --kind formula on nsp search (needs search_cmd porting) wires
retrieve_formula into the CLI; lands with the search command.
Ports search/__init__.py (combined_search/RRF/aggregate) renamed psdwp3->nsp.
New aggregate_lsh_multi (the D14a deviation): groups LSH hits by table keeping
EVERY column where a value appears -- NOT collapsed to a single best column.
The old _aggregate_lsh hid alternative groundings (e.g. 'ablazione' matching both
a boolean flag and a free-text patologia field). aggregate_lsh_multi exposes all
columns so the value-grounding widget lets the reviewer choose the anchor(s).
Within one (table, column) the best-scored value is kept; columns ordered by score.
value_grounded added to DecisionType (records the reviewer's anchor choice).
L1: test_value_grounding (6 tests) -- multi-column exposure, grouping, within-column
best-value, ordering, empty, and the value_grounded decision-type existence.
Deferred: lshindex/ (needs vendor/thoth_lsh) and the L0 test_rrf.py land with the
nsp lsh build command + index-building path; not needed for the pure L1 core here.
The D11 deviation is a single-row pgvector upsert, not a full vectorstore resync.
Ports memory.py + session/{store,artifacts} + textutil (deps of memory), renamed
psdwp3->nsp. Decision import paths rewired from nsp.session.decisions to nsp.decisions
(our A4 port lives at the top level). session/models.py left UNCHANGED to preserve
the Phase-A ThothII additions (D12/D15 author/summary, D14a grounded_values,
D14b concept_formulas).
New in memory.py:
- memory_vector_record_for_decision(records, decision_seq): the single VectorRecord
for a chosen decision (reuses memory_vector_records, filtered to one).
- save_one_memory(records, decision_seq, writer, embedder): embeds one record and
calls writer.upsert_records('memory', [row]) -- NEVER writer.sync (that's the
full-resync, server-side-only path). Returns the upsert count.
L1: test_memory_save_one (5 tests) pins the contract -- single row, one upsert
call, sync never called, None/0 for unknown seq.
Deferred: the full nsp memory save-one CLI command (config/session loading + the
workstation write-guard) lands when memory_cmd.py is ported alongside the other
CLI commands. The pure D11 core is what L1 can honestly cover here.
Ports vectorstore/{rest_client,rest_writer,store,reader,embeddings,records},
evidence/model (leaf dep of records), and cli/_guards (require_vector_write_allowed
workstation write-guard). Renamed psdwp3->nsp, verbatim.
VectorRestClient gains an api_key property so reader/writer clients carry their
distinct keys visibly (spec D11: vector_reader / vector_writer on the same endpoint).
scripts/create_vector_reader_rpc.sql is NEW: the reader RPCs (search_similar,
list_tables) lived server-side in Supabase and were never versioned. Authored now
mirroring the writer allowlist pattern (table allowlist, security definer, revoke
from anon/authenticated, grant to vector_reader only). Writer RPC ported verbatim.
L1: test_vector_dual_key (7 tests) pins the dual-key construction + the workstation
write-guard (exit 4 without writer key).
nsp.cli app registers phase_app (the gate's workflow-fact source). phase meta
--json emits {schema_version, max_phase, phases:[{num,id,name,advance,artifacts_out}]}
from load_workflow() -- the single source of truth. F8/datamart is present (the
exact JS-drift bug in ChironeWp3, PHASE_NAMES truncated to 7, is structurally gone).
Scope: the 12 other command groups land in their porting tasks (A9 ports
db/mschema/_guards; B1 vectorstore; B3 search/lshindex). Eager-importing them now
would break the app on unported deps -- deferred to keep the suite green at each commit.
Genera un documento compatto per fase, derivato da artefatti + effective_decisions,
con byte budget enforced (target <20k token per un 35B/<200k). Mai incorpora
physical.yaml (~190k token, fatale). D15+D16 complementari: il brief delle decisioni
e' effective-aware, quindi post-rollback riflette lo stato corretto (le stale di
fasi > current_phase sono escluse).
6 tests (question+schema, no physical.yaml, budget ok/violato, header fase,
stale-excluded post-rollback). 32 total passing.
Cancella ogni artefatto la cui fase produttrice > target, usando artifacts_out di
workflow.yaml. Risolve il bug latente di ChironeWp3: ctes/*.sql orfani (non piu'
nel piano dopo un re-derive) restavano su disco e bloccavano finalize.
Da chiamare insieme all'append di phase_reopened per mantenere stato coerente.
6 tests (target 4/1/7, empty, missing, orfani CTE). 26 total passing.
The single most important architectural fix vs ChironeWp3: ALL helpers consult
effective_decisions() instead of raw list_decisions(), so the reopen-aware view is
consistent everywhere (fixes the bug where approved_ctes/advance_problems conflated
stale pre-reopen decisions with new ones).
Model (corrected during TDD):
- current_phase folds the audit (excluding retracted) with guard 'n == cur' --
already reopen-aware (old phase_approved:N after reopen to M<N don't advance).
- effective_decisions = decisions whose phase <= current_phase. A sql_approved at
phase 7 is stale when current_phase=4 after a rollback to F4, even if in the ledger.
Rollback to F4 does NOT invalidate decisions of phases 1-3 (they stay effective).
- decision_retracted markers excluded (audit-only).
Also: session/models.py ported (SchemaLinking + Candidate with grounded_values D14a
+ concept_formulas D14b). MAX_PHASE/PHASE_NAMES read from workflow.yaml via
load_workflow() (no duplication). Strada 2: ladder if-phase-N kept for now,
generic prerequisites evaluator (F2 full) deferred.
7 phase tests + 20 total passing.
- ported from ChironeWp3 (22 DecisionType, append-only jsonl, monotonic seq)
- added decision_retracted type + retracts field for step-level rollback (D15):
the retracted decision stays in the audit log, effective_decisions() (Task A5)
will exclude it from the active view
- 5 tests: retract marker + monotonic seq + retracts default + empty session +
literal includes retracted. All 13 harness tests pass.
La prima stesura usava una struttura 'ideale' (relational/vector_db.collection/
embeddings.provider) che non combaciava col modello Config portato da ChironeWp3.
Allineato alla struttura reale (database/rest/vector_rest/vector_write_rest/
vector_db top-level). Aggiunta nota di allineamento + modello delle key D11.