Commit Graph
61 Commits
Author SHA1 Message Date
marcopan 717e5ecced refactor(dwh): adapt direct and REST transports 2026-07-11 20:03:31 +02:00
marcopan a4eb6cc9e5 refactor(dwh): define adapter contract 2026-07-11 19:57:39 +02:00
marcopanandClaude Opus 4.6 1e4bc11418 fix(embed): fast-fail + auto-restart Ollama on solved-search hang
Embeddings timeout was 120s, causing multi-minute hangs when Ollama was
down during F4/F6/F7 solved-search. Now: connect_timeout=5s across all
HTTP clients (REST + Ollama), read_timeout reduced to 30s for embeddings,
and OllamaEmbeddings auto-restarts the server on ConnectionError before
degrading gracefully.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-07-07 19:49:55 +02:00
marcopanandClaude Fable 5 e24b41b156 feat(opt): three efficiency levers for NL→SQL workflow
Lever 1: Join-graph via FK logics in annotations + suggest-fks command
  - TableAnnotation.foreign_keys field stores curated logical FKs (DWH has no FK constraints)
  - tht schema suggest-fks: mine from approved SQL, heuristics (time_key → dim_time),
    same-name discovery + explicit --assume flag for multi-owner PKs
  - mschema renders 【Foreign keys】 section populated; validation in merge.py
  - SKILL.md F4 now reads FKs from mschema-text, no custom data_time_key logic

Lever 2: Context-pack consolidation at kickoff (tht search pack)
  - Single embedding of question, reused for schema + evidence + solved searches
  - One command: tht search pack <question> --session <id> → retrieval_pack.md
  - Graceful degradation when Ollama/vector store unreachable (exit 0, empty sections)
  - SKILL.md F1 prescribes as first call; reduces model thinking turns via pre-retrieval

Lever 3: Phase-summary recap v2 auto-construction from session ledger
  - tht session show --json includes full decisions ledger
  - tht phase meta --json exports 'emits' (substantive decision types per phase)
  - Gate appends deterministic 【Decisioni registrate in questa fase】 section (appendLedgerSection)
  - Model authors only summary + checks; recap table comes from persisted state (exact by construction)
  - SKILL.md Disciplina 6: brief model output, gate fills the rest

Tests: 358 Python (including 10 FK + 3 pack + 1 session-ledger tests) + 111 JS gate tests, all pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 17:43:08 +02:00
marcopanandClaude Fable 5 34aeda006e perf(f1): cache-guard schema introspect + F1 toolbox in skill (~-5 min per session)
Transcript analysis (session 2026-07-06-175012, GLM 5.2) showed F1 at 567s:
182s wasted on a useless `tht schema introspect` (re-introspecting the remote
DWH although physical.yaml was already materialized) plus ~220s of model
thinking inflated by ~7 exploratory turns (--help/find/cat). The actual
searches cost ~15s; reviewer gates (~145s, untouched) are the quality contract.

- schema_cmd.py: introspect now exits 0 with "OK (cache)" in ~1s when
  physical.yaml exists; --refresh forces the real re-introspection.
  Deterministic cross-model guarantee, verified live on psd (163 tables, 1.2s).
- SKILL.md: F1 toolbox (only `tht search find` + `tht schema render`; no
  introspect/--help/filesystem browsing; batch all searches in one turn);
  F4 step 1 is render-only with a one-shot introspect fallback.
- tht-gate.js: `tht schema introspect ... --refresh` added to FORBIDDEN
  (maintenance stays shell-only, never in-session).
- tests: 4 new pytest cases (cache hit placement proven with fake credentials,
  refresh bypass, corrupt-catalog fall-through, render fallback message) and
  2 gate anti-bypass JS cases.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 16:02:45 +02:00
marcopanandClaude Fable 5 a26f16ad79 feat(vector): server-side kinds filter for search_similar (legacy fallback) + graceful solved-search degrade
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 14:23:28 +02:00
marcopanandClaude Fable 5 1f429b6b35 feat(cli): tht memory solved-index / solved-search (question->SQL exemplars)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 13:43:52 +02:00
marcopanandClaude Fable 5 c0324c127a feat(solved): solved_question vector kind + one-row upsert (D11 pattern)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 13:38:10 +02:00
marcopanandClaude Fable 5 bf7e850e6b feat(memory): filter gate-declined candidates from promotion preview
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 13:26:59 +02:00
marcopanandClaude Fable 5 cadbdf177a feat(memory): memory_promoted/memory_promotion_declined decision types (F8 emits)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 13:22:07 +02:00
marcopan 9b4f6b9804 feat(harness): persist CTE preview rows, add tht cte info, --doc for cte plan
WS1 of review-gates-v2: gives the JS gate (WS2) deterministic data to build the
cte_result v2 payload.

- CteTestRecord gains optional preview_rows (JSON-coerced, truncated cells);
  test_cmd populates it from the bounded result rows.
- New read-only `tht cte info <name> --session <id> [--json]`: plan
  index/total, persisted .sql, cte_plan_doc.json entry (if any), last
  CteTestRecord. Exits 1 with a clean stderr message on missing
  session/plan/name/sql.
- `tht cte plan --doc -` validates a chain-doc JSON (ctes[].name must match
  --name, same order) and writes it to cte_plan_doc.json; cte_plan.json stays
  a plain list[str] (load-bearing for tht.phase.next_cte). --doc is optional.
2026-07-07 00:23:50 +02:00
marcopanandMarco Pancotti 3ed570ee60 feat(tht): promoted_columns_for helper (curated column set) 2026-07-06 23:25:00 +02:00
00365cc6e6 feat(tht): sync-schema-linking projects F4 ledger into schema_linking.json
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 23:25:00 +02:00
marcopanandMarco Pancotti 1df5d40d05 feat(tht): column_promoted/column_excluded decision types (F4) 2026-07-06 23:25:00 +02:00
marcopanandMarco Pancotti 9339a8272b feat(tht): schema columns reader (name/description/type/pk) from catalog 2026-07-06 23:25:00 +02:00
marcopanandClaude Opus 4.8 00777922d8 feat(cli): 'tht session set-schema-linking' (file/stdin, validated)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-01 16:44:36 +02:00
marcopanandClaude Opus 4.8 92caaac18d feat(store): set_schema_linking validates then writes the F4 artifact
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-01 16:40:04 +02:00
marcopanandClaude Opus 4.8 47a52170ce feat(cte): add 'tht cte next' — first unapproved plan CTE
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-01 16:29:52 +02:00
marcopanandClaude Opus 4.8 c12bdcd257 feat(harness): derive a 3-5 keyword session name (YAKE, no LLM)
New sessions get a concise Italian-keyword `name` instead of the truncated
question. `tht session new` (when no --name is given) derives it via a new
`_extract_name` helper using YAKE (pure-Python, unsupervised, Italian, no LLM),
dropping generic query verbs and keeping the top keywords in reading order;
falls back to `_summarize` if YAKE is unavailable. `create_session` core keeps
its `name=None` default — the policy lives at the CLI layer.

TDD: tests/test_session_name.py (unit + CliRunner integration). Full harness
suite 269 passed; ruff clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 15:54:54 +02:00
marcopanandClaude Opus 4.8 418187a4ad fix(backend): echo Pi's RPC id so reviewer gates unblock after answer
ctx.ui.input in `pi --mode rpc` correlates extension_ui_response on its own
top-level RPC id (crypto.randomUUID), not the descriptor id the gate carries
in `title`. SessionBridge replied with the descriptor id, so Pi silently
dropped the response and the model never resumed — every reviewer widget hung
after the human answered.

SessionBridge now stores Pi's top-level m.id (pendingPiId) and replies
extension_ui_response{ id: pendingPiId, value: <uiResponse> }; value still
carries the descriptor id so the gate's internal resp.id === descriptor.id
check still holds.

The fake-pi double had masked the bug by forcing m.id == descriptor.id; it now
mirrors real Pi (distinct randomUUID, correlate on it, drop unknown ids), with
a negative regression test. SKILL.md Phase 1 also now steers multi-answer
disambiguation to reviewer_decide (multiselect).

Tests: backend 67/67, tsc clean, fake-pi contract 2/2.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 10:43:49 +02:00
marcopan deb8081bc2 test(harness): end-to-end tht ollama ensure via root app 2026-06-29 19:30:00 +02:00
marcopan 7e19f81173 feat(harness): tht ollama ensure CLI command (hard-fail preflight) 2026-06-29 19:25:23 +02:00
marcopan 9c39ea129e fix(harness): ensure_ollama never raises on start/probe failure (stage server) 2026-06-29 19:19:26 +02:00
marcopan 5f5cac05d5 feat(harness): EmbeddingsConfig bin/start_cmd + ensure_ollama preflight orchestration 2026-06-29 19:15:28 +02:00
marcopan 03053b0716 feat(harness): tht session documents --json read command 2026-06-29 12:25:40 +02:00
marcopan 38cd4889f4 style(harness): move test_session_mutations imports to top, drop unused json 2026-06-29 12:23:48 +02:00
marcopanandClaude Opus 4.8 26337d0e32 feat(harness): tht session set-name/set-group/archive/unarchive/delete commands
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 12:20:43 +02:00
marcopan 60fa665d73 feat(harness): manifest archived/group fields + session mutation helpers 2026-06-29 12:16:52 +02:00
marcopanandClaude Opus 4.8 ec36ee421a feat(backend): POST /sessions applies global settings, body is question-only
workspace/provider/model/thinking now come from getSettings() injected into
sessionRoutes; the request body supplies only question+name. Also teaches
fake_pi_rpc to respond to set_model and set_thinking_level RPC commands so
tests that pass real model settings don't hang.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 16:16:08 +02:00
marcopanandClaude Sonnet 4.6 d630cac3ab feat(task-10): SQL routes + /workspaces + /models + fix sql preview path bug
Harness:
- preview_cmd FILE positional arg made optional; when omitted with --session,
  path is derived via _session_sql_file (mirrors export_cmd) — fixes the
  deferred Task-5 bug where the backend passed sessions/<id>/sql_final.sql
  relative to harnessDir, which broke for workspace-dependent paths.
- New pytest: test_preview_session_no_file_resolves_sql_final

Backend:
- ThtRunner.sqlPreview: drop positional file arg; use --session only
- New routes/sql.ts: POST /sessions/:id/sql/preview + /export
- New routes/meta.ts: GET /workspaces (yaml scan) + GET /models (injectable
  seam + graceful fallback to {models:[]})
- app.ts: register sqlRoutes + metaRoutes; add listModels to BuildAppDeps
- tht-runner.test.ts: add sqlPreview argv assertion (no file path)
- test/routes-sql-meta.test.ts: 9 tests (sql preview/export + meta routes)

Tests: harness 233 passed; backend 29 passed; build clean.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-27 21:31:35 +02:00
marcopanandClaude Sonnet 4.6 17c277bb99 test(harness): fake-pi-rpc protocol double + F1 widget contract golden (D10)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-27 20:36:23 +02:00
marcopanandClaude Sonnet 4.6 5d9a0bb548 feat(harness): manifest provider/model/thinking/name + session new options (BE-6/7)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-27 20:29:40 +02:00
marcopanandClaude Sonnet 4.6 5e0e1cee03 test(harness): pin session list/show json keys + schema alias
- test_list_sessions_required_keys: assert full spec key set
  (adds summary/updated_at/author) so dropping any goes caught.
- test_cli_show_json_valid: assert "schema" in / "db_schema" not in
  data to pin by_alias=True on the alias-sensitive field.

Production code unchanged. 8 passed; full suite 230 passed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-27 20:27:18 +02:00
marcopanandClaude Sonnet 4.6 a10a872a99 feat(cli): tht session list --json + session show --json (Task 7)
- Add _list_sessions(sessions_root) pure helper: scans sessions dir,
  returns list[dict] with id/status/question/summary/created_at/
  updated_at/author, sorted by created_at desc.
- Add `tht session list` command: --json emits pristine JSON array,
  human mode prints one line per session.
- Add --json flag to `tht session show`: emits manifest
  (model_dump by_alias) + phase (current_phase) + has_schema_linking.
- Tests: 8 tests in test_session_list_json.py (TDD red→green).
- Full suite: 230 passed (--ignore=tests/l2).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-27 20:23:46 +02:00
marcopanandClaude Opus 4.8 85e2fdc00a fix(harness): correct truncated signalling for sql preview offset>0
When offset>0 the wrapper added an outer LIMIT N, so run_controlled's
_inject_limit bailed (a LIMIT IS present) and truncated was always False —
AGGrid could never detect more rows. Fix: for offset>0 probe with LIMIT (N+1)
OFFSET M, then compute truncated = len(rows) > N in do_run and slice back to N.
offset==0 path unchanged (delegates to extracted _run_transport helper). JSON
still reports the user's requested limit N and correct truncated. Adds 3 tests
exercising the real do_run offset>0 path (N+1 -> truncated True, N -> False,
offset==0 verbatim). 222/222 passing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 20:18:58 +02:00
marcopanandClaude Sonnet 4.6 7d9769cfff feat(harness): tht sql preview --json + --offset for AGGrid paging (BE-2)
Add inject_limit_offset (tht/execute/limit.py) — pure subquery wrapper that
applies LIMIT/OFFSET non-destructively without clobbering user-supplied LIMITs.
Wire offset param into do_run (pre-processing when offset>0) and add --offset /
--json flags to preview_cmd; JSON mode emits pristine stdout with columns, rows,
execution_ms, truncated, limit, offset. 5 new tests (4 unit + 1 JSON-purity),
219/219 total passing (no regressions).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-27 20:13:20 +02:00
marcopanandClaude Opus 4.8 c4d130828f fix(harness): remediation difetti review — gate↔CLI, D15, D7/D6, D14, robustezza
Implementazione del piano di remediation progressiva sui difetti emersi
dall'analisi dell'harness. Tutto verificato: 214 test Python (incl. L0 su
Postgres reale), 14 test JS del gate, ruff pulito.

Blocco 1 (CRITICA, integrazione gate↔CLI):
- phase advance: gate usa --auto + exit 6; reviewer_confirm kind:phase fa
  advance esplicito che applica i prerequisiti (prima non avanzava per le
  fasi a conferma umana).
- cte plan riceve i --name dal gate (param names); set-question con id
  posizionale; skill `tht search find`; nuovo comando `tht memory save-one`
  con dedup hash client-side in save_one_memory.

Blocco 2 (D15, stato post-rollback):
- campo `phase` su DecisionRecord + effective_decisions phase-aware per i
  subject "a nome" (cte_approved ecc.); _compute_promotions e finalize sulla
  vista effective; finalize confronta col piano CTE effettivo, non glob;
  `decision add --retracts` + comando `decision retract`.

Blocco 3 (D7 read-only + D6 manifest):
- assert_read_only su tutti e quattro i codepath (direct + REST);
- manifest author/summary/updated_at/updated_by/schema_version popolati +
  helper touch_manifest sulle mutazioni.

Blocco 4-5 (D14a/D14b):
- decision_min_phase data-driven via `emits:` in workflow.yaml;
- formula evidence: status auto, search_formulas, gruppo CLI `tht formula`,
  `search find --kind formula`, load_evidence_dir salta i .sql.md.

Blocco 6 (robustezza):
- taskdoc slice promoted_tables + bound enforced; report escaping/bound +
  rsplit note; filtro kind reader REST/direct; conteggio upserted robusto;
  guard REST run_query non-list; LSH disallineato -> LshIndexError.

Blocco 7 (pulizia):
- dead code gate e KIND_TO_TABLE morto rimossi; doc Postgres-only
  (README + connection.py).

Blocco 0 (parziale): test di compatibilità firma gate↔CLI
(tests/integration). Rinviati: fake-Pi runtime completo, artifact-gate da
disco (#23), parità eligibility REST/direct (#28), unificazione
reserved-labels (#30), memory_rejected da deselezione (#33).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 17:16:51 +02:00
marcopan 0a9ebddaec feat(harness): Onda 0b — setup workspace per-cliente + build LSH (D14a validato)
Repo workspace per-cliente tht-workspace-psd/ (git separato): 35 evidence Thoth
frontmatturate (da etl/docs/evidence/, non tutto etl/docs), psd.yaml con path
assoluti ancorati al repo, THT_DOCS_ROOT -> radice repo cliente. README documeta
il deployment shape (2 checkout + .env).

LSH build sul DWH reale (REST, VPN): 75737 valori / 491 colonne eligible -> indice
175M nel repo cliente. tht lsh query 'ablazione' -> 8 colonne (D14a non-collapsing
validato end-to-end: procedure_type, descrizione_procedura, intervento, ...).

Correzioni al piano eseguite durante l'implementazione:
- aggiunto 'tht schema introspect' (prerequisito di lsh build, omesso nel piano)
- uso -c (il cmd ha --config, non --workspace residuo)
- path assoluti nel psd.yaml (paths sono relativi alla CWD, non al file YAML)
- evidence corretta: solo etl/docs/evidence/ (35), non tutto etl/docs (895)

Test L2 riallineato: WORKSPACE -> psd.yaml nel repo cliente (non tht-test.yaml),
index_dir -> paths.indexes/'lsh', nome -> db_schema (non letterale 'datawarehouse').
Bug latente del test (path mismatch) mai emerso prima: ora PASS invece di SKIP.

Verifica: pytest 165 passed, 5 deselected; pytest -m l2 test_value_grounding_real PASS.
2026-06-27 15:31:03 +02:00
marcopan 159207a8f1 feat(harness): arricchisci metadata memory (subject/detail/rationale) — Onda 3.1 TDD
Correzione del gap ereditato (resosi NECESSARIO dal drop del registry, spec 5): il
metadata del VectorRecord memory ora porta subject/detail/rationale oltre a
type/session_id/tables/concepts. pack_metadata li serializza nel jsonb via
**record.metadata. search_similar proietta metadata completo -> la F2 ricostruisce
la decisione direttamente dall'hit, senza lookup registro.

L1: 4 test (subject/detail/rationale presenti, campi esistenti preservati,
no cross-contamination multi-record, save_one_memory propaga il metadata alla riga).
Suite: 165 passed.
2026-06-27 14:10:39 +02:00
marcopan a312746fc8 feat(harness): require_phase_or_exit + phase advance/reopen/show (Onda 1.2)
require_phase_or_exit: guard riscritto vs Workflow (load_workflow().phase_name invece
della costante PHASE_NAMES drift). Exit 1 se la sessione e' sotto soglia. Usato da
cte/decision/datamart cmd.

Comandi phase (portati + adattati al modello ThothII, non copia cieca):
- advance: persiste phase_approved; --auto exit 6 se la fase non e' completa
  (contratto col gate)
- reopen: persiste phase_reopened + teardown_to_phase degli artefatti oltre il target
- show: stato sessione (fase corrente, ultime decisioni)

session_dir helper tenuto qui (mirror di session_cmd) per evitare circular import.
_cfg() fa fallback a THT_WORKSPACE env finche' _load_config_or_exit (Onda 1.4) non
sara' portato.

L1: 4 test require_phase_or_exit (allow at/above, exit below, message con nome fase
dal workflow). Suite: 161 passed.
2026-06-27 13:15:12 +02:00
marcopan 37bb074efe feat(harness): Workflow.schema_linking_phase() (Onda 1.1, TDD)
Aggiunge il metodo che i cmd CLI useranno al posto della vecchia costante
SCHEMA_LINKING_PHASE (drift fix Onda 1). Ritorna il num della fase il cui
artifacts_out contiene schema_linking.json, default 5 se nessuna la dichiara.

L1: 4 test (fase reale F4, posizione arbitraria, default 5, artefatti multipli).
Suite: 157 passed.
2026-06-27 13:09:51 +02:00
marcopan 31552782c0 test(harness): L1 characterization tests per Onda 0 (sqlcheck, ctetest, execute)
44 test L1 sui 3 moduli backend con logica non banale (opzione 2 della user review):

- sqlcheck.validate_sql (16 test): parse/single-statement, read-only enforcement
  (INSERT/UPDATE/DELETE/CREATE/DROP/ALTER/TRUNCATE/GRANT rifiutati, WITH/UNION ok),
  forbidden functions (dblink default blacklist, custom set, allowed not flagged),
  object-existence (tabella inesistente, CTE non flaggata, perimetro promoted warning,
  colonna inesistente con alias). Documenta una limitazione reale: le funzioni
  aggregate specializzate (count/sum/coalesce) NON sono catturate dal name-matcher
  perche' sqlglot modella .name come argomento, non come nome funzione.

- ctetest (14 test): has_trailing_select (semantica controintuitiva: True = violazione),
  last_cte_name, build_test_sql, ledger I/O (load/append roundtrip, JSON-array e
  JSONL tolleranti, corrupt-ledger raise).

- execute._inject_limit (6 test): LIMIT iniettato quando assente (limit+1 per
  troncamento), rispettato quando presente, non iniettato su non-query, UNION/WITH ok.

Suite: 153 passed (109 + 44). Bonus: __psd_probe__ -> __tht_probe__ (riferimento
cliente neutralizzato in ctetest).
2026-06-27 13:04:34 +02:00
marcopan fc5fbe6b65 refactor(harness): renaming prodotto tht (Onda -1)
Thoth (tht) è il prodotto, PSD è il cliente. Nessun riferimento al contesto
clinico nel codice.

Rinomine:
- comando+package nsp→tht (dir nsp/→tht/, 46 import, pyproject entry point)
- gate nsp-gate.js→tht-gate.js (+ rewrite token, relayIfNspFails→relayIfThtFails)
- workspace chirone.{example,test}.yaml→tht.{example,test}.yaml (generici)
- env THOTH_→THT_ (19 var) + NSP_ stragglers (NSP_HARNESS_ROOT, NSP_SESSION)
- commenti/docstring chirone/psdwp3/policlinico neutralizzati ('the reference
  implementation', 'the DWH')

Aggiunto [tool.setuptools.packages.find] include=['tht*'] (necessario: l'auto-
discovery rompeva con tht/ + workspaces/ come top-level multipli).

.env operatore aggiornato in-place (prefissi THT_, valori preservati, gitignored).

Verifica: pytest 109 passed, npm test 14 pass, tht phase meta --json OK, zero
residui nsp/THOTH_/NSP_/chirone nel package.
2026-06-27 10:33:16 +02:00
marcopan e58f6c092e fix(harness): workspace + write-URL + L2 tests per accesso REST reale
Bug trovato provando la connessione reale col .env: il write endpoint vive su un
PATH DEDICATO /vector/write/v1/ (non /vector/v1/), e il modello Config ha write_rest/
vector_write_rest a TOP-LEVEL (non nidificati in vector_db).

- .env.example: aggiunge THOTH_VEC_WRITE_REST_URL (path dedicato del writer, con
  avviso che le due chiavi valgono su path separati).
- workspaces/chirone-test.yaml: riscritto allineato a chirone.example.yaml + config.py
  (vector_rest/vector_write_rest top-level; write_rest punta a THOTH_VEC_WRITE_REST_URL).
- tests/l2/*: corretti gli accessi strutturali (ws.vector_write_rest invece di
  ws.vector_db.write_rest; ws.vector_rest invece di ws.vector_db.rest).
  test_value_grounding_real skip-when-import-fails su nsp.lshindex (modulo deferred da B3).

Verificato end-to-end: save_one_memory (embeddings -> writer REST /vector/write/v1/
-> upsert pgvector -> read-back reader) PASSED. Suite L0+L1: 109 passed. Suite L2:
4 passed, 1 skipped (lshindex deferred).

Nota operativa: THOTH_SSL_CA va lasciato VUOTO sulla workstation (cert GoDaddy
pubblico in certifi). I campi direct-transport (THOTH_DB_*, THOTH_VEC_PASSWORD)
sono obbligatori per il modello ma inutilizzati in transport=rest: riempiti con
dummy nel .env locale (come faceva ChironeWp3).
2026-06-27 08:20:36 +02:00
marcopan 276717005d fix(harness): drop THOTH_SSL_CA from REQUIRED_L2 + isolate profile in workspace test
Two fixes found while unblocking the L2 setup:

1. conftest: THOTH_SSL_CA is NOT an L2 prerequisite. The DWH endpoint presents a
   public cert (*.policlinicosandonato.it, signed by GoDaddy), already in the
   certifi bundle, so the REST clients validate TLS with verify=True -- no CA file
   needed. The ssl_ca line was commented out in ChironeWp3's nsp.yaml too.

2. test_workspace: the profile-default assertion collided with the operator's real
   harness/.env once load_dotenv (D3) started injecting THOTH_PROFILE into the
   process env. The test now dels THOTH_PROFILE to assert the actual *default*
   (server), regardless of what the operator set in .env.

Suite: 109 passed.
2026-06-27 06:45:00 +02:00
marcopan c861a0df9e test(harness): L2 tests -- ablazione session + value grounding + memory save-one (D4, D5)
Pre-release, non-deterministic tests (marker l2, skipped without .env + VPN). They
close the gaps L1 leaves open: real value grounding on the live schema, real memory
save-one upsert to pgvector, and the full GLM 5.2 -> gate conversation on the
'ablazione' question (which exercises D14 value grounding + formula on a multi-
column case + the gate glue L1 cannot reach).

workspaces/chirone-test.yaml points at the remote endpoints (DWH read-only +
pgvector dual-key, TLS self-signed); secrets via ${THOTH_*}.

- test_session_ablazione: precondition checks (workspace loads, env present, pi on
  PATH) + the documented manual run protocol (human-in-the-loop; scripted-answers
  variant is a follow-up). Default run skips cleanly.
- test_value_grounding_real: 'ablazione' grounds to multiple columns on the real
  schema (D14a non-collapsing), needs a built LSH index.
- test_memory_save_one_real: save_one_memory upserts one row via the writer key
  (D11) and search_similar retrieves it via the reader key.

Operator runs before release (pytest -m l2). Default run: 109 passed, 5 skipped.
2026-06-26 23:18:34 +02:00
marcopan 806bc510d7 test(harness): L2 marker + skip-when-no-.env guard (D3, Testing Strategy)
conftest loads harness/.env once (session, autouse) via python-dotenv, and exposes
an l2_env fixture that SKIPS (not fails) when any L2 prerequisite var is missing/
empty: THOTH_DWH_API_KEY, THOTH_VEC_API_KEY, THOTH_VEC_WRITE_API_KEY, THOTH_SSL_CA.
So the default run (pytest = L0+L1, addopts '-m not l2') stays green without .env;
only pytest -m l2 (pre-release, with .env + VPN) exercises them. l0/l2 markers were
registered in A9. tests/l2/ package created for the L2 tests (D4, D5).
2026-06-26 23:16:40 +02:00
marcopan d5c0fffc2a test(harness): L1 session-coherence smoke -- full walk + rollback (D1)
Pure-logic smoke (no LLM, no DB) that builds a synthetic ledger by hand and asserts
the Phase-A substrate stays coherent: full F1->F8 walk reaches terminal phase
(max+1); rollback truncates the effective view (stale phase-7 decision excluded
after reopen to F4) and resets current_phase; teardown deletes artifacts beyond the
target while preserving the target phase's; re-approve after rollback advances
correctly; taskdoc stays under byte budget across all phases; decision_retraction
excludes the retracted seq + the marker itself from effective_decisions.

This is the CI-runnable coherence net for the L2 session test (which exercises the
LLM->gate loop that L1 cannot).
2026-06-26 23:14:57 +02:00
marcopan e6eeb2ae3e test(harness): free-text rationale-capture contract (D13, §4.6)
D13 instructs the model to evaluate Altro/Rifiuta/steering free text in context,
act on it, re-ask if ambiguous, and record the user's words in the decision
rationale. The actual interpretation is model behavior enforced by the skill prose
+ gate, validated at L2; this test pins the RECORDING contract the gate relies on:
free text from 'Altro' round-trips into the decision rationale and survives
persistence, never silently discarded.

L1: test_freetext_interpretation (4 tests) -- Altro text preserved, persistence
roundtrip (exact), multiline steering intact, empty rationale allowed.

The skill prose ('Interpretazione del testo libero') ports with the .pi/ skill
in Phase D.
2026-06-26 23:04:22 +02:00
marcopan f104c015a1 feat(harness): SQL formula evidence -- concept->formula units + retrieval (D14b)
New evidence/formula_store.py: ConceptFormula (concept, columns, sql, status,
sources) as a frontmatter-YAML + SQL-body unit, stored one-file-per-formula under
<formulas>/<slug>-<n>.sql.md. save_formula is append-only (competing drafts and
reviewed versions coexist); retrieve_formula(concept) returns all of them so the
gate can surface candidates and let the reviewer choose.

concept_formula_approved / concept_formula_rejected added to DecisionType
(records the reviewer's choice; approved formulas travel with schema-linking).

L1: test_formula (7 tests) -- retrieval by concept, save/reload roundtrip (SQL
body preserved, frontmatter well-formed), multiple formulas per concept, empty
on no-match / missing dir, decision-type existence, default draft status.

Deferred: --kind formula on nsp search (needs search_cmd porting) wires
retrieve_formula into the CLI; lands with the search command.
2026-06-26 23:03:29 +02:00