commit 561e7aef084f254bb883717027cc1863f6635c0f Author: mptyl Date: Fri Jun 26 22:00:43 2026 +0200 chore: init ThothII repo — gitignore references, baseline docs (PRD, spec, harness plan) diff --git a/.gitignore b/.gitignore new file mode 100644 index 00000000..667fe4f5 --- /dev/null +++ b/.gitignore @@ -0,0 +1,45 @@ +# === macOS === +.DS_Store + +# === Reference / consultation material (local-only, NOT ThothII deliverables) === +ChironeWp3/ +Thoth/ + +# === Visual companion brainstorming artifacts (local-only) === +.superpowers/ + +# === Python === +__pycache__/ +*.pyc +*.pyo +.venv/ +venv/ +*.egg-info/ +dist/ +build/ + +# === Node === +node_modules/ + +# === Secrets — NEVER commit === +.env +*.pem +ca-chain.pem +config/ca-chain.pem + +# === Runtime data (sessions contain PII; indexes are derived) === +harness/sessions/ +harness/indexes/ +backend/sessions/ +backend/indexes/ + +# === Test artifacts === +.pytest_cache/ +.ruff_cache/ +.coverage +htmlcov/ + +# === Editor === +.vscode/ +.idea/ +*.swp diff --git a/.kilo/kilo.jsonc b/.kilo/kilo.jsonc new file mode 100644 index 00000000..cc2fb594 --- /dev/null +++ b/.kilo/kilo.jsonc @@ -0,0 +1,7 @@ +{ + "$schema": "https://app.kilo.ai/config.json", + "indexing": { + "vectorStore": "qdrant", + "model": "sentence-transformers/all-minilm-l12-v2" + } +} \ No newline at end of file diff --git a/docs/superpowers/plans/2026-06-25-harness-implementation.md b/docs/superpowers/plans/2026-06-25-harness-implementation.md new file mode 100644 index 00000000..f59bf21b --- /dev/null +++ b/docs/superpowers/plans/2026-06-25-harness-implementation.md @@ -0,0 +1,1697 @@ +# ThothII — Harness Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Build `harness/` — the self-contained Pi layer (CLI `nsp` Python + gate extension JS + skills markdown + `.pi/`) that runs the 8-phase NL→SQL workflow emitting/consuming widget-descriptor JSON, derived from ChironeWp3 as a validated-and-adapted starting point (not assumed reliable). + +**Architecture:** Port the relevant ChironeWp3 files into `harness/` task-by-task, then adapt each to the ThothII contract: `workflow.yaml` as the single source of workflow truth (F2), an "effective decisions as of pointer" view + `teardown_to_phase` for correct rollback (D15), a per-step task-document generator with enforced bounds for a 35B/<200k model (D16), the dual vector API key + `memory save-one` (D11), free-text interpretation (D13), value/formula grounding (D14), and the gate rewritten to emit/consume the 6-widget descriptor taxonomy (D2/D4). + +**Tech Stack:** Python ≥3.12 (typer, pydantic v2, pyyaml, sqlalchemy 2.0, psycopg2-binary, datasketch, sqlglot, pytest, testcontainers[postgres]); Node/JS for the Pi gate extension; Ollama embeddings (nomic-embed-text-v2-moe); Pi + GLM 5.2 for L2 tests. Testing: two levels (L1 fake + L2 real), see Testing Strategy below. + +**Reference spec:** `docs/superpowers/specs/2026-06-25-thothii-architecture-design.md` (all D1–D16 and §4.*). + +**Source of ported code:** `/Users/mp/projects/ThothII/ChironeWp3/` (read-only reference; never modify it). + +--- + +## Testing Strategy — three levels (L0 testcontainers, L1 fake, L2 real) + +This plan uses a **three-level test strategy**. The split is mandatory because the skill→LLM→tool-call loop cannot be exercised without a real model, and the DB-touching ported code needs a real database to validate (the "not assumed reliable" principle is hollow without it). + +### ⚠️ The honest headline: the core of the system has NO automated regression coverage + +The **skill→LLM→gate loop** — the heart of the system — is covered **only by L2** (manual, non-deterministic, slow, requires credentials+VPN). **There is no automated regression protection on it.** This is a deliberate, conscious choice, not an accidental side-effect: the LLM is non-deterministic and requires a configured Pi + network, so it cannot live in the fast automated loop. **Consequence: agentic-behavior regressions surface at pre-release L2 runs, not at commit.** Accept this and plan L2 runs before any release. A partial automated net for this gap would be a **fake-Pi** that mocks the Pi runtime and lets the gate glue run in CI; that is documented as a **cross-cutting follow-up** (built alongside the backend plan, not in this plan). + +### L0 — testcontainers, local Postgres, runs on every local run (CI) + +**What:** integrity tests of ported DB-touching modules against a real Postgres in a Docker container (testcontainers). No LLM, no remote network. Docker is available on the dev machine, so L0 is always-on locally. + +**Dependencies:** Docker (present in the dev env). No credentials, no VPN. + +**Coverage:** the ported logic that speaks to a DB — `db/connection` (read-only enforcement, exit code 2 if writable), `db/introspect`, `db/sampling`, `db/execute`, `search/RRF` + LSH with real data, `vectorstore/store` direct-load path. This is where "not assumed reliable" gains real teeth for the data layer. + +### L1 — fake data, deterministic, runs on every local run (CI) + +**What:** logic-pure tests with fake data (`tmp_path`, fixtures, mocks). No DB, no LLM, no network. + +**Coverage (honest):** +- Python logic-pure: `workflow.yaml` loading, `effective_decisions`, `teardown_to_phase`, `generate_task_doc`, `aggregate_lsh_multi` (on fake hits), formula store read/write, `decision_retracted`, CLI contract tests (valid input → well-formed JSON). +- **Gate builder functions (pure, in JS, tested in JS)**: the widget-descriptor builders produce the correct JSON given params. Tested in-language (node:test/vitest), no Python↔JS bridge, no Python mirror. + +**Honest limitation (load-bearing):** L1 can test the gate **builders** (pure functions) but **NOT the gate glue** — registration, emission via `ctx.sendRaw`, the no-limbo loop, the anti-bypass hooks. The glue depends on the Pi runtime (`pi.registerTool`, `ctx.sendRaw`, `emitAndWait`) and cannot run without either a real Pi or a fake-Pi mock. **The glue is tested only at L2.** Fuzzy/negative tests on the glue (malformed tool params, out-of-order calls, orphan responses) likewise live at L2; only fuzzy tests on the pure builders live in L1. + +### L2 — real LLM + real remote DB, manual/pre-release, NOT automated + +**What:** end-to-end sessions with GLM 5.2 + the real Chirone datawarehouse and pgvector, reached via REST over VPN. Plus the gate-glue validation (the part L1 cannot reach). + +**Dependencies (all required, skip if missing):** +- LLM: Pi configured locally with GLM 5.2 (already ready). +- DB: the remote Supabase endpoints (see "L2 connection params" below), reachable via VPN. +- `harness/.env` populated with the API keys + CA path. + +**Modes:** +- **Human-in-the-loop (default):** the test runs the harness against GLM 5.2 and the reviewer answers each gate widget via the terminal, prompted by the test. Used for exploratory validation and for exercising the gate glue (which L1 cannot). +- **Pre-scripted answers (opt-in):** for specific cases (free-text "Altro", value grounding "ablazione", rollback), the test supplies canned reviewer answers automatically and asserts the outcome — no human typing. Used to test exact behaviors deterministically with the real model. + +**Coverage (honest):** validates the assumption L1 cannot — that GLM 5.2 produces tool calls the gate accepts, that the skill's prompts lead to the expected interaction shape, that the gate glue handles real tool-call sequences (incl. Altro/Rifiuta/rollback), that value grounding and formula approval surface correctly on the real schema, that `memory save-one` upserts to the real pgvector. Closes the skill→LLM→gate loop AND exercises the gate glue. + +**Honest limitation:** L2 is non-deterministic (the model may behave differently across runs) and slow/costly. It is a pre-release safety net, not a regression gate. A small, curated set of scenarios. + +### L2 connection params (from the operator) + +- **DWH (read-only):** `https://supabase-aritmolab.policlinicosandonato.it/dwh/` — PostgREST, schema `datawarehouse` + `datawarehouse_marts`, role `dwh_reader`. Header `X-API-Key: ${THOTH_DWH_API_KEY}`. +- **pgvector reader:** `https://supabase-aritmolab.policlinicosandonato.it/vector/v1/` — role `vector_reader` (`search_similar`, `list_tables`). Header `X-API-Key: ${THOTH_VEC_API_KEY}` (same value as `THOTH_DWH_API_KEY` — shared reader key). +- **pgvector writer (upsert-only):** same URL as reader, different key → role `vector_writer` (`existing_vector_hashes`, `upsert_vector_records`, no DELETE). Header `X-API-Key: ${THOTH_VEC_WRITE_API_KEY}`. +- **TLS:** self-signed, CA = leaf. On the operator's Mac, a local copy of the CA bundle must be present; `${THOTH_SSL_CA}` points to its local path. (Server path `/etc/nginx/ssl/policlinicosandonato.it.fullchain.crt` is unreachable from outside.) +- **Workspace:** `chirone-test`. +- **Test question (L2):** "dammi la lista dei pazienti che hanno fatto un'ablazione nel 2025" — exercises D14 value grounding + formula on the "ablazione" multi-column case. + +### API key handling (security) + +- All keys live **only** in `harness/.env` (gitignored). Never in code, never in the plan, never committed. +- `.env.example` is committed with variable names and empty values. +- L2 tests load `.env` via `python-dotenv` at startup; if any required var is missing/empty, the L2 test is **skipped with a clear message** (not failed), so L0/L1 never break for missing credentials. +- Tests never log key values. URLs in logs are fine; secrets are masked. +- ⚠️ The operator should rotate the key that appeared in chat transcripts. + +### Test marker convention + +- **L0 tests:** marker `@pytest.mark.l0`. Run always locally (Docker present). Auto-skip if Docker unavailable (rare). File naming: `tests/l0/test_*.py`. +- **L1 tests:** no marker (default, run always). File naming: `test_*.py` under `tests/`. Gate builder tests live in JS (node:test/vitest) under `harness/.pi/extensions/gate/__tests__/` and run via `npm test` alongside pytest. +- **L2 tests:** marker `@pytest.mark.l2`. Skipped automatically when `.env` incomplete. File naming: `tests/l2/test_*.py`. Default run: `pytest -m "not l2"` (L0+L1 only); pre-release: `pytest -m l2` (L2 only). + +--- + +## File Structure + +``` +harness/ +├── .pi/ ← Pi project config (ported + adapted) +│ ├── settings.json +│ ├── prompts/ ← /nuova-domanda, /riprendi-sessione +│ ├── skills/nsp-sessione/ ← SKILL.md + sub-files (adapted: task-doc, D13/D14) +│ └── extensions/ +│ ├── nsp-gate.js ← REWRITTEN: 6-widget descriptor emit/consume (glue, verified at L2) +│ ├── reserved-labels.mjs ← ported +│ └── gate/ +│ ├── builders.js ← NEW: pure widget-builder functions (L1, JS-tested) +│ └── __tests__/ ← node:test golden + fuzzy (L1) +│ ├── builders.test.js +│ └── golden/ +├── nsp/ ← Python package (ported + adapted) +│ ├── __init__.py +│ ├── config.py ← ported + vector_write_rest (D11) +│ ├── workspace.py ← NEW: YAML load + ${VAR} expand (D3 confine) +│ ├── workflow.py ← NEW: reads workflow.yaml (F2) +│ ├── phase.py ← REWRITTEN: data-driven + effective_decisions (D15) +│ ├── teardown.py ← NEW: teardown_to_phase (D15) +│ ├── taskdoc.py ← NEW: per-step task document generator (D16) +│ ├── decisions.py ← ported + decision_retracted (D15) +│ ├── cli/ +│ │ ├── __init__.py ← Typer app `nsp` +│ │ ├── _guards.py ← ported (require_vector_write_allowed etc.) +│ │ ├── workspace_cmd.py ← NEW: nsp workspace list/show +│ │ ├── phase_cmd.py ← adapted: + phase meta --json (F2) +│ │ ├── decision_cmd.py ← adapted: + retraction (D15) +│ │ ├── memory_cmd.py ← adapted: + save-one (D11) +│ │ ├── search_cmd.py ← adapted: + --kind formula (D14b) +│ │ ├── session_cmd.py ← adapted: + consistency check (D15) +│ │ └── (sql_cmd, cte_cmd, db_cmd, lsh_cmd, vector_cmd, evidence_cmd, schema_cmd — ported) +│ ├── db/ ← ported (connection, introspect, sampling, execute) +│ ├── rest/ ← ported (client, execute) +│ ├── search/ ← ported (combined_search, RRF) + formula retrieval (D14b) +│ ├── mschema/ ← ported (models, render, eligibility) +│ ├── vectorstore/ ← ported + dual-key rest_client (D11) +│ ├── evidence/ ← ported + formula kind (D14b) +│ ├── memory.py ← ported +│ └── session/ ← ported (models, store, artifacts) +├── workflow.yaml ← NEW: single source of workflow truth (F2) +├── workspaces/ ← workspace YAML definitions (D3) +│ └── chirone.example.yaml +├── sessions/ ← session persistence (FS, append-only ledger) +├── tests/ +│ ├── conftest.py ← L2 skip-when-no-.env fixtures +│ ├── golden/ ← golden JSON for widget builders (L1) +│ ├── fixtures/ +│ │ └── scenarios/ ← static .jsonl fixtures (synthesized once by build-scenarios) +│ ├── test_workflow.py ← L1 +│ ├── test_phase_effective.py ← L1 +│ ├── test_teardown.py ← L1 +│ ├── test_taskdoc.py ← L1 +│ ├── test_workspace.py ← L1 +│ ├── test_vector_dual_key.py ← L1 +│ ├── test_memory_save_one.py ← L1 (mocked writer) +│ ├── test_freetext_interpretation.py ← L1 (rationale contract) +│ ├── test_db_connection.py ← L0 Livello B (testcontainers: read-only enforcement) +│ ├── test_rest_client.py ← L1 Livello B (mock transport: RPC call) +│ ├── test_mschema_render.py ← L1 Livello B (3 formats on fake mschema) +│ ├── test_mschema_eligibility.py ← L1 Livello B (eligibility rules on fake columns) +│ ├── test_db_introspect.py ← L0 Livello B (testcontainers: known schema) +│ ├── test_db_sampling.py ← L0 Livello B (testcontainers: known data) +│ ├── test_rrf.py ← L0 Livello B (testcontainers: RRF fusion with real data) +│ ├── test_cli_contract.py ← L1 Livello B (each CLI: valid input → well-formed JSON) +│ ├── l0/ ← testcontainers tests (need Docker) +│ │ └── __init__.py +│ └── l2/ ← L2 (real GLM 5.2 + remote DB, @pytest.mark.l2) +│ └── __init__.py +│ └── l2/ +│ ├── test_session_ablazione.py ← L2 (GLM 5.2 + real DWH, human or scripted) +│ ├── test_value_grounding_real.py ← L2 ("ablazione" on real schema) +│ └── test_memory_save_one_real.py ← L2 (real pgvector writer upsert) +├── scripts/ +│ ├── create_vector_writer_rpc.sql ← ported (writer RPC allowlist) +│ └── create_vector_reader_rpc.sql ← NEW (reader RPC, D11 §5.4 note) +├── pyproject.toml +└── .env.example +``` + +--- + +## Phase A — Foundation (cross-cutting infrastructure first) + +Rationale: every feature depends on these. `workflow.yaml` + `effective_decisions` + `teardown` + `taskdoc` are the substrate the gate and the D11/D13/D14 features build on. + +### Task A1: Scaffold the harness project + +**Files:** +- Create: `harness/pyproject.toml` +- Create: `harness/.env.example` +- Create: `harness/nsp/__init__.py` +- Create: `harness/.gitignore` + +- [ ] **Step 1: Create `pyproject.toml`** + +```toml +[project] +name = "nsp" +version = "0.1.0" +description = "ThothII harness — deterministic CLI for the NL→SQL workflow" +requires-python = ">=3.12" +dependencies = [ + "typer>=0.12", + "pydantic>=2.0", + "pyyaml>=6.0", + "sqlalchemy>=2.0", + "psycopg2-binary>=2.9", + "datasketch>=1.6", + "sqlglot>=23.0", + "python-dotenv>=1.0", + "rich>=13.0", + "requests>=2.31", + "tqdm>=4.66", +] + +[project.scripts] +nsp = "nsp.cli:app" + +[project.optional-dependencies] +dev = [ + "pytest>=8.0", + "testcontainers[postgres]>=4.0", + "ruff>=0.5", +] + +[tool.ruff] +line-length = 100 + +[tool.pytest.ini_options] +testpaths = ["tests"] +``` + +- [ ] **Step 2: Create `.env.example`** (port the structure from `ChironeWp3/.env.example`, renamed to `THOTH_*` per §5.4) + +```bash +# ThothII profile: server (full rebuild) | workstation (read REST + optional upsert) +THOTH_PROFILE=server + +# Workspace DB (relational, direct transport) — example: chirone +THOTH_DB_HOST= +THOTH_DB_PORT=5432 +THOTH_DB_USER= +THOTH_DB_PASSWORD= + +# Vector REST — LETTURA (rpc search_similar) sul Supabase remoto +THOTH_VEC_REST_URL=https://host/vector/v1/ +THOTH_VEC_API_KEY= + +# Vector REST — SCRITTURA controllata (upsert only) da client remoti autorizzati +THOTH_VEC_WRITE_API_KEY= + +# Vector direct loading (server-only) +THOTH_VEC_HOST=localhost +THOTH_VEC_PORT=5438 +THOTH_VEC_USER=postgres +THOTH_VEC_PASSWORD= + +# Embeddings (Ollama) +THOTH_OLLAMA_URL=http://localhost:11434 + +# TLS (optional, internal CA) +THOTH_SSL_CA= +``` + +- [ ] **Step 3: Create `.gitignore`** + +``` +__pycache__/ +*.pyc +.env +.venv/ +sessions/ +indexes/ +*.egg-info/ +dist/ +build/ +``` + +- [ ] **Step 4: Create empty `nsp/__init__.py` and install** + +```bash +cd /Users/mp/projects/ThothII/harness +python -m venv .venv && source .venv/bin/activate +pip install -e ".[dev]" +``` + +- [ ] **Step 5: Verify `nsp` command is not yet wired (expected: import error) then commit** + +```bash +nsp --help 2>&1 | head -3 +``` +Expected: error (no `cli` module yet) — this is fine, we add it in Task A4. + +```bash +git add harness/ +git commit -m "feat(harness): scaffold project (pyproject, env, gitignore)" +``` + +--- + +### Task A2: Port config + create `workspace.py` (D3 confine) + +**Files:** +- Create: `harness/nsp/config.py` (ported + adapted from `ChironeWp3/src/psdwp3/config.py`) +- Create: `harness/nsp/workspace.py` +- Create: `harness/workspaces/chirone.example.yaml` +- Test: `harness/tests/test_workspace.py` + +- [ ] **Step 1: Port `config.py`** — copy `ChironeWp3/src/psdwp3/config.py` → `harness/nsp/config.py`. Rename the package prefix `psdwp3` → `nsp` everywhere. Verify it still defines: `DatabaseConfig`, `RestConfig`, `PathsConfig`, `EligibilityConfig`, `LshConfig`, `EvidenceSourcesConfig`, `EmbeddingsConfig`, `VectorConfig`, `SearchConfig`, `ExecutionConfig`, and the `Config` class with `vector_rest` + `vector_write_rest` (both `RestConfig | None`, see ChironeWp3 config.py:154-159) + `_expand_env` (lines 16-32). + +- [ ] **Step 2: Create `workspaces/chirone.example.yaml`** per spec §5.1 (relational + vector_db with rest/write_rest + evidence + embeddings + execution). + +- [ ] **Step 3: Write failing test for workspace loading** + +```python +# tests/test_workspace.py +from pathlib import Path +from nsp.workspace import load_workspace, WorkspaceError + +def test_load_workspace_expands_env_vars(monkeypatch, tmp_path): + monkeypatch.setenv("THOTH_VEC_API_KEY", "secret-reader") + monkeypatch.setenv("THOTH_VEC_WRITE_API_KEY", "secret-writer") + monkeypatch.setenv("THOTH_VEC_REST_URL", "https://example/vector/v1/") + yaml = tmp_path / "w.yaml" + yaml.write_text( + "name: test\n" + "relational:\n db_type: postgres\n transport: rest\n" + " rest: { base_url: 'https://dwh/', api_key: '${THOTH_VEC_API_KEY}' }\n" + "vector_db:\n collection: test_docs\n dim: 768\n" + " rest: { base_url: '${THOTH_VEC_REST_URL}', api_key: '${THOTH_VEC_API_KEY}' }\n" + " write_rest: { base_url: '${THOTH_VEC_REST_URL}', api_key: '${THOTH_VEC_WRITE_API_KEY}' }\n" + "embeddings:\n provider: ollama\n base_url: 'http://localhost:11434'\n model: m\n dim: 768\n" + ) + ws = load_workspace(yaml) + assert ws.vector_db.rest.api_key == "secret-reader" + assert ws.vector_db.write_rest.api_key == "secret-writer" + +def test_load_workspace_missing_env_raises(monkeypatch, tmp_path): + monkeypatch.delenv("THOTH_VEC_API_KEY", raising=False) + yaml = tmp_path / "w.yaml" + yaml.write_text( + "name: test\nrelational: { db_type: postgres, transport: rest, rest: { base_url: 'https://dwh/', api_key: '${THOTH_VEC_API_KEY}' } }\n" + "vector_db: { collection: c, dim: 768, rest: { base_url: 'https://v/', api_key: '${THOTH_VEC_API_KEY}' } }\n" + "embeddings: { provider: ollama, base_url: 'http://x', model: m, dim: 768 }\n" + ) + try: + load_workspace(yaml) + assert False, "should have raised" + except WorkspaceError as e: + assert "THOTH_VEC_API_KEY" in str(e) +``` + +- [ ] **Step 4: Run test, verify fail** + +Run: `cd harness && pytest tests/test_workspace.py -v` +Expected: FAIL — `ModuleNotFoundError: nsp.workspace` + +- [ ] **Step 5: Implement `workspace.py`** + +```python +# nsp/workspace.py +"""Workspace YAML loading — the single boundary for workspace configuration (spec D3). +Reads workspaces/.yaml, expands ${VAR} from env, validates via the Config model. +Future migration to a DB store would replace only this module. +""" +from __future__ import annotations +import os +from pathlib import Path +import yaml +from nsp.config import Config, ConfigError + +class WorkspaceError(Exception): + pass + +def _expand_str(s: str) -> str: + """Expand ${VAR} occurrences in s. Raises if a referenced var is unset.""" + PREFIX = "${" + SUFFIX = "}" + out: list[str] = [] + i = 0 + while i < len(s): + start = s.find(PREFIX, i) + if start == -1: + out.append(s[i:]) + break + out.append(s[i:start]) + end = s.find(SUFFIX, start + len(PREFIX)) + if end == -1: + raise WorkspaceError(f'Sintassi non valida (manca "}}"): {s[start:]}') + var = s[start + len(PREFIX) : end] + if var not in os.environ: + raise WorkspaceError(f"Variabile d'ambiente non definita: {var}") + out.append(os.environ[var]) + i = end + len(SUFFIX) + return "".join(out) + +def _expand_env(obj): + if isinstance(obj, str): + return _expand_str(obj) + if isinstance(obj, dict): + return {k: _expand_env(v) for k, v in obj.items()} + if isinstance(obj, list): + return [_expand_env(v) for v in obj] + return obj + +def load_workspace(path: str | Path) -> Config: + raw = yaml.safe_load(Path(path).read_text()) + try: + return Config.model_validate(_expand_env(raw)) + except ConfigError as e: + raise WorkspaceError(str(e)) from e +``` + +- [ ] **Step 6: Run test, verify pass** + +Run: `pytest tests/test_workspace.py -v` +Expected: 2 PASS + +- [ ] **Step 7: Commit** + +```bash +git add harness/nsp/config.py harness/nsp/workspace.py harness/workspaces/chirone.example.yaml harness/tests/test_workspace.py +git commit -m "feat(harness): port config + workspace.py YAML loader (D3)" +``` + +--- + +### Task A3: Create `workflow.yaml` + `workflow.py` (F2) — single source of truth + +**Files:** +- Create: `harness/workflow.yaml` +- Create: `harness/nsp/workflow.py` +- Test: `harness/tests/test_workflow.py` + +- [ ] **Step 1: Create `workflow.yaml`** (the 8-phase definition from spec §5.3) + +```yaml +# harness/workflow.yaml — single source of workflow truth (spec F2) +schema_version: 1 + +phases: + - id: F1 + name: chiaramento + advance: kind:phase + prerequisites: [] + artifacts_out: [] + - id: F2 + name: memoria + advance: auto_if_empty + prerequisites: [] + artifacts_out: [] + - id: F3 + name: riscrittura + advance: kind:phase + prerequisites: + - decision_exists: question_rewritten + artifacts_out: [question.md] + - id: F4 + name: schema_linking + advance: reviewer_decide + prerequisites: [] + artifacts_out: [schema_linking.json] + - id: F5 + name: sintesi + advance: kind:phase + prerequisites: + - file_validates: [schema_linking.json, SchemaLinking] + artifacts_out: [] + - id: F6 + name: cte + advance: auto_if_empty_or_skipped + prerequisites: + - any: + - decision_subject_exists: [phase_skipped, "phase:6"] + - all_ctes_approved: true + artifacts_out: [cte_plan.json, ctes/, cte_tests.json] + - id: F7 + name: sql_finale + advance: kind:phase + prerequisites: + - decision_exists: sql_approved + artifacts_out: [sql_final.sql] + - id: F8 + name: datamart + advance: reviewer_decide + prerequisites: + - any: + - decision_exists: datamart_requested + - decision_exists: datamart_declined + artifacts_out: [] + +decision_min_phase: auto +max_phase: auto +``` + +- [ ] **Step 2: Write failing test for workflow loading** + +```python +# tests/test_workflow.py +from nsp.workflow import load_workflow, PhaseSpec + +def test_workflow_loads_8_phases(): + wf = load_workflow() + assert len(wf.phases) == 8 + assert wf.phases[0].id == "F1" + assert wf.max_phase == 8 + assert wf.phase_by_num(1).name == "chiaramento" + +def test_decision_min_phase_derived(): + wf = load_workflow() + # question_rewritten is a prerequisite of F3 → min phase 3 + assert wf.decision_min_phase("question_rewritten") == 3 + # sql_approved is a prerequisite of F7 → min phase 7 + assert wf.decision_min_phase("sql_approved") == 7 + # unknown type → phase 1 (default) + assert wf.decision_min_phase("nonexistent_type") == 1 + +def test_phase_name_lookup(): + wf = load_workflow() + assert wf.phase_name(6) == "cte" + assert wf.phase_name(8) == "datamart" # the JS drift bug — F8 must be present + +def test_artifacts_out_per_phase(): + wf = load_workflow() + assert "schema_linking.json" in wf.phase_by_num(4).artifacts_out + assert "sql_final.sql" in wf.phase_by_num(7).artifacts_out +``` + +- [ ] **Step 3: Run test, verify fail** + +Run: `pytest tests/test_workflow.py -v` +Expected: FAIL — `ModuleNotFoundError: nsp.workflow` + +- [ ] **Step 4: Implement `workflow.py`** + +```python +# nsp/workflow.py +"""Reads workflow.yaml — the SINGLE source of workflow truth (spec F2, §5.3). +phase.py, the gate, and the skill all read from here. No more duplicated constants. +""" +from __future__ import annotations +from dataclasses import dataclass, field +from pathlib import Path +import yaml + +_WF_PATH = Path(__file__).resolve().parent.parent / "workflow.yaml" + +@dataclass +class PhaseSpec: + id: str + num: int + name: str + advance: str + prerequisites: list + artifacts_out: list[str] = field(default_factory=list) + +@dataclass +class Workflow: + schema_version: int + phases: list[PhaseSpec] + _decision_min_map: dict[str, int] = field(default_factory=dict) + + @property + def max_phase(self) -> int: + return len(self.phases) + + def phase_by_num(self, n: int) -> PhaseSpec: + return self.phases[n - 1] + + def phase_name(self, n: int) -> str: + return self.phase_by_num(n).name if 1 <= n <= self.max_phase else "?" + + def decision_min_phase(self, decision_type: str) -> int: + # A decision type's min phase = the earliest phase whose prerequisites + # reference it (via decision_exists/decision_subject_exists), else 1. + return self._decision_min_map.get(decision_type, 1) + +def _collect_decision_mins(phases: list[PhaseSpec]) -> dict[str, int]: + """Scan prerequisites for decision_exists / decision_subject_exists mentions.""" + mins: dict[str, int] = {} + def scan(node, phase_num: int): + if isinstance(node, dict): + for k, v in node.items(): + if k in ("decision_exists", "decision_subject_exists"): + if isinstance(v, list): + dtype = v[0] + else: + dtype = v + if dtype not in mins or phase_num < mins[dtype]: + mins[dtype] = phase_num + else: + scan(v, phase_num) + elif isinstance(node, list): + for item in node: + scan(item, phase_num) + for p in phases: + scan(p.prerequisites, p.num) + return mins + +def load_workflow(path: Path | str = _WF_PATH) -> Workflow: + raw = yaml.safe_load(Path(path).read_text()) + phases = [] + for i, p in enumerate(raw["phases"], start=1): + phases.append(PhaseSpec( + id=p["id"], num=i, name=p["name"], + advance=p["advance"], prerequisites=p.get("prerequisites", []), + artifacts_out=p.get("artifacts_out", []), + )) + return Workflow( + schema_version=raw.get("schema_version", 1), + phases=phases, + _decision_min_map=_collect_decision_mins(phases), + ) +``` + +- [ ] **Step 5: Run test, verify pass** + +Run: `pytest tests/test_workflow.py -v` +Expected: 4 PASS + +- [ ] **Step 6: Commit** + +```bash +git add harness/workflow.yaml harness/nsp/workflow.py harness/tests/test_workflow.py +git commit -m "feat(harness): workflow.yaml as single source of truth + workflow.py loader (F2)" +``` + +--- + +### Task A4: Port `decisions.py` + add `decision_retracted` (D15) + +**Files:** +- Create: `harness/nsp/decisions.py` (ported + adapted) +- Test: `harness/tests/test_decisions_retract.py` + +- [ ] **Step 1: Port `decisions.py`** — copy `ChironeWp3/src/psdwp3/session/decisions.py`. Rename package. Add `decision_retracted` to the `DecisionType` Literal. The `DecisionRecord` model gains an optional `retracts: int | None` field (the `decision_seq` being retracted, for step-level rollback). + +- [ ] **Step 2: Write failing test for retraction** + +```python +# tests/test_decisions_retract.py +from pathlib import Path +from nsp.decisions import append_decision, list_decisions, DecisionRecord + +def test_retracted_decision_in_audit_but_marked(tmp_path): + session = tmp_path / "s1" + session.mkdir() + append_decision(session, DecisionRecord(seq=1, ts="t", type="table_promoted", + subject="phase:4", detail="t1", rationale="r")) + append_decision(session, DecisionRecord(seq=2, ts="t", type="decision_retracted", + subject="phase:4", detail="retract", rationale="wrong", retracts=1)) + all_decisions = list_decisions(session) + assert len(all_decisions) == 2 # both in audit + assert all_decisions[1].retracts == 1 +``` + +- [ ] **Step 3: Run, verify fail, implement the `retracts` field + ensure `append_decision` writes it, verify pass.** (The ported `append_decision` already writes all fields; just ensure the new field is serialized. If using pydantic `model_dump`, add `retracts: int | None = None`.) + +Run: `pytest tests/test_decisions_retract.py -v` +Expected after implement: PASS + +- [ ] **Step 4: Commit** + +```bash +git add harness/nsp/decisions.py harness/tests/test_decisions_retract.py +git commit -m "feat(harness): port decisions.py + decision_retracted for step rollback (D15)" +``` + +--- + +### Task A5: Rewrite `phase.py` — data-driven + `effective_decisions` (D15 core fix) + +**Files:** +- Create: `harness/nsp/phase.py` (rewritten) +- Port: `harness/nsp/session/{models,store,artifacts}.py` (needed for SchemaLinking validation) +- Test: `harness/tests/test_phase_effective.py` + +This is the single most important architectural fix (spec §4.8). `phase.py` reads from `workflow.yaml`, and ALL helpers consult `effective_decisions()` instead of raw `list_decisions()`. + +- [ ] **Step 1: Port session models/store/artifacts** — copy `ChironeWp3/src/psdwp3/session/{models.py, store.py, artifacts.py}` → `harness/nsp/session/`. Rename package. These are needed for `SchemaLinking` validation in `advance_problems`. + +- [ ] **Step 2: Write failing test for effective_decisions** + +```python +# tests/test_phase_effective.py +from pathlib import Path +from nsp.phase import current_phase, effective_decisions +from nsp.decisions import append_decision, DecisionRecord + +def _d(session, seq, dtype, subject, **kw): + append_decision(session, DecisionRecord(seq=seq, ts="t", type=dtype, subject=subject, + detail=kw.get("detail", ""), rationale=kw.get("rationale", ""), + retracts=kw.get("retracts"))) + +def test_effective_decisions_excludes_pre_reopen_tail(tmp_path): + s = tmp_path / "s"; s.mkdir() + _d(s, 1, "phase_approved", "phase:1") + _d(s, 2, "phase_approved", "phase:2") + _d(s, 3, "phase_reopened", "phase:1") # rollback to phase 1 + _d(s, 4, "table_promoted", "phase:4") # stale: produced after reopen but for phase 4 (not current) + # After reopen to phase:1, the effective view truncates everything after the last phase_reopened. + eff = effective_decisions(s) + # The reopen at seq 3 means decisions after it (seq 4) are NOT effective, + # and the reopen itself sets the pointer. Effective = decisions up to and incl. reopen, + # then re-walked. The stale table_promoted at seq 4 must be excluded. + types = [d.type for d in eff] + assert "table_promoted" not in types + +def test_current_phase_after_reopen(tmp_path): + s = tmp_path / "s"; s.mkdir() + _d(s, 1, "phase_approved", "phase:1") + _d(s, 2, "phase_approved", "phase:2") + _d(s, 3, "phase_reopened", "phase:1") + assert current_phase(s) == 1 + _d(s, 4, "phase_approved", "phase:1") # re-approve after reopen + assert current_phase(s) == 2 + +def test_retracted_decision_excluded_from_effective(tmp_path): + s = tmp_path / "s"; s.mkdir() + _d(s, 1, "table_promoted", "phase:4", detail="t1") + _d(s, 2, "decision_retracted", "phase:4", retracts=1) + eff = effective_decisions(s) + types = [d.type for d in eff] + assert "table_promoted" not in types # retracted → excluded +``` + +- [ ] **Step 3: Run, verify fail** + +Run: `pytest tests/test_phase_effective.py -v` +Expected: FAIL + +- [ ] **Step 4: Implement `phase.py`** — the core. Port the fold logic from `ChironeWp3/src/psdwp3/session/phase.py` but: + - Read `max_phase`, `phase_name`, `decision_min_phase`, `advance_problems` rules from `workflow.yaml` via `load_workflow()`. + - Add `effective_decisions(session)`: replay the ledger; when hitting a `phase_reopened phase:N`, truncate everything after it in the effective view AND re-walk from phase N. When hitting a `decision_retracted`, exclude the retracted seq. + - Rewrite `advance_problems(phase)`, `approved_ctes`, `substantive_count_current_phase`, `auto_advance_eligible` to consult `effective_decisions()` instead of `list_decisions()`. + +Reference skeleton (the `effective_decisions` core — the load-bearing part): + +```python +# nsp/phase.py (key function) +from nsp.decisions import list_decisions +from nsp.workflow import load_workflow + +def effective_decisions(session_dir) -> list: + """The canonical effective view. ALL helpers MUST use this, not list_decisions(). + Replays the ledger: phase_reopened truncates the tail; decision_retracted excludes the target. + """ + all_d = list_decisions(session_dir) + retracted_seqs = {d.retracts for d in all_d if d.type == "decision_retracted" and d.retracts} + # Find the LAST phase_reopened; everything strictly after it is stale. + last_reopen_idx = -1 + for i, d in enumerate(all_d): + if d.type == "phase_reopened": + last_reopen_idx = i + if last_reopen_idx >= 0: + effective = all_d[: last_reopen_idx + 1] + else: + effective = list(all_d) + # Exclude retracted decisions (by seq), and the retraction records themselves. + return [d for d in effective if d.seq not in retracted_seqs and d.type != "decision_retracted"] +``` + +Implement `current_phase(session_dir)` as the fold over `effective_decisions(session_dir)` (not over raw list). Implement `advance_problems(phase, session_dir)` by evaluating the `prerequisites` list from `workflow.yaml` for that phase against `effective_decisions`. Reuse the prerequisite predicate types: `decision_exists`, `decision_subject_exists`, `file_validates`, `all_ctes_approved`, `any`, `all`. + +- [ ] **Step 5: Run, verify pass** + +Run: `pytest tests/test_phase_effective.py -v` +Expected: 3 PASS + +- [ ] **Step 6: Commit** + +```bash +git add harness/nsp/phase.py harness/nsp/session/ harness/tests/test_phase_effective.py +git commit -m "feat(harness): rewrite phase.py data-driven + effective_decisions (D15 core, F2)" +``` + +--- + +### Task A6: Create `teardown.py` — `teardown_to_phase` (D15 teardown) + +**Files:** +- Create: `harness/nsp/teardown.py` +- Test: `harness/tests/test_teardown.py` + +- [ ] **Step 1: Write failing test** + +```python +# tests/test_teardown.py +from pathlib import Path +from nsp.teardown import teardown_to_phase, TeardownReport + +def test_teardown_to_phase_4_deletes_phase5plus_artifacts(tmp_path): + s = tmp_path / "sess"; s.mkdir() + (s / "schema_linking.json").write_text("{}") # F4 artifact + (s / "cte_plan.json").write_text("[]") # F6 artifact + (s / "ctes").mkdir() + (s / "ctes" / "x.sql").write_text("SELECT 1") # F6 artifact + (s / "sql_final.sql").write_text("SELECT 1") # F7 artifact + report = teardown_to_phase(s, target_phase=4) + assert (s / "schema_linking.json").exists() # F4 preserved (target is 4) + assert not (s / "cte_plan.json").exists() # F6 deleted (>4) + assert not (s / "ctes").exists() # F6 dir deleted + assert not (s / "sql_final.sql").exists() # F7 deleted + assert "cte_plan.json" in report.deleted_files + assert "sql_final.sql" in report.deleted_files + +def test_teardown_records_deleted_in_ledger(tmp_path): + # teardown itself appends a phase_reopened decision (the rollback marker) + # — verifying that is covered by test_phase_effective; here we check the report. + s = tmp_path / "sess"; s.mkdir() + (s / "sql_final.sql").write_text("SELECT 1") + report = teardown_to_phase(s, target_phase=4) + assert len(report.deleted_files) >= 1 +``` + +- [ ] **Step 2: Run, verify fail** + +Run: `pytest tests/test_teardown.py -v` + +- [ ] **Step 3: Implement `teardown.py`** + +```python +# nsp/teardown.py +"""Artifact teardown on rollback (spec D15, §4.8). +Deletes every artifact whose producing phase > target, using the artifacts_out map +from workflow.yaml. Recomputes derived state. Fixes the orphaned-CTE-blocks-finalize bug. +""" +from __future__ import annotations +from dataclasses import dataclass, field +from pathlib import Path +from nsp.workflow import load_workflow + +@dataclass +class TeardownReport: + target_phase: int + deleted_files: list[str] = field(default_factory=list) + +def teardown_to_phase(session_dir: str | Path, target_phase: int) -> TeardownReport: + session_dir = Path(session_dir) + wf = load_workflow() + report = TeardownReport(target_phase=target_phase) + for phase in wf.phases: + if phase.num <= target_phase: + continue + for artifact in phase.artifacts_out: + target = session_dir / artifact.rstrip("/") + if artifact.endswith("/"): + # directory artifact (e.g. ctes/) + if target.exists(): + for f in target.glob("*"): + f.unlink() + report.deleted_files.append(f.name) + target.rmdir() + else: + if target.exists(): + target.unlink() + report.deleted_files.append(artifact) + return report +``` + +- [ ] **Step 4: Run, verify pass** + +Run: `pytest tests/test_teardown.py -v` +Expected: 2 PASS + +- [ ] **Step 5: Commit** + +```bash +git add harness/nsp/teardown.py harness/tests/test_teardown.py +git commit -m "feat(harness): teardown_to_phase — artifact teardown on rollback (D15)" +``` + +--- + +### Task A7: Create `taskdoc.py` — per-step task document generator (D16) + +**Files:** +- Create: `harness/nsp/taskdoc.py` +- Test: `harness/tests/test_taskdoc.py` + +- [ ] **Step 1: Write failing test** + +```python +# tests/test_taskdoc.py +from pathlib import Path +from nsp.taskdoc import generate_task_doc, TaskDoc + +def test_task_doc_includes_question_and_schema_scope_not_full_physical(tmp_path): + s = tmp_path / "sess"; s.mkdir() + (s / "question.md").write_text("# Domanda\nQuanti pazienti?\n## Assunzioni\n- a") + (s / "schema_linking.json").write_text('{"candidates":[{"kind":"table","name":"pazienti"}],"joins":[]}') + # A fatal-sized artifact that must NOT appear in the task doc + (s.parent / "physical.yaml").write_text("x: " + "y" * 800_000) + doc = generate_task_doc(session_dir=s, phase=7, promoted_tables=["pazienti"]) + assert "Quanti pazienti?" in doc.body + assert "pazienti" in doc.body + assert len(doc.body) < 100_000 # bounded — no fatal full-schema read + assert "physical.yaml" not in doc.body # never embedded + +def test_task_doc_byte_budget_enforced(tmp_path): + s = tmp_path / "sess"; s.mkdir() + (s / "question.md").write_text("q") + doc = generate_task_doc(session_dir=s, phase=1, promoted_tables=[]) + assert doc.byte_budget_ok is True +``` + +- [ ] **Step 2: Run, verify fail** + +Run: `pytest tests/test_taskdoc.py -v` + +- [ ] **Step 3: Implement `taskdoc.py`** + +```python +# nsp/taskdoc.py +"""Per-step task document generator (spec D16, §4.9). +Emits a single compact document per phase/step, derived from prior artifacts, +with enforced byte budget (target <20k tokens). NEVER embeds full physical.yaml. +""" +from __future__ import annotations +from dataclasses import dataclass +from pathlib import Path + +MAX_BODY_BYTES = 80_000 # ~20k tokens + +@dataclass +class TaskDoc: + phase: int + body: str + byte_budget_ok: bool + +def generate_task_doc(session_dir: Path | str, phase: int, promoted_tables: list[str] | None = None) -> TaskDoc: + session_dir = Path(session_dir) + parts: list[str] = [] + q = session_dir / "question.md" + if q.exists(): + parts.append("## Domanda\n" + q.read_text()) + sl = session_dir / "schema_linking.json" + if sl.exists() and phase >= 4: + parts.append("## Schema linking (deciso)\n```json\n" + sl.read_text() + "\n```") + # Task header for the phase + parts.append(f"## Task: fase {phase}") + body = "\n\n".join(parts) + return TaskDoc(phase=phase, body=body, byte_budget_ok=len(body.encode()) <= MAX_BODY_BYTES) +``` + +- [ ] **Step 4: Run, verify pass** + +Run: `pytest tests/test_taskdoc.py -v` +Expected: 2 PASS + +- [ ] **Step 5: Commit** + +```bash +git add harness/nsp/taskdoc.py harness/tests/test_taskdoc.py +git commit -m "feat(harness): taskdoc per-step generator with byte budget (D16)" +``` + +--- + +### Task A8: Port the CLI skeleton + `nsp phase meta --json` (F2) + +**Files:** +- Create: `harness/nsp/cli/__init__.py` (Typer app) +- Create: `harness/nsp/cli/phase_cmd.py` (adapted) +- Create: `harness/nsp/cli/workspace_cmd.py` +- Port the command modules (sql, cte, db, lsh, vector, evidence, schema, memory, search, session, decision) — adapt imports +- Test: `harness/tests/test_cli_phase_meta.py` + +**Honest scope:** this task ports the CLI structure + `phase meta` and verifies that ONE command behaves. It does NOT verify the 11 ported command modules — those get contract tests in **Task A9 (Livello B)**. Porting them here without tests would assume them reliable, which the spec forbids. + +- [ ] **Step 1: Port CLI modules** — copy `ChironeWp3/src/psdwp3/cli/*.py` → `harness/nsp/cli/`. Rename package imports. The gate will need `nsp phase meta --json` to avoid mirroring constants. + +- [ ] **Step 2: Write failing test for `nsp phase meta --json`** + +```python +# tests/test_cli_phase_meta.py +import json +from typer.testing import CliRunner +from nsp.cli import app + +runner = CliRunner() + +def test_phase_meta_json_returns_workflow_data(): + result = runner.invoke(app, ["phase", "meta", "--json"]) + assert result.exit_code == 0 + data = json.loads(result.stdout) + assert data["max_phase"] == 8 + assert len(data["phases"]) == 8 + assert data["phases"][7]["name"] == "datamart" # F8 present (fixes the JS drift) + assert "advance" in data["phases"][0] +``` + +- [ ] **Step 3: Run, verify fail, then implement `phase meta` in `phase_cmd.py`** + +The command reads from `load_workflow()` and dumps `{max_phase, phases: [{num,name,advance,artifacts_out}]}` as JSON. + +- [ ] **Step 4: Run, verify pass; commit** + +```bash +git add harness/nsp/cli/ harness/tests/test_cli_phase_meta.py +git commit -m "feat(harness): port CLI structure + nsp phase meta --json (F2, kills JS/Python drift)" +``` + +--- + +### Task A9: Contract tests for ported load-bearing modules — L0 (testcontainers) + L1 (fake data) + +**Rationale:** spec §1 forbids assuming ported code is reliable. The DB-touching modules get **L0 tests** (testcontainers — real Postgres in container, where "not assumed reliable" gains real teeth); the logic modules get **L1 tests** (fake data). Docker is available on the dev machine, so L0 runs locally on every `pytest`. + +**L0 tests (testcontainers, marker `@pytest.mark.l0`):** + +- [ ] **Step 1: `tests/l0/test_db_connection.py`** — read-only enforcement (exit code 2 if writable role), `ping` succeeds on read-only role. + +- [ ] **Step 2: `tests/l0/test_db_introspect.py` + `test_db_sampling.py`** — create a known schema + known rows in the container, verify `introspect` returns them with columns + comments, and `unique_values_for_lsh` returns the expected most-frequent values. + +- [ ] **Step 3: `tests/l0/test_rrf.py`** — load real-ish data into the container (LSH index built from sampled values + a small vector set), verify `combined_search`/`rrf_fuse` fuses them with stable, sensible ranking. This is the one that truly exercises the ported RRF/LSH pipeline. + +**L1 tests (fake data, no marker):** + +- [ ] **Step 4: `tests/test_rest_client.py`** — mock `requests`, verify one RPC call carries the `X-API-Key` header and parses the response. + +- [ ] **Step 5: `tests/test_mschema_render.py`** — the 3 formats (markdown, mschema-text ThothAI, schema-dict) from a fake `PhysicalSchema`. Catches port breaks in the render layer. + +- [ ] **Step 6: `tests/test_mschema_eligibility.py`** — eligibility rules on fake columns (wide-text excluded above threshold). + +- [ ] **Step 7: `tests/test_cli_contract.py`** — one test per ported CLI command (sql, cte, schema, session, decision, lsh, vector, evidence, memory, search, db): valid input → exit 0 + parseable JSON. + +- [ ] **Step 8: Register `l0` marker in `pyproject.toml`** and run all A9 tests. + +```toml +[tool.pytest.ini_options] +testpaths = ["tests"] +markers = [ + "l0: testcontainers tests (need Docker, run locally)", + "l2: end-to-end tests requiring real GLM 5.2 + remote DB (skipped when .env incomplete)", +] +addopts = "-m 'not l2'" # L0 runs by default (Docker present); L2 opt-in +``` + +```bash +cd harness && pytest tests/l0/ tests/test_rest_client.py tests/test_mschema_render.py \ + tests/test_mschema_eligibility.py tests/test_cli_contract.py -v +``` + +- [ ] **Step 9: Fix porting bugs the tests surface; commit** + +```bash +git add harness/tests/l0/ harness/tests/test_rest_client.py harness/tests/test_mschema_*.py \ + harness/tests/test_cli_contract.py harness/nsp/{db,rest,mschema,search}/ harness/pyproject.toml +git commit -m "test(harness): L0 testcontainers + L1 contract tests for ported modules (spec §1)" +``` + +--- + +## Phase B — Feature deviations (D11, D13, D14) + +### Task B1: Dual vector API key — port `vectorstore` + verify writer config (D11) + +**Files:** +- Port: `harness/nsp/vectorstore/{rest_client.py, rest_writer.py, store.py, reader.py, embeddings.py, records.py}` +- Port: `harness/nsp/cli/_guards.py` +- Port: `harness/scripts/create_vector_writer_rpc.sql` +- Create: `harness/scripts/create_vector_reader_rpc.sql` +- Test: `harness/tests/test_vector_dual_key.py` + +- [ ] **Step 1: Port the vectorstore modules** (rest_client, rest_writer, store, reader, embeddings, records) and `_guards.py` from ChironeWp3. These already implement the dual-key model per spec §5.4 — port them, rename package, verify `has_vector_write_rest` checks `.api_key.strip()`. + +- [ ] **Step 2: Write failing test for dual-key client construction** + +```python +# tests/test_vector_dual_key.py +from nsp.config import RestConfig +from nsp.vectorstore.rest_client import VectorRestClient + +def test_reader_and_writer_use_separate_keys(): + reader = VectorRestClient(RestConfig(base_url="https://v/", api_key="K-READ")) + writer = VectorRestClient(RestConfig(base_url="https://v/", api_key="K-WRITE")) + assert reader.api_key == "K-READ" + assert writer.api_key == "K-WRITE" + +def test_has_vector_write_rest_false_for_empty_key(): + from nsp.cli._guards import has_vector_write_rest + from nsp.config import Config, RestConfig + cfg = Config(vector_write_rest=RestConfig(base_url="x", api_key=" ")) + assert has_vector_write_rest(cfg) is False +``` + +- [ ] **Step 3: Run, verify fail, adjust ported code so `VectorRestClient` exposes `api_key` (it stores the config; add a property if missing), verify pass.** + +- [ ] **Step 4: Create `create_vector_reader_rpc.sql`** — author the reader RPC (`search_similar`, `list_tables`) mirroring the writer allowlist pattern, granted to a `vector_reader` role. (The reader RPCs lived server-side in Supabase and were never in the ChironeWp3 repo — spec §5.4 note. Author them now.) + +- [ ] **Step 5: Commit** + +```bash +git add harness/nsp/vectorstore/ harness/nsp/cli/_guards.py harness/scripts/ harness/tests/test_vector_dual_key.py +git commit -m "feat(harness): port vectorstore dual-key + reader RPC (D11, §5.4)" +``` + +--- + +### Task B2: `nsp memory save-one` — targeted upsert via writer key (D11) + +**Files:** +- Modify: `harness/nsp/cli/memory_cmd.py` (add `save-one`) +- Modify: `harness/nsp/memory.py` (add `memory_vector_record_for_decision`) +- Test: `harness/tests/test_memory_save_one.py` + +- [ ] **Step 1: Write failing test** + +```python +# tests/test_memory_save_one.py +from pathlib import Path +from unittest.mock import patch, MagicMock +from typer.testing import CliRunner +from nsp.cli import app + +runner = CliRunner() + +def test_save_one_calls_upsert_with_single_record(tmp_path, monkeypatch): + session = tmp_path / "s"; session.mkdir() + # build a minimal registry + a decision_seq to save + from nsp.decisions import append_decision, DecisionRecord + append_decision(session, DecisionRecord(seq=7, ts="t", type="table_promoted", + subject="phase:4", detail="pazienti", rationale="r")) + mock_writer = MagicMock() + mock_writer.upsert_records.return_value = 1 + with patch("nsp.cli.memory_cmd.open_store", return_value=mock_writer): + result = runner.invoke(app, ["memory", "save-one", "--session", str(session), "--decision-seq", "7"]) + assert result.exit_code == 0 + mock_writer.sync.assert_not_called() # must be a single upsert, NOT full resync + mock_writer.upsert_records.assert_called_once() + args = mock_writer.upsert_records.call_args + assert len(args[0][1]) == 1 # exactly one row + +def test_save_one_refused_without_writer_key(tmp_path, monkeypatch): + monkeypatch.setenv("THOTH_PROFILE", "workstation") + session = tmp_path / "s"; session.mkdir() + # config with no write_rest + result = runner.invoke(app, ["memory", "save-one", "--session", str(session), "--decision-seq", "1"]) + assert result.exit_code == 4 # require_vector_write_allowed gate +``` + +- [ ] **Step 2: Run, verify fail** + +- [ ] **Step 3: Implement `save-one`** in `memory_cmd.py`: + - Guard: `require_vector_write_allowed(cfg, "memory save-one")`. + - Load the registry, find the memory record for `decision_seq`. + - Build a single `VectorRecord` (reuse `memory_vector_records` but filter to one). + - Call `writer.existing_hashes(kinds={"memory"})` + embed (if hash changed) + `writer.upsert_records(table="memory", rows=[one])`. + - Do NOT call `writer.sync` (that's the full-resync path). + +- [ ] **Step 4: Run, verify pass; commit** + +```bash +git add harness/nsp/cli/memory_cmd.py harness/nsp/memory.py harness/tests/test_memory_save_one.py +git commit -m "feat(harness): nsp memory save-one — targeted upsert via writer key (D11)" +``` + +--- + +### Task B3: Value & Schema Linking — `value_grounded` decision + LSH multi-column (D14a) + +**Files:** +- Modify: `harness/nsp/decisions.py` (add `value_grounded`) +- Modify: `harness/nsp/search/__init__.py` (stop collapsing multi-column to single best) +- Modify: `harness/nsp/session/models.py` (`Candidate.grounded_values`) +- Test: `harness/tests/test_value_grounding.py` + +- [ ] **Step 1: Port `search/__init__.py`, `lshindex/`, `db/sampling.py`** from ChironeWp3. Then modify. + +- [ ] **Step 2: Write failing test for multi-column LSH exposure** + +```python +# tests/test_value_grounding.py +from nsp.search import aggregate_lsh_multi + +def test_value_in_multiple_columns_returns_all(): + hits = [ + {"table": "t", "column": "c1", "value": "ablazione", "score": 0.9}, + {"table": "t", "column": "c2", "value": "ablazione", "score": 0.7}, + ] + result = aggregate_lsh_multi(hits) + # NOT collapsed to single best — both columns exposed + cols = {h["column"] for h in result["t"]} + assert cols == {"c1", "c2"} +``` + +- [ ] **Step 3: Run, verify fail, then implement `aggregate_lsh_multi`** (replaces the collapsing `_aggregate_lsh` behavior — keep the old name as a thin wrapper to the new one if other code needs single-best). Add `value_grounded` to DecisionType. Add `grounded_values: list[dict] = []` to `Candidate` in `session/models.py`. + +- [ ] **Step 4: Run, verify pass; commit** + +```bash +git add harness/nsp/search/ harness/nsp/lshindex/ harness/nsp/db/sampling.py harness/nsp/decisions.py harness/nsp/session/models.py harness/tests/test_value_grounding.py +git commit -m "feat(harness): value grounding — multi-column LSH + value_grounded (D14a)" +``` + +--- + +### Task B4: SQL formula evidence — formula kind + retrieval + approval (D14b) + +**Files:** +- Modify: `harness/nsp/evidence/model.py` (activate `tier: concept` / add formula fields) +- Create: `harness/nsp/evidence/formula_store.py` +- Modify: `harness/nsp/cli/search_cmd.py` (`--kind formula`) +- Modify: `harness/nsp/decisions.py` (add `concept_formula_approved`, `concept_formula_rejected`) +- Test: `harness/tests/test_formula.py` + +- [ ] **Step 1: Port `evidence/` modules** from ChironeWp3. + +- [ ] **Step 2: Write failing test for formula retrieval** + +```python +# tests/test_formula.py +from pathlib import Path +from nsp.evidence.formula_store import ConceptFormula, save_formula, retrieve_formula + +def test_formula_retrieval_by_concept(tmp_path): + f = ConceptFormula(concept="fascia pediatrica", columns=["data_nascita"], + sql="CASE WHEN ... END", status="reviewed", sources=["s1"]) + save_formula(tmp_path, f) + results = retrieve_formula(tmp_path, "fascia pediatrica") + assert len(results) == 1 + assert results[0].sql.startswith("CASE WHEN") + assert results[0].concept == "fascia pediatrica" +``` + +- [ ] **Step 3: Run, verify fail, implement `formula_store.py`** (frontmatter YAML + SQL body, concept→formula units). Add `--kind formula` to `search_cmd` that calls `retrieve_formula`. Add the two new decision types. + +- [ ] **Step 4: Run, verify pass; commit** + +```bash +git add harness/nsp/evidence/ harness/nsp/cli/search_cmd.py harness/nsp/decisions.py harness/tests/test_formula.py +git commit -m "feat(harness): SQL formula evidence — concept→formula units + retrieval + approval decisions (D14b)" +``` + +--- + +### Task B5: Free-text interpretation guidance (D13) + +**Files:** +- Modify: `harness/nsp/cli/decision_cmd.py` (ensure free-text from "Altro"/"Rifiuta" is captured in `rationale`) +- Modify: `.pi/skills/nsp-sessione/SKILL.md` (+ sub-files) — port + add the D13 §4.6 guidance +- Test: `harness/tests/test_freetext_interpretation.py` + +- [ ] **Step 1: Write failing test** — a golden-style test that, given a `ui_response` with `control: freetext` (Altro text), verifies the harness records the user's text in the decision `rationale` (not discarded). + +```python +# tests/test_freetext_interpretation.py +from pathlib import Path +from nsp.decisions import append_decision, list_decisions, DecisionRecord + +def test_altro_freetext_recorded_in_rationale(tmp_path): + s = tmp_path / "sess"; s.mkdir() + # The gate, on receiving Altro text, appends a decision whose rationale carries the user's words. + append_decision(s, DecisionRecord(seq=1, ts="t", type="table_promoted", + subject="phase:4", detail="ablazione", + rationale="Utente (Altro): 'ablazione' va cercato anche in patologia, non solo nel flag")) + decisions = list_decisions(s) + assert "patologia" in decisions[0].rationale # user text preserved, not discarded +``` + +- [ ] **Step 2: Port `SKILL.md` + sub-files** from ChironeWp3. Add a section "Interpretazione del testo libero" per spec §4.6: instructs the model to (1) evaluate Altro/Rifiuta/steering text in context, (2) act on it, (3) if ambiguous, re-ask instead of defaulting, (4) record the text in the decision rationale. + +- [ ] **Step 3: Run, verify pass (the test asserts the recording contract; the actual model behavior is enforced by the skill prose + gate, validated by fake-pi golden tests in Task C3); commit** + +```bash +git add harness/.pi/skills/ harness/nsp/cli/decision_cmd.py harness/tests/test_freetext_interpretation.py +git commit -m "feat(harness): free-text interpretation guidance + rationale capture (D13)" +``` + +--- + +## Phase C — Gate → widget-descriptor (D2/D4) + +The gate has two parts with **very different testability**: +- **Builders** (pure functions that turn tool-call params into widget-descriptor JSON) — fully testable in L1, **in JS, in-language**. +- **Glue** (Pi runtime integration: `pi.registerTool`, `ctx.sendRaw`, `emitAndWait`, anti-bypass hooks, no-limbo loop) — **NOT testable in L1**; it depends on the Pi runtime. Tested at L2. + +This phase therefore has: **C1 (builders, L1, JS)** + **C2 (glue implementation, no L1 test — verified at L2)**. There is no fake-LLM driver and no `build-scenarios` in this plan (see cross-cutting follow-up). + +### Task C1: Pure widget-builder functions + JS tests (L1) + +**Files:** +- Create: `harness/.pi/extensions/gate/builders.js` (pure functions, no Pi context) +- Create: `harness/.pi/extensions/gate/__tests__/builders.test.js` (node:test, in-language) +- Create: `harness/.pi/extensions/gate/__tests__/golden/*.json` (golden widget descriptors) +- Create: `harness/package.json` (for `npm test` → `node --test`) + +This task extracts widget-descriptor **construction** into pure, testable functions and tests them in JS directly — no Python↔JS bridge, no Python mirror (a mirror would drift; the golden tests would validate the mirror, not the real gate). + +- [ ] **Step 1: Create `package.json`** for the gate JS tests. + +```json +{ + "name": "thothii-harness-gate", + "private": true, + "scripts": { "test": "node --test .pi/extensions/gate/__tests__/" } +} +``` + +- [ ] **Step 2: Define the pure builder API in `builders.js`.** Each builder takes plain params and returns a plain widget-descriptor object. No `ctx`, no I/O. + +```javascript +// gate/builders.js — pure functions, no Pi context. Each returns a ui_request descriptor (spec §4.1). +function buildSelectRequest({ id, phase, title, options, intro = null, allowOther = true, recommended = null }) { /* ... */ } +function buildMultiselectRequest({ id, phase, title, options, content = null, allowEmpty = false, allowOther = true }) { /* ... */ } +function buildArtifactGate({ id, phase, title, artifact /* {kind, data, version} */, action /* {kind, prompt} */ }) { /* ... */ } +function buildInfoRequest({ phase, level, text }) { /* ... */ } +function buildFreetextRequest({ id, phase, title }) { /* ... */ } +function withChildLinkage(option, widgetSpec) { /* sets option.opens = widgetSpec; returns option */ } +module.exports = { buildSelectRequest, buildMultiselectRequest, buildArtifactGate, buildInfoRequest, buildFreetextRequest, withChildLinkage }; +``` + +- [ ] **Step 3: Write JS golden tests** — one golden file per widget type, one test per widget. Assert exact shape. + +```javascript +// gate/__tests__/builders.test.js +const test = require("node:test"); +const assert = require("node:assert"); +const fs = require("node:fs"); +const path = require("node:path"); +const { buildSelectRequest, buildMultiselectRequest, buildArtifactGate } = require("../builders.js"); + +const GOLDEN = path.join(__dirname, "golden"); + +test("select F1 matches golden", () => { + const result = buildSelectRequest({ id: "u1", phase: "F1", title: "Disambigua 'ablazione'", + options: [{ id: "o1", label: "procedura" }, { id: "o2", label: "patologia" }] }); + const golden = JSON.parse(fs.readFileSync(path.join(GOLDEN, "select_F1.json"))); + assert.equal(result.widget, "select"); + assert.deepEqual(result.reserved, golden.reserved); // back/exit/other + assert.deepEqual(result.options, golden.options); + assert.ok(result.schema_version); +}); + +test("artifact-gate F5 schema matches golden", () => { + const result = buildArtifactGate({ id: "u2", phase: "F5", title: "Schema-linking", + artifact: { kind: "schema_linking", data: { candidates: [] }, version: 1 }, + action: { kind: "confirm", prompt: "Confermi?" } }); + assert.equal(result.widget, "artifact-gate"); + assert.equal(result.artifact.kind, "schema_linking"); + assert.equal(result.action.kind, "confirm"); +}); + +test("multiselect F4 carries content + allowEmpty", () => { + const result = buildMultiselectRequest({ id: "u3", phase: "F4", title: "Tabelle", + options: [], content: { candidates: [] }, allowEmpty: false }); + assert.equal(result.widget, "multiselect"); + assert.equal(result.allow_empty, false); + assert.ok("content" in result); +}); + +test("Altro option carries freetext linkage", () => { + const result = buildSelectRequest({ id: "u4", phase: "F1", title: "x", options: [] }); + const other = result.options.find(o => o.id === "other"); + assert.equal(other.opens.widget, "freetext"); +}); +``` + +- [ ] **Step 4: Create the golden files**, run `npm test`, verify pass. + +```bash +cd harness && npm test +``` + +- [ ] **Step 5: Fuzzy tests on builders (L1, JS)** — malformed params (missing `title`, non-list `options`, empty `options` with `allowEmpty:false`). Assert builders throw clearly or return a well-formed error descriptor, never silently produce a broken widget. + +```javascript +test("select with missing title throws clearly", () => { + assert.throws(() => buildSelectRequest({ id: "u", phase: "F1", options: [] }), /title/); +}); +test("multiselect allowEmpty:false with zero options throws", () => { + assert.throws(() => buildMultiselectRequest({ id: "u", phase: "F4", title: "x", options: [], allowEmpty: false }), /allow_empty|options/); +}); +``` + +- [ ] **Step 6: Commit** + +```bash +git add harness/package.json harness/.pi/extensions/gate/ +git commit -m "feat(harness): pure widget-builder functions + JS golden/fuzzy tests (D2, L1)" +``` + +--- + +### Task C2: Gate glue — rewrite `nsp-gate.js` (implementation; verification at L2) + +**Files:** +- Create: `harness/.pi/extensions/nsp-gate.js` (rewritten from ChironeWp3) +- Port: `harness/.pi/extensions/reserved-labels.mjs` +- **No L1 test** — the glue depends on the Pi runtime and is verified at L2 (Task D4). + +**Honest scope:** this task implements the glue but cannot unit-test it in L1. The glue's correctness is validated end-to-end at L2 with real Pi + GLM 5.2. This is the load-bearing limitation stated in the Testing Strategy headline. Do NOT attempt to mock Pi here — that is the fake-Pi follow-up (cross-cutting), not part of this plan. + +- [ ] **Step 1: Port `reserved-labels.mjs`** (BACK/EXIT/OTHER labels, `isReserved`/`stripReserved`). + +- [ ] **Step 2: Port the anti-bypass hooks verbatim** from the original `nsp-gate.js`: `tool_call` hook blocking `nsp phase advance|reopen`, `nsp decision add`, `nsp cte plan`; protected-file list (`review_decisions.jsonl`, `session_manifest.yaml`, `cte_plan.json`); input-lock; `before_agent_start` kickoff. These are load-bearing (spec D4) and unchanged. + +- [ ] **Step 3: Wire tools to builders + emission.** Each of the 4 tools (`reviewer_select`, `reviewer_decide`, `reviewer_confirm`, `rewrite_question`) calls the matching builder from C1, then emits via `ctx.sendRaw({type:"extension_ui_request", ...})` and awaits the correlated response by `id`. Preserve: no-limbo invariant (Esc/cancel re-presents — never returns undefined), recommended-option marker, reading `SCHEMA_LINKING_PHASE`/`PHASE_NAMES` from `nsp phase meta --json` (no mirroring). + +```javascript +// nsp-gate.js (the glue) — wires C1 builders to the Pi runtime +const { buildSelectRequest, buildMultiselectRequest, buildArtifactGate } = require("./gate/builders.js"); + +module.exports = function (pi) { + pi.registerTool({ + name: "reviewer_select", /* ... */, + async execute(id, params, signal, onUpdate, ctx) { + const widget = buildSelectRequest({ id, phase: await currentPhase(ctx), ...params }); + const resp = await emitAndWait(ctx, widget); // extension_ui_request → wait → response + return handleSelectResponse(resp, params); // incl. Altro→freetext linkage, Back→reopen + } + }); + // ... reviewer_decide, reviewer_confirm, rewrite_question, anti-bypass hooks, kickoff injection +}; +``` + +- [ ] **Step 4: Manual smoke (optional, pre-L2)** — launch `pi --mode rpc` from `harness/` once, confirm the extension loads and `/nuova-domanda` triggers a tool registration (no crash). This is a sanity check, not a regression test. + +- [ ] **Step 5: Commit** (verification happens at L2 in Task D4) + +```bash +git add harness/.pi/extensions/nsp-gate.js harness/.pi/extensions/reserved-labels.mjs +git commit -m "feat(harness): rewrite nsp-gate.js glue wired to builders (D2/D4) — verified at L2" +``` + +--- + +## Phase D — End-to-end validation (L1 smoke + L2 real) + +### Task D1: L1 session-coherence smoke test (pure logic, CI) + +**Files:** +- Create: `harness/tests/test_session_coherence_smoke.py` + +A pure-logic test (no LLM, no DB) that constructs a plausible ledger by hand and asserts the invariants D1-the-L2-version would check. This is the CI-runnable proxy for "full session coherence." + +- [ ] **Step 1: Write the test** — build a synthetic ledger (decisions appended directly, not via LLM) representing a full F1→F8 walk + a rollback F6→F4 + re-derive. Assert: `current_phase` folds correctly at each step, `effective_decisions` excludes the stale tail after rollback, `teardown_to_phase(4)` deletes the right artifacts, `generate_task_doc` stays under byte budget at every phase, and `nsp session consistency` (the post-rollback coherence check) passes. + +```python +# tests/test_session_coherence_smoke.py +from pathlib import Path +from nsp.decisions import append_decision, DecisionRecord +from nsp.phase import current_phase, effective_decisions +from nsp.teardown import teardown_to_phase +from nsp.taskdoc import generate_task_doc + +def test_full_walk_then_rollback_stays_coherent(tmp_path): + s = tmp_path / "sess"; s.mkdir() + # build a synthetic full-session ledger + for n in range(1, 9): + append_decision(s, DecisionRecord(seq=n, ts="t", type="phase_approved", subject=f"phase:{n}", + detail="", rationale="")) + assert current_phase(s) == 9 # max_phase + 1 + # rollback to F4 + append_decision(s, DecisionRecord(seq=9, ts="t", type="phase_reopened", subject="phase:4", + detail="", rationale="")) + report = teardown_to_phase(s, target_phase=4) + assert current_phase(s) == 4 + assert "sql_final.sql" in report.deleted_files or "cte_plan.json" in report.deleted_files + # task docs stay bounded across phases + for ph in range(1, 9): + doc = generate_task_doc(session_dir=s, phase=ph, promoted_tables=["dim_paziente"]) + assert doc.byte_budget_ok, f"phase {ph} task doc over budget" +``` + +- [ ] **Step 2: Run, verify pass; commit** + +```bash +git add harness/tests/test_session_coherence_smoke.py +git commit -m "test(harness): L1 session-coherence smoke (full walk + rollback, pure logic)" +``` + +--- + +### Task D2: `.pi/` config + prompts + theme ported + +**Files:** +- Create: `harness/.pi/settings.json` +- Create: `harness/.pi/themes/` (port theme) +- Create: `harness/.pi/prompts/nuova-domanda.md`, `riprendi-sessione.md` +- No test (config files) + +- [ ] **Step 1: Port** `.pi/settings.json`, theme, and the two prompt files from ChironeWp3. Verify `pi --mode rpc` launched with `cwd=harness/` finds the `.pi/` directory. + +- [ ] **Step 2: Commit** + +```bash +git add harness/.pi/ +git commit -m "feat(harness): port .pi/ config, prompts, theme" +``` + +--- + +### Task D3: L2 conftest + skip-when-no-.env guard + +**Files:** +- Create: `harness/tests/conftest.py` +- Create: `harness/tests/l2/__init__.py` +- Modify: `harness/pyproject.toml` (register `l2` marker) + +- [ ] **Step 1: Register the `l2` marker** in `pyproject.toml`: + +```toml +[tool.pytest.ini_options] +testpaths = ["tests"] +markers = [ + "l2: end-to-end tests requiring real GLM 5.2 + remote DB (skipped when .env incomplete)", +] +addopts = "-m 'not l2'" # CI default: skip L2 unless explicitly requested +``` + +- [ ] **Step 2: Write `conftest.py`** with a session fixture that loads `.env` and a `l2_env` fixture that skips if any required var is missing. + +```python +# tests/conftest.py +import os +from pathlib import Path +import pytest +from dotenv import load_dotenv + +REQUIRED_L2 = ["THOTH_DWH_API_KEY", "THOTH_VEC_API_KEY", "THOTH_VEC_WRITE_API_KEY", "THOTH_SSL_CA"] + +@pytest.fixture(scope="session", autouse=True) +def _load_env(): + load_dotenv(Path(__file__).resolve().parent.parent / ".env") + +@pytest.fixture(scope="session") +def l2_env(): + missing = [v for v in REQUIRED_L2 if not os.environ.get(v, "").strip()] + if missing: + pytest.skip(f"L2 skipped — missing env vars: {', '.join(missing)} (populate harness/.env)") + return True +``` + +- [ ] **Step 3: Commit** + +```bash +git add harness/tests/conftest.py harness/tests/l2/__init__.py harness/pyproject.toml +git commit -m "test(harness): L2 marker + skip-when-no-.env guard (Testing Strategy)" +``` + +--- + +### Task D4: L2 full session — "ablazione" with GLM 5.2 (human-in-the-loop) + +**Files:** +- Create: `harness/tests/l2/test_session_ablazione.py` +- Create: `harness/workspaces/chirone-test.yaml` + +This is the **end-to-end validation that L1 cannot do**: GLM 5.2 driving the real harness against the real DWH + pgvector, on the "ablazione" question (which exercises D14 value grounding + formula on a multi-column case). Default mode = human answers via terminal; scripted-answers mode for repeatability. + +- [ ] **Step 1: Create `workspaces/chirone-test.yaml`** pointing at the remote endpoints per the L2 connection params. Secrets via `${THOTH_*}`. + +```yaml +name: chirone-test +description: "Chirone DWH — L2 test workspace" +relational: + db_type: postgres + transport: rest + rest: + base_url: https://supabase-aritmolab.policlinicosandonato.it/dwh/ + api_key: ${THOTH_DWH_API_KEY} + ssl_ca: ${THOTH_SSL_CA} +vector_db: + collection: chirone_docs + dim: 768 + rest: + base_url: https://supabase-aritmolab.policlinicosandonato.it/vector/v1/ + api_key: ${THOTH_VEC_API_KEY} + ssl_ca: ${THOTH_SSL_CA} + write_rest: + base_url: https://supabase-aritmolab.policlinicosandonato.it/vector/v1/ + api_key: ${THOTH_VEC_WRITE_API_KEY} + ssl_ca: ${THOTH_SSL_CA} +evidence: + source_root: ${EVIDENCE_ROOT} + evidence_dir: evidence/chirone +embeddings: + provider: ollama + base_url: ${THOTH_OLLAMA_URL} + model: nomic-embed-text-v2-moe + dim: 768 + batch_size: 64 +execution: + allow: [cte_test, explain, preview, aggregate, export] + max_preview_rows: 10 + statement_timeout_ms: 5000 + forbidden_functions: [set_config, dblink, dblink_exec, lo_import] +``` + +- [ ] **Step 2: Write the L2 test** — spawns Pi (`pi --mode rpc`, cwd=harness/) with GLM 5.2, runs `/nuova-domanda "dammi la lista dei pazienti che hanno fatto un'ablazione nel 2025"`, and either (a) prompts the reviewer in the terminal for each gate widget (human mode) or (b) feeds canned answers (scripted mode via `--answers-file`). Asserts: a session is produced, it reaches a finalized state, the `value_grounded` and `concept_formula_approved` decisions appear (D14 exercised), `sql_final.sql` is present and read-only-validates against the DWH. + +```python +# tests/l2/test_session_ablazione.py +import subprocess, json +from pathlib import Path +import pytest + +@pytest.mark.l2 +def test_ablazione_session_human_in_loop(l2_env, tmp_path): + """L2: GLM 5.2 + real DWH. Reviewer answers gates via terminal. + Run with: pytest -m l2 tests/l2/test_session_ablazione.py -s + """ + session_dir = tmp_path / "sess" + proc = subprocess.run( + ["pi", "--mode", "rpc"], + cwd="harness/", + # the test harness feeds /nuova-domanda and relays terminal I/O for reviewer answers + ... + timeout=600, + ) + assert (session_dir / "sql_final.sql").exists() + ledger = (session_dir / "review_decisions.jsonl").read_text().splitlines() + types = {json.loads(l)["type"] for l in ledger} + assert "value_grounded" in types or "concept_formula_approved" in types # D14 exercised +``` + +- [ ] **Step 3: Run manually** (`pytest -m l2 -s`), confirm the session finalizes and D14 surfaces. This is non-deterministic and slow — it's a pre-release check, not CI. **Commit the test + workspace; the operator runs it before release.** + +```bash +git add harness/tests/l2/test_session_ablazione.py harness/workspaces/chirone-test.yaml +git commit -m "test(harness): L2 ablazione session with GLM 5.2 + real DWH (D14, human/scripted)" +``` + +--- + +### Task D5: L2 specific behaviors — value grounding (real) + memory save-one (real) + +**Files:** +- Create: `harness/tests/l2/test_value_grounding_real.py` +- Create: `harness/tests/l2/test_memory_save_one_real.py` + +Targeted L2 tests that close specific gaps L1 leaves open: real value grounding on the live schema, and real `memory save-one` upsert to pgvector. + +- [ ] **Step 1: `test_value_grounding_real.py`** — runs `nsp search "ablazione" --kind values --workspace chirone-test`, asserts multiple columns are returned (not collapsed), and that a grounding widget would surface them. Validates D14a on the real schema. + +- [ ] **Step 2: `test_memory_save_one_real.py`** — runs `nsp memory save-one` against `chirone-test` (writer key), asserts the upsert returns `upserted >= 0` (idempotent) and a subsequent `search_similar` finds the memory. Validates D11 end-to-end. + +- [ ] **Step 3: Run manually (`pytest -m l2 -s tests/l2/`); commit** + +```bash +git add harness/tests/l2/test_value_grounding_real.py harness/tests/l2/test_memory_save_one_real.py +git commit -m "test(harness): L2 value grounding (real schema) + memory save-one (real pgvector)" +``` + +--- + +### Task D6: Documentation — README + workflow editing + testing guide + +**Files:** +- Create: `harness/README.md` +- Create: `harness/docs/workflow-editing.md` +- Create: `harness/docs/testing.md` + +- [ ] **Step 1: Write `README.md`** — install, configure `.env` + `workspaces/`, run `nsp`, run L0+L1 tests, run L2 tests (with the CA-bundle copy step). + +- [ ] **Step 2: Write `docs/workflow-editing.md`** — how to edit `workflow.yaml` (add/reorder/merge/skip phases), referencing spec §5.3. + +- [ ] **Step 3: Write `docs/testing.md`** — explain the L0/L1/L2 split honestly: what each level covers and does NOT cover (in plain language: L0 = DB-touching code vs real Postgres in container; L1 = pure logic + gate builders, no DB no LLM; L2 = full system with real GLM 5.2 + remote DB, manual pre-release). State the headline limitation plainly: the LLM→gate conversation has no automated regression coverage. How to run each level (`pytest` for L0+L1, `pytest -m l2` for L2 needing `.env` + VPN + CA bundle). The security note on keys (`.env` gitignored, never logged). + +- [ ] **Step 4: Commit** + +```bash +git add harness/README.md harness/docs/ +git commit -m "docs(harness): README + workflow editing + testing guide (L1/L2)" +``` + +--- + +## Self-Review Notes + +**Spec coverage check** (spec section → task): +- §1 (ChironeWp3 as starting point, NOT assumed reliable): enforced by a **3-tier porting classification**: + - **Tier A (modified by ThothII → full regression test):** phase.py (A5), decisions.py (A4), vectorstore/_guards dual-key (B1), memory_cmd save-one (B2), search multi-column (B3), evidence formula (B4), phase_cmd meta (A8). Each has a focused L1 test for the changed behavior. + - **Tier B (ported unchanged but load-bearing → contract test):** db/connection + introspect + sampling + RRF (**L0 testcontainers**, A9 steps 1-3 — real Postgres, where "not reliable" is tested for real), rest/client + mschema/render + mschema/eligibility + 11 CLI commands (**L1 fake data**, A9 steps 4-7). The L0/L1 split inside Tier B reflects whether the module touches a DB. + - **Tier C (ported unchanged, non-load-bearing → smoke OK):** fetch_ca, output/formatting helpers. Acceptable with smoke. + This makes "not assumed reliable" honest for the load-bearing surface. ✓ +- D1 (3 projects): this plan = harness only; backend/frontend are separate plans. ✓ +- D2/D4 (widget-descriptor gate): **builders** in C1 (L1, JS, in-language — golden + fuzzy); **glue** in C2 (implementation only — verified at L2, NOT in L1, because it depends on the Pi runtime). The limitation is documented in the Testing Strategy headline. ✓ +- D3 (workspace YAML): Task A2. ✓ +- D5 (FS persistence, ledger): Task A4 (decisions port), session models in A5. ✓ +- D6 (auth): out of scope for harness (auth is backend). ✓ (noted) +- D7 (backend SQL read-only): out of scope for harness. ✓ +- D8 (nsp stays Python): all harness tasks are Python (gate glue is JS, per the source). ✓ +- D10 (testing): **three levels** — L0 testcontainers (A9 steps 1-3, runs locally), L1 fake/logic + gate builders (A2-A8, B1-B5, C1, D1; runs locally), L2 real GLM 5.2 + remote DB (D4, D5; pre-release). Honest headline in Testing Strategy: the skill→LLM→gate loop has NO automated regression coverage. ✓ +- D11 (dual key + save-one): B1 (L1 dual-key), B2 (L1 save-one mocked), D5 (L2 save-one real). ✓ +- D12 (deployment model B): harness runs locally — `.env.example` (THOTH_PROFILE), L2 against remote REST over VPN. ✓ +- D13 (free-text interpretation): B5 (L1 rationale contract), D4 (L2 with real model — the glue that records rationale is L2-only). ✓ +- D14a (value grounding): B3 (L1 aggregate_lsh_multi on fake hits), D5 (L2 real schema). ✓ +- D14b (formula): B4 (L1 formula store). ✓ (L2 formula approval rides on D4) +- D15 (rollback): A4 (retract), A5 (effective view), A6 (teardown), D1 (L1 coherence smoke), D4 (L2 glue behavior). ✓ +- D16 (context minimization): A7 (taskdoc, L1). ✓ (per-phase reset in Pi noted as execution detail) + +**Type consistency:** `effective_decisions`, `teardown_to_phase`, `generate_task_doc`, `load_workflow`, `save_formula`, `retrieve_formula`, `aggregate_lsh_multi`, `decision_retracted`, `value_grounded`, `concept_formula_approved/rejected`, builder functions (`buildSelectRequest` etc.) — names match across tasks. ✓ + +**Testing honesty (the load-bearing facts):** +- **The skill→LLM→gate loop has NO automated regression coverage.** It is exercised only at L2 (manual, non-deterministic, pre-release). This is a conscious choice; plan L2 runs before release. +- **The gate glue cannot be unit-tested in L1** — it depends on the Pi runtime (`pi.registerTool`, `ctx.sendRaw`). Only the pure builders are L1-tested (in JS, in-language). +- L0 (testcontainers) tests the DB-touching ported code against a real Postgres — this is where "not assumed reliable" is actually enforced for the data layer. +- L2 tests (D4, D5) are non-deterministic, slow, require credentials + VPN + CA bundle. Skipped automatically when `.env` incomplete. + +**Known follow-ups (not in this plan):** +- **Cross-cutting: a fake-Pi runtime mock** that lets the gate glue run in L1/CI. When built (likely alongside the backend plan, which also needs a fake-Pi to test its RPC client), it would close the gap "gate glue has no L1 test" AND enable a `build-scenarios`/`fake-LLM` replay harness for deterministic gate-behavior tests. This is the single highest-value testing investment not in this plan. It is deferred because it is substantial and shared with the backend. +- The SSE/REST surface the backend will build on (harness speaks JSONL via Pi + `nsp --json`). +- `session consistency` check integration into `nsp session check` (Task A5 stubs the effective view; the full assertion grows during execution). +- Per-phase context reset mechanism in Pi (D16 §4.9 step 3) — depends on Pi's context-management capability, verified when running the gate against real Pi in Task D4. + +This plan is self-contained: at the end, `harness/` runs the full 8-phase workflow emitting/consuming widget-descriptor JSON, validated by L0 (testcontainers) + L1 (logic + gate builders) on every local run, plus L2 (GLM 5.2 + real DB) pre-release. Ready for the backend to spawn and drive. diff --git a/docs/superpowers/specs/2026-06-25-thothii-architecture-design.md b/docs/superpowers/specs/2026-06-25-thothii-architecture-design.md new file mode 100644 index 00000000..f3a36937 --- /dev/null +++ b/docs/superpowers/specs/2026-06-25-thothii-architecture-design.md @@ -0,0 +1,761 @@ +# ThothII — Design dell'architettura + +**Data:** 2026-06-25 +**Stato:** Draft, in attesa di review +**Fonti:** `prd/ThothII-prd.md`, analisi di `ChironeWp3/` (harness funzionante) e `Thoth/thoth_sqldb2/` (modulo DB di riferimento) + +--- + +## 1. Obiettivo + +Costruire un sistema che, a partire da una richiesta in linguaggio naturale, generi SQL eseguibile su un database target, attraverso un workflow human-in-the-loop a 8 fasi orchestrato dal coding harness Pi in modalità RPC. Il sistema sostituisce l'interazione terminale di ChironeWp3 con un'interfaccia React guidata da uno scambio strutturato di JSON. + +ThothII si articola in **tre progetti autonomi** (harness, backend, frontend), sviluppati e testabili in modo indipendente, integrati tramite un contratto JSON esplicito. + +### Posizione su ChironeWp3 (premessa importante) + +ChironeWp3 è il **punto di partenza** dell'harness di ThothII, **non un asset intoccabile o "collaudato al 100%"**. Il suo codice (CLI `nsp` in Python, gate extension in JS, skills markdown, modelli di sessione/workflow) viene **portato dentro il progetto `harness/` di ThothII per essere rivalidato e perfezionato**, non assunto come affidabile per inerzia. Nello sviluppo si applicano quindi, per ogni componente portata: + +- **Lettura critica** del codice portato: si verifica che faccia davvero ciò che lo spec descrive, si individuano rigidità, duplicazioni (come il drift `PHASE_NAMES` già scoperto tra Python e JS), accoppiamenti nascosti, e invarianti sottintesi (come il no-limbo enforcement solo in JS). +- **Aggiornamento** dove ThothII cambia il contratto: il gate passa da TUI a widget-descriptor (D2/D4); `phase.py` passa da ladder `if==N` a `workflow.yaml` data-driven (F2); il modello vector DB passa a doppia key (D11); si aggiunge `nsp memory save-one`. Queste **non sono riusi passivi**: sono modifiche che vanno progettate e testate. +- **Copertura di test** (D10): i golden test su widget-descriptor validano il comportamento portato, anche per le parti "ereditate". Nessun componente viene considerato pronto solo perché proveniva da ChironeWp3. + +Dove lo spec dice "riuso" va inteso come "**punto di partenza da adattare e validare**", non come "codice sicuro da prendere tal quale". Le decisioni che scelgono il riuso lo fanno perché **riducono il rischio rispetto a una riscrittura da zero** — ma il riuso stesso è lavoro di adattamento, non un'assunzione di affidabilità. + +--- + +## 2. Decisioni architetturali ( locked ) + +Le decisioni seguenti sono state prese durante il brainstorming. Ogni voce riporta l'opzione scelta e il perché. + +**D1 — Decomposizione: tre progetti autonomi, harness autosufficiente** +`harness/` contiene tutto il layer Pi: skills markdown, `nsp` CLI Python (codice deterministico), gate extension JS, `.pi/`. `backend/` è puro Node+Fastify. `frontend/` è React/Next/ShadCn/AGGrid. +Perché: punto di partenza ampio da ChironeWp3 (da rivalidare e adattare in `harness/`, non assunto affidabile per inerzia — vedi §1), confini puliti, l'harness resta testabile in isolamento scambiando JSON. + +**D2 — Contratto centrale: widget-descriptor JSON** +L'harness emette e riceve messaggi JSON strutturati (vedi §4) invece di un TUI. Il backend è un traduttore passivo che forwarda questi messaggi tra Pi e frontend. +Perché: mappa 1:1 i tipi di interazione del PRD, è minimale e testabile. + +**D3 — Workspace come YAML** +I workspace (DB relazionale + pgvector + evidence + embeddings) sono definiti in `harness/workspaces/.yaml`. I secret stanno in `.env`, referenziati come `${VAR}`. Il caricamento è isolato nel modulo `harness/nsp/workspace.py`. +Perché: coerente con ChironeWp3, versionabile, testabile; il modulo `workspace.py` è il confine per future migrazioni. + +**D4 — Gate come extension JS dentro Pi** +La logica del gate (anti-bypass, input-lock, presentazione dei widget, iniezione kickoff) resta in un'extension JS dentro Pi, adattata per emettere widget-descriptor via `extension_ui_request`/`extension_ui_response`. Il backend non implementa gate. +Perché: la logica del gate è la parte più delicata di ChironeWp3 (anti-bypass, no-limbo, iniezione kickoff) ed è il punto di partenza più ragionevole — nonostante richieda rivalutazione e adattamento al nuovo contratto widget-descriptor (vedi §1). Il protocollo RPC di Pi è progettato per questo. Spostarla nel backend significherebbe reimplementare l'anti-bypass da zero. + +**D5 — Persistenza su filesystem (identica a ChironeWp3)** +Sessioni e artefatti su `harness/sessions//`. Il ledger delle decisioni (`review_decisions.jsonl`) è append-only ed è la verità. La fase corrente è derivata (chronological fold del ledger). Il backend fa da proxy REST verso il filesystem via `nsp ... --json`. +Perché: il modello di sessione di ChironeWp3 (ledger append-only + fold cronologico) è uno dei pezzi di valore e il punto di partenza più solido — da rivalutare e adattare (es. introducendo `schema_version`, §5.3) ma non da riscrivere da zero nel backend. + +**D6 — Auth a livelli con middleware OIDC pluggabile** +Tre modalità selezionate da config: `none` (utente `dev@local`, **modalità primaria nell'MVP modello B** — l'operatore è sulla propria macchina), `mock` (utente statico da header per test), `oidc` (OIDC standard: Authentik, Entra ID — stessa codepath, rilevante nell'evoluzione ad A "web app centrale"). Le sessioni ThothII sono associate all'utente autenticato (campo `author` nel manifest). +Perché: copre tutti i casi del PRD con una sola codepath OIDC; `none` è la scelta naturale per l'MVP modello B (postazione singolo-operatore) e rende i test dell'harness indipendenti dall'auth. + +**D7 — Backend Node+Fastify+TS con DB ibrido** +`nsp` Python possiede tutta la logica DB deterministica (introspection, value sampling, RRF/LSH, schema-link, validazione/preview durante il workflow). Il backend Node si collega al DB del workspace solo per eseguire lo SQL finale approvato in read-only, alimentare AGGrid e gli export. +Perché: coerente con D1 (harness autosufficiente); risolve informix (driver Node assenti); nessuna riscrittura di thoth_sqldb2 in TS. +**Rischio architetturale (D7):** due codepath di esecuzione SQL vanno mantenute allineate sull'enforcement read-only. Mitigazione: un'unica fonte di verità per "cosa è permesso eseguire" (config `execution.allow` del workspace, condivisa tra `nsp` e backend). + +**D8 — `nsp` CLI resta Python** +Riuso diretto di ChironeWp3 (SQLAlchemy, psycopg2, datasketch/LSH, RRF, schema-link, output `--json`). Il gate JS dentro Pi lo chiama via `bash`. +Perché: partire dal codice Python esistente di ChironeWp3 (SQLAlchemy, psycopg2, datasketch/LSH, RRF, schema-link) riduce il rischio rispetto a una riscrittura TS da zero — ma il codice va comunque portato in `harness/`, letto criticamente, adattato ai nuovi contratti e coperto dai golden test (D10). I "due runtime" nell'harness non sono un problema reale. + +**D9 — Strategia di implementazione: vertical slice per fase del workflow** +Ordine harness → backend → frontend, ma con loop end-to-end precoci: ad ogni ciclo di sviluppo si estendono tutti e tre i layer solo per la **fase del workflow** corrente (F1, poi F2, ecc. — non "fasi di sviluppo" generiche). Feedback continuo, nessun progetto in sospeso. + +**D10 — Test dell'harness: fake-Pi JSONL + golden test** +Un `fake-pi` (script Python o JS) implementa il protocollo RPC di Pi. I test verificano che i widget-descriptor emessi matchino golden JSON salvati. Deterministico, senza LLM né Pi reale nei CI. +Perché: è l'unica che testa il contratto widget-descriptor (la parte nuova) in modo deterministico e ripetibile. + +**D11 — Modello a doppia API key per il vector DB + upsert mirato delle Memory** +Il vector DB (pgvector) è esposto via REST con **due endpoint separati, due API key distinte**: +- **Reader** (`vector_rest`, key `*_API_KEY`): solo `search_similar` / `list_tables`, sola lettura. +- **Writer** (`vector_write_rest`, key `*_WRITE_API_KEY`, opzionale): solo `existing_vector_hashes` / `upsert_vector_records`, upsert + hash sync, **no delete/clear**. + +Entrambi condividono la stessa URL e lo stesso tipo di config (`RestConfig | None`). "Writer assente" = sezione omessa o key vuota. Un unico client HTTP (`VectorRestClient`) parametrizzato dalla config: quale `RestConfig` gli viene passata determina key e allowlist. + +Per salvare una Memory generata durante una sessione (anche remota), ThothII introduce un **nuovo comando mirato `nsp memory save-one `** (deviazione controllata da ChironeWp3, dove il salvataggio avviene solo via `memory index` = resync completo, server-only). `save-one` fa un singolo upsert di un record via writer key, usando le RPC `upsert_vector_records` esistenti, con hash dedup client-side (SHA-256 del content, embedda solo new/changed). È gated da `require_vector_write_allowed` (permesso su workstation **solo se** writer configurato, exit 4 altrimenti). +Perché: il caso d'uso reale è "salvare la memory appena creata in F5" — un resync intero (`memory index`) è sovradimensionato e `memory promote` è server-only. L'upsert mirato è efficiente e abilita il lavoro remoto (scopo esplicito della doppia key). Riuso totale delle RPC writer e del pattern di gating di ChironeWp3; solo il comando `save-one` è nuovo. + +**D12 — Modello di deployment: postazione remota (B) nell'MVP, evoluzione possibile a web app centrale (A)** +L'MVP segue il modello "postazione remota" di ChironeWp3: **tutti e tre i layer girano sulla macchina dell'operatore** in localhost (pattern "app desktop con UI browser", come Jupyter o VS Code server). Il vector DB e il DWH restano centrali, raggiungiti via REST con la doppia key (D11). L'evoluzione futura a "web app centrale" (A) sposta backend+harness su un server centrale e serve il frontend via browser a più utenti; i contratti FE↔BE e interni **non cambiano** — è un riposizionamento di deployment, non una riscrittura. +Perché: corrisponde al modo di lavoro reale attuale (operatori su postazioni dedicate fuori dal server). Mantenere i 3 layer anche in B rende l'evoluzione B→A pulita (nessun refactor dei contratti) e lascia all'harness le sue responsabilità (workflow, gate) e al backend le sue (REST/SSE, job, API stabile) senza mescolarle. + +**D13 — Il testo libero dell'utente va interpretato, non ignorato** +Ogni volta che una risposta permette testo libero (campo `freetext` del widget, opzione "Altro — specifica", motivazione di un rifiuto, steering `!`), l'harness **deve valutare il testo dell'utente cercando di interpretarlo al meglio nel contesto corrente della sessione** (domanda, fase, artefatto mostrato, decisioni già prese), invece di passare alla risposta di default. Questa è una **deviazione comportamentale esplicita** da ChironeWp3, che tende a ignorare il testo libero a favore della risposta di default. +Perché: il revisore che si prende la briga di scrivere testo libero sta comunicando qualcosa che le opzioni predefinite non coprono. Ignorarlo degrada la qualità del risultato (la sua correzione va persa) e la fiducia nell'interazione. Il costo è nel prompt/gate, non in una nuova infrastruttura. Vedi §4.6 per il comportamento atteso. + +**D14 — Gestione esplicita delle richieste incomprensibili (non ambigue): Value/Schema Linking e SQL Formula Evidence** +ChironeWp3 tratta due casi critici in modo inadeguato e vanno **sviluppati esplicitamente** in ThothII come capacità di prima classe dell'harness. Entrambi riguardano la situazione in cui la richiesta dell'utente **non è ambigua** (il revisore sa cosa vuole) ma è **incomprensibile per il modello** senza chiarimenti o evidenze. Vedi §4.7 per il dettaglio dei due casi. +Perché: sono i casi in cui un NL→SQL "silenzioso" produce SQL sbagliato senza che nessuno se ne accorga (il modello indovina la colonna sbagliata per un valore, o inventa una formula per un concetto). Svilupparli è core-value, non optional. Sono punti di **perfezionamento sostanziale** del codice portato da ChironeWp3 (vedi §1). + +**D14a — Value and Schema Linking (value → column grounding).** Quando la domanda cita un valore (es. "ablazione", "DRG 123", "fibrillazione atriale") il cui mapping a colonna/e è poco chiaro per il modello, l'harness deve **chiarire con l'utente** (usando l'indice LSH + RRF, che già restituiscono `table.column → valore → score`) quale colonna/e corrispondono al valore — gestendo esplicitamente il caso multi-colonna e quello in cui il valore richiede una formula (si collega a D14b). In ChironeWp3 la metà retrieval esiste (`nsp search --kind values`) ma manca del tutto la metà workflow: nessun decision type, nessuna istruzione skill, nessun widget, l'aggregazione LSH collassa un valore presente in N colonne a una sola. +Perché: è il caso in cui il modello scrive `WHERE colonna_sbagliata = 'ablazione'` in silenzio. Il revisore che cita un valore lo fa apposta — va confermato il grounding prima che diventi SQL. + +**D14b — SQL functions and formula evidence (concept → formula, reviewer-approved).** Quando la domanda contiene un concetto calcolato (es. "fascia di età pediatrica", "indice di Charlson", "ricovero a 30 giorni", ma anche "ablazione" quando richiede più colonne) che si traduce in una formula SQL su più campi, l'harness deve **proporre la formula e chiedere l'approvazione del revisore** prima che fluisca nel CTE/SQL. Richiede un nuovo tipo di evidenza "formula" (il `tier:"concept"` di ChironeWp3 è definito ma **mai usato**; il contenuto esiste già negli `30-esempi-nlq/*.md` ma come prose, non come unità recuperabile/validabile) e un flusso di approvazione per-concetto (analogico al gate CTE di F6, ma a granularità concetto e potenzialmente riutilizzabile tra sessioni). +Perché: oggi il modello inventa la formula o la legge da prose non strutturata, senza conferma. Un concetto calcolato tradotto male invalida tutta la query. L'approvazione del revisore sul SQL-espressione + colonna/e è load-bearing per la correttezza. + +**D15 — Rollback a tre granularità con teardown completo dei documenti** +Il revisore deve poter tornare indietro a tre livelli: (a) **ripresenta il widget corrente e scarta l'ultima risposta** (granularità step, dentro la stessa fase); (b) **torna all'inizio dello step precedente del workflow** (fase precedente); (c) **torna a uno step specifico** (qualsiasi fase precedente). In ogni caso di rollback, **tutte le scelte fatte dopo il punto di rollback vanno dimenticate e gli artefatti prodotti vanno cancellati**. Questa è una **deviazione sostanziale** da ChironeWp3, dove `phase reopen` fa solo `append_decision` (niente teardown), gli helper aggregano decisioni stale+nuove ignorando il boundary di reopen, e non esiste la granularità step. Vedi §4.8. +Perché: un rollback che lascia artefatti stale e decisioni incoerenti è peggio di niente — il modello/prossimo passo legge uno stato inconsistente (es. CTE orfani che bloccano `finalize`, già verificato come bug latente in ChironeWp3). La correttezza post-rollback è load-bearing. + +**D16 — Minimizzazione del contesto per LLM medio (35B, <200k token) tramite task document per-step** +L'architettura deve far sì che ad ogni passaggio il modello riceva **un singolo documento di task con esattamente le informazioni necessarie per eseguire il task corrente, derivate dagli step precedenti** — non l'intero contesto accumulato nella conversazione. Obiettivo: poter usare un modello di medie dimensioni (35B param, contesto <200k). L'implementazione richiede: (1) un generatore di **task document** che legge gli artefatti precedenti e emette il slice minimale per la fase corrente; (2) **enforcement** che il modello non legga mai artefatti integrali fatali (es. `physical.yaml` = 760KB ≈ 190k token, `report.md` = 344KB — entrambi fatali per un 35B); (3) gestione del contesto della chat (pruning/ricomposizione) perché la conversazione non cresca senza bound. Vedi §4.9. +Perché: la base "context-fresh dagli artefatti via `nsp`" di ChironeWp3 è giusta, ma non compatta il contesto e non enforce i bound — un modello 35B collasserebbe leggendo lo schema fisico integrale. La target hardware impone il constraint; il task document per-step è la soluzione architetturale. + +--- + +## 3. Architettura dei tre progetti e flusso dei dati + +**Modello di deployment (D12, MVP = B "postazione remota"):** i tre layer girano in localhost sulla macchina dell'operatore. Il vector DB e il DWH restano centrali, raggiungibili via REST con la doppia key (D11). L'evoluzione futura ad A ("web app centrale") riposiziona backend+harness su server centrale senza modificare i contratti. + +``` +┌─ POSTAZIONE OPERATORE (localhost, MVP) ─────────────────────────────┐ +│ │ +│ ┌─ FRONTEND (frontend/) ─────────────────────────────────────────┐ │ +│ │ browser → localhost · React + Next.js + ShadCn + AGGrid │ │ +│ │ Consuma SOLO la REST del backend (in localhost). │ │ +│ │ • SSE: stream di eventi (text_delta, ui_request, lifecycle) │ │ +│ │ • POST: risposte utente, azioni (reset fase, export) │ │ +│ └───────────────────────────────▲──────────────────────────────────┘ │ +│ │ HTTP/SSE (JSON, localhost) │ +│ ┌─ BACKEND (backend/) ───────────┴────────────────────────────────┐ │ +│ │ Node.js + Fastify + TypeScript · gira in localhost │ │ +│ │ • Avvia Pi: spawn("pi", ["--mode","rpc"], {cwd: repoRoot}) │ │ +│ │ • RpcClient: LineSplitter (LF only!) + dispatch per id │ │ +│ │ • Traduttore: Pi extension_ui_request → ui_request (FE) │ │ +│ │ FE ui_response → extension_ui_response (Pi) │ │ +│ │ • DB workspace: SOLO SQL finale read-only → AGGrid/export │ │ +│ │ • REST: workspaces, sessions, workflow, artifacts │ │ +│ │ • Auth middleware (D6): none (primaria MVP B) | mock | oidc │ │ +│ └───────────────────────────────▲──────────────────────────────────┘ │ +│ │ JSONL (newline-delimited, LF only)│ +│ ┌─ HARNESS (harness/) ───────────┴────────────────────────────────┐ │ +│ │ ├── .pi/ ← config progetto Pi │ │ +│ │ │ ├── settings.json, themes/ │ │ +│ │ │ ├── prompts/ ← /nuova-domanda, /riprendi │ │ +│ │ │ ├── skills/nsp-sessione/ ← SKILL.md + rewriting/mem/cte/sql│ │ +│ │ │ └── extensions/nsp-gate.js ← GATE: anti-bypass + widget │ │ +│ │ ├── nsp/ (Python package) ← CLI deterministica, parla --json│ │ +│ │ │ ├── cli/ ← command groups (da ChironeWp3) │ │ +│ │ │ ├── workspace.py ← caricamento YAML (confine D3) │ │ +│ │ │ ├── workflow.py ← lettura workflow.yaml (F2) │ │ +│ │ │ ├── db/, rest/, search/, mschema/, vectorstore/, session/ │ │ +│ │ ├── workflow.yaml ← UNICA definizione workflow (F2) │ │ +│ │ ├── workspaces/*.yaml ← definizioni workspace │ │ +│ │ ├── sessions// ← persistenza FS (locale, PII) │ │ +│ │ └── tests/ + fake-pi/ ← golden test (D10) │ │ +│ └──────────────────────────────────────────────────────────────────┘ │ +│ │ +└────────────────────────────────▲─────────────────────────────────────┘ + │ HTTPS (443), doppia API key (D11) +┌─ SERVER / SUPABASE CENTRALE ────┴─────────────────────────────────────┐ +│ /dwh/ PostgREST ──► DWH (schema datawarehouse, read-only) │ +│ /vector/v1 RPC allowlist │ +│ ├── reader (key reader): search_similar, list_tables │ +│ └── writer (key writer): existing_vector_hashes, │ +│ upsert_vector_records (no del)│ +│ ──► pgvector (schema_records/evidence/ │ +│ memory) │ +└──────────────────────────────────────────────────────────────────────┘ +``` + +### Flusso di una decisione (es. "promuovi tabella" in F4) + +1. Il LLM dentro Pi chiama il tool `reviewer_decide`. Il gate `nsp-gate.js` costruisce un widget-descriptor `{type:"ui_request", id:"u42", phase:"F4", widget:"multiselect", options:[...]}`. +2. Il gate lo emette via `ctx.ui.custom`. In RPC mode diventa `extension_ui_request`. +3. Il RpcClient del backend lo riceve, lo forwarda via SSE al frontend come `ui_request`. +4. Il frontend renderizza il widget, l'utente seleziona, il FE fa `POST /sessions/:id/response` con `{type:"ui_response", id:"u42", choices:[...], decision:{type:"table_promoted"}}`. +5. Il backend lo traduce in `extension_ui_response` e lo scrive su stdin di Pi. +6. Il gate lo legge, chiama `nsp decision add` (autorizzato dal suo stesso hook anti-bypass), il ledger si aggiorna, la fase deriva. + +--- + +## 4. Il contratto widget-descriptor (D2) + +Questa è la parte centrale: il "linguaggio" tra harness, backend e frontend. Deriva dalla verifica esaustiva dei pattern di interazione reali di ChironeWp3 (10 pattern, 4 primitive UI), mappati su una tassonomia di 6 widget. + +### Principi di flessibilità + +Il contratto è progettato per accogliere future modalità di interazione: + +- **`widget` è un campo aperto, non un enum chiuso.** I 6 widget sono i `kind` iniziali registrati. Aggiungerne uno richiede definire un nuovo `kind`, il suo payload e il renderer frontend. Non cambia l'infrastruttura (forwarding, correlazione `id`, ledger). +- **Estensibilità per composizione, non per enumerazione.** I comportamenti complessi si ottengono componendo primitive (artefatto+decisione = `artifact-gate`; select+testo se Altro = linkage `option.opens`; multiselect+contesto = `multiselect` con `content`). +- **Versioning del contratto + fallback graceful.** Ogni messaggio porta `schema_version`. Il frontend gestisce i `kind` sconosciuti con un fallback universale: se arriva un widget che non sa renderizzare, mostra il payload come JSON formattato in un box "Widget non supportato (kind: X) — rispondi manualmente". + +### Messaggi harness → backend → frontend + +```jsonc +// ui_request — l'harness chiede qualcosa all'utente +{ + "type": "ui_request", + "id": "u42", // correlazione con la risposta + "schema_version": 1, + "session_id": "2026-06-25-...", + "phase": "F4_schema_linking", // fase del workflow + "title": "Conferma le tabelle per lo schema-linking", + "intro": "Ho selezionato 3 tabelle candidate. Segna quelle da promuovere.", + "widget": "multiselect", // vedi tassonomia §4.1 + "options": [ // per select/multiselect + {"id": "t_pazienti", "label": "pazienti", + "meta": {"signals": {...}}, "selected": true} + ], + "reserved": ["back", "exit", "other"], // opzioni di controllo framework + "timeout_ms": null +} + +// info — messaggio informativo (tipo 1 PRD), non blocca, niente risposta +{ + "type": "info", + "schema_version": 1, + "session_id": "...", + "phase": "F4_schema_linking", + "level": "info", // info|warning|error + "text": "Sto per proporti lo schema-linking..." +} +``` + +### Messaggi frontend → backend → harness + +```jsonc +// ui_response — la risposta dell'utente a una ui_request +{ + "type": "ui_response", + "id": "u42", // correla la ui_request + "kind": "multiselect", + "choices": ["t_pazienti", "t_ricoveri"], + "decision": {"type": "table_promoted"} // mappa sul ledger +} + +// reserved control — "torna indietro"/"esci"/"altro" +{"type": "ui_response", "id": "u42", "control": "back"} +{"type": "ui_response", "id": "u42", "control": "freetext", "text": "..."} +``` + +### 4.1 Tassonomia dei widget (6) + +Ogni voce: copertura, caratteristiche obbligatorie, esempio d'uso. + +**`info`** — fire-and-forget (notifica toast, livello `info`/`warning`/`error`). Non blocca, non richiede risposta. +Usata per: notifiche di sistema (es. "sto per proporti lo schema-linking"), preflight Ollama, warning di input-lock. + +**`select`** — single-pick da una lista di opzioni. +Caratteristiche obbligatorie: escape hatch framework iniettati (`Altro`/`Torna indietro`/`Esci`), marker `option.recommended` evidenziato "(consigliato)", linkage `option.opens` per aprire un widget figlio, invariante no-limbo (Esc/cancel non è mai risposta → loop). +Usata per: disambiguazione F1, conferme sì/no (riapri fase), scelte singole. + +**`multiselect`** — multi-pick con checkbox. +Estensioni richieste: stato iniziale pre-selezionato (`selected[]`), toggle "seleziona/deseleziona tutti", `artifact.content` embed scrollabile (contesto), descrizione del focus (progressive disclosure), `allow_empty: true|false`, `Altro` inline che ritorna `{text, choices}` insieme, side-effect di deselezione (può emettere decision negativa). +Usata per: promozione tabelle F4, selezione memoria F2/F5, evidence accepted/rejected. + +**`freetext`** — testo libero. Sempre figlio di `select`/`artifact-gate` (via `Altro`/`Rifiuta`) o via canale steering ambientale. Il testo va **interpretato** dall'harness nel contesto della sessione, non ignorato a favore di una risposta di default (D13, vedi §4.6). +Usata per: "Altro — specifica", motivazione del rifiuto, steering `!`. + +**`artifact-gate`** — artefatto + lista disposizioni (fusione di artefatto e select del Pattern 3 di ChironeWp3). È il cuore decisionale di F5/F6/F7. +Contenuto: artefatto renderizzato (schema_linking con render umano + JSON intero; cte che legge `ctes/.sql`; sql che legge `sql_final.sql` + esegue preview live; memory, question, result) + disposizioni (`confirm`/`approve_reject`/`view_only`). `Rifiuta` → linkage a `freetext` per la motivazione. +Perché separato da `artifact`: ChironeWp3 non mostra mai un artefatto puro — l'artefatto è sempre accoppiato a una decisione (Approva/Rifiuta/Altro/Torna/Esci). La decisione è load-bearing sul documento. + +**`artifact`** — view-only, nessuna disposizione. Per artefatti da mostrare nel pannello destro senza decisione (COT, thinking, risultato in lettura). + +### 4.2 Linkage annidato + +Un'opzione può dichiarare quale widget apre se scelta: + +```jsonc +ui_request select: + options: [ + { id:"approve", label:"Approva" }, + { id:"other", label:"Altro — specifica…", + opens: { widget:"freetext", title:"Specifica…" } }, + { id:"reject", label:"Rifiuta (rivedi e riprova)", + opens: { widget:"freetext", title:"Motivazione del rifiuto" } }, + { id:"back", label:"Torna indietro", + opens: { widget:"select", title:"A quale fase?", options:[...] } } + ] +``` + +Il frontend, se l'utente sceglie "other", mostra il widget figlio e raccoglie entrambe le risposte, inviandole insieme nella `ui_response`. Il testo raccolto via `Altro` o `Rifiuta` è testo libero che l'harness deve interpretare nel contesto (D13, §4.6), non un'etichetta da archiviare e dimenticare. + +### 4.3 Canale steering ambientale + +Free-text non modale (prefisso `!` in ChironeWp3): il revisore può iniettare testo arbitrario in qualsiasi momento durante una sessione attiva. Non è un widget modale, è un canale stream-level. Modellato come un endpoint `POST /sessions/:id/steer` con `{text}`. Lo steering è il caso più forte del principio D13: il revisore interrompe appositamente per ridirezionare — ignorare il testo o trattarlo come "continua con il default" vanifica lo scopo del canale. + +### 4.4 Evento di sistema: auto-advance silenzioso + +Una fase può chiudersi senza input utente (F2/F6 vuote, `phase_auto_approved`). Non è un widget — è l'assenza di interazione. Modellato come evento SSE distinto `{type:"system_event", event:"auto_advance_silent", phase:"F2"}` perché il revisore possa capire perché una fase è sparita. + +### 4.5 Ledger decisioni (enumerati leggendo il codice ChironeWp3, 22 tipi) + +I `decision.type` del widget-descriptor sono esattamente quelli del `review_decisions.jsonl` di ChironeWp3, verified leggendo `src/psdwp3/session/decisions.py:9-32`: + +- **F1 chiarimento:** `concept_clarified`, `ambiguity_open` +- **F2 memoria:** `memory_rejected` (side-effect di deselezione), `phase_auto_approved` (vuota) +- **F3 riscrittura:** `question_rewritten` +- **F4 schema-linking:** `table_promoted`, `table_excluded`, `column_corrected`, `join_modified`, `evidence_accepted`, `evidence_rejected` +- **F5 sintesi:** `phase_approved` +- **F6 cte:** `phase_skipped`, `cte_approved`, `cte_corrected`*, `cte_rejected`, `phase_auto_approved` +- **F7 sql finale:** `sql_revised`, `sql_approved`, `sql_rejected` +- **F8 datamart:** `datamart_requested`, `datamart_declined` +- **meta cross-fase:** `phase_approved`, `phase_auto_approved`, `phase_reopened`, `phase_skipped` + +(*) `cte_corrected` è nel `Literal` ma non è mai emesso dal gate di ChironeWp3 (reserved/historic). + +### 4.6 Interpretazione del testo libero (D13) + +Questa è una **deviazione comportamentale esplicita da ChironeWp3**, registrata come D13. Il principio: quando l'utente fornisce testo libero, sta comunicando qualcosa che le opzioni predefinite non coprono. L'harness lo valuta e cerca di interpretarlo nel contesto, invece di passare alla risposta di default. + +**Dove si applica (tutti i canali di testo libero):** + +- **`Altro — specifica…`** in `select` e `multiselect` (linkage `option.opens`, §4.2): l'utente descrive una scelta fuori dalle opzioni. L'harness interpreta il testo e, se necessario, lo traduce in una decisione o in una nuova proposta (es. in F1 disambiguazione, il testo può chiarire un concetto non previsto; in F4 può indicare una tabella o una correzione non in lista). +- **`Rifiuta (rivedi e riprova)` → motivazione** (linkage, §4.2): la motivazione del rifiuto non è decorativa — deve guidare la rigenerazione. Rifiutare un CTE con "la join è sbagliata, va su dim_pazienti non fact_ricoveri" deve portare a rivedere proprio quel join, non a riproporre lo stesso CTE. +- **Steering ambientale `!`** (§4.3): è il caso più forte. L'utente interrompe per ridirezionare ("stai escludendo i pazienti pediatrici", "considera solo il 2024"). Ignorarlo o trattarlo come "continua" vanifica lo scopo. + +**Comportamento atteso dell'harness (il "cosa", non il "come"):** + +1. **Riceve** il testo libero nella `ui_response` (campo `text` per Altro/motivazione, o payload dello steering). +2. **Lo valuta nel contesto corrente**: domanda originale (e sua versione riscritta in F3+), fase attiva, artefatto mostrato, decisioni già presenti nel ledger, candidate/elementi in gioco. L'interpretazione è compito del LLM dentro Pi, guidato dalla skill; il gate si limita a consegnargli il testo e a non permettergli di "svignarsela" con un default. +3. **Agisce coerentemente**: applica l'interpretazione — corregge la proposta, aggiunge un vincolo, riapre una fase se serve, riscrive la domanda se il chiarimento lo richiede. Se il testo è ambiguo, **chiede chiarimento** (emette un nuovo widget) invece di indovinare in silenzio o ignorare. +4. **Lascia traccia**: quando il testo libero influenza una decisione, questa va registrata nel ledger con il `rationale` che riporta (anche sintetizzato) il testo dell'utente, così l'origine della decisione è ricostruibile. + +**Cosa NON deve fare (anti-pattern da ChironeWp3 da evitare):** + +- Non trattare il testo libero come etichetta opaca da archiviare e dimenticare. +- Non passare alla risposta/default di default quando c'è testo non vuoto. +- Non "accettare" la direzione dell'utente a parole e poi proseguire con la proposta originaria. +- Non ignorare il testo dello steering trattandolo come conferma generica. + +**Implementazione (dove vive):** la responsabilità è dell'harness — in particolare nella skill `nsp-sessione` (istruzioni al LLM su come trattare il testo libero in ogni fase) e nel gate `nsp-gate.js` (che consegna il testo al modello e, come per gli altri invarianti, non lascia spazio a "svignarsela"). Il backend e il frontend sono solo trasporto: il FE raccoglie il testo, il BE lo forwarda. Nessuna logica di interpretazione fuori dall'harness. È un punto esplicito di **perfezionamento del codice portato da ChironeWp3** (vedi §1), e va coperto dai golden test (D10) con scenari in cui l'utente fornisce testo libero e si verifica che l'harness ne tenga conto. + +### 4.7 Gestione delle richieste incomprensibili (non ambigue) — Value/Schema Linking e SQL Formula Evidence (D14) + +Questa sezione specifica il principio D14: due casi in cui la richiesta dell'utente **non è ambigua** (sa cosa vuole) ma è **incomprensibile per il modello** senza chiarimenti o evidenze. Sono i casi in cui un NL→SQL "silenzioso" produce SQL sbagliato senza allarme. Entrambi sono **sotto-sviluppati in ChironeWp3** e vanno sviluppati come capacità di prima classe. Sono punti di **perfezionamento sostanziale** del codice portato (vedi §1), non riuso passivo. + +**Distinzione chiave — ambiguo vs incomprensibile.** F1 (chiarimento, `concept_clarified`/`ambiguity_open`) gestisce il caso *ambiguo*: il concetto della domanda ammette più letture e il modello chiede quale. I due casi qui sono *incomprensibili*: il modello non sa tradurre un elemento specifico della domanda in schema SQL, anche sapendo cosa vuole l'utente. Sono ortogonali a F1 e oggi cadono tra le fasi. + +#### 4.7.1 Value and Schema Linking (D14a) + +**Il caso.** La domanda cita un **valore** (es. "ablazione", "DRG 123", "fibrillazione atriale", "ricovero in UTIC") che il modello non sa mappare a una colonna. Il valore non è ambiguo (l'utente sa cosa intende), ma il modello non sa *dove vive* nello schema. + +**Cosa esiste in ChironeWp3 (metà retrieval, OK).** L'indice LSH + RRF funziona: `nsp search "" --kind values --json` restituisce già `table.column → valore → score`. Il dato c'è. + +**Cosa manca (metà workflow, da sviluppare).** Tutto il flusso che usa quel dato per chiarire con l'utente: +- Nessun **decision type** per il value-grounding (es. `value_grounded`). +- Nessuna **istruzione nella skill** di: "se la domanda cita un valore, esegui `nsp search --kind values`; se il valore mappa a >1 colonna o a una colonna non ovvia, chiedi conferma del grounding." +- L'**aggregazione LSH collassa un valore presente in N colonne a una sola** (`_aggregate_lsh` in `search/__init__.py` tiene solo il best). Il caso multi-colonna non è rappresentato. +- Nessun **widget dedicato** "conferma valore X → colonna Y". + +**Il caso ablazione prova che il gap è reale e multi-colonna.** "Ablazione" nello schema Chirone si calcola da più punti: il flag `fact_..._ablazione.ablazione_transcatetere`, e/o `fact_see_ablazione_procedura_patologia.patologia = 'ablazione'`, e/o `procedure_type = 'ablazione'`. Il modello che scrive `WHERE ablazione_transcatetere IS TRUE` in silenzio può produrre una query semanticamente diversa da quella voluta. Il grounding del valore *è esso stesso* una decisione di formula (collegamento a §4.7.2). + +**Comportamento atteso in ThothII:** + +1. **Trigger:** durante la riscrittura/analisi (F3/F4), per ogni valore letterale citato nella domanda, l'harness esegue `nsp search "" --kind values`. +2. **Decisione di chiarire:** se il valore mappa a **più di una colonna**, o a una colonna che il modello non avrebbe scelto da solo, o a una colonna che richiede formula (caso §4.7.2), l'harness **presenta un grounding clarification** (widget `select` o `multiselect` che mostra i candidati `colonna → valore → score` con provenance LSH). +3. **Interpretazione del testo libero (D13):** se l'utente sceglie "Altro" e indica una colonna o una formula diversa, l'harness la valuta nel contesto, non la ignora. +4. **Registrazione:** il grounding confermato si registra nel ledger con il nuovo decision type (es. `value_grounded`, `subject` = il valore, `detail` = colonna/e scelta, `rationale` = score + eventuale testo utente) e si riflette in `schema_linking.json` (i `Candidate` oggi non hanno un campo "valore grounded" / "valore di filtro"). +5. **Connessione alla formula:** se il grounding richiede una formula (es. ablazione), si passa al flusso §4.7.2. + +**Nuovi elementi dati:** +- Decision type `value_grounded` (+ `value_grounded_multi` se serve distinguere). +- Estensione di `Candidate` in `schema_linking.json` con campo `grounded_values: [{value, column, score}]` per registrare dove ogni valore citato è stato ancorato. +- Possibilmente un widget dedicato, ma il pattern `select`/`multiselect` con `meta` (provenance + score) basta; non serve un nuovo `kind` di widget. + +**Dove nel workflow:**micro-step tra F3 (riscrittura) e F4 (schema linking), o un sotto-passo esplicito di F4. Nel modello data-driven di `workflow.yaml` (F2), si modella come un prerequisito/attività di una fase, non come fase numerica separata nell'MVP — ma l'hook va chiarito. + +#### 4.7.2 SQL functions and formula evidence (D14b) + +**Il caso.** La domanda contiene un **concetto calcolato** (es. "fascia di età pediatrica", "indice di Charlson", "ricovero a 30 giorni", "ablazione") che si traduce in una **formula SQL su più campi**. Il modello non può "indovinare" la formula: o la recupera da una evidenza di formula, o la sintetizza e la fa approvare. + +**Cosa esiste in ChironeWp3 (contenuto, ma non strutturato).** Gli `artifacts/evidence/30-esempi-nlq/*.md` contengono **SQL completo già scritto** per concetti come ablazione, cardioversione, device. Ma è prose opaca dentro esempi domanda→SQL, non un'unità "concetto → formula" recuperabile e validabile. In particolare: `tier: "concept"` è definito nello schema evidence (`evidence/model.py`) ma **mai usato** — zero file lo impostano, la retrieval non lo filtra. + +**Cosa manca (tutto il layer formula, da sviluppare).** +- Nessun **tipo dato first-class "formula"**. Né evidence né le annotation di mschema possono esprimere "concetto C = espressione SQL E su colonne [c1, c2, …]". +- Nessuna **retrieval per formula**: non esiste `nsp search --kind formula ""`. +- Nessun **decision type** `formula_approved`/`concept_formula_rejected`. +- Nessuna **istruzione skill** di proporre la formula per un concetto e chiederne approvazione prima del CTE. +- Nessun **flusso di approvazione per-concetto**: oggi si approvano CTE/query intere (F6), non formule per concetto. La parola "formula" non appare nel codice. + +**Comportamento atteso in ThothII:** + +1. **Authoring (offline):** un nuovo tipo di evidenza/artefatto per **formule di concetto** — frontmatter `{concept, columns[], sql, status (auto/draft/reviewed), sources}` + corpo SQL. Attiva il `tier:"concept"` morto (o un nuovo `kind`) rendendolo filtrato e recuperabile. Il contenuto esiste già negli `30-esempi-nlq`; va ristrutturato da "esempi domanda→SQL" a "unità concetto→formula". +2. **Retrieval:** `nsp search --kind formula ""` restituisce la/e formula/e candidata(e), con provenance (recuperata vs sintetizzata). +3. **Runtime (per-sessione):** per ogni concetto calcolato nella domanda, l'harness **propone la formula** (recuperata o sintetizzata) e **chiede approvazione del revisore** sul SQL-espressione + colonna/e, *prima* che fluisca nel CTE/SQL. Il widget è `artifact-gate` con `kind:"formula"` (template pronto: il `runArtifactGate` di ChironeWp3 già mostra SQL scrollabile + approve/reject/other). +4. **Registrazione:** decision type `concept_formula_approved`/`concept_formula_rejected`, con `subject` = concetto, `detail` = formula approvata, `rationale` = provenance + eventuale testo utente. +5. **Riutilizzo (chiusura del cerchio con la memoria):** le formule approvate possono essere persistite nello store formule (come `memory.py` fa per le decisioni riusabili) così che un "indice di Charlson" approvato in una sessione venga riutilizzato nelle successive. Da considerare post-MVP. + +**Nuovi elementi dati:** +- Artefatto formula: `concept_formulas/.{yaml,sql}` (frontmatter + SQL) o un `kind:"formula"` in evidence con campo `sql`/`expression`. +- Decision type `concept_formula_approved`, `concept_formula_rejected`. +- Estensione di `schema_linking.json` con `concept_formulas: [{concept, sql, columns, status, source}]`. +- `nsp search --kind formula` (parallelo a `--kind values` e `--kind schema`). +- Opzionale: `nsp formula save-one` (analogico a `memory save-one`, D11) per persistere una formula approvata nello store centrale via writer key. + +**Dove nel workflow:** micro-step di F4/F5 (dopo il grounding dei valori, prima della sintesi dello schema-linking). Nel modello `workflow.yaml` si modella come attività/prerequisito di una fase esistente, non come fase nuova nell'MVP. + +#### 4.7.3 Relazione tra i due casi e con il workflow data-driven + +I due casi sono **complementari e a volte sovrapposti**: "ablazione" è sia un problema di value-grounding (in quale colonna) sia di formula (booleano OR patologia OR tipo). Il design deve gestire il continuum: +- grounding semplice (valore → 1 colonna, confermato), +- grounding multi-colonna (valore → N colonne, selezionate), +- grounding con formula (valore/concetto → formula SQL su più colonne, approvata). + +Tutti e tre registrano una decisione tipizzata nel ledger e si riflettono in `schema_linking.json`, diventando parte dell'artefatto che F5 presenta per la conferma di sintesi. Essendo il workflow data-driven (F2, §5.3), **possono essere introdotti come micro-step senza rinumerare le fasi**: si aggiungono attività/prerequisiti alla fase F4 (o a una nuova F4.5 interna), e si evolve gradualmente. + +**Impatto sui test (D10):** entrambi i casi vanno coperti da golden test — scenari in cui la domanda cita un valore multi-colonna (es. ablazione) e un concetto calcolato (es. fascia pediatrica), verificando che l'harness (a) presenti il grounding/formula gate e (b) registri la decisione corretta. Senza questi test il comportamento "silenzioso" di ChironeWp3 può riaffiorare. + +### 4.8 Rollback a tre granularità con teardown (D15) + +Questa sezione specifica D15. Il requisito: il revisore può tornare indietro a tre livelli, e in ogni caso le scelte fatte dopo il punto di rollback **vanno dimenticate** e gli artefatti prodotti **vanno cancellati**. + +**Stato di ChironeWp3 (3 blocker verificati):** + +- **Nessun teardown dei documenti.** `phase reopen` (`phase_cmd.py:102-127`) fa solo `append_decision`; nessuna I/O su file. Conseguenza: `schema_linking.json`, `ctes/*.sql`, `sql_final.sql` persistono stale dopo il reopen. Bug latente confermato: CTE orfani (non più nel plan riderivato) restano su disco e **bloccano `finalize`** perché itera `glob("*.sql")` richiedendo che ognuno sia testato (`session_cmd.py:192-204`). +- **Gli helper NON sono reopen-aware.** `current_phase` (il fold, `phase.py:36-47`) è corretto, ma tutti gli altri helper leggono il ledger intero ignorando il boundary di reopen: `approved_ctes` (`phase.py:125-126`), `_has_decision`/`_has_decision_subject` (`phase.py:138-143`, usati da `advance_problems`), `build_evidence_entries` (`artifacts.py:30-39`), `_compute_promotions` (`memory.py:135`). Mescolano decisioni stale pre-reopen con quelle nuove. +- **Manca la granularità step.** "Torna indietro" (`doGoBack`, `nsp-gate.js:295-327`) apre sempre il phase-picker → reopen di fase intera. Non esiste "ripresenta il widget corrente, scarta l'ultima risposta" perché il ledger è append-only senza tombstone/retract. + +**Le tre granularità richieste:** + +(a) **Re-ask current widget (scarta l'ultima risposta, stessa fase).** Il revisore risponde male a una domanda e vuole ridarla. Semantica: l'ultima decisione registrata per il widget corrente viene **ritirata** (tombstone/retract nel ledger), il widget viene ripresentato. Non cambia la fase. +(b) **Torna all'inizio dello step precedente del workflow.** Reopen della fase precedente, con teardown di tutti gli artefatti e decisioni da lì in poi. +(c) **Torna a uno step specifico.** Reopen di una fase arbitraria precedente, stesso teardown. + +**Architettura del rollback corretto (cosa serve in ThothII):** + +1. **Vista ledger "effective as of pointer".** Un'unica funzione `effective_decisions(session)` che tutti gli helper consultano, invece degli scan ad-hoc. Implementazione: replay del ledger troncando all'ultima `phase_reopened` (o usando un generation counter / high-water-mark). `approved_ctes`, `advance_problems`, `build_evidence_entries`, `_compute_promotions` e chiunque legga il ledger **deve** passare da qui. È la singola fix architetturale più importante. +2. **Retract/tombstone per la granularità step.** Per (a) serve poter ritirare l'ultima decisione senza cancellare la riga (audit). Introdurre un decision type `decision_retracted` con `subject` = `decision_seq` della decisione ritirata; la vista "effective" lo onora (la decisione ritirata non conta più, ma resta nell'audit). In alternativa, uno slot per-step mutabile distinto dal log di audit — più invasivo, si valuta. +3. **Teardown degli artefatti al reopen.** `phase reopen` deve cancellare gli artefatti prodotti da fasi > target e ricalcolare lo stato derivato. Una funzione `teardown_to_phase(session, target)` che: cancella `schema_linking.json`/`cte_plan.json`/`ctes/*.sql`/`cte_tests.json`/`sql_final.sql`/`evidence.json` a seconda della fase target (la mappa fase→artefatto è nota dal `workflow.yaml`, §5.3 — ogni fase dichiara `artifacts_out`); ricrea quelli della fase target allo stato "vuoto/da produrre"; ricalcola `evidence.json` e l'insieme dei CTE approvati dalla vista effective. Risolve il bug degli orfani. +4. **Widget UI per le tre granularità.** Il widget `select` con linkage `option.opens` (§4.2) le copre: + - "Rispondi di nuovo a questa domanda" → semantica (a), ritira l'ultima decisione del widget corrente, ripresenta. + - "Torna all'inizio della fase precedente" → semantica (b), reopen + teardown. + - "Torna a una fase specifica…" → linkage a un widget `select` phase-picker → semantica (c), reopen + teardown. + Il frontend deve rendere chiara la distinzione e la conseguenza ("questo cancellerà X e Y"). + +**Invariante forte (da enforcement nel gate):** dopo ogni rollback, lo stato di sessione (ledger effective + artefatti su disco + fase derivata) deve essere **coerente** — non deve esistere artefatto stale né decisione stale che conti. Un check `session consistency` (parte di `nsp session check`) lo verifica; se fallisce, il rollback non è completo. + +### 4.9 Minimizzazione del contesto per LLM medio (D16) + +Questa sezione specifica D16. Il requisito: ogni passaggio deve dare al modello **un singolo documento di task con esattamente le informazioni necessarie, derivate dagli step precedenti** — non la conversazione accumulata. Target: modello 35B, contesto <200k. + +**Stato di ChironeWp3 (buona base, ma 3 gap):** + +- **Base giusta:** il design è "context-fresh dagli artefatti via `nsp`" — la skill istruisce comandi per-fase (`nsp search`, `nsp schema render --table`, `nsp memory search`), non replay della chat. Gli artefatti di sessione sono piccoli (KB). +- **Gap 1 — niente task document compilato.** Non esiste un generatore che produce "question + schema-linking-deciso + il tuo task per questa fase" in un singolo documento minimale. Il modello deve auto-assemblare il contesto lanciando i comandi giusti. +- **Gap 2 — niente enforcement dei bound.** Nulla impedisce al modello di leggere artefatti fatali. **`physical.yaml` = 760KB ≈ 190k token, `report.md` = 344KB** — entrambi fatali per un 35B/<200k. La skill steer-a su `mschema-text --table` (scoped, piccolo) ma è solo una raccomandazione. +- **Gap 3 — niente compaction della chat.** Il gate inietta un kickoff one-shot e fa steer, ma **nessun pruning/summarization**. La conversazione cresce senza bound; su 8 fasi un 35B esaurisce la finestra. + +**Architettura del context minimization (cosa serve in ThothII):** + +1. **Generatore di task document per-step.** Un componente `task_doc(session, phase, step)` che legge gli artefatti precedenti e gli input della fase, e emette un **singolo documento compatto** con: la domanda (originale + riscritta), lo schema-linking deciso (solo le tabelle/colonne promosse, non tutto lo schema), i grounding/formula approvati (D14), l'output dei CTE precedenti (riassunto, non intero), e il task specifico della fase/step corrente. Questo documento è **l'input primario del modello per il passo** — non la chat. +2. **Enforcement dei bound (deny-list di letture fatali).** Il gate (o il layer RPC) blocca le letture di artefatti integrali che sforano il budget: `physical.yaml`, `report.md` integrale, e in generale qualsiasi file > soglia (es. 50KB) senza scope. Lo schema arriva al modello **solo** come slice scoped (`mschema-text --table `) compilato nel task document. Questo è l'invariante che protegge il 35B dal collasso. +3. **Compaction della chat (per-fase).** All'inizio di ogni fase, la chat precedente viene **ricompattata**: il modello riparte dal task document della fase + un riassunto minimale delle decisioni chiave delle fasi precedenti (estratto dal ledger effective, §4.8), non dal transcript integrale. Il transcript completo resta accessibile (audit) ma non è nel contesto attivo. Meccanismo: il gate svuota/resetta il contesto attivo del modello all'inizio di ogni fase, consegnando il task document + il brief delle decisioni. (Dettaglio implementativo: dipende dalle capability di Pi di gestire il contesto; da verificare nel piano harness.) + +**Budget indicativo per un 35B / <200k:** +- Task document per-step: target <20k token (schema scoped + stato + task). +- Brief decisioni fasi precedenti: target <5k token. +- Lascia ~175k token per il reasoning del modello sul task — abbondante per un singolo step. + +**Relazione con il rollback (D15):** il task document è generato dalla **vista effective** del ledger (§4.8), quindi post-rollback riflette automaticamente lo stato corretto (decisioni stale escluse). Le due decisioni sono complementari: il rollback produce uno stato coerente, il context minimization lo serve al modello in forma compatta. + +**Impatto sui test (D10):** golden test che verificano (a) il task document di una fase contiene esattamente il slice atteso (non artefatti integrali), (b) una lettura di `physical.yaml` è bloccata, (c) post-rollback il task document esclude le decisioni stale. + +--- + +## 5. Modelli dati + +### 5.1 Workspace — `harness/workspaces/.yaml` + +```yaml +name: chirone +description: "Datawarehouse Policlinico San Donato" + +relational: + db_type: postgres # postgres|mariadb|sqlserver|informix|sqlite + transport: rest # direct|rest (rest=PostgREST, direct=nativo) + # --- per transport: direct --- + host: ${CHIRONE_DB_HOST} + port: 5432 + database: datawarehouse + schema: datawarehouse + user: ${CHIRONE_DB_USER} + password: ${CHIRONE_DB_PASSWORD} + # --- per transport: rest (PostgREST) --- + # rest: { base_url: ${CHIRONE_REST_URL}, api_key: ${CHIRONE_REST_KEY}, ssl_ca: ... } + # --- eventuale tunnel ssh (pattern thoth_sqldb2) --- + # ssh: { enabled: true, host: ..., username: ..., auth: private_key, key_path: ... } + +vector_db: + collection: chirone_docs # tabella pgvector target (schema_records|evidence|memory) + dim: 768 + # --- LOADING diretto (server-only, come ChironeWp3 vector_db): ricostruzione distruttiva --- + # local: + # host: localhost + # port: 5438 + # database: postgres + # schema: vectors + # --- LETTURA via REST remota (rpc search_similar, key reader) --- + rest: + base_url: ${THOTH_VEC_REST_URL} # es. https://host/vector/v1/ + api_key: ${THOTH_VEC_API_KEY} # header X-API-Key, read-only + ssl_ca: ${THOTH_SSL_CA} # opzionale, CA interna + # --- SCRITTURA via REST remota (upsert/hash via RPC allowlist, key writer) --- + # OPZIONALE: assente o key vuota = scrittura non abilitata (solo lettura). + # Abilita nsp memory save-one / vector index-schema su postazione remota. + write_rest: + base_url: ${THOTH_VEC_REST_URL} # stessa URL del reader + api_key: ${THOTH_VEC_WRITE_API_KEY} # key SEPARATA, ruolo vector_writer + ssl_ca: ${THOTH_SSL_CA} + +evidence: + source_root: ${EVIDENCE_ROOT} + evidence_dir: evidence/chirone + +embeddings: + provider: ollama # interfaccia embed(texts)→vectors; 1 impl nell'MVP + base_url: ${OLLAMA_URL} # http://localhost:11434 + model: nomic-embed-text-v2-moe + dim: 768 + batch_size: 64 + +execution: # fonte di verità read-only (D7 rischio) + allow: [cte_test, explain, preview, aggregate, export] + max_preview_rows: 10 + max_export_rows: 10000 + statement_timeout_ms: 5000 + # esempio: funzioni che permettono side-effect o escalation privilegi + forbidden_functions: [set_config, dblink, dblink_exec, lo_import] +``` + +Il modulo `workspace.py` carica + valida + espande `${VAR}` dal `.env`. Una sola fonte di verità per `execution.allow`, condivisa tra `nsp` (validazione durante il workflow) e backend (SQL finale read-only). + +### 5.2 Session — `harness/sessions//` + +Eredita il modello di ChironeWp3: + +``` +sessions// # = YYYY-MM-DD-HHMMSS- +├── session_manifest.yaml # id, created_at, author, status, question, database, schema, schema_version +├── question.md # "# Domanda" + "## Assunzioni" +├── review_decisions.jsonl # VERITÀ: ledger append-only +├── schema_linking.json # candidates[], joins[], excluded[], open_questions[] +├── cte_plan.json # [ "pazienti_base", "ricoveri_recenti", ... ] +├── ctes/.sql # un file per CTE +├── cte_tests.json # CteTestRecord[] +├── sql_final.sql # SQL finale approvato +├── evidence.json # evidence usate/scartate (a finalize) +├── validation_report.md # parsing/read-only/EXPLAIN/preview (a finalize) +└── risultati_.csv # export (su richiesta) +``` + +**Campi nuovi nel manifest per ThothII** (PRD richiede): +- `author`: id utente autenticato (D6) +- `summary`: domanda sintetica +- `updated_at` + `updated_by`: timestamp e autore ultima modifica +- `schema_version`: versione del workflow usato (F2, vedi §5.4) + +**Fase corrente = derivata, non memorizzata.** Chronological fold del ledger (come ChironeWp3 `phase.py`): ogni `phase_approved`/`phase_auto_approved` avanza, `phase_reopened` torna indietro. Il `review_decisions.jsonl` è la verità. + +### 5.3 Workflow — `harness/workflow.yaml` (unica fonte di verità) + +Questa è la principale deviazione architetturale da ChironeWp3, introdotta per abilitare la flessibilità futura del workflow (semplificazione, riordino, fasi opzionali) che il PRD lascia aperta. + +**Problema risolto:** in ChironeWp3 la forma del workflow è codificata in 4 posti indipendenti (`phase.py` con `MAX_PHASE=8`, `PHASE_NAMES`, `SCHEMA_LINKING_PHASE=5`, `DECISION_MIN_PHASE` 16-entry, `advance_problems` ladder `if phase==N`; `nsp-gate.js` con costanti mirrorate — già driftato: il `PHASE_NAMES` JS ha solo 7 entry e manca la fase 8; `SKILL.md` in prose; `finalize` con secondo enforcement point). Cambiare il workflow significa toccare 4 posti in 2 linguaggi. + +**Soluzione:** una sola definizione data-driven. + +```yaml +# harness/workflow.yaml +schema_version: 1 + +phases: + - id: F1 + name: chiarimento + advance: kind:phase # meccanismo di chiusura + prerequisites: [] # data-driven, no ladder if==N + - id: F2 + name: memoria + advance: auto_if_empty # auto-advance se 0 decisioni sostanziose + prerequisites: [] + - id: F3 + name: riscrittura + advance: kind:phase + prerequisites: + - decision_exists: question_rewritten + - id: F4 + name: schema_linking + advance: reviewer_decide + prerequisites: [] + - id: F5 + name: sintesi + advance: kind:phase + prerequisites: + - file_validates: [schema_linking.json, SchemaLinking] + - id: F6 + name: cte + advance: auto_if_empty_or_skipped + prerequisites: + - any: + - decision_subject_exists: [phase_skipped, "phase:6"] + - all_ctes_approved: true + - id: F7 + name: sql_finale + advance: kind:phase + prerequisites: + - decision_exists: sql_approved + - id: F8 + name: datamart + advance: reviewer_decide + prerequisites: + - any: + - decision_exists: datamart_requested + - decision_exists: datamart_declined + +decision_min_phase: auto # DERIVATO dall'ordine delle fasi +max_phase: auto # = len(phases) +``` + +**Come si ottiene la flessibilità:** + +- **`phase.py` legge `workflow.yaml`.** `max_phase = len(phases)` (non più hardcoded). `advance_problems` valuta i `prerequisites` della fase (fine della ladder `if==N`). `decision_min_phase` derivato dalla posizione della fase che emette quel decision type. Nessuna costante duplicata. +- **`nsp-gate.js` non mirrora più niente.** Legge i metadati via un nuovo comando `nsp phase meta --json` (restituisce `max_phase`, `PHASE_NAMES`, `SCHEMA_LINKING_PHASE`, advance strategy per fase). Fine del drift JS/Python (il bug F8 scompare). +- **`SKILL.md` generato o validato vs `workflow.yaml`.** Le sezioni "## Fase N" possono essere generate da `workflow.yaml`; in alternativa un check assicura che skill e yaml siano allineati. +- **`schema_version` nel manifest.** `SessionManifest` porta la versione del workflow usato. La funzione `current_phase` interpreta il ledger secondo la versione, consentendo di evolvere il workflow senza rompere le sessioni esistenti. +- **`finalize` legge gli stessi `prerequisites`.** Un solo enforcement point, non due. + +**Le 4 trasformazioni rese fattibili:** +- Ridurre 8→5 fasi (fondere): edit `workflow.yaml`, i prerequisites data-driven si adattano. +- Riordinare: edit l'ordine in yaml, `decision_min_phase` si ricalcola. +- Aggiungere una fase: aggiungi una entry in yaml. +- Rendere una fase skippable: aggiungi `prerequisites: any: [decision_subject_exists: [phase_skipped, "phase:N"], ...]`. + +**Cosa si porta da ChironeWp3 come punto di partenza** (da rivalutare in `harness/`, non assunto affidabile per inerzia — vedi §1): ledger append-only, fold cronologico come meccanismo (count-agnostic), `decision_seq` come foreign key, reopen generalizzata (a qualsiasi fase precedente), modello memoria (`REUSABLE_TYPES`, `decision_seq`), i 22 `decision.type`. Ciascuno va verificato e coperto dai golden test (D10). + +### 5.4 Vector DB — modello di accesso a doppia API key (D11) + +Il vector DB (pgvector) è raggiunto via REST con **due endpoint separati, due API key distinte**, per permettere alle postazioni remote di salvare Memory senza poter fare operazioni distruttive. Modello derivato dalla verifica del codice ChironeWp3 (`config.py`, `vectorstore/rest_client.py`, `vectorstore/rest_writer.py`, `cli/_guards.py`, `scripts/create_vector_writer_rpc.sql`). + +**Due ruoli, due config:** + +- **Reader** (`vector_db.rest`, config `RestConfig | None`). Header `X-API-Key: ${THOTH_VEC_API_KEY}`. Allowlist RPC: `search_similar`, `list_tables`. Sola lettura. Sempre necessaria per `nsp search` (RRF). +- **Writer** (`vector_db.write_rest`, config `RestConfig | None`, **opzionale**). Header `X-API-Key: ${THOTH_VEC_WRITE_API_KEY}`. Allowlist RPC: `existing_vector_hashes`, `upsert_vector_records`. Upsert + hash sync, **no delete/clear**. + +**Un unico client HTTP, parametrizzato dalla config.** `VectorRestClient(cfg: RestConfig)` è la stessa classe per reader e writer; quale `RestConfig` gli viene passata determina key e allowlist. "Writer realmente configurato" = sezione presente **e** `api_key` non vuota (`has_vector_write_rest` controlla entrambi). + +**Contratto RPC del writer (load-bearing, definito server-side in `scripts/create_vector_writer_rpc.sql`):** + +- `existing_vector_hashes(table_name text, kinds text[]) → table(record_key text, content_hash text)` — restituisce gli hash correnti per il sync differenziale. +- `upsert_vector_records(table_name text, rows jsonb) → jsonb` (`{"upserted": N}`) — `ON CONFLICT (record_key) DO UPDATE`, aggiorna `kind/content_hash/metadata/embedding/indexed_at`. **No delete.** +- Entrambi `SECURITY DEFINER`, `set search_path = public, vectors, extensions`, `REVOKE` da `public`/`anon`/`authenticated`, `GRANT EXECUTE` solo al ruolo `vector_writer`. La key writer mappa su `vector_writer` → **EXECUTE sulle funzioni, nessun DELETE sulle tabelle raw**. +- Tabelle/kinds ammessi: `schema_records` (schema_table, schema_column), `evidence` (evidence), `memory` (memory). + +**Hash dedup client-side.** `content_hash` = SHA-256 del content. Il writer confronta gli hash ricalcolati con `existing_vector_hashes`, embedda solo i record new/changed, li upserta. Idempotente per costruzione. + +**Gating (3 guard functions in `nsp/cli/_guards.py`, punto di partenza da ChironeWp3 da portare e rivalutare in `harness/`):** + +- `require_server_profile(cfg, command)` — operazioni distruttive/server-only (`vector init`, `memory clear`, `memory promote`, `memory update`, `memory delete`): **exit 4** su profilo `workstation` sempre. +- `has_vector_write_rest(cfg)` — predicato "writer realmente configurato" (sezione presente **e** key non vuota). +- `require_vector_write_allowed(cfg, command)` — operazioni di upsert (`vector index-schema`, `evidence index`, `memory index`, e il nuovo `memory save-one`): permesse su `workstation` **solo se** `has_vector_write_rest`, exit 4 altrimenti. Su `server` sempre permesse. + +**Salvataggio mirato delle Memory — nuovo comando `nsp memory save-one` (D11, deviazione controllata).** + +In ChironeWp3 il salvataggio delle Memory su pgvector avviene solo via `memory index` (resync completo dell'intero registro) o `memory promote` (server-only). Per ThothII si introduce `nsp memory save-one `: + +- Esegue un **singolo upsert mirato** (un record) del record di memoria associato a quel `decision_seq`, via writer key. +- Usa le RPC esistenti `existing_vector_hashes` + `upsert_vector_records` (nessuna nuova RPC server-side). +- Hash dedup client-side: embedda solo se il content è cambiato. +- Gated da `require_vector_write_allowed` → funziona su postazione remota se la writer key è configurata. +- Il workflow (F5 sintesi / F2 memoria) lo chiama quando una Memory viene promossa/accettata per la sessione corrente, invece di scatenare un resync completo. + +**Perché `save-one` invece di `memory index`:** il caso d'uso reale è "salvare la memory appena generata in F5", non "resyncare tutto il registro". Un resync intero è sovradimensionato e rallenta il flusso interattivo. `save-one` è efficiente, idempotente (hash dedup), e abilita il lavoro remoto — che è lo scopo esplicito della doppia key. + +**Profilo `workstation` vs `server` (variabile `THOTH_PROFILE`):** + +- `server` (default): ricostruzione distruttiva completa via `vector_db.local` diretto (init, clear, rebuild). Le guard server-only passano. +- `workstation`: blocca init/clear/promote/update/delete (exit 4). Permette upsert (`index-schema`, `evidence index`, `memory index`, `memory save-one`) **solo se** `write_rest` configurato. + +**Variabili d'ambiente (`.env`):** + +- `THOTH_PROFILE` — `server` | `workstation` +- `THOTH_VEC_REST_URL` — URL unica per reader e writer +- `THOTH_VEC_API_KEY` — key reader (search_similar) +- `THOTH_VEC_WRITE_API_KEY` — key writer (upsert), solo se upsert remoto abilitato +- `THOTH_VEC_HOST/PORT/USER/PASSWORD` — `vector_db.local` diretto, server-only (in remoto: segnaposto) +- `THOTH_SSL_CA` — path CA per HTTPS interno (alimenta `rest.ssl_ca` e `write_rest.ssl_ca`) + +--- + +## 6. Frontend / UI + +Le decisioni UI derivano dal PRD, formalizzate durante il brainstorming. + +**Layout — ibrido a 4 zone (Q9-C).** Nav sinistra (funzioni + lista sessioni) | workflow bar orizzontale sopra la chat | chat+input al centro | sidebar destra collassabile con tutti gli artefatti/COT/thinking della sessione. + +**Schema-linking viewer (Q10-A+B).** Mermaid flowchart verticale (default, top-to-bottom) + tabella gerarchica come vista alternativa via toggle. Entrambi vincolati a ≤45 elementi e sviluppo verticale. Commento "perché" inline. + +**CTE / SQL viewer (Q11-A).** Code blocks collassabili (`▾/▸`) con header (nome, n° campi, stato test), SQL formattato + commento per campo. Toggle verticale/orizzontale globale. SQL finale = stesso componente con SELECT espansa e commenti solo su JOIN/WHERE/HAVING/ORDER BY. Evidenziazione sintattica (shiki/highlight.js). Nessun parsing AST richiesto. + +**Pannello risultati (Q12-A).** Contestuale. Numero → grassetto. Lista → AGGrid community con export CSV. Selettore `[10 ▾ / tutti]`. Datamart (dbt/CSV/Excel) come azione separata con conferma pseudoanonimizzazione. + +**Testi lunghi e markdown.** Box con scorrimento orizzontale e verticale. Markdown "mermaid enhanced": formattazione + rendering degli schemi mermaid inclusi. + +--- + +## 7. Strategia di implementazione (D9) + +Vertical slice per fase. Ordine harness → backend → frontend, con loop end-to-end precoci. Schema indicativo dei cicli: + +- **Ciclo 0 (harness isolato):** `nsp` risponde a `nsp search/phase/session` in JSON, più il nuovo `nsp phase meta --json` (F2) che il gate userà per non mirrorare le costanti. Testato con script che mandano JSON a mano + fake-Pi con golden test (D10). +- **Ciclo 1 (harness + backend minimale):** il BE fa spawn di Pi, forwarda 1 widget-descriptor (F1: disambiguazione). Test end-to-end via curl. +- **Ciclo 2 (FE minimale):** aggiungi il frontend minimo (solo F1) per chiudere il loop visivamente. +- **Cicli 3+:** estendi fase per fase (F2 memory, F3 rewrite, F4 schema-link, …) su tutti e tre i layer. + +--- + +## 8. Flessibilità e parametricità — quadro + +Applicato un filtro secco: flessibilità inclusa solo dove (a) il PRD la chiede, oppure (b) risolve una tensione architetturale già emersa, e il costo è basso. + +**Accettate (6):** + +- **F1 — Widget descriptor con `kind` aperto + fallback.** Costo basso (un campo + renderer fallback). Risolve "future modalità di interazione" (richiesta esplicita). Vedi §4. +- **F2 — `workflow.yaml` come unica fonte di verità.** Risolve il drift reale già verificato (JS manca F8). Abilita semplificazione futura. Vedi §5.3. +- **F3 — Workspace YAML parametrico su `db_type` + `transport`.** È il requisito del PRD (postgres/sqlserver/mariadb/informix + REST/SSH/diretto). Vedi §5.1. +- **F4 — Auth middleware pluggabile `none`/`mock`/`oidc`.** PRD chiede esplicitamente Authentik+Entra ID+no-auth. Una sola codepath OIDC. Vedi D6. +- **F5 — `embeddings.provider` nel workspace.** Il layer embeddings è dietro un'interfaccia semplice (`embed(texts)→vectors`, 1 implementazione). Permette swap provider senza toccare RRF. +- **F6 — Modello a doppia API key per il vector DB.** Il PRD richiede lavoro remoto (postazione fuori server) con salvataggio Memory controllato. Due key (reader/writer) con allowlist RPC separate è il modo pulito per abilitarlo senza esporre delete/clear. Vedi §5.4, D11. + +**Rifiutate (8) — over-engineering:** + +- **R1 — Registry di adapter DB pluggabile nel backend Node.** Il backend (D7) tocca i DB solo per SQL finale read-only; serve un driver per il workspace corrente, non un registry eterogeneo. La complessità dei 5 DB vive in `nsp` Python (D8). Duplicare thoth_sqldb2 nel backend è senza valore. +- **R2 — Sistema di plugin per i widget del frontend.** I 6 widget + fallback (F1) coprono il caso. Un framework di plugin è speculativo: nessuna evidenza di bisogno. Un settimo widget si aggiunge come componente React. +- **R3 — Repository pattern / astrazione sulla persistenza session.** Le sessioni sono file su FS (D5), accessi via `nsp`. Avvolgere il FS in `SessionRepository` per swap futuro a DB è YAGNI. `nsp session/store.py` è già il confine. +- **R4 — Event sourcing / message bus interno al backend.** Il backend fa solo da traduttore. Non è un sistema event-driven con molti produttori/consumatori. Una message bus aggiunge complessità senza un secondo consumatore. +- **R5 — Configurabilità dei widget e del layout via configurazione.** Il PRD descrive una UI specifica. Rendere configurabile il layout = costruirla due volte. YAGNI. +- **R6 — Multi-tenancy con workflow diversi per tenant.** L'MVP è mono-utente (D6). Per-tenant workflow è speculativo. Se servirà, "workflow.yaml per-workspace" è un'estensione naturale di F2. +- **R7 — Livello di astrazione sul transport FE↔BE (SSE vs WebSocket vs polling).** SSE basta per lo streaming unidirezionale. Astrarrlo per swap futuro è YAGNI. +- **R8 — SQL builder driver-agnostic nel backend.** Il backend esegue SQL già generato da `nsp`, non lo costruisce. Un query builder è inutile. + +--- + +## 9. Fuori scope (MVP) + +- **Web app centrale multi-utente (modello A, D12):** l'MVP è modello B (postazione remota localhost). L'evoluzione ad A riposiziona backend+harness su server centrale + attiva OIDC (D6); i contratti non cambiano. +- Multi-utente reale con concorrenza (D6 prepara il terreno ma l'MVP è mono-operatore in localhost). +- CRUD workspace via API (D3: i workspace sono YAML; la scrittura è manuale, il backend li espone in lettura). +- Plugin system per DB adapter nel backend (R1), plugin widget FE (R2), repository pattern (R3), event bus (R4), configurabilità layout (R5), multi-tenancy (R6). +- Job runner asincrono per export/preview (Q12-A scelta contestuale, sincrono). +- Form builder UI per widget futuri (R2). + +--- + +## 10. Rischi aperti + +**Rischio D7 — due codepath SQL read-only.** `nsp` (Python) e backend (Node) eseguono entrambi SQL sul workspace. Devono rimanere allineate sull'enforcement read-only. Mitigazione: un'unica fonte di verità (`execution.allow` nel workspace YAML, condivisa). Punto di attenzione nello sviluppo. + +**Rischio fake-Pi — fedeltà del protocollo.** Il fake-Pi (D10) deve riprodurre fedelmente il framing del protocollo RPC di Pi (LF-only, niente `readline`; split su `\n`). Se devia, i golden test non catturano regressioni reali. Mitigazione: basare il fake-Pi sul reference `RpcClient`/`LineSplitter` di ChironeWp3 (`docs/superpowers/plans/2026-06-14-psdwp3-pi-web-console-bridge.md`). + +**Rischio workflow.yaml — curva di adozione.** Introdurre `workflow.yaml` come unica fonte di verità richiede di riscrivere `phase.py` (da ladder `if==N` a evaluation data-driven) e aggiungere `nsp phase meta --json`. È lavoro in più rispetto a "copia ChironeWp3 così com'è", ma è il prezzo della flessibilità futura richiesta. Da bilanciare con D9 (vertical slice): il `workflow.yaml` può entrare gradualmente nei cicli. + +--- + +## 11. Nota sulla struttura dei piani di implementazione + +Questo documento è la **vista d'insieme** dell'architettura ThothII. Copre i tre progetti (harness, backend, frontend) e i loro contratti. La fase di `writing-plans` produrrà **tre piani separati**, uno per progetto, nell'ordine della strategia D9 (prima harness, poi backend, poi frontend). Ogni piano è autonomo e referenzia questo spec per i contratti condivisi (widget-descriptor di §4, modelli dati di §5). I contratti tra i layer sono definiti qui una volta per tutte, così ogni piano può essere implementato e testato in modo indipendente contro il contratto — esattamente come richiesto dal PRD ("tre diversi progetti autonomi"). diff --git a/prd/ThothII-prd.md b/prd/ThothII-prd.md new file mode 100644 index 00000000..83f264e0 --- /dev/null +++ b/prd/ThothII-prd.md @@ -0,0 +1,113 @@ +# ThothII + +Obiettivo del progetto: costruzione di un sistema che permetta, a partire da una richiesta fatta in linguaggio naturale, di generare un SQL eseguibile su un certo database + +ThothII deve essere scritto in Python e Typescript/Javascript ed appoggiarsi al coding harness Pi (http://pi.dev) per l'esecuzione dei task che devono essere delegati a un AI model. + +Il codice va organizzato in tre layer. ognuno dei quali va sviluppato all'interno di una sua cartella: + +1. un frontend in React, che usa NextJs+ShadCn+AGGrid + qualunque altra libreria di frontend adatta allo scopo; +2. un backend scritto in qualunque modo, che si interfaccia con il coding harness Pi per eseguire i task che devono essere delegati a un AI model. Può essere tranquillamente un'applicazione NodeJs+Fastify +3. un harness, cioè un insieme di Typescript/Javascript + Skills in markdown che costituiscano l'harness di Pi usato in modalità rpc. E' già stato scritto un harness, disponibile in ./ChironeWp3, che funziona, ma va riscritto tenendo conto che ThothII non prevede la possibilità di interagire direttamente con il coding harness Pi. Il codice di ChironeWp3 va riscritto tenendo conto che Pi deve rispondere solo im json, in modo che le sue risposte possano essere facilmente parsate dal backend ed esposte nel frontend con i giusti widget. Ciò è una importante variazione. Inoltre + +L'architettura può essere oggetto di discussione durante la fase di brainstorming e design. + +## Punti chiave di discussione + +Per facilitare la discussione, riporto in ordine sparso idee e considerazioni varie, tutte potenziale oggetto di arricchimento durante il brainstorming. + +L'applicazione deve avere al centro il workflow di generazione del SQL a partire da una domanda espressa in linguggio naturale, ma deve anche fornire un frontend per l'esecuzione di attività normalmente fatte tramite CLI su Pi come la scelta del model, del tema, del linguaggio, la selezione, il richiamo e la navigazione di una session, ecc. La lista dei comandi di Pi e della lista del setup del frontend React da implementare sarà oggetto di discussione durante il brainstorming ed il design. + +L'aderenza al workflow delineato in ./ChironeWp3 deve essere stretta in quanto funzionante. Potrà essere oggetto di perfezionamento in versioni future, ma il MVP deeve aderire a quanto sviluppato, a parte eventuali bug fix o evidenti miglioramenti applicabili subito. + +L'applicazione, come già fa quella attualmente sviluppata, può contare su tre risorse disponibili collegandosi al server di produzione: + +- un Supabase contenente il datawarehouse per cui si vuole generare il SQL +- un pgvector, contenuto anch'esso nel Supabase, che contiene gli embeddings dei documenti che descrivono il datawarehouse +- un LLM (qwen 3.6 - 35B) utilizzabile da Pi che gira sulle GPU del server di produzione, ed è quindi gratuito + +all'interno del progetto ./ChironeWp3 vi sono già tutti gli elementi necessari per gestire la connessione col datawarehouse del policlinicosandonato, ma ThothII deve potersi interfacciare con qualunque database e con un pgvector locale nel caso non sia disponibile un pgvector remoto. Per cui deve essere previsto un insime di configurazioni destinate a implementare il concetto di workspace composto da db relazionale + pgvector (locale o remoto) su cui operare prevedendo diverse modalità di accesso (REST, tunnel ssh, accesso diretto) e diverse tipologie di db relazionale (posthres, sqlserver, mariadb ed informix innanzitutto) + +Per quanto riguarda il collegamento ad un database qualunque trovi in ./Thoth/thoth_sqldb2 del codice a cui potersi ispirarsi per l'implementazione di un modulo di connessione a database generico. + +## il workflow + +1. disambiguazione della richiesta fatta +2. recupero delle memory per loro utilizzo nella creazione dello schema-linking +3. riscrittura ed approvazione della domanda riscritta +4. creazione e discussione dello schema-linking +5. sintesi dello schema-linking determinato +6. costruzione e discussione dei CTE +7. generazione e discussione del SQL finale +8. visualizzazione, tramite AGGrid, del risultato ed eventuale generazione del dbt in grado di essere eseguito in un flusso ETL + +Questo workflow è già stato implementato in ./ChironeWp3, ma ThothII deve avere una gestne basata su una UI gestita in REACT appositamente per avere un controllo migliore del workflow e dei documenti intermedi che vengono prodotti. + +Prima di tutto ci deve essere, in ThothII, una pagina in cui si vede il workflow, come una specie di lista numerata, com visualizzazione deegli step terminati e possibilità di reset del processo al punto richiamato. Ogni step del workflow deve produrre un documento che sta alla base dello step successivo, in modo da minimizzare il contesto necessario ad ogni fase. + +## L'interfaccia utente + +l'interfaccia deve mostrare i messaggi che arrivano da PI sia come messaggi che spiegano, sia come richieste di scelta tra diverse opzioni o richieste di informazioni tramite chat. Deve però mostrare solo i messaggi principali, mentre i COT e l'eventuale thinking, se presente, devono essere visibili in finestre dedicate, nel sidebar di destra, a richiesta. + +La deve essere divisa in quattro parti, come fa Codex di OpenAI. Un sidebar di sinistra dove avere link verso determinate funzioni e una lista gestibile di sessioni. Una parte centrale divisa in due: la parte bassa come area di input, ed una parte alta dove vengono mostrati i messaggi da parte del modello e dove vengono proposte le possibili risposte alle domande poste dal modello. Una sidebar di destra dove vengono mostratigli gli artefatti come messaggi, schema-linking, COT ed SQL e dove è possibile visualizzare le COT ed il thinking, se presenti. + +Come si può vedere, in buona parte delle volte che il model in ChironeWp3 interagisce con l'utente lo fa in tre modi: + +1. lo informa di quanto sta per fare (le domande per togliere ambiquità alla richresta, o gli elementi dello schema-linking che sta per proporgli +2. gli fa la domanda vera e propria, proponendo di selezionare una o più risposte, oppure di inserire un testo libero, o di tornare indietro o di interrompere il processo. +3. gli fa vedere il risultato di un blocco di domande-risposte (intero messaggio disambiguato, intero schema-linking, intero SQL generato, ecc.) + +Per ogni tipo di interazione occorre prevedere una specifica forma di widget, ispirata a come fa codex, oppure ZCode o la versione desktop di Claude Code. E ogni output di tipo 3 deve essere correttamente gestito con la parte destra della UI, che deve comparire solo a richiesta, e deve essere possibile salvare ogni output in una pagina separata, con un link che lo riporti alla pagina principale. + +In particolare: + +- i testi lunghi devono essere inseriti in box con scorrimento orizzontale e verticale +- i markdown devono essere mostrati in modo "mermaid enanced", nel senso che devono avere la formattazione e mostrare eventuali schemi mermaid inclusi nel testo +- gli sql devono essere formattati come codice ed essere presentati in modo colorato, con le liste dei campi delle select espansi in modo orizzontale e con un widget che mi permetta di espandere o collassare le varie parti dello statement SQL + +Ovviamente l'interfaccia utente deve permettere la selezione del workspace su cui si vuole operare, che a sua volta permette la determinazione del DB, del VectorDb e della collection associata, delle evidence associate. + +## La visualizzaione dello schema-link + +Lo schema Link richiede una presentazione particolarmente accurata infatti si tratta di presentare un insieme di tabelle correlate tra di loro i campi di queste tabelle selezionati per essere fonte dei dati e il collegamento concettuale che ha portato a scegliere quei campi e quelle tabelle per estrarre le informazioni richieste. Di conseguenza è necessario prevedere una attenta impostazione nella presentazione di queste informazioni per aiutare l'utente a determinare se si tratta di impostazioni corrette o se bisogna applicare delle modifiche per ottenere il migliore dei risultati importante, quindi sarà l'impostazione grafica che dovrà essere data per comunicare queste informazioni includi tra le possibilità l'uso di Mermaid per rappresentare schemi concettuali, ma mantieni sempre l'obiettivo di contenere ad un massimo di 45 elementi da includere nello schema e di sviluppare lo schema in verticale e non in orizzontale per permetterne una maggiore leggibilità + +## l'visualizzazione dei CTE + +I CTE devono essere presentati in modo che sia possibile espandere e collassare le varie parti del CTE, e che sia possibile visualizzare il CTE in modo orizzontale o verticale. + +Inoltre per ogni campo incluso nel CTE deve esserci un commento che indica il contenuto atteso nel campo e le motivazioni per cui è stato selezionato + +## La visualizzazione del SQL Finale +Il seguente finale deve essere presentato in modo che sia leggibile facilmente da parte dell'utente. Quindi tutti i campi delle Select devono essere sviluppati in orizzontale senza preoccuparsi di commentarne il contenuto perché questo è già avvenuto a livello di CTE. Devono essere invece commentati le JOIN le WHERE le HAVING, le ORDER BY e tutti gli altri elementi che sono stati aggiunti a livello di assemblaggio finale + + +## la gestione delle sessioni + +Le sessioni devono essere: + +- listate e gestite nella sidebar di sinistra +- richiamabili con recupero delll'interezza degli artefatti della sessione richiamata e la possibilità di ripartire da un punto del workflow, modificare le risposte ad una domanda e rifare le fasi terminali del processo + +Ogni sessione deve essere evidenziata con un codice autogenerato, l'id github dell'autore, una summary della domanda, la domanda per esteso, u timestamp della data e dell'autore della creazione della sessione, un timestamp con data ed autore dell'ultima modifica. + +## La presentazione dell'esecuzione del SQL prodotto + +L'esecuzione del SQL può produrre un numero, o una lista di elementi presentabili. Nel primo caso il numero deve essere presentato in grassetto, mentre nel secondo caso deve essere usata la libreria AGGrid in versione community per presentare la lista di elementi, con attiva l'opzione di esportazione in csv. Deve essere data la possibilità all'utente di scegliere tra presentare l'intera lista o un estratto di 10 record. + +## La generazione dei Datamart + +Il passo finale previsto dal Workflow deve essere: + +- la generazione di un artefatto utile per essere inserito in un processo ETL. Nel MVP sarebbe il dbt da inserire nel processo ETL implementato in Policlinico San Donato. +- la generazione di altro tipo di artefatto secondo indicazioni indicate nei parametri di workspace; +- una lista in formato CSV dei dati estratti, sia con nomi, cognomi ed ID pseudoanonimizzati, sia con anagrafica in chiaro +- lo stesso tipo di lista ma in formato Excel + + +## Gli artifacts e le altre impostazioni di ChironeWp3 + +Il processo previsto da ThothII si basa, tra le altre cose, sulla presenza di artifacts che comprendono delle Evidence. Prevedere una gestione locale delle evidence, con memorizzazione degi chunk creati e di cui si è fatto l'embedding nel database vettoriale associato al workspace. Ovviamente le evidence devono esssere distinte per workspace. + +## L'autenticazione + +L'applicazione deve prevedere la possibilità di collegarsi via http ad un Identity Manager. Nel MVP deve essere impostata l'autenticazione via Athentik, il quale a sua volta si interfaccia con il sistema di autenticazione del Policlinico San Donato basato su LDAP. Però deve essere anche prevista la possibilità di autenticarsi con un Entra ID. Per cui il sistema deve prevedere la possibilità di collegarsi a più Identity Manager, sostanzialmente tutti OIDC, ma diversi tra loro. Deve però poter operare anche senza autenticazione, sia per facilitare i test e lo sviluppo, sia come condizione potenziale di configurazione anche a sistema sviluppato e ready-for-production \ No newline at end of file