1698 lines
78 KiB
Markdown
1698 lines
78 KiB
Markdown
# ThothII — Harness Implementation Plan
|
||
|
||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||
|
||
**Goal:** Build `harness/` — the self-contained Pi layer (CLI `nsp` Python + gate extension JS + skills markdown + `.pi/`) that runs the 8-phase NL→SQL workflow emitting/consuming widget-descriptor JSON, derived from ChironeWp3 as a validated-and-adapted starting point (not assumed reliable).
|
||
|
||
**Architecture:** Port the relevant ChironeWp3 files into `harness/` task-by-task, then adapt each to the ThothII contract: `workflow.yaml` as the single source of workflow truth (F2), an "effective decisions as of pointer" view + `teardown_to_phase` for correct rollback (D15), a per-step task-document generator with enforced bounds for a 35B/<200k model (D16), the dual vector API key + `memory save-one` (D11), free-text interpretation (D13), value/formula grounding (D14), and the gate rewritten to emit/consume the 6-widget descriptor taxonomy (D2/D4).
|
||
|
||
**Tech Stack:** Python ≥3.12 (typer, pydantic v2, pyyaml, sqlalchemy 2.0, psycopg2-binary, datasketch, sqlglot, pytest, testcontainers[postgres]); Node/JS for the Pi gate extension; Ollama embeddings (nomic-embed-text-v2-moe); Pi + GLM 5.2 for L2 tests. Testing: two levels (L1 fake + L2 real), see Testing Strategy below.
|
||
|
||
**Reference spec:** `docs/superpowers/specs/2026-06-25-thothii-architecture-design.md` (all D1–D16 and §4.*).
|
||
|
||
**Source of ported code:** `/Users/mp/projects/ThothII/ChironeWp3/` (read-only reference; never modify it).
|
||
|
||
---
|
||
|
||
## Testing Strategy — three levels (L0 testcontainers, L1 fake, L2 real)
|
||
|
||
This plan uses a **three-level test strategy**. The split is mandatory because the skill→LLM→tool-call loop cannot be exercised without a real model, and the DB-touching ported code needs a real database to validate (the "not assumed reliable" principle is hollow without it).
|
||
|
||
### ⚠️ The honest headline: the core of the system has NO automated regression coverage
|
||
|
||
The **skill→LLM→gate loop** — the heart of the system — is covered **only by L2** (manual, non-deterministic, slow, requires credentials+VPN). **There is no automated regression protection on it.** This is a deliberate, conscious choice, not an accidental side-effect: the LLM is non-deterministic and requires a configured Pi + network, so it cannot live in the fast automated loop. **Consequence: agentic-behavior regressions surface at pre-release L2 runs, not at commit.** Accept this and plan L2 runs before any release. A partial automated net for this gap would be a **fake-Pi** that mocks the Pi runtime and lets the gate glue run in CI; that is documented as a **cross-cutting follow-up** (built alongside the backend plan, not in this plan).
|
||
|
||
### L0 — testcontainers, local Postgres, runs on every local run (CI)
|
||
|
||
**What:** integrity tests of ported DB-touching modules against a real Postgres in a Docker container (testcontainers). No LLM, no remote network. Docker is available on the dev machine, so L0 is always-on locally.
|
||
|
||
**Dependencies:** Docker (present in the dev env). No credentials, no VPN.
|
||
|
||
**Coverage:** the ported logic that speaks to a DB — `db/connection` (read-only enforcement, exit code 2 if writable), `db/introspect`, `db/sampling`, `db/execute`, `search/RRF` + LSH with real data, `vectorstore/store` direct-load path. This is where "not assumed reliable" gains real teeth for the data layer.
|
||
|
||
### L1 — fake data, deterministic, runs on every local run (CI)
|
||
|
||
**What:** logic-pure tests with fake data (`tmp_path`, fixtures, mocks). No DB, no LLM, no network.
|
||
|
||
**Coverage (honest):**
|
||
- Python logic-pure: `workflow.yaml` loading, `effective_decisions`, `teardown_to_phase`, `generate_task_doc`, `aggregate_lsh_multi` (on fake hits), formula store read/write, `decision_retracted`, CLI contract tests (valid input → well-formed JSON).
|
||
- **Gate builder functions (pure, in JS, tested in JS)**: the widget-descriptor builders produce the correct JSON given params. Tested in-language (node:test/vitest), no Python↔JS bridge, no Python mirror.
|
||
|
||
**Honest limitation (load-bearing):** L1 can test the gate **builders** (pure functions) but **NOT the gate glue** — registration, emission via `ctx.sendRaw`, the no-limbo loop, the anti-bypass hooks. The glue depends on the Pi runtime (`pi.registerTool`, `ctx.sendRaw`, `emitAndWait`) and cannot run without either a real Pi or a fake-Pi mock. **The glue is tested only at L2.** Fuzzy/negative tests on the glue (malformed tool params, out-of-order calls, orphan responses) likewise live at L2; only fuzzy tests on the pure builders live in L1.
|
||
|
||
### L2 — real LLM + real remote DB, manual/pre-release, NOT automated
|
||
|
||
**What:** end-to-end sessions with GLM 5.2 + the real Chirone datawarehouse and pgvector, reached via REST over VPN. Plus the gate-glue validation (the part L1 cannot reach).
|
||
|
||
**Dependencies (all required, skip if missing):**
|
||
- LLM: Pi configured locally with GLM 5.2 (already ready).
|
||
- DB: the remote Supabase endpoints (see "L2 connection params" below), reachable via VPN.
|
||
- `harness/.env` populated with the API keys + CA path.
|
||
|
||
**Modes:**
|
||
- **Human-in-the-loop (default):** the test runs the harness against GLM 5.2 and the reviewer answers each gate widget via the terminal, prompted by the test. Used for exploratory validation and for exercising the gate glue (which L1 cannot).
|
||
- **Pre-scripted answers (opt-in):** for specific cases (free-text "Altro", value grounding "ablazione", rollback), the test supplies canned reviewer answers automatically and asserts the outcome — no human typing. Used to test exact behaviors deterministically with the real model.
|
||
|
||
**Coverage (honest):** validates the assumption L1 cannot — that GLM 5.2 produces tool calls the gate accepts, that the skill's prompts lead to the expected interaction shape, that the gate glue handles real tool-call sequences (incl. Altro/Rifiuta/rollback), that value grounding and formula approval surface correctly on the real schema, that `memory save-one` upserts to the real pgvector. Closes the skill→LLM→gate loop AND exercises the gate glue.
|
||
|
||
**Honest limitation:** L2 is non-deterministic (the model may behave differently across runs) and slow/costly. It is a pre-release safety net, not a regression gate. A small, curated set of scenarios.
|
||
|
||
### L2 connection params (from the operator)
|
||
|
||
- **DWH (read-only):** `https://supabase-aritmolab.policlinicosandonato.it/dwh/` — PostgREST, schema `datawarehouse` + `datawarehouse_marts`, role `dwh_reader`. Header `X-API-Key: ${THOTH_DWH_API_KEY}`.
|
||
- **pgvector reader:** `https://supabase-aritmolab.policlinicosandonato.it/vector/v1/` — role `vector_reader` (`search_similar`, `list_tables`). Header `X-API-Key: ${THOTH_VEC_API_KEY}` (same value as `THOTH_DWH_API_KEY` — shared reader key).
|
||
- **pgvector writer (upsert-only):** same URL as reader, different key → role `vector_writer` (`existing_vector_hashes`, `upsert_vector_records`, no DELETE). Header `X-API-Key: ${THOTH_VEC_WRITE_API_KEY}`.
|
||
- **TLS:** self-signed, CA = leaf. On the operator's Mac, a local copy of the CA bundle must be present; `${THOTH_SSL_CA}` points to its local path. (Server path `/etc/nginx/ssl/policlinicosandonato.it.fullchain.crt` is unreachable from outside.)
|
||
- **Workspace:** `chirone-test`.
|
||
- **Test question (L2):** "dammi la lista dei pazienti che hanno fatto un'ablazione nel 2025" — exercises D14 value grounding + formula on the "ablazione" multi-column case.
|
||
|
||
### API key handling (security)
|
||
|
||
- All keys live **only** in `harness/.env` (gitignored). Never in code, never in the plan, never committed.
|
||
- `.env.example` is committed with variable names and empty values.
|
||
- L2 tests load `.env` via `python-dotenv` at startup; if any required var is missing/empty, the L2 test is **skipped with a clear message** (not failed), so L0/L1 never break for missing credentials.
|
||
- Tests never log key values. URLs in logs are fine; secrets are masked.
|
||
- ⚠️ The operator should rotate the key that appeared in chat transcripts.
|
||
|
||
### Test marker convention
|
||
|
||
- **L0 tests:** marker `@pytest.mark.l0`. Run always locally (Docker present). Auto-skip if Docker unavailable (rare). File naming: `tests/l0/test_*.py`.
|
||
- **L1 tests:** no marker (default, run always). File naming: `test_*.py` under `tests/`. Gate builder tests live in JS (node:test/vitest) under `harness/.pi/extensions/gate/__tests__/` and run via `npm test` alongside pytest.
|
||
- **L2 tests:** marker `@pytest.mark.l2`. Skipped automatically when `.env` incomplete. File naming: `tests/l2/test_*.py`. Default run: `pytest -m "not l2"` (L0+L1 only); pre-release: `pytest -m l2` (L2 only).
|
||
|
||
---
|
||
|
||
## File Structure
|
||
|
||
```
|
||
harness/
|
||
├── .pi/ ← Pi project config (ported + adapted)
|
||
│ ├── settings.json
|
||
│ ├── prompts/ ← /nuova-domanda, /riprendi-sessione
|
||
│ ├── skills/nsp-sessione/ ← SKILL.md + sub-files (adapted: task-doc, D13/D14)
|
||
│ └── extensions/
|
||
│ ├── nsp-gate.js ← REWRITTEN: 6-widget descriptor emit/consume (glue, verified at L2)
|
||
│ ├── reserved-labels.mjs ← ported
|
||
│ └── gate/
|
||
│ ├── builders.js ← NEW: pure widget-builder functions (L1, JS-tested)
|
||
│ └── __tests__/ ← node:test golden + fuzzy (L1)
|
||
│ ├── builders.test.js
|
||
│ └── golden/
|
||
├── nsp/ ← Python package (ported + adapted)
|
||
│ ├── __init__.py
|
||
│ ├── config.py ← ported + vector_write_rest (D11)
|
||
│ ├── workspace.py ← NEW: YAML load + ${VAR} expand (D3 confine)
|
||
│ ├── workflow.py ← NEW: reads workflow.yaml (F2)
|
||
│ ├── phase.py ← REWRITTEN: data-driven + effective_decisions (D15)
|
||
│ ├── teardown.py ← NEW: teardown_to_phase (D15)
|
||
│ ├── taskdoc.py ← NEW: per-step task document generator (D16)
|
||
│ ├── decisions.py ← ported + decision_retracted (D15)
|
||
│ ├── cli/
|
||
│ │ ├── __init__.py ← Typer app `nsp`
|
||
│ │ ├── _guards.py ← ported (require_vector_write_allowed etc.)
|
||
│ │ ├── workspace_cmd.py ← NEW: nsp workspace list/show
|
||
│ │ ├── phase_cmd.py ← adapted: + phase meta --json (F2)
|
||
│ │ ├── decision_cmd.py ← adapted: + retraction (D15)
|
||
│ │ ├── memory_cmd.py ← adapted: + save-one (D11)
|
||
│ │ ├── search_cmd.py ← adapted: + --kind formula (D14b)
|
||
│ │ ├── session_cmd.py ← adapted: + consistency check (D15)
|
||
│ │ └── (sql_cmd, cte_cmd, db_cmd, lsh_cmd, vector_cmd, evidence_cmd, schema_cmd — ported)
|
||
│ ├── db/ ← ported (connection, introspect, sampling, execute)
|
||
│ ├── rest/ ← ported (client, execute)
|
||
│ ├── search/ ← ported (combined_search, RRF) + formula retrieval (D14b)
|
||
│ ├── mschema/ ← ported (models, render, eligibility)
|
||
│ ├── vectorstore/ ← ported + dual-key rest_client (D11)
|
||
│ ├── evidence/ ← ported + formula kind (D14b)
|
||
│ ├── memory.py ← ported
|
||
│ └── session/ ← ported (models, store, artifacts)
|
||
├── workflow.yaml ← NEW: single source of workflow truth (F2)
|
||
├── workspaces/ ← workspace YAML definitions (D3)
|
||
│ └── chirone.example.yaml
|
||
├── sessions/ ← session persistence (FS, append-only ledger)
|
||
├── tests/
|
||
│ ├── conftest.py ← L2 skip-when-no-.env fixtures
|
||
│ ├── golden/ ← golden JSON for widget builders (L1)
|
||
│ ├── fixtures/
|
||
│ │ └── scenarios/ ← static .jsonl fixtures (synthesized once by build-scenarios)
|
||
│ ├── test_workflow.py ← L1
|
||
│ ├── test_phase_effective.py ← L1
|
||
│ ├── test_teardown.py ← L1
|
||
│ ├── test_taskdoc.py ← L1
|
||
│ ├── test_workspace.py ← L1
|
||
│ ├── test_vector_dual_key.py ← L1
|
||
│ ├── test_memory_save_one.py ← L1 (mocked writer)
|
||
│ ├── test_freetext_interpretation.py ← L1 (rationale contract)
|
||
│ ├── test_db_connection.py ← L0 Livello B (testcontainers: read-only enforcement)
|
||
│ ├── test_rest_client.py ← L1 Livello B (mock transport: RPC call)
|
||
│ ├── test_mschema_render.py ← L1 Livello B (3 formats on fake mschema)
|
||
│ ├── test_mschema_eligibility.py ← L1 Livello B (eligibility rules on fake columns)
|
||
│ ├── test_db_introspect.py ← L0 Livello B (testcontainers: known schema)
|
||
│ ├── test_db_sampling.py ← L0 Livello B (testcontainers: known data)
|
||
│ ├── test_rrf.py ← L0 Livello B (testcontainers: RRF fusion with real data)
|
||
│ ├── test_cli_contract.py ← L1 Livello B (each CLI: valid input → well-formed JSON)
|
||
│ ├── l0/ ← testcontainers tests (need Docker)
|
||
│ │ └── __init__.py
|
||
│ └── l2/ ← L2 (real GLM 5.2 + remote DB, @pytest.mark.l2)
|
||
│ └── __init__.py
|
||
│ └── l2/
|
||
│ ├── test_session_ablazione.py ← L2 (GLM 5.2 + real DWH, human or scripted)
|
||
│ ├── test_value_grounding_real.py ← L2 ("ablazione" on real schema)
|
||
│ └── test_memory_save_one_real.py ← L2 (real pgvector writer upsert)
|
||
├── scripts/
|
||
│ ├── create_vector_writer_rpc.sql ← ported (writer RPC allowlist)
|
||
│ └── create_vector_reader_rpc.sql ← NEW (reader RPC, D11 §5.4 note)
|
||
├── pyproject.toml
|
||
└── .env.example
|
||
```
|
||
|
||
---
|
||
|
||
## Phase A — Foundation (cross-cutting infrastructure first)
|
||
|
||
Rationale: every feature depends on these. `workflow.yaml` + `effective_decisions` + `teardown` + `taskdoc` are the substrate the gate and the D11/D13/D14 features build on.
|
||
|
||
### Task A1: Scaffold the harness project
|
||
|
||
**Files:**
|
||
- Create: `harness/pyproject.toml`
|
||
- Create: `harness/.env.example`
|
||
- Create: `harness/nsp/__init__.py`
|
||
- Create: `harness/.gitignore`
|
||
|
||
- [ ] **Step 1: Create `pyproject.toml`**
|
||
|
||
```toml
|
||
[project]
|
||
name = "nsp"
|
||
version = "0.1.0"
|
||
description = "ThothII harness — deterministic CLI for the NL→SQL workflow"
|
||
requires-python = ">=3.12"
|
||
dependencies = [
|
||
"typer>=0.12",
|
||
"pydantic>=2.0",
|
||
"pyyaml>=6.0",
|
||
"sqlalchemy>=2.0",
|
||
"psycopg2-binary>=2.9",
|
||
"datasketch>=1.6",
|
||
"sqlglot>=23.0",
|
||
"python-dotenv>=1.0",
|
||
"rich>=13.0",
|
||
"requests>=2.31",
|
||
"tqdm>=4.66",
|
||
]
|
||
|
||
[project.scripts]
|
||
nsp = "nsp.cli:app"
|
||
|
||
[project.optional-dependencies]
|
||
dev = [
|
||
"pytest>=8.0",
|
||
"testcontainers[postgres]>=4.0",
|
||
"ruff>=0.5",
|
||
]
|
||
|
||
[tool.ruff]
|
||
line-length = 100
|
||
|
||
[tool.pytest.ini_options]
|
||
testpaths = ["tests"]
|
||
```
|
||
|
||
- [ ] **Step 2: Create `.env.example`** (port the structure from `ChironeWp3/.env.example`, renamed to `THOTH_*` per §5.4)
|
||
|
||
```bash
|
||
# ThothII profile: server (full rebuild) | workstation (read REST + optional upsert)
|
||
THOTH_PROFILE=server
|
||
|
||
# Workspace DB (relational, direct transport) — example: chirone
|
||
THOTH_DB_HOST=
|
||
THOTH_DB_PORT=5432
|
||
THOTH_DB_USER=
|
||
THOTH_DB_PASSWORD=
|
||
|
||
# Vector REST — LETTURA (rpc search_similar) sul Supabase remoto
|
||
THOTH_VEC_REST_URL=https://host/vector/v1/
|
||
THOTH_VEC_API_KEY=
|
||
|
||
# Vector REST — SCRITTURA controllata (upsert only) da client remoti autorizzati
|
||
THOTH_VEC_WRITE_API_KEY=
|
||
|
||
# Vector direct loading (server-only)
|
||
THOTH_VEC_HOST=localhost
|
||
THOTH_VEC_PORT=5438
|
||
THOTH_VEC_USER=postgres
|
||
THOTH_VEC_PASSWORD=
|
||
|
||
# Embeddings (Ollama)
|
||
THOTH_OLLAMA_URL=http://localhost:11434
|
||
|
||
# TLS (optional, internal CA)
|
||
THOTH_SSL_CA=
|
||
```
|
||
|
||
- [ ] **Step 3: Create `.gitignore`**
|
||
|
||
```
|
||
__pycache__/
|
||
*.pyc
|
||
.env
|
||
.venv/
|
||
sessions/
|
||
indexes/
|
||
*.egg-info/
|
||
dist/
|
||
build/
|
||
```
|
||
|
||
- [ ] **Step 4: Create empty `nsp/__init__.py` and install**
|
||
|
||
```bash
|
||
cd /Users/mp/projects/ThothII/harness
|
||
python -m venv .venv && source .venv/bin/activate
|
||
pip install -e ".[dev]"
|
||
```
|
||
|
||
- [ ] **Step 5: Verify `nsp` command is not yet wired (expected: import error) then commit**
|
||
|
||
```bash
|
||
nsp --help 2>&1 | head -3
|
||
```
|
||
Expected: error (no `cli` module yet) — this is fine, we add it in Task A4.
|
||
|
||
```bash
|
||
git add harness/
|
||
git commit -m "feat(harness): scaffold project (pyproject, env, gitignore)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task A2: Port config + create `workspace.py` (D3 confine)
|
||
|
||
**Files:**
|
||
- Create: `harness/nsp/config.py` (ported + adapted from `ChironeWp3/src/psdwp3/config.py`)
|
||
- Create: `harness/nsp/workspace.py`
|
||
- Create: `harness/workspaces/chirone.example.yaml`
|
||
- Test: `harness/tests/test_workspace.py`
|
||
|
||
- [ ] **Step 1: Port `config.py`** — copy `ChironeWp3/src/psdwp3/config.py` → `harness/nsp/config.py`. Rename the package prefix `psdwp3` → `nsp` everywhere. Verify it still defines: `DatabaseConfig`, `RestConfig`, `PathsConfig`, `EligibilityConfig`, `LshConfig`, `EvidenceSourcesConfig`, `EmbeddingsConfig`, `VectorConfig`, `SearchConfig`, `ExecutionConfig`, and the `Config` class with `vector_rest` + `vector_write_rest` (both `RestConfig | None`, see ChironeWp3 config.py:154-159) + `_expand_env` (lines 16-32).
|
||
|
||
- [ ] **Step 2: Create `workspaces/chirone.example.yaml`** per spec §5.1 (relational + vector_db with rest/write_rest + evidence + embeddings + execution).
|
||
|
||
- [ ] **Step 3: Write failing test for workspace loading**
|
||
|
||
```python
|
||
# tests/test_workspace.py
|
||
from pathlib import Path
|
||
from nsp.workspace import load_workspace, WorkspaceError
|
||
|
||
def test_load_workspace_expands_env_vars(monkeypatch, tmp_path):
|
||
monkeypatch.setenv("THOTH_VEC_API_KEY", "secret-reader")
|
||
monkeypatch.setenv("THOTH_VEC_WRITE_API_KEY", "secret-writer")
|
||
monkeypatch.setenv("THOTH_VEC_REST_URL", "https://example/vector/v1/")
|
||
yaml = tmp_path / "w.yaml"
|
||
yaml.write_text(
|
||
"name: test\n"
|
||
"relational:\n db_type: postgres\n transport: rest\n"
|
||
" rest: { base_url: 'https://dwh/', api_key: '${THOTH_VEC_API_KEY}' }\n"
|
||
"vector_db:\n collection: test_docs\n dim: 768\n"
|
||
" rest: { base_url: '${THOTH_VEC_REST_URL}', api_key: '${THOTH_VEC_API_KEY}' }\n"
|
||
" write_rest: { base_url: '${THOTH_VEC_REST_URL}', api_key: '${THOTH_VEC_WRITE_API_KEY}' }\n"
|
||
"embeddings:\n provider: ollama\n base_url: 'http://localhost:11434'\n model: m\n dim: 768\n"
|
||
)
|
||
ws = load_workspace(yaml)
|
||
assert ws.vector_db.rest.api_key == "secret-reader"
|
||
assert ws.vector_db.write_rest.api_key == "secret-writer"
|
||
|
||
def test_load_workspace_missing_env_raises(monkeypatch, tmp_path):
|
||
monkeypatch.delenv("THOTH_VEC_API_KEY", raising=False)
|
||
yaml = tmp_path / "w.yaml"
|
||
yaml.write_text(
|
||
"name: test\nrelational: { db_type: postgres, transport: rest, rest: { base_url: 'https://dwh/', api_key: '${THOTH_VEC_API_KEY}' } }\n"
|
||
"vector_db: { collection: c, dim: 768, rest: { base_url: 'https://v/', api_key: '${THOTH_VEC_API_KEY}' } }\n"
|
||
"embeddings: { provider: ollama, base_url: 'http://x', model: m, dim: 768 }\n"
|
||
)
|
||
try:
|
||
load_workspace(yaml)
|
||
assert False, "should have raised"
|
||
except WorkspaceError as e:
|
||
assert "THOTH_VEC_API_KEY" in str(e)
|
||
```
|
||
|
||
- [ ] **Step 4: Run test, verify fail**
|
||
|
||
Run: `cd harness && pytest tests/test_workspace.py -v`
|
||
Expected: FAIL — `ModuleNotFoundError: nsp.workspace`
|
||
|
||
- [ ] **Step 5: Implement `workspace.py`**
|
||
|
||
```python
|
||
# nsp/workspace.py
|
||
"""Workspace YAML loading — the single boundary for workspace configuration (spec D3).
|
||
Reads workspaces/<name>.yaml, expands ${VAR} from env, validates via the Config model.
|
||
Future migration to a DB store would replace only this module.
|
||
"""
|
||
from __future__ import annotations
|
||
import os
|
||
from pathlib import Path
|
||
import yaml
|
||
from nsp.config import Config, ConfigError
|
||
|
||
class WorkspaceError(Exception):
|
||
pass
|
||
|
||
def _expand_str(s: str) -> str:
|
||
"""Expand ${VAR} occurrences in s. Raises if a referenced var is unset."""
|
||
PREFIX = "${"
|
||
SUFFIX = "}"
|
||
out: list[str] = []
|
||
i = 0
|
||
while i < len(s):
|
||
start = s.find(PREFIX, i)
|
||
if start == -1:
|
||
out.append(s[i:])
|
||
break
|
||
out.append(s[i:start])
|
||
end = s.find(SUFFIX, start + len(PREFIX))
|
||
if end == -1:
|
||
raise WorkspaceError(f'Sintassi non valida (manca "}}"): {s[start:]}')
|
||
var = s[start + len(PREFIX) : end]
|
||
if var not in os.environ:
|
||
raise WorkspaceError(f"Variabile d'ambiente non definita: {var}")
|
||
out.append(os.environ[var])
|
||
i = end + len(SUFFIX)
|
||
return "".join(out)
|
||
|
||
def _expand_env(obj):
|
||
if isinstance(obj, str):
|
||
return _expand_str(obj)
|
||
if isinstance(obj, dict):
|
||
return {k: _expand_env(v) for k, v in obj.items()}
|
||
if isinstance(obj, list):
|
||
return [_expand_env(v) for v in obj]
|
||
return obj
|
||
|
||
def load_workspace(path: str | Path) -> Config:
|
||
raw = yaml.safe_load(Path(path).read_text())
|
||
try:
|
||
return Config.model_validate(_expand_env(raw))
|
||
except ConfigError as e:
|
||
raise WorkspaceError(str(e)) from e
|
||
```
|
||
|
||
- [ ] **Step 6: Run test, verify pass**
|
||
|
||
Run: `pytest tests/test_workspace.py -v`
|
||
Expected: 2 PASS
|
||
|
||
- [ ] **Step 7: Commit**
|
||
|
||
```bash
|
||
git add harness/nsp/config.py harness/nsp/workspace.py harness/workspaces/chirone.example.yaml harness/tests/test_workspace.py
|
||
git commit -m "feat(harness): port config + workspace.py YAML loader (D3)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task A3: Create `workflow.yaml` + `workflow.py` (F2) — single source of truth
|
||
|
||
**Files:**
|
||
- Create: `harness/workflow.yaml`
|
||
- Create: `harness/nsp/workflow.py`
|
||
- Test: `harness/tests/test_workflow.py`
|
||
|
||
- [ ] **Step 1: Create `workflow.yaml`** (the 8-phase definition from spec §5.3)
|
||
|
||
```yaml
|
||
# harness/workflow.yaml — single source of workflow truth (spec F2)
|
||
schema_version: 1
|
||
|
||
phases:
|
||
- id: F1
|
||
name: chiaramento
|
||
advance: kind:phase
|
||
prerequisites: []
|
||
artifacts_out: []
|
||
- id: F2
|
||
name: memoria
|
||
advance: auto_if_empty
|
||
prerequisites: []
|
||
artifacts_out: []
|
||
- id: F3
|
||
name: riscrittura
|
||
advance: kind:phase
|
||
prerequisites:
|
||
- decision_exists: question_rewritten
|
||
artifacts_out: [question.md]
|
||
- id: F4
|
||
name: schema_linking
|
||
advance: reviewer_decide
|
||
prerequisites: []
|
||
artifacts_out: [schema_linking.json]
|
||
- id: F5
|
||
name: sintesi
|
||
advance: kind:phase
|
||
prerequisites:
|
||
- file_validates: [schema_linking.json, SchemaLinking]
|
||
artifacts_out: []
|
||
- id: F6
|
||
name: cte
|
||
advance: auto_if_empty_or_skipped
|
||
prerequisites:
|
||
- any:
|
||
- decision_subject_exists: [phase_skipped, "phase:6"]
|
||
- all_ctes_approved: true
|
||
artifacts_out: [cte_plan.json, ctes/, cte_tests.json]
|
||
- id: F7
|
||
name: sql_finale
|
||
advance: kind:phase
|
||
prerequisites:
|
||
- decision_exists: sql_approved
|
||
artifacts_out: [sql_final.sql]
|
||
- id: F8
|
||
name: datamart
|
||
advance: reviewer_decide
|
||
prerequisites:
|
||
- any:
|
||
- decision_exists: datamart_requested
|
||
- decision_exists: datamart_declined
|
||
artifacts_out: []
|
||
|
||
decision_min_phase: auto
|
||
max_phase: auto
|
||
```
|
||
|
||
- [ ] **Step 2: Write failing test for workflow loading**
|
||
|
||
```python
|
||
# tests/test_workflow.py
|
||
from nsp.workflow import load_workflow, PhaseSpec
|
||
|
||
def test_workflow_loads_8_phases():
|
||
wf = load_workflow()
|
||
assert len(wf.phases) == 8
|
||
assert wf.phases[0].id == "F1"
|
||
assert wf.max_phase == 8
|
||
assert wf.phase_by_num(1).name == "chiaramento"
|
||
|
||
def test_decision_min_phase_derived():
|
||
wf = load_workflow()
|
||
# question_rewritten is a prerequisite of F3 → min phase 3
|
||
assert wf.decision_min_phase("question_rewritten") == 3
|
||
# sql_approved is a prerequisite of F7 → min phase 7
|
||
assert wf.decision_min_phase("sql_approved") == 7
|
||
# unknown type → phase 1 (default)
|
||
assert wf.decision_min_phase("nonexistent_type") == 1
|
||
|
||
def test_phase_name_lookup():
|
||
wf = load_workflow()
|
||
assert wf.phase_name(6) == "cte"
|
||
assert wf.phase_name(8) == "datamart" # the JS drift bug — F8 must be present
|
||
|
||
def test_artifacts_out_per_phase():
|
||
wf = load_workflow()
|
||
assert "schema_linking.json" in wf.phase_by_num(4).artifacts_out
|
||
assert "sql_final.sql" in wf.phase_by_num(7).artifacts_out
|
||
```
|
||
|
||
- [ ] **Step 3: Run test, verify fail**
|
||
|
||
Run: `pytest tests/test_workflow.py -v`
|
||
Expected: FAIL — `ModuleNotFoundError: nsp.workflow`
|
||
|
||
- [ ] **Step 4: Implement `workflow.py`**
|
||
|
||
```python
|
||
# nsp/workflow.py
|
||
"""Reads workflow.yaml — the SINGLE source of workflow truth (spec F2, §5.3).
|
||
phase.py, the gate, and the skill all read from here. No more duplicated constants.
|
||
"""
|
||
from __future__ import annotations
|
||
from dataclasses import dataclass, field
|
||
from pathlib import Path
|
||
import yaml
|
||
|
||
_WF_PATH = Path(__file__).resolve().parent.parent / "workflow.yaml"
|
||
|
||
@dataclass
|
||
class PhaseSpec:
|
||
id: str
|
||
num: int
|
||
name: str
|
||
advance: str
|
||
prerequisites: list
|
||
artifacts_out: list[str] = field(default_factory=list)
|
||
|
||
@dataclass
|
||
class Workflow:
|
||
schema_version: int
|
||
phases: list[PhaseSpec]
|
||
_decision_min_map: dict[str, int] = field(default_factory=dict)
|
||
|
||
@property
|
||
def max_phase(self) -> int:
|
||
return len(self.phases)
|
||
|
||
def phase_by_num(self, n: int) -> PhaseSpec:
|
||
return self.phases[n - 1]
|
||
|
||
def phase_name(self, n: int) -> str:
|
||
return self.phase_by_num(n).name if 1 <= n <= self.max_phase else "?"
|
||
|
||
def decision_min_phase(self, decision_type: str) -> int:
|
||
# A decision type's min phase = the earliest phase whose prerequisites
|
||
# reference it (via decision_exists/decision_subject_exists), else 1.
|
||
return self._decision_min_map.get(decision_type, 1)
|
||
|
||
def _collect_decision_mins(phases: list[PhaseSpec]) -> dict[str, int]:
|
||
"""Scan prerequisites for decision_exists / decision_subject_exists mentions."""
|
||
mins: dict[str, int] = {}
|
||
def scan(node, phase_num: int):
|
||
if isinstance(node, dict):
|
||
for k, v in node.items():
|
||
if k in ("decision_exists", "decision_subject_exists"):
|
||
if isinstance(v, list):
|
||
dtype = v[0]
|
||
else:
|
||
dtype = v
|
||
if dtype not in mins or phase_num < mins[dtype]:
|
||
mins[dtype] = phase_num
|
||
else:
|
||
scan(v, phase_num)
|
||
elif isinstance(node, list):
|
||
for item in node:
|
||
scan(item, phase_num)
|
||
for p in phases:
|
||
scan(p.prerequisites, p.num)
|
||
return mins
|
||
|
||
def load_workflow(path: Path | str = _WF_PATH) -> Workflow:
|
||
raw = yaml.safe_load(Path(path).read_text())
|
||
phases = []
|
||
for i, p in enumerate(raw["phases"], start=1):
|
||
phases.append(PhaseSpec(
|
||
id=p["id"], num=i, name=p["name"],
|
||
advance=p["advance"], prerequisites=p.get("prerequisites", []),
|
||
artifacts_out=p.get("artifacts_out", []),
|
||
))
|
||
return Workflow(
|
||
schema_version=raw.get("schema_version", 1),
|
||
phases=phases,
|
||
_decision_min_map=_collect_decision_mins(phases),
|
||
)
|
||
```
|
||
|
||
- [ ] **Step 5: Run test, verify pass**
|
||
|
||
Run: `pytest tests/test_workflow.py -v`
|
||
Expected: 4 PASS
|
||
|
||
- [ ] **Step 6: Commit**
|
||
|
||
```bash
|
||
git add harness/workflow.yaml harness/nsp/workflow.py harness/tests/test_workflow.py
|
||
git commit -m "feat(harness): workflow.yaml as single source of truth + workflow.py loader (F2)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task A4: Port `decisions.py` + add `decision_retracted` (D15)
|
||
|
||
**Files:**
|
||
- Create: `harness/nsp/decisions.py` (ported + adapted)
|
||
- Test: `harness/tests/test_decisions_retract.py`
|
||
|
||
- [ ] **Step 1: Port `decisions.py`** — copy `ChironeWp3/src/psdwp3/session/decisions.py`. Rename package. Add `decision_retracted` to the `DecisionType` Literal. The `DecisionRecord` model gains an optional `retracts: int | None` field (the `decision_seq` being retracted, for step-level rollback).
|
||
|
||
- [ ] **Step 2: Write failing test for retraction**
|
||
|
||
```python
|
||
# tests/test_decisions_retract.py
|
||
from pathlib import Path
|
||
from nsp.decisions import append_decision, list_decisions, DecisionRecord
|
||
|
||
def test_retracted_decision_in_audit_but_marked(tmp_path):
|
||
session = tmp_path / "s1"
|
||
session.mkdir()
|
||
append_decision(session, DecisionRecord(seq=1, ts="t", type="table_promoted",
|
||
subject="phase:4", detail="t1", rationale="r"))
|
||
append_decision(session, DecisionRecord(seq=2, ts="t", type="decision_retracted",
|
||
subject="phase:4", detail="retract", rationale="wrong", retracts=1))
|
||
all_decisions = list_decisions(session)
|
||
assert len(all_decisions) == 2 # both in audit
|
||
assert all_decisions[1].retracts == 1
|
||
```
|
||
|
||
- [ ] **Step 3: Run, verify fail, implement the `retracts` field + ensure `append_decision` writes it, verify pass.** (The ported `append_decision` already writes all fields; just ensure the new field is serialized. If using pydantic `model_dump`, add `retracts: int | None = None`.)
|
||
|
||
Run: `pytest tests/test_decisions_retract.py -v`
|
||
Expected after implement: PASS
|
||
|
||
- [ ] **Step 4: Commit**
|
||
|
||
```bash
|
||
git add harness/nsp/decisions.py harness/tests/test_decisions_retract.py
|
||
git commit -m "feat(harness): port decisions.py + decision_retracted for step rollback (D15)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task A5: Rewrite `phase.py` — data-driven + `effective_decisions` (D15 core fix)
|
||
|
||
**Files:**
|
||
- Create: `harness/nsp/phase.py` (rewritten)
|
||
- Port: `harness/nsp/session/{models,store,artifacts}.py` (needed for SchemaLinking validation)
|
||
- Test: `harness/tests/test_phase_effective.py`
|
||
|
||
This is the single most important architectural fix (spec §4.8). `phase.py` reads from `workflow.yaml`, and ALL helpers consult `effective_decisions()` instead of raw `list_decisions()`.
|
||
|
||
- [ ] **Step 1: Port session models/store/artifacts** — copy `ChironeWp3/src/psdwp3/session/{models.py, store.py, artifacts.py}` → `harness/nsp/session/`. Rename package. These are needed for `SchemaLinking` validation in `advance_problems`.
|
||
|
||
- [ ] **Step 2: Write failing test for effective_decisions**
|
||
|
||
```python
|
||
# tests/test_phase_effective.py
|
||
from pathlib import Path
|
||
from nsp.phase import current_phase, effective_decisions
|
||
from nsp.decisions import append_decision, DecisionRecord
|
||
|
||
def _d(session, seq, dtype, subject, **kw):
|
||
append_decision(session, DecisionRecord(seq=seq, ts="t", type=dtype, subject=subject,
|
||
detail=kw.get("detail", ""), rationale=kw.get("rationale", ""),
|
||
retracts=kw.get("retracts")))
|
||
|
||
def test_effective_decisions_excludes_pre_reopen_tail(tmp_path):
|
||
s = tmp_path / "s"; s.mkdir()
|
||
_d(s, 1, "phase_approved", "phase:1")
|
||
_d(s, 2, "phase_approved", "phase:2")
|
||
_d(s, 3, "phase_reopened", "phase:1") # rollback to phase 1
|
||
_d(s, 4, "table_promoted", "phase:4") # stale: produced after reopen but for phase 4 (not current)
|
||
# After reopen to phase:1, the effective view truncates everything after the last phase_reopened.
|
||
eff = effective_decisions(s)
|
||
# The reopen at seq 3 means decisions after it (seq 4) are NOT effective,
|
||
# and the reopen itself sets the pointer. Effective = decisions up to and incl. reopen,
|
||
# then re-walked. The stale table_promoted at seq 4 must be excluded.
|
||
types = [d.type for d in eff]
|
||
assert "table_promoted" not in types
|
||
|
||
def test_current_phase_after_reopen(tmp_path):
|
||
s = tmp_path / "s"; s.mkdir()
|
||
_d(s, 1, "phase_approved", "phase:1")
|
||
_d(s, 2, "phase_approved", "phase:2")
|
||
_d(s, 3, "phase_reopened", "phase:1")
|
||
assert current_phase(s) == 1
|
||
_d(s, 4, "phase_approved", "phase:1") # re-approve after reopen
|
||
assert current_phase(s) == 2
|
||
|
||
def test_retracted_decision_excluded_from_effective(tmp_path):
|
||
s = tmp_path / "s"; s.mkdir()
|
||
_d(s, 1, "table_promoted", "phase:4", detail="t1")
|
||
_d(s, 2, "decision_retracted", "phase:4", retracts=1)
|
||
eff = effective_decisions(s)
|
||
types = [d.type for d in eff]
|
||
assert "table_promoted" not in types # retracted → excluded
|
||
```
|
||
|
||
- [ ] **Step 3: Run, verify fail**
|
||
|
||
Run: `pytest tests/test_phase_effective.py -v`
|
||
Expected: FAIL
|
||
|
||
- [ ] **Step 4: Implement `phase.py`** — the core. Port the fold logic from `ChironeWp3/src/psdwp3/session/phase.py` but:
|
||
- Read `max_phase`, `phase_name`, `decision_min_phase`, `advance_problems` rules from `workflow.yaml` via `load_workflow()`.
|
||
- Add `effective_decisions(session)`: replay the ledger; when hitting a `phase_reopened phase:N`, truncate everything after it in the effective view AND re-walk from phase N. When hitting a `decision_retracted`, exclude the retracted seq.
|
||
- Rewrite `advance_problems(phase)`, `approved_ctes`, `substantive_count_current_phase`, `auto_advance_eligible` to consult `effective_decisions()` instead of `list_decisions()`.
|
||
|
||
Reference skeleton (the `effective_decisions` core — the load-bearing part):
|
||
|
||
```python
|
||
# nsp/phase.py (key function)
|
||
from nsp.decisions import list_decisions
|
||
from nsp.workflow import load_workflow
|
||
|
||
def effective_decisions(session_dir) -> list:
|
||
"""The canonical effective view. ALL helpers MUST use this, not list_decisions().
|
||
Replays the ledger: phase_reopened truncates the tail; decision_retracted excludes the target.
|
||
"""
|
||
all_d = list_decisions(session_dir)
|
||
retracted_seqs = {d.retracts for d in all_d if d.type == "decision_retracted" and d.retracts}
|
||
# Find the LAST phase_reopened; everything strictly after it is stale.
|
||
last_reopen_idx = -1
|
||
for i, d in enumerate(all_d):
|
||
if d.type == "phase_reopened":
|
||
last_reopen_idx = i
|
||
if last_reopen_idx >= 0:
|
||
effective = all_d[: last_reopen_idx + 1]
|
||
else:
|
||
effective = list(all_d)
|
||
# Exclude retracted decisions (by seq), and the retraction records themselves.
|
||
return [d for d in effective if d.seq not in retracted_seqs and d.type != "decision_retracted"]
|
||
```
|
||
|
||
Implement `current_phase(session_dir)` as the fold over `effective_decisions(session_dir)` (not over raw list). Implement `advance_problems(phase, session_dir)` by evaluating the `prerequisites` list from `workflow.yaml` for that phase against `effective_decisions`. Reuse the prerequisite predicate types: `decision_exists`, `decision_subject_exists`, `file_validates`, `all_ctes_approved`, `any`, `all`.
|
||
|
||
- [ ] **Step 5: Run, verify pass**
|
||
|
||
Run: `pytest tests/test_phase_effective.py -v`
|
||
Expected: 3 PASS
|
||
|
||
- [ ] **Step 6: Commit**
|
||
|
||
```bash
|
||
git add harness/nsp/phase.py harness/nsp/session/ harness/tests/test_phase_effective.py
|
||
git commit -m "feat(harness): rewrite phase.py data-driven + effective_decisions (D15 core, F2)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task A6: Create `teardown.py` — `teardown_to_phase` (D15 teardown)
|
||
|
||
**Files:**
|
||
- Create: `harness/nsp/teardown.py`
|
||
- Test: `harness/tests/test_teardown.py`
|
||
|
||
- [ ] **Step 1: Write failing test**
|
||
|
||
```python
|
||
# tests/test_teardown.py
|
||
from pathlib import Path
|
||
from nsp.teardown import teardown_to_phase, TeardownReport
|
||
|
||
def test_teardown_to_phase_4_deletes_phase5plus_artifacts(tmp_path):
|
||
s = tmp_path / "sess"; s.mkdir()
|
||
(s / "schema_linking.json").write_text("{}") # F4 artifact
|
||
(s / "cte_plan.json").write_text("[]") # F6 artifact
|
||
(s / "ctes").mkdir()
|
||
(s / "ctes" / "x.sql").write_text("SELECT 1") # F6 artifact
|
||
(s / "sql_final.sql").write_text("SELECT 1") # F7 artifact
|
||
report = teardown_to_phase(s, target_phase=4)
|
||
assert (s / "schema_linking.json").exists() # F4 preserved (target is 4)
|
||
assert not (s / "cte_plan.json").exists() # F6 deleted (>4)
|
||
assert not (s / "ctes").exists() # F6 dir deleted
|
||
assert not (s / "sql_final.sql").exists() # F7 deleted
|
||
assert "cte_plan.json" in report.deleted_files
|
||
assert "sql_final.sql" in report.deleted_files
|
||
|
||
def test_teardown_records_deleted_in_ledger(tmp_path):
|
||
# teardown itself appends a phase_reopened decision (the rollback marker)
|
||
# — verifying that is covered by test_phase_effective; here we check the report.
|
||
s = tmp_path / "sess"; s.mkdir()
|
||
(s / "sql_final.sql").write_text("SELECT 1")
|
||
report = teardown_to_phase(s, target_phase=4)
|
||
assert len(report.deleted_files) >= 1
|
||
```
|
||
|
||
- [ ] **Step 2: Run, verify fail**
|
||
|
||
Run: `pytest tests/test_teardown.py -v`
|
||
|
||
- [ ] **Step 3: Implement `teardown.py`**
|
||
|
||
```python
|
||
# nsp/teardown.py
|
||
"""Artifact teardown on rollback (spec D15, §4.8).
|
||
Deletes every artifact whose producing phase > target, using the artifacts_out map
|
||
from workflow.yaml. Recomputes derived state. Fixes the orphaned-CTE-blocks-finalize bug.
|
||
"""
|
||
from __future__ import annotations
|
||
from dataclasses import dataclass, field
|
||
from pathlib import Path
|
||
from nsp.workflow import load_workflow
|
||
|
||
@dataclass
|
||
class TeardownReport:
|
||
target_phase: int
|
||
deleted_files: list[str] = field(default_factory=list)
|
||
|
||
def teardown_to_phase(session_dir: str | Path, target_phase: int) -> TeardownReport:
|
||
session_dir = Path(session_dir)
|
||
wf = load_workflow()
|
||
report = TeardownReport(target_phase=target_phase)
|
||
for phase in wf.phases:
|
||
if phase.num <= target_phase:
|
||
continue
|
||
for artifact in phase.artifacts_out:
|
||
target = session_dir / artifact.rstrip("/")
|
||
if artifact.endswith("/"):
|
||
# directory artifact (e.g. ctes/)
|
||
if target.exists():
|
||
for f in target.glob("*"):
|
||
f.unlink()
|
||
report.deleted_files.append(f.name)
|
||
target.rmdir()
|
||
else:
|
||
if target.exists():
|
||
target.unlink()
|
||
report.deleted_files.append(artifact)
|
||
return report
|
||
```
|
||
|
||
- [ ] **Step 4: Run, verify pass**
|
||
|
||
Run: `pytest tests/test_teardown.py -v`
|
||
Expected: 2 PASS
|
||
|
||
- [ ] **Step 5: Commit**
|
||
|
||
```bash
|
||
git add harness/nsp/teardown.py harness/tests/test_teardown.py
|
||
git commit -m "feat(harness): teardown_to_phase — artifact teardown on rollback (D15)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task A7: Create `taskdoc.py` — per-step task document generator (D16)
|
||
|
||
**Files:**
|
||
- Create: `harness/nsp/taskdoc.py`
|
||
- Test: `harness/tests/test_taskdoc.py`
|
||
|
||
- [ ] **Step 1: Write failing test**
|
||
|
||
```python
|
||
# tests/test_taskdoc.py
|
||
from pathlib import Path
|
||
from nsp.taskdoc import generate_task_doc, TaskDoc
|
||
|
||
def test_task_doc_includes_question_and_schema_scope_not_full_physical(tmp_path):
|
||
s = tmp_path / "sess"; s.mkdir()
|
||
(s / "question.md").write_text("# Domanda\nQuanti pazienti?\n## Assunzioni\n- a")
|
||
(s / "schema_linking.json").write_text('{"candidates":[{"kind":"table","name":"pazienti"}],"joins":[]}')
|
||
# A fatal-sized artifact that must NOT appear in the task doc
|
||
(s.parent / "physical.yaml").write_text("x: " + "y" * 800_000)
|
||
doc = generate_task_doc(session_dir=s, phase=7, promoted_tables=["pazienti"])
|
||
assert "Quanti pazienti?" in doc.body
|
||
assert "pazienti" in doc.body
|
||
assert len(doc.body) < 100_000 # bounded — no fatal full-schema read
|
||
assert "physical.yaml" not in doc.body # never embedded
|
||
|
||
def test_task_doc_byte_budget_enforced(tmp_path):
|
||
s = tmp_path / "sess"; s.mkdir()
|
||
(s / "question.md").write_text("q")
|
||
doc = generate_task_doc(session_dir=s, phase=1, promoted_tables=[])
|
||
assert doc.byte_budget_ok is True
|
||
```
|
||
|
||
- [ ] **Step 2: Run, verify fail**
|
||
|
||
Run: `pytest tests/test_taskdoc.py -v`
|
||
|
||
- [ ] **Step 3: Implement `taskdoc.py`**
|
||
|
||
```python
|
||
# nsp/taskdoc.py
|
||
"""Per-step task document generator (spec D16, §4.9).
|
||
Emits a single compact document per phase/step, derived from prior artifacts,
|
||
with enforced byte budget (target <20k tokens). NEVER embeds full physical.yaml.
|
||
"""
|
||
from __future__ import annotations
|
||
from dataclasses import dataclass
|
||
from pathlib import Path
|
||
|
||
MAX_BODY_BYTES = 80_000 # ~20k tokens
|
||
|
||
@dataclass
|
||
class TaskDoc:
|
||
phase: int
|
||
body: str
|
||
byte_budget_ok: bool
|
||
|
||
def generate_task_doc(session_dir: Path | str, phase: int, promoted_tables: list[str] | None = None) -> TaskDoc:
|
||
session_dir = Path(session_dir)
|
||
parts: list[str] = []
|
||
q = session_dir / "question.md"
|
||
if q.exists():
|
||
parts.append("## Domanda\n" + q.read_text())
|
||
sl = session_dir / "schema_linking.json"
|
||
if sl.exists() and phase >= 4:
|
||
parts.append("## Schema linking (deciso)\n```json\n" + sl.read_text() + "\n```")
|
||
# Task header for the phase
|
||
parts.append(f"## Task: fase {phase}")
|
||
body = "\n\n".join(parts)
|
||
return TaskDoc(phase=phase, body=body, byte_budget_ok=len(body.encode()) <= MAX_BODY_BYTES)
|
||
```
|
||
|
||
- [ ] **Step 4: Run, verify pass**
|
||
|
||
Run: `pytest tests/test_taskdoc.py -v`
|
||
Expected: 2 PASS
|
||
|
||
- [ ] **Step 5: Commit**
|
||
|
||
```bash
|
||
git add harness/nsp/taskdoc.py harness/tests/test_taskdoc.py
|
||
git commit -m "feat(harness): taskdoc per-step generator with byte budget (D16)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task A8: Port the CLI skeleton + `nsp phase meta --json` (F2)
|
||
|
||
**Files:**
|
||
- Create: `harness/nsp/cli/__init__.py` (Typer app)
|
||
- Create: `harness/nsp/cli/phase_cmd.py` (adapted)
|
||
- Create: `harness/nsp/cli/workspace_cmd.py`
|
||
- Port the command modules (sql, cte, db, lsh, vector, evidence, schema, memory, search, session, decision) — adapt imports
|
||
- Test: `harness/tests/test_cli_phase_meta.py`
|
||
|
||
**Honest scope:** this task ports the CLI structure + `phase meta` and verifies that ONE command behaves. It does NOT verify the 11 ported command modules — those get contract tests in **Task A9 (Livello B)**. Porting them here without tests would assume them reliable, which the spec forbids.
|
||
|
||
- [ ] **Step 1: Port CLI modules** — copy `ChironeWp3/src/psdwp3/cli/*.py` → `harness/nsp/cli/`. Rename package imports. The gate will need `nsp phase meta --json` to avoid mirroring constants.
|
||
|
||
- [ ] **Step 2: Write failing test for `nsp phase meta --json`**
|
||
|
||
```python
|
||
# tests/test_cli_phase_meta.py
|
||
import json
|
||
from typer.testing import CliRunner
|
||
from nsp.cli import app
|
||
|
||
runner = CliRunner()
|
||
|
||
def test_phase_meta_json_returns_workflow_data():
|
||
result = runner.invoke(app, ["phase", "meta", "--json"])
|
||
assert result.exit_code == 0
|
||
data = json.loads(result.stdout)
|
||
assert data["max_phase"] == 8
|
||
assert len(data["phases"]) == 8
|
||
assert data["phases"][7]["name"] == "datamart" # F8 present (fixes the JS drift)
|
||
assert "advance" in data["phases"][0]
|
||
```
|
||
|
||
- [ ] **Step 3: Run, verify fail, then implement `phase meta` in `phase_cmd.py`**
|
||
|
||
The command reads from `load_workflow()` and dumps `{max_phase, phases: [{num,name,advance,artifacts_out}]}` as JSON.
|
||
|
||
- [ ] **Step 4: Run, verify pass; commit**
|
||
|
||
```bash
|
||
git add harness/nsp/cli/ harness/tests/test_cli_phase_meta.py
|
||
git commit -m "feat(harness): port CLI structure + nsp phase meta --json (F2, kills JS/Python drift)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task A9: Contract tests for ported load-bearing modules — L0 (testcontainers) + L1 (fake data)
|
||
|
||
**Rationale:** spec §1 forbids assuming ported code is reliable. The DB-touching modules get **L0 tests** (testcontainers — real Postgres in container, where "not assumed reliable" gains real teeth); the logic modules get **L1 tests** (fake data). Docker is available on the dev machine, so L0 runs locally on every `pytest`.
|
||
|
||
**L0 tests (testcontainers, marker `@pytest.mark.l0`):**
|
||
|
||
- [ ] **Step 1: `tests/l0/test_db_connection.py`** — read-only enforcement (exit code 2 if writable role), `ping` succeeds on read-only role.
|
||
|
||
- [ ] **Step 2: `tests/l0/test_db_introspect.py` + `test_db_sampling.py`** — create a known schema + known rows in the container, verify `introspect` returns them with columns + comments, and `unique_values_for_lsh` returns the expected most-frequent values.
|
||
|
||
- [ ] **Step 3: `tests/l0/test_rrf.py`** — load real-ish data into the container (LSH index built from sampled values + a small vector set), verify `combined_search`/`rrf_fuse` fuses them with stable, sensible ranking. This is the one that truly exercises the ported RRF/LSH pipeline.
|
||
|
||
**L1 tests (fake data, no marker):**
|
||
|
||
- [ ] **Step 4: `tests/test_rest_client.py`** — mock `requests`, verify one RPC call carries the `X-API-Key` header and parses the response.
|
||
|
||
- [ ] **Step 5: `tests/test_mschema_render.py`** — the 3 formats (markdown, mschema-text ThothAI, schema-dict) from a fake `PhysicalSchema`. Catches port breaks in the render layer.
|
||
|
||
- [ ] **Step 6: `tests/test_mschema_eligibility.py`** — eligibility rules on fake columns (wide-text excluded above threshold).
|
||
|
||
- [ ] **Step 7: `tests/test_cli_contract.py`** — one test per ported CLI command (sql, cte, schema, session, decision, lsh, vector, evidence, memory, search, db): valid input → exit 0 + parseable JSON.
|
||
|
||
- [ ] **Step 8: Register `l0` marker in `pyproject.toml`** and run all A9 tests.
|
||
|
||
```toml
|
||
[tool.pytest.ini_options]
|
||
testpaths = ["tests"]
|
||
markers = [
|
||
"l0: testcontainers tests (need Docker, run locally)",
|
||
"l2: end-to-end tests requiring real GLM 5.2 + remote DB (skipped when .env incomplete)",
|
||
]
|
||
addopts = "-m 'not l2'" # L0 runs by default (Docker present); L2 opt-in
|
||
```
|
||
|
||
```bash
|
||
cd harness && pytest tests/l0/ tests/test_rest_client.py tests/test_mschema_render.py \
|
||
tests/test_mschema_eligibility.py tests/test_cli_contract.py -v
|
||
```
|
||
|
||
- [ ] **Step 9: Fix porting bugs the tests surface; commit**
|
||
|
||
```bash
|
||
git add harness/tests/l0/ harness/tests/test_rest_client.py harness/tests/test_mschema_*.py \
|
||
harness/tests/test_cli_contract.py harness/nsp/{db,rest,mschema,search}/ harness/pyproject.toml
|
||
git commit -m "test(harness): L0 testcontainers + L1 contract tests for ported modules (spec §1)"
|
||
```
|
||
|
||
---
|
||
|
||
## Phase B — Feature deviations (D11, D13, D14)
|
||
|
||
### Task B1: Dual vector API key — port `vectorstore` + verify writer config (D11)
|
||
|
||
**Files:**
|
||
- Port: `harness/nsp/vectorstore/{rest_client.py, rest_writer.py, store.py, reader.py, embeddings.py, records.py}`
|
||
- Port: `harness/nsp/cli/_guards.py`
|
||
- Port: `harness/scripts/create_vector_writer_rpc.sql`
|
||
- Create: `harness/scripts/create_vector_reader_rpc.sql`
|
||
- Test: `harness/tests/test_vector_dual_key.py`
|
||
|
||
- [ ] **Step 1: Port the vectorstore modules** (rest_client, rest_writer, store, reader, embeddings, records) and `_guards.py` from ChironeWp3. These already implement the dual-key model per spec §5.4 — port them, rename package, verify `has_vector_write_rest` checks `.api_key.strip()`.
|
||
|
||
- [ ] **Step 2: Write failing test for dual-key client construction**
|
||
|
||
```python
|
||
# tests/test_vector_dual_key.py
|
||
from nsp.config import RestConfig
|
||
from nsp.vectorstore.rest_client import VectorRestClient
|
||
|
||
def test_reader_and_writer_use_separate_keys():
|
||
reader = VectorRestClient(RestConfig(base_url="https://v/", api_key="K-READ"))
|
||
writer = VectorRestClient(RestConfig(base_url="https://v/", api_key="K-WRITE"))
|
||
assert reader.api_key == "K-READ"
|
||
assert writer.api_key == "K-WRITE"
|
||
|
||
def test_has_vector_write_rest_false_for_empty_key():
|
||
from nsp.cli._guards import has_vector_write_rest
|
||
from nsp.config import Config, RestConfig
|
||
cfg = Config(vector_write_rest=RestConfig(base_url="x", api_key=" "))
|
||
assert has_vector_write_rest(cfg) is False
|
||
```
|
||
|
||
- [ ] **Step 3: Run, verify fail, adjust ported code so `VectorRestClient` exposes `api_key` (it stores the config; add a property if missing), verify pass.**
|
||
|
||
- [ ] **Step 4: Create `create_vector_reader_rpc.sql`** — author the reader RPC (`search_similar`, `list_tables`) mirroring the writer allowlist pattern, granted to a `vector_reader` role. (The reader RPCs lived server-side in Supabase and were never in the ChironeWp3 repo — spec §5.4 note. Author them now.)
|
||
|
||
- [ ] **Step 5: Commit**
|
||
|
||
```bash
|
||
git add harness/nsp/vectorstore/ harness/nsp/cli/_guards.py harness/scripts/ harness/tests/test_vector_dual_key.py
|
||
git commit -m "feat(harness): port vectorstore dual-key + reader RPC (D11, §5.4)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task B2: `nsp memory save-one` — targeted upsert via writer key (D11)
|
||
|
||
**Files:**
|
||
- Modify: `harness/nsp/cli/memory_cmd.py` (add `save-one`)
|
||
- Modify: `harness/nsp/memory.py` (add `memory_vector_record_for_decision`)
|
||
- Test: `harness/tests/test_memory_save_one.py`
|
||
|
||
- [ ] **Step 1: Write failing test**
|
||
|
||
```python
|
||
# tests/test_memory_save_one.py
|
||
from pathlib import Path
|
||
from unittest.mock import patch, MagicMock
|
||
from typer.testing import CliRunner
|
||
from nsp.cli import app
|
||
|
||
runner = CliRunner()
|
||
|
||
def test_save_one_calls_upsert_with_single_record(tmp_path, monkeypatch):
|
||
session = tmp_path / "s"; session.mkdir()
|
||
# build a minimal registry + a decision_seq to save
|
||
from nsp.decisions import append_decision, DecisionRecord
|
||
append_decision(session, DecisionRecord(seq=7, ts="t", type="table_promoted",
|
||
subject="phase:4", detail="pazienti", rationale="r"))
|
||
mock_writer = MagicMock()
|
||
mock_writer.upsert_records.return_value = 1
|
||
with patch("nsp.cli.memory_cmd.open_store", return_value=mock_writer):
|
||
result = runner.invoke(app, ["memory", "save-one", "--session", str(session), "--decision-seq", "7"])
|
||
assert result.exit_code == 0
|
||
mock_writer.sync.assert_not_called() # must be a single upsert, NOT full resync
|
||
mock_writer.upsert_records.assert_called_once()
|
||
args = mock_writer.upsert_records.call_args
|
||
assert len(args[0][1]) == 1 # exactly one row
|
||
|
||
def test_save_one_refused_without_writer_key(tmp_path, monkeypatch):
|
||
monkeypatch.setenv("THOTH_PROFILE", "workstation")
|
||
session = tmp_path / "s"; session.mkdir()
|
||
# config with no write_rest
|
||
result = runner.invoke(app, ["memory", "save-one", "--session", str(session), "--decision-seq", "1"])
|
||
assert result.exit_code == 4 # require_vector_write_allowed gate
|
||
```
|
||
|
||
- [ ] **Step 2: Run, verify fail**
|
||
|
||
- [ ] **Step 3: Implement `save-one`** in `memory_cmd.py`:
|
||
- Guard: `require_vector_write_allowed(cfg, "memory save-one")`.
|
||
- Load the registry, find the memory record for `decision_seq`.
|
||
- Build a single `VectorRecord` (reuse `memory_vector_records` but filter to one).
|
||
- Call `writer.existing_hashes(kinds={"memory"})` + embed (if hash changed) + `writer.upsert_records(table="memory", rows=[one])`.
|
||
- Do NOT call `writer.sync` (that's the full-resync path).
|
||
|
||
- [ ] **Step 4: Run, verify pass; commit**
|
||
|
||
```bash
|
||
git add harness/nsp/cli/memory_cmd.py harness/nsp/memory.py harness/tests/test_memory_save_one.py
|
||
git commit -m "feat(harness): nsp memory save-one — targeted upsert via writer key (D11)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task B3: Value & Schema Linking — `value_grounded` decision + LSH multi-column (D14a)
|
||
|
||
**Files:**
|
||
- Modify: `harness/nsp/decisions.py` (add `value_grounded`)
|
||
- Modify: `harness/nsp/search/__init__.py` (stop collapsing multi-column to single best)
|
||
- Modify: `harness/nsp/session/models.py` (`Candidate.grounded_values`)
|
||
- Test: `harness/tests/test_value_grounding.py`
|
||
|
||
- [ ] **Step 1: Port `search/__init__.py`, `lshindex/`, `db/sampling.py`** from ChironeWp3. Then modify.
|
||
|
||
- [ ] **Step 2: Write failing test for multi-column LSH exposure**
|
||
|
||
```python
|
||
# tests/test_value_grounding.py
|
||
from nsp.search import aggregate_lsh_multi
|
||
|
||
def test_value_in_multiple_columns_returns_all():
|
||
hits = [
|
||
{"table": "t", "column": "c1", "value": "ablazione", "score": 0.9},
|
||
{"table": "t", "column": "c2", "value": "ablazione", "score": 0.7},
|
||
]
|
||
result = aggregate_lsh_multi(hits)
|
||
# NOT collapsed to single best — both columns exposed
|
||
cols = {h["column"] for h in result["t"]}
|
||
assert cols == {"c1", "c2"}
|
||
```
|
||
|
||
- [ ] **Step 3: Run, verify fail, then implement `aggregate_lsh_multi`** (replaces the collapsing `_aggregate_lsh` behavior — keep the old name as a thin wrapper to the new one if other code needs single-best). Add `value_grounded` to DecisionType. Add `grounded_values: list[dict] = []` to `Candidate` in `session/models.py`.
|
||
|
||
- [ ] **Step 4: Run, verify pass; commit**
|
||
|
||
```bash
|
||
git add harness/nsp/search/ harness/nsp/lshindex/ harness/nsp/db/sampling.py harness/nsp/decisions.py harness/nsp/session/models.py harness/tests/test_value_grounding.py
|
||
git commit -m "feat(harness): value grounding — multi-column LSH + value_grounded (D14a)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task B4: SQL formula evidence — formula kind + retrieval + approval (D14b)
|
||
|
||
**Files:**
|
||
- Modify: `harness/nsp/evidence/model.py` (activate `tier: concept` / add formula fields)
|
||
- Create: `harness/nsp/evidence/formula_store.py`
|
||
- Modify: `harness/nsp/cli/search_cmd.py` (`--kind formula`)
|
||
- Modify: `harness/nsp/decisions.py` (add `concept_formula_approved`, `concept_formula_rejected`)
|
||
- Test: `harness/tests/test_formula.py`
|
||
|
||
- [ ] **Step 1: Port `evidence/` modules** from ChironeWp3.
|
||
|
||
- [ ] **Step 2: Write failing test for formula retrieval**
|
||
|
||
```python
|
||
# tests/test_formula.py
|
||
from pathlib import Path
|
||
from nsp.evidence.formula_store import ConceptFormula, save_formula, retrieve_formula
|
||
|
||
def test_formula_retrieval_by_concept(tmp_path):
|
||
f = ConceptFormula(concept="fascia pediatrica", columns=["data_nascita"],
|
||
sql="CASE WHEN ... END", status="reviewed", sources=["s1"])
|
||
save_formula(tmp_path, f)
|
||
results = retrieve_formula(tmp_path, "fascia pediatrica")
|
||
assert len(results) == 1
|
||
assert results[0].sql.startswith("CASE WHEN")
|
||
assert results[0].concept == "fascia pediatrica"
|
||
```
|
||
|
||
- [ ] **Step 3: Run, verify fail, implement `formula_store.py`** (frontmatter YAML + SQL body, concept→formula units). Add `--kind formula` to `search_cmd` that calls `retrieve_formula`. Add the two new decision types.
|
||
|
||
- [ ] **Step 4: Run, verify pass; commit**
|
||
|
||
```bash
|
||
git add harness/nsp/evidence/ harness/nsp/cli/search_cmd.py harness/nsp/decisions.py harness/tests/test_formula.py
|
||
git commit -m "feat(harness): SQL formula evidence — concept→formula units + retrieval + approval decisions (D14b)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task B5: Free-text interpretation guidance (D13)
|
||
|
||
**Files:**
|
||
- Modify: `harness/nsp/cli/decision_cmd.py` (ensure free-text from "Altro"/"Rifiuta" is captured in `rationale`)
|
||
- Modify: `.pi/skills/nsp-sessione/SKILL.md` (+ sub-files) — port + add the D13 §4.6 guidance
|
||
- Test: `harness/tests/test_freetext_interpretation.py`
|
||
|
||
- [ ] **Step 1: Write failing test** — a golden-style test that, given a `ui_response` with `control: freetext` (Altro text), verifies the harness records the user's text in the decision `rationale` (not discarded).
|
||
|
||
```python
|
||
# tests/test_freetext_interpretation.py
|
||
from pathlib import Path
|
||
from nsp.decisions import append_decision, list_decisions, DecisionRecord
|
||
|
||
def test_altro_freetext_recorded_in_rationale(tmp_path):
|
||
s = tmp_path / "sess"; s.mkdir()
|
||
# The gate, on receiving Altro text, appends a decision whose rationale carries the user's words.
|
||
append_decision(s, DecisionRecord(seq=1, ts="t", type="table_promoted",
|
||
subject="phase:4", detail="ablazione",
|
||
rationale="Utente (Altro): 'ablazione' va cercato anche in patologia, non solo nel flag"))
|
||
decisions = list_decisions(s)
|
||
assert "patologia" in decisions[0].rationale # user text preserved, not discarded
|
||
```
|
||
|
||
- [ ] **Step 2: Port `SKILL.md` + sub-files** from ChironeWp3. Add a section "Interpretazione del testo libero" per spec §4.6: instructs the model to (1) evaluate Altro/Rifiuta/steering text in context, (2) act on it, (3) if ambiguous, re-ask instead of defaulting, (4) record the text in the decision rationale.
|
||
|
||
- [ ] **Step 3: Run, verify pass (the test asserts the recording contract; the actual model behavior is enforced by the skill prose + gate, validated by fake-pi golden tests in Task C3); commit**
|
||
|
||
```bash
|
||
git add harness/.pi/skills/ harness/nsp/cli/decision_cmd.py harness/tests/test_freetext_interpretation.py
|
||
git commit -m "feat(harness): free-text interpretation guidance + rationale capture (D13)"
|
||
```
|
||
|
||
---
|
||
|
||
## Phase C — Gate → widget-descriptor (D2/D4)
|
||
|
||
The gate has two parts with **very different testability**:
|
||
- **Builders** (pure functions that turn tool-call params into widget-descriptor JSON) — fully testable in L1, **in JS, in-language**.
|
||
- **Glue** (Pi runtime integration: `pi.registerTool`, `ctx.sendRaw`, `emitAndWait`, anti-bypass hooks, no-limbo loop) — **NOT testable in L1**; it depends on the Pi runtime. Tested at L2.
|
||
|
||
This phase therefore has: **C1 (builders, L1, JS)** + **C2 (glue implementation, no L1 test — verified at L2)**. There is no fake-LLM driver and no `build-scenarios` in this plan (see cross-cutting follow-up).
|
||
|
||
### Task C1: Pure widget-builder functions + JS tests (L1)
|
||
|
||
**Files:**
|
||
- Create: `harness/.pi/extensions/gate/builders.js` (pure functions, no Pi context)
|
||
- Create: `harness/.pi/extensions/gate/__tests__/builders.test.js` (node:test, in-language)
|
||
- Create: `harness/.pi/extensions/gate/__tests__/golden/*.json` (golden widget descriptors)
|
||
- Create: `harness/package.json` (for `npm test` → `node --test`)
|
||
|
||
This task extracts widget-descriptor **construction** into pure, testable functions and tests them in JS directly — no Python↔JS bridge, no Python mirror (a mirror would drift; the golden tests would validate the mirror, not the real gate).
|
||
|
||
- [ ] **Step 1: Create `package.json`** for the gate JS tests.
|
||
|
||
```json
|
||
{
|
||
"name": "thothii-harness-gate",
|
||
"private": true,
|
||
"scripts": { "test": "node --test .pi/extensions/gate/__tests__/" }
|
||
}
|
||
```
|
||
|
||
- [ ] **Step 2: Define the pure builder API in `builders.js`.** Each builder takes plain params and returns a plain widget-descriptor object. No `ctx`, no I/O.
|
||
|
||
```javascript
|
||
// gate/builders.js — pure functions, no Pi context. Each returns a ui_request descriptor (spec §4.1).
|
||
function buildSelectRequest({ id, phase, title, options, intro = null, allowOther = true, recommended = null }) { /* ... */ }
|
||
function buildMultiselectRequest({ id, phase, title, options, content = null, allowEmpty = false, allowOther = true }) { /* ... */ }
|
||
function buildArtifactGate({ id, phase, title, artifact /* {kind, data, version} */, action /* {kind, prompt} */ }) { /* ... */ }
|
||
function buildInfoRequest({ phase, level, text }) { /* ... */ }
|
||
function buildFreetextRequest({ id, phase, title }) { /* ... */ }
|
||
function withChildLinkage(option, widgetSpec) { /* sets option.opens = widgetSpec; returns option */ }
|
||
module.exports = { buildSelectRequest, buildMultiselectRequest, buildArtifactGate, buildInfoRequest, buildFreetextRequest, withChildLinkage };
|
||
```
|
||
|
||
- [ ] **Step 3: Write JS golden tests** — one golden file per widget type, one test per widget. Assert exact shape.
|
||
|
||
```javascript
|
||
// gate/__tests__/builders.test.js
|
||
const test = require("node:test");
|
||
const assert = require("node:assert");
|
||
const fs = require("node:fs");
|
||
const path = require("node:path");
|
||
const { buildSelectRequest, buildMultiselectRequest, buildArtifactGate } = require("../builders.js");
|
||
|
||
const GOLDEN = path.join(__dirname, "golden");
|
||
|
||
test("select F1 matches golden", () => {
|
||
const result = buildSelectRequest({ id: "u1", phase: "F1", title: "Disambigua 'ablazione'",
|
||
options: [{ id: "o1", label: "procedura" }, { id: "o2", label: "patologia" }] });
|
||
const golden = JSON.parse(fs.readFileSync(path.join(GOLDEN, "select_F1.json")));
|
||
assert.equal(result.widget, "select");
|
||
assert.deepEqual(result.reserved, golden.reserved); // back/exit/other
|
||
assert.deepEqual(result.options, golden.options);
|
||
assert.ok(result.schema_version);
|
||
});
|
||
|
||
test("artifact-gate F5 schema matches golden", () => {
|
||
const result = buildArtifactGate({ id: "u2", phase: "F5", title: "Schema-linking",
|
||
artifact: { kind: "schema_linking", data: { candidates: [] }, version: 1 },
|
||
action: { kind: "confirm", prompt: "Confermi?" } });
|
||
assert.equal(result.widget, "artifact-gate");
|
||
assert.equal(result.artifact.kind, "schema_linking");
|
||
assert.equal(result.action.kind, "confirm");
|
||
});
|
||
|
||
test("multiselect F4 carries content + allowEmpty", () => {
|
||
const result = buildMultiselectRequest({ id: "u3", phase: "F4", title: "Tabelle",
|
||
options: [], content: { candidates: [] }, allowEmpty: false });
|
||
assert.equal(result.widget, "multiselect");
|
||
assert.equal(result.allow_empty, false);
|
||
assert.ok("content" in result);
|
||
});
|
||
|
||
test("Altro option carries freetext linkage", () => {
|
||
const result = buildSelectRequest({ id: "u4", phase: "F1", title: "x", options: [] });
|
||
const other = result.options.find(o => o.id === "other");
|
||
assert.equal(other.opens.widget, "freetext");
|
||
});
|
||
```
|
||
|
||
- [ ] **Step 4: Create the golden files**, run `npm test`, verify pass.
|
||
|
||
```bash
|
||
cd harness && npm test
|
||
```
|
||
|
||
- [ ] **Step 5: Fuzzy tests on builders (L1, JS)** — malformed params (missing `title`, non-list `options`, empty `options` with `allowEmpty:false`). Assert builders throw clearly or return a well-formed error descriptor, never silently produce a broken widget.
|
||
|
||
```javascript
|
||
test("select with missing title throws clearly", () => {
|
||
assert.throws(() => buildSelectRequest({ id: "u", phase: "F1", options: [] }), /title/);
|
||
});
|
||
test("multiselect allowEmpty:false with zero options throws", () => {
|
||
assert.throws(() => buildMultiselectRequest({ id: "u", phase: "F4", title: "x", options: [], allowEmpty: false }), /allow_empty|options/);
|
||
});
|
||
```
|
||
|
||
- [ ] **Step 6: Commit**
|
||
|
||
```bash
|
||
git add harness/package.json harness/.pi/extensions/gate/
|
||
git commit -m "feat(harness): pure widget-builder functions + JS golden/fuzzy tests (D2, L1)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task C2: Gate glue — rewrite `nsp-gate.js` (implementation; verification at L2)
|
||
|
||
**Files:**
|
||
- Create: `harness/.pi/extensions/nsp-gate.js` (rewritten from ChironeWp3)
|
||
- Port: `harness/.pi/extensions/reserved-labels.mjs`
|
||
- **No L1 test** — the glue depends on the Pi runtime and is verified at L2 (Task D4).
|
||
|
||
**Honest scope:** this task implements the glue but cannot unit-test it in L1. The glue's correctness is validated end-to-end at L2 with real Pi + GLM 5.2. This is the load-bearing limitation stated in the Testing Strategy headline. Do NOT attempt to mock Pi here — that is the fake-Pi follow-up (cross-cutting), not part of this plan.
|
||
|
||
- [ ] **Step 1: Port `reserved-labels.mjs`** (BACK/EXIT/OTHER labels, `isReserved`/`stripReserved`).
|
||
|
||
- [ ] **Step 2: Port the anti-bypass hooks verbatim** from the original `nsp-gate.js`: `tool_call` hook blocking `nsp phase advance|reopen`, `nsp decision add`, `nsp cte plan`; protected-file list (`review_decisions.jsonl`, `session_manifest.yaml`, `cte_plan.json`); input-lock; `before_agent_start` kickoff. These are load-bearing (spec D4) and unchanged.
|
||
|
||
- [ ] **Step 3: Wire tools to builders + emission.** Each of the 4 tools (`reviewer_select`, `reviewer_decide`, `reviewer_confirm`, `rewrite_question`) calls the matching builder from C1, then emits via `ctx.sendRaw({type:"extension_ui_request", ...})` and awaits the correlated response by `id`. Preserve: no-limbo invariant (Esc/cancel re-presents — never returns undefined), recommended-option marker, reading `SCHEMA_LINKING_PHASE`/`PHASE_NAMES` from `nsp phase meta --json` (no mirroring).
|
||
|
||
```javascript
|
||
// nsp-gate.js (the glue) — wires C1 builders to the Pi runtime
|
||
const { buildSelectRequest, buildMultiselectRequest, buildArtifactGate } = require("./gate/builders.js");
|
||
|
||
module.exports = function (pi) {
|
||
pi.registerTool({
|
||
name: "reviewer_select", /* ... */,
|
||
async execute(id, params, signal, onUpdate, ctx) {
|
||
const widget = buildSelectRequest({ id, phase: await currentPhase(ctx), ...params });
|
||
const resp = await emitAndWait(ctx, widget); // extension_ui_request → wait → response
|
||
return handleSelectResponse(resp, params); // incl. Altro→freetext linkage, Back→reopen
|
||
}
|
||
});
|
||
// ... reviewer_decide, reviewer_confirm, rewrite_question, anti-bypass hooks, kickoff injection
|
||
};
|
||
```
|
||
|
||
- [ ] **Step 4: Manual smoke (optional, pre-L2)** — launch `pi --mode rpc` from `harness/` once, confirm the extension loads and `/nuova-domanda` triggers a tool registration (no crash). This is a sanity check, not a regression test.
|
||
|
||
- [ ] **Step 5: Commit** (verification happens at L2 in Task D4)
|
||
|
||
```bash
|
||
git add harness/.pi/extensions/nsp-gate.js harness/.pi/extensions/reserved-labels.mjs
|
||
git commit -m "feat(harness): rewrite nsp-gate.js glue wired to builders (D2/D4) — verified at L2"
|
||
```
|
||
|
||
---
|
||
|
||
## Phase D — End-to-end validation (L1 smoke + L2 real)
|
||
|
||
### Task D1: L1 session-coherence smoke test (pure logic, CI)
|
||
|
||
**Files:**
|
||
- Create: `harness/tests/test_session_coherence_smoke.py`
|
||
|
||
A pure-logic test (no LLM, no DB) that constructs a plausible ledger by hand and asserts the invariants D1-the-L2-version would check. This is the CI-runnable proxy for "full session coherence."
|
||
|
||
- [ ] **Step 1: Write the test** — build a synthetic ledger (decisions appended directly, not via LLM) representing a full F1→F8 walk + a rollback F6→F4 + re-derive. Assert: `current_phase` folds correctly at each step, `effective_decisions` excludes the stale tail after rollback, `teardown_to_phase(4)` deletes the right artifacts, `generate_task_doc` stays under byte budget at every phase, and `nsp session consistency` (the post-rollback coherence check) passes.
|
||
|
||
```python
|
||
# tests/test_session_coherence_smoke.py
|
||
from pathlib import Path
|
||
from nsp.decisions import append_decision, DecisionRecord
|
||
from nsp.phase import current_phase, effective_decisions
|
||
from nsp.teardown import teardown_to_phase
|
||
from nsp.taskdoc import generate_task_doc
|
||
|
||
def test_full_walk_then_rollback_stays_coherent(tmp_path):
|
||
s = tmp_path / "sess"; s.mkdir()
|
||
# build a synthetic full-session ledger
|
||
for n in range(1, 9):
|
||
append_decision(s, DecisionRecord(seq=n, ts="t", type="phase_approved", subject=f"phase:{n}",
|
||
detail="", rationale=""))
|
||
assert current_phase(s) == 9 # max_phase + 1
|
||
# rollback to F4
|
||
append_decision(s, DecisionRecord(seq=9, ts="t", type="phase_reopened", subject="phase:4",
|
||
detail="", rationale=""))
|
||
report = teardown_to_phase(s, target_phase=4)
|
||
assert current_phase(s) == 4
|
||
assert "sql_final.sql" in report.deleted_files or "cte_plan.json" in report.deleted_files
|
||
# task docs stay bounded across phases
|
||
for ph in range(1, 9):
|
||
doc = generate_task_doc(session_dir=s, phase=ph, promoted_tables=["dim_paziente"])
|
||
assert doc.byte_budget_ok, f"phase {ph} task doc over budget"
|
||
```
|
||
|
||
- [ ] **Step 2: Run, verify pass; commit**
|
||
|
||
```bash
|
||
git add harness/tests/test_session_coherence_smoke.py
|
||
git commit -m "test(harness): L1 session-coherence smoke (full walk + rollback, pure logic)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task D2: `.pi/` config + prompts + theme ported
|
||
|
||
**Files:**
|
||
- Create: `harness/.pi/settings.json`
|
||
- Create: `harness/.pi/themes/` (port theme)
|
||
- Create: `harness/.pi/prompts/nuova-domanda.md`, `riprendi-sessione.md`
|
||
- No test (config files)
|
||
|
||
- [ ] **Step 1: Port** `.pi/settings.json`, theme, and the two prompt files from ChironeWp3. Verify `pi --mode rpc` launched with `cwd=harness/` finds the `.pi/` directory.
|
||
|
||
- [ ] **Step 2: Commit**
|
||
|
||
```bash
|
||
git add harness/.pi/
|
||
git commit -m "feat(harness): port .pi/ config, prompts, theme"
|
||
```
|
||
|
||
---
|
||
|
||
### Task D3: L2 conftest + skip-when-no-.env guard
|
||
|
||
**Files:**
|
||
- Create: `harness/tests/conftest.py`
|
||
- Create: `harness/tests/l2/__init__.py`
|
||
- Modify: `harness/pyproject.toml` (register `l2` marker)
|
||
|
||
- [ ] **Step 1: Register the `l2` marker** in `pyproject.toml`:
|
||
|
||
```toml
|
||
[tool.pytest.ini_options]
|
||
testpaths = ["tests"]
|
||
markers = [
|
||
"l2: end-to-end tests requiring real GLM 5.2 + remote DB (skipped when .env incomplete)",
|
||
]
|
||
addopts = "-m 'not l2'" # CI default: skip L2 unless explicitly requested
|
||
```
|
||
|
||
- [ ] **Step 2: Write `conftest.py`** with a session fixture that loads `.env` and a `l2_env` fixture that skips if any required var is missing.
|
||
|
||
```python
|
||
# tests/conftest.py
|
||
import os
|
||
from pathlib import Path
|
||
import pytest
|
||
from dotenv import load_dotenv
|
||
|
||
REQUIRED_L2 = ["THOTH_DWH_API_KEY", "THOTH_VEC_API_KEY", "THOTH_VEC_WRITE_API_KEY", "THOTH_SSL_CA"]
|
||
|
||
@pytest.fixture(scope="session", autouse=True)
|
||
def _load_env():
|
||
load_dotenv(Path(__file__).resolve().parent.parent / ".env")
|
||
|
||
@pytest.fixture(scope="session")
|
||
def l2_env():
|
||
missing = [v for v in REQUIRED_L2 if not os.environ.get(v, "").strip()]
|
||
if missing:
|
||
pytest.skip(f"L2 skipped — missing env vars: {', '.join(missing)} (populate harness/.env)")
|
||
return True
|
||
```
|
||
|
||
- [ ] **Step 3: Commit**
|
||
|
||
```bash
|
||
git add harness/tests/conftest.py harness/tests/l2/__init__.py harness/pyproject.toml
|
||
git commit -m "test(harness): L2 marker + skip-when-no-.env guard (Testing Strategy)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task D4: L2 full session — "ablazione" with GLM 5.2 (human-in-the-loop)
|
||
|
||
**Files:**
|
||
- Create: `harness/tests/l2/test_session_ablazione.py`
|
||
- Create: `harness/workspaces/chirone-test.yaml`
|
||
|
||
This is the **end-to-end validation that L1 cannot do**: GLM 5.2 driving the real harness against the real DWH + pgvector, on the "ablazione" question (which exercises D14 value grounding + formula on a multi-column case). Default mode = human answers via terminal; scripted-answers mode for repeatability.
|
||
|
||
- [ ] **Step 1: Create `workspaces/chirone-test.yaml`** pointing at the remote endpoints per the L2 connection params. Secrets via `${THOTH_*}`.
|
||
|
||
```yaml
|
||
name: chirone-test
|
||
description: "Chirone DWH — L2 test workspace"
|
||
relational:
|
||
db_type: postgres
|
||
transport: rest
|
||
rest:
|
||
base_url: https://supabase-aritmolab.policlinicosandonato.it/dwh/
|
||
api_key: ${THOTH_DWH_API_KEY}
|
||
ssl_ca: ${THOTH_SSL_CA}
|
||
vector_db:
|
||
collection: chirone_docs
|
||
dim: 768
|
||
rest:
|
||
base_url: https://supabase-aritmolab.policlinicosandonato.it/vector/v1/
|
||
api_key: ${THOTH_VEC_API_KEY}
|
||
ssl_ca: ${THOTH_SSL_CA}
|
||
write_rest:
|
||
base_url: https://supabase-aritmolab.policlinicosandonato.it/vector/v1/
|
||
api_key: ${THOTH_VEC_WRITE_API_KEY}
|
||
ssl_ca: ${THOTH_SSL_CA}
|
||
evidence:
|
||
source_root: ${EVIDENCE_ROOT}
|
||
evidence_dir: evidence/chirone
|
||
embeddings:
|
||
provider: ollama
|
||
base_url: ${THOTH_OLLAMA_URL}
|
||
model: nomic-embed-text-v2-moe
|
||
dim: 768
|
||
batch_size: 64
|
||
execution:
|
||
allow: [cte_test, explain, preview, aggregate, export]
|
||
max_preview_rows: 10
|
||
statement_timeout_ms: 5000
|
||
forbidden_functions: [set_config, dblink, dblink_exec, lo_import]
|
||
```
|
||
|
||
- [ ] **Step 2: Write the L2 test** — spawns Pi (`pi --mode rpc`, cwd=harness/) with GLM 5.2, runs `/nuova-domanda "dammi la lista dei pazienti che hanno fatto un'ablazione nel 2025"`, and either (a) prompts the reviewer in the terminal for each gate widget (human mode) or (b) feeds canned answers (scripted mode via `--answers-file`). Asserts: a session is produced, it reaches a finalized state, the `value_grounded` and `concept_formula_approved` decisions appear (D14 exercised), `sql_final.sql` is present and read-only-validates against the DWH.
|
||
|
||
```python
|
||
# tests/l2/test_session_ablazione.py
|
||
import subprocess, json
|
||
from pathlib import Path
|
||
import pytest
|
||
|
||
@pytest.mark.l2
|
||
def test_ablazione_session_human_in_loop(l2_env, tmp_path):
|
||
"""L2: GLM 5.2 + real DWH. Reviewer answers gates via terminal.
|
||
Run with: pytest -m l2 tests/l2/test_session_ablazione.py -s
|
||
"""
|
||
session_dir = tmp_path / "sess"
|
||
proc = subprocess.run(
|
||
["pi", "--mode", "rpc"],
|
||
cwd="harness/",
|
||
# the test harness feeds /nuova-domanda and relays terminal I/O for reviewer answers
|
||
...
|
||
timeout=600,
|
||
)
|
||
assert (session_dir / "sql_final.sql").exists()
|
||
ledger = (session_dir / "review_decisions.jsonl").read_text().splitlines()
|
||
types = {json.loads(l)["type"] for l in ledger}
|
||
assert "value_grounded" in types or "concept_formula_approved" in types # D14 exercised
|
||
```
|
||
|
||
- [ ] **Step 3: Run manually** (`pytest -m l2 -s`), confirm the session finalizes and D14 surfaces. This is non-deterministic and slow — it's a pre-release check, not CI. **Commit the test + workspace; the operator runs it before release.**
|
||
|
||
```bash
|
||
git add harness/tests/l2/test_session_ablazione.py harness/workspaces/chirone-test.yaml
|
||
git commit -m "test(harness): L2 ablazione session with GLM 5.2 + real DWH (D14, human/scripted)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task D5: L2 specific behaviors — value grounding (real) + memory save-one (real)
|
||
|
||
**Files:**
|
||
- Create: `harness/tests/l2/test_value_grounding_real.py`
|
||
- Create: `harness/tests/l2/test_memory_save_one_real.py`
|
||
|
||
Targeted L2 tests that close specific gaps L1 leaves open: real value grounding on the live schema, and real `memory save-one` upsert to pgvector.
|
||
|
||
- [ ] **Step 1: `test_value_grounding_real.py`** — runs `nsp search "ablazione" --kind values --workspace chirone-test`, asserts multiple columns are returned (not collapsed), and that a grounding widget would surface them. Validates D14a on the real schema.
|
||
|
||
- [ ] **Step 2: `test_memory_save_one_real.py`** — runs `nsp memory save-one` against `chirone-test` (writer key), asserts the upsert returns `upserted >= 0` (idempotent) and a subsequent `search_similar` finds the memory. Validates D11 end-to-end.
|
||
|
||
- [ ] **Step 3: Run manually (`pytest -m l2 -s tests/l2/`); commit**
|
||
|
||
```bash
|
||
git add harness/tests/l2/test_value_grounding_real.py harness/tests/l2/test_memory_save_one_real.py
|
||
git commit -m "test(harness): L2 value grounding (real schema) + memory save-one (real pgvector)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task D6: Documentation — README + workflow editing + testing guide
|
||
|
||
**Files:**
|
||
- Create: `harness/README.md`
|
||
- Create: `harness/docs/workflow-editing.md`
|
||
- Create: `harness/docs/testing.md`
|
||
|
||
- [ ] **Step 1: Write `README.md`** — install, configure `.env` + `workspaces/`, run `nsp`, run L0+L1 tests, run L2 tests (with the CA-bundle copy step).
|
||
|
||
- [ ] **Step 2: Write `docs/workflow-editing.md`** — how to edit `workflow.yaml` (add/reorder/merge/skip phases), referencing spec §5.3.
|
||
|
||
- [ ] **Step 3: Write `docs/testing.md`** — explain the L0/L1/L2 split honestly: what each level covers and does NOT cover (in plain language: L0 = DB-touching code vs real Postgres in container; L1 = pure logic + gate builders, no DB no LLM; L2 = full system with real GLM 5.2 + remote DB, manual pre-release). State the headline limitation plainly: the LLM→gate conversation has no automated regression coverage. How to run each level (`pytest` for L0+L1, `pytest -m l2` for L2 needing `.env` + VPN + CA bundle). The security note on keys (`.env` gitignored, never logged).
|
||
|
||
- [ ] **Step 4: Commit**
|
||
|
||
```bash
|
||
git add harness/README.md harness/docs/
|
||
git commit -m "docs(harness): README + workflow editing + testing guide (L1/L2)"
|
||
```
|
||
|
||
---
|
||
|
||
## Self-Review Notes
|
||
|
||
**Spec coverage check** (spec section → task):
|
||
- §1 (ChironeWp3 as starting point, NOT assumed reliable): enforced by a **3-tier porting classification**:
|
||
- **Tier A (modified by ThothII → full regression test):** phase.py (A5), decisions.py (A4), vectorstore/_guards dual-key (B1), memory_cmd save-one (B2), search multi-column (B3), evidence formula (B4), phase_cmd meta (A8). Each has a focused L1 test for the changed behavior.
|
||
- **Tier B (ported unchanged but load-bearing → contract test):** db/connection + introspect + sampling + RRF (**L0 testcontainers**, A9 steps 1-3 — real Postgres, where "not reliable" is tested for real), rest/client + mschema/render + mschema/eligibility + 11 CLI commands (**L1 fake data**, A9 steps 4-7). The L0/L1 split inside Tier B reflects whether the module touches a DB.
|
||
- **Tier C (ported unchanged, non-load-bearing → smoke OK):** fetch_ca, output/formatting helpers. Acceptable with smoke.
|
||
This makes "not assumed reliable" honest for the load-bearing surface. ✓
|
||
- D1 (3 projects): this plan = harness only; backend/frontend are separate plans. ✓
|
||
- D2/D4 (widget-descriptor gate): **builders** in C1 (L1, JS, in-language — golden + fuzzy); **glue** in C2 (implementation only — verified at L2, NOT in L1, because it depends on the Pi runtime). The limitation is documented in the Testing Strategy headline. ✓
|
||
- D3 (workspace YAML): Task A2. ✓
|
||
- D5 (FS persistence, ledger): Task A4 (decisions port), session models in A5. ✓
|
||
- D6 (auth): out of scope for harness (auth is backend). ✓ (noted)
|
||
- D7 (backend SQL read-only): out of scope for harness. ✓
|
||
- D8 (nsp stays Python): all harness tasks are Python (gate glue is JS, per the source). ✓
|
||
- D10 (testing): **three levels** — L0 testcontainers (A9 steps 1-3, runs locally), L1 fake/logic + gate builders (A2-A8, B1-B5, C1, D1; runs locally), L2 real GLM 5.2 + remote DB (D4, D5; pre-release). Honest headline in Testing Strategy: the skill→LLM→gate loop has NO automated regression coverage. ✓
|
||
- D11 (dual key + save-one): B1 (L1 dual-key), B2 (L1 save-one mocked), D5 (L2 save-one real). ✓
|
||
- D12 (deployment model B): harness runs locally — `.env.example` (THOTH_PROFILE), L2 against remote REST over VPN. ✓
|
||
- D13 (free-text interpretation): B5 (L1 rationale contract), D4 (L2 with real model — the glue that records rationale is L2-only). ✓
|
||
- D14a (value grounding): B3 (L1 aggregate_lsh_multi on fake hits), D5 (L2 real schema). ✓
|
||
- D14b (formula): B4 (L1 formula store). ✓ (L2 formula approval rides on D4)
|
||
- D15 (rollback): A4 (retract), A5 (effective view), A6 (teardown), D1 (L1 coherence smoke), D4 (L2 glue behavior). ✓
|
||
- D16 (context minimization): A7 (taskdoc, L1). ✓ (per-phase reset in Pi noted as execution detail)
|
||
|
||
**Type consistency:** `effective_decisions`, `teardown_to_phase`, `generate_task_doc`, `load_workflow`, `save_formula`, `retrieve_formula`, `aggregate_lsh_multi`, `decision_retracted`, `value_grounded`, `concept_formula_approved/rejected`, builder functions (`buildSelectRequest` etc.) — names match across tasks. ✓
|
||
|
||
**Testing honesty (the load-bearing facts):**
|
||
- **The skill→LLM→gate loop has NO automated regression coverage.** It is exercised only at L2 (manual, non-deterministic, pre-release). This is a conscious choice; plan L2 runs before release.
|
||
- **The gate glue cannot be unit-tested in L1** — it depends on the Pi runtime (`pi.registerTool`, `ctx.sendRaw`). Only the pure builders are L1-tested (in JS, in-language).
|
||
- L0 (testcontainers) tests the DB-touching ported code against a real Postgres — this is where "not assumed reliable" is actually enforced for the data layer.
|
||
- L2 tests (D4, D5) are non-deterministic, slow, require credentials + VPN + CA bundle. Skipped automatically when `.env` incomplete.
|
||
|
||
**Known follow-ups (not in this plan):**
|
||
- **Cross-cutting: a fake-Pi runtime mock** that lets the gate glue run in L1/CI. When built (likely alongside the backend plan, which also needs a fake-Pi to test its RPC client), it would close the gap "gate glue has no L1 test" AND enable a `build-scenarios`/`fake-LLM` replay harness for deterministic gate-behavior tests. This is the single highest-value testing investment not in this plan. It is deferred because it is substantial and shared with the backend.
|
||
- The SSE/REST surface the backend will build on (harness speaks JSONL via Pi + `nsp --json`).
|
||
- `session consistency` check integration into `nsp session check` (Task A5 stubs the effective view; the full assertion grows during execution).
|
||
- Per-phase context reset mechanism in Pi (D16 §4.9 step 3) — depends on Pi's context-management capability, verified when running the gate against real Pi in Task D4.
|
||
|
||
This plan is self-contained: at the end, `harness/` runs the full 8-phase workflow emitting/consuming widget-descriptor JSON, validated by L0 (testcontainers) + L1 (logic + gate builders) on every local run, plus L2 (GLM 5.2 + real DB) pre-release. Ready for the backend to spawn and drive.
|