78 KiB
ThothII — Harness Implementation Plan
For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (
- [ ]) syntax for tracking.
Goal: Build harness/ — the self-contained Pi layer (CLI nsp Python + gate extension JS + skills markdown + .pi/) that runs the 8-phase NL→SQL workflow emitting/consuming widget-descriptor JSON, derived from ChironeWp3 as a validated-and-adapted starting point (not assumed reliable).
Architecture: Port the relevant ChironeWp3 files into harness/ task-by-task, then adapt each to the ThothII contract: workflow.yaml as the single source of workflow truth (F2), an "effective decisions as of pointer" view + teardown_to_phase for correct rollback (D15), a per-step task-document generator with enforced bounds for a 35B/<200k model (D16), the dual vector API key + memory save-one (D11), free-text interpretation (D13), value/formula grounding (D14), and the gate rewritten to emit/consume the 6-widget descriptor taxonomy (D2/D4).
Tech Stack: Python ≥3.12 (typer, pydantic v2, pyyaml, sqlalchemy 2.0, psycopg2-binary, datasketch, sqlglot, pytest, testcontainers[postgres]); Node/JS for the Pi gate extension; Ollama embeddings (nomic-embed-text-v2-moe); Pi + GLM 5.2 for L2 tests. Testing: two levels (L1 fake + L2 real), see Testing Strategy below.
Reference spec: docs/superpowers/specs/2026-06-25-thothii-architecture-design.md (all D1–D16 and §4.*).
Source of ported code: /Users/mp/projects/ThothII/ChironeWp3/ (read-only reference; never modify it).
Testing Strategy — three levels (L0 testcontainers, L1 fake, L2 real)
This plan uses a three-level test strategy. The split is mandatory because the skill→LLM→tool-call loop cannot be exercised without a real model, and the DB-touching ported code needs a real database to validate (the "not assumed reliable" principle is hollow without it).
⚠️ The honest headline: the core of the system has NO automated regression coverage
The skill→LLM→gate loop — the heart of the system — is covered only by L2 (manual, non-deterministic, slow, requires credentials+VPN). There is no automated regression protection on it. This is a deliberate, conscious choice, not an accidental side-effect: the LLM is non-deterministic and requires a configured Pi + network, so it cannot live in the fast automated loop. Consequence: agentic-behavior regressions surface at pre-release L2 runs, not at commit. Accept this and plan L2 runs before any release. A partial automated net for this gap would be a fake-Pi that mocks the Pi runtime and lets the gate glue run in CI; that is documented as a cross-cutting follow-up (built alongside the backend plan, not in this plan).
L0 — testcontainers, local Postgres, runs on every local run (CI)
What: integrity tests of ported DB-touching modules against a real Postgres in a Docker container (testcontainers). No LLM, no remote network. Docker is available on the dev machine, so L0 is always-on locally.
Dependencies: Docker (present in the dev env). No credentials, no VPN.
Coverage: the ported logic that speaks to a DB — db/connection (read-only enforcement, exit code 2 if writable), db/introspect, db/sampling, db/execute, search/RRF + LSH with real data, vectorstore/store direct-load path. This is where "not assumed reliable" gains real teeth for the data layer.
L1 — fake data, deterministic, runs on every local run (CI)
What: logic-pure tests with fake data (tmp_path, fixtures, mocks). No DB, no LLM, no network.
Coverage (honest):
- Python logic-pure:
workflow.yamlloading,effective_decisions,teardown_to_phase,generate_task_doc,aggregate_lsh_multi(on fake hits), formula store read/write,decision_retracted, CLI contract tests (valid input → well-formed JSON). - Gate builder functions (pure, in JS, tested in JS): the widget-descriptor builders produce the correct JSON given params. Tested in-language (node:test/vitest), no Python↔JS bridge, no Python mirror.
Honest limitation (load-bearing): L1 can test the gate builders (pure functions) but NOT the gate glue — registration, emission via ctx.sendRaw, the no-limbo loop, the anti-bypass hooks. The glue depends on the Pi runtime (pi.registerTool, ctx.sendRaw, emitAndWait) and cannot run without either a real Pi or a fake-Pi mock. The glue is tested only at L2. Fuzzy/negative tests on the glue (malformed tool params, out-of-order calls, orphan responses) likewise live at L2; only fuzzy tests on the pure builders live in L1.
L2 — real LLM + real remote DB, manual/pre-release, NOT automated
What: end-to-end sessions with GLM 5.2 + the real Chirone datawarehouse and pgvector, reached via REST over VPN. Plus the gate-glue validation (the part L1 cannot reach).
Dependencies (all required, skip if missing):
- LLM: Pi configured locally with GLM 5.2 (already ready).
- DB: the remote Supabase endpoints (see "L2 connection params" below), reachable via VPN.
harness/.envpopulated with the API keys + CA path.
Modes:
- Human-in-the-loop (default): the test runs the harness against GLM 5.2 and the reviewer answers each gate widget via the terminal, prompted by the test. Used for exploratory validation and for exercising the gate glue (which L1 cannot).
- Pre-scripted answers (opt-in): for specific cases (free-text "Altro", value grounding "ablazione", rollback), the test supplies canned reviewer answers automatically and asserts the outcome — no human typing. Used to test exact behaviors deterministically with the real model.
Coverage (honest): validates the assumption L1 cannot — that GLM 5.2 produces tool calls the gate accepts, that the skill's prompts lead to the expected interaction shape, that the gate glue handles real tool-call sequences (incl. Altro/Rifiuta/rollback), that value grounding and formula approval surface correctly on the real schema, that memory save-one upserts to the real pgvector. Closes the skill→LLM→gate loop AND exercises the gate glue.
Honest limitation: L2 is non-deterministic (the model may behave differently across runs) and slow/costly. It is a pre-release safety net, not a regression gate. A small, curated set of scenarios.
L2 connection params (from the operator)
- DWH (read-only):
https://supabase-aritmolab.policlinicosandonato.it/dwh/— PostgREST, schemadatawarehouse+datawarehouse_marts, roledwh_reader. HeaderX-API-Key: ${THOTH_DWH_API_KEY}. - pgvector reader:
https://supabase-aritmolab.policlinicosandonato.it/vector/v1/— rolevector_reader(search_similar,list_tables). HeaderX-API-Key: ${THOTH_VEC_API_KEY}(same value asTHOTH_DWH_API_KEY— shared reader key). - pgvector writer (upsert-only): same URL as reader, different key → role
vector_writer(existing_vector_hashes,upsert_vector_records, no DELETE). HeaderX-API-Key: ${THOTH_VEC_WRITE_API_KEY}. - TLS: self-signed, CA = leaf. On the operator's Mac, a local copy of the CA bundle must be present;
${THOTH_SSL_CA}points to its local path. (Server path/etc/nginx/ssl/policlinicosandonato.it.fullchain.crtis unreachable from outside.) - Workspace:
chirone-test. - Test question (L2): "dammi la lista dei pazienti che hanno fatto un'ablazione nel 2025" — exercises D14 value grounding + formula on the "ablazione" multi-column case.
API key handling (security)
- All keys live only in
harness/.env(gitignored). Never in code, never in the plan, never committed. .env.exampleis committed with variable names and empty values.- L2 tests load
.envviapython-dotenvat startup; if any required var is missing/empty, the L2 test is skipped with a clear message (not failed), so L0/L1 never break for missing credentials. - Tests never log key values. URLs in logs are fine; secrets are masked.
- ⚠️ The operator should rotate the key that appeared in chat transcripts.
Test marker convention
- L0 tests: marker
@pytest.mark.l0. Run always locally (Docker present). Auto-skip if Docker unavailable (rare). File naming:tests/l0/test_*.py. - L1 tests: no marker (default, run always). File naming:
test_*.pyundertests/. Gate builder tests live in JS (node:test/vitest) underharness/.pi/extensions/gate/__tests__/and run vianpm testalongside pytest. - L2 tests: marker
@pytest.mark.l2. Skipped automatically when.envincomplete. File naming:tests/l2/test_*.py. Default run:pytest -m "not l2"(L0+L1 only); pre-release:pytest -m l2(L2 only).
File Structure
harness/
├── .pi/ ← Pi project config (ported + adapted)
│ ├── settings.json
│ ├── prompts/ ← /nuova-domanda, /riprendi-sessione
│ ├── skills/nsp-sessione/ ← SKILL.md + sub-files (adapted: task-doc, D13/D14)
│ └── extensions/
│ ├── nsp-gate.js ← REWRITTEN: 6-widget descriptor emit/consume (glue, verified at L2)
│ ├── reserved-labels.mjs ← ported
│ └── gate/
│ ├── builders.js ← NEW: pure widget-builder functions (L1, JS-tested)
│ └── __tests__/ ← node:test golden + fuzzy (L1)
│ ├── builders.test.js
│ └── golden/
├── nsp/ ← Python package (ported + adapted)
│ ├── __init__.py
│ ├── config.py ← ported + vector_write_rest (D11)
│ ├── workspace.py ← NEW: YAML load + ${VAR} expand (D3 confine)
│ ├── workflow.py ← NEW: reads workflow.yaml (F2)
│ ├── phase.py ← REWRITTEN: data-driven + effective_decisions (D15)
│ ├── teardown.py ← NEW: teardown_to_phase (D15)
│ ├── taskdoc.py ← NEW: per-step task document generator (D16)
│ ├── decisions.py ← ported + decision_retracted (D15)
│ ├── cli/
│ │ ├── __init__.py ← Typer app `nsp`
│ │ ├── _guards.py ← ported (require_vector_write_allowed etc.)
│ │ ├── workspace_cmd.py ← NEW: nsp workspace list/show
│ │ ├── phase_cmd.py ← adapted: + phase meta --json (F2)
│ │ ├── decision_cmd.py ← adapted: + retraction (D15)
│ │ ├── memory_cmd.py ← adapted: + save-one (D11)
│ │ ├── search_cmd.py ← adapted: + --kind formula (D14b)
│ │ ├── session_cmd.py ← adapted: + consistency check (D15)
│ │ └── (sql_cmd, cte_cmd, db_cmd, lsh_cmd, vector_cmd, evidence_cmd, schema_cmd — ported)
│ ├── db/ ← ported (connection, introspect, sampling, execute)
│ ├── rest/ ← ported (client, execute)
│ ├── search/ ← ported (combined_search, RRF) + formula retrieval (D14b)
│ ├── mschema/ ← ported (models, render, eligibility)
│ ├── vectorstore/ ← ported + dual-key rest_client (D11)
│ ├── evidence/ ← ported + formula kind (D14b)
│ ├── memory.py ← ported
│ └── session/ ← ported (models, store, artifacts)
├── workflow.yaml ← NEW: single source of workflow truth (F2)
├── workspaces/ ← workspace YAML definitions (D3)
│ └── chirone.example.yaml
├── sessions/ ← session persistence (FS, append-only ledger)
├── tests/
│ ├── conftest.py ← L2 skip-when-no-.env fixtures
│ ├── golden/ ← golden JSON for widget builders (L1)
│ ├── fixtures/
│ │ └── scenarios/ ← static .jsonl fixtures (synthesized once by build-scenarios)
│ ├── test_workflow.py ← L1
│ ├── test_phase_effective.py ← L1
│ ├── test_teardown.py ← L1
│ ├── test_taskdoc.py ← L1
│ ├── test_workspace.py ← L1
│ ├── test_vector_dual_key.py ← L1
│ ├── test_memory_save_one.py ← L1 (mocked writer)
│ ├── test_freetext_interpretation.py ← L1 (rationale contract)
│ ├── test_db_connection.py ← L0 Livello B (testcontainers: read-only enforcement)
│ ├── test_rest_client.py ← L1 Livello B (mock transport: RPC call)
│ ├── test_mschema_render.py ← L1 Livello B (3 formats on fake mschema)
│ ├── test_mschema_eligibility.py ← L1 Livello B (eligibility rules on fake columns)
│ ├── test_db_introspect.py ← L0 Livello B (testcontainers: known schema)
│ ├── test_db_sampling.py ← L0 Livello B (testcontainers: known data)
│ ├── test_rrf.py ← L0 Livello B (testcontainers: RRF fusion with real data)
│ ├── test_cli_contract.py ← L1 Livello B (each CLI: valid input → well-formed JSON)
│ ├── l0/ ← testcontainers tests (need Docker)
│ │ └── __init__.py
│ └── l2/ ← L2 (real GLM 5.2 + remote DB, @pytest.mark.l2)
│ └── __init__.py
│ └── l2/
│ ├── test_session_ablazione.py ← L2 (GLM 5.2 + real DWH, human or scripted)
│ ├── test_value_grounding_real.py ← L2 ("ablazione" on real schema)
│ └── test_memory_save_one_real.py ← L2 (real pgvector writer upsert)
├── scripts/
│ ├── create_vector_writer_rpc.sql ← ported (writer RPC allowlist)
│ └── create_vector_reader_rpc.sql ← NEW (reader RPC, D11 §5.4 note)
├── pyproject.toml
└── .env.example
Phase A — Foundation (cross-cutting infrastructure first)
Rationale: every feature depends on these. workflow.yaml + effective_decisions + teardown + taskdoc are the substrate the gate and the D11/D13/D14 features build on.
Task A1: Scaffold the harness project
Files:
-
Create:
harness/pyproject.toml -
Create:
harness/.env.example -
Create:
harness/nsp/__init__.py -
Create:
harness/.gitignore -
Step 1: Create
pyproject.toml
[project]
name = "nsp"
version = "0.1.0"
description = "ThothII harness — deterministic CLI for the NL→SQL workflow"
requires-python = ">=3.12"
dependencies = [
"typer>=0.12",
"pydantic>=2.0",
"pyyaml>=6.0",
"sqlalchemy>=2.0",
"psycopg2-binary>=2.9",
"datasketch>=1.6",
"sqlglot>=23.0",
"python-dotenv>=1.0",
"rich>=13.0",
"requests>=2.31",
"tqdm>=4.66",
]
[project.scripts]
nsp = "nsp.cli:app"
[project.optional-dependencies]
dev = [
"pytest>=8.0",
"testcontainers[postgres]>=4.0",
"ruff>=0.5",
]
[tool.ruff]
line-length = 100
[tool.pytest.ini_options]
testpaths = ["tests"]
- Step 2: Create
.env.example(port the structure fromChironeWp3/.env.example, renamed toTHOTH_*per §5.4)
# ThothII profile: server (full rebuild) | workstation (read REST + optional upsert)
THOTH_PROFILE=server
# Workspace DB (relational, direct transport) — example: chirone
THOTH_DB_HOST=
THOTH_DB_PORT=5432
THOTH_DB_USER=
THOTH_DB_PASSWORD=
# Vector REST — LETTURA (rpc search_similar) sul Supabase remoto
THOTH_VEC_REST_URL=https://host/vector/v1/
THOTH_VEC_API_KEY=
# Vector REST — SCRITTURA controllata (upsert only) da client remoti autorizzati
THOTH_VEC_WRITE_API_KEY=
# Vector direct loading (server-only)
THOTH_VEC_HOST=localhost
THOTH_VEC_PORT=5438
THOTH_VEC_USER=postgres
THOTH_VEC_PASSWORD=
# Embeddings (Ollama)
THOTH_OLLAMA_URL=http://localhost:11434
# TLS (optional, internal CA)
THOTH_SSL_CA=
- Step 3: Create
.gitignore
__pycache__/
*.pyc
.env
.venv/
sessions/
indexes/
*.egg-info/
dist/
build/
- Step 4: Create empty
nsp/__init__.pyand install
cd /Users/mp/projects/ThothII/harness
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
- Step 5: Verify
nspcommand is not yet wired (expected: import error) then commit
nsp --help 2>&1 | head -3
Expected: error (no cli module yet) — this is fine, we add it in Task A4.
git add harness/
git commit -m "feat(harness): scaffold project (pyproject, env, gitignore)"
Task A2: Port config + create workspace.py (D3 confine)
Files:
-
Create:
harness/nsp/config.py(ported + adapted fromChironeWp3/src/psdwp3/config.py) -
Create:
harness/nsp/workspace.py -
Create:
harness/workspaces/chirone.example.yaml -
Test:
harness/tests/test_workspace.py -
Step 1: Port
config.py— copyChironeWp3/src/psdwp3/config.py→harness/nsp/config.py. Rename the package prefixpsdwp3→nspeverywhere. Verify it still defines:DatabaseConfig,RestConfig,PathsConfig,EligibilityConfig,LshConfig,EvidenceSourcesConfig,EmbeddingsConfig,VectorConfig,SearchConfig,ExecutionConfig, and theConfigclass withvector_rest+vector_write_rest(bothRestConfig | None, see ChironeWp3 config.py:154-159) +_expand_env(lines 16-32). -
Step 2: Create
workspaces/chirone.example.yamlper spec §5.1 (relational + vector_db with rest/write_rest + evidence + embeddings + execution). -
Step 3: Write failing test for workspace loading
# tests/test_workspace.py
from pathlib import Path
from nsp.workspace import load_workspace, WorkspaceError
def test_load_workspace_expands_env_vars(monkeypatch, tmp_path):
monkeypatch.setenv("THOTH_VEC_API_KEY", "secret-reader")
monkeypatch.setenv("THOTH_VEC_WRITE_API_KEY", "secret-writer")
monkeypatch.setenv("THOTH_VEC_REST_URL", "https://example/vector/v1/")
yaml = tmp_path / "w.yaml"
yaml.write_text(
"name: test\n"
"relational:\n db_type: postgres\n transport: rest\n"
" rest: { base_url: 'https://dwh/', api_key: '${THOTH_VEC_API_KEY}' }\n"
"vector_db:\n collection: test_docs\n dim: 768\n"
" rest: { base_url: '${THOTH_VEC_REST_URL}', api_key: '${THOTH_VEC_API_KEY}' }\n"
" write_rest: { base_url: '${THOTH_VEC_REST_URL}', api_key: '${THOTH_VEC_WRITE_API_KEY}' }\n"
"embeddings:\n provider: ollama\n base_url: 'http://localhost:11434'\n model: m\n dim: 768\n"
)
ws = load_workspace(yaml)
assert ws.vector_db.rest.api_key == "secret-reader"
assert ws.vector_db.write_rest.api_key == "secret-writer"
def test_load_workspace_missing_env_raises(monkeypatch, tmp_path):
monkeypatch.delenv("THOTH_VEC_API_KEY", raising=False)
yaml = tmp_path / "w.yaml"
yaml.write_text(
"name: test\nrelational: { db_type: postgres, transport: rest, rest: { base_url: 'https://dwh/', api_key: '${THOTH_VEC_API_KEY}' } }\n"
"vector_db: { collection: c, dim: 768, rest: { base_url: 'https://v/', api_key: '${THOTH_VEC_API_KEY}' } }\n"
"embeddings: { provider: ollama, base_url: 'http://x', model: m, dim: 768 }\n"
)
try:
load_workspace(yaml)
assert False, "should have raised"
except WorkspaceError as e:
assert "THOTH_VEC_API_KEY" in str(e)
- Step 4: Run test, verify fail
Run: cd harness && pytest tests/test_workspace.py -v
Expected: FAIL — ModuleNotFoundError: nsp.workspace
- Step 5: Implement
workspace.py
# nsp/workspace.py
"""Workspace YAML loading — the single boundary for workspace configuration (spec D3).
Reads workspaces/<name>.yaml, expands ${VAR} from env, validates via the Config model.
Future migration to a DB store would replace only this module.
"""
from __future__ import annotations
import os
from pathlib import Path
import yaml
from nsp.config import Config, ConfigError
class WorkspaceError(Exception):
pass
def _expand_str(s: str) -> str:
"""Expand ${VAR} occurrences in s. Raises if a referenced var is unset."""
PREFIX = "${"
SUFFIX = "}"
out: list[str] = []
i = 0
while i < len(s):
start = s.find(PREFIX, i)
if start == -1:
out.append(s[i:])
break
out.append(s[i:start])
end = s.find(SUFFIX, start + len(PREFIX))
if end == -1:
raise WorkspaceError(f'Sintassi non valida (manca "}}"): {s[start:]}')
var = s[start + len(PREFIX) : end]
if var not in os.environ:
raise WorkspaceError(f"Variabile d'ambiente non definita: {var}")
out.append(os.environ[var])
i = end + len(SUFFIX)
return "".join(out)
def _expand_env(obj):
if isinstance(obj, str):
return _expand_str(obj)
if isinstance(obj, dict):
return {k: _expand_env(v) for k, v in obj.items()}
if isinstance(obj, list):
return [_expand_env(v) for v in obj]
return obj
def load_workspace(path: str | Path) -> Config:
raw = yaml.safe_load(Path(path).read_text())
try:
return Config.model_validate(_expand_env(raw))
except ConfigError as e:
raise WorkspaceError(str(e)) from e
- Step 6: Run test, verify pass
Run: pytest tests/test_workspace.py -v
Expected: 2 PASS
- Step 7: Commit
git add harness/nsp/config.py harness/nsp/workspace.py harness/workspaces/chirone.example.yaml harness/tests/test_workspace.py
git commit -m "feat(harness): port config + workspace.py YAML loader (D3)"
Task A3: Create workflow.yaml + workflow.py (F2) — single source of truth
Files:
-
Create:
harness/workflow.yaml -
Create:
harness/nsp/workflow.py -
Test:
harness/tests/test_workflow.py -
Step 1: Create
workflow.yaml(the 8-phase definition from spec §5.3)
# harness/workflow.yaml — single source of workflow truth (spec F2)
schema_version: 1
phases:
- id: F1
name: chiaramento
advance: kind:phase
prerequisites: []
artifacts_out: []
- id: F2
name: memoria
advance: auto_if_empty
prerequisites: []
artifacts_out: []
- id: F3
name: riscrittura
advance: kind:phase
prerequisites:
- decision_exists: question_rewritten
artifacts_out: [question.md]
- id: F4
name: schema_linking
advance: reviewer_decide
prerequisites: []
artifacts_out: [schema_linking.json]
- id: F5
name: sintesi
advance: kind:phase
prerequisites:
- file_validates: [schema_linking.json, SchemaLinking]
artifacts_out: []
- id: F6
name: cte
advance: auto_if_empty_or_skipped
prerequisites:
- any:
- decision_subject_exists: [phase_skipped, "phase:6"]
- all_ctes_approved: true
artifacts_out: [cte_plan.json, ctes/, cte_tests.json]
- id: F7
name: sql_finale
advance: kind:phase
prerequisites:
- decision_exists: sql_approved
artifacts_out: [sql_final.sql]
- id: F8
name: datamart
advance: reviewer_decide
prerequisites:
- any:
- decision_exists: datamart_requested
- decision_exists: datamart_declined
artifacts_out: []
decision_min_phase: auto
max_phase: auto
- Step 2: Write failing test for workflow loading
# tests/test_workflow.py
from nsp.workflow import load_workflow, PhaseSpec
def test_workflow_loads_8_phases():
wf = load_workflow()
assert len(wf.phases) == 8
assert wf.phases[0].id == "F1"
assert wf.max_phase == 8
assert wf.phase_by_num(1).name == "chiaramento"
def test_decision_min_phase_derived():
wf = load_workflow()
# question_rewritten is a prerequisite of F3 → min phase 3
assert wf.decision_min_phase("question_rewritten") == 3
# sql_approved is a prerequisite of F7 → min phase 7
assert wf.decision_min_phase("sql_approved") == 7
# unknown type → phase 1 (default)
assert wf.decision_min_phase("nonexistent_type") == 1
def test_phase_name_lookup():
wf = load_workflow()
assert wf.phase_name(6) == "cte"
assert wf.phase_name(8) == "datamart" # the JS drift bug — F8 must be present
def test_artifacts_out_per_phase():
wf = load_workflow()
assert "schema_linking.json" in wf.phase_by_num(4).artifacts_out
assert "sql_final.sql" in wf.phase_by_num(7).artifacts_out
- Step 3: Run test, verify fail
Run: pytest tests/test_workflow.py -v
Expected: FAIL — ModuleNotFoundError: nsp.workflow
- Step 4: Implement
workflow.py
# nsp/workflow.py
"""Reads workflow.yaml — the SINGLE source of workflow truth (spec F2, §5.3).
phase.py, the gate, and the skill all read from here. No more duplicated constants.
"""
from __future__ import annotations
from dataclasses import dataclass, field
from pathlib import Path
import yaml
_WF_PATH = Path(__file__).resolve().parent.parent / "workflow.yaml"
@dataclass
class PhaseSpec:
id: str
num: int
name: str
advance: str
prerequisites: list
artifacts_out: list[str] = field(default_factory=list)
@dataclass
class Workflow:
schema_version: int
phases: list[PhaseSpec]
_decision_min_map: dict[str, int] = field(default_factory=dict)
@property
def max_phase(self) -> int:
return len(self.phases)
def phase_by_num(self, n: int) -> PhaseSpec:
return self.phases[n - 1]
def phase_name(self, n: int) -> str:
return self.phase_by_num(n).name if 1 <= n <= self.max_phase else "?"
def decision_min_phase(self, decision_type: str) -> int:
# A decision type's min phase = the earliest phase whose prerequisites
# reference it (via decision_exists/decision_subject_exists), else 1.
return self._decision_min_map.get(decision_type, 1)
def _collect_decision_mins(phases: list[PhaseSpec]) -> dict[str, int]:
"""Scan prerequisites for decision_exists / decision_subject_exists mentions."""
mins: dict[str, int] = {}
def scan(node, phase_num: int):
if isinstance(node, dict):
for k, v in node.items():
if k in ("decision_exists", "decision_subject_exists"):
if isinstance(v, list):
dtype = v[0]
else:
dtype = v
if dtype not in mins or phase_num < mins[dtype]:
mins[dtype] = phase_num
else:
scan(v, phase_num)
elif isinstance(node, list):
for item in node:
scan(item, phase_num)
for p in phases:
scan(p.prerequisites, p.num)
return mins
def load_workflow(path: Path | str = _WF_PATH) -> Workflow:
raw = yaml.safe_load(Path(path).read_text())
phases = []
for i, p in enumerate(raw["phases"], start=1):
phases.append(PhaseSpec(
id=p["id"], num=i, name=p["name"],
advance=p["advance"], prerequisites=p.get("prerequisites", []),
artifacts_out=p.get("artifacts_out", []),
))
return Workflow(
schema_version=raw.get("schema_version", 1),
phases=phases,
_decision_min_map=_collect_decision_mins(phases),
)
- Step 5: Run test, verify pass
Run: pytest tests/test_workflow.py -v
Expected: 4 PASS
- Step 6: Commit
git add harness/workflow.yaml harness/nsp/workflow.py harness/tests/test_workflow.py
git commit -m "feat(harness): workflow.yaml as single source of truth + workflow.py loader (F2)"
Task A4: Port decisions.py + add decision_retracted (D15)
Files:
-
Create:
harness/nsp/decisions.py(ported + adapted) -
Test:
harness/tests/test_decisions_retract.py -
Step 1: Port
decisions.py— copyChironeWp3/src/psdwp3/session/decisions.py. Rename package. Adddecision_retractedto theDecisionTypeLiteral. TheDecisionRecordmodel gains an optionalretracts: int | Nonefield (thedecision_seqbeing retracted, for step-level rollback). -
Step 2: Write failing test for retraction
# tests/test_decisions_retract.py
from pathlib import Path
from nsp.decisions import append_decision, list_decisions, DecisionRecord
def test_retracted_decision_in_audit_but_marked(tmp_path):
session = tmp_path / "s1"
session.mkdir()
append_decision(session, DecisionRecord(seq=1, ts="t", type="table_promoted",
subject="phase:4", detail="t1", rationale="r"))
append_decision(session, DecisionRecord(seq=2, ts="t", type="decision_retracted",
subject="phase:4", detail="retract", rationale="wrong", retracts=1))
all_decisions = list_decisions(session)
assert len(all_decisions) == 2 # both in audit
assert all_decisions[1].retracts == 1
- Step 3: Run, verify fail, implement the
retractsfield + ensureappend_decisionwrites it, verify pass. (The portedappend_decisionalready writes all fields; just ensure the new field is serialized. If using pydanticmodel_dump, addretracts: int | None = None.)
Run: pytest tests/test_decisions_retract.py -v
Expected after implement: PASS
- Step 4: Commit
git add harness/nsp/decisions.py harness/tests/test_decisions_retract.py
git commit -m "feat(harness): port decisions.py + decision_retracted for step rollback (D15)"
Task A5: Rewrite phase.py — data-driven + effective_decisions (D15 core fix)
Files:
- Create:
harness/nsp/phase.py(rewritten) - Port:
harness/nsp/session/{models,store,artifacts}.py(needed for SchemaLinking validation) - Test:
harness/tests/test_phase_effective.py
This is the single most important architectural fix (spec §4.8). phase.py reads from workflow.yaml, and ALL helpers consult effective_decisions() instead of raw list_decisions().
-
Step 1: Port session models/store/artifacts — copy
ChironeWp3/src/psdwp3/session/{models.py, store.py, artifacts.py}→harness/nsp/session/. Rename package. These are needed forSchemaLinkingvalidation inadvance_problems. -
Step 2: Write failing test for effective_decisions
# tests/test_phase_effective.py
from pathlib import Path
from nsp.phase import current_phase, effective_decisions
from nsp.decisions import append_decision, DecisionRecord
def _d(session, seq, dtype, subject, **kw):
append_decision(session, DecisionRecord(seq=seq, ts="t", type=dtype, subject=subject,
detail=kw.get("detail", ""), rationale=kw.get("rationale", ""),
retracts=kw.get("retracts")))
def test_effective_decisions_excludes_pre_reopen_tail(tmp_path):
s = tmp_path / "s"; s.mkdir()
_d(s, 1, "phase_approved", "phase:1")
_d(s, 2, "phase_approved", "phase:2")
_d(s, 3, "phase_reopened", "phase:1") # rollback to phase 1
_d(s, 4, "table_promoted", "phase:4") # stale: produced after reopen but for phase 4 (not current)
# After reopen to phase:1, the effective view truncates everything after the last phase_reopened.
eff = effective_decisions(s)
# The reopen at seq 3 means decisions after it (seq 4) are NOT effective,
# and the reopen itself sets the pointer. Effective = decisions up to and incl. reopen,
# then re-walked. The stale table_promoted at seq 4 must be excluded.
types = [d.type for d in eff]
assert "table_promoted" not in types
def test_current_phase_after_reopen(tmp_path):
s = tmp_path / "s"; s.mkdir()
_d(s, 1, "phase_approved", "phase:1")
_d(s, 2, "phase_approved", "phase:2")
_d(s, 3, "phase_reopened", "phase:1")
assert current_phase(s) == 1
_d(s, 4, "phase_approved", "phase:1") # re-approve after reopen
assert current_phase(s) == 2
def test_retracted_decision_excluded_from_effective(tmp_path):
s = tmp_path / "s"; s.mkdir()
_d(s, 1, "table_promoted", "phase:4", detail="t1")
_d(s, 2, "decision_retracted", "phase:4", retracts=1)
eff = effective_decisions(s)
types = [d.type for d in eff]
assert "table_promoted" not in types # retracted → excluded
- Step 3: Run, verify fail
Run: pytest tests/test_phase_effective.py -v
Expected: FAIL
- Step 4: Implement
phase.py— the core. Port the fold logic fromChironeWp3/src/psdwp3/session/phase.pybut:- Read
max_phase,phase_name,decision_min_phase,advance_problemsrules fromworkflow.yamlviaload_workflow(). - Add
effective_decisions(session): replay the ledger; when hitting aphase_reopened phase:N, truncate everything after it in the effective view AND re-walk from phase N. When hitting adecision_retracted, exclude the retracted seq. - Rewrite
advance_problems(phase),approved_ctes,substantive_count_current_phase,auto_advance_eligibleto consulteffective_decisions()instead oflist_decisions().
- Read
Reference skeleton (the effective_decisions core — the load-bearing part):
# nsp/phase.py (key function)
from nsp.decisions import list_decisions
from nsp.workflow import load_workflow
def effective_decisions(session_dir) -> list:
"""The canonical effective view. ALL helpers MUST use this, not list_decisions().
Replays the ledger: phase_reopened truncates the tail; decision_retracted excludes the target.
"""
all_d = list_decisions(session_dir)
retracted_seqs = {d.retracts for d in all_d if d.type == "decision_retracted" and d.retracts}
# Find the LAST phase_reopened; everything strictly after it is stale.
last_reopen_idx = -1
for i, d in enumerate(all_d):
if d.type == "phase_reopened":
last_reopen_idx = i
if last_reopen_idx >= 0:
effective = all_d[: last_reopen_idx + 1]
else:
effective = list(all_d)
# Exclude retracted decisions (by seq), and the retraction records themselves.
return [d for d in effective if d.seq not in retracted_seqs and d.type != "decision_retracted"]
Implement current_phase(session_dir) as the fold over effective_decisions(session_dir) (not over raw list). Implement advance_problems(phase, session_dir) by evaluating the prerequisites list from workflow.yaml for that phase against effective_decisions. Reuse the prerequisite predicate types: decision_exists, decision_subject_exists, file_validates, all_ctes_approved, any, all.
- Step 5: Run, verify pass
Run: pytest tests/test_phase_effective.py -v
Expected: 3 PASS
- Step 6: Commit
git add harness/nsp/phase.py harness/nsp/session/ harness/tests/test_phase_effective.py
git commit -m "feat(harness): rewrite phase.py data-driven + effective_decisions (D15 core, F2)"
Task A6: Create teardown.py — teardown_to_phase (D15 teardown)
Files:
-
Create:
harness/nsp/teardown.py -
Test:
harness/tests/test_teardown.py -
Step 1: Write failing test
# tests/test_teardown.py
from pathlib import Path
from nsp.teardown import teardown_to_phase, TeardownReport
def test_teardown_to_phase_4_deletes_phase5plus_artifacts(tmp_path):
s = tmp_path / "sess"; s.mkdir()
(s / "schema_linking.json").write_text("{}") # F4 artifact
(s / "cte_plan.json").write_text("[]") # F6 artifact
(s / "ctes").mkdir()
(s / "ctes" / "x.sql").write_text("SELECT 1") # F6 artifact
(s / "sql_final.sql").write_text("SELECT 1") # F7 artifact
report = teardown_to_phase(s, target_phase=4)
assert (s / "schema_linking.json").exists() # F4 preserved (target is 4)
assert not (s / "cte_plan.json").exists() # F6 deleted (>4)
assert not (s / "ctes").exists() # F6 dir deleted
assert not (s / "sql_final.sql").exists() # F7 deleted
assert "cte_plan.json" in report.deleted_files
assert "sql_final.sql" in report.deleted_files
def test_teardown_records_deleted_in_ledger(tmp_path):
# teardown itself appends a phase_reopened decision (the rollback marker)
# — verifying that is covered by test_phase_effective; here we check the report.
s = tmp_path / "sess"; s.mkdir()
(s / "sql_final.sql").write_text("SELECT 1")
report = teardown_to_phase(s, target_phase=4)
assert len(report.deleted_files) >= 1
- Step 2: Run, verify fail
Run: pytest tests/test_teardown.py -v
- Step 3: Implement
teardown.py
# nsp/teardown.py
"""Artifact teardown on rollback (spec D15, §4.8).
Deletes every artifact whose producing phase > target, using the artifacts_out map
from workflow.yaml. Recomputes derived state. Fixes the orphaned-CTE-blocks-finalize bug.
"""
from __future__ import annotations
from dataclasses import dataclass, field
from pathlib import Path
from nsp.workflow import load_workflow
@dataclass
class TeardownReport:
target_phase: int
deleted_files: list[str] = field(default_factory=list)
def teardown_to_phase(session_dir: str | Path, target_phase: int) -> TeardownReport:
session_dir = Path(session_dir)
wf = load_workflow()
report = TeardownReport(target_phase=target_phase)
for phase in wf.phases:
if phase.num <= target_phase:
continue
for artifact in phase.artifacts_out:
target = session_dir / artifact.rstrip("/")
if artifact.endswith("/"):
# directory artifact (e.g. ctes/)
if target.exists():
for f in target.glob("*"):
f.unlink()
report.deleted_files.append(f.name)
target.rmdir()
else:
if target.exists():
target.unlink()
report.deleted_files.append(artifact)
return report
- Step 4: Run, verify pass
Run: pytest tests/test_teardown.py -v
Expected: 2 PASS
- Step 5: Commit
git add harness/nsp/teardown.py harness/tests/test_teardown.py
git commit -m "feat(harness): teardown_to_phase — artifact teardown on rollback (D15)"
Task A7: Create taskdoc.py — per-step task document generator (D16)
Files:
-
Create:
harness/nsp/taskdoc.py -
Test:
harness/tests/test_taskdoc.py -
Step 1: Write failing test
# tests/test_taskdoc.py
from pathlib import Path
from nsp.taskdoc import generate_task_doc, TaskDoc
def test_task_doc_includes_question_and_schema_scope_not_full_physical(tmp_path):
s = tmp_path / "sess"; s.mkdir()
(s / "question.md").write_text("# Domanda\nQuanti pazienti?\n## Assunzioni\n- a")
(s / "schema_linking.json").write_text('{"candidates":[{"kind":"table","name":"pazienti"}],"joins":[]}')
# A fatal-sized artifact that must NOT appear in the task doc
(s.parent / "physical.yaml").write_text("x: " + "y" * 800_000)
doc = generate_task_doc(session_dir=s, phase=7, promoted_tables=["pazienti"])
assert "Quanti pazienti?" in doc.body
assert "pazienti" in doc.body
assert len(doc.body) < 100_000 # bounded — no fatal full-schema read
assert "physical.yaml" not in doc.body # never embedded
def test_task_doc_byte_budget_enforced(tmp_path):
s = tmp_path / "sess"; s.mkdir()
(s / "question.md").write_text("q")
doc = generate_task_doc(session_dir=s, phase=1, promoted_tables=[])
assert doc.byte_budget_ok is True
- Step 2: Run, verify fail
Run: pytest tests/test_taskdoc.py -v
- Step 3: Implement
taskdoc.py
# nsp/taskdoc.py
"""Per-step task document generator (spec D16, §4.9).
Emits a single compact document per phase/step, derived from prior artifacts,
with enforced byte budget (target <20k tokens). NEVER embeds full physical.yaml.
"""
from __future__ import annotations
from dataclasses import dataclass
from pathlib import Path
MAX_BODY_BYTES = 80_000 # ~20k tokens
@dataclass
class TaskDoc:
phase: int
body: str
byte_budget_ok: bool
def generate_task_doc(session_dir: Path | str, phase: int, promoted_tables: list[str] | None = None) -> TaskDoc:
session_dir = Path(session_dir)
parts: list[str] = []
q = session_dir / "question.md"
if q.exists():
parts.append("## Domanda\n" + q.read_text())
sl = session_dir / "schema_linking.json"
if sl.exists() and phase >= 4:
parts.append("## Schema linking (deciso)\n```json\n" + sl.read_text() + "\n```")
# Task header for the phase
parts.append(f"## Task: fase {phase}")
body = "\n\n".join(parts)
return TaskDoc(phase=phase, body=body, byte_budget_ok=len(body.encode()) <= MAX_BODY_BYTES)
- Step 4: Run, verify pass
Run: pytest tests/test_taskdoc.py -v
Expected: 2 PASS
- Step 5: Commit
git add harness/nsp/taskdoc.py harness/tests/test_taskdoc.py
git commit -m "feat(harness): taskdoc per-step generator with byte budget (D16)"
Task A8: Port the CLI skeleton + nsp phase meta --json (F2)
Files:
- Create:
harness/nsp/cli/__init__.py(Typer app) - Create:
harness/nsp/cli/phase_cmd.py(adapted) - Create:
harness/nsp/cli/workspace_cmd.py - Port the command modules (sql, cte, db, lsh, vector, evidence, schema, memory, search, session, decision) — adapt imports
- Test:
harness/tests/test_cli_phase_meta.py
Honest scope: this task ports the CLI structure + phase meta and verifies that ONE command behaves. It does NOT verify the 11 ported command modules — those get contract tests in Task A9 (Livello B). Porting them here without tests would assume them reliable, which the spec forbids.
-
Step 1: Port CLI modules — copy
ChironeWp3/src/psdwp3/cli/*.py→harness/nsp/cli/. Rename package imports. The gate will neednsp phase meta --jsonto avoid mirroring constants. -
Step 2: Write failing test for
nsp phase meta --json
# tests/test_cli_phase_meta.py
import json
from typer.testing import CliRunner
from nsp.cli import app
runner = CliRunner()
def test_phase_meta_json_returns_workflow_data():
result = runner.invoke(app, ["phase", "meta", "--json"])
assert result.exit_code == 0
data = json.loads(result.stdout)
assert data["max_phase"] == 8
assert len(data["phases"]) == 8
assert data["phases"][7]["name"] == "datamart" # F8 present (fixes the JS drift)
assert "advance" in data["phases"][0]
- Step 3: Run, verify fail, then implement
phase metainphase_cmd.py
The command reads from load_workflow() and dumps {max_phase, phases: [{num,name,advance,artifacts_out}]} as JSON.
- Step 4: Run, verify pass; commit
git add harness/nsp/cli/ harness/tests/test_cli_phase_meta.py
git commit -m "feat(harness): port CLI structure + nsp phase meta --json (F2, kills JS/Python drift)"
Task A9: Contract tests for ported load-bearing modules — L0 (testcontainers) + L1 (fake data)
Rationale: spec §1 forbids assuming ported code is reliable. The DB-touching modules get L0 tests (testcontainers — real Postgres in container, where "not assumed reliable" gains real teeth); the logic modules get L1 tests (fake data). Docker is available on the dev machine, so L0 runs locally on every pytest.
L0 tests (testcontainers, marker @pytest.mark.l0):
-
Step 1:
tests/l0/test_db_connection.py— read-only enforcement (exit code 2 if writable role),pingsucceeds on read-only role. -
Step 2:
tests/l0/test_db_introspect.py+test_db_sampling.py— create a known schema + known rows in the container, verifyintrospectreturns them with columns + comments, andunique_values_for_lshreturns the expected most-frequent values. -
Step 3:
tests/l0/test_rrf.py— load real-ish data into the container (LSH index built from sampled values + a small vector set), verifycombined_search/rrf_fusefuses them with stable, sensible ranking. This is the one that truly exercises the ported RRF/LSH pipeline.
L1 tests (fake data, no marker):
-
Step 4:
tests/test_rest_client.py— mockrequests, verify one RPC call carries theX-API-Keyheader and parses the response. -
Step 5:
tests/test_mschema_render.py— the 3 formats (markdown, mschema-text ThothAI, schema-dict) from a fakePhysicalSchema. Catches port breaks in the render layer. -
Step 6:
tests/test_mschema_eligibility.py— eligibility rules on fake columns (wide-text excluded above threshold). -
Step 7:
tests/test_cli_contract.py— one test per ported CLI command (sql, cte, schema, session, decision, lsh, vector, evidence, memory, search, db): valid input → exit 0 + parseable JSON. -
Step 8: Register
l0marker inpyproject.tomland run all A9 tests.
[tool.pytest.ini_options]
testpaths = ["tests"]
markers = [
"l0: testcontainers tests (need Docker, run locally)",
"l2: end-to-end tests requiring real GLM 5.2 + remote DB (skipped when .env incomplete)",
]
addopts = "-m 'not l2'" # L0 runs by default (Docker present); L2 opt-in
cd harness && pytest tests/l0/ tests/test_rest_client.py tests/test_mschema_render.py \
tests/test_mschema_eligibility.py tests/test_cli_contract.py -v
- Step 9: Fix porting bugs the tests surface; commit
git add harness/tests/l0/ harness/tests/test_rest_client.py harness/tests/test_mschema_*.py \
harness/tests/test_cli_contract.py harness/nsp/{db,rest,mschema,search}/ harness/pyproject.toml
git commit -m "test(harness): L0 testcontainers + L1 contract tests for ported modules (spec §1)"
Phase B — Feature deviations (D11, D13, D14)
Task B1: Dual vector API key — port vectorstore + verify writer config (D11)
Files:
-
Port:
harness/nsp/vectorstore/{rest_client.py, rest_writer.py, store.py, reader.py, embeddings.py, records.py} -
Port:
harness/nsp/cli/_guards.py -
Port:
harness/scripts/create_vector_writer_rpc.sql -
Create:
harness/scripts/create_vector_reader_rpc.sql -
Test:
harness/tests/test_vector_dual_key.py -
Step 1: Port the vectorstore modules (rest_client, rest_writer, store, reader, embeddings, records) and
_guards.pyfrom ChironeWp3. These already implement the dual-key model per spec §5.4 — port them, rename package, verifyhas_vector_write_restchecks.api_key.strip(). -
Step 2: Write failing test for dual-key client construction
# tests/test_vector_dual_key.py
from nsp.config import RestConfig
from nsp.vectorstore.rest_client import VectorRestClient
def test_reader_and_writer_use_separate_keys():
reader = VectorRestClient(RestConfig(base_url="https://v/", api_key="K-READ"))
writer = VectorRestClient(RestConfig(base_url="https://v/", api_key="K-WRITE"))
assert reader.api_key == "K-READ"
assert writer.api_key == "K-WRITE"
def test_has_vector_write_rest_false_for_empty_key():
from nsp.cli._guards import has_vector_write_rest
from nsp.config import Config, RestConfig
cfg = Config(vector_write_rest=RestConfig(base_url="x", api_key=" "))
assert has_vector_write_rest(cfg) is False
-
Step 3: Run, verify fail, adjust ported code so
VectorRestClientexposesapi_key(it stores the config; add a property if missing), verify pass. -
Step 4: Create
create_vector_reader_rpc.sql— author the reader RPC (search_similar,list_tables) mirroring the writer allowlist pattern, granted to avector_readerrole. (The reader RPCs lived server-side in Supabase and were never in the ChironeWp3 repo — spec §5.4 note. Author them now.) -
Step 5: Commit
git add harness/nsp/vectorstore/ harness/nsp/cli/_guards.py harness/scripts/ harness/tests/test_vector_dual_key.py
git commit -m "feat(harness): port vectorstore dual-key + reader RPC (D11, §5.4)"
Task B2: nsp memory save-one — targeted upsert via writer key (D11)
Files:
-
Modify:
harness/nsp/cli/memory_cmd.py(addsave-one) -
Modify:
harness/nsp/memory.py(addmemory_vector_record_for_decision) -
Test:
harness/tests/test_memory_save_one.py -
Step 1: Write failing test
# tests/test_memory_save_one.py
from pathlib import Path
from unittest.mock import patch, MagicMock
from typer.testing import CliRunner
from nsp.cli import app
runner = CliRunner()
def test_save_one_calls_upsert_with_single_record(tmp_path, monkeypatch):
session = tmp_path / "s"; session.mkdir()
# build a minimal registry + a decision_seq to save
from nsp.decisions import append_decision, DecisionRecord
append_decision(session, DecisionRecord(seq=7, ts="t", type="table_promoted",
subject="phase:4", detail="pazienti", rationale="r"))
mock_writer = MagicMock()
mock_writer.upsert_records.return_value = 1
with patch("nsp.cli.memory_cmd.open_store", return_value=mock_writer):
result = runner.invoke(app, ["memory", "save-one", "--session", str(session), "--decision-seq", "7"])
assert result.exit_code == 0
mock_writer.sync.assert_not_called() # must be a single upsert, NOT full resync
mock_writer.upsert_records.assert_called_once()
args = mock_writer.upsert_records.call_args
assert len(args[0][1]) == 1 # exactly one row
def test_save_one_refused_without_writer_key(tmp_path, monkeypatch):
monkeypatch.setenv("THOTH_PROFILE", "workstation")
session = tmp_path / "s"; session.mkdir()
# config with no write_rest
result = runner.invoke(app, ["memory", "save-one", "--session", str(session), "--decision-seq", "1"])
assert result.exit_code == 4 # require_vector_write_allowed gate
-
Step 2: Run, verify fail
-
Step 3: Implement
save-oneinmemory_cmd.py:- Guard:
require_vector_write_allowed(cfg, "memory save-one"). - Load the registry, find the memory record for
decision_seq. - Build a single
VectorRecord(reusememory_vector_recordsbut filter to one). - Call
writer.existing_hashes(kinds={"memory"})+ embed (if hash changed) +writer.upsert_records(table="memory", rows=[one]). - Do NOT call
writer.sync(that's the full-resync path).
- Guard:
-
Step 4: Run, verify pass; commit
git add harness/nsp/cli/memory_cmd.py harness/nsp/memory.py harness/tests/test_memory_save_one.py
git commit -m "feat(harness): nsp memory save-one — targeted upsert via writer key (D11)"
Task B3: Value & Schema Linking — value_grounded decision + LSH multi-column (D14a)
Files:
-
Modify:
harness/nsp/decisions.py(addvalue_grounded) -
Modify:
harness/nsp/search/__init__.py(stop collapsing multi-column to single best) -
Modify:
harness/nsp/session/models.py(Candidate.grounded_values) -
Test:
harness/tests/test_value_grounding.py -
Step 1: Port
search/__init__.py,lshindex/,db/sampling.pyfrom ChironeWp3. Then modify. -
Step 2: Write failing test for multi-column LSH exposure
# tests/test_value_grounding.py
from nsp.search import aggregate_lsh_multi
def test_value_in_multiple_columns_returns_all():
hits = [
{"table": "t", "column": "c1", "value": "ablazione", "score": 0.9},
{"table": "t", "column": "c2", "value": "ablazione", "score": 0.7},
]
result = aggregate_lsh_multi(hits)
# NOT collapsed to single best — both columns exposed
cols = {h["column"] for h in result["t"]}
assert cols == {"c1", "c2"}
-
Step 3: Run, verify fail, then implement
aggregate_lsh_multi(replaces the collapsing_aggregate_lshbehavior — keep the old name as a thin wrapper to the new one if other code needs single-best). Addvalue_groundedto DecisionType. Addgrounded_values: list[dict] = []toCandidateinsession/models.py. -
Step 4: Run, verify pass; commit
git add harness/nsp/search/ harness/nsp/lshindex/ harness/nsp/db/sampling.py harness/nsp/decisions.py harness/nsp/session/models.py harness/tests/test_value_grounding.py
git commit -m "feat(harness): value grounding — multi-column LSH + value_grounded (D14a)"
Task B4: SQL formula evidence — formula kind + retrieval + approval (D14b)
Files:
-
Modify:
harness/nsp/evidence/model.py(activatetier: concept/ add formula fields) -
Create:
harness/nsp/evidence/formula_store.py -
Modify:
harness/nsp/cli/search_cmd.py(--kind formula) -
Modify:
harness/nsp/decisions.py(addconcept_formula_approved,concept_formula_rejected) -
Test:
harness/tests/test_formula.py -
Step 1: Port
evidence/modules from ChironeWp3. -
Step 2: Write failing test for formula retrieval
# tests/test_formula.py
from pathlib import Path
from nsp.evidence.formula_store import ConceptFormula, save_formula, retrieve_formula
def test_formula_retrieval_by_concept(tmp_path):
f = ConceptFormula(concept="fascia pediatrica", columns=["data_nascita"],
sql="CASE WHEN ... END", status="reviewed", sources=["s1"])
save_formula(tmp_path, f)
results = retrieve_formula(tmp_path, "fascia pediatrica")
assert len(results) == 1
assert results[0].sql.startswith("CASE WHEN")
assert results[0].concept == "fascia pediatrica"
-
Step 3: Run, verify fail, implement
formula_store.py(frontmatter YAML + SQL body, concept→formula units). Add--kind formulatosearch_cmdthat callsretrieve_formula. Add the two new decision types. -
Step 4: Run, verify pass; commit
git add harness/nsp/evidence/ harness/nsp/cli/search_cmd.py harness/nsp/decisions.py harness/tests/test_formula.py
git commit -m "feat(harness): SQL formula evidence — concept→formula units + retrieval + approval decisions (D14b)"
Task B5: Free-text interpretation guidance (D13)
Files:
-
Modify:
harness/nsp/cli/decision_cmd.py(ensure free-text from "Altro"/"Rifiuta" is captured inrationale) -
Modify:
.pi/skills/nsp-sessione/SKILL.md(+ sub-files) — port + add the D13 §4.6 guidance -
Test:
harness/tests/test_freetext_interpretation.py -
Step 1: Write failing test — a golden-style test that, given a
ui_responsewithcontrol: freetext(Altro text), verifies the harness records the user's text in the decisionrationale(not discarded).
# tests/test_freetext_interpretation.py
from pathlib import Path
from nsp.decisions import append_decision, list_decisions, DecisionRecord
def test_altro_freetext_recorded_in_rationale(tmp_path):
s = tmp_path / "sess"; s.mkdir()
# The gate, on receiving Altro text, appends a decision whose rationale carries the user's words.
append_decision(s, DecisionRecord(seq=1, ts="t", type="table_promoted",
subject="phase:4", detail="ablazione",
rationale="Utente (Altro): 'ablazione' va cercato anche in patologia, non solo nel flag"))
decisions = list_decisions(s)
assert "patologia" in decisions[0].rationale # user text preserved, not discarded
-
Step 2: Port
SKILL.md+ sub-files from ChironeWp3. Add a section "Interpretazione del testo libero" per spec §4.6: instructs the model to (1) evaluate Altro/Rifiuta/steering text in context, (2) act on it, (3) if ambiguous, re-ask instead of defaulting, (4) record the text in the decision rationale. -
Step 3: Run, verify pass (the test asserts the recording contract; the actual model behavior is enforced by the skill prose + gate, validated by fake-pi golden tests in Task C3); commit
git add harness/.pi/skills/ harness/nsp/cli/decision_cmd.py harness/tests/test_freetext_interpretation.py
git commit -m "feat(harness): free-text interpretation guidance + rationale capture (D13)"
Phase C — Gate → widget-descriptor (D2/D4)
The gate has two parts with very different testability:
- Builders (pure functions that turn tool-call params into widget-descriptor JSON) — fully testable in L1, in JS, in-language.
- Glue (Pi runtime integration:
pi.registerTool,ctx.sendRaw,emitAndWait, anti-bypass hooks, no-limbo loop) — NOT testable in L1; it depends on the Pi runtime. Tested at L2.
This phase therefore has: C1 (builders, L1, JS) + C2 (glue implementation, no L1 test — verified at L2). There is no fake-LLM driver and no build-scenarios in this plan (see cross-cutting follow-up).
Task C1: Pure widget-builder functions + JS tests (L1)
Files:
- Create:
harness/.pi/extensions/gate/builders.js(pure functions, no Pi context) - Create:
harness/.pi/extensions/gate/__tests__/builders.test.js(node:test, in-language) - Create:
harness/.pi/extensions/gate/__tests__/golden/*.json(golden widget descriptors) - Create:
harness/package.json(fornpm test→node --test)
This task extracts widget-descriptor construction into pure, testable functions and tests them in JS directly — no Python↔JS bridge, no Python mirror (a mirror would drift; the golden tests would validate the mirror, not the real gate).
- Step 1: Create
package.jsonfor the gate JS tests.
{
"name": "thothii-harness-gate",
"private": true,
"scripts": { "test": "node --test .pi/extensions/gate/__tests__/" }
}
- Step 2: Define the pure builder API in
builders.js. Each builder takes plain params and returns a plain widget-descriptor object. Noctx, no I/O.
// gate/builders.js — pure functions, no Pi context. Each returns a ui_request descriptor (spec §4.1).
function buildSelectRequest({ id, phase, title, options, intro = null, allowOther = true, recommended = null }) { /* ... */ }
function buildMultiselectRequest({ id, phase, title, options, content = null, allowEmpty = false, allowOther = true }) { /* ... */ }
function buildArtifactGate({ id, phase, title, artifact /* {kind, data, version} */, action /* {kind, prompt} */ }) { /* ... */ }
function buildInfoRequest({ phase, level, text }) { /* ... */ }
function buildFreetextRequest({ id, phase, title }) { /* ... */ }
function withChildLinkage(option, widgetSpec) { /* sets option.opens = widgetSpec; returns option */ }
module.exports = { buildSelectRequest, buildMultiselectRequest, buildArtifactGate, buildInfoRequest, buildFreetextRequest, withChildLinkage };
- Step 3: Write JS golden tests — one golden file per widget type, one test per widget. Assert exact shape.
// gate/__tests__/builders.test.js
const test = require("node:test");
const assert = require("node:assert");
const fs = require("node:fs");
const path = require("node:path");
const { buildSelectRequest, buildMultiselectRequest, buildArtifactGate } = require("../builders.js");
const GOLDEN = path.join(__dirname, "golden");
test("select F1 matches golden", () => {
const result = buildSelectRequest({ id: "u1", phase: "F1", title: "Disambigua 'ablazione'",
options: [{ id: "o1", label: "procedura" }, { id: "o2", label: "patologia" }] });
const golden = JSON.parse(fs.readFileSync(path.join(GOLDEN, "select_F1.json")));
assert.equal(result.widget, "select");
assert.deepEqual(result.reserved, golden.reserved); // back/exit/other
assert.deepEqual(result.options, golden.options);
assert.ok(result.schema_version);
});
test("artifact-gate F5 schema matches golden", () => {
const result = buildArtifactGate({ id: "u2", phase: "F5", title: "Schema-linking",
artifact: { kind: "schema_linking", data: { candidates: [] }, version: 1 },
action: { kind: "confirm", prompt: "Confermi?" } });
assert.equal(result.widget, "artifact-gate");
assert.equal(result.artifact.kind, "schema_linking");
assert.equal(result.action.kind, "confirm");
});
test("multiselect F4 carries content + allowEmpty", () => {
const result = buildMultiselectRequest({ id: "u3", phase: "F4", title: "Tabelle",
options: [], content: { candidates: [] }, allowEmpty: false });
assert.equal(result.widget, "multiselect");
assert.equal(result.allow_empty, false);
assert.ok("content" in result);
});
test("Altro option carries freetext linkage", () => {
const result = buildSelectRequest({ id: "u4", phase: "F1", title: "x", options: [] });
const other = result.options.find(o => o.id === "other");
assert.equal(other.opens.widget, "freetext");
});
- Step 4: Create the golden files, run
npm test, verify pass.
cd harness && npm test
- Step 5: Fuzzy tests on builders (L1, JS) — malformed params (missing
title, non-listoptions, emptyoptionswithallowEmpty:false). Assert builders throw clearly or return a well-formed error descriptor, never silently produce a broken widget.
test("select with missing title throws clearly", () => {
assert.throws(() => buildSelectRequest({ id: "u", phase: "F1", options: [] }), /title/);
});
test("multiselect allowEmpty:false with zero options throws", () => {
assert.throws(() => buildMultiselectRequest({ id: "u", phase: "F4", title: "x", options: [], allowEmpty: false }), /allow_empty|options/);
});
- Step 6: Commit
git add harness/package.json harness/.pi/extensions/gate/
git commit -m "feat(harness): pure widget-builder functions + JS golden/fuzzy tests (D2, L1)"
Task C2: Gate glue — rewrite nsp-gate.js (implementation; verification at L2)
Files:
- Create:
harness/.pi/extensions/nsp-gate.js(rewritten from ChironeWp3) - Port:
harness/.pi/extensions/reserved-labels.mjs - No L1 test — the glue depends on the Pi runtime and is verified at L2 (Task D4).
Honest scope: this task implements the glue but cannot unit-test it in L1. The glue's correctness is validated end-to-end at L2 with real Pi + GLM 5.2. This is the load-bearing limitation stated in the Testing Strategy headline. Do NOT attempt to mock Pi here — that is the fake-Pi follow-up (cross-cutting), not part of this plan.
-
Step 1: Port
reserved-labels.mjs(BACK/EXIT/OTHER labels,isReserved/stripReserved). -
Step 2: Port the anti-bypass hooks verbatim from the original
nsp-gate.js:tool_callhook blockingnsp phase advance|reopen,nsp decision add,nsp cte plan; protected-file list (review_decisions.jsonl,session_manifest.yaml,cte_plan.json); input-lock;before_agent_startkickoff. These are load-bearing (spec D4) and unchanged. -
Step 3: Wire tools to builders + emission. Each of the 4 tools (
reviewer_select,reviewer_decide,reviewer_confirm,rewrite_question) calls the matching builder from C1, then emits viactx.sendRaw({type:"extension_ui_request", ...})and awaits the correlated response byid. Preserve: no-limbo invariant (Esc/cancel re-presents — never returns undefined), recommended-option marker, readingSCHEMA_LINKING_PHASE/PHASE_NAMESfromnsp phase meta --json(no mirroring).
// nsp-gate.js (the glue) — wires C1 builders to the Pi runtime
const { buildSelectRequest, buildMultiselectRequest, buildArtifactGate } = require("./gate/builders.js");
module.exports = function (pi) {
pi.registerTool({
name: "reviewer_select", /* ... */,
async execute(id, params, signal, onUpdate, ctx) {
const widget = buildSelectRequest({ id, phase: await currentPhase(ctx), ...params });
const resp = await emitAndWait(ctx, widget); // extension_ui_request → wait → response
return handleSelectResponse(resp, params); // incl. Altro→freetext linkage, Back→reopen
}
});
// ... reviewer_decide, reviewer_confirm, rewrite_question, anti-bypass hooks, kickoff injection
};
-
Step 4: Manual smoke (optional, pre-L2) — launch
pi --mode rpcfromharness/once, confirm the extension loads and/nuova-domandatriggers a tool registration (no crash). This is a sanity check, not a regression test. -
Step 5: Commit (verification happens at L2 in Task D4)
git add harness/.pi/extensions/nsp-gate.js harness/.pi/extensions/reserved-labels.mjs
git commit -m "feat(harness): rewrite nsp-gate.js glue wired to builders (D2/D4) — verified at L2"
Phase D — End-to-end validation (L1 smoke + L2 real)
Task D1: L1 session-coherence smoke test (pure logic, CI)
Files:
- Create:
harness/tests/test_session_coherence_smoke.py
A pure-logic test (no LLM, no DB) that constructs a plausible ledger by hand and asserts the invariants D1-the-L2-version would check. This is the CI-runnable proxy for "full session coherence."
- Step 1: Write the test — build a synthetic ledger (decisions appended directly, not via LLM) representing a full F1→F8 walk + a rollback F6→F4 + re-derive. Assert:
current_phasefolds correctly at each step,effective_decisionsexcludes the stale tail after rollback,teardown_to_phase(4)deletes the right artifacts,generate_task_docstays under byte budget at every phase, andnsp session consistency(the post-rollback coherence check) passes.
# tests/test_session_coherence_smoke.py
from pathlib import Path
from nsp.decisions import append_decision, DecisionRecord
from nsp.phase import current_phase, effective_decisions
from nsp.teardown import teardown_to_phase
from nsp.taskdoc import generate_task_doc
def test_full_walk_then_rollback_stays_coherent(tmp_path):
s = tmp_path / "sess"; s.mkdir()
# build a synthetic full-session ledger
for n in range(1, 9):
append_decision(s, DecisionRecord(seq=n, ts="t", type="phase_approved", subject=f"phase:{n}",
detail="", rationale=""))
assert current_phase(s) == 9 # max_phase + 1
# rollback to F4
append_decision(s, DecisionRecord(seq=9, ts="t", type="phase_reopened", subject="phase:4",
detail="", rationale=""))
report = teardown_to_phase(s, target_phase=4)
assert current_phase(s) == 4
assert "sql_final.sql" in report.deleted_files or "cte_plan.json" in report.deleted_files
# task docs stay bounded across phases
for ph in range(1, 9):
doc = generate_task_doc(session_dir=s, phase=ph, promoted_tables=["dim_paziente"])
assert doc.byte_budget_ok, f"phase {ph} task doc over budget"
- Step 2: Run, verify pass; commit
git add harness/tests/test_session_coherence_smoke.py
git commit -m "test(harness): L1 session-coherence smoke (full walk + rollback, pure logic)"
Task D2: .pi/ config + prompts + theme ported
Files:
-
Create:
harness/.pi/settings.json -
Create:
harness/.pi/themes/(port theme) -
Create:
harness/.pi/prompts/nuova-domanda.md,riprendi-sessione.md -
No test (config files)
-
Step 1: Port
.pi/settings.json, theme, and the two prompt files from ChironeWp3. Verifypi --mode rpclaunched withcwd=harness/finds the.pi/directory. -
Step 2: Commit
git add harness/.pi/
git commit -m "feat(harness): port .pi/ config, prompts, theme"
Task D3: L2 conftest + skip-when-no-.env guard
Files:
-
Create:
harness/tests/conftest.py -
Create:
harness/tests/l2/__init__.py -
Modify:
harness/pyproject.toml(registerl2marker) -
Step 1: Register the
l2marker inpyproject.toml:
[tool.pytest.ini_options]
testpaths = ["tests"]
markers = [
"l2: end-to-end tests requiring real GLM 5.2 + remote DB (skipped when .env incomplete)",
]
addopts = "-m 'not l2'" # CI default: skip L2 unless explicitly requested
- Step 2: Write
conftest.pywith a session fixture that loads.envand al2_envfixture that skips if any required var is missing.
# tests/conftest.py
import os
from pathlib import Path
import pytest
from dotenv import load_dotenv
REQUIRED_L2 = ["THOTH_DWH_API_KEY", "THOTH_VEC_API_KEY", "THOTH_VEC_WRITE_API_KEY", "THOTH_SSL_CA"]
@pytest.fixture(scope="session", autouse=True)
def _load_env():
load_dotenv(Path(__file__).resolve().parent.parent / ".env")
@pytest.fixture(scope="session")
def l2_env():
missing = [v for v in REQUIRED_L2 if not os.environ.get(v, "").strip()]
if missing:
pytest.skip(f"L2 skipped — missing env vars: {', '.join(missing)} (populate harness/.env)")
return True
- Step 3: Commit
git add harness/tests/conftest.py harness/tests/l2/__init__.py harness/pyproject.toml
git commit -m "test(harness): L2 marker + skip-when-no-.env guard (Testing Strategy)"
Task D4: L2 full session — "ablazione" with GLM 5.2 (human-in-the-loop)
Files:
- Create:
harness/tests/l2/test_session_ablazione.py - Create:
harness/workspaces/chirone-test.yaml
This is the end-to-end validation that L1 cannot do: GLM 5.2 driving the real harness against the real DWH + pgvector, on the "ablazione" question (which exercises D14 value grounding + formula on a multi-column case). Default mode = human answers via terminal; scripted-answers mode for repeatability.
- Step 1: Create
workspaces/chirone-test.yamlpointing at the remote endpoints per the L2 connection params. Secrets via${THOTH_*}.
name: chirone-test
description: "Chirone DWH — L2 test workspace"
relational:
db_type: postgres
transport: rest
rest:
base_url: https://supabase-aritmolab.policlinicosandonato.it/dwh/
api_key: ${THOTH_DWH_API_KEY}
ssl_ca: ${THOTH_SSL_CA}
vector_db:
collection: chirone_docs
dim: 768
rest:
base_url: https://supabase-aritmolab.policlinicosandonato.it/vector/v1/
api_key: ${THOTH_VEC_API_KEY}
ssl_ca: ${THOTH_SSL_CA}
write_rest:
base_url: https://supabase-aritmolab.policlinicosandonato.it/vector/v1/
api_key: ${THOTH_VEC_WRITE_API_KEY}
ssl_ca: ${THOTH_SSL_CA}
evidence:
source_root: ${EVIDENCE_ROOT}
evidence_dir: evidence/chirone
embeddings:
provider: ollama
base_url: ${THOTH_OLLAMA_URL}
model: nomic-embed-text-v2-moe
dim: 768
batch_size: 64
execution:
allow: [cte_test, explain, preview, aggregate, export]
max_preview_rows: 10
statement_timeout_ms: 5000
forbidden_functions: [set_config, dblink, dblink_exec, lo_import]
- Step 2: Write the L2 test — spawns Pi (
pi --mode rpc, cwd=harness/) with GLM 5.2, runs/nuova-domanda "dammi la lista dei pazienti che hanno fatto un'ablazione nel 2025", and either (a) prompts the reviewer in the terminal for each gate widget (human mode) or (b) feeds canned answers (scripted mode via--answers-file). Asserts: a session is produced, it reaches a finalized state, thevalue_groundedandconcept_formula_approveddecisions appear (D14 exercised),sql_final.sqlis present and read-only-validates against the DWH.
# tests/l2/test_session_ablazione.py
import subprocess, json
from pathlib import Path
import pytest
@pytest.mark.l2
def test_ablazione_session_human_in_loop(l2_env, tmp_path):
"""L2: GLM 5.2 + real DWH. Reviewer answers gates via terminal.
Run with: pytest -m l2 tests/l2/test_session_ablazione.py -s
"""
session_dir = tmp_path / "sess"
proc = subprocess.run(
["pi", "--mode", "rpc"],
cwd="harness/",
# the test harness feeds /nuova-domanda and relays terminal I/O for reviewer answers
...
timeout=600,
)
assert (session_dir / "sql_final.sql").exists()
ledger = (session_dir / "review_decisions.jsonl").read_text().splitlines()
types = {json.loads(l)["type"] for l in ledger}
assert "value_grounded" in types or "concept_formula_approved" in types # D14 exercised
- Step 3: Run manually (
pytest -m l2 -s), confirm the session finalizes and D14 surfaces. This is non-deterministic and slow — it's a pre-release check, not CI. Commit the test + workspace; the operator runs it before release.
git add harness/tests/l2/test_session_ablazione.py harness/workspaces/chirone-test.yaml
git commit -m "test(harness): L2 ablazione session with GLM 5.2 + real DWH (D14, human/scripted)"
Task D5: L2 specific behaviors — value grounding (real) + memory save-one (real)
Files:
- Create:
harness/tests/l2/test_value_grounding_real.py - Create:
harness/tests/l2/test_memory_save_one_real.py
Targeted L2 tests that close specific gaps L1 leaves open: real value grounding on the live schema, and real memory save-one upsert to pgvector.
-
Step 1:
test_value_grounding_real.py— runsnsp search "ablazione" --kind values --workspace chirone-test, asserts multiple columns are returned (not collapsed), and that a grounding widget would surface them. Validates D14a on the real schema. -
Step 2:
test_memory_save_one_real.py— runsnsp memory save-oneagainstchirone-test(writer key), asserts the upsert returnsupserted >= 0(idempotent) and a subsequentsearch_similarfinds the memory. Validates D11 end-to-end. -
Step 3: Run manually (
pytest -m l2 -s tests/l2/); commit
git add harness/tests/l2/test_value_grounding_real.py harness/tests/l2/test_memory_save_one_real.py
git commit -m "test(harness): L2 value grounding (real schema) + memory save-one (real pgvector)"
Task D6: Documentation — README + workflow editing + testing guide
Files:
-
Create:
harness/README.md -
Create:
harness/docs/workflow-editing.md -
Create:
harness/docs/testing.md -
Step 1: Write
README.md— install, configure.env+workspaces/, runnsp, run L0+L1 tests, run L2 tests (with the CA-bundle copy step). -
Step 2: Write
docs/workflow-editing.md— how to editworkflow.yaml(add/reorder/merge/skip phases), referencing spec §5.3. -
Step 3: Write
docs/testing.md— explain the L0/L1/L2 split honestly: what each level covers and does NOT cover (in plain language: L0 = DB-touching code vs real Postgres in container; L1 = pure logic + gate builders, no DB no LLM; L2 = full system with real GLM 5.2 + remote DB, manual pre-release). State the headline limitation plainly: the LLM→gate conversation has no automated regression coverage. How to run each level (pytestfor L0+L1,pytest -m l2for L2 needing.env+ VPN + CA bundle). The security note on keys (.envgitignored, never logged). -
Step 4: Commit
git add harness/README.md harness/docs/
git commit -m "docs(harness): README + workflow editing + testing guide (L1/L2)"
Self-Review Notes
Spec coverage check (spec section → task):
- §1 (ChironeWp3 as starting point, NOT assumed reliable): enforced by a 3-tier porting classification:
- Tier A (modified by ThothII → full regression test): phase.py (A5), decisions.py (A4), vectorstore/_guards dual-key (B1), memory_cmd save-one (B2), search multi-column (B3), evidence formula (B4), phase_cmd meta (A8). Each has a focused L1 test for the changed behavior.
- Tier B (ported unchanged but load-bearing → contract test): db/connection + introspect + sampling + RRF (L0 testcontainers, A9 steps 1-3 — real Postgres, where "not reliable" is tested for real), rest/client + mschema/render + mschema/eligibility + 11 CLI commands (L1 fake data, A9 steps 4-7). The L0/L1 split inside Tier B reflects whether the module touches a DB.
- Tier C (ported unchanged, non-load-bearing → smoke OK): fetch_ca, output/formatting helpers. Acceptable with smoke. This makes "not assumed reliable" honest for the load-bearing surface. ✓
- D1 (3 projects): this plan = harness only; backend/frontend are separate plans. ✓
- D2/D4 (widget-descriptor gate): builders in C1 (L1, JS, in-language — golden + fuzzy); glue in C2 (implementation only — verified at L2, NOT in L1, because it depends on the Pi runtime). The limitation is documented in the Testing Strategy headline. ✓
- D3 (workspace YAML): Task A2. ✓
- D5 (FS persistence, ledger): Task A4 (decisions port), session models in A5. ✓
- D6 (auth): out of scope for harness (auth is backend). ✓ (noted)
- D7 (backend SQL read-only): out of scope for harness. ✓
- D8 (nsp stays Python): all harness tasks are Python (gate glue is JS, per the source). ✓
- D10 (testing): three levels — L0 testcontainers (A9 steps 1-3, runs locally), L1 fake/logic + gate builders (A2-A8, B1-B5, C1, D1; runs locally), L2 real GLM 5.2 + remote DB (D4, D5; pre-release). Honest headline in Testing Strategy: the skill→LLM→gate loop has NO automated regression coverage. ✓
- D11 (dual key + save-one): B1 (L1 dual-key), B2 (L1 save-one mocked), D5 (L2 save-one real). ✓
- D12 (deployment model B): harness runs locally —
.env.example(THOTH_PROFILE), L2 against remote REST over VPN. ✓ - D13 (free-text interpretation): B5 (L1 rationale contract), D4 (L2 with real model — the glue that records rationale is L2-only). ✓
- D14a (value grounding): B3 (L1 aggregate_lsh_multi on fake hits), D5 (L2 real schema). ✓
- D14b (formula): B4 (L1 formula store). ✓ (L2 formula approval rides on D4)
- D15 (rollback): A4 (retract), A5 (effective view), A6 (teardown), D1 (L1 coherence smoke), D4 (L2 glue behavior). ✓
- D16 (context minimization): A7 (taskdoc, L1). ✓ (per-phase reset in Pi noted as execution detail)
Type consistency: effective_decisions, teardown_to_phase, generate_task_doc, load_workflow, save_formula, retrieve_formula, aggregate_lsh_multi, decision_retracted, value_grounded, concept_formula_approved/rejected, builder functions (buildSelectRequest etc.) — names match across tasks. ✓
Testing honesty (the load-bearing facts):
- The skill→LLM→gate loop has NO automated regression coverage. It is exercised only at L2 (manual, non-deterministic, pre-release). This is a conscious choice; plan L2 runs before release.
- The gate glue cannot be unit-tested in L1 — it depends on the Pi runtime (
pi.registerTool,ctx.sendRaw). Only the pure builders are L1-tested (in JS, in-language). - L0 (testcontainers) tests the DB-touching ported code against a real Postgres — this is where "not assumed reliable" is actually enforced for the data layer.
- L2 tests (D4, D5) are non-deterministic, slow, require credentials + VPN + CA bundle. Skipped automatically when
.envincomplete.
Known follow-ups (not in this plan):
- Cross-cutting: a fake-Pi runtime mock that lets the gate glue run in L1/CI. When built (likely alongside the backend plan, which also needs a fake-Pi to test its RPC client), it would close the gap "gate glue has no L1 test" AND enable a
build-scenarios/fake-LLMreplay harness for deterministic gate-behavior tests. This is the single highest-value testing investment not in this plan. It is deferred because it is substantial and shared with the backend. - The SSE/REST surface the backend will build on (harness speaks JSONL via Pi +
nsp --json). session consistencycheck integration intonsp session check(Task A5 stubs the effective view; the full assertion grows during execution).- Per-phase context reset mechanism in Pi (D16 §4.9 step 3) — depends on Pi's context-management capability, verified when running the gate against real Pi in Task D4.
This plan is self-contained: at the end, harness/ runs the full 8-phase workflow emitting/consuming widget-descriptor JSON, validated by L0 (testcontainers) + L1 (logic + gate builders) on every local run, plus L2 (GLM 5.2 + real DB) pre-release. Ready for the backend to spawn and drive.