docs(harness): README + workflow editing + testing guide (D6)

README: install, configure (.env + workspaces/), the workflow, run inside Pi, the
three-level test commands, layout, references.

docs/workflow-editing.md: how to edit workflow.yaml (add/reorder/merge/skip phases,
advance kinds, prerequisite predicates, decision_min_phase derivation, artifacts_out
+ teardown) -- referencing spec §5.3. Emphasizes no mirrored constants (the F2 point).

docs/testing.md: the honest L0/L1/L2 split in plain language -- what each covers and
does NOT. States the headline plainly: the skill->LLM->gate loop has NO automated
regression coverage (L2 only, pre-release). Documents the fake-Pi follow-up as the
gap-closer. Security note on keys (.env gitignored, never logged, rotate leaked keys).
This commit is contained in:
2026-06-26 23:20:06 +02:00
parent c861a0df9e
commit 50d5f9c9cf
3 changed files with 259 additions and 0 deletions
+96
View File
@@ -0,0 +1,96 @@
# Testing the harness — the L0 / L1 / L2 split
The harness is tested at **three levels**. The split is mandatory and honest: the
non-deterministic core (the skill→LLM→gate loop) cannot live in the fast automated
loop, and the DB-touching ported code needs a real database to validate.
## ⚠️ The honest headline
**The skill→LLM→gate loop — the heart of the system — has NO automated regression
coverage.** It is exercised **only at L2** (manual, non-deterministic, slow, requires
credentials + VPN). This is a deliberate, conscious choice: the LLM is non-deterministic
and needs a configured Pi + network, so it cannot run in CI.
**Consequence: agentic-behavior regressions surface at pre-release L2 runs, not at
commit. Accept this and run L2 before any release.**
A partial automated net for this gap would be a **fake-Pi** runtime mock that lets the
gate glue run in CI — documented as the single highest-value cross-cutting follow-up
(built alongside the backend plan, not here).
## L0 — testcontainers, real Postgres (runs locally on every `pytest`)
**What:** integrity tests of the ported DB-touching modules against a real Postgres in
a Docker container. No LLM, no remote network.
**Dependencies:** Docker (present on the dev machine). No credentials, no VPN.
**Coverage:** `db/connection` (read-only enforcement — exit 2 if writable role,
cannot INSERT), `db/introspect` (known schema: tables, columns, types, comments, FKs,
enum, composite PK), `db/sampling` (most-frequent values + truncation reporting),
`mschema/render` + `mschema/eligibility` (the column-eligibility principle), the RRF
pipeline (when the LSH index path lands). This is where "ported code is not assumed
reliable" gains real teeth for the data layer.
**Run:** `pytest` (default; auto-skips if Docker is absent). File naming:
`tests/l0/test_*.py`, marker `@pytest.mark.l0`.
## L1 — fake data, deterministic (runs locally on every `pytest`)
**What:** logic-pure tests with fake data (`tmp_path`, fixtures, mocks). No DB, no LLM,
no network.
**Coverage (honest):**
- Python logic-pure: `workflow.yaml` loading, `effective_decisions`,
`teardown_to_phase`, `generate_task_doc`, `aggregate_lsh_multi` (on fake hits),
`formula_store` read/write, `decision_retracted`, `save_one_memory`, the
rationale-capture contract, the session-coherence smoke, CLI `phase meta --json`.
- **Gate builder functions** (pure, in JS, tested in JS): the widget-descriptor
builders produce the correct JSON given params. Tested in-language (`node --test`),
no Python↔JS bridge, no Python mirror.
**Honest limitation (load-bearing):** L1 can test the gate **builders** (pure functions)
but **NOT the gate glue** — registration, emission via `ctx.sendRaw`, the no-limbo loop,
the anti-bypass hooks. The glue depends on the Pi runtime (`pi.registerTool`,
`ctx.sendRaw`, `emitAndWait`) and cannot run without either a real Pi or a fake-Pi mock.
**The glue is tested only at L2.**
**Run:** `pytest` (default) + `npm test` (gate builders, JS). No marker.
## L2 — real LLM + real remote DB, manual / pre-release (NOT automated)
**What:** end-to-end sessions with GLM 5.2 + the real Chirone DWH + pgvector, reached
via REST over VPN. Plus the gate-glue validation (the part L1 cannot reach).
**Dependencies (all required, skip cleanly if missing):**
- LLM: Pi configured locally with GLM 5.2.
- DB: the remote Supabase endpoints (DWH read-only + pgvector reader/writer), via VPN.
- `harness/.env` populated with the API keys + CA path.
**Coverage (honest):** validates the assumption L1 cannot — that GLM 5.2 produces tool
calls the gate accepts, that the skill's prompts lead to the expected interaction shape,
that the gate glue handles real tool-call sequences (incl. Altro/Rifiuta/rollback),
that value grounding and formula approval surface correctly on the real schema, that
`memory save-one` upserts to the real pgvector. **Closes the skill→LLM→gate loop AND
exercises the gate glue.**
**Honest limitation:** L2 is non-deterministic (the model may behave differently across
runs) and slow/costly. It is a **pre-release safety net, not a regression gate**.
**Run:** `pytest -m l2` (only; default run is `pytest -m 'not l2'`). The `l2_env`
fixture skips each L2 test (not fails) when `.env` is incomplete. File naming:
`tests/l2/test_*.py`, marker `@pytest.mark.l2`.
## How to run each level
```bash
pytest # L0 + L1 (default; addopts '-m not l2')
npm test # gate builders (JS, node --test)
pytest -m l2 # L2 only — pre-release, needs .env + VPN + CA bundle
```
## Security note on keys
All keys live **only** in `harness/.env` (gitignored). Never in code, never committed,
never logged. `.env.example` is committed with variable names and empty values. Tests
mask secrets; URLs in logs are fine. **Rotate any key that appeared in a chat transcript.**
+75
View File
@@ -0,0 +1,75 @@
# Editing the workflow
`harness/workflow.yaml` is the **single source of workflow truth** (spec F2, §5.3).
`phase.py`, the gate (`nsp-gate.js`), and the skill all read from it. There are **no
mirrored constants** in JS or Python — that was the ChironeWp3 drift bug
(`PHASE_NAMES` truncated to 7 entries in JS, F8/datamart silently dropped). Editing
this one file is the only place the workflow changes.
## Structure
```yaml
schema_version: 1
phases:
- id: F3 # stable id (referenced by the gate)
name: riscrittura # human label (nsp phase show / gate UI)
advance: kind:phase # how the phase advances (see below)
prerequisites: # what must hold before advancing (gate checks these)
- decision_exists: question_rewritten
artifacts_out: [question.md] # files produced; used by teardown on rollback (D15)
```
`decision_min_phase: auto` and `max_phase: auto` mean these are derived from the
phase list (don't hard-code them).
## `advance` kinds
| kind | meaning |
|----------------------------|----------------------------------------------------------------|
| `kind:phase` | advances on any phase-approved decision for this phase |
| `auto_if_empty` | auto-advances if no decisions were made (F2 memoria) |
| `auto_if_empty_or_skipped` | auto-advances if empty OR a `phase_skipped` decision exists |
| `reviewer_decide` | requires an explicit reviewer decision (F4, F8) |
## `prerequisites` predicates
A phase is "ready to advance" when ALL its prerequisites hold. Predicates:
- `decision_exists: <type>` — a decision of that type exists (effective view).
- `decision_subject_exists: [<type>, <subject>]` — e.g. a `phase_skipped` for `phase:6`.
- `file_validates: [<artifact>, <model>]` — e.g. `schema_linking.json` parses as a
`SchemaLinking`.
- `all_ctes_approved: true` — all CTE blocks approved (F6).
- `any: [...]` / `all: [...]` — combine predicates.
A decision type's **min phase** is derived: the earliest phase whose prerequisites
reference it. So adding a prerequisite like `decision_exists: my_new_decision` to F5
automatically makes `my_new_decision` valid from phase 5.
## Common edits
### Add a phase
Append a phase to `phases:`. `max_phase` becomes the new count automatically. Give it
an `id`, `name`, `advance`, and (optionally) `prerequisites` / `artifacts_out`.
### Reorder / merge / skip phases
Reorder the list; `max_phase` and `decision_min_phase` recompute. To make a phase
skippable, set `advance: auto_if_empty_or_skipped` and let the gate emit a
`phase_skipped` decision (predicate `decision_subject_exists: [phase_skipped, "phase:N"]`).
### Add an artifact
Add it to a phase's `artifacts_out`. On rollback (`teardown_to_phase`), artifacts
produced by phases **after** the target are deleted; the target phase's artifacts are
preserved. This is the D15 fix for the orphaned-CTE-blocks-finalize bug.
## After editing
```bash
nsp phase meta --json # confirm the new shape (max_phase, phases, artifacts_out)
pytest # L1 coherence smoke re-derives from the new workflow.yaml
```
No JS or Python constants to update — that's the point.