99 lines
4.9 KiB
Markdown
99 lines
4.9 KiB
Markdown
# Testing the harness — the L0 / L1 / L2 split
|
|
|
|
The harness is tested at **three levels**. The split is mandatory and honest: the
|
|
non-deterministic core (the skill→LLM→gate loop) cannot live in the fast automated
|
|
loop, and the DB-touching ported code needs a real database to validate.
|
|
|
|
## ⚠️ The honest headline
|
|
|
|
**The skill→LLM→gate loop — the heart of the system — has NO automated regression
|
|
coverage.** It is exercised **only at L2** (manual, non-deterministic, slow, requires
|
|
credentials + VPN). This is a deliberate, conscious choice: the LLM is non-deterministic
|
|
and needs a configured Pi + network, so it cannot run in CI.
|
|
|
|
**Consequence: agentic-behavior regressions surface at pre-release L2 runs, not at
|
|
commit. Accept this and run L2 before any release.**
|
|
|
|
A partial automated net for this gap would be a **fake-Pi** runtime mock that lets the
|
|
gate glue run in CI — documented as the single highest-value cross-cutting follow-up
|
|
(built alongside the backend plan, not here).
|
|
|
|
## L0 — testcontainers, real Postgres (runs locally on every `pytest`)
|
|
|
|
**What:** integrity tests of the ported DB-touching modules against a real Postgres in
|
|
a Docker container. No LLM, no remote network.
|
|
|
|
**Dependencies:** Docker (present on the dev machine). No credentials, no VPN.
|
|
|
|
**Coverage:** `db/connection` (read-only enforcement — exit 2 if writable role,
|
|
cannot INSERT), `db/introspect` (known schema: tables, columns, types, comments, FKs,
|
|
enum, composite PK), `db/sampling` (most-frequent values + truncation reporting),
|
|
`mschema/render` + `mschema/eligibility` (the column-eligibility principle), the RRF
|
|
pipeline (when the LSH index path lands). This is where "ported code is not assumed
|
|
reliable" gains real teeth for the data layer.
|
|
|
|
**Run:** `pytest` (default; auto-skips if Docker is absent). File naming:
|
|
`tests/l0/test_*.py`, marker `@pytest.mark.l0`.
|
|
|
|
## L1 — fake data, deterministic (runs locally on every `pytest`)
|
|
|
|
**What:** logic-pure tests with fake data (`tmp_path`, fixtures, mocks). No DB, no LLM,
|
|
no network.
|
|
|
|
**Coverage (honest):**
|
|
- Python logic-pure: `workflow.yaml` loading, `effective_decisions`,
|
|
`teardown_to_phase`, `aggregate_lsh_multi` (on fake hits),
|
|
`formula_store` read/write, `decision_retracted`, `save_one_memory`, the
|
|
rationale-capture contract, the session-coherence smoke, CLI `phase meta --json`.
|
|
- **Gate builder functions** (pure, in JS, tested in JS): the widget-descriptor
|
|
builders produce the correct JSON given params. Tested in-language (`node --test`),
|
|
no Python↔JS bridge, no Python mirror.
|
|
|
|
**Honest limitation (load-bearing):** L1 can test the gate **builders** (pure functions)
|
|
but **NOT the gate glue** — registration, emission via `ctx.sendRaw`, the no-limbo loop,
|
|
the anti-bypass hooks. The glue depends on the Pi runtime (`pi.registerTool`,
|
|
`ctx.sendRaw`, `emitAndWait`) and cannot run without either a real Pi or a fake-Pi mock.
|
|
**The glue is tested only at L2.**
|
|
|
|
**Run:** `pytest` (default) + `npm test` (gate builders, JS). No marker.
|
|
|
|
## L2 — real LLM + real remote DB, manual / pre-release (NOT automated)
|
|
|
|
**What:** end-to-end sessions with GLM 5.2 + the real Chirone DWH, plus the internal
|
|
Qdrant/Ollama semantic services started by the ThothII stack. Plus the gate-glue
|
|
validation (the part L1 cannot reach).
|
|
|
|
**Dependencies (all required, skip cleanly if missing):**
|
|
- LLM: Pi configured locally with GLM 5.2.
|
|
- DB: the remote DWH endpoint, via VPN when required.
|
|
- ThothII stack: internal Qdrant and Ollama services reachable from `core`.
|
|
- `harness/.env` populated with the required DWH/model API keys + CA path.
|
|
|
|
**Coverage (honest):** validates the assumption L1 cannot — that GLM 5.2 produces tool
|
|
calls the gate accepts, that the skill's prompts lead to the expected interaction shape,
|
|
that the gate glue handles real tool-call sequences (incl. Altro/Rifiuta/rollback),
|
|
that value grounding and formula approval surface correctly on the real schema, that
|
|
`memory save-one` upserts to the configured semantic store. **Closes the
|
|
skill→LLM→gate loop AND exercises the gate glue.**
|
|
|
|
**Honest limitation:** L2 is non-deterministic (the model may behave differently across
|
|
runs) and slow/costly. It is a **pre-release safety net, not a regression gate**.
|
|
|
|
**Run:** `pytest -m l2` (only; default run is `pytest -m 'not l2'`). The `l2_env`
|
|
fixture skips each L2 test (not fails) when `.env` is incomplete. File naming:
|
|
`tests/l2/test_*.py`, marker `@pytest.mark.l2`.
|
|
|
|
## How to run each level
|
|
|
|
```bash
|
|
pytest # L0 + L1 (default; addopts '-m not l2')
|
|
npm test # gate builders (JS, node --test)
|
|
pytest -m l2 # L2 only — pre-release, needs .env + VPN + CA bundle
|
|
```
|
|
|
|
## Security note on keys
|
|
|
|
All keys live **only** in `harness/.env` (gitignored). Never in code, never committed,
|
|
never logged. `.env.example` is committed with variable names and empty values. Tests
|
|
mask secrets; URLs in logs are fine. **Rotate any key that appeared in a chat transcript.**
|