Files

4.9 KiB

Testing the harness — the L0 / L1 / L2 split

The harness is tested at three levels. The split is mandatory and honest: the non-deterministic core (the skill→LLM→gate loop) cannot live in the fast automated loop, and the DB-touching ported code needs a real database to validate.

⚠️ The honest headline

The skill→LLM→gate loop — the heart of the system — has NO automated regression coverage. It is exercised only at L2 (manual, non-deterministic, slow, requires credentials + VPN). This is a deliberate, conscious choice: the LLM is non-deterministic and needs a configured Pi + network, so it cannot run in CI.

Consequence: agentic-behavior regressions surface at pre-release L2 runs, not at commit. Accept this and run L2 before any release.

A partial automated net for this gap would be a fake-Pi runtime mock that lets the gate glue run in CI — documented as the single highest-value cross-cutting follow-up (built alongside the backend plan, not here).

L0 — testcontainers, real Postgres (runs locally on every pytest)

What: integrity tests of the ported DB-touching modules against a real Postgres in a Docker container. No LLM, no remote network.

Dependencies: Docker (present on the dev machine). No credentials, no VPN.

Coverage: db/connection (read-only enforcement — exit 2 if writable role, cannot INSERT), db/introspect (known schema: tables, columns, types, comments, FKs, enum, composite PK), db/sampling (most-frequent values + truncation reporting), mschema/render + mschema/eligibility (the column-eligibility principle), the RRF pipeline (when the LSH index path lands). This is where "ported code is not assumed reliable" gains real teeth for the data layer.

Run: pytest (default; auto-skips if Docker is absent). File naming: tests/l0/test_*.py, marker @pytest.mark.l0.

L1 — fake data, deterministic (runs locally on every pytest)

What: logic-pure tests with fake data (tmp_path, fixtures, mocks). No DB, no LLM, no network.

Coverage (honest):

  • Python logic-pure: workflow.yaml loading, effective_decisions, teardown_to_phase, aggregate_lsh_multi (on fake hits), decision_retracted, save_one_memory, the rationale-capture contract, the session-coherence smoke, CLI phase meta --json.
  • Gate builder functions (pure, in JS, tested in JS): the widget-descriptor builders produce the correct JSON given params. Tested in-language (node --test), no Python↔JS bridge, no Python mirror.

Honest limitation (load-bearing): L1 can test the gate builders (pure functions) but NOT the gate glue — registration, emission via ctx.sendRaw, the no-limbo loop, the anti-bypass hooks. The glue depends on the Pi runtime (pi.registerTool, ctx.sendRaw, emitAndWait) and cannot run without either a real Pi or a fake-Pi mock. The glue is tested only at L2.

Run: pytest (default) + npm test (gate builders, JS). No marker.

L2 — real LLM + real remote DB, manual / pre-release (NOT automated)

What: end-to-end sessions with GLM 5.2 + the real Chirone DWH, plus the internal Qdrant/Ollama semantic services started by the ThothII stack. Plus the gate-glue validation (the part L1 cannot reach).

Dependencies (all required, skip cleanly if missing):

  • LLM: Pi configured locally with GLM 5.2.
  • DB: the remote DWH endpoint, via VPN when required.
  • ThothII stack: internal Qdrant and Ollama services reachable from core.
  • harness/.env populated with the required DWH/model API keys + CA path.

Coverage (honest): validates the assumption L1 cannot — that GLM 5.2 produces tool calls the gate accepts, that the skill's prompts lead to the expected interaction shape, that the gate glue handles real tool-call sequences (incl. Altro/Rifiuta/rollback), that value grounding and formula approval surface correctly on the real schema, that memory save-one upserts to the configured semantic store. Closes the skill→LLM→gate loop AND exercises the gate glue.

Honest limitation: L2 is non-deterministic (the model may behave differently across runs) and slow/costly. It is a pre-release safety net, not a regression gate.

Run: pytest -m l2 (only; default run is pytest -m 'not l2'). The l2_env fixture skips each L2 test (not fails) when .env is incomplete. File naming: tests/l2/test_*.py, marker @pytest.mark.l2.

How to run each level

pytest            # L0 + L1 (default; addopts '-m not l2')
npm test          # gate builders (JS, node --test)
pytest -m l2      # L2 only — pre-release, needs .env + VPN + CA bundle

Security note on keys

All keys live only in harness/.env (gitignored). Never in code, never committed, never logged. .env.example is committed with variable names and empty values. Tests mask secrets; URLs in logs are fine. Rotate any key that appeared in a chat transcript.