test: add independent P1.1 acceptance and manual tooling

This commit is contained in:
2026-08-11 16:04:44 +02:00
parent 930335a804
commit a22d232aa2
11 changed files with 1923 additions and 0 deletions
+89
View File
@@ -0,0 +1,89 @@
# P1.1 manual acceptance
This walkthrough is the separate human gate for the P1.1 workspace-directory registry.
It is independent from both `.artifacts/p1-integration/**` and `.artifacts/p11-integration/**`.
The helper prepares and serves the lab, but the reviewer performs the registry, Git, UI, export,
render, `tht`, refusal, secret-scan, and cleanup checks and records the verdict.
## Prerequisites
- clean repository checkout with the P1.1 implementation present;
- `node`, `npm`, `git`, `curl`, and `python3` available;
- built production assets:
```bash
npm --prefix backend run build
npm --prefix frontend run build
```
- executable harness CLI at `harness/.venv/bin/tht`;
- free loopback ports `127.0.0.1:8791` and `127.0.0.1:8792`.
## Lifecycle commands
Run from the repository root:
```bash
./scripts/p11-manual-acceptance.sh prepare
./scripts/p11-manual-acceptance.sh serve
./scripts/p11-manual-acceptance.sh stop
./scripts/p11-manual-acceptance.sh cleanup
```
The fixed lab root is:
```text
.artifacts/manual-acceptance/p11/
```
Expected lifecycle behavior:
- `prepare` creates the fixed root, ownership record, bare remote, curator clone, root catalog,
nested filesystem evidence, fixture secrets, request fixtures, generated command scripts, and
`GUIDE.md`; it leaves status `PENDING`, performs no reviewer publish operation, and never writes
`VERDICT.md`.
- `serve` starts the production backend on `127.0.0.1:8791` and a production-built frontend preview
on `127.0.0.1:8792`, recording exact ownership for both.
- `stop` refuses foreign or partial ownership and stops only the two owned loopback processes.
- `cleanup` refuses live state and removes only `.artifacts/manual-acceptance/p11/`.
## Reviewer workflow
After `prepare`, open the generated `.artifacts/manual-acceptance/p11/GUIDE.md` and personally:
1. inspect the catalog, nested descriptor/evidence layout, ownership, and secret-path bindings;
2. serve both surfaces and verify the owned listeners;
3. list `configuration_required` slots;
4. validate and bootstrap-create descriptors exactly once;
5. inspect catalog/descriptor/evidence/docs Git object IDs;
6. retry create/update/delete and verify refusal plus unchanged object IDs;
7. make a curator descriptor+catalog edit, push, pull, and verify the API did not rewrite curator bytes;
8. make an evidence-only commit and inspect the new revision identity;
9. verify the live UI shows read-only existing workspaces and bootstrap-only editing for missing slots;
10. exercise export/import under bootstrap-only rules;
11. render twice, diff the results, and run `tht config check`;
12. run negative catalog/path/secret cases and a bounded secret scan;
13. stop the lab, verify both listeners are gone, write `VERDICT.md`, and only then cleanup if desired.
## Expected outcomes
- `prepare` produces a fresh P1.1-only lab and leaves no `VERDICT.md`.
- `serve` exposes only the owned loopback backend and frontend preview.
- positive API operations succeed once; curator-owned follow-up mutations are refused safely;
- curator Git changes become active only after pull;
- renders are deterministic; `tht config check -c <file>` succeeds;
- secret scans find no canaries outside the fixture-secret boundary;
- after `stop`, nothing remains listening on `127.0.0.1:8791` or `127.0.0.1:8792`.
## Verdict format
The reviewer creates `VERDICT.md` manually. Include:
- reviewer identity;
- UTC timestamp;
- result for each checklist step;
- observations and failure evidence;
- exactly one final line: `manual acceptance: PASS` or `manual acceptance: FAIL`.
Passing `bash scripts/test-p11-manual-acceptance.sh` proves only the tooling/lifecycle guards. It
does not perform or approve manual acceptance.