chore: commit remaining worktree changes

This commit is contained in:
2026-08-26 08:10:37 +02:00
parent ec061c42d4
commit f48196a57f
234 changed files with 146 additions and 61044 deletions
@@ -1,66 +0,0 @@
# Task 9 quality audit — final 5
**Scope:** the two blocking findings from `task9-quality-audit-final4.md` — unbound production
module graph at manual serve, and commit-addressed snapshots accepted without content identity at
render. Manual acceptance remains **PENDING**; no `VERDICT.md` was created.
## Verdict: APPROVED for the two final integrity blockers
### 1. Manual serve binds the complete `backend/dist` module graph, not only `server.js`
`prepare` now builds a post-build manifest of every regular `backend/dist` file
(relative path, size, SHA-256, device, inode) and writes it as an exclusive `0600` record
(`installation/runtime/backend-dist.manifest.json`) inside the owned root; `ownership.json`
records that record's path/device/inode/size/SHA-256. `serve` revalidates the manifest record
identity and bytes, revalidates every distribution file against it (no-follow, single inode,
size and digest), and refuses before spawning. The manifest descriptor is passed to the child on
fd 4 together with the entrypoint on fd 3. The immutable preload parses the manifest, verifies
the entrypoint cross-digest, reads and hash-verifies **every** file at startup, caches the
verified bytes, and its load hook serves **only** those cached bytes for any import below
`backend/dist` (entry URL still served from the bound fd-3 bytes). A same-path regular
replacement of any imported dependency is therefore refused before `RUNNING` (serve-time
validation), refused at child startup (startup verification), or rendered harmless (cached
bytes), and the parent revalidates the full manifest at `RUNNING` publication and at `stop`.
### 2. Renderer binds snapshot content to its commit identity
The generated render command validates the bounded saved read/publish revisions, the
commit-addressed owned snapshot path, the installed Git HEAD, and the bounded
`snapshot.json` manifest of that commit: `head` equals the commit, `files[<id>.yaml]` is the
SHA-256 of the snapshot bytes, the manifest revision binds commit/blob/snapshot path, the saved
revision blob equals the manifest blob, and `git rev-parse <commit>:workspaces/<id>.yaml` plus
`git hash-object` of the snapshot bytes both equal that blob. It passes the expected digest as
`--snapshot-sha256`. The renderer re-reads the bounded `snapshot.json` (`head`,
`files[<id>.yaml]` must equal the carried digest), opens the snapshot once with no-follow
semantics and bounded reads, renders only the digest-verified bytes, re-verifies around lease
publication, releases the lease in `finally`, and publishes no output on any refusal.
## Deterministic regressions added
- static regular replacement of an imported production dependency after `prepare` is refused,
no marker, no accepted PID record, no orphan;
- deterministic dependency check/load swap (`beforeSpawn` rename) is refused by the child's
startup verification, no marker, no PID record, no orphan;
- after `RUNNING`, a same-path regular dependency replacement is never executed: the loader
serves the verified cached bytes (health-visible source stays the original) and the marker is
absent;
- renderer refuses a same-path regular snapshot byte replacement against the carried digest and
manifest, with lease release and no output;
- renderer refuses manifest `head`, `files` digest, expected-digest, missing, and malformed
cases, with lease release and no output;
- wrapper refuses missing manifest, manifest head/digest/revision tampering, saved-revision blob
mismatch, Git blob mismatch, and snapshot-vs-Git-bytes mismatch, and passes the exact
`--snapshot-sha256` on the valid path (stub renderer records arguments).
## Verification
- `bash scripts/test-p1-manual-acceptance.sh` (backend build + both suites): **59 tests, 59
pass, 0 fail**; no `8791/8792` listener and no `--p1-manual-nonce` process remain.
- `npx tsc --noEmit -p .` (backend): PASS.
- Real-repository `prepare` + `cleanup` cycle: 39 distribution files bound, entrypoint
cross-digest verified, owned root fully removed afterwards.
- Diff check: only the seven Task 9 paths are touched; no Task 8 file was modified.
- This report and the implementation contain no fixture secret or canary values.
Manual acceptance remains **PENDING** by design; the walkthrough and human verdict are
unchanged.
-282
View File
@@ -1,282 +0,0 @@
{
"schema": "thothii-task4-certification-v1",
"generated_on": "2026-08-18",
"started_at_utc": "2026-08-18T14:16:40Z",
"ended_at_utc": "2026-08-18T14:20:10Z",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"source_immutability": {
"status": "PASS",
"tracked_changes_after_freeze": false,
"allowed_untracked": [".playwright-cli/", ".thothctl/"]
},
"source_commits": {
"task4_candidate": "b31b27e5845ffd3adf311429367319beaba263c7",
"task1": "d43738eeae6d14bb5e470093058b069a983f5372",
"task2": "5f9a3ae066a060b43a11a959b60a1efadd1c2425",
"task3": "0d8e707533fada938c99eb06f8457150e7ef2b40",
"task3_follow_up": "b31b27e5845ffd3adf311429367319beaba263c7",
"fix_round_1_source": "10cd66fe6a5b484a4dc569326a228c1c5484a5d4",
"fix_round_2_source": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"historical_task15_final": "74b062f1a737103524cbe706346cfd65f87cdfd1"
},
"versions": {
"node_contract": "v24.16.0",
"node_host_default": "v25.6.1",
"go": "go1.26.5",
"pi": "0.80.3"
},
"retained_report": ".superpowers/sdd/2026-08-16-thothii-authentication/task-15-report.md",
"task4_report": ".superpowers/sdd/2026-08-18-thothii-authentication-remediation/task-4-report.md",
"fix_round_2_report": ".superpowers/sdd/2026-08-18-thothii-authentication-remediation/fix-round-2-report.md",
"workflow": {
"run_id": "32147345625",
"url": "https://github.com/mptyl/ThothII/actions/runs/32147345625",
"event": "workflow_dispatch",
"head_sha": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"status": "completed",
"conclusion": "failure",
"windows_job": {
"name": "Windows clone and Compose contract",
"job_id": "95744249248",
"url": "https://github.com/mptyl/ThothII/actions/runs/32147345625/job/95744249248",
"conclusion": "failure",
"native_step": "Run native Windows retained-capability tests",
"native_step_conclusion": "success",
"command": "go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1",
"requested_packages": ["internal/safeio", "internal/backup", "internal/authstorage"],
"executed_packages": ["internal/safeio", "internal/backup", "internal/authstorage"],
"not_executed_packages": [],
"package_results": {
"internal/safeio": "PASS (22.058s)",
"internal/backup": "PASS (7.161s)",
"internal/authstorage": "PASS (16.088s)"
},
"failed_step": "Verify Windows clone contract",
"failure_category": "baseline_powershell_parser",
"failure_detail": "scripts/test-windows-clone-contract.ps1:208 parses $remoteYaml: as an invalid variable reference"
},
"lf_compose_docs_typescript_job": {
"name": "LF, Compose, docs, and TypeScript",
"job_id": "95744249458",
"url": "https://github.com/mptyl/ThothII/actions/runs/32147345625/job/95744249458",
"conclusion": "failure",
"failed_step": "Verify Compose and installation contracts",
"category": "baseline_ci_contract",
"detail": "unified Compose contract passed; test-no-deployment-coupling-scope.sh stopped on TMPDIR: unbound variable",
"downstream_steps": "skipped"
},
"linux_docker_job": {
"name": "Linux Docker deployment and rollback",
"job_id": "95744249354",
"url": "https://github.com/mptyl/ThothII/actions/runs/32147345625/job/95744249354",
"conclusion": "failure",
"failed_step": "Run unified deployment smoke",
"category": "infrastructure_prerequisite",
"detail": "Task 13 smoke failed before deployment because rg is required",
"cleanup": "PASS",
"image_manifest": "not_generated"
},
"windows_docker_startup_job": {
"name": "Native Windows Docker Desktop/WSL2 startup",
"job_id": "95744250450",
"url": "https://github.com/mptyl/ThothII/actions/runs/32147345625/job/95744250450",
"status": "NOT_RUN",
"classification": "BLOCKED",
"workflow_conclusion": "skipped",
"reason": "workflow conditions skipped the job; no Windows Docker Desktop/WSL2 command executed"
}
},
"docker_image_evidence": {
"authentication_smoke": {
"status": "PASS",
"docker_images": [],
"reason": "no_docker_images_exercised"
},
"unified_docker_smoke": {
"status": "FAIL",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"run_id": "32147345625",
"workflow_job_id": "95744249354",
"manifest": ".artifacts/task-15/unified-docker-images.json",
"reason": "workflow attempt stopped before deployment because rg is required",
"cleanup": "PASS",
"images": 0,
"historical": {
"status": "PASS",
"source_commit": "74b062f1a737103524cbe706346cfd65f87cdfd1",
"run_id": "20260818070637-66409-30058",
"manifest_sha256": "9c8dec4546909fd93799dbcf374bcb3a89bc46cfe0fd482472c0cbe757ddf5b6",
"images": 5,
"cleanup": "PASS"
}
}
},
"gates": {
"posix_registry_ownership": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"evidence": "backend Node 24 full suite including local-registry ownership coverage"
},
"stagearchive_unix_retained_capability": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"evidence": "focused safeio/backup tests, Go race suite, and Unix ancestor-swap coverage"
},
"windows_stagearchive_retained_capability": {
"status": "PASS",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"evidence": "native Windows backup package passed, including the two-file shared retained-root staging test"
},
"windows_claim_retained_capability": {
"status": "PASS",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"evidence": "native Windows safeio and authstorage packages passed concurrent claim/consume coverage"
},
"workflow_lf_compose_docs_typescript": {
"status": "FAIL",
"classification": "baseline_ci_contract",
"reason": "TMPDIR was unset after the unified Compose contract passed"
},
"workflow_linux_docker": {
"status": "FAIL",
"classification": "infrastructure_prerequisite",
"reason": "runner did not provide rg; cleanup proof passed and no image manifest was generated"
},
"go_security_build": {
"status": "PASS",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"focused_packages": 3,
"race_packages": 18,
"focused_test": "PASS",
"race": "PASS",
"vet": "PASS",
"host_build": "PASS"
},
"windows_cross_compile": {
"status": "PASS",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"focused_test_packages": 3,
"cli_build": "PASS",
"execution": "cross_compile_only_not_native_execution"
},
"backend_node24": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"node": "v24.16.0",
"files": 76,
"tests": 1092,
"typecheck": "PASS",
"build": "PASS",
"note": "an initial full run had one workspace-registry timeout; focused rerun and complete rerun passed"
},
"frontend_node24": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"node": "v24.16.0",
"files": 61,
"tests": 444,
"typecheck": "PASS",
"build": "PASS"
},
"authentication_and_f1_smoke": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"node": "v24.16.0",
"filtered_e2e": "1 passed",
"sentinel_leak_scan": "PASS"
},
"harness_pytest": {
"status": "FAIL",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"passed": 951,
"failed": 1,
"skipped": 4,
"subtests": 232,
"failure": "test_column_decisions::test_f4_emits_column_types: workflow.yaml not found from harness test cwd"
},
"authentication_docs": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7"
},
"shell_syntax": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7"
},
"authentication_smoke_runtime": {
"status": "PASS",
"node": "v24.16.0",
"sentinel_leak_scan": "PASS"
},
"compose_default": {
"status": "FAIL",
"reason": "required THT_WORKSPACE_GIT_REMOTE was not available"
},
"compose_unified": {
"status": "FAIL",
"reason": "compose.unified.yaml is absent from the frozen source"
},
"unified_docker_smoke": {
"status": "FAIL",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"workflow_run_id": "32147345625",
"reason": "remote workflow attempted the smoke but stopped before deployment because rg is required",
"cleanup": "PASS",
"image_manifest": "not_generated"
},
"ruff": {
"status": "FAIL",
"errors": 192,
"classification": "known_baseline"
},
"mkdocs_strict": {
"status": "NOT_RUN",
"classification": "BLOCKED",
"historical_status": "FAIL",
"historical_warnings": 69
},
"canonical_install_docs": {
"status": "NOT_RUN",
"classification": "BLOCKED",
"historical_status": "FAIL"
},
"workspace_install_docs": {
"status": "NOT_RUN",
"classification": "BLOCKED",
"historical_status": "FAIL"
},
"pi_user_auth_compose": {
"status": "NOT_RUN",
"classification": "BLOCKED",
"historical_status": "FAIL"
},
"deployment_coupling": {
"status": "NOT_RUN",
"classification": "BLOCKED",
"historical_status": "FAIL"
},
"l2": {
"status": "PENDING",
"reason": "configured secret layout unavailable; gate not run after stop"
},
"manual_psd": {
"status": "PENDING",
"reason": "approved real identity/access unavailable; gate not run after stop"
},
"provider_readiness": {
"status": "PENDING",
"reason": "provider prerequisite unavailable; gate not run after stop"
}
},
"review": {
"original_important_findings_resolved": 3,
"fix_round_2_important_lifecycle": "ADDRESSED",
"fix_round_2_minor_windows_diagnostics": "ADDRESSED",
"verdict": "PASS",
"reason": "the lifecycle controller is bounded and cancellation-aware with cancel, bounded join, and lock-release proof; the temporary Windows diagnostic matrix is removed; exact-source native safeio, backup, and authstorage all pass"
},
"remediation_status": "PASS",
"release_complete": false,
"authentication_implementation_complete": true,
"release_readiness": "FAIL",
"release_readiness_pending_external_gates": true
}
@@ -1,54 +0,0 @@
{
"gate": "unified-deployment-smoke",
"status": "pass",
"source_commit": "74b062f1a737103524cbe706346cfd65f87cdfd1",
"run_id": "20260818070637-66409-30058",
"images": [
{
"id": "sha256:2d7b19491c7eb8c119c3cedb390aaeb2ff5593f6fc43ab66c317565560da6d7d",
"roles": [
"compose-runtime",
"fixture-runtime"
],
"repo_digests": [
"sha256:2d7b19491c7eb8c119c3cedb390aaeb2ff5593f6fc43ab66c317565560da6d7d"
]
},
{
"id": "sha256:3b6c31a5d8f8fc58fa3233391b6175bd2fbc793eebb44d5e285ecc6e02e9e687",
"roles": [
"compose-runtime"
],
"repo_digests": [
"sha256:3b6c31a5d8f8fc58fa3233391b6175bd2fbc793eebb44d5e285ecc6e02e9e687"
]
},
{
"id": "sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a",
"roles": [
"compose-runtime"
],
"repo_digests": [
"sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a"
]
},
{
"id": "sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c",
"roles": [
"compose-runtime"
],
"repo_digests": [
"sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c"
]
},
{
"id": "sha256:c3cbe1cc1aa588a64951ac6286e0df7b27fe2e6324b1001c619bb358770c0178",
"roles": [
"rollback-candidate"
],
"repo_digests": [
"sha256:c3cbe1cc1aa588a64951ac6286e0df7b27fe2e6324b1001c619bb358770c0178"
]
}
]
}
-11
View File
@@ -1,11 +0,0 @@
{
"version": "0.0.1",
"configurations": [
{
"name": "replay",
"runtimeExecutable": "node",
"runtimeArgs": ["tools/replay/server.mjs"],
"port": 5333
}
]
}
-2
View File
@@ -27,5 +27,3 @@ coverage/
data/
sessions/
workspace-registry/
# docs/site (mkdocs build) — non necessari nelle immagini
docs/superpowers/plans
-1
View File
@@ -50,7 +50,6 @@ jobs:
bash scripts/test-no-deployment-coupling-scope.sh
bash scripts/test-compose-secret-policy.sh
bash scripts/test-no-deployment-coupling.sh
bash scripts/test-preprocess-compose-config.sh
bash scripts/test-verify-workspace-install-docs.sh
git diff --check
- name: Assert clean checkout before release trust bootstrap
-2
View File
@@ -5,8 +5,6 @@
ChironeWp3/
Thoth/
# === Visual companion brainstorming artifacts (local-only) ===
.superpowers/
.worktrees/
.tht/
-7
View File
@@ -1,7 +0,0 @@
{
"$schema": "https://app.kilo.ai/config.json",
"indexing": {
"vectorStore": "qdrant",
"model": "sentence-transformers/all-minilm-l12-v2"
}
}
@@ -1,105 +0,0 @@
# Task 3 — Diagnostic contract remediation report
Date: 2026-08-04
## Scope
This remediation is limited to the four approved review findings for the workspace diagnostic
extension. It does not add registry routes, change workspace publication, alter session startup,
or expand transport support.
## Changes
1. `RuntimeBindings` now has an explicit `vectorWriter` binding. The new
`resolveRuntimeBindings()` resolves DWH, vector reader, vector writer, and embedding bindings
together. The diagnoser takes the writer credential only from `bindings.vectorWriter`, never
from vector-reader values.
2. Direct PostgreSQL and SSH-tunnelled direct probes accept an absent CA binding while retaining
certificate verification through the runtime system trust store. A supplied CA still uses
verified private-CA trust. REST private-CA refusal is unchanged.
3. A reversible vector probe now requires an authenticated POST declaration with a response map
containing `operation`. The adapter requires the successful JSON response to echo `create` or
`remove` respectively, so an arbitrary 2xx or an upsert-only response cannot activate the
write probe.
4. For DWH and vector REST diagnostics declared with `auth: none`, the resolver no longer
requires an API-key file and the adapter sends no credential. Credential-backed diagnostics
continue to require their local secret file.
## TDD evidence
The first focused RED run failed for the intended missing behavior:
- `resolveRuntimeBindings is not a function` for unauthenticated resolver bindings;
- schema accepted a reversible probe without a response contract; and
- existing diagnostic fixtures rejected the new `response` declaration until schema support was
implemented.
The focused GREEN run passed `43/43` tests across:
- `test/workspaces-bindings.test.ts`
- `test/workspaces-schema.test.ts`
- `test/workspaces-diagnostics.test.ts`
The regression coverage includes resolver-to-diagnoser writer propagation without manually
inserting the writer key into vector-reader bindings, no-CA direct/SSH system-trust requests,
operation-echo validation for create/remove, and `auth: none` bindings without secret files.
## Documentation and design
- `docs/workspace-diagnostic-protocol.md` now documents the verified system-trust fallback,
no-secret `auth: none` behavior, and required reversible response contract.
- `docs/superpowers/specs/2026-08-03-git-workspace-registry-design.md` now records the same
response, CA, SSH, and authentication rules.
## Final verification
The initial sandboxed full suite could not bind its local SSE listener (`listen EPERM:
operation not permitted 127.0.0.1`). It was rerun unchanged with local-listener permission.
```text
backend: npx vitest run
31 test files passed; 329 tests passed
backend: npx tsc --noEmit -p .
exit 0
repository: git diff --check
exit 0
```
Expected test harness stderr from existing Pi/process failure-path tests remained present; no test
failed and no diagnostic secret was emitted.
## Blockers
None.
## Round 2 remediation
The final review found two remaining contract gaps. The binding resolver already treated
`auth: none` as credential-free, but the runtime renderer and diagnostic connector still required
the API-key file. Rendering and connector construction now make that requirement conditional on
the declared REST authentication mode, so a DWH/vector `auth: none` workspace passes resolver,
runtime rendering, and diagnostics with no API-key file.
SSH forwarding previously changed the PostgreSQL connection host to `127.0.0.1` without retaining
the original target for TLS hostname validation. Forwarded probes now carry `SSH_TARGET_HOST` as
`tlsServername` into the PostgreSQL TLS options; private CA and verified system trust behavior are
unchanged.
TDD RED: the new end-to-end no-key test failed at the unconditional runtime
`API_KEY_FILE` requirement, while the SSH test showed no `tlsServername` on the loopback probe or
database-client request. TDD GREEN: the focused backend workspace tests passed `40/40`.
Round 2 final verification:
```text
backend: npx vitest run
31 test files passed; 332 tests passed
backend: npx tsc --noEmit -p .
exit 0
repository: git diff --check
exit 0
```
@@ -1,101 +0,0 @@
# Task 7 report — revision-pinned sessions
## Delivered
- New-session requests may carry `workspaceId`, provider, model, and thinking. The backend
resolves the active operational registry revision, enforces its LLM policy, and persists the
workspace ID/revision with the selected LLM settings.
- The harness manifest and `tht session new` support the optional, backward-compatible
`workspace_id` and `workspace_revision` fields.
- Resume resolves the manifest's retained snapshot, including after later registry publication.
A missing retained revision returns a sanitized `workspace_revision_unavailable` response.
Legacy manifests retain the prior workspace behavior and are marked with a visible warning on
`GET /sessions/:id`.
- `/settings` is now a non-mutating compatibility endpoint: installation defaults remain
readable, while anonymous workspace/provider/model/thinking selections are no longer written
to backend settings or principal preferences.
## TDD evidence
- RED: `npx vitest run test/routes-sessions.test.ts test/routes-settings.test.ts` failed for the
new immutable-snapshot and no-settings-mutation assertions; the manifest test failed because
`new_session_manifest` did not accept workspace revision fields.
- GREEN: `npx vitest run test/tht-runner.test.ts test/routes-sessions.test.ts test/routes-settings.test.ts && npx tsc --noEmit -p .`
completed with 97 passing tests and a clean type check.
- GREEN: `THT_HOME=/private/tmp/thothii-task7-home .venv/bin/pytest tests/test_session_documents.py tests/test_session_mutations.py -q`
completed with 22 passing tests.
- `git diff --check` completed cleanly.
## Review fixes — round 3
- The active registry snapshot that located a session now remains the authorization and mutation
config for response, steer, events, close/delete, archive/group/rename, documents, and detail.
A pruned historical revision cannot block an already-located session's active lifecycle.
- Only Resume resolves the retained pinned descriptor because Pi needs that immutable config to
restart safely. A pruned pin therefore returns the existing sanitized
`workspace_revision_unavailable` 409 solely for Resume.
### Round 3 verification
- RED: with a manifest found through an active registry snapshot and `readPinned` forced to fail,
`POST /sessions/:id/response` returned 409 instead of forwarding the active gate response.
- GREEN: `npx vitest run test/routes-sessions.test.ts test/tht-runner.test.ts test/routes-settings.test.ts && npx tsc --noEmit -p .`
— 102 tests passed with a clean type check. The regression confirms response, close, and delete
use the locating snapshot without calling `readPinned`, while Resume returns a sanitized 409.
- `git diff --check` completed cleanly.
## Review fixes — round 2
- Lifecycle authorization no longer selects the installation-default workspace. The backend now
finds each session by querying every operational registry snapshot with the authenticated
principal, preserving RLS ownership concealment.
- After locating the manifest, durable pinned sessions resolve their retained descriptor before
any lifecycle mutation/reopen. Legacy sessions continue using the locating registry snapshot.
- Session listing aggregates the owner-visible rows from all operational registry snapshots;
detail, response, steer, resume, events, documents, and lifecycle mutations use the same
server-side locator. No route depends on browser-local workspace state.
### Round 2 verification
- RED: the new cross-workspace route integration test created a B session while installation
default A was selected, then demonstrated that `GET /sessions` returned an empty list.
- GREEN: `npx vitest run test/routes-sessions.test.ts test/tht-runner.test.ts test/routes-settings.test.ts && npx tsc --noEmit -p .`
— 101 tests passed with a clean type check. The integration test covers create B, list, detail,
response, and resume through B's pinned descriptor while default A remains configured.
- Full backend suite: 342 tests passed. The remaining 7 tests require binding `127.0.0.1` and
fail in this sandbox with `listen EPERM: operation not permitted`; no application assertion
failed. The focused typecheck above passed.
- `git diff --check` completed cleanly.
## Verification note
The unscoped backend suite was also run. The Task 7 code regressions in `test/tht-runner.test.ts`
were fixed; the remaining failures were existing sandbox restrictions on tests that listen on
`127.0.0.1` (`listen EPERM: operation not permitted` in SSE/e2e health tests), not application
assertions.
## Review fixes — round 1
- Every new session now resolves `workspaceId` through the registry; an omitted value uses the
configured installation default and persists both the resolved ID and revision. Callers cannot
bypass revision pinning by supplying a workspace ID.
- Browser-local preferences now migrate once from the read-only legacy settings response and hold
workspace, provider, model, and thinking. Session creation includes those selections, including
direct entry points that run before the composer mounts. The frontend no longer `PUT`s shared
settings.
- The settings compatibility endpoint honors a stored installation workspace before falling back
to the first workspace configuration.
- Resume rejects finalized and archived sessions before looking up any pinned snapshot, preserving
the read-only response even when a historical snapshot is unavailable.
### Review verification
- RED: the added backend tests failed for omitted-default pinning, read-only resume ordering, and
stored-default precedence; the added frontend preference tests failed because preferences were
neither stored nor included in session requests.
- GREEN: `npx vitest run test/tht-runner.test.ts test/routes-sessions.test.ts test/routes-settings.test.ts && npx tsc --noEmit -p .`
— 100 tests passed with a clean type check.
- GREEN: `npx vitest run && npx tsc -b` — 332 frontend tests passed with a clean type check.
- GREEN: `THT_HOME=/private/tmp/thothii-task7-home .venv/bin/pytest tests/test_session_documents.py tests/test_session_mutations.py -q`
— 22 tests passed (one existing testcontainers deprecation warning).
- `git diff --check` completed cleanly.
@@ -1,73 +0,0 @@
# Task 9 report — Workspace Management CRUD page
## Delivered
- Added the Workspace management dialog, launched from the persistent right sidebar and the
Model activity header without touching live-session/SSE state.
- Added a workspace list/detail editor for General, DWH, Semantic index, LLM policy,
Installation requirements, and Git status/history.
- Added browser-only New, Edit, Duplicate, Save draft, and Delete-draft workflows. A deletion
draft stores only ID and immutable revision references; publication remains a Task 10 action.
- Used closed native controls for languages, engines, transports, distance metrics, embedding
providers, and selectable default models. Free values have client-side, accessible errors.
- Made semantic-index dimensions atomic: one editor field always writes the same value to the
vector-store and embedding contracts.
- Added Validate and Test-on-this-installation actions. They display sanitized code/message
diagnostics only; neither action exposes or stores credentials, secrets, or raw response bodies.
- Explicitly excluded publish, pull, import, and export user flows from this task.
## TDD evidence
- RED: `npx vitest run src/shell/WorkspaceManager.test.tsx src/shell/WorkspaceEditor.test.tsx`
failed because the manager and editor modules did not exist.
- GREEN: focused manager/editor/AppShell coverage passed after the implementation.
- RED: a deletion-draft persistence regression failed with
`Cannot read properties of undefined (reading 'save')` before the sanitized draft store was added.
- GREEN: the draft-store and manager tests passed once deletion intent persisted locally.
## Verification
Executed from `frontend/`:
```text
npx vitest run
50 test files passed, 358 tests passed
npx tsc -b
exit 0
```
`git diff --check` passed before commit. No workspace secret value, secret-file path, raw
diagnostic body, publish call, import flow, or export flow was introduced.
## Fix round 1
### Root causes and fixes
- The original duplicate proposal appended `-copy` and then truncated at 63 characters. For an
already-maximal ID, truncation could remove the suffix and reproduce the immutable source ID.
The proposal now reserves suffix space and falls back to a distinct `-2` suffix when a maximal
source already ends in `-copy`.
- `dwh.timeout_ms` was rendered as a positive numeric field but was absent from the client
validation map. It now has the same immediate accessible error treatment as other numeric
fields, so a rejected save never reaches the manager’s saved-draft toast.
- Registry status, workspace list, and selected-detail React Query failures were rendered as
loading, empty, or unselected states. Each now has a named alert and a retry control, distinct
from its corresponding loading and empty state.
### TDD evidence
- RED: max-length duplication retained the original 63-character ID; the timeout field produced
no alert; and each of the three failed queries had no accessible retry control.
- GREEN: the focused manager/editor tests passed **12/12**, covering a valid changed duplicate
proposal, rejected zero timeout with no save toast, and status/list/detail retry recovery.
### Verification
Executed from `frontend/`:
```text
npx vitest run
50 test files passed, 364 tests passed
npx tsc -b
exit 0
```
@@ -1,180 +0,0 @@
# Task 11 report
Status: completed on 2026-08-08.
## Scope delivered
- Updated operator-facing documentation for the internal Qdrant + Ollama architecture.
- Tightened documentation contract tests to require the current four-service-plus-init topology,
CPU-first/GPU-override guidance, fixed internal model/dimensions, schema-v3 migration wording,
one-collection-per-workspace ownership, and Qdrant backup/restore safety.
- Updated stable repo guidance in `AGENTS.md` and the current snapshot in `PROJECT_STATE.md`.
- Rewrote the workspace diagnostic protocol to the schema-v3/internal-semantic-service contract.
- Updated the memory guide to describe Qdrant as the derived persistent index.
- Updated the runtime secret-bundle guide to remove active vector/embedding secret guidance.
## Files changed
- `README.md`
- `AGENTS.md`
- `PROJECT_STATE.md`
- `docs/install/local-workspace-registry.md`
- `docs/install/server-workspace-registry.md`
- `docs/installazione-docker-4-contesti.md`
- `docs/workspace-diagnostic-protocol.md`
- `docs/gestione-memory.md`
- `deploy/secrets/README.md`
- `scripts/verify-workspace-install-docs.sh`
- `scripts/test-verify-workspace-install-docs.sh`
## Verification
Fresh successful runs:
```sh
./scripts/test-verify-workspace-install-docs.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
git diff --check
```
Key outcomes:
- internal semantic infrastructure documentation contract passed
- all existing install/manual fixture contracts still passed
- diff hygiene passed with no whitespace/errors
## Self-review notes
- The updated docs now match the code-backed Compose topology: `frontend`, `core`, `qdrant`,
`embedding`, and `embedding-model-init`.
- Active manuals no longer instruct operators to configure external vector or embedding runtime
endpoints/secrets.
- Qdrant backup/restore wording now matches the helper scripts' exact confirmation and rollback
behavior.
- Legacy descriptor handling is documented as explicit schema-v3 migration only; no silent
semantic-data migration is claimed.
## Residual concerns
- The broader repository still contains historical design/spec material that references older
pgvector/external-embedding architecture; this task intentionally updated operator/current-state
documentation and the corresponding contract tests, not historical planning documents.
## Fix round 1/5 — 2026-08-08
Addressed reviewer findings:
- Moved superseded rollout/state blocks in `PROJECT_STATE.md` behind an explicit
`## Historical snapshots and archived reference notes` boundary.
- Renamed superseded snapshot headings so historical notes no longer present as active `LIVE`
state.
- Added a current-state regression that rejects contradictory active blocks (for example:
schema-v2 operational, two-service active stack, or external vector/embedding runtime claims
before the historical boundary).
- Refactored new internal-semantic doc checks away from exact-sentence coupling:
- parse `compose.yaml` structurally with YAML;
- parse workspace examples structurally with YAML;
- inspect backup/restore stable usage interface;
- keep targeted forbidden-term checks for active docs while allowing historical sections;
- use regex/concept checks for prose.
Evidence:
```sh
./scripts/test-verify-workspace-install-docs.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
git diff --check
```
Observed RED before the fix:
```text
PROJECT_STATE.md: missing Historical snapshots boundary
```
## Fix round 2/5 — 2026-08-08
Addressed reviewer findings:
- Renamed every historical `PROJECT_STATE.md` heading after the historical boundary so no heading
level uses `LIVE` or current-state semantics there.
- Strengthened the historical-boundary regression to reject any Markdown heading level
(`#` through `######`) containing `LIVE` or current-state wording after the boundary.
- Added a fixture with a `### ... — LIVE ...` historical heading to prove RED then GREEN.
- Replaced remaining exact phrase checks with concept/semantic validation for:
- one-workspace/one-collection ownership;
- external boundary (DWH/LLM external; vector/embedding internal);
- the Italian compact install note.
- Added paraphrase fixtures that pass and omission/inversion fixtures that fail.
Evidence:
```sh
./scripts/test-verify-workspace-install-docs.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
git diff --check
```
## Fix round 4/5 — 2026-08-08
Addressed reviewer finding:
- Eliminated semantic-index verifier/test contract drift by extracting the production
semantic-index ownership row matcher into `semantic_index_relationship_spec` and reusing it in
the fixture-level paraphrase, omission, and scattered-token checks.
- Kept the relationship constrained to one structured Markdown table row via
`verify_markdown_table_relationships`; the scattered-token fixture still removes the row and
appends the same words outside the table, where it must be rejected.
- Added a direct regression that copies the repository docs into an isolated root, applies the
accepted paraphrase “A workspace keeps exactly one Qdrant collection reserved for itself”, and
runs that root's actual `scripts/verify-workspace-install-docs.sh --fixtures-only` instead of a
separate temporary spec.
Observed RED before the fix:
```text
production verifier rejected the accepted semantic-index paraphrase
local workspace manual: missing relationship in 'Semantic index ownership contract': {'scope': 'workspace semantic index', 'ownership rule': '(each|one|single).*(workspace).*(single|one).*(Qdrant).*(collection)|(each workspace reserves a single qdrant collection)', 'isolation rule': 'schema.*evidence.*memory.*(one|that).*(collection).*(kind|payload)'}
```
Evidence:
```sh
./scripts/test-verify-workspace-install-docs.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
git diff --check
```
Observed RED during this round:
```text
PROJECT_STATE.md: historical section still contains active/live heading markers
compact manual paraphrase lacks required pattern: (esterni solo|solo esterni|restano esterni)
```
## Fix round 3/5 — 2026-08-08
Addressed reviewer findings:
- Added table-driven historical-heading fixtures for every Markdown heading level `#` through
`######`; all are rejected after the historical boundary when they contain `LIVE`/current-state
semantics.
- Added small structured ownership tables to the active local/server manuals and to the compact
Italian operator note.
- Added small structured semantic-index ownership tables to the active local/server manuals.
- Replaced the remaining scattered-token relationship checks with explicit structured-section
parsing:
- architecture ownership rows map DWH → external, LLM → external, Qdrant → internal,
Ollama embedding → internal;
- semantic-index ownership rows localize the one-workspace/one-collection contract and the
schema/Evidence/Memory isolation rule.
- Added adversarial fixtures that fail when the same tokens are merely scattered in free text.
- Added structured paraphrase fixtures that pass and omission/inversion fixtures that fail.
Evidence:
```sh
./scripts/test-verify-workspace-install-docs.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
git diff --check
```
@@ -1,43 +0,0 @@
# Task 12 Report — Remove unreachable pgvector runtime code
Status: completed
Summary:
- Proved the retired pgvector runtime had no remaining operational adapter call sites after migration by re-running the required grep; only the packaging assertion still mentions `migrations/vector`.
- Removed the obsolete pgvector/HTTP/direct vector runtime modules, vector SQL migrations, and their affected runtime tests.
- Kept the operational semantic path on Qdrant and migrated the remaining runtime callers to that path.
- Kept `psycopg2-binary` because DWH direct PostgreSQL and session PostgreSQL code still depend on it.
Implementation notes:
- Extracted shared collection/kind validation into `harness/tht/adapters/vector/_shared.py` so `QdrantVectorStore` no longer depends on the deleted pgvector module.
- Simplified `build_vector_store()` to return only `QdrantVectorStore`.
- Migrated vector/evidence/memory CLI paths away from legacy pgvector loaders and REST vector clients.
- Updated packaging coverage so the built wheel asserts session SQL migrations are present and vector SQL migrations are absent.
Verification:
- `cd harness && .venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py tests/test_semantic_kind_isolation.py tests/test_vector_migration_packaging.py -q`
- `cd harness && .venv/bin/pytest tests/test_adapter_factory.py tests/test_solved_search_cli.py -q`
- `cd harness && .venv/bin/python -c "import tht.cli, tht.adapters.factory, tht.adapters.vector, tht.vectorstore.reader"`
- `cd harness && uv build`
- `harness/.venv/bin/ruff check harness/tests/test_adapter_factory.py harness/tests/test_solved_search_cli.py harness/tests/test_vector_migration_packaging.py harness/tests/test_vector_port_contract.py harness/tht/adapters/factory.py harness/tht/adapters/vector/__init__.py harness/tht/adapters/vector/_shared.py harness/tht/adapters/vector/qdrant.py harness/tht/cli/evidence_cmd.py harness/tht/cli/memory_cmd.py harness/tht/cli/search_cmd.py harness/tht/cli/vector_cmd.py harness/tht/solved.py harness/tht/vectorstore/reader.py`
- `git diff --check`
Notes / concerns:
- Repository-wide `harness/.venv/bin/ruff check .` still reports many pre-existing findings outside this task’s touched files; it is not clean on this branch baseline.
- Some legacy config compatibility parsing still exists outside the deleted runtime path. This task removed the unreachable runtime/migration code without broad config-schema refactoring.
## Fix round 1 evidence
Changes:
- Removed dead `vector migrate` registration from `harness/tht/cli/__init__.py` and deleted `harness/tht/cli/vector_migrate_cmd.py`.
- Added CLI regressions proving `vector migrate` is absent while `vector init` and `vector index-schema` remain available.
- Restored the accidentally removed non-vector regressions by moving report coverage into `harness/tests/test_report.py` and restoring the taskdoc promoted-table slicing check in `harness/tests/test_taskdoc.py`.
- Reworded surviving active help/docstrings away from pgvector-specific wording in the touched Qdrant-backed command surface.
Verification:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py tests/test_report.py tests/test_taskdoc.py tests/test_vector_migration_packaging.py -q`
- `cd harness && .venv/bin/python -c "from typer.testing import CliRunner; from tht.cli import app; r=CliRunner().invoke(app, ['vector','--help']); assert r.exit_code == 0, r.output; assert 'migrate' not in r.output; r=CliRunner().invoke(app, ['vector','migrate','--help']); assert r.exit_code != 0, r.output; print('cli-help-ok')"`
- `cd harness && .venv/bin/python -c "import tht.cli, tht.cli.vector_cmd, tht.report, tht.taskdoc; print('imports-ok')"`
- `cd harness && uv build`
- `harness/.venv/bin/ruff check harness/tests/test_qdrant_cli_commands.py harness/tests/test_report.py harness/tests/test_taskdoc.py harness/tests/test_vector_migration_packaging.py harness/tht/cli/__init__.py harness/tht/cli/search_cmd.py harness/tht/cli/vector_cmd.py harness/tht/cli/memory_cmd.py harness/tht/solved.py`
- `git diff --check`
@@ -1,175 +0,0 @@
# Task 13 Implementation Report
## Status
DONE_WITH_CONCERNS
## Changes
- Updated stale harness/backend/frontend tests and fixtures to the Task 13 internal Qdrant/Ollama contract.
- Made `deploy/workspaces/psd.yaml.example` generic while preserving schema-v3 Qdrant/Ollama shape.
- Fixed `scripts/workspace-registry-smoke.sh` to pass the required legacy migration `--collection` and prove exact Docker cleanup, including its smoke image.
- Updated `PROJECT_STATE.md` with only evidence observed in this run.
Changed files:
- `PROJECT_STATE.md`
- `backend/test/routes-workspaces.test.ts`
- `backend/test/workspace-runtime-handoff.test.ts`
- `backend/test/workspaces-contracts.test.ts`
- `backend/test/workspaces-git-repository.test.ts`
- `deploy/workspaces/psd.yaml.example`
- `frontend/src/shell/NewSessionDialog.test.tsx`
- `harness/tests/test_adapter_command_regressions.py`
- `harness/tests/test_workspace.py`
- `scripts/task13-runtime-fixture-check.ts`
- `scripts/test-verify-workspace-install-docs.sh`
- `scripts/workspace-registry-smoke.sh`
## Verification
Deterministic gates:
- `cd harness && .venv/bin/pytest -q && .venv/bin/ruff check .`
- Initial red: 2 harness pytest failures.
- After fixture fixes: harness pytest passed `819 passed, 4 deselected, 74 warnings in 27.73s`.
- Ruff still failed with `Found 220 errors`; treated as existing unrelated debt.
- Touched harness files verified clean with `cd harness && .venv/bin/ruff check tests/test_adapter_command_regressions.py tests/test_workspace.py && .venv/bin/pytest -q tests/test_adapter_command_regressions.py::test_solved_index_writes_through_writer_only_factory_store tests/test_workspace.py::test_load_workspace_expands_env_vars`: `All checks passed!` and `2 passed, 2 warnings in 0.14s`.
- `cd backend && npx vitest run && npx tsc --noEmit -p . && npm run build`
- Initial red: 4 backend Vitest failures.
- After fixes: `Test Files 39 passed (39)`, `Tests 464 passed (464)`, TypeScript passed, build passed.
- `cd frontend && npx vitest run && npx tsc -b && npm run build`
- Initial red: 1 frontend Vitest failure.
- After fix: frontend Vitest passed `374/374`, TypeScript passed, build passed with Vite `built in 6.55s`.
- `git diff --check`
- Passed with no output.
Focused reruns:
- `cd backend && npx vitest run test/workspaces-migrate-legacy.test.ts test/workspaces-contracts.test.ts test/routes-workspaces.test.ts test/workspace-runtime-handoff.test.ts test/workspaces-git-repository.test.ts && cd .. && ./scripts/test-no-deployment-coupling.sh && ./scripts/verify-workspace-install-docs.sh --fixtures-only && git diff --check`
- `Test Files 5 passed (5)`, `Tests 35 passed (35)`.
- Coupling guard passed: `no active retired deployment or external semantic coupling found.`
- Install docs fixtures passed through `relative secret-source fixture rejected passed`.
Deployment contracts:
- `./scripts/test-default-compose.sh && ./scripts/test-unified-compose.sh && ./scripts/test-internal-semantic-compose.sh && ./scripts/test-no-deployment-coupling.sh && ./scripts/test-compose-secret-policy.sh && ./scripts/verify-workspace-install-docs.sh --fixtures-only`
- Passed. Output included:
- `default Compose contract passed.`
- `unified Compose contract passed.`
- `internal semantic Compose/script contracts passed.`
- `no active retired deployment or external semantic coupling found.`
- `Compose secret policy passed.`
- install-doc fixture checks through `relative secret-source fixture rejected passed`.
Docker smokes:
- `/usr/bin/time -p ./scripts/internal-semantic-smoke.sh`
- Passed: `Task 13 internal semantic smoke passed.`
- Cleanup proof: `no labeled containers, volumes, networks, or images remain for 20260808200245-83368-17823.`
- Duration: `real 217.34`.
- `/usr/bin/time -p ./scripts/workspace-registry-smoke.sh`
- Initial red: `usage: migrate-legacy --input <legacy-workspace.yaml> --output <repository-root> --collection <qdrant-collection> [--id <workspace-id>]`.
- After fix: `workspace registry smoke passed`.
- Cleanup proof: `no compose containers, volumes, networks, or image remain for thoth-workspace-registry-smoke-89671.`
- Duration: `real 9.93`.
- `/usr/bin/time -p ./scripts/unified-deployment-smoke.sh`
- Passed: `Task 13 full deployment smoke passed.`
- Cleanup proof: `no labeled containers, volumes, networks, or images remain for 20260808200706-85638-13391.`
- Duration: `real 125.57`.
- `/usr/bin/time -p ./scripts/thothctl-update-smoke.sh`
- Passed: `Task 13 update deployment smoke passed.`
- Cleanup proof: `no labeled containers, volumes, networks, or images remain for 20260808200918-87340-10404.`
- Duration: `real 85.40`.
- `/usr/bin/time -p ./scripts/server-deployment-smoke.sh`
- Passed: `Task 13 Linux server deployment smoke passed.`
- Cleanup proof: `no labeled containers, volumes, networks, or images remain for 20260808201047-88645-20675.`
- Duration: `real 55.99`.
Final audit:
- `rg -n "pgvector|local-vector|THT_VECTOR_|EMBEDDING_BASE_URL|openai_compatible|ollama_compatible" . --glob '!docs/plans/**' --glob '!docs/superpowers/**' --glob '!**/node_modules/**' --glob '!**/.venv/**' --glob '!**/.git/**'`
- Returned matches in legacy schema-v1/v2 support, migration tests, negative guards, historical notes, and older harness docs/code.
- This remains a concern: the audit is not clean under the brief's strict expected outcome.
- `git status --short`
- Before report/commit, contained only intentional Task 13 changes.
## Image and Host Evidence
- Host CPU: `Apple M4 Pro`.
- Host OS: `Darwin MacProM4-di-Marco.local 25.5.0 Darwin Kernel Version 25.5.0: Tue Jun 9 22:28:34 PDT 2026; root:xnu-12377.121.10~1/RELEASE_ARM64_T6041 arm64`.
- Docker server: `29.6.2 linux/arm64`.
- Verified pinned images:
- `qdrant/qdrant:v1.18.2@sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c`.
- `ollama/ollama:0.32.0@sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a`.
- Workspace registry smoke ephemeral image:
- Manifest list: `sha256:4d056bf2cb38d0e8ede91fbf121df1f9f18caee0d401581618ccef9ed8a55e73`.
- Config: `sha256:613f8fb28c0517adee4085f41bc447f2c3813b0fdbb7b26624bfb4cb192b6fd8`.
- Removed during cleanup.
## Manual Gates
- GPU exposure gate (`THOTH_ENABLE_EMBEDDING_GPU=1` on Linux): not executed in this run.
- Windows Docker Desktop startup/manual job: not executed in this run.
## Commits
- `4e810af` (`test: align qdrant ollama verification fixtures`)
- `7c09b98` (`docs: record qdrant ollama verification`)
## Known Limitations
- Broad harness Ruff remains existing unrelated debt: `Found 220 errors`.
- Final active-reference audit is not clean; it still finds legacy/negative-guard references outside explicit migration fixture files.
- Ephemeral Task 13 core/frontend image IDs from `internal-semantic-smoke.sh`, `unified-deployment-smoke.sh`, `thothctl-update-smoke.sh`, and `server-deployment-smoke.sh` were removed by exact cleanup and were not emitted in stdout; pinned Qdrant/Ollama digests and the workspace-registry smoke image digest were captured.
## Fix Round 1 — reviewer findings
Status: DONE
Changes:
- `scripts/workspace-registry-smoke.sh` now derives the smoke image reference from the already unique Compose project instead of using the global tag `thothii-workspace-registry-smoke:local`.
- The workspace-registry cleanup helpers remove and verify only the exact per-run image reference, plus Compose resources labeled with the exact project.
- Added deterministic self-test coverage in `backend/test/workspaces-migrate-legacy.test.ts` via `WORKSPACE_REGISTRY_SMOKE_SELF_TEST=image-cleanup-identity`; it stubs Docker and fails if cleanup touches same-repository foreign tags such as `:local` or another project tag.
- Updated active harness/testing/PRD docs and Python comments that still described the current semantic store as pgvector/vectordb. Preserved schema-v1/v2 and harness legacy compatibility fixtures.
- Updated `PROJECT_STATE.md` with fix-round smoke evidence and a precise, non-overclaiming audit limitation.
Focused verification:
- `cd backend && npx vitest run test/workspaces-migrate-legacy.test.ts`
- Passed: `7 passed`.
- `cd harness && .venv/bin/pytest -q tests/test_memory_save_one.py tests/test_adapter_command_regressions.py tests/test_solved_search_cli.py tests/test_search_pack.py`
- Passed: `22 passed, 14 warnings`.
- `cd harness && .venv/bin/ruff check tht/memory.py tht/search/__init__.py tht/workspace.py tht/vectorstore/store.py tests/test_memory_save_one.py tests/test_adapter_command_regressions.py tests/test_solved_search_cli.py`
- Passed: `All checks passed!`
- `bash -n scripts/workspace-registry-smoke.sh && WORKSPACE_REGISTRY_SMOKE_SELF_TEST=image-cleanup-identity bash scripts/workspace-registry-smoke.sh`
- Passed: `workspace registry smoke image cleanup identity self-test passed`.
- `./scripts/test-no-deployment-coupling.sh`
- Passed: `no active retired deployment or external semantic coupling found.`
- `./scripts/verify-workspace-install-docs.sh --fixtures-only`
- Passed through `relative secret-source fixture rejected passed`.
- `cd backend && npx tsc --noEmit -p .`
- Passed with no output.
- `/usr/bin/time -p ./scripts/workspace-registry-smoke.sh`
- Passed: `workspace registry smoke passed`.
- Built exact per-run tag: `thothii-workspace-registry-smoke:thoth-workspace-registry-smoke-thoth-workspace-registry-smoke-10vi3a-19157`.
- Manifest list: `sha256:715b943057929418cad4aa71806d9edbaf823555d19bda6b875297617463fd4a`.
- Config: `sha256:a566521981e08958aae9a12bfc7803bb5f3f835536b4bb8c39df8fcf26063161`.
- Cleanup proof: `no compose containers, volumes, networks, or image remain for thoth-workspace-registry-smoke-thoth-workspace-registry-smoke-10vi3a-19157.`
- Duration: `real 42.06`.
Fix-round audit command:
- `rg -n "pgvector|local-vector|THT_VECTOR_|EMBEDDING_BASE_URL|openai_compatible|ollama_compatible" . --glob '!docs/plans/**' --glob '!docs/superpowers/**' --glob '!**/node_modules/**' --glob '!**/.venv/**' --glob '!**/.git/**'`
Categorized remaining hits:
- Backend legacy parser/migration compatibility, kept deliberately non-operational for schema-v1/v2 descriptors: `backend/src/workspaces/schema.ts`, `types.ts`, `migrate-legacy.ts`, `runtime-renderer.ts`, `bindings.ts`, `contracts.ts`, `diagnostics.ts`.
- Backend negative guards and legacy fixture tests: `backend/test/workspaces-schema.test.ts`, `workspaces-migrate-v2-qdrant.test.ts`, `workspace-registry.test.ts`, `workspace-runtime-renderer.test.ts`, `workspaces-bindings.test.ts`, `workspaces-contracts.test.ts`, `workspaces-diagnostics.test.ts`, `workspaces-git-repository.test.ts`, `routes-workspaces.test.ts`, `routes-sessions.test.ts`, `provider-credentials.test.ts`.
- Secret/env scrub guards for retired variables: `backend/src/config.ts`, `backend/src/config/secret-bundle.ts`, `backend/src/pi/provider-credentials.ts`, `scripts/compose-with-preflight.sh`, `scripts/test-external-compose-lifecycle.sh`.
- Deployment negative guards and fixture-scope tests: `scripts/test-no-deployment-coupling.sh`, `scripts/test-no-deployment-coupling-scope.sh`, `scripts/test-preprocess-compose-config.sh`, `scripts/test-verify-workspace-install-docs.sh`, `scripts/verify-workspace-install-docs.sh`, `scripts/vector-rotate-bootstrap-password.sh`.
- Harness legacy config compatibility and fixtures: `harness/tht/config.py`, `harness/tht/config_compat.py`, `harness/tests/test_config_resources.py`, `harness/tests/l2/test_session_ablazione.py`, `harness/workspaces/tht.example.yaml`, `harness/workspaces/tht-test.yaml`.
- Retained off-repository migration SQL fixtures: `harness/scripts/create_vector_reader_rpc.sql`, `harness/scripts/create_vector_writer_rpc.sql`.
- Historical/reference notes, not active operator contracts: `brain/codebase/datamart-builder-deployment-gotchas.md`, `PROJECT_STATE.md`.
- Gitignored task report self-reference: `.superpowers/sdd/2026-08-08-internal-qdrant-ollama/task-13-implementation.md`.
@@ -1,165 +0,0 @@
Task 2 report — Make collection ownership unique in the Git registry
Summary
- Implemented unique Qdrant collection ownership enforcement during registry snapshot activation.
- Registry session revision leases now reject `migration_required` descriptors.
- Legacy migration now requires an explicit target collection and emits schema v3 descriptors.
- Preserved active snapshot rollback behavior on invalid pulled snapshots.
RED evidence
Focused RED command from the brief:
```bash
cd backend
npx vitest run test/workspace-registry.test.ts test/workspaces-migrate-legacy.test.ts \
-t "collection|migration_required"
```
Observed failures before implementation:
- `rejects duplicate schema v3 collection ownership and keeps the previous active snapshot`
- `registry.pull()` resolved instead of rejecting.
- `does not acquire a session revision lease for a migration_required workspace`
- `acquireSessionRevision()` resolved instead of rejecting.
- `migrates a legacy descriptor only with an explicit target collection into schema v3`
- received schema version `1` instead of `3`.
- `requires an explicit target collection for legacy migration`
- migration did not throw without a collection.
GREEN evidence
Focused GREEN command from the brief:
```bash
cd backend
npx vitest run test/workspace-registry.test.ts test/workspaces-migrate-legacy.test.ts \
-t "collection|migration_required"
```
Fresh result after implementation:
- 2 files passed
- 4 tests passed
- 0 failures
Additional verification run after final cleanup:
```bash
cd backend
npx vitest run test/routes-workspaces.test.ts
npx vitest run
npx tsc --noEmit -p .
git diff --check
```
Fresh results:
- `test/routes-workspaces.test.ts`: 7 passed
- full backend Vitest: 39 files passed, 454 tests passed
- backend typecheck: passed
- `git diff --check`: passed
Changed files
- `backend/src/workspaces/registry.ts`
- `backend/src/workspaces/migrate-legacy.ts`
- `backend/test/workspace-registry.test.ts`
- `backend/test/workspaces-migrate-legacy.test.ts`
- `backend/test/routes-workspaces.test.ts`
Why one extra file changed
- `backend/test/routes-workspaces.test.ts` needed updating because Task 1 made schema v3 the only operational descriptor shape, and the route test still assumed the old pre-Task-3 runtime behavior. Updating that expectation was necessary to keep the required backend suite verification meaningful.
Implementation notes
- Duplicate collection detection is enforced only for operational schema v3 descriptors by tracking `collection -> workspaceId` during activation.
- Duplicate failures are sanitized back to `workspace_invalid` / `Workspace repository content is invalid`.
- `acquireSessionRevision()` now fails closed for `migration_required` revisions.
- Legacy migration CLI now requires `--collection <qdrant-collection>`.
- Legacy migration output is schema v3 with the fixed internal semantic contract:
- `vector_store.engine = qdrant`
- explicit `collection`
- embedding provider `ollama_internal`
- embedding model `qwen3-embedding:0.6b`
self-review
- Confirmed invalid pulled snapshots do not replace the previous active snapshot.
- Confirmed duplicate collection enforcement does not affect legacy migration-required descriptors.
- Confirmed create/update publication tests still pass with unique per-workspace collections.
- Confirmed no JSON stdout contract regressions in the migration CLI.
- Kept runtime/data mutation scope descriptor-only; no user workspace repo or Qdrant data changes.
Concerns
- No code concerns remaining for Task 2.
- One deliberate scope exception: a route test was updated to align with the already-established Task 1 / Task 3 fail-closed contract.
Fix round 1
Scope
- Restored meaningful route-level diagnoser coverage without reopening schema-v3 semantic runtime paths.
- Added direct schema-v2 registry coverage for `migration_required` listing and lease rejection.
Covering test files
- `backend/test/routes-workspaces.test.ts`
- `backend/test/workspace-registry.test.ts`
RED command and output
Command:
```bash
cd backend
npx vitest run test/routes-workspaces.test.ts test/workspace-registry.test.ts
```
Observed result on top of `76bc94d` after adding the restored/new assertions:
- 2 files passed
- 37 tests passed
- 0 failures
Why no RED appeared:
- The review items exposed missing/weakened coverage, not a production behavior bug.
- `/workspaces/:id/test` already reaches the diagnoser for resolvable legacy v2 descriptors.
- Schema-v3 `/workspaces/:id/test` already fails closed before diagnoser entry.
- Schema-v2 descriptors were already listed as `migration_required` and already rejected by `acquireSessionRevision()`.
GREEN command and output
Command:
```bash
cd backend
npx vitest run test/routes-workspaces.test.ts test/workspace-registry.test.ts
npx tsc --noEmit -p .
```
Fresh results:
- covering tests: 2 files passed, 37 tests passed
- backend typecheck: passed
Changed files
- `backend/test/routes-workspaces.test.ts`
- `backend/test/workspace-registry.test.ts`
- `.superpowers/sdd/2026-08-08-internal-qdrant-ollama/task-2-report.md`
What changed
- Split route coverage so `POST /workspaces/validate` still checks canonical validation independently.
- Restored route-level diagnoser coverage through a migration-required schema-v2 descriptor with resolvable legacy bindings.
- Added an explicit schema-v3 fail-closed regression for `POST /workspaces/:id/test`.
- Added a direct schema-v2 registry regression proving `list()` returns `migration_required` and `acquireSessionRevision()` rejects it.
Concerns
- No production concerns. This round only tightened coverage and corrected the weakened test expectation.
@@ -1,132 +0,0 @@
# Task 3 report — Remove external semantic bindings and render internal endpoints
Date: 2026-08-08
## Scope
Implemented backend-owned schema-v3 semantic runtime rendering so workspace descriptors and installation contracts remain free of external Qdrant/Ollama endpoints and credentials, while DWH bindings stay unchanged.
## RED evidence
Focused RED command:
`cd backend && npx vitest run test/workspaces-contracts.test.ts test/workspaces-bindings.test.ts test/workspace-runtime-renderer.test.ts test/config.test.ts`
Observed failures before implementation:
- `config.test.ts`
- missing `internalQdrantUrl`
- missing `internalEmbeddingUrl`
- `workspaces-bindings.test.ts`
- schema v3 semantic binding resolution threw unsupported errors
- `workspace-runtime-renderer.test.ts`
- schema v3 runtime rendering threw `Schema version 3 runtime rendering is unsupported until the internal semantic runtime is implemented`
## GREEN evidence
Focused GREEN command:
`cd backend && npx vitest run test/workspaces-contracts.test.ts test/workspaces-bindings.test.ts test/workspace-runtime-renderer.test.ts test/config.test.ts`
Result:
- 4 test files passed
- 36 tests passed
Typecheck:
`cd backend && npx tsc --noEmit -p .`
Result:
- passed
Hygiene:
- `git diff --check` passed
## Files changed
Listed-task files changed:
- `backend/src/config.ts`
- `backend/src/workspaces/bindings.ts`
- `backend/src/workspaces/runtime-renderer.ts`
- `backend/test/config.test.ts`
- `backend/test/workspace-runtime-renderer.test.ts`
- `backend/test/workspaces-bindings.test.ts`
- `backend/test/workspaces-contracts.test.ts`
Listed-task files inspected but not changed:
- `backend/src/workspaces/contracts.ts`
Unavoidable additional wiring changes:
- `backend/src/app.ts`
- `backend/src/tht/tht-runner.ts`
Reason: the new typed internal semantic runtime config had to flow from backend config into ephemeral harness config rendering at runtime.
## Behavior delivered
- schema-v3 installation contract exposes DWH bindings only
- schema-v3 binding resolution ignores external semantic env vars instead of sourcing runtime semantics from them
- runtime rendering for schema v3 emits backend-owned internal semantic endpoints:
- Qdrant: `http://qdrant:6333`
- Embedding: `http://embedding:11434`
- Model: `qwen3-embedding:0.6b`
- Dimensions: `1024`
- internal semantic URLs are validated to allow only `qdrant` / `embedding` / `localhost` / loopback hosts
- DWH transport/runtime behavior remains unchanged
## Self-review
- Confirmed schema-v3 contracts/docs no longer advertise VECTOR or EMBEDDING installation variables.
- Confirmed schema-v3 runtime output ignores injected external semantic endpoints from env bindings.
- Confirmed semantic endpoints are rendered only in the ephemeral backend-owned harness config path.
- Confirmed type wiring is explicit from `AppConfig` → `ThtRunner` → runtime renderer.
## Concerns
- Host validation currently permits both `http` and `https` on the allowed internal hosts. That keeps the configuration flexible, but if the installation contract intended `http` only, that restriction is not enforced here.
## Fix round 1/5
Scope:
- moved schema-v3 internal embeddings under `resources.embeddings`
- enforced `http`-only internal semantic URLs
RED evidence:
`cd backend && npx vitest run test/workspace-runtime-renderer.test.ts test/config.test.ts`
Observed failures on `bc8afe0`:
- `workspace-runtime-renderer.test.ts`
- schema-v3 output omitted `resources.embeddings`
- schema-v3 still exposed top-level `embeddings`
- `config.test.ts`
- `https://qdrant:6333` was accepted
GREEN evidence:
`cd backend && npx vitest run test/workspace-runtime-renderer.test.ts test/config.test.ts`
Result:
- 2 test files passed
- 16 tests passed
Typecheck:
`cd backend && npx tsc --noEmit -p .`
Result:
- passed
Updated concerns:
- none for this round beyond future tightening if exact-port rejection is later requested explicitly.
@@ -1,149 +0,0 @@
# Task 4 Report — Narrow harness embedding configuration to internal Ollama
## Status
Implemented on 2026-08-08 in `/Users/mp/projects/ThothII/.worktrees/git-workspace-registry`.
## RED evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py -q
```
Observed before implementation:
- exit code `1`
- `10 failed, 10 passed`
- failures proved the missing `OllamaInternalEmbeddings` client and missing internal-only config validation
Representative failures:
- `ImportError: cannot import name 'OllamaInternalEmbeddings'`
- `AttributeError: 'EmbeddingsConfig' object has no attribute 'provider'`
- config tests `DID NOT RAISE ConfigError` for external provider, API key, and non-private base URL
## GREEN evidence
Focused behavior suite:
```bash
cd harness
./.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py -q
```
- exit code `0`
- `20 passed`
Relevant harness verification:
```bash
cd harness
./.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py tests/test_ollama_ensure.py -q
```
- exit code `0`
- `36 passed, 2 warnings`
Changed-file lint:
```bash
cd harness
./.venv/bin/ruff check tht/config.py tht/config_compat.py tht/vectorstore/embeddings.py tht/cli/ollama_cmd.py tests/test_config_resources.py tests/test_internal_embeddings.py
```
- exit code `0`
- `All checks passed!`
Patch hygiene:
```bash
git diff --check
```
- exit code `0`
## What changed
- translated schema-v3 `resources.embeddings` into the harness-compatible embedding config view
- validated the internal embedding contract only for that runtime-owned `resources.embeddings` path:
- provider must be `ollama_internal`
- model must be `qwen3-embedding:0.6b`
- dimensions must be `1024`
- base URL must be `http://embedding:11434` or loopback HTTP on port `11434`
- extra fields like `api_key` are rejected
- replaced the active embed client with `OllamaInternalEmbeddings`, using one bounded `/api/embed` request per batch
- removed task/query prefix rewriting from the active embedding path
- validated response count, vector dimension, and finite numeric values before returning embeddings
- kept `tht ollama ensure --json` stdout pristine while warming through the internal client
## Self-review
- kept changes inside the brief-listed files
- preserved DWH and session-persistence behavior
- preserved the legacy `OllamaEmbeddings` import path as an alias to avoid unrelated call-site churn
## Concerns
- the focused harness verification still emits two pre-existing warnings:
- `DeprecationWarning` from `testcontainers.postgres`
- `FutureWarning` because `resources` currently flows through the legacy config translation path
## Fix round 1 — 2026-08-08
### Findings addressed
- HIGH: external top-level `embeddings` remained an operational fallback and could still load
- MEDIUM: non-object embed JSON payloads escaped as raw `AttributeError`
### RED evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py tests/test_ollama_ensure.py -q
```
Observed before the fix:
- exit code `1`
- `2 failed, 36 passed, 2 warnings`
Representative failures:
- `AttributeError: 'list' object has no attribute 'get'` from `response.json()` returning a JSON array
- `Failed: DID NOT RAISE ConfigError` for top-level external `embeddings.provider=openai_compatible`
### GREEN evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py tests/test_ollama_ensure.py -q
```
Observed after the fix:
- exit code `0`
- `38 passed, 2 warnings`
Touched-file lint:
```bash
cd harness
./.venv/bin/ruff check tht/config.py tht/vectorstore/embeddings.py tests/test_internal_embeddings.py tests/test_config_resources.py
```
- exit code `0`
- `All checks passed!`
### Minimal fix
- validated the final active `cfg.embeddings` contract after config loading, so legacy top-level
embedding inputs now fail explicitly unless they exactly match the internal Ollama contract
- converted non-mapping embed JSON payloads into controlled `EmbeddingsError` failures with
sanitized diagnostics instead of raw attribute errors
@@ -1,170 +0,0 @@
# Task 5 Report — Implement the Qdrant VectorStore adapter
## Status
Implemented on 2026-08-08 in `/Users/mp/projects/ThothII/.worktrees/git-workspace-registry`.
## RED evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
```
Observed before implementation:
- exit code `2`
- collection failed during import because the adapter did not exist yet
Representative failures:
- `ModuleNotFoundError: No module named 'tht.adapters.vector.qdrant'`
## GREEN evidence
Focused behavior suite:
```bash
cd harness
./.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
```
- exit code `0`
- `31 passed, 1 warning`
Touched-file lint:
```bash
cd harness
./.venv/bin/ruff check tht/adapters/vector/qdrant.py tht/adapters/vector/__init__.py \
tht/ports/vector.py tht/vectorstore/records.py tht/vectorstore/store.py \
tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py
```
- exit code `0`
- `All checks passed!`
Patch hygiene:
```bash
git diff --check
```
- exit code `0`
## What changed
- added `QdrantVectorStore` with direct `requests`-based REST calls for:
- `GET /collections/{collection}`
- `PUT /collections/{collection}`
- `PUT /collections/{collection}/index`
- `PUT /collections/{collection}/points?wait=true`
- `POST /collections/{collection}/points/query`
- `POST /collections/{collection}/points/scroll`
- `POST /collections/{collection}/points/delete?wait=true`
- implemented idempotent collection provisioning for `1024` dimensions and `Cosine` distance
- created deterministic UUIDv5 point IDs from workspace, semantic kind, and canonical record key
- preserved canonical record identity and only upserted/deleted points matching the exact workspace
and generation filters
- added Qdrant payload helpers so stored payloads carry:
- `workspace_id`
- grouped semantic `kind` (`schema`, `evidence`, `memory`)
- original `record_kind`
- canonical `record_key`
- `content_hash`
- existing Thoth metadata fields
- mapped Qdrant payloads back into existing `VectorHit` objects without losing the original
Thoth kind
- exported the new adapter from the public vector adapter package and added focused contract tests
- sanitized timeout and malformed-response failures so CLI-facing callers do not leak raw endpoint
details
## Self-review
- confirmed collection mismatch fails without any delete/recreate path
- confirmed every query/scroll/delete operation includes a workspace filter
- confirmed the adapter never deletes or rewrites unrelated Qdrant points
- added keyword payload indexes for all filter-critical fields used here, including `document_id`
for exact Evidence filtering
## Concerns
- the requested `adversarial-review` skill could not run its full external reviewer flow in this
environment because the skill’s referenced `brain/` files are missing at
`/Users/mp/.agents/skills/adversarial-review`; I performed a manual adversarial self-review
instead
- the focused suite still emits one pre-existing warning from `testcontainers.postgres`
## Fix round 1 — 2026-08-08
### Findings addressed
- IMPORTANT: metadata collisions could override canonical Qdrant payload identity fields and break
workspace isolation
- IMPORTANT: scroll-based operations only read the first page and did not follow
`next_page_offset`, making `existing_hashes`, `list_evidence_generations`, and delete counts
inexact beyond one page
### RED evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
```
Observed before the fix:
- exit code `1`
- `2 failed, 31 passed, 1 warning`
Representative failures:
- `assert payload["workspace_id"] == "demo"` failed because colliding `record.metadata`
overwrote canonical payload fields
- paginated scroll test missed later pages, so `existing_hashes` and generation cleanup counts
were incomplete
### GREEN evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
```
Observed after the fix:
- exit code `0`
- `33 passed, 1 warning`
Touched-file lint:
```bash
cd harness
./.venv/bin/ruff check tht/adapters/vector/qdrant.py tht/vectorstore/records.py \
tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py
```
- exit code `0`
- `All checks passed!`
Patch hygiene:
```bash
git diff --check
```
- exit code `0`
### Minimal fix
- made `qdrant_payload` apply canonical fields after `record.metadata` so workspace ID, semantic
kind, original record kind, canonical record key, and content hash cannot be overridden by
metadata collisions
- paginated `_scroll` until `next_page_offset` is absent, sent the returned `offset` back on the
next request, and reject repeated offsets as malformed to avoid infinite loops
@@ -1,144 +0,0 @@
# Task 6 Report
Date: 2026-08-08
Status: implemented and verified
Summary:
- Added schema-v3 Qdrant runtime support to the harness config/resource layer and vector factory.
- Made Qdrant payloads carry `workspace_id` and `workspace_revision` on every point.
- Routed schema and memory bulk indexing through the transport-neutral vector port with canonical hash-based dedup.
- Kept Evidence canonical on filesystem and Memory canonical in JSONL; Qdrant remains derived/rebuildable.
- Added focused tests for semantic-kind isolation, shared identity fields, search-pack kind boundaries, and the schema-v3 factory/config path.
Files changed:
- `harness/tht/config.py`
- `harness/tht/config_compat.py`
- `harness/tht/adapters/factory.py`
- `harness/tht/adapters/vector/qdrant.py`
- `harness/tht/vectorstore/records.py`
- `harness/tht/cli/vector_cmd.py`
- `harness/tht/cli/memory_cmd.py`
- `harness/tests/test_semantic_kind_isolation.py`
- `harness/tests/test_memory_save_one.py`
- `harness/tests/test_search_pack.py`
- `harness/tests/test_qdrant_vector_store.py`
- `harness/tests/test_adapter_factory.py`
- `harness/tests/test_config_resources.py`
Verification:
- Focused RED/GREEN task suite:
- `cd harness && .venv/bin/pytest tests/test_semantic_kind_isolation.py tests/test_memory_save_one.py tests/test_search_pack.py -q`
- Relevant harness suite:
- `cd harness && .venv/bin/pytest tests/test_semantic_kind_isolation.py tests/test_memory_save_one.py tests/test_search_pack.py tests/test_qdrant_vector_store.py tests/test_adapter_factory.py tests/test_config_resources.py tests/test_vector_port_contract.py tests/test_corpus_pipeline.py -q`
- Result: `131 passed`
- Changed-file Ruff:
- `cd harness && .venv/bin/ruff check tht/vectorstore/records.py tht/adapters/vector/qdrant.py tht/config_compat.py tht/config.py tht/adapters/factory.py tht/cli/vector_cmd.py tht/cli/memory_cmd.py tests/test_memory_save_one.py tests/test_search_pack.py tests/test_semantic_kind_isolation.py tests/test_qdrant_vector_store.py tests/test_adapter_factory.py tests/test_config_resources.py`
- Result: clean
Concerns / follow-up:
- `memory clear` still retains its older direct-vector assumptions and was not expanded in this task because the brief focused on canonical builders and schema/evidence/memory routing through the active Qdrant path.
- The relevant suite still emits pre-existing warnings (legacy config deprecation in older fixtures, plus existing Pydantic serializer warnings in corpus tests), but they are not introduced by this task.
## Fix round 1 (2026-08-08)
Scope:
- Fixed qdrant-only schema-v3 command gating for `vector index-schema`, `memory promote`, and `memory index`.
- Replaced `memory clear`'s direct-pgvector-only path with vector-port deletion by kind.
- Added focused qdrant-only CLI regression tests and refreshed older CLI fixtures to the enforced internal embedding contract.
RED evidence:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py -q`
- Initial result against commit `5e39cfa`: `4 failed`
- Failure signatures:
- `ERRORE: sezioni mancanti nel workspace yaml: vector_db o vector_write_rest.`
- `ERRORE: sezioni mancanti nel workspace yaml: vector_db.`
GREEN evidence:
- Focused fix suite:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py tests/test_qdrant_vector_store.py tests/test_adapter_factory.py tests/test_config_resources.py tests/test_memory_save_one.py tests/test_search_pack.py -q`
- Result: `51 passed`
- Relevant broader vector/memory/schema/search suite:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py tests/test_qdrant_vector_store.py tests/test_adapter_factory.py tests/test_config_resources.py tests/test_memory_save_one.py tests/test_search_pack.py tests/test_vector_port_contract.py tests/test_adapter_command_regressions.py tests/test_solved_search_cli.py tests/test_schema_introspect_guard.py tests/test_semantic_kind_isolation.py tests/test_corpus_pipeline.py -q`
- Result: `154 passed`
- Ruff on the fix surface:
- `cd harness && .venv/bin/ruff check tht/ports/vector.py tht/adapters/vector/qdrant.py tht/adapters/vector/pgvector.py tht/adapters/vector/thoth_http.py tht/vectorstore/rest_client.py tht/cli/vector_cmd.py tht/cli/memory_cmd.py tests/test_qdrant_cli_commands.py tests/test_solved_search_cli.py`
- Result: clean
Notes:
- `memory clear` now deletes derived `kind=memory` points through the configured writable vector store, while leaving the JSONL registry as the source of truth until the registry file is removed by the command.
- The broader suite still carries the same pre-existing warnings noted above; this fix round did not add new warnings or failures.
## Fix round 2 (2026-08-08)
Scope:
- Removed the accidental HTTP writer `delete_kinds` capability expansion from `ThothHttpVectorStore` and `VectorRestClient`.
- Reworked `memory clear` so schema-v3 Qdrant uses scoped `kind=memory` deletion, while legacy transports keep the pre-task direct-sync path instead of advertising a nonexistent RPC.
- Tightened the qdrant-only memory-clear regression to assert the exact `("memory", ["memory"])` delete scope.
RED evidence:
- Re-review found a transport contract mismatch in fix round 1:
- `ThothHttpVectorStore` exposed `delete_kinds(...)`
- `VectorRestClient` exposed `delete_kinds(...)`
- but the legacy HTTP writer migration only allowlists `delete_vector_generation`, not `delete_vector_kinds`
- The new regressions added in this round capture that mismatch and the missing qdrant delete-scope assertion:
- `tests/test_vector_port_contract.py::test_http_store_supports_writer_without_reader`
- `tests/l0/test_vector_adapter_parity.py::test_http_rest_client_does_not_advertise_nonexistent_delete_kinds_rpc`
- `tests/test_qdrant_cli_commands.py::test_memory_clear_accepts_qdrant_only_runtime_config`
GREEN evidence:
- Focused regression suite:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py tests/test_vector_port_contract.py tests/l0/test_vector_adapter_parity.py tests/test_adapter_command_regressions.py -q`
- Result: `53 passed`
- Broader relevant vector/memory/search suite:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py tests/test_adapter_command_regressions.py tests/test_vector_port_contract.py tests/l0/test_vector_adapter_parity.py tests/test_solved_search_cli.py tests/test_qdrant_vector_store.py tests/test_search_similar_kinds.py tests/test_corpus_pipeline.py -q`
- Result: `135 passed`
- Ruff on the changed fix surface:
- `cd harness && .venv/bin/ruff check tht/cli/memory_cmd.py tht/ports/vector.py tht/adapters/vector/thoth_http.py tht/vectorstore/rest_client.py tests/test_qdrant_cli_commands.py tests/test_vector_port_contract.py tests/l0/test_vector_adapter_parity.py`
- Result: clean
Notes:
- Legacy HTTP/vector-rest deployments do not gain a new destructive RPC surface from this fix; they keep their previous behavior and continue to fail closed for unsupported cleanup.
- The broader suite still emits the same pre-existing deprecation and serializer warnings already noted above; this round did not introduce new warnings.
## Fix round 3 (2026-08-08)
Scope:
- Added an adapter-level Qdrant regression for mixed semantic kinds within one workspace plus a second workspace memory point.
- Verified that `delete_kinds("memory", ["memory"])` emits the real adapter filter with both `workspace_id=demo` and `record_kind=memory`.
- Verified that non-memory semantic kinds in the same workspace and memory from another workspace survive the delete.
RED evidence:
- Re-review identified a test gap rather than a confirmed runtime bug:
- existing coverage asserted only the CLI mock call shape for qdrant memory clear
- there was no adapter-level regression proving the real Qdrant delete filter and resulting fake-Qdrant state across mixed semantic kinds/workspaces
- Added regression:
- `tests/test_qdrant_vector_store.py::test_delete_kinds_is_workspace_scoped_and_preserves_other_semantic_kinds`
GREEN evidence:
- Requested focused suite:
- `cd harness && .venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_qdrant_cli_commands.py tests/test_semantic_kind_isolation.py -q`
- Result: `18 passed`
- Ruff on changed files:
- `cd harness && .venv/bin/ruff check tests/test_qdrant_vector_store.py`
- Result: clean
Notes:
- This round required no production change; the new adapter regression passed against the existing Qdrant implementation.
- The focused suite still emits the same pre-existing `testcontainers.postgres` deprecation warning from `tests/conftest.py`; no new warnings were introduced.
@@ -1,98 +0,0 @@
# Task 7 report — mandatory Qdrant and Ollama Compose services
Date: 2026-08-08
Status: completed
Summary:
- Added mandatory private `qdrant`, `embedding`, and `embedding-model-init` services to the base Compose stack.
- Pinned Qdrant `v1.18.2` and Ollama `0.32.0` by immutable multi-arch digest.
- Persisted Qdrant storage in `qdrant-data` and Ollama model cache in `embedding-models`.
- Wired `core` to fixed internal semantic endpoints:
- `THT_INTERNAL_QDRANT_URL=http://qdrant:6333`
- `THT_INTERNAL_EMBEDDING_URL=http://embedding:11434`
- `THT_INTERNAL_EMBEDDING_MODEL=qwen3-embedding:0.6b`
- `THT_INTERNAL_EMBEDDING_DIMENSIONS=1024`
- Removed external vector / embedding endpoint requirements from the local and server env examples.
- Added an idempotent Ollama model bootstrap script that:
- waits up to a bounded deadline for `/api/tags`
- skips `ollama pull` when the model is already cached
- pulls `qwen3-embedding:0.6b` only when needed
- verifies the model appears in `/api/tags` after pull
- Added optional GPU override file `deploy/compose.embedding-gpu.yaml`; base Compose remains CPU-only.
- Updated `scripts/run-stack.sh` so the GPU override is included only when `THOTH_ENABLE_EMBEDDING_GPU=1`.
Verification:
- RED confirmed before implementation:
- `./scripts/test-default-compose.sh` failed on missing required services.
- `./scripts/test-unified-compose.sh` failed on missing required services.
- `./scripts/test-internal-semantic-compose.sh` failed because the GPU override file did not exist.
- GREEN after implementation:
- `./scripts/test-default-compose.sh`
- `./scripts/test-unified-compose.sh`
- `./scripts/test-internal-semantic-compose.sh`
- `git diff --check`
- Additional shell verification:
- `scripts/run-stack.sh --wait` includes only base + local Compose files by default.
- `THOTH_ENABLE_EMBEDDING_GPU=1 scripts/run-stack.sh --wait` adds `deploy/compose.embedding-gpu.yaml`.
Resolved image digests:
- `qdrant/qdrant:v1.18.2@sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c`
- `ollama/ollama:0.32.0@sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a`
Self-review:
- The first bootstrap-script draft depended on tools not guaranteed inside the Ollama image. This was corrected after image inspection; the final script uses only confirmed image tools (`bash`, `ollama`, `grep`) plus raw HTTP over `/dev/tcp`.
- The server overlay intentionally replaces most named core volumes with bind mounts, so the unified contract was tightened to require named semantic-cache volumes there while preserving the local/base named-volume checks.
Concerns:
- The model bootstrap waits for Ollama readiness and verifies cache state, but the first real cold-start will still take time to download `qwen3-embedding:0.6b`.
- The GPU override requests generic Docker GPU capability only; actual GPU availability remains host/runtime dependent and intentionally stays opt-in.
## Fix round 1 / 5 — 2026-08-08
Rulings applied:
- Kept the Task 1 boundary intact: schema-v3 remains the only operational workspace descriptor shape.
- Did not restore any external semantic fallback for schema-v2 live sessions.
- Treated `PROJECT_STATE.md` as stale documentation for this point, not runtime truth.
Focused schema-v2 evidence:
- Re-ran the existing targeted registry test:
- `cd backend && npx vitest run test/workspace-registry.test.ts -t "lists a schema v2 descriptor as migration_required and refuses to acquire it"`
- Result: pass.
- Evidence from that test:
- schema-v2 descriptors list as `migration_required`
- `acquireSessionRevision("psd-clinical")` rejects with `code: "workspace_invalid"`
- Conclusion: schema-v2 acquisition remains blocked; no external semantic fallback was reintroduced.
Contract consistency fixes:
- Updated `harness/tests/test_local_compose_contract.py` to assert the mandatory internal semantic stack, fixed internal core semantic env, private-service topology, persistent volumes, and Ollama health/dependency contract.
- Updated shell Compose contracts to require:
- Ollama healthcheck on `embedding`
- `embedding-model-init` dependency on `embedding: service_healthy`
- Updated `scripts/unified-deployment-smoke.sh` rendered-contract helper to expect the mandatory internal semantic topology and internal semantic env names, and to reject retired external semantic bindings.
- Updated `scripts/test-task13-runtime-fixtures.sh` to exercise `task13_assert_rendered_contract` for both local and server fixture renders.
Fix round 1 verification:
- RED before implementation:
- `cd harness && .venv/bin/pytest tests/test_local_compose_contract.py -q` failed because `embedding` had no healthcheck.
- `./scripts/test-default-compose.sh` failed because `embedding` had no healthcheck.
- `./scripts/test-unified-compose.sh` failed because `embedding` had no healthcheck.
- `./scripts/test-task13-runtime-fixtures.sh local` failed because `unified-deployment-smoke.sh` still expected `core,frontend`.
- GREEN after implementation:
- `./scripts/test-default-compose.sh`
- `./scripts/test-unified-compose.sh`
- `./scripts/test-internal-semantic-compose.sh`
- `cd harness && .venv/bin/pytest tests/test_local_compose_contract.py -q`
- `./scripts/test-task13-runtime-fixtures.sh local`
- `./scripts/test-task13-runtime-fixtures.sh server`
- `cd backend && npx vitest run test/workspace-registry.test.ts -t "lists a schema v2 descriptor as migration_required and refuses to acquire it"`
- `docker compose --env-file deploy/env/local.env.example -f compose.yaml -f deploy/compose.local.yaml config --format json`
@@ -1,65 +0,0 @@
Status: completed on August 8, 2026.
Summary:
- Updated the frontend workspace contract from schema v2 editing to schema v3 publishing.
- Kept only `semantic_index.vector_store.collection` editable; rendered qdrant / internal Ollama semantic values as fixed read-only architecture values.
- Removed external vector transport / endpoint / credential / embedding diagnostics branches from frontend draft sanitization, conflict parsing, and editor UI.
- Added a migration-required banner in workspace management and blocked `migration_required` workspaces from new-session selection.
- Aligned the example workspace YAML comments with the fixed internal qdrant/Ollama architecture.
Files changed:
- `frontend/src/api/workspaces.ts`
- `frontend/src/api/workspaces.test.ts`
- `frontend/src/workspaces/drafts.ts`
- `frontend/src/workspaces/drafts.test.ts`
- `frontend/src/shell/WorkspaceEditor.tsx`
- `frontend/src/shell/WorkspaceEditor.test.tsx`
- `frontend/src/shell/WorkspaceManager.tsx`
- `frontend/src/shell/WorkspaceManager.test.tsx`
- `frontend/src/shell/WorkspacePublishDialog.test.tsx`
- `frontend/src/api/sessions.ts`
- `frontend/src/shell/SteerInput.tsx`
- `frontend/src/shell/SteerInput.test.tsx`
- `deploy/workspaces/example.yaml`
- `deploy/workspaces/psd.yaml.example`
Verification:
- `cd frontend && npx vitest run src/shell/SteerInput.test.tsx src/shell/WorkspaceEditor.test.tsx src/shell/WorkspaceManager.test.tsx src/shell/WorkspacePublishDialog.test.tsx src/workspaces/drafts.test.ts src/api/workspaces.test.ts`
- Result: 6 files passed, 59 tests passed.
- `cd frontend && npx tsc -b`
- Result: passed.
- `git diff --check`
- Result: passed.
Self-review:
- The frontend now publishes the exact schema v3 semantic shape and no longer persists legacy semantic transport/credential branches.
- Migration-required workspaces are visible in management with an explicit banner and are excluded from the composer workspace selector.
- One dependent test file outside the original brief list (`WorkspacePublishDialog.test.tsx`) and the composer/session-selection path (`api/sessions.ts`, `SteerInput.tsx`, related test) were updated because they were directly coupled to the old v2 semantic/edit-selection behavior.
Concerns:
- The composer still retains backward-compatible behavior for summaries that omit `revision` entirely; only explicit `revision.state === "migration_required"` is blocked. That matches the current mixed-test environment, but once summary responses are guaranteed to include `revision`, that fallback may be removable.
Fix round 1/5 — August 8, 2026
Summary:
- Made missing or invalid workspace summaries fail safe in frontend session creation and composer selection instead of falling open as legacy.
- Added an actionable unavailable message in workspace management for incomplete summaries with no canonical revision.
- Replaced the old runtime-oriented example descriptor files with exact backend WorkspaceV3 descriptor YAML.
Additional files changed:
- `frontend/src/api/sessions.test.ts`
- `backend/test/workspaces-schema.test.ts`
Fix-round verification:
- `cd frontend && npx vitest run src/api/sessions.test.ts src/shell/SteerInput.test.tsx src/shell/WorkspaceManager.test.tsx src/shell/WorkspaceEditor.test.tsx src/shell/WorkspacePublishDialog.test.tsx src/workspaces/drafts.test.ts src/api/workspaces.test.ts`
- Result: 7 files passed, 73 tests passed.
- `cd frontend && npx tsc -b`
- Result: passed.
- `cd backend && npx vitest run test/workspaces-schema.test.ts`
- Result: 1 file passed, 17 tests passed.
- `git diff --check`
- Result: passed.
Notes:
- Missing `revision` in a workspace summary now fails with the same session/composer safety posture as `migration_required`, using the existing safe workspace-policy error for session creation and an explicit unavailable message in workspace management.
- The committed example files now validate as actual schema-v3 descriptors instead of deployment/runtime templates with forbidden semantic endpoint fields.
@@ -1,83 +0,0 @@
# Task 5 Report — `tht setup` lifecycle orchestration
## Status
Completed. `tht setup` now validates the checkout and host prerequisites, creates or validates
the non-secret installation files, validates Compose, and by default builds, starts, health-checks,
and verifies the installation. `tht setup --configure-only` stops immediately after successful
Compose rendering.
## Implementation
- Added `setup.Run`, with an ordered host preflight: project/worktree discovery, Docker Engine,
Docker Compose, supported architecture, and LF line-ending checks.
- Reused `config.Installation.ComposeArgs` for all Compose calls and added a narrow
`compose.InstallationRunner` adapter for Pi diagnostics; no shell command construction was added
to the top-level CLI parser.
- Default setup performs `compose build`, `compose up --detach --remove-orphans`, bounded polling
for `core`, `frontend`, `qdrant`, `embedding`, and `embedding-model-init`, then aggregate volume
diagnostics and `pi.Doctor`.
- Health timeout errors identify the last failing service and preserve containers for diagnosis,
with `tht logs <service>` and `tht status` guidance.
- Completion output includes the frontend URL, selected descriptor, and next action.
## TDD evidence
The initial focused test run failed because `setup.Run` did not exist. Tests were then written
against a fake Compose runner before the orchestration was implemented. They cover the complete
ordered flow, configure-only stop, preflight failure before writing configuration, health retry,
timeout guidance, and CLI default versus `--configure-only` dispatch.
## Verification
Executed from `tools/tht`:
```bash
go test ./internal/setup ./internal/compose ./cmd/tht -run 'TestRun|TestSetupCommand|TestInstallationRunner' -count=1
go test ./internal/setup ./internal/compose ./cmd/tht -count=1
go test ./...
git diff --check
```
All commands passed. No actual Docker build, container start, live-stack restart, system
installation, Pi configuration edit, or documentation rewrite was performed.
## Commit
`feat(setup): build start and verify ThothII` (this report is included in that commit).
## Concerns
- The bounded health wait is verified with fakes only, as required for this task. Real Docker
lifecycle verification belongs to the later live acceptance task.
- The existing aggregate `tht doctor` command remains a separate implementation; Task 5 performs
its equivalent setup-time prerequisite checks plus `pi.Doctor` without invoking a nested CLI
process.
## Fix round 1
The independent review identified three gaps. All were reproduced with RED tests before the
production change:
- A rendered Compose document containing any one volume was accepted. `requireVolumes` now
requires `settings`, `pi-state`, `workspace-registry`, `workspace-secrets`, `sessions`,
`qdrant-data`, and `embedding-models`; tests reject each individual omission and an
unrelated-only volume set.
- Failures after `compose up` could return without recovery instructions. A single recovery
wrapper now preserves the underlying error while adding the retained-container, `tht logs
<service>`, and `tht status` guidance for failed `up`, health, aggregate doctor, and Pi doctor
phases. Focused tests also prove build failure stops before attempting startup.
- LF inspection previously walked the full checkout. It now inspects only `compose.yaml`,
`deploy/`, and `docker/`; a test proves CRLF content under `node_modules/` is ignored.
Verification added for this round:
```bash
go test ./internal/setup -run 'TestRequireVolumes|TestRun(BuildFailure|UpFailure|AggregateDoctorFailure|PiDoctorFailure|IgnoresIrrelevant|TimesOut)' -count=1
go test ./internal/setup -count=1
```
Both passed before the final full-suite verification. No Docker or live operation was run.
Implementation commit evidence: `ea70cc95b04532043744a9de6c5912e30a214595` —
`fix(setup): harden verification and recovery`.
@@ -1,55 +0,0 @@
# Task 6 — Version, aggregate doctor, and build-aware start
Status: complete.
Implemented the host-side `tht version`, aggregate `tht doctor [--json]`, and `tht start [--build]` contracts.
- `version` is descriptor-free and reports semantic version, commit, build time, OS, and architecture.
- `doctor` emits typed, redacted checks for descriptor state, Docker/Compose, rendered volumes, file permissions, service health, workspace registry, the container-local workflow doctor, and Pi doctor. Its JSON mode writes exactly one JSON document to stdout.
- The Python workflow doctor is invoked only as `docker compose exec -T core tht doctor --json` after core is running.
- `start` uses the shared lifecycle service: default `up → health`; `--build` is `build → up → health`.
- `setup` now reuses the shared lifecycle and aggregate diagnostics rather than keeping parallel health/volume implementations.
Verification performed without live Docker/container commands:
```bash
cd tools/tht
go test ./internal/version ./internal/doctor ./internal/service ./cmd/tht \
-run 'TestVersion|TestDoctor|TestStart|TestCurrent|TestRun' -count=1
go test ./internal/setup -count=1 -run 'TestRun' -v
go test ./... -count=1
git diff --check
```
All completed successfully. The intentionally fake runner coverage includes unavailable Docker,
stopped/running core, workflow failure redaction, pristine JSON output, and start ordering.
Concerns: no live Docker validation or host installation was run, by explicit task constraint.
## Fix round 1
Completed the independent-review follow-up without live Docker operations.
- `workspace-registry` now executes a container-local, read-only Node validation of
`/data/workspace-registry/state/active.json` and every declared snapshot descriptor. It no
longer passes merely because Compose declares a volume.
- Host file permissions are checked before Docker/Compose availability and therefore remain
visible as failures when Docker is unavailable.
- Separate typed, bounded HTTP probes verify core (`curl --max-time 5`) and frontend
(`wget -T 5`) reachability, independently of Compose health. The probe is injectable in tests.
- The successful report tests assert the stable full checklist:
`descriptor`, `files`, `docker`, `compose`, `configuration`, `services`, `core-http`,
`frontend-http`, `workspace-registry`, `workflow`, `pi`.
Additional verification:
```bash
cd tools/tht
go test ./internal/doctor -run 'TestRun(ChecksUnsafeFilesEvenWhenDockerIsUnavailable|FailsAnInvalidContainerLocalRegistryState|ReportsEachHTTPReachabilityProbeFailure|UsesOnlyContainerLocalWorkflowAndPiDiagnosticsWhenCoreRuns)' -count=1 -v
go test ./internal/doctor ./internal/setup ./internal/service ./cmd/tht -count=1
go test ./... -count=1
git diff --check
```
All passed with fake runners/probes only. No live container, HTTP endpoint, or host installation
was touched.
@@ -1,163 +0,0 @@
# Task 15 retained release-gate report — fix round 5 (sanitized)
## Final-review fix-round-2 addendum — frozen source `2a9359071257f9b8a71d36ec2bbb25b161003f81`
This addendum supersedes the fix-round-1 addendum for current authentication remediation status
while preserving the fix-round-5 material below as historical provenance.
- Authentication remediation status: `PASS`. The three original remediation Important findings
remain `RESOLVED`; the fix-round-2 fully bounded lifecycle Important is `ADDRESSED`; and the
temporary Windows diagnostic-matrix Minor is `ADDRESSED`.
- Overall branch/release readiness is separately `FAIL`, with unavailable external/manual gates
`PENDING`.
- Completed exact-source workflow run `32147345625` concluded `failure` on baseline release jobs.
Its `Windows clone and Compose contract` job (`95744249248`) executed the unfiltered command
`go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1`; the native step
passed all three packages: safeio `22.058s`, backup `7.161s`, authstorage `16.088s`.
- The Windows job failed only afterward in the baseline clone-contract script at
`scripts/test-windows-clone-contract.ps1:208`, where PowerShell rejects the undelimited
`$remoteYaml:` variable reference.
- `LF, Compose, docs, and TypeScript` job `95744249458` reproduced the baseline unset-`TMPDIR`
failure after unified Compose passed. Linux Docker job `95744249354` reproduced the missing-`rg`
prerequisite failure; cleanup passed and no image manifest was generated.
- The skipped Windows Docker Desktop/WSL2 job is recorded as `NOT_RUN` / `BLOCKED`, not FAIL.
Downstream commands skipped after executed baseline failures use the same classification. The
matrix contains an explicit native `windows_stagearchive_retained_capability` PASS row.
- Historical Node/auth/browser/docs PASS and harness/Ruff/Compose FAIL evidence remains bound to
its recorded source where not rerun. L2, real PSD/manual acceptance, and provider readiness
remain `PENDING`.
- Current machine-readable evidence and the requested Task 4 report are recorded in
`.artifacts/task-15/automated-gates.json` and
`.superpowers/sdd/2026-08-18-thothii-authentication-remediation/task-4-report.md`.
- The full fix-round-2 RED/GREEN and finding disposition is recorded in
`.superpowers/sdd/2026-08-18-thothii-authentication-remediation/fix-round-2-report.md`.
- Current automated-gates SHA-256:
`6c516db5c2064c4a4a2e5f25961b993cd4a8fe020bbbb822fbac7faa0c119599`.
- Historical unified Docker manifest SHA-256: `9c8dec4546909fd93799dbcf374bcb3a89bc46cfe0fd482472c0cbe757ddf5b6`.
The complete sanitized Task 4 matrix and the separate remediation/release verdicts are in the
requested Task 4 report.
- Final tested source commit: `74b062f1a737103524cbe706346cfd65f87cdfd1`.
- Historical retained source commits: fix-round-2 `fe190e7046acc173f510dddcb32f46ed142858c1`,
maintenance follow-up `4d230b87afdcd24f02264f8f937c8628b92db05a`, prior final Docker
source `e20bf33e2a00102192e5be66b178037aeca3a7b1`, and fix-round-4 streamed
archive privacy `54698e73400a54ce7c3e6c10099e14eb471ce8b9`.
- Versions: Node contract `v24.16.0`; host default Node `v25.6.1`; Go `go1.26.5`;
Pi `0.80.3`.
- Historical automated gate artifact: `.artifacts/task-15/automated-gates.json`;
SHA-256 `7d9ec93af15510605f1aa7179b26a7ee46d78122f647854300f7a9922057a63f`.
- Docker image manifest: `.artifacts/task-15/unified-docker-images.json`;
SHA-256 `9c8dec4546909fd93799dbcf374bcb3a89bc46cfe0fd482472c0cbe757ddf5b6`.
## Fix-round-5 evidence
- PASS, RED then GREEN: `TestCreateCanonicalNewPrivateFileUsesPinnedParentAfterAncestorSwap`
first failed because the creator had not retained its parent before creation. It now opens every
Unix ancestor once, creates the leaf with `openat(O_NOFOLLOW|O_CREAT|O_EXCL)`, applies and checks
`0600` by descriptor (`fchmod`/`fstat`), and uses `unlinkat` for creator failure cleanup. The
deterministic test moves the opened parent, replaces its lexical name with an outside symlink,
validates the archive under the moved original parent, and proves no outside archive was written.
- PASS: the Windows implementation uses NT `RootDirectory`-relative traversal for every component
after the volume root and for final file creation. The retained final parent receives only the
required child-create right (`FILE_WRITE_DATA` for a file, `FILE_APPEND_DATA` for a directory),
reparse points are rejected, and the owner-only protected DACL is installed in the same
`NtCreateFile` operation. The native-Windows test attempts the pre-create parent swap and calls
`safeio.ValidatePrivateRegular`; it is compiled but not executed on this host.
- PASS: `go test ./internal/safeio ./internal/backup -count=1`, `go test -race ./...` across
`18` packages, `go vet ./...`, and a native host `tht` CLI build. Existing StageArchive
capacity, lifecycle, rollback, streaming, and cleanup tests remain passing.
- PASS, compile-only: Windows amd64 static test/build compilation across `18` packages, including
the retained-handle Windows tests. No Windows executable was run; native execution remains
PENDING and is not inferred from compilation.
- PASS on Node `v24.16.0`: the hermetic OIDC/F1 authentication browser smoke passed all current
`8` checks in `frontend/e2e/auth.spec.ts` and `frontend/e2e/f1.spec.ts`; the runtime sentinel
leak scan passed.
- PASS: shell syntax, unified-smoke safety self-test, default Compose contract, unified Compose
contract, and Compose secret-policy contract.
- PASS: final unified Docker deployment smoke run `20260818070637-66409-30058`, bound exactly to
source `74b062f1a737103524cbe706346cfd65f87cdfd1`. It exercised maintenance-auth isolation,
restore, registry lifecycle, bad-candidate rollback, image revalidation, and task-scoped cleanup.
## Sanitized final unified Docker output
```text
== Build and start isolated local Compose distribution ==
== Recreate offline and retain the validated registry snapshot ==
== Pull a valid catalog+descriptor metadata update ==
== Pull a content-only Git Evidence update ==
== Reject catalog/descriptor metadata mismatch and retain the valid snapshot ==
== Reject orphan descriptor directories not listed in the catalog ==
== Reject the retired flat workspace layout and retain the valid snapshot ==
== Inject a bad pinned Pi candidate and prove automatic rollback ==
Task 13 full deployment smoke passed.
Task 13 cleanup proof: no labeled containers, volumes, networks, or images remain for 20260818070637-66409-30058.
```
## Sanitized Docker image identities
- `sha256:2d7b19491c7eb8c119c3cedb390aaeb2ff5593f6fc43ab66c317565560da6d7d`;
roles `compose-runtime`, `fixture-runtime`.
- `sha256:3b6c31a5d8f8fc58fa3233391b6175bd2fbc793eebb44d5e285ecc6e02e9e687`;
role `compose-runtime`.
- `sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a`;
role `compose-runtime`.
- `sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c`;
role `compose-runtime`.
- `sha256:c3cbe1cc1aa588a64951ac6286e0df7b27fe2e6324b1001c619bb358770c0178`;
role `rollback-candidate`.
For each image, the retained repository-digest component equals the listed image digest. Registry
names and credentials are deliberately omitted.
## Complete observed matrix
- PASS: Task 13 lifecycle carry-ins; retained-handle owner-private restore staging; provider fixture
round-one `6/6`; backend Node 24 round-one suite `75 files / 1081 tests`; frontend Node 24
round-one suite `61 files / 444 tests`; current Node 24 authentication/F1 browser smoke `8/8`;
final-source Go race/build `18 packages`; Windows static cross-compile `18 packages`; harness
round-one suite `921 passed / 4 L2 deselected`; authentication docs round-one gate; shell/Compose
contracts; final unified Docker smoke; five-image traceability; and Docker cleanup.
- FAIL: Ruff `192` known-baseline errors; MkDocs strict `69` known-baseline warnings; existing
canonical/workspace install wording checks; existing Pi model-policy check; deployment-coupling
scan against preserved ignored private material.
- PENDING: native Windows execution because required host prerequisites are unavailable; L2 because
the configured secret layout is unavailable; real PSD/manual acceptance because no real
identity/access is available; isolated provider readiness because an unrelated host port is
occupied.
## Final Task 15 review after fix round 5
The fresh Terra review verdict is **CHANGES REQUIRED**. The five-round breaker is exhausted; no
sixth implementation round was started. Two Important findings remain:
- `StageArchive` does not retain the opaque parent/directory capability through the complete
stream and `Close` lifecycle. Staging-directory creation and final cleanup still use pathname
operations, so an ancestor swap after creation can strand the secret-bearing archive or redirect
cleanup. Deterministic StageArchive swap-and-cleanup coverage is still required on Unix and
native Windows.
- Windows claim removal closes its validated retained parent handles before calling pathname-based
`DeleteFile`. Removal must instead remain handle-relative (or delete through the opened handle),
with a native-Windows ancestor-swap test.
The focused/full Go, cross-compile, Node 24, browser, Compose, Docker lifecycle, image-traceability,
and cleanup results above remain valid evidence for source `74b062f1a737103524cbe706346cfd65f87cdfd1`.
They do not override the final code-review verdict. Native Windows execution remains PENDING.
The authentication feature is **not implementation-complete or release-complete** while these code
findings and the required FAIL/PENDING gates remain. No secret values, real identities, internal
endpoints, or registry names are retained.
## Final whole-branch review
The final read-only Terra review of `351361f..39b5453` also returned **CHANGES REQUIRED** and found
one additional Important issue: the POSIX local-user registry validates file type, link count, and
mode for `users.yaml` and its parent directory, but does not require ownership by the effective UID.
A foreign-owned `0600` registry inside a runtime-owned `0700` directory can remain writable by the
foreign owner and be used to alter credentials or grant the administrator role. The registry must
enforce effective-UID ownership on every POSIX `lstat`/`fstat` path and add foreign-owner rejection
coverage.
No new Critical issue or load-bearing Minor issue was found. The branch is **not ready to merge**:
this ownership defect and the two retained-capability cleanup defects above require fixes and renewed
review, independently of the remaining FAIL/PENDING release gates.
@@ -1,162 +0,0 @@
# Final-review fix round 1 report (sanitized)
## Verdict
- Base: `fa499a9bdd37011833691b0f447470d8b7e8a3a6`.
- Final frozen source: `10cd66fe6a5b484a4dc569326a228c1c5484a5d4` on
`feat/thoth-auth`.
- Authentication remediation: **PASS / ADDRESSED**. All four final-review Important findings are
resolved relative to the remediation brief.
- Terra Minor evidence corrections: **ADDRESSED**.
- Branch/release readiness: **FAIL**. The completed exact-source workflow still contains executed
baseline clone-contract, LF/Compose, and Linux Docker failures. Unavailable external/manual
gates remain **PENDING**.
- Source and evidence remain separate commits. No workflow was dispatched from the evidence-only
phase.
## Finding disposition
| Finding | Disposition | Evidence |
|---|---|---|
| Important 1 — exhaustive Windows cleanup | RESOLVED | Cleanup now attempts close/delete/validation operations in deterministic order and returns sanitized `ErrUnsafeFile` after aggregating failures. `TestWindowsPrivateRegularCleanupClosesAfterDeleteDispositionFailure` and `TestWindowsClaimCleanupAttemptsLaterOperationsAfterEarlierFailure` cover the non-short-circuit contract. Global no-delete sharing remains unchanged. |
| Important 2 — usable native Windows authority | RESOLVED | Owner-only descriptors use the current user SID, protected/non-defaulted DACL semantics, valid NT attributes/access masks, self-relative creation descriptors, and semantic full-control validation. Equal-or-stronger Windows fixture adaptations retain no-delete handles instead of weakening ACL/identity checks. The final native three-package gate passes. |
| Important 3 — restore-test deadlock | RESOLVED | Lifecycle-stage release observes the buffered worker outcome, uses a bounded/cancellable release, reports premature completion directly, and never waits indefinitely on `done`. `TestReleaseLifecycleStageReturnsPrematureWorkerOutcome` and the lifecycle-lock terminal-cleanup test are green. |
| Important 4 — complete native package gate | RESOLVED | Workflow and remediation plan both use the exact unfiltered command `go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1`. Final logs prove all three packages executed natively. |
| Minor — non-executed gate classification | RESOLVED | Non-executed/skipped commands are `NOT_RUN` / `BLOCKED`; `FAIL` is reserved for commands that ran and failed. Historical results remain separately labelled. |
| Minor — explicit Windows StageArchive row | RESOLVED | `.artifacts/task-15/automated-gates.json` contains `windows_stagearchive_retained_capability` = PASS, bound to the final source and native backup result. |
Additional failures exposed by the required unfiltered gate were fixed without narrowing the
workflow: Windows secret-bearing archive reservation is protected before use; StageArchive shares
one retained root capability across both staged files; claim/consume transitions serialize the
complete public validation and retained-handle operation while preserving ACL, hard-link identity,
reparse rejection, and no-delete invariants.
## RED → GREEN record
### Initial RED
- Run `32122302381`:
https://github.com/mptyl/ThothII/actions/runs/32122302381
- Source: `b31b27e5845ffd3adf311429367319beaba263c7`.
- Windows job: `95665197885`.
- Result: native `safeio`/`backup` failure, including the 10-minute restore lifecycle timeout;
`authstorage` was absent from the command. This established the RED for Important 2–4 and the
required native authority.
- Cleanup failure-injection tests added for Important 1 first exposed the short-circuit behavior
before the implementation was changed.
### Final concurrency RED
- Run `32140481263`:
https://github.com/mptyl/ThothII/actions/runs/32140481263
- Source: `b48e9e9189dd0e8083db9bd0378704524e670edb`.
- Windows job: `95721724645`.
- Native results: backup PASS (`20.757s`), authstorage PASS (`104.180s`), safeio FAIL
(`63.502s`). The only failures were:
- `TestCanonicalPrivateClaimWaitsForRetainedRemoveOperation`: the concurrent claim returned
`false, unsafe file` before retained removal completed;
- `TestCanonicalPrivateClaimConsumeHasOneConcurrentWinner`: iteration 8 returned `unsafe file`.
- Diagnosis: the process mutex started below `validateClaimPaths`; a concurrent caller could fail
while reopening the retained no-delete directory before reaching the lock.
### GREEN implementation and local gates
The lock boundary was moved to the three public claim/read/remove APIs, covering validation,
relative operation, and handle close. The Unix implementation uses a no-op boundary and retains its
existing descriptor-relative semantics.
Final-source local commands passed:
```text
go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1
go test -race ./...
go vet ./...
go build -o /tmp/thothii-tht-host ./cmd/tht
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go build -o /tmp/thothii-tht-windows.exe ./cmd/tht
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go test -c ... ./internal/{safeio,backup,authstorage}
```
- Focused host package times: safeio `8.750s`, backup `8.378s`, authstorage `8.854s`.
- Race suite and vet: PASS.
- Host CLI: Mach-O arm64; Windows CLI and all three Windows test binaries: PE32+ x86-64.
- Cross-compilation remains compile-only and is not used as native proof.
## Exact-source native certification
- Run: `32141428407`
- URL: https://github.com/mptyl/ThothII/actions/runs/32141428407
- Event/status/conclusion: `workflow_dispatch` / `completed` / `failure`.
- Head SHA: `10cd66fe6a5b484a4dc569326a228c1c5484a5d4` — exact final source match.
- Windows job: `Windows clone and Compose contract`, job `95724751282`:
https://github.com/mptyl/ThothII/actions/runs/32141428407/job/95724751282
- Native step: `Run native Windows retained-capability tests` — **PASS**.
- Exact command: `go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1`.
- Native package results:
- safeio PASS (`8.230s`);
- backup PASS (`5.195s`);
- authstorage PASS (`8.383s`).
- Job conclusion: `failure` only because the following `Verify Windows clone contract` baseline
step failed with a PowerShell `ParserError` at
`scripts/test-windows-clone-contract.ps1:208`; `$remoteYaml:` is not delimited before `:`.
## Remaining branch/release blockers
| Gate | Classification | Exact outcome |
|---|---|---|
| Windows native authentication packages | PASS | All three required packages executed on final source. |
| Windows clone contract | FAIL / baseline | Executed after native PASS; PowerShell parser error at line 208. |
| LF, Compose, docs, and TypeScript | FAIL / baseline CI contract | Job `95724751205`; unified Compose passed, then `test-no-deployment-coupling-scope.sh` failed because `TMPDIR` was unset. Downstream skipped commands are `NOT_RUN` / `BLOCKED`. |
| Linux Docker deployment and rollback | FAIL / infrastructure prerequisite | Job `95724751356`; executed smoke stopped because `rg` was unavailable. Cleanup proof passed; no new image manifest was generated. |
| Native Windows Docker Desktop/WSL2 startup | NOT_RUN / BLOCKED | Job `95724752028` was skipped by workflow conditions; no Docker/WSL2 command executed. |
| Harness/Ruff/other historical baseline gates | FAIL | Retained with their recorded source and results; not rewritten as final-source proof. |
| L2, real PSD/manual acceptance, provider readiness | PENDING | Required secrets, identity/access, or provider prerequisites remain unavailable. |
The historical Docker image manifest remains bound to source
`74b062f1a737103524cbe706346cfd65f87cdfd1`; it was not reused as proof for the final source.
## Principal source commits
- `cd5f505` — exhaustive cleanup, Windows authority foundation, restore deadlock tests/fix, and
complete workflow/plan package command.
- `a0e05ad` through `b6396e6` — effective full-control DACL semantics, valid NT attributes/access,
self-relative descriptors, retained no-delete fixture ordering, and Windows installation fixture
protection.
- `824245d` — preserve existing lifecycle ACL trees instead of mutating inherited authority.
- `455fffb`, `2d1670e`, `c01482c`, `9fc1a15` — concurrent claim/consume and settled-loss handling.
- `6474118` — one retained StageArchive root capability shared across staged files.
- `feee4ee` — unified Windows path wrappers on the retained primitive.
- `b261dd4` — bounded private-root sharing contention handling.
- `b48e9e9` — deterministic retained-remove concurrency regression and claim-operation lock.
- `10cd66f` — final lock boundary includes public path validation; frozen source.
## Files changed
Source changes relative to the fix-round base:
- `.github/workflows/deployment.yml`;
- `docs/superpowers/plans/2026-08-18-thothii-authentication-remediation.md`;
- `tools/tht/internal/authstorage/storage_test.go`;
- `tools/tht/internal/backup/{create.go,create_test.go,fixture_security_unix_test.go,fixture_security_windows_test.go,preflight.go,preflight_test.go,preflight_windows_test.go,restore.go,restore_test.go}`;
- `tools/tht/internal/safeio/{claim_unix.go,claim_windows.go,claim_windows_test.go,files.go,files_test.go,private_root_windows.go,private_windows.go,private_windows_test.go}`.
Evidence/status changes are restricted to:
- `.artifacts/task-15/automated-gates.json`;
- `.superpowers/sdd/2026-08-18-thothii-authentication-remediation/task-4-report.md`;
- `.superpowers/sdd/2026-08-18-thothii-authentication-remediation/fix-round-1-report.md`;
- `.superpowers/sdd/2026-08-16-thothii-authentication/task-15-report.md`;
- `PROJECT_STATE.md`.
Machine-readable evidence SHA-256:
`5c110b7b2607693de078def441b10290c5a29024c83b7e5a0ced894b72b7507f`.
## Git and protection status
- The evidence commit contains only the five evidence/status files listed above; no source is
changed after frozen source `10cd66fe6a5b484a4dc569326a228c1c5484a5d4`.
- After the evidence commit and push, the intended status is synchronized
`feat/thoth-auth...origin/feat/thoth-auth` with only protected untracked `.playwright-cli/` and
`.thothctl/`.
- `AGENTS.md`, `CLAUDE.md`, and `docs/agents/` are untouched. No generated `tools/tht/tht` exists.
- Evidence commit SHA is reported externally after commit creation because a commit cannot contain
its own final hash.
@@ -1,133 +0,0 @@
# Final-review fix round 2 report (sanitized)
## Verdict
- Base evidence head: `0f762ad6b67675356389cc546421a1c46ad5a736`.
- Frozen source: `2a9359071257f9b8a71d36ec2bbb25b161003f81` on `feat/thoth-auth`.
- Authentication remediation: **PASS**.
- Three original remediation Important findings: **RESOLVED**.
- Fix-round-2 bounded lifecycle Important: **ADDRESSED**.
- Fix-round-2 temporary Windows diagnostics Minor: **ADDRESSED**.
- Release readiness: **FAIL** for executed unrelated baseline gates, with unavailable
external/manual gates separately **PENDING**.
- Source and evidence are separate commits. The evidence-only phase changed no source or tests and
dispatched no workflow.
## Finding disposition
| Finding | Disposition | Evidence |
|---|---|---|
| Original Important — POSIX local-registry ownership | RESOLVED | Effective-UID ownership enforcement and its Node 24 coverage remain green at their recorded source. Fix round 2 did not alter this boundary. |
| Original Important — retained-capability StageArchive lifecycle | RESOLVED | Native Windows `internal/backup` passed on the exact source, preserving the retained-root staging and cleanup coverage. |
| Original Important — handle-relative Windows claim removal | RESOLVED | Native Windows `internal/safeio` and `internal/authstorage` passed on the exact source, including retained claim/consume coverage. |
| Fix-round-2 Important — fully bounded restore lifecycle test | ADDRESSED | Gate publication and release are context-aware; stage, outcome, admission, checkpoint, and verification waits are bounded; aborts cancel, safely release, bounded-join, then assert lock-free. The deterministic withheld-gate test proves prompt timeout/cancellation, worker join, and eventual lock release. |
| Fix-round-2 Minor — temporary Windows diagnostic matrix | ADDRESSED | `windowsRelativeOpenMatrix` and its diagnostic-only call/import were removed. Owner-only DACL shape, NT access normalization, full-control, cleanup, and retained no-delete tests remain. |
The round-1 restore lifecycle finding was broadened by the scoped round-2 review: bounded release
alone was insufficient while stage publication, gate waits, and nearby outcome/admission waits
could still outlive a controller abort. The round-2 implementation closes that broader test
orchestration gap without changing production authentication semantics.
## RED → GREEN record
### RED
The deterministic withheld-gate regression was introduced first and run without relying on a
global ten-minute package timeout:
```text
go test ./internal/backup -run '^TestRestoreLifecycleCancellationJoinsWithWithheldGate$' -count=1
```
It failed in approximately `0.64s` with:
```text
cancelled restore worker did not join within the bounded deadline
```
This proved that cancellation did not yet unblock and join a worker retained at the lifecycle
gate.
### GREEN and refactor
- The gate uses a cancellation source shared by controller and worker. Both publication and
release are `select`-based and cancellation-aware.
- Shared bounded helpers cover stage, outcome, error, signal, release, and admission waits.
- Abort cleanup is ordered: cancel, cancel the controller gate when distinct, safely release a
pending gate, bounded-join the worker, then prove the lifecycle lock is free.
- Premature worker outcomes retain and surface their original error.
- The existing success, recovery, maintenance-barrier, stale-checkpoint, and verification
assertions remain active.
Final local gates on the frozen source:
```text
go test ./internal/backup -run '^(TestRestoreLifecycleCancellationJoinsWithWithheldGate|TestReleaseLifecycleStage|TestRestoreLifecycleLockExcludesCompetingTransactionsUntilTerminalCleanup|TestRestoreCannotApplyAStaleCheckpointOverAnInterleavedRestore|TestRestoreKeepsAdmissionBarrierActiveUntilVerificationCommits)$' -count=1
go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1
go test ./... -count=1
go test -race ./...
go vet ./...
go build -o /tmp/thothii-tht-host-fix-round-2 ./cmd/tht
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go test -c ./internal/safeio -o /tmp/tht-safeio-fix-round-2-windows.test.exe
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go test -c ./internal/backup -o /tmp/tht-backup-fix-round-2-windows.test.exe
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go test -c ./internal/authstorage -o /tmp/tht-authstorage-fix-round-2-windows.test.exe
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go build -o /tmp/thothii-tht-fix-round-2-windows.exe ./cmd/tht
```
All commands passed. The final focused lifecycle run completed in `0.672s`; the full security
package run passed safeio, backup, and authstorage; race, vet, host build, Windows test-package
cross-compiles, and Windows CLI cross-compile also passed. Cross-compilation is recorded only as
compile evidence and is not used as native authority.
## Exact-source native certification
- Controller-authorized run: `32147345625` —
https://github.com/mptyl/ThothII/actions/runs/32147345625.
- Event/status/conclusion: `workflow_dispatch` / `completed` / `failure`.
- Head SHA: `2a9359071257f9b8a71d36ec2bbb25b161003f81`, exactly matching the frozen source.
- Windows job: `Windows clone and Compose contract`, job `95744249248` —
https://github.com/mptyl/ThothII/actions/runs/32147345625/job/95744249248.
- Native step: `Run native Windows retained-capability tests` — **PASS**.
- Exact unfiltered command:
`go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1`.
- Native package results:
- `internal/safeio` PASS (`22.058s`);
- `internal/backup` PASS (`7.161s`);
- `internal/authstorage` PASS (`16.088s`).
The Windows job failed only in the following baseline clone-contract step. PowerShell reported a
parser error at `scripts/test-windows-clone-contract.ps1:208` because `$remoteYaml:` is not a
delimited variable reference. This later failure does not alter the successful native Go step.
## Separate release-readiness verdict
| Gate | Classification | Exact outcome |
|---|---|---|
| Authentication remediation | PASS | Source and exact-source native three-package authority are green. |
| Windows clone contract | FAIL / baseline | Job `95744249248`; parser error at `scripts/test-windows-clone-contract.ps1:208`, after native PASS. |
| LF, Compose, docs, and TypeScript | FAIL / baseline CI contract | Job `95744249458`; unified Compose passed, then the existing unset-`TMPDIR` failure stopped the contract step. Downstream commands were skipped. |
| Linux Docker deployment and rollback | FAIL / infrastructure prerequisite | Job `95744249354`; the existing missing-`rg` prerequisite stopped the smoke before deployment. Cleanup passed and no new image manifest was generated. |
| Native Windows Docker Desktop/WSL2 startup | NOT_RUN / BLOCKED | Job `95744250450` was skipped by workflow conditions; no native Docker/WSL2 command ran. |
| L2, real PSD/manual acceptance, provider readiness | PENDING | Required secrets, identity/access, or provider prerequisites remain unavailable. |
Executed failures remain `FAIL`; skipped commands are `NOT_RUN` / `BLOCKED`; unavailable external
gates remain `PENDING`. Therefore remediation PASS does not imply release readiness PASS.
## Evidence and protection status
- Machine-readable evidence: `.artifacts/task-15/automated-gates.json`; SHA-256
`6c516db5c2064c4a4a2e5f25961b993cd4a8fe020bbbb822fbac7faa0c119599`.
- Current Task 4 report:
`.superpowers/sdd/2026-08-18-thothii-authentication-remediation/task-4-report.md`.
- Retained Task 15 report:
`.superpowers/sdd/2026-08-16-thothii-authentication/task-15-report.md`.
- Project snapshot: `PROJECT_STATE.md`.
- Historical Docker evidence remains bound to its recorded older source and is not reused as proof
for `2a9359071257f9b8a71d36ec2bbb25b161003f81`.
- `.playwright-cli/` and `.thothctl/` remain protected and untracked. No source/test file,
instruction file, workflow, or `docs/agents/` content changed in this evidence phase.
- The separate evidence commit SHA is reported after commit creation because a commit cannot
contain its own final hash.
No credentials, tokens, internal endpoints, identities, registry names, raw environments, or
browser traces are retained in this report.
@@ -1,131 +0,0 @@
# Task 4 authentication remediation recertification (sanitized)
## Fix-round-2 recertification — remediation PASS
- Exact source: `2a9359071257f9b8a71d36ec2bbb25b161003f81` on `feat/thoth-auth`.
- Authorized workflow: completed run `32147345625`,
https://github.com/mptyl/ThothII/actions/runs/32147345625, exact matching head SHA.
- Native job: `Windows clone and Compose contract`, job `95744249248`.
- Required native step: `Run native Windows retained-capability tests` — **PASS**.
- Exact unfiltered command:
`go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1`.
- Package evidence: `internal/safeio` PASS (`22.058s`), `internal/backup` PASS (`7.161s`),
`internal/authstorage` PASS (`16.088s`). This includes explicit native Windows
StageArchive retained-capability and concurrent claim-consume coverage.
- The later `Verify Windows clone contract` step failed independently at
`scripts/test-windows-clone-contract.ps1:208`: PowerShell parsed `$remoteYaml:` as an invalid
variable reference. This baseline deployment-contract failure does not change the native Go
package result.
- The optional `Native Windows Docker Desktop/WSL2 startup` job was skipped by workflow
conditions. It is `NOT_RUN` / `BLOCKED`, because no Docker Desktop/WSL2 command executed.
- The workflow reached `completed` with conclusion `failure`: the native authentication step is
PASS, while the later clone-contract, LF/Compose, and Linux Docker baseline steps are FAIL.
- Existing LF/Compose job `95744249458` and Linux Docker job `95744249354` failures repeated
before downstream work. Skipped commands are `NOT_RUN` / `BLOCKED`, not executed failures.
External L2/PSD/provider gates remain `PENDING`.
Finding disposition is explicit: the three original remediation Important findings remain
**RESOLVED**; the fix-round-2 lifecycle Important is **ADDRESSED**; and the temporary Windows
diagnostic-matrix Minor is **ADDRESSED**. Authentication remediation is **PASS**. This does not
change overall release readiness: executed baseline gates remain **FAIL**, while unavailable
external/manual gates remain **PENDING**.
The section below is retained as historical evidence for the pre-fix frozen source.
## Historical pre-fix result
- Frozen source under test: `b31b27e5845ffd3adf311429367319beaba263c7` on `feat/thoth-auth`.
- Freeze check: PASS. No tracked source changed during certification. The only untracked paths
retained are `.playwright-cli/` and `.thothctl/`.
- Certification window: `2026-08-18T09:26Z` to `2026-08-18T09:48:36Z` (UTC; the start marker is
minute-precision because no earlier second-level operator timestamp was captured).
- Overall result: `FAIL` / `CHANGES_REQUIRED`. The three Important findings are not closed and
authentication is not implementation-complete or release-complete.
## Local gate matrix
| Gate | Result | Sanitized evidence |
|---|---|---|
| Go focused security tests | PASS | `safeio`, `backup`, and `authstorage`; 3 packages |
| Go race/vet/host build | PASS | 18 race-tested packages; vet and host CLI build exit 0 |
| Windows amd64 cross-compile | PASS | focused safeio/backup test binaries and CLI build; compile-only |
| POSIX registry ownership | PASS | Node 24 backend suite includes local-registry ownership coverage |
| Unix StageArchive retained capability | PASS | focused safeio/backup and race coverage passed on host |
| Backend Node 24 | PASS | 76 files / 1092 tests; typecheck and build passed |
| Frontend Node 24 | PASS | 61 files / 444 tests; typecheck and build passed |
| Authentication/F1 browser smoke | PASS | Node `v24.16.0`; filtered E2E 1 passed; sentinel scan passed |
| Harness pytest | FAIL | 951 passed, 1 failed, 4 skipped, 232 subtests; `test_f4_emits_column_types` could not find `workflow.yaml` from its test cwd |
| Ruff | FAIL | 192 errors; known baseline |
| Authentication docs smoke | PASS | required-term and forbidden-word checks passed |
| Shell syntax | PASS | `bash -n scripts/*.sh` |
| Default Compose contract | FAIL | required `THT_WORKSPACE_GIT_REMOTE` was unavailable |
| Unified Compose contract | FAIL | `compose.unified.yaml` is absent from the frozen source |
| Unified Docker smoke | FAIL | workflow attempted it on the frozen SHA but stopped before deployment because `rg` was unavailable; cleanup proof passed and no new image manifest was generated |
| L2 / PSD manual / provider readiness | PENDING | required external secrets, identities/access, or provider prerequisites unavailable/not reached |
The first full backend Vitest attempt had one workspace-registry timeout. The focused test and a
fresh complete rerun passed, so the current backend result above is the fresh complete rerun.
## Native Windows authority
The authorized dispatch was bound to the frozen SHA:
- Run: `32122302381`
- URL: https://github.com/mptyl/ThothII/actions/runs/32122302381
- Head SHA: `b31b27e5845ffd3adf311429367319beaba263c7`
- Workflow conclusion: `failure`
- Job: `Windows clone and Compose contract`, job `95665197885`
- Job URL: https://github.com/mptyl/ThothII/actions/runs/32122302381/job/95665197885
- Native step: `Run native Windows retained-capability tests` — `failure`
- Executed command: `go test ./internal/safeio ./internal/backup -count=1`
- Observed focused failures include `TestRemoveCanonicalPrivateClaimRetainsParentDuringDeletion`
and `TestRemoveCanonicalPrivateClaimPreservesOrphan`.
- The backup package timed out in
`TestRestoreLifecycleLockExcludesCompetingTransactionsUntilTerminalCleanup` after `10m0s`.
- Additional backup failures included retained-staging `unsafe file` results, Windows temporary-file
cleanup reporting that a file was still in use, and fixture cases that could not read external
secret declarations. The first two categories are remediation/security-boundary failures; the
fixture declaration failures are recorded as an accompanying CI-fixture issue.
- `internal/authstorage` was not requested by the frozen workflow step and therefore has no native
Windows execution evidence. Cross-compilation does not substitute for this gate.
This native failure is the blocking gate. No source fix was attempted, and no later Docker smoke
was run locally after the failure.
## Other workflow failures
- `LF, Compose, docs, and TypeScript` (job `95665197839`) failed in
`Verify Compose and installation contracts` after the unified Compose contract itself passed.
`test-no-deployment-coupling-scope.sh` aborted on `TMPDIR: unbound variable`; this is classified
as a baseline/CI contract prerequisite, and later docs/TypeScript steps were skipped.
- `Linux Docker deployment and rollback` (job `95665197846`) failed before deployment because the
runner did not provide `rg` (`Task 13 smoke failed: rg is required`). The sanitized cleanup proof
passed and no Docker image manifest was generated. This is classified as an infrastructure
prerequisite failure, not as evidence of a remediation regression.
## Evidence and provenance
- Current machine-readable matrix: `.artifacts/task-15/automated-gates.json`; SHA-256
`6c516db5c2064c4a4a2e5f25961b993cd4a8fe020bbbb822fbac7faa0c119599`.
- Current requested report: this file (SHA-256 recorded after the evidence commit if needed for
external indexing).
- Current fix-round report:
`.superpowers/sdd/2026-08-18-thothii-authentication-remediation/fix-round-2-report.md`.
- Historical Docker image manifest: `.artifacts/task-15/unified-docker-images.json`, unchanged
because no new immutable-source Docker smoke ran. Its retained historical SHA-256 is
`9c8dec4546909fd93799dbcf374bcb3a89bc46cfe0fd482472c0cbe757ddf5b6`, bound to historical source
`74b062f1a737103524cbe706346cfd65f87cdfd1`, not to this Task 4 candidate.
- The historical Task 15 report remains provenance for earlier source SHAs; its current addendum
records this recertification separately.
No credentials, tokens, internal endpoints, provider identities, registry names, raw environments,
or browser traces are retained here.
## Separate verdicts
- Three Important findings: `CHANGES_REQUIRED`. Native Windows retained-capability authority
failed, and the frozen workflow omits the required `authstorage` package from its native command.
- Overall release readiness: `FAIL` with additional `PENDING` gates. The native Windows remediation
gate failed; the remote Docker attempt failed on a missing runner prerequisite; existing
Ruff/harness/Compose failures and external/manual prerequisites remain unresolved; and no
successful new unified Docker image evidence exists.
@@ -1,175 +0,0 @@
# Final-review fix report — Evidence #43–#46
Date: 2026-08-25
Base ThothII revision: `d4818c8cc33b8b11204377ff3cc65c6c9ee1425e`
Binding inputs:
- requirements: `docs/plans/2026-08-24-evidence-restructuring.md`;
- approved design: `docs/plans/2026-08-24-evidence-restructuring-design.md`;
- final review: `.superpowers/sdd/2026-08-24-evidence-restructuring/final-review.md`.
No PSD path was read or mutated during this fix wave. No issue was closed.
## Verdict
All three Important findings are fixed. Candidate evaluation remains mandatory for schema-v2
filesystem curated corpora and is absent for legacy v1, HTTP-only, and S3-only acquisition.
Reviewed formula migration now fails closed unless the caller supplies the exact original source,
the source parses back to the same formula, and every supporting excerpt is present after the
canonical mechanical normalization. Current owner-gate summaries explicitly distinguish the
immutable 225-test owner-gate observation from the expanded 230-test final-review verification and
place stale 222/224-era statements under an explicitly superseded historical section.
## RED/GREEN record
### Finding 1 — evaluator compatibility boundary
RED:
```text
cd harness
.venv/bin/pytest -q tests/test_preprocess_cli.py tests/test_evidence_formula_migration.py \
tests/test_evidence_restructuring_fixture.py \
-k 'candidate_evaluation_is_not_required or runtime_identity'
6 failed, 18 deselected, 1 warning
```
The v1 filesystem run still received a callable evaluator; v1 HTTP/S3 and v2 HTTP/S3 attempted to
resolve a filesystem evaluation root before the pipeline could run.
GREEN:
```text
6 passed, 9 deselected, 1 warning
```
`_requires_candidate_evaluation()` is now the single boundary used by both curated-corpus
validation and evaluator construction. It returns true only for `schema_version: 2` with a
filesystem source (including an explicit filesystem `source_root`). The existing integrated v2
candidate test remains the positive proof: it validates the corpus, performs all nine
candidate-generation searches while inactive, publishes only after PASS, and compensates a failed
candidate.
### Finding 2 — formula provenance
RED:
```text
cd harness
.venv/bin/pytest -q tests/test_evidence_formula_migration.py \
tests/test_evidence_restructuring_fixture.py
5 failed, 4 passed, 1 warning
```
The converter rejected the new `source_content` argument, returned clean Curated Evidence when no
original source was supplied, and the documentation consistency assertion still failed.
GREEN:
```text
cd harness
.venv/bin/pytest -q tests/test_evidence_formula_migration.py tests/test_formula.py \
tests/test_formula_wiring.py tests/test_preprocess_cli.py \
tests/test_evidence_candidate_publication.py
39 passed, 1 warning
```
The migration now:
1. returns `None` unchanged for `auto` and `draft` formulas;
2. requires original source content for a reviewed formula;
3. normalizes it with the authoring validator's `normalize_source_text()`;
4. parses it and requires equality with the supplied `ConceptFormula`;
5. requires one or more legacy provenance notes and verifies each note in normalized source text;
6. hashes the verified normalized original source, never `formula.dump()`;
7. returns `LegacyFormulaMigrationFailure(code="legacy_formula_requires_manual_review")` with
bounded problem codes for missing, invalid, mismatched, or excerpt-incomplete provenance.
The regression matrix covers exact original formatting, deterministic digest, absent original
source, absent excerpts, a YAML-folded excerpt that cannot validate literally, mismatched original
formula content, invalid full-query Formula Evidence, stable path-derived IDs, and unreviewed
session-only behavior.
### Finding 3 — current owner-gate record
RED:
```text
cd harness
.venv/bin/pytest -q tests/test_evidence_restructuring_fixture.py \
-k current_summaries
1 failed, 3 deselected, 1 warning
```
The current report did not contain the fresh expanded-suite result. Earlier RED also showed that
the leading report still lacked an authoritative read-only PSD summary and `PROJECT_STATE.md`
still said 222.
GREEN:
```text
4 passed, 1 warning
```
The task report now begins with one authoritative current record and moves all earlier scope/count
statements below `Historical record — explicitly superseded`. The manual and `PROJECT_STATE.md`
record both immutable observations without conflating them: 225 passed at the reviewed `d4818c8`
owner gate; 230 passed after five final-review regressions entered the selected acceptance files.
Both summaries describe the authorized PSD package as read-only, with 35 moveable sources plus the
retained README, narrow no-payload/no-vector observations, no secrets, and no mutation. Issue #47
still owns migration and manual acceptance.
## Requirement mapping
| Requirement | Implementation evidence | Verification |
| --- | --- | --- |
| v1 filesystem remains viable without `evaluation.yaml` or a new curated tree | `_requires_candidate_evaluation()` returns false for schema v1; `run_from_config()` passes `None` | v1 runtime identity test and parameterized filesystem case |
| HTTP/S3 preserve existing behavior | evaluator construction returns `None` for HTTP-only/S3-only configs in schema v1 and v2 | four parameterized HTTP/S3 cases |
| v2 filesystem evaluation remains mandatory | shared predicate drives validation and evaluator; no permissive missing-fixture path was added | integrated inactive-candidate publication/failure-compensation test |
| formula digest represents verified normalized original | caller supplies `source_content`; authoring normalization computes digest | literal SHA-256 assertion from independently normalized original fixture |
| formula and source cannot silently mismatch | original source is parsed and compared with the supplied model | `original_source_mismatch` regression |
| supporting excerpts are real and nonempty | notes are required and each normalized note must occur in normalized source | missing and folded/unverifiable excerpt regressions |
| unverifiable migration requires human resolution | every provenance failure returns the existing bounded manual-review result | exact code/problem assertions preserve original path and formula |
| current acceptance record is unambiguous | authoritative current sections plus explicit historical supersession | owner-gate document consistency test |
| PSD scope remains read-only | no PSD tool/path used in fix; current docs retain issue #47 authorization gate | diff review and acceptance isolation probe |
## Verification evidence
```text
focused Evidence compatibility/formula/candidate: 39 passed, 1 warning
owner-gate document contract: 4 passed, 1 warning
harness full: 1093 passed, 4 deselected, 53 warnings
harness Ruff: All checks passed!
acceptance runner: 230 passed, 1 warning; PASS
backend TypeScript: PASS
native tht build: PASS
git diff --check: PASS
```
The full backend Vitest gate remains outside this patch's modified files and reported 13 failures:
the ten previously recorded auth-runtime projection failures, the recorded Node 25 versus Node 24
contract failure, the recorded Argon2 401/429 timing assertion, plus one Windows timeout-helper
marker failure that is timing-sensitive and was not part of the prior stable 12-failure baseline.
The native Go suite with `-timeout 20s` reproduced the recorded authconfig timeout and backup/setup
failures. These unrelated failures were not repaired or hidden; backend TypeScript and native build
both pass.
## SHA-256 artifact hashes
```text
193f9e83634c17160dc5f589ee392f6cefbca2e24efa0770181500e7dce008a7 PROJECT_STATE.md
4f9cfb213fdfbab481fcad9eeb0002ab0bd69d83d7ef0c459c5fccd5f5b44048 docs/testing/evidence-restructuring-manual.md
e47862d4bf9b846f059ef133b77a6311fbeb297671843bfeda0d0445eb2e4933 harness/tht/cli/preprocess_cmd.py
4017f66dc6d3aeaa4f0294bbca581b8d447e3b60bd8daff4efff08048dc43d4b harness/tht/evidence/formula_store.py
35bd73f6d0759cf0c8cb15a9125cc97b5dda64eb02417eda9dd1a9c82ed8a558 harness/tests/test_preprocess_cli.py
a0affbf9213c41c95d5fea8e596ef37ae21836e0df9bdfa8f6adda95d9238fb6 harness/tests/test_evidence_formula_migration.py
6ed7d7679c860c6edc54cbcfc2a04db89c7d32ecab80c7daa72be96cdc14f164 harness/tests/test_formula.py
26d886074244be7cf6ca084f859dcbc0c0ade74c302077a506198a2e96bfc07e harness/tests/test_evidence_restructuring_fixture.py
62743e885230823f8942d14feadae511b6d09b5d391807a04d10a164f71dd919 .superpowers/sdd/2026-08-24-evidence-restructuring/task-12-report.md
e753674bb55d4880448a51e7fbdb3ccf506800ecc5fd3b73f59ed655cf9b63c7 tracked working-tree diff before this report
```
The final commit hash is reported with the completed handoff because a commit cannot include its
own hash without changing itself.
@@ -1,276 +0,0 @@
# Task 12 report — owner migration gate (#46)
## Current owner-gate record
This section supersedes every historical section below. The owner-gate run at `d4818c8` observed
**225 passed, 1 warning**. The final-review fix added five regressions to the selected files; its
fresh run of `bash scripts/evidence-restructuring-acceptance.sh` observed **230 passed, 1 warning**.
Both runs remained hermetic: they used only an isolated temporary authoring workspace and did not
receive or mutate a PSD path.
The authorized read-only PSD owner-gate package inspected the clean repository at immutable
commit `47516f85b4db4a67cfa8a86cea4cb2e7b98c5813`. It recorded 35 moveable source documents plus
the retained `psd-clinical/evidence/README.md` (36 current Evidence files total), and the narrow
legacy Qdrant baseline: 163 `schema_table`, 2,275 `schema_column`, 2 `memory`, and 1
`solved_question`. The replacement observations used exact filtered counts and at most three IDs
per kind with `with_payload:false` and `with_vector:false`. PSD Git status remained clean; no
secret was read and no external mutation, branch, migration, activation, commit, or push occurred.
The real Pi call, PSD migration, human Git review, authoritative pre/post vector inspection,
activation, and complete manual walkthrough remain pending in issue #47 until the owner
authorizes the migration window and exact target branch. The current package is read-only
preparation, not migration or manual acceptance.
## Historical record — explicitly superseded
Everything below preserves the sequence of earlier Task 12 runs for audit history only. Counts,
scope statements, and pending-inventory wording below are not current; the current owner-gate
record above and `docs/testing/evidence-restructuring-manual.md` are authoritative.
### Initial scope and boundary
Implemented only ThothII-local artifacts:
- `harness/tests/fixtures/evidence_authoring/poorly_structured.md` supplies prose, a
rough list, enum values, URL, ambiguity, and a PostgreSQL-expression candidate.
- `scripts/evidence-restructuring-acceptance.sh` uses a fake restructurer and an
isolated `mktemp` authoring workspace. It never receives, discovers, or accesses a
PSD path; it also asserts that the ThothII worktree status is unchanged.
- `docs/testing/evidence-restructuring-manual.md` is the owner-gate package and manual
procedure. It intentionally leaves the PSD 36-path inventory, PSD commit, and
before-counts pending authorization.
- `PROJECT_STATE.md` records only the observed hermetic result and says that issue #47
is pending.
No external PSD authoring repository was accessed, modified, staged, migrated, or
inspected. No GitHub issue was closed.
### Initial TDD record
RED:
```text
cd harness && .venv/bin/pytest -q tests/test_evidence_restructuring_fixture.py
FAILED: FileNotFoundError for fixtures/evidence_authoring/poorly_structured.md
```
GREEN:
```text
cd harness && .venv/bin/pytest -q tests/test_evidence_restructuring_fixture.py
1 passed
cd harness && .venv/bin/ruff check .
All checks passed!
```
### Initial hermetic acceptance evidence
```text
bash scripts/evidence-restructuring-acceptance.sh
PASS hermetic fake-restructurer: split, review, Git recovery/diff, no-op, dirty, upgrade, orphan
PASS isolation: only a temporary workspace was supplied; no external PSD path was read or written
222 passed, 1 warning
PASS evidence restructuring automated acceptance
```
The runner directly proves typed splitting; visible unresolved-review and orphan
validation failures; one-source manifest membership; recovery of the committed curated
baseline and a visible Git diff; unchanged no-op; dirty-state refusal; no-write
pipeline mismatch refusal and explicit all-source upgrade. Its selected suites cover
all eight typed kinds, no-tool/no-session Pi invocation, formula acceptance/rejection,
atomic semantic chunking and the 4,000-character policy, candidate evaluation and
inactive-generation failure behavior, deterministic query rendering, hybrid branch and
fused ranks, fail-closed search, and the pinned Qdrant Italian-BM25 L0 contract.
### Initial required local gates
| Gate | Result |
| --- | --- |
| `harness/.venv/bin/pytest -q` | PASS — 1080 passed, 4 deselected, 53 warnings |
| `harness/.venv/bin/ruff check .` | PASS |
| `backend/npx tsc --noEmit -p .` | PASS |
| `backend/npx vitest run` | FAIL — pre-existing failures listed below |
| `tools/tht/go build ./cmd/tht` | PASS |
| `tools/tht/go test ./...` | FAIL / timeout — pre-existing failures listed below |
| `bash scripts/evidence-restructuring-acceptance.sh` | PASS — 222 passed, 1 warning |
#### Baseline proof for unrelated failures
The immutable pre-task source `0caa747` was checked out to a temporary detached
worktree. Its targeted backend run reproduced the same 12 stable failures:
- all ten failures in `test/auth-runtime-projection.test.ts` (`authentication runtime
projection is invalid`);
- `test/health.test.ts` expects Node 24 but the host runs Node 25;
- `test/auth-routes-local.test.ts` expects the third Argon2 request to receive 429 but
receives 401.
The native host suite was run first unbounded and remained silent for more than seven
minutes, then was rerun with `go test ./... -timeout 20s`. Both the task worktree and
the detached `0caa747` baseline reproduce the same failures:
- `internal/authconfig: TestRunProjectedMutationHoldsOuterLockAcrossCanonicalAndProjection`
times out;
- `internal/backup` restore/auth publication assertions fail;
- `internal/setup` projected-server/root assertions fail.
These packages and backend files were not changed by Task 12. The suite failures are
therefore recorded as pre-existing concerns and were not repaired or hidden.
### Initial pending owner decision
Focused ThothII commit: `fc83d29b5836b4db173e0874689426ecf2da526f`
(`test(evidence): record restructuring acceptance`). Issue #46 is ready for the owner
migration authorization gate. Issue #47 remains pending: the owner must authorize the
migration window and exact PSD branch before any external repository inspection,
36-file inventory, migration, validation, activation, or manual PSD acceptance work
begins.
### Fix round 1 evidence (Task 12 review)
#### TDD
RED:
```text
cd harness && .venv/bin/pytest -q tests/test_evidence_candidate_publication.py tests/test_evidence_restructuring_fixture.py
FAILED: acceptance runner did not include test_evidence_candidate_publication.py
```
The first integrated-test draft also exposed that the real configuration refuses a
non-1024 embedding dimension; the fake was corrected to model the production contract
instead of bypassing configuration loading.
GREEN:
```text
cd harness && .venv/bin/pytest -q tests/test_evidence_candidate_publication.py tests/test_evidence_restructuring_fixture.py
3 passed, 1 warning
cd harness && .venv/bin/ruff check tests/test_evidence_candidate_publication.py tests/test_evidence_restructuring_fixture.py
All checks passed!
bash scripts/evidence-restructuring-acceptance.sh
224 passed, 1 warning
```
The new integrated test invokes the real v2 canonical-corpus validator, the real
`preprocess_cmd._candidate_evaluator`, and `CorpusPipeline`, with an observable store,
vector writer, and searcher. It records nine exact-generation branch searches (dense,
BM25, fused for lexical, semantic, and mixed queries), asserts that the candidate is
not ACTIVE during any search, publishes only after the PASS report, then proves a
failed candidate remains inactive and is compensated. It also asserts the 4,000 default,
rendered fragment bound, and rejection/absence of parallel vector-size aliases.
The acceptance probe now creates a real temporary Git repository, commits its initial
and proposed corpus, makes `evidence/curated` genuinely dirty, and calls
`prepare_workspace_evidence` with the default Git-status detection path. It restores
the temporary file before continuing; no injected `git_status` is used by the dirty
case, no-op, pipeline mismatch, or explicit upgrade checks.
#### Authorized read-only PSD preparation
The PSD repository was inspected read-only at immutable commit
`47516f85b4db4a67cfa8a86cea4cb2e7b98c5813`. `git status --porcelain` was empty before
and after, `git diff --quiet` succeeded after, and no PSD secret, worktree content
beyond the listed source names, or mutation command was used. The exact 36 paths,
concrete rollback command, and Qdrant baseline are now in
`docs/testing/evidence-restructuring-manual.md`.
Read-only Qdrant observation for `127.0.0.1:6333`, collection `psd-clinical`:
`schema_table=163`, `schema_column=2275`, `memory=2`, `solved_question=1`; the manual
records representative IDs. The standard `tht workspace vector inspect --json` command
was attempted, but its temporary maintenance container stopped before Qdrant access
because production auth configuration was unavailable. No secret was read to bypass
that guard; the owner-gate package explicitly requires repeating the contract command
immediately before authorized preprocessing.
Fresh no-write verification after the inspection:
```text
status_lines=0 diff_quiet_exit=0
head=47516f85b4db4a67cfa8a86cea4cb2e7b98c5813
```
Fresh required-gate evidence for this fix round:
```text
harness pytest: 1082 passed, 4 deselected, 53 warnings
harness ruff: PASS
acceptance runner: 224 passed, 1 warning
backend tsc: PASS
backend vitest: 12 failures, exactly reproduced at baseline 0caa747
Go build: PASS
Go test -timeout 20s: authconfig timeout plus backup/setup failures, exactly reproduced at baseline 0caa747
```
Fix-round commit: `8542124f2737354b7ae35973933fc38ca85a31c1`
(`test(evidence): harden owner gate acceptance`).
### Fix round 2 evidence (Task 12 re-review)
#### TDD
RED:
```text
cd harness && .venv/bin/pytest -q tests/test_evidence_restructuring_fixture.py
1 failed, 2 passed, 1 warning
```
The new manual-contract test failed because the package still described 36 source
documents and retained the old broad Qdrant request.
GREEN: the package now distinguishes exactly 35 moveable source documents from the
retained `psd-clinical/evidence/README.md` (36 current evidence files total). The
authorized migration instruction moves exactly the 35 listed documents to
`evidence/source/` and retains the README at its current path; the owner-only rollback
restores the immutable pre-migration SHA.
#### Narrow replacement Qdrant baseline
The prior `with_payload:true` full-collection scroll was out-of-scope and is withdrawn
as evidence. It is not a permitted fallback, and this report makes no claim that it
did not materialize payloads.
On 2026-08-25, the replacement baseline sent, for each of `schema_table`,
`schema_column`, `memory`, and `solved_question`:
```text
POST /collections/psd-clinical/points/count
{"filter":{"must":[{"key":"record_kind","match":{"value":"<kind>"}}]},"exact":true}
POST /collections/psd-clinical/points/scroll
{"filter":{"must":[{"key":"record_kind","match":{"value":"<kind>"}}]},"limit":3,"with_payload":false,"with_vector":false}
```
The resulting no-payload/no-vector observations were:
```text
schema_table=163: 01bc2535-24d6-5722-a58b-64a122b90b36,
024ddf60-c80d-5223-ac06-9247e7de7027, 086e00c8-b4a9-5063-b82d-e5a67add29cd
schema_column=2275: 00126cc1-7564-521a-a084-c2d670263258,
00365200-2c55-5bb4-86bb-87dd2d1bb529, 004304bc-b543-5b2d-8b40-f18da8e82af7
memory=2: 8d5cd772-563a-5e22-b764-2ca76cf6efca,
db74457a-3de8-5b95-9a31-d28a1ecf8141
solved_question=1: 2b6bb7d2-1a35-5f49-bd8d-b0cdcb98459a
```
No secret was read, no PSD mutation/stage/branch/commit/migration/activation/push command
was invoked, and the PSD repository remained at
`47516f85b4db4a67cfa8a86cea4cb2e7b98c5813` with empty porcelain status and a successful
`git diff --quiet` after the replacement queries. The standard `tht workspace vector
inspect --json` attempt remains blocked by missing production auth and must be repeated
successfully immediately before any owner-authorized preprocessing.
Focused correction gates:
```text
cd harness && .venv/bin/pytest -q tests/test_evidence_restructuring_fixture.py
3 passed, 1 warning
cd harness && .venv/bin/ruff check tests/test_evidence_restructuring_fixture.py
All checks passed!
bash scripts/evidence-restructuring-acceptance.sh
225 passed, 1 warning
```
Fix-round commit: `d4818c8cc33b8b11204377ff3cc65c6c9ee1425e`
(`docs(evidence): narrow owner gate baseline`).
@@ -1,137 +0,0 @@
# Adapter Foundations final-review fix report
Date: 2026-07-11
Branch: `codex/portable-deployment`
Worktree: `/Users/mp/projects/ThothII/.worktrees/portable-deployment`
Binding findings: `.superpowers/sdd/adapter-final-review-findings.md`
## Outcome
All seven final-review findings are addressed as one coherent adapter-foundations change:
1. HTTP vector reader and writer clients are independently optional. Capabilities reflect the
configured side; writer-only new and legacy configurations build successfully for targeted
writes; search without a reader raises public `VectorReadUnavailable`.
2. `VectorHealth` now reports read/write configured and reachable state independently, preserves
side-specific errors, and reports expected/observed embedding dimensions plus compatibility.
HTTP diagnostics cover read-only, write-only, both-up, and writer-down cases. Direct health
exposes its configured expected dimension without adding schema or migration work.
3. `ThothRestDwhAdapter` accepts `DatabaseIdentityConfig`, matching its resource contract.
4. Both vector adapters reject bools, floats, zero, and negative search limits using one exact
positive-integer guard.
5. Port tests explicitly cover public exports and frozen capability records.
6. A real `tht` subprocess test proves one legacy deprecation warning per config load on stderr
while JSON stdout remains parseable and uncontaminated.
7. The adapter plan and SDD progress explicitly constrain `build_vector_loader` to transitional
bulk sync and schedule its removal/migration in the local pgvector plan. Targeted memory and
solved-question writes remain on `build_vector_store(..., require_write=True)`.
No pgvector schema or migration changes were made.
## Files changed
- `harness/tht/ports/vector.py`
- `harness/tht/ports/__init__.py`
- `harness/tht/adapters/vector/thoth_http.py`
- `harness/tht/adapters/vector/legacy_direct.py`
- `harness/tht/adapters/factory.py`
- `harness/tht/adapters/dwh/thoth_rest.py`
- `harness/tests/test_vector_port_contract.py`
- `harness/tests/test_adapter_factory.py`
- `harness/tests/test_config_resources.py`
- `harness/tests/test_config_legacy_compat.py`
- `harness/tests/test_adapter_command_regressions.py`
- `harness/tests/test_dwh_port_contract.py`
- `docs/superpowers/plans/2026-07-11-adapter-foundations.md`
- `.superpowers/sdd/progress.md`
- `.superpowers/sdd/adapter-final-fix-report.md`
## TDD and verification evidence
RED:
```text
cd harness && .venv/bin/pytest tests/test_vector_port_contract.py \
tests/test_adapter_factory.py tests/test_config_resources.py \
tests/test_config_legacy_compat.py -q
```
Result: collection failed as expected because `VectorReadUnavailable` did not exist. After the
initial implementation, the same command exposed two expected contract/test-harness corrections:
dimension mismatch makes aggregate health unhealthy, and the installed CLI entry point is `tht`
rather than `python -m tht.cli`.
GREEN, covering adapter/config/command regressions:
```text
cd harness && .venv/bin/pytest tests/test_vector_port_contract.py \
tests/test_adapter_factory.py tests/test_config_resources.py \
tests/test_config_legacy_compat.py tests/test_adapter_command_regressions.py \
tests/test_dwh_port_contract.py tests/test_memory_save_one.py \
tests/test_solved_question.py tests/test_search_similar_kinds.py \
tests/test_vector_dual_key.py -q
```
Result: `66 passed in 0.45s`.
Docker availability:
```text
docker info --format '{{.ServerVersion}}'
```
Result: `29.4.1` (available; command required Docker socket access).
Full repository-default non-L2 harness suite, with Docker available for L0 tests:
```text
cd harness && .venv/bin/pytest -q
```
Result: `433 passed, 5 deselected, 17 warnings in 9.14s`. The five deselections are the configured
L2/live-service tests. Warnings are existing legacy-workspace `FutureWarning` emissions.
Scoped lint and diff hygiene:
```text
cd harness && .venv/bin/ruff check tht/ports tht/adapters \
tests/test_vector_port_contract.py tests/test_adapter_factory.py \
tests/test_config_resources.py tests/test_config_legacy_compat.py \
tests/test_adapter_command_regressions.py tests/test_dwh_port_contract.py
git diff --check
```
Result: `All checks passed!`; `git diff --check` produced no output.
## Commit
Commit subject: `fix(adapter): close final foundation review`
The report is part of that same final commit. A Git object cannot contain its own SHA without
changing that SHA; the exact resulting commit ID is therefore recorded in the task handoff from
`git rev-parse HEAD` after creation.
## Self-review
- Reader/writer separation is preserved: search dereferences only `_reader`; hashes/upsert only
`_writer`; health probes each configured client independently and never substitutes one result
for the other.
- Writer failure contributes to aggregate `ok=False`, even when the reader succeeds.
- Dimension compatibility is derived only from configured embedding dimension and existing
`list_tables` metadata. Missing metadata remains `None`, not a guessed success/failure.
- The shared limit guard uses `type(limit) is int`, intentionally rejecting Python booleans and
numeric coercions before either adapter reaches its transport.
- Existing JSON/CLI behavior is preserved; the subprocess regression parses stdout as JSON and
counts exactly one deprecation marker on stderr.
- Scope remains adapter foundations. No vector DDL, schema initialization, or migration work was
introduced.
## Concerns / follow-up
- Write reachability uses the existing `list_tables` diagnostic on the separately authenticated
writer client. Deployments must allow that non-mutating diagnostic RPC to the writer credential;
failures are intentionally visible rather than hidden by reader success.
- Existing legacy-workspace tests emit 17 `FutureWarning`s in the full suite. This wave pins the
required production stderr behavior but does not migrate unrelated test fixtures.
- `build_vector_loader` remains transitional technical debt only for bulk sync, explicitly assigned
to `2026-07-11-local-pgvector-profile.md`.
-206
View File
@@ -1,206 +0,0 @@
# Container Packaging Task 3 Report
## Status
Implemented the multi-stage core application image, non-root runtime, pinned Pi installation,
container entrypoint, context exclusions, and an in-image health smoke test.
## TDD / Build Evidence
Initial RED:
```text
docker build -f docker/core.Dockerfile -t thothii-core:test .
ERROR: failed to build: resolve : lstat docker: no such file or directory
```
The first sandboxed attempt could not access the Docker socket; the authorized rerun reached the
builder and failed for the expected reason: the Dockerfile did not exist.
GREEN build:
```text
sh -n docker/core-entrypoint.sh docker/smoke/core-smoke.sh
docker build --progress=plain -f docker/core.Dockerfile -t thothii-core:test .
```
Result: shell syntax exited 0; Docker build exited 0. A final rebuild after tightening
`.dockerignore` also exited 0 and transferred only 17.60 kB of changed context (the initial clean
build transferred 1.02 MB).
## Runtime and Entrypoints
- Runtime user is `10001:10001` (`thoth`), never root.
- Runtime contains Node `v22.19.0` and Python `3.12.13`. Python 3.12 is intentional because the
harness declares `requires-python = ">=3.12"` and also satisfies the deployment floor of 3.11+.
- Pi is installed exactly as `@earendil-works/pi-coding-agent@0.80.3`; its build-time and runtime
version probes both reported `0.80.3`.
- `server` starts `/app/backend/dist/server.js`; `doctor` routes to `tht doctor`; `preprocess`
routes to the future-facing `tht preprocess` command; explicit `tht ...` and arbitrary CLI
arguments route to the installed `tht` binary.
- The gate extension's `typebox` runtime dependency is installed from the harness lockfile.
## Smoke and Diagnostic Results
```text
docker run --rm thothii-core:test doctor
config: error - configuration is invalid or unreadable
data_root: ok
```
Result: expected exit 1 for absent mounted workspace configuration, with no traceback and no
secret-bearing validation detail.
```text
docker run --rm --entrypoint /app/docker/smoke/core-smoke.sh thothii-core:test
backend listening on http://127.0.0.1:8787
v22.19.0
Python 3.12.13
core smoke: ok
```
Result: exit 0. The script asserted non-root execution, `tht --help`, `pi --version`, runtime
version floors, and `GET /health` through curl. Fastify's returned display address was loopback;
the inspected container environment is `HOST=0.0.0.0`, and the compiled server passes that value
to `app.listen`.
```text
docker run --rm thothii-core:test tht --version
0.1.0
```
Result: arbitrary `tht` entrypoint exited 0.
An explicit runtime assertion checked UID 10001, exact Node and Pi versions, Python 3.11+, and the
absence of `/app/harness/.env` and `/app/harness/workspaces`; it exited 0.
## Image Size and Containment Inspection
```text
docker image inspect thothii-core:test --format '{{.Size}} {{json .Config.User}} {{json .Config.Env}}'
221419008 "10001:10001" [...runtime paths and version metadata only...]
```
Image size: **221,419,008 bytes** (about 211.2 MiB).
`docker history --no-trunc thothii-core:test` was inspected. It contains only Dockerfile commands,
the pinned public package name/version, base-image metadata, and non-sensitive runtime variables;
no credentials or customer paths were found. An in-image filename scan found only
`/app/harness/.pi/settings.json` among `.env`, key/certificate, and settings-name candidates; that
tracked Pi file contains theme/startup preferences, not secrets. The build asserts `.env` and
workspace directories are absent.
`.dockerignore` excludes VCS/agent state, all environment files except examples, package-manager
credential files, SSH/private-key and certificate formats, local virtualenvs/node_modules/caches,
backend runtime data, customer workspaces, sessions, artifacts, indexes, corpus, and deployment
mount content.
## Self-review
- `git diff --check` is clean.
- Entrypoint processes use `exec`, preserving container signal handling.
- Backend production dependencies are pruned; TypeScript build tools remain in the build stage.
- The writable `/data` root is owned by UID 10001; application payload remains root-owned and
read-only to the runtime user.
- CA certificates and curl are present for HTTPS integrations and health probing.
- No existing source, customer workspace, secret, or unrelated progress-ledger change is included
in the task commit.
## Concerns
- The `tht preprocess` command is deliberately a future-facing routing contract; its CLI group is
scheduled in the Evidence/preprocessing plan and is not implemented in the current harness.
- Python dependencies are range-resolved because the existing harness has no Python lockfile. The
Pi package, Node runtime, and package-lock-backed Node dependency sets are pinned/reproducible.
- The image was built and smoked on Docker Desktop arm64. The chosen official multi-arch base
images and Pi package are architecture-neutral at the package level, but amd64 still needs a CI
build/smoke before being advertised as verified.
## Reproducibility Review Fix
The original image pinned Pi's direct version in the Dockerfile but resolved its transitives at
build time, and pip resolved all harness dependencies from ranges. Both paths now consume committed
locks.
### Lock generation
Pi uses the minimal `docker/pi-runtime/package.json` and its committed npm v3 lock. It was generated
with:
```text
npm install --package-lock-only --ignore-scripts --no-audit --no-fund \
--prefix docker/pi-runtime
```
The package manifest specifies exact `@earendil-works/pi-coding-agent` version `0.80.3`; a lock
inspection confirmed that same resolved package version. Docker installs it with:
```text
npm ci --omit=dev --ignore-scripts --no-audit --no-fund
```
The Python lock was generated directly from the harness production metadata plus one explicit,
pinned PEP 517 build-backend input—not from a host `pip freeze`:
```text
uv pip compile harness/pyproject.toml docker/python-runtime/build-requirements.in \
--universal \
--python-version 3.12 \
--no-emit-package tht \
--generate-hashes \
--custom-compile-command \
'uv pip compile harness/pyproject.toml docker/python-runtime/build-requirements.in --universal --python-version 3.12 --no-emit-package tht --generate-hashes --output-file docker/python-runtime/requirements.lock' \
--output-file docker/python-runtime/requirements.lock
```
`pytest`, `ruff`, and `testcontainers` are absent. All production direct and transitive packages
are exact and hashed. `setuptools==80.9.0` is explicit so the local harness install can use
`--no-build-isolation` without an unpinned build-time resolution. Refresh instructions are in
`docker/LOCKS.md`.
### No-cache rebuild and verification
Final build command:
```text
docker build --no-cache -f docker/core.Dockerfile -t thothii-core:test .
```
Result: exit 0. The logs showed Pi `0.80.3`, Node `v22.19.0`, a hash-enforced Python dependency
install, explicit `setuptools==80.9.0`, and a non-isolated local `tht` wheel build. No isolated
build-dependency download occurred.
Fresh runtime checks:
```text
docker run --rm --entrypoint /app/docker/smoke/core-smoke.sh thothii-core:test
backend listening on http://127.0.0.1:8787
v22.19.0
Python 3.12.13
core smoke: ok
docker run --rm thothii-core:test tht --version
0.1.0
/opt/venv/bin/pip check
No broken requirements found.
```
An in-container package inspection reconfirmed Pi `0.80.3`. Non-root UID, runtime version floors,
doctor's expected concise exit 1/no traceback, `/health`, and arbitrary `tht` routing all passed.
The full filename containment scan found no `.env`, PEM, private-key, P12, or PFX file in `/app`;
`/app/harness/workspaces` remains absent. Image environment and `docker history --no-trunc` were
re-inspected and contain only public package/build commands and non-sensitive runtime metadata.
Final locked image size:
```text
220003986 10001:10001
```
That is **220,003,986 bytes** (about 209.8 MiB), 1,415,022 bytes smaller than the original image.
Remaining concern: the universal lock is resolved for Python 3.12 and includes hashes/markers for
all supported platforms, but only Linux arm64 has been built and smoked locally; amd64 remains a CI
verification gate.
@@ -1,82 +0,0 @@
# Container Packaging Task 4 Report
## Status
Implemented and verified runtime-configured frontend packaging.
## Changes
- Added the browser runtime contract `window.__THOTHII_CONFIG__.backendBaseUrl`.
- Loaded `/config.js` before the Vite module entrypoint.
- Made runtime configuration take precedence while preserving `VITE_BACKEND_URL` and the
existing `http://localhost:8787` client default for development and tests.
- Added a multi-stage frontend image that builds with Node and serves static assets as
unprivileged UID/GID `101:101` with nginx on port 8080.
- Added startup-time `BACKEND_BASE_URL` substitution (default `/api`).
- Added `/api/` reverse proxying to `core:8787`, SPA fallback, no-cache runtime config,
and SSE-safe proxy settings (`proxy_buffering off`, `proxy_cache off`, one-hour read timeout).
## TDD evidence
- RED: `npx vitest run src/api/runtime-config.test.ts` failed because
`./runtime-config` did not exist.
- GREEN: targeted runtime config suite passed (3 tests after preserving the legacy client
default).
## Verification
- `cd frontend && npx vitest run --reporter=dot && npx tsc -b && npm run build` — exit 0
(40 test files, 185 tests; TypeScript and Vite production build passed).
- `docker build -f docker/frontend.Dockerfile -t thothii-frontend:test .` — success.
- Image metadata reports `USER 101:101`.
- Two-container isolated-network smoke:
- `/config.js` returned `window.__THOTHII_CONFIG__ = { backendBaseUrl: "/api" };`
- `/api/health` proxied to the core image and returned `{"status":"ok"}`.
- an unknown nested route returned the SPA `index.html`.
- active nginx config contained `proxy_buffering off`, `proxy_cache off`, and
`proxy_read_timeout 1h`.
- `/config.js` returned `Cache-Control: no-store`.
- `sh -n docker/frontend-entrypoint.sh` and `git diff --check` — exit 0.
## Secret-leakage inspection
- `.dockerignore` excludes `.env*` (except examples), credentials/key formats, dependency
trees, build outputs, backend data, and deployment data.
- The runtime web root contained no `.env*`, `.pem`, `.key`, `.p12`, or `.pfx` files.
- Image history contained build/package instructions only; no secret build arguments or
credential values were introduced by this task.
## Self-review / concerns
- nginx resolves the `core` hostname at startup, matching the planned Compose service name;
standalone runs therefore need a reachable network alias named `core`.
- Existing frontend test warnings (React refs/act, MSW unmatched incidental requests, Vite
chunk-size warnings) remain; they did not fail the requested gates and are unrelated to
this task.
- `.superpowers/sdd/progress.md` was already modified by the orchestrator and was intentionally
excluded from this task's commit.
## P1 review fixes
Follow-up commit work addressed both review findings:
- Runtime configuration is now produced with `jq -cn --arg`, so `BACKEND_BASE_URL` is encoded
by a real JSON serializer rather than interpolated into JavaScript by `sed`.
- The image includes `frontend-config-smoke`, which strips only the fixed assignment wrapper,
parses the remaining JSON with `jq`, requires exactly the `backendBaseUrl` key, and compares
the decoded value to the environment input.
- The hostile smoke passed with quotes, backslashes, a literal newline, ampersand, pipe, and
`"; globalThis.PWNED=true; //` in the value. A breakout would leave non-JSON trailing input
and fail parsing.
- Added `joinBackendPath`, shared by API fetch and EventSource creation. It removes duplicate
boundary slashes for relative and absolute bases while keeping empty and `/` bases rooted.
Follow-up verification:
- RED: six join cases failed with `joinBackendPath is not a function` before implementation.
- Targeted: runtime config, API client, and EventSource suites — 14 tests passed.
- Full frontend gate — exit 0 (40 test files, 191 tests, TypeScript, Vite build).
- Rebuilt `thothii-frontend:test` successfully.
- Hostile config image smoke — `frontend runtime config smoke: ok`.
- Rebuilt two-container smoke — default `/api` config, proxied `/api/health`, SPA fallback,
and SSE-safe nginx directives all passed.
@@ -1,98 +0,0 @@
# Evidence / Preprocessing Task 1 Report
## Outcome
Implemented the additive Evidence source port and canonical corpus records. Existing evidence,
search, vector, and session runtime code is unchanged.
## Contract
- `EvidenceSource` is a runtime-checkable protocol with `discover` and `acquire` operations.
- `SourceObject` and `AcquiredDocument` are frozen, reject extra fields, use independent metadata
defaults, and restrict metadata to Pydantic `JsonValue` values.
- `CanonicalDocument`, `CanonicalChunk`, and `CorpusManifest` are frozen and reject extra fields.
- Provenance includes stable source IDs, canonical URIs, fingerprints, modification time, and
content hashes.
- Pipeline versions are recorded on documents, chunks, and manifests. Manifests also carry schema
version, optional publish ID/vector generation, and paired embedding model/dimension fields.
- Credential-like metadata keys are rejected recursively. Credentials are not model fields and
therefore cannot enter serialized canonical artifacts through extras.
## TDD evidence
The initial focused run failed during collection because `tht.ports.evidence` and `tht.corpus`
did not exist. After implementation, the focused suite passed.
## Verification
- Focused models/protocol tests: 13 passed.
- Harness excluding Docker-backed L0 and the network-dependent wheel packaging test: 444 passed,
5 deselected.
- Focused Ruff: passed.
- Full-repository Ruff remains blocked by 34 pre-existing findings outside the task files.
- An unrestricted `pytest -q` attempt reached 453 passed and 5 deselected, but reported 47 Docker
setup errors plus 4 Docker parity failures because the sandbox cannot access the Docker socket;
the wheel packaging test also failed because its isolated `uv build` needs unavailable network.
## Concerns / follow-up
- Pydantic's `frozen=True` prevents model field reassignment but does not recursively freeze list
and dict contents. `default_factory` prevents shared mutable defaults. Later pipeline stages should
treat these value objects as immutable and construct replacements rather than mutate collections.
- The adapter and normalization tasks should preserve the credential-free boundary by passing only
these records beyond acquisition.
## Review hardening follow-up
All six binding review areas were addressed in a separate TDD pass:
- JSON metadata is recursively converted to immutable `FrozenDict`/tuple values while retaining
stable object/array JSON serialization. Manifest document and chunk collections are tuples.
- Secret-key matching now normalizes camelCase and punctuation. It rejects credential-specific
names (passwords, API keys, access/refresh tokens, client/private keys, session cookies and
authorization) recursively, while deliberate benign labels such as generic `token` and `secret`
remain valid.
- Canonical URIs require a scheme and reject userinfo or credential-bearing query parameters.
- Namespaced IDs, SHA-256 content hashes, timezone-aware UTC timestamps, embedding/vector
compatibility, unique IDs, chunk referential/provenance integrity, contiguous per-document
ordinals and pipeline-version consistency are validated. Nested Pydantic instances are always
revalidated so `model_copy(update=...)` cannot bypass a manifest boundary.
- Acquired arbitrary bytes have explicit base64 JSON encoding and validation, covered by a JSON
round-trip test.
- `EvidenceSourceError` classifies transient/retryable versus permanent failures and exposes only
recursively immutable, credential-screened JSON details.
Follow-up verification:
- Focused contract suite: 39 passed.
- Focused Ruff: passed.
- Harness excluding Docker-backed L0 and the network-dependent wheel packaging test: 470 passed,
5 deselected.
- Fresh unrestricted harness attempt: 479 passed, 5 deselected; the same environmental boundary
remains (47 Docker socket setup errors, four Docker parity failures, one isolated `uv build`
network failure).
## Final blocker follow-up
The remaining four contract blockers were closed in a third TDD cycle:
- `EvidenceSourceError` now always exposes the fixed public message/`args` value `evidence source
operation failed`; caller diagnostics are not retained. Category, details and args cannot be
reassigned, details remain recursively frozen and credential-screened, and an original exception
is available only when callers use standard exception chaining.
- Canonical document/chunk provenance stores only URI scheme, authority and path. Userinfo is
rejected; query strings and fragments are removed unconditionally, including AWS `X-Amz-*`, SAS
`sig`, and fragment token material.
- Binding model bases override Pydantic's unchecked `model_copy(update=...)`: merged values always
pass full field/model validation, so invalid copied records and top-level manifests fail.
- A canonical document/chunk `content_hash` must equal SHA-256 of the exact stored text encoded as
UTF-8. This establishes the normalization boundary explicitly: line-ending/frontmatter/text
normalization happens before model construction; the canonical models never rewrite content.
Final follow-up verification:
- Focused contract suite: 45 passed.
- Focused Ruff: passed.
- Harness excluding Docker-backed L0 and network-dependent packaging: 476 passed, 5 deselected.
- Fresh unrestricted harness attempt: 486 passed, 5 deselected, with the unchanged environmental
failures (47 Docker setup errors, four Docker parity failures, one isolated `uv build` failure).
@@ -1,65 +0,0 @@
# Evidence Task 2 Report
## Status
Implemented filesystem and explicit-manifest HTTP Evidence source adapters, typed source
configuration with legacy compatibility, and factory construction.
## Delivered behavior
- Filesystem discovery is deterministic and rooted at a strict canonical directory.
- Symlink/path escapes are rejected before content is exposed.
- Discovery hashing and acquisition reads enforce a configurable byte limit.
- Filesystem fingerprints are content SHA-256 values; stable IDs derive from relative paths.
- HTTP accepts only explicit `http`/`https` manifest entries and keeps transport URLs private.
- HTTP provenance strips query strings/fragments, while config and adapter representations hide
signed or secret-bearing transport URLs.
- HTTP acquisition uses separate connect/read timeouts, streaming byte limits, bounded redirects,
private redirect rejection, and safe transient/permanent error classification.
- HTTP fingerprints prefer a deterministic ETag digest, then Last-Modified, then content SHA-256.
- `build_evidence_sources(cfg)` supports both typed `evidence.sources` entries and the legacy
`source_root` plus `evidence_dir` filesystem configuration.
## TDD and verification
- RED: focused tests initially failed during collection because the adapter package did not exist.
- GREEN: `15 passed` for filesystem, HTTP, and resource-config tests.
- Full harness: `548 passed, 5 deselected`.
- Changed-file Ruff: clean.
- Repository-wide Ruff remains non-clean due to 34 pre-existing findings in unrelated test files;
no unrelated lint files were modified.
## Notes
The approved `SourceObject` namespace grammar does not permit raw quoted ETags such as
`etag:"abc"`. The adapter therefore uses `etag:<sha256-of-opaque-etag>`: it preserves ETag-based
change identity without weakening the canonical contract or exposing validator contents.
## Review hardening follow-up
Four review findings were closed in a separate follow-up commit:
- Filesystem access now anchors a persistent descriptor at the canonical root and walks each
component with `openat` semantics (`dir_fd`, `O_NOFOLLOW`, and `O_DIRECTORY`). The regular-file
check, bounded read, metadata, and hash all use the opened descriptor. Acquisition reopens by
the same path-safe mechanism and rejects a changed fingerprint. Deterministic tests swap both a
leaf and an ancestor to symlinks at open time.
- HTTP network policy defaults to public hosts only. Initial URLs and every redirect reject
userinfo, mixed public/private IPv4/IPv6 answers fail closed, and the connected peer must be a
public member of the previously validated DNS answer set before any body bytes are consumed.
Explicit `allow_private_hosts: true` is required for trusted private deployments and local tests.
- Every HTTP response is closed in a `finally` block, including redirects, status failures,
policy failures, oversized bodies, and mid-stream exceptions.
- ETag and Last-Modified values remain adapter-internal. Repeated discovery and acquisition send
conditional headers; a 304 reuses only previously verified cached bytes and identity. The LRU
content cache has an explicit byte bound (`max_cache_bytes`). Validators are not forwarded
across redirect origins.
### Conditional cache binding correction
The conditional cache now binds bytes and validators to both the canonical provenance key and the
exact final effective representation URL. Redirect traversal recomputes request headers per hop:
validators are sent only when that exact URL matches the cached final URL, never merely because a
redirect retains an origin. A same-origin path change therefore downloads and replaces the body.
The adapter accepts 304 only when the exact request carried a bound ETag or Last-Modified validator;
unsolicited and cross-origin 304 responses are permanent protocol errors.
@@ -1,59 +0,0 @@
# Evidence Task 3 — deterministic normalization and chunking
## Outcome
- Added pure `normalize(acquired, pipeline_version)` and `chunk(document, policy)` transforms.
- Normalization enforces UTF-8 (including UTF-8 BOM), a 10 MiB input ceiling, LF line endings,
NFC Unicode, safe YAML frontmatter extraction, canonical provenance URIs, and hashes the exact
canonical UTF-8 text stored on the document.
- Undecodable, unsupported-charset, oversized, and invalid-frontmatter inputs fail explicitly;
byte content is never truncated.
- Chunking uses a versioned immutable policy, paragraph/word boundaries with deterministic
character-count hard splits for long tokens, contiguous ordinals, provenance metadata, exact
per-chunk hashes, and IDs derived from document hash + ordinal + policy version.
- Empty documents produce no chunks. Non-ASCII, CRLF equivalence, repeatability, policy changes,
duplicate-content ordinal collisions, and max-character limits are covered by tests.
## TDD evidence
- Initial focused test run failed during collection because both transform modules were absent.
- The EOF-frontmatter edge test was separately observed failing before its implementation.
- Final focused verification: `12 passed`.
## Verification
- `cd harness && .venv/bin/pytest tests/test_corpus_normalize.py tests/test_corpus_chunk.py -q`
— **12 passed**.
- `cd harness && .venv/bin/pytest -q` — **573 passed, 5 deselected**. The sandboxed attempt could
not access Docker; the approved rerun with local Docker access passed.
- Targeted Ruff over all four implementation/test files — **clean**.
- Full `cd harness && .venv/bin/ruff check .` — reports **34 pre-existing errors** in unrelated
legacy tests (unused imports and existing E702 semicolon lines); none are in Task 3 files.
## Concerns
- The 10 MiB normalization ceiling is deliberately explicit and independent of adapter download
limits. If deployment policy needs a different ceiling, it should become a versioned pipeline
configuration before ingestion is wired.
- Character limits use Python Unicode code points (`len`), not UTF-8 bytes or tokenizer tokens;
this is recorded in the chunk-policy metadata and tested with non-ASCII content.
## Review hardening follow-up
- Chunk IDs now bind the canonical document identity, document content hash, ordinal, chunk hash,
and a canonical SHA-256 fingerprint of every `ChunkPolicy` field. Identical content in separate
documents and same-version policies with different limits cannot collide.
- Boundary-aware slicing now retains separators in the slices. Concatenating every chunk exactly
reconstructs the canonical document for repeated spaces, tabs, blank lines, Markdown hard
breaks, fenced code, whitespace-only input, Unicode, and overlong tokens; every slice remains
within `max_chars`.
- Frontmatter uses a bounded `SafeLoader` variant: duplicate keys, anchors/aliases, structures
deeper than 20 nodes, and documents larger than 1000 composed nodes are rejected. YAML parse,
JSON type, credential-safety, and resulting canonical-model errors attributable to frontmatter
map to `PermanentNormalizationError(reason="invalid_frontmatter")`; invalid pipeline policy
remains a programmer-facing `ValueError`.
- Follow-up TDD evidence: the expanded focused suite first reported 11 expected failures against
the prior implementation, then passed **45/45** across normalization, chunking, and manifest
invariants.
- Follow-up full verification: **586 passed, 5 deselected**. Targeted Ruff is clean. Full Ruff
continues to report the same **34 unrelated pre-existing** violations in legacy tests.
-107
View File
@@ -1,107 +0,0 @@
# Evidence Task 4 — shared job envelope
Status: complete
## Delivered
- Immutable `JobSpec`, `JobRun`, `JobReport`, per-stage state, sanitized error, and UTC
timestamp records.
- `run_job(spec, stages)` with a durable checkpoint at job start, before and after every stage,
and at terminal state. Successful stages are skipped when a prior run is resumed.
- Atomic JSON checkpoint/report replacement using a unique same-directory temporary file,
file `fsync`, atomic `os.replace`, and parent-directory `fsync`.
- Public reports contain fixed operational fields only. Workspace paths, stage return values,
exception messages, source content, credentials, and arbitrary metadata are not serialized.
- `WorkspaceJobLock` uses non-blocking kernel `flock` on a stable workspace/job-specific inode.
Locks are released by the kernel on process exit; lock files are never removed based on PID,
avoiding stale-lock and PID-reuse deletion races. Evidence and DWH use distinct lock files.
- Dry-run intent is immutable in the spec/report and exposed to every stage through `JobContext`.
## TDD evidence
Initial focused collection failed because `tht.jobs` did not exist. Tests then drove:
- failure, sanitized reporting, resume, and idempotent successful-stage skipping;
- corrupt-checkpoint refusal before stage execution;
- JSON schema and path/secret/PII exclusion;
- dry-run propagation and ordered aware timestamps;
- multiprocessing exclusion, distinct Evidence/DWH jobs, traversal rejection, and recovery after
a lock-owning process crashes.
Final focused result:
```text
11 passed in 0.42s
```
## Verification
```text
cd harness && .venv/bin/pytest -q
597 passed, 5 deselected, 17 warnings in 28.45s
cd harness && .venv/bin/ruff check tht/jobs tests/test_job_runner.py tests/test_job_locking.py
All checks passed!
```
The full Ruff invocation was also run. It reports 34 pre-existing violations in unrelated legacy
tests; no Task 4 file is among them. L2 tests remain deselected by the repository configuration.
## Operational notes
- `fcntl.flock` intentionally targets the supported Linux/macOS deployment environments; it is not
a Windows locking implementation.
- The envelope does not publish or mutate an active corpus. Later pipeline stages must use
`JobContext.run_dir` for staging and perform their own final atomic publish only after validation.
- A dry run is an execution mode foundation: the runner exposes and records it; individual stages
remain responsible for suppressing external mutations.
## Review hardening follow-up
Four post-implementation findings were fixed test-first:
1. Resume compatibility is now a canonical SHA-256 fingerprint over checkpoint schema version,
hashed workspace identity, job type, dry-run mode, explicit spec/pipeline versions,
configuration/input fingerprints, and the exact ordered explicit `stage_ids`. Any insertion,
removal, reorder, mode, identity, version, config, or input change rejects resume before a stage
executes. Omitting `resume_run_id` remains the explicit safe path for a new run.
2. Lock traversal now uses directory file descriptors with `O_DIRECTORY` and `O_NOFOLLOW`.
Lock files use `O_NOFOLLOW | O_CLOEXEC`; `fstat` requires a regular file owned by the current
UID with one link, and permissions are forced to `0600` (`0700` for private directories).
Pre-existing lock-file and lock-directory symlinks are rejected.
3. Stage failures now serialize only the fixed safe tuple `internal` / `stage_exception` /
`stage execution failed`. Neither exception class names nor messages are inspected for output;
a hostile exception-name/message regression test proves a terminal failed report is retained.
4. Job/run directory creation is no-follow, owner-checked, private, and durable. Each newly created
parent is fsynced, the run directory is fsynced before the first atomic file write, and the
existing file-fsync → replace → directory-fsync ordering has an explicit regression test.
Follow-up verification:
```text
focused job/lock suite: 27 passed in 0.45s
full harness suite: 613 passed, 5 deselected, 17 warnings in 29.65s
Task 4 scoped Ruff: All checks passed
```
Repository-wide Ruff continues to report the same 34 unrelated pre-existing legacy-test findings.
## Final resume-integrity fix
Resume is now read-only until the source checkpoint proves trustworthy. The runner loads the source
before allocating a new run ID or directory, validates the exact stage state/timestamp/error ledger,
rejects duplicate stage identifiers, and recomputes compatibility from every persisted compatibility
field plus the exact ordered persisted stage IDs. It first requires the stored fingerprint to match
that recomputation, then compares the trusted recomputation with the requested job fingerprint.
Valid-JSON tampering tests cover removed, inserted/duplicated, reordered, and substituted stages;
input-field and stored-fingerprint changes; and invalid stage-state shapes. Every rejection occurs
before stage execution and asserts that the runs directory contains no orphan allocation.
Final verification:
```text
focused job/lock suite: 34 passed in 0.56s
full harness suite: 620 passed, 5 deselected, 17 warnings in 27.42s
Task 4 scoped Ruff: All checks passed
```
@@ -1,82 +0,0 @@
# Evidence Task 5 report
## Outcome
Implemented an incremental Evidence corpus pipeline with immutable materialized generations,
generation-scoped vector records, and an fsynced atomic `ACTIVE` pointer. Runtime Evidence
artifact lookup reads the active canonical manifest and keeps a legacy source-tree fallback only
when no corpus has been published.
The CLI is available as `tht preprocess evidence [--dry-run] [--resume RUN_ID] [--json]`.
JSON success and failure output is pristine and failure details are sanitized.
## Safety and failure model
- A workspace writer lock serializes preprocess writers; readers never take the lock.
- Generation directories, manifests, materialized files, locks, and `ACTIVE` reject symlink/path
escape cases and use owner-only durable writes.
- Vector records use generation-specific keys and metadata. The active manifest maps each active
document to its valid vector generation, allowing unchanged documents to retain their vectors.
- Runtime retrieval admits only active document IDs and their manifest-selected generations.
Removed documents and partial writes from failed generations are therefore unreachable.
- Embedding count and dimension checks occur before vector upsert; vector write count is checked
before staging/publish. Any failure leaves `ACTIVE` unchanged.
- Dry runs perform discovery/fingerprint planning only and never acquire, embed, write vectors, or
publish. Fully unchanged runs return the active generation without creating a replacement.
- Resume can safely retry idempotent generation-scoped upserts and publish an already staged,
compatibility-checked generation after a crash between staging and pointer replacement.
## TDD evidence
Initial focused collection failed because `tht.corpus.pipeline` and `tht.corpus.store` did not
exist. The implemented suite covers incremental skips, removals, model/policy rebuilds, acquire and
partial-vector failures, dry-run isolation, dimension validation, atomic reader snapshots, pointer
validation, symlink defense, and pristine CLI JSON.
Fresh focused verification:
```text
18 passed, 3 warnings in 0.39s
```
Command:
```text
.venv/bin/pytest tests/test_corpus_pipeline.py tests/test_corpus_publish.py \
tests/test_preprocess_cli.py tests/test_search_pack.py tests/test_session_documents.py -q
```
Scoped Ruff: `All checks passed!`
Broader non-Docker/non-packaging run reached `560 passed, 5 deselected`; ten pre-existing HTTP
adapter tests could not bind localhost under the sandbox. The complete suite reached `570 passed,
5 deselected`, with the remaining failures/errors caused by denied Docker socket, localhost bind,
and offline wheel-build access. No task-focused test failed.
## Remaining operational gate
Live pgvector integration needs Docker or an authorized local pgvector endpoint. The compensation
strategy is logical isolation rather than destructive cleanup because the shared `VectorStore`
port intentionally exposes no delete/transaction API; unreachable failed generations can be
garbage-collected by a future maintenance job.
## Review integration wave
Added an enforceable `metadata_filter` vector-port contract and capability flags. Direct pgvector
places exact Evidence generation/document predicates in SQL before `LIMIT`; HTTP sends the same
filter to the RPC and deliberately does not use the legacy 404 fallback. The reader RPC script now
validates and applies that filter. Normal Evidence search and search-pack use an ACTIVE-aware
searcher that groups active documents by generation, executes complete server-filtered searches,
and merges the results.
Added exact-generation Evidence cleanup to direct and HTTP writers plus the allowlisted writer RPC.
Pipeline failures compensate both staged filesystem state and vector writes; cleanup failures stay
sanitized and ACTIVE filtering remains the exposure boundary. Corpus-present session artifact
resolution now fails closed on corrupt/missing ACTIVE rather than falling through to source files.
Focused review-wave verification: 45 passed, scoped Ruff clean. A mocked REST regression proves
the exact filter payload and fail-closed legacy 404 behavior.
Still outstanding from the expanded review request: Task-4 JobRunner stage-by-stage integration,
published-generation retention/garbage collection, same-fd `dirfd` materialized-file reads, and
live local pgvector integration could not be completed in this wave.
@@ -1,94 +0,0 @@
# Evidence Task 5B implementation report
## Status
Integrated Evidence preprocessing with the Task 4 `JobRunner`. The CLI now accepts only a
32-character JobRunner run ID for `--resume`; generation IDs remain outputs. Runs persist the
exact ordered stages `discover`, `acquire_normalize_chunk`, `embed`, `vector_upsert`,
`stage_validate`, `publish`, and `retention_cleanup`.
Successful-stage artifacts are copied into the new resume run before execution, allowing later
stages to continue without rediscovery, acquisition, normalization, chunking, or embedding.
Job compatibility includes workspace, configuration, discovered-input, pipeline, embedding, and
chunk-policy fingerprints. Generation-specific filesystem/vector compensation is retained, and a
compensated generation is rotated before retry. `ACTIVE` is mutated only by `publish`.
Dry-run executes discovery/planning and makes every side-effecting stage a no-op. JSON output is
pristine and includes the JobRunner `run_id`, `resumed_from`, generation, plan, and publish status.
## TDD evidence
- RED: run-ID rejection and resume-artifact tests failed because generation IDs reached
configuration and resume runs had empty artifact directories.
- GREEN: the two regression tests passed after strict CLI validation and durable artifact carryover.
- Added pipeline job-plan and dry-run counting-fake coverage; both passed.
## Fresh verification
- Focused integration/search suite: `62 passed, 4 warnings`.
- Available harness suite excluding sandbox-blocked Docker, loopback HTTP-server, and networked
wheel-build tests: `559 passed, 5 deselected, 18 warnings`.
- Scoped Ruff: `All checks passed!`.
- `git diff --check`: clean.
## Environment limitations and concerns
The literal full harness invocation cannot complete in the managed sandbox: Docker socket access,
loopback HTTP test servers, and the `uv build` dependency resolution path are denied. It reached
`575 passed, 5 deselected` before those environment errors. The available-suite rerun above is
green.
One pre-existing Pydantic serialization warning is exposed by the new end-to-end job test when
canonical metadata contains frozen tuple values; it does not contaminate CLI stdout. Retention is
an explicit stable no-op until a retention policy is configured.
## Review fix wave — crash consistency and artifact integrity
Addressed all five follow-up findings:
- `JobRunner` now supports a test-only post-call/pre-checkpoint fault hook. Each stage seals a
canonical artifact manifest containing required flat filenames, SHA-256, byte size, producer
stage, and the full spec compatibility fingerprint. Resume validates the checkpoint and every
sealed artifact before allocating/copying a new run, rejecting missing, tampered, extra, nested,
or symlinked state. A sealed `running` stage is promoted after a simulated process crash; a
sealed `failed` stage is deliberately retried.
- Vector intent (exact record IDs and content hashes) is sealed before upsert. Execution reconciles
`existing_hashes` and writes only missing/mismatched rows. Crash-after-effect tests prove no
duplicate acquire, embed, or vector upsert.
- Raw upsert, stage, recovery-upsert, recovery-stage, and publish exceptions compensate the exact
generation. Compensation markers survive failed checkpoints; resume rotates the generation,
refreshes generation-bound artifacts, reconciles vectors, and stages idempotently.
- `CorpusStore.publish` is idempotent and failure-atomic. If replace succeeds but directory fsync
fails, it restores the previous `ACTIVE` value (or removes a newly created pointer), fsyncs the
rollback, and re-raises. Pipeline cleanup refuses to discard a generation referenced by ACTIVE.
- Added crash/resume coverage after all seven ordered stages; corrupt/missing plan, manifest, and
embeddings; unsafe extra paths; nonexistent run IDs; raw vector/stage failures; and post-replace
ACTIVE rollback.
Fresh fix-wave verification:
- Focused jobs/corpus/CLI/search suite: `82 passed, 17 warnings`.
- Available harness suite (same sandbox exclusions described above):
`579 passed, 5 deselected, 31 warnings`.
- Scoped Ruff and `git diff --check`: clean.
## Final P1 fix — effect state and checkpoint-bound manifest roots
- Stage checkpoints now distinguish `intent` from `completed`. Vector intent is atomically sealed
and checkpointed before upsert. A process-level `BaseException` after a partial multi-record
write leaves the stage `running/intent`; resume never promotes it and instead reconciles
`existing_hashes`, writing only the missing records. The completed state is persisted only after
reconciliation returns successfully.
- Every stage now persists its completed artifact state while still `running`, before the
post-call fault hook. The checkpoint binds the SHA-256 of canonical `artifact-manifest.json`,
effect state, exact producer stage, and exact required-file mapping. Resume validates this root
and all bindings before promotion or copying.
- Added process-interruption coverage proving the already-written vector record is not submitted
twice, remaining records are written, and publish completes only after reconciliation. Added
coordinated artifact/manifest, spec-binding, and producer-binding tamper rejection tests.
Fresh verification:
- Focused jobs/corpus/CLI/search suite: `86 passed, 18 warnings`.
- Available broad harness suite: `583 passed, 5 deselected, 32 warnings`.
- Scoped Ruff and `git diff --check`: clean.
-188
View File
@@ -1,188 +0,0 @@
# Evidence Task 5C report
## Delivered
- Added `vector.retain_published_generations` (default `3`, validation minimum `1`).
- Retention runs only after publication. It keeps ACTIVE, the newest configured generations,
and generations referenced by running or resumable failed job checkpoints.
- Cleanup deletes the exact Evidence generation from the vector store before removing its
immutable filesystem directory. Vector failures retain filesystem metadata for retry and
produce credential-free partial reports.
- Added idempotent `tht preprocess evidence gc [--dry-run] --json` reconciliation with pristine
JSON output.
- Materialized document reads now open generation/documents components with directory file
descriptors and `O_NOFOLLOW`, require a regular file owned by the process with one link, and
hash the bytes read from the same descriptor against the canonical manifest.
- HTTP generation deletion is pinned to `delete_vector_generation` with exact
table/kind/generation arguments. Legacy 404 responses fail closed with an actionable,
sanitized migration message.
## Evidence
- Focused retention, safe-read, CLI, and HTTP contract tests: `51 passed` (Docker-backed direct
parametrizations excluded from that focused invocation).
- Real Docker pgvector adapter suites: `33 passed`.
- Full harness suite, including Docker-backed tests: `668 passed, 5 deselected`.
- Changed-file Ruff: clean.
- `git diff --check`: clean.
The five deselected tests are the repository's opt-in `l2` tests requiring external services;
they are not local pgvector tests. Test output retains pre-existing Pydantic serialization and
legacy-config deprecation warnings.
## Review fix wave
- Publication is now explicit and durable (`PUBLISHED` marker). Retention candidates require a
valid generation manifest and publication marker (ACTIVE remains backward-compatible), so
staged and malformed directories neither consume retention slots nor become deletion targets.
- The policy retains ACTIVE plus exactly `N-1` newest rollback publications, ordered by durable
publication time and generation id. Running and failed-resumable JobRunner checkpoints protect
every referenced plan generation.
- `VectorStore` now exposes exact Evidence generation inventory. Direct pgvector uses a constrained
`SELECT DISTINCT` over `kind='evidence'` and `metadata.vector_generation`; HTTP uses the
allowlisted `list_evidence_generations` RPC and fails closed on legacy 404. The writer RPC SQL,
revokes, and grants are packaged in `create_vector_writer_rpc.sql`.
- Explicit GC reconciles the union of published filesystem generations and vector-only orphans,
preserving vector-before-filesystem deletion and retry semantics.
- `run_as_job` holds the same corpus writer lock across checkpoint recovery, staging, publish, and
retention. Explicit GC already uses this lock, serializing candidate snapshots with publishers.
- Session artifact consumers no longer receive the corpus source path after validation. They get
an owned, read-only copy atomically written from the bytes read and hash-validated on the same
descriptor.
Fresh verification after the fix wave: full harness `672 passed, 5 deselected`; Docker pgvector,
HTTP parity, and migration suites `43 passed`; exact direct inventory/delete integration `1 passed`;
changed-file Ruff and `git diff --check` clean.
## Final hardening verification
- Canonical generation validation is exact (`^gen:[0-9a-f]{32}$`) before HTTP/direct deletion;
malformed HTTP inventory rows fail closed rather than entering the GC candidate set.
- Added explicit protection coverage for running and failed-resumable JobRunner checkpoints, plus
a second-GC idempotence assertion for vector-only orphan reconciliation.
- Added deterministic concurrent locking coverage: a job paused after discovery retains the corpus
writer lock, explicit GC blocks, then completes after publication without deleting the active run.
- Added a descriptor-race regression: replacing the corpus pathname immediately after `read(2)`
leaves the atomically materialized session-owned copy byte-for-byte equal to the validated ACTIVE
document and its manifest hash.
Final fresh evidence: Docker pgvector/HTTP/migration suites `48 passed`; full harness `680 passed,
5 external L2 deselected`; changed-file Ruff and `git diff --check` clean.
## Integrated Task 5 dependency fixes
- GC now distinguishes filesystem retention from vector dependencies. ACTIVE and the newest
`N-1` published manifests keep their directories; every exact generation in their
`document_generations` maps remains vector-protected even after its old publication directory is
evicted. Job-protected manifests receive the same dependency treatment.
- The real four-publication Docker lifecycle now includes an unchanged document whose vectors come
from the first generation. With retention `N=2`, only the final two publication directories remain
while the first generation's vectors remain searchable from ACTIVE and survive restart/explicit GC.
- Evidence lookup is always wrapped by the ACTIVE-aware searcher. With no corpus/ACTIVE, Evidence
returns no rows and search packs cannot expose legacy vectors; non-Evidence kinds are unchanged.
- Session artifact resolution holds the corpus writer lock, snapshots the active manifest once, and
materializes bytes using that exact `manifest_id`, preventing a concurrent publish/retain-1 GC from
changing or deleting the selected source generation.
Focused unit tests, the updated real Docker lifecycle, changed-file Ruff, and `git diff --check` pass.
The final full harness invocation completed with exit code 0, including the concurrently added DWH
JobRunner tests.
## Final ACTIVE search review fixes
- `ActiveEvidenceSearcher` now treats default (`kinds=None`) and mixed-kind searches as explicit
split queries: non-Evidence kinds are queried separately, while Evidence is queried only with
ACTIVE manifest generation/document predicates applied server-side before every limit.
- Results are merged deterministically by descending similarity then stable id and truncated once
to the caller's global `top_n`. Pure non-Evidence searches retain their original delegate path.
- The corpus writer lock now covers manifest snapshot construction and all corresponding vector
queries, preventing retain-1 publication/GC from switching or deleting generations mid-search.
- Removed the public post-LIMIT `active_evidence_hits` helper; no public Evidence path performs
client filtering after limit.
Focused default/mixed/no-ACTIVE/search-pack tests pass, the real Docker pgvector lifecycle passes,
and the final full harness plus scoped Ruff/diff invocation completed with exit code 0.
## Workspace-scoped Evidence isolation
- Evidence manifests, vector metadata, and record keys now carry the stable JobRunner workspace id
derived from the configured workspace identity (config stem), never credentials or absolute paths.
- Every ACTIVE server-side predicate includes `workspace_id`. Legacy unscoped rows therefore fail
closed and cannot appear in Evidence results.
- Vector generation inventory and deletion require the workspace namespace across the port, direct
pgvector adapter, HTTP client/adapter, and allowlisted RPC SQL. Legacy unscoped RPC overloads are
explicitly dropped during migration; destructive SQL matches collection, kind, generation, and
workspace together.
- GC recovers the persisted namespace from ACTIVE for explicit/restarted cleanup and can only list
or delete that workspace's generations. Real shared-pgvector coverage proves deleting a generation
for workspace A preserves the same generation in workspace B.
- `PipelineResult.model_dump` now serializes fields explicitly instead of `dataclasses.asdict`,
avoiding deepcopy of immutable `FrozenDict` metadata while preserving pristine JSON CLI output.
Final focused verification: `89 passed` across corpus/CLI JSON, direct/HTTP parity, migrations, and
real Docker pgvector lifecycle; scoped Ruff and `git diff --check` clean. A contemporaneous full-suite
run reached unrelated Task 6 immutable-file tamper tests; those files were deliberately not changed.
## Immutable corpus/workspace binding
- A corpus root becomes bound to the workspace id persisted in its ACTIVE manifest. Job, non-job,
explicit GC, and ACTIVE search entry points compare the configured namespace before discovery,
vector access, staging, deletion, or ACTIVE mutation.
- Reusing the same paths after renaming a workspace now fails closed with a typed/sanitized message:
use a new corpus root or perform an intentional explicit rebuild. Unscoped legacy manifests also
fail this ownership check.
- Tests prove unchanged-document reuse cannot silently mix workspace A vectors into a workspace B
manifest, and that mismatched job, GC, and search paths perform no vector/filesystem mutations.
Focused workspace-binding, search-pack, preprocess JSON, and scoped Ruff/diff tests pass.
Compatibility follow-up: direct/internal `CorpusPipeline` instances now distinguish an omitted
workspace identity from an explicit config/job identity. An unbound instance adopts the persisted
ACTIVE owner (or `default` only for a brand-new direct corpus), preserving safe resume/GC tests and
the real pgvector lifecycle. Explicit config/job identities still fail closed on any mismatch. The
two reported regressions, workspace mismatch guards, real Docker lifecycle, scoped Ruff/diff, and
the full harness suite all pass.
Final fail-closed follow-up: persisted ACTIVE ownership is now validated under the corpus lock before
every configured search delegate, including default, mixed, pack, and non-Evidence-only operations.
Malformed or missing `metadata.workspace_id` is intrinsically rejected even for unbound direct
callers; source discovery, vector operations, GC, files, and ACTIVE remain untouched. Focused tests,
real Docker lifecycle, scoped Ruff/diff, and the full harness regression run pass.
Final lock/preflight follow-up: `CorpusPipeline.gc()` now acquires the corpus writer lock itself for
ownership validation through vector/filesystem cleanup. The store lock is thread-reentrant so nested
job retention is safe without weakening cross-thread/process exclusion; the CLI wrapper no longer
double-locks. Search find/pack performs locked corpus ownership preflight immediately after config
load, before DWH leasing, vector/searcher factories, embeddings, or schema work. Focused concurrency
and fail-closed tests, real Docker lifecycle, scoped Ruff/diff, and the full harness pass.
## Compact public Evidence reports
- Public `PipelineResult.model_dump()` is now a bounded operational envelope: terminal status,
run/resume/publication/generation/manifest identifiers, capped changed/unchanged/removed source
identifiers, and aggregate document/chunk counts. Full manifests, bodies, and metadata remain
internal/on disk and are never serialized to CLI stdout.
- `tht preprocess evidence` exits `1` for any durable terminal status other than `succeeded` in
both JSON and text modes. JSON stdout remains one pristine sanitized object; text mode emits one
compact stderr error without traceback, exception identity, evidence content, or credentials.
- Tests cover a real failed acquisition job, sensitive evidence content, capped thousand-item
summaries, bounded report size, and smoke-compatible changed/unchanged fields.
Focused tests and scoped Ruff/diff pass. The contemporaneous full suite reaches an unrelated Task 6
DWH snapshot fixture missing its newly required workspace identity.
### Safe result representation and exact text totals
- `PipelineResult.manifest` is explicitly excluded from dataclass representation and the custom
representation is fixed-size operational data only. It omits manifest ids, documents, chunks,
content, metadata, and errors; `str(result)` inherits the same safe representation.
- Text-mode Evidence success output reads the uncapped aggregate totals from `payload["counts"]`
rather than the intentionally capped identifier arrays.
- Regression coverage builds a thousand-document/chunk manifest containing content and
credential-like metadata secrets, checks bounded `repr`/`str`, and verifies exact totals above
the 100-item public-array cap.
Focused Evidence verification passes (`67 passed`), and scoped Ruff is clean. The full harness run
is not green in this sandbox: Docker-backed tests cannot access the daemon, wheel packaging cannot
use the restricted build environment, and concurrent Task 6 DWH binding changes currently fail two
DWH tests. None of those failures touch the Evidence files in this follow-up.
@@ -1,50 +0,0 @@
# Evidence Task 5D — Real pgvector lifecycle gate
## Status
Complete. The Docker-backed L0 gate uses one persistent `pgvector/pgvector:pg16`
database and the production migrations, direct reader/writer `PgVectorStore`,
`CorpusStore`, `CorpusPipeline.run_as_job`/JobRunner, ACTIVE Evidence retrieval,
search-pack fusion, owned session artifact copy, retention, and explicit GC.
## Lifecycle covered
- Four real corpus publications with retention set to two generations.
- A higher-similarity stale vector proves ACTIVE metadata filtering happens before LIMIT
for normal Evidence retrieval and the search-pack fusion path.
- A removed source is absent from ACTIVE retrieval and cannot be copied to a session.
- An injected process death occurs after one real committed vector upsert. Resume uses the
real run ID, preserves that record, fills the missing records, and produces no duplicate keys.
- Database engines and direct store objects are disposed/recreated before persisted ACTIVE
retrieval is checked again.
- An exact canonical vector-only orphan generation is discovered and removed by explicit GC.
- Filesystem and vector inventories converge exactly to ACTIVE plus one rollback; a second GC
is a no-op.
- Owned session artifact bytes and SHA-256 match the ACTIVE canonical document.
## Production bug found and fixed
Production migration `003_roles.sql` intentionally restricted `vector_writer`, but omitted
the privileges used by the production generation lifecycle: `SELECT(metadata)` for inventory
and `DELETE` for cleanup on `vectors.evidence`. Consequently a real job published successfully
and then failed in `retention_cleanup` on its first run.
Added versioned migration `004_evidence_generation_gc.sql` granting only those two Evidence
generation-management privileges. Runtime application code was not redesigned.
## Verification
- Target lifecycle: `1 passed` (Docker-backed).
- Full harness: `681 passed, 5 deselected`.
- Scoped Ruff: passed.
- `git diff --check`: passed.
The existing Pydantic serialization and legacy-workspace deprecation warnings remain unchanged.
## Follow-up assertion correction
The removal phase now retains the removed canonical document ID/ref before publication and
asserts both fields are absent from post-resume ACTIVE Evidence hits. It reruns the real
search-pack fusion after removal, proves active fourth-generation content is positively
returned in both paths, and proves the removed content remains absent. The owned session
artifact lookup for the retained removed ID remains empty.
@@ -1,49 +0,0 @@
# Evidence Task 6 — final fd-anchored DWH correction
All DWH generation state below `.tht-dwh` is now accessed relative to the directory descriptor
retained by the shared/exclusive generation lease. ACTIVE reads, atomic temp writes, replacement,
fsync, and rollback use `openat`/`replaceat` operations. Generation staging, validation,
reconciliation, resume checks, retention classification, and recursive deletion likewise use owned
root/generations/candidate descriptors with `O_NOFOLLOW`; locked operations no longer reopen
generation paths through `workspace_root`.
Portable reader snapshots are copied from validated generation file descriptors into private 0700
process-owned temporary directories while the shared lease is held. This avoids Linux-only
`/proc/self/fd` paths and prevents a renamed/replaced `.tht-dwh` pathname from redirecting later
schema or LSH reads. Lease-scoped copies are removed on exit and standalone snapshots are removed
at process exit.
Deterministic adversarial tests rename the DWH root after lease acquisition during ACTIVE reads,
ACTIVE publication, and retention cleanup. Each test proves the replacement tree is never read,
written, or deleted; the descriptor-pinned original either completes consistently or fails closed.
Existing owner binding, legacy rejection, crash reconciliation, resume, atomic rollback, retention,
and reader/writer exclusion behavior remains covered.
## Final review correction
Snapshot materialization now reads the manifest and every owned artifact exactly once through the
already-open generation descriptor, validates each hash against those exact bytes, and writes the
same byte objects to the private snapshot. A deterministic second-read mutation test proves hostile
pickle bytes can neither pass validation nor enter the snapshot. Reconciliation closes the ACTIVE
generation descriptor in a `finally` block on matches, mismatches, and exceptions. Pipeline-owned
snapshot directories are removed and deregistered after `run_job` on both successful and failed
runs, preventing repeated pipeline use from accumulating temporary directories or registry entries.
The cleanup boundary now begins immediately after snapshot materialization. Resume checkpoint
validation and `JobSpec` construction are guarded by the same release routine as `run_job`, so
corrupt/mismatched resume state or constructor failure clears the pipeline holder, removes the
private directory, and restores the snapshot registry to its prior state before propagating.
## Shipped preprocessing startup contract
Local-vector preprocessing now uses a dedicated Compose override. Both one-shot jobs depend on a
successfully completed `vector-migrate`, whose transitive chain waits for database health and role
reconciliation. The generic preprocessing overlay remains independently renderable and contains no
local-vector services or password secrets. README commands include the local override and build the
job image before running.
The real clean-project smoke no longer injects dependencies or manually starts, reconciles, or
migrates PostgreSQL. Its first shipped `compose run preprocess-evidence` demonstrably creates the
database, waits for health, runs reconciliation and migration, then runs the Evidence job. Unchanged
rerun, changed-source publish, DWH preprocessing, ACTIVE verification, and injected-failure cleanup
all pass through the same shipped dependency path.
@@ -1,93 +0,0 @@
# Evidence preprocessing Task 7 report
Implemented the S3-compatible Evidence adapter, explicit preprocessing Compose overlay, and
operational gates.
- S3 discovery uses bounded paginator pages, page size, and total objects; acquisition enforces a
byte ceiling and always closes streaming bodies.
- Provenance is canonical `s3://bucket/key`. Versioned objects use `s3-version:<version>`;
unversioned objects use a hashed exact ETag, and acquisition refuses validator drift.
- The adapter uses boto3/botocore rather than custom signing. TLS verification is enabled by
default. Custom HTTP and private endpoints require independent explicit opt-ins; endpoint
userinfo is rejected and public custom endpoints are DNS-policy checked.
- Access, secret, and session credentials support file-secret resolution into masked `SecretStr`
config fields. They are never emitted in provenance, reports, errors, or Compose environment.
- `deploy/compose.preprocess.yaml` provides separate one-shot Evidence and DWH jobs and is inert
unless explicitly included with the `preprocess` profile.
- `scripts/preprocess-smoke.sh` verifies both services render without secret material and pins an
unchanged rerun plus a modified generation through deterministic pipeline tests.
Verification: focused S3/HTTP/filesystem/config tests 34 passed; operational smoke 2 passed; core
image with locked boto3 extra built; full harness 702 passed, 5 deselected; scoped Ruff and diff
checks passed.
Operational risk: custom S3-compatible endpoints remain part of the deployment trust boundary.
Private endpoint access must be explicitly enabled and should be restricted by container egress
policy in production. S3 list consistency semantics are provider-defined; version IDs are preferred
over ETags wherever bucket versioning is available.
## Review correction
The Compose overlay now uses committed, purpose-built Evidence and DWH workspace files with
job-specific dependencies. Its services create their lock roots and mount only the vector secrets
they consume. The operational smoke is a real isolated Compose project: real pgvector migrations,
a deterministic in-project embeddings endpoint, actual Evidence CLI JSON across initial/unchanged/
mutated runs, exact ACTIVE verification, an actual DWH introspection job, and owned cleanup.
S3 custom endpoints now fail closed unless declared trusted; HTTP and private loopback endpoints
need additional independent opt-ins. Boto uses forced path-style addressing. Custom endpoints reject
userinfo, query, fragment, and non-root paths. Buckets use strict DNS syntax; listed keys must remain
under prefix and within the S3 byte bound; validators must be nonempty/bounded. Because
ListObjectsV2 does not provide version IDs, discovery honestly fingerprints the exact ETag and
acquisition rejects ETag drift.
Final correction verification: S3/config focused 20 passed; full harness 721 passed, 5 deselected;
real Compose smoke and image build passed; scoped Ruff, shell syntax, and diff checks passed.
## Final security review correction
Literal non-global IPv4/IPv6 endpoints now require the private-endpoint opt-in without claiming DNS
pinning for hostnames. Pagination uses explicit continuation requests and never fetches page
`max_pages + 1`. IP-shaped buckets, leading-slash prefixes, empty/overlong/control-character keys,
and absent validators fail closed. Acquisition accepts only the exact stored `SourceObject` and
compares the response ETag with the stored discovery validator. The real smoke snapshots generation
directory counts after every run and has an injected-failure cleanup mode; cleanup fails if Compose
down fails or any owned container, volume, or network remains.
The canonical smoke correction counts only root-level `corpus/gen-<32 hex>` directories. It exposed
that the durable job path still published an empty unchanged generation; the pipeline now returns
the existing ACTIVE generation without staging a directory when compatibility and all source
fingerprints are unchanged. The smoke therefore proves directory deltas `+1`, `+0`, `+1`.
Failure injection runs a real exit-97 command after resources exist and reaches the EXIT trap.
Cleanup aggregates Compose-down, residual container/volume/network, and temp-directory failures
while preserving the original failure status. S3 prefixes are validated before any client request
for leading slash, UTF-8 byte length, controls, and DEL.
## Canonical unchanged-run correction
The durable job now persists a deterministic source snapshot keyed by source identity. Each entry
binds canonical URI, exact source fingerprint, UTC modification time, canonical immutable metadata,
and explicit media type and size contract fields. The manifest also binds document-to-source
provenance, supplied config/input fingerprints, compatibility, embedding settings, and pipeline and
chunk-policy versions.
An unchanged run reuses ACTIVE only when ownership, bindings, the complete snapshot, document
provenance, materialized document hashes, and every required vector ID/content hash match exactly.
Snapshot changes rebuild only the affected sources; job input/config changes publish a new manifest
while retaining valid stable vector-generation dependencies. Missing or corrupt legacy contract
metadata, documents, or vectors fails closed and rebuilds. The Compose smoke now explicitly expects
the unchanged no-op to report `published=false` while proving generation deltas `+1`, `+0`, `+1`.
## Corrupt ACTIVE reconstruction correction
ACTIVE reuse now reconstructs each source contract from the persisted discovery snapshot and checks
the deterministic document identity, canonical URI, source fingerprint, UTC modification time,
source metadata, applicable media type, content hash, and pipeline identity against the owned
materialized document. The persisted document-source map carries the same exact binding.
Chunks are recomputed under the current chunk policy and must match the manifest exactly in count,
order, IDs, ordinals, content, hashes, linkage, provenance, and policy metadata. Vector health must
report the configured dimension, and every recomputed chunk must have its generation-scoped vector
ID with the exact content hash. Missing, altered, or extra chunks and corrupt document or vector
contracts therefore disable the no-op and rebuild, while a valid unchanged run still performs no
source acquisition.
@@ -1,16 +0,0 @@
# Model provider credential boundary
The backend accepts only an absolute `THT_MODEL_API_KEY_FILE` reference. `PiProcessManager` reads
and validates it afresh before each hosted-provider spawn, rejects symlinks, non-regular/hard-linked,
empty, whitespace-containing, oversized, unreadable, or permissively-mode files, and accepts Docker
0444 secrets only beneath `/run/secrets`. Failures are sanitized and occur before child creation.
Provider names are normalized and mapped to Pi-recognized variables. The child environment removes
the generic path, deprecated `PI_PROVIDER_API_KEY`, and all unselected known provider keys before
injecting only the selected key. Values never enter argv, settings, health, or diagnostics. Local
providers remain keyless and unknown hosted providers fail closed.
The production Compose overlay mounts `model_api_key` read-only and points the backend at its file;
the deployment render smoke proves the value is absent from rendered configuration. Entrypoint,
root README, Pi configuration guide, environment example, and secrets operator guide document the
new contract and reject the legacy generic value variable.
@@ -1,59 +0,0 @@
# Local pgvector whole-plan final fix report
## Outcome
All four binding final-review findings are closed.
1. `PgVectorStore.health()` checks namespace `USAGE` independently for reader and writer
before inspecting vector types. Real PostgreSQL tests revoke only schema `USAGE`, prove both
health sides false and operations unavailable, then grant it back and prove recovery.
2. Direct reader/writer passwords use workspace `password_file` references. Compose mounts the
two files read-only into core and exposes only `_FILE` paths. Rendered Compose and live
`docker inspect` checks prove secret contents are absent.
3. Direct search failures map to `VectorReadUnavailable`; hash/upsert failures map to
`VectorWriteUnavailable`. Messages are fixed and sanitized, original exceptions remain chained,
and upsert rollback is preserved.
4. The shared secret policy uses Linux `stat -c` with macOS `stat -f` fallback. Host files permit
only `0600`/`0400`; Docker's read-only `0444` is accepted only beneath `/run/secrets`. Tests and
operator docs pin this exact policy.
## TDD evidence
The new config, mode, schema-usage, unavailable-connection, and permission regressions failed
before their implementations. The first live secret-policy run also caught GNU `stat -f` accepting
an incompatible format invocation; detection now tries the native Linux form first. The next live
run caught smoke-generated rotation fixtures at `0644`; fixtures now model the documented host
policy.
## Verification
- Real direct pgvector + HTTP parity: `31 passed`.
- Full harness from `harness/`: `493 passed, 5 deselected`.
- Live `local-vector` rotation, restart persistence, inspect boundary, and backup/restore: pass.
- Core image vector migration discovery/status smoke: pass.
- External and local Compose deployment security contracts: pass.
- Config/port focused suite: `26 passed`.
- Secret policy, bootstrap rotation, and backup/restore safety scripts: pass.
- Changed Python Ruff, shell syntax, and `git diff --check`: pass.
One attempted full-harness invocation from the repository root produced a path-dependent failure
in an existing test that opens `workflow.yaml` relative to CWD. It was immediately rerun using the
documented `cd harness && .venv/bin/pytest -q` command and passed completely.
## Operational notes
Workspace files contain file paths, never direct passwords. Secret contents necessarily exist in
the in-process validated `DatabaseConfig` used to establish PostgreSQL connections, but are not
serialized by doctor/Compose/inspect paths. Docker Desktop file-backed secrets may appear as bind
mounts; the safe runtime exception is therefore based on the read-only service mount location
`/run/secrets`, while source files remain owner-only on the host.
## External-profile regression follow-up
Local pgvector is now an explicit `deploy/compose.local-vector.yaml` overlay. The base Compose and
production external override contain no direct vector password declarations, mounts, or `_FILE`
variables, so external deployments do not resolve or require local password files. A real lifecycle
gate unsets all local secret-file variables, renders external config, builds and starts core, waits
for health, and inspects the live container for absence of local direct-vector secret paths. The
local overlay retains its live inspect assertion (paths present, values absent), rotation, restart
persistence, and transactional backup/restore drill.
@@ -1,95 +0,0 @@
# Local pgvector Task 1 report
## Status
Implemented the direct `PgVectorStore` behind the transport-neutral `VectorStore` port.
The adapter uses separate optional reader and writer database configurations, derives
capabilities from configured authority, validates strict positive search limits, filters kinds
in SQL before limiting, and merges multi-collection results by cosine similarity.
All collection identifiers are selected from the fixed `schema_records`, `evidence`, and
`memory` allowlist and composed with `psycopg2.sql.Identifier`. Values, vectors, kinds, hashes,
and limits remain bound parameters. Collection/kind mismatches fail with `VectorStoreError`.
Upserts preserve the canonical metadata shape, use `record_key` conflict semantics, update the
transport hash and embedding, and leave semantic metadata fields intact. Health probes reader
and writer independently and reports observed `vector(N)` dimensions against the configured
embedding dimension.
## Configuration and factory
`pgvector_direct` now accepts explicit optional `reader` and `writer` `DatabaseConfig` entries.
The former `connection` entry remains supported as a deprecated read-only compatibility path.
`build_vector_store(..., require_write=True)` accepts writer-only direct configurations and
fails early when no explicit writer is present.
The transitional `build_vector_loader` bulk-sync path remains in place. It uses an explicit
direct writer when present, or the legacy `connection`; it deliberately does not treat a new
reader-only credential as writable. No production schema migration was added.
## TDD and verification
- RED: the new tests initially failed at collection because `PgVectorStore` did not exist.
- Docker L0 pgvector tests: `11 passed`.
- Direct + HTTP parity/factory/config focus: `51 passed`.
- Full harness: `461 passed, 5 deselected`.
- Changed-file Ruff lint: clean.
- Changed-file Ruff format check: clean.
- `git diff --check`: clean.
The repository-wide `ruff check .` still reports 34 pre-existing test-file findings outside
Task 1; none are in changed files. The full pytest suite emits 17 existing legacy-config
deprecation warnings.
## Scope and concerns
- Test fixtures create only the three existing vector tables needed to exercise the adapter;
migration/versioning remains Task 2.
- The legacy single `connection` form stays read-only through the public port, matching its
previous adapter behavior, while remaining available to the explicitly documented bulk-loader
transition.
## Review fix wave
The Task 1 review findings were addressed in a follow-up TDD cycle:
- Search now validates requested kinds against the global known-kind set, intersects valid kinds
with each collection, and skips unrelated collections. A direct-versus-HTTP parity test covers
the multi-collection case.
- Health requires all three allowlisted tables, an `embedding vector(N)` column on every table,
the expected dimension on every table, and the appropriate read or write table privileges for
each configured side. Empty and partial schemas return deterministic, credential-free details;
unexpected database failures expose only their exception class.
- The Docker L0 fixture now provisions separate least-privilege reader and writer roles. Tests
prove the reader cannot insert, the writer cannot execute the cosine-search SELECT, and the
adapter still routes search to the reader and upsert/hash operations to the writer. Direct
upsert uses an atomic `INSERT ... ON CONFLICT DO NOTHING` followed by `UPDATE` for an existing
key, avoiding broad SELECT authority while retaining conflict-safe hash/upsert semantics.
Fresh verification after the fix wave:
- Docker L0 + HTTP port/search parity: `42 passed` (earlier checkpoint); the final L0 file has
`16 passed` including the stricter raw-role search denial.
- Expanded focused adapter/config suite: `56 passed`.
- Full harness: `466 passed, 5 deselected`.
- Changed-file Ruff lint/format and `git diff --check`: clean.
## Sequence privilege health follow-up
Writer health now resolves the real serial/identity sequence for the `id` column of every
required collection using `pg_get_serial_sequence`. It requires `USAGE` on each resolved
sequence, which is the privilege used by the adapter's implicit `nextval`; sequence `SELECT` is
not required because no adapter operation reads sequence state.
The Docker fixture includes a writer role with complete table/hash-column authority but no
sequence grant. Its health is deterministically unhealthy and a new-key upsert fails. Granting
only sequence `USAGE` makes health green and the same port upsert succeeds. Sequence discovery is
guarded for partial schemas so a missing `id` column produces the existing sanitized schema
diagnostic instead of a PostgreSQL error.
Fresh verification for this follow-up:
- Docker pgvector L0 after formatting: `17 passed`.
- Expanded focused adapter/config/parity suite: `57 passed`.
- Full harness: `467 passed, 5 deselected`.
- Changed-file Ruff lint/format and `git diff --check`: clean.
@@ -1,82 +0,0 @@
# Local pgvector Task 2 report
## Outcome
Implemented ordered, idempotent production migrations and the `tht vector migrate`
interface, including `tht vector migrate --status --json` with pristine JSON output.
## Implementation
- `001_extensions.sql` installs pgvector.
- `002_schema_tables.sql` creates `vectors.schema_records`, `vectors.evidence`, and
`vectors.memory` with the `VectorWriteRecord` columns and `vector(768)` embeddings.
- `003_roles.sql` creates passwordless `NOLOGIN` reader/writer roles. Deployments inject
credentials (or grant these roles to separately-created login roles); no production secret
is stored in the repository.
- Reader authority is schema usage plus table `SELECT`.
- Writer authority is schema usage, table `INSERT`/`UPDATE`, narrow hash-probe column `SELECT`,
and sequence `USAGE`. It has no `DELETE`, broad row `SELECT`, DDL, or ownership authority.
- The migration runner discovers ordered SQL files, records SHA-256 checksums in
`public.tht_vector_migrations`, serializes runners with a transaction-scoped advisory lock,
and applies the full pending batch in one transaction.
- Status distinguishes applied, pending, and checksum-drifted migrations. Apply refuses drift.
A failed migration rolls back both prior migrations in that batch and ledger writes.
## TDD evidence
RED was observed with a real `pgvector/pgvector:pg16` testcontainer: 6 failures for the missing
module, missing command, and missing schema.
GREEN verification:
- Focused migration + direct adapter integration: `23 passed`.
- Full harness from the documented `harness/` cwd: `473 passed, 5 deselected`.
- Targeted Ruff (`tht` plus the new L0 test): clean.
- `git diff --check`: clean.
The new L0 coverage exercises clean install, idempotent rerun, pristine JSON status, checksum
drift, transaction rollback, exact tables/columns/dimensions, role isolation, sequence authority,
and the real `PgVectorStore.health()` plus `VectorWriteRecord` upsert path.
## Existing repository lint baseline
The requested full `ruff check .` was run. It reports 34 pre-existing violations in unrelated
test files (unused imports and one-line semicolon statements). None are in Task 2 files; changing
them would exceed this task's scope. The complete harness test gate is green.
## Self-review
No unresolved Task 2 correctness concern found. One deliberate contract choice is worth noting:
writer `INSERT` and `UPDATE` are table-level because the approved direct adapter health probe uses
`has_table_privilege` for those authorities. Least privilege is retained by withholding broad
`SELECT`, `DELETE`, DDL, ownership, and credentials.
## Review fix wave
The post-implementation review found four production-boundary gaps. They are fixed as follows:
- Migration SQL now ships inside the `tht` wheel (`tht/migrations/vector`) via explicit
setuptools package-data and is discovered through `importlib.resources`, rather than relying on
a source-checkout-relative directory.
- Both status and apply reject ledger versions absent from the installed manifest, including
nonnumeric future version labels. This treats a binary/database downgrade as drift instead of
silently reporting a healthy state.
- Migration files are ordered by parsed integer version; spellings such as `2` and `02` are
rejected as duplicate versions.
- Every migration transaction pins `search_path` locally to `pg_catalog, pg_temp`; catalog calls
and the ledger are schema-qualified. pgvector is installed into the locked `vectors` schema,
tables use `vectors.vector`, and `PgVectorStore` qualifies vector casts and the cosine operator.
A hostile admin default path with a writable shadow schema cannot redirect migration objects.
- The core image build asserts CLI discovery. Image verification now starts an ephemeral pgvector
database, runs the installed image's migration command, and compares pristine apply/status JSON.
Additional verification after the fix wave:
- Focused migration, adapter, hostile-path, and wheel suite: `27 passed`.
- Full harness: `477 passed, 5 deselected`.
- Production core image build: passed, including build-time CLI discovery.
- Core-image apply/status smoke against `pgvector/pgvector:pg16`: passed.
- Changed production and test files: Ruff clean; `git diff --check` clean.
- Full Ruff remains at the same 34 pre-existing unrelated test-file findings documented above.
No dependency changed, so the committed Python requirements lock did not require regeneration.
-133
View File
@@ -1,133 +0,0 @@
# Task 3 report — optional local pgvector profile
## Status
Implemented and verified the `local-vector` Compose profile.
- `vector-db` uses pgvector 0.8.5 on PostgreSQL 16, pinned to the official multi-arch
manifest digest.
- `vector_data` is a project-scoped named volume and is not shared with application data.
- database readiness gates the packaged one-shot `vector-migrate` job; core declares the
migration completion dependency while remaining usable in the pre-existing external profile.
- bootstrap, migrator, reader, and writer identities are distinct. Bootstrap and migration
credentials are supplied as Compose secrets; the application receives only reader/writer
credentials.
- `deploy/workspaces/local-vector.yaml` selects `pgvector_direct` with separate reader and
writer connections.
- the base loopback port binding, `AUTH_MODE=none`, and `THOTH_PUBLIC_EXPOSURE=false` defaults
are unchanged.
## Red/green evidence
The initial Compose contract did not list `vector-db`, as required by the brief. The first real
smoke then failed migration 002 because bootstrap installed the vector extension in `public`.
The bootstrap was corrected to create the `vectors` schema under the migration owner and install
the extension there. A clean-volume rerun passed.
## Verification
- `./scripts/local-vector-smoke.sh`: PASS
- isolated generated Compose project and credentials
- clean migration plus idempotent status rerun
- reader/writer privilege health
- one-record upsert and similarity search
- restart of both `core` and `vector-db`
- persisted search result after restart
- project-only volume cleanup
- `./scripts/test-container-deployment.sh`: PASS
- `./scripts/test-backend-url-policy.sh`: PASS
- `docker compose --profile local-vector config --quiet`: PASS
- harness: 477 passed, 5 deselected
- backend: 84 passed; TypeScript typecheck PASS
- frontend: 226 passed; TypeScript typecheck PASS
- `git diff --check`: PASS
## Self-review / concerns
- Compose cannot make a dependency required only under one profile. The core dependency uses
`required: false` so the established `external` profile does not activate local infrastructure;
under `local-vector`, `compose up --wait` still fails if `vector-migrate` exits nonzero, and the
smoke verifies that successful migration precedes the healthy stack.
- Reader/writer passwords are injected into core environment variables because Compose service
attributes cannot be conditional by profile. Bootstrap and migrator credentials remain
file-backed secrets and are never exposed to core.
- The smoke intentionally refuses the operator project name `thothii` and removes only its unique
project namespace and volumes.
## Follow-up hardening — credential reconciliation and cleanup ownership
Review findings were resolved in a separate follow-up:
- Replaced fresh-volume-only initialization with `vector-reconcile`, an idempotent one-shot that
runs after database health and before `vector-migrate`. It authenticates with only the bootstrap
admin secret, safely creates missing identities, reconciles role attributes and passwords on
existing volumes, restores memberships/ownership, and leaves vector data untouched.
- The migrator is explicitly `NOSUPERUSER NOCREATEDB NOCREATEROLE`. Schema/database ownership is
sufficient for all packaged migrations because reconciliation creates the two group roles first.
- The live smoke rotates migrator, reader, and writer secrets on the same populated volume, rejects
the old reader credential, reruns migrations, recreates core with the new runtime credentials,
and retrieves the record written before rotation and again after database/core restart.
- Smoke project names are no longer caller-controlled. Each run creates a unique namespace and
ownership token. Containers, networks, and volumes carry the ownership label; preflight refuses
any collision and cleanup verifies every discovered resource before `down --volumes`.
- Added a dynamic fake-Docker contract suite for caller override, collision, and mismatched cleanup
labels, plus a real-Docker collision probe using a unique labeled volume.
Follow-up verification:
- `./scripts/local-vector-smoke.sh`: PASS, including live secret rotation and persisted retrieval
- `./scripts/test-local-vector-smoke-safety.sh`: PASS
- `./scripts/test-local-vector-smoke-live-collision.sh`: PASS
- harness: 477 passed, 5 deselected
- backend: 84 passed; TypeScript typecheck PASS
- frontend: 226 passed; TypeScript typecheck PASS
- Compose security, backend URL, config, shell syntax, and diff checks: PASS
Remaining operational constraint: the bootstrap admin secret must continue to match the PostgreSQL
bootstrap account stored in the volume. Runtime migrator/reader/writer rotation is supported without
data deletion; bootstrap-account password rotation is a distinct database-administration operation.
## Final hardening — bootstrap account rotation
The remaining operational constraint is now covered by
`scripts/vector-rotate-bootstrap-password.sh OLD_SECRET_FILE NEW_SECRET_FILE`:
- It does not rely on `POSTGRES_PASSWORD_FILE` after initialization.
- It pre-stages the deployment-file replacement in the same directory, authenticates to the live
database with the explicit old file, and changes only the authenticated bootstrap role.
- Passwords are passed as connection parameters and rendered with psycopg2 SQL composition, so
shell and SQL metacharacters are not interpolated.
- A second connection must authenticate with the new password before the command succeeds. If that
verification fails, the still-open old connection restores the old database password.
- Only after verified database login does an atomic rename replace the current deployment secret.
Wrong-old authentication and verification failures leave deployment configuration unchanged.
Final live smoke evidence on one existing `vector_data` volume:
- wrong-old bootstrap rotation rejected; current deployment secret unchanged
- bootstrap password with quote characters rotated successfully
- old bootstrap login rejected and new login accepted
- `vector-reconcile`, packaged migrations, and core health passed afterward
- the vector record written before rotation remained searchable after rotation and after a further
database/core restart
Final tests:
- `./scripts/test-vector-bootstrap-rotation.sh`: PASS
- `./scripts/local-vector-smoke.sh`: PASS with negative and positive live bootstrap rotation
- existing local-vector collision/safety and Compose deployment contracts: PASS
## Final identity and secret-policy alignment
- `THT_VECTOR_BOOTSTRAP_USER` is now passed through core as well as vector-db and reconciliation,
so the rotation helper uses the authoritative configured role instead of defaulting to `postgres`.
- Rotation and reconciliation source the same raw-file `secret-policy.sh`: non-empty and no
whitespace, including trailing newlines. Rotation validates both files before Docker,
PostgreSQL, or atomic replacement staging; `test-vector-secret-policy.sh` pins empty, newline,
internal-space, and valid metacharacter cases.
- Fake-Docker tests prove a non-default identity reaches the helper path and whitespace rejection
performs no Docker call and creates no staged replacement.
- The real smoke runs the entire stack as `thoth_bootstrap_smoke`. Its whitespace-negative case
leaves the deployment file unchanged and proves the existing database login still succeeds;
non-default-account bootstrap rotation, reconciliation, migration, core health, restart, and
persisted retrieval all pass.
@@ -1,94 +0,0 @@
# Local pgvector Task 4 report
## Outcome
Implemented adapter parity gates and an operator-safe custom-format backup/restore workflow.
- Direct and HTTP stores now share validation, configured-dimension rejection, and deterministic
similarity ordering with record ID as the tie-break.
- The parity fixture exercises identical records through real pgvector and the HTTP RPC contract:
kind filtering, ordering, hashes, replacement upserts, invalid collection/kind errors, and query
plus write dimensions.
- Backup explicitly allowlists the three vector tables and migration ledger, refuses overwrite,
writes through a partial file, and uses a custom compressed archive.
- Restore requires explicit active-source and target coordinates. It compares PostgreSQL system
identifier plus database OID (robust across DNS aliases), refuses the active database, checks for
an empty target unless force is explicit, and restores with exit-on-error.
- Passwords are accepted only through validated secret files, converted to private temporary
`PGPASSFILE`s, and never placed in command arguments or success/error logs.
- Role passwords/login identities are deliberately not dumped. The target must have the approved
passwordless group roles and pgvector extension reconciled before restore; archived ACLs restore
the reader/writer grants.
## TDD and semantic alignment
The first parity run exposed the intended HTTP differences: it accepted unknown collections and
wrong dimensions. Direct pgvector also had no stable order for equal cosine distance. The adapters
were aligned, and the final focused real-pgvector gate passed: **25 passed**.
The first recovery run caught an incorrect probe username before restore. The second caught an
intersection between `pg_dump --schema` and the explicit public ledger table. The third confirmed
the archive contents but caught missing target group roles. Each defect was corrected and the
complete drill was rerun from a fresh generated project.
## Live recovery smoke
`./scripts/local-vector-smoke.sh --backup-restore`: **PASS**.
- generated/owned source Compose project and source `vector_data`
- distinct restore container and distinct named restore volume
- migration and role health, secret rotation, restart persistence
- real custom backup, then deliberate mutation of the active source record
- same-database identity guard evaluated before restore
- restore into the separate target only
- restored hash equals the pre-mutation backup, proving retrieval parity
- migration ledger has all three applied versions
- all three restored embedding columns report `vectors.vector(768)`
- ownership-checked cleanup; the active operator project/volume is never addressed
## Verification
- parity + direct adapter: 25 passed
- full harness: 485 passed, 5 deselected
- changed Python files: Ruff clean
- shell syntax: clean
- `git diff --check`: clean
- full Ruff: unchanged repository baseline of 34 unrelated pre-existing test-file violations
## Self-review and operational constraints
The restore account must be able to read `pg_control_system()` for the robust cluster-identity
comparison and create/restore the selected objects. This is intentionally an administrative
recovery operation, not a runtime reader/writer action. `--force-nonempty` is explicit but still
uses `pg_restore --clean --if-exists`; operators should prefer a new database/volume and validate
migration status, health, and known retrieval before endpoint cutover.
## Post-review hardening
All five final review findings were addressed in a follow-up commit:
- Restore now requires a physically separate PostgreSQL cluster and refuses any equal
`system_identifier`, independent of database OID or hostname.
- `pg_restore` combines `--single-transaction` with `--exit-on-error`. The live drill creates an
existing vector sentinel, deliberately fails late during a forced restore, and proves the
original sentinel row/hash remains unchanged before performing the successful restore.
- Backup uses a mode-0600 `mktemp` in the output directory, atomically renames it, and cleans only
that owned path. A fake-command test pins symlink-clobber resistance and preserves an adversarial
legacy `.partial` symlink and its target.
- HTTP parity now traverses the real `VectorRestClient` transport boundary. It asserts RPC URL/key
and kinds payloads, legacy 404 fallback, response conversion, malformed metadata tolerance, and
canonical `VectorRestError` to `VectorStoreError` mapping.
- The restored target runs role/secret reconciliation and a real `PgVectorStore` with separate
reader/writer logins. Health, known-record search, writer upsert, hash probe, schema/table/column/
sequence authority, and 768-dimensional compatibility are therefore verified through the
production adapter. Reconciliation now restores group-role schema `USAGE`, which table-selected
archives cannot carry.
### Atomic no-replace backup publication
The final publication review is also closed. The private same-directory archive is published with
an atomic hard-link create rather than rename-overwrite semantics. If any process creates the final
file or symlink after preflight but before publication, `ln` fails with `EEXIST`, the backup exits
nonzero, the concurrent destination remains byte-for-byte intact, and the trap removes only the
randomly named temporary archive owned by this invocation. The fake `pg_dump` safety test creates
that destination immediately before returning and pins the failure and cleanup behavior.
-929
View File
@@ -1,929 +0,0 @@
# Pre-deployment Fix Wave Report
Date: 2026-07-14
Worktree: `/home/chirone/ThothII/.worktrees/activity-log-cte-layout`
Base: `e5366d14a6da8fb331d94be60b8929cefb1fe3e0`
## Outcome
All three reviewed findings are implemented in one coherent backend/frontend wave:
1. Resume leaves the prior selection, Zustand state, document panel, and EventSource untouched
until `POST /resume` succeeds. Cold Resume changes state and reconnects only after backend
clear/rebind; already-active same-session Resume preserves the existing binding; failure is a
no-op apart from the fixed toast.
2. SSE uses monotonically increasing per-session ids, cursor-filtered replay, native and manual
reconnect cursors, id continuity across `hub.clear`, and descriptor-id pending-gate
idempotence at both backend and frontend layers.
3. Generic Pi system events and readiness errors are projected through explicit public
allowlists. Sentinel URLs, paths, tokens, stderr, commands, and extra fields do not reach HTTP
or SSE.
No harness, workflow, persistence, model, CTE viewer, CTE card, or shared Card file changed.
## Interfaces
- Frontend `resumeSession(id)` now returns
`Promise<{ id: string; alreadyActive: boolean }>` via `ResumeSessionResult`.
- Backend successful Resume always returns the same shape:
- running/waiting runtime: `{ id, alreadyActive: true }`
- validated cold runtime: `{ id, alreadyActive: false }`
- `SseHub.publish(sessionId, event, data): number` returns the assigned SSE id.
- `SseHub.subscribe(sessionId, send, { afterId, pending })` calls
`send(event, data, id)` for replay/live frames with `id > afterId`.
- `GET /sessions/:id/events` accepts native `Last-Event-ID` and manual
`?lastEventId=<integer>`; when both are valid it uses the greater cursor.
- Every emitted SSE frame is `id: <n>\nevent: <name>\ndata: <json>\n\n`.
- Public readiness failure is exactly:
`Session services are not ready. Check configuration and connectivity, then try again.`
- Generic Pi system events are exactly `{ type: "system_event", event }`, and `event` must be a
non-empty string.
## Files
Backend production:
- `backend/src/bridge/session-bridge.ts`
- `backend/src/routes/sessions.ts`
- `backend/src/sse/sse-hub.ts`
Backend tests:
- `backend/test/routes-sessions.test.ts`
- `backend/test/session-bridge.test.ts`
- `backend/test/sse-hub.test.ts`
- `backend/test/sse-route.test.ts` (new)
Frontend production/support:
- `frontend/src/api/sessions.ts`
- `frontend/src/api/types.ts`
- `frontend/src/shell/AppShell.tsx`
- `frontend/src/store/sessionStore.ts`
- `frontend/src/stream/useSessionStream.ts`
- `frontend/src/test/fakeEventSource.ts`
Frontend tests:
- `frontend/src/api/sessions.test.ts`
- `frontend/src/shell/AppShell.session-mgmt.test.tsx`
- `frontend/src/store/sessionStore.test.ts`
- `frontend/src/stream/useSessionStream.test.tsx`
## TDD RED/GREEN evidence
### 1. Backend Resume result and client-boundary allowlists
RED command:
```text
cd backend && npx vitest run test/routes-sessions.test.ts test/session-bridge.test.ts
```
RED output (exit 1):
```text
Test Files 2 failed (2)
Tests 6 failed | 35 passed (41)
expected { id: 's1' } to deeply equal { id: 's1', alreadyActive: false }
expected raw readiness URL/token/path to equal the fixed public message
expected three raw generic system events to equal [{ type: 'system_event', event: 'session_exit' }]
```
GREEN command:
```text
cd backend && npx vitest run test/routes-sessions.test.ts test/session-bridge.test.ts
```
GREEN output (exit 0):
```text
✓ test/session-bridge.test.ts (14 tests)
✓ test/routes-sessions.test.ts (27 tests)
Test Files 2 passed (2)
Tests 41 passed (41)
```
### 2. Backend exact-once SseHub and route framing
RED command:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/sse-route.test.ts
```
RED output (exit 1):
```text
Test Files 2 failed (2)
Tests 7 failed (7)
expected [undefined, undefined, undefined] to deeply equal [1, 2, 3]
expected unconditional replay not to contain "one" / "two"
expected one buffered pending gate, received replay plus a second pending emission
```
GREEN command:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/sse-route.test.ts
```
GREEN output (exit 0):
```text
✓ test/sse-hub.test.ts (4 tests)
✓ test/sse-route.test.ts (3 tests)
Test Files 2 passed (2)
Tests 7 passed (7)
```
### 3. Frontend cursor tracking and gate idempotence
RED command:
```text
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx src/store/sessionStore.test.ts
```
RED output (exit 1):
```text
Test Files 2 failed (2)
Tests 2 failed | 28 passed (30)
expected /sessions/s1/events to be /sessions/s1/events?lastEventId=7
expected duplicate gate pendingWidget to remain null, received gate-1
```
GREEN command:
```text
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx src/store/sessionStore.test.ts
```
GREEN output (exit 0):
```text
✓ src/store/sessionStore.test.ts (23 tests)
✓ src/stream/useSessionStream.test.tsx (7 tests)
Test Files 2 passed (2)
Tests 30 passed (30)
```
### 4. Frontend typed Resume and AppShell ordering/preservation
Typed API RED command:
```text
cd frontend && npx tsc -b
```
Typed API RED output (exit 1):
```text
src/api/sessions.test.ts(43,9): error TS2322: Type 'void' is not assignable to type
'{ id: string; alreadyActive: boolean; }'.
```
Lifecycle RED command:
```text
cd frontend && npx vitest run src/api/sessions.test.ts src/shell/AppShell.session-mgmt.test.tsx
```
Lifecycle RED output (exit 1):
```text
✓ src/api/sessions.test.ts (9 tests)
❯ src/shell/AppShell.session-mgmt.test.tsx (15 tests | 4 failed)
Test Files 1 failed | 1 passed (2)
Tests 4 failed | 20 passed (24)
already-active same-session Resume created two EventSources instead of one
deferred cold Resume closed the document panel before POST completion
failed same-session Resume closed the prior EventSource
failed Resume with no active session opened an EventSource
```
GREEN commands:
```text
cd frontend && npx vitest run src/api/sessions.test.ts src/shell/AppShell.session-mgmt.test.tsx
cd frontend && npx tsc -b
```
GREEN output (exit 0):
```text
✓ src/api/sessions.test.ts (9 tests)
✓ src/shell/AppShell.session-mgmt.test.tsx (15 tests)
Test Files 2 passed (2)
Tests 24 passed (24)
TypeScript: no output, exit 0
```
The AppShell cold-reconnect test additionally proves that the old source accepts an event while
Resume is pending, the replacement URL carries `lastEventId=8`, the replacement receives one
post-resume transcript/activity row, and two deliveries of the same descriptor id yield one gate.
## Affected verification
Backend command:
```text
cd backend && npx vitest run test/routes-sessions.test.ts test/session-bridge.test.ts \
test/sse-hub.test.ts test/sse-route.test.ts test/health.test.ts test/e2e-f1.test.ts
```
Output (exit 0):
```text
Test Files 6 passed (6)
Tests 52 passed (52)
```
Backend typecheck:
```text
cd backend && npx tsc --noEmit -p .
```
Output: no output, exit 0.
Frontend command:
```text
cd frontend && npx vitest run src/api/sessions.test.ts src/store/sessionStore.test.ts \
src/stream/useSessionStream.test.tsx src/shell/AppShell.session-mgmt.test.tsx \
src/shell/CentralStatus.test.tsx src/shell/ModelActivityPanel.test.tsx \
src/shell/f1-loop.test.tsx src/shell/AppShell.new-session.test.tsx
```
Output (exit 0):
```text
Test Files 8 passed (8)
Tests 74 passed (74)
```
Frontend typecheck:
```text
cd frontend && npx tsc -b
```
Output: no output, exit 0.
## Full verification
Backend full suite:
```text
cd backend && npx vitest run
```
```text
Test Files 22 passed (22)
Tests 177 passed (177)
```
Frontend full suite:
```text
cd frontend && npx vitest run
```
```text
Test Files 43 passed (43)
Tests 271 passed (271)
```
Backend production build:
```text
cd backend && npm run build
> tsc -p tsconfig.json
exit 0
```
Frontend production build:
```text
cd frontend && npm run build
> tsc -b && vite build
✓ 4835 modules transformed.
✓ built in 8.25s
exit 0
```
## Integrated re-review closure (2026-07-15)
This section supersedes the earlier cold same-session assertion that the replacement URL carries
`lastEventId=8`. That behavior was correct only while the backend process and its in-memory id
sequence survived. A restarted backend begins a fresh sequence, so a successful cold Resume now
explicitly discards the browser's cursor before replacing the EventSource.
All four integrated re-review findings are closed:
1. `AppShell` passes a dedicated cursor-reset epoch to `useSessionStream`. A cold same-session
Resume increments it only after `alreadyActive: false`; a high cursor such as `901` is omitted
from the replacement URL and fresh low-id events/gates are consumed. An already-active
same-session Resume still preserves its source, cursor, and store.
2. `useSessionStream` no longer mutates the cursor ref during render. Effect setup resets cursor
state on session/reset-epoch changes, callbacks are guarded by a captured active-source
identity, and cleanup clears only its own active identity. A queued event from the replaced
source cannot write the new store or poison its next reconnect URL.
3. Backend Resume is serialized per session and rechecks runtime state inside the lock. Manifest,
readiness, and reopen validation precede the transport commit. Idle/failed replacement creates
and binds the new runtime before `hub.clear`, which occurs synchronously immediately before the
first `Resuming session` publish. Reopen/create failure returns exactly
`Session could not be resumed. Check configuration and connectivity, then try again.`, keeps the
prior hub buffer/subscribers attached, and does not expose exception sentinels. Concurrent calls
perform one cold start and the waiter returns `alreadyActive: true`.
4. `SseHub.forget(id)` removes subscribers, buffered events, and the last id. Permanent session
DELETE invokes it after disk deletion; ordinary close and Resume continue to use `clear`, which
preserves the id sequence.
### Re-review files
Production:
- `backend/src/pi/pi-process-manager.ts`
- `backend/src/routes/sessions.ts`
- `backend/src/sse/sse-hub.ts`
- `frontend/src/shell/AppShell.tsx`
- `frontend/src/stream/useSessionStream.ts`
Tests/support:
- `backend/test/pi-process-manager.test.ts`
- `backend/test/routes-sessions.test.ts`
- `backend/test/sse-hub.test.ts`
- `frontend/src/shell/AppShell.session-mgmt.test.tsx`
- `frontend/src/stream/useSessionStream.test.tsx`
- `frontend/src/test/fakeEventSource.ts`
### Re-review TDD RED/GREEN evidence
Frontend RED command:
```text
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx \
src/shell/AppShell.session-mgmt.test.tsx
```
RED output (exit 1):
```text
Test Files 2 failed (2)
Tests 3 failed | 21 passed (24)
reset epoch: expected the old source to close, received false
cold same-session: expected /sessions/s1/events, received ?lastEventId=901
stale source: expected an empty transcript, received "stale session one"
```
Frontend GREEN command:
```text
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx \
src/shell/AppShell.session-mgmt.test.tsx
cd frontend && npx tsc -b
```
GREEN output (exit 0):
```text
Test Files 2 passed (2)
Tests 24 passed (24)
TypeScript: no output, exit 0
```
Backend RED command:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/routes-sessions.test.ts
```
RED output (exit 1):
```text
Test Files 2 failed (2)
Tests 8 failed | 28 passed (36)
three Resume ordering assertions observed clear before reopen/create
reopen and create sentinels escaped as raw HTTP 500 responses
the concurrent waiter cold-started again instead of returning alreadyActive: true
SseHub.forget was absent and DELETE did not invoke permanent cleanup
```
Backend GREEN command:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/routes-sessions.test.ts
cd backend && npx tsc --noEmit -p .
```
GREEN output (exit 0):
```text
Test Files 2 passed (2)
Tests 36 passed (36)
TypeScript: no output, exit 0
```
The failure tests publish a post-failure probe through the same hub and prove that a subscriber
attached before either reopen or create rejection still receives it. The concurrency test overlaps
two same-id requests behind a deferred reopen and proves one manifest/readiness/reopen/create/clear
sequence.
### Initial re-review verification (before independent-review hardening)
```text
cd backend && npx vitest run
Test Files 22 passed (22)
Tests 182 passed (182)
cd frontend && npx vitest run
Test Files 43 passed (43)
Tests 273 passed (273)
cd backend && npm run build
> tsc -p tsconfig.json
exit 0
cd frontend && npm run build
> tsc -b && vite build
✓ 4835 modules transformed.
✓ built in 8.46s
exit 0
```
`git diff --check` produced no output (exit 0). The frontend build retains its pre-existing
large-chunk warning; no new build or type errors were introduced.
Final whitespace verification:
```text
git diff --check
no output, exit 0
```
## Self-review
- Resume sequencing: reopen and runtime binding precede backend `clear` and HTTP success; frontend
state mutation and cursor-reset epoch follow it. Failure catch only emits fixed UI copy.
- Already active: same-session returns before reset/generation/manifest repaint; different session
resets the single-session store and binds the new id only after success.
- SSE exact-once: ids are transport identity, not content hashes; replay is strictly `id > cursor`;
`clear` retains the counter; pending gate matching uses only descriptor id.
- Cursor behavior: hook tracks `MessageEvent.lastEventId`, carries it only to an ordinary same-id
generation, and resets it on session-id/cold-runtime epoch change. Native EventSource reconnect
remains supported by the route header.
- Gate defense: the Zustand set survives pending clear but resets with the session store.
- Client boundary: raw `ensure.error` is unused in public responses; generic Pi system events are
reconstructed rather than spread; frontend type mirrors the two-field event.
- Scope: `git diff` contains no CTE/Card/harness/workflow/persistence/model changes. Four pre-existing
modified `.superpowers/sdd/{progress,task-2-report,task-3-report,task-4-report}.md` files are user
work and are excluded from staging.
## Remaining concerns
- The 200-event SSE ring limit remains intentional. A brand-new page can reconstruct only retained
backlog; an in-memory same-session reconnect is exact-once from its cursor.
- Per-session sequence counters remain in backend memory after `clear` by design so later in-process
cold same-id Resume cannot reuse ids. Permanent DELETE removes the counter via `forget`.
- The Delete-then-Resume adversarial route test proves the deleted session is not resurrected but
currently receives the runner's generic HTTP 500 when `sessionShow` can no longer find it. A
future API cleanup can normalize that missing-session response to 404 or 409.
- Frontend tests still print pre-existing MSW unhandled-request and React ref/`act` warnings even
though all 276 tests pass. The frontend production build still reports pre-existing large chunk
warnings. Neither warning class was introduced or expanded by this change.
- No live Pi/DWH smoke was run; this wave changes only REST/SSE/frontend lifecycle boundaries and
is covered by fake-Pi, live Fastify SSE, component, full-suite, typecheck, and production-build
gates.
## Independent-review hardening
The required independent review was run repeatedly against the uncommitted diff. Its first pass
found four Important lifecycle edges beyond the integrated findings: queued old-runtime callbacks,
post-spawn construction cleanup, concurrent frontend Resume completions, and the passive-effect
commit window. Its second pass confirmed those fixes and identified one remaining Important
retention issue in the new runtime-identity map. The final pass reported no Critical, Important, or
Minor findings and assessed the diff ready to merge.
The resulting hardening is:
- Runtime bridge callbacks are gated by the bound runtime identity. Replacement, close, and DELETE
invalidate the old identity, so queued old events cannot publish or call `failSession`. An active
runtime removed by the manager can still publish its complete public failure sequence; after the
terminal unmanaged `agent_end`, its binding is released and later events are rejected.
- `PiProcessManager` kills the spawned child and removes any registered map entry if either
spawn-boundary stderr setup or later RPC/bridge/map initialization throws.
- Resume completion compares against synchronously maintained current active-session identity.
Concurrent `alreadyActive: false` then `alreadyActive: true` results preserve the cold source,
cursor, store, and replayed gate.
- Stream source replacement uses a layout effect. A deterministic later-layout-effect test delivers
a queued old event inside the former commit-to-passive-cleanup window and proves it is ignored.
- Cursor tests cover both a restarted backend's fresh low ids and an in-process hub's preserved high
ids followed by a cursor-bearing ordinary reconnect.
### Hardening TDD RED/GREEN evidence
Backend identity/construction RED command:
```text
cd backend && npx vitest run test/pi-process-manager.test.ts test/routes-sessions.test.ts
```
```text
Test Files 2 failed (2)
Tests 3 failed | 69 passed (72)
post-spawn reader initialization did not kill the child
replaced and deleted runtime callbacks still called failSession/published
```
Additional spawn-boundary and terminal-release RED checks:
```text
cd backend && npx vitest run test/pi-process-manager.test.ts \
-t "spawn boundary initialization"
Tests 1 failed | 38 skipped (39)
cd backend && npx vitest run test/routes-sessions.test.ts -t "terminal sequence"
Tests 1 failed | 34 skipped (35)
```
Frontend concurrency/layout RED command:
```text
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx \
src/shell/AppShell.session-mgmt.test.tsx
```
```text
Test Files 2 failed (2)
Tests 2 failed | 25 passed (27)
the later-layout-effect event wrote "commit-window stale text"
the false→true completion pair erased pending gate "cold-gate"
```
Final focused GREEN commands:
```text
cd backend && npx vitest run test/pi-process-manager.test.ts \
test/routes-sessions.test.ts test/sse-hub.test.ts
cd backend && npx tsc --noEmit -p .
Test Files 3 passed (3)
Tests 79 passed (79)
TypeScript: no output, exit 0
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx \
src/shell/AppShell.session-mgmt.test.tsx
cd frontend && npx tsc -b
Test Files 2 passed (2)
Tests 27 passed (27)
TypeScript: no output, exit 0
```
### Final full verification after review hardening
```text
cd backend && npx vitest run
Test Files 22 passed (22)
Tests 188 passed (188)
cd frontend && npx vitest run
Test Files 43 passed (43)
Tests 276 passed (276)
cd backend && npm run build
> tsc -p tsconfig.json
exit 0
cd frontend && npm run build
> tsc -b && vite build
✓ 4835 modules transformed.
✓ built in 8.47s
exit 0
```
The final frontend run retains the repository's pre-existing MSW/ref/`act` warnings, and the build
retains the pre-existing large-chunk warning. No test, typecheck, or build failures remain.
## Stale-bootstrap, lifecycle-lock, and competing-Resume hardening
Date: 2026-07-15
Base: `08b1f4909e8eb7538156cecc2e7a6cafb46ddfc7`
This follow-up closes asynchronous identity/order and multi-client transport gaps found in the
pre-deployment review:
- `PiProcessManager.teardownIfCurrent(id, runtime)` makes teardown an identity-checked operation.
Bootstrap re-checks identity after configuration/retrieval and before both the public
`Starting model` event and model start. Its failure continuation acquires the same session
lifecycle lock, claims only its own runtime identity, and holds serialization through persisted
failure and the public terminal sequence. A continuation left behind by Close or DELETE cannot
target a replacement or recreate forgotten SSE state.
- The former Resume-only promise tail is now a per-session lifecycle lock shared by Resume, Close,
and DELETE. Each route reads the current runtime inside the lock immediately before replacement
or removal and uses identity-checked teardown. Deferred route tests prove both orderings:
Resume then Close/Delete finishes removed with no post-removal bootstrap event; Close then Resume
creates only after Close completes; DELETE then Resume cannot recreate a deleted session.
- AppShell assigns each Resume invocation a monotonic token and records the latest target. A
completion for a different, superseding session id cannot reset the store, select a source, close
the panel, or repaint phase from a late manifest. Same-id invocations are per-target single-flight
operations through the POST and local binding commit: repeated pre-commit clicks update the
shared operation's latest token but issue no second POST or commit path. The operation becomes
joinable again before its manifest fetch, whose repaint remains token/id/selection guarded. Start
new, Stop, streamed session exit, and active-session deletion invalidate pending Resume work.
This prevents stale-source preservation and reverse/non-Resume intent overwrite without allowing
a slow manifest to suppress a later explicit rebind.
- `SseHub` subscriber registrations now carry idempotent transport-close callbacks. `clear` and
`forget` snapshot and actively close every response before discarding runtime transport state;
callback-driven unsubscription during that iteration is safe. The SSE route ends its response so
native EventSource reconnects with `Last-Event-ID`. Post-clear events retain monotonic ids and are
buffered for replay; `forget` additionally resets the id state.
Production files:
- `backend/src/pi/pi-process-manager.ts`
- `backend/src/routes/sessions.ts`
- `backend/src/sse/sse-hub.ts`
- `frontend/src/shell/AppShell.tsx`
Regression tests:
- `backend/test/pi-process-manager.test.ts`
- `backend/test/routes-sessions.test.ts`
- `backend/test/sse-hub.test.ts`
- `backend/test/sse-route.test.ts`
- `frontend/src/shell/AppShell.session-mgmt.test.tsx`
### TDD RED/GREEN evidence
Runtime identity API RED:
```text
cd backend && npx vitest run test/pi-process-manager.test.ts -t "identity-checked teardown"
Test Files 1 failed (1)
Tests 1 failed | 39 skipped (40)
TypeError: mgr.teardownIfCurrent is not a function
```
Runtime identity API GREEN:
```text
Test Files 1 passed (1)
Tests 1 passed | 39 skipped (40)
```
Deferred bootstrap RED:
```text
cd backend && npx vitest run test/routes-sessions.test.ts \
-t "stale bootstrap|bootstrap that"
Test Files 1 failed (1)
Tests 6 failed | 35 skipped (41)
close/delete + replacement: stale continuation removed the replacement runtime
delete without replacement: stale continuation called failSession after forget
```
Deferred bootstrap GREEN:
```text
Test Files 1 passed (1)
Tests 6 passed | 35 skipped (41)
```
Shared lifecycle ordering RED:
```text
cd backend && npx vitest run test/routes-sessions.test.ts \
-t "Resume followed|Close followed|Delete followed"
Test Files 1 failed (1)
Tests 4 failed | 41 skipped (45)
All four deferred assertions observed the competing route settle before the first lifecycle
operation released.
```
Bootstrap plus lifecycle GREEN:
```text
cd backend && npx vitest run test/routes-sessions.test.ts \
-t "Resume followed|Close followed|Delete followed|stale bootstrap|bootstrap that"
Test Files 1 passed (1)
Tests 10 passed | 35 skipped (45)
```
Competing frontend Resume RED:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "competing Resume|stale Resume manifest"
Test Files 1 failed (1)
Tests 2 failed | 16 skipped (18)
reverse POST completion opened a second, stale EventSource
late s1 manifest repainted the selected s3 phase from F3 to F7
```
Competing and same-id Resume GREEN:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "competing Resume|stale Resume manifest|false then true"
Test Files 1 passed (1)
Tests 3 passed | 15 skipped (18)
```
### Independent-review hardening RED/GREEN
The first final review reported no Critical findings and three Important edge cases: bootstrap
could start during an in-progress Close; bootstrap-owned failure was persisted twice; and an older
same-id result could overwrite newer state. The integrated reviewer also required non-Resume
navigation to invalidate pending Resume work. The final main review tightened the same-ID contract
to true single-flight so a second same-target click cannot preserve a dead pre-restart source.
Backend review RED:
```text
cd backend && npx vitest run test/routes-sessions.test.ts \
-t "Close suppresses|bootstrap failure persists once"
Test Files 1 failed (1)
Tests 2 failed | 45 skipped (47)
deferred configure started Pi while closeSession was still pending
bootstrap/public failure called failSession twice
```
Backend review GREEN:
```text
Test Files 1 passed (1)
Tests 2 passed | 45 skipped (47)
```
Same-id single-flight RED:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "share one cold request"
Test Files 1 failed (1)
Tests 1 failed | 18 skipped (19)
two concurrent same-ID invocations issued two cold POSTs (three total including initial activation)
```
Non-Resume invalidation RED:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "starting a new question invalidates"
Test Files 1 failed (1)
Tests 1 failed | 19 skipped (20)
the late Resume opened an EventSource after Start new returned to the landing state
```
Frontend review GREEN:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "share one cold request|competing Resume|stale Resume manifest|starting a new question invalidates"
Test Files 1 passed (1)
Tests 4 passed | 15 skipped (19)
```
Post-commit single-flight lifetime RED:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "releases same-id single-flight"
Test Files 1 failed (1)
Tests 1 failed | 19 skipped (20)
s1 committed and waited on its manifest; after s3 superseded it, a new s1 Resume reused the old
operation and issued no second s1 POST (expected 2, received 1).
```
Same-id and manifest lifetime GREEN:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx -t "same-id|manifest"
Test Files 1 passed (1)
Tests 4 passed | 16 skipped (20)
cd frontend && npx tsc -b
no output, exit 0
```
Multi-client SSE disconnect RED:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/sse-route.test.ts
Test Files 2 failed (2)
Tests 3 failed | 6 passed (9)
clear/forget invoked zero of two registered close callbacks, and two live HTTP SSE responses timed
out instead of reaching EOF after clear.
```
Multi-client SSE disconnect GREEN:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/sse-route.test.ts
Test Files 2 passed (2)
Tests 9 passed (9)
cd backend && npx tsc --noEmit -p .
no output, exit 0
```
The Hub tests use two subscribers whose close callbacks immediately unsubscribe themselves, proving
safe snapshot iteration and exactly-once closure. The live-route test opens two HTTP streams, proves
both receive EOF on clear, publishes a new event and gate, then reconnects after id 1 and replays
exactly ids 2 and 3. The forget test closes both subscribers and proves the next id resets to 1.
Close now removes the observed runtime identity before awaiting persistence. Failure persistence is
claimed once per runtime and lifecycle-serialized; bootstrap's public `session_failed` cannot start
a duplicate. A per-target in-flight map owns the only same-ID POST and commit while its mutable
latest token keeps s1→s2→s1 ordering correct; it is removed immediately after the binding commit,
before awaiting the independently guarded manifest. One shared invalidation helper is called when
active deletion, streamed exit, Start new, or Stop begins.
### Focused verification
```text
cd backend && npx vitest run test/routes-sessions.test.ts test/pi-process-manager.test.ts \
test/sse-hub.test.ts test/sse-route.test.ts
Test Files 4 passed (4)
Tests 96 passed (96)
cd backend && npx tsc --noEmit -p .
no output, exit 0
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
src/shell/AppShell.new-session.test.tsx src/stream/useSessionStream.test.tsx
Test Files 3 passed (3)
Tests 36 passed (36)
cd frontend && npx tsc -b
no output, exit 0
```
### Full verification
```text
cd backend && npx vitest run
Test Files 22 passed (22)
Tests 202 passed (202)
cd frontend && npx vitest run
Test Files 43 passed (43)
Tests 280 passed (280)
cd backend && npm run build
> tsc -p tsconfig.json
exit 0
cd frontend && npm run build
> tsc -b && vite build
✓ 4835 modules transformed.
✓ built in 8.47s
exit 0
```
The frontend suite/build retain the previously documented MSW, React ref/`act`, experimental type
stripping, and large-chunk warnings. No warning class was introduced by this wave. No harness,
workflow, persistence, SQL/CTE viewer, model-selection, or deployment file changed. The four
pre-existing modified `.superpowers/sdd/{progress,task-2-report,task-3-report,task-4-report}.md`
files remain excluded from staging.
### Final independent-review verdict
After the multi-client transport fix, the independent reviewer reported no Critical, Important, or
Minor findings. Its own focused verification passed 96 backend transport/lifecycle tests, 31
frontend Resume/stream tests, both TypeScript checks, and `git diff --check`. Final assessment:
**Ready to deploy: Yes.**
-53
View File
@@ -1,53 +0,0 @@
# DWH REST per-installation authentication SDD progress
Plan: `docs/superpowers/plans/2026-08-20-dwh-rest-per-installation-auth.md`
Branch: `feat/dwh-rest-installation-auth`
Worktree: `/home/chirone/ThothII-next/.worktrees/dwh-rest-installation-auth`
Baseline: workspace docs PASS; Go unavailable on host (use containerized Go 1.26.5); default Compose pre-existing path-sensitive false positive under `/home/chirone`.
Task 1: complete (commits 3bc84b0..1e82fd3, review clean after bounded-digest fix wave).
Task 1 plan note: the dotted-import grep was resolved by the exact module assertion in `09290a0`; `go list -m all` confirms the DWH module only.
Task 2: complete (commits 541ef45, 971a0e6, e90a1a1; independent vet fix d2415b5; review clean after bounded-read, expiry/JSON, same-Store, and cross-Store synchronization fix waves).
Task 3: complete (commits ebb360f, 055dcab, b1079bd; independent review clean after CLI grammar, metadata validation, and ambiguous-publication cleanup fix waves).
Task 4: complete (commits 943f809, 419c344, 134dc19; independent review PASS after legacy multiplicity, fail-closed handler/socket, SGID 2750, O_RDONLY shared lock, OpenReadOnly, and safe socket-parent waves). Task 5 must precreate `.writer.lock` as `0640 root:dwh-auth`; add a non-owner group/cross-process integration proof when packaging permits.
Task 5: complete (commits 87606c7, 62ec29f; independent review PASS after auth-subrequest header isolation and target-Nginx duplicate-header verification). No runtime installation or service/Nginx mutation performed.
Task 6: complete (commits 1b18a0f, d0f7e04; independent review PASS after regex/duplicate bypass, exact-PID TCP, and bounded-cleanup hardening). Minor for final review: remove or rename the redundant legacy `negative_postgrest_bypass` fixture and align the historical fixture-count prose if useful. No active Nginx/runtime mutation performed.
Task 7: complete (commits 7b9b8b3, 707c13d, f616aab, 7b86ea9, 7fe5316; Terra review PASS). Documentation and rollout-contract alignment completed; no runtime mutation performed.
Task 8: complete at frozen SHA `6499d24892b4383ac492579303e766cfb51fe44e` (fix commits `09290a0`, `6499d24`; Terra review PASS). Focused, portability, scanner, DWH Go, and two full tools/tht matrix runs PASS; evidence recorded at `.artifacts/dwh-auth/source-verification.md`. Broad coupling remains `BASELINE_RED` debt; immutable paths remain unchanged.
Task 9: PASS at frozen SHA `0c4ff3750d3ecd3fc514e50e511cf7475fbe0446`. Built and installed the exact local candidate, enabled and started `dwh-auth`, imported the protected legacy credential under public ID `legacy-shared`, and created `psd-mac-primary` with public key ID `oNPdOfoH7ypLtVb1`. Registry check, AF_UNIX-only listener, v1/legacy `204`, random/missing `401`, bounded journal scan, installed-file hashes, and `nginx -t` all PASS. Protected report: `/root/dwh-auth-provision/gate9-20260821T054514Z.report.md`, SHA-256 `65ce1e0d8be5f74355eca2b1dca901da16f2864f68eafb7dede9d23ef36b82d5`. Nginx was not changed or reloaded; the legacy key remains active; the old stack was not changed or stopped. Terra final review: PASS with no Critical or Important findings.
Task 10: NOT STARTED and requires a second explicit authorization. Public Nginx cutover, Mac-key delivery/configuration, and legacy revocation have not occurred. Activity 1 remains `IN_DISCUSSION`; external deployment remains `SURVEY_NO_GO`.
Final clarification (bookkeeping): initial authorization at `6499d24` stopped before installation because the protected legacy file was missing and a journal-scan finding remained. A secret-safe legacy file was prepared without emit/hash; Nginx metadata remained unchanged and `nginx -t` PASS. Fix commits `6fb4886`, `dee0f9c`, `0c4ff37` received Terra PASS, followed by a detached complete re-freeze PASS at full `0c4ff3750d3ecd3fc514e50e511cf7475fbe0446`. The owner then explicitly authorized Gate 9 at that exact SHA; Gate 9 completed as recorded above. Task 10 remains a separate gate.
## Project A authentication runtime projection
Plan: `docs/superpowers/plans/2026-08-21-project-a-server-auth-runtime-projection.md`
Plan commit: `64f46c7019e11a74dae35a7cdb447cd881e17061`
Implementation baseline: `64f46c7019e11a74dae35a7cdb447cd881e17061`
Runtime constraint: source, synthetic tests, and documentation only; Project A, `/srv`, Nginx,
the legacy stack, and shared services remain untouched.
Task 1: complete (commits `8a8f2c2`, `f9e2950`, `de86760`; independent Terra review PASS after descriptor-relative rewrite, full-history validation, crash recovery, destructive replacement guards, deterministic failure seams, and interrupted-retention recovery).
Task 2: complete (commits `05f8615`, `da2f4a6`, `1e2c4e6`; independent review PASS after exact GID enforcement, bounded descriptor-bound namespace enumeration, strict trailing-slash parity, OIDC coverage, and complete one-retry `CURRENT` publication linearization).
Task 3: complete (commit `903c0b4`; independent Terra review PASS after retained-FD outer locking, cancellable runtime/canonical waits, public transaction-context propagation, deterministic swap/metadata/creator-race tests, and fail-closed pre/post-commit error handling).
Task 4: complete (commit `3d9a9f0`; independent Terra review PASS after moving the Linux/root restore gate before secret-bearing checkpoint creation). Exact focused Go, race, vet, Node 24 focused/full, TypeScript, projected Compose, secret-policy, Windows backup compile, and full serialized tools/tht gates PASS. Historical canonical/unified Compose failures were reproduced as baseline-only documentation/path coupling failures and were not weakened.
Task 5: complete (commit `ef7ae70`; independent Terra review PASS after read-only Vitest gate repair,
remote-Docker/context hardening, and adversarial canonical-mount/`sudo printenv` verifier fixes).
All 14 cross-layer acceptance cases, documentation verifiers, shell syntax, full Go/race/vet,
backend Node 24 Vitest/typecheck, projected Compose, and secret-policy gates PASS. The three known
default/canonical/unified Compose policy failures were reproduced at `a21e2c1` and remain
unmodified baseline debt. Project A has not been started; applying the descriptor or any runtime
root under `/srv/thothii` still requires a new explicit authorization.
Final Project A source review: PASS for `a21e2c1..ef7ae70`; independent Terra review found no
remaining Critical or Important issue after validating restore admission/order, backend path and
identity controls, transaction cancellation, lifecycle gating, portability, dependency scope, and
redaction. Pre-live stop boundary remains in force.
-325
View File
@@ -1,325 +0,0 @@
# Task 2 report — protected atomic registry
## Scope and commit
- Commit: `541ef45 feat: add protected DWH credential registry`
- Committed files only:
- `tools/dwh-auth/internal/securefile/securefile_linux.go`
- `tools/dwh-auth/internal/securefile/securefile_linux_test.go`
- `tools/dwh-auth/internal/registry/store.go`
- `tools/dwh-auth/internal/registry/store_test.go`
- No server, Nginx, systemd, Docker stack, real registry, secrets, or legacy ThothII files were
read or changed. Tests use `t.TempDir` and synthetic record digests only.
## TDD evidence
All Go commands ran in the required official `golang:1.26.5` container with only this linked
worktree bind-mounted at `/work`. The container image reports `go version go1.26.5 linux/amd64`.
### RED
Before either Task 2 production file existed, the focused command was run inside the container:
```text
go test ./internal/securefile ./internal/registry -count=1
```
It failed non-zero for the expected absent implementation symbols, including `undefined: OpenDir`,
`undefined: ReadSecret`, `undefined: Open`, `undefined: State`, `undefined: PublicRecord`, and
`undefined: Store`.
### GREEN
After the minimal implementation and formatting:
```text
go test ./internal/securefile ./internal/registry -count=1
```
Result:
```text
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry
```
### Race verification
The required race command completed successfully:
```text
go test -race ./internal/securefile ./internal/registry -count=1
```
Result:
```text
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 1.027s
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 1.179s
```
Additional scoped verification:
```text
go vet ./internal/securefile ./internal/registry
go test ./... -count=1
git diff --cached --check
```
The Task 2 vet command completed with no findings; all four DWH-auth packages passed the full
module test run; the staged-diff check completed with no output.
## Delivered behavior
- `securefile` is Linux-only and traverses absolute paths through descriptor-anchored
`syscall.Open`/`Openat` calls with `O_NOFOLLOW|O_CLOEXEC`; protected roots, child directories,
records, and secret files are regular/directories only and are checked against `Lstat` after
`Fstat`.
- Protected reads reject special, group-writable, or world-writable modes, cap record reads at
4096 bytes, read at most one extra byte, and reject file-size changes or short/partial reads.
Secret ingress additionally requires exact `0600`.
- Secret output uses `O_CREAT|O_EXCL|O_NOFOLLOW`, exact `0600`, and an absolute protected parent.
- `registry.Open` creates protected `active` and `revoked` subdirectories under an existing safe
root. Record enumeration rejects unexpected entries, unsafe files, symlinks, oversized files,
bad filenames, malformed JSON, unknown JSON fields, duplicate JSON fields, and trailing JSON.
- `Add` validates Task 1 records, writes canonical JSON plus one newline through an exclusive
temporary file, sets final mode `0640`, syncs the file, renames under a protected per-root writer
lock, then syncs the directory.
- `Revoke` writes and syncs a valid revoked record before unlinking and syncing the active record.
`Find` checks revoked first; `List` resolves an active/revoked overlap to the revoked public
record. `FindLegacy` scans fail-closed and permits only the reserved legacy record state.
- `PublicRecord` deliberately omits `secret_sha256`; the redaction is regression-tested.
## Security-test coverage
- protected normal files and canonical record publication;
- symlinked roots, registry directories, records, and secret input;
- unsafe root/directory/record/secret modes;
- bounded/oversized record input;
- unknown, duplicate, trailing, and partial JSON;
- filename mismatch and multiple legacy-record integrity failures;
- revoked-state precedence when both active and revoked files exist;
- concurrent adds and concurrent reads during revocation, including the race detector.
## Self-review
Reviewed all syscall, path, mode, and error paths after the final race run:
- Directory traversal never follows a supplied component; later operations use retained directory
descriptors, not re-opened untrusted prefixes.
- `Fstat` validates the opened object and `Lstat` must identify the same inode/device; the direct
child name grammar refuses separators, dot components, and NUL.
- File validation occurs before and after reads; mode/type/size checks fail closed. Directory
listing obtains a fresh `openat(dirfd, ".")` descriptor so scans do not share a mutable directory
offset.
- Writer serialization protects the check-then-rename no-replace sequence. Failed temporary
cleanup leaves an unexpected entry that later scans reject rather than silently accepting it.
- State-specific validation rejects revocation metadata in active records and requires it in
revoked records. Revoked files are consulted before active files so interruption after revoked
publication cannot reactivate a credential.
- All functionality uses only Go standard-library packages and Linux `syscall`; no CGO, SQLite,
or third-party module was added.
## Concerns
- The optional whole-module `go vet ./...` reports a pre-existing Task 1 test warning at
`internal/credential/credential_test.go:86` (`append` with no variadic values). The identical
line is present in approved HEAD `1e82fd3`, outside this task’s authorized files. Focused Task 2
vet passes, and all module tests pass.
- The official image's login shell resets `PATH` and hides `/usr/local/go/bin`; all evidence uses
direct `go`/`gofmt` container entrypoints, which preserves the image’s Go 1.26.5 environment.
- The generic `apply_patch` helper intermittently failed before file access with a sandbox network
namespace error. Exact scoped corrections were applied through the shared worktree workflow;
this did not affect the final staged file set or verification evidence.
## Review remediation — 2026-08-21
### Scope and fix commit
- Review-fix commit: `971a0e6 fix: harden DWH credential registry reads`.
- Committed files only:
- `tools/dwh-auth/internal/registry/store.go`
- `tools/dwh-auth/internal/registry/store_test.go`
- The separate Task 1 vet correction is the independent preceding commit `d2415b5`; it is not
included in this Task 2 fix commit. No filesystem primitive, server, Nginx, service, registry,
secret, Docker stack, or legacy ThothII file was changed.
### Strict TDD evidence
All commands again used the official `golang:1.26.5` image with only this linked worktree mounted
at `/work`.
#### RED
The first focused command was run after the new regression tests and before production changes:
```text
go test ./internal/securefile ./internal/registry -count=1
```
It failed as intended. The three case-variant aliases (`SECRET_SHA256`, `Secret_SHA256`, and
`Schema_Version`) were accepted; past expiry returned active records from both `Find` and
`FindLegacy`; a revocation snapshot let readers return active data before publication; a temporary
file let `List`, `Check`, and `FindLegacy` observe false integrity failures; and the original
concurrent-read regression observed `ErrNotFound` during revocation.
The deterministic exact-expiry test was then added before the clock implementation. Its focused
run failed as intended with:
```text
internal/registry/store_test.go:572:10: store.now undefined
```
#### GREEN and verification
After the minimum implementation and `gofmt`:
```text
go test ./internal/securefile ./internal/registry -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 0.014s
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 0.684s
go test -race ./internal/securefile ./internal/registry -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 1.022s
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 1.711s
go vet ./internal/securefile ./internal/registry
```
The focused vet output was empty (success). Additional final checks passed:
```text
go test ./... -count=1
ok internal/credential
ok internal/record
ok internal/registry
ok internal/securefile
go vet ./...
go test -race ./internal/registry -count=10
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 8.127s
git diff --cached --check
```
### Remediated security invariants
- `Find` and `FindLegacy` now deny an active record when `ExpiresAt <= now.UTC()`, returning the
existing non-disclosing `ErrNotFound`. The unexported per-Store `now` function is the minimal
deterministic clock seam; past, exact-equality, and future cases are covered for v1 and legacy
records. A revoked record is still consulted before expiry and therefore remains authoritative.
- One per-Store `sync.RWMutex` creates an in-process consistent snapshot. `Add` and `Revoke` hold
it exclusively for their full writer-lock lifetime, including temporary-file publication and
revoked-then-active removal. `Find`, `FindLegacy`, `List`, and `Check` hold a shared lock; their
bodies delegate only to unlocked helpers, preventing nested-lock deadlocks. `Close` also takes
the exclusive lock before closing descriptors.
- Deterministic regression tests hold the writer path at the revocation publication/unlink and
temporary-file stages. They prove public readers wait, then see either the final revoked state or
a clean directory, eliminating the `Names`-to-load/unlink and temporary-entry false failures
within the Store contract.
- Before struct decoding, the outer record JSON object now requires exactly spelled keys from the
schema allowlist and rejects duplicate literal keys. `Decoder.DisallowUnknownFields`, recursive
duplicate detection, trailing-value rejection, record validation, filename matching, no-follow
reads, modes, and durability ordering remain intact.
### Self-review and concerns
- Reviewed the new lock boundaries, error returns, revoked-first ordering, clock fallback,
JSON-token consumption, and every unchanged `securefile` syscall/path/mode boundary. The change
adds only standard-library `sync`; it does not relax existing fail-closed behavior.
- The synchronized snapshot is intentionally per `Store`, matching the requested in-process
contract. The existing protected advisory lock continues to serialize writers across Store
instances/processes; no cross-process reader snapshot is claimed by this fix.
- The historical whole-module vet concern in the original Task 2 report is now resolved by the
independent Task 1 commit `d2415b5`; complete module vet passes in the final evidence above.
## Cross-Store snapshot remediation — 2026-08-21
### Scope and TDD evidence
This third Task 2 fix wave changes only the protected lock primitive and registry snapshot code:
- `tools/dwh-auth/internal/securefile/securefile_linux.go`
- `tools/dwh-auth/internal/securefile/securefile_linux_test.go`
- `tools/dwh-auth/internal/registry/store.go`
- `tools/dwh-auth/internal/registry/store_test.go`
All commands used the official `golang:1.26.5` image with only this linked worktree mounted at
`/work`.
The test-only red patch initially tried to inspect the unexported `securefile.Dir.fd` through the
registry package and therefore did not compile. That assertion was removed without production
changes: the registry tests still create writer Store A and reader Store B through two independent
`Open(root)` calls, while the securefile test proves separate descriptors directly in its own
package. The subsequent behavioral RED run, before the production change, was:
```text
go test ./internal/securefile ./internal/registry -count=1
FAIL TestLockSharedAllowsReadersAndBlocksExclusiveWriter: Dir lacks shared advisory locking
FAIL TestCrossStoreReadersWaitAcrossRevokePublicationAndUnlink:
Find, List, Check, and FindLegacy completed during Store A's revocation snapshot
FAIL TestCrossStoreScanReadersWaitForWriterTemporaryFile:
Store B's List, Check, and FindLegacy observed `.tmp-regression`
```
### Delivered synchronization contract
- `securefile.Dir.LockShared` now acquires `LOCK_SH` on the same protected, no-follow, exact-0600
root lock file used by `Lock`, which continues to acquire `LOCK_EX`. The lock file is still
opened/created, mode-validated, inode-checked, and closed through the existing Linux syscall
path.
- Every public snapshot reader (`Find`, `FindLegacy`, `List`, and `Check`) takes its Store
`RLock`, then a shared advisory lock on root `.writer.lock`, and retains both through the whole
revoked/active lookup or directory scan/load. `Add` and `Revoke` retain Store `Lock`, then the
same root lock under `LOCK_EX`, over their full operation.
- The lock order is universally Store mutex then root advisory lock. Public methods delegate only
to unlocked helpers, so neither reader nor writer paths recursively acquire the Store mutex.
`Close` retains its exclusive Store mutex, preventing descriptor closure from racing any locked
reader or writer.
- Revoked-first precedence, expiry denial, exact JSON validation, no-follow checks, record modes,
temporary-file durability, and all previous behavior remain unchanged.
### GREEN and repeated verification
```text
go test ./internal/securefile ./internal/registry -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 0.019s
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 0.711s
go test -race ./internal/securefile ./internal/registry -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 1.032s
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 1.748s
go vet ./internal/securefile ./internal/registry
go test ./... -count=1
ok internal/credential
ok internal/record
ok internal/registry
ok internal/securefile
go vet ./...
go test -race ./internal/registry \
-run 'TestCrossStoreReadersWaitAcrossRevokePublicationAndUnlink|TestCrossStoreScanReadersWaitForWriterTemporaryFile' -count=20
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 9.880s
go test -race ./internal/securefile \
-run TestLockSharedAllowsReadersAndBlocksExclusiveWriter -count=20
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 1.063s
git diff --check
```
Both vet commands and the whitespace check produced no output. The securefile regression opens
three protected directory descriptors, proves they are distinct, permits two independent shared
holders, proves a third descriptor cannot take `LOCK_EX|LOCK_NB`, then proves exclusive acquisition
succeeds after shared release. The registry regressions deterministically block Store B readers
while Store A holds the exclusive root lock and verify only final revoked/clean states afterward.
### Self-review and concerns
- Reviewed lock creation/reopen races, no-follow flags, exact lock-file mode validation, lock
release, descriptor lifetime, lock ordering, error wrapping, and the unlocked-helper call graph.
No public reader invokes another public reader or writer while holding a Store lock.
- Advisory synchronization necessarily covers cooperating registry Store instances/processes;
arbitrary external filesystem mutation remains fail-closed through the existing integrity
checks rather than being silently accepted.
- No known concerns within the registry's cooperating-process contract.
-129
View File
@@ -1,129 +0,0 @@
# Task 3 report — secret-safe dwh-auth administrative CLI
## Scope
- Added `tools/dwh-auth/internal/command/command.go`, its command tests, and
`tools/dwh-auth/cmd/dwh-auth/main.go`.
- The CLI accepts only the frozen Task 3 grammar: key create/import/list/status/revoke,
registry check, and the reserved serve invocation.
- Existing Task 1–2 APIs are consumed without modifying their files.
- No server, Nginx, systemd, Compose, portable `tht`, real registry, real secret, or legacy
stack was accessed or changed.
## TDD evidence
Tests were written before `Run` existed. In the official `golang:1.26.5` container, mounted
against only the dedicated worktree, the focused RED run was:
```text
go test ./internal/command -count=1
internal/command/command_test.go:214:10: undefined: Run
FAIL
```
After implementation and formatting:
```text
go test ./internal/command -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/command
go test -race ./internal/command -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/command
go test ./... -count=1
ok internal/command, internal/credential, internal/record, internal/registry, internal/securefile
go test -race ./... -count=1
ok internal/command, internal/credential, internal/record, internal/registry, internal/securefile
go vet ./...
```
## Contract coverage
- Create generates the Task 1 canonical credential, writes it once through the protected
exclusive `0600` output primitive, syncs/closes it before registry publication, and emits
only `created key_id=... installation_id=... output=...`.
- Existing output is never overwritten. Publication failure attempts compensating removal;
cleanup uncertainty returns exit 4 and reports only the output path.
- Legacy import requires `--legacy-raw`, the reserved `legacy-shared` installation ID, and an
absolute exact-`0600` source. It verifies the opaque value, never changes its source, and
stores only its digest.
- List/status expose `PublicRecord` data only; JSON is written as pristine JSON with no digest.
Revoke requires a non-empty reason and reports only its public key ID.
- Relative paths, malformed/unknown flags, duplicate options, invalid IDs/metadata/expiry,
and missing required values return exit 2. Missing status/revoke keys return exit 3.
Registry/filesystem/integrity failures return exit 4.
- Diagnostics are fixed redacted strings. Tests use a sentinel secret and assert it is absent
from stdout/stderr, list/status/check JSON, import output, and unsafe/integrity failures.
## Secret-redaction evidence
The command never prints credential contents or digests. It does not use environment fallback,
interactive stdin, or `flag` diagnostics that echo argument values. The sentinel appears only
in synthetic temporary test input and an integrity-fixture file; all command output assertions
confirm it is absent. The registry’s existing `PublicRecord` contract omits `secret_sha256`.
## Concerns
- `serve` is grammar-reserved and returns a redacted exit-4 unavailable response; Task 4 owns
the Unix-socket service implementation and will wire this dispatch.
- The earlier concern about UTF-8 metadata hardening is superseded by `055dcab`: metadata now
rejects invalid UTF-8 and Unicode controls before generation/output. The nil-safe cleanup note
remains non-blocking and outside this review wave.
## Review-fix wave
Review findings were addressed in separate commit `055dcab`. Regression tests were added first. The focused RED run in the official Go 1.26.5 container failed on intentionally absent seams:
```text
undefined: nowUTC
undefined: addRecord
undefined: closeStore
FAIL github.com/aritmolab/thothii/tools/dwh-auth/internal/command
```
The fix rejects embedded canonical v1 credentials in description/revocation reason without echoing metadata, validates UTF-8/Unicode controls and expiry against one captured UTC creation time before generation/output, reserves exactly `serve --registry-root ABS --socket ABS`, and makes publication cleanup depend on a definitive registry lookup. Output is retained after publication or close ambiguity, with path-only recovery guidance.
Review-fix verification in Go 1.26.5:
```text
go test ./internal/command -count=1 PASS
go test -race ./internal/command -count=1 PASS
go test ./... -count=1 PASS
go test -race ./... -count=1 PASS
go vet ./... PASS
git diff --check PASS
```
New tests cover synthetic canonical credentials embedded with prefix/suffix, invalid UTF-8, C1 Unicode controls, past/equal/future expiry, exact serve ordering, deterministic pre-/post-publication and close-failure seams, and sentinel absence from stdout/stderr/list/status JSON.
## Cleanup snapshot review-fix wave
The second re-review added two regression tests before implementation. The RED run in the
official Go 1.26.5 container showed the old `Find` proof incorrectly treated both cases as
cleanup-safe:
```text
FAIL TestCreateRetainsOutputWhenSnapshotFindsUnrelatedIntegrityFailure
corrupt snapshot result = (4, "", "integrity failure\n")
FAIL TestCreateRetainsOutputWhenFailedPublicationRecordIsExpired
expired publication result = (4, "", "integrity failure\n")
```
Commit `b1079bd fix: retain DWH key output on ambiguous publication` replaces the `Find` proof
with a complete `Store.List()` snapshot. It removes generated output only when the snapshot
succeeds, the generated key ID is absent, and `Store.Close()` succeeds. Any unrelated integrity
error, active/revoked/expired record, or close error retains the output and emits only path-based
recovery guidance. The clean pre-publication failure path still removes the output.
Final cleanup-wave verification in Go 1.26.5:
```text
go test ./internal/command -count=1 PASS
go test -race ./internal/command -count=1 PASS
go test ./... -count=1 PASS
go test -race ./... -count=1 PASS
go vet ./... PASS
git diff --check PASS
```
-60
View File
@@ -1,60 +0,0 @@
# Task 4 — Unix-socket DWH verification report
## Scope
Implemented the standalone Linux verifier at `tools/dwh-auth/internal/service` and wired the exact command:
```text
dwh-auth serve --registry-root ABSOLUTE_CANONICAL --socket ABSOLUTE_CANONICAL
```
The service accepts only `GET /verify`. It returns empty `204` responses with `X-DWH-Key-ID` for verified v1 or reserved legacy credentials; credential failures are generic empty `401` responses, and registry/integrity faults are empty `503` responses. Other paths/methods return empty `404`/`405`.
## Security decisions
- Exactly one `X-API-Key` header, maximum 128 bytes.
- Strict `thtdwh_v1.` parsing precedes legacy lookup; non-v1 values alone may use the reserved legacy record.
- Registry integrity is checked before every verification request, so unrelated malformed/unsafe records fail closed with `503`.
- Logs emit only timestamp, decision code, and (when safely parsed or verified) public key ID; test sentinels prove no key, digest, description, or query value is emitted.
- `serve` validates canonical absolute paths, performs startup `Store.Check`, and reports service startup errors as non-secret `integrity failure`.
- Socket collisions that are regular files, directories, symlinks, live sockets, or foreign-owned stale sockets are refused. Only an owned stale Unix socket after `ECONNREFUSED` can be reclaimed.
- Published sockets are mode `0660`; cancellation calls graceful shutdown and removes only a revalidated same-device/same-inode owned socket. A test seam proves a changed path is retained rather than unlinked.
## Required supporting security fix
Commit `943f809` (`fix: reject duplicate legacy DWH records`) tightens the Task 2 registry contract: a synthetically valid active plus revoked legacy pair is now an integrity failure. It is intentionally separate from the Task 4 commit.
## TDD evidence
RED was observed for the missing handler, listener/configuration API, CLI wiring, unrelated-registry corruption, active+revoked legacy state, and cleanup replacement race. Each increment was then implemented minimally and rerun GREEN.
## Verification
All commands were executed in official `golang:1.26.5`, with only this worktree mounted:
```text
gofmt -w cmd internal/command internal/service
go test ./internal/service ./internal/command -count=1
go test ./... -count=1
go test -race ./... -count=1
go vet ./...
git diff --check
```
All passed. A dependency scan also found no third-party Go dependencies.
## Scope boundary
No Nginx, systemd, real Unix socket, real registry, credential, legacy stack, or external service was changed. All test data was synthetic and temporary.
## Follow-up hardening: runtime read-only registry and socket parent
The Task 5 storage contract uses `root:dwh-auth` SGID directories (`2750`) and a service account with read-only group access. The original registry reader path was incompatible because shared locks were opened `O_RDWR` and lazily created as `0600`; secure-directory validation also rejected SGID.
The runtime path now uses `registry.OpenReadOnly`: it opens only preprovisioned root, `active`, `revoked`, and `.writer.lock` paths, and rejects `Add`/`Revoke`. The administrative `Open` path bootstraps the lock through the exclusive writer path. Shared lock acquisition opens the existing `root:dwh-auth 0640` lock `O_RDONLY` with `LOCK_SH`; writer acquisition remains `O_RDWR` with `LOCK_EX`, preserving cross-process snapshot exclusion. Secure directories allow SGID but still reject setuid, sticky, group-write, and world-write bits.
Task 5 must create `.writer.lock` as `0640 root:dwh-auth` alongside the `2750 root:dwh-auth` registry directories before the service starts.
The socket parent must be a canonical non-symlink directory owned by the service EUID and not group/world writable. This removes the bind-to-chmod and path-replacement exposure from other principals. The remaining POSIX path race is bounded to trusted processes sharing the service EUID inside that non-contendible parent.
Additional verification (official `golang:1.26.5`, worktree only): focused securefile/registry/service/command tests, full tests, full race tests, vet, plus ten race repetitions each for cross-store snapshot readers, `OpenReadOnly`, and listener tests: all PASS.
-71
View File
@@ -1,71 +0,0 @@
# Task 5 report — backend principal enforcement
## RED
Added backend route/auth tests before implementation. The initial focused run failed in
seven new assertions: `getPrincipal` did not exist, upstream requests still required the
legacy identity header, foreign session/SSE routes were not hidden, admin scope was not
enforced, new sessions had no trusted principal binding, and settings were global.
## GREEN
- Focused backend suite: `66 passed` across auth, sessions, SSE, and settings tests.
- Complete backend Vitest suite: `209 passed` across `22` files.
- `npx tsc --noEmit -p .`, `npm run build`, `git diff --check`, and changed Python
source Ruff all exit successfully.
- Harness targeted repository/local/migration tests and Python bytecode compilation exit
successfully. The new `tht session preferences get|set` commands are registered and
expose the expected Typer help. A direct local CLI preference smoke was not run because
the checked-in local workspace requires unavailable `THT_DB_HOST` configuration.
## Route and child-process coverage
- `GET /me` returns the request `PrincipalContext`; upstream accepts only the portal's
normalized `X-Thoth-*` identity tuple, with the legacy header ignored. Local mode uses
the same stable `THT_HOME`/`~/.thothii/identity.json` UUID contract as the harness.
- All session operations are principal-scoped: list (`mine` and admin-only `all`), show,
create, resume, close, delete, rename, group, archive, unarchive, documents, reviewer
response, steer, SQL preview/export, and SSE. Missing and foreign sessions are 404;
absent upstream identity is 401. SSE is authorized before response headers or hub
subscription, so a rejected request cannot attach to a live stream.
- New/resumed Pi runtimes and every route-spawned `tht` process receive
`THT_PRINCIPAL_ISSUER`, `THT_PRINCIPAL_SUBJECT`, optional display name, and admin flag.
The readiness `tht` child is also principal-bound.
- Settings use asynchronous repository-backed `tht session preferences get|set` in the
production runner, which isolates preferences by principal. The legacy settings file is
retained only as an injected-runner compatibility fallback for existing isolated tests.
- Repository/settings authorization failures map to 503 before model startup. SQL execution
errors remain 500 after authorization, preserving the prior API distinction.
## Self-review and concerns
- Confirmed the Task 4 portal emits lowercase `true`/`false` for the admin header; the
parser accepts that exact normalized form plus the repository's existing `1`/`0`
compatibility form, and rejects all other values.
- The harness principal resolver is the ownership authority; the backend never accepts an
owner supplied in request bodies. Its route guards use a repository-scoped `session show`
before every session resource operation.
- Existing dependency-injected route fakes without `sessionShow` retain a narrow test seam;
production `ThtRunner` always has that method, so deployed requests cannot bypass the
repository authorization check.
## Review follow-up
### RED
Focused regressions initially failed exactly at the three review findings: stale ambient
display names survived into both `tht` and Pi child environments; mutation/document runner
methods dropped the selected workspace; and `expandLocalHome` did not exist.
### GREEN
- Child environments now remove all four `THT_PRINCIPAL_*` keys from their cloned base
environment before applying the exact request principal. Regression tests prove an absent
display name does not inherit a stale ambient value in either child path.
- `setName`, `setGroup`, `archive`, `unarchive`, and `documents` now take and retain an
optional workspace. The rename route regression proves `session show` authorization and
the mutation use the same non-default workspace.
- Local principal paths expand `~`/`~/...`; existing local home and identity file modes are
repaired to POSIX `0700`/`0600` when applicable, with Windows left unchanged.
- Focused suite: `74 passed`; full backend suite: `213 passed` across `22` files, followed by
TypeScript typecheck, production build, and diff check.
-298
View File
@@ -1,298 +0,0 @@
# Task 6 — Frontend identity and administrator UX report
## RED
- Added API tests for the `/me` principal call and `mine`/`all` session-list scopes.
- Added component tests for regular-user scope, admin scope switching, owner labels,
administrator banner, foreign-owner delete confirmation, and foreign-owner archive
confirmation.
- Initial focused run: 7 expected failures (missing `getMe`, missing scope query,
missing owner label/admin controls, and missing foreign-action confirmation).
- The archive-confirmation regression was also run separately before its implementation
and failed because `window.confirm` was not called.
## GREEN
- `npx vitest run src/api/sessions.test.ts src/shell/NavSessions.test.tsx src/shell/AppShell.session-mgmt.test.tsx`
— passed (47 tests before the archive follow-up; the focused archive regression then passed).
- `npm test` — passed: 44 files / 305 tests.
- `npx tsc -b` — passed.
- `npm run build` — passed.
- `git diff --check` — passed.
- `npm run e2e` reached Playwright but could not run: the environment has no Chromium
executable at Playwright's configured cache path. No application test failure was reported.
## Files changed
- `frontend/src/api/types.ts`: typed principal and session scope contracts.
- `frontend/src/api/sessions.ts`: typed `/me` API call; scoped listing defaults to `mine`.
- `frontend/src/shell/AppShell.tsx`: identity query, admin-only session scope selector and
banner, owner-aware destructive action confirmations.
- `frontend/src/shell/NavSessions.tsx`: owner labels in the all-sessions view.
- `frontend/src/api/sessions.test.ts`, `frontend/src/shell/NavSessions.test.tsx`, and
`frontend/src/shell/AppShell.session-mgmt.test.tsx`: contract and UX coverage.
## Self-review
- Regular users remain fail-closed on `mine`; no administrator control renders without
`principal.isAdmin`.
- The all-sessions view includes owner labels (including `Unknown` for legacy records).
- Delete confirmation preserves the pre-existing select-all behavior and adds confirmation
for foreign/unknown owners. Foreign archive now also requires an explicit browser
confirmation; existing Stop & save already has its confirmation dialog.
- A read-only review found no critical, important, or minor issues. The archive guard was
added after that review in response to the requirement to cover every destructive rail
action, and has its own RED/GREEN regression plus the final full verification above.
## Concerns
- E2E remains environment-blocked until the Playwright Chromium browser is installed.
- Existing Vitest runs emit pre-existing MSW unmatched-request and dialog-ref warnings; all
assertions pass and this task does not modify those shared test/UI primitives.
## Review remediation
- A post-commit review correctly identified that matching `displayName` must never establish
ownership. The predicate now skips confirmation only when `session.author` exactly equals
`principal.subject`; all display-name matches and missing authors are conservative
cross-owner actions.
- Added RED/GREEN regressions where two principals share display name `Alice` but have distinct
subjects: both delete (with another session present, so select-all cannot mask the guard) and
archive require confirmation.
- Added `aria-pressed` to the My sessions / All sessions controls and asserts their selected state
before and after switching.
- Remediation verification: focused regressions passed; full frontend Vitest (44 files / 305
tests), `npx tsc -b`, `npm run build`, and `git diff --check` all passed.
---
# DWH authentication Task 6 — Nginx and CI gate report
## Scope
Added only the two DWH-auth Nginx gates and the `dwh-auth-linux` deployment workflow job:
- `scripts/test-dwh-auth-nginx-contract.sh`
- `scripts/test-dwh-auth-nginx-integration.sh`
- `.github/workflows/deployment.yml`
This report deliberately remains unstaged. The pre-existing frontend Task 6 report above is
preserved rather than overwritten.
## TDD RED
The structural gate was written before any Task 5 template change. Those templates already met
the approved contract, so the behavioral RED was obtained by copying them into one exact temporary
root and removing only the effective `/dwh/` `auth_request` directive. The new checker failed as
required, with no credential material in output:
```text
case=source_contract status=FAIL
```
The runtime gate was also first invoked before its file existed:
```text
bash: scripts/test-dwh-auth-nginx-integration.sh: No such file or directory
```
The CI-job RED check found no `dwh-auth-linux` job in `deployment.yml`. No production template was
modified: the tests prove the existing Task 5 template contract instead of weakening it.
## GREEN
Shell syntax and workflow YAML were checked with:
```text
bash -n scripts/test-dwh-auth-nginx-contract.sh scripts/test-dwh-auth-nginx-integration.sh
python3 -c import-yaml-and-safe-load
```
The structural gate passed its source contract plus these 13 real copied-and-mutated Nginx fixtures:
```text
missing_auth_request
missing_proxy_method
missing_proxy_body
missing_proxy_header_isolation
missing_content_length_clear
missing_verifier_key_forward
missing_upstream_key_clear
missing_failure_mapping
public_verifier
tcp_authenticator
postgrest_bypass
failure_mapped_to_success
full_secret_rate_key
```
Each test mutates an effective, not comment-only, directive and requires the checker to reject it.
The source test and all 13 fixture tests emitted `case=... status=PASS`, followed by
`case=summary status=PASS`.
The isolated Nginx 1.24 smoke passed these sanitized cases:
```text
nginx_1_24
build_dwh_auth
registry_setup
verifier_start
synthetic_upstreams
composite_nginx_config
nginx_start
auth_socket_unix_only
verifier_not_public
valid_v1
valid_legacy
invalid_key
revoked_key
expired_key
duplicate_v1
duplicate_legacy
stopped_verifier
header_and_path_isolation
summary
```
It builds with the pinned official Go 1.26.5 image when the host Go binary is absent, creates only
synthetic v1, legacy, revoked, and expired credentials in a `0700` `/tmp` root, runs both Nginx and
the verifier on explicit temporary Unix sockets, and uses a loopback-only marker backend. Its output
is strictly `case` and `status`; keys, values, and digests remain only in the exact temporary root
and are removed by the trap.
`nginx -t` passed against the complete generated configuration. The marker proves that successful
`/dwh/?keep=exact&second=two` reaches the upstream unchanged, while neither the client API key nor
client or verifier `X-DWH-Key-ID` reaches it. A Unix forwarding probe proves that the verifier sees
only `X-API-Key`, with Cookie, Authorization, and spoofed audit ID absent. Duplicate v1 and ordinary
legacy headers return 401 through Nginx; a stopped verifier returns 503.
The final local equivalent of the four CI commands passed:
```text
Docker Go 1.26.5: go test -race ./... -count=1 and go vet ./...
bash scripts/test-dwh-auth-build-contract.sh
bash scripts/test-dwh-auth-nginx-contract.sh
bash scripts/test-dwh-auth-nginx-integration.sh
```
The Go race suite passed for command, credential, record, registry, securefile, and service;
`go vet` was silent; the build contract passed; both Nginx gates reached their summaries.
## CI contract
The new job uses `actions/checkout` with `persist-credentials: false`, pins Go 1.26.5 with cache
keyed on `tools/dwh-auth/go.mod`, installs `nginx-light`, and runs exactly the four required commands.
Existing jobs were not altered.
## Self-review
- The template tests parse normalized effective directives, so commented-out declarations cannot
satisfy the gate.
- The authentication socket is configured as `http://unix:...:/verify`, is observed by `ss -xl`,
and Nginx itself listens only on a temporary Unix socket; neither test starts a public listener.
- All spawned processes are registered by PID; cleanup signals only those PIDs and deletes only the
exact `mktemp` root after a guarded path check.
- The verifier, marker, registry, Nginx prefix, PID, logs, config, and sockets all reside beneath
that root. No `/etc`, systemd, active Nginx config, stack, legacy route, or real registry/key is
read or changed.
- Task 5 templates were not modified because the structural and runtime tests passed unchanged.
## Concern
The sandbox `apply_patch` helper repeatedly failed with `bwrap: loopback: Failed RTM_NEWADDR:
Operation not permitted`. A narrowly scoped fallback editor was used only for the workflow and the
Nginx-version assertion. Its first workflow insertion interpreted the action-reference at signs;
the two malformed values were immediately corrected and all final YAML, exact-string, syntax, and
four-command checks were rerun. No remaining product concern is known; the integration gate requires
Nginx 1.24 and Python 3, both supplied by the specified Ubuntu CI runner.
---
# DWH authentication Task 6 — review remediation wave
## Review findings and RED evidence
The three review findings were reproduced against the Task 6 commit before their corresponding
hardening was accepted.
1. The contract checker originally selected only the first matching `/dwh/` location. A real copied
fixture appended this competing location without authentication:
```nginx
location ~ ^/dwh/ {
proxy_pass http://127.0.0.1:3001;
}
```
The first run reached the new check and failed as required:
```text
case=negative_postgrest_regex_bypass status=FAIL
```
2. The previous process stop sent TERM and immediately used an unbounded `wait`. A synthetic Python
child ignored TERM; the RED run used one exact short-lived watchdog only to prevent a test hang and
produced:
```text
case=cleanup_term_ignored_bounded status=FAIL
```
3. The TCP detector has a positive-control regression. A scratch copy of the integration script
replaced its `ss -ltnpH` detector with `return 1`; its known loopback listener was then not
detected and the run failed with:
```text
case=tcp_listener_detector_positive status=FAIL
```
All RED fixtures and the scratch script used an exact temporary path and were removed. No template,
service, workflow, key, or active Nginx configuration was changed.
## GREEN changes
- `location_declarations` consumes normalized, comment-stripped effective lines and `check_templates`
requires exactly one each of the only approved locations: verifier, unavailable named location, and
`/dwh/`. It therefore rejects both any extra intercepting location and a duplicate. The real regex
bypass and a new real duplicate `/dwh/` bypass fixture both pass by being rejected.
- `tcp_listener_for_pid` uses `ss -ltnpH` and a PID-bound match. The integration gate starts a
loopback-only synthetic listener, proves the detector sees that exact PID, stops and deregisters it,
then proves the verifier PID has no TCP listener while its Unix socket remains present.
- `stop_registered_pid` now sends TERM, polls for exit or zombie for a bounded deadline, sends KILL
if required, polls a second bounded deadline, and only reaps a direct child after terminal state is
proved. Explicit stops deregister their PID. The cleanup loop invokes that bounded operation only
for recorded PIDs and removes only its guarded temporary root.
- The synthetic child that ignores TERM is killed by the bounded path, must no longer answer to
`kill -0`, must not remain registered, and must finish within three seconds. Final gate output is
restricted to `case` and `status` lines.
## GREEN verification
```text
bash -n scripts/test-dwh-auth-nginx-contract.sh scripts/test-dwh-auth-nginx-integration.sh
Docker Go 1.26.5: go test -race ./... -count=1 and go vet ./...
bash scripts/test-dwh-auth-build-contract.sh
bash scripts/test-dwh-auth-nginx-contract.sh
gate contract: source plus 15 negative fixtures PASS, then summary PASS
bash scripts/test-dwh-auth-nginx-integration.sh
gate integration: 20 named cases PASS, then summary PASS
git diff --check
```
The integration cases include `cleanup_term_ignored_bounded`,
`tcp_listener_detector_positive`, `auth_socket_unix_only`, all existing credential decisions,
composite Nginx syntax, and stopped-verifier 503 behavior. Go race tests passed for command,
credential, record, registry, securefile, and service; vet and both diff checks were silent.
## Self-review and concern
The new location parser rejects comment-only and non-exact declarations because it operates on the
same normalized effective representation used by the rest of the contract. The TCP positive control
binds only `127.0.0.1` on a kernel-selected temporary port and is stopped through the same exact-PID
path under test. The bounded cleanup avoids arbitrary process lookup or broad signaling.
The environment still intermittently rejects `apply_patch` with the sandbox loopback error noted in
the original report; only narrowly scoped fallback edits to the two authorized scripts were used and
all final gates were rerun. No remaining review concern is known.
-90
View File
@@ -1,90 +0,0 @@
# Task 7 — report
## RED
- Creato `scripts/test-verify-dwh-auth-docs.sh` con fixture positiva e fixture negative per
credenziale/digest sintetici, TLS insicuro, segreto in env/argv, mode world-readable, cattura
Nginx e coupling Compose.
- Eseguito `bash scripts/test-verify-dwh-auth-docs.sh` prima del verificatore: `case=verifier_missing status=FAIL`.
## GREEN
- Aggiunti manuali server, client, TLS, runbook PSD, collaudo ed evidenza sanitizzata; collegati
manuali locali/server, setup PSD, guida, indice e nav MkDocs.
- Eseguiti: `bash -n scripts/verify-dwh-auth-docs.sh scripts/test-verify-dwh-auth-docs.sh`,
`bash scripts/test-verify-dwh-auth-docs.sh`, `bash scripts/verify-dwh-auth-docs.sh`,
`bash scripts/test-verify-workspace-install-docs.sh`, `bash scripts/auth-docs-smoke.sh`.
- Tutti gli output finali sono PASS; il nuovo gate esercita una fixture positiva e nove negative.
## Self-review
- Verificati path/owner/mode: registry 2750, lock/record 0640, socket 0660.
- Verificata separazione: chiavi solo `rest_api`; PSD server `postgres_direct`; Mac/remoti REST;
nessun lifecycle Compose per `dwh-auth`.
- Verificati TLS `.it`/SAN, `.com` non coperto, `TLS_CA_FILE`, fingerprint fuori banda, rinnovo e
assenza di bypass.
- Verificati due gate Task 9–10, evidenze solo metadati e nessuna migrazione di sessioni/index/cache legacy.
## Concern
- Nessuna mutazione PSD/Nginx/systemd/registry o lettura di segreti è stata eseguita. I comandi del
runbook restano condizionati alle autorizzazioni separate dei Task 9 e 10.
## Review fix — RED/GREEN
### RED review
- La fixture `sudo nginx -T` ha prodotto il rifiuto `case=sudo_raw_nginx_capture status=FAIL` prima della correzione del gate.
- La fixture header legacy opaco ha prodotto `case=opaque_legacy_header_literal status=FAIL` prima della correzione del gate.
- Dopo avere riallineato le label UI nei manuali, `bash scripts/test-verify-workspace-install-docs.sh` ha prodotto `server-workspace-registry.md: curator flow missing registry rule`: il verifier cercava ancora le due label precedenti. Il test sulla base HEAD e il diff hanno confermato la causa.
### GREEN review
- Il gate DWH ora rifiuta anche header opaco, digest JSON quotato, `export` di API key, `curl --header` e `-H`, `sudo nginx -T`, raw diff e Compose; le mutation fixture coprono label, PSD direct/Mac REST/CA, socket e flag REST.
- Il runbook non prescrive raw diff o dump: solo checker strutturale e secret scan con metadati e PASS/FAIL. Il piano Task 10 adotta la stessa regola.
- Il template `psd-local` resta `rest_api` solo Mac/local/remota; il server PSD Project A resta `postgres_direct` con binding separato. La CA privata e `TLS_CA_FILE` sono obbligatori salvo trust approvato equivalente.
- Le procedure server ora coprono backup manifest protetto, restore, curl config 0600 senza segreto in argv/env/output, Unix 204/401, HTTPS 2xx/401, 503 bounded con trap, journal PASS/FAIL e retention alla disinstallazione.
- Il verifier workspace-install e entrambi i manuali registry usano ora le quattro label effettive: `Validate workspace source`, `Test workspace connections`, `Save entered secrets`, `Forget stored value`.
### Final verification review
- PASS: `bash scripts/test-verify-dwh-auth-docs.sh`.
- PASS: `bash scripts/verify-dwh-auth-docs.sh`.
- PASS: `bash scripts/test-verify-workspace-install-docs.sh` (fixture complete).
- PASS: `bash scripts/auth-docs-smoke.sh`.
- PASS: `bash -n scripts/verify-dwh-auth-docs.sh scripts/test-verify-dwh-auth-docs.sh` e `git diff --check`.
### Review concern
- Nessuna mutazione runtime e nessun segreto reale sono stati letti. I soli comandi server documentati restano soggetti ai gate autorizzativi Task 9 e Task 10.
## Review fix wave 2 — RED/GREEN
### RED wave 2
- Prima della correzione del proxy, `bash scripts/test-dwh-auth-build-contract.sh` ha fallito il contratto di preservazione path e `bash scripts/test-dwh-auth-nginx-integration.sh` ha chiuso con `case=header_and_path_isolation status=FAIL`: il prefisso `/dwh` arrivava a PostgREST invece di essere rimosso.
- Prima delle procedure finali, il gate docs ha rifiutato il path chiave non deterministico e la fixture curl con header legacy opaco ha dato `case=header_file_curl_synthetic status=FAIL` perché il valore non veniva confrontato esattamente.
- Le mutation fixture hanno catturato l'estrazione tar sul registro attivo e i rename non protetti. Dopo l'inasprimento finale del gate, la sorgente ha dato `dwh-auth docs: restore must stage/check then use guarded same-filesystem renames` finché mancava il controllo fail-closed del candidato.
- Il RED finale dello scanner journal è stato `dwh-auth docs: docs/install/dwh-auth-server.md lacks required topic: sys.argv[2:]`: il gate esige la lettura byte-esatta di v1 e legacy e un `journalctl` che fallisca chiuso.
### GREEN wave 2
- Commit `f616aab fix: preserve PostgREST RPC path through DWH proxy`: `proxy_pass` termina con `/`; il contratto e l'integrazione verificano `/dwh/rpc/ping?x` verso `/rpc/ping?x`.
- Il runbook usa un singolo file chiave v1, header file `0600` passati solo con `curl --header @file`, socket 204 dual-key, HTTPS 2xx pre/post per v1 e 401 post-revoca per legacy `legacy-shared`.
- Restore protetto: staging sul filesystem `/var/lib`, check candidato, `mv -T --` guardato per ogni publish/rollback e pre-restore conservato. Backup/manifest restano root-only `0600` su storage cifrato approvato.
- Lo scanner journal esegue `journalctl` in un unico processo Python root, sopprime stderr, controlla return code e bytes esatti di entrambe le chiavi senza emettere journal o segreti; la shell mostra solo PASS/FAIL.
- Il verifier rifiuta `curl --config`, header in argv, raw Nginx/diff, TLS insicuro, segreti env, mode insicuri e Compose. Le fixture mutano path chiave, ID legacy, header/legacy probes, restore, journal, codici HTTPS e label UI.
### Final verification wave 2
- PASS: `bash scripts/test-dwh-auth-build-contract.sh`.
- PASS: `bash scripts/test-dwh-auth-nginx-contract.sh`.
- PASS: `bash scripts/test-dwh-auth-nginx-integration.sh`.
- PASS: `bash scripts/test-verify-dwh-auth-docs.sh` e `bash scripts/verify-dwh-auth-docs.sh`.
- PASS: `bash scripts/test-verify-workspace-install-docs.sh` e `bash scripts/auth-docs-smoke.sh`.
- PASS: `bash -n` sugli otto gate shell e `git diff --check`.
### Review concern wave 2
- Nessuna configurazione protetta, chiave reale, Nginx, systemd o stack PSD è stata letta o mutata. Le procedure privilegiate restano istruzioni condizionate ai Gate 9–10; la verifica degli owner/mode reali è un'attività del rollout autorizzato, non di questo task documentale.
+3 -1
View File
@@ -7,7 +7,9 @@ This file provides guidance to Codex (Codex.ai/code) when working with code in t
Read [PROJECT_STATE.md](PROJECT_STATE.md) for the current-state snapshot: what was last
built, pending manual gates, workspace/secret layout, and design-doc locations. This file
holds the stable commands + architecture mental model; PROJECT_STATE.md holds the evolving
detail. Design history lives in `docs/superpowers/specs/` and `docs/superpowers/plans/`.
detail. Current architecture and contracts live in `docs/architecture/`, `docs/contracts/`,
and `docs/evidence.md`; durable design decisions live in `docs/adr/`. Git history is the source
for superseded designs and implementation plans.
## Commands
-105
View File
@@ -1,105 +0,0 @@
# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## Start here
Read [PROJECT_STATE.md](PROJECT_STATE.md) for the current-state snapshot: what was last
built, pending manual gates, workspace/secret layout, and design-doc locations. This file
holds the stable commands + architecture mental model; PROJECT_STATE.md holds the evolving
detail. Design history lives in `docs/superpowers/specs/` and `docs/superpowers/plans/`.
## Agent skills
### Issue tracker
Issues and specifications are tracked in GitHub Issues for `mptyl/ThothII`.
See `docs/agents/issue-tracker.md`.
### Triage labels
Use the standard Matt Pocock triage roles and their corresponding GitHub labels.
See `docs/agents/triage-labels.md`.
### Domain documentation
This repository uses a single-context domain layout: `CONTEXT.md` at the
repository root, with repository-wide ADRs stored under `docs/adr/`.
See `docs/agents/domain.md`.
## Commands
The repo has three independently-built layers. Run the **full stack** (real Pi + DWH, needs
VPN + `harness/.env` + `pi` on PATH) with `./scripts/run-stack.sh` (frontend :5173 → backend :8787).
**harness/** (Python `tht` CLI + Pi gate extension)
- Install: `cd harness && python -m venv .venv && pip install -e ".[dev]"` (puts `tht` on PATH)
- Test: `.venv/bin/pytest -q` — `l2` (real GLM + remote DB) is opt-in via `addopts = -m 'not l2'`; `l0` (testcontainers) needs Docker
- Single test: `.venv/bin/pytest tests/test_session_mutations.py::test_set_name -v` (or `-k <pattern>`); include e2e with `-m l2`
- Lint: `.venv/bin/ruff check .` (line-length 100)
**backend/** (Fastify + TypeScript, vitest)
- Dev: `npm run dev` (tsx watch `src/server.ts`) · Build: `npm run build` (tsc → `dist/`)
- Test: `npx vitest run` · Single: `npx vitest run test/routes-sessions.test.ts -t "rename"`
- Typecheck: `npx tsc --noEmit -p .` (vitest does NOT type-check — run this before committing)
**frontend/** (React 18 + Vite + vitest)
- Dev: `npm run dev` (Vite; set `VITE_BACKEND_URL`) · Build: `npm run build`
- Test: `npx vitest run` · Single: `npx vitest run src/shell/NavSessions.test.tsx`
- Typecheck: `npx tsc -b` · E2E: `npm run e2e` (Playwright)
No ESLint on the TS layers — `tsc` is the gate. Tests use vitest + MSW (no network).
## Architecture (the parts that need multiple files to see)
```
frontend (React/SSE) → backend (Fastify) → pi --mode rpc → tht/harness → DWH (read-only)
```
- **The harness owns the workflow and all persistence.** `tht` (Python) is a deterministic
CLI; `harness/.pi/extensions/tht-gate.js` is a Pi extension that drives an **8-phase
NL→SQL workflow**. The single source of workflow truth is `harness/workflow.yaml`; the
orchestration rules the model must follow are `harness/.pi/skills/tht-sessione/SKILL.md`.
"Current phase" is computed by folding the decision ledger (`harness/tht/phase.py`), not
stored — read it before reasoning about phase logic.
- **Persistence = phase documents, NOT chat.** A session is a directory under the workspace's
`sessions/` path: `session_manifest.yaml` + per-phase artifacts (`question.md`,
`schema_linking.json`, `sql_final.sql`, …) + `review_decisions.jsonl`. The contract
(SKILL.md): *"the persisted state is the truth — what is not recorded did not happen."*
There is no verbatim transcript store. A resumed Pi process rebuilds context from
`tht session show <id>` + the on-disk artifacts.
- **The backend is a thin bridge with no database of its own.** `ThtRunner` shells `tht`
subcommands; `PiProcessManager` runs one Pi child per session and bridges its RPC stream;
`SessionBridge` maps Pi RPC events → client events (`ui_request`/`text_delta`/`info`);
`SseHub` fans them out over SSE to the browser. Persistence belongs to the HARNESS, which
selects the session repository from the workspace config (`harness/tht/session/repository.py`):
filesystem by default, **PostgreSQL when `session_storage` is configured** (server/portable
deployment). Settings flow through harness preferences (`tht session preferences`) with
`backend/data/settings.json` only as the file fallback for injected runners/tests.
- **Human-in-the-loop gate contract.** The model proposes; a human reviewer decides at gates
via widgets (`reviewer_select` = single pick — a chosen option carrying a `decision` payload
auto-confirms/persists directly, an option without one only asks; `reviewer_decide` = multiselect,
each choice IS a decision; `reviewer_confirm` = artifact/phase gate). The frontend renders these
widget-descriptors (`src/widgets/` registry) and the live transcript is rebuilt in-memory
from the SSE stream (`src/store/sessionStore.ts`) — it is not persisted.
## Project-specific gotchas
- **`tht`'s `-c`/`--config` is a PER-COMMAND option** — it must follow the subcommand, never
precede it (`ThtRunner.buildArgv` enforces this; prepending caused live 500s).
- **`--json` output must be pristine** (only valid JSON on stdout) — used as a machine contract.
- **UI strings are English; document *content* stays the workspace language** (Italian for
`psd`) because it's the real data. Only chrome/labels are English.
- **Workspaces** (`harness/workspaces/*.yaml`) set the DB target and **absolute**
`paths.sessions/artifacts/indexes` — for `psd` these point at a *separate, uncommitted* repo
(`tht-workspace-psd/`). Secrets live ONLY in `harness/.env` (gitignored).
- **Settings are global** (workspace/provider/model/thinking, persisted via harness
preferences — `backend/data/settings.json` is only the fallback); the New-session form is
question-only.
- **Resume**: a resumable session re-enters at its last incomplete phase. The backend refuses
resume with 409 when `finalized` or `archived`, and `PiProcessManager.spawnFor` must send
`/riprendi-sessione <id>` (resume mode) vs `/nuova-domanda` (new) — sending the wrong prompt
silently turns a resume into a new question.
+82 -1313
View File
File diff suppressed because it is too large Load Diff
+7 -9
View File
@@ -202,20 +202,18 @@ Startup mode adds bounded image build/two-service health startup, installation-a
status, stopped-container-aware ownership checks, and exact cleanup. The ordinary hosted Windows
job remains deterministic and does not claim Docker startup.
## Preprocessing jobs and S3 Evidence
## Workspace preprocessing and S3 Evidence
The included preprocessing services reuse the internal Qdrant/Ollama stack. Mount Evidence at
`/data/source/evidence`, then run the explicit preprocessing preset:
Run preprocessing through the native host CLI and the installation descriptor:
```sh
docker compose --env-file deploy/env/local.env \
-f compose.yaml -f deploy/compose.local.yaml \
-f deploy/compose.preprocess.yaml --profile preprocess run --rm preprocess-evidence
tht --installation /absolute/path/thothii-installation.yaml workspace preprocess evidence
tht --installation /absolute/path/thothii-installation.yaml workspace preprocess dwh
```
Replace the final service with `preprocess-dwh` when required. The overlay makes each job wait for the internal Qdrant
service health checks and embedding model initialization; no separate semantic-service startup is
required.
The CLI starts the profile-gated `workspace-maintenance` service and enforces the workspace,
secret, Qdrant, and embedding contracts. See [Evidence](docs/evidence.md) and the
[workspace preprocessing CLI contract](docs/contracts/workspace-preprocessing-cli.md).
S3 Evidence uses the optional `tht[s3]` dependency and canonical `s3://bucket/key` provenance.
AWS endpoints are used when no custom URL is supplied. Every custom endpoint is an explicit egress
@@ -15,9 +15,6 @@ const allowedKinds = new Set(["policy_text", "workspace_descriptor", "deployment
// opener line through the closer line (including physical line endings). These
// blocks are reviewed non-workspace runtime/config generation, not semantic proof.
const reviewedExpandableBlocks = new Map([
["scripts/preprocess-smoke.sh", [
{ sha256: "fc530dc721c946644ab6552bbd46b7918d6c5f11f06f3495b6ea1fcda819b38d", rationale: "Generates the reviewed preprocess Compose override." },
]],
["scripts/test-dwh-auth-nginx-integration.sh", [
{ sha256: "ead57234ad3520b5c7d4262b772957cbc7b9589da4f35fb17b160f948eb2ac7b", rationale: "Generates the reviewed isolated Nginx integration configuration." },
]],
@@ -624,7 +624,6 @@ test("an in-band marker cannot authorize expandable content", async (t) => {
test("current exact reviewed expandable blocks pass only at their trusted paths", async (t) => {
const reviewedPaths = [
"scripts/preprocess-smoke.sh",
"scripts/test-server-pi-state-topology.sh",
"scripts/test-vector-backup-restore-safety.sh",
"scripts/test-windows-clone-contract.ps1",
@@ -636,19 +635,6 @@ test("current exact reviewed expandable blocks pass only at their trusted paths"
root: repositoryRoot,
entries: reviewedPaths.map((path) => entry("deployment_script", path)),
});
const root = await fixture(t);
const original = await readFile(join(repositoryRoot, "scripts/preprocess-smoke.sh"), "utf8");
await put(root, "scripts/copied-preprocess.sh", original);
await assert.rejects(
verifyEntries({ root, entries: [entry("deployment_script", "scripts/copied-preprocess.sh")] }),
/exact-content reviewed allowlist/,
);
await put(root, "scripts/preprocess-smoke.sh", original.replace('$tmp/smoke.yaml', '$tmp/other.yaml'));
await assert.rejects(
verifyEntries({ root, entries: [entry("deployment_script", "scripts/preprocess-smoke.sh")] }),
/exact-content reviewed allowlist/,
);
});
test("PowerShell tokenizer ignores opener text in comments and ordinary strings", async (t) => {
@@ -1,14 +0,0 @@
# Datamart Builder deployment gotchas
- Il percorso pubblico attraversa due reverse proxy: nginx host → nginx Omics Portal → core ThothII.
- Per SSE, ogni livello deve disabilitare `proxy_buffering`, `proxy_request_buffering` e cache, usare HTTP/1.1, timeout lunghi e propagare `X-Accel-Buffering: no`.
- Il core aggiunge `X-Accel-Buffering: no` alla risposta EventSource; Omics Portal lo riaggiunge esplicitamente per i proxy a monte.
- Una sonda utile deve attraversare il portale autenticato e misurare l'arrivo degli header `200 text/event-stream`, non solo interrogare il core nel network Docker.
- La configurazione nginx host attiva vive in `/etc/nginx/sites-available/policlinicosandonato`; validare con `nginx -t` prima del reload.
- Django può tenere in memoria il manifest Vite per worker: dopo un rebuild del frontend riavviare anche i worker Omics Portal, altrimenti richieste diverse possono produrre hash asset vecchi e nuovi.
- Nel profilo server legacy, `vector_db` è una connessione pgvector RW condivisa; il factory può riusarla come writer solo con `profile=server`. La workstation senza writer REST deve restare read-only.
- Non convertire in-place `local.yaml` da chiavi legacy a risorse moderne durante un incident fix: cambia il binding del workspace e può invalidare gli snapshot DWH attivi.
- `session_storage` è configurazione di persistenza delle sessioni, non degli artefatti DWH: escluderla dall'impronta in `config_dwh_binding`, altrimenti l'attivazione di profili utente rende incompatibile un indice DWH già valido.
- Con profili per utente, un profilo privato vuoto deve essere inizializzato una sola volta dai settings legacy completi; trattare `{}` come settings effettivi fa creare sessioni senza provider/modello e lo spawn Pi fallisce prima di partire.
- Regola operativa concordata: dopo ogni modifica o enhancement, ricostruire e ricreare tutti i container ThothII coinvolti prima dell'handoff, quindi verificare che siano healthy affinché l'utente possa testare live. Il push Git non aggiorna le immagini Docker automaticamente.
-11
View File
@@ -1,11 +0,0 @@
# Pi model selection
- `get_available_models` restituisce tutti i modelli con autenticazione/configurazione disponibile; non applica `settings.json.enabledModels`.
- Il selettore web usa l'intersezione esatta tra catalogo RPC e `enabledModels`, preservando l'ordine configurato e confrontando `provider/model`.
- In produzione `PI_PROVIDER` può essere assente: il provider corrente vive in `/data/settings/settings.json`; l'enumerazione Pi deve quindi essere provider-neutral e usare il profilo montato per l'autenticazione.
- Lo spawn di enumerazione deve comunque passare da `buildPiChildEnv({})` per rimuovere credenziali ambientali e metadati dei secret.
- `local-qwen` è un provider custom locale esplicito; non richiede la chiave generica dei provider hosted. Provider sconosciuti e composti restano fail-closed.
- Il profilo Pi live è montato da `/home/chirone/thothii-data/pi-config` a `/home/thoth/.pi`; non leggere né stampare mai i valori di `auth.json`.
- Il provider live `local-qwen` usa `http://localllm-vllm:8000/v1` e il modello `qwen3.6-35b-a3b`; il file montato conserva owner/mode e non contiene modifiche alle credenziali.
- In Compose solo `core` entra nella rete esterna `localllm_default`, oltre alla rete del portale; `frontend` non deve avere accesso diretto alla rete del modello.
- La connettività è verificata dal namespace di `core`: catalogo `/v1/models`, modello atteso presente e completion reale non vuota, senza stampare il testo della risposta.
-20
View File
@@ -1,20 +0,0 @@
# PSD DWH transport
- Il workspace PSD supporta `rest_api` e `postgres_direct`; il trasporto è scelto dall'installazione.
- Il Mac usa REST/PostgREST; il ThothII installato sul server PSD usa PostgreSQL diretto.
- Il server deve usare un ruolo DWH dedicato: `USAGE` sullo schema e `SELECT` soltanto, verificati sui grant reali.
- La difesa read-only applicativa aggiunge: SQL strutturalmente SELECT-only, transazione `READ ONLY`, timeout, limite e rollback.
- La credenziale DWH REST `X-API-Key` non deve essere fornita al ThothII server.
- REST è il trasporto stabile per il Mac e per future installazioni remote che non possono aprire tunnel SSH.
- L'autenticazione REST target deve dare a ogni installazione un'identità separata, revocabile e auditabile; una chiave globale condivisa non scala.
- La rotazione della chiave globale esposta richiede una finestra dual-key che preservi il Mac prima della revoca.
- Il certificato REST corrente resta invariato: è self-issued e viene presentato anche dall'endpoint esterno; i client che non lo considerano già trusted richiedono `TLS_CA_FILE`.
- Il manuale deve coprire consegna e fingerprint della CA, scadenza/rinnovo coordinato e il fatto che i SAN correnti coprono `.it`, non `.com`.
- Non esiste un ambiente di test PSD: le rotazioni devono usare backup, dual-key, probe read-only e rollback sull'endpoint di produzione.
- Supabase Studio è un pannello amministrativo loopback, non un data-plane o un trasporto per ThothII.
- Le credenziali DWH, session storage e amministrazione Supabase devono restare separate.
- Il collegamento container→PostgreSQL deve usare un endpoint host/rete esplicito e ristretto; il loopback dell'host non è il loopback del container.
- Il nightly ETL PSD delle 03:00 usa PostgreSQL diretto; non è un consumer della route REST `/dwh/`.
- Un preprocessing REST Thoth genera `1 + 3T + Ct + Ce` richieste: una lista tabelle, tre RPC per tabella e due famiglie di campionamento testuale.
- Per PSD nel run 2026-08-13: `T=163` e `Ct+Ce=1180`, quindi 1670 richieste per ciclo; undici rerun spiegano 18.370 richieste.
- I repository server esistenti contengono materiale sensibile hardcoded: non copiarlo; inventariare, rimuovere dal tracking e ruotare i segreti coinvolti.
-44
View File
@@ -1,44 +0,0 @@
# ThothII workflow UI contracts
- Pi reasoning arrives as nested `message_update.assistantMessageEvent.type = thinking_delta`.
The backend maps it to the named SSE event `activity_delta`; the frontend EventSource must
explicitly subscribe to that name. Model activity is separate from final `text_delta` output.
- Session create/resume must preserve the configured or persisted thinking level. Forcing
`thinking: off` disables the upstream signal and makes the activity panel legitimately empty.
- A join-only `reviewer_decide` proposal is one complete, atomic join set. The read-only
`join-review` widget persists every proposed join on Continue; `Other — specify` persists none
and requires the model to propose the complete corrected set again. Ledger read, sequence
assignment, and atomic replacement share a per-session cross-process writer lock.
- A v2 phase summary accepts `open_questions?: string[]`. Validate this at the gate boundary and
normalize legacy malformed entries defensively in the viewer so one object cannot crash React.
- `SqlViewer`'s horizontal/vertical layout control is meaningful only with multiple SQL blocks;
hide it for the single CTE result shown by `CteResultViewer`.
- A Pi turn is `idle`, `running`, `waiting`, or `failed`. A reviewer gate/request moves it to
`waiting`; the reviewer response and steering move it back to `running`; provider failures and
unexpected Pi child exits mark it `failed` without forwarding raw failure detail.
- Resume preserves `running`/`waiting` runtimes. After manifest/readiness validation, every cold
path clears old SSE buffer/subscribers before reopen—including when a crashed child has already
left no runtime—and reuses persisted provider/model/thinking. Failed validation does not clear.
- A successful Resume of the already selected session increments the stream generation so React
closes the old EventSource and opens the same session URL again. Failed Resume must not reconnect.
- SSE endpoints are intentionally keep-alive. Browser cleanup and one-off probes must explicitly
close the EventSource or cancel/abort the response reader after their terminal event.
- `activityLog` remains the complete in-memory chronological fold. The left Model activity panel
default-denies every kind except prompt, thinking, and assistant, presenting them as Question,
Reasoning, and Response in source order; status, tool, gate, lifecycle, and unknown kinds stay
hidden. Its desktop width is pointer/keyboard resizable from 288–576 px while preserving 512 px
centrally, persists globally in localStorage, and becomes an overlay drawer below `lg`.
- F6 CTE cards keep their semantic structure, responsive grids, 12/16 px lateral padding, and 8 px
header/content edge padding. Internal section gaps are 12 px, headings/dividers use 4 px spacing,
and table/filter rows use 4 px vertical padding with compact line heights.
- Pi tool events may cross the backend/client boundary only as call id, tool name, and
`running`/`completed`/`failed` status. Tool updates, arguments, partial/final results, commands,
raw output, and raw errors remain server-side.
- The frontend uses Tailwind CSS 3.4. Shared primitives must use concrete Tailwind 3-compatible
spacing utilities; Tailwind 4 custom-spacing syntax can compile to no effective padding here.
- A native EventSource can retain a `Last-Event-ID` from an older backend process. If that cursor
is newer than every id produced by the current `SseHub` generation, treat it as stale and replay
the fresh generation from id 0 instead of suppressing all new low-id events.
- Batch delete mutates client selection, active-session, panel, and Resume intent only for ids whose
DELETE actually succeeded. A failed active delete preserves its live binding, and deleting an
unrelated session must not invalidate a concurrent Resume targeting another id.
-6
View File
@@ -1,6 +0,0 @@
# Brain
- [[codebase/datamart-builder-deployment-gotchas]]
- [[codebase/pi-model-selection]]
- [[codebase/psd-dwh-transport]]
- [[codebase/workflow-ui-contracts]]
-44
View File
@@ -1,44 +0,0 @@
# Retired for operator use: this profile remains only as a non-public engine-fixture path.
# It exercises the legacy preprocessing fixtures and must not become a second operator interface.
services:
preprocess-evidence:
image: thothii-core:local
profiles: [preprocess]
build:
context: .
dockerfile: docker/core.Dockerfile
entrypoint: [sh, -ec]
command: ["mkdir -p /data/workspaces/preprocess-evidence && exec /app/docker/core-entrypoint.sh preprocess evidence --json -c /app/harness/workspaces/preprocess-evidence.yaml"]
environment:
THT_DATA_ROOT: /data
THT_SECRETS_FILE: /run/secrets/thothii.secrets
secrets: [{source: thothii_secrets, target: thothii.secrets}]
volumes:
- thoth_data:/data
- ./deploy/workspaces:/app/harness/workspaces:ro
restart: "no"
depends_on:
qdrant:
condition: service_healthy
embedding-model-init:
condition: service_completed_successfully
preprocess-dwh:
image: thothii-core:local
profiles: [preprocess]
build:
context: .
dockerfile: docker/core.Dockerfile
entrypoint: [sh, -ec]
command: ["mkdir -p /data/workspaces/preprocess-dwh && exec /app/docker/core-entrypoint.sh preprocess dwh --steps introspect --json -c /app/harness/workspaces/preprocess-dwh.yaml"]
environment:
THT_DATA_ROOT: /data
THT_SECRETS_FILE: /run/secrets/thothii.secrets
secrets: [{source: thothii_secrets, target: thothii.secrets}]
volumes:
- thoth_data:/data
- ./deploy/workspaces:/app/harness/workspaces:ro
restart: "no"
volumes:
thoth_data:
-11
View File
@@ -1,11 +0,0 @@
language: en
dwh:
type: postgres_direct
connection:
host: "${THT_PREPROCESS_DWH_HOST:-dwh}"
port: "${THT_PREPROCESS_DWH_PORT:-5432}"
database: "${THT_PREPROCESS_DWH_DATABASE:-warehouse}"
schema: "${THT_PREPROCESS_DWH_SCHEMA:-public}"
user: "${THT_PREPROCESS_DWH_USER:-thoth_reader}"
password_file: "${THT_PREPROCESS_DWH_PASSWORD_FILE:-/run/secrets/preprocess-dwh-password}"
roots: {artifacts: artifacts, indexes: indexes, sessions: sessions}
@@ -1,22 +0,0 @@
language: en
dwh:
type: postgres_direct
connection:
host: "${THT_PREPROCESS_DWH_HOST:-unused}"
port: "${THT_PREPROCESS_DWH_PORT:-5432}"
database: "${THT_PREPROCESS_DWH_DATABASE:-unused}"
schema: "${THT_PREPROCESS_DWH_SCHEMA:-public}"
user: "${THT_PREPROCESS_DWH_USER:-unused}"
password_file: "${THT_PREPROCESS_DWH_PASSWORD_FILE:-/run/secrets/preprocess-dwh-password}"
vectors:
type: qdrant
base_url: http://qdrant:6333
collection: preprocess-evidence
roots: {artifacts: artifacts, indexes: indexes, sessions: sessions}
evidence: {source_root: /data/source, evidence_dir: evidence}
embeddings:
provider: ollama_internal
base_url: http://embedding:11434
model: qwen3-embedding:0.6b
dim: 1024
batch_size: 32
+5 -2
View File
@@ -1,6 +1,9 @@
# Panoramica dell'architettura
> Sintesi ad uso documentazione. Per il dettaglio storico delle decisioni di design vedi le [Specifiche di Design](../superpowers/specs/2026-06-25-thothii-architecture-design.md) e i [Piani di Implementazione](../superpowers/plans/2026-06-25-harness-implementation.md). Per lo stato corrente del progetto (gate manuali pendenti, layout workspace/secret) vedi `PROJECT_STATE.md` nella radice del repo.
> Sintesi ad uso documentazione. Per il dettaglio dei moduli e dei flussi vedi
> [Componenti, moduli e flussi](components.md); i contratti correnti sono in `docs/contracts/`
> e le decisioni durevoli in `docs/adr/`. Per lo stato corrente del progetto (gate manuali
> pendenti, layout workspace/secret) vedi `PROJECT_STATE.md` nella radice del repo.
ThothII è un **datamart builder human-in-the-loop**: trasforma una domanda in linguaggio naturale in SQL validato (ed eventualmente un datamart dbt) attraverso un **workflow deterministico a 8 fasi NL→SQL**, in cui il modello *propone* e un revisore umano *decide* ai gate.
@@ -76,4 +79,4 @@ Lo stack locale si avvia con `./scripts/run-stack.sh`, dopo aver creato
`deploy/env/local.env` da `deploy/env/local.env.example`. Il core Compose include Pi; DWH,
vector DB, embedding e LLM sono endpoint esterni configurati nel file locale.
Comandi per singolo layer, test, lint: vedi il file `CLAUDE.md` nella radice del repo (guida operativa per Claude Code, tenuta sincronizzata con questa pagina).
Comandi per singolo layer, test e lint: vedi `AGENTS.md` nella radice del repo.
+6 -11
View File
@@ -55,21 +55,16 @@ Una CA privata PEM resta esterna al bundle e va montata con un override Compose
## Preprocessing
I job di preprocessing usano gli stessi servizi interni Qdrant/Ollama:
Il preprocessing passa dal CLI host nativo e dal descrittore dell'installazione:
```sh
docker compose --env-file deploy/env/local.env \
-f compose.yaml -f deploy/compose.local.yaml \
-f deploy/compose.preprocess.yaml --profile preprocess run --rm preprocess-evidence
tht --installation /percorso/assoluto/thothii-installation.yaml workspace preprocess evidence
tht --installation /percorso/assoluto/thothii-installation.yaml workspace preprocess dwh
```
Per introspezione DWH:
```sh
docker compose --env-file deploy/env/local.env \
-f compose.yaml -f deploy/compose.local.yaml \
-f deploy/compose.preprocess.yaml --profile preprocess run --rm preprocess-dwh
```
Il CLI esegue il servizio profile-gated `workspace-maintenance`. I dettagli sono nel
[contratto del preprocessing](contracts/workspace-preprocessing-cli.md) e nella guida
[Evidence](evidence.md).
## Server
+1 -1
View File
@@ -1,6 +1,6 @@
# PSD — rollout controllato DWH REST
Questo runbook rispecchia i Task 9–10 del [piano](../superpowers/plans/2026-08-20-dwh-rest-per-installation-auth.md). Gate A e la parte dual-key di Gate B sono stati eseguiti con autorizzazioni separate. L'emendamento del proprietario del 2026-08-21 rinvia collaudo Mac, osservazione e revoca a prima di Project B; non autorizza ulteriori mutazioni.
Questo runbook completa il [piano di accettazione autenticazione](../plans/2026-08-18-thothii-authentication-acceptance-and-psd-deployment.md) e il [programma di deployment PSD](../plans/2026-08-20-psd-server-deployment-program.md). Gate A e la parte dual-key di Gate B sono stati eseguiti con autorizzazioni separate. L'emendamento del proprietario del 2026-08-21 rinvia collaudo Mac, osservazione e revoca a prima di Project B; non autorizza ulteriori mutazioni.
## Invarianti
@@ -11,7 +11,7 @@ Fonti normative:
- `docs/operations/psd-server-sol-orchestration-prompt.md`
- `docs/plans/2026-08-20-psd-server-deployment-program.md`
- `docs/plans/2026-08-20-psd-server-survey.md`
- `docs/superpowers/specs/2026-08-20-psd-survey-remediation-checklist-design.md`
- `docs/plans/2026-08-20-psd-server-deployment-program-design.md`
Questo documento non autorizza modifiche a server, servizi, database, Nginx, load balancer,
Authentik, Aritmolab o repository esterni.
@@ -149,9 +149,9 @@ Authentik, Aritmolab o repository esterni.
- il proprietario conferma che le sessioni, gli indici e le cache del vecchio ThothII erano solo
test e non richiedono migrazione o attività di salvaguardia dedicate. Restano necessari il
rollback della route DWH condivisa e il rispetto del gate esplicito prima di fermare lo stack;
- il proprietario ha revisionato e approvato la specifica scritta. Il piano eseguibile è in
`docs/superpowers/plans/2026-08-20-dwh-rest-per-installation-auth.md`; separa sviluppo e
verifica del componente dai due gate espliciti di mutazione PSD;
- il proprietario ha revisionato e approvato la specifica scritta. Il runbook eseguibile è in
`docs/operations/psd-dwh-auth-rollout.md`; separa sviluppo e verifica del componente dai due
gate espliciti di mutazione PSD;
- la credenziale condivisa corrente, priva del nuovo identificativo pubblico, sarà l’unico record
temporaneo `legacy_raw` con ID `legacy-shared`. Dopo la revoca non saranno accettate chiavi
prive del formato versionato per installazione.
@@ -1,19 +0,0 @@
# Button Press Feedback Design
## Context
Buttons currently change on hover, but many provide little or no visible acknowledgment while the pointer is pressed. The interface mixes a shared Base UI button with native buttons, so changing only the shared component would leave inconsistent behavior.
## Chosen interaction
Apply one CSS press vocabulary to every enabled native button. During `:active`, the button compresses to `scale(0.97)`, loses raised shadow, and receives a restrained brightness change. The transition lasts 140 ms and uses an ease-out-quint curve (`cubic-bezier(0.22, 1, 0.36, 1)`). This reads as a physical press without bounce, ripple, layout movement, or JavaScript state.
The existing one-pixel translation on the shared button is removed so shared and native buttons do not combine two motion patterns.
## Accessibility
Under `prefers-reduced-motion: reduce`, scale is disabled. The brightness and shadow change remain, providing a clear pressed state without kinetic motion. Disabled buttons receive no press treatment.
## Verification
A focused contract test checks the global enabled-button selector, scale, timing, easing, and reduced-motion override. The full frontend suite and typecheck guard regressions. Playwright then holds a real button in the active state and confirms its computed transform, followed by a visual screenshot/snapshot check.
@@ -1,72 +0,0 @@
# Button Press Feedback Implementation Plan
> **For Claude:** REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
**Goal:** Give every enabled button immediate, consistent click acknowledgment while preserving a non-kinetic reduced-motion alternative.
**Architecture:** Define the interaction once in the global Tailwind base layer so both Base UI and native buttons inherit it. Remove the shared button's older translation-only active state to avoid compounded transforms. Verify the CSS contract first, then exercise the real interaction in Playwright.
**Tech Stack:** React 18, Tailwind CSS 3, Vitest, Playwright CLI.
---
### Task 1: Specify the global press contract
**Files:**
- Create: `frontend/src/button-press-feedback.test.ts`
- Test: `frontend/src/button-press-feedback.test.ts`
**Step 1: Write the failing test**
Read `src/index.css` and assert the enabled-button active selector, `scale(0.97)`, 140 ms duration, ease-out-quint curve, disabled exclusion, and reduced-motion transform override. Assert that `components/ui/button.tsx` no longer contains the legacy translation active class.
**Step 2: Run test to verify it fails**
Run: `npx vitest run src/button-press-feedback.test.ts`
Expected: FAIL because the global press rules do not exist and the shared button still uses translation.
### Task 2: Implement the press feedback
**Files:**
- Modify: `frontend/src/index.css`
- Modify: `frontend/src/components/ui/button.tsx`
- Test: `frontend/src/button-press-feedback.test.ts`
**Step 1: Add the minimal CSS**
Add a global enabled-button transition and active state using only transform, filter, and shadow. Add a `prefers-reduced-motion` override that removes scale while preserving non-kinetic contrast feedback.
**Step 2: Remove the legacy shared-button translation**
Delete `active:not-aria-[haspopup]:translate-y-px` from the shared variant base string.
**Step 3: Run the focused test**
Run: `npx vitest run src/button-press-feedback.test.ts`
Expected: PASS.
### Task 3: Verify regressions and real-browser behavior
**Files:**
- Verify: `frontend/src/index.css`
- Verify: `frontend/src/components/ui/button.tsx`
**Step 1: Run frontend verification**
Run: `npx vitest run`
Run: `npx tsc -b`
Expected: all tests pass and typecheck exits 0.
**Step 2: Verify in Playwright**
Open `http://localhost:5173`, hold pointer-down on an enabled button, and inspect its computed transform and filter before release. Repeat with reduced motion emulation and confirm transform remains `none` while contrast feedback remains.
**Step 3: Review the final diff**
Run: `git diff --check` and inspect `git diff --stat`.
Expected: no whitespace errors and only the intended product/design, CSS, component, and test files changed.
@@ -1,773 +0,0 @@
# Internal Qdrant and Ollama Implementation Plan
> **Historical nomenclature:** this plan predates the native host CLI convergence. The current
> operator command is `tht`; any older `thothctl` smoke-script or rollback wording below is retained
> only as historical evidence.
> **For Claude:** REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
**Goal:** Make Qdrant and Ollama mandatory internal ThothII services while keeping the analytical
DWH external and associating each workspace with one Qdrant collection for schema, Evidence, and
Memory embeddings.
**Architecture:** Introduce workspace schema v3, preserve v1/v2 only as migration inputs, and keep
the existing harness `VectorStore` port behind a new Qdrant REST adapter. Base Compose owns Qdrant,
Ollama, their persistent volumes, and model initialization; workspace descriptors contain semantic
identity but no vector/embedding endpoints or credentials.
**Tech Stack:** TypeScript/Fastify/Zod, Python 3.12/Pydantic/requests, React 18, Docker Compose,
Qdrant REST API, Ollama `/api/embed`, Vitest, pytest.
---
## Guardrails
- Apply `@superpowers:test-driven-development` to every behavior change: add one focused failing
test, observe the expected failure, implement the minimum, and rerun the focused test.
- Do not run broad suites until the corresponding code/config changes exist; this preserves the
requested ordering while still using TDD.
- Preserve the external DWH connector contract and session persistence model.
- Do not retain an operational fallback to pgvector or an external embedding endpoint.
- Do not delete or rewrite user workspace repositories or Qdrant data. Migration is descriptor-only;
semantic data is rebuilt explicitly.
- Commit after each task only when focused tests are green.
### Task 1: Define workspace schema v3
**Files:**
- Modify: `backend/src/workspaces/schema.ts`
- Modify: `backend/src/workspaces/types.ts`
- Modify: `backend/test/workspaces-schema.test.ts`
- Modify: `backend/test/workspaces-migrate-legacy.test.ts`
- Create: `backend/src/workspaces/migrate-v2-qdrant.ts`
- Create: `backend/test/workspaces-migrate-v2-qdrant.test.ts`
**Step 1: Write the failing schema tests**
Add tests proving that schema v3 accepts only this semantic shape:
```ts
const semantic_index = {
vector_store: {
engine: "qdrant",
collection: "psd-clinical",
dimensions: 1024,
distance: "cosine",
},
embedding: {
provider: "ollama_internal",
model: "qwen3-embedding:0.6b",
dimensions: 1024,
},
};
```
Add separate rejection cases for `pgvector`, `supported_transports`, external embedding providers,
non-1024 dimensions, non-cosine distance, and unknown fields. Assert v1/v2 remain parseable as
legacy descriptors but `isOperationalWorkspace()` returns false.
**Step 2: Run the tests and verify RED**
Run:
```bash
cd backend
npx vitest run test/workspaces-schema.test.ts test/workspaces-migrate-v2-qdrant.test.ts
```
Expected: failure because schema version 3 and `migrateWorkspaceV2ToV3` do not exist.
**Step 3: Implement the minimum schema and migration**
Add `QdrantVectorStore`, `InternalEmbedding`, and `WorkspaceV3` types. Replace the operational type
guard with schema-v3-only semantics. Implement:
```ts
export function migrateWorkspaceV2ToV3(
legacy: WorkspaceV2,
collection: string,
): WorkspaceV3 {
return validateOperationalWorkspace({
workspace: { ...legacy.workspace, schema_version: 3 },
dwh: legacy.dwh,
semantic_index: {
vector_store: {
engine: "qdrant",
collection,
dimensions: 1024,
distance: "cosine",
},
embedding: {
provider: "ollama_internal",
model: "qwen3-embedding:0.6b",
dimensions: 1024,
},
},
llm_policy: legacy.llm_policy,
...(legacy.diagnostics?.dwh_rest
? { diagnostics: { dwh_rest: legacy.diagnostics.dwh_rest } }
: {}),
});
}
```
Do not copy vector/embedding diagnostics or transports.
**Step 4: Verify GREEN**
Run the command from Step 2. Expected: all selected tests pass.
**Step 5: Commit**
```bash
git add backend/src/workspaces/schema.ts backend/src/workspaces/types.ts \
backend/src/workspaces/migrate-v2-qdrant.ts backend/test/workspaces-schema.test.ts \
backend/test/workspaces-migrate-legacy.test.ts backend/test/workspaces-migrate-v2-qdrant.test.ts
git commit -m "feat: define internal semantic workspace schema"
```
### Task 2: Make collection ownership unique in the Git registry
**Files:**
- Modify: `backend/src/workspaces/registry.ts`
- Modify: `backend/src/workspaces/migrate-legacy.ts`
- Modify: `backend/test/workspace-registry.test.ts`
- Modify: `backend/test/workspaces-migrate-legacy.test.ts`
**Step 1: Write failing registry tests**
Add fixtures with two schema-v3 workspaces claiming `collection: shared`. Assert snapshot activation
fails with `workspace_invalid` and retains the previous active snapshot. Assert v1/v2 entries are
listed as `migration_required` and cannot be acquired with `acquireSessionRevision()`.
**Step 2: Verify RED**
```bash
cd backend
npx vitest run test/workspace-registry.test.ts test/workspaces-migrate-legacy.test.ts \
-t "collection|migration_required"
```
Expected: duplicate collections are currently accepted and v2 is currently operational.
**Step 3: Implement uniqueness and migration state**
During snapshot validation, build `Map<collection, workspaceId>` for operational descriptors and
raise a sanitized `workspace_invalid` error on a duplicate. Update migration output and CLI wording
to require an explicit target collection and schema v3.
**Step 4: Verify GREEN and commit**
```bash
cd backend
npx vitest run test/workspace-registry.test.ts test/workspaces-migrate-legacy.test.ts \
-t "collection|migration_required"
cd ..
git add backend/src/workspaces/registry.ts backend/src/workspaces/migrate-legacy.ts \
backend/test/workspace-registry.test.ts backend/test/workspaces-migrate-legacy.test.ts
git commit -m "feat: reserve one qdrant collection per workspace"
```
### Task 3: Remove external semantic bindings and render internal endpoints
**Files:**
- Modify: `backend/src/workspaces/contracts.ts`
- Modify: `backend/src/workspaces/bindings.ts`
- Modify: `backend/src/workspaces/runtime-renderer.ts`
- Modify: `backend/src/config.ts`
- Modify: `backend/test/workspaces-contracts.test.ts`
- Modify: `backend/test/workspaces-bindings.test.ts`
- Modify: `backend/test/workspace-runtime-renderer.test.ts`
- Modify: `backend/test/config.test.ts`
**Step 1: Write failing contract tests**
Assert schema-v3 installation contracts contain DWH variables only. Assert environment variables
matching `*_VECTOR_*`, `*_EMBEDDING_BASE_URL`, or semantic API-key suffixes are ignored/rejected.
Assert the rendered harness config always contains:
```yaml
resources:
vector:
engine: qdrant
base_url: http://qdrant:6333
collection: psd-clinical
embeddings:
provider: ollama_internal
base_url: http://embedding:11434
model: qwen3-embedding:0.6b
dimensions: 1024
```
**Step 2: Verify RED**
```bash
cd backend
npx vitest run test/workspaces-contracts.test.ts test/workspaces-bindings.test.ts \
test/workspace-runtime-renderer.test.ts test/config.test.ts
```
Expected: current contracts require external vector and embedding bindings.
**Step 3: Implement internal runtime configuration**
Add typed backend config fields with Compose defaults:
```ts
internalQdrantUrl: "http://qdrant:6333"
internalEmbeddingUrl: "http://embedding:11434"
internalEmbeddingModel: "qwen3-embedding:0.6b"
internalEmbeddingDimensions: 1024
```
Accept only `qdrant`, `embedding`, `localhost`, or loopback hosts. Keep these values out of Git
workspace descriptors, API payloads, and generated installation docs. Render them into the
ephemeral backend-owned harness config after descriptor validation.
**Step 4: Verify GREEN and commit**
Run Step 2, then:
```bash
git add backend/src/config.ts backend/src/workspaces/contracts.ts backend/src/workspaces/bindings.ts \
backend/src/workspaces/runtime-renderer.ts backend/test/config.test.ts \
backend/test/workspaces-contracts.test.ts backend/test/workspaces-bindings.test.ts \
backend/test/workspace-runtime-renderer.test.ts
git commit -m "feat: render private semantic service endpoints"
```
### Task 4: Narrow harness embedding configuration to internal Ollama
**Files:**
- Modify: `harness/tht/config.py`
- Modify: `harness/tht/config_compat.py`
- Modify: `harness/tht/vectorstore/embeddings.py`
- Modify: `harness/tht/cli/ollama_cmd.py`
- Modify: `harness/tests/test_config_resources.py`
- Create: `harness/tests/test_internal_embeddings.py`
**Step 1: Write failing embedding tests**
Use a fake `requests.Session` to prove `OllamaInternalEmbeddings.embed()` calls `/api/embed` with
model and batch input, returns 1024-dimensional finite vectors, and rejects count/dimension/NaN
mismatches. Add config tests rejecting external providers, API keys, and non-private base URLs.
**Step 2: Verify RED**
```bash
cd harness
.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py -q
```
Expected: `OllamaInternalEmbeddings` and internal-only config do not exist.
**Step 3: Implement the client**
Implement one bounded `/api/embed` request per batch:
```python
response = self._session.post(
f"{self.base_url}/api/embed",
json={"model": self.model, "input": texts},
timeout=self.timeout,
)
```
Validate response shape before returning any vector. Keep retry behavior bounded and sanitize URLs
and response bodies from raised errors.
**Step 4: Verify GREEN and commit**
```bash
cd harness
.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py -q
cd ..
git add harness/tht/config.py harness/tht/config_compat.py harness/tht/vectorstore/embeddings.py \
harness/tht/cli/ollama_cmd.py harness/tests/test_config_resources.py \
harness/tests/test_internal_embeddings.py
git commit -m "feat: use internal ollama embeddings"
```
### Task 5: Implement the Qdrant VectorStore adapter
**Files:**
- Create: `harness/tht/adapters/vector/qdrant.py`
- Modify: `harness/tht/adapters/vector/__init__.py`
- Modify: `harness/tht/ports/vector.py`
- Modify: `harness/tht/vectorstore/records.py`
- Modify: `harness/tht/vectorstore/store.py`
- Create: `harness/tests/test_qdrant_vector_store.py`
- Modify: `harness/tests/test_vector_port_contract.py`
**Step 1: Write failing adapter tests**
Test a real adapter against a deterministic fake HTTP server. Cover:
- idempotent collection create with 1024/Cosine;
- mismatch fails without delete/recreate;
- keyword payload-index creation;
- deterministic UUIDv5 point IDs;
- upsert payload for `schema`, `evidence`, and `memory`;
- query filtered by workspace and allowed kinds;
- `existing_hashes`, exact Evidence generation list/delete, and health;
- sanitized timeouts and malformed responses.
The point ID helper must satisfy:
```python
def point_id(workspace_id: str, kind: str, record_key: str) -> str:
return str(uuid5(NAMESPACE_URL, f"thothii:{workspace_id}:{kind}:{record_key}"))
```
**Step 2: Verify RED**
```bash
cd harness
.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
```
Expected: import failure for the Qdrant adapter.
**Step 3: Implement minimal REST mappings**
Use existing `requests` dependency and these endpoints:
```text
GET /collections/{collection}
PUT /collections/{collection}
PUT /collections/{collection}/index
PUT /collections/{collection}/points?wait=true
POST /collections/{collection}/points/query
POST /collections/{collection}/points/scroll
POST /collections/{collection}/points/delete?wait=true
```
Every operation must include the workspace filter even though the collection is workspace-owned.
Map Qdrant scores and payloads back into existing `VectorHit` objects.
**Step 4: Verify GREEN and commit**
```bash
cd harness
.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
cd ..
git add harness/tht/adapters/vector/qdrant.py harness/tht/adapters/vector/__init__.py \
harness/tht/ports/vector.py harness/tht/vectorstore/records.py \
harness/tht/vectorstore/store.py harness/tests/test_qdrant_vector_store.py \
harness/tests/test_vector_port_contract.py
git commit -m "feat: add qdrant vector adapter"
```
### Task 6: Wire schema, Evidence, and Memory through Qdrant
**Files:**
- Modify: `harness/tht/vectorstore/reader.py`
- Modify: `harness/tht/cli/vector_cmd.py`
- Modify: `harness/tht/cli/memory_cmd.py`
- Modify: `harness/tht/corpus/pipeline.py`
- Modify: `harness/tht/search/evidence.py`
- Modify: `harness/tht/cli/schema_cmd.py`
- Modify: `harness/tests/test_memory_save_one.py`
- Modify: `harness/tests/test_search_pack.py`
- Create: `harness/tests/test_semantic_kind_isolation.py`
**Step 1: Write failing integration tests**
Use an in-memory fake implementing the `VectorStore` port. Assert:
- schema records use `kind=schema`;
- corpus records use `kind=evidence` and exact generation;
- approved memories use `kind=memory`;
- search pack requests only its allowed kind set;
- all three paths share `workspace_id`, `workspace_revision`, hashing, and point-key construction;
- retries do not duplicate points.
**Step 2: Verify RED**
```bash
cd harness
.venv/bin/pytest tests/test_semantic_kind_isolation.py tests/test_memory_save_one.py \
tests/test_search_pack.py -q
```
Expected: current factories select pgvector/HTTP adapters and payloads lack the v3 identity fields.
**Step 3: Wire the adapter**
Make schema-v3 `qdrant` the only operational vector factory branch. Reuse the current canonical
record builders; add only missing identity fields. Keep the JSONL Memory registry and filesystem
Evidence corpus as sources of truth.
**Step 4: Verify GREEN and commit**
Run Step 2, then commit the listed files with:
```bash
git commit -m "feat: index semantic records in qdrant"
```
### Task 7: Add mandatory Qdrant and Ollama Compose services
**Files:**
- Modify: `compose.yaml`
- Create: `deploy/compose.embedding-gpu.yaml`
- Create: `docker/embedding-model-init.sh`
- Modify: `docker/core.Dockerfile`
- Modify: `deploy/env/local.env.example`
- Modify: `deploy/env/server.env.example`
- Modify: `scripts/run-stack.sh`
- Modify: `scripts/test-default-compose.sh`
- Modify: `scripts/test-unified-compose.sh`
- Create: `scripts/test-internal-semantic-compose.sh`
**Step 1: Write failing Compose contract tests**
Assert the rendered base profile has `core`, `frontend`, `qdrant`, `embedding`, and
`embedding-model-init`; private services have no published ports; persistent volumes exist; core
depends on Qdrant health and successful model init; no external vector/embedding binding is required.
Also assert all service images use version plus immutable digest. Resolve and record supported
multi-architecture digests for Qdrant v1.18.x and Ollama v0.32.x during implementation:
```bash
docker buildx imagetools inspect qdrant/qdrant:v1.18.2
docker buildx imagetools inspect ollama/ollama:0.32.0
```
**Step 2: Verify RED**
```bash
./scripts/test-default-compose.sh
./scripts/test-unified-compose.sh
./scripts/test-internal-semantic-compose.sh
```
Expected: required services and volumes are absent.
**Step 3: Implement the services**
`embedding-model-init.sh` must wait with a bounded deadline, call `ollama pull` for the exact model,
and verify it appears in `/api/tags`. The Qdrant healthcheck uses its HTTP health endpoint. The CPU
base has no device reservation; the GPU override adds only the supported device stanza.
**Step 4: Verify GREEN and commit**
Run Step 2, then:
```bash
git add compose.yaml deploy/compose.embedding-gpu.yaml docker/embedding-model-init.sh \
docker/core.Dockerfile deploy/env/local.env.example deploy/env/server.env.example \
scripts/run-stack.sh scripts/test-default-compose.sh scripts/test-unified-compose.sh \
scripts/test-internal-semantic-compose.sh
git commit -m "feat: run qdrant and ollama inside thothii"
```
### Task 8: Retire pgvector deployment and external semantic connectors
**Files:**
- Delete: `deploy/compose.local-vector.yaml`
- Delete: `deploy/compose.preprocess-local-vector.yaml`
- Delete: `deploy/sql/20-vector-roles.sql`
- Delete: `deploy/vector/reconcile-roles.sh`
- Delete: `deploy/vector/rotate-bootstrap-password.py`
- Delete: `deploy/vector/secret-policy.sh`
- Delete: `deploy/vector/vector-db-entrypoint.sh`
- Delete: `scripts/local-vector-smoke.sh`
- Delete: `scripts/test-local-vector-smoke-safety.sh`
- Delete: `scripts/test-local-vector-smoke-live-collision.sh`
- Delete: `scripts/test-vector-bootstrap-rotation.sh`
- Delete: `scripts/test-vector-migration-image.sh`
- Delete: `scripts/test-vector-secret-policy.sh`
- Modify: `scripts/test-no-deployment-coupling.sh`
- Modify: `scripts/test-no-deployment-coupling-scope.sh`
- Modify: `scripts/test-compose-secret-policy.sh`
- Modify: `.github/workflows/deployment.yml`
**Step 1: Write the failing coupling test**
Teach the coupling gate to reject active `pgvector`, `local-vector`, `THT_VECTOR_*`, workspace
embedding URLs/API keys, and external vector transports while allowing historical specs and the
explicit descriptor migration module.
**Step 2: Verify RED**
```bash
./scripts/test-no-deployment-coupling-scope.sh
./scripts/test-no-deployment-coupling.sh
./scripts/test-compose-secret-policy.sh
```
Expected: active pgvector deployment paths are reported.
**Step 3: Remove the retired paths and update CI**
Remove only repository deployment machinery. Retain harness pgvector code temporarily only if it
is needed to read/export legacy data during migration; it must not be reachable from schema v3 or
Compose. Remove it in a follow-up task once migration fixtures no longer import it.
**Step 4: Verify GREEN and commit**
Run Step 2 and the workflow fixture tests, then commit all deletions and modifications:
```bash
git add -A deploy scripts .github/workflows/deployment.yml
git commit -m "refactor: retire external vector deployment"
```
### Task 9: Update frontend workspace editing and examples
**Files:**
- Modify: `frontend/src/api/workspaces.ts`
- Modify: `frontend/src/shell/WorkspaceEditor.tsx`
- Modify: `frontend/src/shell/WorkspaceEditor.test.tsx`
- Modify: `frontend/src/shell/WorkspaceManager.test.tsx`
- Modify: `frontend/src/api/workspaces.test.ts`
- Modify: `frontend/src/workspaces/drafts.test.ts`
- Modify: `deploy/workspaces/example.yaml`
- Modify: `deploy/workspaces/psd.yaml.example`
**Step 1: Write failing UI tests**
Assert editor/preview show Qdrant collection and fixed internal embedding model, expose no vector
endpoint/credential fields, and publish schema v3. Assert legacy descriptors display a migration
banner and cannot be selected for a new session.
**Step 2: Verify RED**
```bash
cd frontend
npx vitest run src/shell/WorkspaceEditor.test.tsx src/shell/WorkspaceManager.test.tsx \
src/api/workspaces.test.ts src/workspaces/drafts.test.ts
```
Expected: fixtures and controls still use pgvector/external embedding.
**Step 3: Implement fixed semantic controls**
Collection remains editable and validated. Engine, provider, model, dimensions, and distance render
as fixed architecture values. Remove external semantic diagnostics from drafts and publish payloads.
**Step 4: Verify GREEN and commit**
Run Step 2, then commit the listed files with:
```bash
git commit -m "feat: edit qdrant workspace collections"
```
### Task 10: Add a real internal semantic smoke
**Files:**
- Create: `scripts/internal-semantic-smoke.sh`
- Modify: `scripts/unified-deployment-smoke.sh`
- Modify: `scripts/server-deployment-smoke.sh`
- Modify: `scripts/task13-runtime-fixture-check.ts`
- Modify: `scripts/test-task13-runtime-fixtures.sh`
**Step 1: Write failing smoke fixture assertions**
The fixture must require private Qdrant/Ollama services, model volume, Qdrant volume, fixed internal
URLs, and no host ports. It must reject wrong service names, external URLs, collection reuse, and
dimension changes.
**Step 2: Verify RED**
```bash
./scripts/test-task13-runtime-fixtures.sh local
./scripts/test-task13-runtime-fixtures.sh server
```
Expected: current fixture expects the two-service topology.
**Step 3: Implement the live smoke**
Using disposable volumes and a fixture workspace, start the stack on CPU, wait for the model, ensure
the collection, embed one record of each kind, query each kind with filters, restart offline, and
prove all points and the model remain available. Cleanup must remain exact and must not prune global
Docker resources.
**Step 4: Verify GREEN and commit**
```bash
./scripts/test-task13-runtime-fixtures.sh local
./scripts/test-task13-runtime-fixtures.sh server
./scripts/internal-semantic-smoke.sh
git add scripts/internal-semantic-smoke.sh scripts/unified-deployment-smoke.sh \
scripts/server-deployment-smoke.sh scripts/task13-runtime-fixture-check.ts \
scripts/test-task13-runtime-fixtures.sh
git commit -m "test: cover internal semantic services"
```
### Task 11: Update operator documentation and state
**Files:**
- Modify: `README.md`
- Modify: `AGENTS.md`
- Modify: `PROJECT_STATE.md`
- Modify: `docs/install/local-workspace-registry.md`
- Modify: `docs/install/server-workspace-registry.md`
- Modify: `docs/installazione-docker-4-contesti.md`
- Modify: `docs/workspace-diagnostic-protocol.md`
- Modify: `docs/gestione-memory.md`
- Modify: `deploy/secrets/README.md`
- Modify: `scripts/verify-workspace-install-docs.sh`
- Modify: `scripts/test-verify-workspace-install-docs.sh`
**Step 1: Write failing documentation contract assertions**
Require the four-service topology, CPU/GPU behavior, volume backup/restore, schema-v3 migration,
Qdrant collection ownership, and removal of external vector/embedding variables from active manuals.
**Step 2: Verify RED**
```bash
./scripts/test-verify-workspace-install-docs.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
```
Expected: manuals still describe external pgvector/embedding and a two-service mandatory stack.
**Step 3: Update documentation**
Document Qdrant as a derived but persistent index, Ollama model cache behavior, CPU-first startup,
optional GPU override, snapshot/restore, explicit legacy migration, and the fact that only the DWH
and LLM remain external application endpoints.
**Step 4: Verify GREEN and commit**
Run Step 2, then:
```bash
git add README.md AGENTS.md PROJECT_STATE.md docs deploy/secrets/README.md \
scripts/verify-workspace-install-docs.sh scripts/test-verify-workspace-install-docs.sh
git commit -m "docs: document internal semantic infrastructure"
```
### Task 12: Remove unreachable pgvector runtime code
**Files:**
- Delete: `harness/tht/adapters/vector/pgvector.py`
- Delete: `harness/tht/adapters/vector/legacy_direct.py`
- Delete: `harness/tht/adapters/vector/thoth_http.py`
- Delete: `harness/tht/vectorstore/rest_client.py`
- Delete: `harness/tht/vectorstore/rest_writer.py`
- Delete: `harness/tht/migrations/vector/001_extensions.sql`
- Delete: `harness/tht/migrations/vector/002_schema_tables.sql`
- Delete: `harness/tht/migrations/vector/003_roles.sql`
- Delete: `harness/tht/migrations/vector/004_evidence_generation_gc.sql`
- Modify: `harness/pyproject.toml`
- Modify/Delete: affected pgvector and migration tests under `harness/tests/l0/`
**Step 1: Prove the code is unreachable**
```bash
rg -n "PgVectorStore|ThothHttpVectorStore|LegacyDirectVectorStore|migrations/vector" \
harness backend frontend compose.yaml deploy scripts docker docs \
--glob '!docs/plans/**' --glob '!docs/superpowers/**'
```
Expected before cleanup: matches only in the files scheduled for deletion and legacy tests. If an
operational call site remains, stop and migrate it before deleting anything.
**Step 2: Delete obsolete runtime and tests**
Retain descriptor migration tests, but remove PostgreSQL vector runtime/migration packaging tests.
Remove `psycopg2-binary` only if the DWH/session PostgreSQL paths do not need it; otherwise keep it.
**Step 3: Verify focused imports and packaging**
```bash
cd harness
.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py \
tests/test_semantic_kind_isolation.py tests/test_vector_migration_packaging.py -q
python -m build
```
Expected: Qdrant tests pass and the wheel contains no pgvector migrations. Adjust the packaging test
to assert Qdrant has no SQL migration payload.
**Step 4: Commit**
```bash
git add -A harness
git commit -m "refactor: remove pgvector runtime"
```
### Task 13: Run complete verification
**Files:**
- Modify only if a genuine regression is discovered.
**Step 1: Deterministic layer gates**
```bash
cd harness && .venv/bin/pytest -q && .venv/bin/ruff check .
cd ../backend && npx vitest run && npx tsc --noEmit -p . && npm run build
cd ../frontend && npx vitest run && npx tsc -b && npm run build
cd .. && git diff --check
```
Expected: all gates pass. Existing unrelated Ruff debt must be reported separately if it remains;
new/modified files must be Ruff-clean.
**Step 2: Deployment contracts**
```bash
./scripts/test-default-compose.sh
./scripts/test-unified-compose.sh
./scripts/test-internal-semantic-compose.sh
./scripts/test-no-deployment-coupling.sh
./scripts/test-compose-secret-policy.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
```
Expected: all pass without external vector/embedding settings.
**Step 3: Docker smokes**
```bash
./scripts/internal-semantic-smoke.sh
./scripts/workspace-registry-smoke.sh
./scripts/unified-deployment-smoke.sh
./scripts/thothctl-update-smoke.sh
./scripts/server-deployment-smoke.sh
```
Expected: CPU semantic smoke passes, persistence survives offline restart, and every script proves
exact cleanup. Investigate the previously observed `thothctl` rollback failure independently if it
recurs; do not weaken the new semantic gate to hide it.
**Step 4: Final audit**
```bash
rg -n "pgvector|local-vector|THT_VECTOR_|EMBEDDING_BASE_URL|openai_compatible|ollama_compatible" \
. --glob '!docs/plans/**' --glob '!docs/superpowers/**' --glob '!**/node_modules/**' \
--glob '!**/.venv/**' --glob '!**/.git/**'
git status --short
```
Expected: no active operational references; only explicit legacy descriptor migration fixtures may
remain. Worktree contains only intentional changes.
**Step 5: Commit verification metadata**
Update `PROJECT_STATE.md` with exact counts, image digests, smoke durations, CPU hardware, and any
manual GPU/Windows gates. Commit only verified claims:
```bash
git add PROJECT_STATE.md
git commit -m "docs: record qdrant ollama verification"
```
@@ -1,447 +0,0 @@
# Piano di implementazione: workspace descriptor esclusivamente schema v3
> **Per gli agenti esecutori:** SUB-SKILL OBBLIGATORIA: usare `superpowers:subagent-driven-development` (raccomandata) oppure `superpowers:executing-plans`, procedendo task per task con TDD e review tra i task.
**Obiettivo:** rimuovere dal prodotto ogni capacità di leggere, migrare, rendere operativo o presentare workspace descriptor schema v1/v2. Il solo descriptor accettato diventa schema v3. Restano intatti i formati versionati non correlati e gli state file del registry già prodotti da versioni recenti con revisioni v3.
**Architettura:** parser, registry, renderer, diagnostica, route e frontend convergono su un solo tipo `WorkspaceV3`. Il campo pubblico `WorkspaceRevision.state` scompare. Un decoder privato normalizza in memoria gli state file già scritti con `state: "operational"`, elimina quel campo prima di qualsiasi uso/API e rifiuta ogni combinazione non-v3 o incoerente. I build backend diventano clean-first, così la cancellazione dei migratori sorgente implica anche la loro assenza da `dist` e dall'immagine core.
**Tech stack:** TypeScript 5, Zod 4, Fastify 5, React 18, Vitest, Node.js 22, Bash/PowerShell, Git e Docker Compose.
**Stato:** piano revisionato dopo review indipendente. La sua approvazione non autorizza l'implementazione; attendere un esplicito ordine separato.
---
## Decisioni confermate
1. Nessun workspace v1/v2 reale deve essere preservato o migrato.
2. Eliminare `migrate-legacy.ts`, `migrate-v2-qdrant.ts` e le relative interfacce CLI.
3. Eliminare il campo `state` dal tipo/API `WorkspaceRevision` e da tutti i nuovi state/manifest del registry.
4. Descriptor v1/v2 presenti in Git o negli snapshot vengono rifiutati, senza conversione automatica.
5. Non toccare i documenti storici sotto `docs/superpowers/` e i vecchi piani; possono descrivere decisioni passate.
6. Non iniziare P2 finché P1 non dispone di nuova evidenza automatica e di una nuova decisione manuale esplicita.
## Confini da non oltrepassare
Questa rimozione riguarda soltanto il **workspace descriptor**. Non eliminare o rinominare:
- `schemaVersion`/`schema_version` di bundle ZIP, report, job, ledger, manifest di sessione o artifact di fase;
- `RevisionLeaseRecord.state` (`creating`/`persisted`), maintenance state, process state o UI state non collegati a `WorkspaceRevision`;
- `migration_required` usato nei futuri piani P3–P6 per ownership DWH, punti semantici revisionless o altre migrazioni non-descriptor;
- `allowLegacy` del frontend sessioni, che significa “sessione senza revisione workspace” e non descriptor v1/v2;
- documenti storici o report conservati.
L'unica compatibilità legacy mantenuta nel codice è il decoder privato degli state file già scritti con il campo revisionale `state: "operational"`. Non costituisce supporto a descriptor v1/v2.
## Contratto v3-only
- `WorkspaceDescriptor`, `CanonicalWorkspace` e `WorkspaceV3` rappresentano la stessa forma v3; mantenere gli alias soltanto quando migliorano la semantica dei confini.
- `parseWorkspaceYaml` e `validateWorkspaceDescriptor` accettano esclusivamente `workspace.schema_version === 3`.
- v1/v2 generano l'errore pubblico già sanitizzato `workspace_invalid`; non usare più il messaggio o lo stato `migration_required` per i descriptor.
- Un'attivazione Git contenente anche un solo descriptor non-v3 fallisce interamente e conserva il precedente active state.
- Le revisioni restituite dalle API contengono esattamente `id`, `commit`, `blob`, `snapshotPath`, senza `state`.
- Nuovi `active.json` e `snapshot.json` non contengono `state` nelle revisioni.
## Compatibilità degli state file esistenti
Definire due decoder stretti e distinti:
```ts
interface StoredWorkspaceRevision {
id: string;
commit: string;
blob: string;
snapshotPath: string;
state?: "operational"; // solo input compatibile; mai restituito
}
interface WorkspaceRevision {
id: string;
commit: string;
blob: string;
snapshotPath: string;
}
```
Regole:
1. `active.json` accetta soltanto `{head,revisions}`; `snapshot.json` soltanto `{head,revisions,files}`.
2. Ogni revision object accetta soltanto i quattro campi correnti più l'opzionale vecchio `state: "operational"`.
3. `state: "migration_required"`, qualsiasi altro valore o campo sconosciuto è rifiutato.
4. Il decoder ricostruisce un nuovo oggetto `WorkspaceRevision`; non restituisce mai l'oggetto JSON originale.
5. Active state e snapshot manifest vengono confrontati dopo la normalizzazione.
6. L'integrità continua a validare path, commit, blob, digest, descriptor v3 e Evidence context.
7. La lettura non modifica snapshot storici. La successiva attivazione riscrive `active.json` nel formato corrente; tutti i nuovi snapshot sono state-free.
8. Un vecchio file già privo di `state` è naturalmente il formato corrente, ma il relativo descriptor deve comunque essere v3.
## Mappa completa dei file
### Backend produttivo
- `backend/src/workspaces/schema.ts`
- `backend/src/workspaces/types.ts`
- `backend/src/workspaces/runtime-renderer.ts`
- `backend/src/workspaces/contracts.ts`
- `backend/src/workspaces/diagnostics.ts`
- `backend/src/workspaces/bindings.ts`
- `backend/src/workspaces/registry.ts`
- `backend/src/routes/workspaces.ts`
- `backend/src/routes/sessions.ts`
- `backend/src/routes/sql.ts`
- Eliminare `backend/src/workspaces/migrate-legacy.ts`
- Eliminare `backend/src/workspaces/migrate-v2-qdrant.ts`
### Build e tooling P1
- `backend/package.json`
- Creare `backend/scripts/clean-dist.mjs`
- Creare un test Node per il clean build
- `backend/scripts/p1-manual-acceptance.mjs`
- `backend/scripts/p1-manual-acceptance.test.mjs`
- `backend/scripts/p1-render-snapshot.test.mjs`
### Frontend
- `frontend/src/api/workspaces.ts`
- `frontend/src/api/sessions.ts`
- `frontend/src/shell/SteerInput.tsx`
- `frontend/src/shell/WorkspaceManager.tsx`
- Test/fixture in `api`, `SteerInput`, `WorkspaceManager`, `NewSessionDialog`, `WorkspacePublishDialog` e `drafts`.
### Deploy, fixture e verificatori
- `scripts/workspace-registry-smoke.sh`
- Creare `scripts/fixtures/workspace-registry-smoke.yaml`
- `scripts/test-no-deployment-coupling-scope.sh`
- `scripts/test-windows-clone-contract.ps1`
- `scripts/verify-workspace-install-docs.sh`
- `scripts/test-verify-workspace-install-docs.sh`
### Documentazione corrente
- `README.md`
- sezione corrente di `PROJECT_STATE.md`, prima di `## Historical snapshots`
- `docs/workspace-diagnostic-protocol.md`
- `docs/install/local-workspace-registry.md`
- `docs/install/server-workspace-registry.md`
---
### Task 0: Congelare scope e baseline prima delle modifiche
**File:** nessuna modifica produttiva.
- [ ] Registrare `BASE_SHA=$(git rev-parse HEAD)` e verificare che gli altri piani non vengano inclusi nei commit di implementazione.
- [ ] Salvare l'inventario iniziale dei simboli descriptor-legacy:
```bash
git grep -nE 'WorkspaceV1|WorkspaceV2|LegacyWorkspace|migration_required|migrate-legacy|migrateWorkspaceV1ToV2|migrateWorkspaceV2ToV3' -- \
backend/src backend/test backend/scripts frontend/src scripts README.md PROJECT_STATE.md docs/install docs/workspace-diagnostic-protocol.md
```
- [ ] Classificare ogni risultato come descriptor legacy, compatibility decoder previsto, contratto diverso o documento storico.
- [ ] Verificare nei registry/installazioni disponibili che i descriptor attivi siano v3; questa è una precondizione di deploy, non un migratore.
- [ ] Non procedere se il worktree contiene modifiche applicative non attribuibili a questo piano.
### Task 1: Scrivere i test RED del contratto v3-only
**File:**
- `backend/test/workspaces-schema.test.ts`
- `backend/test/workspace-registry.test.ts`
- `backend/test/routes-workspaces.test.ts`
- [ ] Aggiungere test che `parseWorkspaceYaml`, `validateWorkspaceDescriptor` e le route validate/publish rifiutino esplicitamente v1 e v2.
- [ ] Aggiungere test registry per:
- bootstrap pulito con solo v1/v2: fallimento, nessun `active.json` pubblicato;
- repository misto v3+v2: attivazione atomica rifiutata;
- pull che introduce v1/v2: precedente active state ancora leggibile;
- retained snapshot contenente descriptor non-v3: rifiuto fail-closed;
- risposta API state-free.
- [ ] Eseguire:
```bash
cd backend
npx vitest run test/workspaces-schema.test.ts test/workspace-registry.test.ts test/routes-workspaces.test.ts
```
Atteso: RED per i nuovi requisiti, non errori di fixture casuali.
### Task 2: Rendere lo schema backend esclusivamente v3
**File:**
- `backend/src/workspaces/schema.ts`
- `backend/src/workspaces/types.ts`
- test del Task 1
- [ ] Eliminare `WorkspaceV1`, `WorkspaceV2`, `LegacyWorkspace`, relativi Zod schema e `migrateWorkspaceV1ToV2`.
- [ ] Rendere `WorkspaceDescriptorSchema = WorkspaceV3Schema`.
- [ ] Eliminare `validateCanonicalWorkspace`, aggiornando **tutti** i chiamanti in `routes/workspaces.ts`, incluso il chiamante attualmente oltre quelli elencati nel vecchio piano.
- [ ] Eliminare `isCanonicalWorkspace`/`isOperationalWorkspace` dopo aver sostituito i rami condizionali con validazione v3 diretta.
- [ ] Conservare test negativi v1/v2; non cancellare le sole prove che impediscono una regressione futura.
- [ ] Eseguire test focalizzati e typecheck.
- [ ] Commit: `refactor: make workspace descriptors schema v3 only`.
### Task 3: Normalizzare in sicurezza active state e snapshot manifest
**File:**
- `backend/src/workspaces/registry.ts`
- `backend/test/workspace-registry.test.ts`
- [ ] Scrivere RED per state/manifest con:
- campo assente;
- vecchio `state: "operational"`;
- `state: "migration_required"`;
- valore sconosciuto;
- campo extra;
- active state e manifest con formati misti;
- snapshot attivo, storico e fallback offline.
- [ ] Rimuovere `state` da `WorkspaceRevision` e da tutti i nuovi writer.
- [ ] Sostituire cast e vecchie migrazioni con decoder stretti che restituiscono oggetti normalizzati state-free.
- [ ] Rimuovere `LegacyWorkspaceRevision`, `LegacyActiveState`, `LegacySnapshotManifest`, `deriveStateFromLegacyRevisions`, `migrateLegacyActiveState`, `migrateLegacySnapshotManifest`, `sameLegacyRevisions` e le condizioni operative basate su `state`.
- [ ] Mantenere tutti i controlli di integrità e far validare ogni YAML come v3.
- [ ] Provare che list/read/API non riemettono il vecchio campo anche immediatamente dopo un restart, prima di una nuova attivazione.
- [ ] Commit: `refactor: remove workspace revision state`.
### Task 4: Eliminare i rami v1/v2 da renderer, contracts, bindings e diagnostica
**File:**
- `backend/src/workspaces/runtime-renderer.ts`
- `backend/src/workspaces/contracts.ts`
- `backend/src/workspaces/diagnostics.ts`
- `backend/src/workspaces/bindings.ts`
- relativi test
- [ ] Scrivere/aggiornare test RED che accettano v3 e rifiutano input non-v3 al confine, senza renderer/diagnoser legacy.
- [ ] Eliminare il renderer v2/pgvector e i rami v1.
- [ ] Eliminare variabili contract e diagnostica solamente v2.
- [ ] Semplificare bindings dopo la validazione v3, senza indebolire validazione secrets/trasporti.
- [ ] Eseguire i test focalizzati:
```bash
cd backend
npx vitest run \
test/workspace-runtime-renderer.test.ts \
test/workspaces-contracts.test.ts \
test/workspaces-diagnostics.test.ts \
test/workspaces-bindings.test.ts \
test/workspace-runtime-handoff.test.ts
```
- [ ] Commit: `refactor: remove legacy workspace runtime branches`.
### Task 5: Rimuovere migratori senza perdere test di deployment non correlati
**File:**
- Eliminare i due migratori e i test esclusivamente di migrazione.
- Creare/spostare in un test dedicato le prove deployment presenti in `workspaces-migrate-legacy.test.ts:81-114`.
- [ ] Prima di eliminare `workspaces-migrate-legacy.test.ts`, spostare in un file con nome coerente:
- volume registry durevole e mount Git read-only;
- contratto Dockerfile;
- fallback offline smoke;
- self-test di cleanup dell'immagine per-run.
- [ ] Eliminare `migrate-legacy.ts`, `migrate-v2-qdrant.ts` e i test di trasformazione.
- [ ] Conservare un fixture v2 soltanto nei test negativi di rifiuto.
- [ ] Eseguire i nuovi test deployment e il typecheck.
- [ ] Commit: `refactor: remove workspace migration utilities`.
### Task 6: Aggiornare tutte le route backend e il tooling P1
**File:**
- `backend/src/routes/workspaces.ts`
- `backend/src/routes/sessions.ts`
- `backend/src/routes/sql.ts`
- test route inclusi `routes-sql-meta.test.ts`
- `backend/scripts/p1-manual-acceptance.mjs`
- test manual/render P1
- [ ] Rimuovere filtri/gate `revision.state` da tutte le route. La garanzia deriva dal registry v3-only.
- [ ] Aggiornare mock/fixture `WorkspaceRevision` in tutti i test backend.
- [ ] Aggiornare il validatore del manifest P1 manuale affinché richieda esattamente la revisione state-free.
- [ ] Aggiornare i fixture `p1-manual-acceptance.test.mjs` e `p1-render-snapshot.test.mjs`.
- [ ] Aggiungere un test JS specifico che rifiuti manifest con revisioni malformate senza reintrodurre `migration_required`.
- [ ] Eseguire:
```bash
cd backend
npx vitest run test/routes-workspaces.test.ts test/routes-sessions.test.ts test/routes-sql-meta.test.ts
cd ..
node --test --test-concurrency=1 \
backend/scripts/p1-manual-acceptance.test.mjs \
backend/scripts/p1-render-snapshot.test.mjs
```
- [ ] Commit: `refactor: remove workspace revision state consumers`.
### Task 7: Rendere il build backend clean-first
**File:**
- `backend/package.json`
- Creare `backend/scripts/clean-dist.mjs`
- Creare test Node del clean build
- [ ] Scrivere RED: creare un file sentinella in `backend/dist/workspaces/`, eseguire il clean/build e verificare che non sopravviva.
- [ ] Implementare la pulizia con API Node multipiattaforma, non con `rm -rf` nella npm script.
- [ ] Fare eseguire il clean prima di `tsc` da `npm run build`.
- [ ] Verificare dopo il build:
```bash
test ! -e backend/dist/workspaces/migrate-legacy.js
test ! -e backend/dist/workspaces/migrate-v2-qdrant.js
```
- [ ] Costruire l'immagine core in un contesto pulito e verificare che i due moduli non esistano nell'immagine.
- [ ] Verificare che i manifest di integrità P1 continuino a legare l'intero nuovo `dist`.
- [ ] Commit: `build: remove stale backend distribution files`.
### Task 8: Aggiornare frontend e contratto API state-free
**File:**
- `frontend/src/api/workspaces.ts`
- `frontend/src/api/sessions.ts`
- `frontend/src/shell/SteerInput.tsx`
- `frontend/src/shell/WorkspaceManager.tsx`
- test/fixture frontend correlati
- [ ] Scrivere/aggiornare test per revisioni senza `state` e risposta non-v3 rifiutata al confine workspace.
- [ ] Eliminare `state` dal tipo e dal parser revisionale.
- [ ] Rimuovere gate/banner/filtro `migration_required` e anche la visualizzazione `record.revision.state`.
- [ ] Mantenere `allowLegacy` per sessioni senza revisione.
- [ ] Aggiornare fixture in:
- `api/workspaces.test.ts`, `api/sessions.test.ts`;
- `SteerInput.test.tsx`, `WorkspaceManager.test.tsx`;
- `NewSessionDialog.test.tsx`, `WorkspacePublishDialog.test.tsx`;
- `drafts.test.ts`, mantenendo il test negativo di schema non-3.
- [ ] Documentare che core e frontend devono essere aggiornati insieme; il parser nuovo non usa più `state`.
- [ ] Eseguire typecheck e suite frontend.
- [ ] Commit: `refactor: remove legacy workspace UI state`.
### Task 9: Sostituire fixture e smoke con descriptor v3 completi
**File:**
- `scripts/workspace-registry-smoke.sh`
- Creare `scripts/fixtures/workspace-registry-smoke.yaml`
- `scripts/test-no-deployment-coupling-scope.sh`
- `scripts/test-windows-clone-contract.ps1`
- test deployment spostati nel Task 5
- [ ] Creare un descriptor v3 completo `id: local`, collection `local`, embedding interno 1024/cosine, LLM policy e diagnostica DWH; omettere Evidence per non richiedere un tree Git nello smoke registry.
- [ ] Validare il fixture con il parser produttivo in un test backend.
- [ ] Copiare il fixture nello seed repository e rimuovere sia l'invocazione del migratore sia il build backend ormai inutile allo smoke.
- [ ] Nel test Windows non cambiare soltanto il numero di versione: fornire il contratto v3 completo mantenendo lo scopo path-with-spaces/clone.
- [ ] Aggiornare il fixture dello scope coupling senza indebolire l'assenza-gate.
- [ ] Eseguire test shell focalizzati e, con Docker disponibile, lo smoke reale senza retry.
- [ ] Commit: `test: replace legacy workspace deployment fixtures`.
### Task 10: Aggiornare documentazione corrente e relativi verifier
**File:**
- documenti/verifier indicati nella mappa
- [ ] Aggiornare README e soltanto la sezione corrente di `PROJECT_STATE.md`; non riscrivere gli snapshot storici.
- [ ] Eliminare procedure di migrazione v1/v2 dai manuali local/server e dal protocollo diagnostico.
- [ ] Modificare `verify-workspace-install-docs.sh` perché richieda “schema v3 only” e l'assenza di `migration_required` nella documentazione corrente.
- [ ] Aggiornare i fixture negativi del test del verifier.
- [ ] Non cambiare gli usi di `migration_required` nei piani P3–P6 relativi a ownership/artifact diversi.
- [ ] Eseguire:
```bash
bash scripts/test-verify-workspace-install-docs.sh
bash scripts/verify-workspace-install-docs.sh --fixtures-only
```
- [ ] Commit: `docs: make schema v3 the only workspace contract`.
### Task 11: Eseguire absence gate e suite complete
- [ ] Eseguire backend clean build, typecheck e test:
```bash
cd backend
npm run build
npx tsc --noEmit -p .
npx vitest run
```
- [ ] Eseguire frontend:
```bash
cd frontend
npx tsc -b
npx vitest run
npm run build
```
- [ ] Eseguire script/verifier interessati, incluso lo smoke Docker obbligatorio se l'ambiente dispone di Docker. Non lasciarlo “opzionale” in una consegna che modifica lo smoke.
- [ ] Eseguire `git diff --check`.
- [ ] Eseguire l'absence gate ristretto:
```bash
git grep -nE 'WorkspaceV1|WorkspaceV2|LegacyWorkspace|migrateWorkspaceV1ToV2|migrateWorkspaceV2ToV3' -- \
backend/src frontend/src scripts && exit 1 || true
git grep -nE 'migration_required|migrate-legacy|migrate-v2-qdrant' -- \
backend/src backend/scripts frontend/src scripts README.md docs/install docs/workspace-diagnostic-protocol.md && exit 1 || true
test ! -e backend/dist/workspaces/migrate-legacy.js
test ! -e backend/dist/workspaces/migrate-v2-qdrant.js
```
Nota: trasformare questi esempi in uno script con allowlist esplicita; non affidarsi a `&& exit 1 || true`, che può mascherare errori di esecuzione. Lo script deve distinguere “nessun match” da errore Git/I/O.
- [ ] Ispezionare il diff per assicurarsi che nessun formato non-descriptor sia stato modificato.
### Task 12: Rigenerare l'evidenza automatica P1
- [ ] Partire dal commit sorgente finale pulito.
- [ ] Eseguire una sola integrazione completa, senza retry automatico:
```bash
./scripts/p1-acceptance.sh integration --keep
```
- [ ] Verificare report JSON/Markdown, hash dichiarati, manifest sorgente/dist, secret scan, ownership cleanup e porte chiuse.
- [ ] Aggiornare `PROJECT_STATE.md` con il nuovo commit/tree/report e con stati distinti:
```text
automated integration: PASS
manual acceptance: PENDING
```
- [ ] Committare soltanto lo stato tracciato, mai `.artifacts`.
- [ ] Non riusare l'evidenza precedente legata a `c733896`.
### Task 13: Riaprire e chiudere il gate manuale P1
- [ ] Preparare un ambiente manuale nuovo:
```bash
./scripts/p1-manual-acceptance.sh prepare
./scripts/p1-manual-acceptance.sh serve
```
- [ ] Il reviewer segue integralmente il nuovo `GUIDE.md`, verificando anche che revisioni/API/manifest siano state-free e che v1/v2 siano rifiutati senza mutazione.
- [ ] Arrestare il server e verificare porte/processi:
```bash
./scripts/p1-manual-acceptance.sh stop
```
- [ ] Solo il reviewer crea `VERDICT.md` e decide PASS/FAIL.
- [ ] Se PASS, aggiornare `PROJECT_STATE.md` e committare `docs: record schema-v3-only P1 acceptance`.
- [ ] Pulire il lab soltanto dopo conferma del reviewer.
- [ ] **STOP:** non iniziare P2 finché il reviewer non approva esplicitamente il nuovo P1.
---
## Criteri finali di accettazione
1. Nessun descriptor v1/v2 viene parsato, pubblicato, attivato, renderizzato, diagnosticato o mostrato.
2. I vecchi state file di revisioni v3 con `state: "operational"` continuano a caricarsi, ma API e nuovi file sono state-free.
3. Descriptor non-v3 o state incoerenti falliscono senza sostituire il precedente active state.
4. Nessun migratore sopravvive in sorgenti, `dist`, immagine core, script o documentazione corrente.
5. I formati versionati non collegati ai workspace descriptor sono invariati.
6. Backend, frontend, verifier, smoke e build interessati sono verdi.
7. Una nuova integrazione P1 è PASS al commit finale.
8. La nuova acceptance manuale P1 è decisa esplicitamente dal reviewer.
9. P2 resta non iniziato fino a ulteriore autorizzazione.
@@ -1,401 +0,0 @@
# Read-only Workspace Runtime Secrets Implementation Plan
> **Historical nomenclature:** this plan predates the native host CLI convergence. References to
> `thothctl` and `tools/thothctl` describe the implementation snapshot from which this plan was
> written; current operator commands and paths use native `tht` and `tools/tht`.
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
**Goal:** Make workspace consumption strictly read-only while adding installation-scoped Git identity and persistent GUI-managed runtime secrets.
**Architecture:** Git remains the source of truth and is fetched into an application-owned checkout; complete candidate commits are validated before atomic activation and the backend has no Git write path. Runtime connector credentials are discovered from trusted connector contracts, stored as authenticated ciphertext by a backend vault, and materialized only for the lifetime of diagnostics or runtime leases. The browser exposes repository/readiness status and write-only secret forms without workspace persistence.
**Tech Stack:** Fastify, TypeScript, Node.js crypto/filesystem, React 18, TanStack Query, Vitest, Go `thothctl`, Docker Compose.
---
### Task 1: Freeze the Git repository boundary to read-only
**Files:**
- Modify: `backend/src/workspaces/types.ts`
- Modify: `backend/src/workspaces/git-repository.ts`
- Modify: `backend/src/workspaces/registry.ts`
- Modify: `backend/test/workspaces-git-repository.test.ts`
- Modify: `backend/test/workspace-registry.test.ts`
- Modify: `backend/test/workspace-registry-deployment.test.ts`
**Step 1: Write failing tests**
Add tests proving that pull never configures a Git author, writes generated files, commits, or pushes; that a malformed candidate leaves the prior active snapshot intact; and that a missing catalog descriptor rejects the whole candidate instead of producing a bootstrap slot.
**Step 2: Run the focused tests**
Run: `cd backend && npx vitest run test/workspaces-git-repository.test.ts test/workspace-registry.test.ts test/workspace-registry-deployment.test.ts`
Expected: FAIL on write/publish behavior and missing-descriptor semantics.
**Step 3: Implement the read-only boundary**
Remove `gitAuthorName`, `gitAuthorEmail`, mutation helpers, generated-document reconciliation, publish/conflict types, and bootstrap-slot activation. `pull()` must fetch, validate the complete commit in a candidate snapshot, and replace active state only after validation succeeds.
**Step 4: Run focused tests**
Run the command from Step 2.
Expected: PASS.
**Step 5: Commit**
```bash
git add backend/src/workspaces backend/test/workspaces-git-repository.test.ts backend/test/workspace-registry.test.ts backend/test/workspace-registry-deployment.test.ts
git commit -m "refactor: make workspace repository strictly read only"
```
### Task 2: Remove publishing and bundle HTTP contracts
**Files:**
- Modify: `backend/src/routes/workspaces.ts`
- Modify: `backend/test/routes-workspaces.test.ts`
- Modify: `backend/test/workspaces-runtime-v3-boundaries.test.ts`
- Modify: `backend/src/config.ts`
- Modify: `backend/test/workspaces-config.test.ts`
**Step 1: Write failing route tests**
Assert `POST /workspaces/publish`, `GET /workspaces/:id/export`, and `POST /workspaces/import` return 404 and that the backend no longer registers multipart or ZIP handling. Assert configuration no longer accepts Git author or bundle-limit settings as workspace-registry fields.
**Step 2: Run tests and observe failure**
Run: `cd backend && npx vitest run test/routes-workspaces.test.ts test/workspaces-config.test.ts test/workspaces-runtime-v3-boundaries.test.ts`
Expected: FAIL because mutation and bundle routes still exist.
**Step 3: Remove the mutation surface**
Delete publish/import/export schemas and helpers, remove `multipart`, `yauzl`, and `yazl` usage from the route, and simplify safe workspace errors to read/validate/sync errors.
**Step 4: Run tests**
Run the command from Step 2 plus `cd backend && npx tsc --noEmit -p .`.
Expected: PASS.
**Step 5: Commit**
```bash
git add backend/src backend/test package.json package-lock.json
git commit -m "refactor: remove workspace publishing and bundles"
```
### Task 3: Expose a sanitized installation repository identity
**Files:**
- Modify: `backend/src/workspaces/git-repository.ts`
- Modify: `backend/src/routes/workspaces.ts`
- Modify: `backend/test/workspaces-git-repository.test.ts`
- Modify: `backend/test/routes-workspaces.test.ts`
- Modify: `tools/thothctl/internal/config/installation.go`
- Modify: `tools/thothctl/internal/config/installation_test.go`
- Modify: `deploy/psd/thothii-installation.yaml.example`
- Modify: `docs/install/examples/thothii-installation.local.yaml`
- Modify: `docs/install/examples/thothii-installation.server.yaml`
**Step 1: Write failing parser and status tests**
Cover HTTPS, SSH URL, and SCP-style remotes; reject embedded user-info for HTTPS; return only `host`, `repository`, `branch`, and `transport`; never return a token, key path, or raw credential-bearing URL. Add installation-descriptor tests for a required `workspaceRepository` block and exactly one read-only transport.
**Step 2: Run focused tests**
Run: `cd backend && npx vitest run test/workspaces-git-repository.test.ts test/routes-workspaces.test.ts && cd ../tools/thothctl && go test ./internal/config`
Expected: FAIL because repository identity and typed installation configuration do not exist.
**Step 3: Implement safe normalization and installation validation**
Add the normalized identity to registry status. Extend `thothii-installation.yaml` with remote, branch, and SSH/HTTPS access metadata, validate it against the selected Compose override and environment without reading or returning secret values, and retain the existing environment rendering boundary.
**Step 4: Run focused tests**
Run the command from Step 2.
Expected: PASS.
**Step 5: Commit**
```bash
git add backend tools/thothctl deploy docs/install/examples
git commit -m "feat: declare workspace repository in installation config"
```
### Task 4: Add the persistent encrypted workspace secret store
**Files:**
- Create: `backend/src/workspaces/secret-store.ts`
- Create: `backend/test/workspace-secret-store.test.ts`
- Modify: `backend/src/config.ts`
- Modify: `backend/src/app.ts`
- Modify: `compose.yaml`
- Modify: `deploy/compose.local.yaml`
- Modify: `deploy/compose.server.yaml`
**Step 1: Write failing vault tests**
Test first-start initialization, atomic blind replacement, deletion, enumeration by configured ID only, AES-256-GCM ciphertext with installation/workspace/field associated data, corruption failure, restrictive files/directories, size limits, and absence of plaintext in persistent bytes.
**Step 2: Run the vault test**
Run: `cd backend && npx vitest run test/workspace-secret-store.test.ts`
Expected: FAIL because `WorkspaceSecretStore` does not exist.
**Step 3: Implement the vault**
Create an injectable `WorkspaceSecretStore` backed by an application-managed data root. Persist a versioned encrypted document atomically, generate or load the installation vault key in the private control area, expose only `has`, `put`, `delete`, and scoped materialization operations, and never add a plaintext read API.
**Step 4: Run tests and typecheck**
Run: `cd backend && npx vitest run test/workspace-secret-store.test.ts && npx tsc --noEmit -p .`
Expected: PASS.
**Step 5: Commit**
```bash
git add backend compose.yaml deploy
git commit -m "feat: persist encrypted workspace runtime secrets"
```
### Task 5: Derive connector requirements and integrate temporary materialization
**Files:**
- Create: `backend/src/workspaces/secret-requirements.ts`
- Create: `backend/test/workspace-secret-requirements.test.ts`
- Modify: `backend/src/workspaces/bindings.ts`
- Modify: `backend/src/workspaces/runtime-config-lease.ts`
- Modify: `backend/src/tht/tht-runner.ts`
- Modify: `backend/src/app.ts`
- Modify: `backend/test/workspace-runtime-config-lease.test.ts`
- Modify: `backend/test/workspace-runtime-handoff.test.ts`
- Modify: `backend/test/workspaces-bindings.test.ts`
**Step 1: Write failing requirement and lifecycle tests**
Cover PostgreSQL password, REST bearer API key, unauthenticated REST, SSH private key/password, signed HTTP Evidence, and static S3 credentials. Assert temporary files are restrictive, live for exactly one diagnostic/runtime lease, disappear on release and error, and are never persisted in the encrypted vault document.
**Step 2: Run focused tests**
Run: `cd backend && npx vitest run test/workspace-secret-requirements.test.ts test/workspaces-bindings.test.ts test/workspace-runtime-config-lease.test.ts test/workspace-runtime-handoff.test.ts`
Expected: FAIL because requirements still come from installation secret-file paths.
**Step 3: Implement dynamic requirement resolution**
Use the selected DWH transport and Evidence authentication contract to map trusted installation-contract suffixes to stable GUI requirement IDs. Overlay materialized temporary file paths only while resolving existing file-oriented connectors, and attach cleanup to every runtime lease.
**Step 4: Run tests and typecheck**
Run the command from Step 2 plus `cd backend && npx tsc --noEmit -p .`.
Expected: PASS.
**Step 5: Commit**
```bash
git add backend/src backend/test
git commit -m "feat: resolve workspace secrets from connector requirements"
```
### Task 6: Add write-only workspace secret and readiness APIs
**Files:**
- Modify: `backend/src/routes/workspaces.ts`
- Modify: `backend/src/app.ts`
- Modify: `backend/src/workspaces/types.ts`
- Modify: `backend/test/routes-workspaces.test.ts`
**Step 1: Write failing API tests**
Test `GET /workspaces/:id/runtime-configuration`, blind `PUT /workspaces/:id/secrets`, and `DELETE /workspaces/:id/secrets/:requirementId`. Assert strict bodies, limits, unknown-ID rejection, status-only responses, diagnostic invalidation, and `configuration_required`/`ready` state transitions.
**Step 2: Run tests**
Run: `cd backend && npx vitest run test/routes-workspaces.test.ts`
Expected: FAIL because the routes do not exist.
**Step 3: Implement the routes and readiness projection**
Inject the secret store into workspace routes and runtime support. Compute per-workspace readiness from active descriptor, current requirement set, configured IDs, and diagnostic generation. Materialize values only inside the diagnostic request and always clean up.
**Step 4: Run backend gates**
Run: `cd backend && npx vitest run && npx tsc --noEmit -p . && npm run build`.
Expected: PASS.
**Step 5: Commit**
```bash
git add backend
git commit -m "feat: manage runtime workspace secrets through the API"
```
### Task 7: Replace workspace management with the two-level read-only UI
**Files:**
- Modify: `frontend/src/api/workspaces.ts`
- Modify: `frontend/src/api/workspaces.test.ts`
- Modify: `frontend/src/shell/WorkspaceManager.tsx`
- Modify: `frontend/src/shell/WorkspaceManager.test.tsx`
- Delete: `frontend/src/shell/WorkspacePublishDialog.tsx`
- Delete: corresponding publish-dialog tests
- Modify/Delete: `frontend/src/shell/WorkspaceEditor.tsx` and bootstrap-only tests as references permit
- Modify: `frontend/src/workspaces/drafts.ts`
- Modify: `frontend/src/workspaces/drafts.test.ts`
**Step 1: Write failing UI/API tests**
Assert the dialog uses at least 60% viewport width and height, shows general repository concepts and exact button consequences at level 1, gates workspace-specific controls on selection, renders requirement explanations and write-only fields at level 2, and has no create/edit/publish/import/export/bundle controls.
**Step 2: Run focused tests**
Run: `cd frontend && npx vitest run src/api/workspaces.test.ts src/shell/WorkspaceManager.test.tsx src/workspaces/drafts.test.ts`
Expected: FAIL on the old draft/publish interface.
**Step 3: Implement the read-only interface**
Replace bootstrap editor state with repository status, selection, validation/readiness details, dynamic secret fields, blind save/forget actions, and connection test. Remove workspace draft persistence and clear secret field component state after submit/close.
**Step 4: Run focused tests and typecheck**
Run the command from Step 2 plus `cd frontend && npx tsc -b`.
Expected: PASS.
**Step 5: Commit**
```bash
git add frontend
git commit -m "feat: add read-only workspace and secret management UI"
```
### Task 8: Remove browser-persisted workspace preferences
**Files:**
- Modify: `frontend/src/workspaces/preferences.ts`
- Modify: `frontend/src/workspaces/preferences.test.ts`
- Modify: `frontend/src/api/sessions.ts`
- Modify: `frontend/src/api/sessions.test.ts`
- Modify: `frontend/src/shell/SteerInput.tsx`
- Modify: `frontend/src/shell/SteerInput.test.tsx`
**Step 1: Write failing persistence-boundary tests**
Assert workspace/model/thinking choices are kept only in current application memory or saved through the existing backend settings API, and that no workspace code calls `localStorage`.
**Step 2: Run focused tests**
Run: `cd frontend && npx vitest run src/workspaces/preferences.test.ts src/api/sessions.test.ts src/shell/SteerInput.test.tsx`
Expected: FAIL because preferences still use browser storage.
**Step 3: Implement ephemeral preferences**
Replace the storage adapter with an in-memory external store seeded from backend settings. Preserve concurrent workspace-policy gates and session request determinism without persisting selections in the browser.
**Step 4: Run frontend gates**
Run: `cd frontend && npx vitest run && npx tsc -b && npm run build`.
Expected: PASS.
**Step 5: Commit**
```bash
git add frontend
git commit -m "refactor: stop persisting workspace state in the browser"
```
### Task 9: Update deployment contracts and documentation
**Files:**
- Modify: `compose.yaml`
- Modify: `deploy/compose.git-ssh.yaml`
- Modify: `deploy/compose.git-https.yaml`
- Modify: `deploy/workspace-registry.env.example`
- Modify: `deploy/psd/operator.env.example`
- Modify: `docs/install/local-workspace-registry.md`
- Modify: `docs/install/server-workspace-registry.md`
- Modify: `docs/guida-utente.md`
- Modify: `scripts/verify-workspace-install-docs.sh`
- Modify: `scripts/workspace-registry-smoke.sh`
**Step 1: Update executable contract tests first**
Require read-only Git wording and configuration, repository identity visibility, vault persistence,
and absence of author/push/bundle/browser-secret instructions.
**Step 2: Run contract tests and observe failure**
Run: `bash scripts/verify-workspace-install-docs.sh`
Expected: FAIL against the old manuals and examples.
**Step 3: Update deployment and manuals**
Remove Git author settings and write-oriented documentation. Document installation Git bootstrap,
GUI runtime-secret completion, platform-neutral application storage, rotation/forget flows, and
candidate validation semantics.
**Step 4: Run contract and Go gates**
Run: `bash scripts/verify-workspace-install-docs.sh && cd tools/thothctl && go test ./...`
Expected: PASS.
**Step 5: Commit**
```bash
git add compose.yaml deploy docs scripts tools/thothctl
git commit -m "docs: describe read-only workspace runtime configuration"
```
### Task 10: Full verification and deployed-container refresh
**Files:**
- Modify only files needed to fix failures found by verification.
**Step 1: Run static and unit gates**
```bash
cd backend && npx vitest run && npx tsc --noEmit -p . && npm run build
cd ../frontend && npx vitest run && npx tsc -b && npm run build
cd ../harness && .venv/bin/pytest -q
cd ../tools/thothctl && go test ./...
```
Expected: all gates PASS.
**Step 2: Run deployment contract gates**
Run: `bash scripts/verify-workspace-install-docs.sh` and the focused workspace registry smoke appropriate to the configured installation.
Expected: PASS without Git writes or secret disclosure.
**Step 3: Inspect the final diff and secret scan**
Run: `git diff --check`, inspect `git status --short`, and search active code/config for removed publish, bundle, Git author, and workspace-localStorage contracts.
Expected: no whitespace errors, no accidental secrets, and only intended changes.
**Step 4: Rebuild and restart affected services**
Use the installation-aware `thothctl` lifecycle for the configured installation to rebuild/restart `core` and `frontend`, then verify health and repository status. Do not restart if no valid local installation descriptor is available; report that external gate explicitly.
**Step 5: Commit verification fixes**
```bash
git add <only-files-changed-for-verification>
git commit -m "test: verify read-only workspace secret flow"
```
@@ -1,217 +0,0 @@
# Evidence canonica — struttura tipizzata per disambiguazione, schema linking e SQL
> **Superseded (2026-08-24).** Questo documento conserva la storia della prima
> proposta. Il disegno approvato è
> [`2026-08-24-evidence-restructuring-design.md`](2026-08-24-evidence-restructuring-design.md)
> e il relativo piano esecutivo è
> [`2026-08-24-evidence-restructuring.md`](2026-08-24-evidence-restructuring.md).
## Contesto e decisioni prese
ThothII ha già due livelli separati che non si parlano:
- **`EvidenceDoc`** (`harness/tht/evidence/model.py`) — runtime: frontmatter piatto
(`id/title/tier/status/tables/concepts/sources`) + `body` markdown libero. `tier`
distingue solo `structural|concept`, insufficiente rispetto ai 6 tipi reali del corpus.
- **Pipeline corpus** (`harness/tht/corpus/`) — canonizzazione *tecnica* (hash,
provenienza, chunking, vettorizzazione), agnostica rispetto al tipo di evidenza.
Il corpus reale (`ChironeWp3/artifacts/evidence/`) ha una tassonomia implicita in 6
directory (`00-glossario`, `10-domini-clinici`, `20-valori-enum`, `30-esempi-nlq`,
`40-mapping-semantico`, `50-metadati-normalizzazione`) ma:
- frontmatter piatto, `tier` inadeguato;
- file già rotti (`--` invece di `---`, bullet `•⁠ ⁠` invece di `-`) che `EvidenceDoc.parse`
rifiuterebbe;
- al retrieval la struttura si perde: `tht search pack` proietta solo `title` + 400 char.
**Decisioni (confermate nel brainstorming):**
1. **Obiettivo**: strutturare il *contenuto* runtime, non toccare la pipeline corpus.
2. **Forma**: ibrido — frontmatter tipizzato + sezioni canoniche per `kind`.
3. **Tassonomia**: doppia — `kind` (contenuto) + `applies_to` (destinazioni:
`disambiguation`, `rewriting`, `schema_linking`, `sql_generation`, `memory`).
4. **Anchors come fonte primaria** per il value-grounding in F4; LSH solo fallback.
5. **Approccio**: contratto in `tht` + authoring sottile (niente app separata, niente
client LLM diretto in `tht`).
6. **Modello sorgente→canonizzato** (non sovrascrittura in place): file umani in
`source_root/evidence/`, canonizzati derivati in `artifacts/evidence/`, coerente con
l'esistente `tht evidence extract`.
7. **Review**: batch con diff aggregato (approvazione per file).
8. **Rielaborazione**: LLM-assistita via Pi (sessione di manutenzione + gate), con parte
deterministica in `tht`.
## Modello canonico v2
### `CanonicalEvidence` (frontmatter tipizzato)
```yaml
---
schema_version: 2
id: ev-dom-ablazione-see
title: Dominio Ablazione e SEE
kind: domain # glossario | domain | enum | example | mapping | normalization
applies_to: # destinazioni d'uso
- disambiguation # F1
- schema_linking # F4
- sql_generation # F6/F7
status: reviewed
language: it
concepts:
- term: ablazione
synonyms: [ablazione transcatetere, SEE, studio elettrofisiologico]
- term: fibrillazione atriale
synonyms: [FA]
tables:
- name: datawarehouse.fact_studio_elettrofisiologico_endocavitario_ablazione
role: fact
columns: [ablazione_transcatetere, cod_paz, num]
anchors:
- column: datawarehouse.fact_see_ablazione_procedura_patologia.patologia
value: ablazione
match: exact
sources: [...]
# provenienza del derivato (aggiunta dal canonicalize, non dall'umano):
source_file: 10-domini-clinici/ablazione.md
source_fingerprint: sha256:...
canonicalized_at: 2026-08-18T...
---
```
Il `body` resta markdown, ma con **titoli di sezione canonici per `kind`**:
- `glossario`: `## Definizione`
- `domain`: `## Cosa rappresenta`, `## Schema a stella`, `## Granularità`, `## Domande di business tipiche`
- `enum`: `## Valori ammessi`
- `example`: `## Domanda → SQL` + `## Nota clinica`
- `mapping`: `## Trasformazioni disponibili`
- `normalization`: `## Regole di normalizzazione`
Le sezioni canoniche permettono al retrieval di estrarre solo la sezione rilevante per
fase invece di 400 caratteri generici.
### `kind` vs `applies_to`
- `kind` descrive il *contenuto* (i 6 tipi già impliciti).
- `applies_to` descrive *quando usarlo*. Esempi: `enum-tipo-intervento` è `kind: enum` ma
`applies_to: [disambiguation, sql_generation]`; un `example` è
`applies_to: [sql_generation, memory]`.
## Modello a 3 livelli
```
source_root/evidence/*.md (umano, libero, può essere sporco)
│ tht evidence canonicalize (deterministico + LLM via Pi + review batch)
▼
artifacts/evidence/*.md (canonico, validato, derivato, rigenerabile)
│ tht evidence index (esistente)
▼
corpus / vector store (vettorializzato dal canonico, mai dal sorgente)
```
Il canonizzato porta `source_file` + `source_fingerprint`: se il sorgente cambia, la
canonizzazione ripropone il diff; se invariato, no-op.
## Componenti da creare/modificare
### 1. `harness/tht/evidence/model.py` (modifica)
- Aggiungere `CanonicalEvidence` (pydantic, v2) con `schema_version`, `kind`,
`applies_to`, `concepts[]` (`{term, synonyms[]}`), `tables[]`
(`{name, role, columns[]}`), `anchors[]` (`{column, value, match}`), `source_file`,
`source_fingerprint`, `canonicalized_at`.
- Validatore strict: `kind` ammesso, `applies_to` ammesso, `anchors[].column` deve
essere `schema.colonna` ben formato, `match` in `{exact, contains, regex, substring}`.
- `EvidenceDoc` v1 resta per leggere il sorgente e per retro-compatibilità dei vecchi
canonizzati; `CanonicalEvidence` è un modello separato, non una sottoclasse.
- Validatore che garantisce `source_fingerprint` = sha256 del contenuto sorgente.
### 2. `harness/tht/evidence/lint.py` (nuovo, deterministico)
- `lint_source(path) -> list[Diagnostic]`: frontmatter rotto, `tier`/`status` non validi,
bullet Unicode, campi mancanti, `tables` non qualificate con schema, concetti senza
sinonimi, sezioni non canoniche per il `kind` atteso.
- Exit code 0 se nessun errore, 1 se warning, 2 se errori. Nessun LLM.
### 3. `harness/tht/evidence/canonicalize.py` (nuovo)
- `plan(sources, artifacts) -> CanonicalizePlan`: confronta fingerprint dei sorgenti con
i canonizzati esistenti, produce la lista dei file da (ri)canonizzare.
- `render_proposal(source, llm_completions) -> CanonicalEvidence`: assembla il
canonizzato da parsing deterministico + campi semantici forniti da Pi.
- `apply(plan, approved_ids) -> None`: scrive solo i canonizzati approvati in
`artifacts/evidence/`, atomico per file, mantiene la gerarchia per dominio.
- `diff(source, candidate) -> str`: diff markdown per il gate.
### 4. `harness/tht/vectorstore/records.py` (modifica)
- `evidence_records()`: metadata ora include `kind`, `applies_to`, `anchors` (non solo
`status/tier/tables/concepts`).
- Il `content` di ogni record resta title+body, ma per file con sezioni canoniche si
indicizza anche un record per sezione (id `evidence:<id>:<sezione>`) così il retrieval
può restringere per sezione oltre che per documento.
### 5. `harness/tht/search/__init__.py` + `cli/search_cmd.py` (modifica)
- `SearchResult` e `combined_search` accettano `applies_to` come filtro metadata.
- `tht search pack`: la sezione "Evidence rilevanti" ora proietta `kind` + sezione
canonica pertinente (non solo excerpt generico); usa `applies_to` per non mischiare
destinazioni.
- `tht search find --kind evidence --applies-to <dest>` per il retrieval mirato.
### 6. `harness/tht/cli/evidence_cmd.py` (modifica)
- Nuovi comandi:
- `tht evidence lint --source <path|dir>` — diagnostica deterministica.
- `tht evidence canonicalize plan --source-root <dir> --json` — piano di
(ri)canonizzazione.
- `tht evidence canonicalize apply --plan <file> --approved <ids> --json` — applica.
- `--json` sempre pristine (solo JSON su stdout), come da contratto di progetto.
### 7. `harness/.pi/extensions/tht-gate.js` (modifica)
- Nuovo tool `reviewer_evidence_batch`: riceve il piano + i diff aggregati e presenta un
multiselect per file (approva/rifiuta/richiedi modifica). Le scelte approvate vengono
persistite tramite `tht evidence canonicalize apply`.
- Aggiungere `tht evidence canonicalize apply` alla FORBIDDEN anti-bypass list (il
modello non può applicare canonizzazioni senza review).
### 8. Sessione di manutenzione Pi (nuova skill o sezione SKILL)
- Una skill dedicata `tht-evidence-canonicalize` (o un comando slash nel gate) descrive
la sessione di manutenzione: il modello legge i sorgenti marcati dal piano, propone i
campi semantici (`kind`, `applies_to`, `concepts` con sinonimi, `tables` con ruolo,
`anchors`) e li invia al `reviewer_evidence_batch`.
- Il gate è l'unico canale di review; il modello non scrive mai direttamente
`artifacts/evidence/`.
### 9. `harness/.pi/skills/tht-sessione/SKILL.md` (modifica)
- Aggiornare i punti F1/F4/F6-F7 dove il modello legge evidence: spiegare che l'evidence
canonica espone `kind`/`applies_to`/`anchors`, che gli `anchors` sono la fonte primaria
per il value-grounding in F4, e che le sezioni canoniche vanno citate per destinazione.
- Aggiungere il nuovo tool `reviewer_evidence_batch` alla lista dei tool disponibili.
## Migrazione del corpus (incrementale, v1+v2 convivono)
- `CanonicalEvidence` (v2) convive con `EvidenceDoc` (v1): `lint` segnala i non-canonici
ma nulla si rompe; `load_evidence_dir` carica entrambi.
- Convertire a lotti, partendo da `20-valori-enum` e `10-domini-clinici` (impatto
maggiore su F1/F4, anchors più ricchi).
- Il canonizzato derivato in `artifacts/evidence/` sostituisce progressivamente il file
sorgente mirrorato; finché un sorgente non è canonizzato, `extract` continua a
rispecchiare il v1 come oggi.
## Ordine di implementazione
1. `CanonicalEvidence` + `lint` (TDD, nessun LLM).
2. `canonicalize.py` (plan/apply/diff) + comandi CLI (TDD).
3. `records.py` + filtro `applies_to` in `search` + proiezione pack (TDD).
4. Gate `reviewer_evidence_batch` + skill manutenzione Pi.
5. Aggiornamento `SKILL.md` tht-sessione.
6. Conversione primo lotto (`enum` + `domain`) con review batch.
## Verifica
- `cd harness && .venv/bin/pytest -q` — test nuovi per model/lint/canonicalize/records/search.
- `node --test` sul gate per `reviewer_evidence_batch` e anti-bypass.
- Su un campione del corpus: `tht evidence lint` riporta i file rotti (frontmatter `--`,
bullet Unicode) senza crash; `tht evidence canonicalize plan --json` produce piano
corretto; `apply` con fingerprint invariato è no-op.
- `tht search pack "<domanda ablazione>"` mostra evidence con `kind`/sezione pertinente e
gli anchor presenti; `tht search find --kind evidence --applies-to schema_linking`
restituisce solo evidence pertinenti.
- Live: una sessione F4 su psd usa un anchor canonico per il value-grounding invece del
solo LSH (verificabile dal `reviewer_decide` con opzioni `value_grounded` provenienti
dall'evidence, non dal ranking LSH).
- `git diff --check` e typecheck/ruff puliti.
@@ -1,6 +1,6 @@
# ThothII Authentication Acceptance and PSD Deployment Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans to execute this plan task-by-task.
> **For agentic workers:** execute one task at a time, record its completion criteria, and stop at every documented approval boundary.
**Goal:** Validate local and OIDC authentication on macOS, deploy the exact feat/thoth-auth candidate to the Aritmolab/PSD server before merging it into main, and complete end-to-end acceptance with remote Authentik.
@@ -1,92 +0,0 @@
# Tht Documentation Convergence Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
**Goal:** Align current documentation and documentation smoke checks with the converged native host CLI `tht`, while preserving historical references only where they describe past decisions or evidence.
**Architecture:** Treat `tools/tht/cmd/tht/main.go` as the canonical host CLI surface for installation, authentication, diagnostics, lifecycle, and workspace operations. Keep the Python `harness/.venv/bin/tht` distinction explicit for the workflow runtime, and update current operator/test instructions to invoke the native `tht` with `--installation`.
**Tech Stack:** Markdown documentation, shell smoke tests, Go CLI command surface, repository search-based verification.
---
### Task 1: Classify current and historical legacy CLI references
**Files:**
- Inspect: `README.md`, `PROJECT_STATE.md`, `AGENTS.md`, `docs/**`, `scripts/**`
- Reference: `tools/tht/cmd/tht/main.go`
**Step 1:** Build a complete occurrence inventory with a case-insensitive search for the former host CLI name and classify every match.
**Step 2:** Classify each occurrence as current operator documentation, documentation smoke expectation, executable/script contract, or historical design/evidence.
**Step 3:** Record the classification in the implementation notes before editing.
### Task 2: Update canonical operator and installation documentation
**Files:**
- Modify: `README.md`
- Modify: `AGENTS.md`
- Modify: `PROJECT_STATE.md`
- Modify: `docs/guida-utente.md`
- Modify: `docs/contracts/workspace-preprocessing-cli.md`
- Rename/update: `docs/contracts/tht-pi.md` as the current `tht` Pi contract
- Modify: relevant installation and architecture pages that expose operator commands
**Step 1:** Replace current host/operator invocations with `tht --installation ...`.
**Step 2:** Document the distinction between the native host CLI `tht` and the Python harness CLI invoked by the backend/runtime.
**Step 3:** Update command examples for `start`, `status`, `doctor`, `auth`, `workspace`, and `pi`.
**Step 4:** Add a short historical note only where a document must explain the former name.
### Task 3: Rewrite authentication acceptance and manual test instructions
**Files:**
- Modify: `docs/testing/authentication-manual-acceptance.md`
- Modify: `docs/plans/2026-08-18-thothii-authentication-acceptance-and-psd-deployment.md`
- Modify: `docs/install/authentication-local.md`
- Modify: `docs/install/authentication-oidc.md`
- Modify: `docs/install/authentik.md`
**Step 1:** Make `tht auth status`, `tht auth check`, `tht auth check --interactive`, and `tht doctor --json` the canonical terminal preflight.
**Step 2:** Use `tht status`, `tht start`, and `tht workspace inspect --workspace psd-clinical --json` for PSD deployment checks.
**Step 3:** Clarify that the P8 L2 gate is authentication-to-application integration through the first reviewer gate.
**Step 4:** Retain the prior functional test suite as a baseline and add only the authentication boundary smoke required for this acceptance.
### Task 4: Align documentation smoke tests
**Files:**
- Modify: `scripts/auth-docs-smoke.sh`
- Modify: `scripts/test-auth-docs-smoke.sh`
- Inspect/update: any current smoke script whose user-facing command examples still require the legacy CLI name
**Step 1:** Replace forbidden/current command assertions with `tht` equivalents.
**Step 2:** Preserve negative checks for obsolete authentication CLI wording.
**Step 3:** Run the positive and negative documentation fixtures.
### Task 5: Preserve or annotate historical material
**Files:**
- Inspect the historical discovery specification for context, without treating it as current operator documentation.
- Inspect: dated reports and archived acceptance scripts
**Step 1:** Do not rewrite historical titles, commit evidence, or old implementation names solely to erase history.
**Step 2:** Add a concise “historical nomenclature” note where an archived document could otherwise be mistaken for current instructions.
### Task 6: Verify the convergence
**Step 1:** Run `scripts/auth-docs-smoke.sh` and `scripts/test-auth-docs-smoke.sh`.
**Step 2:** Search active documentation for remaining legacy CLI references.
**Step 3:** Confirm every remaining match is either an explicit historical note, an ignored runtime directory name, or a non-document executable compatibility artifact.
**Step 4:** Run `git diff --check` and report the exact files changed plus any intentionally retained historical references.
@@ -1,6 +1,6 @@
# PSD Server Deployment Program Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
> **For agentic workers:** execute one task at a time, record its completion criteria, and stop at every documented approval boundary.
**Goal:** Replace the legacy PSD ThothII installation, prove the replacement with local authentication, and then integrate the accepted release with Supabase, Authentik, Nginx, the load balancer, and the Aritmolab sidebar.
@@ -1,6 +1,6 @@
# PSD Server Project A Standalone Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
> **For agentic workers:** execute one task at a time, record its completion criteria, and stop at every documented approval boundary.
**Goal:** Install a clean PSD ThothII stack with local authentication, direct read-only DWH access, internal Qdrant/Ollama, rebuilt preprocessing, and one completed F1-F8 work session.
@@ -1,6 +1,6 @@
# PSD Server Project B Authentik Integration Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
> **For agentic workers:** execute one task at a time, record its completion criteria, and stop at every documented approval boundary.
**Goal:** Convert the accepted Project A installation to public OIDC mode, store owned work sessions in the existing Supabase database's `thoth_sessions` schema, and restore the established Aritmolab-sidebar user journey through the load balancer and Nginx.
+1 -1
View File
@@ -1,6 +1,6 @@
# PSD Server Survey Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
> **For agentic workers:** execute one task at a time, record its completion criteria, and stop at every documented approval boundary.
**Goal:** Produce a non-mutating, redacted survey of the PSD server that resolves every path, owner, network boundary, credential location, and rollback prerequisite needed by Projects A and B.
@@ -1,7 +1,7 @@
# Ristrutturazione delle Evidence — disegno approvato
**Stato:** approvato il 24 agosto 2026
**Sostituisce:** `docs/plans/2026-08-18-evidence-canonica-design.md`
**Sostituisce:** il precedente disegno di Evidence canonica, disponibile nella storia Git
**Ambito:** authoring, revisione, pubblicazione, indicizzazione e uso runtime delle Evidence
## 1. Obiettivo
File diff suppressed because it is too large Load Diff
@@ -1,11 +1,11 @@
# PRD — Preprocessing per-workspace su ThothII (Qdrant + Git workspace registry)
**Status:** PRD in revisione — decisioni D1–D9 chiuse il 2026-08-09; i piani P1–P10 partiranno solo dopo
revisione e conferma del proprietario
**Status:** baseline storica dei requisiti — implementazione completata; per i contratti correnti vedere
`docs/contracts/workspace-preprocessing-cli.md`, `docs/contracts/workspace-evidence-v3.md` e
`docs/evidence.md`
**Data:** 2026-08-09
**Autore:** analisi dello stato attuale (branch `codex/git-workspace-registry`) + decisioni con il proprietario
**Uso:** riferimento stabile di requisiti e decisioni; da ogni punto nascerà un piano separato in
`docs/superpowers/plans/` (sez. 11) — questo documento non è un piano di lavoro
**Uso:** riferimento stabile delle decisioni originarie; questo documento non è un piano operativo
---
@@ -14,8 +14,8 @@ revisione e conferma del proprietario
> descriptor `<id>/workspace.yaml`, evidence embedded `<id>/evidence/`, annotazioni FK curate
> `<id>/schema/annotations.yaml` (P5), docs generate `workspace-docs/<id>/`. I vecchi percorsi
> piatti (`workspaces/<id>.yaml`, `workspace-content/<id>/evidence/`) sono superseded; le uniche
> occorrenze rimaste sono storiche (changelog/revisioni). Vedi
> `docs/superpowers/plans/2026-08-11-p2-p6-adaptation-to-p1-1-registry.md`.
> occorrenze rimaste sono storiche (changelog/revisioni). Il contratto corrente è descritto in
> `docs/contracts/workspace-evidence-v3.md`.
---
@@ -34,10 +34,9 @@ scrittura via REST dedicato o loading diretto) a un'architettura con:
`resources.embeddings` + `roots` sotto `<dataRoot>/sessions/<wsId>/`).
La **macchina di preprocessing** (comandi, job a generazioni con publish atomico, adapter Qdrant, corpus
evidence, FK, memory) **esiste già ed è testata**: `tht preprocess evidence|dwh`, `tht vector init|index-schema`,
`tht schema introspect|suggest-fks|check`, `tht evidence extract|index`, `tht lsh build`, memory/solved.
Esistono job Compose fixture (`deploy/compose.preprocess.yaml` + `deploy/workspaces/preprocess-{dwh,evidence}.yaml`)
e uno smoke (`scripts/preprocess-smoke.sh`).
evidence, FK, memory) **esiste ed è testata**. La superficie operativa corrente è il comando host
`tht --installation ... workspace preprocess ...`, che esegue il servizio profile-gated
`workspace-maintenance`; le vecchie fixture Compose dedicate sono state ritirate.
### Il problema
@@ -210,8 +209,7 @@ documentato e verificato da smoke end-to-end.
verifica → uso → backup/restore).
- RF8.2 La documentazione di progetto spiega **cos'è `.tht-dwh`** (generazioni, `OWNER.json`, `ACTIVE`,
vincolo di fingerprint) in modo comprensibile per l'operatore (D3).
- RF8.3 Smoke end-to-end automatico (workspace nuovo → tutto il ciclo → sessione reale → cleanup) che
sostituisce/completa `preprocess-smoke.sh` (oggi solo fixture).
- RF8.3 Smoke end-to-end automatico (workspace nuovo → tutto il ciclo → sessione reale → cleanup).
- RF8.4 Backup/restore coprono Qdrant (già `vector-backup.sh`/`vector-restore.sh`), corpus, `.tht-dwh` e
registry.
- RF8.5 Ogni piano successivo traduce il proprio risultato operativo in un **process goal** verificabile da
@@ -332,7 +330,7 @@ Ogni piano tecnico riporta, adattandoli al proprio scope:
keyword-index.
8. Rerun del preprocessing: `unchanged` (nessun duplicato); modifica di un'evidence → nuova generazione,
ACTIVE aggiornato, vecchie generazioni in GC.
9. Smoke end-to-end automatico verde in CI con cleanup esatto (stile `preprocess-smoke.sh`).
9. Smoke end-to-end automatico verde in CI con cleanup esatto.
10. Migrazione PSD documentata e provata almeno in dry-run (re-introspection o riuso catalogo + re-embedding).
11. Ogni piano tecnico successivo include un process goal automatico completo per il proprio scope e un
walkthrough manuale quando utile, oppure documenta l'inevitabile eccezione umana secondo S4.
@@ -417,12 +415,11 @@ Ogni piano tecnico riporta, adattandoli al proprio scope:
- Default invariati (`retain_published_generations: 3`, chunk 4000 char), configurabili per-workspace via
la sezione `evidence`/policy del descriptor (D1).
## 11. Mappa dei piani (uno per punto del PRD)
## 11. Mappa storica dell'implementazione
Questo PRD non diventa un unico piano: **ogni decisione/requisito produce un piano separato (P1–P10)** in
`docs/superpowers/plans/`, eseguibile in sequenza o come workstream indipendenti. Il PRD resta il
riferimento stabile (requisiti + decisioni); ogni piano cita il punto di origine e i criteri di
accettazione applicabili (sez. 9) e adotta lo standard integration-first (sez. 8).
L'implementazione è stata suddivisa nei workstream P1–P10 riportati sotto. I piani esecutivi
superati sono disponibili nella storia Git; questa tabella conserva soltanto la relazione tra
requisiti, dipendenze e risultati attesi.
| Piano | Punto PRD | Contenuto sintetico | Dipende da |
| --- | --- | --- | --- |
@@ -532,12 +529,11 @@ traccia separatamente implementazione, automated integration e manual acceptance
- Stato attuale: `PROJECT_STATE.md` (sezioni "Internal Qdrant + Ollama semantic infrastructure", snapshot
registry) e `AGENTS.md`.
- Design architettura semantica: `docs/plans/2026-08-08-internal-qdrant-ollama-design.md` e relativo piano.
- Registry: `docs/superpowers/specs/2026-08-03-git-workspace-registry-design.md`, manuali
- Registry e Evidence: `docs/contracts/workspace-evidence-v3.md`, `docs/evidence.md`, manuali
`docs/install/local-workspace-registry.md` / `server-workspace-registry.md`.
- Motore preprocessing: `harness/tht/cli/preprocess_cmd.py`, `harness/tht/corpus/pipeline.py`,
`harness/tht/jobs/dwh_pipeline.py`, `harness/tht/adapters/vector/qdrant.py`,
`harness/tht/vectorstore/records.py`, `harness/tht/cli/{vector,schema,evidence,memory}_cmd.py`.
- Fixture attuali: `deploy/compose.preprocess.yaml`, `deploy/workspaces/preprocess-{dwh,evidence}.yaml`,
`scripts/preprocess-smoke.sh`.
- Superficie operativa: `tools/tht/` e `docs/contracts/workspace-preprocessing-cli.md`.
- Ammissione runtime: `backend/src/tht/tht-runner.ts` (`qdrantEnsure`/`ollamaEnsure`),
`backend/src/workspaces/runtime-renderer.ts`.
@@ -1,171 +0,0 @@
# Audit critico dei comandi `tht`
Data: 2026-08-15
## Scopo
Questo audit valuta tutti i comandi terminali registrati dall'attuale CLI Python `tht` prima di
unificare la CLI di ThothII sotto un solo eseguibile pubblico. La valutazione incrocia:
- il contratto del workflow Pi in `harness/.pi/skills/tht-sessione/SKILL.md`;
- le invocazioni reali del gate in `harness/.pi/extensions/tht-gate.js`;
- le invocazioni del backend in `backend/src/tht/tht-runner.ts`;
- i job operatore in `backend/src/workspaces/preprocessing-service.ts`;
- il migratore in `docker/session-migrate.sh`;
- test e documentazione esistenti.
L'inventario autorevole contiene 76 comandi Typer più il comando callback `doctor`: 77 comandi
terminali complessivi.
## Legenda
- **WF — intoccabile workflow**: chiamato da Pi, dal gate o dal contratto delle otto fasi. Va
conservato con semantica, output JSON ed exit code compatibili. Non deve necessariamente apparire
nell'help ordinario dell'utente.
- **PL — intoccabile piattaforma**: chiamato dal backend, dai job workspace o dal deployment. Anche
questo è un contratto interno, non necessariamente un comando da mostrare all'utente.
- **ADV — mantenere avanzato**: non è nel flusso automatico, ma offre una capacità amministrativa o
di recupero che sarebbe imprudente perdere. Va nascosto dall'help base.
- **ACCORPA**: la capacità serve, ma non merita un comando autonomo.
- **RIMUOVI**: il comando non ha chiamanti reali ed è duplicato, superato, pericoloso o incompleto.
L'eventuale logica riutilizzata da altri flussi resta una libreria interna.
## Risultato sintetico
| Esito | Numero | Conseguenza |
|---|---:|---|
| WF o PL, intoccabili | 55 | Conservare il contratto; nascondere i primitivi tecnici dall'help base |
| ADV o ACCORPA | 8 | Conservare la capacità riducendo la superficie UX |
| RIMUOVI | 14 | Eliminare il comando dalla nuova CLI |
| **Totale** | **77** | Una sola CLI pubblica molto più semplice, senza riscrivere il workflow vivo |
## Matrice completa
### Diagnostica, configurazione e dipendenze
| Comando | Valutazione |
|---|---|
| `config check` | **ACCORPA** in `tht doctor`: la validazione della configurazione serve, ma due preflight distinti confondono l'utente. |
| `doctor` | **ACCORPA/MANTIENI pubblico** come unico `tht doctor`, includendo controlli host, Compose, storage e configurazione runtime. |
| `db ping` | **PL**: il backend lo usa per rifiutare correttamente una nuova sessione quando il DWH non è raggiungibile o non è read-only. Interno. |
| `db fetch-ca` | **ACCORPA** in `tht setup` o nella configurazione workspace: utile per TLS, ma non giustifica un comando isolato. |
| `ollama ensure` | **PL**: preflight automatico dell'embedder usato dal backend. Interno. |
### Fasi e decision ledger
| Comando | Valutazione |
|---|---|
| `phase advance` | **WF**: il gate lo usa per avanzare solo dopo la decisione umana. Primitivo anti-bypass, quindi interno. |
| `phase meta` | **WF**: fornisce al gate la definizione data-driven delle fasi e dei tipi di decisione. Interno. |
| `phase reopen` | **WF**: è il percorso canonico per tornare a una fase precedente e invalidare deterministicamente gli artefatti successivi. |
| `phase show` | **WF**: il gate lo usa per calcolare la fase corrente. Interno. |
| `decision add` | **WF**: persistenza fondamentale delle decisioni del reviewer. Solo gate, non shell utente. |
| `decision add-batch` | **WF**: scrittura atomica delle decisioni multiple. Evita ledger parziali. |
| `decision add-join-set` | **WF**: sostituzione atomica dell'intero insieme di join. |
| `decision list` | **RIMUOVI**: nessun chiamante; `session show --json` contiene già il ledger necessario. |
| `decision retract` | **RIMUOVI** dalla CLI: nessun flusso vivo lo invoca e `phase reopen` è il percorso di correzione supportato. La semantica tombstone può restare nel dominio finché utile. |
### Sessioni
| Comando | Valutazione |
|---|---|
| `session archive` | **PL**: usato dalla gestione sessioni del backend. |
| `session check` | **WF**: gate oggettivo della fase 5; verifica decisioni e schema linking. |
| `session close` | **PL**: usato dal backend. |
| `session delete` | **PL**: usato dal backend con i relativi controlli applicativi. |
| `session documents` | **WF/PL**: ricostruisce il contesto persistito e alimenta sia Pi sia la GUI. |
| `session fail` | **PL**: usato dal backend per rappresentare il fallimento terminale. |
| `session finalize` | **WF**: chiusura deterministica della fase finale e indicizzazione della domanda risolta. |
| `session list` | **PL**: alimenta la lista sessioni della GUI. |
| `session migrate` | **PL**: eseguito dal servizio one-shot di migrazione server; resta interno dietro `tht sessions migrate`. |
| `session new` | **WF/PL**: crea la persistenza iniziale della domanda; il backend dipende dal JSON restituito. |
| `session preferences get` | **PL**: lettura delle preferenze applicative. Interno. |
| `session preferences set` | **PL**: scrittura delle preferenze applicative. Interno. |
| `session reopen` | **PL**: riapertura dello stato terminale esposta dalla gestione sessioni. |
| `session retrieval-pack` | **WF**: legge il retrieval pack già persistito per il kickoff di Pi. Distinto da `search pack`, che lo costruisce. |
| `session set-group` | **PL**: rinomina il raggruppamento dalla GUI. |
| `session set-name` | **PL**: rinomina la sessione dalla GUI. |
| `session set-question` | **WF**: persiste deterministicamente domanda riscritta e assunzioni. Solo gate. |
| `session set-schema-linking` | **WF**: valida e scrive `schema_linking.json`. Solo gate. |
| `session show` | **WF/PL**: fonte compatta dello stato persistito per resume, gate e backend. |
| `session sync-schema-linking` | **WF**: riproietta deterministicamente il ledger nello schema linking. |
| `session unarchive` | **PL**: usato dalla gestione sessioni del backend. |
### Schema e retrieval
| Comando | Valutazione |
|---|---|
| `schema check` | **PL**: validazione delle annotazioni curate nel workflow workspace. |
| `schema columns` | **WF**: il gate usa il catalogo colonne per validare e correggere il linking. |
| `schema introspect` | **WF**: fallback previsto dal contratto quando manca lo schema fisico; la modalità refresh resta manutenzione. |
| `schema render` | **WF**: produce il contesto mschema usato dal modello. |
| `schema suggest-fks` | **PL**: comando del flusso operatore per le annotazioni FK curate. |
| `search find` | **WF**: ricerca mirata di evidence, valori e formule durante le fasi. |
| `search pack` | **WF/PL**: costruisce e persiste il contesto iniziale F1; usato anche dal backend. |
### CTE, SQL e datamart
| Comando | Valutazione |
|---|---|
| `cte info` | **WF**: restituisce SQL persistito, posizione nel piano e ultimo test. |
| `cte list` | **RIMUOVI**: nessun chiamante o test; `cte plan`, `cte info` e `session documents` coprono il bisogno. |
| `cte next` | **WF**: il gate determina il prossimo CTE da revisionare. |
| `cte plan` | **WF**: persiste l'ordine completo dei CTE. |
| `cte save` | **WF**: tool deterministico di scrittura usato dal gate. |
| `cte test` | **WF**: verifica read-only dei CTE prevista esplicitamente dal contratto. |
| `sql validate` | **WF**: validazione strutturale e read-only prima dell'esecuzione. |
| `sql preview` | **WF/PL**: preview controllata usata dal modello e dalla GUI. |
| `sql set-final` | **WF**: unica scrittura canonica di `sql_final.sql` attraverso il repository di sessione. |
| `sql export` | **PL**: esportazione richiesta dalla GUI. |
| `sql explain` | **RIMUOVI**: nessun chiamante, test o requisito nel workflow corrente. Si reintroduce solo con un vero passo di analisi del piano. |
| `sql save` | **RIMUOVI**: duplica `set-final` ed `export` e permette un percorso di scrittura non usato. |
| `datamart generate` | **WF**: fase 8 del workflow. |
### Memory
| Comando | Valutazione |
|---|---|
| `memory promote` | **WF**: preview dei candidati di promozione usata dal gate. |
| `memory save-one` | **WF**: persistenza atomica della singola memory approvata. |
| `memory search` | **WF**: recupero delle memory riutilizzabili nella fase 2. |
| `memory solved-index` | **WF**: recupero manuale previsto se l'indicizzazione al finalize fallisce. |
| `memory solved-search` | **WF**: recupero di domande risolte simili nelle fasi successive. |
| `memory list` | **ADV**: mantenere per amministrare record errati, ma fuori dall'help base. |
| `memory show` | **ADV**: mantenere insieme a `list` per ispezione puntuale. |
| `memory update` | **ADV**: mantenere per correggere il merito di una memory senza alterarne la provenienza. |
| `memory delete` | **ADV**: mantenere come rimedio selettivo; richiede conferma esplicita nella nuova CLI. |
| `memory index` | **ADV**: utile come riparazione/full-resync, ma va presentato come manutenzione e non come uso normale. |
| `memory clear` | **RIMUOVI**: distruzione globale non usata; confligge con una UX sicura di backup/ripristino. |
| `memory migrate` | **RIMUOVI**: migrazione legacy una tantum senza dati di produzione da preservare. |
### Preprocessing, evidence e indici
| Comando | Valutazione |
|---|---|
| `preprocess dwh` | **PL**: pipeline canonica usata da `tht workspace preprocess dwh/run`. |
| `preprocess evidence` | **PL**: pipeline canonica usata da `tht workspace preprocess evidence/run`. |
| `vector index-schema` | **PL**: indicizzazione schema usata dal workflow workspace. |
| `evidence extract` | **RIMUOVI**: primitivo superato dalla pipeline versionata `preprocess evidence`. Conservare soltanto la logica riusata. |
| `evidence index` | **RIMUOVI**: primitivo superato dalla stessa pipeline versionata. |
| `lsh build` | **RIMUOVI** come comando: è già uno step di `preprocess dwh`; il builder resta interno. |
| `lsh query` | **RIMUOVI**: probe visuale senza chiamanti, test o documentazione operativa. La ricerca applicativa passa da `search find`. |
| `vector init` | **RIMUOVI**: il controllo di Qdrant/embedder è ormai coperto dal reconciler di collezione, da `ollama ensure` e dal nuovo `tht doctor`. |
### Formule di concetto
| Comando | Valutazione |
|---|---|
| `formula save` | **RIMUOVI** dalla CLI corrente: nessun chiamante, test o flusso di approvazione lo usa. Conservare il formato/store e la lettura tramite `search find --kind formula`. |
| `formula list` | **RIMUOVI**: stesso sottosistema incompleto. Un futuro flusso di curation dovrà progettare insieme creazione, approvazione, elenco e modifica. |
## Conseguenza per la nuova CLI unica
La semplificazione migliore non consiste nel rinominare tutti i 55 contratti vivi o nel mostrarli
all'utente. Consiste nel mantenere un unico eseguibile `tht` con due livelli di visibilità:
1. l'help ordinario mostra soltanto setup, lifecycle, backup/restore, Pi e workspace;
2. i contratti WF/PL restano invocabili dallo stesso eseguibile, ma sono interni/nascosti e usati da
backend, gate e job one-shot.
In questo modo l'utente vede una CLI piccola, mentre il workflow non subisce una riscrittura inutile
e rischiosa. Non serve un secondo eseguibile né un alias `thothctl`.
@@ -1,148 +0,0 @@
# Proposta maintain-erase-enhance per i comandi `tht`
Data: 2026-08-15
## Criterio
- **MAINTAIN**: il comando resta disponibile senza modifiche sostanziali. Come richiesto, non viene
aggiunta una motivazione.
- **ERASE**: il comando viene eliminato dalla nuova CLI; la motivazione indica la duplicazione, il
superamento o l'assenza di un utilizzo reale.
- **ENHANCE**: la capacità viene mantenuta, ma il comando viene migliorato, accorpato o reso più
sicuro. La proposta indica l'intervento.
La proposta copre tutti i 77 comandi terminali dell'attuale CLI Python.
## Sintesi
| Proposta | Numero |
|---|---:|
| MAINTAIN | 55 |
| ENHANCE | 8 |
| ERASE | 14 |
| **Totale** | **77** |
## Lista completa
### Diagnostica, configurazione e dipendenze
| Comando | Proposta |
|---|---|
| `config check` | **ENHANCE** — incorporare la validazione nel comando pubblico `tht doctor`, mantenendo una funzione interna riutilizzabile e l'output strutturato. Evita due preflight sovrapposti. |
| `doctor` | **ENHANCE** — farne l'unica diagnostica multilivello: installazione, descriptor, Compose, storage, configurazione runtime, DWH, Pi, Qdrant ed embedder. Deve offrire output umano e `--json`, senza mutare lo stato. |
| `db ping` | **MAINTAIN** |
| `db fetch-ca` | **ENHANCE** — integrarlo nel setup guidato del workspace, mostrando endpoint e fingerprint prima della conferma. Può restare disponibile come operazione TLS avanzata, ma non come passaggio manuale obbligatorio. |
| `ollama ensure` | **MAINTAIN** |
### Fasi e decision ledger
| Comando | Proposta |
|---|---|
| `phase advance` | **MAINTAIN** |
| `phase meta` | **MAINTAIN** |
| `phase reopen` | **MAINTAIN** |
| `phase show` | **MAINTAIN** |
| `decision add` | **MAINTAIN** |
| `decision add-batch` | **MAINTAIN** |
| `decision add-join-set` | **MAINTAIN** |
| `decision list` | **ERASE** — non ha chiamanti reali e duplica il ledger già restituito da `session show --json`. |
| `decision retract` | **ERASE** — non è invocato dal workflow corrente; `phase reopen` è il percorso supportato per correggere e invalidare deterministicamente le decisioni. La semantica tombstone può restare nel dominio. |
### Sessioni
| Comando | Proposta |
|---|---|
| `session archive` | **MAINTAIN** |
| `session check` | **MAINTAIN** |
| `session close` | **MAINTAIN** |
| `session delete` | **MAINTAIN** |
| `session documents` | **MAINTAIN** |
| `session fail` | **MAINTAIN** |
| `session finalize` | **MAINTAIN** |
| `session list` | **MAINTAIN** |
| `session migrate` | **MAINTAIN** |
| `session new` | **MAINTAIN** |
| `session preferences get` | **MAINTAIN** |
| `session preferences set` | **MAINTAIN** |
| `session reopen` | **MAINTAIN** |
| `session retrieval-pack` | **MAINTAIN** |
| `session set-group` | **MAINTAIN** |
| `session set-name` | **MAINTAIN** |
| `session set-question` | **MAINTAIN** |
| `session set-schema-linking` | **MAINTAIN** |
| `session show` | **MAINTAIN** |
| `session sync-schema-linking` | **MAINTAIN** |
| `session unarchive` | **MAINTAIN** |
### Schema e retrieval
| Comando | Proposta |
|---|---|
| `schema check` | **MAINTAIN** |
| `schema columns` | **MAINTAIN** |
| `schema introspect` | **MAINTAIN** |
| `schema render` | **MAINTAIN** |
| `schema suggest-fks` | **MAINTAIN** |
| `search find` | **MAINTAIN** |
| `search pack` | **MAINTAIN** |
### CTE, SQL e datamart
| Comando | Proposta |
|---|---|
| `cte info` | **MAINTAIN** |
| `cte list` | **ERASE** — non ha chiamanti o test e sovrappone informazioni già disponibili con `cte plan`, `cte info` e `session documents`. |
| `cte next` | **MAINTAIN** |
| `cte plan` | **MAINTAIN** |
| `cte save` | **MAINTAIN** |
| `cte test` | **MAINTAIN** |
| `sql validate` | **MAINTAIN** |
| `sql preview` | **MAINTAIN** |
| `sql set-final` | **MAINTAIN** |
| `sql export` | **MAINTAIN** |
| `sql explain` | **ERASE** — non è usato né testato dal workflow attuale. Va reintrodotto soltanto se l'analisi del piano diventa un passo esplicito del processo. |
| `sql save` | **ERASE** — duplica `sql set-final` e `sql export` e introduce un percorso di scrittura non utilizzato. |
| `datamart generate` | **MAINTAIN** |
### Memory
| Comando | Proposta |
|---|---|
| `memory promote` | **MAINTAIN** |
| `memory save-one` | **MAINTAIN** |
| `memory search` | **MAINTAIN** |
| `memory solved-index` | **MAINTAIN** |
| `memory solved-search` | **MAINTAIN** |
| `memory list` | **ENHANCE** — trasformarlo in una vista amministrativa paginata, con filtri, provenienza, stato e output `--json`; non mostrarlo nell'help base. |
| `memory show` | **ENHANCE** — mostrare provenienza immutabile, decisione sorgente, stato dell'indice e riferimenti necessari a una correzione consapevole. |
| `memory update` | **ENHANCE** — limitare l'aggiornamento ai campi modificabili, mostrare un diff prima della conferma e impedire modifiche alla provenienza. |
| `memory delete` | **ENHANCE** — richiedere identificatore esatto e conferma esplicita, mostrare l'impatto e verificare la rimozione coerente da registro e indice. |
| `memory index` | **ENHANCE** — riposizionarlo come comando di repair: prima rileva il drift, poi ricostruisce soltanto con conferma e verifica finale. Non deve sembrare un'operazione ordinaria. |
| `memory clear` | **ERASE** — cancellazione globale non usata e troppo facile da eseguire per errore; backup/ripristino e cancellazione selettiva sono percorsi più sicuri. |
| `memory migrate` | **ERASE** — migrazione legacy una tantum; non esistono dati di produzione da preservare e la nuova architettura può partire direttamente dal formato corrente. |
### Preprocessing, evidence e indici
| Comando | Proposta |
|---|---|
| `preprocess dwh` | **MAINTAIN** |
| `preprocess evidence` | **MAINTAIN** |
| `vector index-schema` | **MAINTAIN** |
| `evidence extract` | **ERASE** — è un primitivo superato dalla pipeline versionata `preprocess evidence`; l'eventuale logica condivisa resta interna. |
| `evidence index` | **ERASE** — è un secondo primitivo superato dalla stessa pipeline, che già gestisce materializzazione, indicizzazione, versionamento e resume. |
| `lsh build` | **ERASE** — la costruzione LSH è già uno step di `preprocess dwh`; mantenere due ingressi permette esecuzioni parziali incoerenti. |
| `lsh query` | **ERASE** — probe visuale senza chiamanti, test o documentazione operativa; il workflow usa `search find`. |
| `vector init` | **ERASE** — il controllo di Qdrant ed embedder è già coperto dal reconciler della collezione, da `ollama ensure` e dal nuovo `tht doctor`. |
### Formule di concetto
| Comando | Proposta |
|---|---|
| `formula save` | **ERASE** — non ha chiamanti, test o un flusso di approvazione completo. Il formato e lo store possono restare disponibili alla ricerca finché non viene progettata una vera curation. |
| `formula list` | **ERASE** — appartiene allo stesso sottosistema incompleto; un futuro flusso deve progettare insieme creazione, approvazione, elenco, modifica e cancellazione. |
## Impatto sulla UX
I 55 comandi `MAINTAIN` comprendono molti contratti macchina intoccabili. Mantenerli non implica
mostrarli tutti nell'help principale. La futura CLI unica può conservare gli stessi percorsi per
backend, gate e job, mostrando all'utente soltanto i gruppi operativi di primo livello.
-126
View File
@@ -1,126 +0,0 @@
# L2 Run Report — 2026-06-27 (sessione cardioversione + ablazione)
> Esito della prima sessione L2 end-to-end dopo il porting CLI+skill (Onda -1→4 +
> Skill + 0b). Sessione non-deterministica, esito informativo non bloccante per il
> "done" del porting codice (come da piano L2.1, riga 1187).
## Setup al momento del run
- Pi: `@earendil-works/pi-coding-agent`, provider `zai`, model `glm-5.2` (default).
- Harness: Onda -1→4 + Skill committate; Onda 0b (workspace cliente `tht-workspace-psd`,
indice LSH 75737 valori, evidence 35).
- `.env` popolato, VPN OK, DWH REST 200, Ollama UP.
- Pre-run fix applicati in questa sessione: `_YamlModel.to_yaml` (bloccava `session new`),
`config/tht.yaml` symlink al workspace cliente (il gate chiama `tht` senza `-c`).
## Come è partita la sessione
Lancio `pi --mode rpc` + `/nuova-domanda "<cardioversione + ablazione same-year>"`.
**Nota critica su `--mode rpc`:** la TUI interattiva di Pi (`pi` senza `--mode`) **non
renderizza** i widget `extension_ui_request` del gate (gestiti solo in `modes/rpc/`).
La modalità RPC emette i widget come JSONL su stdio per un client esterno — che non
esiste ancora in ThothII. Il run è stato possibile solo perché il modello, non vedendo
UI, ha operato via shell/tool fino al blocco fatale (vedi bug #4).
## Cosa ha fatto il modello (transcript: 117 eventi, 365KB)
Sessione Pi: `~/.pi/agent/sessions/--Users-mp-projects-ThothII-harness--/2026-06-27T13-44-51...jsonl`.
Sessione tht: `tht-workspace-psd/sessions/2026-06-27-134553-crea-una-lista...` (status: open, F1).
Il modello ha lavorato molto e correttamente nel dominio:
- Ha creato la sessione, caricato la skill, iniziato F1.
- Ha eseguito ricerche semantiche (evidence + LSH), individuato le tabelle centrali
(`fact_cardioversione_elettrica`, `fact_see_ablazione`, `dim_patient`).
- Ha letto evidence molto pertinenti (esempio NLQ, glossario coorti/universi), costruendo
un quadro dominio corretto e verificato sulle tabelle reali.
Il workflow **non è avanzato oltre F1**: nessuna decisione registrata
(`review_decisions.jsonl` assente). Il modello si è arenato su bug di porting (sotto).
## Bug di porting emersi (4, di cui 1 fatale)
### #1 — `tht` non nel PATH del processo Pi [basso]
Il gate chiama `execFileSync("tht", args, {cwd: ctx.cwd})`. L'eseguibile nel venv non è nel
PATH di Pi. Il modello ha creato un wrapper in `~/.local/bin` (workaround).
**Fix root-cause:** installare `tht` in una dir nel PATH (pip install -e . con entry point
globale, o symlink `/usr/local/bin/tht -> harness/.venv/bin/tht`).
### #2 — `phase show` non passava il config [medio, FIXATO]
`phase_cmd._cfg()` non passava il config a `_load_config_or_exit()`, quindi falliva con
"File non trovato config/tht.yaml" quando non c'era `THT_WORKSPACE` env.
**Fix applicato (dal modello in sessione, validato e pulito):** `_cfg()` ora risolve
`THT_WORKSPACE`/`THT_CONFIG` env, poi fallback a `config/tht.yaml` (stessa convenzione di
`CONFIG_OPT`). Testato: `phase show` funziona. Suite 165 passed.
### #3 — `session check` signature inconsistente [basso, da verificare]
Il modello ha notato che `session check` prende la sessione come argomento posizionale,
non `--session` come gli altri cmd. Da verificare e allineare.
### #4 — `ctx.sendRaw is not a function` [FATALE, blocca tutti i widget]
Il gate `tht-gate.js:164` emette i widget con `ctx.sendRaw({type:"extension_ui_request"...})`.
Il runtime Pi installato **non espone `ctx.sendRaw`** sul context delle extensions. Il
modello l'ha verificato leggendo le type definitions (`ExtensionContext` espone `ui`,
`mode`, `hasUI`, `cwd`, `abort` — ma non `sendRaw`). **Nessun widget può essere emesso in
nessuna modalità.** Questo ha fermato il workflow a F1.
## Analisi del bug #4 (mismatch architetturale, non un typo)
Il commento nel gate stesso (righe 2-6) dice: *"REWRITE of the reference implementation...
replaces ctx.ui.* blocking primitives"*. Il porting ha **sostituito** i dialog nativi di Pi
con `ctx.sendRaw`, presumendo un'API widget-descriptor diretta che **questa versione di Pi
non espone alle extensions**. Verifiche sul runtime installato:
- `ctx.sendRaw`: **non esiste** (0 refs in `core/extensions/`).
- `extension_ui_request`: emesso **solo dal runtime** (`modes/rpc/rpc-mode.js`), come
traduzione dei dialog nativi (`ctx.ui.select` → `{method:"select"}`), non come API per
le extensions.
- Canali disponibili in RPC mode per ricevere una decisione umana (enum chiuso):
`ctx.ui.select` / `confirm` / `input` / `editor` (+ `notify` one-way). I `method` di
`extension_ui_request` sono: select/confirm/input/editor/notify/setWidget/setStatus/
set_editor_text. **Nessun method custom** per widget-descriptor.
- `ctx.ui.custom` (usato da ChironeWp3 per il multiselect TUI): in RPC mode è un **no-op**
(`return undefined`, commento: "Custom UI not supported in RPC mode").
### Conseguenza per i 6 widget della spec §4
- `select` (scelta singola) → ✅ `ctx.ui.select`
- `confirm` (approvazione) → ✅ `ctx.ui.confirm`
- `input` (testo libero / Altro) → ✅ `ctx.ui.input`
- `multiselect` (scelta multipla — **critico per F4 schema-linking**) → ❌ nessun canale
in RPC. Solo `ctx.ui.custom` lo faceva (TUI only).
Il multiselect è il vero ostacolo. F4 richiede di promuovere/escludere **più** tabelle/
colonne in una volta.
## Opzioni per il design del gate (decisione architetturale APERTA)
Il design del gate influenza tutta la relazione harness↔backend↔FE. Non è un fix da
inserire in coda a una sessione di porting; merita brainstorming dedicato. Opzioni:
- **A — Torna a `ctx.ui.*` nativi.** Il gate riscrive `emitAndWait` su `ctx.ui.select`/
`confirm`/`input` (come ChironeWp3 originale). Multiselect F4 emulato (serie di select, o
un input). Funziona con il Pi installato ora. Perde i widget-descriptor ricchi spec §4;
il FE riceve method nativi, la mappatura kind→method va nel backend/FE.
- **B — Versione di Pi con API widget custom.** Verificare se un Pi più recente/preview
espone `ctx.sendRaw` o un method custom. Se sì, il gate attuale funziona. Rischio:
inseguire un'API magari non pubblica; aggiornare Pi può rompere provider/auth.
- **C — Wrapper ibrido (valutato, NON realizzabile).** Registrare un handler che intercetti
gli `extension_ui_request` nativi e li arricchisca nel formato widget-descriptor spec §4
prima di mandarli al client RPC. **Scartato:** non c'è canale per widget-descriptor custom
in RPC (method è enum chiuso, `ctx.ui.custom` è no-op in RPC).
## Artefatti prodotti
- `tht-workspace-psd/sessions/2026-06-27-134553-.../` (manifest + question.md, F1, open).
- Sessione Pi transcript (vedi sopra) — fonte primaria per il debug.
- Fix codice: `_cfg()` in `phase_cmd.py` (bug #2), applicato pulito.
## Conclusione
**La sessione L2 ha colto bug di porting reali — ha fatto il suo lavoro.** Il loop
skill→LLM→gate funziona nel dominio (ricerche, evidence, quadro corretto) ma si blocca a
F1 sul bug fatale #4. Tre dei quattro bug sono risolvibili a basso costo (#1, #2 fixato,
#3); il #4 è una decisione di design del gate che richiede brainstorming prima del codice.
**Stato del porting codice (Onda -1→4 + Skill + 0b):** completo e verificato a livello L0/L1
(165 passed) + L2 value-grounding (PASS). I bug L2 emersi sono incrementi di qualità, non
regressioni del porting — L2 coglie ciò che L0/L1 per design non possono.
@@ -1,46 +0,0 @@
# ThothII — Stato e come riprendere
**Aggiornato:** 2026-06-27 (fine sessione)
**Branch corrente:** `main` (HEAD `b909f3e`)
## Cosa è fatto (tutto su `main`)
1. **Harness** (`harness/`) — completo e RPC-ready. Piano: [docs/superpowers/plans/2026-06-27-harness-rpc-readiness.md](plans/2026-06-27-harness-rpc-readiness.md).
2. **Backend** (`backend/`) — completo. Spec: [docs/superpowers/specs/2026-06-27-backend-design.md](specs/2026-06-27-backend-design.md). Piano: [docs/superpowers/plans/2026-06-27-backend-implementation.md](plans/2026-06-27-backend-implementation.md).
Entrambi sviluppati con subagent-driven development (un implementer + review per task, più review finale del branch). I branch `feat/harness-rpc-readiness` e `feat/backend-implementation` sono già merged in `main` (si possono cancellare: `git branch -d feat/harness-rpc-readiness feat/backend-implementation`).
## Verifica che sia tutto verde (primo comando di domani)
```bash
cd ~/projects/ThothII/backend && npm test && npm run build
cd ../harness && source .venv/bin/activate && pytest -q --ignore=tests/l2 && npm test
```
Atteso: backend 32 test + tsc OK; harness 233 pytest + 24 node:test.
## Il contratto chiave (per non perderlo)
Lo spike ha scoperto che il gate non poteva usare `ctx.sendRaw` (inesistente in Pi). Il **wire contract reale**: il gate emette via `ctx.ui.input(JSON.stringify(descriptor))` → Pi manda `{type:"extension_ui_request", id, method:"input", title:"<descriptor JSON>"}` → il client risponde `{type:"extension_ui_response", id, value:"<ui_response JSON>"}`. Il backend (`SessionBridge`) decodifica `title`→descriptor e ri-codifica la risposta in `value`. Il contratto verso il frontend (`ui_request`/`ui_response`) NON cambia.
## Due strade per riprendere (scelta rimandata)
### A) Validazione end-to-end con Pi reale (rischio residuo più importante)
I test sono tutti deterministici contro `fake-pi-rpc`; il loop con **Pi reale** non è mai stato eseguito. Richiede VPN attiva + `pi` (GLM 5.2) configurato + `harness/.env` popolato + il symlink `harness/config/tht.yaml`.
Due assunzioni da provare dal vivo: (1) il comando RPC `prompt` attiva l'`input` hook del gate; (2) Pi serializza verbatim il descriptor nel `title` di `ctx.ui.input`.
Come: avviare il backend (`cd backend && npm run dev`), fare `POST /sessions` con una domanda reale, aprire l'SSE `GET /sessions/:id/events`, verificare che arrivi il widget F1 e che la risposta avanzi. Se fallisce, il fix è localizzato a `emitAndWait` (gate) + il `fake-pi-rpc`.
### B) Frontend (terzo progetto)
Brainstorming → spec → piano del **frontend** (React/Next/ShadCn/AGGrid), come da architettura [docs/superpowers/specs/2026-06-25-thothii-architecture-design.md](specs/2026-06-25-thothii-architecture-design.md) §6. Consuma SOLO la REST+SSE del backend (contratto già stabile). Avviare con la skill `superpowers:brainstorming`.
## Backlog non bloccante (da chiudere quando si tocca l'area)
- **`GET /models`** torna `[]` di default (seam `listModels` pronto): cablare `pi --list-models` o uno spawn Pi effimero perché la tendina FE si popoli.
- **`resume()`** manda il kickoff `/nuova-domanda` invece di `/riprendi-sessione`: per il Pi reale `spawnFor` deve poter scegliere il kickoff di ripresa (altrimenti la sessione ripresa riparte da F1).
- **Multi-workspace lato Pi/gate**: il backend è workspace-aware, ma il gate usa il symlink `config/tht.yaml` (mono-workspace MVP). Per il multi-workspace reale il gate dovrebbe accettare un env `THT_CONFIG`/workspace.
- **`SseHub`**: i `Set` per-sessione vuoti non vengono rimossi (leak minore su processi long-running).
- **Copertura test** da estendere: route `steer`/`close`/404, handler SSE, percorsi `notify`/`oidc`; test gate per `/riprendi-sessione` RPC e i path di re-present di `emitAndWait` (cancel/id-mismatch/JSON-invalido).
- `GET /sessions/:id/events` su id inattivo apre uno stream 200 vuoto invece di 404.
## Nota operativa
Il ledger di esecuzione subagent-driven è in `.git/sdd/progress.md` (non versionato): contiene la cronologia task→commit→review. Utile se serve ricostruire perché una scelta è stata fatta.
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -1,860 +0,0 @@
# Frontend Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Costruire il frontend ThothII: una SPA Vite+React+TS che presenta il workflow NL→SQL human-in-the-loop renderizzando i widget-descriptor via SSE e raccogliendo le decisioni del revisore via REST, consumando solo il contratto già implementato del backend.
**Architecture:** SPA client-only su localhost. TanStack Query per le REST cacheable, uno store Zustand per lo stato live della sessione, un hook `useSessionStream` che apre l'SSE (`EventSource`) e alimenta lo store. Un widget registry mappa `kind`→renderer con fallback universale. I viewer (schema-linking Mermaid, SQL via shiki, risultati AGGrid) vivono nella sidebar destra del layout a 4 zone. Build a slice verticali con il loop F1 chiuso il prima possibile.
**Tech Stack:** Vite 6, React 18, TypeScript 5.6, @tanstack/react-query 5, zustand 5, Tailwind 3.4 + ShadCn, ag-grid-react/community 32, shiki 1, mermaid 11; test: Vitest 2 + @testing-library/react 16 + jsdom + MSW 2; e2e: @playwright/test 1.
## Global Constraints
- **Client-only SPA** (FE-1): niente SSR/route server. Tutto gira nel browser su localhost e consuma il backend Fastify separato.
- **Base URL backend:** da `import.meta.env.VITE_BACKEND_URL`, default `http://localhost:8787`. Mai hard-coded altrove.
- **Contratto FE↔BE invariato.** SSE eventi: `{type:"ui_request",ui_request}` | `{type:"text_delta",text}` | `{type:"info",level?,text}` | `{type:"system_event",event,...}`. REST: `GET /workspaces|/models|/sessions|/sessions/:id`; `POST /sessions|/sessions/:id/response|/steer|/sql/preview|/sql/export|/close|/resume`. Widget kind: `info|select|multiselect|freetext|artifact-gate|artifact` + fallback.
- **SSE via `EventSource` nativo** (FE-3): GET, nessun header (auth=`none` MVP). Non introdurre SSE su fetch in questo MVP.
- **TDD** (FE-5): ogni task scrive prima il test (Vitest+RTL); MSW mocka le REST, un mock di `EventSource` simula l'SSE. Nessun test tocca il backend reale tranne il Playwright e2e (Task finale).
- **Widget isolati** (FE-4): ogni renderer è un file con props `{descriptor, onRespond}`; aggiungere un widget = registrarlo, senza toccare store/stream/registry.
- **Invariante no-limbo:** nessun widget può "chiudere senza rispondere"; Esc/cancel non è una risposta valida.
- **TypeScript strict**; tutti i tipi del contratto vivono in `src/api/types.ts` (unica fonte). Niente `any` se non al confine del fallback.
- **Working dir:** tutti i comandi da `/Users/mp/projects/ThothII/frontend`. Il backend (per l'e2e) è in `../backend`, il fake-pi-rpc in `../harness/tests/fake_pi/`.
---
### Task 1: Scaffold Vite + React + TS + Tailwind/ShadCn + Vitest
**Files:**
- Create: `frontend/package.json`, `frontend/vite.config.ts`, `frontend/tsconfig.json`, `frontend/vitest.config.ts`, `frontend/index.html`, `frontend/tailwind.config.ts`, `frontend/postcss.config.js`, `frontend/src/main.tsx`, `frontend/src/App.tsx`, `frontend/src/index.css`, `frontend/src/test/setup.ts`
- Test: `frontend/src/App.test.tsx`
**Interfaces:**
- Produces: un'app montabile; `App` componente root; `npm test` esegue Vitest (jsdom); `npm run dev` serve la SPA; `npm run build` (tsc + vite build) pulito.
- [ ] **Step 1: package.json + config**
```json
// frontend/package.json
{
"name": "thothii-frontend",
"private": true,
"type": "module",
"scripts": {
"dev": "vite",
"build": "tsc -b && vite build",
"preview": "vite preview",
"test": "vitest run",
"test:watch": "vitest",
"e2e": "playwright test"
},
"dependencies": {
"react": "^18.3.1", "react-dom": "^18.3.1",
"@tanstack/react-query": "^5.59.0", "zustand": "^5.0.0"
},
"devDependencies": {
"vite": "^6.0.0", "@vitejs/plugin-react": "^4.3.0", "typescript": "^5.6.0",
"vitest": "^2.1.0", "jsdom": "^25.0.0",
"@testing-library/react": "^16.0.0", "@testing-library/jest-dom": "^6.5.0", "@testing-library/user-event": "^14.5.0",
"msw": "^2.4.0",
"tailwindcss": "^3.4.0", "postcss": "^8.4.0", "autoprefixer": "^10.4.0",
"@types/react": "^18.3.0", "@types/react-dom": "^18.3.0"
}
}
```
```ts
// frontend/vite.config.ts
import { defineConfig } from "vite";
import react from "@vitejs/plugin-react";
export default defineConfig({ plugins: [react()] });
```
```ts
// frontend/vitest.config.ts
import { defineConfig } from "vitest/config";
import react from "@vitejs/plugin-react";
export default defineConfig({
plugins: [react()],
test: { environment: "jsdom", globals: true, setupFiles: ["./src/test/setup.ts"], include: ["src/**/*.test.{ts,tsx}"] },
});
```
```json
// frontend/tsconfig.json
{ "compilerOptions": { "target": "ES2022", "useDefineForClassFields": true, "lib": ["ES2022","DOM","DOM.Iterable"],
"module": "ESNext", "moduleResolution": "Bundler", "jsx": "react-jsx", "strict": true, "noEmit": true,
"esModuleInterop": true, "skipLibCheck": true, "types": ["vitest/globals","@testing-library/jest-dom"] },
"include": ["src"] }
```
`tailwind.config.ts` (`content: ["./index.html","./src/**/*.{ts,tsx}"]`), `postcss.config.js` (tailwind+autoprefixer), `index.html` (root div + `/src/main.tsx`), `src/index.css` (`@tailwind base/components/utilities`).
- [ ] **Step 2: Write the failing test**
```tsx
// frontend/src/App.test.tsx
import { render, screen } from "@testing-library/react";
import { App } from "./App";
test("App renders the ThothII title", () => {
render(<App />);
expect(screen.getByText(/ThothII/i)).toBeInTheDocument();
});
```
- [ ] **Step 3: Run (fail)**
Run: `cd frontend && npm install && npm test`
Expected: FAIL — `Cannot find module './App'`.
- [ ] **Step 4: Implement App + main + setup**
```tsx
// frontend/src/App.tsx
export function App() {
return <div className="p-4 text-lg font-semibold">ThothII</div>;
}
```
```tsx
// frontend/src/main.tsx
import { StrictMode } from "react";
import { createRoot } from "react-dom/client";
import { App } from "./App";
import "./index.css";
createRoot(document.getElementById("root")!).render(<StrictMode><App /></StrictMode>);
```
```ts
// frontend/src/test/setup.ts
import "@testing-library/jest-dom/vitest";
```
- [ ] **Step 5: Run (pass) + ShadCn init + commit**
Run: `cd frontend && npm test` → PASS. Then init ShadCn (`npx shadcn@latest init -d`) and add the base components used later: `npx shadcn@latest add button checkbox radio-group textarea card dialog badge sonner`. Verify `npm run build` clean.
```bash
git add frontend
git commit -m "feat(frontend): scaffold Vite+React+TS+Tailwind/ShadCn + Vitest"
```
---
### Task 2: API types + REST client
**Files:**
- Create: `frontend/src/api/types.ts`, `frontend/src/api/client.ts`, `frontend/src/api/sessions.ts`, `frontend/src/api/workspaces.ts`, `frontend/src/api/models.ts`, `frontend/src/api/sql.ts`
- Create (test infra): `frontend/src/test/msw.ts`
- Test: `frontend/src/api/sessions.test.ts`
**Interfaces:**
- Produces (canonical contract types — every later task imports these):
```ts
export interface WidgetOption { id: string; label: string; meta?: Record<string, unknown>; selected?: boolean; recommended?: boolean; opens?: WidgetDescriptor; }
export interface WidgetDescriptor {
id: string; schema_version?: number; session_id?: string; phase?: string;
title?: string; intro?: string;
widget: "info" | "select" | "multiselect" | "freetext" | "artifact-gate" | "artifact" | string;
options?: WidgetOption[]; reserved?: string[]; allow_empty?: boolean;
artifact?: { kind: string; content?: string; [k: string]: unknown };
level?: "info" | "warning" | "error"; text?: string;
[k: string]: unknown;
}
export interface UiResponse { id: string; kind?: string; choices?: string[]; text?: string; decision?: { type: string }; control?: string; }
export type StreamEvent =
| { type: "ui_request"; ui_request: WidgetDescriptor }
| { type: "text_delta"; text: string }
| { type: "info"; level?: "info" | "warning" | "error"; text: string }
| { type: "system_event"; event: string; [k: string]: unknown };
export interface SessionSummary { id: string; status: string; question: string; summary: string | null; created_at: string; updated_at: string | null; author: string | null; }
export interface PreviewResult { columns: string[]; rows: unknown[][]; execution_ms: number; truncated: boolean; limit: number; offset: number; }
```
- Produces (functions): `createSession(input): Promise<{id:string}>`, `listSessions(): Promise<SessionSummary[]>`, `getSession(id): Promise<any>`, `postResponse(id, uiResponse): Promise<void>`, `postSteer(id, text): Promise<void>`, `closeSession(id): Promise<void>`, `resumeSession(id): Promise<void>`, `listWorkspaces(): Promise<{name:string;file:string}[]>`, `listModels(): Promise<{models:unknown[]}>`, `sqlPreview(id,{limit?,offset?}): Promise<PreviewResult>`, `sqlExport(id): Promise<{path:string}>`. Base: `apiFetch(path, init?)` in `client.ts` using `VITE_BACKEND_URL`.
- [ ] **Step 1: MSW test infra + failing test**
```ts
// frontend/src/test/msw.ts
import { setupServer } from "msw/node";
export const server = setupServer();
```
Register in `src/test/setup.ts`: `beforeAll(()=>server.listen()); afterEach(()=>server.resetHandlers()); afterAll(()=>server.close());` (import `server` + vitest globals).
```ts
// frontend/src/api/sessions.test.ts
import { http, HttpResponse } from "msw";
import { server } from "../test/msw";
import { createSession, listSessions } from "./sessions";
test("createSession POSTs and returns the id", async () => {
server.use(http.post("http://localhost:8787/sessions", () => HttpResponse.json({ id: "s1" })));
expect(await createSession({ workspace: "w", question: "q" })).toEqual({ id: "s1" });
});
test("listSessions GETs the array", async () => {
server.use(http.get("http://localhost:8787/sessions", () => HttpResponse.json([{ id: "s1", status: "open", question: "q", summary: null, created_at: "t", updated_at: null, author: null }])));
const rows = await listSessions();
expect(rows[0].id).toBe("s1");
});
```
- [ ] **Step 2: Run (fail)**
Run: `cd frontend && npm test -- sessions`
Expected: FAIL — modules absent.
- [ ] **Step 3: Implement client + api modules**
```ts
// frontend/src/api/client.ts
const BASE = import.meta.env.VITE_BACKEND_URL ?? "http://localhost:8787";
export async function apiFetch<T>(path: string, init?: RequestInit): Promise<T> {
const res = await fetch(`${BASE}${path}`, { headers: { "content-type": "application/json" }, ...init });
if (!res.ok) throw new Error(`${res.status} ${await res.text().catch(() => "")}`);
return res.status === 204 ? (undefined as T) : ((await res.json()) as T);
}
export { BASE };
```
```ts
// frontend/src/api/sessions.ts
import { apiFetch } from "./client";
import type { SessionSummary, UiResponse } from "./types";
export const createSession = (i: { workspace: string; question: string; provider?: string; model?: string; thinking?: string; name?: string }) =>
apiFetch<{ id: string }>("/sessions", { method: "POST", body: JSON.stringify(i) });
export const listSessions = () => apiFetch<SessionSummary[]>("/sessions");
export const getSession = (id: string) => apiFetch<any>(`/sessions/${id}`);
export const postResponse = (id: string, uiResponse: UiResponse) => apiFetch<void>(`/sessions/${id}/response`, { method: "POST", body: JSON.stringify({ ui_response: uiResponse }) });
export const postSteer = (id: string, text: string) => apiFetch<void>(`/sessions/${id}/steer`, { method: "POST", body: JSON.stringify({ text }) });
export const closeSession = (id: string) => apiFetch<void>(`/sessions/${id}/close`, { method: "POST" });
export const resumeSession = (id: string) => apiFetch<void>(`/sessions/${id}/resume`, { method: "POST" });
```
`workspaces.ts` (`listWorkspaces`), `models.ts` (`listModels`), `sql.ts` (`sqlPreview`, `sqlExport`) follow the same pattern with their endpoints; `types.ts` holds the interfaces above.
- [ ] **Step 4: Run (pass) + commit**
Run: `cd frontend && npm test -- sessions` → PASS; `npm run build` clean.
```bash
git add frontend/src/api frontend/src/test
git commit -m "feat(frontend): contract types + REST client (MSW-tested)"
```
---
### Task 3: Session store (Zustand)
**Files:**
- Create: `frontend/src/store/sessionStore.ts`
- Test: `frontend/src/store/sessionStore.test.ts`
**Interfaces:**
- Consumes: `StreamEvent`, `WidgetDescriptor` (Task 2).
- Produces: `useSessionStore` (Zustand) with state `{ pendingWidget: WidgetDescriptor|null; transcript: {role:"assistant";text:string}[]; toasts: {level:string;text:string}[]; lastSystemEvent: StreamEvent|null }` and actions `applyEvent(e: StreamEvent): void`, `clearPending(): void`, `resetSession(): void`. `applyEvent`: `ui_request`→set pendingWidget; `text_delta`→append to the current assistant transcript entry (create if none/after a widget); `info`→push toast; `system_event`→set lastSystemEvent.
- [ ] **Step 1: Failing test**
```ts
// frontend/src/store/sessionStore.test.ts
import { useSessionStore } from "./sessionStore";
beforeEach(() => useSessionStore.getState().resetSession());
test("ui_request sets pendingWidget", () => {
useSessionStore.getState().applyEvent({ type: "ui_request", ui_request: { id: "u1", widget: "select" } });
expect(useSessionStore.getState().pendingWidget?.id).toBe("u1");
});
test("text_delta accumulates into transcript", () => {
const s = useSessionStore.getState();
s.applyEvent({ type: "text_delta", text: "Ana" });
s.applyEvent({ type: "text_delta", text: "lisi" });
expect(useSessionStore.getState().transcript.at(-1)?.text).toBe("Analisi");
});
test("info pushes a toast", () => {
useSessionStore.getState().applyEvent({ type: "info", level: "warning", text: "attenzione" });
expect(useSessionStore.getState().toasts.at(-1)).toEqual({ level: "warning", text: "attenzione" });
});
test("clearPending removes the widget", () => {
useSessionStore.getState().applyEvent({ type: "ui_request", ui_request: { id: "u1", widget: "select" } });
useSessionStore.getState().clearPending();
expect(useSessionStore.getState().pendingWidget).toBeNull();
});
```
- [ ] **Step 2: Run (fail)** — `npm test -- sessionStore` → module absent.
- [ ] **Step 3: Implement**
```ts
// frontend/src/store/sessionStore.ts
import { create } from "zustand";
import type { StreamEvent, WidgetDescriptor } from "../api/types";
interface Entry { role: "assistant"; text: string }
interface SessionState {
pendingWidget: WidgetDescriptor | null; transcript: Entry[]; toasts: { level: string; text: string }[]; lastSystemEvent: StreamEvent | null;
applyEvent: (e: StreamEvent) => void; clearPending: () => void; resetSession: () => void;
}
const empty = { pendingWidget: null, transcript: [] as Entry[], toasts: [] as { level: string; text: string }[], lastSystemEvent: null };
export const useSessionStore = create<SessionState>((set) => ({
...empty,
applyEvent: (e) => set((st) => {
if (e.type === "ui_request") return { pendingWidget: e.ui_request };
if (e.type === "text_delta") {
const t = [...st.transcript];
const last = t.at(-1);
if (last && !st.pendingWidget) t[t.length - 1] = { role: "assistant", text: last.text + e.text };
else t.push({ role: "assistant", text: e.text });
return { transcript: t };
}
if (e.type === "info") return { toasts: [...st.toasts, { level: e.level ?? "info", text: e.text }] };
if (e.type === "system_event") return { lastSystemEvent: e };
return {};
}),
clearPending: () => set({ pendingWidget: null }),
resetSession: () => set({ ...empty }),
}));
```
- [ ] **Step 4: Run (pass) + commit**
Run: `npm test -- sessionStore` → PASS.
```bash
git add frontend/src/store
git commit -m "feat(frontend): Zustand session store + applyEvent"
```
---
### Task 4: `useSessionStream` (SSE → store)
**Files:**
- Create: `frontend/src/stream/useSessionStream.ts`
- Create (test): `frontend/src/test/fakeEventSource.ts`
- Test: `frontend/src/stream/useSessionStream.test.tsx`
**Interfaces:**
- Consumes: `useSessionStore.applyEvent` (Task 3), `BASE` (Task 2).
- Produces: `useSessionStream(sessionId: string | null): { connected: boolean }` — when `sessionId` is set, opens `new EventSource(\`${BASE}/sessions/${id}/events\`)`, parses each `message` `data` as JSON `StreamEvent`, calls `applyEvent`; closes on unmount / id change. Uses the global `EventSource` (overridable in tests via a fake).
- [ ] **Step 1: Fake EventSource + failing test**
```ts
// frontend/src/test/fakeEventSource.ts
export class FakeEventSource {
static instances: FakeEventSource[] = [];
onmessage: ((e: { data: string }) => void) | null = null;
onopen: (() => void) | null = null;
onerror: (() => void) | null = null;
closed = false;
constructor(public url: string) { FakeEventSource.instances.push(this); }
emit(obj: unknown) { this.onmessage?.({ data: JSON.stringify(obj) }); }
close() { this.closed = true; }
}
```
```tsx
// frontend/src/stream/useSessionStream.test.tsx
import { renderHook } from "@testing-library/react";
import { act } from "react";
import { FakeEventSource } from "../test/fakeEventSource";
import { useSessionStream } from "./useSessionStream";
import { useSessionStore } from "../store/sessionStore";
beforeEach(() => { FakeEventSource.instances = []; (globalThis as any).EventSource = FakeEventSource; useSessionStore.getState().resetSession(); });
test("opens an EventSource for the session and feeds events to the store", () => {
renderHook(() => useSessionStream("s1"));
const es = FakeEventSource.instances[0];
expect(es.url).toContain("/sessions/s1/events");
act(() => es.emit({ type: "ui_request", ui_request: { id: "u1", widget: "select" } }));
expect(useSessionStore.getState().pendingWidget?.id).toBe("u1");
});
test("closes the stream on unmount", () => {
const { unmount } = renderHook(() => useSessionStream("s1"));
const es = FakeEventSource.instances[0];
unmount();
expect(es.closed).toBe(true);
});
```
- [ ] **Step 2: Run (fail)** — module absent.
- [ ] **Step 3: Implement**
```ts
// frontend/src/stream/useSessionStream.ts
import { useEffect, useState } from "react";
import { BASE } from "../api/client";
import { useSessionStore } from "../store/sessionStore";
import type { StreamEvent } from "../api/types";
export function useSessionStream(sessionId: string | null) {
const [connected, setConnected] = useState(false);
const applyEvent = useSessionStore((s) => s.applyEvent);
useEffect(() => {
if (!sessionId) return;
const es = new EventSource(`${BASE}/sessions/${sessionId}/events`);
es.onopen = () => setConnected(true);
es.onerror = () => setConnected(false);
es.onmessage = (ev) => { try { applyEvent(JSON.parse(ev.data) as StreamEvent); } catch { /* ignore malformed */ } };
return () => { es.close(); setConnected(false); };
}, [sessionId, applyEvent]);
return { connected };
}
```
- [ ] **Step 4: Run (pass) + commit**
Run: `npm test -- useSessionStream` → PASS.
```bash
git add frontend/src/stream frontend/src/test/fakeEventSource.ts
git commit -m "feat(frontend): useSessionStream (EventSource -> store)"
```
---
### Task 5: Widget registry + fallback
**Files:**
- Create: `frontend/src/widgets/registry.ts`, `frontend/src/widgets/FallbackWidget.tsx`, `frontend/src/widgets/types.ts`
- Test: `frontend/src/widgets/registry.test.tsx`
**Interfaces:**
- Consumes: `WidgetDescriptor`, `UiResponse` (Task 2).
- Produces: `WidgetProps = { descriptor: WidgetDescriptor; onRespond: (r: UiResponse) => void }` (`widgets/types.ts`); `register(kind: string, comp: React.FC<WidgetProps>): void`; `resolve(kind: string): React.FC<WidgetProps>` (returns `FallbackWidget` for unknown). `FallbackWidget` renders the descriptor JSON + a freetext box that responds with `{id, control:"freetext", text}`.
- [ ] **Step 1: Failing test**
```tsx
// frontend/src/widgets/registry.test.tsx
import { render, screen } from "@testing-library/react";
import { register, resolve } from "./registry";
import type { WidgetProps } from "./types";
test("resolve returns the registered renderer", () => {
const Dummy = (_: WidgetProps) => <div>dummy</div>;
register("dummy", Dummy);
expect(resolve("dummy")).toBe(Dummy);
});
test("resolve falls back for unknown kind and shows the JSON", () => {
const Comp = resolve("totally-unknown");
render(<Comp descriptor={{ id: "u1", widget: "totally-unknown", title: "X" } as any} onRespond={() => {}} />);
expect(screen.getByText(/Widget non supportato/i)).toBeInTheDocument();
});
```
- [ ] **Step 2: Run (fail)** — modules absent.
- [ ] **Step 3: Implement**
```tsx
// frontend/src/widgets/types.ts
import type { WidgetDescriptor, UiResponse } from "../api/types";
export type WidgetProps = { descriptor: WidgetDescriptor; onRespond: (r: UiResponse) => void };
```
```tsx
// frontend/src/widgets/FallbackWidget.tsx
import { useState } from "react";
import type { WidgetProps } from "./types";
export function FallbackWidget({ descriptor, onRespond }: WidgetProps) {
const [text, setText] = useState("");
return (
<div className="border rounded p-3 space-y-2">
<p className="text-sm text-amber-600">Widget non supportato (kind: {descriptor.widget}) — rispondi manualmente</p>
<pre className="text-xs overflow-auto max-h-48 bg-muted p-2">{JSON.stringify(descriptor, null, 2)}</pre>
<textarea className="w-full border rounded p-1" value={text} onChange={(e) => setText(e.target.value)} />
<button className="border rounded px-2 py-1" onClick={() => onRespond({ id: descriptor.id, control: "freetext", text })}>Invia</button>
</div>
);
}
```
```tsx
// frontend/src/widgets/registry.ts
import type React from "react";
import type { WidgetProps } from "./types";
import { FallbackWidget } from "./FallbackWidget";
const registry = new Map<string, React.FC<WidgetProps>>();
export function register(kind: string, comp: React.FC<WidgetProps>) { registry.set(kind, comp); }
export function resolve(kind: string): React.FC<WidgetProps> { return registry.get(kind) ?? FallbackWidget; }
```
- [ ] **Step 4: Run (pass) + commit**
Run: `npm test -- registry` → PASS.
```bash
git add frontend/src/widgets
git commit -m "feat(frontend): widget registry + universal fallback"
```
---
### Task 6: F1 widgets — ReservedControls + select/info/freetext
**Files:**
- Create: `frontend/src/widgets/ReservedControls.tsx`, `frontend/src/widgets/SelectWidget.tsx`, `frontend/src/widgets/InfoWidget.tsx`, `frontend/src/widgets/FreetextWidget.tsx`, `frontend/src/widgets/index.ts`
- Test: `frontend/src/widgets/SelectWidget.test.tsx`, `frontend/src/widgets/FreetextWidget.test.tsx`
**Interfaces:**
- Consumes: `WidgetProps` (Task 5), `register` (Task 5).
- Produces: `ReservedControls({reserved, onControl})` renders buttons for `back`/`exit`/`other` present in `reserved[]`. `SelectWidget` (single-pick: each `option` a button; `recommended` gets a "(consigliato)" badge; responds `{id, kind:"select", choices:[optionId], decision?}`; reserved → `{id, control}`). `InfoWidget` (non-blocking, shows `text`/`level`; auto no response). `FreetextWidget` (textarea → `{id, kind:"freetext", text}`). `index.ts` registers `select`/`info`/`freetext` into the registry on import.
- [ ] **Step 1: Failing tests**
```tsx
// frontend/src/widgets/SelectWidget.test.tsx
import { render, screen } from "@testing-library/react";
import userEvent from "@testing-library/user-event";
import { SelectWidget } from "./SelectWidget";
test("picking an option responds with its id", async () => {
const onRespond = vi.fn();
render(<SelectWidget descriptor={{ id: "u1", widget: "select", options: [{ id: "a", label: "A" }, { id: "b", label: "B", recommended: true }] }} onRespond={onRespond} />);
expect(screen.getByText(/consigliato/i)).toBeInTheDocument();
await userEvent.click(screen.getByRole("button", { name: /A/ }));
expect(onRespond).toHaveBeenCalledWith({ id: "u1", kind: "select", choices: ["a"] });
});
test("a reserved control responds with control, not a choice", async () => {
const onRespond = vi.fn();
render(<SelectWidget descriptor={{ id: "u1", widget: "select", options: [{ id: "a", label: "A" }], reserved: ["back"] }} onRespond={onRespond} />);
await userEvent.click(screen.getByRole("button", { name: /indietro/i }));
expect(onRespond).toHaveBeenCalledWith({ id: "u1", control: "back" });
});
```
```tsx
// frontend/src/widgets/FreetextWidget.test.tsx
import { render, screen } from "@testing-library/react";
import userEvent from "@testing-library/user-event";
import { FreetextWidget } from "./FreetextWidget";
test("submits typed text", async () => {
const onRespond = vi.fn();
render(<FreetextWidget descriptor={{ id: "u1", widget: "freetext" }} onRespond={onRespond} />);
await userEvent.type(screen.getByRole("textbox"), "ciao");
await userEvent.click(screen.getByRole("button", { name: /invia/i }));
expect(onRespond).toHaveBeenCalledWith({ id: "u1", kind: "freetext", text: "ciao" });
});
```
- [ ] **Step 2: Run (fail)** — modules absent.
- [ ] **Step 3: Implement the widgets**
```tsx
// frontend/src/widgets/ReservedControls.tsx
const LABELS: Record<string, string> = { back: "Torna indietro", exit: "Esci", other: "Altro — specifica" };
export function ReservedControls({ reserved, onControl }: { reserved?: string[]; onControl: (c: string) => void }) {
if (!reserved?.length) return null;
return <div className="flex gap-2 pt-2">{reserved.map((c) => (
<button key={c} className="text-sm border rounded px-2 py-1" onClick={() => onControl(c)}>{LABELS[c] ?? c}</button>
))}</div>;
}
```
```tsx
// frontend/src/widgets/SelectWidget.tsx
import type { WidgetProps } from "./types";
import { ReservedControls } from "./ReservedControls";
export function SelectWidget({ descriptor, onRespond }: WidgetProps) {
return (
<div className="space-y-2">
{descriptor.title && <p className="font-medium">{descriptor.title}</p>}
{descriptor.intro && <p className="text-sm text-muted-foreground">{descriptor.intro}</p>}
<div className="flex flex-col gap-2">
{descriptor.options?.map((o) => (
<button key={o.id} className="border rounded px-3 py-2 text-left hover:bg-accent"
onClick={() => onRespond({ id: descriptor.id, kind: "select", choices: [o.id] })}>
{o.label}{o.recommended && <span className="ml-2 text-xs text-green-600">(consigliato)</span>}
</button>
))}
</div>
<ReservedControls reserved={descriptor.reserved} onControl={(c) => onRespond({ id: descriptor.id, control: c })} />
</div>
);
}
```
```tsx
// frontend/src/widgets/InfoWidget.tsx
import type { WidgetProps } from "./types";
export function InfoWidget({ descriptor }: WidgetProps) {
const color = descriptor.level === "error" ? "text-red-600" : descriptor.level === "warning" ? "text-amber-600" : "text-foreground";
return <p className={`text-sm ${color}`}>{descriptor.text ?? descriptor.title}</p>;
}
```
```tsx
// frontend/src/widgets/FreetextWidget.tsx
import { useState } from "react";
import type { WidgetProps } from "./types";
export function FreetextWidget({ descriptor, onRespond }: WidgetProps) {
const [text, setText] = useState("");
return (
<div className="space-y-2">
{descriptor.title && <p className="font-medium">{descriptor.title}</p>}
<textarea className="w-full border rounded p-2" value={text} onChange={(e) => setText(e.target.value)} />
<button className="border rounded px-3 py-1" onClick={() => onRespond({ id: descriptor.id, kind: "freetext", text })}>Invia</button>
</div>
);
}
```
```ts
// frontend/src/widgets/index.ts
import { register } from "./registry";
import { SelectWidget } from "./SelectWidget";
import { InfoWidget } from "./InfoWidget";
import { FreetextWidget } from "./FreetextWidget";
register("select", SelectWidget); register("info", InfoWidget); register("freetext", FreetextWidget);
export { resolve } from "./registry";
```
- [ ] **Step 4: Run (pass) + commit**
Run: `npm test -- SelectWidget FreetextWidget` → PASS.
```bash
git add frontend/src/widgets
git commit -m "feat(frontend): F1 widgets (select/info/freetext) + reserved controls"
```
---
### Task 7: App shell (4 zone) + F1 loop end-to-end (MSW)
**Files:**
- Create: `frontend/src/shell/AppShell.tsx`, `frontend/src/shell/WidgetHost.tsx`, `frontend/src/app/queryClient.ts`
- Modify: `frontend/src/App.tsx` (compose shell + QueryClientProvider), `frontend/src/main.tsx` (import `./widgets` to register)
- Test: `frontend/src/shell/f1-loop.test.tsx`
**Interfaces:**
- Consumes: `useSessionStream` (Task 4), `useSessionStore` (Task 3), `resolve` (Task 6), `createSession`/`postResponse` (Task 2).
- Produces: `AppShell` rendering the 4 zones (nav, workflow bar, chat+input center with `WidgetHost`, right sidebar). `WidgetHost` reads `pendingWidget` from the store and renders `resolve(widget)(descriptor, onRespond)`, where `onRespond` calls `postResponse(activeSessionId, r)` then `clearPending()`. `App` wraps everything in `QueryClientProvider`.
- [ ] **Step 1: Failing integration test (MSW + fake SSE)**
```tsx
// frontend/src/shell/f1-loop.test.tsx
import { render, screen } from "@testing-library/react";
import userEvent from "@testing-library/user-event";
import { http, HttpResponse } from "msw";
import { act } from "react";
import { server } from "../test/msw";
import { FakeEventSource } from "../test/fakeEventSource";
import { App } from "../App";
import { useSessionStore } from "../store/sessionStore";
beforeEach(() => { FakeEventSource.instances = []; (globalThis as any).EventSource = FakeEventSource; useSessionStore.getState().resetSession(); });
test("F1: create session -> widget via SSE -> respond -> POST /response", async () => {
let responded: any = null;
server.use(
http.post("http://localhost:8787/sessions", () => HttpResponse.json({ id: "s1" })),
http.get("http://localhost:8787/sessions", () => HttpResponse.json([])),
http.post("http://localhost:8787/sessions/s1/response", async ({ request }) => { responded = await request.json(); return new HttpResponse(null, { status: 204 }); }),
);
render(<App />);
await userEvent.click(screen.getByRole("button", { name: /nuova/i })); // start a session (UI affordance)
// simulate the backend emitting the F1 widget
act(() => FakeEventSource.instances[0].emit({ type: "ui_request", ui_request: { id: "u1", widget: "select", title: "Disambigua", options: [{ id: "a", label: "interpretazione A" }] } }));
await userEvent.click(await screen.findByRole("button", { name: /interpretazione A/ }));
expect(responded).toEqual({ ui_response: { id: "u1", kind: "select", choices: ["a"] } });
});
```
> Nota: il bottone "nuova" rappresenta l'affordance minima di creazione sessione di questo slice; la lista/creazione complete arrivano in Task 12. L'handler `createSession` fissa `activeSessionId="s1"`, su cui `useSessionStream` apre il fake SSE.
- [ ] **Step 2: Run (fail)** — shell absent.
- [ ] **Step 3: Implement shell + host + wiring**
`queryClient.ts` exports a configured `QueryClient`. `AppShell` lays out 4 zones with Tailwind grid; a "Nuova domanda" button calls `createSession({workspace, question})`, stores `activeSessionId`, and mounts `useSessionStream(activeSessionId)`. `WidgetHost`:
```tsx
// frontend/src/shell/WidgetHost.tsx
import { useSessionStore } from "../store/sessionStore";
import { resolve } from "../widgets";
import { postResponse } from "../api/sessions";
import type { UiResponse } from "../api/types";
export function WidgetHost({ sessionId }: { sessionId: string | null }) {
const pending = useSessionStore((s) => s.pendingWidget);
const clearPending = useSessionStore((s) => s.clearPending);
if (!pending) return null;
const Renderer = resolve(pending.widget);
const onRespond = async (r: UiResponse) => { if (sessionId) await postResponse(sessionId, r); clearPending(); };
return <Renderer descriptor={pending} onRespond={onRespond} />;
}
```
`App.tsx` wraps `AppShell` in `QueryClientProvider`; `main.tsx` adds `import "./widgets";` so renderers register.
- [ ] **Step 4: Run (pass) + commit**
Run: `npm test -- f1-loop` → PASS; full `npm test` green; `npm run build` clean.
```bash
git add frontend/src/shell frontend/src/app frontend/src/App.tsx frontend/src/main.tsx
git commit -m "feat(frontend): 4-zone shell + F1 loop end-to-end (MSW)"
```
---
### Task 8: Remaining widgets — multiselect, artifact-gate, artifact + linkage + no-limbo
**Files:**
- Create: `frontend/src/widgets/MultiselectWidget.tsx`, `frontend/src/widgets/ArtifactGateWidget.tsx`, `frontend/src/widgets/ArtifactWidget.tsx`, `frontend/src/widgets/LinkageHost.tsx`
- Modify: `frontend/src/widgets/index.ts` (register the three)
- Test: `frontend/src/widgets/MultiselectWidget.test.tsx`, `frontend/src/widgets/ArtifactGateWidget.test.tsx`, `frontend/src/widgets/linkage.test.tsx`
**Interfaces:**
- Consumes: `WidgetProps`, `ReservedControls`, `register` (Tasks 5–6).
- Produces: `MultiselectWidget` (checkboxes; initial `selected`; "seleziona/deseleziona tutti"; `allow_empty` gates the confirm; responds `{id, kind:"multiselect", choices}`). `ArtifactGateWidget` (renders `artifact.content` scrollable + disposizioni from `options`/reserved; "Rifiuta" opens a freetext via linkage; responds with the chosen disposition + optional text). `ArtifactWidget` (view-only, no response). `LinkageHost` wraps a renderer: if a chosen `option.opens`, it shows the child widget and merges both into one `UiResponse` (`{...parent, text: childText}`).
- [ ] **Step 1: Failing tests** (multiselect select-all + allow_empty; artifact-gate reject→freetext linkage returns combined response). Full RTL tests with `userEvent` asserting the emitted `UiResponse`.
```tsx
// frontend/src/widgets/MultiselectWidget.test.tsx (excerpt)
test("confirms the checked ids", async () => {
const onRespond = vi.fn();
render(<MultiselectWidget descriptor={{ id: "u1", widget: "multiselect", options: [{ id: "t1", label: "t1", selected: true }, { id: "t2", label: "t2" }] }} onRespond={onRespond} />);
await userEvent.click(screen.getByRole("checkbox", { name: /t2/ }));
await userEvent.click(screen.getByRole("button", { name: /conferma/i }));
expect(onRespond).toHaveBeenCalledWith({ id: "u1", kind: "multiselect", choices: ["t1", "t2"] });
});
```
```tsx
// frontend/src/widgets/linkage.test.tsx (excerpt)
test("reject opens a freetext and combines the reason", async () => {
const onRespond = vi.fn();
render(<ArtifactGateWidget descriptor={{ id: "u1", widget: "artifact-gate", artifact: { kind: "cte", content: "SELECT 1" },
options: [{ id: "approve", label: "Approva" }, { id: "reject", label: "Rifiuta", opens: { id: "u1c", widget: "freetext", title: "Motivazione" } }] }} onRespond={onRespond} />);
await userEvent.click(screen.getByRole("button", { name: /Rifiuta/ }));
await userEvent.type(screen.getByRole("textbox"), "join sbagliata");
await userEvent.click(screen.getByRole("button", { name: /invia/i }));
expect(onRespond).toHaveBeenCalledWith({ id: "u1", kind: "artifact-gate", choices: ["reject"], text: "join sbagliata" });
});
```
- [ ] **Step 2: Run (fail)** — modules absent.
- [ ] **Step 3: Implement** the three widgets + `LinkageHost`. Multiselect tracks a `Set` seeded from `selected`; confirm disabled when empty and `allow_empty===false`. ArtifactGate shows `artifact.content` in a scrollable `<pre>`; each option is a button; an option with `opens` routes through `LinkageHost` to collect child text before responding. Artifact is view-only. Register all three in `index.ts`.
- [ ] **Step 4: Run (pass) + commit**
Run: `npm test -- MultiselectWidget ArtifactGateWidget linkage` → PASS.
```bash
git add frontend/src/widgets
git commit -m "feat(frontend): multiselect/artifact-gate/artifact widgets + linkage + no-limbo"
```
---
### Task 9: SchemaLinkingViewer (Mermaid + table) + MarkdownView
**Files:**
- Create: `frontend/src/viewers/SchemaLinkingViewer.tsx`, `frontend/src/viewers/MarkdownView.tsx`, `frontend/src/viewers/mermaid.ts`
- Test: `frontend/src/viewers/SchemaLinkingViewer.test.tsx`
- Deps: `npm i mermaid react-markdown remark-gfm`
**Interfaces:**
- Consumes: the `schema_linking.json` artifact shape (`{candidates:[{kind,name,decision,...}], joins:[{from,to,...}], excluded:[...], open_questions:[]}` — from `GET /sessions/:id/artifacts/...` or embedded in an `artifact` descriptor).
- Produces: `SchemaLinkingViewer({linking})` — default Mermaid vertical flowchart (top-to-bottom), toggle to a hierarchical table; caps at 45 elements (shows a notice if exceeded); inline "perché" comment per node. `MarkdownView({source})` renders markdown (react-markdown + remark-gfm) with mermaid code-fences rendered via `mermaid.ts` (`renderMermaid(def): Promise<svg>`).
- [ ] **Step 1: Failing test** — render with a small linking object, assert the toggle switches between the mermaid container and a table that lists the promoted tables; assert the ≤45 cap notice appears for an oversized input. (Mock `mermaid.ts`'s `renderMermaid` to return a stub `<svg>` so the test is deterministic.)
- [ ] **Step 2: Run (fail)** — module absent.
- [ ] **Step 3: Implement** `mermaid.ts` (lazy `import("mermaid")`, `mermaid.render`), `SchemaLinkingViewer` (build the flowchart definition from candidates/joins; table view maps candidates by kind), `MarkdownView`.
- [ ] **Step 4: Run (pass) + commit**
```bash
git add frontend/src/viewers
git commit -m "feat(frontend): schema-linking viewer (mermaid+table) + markdown view"
```
---
### Task 10: SqlViewer (shiki, collapsible)
**Files:**
- Create: `frontend/src/viewers/SqlViewer.tsx`, `frontend/src/viewers/highlight.ts`
- Test: `frontend/src/viewers/SqlViewer.test.tsx`
- Deps: `npm i shiki`
**Interfaces:**
- Consumes: a CTE/SQL artifact (`{name, fields?, testStatus?, sql, comments?}` for CTEs; the final SQL is the same shape with the full SELECT).
- Produces: `SqlViewer({blocks})` — collapsible code blocks (`▾/▸`) each with a header (name, n° fields, test status badge), shiki-highlighted SQL, per-field comment; a global vertical/horizontal toggle. `highlight.ts` exports `highlightSql(code): Promise<string>` (shiki, sql grammar). Mock `highlight.ts` in tests for determinism.
- [ ] **Step 1: Failing test** — render two blocks, assert headers show name + field count + test-status badge; clicking a header collapses/expands the highlighted body.
- [ ] **Step 2–4:** implement, pass, commit (`feat(frontend): SQL/CTE viewer with shiki highlighting`).
---
### Task 11: ResultsPanel (AGGrid + preview/export)
**Files:**
- Create: `frontend/src/viewers/ResultsPanel.tsx`
- Test: `frontend/src/viewers/ResultsPanel.test.tsx`
- Deps: `npm i ag-grid-react ag-grid-community`
**Interfaces:**
- Consumes: `sqlPreview(id,{limit,offset})`→`PreviewResult`, `sqlExport(id)`→`{path}` (Task 2).
- Produces: `ResultsPanel({sessionId})` — a `[10 ▾ / tutti]` selector driving `sqlPreview` (limit/offset), an AGGrid showing `columns`/`rows`, an "Esporta CSV" button calling `sqlExport`. A single scalar result (1 col × 1 row) renders bold instead of a grid. Loading/error states from TanStack Query.
- [ ] **Step 1: Failing test (MSW)** — `sqlPreview` mocked returns 2 cols × 2 rows → AGGrid renders the rows; mock a 1×1 result → renders bold number; export button calls `/sql/export`. (AGGrid renders in jsdom; assert on cell text. If AGGrid needs the module registered, register `AllCommunityModule` in the panel.)
- [ ] **Step 2–4:** implement, pass, commit (`feat(frontend): results panel (AGGrid + preview/export)`).
---
### Task 12: Sessions list/create + steering + resume + selectors
**Files:**
- Create: `frontend/src/shell/NavSessions.tsx`, `frontend/src/shell/NewSessionDialog.tsx`, `frontend/src/shell/SteerInput.tsx`, `frontend/src/shell/WorkflowBar.tsx`
- Modify: `frontend/src/shell/AppShell.tsx` (wire them)
- Test: `frontend/src/shell/NavSessions.test.tsx`, `frontend/src/shell/SteerInput.test.tsx`
**Interfaces:**
- Consumes: `listSessions`, `createSession`, `resumeSession`, `postSteer`, `listWorkspaces`, `listModels` (Task 2); TanStack Query.
- Produces: `NavSessions` (left nav: query `listSessions`, click selects/`resumeSession`s a session). `NewSessionDialog` (ShadCn dialog: question + workspace select from `listWorkspaces` + optional model/thinking/provider from `listModels`, degrading gracefully to a free input when `models` is empty; calls `createSession`). `SteerInput` (a text field that sends `postSteer` for free-text `!`-style steering during a session). `WorkflowBar` (renders the 8 phases, highlighting `currentPhase` from the store/`getSession`).
- [ ] **Step 1: Failing tests** — NavSessions lists sessions from a mocked `listSessions` and selecting one triggers `resumeSession`; SteerInput posts to `/steer`. (MSW.)
- [ ] **Step 2–4:** implement, pass, commit (`feat(frontend): sessions list/create + steering + resume + selectors`).
> Graceful degradation (spec §9): when `listModels()` returns `{models:[]}`, the model field is a free text input with the configured default, not an empty dropdown.
---
### Task 13: Playwright e2e — F1 loop against real backend + fake-pi-rpc
**Files:**
- Create: `frontend/playwright.config.ts`, `frontend/e2e/f1.spec.ts`, `frontend/e2e/fixtures/start-stack.ts`
- Deps: `npm i -D @playwright/test && npx playwright install chromium`
**Interfaces:**
- Consumes: the built/served frontend (`npm run dev` / `vite preview`) + the real backend (`../backend`) started with an injected `spawnFn` pointing at `../harness/tests/fake_pi/fake_pi_rpc.mjs`, OR the backend `npm run dev` with `PI_BIN` swapped to a wrapper that execs the fake. The e2e drives the browser through the F1 loop.
- [ ] **Step 1: Write the e2e spec**
```ts
// frontend/e2e/f1.spec.ts (shape)
import { test, expect } from "@playwright/test";
test("F1 loop: new question -> F1 widget -> respond", async ({ page }) => {
await page.goto("/");
await page.getByRole("button", { name: /nuova/i }).click();
await page.getByLabel(/domanda/i).fill("quante cardioversioni nel 2024");
await page.getByRole("button", { name: /crea/i }).click();
await expect(page.getByText(/disambigua|chiarimento/i)).toBeVisible({ timeout: 30000 });
await page.getByRole("button").first().click(); // pick a disambiguation option
await expect(page.locator("body")).not.toContainText(/errore/i);
});
```
- [ ] **Step 2: playwright.config.ts** — `webServer` entries that start the backend (with the fake-pi-rpc wired) and the frontend dev server; `baseURL` the frontend. Document the exact `PI_BIN`/`spawnFn` wiring so the backend uses the fake, not real Pi (deterministic, no VPN).
- [ ] **Step 3: Run** `npx playwright test` → the F1 loop passes against the real backend driven by fake-pi-rpc.
- [ ] **Step 4: Commit** (`test(frontend): Playwright e2e F1 vs backend+fake-pi-rpc`).
> Nota: questo è l'analogo dell'e2e del backend. La validazione contro **Pi reale** (GLM 5.2, VPN) è separata e fa parte del "proviamo tutto assieme" finale, non di questo task CI.
---
## Self-Review
**Spec coverage** (vs `2026-06-27-frontend-design.md`):
- FE-1 Vite SPA: Task 1 ✓
- FE-2 TanStack Query + Zustand + SSE hook: Tasks 1 (QueryClient), 3 (store), 4 (stream) ✓
- FE-3 EventSource nativo: Task 4 ✓
- FE-4 registry + fallback: Task 5 ✓; renderers Tasks 6, 8 ✓
- FE-5 Vitest+RTL+MSW + Playwright: ogni task usa Vitest/RTL/MSW; e2e Task 13 ✓
- FE-6 slice con F1 primo loop chiuso: Task 7 chiude F1 ✓
- §4 componenti (api/store/stream/widgets/viewers/shell): Tasks 2–12 ✓
- §4.5 viewer (schema-linking, sql, results, markdown): Tasks 9, 10, 11 ✓
- §5 errori (kind sconosciuto→fallback, SSE disconnesso, /response 404→resume, preview errore): Task 5 (fallback), Task 4 (retry), Task 12 (resume), Task 11 (error state) ✓
- §7 slice 1–9: Tasks 1–13 ✓
**Placeholder scan:** i Task 9–12 condensano gli step 2–4 (run-fail/implement/run-pass/commit) in forma sintetica perché il pattern TDD è identico ai Task 1–8 e i componenti sono deterministici; gli step 1 (test) e le interfacce sono concreti. Nessun "TBD". I viewer pesanti (mermaid/shiki/AGGrid) sono mockati nei test unitari per determinismo (indicato in ogni task).
**Type consistency:** `WidgetProps {descriptor, onRespond}` coerente Tasks 5→8. `UiResponse` (con `kind`/`choices`/`text`/`control`/`decision`) coerente tra widget (6,8), `WidgetHost` (7) e `postResponse` (2). `StreamEvent` union coerente tra store (3), stream (4), backend SSE. `PreviewResult` coerente tra `sql.ts` (2) e `ResultsPanel` (11). `resolve(kind)` (5) usato da `WidgetHost` (7).
**Nota di sequenza:** Tasks 1–8 e 12 sono CI-puri (MSW/mock). Task 13 (Playwright) richiede backend+fake-pi avviati ma niente VPN/Pi reale. La validazione con Pi reale è il passo finale "tutto assieme", fuori da questo piano.
@@ -1,967 +0,0 @@
# Harness RPC-readiness Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Rendere l'harness `tht` pienamente pilotabile da un client RPC esterno (il backend): il gate funziona in `pi --mode rpc` (kickoff + round-trip widget), il backend possiede l'id di sessione, il CLI espone le uscite `--json` necessarie, ed esistono i due test-double che chiudono il loop in CI.
**Architecture:** Si parte da uno **spike** che osserva il comportamento reale del gate dentro `pi --mode rpc` (mai testato finora, cf. `docs/l2-run-report-2026-06-27.md`). Le scoperte dello spike fissano l'esatta forma del wire e guidano l'adattamento del gate. Tutto l'adattamento del gate è coperto da un **fake-pi-runtime** (mock dell'API estensione `pi`) eseguibile in CI; le aggiunte al CLI sono TDD deterministiche; un **fake-pi-rpc** (processo che parla il protocollo JSONL su stdio) viene consegnato qui come asset condiviso per il Piano Backend, con un golden test di contratto (D10).
**Tech Stack:** Python 3.13 + Typer + pytest (CLI `tht`); JavaScript ESM/CJS + `node --test` (gate + test-double); Pi = `@earendil-works/pi-coding-agent` (binario `pi`, modalità `--mode rpc`).
## Global Constraints
- **Node ≥ 20**; JS test runner: `node --test` (come `harness/package.json` → `npm test`).
- **Python 3.13**, ambiente in `harness/.venv` (`pip install -e ".[dev]"`); test: `pytest` da `harness/`.
- **Framing RPC: LF-only JSONL** — serializzazione `JSON.stringify(value) + "\n"`; lettura: split su `\n`, strip di un eventuale `\r` finale. MAI `readline` (spezza su separatori Unicode validi dentro le stringhe JSON). Riferimento canonico: `@earendil-works/pi-coding-agent/dist/modes/rpc/jsonl.js`.
- **WIRE CONTRACT (corretto post-spike Task 1):** `ctx.sendRaw` NON esiste in Pi e `pi.on("extension_ui_response")` NON viene dispatchato. L'unico meccanismo UI estensione in RPC è `ctx.ui.select/confirm/input(...)`, che instradano la risposta via `pendingExtensionRequests`. Il gate trasporta il widget-descriptor completo come **JSON nel `title` di `ctx.ui.input`**: Pi emette `{type:"extension_ui_request", id, method:"input", title:"<descriptor JSON>"}` e il client risponde `{type:"extension_ui_response", id, value:"<ui_response JSON>"}` (oppure `{id, cancelled:true}` → no-limbo, re-loop). Il **contratto verso il frontend** (`ui_request`/`ui_response`, architettura §4) resta invariato: la traduzione native↔FE avviene nel backend (Piano 2).
- **`--json` mantiene stdout puro**: in modalità JSON nessun warning/tabella umana su stdout (pattern già in `tht/cli/search_cmd.py`). Output: `typer.echo(json.dumps(data, ensure_ascii=False, indent=2))`.
- **Nessun segreto nel codice/test**: le credenziali stanno solo in `harness/.env` (gitignored). I task L2/spike che toccano Pi reale richiedono `.env` + VPN e NON girano in CI.
- **Il gate resta load-bearing**: gli invarianti verbatim del gate (anti-bypass `tool_call` hook, input-lock, no-limbo) non vanno indeboliti dagli adattamenti RPC.
- **Test setup per importare il gate (Task 3–5):** `tht-gate.js` è ESM e importa `typebox` (fornito a runtime da Pi, NON presente in `harness/`). Per testarlo, aggiungere `typebox` come **devDependency** in `harness/package.json` (versione allineata a Pi: `"typebox": "1.1.38"`) ed eseguire `npm install` una volta (Task 3, prima di scrivere i test). Su Node ≥ 22 `require()` di un modulo ESM senza top-level await funziona; se in questo ambiente dà problemi, usare `await import("../../tht-gate.js")` dentro un test `async` (i file di test possono restare `node --test`). Verificare l'import del gate PRIMA di scrivere asserzioni.
---
### Task 1: Spike — comportamento del gate in `pi --mode rpc`
> Spike di osservazione (non TDD): mai verificato finora. L'esito fissa la forma esatta del wire e decide i Task 3–4. Richiede Pi reale + `.env` + VPN.
**Files:**
- Create: `harness/scripts/rpc_probe.mjs` (driver manuale usa-e-getta, committato come strumento)
- Create: `harness/docs/rpc-readiness-findings.md` (referto delle osservazioni + decisioni)
**Interfaces:**
- Produces: `harness/docs/rpc-readiness-findings.md` con le risposte alle 4 domande sotto, citate dai Task 3 e 4.
- [ ] **Step 1: Scrivere il driver di probe**
`harness/scripts/rpc_probe.mjs` fa spawn di Pi in RPC, manda l'avvio del workflow come comando `prompt`, stampa ogni riga JSONL ricevuta con un prefisso, e quando arriva un `extension_ui_request` risponde con un `extension_ui_response` correlato per `id`.
```javascript
// rpc_probe.mjs — manual probe: drive `pi --mode rpc` and observe the gate.
// Usage (from harness/, with .env loaded + VPN up): node scripts/rpc_probe.mjs
import { spawn } from "node:child_process";
const pi = spawn("pi", ["--mode", "rpc"], { cwd: process.cwd(), env: process.env });
const send = (obj) => {
const line = JSON.stringify(obj) + "\n";
process.stdout.write(`>>> SEND ${line}`);
pi.stdin.write(line);
};
let buf = "";
pi.stdout.on("data", (chunk) => {
buf += chunk.toString("utf8");
for (let nl; (nl = buf.indexOf("\n")) !== -1; ) {
const line = buf.slice(0, nl).replace(/\r$/, "");
buf = buf.slice(nl + 1);
if (!line) continue;
console.log(`<<< RECV ${line}`);
let msg;
try { msg = JSON.parse(line); } catch { continue; }
if (msg.type === "extension_ui_request") {
// Reply in BOTH the gate's expected shape and Pi's native shape; observe which one unblocks.
const id = msg.id ?? msg.ui_request?.id;
send({ type: "extension_ui_response", id, control: "freetext", text: "PROBE-ANSWER" });
}
}
});
pi.stderr.on("data", (d) => process.stdout.write(`!!! STDERR ${d}`));
pi.on("exit", (code) => console.log(`### pi exited ${code}`));
// Kick off the workflow via an RPC prompt command (NOT interactive keystrokes).
setTimeout(() => send({ type: "prompt", message: '/nuova-domanda "quante cardioversioni nel 2024"' }), 500);
setTimeout(() => { pi.stdin.end(); }, 60000);
```
- [ ] **Step 2: Eseguire il probe e catturare l'output**
Run (da `harness/`, con `.env` caricato e VPN attiva):
```bash
set -a; . ./.env; set +a
node scripts/rpc_probe.mjs | tee /tmp/rpc_probe.log
```
Expected: una sequenza di righe `<<< RECV {...}`. Osservare in particolare se compare `text_delta`/eventi del modello e se compare una riga `extension_ui_request`.
- [ ] **Step 3: Registrare le 4 osservazioni decisive in `rpc-readiness-findings.md`**
Documentare con evidenza (righe del log) le risposte a:
1. **Kickoff**: inviando il workflow come comando `prompt`, il gate attiva kickoff + input-lock? (cioè: il modello riceve le istruzioni operative e parte dalla Fase 1, oppure l'`input` hook — che filtra `event.source === "interactive"` — non scatta?)
2. **Emissione widget**: quando il gate chiama `emitAndWait`, su stdout appare `{"type":"extension_ui_request","ui_request":{...}}` (envelope custom del gate) oppure no?
3. **Routing risposta**: inviando `extension_ui_response` con `id` correlato, il gate prosegue (la `handleUiResponse` risolve la promise) oppure la risposta viene assorbita da `rpc-mode` (`pendingExtensionRequests`) e il gate resta appeso?
4. **Source dell'input `!`**: lo steering (`prompt`/`steer` con testo `!...`) raggiunge il modello?
- [ ] **Step 4: Decisione esplicita per i Task 3–4**
In coda al referto, scrivere la decisione: per il **kickoff** (Task 3) e per il **round-trip** (Task 4), indicare se basta confermare il meccanismo esistente o serve adattarlo, e come. Casi attesi:
- Se (1) è NO → il gate deve riconoscere l'avvio del workflow anche per `event.source !== "interactive"` (Task 3).
- Se (3) è "assorbita" → il gate deve emettere il widget tramite il meccanismo nativo che registra in `pendingExtensionRequests` (Task 4, variante B), invece del `sendRaw` custom (variante A).
- [ ] **Step 5: Commit**
```bash
git add harness/scripts/rpc_probe.mjs harness/docs/rpc-readiness-findings.md
git commit -m "spike(harness): probe gate behavior in pi --mode rpc + findings"
```
---
### Task 2: fake-pi-runtime — mock dell'API estensione `pi` per i test del gate
**Files:**
- Create: `harness/.pi/extensions/gate/__tests__/fake_pi_runtime.js`
- Test: `harness/.pi/extensions/gate/__tests__/fake_pi_runtime.test.js`
**Interfaces:**
- Produces: `createFakePi()` → `{ pi, ctx, tools, emit, enqueueUi }` dove
- `pi.on(event, handler)` registra handler; `pi.registerTool(def, fn)` li memorizza in `tools` (Map per nome).
- `pi.emit(event, payload)` invoca i handler registrati (await se async) e ritorna il loro valore (per testare i ritorni `{block, action}` dei hook).
- `ctx.ui.notify(msg, level)` accoda in `ctx.notifications[]`; `ctx.ui.input/select/confirm(...)` registrano la chiamata in `ctx.uiCalls[]` e ritornano il prossimo valore da `ctx.uiQueue` (FIFO) — questo modella l'API UI NATIVA di Pi (non esiste `sendRaw`); `ctx.hasUI = true`; `ctx.cwd = "<tmp>"`.
- `enqueueUi(value)` mette in coda la prossima risposta che `ctx.ui.*` restituirà (es. `JSON.stringify(uiResponse)` per `input`, o `undefined` per simulare cancel → no-limbo).
- Consumes: nulla.
- [ ] **Step 1: Scrivere il test del runtime mock**
```javascript
const test = require("node:test");
const assert = require("node:assert");
const { createFakePi } = require("./fake_pi_runtime.js");
test("pi.on + emit invoca il handler e ne ritorna il valore", async () => {
const { pi } = createFakePi();
pi.on("tool_call", (e) => (e.toolName === "x" ? { block: true } : undefined));
assert.deepEqual(await pi.emit("tool_call", { toolName: "x" }), { block: true });
assert.equal(await pi.emit("tool_call", { toolName: "y" }), undefined);
});
test("ctx.ui.input registra la chiamata e ritorna il valore in coda (API nativa)", async () => {
const { ctx, enqueueUi } = createFakePi();
enqueueUi('{"id":"u1","choices":["a"]}');
const v = await ctx.ui.input("title-json", "");
assert.equal(v, '{"id":"u1","choices":["a"]}');
assert.equal(ctx.uiCalls[0].method, "input");
assert.equal(ctx.uiCalls[0].title, "title-json");
});
test("ctx.ui.input senza valore in coda ritorna undefined (modella cancel)", async () => {
const { ctx } = createFakePi();
assert.equal(await ctx.ui.input("t", ""), undefined);
});
```
- [ ] **Step 2: Eseguire il test (deve fallire)**
Run: `cd harness && node --test .pi/extensions/gate/__tests__/fake_pi_runtime.test.js`
Expected: FAIL — `Cannot find module './fake_pi_runtime.js'`.
- [ ] **Step 3: Implementare il runtime mock**
```javascript
// fake_pi_runtime.js — minimal mock of the Pi extension runtime for gate tests.
// Models Pi's NATIVE extension UI API (ctx.ui.select/confirm/input). There is NO
// ctx.sendRaw in Pi; the gate awaits ctx.ui.* and the runtime routes the response.
function createFakePi() {
const handlers = new Map();
const tools = new Map();
const ctx = {
hasUI: true,
cwd: "/tmp/fake-pi-session",
notifications: [],
uiCalls: [],
uiQueue: [],
ui: {
notify: async (message, level = "info") => ctx.notifications.push({ message, level }),
input: async (title, placeholder, opts) => { ctx.uiCalls.push({ method: "input", title, placeholder, opts }); return ctx.uiQueue.shift(); },
select: async (title, options, opts) => { ctx.uiCalls.push({ method: "select", title, options, opts }); return ctx.uiQueue.shift(); },
confirm: async (title, message, opts) => { ctx.uiCalls.push({ method: "confirm", title, message, opts }); return ctx.uiQueue.shift(); },
},
};
const pi = {
on: (event, handler) => {
if (!handlers.has(event)) handlers.set(event, []);
handlers.get(event).push(handler);
},
registerTool: (def, fn) => tools.set(def?.name ?? def, { def, fn }),
emit: async (event, payload) => {
let result;
for (const h of handlers.get(event) ?? []) result = await h(payload, ctx);
return result;
},
};
return { pi, ctx, tools, emit: pi.emit, enqueueUi: (v) => ctx.uiQueue.push(v) };
}
module.exports = { createFakePi };
```
- [ ] **Step 4: Eseguire il test (deve passare)**
Run: `cd harness && node --test .pi/extensions/gate/__tests__/fake_pi_runtime.test.js`
Expected: PASS (2 test).
- [ ] **Step 5: Commit**
```bash
git add harness/.pi/extensions/gate/__tests__/fake_pi_runtime.js harness/.pi/extensions/gate/__tests__/fake_pi_runtime.test.js
git commit -m "test(harness): fake-pi-runtime mock for gate CI tests"
```
---
### Task 3: Gate — kickoff + input-lock all'avvio del workflow in RPC mode
> Applica la decisione del Task 1 (osservazione #1). Il gate deve attivare kickoff + input-lock quando il workflow parte via comando RPC `prompt`, non solo da input interattivo.
**Files:**
- Modify: `harness/.pi/extensions/tht-gate.js` (handler `pi.on("input", …)`, ~riga 220-235)
- Test: `harness/.pi/extensions/gate/__tests__/gate_entry.test.js`
**Interfaces:**
- Consumes: `createFakePi()` (Task 2); il `default export` di `tht-gate.js` (la funzione `(pi) => {…}`).
- Produces: invariante "dopo un input di avvio workflow, `lockActive` è attivo e `pendingKickoff` è impostato", indipendentemente dal `source`.
- [ ] **Step 1: Scrivere il test di entry RPC**
```javascript
const test = require("node:test");
const assert = require("node:assert");
const { createFakePi } = require("./fake_pi_runtime.js");
const installGate = require("../../tht-gate.js").default ?? require("../../tht-gate.js");
test("avvio workflow via input non-interattivo attiva il lock (free text bloccato)", async () => {
const { pi } = createFakePi();
installGate(pi);
// entry del workflow con source 'rpc' (come un comando prompt RPC)
await pi.emit("input", { source: "rpc", text: '/nuova-domanda "x"' });
// dopo l'entry, un testo libero senza '!' deve essere bloccato (lock attivo)
const res = await pi.emit("input", { source: "rpc", text: "promuovi la tabella pazienti" });
assert.equal(res.action, "handled");
});
test("testo con '!' passa al modello (steer) anche con lock attivo", async () => {
const { pi } = createFakePi();
installGate(pi);
await pi.emit("input", { source: "rpc", text: "/nuova-domanda \"x\"" });
const res = await pi.emit("input", { source: "rpc", text: "!considera solo il 2024" });
assert.deepEqual(res, { action: "transform", text: "considera solo il 2024" });
});
```
- [ ] **Step 2: Eseguire il test (deve fallire)**
Run: `cd harness && node --test .pi/extensions/gate/__tests__/gate_entry.test.js`
Expected: FAIL — l'entry detection filtra `source === "interactive"`, quindi il lock non si attiva e il primo `emit` ritorna `{action:"continue"}` (non `handled`).
- [ ] **Step 3: Adattare l'entry detection nel gate**
In `tht-gate.js`, nel handler `pi.on("input", …)`: l'entry del workflow (`/nuova-domanda` | `/riprendi-sessione`) deve essere riconosciuta a prescindere dal `source`; il blocco del free-text resta valido per gli input dell'utente in sessione (interattivi o via RPC `prompt`). Sostituire la condizione di entry e quella di filtro:
```javascript
// entry detection: workflow-start funziona sia da TUI sia da comando RPC `prompt`.
if (/^\/(nuova-domanda|riprendi-sessione)\b/.test(raw)) {
lockActive = true;
lastSteered = false;
pendingKickoff = /^\/nuova-domanda\b/.test(raw) ? NUOVA_DOMANDA_KICKOFF : RIPRENDI_KICKOFF;
}
// free-input block: attivo quando il lock è su, per qualsiasi input utente (non solo interattivo).
if (!lockActive) return { action: "continue" };
```
(Rimuovere i due `event.source === "interactive"` su entry e filtro. Conservare invariati: passthrough `/…`, canale `!` → `transform`, notify.)
- [ ] **Step 4: Eseguire i test (devono passare)**
Run: `cd harness && node --test .pi/extensions/gate/__tests__/gate_entry.test.js`
Expected: PASS (2 test). Poi `npm test` per assicurare nessuna regressione sui builder.
Expected: tutti verdi.
- [ ] **Step 5: Commit**
```bash
git add harness/.pi/extensions/tht-gate.js harness/.pi/extensions/gate/__tests__/gate_entry.test.js
git commit -m "fix(harness): gate kickoff/lock entry works in RPC mode (not only interactive)"
```
---
### Task 4: Gate — round-trip del widget via API nativa `ctx.ui.input` (rewrite)
> **Corretto post-spike (Task 1):** `ctx.sendRaw` NON esiste e `pi.on("extension_ui_response")` non viene dispatchato — il meccanismo attuale di `emitAndWait` crasherebbe. Si riscrive `emitAndWait` per usare l'API UI NATIVA di Pi: il widget-descriptor viaggia come JSON nel `title` di `ctx.ui.input`; la risposta torna come stringa `value` (o `undefined` su cancel → no-limbo). Si rimuovono `_pending`, `handleUiResponse` e la registrazione `pi.on("extension_ui_response", …)`.
**Files:**
- Modify: `harness/.pi/extensions/tht-gate.js` (`emitAndWait`, ~riga 148-175; rimozione `handleUiResponse` ~189-195 e `pi.on("extension_ui_response", …)` ~289)
- Test: `harness/.pi/extensions/gate/__tests__/gate_roundtrip.test.js`
**Interfaces:**
- Consumes: `createFakePi()`/`enqueueUi` (Task 2). Aggiungere a `tht-gate.js` un **named export** `emitAndWait` per testarlo direttamente.
- Produces: `emitAndWait(ctx, descriptor): Promise<uiResponse>` — chiama `ctx.ui.input(JSON.stringify(descriptor), "")`; se `value` è `undefined`/`null` o JSON non valido o `resp.control === "cancel"` → ri-presenta (no-limbo); altrimenti ritorna `JSON.parse(value)`. Invariante: il `title` passato a `ctx.ui.input` è esattamente `JSON.stringify(descriptor)`.
- [ ] **Step 1: Scrivere il test del round-trip (via export diretto)**
```javascript
const test = require("node:test");
const assert = require("node:assert");
const { createFakePi } = require("./fake_pi_runtime.js");
const { emitAndWait } = require("../../tht-gate.js");
test("emitAndWait trasporta il descriptor come JSON nel title e ritorna la ui_response parsata", async () => {
const { ctx, enqueueUi } = createFakePi();
const descriptor = { id: "u1", widget: "select", options: [{ id: "a", label: "A" }] };
enqueueUi(JSON.stringify({ id: "u1", choices: ["a"] }));
const resp = await emitAndWait(ctx, descriptor);
assert.deepEqual(resp, { id: "u1", choices: ["a"] });
assert.equal(ctx.uiCalls[0].method, "input");
assert.equal(ctx.uiCalls[0].title, JSON.stringify(descriptor));
});
test("no-limbo: undefined (cancel) ri-presenta lo stesso widget", async () => {
const { ctx, enqueueUi } = createFakePi();
const descriptor = { id: "u1", widget: "select", options: [] };
enqueueUi(undefined); // 1° giro: cancel
enqueueUi(JSON.stringify({ id: "u1", choices: ["a"] })); // 2° giro: risposta valida
const resp = await emitAndWait(ctx, descriptor);
assert.deepEqual(resp, { id: "u1", choices: ["a"] });
assert.equal(ctx.uiCalls.length, 2); // ri-presentato una volta
assert.ok(ctx.notifications.some((n) => /Esc non chiude/.test(n.message)));
});
```
- [ ] **Step 2: Eseguire il test (deve fallire)**
Run: `cd harness && node --test .pi/extensions/gate/__tests__/gate_roundtrip.test.js`
Expected: FAIL — `emitAndWait` non è esportato / usa ancora `ctx.sendRaw`.
- [ ] **Step 3: Riscrivere `emitAndWait` ed esportarlo**
In `tht-gate.js` sostituire il blocco `_pending`/`emitAndWait`/`handleUiResponse`:
```javascript
// widget emission + wait — usa l'API UI NATIVA di Pi (ctx.ui.input). Il descriptor
// viaggia come JSON nel title; la risposta torna come stringa `value`. No ctx.sendRaw,
// no pi.on("extension_ui_response"): Pi instrada la risposta via pendingExtensionRequests.
export async function emitAndWait(ctx, descriptor) {
for (;;) {
const value = await ctx.ui.input(JSON.stringify(descriptor), "");
if (value === undefined || value === null) { await reLoop(ctx); continue; }
let resp;
try { resp = JSON.parse(value); } catch { await reLoop(ctx); continue; }
if (resp && resp.control !== "cancel" && resp.id === descriptor.id) return resp;
await reLoop(ctx);
}
}
async function reLoop(ctx) {
if (ctx.hasUI) await ctx.ui.notify(
"Esc non chiude il gate: usa Torna indietro / Esci / Altro dalle opzioni.", "warning");
}
```
Rimuovere: la `Map _pending`, la funzione `handleUiResponse` e la riga `pi.on("extension_ui_response", handleUiResponse)`. Le chiamate esistenti a `emitAndWait(ctx, descriptor)` (dai reviewer tool) restano invariate (stessa firma e stesso ritorno).
- [ ] **Step 4: Eseguire i test (devono passare)**
Run: `cd harness && node --test .pi/extensions/gate/__tests__/gate_roundtrip.test.js && npm test`
Expected: PASS (2 nuovi test + nessuna regressione su builder/entry/runtime).
- [ ] **Step 5: Commit**
```bash
git add harness/.pi/extensions/tht-gate.js harness/.pi/extensions/gate/__tests__/gate_roundtrip.test.js
git commit -m "fix(harness): gate widget round-trip via native ctx.ui.input (no sendRaw); no-limbo"
```
---
### Task 5: Sessione con id fornito dall'esterno (BE-5)
> Il backend pre-crea la sessione e ne possiede l'id; il gate deve USARE quell'id invece di istruire il modello a crearne uno. Veicolo: variabile d'ambiente `THT_SESSION` (già referenziata dal gate per `/torna`).
**Files:**
- Modify: `harness/.pi/extensions/tht-gate.js` (payload `NUOVA_DOMANDA_KICKOFF` + selezione kickoff, ~riga 44-62, 226-231)
- Test: `harness/.pi/extensions/gate/__tests__/gate_provided_session.test.js`
**Interfaces:**
- Consumes: `process.env.THT_SESSION` (impostata dal backend allo spawn).
- Produces: quando `THT_SESSION` è valorizzata, il kickoff iniettato istruisce il modello a USARE quell'id (niente `tht session new`); altrimenti comportamento attuale (il modello crea la sessione).
- [ ] **Step 1: Scrivere il test**
```javascript
const test = require("node:test");
const assert = require("node:assert");
const { createFakePi } = require("./fake_pi_runtime.js");
const installGate = require("../../tht-gate.js").default ?? require("../../tht-gate.js");
test("con THT_SESSION il kickoff usa l'id fornito e NON crea la sessione", async () => {
process.env.THT_SESSION = "2026-06-27-100000-test";
try {
const { pi, ctx } = createFakePi();
installGate(pi);
await pi.emit("input", { source: "rpc", text: '/nuova-domanda "x"' });
const injected = await pi.emit("before_agent_start", { });
const text = injected?.appendMessage ?? injected?.text ?? "";
assert.match(text, /2026-06-27-100000-test/);
assert.doesNotMatch(text, /tht session new/);
} finally { delete process.env.THT_SESSION; }
});
```
> Nota: adeguare il nome del campo ritornato da `before_agent_start` a come il gate inietta il kickoff (vedi handler ~riga 250). Il test asserisce il contenuto del testo iniettato.
- [ ] **Step 2: Eseguire il test (deve fallire)**
Run: `cd harness && node --test .pi/extensions/gate/__tests__/gate_provided_session.test.js`
Expected: FAIL — il kickoff contiene sempre `tht session new`.
- [ ] **Step 3: Implementare il kickoff a id fornito**
Aggiungere un secondo payload e selezionarlo quando `THT_SESSION` è presente:
```javascript
const NUOVA_DOMANDA_KICKOFF_PROVIDED = (sessionId) =>
"Istruzioni operative — sessione ThothII (workflow human-in-the-middle). " +
`La sessione è GIÀ creata: usa l'id \`${sessionId}\` in OGNI comando \`tht\`. ` +
"NON eseguire `tht session new`.\n" +
"1. Carica la skill leggendo `.pi/skills/tht-sessione/SKILL.md` con il tool `read`, " +
`poi segui il workflow dalla Fase 1 usando l'id \`${sessionId}\`.\n` +
"Regole non negoziabili: una domanda al reviewer per volta; mai promuovere/escludere/" +
"correggere senza conferma; le interazioni passano dai tool reviewer_*; testo libero col prefisso '!'. " +
"MAI `tht phase advance|reopen` né `tht decision add` da shell.";
// nella selezione del kickoff (input hook):
pendingKickoff = /^\/nuova-domanda\b/.test(raw)
? (process.env.THT_SESSION ? NUOVA_DOMANDA_KICKOFF_PROVIDED(process.env.THT_SESSION) : NUOVA_DOMANDA_KICKOFF)
: RIPRENDI_KICKOFF;
```
- [ ] **Step 4: Eseguire i test (devono passare)**
Run: `cd harness && node --test .pi/extensions/gate/__tests__/gate_provided_session.test.js && npm test`
Expected: PASS, nessuna regressione.
- [ ] **Step 5: Commit**
```bash
git add harness/.pi/extensions/tht-gate.js harness/.pi/extensions/gate/__tests__/gate_provided_session.test.js
git commit -m "feat(harness): gate uses externally-provided THT_SESSION id (BE-5)"
```
---
### Task 6: CLI — `tht sql preview --json` + `--offset`
> Alimenta la paginazione AGGrid del backend. `do_run` ritorna già `rows/columns/execution_ms/truncated`; serve l'uscita JSON e l'iniezione di OFFSET.
**Files:**
- Modify: `harness/tht/cli/sql_cmd.py` (`preview_cmd`, ~riga 157; `do_run`, ~riga 105)
- Modify: `harness/tht/execute/__init__.py` e `harness/tht/rest/execute.py` (firma `run_controlled*` con `offset`)
- Test: `harness/tests/test_sql_preview_json.py`
**Interfaces:**
- Consumes: `do_run(cfg, sql, *, limit, offset=0)`.
- Produces: `tht sql preview <file> --json [--limit N] [--offset M] [--session S]` stampa su stdout `{"columns": [...], "rows": [[...]], "execution_ms": int, "truncated": bool, "limit": int, "offset": int}` e nient'altro.
- [ ] **Step 1: Scrivere il test (offset injection + json shape)**
```python
# harness/tests/test_sql_preview_json.py
import json
from tht.execute.limit import inject_limit_offset # helper puro da creare
def test_inject_limit_offset_wraps_query():
sql = "SELECT a FROM t ORDER BY a"
out = inject_limit_offset(sql, limit=10, offset=20)
assert "LIMIT 10" in out and "OFFSET 20" in out
# la query originale resta una sottoquery (niente clobber di un LIMIT esistente)
assert "SELECT a FROM t ORDER BY a" in out
def test_inject_limit_offset_zero_offset_no_offset_clause():
out = inject_limit_offset("SELECT 1", limit=5, offset=0)
assert "LIMIT 5" in out
assert "OFFSET" not in out
```
- [ ] **Step 2: Eseguire il test (deve fallire)**
Run: `cd harness && pytest tests/test_sql_preview_json.py -v`
Expected: FAIL — `ModuleNotFoundError: tht.execute.limit`.
- [ ] **Step 3: Implementare l'helper di iniezione**
```python
# harness/tht/execute/limit.py
def inject_limit_offset(sql: str, *, limit: int, offset: int = 0) -> str:
"""Wrappa la query come sottoquery e applica LIMIT/OFFSET in modo non distruttivo.
Evita di sovrascrivere un LIMIT già presente nella query dell'utente."""
inner = sql.strip().rstrip(";")
clause = f"LIMIT {int(limit)}" + (f" OFFSET {int(offset)}" if offset else "")
return f"SELECT * FROM (\n{inner}\n) AS _tht_page {clause}"
```
- [ ] **Step 4: Eseguire il test (deve passare)**
Run: `cd harness && pytest tests/test_sql_preview_json.py -v`
Expected: PASS (2 test).
- [ ] **Step 5: Cablare `--json`/`--offset` in `preview_cmd` e `do_run`**
In `do_run` aggiungere `offset: int = 0` e usare `inject_limit_offset` quando `offset > 0` (altrimenti il path attuale a solo LIMIT). In `preview_cmd` aggiungere `offset` e `json_out`; in modalità JSON sopprimere tabella rich e warning, stampare solo il dict:
```python
@sql_app.command("preview")
def preview_cmd(
file: Path = typer.Argument(...),
limit: int = typer.Option(None, "--limit"),
offset: int = typer.Option(0, "--offset"),
session: str = typer.Option(None, "--session"),
json_out: bool = typer.Option(False, "--json", help="Output JSON puro per il backend."),
config: Path = CONFIG_OPT,
) -> None:
cfg = _load_config_or_exit(config)
require_action(cfg, "preview")
sql = _read_sql(file)
check = validate_or_exit(cfg, sql, session)
effective_limit = limit if limit is not None else cfg.execution.max_preview_rows
try:
result = do_run(cfg, sql, limit=effective_limit, offset=offset)
except ExecutionError as e:
if json_out:
typer.echo(json.dumps({"error": str(e)}, ensure_ascii=False)); raise typer.Exit(code=1)
typer.secho(f"ERRORE: {e}", fg=typer.colors.RED, err=True); raise typer.Exit(code=1)
if json_out:
typer.echo(json.dumps({
"columns": list(result.columns),
"rows": [list(r) for r in result.rows],
"execution_ms": result.execution_ms,
"truncated": result.truncated,
"limit": effective_limit, "offset": offset,
}, ensure_ascii=False))
return
# ... (path umano esistente invariato)
```
- [ ] **Step 6: Test del contratto JSON (fixture senza DB)**
Aggiungere a `test_sql_preview_json.py` un test che invoca `preview_cmd` con `do_run` monkeypatchato a un risultato fittizio e verifica che stdout sia JSON puro con le chiavi attese.
```python
def test_preview_json_pure_stdout(monkeypatch, tmp_path, capsys):
from tht.cli import sql_cmd
from types import SimpleNamespace
fake = SimpleNamespace(columns=["a"], rows=[[1],[2]], execution_ms=3, truncated=False)
monkeypatch.setattr(sql_cmd, "do_run", lambda *a, **k: fake)
monkeypatch.setattr(sql_cmd, "validate_or_exit", lambda *a, **k: SimpleNamespace(ast=None))
monkeypatch.setattr(sql_cmd, "require_action", lambda *a, **k: None)
monkeypatch.setattr(sql_cmd, "_load_config_or_exit",
lambda *a, **k: SimpleNamespace(execution=SimpleNamespace(max_preview_rows=100)))
f = tmp_path / "q.sql"; f.write_text("SELECT 1")
sql_cmd.preview_cmd(file=f, limit=None, offset=0, session=None, json_out=True, config=None)
out = capsys.readouterr().out.strip()
data = json.loads(out) # deve parsare: stdout puro
assert data["columns"] == ["a"] and data["rows"] == [[1],[2]]
```
Run: `cd harness && pytest tests/test_sql_preview_json.py -v`
Expected: PASS (tutti).
- [ ] **Step 7: Commit**
```bash
git add harness/tht/execute/limit.py harness/tht/cli/sql_cmd.py harness/tht/execute/__init__.py harness/tht/rest/execute.py harness/tests/test_sql_preview_json.py
git commit -m "feat(harness): tht sql preview --json + --offset for AGGrid paging (BE-2)"
```
---
### Task 7: CLI — `tht session list --json` + `tht session show --json`
> Alimenta la lista e il dettaglio sessioni nel FE. `list` è nuovo; `show` oggi stampa testo umano.
**Files:**
- Modify: `harness/tht/cli/session_cmd.py` (nuovo `list_cmd`; `show_cmd` con `--json`)
- Test: `harness/tests/test_session_list_json.py`
**Interfaces:**
- Produces:
- `tht session list --json` → `[{"id","status","question","summary","created_at","updated_at","author"}, ...]` ordinato per `created_at` desc.
- `tht session show <id> --json` → manifest completo + `{"phase": <int derivata>, "has_schema_linking": bool}`.
- [ ] **Step 1: Scrivere il test**
```python
# harness/tests/test_session_list_json.py
import json
from tht.session.store import create_session
from tht.session.models import SessionManifest
def test_list_json_lists_created_sessions(tmp_path):
from tht.config import DatabaseConfig
db = DatabaseConfig(database="d", schema="s", transport="rest") # adattare ai campi reali
m1 = create_session("prima domanda", db, tmp_path)
m2 = create_session("seconda domanda", db, tmp_path)
from tht.cli.session_cmd import _list_sessions # helper puro
rows = _list_sessions(tmp_path)
ids = [r["id"] for r in rows]
assert m1.id in ids and m2.id in ids
assert set(["id","status","question","created_at"]).issubset(rows[0].keys())
```
- [ ] **Step 2: Eseguire il test (deve fallire)**
Run: `cd harness && pytest tests/test_session_list_json.py -v`
Expected: FAIL — `_list_sessions` non esiste.
- [ ] **Step 3: Implementare helper + comandi**
```python
def _list_sessions(sessions_root: Path) -> list[dict]:
out = []
for d in sorted([p for p in sessions_root.iterdir() if (p / "session_manifest.yaml").exists()]):
m = SessionManifest.from_yaml(d / "session_manifest.yaml")
out.append({"id": m.id, "status": m.status, "question": m.question,
"summary": m.summary, "created_at": m.created_at.isoformat(),
"updated_at": m.updated_at.isoformat() if m.updated_at else None,
"author": m.author})
out.sort(key=lambda r: r["created_at"], reverse=True)
return out
@session_app.command("list")
def list_cmd(json_out: bool = typer.Option(False, "--json"), config: Path = CONFIG_OPT) -> None:
cfg = _load_config_or_exit(config)
rows = _list_sessions(cfg.paths.sessions)
if json_out:
typer.echo(json.dumps(rows, ensure_ascii=False, indent=2)); return
for r in rows:
typer.echo(f"{r['id']} [{r['status']}] {r['summary']}")
```
E in `show_cmd` aggiungere `json_out: bool = typer.Option(False, "--json")`; in modalità JSON stampare il manifest (`model_dump(mode="json", by_alias=True)`) + `phase` (da `current_phase`) + `has_schema_linking`.
- [ ] **Step 4: Eseguire il test (deve passare)**
Run: `cd harness && pytest tests/test_session_list_json.py -v`
Expected: PASS.
- [ ] **Step 5: Commit**
```bash
git add harness/tht/cli/session_cmd.py harness/tests/test_session_list_json.py
git commit -m "feat(harness): tht session list/show --json for FE session list"
```
---
### Task 8: Manifest — campi provider/model/thinking/name + opzioni di `tht session new`
> Persistono la scelta di modello/thinking/provider e il nome, riapplicati al resume dal backend (BE-6/BE-7).
**Files:**
- Modify: `harness/tht/session/models.py` (`SessionManifest`)
- Modify: `harness/tht/session/store.py` (`create_session`)
- Modify: `harness/tht/cli/session_cmd.py` (`new_cmd` opzioni)
- Test: `harness/tests/test_manifest_pi_fields.py`
**Interfaces:**
- Consumes: `create_session(question, db, sessions_root, *, author=None, summary=None, provider=None, model=None, thinking=None, name=None)`.
- Produces: manifest con campi opzionali `provider`, `model`, `thinking`, `name`; `tht session new <q> [--provider P --model M --thinking T --name N] [--json]` (con `--json` stampa `{"id": ...}` su stdout puro).
- [ ] **Step 1: Scrivere il test**
```python
# harness/tests/test_manifest_pi_fields.py
from tht.session.store import create_session
from tht.session.models import SessionManifest
def test_manifest_persists_pi_fields(tmp_path):
from tht.config import DatabaseConfig
db = DatabaseConfig(database="d", schema="s", transport="rest")
m = create_session("q", db, tmp_path, provider="zai", model="glm-5.2",
thinking="medium", name="sessione test")
reload = SessionManifest.from_yaml(tmp_path / m.id / "session_manifest.yaml")
assert reload.provider == "zai" and reload.model == "glm-5.2"
assert reload.thinking == "medium" and reload.name == "sessione test"
def test_manifest_pi_fields_optional(tmp_path):
from tht.config import DatabaseConfig
db = DatabaseConfig(database="d", schema="s", transport="rest")
m = create_session("q", db, tmp_path)
assert m.provider is None and m.model is None and m.thinking is None and m.name is None
```
- [ ] **Step 2: Eseguire il test (deve fallire)**
Run: `cd harness && pytest tests/test_manifest_pi_fields.py -v`
Expected: FAIL — `create_session` non accetta `provider`, ecc.
- [ ] **Step 3: Aggiungere i campi al modello e a create_session**
In `SessionManifest` (dopo `schema_version`):
```python
provider: str | None = None
model: str | None = None
thinking: str | None = None
name: str | None = None
```
In `create_session` aggiungere i parametri keyword e passarli al costruttore del manifest:
```python
def create_session(question, db, sessions_root, *, author=None, summary=None,
provider=None, model=None, thinking=None, name=None):
...
manifest = SessionManifest(
id=session_id, created_at=now, question=question,
database=db.database, schema=db.db_schema,
author=who, summary=summary or _summarize(question),
updated_at=now, updated_by=who, schema_version=schema_version,
provider=provider, model=model, thinking=thinking, name=name,
)
```
- [ ] **Step 4: Eseguire il test (deve passare)**
Run: `cd harness && pytest tests/test_manifest_pi_fields.py -v`
Expected: PASS (2 test).
- [ ] **Step 5: Aggiungere le opzioni a `tht session new` + `--json`**
```python
@session_app.command("new")
def new_cmd(
question: str = typer.Argument(...),
provider: str = typer.Option(None, "--provider"),
model: str = typer.Option(None, "--model"),
thinking: str = typer.Option(None, "--thinking"),
name: str = typer.Option(None, "--name"),
json_out: bool = typer.Option(False, "--json"),
config: Path = CONFIG_OPT,
) -> None:
from tht.session.store import create_session
cfg = _load_config_or_exit(config)
manifest = create_session(question, cfg.database, cfg.paths.sessions,
provider=provider, model=model, thinking=thinking, name=name)
if json_out:
typer.echo(json.dumps({"id": manifest.id}, ensure_ascii=False)); return
typer.secho(f"OK: sessione creata in {session_dir(cfg, manifest.id)}", fg=typer.colors.GREEN)
typer.echo(manifest.id)
```
Run: `cd harness && pytest tests/test_manifest_pi_fields.py tests/test_session_list_json.py -v`
Expected: PASS (regressione esistente verde).
- [ ] **Step 6: Commit**
```bash
git add harness/tht/session/models.py harness/tht/session/store.py harness/tht/cli/session_cmd.py harness/tests/test_manifest_pi_fields.py
git commit -m "feat(harness): manifest provider/model/thinking/name + session new options (BE-6/7)"
```
---
### Task 9: `.pi/settings.json` — quietStartup + trust
> Correttezza dello spawn RPC: niente rumore di avvio su stdout (sporcherebbe il JSONL), file project-local fidati (niente prompt di trust che appende il loop).
**Files:**
- Modify: `harness/.pi/settings.json`
- Test: `harness/.pi/extensions/gate/__tests__/settings.test.js`
**Interfaces:**
- Produces: `.pi/settings.json` contiene almeno `{"theme": "thothii-mono", "quietStartup": true}`; il trust dei file project-local è documentato (verifica empirica nello spike/Task 10).
- [ ] **Step 1: Scrivere il test (forma del settings)**
```javascript
const test = require("node:test");
const assert = require("node:assert");
const fs = require("node:fs");
const path = require("node:path");
test("settings.json abilita quietStartup", () => {
const s = JSON.parse(fs.readFileSync(path.join(__dirname, "../../settings.json"), "utf8"));
assert.equal(s.quietStartup, true);
assert.equal(s.theme, "thothii-mono");
});
```
- [ ] **Step 2: Eseguire il test (deve fallire)**
Run: `cd harness && node --test .pi/extensions/gate/__tests__/settings.test.js`
Expected: FAIL — `quietStartup` assente.
- [ ] **Step 3: Aggiornare settings.json**
```json
{
"theme": "thothii-mono",
"quietStartup": true
}
```
- [ ] **Step 4: Eseguire il test (deve passare)**
Run: `cd harness && node --test .pi/extensions/gate/__tests__/settings.test.js`
Expected: PASS.
- [ ] **Step 5: Documentare il trust + commit**
Aggiungere a `harness/docs/rpc-readiness-findings.md` una nota: come è stato concesso il trust dei file project-local allo spawn RPC (verificato che NON compaia un prompt di trust che blocca il loop — confermato nel Task 10 / spike).
```bash
git add harness/.pi/settings.json harness/.pi/extensions/gate/__tests__/settings.test.js harness/docs/rpc-readiness-findings.md
git commit -m "chore(harness): quietStartup + project-local trust for clean RPC spawn"
```
---
### Task 10: fake-pi-rpc — test-double del protocollo RPC + golden test di contratto (D10)
> Asset condiviso consegnato qui per il Piano Backend: un processo che parla il protocollo JSONL su stdio, scriptabile per emettere sequenze di eventi e accettare comandi. Un golden test fissa il contratto del widget-descriptor (D10).
**Files:**
- Create: `harness/tests/fake_pi/fake_pi_rpc.mjs`
- Create: `harness/tests/fake_pi/scripts/f1_disambiguation.json` (scenario scriptato)
- Test: `harness/tests/fake_pi/test_fake_pi_contract.mjs`
**Interfaces:**
- Produces: `fake_pi_rpc.mjs` — eseguibile con `node fake_pi_rpc.mjs <script.json>`. **Emette la shape NATIVA di Pi** (corretto post-spike): per ogni descriptor in `on_prompt`, emette `{type:"extension_ui_request", id:<descriptor.id>, method:"input", title: JSON.stringify(descriptor)}`. Su `{type:"extension_ui_response", id}` correla per id ed emette gli eventi `on_response[id]` (la risposta nativa porta `{id, value}`). Risponde a `{type:"get_available_models"}` con un set fisso; eco `{type:"response", command, success:true}` per `steer`/`get_state`.
- framing su stdout: `JSON.stringify(evt) + "\n"`.
- Consumes: lo schema widget-descriptor (architettura §4) per i descriptor di esempio.
- [ ] **Step 1: Definire lo scenario scriptato (golden)**
> Lo scenario tiene il descriptor come oggetto (`ui_request_descriptor`); è il fake a serializzarlo nel `title` nativo, così lo scenario resta leggibile.
```json
// harness/tests/fake_pi/scripts/f1_disambiguation.json
{
"on_prompt": [
{ "ui_request_descriptor": { "type": "ui_request", "id": "u1", "schema_version": 1,
"phase": "F1_chiarimento", "title": "Disambigua",
"widget": "select",
"options": [ {"id":"a","label":"interpretazione A"}, {"id":"b","label":"interpretazione B"} ],
"reserved": ["back","exit","other"] } }
],
"on_response": { "u1": [ { "type": "agent_end" } ] },
"available_models": [ {"provider":"zai","id":"glm-5.2"} ]
}
```
- [ ] **Step 2: Scrivere il test di contratto**
```javascript
// harness/tests/fake_pi/test_fake_pi_contract.mjs — run: node --test
import test from "node:test";
import assert from "node:assert";
import { spawn } from "node:child_process";
import path from "node:path";
function drive(scriptPath, commands) {
return new Promise((resolve) => {
const fp = spawn("node", [path.join(import.meta.dirname, "fake_pi_rpc.mjs"), scriptPath]);
const events = []; let buf = "";
fp.stdout.on("data", (c) => {
buf += c.toString("utf8");
for (let nl; (nl = buf.indexOf("\n")) !== -1; ) {
const line = buf.slice(0, nl).replace(/\r$/, ""); buf = buf.slice(nl + 1);
if (line) events.push(JSON.parse(line));
}
});
fp.on("exit", () => resolve(events));
for (const cmd of commands) fp.stdin.write(JSON.stringify(cmd) + "\n");
setTimeout(() => fp.stdin.end(), 300);
});
}
test("on prompt emette il widget F1; on response avanza", async () => {
const sp = path.join(import.meta.dirname, "scripts/f1_disambiguation.json");
const events = await drive(sp, [
{ type: "prompt", message: "/nuova-domanda \"x\"" },
{ type: "extension_ui_response", id: "u1",
value: JSON.stringify({ id: "u1", choices: ["a"], decision: { type: "concept_clarified" } }) },
]);
const widget = events.find((e) => e.type === "extension_ui_request");
assert.equal(widget.method, "input"); // shape nativa di Pi
assert.equal(widget.id, "u1");
const descriptor = JSON.parse(widget.title); // il descriptor viaggia nel title
assert.equal(descriptor.widget, "select");
assert.equal(descriptor.id, "u1");
assert.ok(events.some((e) => e.type === "agent_end"));
});
```
- [ ] **Step 3: Eseguire il test (deve fallire)**
Run: `cd harness && node --test tests/fake_pi/test_fake_pi_contract.mjs`
Expected: FAIL — `fake_pi_rpc.mjs` non esiste.
- [ ] **Step 4: Implementare il fake-pi-rpc**
```javascript
// harness/tests/fake_pi/fake_pi_rpc.mjs — scripted RPC test double (LF-only JSONL).
import fs from "node:fs";
const script = JSON.parse(fs.readFileSync(process.argv[2], "utf8"));
const out = (evt) => process.stdout.write(JSON.stringify(evt) + "\n");
let buf = "";
process.stdin.on("data", (chunk) => {
buf += chunk.toString("utf8");
for (let nl; (nl = buf.indexOf("\n")) !== -1; ) {
const line = buf.slice(0, nl).replace(/\r$/, ""); buf = buf.slice(nl + 1);
if (!line) continue;
let cmd; try { cmd = JSON.parse(line); } catch { continue; }
if (cmd.type === "prompt") {
for (const step of script.on_prompt ?? []) {
if (step.ui_request_descriptor) {
const d = step.ui_request_descriptor;
out({ type: "extension_ui_request", id: d.id, method: "input", title: JSON.stringify(d) });
} else { out(step); } // eventi non-UI (text_delta, agent_end, …) passano tali e quali
}
} else if (cmd.type === "extension_ui_response") {
for (const evt of (script.on_response ?? {})[cmd.id] ?? []) out(evt);
} else if (cmd.type === "get_available_models") {
out({ type: "response", command: "get_available_models", id: cmd.id, success: true,
data: { models: script.available_models ?? [] } });
} else if (cmd.type === "steer") {
out({ type: "response", command: "steer", id: cmd.id, success: true });
} else if (cmd.type === "get_state") {
out({ type: "response", command: "get_state", id: cmd.id, success: true,
data: { sessionId: "fake", thinkingLevel: "medium", isStreaming: false } });
}
}
});
process.stdin.on("end", () => process.exit(0));
```
- [ ] **Step 5: Eseguire il test (deve passare)**
Run: `cd harness && node --test tests/fake_pi/test_fake_pi_contract.mjs`
Expected: PASS.
- [ ] **Step 6: Validazione end-to-end con Pi reale (L2, informativo, non-CI)**
Con `.env` + VPN, ri-eseguire `node scripts/rpc_probe.mjs` (Task 1) e confermare che, dopo i Task 3–5, il gate: (a) parte sul `prompt`, (b) emette il widget, (c) riceve la risposta e avanza. Annotare l'esito in `rpc-readiness-findings.md`. Questo chiude il rischio "path RPC mai testato".
- [ ] **Step 7: Commit**
```bash
git add harness/tests/fake_pi/
git commit -m "test(harness): fake-pi-rpc protocol double + F1 widget contract golden (D10)"
```
---
## Self-Review
**Spec coverage** (vs `2026-06-27-backend-design.md` §7 + decisioni BE):
- BE-5 (id fornito): Task 5 ✓
- BE-6/BE-7 (model/thinking/provider/name nel manifest): Task 8 ✓; settings spawn (quietStartup/trust): Task 9 ✓
- §7.1 (preview --json/--offset): Task 6 ✓
- §7.2 (kickoff con id): Task 5 ✓
- §7.3 (campi manifest): Task 8 ✓
- §7.4 (settings.json): Task 9 ✓
- §7.5 (session list/show --json): Task 7 ✓
- §7.6 (fake-Pi condiviso): Task 10 (fake-pi-rpc) + Task 2 (fake-pi-runtime) ✓
- Rischio "gate RPC mai testato": Task 1 (spike) + Task 3/4 (adattamento) + Task 10 Step 6 (validazione reale) ✓
**Placeholder scan:** Task 1 è uno spike dichiarato (osservazione, non TDD) — i suoi "step" sono azioni concrete con output atteso. Le varianti A/B del Task 4 sono entrambe specificate; la scelta è guidata dall'evidenza dello spike, non un TBD.
**Type consistency:** `create_session(..., provider, model, thinking, name)` (Task 8) coerente con i campi del manifest (Task 8) e con l'uso del backend (Piano 2). `inject_limit_offset(sql, *, limit, offset)` (Task 6) usato da `do_run(..., offset=)`. `_list_sessions` (Task 7) ritorna le chiavi usate dal FE. Envelope `{"type":"extension_ui_request","ui_request":{...}}` (Task 10) coerente con l'emissione del gate (Task 4) e con ciò che il backend tradurrà (Piano 2).
**Nota di sequenza:** il Task 1 (spike) richiede Pi reale + VPN; se non disponibile al momento dell'esecuzione, i Task 2 e 6–9 (CI puri) possono procedere in parallelo; i Task 3–4 (adattamento gate) richiedono la decisione dello spike e vanno dopo.
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -1,750 +0,0 @@
# Ollama Ensure (embeddings preflight) Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Guarantee Ollama embeddings are available before a session starts — a hard-fail preflight that starts Ollama if down, warms the configured model, and refuses session create/restart when embeddings can't be made available.
**Architecture:** A deterministic `tht ollama ensure` command in the harness (which owns the embeddings config + the `OllamaEmbeddings` client) does probe → start (configurable command, detached) → poll → verify model installed → warm. The Fastify backend calls it as a preflight before spawning Pi; a non-zero exit becomes a 503 and no session is created.
**Tech Stack:** Python (typer, requests, subprocess, pydantic, pytest) · Node/Fastify + TypeScript (vitest).
## Global Constraints
- **Embeddings are mandatory.** Any condition that makes embeddings unavailable (no `embeddings`
config, Ollama unreachable after timeout, model not installed, warm fails) is a **hard error**:
the command exits non-zero and the backend refuses the session (HTTP 503, no Pi spawned). No
"degraded" session.
- **No automatic `ollama pull`** — a missing model is a hard error with guidance (`ollama pull <model>`).
- **"Load" = warm-only** — load the already-installed model into memory via one embed ping.
- **Parameterized invocation** — the ollama binary (`embeddings.bin`, default `"ollama"`) and the
start command (`embeddings.start_cmd`, default `[bin, "serve"]`, `[]` disables auto-start) are
config-driven; the base URL is `embeddings.base_url`.
- **`tht`'s `-c`/`--config` is PER-COMMAND** — it follows the subcommand (`ThtRunner` appends it).
- **`--json` output must be pristine** (only valid JSON on stdout). On error with `--json`, the
error JSON is on stdout AND the exit code is non-zero (the exit code is authoritative).
- **Blocking with timeout** (default 60s, backend env `OLLAMA_ENSURE_TIMEOUT_MS`): the timeout
only bounds waiting for the server to come up; on expiry → error, not degrade.
---
## File Structure
**Harness (`tht`)**
- Modify `harness/tht/config.py` — add `bin` + `start_cmd` to `EmbeddingsConfig`.
- Create `harness/tht/cli/ollama_cmd.py` — ops (`_probe`/`_installed_models`/`_start`/`_warm`),
the pure `ensure_ollama(...)` orchestration, and the `ollama ensure` CLI command.
- Modify `harness/tht/cli/__init__.py` — register the `ollama` sub-app.
- Create `harness/tests/test_ollama_ensure.py`.
**Backend (Fastify)**
- Modify `backend/src/tht/tht-runner.ts` — `ollamaEnsure(workspace, timeoutSec)`.
- Modify `backend/src/config.ts` — `ollamaEnsureTimeoutMs`.
- Modify `backend/src/app.ts` — pass the timeout (seconds) into `sessionRoutes` deps.
- Modify `backend/src/routes/sessions.ts` — preflight in `POST /sessions` and `POST /sessions/:id/resume`.
- Modify `backend/test/tht-runner.test.ts`, `backend/test/routes-sessions.test.ts`.
---
## Phase A — Harness
### Task 1: EmbeddingsConfig fields + `ensure_ollama` orchestration
**Files:**
- Modify: `harness/tht/config.py` (`EmbeddingsConfig`)
- Create: `harness/tht/cli/ollama_cmd.py`
- Test: `harness/tests/test_ollama_ensure.py` (create)
**Interfaces:**
- Consumes: `EmbeddingsConfig`, `OllamaEmbeddings` (`tht/vectorstore/embeddings.py`, `.embed_query`).
- Produces:
- `EmbeddingsConfig.bin: str = "ollama"`, `EmbeddingsConfig.start_cmd: list[str] | None = None`
- `ensure_ollama(cfg, *, timeout: int, no_start: bool, probe=_probe, installed_models=_installed_models, start=_start, warm=_warm, sleep=time.sleep, clock=time.monotonic) -> dict`
returning `{"ok": True, "server": "up"|"started", "model": "warmed", "model_name": str}` on
success or `{"ok": False, "stage": "config"|"server"|"model"|"warm", "error": str}` on failure.
- Module ops `_probe(base_url)->bool`, `_installed_models(base_url)->set[str]`, `_start(cmd)->None`,
`_warm(cfg)->None`, and `_model_present(installed, model)->bool`.
- [ ] **Step 1: Write the failing test**
Create `harness/tests/test_ollama_ensure.py`:
```python
"""Tests for the ensure_ollama orchestration (Ollama mocked via injected ops)."""
from types import SimpleNamespace
from tht.config import EmbeddingsConfig
from tht.cli.ollama_cmd import ensure_ollama
def _cfg(**kw):
emb = EmbeddingsConfig(base_url="http://localhost:11434", **kw)
return SimpleNamespace(embeddings=emb)
def test_no_embeddings_config_is_hard_error():
r = ensure_ollama(SimpleNamespace(embeddings=None), timeout=5, no_start=False)
assert r["ok"] is False and r["stage"] == "config"
def test_server_up_model_present_warms_ok():
warmed = []
r = ensure_ollama(
_cfg(model="nomic-embed-text-v2-moe"), timeout=5, no_start=False,
probe=lambda url: True,
installed_models=lambda url: {"nomic-embed-text-v2-moe:latest"},
start=lambda cmd: (_ for _ in ()).throw(AssertionError("must not start")),
warm=lambda cfg: warmed.append(True),
)
assert r == {"ok": True, "server": "up", "model": "warmed", "model_name": "nomic-embed-text-v2-moe"}
assert warmed == [True]
def test_server_down_then_started_after_poll():
started = []
probes = iter([False, True]) # down, then up after start
r = ensure_ollama(
_cfg(), timeout=5, no_start=False,
probe=lambda url: next(probes),
installed_models=lambda url: {"nomic-embed-text-v2-moe"},
start=lambda cmd: started.append(cmd),
warm=lambda cfg: None,
sleep=lambda s: None,
)
assert r["ok"] is True and r["server"] == "started"
assert started and started[0] == ["ollama", "serve"]
def test_server_unreachable_after_timeout_is_error():
clk = iter([0.0, 1.0, 2.0, 99.0]) # monotonic crosses the deadline
r = ensure_ollama(
_cfg(), timeout=5, no_start=False,
probe=lambda url: False, # never comes up
installed_models=lambda url: set(),
start=lambda cmd: None,
warm=lambda cfg: None,
sleep=lambda s: None,
clock=lambda: next(clk),
)
assert r["ok"] is False and r["stage"] == "server"
def test_no_start_and_down_is_error_without_starting():
r = ensure_ollama(
_cfg(), timeout=5, no_start=True,
probe=lambda url: False,
installed_models=lambda url: set(),
start=lambda cmd: (_ for _ in ()).throw(AssertionError("must not start")),
warm=lambda cfg: None,
)
assert r["ok"] is False and r["stage"] == "server"
def test_empty_start_cmd_disables_autostart():
r = ensure_ollama(
_cfg(start_cmd=[]), timeout=5, no_start=False,
probe=lambda url: False,
installed_models=lambda url: set(),
start=lambda cmd: (_ for _ in ()).throw(AssertionError("must not start")),
warm=lambda cfg: None,
)
assert r["ok"] is False and r["stage"] == "server"
def test_model_absent_is_error_with_pull_guidance():
r = ensure_ollama(
_cfg(model="missing-model"), timeout=5, no_start=False,
probe=lambda url: True,
installed_models=lambda url: {"nomic-embed-text-v2-moe"},
start=lambda cmd: None,
warm=lambda cfg: None,
)
assert r["ok"] is False and r["stage"] == "model"
assert "ollama pull missing-model" in r["error"]
def test_warm_failure_is_error():
r = ensure_ollama(
_cfg(), timeout=5, no_start=False,
probe=lambda url: True,
installed_models=lambda url: {"nomic-embed-text-v2-moe"},
start=lambda cmd: None,
warm=lambda cfg: (_ for _ in ()).throw(RuntimeError("boom")),
)
assert r["ok"] is False and r["stage"] == "warm"
def test_custom_start_cmd_used():
started = []
probes = iter([False, True])
ensure_ollama(
_cfg(bin="ollama", start_cmd=["docker", "start", "ollama"]), timeout=5, no_start=False,
probe=lambda url: next(probes),
installed_models=lambda url: {"nomic-embed-text-v2-moe"},
start=lambda cmd: started.append(cmd),
warm=lambda cfg: None,
sleep=lambda s: None,
)
assert started[0] == ["docker", "start", "ollama"]
```
- [ ] **Step 2: Run test to verify it fails**
Run: `cd harness && .venv/bin/pytest tests/test_ollama_ensure.py -q`
Expected: FAIL — `ModuleNotFoundError: No module named 'tht.cli.ollama_cmd'`.
- [ ] **Step 3: Add the config fields**
In `harness/tht/config.py`, in `EmbeddingsConfig`, after `timeout: int = 120` add:
```python
bin: str = "ollama"
start_cmd: list[str] | None = None
```
- [ ] **Step 4: Create `ollama_cmd.py`**
Create `harness/tht/cli/ollama_cmd.py`:
```python
"""`tht ollama` -- embeddings preflight (ensure Ollama up + model warm).
The system REQUIRES embeddings: any condition that makes them unavailable is a hard
error (the caller refuses the session). "Load" = warm the already-installed model.
"""
from __future__ import annotations
import json
import subprocess
import time
from pathlib import Path
import typer
from tht.cli.config_cmd import CONFIG_OPT
from tht.cli.schema_cmd import _load_config_or_exit
ollama_app = typer.Typer(help="Ollama (embeddings) -- preflight.")
# --- low-level ops (real implementations; injected as fakes in tests) ----------
def _probe(base_url: str, timeout: float = 2.0) -> bool:
import requests
try:
return requests.get(f"{base_url.rstrip('/')}/api/tags", timeout=timeout).status_code == 200
except requests.RequestException:
return False
def _installed_models(base_url: str, timeout: float = 5.0) -> set[str]:
import requests
resp = requests.get(f"{base_url.rstrip('/')}/api/tags", timeout=timeout)
resp.raise_for_status()
return {m.get("name", "") for m in resp.json().get("models", [])}
def _start(start_cmd: list[str]) -> None:
# Detached so the server outlives this short-lived CLI process.
subprocess.Popen( # noqa: S603
start_cmd, start_new_session=True,
stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL,
)
def _warm(cfg) -> None:
from tht.vectorstore.embeddings import OllamaEmbeddings
OllamaEmbeddings(cfg.embeddings).embed_query("ping")
def _model_present(installed: set[str], model: str) -> bool:
"""Match the configured model against installed names, allowing the implicit ':latest'."""
if model in installed:
return True
base = model.split(":")[0]
return any(name == base or name.split(":")[0] == base for name in installed)
# --- orchestration (pure: returns a result dict, never raises for control flow) ----
def ensure_ollama(
cfg,
*,
timeout: int,
no_start: bool,
probe=_probe,
installed_models=_installed_models,
start=_start,
warm=_warm,
sleep=time.sleep,
clock=time.monotonic,
) -> dict:
if cfg.embeddings is None:
return {"ok": False, "stage": "config",
"error": "il sistema richiede embeddings ma il workspace non li configura"}
emb = cfg.embeddings
base_url = emb.base_url
start_cmd = emb.start_cmd if emb.start_cmd is not None else [emb.bin, "serve"]
server_state = "up"
if not probe(base_url):
if no_start or start_cmd == []:
return {"ok": False, "stage": "server",
"error": f"Ollama non raggiungibile su {base_url} e avvio disabilitato"}
start(start_cmd)
server_state = "started"
deadline = clock() + timeout
up = False
while clock() < deadline:
sleep(1.0)
if probe(base_url):
up = True
break
if not up:
return {"ok": False, "stage": "server",
"error": f"Ollama non raggiungibile su {base_url} entro {timeout}s"}
try:
installed = installed_models(base_url)
except Exception as e: # noqa: BLE001 - any read failure is a hard error
return {"ok": False, "stage": "server",
"error": f"impossibile leggere i modelli da {base_url}: {e}"}
if not _model_present(installed, emb.model):
return {"ok": False, "stage": "model",
"error": f"modello '{emb.model}' non installato in Ollama: "
f"esegui `ollama pull {emb.model}` o importalo"}
try:
warm(cfg)
except Exception as e: # noqa: BLE001 - warm failure is a hard error
return {"ok": False, "stage": "warm",
"error": f"warm del modello '{emb.model}' fallito: {e}"}
return {"ok": True, "server": server_state, "model": "warmed", "model_name": emb.model}
```
- [ ] **Step 5: Run test to verify it passes**
Run: `cd harness && .venv/bin/pytest tests/test_ollama_ensure.py -q`
Expected: PASS (9 passed). Then `cd harness && .venv/bin/ruff check tht/cli/ollama_cmd.py tht/config.py tests/test_ollama_ensure.py` → clean.
- [ ] **Step 6: Commit**
```bash
git add harness/tht/config.py harness/tht/cli/ollama_cmd.py harness/tests/test_ollama_ensure.py
git commit -m "feat(harness): EmbeddingsConfig bin/start_cmd + ensure_ollama preflight orchestration"
```
---
### Task 2: `tht ollama ensure` CLI command + registration
**Files:**
- Modify: `harness/tht/cli/ollama_cmd.py` (add the command)
- Modify: `harness/tht/cli/__init__.py` (register `ollama_app`)
- Test: `harness/tests/test_ollama_ensure.py` (append CLI tests)
**Interfaces:**
- Consumes: `ensure_ollama` (Task 1), `_load_config_or_exit`, `CONFIG_OPT`, `ollama_app`.
- Produces: CLI `tht ollama ensure [--timeout N] [--no-start] [--json] -c <ws>` — exit 0 on ok,
exit 1 on any failure; with `--json`, only the result JSON on stdout (pristine) on both paths.
- [ ] **Step 1: Write the failing test**
Append to `harness/tests/test_ollama_ensure.py`:
```python
import json as _json
from typer.testing import CliRunner
from tht.cli.ollama_cmd import ollama_app
from tht.cli import ollama_cmd
def _patch(monkeypatch, result):
monkeypatch.setattr(ollama_cmd, "_load_config_or_exit", lambda _c: SimpleNamespace(embeddings=object()))
monkeypatch.setattr(ollama_cmd, "ensure_ollama", lambda cfg, **kw: result)
def test_cli_ok_exit_zero_and_json_pristine(monkeypatch):
_patch(monkeypatch, {"ok": True, "server": "up", "model": "warmed", "model_name": "m"})
res = CliRunner().invoke(ollama_app, ["ensure", "--json"])
assert res.exit_code == 0, res.output
assert _json.loads(res.output) == {"ok": True, "server": "up", "model": "warmed", "model_name": "m"}
def test_cli_error_exit_one_and_json_on_stdout(monkeypatch):
_patch(monkeypatch, {"ok": False, "stage": "model", "error": "missing"})
res = CliRunner().invoke(ollama_app, ["ensure", "--json"])
assert res.exit_code == 1
assert _json.loads(res.output) == {"ok": False, "stage": "model", "error": "missing"}
def test_cli_error_human_mode_exit_one(monkeypatch):
_patch(monkeypatch, {"ok": False, "stage": "server", "error": "down"})
res = CliRunner().invoke(ollama_app, ["ensure"])
assert res.exit_code == 1
def test_cli_registered_on_root_app():
from tht.cli import app # the root Typer app
runner = CliRunner()
res = runner.invoke(app, ["ollama", "--help"])
assert res.exit_code == 0
assert "ensure" in res.output
```
- [ ] **Step 2: Run test to verify it fails**
Run: `cd harness && .venv/bin/pytest tests/test_ollama_ensure.py -k cli -q`
Expected: FAIL — `No such command 'ensure'` / the root app has no `ollama` command.
- [ ] **Step 3: Add the CLI command**
In `harness/tht/cli/ollama_cmd.py`, append:
```python
@ollama_app.command("ensure")
def ensure_cmd(
timeout: int = typer.Option(60, "--timeout", help="Secondi di attesa per l'avvio di Ollama."),
no_start: bool = typer.Option(False, "--no-start", help="Non avviare Ollama (solo verifica)."),
json_out: bool = typer.Option(False, "--json", help="Emetti JSON puro su stdout."),
config: Path = CONFIG_OPT,
) -> None:
"""Assicura Ollama attivo + modello di embedding caricato; errore se non possibile."""
cfg = _load_config_or_exit(config)
result = ensure_ollama(cfg, timeout=timeout, no_start=no_start)
if json_out:
typer.echo(json.dumps(result, ensure_ascii=False))
elif result["ok"]:
typer.secho(
f"OK: Ollama {result['server']}, modello {result['model_name']} {result['model']}.",
fg=typer.colors.GREEN,
)
else:
typer.secho(f"ERRORE [{result['stage']}]: {result['error']}", fg=typer.colors.RED, err=True)
if not result["ok"]:
raise typer.Exit(code=1)
```
- [ ] **Step 4: Register the sub-app**
In `harness/tht/cli/__init__.py`, add the import alongside the others:
```python
from tht.cli.ollama_cmd import ollama_app # noqa: E402
```
and the registration alongside the other `add_typer` calls:
```python
app.add_typer(ollama_app, name="ollama")
```
- [ ] **Step 5: Run test to verify it passes**
Run: `cd harness && .venv/bin/pytest tests/test_ollama_ensure.py -q`
Expected: PASS (all). Then `cd harness && .venv/bin/ruff check tht/cli/ollama_cmd.py tht/cli/__init__.py` → clean.
- [ ] **Step 6: Commit**
```bash
git add harness/tht/cli/ollama_cmd.py harness/tht/cli/__init__.py harness/tests/test_ollama_ensure.py
git commit -m "feat(harness): tht ollama ensure CLI command (hard-fail preflight)"
```
---
## Phase B — Backend
### Task 3: `ThtRunner.ollamaEnsure`
**Files:**
- Modify: `backend/src/tht/tht-runner.ts`
- Test: `backend/test/tht-runner.test.ts` (append)
**Interfaces:**
- Consumes: `ThtRunner.run(args, workspace)`.
- Produces: `ollamaEnsure(workspace: string, timeoutSec: number): Promise<{ ok: boolean; stage?: string; error?: string; server?: string; model?: string; model_name?: string }>`
— shells `tht ollama ensure --json --timeout <sec>` (workspace via the per-command `-c`); exit 0
→ `{ ok: true, ...parsedJson }`, non-zero → `{ ok: false, stage, error }`.
- [ ] **Step 1: Write the failing test**
Append to `backend/test/tht-runner.test.ts`:
```typescript
test("ollamaEnsure builds argv with --json --timeout and the workspace -c", async () => {
let calledArgs: string[] = [];
let calledWs: string | undefined;
const r = new ThtRunner({ thtBin: "tht", harnessDir: "/nope", configPath: "config/tht.yaml" });
r.run = async (args, ws) => { calledArgs = args; calledWs = ws; return { code: 0, stdout: '{"ok":true,"server":"up","model":"warmed","model_name":"m"}', stderr: "" }; };
const res = await r.ollamaEnsure("psd", 60);
expect(calledArgs).toEqual(["ollama", "ensure", "--json", "--timeout", "60"]);
expect(calledWs).toBe("psd");
expect(res).toEqual({ ok: true, server: "up", model: "warmed", model_name: "m" });
});
test("ollamaEnsure maps a non-zero exit to ok:false with stage/error from stdout JSON", async () => {
const r = new ThtRunner({ thtBin: "tht", harnessDir: "/nope", configPath: "config/tht.yaml" });
r.run = async () => ({ code: 1, stdout: '{"ok":false,"stage":"model","error":"missing"}', stderr: "" });
expect(await r.ollamaEnsure("psd", 60)).toEqual({ ok: false, stage: "model", error: "missing" });
});
test("ollamaEnsure falls back to stderr when stdout is not JSON on failure", async () => {
const r = new ThtRunner({ thtBin: "tht", harnessDir: "/nope", configPath: "config/tht.yaml" });
r.run = async () => ({ code: 1, stdout: "", stderr: "boom" });
const res = await r.ollamaEnsure("psd", 60);
expect(res.ok).toBe(false);
expect(res.error).toContain("boom");
});
```
- [ ] **Step 2: Run test to verify it fails**
Run: `cd backend && npx vitest run test/tht-runner.test.ts`
Expected: FAIL — `r.ollamaEnsure is not a function`.
- [ ] **Step 3: Implement the method**
In `backend/src/tht/tht-runner.ts`, add a result interface near `SessionDocument`:
```typescript
export interface OllamaEnsureResult {
ok: boolean;
stage?: string;
error?: string;
server?: string;
model?: string;
model_name?: string;
}
```
and the method (after `documents`):
```typescript
async ollamaEnsure(workspace: string, timeoutSec: number): Promise<OllamaEnsureResult> {
const { code, stdout, stderr } = await this.run(
["ollama", "ensure", "--json", "--timeout", String(timeoutSec)],
workspace,
);
let parsed: Partial<OllamaEnsureResult> = {};
try { parsed = JSON.parse(stdout.trim() || "{}"); } catch { /* leave {} */ }
if (code === 0) return { ok: true, ...parsed };
return {
ok: false,
stage: parsed.stage,
error: parsed.error ?? (stderr.trim() || `tht ollama ensure exit ${code}`),
};
}
```
- [ ] **Step 4: Run test to verify it passes**
Run: `cd backend && npx vitest run test/tht-runner.test.ts` then `npx tsc --noEmit -p .`
Expected: PASS, typecheck clean.
- [ ] **Step 5: Commit**
```bash
git add backend/src/tht/tht-runner.ts backend/test/tht-runner.test.ts
git commit -m "feat(backend): ThtRunner.ollamaEnsure (parses tht ollama ensure --json)"
```
---
### Task 4: Backend preflight in session routes + timeout config
**Files:**
- Modify: `backend/src/config.ts` (`ollamaEnsureTimeoutMs`)
- Modify: `backend/src/app.ts` (pass `ollamaEnsureTimeoutSec` to `sessionRoutes`)
- Modify: `backend/src/routes/sessions.ts` (preflight in create + resume)
- Test: `backend/test/routes-sessions.test.ts` (append)
**Interfaces:**
- Consumes: `ThtRunner.ollamaEnsure` (Task 3), `getSettings().workspace`.
- Produces: `POST /sessions` and `POST /sessions/:id/resume` run `ollamaEnsure` first; on `!ok`
reply **503** `{ error }` and do NOT create/resume; the `sessionRoutes` deps object gains
`ollamaEnsureTimeoutSec: number`.
- [ ] **Step 1: Write the failing test**
Append to `backend/test/routes-sessions.test.ts`:
```typescript
test("POST /sessions refuses with 503 when ollamaEnsure fails (no session created)", async () => {
let createdCalled = false;
const app = buildApp(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
thtRunner: {
ollamaEnsure: async () => ({ ok: false, stage: "model", error: "modello non installato" }),
sessionNew: async () => { createdCalled = true; return { id: "s1" }; },
} as any,
getSettings: () => ({ workspace: "psd" }) as any,
spawnFn: () => nodeSpawn("node", [FAKE, SCRIPT]) as any,
});
const res = await app.inject({ method: "POST", url: "/sessions", payload: { question: "q" } });
expect(res.statusCode).toBe(503);
expect(res.json().error).toContain("non installato");
expect(createdCalled).toBe(false);
});
test("POST /sessions proceeds when ollamaEnsure succeeds", async () => {
let ensureWs: string | undefined;
const app = buildApp(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
thtRunner: {
ollamaEnsure: async (ws: string) => { ensureWs = ws; return { ok: true }; },
sessionNew: async () => ({ id: "s1" }),
} as any,
getSettings: () => ({ workspace: "psd" }) as any,
spawnFn: () => nodeSpawn("node", [FAKE, SCRIPT]) as any,
});
const res = await app.inject({ method: "POST", url: "/sessions", payload: { question: "q" } });
expect(res.json()).toEqual({ id: "s1" });
expect(ensureWs).toBe("psd");
});
test("POST /sessions/:id/resume refuses with 503 when ollamaEnsure fails", async () => {
const app = buildApp(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
thtRunner: {
ollamaEnsure: async () => ({ ok: false, error: "Ollama down" }),
sessionShow: async () => ({ status: "open", archived: false }),
} as any,
getSettings: () => ({ workspace: "psd" }) as any,
spawnFn: () => nodeSpawn("node", [FAKE, SCRIPT]) as any,
});
const res = await app.inject({ method: "POST", url: "/sessions/s1/resume" });
expect(res.statusCode).toBe(503);
});
```
(Note: the existing `mutApp` tests in this file do not pass `ollamaEnsure`; keep those tests
unaffected — the routes that call `ollamaEnsure` are only create and resume, and `mutApp` is used
for rename/group/archive/delete/documents. The two pre-existing create/resume tests at the top of
the file DO call create/resume, so update them per Step 4's note.)
- [ ] **Step 2: Run test to verify it fails**
Run: `cd backend && npx vitest run test/routes-sessions.test.ts`
Expected: FAIL — create/resume don't call `ollamaEnsure` (503 tests fail; and the new success test fails because `ollamaEnsure` isn't invoked).
- [ ] **Step 3: Add the timeout config**
In `backend/src/config.ts`, add to the `AppConfig` interface:
```typescript
ollamaEnsureTimeoutMs: number;
```
and in `loadConfig`'s returned object:
```typescript
ollamaEnsureTimeoutMs: Number(env.OLLAMA_ENSURE_TIMEOUT_MS ?? 60000),
```
- [ ] **Step 4: Wire the preflight into the routes**
In `backend/src/app.ts`, pass the timeout (seconds) into `sessionRoutes`. Change the call:
```typescript
sessionRoutes(app, { mgr, tht: tht as ThtRunner, hub, getSettings });
```
to:
```typescript
sessionRoutes(app, {
mgr, tht: tht as ThtRunner, hub, getSettings,
ollamaEnsureTimeoutSec: Math.round(config.ollamaEnsureTimeoutMs / 1000),
});
```
In `backend/src/routes/sessions.ts`, extend the deps type and add the preflight. Change the
function signature's deps to include the timeout:
```typescript
export function sessionRoutes(
app: FastifyInstance,
d: { mgr: PiProcessManager; tht: ThtRunner; hub: SseHub; getSettings: () => Settings; ollamaEnsureTimeoutSec: number },
) {
```
At the **start** of the `POST /sessions` handler (before `d.tht.sessionNew`):
```typescript
app.post("/sessions", async (req, reply) => {
const b = req.body as { question: string; name?: string };
const s = d.getSettings();
const ensure = await d.tht.ollamaEnsure(s.workspace, d.ollamaEnsureTimeoutSec);
if (!ensure.ok) return reply.code(503).send({ error: ensure.error ?? "Ollama/embeddings non disponibili" });
// ... existing sessionNew + spawnFor unchanged ...
```
At the **start** of the `POST /sessions/:id/resume` handler (before the manifest read / guard):
```typescript
app.post("/sessions/:id/resume", async (req, reply) => {
const id = (req.params as any).id;
const ensure = await d.tht.ollamaEnsure(d.getSettings().workspace, d.ollamaEnsureTimeoutSec);
if (!ensure.ok) return reply.code(503).send({ error: ensure.error ?? "Ollama/embeddings non disponibili" });
// ... existing manifest read + finalized/archived guard + mgr.resume unchanged ...
```
**Note — keep the pre-existing tests green.** The preflight now runs before any create/resume
logic, so every injected `thtRunner` double that hits `POST /sessions` or `/resume` needs an
`ollamaEnsure`, else it throws `d.tht.ollamaEnsure is not a function`. Two edits cover all cases:
1. **Give `mutApp` a default `ollamaEnsure`** so its tests (including the resume-guard 409 tests
from the session-management feature) pass the preflight and still reach their logic. Change the
helper so the injected runner is `{ ollamaEnsure: async () => ({ ok: true }), ...thtRunner }`:
```typescript
function mutApp(thtRunner: any) {
return buildApp(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
thtRunner: { ollamaEnsure: async () => ({ ok: true }), ...thtRunner },
getSettings: () => ({ workspace: "w" }) as any,
spawnFn: () => nodeSpawn("node", [FAKE, SCRIPT]) as any,
});
}
```
2. **Add `ollamaEnsure: async () => ({ ok: true })`** to the two top-of-file inline doubles that
POST `/sessions`: the "POST /sessions usa i settings…" test and the
"POST /sessions/:id/response inoltra al bridge…" test.
- [ ] **Step 5: Run test to verify it passes**
Run: `cd backend && npx vitest run test/routes-sessions.test.ts` then `npx tsc --noEmit -p .`
Expected: PASS (new 503/success tests + the updated pre-existing tests), typecheck clean.
- [ ] **Step 6: Run the full backend suite (catch regressions)**
Run: `cd backend && npx vitest run`
Expected: PASS. The resume-guard 409 tests from the session-management feature use `mutApp`, so the
default `ollamaEnsure` added in Step 4 lets the preflight pass and the 409 guard is still reached.
If any other test that POSTs create/resume was missed, give its `thtRunner` double an
`ollamaEnsure: async () => ({ ok: true })`.
- [ ] **Step 7: Commit**
```bash
git add backend/src/config.ts backend/src/app.ts backend/src/routes/sessions.ts backend/test/routes-sessions.test.ts
git commit -m "feat(backend): Ollama embeddings preflight on session create/resume (503 hard-fail)"
```
---
## Final verification
- [ ] **Harness:** `cd harness && .venv/bin/pytest tests/test_ollama_ensure.py -q` → all pass; `.venv/bin/ruff check tht/cli/ollama_cmd.py` → clean.
- [ ] **Backend:** `cd backend && npx vitest run` → all pass; `npx tsc --noEmit -p .` → clean.
- [ ] **Manual smoke (optional, needs the stack):** with Ollama stopped, `tht ollama ensure --json -c workspaces/psd.yaml` starts it and warms the model (exit 0); with the model uninstalled it exits 1 with `ollama pull` guidance; creating a session via the UI while Ollama is down returns a 503 with the diagnostic.
## Spec coverage check
- `bin`/`start_cmd` parameterization → Task 1.
- Hard-fail matrix (config/server/model/warm) → Task 1 (`ensure_ollama`) + Task 2 (exit codes).
- `tht ollama ensure` command, pristine `--json`, `--no-start` → Task 2.
- `ThtRunner.ollamaEnsure` parsing + non-zero mapping → Task 3.
- Preflight-before-spawn, 503, no degraded session, timeout env → Task 4.
- Warm-only / no auto-pull → enforced by `ensure_ollama` (model-absent = error) — Task 1.
- Frontend: none (the 503 reaches the existing error path) — no task, by design.

Some files were not shown because too many files have changed in this diff Show More