merge: reconcile GitHub main into canonical Gitea main
Publish documentation / publish (push) Successful in 33s

This commit is contained in:
Codex
2026-08-26 13:36:33 +02:00
521 changed files with 15847 additions and 68853 deletions
@@ -1,66 +0,0 @@
# Task 9 quality audit — final 5
**Scope:** the two blocking findings from `task9-quality-audit-final4.md` — unbound production
module graph at manual serve, and commit-addressed snapshots accepted without content identity at
render. Manual acceptance remains **PENDING**; no `VERDICT.md` was created.
## Verdict: APPROVED for the two final integrity blockers
### 1. Manual serve binds the complete `backend/dist` module graph, not only `server.js`
`prepare` now builds a post-build manifest of every regular `backend/dist` file
(relative path, size, SHA-256, device, inode) and writes it as an exclusive `0600` record
(`installation/runtime/backend-dist.manifest.json`) inside the owned root; `ownership.json`
records that record's path/device/inode/size/SHA-256. `serve` revalidates the manifest record
identity and bytes, revalidates every distribution file against it (no-follow, single inode,
size and digest), and refuses before spawning. The manifest descriptor is passed to the child on
fd 4 together with the entrypoint on fd 3. The immutable preload parses the manifest, verifies
the entrypoint cross-digest, reads and hash-verifies **every** file at startup, caches the
verified bytes, and its load hook serves **only** those cached bytes for any import below
`backend/dist` (entry URL still served from the bound fd-3 bytes). A same-path regular
replacement of any imported dependency is therefore refused before `RUNNING` (serve-time
validation), refused at child startup (startup verification), or rendered harmless (cached
bytes), and the parent revalidates the full manifest at `RUNNING` publication and at `stop`.
### 2. Renderer binds snapshot content to its commit identity
The generated render command validates the bounded saved read/publish revisions, the
commit-addressed owned snapshot path, the installed Git HEAD, and the bounded
`snapshot.json` manifest of that commit: `head` equals the commit, `files[<id>.yaml]` is the
SHA-256 of the snapshot bytes, the manifest revision binds commit/blob/snapshot path, the saved
revision blob equals the manifest blob, and `git rev-parse <commit>:workspaces/<id>.yaml` plus
`git hash-object` of the snapshot bytes both equal that blob. It passes the expected digest as
`--snapshot-sha256`. The renderer re-reads the bounded `snapshot.json` (`head`,
`files[<id>.yaml]` must equal the carried digest), opens the snapshot once with no-follow
semantics and bounded reads, renders only the digest-verified bytes, re-verifies around lease
publication, releases the lease in `finally`, and publishes no output on any refusal.
## Deterministic regressions added
- static regular replacement of an imported production dependency after `prepare` is refused,
no marker, no accepted PID record, no orphan;
- deterministic dependency check/load swap (`beforeSpawn` rename) is refused by the child's
startup verification, no marker, no PID record, no orphan;
- after `RUNNING`, a same-path regular dependency replacement is never executed: the loader
serves the verified cached bytes (health-visible source stays the original) and the marker is
absent;
- renderer refuses a same-path regular snapshot byte replacement against the carried digest and
manifest, with lease release and no output;
- renderer refuses manifest `head`, `files` digest, expected-digest, missing, and malformed
cases, with lease release and no output;
- wrapper refuses missing manifest, manifest head/digest/revision tampering, saved-revision blob
mismatch, Git blob mismatch, and snapshot-vs-Git-bytes mismatch, and passes the exact
`--snapshot-sha256` on the valid path (stub renderer records arguments).
## Verification
- `bash scripts/test-p1-manual-acceptance.sh` (backend build + both suites): **59 tests, 59
pass, 0 fail**; no `8791/8792` listener and no `--p1-manual-nonce` process remain.
- `npx tsc --noEmit -p .` (backend): PASS.
- Real-repository `prepare` + `cleanup` cycle: 39 distribution files bound, entrypoint
cross-digest verified, owned root fully removed afterwards.
- Diff check: only the seven Task 9 paths are touched; no Task 8 file was modified.
- This report and the implementation contain no fixture secret or canary values.
Manual acceptance remains **PENDING** by design; the walkthrough and human verdict are
unchanged.
-282
View File
@@ -1,282 +0,0 @@
{
"schema": "thothii-task4-certification-v1",
"generated_on": "2026-08-18",
"started_at_utc": "2026-08-18T14:16:40Z",
"ended_at_utc": "2026-08-18T14:20:10Z",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"source_immutability": {
"status": "PASS",
"tracked_changes_after_freeze": false,
"allowed_untracked": [".playwright-cli/", ".thothctl/"]
},
"source_commits": {
"task4_candidate": "b31b27e5845ffd3adf311429367319beaba263c7",
"task1": "d43738eeae6d14bb5e470093058b069a983f5372",
"task2": "5f9a3ae066a060b43a11a959b60a1efadd1c2425",
"task3": "0d8e707533fada938c99eb06f8457150e7ef2b40",
"task3_follow_up": "b31b27e5845ffd3adf311429367319beaba263c7",
"fix_round_1_source": "10cd66fe6a5b484a4dc569326a228c1c5484a5d4",
"fix_round_2_source": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"historical_task15_final": "74b062f1a737103524cbe706346cfd65f87cdfd1"
},
"versions": {
"node_contract": "v24.16.0",
"node_host_default": "v25.6.1",
"go": "go1.26.5",
"pi": "0.80.3"
},
"retained_report": ".superpowers/sdd/2026-08-16-thothii-authentication/task-15-report.md",
"task4_report": ".superpowers/sdd/2026-08-18-thothii-authentication-remediation/task-4-report.md",
"fix_round_2_report": ".superpowers/sdd/2026-08-18-thothii-authentication-remediation/fix-round-2-report.md",
"workflow": {
"run_id": "32147345625",
"url": "https://github.com/mptyl/ThothII/actions/runs/32147345625",
"event": "workflow_dispatch",
"head_sha": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"status": "completed",
"conclusion": "failure",
"windows_job": {
"name": "Windows clone and Compose contract",
"job_id": "95744249248",
"url": "https://github.com/mptyl/ThothII/actions/runs/32147345625/job/95744249248",
"conclusion": "failure",
"native_step": "Run native Windows retained-capability tests",
"native_step_conclusion": "success",
"command": "go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1",
"requested_packages": ["internal/safeio", "internal/backup", "internal/authstorage"],
"executed_packages": ["internal/safeio", "internal/backup", "internal/authstorage"],
"not_executed_packages": [],
"package_results": {
"internal/safeio": "PASS (22.058s)",
"internal/backup": "PASS (7.161s)",
"internal/authstorage": "PASS (16.088s)"
},
"failed_step": "Verify Windows clone contract",
"failure_category": "baseline_powershell_parser",
"failure_detail": "scripts/test-windows-clone-contract.ps1:208 parses $remoteYaml: as an invalid variable reference"
},
"lf_compose_docs_typescript_job": {
"name": "LF, Compose, docs, and TypeScript",
"job_id": "95744249458",
"url": "https://github.com/mptyl/ThothII/actions/runs/32147345625/job/95744249458",
"conclusion": "failure",
"failed_step": "Verify Compose and installation contracts",
"category": "baseline_ci_contract",
"detail": "unified Compose contract passed; test-no-deployment-coupling-scope.sh stopped on TMPDIR: unbound variable",
"downstream_steps": "skipped"
},
"linux_docker_job": {
"name": "Linux Docker deployment and rollback",
"job_id": "95744249354",
"url": "https://github.com/mptyl/ThothII/actions/runs/32147345625/job/95744249354",
"conclusion": "failure",
"failed_step": "Run unified deployment smoke",
"category": "infrastructure_prerequisite",
"detail": "Task 13 smoke failed before deployment because rg is required",
"cleanup": "PASS",
"image_manifest": "not_generated"
},
"windows_docker_startup_job": {
"name": "Native Windows Docker Desktop/WSL2 startup",
"job_id": "95744250450",
"url": "https://github.com/mptyl/ThothII/actions/runs/32147345625/job/95744250450",
"status": "NOT_RUN",
"classification": "BLOCKED",
"workflow_conclusion": "skipped",
"reason": "workflow conditions skipped the job; no Windows Docker Desktop/WSL2 command executed"
}
},
"docker_image_evidence": {
"authentication_smoke": {
"status": "PASS",
"docker_images": [],
"reason": "no_docker_images_exercised"
},
"unified_docker_smoke": {
"status": "FAIL",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"run_id": "32147345625",
"workflow_job_id": "95744249354",
"manifest": ".artifacts/task-15/unified-docker-images.json",
"reason": "workflow attempt stopped before deployment because rg is required",
"cleanup": "PASS",
"images": 0,
"historical": {
"status": "PASS",
"source_commit": "74b062f1a737103524cbe706346cfd65f87cdfd1",
"run_id": "20260818070637-66409-30058",
"manifest_sha256": "9c8dec4546909fd93799dbcf374bcb3a89bc46cfe0fd482472c0cbe757ddf5b6",
"images": 5,
"cleanup": "PASS"
}
}
},
"gates": {
"posix_registry_ownership": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"evidence": "backend Node 24 full suite including local-registry ownership coverage"
},
"stagearchive_unix_retained_capability": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"evidence": "focused safeio/backup tests, Go race suite, and Unix ancestor-swap coverage"
},
"windows_stagearchive_retained_capability": {
"status": "PASS",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"evidence": "native Windows backup package passed, including the two-file shared retained-root staging test"
},
"windows_claim_retained_capability": {
"status": "PASS",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"evidence": "native Windows safeio and authstorage packages passed concurrent claim/consume coverage"
},
"workflow_lf_compose_docs_typescript": {
"status": "FAIL",
"classification": "baseline_ci_contract",
"reason": "TMPDIR was unset after the unified Compose contract passed"
},
"workflow_linux_docker": {
"status": "FAIL",
"classification": "infrastructure_prerequisite",
"reason": "runner did not provide rg; cleanup proof passed and no image manifest was generated"
},
"go_security_build": {
"status": "PASS",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"focused_packages": 3,
"race_packages": 18,
"focused_test": "PASS",
"race": "PASS",
"vet": "PASS",
"host_build": "PASS"
},
"windows_cross_compile": {
"status": "PASS",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"focused_test_packages": 3,
"cli_build": "PASS",
"execution": "cross_compile_only_not_native_execution"
},
"backend_node24": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"node": "v24.16.0",
"files": 76,
"tests": 1092,
"typecheck": "PASS",
"build": "PASS",
"note": "an initial full run had one workspace-registry timeout; focused rerun and complete rerun passed"
},
"frontend_node24": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"node": "v24.16.0",
"files": 61,
"tests": 444,
"typecheck": "PASS",
"build": "PASS"
},
"authentication_and_f1_smoke": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"node": "v24.16.0",
"filtered_e2e": "1 passed",
"sentinel_leak_scan": "PASS"
},
"harness_pytest": {
"status": "FAIL",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"passed": 951,
"failed": 1,
"skipped": 4,
"subtests": 232,
"failure": "test_column_decisions::test_f4_emits_column_types: workflow.yaml not found from harness test cwd"
},
"authentication_docs": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7"
},
"shell_syntax": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7"
},
"authentication_smoke_runtime": {
"status": "PASS",
"node": "v24.16.0",
"sentinel_leak_scan": "PASS"
},
"compose_default": {
"status": "FAIL",
"reason": "required THT_WORKSPACE_GIT_REMOTE was not available"
},
"compose_unified": {
"status": "FAIL",
"reason": "compose.unified.yaml is absent from the frozen source"
},
"unified_docker_smoke": {
"status": "FAIL",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"workflow_run_id": "32147345625",
"reason": "remote workflow attempted the smoke but stopped before deployment because rg is required",
"cleanup": "PASS",
"image_manifest": "not_generated"
},
"ruff": {
"status": "FAIL",
"errors": 192,
"classification": "known_baseline"
},
"mkdocs_strict": {
"status": "NOT_RUN",
"classification": "BLOCKED",
"historical_status": "FAIL",
"historical_warnings": 69
},
"canonical_install_docs": {
"status": "NOT_RUN",
"classification": "BLOCKED",
"historical_status": "FAIL"
},
"workspace_install_docs": {
"status": "NOT_RUN",
"classification": "BLOCKED",
"historical_status": "FAIL"
},
"pi_user_auth_compose": {
"status": "NOT_RUN",
"classification": "BLOCKED",
"historical_status": "FAIL"
},
"deployment_coupling": {
"status": "NOT_RUN",
"classification": "BLOCKED",
"historical_status": "FAIL"
},
"l2": {
"status": "PENDING",
"reason": "configured secret layout unavailable; gate not run after stop"
},
"manual_psd": {
"status": "PENDING",
"reason": "approved real identity/access unavailable; gate not run after stop"
},
"provider_readiness": {
"status": "PENDING",
"reason": "provider prerequisite unavailable; gate not run after stop"
}
},
"review": {
"original_important_findings_resolved": 3,
"fix_round_2_important_lifecycle": "ADDRESSED",
"fix_round_2_minor_windows_diagnostics": "ADDRESSED",
"verdict": "PASS",
"reason": "the lifecycle controller is bounded and cancellation-aware with cancel, bounded join, and lock-release proof; the temporary Windows diagnostic matrix is removed; exact-source native safeio, backup, and authstorage all pass"
},
"remediation_status": "PASS",
"release_complete": false,
"authentication_implementation_complete": true,
"release_readiness": "FAIL",
"release_readiness_pending_external_gates": true
}
@@ -1,54 +0,0 @@
{
"gate": "unified-deployment-smoke",
"status": "pass",
"source_commit": "74b062f1a737103524cbe706346cfd65f87cdfd1",
"run_id": "20260818070637-66409-30058",
"images": [
{
"id": "sha256:2d7b19491c7eb8c119c3cedb390aaeb2ff5593f6fc43ab66c317565560da6d7d",
"roles": [
"compose-runtime",
"fixture-runtime"
],
"repo_digests": [
"sha256:2d7b19491c7eb8c119c3cedb390aaeb2ff5593f6fc43ab66c317565560da6d7d"
]
},
{
"id": "sha256:3b6c31a5d8f8fc58fa3233391b6175bd2fbc793eebb44d5e285ecc6e02e9e687",
"roles": [
"compose-runtime"
],
"repo_digests": [
"sha256:3b6c31a5d8f8fc58fa3233391b6175bd2fbc793eebb44d5e285ecc6e02e9e687"
]
},
{
"id": "sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a",
"roles": [
"compose-runtime"
],
"repo_digests": [
"sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a"
]
},
{
"id": "sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c",
"roles": [
"compose-runtime"
],
"repo_digests": [
"sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c"
]
},
{
"id": "sha256:c3cbe1cc1aa588a64951ac6286e0df7b27fe2e6324b1001c619bb358770c0178",
"roles": [
"rollback-candidate"
],
"repo_digests": [
"sha256:c3cbe1cc1aa588a64951ac6286e0df7b27fe2e6324b1001c619bb358770c0178"
]
}
]
}
-11
View File
@@ -1,11 +0,0 @@
{
"version": "0.0.1",
"configurations": [
{
"name": "replay",
"runtimeExecutable": "node",
"runtimeArgs": ["tools/replay/server.mjs"],
"port": 5333
}
]
}
-2
View File
@@ -27,5 +27,3 @@ coverage/
data/
sessions/
workspace-registry/
# docs/site (mkdocs build) — non necessari nelle immagini
docs/superpowers/plans
+24 -1
View File
@@ -36,6 +36,10 @@ jobs:
with:
node-version: "24.16.0"
package-manager-cache: false
- name: Install release gate prerequisites
run: |
sudo apt-get update
sudo apt-get install --yes --no-install-recommends ripgrep
- name: Verify shell syntax and LF policy
run: |
git ls-files -z '*.sh' | xargs -0 -n1 bash -n
@@ -46,7 +50,6 @@ jobs:
bash scripts/test-no-deployment-coupling-scope.sh
bash scripts/test-compose-secret-policy.sh
bash scripts/test-no-deployment-coupling.sh
bash scripts/test-preprocess-compose-config.sh
bash scripts/test-verify-workspace-install-docs.sh
git diff --check
- name: Assert clean checkout before release trust bootstrap
@@ -60,6 +63,14 @@ jobs:
run: |
bash scripts/test-server-pi-state-topology.sh
bash scripts/unified-deployment-smoke.sh --self-test
- name: Install harness CLI for backend integration tests
working-directory: harness
run: |
python3 -m venv .venv
.venv/bin/python -m pip install -e .
- name: Install backend dependencies
working-directory: backend
run: npm ci
- name: Test and type-check backend
working-directory: backend
run: |
@@ -140,11 +151,23 @@ jobs:
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
- name: Install release gate prerequisites
run: |
sudo apt-get update
sudo apt-get install --yes --no-install-recommends ripgrep
- name: Reclaim unused hosted-runner space
run: bash scripts/prepare-linux-docker-runner.sh
- name: Run unified deployment smoke
env:
TASK13_IMAGE_EVIDENCE_OUTPUT: ${{ runner.temp }}/task13-images.json
run: timeout --signal=TERM --kill-after=45s 32m bash scripts/unified-deployment-smoke.sh
- name: Run tht update smoke
env:
TASK13_IMAGE_EVIDENCE_OUTPUT: ${{ runner.temp }}/task13-images.json
run: timeout --signal=TERM --kill-after=45s 32m bash scripts/tht-update-smoke.sh
- name: Run Linux server deployment smoke
env:
TASK13_IMAGE_EVIDENCE_OUTPUT: ${{ runner.temp }}/task13-images.json
run: timeout --signal=TERM --kill-after=45s 32m bash scripts/server-deployment-smoke.sh
windows-clone:
-2
View File
@@ -5,8 +5,6 @@
ChironeWp3/
Thoth/
# === Visual companion brainstorming artifacts (local-only) ===
.superpowers/
.worktrees/
.tht/
+203
View File
@@ -0,0 +1,203 @@
{
"schemaVersion": 2,
"generatedAt": "2026-08-26T10:10:15.926Z",
"title": "Design System: ThothII",
"extensions": {
"colorMeta": {
"instrument-red": {
"role": "primary",
"displayName": "Instrument Red",
"canonical": "oklch(55.87% 0.1881 23.2)",
"tonalRamp": ["oklch(15% 0.07 23.2)", "oklch(28% 0.12 23.2)", "oklch(42% 0.16 23.2)", "oklch(56% 0.1881 23.2)", "oklch(68% 0.17 23.2)", "oklch(78% 0.13 23.2)", "oklch(88% 0.07 23.2)", "oklch(95% 0.03 23.2)"]
},
"instrument-red-hover": {
"role": "primary",
"displayName": "Instrument Red Pressed",
"canonical": "oklch(50.95% 0.1812 24.1)",
"tonalRamp": ["oklch(15% 0.07 24.1)", "oklch(28% 0.12 24.1)", "oklch(42% 0.16 24.1)", "oklch(51% 0.1812 24.1)", "oklch(68% 0.16 24.1)", "oklch(78% 0.12 24.1)", "oklch(88% 0.07 24.1)", "oklch(95% 0.03 24.1)"]
},
"porcelain-background": {
"role": "neutral",
"displayName": "Porcelain Background",
"canonical": "oklch(99.18% 0.0011 17.2)",
"tonalRamp": ["oklch(15% 0.0011 17.2)", "oklch(28% 0.0011 17.2)", "oklch(42% 0.0011 17.2)", "oklch(56% 0.0011 17.2)", "oklch(68% 0.0011 17.2)", "oklch(78% 0.0011 17.2)", "oklch(88% 0.0011 17.2)", "oklch(95% 0.0011 17.2)"]
},
"porcelain-card": {
"role": "neutral",
"displayName": "Porcelain Card",
"canonical": "oklch(99.85% 0.0006 17.2)",
"tonalRamp": ["oklch(15% 0.0006 17.2)", "oklch(28% 0.0006 17.2)", "oklch(42% 0.0006 17.2)", "oklch(56% 0.0006 17.2)", "oklch(68% 0.0006 17.2)", "oklch(78% 0.0006 17.2)", "oklch(88% 0.0006 17.2)", "oklch(95% 0.0006 17.2)"]
},
"warm-surface": {
"role": "neutral",
"displayName": "Warm Surface",
"canonical": "oklch(97.09% 0.0011 17.2)",
"tonalRamp": ["oklch(15% 0.0011 17.2)", "oklch(28% 0.0011 17.2)", "oklch(42% 0.0011 17.2)", "oklch(56% 0.0011 17.2)", "oklch(68% 0.0011 17.2)", "oklch(78% 0.0011 17.2)", "oklch(88% 0.0011 17.2)", "oklch(95% 0.0011 17.2)"]
},
"sunken-surface": {
"role": "neutral",
"displayName": "Sunken Surface",
"canonical": "oklch(94.08% 0.0011 17.2)",
"tonalRamp": ["oklch(15% 0.0011 17.2)", "oklch(28% 0.0011 17.2)", "oklch(42% 0.0011 17.2)", "oklch(56% 0.0011 17.2)", "oklch(68% 0.0011 17.2)", "oklch(78% 0.0011 17.2)", "oklch(88% 0.0011 17.2)", "oklch(95% 0.0011 17.2)"]
},
"warm-graphite": {
"role": "neutral",
"displayName": "Warm Graphite",
"canonical": "oklch(26.78% 0.0097 355.6)",
"tonalRamp": ["oklch(15% 0.0097 355.6)", "oklch(28% 0.0097 355.6)", "oklch(42% 0.0097 355.6)", "oklch(56% 0.0097 355.6)", "oklch(68% 0.008 355.6)", "oklch(78% 0.006 355.6)", "oklch(88% 0.004 355.6)", "oklch(95% 0.002 355.6)"]
},
"muted-graphite": {
"role": "neutral",
"displayName": "Muted Graphite",
"canonical": "oklch(51.33% 0.0088 345.6)",
"tonalRamp": ["oklch(15% 0.0088 345.6)", "oklch(28% 0.0088 345.6)", "oklch(42% 0.0088 345.6)", "oklch(56% 0.0088 345.6)", "oklch(68% 0.007 345.6)", "oklch(78% 0.005 345.6)", "oklch(88% 0.003 345.6)", "oklch(95% 0.002 345.6)"]
},
"quiet-border": {
"role": "neutral",
"displayName": "Quiet Border",
"canonical": "oklch(90.93% 0.0035 354.7)",
"tonalRamp": ["oklch(15% 0.0035 354.7)", "oklch(28% 0.0035 354.7)", "oklch(42% 0.0035 354.7)", "oklch(56% 0.0035 354.7)", "oklch(68% 0.0035 354.7)", "oklch(78% 0.0035 354.7)", "oklch(88% 0.003 354.7)", "oklch(95% 0.002 354.7)"]
},
"success-mint": {
"role": "secondary",
"displayName": "Success Mint",
"canonical": "oklch(75.77% 0.1581 165)",
"tonalRamp": ["oklch(15% 0.06 165)", "oklch(28% 0.1 165)", "oklch(42% 0.14 165)", "oklch(56% 0.1581 165)", "oklch(68% 0.15 165)", "oklch(78% 0.12 165)", "oklch(88% 0.07 165)", "oklch(95% 0.03 165)"]
},
"warning-amber": {
"role": "tertiary",
"displayName": "Warning Amber",
"canonical": "oklch(85.23% 0.1386 78.9)",
"tonalRamp": ["oklch(15% 0.05 78.9)", "oklch(28% 0.09 78.9)", "oklch(42% 0.12 78.9)", "oklch(56% 0.1386 78.9)", "oklch(68% 0.13 78.9)", "oklch(78% 0.1 78.9)", "oklch(88% 0.06 78.9)", "oklch(95% 0.025 78.9)"]
},
"information-blue": {
"role": "tertiary",
"displayName": "Information Blue",
"canonical": "oklch(70.35% 0.1128 221.3)",
"tonalRamp": ["oklch(15% 0.045 221.3)", "oklch(28% 0.075 221.3)", "oklch(42% 0.1 221.3)", "oklch(56% 0.1128 221.3)", "oklch(68% 0.105 221.3)", "oklch(78% 0.08 221.3)", "oklch(88% 0.045 221.3)", "oklch(95% 0.02 221.3)"]
}
},
"typographyMeta": {
"display": {"displayName": "Display", "purpose": "Authentication and exceptional page-level statements only."},
"headline": {"displayName": "Headline", "purpose": "Major page and persisted artifact titles."},
"title": {"displayName": "Title", "purpose": "Panel and document section hierarchy."},
"body": {"displayName": "Body", "purpose": "Operational prose and sustained reading."},
"control": {"displayName": "Control", "purpose": "Buttons, inputs, tabs, and compact actions."},
"label": {"displayName": "Machine Label", "purpose": "Uppercase metadata and machine-oriented micro-labels."}
},
"shadows": [
{"name": "contact", "value": "0 1px 2px oklch(var(--shadow-tint) / 0.05)", "purpose": "Contact shadow for controls and code blocks."},
{"name": "panel", "value": "0 1px 2px oklch(var(--shadow-tint) / 0.05), 0 2px 6px -1px oklch(var(--shadow-tint) / 0.05)", "purpose": "Small structural lift for selected cards."},
{"name": "overlay", "value": "0 2px 4px -2px oklch(var(--shadow-tint) / 0.06), 0 12px 32px -8px oklch(var(--shadow-tint) / 0.1)", "purpose": "Broad low-opacity lift for dialogs and floating layers."}
],
"motion": [
{"name": "control-feedback", "value": "140ms cubic-bezier(0.22, 1, 0.36, 1)", "purpose": "Button hover, focus, and press feedback."},
{"name": "overlay-transition", "value": "100ms ease-out", "purpose": "Dialog fade and scale transitions."},
{"name": "activity-pulse", "value": "1.5s ease-in-out infinite", "purpose": "Live model activity only; disabled for reduced motion."}
],
"breakpoints": [
{"name": "sm", "value": "640px"},
{"name": "lg", "value": "1024px"}
]
},
"components": [
{
"name": "Primary Button",
"kind": "button",
"refersTo": "button-primary",
"description": "The authoritative action for the current workflow step.",
"html": "<button class=\"ds-button-primary\">Confirm review</button>",
"css": ".ds-button-primary { display:inline-flex; align-items:center; justify-content:center; height:32px; padding:0 14px; border:1px solid transparent; border-radius:8px; background:oklch(var(--primary)); color:oklch(var(--primary-foreground)); font:600 14px/1.25 var(--font-sans); letter-spacing:0.005em; box-shadow:var(--shadow-xs); transition:color 140ms cubic-bezier(0.22,1,0.36,1),background-color 140ms cubic-bezier(0.22,1,0.36,1),box-shadow 140ms cubic-bezier(0.22,1,0.36,1),transform 140ms cubic-bezier(0.22,1,0.36,1); } .ds-button-primary:hover { background:oklch(var(--primary-hover)); } .ds-button-primary:focus-visible { outline:3px solid oklch(var(--ring)/0.25); outline-offset:2px; } .ds-button-primary:active { transform:scale(0.97); box-shadow:none; }"
},
{
"name": "Outline Button",
"kind": "button",
"refersTo": "button-secondary",
"description": "A compact secondary action that preserves the primary action hierarchy.",
"html": "<button class=\"ds-button-outline\">Inspect details</button>",
"css": ".ds-button-outline { display:inline-flex; align-items:center; justify-content:center; height:32px; padding:0 14px; border:1px solid oklch(var(--border)); border-radius:8px; background:oklch(var(--card)); color:oklch(var(--foreground)); font:600 14px/1.25 var(--font-sans); box-shadow:var(--shadow-xs); transition:background-color 140ms cubic-bezier(0.22,1,0.36,1),transform 140ms cubic-bezier(0.22,1,0.36,1); } .ds-button-outline:hover { background:oklch(var(--muted)); } .ds-button-outline:focus-visible { outline:3px solid oklch(var(--ring)/0.25); outline-offset:2px; } .ds-button-outline:active { transform:scale(0.97); box-shadow:none; }"
},
{
"name": "Status Badge",
"kind": "chip",
"refersTo": "badge-primary",
"description": "A compact state label that always carries readable text.",
"html": "<span class=\"ds-status-badge\">Ready for review</span>",
"css": ".ds-status-badge { display:inline-flex; align-items:center; height:20px; padding:2px 8px; border:1px solid transparent; border-radius:6px; background:oklch(var(--primary)); color:oklch(var(--primary-foreground)); font:600 12px/1.25 var(--font-sans); white-space:nowrap; } .ds-status-badge:focus-visible { outline:3px solid oklch(var(--ring)/0.5); outline-offset:2px; }"
},
{
"name": "Text Field",
"kind": "input",
"refersTo": "input-default",
"description": "A readable operational field with an explicit focus state.",
"html": "<input class=\"ds-text-field\" value=\"Fascia pediatrica\" aria-label=\"Session name\">",
"css": ".ds-text-field { width:280px; height:40px; padding:0 12px; border:1px solid oklch(var(--input)); border-radius:8px; background:oklch(var(--background)); color:oklch(var(--foreground)); font:400 14px/1.5 var(--font-sans); outline:none; } .ds-text-field:hover { border-color:oklch(var(--muted-foreground)/0.65); } .ds-text-field:focus-visible { border-color:oklch(var(--ring)); box-shadow:0 0 0 3px oklch(var(--ring)/0.25); } .ds-text-field:disabled { opacity:0.5; cursor:not-allowed; }"
},
{
"name": "Work Card",
"kind": "card",
"refersTo": "card-default",
"description": "A single-level container for a coherent review surface.",
"html": "<section class=\"ds-work-card\"><h3>Schema linking</h3><p>Review the linked tables and columns before continuing.</p></section>",
"css": ".ds-work-card { width:320px; padding:16px; border:1px solid oklch(var(--border)/0.7); border-radius:12px; background:oklch(var(--card)); color:oklch(var(--card-foreground)); box-shadow:var(--shadow-sm); } .ds-work-card h3 { margin:0 0 8px; font:500 16px/1.35 var(--font-heading); letter-spacing:-0.01em; } .ds-work-card p { margin:0; color:oklch(var(--muted-foreground)); font:400 14px/1.6 var(--font-sans); } .ds-work-card:focus-within { outline:3px solid oklch(var(--ring)/0.25); outline-offset:2px; }"
},
{
"name": "Session Navigation Item",
"kind": "nav",
"description": "A dense session row with restrained hover and active hierarchy.",
"html": "<button class=\"ds-session-item\"><span class=\"ds-session-dot\"></span><span><strong>Patient cohorts</strong><small>Schema linking</small></span></button>",
"css": ".ds-session-item { display:flex; width:260px; align-items:center; gap:8px; padding:4px 8px; border:0; border-radius:8px; background:transparent; color:oklch(var(--foreground)); text-align:left; font-family:var(--font-sans); transition:background-color 140ms cubic-bezier(0.22,1,0.36,1); } .ds-session-item:hover,.ds-session-item[aria-current=\"page\"] { background:oklch(var(--accent)); } .ds-session-item:focus-visible { outline:2px solid oklch(var(--ring)/0.4); outline-offset:1px; } .ds-session-dot { width:6px; height:6px; flex:none; border-radius:9999px; background:oklch(var(--success)); } .ds-session-item strong,.ds-session-item small { display:block; } .ds-session-item strong { font-size:13px; font-weight:600; } .ds-session-item small { margin-top:2px; color:oklch(var(--muted-foreground)); font-size:11px; }"
},
{
"name": "Curated Evidence Document",
"kind": "custom",
"description": "The table-free reading hierarchy for persisted evidence.",
"html": "<article class=\"ds-evidence\"><h2>Fascia pediatrica</h2><div class=\"ds-evidence-summary\"><strong>Dominio</strong> · Italiano<br><span>Scopi: Disambiguazione · Generazione SQL</span></div><h3>Ambito di applicazione</h3><ul><li>fascia pediatrica</li><li>paziente minore</li></ul><h3>Regola</h3><p>La fascia pediatrica comprende i pazienti con età inferiore a 18 anni.</p><details><summary>Dettagli tecnici e provenienza</summary><code>evidence:fascia-pediatrica</code></details></article>",
"css": ".ds-evidence { max-width:70ch; color:oklch(var(--foreground)); font:400 15px/1.65 var(--font-sans); } .ds-evidence h2,.ds-evidence h3 { font-family:var(--font-heading); letter-spacing:-0.01em; } .ds-evidence h2 { margin:0 0 16px; font-size:24px; } .ds-evidence h3 { margin:24px 0 8px; font-size:18px; } .ds-evidence-summary { padding:12px 14px; border:1px solid oklch(var(--border)); border-radius:8px; background:oklch(var(--muted)); color:oklch(var(--muted-foreground)); } .ds-evidence-summary strong { color:oklch(var(--foreground)); } .ds-evidence ul { padding-left:20px; } .ds-evidence details { margin-top:24px; padding:10px 12px; border:1px solid oklch(var(--border)); border-radius:8px; background:oklch(var(--card)); } .ds-evidence summary { cursor:pointer; font-weight:600; } .ds-evidence code { font-family:var(--font-mono); }"
}
],
"narrative": {
"northStar": "The Clinical Workbench",
"overview": "ThothII should feel like a well-kept clinical workbench: warm enough for sustained reading, exact enough for consequential review, and quiet enough that evidence, state, and decisions remain in the foreground. The visual system is calm, precise, and trustworthy. It uses familiar product patterns, restrained color, and deliberate density instead of decorative spectacle.\n\nThe primary physical scene is an analyst reviewing persisted evidence and SQL on a large monitor in a well-lit working environment. This makes the warm light theme the default. The supported dark theme serves lower-light work without becoming a separate neon aesthetic. Both themes preserve the same hierarchy and semantic roles.\n\nThe system rejects generic SaaS ornament, conspicuous ripples, bounce or elastic motion, long choreographed transitions, and effects that compete with the analytical task. Controls should feel disciplined and tactile, never playful, sluggish, or visually unstable.",
"keyCharacteristics": [
"Warm, restrained surfaces with one scarce red accent.",
"Editorial headings paired with highly legible operational body text.",
"Dense information organized through hierarchy, rhythm, and progressive disclosure.",
"Persisted artifacts and reviewer decisions presented as the visual source of truth.",
"Fast state feedback with reduced-motion parity."
],
"rules": [
{"name": "The Workbench Rule", "body": "Every visual element must support inspection, action, state, or provenance. Decoration without an operational purpose is forbidden.", "section": "overview"},
{"name": "The Persisted Truth Rule", "body": "Persisted artifacts and reviewer decisions receive stronger hierarchy than transient model narration.", "section": "overview"},
{"name": "The Density with Rhythm Rule", "body": "Preserve information density, but vary spacing between groups so users can scan structure without adding nested containers.", "section": "overview"},
{"name": "The One Voice Rule", "body": "Instrument Red should occupy no more than roughly ten percent of a screen. Its rarity is what makes it authoritative.", "section": "colors"},
{"name": "The State Has a Name Rule", "body": "Success, warning, information, and destructive colors are reserved for their named states. Color is never the only state indicator.", "section": "colors"},
{"name": "The Three Registers Rule", "body": "Serif means authority, sans means interaction and reading, mono means machine identity. Do not exchange these roles for novelty.", "section": "typography"},
{"name": "The Read Once Rule", "body": "A heading, label, and body must be distinguishable on first glance through size and weight. Do not repeat headings in explanatory copy.", "section": "typography"},
{"name": "The Flat by Default Rule", "body": "A resting surface has no shadow unless it is physically above another surface. If every panel floats, none of them has hierarchy.", "section": "elevation"},
{"name": "The Borders Structure, Shadows Elevate Rule", "body": "Never use shadow as a substitute for grouping or a border as a decorative accent.", "section": "elevation"},
{"name": "The Review Surface Rule", "body": "The visible Markdown must be readable without understanding the machine contract. Technical metadata belongs in progressive disclosure, not above the title.", "section": "components"}
],
"dos": [
"Do make every state change unmistakable without interrupting flow.",
"Do use Instrument Red only for primary action, current selection, focus identity, or explicit destructive meaning.",
"Do preserve information density with headings, rhythm, and progressive disclosure.",
"Do keep keyboard focus explicit and pair color with text, shape, icon, or position.",
"Do respect prefers-reduced-motion while preserving immediate non-kinetic feedback.",
"Do use English for interface chrome and the workspace language for persisted document content.",
"Do render curated metadata and scope as Markdown prose or lists, never as a frontmatter table."
],
"donts": [
"Don't add generic SaaS ornament, conspicuous ripples, bounce or elastic motion, long choreographed transitions, or effects that compete with the analytical task.",
"Don't make controls feel playful, sluggish, or visually unstable.",
"Don't use gradient text, decorative glassmorphism, or full-saturation accents on inactive states.",
"Don't use a colored side stripe greater than one pixel on cards, callouts, list items, or blockquotes. Use a full border, tonal background, icon, or heading instead.",
"Don't nest cards or wrap every section in a container.",
"Don't use a modal before exhausting inline or progressive alternatives.",
"Don't use tables for applies_to, metadata, enum values, or other one-dimensional content.",
"Don't use color as the sole carrier of success, warning, error, selection, or progress.",
"Don't use display typography for buttons, labels, or data.",
"Don't add em dashes to interface copy. Use commas, colons, semicolons, or parentheses."
]
}
}
-7
View File
@@ -1,7 +0,0 @@
{
"$schema": "https://app.kilo.ai/config.json",
"indexing": {
"vectorStore": "qdrant",
"model": "sentence-transformers/all-minilm-l12-v2"
}
}
@@ -1,105 +0,0 @@
# Task 3 — Diagnostic contract remediation report
Date: 2026-08-04
## Scope
This remediation is limited to the four approved review findings for the workspace diagnostic
extension. It does not add registry routes, change workspace publication, alter session startup,
or expand transport support.
## Changes
1. `RuntimeBindings` now has an explicit `vectorWriter` binding. The new
`resolveRuntimeBindings()` resolves DWH, vector reader, vector writer, and embedding bindings
together. The diagnoser takes the writer credential only from `bindings.vectorWriter`, never
from vector-reader values.
2. Direct PostgreSQL and SSH-tunnelled direct probes accept an absent CA binding while retaining
certificate verification through the runtime system trust store. A supplied CA still uses
verified private-CA trust. REST private-CA refusal is unchanged.
3. A reversible vector probe now requires an authenticated POST declaration with a response map
containing `operation`. The adapter requires the successful JSON response to echo `create` or
`remove` respectively, so an arbitrary 2xx or an upsert-only response cannot activate the
write probe.
4. For DWH and vector REST diagnostics declared with `auth: none`, the resolver no longer
requires an API-key file and the adapter sends no credential. Credential-backed diagnostics
continue to require their local secret file.
## TDD evidence
The first focused RED run failed for the intended missing behavior:
- `resolveRuntimeBindings is not a function` for unauthenticated resolver bindings;
- schema accepted a reversible probe without a response contract; and
- existing diagnostic fixtures rejected the new `response` declaration until schema support was
implemented.
The focused GREEN run passed `43/43` tests across:
- `test/workspaces-bindings.test.ts`
- `test/workspaces-schema.test.ts`
- `test/workspaces-diagnostics.test.ts`
The regression coverage includes resolver-to-diagnoser writer propagation without manually
inserting the writer key into vector-reader bindings, no-CA direct/SSH system-trust requests,
operation-echo validation for create/remove, and `auth: none` bindings without secret files.
## Documentation and design
- `docs/workspace-diagnostic-protocol.md` now documents the verified system-trust fallback,
no-secret `auth: none` behavior, and required reversible response contract.
- `docs/superpowers/specs/2026-08-03-git-workspace-registry-design.md` now records the same
response, CA, SSH, and authentication rules.
## Final verification
The initial sandboxed full suite could not bind its local SSE listener (`listen EPERM:
operation not permitted 127.0.0.1`). It was rerun unchanged with local-listener permission.
```text
backend: npx vitest run
31 test files passed; 329 tests passed
backend: npx tsc --noEmit -p .
exit 0
repository: git diff --check
exit 0
```
Expected test harness stderr from existing Pi/process failure-path tests remained present; no test
failed and no diagnostic secret was emitted.
## Blockers
None.
## Round 2 remediation
The final review found two remaining contract gaps. The binding resolver already treated
`auth: none` as credential-free, but the runtime renderer and diagnostic connector still required
the API-key file. Rendering and connector construction now make that requirement conditional on
the declared REST authentication mode, so a DWH/vector `auth: none` workspace passes resolver,
runtime rendering, and diagnostics with no API-key file.
SSH forwarding previously changed the PostgreSQL connection host to `127.0.0.1` without retaining
the original target for TLS hostname validation. Forwarded probes now carry `SSH_TARGET_HOST` as
`tlsServername` into the PostgreSQL TLS options; private CA and verified system trust behavior are
unchanged.
TDD RED: the new end-to-end no-key test failed at the unconditional runtime
`API_KEY_FILE` requirement, while the SSH test showed no `tlsServername` on the loopback probe or
database-client request. TDD GREEN: the focused backend workspace tests passed `40/40`.
Round 2 final verification:
```text
backend: npx vitest run
31 test files passed; 332 tests passed
backend: npx tsc --noEmit -p .
exit 0
repository: git diff --check
exit 0
```
@@ -1,101 +0,0 @@
# Task 7 report — revision-pinned sessions
## Delivered
- New-session requests may carry `workspaceId`, provider, model, and thinking. The backend
resolves the active operational registry revision, enforces its LLM policy, and persists the
workspace ID/revision with the selected LLM settings.
- The harness manifest and `tht session new` support the optional, backward-compatible
`workspace_id` and `workspace_revision` fields.
- Resume resolves the manifest's retained snapshot, including after later registry publication.
A missing retained revision returns a sanitized `workspace_revision_unavailable` response.
Legacy manifests retain the prior workspace behavior and are marked with a visible warning on
`GET /sessions/:id`.
- `/settings` is now a non-mutating compatibility endpoint: installation defaults remain
readable, while anonymous workspace/provider/model/thinking selections are no longer written
to backend settings or principal preferences.
## TDD evidence
- RED: `npx vitest run test/routes-sessions.test.ts test/routes-settings.test.ts` failed for the
new immutable-snapshot and no-settings-mutation assertions; the manifest test failed because
`new_session_manifest` did not accept workspace revision fields.
- GREEN: `npx vitest run test/tht-runner.test.ts test/routes-sessions.test.ts test/routes-settings.test.ts && npx tsc --noEmit -p .`
completed with 97 passing tests and a clean type check.
- GREEN: `THT_HOME=/private/tmp/thothii-task7-home .venv/bin/pytest tests/test_session_documents.py tests/test_session_mutations.py -q`
completed with 22 passing tests.
- `git diff --check` completed cleanly.
## Review fixes — round 3
- The active registry snapshot that located a session now remains the authorization and mutation
config for response, steer, events, close/delete, archive/group/rename, documents, and detail.
A pruned historical revision cannot block an already-located session's active lifecycle.
- Only Resume resolves the retained pinned descriptor because Pi needs that immutable config to
restart safely. A pruned pin therefore returns the existing sanitized
`workspace_revision_unavailable` 409 solely for Resume.
### Round 3 verification
- RED: with a manifest found through an active registry snapshot and `readPinned` forced to fail,
`POST /sessions/:id/response` returned 409 instead of forwarding the active gate response.
- GREEN: `npx vitest run test/routes-sessions.test.ts test/tht-runner.test.ts test/routes-settings.test.ts && npx tsc --noEmit -p .`
— 102 tests passed with a clean type check. The regression confirms response, close, and delete
use the locating snapshot without calling `readPinned`, while Resume returns a sanitized 409.
- `git diff --check` completed cleanly.
## Review fixes — round 2
- Lifecycle authorization no longer selects the installation-default workspace. The backend now
finds each session by querying every operational registry snapshot with the authenticated
principal, preserving RLS ownership concealment.
- After locating the manifest, durable pinned sessions resolve their retained descriptor before
any lifecycle mutation/reopen. Legacy sessions continue using the locating registry snapshot.
- Session listing aggregates the owner-visible rows from all operational registry snapshots;
detail, response, steer, resume, events, documents, and lifecycle mutations use the same
server-side locator. No route depends on browser-local workspace state.
### Round 2 verification
- RED: the new cross-workspace route integration test created a B session while installation
default A was selected, then demonstrated that `GET /sessions` returned an empty list.
- GREEN: `npx vitest run test/routes-sessions.test.ts test/tht-runner.test.ts test/routes-settings.test.ts && npx tsc --noEmit -p .`
— 101 tests passed with a clean type check. The integration test covers create B, list, detail,
response, and resume through B's pinned descriptor while default A remains configured.
- Full backend suite: 342 tests passed. The remaining 7 tests require binding `127.0.0.1` and
fail in this sandbox with `listen EPERM: operation not permitted`; no application assertion
failed. The focused typecheck above passed.
- `git diff --check` completed cleanly.
## Verification note
The unscoped backend suite was also run. The Task 7 code regressions in `test/tht-runner.test.ts`
were fixed; the remaining failures were existing sandbox restrictions on tests that listen on
`127.0.0.1` (`listen EPERM: operation not permitted` in SSE/e2e health tests), not application
assertions.
## Review fixes — round 1
- Every new session now resolves `workspaceId` through the registry; an omitted value uses the
configured installation default and persists both the resolved ID and revision. Callers cannot
bypass revision pinning by supplying a workspace ID.
- Browser-local preferences now migrate once from the read-only legacy settings response and hold
workspace, provider, model, and thinking. Session creation includes those selections, including
direct entry points that run before the composer mounts. The frontend no longer `PUT`s shared
settings.
- The settings compatibility endpoint honors a stored installation workspace before falling back
to the first workspace configuration.
- Resume rejects finalized and archived sessions before looking up any pinned snapshot, preserving
the read-only response even when a historical snapshot is unavailable.
### Review verification
- RED: the added backend tests failed for omitted-default pinning, read-only resume ordering, and
stored-default precedence; the added frontend preference tests failed because preferences were
neither stored nor included in session requests.
- GREEN: `npx vitest run test/tht-runner.test.ts test/routes-sessions.test.ts test/routes-settings.test.ts && npx tsc --noEmit -p .`
— 100 tests passed with a clean type check.
- GREEN: `npx vitest run && npx tsc -b` — 332 frontend tests passed with a clean type check.
- GREEN: `THT_HOME=/private/tmp/thothii-task7-home .venv/bin/pytest tests/test_session_documents.py tests/test_session_mutations.py -q`
— 22 tests passed (one existing testcontainers deprecation warning).
- `git diff --check` completed cleanly.
@@ -1,73 +0,0 @@
# Task 9 report — Workspace Management CRUD page
## Delivered
- Added the Workspace management dialog, launched from the persistent right sidebar and the
Model activity header without touching live-session/SSE state.
- Added a workspace list/detail editor for General, DWH, Semantic index, LLM policy,
Installation requirements, and Git status/history.
- Added browser-only New, Edit, Duplicate, Save draft, and Delete-draft workflows. A deletion
draft stores only ID and immutable revision references; publication remains a Task 10 action.
- Used closed native controls for languages, engines, transports, distance metrics, embedding
providers, and selectable default models. Free values have client-side, accessible errors.
- Made semantic-index dimensions atomic: one editor field always writes the same value to the
vector-store and embedding contracts.
- Added Validate and Test-on-this-installation actions. They display sanitized code/message
diagnostics only; neither action exposes or stores credentials, secrets, or raw response bodies.
- Explicitly excluded publish, pull, import, and export user flows from this task.
## TDD evidence
- RED: `npx vitest run src/shell/WorkspaceManager.test.tsx src/shell/WorkspaceEditor.test.tsx`
failed because the manager and editor modules did not exist.
- GREEN: focused manager/editor/AppShell coverage passed after the implementation.
- RED: a deletion-draft persistence regression failed with
`Cannot read properties of undefined (reading 'save')` before the sanitized draft store was added.
- GREEN: the draft-store and manager tests passed once deletion intent persisted locally.
## Verification
Executed from `frontend/`:
```text
npx vitest run
50 test files passed, 358 tests passed
npx tsc -b
exit 0
```
`git diff --check` passed before commit. No workspace secret value, secret-file path, raw
diagnostic body, publish call, import flow, or export flow was introduced.
## Fix round 1
### Root causes and fixes
- The original duplicate proposal appended `-copy` and then truncated at 63 characters. For an
already-maximal ID, truncation could remove the suffix and reproduce the immutable source ID.
The proposal now reserves suffix space and falls back to a distinct `-2` suffix when a maximal
source already ends in `-copy`.
- `dwh.timeout_ms` was rendered as a positive numeric field but was absent from the client
validation map. It now has the same immediate accessible error treatment as other numeric
fields, so a rejected save never reaches the manager’s saved-draft toast.
- Registry status, workspace list, and selected-detail React Query failures were rendered as
loading, empty, or unselected states. Each now has a named alert and a retry control, distinct
from its corresponding loading and empty state.
### TDD evidence
- RED: max-length duplication retained the original 63-character ID; the timeout field produced
no alert; and each of the three failed queries had no accessible retry control.
- GREEN: the focused manager/editor tests passed **12/12**, covering a valid changed duplicate
proposal, rejected zero timeout with no save toast, and status/list/detail retry recovery.
### Verification
Executed from `frontend/`:
```text
npx vitest run
50 test files passed, 364 tests passed
npx tsc -b
exit 0
```
@@ -1,180 +0,0 @@
# Task 11 report
Status: completed on 2026-08-08.
## Scope delivered
- Updated operator-facing documentation for the internal Qdrant + Ollama architecture.
- Tightened documentation contract tests to require the current four-service-plus-init topology,
CPU-first/GPU-override guidance, fixed internal model/dimensions, schema-v3 migration wording,
one-collection-per-workspace ownership, and Qdrant backup/restore safety.
- Updated stable repo guidance in `AGENTS.md` and the current snapshot in `PROJECT_STATE.md`.
- Rewrote the workspace diagnostic protocol to the schema-v3/internal-semantic-service contract.
- Updated the memory guide to describe Qdrant as the derived persistent index.
- Updated the runtime secret-bundle guide to remove active vector/embedding secret guidance.
## Files changed
- `README.md`
- `AGENTS.md`
- `PROJECT_STATE.md`
- `docs/install/local-workspace-registry.md`
- `docs/install/server-workspace-registry.md`
- `docs/installazione-docker-4-contesti.md`
- `docs/workspace-diagnostic-protocol.md`
- `docs/gestione-memory.md`
- `deploy/secrets/README.md`
- `scripts/verify-workspace-install-docs.sh`
- `scripts/test-verify-workspace-install-docs.sh`
## Verification
Fresh successful runs:
```sh
./scripts/test-verify-workspace-install-docs.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
git diff --check
```
Key outcomes:
- internal semantic infrastructure documentation contract passed
- all existing install/manual fixture contracts still passed
- diff hygiene passed with no whitespace/errors
## Self-review notes
- The updated docs now match the code-backed Compose topology: `frontend`, `core`, `qdrant`,
`embedding`, and `embedding-model-init`.
- Active manuals no longer instruct operators to configure external vector or embedding runtime
endpoints/secrets.
- Qdrant backup/restore wording now matches the helper scripts' exact confirmation and rollback
behavior.
- Legacy descriptor handling is documented as explicit schema-v3 migration only; no silent
semantic-data migration is claimed.
## Residual concerns
- The broader repository still contains historical design/spec material that references older
pgvector/external-embedding architecture; this task intentionally updated operator/current-state
documentation and the corresponding contract tests, not historical planning documents.
## Fix round 1/5 — 2026-08-08
Addressed reviewer findings:
- Moved superseded rollout/state blocks in `PROJECT_STATE.md` behind an explicit
`## Historical snapshots and archived reference notes` boundary.
- Renamed superseded snapshot headings so historical notes no longer present as active `LIVE`
state.
- Added a current-state regression that rejects contradictory active blocks (for example:
schema-v2 operational, two-service active stack, or external vector/embedding runtime claims
before the historical boundary).
- Refactored new internal-semantic doc checks away from exact-sentence coupling:
- parse `compose.yaml` structurally with YAML;
- parse workspace examples structurally with YAML;
- inspect backup/restore stable usage interface;
- keep targeted forbidden-term checks for active docs while allowing historical sections;
- use regex/concept checks for prose.
Evidence:
```sh
./scripts/test-verify-workspace-install-docs.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
git diff --check
```
Observed RED before the fix:
```text
PROJECT_STATE.md: missing Historical snapshots boundary
```
## Fix round 2/5 — 2026-08-08
Addressed reviewer findings:
- Renamed every historical `PROJECT_STATE.md` heading after the historical boundary so no heading
level uses `LIVE` or current-state semantics there.
- Strengthened the historical-boundary regression to reject any Markdown heading level
(`#` through `######`) containing `LIVE` or current-state wording after the boundary.
- Added a fixture with a `### ... — LIVE ...` historical heading to prove RED then GREEN.
- Replaced remaining exact phrase checks with concept/semantic validation for:
- one-workspace/one-collection ownership;
- external boundary (DWH/LLM external; vector/embedding internal);
- the Italian compact install note.
- Added paraphrase fixtures that pass and omission/inversion fixtures that fail.
Evidence:
```sh
./scripts/test-verify-workspace-install-docs.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
git diff --check
```
## Fix round 4/5 — 2026-08-08
Addressed reviewer finding:
- Eliminated semantic-index verifier/test contract drift by extracting the production
semantic-index ownership row matcher into `semantic_index_relationship_spec` and reusing it in
the fixture-level paraphrase, omission, and scattered-token checks.
- Kept the relationship constrained to one structured Markdown table row via
`verify_markdown_table_relationships`; the scattered-token fixture still removes the row and
appends the same words outside the table, where it must be rejected.
- Added a direct regression that copies the repository docs into an isolated root, applies the
accepted paraphrase “A workspace keeps exactly one Qdrant collection reserved for itself”, and
runs that root's actual `scripts/verify-workspace-install-docs.sh --fixtures-only` instead of a
separate temporary spec.
Observed RED before the fix:
```text
production verifier rejected the accepted semantic-index paraphrase
local workspace manual: missing relationship in 'Semantic index ownership contract': {'scope': 'workspace semantic index', 'ownership rule': '(each|one|single).*(workspace).*(single|one).*(Qdrant).*(collection)|(each workspace reserves a single qdrant collection)', 'isolation rule': 'schema.*evidence.*memory.*(one|that).*(collection).*(kind|payload)'}
```
Evidence:
```sh
./scripts/test-verify-workspace-install-docs.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
git diff --check
```
Observed RED during this round:
```text
PROJECT_STATE.md: historical section still contains active/live heading markers
compact manual paraphrase lacks required pattern: (esterni solo|solo esterni|restano esterni)
```
## Fix round 3/5 — 2026-08-08
Addressed reviewer findings:
- Added table-driven historical-heading fixtures for every Markdown heading level `#` through
`######`; all are rejected after the historical boundary when they contain `LIVE`/current-state
semantics.
- Added small structured ownership tables to the active local/server manuals and to the compact
Italian operator note.
- Added small structured semantic-index ownership tables to the active local/server manuals.
- Replaced the remaining scattered-token relationship checks with explicit structured-section
parsing:
- architecture ownership rows map DWH → external, LLM → external, Qdrant → internal,
Ollama embedding → internal;
- semantic-index ownership rows localize the one-workspace/one-collection contract and the
schema/Evidence/Memory isolation rule.
- Added adversarial fixtures that fail when the same tokens are merely scattered in free text.
- Added structured paraphrase fixtures that pass and omission/inversion fixtures that fail.
Evidence:
```sh
./scripts/test-verify-workspace-install-docs.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
git diff --check
```
@@ -1,43 +0,0 @@
# Task 12 Report — Remove unreachable pgvector runtime code
Status: completed
Summary:
- Proved the retired pgvector runtime had no remaining operational adapter call sites after migration by re-running the required grep; only the packaging assertion still mentions `migrations/vector`.
- Removed the obsolete pgvector/HTTP/direct vector runtime modules, vector SQL migrations, and their affected runtime tests.
- Kept the operational semantic path on Qdrant and migrated the remaining runtime callers to that path.
- Kept `psycopg2-binary` because DWH direct PostgreSQL and session PostgreSQL code still depend on it.
Implementation notes:
- Extracted shared collection/kind validation into `harness/tht/adapters/vector/_shared.py` so `QdrantVectorStore` no longer depends on the deleted pgvector module.
- Simplified `build_vector_store()` to return only `QdrantVectorStore`.
- Migrated vector/evidence/memory CLI paths away from legacy pgvector loaders and REST vector clients.
- Updated packaging coverage so the built wheel asserts session SQL migrations are present and vector SQL migrations are absent.
Verification:
- `cd harness && .venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py tests/test_semantic_kind_isolation.py tests/test_vector_migration_packaging.py -q`
- `cd harness && .venv/bin/pytest tests/test_adapter_factory.py tests/test_solved_search_cli.py -q`
- `cd harness && .venv/bin/python -c "import tht.cli, tht.adapters.factory, tht.adapters.vector, tht.vectorstore.reader"`
- `cd harness && uv build`
- `harness/.venv/bin/ruff check harness/tests/test_adapter_factory.py harness/tests/test_solved_search_cli.py harness/tests/test_vector_migration_packaging.py harness/tests/test_vector_port_contract.py harness/tht/adapters/factory.py harness/tht/adapters/vector/__init__.py harness/tht/adapters/vector/_shared.py harness/tht/adapters/vector/qdrant.py harness/tht/cli/evidence_cmd.py harness/tht/cli/memory_cmd.py harness/tht/cli/search_cmd.py harness/tht/cli/vector_cmd.py harness/tht/solved.py harness/tht/vectorstore/reader.py`
- `git diff --check`
Notes / concerns:
- Repository-wide `harness/.venv/bin/ruff check .` still reports many pre-existing findings outside this task’s touched files; it is not clean on this branch baseline.
- Some legacy config compatibility parsing still exists outside the deleted runtime path. This task removed the unreachable runtime/migration code without broad config-schema refactoring.
## Fix round 1 evidence
Changes:
- Removed dead `vector migrate` registration from `harness/tht/cli/__init__.py` and deleted `harness/tht/cli/vector_migrate_cmd.py`.
- Added CLI regressions proving `vector migrate` is absent while `vector init` and `vector index-schema` remain available.
- Restored the accidentally removed non-vector regressions by moving report coverage into `harness/tests/test_report.py` and restoring the taskdoc promoted-table slicing check in `harness/tests/test_taskdoc.py`.
- Reworded surviving active help/docstrings away from pgvector-specific wording in the touched Qdrant-backed command surface.
Verification:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py tests/test_report.py tests/test_taskdoc.py tests/test_vector_migration_packaging.py -q`
- `cd harness && .venv/bin/python -c "from typer.testing import CliRunner; from tht.cli import app; r=CliRunner().invoke(app, ['vector','--help']); assert r.exit_code == 0, r.output; assert 'migrate' not in r.output; r=CliRunner().invoke(app, ['vector','migrate','--help']); assert r.exit_code != 0, r.output; print('cli-help-ok')"`
- `cd harness && .venv/bin/python -c "import tht.cli, tht.cli.vector_cmd, tht.report, tht.taskdoc; print('imports-ok')"`
- `cd harness && uv build`
- `harness/.venv/bin/ruff check harness/tests/test_qdrant_cli_commands.py harness/tests/test_report.py harness/tests/test_taskdoc.py harness/tests/test_vector_migration_packaging.py harness/tht/cli/__init__.py harness/tht/cli/search_cmd.py harness/tht/cli/vector_cmd.py harness/tht/cli/memory_cmd.py harness/tht/solved.py`
- `git diff --check`
@@ -1,175 +0,0 @@
# Task 13 Implementation Report
## Status
DONE_WITH_CONCERNS
## Changes
- Updated stale harness/backend/frontend tests and fixtures to the Task 13 internal Qdrant/Ollama contract.
- Made `deploy/workspaces/psd.yaml.example` generic while preserving schema-v3 Qdrant/Ollama shape.
- Fixed `scripts/workspace-registry-smoke.sh` to pass the required legacy migration `--collection` and prove exact Docker cleanup, including its smoke image.
- Updated `PROJECT_STATE.md` with only evidence observed in this run.
Changed files:
- `PROJECT_STATE.md`
- `backend/test/routes-workspaces.test.ts`
- `backend/test/workspace-runtime-handoff.test.ts`
- `backend/test/workspaces-contracts.test.ts`
- `backend/test/workspaces-git-repository.test.ts`
- `deploy/workspaces/psd.yaml.example`
- `frontend/src/shell/NewSessionDialog.test.tsx`
- `harness/tests/test_adapter_command_regressions.py`
- `harness/tests/test_workspace.py`
- `scripts/task13-runtime-fixture-check.ts`
- `scripts/test-verify-workspace-install-docs.sh`
- `scripts/workspace-registry-smoke.sh`
## Verification
Deterministic gates:
- `cd harness && .venv/bin/pytest -q && .venv/bin/ruff check .`
- Initial red: 2 harness pytest failures.
- After fixture fixes: harness pytest passed `819 passed, 4 deselected, 74 warnings in 27.73s`.
- Ruff still failed with `Found 220 errors`; treated as existing unrelated debt.
- Touched harness files verified clean with `cd harness && .venv/bin/ruff check tests/test_adapter_command_regressions.py tests/test_workspace.py && .venv/bin/pytest -q tests/test_adapter_command_regressions.py::test_solved_index_writes_through_writer_only_factory_store tests/test_workspace.py::test_load_workspace_expands_env_vars`: `All checks passed!` and `2 passed, 2 warnings in 0.14s`.
- `cd backend && npx vitest run && npx tsc --noEmit -p . && npm run build`
- Initial red: 4 backend Vitest failures.
- After fixes: `Test Files 39 passed (39)`, `Tests 464 passed (464)`, TypeScript passed, build passed.
- `cd frontend && npx vitest run && npx tsc -b && npm run build`
- Initial red: 1 frontend Vitest failure.
- After fix: frontend Vitest passed `374/374`, TypeScript passed, build passed with Vite `built in 6.55s`.
- `git diff --check`
- Passed with no output.
Focused reruns:
- `cd backend && npx vitest run test/workspaces-migrate-legacy.test.ts test/workspaces-contracts.test.ts test/routes-workspaces.test.ts test/workspace-runtime-handoff.test.ts test/workspaces-git-repository.test.ts && cd .. && ./scripts/test-no-deployment-coupling.sh && ./scripts/verify-workspace-install-docs.sh --fixtures-only && git diff --check`
- `Test Files 5 passed (5)`, `Tests 35 passed (35)`.
- Coupling guard passed: `no active retired deployment or external semantic coupling found.`
- Install docs fixtures passed through `relative secret-source fixture rejected passed`.
Deployment contracts:
- `./scripts/test-default-compose.sh && ./scripts/test-unified-compose.sh && ./scripts/test-internal-semantic-compose.sh && ./scripts/test-no-deployment-coupling.sh && ./scripts/test-compose-secret-policy.sh && ./scripts/verify-workspace-install-docs.sh --fixtures-only`
- Passed. Output included:
- `default Compose contract passed.`
- `unified Compose contract passed.`
- `internal semantic Compose/script contracts passed.`
- `no active retired deployment or external semantic coupling found.`
- `Compose secret policy passed.`
- install-doc fixture checks through `relative secret-source fixture rejected passed`.
Docker smokes:
- `/usr/bin/time -p ./scripts/internal-semantic-smoke.sh`
- Passed: `Task 13 internal semantic smoke passed.`
- Cleanup proof: `no labeled containers, volumes, networks, or images remain for 20260808200245-83368-17823.`
- Duration: `real 217.34`.
- `/usr/bin/time -p ./scripts/workspace-registry-smoke.sh`
- Initial red: `usage: migrate-legacy --input <legacy-workspace.yaml> --output <repository-root> --collection <qdrant-collection> [--id <workspace-id>]`.
- After fix: `workspace registry smoke passed`.
- Cleanup proof: `no compose containers, volumes, networks, or image remain for thoth-workspace-registry-smoke-89671.`
- Duration: `real 9.93`.
- `/usr/bin/time -p ./scripts/unified-deployment-smoke.sh`
- Passed: `Task 13 full deployment smoke passed.`
- Cleanup proof: `no labeled containers, volumes, networks, or images remain for 20260808200706-85638-13391.`
- Duration: `real 125.57`.
- `/usr/bin/time -p ./scripts/thothctl-update-smoke.sh`
- Passed: `Task 13 update deployment smoke passed.`
- Cleanup proof: `no labeled containers, volumes, networks, or images remain for 20260808200918-87340-10404.`
- Duration: `real 85.40`.
- `/usr/bin/time -p ./scripts/server-deployment-smoke.sh`
- Passed: `Task 13 Linux server deployment smoke passed.`
- Cleanup proof: `no labeled containers, volumes, networks, or images remain for 20260808201047-88645-20675.`
- Duration: `real 55.99`.
Final audit:
- `rg -n "pgvector|local-vector|THT_VECTOR_|EMBEDDING_BASE_URL|openai_compatible|ollama_compatible" . --glob '!docs/plans/**' --glob '!docs/superpowers/**' --glob '!**/node_modules/**' --glob '!**/.venv/**' --glob '!**/.git/**'`
- Returned matches in legacy schema-v1/v2 support, migration tests, negative guards, historical notes, and older harness docs/code.
- This remains a concern: the audit is not clean under the brief's strict expected outcome.
- `git status --short`
- Before report/commit, contained only intentional Task 13 changes.
## Image and Host Evidence
- Host CPU: `Apple M4 Pro`.
- Host OS: `Darwin MacProM4-di-Marco.local 25.5.0 Darwin Kernel Version 25.5.0: Tue Jun 9 22:28:34 PDT 2026; root:xnu-12377.121.10~1/RELEASE_ARM64_T6041 arm64`.
- Docker server: `29.6.2 linux/arm64`.
- Verified pinned images:
- `qdrant/qdrant:v1.18.2@sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c`.
- `ollama/ollama:0.32.0@sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a`.
- Workspace registry smoke ephemeral image:
- Manifest list: `sha256:4d056bf2cb38d0e8ede91fbf121df1f9f18caee0d401581618ccef9ed8a55e73`.
- Config: `sha256:613f8fb28c0517adee4085f41bc447f2c3813b0fdbb7b26624bfb4cb192b6fd8`.
- Removed during cleanup.
## Manual Gates
- GPU exposure gate (`THOTH_ENABLE_EMBEDDING_GPU=1` on Linux): not executed in this run.
- Windows Docker Desktop startup/manual job: not executed in this run.
## Commits
- `4e810af` (`test: align qdrant ollama verification fixtures`)
- `7c09b98` (`docs: record qdrant ollama verification`)
## Known Limitations
- Broad harness Ruff remains existing unrelated debt: `Found 220 errors`.
- Final active-reference audit is not clean; it still finds legacy/negative-guard references outside explicit migration fixture files.
- Ephemeral Task 13 core/frontend image IDs from `internal-semantic-smoke.sh`, `unified-deployment-smoke.sh`, `thothctl-update-smoke.sh`, and `server-deployment-smoke.sh` were removed by exact cleanup and were not emitted in stdout; pinned Qdrant/Ollama digests and the workspace-registry smoke image digest were captured.
## Fix Round 1 — reviewer findings
Status: DONE
Changes:
- `scripts/workspace-registry-smoke.sh` now derives the smoke image reference from the already unique Compose project instead of using the global tag `thothii-workspace-registry-smoke:local`.
- The workspace-registry cleanup helpers remove and verify only the exact per-run image reference, plus Compose resources labeled with the exact project.
- Added deterministic self-test coverage in `backend/test/workspaces-migrate-legacy.test.ts` via `WORKSPACE_REGISTRY_SMOKE_SELF_TEST=image-cleanup-identity`; it stubs Docker and fails if cleanup touches same-repository foreign tags such as `:local` or another project tag.
- Updated active harness/testing/PRD docs and Python comments that still described the current semantic store as pgvector/vectordb. Preserved schema-v1/v2 and harness legacy compatibility fixtures.
- Updated `PROJECT_STATE.md` with fix-round smoke evidence and a precise, non-overclaiming audit limitation.
Focused verification:
- `cd backend && npx vitest run test/workspaces-migrate-legacy.test.ts`
- Passed: `7 passed`.
- `cd harness && .venv/bin/pytest -q tests/test_memory_save_one.py tests/test_adapter_command_regressions.py tests/test_solved_search_cli.py tests/test_search_pack.py`
- Passed: `22 passed, 14 warnings`.
- `cd harness && .venv/bin/ruff check tht/memory.py tht/search/__init__.py tht/workspace.py tht/vectorstore/store.py tests/test_memory_save_one.py tests/test_adapter_command_regressions.py tests/test_solved_search_cli.py`
- Passed: `All checks passed!`
- `bash -n scripts/workspace-registry-smoke.sh && WORKSPACE_REGISTRY_SMOKE_SELF_TEST=image-cleanup-identity bash scripts/workspace-registry-smoke.sh`
- Passed: `workspace registry smoke image cleanup identity self-test passed`.
- `./scripts/test-no-deployment-coupling.sh`
- Passed: `no active retired deployment or external semantic coupling found.`
- `./scripts/verify-workspace-install-docs.sh --fixtures-only`
- Passed through `relative secret-source fixture rejected passed`.
- `cd backend && npx tsc --noEmit -p .`
- Passed with no output.
- `/usr/bin/time -p ./scripts/workspace-registry-smoke.sh`
- Passed: `workspace registry smoke passed`.
- Built exact per-run tag: `thothii-workspace-registry-smoke:thoth-workspace-registry-smoke-thoth-workspace-registry-smoke-10vi3a-19157`.
- Manifest list: `sha256:715b943057929418cad4aa71806d9edbaf823555d19bda6b875297617463fd4a`.
- Config: `sha256:a566521981e08958aae9a12bfc7803bb5f3f835536b4bb8c39df8fcf26063161`.
- Cleanup proof: `no compose containers, volumes, networks, or image remain for thoth-workspace-registry-smoke-thoth-workspace-registry-smoke-10vi3a-19157.`
- Duration: `real 42.06`.
Fix-round audit command:
- `rg -n "pgvector|local-vector|THT_VECTOR_|EMBEDDING_BASE_URL|openai_compatible|ollama_compatible" . --glob '!docs/plans/**' --glob '!docs/superpowers/**' --glob '!**/node_modules/**' --glob '!**/.venv/**' --glob '!**/.git/**'`
Categorized remaining hits:
- Backend legacy parser/migration compatibility, kept deliberately non-operational for schema-v1/v2 descriptors: `backend/src/workspaces/schema.ts`, `types.ts`, `migrate-legacy.ts`, `runtime-renderer.ts`, `bindings.ts`, `contracts.ts`, `diagnostics.ts`.
- Backend negative guards and legacy fixture tests: `backend/test/workspaces-schema.test.ts`, `workspaces-migrate-v2-qdrant.test.ts`, `workspace-registry.test.ts`, `workspace-runtime-renderer.test.ts`, `workspaces-bindings.test.ts`, `workspaces-contracts.test.ts`, `workspaces-diagnostics.test.ts`, `workspaces-git-repository.test.ts`, `routes-workspaces.test.ts`, `routes-sessions.test.ts`, `provider-credentials.test.ts`.
- Secret/env scrub guards for retired variables: `backend/src/config.ts`, `backend/src/config/secret-bundle.ts`, `backend/src/pi/provider-credentials.ts`, `scripts/compose-with-preflight.sh`, `scripts/test-external-compose-lifecycle.sh`.
- Deployment negative guards and fixture-scope tests: `scripts/test-no-deployment-coupling.sh`, `scripts/test-no-deployment-coupling-scope.sh`, `scripts/test-preprocess-compose-config.sh`, `scripts/test-verify-workspace-install-docs.sh`, `scripts/verify-workspace-install-docs.sh`, `scripts/vector-rotate-bootstrap-password.sh`.
- Harness legacy config compatibility and fixtures: `harness/tht/config.py`, `harness/tht/config_compat.py`, `harness/tests/test_config_resources.py`, `harness/tests/l2/test_session_ablazione.py`, `harness/workspaces/tht.example.yaml`, `harness/workspaces/tht-test.yaml`.
- Retained off-repository migration SQL fixtures: `harness/scripts/create_vector_reader_rpc.sql`, `harness/scripts/create_vector_writer_rpc.sql`.
- Historical/reference notes, not active operator contracts: `brain/codebase/datamart-builder-deployment-gotchas.md`, `PROJECT_STATE.md`.
- Gitignored task report self-reference: `.superpowers/sdd/2026-08-08-internal-qdrant-ollama/task-13-implementation.md`.
@@ -1,165 +0,0 @@
Task 2 report — Make collection ownership unique in the Git registry
Summary
- Implemented unique Qdrant collection ownership enforcement during registry snapshot activation.
- Registry session revision leases now reject `migration_required` descriptors.
- Legacy migration now requires an explicit target collection and emits schema v3 descriptors.
- Preserved active snapshot rollback behavior on invalid pulled snapshots.
RED evidence
Focused RED command from the brief:
```bash
cd backend
npx vitest run test/workspace-registry.test.ts test/workspaces-migrate-legacy.test.ts \
-t "collection|migration_required"
```
Observed failures before implementation:
- `rejects duplicate schema v3 collection ownership and keeps the previous active snapshot`
- `registry.pull()` resolved instead of rejecting.
- `does not acquire a session revision lease for a migration_required workspace`
- `acquireSessionRevision()` resolved instead of rejecting.
- `migrates a legacy descriptor only with an explicit target collection into schema v3`
- received schema version `1` instead of `3`.
- `requires an explicit target collection for legacy migration`
- migration did not throw without a collection.
GREEN evidence
Focused GREEN command from the brief:
```bash
cd backend
npx vitest run test/workspace-registry.test.ts test/workspaces-migrate-legacy.test.ts \
-t "collection|migration_required"
```
Fresh result after implementation:
- 2 files passed
- 4 tests passed
- 0 failures
Additional verification run after final cleanup:
```bash
cd backend
npx vitest run test/routes-workspaces.test.ts
npx vitest run
npx tsc --noEmit -p .
git diff --check
```
Fresh results:
- `test/routes-workspaces.test.ts`: 7 passed
- full backend Vitest: 39 files passed, 454 tests passed
- backend typecheck: passed
- `git diff --check`: passed
Changed files
- `backend/src/workspaces/registry.ts`
- `backend/src/workspaces/migrate-legacy.ts`
- `backend/test/workspace-registry.test.ts`
- `backend/test/workspaces-migrate-legacy.test.ts`
- `backend/test/routes-workspaces.test.ts`
Why one extra file changed
- `backend/test/routes-workspaces.test.ts` needed updating because Task 1 made schema v3 the only operational descriptor shape, and the route test still assumed the old pre-Task-3 runtime behavior. Updating that expectation was necessary to keep the required backend suite verification meaningful.
Implementation notes
- Duplicate collection detection is enforced only for operational schema v3 descriptors by tracking `collection -> workspaceId` during activation.
- Duplicate failures are sanitized back to `workspace_invalid` / `Workspace repository content is invalid`.
- `acquireSessionRevision()` now fails closed for `migration_required` revisions.
- Legacy migration CLI now requires `--collection <qdrant-collection>`.
- Legacy migration output is schema v3 with the fixed internal semantic contract:
- `vector_store.engine = qdrant`
- explicit `collection`
- embedding provider `ollama_internal`
- embedding model `qwen3-embedding:0.6b`
self-review
- Confirmed invalid pulled snapshots do not replace the previous active snapshot.
- Confirmed duplicate collection enforcement does not affect legacy migration-required descriptors.
- Confirmed create/update publication tests still pass with unique per-workspace collections.
- Confirmed no JSON stdout contract regressions in the migration CLI.
- Kept runtime/data mutation scope descriptor-only; no user workspace repo or Qdrant data changes.
Concerns
- No code concerns remaining for Task 2.
- One deliberate scope exception: a route test was updated to align with the already-established Task 1 / Task 3 fail-closed contract.
Fix round 1
Scope
- Restored meaningful route-level diagnoser coverage without reopening schema-v3 semantic runtime paths.
- Added direct schema-v2 registry coverage for `migration_required` listing and lease rejection.
Covering test files
- `backend/test/routes-workspaces.test.ts`
- `backend/test/workspace-registry.test.ts`
RED command and output
Command:
```bash
cd backend
npx vitest run test/routes-workspaces.test.ts test/workspace-registry.test.ts
```
Observed result on top of `76bc94d` after adding the restored/new assertions:
- 2 files passed
- 37 tests passed
- 0 failures
Why no RED appeared:
- The review items exposed missing/weakened coverage, not a production behavior bug.
- `/workspaces/:id/test` already reaches the diagnoser for resolvable legacy v2 descriptors.
- Schema-v3 `/workspaces/:id/test` already fails closed before diagnoser entry.
- Schema-v2 descriptors were already listed as `migration_required` and already rejected by `acquireSessionRevision()`.
GREEN command and output
Command:
```bash
cd backend
npx vitest run test/routes-workspaces.test.ts test/workspace-registry.test.ts
npx tsc --noEmit -p .
```
Fresh results:
- covering tests: 2 files passed, 37 tests passed
- backend typecheck: passed
Changed files
- `backend/test/routes-workspaces.test.ts`
- `backend/test/workspace-registry.test.ts`
- `.superpowers/sdd/2026-08-08-internal-qdrant-ollama/task-2-report.md`
What changed
- Split route coverage so `POST /workspaces/validate` still checks canonical validation independently.
- Restored route-level diagnoser coverage through a migration-required schema-v2 descriptor with resolvable legacy bindings.
- Added an explicit schema-v3 fail-closed regression for `POST /workspaces/:id/test`.
- Added a direct schema-v2 registry regression proving `list()` returns `migration_required` and `acquireSessionRevision()` rejects it.
Concerns
- No production concerns. This round only tightened coverage and corrected the weakened test expectation.
@@ -1,132 +0,0 @@
# Task 3 report — Remove external semantic bindings and render internal endpoints
Date: 2026-08-08
## Scope
Implemented backend-owned schema-v3 semantic runtime rendering so workspace descriptors and installation contracts remain free of external Qdrant/Ollama endpoints and credentials, while DWH bindings stay unchanged.
## RED evidence
Focused RED command:
`cd backend && npx vitest run test/workspaces-contracts.test.ts test/workspaces-bindings.test.ts test/workspace-runtime-renderer.test.ts test/config.test.ts`
Observed failures before implementation:
- `config.test.ts`
- missing `internalQdrantUrl`
- missing `internalEmbeddingUrl`
- `workspaces-bindings.test.ts`
- schema v3 semantic binding resolution threw unsupported errors
- `workspace-runtime-renderer.test.ts`
- schema v3 runtime rendering threw `Schema version 3 runtime rendering is unsupported until the internal semantic runtime is implemented`
## GREEN evidence
Focused GREEN command:
`cd backend && npx vitest run test/workspaces-contracts.test.ts test/workspaces-bindings.test.ts test/workspace-runtime-renderer.test.ts test/config.test.ts`
Result:
- 4 test files passed
- 36 tests passed
Typecheck:
`cd backend && npx tsc --noEmit -p .`
Result:
- passed
Hygiene:
- `git diff --check` passed
## Files changed
Listed-task files changed:
- `backend/src/config.ts`
- `backend/src/workspaces/bindings.ts`
- `backend/src/workspaces/runtime-renderer.ts`
- `backend/test/config.test.ts`
- `backend/test/workspace-runtime-renderer.test.ts`
- `backend/test/workspaces-bindings.test.ts`
- `backend/test/workspaces-contracts.test.ts`
Listed-task files inspected but not changed:
- `backend/src/workspaces/contracts.ts`
Unavoidable additional wiring changes:
- `backend/src/app.ts`
- `backend/src/tht/tht-runner.ts`
Reason: the new typed internal semantic runtime config had to flow from backend config into ephemeral harness config rendering at runtime.
## Behavior delivered
- schema-v3 installation contract exposes DWH bindings only
- schema-v3 binding resolution ignores external semantic env vars instead of sourcing runtime semantics from them
- runtime rendering for schema v3 emits backend-owned internal semantic endpoints:
- Qdrant: `http://qdrant:6333`
- Embedding: `http://embedding:11434`
- Model: `qwen3-embedding:0.6b`
- Dimensions: `1024`
- internal semantic URLs are validated to allow only `qdrant` / `embedding` / `localhost` / loopback hosts
- DWH transport/runtime behavior remains unchanged
## Self-review
- Confirmed schema-v3 contracts/docs no longer advertise VECTOR or EMBEDDING installation variables.
- Confirmed schema-v3 runtime output ignores injected external semantic endpoints from env bindings.
- Confirmed semantic endpoints are rendered only in the ephemeral backend-owned harness config path.
- Confirmed type wiring is explicit from `AppConfig` → `ThtRunner` → runtime renderer.
## Concerns
- Host validation currently permits both `http` and `https` on the allowed internal hosts. That keeps the configuration flexible, but if the installation contract intended `http` only, that restriction is not enforced here.
## Fix round 1/5
Scope:
- moved schema-v3 internal embeddings under `resources.embeddings`
- enforced `http`-only internal semantic URLs
RED evidence:
`cd backend && npx vitest run test/workspace-runtime-renderer.test.ts test/config.test.ts`
Observed failures on `bc8afe0`:
- `workspace-runtime-renderer.test.ts`
- schema-v3 output omitted `resources.embeddings`
- schema-v3 still exposed top-level `embeddings`
- `config.test.ts`
- `https://qdrant:6333` was accepted
GREEN evidence:
`cd backend && npx vitest run test/workspace-runtime-renderer.test.ts test/config.test.ts`
Result:
- 2 test files passed
- 16 tests passed
Typecheck:
`cd backend && npx tsc --noEmit -p .`
Result:
- passed
Updated concerns:
- none for this round beyond future tightening if exact-port rejection is later requested explicitly.
@@ -1,149 +0,0 @@
# Task 4 Report — Narrow harness embedding configuration to internal Ollama
## Status
Implemented on 2026-08-08 in `/Users/mp/projects/ThothII/.worktrees/git-workspace-registry`.
## RED evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py -q
```
Observed before implementation:
- exit code `1`
- `10 failed, 10 passed`
- failures proved the missing `OllamaInternalEmbeddings` client and missing internal-only config validation
Representative failures:
- `ImportError: cannot import name 'OllamaInternalEmbeddings'`
- `AttributeError: 'EmbeddingsConfig' object has no attribute 'provider'`
- config tests `DID NOT RAISE ConfigError` for external provider, API key, and non-private base URL
## GREEN evidence
Focused behavior suite:
```bash
cd harness
./.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py -q
```
- exit code `0`
- `20 passed`
Relevant harness verification:
```bash
cd harness
./.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py tests/test_ollama_ensure.py -q
```
- exit code `0`
- `36 passed, 2 warnings`
Changed-file lint:
```bash
cd harness
./.venv/bin/ruff check tht/config.py tht/config_compat.py tht/vectorstore/embeddings.py tht/cli/ollama_cmd.py tests/test_config_resources.py tests/test_internal_embeddings.py
```
- exit code `0`
- `All checks passed!`
Patch hygiene:
```bash
git diff --check
```
- exit code `0`
## What changed
- translated schema-v3 `resources.embeddings` into the harness-compatible embedding config view
- validated the internal embedding contract only for that runtime-owned `resources.embeddings` path:
- provider must be `ollama_internal`
- model must be `qwen3-embedding:0.6b`
- dimensions must be `1024`
- base URL must be `http://embedding:11434` or loopback HTTP on port `11434`
- extra fields like `api_key` are rejected
- replaced the active embed client with `OllamaInternalEmbeddings`, using one bounded `/api/embed` request per batch
- removed task/query prefix rewriting from the active embedding path
- validated response count, vector dimension, and finite numeric values before returning embeddings
- kept `tht ollama ensure --json` stdout pristine while warming through the internal client
## Self-review
- kept changes inside the brief-listed files
- preserved DWH and session-persistence behavior
- preserved the legacy `OllamaEmbeddings` import path as an alias to avoid unrelated call-site churn
## Concerns
- the focused harness verification still emits two pre-existing warnings:
- `DeprecationWarning` from `testcontainers.postgres`
- `FutureWarning` because `resources` currently flows through the legacy config translation path
## Fix round 1 — 2026-08-08
### Findings addressed
- HIGH: external top-level `embeddings` remained an operational fallback and could still load
- MEDIUM: non-object embed JSON payloads escaped as raw `AttributeError`
### RED evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py tests/test_ollama_ensure.py -q
```
Observed before the fix:
- exit code `1`
- `2 failed, 36 passed, 2 warnings`
Representative failures:
- `AttributeError: 'list' object has no attribute 'get'` from `response.json()` returning a JSON array
- `Failed: DID NOT RAISE ConfigError` for top-level external `embeddings.provider=openai_compatible`
### GREEN evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py tests/test_ollama_ensure.py -q
```
Observed after the fix:
- exit code `0`
- `38 passed, 2 warnings`
Touched-file lint:
```bash
cd harness
./.venv/bin/ruff check tht/config.py tht/vectorstore/embeddings.py tests/test_internal_embeddings.py tests/test_config_resources.py
```
- exit code `0`
- `All checks passed!`
### Minimal fix
- validated the final active `cfg.embeddings` contract after config loading, so legacy top-level
embedding inputs now fail explicitly unless they exactly match the internal Ollama contract
- converted non-mapping embed JSON payloads into controlled `EmbeddingsError` failures with
sanitized diagnostics instead of raw attribute errors
@@ -1,170 +0,0 @@
# Task 5 Report — Implement the Qdrant VectorStore adapter
## Status
Implemented on 2026-08-08 in `/Users/mp/projects/ThothII/.worktrees/git-workspace-registry`.
## RED evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
```
Observed before implementation:
- exit code `2`
- collection failed during import because the adapter did not exist yet
Representative failures:
- `ModuleNotFoundError: No module named 'tht.adapters.vector.qdrant'`
## GREEN evidence
Focused behavior suite:
```bash
cd harness
./.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
```
- exit code `0`
- `31 passed, 1 warning`
Touched-file lint:
```bash
cd harness
./.venv/bin/ruff check tht/adapters/vector/qdrant.py tht/adapters/vector/__init__.py \
tht/ports/vector.py tht/vectorstore/records.py tht/vectorstore/store.py \
tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py
```
- exit code `0`
- `All checks passed!`
Patch hygiene:
```bash
git diff --check
```
- exit code `0`
## What changed
- added `QdrantVectorStore` with direct `requests`-based REST calls for:
- `GET /collections/{collection}`
- `PUT /collections/{collection}`
- `PUT /collections/{collection}/index`
- `PUT /collections/{collection}/points?wait=true`
- `POST /collections/{collection}/points/query`
- `POST /collections/{collection}/points/scroll`
- `POST /collections/{collection}/points/delete?wait=true`
- implemented idempotent collection provisioning for `1024` dimensions and `Cosine` distance
- created deterministic UUIDv5 point IDs from workspace, semantic kind, and canonical record key
- preserved canonical record identity and only upserted/deleted points matching the exact workspace
and generation filters
- added Qdrant payload helpers so stored payloads carry:
- `workspace_id`
- grouped semantic `kind` (`schema`, `evidence`, `memory`)
- original `record_kind`
- canonical `record_key`
- `content_hash`
- existing Thoth metadata fields
- mapped Qdrant payloads back into existing `VectorHit` objects without losing the original
Thoth kind
- exported the new adapter from the public vector adapter package and added focused contract tests
- sanitized timeout and malformed-response failures so CLI-facing callers do not leak raw endpoint
details
## Self-review
- confirmed collection mismatch fails without any delete/recreate path
- confirmed every query/scroll/delete operation includes a workspace filter
- confirmed the adapter never deletes or rewrites unrelated Qdrant points
- added keyword payload indexes for all filter-critical fields used here, including `document_id`
for exact Evidence filtering
## Concerns
- the requested `adversarial-review` skill could not run its full external reviewer flow in this
environment because the skill’s referenced `brain/` files are missing at
`/Users/mp/.agents/skills/adversarial-review`; I performed a manual adversarial self-review
instead
- the focused suite still emits one pre-existing warning from `testcontainers.postgres`
## Fix round 1 — 2026-08-08
### Findings addressed
- IMPORTANT: metadata collisions could override canonical Qdrant payload identity fields and break
workspace isolation
- IMPORTANT: scroll-based operations only read the first page and did not follow
`next_page_offset`, making `existing_hashes`, `list_evidence_generations`, and delete counts
inexact beyond one page
### RED evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
```
Observed before the fix:
- exit code `1`
- `2 failed, 31 passed, 1 warning`
Representative failures:
- `assert payload["workspace_id"] == "demo"` failed because colliding `record.metadata`
overwrote canonical payload fields
- paginated scroll test missed later pages, so `existing_hashes` and generation cleanup counts
were incomplete
### GREEN evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
```
Observed after the fix:
- exit code `0`
- `33 passed, 1 warning`
Touched-file lint:
```bash
cd harness
./.venv/bin/ruff check tht/adapters/vector/qdrant.py tht/vectorstore/records.py \
tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py
```
- exit code `0`
- `All checks passed!`
Patch hygiene:
```bash
git diff --check
```
- exit code `0`
### Minimal fix
- made `qdrant_payload` apply canonical fields after `record.metadata` so workspace ID, semantic
kind, original record kind, canonical record key, and content hash cannot be overridden by
metadata collisions
- paginated `_scroll` until `next_page_offset` is absent, sent the returned `offset` back on the
next request, and reject repeated offsets as malformed to avoid infinite loops
@@ -1,144 +0,0 @@
# Task 6 Report
Date: 2026-08-08
Status: implemented and verified
Summary:
- Added schema-v3 Qdrant runtime support to the harness config/resource layer and vector factory.
- Made Qdrant payloads carry `workspace_id` and `workspace_revision` on every point.
- Routed schema and memory bulk indexing through the transport-neutral vector port with canonical hash-based dedup.
- Kept Evidence canonical on filesystem and Memory canonical in JSONL; Qdrant remains derived/rebuildable.
- Added focused tests for semantic-kind isolation, shared identity fields, search-pack kind boundaries, and the schema-v3 factory/config path.
Files changed:
- `harness/tht/config.py`
- `harness/tht/config_compat.py`
- `harness/tht/adapters/factory.py`
- `harness/tht/adapters/vector/qdrant.py`
- `harness/tht/vectorstore/records.py`
- `harness/tht/cli/vector_cmd.py`
- `harness/tht/cli/memory_cmd.py`
- `harness/tests/test_semantic_kind_isolation.py`
- `harness/tests/test_memory_save_one.py`
- `harness/tests/test_search_pack.py`
- `harness/tests/test_qdrant_vector_store.py`
- `harness/tests/test_adapter_factory.py`
- `harness/tests/test_config_resources.py`
Verification:
- Focused RED/GREEN task suite:
- `cd harness && .venv/bin/pytest tests/test_semantic_kind_isolation.py tests/test_memory_save_one.py tests/test_search_pack.py -q`
- Relevant harness suite:
- `cd harness && .venv/bin/pytest tests/test_semantic_kind_isolation.py tests/test_memory_save_one.py tests/test_search_pack.py tests/test_qdrant_vector_store.py tests/test_adapter_factory.py tests/test_config_resources.py tests/test_vector_port_contract.py tests/test_corpus_pipeline.py -q`
- Result: `131 passed`
- Changed-file Ruff:
- `cd harness && .venv/bin/ruff check tht/vectorstore/records.py tht/adapters/vector/qdrant.py tht/config_compat.py tht/config.py tht/adapters/factory.py tht/cli/vector_cmd.py tht/cli/memory_cmd.py tests/test_memory_save_one.py tests/test_search_pack.py tests/test_semantic_kind_isolation.py tests/test_qdrant_vector_store.py tests/test_adapter_factory.py tests/test_config_resources.py`
- Result: clean
Concerns / follow-up:
- `memory clear` still retains its older direct-vector assumptions and was not expanded in this task because the brief focused on canonical builders and schema/evidence/memory routing through the active Qdrant path.
- The relevant suite still emits pre-existing warnings (legacy config deprecation in older fixtures, plus existing Pydantic serializer warnings in corpus tests), but they are not introduced by this task.
## Fix round 1 (2026-08-08)
Scope:
- Fixed qdrant-only schema-v3 command gating for `vector index-schema`, `memory promote`, and `memory index`.
- Replaced `memory clear`'s direct-pgvector-only path with vector-port deletion by kind.
- Added focused qdrant-only CLI regression tests and refreshed older CLI fixtures to the enforced internal embedding contract.
RED evidence:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py -q`
- Initial result against commit `5e39cfa`: `4 failed`
- Failure signatures:
- `ERRORE: sezioni mancanti nel workspace yaml: vector_db o vector_write_rest.`
- `ERRORE: sezioni mancanti nel workspace yaml: vector_db.`
GREEN evidence:
- Focused fix suite:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py tests/test_qdrant_vector_store.py tests/test_adapter_factory.py tests/test_config_resources.py tests/test_memory_save_one.py tests/test_search_pack.py -q`
- Result: `51 passed`
- Relevant broader vector/memory/schema/search suite:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py tests/test_qdrant_vector_store.py tests/test_adapter_factory.py tests/test_config_resources.py tests/test_memory_save_one.py tests/test_search_pack.py tests/test_vector_port_contract.py tests/test_adapter_command_regressions.py tests/test_solved_search_cli.py tests/test_schema_introspect_guard.py tests/test_semantic_kind_isolation.py tests/test_corpus_pipeline.py -q`
- Result: `154 passed`
- Ruff on the fix surface:
- `cd harness && .venv/bin/ruff check tht/ports/vector.py tht/adapters/vector/qdrant.py tht/adapters/vector/pgvector.py tht/adapters/vector/thoth_http.py tht/vectorstore/rest_client.py tht/cli/vector_cmd.py tht/cli/memory_cmd.py tests/test_qdrant_cli_commands.py tests/test_solved_search_cli.py`
- Result: clean
Notes:
- `memory clear` now deletes derived `kind=memory` points through the configured writable vector store, while leaving the JSONL registry as the source of truth until the registry file is removed by the command.
- The broader suite still carries the same pre-existing warnings noted above; this fix round did not add new warnings or failures.
## Fix round 2 (2026-08-08)
Scope:
- Removed the accidental HTTP writer `delete_kinds` capability expansion from `ThothHttpVectorStore` and `VectorRestClient`.
- Reworked `memory clear` so schema-v3 Qdrant uses scoped `kind=memory` deletion, while legacy transports keep the pre-task direct-sync path instead of advertising a nonexistent RPC.
- Tightened the qdrant-only memory-clear regression to assert the exact `("memory", ["memory"])` delete scope.
RED evidence:
- Re-review found a transport contract mismatch in fix round 1:
- `ThothHttpVectorStore` exposed `delete_kinds(...)`
- `VectorRestClient` exposed `delete_kinds(...)`
- but the legacy HTTP writer migration only allowlists `delete_vector_generation`, not `delete_vector_kinds`
- The new regressions added in this round capture that mismatch and the missing qdrant delete-scope assertion:
- `tests/test_vector_port_contract.py::test_http_store_supports_writer_without_reader`
- `tests/l0/test_vector_adapter_parity.py::test_http_rest_client_does_not_advertise_nonexistent_delete_kinds_rpc`
- `tests/test_qdrant_cli_commands.py::test_memory_clear_accepts_qdrant_only_runtime_config`
GREEN evidence:
- Focused regression suite:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py tests/test_vector_port_contract.py tests/l0/test_vector_adapter_parity.py tests/test_adapter_command_regressions.py -q`
- Result: `53 passed`
- Broader relevant vector/memory/search suite:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py tests/test_adapter_command_regressions.py tests/test_vector_port_contract.py tests/l0/test_vector_adapter_parity.py tests/test_solved_search_cli.py tests/test_qdrant_vector_store.py tests/test_search_similar_kinds.py tests/test_corpus_pipeline.py -q`
- Result: `135 passed`
- Ruff on the changed fix surface:
- `cd harness && .venv/bin/ruff check tht/cli/memory_cmd.py tht/ports/vector.py tht/adapters/vector/thoth_http.py tht/vectorstore/rest_client.py tests/test_qdrant_cli_commands.py tests/test_vector_port_contract.py tests/l0/test_vector_adapter_parity.py`
- Result: clean
Notes:
- Legacy HTTP/vector-rest deployments do not gain a new destructive RPC surface from this fix; they keep their previous behavior and continue to fail closed for unsupported cleanup.
- The broader suite still emits the same pre-existing deprecation and serializer warnings already noted above; this round did not introduce new warnings.
## Fix round 3 (2026-08-08)
Scope:
- Added an adapter-level Qdrant regression for mixed semantic kinds within one workspace plus a second workspace memory point.
- Verified that `delete_kinds("memory", ["memory"])` emits the real adapter filter with both `workspace_id=demo` and `record_kind=memory`.
- Verified that non-memory semantic kinds in the same workspace and memory from another workspace survive the delete.
RED evidence:
- Re-review identified a test gap rather than a confirmed runtime bug:
- existing coverage asserted only the CLI mock call shape for qdrant memory clear
- there was no adapter-level regression proving the real Qdrant delete filter and resulting fake-Qdrant state across mixed semantic kinds/workspaces
- Added regression:
- `tests/test_qdrant_vector_store.py::test_delete_kinds_is_workspace_scoped_and_preserves_other_semantic_kinds`
GREEN evidence:
- Requested focused suite:
- `cd harness && .venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_qdrant_cli_commands.py tests/test_semantic_kind_isolation.py -q`
- Result: `18 passed`
- Ruff on changed files:
- `cd harness && .venv/bin/ruff check tests/test_qdrant_vector_store.py`
- Result: clean
Notes:
- This round required no production change; the new adapter regression passed against the existing Qdrant implementation.
- The focused suite still emits the same pre-existing `testcontainers.postgres` deprecation warning from `tests/conftest.py`; no new warnings were introduced.
@@ -1,98 +0,0 @@
# Task 7 report — mandatory Qdrant and Ollama Compose services
Date: 2026-08-08
Status: completed
Summary:
- Added mandatory private `qdrant`, `embedding`, and `embedding-model-init` services to the base Compose stack.
- Pinned Qdrant `v1.18.2` and Ollama `0.32.0` by immutable multi-arch digest.
- Persisted Qdrant storage in `qdrant-data` and Ollama model cache in `embedding-models`.
- Wired `core` to fixed internal semantic endpoints:
- `THT_INTERNAL_QDRANT_URL=http://qdrant:6333`
- `THT_INTERNAL_EMBEDDING_URL=http://embedding:11434`
- `THT_INTERNAL_EMBEDDING_MODEL=qwen3-embedding:0.6b`
- `THT_INTERNAL_EMBEDDING_DIMENSIONS=1024`
- Removed external vector / embedding endpoint requirements from the local and server env examples.
- Added an idempotent Ollama model bootstrap script that:
- waits up to a bounded deadline for `/api/tags`
- skips `ollama pull` when the model is already cached
- pulls `qwen3-embedding:0.6b` only when needed
- verifies the model appears in `/api/tags` after pull
- Added optional GPU override file `deploy/compose.embedding-gpu.yaml`; base Compose remains CPU-only.
- Updated `scripts/run-stack.sh` so the GPU override is included only when `THOTH_ENABLE_EMBEDDING_GPU=1`.
Verification:
- RED confirmed before implementation:
- `./scripts/test-default-compose.sh` failed on missing required services.
- `./scripts/test-unified-compose.sh` failed on missing required services.
- `./scripts/test-internal-semantic-compose.sh` failed because the GPU override file did not exist.
- GREEN after implementation:
- `./scripts/test-default-compose.sh`
- `./scripts/test-unified-compose.sh`
- `./scripts/test-internal-semantic-compose.sh`
- `git diff --check`
- Additional shell verification:
- `scripts/run-stack.sh --wait` includes only base + local Compose files by default.
- `THOTH_ENABLE_EMBEDDING_GPU=1 scripts/run-stack.sh --wait` adds `deploy/compose.embedding-gpu.yaml`.
Resolved image digests:
- `qdrant/qdrant:v1.18.2@sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c`
- `ollama/ollama:0.32.0@sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a`
Self-review:
- The first bootstrap-script draft depended on tools not guaranteed inside the Ollama image. This was corrected after image inspection; the final script uses only confirmed image tools (`bash`, `ollama`, `grep`) plus raw HTTP over `/dev/tcp`.
- The server overlay intentionally replaces most named core volumes with bind mounts, so the unified contract was tightened to require named semantic-cache volumes there while preserving the local/base named-volume checks.
Concerns:
- The model bootstrap waits for Ollama readiness and verifies cache state, but the first real cold-start will still take time to download `qwen3-embedding:0.6b`.
- The GPU override requests generic Docker GPU capability only; actual GPU availability remains host/runtime dependent and intentionally stays opt-in.
## Fix round 1 / 5 — 2026-08-08
Rulings applied:
- Kept the Task 1 boundary intact: schema-v3 remains the only operational workspace descriptor shape.
- Did not restore any external semantic fallback for schema-v2 live sessions.
- Treated `PROJECT_STATE.md` as stale documentation for this point, not runtime truth.
Focused schema-v2 evidence:
- Re-ran the existing targeted registry test:
- `cd backend && npx vitest run test/workspace-registry.test.ts -t "lists a schema v2 descriptor as migration_required and refuses to acquire it"`
- Result: pass.
- Evidence from that test:
- schema-v2 descriptors list as `migration_required`
- `acquireSessionRevision("psd-clinical")` rejects with `code: "workspace_invalid"`
- Conclusion: schema-v2 acquisition remains blocked; no external semantic fallback was reintroduced.
Contract consistency fixes:
- Updated `harness/tests/test_local_compose_contract.py` to assert the mandatory internal semantic stack, fixed internal core semantic env, private-service topology, persistent volumes, and Ollama health/dependency contract.
- Updated shell Compose contracts to require:
- Ollama healthcheck on `embedding`
- `embedding-model-init` dependency on `embedding: service_healthy`
- Updated `scripts/unified-deployment-smoke.sh` rendered-contract helper to expect the mandatory internal semantic topology and internal semantic env names, and to reject retired external semantic bindings.
- Updated `scripts/test-task13-runtime-fixtures.sh` to exercise `task13_assert_rendered_contract` for both local and server fixture renders.
Fix round 1 verification:
- RED before implementation:
- `cd harness && .venv/bin/pytest tests/test_local_compose_contract.py -q` failed because `embedding` had no healthcheck.
- `./scripts/test-default-compose.sh` failed because `embedding` had no healthcheck.
- `./scripts/test-unified-compose.sh` failed because `embedding` had no healthcheck.
- `./scripts/test-task13-runtime-fixtures.sh local` failed because `unified-deployment-smoke.sh` still expected `core,frontend`.
- GREEN after implementation:
- `./scripts/test-default-compose.sh`
- `./scripts/test-unified-compose.sh`
- `./scripts/test-internal-semantic-compose.sh`
- `cd harness && .venv/bin/pytest tests/test_local_compose_contract.py -q`
- `./scripts/test-task13-runtime-fixtures.sh local`
- `./scripts/test-task13-runtime-fixtures.sh server`
- `cd backend && npx vitest run test/workspace-registry.test.ts -t "lists a schema v2 descriptor as migration_required and refuses to acquire it"`
- `docker compose --env-file deploy/env/local.env.example -f compose.yaml -f deploy/compose.local.yaml config --format json`
@@ -1,65 +0,0 @@
Status: completed on August 8, 2026.
Summary:
- Updated the frontend workspace contract from schema v2 editing to schema v3 publishing.
- Kept only `semantic_index.vector_store.collection` editable; rendered qdrant / internal Ollama semantic values as fixed read-only architecture values.
- Removed external vector transport / endpoint / credential / embedding diagnostics branches from frontend draft sanitization, conflict parsing, and editor UI.
- Added a migration-required banner in workspace management and blocked `migration_required` workspaces from new-session selection.
- Aligned the example workspace YAML comments with the fixed internal qdrant/Ollama architecture.
Files changed:
- `frontend/src/api/workspaces.ts`
- `frontend/src/api/workspaces.test.ts`
- `frontend/src/workspaces/drafts.ts`
- `frontend/src/workspaces/drafts.test.ts`
- `frontend/src/shell/WorkspaceEditor.tsx`
- `frontend/src/shell/WorkspaceEditor.test.tsx`
- `frontend/src/shell/WorkspaceManager.tsx`
- `frontend/src/shell/WorkspaceManager.test.tsx`
- `frontend/src/shell/WorkspacePublishDialog.test.tsx`
- `frontend/src/api/sessions.ts`
- `frontend/src/shell/SteerInput.tsx`
- `frontend/src/shell/SteerInput.test.tsx`
- `deploy/workspaces/example.yaml`
- `deploy/workspaces/psd.yaml.example`
Verification:
- `cd frontend && npx vitest run src/shell/SteerInput.test.tsx src/shell/WorkspaceEditor.test.tsx src/shell/WorkspaceManager.test.tsx src/shell/WorkspacePublishDialog.test.tsx src/workspaces/drafts.test.ts src/api/workspaces.test.ts`
- Result: 6 files passed, 59 tests passed.
- `cd frontend && npx tsc -b`
- Result: passed.
- `git diff --check`
- Result: passed.
Self-review:
- The frontend now publishes the exact schema v3 semantic shape and no longer persists legacy semantic transport/credential branches.
- Migration-required workspaces are visible in management with an explicit banner and are excluded from the composer workspace selector.
- One dependent test file outside the original brief list (`WorkspacePublishDialog.test.tsx`) and the composer/session-selection path (`api/sessions.ts`, `SteerInput.tsx`, related test) were updated because they were directly coupled to the old v2 semantic/edit-selection behavior.
Concerns:
- The composer still retains backward-compatible behavior for summaries that omit `revision` entirely; only explicit `revision.state === "migration_required"` is blocked. That matches the current mixed-test environment, but once summary responses are guaranteed to include `revision`, that fallback may be removable.
Fix round 1/5 — August 8, 2026
Summary:
- Made missing or invalid workspace summaries fail safe in frontend session creation and composer selection instead of falling open as legacy.
- Added an actionable unavailable message in workspace management for incomplete summaries with no canonical revision.
- Replaced the old runtime-oriented example descriptor files with exact backend WorkspaceV3 descriptor YAML.
Additional files changed:
- `frontend/src/api/sessions.test.ts`
- `backend/test/workspaces-schema.test.ts`
Fix-round verification:
- `cd frontend && npx vitest run src/api/sessions.test.ts src/shell/SteerInput.test.tsx src/shell/WorkspaceManager.test.tsx src/shell/WorkspaceEditor.test.tsx src/shell/WorkspacePublishDialog.test.tsx src/workspaces/drafts.test.ts src/api/workspaces.test.ts`
- Result: 7 files passed, 73 tests passed.
- `cd frontend && npx tsc -b`
- Result: passed.
- `cd backend && npx vitest run test/workspaces-schema.test.ts`
- Result: 1 file passed, 17 tests passed.
- `git diff --check`
- Result: passed.
Notes:
- Missing `revision` in a workspace summary now fails with the same session/composer safety posture as `migration_required`, using the existing safe workspace-policy error for session creation and an explicit unavailable message in workspace management.
- The committed example files now validate as actual schema-v3 descriptors instead of deployment/runtime templates with forbidden semantic endpoint fields.
@@ -1,83 +0,0 @@
# Task 5 Report — `tht setup` lifecycle orchestration
## Status
Completed. `tht setup` now validates the checkout and host prerequisites, creates or validates
the non-secret installation files, validates Compose, and by default builds, starts, health-checks,
and verifies the installation. `tht setup --configure-only` stops immediately after successful
Compose rendering.
## Implementation
- Added `setup.Run`, with an ordered host preflight: project/worktree discovery, Docker Engine,
Docker Compose, supported architecture, and LF line-ending checks.
- Reused `config.Installation.ComposeArgs` for all Compose calls and added a narrow
`compose.InstallationRunner` adapter for Pi diagnostics; no shell command construction was added
to the top-level CLI parser.
- Default setup performs `compose build`, `compose up --detach --remove-orphans`, bounded polling
for `core`, `frontend`, `qdrant`, `embedding`, and `embedding-model-init`, then aggregate volume
diagnostics and `pi.Doctor`.
- Health timeout errors identify the last failing service and preserve containers for diagnosis,
with `tht logs <service>` and `tht status` guidance.
- Completion output includes the frontend URL, selected descriptor, and next action.
## TDD evidence
The initial focused test run failed because `setup.Run` did not exist. Tests were then written
against a fake Compose runner before the orchestration was implemented. They cover the complete
ordered flow, configure-only stop, preflight failure before writing configuration, health retry,
timeout guidance, and CLI default versus `--configure-only` dispatch.
## Verification
Executed from `tools/tht`:
```bash
go test ./internal/setup ./internal/compose ./cmd/tht -run 'TestRun|TestSetupCommand|TestInstallationRunner' -count=1
go test ./internal/setup ./internal/compose ./cmd/tht -count=1
go test ./...
git diff --check
```
All commands passed. No actual Docker build, container start, live-stack restart, system
installation, Pi configuration edit, or documentation rewrite was performed.
## Commit
`feat(setup): build start and verify ThothII` (this report is included in that commit).
## Concerns
- The bounded health wait is verified with fakes only, as required for this task. Real Docker
lifecycle verification belongs to the later live acceptance task.
- The existing aggregate `tht doctor` command remains a separate implementation; Task 5 performs
its equivalent setup-time prerequisite checks plus `pi.Doctor` without invoking a nested CLI
process.
## Fix round 1
The independent review identified three gaps. All were reproduced with RED tests before the
production change:
- A rendered Compose document containing any one volume was accepted. `requireVolumes` now
requires `settings`, `pi-state`, `workspace-registry`, `workspace-secrets`, `sessions`,
`qdrant-data`, and `embedding-models`; tests reject each individual omission and an
unrelated-only volume set.
- Failures after `compose up` could return without recovery instructions. A single recovery
wrapper now preserves the underlying error while adding the retained-container, `tht logs
<service>`, and `tht status` guidance for failed `up`, health, aggregate doctor, and Pi doctor
phases. Focused tests also prove build failure stops before attempting startup.
- LF inspection previously walked the full checkout. It now inspects only `compose.yaml`,
`deploy/`, and `docker/`; a test proves CRLF content under `node_modules/` is ignored.
Verification added for this round:
```bash
go test ./internal/setup -run 'TestRequireVolumes|TestRun(BuildFailure|UpFailure|AggregateDoctorFailure|PiDoctorFailure|IgnoresIrrelevant|TimesOut)' -count=1
go test ./internal/setup -count=1
```
Both passed before the final full-suite verification. No Docker or live operation was run.
Implementation commit evidence: `ea70cc95b04532043744a9de6c5912e30a214595` —
`fix(setup): harden verification and recovery`.
@@ -1,55 +0,0 @@
# Task 6 — Version, aggregate doctor, and build-aware start
Status: complete.
Implemented the host-side `tht version`, aggregate `tht doctor [--json]`, and `tht start [--build]` contracts.
- `version` is descriptor-free and reports semantic version, commit, build time, OS, and architecture.
- `doctor` emits typed, redacted checks for descriptor state, Docker/Compose, rendered volumes, file permissions, service health, workspace registry, the container-local workflow doctor, and Pi doctor. Its JSON mode writes exactly one JSON document to stdout.
- The Python workflow doctor is invoked only as `docker compose exec -T core tht doctor --json` after core is running.
- `start` uses the shared lifecycle service: default `up → health`; `--build` is `build → up → health`.
- `setup` now reuses the shared lifecycle and aggregate diagnostics rather than keeping parallel health/volume implementations.
Verification performed without live Docker/container commands:
```bash
cd tools/tht
go test ./internal/version ./internal/doctor ./internal/service ./cmd/tht \
-run 'TestVersion|TestDoctor|TestStart|TestCurrent|TestRun' -count=1
go test ./internal/setup -count=1 -run 'TestRun' -v
go test ./... -count=1
git diff --check
```
All completed successfully. The intentionally fake runner coverage includes unavailable Docker,
stopped/running core, workflow failure redaction, pristine JSON output, and start ordering.
Concerns: no live Docker validation or host installation was run, by explicit task constraint.
## Fix round 1
Completed the independent-review follow-up without live Docker operations.
- `workspace-registry` now executes a container-local, read-only Node validation of
`/data/workspace-registry/state/active.json` and every declared snapshot descriptor. It no
longer passes merely because Compose declares a volume.
- Host file permissions are checked before Docker/Compose availability and therefore remain
visible as failures when Docker is unavailable.
- Separate typed, bounded HTTP probes verify core (`curl --max-time 5`) and frontend
(`wget -T 5`) reachability, independently of Compose health. The probe is injectable in tests.
- The successful report tests assert the stable full checklist:
`descriptor`, `files`, `docker`, `compose`, `configuration`, `services`, `core-http`,
`frontend-http`, `workspace-registry`, `workflow`, `pi`.
Additional verification:
```bash
cd tools/tht
go test ./internal/doctor -run 'TestRun(ChecksUnsafeFilesEvenWhenDockerIsUnavailable|FailsAnInvalidContainerLocalRegistryState|ReportsEachHTTPReachabilityProbeFailure|UsesOnlyContainerLocalWorkflowAndPiDiagnosticsWhenCoreRuns)' -count=1 -v
go test ./internal/doctor ./internal/setup ./internal/service ./cmd/tht -count=1
go test ./... -count=1
git diff --check
```
All passed with fake runners/probes only. No live container, HTTP endpoint, or host installation
was touched.
@@ -1,163 +0,0 @@
# Task 15 retained release-gate report — fix round 5 (sanitized)
## Final-review fix-round-2 addendum — frozen source `2a9359071257f9b8a71d36ec2bbb25b161003f81`
This addendum supersedes the fix-round-1 addendum for current authentication remediation status
while preserving the fix-round-5 material below as historical provenance.
- Authentication remediation status: `PASS`. The three original remediation Important findings
remain `RESOLVED`; the fix-round-2 fully bounded lifecycle Important is `ADDRESSED`; and the
temporary Windows diagnostic-matrix Minor is `ADDRESSED`.
- Overall branch/release readiness is separately `FAIL`, with unavailable external/manual gates
`PENDING`.
- Completed exact-source workflow run `32147345625` concluded `failure` on baseline release jobs.
Its `Windows clone and Compose contract` job (`95744249248`) executed the unfiltered command
`go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1`; the native step
passed all three packages: safeio `22.058s`, backup `7.161s`, authstorage `16.088s`.
- The Windows job failed only afterward in the baseline clone-contract script at
`scripts/test-windows-clone-contract.ps1:208`, where PowerShell rejects the undelimited
`$remoteYaml:` variable reference.
- `LF, Compose, docs, and TypeScript` job `95744249458` reproduced the baseline unset-`TMPDIR`
failure after unified Compose passed. Linux Docker job `95744249354` reproduced the missing-`rg`
prerequisite failure; cleanup passed and no image manifest was generated.
- The skipped Windows Docker Desktop/WSL2 job is recorded as `NOT_RUN` / `BLOCKED`, not FAIL.
Downstream commands skipped after executed baseline failures use the same classification. The
matrix contains an explicit native `windows_stagearchive_retained_capability` PASS row.
- Historical Node/auth/browser/docs PASS and harness/Ruff/Compose FAIL evidence remains bound to
its recorded source where not rerun. L2, real PSD/manual acceptance, and provider readiness
remain `PENDING`.
- Current machine-readable evidence and the requested Task 4 report are recorded in
`.artifacts/task-15/automated-gates.json` and
`.superpowers/sdd/2026-08-18-thothii-authentication-remediation/task-4-report.md`.
- The full fix-round-2 RED/GREEN and finding disposition is recorded in
`.superpowers/sdd/2026-08-18-thothii-authentication-remediation/fix-round-2-report.md`.
- Current automated-gates SHA-256:
`6c516db5c2064c4a4a2e5f25961b993cd4a8fe020bbbb822fbac7faa0c119599`.
- Historical unified Docker manifest SHA-256: `9c8dec4546909fd93799dbcf374bcb3a89bc46cfe0fd482472c0cbe757ddf5b6`.
The complete sanitized Task 4 matrix and the separate remediation/release verdicts are in the
requested Task 4 report.
- Final tested source commit: `74b062f1a737103524cbe706346cfd65f87cdfd1`.
- Historical retained source commits: fix-round-2 `fe190e7046acc173f510dddcb32f46ed142858c1`,
maintenance follow-up `4d230b87afdcd24f02264f8f937c8628b92db05a`, prior final Docker
source `e20bf33e2a00102192e5be66b178037aeca3a7b1`, and fix-round-4 streamed
archive privacy `54698e73400a54ce7c3e6c10099e14eb471ce8b9`.
- Versions: Node contract `v24.16.0`; host default Node `v25.6.1`; Go `go1.26.5`;
Pi `0.80.3`.
- Historical automated gate artifact: `.artifacts/task-15/automated-gates.json`;
SHA-256 `7d9ec93af15510605f1aa7179b26a7ee46d78122f647854300f7a9922057a63f`.
- Docker image manifest: `.artifacts/task-15/unified-docker-images.json`;
SHA-256 `9c8dec4546909fd93799dbcf374bcb3a89bc46cfe0fd482472c0cbe757ddf5b6`.
## Fix-round-5 evidence
- PASS, RED then GREEN: `TestCreateCanonicalNewPrivateFileUsesPinnedParentAfterAncestorSwap`
first failed because the creator had not retained its parent before creation. It now opens every
Unix ancestor once, creates the leaf with `openat(O_NOFOLLOW|O_CREAT|O_EXCL)`, applies and checks
`0600` by descriptor (`fchmod`/`fstat`), and uses `unlinkat` for creator failure cleanup. The
deterministic test moves the opened parent, replaces its lexical name with an outside symlink,
validates the archive under the moved original parent, and proves no outside archive was written.
- PASS: the Windows implementation uses NT `RootDirectory`-relative traversal for every component
after the volume root and for final file creation. The retained final parent receives only the
required child-create right (`FILE_WRITE_DATA` for a file, `FILE_APPEND_DATA` for a directory),
reparse points are rejected, and the owner-only protected DACL is installed in the same
`NtCreateFile` operation. The native-Windows test attempts the pre-create parent swap and calls
`safeio.ValidatePrivateRegular`; it is compiled but not executed on this host.
- PASS: `go test ./internal/safeio ./internal/backup -count=1`, `go test -race ./...` across
`18` packages, `go vet ./...`, and a native host `tht` CLI build. Existing StageArchive
capacity, lifecycle, rollback, streaming, and cleanup tests remain passing.
- PASS, compile-only: Windows amd64 static test/build compilation across `18` packages, including
the retained-handle Windows tests. No Windows executable was run; native execution remains
PENDING and is not inferred from compilation.
- PASS on Node `v24.16.0`: the hermetic OIDC/F1 authentication browser smoke passed all current
`8` checks in `frontend/e2e/auth.spec.ts` and `frontend/e2e/f1.spec.ts`; the runtime sentinel
leak scan passed.
- PASS: shell syntax, unified-smoke safety self-test, default Compose contract, unified Compose
contract, and Compose secret-policy contract.
- PASS: final unified Docker deployment smoke run `20260818070637-66409-30058`, bound exactly to
source `74b062f1a737103524cbe706346cfd65f87cdfd1`. It exercised maintenance-auth isolation,
restore, registry lifecycle, bad-candidate rollback, image revalidation, and task-scoped cleanup.
## Sanitized final unified Docker output
```text
== Build and start isolated local Compose distribution ==
== Recreate offline and retain the validated registry snapshot ==
== Pull a valid catalog+descriptor metadata update ==
== Pull a content-only Git Evidence update ==
== Reject catalog/descriptor metadata mismatch and retain the valid snapshot ==
== Reject orphan descriptor directories not listed in the catalog ==
== Reject the retired flat workspace layout and retain the valid snapshot ==
== Inject a bad pinned Pi candidate and prove automatic rollback ==
Task 13 full deployment smoke passed.
Task 13 cleanup proof: no labeled containers, volumes, networks, or images remain for 20260818070637-66409-30058.
```
## Sanitized Docker image identities
- `sha256:2d7b19491c7eb8c119c3cedb390aaeb2ff5593f6fc43ab66c317565560da6d7d`;
roles `compose-runtime`, `fixture-runtime`.
- `sha256:3b6c31a5d8f8fc58fa3233391b6175bd2fbc793eebb44d5e285ecc6e02e9e687`;
role `compose-runtime`.
- `sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a`;
role `compose-runtime`.
- `sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c`;
role `compose-runtime`.
- `sha256:c3cbe1cc1aa588a64951ac6286e0df7b27fe2e6324b1001c619bb358770c0178`;
role `rollback-candidate`.
For each image, the retained repository-digest component equals the listed image digest. Registry
names and credentials are deliberately omitted.
## Complete observed matrix
- PASS: Task 13 lifecycle carry-ins; retained-handle owner-private restore staging; provider fixture
round-one `6/6`; backend Node 24 round-one suite `75 files / 1081 tests`; frontend Node 24
round-one suite `61 files / 444 tests`; current Node 24 authentication/F1 browser smoke `8/8`;
final-source Go race/build `18 packages`; Windows static cross-compile `18 packages`; harness
round-one suite `921 passed / 4 L2 deselected`; authentication docs round-one gate; shell/Compose
contracts; final unified Docker smoke; five-image traceability; and Docker cleanup.
- FAIL: Ruff `192` known-baseline errors; MkDocs strict `69` known-baseline warnings; existing
canonical/workspace install wording checks; existing Pi model-policy check; deployment-coupling
scan against preserved ignored private material.
- PENDING: native Windows execution because required host prerequisites are unavailable; L2 because
the configured secret layout is unavailable; real PSD/manual acceptance because no real
identity/access is available; isolated provider readiness because an unrelated host port is
occupied.
## Final Task 15 review after fix round 5
The fresh Terra review verdict is **CHANGES REQUIRED**. The five-round breaker is exhausted; no
sixth implementation round was started. Two Important findings remain:
- `StageArchive` does not retain the opaque parent/directory capability through the complete
stream and `Close` lifecycle. Staging-directory creation and final cleanup still use pathname
operations, so an ancestor swap after creation can strand the secret-bearing archive or redirect
cleanup. Deterministic StageArchive swap-and-cleanup coverage is still required on Unix and
native Windows.
- Windows claim removal closes its validated retained parent handles before calling pathname-based
`DeleteFile`. Removal must instead remain handle-relative (or delete through the opened handle),
with a native-Windows ancestor-swap test.
The focused/full Go, cross-compile, Node 24, browser, Compose, Docker lifecycle, image-traceability,
and cleanup results above remain valid evidence for source `74b062f1a737103524cbe706346cfd65f87cdfd1`.
They do not override the final code-review verdict. Native Windows execution remains PENDING.
The authentication feature is **not implementation-complete or release-complete** while these code
findings and the required FAIL/PENDING gates remain. No secret values, real identities, internal
endpoints, or registry names are retained.
## Final whole-branch review
The final read-only Terra review of `351361f..39b5453` also returned **CHANGES REQUIRED** and found
one additional Important issue: the POSIX local-user registry validates file type, link count, and
mode for `users.yaml` and its parent directory, but does not require ownership by the effective UID.
A foreign-owned `0600` registry inside a runtime-owned `0700` directory can remain writable by the
foreign owner and be used to alter credentials or grant the administrator role. The registry must
enforce effective-UID ownership on every POSIX `lstat`/`fstat` path and add foreign-owner rejection
coverage.
No new Critical issue or load-bearing Minor issue was found. The branch is **not ready to merge**:
this ownership defect and the two retained-capability cleanup defects above require fixes and renewed
review, independently of the remaining FAIL/PENDING release gates.
@@ -1,162 +0,0 @@
# Final-review fix round 1 report (sanitized)
## Verdict
- Base: `fa499a9bdd37011833691b0f447470d8b7e8a3a6`.
- Final frozen source: `10cd66fe6a5b484a4dc569326a228c1c5484a5d4` on
`feat/thoth-auth`.
- Authentication remediation: **PASS / ADDRESSED**. All four final-review Important findings are
resolved relative to the remediation brief.
- Terra Minor evidence corrections: **ADDRESSED**.
- Branch/release readiness: **FAIL**. The completed exact-source workflow still contains executed
baseline clone-contract, LF/Compose, and Linux Docker failures. Unavailable external/manual
gates remain **PENDING**.
- Source and evidence remain separate commits. No workflow was dispatched from the evidence-only
phase.
## Finding disposition
| Finding | Disposition | Evidence |
|---|---|---|
| Important 1 — exhaustive Windows cleanup | RESOLVED | Cleanup now attempts close/delete/validation operations in deterministic order and returns sanitized `ErrUnsafeFile` after aggregating failures. `TestWindowsPrivateRegularCleanupClosesAfterDeleteDispositionFailure` and `TestWindowsClaimCleanupAttemptsLaterOperationsAfterEarlierFailure` cover the non-short-circuit contract. Global no-delete sharing remains unchanged. |
| Important 2 — usable native Windows authority | RESOLVED | Owner-only descriptors use the current user SID, protected/non-defaulted DACL semantics, valid NT attributes/access masks, self-relative creation descriptors, and semantic full-control validation. Equal-or-stronger Windows fixture adaptations retain no-delete handles instead of weakening ACL/identity checks. The final native three-package gate passes. |
| Important 3 — restore-test deadlock | RESOLVED | Lifecycle-stage release observes the buffered worker outcome, uses a bounded/cancellable release, reports premature completion directly, and never waits indefinitely on `done`. `TestReleaseLifecycleStageReturnsPrematureWorkerOutcome` and the lifecycle-lock terminal-cleanup test are green. |
| Important 4 — complete native package gate | RESOLVED | Workflow and remediation plan both use the exact unfiltered command `go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1`. Final logs prove all three packages executed natively. |
| Minor — non-executed gate classification | RESOLVED | Non-executed/skipped commands are `NOT_RUN` / `BLOCKED`; `FAIL` is reserved for commands that ran and failed. Historical results remain separately labelled. |
| Minor — explicit Windows StageArchive row | RESOLVED | `.artifacts/task-15/automated-gates.json` contains `windows_stagearchive_retained_capability` = PASS, bound to the final source and native backup result. |
Additional failures exposed by the required unfiltered gate were fixed without narrowing the
workflow: Windows secret-bearing archive reservation is protected before use; StageArchive shares
one retained root capability across both staged files; claim/consume transitions serialize the
complete public validation and retained-handle operation while preserving ACL, hard-link identity,
reparse rejection, and no-delete invariants.
## RED → GREEN record
### Initial RED
- Run `32122302381`:
https://github.com/mptyl/ThothII/actions/runs/32122302381
- Source: `b31b27e5845ffd3adf311429367319beaba263c7`.
- Windows job: `95665197885`.
- Result: native `safeio`/`backup` failure, including the 10-minute restore lifecycle timeout;
`authstorage` was absent from the command. This established the RED for Important 2–4 and the
required native authority.
- Cleanup failure-injection tests added for Important 1 first exposed the short-circuit behavior
before the implementation was changed.
### Final concurrency RED
- Run `32140481263`:
https://github.com/mptyl/ThothII/actions/runs/32140481263
- Source: `b48e9e9189dd0e8083db9bd0378704524e670edb`.
- Windows job: `95721724645`.
- Native results: backup PASS (`20.757s`), authstorage PASS (`104.180s`), safeio FAIL
(`63.502s`). The only failures were:
- `TestCanonicalPrivateClaimWaitsForRetainedRemoveOperation`: the concurrent claim returned
`false, unsafe file` before retained removal completed;
- `TestCanonicalPrivateClaimConsumeHasOneConcurrentWinner`: iteration 8 returned `unsafe file`.
- Diagnosis: the process mutex started below `validateClaimPaths`; a concurrent caller could fail
while reopening the retained no-delete directory before reaching the lock.
### GREEN implementation and local gates
The lock boundary was moved to the three public claim/read/remove APIs, covering validation,
relative operation, and handle close. The Unix implementation uses a no-op boundary and retains its
existing descriptor-relative semantics.
Final-source local commands passed:
```text
go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1
go test -race ./...
go vet ./...
go build -o /tmp/thothii-tht-host ./cmd/tht
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go build -o /tmp/thothii-tht-windows.exe ./cmd/tht
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go test -c ... ./internal/{safeio,backup,authstorage}
```
- Focused host package times: safeio `8.750s`, backup `8.378s`, authstorage `8.854s`.
- Race suite and vet: PASS.
- Host CLI: Mach-O arm64; Windows CLI and all three Windows test binaries: PE32+ x86-64.
- Cross-compilation remains compile-only and is not used as native proof.
## Exact-source native certification
- Run: `32141428407`
- URL: https://github.com/mptyl/ThothII/actions/runs/32141428407
- Event/status/conclusion: `workflow_dispatch` / `completed` / `failure`.
- Head SHA: `10cd66fe6a5b484a4dc569326a228c1c5484a5d4` — exact final source match.
- Windows job: `Windows clone and Compose contract`, job `95724751282`:
https://github.com/mptyl/ThothII/actions/runs/32141428407/job/95724751282
- Native step: `Run native Windows retained-capability tests` — **PASS**.
- Exact command: `go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1`.
- Native package results:
- safeio PASS (`8.230s`);
- backup PASS (`5.195s`);
- authstorage PASS (`8.383s`).
- Job conclusion: `failure` only because the following `Verify Windows clone contract` baseline
step failed with a PowerShell `ParserError` at
`scripts/test-windows-clone-contract.ps1:208`; `$remoteYaml:` is not delimited before `:`.
## Remaining branch/release blockers
| Gate | Classification | Exact outcome |
|---|---|---|
| Windows native authentication packages | PASS | All three required packages executed on final source. |
| Windows clone contract | FAIL / baseline | Executed after native PASS; PowerShell parser error at line 208. |
| LF, Compose, docs, and TypeScript | FAIL / baseline CI contract | Job `95724751205`; unified Compose passed, then `test-no-deployment-coupling-scope.sh` failed because `TMPDIR` was unset. Downstream skipped commands are `NOT_RUN` / `BLOCKED`. |
| Linux Docker deployment and rollback | FAIL / infrastructure prerequisite | Job `95724751356`; executed smoke stopped because `rg` was unavailable. Cleanup proof passed; no new image manifest was generated. |
| Native Windows Docker Desktop/WSL2 startup | NOT_RUN / BLOCKED | Job `95724752028` was skipped by workflow conditions; no Docker/WSL2 command executed. |
| Harness/Ruff/other historical baseline gates | FAIL | Retained with their recorded source and results; not rewritten as final-source proof. |
| L2, real PSD/manual acceptance, provider readiness | PENDING | Required secrets, identity/access, or provider prerequisites remain unavailable. |
The historical Docker image manifest remains bound to source
`74b062f1a737103524cbe706346cfd65f87cdfd1`; it was not reused as proof for the final source.
## Principal source commits
- `cd5f505` — exhaustive cleanup, Windows authority foundation, restore deadlock tests/fix, and
complete workflow/plan package command.
- `a0e05ad` through `b6396e6` — effective full-control DACL semantics, valid NT attributes/access,
self-relative descriptors, retained no-delete fixture ordering, and Windows installation fixture
protection.
- `824245d` — preserve existing lifecycle ACL trees instead of mutating inherited authority.
- `455fffb`, `2d1670e`, `c01482c`, `9fc1a15` — concurrent claim/consume and settled-loss handling.
- `6474118` — one retained StageArchive root capability shared across staged files.
- `feee4ee` — unified Windows path wrappers on the retained primitive.
- `b261dd4` — bounded private-root sharing contention handling.
- `b48e9e9` — deterministic retained-remove concurrency regression and claim-operation lock.
- `10cd66f` — final lock boundary includes public path validation; frozen source.
## Files changed
Source changes relative to the fix-round base:
- `.github/workflows/deployment.yml`;
- `docs/superpowers/plans/2026-08-18-thothii-authentication-remediation.md`;
- `tools/tht/internal/authstorage/storage_test.go`;
- `tools/tht/internal/backup/{create.go,create_test.go,fixture_security_unix_test.go,fixture_security_windows_test.go,preflight.go,preflight_test.go,preflight_windows_test.go,restore.go,restore_test.go}`;
- `tools/tht/internal/safeio/{claim_unix.go,claim_windows.go,claim_windows_test.go,files.go,files_test.go,private_root_windows.go,private_windows.go,private_windows_test.go}`.
Evidence/status changes are restricted to:
- `.artifacts/task-15/automated-gates.json`;
- `.superpowers/sdd/2026-08-18-thothii-authentication-remediation/task-4-report.md`;
- `.superpowers/sdd/2026-08-18-thothii-authentication-remediation/fix-round-1-report.md`;
- `.superpowers/sdd/2026-08-16-thothii-authentication/task-15-report.md`;
- `PROJECT_STATE.md`.
Machine-readable evidence SHA-256:
`5c110b7b2607693de078def441b10290c5a29024c83b7e5a0ced894b72b7507f`.
## Git and protection status
- The evidence commit contains only the five evidence/status files listed above; no source is
changed after frozen source `10cd66fe6a5b484a4dc569326a228c1c5484a5d4`.
- After the evidence commit and push, the intended status is synchronized
`feat/thoth-auth...origin/feat/thoth-auth` with only protected untracked `.playwright-cli/` and
`.thothctl/`.
- `AGENTS.md`, `CLAUDE.md`, and `docs/agents/` are untouched. No generated `tools/tht/tht` exists.
- Evidence commit SHA is reported externally after commit creation because a commit cannot contain
its own final hash.
@@ -1,133 +0,0 @@
# Final-review fix round 2 report (sanitized)
## Verdict
- Base evidence head: `0f762ad6b67675356389cc546421a1c46ad5a736`.
- Frozen source: `2a9359071257f9b8a71d36ec2bbb25b161003f81` on `feat/thoth-auth`.
- Authentication remediation: **PASS**.
- Three original remediation Important findings: **RESOLVED**.
- Fix-round-2 bounded lifecycle Important: **ADDRESSED**.
- Fix-round-2 temporary Windows diagnostics Minor: **ADDRESSED**.
- Release readiness: **FAIL** for executed unrelated baseline gates, with unavailable
external/manual gates separately **PENDING**.
- Source and evidence are separate commits. The evidence-only phase changed no source or tests and
dispatched no workflow.
## Finding disposition
| Finding | Disposition | Evidence |
|---|---|---|
| Original Important — POSIX local-registry ownership | RESOLVED | Effective-UID ownership enforcement and its Node 24 coverage remain green at their recorded source. Fix round 2 did not alter this boundary. |
| Original Important — retained-capability StageArchive lifecycle | RESOLVED | Native Windows `internal/backup` passed on the exact source, preserving the retained-root staging and cleanup coverage. |
| Original Important — handle-relative Windows claim removal | RESOLVED | Native Windows `internal/safeio` and `internal/authstorage` passed on the exact source, including retained claim/consume coverage. |
| Fix-round-2 Important — fully bounded restore lifecycle test | ADDRESSED | Gate publication and release are context-aware; stage, outcome, admission, checkpoint, and verification waits are bounded; aborts cancel, safely release, bounded-join, then assert lock-free. The deterministic withheld-gate test proves prompt timeout/cancellation, worker join, and eventual lock release. |
| Fix-round-2 Minor — temporary Windows diagnostic matrix | ADDRESSED | `windowsRelativeOpenMatrix` and its diagnostic-only call/import were removed. Owner-only DACL shape, NT access normalization, full-control, cleanup, and retained no-delete tests remain. |
The round-1 restore lifecycle finding was broadened by the scoped round-2 review: bounded release
alone was insufficient while stage publication, gate waits, and nearby outcome/admission waits
could still outlive a controller abort. The round-2 implementation closes that broader test
orchestration gap without changing production authentication semantics.
## RED → GREEN record
### RED
The deterministic withheld-gate regression was introduced first and run without relying on a
global ten-minute package timeout:
```text
go test ./internal/backup -run '^TestRestoreLifecycleCancellationJoinsWithWithheldGate$' -count=1
```
It failed in approximately `0.64s` with:
```text
cancelled restore worker did not join within the bounded deadline
```
This proved that cancellation did not yet unblock and join a worker retained at the lifecycle
gate.
### GREEN and refactor
- The gate uses a cancellation source shared by controller and worker. Both publication and
release are `select`-based and cancellation-aware.
- Shared bounded helpers cover stage, outcome, error, signal, release, and admission waits.
- Abort cleanup is ordered: cancel, cancel the controller gate when distinct, safely release a
pending gate, bounded-join the worker, then prove the lifecycle lock is free.
- Premature worker outcomes retain and surface their original error.
- The existing success, recovery, maintenance-barrier, stale-checkpoint, and verification
assertions remain active.
Final local gates on the frozen source:
```text
go test ./internal/backup -run '^(TestRestoreLifecycleCancellationJoinsWithWithheldGate|TestReleaseLifecycleStage|TestRestoreLifecycleLockExcludesCompetingTransactionsUntilTerminalCleanup|TestRestoreCannotApplyAStaleCheckpointOverAnInterleavedRestore|TestRestoreKeepsAdmissionBarrierActiveUntilVerificationCommits)$' -count=1
go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1
go test ./... -count=1
go test -race ./...
go vet ./...
go build -o /tmp/thothii-tht-host-fix-round-2 ./cmd/tht
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go test -c ./internal/safeio -o /tmp/tht-safeio-fix-round-2-windows.test.exe
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go test -c ./internal/backup -o /tmp/tht-backup-fix-round-2-windows.test.exe
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go test -c ./internal/authstorage -o /tmp/tht-authstorage-fix-round-2-windows.test.exe
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go build -o /tmp/thothii-tht-fix-round-2-windows.exe ./cmd/tht
```
All commands passed. The final focused lifecycle run completed in `0.672s`; the full security
package run passed safeio, backup, and authstorage; race, vet, host build, Windows test-package
cross-compiles, and Windows CLI cross-compile also passed. Cross-compilation is recorded only as
compile evidence and is not used as native authority.
## Exact-source native certification
- Controller-authorized run: `32147345625` —
https://github.com/mptyl/ThothII/actions/runs/32147345625.
- Event/status/conclusion: `workflow_dispatch` / `completed` / `failure`.
- Head SHA: `2a9359071257f9b8a71d36ec2bbb25b161003f81`, exactly matching the frozen source.
- Windows job: `Windows clone and Compose contract`, job `95744249248` —
https://github.com/mptyl/ThothII/actions/runs/32147345625/job/95744249248.
- Native step: `Run native Windows retained-capability tests` — **PASS**.
- Exact unfiltered command:
`go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1`.
- Native package results:
- `internal/safeio` PASS (`22.058s`);
- `internal/backup` PASS (`7.161s`);
- `internal/authstorage` PASS (`16.088s`).
The Windows job failed only in the following baseline clone-contract step. PowerShell reported a
parser error at `scripts/test-windows-clone-contract.ps1:208` because `$remoteYaml:` is not a
delimited variable reference. This later failure does not alter the successful native Go step.
## Separate release-readiness verdict
| Gate | Classification | Exact outcome |
|---|---|---|
| Authentication remediation | PASS | Source and exact-source native three-package authority are green. |
| Windows clone contract | FAIL / baseline | Job `95744249248`; parser error at `scripts/test-windows-clone-contract.ps1:208`, after native PASS. |
| LF, Compose, docs, and TypeScript | FAIL / baseline CI contract | Job `95744249458`; unified Compose passed, then the existing unset-`TMPDIR` failure stopped the contract step. Downstream commands were skipped. |
| Linux Docker deployment and rollback | FAIL / infrastructure prerequisite | Job `95744249354`; the existing missing-`rg` prerequisite stopped the smoke before deployment. Cleanup passed and no new image manifest was generated. |
| Native Windows Docker Desktop/WSL2 startup | NOT_RUN / BLOCKED | Job `95744250450` was skipped by workflow conditions; no native Docker/WSL2 command ran. |
| L2, real PSD/manual acceptance, provider readiness | PENDING | Required secrets, identity/access, or provider prerequisites remain unavailable. |
Executed failures remain `FAIL`; skipped commands are `NOT_RUN` / `BLOCKED`; unavailable external
gates remain `PENDING`. Therefore remediation PASS does not imply release readiness PASS.
## Evidence and protection status
- Machine-readable evidence: `.artifacts/task-15/automated-gates.json`; SHA-256
`6c516db5c2064c4a4a2e5f25961b993cd4a8fe020bbbb822fbac7faa0c119599`.
- Current Task 4 report:
`.superpowers/sdd/2026-08-18-thothii-authentication-remediation/task-4-report.md`.
- Retained Task 15 report:
`.superpowers/sdd/2026-08-16-thothii-authentication/task-15-report.md`.
- Project snapshot: `PROJECT_STATE.md`.
- Historical Docker evidence remains bound to its recorded older source and is not reused as proof
for `2a9359071257f9b8a71d36ec2bbb25b161003f81`.
- `.playwright-cli/` and `.thothctl/` remain protected and untracked. No source/test file,
instruction file, workflow, or `docs/agents/` content changed in this evidence phase.
- The separate evidence commit SHA is reported after commit creation because a commit cannot
contain its own final hash.
No credentials, tokens, internal endpoints, identities, registry names, raw environments, or
browser traces are retained in this report.
@@ -1,131 +0,0 @@
# Task 4 authentication remediation recertification (sanitized)
## Fix-round-2 recertification — remediation PASS
- Exact source: `2a9359071257f9b8a71d36ec2bbb25b161003f81` on `feat/thoth-auth`.
- Authorized workflow: completed run `32147345625`,
https://github.com/mptyl/ThothII/actions/runs/32147345625, exact matching head SHA.
- Native job: `Windows clone and Compose contract`, job `95744249248`.
- Required native step: `Run native Windows retained-capability tests` — **PASS**.
- Exact unfiltered command:
`go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1`.
- Package evidence: `internal/safeio` PASS (`22.058s`), `internal/backup` PASS (`7.161s`),
`internal/authstorage` PASS (`16.088s`). This includes explicit native Windows
StageArchive retained-capability and concurrent claim-consume coverage.
- The later `Verify Windows clone contract` step failed independently at
`scripts/test-windows-clone-contract.ps1:208`: PowerShell parsed `$remoteYaml:` as an invalid
variable reference. This baseline deployment-contract failure does not change the native Go
package result.
- The optional `Native Windows Docker Desktop/WSL2 startup` job was skipped by workflow
conditions. It is `NOT_RUN` / `BLOCKED`, because no Docker Desktop/WSL2 command executed.
- The workflow reached `completed` with conclusion `failure`: the native authentication step is
PASS, while the later clone-contract, LF/Compose, and Linux Docker baseline steps are FAIL.
- Existing LF/Compose job `95744249458` and Linux Docker job `95744249354` failures repeated
before downstream work. Skipped commands are `NOT_RUN` / `BLOCKED`, not executed failures.
External L2/PSD/provider gates remain `PENDING`.
Finding disposition is explicit: the three original remediation Important findings remain
**RESOLVED**; the fix-round-2 lifecycle Important is **ADDRESSED**; and the temporary Windows
diagnostic-matrix Minor is **ADDRESSED**. Authentication remediation is **PASS**. This does not
change overall release readiness: executed baseline gates remain **FAIL**, while unavailable
external/manual gates remain **PENDING**.
The section below is retained as historical evidence for the pre-fix frozen source.
## Historical pre-fix result
- Frozen source under test: `b31b27e5845ffd3adf311429367319beaba263c7` on `feat/thoth-auth`.
- Freeze check: PASS. No tracked source changed during certification. The only untracked paths
retained are `.playwright-cli/` and `.thothctl/`.
- Certification window: `2026-08-18T09:26Z` to `2026-08-18T09:48:36Z` (UTC; the start marker is
minute-precision because no earlier second-level operator timestamp was captured).
- Overall result: `FAIL` / `CHANGES_REQUIRED`. The three Important findings are not closed and
authentication is not implementation-complete or release-complete.
## Local gate matrix
| Gate | Result | Sanitized evidence |
|---|---|---|
| Go focused security tests | PASS | `safeio`, `backup`, and `authstorage`; 3 packages |
| Go race/vet/host build | PASS | 18 race-tested packages; vet and host CLI build exit 0 |
| Windows amd64 cross-compile | PASS | focused safeio/backup test binaries and CLI build; compile-only |
| POSIX registry ownership | PASS | Node 24 backend suite includes local-registry ownership coverage |
| Unix StageArchive retained capability | PASS | focused safeio/backup and race coverage passed on host |
| Backend Node 24 | PASS | 76 files / 1092 tests; typecheck and build passed |
| Frontend Node 24 | PASS | 61 files / 444 tests; typecheck and build passed |
| Authentication/F1 browser smoke | PASS | Node `v24.16.0`; filtered E2E 1 passed; sentinel scan passed |
| Harness pytest | FAIL | 951 passed, 1 failed, 4 skipped, 232 subtests; `test_f4_emits_column_types` could not find `workflow.yaml` from its test cwd |
| Ruff | FAIL | 192 errors; known baseline |
| Authentication docs smoke | PASS | required-term and forbidden-word checks passed |
| Shell syntax | PASS | `bash -n scripts/*.sh` |
| Default Compose contract | FAIL | required `THT_WORKSPACE_GIT_REMOTE` was unavailable |
| Unified Compose contract | FAIL | `compose.unified.yaml` is absent from the frozen source |
| Unified Docker smoke | FAIL | workflow attempted it on the frozen SHA but stopped before deployment because `rg` was unavailable; cleanup proof passed and no new image manifest was generated |
| L2 / PSD manual / provider readiness | PENDING | required external secrets, identities/access, or provider prerequisites unavailable/not reached |
The first full backend Vitest attempt had one workspace-registry timeout. The focused test and a
fresh complete rerun passed, so the current backend result above is the fresh complete rerun.
## Native Windows authority
The authorized dispatch was bound to the frozen SHA:
- Run: `32122302381`
- URL: https://github.com/mptyl/ThothII/actions/runs/32122302381
- Head SHA: `b31b27e5845ffd3adf311429367319beaba263c7`
- Workflow conclusion: `failure`
- Job: `Windows clone and Compose contract`, job `95665197885`
- Job URL: https://github.com/mptyl/ThothII/actions/runs/32122302381/job/95665197885
- Native step: `Run native Windows retained-capability tests` — `failure`
- Executed command: `go test ./internal/safeio ./internal/backup -count=1`
- Observed focused failures include `TestRemoveCanonicalPrivateClaimRetainsParentDuringDeletion`
and `TestRemoveCanonicalPrivateClaimPreservesOrphan`.
- The backup package timed out in
`TestRestoreLifecycleLockExcludesCompetingTransactionsUntilTerminalCleanup` after `10m0s`.
- Additional backup failures included retained-staging `unsafe file` results, Windows temporary-file
cleanup reporting that a file was still in use, and fixture cases that could not read external
secret declarations. The first two categories are remediation/security-boundary failures; the
fixture declaration failures are recorded as an accompanying CI-fixture issue.
- `internal/authstorage` was not requested by the frozen workflow step and therefore has no native
Windows execution evidence. Cross-compilation does not substitute for this gate.
This native failure is the blocking gate. No source fix was attempted, and no later Docker smoke
was run locally after the failure.
## Other workflow failures
- `LF, Compose, docs, and TypeScript` (job `95665197839`) failed in
`Verify Compose and installation contracts` after the unified Compose contract itself passed.
`test-no-deployment-coupling-scope.sh` aborted on `TMPDIR: unbound variable`; this is classified
as a baseline/CI contract prerequisite, and later docs/TypeScript steps were skipped.
- `Linux Docker deployment and rollback` (job `95665197846`) failed before deployment because the
runner did not provide `rg` (`Task 13 smoke failed: rg is required`). The sanitized cleanup proof
passed and no Docker image manifest was generated. This is classified as an infrastructure
prerequisite failure, not as evidence of a remediation regression.
## Evidence and provenance
- Current machine-readable matrix: `.artifacts/task-15/automated-gates.json`; SHA-256
`6c516db5c2064c4a4a2e5f25961b993cd4a8fe020bbbb822fbac7faa0c119599`.
- Current requested report: this file (SHA-256 recorded after the evidence commit if needed for
external indexing).
- Current fix-round report:
`.superpowers/sdd/2026-08-18-thothii-authentication-remediation/fix-round-2-report.md`.
- Historical Docker image manifest: `.artifacts/task-15/unified-docker-images.json`, unchanged
because no new immutable-source Docker smoke ran. Its retained historical SHA-256 is
`9c8dec4546909fd93799dbcf374bcb3a89bc46cfe0fd482472c0cbe757ddf5b6`, bound to historical source
`74b062f1a737103524cbe706346cfd65f87cdfd1`, not to this Task 4 candidate.
- The historical Task 15 report remains provenance for earlier source SHAs; its current addendum
records this recertification separately.
No credentials, tokens, internal endpoints, provider identities, registry names, raw environments,
or browser traces are retained here.
## Separate verdicts
- Three Important findings: `CHANGES_REQUIRED`. Native Windows retained-capability authority
failed, and the frozen workflow omits the required `authstorage` package from its native command.
- Overall release readiness: `FAIL` with additional `PENDING` gates. The native Windows remediation
gate failed; the remote Docker attempt failed on a missing runner prerequisite; existing
Ruff/harness/Compose failures and external/manual prerequisites remain unresolved; and no
successful new unified Docker image evidence exists.
@@ -1,137 +0,0 @@
# Adapter Foundations final-review fix report
Date: 2026-07-11
Branch: `codex/portable-deployment`
Worktree: `/Users/mp/projects/ThothII/.worktrees/portable-deployment`
Binding findings: `.superpowers/sdd/adapter-final-review-findings.md`
## Outcome
All seven final-review findings are addressed as one coherent adapter-foundations change:
1. HTTP vector reader and writer clients are independently optional. Capabilities reflect the
configured side; writer-only new and legacy configurations build successfully for targeted
writes; search without a reader raises public `VectorReadUnavailable`.
2. `VectorHealth` now reports read/write configured and reachable state independently, preserves
side-specific errors, and reports expected/observed embedding dimensions plus compatibility.
HTTP diagnostics cover read-only, write-only, both-up, and writer-down cases. Direct health
exposes its configured expected dimension without adding schema or migration work.
3. `ThothRestDwhAdapter` accepts `DatabaseIdentityConfig`, matching its resource contract.
4. Both vector adapters reject bools, floats, zero, and negative search limits using one exact
positive-integer guard.
5. Port tests explicitly cover public exports and frozen capability records.
6. A real `tht` subprocess test proves one legacy deprecation warning per config load on stderr
while JSON stdout remains parseable and uncontaminated.
7. The adapter plan and SDD progress explicitly constrain `build_vector_loader` to transitional
bulk sync and schedule its removal/migration in the local pgvector plan. Targeted memory and
solved-question writes remain on `build_vector_store(..., require_write=True)`.
No pgvector schema or migration changes were made.
## Files changed
- `harness/tht/ports/vector.py`
- `harness/tht/ports/__init__.py`
- `harness/tht/adapters/vector/thoth_http.py`
- `harness/tht/adapters/vector/legacy_direct.py`
- `harness/tht/adapters/factory.py`
- `harness/tht/adapters/dwh/thoth_rest.py`
- `harness/tests/test_vector_port_contract.py`
- `harness/tests/test_adapter_factory.py`
- `harness/tests/test_config_resources.py`
- `harness/tests/test_config_legacy_compat.py`
- `harness/tests/test_adapter_command_regressions.py`
- `harness/tests/test_dwh_port_contract.py`
- `docs/superpowers/plans/2026-07-11-adapter-foundations.md`
- `.superpowers/sdd/progress.md`
- `.superpowers/sdd/adapter-final-fix-report.md`
## TDD and verification evidence
RED:
```text
cd harness && .venv/bin/pytest tests/test_vector_port_contract.py \
tests/test_adapter_factory.py tests/test_config_resources.py \
tests/test_config_legacy_compat.py -q
```
Result: collection failed as expected because `VectorReadUnavailable` did not exist. After the
initial implementation, the same command exposed two expected contract/test-harness corrections:
dimension mismatch makes aggregate health unhealthy, and the installed CLI entry point is `tht`
rather than `python -m tht.cli`.
GREEN, covering adapter/config/command regressions:
```text
cd harness && .venv/bin/pytest tests/test_vector_port_contract.py \
tests/test_adapter_factory.py tests/test_config_resources.py \
tests/test_config_legacy_compat.py tests/test_adapter_command_regressions.py \
tests/test_dwh_port_contract.py tests/test_memory_save_one.py \
tests/test_solved_question.py tests/test_search_similar_kinds.py \
tests/test_vector_dual_key.py -q
```
Result: `66 passed in 0.45s`.
Docker availability:
```text
docker info --format '{{.ServerVersion}}'
```
Result: `29.4.1` (available; command required Docker socket access).
Full repository-default non-L2 harness suite, with Docker available for L0 tests:
```text
cd harness && .venv/bin/pytest -q
```
Result: `433 passed, 5 deselected, 17 warnings in 9.14s`. The five deselections are the configured
L2/live-service tests. Warnings are existing legacy-workspace `FutureWarning` emissions.
Scoped lint and diff hygiene:
```text
cd harness && .venv/bin/ruff check tht/ports tht/adapters \
tests/test_vector_port_contract.py tests/test_adapter_factory.py \
tests/test_config_resources.py tests/test_config_legacy_compat.py \
tests/test_adapter_command_regressions.py tests/test_dwh_port_contract.py
git diff --check
```
Result: `All checks passed!`; `git diff --check` produced no output.
## Commit
Commit subject: `fix(adapter): close final foundation review`
The report is part of that same final commit. A Git object cannot contain its own SHA without
changing that SHA; the exact resulting commit ID is therefore recorded in the task handoff from
`git rev-parse HEAD` after creation.
## Self-review
- Reader/writer separation is preserved: search dereferences only `_reader`; hashes/upsert only
`_writer`; health probes each configured client independently and never substitutes one result
for the other.
- Writer failure contributes to aggregate `ok=False`, even when the reader succeeds.
- Dimension compatibility is derived only from configured embedding dimension and existing
`list_tables` metadata. Missing metadata remains `None`, not a guessed success/failure.
- The shared limit guard uses `type(limit) is int`, intentionally rejecting Python booleans and
numeric coercions before either adapter reaches its transport.
- Existing JSON/CLI behavior is preserved; the subprocess regression parses stdout as JSON and
counts exactly one deprecation marker on stderr.
- Scope remains adapter foundations. No vector DDL, schema initialization, or migration work was
introduced.
## Concerns / follow-up
- Write reachability uses the existing `list_tables` diagnostic on the separately authenticated
writer client. Deployments must allow that non-mutating diagnostic RPC to the writer credential;
failures are intentionally visible rather than hidden by reader success.
- Existing legacy-workspace tests emit 17 `FutureWarning`s in the full suite. This wave pins the
required production stderr behavior but does not migrate unrelated test fixtures.
- `build_vector_loader` remains transitional technical debt only for bulk sync, explicitly assigned
to `2026-07-11-local-pgvector-profile.md`.
-206
View File
@@ -1,206 +0,0 @@
# Container Packaging Task 3 Report
## Status
Implemented the multi-stage core application image, non-root runtime, pinned Pi installation,
container entrypoint, context exclusions, and an in-image health smoke test.
## TDD / Build Evidence
Initial RED:
```text
docker build -f docker/core.Dockerfile -t thothii-core:test .
ERROR: failed to build: resolve : lstat docker: no such file or directory
```
The first sandboxed attempt could not access the Docker socket; the authorized rerun reached the
builder and failed for the expected reason: the Dockerfile did not exist.
GREEN build:
```text
sh -n docker/core-entrypoint.sh docker/smoke/core-smoke.sh
docker build --progress=plain -f docker/core.Dockerfile -t thothii-core:test .
```
Result: shell syntax exited 0; Docker build exited 0. A final rebuild after tightening
`.dockerignore` also exited 0 and transferred only 17.60 kB of changed context (the initial clean
build transferred 1.02 MB).
## Runtime and Entrypoints
- Runtime user is `10001:10001` (`thoth`), never root.
- Runtime contains Node `v22.19.0` and Python `3.12.13`. Python 3.12 is intentional because the
harness declares `requires-python = ">=3.12"` and also satisfies the deployment floor of 3.11+.
- Pi is installed exactly as `@earendil-works/pi-coding-agent@0.80.3`; its build-time and runtime
version probes both reported `0.80.3`.
- `server` starts `/app/backend/dist/server.js`; `doctor` routes to `tht doctor`; `preprocess`
routes to the future-facing `tht preprocess` command; explicit `tht ...` and arbitrary CLI
arguments route to the installed `tht` binary.
- The gate extension's `typebox` runtime dependency is installed from the harness lockfile.
## Smoke and Diagnostic Results
```text
docker run --rm thothii-core:test doctor
config: error - configuration is invalid or unreadable
data_root: ok
```
Result: expected exit 1 for absent mounted workspace configuration, with no traceback and no
secret-bearing validation detail.
```text
docker run --rm --entrypoint /app/docker/smoke/core-smoke.sh thothii-core:test
backend listening on http://127.0.0.1:8787
v22.19.0
Python 3.12.13
core smoke: ok
```
Result: exit 0. The script asserted non-root execution, `tht --help`, `pi --version`, runtime
version floors, and `GET /health` through curl. Fastify's returned display address was loopback;
the inspected container environment is `HOST=0.0.0.0`, and the compiled server passes that value
to `app.listen`.
```text
docker run --rm thothii-core:test tht --version
0.1.0
```
Result: arbitrary `tht` entrypoint exited 0.
An explicit runtime assertion checked UID 10001, exact Node and Pi versions, Python 3.11+, and the
absence of `/app/harness/.env` and `/app/harness/workspaces`; it exited 0.
## Image Size and Containment Inspection
```text
docker image inspect thothii-core:test --format '{{.Size}} {{json .Config.User}} {{json .Config.Env}}'
221419008 "10001:10001" [...runtime paths and version metadata only...]
```
Image size: **221,419,008 bytes** (about 211.2 MiB).
`docker history --no-trunc thothii-core:test` was inspected. It contains only Dockerfile commands,
the pinned public package name/version, base-image metadata, and non-sensitive runtime variables;
no credentials or customer paths were found. An in-image filename scan found only
`/app/harness/.pi/settings.json` among `.env`, key/certificate, and settings-name candidates; that
tracked Pi file contains theme/startup preferences, not secrets. The build asserts `.env` and
workspace directories are absent.
`.dockerignore` excludes VCS/agent state, all environment files except examples, package-manager
credential files, SSH/private-key and certificate formats, local virtualenvs/node_modules/caches,
backend runtime data, customer workspaces, sessions, artifacts, indexes, corpus, and deployment
mount content.
## Self-review
- `git diff --check` is clean.
- Entrypoint processes use `exec`, preserving container signal handling.
- Backend production dependencies are pruned; TypeScript build tools remain in the build stage.
- The writable `/data` root is owned by UID 10001; application payload remains root-owned and
read-only to the runtime user.
- CA certificates and curl are present for HTTPS integrations and health probing.
- No existing source, customer workspace, secret, or unrelated progress-ledger change is included
in the task commit.
## Concerns
- The `tht preprocess` command is deliberately a future-facing routing contract; its CLI group is
scheduled in the Evidence/preprocessing plan and is not implemented in the current harness.
- Python dependencies are range-resolved because the existing harness has no Python lockfile. The
Pi package, Node runtime, and package-lock-backed Node dependency sets are pinned/reproducible.
- The image was built and smoked on Docker Desktop arm64. The chosen official multi-arch base
images and Pi package are architecture-neutral at the package level, but amd64 still needs a CI
build/smoke before being advertised as verified.
## Reproducibility Review Fix
The original image pinned Pi's direct version in the Dockerfile but resolved its transitives at
build time, and pip resolved all harness dependencies from ranges. Both paths now consume committed
locks.
### Lock generation
Pi uses the minimal `docker/pi-runtime/package.json` and its committed npm v3 lock. It was generated
with:
```text
npm install --package-lock-only --ignore-scripts --no-audit --no-fund \
--prefix docker/pi-runtime
```
The package manifest specifies exact `@earendil-works/pi-coding-agent` version `0.80.3`; a lock
inspection confirmed that same resolved package version. Docker installs it with:
```text
npm ci --omit=dev --ignore-scripts --no-audit --no-fund
```
The Python lock was generated directly from the harness production metadata plus one explicit,
pinned PEP 517 build-backend input—not from a host `pip freeze`:
```text
uv pip compile harness/pyproject.toml docker/python-runtime/build-requirements.in \
--universal \
--python-version 3.12 \
--no-emit-package tht \
--generate-hashes \
--custom-compile-command \
'uv pip compile harness/pyproject.toml docker/python-runtime/build-requirements.in --universal --python-version 3.12 --no-emit-package tht --generate-hashes --output-file docker/python-runtime/requirements.lock' \
--output-file docker/python-runtime/requirements.lock
```
`pytest`, `ruff`, and `testcontainers` are absent. All production direct and transitive packages
are exact and hashed. `setuptools==80.9.0` is explicit so the local harness install can use
`--no-build-isolation` without an unpinned build-time resolution. Refresh instructions are in
`docker/LOCKS.md`.
### No-cache rebuild and verification
Final build command:
```text
docker build --no-cache -f docker/core.Dockerfile -t thothii-core:test .
```
Result: exit 0. The logs showed Pi `0.80.3`, Node `v22.19.0`, a hash-enforced Python dependency
install, explicit `setuptools==80.9.0`, and a non-isolated local `tht` wheel build. No isolated
build-dependency download occurred.
Fresh runtime checks:
```text
docker run --rm --entrypoint /app/docker/smoke/core-smoke.sh thothii-core:test
backend listening on http://127.0.0.1:8787
v22.19.0
Python 3.12.13
core smoke: ok
docker run --rm thothii-core:test tht --version
0.1.0
/opt/venv/bin/pip check
No broken requirements found.
```
An in-container package inspection reconfirmed Pi `0.80.3`. Non-root UID, runtime version floors,
doctor's expected concise exit 1/no traceback, `/health`, and arbitrary `tht` routing all passed.
The full filename containment scan found no `.env`, PEM, private-key, P12, or PFX file in `/app`;
`/app/harness/workspaces` remains absent. Image environment and `docker history --no-trunc` were
re-inspected and contain only public package/build commands and non-sensitive runtime metadata.
Final locked image size:
```text
220003986 10001:10001
```
That is **220,003,986 bytes** (about 209.8 MiB), 1,415,022 bytes smaller than the original image.
Remaining concern: the universal lock is resolved for Python 3.12 and includes hashes/markers for
all supported platforms, but only Linux arm64 has been built and smoked locally; amd64 remains a CI
verification gate.
@@ -1,82 +0,0 @@
# Container Packaging Task 4 Report
## Status
Implemented and verified runtime-configured frontend packaging.
## Changes
- Added the browser runtime contract `window.__THOTHII_CONFIG__.backendBaseUrl`.
- Loaded `/config.js` before the Vite module entrypoint.
- Made runtime configuration take precedence while preserving `VITE_BACKEND_URL` and the
existing `http://localhost:8787` client default for development and tests.
- Added a multi-stage frontend image that builds with Node and serves static assets as
unprivileged UID/GID `101:101` with nginx on port 8080.
- Added startup-time `BACKEND_BASE_URL` substitution (default `/api`).
- Added `/api/` reverse proxying to `core:8787`, SPA fallback, no-cache runtime config,
and SSE-safe proxy settings (`proxy_buffering off`, `proxy_cache off`, one-hour read timeout).
## TDD evidence
- RED: `npx vitest run src/api/runtime-config.test.ts` failed because
`./runtime-config` did not exist.
- GREEN: targeted runtime config suite passed (3 tests after preserving the legacy client
default).
## Verification
- `cd frontend && npx vitest run --reporter=dot && npx tsc -b && npm run build` — exit 0
(40 test files, 185 tests; TypeScript and Vite production build passed).
- `docker build -f docker/frontend.Dockerfile -t thothii-frontend:test .` — success.
- Image metadata reports `USER 101:101`.
- Two-container isolated-network smoke:
- `/config.js` returned `window.__THOTHII_CONFIG__ = { backendBaseUrl: "/api" };`
- `/api/health` proxied to the core image and returned `{"status":"ok"}`.
- an unknown nested route returned the SPA `index.html`.
- active nginx config contained `proxy_buffering off`, `proxy_cache off`, and
`proxy_read_timeout 1h`.
- `/config.js` returned `Cache-Control: no-store`.
- `sh -n docker/frontend-entrypoint.sh` and `git diff --check` — exit 0.
## Secret-leakage inspection
- `.dockerignore` excludes `.env*` (except examples), credentials/key formats, dependency
trees, build outputs, backend data, and deployment data.
- The runtime web root contained no `.env*`, `.pem`, `.key`, `.p12`, or `.pfx` files.
- Image history contained build/package instructions only; no secret build arguments or
credential values were introduced by this task.
## Self-review / concerns
- nginx resolves the `core` hostname at startup, matching the planned Compose service name;
standalone runs therefore need a reachable network alias named `core`.
- Existing frontend test warnings (React refs/act, MSW unmatched incidental requests, Vite
chunk-size warnings) remain; they did not fail the requested gates and are unrelated to
this task.
- `.superpowers/sdd/progress.md` was already modified by the orchestrator and was intentionally
excluded from this task's commit.
## P1 review fixes
Follow-up commit work addressed both review findings:
- Runtime configuration is now produced with `jq -cn --arg`, so `BACKEND_BASE_URL` is encoded
by a real JSON serializer rather than interpolated into JavaScript by `sed`.
- The image includes `frontend-config-smoke`, which strips only the fixed assignment wrapper,
parses the remaining JSON with `jq`, requires exactly the `backendBaseUrl` key, and compares
the decoded value to the environment input.
- The hostile smoke passed with quotes, backslashes, a literal newline, ampersand, pipe, and
`"; globalThis.PWNED=true; //` in the value. A breakout would leave non-JSON trailing input
and fail parsing.
- Added `joinBackendPath`, shared by API fetch and EventSource creation. It removes duplicate
boundary slashes for relative and absolute bases while keeping empty and `/` bases rooted.
Follow-up verification:
- RED: six join cases failed with `joinBackendPath is not a function` before implementation.
- Targeted: runtime config, API client, and EventSource suites — 14 tests passed.
- Full frontend gate — exit 0 (40 test files, 191 tests, TypeScript, Vite build).
- Rebuilt `thothii-frontend:test` successfully.
- Hostile config image smoke — `frontend runtime config smoke: ok`.
- Rebuilt two-container smoke — default `/api` config, proxied `/api/health`, SPA fallback,
and SSE-safe nginx directives all passed.
@@ -1,98 +0,0 @@
# Evidence / Preprocessing Task 1 Report
## Outcome
Implemented the additive Evidence source port and canonical corpus records. Existing evidence,
search, vector, and session runtime code is unchanged.
## Contract
- `EvidenceSource` is a runtime-checkable protocol with `discover` and `acquire` operations.
- `SourceObject` and `AcquiredDocument` are frozen, reject extra fields, use independent metadata
defaults, and restrict metadata to Pydantic `JsonValue` values.
- `CanonicalDocument`, `CanonicalChunk`, and `CorpusManifest` are frozen and reject extra fields.
- Provenance includes stable source IDs, canonical URIs, fingerprints, modification time, and
content hashes.
- Pipeline versions are recorded on documents, chunks, and manifests. Manifests also carry schema
version, optional publish ID/vector generation, and paired embedding model/dimension fields.
- Credential-like metadata keys are rejected recursively. Credentials are not model fields and
therefore cannot enter serialized canonical artifacts through extras.
## TDD evidence
The initial focused run failed during collection because `tht.ports.evidence` and `tht.corpus`
did not exist. After implementation, the focused suite passed.
## Verification
- Focused models/protocol tests: 13 passed.
- Harness excluding Docker-backed L0 and the network-dependent wheel packaging test: 444 passed,
5 deselected.
- Focused Ruff: passed.
- Full-repository Ruff remains blocked by 34 pre-existing findings outside the task files.
- An unrestricted `pytest -q` attempt reached 453 passed and 5 deselected, but reported 47 Docker
setup errors plus 4 Docker parity failures because the sandbox cannot access the Docker socket;
the wheel packaging test also failed because its isolated `uv build` needs unavailable network.
## Concerns / follow-up
- Pydantic's `frozen=True` prevents model field reassignment but does not recursively freeze list
and dict contents. `default_factory` prevents shared mutable defaults. Later pipeline stages should
treat these value objects as immutable and construct replacements rather than mutate collections.
- The adapter and normalization tasks should preserve the credential-free boundary by passing only
these records beyond acquisition.
## Review hardening follow-up
All six binding review areas were addressed in a separate TDD pass:
- JSON metadata is recursively converted to immutable `FrozenDict`/tuple values while retaining
stable object/array JSON serialization. Manifest document and chunk collections are tuples.
- Secret-key matching now normalizes camelCase and punctuation. It rejects credential-specific
names (passwords, API keys, access/refresh tokens, client/private keys, session cookies and
authorization) recursively, while deliberate benign labels such as generic `token` and `secret`
remain valid.
- Canonical URIs require a scheme and reject userinfo or credential-bearing query parameters.
- Namespaced IDs, SHA-256 content hashes, timezone-aware UTC timestamps, embedding/vector
compatibility, unique IDs, chunk referential/provenance integrity, contiguous per-document
ordinals and pipeline-version consistency are validated. Nested Pydantic instances are always
revalidated so `model_copy(update=...)` cannot bypass a manifest boundary.
- Acquired arbitrary bytes have explicit base64 JSON encoding and validation, covered by a JSON
round-trip test.
- `EvidenceSourceError` classifies transient/retryable versus permanent failures and exposes only
recursively immutable, credential-screened JSON details.
Follow-up verification:
- Focused contract suite: 39 passed.
- Focused Ruff: passed.
- Harness excluding Docker-backed L0 and the network-dependent wheel packaging test: 470 passed,
5 deselected.
- Fresh unrestricted harness attempt: 479 passed, 5 deselected; the same environmental boundary
remains (47 Docker socket setup errors, four Docker parity failures, one isolated `uv build`
network failure).
## Final blocker follow-up
The remaining four contract blockers were closed in a third TDD cycle:
- `EvidenceSourceError` now always exposes the fixed public message/`args` value `evidence source
operation failed`; caller diagnostics are not retained. Category, details and args cannot be
reassigned, details remain recursively frozen and credential-screened, and an original exception
is available only when callers use standard exception chaining.
- Canonical document/chunk provenance stores only URI scheme, authority and path. Userinfo is
rejected; query strings and fragments are removed unconditionally, including AWS `X-Amz-*`, SAS
`sig`, and fragment token material.
- Binding model bases override Pydantic's unchecked `model_copy(update=...)`: merged values always
pass full field/model validation, so invalid copied records and top-level manifests fail.
- A canonical document/chunk `content_hash` must equal SHA-256 of the exact stored text encoded as
UTF-8. This establishes the normalization boundary explicitly: line-ending/frontmatter/text
normalization happens before model construction; the canonical models never rewrite content.
Final follow-up verification:
- Focused contract suite: 45 passed.
- Focused Ruff: passed.
- Harness excluding Docker-backed L0 and network-dependent packaging: 476 passed, 5 deselected.
- Fresh unrestricted harness attempt: 486 passed, 5 deselected, with the unchanged environmental
failures (47 Docker setup errors, four Docker parity failures, one isolated `uv build` failure).
@@ -1,65 +0,0 @@
# Evidence Task 2 Report
## Status
Implemented filesystem and explicit-manifest HTTP Evidence source adapters, typed source
configuration with legacy compatibility, and factory construction.
## Delivered behavior
- Filesystem discovery is deterministic and rooted at a strict canonical directory.
- Symlink/path escapes are rejected before content is exposed.
- Discovery hashing and acquisition reads enforce a configurable byte limit.
- Filesystem fingerprints are content SHA-256 values; stable IDs derive from relative paths.
- HTTP accepts only explicit `http`/`https` manifest entries and keeps transport URLs private.
- HTTP provenance strips query strings/fragments, while config and adapter representations hide
signed or secret-bearing transport URLs.
- HTTP acquisition uses separate connect/read timeouts, streaming byte limits, bounded redirects,
private redirect rejection, and safe transient/permanent error classification.
- HTTP fingerprints prefer a deterministic ETag digest, then Last-Modified, then content SHA-256.
- `build_evidence_sources(cfg)` supports both typed `evidence.sources` entries and the legacy
`source_root` plus `evidence_dir` filesystem configuration.
## TDD and verification
- RED: focused tests initially failed during collection because the adapter package did not exist.
- GREEN: `15 passed` for filesystem, HTTP, and resource-config tests.
- Full harness: `548 passed, 5 deselected`.
- Changed-file Ruff: clean.
- Repository-wide Ruff remains non-clean due to 34 pre-existing findings in unrelated test files;
no unrelated lint files were modified.
## Notes
The approved `SourceObject` namespace grammar does not permit raw quoted ETags such as
`etag:"abc"`. The adapter therefore uses `etag:<sha256-of-opaque-etag>`: it preserves ETag-based
change identity without weakening the canonical contract or exposing validator contents.
## Review hardening follow-up
Four review findings were closed in a separate follow-up commit:
- Filesystem access now anchors a persistent descriptor at the canonical root and walks each
component with `openat` semantics (`dir_fd`, `O_NOFOLLOW`, and `O_DIRECTORY`). The regular-file
check, bounded read, metadata, and hash all use the opened descriptor. Acquisition reopens by
the same path-safe mechanism and rejects a changed fingerprint. Deterministic tests swap both a
leaf and an ancestor to symlinks at open time.
- HTTP network policy defaults to public hosts only. Initial URLs and every redirect reject
userinfo, mixed public/private IPv4/IPv6 answers fail closed, and the connected peer must be a
public member of the previously validated DNS answer set before any body bytes are consumed.
Explicit `allow_private_hosts: true` is required for trusted private deployments and local tests.
- Every HTTP response is closed in a `finally` block, including redirects, status failures,
policy failures, oversized bodies, and mid-stream exceptions.
- ETag and Last-Modified values remain adapter-internal. Repeated discovery and acquisition send
conditional headers; a 304 reuses only previously verified cached bytes and identity. The LRU
content cache has an explicit byte bound (`max_cache_bytes`). Validators are not forwarded
across redirect origins.
### Conditional cache binding correction
The conditional cache now binds bytes and validators to both the canonical provenance key and the
exact final effective representation URL. Redirect traversal recomputes request headers per hop:
validators are sent only when that exact URL matches the cached final URL, never merely because a
redirect retains an origin. A same-origin path change therefore downloads and replaces the body.
The adapter accepts 304 only when the exact request carried a bound ETag or Last-Modified validator;
unsolicited and cross-origin 304 responses are permanent protocol errors.
@@ -1,59 +0,0 @@
# Evidence Task 3 — deterministic normalization and chunking
## Outcome
- Added pure `normalize(acquired, pipeline_version)` and `chunk(document, policy)` transforms.
- Normalization enforces UTF-8 (including UTF-8 BOM), a 10 MiB input ceiling, LF line endings,
NFC Unicode, safe YAML frontmatter extraction, canonical provenance URIs, and hashes the exact
canonical UTF-8 text stored on the document.
- Undecodable, unsupported-charset, oversized, and invalid-frontmatter inputs fail explicitly;
byte content is never truncated.
- Chunking uses a versioned immutable policy, paragraph/word boundaries with deterministic
character-count hard splits for long tokens, contiguous ordinals, provenance metadata, exact
per-chunk hashes, and IDs derived from document hash + ordinal + policy version.
- Empty documents produce no chunks. Non-ASCII, CRLF equivalence, repeatability, policy changes,
duplicate-content ordinal collisions, and max-character limits are covered by tests.
## TDD evidence
- Initial focused test run failed during collection because both transform modules were absent.
- The EOF-frontmatter edge test was separately observed failing before its implementation.
- Final focused verification: `12 passed`.
## Verification
- `cd harness && .venv/bin/pytest tests/test_corpus_normalize.py tests/test_corpus_chunk.py -q`
— **12 passed**.
- `cd harness && .venv/bin/pytest -q` — **573 passed, 5 deselected**. The sandboxed attempt could
not access Docker; the approved rerun with local Docker access passed.
- Targeted Ruff over all four implementation/test files — **clean**.
- Full `cd harness && .venv/bin/ruff check .` — reports **34 pre-existing errors** in unrelated
legacy tests (unused imports and existing E702 semicolon lines); none are in Task 3 files.
## Concerns
- The 10 MiB normalization ceiling is deliberately explicit and independent of adapter download
limits. If deployment policy needs a different ceiling, it should become a versioned pipeline
configuration before ingestion is wired.
- Character limits use Python Unicode code points (`len`), not UTF-8 bytes or tokenizer tokens;
this is recorded in the chunk-policy metadata and tested with non-ASCII content.
## Review hardening follow-up
- Chunk IDs now bind the canonical document identity, document content hash, ordinal, chunk hash,
and a canonical SHA-256 fingerprint of every `ChunkPolicy` field. Identical content in separate
documents and same-version policies with different limits cannot collide.
- Boundary-aware slicing now retains separators in the slices. Concatenating every chunk exactly
reconstructs the canonical document for repeated spaces, tabs, blank lines, Markdown hard
breaks, fenced code, whitespace-only input, Unicode, and overlong tokens; every slice remains
within `max_chars`.
- Frontmatter uses a bounded `SafeLoader` variant: duplicate keys, anchors/aliases, structures
deeper than 20 nodes, and documents larger than 1000 composed nodes are rejected. YAML parse,
JSON type, credential-safety, and resulting canonical-model errors attributable to frontmatter
map to `PermanentNormalizationError(reason="invalid_frontmatter")`; invalid pipeline policy
remains a programmer-facing `ValueError`.
- Follow-up TDD evidence: the expanded focused suite first reported 11 expected failures against
the prior implementation, then passed **45/45** across normalization, chunking, and manifest
invariants.
- Follow-up full verification: **586 passed, 5 deselected**. Targeted Ruff is clean. Full Ruff
continues to report the same **34 unrelated pre-existing** violations in legacy tests.
-107
View File
@@ -1,107 +0,0 @@
# Evidence Task 4 — shared job envelope
Status: complete
## Delivered
- Immutable `JobSpec`, `JobRun`, `JobReport`, per-stage state, sanitized error, and UTC
timestamp records.
- `run_job(spec, stages)` with a durable checkpoint at job start, before and after every stage,
and at terminal state. Successful stages are skipped when a prior run is resumed.
- Atomic JSON checkpoint/report replacement using a unique same-directory temporary file,
file `fsync`, atomic `os.replace`, and parent-directory `fsync`.
- Public reports contain fixed operational fields only. Workspace paths, stage return values,
exception messages, source content, credentials, and arbitrary metadata are not serialized.
- `WorkspaceJobLock` uses non-blocking kernel `flock` on a stable workspace/job-specific inode.
Locks are released by the kernel on process exit; lock files are never removed based on PID,
avoiding stale-lock and PID-reuse deletion races. Evidence and DWH use distinct lock files.
- Dry-run intent is immutable in the spec/report and exposed to every stage through `JobContext`.
## TDD evidence
Initial focused collection failed because `tht.jobs` did not exist. Tests then drove:
- failure, sanitized reporting, resume, and idempotent successful-stage skipping;
- corrupt-checkpoint refusal before stage execution;
- JSON schema and path/secret/PII exclusion;
- dry-run propagation and ordered aware timestamps;
- multiprocessing exclusion, distinct Evidence/DWH jobs, traversal rejection, and recovery after
a lock-owning process crashes.
Final focused result:
```text
11 passed in 0.42s
```
## Verification
```text
cd harness && .venv/bin/pytest -q
597 passed, 5 deselected, 17 warnings in 28.45s
cd harness && .venv/bin/ruff check tht/jobs tests/test_job_runner.py tests/test_job_locking.py
All checks passed!
```
The full Ruff invocation was also run. It reports 34 pre-existing violations in unrelated legacy
tests; no Task 4 file is among them. L2 tests remain deselected by the repository configuration.
## Operational notes
- `fcntl.flock` intentionally targets the supported Linux/macOS deployment environments; it is not
a Windows locking implementation.
- The envelope does not publish or mutate an active corpus. Later pipeline stages must use
`JobContext.run_dir` for staging and perform their own final atomic publish only after validation.
- A dry run is an execution mode foundation: the runner exposes and records it; individual stages
remain responsible for suppressing external mutations.
## Review hardening follow-up
Four post-implementation findings were fixed test-first:
1. Resume compatibility is now a canonical SHA-256 fingerprint over checkpoint schema version,
hashed workspace identity, job type, dry-run mode, explicit spec/pipeline versions,
configuration/input fingerprints, and the exact ordered explicit `stage_ids`. Any insertion,
removal, reorder, mode, identity, version, config, or input change rejects resume before a stage
executes. Omitting `resume_run_id` remains the explicit safe path for a new run.
2. Lock traversal now uses directory file descriptors with `O_DIRECTORY` and `O_NOFOLLOW`.
Lock files use `O_NOFOLLOW | O_CLOEXEC`; `fstat` requires a regular file owned by the current
UID with one link, and permissions are forced to `0600` (`0700` for private directories).
Pre-existing lock-file and lock-directory symlinks are rejected.
3. Stage failures now serialize only the fixed safe tuple `internal` / `stage_exception` /
`stage execution failed`. Neither exception class names nor messages are inspected for output;
a hostile exception-name/message regression test proves a terminal failed report is retained.
4. Job/run directory creation is no-follow, owner-checked, private, and durable. Each newly created
parent is fsynced, the run directory is fsynced before the first atomic file write, and the
existing file-fsync → replace → directory-fsync ordering has an explicit regression test.
Follow-up verification:
```text
focused job/lock suite: 27 passed in 0.45s
full harness suite: 613 passed, 5 deselected, 17 warnings in 29.65s
Task 4 scoped Ruff: All checks passed
```
Repository-wide Ruff continues to report the same 34 unrelated pre-existing legacy-test findings.
## Final resume-integrity fix
Resume is now read-only until the source checkpoint proves trustworthy. The runner loads the source
before allocating a new run ID or directory, validates the exact stage state/timestamp/error ledger,
rejects duplicate stage identifiers, and recomputes compatibility from every persisted compatibility
field plus the exact ordered persisted stage IDs. It first requires the stored fingerprint to match
that recomputation, then compares the trusted recomputation with the requested job fingerprint.
Valid-JSON tampering tests cover removed, inserted/duplicated, reordered, and substituted stages;
input-field and stored-fingerprint changes; and invalid stage-state shapes. Every rejection occurs
before stage execution and asserts that the runs directory contains no orphan allocation.
Final verification:
```text
focused job/lock suite: 34 passed in 0.56s
full harness suite: 620 passed, 5 deselected, 17 warnings in 27.42s
Task 4 scoped Ruff: All checks passed
```
@@ -1,82 +0,0 @@
# Evidence Task 5 report
## Outcome
Implemented an incremental Evidence corpus pipeline with immutable materialized generations,
generation-scoped vector records, and an fsynced atomic `ACTIVE` pointer. Runtime Evidence
artifact lookup reads the active canonical manifest and keeps a legacy source-tree fallback only
when no corpus has been published.
The CLI is available as `tht preprocess evidence [--dry-run] [--resume RUN_ID] [--json]`.
JSON success and failure output is pristine and failure details are sanitized.
## Safety and failure model
- A workspace writer lock serializes preprocess writers; readers never take the lock.
- Generation directories, manifests, materialized files, locks, and `ACTIVE` reject symlink/path
escape cases and use owner-only durable writes.
- Vector records use generation-specific keys and metadata. The active manifest maps each active
document to its valid vector generation, allowing unchanged documents to retain their vectors.
- Runtime retrieval admits only active document IDs and their manifest-selected generations.
Removed documents and partial writes from failed generations are therefore unreachable.
- Embedding count and dimension checks occur before vector upsert; vector write count is checked
before staging/publish. Any failure leaves `ACTIVE` unchanged.
- Dry runs perform discovery/fingerprint planning only and never acquire, embed, write vectors, or
publish. Fully unchanged runs return the active generation without creating a replacement.
- Resume can safely retry idempotent generation-scoped upserts and publish an already staged,
compatibility-checked generation after a crash between staging and pointer replacement.
## TDD evidence
Initial focused collection failed because `tht.corpus.pipeline` and `tht.corpus.store` did not
exist. The implemented suite covers incremental skips, removals, model/policy rebuilds, acquire and
partial-vector failures, dry-run isolation, dimension validation, atomic reader snapshots, pointer
validation, symlink defense, and pristine CLI JSON.
Fresh focused verification:
```text
18 passed, 3 warnings in 0.39s
```
Command:
```text
.venv/bin/pytest tests/test_corpus_pipeline.py tests/test_corpus_publish.py \
tests/test_preprocess_cli.py tests/test_search_pack.py tests/test_session_documents.py -q
```
Scoped Ruff: `All checks passed!`
Broader non-Docker/non-packaging run reached `560 passed, 5 deselected`; ten pre-existing HTTP
adapter tests could not bind localhost under the sandbox. The complete suite reached `570 passed,
5 deselected`, with the remaining failures/errors caused by denied Docker socket, localhost bind,
and offline wheel-build access. No task-focused test failed.
## Remaining operational gate
Live pgvector integration needs Docker or an authorized local pgvector endpoint. The compensation
strategy is logical isolation rather than destructive cleanup because the shared `VectorStore`
port intentionally exposes no delete/transaction API; unreachable failed generations can be
garbage-collected by a future maintenance job.
## Review integration wave
Added an enforceable `metadata_filter` vector-port contract and capability flags. Direct pgvector
places exact Evidence generation/document predicates in SQL before `LIMIT`; HTTP sends the same
filter to the RPC and deliberately does not use the legacy 404 fallback. The reader RPC script now
validates and applies that filter. Normal Evidence search and search-pack use an ACTIVE-aware
searcher that groups active documents by generation, executes complete server-filtered searches,
and merges the results.
Added exact-generation Evidence cleanup to direct and HTTP writers plus the allowlisted writer RPC.
Pipeline failures compensate both staged filesystem state and vector writes; cleanup failures stay
sanitized and ACTIVE filtering remains the exposure boundary. Corpus-present session artifact
resolution now fails closed on corrupt/missing ACTIVE rather than falling through to source files.
Focused review-wave verification: 45 passed, scoped Ruff clean. A mocked REST regression proves
the exact filter payload and fail-closed legacy 404 behavior.
Still outstanding from the expanded review request: Task-4 JobRunner stage-by-stage integration,
published-generation retention/garbage collection, same-fd `dirfd` materialized-file reads, and
live local pgvector integration could not be completed in this wave.
@@ -1,94 +0,0 @@
# Evidence Task 5B implementation report
## Status
Integrated Evidence preprocessing with the Task 4 `JobRunner`. The CLI now accepts only a
32-character JobRunner run ID for `--resume`; generation IDs remain outputs. Runs persist the
exact ordered stages `discover`, `acquire_normalize_chunk`, `embed`, `vector_upsert`,
`stage_validate`, `publish`, and `retention_cleanup`.
Successful-stage artifacts are copied into the new resume run before execution, allowing later
stages to continue without rediscovery, acquisition, normalization, chunking, or embedding.
Job compatibility includes workspace, configuration, discovered-input, pipeline, embedding, and
chunk-policy fingerprints. Generation-specific filesystem/vector compensation is retained, and a
compensated generation is rotated before retry. `ACTIVE` is mutated only by `publish`.
Dry-run executes discovery/planning and makes every side-effecting stage a no-op. JSON output is
pristine and includes the JobRunner `run_id`, `resumed_from`, generation, plan, and publish status.
## TDD evidence
- RED: run-ID rejection and resume-artifact tests failed because generation IDs reached
configuration and resume runs had empty artifact directories.
- GREEN: the two regression tests passed after strict CLI validation and durable artifact carryover.
- Added pipeline job-plan and dry-run counting-fake coverage; both passed.
## Fresh verification
- Focused integration/search suite: `62 passed, 4 warnings`.
- Available harness suite excluding sandbox-blocked Docker, loopback HTTP-server, and networked
wheel-build tests: `559 passed, 5 deselected, 18 warnings`.
- Scoped Ruff: `All checks passed!`.
- `git diff --check`: clean.
## Environment limitations and concerns
The literal full harness invocation cannot complete in the managed sandbox: Docker socket access,
loopback HTTP test servers, and the `uv build` dependency resolution path are denied. It reached
`575 passed, 5 deselected` before those environment errors. The available-suite rerun above is
green.
One pre-existing Pydantic serialization warning is exposed by the new end-to-end job test when
canonical metadata contains frozen tuple values; it does not contaminate CLI stdout. Retention is
an explicit stable no-op until a retention policy is configured.
## Review fix wave — crash consistency and artifact integrity
Addressed all five follow-up findings:
- `JobRunner` now supports a test-only post-call/pre-checkpoint fault hook. Each stage seals a
canonical artifact manifest containing required flat filenames, SHA-256, byte size, producer
stage, and the full spec compatibility fingerprint. Resume validates the checkpoint and every
sealed artifact before allocating/copying a new run, rejecting missing, tampered, extra, nested,
or symlinked state. A sealed `running` stage is promoted after a simulated process crash; a
sealed `failed` stage is deliberately retried.
- Vector intent (exact record IDs and content hashes) is sealed before upsert. Execution reconciles
`existing_hashes` and writes only missing/mismatched rows. Crash-after-effect tests prove no
duplicate acquire, embed, or vector upsert.
- Raw upsert, stage, recovery-upsert, recovery-stage, and publish exceptions compensate the exact
generation. Compensation markers survive failed checkpoints; resume rotates the generation,
refreshes generation-bound artifacts, reconciles vectors, and stages idempotently.
- `CorpusStore.publish` is idempotent and failure-atomic. If replace succeeds but directory fsync
fails, it restores the previous `ACTIVE` value (or removes a newly created pointer), fsyncs the
rollback, and re-raises. Pipeline cleanup refuses to discard a generation referenced by ACTIVE.
- Added crash/resume coverage after all seven ordered stages; corrupt/missing plan, manifest, and
embeddings; unsafe extra paths; nonexistent run IDs; raw vector/stage failures; and post-replace
ACTIVE rollback.
Fresh fix-wave verification:
- Focused jobs/corpus/CLI/search suite: `82 passed, 17 warnings`.
- Available harness suite (same sandbox exclusions described above):
`579 passed, 5 deselected, 31 warnings`.
- Scoped Ruff and `git diff --check`: clean.
## Final P1 fix — effect state and checkpoint-bound manifest roots
- Stage checkpoints now distinguish `intent` from `completed`. Vector intent is atomically sealed
and checkpointed before upsert. A process-level `BaseException` after a partial multi-record
write leaves the stage `running/intent`; resume never promotes it and instead reconciles
`existing_hashes`, writing only the missing records. The completed state is persisted only after
reconciliation returns successfully.
- Every stage now persists its completed artifact state while still `running`, before the
post-call fault hook. The checkpoint binds the SHA-256 of canonical `artifact-manifest.json`,
effect state, exact producer stage, and exact required-file mapping. Resume validates this root
and all bindings before promotion or copying.
- Added process-interruption coverage proving the already-written vector record is not submitted
twice, remaining records are written, and publish completes only after reconciliation. Added
coordinated artifact/manifest, spec-binding, and producer-binding tamper rejection tests.
Fresh verification:
- Focused jobs/corpus/CLI/search suite: `86 passed, 18 warnings`.
- Available broad harness suite: `583 passed, 5 deselected, 32 warnings`.
- Scoped Ruff and `git diff --check`: clean.
-188
View File
@@ -1,188 +0,0 @@
# Evidence Task 5C report
## Delivered
- Added `vector.retain_published_generations` (default `3`, validation minimum `1`).
- Retention runs only after publication. It keeps ACTIVE, the newest configured generations,
and generations referenced by running or resumable failed job checkpoints.
- Cleanup deletes the exact Evidence generation from the vector store before removing its
immutable filesystem directory. Vector failures retain filesystem metadata for retry and
produce credential-free partial reports.
- Added idempotent `tht preprocess evidence gc [--dry-run] --json` reconciliation with pristine
JSON output.
- Materialized document reads now open generation/documents components with directory file
descriptors and `O_NOFOLLOW`, require a regular file owned by the process with one link, and
hash the bytes read from the same descriptor against the canonical manifest.
- HTTP generation deletion is pinned to `delete_vector_generation` with exact
table/kind/generation arguments. Legacy 404 responses fail closed with an actionable,
sanitized migration message.
## Evidence
- Focused retention, safe-read, CLI, and HTTP contract tests: `51 passed` (Docker-backed direct
parametrizations excluded from that focused invocation).
- Real Docker pgvector adapter suites: `33 passed`.
- Full harness suite, including Docker-backed tests: `668 passed, 5 deselected`.
- Changed-file Ruff: clean.
- `git diff --check`: clean.
The five deselected tests are the repository's opt-in `l2` tests requiring external services;
they are not local pgvector tests. Test output retains pre-existing Pydantic serialization and
legacy-config deprecation warnings.
## Review fix wave
- Publication is now explicit and durable (`PUBLISHED` marker). Retention candidates require a
valid generation manifest and publication marker (ACTIVE remains backward-compatible), so
staged and malformed directories neither consume retention slots nor become deletion targets.
- The policy retains ACTIVE plus exactly `N-1` newest rollback publications, ordered by durable
publication time and generation id. Running and failed-resumable JobRunner checkpoints protect
every referenced plan generation.
- `VectorStore` now exposes exact Evidence generation inventory. Direct pgvector uses a constrained
`SELECT DISTINCT` over `kind='evidence'` and `metadata.vector_generation`; HTTP uses the
allowlisted `list_evidence_generations` RPC and fails closed on legacy 404. The writer RPC SQL,
revokes, and grants are packaged in `create_vector_writer_rpc.sql`.
- Explicit GC reconciles the union of published filesystem generations and vector-only orphans,
preserving vector-before-filesystem deletion and retry semantics.
- `run_as_job` holds the same corpus writer lock across checkpoint recovery, staging, publish, and
retention. Explicit GC already uses this lock, serializing candidate snapshots with publishers.
- Session artifact consumers no longer receive the corpus source path after validation. They get
an owned, read-only copy atomically written from the bytes read and hash-validated on the same
descriptor.
Fresh verification after the fix wave: full harness `672 passed, 5 deselected`; Docker pgvector,
HTTP parity, and migration suites `43 passed`; exact direct inventory/delete integration `1 passed`;
changed-file Ruff and `git diff --check` clean.
## Final hardening verification
- Canonical generation validation is exact (`^gen:[0-9a-f]{32}$`) before HTTP/direct deletion;
malformed HTTP inventory rows fail closed rather than entering the GC candidate set.
- Added explicit protection coverage for running and failed-resumable JobRunner checkpoints, plus
a second-GC idempotence assertion for vector-only orphan reconciliation.
- Added deterministic concurrent locking coverage: a job paused after discovery retains the corpus
writer lock, explicit GC blocks, then completes after publication without deleting the active run.
- Added a descriptor-race regression: replacing the corpus pathname immediately after `read(2)`
leaves the atomically materialized session-owned copy byte-for-byte equal to the validated ACTIVE
document and its manifest hash.
Final fresh evidence: Docker pgvector/HTTP/migration suites `48 passed`; full harness `680 passed,
5 external L2 deselected`; changed-file Ruff and `git diff --check` clean.
## Integrated Task 5 dependency fixes
- GC now distinguishes filesystem retention from vector dependencies. ACTIVE and the newest
`N-1` published manifests keep their directories; every exact generation in their
`document_generations` maps remains vector-protected even after its old publication directory is
evicted. Job-protected manifests receive the same dependency treatment.
- The real four-publication Docker lifecycle now includes an unchanged document whose vectors come
from the first generation. With retention `N=2`, only the final two publication directories remain
while the first generation's vectors remain searchable from ACTIVE and survive restart/explicit GC.
- Evidence lookup is always wrapped by the ACTIVE-aware searcher. With no corpus/ACTIVE, Evidence
returns no rows and search packs cannot expose legacy vectors; non-Evidence kinds are unchanged.
- Session artifact resolution holds the corpus writer lock, snapshots the active manifest once, and
materializes bytes using that exact `manifest_id`, preventing a concurrent publish/retain-1 GC from
changing or deleting the selected source generation.
Focused unit tests, the updated real Docker lifecycle, changed-file Ruff, and `git diff --check` pass.
The final full harness invocation completed with exit code 0, including the concurrently added DWH
JobRunner tests.
## Final ACTIVE search review fixes
- `ActiveEvidenceSearcher` now treats default (`kinds=None`) and mixed-kind searches as explicit
split queries: non-Evidence kinds are queried separately, while Evidence is queried only with
ACTIVE manifest generation/document predicates applied server-side before every limit.
- Results are merged deterministically by descending similarity then stable id and truncated once
to the caller's global `top_n`. Pure non-Evidence searches retain their original delegate path.
- The corpus writer lock now covers manifest snapshot construction and all corresponding vector
queries, preventing retain-1 publication/GC from switching or deleting generations mid-search.
- Removed the public post-LIMIT `active_evidence_hits` helper; no public Evidence path performs
client filtering after limit.
Focused default/mixed/no-ACTIVE/search-pack tests pass, the real Docker pgvector lifecycle passes,
and the final full harness plus scoped Ruff/diff invocation completed with exit code 0.
## Workspace-scoped Evidence isolation
- Evidence manifests, vector metadata, and record keys now carry the stable JobRunner workspace id
derived from the configured workspace identity (config stem), never credentials or absolute paths.
- Every ACTIVE server-side predicate includes `workspace_id`. Legacy unscoped rows therefore fail
closed and cannot appear in Evidence results.
- Vector generation inventory and deletion require the workspace namespace across the port, direct
pgvector adapter, HTTP client/adapter, and allowlisted RPC SQL. Legacy unscoped RPC overloads are
explicitly dropped during migration; destructive SQL matches collection, kind, generation, and
workspace together.
- GC recovers the persisted namespace from ACTIVE for explicit/restarted cleanup and can only list
or delete that workspace's generations. Real shared-pgvector coverage proves deleting a generation
for workspace A preserves the same generation in workspace B.
- `PipelineResult.model_dump` now serializes fields explicitly instead of `dataclasses.asdict`,
avoiding deepcopy of immutable `FrozenDict` metadata while preserving pristine JSON CLI output.
Final focused verification: `89 passed` across corpus/CLI JSON, direct/HTTP parity, migrations, and
real Docker pgvector lifecycle; scoped Ruff and `git diff --check` clean. A contemporaneous full-suite
run reached unrelated Task 6 immutable-file tamper tests; those files were deliberately not changed.
## Immutable corpus/workspace binding
- A corpus root becomes bound to the workspace id persisted in its ACTIVE manifest. Job, non-job,
explicit GC, and ACTIVE search entry points compare the configured namespace before discovery,
vector access, staging, deletion, or ACTIVE mutation.
- Reusing the same paths after renaming a workspace now fails closed with a typed/sanitized message:
use a new corpus root or perform an intentional explicit rebuild. Unscoped legacy manifests also
fail this ownership check.
- Tests prove unchanged-document reuse cannot silently mix workspace A vectors into a workspace B
manifest, and that mismatched job, GC, and search paths perform no vector/filesystem mutations.
Focused workspace-binding, search-pack, preprocess JSON, and scoped Ruff/diff tests pass.
Compatibility follow-up: direct/internal `CorpusPipeline` instances now distinguish an omitted
workspace identity from an explicit config/job identity. An unbound instance adopts the persisted
ACTIVE owner (or `default` only for a brand-new direct corpus), preserving safe resume/GC tests and
the real pgvector lifecycle. Explicit config/job identities still fail closed on any mismatch. The
two reported regressions, workspace mismatch guards, real Docker lifecycle, scoped Ruff/diff, and
the full harness suite all pass.
Final fail-closed follow-up: persisted ACTIVE ownership is now validated under the corpus lock before
every configured search delegate, including default, mixed, pack, and non-Evidence-only operations.
Malformed or missing `metadata.workspace_id` is intrinsically rejected even for unbound direct
callers; source discovery, vector operations, GC, files, and ACTIVE remain untouched. Focused tests,
real Docker lifecycle, scoped Ruff/diff, and the full harness regression run pass.
Final lock/preflight follow-up: `CorpusPipeline.gc()` now acquires the corpus writer lock itself for
ownership validation through vector/filesystem cleanup. The store lock is thread-reentrant so nested
job retention is safe without weakening cross-thread/process exclusion; the CLI wrapper no longer
double-locks. Search find/pack performs locked corpus ownership preflight immediately after config
load, before DWH leasing, vector/searcher factories, embeddings, or schema work. Focused concurrency
and fail-closed tests, real Docker lifecycle, scoped Ruff/diff, and the full harness pass.
## Compact public Evidence reports
- Public `PipelineResult.model_dump()` is now a bounded operational envelope: terminal status,
run/resume/publication/generation/manifest identifiers, capped changed/unchanged/removed source
identifiers, and aggregate document/chunk counts. Full manifests, bodies, and metadata remain
internal/on disk and are never serialized to CLI stdout.
- `tht preprocess evidence` exits `1` for any durable terminal status other than `succeeded` in
both JSON and text modes. JSON stdout remains one pristine sanitized object; text mode emits one
compact stderr error without traceback, exception identity, evidence content, or credentials.
- Tests cover a real failed acquisition job, sensitive evidence content, capped thousand-item
summaries, bounded report size, and smoke-compatible changed/unchanged fields.
Focused tests and scoped Ruff/diff pass. The contemporaneous full suite reaches an unrelated Task 6
DWH snapshot fixture missing its newly required workspace identity.
### Safe result representation and exact text totals
- `PipelineResult.manifest` is explicitly excluded from dataclass representation and the custom
representation is fixed-size operational data only. It omits manifest ids, documents, chunks,
content, metadata, and errors; `str(result)` inherits the same safe representation.
- Text-mode Evidence success output reads the uncapped aggregate totals from `payload["counts"]`
rather than the intentionally capped identifier arrays.
- Regression coverage builds a thousand-document/chunk manifest containing content and
credential-like metadata secrets, checks bounded `repr`/`str`, and verifies exact totals above
the 100-item public-array cap.
Focused Evidence verification passes (`67 passed`), and scoped Ruff is clean. The full harness run
is not green in this sandbox: Docker-backed tests cannot access the daemon, wheel packaging cannot
use the restricted build environment, and concurrent Task 6 DWH binding changes currently fail two
DWH tests. None of those failures touch the Evidence files in this follow-up.
@@ -1,50 +0,0 @@
# Evidence Task 5D — Real pgvector lifecycle gate
## Status
Complete. The Docker-backed L0 gate uses one persistent `pgvector/pgvector:pg16`
database and the production migrations, direct reader/writer `PgVectorStore`,
`CorpusStore`, `CorpusPipeline.run_as_job`/JobRunner, ACTIVE Evidence retrieval,
search-pack fusion, owned session artifact copy, retention, and explicit GC.
## Lifecycle covered
- Four real corpus publications with retention set to two generations.
- A higher-similarity stale vector proves ACTIVE metadata filtering happens before LIMIT
for normal Evidence retrieval and the search-pack fusion path.
- A removed source is absent from ACTIVE retrieval and cannot be copied to a session.
- An injected process death occurs after one real committed vector upsert. Resume uses the
real run ID, preserves that record, fills the missing records, and produces no duplicate keys.
- Database engines and direct store objects are disposed/recreated before persisted ACTIVE
retrieval is checked again.
- An exact canonical vector-only orphan generation is discovered and removed by explicit GC.
- Filesystem and vector inventories converge exactly to ACTIVE plus one rollback; a second GC
is a no-op.
- Owned session artifact bytes and SHA-256 match the ACTIVE canonical document.
## Production bug found and fixed
Production migration `003_roles.sql` intentionally restricted `vector_writer`, but omitted
the privileges used by the production generation lifecycle: `SELECT(metadata)` for inventory
and `DELETE` for cleanup on `vectors.evidence`. Consequently a real job published successfully
and then failed in `retention_cleanup` on its first run.
Added versioned migration `004_evidence_generation_gc.sql` granting only those two Evidence
generation-management privileges. Runtime application code was not redesigned.
## Verification
- Target lifecycle: `1 passed` (Docker-backed).
- Full harness: `681 passed, 5 deselected`.
- Scoped Ruff: passed.
- `git diff --check`: passed.
The existing Pydantic serialization and legacy-workspace deprecation warnings remain unchanged.
## Follow-up assertion correction
The removal phase now retains the removed canonical document ID/ref before publication and
asserts both fields are absent from post-resume ACTIVE Evidence hits. It reruns the real
search-pack fusion after removal, proves active fourth-generation content is positively
returned in both paths, and proves the removed content remains absent. The owned session
artifact lookup for the retained removed ID remains empty.
@@ -1,49 +0,0 @@
# Evidence Task 6 — final fd-anchored DWH correction
All DWH generation state below `.tht-dwh` is now accessed relative to the directory descriptor
retained by the shared/exclusive generation lease. ACTIVE reads, atomic temp writes, replacement,
fsync, and rollback use `openat`/`replaceat` operations. Generation staging, validation,
reconciliation, resume checks, retention classification, and recursive deletion likewise use owned
root/generations/candidate descriptors with `O_NOFOLLOW`; locked operations no longer reopen
generation paths through `workspace_root`.
Portable reader snapshots are copied from validated generation file descriptors into private 0700
process-owned temporary directories while the shared lease is held. This avoids Linux-only
`/proc/self/fd` paths and prevents a renamed/replaced `.tht-dwh` pathname from redirecting later
schema or LSH reads. Lease-scoped copies are removed on exit and standalone snapshots are removed
at process exit.
Deterministic adversarial tests rename the DWH root after lease acquisition during ACTIVE reads,
ACTIVE publication, and retention cleanup. Each test proves the replacement tree is never read,
written, or deleted; the descriptor-pinned original either completes consistently or fails closed.
Existing owner binding, legacy rejection, crash reconciliation, resume, atomic rollback, retention,
and reader/writer exclusion behavior remains covered.
## Final review correction
Snapshot materialization now reads the manifest and every owned artifact exactly once through the
already-open generation descriptor, validates each hash against those exact bytes, and writes the
same byte objects to the private snapshot. A deterministic second-read mutation test proves hostile
pickle bytes can neither pass validation nor enter the snapshot. Reconciliation closes the ACTIVE
generation descriptor in a `finally` block on matches, mismatches, and exceptions. Pipeline-owned
snapshot directories are removed and deregistered after `run_job` on both successful and failed
runs, preventing repeated pipeline use from accumulating temporary directories or registry entries.
The cleanup boundary now begins immediately after snapshot materialization. Resume checkpoint
validation and `JobSpec` construction are guarded by the same release routine as `run_job`, so
corrupt/mismatched resume state or constructor failure clears the pipeline holder, removes the
private directory, and restores the snapshot registry to its prior state before propagating.
## Shipped preprocessing startup contract
Local-vector preprocessing now uses a dedicated Compose override. Both one-shot jobs depend on a
successfully completed `vector-migrate`, whose transitive chain waits for database health and role
reconciliation. The generic preprocessing overlay remains independently renderable and contains no
local-vector services or password secrets. README commands include the local override and build the
job image before running.
The real clean-project smoke no longer injects dependencies or manually starts, reconciles, or
migrates PostgreSQL. Its first shipped `compose run preprocess-evidence` demonstrably creates the
database, waits for health, runs reconciliation and migration, then runs the Evidence job. Unchanged
rerun, changed-source publish, DWH preprocessing, ACTIVE verification, and injected-failure cleanup
all pass through the same shipped dependency path.
@@ -1,93 +0,0 @@
# Evidence preprocessing Task 7 report
Implemented the S3-compatible Evidence adapter, explicit preprocessing Compose overlay, and
operational gates.
- S3 discovery uses bounded paginator pages, page size, and total objects; acquisition enforces a
byte ceiling and always closes streaming bodies.
- Provenance is canonical `s3://bucket/key`. Versioned objects use `s3-version:<version>`;
unversioned objects use a hashed exact ETag, and acquisition refuses validator drift.
- The adapter uses boto3/botocore rather than custom signing. TLS verification is enabled by
default. Custom HTTP and private endpoints require independent explicit opt-ins; endpoint
userinfo is rejected and public custom endpoints are DNS-policy checked.
- Access, secret, and session credentials support file-secret resolution into masked `SecretStr`
config fields. They are never emitted in provenance, reports, errors, or Compose environment.
- `deploy/compose.preprocess.yaml` provides separate one-shot Evidence and DWH jobs and is inert
unless explicitly included with the `preprocess` profile.
- `scripts/preprocess-smoke.sh` verifies both services render without secret material and pins an
unchanged rerun plus a modified generation through deterministic pipeline tests.
Verification: focused S3/HTTP/filesystem/config tests 34 passed; operational smoke 2 passed; core
image with locked boto3 extra built; full harness 702 passed, 5 deselected; scoped Ruff and diff
checks passed.
Operational risk: custom S3-compatible endpoints remain part of the deployment trust boundary.
Private endpoint access must be explicitly enabled and should be restricted by container egress
policy in production. S3 list consistency semantics are provider-defined; version IDs are preferred
over ETags wherever bucket versioning is available.
## Review correction
The Compose overlay now uses committed, purpose-built Evidence and DWH workspace files with
job-specific dependencies. Its services create their lock roots and mount only the vector secrets
they consume. The operational smoke is a real isolated Compose project: real pgvector migrations,
a deterministic in-project embeddings endpoint, actual Evidence CLI JSON across initial/unchanged/
mutated runs, exact ACTIVE verification, an actual DWH introspection job, and owned cleanup.
S3 custom endpoints now fail closed unless declared trusted; HTTP and private loopback endpoints
need additional independent opt-ins. Boto uses forced path-style addressing. Custom endpoints reject
userinfo, query, fragment, and non-root paths. Buckets use strict DNS syntax; listed keys must remain
under prefix and within the S3 byte bound; validators must be nonempty/bounded. Because
ListObjectsV2 does not provide version IDs, discovery honestly fingerprints the exact ETag and
acquisition rejects ETag drift.
Final correction verification: S3/config focused 20 passed; full harness 721 passed, 5 deselected;
real Compose smoke and image build passed; scoped Ruff, shell syntax, and diff checks passed.
## Final security review correction
Literal non-global IPv4/IPv6 endpoints now require the private-endpoint opt-in without claiming DNS
pinning for hostnames. Pagination uses explicit continuation requests and never fetches page
`max_pages + 1`. IP-shaped buckets, leading-slash prefixes, empty/overlong/control-character keys,
and absent validators fail closed. Acquisition accepts only the exact stored `SourceObject` and
compares the response ETag with the stored discovery validator. The real smoke snapshots generation
directory counts after every run and has an injected-failure cleanup mode; cleanup fails if Compose
down fails or any owned container, volume, or network remains.
The canonical smoke correction counts only root-level `corpus/gen-<32 hex>` directories. It exposed
that the durable job path still published an empty unchanged generation; the pipeline now returns
the existing ACTIVE generation without staging a directory when compatibility and all source
fingerprints are unchanged. The smoke therefore proves directory deltas `+1`, `+0`, `+1`.
Failure injection runs a real exit-97 command after resources exist and reaches the EXIT trap.
Cleanup aggregates Compose-down, residual container/volume/network, and temp-directory failures
while preserving the original failure status. S3 prefixes are validated before any client request
for leading slash, UTF-8 byte length, controls, and DEL.
## Canonical unchanged-run correction
The durable job now persists a deterministic source snapshot keyed by source identity. Each entry
binds canonical URI, exact source fingerprint, UTC modification time, canonical immutable metadata,
and explicit media type and size contract fields. The manifest also binds document-to-source
provenance, supplied config/input fingerprints, compatibility, embedding settings, and pipeline and
chunk-policy versions.
An unchanged run reuses ACTIVE only when ownership, bindings, the complete snapshot, document
provenance, materialized document hashes, and every required vector ID/content hash match exactly.
Snapshot changes rebuild only the affected sources; job input/config changes publish a new manifest
while retaining valid stable vector-generation dependencies. Missing or corrupt legacy contract
metadata, documents, or vectors fails closed and rebuilds. The Compose smoke now explicitly expects
the unchanged no-op to report `published=false` while proving generation deltas `+1`, `+0`, `+1`.
## Corrupt ACTIVE reconstruction correction
ACTIVE reuse now reconstructs each source contract from the persisted discovery snapshot and checks
the deterministic document identity, canonical URI, source fingerprint, UTC modification time,
source metadata, applicable media type, content hash, and pipeline identity against the owned
materialized document. The persisted document-source map carries the same exact binding.
Chunks are recomputed under the current chunk policy and must match the manifest exactly in count,
order, IDs, ordinals, content, hashes, linkage, provenance, and policy metadata. Vector health must
report the configured dimension, and every recomputed chunk must have its generation-scoped vector
ID with the exact content hash. Missing, altered, or extra chunks and corrupt document or vector
contracts therefore disable the no-op and rebuild, while a valid unchanged run still performs no
source acquisition.
@@ -1,16 +0,0 @@
# Model provider credential boundary
The backend accepts only an absolute `THT_MODEL_API_KEY_FILE` reference. `PiProcessManager` reads
and validates it afresh before each hosted-provider spawn, rejects symlinks, non-regular/hard-linked,
empty, whitespace-containing, oversized, unreadable, or permissively-mode files, and accepts Docker
0444 secrets only beneath `/run/secrets`. Failures are sanitized and occur before child creation.
Provider names are normalized and mapped to Pi-recognized variables. The child environment removes
the generic path, deprecated `PI_PROVIDER_API_KEY`, and all unselected known provider keys before
injecting only the selected key. Values never enter argv, settings, health, or diagnostics. Local
providers remain keyless and unknown hosted providers fail closed.
The production Compose overlay mounts `model_api_key` read-only and points the backend at its file;
the deployment render smoke proves the value is absent from rendered configuration. Entrypoint,
root README, Pi configuration guide, environment example, and secrets operator guide document the
new contract and reject the legacy generic value variable.
@@ -1,59 +0,0 @@
# Local pgvector whole-plan final fix report
## Outcome
All four binding final-review findings are closed.
1. `PgVectorStore.health()` checks namespace `USAGE` independently for reader and writer
before inspecting vector types. Real PostgreSQL tests revoke only schema `USAGE`, prove both
health sides false and operations unavailable, then grant it back and prove recovery.
2. Direct reader/writer passwords use workspace `password_file` references. Compose mounts the
two files read-only into core and exposes only `_FILE` paths. Rendered Compose and live
`docker inspect` checks prove secret contents are absent.
3. Direct search failures map to `VectorReadUnavailable`; hash/upsert failures map to
`VectorWriteUnavailable`. Messages are fixed and sanitized, original exceptions remain chained,
and upsert rollback is preserved.
4. The shared secret policy uses Linux `stat -c` with macOS `stat -f` fallback. Host files permit
only `0600`/`0400`; Docker's read-only `0444` is accepted only beneath `/run/secrets`. Tests and
operator docs pin this exact policy.
## TDD evidence
The new config, mode, schema-usage, unavailable-connection, and permission regressions failed
before their implementations. The first live secret-policy run also caught GNU `stat -f` accepting
an incompatible format invocation; detection now tries the native Linux form first. The next live
run caught smoke-generated rotation fixtures at `0644`; fixtures now model the documented host
policy.
## Verification
- Real direct pgvector + HTTP parity: `31 passed`.
- Full harness from `harness/`: `493 passed, 5 deselected`.
- Live `local-vector` rotation, restart persistence, inspect boundary, and backup/restore: pass.
- Core image vector migration discovery/status smoke: pass.
- External and local Compose deployment security contracts: pass.
- Config/port focused suite: `26 passed`.
- Secret policy, bootstrap rotation, and backup/restore safety scripts: pass.
- Changed Python Ruff, shell syntax, and `git diff --check`: pass.
One attempted full-harness invocation from the repository root produced a path-dependent failure
in an existing test that opens `workflow.yaml` relative to CWD. It was immediately rerun using the
documented `cd harness && .venv/bin/pytest -q` command and passed completely.
## Operational notes
Workspace files contain file paths, never direct passwords. Secret contents necessarily exist in
the in-process validated `DatabaseConfig` used to establish PostgreSQL connections, but are not
serialized by doctor/Compose/inspect paths. Docker Desktop file-backed secrets may appear as bind
mounts; the safe runtime exception is therefore based on the read-only service mount location
`/run/secrets`, while source files remain owner-only on the host.
## External-profile regression follow-up
Local pgvector is now an explicit `deploy/compose.local-vector.yaml` overlay. The base Compose and
production external override contain no direct vector password declarations, mounts, or `_FILE`
variables, so external deployments do not resolve or require local password files. A real lifecycle
gate unsets all local secret-file variables, renders external config, builds and starts core, waits
for health, and inspects the live container for absence of local direct-vector secret paths. The
local overlay retains its live inspect assertion (paths present, values absent), rotation, restart
persistence, and transactional backup/restore drill.
@@ -1,95 +0,0 @@
# Local pgvector Task 1 report
## Status
Implemented the direct `PgVectorStore` behind the transport-neutral `VectorStore` port.
The adapter uses separate optional reader and writer database configurations, derives
capabilities from configured authority, validates strict positive search limits, filters kinds
in SQL before limiting, and merges multi-collection results by cosine similarity.
All collection identifiers are selected from the fixed `schema_records`, `evidence`, and
`memory` allowlist and composed with `psycopg2.sql.Identifier`. Values, vectors, kinds, hashes,
and limits remain bound parameters. Collection/kind mismatches fail with `VectorStoreError`.
Upserts preserve the canonical metadata shape, use `record_key` conflict semantics, update the
transport hash and embedding, and leave semantic metadata fields intact. Health probes reader
and writer independently and reports observed `vector(N)` dimensions against the configured
embedding dimension.
## Configuration and factory
`pgvector_direct` now accepts explicit optional `reader` and `writer` `DatabaseConfig` entries.
The former `connection` entry remains supported as a deprecated read-only compatibility path.
`build_vector_store(..., require_write=True)` accepts writer-only direct configurations and
fails early when no explicit writer is present.
The transitional `build_vector_loader` bulk-sync path remains in place. It uses an explicit
direct writer when present, or the legacy `connection`; it deliberately does not treat a new
reader-only credential as writable. No production schema migration was added.
## TDD and verification
- RED: the new tests initially failed at collection because `PgVectorStore` did not exist.
- Docker L0 pgvector tests: `11 passed`.
- Direct + HTTP parity/factory/config focus: `51 passed`.
- Full harness: `461 passed, 5 deselected`.
- Changed-file Ruff lint: clean.
- Changed-file Ruff format check: clean.
- `git diff --check`: clean.
The repository-wide `ruff check .` still reports 34 pre-existing test-file findings outside
Task 1; none are in changed files. The full pytest suite emits 17 existing legacy-config
deprecation warnings.
## Scope and concerns
- Test fixtures create only the three existing vector tables needed to exercise the adapter;
migration/versioning remains Task 2.
- The legacy single `connection` form stays read-only through the public port, matching its
previous adapter behavior, while remaining available to the explicitly documented bulk-loader
transition.
## Review fix wave
The Task 1 review findings were addressed in a follow-up TDD cycle:
- Search now validates requested kinds against the global known-kind set, intersects valid kinds
with each collection, and skips unrelated collections. A direct-versus-HTTP parity test covers
the multi-collection case.
- Health requires all three allowlisted tables, an `embedding vector(N)` column on every table,
the expected dimension on every table, and the appropriate read or write table privileges for
each configured side. Empty and partial schemas return deterministic, credential-free details;
unexpected database failures expose only their exception class.
- The Docker L0 fixture now provisions separate least-privilege reader and writer roles. Tests
prove the reader cannot insert, the writer cannot execute the cosine-search SELECT, and the
adapter still routes search to the reader and upsert/hash operations to the writer. Direct
upsert uses an atomic `INSERT ... ON CONFLICT DO NOTHING` followed by `UPDATE` for an existing
key, avoiding broad SELECT authority while retaining conflict-safe hash/upsert semantics.
Fresh verification after the fix wave:
- Docker L0 + HTTP port/search parity: `42 passed` (earlier checkpoint); the final L0 file has
`16 passed` including the stricter raw-role search denial.
- Expanded focused adapter/config suite: `56 passed`.
- Full harness: `466 passed, 5 deselected`.
- Changed-file Ruff lint/format and `git diff --check`: clean.
## Sequence privilege health follow-up
Writer health now resolves the real serial/identity sequence for the `id` column of every
required collection using `pg_get_serial_sequence`. It requires `USAGE` on each resolved
sequence, which is the privilege used by the adapter's implicit `nextval`; sequence `SELECT` is
not required because no adapter operation reads sequence state.
The Docker fixture includes a writer role with complete table/hash-column authority but no
sequence grant. Its health is deterministically unhealthy and a new-key upsert fails. Granting
only sequence `USAGE` makes health green and the same port upsert succeeds. Sequence discovery is
guarded for partial schemas so a missing `id` column produces the existing sanitized schema
diagnostic instead of a PostgreSQL error.
Fresh verification for this follow-up:
- Docker pgvector L0 after formatting: `17 passed`.
- Expanded focused adapter/config/parity suite: `57 passed`.
- Full harness: `467 passed, 5 deselected`.
- Changed-file Ruff lint/format and `git diff --check`: clean.
@@ -1,82 +0,0 @@
# Local pgvector Task 2 report
## Outcome
Implemented ordered, idempotent production migrations and the `tht vector migrate`
interface, including `tht vector migrate --status --json` with pristine JSON output.
## Implementation
- `001_extensions.sql` installs pgvector.
- `002_schema_tables.sql` creates `vectors.schema_records`, `vectors.evidence`, and
`vectors.memory` with the `VectorWriteRecord` columns and `vector(768)` embeddings.
- `003_roles.sql` creates passwordless `NOLOGIN` reader/writer roles. Deployments inject
credentials (or grant these roles to separately-created login roles); no production secret
is stored in the repository.
- Reader authority is schema usage plus table `SELECT`.
- Writer authority is schema usage, table `INSERT`/`UPDATE`, narrow hash-probe column `SELECT`,
and sequence `USAGE`. It has no `DELETE`, broad row `SELECT`, DDL, or ownership authority.
- The migration runner discovers ordered SQL files, records SHA-256 checksums in
`public.tht_vector_migrations`, serializes runners with a transaction-scoped advisory lock,
and applies the full pending batch in one transaction.
- Status distinguishes applied, pending, and checksum-drifted migrations. Apply refuses drift.
A failed migration rolls back both prior migrations in that batch and ledger writes.
## TDD evidence
RED was observed with a real `pgvector/pgvector:pg16` testcontainer: 6 failures for the missing
module, missing command, and missing schema.
GREEN verification:
- Focused migration + direct adapter integration: `23 passed`.
- Full harness from the documented `harness/` cwd: `473 passed, 5 deselected`.
- Targeted Ruff (`tht` plus the new L0 test): clean.
- `git diff --check`: clean.
The new L0 coverage exercises clean install, idempotent rerun, pristine JSON status, checksum
drift, transaction rollback, exact tables/columns/dimensions, role isolation, sequence authority,
and the real `PgVectorStore.health()` plus `VectorWriteRecord` upsert path.
## Existing repository lint baseline
The requested full `ruff check .` was run. It reports 34 pre-existing violations in unrelated
test files (unused imports and one-line semicolon statements). None are in Task 2 files; changing
them would exceed this task's scope. The complete harness test gate is green.
## Self-review
No unresolved Task 2 correctness concern found. One deliberate contract choice is worth noting:
writer `INSERT` and `UPDATE` are table-level because the approved direct adapter health probe uses
`has_table_privilege` for those authorities. Least privilege is retained by withholding broad
`SELECT`, `DELETE`, DDL, ownership, and credentials.
## Review fix wave
The post-implementation review found four production-boundary gaps. They are fixed as follows:
- Migration SQL now ships inside the `tht` wheel (`tht/migrations/vector`) via explicit
setuptools package-data and is discovered through `importlib.resources`, rather than relying on
a source-checkout-relative directory.
- Both status and apply reject ledger versions absent from the installed manifest, including
nonnumeric future version labels. This treats a binary/database downgrade as drift instead of
silently reporting a healthy state.
- Migration files are ordered by parsed integer version; spellings such as `2` and `02` are
rejected as duplicate versions.
- Every migration transaction pins `search_path` locally to `pg_catalog, pg_temp`; catalog calls
and the ledger are schema-qualified. pgvector is installed into the locked `vectors` schema,
tables use `vectors.vector`, and `PgVectorStore` qualifies vector casts and the cosine operator.
A hostile admin default path with a writable shadow schema cannot redirect migration objects.
- The core image build asserts CLI discovery. Image verification now starts an ephemeral pgvector
database, runs the installed image's migration command, and compares pristine apply/status JSON.
Additional verification after the fix wave:
- Focused migration, adapter, hostile-path, and wheel suite: `27 passed`.
- Full harness: `477 passed, 5 deselected`.
- Production core image build: passed, including build-time CLI discovery.
- Core-image apply/status smoke against `pgvector/pgvector:pg16`: passed.
- Changed production and test files: Ruff clean; `git diff --check` clean.
- Full Ruff remains at the same 34 pre-existing unrelated test-file findings documented above.
No dependency changed, so the committed Python requirements lock did not require regeneration.
-133
View File
@@ -1,133 +0,0 @@
# Task 3 report — optional local pgvector profile
## Status
Implemented and verified the `local-vector` Compose profile.
- `vector-db` uses pgvector 0.8.5 on PostgreSQL 16, pinned to the official multi-arch
manifest digest.
- `vector_data` is a project-scoped named volume and is not shared with application data.
- database readiness gates the packaged one-shot `vector-migrate` job; core declares the
migration completion dependency while remaining usable in the pre-existing external profile.
- bootstrap, migrator, reader, and writer identities are distinct. Bootstrap and migration
credentials are supplied as Compose secrets; the application receives only reader/writer
credentials.
- `deploy/workspaces/local-vector.yaml` selects `pgvector_direct` with separate reader and
writer connections.
- the base loopback port binding, `AUTH_MODE=none`, and `THOTH_PUBLIC_EXPOSURE=false` defaults
are unchanged.
## Red/green evidence
The initial Compose contract did not list `vector-db`, as required by the brief. The first real
smoke then failed migration 002 because bootstrap installed the vector extension in `public`.
The bootstrap was corrected to create the `vectors` schema under the migration owner and install
the extension there. A clean-volume rerun passed.
## Verification
- `./scripts/local-vector-smoke.sh`: PASS
- isolated generated Compose project and credentials
- clean migration plus idempotent status rerun
- reader/writer privilege health
- one-record upsert and similarity search
- restart of both `core` and `vector-db`
- persisted search result after restart
- project-only volume cleanup
- `./scripts/test-container-deployment.sh`: PASS
- `./scripts/test-backend-url-policy.sh`: PASS
- `docker compose --profile local-vector config --quiet`: PASS
- harness: 477 passed, 5 deselected
- backend: 84 passed; TypeScript typecheck PASS
- frontend: 226 passed; TypeScript typecheck PASS
- `git diff --check`: PASS
## Self-review / concerns
- Compose cannot make a dependency required only under one profile. The core dependency uses
`required: false` so the established `external` profile does not activate local infrastructure;
under `local-vector`, `compose up --wait` still fails if `vector-migrate` exits nonzero, and the
smoke verifies that successful migration precedes the healthy stack.
- Reader/writer passwords are injected into core environment variables because Compose service
attributes cannot be conditional by profile. Bootstrap and migrator credentials remain
file-backed secrets and are never exposed to core.
- The smoke intentionally refuses the operator project name `thothii` and removes only its unique
project namespace and volumes.
## Follow-up hardening — credential reconciliation and cleanup ownership
Review findings were resolved in a separate follow-up:
- Replaced fresh-volume-only initialization with `vector-reconcile`, an idempotent one-shot that
runs after database health and before `vector-migrate`. It authenticates with only the bootstrap
admin secret, safely creates missing identities, reconciles role attributes and passwords on
existing volumes, restores memberships/ownership, and leaves vector data untouched.
- The migrator is explicitly `NOSUPERUSER NOCREATEDB NOCREATEROLE`. Schema/database ownership is
sufficient for all packaged migrations because reconciliation creates the two group roles first.
- The live smoke rotates migrator, reader, and writer secrets on the same populated volume, rejects
the old reader credential, reruns migrations, recreates core with the new runtime credentials,
and retrieves the record written before rotation and again after database/core restart.
- Smoke project names are no longer caller-controlled. Each run creates a unique namespace and
ownership token. Containers, networks, and volumes carry the ownership label; preflight refuses
any collision and cleanup verifies every discovered resource before `down --volumes`.
- Added a dynamic fake-Docker contract suite for caller override, collision, and mismatched cleanup
labels, plus a real-Docker collision probe using a unique labeled volume.
Follow-up verification:
- `./scripts/local-vector-smoke.sh`: PASS, including live secret rotation and persisted retrieval
- `./scripts/test-local-vector-smoke-safety.sh`: PASS
- `./scripts/test-local-vector-smoke-live-collision.sh`: PASS
- harness: 477 passed, 5 deselected
- backend: 84 passed; TypeScript typecheck PASS
- frontend: 226 passed; TypeScript typecheck PASS
- Compose security, backend URL, config, shell syntax, and diff checks: PASS
Remaining operational constraint: the bootstrap admin secret must continue to match the PostgreSQL
bootstrap account stored in the volume. Runtime migrator/reader/writer rotation is supported without
data deletion; bootstrap-account password rotation is a distinct database-administration operation.
## Final hardening — bootstrap account rotation
The remaining operational constraint is now covered by
`scripts/vector-rotate-bootstrap-password.sh OLD_SECRET_FILE NEW_SECRET_FILE`:
- It does not rely on `POSTGRES_PASSWORD_FILE` after initialization.
- It pre-stages the deployment-file replacement in the same directory, authenticates to the live
database with the explicit old file, and changes only the authenticated bootstrap role.
- Passwords are passed as connection parameters and rendered with psycopg2 SQL composition, so
shell and SQL metacharacters are not interpolated.
- A second connection must authenticate with the new password before the command succeeds. If that
verification fails, the still-open old connection restores the old database password.
- Only after verified database login does an atomic rename replace the current deployment secret.
Wrong-old authentication and verification failures leave deployment configuration unchanged.
Final live smoke evidence on one existing `vector_data` volume:
- wrong-old bootstrap rotation rejected; current deployment secret unchanged
- bootstrap password with quote characters rotated successfully
- old bootstrap login rejected and new login accepted
- `vector-reconcile`, packaged migrations, and core health passed afterward
- the vector record written before rotation remained searchable after rotation and after a further
database/core restart
Final tests:
- `./scripts/test-vector-bootstrap-rotation.sh`: PASS
- `./scripts/local-vector-smoke.sh`: PASS with negative and positive live bootstrap rotation
- existing local-vector collision/safety and Compose deployment contracts: PASS
## Final identity and secret-policy alignment
- `THT_VECTOR_BOOTSTRAP_USER` is now passed through core as well as vector-db and reconciliation,
so the rotation helper uses the authoritative configured role instead of defaulting to `postgres`.
- Rotation and reconciliation source the same raw-file `secret-policy.sh`: non-empty and no
whitespace, including trailing newlines. Rotation validates both files before Docker,
PostgreSQL, or atomic replacement staging; `test-vector-secret-policy.sh` pins empty, newline,
internal-space, and valid metacharacter cases.
- Fake-Docker tests prove a non-default identity reaches the helper path and whitespace rejection
performs no Docker call and creates no staged replacement.
- The real smoke runs the entire stack as `thoth_bootstrap_smoke`. Its whitespace-negative case
leaves the deployment file unchanged and proves the existing database login still succeeds;
non-default-account bootstrap rotation, reconciliation, migration, core health, restart, and
persisted retrieval all pass.
@@ -1,94 +0,0 @@
# Local pgvector Task 4 report
## Outcome
Implemented adapter parity gates and an operator-safe custom-format backup/restore workflow.
- Direct and HTTP stores now share validation, configured-dimension rejection, and deterministic
similarity ordering with record ID as the tie-break.
- The parity fixture exercises identical records through real pgvector and the HTTP RPC contract:
kind filtering, ordering, hashes, replacement upserts, invalid collection/kind errors, and query
plus write dimensions.
- Backup explicitly allowlists the three vector tables and migration ledger, refuses overwrite,
writes through a partial file, and uses a custom compressed archive.
- Restore requires explicit active-source and target coordinates. It compares PostgreSQL system
identifier plus database OID (robust across DNS aliases), refuses the active database, checks for
an empty target unless force is explicit, and restores with exit-on-error.
- Passwords are accepted only through validated secret files, converted to private temporary
`PGPASSFILE`s, and never placed in command arguments or success/error logs.
- Role passwords/login identities are deliberately not dumped. The target must have the approved
passwordless group roles and pgvector extension reconciled before restore; archived ACLs restore
the reader/writer grants.
## TDD and semantic alignment
The first parity run exposed the intended HTTP differences: it accepted unknown collections and
wrong dimensions. Direct pgvector also had no stable order for equal cosine distance. The adapters
were aligned, and the final focused real-pgvector gate passed: **25 passed**.
The first recovery run caught an incorrect probe username before restore. The second caught an
intersection between `pg_dump --schema` and the explicit public ledger table. The third confirmed
the archive contents but caught missing target group roles. Each defect was corrected and the
complete drill was rerun from a fresh generated project.
## Live recovery smoke
`./scripts/local-vector-smoke.sh --backup-restore`: **PASS**.
- generated/owned source Compose project and source `vector_data`
- distinct restore container and distinct named restore volume
- migration and role health, secret rotation, restart persistence
- real custom backup, then deliberate mutation of the active source record
- same-database identity guard evaluated before restore
- restore into the separate target only
- restored hash equals the pre-mutation backup, proving retrieval parity
- migration ledger has all three applied versions
- all three restored embedding columns report `vectors.vector(768)`
- ownership-checked cleanup; the active operator project/volume is never addressed
## Verification
- parity + direct adapter: 25 passed
- full harness: 485 passed, 5 deselected
- changed Python files: Ruff clean
- shell syntax: clean
- `git diff --check`: clean
- full Ruff: unchanged repository baseline of 34 unrelated pre-existing test-file violations
## Self-review and operational constraints
The restore account must be able to read `pg_control_system()` for the robust cluster-identity
comparison and create/restore the selected objects. This is intentionally an administrative
recovery operation, not a runtime reader/writer action. `--force-nonempty` is explicit but still
uses `pg_restore --clean --if-exists`; operators should prefer a new database/volume and validate
migration status, health, and known retrieval before endpoint cutover.
## Post-review hardening
All five final review findings were addressed in a follow-up commit:
- Restore now requires a physically separate PostgreSQL cluster and refuses any equal
`system_identifier`, independent of database OID or hostname.
- `pg_restore` combines `--single-transaction` with `--exit-on-error`. The live drill creates an
existing vector sentinel, deliberately fails late during a forced restore, and proves the
original sentinel row/hash remains unchanged before performing the successful restore.
- Backup uses a mode-0600 `mktemp` in the output directory, atomically renames it, and cleans only
that owned path. A fake-command test pins symlink-clobber resistance and preserves an adversarial
legacy `.partial` symlink and its target.
- HTTP parity now traverses the real `VectorRestClient` transport boundary. It asserts RPC URL/key
and kinds payloads, legacy 404 fallback, response conversion, malformed metadata tolerance, and
canonical `VectorRestError` to `VectorStoreError` mapping.
- The restored target runs role/secret reconciliation and a real `PgVectorStore` with separate
reader/writer logins. Health, known-record search, writer upsert, hash probe, schema/table/column/
sequence authority, and 768-dimensional compatibility are therefore verified through the
production adapter. Reconciliation now restores group-role schema `USAGE`, which table-selected
archives cannot carry.
### Atomic no-replace backup publication
The final publication review is also closed. The private same-directory archive is published with
an atomic hard-link create rather than rename-overwrite semantics. If any process creates the final
file or symlink after preflight but before publication, `ln` fails with `EEXIST`, the backup exits
nonzero, the concurrent destination remains byte-for-byte intact, and the trap removes only the
randomly named temporary archive owned by this invocation. The fake `pg_dump` safety test creates
that destination immediately before returning and pins the failure and cleanup behavior.
-929
View File
@@ -1,929 +0,0 @@
# Pre-deployment Fix Wave Report
Date: 2026-07-14
Worktree: `/home/chirone/ThothII/.worktrees/activity-log-cte-layout`
Base: `e5366d14a6da8fb331d94be60b8929cefb1fe3e0`
## Outcome
All three reviewed findings are implemented in one coherent backend/frontend wave:
1. Resume leaves the prior selection, Zustand state, document panel, and EventSource untouched
until `POST /resume` succeeds. Cold Resume changes state and reconnects only after backend
clear/rebind; already-active same-session Resume preserves the existing binding; failure is a
no-op apart from the fixed toast.
2. SSE uses monotonically increasing per-session ids, cursor-filtered replay, native and manual
reconnect cursors, id continuity across `hub.clear`, and descriptor-id pending-gate
idempotence at both backend and frontend layers.
3. Generic Pi system events and readiness errors are projected through explicit public
allowlists. Sentinel URLs, paths, tokens, stderr, commands, and extra fields do not reach HTTP
or SSE.
No harness, workflow, persistence, model, CTE viewer, CTE card, or shared Card file changed.
## Interfaces
- Frontend `resumeSession(id)` now returns
`Promise<{ id: string; alreadyActive: boolean }>` via `ResumeSessionResult`.
- Backend successful Resume always returns the same shape:
- running/waiting runtime: `{ id, alreadyActive: true }`
- validated cold runtime: `{ id, alreadyActive: false }`
- `SseHub.publish(sessionId, event, data): number` returns the assigned SSE id.
- `SseHub.subscribe(sessionId, send, { afterId, pending })` calls
`send(event, data, id)` for replay/live frames with `id > afterId`.
- `GET /sessions/:id/events` accepts native `Last-Event-ID` and manual
`?lastEventId=<integer>`; when both are valid it uses the greater cursor.
- Every emitted SSE frame is `id: <n>\nevent: <name>\ndata: <json>\n\n`.
- Public readiness failure is exactly:
`Session services are not ready. Check configuration and connectivity, then try again.`
- Generic Pi system events are exactly `{ type: "system_event", event }`, and `event` must be a
non-empty string.
## Files
Backend production:
- `backend/src/bridge/session-bridge.ts`
- `backend/src/routes/sessions.ts`
- `backend/src/sse/sse-hub.ts`
Backend tests:
- `backend/test/routes-sessions.test.ts`
- `backend/test/session-bridge.test.ts`
- `backend/test/sse-hub.test.ts`
- `backend/test/sse-route.test.ts` (new)
Frontend production/support:
- `frontend/src/api/sessions.ts`
- `frontend/src/api/types.ts`
- `frontend/src/shell/AppShell.tsx`
- `frontend/src/store/sessionStore.ts`
- `frontend/src/stream/useSessionStream.ts`
- `frontend/src/test/fakeEventSource.ts`
Frontend tests:
- `frontend/src/api/sessions.test.ts`
- `frontend/src/shell/AppShell.session-mgmt.test.tsx`
- `frontend/src/store/sessionStore.test.ts`
- `frontend/src/stream/useSessionStream.test.tsx`
## TDD RED/GREEN evidence
### 1. Backend Resume result and client-boundary allowlists
RED command:
```text
cd backend && npx vitest run test/routes-sessions.test.ts test/session-bridge.test.ts
```
RED output (exit 1):
```text
Test Files 2 failed (2)
Tests 6 failed | 35 passed (41)
expected { id: 's1' } to deeply equal { id: 's1', alreadyActive: false }
expected raw readiness URL/token/path to equal the fixed public message
expected three raw generic system events to equal [{ type: 'system_event', event: 'session_exit' }]
```
GREEN command:
```text
cd backend && npx vitest run test/routes-sessions.test.ts test/session-bridge.test.ts
```
GREEN output (exit 0):
```text
✓ test/session-bridge.test.ts (14 tests)
✓ test/routes-sessions.test.ts (27 tests)
Test Files 2 passed (2)
Tests 41 passed (41)
```
### 2. Backend exact-once SseHub and route framing
RED command:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/sse-route.test.ts
```
RED output (exit 1):
```text
Test Files 2 failed (2)
Tests 7 failed (7)
expected [undefined, undefined, undefined] to deeply equal [1, 2, 3]
expected unconditional replay not to contain "one" / "two"
expected one buffered pending gate, received replay plus a second pending emission
```
GREEN command:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/sse-route.test.ts
```
GREEN output (exit 0):
```text
✓ test/sse-hub.test.ts (4 tests)
✓ test/sse-route.test.ts (3 tests)
Test Files 2 passed (2)
Tests 7 passed (7)
```
### 3. Frontend cursor tracking and gate idempotence
RED command:
```text
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx src/store/sessionStore.test.ts
```
RED output (exit 1):
```text
Test Files 2 failed (2)
Tests 2 failed | 28 passed (30)
expected /sessions/s1/events to be /sessions/s1/events?lastEventId=7
expected duplicate gate pendingWidget to remain null, received gate-1
```
GREEN command:
```text
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx src/store/sessionStore.test.ts
```
GREEN output (exit 0):
```text
✓ src/store/sessionStore.test.ts (23 tests)
✓ src/stream/useSessionStream.test.tsx (7 tests)
Test Files 2 passed (2)
Tests 30 passed (30)
```
### 4. Frontend typed Resume and AppShell ordering/preservation
Typed API RED command:
```text
cd frontend && npx tsc -b
```
Typed API RED output (exit 1):
```text
src/api/sessions.test.ts(43,9): error TS2322: Type 'void' is not assignable to type
'{ id: string; alreadyActive: boolean; }'.
```
Lifecycle RED command:
```text
cd frontend && npx vitest run src/api/sessions.test.ts src/shell/AppShell.session-mgmt.test.tsx
```
Lifecycle RED output (exit 1):
```text
✓ src/api/sessions.test.ts (9 tests)
❯ src/shell/AppShell.session-mgmt.test.tsx (15 tests | 4 failed)
Test Files 1 failed | 1 passed (2)
Tests 4 failed | 20 passed (24)
already-active same-session Resume created two EventSources instead of one
deferred cold Resume closed the document panel before POST completion
failed same-session Resume closed the prior EventSource
failed Resume with no active session opened an EventSource
```
GREEN commands:
```text
cd frontend && npx vitest run src/api/sessions.test.ts src/shell/AppShell.session-mgmt.test.tsx
cd frontend && npx tsc -b
```
GREEN output (exit 0):
```text
✓ src/api/sessions.test.ts (9 tests)
✓ src/shell/AppShell.session-mgmt.test.tsx (15 tests)
Test Files 2 passed (2)
Tests 24 passed (24)
TypeScript: no output, exit 0
```
The AppShell cold-reconnect test additionally proves that the old source accepts an event while
Resume is pending, the replacement URL carries `lastEventId=8`, the replacement receives one
post-resume transcript/activity row, and two deliveries of the same descriptor id yield one gate.
## Affected verification
Backend command:
```text
cd backend && npx vitest run test/routes-sessions.test.ts test/session-bridge.test.ts \
test/sse-hub.test.ts test/sse-route.test.ts test/health.test.ts test/e2e-f1.test.ts
```
Output (exit 0):
```text
Test Files 6 passed (6)
Tests 52 passed (52)
```
Backend typecheck:
```text
cd backend && npx tsc --noEmit -p .
```
Output: no output, exit 0.
Frontend command:
```text
cd frontend && npx vitest run src/api/sessions.test.ts src/store/sessionStore.test.ts \
src/stream/useSessionStream.test.tsx src/shell/AppShell.session-mgmt.test.tsx \
src/shell/CentralStatus.test.tsx src/shell/ModelActivityPanel.test.tsx \
src/shell/f1-loop.test.tsx src/shell/AppShell.new-session.test.tsx
```
Output (exit 0):
```text
Test Files 8 passed (8)
Tests 74 passed (74)
```
Frontend typecheck:
```text
cd frontend && npx tsc -b
```
Output: no output, exit 0.
## Full verification
Backend full suite:
```text
cd backend && npx vitest run
```
```text
Test Files 22 passed (22)
Tests 177 passed (177)
```
Frontend full suite:
```text
cd frontend && npx vitest run
```
```text
Test Files 43 passed (43)
Tests 271 passed (271)
```
Backend production build:
```text
cd backend && npm run build
> tsc -p tsconfig.json
exit 0
```
Frontend production build:
```text
cd frontend && npm run build
> tsc -b && vite build
✓ 4835 modules transformed.
✓ built in 8.25s
exit 0
```
## Integrated re-review closure (2026-07-15)
This section supersedes the earlier cold same-session assertion that the replacement URL carries
`lastEventId=8`. That behavior was correct only while the backend process and its in-memory id
sequence survived. A restarted backend begins a fresh sequence, so a successful cold Resume now
explicitly discards the browser's cursor before replacing the EventSource.
All four integrated re-review findings are closed:
1. `AppShell` passes a dedicated cursor-reset epoch to `useSessionStream`. A cold same-session
Resume increments it only after `alreadyActive: false`; a high cursor such as `901` is omitted
from the replacement URL and fresh low-id events/gates are consumed. An already-active
same-session Resume still preserves its source, cursor, and store.
2. `useSessionStream` no longer mutates the cursor ref during render. Effect setup resets cursor
state on session/reset-epoch changes, callbacks are guarded by a captured active-source
identity, and cleanup clears only its own active identity. A queued event from the replaced
source cannot write the new store or poison its next reconnect URL.
3. Backend Resume is serialized per session and rechecks runtime state inside the lock. Manifest,
readiness, and reopen validation precede the transport commit. Idle/failed replacement creates
and binds the new runtime before `hub.clear`, which occurs synchronously immediately before the
first `Resuming session` publish. Reopen/create failure returns exactly
`Session could not be resumed. Check configuration and connectivity, then try again.`, keeps the
prior hub buffer/subscribers attached, and does not expose exception sentinels. Concurrent calls
perform one cold start and the waiter returns `alreadyActive: true`.
4. `SseHub.forget(id)` removes subscribers, buffered events, and the last id. Permanent session
DELETE invokes it after disk deletion; ordinary close and Resume continue to use `clear`, which
preserves the id sequence.
### Re-review files
Production:
- `backend/src/pi/pi-process-manager.ts`
- `backend/src/routes/sessions.ts`
- `backend/src/sse/sse-hub.ts`
- `frontend/src/shell/AppShell.tsx`
- `frontend/src/stream/useSessionStream.ts`
Tests/support:
- `backend/test/pi-process-manager.test.ts`
- `backend/test/routes-sessions.test.ts`
- `backend/test/sse-hub.test.ts`
- `frontend/src/shell/AppShell.session-mgmt.test.tsx`
- `frontend/src/stream/useSessionStream.test.tsx`
- `frontend/src/test/fakeEventSource.ts`
### Re-review TDD RED/GREEN evidence
Frontend RED command:
```text
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx \
src/shell/AppShell.session-mgmt.test.tsx
```
RED output (exit 1):
```text
Test Files 2 failed (2)
Tests 3 failed | 21 passed (24)
reset epoch: expected the old source to close, received false
cold same-session: expected /sessions/s1/events, received ?lastEventId=901
stale source: expected an empty transcript, received "stale session one"
```
Frontend GREEN command:
```text
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx \
src/shell/AppShell.session-mgmt.test.tsx
cd frontend && npx tsc -b
```
GREEN output (exit 0):
```text
Test Files 2 passed (2)
Tests 24 passed (24)
TypeScript: no output, exit 0
```
Backend RED command:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/routes-sessions.test.ts
```
RED output (exit 1):
```text
Test Files 2 failed (2)
Tests 8 failed | 28 passed (36)
three Resume ordering assertions observed clear before reopen/create
reopen and create sentinels escaped as raw HTTP 500 responses
the concurrent waiter cold-started again instead of returning alreadyActive: true
SseHub.forget was absent and DELETE did not invoke permanent cleanup
```
Backend GREEN command:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/routes-sessions.test.ts
cd backend && npx tsc --noEmit -p .
```
GREEN output (exit 0):
```text
Test Files 2 passed (2)
Tests 36 passed (36)
TypeScript: no output, exit 0
```
The failure tests publish a post-failure probe through the same hub and prove that a subscriber
attached before either reopen or create rejection still receives it. The concurrency test overlaps
two same-id requests behind a deferred reopen and proves one manifest/readiness/reopen/create/clear
sequence.
### Initial re-review verification (before independent-review hardening)
```text
cd backend && npx vitest run
Test Files 22 passed (22)
Tests 182 passed (182)
cd frontend && npx vitest run
Test Files 43 passed (43)
Tests 273 passed (273)
cd backend && npm run build
> tsc -p tsconfig.json
exit 0
cd frontend && npm run build
> tsc -b && vite build
✓ 4835 modules transformed.
✓ built in 8.46s
exit 0
```
`git diff --check` produced no output (exit 0). The frontend build retains its pre-existing
large-chunk warning; no new build or type errors were introduced.
Final whitespace verification:
```text
git diff --check
no output, exit 0
```
## Self-review
- Resume sequencing: reopen and runtime binding precede backend `clear` and HTTP success; frontend
state mutation and cursor-reset epoch follow it. Failure catch only emits fixed UI copy.
- Already active: same-session returns before reset/generation/manifest repaint; different session
resets the single-session store and binds the new id only after success.
- SSE exact-once: ids are transport identity, not content hashes; replay is strictly `id > cursor`;
`clear` retains the counter; pending gate matching uses only descriptor id.
- Cursor behavior: hook tracks `MessageEvent.lastEventId`, carries it only to an ordinary same-id
generation, and resets it on session-id/cold-runtime epoch change. Native EventSource reconnect
remains supported by the route header.
- Gate defense: the Zustand set survives pending clear but resets with the session store.
- Client boundary: raw `ensure.error` is unused in public responses; generic Pi system events are
reconstructed rather than spread; frontend type mirrors the two-field event.
- Scope: `git diff` contains no CTE/Card/harness/workflow/persistence/model changes. Four pre-existing
modified `.superpowers/sdd/{progress,task-2-report,task-3-report,task-4-report}.md` files are user
work and are excluded from staging.
## Remaining concerns
- The 200-event SSE ring limit remains intentional. A brand-new page can reconstruct only retained
backlog; an in-memory same-session reconnect is exact-once from its cursor.
- Per-session sequence counters remain in backend memory after `clear` by design so later in-process
cold same-id Resume cannot reuse ids. Permanent DELETE removes the counter via `forget`.
- The Delete-then-Resume adversarial route test proves the deleted session is not resurrected but
currently receives the runner's generic HTTP 500 when `sessionShow` can no longer find it. A
future API cleanup can normalize that missing-session response to 404 or 409.
- Frontend tests still print pre-existing MSW unhandled-request and React ref/`act` warnings even
though all 276 tests pass. The frontend production build still reports pre-existing large chunk
warnings. Neither warning class was introduced or expanded by this change.
- No live Pi/DWH smoke was run; this wave changes only REST/SSE/frontend lifecycle boundaries and
is covered by fake-Pi, live Fastify SSE, component, full-suite, typecheck, and production-build
gates.
## Independent-review hardening
The required independent review was run repeatedly against the uncommitted diff. Its first pass
found four Important lifecycle edges beyond the integrated findings: queued old-runtime callbacks,
post-spawn construction cleanup, concurrent frontend Resume completions, and the passive-effect
commit window. Its second pass confirmed those fixes and identified one remaining Important
retention issue in the new runtime-identity map. The final pass reported no Critical, Important, or
Minor findings and assessed the diff ready to merge.
The resulting hardening is:
- Runtime bridge callbacks are gated by the bound runtime identity. Replacement, close, and DELETE
invalidate the old identity, so queued old events cannot publish or call `failSession`. An active
runtime removed by the manager can still publish its complete public failure sequence; after the
terminal unmanaged `agent_end`, its binding is released and later events are rejected.
- `PiProcessManager` kills the spawned child and removes any registered map entry if either
spawn-boundary stderr setup or later RPC/bridge/map initialization throws.
- Resume completion compares against synchronously maintained current active-session identity.
Concurrent `alreadyActive: false` then `alreadyActive: true` results preserve the cold source,
cursor, store, and replayed gate.
- Stream source replacement uses a layout effect. A deterministic later-layout-effect test delivers
a queued old event inside the former commit-to-passive-cleanup window and proves it is ignored.
- Cursor tests cover both a restarted backend's fresh low ids and an in-process hub's preserved high
ids followed by a cursor-bearing ordinary reconnect.
### Hardening TDD RED/GREEN evidence
Backend identity/construction RED command:
```text
cd backend && npx vitest run test/pi-process-manager.test.ts test/routes-sessions.test.ts
```
```text
Test Files 2 failed (2)
Tests 3 failed | 69 passed (72)
post-spawn reader initialization did not kill the child
replaced and deleted runtime callbacks still called failSession/published
```
Additional spawn-boundary and terminal-release RED checks:
```text
cd backend && npx vitest run test/pi-process-manager.test.ts \
-t "spawn boundary initialization"
Tests 1 failed | 38 skipped (39)
cd backend && npx vitest run test/routes-sessions.test.ts -t "terminal sequence"
Tests 1 failed | 34 skipped (35)
```
Frontend concurrency/layout RED command:
```text
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx \
src/shell/AppShell.session-mgmt.test.tsx
```
```text
Test Files 2 failed (2)
Tests 2 failed | 25 passed (27)
the later-layout-effect event wrote "commit-window stale text"
the false→true completion pair erased pending gate "cold-gate"
```
Final focused GREEN commands:
```text
cd backend && npx vitest run test/pi-process-manager.test.ts \
test/routes-sessions.test.ts test/sse-hub.test.ts
cd backend && npx tsc --noEmit -p .
Test Files 3 passed (3)
Tests 79 passed (79)
TypeScript: no output, exit 0
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx \
src/shell/AppShell.session-mgmt.test.tsx
cd frontend && npx tsc -b
Test Files 2 passed (2)
Tests 27 passed (27)
TypeScript: no output, exit 0
```
### Final full verification after review hardening
```text
cd backend && npx vitest run
Test Files 22 passed (22)
Tests 188 passed (188)
cd frontend && npx vitest run
Test Files 43 passed (43)
Tests 276 passed (276)
cd backend && npm run build
> tsc -p tsconfig.json
exit 0
cd frontend && npm run build
> tsc -b && vite build
✓ 4835 modules transformed.
✓ built in 8.47s
exit 0
```
The final frontend run retains the repository's pre-existing MSW/ref/`act` warnings, and the build
retains the pre-existing large-chunk warning. No test, typecheck, or build failures remain.
## Stale-bootstrap, lifecycle-lock, and competing-Resume hardening
Date: 2026-07-15
Base: `08b1f4909e8eb7538156cecc2e7a6cafb46ddfc7`
This follow-up closes asynchronous identity/order and multi-client transport gaps found in the
pre-deployment review:
- `PiProcessManager.teardownIfCurrent(id, runtime)` makes teardown an identity-checked operation.
Bootstrap re-checks identity after configuration/retrieval and before both the public
`Starting model` event and model start. Its failure continuation acquires the same session
lifecycle lock, claims only its own runtime identity, and holds serialization through persisted
failure and the public terminal sequence. A continuation left behind by Close or DELETE cannot
target a replacement or recreate forgotten SSE state.
- The former Resume-only promise tail is now a per-session lifecycle lock shared by Resume, Close,
and DELETE. Each route reads the current runtime inside the lock immediately before replacement
or removal and uses identity-checked teardown. Deferred route tests prove both orderings:
Resume then Close/Delete finishes removed with no post-removal bootstrap event; Close then Resume
creates only after Close completes; DELETE then Resume cannot recreate a deleted session.
- AppShell assigns each Resume invocation a monotonic token and records the latest target. A
completion for a different, superseding session id cannot reset the store, select a source, close
the panel, or repaint phase from a late manifest. Same-id invocations are per-target single-flight
operations through the POST and local binding commit: repeated pre-commit clicks update the
shared operation's latest token but issue no second POST or commit path. The operation becomes
joinable again before its manifest fetch, whose repaint remains token/id/selection guarded. Start
new, Stop, streamed session exit, and active-session deletion invalidate pending Resume work.
This prevents stale-source preservation and reverse/non-Resume intent overwrite without allowing
a slow manifest to suppress a later explicit rebind.
- `SseHub` subscriber registrations now carry idempotent transport-close callbacks. `clear` and
`forget` snapshot and actively close every response before discarding runtime transport state;
callback-driven unsubscription during that iteration is safe. The SSE route ends its response so
native EventSource reconnects with `Last-Event-ID`. Post-clear events retain monotonic ids and are
buffered for replay; `forget` additionally resets the id state.
Production files:
- `backend/src/pi/pi-process-manager.ts`
- `backend/src/routes/sessions.ts`
- `backend/src/sse/sse-hub.ts`
- `frontend/src/shell/AppShell.tsx`
Regression tests:
- `backend/test/pi-process-manager.test.ts`
- `backend/test/routes-sessions.test.ts`
- `backend/test/sse-hub.test.ts`
- `backend/test/sse-route.test.ts`
- `frontend/src/shell/AppShell.session-mgmt.test.tsx`
### TDD RED/GREEN evidence
Runtime identity API RED:
```text
cd backend && npx vitest run test/pi-process-manager.test.ts -t "identity-checked teardown"
Test Files 1 failed (1)
Tests 1 failed | 39 skipped (40)
TypeError: mgr.teardownIfCurrent is not a function
```
Runtime identity API GREEN:
```text
Test Files 1 passed (1)
Tests 1 passed | 39 skipped (40)
```
Deferred bootstrap RED:
```text
cd backend && npx vitest run test/routes-sessions.test.ts \
-t "stale bootstrap|bootstrap that"
Test Files 1 failed (1)
Tests 6 failed | 35 skipped (41)
close/delete + replacement: stale continuation removed the replacement runtime
delete without replacement: stale continuation called failSession after forget
```
Deferred bootstrap GREEN:
```text
Test Files 1 passed (1)
Tests 6 passed | 35 skipped (41)
```
Shared lifecycle ordering RED:
```text
cd backend && npx vitest run test/routes-sessions.test.ts \
-t "Resume followed|Close followed|Delete followed"
Test Files 1 failed (1)
Tests 4 failed | 41 skipped (45)
All four deferred assertions observed the competing route settle before the first lifecycle
operation released.
```
Bootstrap plus lifecycle GREEN:
```text
cd backend && npx vitest run test/routes-sessions.test.ts \
-t "Resume followed|Close followed|Delete followed|stale bootstrap|bootstrap that"
Test Files 1 passed (1)
Tests 10 passed | 35 skipped (45)
```
Competing frontend Resume RED:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "competing Resume|stale Resume manifest"
Test Files 1 failed (1)
Tests 2 failed | 16 skipped (18)
reverse POST completion opened a second, stale EventSource
late s1 manifest repainted the selected s3 phase from F3 to F7
```
Competing and same-id Resume GREEN:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "competing Resume|stale Resume manifest|false then true"
Test Files 1 passed (1)
Tests 3 passed | 15 skipped (18)
```
### Independent-review hardening RED/GREEN
The first final review reported no Critical findings and three Important edge cases: bootstrap
could start during an in-progress Close; bootstrap-owned failure was persisted twice; and an older
same-id result could overwrite newer state. The integrated reviewer also required non-Resume
navigation to invalidate pending Resume work. The final main review tightened the same-ID contract
to true single-flight so a second same-target click cannot preserve a dead pre-restart source.
Backend review RED:
```text
cd backend && npx vitest run test/routes-sessions.test.ts \
-t "Close suppresses|bootstrap failure persists once"
Test Files 1 failed (1)
Tests 2 failed | 45 skipped (47)
deferred configure started Pi while closeSession was still pending
bootstrap/public failure called failSession twice
```
Backend review GREEN:
```text
Test Files 1 passed (1)
Tests 2 passed | 45 skipped (47)
```
Same-id single-flight RED:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "share one cold request"
Test Files 1 failed (1)
Tests 1 failed | 18 skipped (19)
two concurrent same-ID invocations issued two cold POSTs (three total including initial activation)
```
Non-Resume invalidation RED:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "starting a new question invalidates"
Test Files 1 failed (1)
Tests 1 failed | 19 skipped (20)
the late Resume opened an EventSource after Start new returned to the landing state
```
Frontend review GREEN:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "share one cold request|competing Resume|stale Resume manifest|starting a new question invalidates"
Test Files 1 passed (1)
Tests 4 passed | 15 skipped (19)
```
Post-commit single-flight lifetime RED:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "releases same-id single-flight"
Test Files 1 failed (1)
Tests 1 failed | 19 skipped (20)
s1 committed and waited on its manifest; after s3 superseded it, a new s1 Resume reused the old
operation and issued no second s1 POST (expected 2, received 1).
```
Same-id and manifest lifetime GREEN:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx -t "same-id|manifest"
Test Files 1 passed (1)
Tests 4 passed | 16 skipped (20)
cd frontend && npx tsc -b
no output, exit 0
```
Multi-client SSE disconnect RED:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/sse-route.test.ts
Test Files 2 failed (2)
Tests 3 failed | 6 passed (9)
clear/forget invoked zero of two registered close callbacks, and two live HTTP SSE responses timed
out instead of reaching EOF after clear.
```
Multi-client SSE disconnect GREEN:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/sse-route.test.ts
Test Files 2 passed (2)
Tests 9 passed (9)
cd backend && npx tsc --noEmit -p .
no output, exit 0
```
The Hub tests use two subscribers whose close callbacks immediately unsubscribe themselves, proving
safe snapshot iteration and exactly-once closure. The live-route test opens two HTTP streams, proves
both receive EOF on clear, publishes a new event and gate, then reconnects after id 1 and replays
exactly ids 2 and 3. The forget test closes both subscribers and proves the next id resets to 1.
Close now removes the observed runtime identity before awaiting persistence. Failure persistence is
claimed once per runtime and lifecycle-serialized; bootstrap's public `session_failed` cannot start
a duplicate. A per-target in-flight map owns the only same-ID POST and commit while its mutable
latest token keeps s1→s2→s1 ordering correct; it is removed immediately after the binding commit,
before awaiting the independently guarded manifest. One shared invalidation helper is called when
active deletion, streamed exit, Start new, or Stop begins.
### Focused verification
```text
cd backend && npx vitest run test/routes-sessions.test.ts test/pi-process-manager.test.ts \
test/sse-hub.test.ts test/sse-route.test.ts
Test Files 4 passed (4)
Tests 96 passed (96)
cd backend && npx tsc --noEmit -p .
no output, exit 0
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
src/shell/AppShell.new-session.test.tsx src/stream/useSessionStream.test.tsx
Test Files 3 passed (3)
Tests 36 passed (36)
cd frontend && npx tsc -b
no output, exit 0
```
### Full verification
```text
cd backend && npx vitest run
Test Files 22 passed (22)
Tests 202 passed (202)
cd frontend && npx vitest run
Test Files 43 passed (43)
Tests 280 passed (280)
cd backend && npm run build
> tsc -p tsconfig.json
exit 0
cd frontend && npm run build
> tsc -b && vite build
✓ 4835 modules transformed.
✓ built in 8.47s
exit 0
```
The frontend suite/build retain the previously documented MSW, React ref/`act`, experimental type
stripping, and large-chunk warnings. No warning class was introduced by this wave. No harness,
workflow, persistence, SQL/CTE viewer, model-selection, or deployment file changed. The four
pre-existing modified `.superpowers/sdd/{progress,task-2-report,task-3-report,task-4-report}.md`
files remain excluded from staging.
### Final independent-review verdict
After the multi-client transport fix, the independent reviewer reported no Critical, Important, or
Minor findings. Its own focused verification passed 96 backend transport/lifecycle tests, 31
frontend Resume/stream tests, both TypeScript checks, and `git diff --check`. Final assessment:
**Ready to deploy: Yes.**
-53
View File
@@ -1,53 +0,0 @@
# DWH REST per-installation authentication SDD progress
Plan: `docs/superpowers/plans/2026-08-20-dwh-rest-per-installation-auth.md`
Branch: `feat/dwh-rest-installation-auth`
Worktree: `/home/chirone/ThothII-next/.worktrees/dwh-rest-installation-auth`
Baseline: workspace docs PASS; Go unavailable on host (use containerized Go 1.26.5); default Compose pre-existing path-sensitive false positive under `/home/chirone`.
Task 1: complete (commits 3bc84b0..1e82fd3, review clean after bounded-digest fix wave).
Task 1 plan note: the dotted-import grep was resolved by the exact module assertion in `09290a0`; `go list -m all` confirms the DWH module only.
Task 2: complete (commits 541ef45, 971a0e6, e90a1a1; independent vet fix d2415b5; review clean after bounded-read, expiry/JSON, same-Store, and cross-Store synchronization fix waves).
Task 3: complete (commits ebb360f, 055dcab, b1079bd; independent review clean after CLI grammar, metadata validation, and ambiguous-publication cleanup fix waves).
Task 4: complete (commits 943f809, 419c344, 134dc19; independent review PASS after legacy multiplicity, fail-closed handler/socket, SGID 2750, O_RDONLY shared lock, OpenReadOnly, and safe socket-parent waves). Task 5 must precreate `.writer.lock` as `0640 root:dwh-auth`; add a non-owner group/cross-process integration proof when packaging permits.
Task 5: complete (commits 87606c7, 62ec29f; independent review PASS after auth-subrequest header isolation and target-Nginx duplicate-header verification). No runtime installation or service/Nginx mutation performed.
Task 6: complete (commits 1b18a0f, d0f7e04; independent review PASS after regex/duplicate bypass, exact-PID TCP, and bounded-cleanup hardening). Minor for final review: remove or rename the redundant legacy `negative_postgrest_bypass` fixture and align the historical fixture-count prose if useful. No active Nginx/runtime mutation performed.
Task 7: complete (commits 7b9b8b3, 707c13d, f616aab, 7b86ea9, 7fe5316; Terra review PASS). Documentation and rollout-contract alignment completed; no runtime mutation performed.
Task 8: complete at frozen SHA `6499d24892b4383ac492579303e766cfb51fe44e` (fix commits `09290a0`, `6499d24`; Terra review PASS). Focused, portability, scanner, DWH Go, and two full tools/tht matrix runs PASS; evidence recorded at `.artifacts/dwh-auth/source-verification.md`. Broad coupling remains `BASELINE_RED` debt; immutable paths remain unchanged.
Task 9: PASS at frozen SHA `0c4ff3750d3ecd3fc514e50e511cf7475fbe0446`. Built and installed the exact local candidate, enabled and started `dwh-auth`, imported the protected legacy credential under public ID `legacy-shared`, and created `psd-mac-primary` with public key ID `oNPdOfoH7ypLtVb1`. Registry check, AF_UNIX-only listener, v1/legacy `204`, random/missing `401`, bounded journal scan, installed-file hashes, and `nginx -t` all PASS. Protected report: `/root/dwh-auth-provision/gate9-20260821T054514Z.report.md`, SHA-256 `65ce1e0d8be5f74355eca2b1dca901da16f2864f68eafb7dede9d23ef36b82d5`. Nginx was not changed or reloaded; the legacy key remains active; the old stack was not changed or stopped. Terra final review: PASS with no Critical or Important findings.
Task 10: NOT STARTED and requires a second explicit authorization. Public Nginx cutover, Mac-key delivery/configuration, and legacy revocation have not occurred. Activity 1 remains `IN_DISCUSSION`; external deployment remains `SURVEY_NO_GO`.
Final clarification (bookkeeping): initial authorization at `6499d24` stopped before installation because the protected legacy file was missing and a journal-scan finding remained. A secret-safe legacy file was prepared without emit/hash; Nginx metadata remained unchanged and `nginx -t` PASS. Fix commits `6fb4886`, `dee0f9c`, `0c4ff37` received Terra PASS, followed by a detached complete re-freeze PASS at full `0c4ff3750d3ecd3fc514e50e511cf7475fbe0446`. The owner then explicitly authorized Gate 9 at that exact SHA; Gate 9 completed as recorded above. Task 10 remains a separate gate.
## Project A authentication runtime projection
Plan: `docs/superpowers/plans/2026-08-21-project-a-server-auth-runtime-projection.md`
Plan commit: `64f46c7019e11a74dae35a7cdb447cd881e17061`
Implementation baseline: `64f46c7019e11a74dae35a7cdb447cd881e17061`
Runtime constraint: source, synthetic tests, and documentation only; Project A, `/srv`, Nginx,
the legacy stack, and shared services remain untouched.
Task 1: complete (commits `8a8f2c2`, `f9e2950`, `de86760`; independent Terra review PASS after descriptor-relative rewrite, full-history validation, crash recovery, destructive replacement guards, deterministic failure seams, and interrupted-retention recovery).
Task 2: complete (commits `05f8615`, `da2f4a6`, `1e2c4e6`; independent review PASS after exact GID enforcement, bounded descriptor-bound namespace enumeration, strict trailing-slash parity, OIDC coverage, and complete one-retry `CURRENT` publication linearization).
Task 3: complete (commit `903c0b4`; independent Terra review PASS after retained-FD outer locking, cancellable runtime/canonical waits, public transaction-context propagation, deterministic swap/metadata/creator-race tests, and fail-closed pre/post-commit error handling).
Task 4: complete (commit `3d9a9f0`; independent Terra review PASS after moving the Linux/root restore gate before secret-bearing checkpoint creation). Exact focused Go, race, vet, Node 24 focused/full, TypeScript, projected Compose, secret-policy, Windows backup compile, and full serialized tools/tht gates PASS. Historical canonical/unified Compose failures were reproduced as baseline-only documentation/path coupling failures and were not weakened.
Task 5: complete (commit `ef7ae70`; independent Terra review PASS after read-only Vitest gate repair,
remote-Docker/context hardening, and adversarial canonical-mount/`sudo printenv` verifier fixes).
All 14 cross-layer acceptance cases, documentation verifiers, shell syntax, full Go/race/vet,
backend Node 24 Vitest/typecheck, projected Compose, and secret-policy gates PASS. The three known
default/canonical/unified Compose policy failures were reproduced at `a21e2c1` and remain
unmodified baseline debt. Project A has not been started; applying the descriptor or any runtime
root under `/srv/thothii` still requires a new explicit authorization.
Final Project A source review: PASS for `a21e2c1..ef7ae70`; independent Terra review found no
remaining Critical or Important issue after validating restore admission/order, backend path and
identity controls, transaction cancellation, lifecycle gating, portability, dependency scope, and
redaction. Pre-live stop boundary remains in force.
-325
View File
@@ -1,325 +0,0 @@
# Task 2 report — protected atomic registry
## Scope and commit
- Commit: `541ef45 feat: add protected DWH credential registry`
- Committed files only:
- `tools/dwh-auth/internal/securefile/securefile_linux.go`
- `tools/dwh-auth/internal/securefile/securefile_linux_test.go`
- `tools/dwh-auth/internal/registry/store.go`
- `tools/dwh-auth/internal/registry/store_test.go`
- No server, Nginx, systemd, Docker stack, real registry, secrets, or legacy ThothII files were
read or changed. Tests use `t.TempDir` and synthetic record digests only.
## TDD evidence
All Go commands ran in the required official `golang:1.26.5` container with only this linked
worktree bind-mounted at `/work`. The container image reports `go version go1.26.5 linux/amd64`.
### RED
Before either Task 2 production file existed, the focused command was run inside the container:
```text
go test ./internal/securefile ./internal/registry -count=1
```
It failed non-zero for the expected absent implementation symbols, including `undefined: OpenDir`,
`undefined: ReadSecret`, `undefined: Open`, `undefined: State`, `undefined: PublicRecord`, and
`undefined: Store`.
### GREEN
After the minimal implementation and formatting:
```text
go test ./internal/securefile ./internal/registry -count=1
```
Result:
```text
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry
```
### Race verification
The required race command completed successfully:
```text
go test -race ./internal/securefile ./internal/registry -count=1
```
Result:
```text
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 1.027s
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 1.179s
```
Additional scoped verification:
```text
go vet ./internal/securefile ./internal/registry
go test ./... -count=1
git diff --cached --check
```
The Task 2 vet command completed with no findings; all four DWH-auth packages passed the full
module test run; the staged-diff check completed with no output.
## Delivered behavior
- `securefile` is Linux-only and traverses absolute paths through descriptor-anchored
`syscall.Open`/`Openat` calls with `O_NOFOLLOW|O_CLOEXEC`; protected roots, child directories,
records, and secret files are regular/directories only and are checked against `Lstat` after
`Fstat`.
- Protected reads reject special, group-writable, or world-writable modes, cap record reads at
4096 bytes, read at most one extra byte, and reject file-size changes or short/partial reads.
Secret ingress additionally requires exact `0600`.
- Secret output uses `O_CREAT|O_EXCL|O_NOFOLLOW`, exact `0600`, and an absolute protected parent.
- `registry.Open` creates protected `active` and `revoked` subdirectories under an existing safe
root. Record enumeration rejects unexpected entries, unsafe files, symlinks, oversized files,
bad filenames, malformed JSON, unknown JSON fields, duplicate JSON fields, and trailing JSON.
- `Add` validates Task 1 records, writes canonical JSON plus one newline through an exclusive
temporary file, sets final mode `0640`, syncs the file, renames under a protected per-root writer
lock, then syncs the directory.
- `Revoke` writes and syncs a valid revoked record before unlinking and syncing the active record.
`Find` checks revoked first; `List` resolves an active/revoked overlap to the revoked public
record. `FindLegacy` scans fail-closed and permits only the reserved legacy record state.
- `PublicRecord` deliberately omits `secret_sha256`; the redaction is regression-tested.
## Security-test coverage
- protected normal files and canonical record publication;
- symlinked roots, registry directories, records, and secret input;
- unsafe root/directory/record/secret modes;
- bounded/oversized record input;
- unknown, duplicate, trailing, and partial JSON;
- filename mismatch and multiple legacy-record integrity failures;
- revoked-state precedence when both active and revoked files exist;
- concurrent adds and concurrent reads during revocation, including the race detector.
## Self-review
Reviewed all syscall, path, mode, and error paths after the final race run:
- Directory traversal never follows a supplied component; later operations use retained directory
descriptors, not re-opened untrusted prefixes.
- `Fstat` validates the opened object and `Lstat` must identify the same inode/device; the direct
child name grammar refuses separators, dot components, and NUL.
- File validation occurs before and after reads; mode/type/size checks fail closed. Directory
listing obtains a fresh `openat(dirfd, ".")` descriptor so scans do not share a mutable directory
offset.
- Writer serialization protects the check-then-rename no-replace sequence. Failed temporary
cleanup leaves an unexpected entry that later scans reject rather than silently accepting it.
- State-specific validation rejects revocation metadata in active records and requires it in
revoked records. Revoked files are consulted before active files so interruption after revoked
publication cannot reactivate a credential.
- All functionality uses only Go standard-library packages and Linux `syscall`; no CGO, SQLite,
or third-party module was added.
## Concerns
- The optional whole-module `go vet ./...` reports a pre-existing Task 1 test warning at
`internal/credential/credential_test.go:86` (`append` with no variadic values). The identical
line is present in approved HEAD `1e82fd3`, outside this task’s authorized files. Focused Task 2
vet passes, and all module tests pass.
- The official image's login shell resets `PATH` and hides `/usr/local/go/bin`; all evidence uses
direct `go`/`gofmt` container entrypoints, which preserves the image’s Go 1.26.5 environment.
- The generic `apply_patch` helper intermittently failed before file access with a sandbox network
namespace error. Exact scoped corrections were applied through the shared worktree workflow;
this did not affect the final staged file set or verification evidence.
## Review remediation — 2026-08-21
### Scope and fix commit
- Review-fix commit: `971a0e6 fix: harden DWH credential registry reads`.
- Committed files only:
- `tools/dwh-auth/internal/registry/store.go`
- `tools/dwh-auth/internal/registry/store_test.go`
- The separate Task 1 vet correction is the independent preceding commit `d2415b5`; it is not
included in this Task 2 fix commit. No filesystem primitive, server, Nginx, service, registry,
secret, Docker stack, or legacy ThothII file was changed.
### Strict TDD evidence
All commands again used the official `golang:1.26.5` image with only this linked worktree mounted
at `/work`.
#### RED
The first focused command was run after the new regression tests and before production changes:
```text
go test ./internal/securefile ./internal/registry -count=1
```
It failed as intended. The three case-variant aliases (`SECRET_SHA256`, `Secret_SHA256`, and
`Schema_Version`) were accepted; past expiry returned active records from both `Find` and
`FindLegacy`; a revocation snapshot let readers return active data before publication; a temporary
file let `List`, `Check`, and `FindLegacy` observe false integrity failures; and the original
concurrent-read regression observed `ErrNotFound` during revocation.
The deterministic exact-expiry test was then added before the clock implementation. Its focused
run failed as intended with:
```text
internal/registry/store_test.go:572:10: store.now undefined
```
#### GREEN and verification
After the minimum implementation and `gofmt`:
```text
go test ./internal/securefile ./internal/registry -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 0.014s
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 0.684s
go test -race ./internal/securefile ./internal/registry -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 1.022s
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 1.711s
go vet ./internal/securefile ./internal/registry
```
The focused vet output was empty (success). Additional final checks passed:
```text
go test ./... -count=1
ok internal/credential
ok internal/record
ok internal/registry
ok internal/securefile
go vet ./...
go test -race ./internal/registry -count=10
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 8.127s
git diff --cached --check
```
### Remediated security invariants
- `Find` and `FindLegacy` now deny an active record when `ExpiresAt <= now.UTC()`, returning the
existing non-disclosing `ErrNotFound`. The unexported per-Store `now` function is the minimal
deterministic clock seam; past, exact-equality, and future cases are covered for v1 and legacy
records. A revoked record is still consulted before expiry and therefore remains authoritative.
- One per-Store `sync.RWMutex` creates an in-process consistent snapshot. `Add` and `Revoke` hold
it exclusively for their full writer-lock lifetime, including temporary-file publication and
revoked-then-active removal. `Find`, `FindLegacy`, `List`, and `Check` hold a shared lock; their
bodies delegate only to unlocked helpers, preventing nested-lock deadlocks. `Close` also takes
the exclusive lock before closing descriptors.
- Deterministic regression tests hold the writer path at the revocation publication/unlink and
temporary-file stages. They prove public readers wait, then see either the final revoked state or
a clean directory, eliminating the `Names`-to-load/unlink and temporary-entry false failures
within the Store contract.
- Before struct decoding, the outer record JSON object now requires exactly spelled keys from the
schema allowlist and rejects duplicate literal keys. `Decoder.DisallowUnknownFields`, recursive
duplicate detection, trailing-value rejection, record validation, filename matching, no-follow
reads, modes, and durability ordering remain intact.
### Self-review and concerns
- Reviewed the new lock boundaries, error returns, revoked-first ordering, clock fallback,
JSON-token consumption, and every unchanged `securefile` syscall/path/mode boundary. The change
adds only standard-library `sync`; it does not relax existing fail-closed behavior.
- The synchronized snapshot is intentionally per `Store`, matching the requested in-process
contract. The existing protected advisory lock continues to serialize writers across Store
instances/processes; no cross-process reader snapshot is claimed by this fix.
- The historical whole-module vet concern in the original Task 2 report is now resolved by the
independent Task 1 commit `d2415b5`; complete module vet passes in the final evidence above.
## Cross-Store snapshot remediation — 2026-08-21
### Scope and TDD evidence
This third Task 2 fix wave changes only the protected lock primitive and registry snapshot code:
- `tools/dwh-auth/internal/securefile/securefile_linux.go`
- `tools/dwh-auth/internal/securefile/securefile_linux_test.go`
- `tools/dwh-auth/internal/registry/store.go`
- `tools/dwh-auth/internal/registry/store_test.go`
All commands used the official `golang:1.26.5` image with only this linked worktree mounted at
`/work`.
The test-only red patch initially tried to inspect the unexported `securefile.Dir.fd` through the
registry package and therefore did not compile. That assertion was removed without production
changes: the registry tests still create writer Store A and reader Store B through two independent
`Open(root)` calls, while the securefile test proves separate descriptors directly in its own
package. The subsequent behavioral RED run, before the production change, was:
```text
go test ./internal/securefile ./internal/registry -count=1
FAIL TestLockSharedAllowsReadersAndBlocksExclusiveWriter: Dir lacks shared advisory locking
FAIL TestCrossStoreReadersWaitAcrossRevokePublicationAndUnlink:
Find, List, Check, and FindLegacy completed during Store A's revocation snapshot
FAIL TestCrossStoreScanReadersWaitForWriterTemporaryFile:
Store B's List, Check, and FindLegacy observed `.tmp-regression`
```
### Delivered synchronization contract
- `securefile.Dir.LockShared` now acquires `LOCK_SH` on the same protected, no-follow, exact-0600
root lock file used by `Lock`, which continues to acquire `LOCK_EX`. The lock file is still
opened/created, mode-validated, inode-checked, and closed through the existing Linux syscall
path.
- Every public snapshot reader (`Find`, `FindLegacy`, `List`, and `Check`) takes its Store
`RLock`, then a shared advisory lock on root `.writer.lock`, and retains both through the whole
revoked/active lookup or directory scan/load. `Add` and `Revoke` retain Store `Lock`, then the
same root lock under `LOCK_EX`, over their full operation.
- The lock order is universally Store mutex then root advisory lock. Public methods delegate only
to unlocked helpers, so neither reader nor writer paths recursively acquire the Store mutex.
`Close` retains its exclusive Store mutex, preventing descriptor closure from racing any locked
reader or writer.
- Revoked-first precedence, expiry denial, exact JSON validation, no-follow checks, record modes,
temporary-file durability, and all previous behavior remain unchanged.
### GREEN and repeated verification
```text
go test ./internal/securefile ./internal/registry -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 0.019s
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 0.711s
go test -race ./internal/securefile ./internal/registry -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 1.032s
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 1.748s
go vet ./internal/securefile ./internal/registry
go test ./... -count=1
ok internal/credential
ok internal/record
ok internal/registry
ok internal/securefile
go vet ./...
go test -race ./internal/registry \
-run 'TestCrossStoreReadersWaitAcrossRevokePublicationAndUnlink|TestCrossStoreScanReadersWaitForWriterTemporaryFile' -count=20
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 9.880s
go test -race ./internal/securefile \
-run TestLockSharedAllowsReadersAndBlocksExclusiveWriter -count=20
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 1.063s
git diff --check
```
Both vet commands and the whitespace check produced no output. The securefile regression opens
three protected directory descriptors, proves they are distinct, permits two independent shared
holders, proves a third descriptor cannot take `LOCK_EX|LOCK_NB`, then proves exclusive acquisition
succeeds after shared release. The registry regressions deterministically block Store B readers
while Store A holds the exclusive root lock and verify only final revoked/clean states afterward.
### Self-review and concerns
- Reviewed lock creation/reopen races, no-follow flags, exact lock-file mode validation, lock
release, descriptor lifetime, lock ordering, error wrapping, and the unlocked-helper call graph.
No public reader invokes another public reader or writer while holding a Store lock.
- Advisory synchronization necessarily covers cooperating registry Store instances/processes;
arbitrary external filesystem mutation remains fail-closed through the existing integrity
checks rather than being silently accepted.
- No known concerns within the registry's cooperating-process contract.
-129
View File
@@ -1,129 +0,0 @@
# Task 3 report — secret-safe dwh-auth administrative CLI
## Scope
- Added `tools/dwh-auth/internal/command/command.go`, its command tests, and
`tools/dwh-auth/cmd/dwh-auth/main.go`.
- The CLI accepts only the frozen Task 3 grammar: key create/import/list/status/revoke,
registry check, and the reserved serve invocation.
- Existing Task 1–2 APIs are consumed without modifying their files.
- No server, Nginx, systemd, Compose, portable `tht`, real registry, real secret, or legacy
stack was accessed or changed.
## TDD evidence
Tests were written before `Run` existed. In the official `golang:1.26.5` container, mounted
against only the dedicated worktree, the focused RED run was:
```text
go test ./internal/command -count=1
internal/command/command_test.go:214:10: undefined: Run
FAIL
```
After implementation and formatting:
```text
go test ./internal/command -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/command
go test -race ./internal/command -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/command
go test ./... -count=1
ok internal/command, internal/credential, internal/record, internal/registry, internal/securefile
go test -race ./... -count=1
ok internal/command, internal/credential, internal/record, internal/registry, internal/securefile
go vet ./...
```
## Contract coverage
- Create generates the Task 1 canonical credential, writes it once through the protected
exclusive `0600` output primitive, syncs/closes it before registry publication, and emits
only `created key_id=... installation_id=... output=...`.
- Existing output is never overwritten. Publication failure attempts compensating removal;
cleanup uncertainty returns exit 4 and reports only the output path.
- Legacy import requires `--legacy-raw`, the reserved `legacy-shared` installation ID, and an
absolute exact-`0600` source. It verifies the opaque value, never changes its source, and
stores only its digest.
- List/status expose `PublicRecord` data only; JSON is written as pristine JSON with no digest.
Revoke requires a non-empty reason and reports only its public key ID.
- Relative paths, malformed/unknown flags, duplicate options, invalid IDs/metadata/expiry,
and missing required values return exit 2. Missing status/revoke keys return exit 3.
Registry/filesystem/integrity failures return exit 4.
- Diagnostics are fixed redacted strings. Tests use a sentinel secret and assert it is absent
from stdout/stderr, list/status/check JSON, import output, and unsafe/integrity failures.
## Secret-redaction evidence
The command never prints credential contents or digests. It does not use environment fallback,
interactive stdin, or `flag` diagnostics that echo argument values. The sentinel appears only
in synthetic temporary test input and an integrity-fixture file; all command output assertions
confirm it is absent. The registry’s existing `PublicRecord` contract omits `secret_sha256`.
## Concerns
- `serve` is grammar-reserved and returns a redacted exit-4 unavailable response; Task 4 owns
the Unix-socket service implementation and will wire this dispatch.
- The earlier concern about UTF-8 metadata hardening is superseded by `055dcab`: metadata now
rejects invalid UTF-8 and Unicode controls before generation/output. The nil-safe cleanup note
remains non-blocking and outside this review wave.
## Review-fix wave
Review findings were addressed in separate commit `055dcab`. Regression tests were added first. The focused RED run in the official Go 1.26.5 container failed on intentionally absent seams:
```text
undefined: nowUTC
undefined: addRecord
undefined: closeStore
FAIL github.com/aritmolab/thothii/tools/dwh-auth/internal/command
```
The fix rejects embedded canonical v1 credentials in description/revocation reason without echoing metadata, validates UTF-8/Unicode controls and expiry against one captured UTC creation time before generation/output, reserves exactly `serve --registry-root ABS --socket ABS`, and makes publication cleanup depend on a definitive registry lookup. Output is retained after publication or close ambiguity, with path-only recovery guidance.
Review-fix verification in Go 1.26.5:
```text
go test ./internal/command -count=1 PASS
go test -race ./internal/command -count=1 PASS
go test ./... -count=1 PASS
go test -race ./... -count=1 PASS
go vet ./... PASS
git diff --check PASS
```
New tests cover synthetic canonical credentials embedded with prefix/suffix, invalid UTF-8, C1 Unicode controls, past/equal/future expiry, exact serve ordering, deterministic pre-/post-publication and close-failure seams, and sentinel absence from stdout/stderr/list/status JSON.
## Cleanup snapshot review-fix wave
The second re-review added two regression tests before implementation. The RED run in the
official Go 1.26.5 container showed the old `Find` proof incorrectly treated both cases as
cleanup-safe:
```text
FAIL TestCreateRetainsOutputWhenSnapshotFindsUnrelatedIntegrityFailure
corrupt snapshot result = (4, "", "integrity failure\n")
FAIL TestCreateRetainsOutputWhenFailedPublicationRecordIsExpired
expired publication result = (4, "", "integrity failure\n")
```
Commit `b1079bd fix: retain DWH key output on ambiguous publication` replaces the `Find` proof
with a complete `Store.List()` snapshot. It removes generated output only when the snapshot
succeeds, the generated key ID is absent, and `Store.Close()` succeeds. Any unrelated integrity
error, active/revoked/expired record, or close error retains the output and emits only path-based
recovery guidance. The clean pre-publication failure path still removes the output.
Final cleanup-wave verification in Go 1.26.5:
```text
go test ./internal/command -count=1 PASS
go test -race ./internal/command -count=1 PASS
go test ./... -count=1 PASS
go test -race ./... -count=1 PASS
go vet ./... PASS
git diff --check PASS
```
-60
View File
@@ -1,60 +0,0 @@
# Task 4 — Unix-socket DWH verification report
## Scope
Implemented the standalone Linux verifier at `tools/dwh-auth/internal/service` and wired the exact command:
```text
dwh-auth serve --registry-root ABSOLUTE_CANONICAL --socket ABSOLUTE_CANONICAL
```
The service accepts only `GET /verify`. It returns empty `204` responses with `X-DWH-Key-ID` for verified v1 or reserved legacy credentials; credential failures are generic empty `401` responses, and registry/integrity faults are empty `503` responses. Other paths/methods return empty `404`/`405`.
## Security decisions
- Exactly one `X-API-Key` header, maximum 128 bytes.
- Strict `thtdwh_v1.` parsing precedes legacy lookup; non-v1 values alone may use the reserved legacy record.
- Registry integrity is checked before every verification request, so unrelated malformed/unsafe records fail closed with `503`.
- Logs emit only timestamp, decision code, and (when safely parsed or verified) public key ID; test sentinels prove no key, digest, description, or query value is emitted.
- `serve` validates canonical absolute paths, performs startup `Store.Check`, and reports service startup errors as non-secret `integrity failure`.
- Socket collisions that are regular files, directories, symlinks, live sockets, or foreign-owned stale sockets are refused. Only an owned stale Unix socket after `ECONNREFUSED` can be reclaimed.
- Published sockets are mode `0660`; cancellation calls graceful shutdown and removes only a revalidated same-device/same-inode owned socket. A test seam proves a changed path is retained rather than unlinked.
## Required supporting security fix
Commit `943f809` (`fix: reject duplicate legacy DWH records`) tightens the Task 2 registry contract: a synthetically valid active plus revoked legacy pair is now an integrity failure. It is intentionally separate from the Task 4 commit.
## TDD evidence
RED was observed for the missing handler, listener/configuration API, CLI wiring, unrelated-registry corruption, active+revoked legacy state, and cleanup replacement race. Each increment was then implemented minimally and rerun GREEN.
## Verification
All commands were executed in official `golang:1.26.5`, with only this worktree mounted:
```text
gofmt -w cmd internal/command internal/service
go test ./internal/service ./internal/command -count=1
go test ./... -count=1
go test -race ./... -count=1
go vet ./...
git diff --check
```
All passed. A dependency scan also found no third-party Go dependencies.
## Scope boundary
No Nginx, systemd, real Unix socket, real registry, credential, legacy stack, or external service was changed. All test data was synthetic and temporary.
## Follow-up hardening: runtime read-only registry and socket parent
The Task 5 storage contract uses `root:dwh-auth` SGID directories (`2750`) and a service account with read-only group access. The original registry reader path was incompatible because shared locks were opened `O_RDWR` and lazily created as `0600`; secure-directory validation also rejected SGID.
The runtime path now uses `registry.OpenReadOnly`: it opens only preprovisioned root, `active`, `revoked`, and `.writer.lock` paths, and rejects `Add`/`Revoke`. The administrative `Open` path bootstraps the lock through the exclusive writer path. Shared lock acquisition opens the existing `root:dwh-auth 0640` lock `O_RDONLY` with `LOCK_SH`; writer acquisition remains `O_RDWR` with `LOCK_EX`, preserving cross-process snapshot exclusion. Secure directories allow SGID but still reject setuid, sticky, group-write, and world-write bits.
Task 5 must create `.writer.lock` as `0640 root:dwh-auth` alongside the `2750 root:dwh-auth` registry directories before the service starts.
The socket parent must be a canonical non-symlink directory owned by the service EUID and not group/world writable. This removes the bind-to-chmod and path-replacement exposure from other principals. The remaining POSIX path race is bounded to trusted processes sharing the service EUID inside that non-contendible parent.
Additional verification (official `golang:1.26.5`, worktree only): focused securefile/registry/service/command tests, full tests, full race tests, vet, plus ten race repetitions each for cross-store snapshot readers, `OpenReadOnly`, and listener tests: all PASS.
-71
View File
@@ -1,71 +0,0 @@
# Task 5 report — backend principal enforcement
## RED
Added backend route/auth tests before implementation. The initial focused run failed in
seven new assertions: `getPrincipal` did not exist, upstream requests still required the
legacy identity header, foreign session/SSE routes were not hidden, admin scope was not
enforced, new sessions had no trusted principal binding, and settings were global.
## GREEN
- Focused backend suite: `66 passed` across auth, sessions, SSE, and settings tests.
- Complete backend Vitest suite: `209 passed` across `22` files.
- `npx tsc --noEmit -p .`, `npm run build`, `git diff --check`, and changed Python
source Ruff all exit successfully.
- Harness targeted repository/local/migration tests and Python bytecode compilation exit
successfully. The new `tht session preferences get|set` commands are registered and
expose the expected Typer help. A direct local CLI preference smoke was not run because
the checked-in local workspace requires unavailable `THT_DB_HOST` configuration.
## Route and child-process coverage
- `GET /me` returns the request `PrincipalContext`; upstream accepts only the portal's
normalized `X-Thoth-*` identity tuple, with the legacy header ignored. Local mode uses
the same stable `THT_HOME`/`~/.thothii/identity.json` UUID contract as the harness.
- All session operations are principal-scoped: list (`mine` and admin-only `all`), show,
create, resume, close, delete, rename, group, archive, unarchive, documents, reviewer
response, steer, SQL preview/export, and SSE. Missing and foreign sessions are 404;
absent upstream identity is 401. SSE is authorized before response headers or hub
subscription, so a rejected request cannot attach to a live stream.
- New/resumed Pi runtimes and every route-spawned `tht` process receive
`THT_PRINCIPAL_ISSUER`, `THT_PRINCIPAL_SUBJECT`, optional display name, and admin flag.
The readiness `tht` child is also principal-bound.
- Settings use asynchronous repository-backed `tht session preferences get|set` in the
production runner, which isolates preferences by principal. The legacy settings file is
retained only as an injected-runner compatibility fallback for existing isolated tests.
- Repository/settings authorization failures map to 503 before model startup. SQL execution
errors remain 500 after authorization, preserving the prior API distinction.
## Self-review and concerns
- Confirmed the Task 4 portal emits lowercase `true`/`false` for the admin header; the
parser accepts that exact normalized form plus the repository's existing `1`/`0`
compatibility form, and rejects all other values.
- The harness principal resolver is the ownership authority; the backend never accepts an
owner supplied in request bodies. Its route guards use a repository-scoped `session show`
before every session resource operation.
- Existing dependency-injected route fakes without `sessionShow` retain a narrow test seam;
production `ThtRunner` always has that method, so deployed requests cannot bypass the
repository authorization check.
## Review follow-up
### RED
Focused regressions initially failed exactly at the three review findings: stale ambient
display names survived into both `tht` and Pi child environments; mutation/document runner
methods dropped the selected workspace; and `expandLocalHome` did not exist.
### GREEN
- Child environments now remove all four `THT_PRINCIPAL_*` keys from their cloned base
environment before applying the exact request principal. Regression tests prove an absent
display name does not inherit a stale ambient value in either child path.
- `setName`, `setGroup`, `archive`, `unarchive`, and `documents` now take and retain an
optional workspace. The rename route regression proves `session show` authorization and
the mutation use the same non-default workspace.
- Local principal paths expand `~`/`~/...`; existing local home and identity file modes are
repaired to POSIX `0700`/`0600` when applicable, with Windows left unchanged.
- Focused suite: `74 passed`; full backend suite: `213 passed` across `22` files, followed by
TypeScript typecheck, production build, and diff check.
-298
View File
@@ -1,298 +0,0 @@
# Task 6 — Frontend identity and administrator UX report
## RED
- Added API tests for the `/me` principal call and `mine`/`all` session-list scopes.
- Added component tests for regular-user scope, admin scope switching, owner labels,
administrator banner, foreign-owner delete confirmation, and foreign-owner archive
confirmation.
- Initial focused run: 7 expected failures (missing `getMe`, missing scope query,
missing owner label/admin controls, and missing foreign-action confirmation).
- The archive-confirmation regression was also run separately before its implementation
and failed because `window.confirm` was not called.
## GREEN
- `npx vitest run src/api/sessions.test.ts src/shell/NavSessions.test.tsx src/shell/AppShell.session-mgmt.test.tsx`
— passed (47 tests before the archive follow-up; the focused archive regression then passed).
- `npm test` — passed: 44 files / 305 tests.
- `npx tsc -b` — passed.
- `npm run build` — passed.
- `git diff --check` — passed.
- `npm run e2e` reached Playwright but could not run: the environment has no Chromium
executable at Playwright's configured cache path. No application test failure was reported.
## Files changed
- `frontend/src/api/types.ts`: typed principal and session scope contracts.
- `frontend/src/api/sessions.ts`: typed `/me` API call; scoped listing defaults to `mine`.
- `frontend/src/shell/AppShell.tsx`: identity query, admin-only session scope selector and
banner, owner-aware destructive action confirmations.
- `frontend/src/shell/NavSessions.tsx`: owner labels in the all-sessions view.
- `frontend/src/api/sessions.test.ts`, `frontend/src/shell/NavSessions.test.tsx`, and
`frontend/src/shell/AppShell.session-mgmt.test.tsx`: contract and UX coverage.
## Self-review
- Regular users remain fail-closed on `mine`; no administrator control renders without
`principal.isAdmin`.
- The all-sessions view includes owner labels (including `Unknown` for legacy records).
- Delete confirmation preserves the pre-existing select-all behavior and adds confirmation
for foreign/unknown owners. Foreign archive now also requires an explicit browser
confirmation; existing Stop & save already has its confirmation dialog.
- A read-only review found no critical, important, or minor issues. The archive guard was
added after that review in response to the requirement to cover every destructive rail
action, and has its own RED/GREEN regression plus the final full verification above.
## Concerns
- E2E remains environment-blocked until the Playwright Chromium browser is installed.
- Existing Vitest runs emit pre-existing MSW unmatched-request and dialog-ref warnings; all
assertions pass and this task does not modify those shared test/UI primitives.
## Review remediation
- A post-commit review correctly identified that matching `displayName` must never establish
ownership. The predicate now skips confirmation only when `session.author` exactly equals
`principal.subject`; all display-name matches and missing authors are conservative
cross-owner actions.
- Added RED/GREEN regressions where two principals share display name `Alice` but have distinct
subjects: both delete (with another session present, so select-all cannot mask the guard) and
archive require confirmation.
- Added `aria-pressed` to the My sessions / All sessions controls and asserts their selected state
before and after switching.
- Remediation verification: focused regressions passed; full frontend Vitest (44 files / 305
tests), `npx tsc -b`, `npm run build`, and `git diff --check` all passed.
---
# DWH authentication Task 6 — Nginx and CI gate report
## Scope
Added only the two DWH-auth Nginx gates and the `dwh-auth-linux` deployment workflow job:
- `scripts/test-dwh-auth-nginx-contract.sh`
- `scripts/test-dwh-auth-nginx-integration.sh`
- `.github/workflows/deployment.yml`
This report deliberately remains unstaged. The pre-existing frontend Task 6 report above is
preserved rather than overwritten.
## TDD RED
The structural gate was written before any Task 5 template change. Those templates already met
the approved contract, so the behavioral RED was obtained by copying them into one exact temporary
root and removing only the effective `/dwh/` `auth_request` directive. The new checker failed as
required, with no credential material in output:
```text
case=source_contract status=FAIL
```
The runtime gate was also first invoked before its file existed:
```text
bash: scripts/test-dwh-auth-nginx-integration.sh: No such file or directory
```
The CI-job RED check found no `dwh-auth-linux` job in `deployment.yml`. No production template was
modified: the tests prove the existing Task 5 template contract instead of weakening it.
## GREEN
Shell syntax and workflow YAML were checked with:
```text
bash -n scripts/test-dwh-auth-nginx-contract.sh scripts/test-dwh-auth-nginx-integration.sh
python3 -c import-yaml-and-safe-load
```
The structural gate passed its source contract plus these 13 real copied-and-mutated Nginx fixtures:
```text
missing_auth_request
missing_proxy_method
missing_proxy_body
missing_proxy_header_isolation
missing_content_length_clear
missing_verifier_key_forward
missing_upstream_key_clear
missing_failure_mapping
public_verifier
tcp_authenticator
postgrest_bypass
failure_mapped_to_success
full_secret_rate_key
```
Each test mutates an effective, not comment-only, directive and requires the checker to reject it.
The source test and all 13 fixture tests emitted `case=... status=PASS`, followed by
`case=summary status=PASS`.
The isolated Nginx 1.24 smoke passed these sanitized cases:
```text
nginx_1_24
build_dwh_auth
registry_setup
verifier_start
synthetic_upstreams
composite_nginx_config
nginx_start
auth_socket_unix_only
verifier_not_public
valid_v1
valid_legacy
invalid_key
revoked_key
expired_key
duplicate_v1
duplicate_legacy
stopped_verifier
header_and_path_isolation
summary
```
It builds with the pinned official Go 1.26.5 image when the host Go binary is absent, creates only
synthetic v1, legacy, revoked, and expired credentials in a `0700` `/tmp` root, runs both Nginx and
the verifier on explicit temporary Unix sockets, and uses a loopback-only marker backend. Its output
is strictly `case` and `status`; keys, values, and digests remain only in the exact temporary root
and are removed by the trap.
`nginx -t` passed against the complete generated configuration. The marker proves that successful
`/dwh/?keep=exact&second=two` reaches the upstream unchanged, while neither the client API key nor
client or verifier `X-DWH-Key-ID` reaches it. A Unix forwarding probe proves that the verifier sees
only `X-API-Key`, with Cookie, Authorization, and spoofed audit ID absent. Duplicate v1 and ordinary
legacy headers return 401 through Nginx; a stopped verifier returns 503.
The final local equivalent of the four CI commands passed:
```text
Docker Go 1.26.5: go test -race ./... -count=1 and go vet ./...
bash scripts/test-dwh-auth-build-contract.sh
bash scripts/test-dwh-auth-nginx-contract.sh
bash scripts/test-dwh-auth-nginx-integration.sh
```
The Go race suite passed for command, credential, record, registry, securefile, and service;
`go vet` was silent; the build contract passed; both Nginx gates reached their summaries.
## CI contract
The new job uses `actions/checkout` with `persist-credentials: false`, pins Go 1.26.5 with cache
keyed on `tools/dwh-auth/go.mod`, installs `nginx-light`, and runs exactly the four required commands.
Existing jobs were not altered.
## Self-review
- The template tests parse normalized effective directives, so commented-out declarations cannot
satisfy the gate.
- The authentication socket is configured as `http://unix:...:/verify`, is observed by `ss -xl`,
and Nginx itself listens only on a temporary Unix socket; neither test starts a public listener.
- All spawned processes are registered by PID; cleanup signals only those PIDs and deletes only the
exact `mktemp` root after a guarded path check.
- The verifier, marker, registry, Nginx prefix, PID, logs, config, and sockets all reside beneath
that root. No `/etc`, systemd, active Nginx config, stack, legacy route, or real registry/key is
read or changed.
- Task 5 templates were not modified because the structural and runtime tests passed unchanged.
## Concern
The sandbox `apply_patch` helper repeatedly failed with `bwrap: loopback: Failed RTM_NEWADDR:
Operation not permitted`. A narrowly scoped fallback editor was used only for the workflow and the
Nginx-version assertion. Its first workflow insertion interpreted the action-reference at signs;
the two malformed values were immediately corrected and all final YAML, exact-string, syntax, and
four-command checks were rerun. No remaining product concern is known; the integration gate requires
Nginx 1.24 and Python 3, both supplied by the specified Ubuntu CI runner.
---
# DWH authentication Task 6 — review remediation wave
## Review findings and RED evidence
The three review findings were reproduced against the Task 6 commit before their corresponding
hardening was accepted.
1. The contract checker originally selected only the first matching `/dwh/` location. A real copied
fixture appended this competing location without authentication:
```nginx
location ~ ^/dwh/ {
proxy_pass http://127.0.0.1:3001;
}
```
The first run reached the new check and failed as required:
```text
case=negative_postgrest_regex_bypass status=FAIL
```
2. The previous process stop sent TERM and immediately used an unbounded `wait`. A synthetic Python
child ignored TERM; the RED run used one exact short-lived watchdog only to prevent a test hang and
produced:
```text
case=cleanup_term_ignored_bounded status=FAIL
```
3. The TCP detector has a positive-control regression. A scratch copy of the integration script
replaced its `ss -ltnpH` detector with `return 1`; its known loopback listener was then not
detected and the run failed with:
```text
case=tcp_listener_detector_positive status=FAIL
```
All RED fixtures and the scratch script used an exact temporary path and were removed. No template,
service, workflow, key, or active Nginx configuration was changed.
## GREEN changes
- `location_declarations` consumes normalized, comment-stripped effective lines and `check_templates`
requires exactly one each of the only approved locations: verifier, unavailable named location, and
`/dwh/`. It therefore rejects both any extra intercepting location and a duplicate. The real regex
bypass and a new real duplicate `/dwh/` bypass fixture both pass by being rejected.
- `tcp_listener_for_pid` uses `ss -ltnpH` and a PID-bound match. The integration gate starts a
loopback-only synthetic listener, proves the detector sees that exact PID, stops and deregisters it,
then proves the verifier PID has no TCP listener while its Unix socket remains present.
- `stop_registered_pid` now sends TERM, polls for exit or zombie for a bounded deadline, sends KILL
if required, polls a second bounded deadline, and only reaps a direct child after terminal state is
proved. Explicit stops deregister their PID. The cleanup loop invokes that bounded operation only
for recorded PIDs and removes only its guarded temporary root.
- The synthetic child that ignores TERM is killed by the bounded path, must no longer answer to
`kill -0`, must not remain registered, and must finish within three seconds. Final gate output is
restricted to `case` and `status` lines.
## GREEN verification
```text
bash -n scripts/test-dwh-auth-nginx-contract.sh scripts/test-dwh-auth-nginx-integration.sh
Docker Go 1.26.5: go test -race ./... -count=1 and go vet ./...
bash scripts/test-dwh-auth-build-contract.sh
bash scripts/test-dwh-auth-nginx-contract.sh
gate contract: source plus 15 negative fixtures PASS, then summary PASS
bash scripts/test-dwh-auth-nginx-integration.sh
gate integration: 20 named cases PASS, then summary PASS
git diff --check
```
The integration cases include `cleanup_term_ignored_bounded`,
`tcp_listener_detector_positive`, `auth_socket_unix_only`, all existing credential decisions,
composite Nginx syntax, and stopped-verifier 503 behavior. Go race tests passed for command,
credential, record, registry, securefile, and service; vet and both diff checks were silent.
## Self-review and concern
The new location parser rejects comment-only and non-exact declarations because it operates on the
same normalized effective representation used by the rest of the contract. The TCP positive control
binds only `127.0.0.1` on a kernel-selected temporary port and is stopped through the same exact-PID
path under test. The bounded cleanup avoids arbitrary process lookup or broad signaling.
The environment still intermittently rejects `apply_patch` with the sandbox loopback error noted in
the original report; only narrowly scoped fallback edits to the two authorized scripts were used and
all final gates were rerun. No remaining review concern is known.
-90
View File
@@ -1,90 +0,0 @@
# Task 7 — report
## RED
- Creato `scripts/test-verify-dwh-auth-docs.sh` con fixture positiva e fixture negative per
credenziale/digest sintetici, TLS insicuro, segreto in env/argv, mode world-readable, cattura
Nginx e coupling Compose.
- Eseguito `bash scripts/test-verify-dwh-auth-docs.sh` prima del verificatore: `case=verifier_missing status=FAIL`.
## GREEN
- Aggiunti manuali server, client, TLS, runbook PSD, collaudo ed evidenza sanitizzata; collegati
manuali locali/server, setup PSD, guida, indice e nav MkDocs.
- Eseguiti: `bash -n scripts/verify-dwh-auth-docs.sh scripts/test-verify-dwh-auth-docs.sh`,
`bash scripts/test-verify-dwh-auth-docs.sh`, `bash scripts/verify-dwh-auth-docs.sh`,
`bash scripts/test-verify-workspace-install-docs.sh`, `bash scripts/auth-docs-smoke.sh`.
- Tutti gli output finali sono PASS; il nuovo gate esercita una fixture positiva e nove negative.
## Self-review
- Verificati path/owner/mode: registry 2750, lock/record 0640, socket 0660.
- Verificata separazione: chiavi solo `rest_api`; PSD server `postgres_direct`; Mac/remoti REST;
nessun lifecycle Compose per `dwh-auth`.
- Verificati TLS `.it`/SAN, `.com` non coperto, `TLS_CA_FILE`, fingerprint fuori banda, rinnovo e
assenza di bypass.
- Verificati due gate Task 9–10, evidenze solo metadati e nessuna migrazione di sessioni/index/cache legacy.
## Concern
- Nessuna mutazione PSD/Nginx/systemd/registry o lettura di segreti è stata eseguita. I comandi del
runbook restano condizionati alle autorizzazioni separate dei Task 9 e 10.
## Review fix — RED/GREEN
### RED review
- La fixture `sudo nginx -T` ha prodotto il rifiuto `case=sudo_raw_nginx_capture status=FAIL` prima della correzione del gate.
- La fixture header legacy opaco ha prodotto `case=opaque_legacy_header_literal status=FAIL` prima della correzione del gate.
- Dopo avere riallineato le label UI nei manuali, `bash scripts/test-verify-workspace-install-docs.sh` ha prodotto `server-workspace-registry.md: curator flow missing registry rule`: il verifier cercava ancora le due label precedenti. Il test sulla base HEAD e il diff hanno confermato la causa.
### GREEN review
- Il gate DWH ora rifiuta anche header opaco, digest JSON quotato, `export` di API key, `curl --header` e `-H`, `sudo nginx -T`, raw diff e Compose; le mutation fixture coprono label, PSD direct/Mac REST/CA, socket e flag REST.
- Il runbook non prescrive raw diff o dump: solo checker strutturale e secret scan con metadati e PASS/FAIL. Il piano Task 10 adotta la stessa regola.
- Il template `psd-local` resta `rest_api` solo Mac/local/remota; il server PSD Project A resta `postgres_direct` con binding separato. La CA privata e `TLS_CA_FILE` sono obbligatori salvo trust approvato equivalente.
- Le procedure server ora coprono backup manifest protetto, restore, curl config 0600 senza segreto in argv/env/output, Unix 204/401, HTTPS 2xx/401, 503 bounded con trap, journal PASS/FAIL e retention alla disinstallazione.
- Il verifier workspace-install e entrambi i manuali registry usano ora le quattro label effettive: `Validate workspace source`, `Test workspace connections`, `Save entered secrets`, `Forget stored value`.
### Final verification review
- PASS: `bash scripts/test-verify-dwh-auth-docs.sh`.
- PASS: `bash scripts/verify-dwh-auth-docs.sh`.
- PASS: `bash scripts/test-verify-workspace-install-docs.sh` (fixture complete).
- PASS: `bash scripts/auth-docs-smoke.sh`.
- PASS: `bash -n scripts/verify-dwh-auth-docs.sh scripts/test-verify-dwh-auth-docs.sh` e `git diff --check`.
### Review concern
- Nessuna mutazione runtime e nessun segreto reale sono stati letti. I soli comandi server documentati restano soggetti ai gate autorizzativi Task 9 e Task 10.
## Review fix wave 2 — RED/GREEN
### RED wave 2
- Prima della correzione del proxy, `bash scripts/test-dwh-auth-build-contract.sh` ha fallito il contratto di preservazione path e `bash scripts/test-dwh-auth-nginx-integration.sh` ha chiuso con `case=header_and_path_isolation status=FAIL`: il prefisso `/dwh` arrivava a PostgREST invece di essere rimosso.
- Prima delle procedure finali, il gate docs ha rifiutato il path chiave non deterministico e la fixture curl con header legacy opaco ha dato `case=header_file_curl_synthetic status=FAIL` perché il valore non veniva confrontato esattamente.
- Le mutation fixture hanno catturato l'estrazione tar sul registro attivo e i rename non protetti. Dopo l'inasprimento finale del gate, la sorgente ha dato `dwh-auth docs: restore must stage/check then use guarded same-filesystem renames` finché mancava il controllo fail-closed del candidato.
- Il RED finale dello scanner journal è stato `dwh-auth docs: docs/install/dwh-auth-server.md lacks required topic: sys.argv[2:]`: il gate esige la lettura byte-esatta di v1 e legacy e un `journalctl` che fallisca chiuso.
### GREEN wave 2
- Commit `f616aab fix: preserve PostgREST RPC path through DWH proxy`: `proxy_pass` termina con `/`; il contratto e l'integrazione verificano `/dwh/rpc/ping?x` verso `/rpc/ping?x`.
- Il runbook usa un singolo file chiave v1, header file `0600` passati solo con `curl --header @file`, socket 204 dual-key, HTTPS 2xx pre/post per v1 e 401 post-revoca per legacy `legacy-shared`.
- Restore protetto: staging sul filesystem `/var/lib`, check candidato, `mv -T --` guardato per ogni publish/rollback e pre-restore conservato. Backup/manifest restano root-only `0600` su storage cifrato approvato.
- Lo scanner journal esegue `journalctl` in un unico processo Python root, sopprime stderr, controlla return code e bytes esatti di entrambe le chiavi senza emettere journal o segreti; la shell mostra solo PASS/FAIL.
- Il verifier rifiuta `curl --config`, header in argv, raw Nginx/diff, TLS insicuro, segreti env, mode insicuri e Compose. Le fixture mutano path chiave, ID legacy, header/legacy probes, restore, journal, codici HTTPS e label UI.
### Final verification wave 2
- PASS: `bash scripts/test-dwh-auth-build-contract.sh`.
- PASS: `bash scripts/test-dwh-auth-nginx-contract.sh`.
- PASS: `bash scripts/test-dwh-auth-nginx-integration.sh`.
- PASS: `bash scripts/test-verify-dwh-auth-docs.sh` e `bash scripts/verify-dwh-auth-docs.sh`.
- PASS: `bash scripts/test-verify-workspace-install-docs.sh` e `bash scripts/auth-docs-smoke.sh`.
- PASS: `bash -n` sugli otto gate shell e `git diff --check`.
### Review concern wave 2
- Nessuna configurazione protetta, chiave reale, Nginx, systemd o stack PSD è stata letta o mutata. Le procedure privilegiate restano istruzioni condizionate ai Gate 9–10; la verifica degli owner/mode reali è un'attività del rollout autorizzato, non di questo task documentale.
+3 -1
View File
@@ -7,7 +7,9 @@ This file provides guidance to Codex (Codex.ai/code) when working with code in t
Read [PROJECT_STATE.md](PROJECT_STATE.md) for the current-state snapshot: what was last
built, pending manual gates, workspace/secret layout, and design-doc locations. This file
holds the stable commands + architecture mental model; PROJECT_STATE.md holds the evolving
detail. Design history lives in `docs/superpowers/specs/` and `docs/superpowers/plans/`.
detail. Current architecture and contracts live in `docs/architecture/`, `docs/contracts/`,
and `docs/evidence.md`; durable design decisions live in `docs/adr/`. Git history is the source
for superseded designs and implementation plans.
## Commands
-87
View File
@@ -1,87 +0,0 @@
# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## Start here
Read [PROJECT_STATE.md](PROJECT_STATE.md) for the current-state snapshot: what was last
built, pending manual gates, workspace/secret layout, and design-doc locations. This file
holds the stable commands + architecture mental model; PROJECT_STATE.md holds the evolving
detail. Design history lives in `docs/superpowers/specs/` and `docs/superpowers/plans/`.
## Commands
The repo has three independently-built layers. Run the **full stack** (real Pi + DWH, needs
VPN + `harness/.env` + `pi` on PATH) with `./scripts/run-stack.sh` (frontend :5173 → backend :8787).
**harness/** (Python `tht` CLI + Pi gate extension)
- Install: `cd harness && python -m venv .venv && pip install -e ".[dev]"` (puts `tht` on PATH)
- Test: `.venv/bin/pytest -q` — `l2` (real GLM + remote DB) is opt-in via `addopts = -m 'not l2'`; `l0` (testcontainers) needs Docker
- Single test: `.venv/bin/pytest tests/test_session_mutations.py::test_set_name -v` (or `-k <pattern>`); include e2e with `-m l2`
- Lint: `.venv/bin/ruff check .` (line-length 100)
**backend/** (Fastify + TypeScript, vitest)
- Dev: `npm run dev` (tsx watch `src/server.ts`) · Build: `npm run build` (tsc → `dist/`)
- Test: `npx vitest run` · Single: `npx vitest run test/routes-sessions.test.ts -t "rename"`
- Typecheck: `npx tsc --noEmit -p .` (vitest does NOT type-check — run this before committing)
**frontend/** (React 18 + Vite + vitest)
- Dev: `npm run dev` (Vite; set `VITE_BACKEND_URL`) · Build: `npm run build`
- Test: `npx vitest run` · Single: `npx vitest run src/shell/NavSessions.test.tsx`
- Typecheck: `npx tsc -b` · E2E: `npm run e2e` (Playwright)
No ESLint on the TS layers — `tsc` is the gate. Tests use vitest + MSW (no network).
## Architecture (the parts that need multiple files to see)
```
frontend (React/SSE) → backend (Fastify) → pi --mode rpc → tht/harness → DWH (read-only)
```
- **The harness owns the workflow and all persistence.** `tht` (Python) is a deterministic
CLI; `harness/.pi/extensions/tht-gate.js` is a Pi extension that drives an **8-phase
NL→SQL workflow**. The single source of workflow truth is `harness/workflow.yaml`; the
orchestration rules the model must follow are `harness/.pi/skills/tht-sessione/SKILL.md`.
"Current phase" is computed by folding the decision ledger (`harness/tht/phase.py`), not
stored — read it before reasoning about phase logic.
- **Persistence = phase documents, NOT chat.** A session is a directory under the workspace's
`sessions/` path: `session_manifest.yaml` + per-phase artifacts (`question.md`,
`schema_linking.json`, `sql_final.sql`, …) + `review_decisions.jsonl`. The contract
(SKILL.md): *"the persisted state is the truth — what is not recorded did not happen."*
There is no verbatim transcript store. A resumed Pi process rebuilds context from
`tht session show <id>` + the on-disk artifacts.
- **The backend is a thin bridge with no database of its own.** `ThtRunner` shells `tht`
subcommands; `PiProcessManager` runs one Pi child per session and bridges its RPC stream;
`SessionBridge` maps Pi RPC events → client events (`ui_request`/`text_delta`/`info`);
`SseHub` fans them out over SSE to the browser. Persistence belongs to the HARNESS, which
selects the session repository from the workspace config (`harness/tht/session/repository.py`):
filesystem by default, **PostgreSQL when `session_storage` is configured** (server/portable
deployment). Settings flow through harness preferences (`tht session preferences`) with
`backend/data/settings.json` only as the file fallback for injected runners/tests.
- **Human-in-the-loop gate contract.** The model proposes; a human reviewer decides at gates
via widgets (`reviewer_select` = single pick — a chosen option carrying a `decision` payload
auto-confirms/persists directly, an option without one only asks; `reviewer_decide` = multiselect,
each choice IS a decision; `reviewer_confirm` = artifact/phase gate). The frontend renders these
widget-descriptors (`src/widgets/` registry) and the live transcript is rebuilt in-memory
from the SSE stream (`src/store/sessionStore.ts`) — it is not persisted.
## Project-specific gotchas
- **`tht`'s `-c`/`--config` is a PER-COMMAND option** — it must follow the subcommand, never
precede it (`ThtRunner.buildArgv` enforces this; prepending caused live 500s).
- **`--json` output must be pristine** (only valid JSON on stdout) — used as a machine contract.
- **UI strings are English; document *content* stays the workspace language** (Italian for
`psd`) because it's the real data. Only chrome/labels are English.
- **Workspaces** (`harness/workspaces/*.yaml`) set the DB target and **absolute**
`paths.sessions/artifacts/indexes` — for `psd` these point at a *separate, uncommitted* repo
(`tht-workspace-psd/`). Secrets live ONLY in `harness/.env` (gitignored).
- **Settings are global** (workspace/provider/model/thinking, persisted via harness
preferences — `backend/data/settings.json` is only the fallback); the New-session form is
question-only.
- **Resume**: a resumable session re-enters at its last incomplete phase. The backend refuses
resume with 409 when `finalized` or `archived`, and `PiProcessManager.spawnFor` must send
`/riprendi-sessione <id>` (resume mode) vs `/nuova-domanda` (new) — sending the wrong prompt
silently turns a resume into a new question.
+153
View File
@@ -91,6 +91,159 @@ correzione successiva crea una nuova sessione derivata, collegata a quella prece
dopo la finalizzazione. Non può modificare il ledger, gli artifact canonici o lo stato
terminale della sessione.
## Evidence
**Evidence Module** — Il modulo autonomo che possiede la preparazione delle Evidence e
la loro consultazione durante il workflow. La preparazione avviene fuori dalle singole
sessioni; il workflow usa soltanto contenuti già pubblicati. A runtime contribuisce agli
stage semantici esistenti, senza diventare uno stage visibile e senza modificare ledger,
artifact o stato del workflow.
**Source Evidence** — Un documento originale del workspace, conservato senza modifiche
come riferimento umano e origine della successiva ristrutturazione.
**Evidence Unit** — La più piccola unità semantica coerente, revisionabile e ricercabile
derivata da una sola Source Evidence. Possiede un identificatore stabile indipendente
dal kind, assegnato una volta nella forma `evidence:<slug>`; fonti diverse non vengono
fuse automaticamente.
**Evidence kind** — La categoria semantica di una Evidence Unit, che ne determina i
campi specifici e ne orienta l'uso. Ogni unità ha un solo kind primario; i tipi iniziali
sono `glossary`, `domain`, `enum`, `example`, `mapping`, `normalization`, `formula` e
`reference`.
**Glossary Evidence** — Una Evidence Unit che definisce il significato linguistico, i
sinonimi o le varianti di un termine.
**Domain Evidence** — Una Evidence Unit che esprime una regola o un vincolo del dominio
non rappresentato da un kind più specifico.
**Enum Evidence** — Una Evidence Unit che collega un insieme finito di valori
memorizzati ai relativi significati.
**Example Evidence** — Una Evidence Unit che associa un input o una domanda alla sua
interpretazione o al risultato atteso.
**Mapping Evidence** — Una Evidence Unit che collega un concetto logico agli elementi
del relativo schema fisico.
**Normalization Evidence** — Una Evidence Unit che descrive la trasformazione di una
rappresentazione in una forma canonica.
**Formula Evidence** — Una Evidence Unit che contiene una singola espressione PostgreSQL
componibile e ne dichiara gli input. Una query SQL completa non è una Formula Evidence.
**Reference Evidence** — Una Evidence Unit che rappresenta un collegamento esterno da
restituire come contenuto autonomo, anziché come semplice provenienza.
**Evidence purpose** — La destinazione dichiarata di una Evidence Unit nel workflow:
disambiguation, rewriting, schema linking o SQL generation. È distinta dall'Evidence
kind: il tipo descrive cosa contiene, il purpose quando può essere utile; durante la
ricerca il purpose richiesto è un filtro obbligatorio. Il recupero di esperienze e
soluzioni precedenti appartiene al Memory Module e non è un Evidence purpose.
**Evidence Search Outcome** — Il risultato tipizzato di una consultazione del modulo
Evidence. Distingue una ricerca disponibile, che può legittimamente non trovare
corrispondenze, da un'indisponibilità tecnica che impedisce allo stage chiamante di
avanzare fino a un retry riuscito.
**Evidence receipt** — La traccia minima di una consultazione disponibile conservata
nella sessione: stage semantico, purpose, generazione interrogata e identificatori delle
Evidence restituite. Non duplica il contenuto delle Evidence.
**Curated Evidence** — Una o più Evidence Unit ristrutturate a partire da una Source
Evidence e conservate nel repository del workspace come proposte per la revisione
umana. Git conserva la versione precedente e rende visibile ogni modifica; una Curated
Evidence non è ancora contenuto autorevole del runtime.
**Published Evidence** — Le Curated Evidence valide appartenenti alla revisione attiva
del workspace e alla generazione Evidence pubblicata. L'approvazione umana precede
l'attivazione, ma non viene duplicata come stato nel manifest.
**Evidence Index** — La proiezione ricercabile e ricostruibile delle Published Evidence.
Accelera il recupero delle informazioni, ma non è una fonte di verità.
**Evidence preparation** — Il processo di authoring che trasforma Source Evidence in
Curated Evidence mediante estrazione e normalizzazione deterministiche, una singola
ristrutturazione assistita dal modello e una validazione finale deterministica. Nella
prima versione accetta Markdown o testo UTF-8 e non acquisisce automaticamente il
contenuto di URL o documenti esterni. Prepara l'intero insieme delle modifiche in
un'area temporanea e lo applica atomicamente soltanto se tutti gli output sono validi;
non ritenta automaticamente una chiamata al modello fallita.
**Supporting excerpt** — Un breve estratto presente nel Source Evidence che sostiene
una Evidence Unit. Il sistema ne verifica deterministicamente la presenza dopo la
normalizzazione meccanica; il curatore resta responsabile di verificarne la sufficienza
semantica.
**Evidence resolution** — L'operazione esplicita con cui un curatore ritira una
Evidence Unit oppure la ricollega a un Source Evidence esistente. Aggiorna documento e
manifest insieme, lascia un diff Git revisionabile e non pubblica né crea commit.
**Review item** — Un blocco di revisione descritto da codice stabile, messaggio umano e
campo opzionale. Finché viene mantenuto nell'Evidence Unit, ne impedisce la
pubblicazione; la sua storia è conservata da Git, non da uno stato interno all'item.
**Retirement candidate** — Una Curated Evidence che il Source Evidence esistente non
sostiene più. Rimane visibile con un Review item e blocca la pubblicazione finché il
curatore non la elimina oppure la rende nuovamente coerente con il sorgente.
**Evidence evaluation set** — Un piccolo insieme versionato di domande rappresentative
e relativi risultati attesi. La baseline è accettabile quando ogni domanda recupera
almeno un risultato atteso nei primi dieci risultati della fusione RRF; il risultato
nei primi cinque è informativo. Comprende almeno un caso lessicale, uno semantico e uno
misto e conserva, a fini diagnostici, le posizioni dense, BM25 e fused.
**Candidate Evidence Generation** — Una generazione completa dell'Evidence Index che
può essere valutata ma non è ancora visibile alle sessioni. Diventa attiva soltanto se
supera l'Evidence evaluation set.
**Evidence manifest** — Il file versionato e gestito dal sistema che collega ogni
Source Evidence al suo hash e alle Evidence Unit derivate. Conserva gli identificatori
stabili, permette l'elaborazione incrementale e segnala le unità rimaste orfane senza
cancellarle automaticamente.
**Orphaned Evidence Unit** — Una Curated Evidence il cui Source Evidence non esiste più.
Rimane disponibile per la revisione, ma blocca la pubblicazione finché non viene
eliminata, ricollegata oppure ne viene ripristinato il sorgente.
**Evidence Fragment** — Una proiezione ricercabile di una sezione semanticamente
coerente di una Published Evidence. La divisione segue intestazioni e confini di
paragrafo; formule, coppie valore/significato, mapping, regole e URL non vengono mai
tagliati. Il testo completo reso per il frammento usa il solo limite esistente
`max_chunk_chars`, pari per default a 4.000 caratteri; un elemento atomico troppo grande
produce un Review item bloccante. Qdrant indicizza i frammenti, mentre l'Evidence Module
li raggruppa per Evidence Unit.
**Evidence Result** — La rappresentazione di una singola Evidence Unit restituita dalla
ricerca con metadati, migliori estratti, provenienza e riferimento al documento completo.
**Hybrid Evidence retrieval** — La ricerca che combina in Qdrant una graduatoria
semantica dense e una graduatoria lessicale BM25 sparse mediante Reciprocal Rank
Fusion. I metadati tipizzati restringono o orientano i risultati senza creare una
collezione separata per ogni Evidence kind.
**Evidence query text** — La rappresentazione deterministica condivisa dalla ricerca
dense e BM25: domanda originale, concetti, tabelle e colonne in ordine fisso. I campi
vuoti sono omessi; domanda e contesto ricevono soltanto normalizzazione Unicode NFC,
conversione degli a-capo e rimozione degli spazi esterni. Gli elementi contestuali sono
poi deduplicati e ordinati senza conversione delle maiuscole, mentre punteggiatura e
spazi interni della domanda non vengono riscritti.
**Additive BM25 upgrade** — L'estensione non distruttiva della collezione semantica di
un workspace che conserva il vettore dense predefinito e aggiunge il solo vettore
sparse `bm25`. Soltanto gli Evidence Fragment ricevono valori BM25; Schema e Memory
mantengono invariati dati e ricerca dense.
**Formula proposal** — Una formula individuata durante una sessione e conservata come
artefatto della sessione. Non diventa Published Evidence finché non viene importata,
revisionata e approvata nel repository del workspace.
**Fail-closed Evidence retrieval** — Il comportamento per cui un indice assente,
incompatibile o non aggiornato produce nessuna Evidence e un avviso esplicito. Il
workflow può continuare, ma non usa mai silenziosamente contenuti di una revisione
precedente o di un altro workspace.
## Catalogo dei metadati
**Workspace Database** — Il database associato a un workspace, considerato nella sua
+327
View File
@@ -0,0 +1,327 @@
---
name: ThothII
description: "A calm, precise clinical analytics workbench for traceable and reviewable SQL workflows."
colors:
instrument-red: "oklch(55.87% 0.1881 23.2)"
instrument-red-hover: "oklch(50.95% 0.1812 24.1)"
porcelain-background: "oklch(99.18% 0.0011 17.2)"
porcelain-card: "oklch(99.85% 0.0006 17.2)"
warm-surface: "oklch(97.09% 0.0011 17.2)"
sunken-surface: "oklch(94.08% 0.0011 17.2)"
warm-graphite: "oklch(26.78% 0.0097 355.6)"
muted-graphite: "oklch(51.33% 0.0088 345.6)"
quiet-border: "oklch(90.93% 0.0035 354.7)"
success-mint: "oklch(75.77% 0.1581 165)"
warning-amber: "oklch(85.23% 0.1386 78.9)"
information-blue: "oklch(70.35% 0.1128 221.3)"
typography:
display:
fontFamily: "Fraunces, Source Serif Pro, Georgia, Times New Roman, serif"
fontSize: "3rem"
fontWeight: 600
lineHeight: 1.03
letterSpacing: "-0.025em"
headline:
fontFamily: "Fraunces, Source Serif Pro, Georgia, Times New Roman, serif"
fontSize: "1.875rem"
fontWeight: 600
lineHeight: 1.15
letterSpacing: "-0.015em"
title:
fontFamily: "Fraunces, Source Serif Pro, Georgia, Times New Roman, serif"
fontSize: "1.2rem"
fontWeight: 600
lineHeight: 1.25
letterSpacing: "-0.01em"
body:
fontFamily: "Manrope, -apple-system, BlinkMacSystemFont, Segoe UI, system-ui, Arial, sans-serif"
fontSize: "0.9375rem"
fontWeight: 400
lineHeight: 1.65
letterSpacing: "normal"
control:
fontFamily: "Manrope, -apple-system, BlinkMacSystemFont, Segoe UI, system-ui, Arial, sans-serif"
fontSize: "0.875rem"
fontWeight: 600
lineHeight: 1.25
letterSpacing: "0.005em"
label:
fontFamily: "ui-monospace, SF Mono, Cascadia Code, Menlo, Consolas, monospace"
fontSize: "0.6875rem"
fontWeight: 600
lineHeight: 1.25
letterSpacing: "0.06em"
rounded:
xs: "4px"
sm: "6px"
md: "8px"
lg: "12px"
xl: "16px"
full: "9999px"
spacing:
xs: "4px"
sm: "8px"
md: "16px"
lg: "24px"
xl: "32px"
components:
button-primary:
backgroundColor: "{colors.instrument-red}"
textColor: "{colors.porcelain-background}"
typography: "{typography.control}"
rounded: "{rounded.md}"
padding: "0 14px"
height: "32px"
button-primary-hover:
backgroundColor: "{colors.instrument-red-hover}"
textColor: "{colors.porcelain-background}"
typography: "{typography.control}"
rounded: "{rounded.md}"
padding: "0 14px"
height: "32px"
button-secondary:
backgroundColor: "{colors.porcelain-card}"
textColor: "{colors.warm-graphite}"
typography: "{typography.control}"
rounded: "{rounded.md}"
padding: "0 14px"
height: "32px"
input-default:
backgroundColor: "{colors.porcelain-background}"
textColor: "{colors.warm-graphite}"
typography: "{typography.body}"
rounded: "{rounded.md}"
padding: "0 12px"
height: "40px"
card-default:
backgroundColor: "{colors.porcelain-card}"
textColor: "{colors.warm-graphite}"
rounded: "{rounded.lg}"
padding: "16px"
badge-primary:
backgroundColor: "{colors.instrument-red}"
textColor: "{colors.porcelain-background}"
typography: "{typography.control}"
rounded: "{rounded.sm}"
padding: "2px 8px"
height: "20px"
---
# Design System: ThothII
## Overview
**Creative North Star: "The Clinical Workbench"**
ThothII should feel like a well-kept clinical workbench: warm enough for sustained reading, exact
enough for consequential review, and quiet enough that evidence, state, and decisions remain in the
foreground. The visual system is calm, precise, and trustworthy. It uses familiar product patterns,
restrained color, and deliberate density instead of decorative spectacle.
The primary physical scene is an analyst reviewing persisted evidence and SQL on a large monitor in
a well-lit working environment. This makes the warm light theme the default. The supported dark
theme serves lower-light work without becoming a separate neon aesthetic. Both themes preserve the
same hierarchy and semantic roles.
The system rejects generic SaaS ornament, conspicuous ripples, bounce or elastic motion, long
choreographed transitions, and effects that compete with the analytical task. Controls should feel
disciplined and tactile, never playful, sluggish, or visually unstable.
**Key Characteristics:**
- Warm, restrained surfaces with one scarce red accent.
- Editorial headings paired with highly legible operational body text.
- Dense information organized through hierarchy, rhythm, and progressive disclosure.
- Persisted artifacts and reviewer decisions presented as the visual source of truth.
- Fast state feedback with reduced-motion parity.
**The Workbench Rule.** Every visual element must support inspection, action, state, or provenance.
Decoration without an operational purpose is forbidden.
**The Persisted Truth Rule.** Persisted artifacts and reviewer decisions receive stronger hierarchy
than transient model narration.
**The Density with Rhythm Rule.** Preserve information density, but vary spacing between groups so
users can scan structure without adding nested containers.
## Colors
The palette combines warm porcelain surfaces, warm graphite text, and an instrument red used only
for action, focus, and important state. OKLCH values in the frontmatter are normative because the
frontend uses OKLCH tokens directly.
### Primary
- **Instrument Red** (`instrument-red`): primary actions, current selection, focus identity, and
destructive meaning where the context already makes the action explicit.
- **Instrument Red Pressed** (`instrument-red-hover`): hover and active emphasis for the primary
action family.
### Neutral
- **Porcelain Background** (`porcelain-background`): the main canvas.
- **Porcelain Card** (`porcelain-card`): lifted panels, cards, and popovers.
- **Warm Surface** (`warm-surface`): sidebars, secondary controls, and muted regions.
- **Sunken Surface** (`sunken-surface`): selected rows, quiet emphasis, and inset regions.
- **Warm Graphite** (`warm-graphite`): primary text and high-confidence labels.
- **Muted Graphite** (`muted-graphite`): descriptions, timestamps, and secondary metadata.
- **Quiet Border** (`quiet-border`): structural boundaries, input outlines, and dividers.
### Semantic
- **Success Mint** (`success-mint`): completed and ready states.
- **Warning Amber** (`warning-amber`): waiting, attention, and in-progress states.
- **Information Blue** (`information-blue`): informational state when red would imply action.
The dark theme keeps the same semantic mapping with neutral near-black surfaces and a slightly
lighter red accent. Do not introduce a second visual identity for dark mode.
**The One Voice Rule.** Instrument Red should occupy no more than roughly ten percent of a screen.
Its rarity is what makes it authoritative.
**The State Has a Name Rule.** Success, warning, information, and destructive colors are reserved
for their named states. Color is never the only state indicator.
## Typography
**Display Font:** Fraunces, with Source Serif Pro, Georgia, and Times New Roman fallbacks
**Body Font:** Manrope, with native system sans-serif fallbacks
**Label/Mono Font:** SF Mono or Cascadia Code, with Menlo and Consolas fallbacks
**Character:** Fraunces gives persisted artifacts and key headings editorial authority. Manrope
keeps dense controls and prose calm and readable. The mono register separates machine identity,
metadata, SQL, identifiers, and micro-labels from natural-language content.
### Hierarchy
- **Display** (600, `3rem`, `1.03`): authentication and exceptional page-level statements only.
- **Headline** (600, `1.875rem`, `1.15`): major page or artifact titles.
- **Title** (600, `1.2rem`, `1.25`): panel and document section hierarchy.
- **Body** (400, `0.9375rem`, `1.65`): operational prose, with a target line length of 65 to 75
characters where the surface controls width.
- **Control** (600, `0.875rem`, `1.25`): buttons, inputs, tabs, and compact actions.
- **Label** (600, `0.6875rem`, `0.06em` tracking): uppercase micro-labels, state metadata, and panel
headers. Labels use the mono family.
Typography uses fixed sizes. Responsive changes happen at structural breakpoints, not through fluid
type scaling. Numeric data and identifiers use tabular numerals where comparison matters.
**The Three Registers Rule.** Serif means authority, sans means interaction and reading, mono means
machine identity. Do not exchange these roles for novelty.
**The Read Once Rule.** A heading, label, and body must be distinguishable on first glance through
size and weight. Do not repeat headings in explanatory copy.
## Elevation
The system is flat by default and layered when necessary. Borders mark structure. Warm, diffuse
shadows mark actual elevation for popovers, dialogs, and selected containers. Tonal layering should
solve most hierarchy before a shadow is introduced.
### Shadow Vocabulary
- **Contact Shadow** (`--shadow-xs`): a one-pixel contact shadow for controls and code blocks.
- **Panel Shadow** (`--shadow-sm`): a small two-stage shadow for cards that need separation from the
canvas.
- **Overlay Shadow** (`--shadow-md`): a broad, low-opacity shadow for dialogs and floating layers.
Focus uses an explicit three-pixel ring. Waiting-for-input state may use a success-tinted ring, but
must retain a textual or structural cue. Motion for button state changes lasts `140ms` with
`cubic-bezier(0.22, 1, 0.36, 1)`. Dialog transitions last `100ms`. Activity pulses may run at
`1.5s`, and must be disabled under `prefers-reduced-motion`.
**The Flat by Default Rule.** A resting surface has no shadow unless it is physically above another
surface. If every panel floats, none of them has hierarchy.
**The Borders Structure, Shadows Elevate Rule.** Never use shadow as a substitute for grouping or a
border as a decorative accent.
## Components
Components are familiar, compact, and state-complete. Every interactive primitive must define
default, hover, focus, active, disabled, loading, and error behavior where those states apply.
### Buttons
- **Shape:** gently curved rectangle (`8px`) with a one-pixel transparent or structural border.
- **Primary:** Instrument Red, porcelain text, `32px` default height, and `14px` horizontal padding.
- **Hover / Focus:** shift to Instrument Red Pressed; show a three-pixel focus ring at 25 percent
opacity. Active state scales to `0.97` for `140ms` and removes elevation.
- **Secondary / Outline:** porcelain card surface, Quiet Border, Warm Graphite text, and a Warm
Surface hover.
- **Ghost:** transparent at rest, Warm Surface on hover. Use only where surrounding structure makes
the hit target obvious.
### Badges and Status Indicators
- **Style:** compact (`20px` height), gently curved (`6px`), and semibold.
- **State:** pair semantic color with text, icon, or position. A colored dot alone is insufficient
when the state affects workflow decisions.
### Cards and Containers
- **Corner Style:** softly rounded (`12px`), with `16px` default internal padding.
- **Background:** Porcelain Card over Porcelain Background or Warm Surface.
- **Shadow Strategy:** Panel Shadow only when the card must read as elevated.
- **Border:** one-pixel Quiet Border at partial opacity.
- **Nesting:** nested cards are forbidden. Use headings, dividers, spacing, or tonal regions.
### Inputs and Fields
- **Style:** `40px` height, `8px` corners, Porcelain Background, Quiet Border, and Manrope body text.
- **Focus:** three-pixel Instrument Red ring with a clear border shift.
- **Error / Disabled:** errors combine destructive color with explanatory text; disabled controls
retain readable contrast and use 50 percent opacity.
### Navigation
- **Style:** compact session rows use `8px` corners and restrained vertical padding.
- **Default / Hover / Active:** transparent at rest, Sunken Surface on hover, and the same surface
with stronger text weight when active.
- **Responsive:** collapse navigation structurally at the application breakpoint. Do not shrink
labels into illegibility.
### Curated Evidence Documents
Curated evidence follows a fixed reading order: title, compact type and purpose summary, scope,
typed content, supporting excerpts, review items, then collapsed technical provenance. Machine
metadata stays in invisible comments so GitHub Preview shows only the reviewable document.
`applies_to` is rendered as “Ambito di applicazione” with separate bullet lists for concepts,
tables, and columns. Enum values also use lists. Tables are forbidden for metadata, scope, or any
one-dimensional collection; reserve tables for genuinely two-dimensional datasets. Long machine
identifiers use inline code. SQL uses fenced code. Supporting excerpts use blockquotes.
**The Review Surface Rule.** The visible Markdown must be readable without understanding the
machine contract. Technical metadata belongs in progressive disclosure, not above the title.
## Do's and Don'ts
### Do:
- **Do** make every state change unmistakable without interrupting flow.
- **Do** use Instrument Red only for primary action, current selection, focus identity, or explicit
destructive meaning.
- **Do** preserve information density with headings, rhythm, and progressive disclosure.
- **Do** keep keyboard focus explicit and pair color with text, shape, icon, or position.
- **Do** respect `prefers-reduced-motion` while preserving immediate non-kinetic feedback.
- **Do** use English for interface chrome and the workspace language for persisted document content.
- **Do** render curated metadata and scope as Markdown prose or lists, never as a frontmatter table.
- **Do** break long curated rules into paragraphs, labelled subsections, and lists at existing
punctuation boundaries while preserving the exact canonical text for machines.
### Don't:
- **Don't** add generic SaaS ornament, conspicuous ripples, bounce or elastic motion, long
choreographed transitions, or effects that compete with the analytical task.
- **Don't** make controls feel playful, sluggish, or visually unstable.
- **Don't** use gradient text, decorative glassmorphism, or full-saturation accents on inactive
states.
- **Don't** use a colored side stripe greater than one pixel on cards, callouts, list items, or
blockquotes. Use a full border, tonal background, icon, or heading instead.
- **Don't** nest cards or wrap every section in a container.
- **Don't** use a modal before exhausting inline or progressive alternatives.
- **Don't** use tables for `applies_to`, metadata, enum values, or other one-dimensional content.
- **Don't** use color as the sole carrier of success, warning, error, selection, or progress.
- **Don't** use display typography for buttons, labels, or data.
- **Don't** add em dashes to interface copy. Use commas, colons, semicolons, or parentheses.
+89 -1259
View File
File diff suppressed because it is too large Load Diff
+7 -9
View File
@@ -202,20 +202,18 @@ Startup mode adds bounded image build/two-service health startup, installation-a
status, stopped-container-aware ownership checks, and exact cleanup. The ordinary hosted Windows
job remains deterministic and does not claim Docker startup.
## Preprocessing jobs and S3 Evidence
## Workspace preprocessing and S3 Evidence
The included preprocessing services reuse the internal Qdrant/Ollama stack. Mount Evidence at
`/data/source/evidence`, then run the explicit preprocessing preset:
Run preprocessing through the native host CLI and the installation descriptor:
```sh
docker compose --env-file deploy/env/local.env \
-f compose.yaml -f deploy/compose.local.yaml \
-f deploy/compose.preprocess.yaml --profile preprocess run --rm preprocess-evidence
tht --installation /absolute/path/thothii-installation.yaml workspace preprocess evidence
tht --installation /absolute/path/thothii-installation.yaml workspace preprocess dwh
```
Replace the final service with `preprocess-dwh` when required. The overlay makes each job wait for the internal Qdrant
service health checks and embedding model initialization; no separate semantic-service startup is
required.
The CLI starts the profile-gated `workspace-maintenance` service and enforces the workspace,
secret, Qdrant, and embedding contracts. See [Evidence](docs/evidence.md) and the
[workspace preprocessing CLI contract](docs/contracts/workspace-preprocessing-cli.md).
S3 Evidence uses the optional `tht[s3]` dependency and canonical `s3://bucket/key` provenance.
AWS endpoints are used when no custom URL is supplied. Every custom endpoint is an explicit egress
@@ -15,28 +15,32 @@ const allowedKinds = new Set(["policy_text", "workspace_descriptor", "deployment
// opener line through the closer line (including physical line endings). These
// blocks are reviewed non-workspace runtime/config generation, not semantic proof.
const reviewedExpandableBlocks = new Map([
["scripts/preprocess-smoke.sh", [
{ sha256: "fc530dc721c946644ab6552bbd46b7918d6c5f11f06f3495b6ea1fcda819b38d", rationale: "Generates the reviewed preprocess Compose override." },
["scripts/test-dwh-auth-nginx-integration.sh", [
{ sha256: "ead57234ad3520b5c7d4262b772957cbc7b9589da4f35fb17b160f948eb2ac7b", rationale: "Generates the reviewed isolated Nginx integration configuration." },
]],
["scripts/test-install-tht.sh", [
{ sha256: "37f18ce7ce93cb8b84f3b3708462cc16d50fdc7bab22836c382dbacf8382f05f", rationale: "Generates the reviewed synthetic tht installer artifact." },
]],
["scripts/test-server-pi-state-topology.sh", [
{ sha256: "6f746f7e8442b0a6ea0e216607a6a923d94b24cd8fa17fa2d1dac56e6f14f7ef", rationale: "Generates the isolated server topology test environment." },
{ sha256: "a9ab86c9408b22afb87f10f615570bc7ec09bc3e7eb8131f6d835afba9a6b1d8", rationale: "Generates the isolated server topology test environment, including its authentication configuration root." },
]],
["scripts/test-vector-backup-restore-safety.sh", [
{ sha256: "40b8a10a3c06aaa98e324fbf688b7d1f5cead330d7ba7eef98e06256d412a85a", rationale: "Generates the reviewed restore safety manifest." },
]],
["scripts/test-windows-clone-contract.ps1", [
{ sha256: "80f4880576a0679cb58e7b92600e7a90550c93c254553a2d4b299539f9ff0bcf", rationale: "Generates reviewed Windows clone test configuration." },
{ sha256: "6166294bdc8a8bf6436ad402bcbf7cae0f3b67dc6051cecfcca79267a62b082c", rationale: "Same reviewed block in the repository-required CRLF checkout representation." },
{ sha256: "3216201d59400ed7d1ec23e536634b8235a2e78b336e45b4dc598624920f0057", rationale: "Generates reviewed Windows clone test configuration." },
{ sha256: "a4044bb38b27e8120e90d65a0695fe0afd7757c067ae8dd67f170edf569a1de0", rationale: "Same reviewed block in the repository-required CRLF checkout representation." },
{ sha256: "3204f772d33cad42bcac99191507051aefb2c91d2935bec6698b956e44f9bf45", rationale: "Generates reviewed Windows clone test configuration with its authentication configuration root." },
{ sha256: "f4814d842a7502b7ef30fd6b224d5cb17b0ffd6fb2367c41c49ac16587536d93", rationale: "Same reviewed block in the repository-required CRLF checkout representation." },
{ sha256: "6f25ce3b58cea47b74fe9319ed917d8089a2fb334bc0d469daa7e1f10865d870", rationale: "Generates the reviewed Windows Compose override for the canonical service topology." },
{ sha256: "45a3cf19f7ce697b858b63d27a4edc7fefa2414d0408e7b6d72a65c86d314f5b", rationale: "Same reviewed Compose override in the repository-required CRLF checkout representation." },
{ sha256: "5d0d1a3fc45e99b3aacaf4ee5dd09a6bee1937784375dfe4bcfaa4ae32cfb9de", rationale: "Generates reviewed Windows clone test configuration." },
{ sha256: "b903e5dae953ae1372f1a5276f12a92ed3dd632b897f3afe5e00c646d90a1b42", rationale: "Same reviewed block in the repository-required CRLF checkout representation." },
]],
["scripts/unified-deployment-smoke.sh", [
{ sha256: "ca0c17d9ff8dc0fbe018fc1c5510eb33bc667a936fbe44a9be2d311390576825", rationale: "Generates reviewed Task 13 runtime configuration." },
{ sha256: "31ec00cc315b52da4a3bb6e3fba2d40aef29cdcd090bbc5d14c31f1aebbcfd04", rationale: "Generates reviewed Task 13 runtime configuration." },
{ sha256: "d92822815357ce3424e1a6eb43923df2b37b4fd93a3b5465ee9dfc69559ab0ed", rationale: "Generates reviewed Task 13 runtime configuration." },
{ sha256: "d6b8b7b951936c0452a485e9ee3b18a61251556581d6f7a2ce66f994b5700695", rationale: "Generates reviewed Task 13 runtime configuration." },
{ sha256: "1d60bf140165a8fabfa0c3729e776136904717e67becf3e0ab68c70d8e37847e", rationale: "Generates reviewed Task 13 runtime configuration." },
{ sha256: "36d3d8a2362dbdc4fad90948d6c227586d749f56b9a4bc5b6b5a91bcbec6407b", rationale: "Generates the reviewed local Task 13 Compose override." },
{ sha256: "c556f7d910d0788e219b042957e6b307cb9925b43920c680535d0d3a6dcbdb25", rationale: "Generates the reviewed local Task 13 installation descriptor." },
{ sha256: "526006fa6d48a8080b3834723630c64de5005a67243e944ebf1da15212b4d654", rationale: "Generates the reviewed server Task 13 Compose override." },
{ sha256: "c57ae2205c21ead0c2015a353aaabb948fa4ddd9b78a2cdcdb71f48cf2db742d", rationale: "Generates the reviewed projected-auth server Task 13 installation descriptor." },
]],
["scripts/vector-backup.sh", [
{ sha256: "571899db49dfdcec8107fbe1e0a86a61e7581979d3c4c248c20546843e275bcf", rationale: "Generates the reviewed backup manifest inside the helper command." },
@@ -624,7 +624,6 @@ test("an in-band marker cannot authorize expandable content", async (t) => {
test("current exact reviewed expandable blocks pass only at their trusted paths", async (t) => {
const reviewedPaths = [
"scripts/preprocess-smoke.sh",
"scripts/test-server-pi-state-topology.sh",
"scripts/test-vector-backup-restore-safety.sh",
"scripts/test-windows-clone-contract.ps1",
@@ -636,19 +635,6 @@ test("current exact reviewed expandable blocks pass only at their trusted paths"
root: repositoryRoot,
entries: reviewedPaths.map((path) => entry("deployment_script", path)),
});
const root = await fixture(t);
const original = await readFile(join(repositoryRoot, "scripts/preprocess-smoke.sh"), "utf8");
await put(root, "scripts/copied-preprocess.sh", original);
await assert.rejects(
verifyEntries({ root, entries: [entry("deployment_script", "scripts/copied-preprocess.sh")] }),
/exact-content reviewed allowlist/,
);
await put(root, "scripts/preprocess-smoke.sh", original.replace('$tmp/smoke.yaml', '$tmp/other.yaml'));
await assert.rejects(
verifyEntries({ root, entries: [entry("deployment_script", "scripts/preprocess-smoke.sh")] }),
/exact-content reviewed allowlist/,
);
});
test("PowerShell tokenizer ignores opener text in comments and ordinary strings", async (t) => {
+1 -1
View File
@@ -581,7 +581,7 @@ function load(root: string): LoadedAuthConfig {
);
return {
value: selectedGeneration.value,
revision: `sha256:${selected.generation}`,
revision: selected.generation,
sourcePath: join(generationsPath, selected.generation, "auth.yaml"),
runtimeProjection: snapshot(
selected.generation,
+15 -2
View File
@@ -4,6 +4,15 @@ const GENERIC_MODEL_FAILURE =
"Model request failed. Check provider connectivity, then Resume the session.";
const SUBSCRIPTION_MODEL_FAILURE =
"The selected model is unavailable for the current subscription. Choose another model and start a new session.";
const PHASE_STARTED_NOTIFICATION_PREFIX = "__tht_phase_started__:";
function phaseStartedNotification(message: unknown): string | null {
if (typeof message !== "string" || !message.startsWith(PHASE_STARTED_NOTIFICATION_PREFIX)) {
return null;
}
const phase = message.slice(PHASE_STARTED_NOTIFICATION_PREFIX.length);
return /^F[1-8]$/.test(phase) ? phase : "";
}
function safeModelFailure(error: unknown): string {
const detail = typeof error === "string" ? error : "";
@@ -36,7 +45,7 @@ export type ClientEvent =
| { type: "activity_event"; activity: ToolActivity }
| { type: "usage"; usage: TokenUsage }
| { type: "info"; [k: string]: any }
| { type: "system_event"; event: string };
| { type: "system_event"; event: string; phase?: string };
export type TurnState = "idle" | "running" | "waiting" | "failed";
@@ -71,7 +80,11 @@ export class SessionBridge {
});
}
} else if (m.type === "extension_ui_request" && m.method === "notify") {
this.fan({ type: "info", level: m.notifyType ?? "info", text: m.message ?? "" });
const phase = phaseStartedNotification(m.message);
if (phase) this.fan({ type: "system_event", event: "phase_started", phase });
else if (phase === null) {
this.fan({ type: "info", level: m.notifyType ?? "info", text: m.message ?? "" });
}
} else if (m.type === "message_update" && m.assistantMessageEvent?.type === "text_delta") {
this.fan({ type: "text_delta", text: m.assistantMessageEvent.delta ?? "" });
} else if (m.type === "message_update" && m.assistantMessageEvent?.type === "thinking_delta") {
+7 -3
View File
@@ -176,7 +176,11 @@ function positiveDimension(value: string | undefined, fallback: number): number
return parsed;
}
export function loadConfig(env: Record<string, string | undefined>): AppConfig {
export function loadConfig(
env: Record<string, string | undefined>,
options: { surface?: "application" | "workspace-maintenance" } = {},
): AppConfig {
const applicationSurface = options.surface !== "workspace-maintenance";
const defaultAuthConfigFile = "/run/thothii-auth/auth.yaml";
const authConfigFile = absoluteAuthPath(env.THT_AUTH_CONFIG_FILE ?? defaultAuthConfigFile, "file");
const authStateRoot = absoluteAuthPath(env.THT_AUTH_STATE_ROOT ?? "/data/auth", "state root");
@@ -211,13 +215,13 @@ export function loadConfig(env: Record<string, string | undefined>): AppConfig {
throw new Error(`unsupported AUTH_MODE=${requestedMode}; use none, mock, or upstream`);
}
const nodeEnvironment = env.NODE_ENV ?? process.env.NODE_ENV;
if ((requestedMode === "none" || requestedMode === "mock")
if (applicationSurface && (requestedMode === "none" || requestedMode === "mock")
&& nodeEnvironment !== "development" && nodeEnvironment !== "test") {
throw new Error("production requires auth.yaml or AUTH_MODE=upstream");
}
authMode = requestedMode as "none" | "mock" | "upstream";
}
const publicExposure = env.THOTH_PUBLIC_EXPOSURE === "true";
const publicExposure = applicationSurface && env.THOTH_PUBLIC_EXPOSURE === "true";
if (publicExposure && authMode !== "oidc" && authMode !== "upstream") {
throw new Error("public exposure requires AUTH_MODE=upstream or configured OIDC behind a trusted proxy");
}
+3 -2
View File
@@ -1,6 +1,6 @@
import { readFileSync } from "node:fs";
import { homedir } from "node:os";
import { join } from "node:path";
import { join, resolve } from "node:path";
export interface PiEnabledModelsResult {
ids: string[];
@@ -43,7 +43,8 @@ function isExactCompositeId(value: unknown): value is string {
export function loadPiEnabledModels(opts: LoadOptions): PiEnabledModelsResult {
const warnings: string[] = [];
const read = opts.read ?? ((path: string) => readFileSync(path, "utf8"));
const agentDir = opts.agentDir ?? join(homedir(), ".pi", "agent");
const agentDir = opts.agentDir
?? resolve(process.env.PI_CODING_AGENT_DIR ?? join(homedir(), ".pi", "agent"));
const globalPath = join(agentDir, "settings.json");
const projectPath = join(opts.harnessDir, ".pi", "settings.json");
const globalSettings = readSettings(globalPath, false, read, warnings);
+30
View File
@@ -56,6 +56,34 @@ export function validateDeclarativePiConfig(raw: string): void {
assertDeclarativePiConfig(parsePiConfigJson(raw));
}
/** Return the selected provider's declarative apiKey value without knowing provider IDs in code. */
export function configuredPiProviderApiKey(
raw: string | undefined,
provider: string | undefined,
): string | undefined {
if (raw === undefined || provider === undefined) return undefined;
const parsed = parsePiConfigJson(raw);
if (!parsed || typeof parsed !== "object" || Array.isArray(parsed)) {
throw new PiManagedConfigError();
}
const providers = (parsed as { providers?: unknown }).providers;
if (!providers || typeof providers !== "object" || Array.isArray(providers)) {
throw new PiManagedConfigError();
}
const entry = Object.entries(providers as Record<string, unknown>)
.find(([id]) => id.trim().toLowerCase() === provider.trim().toLowerCase());
if (!entry) return undefined;
const config = entry[1];
if (!config || typeof config !== "object" || Array.isArray(config)) {
throw new PiManagedConfigError();
}
assertDeclarativePiConfig(config);
const apiKey = (config as { apiKey?: unknown }).apiKey;
if (apiKey === undefined) return undefined;
if (typeof apiKey !== "string" || apiKey.length === 0) throw new PiManagedConfigError();
return apiKey;
}
function configuredPiAgentDir(): string {
return resolve(process.env.PI_CODING_AGENT_DIR ?? join(homedir(), ".pi", "agent"));
}
@@ -102,6 +130,7 @@ export function readConfiguredPiAgentFile(
export interface PiRuntimeAgentSnapshot {
agentDir: string;
sessionDir: string;
models?: string;
cleanup: () => void;
}
@@ -152,6 +181,7 @@ export function createPiRuntimeAgentSnapshot(): PiRuntimeAgentSnapshot {
return {
agentDir: snapshotDir,
sessionDir: process.env.PI_CODING_AGENT_SESSION_DIR || join(sourceAgentDir, "sessions"),
models,
cleanup: () => {
if (cleaned) return;
cleaned = true;
+6
View File
@@ -9,8 +9,10 @@ import {
} from "../settings/settings-store.js";
import type { PiModel } from "./list-models.js";
import {
configuredPiProviderApiKey,
PI_MANAGED_CONFIG_ERROR_MESSAGE,
isPiManagedConfigError,
readConfiguredPiAgentFile,
} from "./managed-config.js";
import { createPiProviderSmoke, type PiProviderSmoke } from "./provider-smoke.js";
import { loadPiAuthProviders } from "./auth-providers.js";
@@ -119,6 +121,10 @@ export function createPiManagement(config: AppConfig, deps: PiManagementDeps): P
authProviders: loadPiAuthProviders(),
resolveCredentialValue: () => secretValue(config, "THT_MODEL_API_KEY"),
credentialFile: config.modelApiKeyFile,
configuredApiKey: configuredPiProviderApiKey(
readConfiguredPiAgentFile("models.json", true),
provider,
),
});
} catch {
return "missing";
+5 -1
View File
@@ -7,7 +7,10 @@ import { buildPiChildEnv, canonicalPiProvider } from "./provider-credentials.js"
import { loadPiAuthProviders } from "./auth-providers.js";
import { secretValue } from "../config/secret-bundle.js";
import { clearPrincipalEnvironment, principalEnvironment, type PrincipalContext } from "../auth/principal.js";
import { createPiRuntimeAgentSnapshot } from "./managed-config.js";
import {
configuredPiProviderApiKey,
createPiRuntimeAgentSnapshot,
} from "./managed-config.js";
export interface SessionRuntime {
rpc: RpcClient;
@@ -81,6 +84,7 @@ export class PiProcessManager {
authProviders: this.loadAuthProviders(agent.agentDir),
credentialValue: secretValue(this.cfg, "THT_MODEL_API_KEY"),
credentialFile: this.cfg.modelApiKeyFile,
configuredApiKey: configuredPiProviderApiKey(agent.models, provider),
additions: { THT_SESSION: sessionId, THT_AUTHOR: author },
});
env.PI_CODING_AGENT_DIR = agent.agentDir;
+23 -7
View File
@@ -42,9 +42,14 @@ const PROVIDER_KEY_ENV: Readonly<Record<string, string>> = {
const COMPOUND_PROVIDERS = new Set([
"amazon-bedrock", "azure-openai-responses", "cloudflare-ai-gateway", "cloudflare-workers-ai",
]);
const LOCAL_PROVIDERS = new Set([
"ollama", "lmstudio", "local", "aritmolab", "local-qwen", "faux",
]);
function configuredCredentialEnv(apiKey: string | undefined): string | null | undefined {
if (apiKey === undefined) return undefined;
if (!apiKey.startsWith("$")) return null;
const matched = /^\$(?:\{([A-Z][A-Z0-9_]*(?:API_KEY|TOKEN))\}|([A-Z][A-Z0-9_]*(?:API_KEY|TOKEN)))$/.exec(apiKey);
if (!matched) throw new Error("model provider credential is unavailable");
return matched[1] ?? matched[2];
}
export function canonicalPiProvider(provider: string | undefined): string | undefined {
const value = provider?.trim().toLowerCase();
@@ -106,6 +111,8 @@ export function buildPiChildEnv(opts: {
credentialFile?: string;
additions?: NodeJS.ProcessEnv;
credentialValue?: string;
/** Exact apiKey declaration from the selected provider in models.json. */
configuredApiKey?: string;
fsOps?: CredentialFsOps;
/**
* Providers pi can authenticate from its own auth store. For these, the single
@@ -147,17 +154,22 @@ export function buildPiChildEnv(opts: {
delete env.THT_SSL_CA_FILE;
for (const name of PI_0803_CREDENTIAL_ENV_NAMES) delete env[name];
const provider = canonicalPiProvider(opts.provider);
const configuredEnv = configuredCredentialEnv(opts.configuredApiKey);
if (configuredEnv) delete env[configuredEnv];
if (provider && COMPOUND_PROVIDERS.has(provider)) {
throw new Error(
"compound credential bundles are unsupported by THT_MODEL_API_KEY_FILE; "
+ "dedicated provider configuration is required",
);
}
if (provider && !LOCAL_PROVIDERS.has(provider)) {
if (provider) {
// pi self-authenticates this provider from its own auth store; injecting the
// single managed key here would force one provider's key onto another.
if (opts.authProviders?.has(provider)) return env;
const envName = PROVIDER_KEY_ENV[provider];
// A literal apiKey is entirely owned by models.json (commonly a non-secret
// placeholder for a local OpenAI-compatible endpoint) and needs no managed key.
if (configuredEnv === null) return env;
const envName = configuredEnv ?? PROVIDER_KEY_ENV[provider];
if (!envName || (!opts.credentialFile && opts.credentialValue === undefined)) {
throw new Error("model provider credential is unavailable");
}
@@ -169,7 +181,7 @@ export function buildPiChildEnv(opts: {
}
else if (opts.credentialFile) env[envName] = readCredential(opts.credentialFile, opts.fsOps ?? realFs);
else throw new Error("model provider credential is unavailable");
} else if (opts.credentialFile && !provider) {
} else if (opts.credentialFile) {
throw new Error("model provider credential is unavailable");
}
return env;
@@ -181,11 +193,14 @@ export function piProviderCredentialStatus(opts: {
credentialFile?: string;
resolveCredentialValue?: () => string | undefined;
authProviders?: ReadonlySet<string>;
configuredApiKey?: string;
fsOps?: CredentialFsOps;
}): PiCredentialStatus {
const provider = canonicalPiProvider(opts.provider);
if (!provider || LOCAL_PROVIDERS.has(provider)) return "missing";
if (!provider) return "missing";
if (opts.authProviders?.has(provider)) return "present";
const configuredEnv = configuredCredentialEnv(opts.configuredApiKey);
if (configuredEnv === null) return "missing";
try {
buildPiChildEnv({
ambient: {},
@@ -193,6 +208,7 @@ export function piProviderCredentialStatus(opts: {
credentialFile: opts.credentialFile,
credentialValue: opts.resolveCredentialValue?.(),
authProviders: opts.authProviders,
configuredApiKey: opts.configuredApiKey,
fsOps: opts.fsOps,
});
return "present";
+5 -3
View File
@@ -11,6 +11,7 @@ import { buildPiChildEnv, canonicalPiProvider } from "./provider-credentials.js"
import type { PiReasoning } from "./management.js";
import {
PiManagedConfigError,
configuredPiProviderApiKey,
isPiManagedConfigError,
parsePiConfigJson,
readConfiguredPiAgentFile,
@@ -65,11 +66,15 @@ export function createPiProviderSmoke(
const canonicalProvider = canonicalPiProvider(provider);
if (!canonicalProvider || timeoutMs <= 0) throw providerFailure();
const configuredAuthProviders = authProviders();
const configuredModels = options.readModelsStore
? options.readModelsStore()
: readConfiguredPiAgentFile("models.json", true);
const env = buildPiChildEnv({
provider: canonicalProvider,
authProviders: configuredAuthProviders,
credentialValue: secretValue(config, "THT_MODEL_API_KEY"),
credentialFile: config.modelApiKeyFile,
configuredApiKey: configuredPiProviderApiKey(configuredModels, canonicalProvider),
});
clearPrincipalEnvironment(env);
delete env.THT_DATA_ROOT;
@@ -90,9 +95,6 @@ export function createPiProviderSmoke(
);
writeDeclarativeAgentConfig(join(isolatedAgentDir, "auth.json"), authStore);
}
const configuredModels = options.readModelsStore
? options.readModelsStore()
: readConfiguredPiAgentFile("models.json", true);
if (configuredModels !== undefined) {
const modelsStore = selectedProviderModelsStore(
configuredModels,
+2 -2
View File
@@ -20,7 +20,7 @@ import {
validateOperationalWorkspace,
type WorkspaceDescriptor,
} from "../workspaces/schema.js";
import { reconcileCollection } from "../workspaces/qdrant-collection.js";
import { reconcileCollection, type CollectionMode } from "../workspaces/qdrant-collection.js";
import type { WorkspaceSecretStore } from "../workspaces/secret-store.js";
export interface ThtConfig extends SecretBundleConfig {
@@ -575,7 +575,7 @@ export class ThtRunner {
async qdrantEnsure(
workspace: WorkspaceDescriptor,
timeoutSec: number,
mode: "self_heal" | "require_existing" = "require_existing",
mode: CollectionMode = "require_existing",
): Promise<QdrantEnsureResult> {
let descriptor;
try {
+5 -1
View File
@@ -212,7 +212,7 @@ async function readSessionInventory(dataRoot: string, workspaceId: string): Prom
}
function createProductionService(): WorkspacePreprocessingService {
const config = loadConfig(process.env);
const config = loadConfig(process.env, { surface: "workspace-maintenance" });
const registry = new WorkspaceRegistry(config.workspaceRegistry);
const workspaceSecretStore = new WorkspaceSecretStore({
root: config.workspaceSecretStoreRoot,
@@ -307,6 +307,10 @@ function createProductionService(): WorkspacePreprocessingService {
const result = await runner.qdrantEnsure(workspace, 30);
return result.ok ? { ok: true as const } : { ok: false as const, code: result.code ?? "workspace_not_activatable" };
},
evidencePreflight: async (workspace) => {
const result = await runner.qdrantEnsure(workspace, 30, "evidence_maintenance");
return result.ok ? { ok: true as const } : { ok: false as const, code: result.code ?? "workspace_not_activatable" };
},
});
}
@@ -9,7 +9,7 @@ import {
writeFileSync,
} from "node:fs";
import { dirname, isAbsolute, join } from "node:path";
import { GitWorkspaceRepository } from "./git-repository.js";
import type { GitWorkspaceRepository } from "../git-repository.js";
export interface EvidenceMaterializationLimits {
maxEntries: number;
@@ -81,7 +81,10 @@ function writeExclusiveNoFollow(path: string, contents: Buffer, mode: number): v
}
export interface MaterializeEvidenceTreeOptions {
repository: GitWorkspaceRepository;
repository: Pick<
GitWorkspaceRepository,
"evidenceTreeObjects" | "evidenceTreeId" | "gitObjectSize" | "evidenceBlobBytes"
>;
revision: string;
id: string;
/** The workspace directory (e.g. `<staging>/<id>`) that will receive `evidence/` and the manifest. */
@@ -0,0 +1,177 @@
import { isIP } from "node:net";
import type { WorkspaceDescriptor } from "../schema.js";
type EvidenceConfig = WorkspaceDescriptor["evidence"];
type SemanticFailureCode = "workspace_not_activatable" | "semantic_index_incompatible";
export interface EvidenceJobState {
runId: string;
completedStages: string[];
childRuns: Record<string, string>;
}
export interface EvidencePreprocessingDependencies {
runStage(argv: string[]): Promise<Record<string, unknown>>;
persistJob(): void;
evidencePreflight(): Promise<{ ok: true } | { ok: false; code: SemanticFailureCode }>;
requireRunId(value: unknown): string;
numberRecord(value: unknown): Record<string, number> | undefined;
}
export interface EvidencePreprocessingRequest {
evidence: EvidenceConfig;
job: EvidenceJobState;
dryRun?: boolean;
httpPrivateHostAllowlist?: readonly string[];
}
export interface EvidencePreprocessingOutcome {
status: "succeeded" | "unchanged" | "dry_run" | "failed";
code: "ok" | "egress_policy_refused" | SemanticFailureCode;
runId?: string;
childRuns?: Record<string, string>;
completedStages?: string[];
counts?: Record<string, number>;
warnings?: string[];
}
function isPrivateHost(hostname: string): boolean {
if (hostname === "localhost" || hostname === "metadata.google.internal") return true;
const address = isIP(hostname);
if (address === 4) {
if (/^127\./.test(hostname) || /^10\./.test(hostname) || /^192\.168\./.test(hostname)) {
return true;
}
if (/^169\.254\./.test(hostname) || /^0\./.test(hostname)) return true;
const match = /^172\.(\d+)\./.exec(hostname);
return Boolean(match && Number(match[1]) >= 16 && Number(match[1]) <= 31);
}
if (address === 6) {
const normalized = hostname.toLowerCase();
return normalized === "::1"
|| normalized.startsWith("fe80:")
|| normalized.startsWith("fd")
|| normalized.startsWith("fc");
}
return hostname.endsWith(".internal");
}
function evidencePolicy(
evidence: EvidenceConfig,
httpPrivateHostAllowlist?: readonly string[],
): EvidencePreprocessingOutcome | undefined {
if (!evidence || evidence.source.type === "filesystem") return undefined;
if (evidence.source.type === "http") {
for (const value of evidence.source.uris) {
const host = new URL(value).hostname;
if (
isPrivateHost(host)
&& !(evidence.source.allow_private_hosts && httpPrivateHostAllowlist?.includes(host))
) {
return { status: "failed", code: "egress_policy_refused" };
}
}
return undefined;
}
if (
evidence.source.endpoint_url !== undefined
|| evidence.source.credentials === "ambient"
|| evidence.source.allow_private_endpoint
|| evidence.source.allow_insecure_endpoint
) {
return { status: "failed", code: "egress_policy_refused" };
}
return undefined;
}
function jobResult(job: EvidenceJobState): Pick<
EvidencePreprocessingOutcome,
"runId" | "childRuns" | "completedStages"
> {
return {
runId: job.runId,
childRuns: { ...job.childRuns },
completedStages: [...job.completedStages],
};
}
async function runEvidenceStage(
request: EvidencePreprocessingRequest,
deps: EvidencePreprocessingDependencies,
): Promise<EvidencePreprocessingOutcome> {
const payload = await deps.runStage([
"preprocess",
"evidence",
...(request.dryRun ? ["--dry-run"] : []),
...(request.job.childRuns.evidence
? ["--resume", request.job.childRuns.evidence]
: []),
"--json",
"-c",
"/dev/fd/3",
]);
if (typeof payload.run_id === "string") {
request.job.childRuns.evidence = deps.requireRunId(payload.run_id);
}
if (!request.dryRun && !request.job.completedStages.includes("evidence")) {
request.job.completedStages.push("evidence");
}
deps.persistJob();
return {
status: request.dryRun ? "dry_run" : "succeeded",
code: "ok",
...jobResult(request.job),
counts: deps.numberRecord(payload.counts),
};
}
export async function preprocessEvidence(
request: EvidencePreprocessingRequest,
deps: EvidencePreprocessingDependencies,
): Promise<EvidencePreprocessingOutcome> {
if (!request.evidence) {
return {
status: "unchanged",
code: "ok",
warnings: ["workspace has no Evidence source"],
};
}
const policy = evidencePolicy(request.evidence, request.httpPrivateHostAllowlist);
if (policy) return policy;
const preflight = await deps.evidencePreflight();
if (!preflight.ok) {
return { status: "failed", code: preflight.code, runId: request.job.runId };
}
if (request.job.completedStages.includes("evidence") && !request.dryRun) {
return {
status: "unchanged",
code: "ok",
runId: request.job.runId,
completedStages: [...request.job.completedStages],
};
}
return await runEvidenceStage(request, deps);
}
export async function continueEvidencePreprocessing(
request: Omit<EvidencePreprocessingRequest, "dryRun"> & {
priorCounts?: Record<string, number>;
},
deps: EvidencePreprocessingDependencies,
): Promise<EvidencePreprocessingOutcome> {
if (!request.evidence) {
return {
status: "succeeded",
code: "ok",
...jobResult(request.job),
warnings: ["workspace has no Evidence source"],
...(request.priorCounts ? { counts: request.priorCounts } : {}),
};
}
const policy = evidencePolicy(request.evidence, request.httpPrivateHostAllowlist);
if (policy) return { ...policy, ...jobResult(request.job) };
if (!request.job.completedStages.includes("evidence")) {
return await runEvidenceStage(request, deps);
}
return { status: "unchanged", code: "ok", ...jobResult(request.job) };
}
+50 -128
View File
@@ -1,7 +1,12 @@
import { createHash, randomBytes } from "node:crypto";
import { readdirSync, readFileSync, rmSync, writeFileSync, mkdirSync } from "node:fs";
import { isIP } from "node:net";
import { join } from "node:path";
import {
continueEvidencePreprocessing,
preprocessEvidence as runEvidencePreprocessing,
type EvidencePreprocessingDependencies,
type EvidencePreprocessingOutcome,
} from "./evidence/preprocessing.js";
import type { WorkspaceDescriptor } from "./schema.js";
import {
PreprocessingStateStore,
@@ -68,6 +73,9 @@ export interface WorkspacePreprocessingServiceDeps {
semanticPreflight(workspace: WorkspaceDescriptor): Promise<
{ ok: true } | { ok: false; code: "workspace_not_activatable" | "semantic_index_incompatible" }
>;
evidencePreflight(workspace: WorkspaceDescriptor): Promise<
{ ok: true } | { ok: false; code: "workspace_not_activatable" | "semantic_index_incompatible" }
>;
httpPrivateHostAllowlist?: readonly string[];
}
@@ -104,26 +112,6 @@ function baseResult(
};
}
function isPrivateHost(hostname: string): boolean {
if (hostname === "localhost" || hostname === "metadata.google.internal") return true;
const address = isIP(hostname);
if (address === 4) {
if (/^127\./.test(hostname) || /^10\./.test(hostname) || /^192\.168\./.test(hostname)) return true;
if (/^169\.254\./.test(hostname) || /^0\./.test(hostname)) return true;
const match = /^172\.(\d+)\./.exec(hostname);
return Boolean(match && Number(match[1]) >= 16 && Number(match[1]) <= 31);
}
if (address === 6) {
const normalized = hostname.toLowerCase();
return normalized === "::1" || normalized.startsWith("fe80:") || normalized.startsWith("fd") || normalized.startsWith("fc");
}
return hostname.endsWith(".internal");
}
function noEvidenceWarning(workspace: WorkspaceDescriptor): string[] {
return workspace.evidence === undefined ? ["workspace has no Evidence source"] : [];
}
export class WorkspacePreprocessingService {
constructor(private readonly deps: WorkspacePreprocessingServiceDeps) {}
@@ -332,36 +320,16 @@ export class WorkspacePreprocessingService {
async preprocessEvidence(options: { workspaceId: string; dryRun?: boolean; resumeRunId?: string }): Promise<WorkspaceOperationResult> {
const scope = await this.startRun(options.workspaceId, "preprocess evidence", options.resumeRunId);
if (scope.runtime.workspace.evidence === undefined) {
return baseResult(scope.runtime, "preprocess evidence", "unchanged", "ok", {
warnings: noEvidenceWarning(scope.runtime.workspace),
});
}
const policy = this.evidencePolicy(scope.runtime.workspace);
if (policy !== undefined) return baseResult(scope.runtime, "preprocess evidence", policy.status, policy.code, { warnings: policy.warnings });
const semantic = await this.deps.semanticPreflight(scope.runtime.workspace);
if (!semantic.ok) return baseResult(scope.runtime, "preprocess evidence", "failed", semantic.code, { runId: scope.job.runId });
if (scope.job.completedStages.includes("evidence") && !options.dryRun) {
return baseResult(scope.runtime, "preprocess evidence", "unchanged", "ok", {
runId: scope.job.runId,
completedStages: [...scope.job.completedStages],
});
}
const payload = await this.runJsonStage(scope.runtime, [
"preprocess", "evidence",
...(options.dryRun ? ["--dry-run"] : []),
...(scope.job.childRuns.evidence ? ["--resume", scope.job.childRuns.evidence] : []),
"--json", "-c", "/dev/fd/3",
]);
if (typeof payload.run_id === "string") scope.job.childRuns.evidence = this.requireRunId(payload.run_id);
if (!options.dryRun && !scope.job.completedStages.includes("evidence")) scope.job.completedStages.push("evidence");
this.state(scope.runtime.workspaceId).writeJob(scope.job);
return baseResult(scope.runtime, "preprocess evidence", options.dryRun ? "dry_run" : "succeeded", "ok", {
runId: scope.job.runId,
childRuns: { ...scope.job.childRuns },
completedStages: [...scope.job.completedStages],
counts: this.numberRecord(payload.counts),
});
const outcome = await runEvidencePreprocessing(
{
evidence: scope.runtime.workspace.evidence,
job: scope.job,
dryRun: options.dryRun,
httpPrivateHostAllowlist: this.deps.httpPrivateHostAllowlist,
},
this.evidenceDependencies(scope),
);
return this.evidenceResult(scope, "preprocess evidence", outcome);
}
async run(options: { workspaceId: string; resumeRunId?: string }): Promise<WorkspaceOperationResult> {
@@ -401,59 +369,23 @@ export class WorkspacePreprocessingService {
}
const semantic = await this.deps.semanticPreflight(scope.runtime.workspace);
if (!semantic.ok) return baseResult(scope.runtime, "preprocess run", "failed", semantic.code, { runId: scope.job.runId });
let schemaCounts: Record<string, number> | undefined;
if (!scope.job.completedStages.includes("schema_index")) {
const payload = await this.runJsonStage(scope.runtime, ["vector", "index-schema", "--json", "-c", "/dev/fd/3"]);
scope.job.completedStages.push("schema_index");
this.state(scope.runtime.workspaceId).writeJob(scope.job);
const warnings = noEvidenceWarning(scope.runtime.workspace);
if (scope.runtime.workspace.evidence === undefined) {
return baseResult(scope.runtime, "preprocess run", "succeeded", "ok", {
runId: scope.job.runId,
childRuns: { ...scope.job.childRuns },
completedStages: [...scope.job.completedStages],
counts: this.numberRecord(payload.counts),
warnings,
});
}
schemaCounts = this.numberRecord(payload.counts);
}
if (scope.runtime.workspace.evidence === undefined) {
return baseResult(scope.runtime, "preprocess run", "succeeded", "ok", {
runId: scope.job.runId,
childRuns: { ...scope.job.childRuns },
completedStages: [...scope.job.completedStages],
warnings: noEvidenceWarning(scope.runtime.workspace),
});
}
const policy = this.evidencePolicy(scope.runtime.workspace);
if (policy !== undefined) {
return baseResult(scope.runtime, "preprocess run", policy.status, policy.code, {
runId: scope.job.runId,
childRuns: { ...scope.job.childRuns },
completedStages: [...scope.job.completedStages],
warnings: policy.warnings,
});
}
if (!scope.job.completedStages.includes("evidence")) {
const payload = await this.runJsonStage(scope.runtime, [
"preprocess", "evidence",
...(scope.job.childRuns.evidence ? ["--resume", scope.job.childRuns.evidence] : []),
"--json", "-c", "/dev/fd/3",
]);
if (typeof payload.run_id === "string") scope.job.childRuns.evidence = this.requireRunId(payload.run_id);
scope.job.completedStages.push("evidence");
this.state(scope.runtime.workspaceId).writeJob(scope.job);
return baseResult(scope.runtime, "preprocess run", "succeeded", "ok", {
runId: scope.job.runId,
childRuns: { ...scope.job.childRuns },
completedStages: [...scope.job.completedStages],
counts: this.numberRecord(payload.counts),
});
}
return baseResult(scope.runtime, "preprocess run", "unchanged", "ok", {
runId: scope.job.runId,
childRuns: { ...scope.job.childRuns },
completedStages: [...scope.job.completedStages],
});
const outcome = await continueEvidencePreprocessing(
{
evidence: scope.runtime.workspace.evidence,
job: scope.job,
httpPrivateHostAllowlist: this.deps.httpPrivateHostAllowlist,
priorCounts: schemaCounts,
},
this.evidenceDependencies(scope),
);
return this.evidenceResult(scope, "preprocess run", outcome);
}
private async startRun(workspaceId: string, operation: string, resumeRunId?: string): Promise<RunScope> {
@@ -476,6 +408,25 @@ export class WorkspacePreprocessingService {
return new PreprocessingStateStore({ dataRoot: this.deps.dataRoot, workspaceId });
}
private evidenceDependencies(scope: RunScope): EvidencePreprocessingDependencies {
return {
runStage: async (argv) => await this.runJsonStage(scope.runtime, argv),
persistJob: () => this.state(scope.runtime.workspaceId).writeJob(scope.job),
evidencePreflight: async () => await this.deps.evidencePreflight(scope.runtime.workspace),
requireRunId: (value) => this.requireRunId(value),
numberRecord: (value) => this.numberRecord(value),
};
}
private evidenceResult(
scope: RunScope,
operation: "preprocess evidence" | "preprocess run",
outcome: EvidencePreprocessingOutcome,
): WorkspaceOperationResult {
const { status, code, ...extra } = outcome;
return baseResult(scope.runtime, operation, status, code, extra);
}
private async runSuggestStage(
scope: RunScope,
fromSql: ReadonlyArray<{ name: string; sql: string }>,
@@ -581,33 +532,4 @@ export class WorkspacePreprocessingService {
return undefined;
}
private evidencePolicy(workspace: WorkspaceDescriptor): {
status: WorkspaceOperationResult["status"];
code: WorkspaceOperationResult["code"];
warnings?: string[];
} | undefined {
const evidence = workspace.evidence;
if (!evidence) return undefined;
// P6: filesystem Evidence is materialized from the pinned commit at activation, so the
// engine may proceed directly against the immutable revision content root.
if (evidence.source.type === "filesystem") return undefined;
if (evidence.source.type === "http") {
for (const value of evidence.source.uris) {
const host = new URL(value).hostname;
if (isPrivateHost(host) && !(evidence.source.allow_private_hosts && this.deps.httpPrivateHostAllowlist?.includes(host))) {
return { status: "failed", code: "egress_policy_refused" };
}
}
return undefined;
}
if (
evidence.source.endpoint_url !== undefined
|| evidence.source.credentials === "ambient"
|| evidence.source.allow_private_endpoint
|| evidence.source.allow_insecure_endpoint
) {
return { status: "failed", code: "egress_policy_refused" };
}
return undefined;
}
}
+53 -8
View File
@@ -3,12 +3,12 @@ export const QDRANT_REQUIRED_INDEXES = Object.freeze([
"record_kind", "vector_generation", "workspace_id", "workspace_revision",
]);
export type CollectionMode = "self_heal" | "require_existing";
export type CollectionMode = "self_heal" | "require_existing" | "evidence_maintenance";
export interface CollectionCheck {
ok: boolean;
code?: "semantic_index_incompatible" | "workspace_not_activatable";
state?: "ready" | "created" | "repaired";
state?: "ready" | "created" | "repaired" | "upgraded";
}
export interface ReconcileCollectionOptions {
@@ -52,12 +52,43 @@ async function createCollection(opts: ReconcileCollectionOptions, request: typeo
const res = await request(qdrantUrl(opts.baseUrl, `/collections/${encodeURIComponent(opts.collection)}`), {
method: "PUT",
headers: { "content-type": "application/json" },
body: JSON.stringify({ vectors: { size: opts.dimensions, distance: qdrantDistance(opts.distance) } }),
body: JSON.stringify({
vectors: { size: opts.dimensions, distance: qdrantDistance(opts.distance) },
...(opts.mode === "evidence_maintenance" ? { sparse_vectors: { bm25: { modifier: "idf" } } } : {}),
}),
signal: opts.signal,
});
if (!res.ok && res.status !== 409) throw new Error("qdrant collection creation failed");
}
type EvidenceSparseCompatibility = "compatible" | "upgradeable" | "incompatible";
function evidenceSparseCompatibility(info: any): EvidenceSparseCompatibility {
const sparseVectors = info?.config?.params?.sparse_vectors;
if (sparseVectors === undefined) return "upgradeable";
if (!sparseVectors || typeof sparseVectors !== "object" || Array.isArray(sparseVectors)) {
return "incompatible";
}
const bm25 = sparseVectors.bm25;
if (bm25 === undefined) return "upgradeable";
return typeof bm25 === "object" && bm25 !== null && !Array.isArray(bm25)
&& typeof bm25.modifier === "string" && bm25.modifier.toLowerCase() === "idf"
? "compatible" : "incompatible";
}
async function createBm25Vector(opts: ReconcileCollectionOptions, request: typeof fetch): Promise<void> {
const res = await request(
qdrantUrl(opts.baseUrl, `/collections/${encodeURIComponent(opts.collection)}/vectors/bm25`),
{
method: "PUT",
headers: { "content-type": "application/json" },
body: JSON.stringify({ sparse: { modifier: "idf" } }),
signal: opts.signal,
},
);
if (!res.ok && res.status !== 409) throw new Error("qdrant BM25 vector creation failed");
}
async function createIndex(opts: ReconcileCollectionOptions, field: string, request: typeof fetch): Promise<void> {
const res = await request(qdrantUrl(opts.baseUrl, `/collections/${encodeURIComponent(opts.collection)}/index`), {
method: "PUT",
@@ -68,13 +99,14 @@ async function createIndex(opts: ReconcileCollectionOptions, field: string, requ
if (!res.ok && res.status !== 409) throw new Error("qdrant index creation failed");
}
/** Reconcile a Qdrant collection: self-heal creates missing collections/indexes; require_existing
* only validates and refuses incompatible contracts (never mutates). */
/** Reconcile a Qdrant collection. Only Evidence maintenance may add the BM25 sparse vector;
* session admission remains limited to the existing dense/index self-heal behavior. */
export async function reconcileCollection(opts: ReconcileCollectionOptions): Promise<CollectionCheck> {
const request = opts.request ?? fetch;
const allowsMutation = opts.mode === "self_heal" || opts.mode === "evidence_maintenance";
let info = await collectionInfo(opts, request);
if (info === undefined) {
if (opts.mode !== "self_heal") return { ok: false, code: "semantic_index_incompatible" };
if (!allowsMutation) return { ok: false, code: "semantic_index_incompatible" };
await createCollection(opts, request);
// Tolerate an already-compatible concurrent creator: re-read the final state.
info = await collectionInfo(opts, request);
@@ -83,9 +115,22 @@ export async function reconcileCollection(opts: ReconcileCollectionOptions): Pro
if (!vectorCompatibility(info, opts)) {
return { ok: false, code: "semantic_index_incompatible" };
}
let bm25Added = false;
if (opts.mode === "evidence_maintenance") {
const sparse = evidenceSparseCompatibility(info);
if (sparse === "incompatible") return { ok: false, code: "semantic_index_incompatible" };
if (sparse === "upgradeable") {
await createBm25Vector(opts, request);
info = await collectionInfo(opts, request);
if (!vectorCompatibility(info, opts) || evidenceSparseCompatibility(info) !== "compatible") {
return { ok: false, code: "semantic_index_incompatible" };
}
bm25Added = true;
}
}
const missing = await missingIndexes(opts, info);
if (missing.length > 0) {
if (opts.mode !== "self_heal") return { ok: false, code: "semantic_index_incompatible" };
if (!allowsMutation) return { ok: false, code: "semantic_index_incompatible" };
for (const field of missing) await createIndex(opts, field, request);
// Qdrant payload indexes become visible asynchronously: poll until the
// contract is complete or a bounded deadline passes (fail closed).
@@ -101,5 +146,5 @@ export async function reconcileCollection(opts: ReconcileCollectionOptions): Pro
}
return { ok: false, code: "semantic_index_incompatible" };
}
return { ok: true, state: "ready" };
return { ok: true, state: bm25Added ? "upgraded" : "ready" };
}
+1 -1
View File
@@ -5,7 +5,7 @@ import { isAbsolute, join } from "node:path";
import { buildInstallationContract, renderWorkspaceDocs } from "./contracts.js";
import { parseAnnotationsYaml } from "./annotations.js";
import { syncAnnotations } from "./annotations-sync.js";
import { materializeEvidenceTree } from "./evidence-materialization.js";
import { materializeEvidenceTree } from "./evidence/materialization.js";
import { assertCatalogMatchesDescriptor, parseWorkspaceCatalogYaml, type WorkspaceCatalog, type WorkspaceCatalogEntry } from "./catalog.js";
import {
GitWorkspaceRepository,
+4 -1
View File
@@ -174,7 +174,10 @@ function renderEvidence(
}
return {
evidence: { sources: [renderedSource] },
evidence: {
...(workspace.evidence.schema_version === 2 ? { schema_version: 2 } : {}),
sources: [renderedSource],
},
vector: {
max_chunk_chars: workspace.evidence.policy.max_chunk_chars,
retain_published_generations: workspace.evidence.policy.retain_published_generations,
+31 -3
View File
@@ -115,6 +115,7 @@ export type EvidenceSource =
};
export interface WorkspaceEvidence {
schema_version: 1 | 2;
source: EvidenceSource;
policy: EvidencePolicy;
}
@@ -246,10 +247,10 @@ const evidencePattern = z.string().refine(isSafeEvidencePattern, {
const filesystemEvidenceSourceSchema = z.object({
type: z.literal("filesystem"),
uri: z.string(),
patterns: z.array(evidencePattern).min(1).default(["**/*.md"]),
patterns: z.array(evidencePattern).min(1).optional(),
max_bytes: positiveSafeInteger.default(10 * 1024 * 1024),
}).strict().superRefine((source, context) => {
if (new Set(source.patterns).size !== source.patterns.length) {
if (source.patterns !== undefined && new Set(source.patterns).size !== source.patterns.length) {
context.addIssue({ code: "custom", path: ["patterns"], message: "evidence patterns must not repeat" });
}
});
@@ -323,12 +324,39 @@ const evidencePolicySchema = z.object({
retain_published_generations: positiveSafeInteger.default(3),
}).strict();
const workspaceEvidenceSchema = z.object({
schema_version: z.union([z.literal(1), z.literal(2)]).default(1),
source: evidenceSourceSchema,
policy: evidencePolicySchema.default({
max_chunk_chars: 4_000,
retain_published_generations: 3,
}),
}).strict();
}).strict().superRefine((evidence, context) => {
if (evidence.schema_version !== 2 || evidence.source.type !== "filesystem") return;
const patterns = evidence.source.patterns ?? ["curated/**/*.md"];
const selectsSource = patterns.some((pattern) => pattern === "source" || pattern.startsWith("source/"));
const selectsCurated = patterns.some((pattern) => pattern === "curated" || pattern.startsWith("curated/"));
if (selectsSource && selectsCurated) {
context.addIssue({
code: "custom",
path: ["source", "patterns"],
message: "schema-versioned filesystem Evidence patterns cannot span source and curated",
});
} else if (patterns.length !== 1 || patterns[0] !== "curated/**/*.md") {
context.addIssue({
code: "custom",
path: ["source", "patterns"],
message: "schema-versioned filesystem Evidence patterns must be exactly curated/**/*.md",
});
}
}).transform((evidence) => ({
...evidence,
source: evidence.source.type !== "filesystem" || evidence.source.patterns !== undefined
? evidence.source
: {
...evidence.source,
patterns: evidence.schema_version === 2 ? ["curated/**/*.md"] : ["**/*.md"],
},
}));
function unique<T>(values: readonly T[], context: z.RefinementCtx, path: PropertyKey[]) {
if (new Set(values).size !== values.length) {
+24 -23
View File
@@ -70,6 +70,7 @@ const passwordHash =
"$argon2id$v=19$m=65536,t=3,p=1$AAECAwQFBgcICQoLDA0ODw$DRo8ZSPI8G5OCvnFFapbVEjP69aDjy1Sw9i2743cPC4";
const userId = "6ba7b810-9dad-4ed1-80b4-00c04fd430c8";
const roots: string[] = [];
const linuxTest = test.runIf(process.platform === "linux");
afterEach(() => {
fsHook.path = undefined;
@@ -276,13 +277,13 @@ function expectDenied(operation: () => unknown): void {
}
}
test("loads ready projection as one immutable auth and local-users snapshot", () => {
linuxTest("loads ready projection as one immutable auth and local-users snapshot", () => {
const root = projectionRoot();
const fixture = localProjectionFixture("synthetic-user", passwordHash);
const generation = writeReadyProjection(root, fixture);
const loaded = createProjectedAuthenticationConfigProvider(root).current();
expect(loaded.revision).toBe(`sha256:${generation}`);
expect(loaded.revision).toBe(generation);
expect(loaded.sourcePath).toBe(
join(root, "generations", generation, "auth.yaml"),
);
@@ -298,7 +299,7 @@ test("loads ready projection as one immutable auth and local-users snapshot", ()
).toBe(true);
});
test("loads a complete OIDC projection without a users snapshot", () => {
linuxTest("loads a complete OIDC projection without a users snapshot", () => {
const root = projectionRoot();
const generation = writeReadyOidcProjection(root);
const loaded = createProjectedAuthenticationConfigProvider(root).current();
@@ -309,7 +310,7 @@ test("loads a complete OIDC projection without a users snapshot", () => {
});
});
test("loadConfig selects an immutable projected local provider and its in-memory registry", async () => {
linuxTest("loadConfig selects an immutable projected local provider and its in-memory registry", async () => {
const root = projectionRoot();
writeReadyProjection(root, localProjectionFixture("projected-user", passwordHash));
@@ -323,7 +324,7 @@ test("loadConfig selects an immutable projected local provider and its in-memory
expect(await registry?.findByUsername("PROJECTED-USER")).toMatchObject({ username: "projected-user" });
});
test("loadConfig selects an immutable projected OIDC provider without direct-file fallback", () => {
linuxTest("loadConfig selects an immutable projected OIDC provider without direct-file fallback", () => {
const root = projectionRoot();
writeReadyOidcProjection(root);
@@ -338,7 +339,7 @@ test("loadConfig selects an immutable projected OIDC provider without direct-fil
});
});
test("rejects a trailing-slash runtime root", () => {
linuxTest("rejects a trailing-slash runtime root", () => {
const root = projectionRoot();
writeReadyProjection(
root,
@@ -349,7 +350,7 @@ test("rejects a trailing-slash runtime root", () => {
);
});
test.each([
linuxTest.each([
["missing", undefined],
[
"blocked",
@@ -381,7 +382,7 @@ test.each([
);
});
test.each([
linuxTest.each([
"root traversal",
"CURRENT symlink",
"CURRENT hardlink",
@@ -436,7 +437,7 @@ test.runIf(process.geteuid?.() === 0)(
},
);
test("rejects a foreign group with the correct owner", () => {
linuxTest("rejects a foreign group with the correct owner", () => {
const root = projectionRoot();
writeReadyProjection(
root,
@@ -449,7 +450,7 @@ test("rejects a foreign group with the correct owner", () => {
);
});
test("enumerates closed namespaces without path-based readdirSync", () => {
linuxTest("enumerates closed namespaces without path-based readdirSync", () => {
const root = projectionRoot();
const generation = writeReadyProjection(
root,
@@ -462,7 +463,7 @@ test("enumerates closed namespaces without path-based readdirSync", () => {
).toBe(generation);
});
test("rejects a symlinked runtime root", () => {
linuxTest("rejects a symlinked runtime root", () => {
const root = projectionRoot();
writeReadyProjection(
root,
@@ -476,7 +477,7 @@ test("rejects a symlinked runtime root", () => {
);
});
test.each([
linuxTest.each([
"root",
"generations",
"selected generation",
@@ -514,7 +515,7 @@ test.each([
);
});
test.each(["manifest", "generation", "size", "digest"])(
linuxTest.each(["manifest", "generation", "size", "digest"])(
"rejects changed %s integrity data without secret disclosure",
(kind) => {
const root = projectionRoot();
@@ -542,7 +543,7 @@ test.each(["manifest", "generation", "size", "digest"])(
},
);
test("switches atomically to a later complete generation", () => {
linuxTest("switches atomically to a later complete generation", () => {
const root = projectionRoot();
const first = writeReadyProjection(
root,
@@ -564,7 +565,7 @@ test("switches atomically to a later complete generation", () => {
expect(provider.current().runtimeProjection?.generation).toBe(second);
});
test("retries once when CURRENT is atomically replaced between lstat and open", () => {
linuxTest("retries once when CURRENT is atomically replaced between lstat and open", () => {
const root = projectionRoot();
const first = writeReadyProjection(
root,
@@ -591,7 +592,7 @@ test("retries once when CURRENT is atomically replaced between lstat and open",
).toBe(second);
});
test("retries once when CURRENT is replaced after the final identity read", () => {
linuxTest("retries once when CURRENT is replaced after the final identity read", () => {
const root = projectionRoot();
const first = writeReadyProjection(
root,
@@ -628,7 +629,7 @@ test("retries once when CURRENT is replaced after the final identity read", () =
expect(observations).toBe(3);
});
test("retries once when CURRENT is replaced between root descriptor and path observations", () => {
linuxTest("retries once when CURRENT is replaced between root descriptor and path observations", () => {
const root = projectionRoot();
const first = writeReadyProjection(
root,
@@ -678,7 +679,7 @@ test("retries once when CURRENT is replaced between root descriptor and path obs
expect(replacedCurrent).toBe(true);
});
test("rejects a second CURRENT replacement after the one permitted retry", () => {
linuxTest("rejects a second CURRENT replacement after the one permitted retry", () => {
const root = projectionRoot();
const first = writeReadyProjection(
root,
@@ -717,7 +718,7 @@ test("rejects a second CURRENT replacement after the one permitted retry", () =>
);
});
test("fails deterministically when generations is replaced during a load", () => {
linuxTest("fails deterministically when generations is replaced during a load", () => {
const root = projectionRoot();
const generation = writeReadyProjection(
root,
@@ -748,7 +749,7 @@ test("fails deterministically when generations is replaced during a load", () =>
);
});
test("fails deterministically when the selected generation directory is replaced during a load", () => {
linuxTest("fails deterministically when the selected generation directory is replaced during a load", () => {
const root = projectionRoot();
const generation = writeReadyProjection(
root,
@@ -780,7 +781,7 @@ test("fails deterministically when the selected generation directory is replaced
);
});
test.each(["corrupt", "symlink"])(
linuxTest.each(["corrupt", "symlink"])(
"rejects a %s retained predecessor generation",
(kind) => {
const root = projectionRoot();
@@ -819,7 +820,7 @@ test.each(["corrupt", "symlink"])(
},
);
test("has no direct-file fallback when CURRENT is absent", () => {
linuxTest("has no direct-file fallback when CURRENT is absent", () => {
const root = projectionRoot();
const generation = writeReadyProjection(
root,
@@ -839,7 +840,7 @@ test("has no direct-file fallback when CURRENT is absent", () => {
);
});
test("in-flight snapshot authenticates A after selection B and deletion A, while a new load sees B", async () => {
linuxTest("in-flight snapshot authenticates A after selection B and deletion A, while a new load sees B", async () => {
const root = projectionRoot();
const first = writeReadyProjection(
root,
+11
View File
@@ -97,6 +97,17 @@ test("loadConfig allows none and mock only outside production when auth.yaml is
.toThrow("production requires auth.yaml or AUTH_MODE=upstream");
});
test("workspace maintenance loads production configuration without an authentication surface", () => {
expect(loadConfig(
{ NODE_ENV: "production", THOTH_PUBLIC_EXPOSURE: "true" },
{ surface: "workspace-maintenance" },
)).toMatchObject({
authMode: "none",
authentication: undefined,
publicExposure: false,
});
});
test("local Compose profiles explicitly select the development auth environment", () => {
for (const profile of ["../../deploy/compose.local.yaml", "../../docker-compose.dev.yml"]) {
expect(readFileSync(new URL(profile, import.meta.url), "utf8")).toMatch(/NODE_ENV:\s*development/);
+15 -1
View File
@@ -1,4 +1,4 @@
import { expect, test } from "vitest";
import { expect, test, vi } from "vitest";
import { mkdirSync, mkdtempSync, rmSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
@@ -35,6 +35,20 @@ test("loads exact global enabledModels in configured order", () => {
} finally { rmSync(f.root, { recursive: true, force: true }); }
});
test("uses the configured Pi agent directory when no explicit directory is passed", () => {
const f = fixture({ enabledModels: ["zai/glm-5.2"] });
vi.stubEnv("PI_CODING_AGENT_DIR", f.agentDir);
try {
expect(loadPiEnabledModels({ harnessDir: f.harnessDir })).toMatchObject({
ids: ["zai/glm-5.2"],
warnings: [],
});
} finally {
vi.unstubAllEnvs();
rmSync(f.root, { recursive: true, force: true });
}
});
test("project enabledModels overrides global enabledModels", () => {
const f = fixture(
{ enabledModels: ["zai/glm-5.2", "zai/glm-5v-turbo"] },
+28
View File
@@ -0,0 +1,28 @@
import { expect, test } from "vitest";
import {
PI_MANAGED_CONFIG_ERROR_MESSAGE,
configuredPiProviderApiKey,
} from "../src/pi/managed-config.js";
test("provider credential declarations are selected from models.json by provider ID", () => {
const raw = JSON.stringify({
providers: {
hosted: { apiKey: "$HOSTED_API_KEY", models: [{ id: "one" }] },
local: { apiKey: "local", models: [{ id: "two" }] },
},
});
expect(configuredPiProviderApiKey(raw, "HOSTED")).toBe("$HOSTED_API_KEY");
expect(configuredPiProviderApiKey(raw, "local")).toBe("local");
expect(configuredPiProviderApiKey(raw, "missing")).toBeUndefined();
});
test.each([
JSON.stringify({ providers: [] }),
JSON.stringify({ providers: { local: "invalid" } }),
JSON.stringify({ providers: { local: { apiKey: 42 } } }),
JSON.stringify({ providers: { local: { apiKey: "!must-not-run" } } }),
])("invalid declarative provider credential configuration fails closed", (raw) => {
expect(() => configuredPiProviderApiKey(raw, "local"))
.toThrow(PI_MANAGED_CONFIG_ERROR_MESSAGE);
});
+4 -1
View File
@@ -198,7 +198,10 @@ test("smoke uses the configured timeout and reports a sanitized timeout", async
message: "Pi smoke check timed out",
checkedAt: "2026-08-05T10:00:00.000Z",
});
expect(calls).toEqual([{ command: "/usr/local/bin/pi", args: ["--version"], timeout: 750 }]);
expect(calls).toHaveLength(1);
expect(calls[0]).toMatchObject({ command: "/usr/local/bin/pi", args: ["--version"] });
expect(calls[0]!.timeout).toBeGreaterThan(0);
expect(calls[0]!.timeout).toBeLessThanOrEqual(750);
});
// Catches a smoke endpoint that validates only the Pi binary/model catalogue and never makes a
+17 -3
View File
@@ -22,7 +22,7 @@ const SAFE_AUTH = '{"deepseek":{"type":"api_key","key":"safe-token"}}\n';
const SAFE_MODELS = [
"{",
' "providers": {',
' "local-qwen": {"baseUrl":"http://model.invalid/v1","models":[{"id":"qwen"}]}',
' "local-qwen": {"baseUrl":"http://model.invalid/v1","apiKey":"local","models":[{"id":"qwen"}]}',
" }",
"}",
"",
@@ -670,9 +670,22 @@ test.each([["OpenAI", "openai"], ["gemini", "google"]])(
},
);
test.each(["ollama", "local-qwen"])(
"local provider %s spawns without a model key and scrubs ambient credentials",
test.each(["installation-local", "private-compatible"])(
"provider %s configured with a literal apiKey spawns without a managed key",
async (provider) => {
const root = mkdtempSync(path.join(tmpdir(), "tht-local-provider-"));
const agentDir = path.join(root, "agent");
mkdirSync(agentDir, { mode: 0o700 });
writeFileSync(path.join(agentDir, "models.json"), JSON.stringify({
providers: {
[provider]: {
baseUrl: "http://model.invalid/v1",
apiKey: "local",
models: [{ id: "model" }],
},
},
}), { mode: 0o600 });
vi.stubEnv("PI_CODING_AGENT_DIR", agentDir);
vi.stubEnv("PI_PROVIDER_API_KEY", "ambient-secret");
vi.stubEnv("THT_MODEL_API_KEY_FILE", "/ambient/secret-path");
vi.stubEnv("OPENAI_API_KEY", "unselected-provider-secret");
@@ -690,6 +703,7 @@ test.each(["ollama", "local-qwen"])(
} finally {
mgr.teardown(`local-session-${provider}`);
vi.unstubAllEnvs();
rmSync(root, { recursive: true, force: true });
}
},
);
+27 -4
View File
@@ -117,20 +117,41 @@ test("single-key providers scrub ambient compound companions before injecting th
expect(env).not.toHaveProperty("CLOUDFLARE_GATEWAY_ID");
});
test("local-qwen is an explicit local provider and needs no generic key", () => {
test("a provider with a literal apiKey in models.json needs no code-level provider exception", () => {
const env = buildPiChildEnv({
ambient: {
PI_PROVIDER_API_KEY: "must-not-leak",
OPENAI_API_KEY: "must-not-leak",
THT_MODEL_API_KEY_FILE: "/must/not/leak",
},
provider: "local-qwen",
provider: "installation-local",
configuredApiKey: "local",
});
expect(env).not.toHaveProperty("PI_PROVIDER_API_KEY");
expect(env).not.toHaveProperty("OPENAI_API_KEY");
expect(env).not.toHaveProperty("THT_MODEL_API_KEY_FILE");
});
test.each(["$PRIVATE_PROVIDER_API_KEY", "${PRIVATE_PROVIDER_API_KEY}"])(
"a custom provider credential target is derived from models.json: %s",
(configuredApiKey) => {
const env = buildPiChildEnv({
ambient: { PRIVATE_PROVIDER_API_KEY: "stale" },
provider: "private-provider",
configuredApiKey,
credentialValue: "selected-secret",
});
expect(env.PRIVATE_PROVIDER_API_KEY).toBe("selected-secret");
},
);
test("a custom provider cannot redirect a managed credential into a process-control variable", () => {
expect(() => buildPiChildEnv({
ambient: {}, provider: "private-provider", configuredApiKey: "$PATH",
credentialValue: "selected-secret",
})).toThrow("model provider credential is unavailable");
});
test("credential status reports only present or missing without treating local providers as credentialed", () => {
expect(piProviderCredentialStatus({
provider: "deepseek",
@@ -139,7 +160,8 @@ test("credential status reports only present or missing without treating local p
})).toBe("present");
expect(piProviderCredentialStatus({ provider: "deepseek" })).toBe("missing");
expect(piProviderCredentialStatus({
provider: "local-qwen",
provider: "installation-local",
configuredApiKey: "local",
resolveCredentialValue: () => "must-not-be-returned",
})).toBe("missing");
});
@@ -159,7 +181,8 @@ test("credential status never resolves the generic secret for auth-store or loca
resolveCredentialValue: unreadableSecret,
})).toBe("present");
expect(piProviderCredentialStatus({
provider: "local-qwen",
provider: "installation-local",
configuredApiKey: "local",
resolveCredentialValue: unreadableSecret,
})).toBe("missing");
expect(secretReads).toBe(0);
+84
View File
@@ -28,6 +28,90 @@ const compatible = (size = 1024, distance = "Cosine", schema = payloadSchema) =>
payload_schema: schema,
});
test("Evidence maintenance adds an absent BM25 vector without changing the dense contract", async () => {
const info = compatible();
const requests: Array<{ url: string; init?: any }> = [];
const request = async (url: string, init?: any) => {
requests.push({ url, init });
if (init?.method === "PUT" && /\/vectors\/bm25$/.test(url)) {
expect(JSON.parse(String(init.body))).toEqual({ sparse: { modifier: "idf" } });
(info.config.params as any).sparse_vectors = { bm25: { modifier: "idf" } };
return { status: 200, ok: true, json: async () => ({}) } as any;
}
return { status: 200, ok: true, json: async () => ({ result: info }) } as any;
};
const result = await reconcileCollection({
baseUrl: "http://qdrant:6333", collection: "c", dimensions: 1024, distance: "cosine",
mode: "evidence_maintenance", request,
});
expect(result).toEqual({ ok: true, state: "upgraded" });
expect(info.config.params.vectors).toEqual({ size: 1024, distance: "Cosine" });
expect(requests.filter(({ init }) => init?.method === "PUT")).toHaveLength(1);
expect(requests[1]?.url).toBe("http://qdrant:6333/collections/c/vectors/bm25");
});
test("Evidence maintenance creates a missing collection with both required vector contracts", async () => {
let info: any;
let createdBody: any;
const request = async (url: string, init?: any) => {
if (init?.method === "PUT") {
createdBody = JSON.parse(String(init.body));
info = { config: { params: createdBody }, payload_schema: payloadSchema };
return { status: 200, ok: true, json: async () => ({}) } as any;
}
if (info === undefined) return { status: 404, ok: false, json: async () => ({}) } as any;
return { status: 200, ok: true, json: async () => ({ result: info }) } as any;
};
const result = await reconcileCollection({
baseUrl: "http://qdrant:6333", collection: "c", dimensions: 1024, distance: "cosine",
mode: "evidence_maintenance", request,
});
expect(result).toEqual({ ok: true, state: "ready" });
expect(createdBody).toEqual({
vectors: { size: 1024, distance: "Cosine" },
sparse_vectors: { bm25: { modifier: "idf" } },
});
});
test("Evidence maintenance refuses an incompatible BM25 definition without mutating", async () => {
const info = compatible();
(info.config.params as any).sparse_vectors = { bm25: { modifier: "none" } };
const requests: Array<{ url: string; init?: any }> = [];
const request = async (url: string, init?: any) => {
requests.push({ url, init });
return { status: 200, ok: true, json: async () => ({ result: info }) } as any;
};
const result = await reconcileCollection({
baseUrl: "http://qdrant:6333", collection: "c", dimensions: 1024, distance: "cosine",
mode: "evidence_maintenance", request,
});
expect(result).toEqual({ ok: false, code: "semantic_index_incompatible" });
expect(requests.filter(({ init }) => init?.method === "PUT")).toEqual([]);
});
test("ordinary session reconciliation does not add BM25", async () => {
const info = compatible();
const requests: Array<{ url: string; init?: any }> = [];
const request = async (url: string, init?: any) => {
requests.push({ url, init });
return { status: 200, ok: true, json: async () => ({ result: info }) } as any;
};
const result = await reconcileCollection({
baseUrl: "http://qdrant:6333", collection: "c", dimensions: 1024, distance: "cosine",
mode: "self_heal", request,
});
expect(result).toEqual({ ok: true, state: "ready" });
expect(requests.filter(({ url }) => /\/vectors\/bm25$/.test(url))).toEqual([]);
});
test("self-heal creates a missing compatible collection", async () => {
const r = await reconcileCollection({
baseUrl: "http://qdrant:6333", collection: "c", dimensions: 1024, distance: "cosine",
+23
View File
@@ -42,6 +42,29 @@ test("real Pi thinking_delta becomes a dedicated activity_delta to the FE", () =
expect(seen).toEqual([{ type: "activity_delta", text: "Valuto le ambiguità" }]);
});
test("a reserved phase notification becomes a structured phase_started event", () => {
const { rpc, fire } = fakeRpc();
const bridge = new SessionBridge(rpc);
const seen: any[] = [];
bridge.onClientEvent((event) => seen.push(event));
fire({
type: "extension_ui_request",
method: "notify",
notifyType: "info",
message: "__tht_phase_started__:F2",
});
fire({
type: "extension_ui_request",
method: "notify",
notifyType: "info",
message: "__tht_phase_started__:F9_DO_NOT_FORWARD",
});
expect(seen).toEqual([{ type: "system_event", event: "phase_started", phase: "F2" }]);
expect(JSON.stringify(seen)).not.toContain("DO_NOT_FORWARD");
});
test("assistant message_end exposes sanitized token usage with the configured context window", () => {
const { rpc, fire } = fakeRpc();
const bridge = new SessionBridge(rpc);
+4 -4
View File
@@ -46,9 +46,9 @@ function realChildBridge(
bridge: factory({
thtExecutable: pathStyle === "windows" ? "C:\\tht.exe" : launcher,
spawnChild: (_executable, args, options) => spawn(launcher, [...args], options),
// Leave enough startup headroom for a real child under a busy CI host while retaining a
// sub-1.5-second bound from request start through final settlement.
deadlinesForTest: { timeoutMs: 750, terminationGraceMs: 50, finalSettlementMs: 500 },
// Keep this stricter than the five-second production timeout without assuming that a
// real Node child can always start within 750 ms on a busy shared runner.
deadlinesForTest: { timeoutMs: 2_000, terminationGraceMs: 50, finalSettlementMs: 500 },
...(mode === "stdin" ? {
beforeInputForTest: async () => {
await waitForMarker(marker, "stdin-closed");
@@ -585,6 +585,6 @@ describe("Windows auth-storage bridge", () => {
await expect(outcome).resolves.toMatchObject({ message: "auth_session_store_invalid" });
await waitForMarker(marker, "terminated");
expect(Date.now() - startedAt).toBeLessThan(1_500);
expect(Date.now() - startedAt).toBeLessThan(3_000);
}, 5_000);
});
@@ -161,6 +161,7 @@ function fixture(workspace = baseWorkspace) {
runChild,
listSessions: async () => [],
semanticPreflight: async () => ({ ok: true }),
evidencePreflight: async () => ({ ok: true }),
});
return { dataRoot, runChild, requests, service };
}
@@ -283,6 +284,7 @@ test("index schema fails closed when semantic preflight refuses the collection",
runChild,
listSessions: async () => [],
semanticPreflight: async () => ({ ok: false, code: "semantic_index_incompatible" }),
evidencePreflight: async () => ({ ok: true }),
});
const result = await service.indexSchema({ workspaceId: "psd-clinical" });
@@ -309,6 +311,7 @@ test("filesystem Evidence proceeds after materialization and private HTTP hosts
runChild: vi.fn(),
listSessions: async () => [],
semanticPreflight: async () => ({ ok: true }),
evidencePreflight: async () => ({ ok: true }),
httpPrivateHostAllowlist: ["metadata.internal"],
});
@@ -513,6 +516,7 @@ test("vector rebuild recreates the full collection contract including keyword in
runChild: vi.fn(),
listSessions: async () => [],
semanticPreflight: async () => ({ ok: true }),
evidencePreflight: async () => ({ ok: true }),
});
// replace global fetch used by vectorRebuild/reconcileCollection
const original = globalThis.fetch;
@@ -293,6 +293,9 @@ test("Windows clone contract copies the shared complete schema v3 descriptor int
expect(windows).toContain('thoth-workspaces.yaml');
expect(windows).toContain('$workspaceDestination = Join-Path $workspaceDirectory "workspace.yaml"');
expect(windows).toContain('Join-Path $workspaceEvidence "guide.md"');
expect(windows).toContain('THT_AUTH_CONFIG_ROOT=$authConfigRoot');
expect(windows).toContain('"core,embedding,embedding-model-init,frontend,qdrant"');
expect(windows).toContain('"core,embedding,frontend,qdrant"');
expect(windows).not.toContain('schema_version: 3');
expect(descriptor).toMatchObject({
workspace: {
@@ -139,8 +139,14 @@ function evidenceSecretFile(name: string, contents: string): string {
return path;
}
function evidenceWorkspace(source: Record<string, unknown>, policy?: Record<string, unknown>) {
return parseWorkspaceYaml(`${canonicalEvidenceWorkspace}\nevidence:\n source: ${JSON.stringify(source)}${
function evidenceWorkspace(
source: Record<string, unknown>,
policy?: Record<string, unknown>,
evidenceSchemaVersion?: number,
) {
return parseWorkspaceYaml(`${canonicalEvidenceWorkspace}\nevidence:${
evidenceSchemaVersion === undefined ? "" : `\n schema_version: ${evidenceSchemaVersion}`
}\n source: ${JSON.stringify(source)}${
policy === undefined ? "" : `\n policy: ${JSON.stringify(policy)}`
}\n`);
}
@@ -180,9 +186,10 @@ function evidenceRender(
source: Record<string, unknown>,
evidenceBinding: RuntimeBindings["evidence"] = { missing: [], values: {} },
policy?: Record<string, unknown>,
evidenceSchemaVersion?: number,
) {
return renderRuntimeConfig(
evidenceWorkspace(source, policy),
evidenceWorkspace(source, policy, evidenceSchemaVersion),
{ ...directBindings, evidence: evidenceBinding },
paths,
evidenceContext,
@@ -195,15 +202,16 @@ test("renders filesystem Evidence below the immutable revision content root with
const yaml = evidenceRender({
type: "filesystem",
uri: "psd-clinical/evidence",
});
}, undefined, undefined, 2);
const rendered = parse(yaml);
expect(rendered.runtime_identity.workspace_revision).toBe(evidenceRevision);
expect(rendered.evidence).toEqual({
schema_version: 2,
sources: [{
type: "filesystem",
root: `/srv/registry/snapshots/${evidenceRevision}/psd-clinical/evidence`,
patterns: ["**/*.md"],
patterns: ["curated/**/*.md"],
max_bytes: 10_485_760,
}],
});
+29 -5
View File
@@ -323,16 +323,40 @@ test("parallel contenders recover a stale lock file without overlapping critical
const second = new WorkspaceRepositoryLock(locks);
let active = 0;
let maximum = 0;
const critical = async () => {
let firstEntered!: () => void;
const entered = new Promise<void>((resolve) => { firstEntered = resolve; });
let releaseFirst!: () => void;
const held = new Promise<void>((resolve) => { releaseFirst = resolve; });
const firstRun = first.run(async () => {
active += 1;
maximum = Math.max(maximum, active);
firstEntered();
await held;
active -= 1;
});
await entered;
const secondRun = second.run(async () => {
active += 1;
maximum = Math.max(maximum, active);
await new Promise((resolve) => setTimeout(resolve, 25));
active -= 1;
};
});
const results = await Promise.allSettled([first.run(critical), second.run(critical)]);
let timeout!: ReturnType<typeof setTimeout>;
const contender = await Promise.race([
secondRun.then(
() => ({ status: "fulfilled" as const }),
() => ({ status: "rejected" as const }),
),
new Promise<{ status: "timed-out" }>((resolve) => {
timeout = setTimeout(() => resolve({ status: "timed-out" }), 2_000);
}),
]);
clearTimeout(timeout);
releaseFirst();
await firstRun;
await secondRun.catch(() => undefined);
expect(results.filter((result) => result.status === "fulfilled")).toHaveLength(1);
expect(results.filter((result) => result.status === "rejected")).toHaveLength(1);
expect(contender.status).toBe("rejected");
expect(maximum).toBe(1);
});
+72
View File
@@ -390,6 +390,78 @@ test("applies filesystem and policy defaults to the canonical descriptor", () =>
});
});
test("defaults schema-versioned filesystem Evidence to curated documents only", () => {
const parsed = validateWorkspaceDescriptor({
...withEvidence({
type: "filesystem",
uri: "psd-clinical/evidence",
}),
evidence: {
schema_version: 2,
source: { type: "filesystem", uri: "psd-clinical/evidence" },
},
});
expect(parsed.evidence).toMatchObject({
schema_version: 2,
source: { patterns: ["curated/**/*.md"] },
});
});
test("rejects a schema-versioned Evidence layout that mixes source and curated runtime patterns", () => {
expectSafeEvidenceError({
...withEvidence({
type: "filesystem",
uri: "psd-clinical/evidence",
}),
evidence: {
schema_version: 2,
source: {
type: "filesystem",
uri: "psd-clinical/evidence",
patterns: ["source/**/*.md", "curated/**/*.md"],
},
},
}, /source.*curated|curated.*source/i);
});
test("rejects a schema-versioned Evidence layout that acquires source documents at runtime", () => {
expectSafeEvidenceError({
...withEvidence({
type: "filesystem",
uri: "psd-clinical/evidence",
}),
evidence: {
schema_version: 2,
source: {
type: "filesystem",
uri: "psd-clinical/evidence",
patterns: ["source/**/*.md"],
},
},
}, /curated/i);
});
test.each(["curated/**/*.yaml", "curated/**"])(
"rejects a schema-versioned Evidence layout that uses the non-canonical curated pattern %s",
(pattern) => {
expectSafeEvidenceError({
...withEvidence({
type: "filesystem",
uri: "psd-clinical/evidence",
}),
evidence: {
schema_version: 2,
source: {
type: "filesystem",
uri: "psd-clinical/evidence",
patterns: [pattern],
},
},
}, /curated\/\*\*\/\*\.md/i);
},
);
test("keeps evidence optional on schema v3", () => {
expect(validateWorkspaceDescriptor(validWorkspaceObject())).not.toHaveProperty("evidence");
});
@@ -0,0 +1,34 @@
import { readdirSync, readFileSync } from "node:fs";
import { basename, join } from "node:path";
import { fileURLToPath } from "node:url";
import ts from "typescript";
import { expect, test } from "vitest";
const evidenceRoot = fileURLToPath(new URL("../../../src/workspaces/evidence/", import.meta.url));
const coreInfrastructure = new Set([
"preprocessing-service",
"qdrant-collection",
"registry",
]);
test("Evidence modules do not import core-owned registry or Qdrant lifecycle", () => {
const violations: Array<{ file: string; dependency: string }> = [];
for (const file of readdirSync(evidenceRoot).filter((name) => name.endsWith(".ts"))) {
const source = ts.createSourceFile(
file,
readFileSync(join(evidenceRoot, file), "utf8"),
ts.ScriptTarget.Latest,
true,
ts.ScriptKind.TS,
);
for (const statement of source.statements) {
if (!ts.isImportDeclaration(statement) || !ts.isStringLiteral(statement.moduleSpecifier)) {
continue;
}
const dependency = basename(statement.moduleSpecifier.text).replace(/\.js$/, "");
if (coreInfrastructure.has(dependency)) violations.push({ file, dependency });
}
}
expect(violations).toEqual([]);
});
@@ -4,9 +4,9 @@ import { tmpdir } from "node:os";
import { join } from "node:path";
import { promisify } from "node:util";
import { afterEach, expect, test } from "vitest";
import { GitWorkspaceRepository } from "../src/workspaces/git-repository.js";
import { materializeEvidenceTree } from "../src/workspaces/evidence-materialization.js";
import type { WorkspaceRegistryConfig } from "../src/workspaces/types.js";
import { GitWorkspaceRepository } from "../../../src/workspaces/git-repository.js";
import { materializeEvidenceTree } from "../../../src/workspaces/evidence/materialization.js";
import type { WorkspaceRegistryConfig } from "../../../src/workspaces/types.js";
const runFile = promisify(execFile);
const temporaryRoots: string[] = [];
@@ -0,0 +1,160 @@
import { expect, test, vi } from "vitest";
import {
continueEvidencePreprocessing,
preprocessEvidence,
type EvidenceJobState,
type EvidencePreprocessingDependencies,
} from "../../../src/workspaces/evidence/preprocessing.js";
import type { WorkspaceDescriptor } from "../../../src/workspaces/schema.js";
type EvidenceConfig = NonNullable<WorkspaceDescriptor["evidence"]>;
const filesystemEvidence = {
source: { type: "filesystem", uri: "research/evidence" },
} as EvidenceConfig;
const privateHttpEvidence = {
source: {
type: "http",
uris: ["http://127.0.0.1/private.md"],
authentication: "none",
connect_timeout_ms: 1000,
read_timeout_ms: 2000,
max_bytes: 100,
max_redirects: 0,
allow_private_hosts: true,
max_cache_bytes: 100,
},
} as EvidenceConfig;
function job(overrides: Partial<EvidenceJobState> = {}): EvidenceJobState {
return {
runId: "a".repeat(32),
childRuns: {},
completedStages: [],
...overrides,
};
}
function dependencies(payload: Record<string, unknown> = {}): EvidencePreprocessingDependencies & {
runStage: ReturnType<typeof vi.fn>;
persistJob: ReturnType<typeof vi.fn>;
evidencePreflight: ReturnType<typeof vi.fn>;
} {
return {
runStage: vi.fn(async () => payload),
persistJob: vi.fn(),
evidencePreflight: vi.fn(async () => ({ ok: true as const })),
requireRunId(value) {
if (typeof value !== "string" || !/^[0-9a-f]{32}$/.test(value)) {
throw new Error("child run id is invalid");
}
return value;
},
numberRecord(value) {
if (!value || typeof value !== "object" || Array.isArray(value)) return undefined;
return Object.fromEntries(
Object.entries(value as Record<string, unknown>).map(([key, nested]) => [key, Number(nested)]),
);
},
};
}
test("Evidence maintenance preflights the additive BM25 contract before starting its stage", async () => {
const state = job({ childRuns: { evidence: "b".repeat(32) } });
const deps = dependencies({ run_id: "c".repeat(32), counts: { added: 2 } });
const result = await preprocessEvidence(
{ evidence: filesystemEvidence, job: state, dryRun: false },
deps,
);
expect(deps.evidencePreflight).toHaveBeenCalledOnce();
expect(deps.runStage).toHaveBeenCalledWith([
"preprocess", "evidence", "--resume", "b".repeat(32), "--json", "-c", "/dev/fd/3",
]);
expect(deps.persistJob).toHaveBeenCalledOnce();
expect(state).toMatchObject({
childRuns: { evidence: "c".repeat(32) },
completedStages: ["evidence"],
});
expect(result).toEqual({
status: "succeeded",
code: "ok",
runId: "a".repeat(32),
childRuns: { evidence: "c".repeat(32) },
completedStages: ["evidence"],
counts: { added: 2 },
});
});
test("owns Evidence egress refusal before shared semantic infrastructure", async () => {
const deps = dependencies();
const result = await preprocessEvidence(
{
evidence: privateHttpEvidence,
job: job(),
httpPrivateHostAllowlist: ["metadata.internal"],
},
deps,
);
expect(result).toEqual({ status: "failed", code: "egress_policy_refused" });
expect(deps.runStage).not.toHaveBeenCalled();
expect(deps.persistJob).not.toHaveBeenCalled();
});
test("projects aggregate no-Evidence and completed-stage outcomes without rerunning", async () => {
const deps = dependencies();
const noEvidence = await continueEvidencePreprocessing(
{
evidence: undefined,
job: job({ completedStages: ["dwh", "schema_index"] }),
priorCounts: { added: 2 },
},
deps,
);
const completed = await continueEvidencePreprocessing(
{
evidence: filesystemEvidence,
job: job({ completedStages: ["dwh", "schema_index", "evidence"] }),
},
deps,
);
expect(noEvidence).toMatchObject({
status: "succeeded",
code: "ok",
warnings: ["workspace has no Evidence source"],
counts: { added: 2 },
});
expect(completed).toMatchObject({
status: "unchanged",
code: "ok",
completedStages: ["dwh", "schema_index", "evidence"],
});
expect(deps.runStage).not.toHaveBeenCalled();
expect(deps.persistJob).not.toHaveBeenCalled();
});
test("preserves the narrow standalone projection for an already completed Evidence stage", async () => {
const deps = dependencies();
const result = await preprocessEvidence(
{
evidence: filesystemEvidence,
job: job({ childRuns: { evidence: "b".repeat(32) }, completedStages: ["evidence"] }),
},
deps,
);
expect(result).toEqual({
status: "unchanged",
code: "ok",
runId: "a".repeat(32),
completedStages: ["evidence"],
});
expect(deps.runStage).not.toHaveBeenCalled();
expect(deps.persistJob).not.toHaveBeenCalled();
});
@@ -1,14 +0,0 @@
# Datamart Builder deployment gotchas
- Il percorso pubblico attraversa due reverse proxy: nginx host → nginx Omics Portal → core ThothII.
- Per SSE, ogni livello deve disabilitare `proxy_buffering`, `proxy_request_buffering` e cache, usare HTTP/1.1, timeout lunghi e propagare `X-Accel-Buffering: no`.
- Il core aggiunge `X-Accel-Buffering: no` alla risposta EventSource; Omics Portal lo riaggiunge esplicitamente per i proxy a monte.
- Una sonda utile deve attraversare il portale autenticato e misurare l'arrivo degli header `200 text/event-stream`, non solo interrogare il core nel network Docker.
- La configurazione nginx host attiva vive in `/etc/nginx/sites-available/policlinicosandonato`; validare con `nginx -t` prima del reload.
- Django può tenere in memoria il manifest Vite per worker: dopo un rebuild del frontend riavviare anche i worker Omics Portal, altrimenti richieste diverse possono produrre hash asset vecchi e nuovi.
- Nel profilo server legacy, `vector_db` è una connessione pgvector RW condivisa; il factory può riusarla come writer solo con `profile=server`. La workstation senza writer REST deve restare read-only.
- Non convertire in-place `local.yaml` da chiavi legacy a risorse moderne durante un incident fix: cambia il binding del workspace e può invalidare gli snapshot DWH attivi.
- `session_storage` è configurazione di persistenza delle sessioni, non degli artefatti DWH: escluderla dall'impronta in `config_dwh_binding`, altrimenti l'attivazione di profili utente rende incompatibile un indice DWH già valido.
- Con profili per utente, un profilo privato vuoto deve essere inizializzato una sola volta dai settings legacy completi; trattare `{}` come settings effettivi fa creare sessioni senza provider/modello e lo spawn Pi fallisce prima di partire.
- Regola operativa concordata: dopo ogni modifica o enhancement, ricostruire e ricreare tutti i container ThothII coinvolti prima dell'handoff, quindi verificare che siano healthy affinché l'utente possa testare live. Il push Git non aggiorna le immagini Docker automaticamente.

Some files were not shown because too many files have changed in this diff Show More