206 Commits
Author SHA1 Message Date
User 2f53512e4d docs: record server release and hand off remaining acceptance checks
Publish documentation / publish (push) Successful in 32s
2026-09-27 00:41:37 +02:00
Codex 497ab84031 docs: pin server handoff to released main revision
Publish documentation / publish (push) Successful in 33s
2026-09-26 16:42:36 +02:00
Codex 0d2e573e0d fix(ui): reset session view on stop and exit 2026-09-26 16:40:37 +02:00
Codex bd416f7327 Fix new-question landing and question-language HITL
Publish documentation / publish (push) Successful in 34s
Reset the activity panel when starting a new question so the landing navigation is restored. Detect and persist the original question language, pass it through runtime and widget descriptors, and scope HITL controls to that language.

Validated with gate, session, backend and frontend tests, TypeScript checks, Ruff and strict docs build. Rebuilt and restarted local core/frontend; both healthy and serving HTTP successfully.
2026-09-21 19:47:22 +02:00
Codex 23e52c80de Record verified Qwen documentation publication
Publish documentation / publish (push) Successful in 23s
2026-09-21 16:27:10 +02:00
Codex 84084bba37 Fix Qwen session tool calls and expose thinking compatibility
Publish documentation / publish (push) Successful in 30s
2026-09-21 16:23:51 +02:00
pinoricci1956 efd7d788d9 correzione scroller verticale pagina di configurazione catalogo 2026-09-16 11:59:51 +02:00
Codex b1c510a097 fix(docs): preserve theme assets in deny-by-default publication
Publish documentation / publish (push) Successful in 36s
2026-09-16 09:36:30 +02:00
Codex 5f3680a0fb docs: record verified live manual publication
Publish documentation / publish (push) Successful in 28s
2026-09-15 14:39:46 +02:00
Codex 4ff91e8d6e docs: consolidate historical records and verify public manual publication
Publish documentation / publish (push) Successful in 29s
2026-09-15 14:37:29 +02:00
Codex 5f3a7f5975 docs: record stale public site publication blocker
Publish documentation / publish (push) Successful in 23s
2026-09-15 10:28:49 +02:00
Codex 043ffdfad6 docs: separate public manual from internal project documentation
Publish documentation / publish (push) Successful in 27s
2026-09-15 10:26:35 +02:00
Codex 6a4634dcf1 Merge manual standalone installation documentation
Publish documentation / publish (push) Successful in 35s
2026-09-15 10:06:33 +02:00
Codex 84804be9f8 docs: publish bilingual manual standalone installation guides 2026-09-15 10:06:28 +02:00
User c3caba94dd fix(ui): expand session dialogs and repeat confirmation actions
Publish documentation / publish (push) Successful in 24s
2026-09-14 18:13:53 +02:00
User b1723c34c4 docs: record server rollout of session and memory fixes
Publish documentation / publish (push) Successful in 32s
2026-09-14 17:25:23 +02:00
User d6cdffea62 fix: keep embedded session controls visible and handle empty memory
Publish documentation / publish (push) Successful in 34s
Cap the embedded shell at its portal container height so steering and stop controls remain accessible. Skip vector retrieval for an empty authoritative Memory archive and compute SQL-rule embeddings lazily.

Validated with 54 Memory tests, 90 frontend tests, five browser scenarios, frontend and Docker builds, and a read-only comparison against the real empty Memory archive.
2026-09-14 17:15:00 +02:00
Codex 49333a2d35 Merge full and embedded shell, administration UI and server handoff
Publish documentation / publish (push) Successful in 1m21s
2026-09-14 15:08:47 +02:00
Codex b006b94479 docs: prepare server Codex deployment handoff for ThothII and Omics 2026-09-14 15:08:46 +02:00
Codex bdcd8fcd28 fix(ui): open one session accordion panel at a time 2026-09-14 01:09:45 +02:00
Codex cf90c1bd51 fix(ui): collapse session lists and scope selection controls to panels 2026-09-13 18:12:27 +02:00
Codex 571a4bcaa2 fix(ui): show workspace readiness dot and bounded session accordions 2026-09-13 17:39:33 +02:00
Codex 9051463654 docs: document full and embedded rendering with server authentication 2026-09-13 17:28:10 +02:00
Codex 26c5605ff7 fix(ui): unify Memory and Evidence reading typography 2026-09-13 17:07:37 +02:00
Codex 3535fda958 fix(ui): improve knowledge reading and add isolated formatting examples 2026-09-13 16:52:53 +02:00
Codex 2953f6b608 fix(ui): unify Session navigation and restore uniform tab borders 2026-09-13 16:22:20 +02:00
Codex 45db3a239b fix(ui): simplify login and suppress pointer focus ring on locale select 2026-09-13 15:57:57 +02:00
Codex 7d826e46c0 fix(ui): match Omics header and compact workspace layout 2026-09-13 15:36:08 +02:00
Codex 648434a32e docs: record approved Omics GitHub to PSD relay handoff 2026-09-13 15:22:30 +02:00
Codex 023b822f83 Merge visual review into full shell and preserve bilingual layout 2026-09-13 14:55:58 +02:00
Codex d8a29bfbdd Add full shell, replaceable Omics adapter and bilingual interaction
Implement approved specification #32 and tickets #33-#37. Keep host authentication server-verified and pin session interaction language. Compile scoped base selectors for browser compatibility and retain full gutters during CSS pruning.
2026-09-13 14:26:39 +02:00
Codex a59624a68f style: align global context panel and simplify selection copy 2026-09-13 10:43:42 +02:00
Codex 5af4408194 style: unify database block gutters and content alignment 2026-09-13 10:32:55 +02:00
Codex e088abd60a style: align catalog status with summary grid 2026-09-13 01:44:48 +02:00
Codex 803e9e9201 fix: remeasure composer after hidden Core becomes visible 2026-09-13 01:39:19 +02:00
Codex 8c81996896 style: place catalog status indicators after their labels 2026-09-13 01:23:46 +02:00
Codex eed398e569 style: restore prominent ThothII application wordmarks 2026-09-13 01:18:06 +02:00
Codex c8d276ddc6 style: unify workbench typography and prepare isolated visual review 2026-09-12 22:51:58 +02:00
Codex 2d1b714ebe fix: resolve admin issue review findings and record verification 2026-09-12 18:24:23 +02:00
Codex c7e5f295e6 fix: address administration layout and navigation issues #28 #29 #30 #31 2026-09-12 18:16:52 +02:00
Codex f52bf22e05 feat: establish unified administration and model context baseline 2026-09-12 18:03:15 +02:00
Codex 840344706f prototype: restore original Core within context shelf alternatives 2026-09-12 14:00:52 +02:00
Codex 41b9fed4d5 prototype: refine context shelf with remembered defaults and session tabs 2026-09-12 12:29:02 +02:00
Codex 36bf659ea9 prototype: explore global context with one operation at a time 2026-09-12 11:53:52 +02:00
Codex 4a67d60233 prototype: revisit five administration workflows after design review 2026-09-10 19:40:21 +02:00
Codex debb63d87b prototype unified administration page layouts 2026-09-10 16:48:25 +02:00
Codex 3943022a97 Merge remote-tracking branch 'origin/main'
# Conflicts:
#	mkdocs.yml
2026-09-10 12:57:43 +02:00
Codex f5ec2d9313 docs: track security evidence and research notes 2026-09-10 12:53:13 +02:00
Codex 82e2c91f42 feat: implement memory and evidence administration with guided repairs
Publish documentation / publish (push) Successful in 1m27s
Add PostgreSQL-backed memory, editable evidence with source review and activation, and human-approved archive repairs across the harness, API, and UI. Include migrations, deployment support, regression coverage, and validation documentation.

Refresh permissions from validated session roles so existing administrator logins can access newly deployed archive management features.
2026-09-10 10:31:34 +02:00
User 8fe526dd6e fix(frontend): show session catalog in pi management 2026-09-08 14:58:18 +02:00
User e68e80a33d fix(frontend): prevent catalog header overlap 2026-09-08 14:35:36 +02:00
Codex 818563c408 fix(core): keep workspace runtime available 2026-09-08 13:37:10 +02:00
Codex 50c546e42d fix(frontend): accept catalog-owned workspace descriptors 2026-09-07 15:01:31 +02:00
Codex 651a5c7902 fix(frontend): isolate administration rail from portal CSS 2026-09-07 11:05:38 +02:00
User 28db30bd78 fix(ops): make server diagnostics release-safe 2026-09-07 01:15:28 +02:00
Codex cffa60772e feat: complete catalog-driven preprocessing
Publish documentation / publish (push) Successful in 2m12s
2026-09-06 17:49:35 +02:00
marcopan 8707ae1d46 fix(frontend): widen management work-area panels 2026-09-05 10:43:49 +02:00
Codex ad744f0212 docs: add guarded server upgrade runbook
Publish documentation / publish (push) Successful in 1m20s
2026-09-04 17:37:26 +02:00
Codex eba6148511 test: stabilize pre-deployment gates
Remove the redundant timing-dependent native Argon2 concurrency test while retaining native vector coverage and deterministic limiter coverage. Refresh stale deployment and browser contracts, make release scripts portable across Bash/macOS, and update production dependency locks for resolved security advisories.
2026-09-04 16:15:35 +02:00
Codex 7b1d69a65b feat: complete catalog sensitivity enhancements 2026-09-04 15:11:18 +02:00
Codex b891246664 docs: add PSD CPU NER benchmark 2026-09-03 10:59:09 +02:00
Codex 8e778b9edb feat: sample sensitive columns progressively 2026-09-03 10:25:05 +02:00
Codex f114d0065a feat: classify sensitive columns locally 2026-09-03 02:11:13 +02:00
Codex 7b87e95427 fix: refresh catalog after hidden sync completion 2026-09-02 23:23:00 +02:00
Codex a50475d687 chore: ignore generated deployment projections 2026-09-02 20:35:08 +02:00
Codex a6a5bf2036 fix: harden model catalog projections 2026-09-02 19:25:01 +02:00
Codex ce4c31a6fb docs: align restore guidance with workspace v4 2026-09-02 18:47:57 +02:00
Codex 538dc8ef56 test: name workspace schema v4 gate 2026-09-02 18:46:49 +02:00
Codex 7b7927bfe5 feat: unify installation model catalog 2026-09-02 18:45:33 +02:00
Codex ae053961a3 feat: refine metadata catalog workflows 2026-09-02 15:58:23 +02:00
Codex 4531746038 feat: refine metadata catalog workflows 2026-09-02 11:38:47 +02:00
Codex 076c9742c5 feat: consolidate database management work
Add catalog-owned logical relationships and runtime snapshots, extend the database-management UI and validation coverage, and document the updated operational workflow.

Keep active sensitive-generation status in a tooltip and indicator, and update the layout E2E to follow the history action in its new database-scoped location.
2026-09-01 14:46:55 +02:00
Codex f586152636 fix: close catalog review gaps
Publish documentation / publish (push) Successful in 38s
2026-08-31 16:28:42 +02:00
Codex 64fbe642ef test: retarget schema v3 documentation gate 2026-08-31 16:07:01 +02:00
Codex 74d5a3c0c8 test: refresh reviewed deployment fixture digest 2026-08-31 16:02:50 +02:00
Codex 7a0697e562 merge: preserve presentation source history 2026-08-31 16:00:10 +02:00
Codex ea903b5bd3 merge: preserve superseded root documents in history 2026-08-31 16:00:08 +02:00
Codex 218e50f124 tools: preserve HTML deck exporter 2026-08-31 15:59:51 +02:00
Codex e82fe8d322 test: align release gates with installation config 2026-08-31 15:59:39 +02:00
Codex b4b97436e1 docs: reconcile catalog run history and fleet state 2026-08-31 15:59:26 +02:00
Codex ded66fde9c chore: preserve database management prototype 2026-08-31 15:59:08 +02:00
Codex 9c697dc062 feat: complete catalog fleet management workflow 2026-08-31 15:58:43 +02:00
Codex 919d408c81 chore: preserve presentation sources and exporter 2026-08-31 15:35:35 +02:00
Codex fa7380b3a3 chore: preserve root worktree documents 2026-08-31 15:34:55 +02:00
Codex 866aee4249 docs: revalidate metadata catalog research 2026-08-31 14:54:52 +02:00
marcopanandCodex 16b7477861 docs: trace metadata publication and qdrant seams 2026-08-31 14:43:23 +02:00
marcopanandCodex 7cbbfbce91 docs: assess catalog postgres deployment constraints 2026-08-31 14:43:23 +02:00
marcopanandCodex 374e8aabc8 docs: inventory legacy metadata capabilities 2026-08-31 14:43:23 +02:00
Codex 57928347b3 docs: track ThothII conference presentation 2026-08-31 14:26:30 +02:00
Codex e2b88ce3c2 feat: add password visibility toggle 2026-08-30 18:25:27 +02:00
Codex cb40c09d9a fix: scope and batch sensitive suggestions 2026-08-30 17:12:20 +02:00
Codex 0736983bc5 feat: protect sensitive catalog samples 2026-08-30 12:14:23 +02:00
Codex 6278ee9d81 test: align workspace documentation verifier 2026-08-29 20:12:25 +02:00
Codex 0ce05869cf docs: reorganize operational documentation 2026-08-29 20:08:48 +02:00
Codex d504b1def1 docs(testing): record AI description acceptance 2026-08-29 16:47:13 +02:00
Codex 376dd5a09d feat: add AI catalog description generation 2026-08-29 16:42:56 +02:00
Codex b0afba81ca build(docs): lock MkDocs toolchain 2026-08-29 15:41:02 +02:00
Codex 7f968359ac test: align local compose service contract 2026-08-29 15:41:02 +02:00
Codex 58ee9cffe4 feat: add metadata catalog cleanup commands 2026-08-28 00:37:43 +02:00
Codex 79c4c925b5 feat: implement metadata catalog database management 2026-08-27 22:43:54 +02:00
Codex 705af3aeb2 fix(frontend): restrict management controls to admins 2026-08-26 20:40:15 +02:00
Codex 9189450fa9 fix(frontend): match database management button to workspace management 2026-08-26 20:12:40 +02:00
Codex 410cb547d2 fix(frontend): show database management entry to local users 2026-08-26 20:08:25 +02:00
Codex a701b19a03 feat(frontend): add database management surface 2026-08-26 18:00:47 +02:00
Codex 11f78c4467 merge: reconcile GitHub main into canonical Gitea main
Publish documentation / publish (push) Successful in 33s
2026-08-26 13:36:33 +02:00
Codex 8b5892a5a1 Merge remote-tracking branch 'origin/main' into codex/evidence-restructuring 2026-08-26 12:46:24 +02:00
Codex 9898726069 feat(evidence): structure v3 domain rules for review 2026-08-26 12:38:30 +02:00
Codex 9d4f994d3e feat(evidence): add table-free v3 and design guidance 2026-08-26 12:15:40 +02:00
Codex 38f02cfd08 feat: complete evidence restructuring worktree 2026-08-26 11:39:02 +02:00
Codex e910c7d49c docs: use English MkDocs navigation
Publish documentation / publish (push) Successful in 37s
2026-08-26 10:59:59 +02:00
Codex 7d32bb1e74 docs: publish English public documentation
Publish documentation / publish (push) Successful in 43s
2026-08-26 10:54:44 +02:00
marcopan a54d4769dd docs: focus public documentation on product usage 2026-08-26 10:15:07 +02:00
marcopan 23bc2f6555 docs: add Mermaid architecture diagrams 2026-08-26 10:00:50 +02:00
marcopan c1290c782c docs: fix Mermaid evidence flowchart syntax 2026-08-26 09:53:19 +02:00
marcopan a8cde2217f docs: expose CLI and evidence pages in navigation 2026-08-26 09:47:48 +02:00
marcopan 29326b064f docs: set Gitea publication URLs 2026-08-26 09:33:35 +02:00
marcopan d991dc2fd1 ci: publish MkDocs with Gitea Actions 2026-08-26 09:24:31 +02:00
marcopan f48196a57f chore: commit remaining worktree changes 2026-08-26 08:10:37 +02:00
marcopan ec061c42d4 docs: add architecture diagrams and evidence guide 2026-08-26 08:09:06 +02:00
Marco PancottiandGitHub 7d0d2b2edc Merge pull request #48 from mptyl/codex/evidence-restructuring
Fix Evidence authoring and complete #47 gates
2026-08-26 04:55:39 +02:00
marcopan dbd7787573 docs record Unix restore ownership repair 2026-08-26 04:26:22 +02:00
marcopan 71a42fbe80 fix restore preserve Unix file ownership 2026-08-26 04:25:55 +02:00
marcopan f16c248019 docs: record private registry fingerprint repair 2026-08-26 03:55:32 +02:00
marcopan a0620ffffa fix(ci): fingerprint private server registry as root 2026-08-26 03:55:18 +02:00
marcopan d638473c12 docs: record privileged server compose repair 2026-08-26 03:53:25 +02:00
marcopan 9a62fce1ad fix(ci): run server compose through privileged surface 2026-08-26 03:53:06 +02:00
marcopan dfc605da78 docs: record server restore ownership repair 2026-08-26 03:26:39 +02:00
marcopan 663c60dc3e fix(ci): handle root-owned server restore artifacts 2026-08-26 03:26:17 +02:00
marcopan 3af59cecbf docs: record projected revision repair 2026-08-26 02:58:54 +02:00
marcopan 8b524b4314 fix(auth): expose raw projected config revision 2026-08-26 02:58:33 +02:00
marcopan 737ec879d9 docs: record server status smoke repair 2026-08-26 02:33:56 +02:00
marcopan 0df95e337e fix(ci): assert projected server auth status 2026-08-26 02:33:39 +02:00
marcopan 52caabf83b docs: record projected auth smoke repair 2026-08-26 02:12:09 +02:00
marcopan a3259ced99 fix(ci): verify projected server auth layout 2026-08-26 02:11:36 +02:00
marcopan 20341a7124 docs(evidence): record server secret projection repair 2026-08-26 01:46:10 +02:00
marcopan 2b6bb058d8 fix(deploy): project server secrets for core uid 2026-08-26 01:45:57 +02:00
marcopan a1f1a63c88 docs(evidence): record server workspace smoke repair 2026-08-26 01:20:07 +02:00
marcopan 73efeb7f3d fix(deploy): make server workspace fixture readable 2026-08-26 01:19:53 +02:00
marcopan 5877adcfac docs(evidence): record session CA smoke repair 2026-08-26 00:56:44 +02:00
marcopan 287ce91e67 fix(deploy): seed session CA in local smoke 2026-08-26 00:56:30 +02:00
marcopan 2ae11f26f2 docs(evidence): record final Linux smoke repair 2026-08-26 00:44:59 +02:00
marcopan 12d257056f fix(deploy): project server session secrets in smoke 2026-08-26 00:44:30 +02:00
marcopan 4c2c9b50a9 docs(evidence): record final deployment repair 2026-08-26 00:12:46 +02:00
marcopan e51a6a2253 fix(deploy): prepare server auth projection root 2026-08-26 00:12:25 +02:00
marcopan 944b0edf7a docs(evidence): record PSD real acceptance 2026-08-25 23:50:00 +02:00
marcopan d49c644b61 fix(deploy): prepare server auth root before configure 2026-08-25 23:46:32 +02:00
marcopan 66f9fa2821 test(workspaces): make lock contention deterministic 2026-08-25 21:39:35 +02:00
marcopan 4499869c75 fix(ci): trust projected server smoke descriptor 2026-08-25 21:36:47 +02:00
marcopan 886faed886 fix(ci): project server auth in deployment smoke 2026-08-25 21:26:07 +02:00
marcopan 4606ec19a9 fix(evidence): support nested workspace roots 2026-08-25 20:54:23 +02:00
marcopan 6afb5d242b fix(ci): isolate generated smoke evidence 2026-08-25 20:30:17 +02:00
marcopan 97c6788f16 fix(ci): reclaim space for semantic backup smoke 2026-08-25 20:12:34 +02:00
marcopan a6655a5e2c fix(ci): preserve secure secret mount ownership 2026-08-25 19:56:48 +02:00
marcopan fb4a1fa25b fix(ci): complete Pi provider fixture metadata 2026-08-25 19:23:57 +02:00
marcopan a8a80b9846 fix(ci): pin maintenance smoke image 2026-08-25 19:10:06 +02:00
marcopan 73b784a176 fix(ci): project private application secrets 2026-08-25 18:47:49 +02:00
marcopan fd878b8c3e fix(workspace): isolate maintenance authentication surface 2026-08-25 18:37:39 +02:00
marcopan db375298d0 fix(test): hermetically exercise auth workflows 2026-08-25 18:27:41 +02:00
marcopan 2c2bf1940b fix(ci): project private runtime fixtures 2026-08-25 18:04:13 +02:00
marcopan abfcfb0e06 fix(ci): use supported Compose create syntax 2026-08-25 17:37:48 +02:00
marcopan f0e78b1ce3 fix(ci): project local auth for core runtime 2026-08-25 17:35:29 +02:00
marcopan 8980c35198 fix(ci): align Windows Compose topology 2026-08-25 17:14:42 +02:00
marcopan 6d0cb6d997 fix(ci): provision integration fixtures 2026-08-25 17:07:51 +02:00
marcopan 06b26cf66f fix(ci): refresh reviewed deployment blocks 2026-08-25 16:55:20 +02:00
marcopan 5b6fec939a fix(ci): isolate npm release configuration 2026-08-25 16:47:56 +02:00
marcopan 3e106d5262 fix(ci): restore deployment release gates 2026-08-25 16:45:56 +02:00
marcopan 8ba87b68dc fix(evidence): stabilize real Pi authoring 2026-08-25 15:21:59 +02:00
marcopan 610ae8c85a fix(auth): make local verification portable
Keep upstream identity visible while limiting logout to local auth. Inject the restore privilege gate so the deterministic core tests do not depend on the host OS, and confine descriptor-backed projection tests to Linux. Accept the real remaining Pi timeout budget instead of an exact millisecond.
2026-08-25 10:50:30 +02:00
marcopan e9c65ef2db fix(evidence): resolve final review findings 2026-08-25 03:03:11 +02:00
marcopan d4818c8cc3 docs(evidence): narrow owner gate baseline 2026-08-25 02:43:01 +02:00
marcopan 8542124f27 test(evidence): harden owner gate acceptance 2026-08-25 02:34:10 +02:00
marcopan fc83d29b58 test(evidence): record restructuring acceptance 2026-08-25 02:17:05 +02:00
marcopan 0caa747914 fix(evidence): validate curated runtime corpus 2026-08-25 01:59:27 +02:00
marcopan 9f0184388d docs(evidence): publish curated-only workspace contract 2026-08-25 01:52:27 +02:00
marcopan 619ac2e141 feat(evidence): evaluate retrieval with a small fixture 2026-08-25 01:44:12 +02:00
marcopan dcb5acc312 fix(evidence): preserve legacy formula migration boundaries 2026-08-25 01:28:54 +02:00
marcopan cc30148b69 refactor(evidence): unify formulas with typed evidence 2026-08-25 01:23:04 +02:00
marcopan d5d65f3659 fix(evidence): index only curated legacy documents 2026-08-25 01:02:41 +02:00
marcopan f1a9b567ba feat(evidence): contribute to semantic stages (#42) 2026-08-24 22:05:30 +02:00
marcopan 3420c57c8b feat(evidence): build semantic fragments from typed units 2026-08-24 21:33:48 +02:00
marcopan 8adc085746 feat(evidence): resolve curated evidence explicitly 2026-08-24 20:52:56 +02:00
marcopan f5c7cc6198 feat(evidence): prepare curated evidence incrementally 2026-08-24 20:08:27 +02:00
marcopan 6176410f42 fix(evidence): fail closed when BM25 is unavailable 2026-08-24 18:21:21 +02:00
marcopan 29d41ac258 feat(evidence): use server-side Qdrant BM25 retrieval 2026-08-24 18:15:33 +02:00
marcopan 0e9add09a9 feat(evidence): add Qdrant BM25 vector in place 2026-08-24 18:01:46 +02:00
marcopan ae0976a4aa feat(evidence): validate canonical curated corpus 2026-08-24 17:38:22 +02:00
marcopan 5c6228f8c2 docs(evidence): finalize ticketed restructuring specification 2026-08-24 17:11:37 +02:00
marcopan d970e10264 docs(evidence): finalize restructuring design 2026-08-24 14:57:53 +02:00
marcopan 6062cb010e Chiusura fase di ristrutturazione e modularizzazione del workflow per favorire sviluppo modulare 2026-08-24 13:32:20 +02:00
marcopan fa2298653b fix(ui): stream phase progress without gates 2026-08-24 12:10:56 +02:00
marcopan 36a7a0ab33 refactor(workflow): contract shared core (#33) 2026-08-24 03:12:00 +02:00
marcopan f375515dc0 refactor(evidence): extract TypeScript lifecycle (#32) 2026-08-24 02:50:56 +02:00
marcopan 840848df94 refactor(evidence): remove legacy Python layout (#31) 2026-08-24 02:36:30 +02:00
marcopan 44f1efa5ba refactor(evidence): migrate acquisition and preprocessing (#30) 2026-08-24 02:25:11 +02:00
marcopan d5e78febd3 refactor(evidence): migrate runtime consumption (#29) 2026-08-24 02:10:06 +02:00
marcopan 8b63715b56 refactor(evidence): add cohesive Python facade (#28) 2026-08-24 02:02:57 +02:00
marcopan 1e459b073e refactor(disambiguation): own F1 clarification policy (#27) 2026-08-24 01:55:42 +02:00
marcopan 2ed55ef131 refactor(disambiguation): extract F3 rewrite path (#26) 2026-08-24 01:44:40 +02:00
marcopan eccf6212f1 refactor(pi): generate modular session instructions (#25) 2026-08-24 01:36:08 +02:00
marcopan 4a654de84a refactor(memory): own solved-question lifecycle (#24) 2026-08-24 01:23:39 +02:00
marcopan 93fe0d733b refactor(memory): extract F2 recall path (#23) 2026-08-24 01:08:15 +02:00
marcopan beac2e80f4 refactor(memory): extract F8 promotion gate (#22) 2026-08-24 00:51:53 +02:00
marcopan d15bb59c3d test(workflow): complete observable baseline (#21) 2026-08-24 00:40:11 +02:00
marcopan dc9726cb35 test(workflow): freeze observable contracts (#21) 2026-08-24 00:25:17 +02:00
marcopan b5db0cd3c1 docs: define modular workflow domain semantics
Deployment release gate / Linux Docker deployment and rollback (push) Canceled after 0s
Deployment release gate / Windows clone and Compose contract (push) Canceled after 0s
Deployment release gate / LF, Compose, docs, and TypeScript (push) Canceled after 0s
Deployment release gate / Hermetic authentication browser gate (push) Canceled after 0s
Deployment release gate / DWH authentication Nginx gate (push) Canceled after 0s
Deployment release gate / Native Windows Docker Desktop/WSL2 startup (push) Canceled after 0s
2026-08-23 14:03:20 +02:00
User bc1b5e79ca docs: update checkout path to Thoth 2026-08-22 22:50:36 +02:00
1228 changed files with 145648 additions and 82673 deletions
@@ -1,66 +0,0 @@
# Task 9 quality audit — final 5
**Scope:** the two blocking findings from `task9-quality-audit-final4.md` — unbound production
module graph at manual serve, and commit-addressed snapshots accepted without content identity at
render. Manual acceptance remains **PENDING**; no `VERDICT.md` was created.
## Verdict: APPROVED for the two final integrity blockers
### 1. Manual serve binds the complete `backend/dist` module graph, not only `server.js`
`prepare` now builds a post-build manifest of every regular `backend/dist` file
(relative path, size, SHA-256, device, inode) and writes it as an exclusive `0600` record
(`installation/runtime/backend-dist.manifest.json`) inside the owned root; `ownership.json`
records that record's path/device/inode/size/SHA-256. `serve` revalidates the manifest record
identity and bytes, revalidates every distribution file against it (no-follow, single inode,
size and digest), and refuses before spawning. The manifest descriptor is passed to the child on
fd 4 together with the entrypoint on fd 3. The immutable preload parses the manifest, verifies
the entrypoint cross-digest, reads and hash-verifies **every** file at startup, caches the
verified bytes, and its load hook serves **only** those cached bytes for any import below
`backend/dist` (entry URL still served from the bound fd-3 bytes). A same-path regular
replacement of any imported dependency is therefore refused before `RUNNING` (serve-time
validation), refused at child startup (startup verification), or rendered harmless (cached
bytes), and the parent revalidates the full manifest at `RUNNING` publication and at `stop`.
### 2. Renderer binds snapshot content to its commit identity
The generated render command validates the bounded saved read/publish revisions, the
commit-addressed owned snapshot path, the installed Git HEAD, and the bounded
`snapshot.json` manifest of that commit: `head` equals the commit, `files[<id>.yaml]` is the
SHA-256 of the snapshot bytes, the manifest revision binds commit/blob/snapshot path, the saved
revision blob equals the manifest blob, and `git rev-parse <commit>:workspaces/<id>.yaml` plus
`git hash-object` of the snapshot bytes both equal that blob. It passes the expected digest as
`--snapshot-sha256`. The renderer re-reads the bounded `snapshot.json` (`head`,
`files[<id>.yaml]` must equal the carried digest), opens the snapshot once with no-follow
semantics and bounded reads, renders only the digest-verified bytes, re-verifies around lease
publication, releases the lease in `finally`, and publishes no output on any refusal.
## Deterministic regressions added
- static regular replacement of an imported production dependency after `prepare` is refused,
no marker, no accepted PID record, no orphan;
- deterministic dependency check/load swap (`beforeSpawn` rename) is refused by the child's
startup verification, no marker, no PID record, no orphan;
- after `RUNNING`, a same-path regular dependency replacement is never executed: the loader
serves the verified cached bytes (health-visible source stays the original) and the marker is
absent;
- renderer refuses a same-path regular snapshot byte replacement against the carried digest and
manifest, with lease release and no output;
- renderer refuses manifest `head`, `files` digest, expected-digest, missing, and malformed
cases, with lease release and no output;
- wrapper refuses missing manifest, manifest head/digest/revision tampering, saved-revision blob
mismatch, Git blob mismatch, and snapshot-vs-Git-bytes mismatch, and passes the exact
`--snapshot-sha256` on the valid path (stub renderer records arguments).
## Verification
- `bash scripts/test-p1-manual-acceptance.sh` (backend build + both suites): **59 tests, 59
pass, 0 fail**; no `8791/8792` listener and no `--p1-manual-nonce` process remain.
- `npx tsc --noEmit -p .` (backend): PASS.
- Real-repository `prepare` + `cleanup` cycle: 39 distribution files bound, entrypoint
cross-digest verified, owned root fully removed afterwards.
- Diff check: only the seven Task 9 paths are touched; no Task 8 file was modified.
- This report and the implementation contain no fixture secret or canary values.
Manual acceptance remains **PENDING** by design; the walkthrough and human verdict are
unchanged.
-282
View File
@@ -1,282 +0,0 @@
{
"schema": "thothii-task4-certification-v1",
"generated_on": "2026-08-18",
"started_at_utc": "2026-08-18T14:16:40Z",
"ended_at_utc": "2026-08-18T14:20:10Z",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"source_immutability": {
"status": "PASS",
"tracked_changes_after_freeze": false,
"allowed_untracked": [".playwright-cli/", ".thothctl/"]
},
"source_commits": {
"task4_candidate": "b31b27e5845ffd3adf311429367319beaba263c7",
"task1": "d43738eeae6d14bb5e470093058b069a983f5372",
"task2": "5f9a3ae066a060b43a11a959b60a1efadd1c2425",
"task3": "0d8e707533fada938c99eb06f8457150e7ef2b40",
"task3_follow_up": "b31b27e5845ffd3adf311429367319beaba263c7",
"fix_round_1_source": "10cd66fe6a5b484a4dc569326a228c1c5484a5d4",
"fix_round_2_source": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"historical_task15_final": "74b062f1a737103524cbe706346cfd65f87cdfd1"
},
"versions": {
"node_contract": "v24.16.0",
"node_host_default": "v25.6.1",
"go": "go1.26.5",
"pi": "0.80.3"
},
"retained_report": ".superpowers/sdd/2026-08-16-thothii-authentication/task-15-report.md",
"task4_report": ".superpowers/sdd/2026-08-18-thothii-authentication-remediation/task-4-report.md",
"fix_round_2_report": ".superpowers/sdd/2026-08-18-thothii-authentication-remediation/fix-round-2-report.md",
"workflow": {
"run_id": "32147345625",
"url": "https://github.com/mptyl/ThothII/actions/runs/32147345625",
"event": "workflow_dispatch",
"head_sha": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"status": "completed",
"conclusion": "failure",
"windows_job": {
"name": "Windows clone and Compose contract",
"job_id": "95744249248",
"url": "https://github.com/mptyl/ThothII/actions/runs/32147345625/job/95744249248",
"conclusion": "failure",
"native_step": "Run native Windows retained-capability tests",
"native_step_conclusion": "success",
"command": "go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1",
"requested_packages": ["internal/safeio", "internal/backup", "internal/authstorage"],
"executed_packages": ["internal/safeio", "internal/backup", "internal/authstorage"],
"not_executed_packages": [],
"package_results": {
"internal/safeio": "PASS (22.058s)",
"internal/backup": "PASS (7.161s)",
"internal/authstorage": "PASS (16.088s)"
},
"failed_step": "Verify Windows clone contract",
"failure_category": "baseline_powershell_parser",
"failure_detail": "scripts/test-windows-clone-contract.ps1:208 parses $remoteYaml: as an invalid variable reference"
},
"lf_compose_docs_typescript_job": {
"name": "LF, Compose, docs, and TypeScript",
"job_id": "95744249458",
"url": "https://github.com/mptyl/ThothII/actions/runs/32147345625/job/95744249458",
"conclusion": "failure",
"failed_step": "Verify Compose and installation contracts",
"category": "baseline_ci_contract",
"detail": "unified Compose contract passed; test-no-deployment-coupling-scope.sh stopped on TMPDIR: unbound variable",
"downstream_steps": "skipped"
},
"linux_docker_job": {
"name": "Linux Docker deployment and rollback",
"job_id": "95744249354",
"url": "https://github.com/mptyl/ThothII/actions/runs/32147345625/job/95744249354",
"conclusion": "failure",
"failed_step": "Run unified deployment smoke",
"category": "infrastructure_prerequisite",
"detail": "Task 13 smoke failed before deployment because rg is required",
"cleanup": "PASS",
"image_manifest": "not_generated"
},
"windows_docker_startup_job": {
"name": "Native Windows Docker Desktop/WSL2 startup",
"job_id": "95744250450",
"url": "https://github.com/mptyl/ThothII/actions/runs/32147345625/job/95744250450",
"status": "NOT_RUN",
"classification": "BLOCKED",
"workflow_conclusion": "skipped",
"reason": "workflow conditions skipped the job; no Windows Docker Desktop/WSL2 command executed"
}
},
"docker_image_evidence": {
"authentication_smoke": {
"status": "PASS",
"docker_images": [],
"reason": "no_docker_images_exercised"
},
"unified_docker_smoke": {
"status": "FAIL",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"run_id": "32147345625",
"workflow_job_id": "95744249354",
"manifest": ".artifacts/task-15/unified-docker-images.json",
"reason": "workflow attempt stopped before deployment because rg is required",
"cleanup": "PASS",
"images": 0,
"historical": {
"status": "PASS",
"source_commit": "74b062f1a737103524cbe706346cfd65f87cdfd1",
"run_id": "20260818070637-66409-30058",
"manifest_sha256": "9c8dec4546909fd93799dbcf374bcb3a89bc46cfe0fd482472c0cbe757ddf5b6",
"images": 5,
"cleanup": "PASS"
}
}
},
"gates": {
"posix_registry_ownership": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"evidence": "backend Node 24 full suite including local-registry ownership coverage"
},
"stagearchive_unix_retained_capability": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"evidence": "focused safeio/backup tests, Go race suite, and Unix ancestor-swap coverage"
},
"windows_stagearchive_retained_capability": {
"status": "PASS",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"evidence": "native Windows backup package passed, including the two-file shared retained-root staging test"
},
"windows_claim_retained_capability": {
"status": "PASS",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"evidence": "native Windows safeio and authstorage packages passed concurrent claim/consume coverage"
},
"workflow_lf_compose_docs_typescript": {
"status": "FAIL",
"classification": "baseline_ci_contract",
"reason": "TMPDIR was unset after the unified Compose contract passed"
},
"workflow_linux_docker": {
"status": "FAIL",
"classification": "infrastructure_prerequisite",
"reason": "runner did not provide rg; cleanup proof passed and no image manifest was generated"
},
"go_security_build": {
"status": "PASS",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"focused_packages": 3,
"race_packages": 18,
"focused_test": "PASS",
"race": "PASS",
"vet": "PASS",
"host_build": "PASS"
},
"windows_cross_compile": {
"status": "PASS",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"focused_test_packages": 3,
"cli_build": "PASS",
"execution": "cross_compile_only_not_native_execution"
},
"backend_node24": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"node": "v24.16.0",
"files": 76,
"tests": 1092,
"typecheck": "PASS",
"build": "PASS",
"note": "an initial full run had one workspace-registry timeout; focused rerun and complete rerun passed"
},
"frontend_node24": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"node": "v24.16.0",
"files": 61,
"tests": 444,
"typecheck": "PASS",
"build": "PASS"
},
"authentication_and_f1_smoke": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"node": "v24.16.0",
"filtered_e2e": "1 passed",
"sentinel_leak_scan": "PASS"
},
"harness_pytest": {
"status": "FAIL",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7",
"passed": 951,
"failed": 1,
"skipped": 4,
"subtests": 232,
"failure": "test_column_decisions::test_f4_emits_column_types: workflow.yaml not found from harness test cwd"
},
"authentication_docs": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7"
},
"shell_syntax": {
"status": "PASS",
"source_commit": "b31b27e5845ffd3adf311429367319beaba263c7"
},
"authentication_smoke_runtime": {
"status": "PASS",
"node": "v24.16.0",
"sentinel_leak_scan": "PASS"
},
"compose_default": {
"status": "FAIL",
"reason": "required THT_WORKSPACE_GIT_REMOTE was not available"
},
"compose_unified": {
"status": "FAIL",
"reason": "compose.unified.yaml is absent from the frozen source"
},
"unified_docker_smoke": {
"status": "FAIL",
"source_commit": "2a9359071257f9b8a71d36ec2bbb25b161003f81",
"workflow_run_id": "32147345625",
"reason": "remote workflow attempted the smoke but stopped before deployment because rg is required",
"cleanup": "PASS",
"image_manifest": "not_generated"
},
"ruff": {
"status": "FAIL",
"errors": 192,
"classification": "known_baseline"
},
"mkdocs_strict": {
"status": "NOT_RUN",
"classification": "BLOCKED",
"historical_status": "FAIL",
"historical_warnings": 69
},
"canonical_install_docs": {
"status": "NOT_RUN",
"classification": "BLOCKED",
"historical_status": "FAIL"
},
"workspace_install_docs": {
"status": "NOT_RUN",
"classification": "BLOCKED",
"historical_status": "FAIL"
},
"pi_user_auth_compose": {
"status": "NOT_RUN",
"classification": "BLOCKED",
"historical_status": "FAIL"
},
"deployment_coupling": {
"status": "NOT_RUN",
"classification": "BLOCKED",
"historical_status": "FAIL"
},
"l2": {
"status": "PENDING",
"reason": "configured secret layout unavailable; gate not run after stop"
},
"manual_psd": {
"status": "PENDING",
"reason": "approved real identity/access unavailable; gate not run after stop"
},
"provider_readiness": {
"status": "PENDING",
"reason": "provider prerequisite unavailable; gate not run after stop"
}
},
"review": {
"original_important_findings_resolved": 3,
"fix_round_2_important_lifecycle": "ADDRESSED",
"fix_round_2_minor_windows_diagnostics": "ADDRESSED",
"verdict": "PASS",
"reason": "the lifecycle controller is bounded and cancellation-aware with cancel, bounded join, and lock-release proof; the temporary Windows diagnostic matrix is removed; exact-source native safeio, backup, and authstorage all pass"
},
"remediation_status": "PASS",
"release_complete": false,
"authentication_implementation_complete": true,
"release_readiness": "FAIL",
"release_readiness_pending_external_gates": true
}
@@ -1,54 +0,0 @@
{
"gate": "unified-deployment-smoke",
"status": "pass",
"source_commit": "74b062f1a737103524cbe706346cfd65f87cdfd1",
"run_id": "20260818070637-66409-30058",
"images": [
{
"id": "sha256:2d7b19491c7eb8c119c3cedb390aaeb2ff5593f6fc43ab66c317565560da6d7d",
"roles": [
"compose-runtime",
"fixture-runtime"
],
"repo_digests": [
"sha256:2d7b19491c7eb8c119c3cedb390aaeb2ff5593f6fc43ab66c317565560da6d7d"
]
},
{
"id": "sha256:3b6c31a5d8f8fc58fa3233391b6175bd2fbc793eebb44d5e285ecc6e02e9e687",
"roles": [
"compose-runtime"
],
"repo_digests": [
"sha256:3b6c31a5d8f8fc58fa3233391b6175bd2fbc793eebb44d5e285ecc6e02e9e687"
]
},
{
"id": "sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a",
"roles": [
"compose-runtime"
],
"repo_digests": [
"sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a"
]
},
{
"id": "sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c",
"roles": [
"compose-runtime"
],
"repo_digests": [
"sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c"
]
},
{
"id": "sha256:c3cbe1cc1aa588a64951ac6286e0df7b27fe2e6324b1001c619bb358770c0178",
"roles": [
"rollback-candidate"
],
"repo_digests": [
"sha256:c3cbe1cc1aa588a64951ac6286e0df7b27fe2e6324b1001c619bb358770c0178"
]
}
]
}
-11
View File
@@ -1,11 +0,0 @@
{
"version": "0.0.1",
"configurations": [
{
"name": "replay",
"runtimeExecutable": "node",
"runtimeArgs": ["tools/replay/server.mjs"],
"port": 5333
}
]
}
+7 -2
View File
@@ -4,6 +4,8 @@
**/__pycache__
**/.pytest_cache
**/dist
frontend/prototypes/
frontend/vite.database-management-prototype.config.ts
**/*.pyc
.git
.worktrees
@@ -15,6 +17,7 @@
!deploy/env/*.env.example
deploy/thothii.env
deploy/secrets/
deploy/psd/
harness/workspaces/*.yaml
!harness/workspaces/local.yaml
!harness/workspaces/tht.example.yaml
@@ -27,5 +30,7 @@ coverage/
data/
sessions/
workspace-registry/
# docs/site (mkdocs build) — non necessari nelle immagini
docs/superpowers/plans
.tht/
deploy/local/
+1
View File
@@ -9,4 +9,5 @@ Dockerfile* text eol=lf
*.tsx text eol=lf
*.py text eol=lf
*.md text eol=lf
*.pptx binary
*.ps1 text eol=crlf
+71
View File
@@ -0,0 +1,71 @@
name: Publish documentation
on:
push:
branches:
- main
paths:
- "docs/**"
- "mkdocs.yml"
- "scripts/build-docs.sh"
- "scripts/verify-public-docs.py"
- "scripts/test-verify-public-docs.py"
- "scripts/verify-auth-docs.py"
- "scripts/test-verify-auth-docs.py"
- "docs/requirements.txt"
- ".gitea/workflows/publish-docs.yml"
workflow_dispatch:
permissions:
contents: write
concurrency:
group: documentation
cancel-in-progress: true
jobs:
publish:
runs-on: ubuntu-latest
env:
GITEA_TOKEN: ${{ secrets.GITEA_TOKEN }}
REPOSITORY_URL: ${{ gitea.server_url }}/${{ gitea.repository }}.git
steps:
- name: Checkout documentation source
uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.x"
cache: pip
cache-dependency-path: docs/requirements.lock
- name: Install MkDocs dependencies
run: python -m pip install -r docs/requirements.lock
- name: Test public documentation boundary
run: python scripts/test-verify-public-docs.py
- name: Test current authentication documentation
run: |
python scripts/verify-auth-docs.py auth
python scripts/verify-auth-docs.py dwh
python scripts/test-verify-auth-docs.py auth
python scripts/test-verify-auth-docs.py dwh
- name: Build documentation
run: |
mkdocs build --strict
python scripts/verify-public-docs.py
- name: Publish generated site to the pages branch
working-directory: site
run: |
git init
git config user.name "Gitea Actions"
git config user.email "actions@${{ gitea.server_url }}"
git add --all
git commit --message "Publish documentation for ${{ gitea.sha }}"
git -c http.extraheader="Authorization: token ${GITEA_TOKEN}" \
push --force "${REPOSITORY_URL}" HEAD:pages
+24 -1
View File
@@ -36,6 +36,10 @@ jobs:
with:
node-version: "24.16.0"
package-manager-cache: false
- name: Install release gate prerequisites
run: |
sudo apt-get update
sudo apt-get install --yes --no-install-recommends ripgrep
- name: Verify shell syntax and LF policy
run: |
git ls-files -z '*.sh' | xargs -0 -n1 bash -n
@@ -46,7 +50,6 @@ jobs:
bash scripts/test-no-deployment-coupling-scope.sh
bash scripts/test-compose-secret-policy.sh
bash scripts/test-no-deployment-coupling.sh
bash scripts/test-preprocess-compose-config.sh
bash scripts/test-verify-workspace-install-docs.sh
git diff --check
- name: Assert clean checkout before release trust bootstrap
@@ -60,6 +63,14 @@ jobs:
run: |
bash scripts/test-server-pi-state-topology.sh
bash scripts/unified-deployment-smoke.sh --self-test
- name: Install harness CLI for backend integration tests
working-directory: harness
run: |
python3 -m venv .venv
.venv/bin/python -m pip install -e .
- name: Install backend dependencies
working-directory: backend
run: npm ci
- name: Test and type-check backend
working-directory: backend
run: |
@@ -140,11 +151,23 @@ jobs:
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
- name: Install release gate prerequisites
run: |
sudo apt-get update
sudo apt-get install --yes --no-install-recommends ripgrep
- name: Reclaim unused hosted-runner space
run: bash scripts/prepare-linux-docker-runner.sh
- name: Run unified deployment smoke
env:
TASK13_IMAGE_EVIDENCE_OUTPUT: ${{ runner.temp }}/task13-images.json
run: timeout --signal=TERM --kill-after=45s 32m bash scripts/unified-deployment-smoke.sh
- name: Run tht update smoke
env:
TASK13_IMAGE_EVIDENCE_OUTPUT: ${{ runner.temp }}/task13-images.json
run: timeout --signal=TERM --kill-after=45s 32m bash scripts/tht-update-smoke.sh
- name: Run Linux server deployment smoke
env:
TASK13_IMAGE_EVIDENCE_OUTPUT: ${{ runner.temp }}/task13-images.json
run: timeout --signal=TERM --kill-after=45s 32m bash scripts/server-deployment-smoke.sh
windows-clone:
+3 -2
View File
@@ -5,8 +5,6 @@
ChironeWp3/
Thoth/
# === Visual companion brainstorming artifacts (local-only) ===
.superpowers/
.worktrees/
.tht/
@@ -37,6 +35,7 @@ config/ca-chain.pem
# ThothII deployment configuration and secret values (keep only the README tracked)
deploy/.env
deploy/env/local.env
deploy/compose.connector-secrets.local.yaml
deploy/compose.psd-local.yaml
deploy/workspaces/psd.yaml
@@ -47,6 +46,8 @@ deploy/secrets/*
# Per-installation configuration generated by `tht setup` (examples stay tracked).
deploy/*/thothii-installation.yaml
deploy/*/operator.env
deploy/*/auth/
deploy/*/generated/
deploy/*/secrets/*
!deploy/*/secrets/.gitkeep
!deploy/*/secrets/*.example
+203
View File
@@ -0,0 +1,203 @@
{
"schemaVersion": 2,
"generatedAt": "2026-08-26T10:10:15.926Z",
"title": "Design System: ThothII",
"extensions": {
"colorMeta": {
"instrument-red": {
"role": "primary",
"displayName": "Instrument Red",
"canonical": "oklch(55.87% 0.1881 23.2)",
"tonalRamp": ["oklch(15% 0.07 23.2)", "oklch(28% 0.12 23.2)", "oklch(42% 0.16 23.2)", "oklch(56% 0.1881 23.2)", "oklch(68% 0.17 23.2)", "oklch(78% 0.13 23.2)", "oklch(88% 0.07 23.2)", "oklch(95% 0.03 23.2)"]
},
"instrument-red-hover": {
"role": "primary",
"displayName": "Instrument Red Pressed",
"canonical": "oklch(50.95% 0.1812 24.1)",
"tonalRamp": ["oklch(15% 0.07 24.1)", "oklch(28% 0.12 24.1)", "oklch(42% 0.16 24.1)", "oklch(51% 0.1812 24.1)", "oklch(68% 0.16 24.1)", "oklch(78% 0.12 24.1)", "oklch(88% 0.07 24.1)", "oklch(95% 0.03 24.1)"]
},
"porcelain-background": {
"role": "neutral",
"displayName": "Porcelain Background",
"canonical": "oklch(99.18% 0.0011 17.2)",
"tonalRamp": ["oklch(15% 0.0011 17.2)", "oklch(28% 0.0011 17.2)", "oklch(42% 0.0011 17.2)", "oklch(56% 0.0011 17.2)", "oklch(68% 0.0011 17.2)", "oklch(78% 0.0011 17.2)", "oklch(88% 0.0011 17.2)", "oklch(95% 0.0011 17.2)"]
},
"porcelain-card": {
"role": "neutral",
"displayName": "Porcelain Card",
"canonical": "oklch(99.85% 0.0006 17.2)",
"tonalRamp": ["oklch(15% 0.0006 17.2)", "oklch(28% 0.0006 17.2)", "oklch(42% 0.0006 17.2)", "oklch(56% 0.0006 17.2)", "oklch(68% 0.0006 17.2)", "oklch(78% 0.0006 17.2)", "oklch(88% 0.0006 17.2)", "oklch(95% 0.0006 17.2)"]
},
"warm-surface": {
"role": "neutral",
"displayName": "Warm Surface",
"canonical": "oklch(97.09% 0.0011 17.2)",
"tonalRamp": ["oklch(15% 0.0011 17.2)", "oklch(28% 0.0011 17.2)", "oklch(42% 0.0011 17.2)", "oklch(56% 0.0011 17.2)", "oklch(68% 0.0011 17.2)", "oklch(78% 0.0011 17.2)", "oklch(88% 0.0011 17.2)", "oklch(95% 0.0011 17.2)"]
},
"sunken-surface": {
"role": "neutral",
"displayName": "Sunken Surface",
"canonical": "oklch(94.08% 0.0011 17.2)",
"tonalRamp": ["oklch(15% 0.0011 17.2)", "oklch(28% 0.0011 17.2)", "oklch(42% 0.0011 17.2)", "oklch(56% 0.0011 17.2)", "oklch(68% 0.0011 17.2)", "oklch(78% 0.0011 17.2)", "oklch(88% 0.0011 17.2)", "oklch(95% 0.0011 17.2)"]
},
"warm-graphite": {
"role": "neutral",
"displayName": "Warm Graphite",
"canonical": "oklch(26.78% 0.0097 355.6)",
"tonalRamp": ["oklch(15% 0.0097 355.6)", "oklch(28% 0.0097 355.6)", "oklch(42% 0.0097 355.6)", "oklch(56% 0.0097 355.6)", "oklch(68% 0.008 355.6)", "oklch(78% 0.006 355.6)", "oklch(88% 0.004 355.6)", "oklch(95% 0.002 355.6)"]
},
"muted-graphite": {
"role": "neutral",
"displayName": "Muted Graphite",
"canonical": "oklch(51.33% 0.0088 345.6)",
"tonalRamp": ["oklch(15% 0.0088 345.6)", "oklch(28% 0.0088 345.6)", "oklch(42% 0.0088 345.6)", "oklch(56% 0.0088 345.6)", "oklch(68% 0.007 345.6)", "oklch(78% 0.005 345.6)", "oklch(88% 0.003 345.6)", "oklch(95% 0.002 345.6)"]
},
"quiet-border": {
"role": "neutral",
"displayName": "Quiet Border",
"canonical": "oklch(90.93% 0.0035 354.7)",
"tonalRamp": ["oklch(15% 0.0035 354.7)", "oklch(28% 0.0035 354.7)", "oklch(42% 0.0035 354.7)", "oklch(56% 0.0035 354.7)", "oklch(68% 0.0035 354.7)", "oklch(78% 0.0035 354.7)", "oklch(88% 0.003 354.7)", "oklch(95% 0.002 354.7)"]
},
"success-mint": {
"role": "secondary",
"displayName": "Success Mint",
"canonical": "oklch(75.77% 0.1581 165)",
"tonalRamp": ["oklch(15% 0.06 165)", "oklch(28% 0.1 165)", "oklch(42% 0.14 165)", "oklch(56% 0.1581 165)", "oklch(68% 0.15 165)", "oklch(78% 0.12 165)", "oklch(88% 0.07 165)", "oklch(95% 0.03 165)"]
},
"warning-amber": {
"role": "tertiary",
"displayName": "Warning Amber",
"canonical": "oklch(85.23% 0.1386 78.9)",
"tonalRamp": ["oklch(15% 0.05 78.9)", "oklch(28% 0.09 78.9)", "oklch(42% 0.12 78.9)", "oklch(56% 0.1386 78.9)", "oklch(68% 0.13 78.9)", "oklch(78% 0.1 78.9)", "oklch(88% 0.06 78.9)", "oklch(95% 0.025 78.9)"]
},
"information-blue": {
"role": "tertiary",
"displayName": "Information Blue",
"canonical": "oklch(70.35% 0.1128 221.3)",
"tonalRamp": ["oklch(15% 0.045 221.3)", "oklch(28% 0.075 221.3)", "oklch(42% 0.1 221.3)", "oklch(56% 0.1128 221.3)", "oklch(68% 0.105 221.3)", "oklch(78% 0.08 221.3)", "oklch(88% 0.045 221.3)", "oklch(95% 0.02 221.3)"]
}
},
"typographyMeta": {
"display": {"displayName": "Display", "purpose": "Authentication and exceptional page-level statements only."},
"headline": {"displayName": "Headline", "purpose": "Major page and persisted artifact titles."},
"title": {"displayName": "Title", "purpose": "Panel and document section hierarchy."},
"body": {"displayName": "Body", "purpose": "Operational prose and sustained reading."},
"control": {"displayName": "Control", "purpose": "Buttons, inputs, tabs, and compact actions."},
"label": {"displayName": "Machine Label", "purpose": "Uppercase metadata and machine-oriented micro-labels."}
},
"shadows": [
{"name": "contact", "value": "0 1px 2px oklch(var(--shadow-tint) / 0.05)", "purpose": "Contact shadow for controls and code blocks."},
{"name": "panel", "value": "0 1px 2px oklch(var(--shadow-tint) / 0.05), 0 2px 6px -1px oklch(var(--shadow-tint) / 0.05)", "purpose": "Small structural lift for selected cards."},
{"name": "overlay", "value": "0 2px 4px -2px oklch(var(--shadow-tint) / 0.06), 0 12px 32px -8px oklch(var(--shadow-tint) / 0.1)", "purpose": "Broad low-opacity lift for dialogs and floating layers."}
],
"motion": [
{"name": "control-feedback", "value": "140ms cubic-bezier(0.22, 1, 0.36, 1)", "purpose": "Button hover, focus, and press feedback."},
{"name": "overlay-transition", "value": "100ms ease-out", "purpose": "Dialog fade and scale transitions."},
{"name": "activity-pulse", "value": "1.5s ease-in-out infinite", "purpose": "Live model activity only; disabled for reduced motion."}
],
"breakpoints": [
{"name": "sm", "value": "640px"},
{"name": "lg", "value": "1024px"}
]
},
"components": [
{
"name": "Primary Button",
"kind": "button",
"refersTo": "button-primary",
"description": "The authoritative action for the current workflow step.",
"html": "<button class=\"ds-button-primary\">Confirm review</button>",
"css": ".ds-button-primary { display:inline-flex; align-items:center; justify-content:center; height:32px; padding:0 14px; border:1px solid transparent; border-radius:8px; background:oklch(var(--primary)); color:oklch(var(--primary-foreground)); font:600 14px/1.25 var(--font-sans); letter-spacing:0.005em; box-shadow:var(--shadow-xs); transition:color 140ms cubic-bezier(0.22,1,0.36,1),background-color 140ms cubic-bezier(0.22,1,0.36,1),box-shadow 140ms cubic-bezier(0.22,1,0.36,1),transform 140ms cubic-bezier(0.22,1,0.36,1); } .ds-button-primary:hover { background:oklch(var(--primary-hover)); } .ds-button-primary:focus-visible { outline:3px solid oklch(var(--ring)/0.25); outline-offset:2px; } .ds-button-primary:active { transform:scale(0.97); box-shadow:none; }"
},
{
"name": "Outline Button",
"kind": "button",
"refersTo": "button-secondary",
"description": "A compact secondary action that preserves the primary action hierarchy.",
"html": "<button class=\"ds-button-outline\">Inspect details</button>",
"css": ".ds-button-outline { display:inline-flex; align-items:center; justify-content:center; height:32px; padding:0 14px; border:1px solid oklch(var(--border)); border-radius:8px; background:oklch(var(--card)); color:oklch(var(--foreground)); font:600 14px/1.25 var(--font-sans); box-shadow:var(--shadow-xs); transition:background-color 140ms cubic-bezier(0.22,1,0.36,1),transform 140ms cubic-bezier(0.22,1,0.36,1); } .ds-button-outline:hover { background:oklch(var(--muted)); } .ds-button-outline:focus-visible { outline:3px solid oklch(var(--ring)/0.25); outline-offset:2px; } .ds-button-outline:active { transform:scale(0.97); box-shadow:none; }"
},
{
"name": "Status Badge",
"kind": "chip",
"refersTo": "badge-primary",
"description": "A compact state label that always carries readable text.",
"html": "<span class=\"ds-status-badge\">Ready for review</span>",
"css": ".ds-status-badge { display:inline-flex; align-items:center; height:20px; padding:2px 8px; border:1px solid transparent; border-radius:6px; background:oklch(var(--primary)); color:oklch(var(--primary-foreground)); font:600 12px/1.25 var(--font-sans); white-space:nowrap; } .ds-status-badge:focus-visible { outline:3px solid oklch(var(--ring)/0.5); outline-offset:2px; }"
},
{
"name": "Text Field",
"kind": "input",
"refersTo": "input-default",
"description": "A readable operational field with an explicit focus state.",
"html": "<input class=\"ds-text-field\" value=\"Fascia pediatrica\" aria-label=\"Session name\">",
"css": ".ds-text-field { width:280px; height:40px; padding:0 12px; border:1px solid oklch(var(--input)); border-radius:8px; background:oklch(var(--background)); color:oklch(var(--foreground)); font:400 14px/1.5 var(--font-sans); outline:none; } .ds-text-field:hover { border-color:oklch(var(--muted-foreground)/0.65); } .ds-text-field:focus-visible { border-color:oklch(var(--ring)); box-shadow:0 0 0 3px oklch(var(--ring)/0.25); } .ds-text-field:disabled { opacity:0.5; cursor:not-allowed; }"
},
{
"name": "Work Card",
"kind": "card",
"refersTo": "card-default",
"description": "A single-level container for a coherent review surface.",
"html": "<section class=\"ds-work-card\"><h3>Schema linking</h3><p>Review the linked tables and columns before continuing.</p></section>",
"css": ".ds-work-card { width:320px; padding:16px; border:1px solid oklch(var(--border)/0.7); border-radius:12px; background:oklch(var(--card)); color:oklch(var(--card-foreground)); box-shadow:var(--shadow-sm); } .ds-work-card h3 { margin:0 0 8px; font:500 16px/1.35 var(--font-heading); letter-spacing:-0.01em; } .ds-work-card p { margin:0; color:oklch(var(--muted-foreground)); font:400 14px/1.6 var(--font-sans); } .ds-work-card:focus-within { outline:3px solid oklch(var(--ring)/0.25); outline-offset:2px; }"
},
{
"name": "Session Navigation Item",
"kind": "nav",
"description": "A dense session row with restrained hover and active hierarchy.",
"html": "<button class=\"ds-session-item\"><span class=\"ds-session-dot\"></span><span><strong>Patient cohorts</strong><small>Schema linking</small></span></button>",
"css": ".ds-session-item { display:flex; width:260px; align-items:center; gap:8px; padding:4px 8px; border:0; border-radius:8px; background:transparent; color:oklch(var(--foreground)); text-align:left; font-family:var(--font-sans); transition:background-color 140ms cubic-bezier(0.22,1,0.36,1); } .ds-session-item:hover,.ds-session-item[aria-current=\"page\"] { background:oklch(var(--accent)); } .ds-session-item:focus-visible { outline:2px solid oklch(var(--ring)/0.4); outline-offset:1px; } .ds-session-dot { width:6px; height:6px; flex:none; border-radius:9999px; background:oklch(var(--success)); } .ds-session-item strong,.ds-session-item small { display:block; } .ds-session-item strong { font-size:13px; font-weight:600; } .ds-session-item small { margin-top:2px; color:oklch(var(--muted-foreground)); font-size:11px; }"
},
{
"name": "Curated Evidence Document",
"kind": "custom",
"description": "The table-free reading hierarchy for persisted evidence.",
"html": "<article class=\"ds-evidence\"><h2>Fascia pediatrica</h2><div class=\"ds-evidence-summary\"><strong>Dominio</strong> · Italiano<br><span>Scopi: Disambiguazione · Generazione SQL</span></div><h3>Ambito di applicazione</h3><ul><li>fascia pediatrica</li><li>paziente minore</li></ul><h3>Regola</h3><p>La fascia pediatrica comprende i pazienti con età inferiore a 18 anni.</p><details><summary>Dettagli tecnici e provenienza</summary><code>evidence:fascia-pediatrica</code></details></article>",
"css": ".ds-evidence { max-width:70ch; color:oklch(var(--foreground)); font:400 15px/1.65 var(--font-sans); } .ds-evidence h2,.ds-evidence h3 { font-family:var(--font-heading); letter-spacing:-0.01em; } .ds-evidence h2 { margin:0 0 16px; font-size:24px; } .ds-evidence h3 { margin:24px 0 8px; font-size:18px; } .ds-evidence-summary { padding:12px 14px; border:1px solid oklch(var(--border)); border-radius:8px; background:oklch(var(--muted)); color:oklch(var(--muted-foreground)); } .ds-evidence-summary strong { color:oklch(var(--foreground)); } .ds-evidence ul { padding-left:20px; } .ds-evidence details { margin-top:24px; padding:10px 12px; border:1px solid oklch(var(--border)); border-radius:8px; background:oklch(var(--card)); } .ds-evidence summary { cursor:pointer; font-weight:600; } .ds-evidence code { font-family:var(--font-mono); }"
}
],
"narrative": {
"northStar": "The Clinical Workbench",
"overview": "ThothII should feel like a well-kept clinical workbench: warm enough for sustained reading, exact enough for consequential review, and quiet enough that evidence, state, and decisions remain in the foreground. The visual system is calm, precise, and trustworthy. It uses familiar product patterns, restrained color, and deliberate density instead of decorative spectacle.\n\nThe primary physical scene is an analyst reviewing persisted evidence and SQL on a large monitor in a well-lit working environment. This makes the warm light theme the default. The supported dark theme serves lower-light work without becoming a separate neon aesthetic. Both themes preserve the same hierarchy and semantic roles.\n\nThe system rejects generic SaaS ornament, conspicuous ripples, bounce or elastic motion, long choreographed transitions, and effects that compete with the analytical task. Controls should feel disciplined and tactile, never playful, sluggish, or visually unstable.",
"keyCharacteristics": [
"Warm, restrained surfaces with one scarce red accent.",
"Editorial headings paired with highly legible operational body text.",
"Dense information organized through hierarchy, rhythm, and progressive disclosure.",
"Persisted artifacts and reviewer decisions presented as the visual source of truth.",
"Fast state feedback with reduced-motion parity."
],
"rules": [
{"name": "The Workbench Rule", "body": "Every visual element must support inspection, action, state, or provenance. Decoration without an operational purpose is forbidden.", "section": "overview"},
{"name": "The Persisted Truth Rule", "body": "Persisted artifacts and reviewer decisions receive stronger hierarchy than transient model narration.", "section": "overview"},
{"name": "The Density with Rhythm Rule", "body": "Preserve information density, but vary spacing between groups so users can scan structure without adding nested containers.", "section": "overview"},
{"name": "The One Voice Rule", "body": "Instrument Red should occupy no more than roughly ten percent of a screen. Its rarity is what makes it authoritative.", "section": "colors"},
{"name": "The State Has a Name Rule", "body": "Success, warning, information, and destructive colors are reserved for their named states. Color is never the only state indicator.", "section": "colors"},
{"name": "The Three Registers Rule", "body": "Serif means authority, sans means interaction and reading, mono means machine identity. Do not exchange these roles for novelty.", "section": "typography"},
{"name": "The Read Once Rule", "body": "A heading, label, and body must be distinguishable on first glance through size and weight. Do not repeat headings in explanatory copy.", "section": "typography"},
{"name": "The Flat by Default Rule", "body": "A resting surface has no shadow unless it is physically above another surface. If every panel floats, none of them has hierarchy.", "section": "elevation"},
{"name": "The Borders Structure, Shadows Elevate Rule", "body": "Never use shadow as a substitute for grouping or a border as a decorative accent.", "section": "elevation"},
{"name": "The Review Surface Rule", "body": "The visible Markdown must be readable without understanding the machine contract. Technical metadata belongs in progressive disclosure, not above the title.", "section": "components"}
],
"dos": [
"Do make every state change unmistakable without interrupting flow.",
"Do use Instrument Red only for primary action, current selection, focus identity, or explicit destructive meaning.",
"Do preserve information density with headings, rhythm, and progressive disclosure.",
"Do keep keyboard focus explicit and pair color with text, shape, icon, or position.",
"Do respect prefers-reduced-motion while preserving immediate non-kinetic feedback.",
"Do use English for interface chrome and the workspace language for persisted document content.",
"Do render curated metadata and scope as Markdown prose or lists, never as a frontmatter table."
],
"donts": [
"Don't add generic SaaS ornament, conspicuous ripples, bounce or elastic motion, long choreographed transitions, or effects that compete with the analytical task.",
"Don't make controls feel playful, sluggish, or visually unstable.",
"Don't use gradient text, decorative glassmorphism, or full-saturation accents on inactive states.",
"Don't use a colored side stripe greater than one pixel on cards, callouts, list items, or blockquotes. Use a full border, tonal background, icon, or heading instead.",
"Don't nest cards or wrap every section in a container.",
"Don't use a modal before exhausting inline or progressive alternatives.",
"Don't use tables for applies_to, metadata, enum values, or other one-dimensional content.",
"Don't use color as the sole carrier of success, warning, error, selection, or progress.",
"Don't use display typography for buttons, labels, or data.",
"Don't add em dashes to interface copy. Use commas, colons, semicolons, or parentheses."
]
}
}
-7
View File
@@ -1,7 +0,0 @@
{
"$schema": "https://app.kilo.ai/config.json",
"indexing": {
"vectorStore": "qdrant",
"model": "sentence-transformers/all-minilm-l12-v2"
}
}
@@ -1,105 +0,0 @@
# Task 3 — Diagnostic contract remediation report
Date: 2026-08-04
## Scope
This remediation is limited to the four approved review findings for the workspace diagnostic
extension. It does not add registry routes, change workspace publication, alter session startup,
or expand transport support.
## Changes
1. `RuntimeBindings` now has an explicit `vectorWriter` binding. The new
`resolveRuntimeBindings()` resolves DWH, vector reader, vector writer, and embedding bindings
together. The diagnoser takes the writer credential only from `bindings.vectorWriter`, never
from vector-reader values.
2. Direct PostgreSQL and SSH-tunnelled direct probes accept an absent CA binding while retaining
certificate verification through the runtime system trust store. A supplied CA still uses
verified private-CA trust. REST private-CA refusal is unchanged.
3. A reversible vector probe now requires an authenticated POST declaration with a response map
containing `operation`. The adapter requires the successful JSON response to echo `create` or
`remove` respectively, so an arbitrary 2xx or an upsert-only response cannot activate the
write probe.
4. For DWH and vector REST diagnostics declared with `auth: none`, the resolver no longer
requires an API-key file and the adapter sends no credential. Credential-backed diagnostics
continue to require their local secret file.
## TDD evidence
The first focused RED run failed for the intended missing behavior:
- `resolveRuntimeBindings is not a function` for unauthenticated resolver bindings;
- schema accepted a reversible probe without a response contract; and
- existing diagnostic fixtures rejected the new `response` declaration until schema support was
implemented.
The focused GREEN run passed `43/43` tests across:
- `test/workspaces-bindings.test.ts`
- `test/workspaces-schema.test.ts`
- `test/workspaces-diagnostics.test.ts`
The regression coverage includes resolver-to-diagnoser writer propagation without manually
inserting the writer key into vector-reader bindings, no-CA direct/SSH system-trust requests,
operation-echo validation for create/remove, and `auth: none` bindings without secret files.
## Documentation and design
- `docs/workspace-diagnostic-protocol.md` now documents the verified system-trust fallback,
no-secret `auth: none` behavior, and required reversible response contract.
- `docs/superpowers/specs/2026-08-03-git-workspace-registry-design.md` now records the same
response, CA, SSH, and authentication rules.
## Final verification
The initial sandboxed full suite could not bind its local SSE listener (`listen EPERM:
operation not permitted 127.0.0.1`). It was rerun unchanged with local-listener permission.
```text
backend: npx vitest run
31 test files passed; 329 tests passed
backend: npx tsc --noEmit -p .
exit 0
repository: git diff --check
exit 0
```
Expected test harness stderr from existing Pi/process failure-path tests remained present; no test
failed and no diagnostic secret was emitted.
## Blockers
None.
## Round 2 remediation
The final review found two remaining contract gaps. The binding resolver already treated
`auth: none` as credential-free, but the runtime renderer and diagnostic connector still required
the API-key file. Rendering and connector construction now make that requirement conditional on
the declared REST authentication mode, so a DWH/vector `auth: none` workspace passes resolver,
runtime rendering, and diagnostics with no API-key file.
SSH forwarding previously changed the PostgreSQL connection host to `127.0.0.1` without retaining
the original target for TLS hostname validation. Forwarded probes now carry `SSH_TARGET_HOST` as
`tlsServername` into the PostgreSQL TLS options; private CA and verified system trust behavior are
unchanged.
TDD RED: the new end-to-end no-key test failed at the unconditional runtime
`API_KEY_FILE` requirement, while the SSH test showed no `tlsServername` on the loopback probe or
database-client request. TDD GREEN: the focused backend workspace tests passed `40/40`.
Round 2 final verification:
```text
backend: npx vitest run
31 test files passed; 332 tests passed
backend: npx tsc --noEmit -p .
exit 0
repository: git diff --check
exit 0
```
@@ -1,101 +0,0 @@
# Task 7 report — revision-pinned sessions
## Delivered
- New-session requests may carry `workspaceId`, provider, model, and thinking. The backend
resolves the active operational registry revision, enforces its LLM policy, and persists the
workspace ID/revision with the selected LLM settings.
- The harness manifest and `tht session new` support the optional, backward-compatible
`workspace_id` and `workspace_revision` fields.
- Resume resolves the manifest's retained snapshot, including after later registry publication.
A missing retained revision returns a sanitized `workspace_revision_unavailable` response.
Legacy manifests retain the prior workspace behavior and are marked with a visible warning on
`GET /sessions/:id`.
- `/settings` is now a non-mutating compatibility endpoint: installation defaults remain
readable, while anonymous workspace/provider/model/thinking selections are no longer written
to backend settings or principal preferences.
## TDD evidence
- RED: `npx vitest run test/routes-sessions.test.ts test/routes-settings.test.ts` failed for the
new immutable-snapshot and no-settings-mutation assertions; the manifest test failed because
`new_session_manifest` did not accept workspace revision fields.
- GREEN: `npx vitest run test/tht-runner.test.ts test/routes-sessions.test.ts test/routes-settings.test.ts && npx tsc --noEmit -p .`
completed with 97 passing tests and a clean type check.
- GREEN: `THT_HOME=/private/tmp/thothii-task7-home .venv/bin/pytest tests/test_session_documents.py tests/test_session_mutations.py -q`
completed with 22 passing tests.
- `git diff --check` completed cleanly.
## Review fixes — round 3
- The active registry snapshot that located a session now remains the authorization and mutation
config for response, steer, events, close/delete, archive/group/rename, documents, and detail.
A pruned historical revision cannot block an already-located session's active lifecycle.
- Only Resume resolves the retained pinned descriptor because Pi needs that immutable config to
restart safely. A pruned pin therefore returns the existing sanitized
`workspace_revision_unavailable` 409 solely for Resume.
### Round 3 verification
- RED: with a manifest found through an active registry snapshot and `readPinned` forced to fail,
`POST /sessions/:id/response` returned 409 instead of forwarding the active gate response.
- GREEN: `npx vitest run test/routes-sessions.test.ts test/tht-runner.test.ts test/routes-settings.test.ts && npx tsc --noEmit -p .`
— 102 tests passed with a clean type check. The regression confirms response, close, and delete
use the locating snapshot without calling `readPinned`, while Resume returns a sanitized 409.
- `git diff --check` completed cleanly.
## Review fixes — round 2
- Lifecycle authorization no longer selects the installation-default workspace. The backend now
finds each session by querying every operational registry snapshot with the authenticated
principal, preserving RLS ownership concealment.
- After locating the manifest, durable pinned sessions resolve their retained descriptor before
any lifecycle mutation/reopen. Legacy sessions continue using the locating registry snapshot.
- Session listing aggregates the owner-visible rows from all operational registry snapshots;
detail, response, steer, resume, events, documents, and lifecycle mutations use the same
server-side locator. No route depends on browser-local workspace state.
### Round 2 verification
- RED: the new cross-workspace route integration test created a B session while installation
default A was selected, then demonstrated that `GET /sessions` returned an empty list.
- GREEN: `npx vitest run test/routes-sessions.test.ts test/tht-runner.test.ts test/routes-settings.test.ts && npx tsc --noEmit -p .`
— 101 tests passed with a clean type check. The integration test covers create B, list, detail,
response, and resume through B's pinned descriptor while default A remains configured.
- Full backend suite: 342 tests passed. The remaining 7 tests require binding `127.0.0.1` and
fail in this sandbox with `listen EPERM: operation not permitted`; no application assertion
failed. The focused typecheck above passed.
- `git diff --check` completed cleanly.
## Verification note
The unscoped backend suite was also run. The Task 7 code regressions in `test/tht-runner.test.ts`
were fixed; the remaining failures were existing sandbox restrictions on tests that listen on
`127.0.0.1` (`listen EPERM: operation not permitted` in SSE/e2e health tests), not application
assertions.
## Review fixes — round 1
- Every new session now resolves `workspaceId` through the registry; an omitted value uses the
configured installation default and persists both the resolved ID and revision. Callers cannot
bypass revision pinning by supplying a workspace ID.
- Browser-local preferences now migrate once from the read-only legacy settings response and hold
workspace, provider, model, and thinking. Session creation includes those selections, including
direct entry points that run before the composer mounts. The frontend no longer `PUT`s shared
settings.
- The settings compatibility endpoint honors a stored installation workspace before falling back
to the first workspace configuration.
- Resume rejects finalized and archived sessions before looking up any pinned snapshot, preserving
the read-only response even when a historical snapshot is unavailable.
### Review verification
- RED: the added backend tests failed for omitted-default pinning, read-only resume ordering, and
stored-default precedence; the added frontend preference tests failed because preferences were
neither stored nor included in session requests.
- GREEN: `npx vitest run test/tht-runner.test.ts test/routes-sessions.test.ts test/routes-settings.test.ts && npx tsc --noEmit -p .`
— 100 tests passed with a clean type check.
- GREEN: `npx vitest run && npx tsc -b` — 332 frontend tests passed with a clean type check.
- GREEN: `THT_HOME=/private/tmp/thothii-task7-home .venv/bin/pytest tests/test_session_documents.py tests/test_session_mutations.py -q`
— 22 tests passed (one existing testcontainers deprecation warning).
- `git diff --check` completed cleanly.
@@ -1,73 +0,0 @@
# Task 9 report — Workspace Management CRUD page
## Delivered
- Added the Workspace management dialog, launched from the persistent right sidebar and the
Model activity header without touching live-session/SSE state.
- Added a workspace list/detail editor for General, DWH, Semantic index, LLM policy,
Installation requirements, and Git status/history.
- Added browser-only New, Edit, Duplicate, Save draft, and Delete-draft workflows. A deletion
draft stores only ID and immutable revision references; publication remains a Task 10 action.
- Used closed native controls for languages, engines, transports, distance metrics, embedding
providers, and selectable default models. Free values have client-side, accessible errors.
- Made semantic-index dimensions atomic: one editor field always writes the same value to the
vector-store and embedding contracts.
- Added Validate and Test-on-this-installation actions. They display sanitized code/message
diagnostics only; neither action exposes or stores credentials, secrets, or raw response bodies.
- Explicitly excluded publish, pull, import, and export user flows from this task.
## TDD evidence
- RED: `npx vitest run src/shell/WorkspaceManager.test.tsx src/shell/WorkspaceEditor.test.tsx`
failed because the manager and editor modules did not exist.
- GREEN: focused manager/editor/AppShell coverage passed after the implementation.
- RED: a deletion-draft persistence regression failed with
`Cannot read properties of undefined (reading 'save')` before the sanitized draft store was added.
- GREEN: the draft-store and manager tests passed once deletion intent persisted locally.
## Verification
Executed from `frontend/`:
```text
npx vitest run
50 test files passed, 358 tests passed
npx tsc -b
exit 0
```
`git diff --check` passed before commit. No workspace secret value, secret-file path, raw
diagnostic body, publish call, import flow, or export flow was introduced.
## Fix round 1
### Root causes and fixes
- The original duplicate proposal appended `-copy` and then truncated at 63 characters. For an
already-maximal ID, truncation could remove the suffix and reproduce the immutable source ID.
The proposal now reserves suffix space and falls back to a distinct `-2` suffix when a maximal
source already ends in `-copy`.
- `dwh.timeout_ms` was rendered as a positive numeric field but was absent from the client
validation map. It now has the same immediate accessible error treatment as other numeric
fields, so a rejected save never reaches the manager’s saved-draft toast.
- Registry status, workspace list, and selected-detail React Query failures were rendered as
loading, empty, or unselected states. Each now has a named alert and a retry control, distinct
from its corresponding loading and empty state.
### TDD evidence
- RED: max-length duplication retained the original 63-character ID; the timeout field produced
no alert; and each of the three failed queries had no accessible retry control.
- GREEN: the focused manager/editor tests passed **12/12**, covering a valid changed duplicate
proposal, rejected zero timeout with no save toast, and status/list/detail retry recovery.
### Verification
Executed from `frontend/`:
```text
npx vitest run
50 test files passed, 364 tests passed
npx tsc -b
exit 0
```
@@ -1,180 +0,0 @@
# Task 11 report
Status: completed on 2026-08-08.
## Scope delivered
- Updated operator-facing documentation for the internal Qdrant + Ollama architecture.
- Tightened documentation contract tests to require the current four-service-plus-init topology,
CPU-first/GPU-override guidance, fixed internal model/dimensions, schema-v3 migration wording,
one-collection-per-workspace ownership, and Qdrant backup/restore safety.
- Updated stable repo guidance in `AGENTS.md` and the current snapshot in `PROJECT_STATE.md`.
- Rewrote the workspace diagnostic protocol to the schema-v3/internal-semantic-service contract.
- Updated the memory guide to describe Qdrant as the derived persistent index.
- Updated the runtime secret-bundle guide to remove active vector/embedding secret guidance.
## Files changed
- `README.md`
- `AGENTS.md`
- `PROJECT_STATE.md`
- `docs/install/local-workspace-registry.md`
- `docs/install/server-workspace-registry.md`
- `docs/installazione-docker-4-contesti.md`
- `docs/workspace-diagnostic-protocol.md`
- `docs/gestione-memory.md`
- `deploy/secrets/README.md`
- `scripts/verify-workspace-install-docs.sh`
- `scripts/test-verify-workspace-install-docs.sh`
## Verification
Fresh successful runs:
```sh
./scripts/test-verify-workspace-install-docs.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
git diff --check
```
Key outcomes:
- internal semantic infrastructure documentation contract passed
- all existing install/manual fixture contracts still passed
- diff hygiene passed with no whitespace/errors
## Self-review notes
- The updated docs now match the code-backed Compose topology: `frontend`, `core`, `qdrant`,
`embedding`, and `embedding-model-init`.
- Active manuals no longer instruct operators to configure external vector or embedding runtime
endpoints/secrets.
- Qdrant backup/restore wording now matches the helper scripts' exact confirmation and rollback
behavior.
- Legacy descriptor handling is documented as explicit schema-v3 migration only; no silent
semantic-data migration is claimed.
## Residual concerns
- The broader repository still contains historical design/spec material that references older
pgvector/external-embedding architecture; this task intentionally updated operator/current-state
documentation and the corresponding contract tests, not historical planning documents.
## Fix round 1/5 — 2026-08-08
Addressed reviewer findings:
- Moved superseded rollout/state blocks in `PROJECT_STATE.md` behind an explicit
`## Historical snapshots and archived reference notes` boundary.
- Renamed superseded snapshot headings so historical notes no longer present as active `LIVE`
state.
- Added a current-state regression that rejects contradictory active blocks (for example:
schema-v2 operational, two-service active stack, or external vector/embedding runtime claims
before the historical boundary).
- Refactored new internal-semantic doc checks away from exact-sentence coupling:
- parse `compose.yaml` structurally with YAML;
- parse workspace examples structurally with YAML;
- inspect backup/restore stable usage interface;
- keep targeted forbidden-term checks for active docs while allowing historical sections;
- use regex/concept checks for prose.
Evidence:
```sh
./scripts/test-verify-workspace-install-docs.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
git diff --check
```
Observed RED before the fix:
```text
PROJECT_STATE.md: missing Historical snapshots boundary
```
## Fix round 2/5 — 2026-08-08
Addressed reviewer findings:
- Renamed every historical `PROJECT_STATE.md` heading after the historical boundary so no heading
level uses `LIVE` or current-state semantics there.
- Strengthened the historical-boundary regression to reject any Markdown heading level
(`#` through `######`) containing `LIVE` or current-state wording after the boundary.
- Added a fixture with a `### ... — LIVE ...` historical heading to prove RED then GREEN.
- Replaced remaining exact phrase checks with concept/semantic validation for:
- one-workspace/one-collection ownership;
- external boundary (DWH/LLM external; vector/embedding internal);
- the Italian compact install note.
- Added paraphrase fixtures that pass and omission/inversion fixtures that fail.
Evidence:
```sh
./scripts/test-verify-workspace-install-docs.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
git diff --check
```
## Fix round 4/5 — 2026-08-08
Addressed reviewer finding:
- Eliminated semantic-index verifier/test contract drift by extracting the production
semantic-index ownership row matcher into `semantic_index_relationship_spec` and reusing it in
the fixture-level paraphrase, omission, and scattered-token checks.
- Kept the relationship constrained to one structured Markdown table row via
`verify_markdown_table_relationships`; the scattered-token fixture still removes the row and
appends the same words outside the table, where it must be rejected.
- Added a direct regression that copies the repository docs into an isolated root, applies the
accepted paraphrase “A workspace keeps exactly one Qdrant collection reserved for itself”, and
runs that root's actual `scripts/verify-workspace-install-docs.sh --fixtures-only` instead of a
separate temporary spec.
Observed RED before the fix:
```text
production verifier rejected the accepted semantic-index paraphrase
local workspace manual: missing relationship in 'Semantic index ownership contract': {'scope': 'workspace semantic index', 'ownership rule': '(each|one|single).*(workspace).*(single|one).*(Qdrant).*(collection)|(each workspace reserves a single qdrant collection)', 'isolation rule': 'schema.*evidence.*memory.*(one|that).*(collection).*(kind|payload)'}
```
Evidence:
```sh
./scripts/test-verify-workspace-install-docs.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
git diff --check
```
Observed RED during this round:
```text
PROJECT_STATE.md: historical section still contains active/live heading markers
compact manual paraphrase lacks required pattern: (esterni solo|solo esterni|restano esterni)
```
## Fix round 3/5 — 2026-08-08
Addressed reviewer findings:
- Added table-driven historical-heading fixtures for every Markdown heading level `#` through
`######`; all are rejected after the historical boundary when they contain `LIVE`/current-state
semantics.
- Added small structured ownership tables to the active local/server manuals and to the compact
Italian operator note.
- Added small structured semantic-index ownership tables to the active local/server manuals.
- Replaced the remaining scattered-token relationship checks with explicit structured-section
parsing:
- architecture ownership rows map DWH → external, LLM → external, Qdrant → internal,
Ollama embedding → internal;
- semantic-index ownership rows localize the one-workspace/one-collection contract and the
schema/Evidence/Memory isolation rule.
- Added adversarial fixtures that fail when the same tokens are merely scattered in free text.
- Added structured paraphrase fixtures that pass and omission/inversion fixtures that fail.
Evidence:
```sh
./scripts/test-verify-workspace-install-docs.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
git diff --check
```
@@ -1,43 +0,0 @@
# Task 12 Report — Remove unreachable pgvector runtime code
Status: completed
Summary:
- Proved the retired pgvector runtime had no remaining operational adapter call sites after migration by re-running the required grep; only the packaging assertion still mentions `migrations/vector`.
- Removed the obsolete pgvector/HTTP/direct vector runtime modules, vector SQL migrations, and their affected runtime tests.
- Kept the operational semantic path on Qdrant and migrated the remaining runtime callers to that path.
- Kept `psycopg2-binary` because DWH direct PostgreSQL and session PostgreSQL code still depend on it.
Implementation notes:
- Extracted shared collection/kind validation into `harness/tht/adapters/vector/_shared.py` so `QdrantVectorStore` no longer depends on the deleted pgvector module.
- Simplified `build_vector_store()` to return only `QdrantVectorStore`.
- Migrated vector/evidence/memory CLI paths away from legacy pgvector loaders and REST vector clients.
- Updated packaging coverage so the built wheel asserts session SQL migrations are present and vector SQL migrations are absent.
Verification:
- `cd harness && .venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py tests/test_semantic_kind_isolation.py tests/test_vector_migration_packaging.py -q`
- `cd harness && .venv/bin/pytest tests/test_adapter_factory.py tests/test_solved_search_cli.py -q`
- `cd harness && .venv/bin/python -c "import tht.cli, tht.adapters.factory, tht.adapters.vector, tht.vectorstore.reader"`
- `cd harness && uv build`
- `harness/.venv/bin/ruff check harness/tests/test_adapter_factory.py harness/tests/test_solved_search_cli.py harness/tests/test_vector_migration_packaging.py harness/tests/test_vector_port_contract.py harness/tht/adapters/factory.py harness/tht/adapters/vector/__init__.py harness/tht/adapters/vector/_shared.py harness/tht/adapters/vector/qdrant.py harness/tht/cli/evidence_cmd.py harness/tht/cli/memory_cmd.py harness/tht/cli/search_cmd.py harness/tht/cli/vector_cmd.py harness/tht/solved.py harness/tht/vectorstore/reader.py`
- `git diff --check`
Notes / concerns:
- Repository-wide `harness/.venv/bin/ruff check .` still reports many pre-existing findings outside this task’s touched files; it is not clean on this branch baseline.
- Some legacy config compatibility parsing still exists outside the deleted runtime path. This task removed the unreachable runtime/migration code without broad config-schema refactoring.
## Fix round 1 evidence
Changes:
- Removed dead `vector migrate` registration from `harness/tht/cli/__init__.py` and deleted `harness/tht/cli/vector_migrate_cmd.py`.
- Added CLI regressions proving `vector migrate` is absent while `vector init` and `vector index-schema` remain available.
- Restored the accidentally removed non-vector regressions by moving report coverage into `harness/tests/test_report.py` and restoring the taskdoc promoted-table slicing check in `harness/tests/test_taskdoc.py`.
- Reworded surviving active help/docstrings away from pgvector-specific wording in the touched Qdrant-backed command surface.
Verification:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py tests/test_report.py tests/test_taskdoc.py tests/test_vector_migration_packaging.py -q`
- `cd harness && .venv/bin/python -c "from typer.testing import CliRunner; from tht.cli import app; r=CliRunner().invoke(app, ['vector','--help']); assert r.exit_code == 0, r.output; assert 'migrate' not in r.output; r=CliRunner().invoke(app, ['vector','migrate','--help']); assert r.exit_code != 0, r.output; print('cli-help-ok')"`
- `cd harness && .venv/bin/python -c "import tht.cli, tht.cli.vector_cmd, tht.report, tht.taskdoc; print('imports-ok')"`
- `cd harness && uv build`
- `harness/.venv/bin/ruff check harness/tests/test_qdrant_cli_commands.py harness/tests/test_report.py harness/tests/test_taskdoc.py harness/tests/test_vector_migration_packaging.py harness/tht/cli/__init__.py harness/tht/cli/search_cmd.py harness/tht/cli/vector_cmd.py harness/tht/cli/memory_cmd.py harness/tht/solved.py`
- `git diff --check`
@@ -1,175 +0,0 @@
# Task 13 Implementation Report
## Status
DONE_WITH_CONCERNS
## Changes
- Updated stale harness/backend/frontend tests and fixtures to the Task 13 internal Qdrant/Ollama contract.
- Made `deploy/workspaces/psd.yaml.example` generic while preserving schema-v3 Qdrant/Ollama shape.
- Fixed `scripts/workspace-registry-smoke.sh` to pass the required legacy migration `--collection` and prove exact Docker cleanup, including its smoke image.
- Updated `PROJECT_STATE.md` with only evidence observed in this run.
Changed files:
- `PROJECT_STATE.md`
- `backend/test/routes-workspaces.test.ts`
- `backend/test/workspace-runtime-handoff.test.ts`
- `backend/test/workspaces-contracts.test.ts`
- `backend/test/workspaces-git-repository.test.ts`
- `deploy/workspaces/psd.yaml.example`
- `frontend/src/shell/NewSessionDialog.test.tsx`
- `harness/tests/test_adapter_command_regressions.py`
- `harness/tests/test_workspace.py`
- `scripts/task13-runtime-fixture-check.ts`
- `scripts/test-verify-workspace-install-docs.sh`
- `scripts/workspace-registry-smoke.sh`
## Verification
Deterministic gates:
- `cd harness && .venv/bin/pytest -q && .venv/bin/ruff check .`
- Initial red: 2 harness pytest failures.
- After fixture fixes: harness pytest passed `819 passed, 4 deselected, 74 warnings in 27.73s`.
- Ruff still failed with `Found 220 errors`; treated as existing unrelated debt.
- Touched harness files verified clean with `cd harness && .venv/bin/ruff check tests/test_adapter_command_regressions.py tests/test_workspace.py && .venv/bin/pytest -q tests/test_adapter_command_regressions.py::test_solved_index_writes_through_writer_only_factory_store tests/test_workspace.py::test_load_workspace_expands_env_vars`: `All checks passed!` and `2 passed, 2 warnings in 0.14s`.
- `cd backend && npx vitest run && npx tsc --noEmit -p . && npm run build`
- Initial red: 4 backend Vitest failures.
- After fixes: `Test Files 39 passed (39)`, `Tests 464 passed (464)`, TypeScript passed, build passed.
- `cd frontend && npx vitest run && npx tsc -b && npm run build`
- Initial red: 1 frontend Vitest failure.
- After fix: frontend Vitest passed `374/374`, TypeScript passed, build passed with Vite `built in 6.55s`.
- `git diff --check`
- Passed with no output.
Focused reruns:
- `cd backend && npx vitest run test/workspaces-migrate-legacy.test.ts test/workspaces-contracts.test.ts test/routes-workspaces.test.ts test/workspace-runtime-handoff.test.ts test/workspaces-git-repository.test.ts && cd .. && ./scripts/test-no-deployment-coupling.sh && ./scripts/verify-workspace-install-docs.sh --fixtures-only && git diff --check`
- `Test Files 5 passed (5)`, `Tests 35 passed (35)`.
- Coupling guard passed: `no active retired deployment or external semantic coupling found.`
- Install docs fixtures passed through `relative secret-source fixture rejected passed`.
Deployment contracts:
- `./scripts/test-default-compose.sh && ./scripts/test-unified-compose.sh && ./scripts/test-internal-semantic-compose.sh && ./scripts/test-no-deployment-coupling.sh && ./scripts/test-compose-secret-policy.sh && ./scripts/verify-workspace-install-docs.sh --fixtures-only`
- Passed. Output included:
- `default Compose contract passed.`
- `unified Compose contract passed.`
- `internal semantic Compose/script contracts passed.`
- `no active retired deployment or external semantic coupling found.`
- `Compose secret policy passed.`
- install-doc fixture checks through `relative secret-source fixture rejected passed`.
Docker smokes:
- `/usr/bin/time -p ./scripts/internal-semantic-smoke.sh`
- Passed: `Task 13 internal semantic smoke passed.`
- Cleanup proof: `no labeled containers, volumes, networks, or images remain for 20260808200245-83368-17823.`
- Duration: `real 217.34`.
- `/usr/bin/time -p ./scripts/workspace-registry-smoke.sh`
- Initial red: `usage: migrate-legacy --input <legacy-workspace.yaml> --output <repository-root> --collection <qdrant-collection> [--id <workspace-id>]`.
- After fix: `workspace registry smoke passed`.
- Cleanup proof: `no compose containers, volumes, networks, or image remain for thoth-workspace-registry-smoke-89671.`
- Duration: `real 9.93`.
- `/usr/bin/time -p ./scripts/unified-deployment-smoke.sh`
- Passed: `Task 13 full deployment smoke passed.`
- Cleanup proof: `no labeled containers, volumes, networks, or images remain for 20260808200706-85638-13391.`
- Duration: `real 125.57`.
- `/usr/bin/time -p ./scripts/thothctl-update-smoke.sh`
- Passed: `Task 13 update deployment smoke passed.`
- Cleanup proof: `no labeled containers, volumes, networks, or images remain for 20260808200918-87340-10404.`
- Duration: `real 85.40`.
- `/usr/bin/time -p ./scripts/server-deployment-smoke.sh`
- Passed: `Task 13 Linux server deployment smoke passed.`
- Cleanup proof: `no labeled containers, volumes, networks, or images remain for 20260808201047-88645-20675.`
- Duration: `real 55.99`.
Final audit:
- `rg -n "pgvector|local-vector|THT_VECTOR_|EMBEDDING_BASE_URL|openai_compatible|ollama_compatible" . --glob '!docs/plans/**' --glob '!docs/superpowers/**' --glob '!**/node_modules/**' --glob '!**/.venv/**' --glob '!**/.git/**'`
- Returned matches in legacy schema-v1/v2 support, migration tests, negative guards, historical notes, and older harness docs/code.
- This remains a concern: the audit is not clean under the brief's strict expected outcome.
- `git status --short`
- Before report/commit, contained only intentional Task 13 changes.
## Image and Host Evidence
- Host CPU: `Apple M4 Pro`.
- Host OS: `Darwin MacProM4-di-Marco.local 25.5.0 Darwin Kernel Version 25.5.0: Tue Jun 9 22:28:34 PDT 2026; root:xnu-12377.121.10~1/RELEASE_ARM64_T6041 arm64`.
- Docker server: `29.6.2 linux/arm64`.
- Verified pinned images:
- `qdrant/qdrant:v1.18.2@sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c`.
- `ollama/ollama:0.32.0@sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a`.
- Workspace registry smoke ephemeral image:
- Manifest list: `sha256:4d056bf2cb38d0e8ede91fbf121df1f9f18caee0d401581618ccef9ed8a55e73`.
- Config: `sha256:613f8fb28c0517adee4085f41bc447f2c3813b0fdbb7b26624bfb4cb192b6fd8`.
- Removed during cleanup.
## Manual Gates
- GPU exposure gate (`THOTH_ENABLE_EMBEDDING_GPU=1` on Linux): not executed in this run.
- Windows Docker Desktop startup/manual job: not executed in this run.
## Commits
- `4e810af` (`test: align qdrant ollama verification fixtures`)
- `7c09b98` (`docs: record qdrant ollama verification`)
## Known Limitations
- Broad harness Ruff remains existing unrelated debt: `Found 220 errors`.
- Final active-reference audit is not clean; it still finds legacy/negative-guard references outside explicit migration fixture files.
- Ephemeral Task 13 core/frontend image IDs from `internal-semantic-smoke.sh`, `unified-deployment-smoke.sh`, `thothctl-update-smoke.sh`, and `server-deployment-smoke.sh` were removed by exact cleanup and were not emitted in stdout; pinned Qdrant/Ollama digests and the workspace-registry smoke image digest were captured.
## Fix Round 1 — reviewer findings
Status: DONE
Changes:
- `scripts/workspace-registry-smoke.sh` now derives the smoke image reference from the already unique Compose project instead of using the global tag `thothii-workspace-registry-smoke:local`.
- The workspace-registry cleanup helpers remove and verify only the exact per-run image reference, plus Compose resources labeled with the exact project.
- Added deterministic self-test coverage in `backend/test/workspaces-migrate-legacy.test.ts` via `WORKSPACE_REGISTRY_SMOKE_SELF_TEST=image-cleanup-identity`; it stubs Docker and fails if cleanup touches same-repository foreign tags such as `:local` or another project tag.
- Updated active harness/testing/PRD docs and Python comments that still described the current semantic store as pgvector/vectordb. Preserved schema-v1/v2 and harness legacy compatibility fixtures.
- Updated `PROJECT_STATE.md` with fix-round smoke evidence and a precise, non-overclaiming audit limitation.
Focused verification:
- `cd backend && npx vitest run test/workspaces-migrate-legacy.test.ts`
- Passed: `7 passed`.
- `cd harness && .venv/bin/pytest -q tests/test_memory_save_one.py tests/test_adapter_command_regressions.py tests/test_solved_search_cli.py tests/test_search_pack.py`
- Passed: `22 passed, 14 warnings`.
- `cd harness && .venv/bin/ruff check tht/memory.py tht/search/__init__.py tht/workspace.py tht/vectorstore/store.py tests/test_memory_save_one.py tests/test_adapter_command_regressions.py tests/test_solved_search_cli.py`
- Passed: `All checks passed!`
- `bash -n scripts/workspace-registry-smoke.sh && WORKSPACE_REGISTRY_SMOKE_SELF_TEST=image-cleanup-identity bash scripts/workspace-registry-smoke.sh`
- Passed: `workspace registry smoke image cleanup identity self-test passed`.
- `./scripts/test-no-deployment-coupling.sh`
- Passed: `no active retired deployment or external semantic coupling found.`
- `./scripts/verify-workspace-install-docs.sh --fixtures-only`
- Passed through `relative secret-source fixture rejected passed`.
- `cd backend && npx tsc --noEmit -p .`
- Passed with no output.
- `/usr/bin/time -p ./scripts/workspace-registry-smoke.sh`
- Passed: `workspace registry smoke passed`.
- Built exact per-run tag: `thothii-workspace-registry-smoke:thoth-workspace-registry-smoke-thoth-workspace-registry-smoke-10vi3a-19157`.
- Manifest list: `sha256:715b943057929418cad4aa71806d9edbaf823555d19bda6b875297617463fd4a`.
- Config: `sha256:a566521981e08958aae9a12bfc7803bb5f3f835536b4bb8c39df8fcf26063161`.
- Cleanup proof: `no compose containers, volumes, networks, or image remain for thoth-workspace-registry-smoke-thoth-workspace-registry-smoke-10vi3a-19157.`
- Duration: `real 42.06`.
Fix-round audit command:
- `rg -n "pgvector|local-vector|THT_VECTOR_|EMBEDDING_BASE_URL|openai_compatible|ollama_compatible" . --glob '!docs/plans/**' --glob '!docs/superpowers/**' --glob '!**/node_modules/**' --glob '!**/.venv/**' --glob '!**/.git/**'`
Categorized remaining hits:
- Backend legacy parser/migration compatibility, kept deliberately non-operational for schema-v1/v2 descriptors: `backend/src/workspaces/schema.ts`, `types.ts`, `migrate-legacy.ts`, `runtime-renderer.ts`, `bindings.ts`, `contracts.ts`, `diagnostics.ts`.
- Backend negative guards and legacy fixture tests: `backend/test/workspaces-schema.test.ts`, `workspaces-migrate-v2-qdrant.test.ts`, `workspace-registry.test.ts`, `workspace-runtime-renderer.test.ts`, `workspaces-bindings.test.ts`, `workspaces-contracts.test.ts`, `workspaces-diagnostics.test.ts`, `workspaces-git-repository.test.ts`, `routes-workspaces.test.ts`, `routes-sessions.test.ts`, `provider-credentials.test.ts`.
- Secret/env scrub guards for retired variables: `backend/src/config.ts`, `backend/src/config/secret-bundle.ts`, `backend/src/pi/provider-credentials.ts`, `scripts/compose-with-preflight.sh`, `scripts/test-external-compose-lifecycle.sh`.
- Deployment negative guards and fixture-scope tests: `scripts/test-no-deployment-coupling.sh`, `scripts/test-no-deployment-coupling-scope.sh`, `scripts/test-preprocess-compose-config.sh`, `scripts/test-verify-workspace-install-docs.sh`, `scripts/verify-workspace-install-docs.sh`, `scripts/vector-rotate-bootstrap-password.sh`.
- Harness legacy config compatibility and fixtures: `harness/tht/config.py`, `harness/tht/config_compat.py`, `harness/tests/test_config_resources.py`, `harness/tests/l2/test_session_ablazione.py`, `harness/workspaces/tht.example.yaml`, `harness/workspaces/tht-test.yaml`.
- Retained off-repository migration SQL fixtures: `harness/scripts/create_vector_reader_rpc.sql`, `harness/scripts/create_vector_writer_rpc.sql`.
- Historical/reference notes, not active operator contracts: `brain/codebase/datamart-builder-deployment-gotchas.md`, `PROJECT_STATE.md`.
- Gitignored task report self-reference: `.superpowers/sdd/2026-08-08-internal-qdrant-ollama/task-13-implementation.md`.
@@ -1,165 +0,0 @@
Task 2 report — Make collection ownership unique in the Git registry
Summary
- Implemented unique Qdrant collection ownership enforcement during registry snapshot activation.
- Registry session revision leases now reject `migration_required` descriptors.
- Legacy migration now requires an explicit target collection and emits schema v3 descriptors.
- Preserved active snapshot rollback behavior on invalid pulled snapshots.
RED evidence
Focused RED command from the brief:
```bash
cd backend
npx vitest run test/workspace-registry.test.ts test/workspaces-migrate-legacy.test.ts \
-t "collection|migration_required"
```
Observed failures before implementation:
- `rejects duplicate schema v3 collection ownership and keeps the previous active snapshot`
- `registry.pull()` resolved instead of rejecting.
- `does not acquire a session revision lease for a migration_required workspace`
- `acquireSessionRevision()` resolved instead of rejecting.
- `migrates a legacy descriptor only with an explicit target collection into schema v3`
- received schema version `1` instead of `3`.
- `requires an explicit target collection for legacy migration`
- migration did not throw without a collection.
GREEN evidence
Focused GREEN command from the brief:
```bash
cd backend
npx vitest run test/workspace-registry.test.ts test/workspaces-migrate-legacy.test.ts \
-t "collection|migration_required"
```
Fresh result after implementation:
- 2 files passed
- 4 tests passed
- 0 failures
Additional verification run after final cleanup:
```bash
cd backend
npx vitest run test/routes-workspaces.test.ts
npx vitest run
npx tsc --noEmit -p .
git diff --check
```
Fresh results:
- `test/routes-workspaces.test.ts`: 7 passed
- full backend Vitest: 39 files passed, 454 tests passed
- backend typecheck: passed
- `git diff --check`: passed
Changed files
- `backend/src/workspaces/registry.ts`
- `backend/src/workspaces/migrate-legacy.ts`
- `backend/test/workspace-registry.test.ts`
- `backend/test/workspaces-migrate-legacy.test.ts`
- `backend/test/routes-workspaces.test.ts`
Why one extra file changed
- `backend/test/routes-workspaces.test.ts` needed updating because Task 1 made schema v3 the only operational descriptor shape, and the route test still assumed the old pre-Task-3 runtime behavior. Updating that expectation was necessary to keep the required backend suite verification meaningful.
Implementation notes
- Duplicate collection detection is enforced only for operational schema v3 descriptors by tracking `collection -> workspaceId` during activation.
- Duplicate failures are sanitized back to `workspace_invalid` / `Workspace repository content is invalid`.
- `acquireSessionRevision()` now fails closed for `migration_required` revisions.
- Legacy migration CLI now requires `--collection <qdrant-collection>`.
- Legacy migration output is schema v3 with the fixed internal semantic contract:
- `vector_store.engine = qdrant`
- explicit `collection`
- embedding provider `ollama_internal`
- embedding model `qwen3-embedding:0.6b`
self-review
- Confirmed invalid pulled snapshots do not replace the previous active snapshot.
- Confirmed duplicate collection enforcement does not affect legacy migration-required descriptors.
- Confirmed create/update publication tests still pass with unique per-workspace collections.
- Confirmed no JSON stdout contract regressions in the migration CLI.
- Kept runtime/data mutation scope descriptor-only; no user workspace repo or Qdrant data changes.
Concerns
- No code concerns remaining for Task 2.
- One deliberate scope exception: a route test was updated to align with the already-established Task 1 / Task 3 fail-closed contract.
Fix round 1
Scope
- Restored meaningful route-level diagnoser coverage without reopening schema-v3 semantic runtime paths.
- Added direct schema-v2 registry coverage for `migration_required` listing and lease rejection.
Covering test files
- `backend/test/routes-workspaces.test.ts`
- `backend/test/workspace-registry.test.ts`
RED command and output
Command:
```bash
cd backend
npx vitest run test/routes-workspaces.test.ts test/workspace-registry.test.ts
```
Observed result on top of `76bc94d` after adding the restored/new assertions:
- 2 files passed
- 37 tests passed
- 0 failures
Why no RED appeared:
- The review items exposed missing/weakened coverage, not a production behavior bug.
- `/workspaces/:id/test` already reaches the diagnoser for resolvable legacy v2 descriptors.
- Schema-v3 `/workspaces/:id/test` already fails closed before diagnoser entry.
- Schema-v2 descriptors were already listed as `migration_required` and already rejected by `acquireSessionRevision()`.
GREEN command and output
Command:
```bash
cd backend
npx vitest run test/routes-workspaces.test.ts test/workspace-registry.test.ts
npx tsc --noEmit -p .
```
Fresh results:
- covering tests: 2 files passed, 37 tests passed
- backend typecheck: passed
Changed files
- `backend/test/routes-workspaces.test.ts`
- `backend/test/workspace-registry.test.ts`
- `.superpowers/sdd/2026-08-08-internal-qdrant-ollama/task-2-report.md`
What changed
- Split route coverage so `POST /workspaces/validate` still checks canonical validation independently.
- Restored route-level diagnoser coverage through a migration-required schema-v2 descriptor with resolvable legacy bindings.
- Added an explicit schema-v3 fail-closed regression for `POST /workspaces/:id/test`.
- Added a direct schema-v2 registry regression proving `list()` returns `migration_required` and `acquireSessionRevision()` rejects it.
Concerns
- No production concerns. This round only tightened coverage and corrected the weakened test expectation.
@@ -1,132 +0,0 @@
# Task 3 report — Remove external semantic bindings and render internal endpoints
Date: 2026-08-08
## Scope
Implemented backend-owned schema-v3 semantic runtime rendering so workspace descriptors and installation contracts remain free of external Qdrant/Ollama endpoints and credentials, while DWH bindings stay unchanged.
## RED evidence
Focused RED command:
`cd backend && npx vitest run test/workspaces-contracts.test.ts test/workspaces-bindings.test.ts test/workspace-runtime-renderer.test.ts test/config.test.ts`
Observed failures before implementation:
- `config.test.ts`
- missing `internalQdrantUrl`
- missing `internalEmbeddingUrl`
- `workspaces-bindings.test.ts`
- schema v3 semantic binding resolution threw unsupported errors
- `workspace-runtime-renderer.test.ts`
- schema v3 runtime rendering threw `Schema version 3 runtime rendering is unsupported until the internal semantic runtime is implemented`
## GREEN evidence
Focused GREEN command:
`cd backend && npx vitest run test/workspaces-contracts.test.ts test/workspaces-bindings.test.ts test/workspace-runtime-renderer.test.ts test/config.test.ts`
Result:
- 4 test files passed
- 36 tests passed
Typecheck:
`cd backend && npx tsc --noEmit -p .`
Result:
- passed
Hygiene:
- `git diff --check` passed
## Files changed
Listed-task files changed:
- `backend/src/config.ts`
- `backend/src/workspaces/bindings.ts`
- `backend/src/workspaces/runtime-renderer.ts`
- `backend/test/config.test.ts`
- `backend/test/workspace-runtime-renderer.test.ts`
- `backend/test/workspaces-bindings.test.ts`
- `backend/test/workspaces-contracts.test.ts`
Listed-task files inspected but not changed:
- `backend/src/workspaces/contracts.ts`
Unavoidable additional wiring changes:
- `backend/src/app.ts`
- `backend/src/tht/tht-runner.ts`
Reason: the new typed internal semantic runtime config had to flow from backend config into ephemeral harness config rendering at runtime.
## Behavior delivered
- schema-v3 installation contract exposes DWH bindings only
- schema-v3 binding resolution ignores external semantic env vars instead of sourcing runtime semantics from them
- runtime rendering for schema v3 emits backend-owned internal semantic endpoints:
- Qdrant: `http://qdrant:6333`
- Embedding: `http://embedding:11434`
- Model: `qwen3-embedding:0.6b`
- Dimensions: `1024`
- internal semantic URLs are validated to allow only `qdrant` / `embedding` / `localhost` / loopback hosts
- DWH transport/runtime behavior remains unchanged
## Self-review
- Confirmed schema-v3 contracts/docs no longer advertise VECTOR or EMBEDDING installation variables.
- Confirmed schema-v3 runtime output ignores injected external semantic endpoints from env bindings.
- Confirmed semantic endpoints are rendered only in the ephemeral backend-owned harness config path.
- Confirmed type wiring is explicit from `AppConfig` → `ThtRunner` → runtime renderer.
## Concerns
- Host validation currently permits both `http` and `https` on the allowed internal hosts. That keeps the configuration flexible, but if the installation contract intended `http` only, that restriction is not enforced here.
## Fix round 1/5
Scope:
- moved schema-v3 internal embeddings under `resources.embeddings`
- enforced `http`-only internal semantic URLs
RED evidence:
`cd backend && npx vitest run test/workspace-runtime-renderer.test.ts test/config.test.ts`
Observed failures on `bc8afe0`:
- `workspace-runtime-renderer.test.ts`
- schema-v3 output omitted `resources.embeddings`
- schema-v3 still exposed top-level `embeddings`
- `config.test.ts`
- `https://qdrant:6333` was accepted
GREEN evidence:
`cd backend && npx vitest run test/workspace-runtime-renderer.test.ts test/config.test.ts`
Result:
- 2 test files passed
- 16 tests passed
Typecheck:
`cd backend && npx tsc --noEmit -p .`
Result:
- passed
Updated concerns:
- none for this round beyond future tightening if exact-port rejection is later requested explicitly.
@@ -1,149 +0,0 @@
# Task 4 Report — Narrow harness embedding configuration to internal Ollama
## Status
Implemented on 2026-08-08 in `/Users/mp/projects/ThothII/.worktrees/git-workspace-registry`.
## RED evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py -q
```
Observed before implementation:
- exit code `1`
- `10 failed, 10 passed`
- failures proved the missing `OllamaInternalEmbeddings` client and missing internal-only config validation
Representative failures:
- `ImportError: cannot import name 'OllamaInternalEmbeddings'`
- `AttributeError: 'EmbeddingsConfig' object has no attribute 'provider'`
- config tests `DID NOT RAISE ConfigError` for external provider, API key, and non-private base URL
## GREEN evidence
Focused behavior suite:
```bash
cd harness
./.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py -q
```
- exit code `0`
- `20 passed`
Relevant harness verification:
```bash
cd harness
./.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py tests/test_ollama_ensure.py -q
```
- exit code `0`
- `36 passed, 2 warnings`
Changed-file lint:
```bash
cd harness
./.venv/bin/ruff check tht/config.py tht/config_compat.py tht/vectorstore/embeddings.py tht/cli/ollama_cmd.py tests/test_config_resources.py tests/test_internal_embeddings.py
```
- exit code `0`
- `All checks passed!`
Patch hygiene:
```bash
git diff --check
```
- exit code `0`
## What changed
- translated schema-v3 `resources.embeddings` into the harness-compatible embedding config view
- validated the internal embedding contract only for that runtime-owned `resources.embeddings` path:
- provider must be `ollama_internal`
- model must be `qwen3-embedding:0.6b`
- dimensions must be `1024`
- base URL must be `http://embedding:11434` or loopback HTTP on port `11434`
- extra fields like `api_key` are rejected
- replaced the active embed client with `OllamaInternalEmbeddings`, using one bounded `/api/embed` request per batch
- removed task/query prefix rewriting from the active embedding path
- validated response count, vector dimension, and finite numeric values before returning embeddings
- kept `tht ollama ensure --json` stdout pristine while warming through the internal client
## Self-review
- kept changes inside the brief-listed files
- preserved DWH and session-persistence behavior
- preserved the legacy `OllamaEmbeddings` import path as an alias to avoid unrelated call-site churn
## Concerns
- the focused harness verification still emits two pre-existing warnings:
- `DeprecationWarning` from `testcontainers.postgres`
- `FutureWarning` because `resources` currently flows through the legacy config translation path
## Fix round 1 — 2026-08-08
### Findings addressed
- HIGH: external top-level `embeddings` remained an operational fallback and could still load
- MEDIUM: non-object embed JSON payloads escaped as raw `AttributeError`
### RED evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py tests/test_ollama_ensure.py -q
```
Observed before the fix:
- exit code `1`
- `2 failed, 36 passed, 2 warnings`
Representative failures:
- `AttributeError: 'list' object has no attribute 'get'` from `response.json()` returning a JSON array
- `Failed: DID NOT RAISE ConfigError` for top-level external `embeddings.provider=openai_compatible`
### GREEN evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py tests/test_ollama_ensure.py -q
```
Observed after the fix:
- exit code `0`
- `38 passed, 2 warnings`
Touched-file lint:
```bash
cd harness
./.venv/bin/ruff check tht/config.py tht/vectorstore/embeddings.py tests/test_internal_embeddings.py tests/test_config_resources.py
```
- exit code `0`
- `All checks passed!`
### Minimal fix
- validated the final active `cfg.embeddings` contract after config loading, so legacy top-level
embedding inputs now fail explicitly unless they exactly match the internal Ollama contract
- converted non-mapping embed JSON payloads into controlled `EmbeddingsError` failures with
sanitized diagnostics instead of raw attribute errors
@@ -1,170 +0,0 @@
# Task 5 Report — Implement the Qdrant VectorStore adapter
## Status
Implemented on 2026-08-08 in `/Users/mp/projects/ThothII/.worktrees/git-workspace-registry`.
## RED evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
```
Observed before implementation:
- exit code `2`
- collection failed during import because the adapter did not exist yet
Representative failures:
- `ModuleNotFoundError: No module named 'tht.adapters.vector.qdrant'`
## GREEN evidence
Focused behavior suite:
```bash
cd harness
./.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
```
- exit code `0`
- `31 passed, 1 warning`
Touched-file lint:
```bash
cd harness
./.venv/bin/ruff check tht/adapters/vector/qdrant.py tht/adapters/vector/__init__.py \
tht/ports/vector.py tht/vectorstore/records.py tht/vectorstore/store.py \
tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py
```
- exit code `0`
- `All checks passed!`
Patch hygiene:
```bash
git diff --check
```
- exit code `0`
## What changed
- added `QdrantVectorStore` with direct `requests`-based REST calls for:
- `GET /collections/{collection}`
- `PUT /collections/{collection}`
- `PUT /collections/{collection}/index`
- `PUT /collections/{collection}/points?wait=true`
- `POST /collections/{collection}/points/query`
- `POST /collections/{collection}/points/scroll`
- `POST /collections/{collection}/points/delete?wait=true`
- implemented idempotent collection provisioning for `1024` dimensions and `Cosine` distance
- created deterministic UUIDv5 point IDs from workspace, semantic kind, and canonical record key
- preserved canonical record identity and only upserted/deleted points matching the exact workspace
and generation filters
- added Qdrant payload helpers so stored payloads carry:
- `workspace_id`
- grouped semantic `kind` (`schema`, `evidence`, `memory`)
- original `record_kind`
- canonical `record_key`
- `content_hash`
- existing Thoth metadata fields
- mapped Qdrant payloads back into existing `VectorHit` objects without losing the original
Thoth kind
- exported the new adapter from the public vector adapter package and added focused contract tests
- sanitized timeout and malformed-response failures so CLI-facing callers do not leak raw endpoint
details
## Self-review
- confirmed collection mismatch fails without any delete/recreate path
- confirmed every query/scroll/delete operation includes a workspace filter
- confirmed the adapter never deletes or rewrites unrelated Qdrant points
- added keyword payload indexes for all filter-critical fields used here, including `document_id`
for exact Evidence filtering
## Concerns
- the requested `adversarial-review` skill could not run its full external reviewer flow in this
environment because the skill’s referenced `brain/` files are missing at
`/Users/mp/.agents/skills/adversarial-review`; I performed a manual adversarial self-review
instead
- the focused suite still emits one pre-existing warning from `testcontainers.postgres`
## Fix round 1 — 2026-08-08
### Findings addressed
- IMPORTANT: metadata collisions could override canonical Qdrant payload identity fields and break
workspace isolation
- IMPORTANT: scroll-based operations only read the first page and did not follow
`next_page_offset`, making `existing_hashes`, `list_evidence_generations`, and delete counts
inexact beyond one page
### RED evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
```
Observed before the fix:
- exit code `1`
- `2 failed, 31 passed, 1 warning`
Representative failures:
- `assert payload["workspace_id"] == "demo"` failed because colliding `record.metadata`
overwrote canonical payload fields
- paginated scroll test missed later pages, so `existing_hashes` and generation cleanup counts
were incomplete
### GREEN evidence
Command:
```bash
cd harness
./.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
```
Observed after the fix:
- exit code `0`
- `33 passed, 1 warning`
Touched-file lint:
```bash
cd harness
./.venv/bin/ruff check tht/adapters/vector/qdrant.py tht/vectorstore/records.py \
tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py
```
- exit code `0`
- `All checks passed!`
Patch hygiene:
```bash
git diff --check
```
- exit code `0`
### Minimal fix
- made `qdrant_payload` apply canonical fields after `record.metadata` so workspace ID, semantic
kind, original record kind, canonical record key, and content hash cannot be overridden by
metadata collisions
- paginated `_scroll` until `next_page_offset` is absent, sent the returned `offset` back on the
next request, and reject repeated offsets as malformed to avoid infinite loops
@@ -1,144 +0,0 @@
# Task 6 Report
Date: 2026-08-08
Status: implemented and verified
Summary:
- Added schema-v3 Qdrant runtime support to the harness config/resource layer and vector factory.
- Made Qdrant payloads carry `workspace_id` and `workspace_revision` on every point.
- Routed schema and memory bulk indexing through the transport-neutral vector port with canonical hash-based dedup.
- Kept Evidence canonical on filesystem and Memory canonical in JSONL; Qdrant remains derived/rebuildable.
- Added focused tests for semantic-kind isolation, shared identity fields, search-pack kind boundaries, and the schema-v3 factory/config path.
Files changed:
- `harness/tht/config.py`
- `harness/tht/config_compat.py`
- `harness/tht/adapters/factory.py`
- `harness/tht/adapters/vector/qdrant.py`
- `harness/tht/vectorstore/records.py`
- `harness/tht/cli/vector_cmd.py`
- `harness/tht/cli/memory_cmd.py`
- `harness/tests/test_semantic_kind_isolation.py`
- `harness/tests/test_memory_save_one.py`
- `harness/tests/test_search_pack.py`
- `harness/tests/test_qdrant_vector_store.py`
- `harness/tests/test_adapter_factory.py`
- `harness/tests/test_config_resources.py`
Verification:
- Focused RED/GREEN task suite:
- `cd harness && .venv/bin/pytest tests/test_semantic_kind_isolation.py tests/test_memory_save_one.py tests/test_search_pack.py -q`
- Relevant harness suite:
- `cd harness && .venv/bin/pytest tests/test_semantic_kind_isolation.py tests/test_memory_save_one.py tests/test_search_pack.py tests/test_qdrant_vector_store.py tests/test_adapter_factory.py tests/test_config_resources.py tests/test_vector_port_contract.py tests/test_corpus_pipeline.py -q`
- Result: `131 passed`
- Changed-file Ruff:
- `cd harness && .venv/bin/ruff check tht/vectorstore/records.py tht/adapters/vector/qdrant.py tht/config_compat.py tht/config.py tht/adapters/factory.py tht/cli/vector_cmd.py tht/cli/memory_cmd.py tests/test_memory_save_one.py tests/test_search_pack.py tests/test_semantic_kind_isolation.py tests/test_qdrant_vector_store.py tests/test_adapter_factory.py tests/test_config_resources.py`
- Result: clean
Concerns / follow-up:
- `memory clear` still retains its older direct-vector assumptions and was not expanded in this task because the brief focused on canonical builders and schema/evidence/memory routing through the active Qdrant path.
- The relevant suite still emits pre-existing warnings (legacy config deprecation in older fixtures, plus existing Pydantic serializer warnings in corpus tests), but they are not introduced by this task.
## Fix round 1 (2026-08-08)
Scope:
- Fixed qdrant-only schema-v3 command gating for `vector index-schema`, `memory promote`, and `memory index`.
- Replaced `memory clear`'s direct-pgvector-only path with vector-port deletion by kind.
- Added focused qdrant-only CLI regression tests and refreshed older CLI fixtures to the enforced internal embedding contract.
RED evidence:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py -q`
- Initial result against commit `5e39cfa`: `4 failed`
- Failure signatures:
- `ERRORE: sezioni mancanti nel workspace yaml: vector_db o vector_write_rest.`
- `ERRORE: sezioni mancanti nel workspace yaml: vector_db.`
GREEN evidence:
- Focused fix suite:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py tests/test_qdrant_vector_store.py tests/test_adapter_factory.py tests/test_config_resources.py tests/test_memory_save_one.py tests/test_search_pack.py -q`
- Result: `51 passed`
- Relevant broader vector/memory/schema/search suite:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py tests/test_qdrant_vector_store.py tests/test_adapter_factory.py tests/test_config_resources.py tests/test_memory_save_one.py tests/test_search_pack.py tests/test_vector_port_contract.py tests/test_adapter_command_regressions.py tests/test_solved_search_cli.py tests/test_schema_introspect_guard.py tests/test_semantic_kind_isolation.py tests/test_corpus_pipeline.py -q`
- Result: `154 passed`
- Ruff on the fix surface:
- `cd harness && .venv/bin/ruff check tht/ports/vector.py tht/adapters/vector/qdrant.py tht/adapters/vector/pgvector.py tht/adapters/vector/thoth_http.py tht/vectorstore/rest_client.py tht/cli/vector_cmd.py tht/cli/memory_cmd.py tests/test_qdrant_cli_commands.py tests/test_solved_search_cli.py`
- Result: clean
Notes:
- `memory clear` now deletes derived `kind=memory` points through the configured writable vector store, while leaving the JSONL registry as the source of truth until the registry file is removed by the command.
- The broader suite still carries the same pre-existing warnings noted above; this fix round did not add new warnings or failures.
## Fix round 2 (2026-08-08)
Scope:
- Removed the accidental HTTP writer `delete_kinds` capability expansion from `ThothHttpVectorStore` and `VectorRestClient`.
- Reworked `memory clear` so schema-v3 Qdrant uses scoped `kind=memory` deletion, while legacy transports keep the pre-task direct-sync path instead of advertising a nonexistent RPC.
- Tightened the qdrant-only memory-clear regression to assert the exact `("memory", ["memory"])` delete scope.
RED evidence:
- Re-review found a transport contract mismatch in fix round 1:
- `ThothHttpVectorStore` exposed `delete_kinds(...)`
- `VectorRestClient` exposed `delete_kinds(...)`
- but the legacy HTTP writer migration only allowlists `delete_vector_generation`, not `delete_vector_kinds`
- The new regressions added in this round capture that mismatch and the missing qdrant delete-scope assertion:
- `tests/test_vector_port_contract.py::test_http_store_supports_writer_without_reader`
- `tests/l0/test_vector_adapter_parity.py::test_http_rest_client_does_not_advertise_nonexistent_delete_kinds_rpc`
- `tests/test_qdrant_cli_commands.py::test_memory_clear_accepts_qdrant_only_runtime_config`
GREEN evidence:
- Focused regression suite:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py tests/test_vector_port_contract.py tests/l0/test_vector_adapter_parity.py tests/test_adapter_command_regressions.py -q`
- Result: `53 passed`
- Broader relevant vector/memory/search suite:
- `cd harness && .venv/bin/pytest tests/test_qdrant_cli_commands.py tests/test_adapter_command_regressions.py tests/test_vector_port_contract.py tests/l0/test_vector_adapter_parity.py tests/test_solved_search_cli.py tests/test_qdrant_vector_store.py tests/test_search_similar_kinds.py tests/test_corpus_pipeline.py -q`
- Result: `135 passed`
- Ruff on the changed fix surface:
- `cd harness && .venv/bin/ruff check tht/cli/memory_cmd.py tht/ports/vector.py tht/adapters/vector/thoth_http.py tht/vectorstore/rest_client.py tests/test_qdrant_cli_commands.py tests/test_vector_port_contract.py tests/l0/test_vector_adapter_parity.py`
- Result: clean
Notes:
- Legacy HTTP/vector-rest deployments do not gain a new destructive RPC surface from this fix; they keep their previous behavior and continue to fail closed for unsupported cleanup.
- The broader suite still emits the same pre-existing deprecation and serializer warnings already noted above; this round did not introduce new warnings.
## Fix round 3 (2026-08-08)
Scope:
- Added an adapter-level Qdrant regression for mixed semantic kinds within one workspace plus a second workspace memory point.
- Verified that `delete_kinds("memory", ["memory"])` emits the real adapter filter with both `workspace_id=demo` and `record_kind=memory`.
- Verified that non-memory semantic kinds in the same workspace and memory from another workspace survive the delete.
RED evidence:
- Re-review identified a test gap rather than a confirmed runtime bug:
- existing coverage asserted only the CLI mock call shape for qdrant memory clear
- there was no adapter-level regression proving the real Qdrant delete filter and resulting fake-Qdrant state across mixed semantic kinds/workspaces
- Added regression:
- `tests/test_qdrant_vector_store.py::test_delete_kinds_is_workspace_scoped_and_preserves_other_semantic_kinds`
GREEN evidence:
- Requested focused suite:
- `cd harness && .venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_qdrant_cli_commands.py tests/test_semantic_kind_isolation.py -q`
- Result: `18 passed`
- Ruff on changed files:
- `cd harness && .venv/bin/ruff check tests/test_qdrant_vector_store.py`
- Result: clean
Notes:
- This round required no production change; the new adapter regression passed against the existing Qdrant implementation.
- The focused suite still emits the same pre-existing `testcontainers.postgres` deprecation warning from `tests/conftest.py`; no new warnings were introduced.
@@ -1,98 +0,0 @@
# Task 7 report — mandatory Qdrant and Ollama Compose services
Date: 2026-08-08
Status: completed
Summary:
- Added mandatory private `qdrant`, `embedding`, and `embedding-model-init` services to the base Compose stack.
- Pinned Qdrant `v1.18.2` and Ollama `0.32.0` by immutable multi-arch digest.
- Persisted Qdrant storage in `qdrant-data` and Ollama model cache in `embedding-models`.
- Wired `core` to fixed internal semantic endpoints:
- `THT_INTERNAL_QDRANT_URL=http://qdrant:6333`
- `THT_INTERNAL_EMBEDDING_URL=http://embedding:11434`
- `THT_INTERNAL_EMBEDDING_MODEL=qwen3-embedding:0.6b`
- `THT_INTERNAL_EMBEDDING_DIMENSIONS=1024`
- Removed external vector / embedding endpoint requirements from the local and server env examples.
- Added an idempotent Ollama model bootstrap script that:
- waits up to a bounded deadline for `/api/tags`
- skips `ollama pull` when the model is already cached
- pulls `qwen3-embedding:0.6b` only when needed
- verifies the model appears in `/api/tags` after pull
- Added optional GPU override file `deploy/compose.embedding-gpu.yaml`; base Compose remains CPU-only.
- Updated `scripts/run-stack.sh` so the GPU override is included only when `THOTH_ENABLE_EMBEDDING_GPU=1`.
Verification:
- RED confirmed before implementation:
- `./scripts/test-default-compose.sh` failed on missing required services.
- `./scripts/test-unified-compose.sh` failed on missing required services.
- `./scripts/test-internal-semantic-compose.sh` failed because the GPU override file did not exist.
- GREEN after implementation:
- `./scripts/test-default-compose.sh`
- `./scripts/test-unified-compose.sh`
- `./scripts/test-internal-semantic-compose.sh`
- `git diff --check`
- Additional shell verification:
- `scripts/run-stack.sh --wait` includes only base + local Compose files by default.
- `THOTH_ENABLE_EMBEDDING_GPU=1 scripts/run-stack.sh --wait` adds `deploy/compose.embedding-gpu.yaml`.
Resolved image digests:
- `qdrant/qdrant:v1.18.2@sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c`
- `ollama/ollama:0.32.0@sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a`
Self-review:
- The first bootstrap-script draft depended on tools not guaranteed inside the Ollama image. This was corrected after image inspection; the final script uses only confirmed image tools (`bash`, `ollama`, `grep`) plus raw HTTP over `/dev/tcp`.
- The server overlay intentionally replaces most named core volumes with bind mounts, so the unified contract was tightened to require named semantic-cache volumes there while preserving the local/base named-volume checks.
Concerns:
- The model bootstrap waits for Ollama readiness and verifies cache state, but the first real cold-start will still take time to download `qwen3-embedding:0.6b`.
- The GPU override requests generic Docker GPU capability only; actual GPU availability remains host/runtime dependent and intentionally stays opt-in.
## Fix round 1 / 5 — 2026-08-08
Rulings applied:
- Kept the Task 1 boundary intact: schema-v3 remains the only operational workspace descriptor shape.
- Did not restore any external semantic fallback for schema-v2 live sessions.
- Treated `PROJECT_STATE.md` as stale documentation for this point, not runtime truth.
Focused schema-v2 evidence:
- Re-ran the existing targeted registry test:
- `cd backend && npx vitest run test/workspace-registry.test.ts -t "lists a schema v2 descriptor as migration_required and refuses to acquire it"`
- Result: pass.
- Evidence from that test:
- schema-v2 descriptors list as `migration_required`
- `acquireSessionRevision("psd-clinical")` rejects with `code: "workspace_invalid"`
- Conclusion: schema-v2 acquisition remains blocked; no external semantic fallback was reintroduced.
Contract consistency fixes:
- Updated `harness/tests/test_local_compose_contract.py` to assert the mandatory internal semantic stack, fixed internal core semantic env, private-service topology, persistent volumes, and Ollama health/dependency contract.
- Updated shell Compose contracts to require:
- Ollama healthcheck on `embedding`
- `embedding-model-init` dependency on `embedding: service_healthy`
- Updated `scripts/unified-deployment-smoke.sh` rendered-contract helper to expect the mandatory internal semantic topology and internal semantic env names, and to reject retired external semantic bindings.
- Updated `scripts/test-task13-runtime-fixtures.sh` to exercise `task13_assert_rendered_contract` for both local and server fixture renders.
Fix round 1 verification:
- RED before implementation:
- `cd harness && .venv/bin/pytest tests/test_local_compose_contract.py -q` failed because `embedding` had no healthcheck.
- `./scripts/test-default-compose.sh` failed because `embedding` had no healthcheck.
- `./scripts/test-unified-compose.sh` failed because `embedding` had no healthcheck.
- `./scripts/test-task13-runtime-fixtures.sh local` failed because `unified-deployment-smoke.sh` still expected `core,frontend`.
- GREEN after implementation:
- `./scripts/test-default-compose.sh`
- `./scripts/test-unified-compose.sh`
- `./scripts/test-internal-semantic-compose.sh`
- `cd harness && .venv/bin/pytest tests/test_local_compose_contract.py -q`
- `./scripts/test-task13-runtime-fixtures.sh local`
- `./scripts/test-task13-runtime-fixtures.sh server`
- `cd backend && npx vitest run test/workspace-registry.test.ts -t "lists a schema v2 descriptor as migration_required and refuses to acquire it"`
- `docker compose --env-file deploy/env/local.env.example -f compose.yaml -f deploy/compose.local.yaml config --format json`
@@ -1,65 +0,0 @@
Status: completed on August 8, 2026.
Summary:
- Updated the frontend workspace contract from schema v2 editing to schema v3 publishing.
- Kept only `semantic_index.vector_store.collection` editable; rendered qdrant / internal Ollama semantic values as fixed read-only architecture values.
- Removed external vector transport / endpoint / credential / embedding diagnostics branches from frontend draft sanitization, conflict parsing, and editor UI.
- Added a migration-required banner in workspace management and blocked `migration_required` workspaces from new-session selection.
- Aligned the example workspace YAML comments with the fixed internal qdrant/Ollama architecture.
Files changed:
- `frontend/src/api/workspaces.ts`
- `frontend/src/api/workspaces.test.ts`
- `frontend/src/workspaces/drafts.ts`
- `frontend/src/workspaces/drafts.test.ts`
- `frontend/src/shell/WorkspaceEditor.tsx`
- `frontend/src/shell/WorkspaceEditor.test.tsx`
- `frontend/src/shell/WorkspaceManager.tsx`
- `frontend/src/shell/WorkspaceManager.test.tsx`
- `frontend/src/shell/WorkspacePublishDialog.test.tsx`
- `frontend/src/api/sessions.ts`
- `frontend/src/shell/SteerInput.tsx`
- `frontend/src/shell/SteerInput.test.tsx`
- `deploy/workspaces/example.yaml`
- `deploy/workspaces/psd.yaml.example`
Verification:
- `cd frontend && npx vitest run src/shell/SteerInput.test.tsx src/shell/WorkspaceEditor.test.tsx src/shell/WorkspaceManager.test.tsx src/shell/WorkspacePublishDialog.test.tsx src/workspaces/drafts.test.ts src/api/workspaces.test.ts`
- Result: 6 files passed, 59 tests passed.
- `cd frontend && npx tsc -b`
- Result: passed.
- `git diff --check`
- Result: passed.
Self-review:
- The frontend now publishes the exact schema v3 semantic shape and no longer persists legacy semantic transport/credential branches.
- Migration-required workspaces are visible in management with an explicit banner and are excluded from the composer workspace selector.
- One dependent test file outside the original brief list (`WorkspacePublishDialog.test.tsx`) and the composer/session-selection path (`api/sessions.ts`, `SteerInput.tsx`, related test) were updated because they were directly coupled to the old v2 semantic/edit-selection behavior.
Concerns:
- The composer still retains backward-compatible behavior for summaries that omit `revision` entirely; only explicit `revision.state === "migration_required"` is blocked. That matches the current mixed-test environment, but once summary responses are guaranteed to include `revision`, that fallback may be removable.
Fix round 1/5 — August 8, 2026
Summary:
- Made missing or invalid workspace summaries fail safe in frontend session creation and composer selection instead of falling open as legacy.
- Added an actionable unavailable message in workspace management for incomplete summaries with no canonical revision.
- Replaced the old runtime-oriented example descriptor files with exact backend WorkspaceV3 descriptor YAML.
Additional files changed:
- `frontend/src/api/sessions.test.ts`
- `backend/test/workspaces-schema.test.ts`
Fix-round verification:
- `cd frontend && npx vitest run src/api/sessions.test.ts src/shell/SteerInput.test.tsx src/shell/WorkspaceManager.test.tsx src/shell/WorkspaceEditor.test.tsx src/shell/WorkspacePublishDialog.test.tsx src/workspaces/drafts.test.ts src/api/workspaces.test.ts`
- Result: 7 files passed, 73 tests passed.
- `cd frontend && npx tsc -b`
- Result: passed.
- `cd backend && npx vitest run test/workspaces-schema.test.ts`
- Result: 1 file passed, 17 tests passed.
- `git diff --check`
- Result: passed.
Notes:
- Missing `revision` in a workspace summary now fails with the same session/composer safety posture as `migration_required`, using the existing safe workspace-policy error for session creation and an explicit unavailable message in workspace management.
- The committed example files now validate as actual schema-v3 descriptors instead of deployment/runtime templates with forbidden semantic endpoint fields.
@@ -1,83 +0,0 @@
# Task 5 Report — `tht setup` lifecycle orchestration
## Status
Completed. `tht setup` now validates the checkout and host prerequisites, creates or validates
the non-secret installation files, validates Compose, and by default builds, starts, health-checks,
and verifies the installation. `tht setup --configure-only` stops immediately after successful
Compose rendering.
## Implementation
- Added `setup.Run`, with an ordered host preflight: project/worktree discovery, Docker Engine,
Docker Compose, supported architecture, and LF line-ending checks.
- Reused `config.Installation.ComposeArgs` for all Compose calls and added a narrow
`compose.InstallationRunner` adapter for Pi diagnostics; no shell command construction was added
to the top-level CLI parser.
- Default setup performs `compose build`, `compose up --detach --remove-orphans`, bounded polling
for `core`, `frontend`, `qdrant`, `embedding`, and `embedding-model-init`, then aggregate volume
diagnostics and `pi.Doctor`.
- Health timeout errors identify the last failing service and preserve containers for diagnosis,
with `tht logs <service>` and `tht status` guidance.
- Completion output includes the frontend URL, selected descriptor, and next action.
## TDD evidence
The initial focused test run failed because `setup.Run` did not exist. Tests were then written
against a fake Compose runner before the orchestration was implemented. They cover the complete
ordered flow, configure-only stop, preflight failure before writing configuration, health retry,
timeout guidance, and CLI default versus `--configure-only` dispatch.
## Verification
Executed from `tools/tht`:
```bash
go test ./internal/setup ./internal/compose ./cmd/tht -run 'TestRun|TestSetupCommand|TestInstallationRunner' -count=1
go test ./internal/setup ./internal/compose ./cmd/tht -count=1
go test ./...
git diff --check
```
All commands passed. No actual Docker build, container start, live-stack restart, system
installation, Pi configuration edit, or documentation rewrite was performed.
## Commit
`feat(setup): build start and verify ThothII` (this report is included in that commit).
## Concerns
- The bounded health wait is verified with fakes only, as required for this task. Real Docker
lifecycle verification belongs to the later live acceptance task.
- The existing aggregate `tht doctor` command remains a separate implementation; Task 5 performs
its equivalent setup-time prerequisite checks plus `pi.Doctor` without invoking a nested CLI
process.
## Fix round 1
The independent review identified three gaps. All were reproduced with RED tests before the
production change:
- A rendered Compose document containing any one volume was accepted. `requireVolumes` now
requires `settings`, `pi-state`, `workspace-registry`, `workspace-secrets`, `sessions`,
`qdrant-data`, and `embedding-models`; tests reject each individual omission and an
unrelated-only volume set.
- Failures after `compose up` could return without recovery instructions. A single recovery
wrapper now preserves the underlying error while adding the retained-container, `tht logs
<service>`, and `tht status` guidance for failed `up`, health, aggregate doctor, and Pi doctor
phases. Focused tests also prove build failure stops before attempting startup.
- LF inspection previously walked the full checkout. It now inspects only `compose.yaml`,
`deploy/`, and `docker/`; a test proves CRLF content under `node_modules/` is ignored.
Verification added for this round:
```bash
go test ./internal/setup -run 'TestRequireVolumes|TestRun(BuildFailure|UpFailure|AggregateDoctorFailure|PiDoctorFailure|IgnoresIrrelevant|TimesOut)' -count=1
go test ./internal/setup -count=1
```
Both passed before the final full-suite verification. No Docker or live operation was run.
Implementation commit evidence: `ea70cc95b04532043744a9de6c5912e30a214595` —
`fix(setup): harden verification and recovery`.
@@ -1,55 +0,0 @@
# Task 6 — Version, aggregate doctor, and build-aware start
Status: complete.
Implemented the host-side `tht version`, aggregate `tht doctor [--json]`, and `tht start [--build]` contracts.
- `version` is descriptor-free and reports semantic version, commit, build time, OS, and architecture.
- `doctor` emits typed, redacted checks for descriptor state, Docker/Compose, rendered volumes, file permissions, service health, workspace registry, the container-local workflow doctor, and Pi doctor. Its JSON mode writes exactly one JSON document to stdout.
- The Python workflow doctor is invoked only as `docker compose exec -T core tht doctor --json` after core is running.
- `start` uses the shared lifecycle service: default `up → health`; `--build` is `build → up → health`.
- `setup` now reuses the shared lifecycle and aggregate diagnostics rather than keeping parallel health/volume implementations.
Verification performed without live Docker/container commands:
```bash
cd tools/tht
go test ./internal/version ./internal/doctor ./internal/service ./cmd/tht \
-run 'TestVersion|TestDoctor|TestStart|TestCurrent|TestRun' -count=1
go test ./internal/setup -count=1 -run 'TestRun' -v
go test ./... -count=1
git diff --check
```
All completed successfully. The intentionally fake runner coverage includes unavailable Docker,
stopped/running core, workflow failure redaction, pristine JSON output, and start ordering.
Concerns: no live Docker validation or host installation was run, by explicit task constraint.
## Fix round 1
Completed the independent-review follow-up without live Docker operations.
- `workspace-registry` now executes a container-local, read-only Node validation of
`/data/workspace-registry/state/active.json` and every declared snapshot descriptor. It no
longer passes merely because Compose declares a volume.
- Host file permissions are checked before Docker/Compose availability and therefore remain
visible as failures when Docker is unavailable.
- Separate typed, bounded HTTP probes verify core (`curl --max-time 5`) and frontend
(`wget -T 5`) reachability, independently of Compose health. The probe is injectable in tests.
- The successful report tests assert the stable full checklist:
`descriptor`, `files`, `docker`, `compose`, `configuration`, `services`, `core-http`,
`frontend-http`, `workspace-registry`, `workflow`, `pi`.
Additional verification:
```bash
cd tools/tht
go test ./internal/doctor -run 'TestRun(ChecksUnsafeFilesEvenWhenDockerIsUnavailable|FailsAnInvalidContainerLocalRegistryState|ReportsEachHTTPReachabilityProbeFailure|UsesOnlyContainerLocalWorkflowAndPiDiagnosticsWhenCoreRuns)' -count=1 -v
go test ./internal/doctor ./internal/setup ./internal/service ./cmd/tht -count=1
go test ./... -count=1
git diff --check
```
All passed with fake runners/probes only. No live container, HTTP endpoint, or host installation
was touched.
@@ -1,163 +0,0 @@
# Task 15 retained release-gate report — fix round 5 (sanitized)
## Final-review fix-round-2 addendum — frozen source `2a9359071257f9b8a71d36ec2bbb25b161003f81`
This addendum supersedes the fix-round-1 addendum for current authentication remediation status
while preserving the fix-round-5 material below as historical provenance.
- Authentication remediation status: `PASS`. The three original remediation Important findings
remain `RESOLVED`; the fix-round-2 fully bounded lifecycle Important is `ADDRESSED`; and the
temporary Windows diagnostic-matrix Minor is `ADDRESSED`.
- Overall branch/release readiness is separately `FAIL`, with unavailable external/manual gates
`PENDING`.
- Completed exact-source workflow run `32147345625` concluded `failure` on baseline release jobs.
Its `Windows clone and Compose contract` job (`95744249248`) executed the unfiltered command
`go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1`; the native step
passed all three packages: safeio `22.058s`, backup `7.161s`, authstorage `16.088s`.
- The Windows job failed only afterward in the baseline clone-contract script at
`scripts/test-windows-clone-contract.ps1:208`, where PowerShell rejects the undelimited
`$remoteYaml:` variable reference.
- `LF, Compose, docs, and TypeScript` job `95744249458` reproduced the baseline unset-`TMPDIR`
failure after unified Compose passed. Linux Docker job `95744249354` reproduced the missing-`rg`
prerequisite failure; cleanup passed and no image manifest was generated.
- The skipped Windows Docker Desktop/WSL2 job is recorded as `NOT_RUN` / `BLOCKED`, not FAIL.
Downstream commands skipped after executed baseline failures use the same classification. The
matrix contains an explicit native `windows_stagearchive_retained_capability` PASS row.
- Historical Node/auth/browser/docs PASS and harness/Ruff/Compose FAIL evidence remains bound to
its recorded source where not rerun. L2, real PSD/manual acceptance, and provider readiness
remain `PENDING`.
- Current machine-readable evidence and the requested Task 4 report are recorded in
`.artifacts/task-15/automated-gates.json` and
`.superpowers/sdd/2026-08-18-thothii-authentication-remediation/task-4-report.md`.
- The full fix-round-2 RED/GREEN and finding disposition is recorded in
`.superpowers/sdd/2026-08-18-thothii-authentication-remediation/fix-round-2-report.md`.
- Current automated-gates SHA-256:
`6c516db5c2064c4a4a2e5f25961b993cd4a8fe020bbbb822fbac7faa0c119599`.
- Historical unified Docker manifest SHA-256: `9c8dec4546909fd93799dbcf374bcb3a89bc46cfe0fd482472c0cbe757ddf5b6`.
The complete sanitized Task 4 matrix and the separate remediation/release verdicts are in the
requested Task 4 report.
- Final tested source commit: `74b062f1a737103524cbe706346cfd65f87cdfd1`.
- Historical retained source commits: fix-round-2 `fe190e7046acc173f510dddcb32f46ed142858c1`,
maintenance follow-up `4d230b87afdcd24f02264f8f937c8628b92db05a`, prior final Docker
source `e20bf33e2a00102192e5be66b178037aeca3a7b1`, and fix-round-4 streamed
archive privacy `54698e73400a54ce7c3e6c10099e14eb471ce8b9`.
- Versions: Node contract `v24.16.0`; host default Node `v25.6.1`; Go `go1.26.5`;
Pi `0.80.3`.
- Historical automated gate artifact: `.artifacts/task-15/automated-gates.json`;
SHA-256 `7d9ec93af15510605f1aa7179b26a7ee46d78122f647854300f7a9922057a63f`.
- Docker image manifest: `.artifacts/task-15/unified-docker-images.json`;
SHA-256 `9c8dec4546909fd93799dbcf374bcb3a89bc46cfe0fd482472c0cbe757ddf5b6`.
## Fix-round-5 evidence
- PASS, RED then GREEN: `TestCreateCanonicalNewPrivateFileUsesPinnedParentAfterAncestorSwap`
first failed because the creator had not retained its parent before creation. It now opens every
Unix ancestor once, creates the leaf with `openat(O_NOFOLLOW|O_CREAT|O_EXCL)`, applies and checks
`0600` by descriptor (`fchmod`/`fstat`), and uses `unlinkat` for creator failure cleanup. The
deterministic test moves the opened parent, replaces its lexical name with an outside symlink,
validates the archive under the moved original parent, and proves no outside archive was written.
- PASS: the Windows implementation uses NT `RootDirectory`-relative traversal for every component
after the volume root and for final file creation. The retained final parent receives only the
required child-create right (`FILE_WRITE_DATA` for a file, `FILE_APPEND_DATA` for a directory),
reparse points are rejected, and the owner-only protected DACL is installed in the same
`NtCreateFile` operation. The native-Windows test attempts the pre-create parent swap and calls
`safeio.ValidatePrivateRegular`; it is compiled but not executed on this host.
- PASS: `go test ./internal/safeio ./internal/backup -count=1`, `go test -race ./...` across
`18` packages, `go vet ./...`, and a native host `tht` CLI build. Existing StageArchive
capacity, lifecycle, rollback, streaming, and cleanup tests remain passing.
- PASS, compile-only: Windows amd64 static test/build compilation across `18` packages, including
the retained-handle Windows tests. No Windows executable was run; native execution remains
PENDING and is not inferred from compilation.
- PASS on Node `v24.16.0`: the hermetic OIDC/F1 authentication browser smoke passed all current
`8` checks in `frontend/e2e/auth.spec.ts` and `frontend/e2e/f1.spec.ts`; the runtime sentinel
leak scan passed.
- PASS: shell syntax, unified-smoke safety self-test, default Compose contract, unified Compose
contract, and Compose secret-policy contract.
- PASS: final unified Docker deployment smoke run `20260818070637-66409-30058`, bound exactly to
source `74b062f1a737103524cbe706346cfd65f87cdfd1`. It exercised maintenance-auth isolation,
restore, registry lifecycle, bad-candidate rollback, image revalidation, and task-scoped cleanup.
## Sanitized final unified Docker output
```text
== Build and start isolated local Compose distribution ==
== Recreate offline and retain the validated registry snapshot ==
== Pull a valid catalog+descriptor metadata update ==
== Pull a content-only Git Evidence update ==
== Reject catalog/descriptor metadata mismatch and retain the valid snapshot ==
== Reject orphan descriptor directories not listed in the catalog ==
== Reject the retired flat workspace layout and retain the valid snapshot ==
== Inject a bad pinned Pi candidate and prove automatic rollback ==
Task 13 full deployment smoke passed.
Task 13 cleanup proof: no labeled containers, volumes, networks, or images remain for 20260818070637-66409-30058.
```
## Sanitized Docker image identities
- `sha256:2d7b19491c7eb8c119c3cedb390aaeb2ff5593f6fc43ab66c317565560da6d7d`;
roles `compose-runtime`, `fixture-runtime`.
- `sha256:3b6c31a5d8f8fc58fa3233391b6175bd2fbc793eebb44d5e285ecc6e02e9e687`;
role `compose-runtime`.
- `sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a`;
role `compose-runtime`.
- `sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c`;
role `compose-runtime`.
- `sha256:c3cbe1cc1aa588a64951ac6286e0df7b27fe2e6324b1001c619bb358770c0178`;
role `rollback-candidate`.
For each image, the retained repository-digest component equals the listed image digest. Registry
names and credentials are deliberately omitted.
## Complete observed matrix
- PASS: Task 13 lifecycle carry-ins; retained-handle owner-private restore staging; provider fixture
round-one `6/6`; backend Node 24 round-one suite `75 files / 1081 tests`; frontend Node 24
round-one suite `61 files / 444 tests`; current Node 24 authentication/F1 browser smoke `8/8`;
final-source Go race/build `18 packages`; Windows static cross-compile `18 packages`; harness
round-one suite `921 passed / 4 L2 deselected`; authentication docs round-one gate; shell/Compose
contracts; final unified Docker smoke; five-image traceability; and Docker cleanup.
- FAIL: Ruff `192` known-baseline errors; MkDocs strict `69` known-baseline warnings; existing
canonical/workspace install wording checks; existing Pi model-policy check; deployment-coupling
scan against preserved ignored private material.
- PENDING: native Windows execution because required host prerequisites are unavailable; L2 because
the configured secret layout is unavailable; real PSD/manual acceptance because no real
identity/access is available; isolated provider readiness because an unrelated host port is
occupied.
## Final Task 15 review after fix round 5
The fresh Terra review verdict is **CHANGES REQUIRED**. The five-round breaker is exhausted; no
sixth implementation round was started. Two Important findings remain:
- `StageArchive` does not retain the opaque parent/directory capability through the complete
stream and `Close` lifecycle. Staging-directory creation and final cleanup still use pathname
operations, so an ancestor swap after creation can strand the secret-bearing archive or redirect
cleanup. Deterministic StageArchive swap-and-cleanup coverage is still required on Unix and
native Windows.
- Windows claim removal closes its validated retained parent handles before calling pathname-based
`DeleteFile`. Removal must instead remain handle-relative (or delete through the opened handle),
with a native-Windows ancestor-swap test.
The focused/full Go, cross-compile, Node 24, browser, Compose, Docker lifecycle, image-traceability,
and cleanup results above remain valid evidence for source `74b062f1a737103524cbe706346cfd65f87cdfd1`.
They do not override the final code-review verdict. Native Windows execution remains PENDING.
The authentication feature is **not implementation-complete or release-complete** while these code
findings and the required FAIL/PENDING gates remain. No secret values, real identities, internal
endpoints, or registry names are retained.
## Final whole-branch review
The final read-only Terra review of `351361f..39b5453` also returned **CHANGES REQUIRED** and found
one additional Important issue: the POSIX local-user registry validates file type, link count, and
mode for `users.yaml` and its parent directory, but does not require ownership by the effective UID.
A foreign-owned `0600` registry inside a runtime-owned `0700` directory can remain writable by the
foreign owner and be used to alter credentials or grant the administrator role. The registry must
enforce effective-UID ownership on every POSIX `lstat`/`fstat` path and add foreign-owner rejection
coverage.
No new Critical issue or load-bearing Minor issue was found. The branch is **not ready to merge**:
this ownership defect and the two retained-capability cleanup defects above require fixes and renewed
review, independently of the remaining FAIL/PENDING release gates.
@@ -1,162 +0,0 @@
# Final-review fix round 1 report (sanitized)
## Verdict
- Base: `fa499a9bdd37011833691b0f447470d8b7e8a3a6`.
- Final frozen source: `10cd66fe6a5b484a4dc569326a228c1c5484a5d4` on
`feat/thoth-auth`.
- Authentication remediation: **PASS / ADDRESSED**. All four final-review Important findings are
resolved relative to the remediation brief.
- Terra Minor evidence corrections: **ADDRESSED**.
- Branch/release readiness: **FAIL**. The completed exact-source workflow still contains executed
baseline clone-contract, LF/Compose, and Linux Docker failures. Unavailable external/manual
gates remain **PENDING**.
- Source and evidence remain separate commits. No workflow was dispatched from the evidence-only
phase.
## Finding disposition
| Finding | Disposition | Evidence |
|---|---|---|
| Important 1 — exhaustive Windows cleanup | RESOLVED | Cleanup now attempts close/delete/validation operations in deterministic order and returns sanitized `ErrUnsafeFile` after aggregating failures. `TestWindowsPrivateRegularCleanupClosesAfterDeleteDispositionFailure` and `TestWindowsClaimCleanupAttemptsLaterOperationsAfterEarlierFailure` cover the non-short-circuit contract. Global no-delete sharing remains unchanged. |
| Important 2 — usable native Windows authority | RESOLVED | Owner-only descriptors use the current user SID, protected/non-defaulted DACL semantics, valid NT attributes/access masks, self-relative creation descriptors, and semantic full-control validation. Equal-or-stronger Windows fixture adaptations retain no-delete handles instead of weakening ACL/identity checks. The final native three-package gate passes. |
| Important 3 — restore-test deadlock | RESOLVED | Lifecycle-stage release observes the buffered worker outcome, uses a bounded/cancellable release, reports premature completion directly, and never waits indefinitely on `done`. `TestReleaseLifecycleStageReturnsPrematureWorkerOutcome` and the lifecycle-lock terminal-cleanup test are green. |
| Important 4 — complete native package gate | RESOLVED | Workflow and remediation plan both use the exact unfiltered command `go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1`. Final logs prove all three packages executed natively. |
| Minor — non-executed gate classification | RESOLVED | Non-executed/skipped commands are `NOT_RUN` / `BLOCKED`; `FAIL` is reserved for commands that ran and failed. Historical results remain separately labelled. |
| Minor — explicit Windows StageArchive row | RESOLVED | `.artifacts/task-15/automated-gates.json` contains `windows_stagearchive_retained_capability` = PASS, bound to the final source and native backup result. |
Additional failures exposed by the required unfiltered gate were fixed without narrowing the
workflow: Windows secret-bearing archive reservation is protected before use; StageArchive shares
one retained root capability across both staged files; claim/consume transitions serialize the
complete public validation and retained-handle operation while preserving ACL, hard-link identity,
reparse rejection, and no-delete invariants.
## RED → GREEN record
### Initial RED
- Run `32122302381`:
https://github.com/mptyl/ThothII/actions/runs/32122302381
- Source: `b31b27e5845ffd3adf311429367319beaba263c7`.
- Windows job: `95665197885`.
- Result: native `safeio`/`backup` failure, including the 10-minute restore lifecycle timeout;
`authstorage` was absent from the command. This established the RED for Important 2–4 and the
required native authority.
- Cleanup failure-injection tests added for Important 1 first exposed the short-circuit behavior
before the implementation was changed.
### Final concurrency RED
- Run `32140481263`:
https://github.com/mptyl/ThothII/actions/runs/32140481263
- Source: `b48e9e9189dd0e8083db9bd0378704524e670edb`.
- Windows job: `95721724645`.
- Native results: backup PASS (`20.757s`), authstorage PASS (`104.180s`), safeio FAIL
(`63.502s`). The only failures were:
- `TestCanonicalPrivateClaimWaitsForRetainedRemoveOperation`: the concurrent claim returned
`false, unsafe file` before retained removal completed;
- `TestCanonicalPrivateClaimConsumeHasOneConcurrentWinner`: iteration 8 returned `unsafe file`.
- Diagnosis: the process mutex started below `validateClaimPaths`; a concurrent caller could fail
while reopening the retained no-delete directory before reaching the lock.
### GREEN implementation and local gates
The lock boundary was moved to the three public claim/read/remove APIs, covering validation,
relative operation, and handle close. The Unix implementation uses a no-op boundary and retains its
existing descriptor-relative semantics.
Final-source local commands passed:
```text
go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1
go test -race ./...
go vet ./...
go build -o /tmp/thothii-tht-host ./cmd/tht
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go build -o /tmp/thothii-tht-windows.exe ./cmd/tht
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go test -c ... ./internal/{safeio,backup,authstorage}
```
- Focused host package times: safeio `8.750s`, backup `8.378s`, authstorage `8.854s`.
- Race suite and vet: PASS.
- Host CLI: Mach-O arm64; Windows CLI and all three Windows test binaries: PE32+ x86-64.
- Cross-compilation remains compile-only and is not used as native proof.
## Exact-source native certification
- Run: `32141428407`
- URL: https://github.com/mptyl/ThothII/actions/runs/32141428407
- Event/status/conclusion: `workflow_dispatch` / `completed` / `failure`.
- Head SHA: `10cd66fe6a5b484a4dc569326a228c1c5484a5d4` — exact final source match.
- Windows job: `Windows clone and Compose contract`, job `95724751282`:
https://github.com/mptyl/ThothII/actions/runs/32141428407/job/95724751282
- Native step: `Run native Windows retained-capability tests` — **PASS**.
- Exact command: `go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1`.
- Native package results:
- safeio PASS (`8.230s`);
- backup PASS (`5.195s`);
- authstorage PASS (`8.383s`).
- Job conclusion: `failure` only because the following `Verify Windows clone contract` baseline
step failed with a PowerShell `ParserError` at
`scripts/test-windows-clone-contract.ps1:208`; `$remoteYaml:` is not delimited before `:`.
## Remaining branch/release blockers
| Gate | Classification | Exact outcome |
|---|---|---|
| Windows native authentication packages | PASS | All three required packages executed on final source. |
| Windows clone contract | FAIL / baseline | Executed after native PASS; PowerShell parser error at line 208. |
| LF, Compose, docs, and TypeScript | FAIL / baseline CI contract | Job `95724751205`; unified Compose passed, then `test-no-deployment-coupling-scope.sh` failed because `TMPDIR` was unset. Downstream skipped commands are `NOT_RUN` / `BLOCKED`. |
| Linux Docker deployment and rollback | FAIL / infrastructure prerequisite | Job `95724751356`; executed smoke stopped because `rg` was unavailable. Cleanup proof passed; no new image manifest was generated. |
| Native Windows Docker Desktop/WSL2 startup | NOT_RUN / BLOCKED | Job `95724752028` was skipped by workflow conditions; no Docker/WSL2 command executed. |
| Harness/Ruff/other historical baseline gates | FAIL | Retained with their recorded source and results; not rewritten as final-source proof. |
| L2, real PSD/manual acceptance, provider readiness | PENDING | Required secrets, identity/access, or provider prerequisites remain unavailable. |
The historical Docker image manifest remains bound to source
`74b062f1a737103524cbe706346cfd65f87cdfd1`; it was not reused as proof for the final source.
## Principal source commits
- `cd5f505` — exhaustive cleanup, Windows authority foundation, restore deadlock tests/fix, and
complete workflow/plan package command.
- `a0e05ad` through `b6396e6` — effective full-control DACL semantics, valid NT attributes/access,
self-relative descriptors, retained no-delete fixture ordering, and Windows installation fixture
protection.
- `824245d` — preserve existing lifecycle ACL trees instead of mutating inherited authority.
- `455fffb`, `2d1670e`, `c01482c`, `9fc1a15` — concurrent claim/consume and settled-loss handling.
- `6474118` — one retained StageArchive root capability shared across staged files.
- `feee4ee` — unified Windows path wrappers on the retained primitive.
- `b261dd4` — bounded private-root sharing contention handling.
- `b48e9e9` — deterministic retained-remove concurrency regression and claim-operation lock.
- `10cd66f` — final lock boundary includes public path validation; frozen source.
## Files changed
Source changes relative to the fix-round base:
- `.github/workflows/deployment.yml`;
- `docs/superpowers/plans/2026-08-18-thothii-authentication-remediation.md`;
- `tools/tht/internal/authstorage/storage_test.go`;
- `tools/tht/internal/backup/{create.go,create_test.go,fixture_security_unix_test.go,fixture_security_windows_test.go,preflight.go,preflight_test.go,preflight_windows_test.go,restore.go,restore_test.go}`;
- `tools/tht/internal/safeio/{claim_unix.go,claim_windows.go,claim_windows_test.go,files.go,files_test.go,private_root_windows.go,private_windows.go,private_windows_test.go}`.
Evidence/status changes are restricted to:
- `.artifacts/task-15/automated-gates.json`;
- `.superpowers/sdd/2026-08-18-thothii-authentication-remediation/task-4-report.md`;
- `.superpowers/sdd/2026-08-18-thothii-authentication-remediation/fix-round-1-report.md`;
- `.superpowers/sdd/2026-08-16-thothii-authentication/task-15-report.md`;
- `PROJECT_STATE.md`.
Machine-readable evidence SHA-256:
`5c110b7b2607693de078def441b10290c5a29024c83b7e5a0ced894b72b7507f`.
## Git and protection status
- The evidence commit contains only the five evidence/status files listed above; no source is
changed after frozen source `10cd66fe6a5b484a4dc569326a228c1c5484a5d4`.
- After the evidence commit and push, the intended status is synchronized
`feat/thoth-auth...origin/feat/thoth-auth` with only protected untracked `.playwright-cli/` and
`.thothctl/`.
- `AGENTS.md`, `CLAUDE.md`, and `docs/agents/` are untouched. No generated `tools/tht/tht` exists.
- Evidence commit SHA is reported externally after commit creation because a commit cannot contain
its own final hash.
@@ -1,133 +0,0 @@
# Final-review fix round 2 report (sanitized)
## Verdict
- Base evidence head: `0f762ad6b67675356389cc546421a1c46ad5a736`.
- Frozen source: `2a9359071257f9b8a71d36ec2bbb25b161003f81` on `feat/thoth-auth`.
- Authentication remediation: **PASS**.
- Three original remediation Important findings: **RESOLVED**.
- Fix-round-2 bounded lifecycle Important: **ADDRESSED**.
- Fix-round-2 temporary Windows diagnostics Minor: **ADDRESSED**.
- Release readiness: **FAIL** for executed unrelated baseline gates, with unavailable
external/manual gates separately **PENDING**.
- Source and evidence are separate commits. The evidence-only phase changed no source or tests and
dispatched no workflow.
## Finding disposition
| Finding | Disposition | Evidence |
|---|---|---|
| Original Important — POSIX local-registry ownership | RESOLVED | Effective-UID ownership enforcement and its Node 24 coverage remain green at their recorded source. Fix round 2 did not alter this boundary. |
| Original Important — retained-capability StageArchive lifecycle | RESOLVED | Native Windows `internal/backup` passed on the exact source, preserving the retained-root staging and cleanup coverage. |
| Original Important — handle-relative Windows claim removal | RESOLVED | Native Windows `internal/safeio` and `internal/authstorage` passed on the exact source, including retained claim/consume coverage. |
| Fix-round-2 Important — fully bounded restore lifecycle test | ADDRESSED | Gate publication and release are context-aware; stage, outcome, admission, checkpoint, and verification waits are bounded; aborts cancel, safely release, bounded-join, then assert lock-free. The deterministic withheld-gate test proves prompt timeout/cancellation, worker join, and eventual lock release. |
| Fix-round-2 Minor — temporary Windows diagnostic matrix | ADDRESSED | `windowsRelativeOpenMatrix` and its diagnostic-only call/import were removed. Owner-only DACL shape, NT access normalization, full-control, cleanup, and retained no-delete tests remain. |
The round-1 restore lifecycle finding was broadened by the scoped round-2 review: bounded release
alone was insufficient while stage publication, gate waits, and nearby outcome/admission waits
could still outlive a controller abort. The round-2 implementation closes that broader test
orchestration gap without changing production authentication semantics.
## RED → GREEN record
### RED
The deterministic withheld-gate regression was introduced first and run without relying on a
global ten-minute package timeout:
```text
go test ./internal/backup -run '^TestRestoreLifecycleCancellationJoinsWithWithheldGate$' -count=1
```
It failed in approximately `0.64s` with:
```text
cancelled restore worker did not join within the bounded deadline
```
This proved that cancellation did not yet unblock and join a worker retained at the lifecycle
gate.
### GREEN and refactor
- The gate uses a cancellation source shared by controller and worker. Both publication and
release are `select`-based and cancellation-aware.
- Shared bounded helpers cover stage, outcome, error, signal, release, and admission waits.
- Abort cleanup is ordered: cancel, cancel the controller gate when distinct, safely release a
pending gate, bounded-join the worker, then prove the lifecycle lock is free.
- Premature worker outcomes retain and surface their original error.
- The existing success, recovery, maintenance-barrier, stale-checkpoint, and verification
assertions remain active.
Final local gates on the frozen source:
```text
go test ./internal/backup -run '^(TestRestoreLifecycleCancellationJoinsWithWithheldGate|TestReleaseLifecycleStage|TestRestoreLifecycleLockExcludesCompetingTransactionsUntilTerminalCleanup|TestRestoreCannotApplyAStaleCheckpointOverAnInterleavedRestore|TestRestoreKeepsAdmissionBarrierActiveUntilVerificationCommits)$' -count=1
go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1
go test ./... -count=1
go test -race ./...
go vet ./...
go build -o /tmp/thothii-tht-host-fix-round-2 ./cmd/tht
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go test -c ./internal/safeio -o /tmp/tht-safeio-fix-round-2-windows.test.exe
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go test -c ./internal/backup -o /tmp/tht-backup-fix-round-2-windows.test.exe
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go test -c ./internal/authstorage -o /tmp/tht-authstorage-fix-round-2-windows.test.exe
GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go build -o /tmp/thothii-tht-fix-round-2-windows.exe ./cmd/tht
```
All commands passed. The final focused lifecycle run completed in `0.672s`; the full security
package run passed safeio, backup, and authstorage; race, vet, host build, Windows test-package
cross-compiles, and Windows CLI cross-compile also passed. Cross-compilation is recorded only as
compile evidence and is not used as native authority.
## Exact-source native certification
- Controller-authorized run: `32147345625` —
https://github.com/mptyl/ThothII/actions/runs/32147345625.
- Event/status/conclusion: `workflow_dispatch` / `completed` / `failure`.
- Head SHA: `2a9359071257f9b8a71d36ec2bbb25b161003f81`, exactly matching the frozen source.
- Windows job: `Windows clone and Compose contract`, job `95744249248` —
https://github.com/mptyl/ThothII/actions/runs/32147345625/job/95744249248.
- Native step: `Run native Windows retained-capability tests` — **PASS**.
- Exact unfiltered command:
`go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1`.
- Native package results:
- `internal/safeio` PASS (`22.058s`);
- `internal/backup` PASS (`7.161s`);
- `internal/authstorage` PASS (`16.088s`).
The Windows job failed only in the following baseline clone-contract step. PowerShell reported a
parser error at `scripts/test-windows-clone-contract.ps1:208` because `$remoteYaml:` is not a
delimited variable reference. This later failure does not alter the successful native Go step.
## Separate release-readiness verdict
| Gate | Classification | Exact outcome |
|---|---|---|
| Authentication remediation | PASS | Source and exact-source native three-package authority are green. |
| Windows clone contract | FAIL / baseline | Job `95744249248`; parser error at `scripts/test-windows-clone-contract.ps1:208`, after native PASS. |
| LF, Compose, docs, and TypeScript | FAIL / baseline CI contract | Job `95744249458`; unified Compose passed, then the existing unset-`TMPDIR` failure stopped the contract step. Downstream commands were skipped. |
| Linux Docker deployment and rollback | FAIL / infrastructure prerequisite | Job `95744249354`; the existing missing-`rg` prerequisite stopped the smoke before deployment. Cleanup passed and no new image manifest was generated. |
| Native Windows Docker Desktop/WSL2 startup | NOT_RUN / BLOCKED | Job `95744250450` was skipped by workflow conditions; no native Docker/WSL2 command ran. |
| L2, real PSD/manual acceptance, provider readiness | PENDING | Required secrets, identity/access, or provider prerequisites remain unavailable. |
Executed failures remain `FAIL`; skipped commands are `NOT_RUN` / `BLOCKED`; unavailable external
gates remain `PENDING`. Therefore remediation PASS does not imply release readiness PASS.
## Evidence and protection status
- Machine-readable evidence: `.artifacts/task-15/automated-gates.json`; SHA-256
`6c516db5c2064c4a4a2e5f25961b993cd4a8fe020bbbb822fbac7faa0c119599`.
- Current Task 4 report:
`.superpowers/sdd/2026-08-18-thothii-authentication-remediation/task-4-report.md`.
- Retained Task 15 report:
`.superpowers/sdd/2026-08-16-thothii-authentication/task-15-report.md`.
- Project snapshot: `PROJECT_STATE.md`.
- Historical Docker evidence remains bound to its recorded older source and is not reused as proof
for `2a9359071257f9b8a71d36ec2bbb25b161003f81`.
- `.playwright-cli/` and `.thothctl/` remain protected and untracked. No source/test file,
instruction file, workflow, or `docs/agents/` content changed in this evidence phase.
- The separate evidence commit SHA is reported after commit creation because a commit cannot
contain its own final hash.
No credentials, tokens, internal endpoints, identities, registry names, raw environments, or
browser traces are retained in this report.
@@ -1,131 +0,0 @@
# Task 4 authentication remediation recertification (sanitized)
## Fix-round-2 recertification — remediation PASS
- Exact source: `2a9359071257f9b8a71d36ec2bbb25b161003f81` on `feat/thoth-auth`.
- Authorized workflow: completed run `32147345625`,
https://github.com/mptyl/ThothII/actions/runs/32147345625, exact matching head SHA.
- Native job: `Windows clone and Compose contract`, job `95744249248`.
- Required native step: `Run native Windows retained-capability tests` — **PASS**.
- Exact unfiltered command:
`go test ./internal/safeio ./internal/backup ./internal/authstorage -count=1`.
- Package evidence: `internal/safeio` PASS (`22.058s`), `internal/backup` PASS (`7.161s`),
`internal/authstorage` PASS (`16.088s`). This includes explicit native Windows
StageArchive retained-capability and concurrent claim-consume coverage.
- The later `Verify Windows clone contract` step failed independently at
`scripts/test-windows-clone-contract.ps1:208`: PowerShell parsed `$remoteYaml:` as an invalid
variable reference. This baseline deployment-contract failure does not change the native Go
package result.
- The optional `Native Windows Docker Desktop/WSL2 startup` job was skipped by workflow
conditions. It is `NOT_RUN` / `BLOCKED`, because no Docker Desktop/WSL2 command executed.
- The workflow reached `completed` with conclusion `failure`: the native authentication step is
PASS, while the later clone-contract, LF/Compose, and Linux Docker baseline steps are FAIL.
- Existing LF/Compose job `95744249458` and Linux Docker job `95744249354` failures repeated
before downstream work. Skipped commands are `NOT_RUN` / `BLOCKED`, not executed failures.
External L2/PSD/provider gates remain `PENDING`.
Finding disposition is explicit: the three original remediation Important findings remain
**RESOLVED**; the fix-round-2 lifecycle Important is **ADDRESSED**; and the temporary Windows
diagnostic-matrix Minor is **ADDRESSED**. Authentication remediation is **PASS**. This does not
change overall release readiness: executed baseline gates remain **FAIL**, while unavailable
external/manual gates remain **PENDING**.
The section below is retained as historical evidence for the pre-fix frozen source.
## Historical pre-fix result
- Frozen source under test: `b31b27e5845ffd3adf311429367319beaba263c7` on `feat/thoth-auth`.
- Freeze check: PASS. No tracked source changed during certification. The only untracked paths
retained are `.playwright-cli/` and `.thothctl/`.
- Certification window: `2026-08-18T09:26Z` to `2026-08-18T09:48:36Z` (UTC; the start marker is
minute-precision because no earlier second-level operator timestamp was captured).
- Overall result: `FAIL` / `CHANGES_REQUIRED`. The three Important findings are not closed and
authentication is not implementation-complete or release-complete.
## Local gate matrix
| Gate | Result | Sanitized evidence |
|---|---|---|
| Go focused security tests | PASS | `safeio`, `backup`, and `authstorage`; 3 packages |
| Go race/vet/host build | PASS | 18 race-tested packages; vet and host CLI build exit 0 |
| Windows amd64 cross-compile | PASS | focused safeio/backup test binaries and CLI build; compile-only |
| POSIX registry ownership | PASS | Node 24 backend suite includes local-registry ownership coverage |
| Unix StageArchive retained capability | PASS | focused safeio/backup and race coverage passed on host |
| Backend Node 24 | PASS | 76 files / 1092 tests; typecheck and build passed |
| Frontend Node 24 | PASS | 61 files / 444 tests; typecheck and build passed |
| Authentication/F1 browser smoke | PASS | Node `v24.16.0`; filtered E2E 1 passed; sentinel scan passed |
| Harness pytest | FAIL | 951 passed, 1 failed, 4 skipped, 232 subtests; `test_f4_emits_column_types` could not find `workflow.yaml` from its test cwd |
| Ruff | FAIL | 192 errors; known baseline |
| Authentication docs smoke | PASS | required-term and forbidden-word checks passed |
| Shell syntax | PASS | `bash -n scripts/*.sh` |
| Default Compose contract | FAIL | required `THT_WORKSPACE_GIT_REMOTE` was unavailable |
| Unified Compose contract | FAIL | `compose.unified.yaml` is absent from the frozen source |
| Unified Docker smoke | FAIL | workflow attempted it on the frozen SHA but stopped before deployment because `rg` was unavailable; cleanup proof passed and no new image manifest was generated |
| L2 / PSD manual / provider readiness | PENDING | required external secrets, identities/access, or provider prerequisites unavailable/not reached |
The first full backend Vitest attempt had one workspace-registry timeout. The focused test and a
fresh complete rerun passed, so the current backend result above is the fresh complete rerun.
## Native Windows authority
The authorized dispatch was bound to the frozen SHA:
- Run: `32122302381`
- URL: https://github.com/mptyl/ThothII/actions/runs/32122302381
- Head SHA: `b31b27e5845ffd3adf311429367319beaba263c7`
- Workflow conclusion: `failure`
- Job: `Windows clone and Compose contract`, job `95665197885`
- Job URL: https://github.com/mptyl/ThothII/actions/runs/32122302381/job/95665197885
- Native step: `Run native Windows retained-capability tests` — `failure`
- Executed command: `go test ./internal/safeio ./internal/backup -count=1`
- Observed focused failures include `TestRemoveCanonicalPrivateClaimRetainsParentDuringDeletion`
and `TestRemoveCanonicalPrivateClaimPreservesOrphan`.
- The backup package timed out in
`TestRestoreLifecycleLockExcludesCompetingTransactionsUntilTerminalCleanup` after `10m0s`.
- Additional backup failures included retained-staging `unsafe file` results, Windows temporary-file
cleanup reporting that a file was still in use, and fixture cases that could not read external
secret declarations. The first two categories are remediation/security-boundary failures; the
fixture declaration failures are recorded as an accompanying CI-fixture issue.
- `internal/authstorage` was not requested by the frozen workflow step and therefore has no native
Windows execution evidence. Cross-compilation does not substitute for this gate.
This native failure is the blocking gate. No source fix was attempted, and no later Docker smoke
was run locally after the failure.
## Other workflow failures
- `LF, Compose, docs, and TypeScript` (job `95665197839`) failed in
`Verify Compose and installation contracts` after the unified Compose contract itself passed.
`test-no-deployment-coupling-scope.sh` aborted on `TMPDIR: unbound variable`; this is classified
as a baseline/CI contract prerequisite, and later docs/TypeScript steps were skipped.
- `Linux Docker deployment and rollback` (job `95665197846`) failed before deployment because the
runner did not provide `rg` (`Task 13 smoke failed: rg is required`). The sanitized cleanup proof
passed and no Docker image manifest was generated. This is classified as an infrastructure
prerequisite failure, not as evidence of a remediation regression.
## Evidence and provenance
- Current machine-readable matrix: `.artifacts/task-15/automated-gates.json`; SHA-256
`6c516db5c2064c4a4a2e5f25961b993cd4a8fe020bbbb822fbac7faa0c119599`.
- Current requested report: this file (SHA-256 recorded after the evidence commit if needed for
external indexing).
- Current fix-round report:
`.superpowers/sdd/2026-08-18-thothii-authentication-remediation/fix-round-2-report.md`.
- Historical Docker image manifest: `.artifacts/task-15/unified-docker-images.json`, unchanged
because no new immutable-source Docker smoke ran. Its retained historical SHA-256 is
`9c8dec4546909fd93799dbcf374bcb3a89bc46cfe0fd482472c0cbe757ddf5b6`, bound to historical source
`74b062f1a737103524cbe706346cfd65f87cdfd1`, not to this Task 4 candidate.
- The historical Task 15 report remains provenance for earlier source SHAs; its current addendum
records this recertification separately.
No credentials, tokens, internal endpoints, provider identities, registry names, raw environments,
or browser traces are retained here.
## Separate verdicts
- Three Important findings: `CHANGES_REQUIRED`. Native Windows retained-capability authority
failed, and the frozen workflow omits the required `authstorage` package from its native command.
- Overall release readiness: `FAIL` with additional `PENDING` gates. The native Windows remediation
gate failed; the remote Docker attempt failed on a missing runner prerequisite; existing
Ruff/harness/Compose failures and external/manual prerequisites remain unresolved; and no
successful new unified Docker image evidence exists.
@@ -1,137 +0,0 @@
# Adapter Foundations final-review fix report
Date: 2026-07-11
Branch: `codex/portable-deployment`
Worktree: `/Users/mp/projects/ThothII/.worktrees/portable-deployment`
Binding findings: `.superpowers/sdd/adapter-final-review-findings.md`
## Outcome
All seven final-review findings are addressed as one coherent adapter-foundations change:
1. HTTP vector reader and writer clients are independently optional. Capabilities reflect the
configured side; writer-only new and legacy configurations build successfully for targeted
writes; search without a reader raises public `VectorReadUnavailable`.
2. `VectorHealth` now reports read/write configured and reachable state independently, preserves
side-specific errors, and reports expected/observed embedding dimensions plus compatibility.
HTTP diagnostics cover read-only, write-only, both-up, and writer-down cases. Direct health
exposes its configured expected dimension without adding schema or migration work.
3. `ThothRestDwhAdapter` accepts `DatabaseIdentityConfig`, matching its resource contract.
4. Both vector adapters reject bools, floats, zero, and negative search limits using one exact
positive-integer guard.
5. Port tests explicitly cover public exports and frozen capability records.
6. A real `tht` subprocess test proves one legacy deprecation warning per config load on stderr
while JSON stdout remains parseable and uncontaminated.
7. The adapter plan and SDD progress explicitly constrain `build_vector_loader` to transitional
bulk sync and schedule its removal/migration in the local pgvector plan. Targeted memory and
solved-question writes remain on `build_vector_store(..., require_write=True)`.
No pgvector schema or migration changes were made.
## Files changed
- `harness/tht/ports/vector.py`
- `harness/tht/ports/__init__.py`
- `harness/tht/adapters/vector/thoth_http.py`
- `harness/tht/adapters/vector/legacy_direct.py`
- `harness/tht/adapters/factory.py`
- `harness/tht/adapters/dwh/thoth_rest.py`
- `harness/tests/test_vector_port_contract.py`
- `harness/tests/test_adapter_factory.py`
- `harness/tests/test_config_resources.py`
- `harness/tests/test_config_legacy_compat.py`
- `harness/tests/test_adapter_command_regressions.py`
- `harness/tests/test_dwh_port_contract.py`
- `docs/superpowers/plans/2026-07-11-adapter-foundations.md`
- `.superpowers/sdd/progress.md`
- `.superpowers/sdd/adapter-final-fix-report.md`
## TDD and verification evidence
RED:
```text
cd harness && .venv/bin/pytest tests/test_vector_port_contract.py \
tests/test_adapter_factory.py tests/test_config_resources.py \
tests/test_config_legacy_compat.py -q
```
Result: collection failed as expected because `VectorReadUnavailable` did not exist. After the
initial implementation, the same command exposed two expected contract/test-harness corrections:
dimension mismatch makes aggregate health unhealthy, and the installed CLI entry point is `tht`
rather than `python -m tht.cli`.
GREEN, covering adapter/config/command regressions:
```text
cd harness && .venv/bin/pytest tests/test_vector_port_contract.py \
tests/test_adapter_factory.py tests/test_config_resources.py \
tests/test_config_legacy_compat.py tests/test_adapter_command_regressions.py \
tests/test_dwh_port_contract.py tests/test_memory_save_one.py \
tests/test_solved_question.py tests/test_search_similar_kinds.py \
tests/test_vector_dual_key.py -q
```
Result: `66 passed in 0.45s`.
Docker availability:
```text
docker info --format '{{.ServerVersion}}'
```
Result: `29.4.1` (available; command required Docker socket access).
Full repository-default non-L2 harness suite, with Docker available for L0 tests:
```text
cd harness && .venv/bin/pytest -q
```
Result: `433 passed, 5 deselected, 17 warnings in 9.14s`. The five deselections are the configured
L2/live-service tests. Warnings are existing legacy-workspace `FutureWarning` emissions.
Scoped lint and diff hygiene:
```text
cd harness && .venv/bin/ruff check tht/ports tht/adapters \
tests/test_vector_port_contract.py tests/test_adapter_factory.py \
tests/test_config_resources.py tests/test_config_legacy_compat.py \
tests/test_adapter_command_regressions.py tests/test_dwh_port_contract.py
git diff --check
```
Result: `All checks passed!`; `git diff --check` produced no output.
## Commit
Commit subject: `fix(adapter): close final foundation review`
The report is part of that same final commit. A Git object cannot contain its own SHA without
changing that SHA; the exact resulting commit ID is therefore recorded in the task handoff from
`git rev-parse HEAD` after creation.
## Self-review
- Reader/writer separation is preserved: search dereferences only `_reader`; hashes/upsert only
`_writer`; health probes each configured client independently and never substitutes one result
for the other.
- Writer failure contributes to aggregate `ok=False`, even when the reader succeeds.
- Dimension compatibility is derived only from configured embedding dimension and existing
`list_tables` metadata. Missing metadata remains `None`, not a guessed success/failure.
- The shared limit guard uses `type(limit) is int`, intentionally rejecting Python booleans and
numeric coercions before either adapter reaches its transport.
- Existing JSON/CLI behavior is preserved; the subprocess regression parses stdout as JSON and
counts exactly one deprecation marker on stderr.
- Scope remains adapter foundations. No vector DDL, schema initialization, or migration work was
introduced.
## Concerns / follow-up
- Write reachability uses the existing `list_tables` diagnostic on the separately authenticated
writer client. Deployments must allow that non-mutating diagnostic RPC to the writer credential;
failures are intentionally visible rather than hidden by reader success.
- Existing legacy-workspace tests emit 17 `FutureWarning`s in the full suite. This wave pins the
required production stderr behavior but does not migrate unrelated test fixtures.
- `build_vector_loader` remains transitional technical debt only for bulk sync, explicitly assigned
to `2026-07-11-local-pgvector-profile.md`.
-206
View File
@@ -1,206 +0,0 @@
# Container Packaging Task 3 Report
## Status
Implemented the multi-stage core application image, non-root runtime, pinned Pi installation,
container entrypoint, context exclusions, and an in-image health smoke test.
## TDD / Build Evidence
Initial RED:
```text
docker build -f docker/core.Dockerfile -t thothii-core:test .
ERROR: failed to build: resolve : lstat docker: no such file or directory
```
The first sandboxed attempt could not access the Docker socket; the authorized rerun reached the
builder and failed for the expected reason: the Dockerfile did not exist.
GREEN build:
```text
sh -n docker/core-entrypoint.sh docker/smoke/core-smoke.sh
docker build --progress=plain -f docker/core.Dockerfile -t thothii-core:test .
```
Result: shell syntax exited 0; Docker build exited 0. A final rebuild after tightening
`.dockerignore` also exited 0 and transferred only 17.60 kB of changed context (the initial clean
build transferred 1.02 MB).
## Runtime and Entrypoints
- Runtime user is `10001:10001` (`thoth`), never root.
- Runtime contains Node `v22.19.0` and Python `3.12.13`. Python 3.12 is intentional because the
harness declares `requires-python = ">=3.12"` and also satisfies the deployment floor of 3.11+.
- Pi is installed exactly as `@earendil-works/pi-coding-agent@0.80.3`; its build-time and runtime
version probes both reported `0.80.3`.
- `server` starts `/app/backend/dist/server.js`; `doctor` routes to `tht doctor`; `preprocess`
routes to the future-facing `tht preprocess` command; explicit `tht ...` and arbitrary CLI
arguments route to the installed `tht` binary.
- The gate extension's `typebox` runtime dependency is installed from the harness lockfile.
## Smoke and Diagnostic Results
```text
docker run --rm thothii-core:test doctor
config: error - configuration is invalid or unreadable
data_root: ok
```
Result: expected exit 1 for absent mounted workspace configuration, with no traceback and no
secret-bearing validation detail.
```text
docker run --rm --entrypoint /app/docker/smoke/core-smoke.sh thothii-core:test
backend listening on http://127.0.0.1:8787
v22.19.0
Python 3.12.13
core smoke: ok
```
Result: exit 0. The script asserted non-root execution, `tht --help`, `pi --version`, runtime
version floors, and `GET /health` through curl. Fastify's returned display address was loopback;
the inspected container environment is `HOST=0.0.0.0`, and the compiled server passes that value
to `app.listen`.
```text
docker run --rm thothii-core:test tht --version
0.1.0
```
Result: arbitrary `tht` entrypoint exited 0.
An explicit runtime assertion checked UID 10001, exact Node and Pi versions, Python 3.11+, and the
absence of `/app/harness/.env` and `/app/harness/workspaces`; it exited 0.
## Image Size and Containment Inspection
```text
docker image inspect thothii-core:test --format '{{.Size}} {{json .Config.User}} {{json .Config.Env}}'
221419008 "10001:10001" [...runtime paths and version metadata only...]
```
Image size: **221,419,008 bytes** (about 211.2 MiB).
`docker history --no-trunc thothii-core:test` was inspected. It contains only Dockerfile commands,
the pinned public package name/version, base-image metadata, and non-sensitive runtime variables;
no credentials or customer paths were found. An in-image filename scan found only
`/app/harness/.pi/settings.json` among `.env`, key/certificate, and settings-name candidates; that
tracked Pi file contains theme/startup preferences, not secrets. The build asserts `.env` and
workspace directories are absent.
`.dockerignore` excludes VCS/agent state, all environment files except examples, package-manager
credential files, SSH/private-key and certificate formats, local virtualenvs/node_modules/caches,
backend runtime data, customer workspaces, sessions, artifacts, indexes, corpus, and deployment
mount content.
## Self-review
- `git diff --check` is clean.
- Entrypoint processes use `exec`, preserving container signal handling.
- Backend production dependencies are pruned; TypeScript build tools remain in the build stage.
- The writable `/data` root is owned by UID 10001; application payload remains root-owned and
read-only to the runtime user.
- CA certificates and curl are present for HTTPS integrations and health probing.
- No existing source, customer workspace, secret, or unrelated progress-ledger change is included
in the task commit.
## Concerns
- The `tht preprocess` command is deliberately a future-facing routing contract; its CLI group is
scheduled in the Evidence/preprocessing plan and is not implemented in the current harness.
- Python dependencies are range-resolved because the existing harness has no Python lockfile. The
Pi package, Node runtime, and package-lock-backed Node dependency sets are pinned/reproducible.
- The image was built and smoked on Docker Desktop arm64. The chosen official multi-arch base
images and Pi package are architecture-neutral at the package level, but amd64 still needs a CI
build/smoke before being advertised as verified.
## Reproducibility Review Fix
The original image pinned Pi's direct version in the Dockerfile but resolved its transitives at
build time, and pip resolved all harness dependencies from ranges. Both paths now consume committed
locks.
### Lock generation
Pi uses the minimal `docker/pi-runtime/package.json` and its committed npm v3 lock. It was generated
with:
```text
npm install --package-lock-only --ignore-scripts --no-audit --no-fund \
--prefix docker/pi-runtime
```
The package manifest specifies exact `@earendil-works/pi-coding-agent` version `0.80.3`; a lock
inspection confirmed that same resolved package version. Docker installs it with:
```text
npm ci --omit=dev --ignore-scripts --no-audit --no-fund
```
The Python lock was generated directly from the harness production metadata plus one explicit,
pinned PEP 517 build-backend input—not from a host `pip freeze`:
```text
uv pip compile harness/pyproject.toml docker/python-runtime/build-requirements.in \
--universal \
--python-version 3.12 \
--no-emit-package tht \
--generate-hashes \
--custom-compile-command \
'uv pip compile harness/pyproject.toml docker/python-runtime/build-requirements.in --universal --python-version 3.12 --no-emit-package tht --generate-hashes --output-file docker/python-runtime/requirements.lock' \
--output-file docker/python-runtime/requirements.lock
```
`pytest`, `ruff`, and `testcontainers` are absent. All production direct and transitive packages
are exact and hashed. `setuptools==80.9.0` is explicit so the local harness install can use
`--no-build-isolation` without an unpinned build-time resolution. Refresh instructions are in
`docker/LOCKS.md`.
### No-cache rebuild and verification
Final build command:
```text
docker build --no-cache -f docker/core.Dockerfile -t thothii-core:test .
```
Result: exit 0. The logs showed Pi `0.80.3`, Node `v22.19.0`, a hash-enforced Python dependency
install, explicit `setuptools==80.9.0`, and a non-isolated local `tht` wheel build. No isolated
build-dependency download occurred.
Fresh runtime checks:
```text
docker run --rm --entrypoint /app/docker/smoke/core-smoke.sh thothii-core:test
backend listening on http://127.0.0.1:8787
v22.19.0
Python 3.12.13
core smoke: ok
docker run --rm thothii-core:test tht --version
0.1.0
/opt/venv/bin/pip check
No broken requirements found.
```
An in-container package inspection reconfirmed Pi `0.80.3`. Non-root UID, runtime version floors,
doctor's expected concise exit 1/no traceback, `/health`, and arbitrary `tht` routing all passed.
The full filename containment scan found no `.env`, PEM, private-key, P12, or PFX file in `/app`;
`/app/harness/workspaces` remains absent. Image environment and `docker history --no-trunc` were
re-inspected and contain only public package/build commands and non-sensitive runtime metadata.
Final locked image size:
```text
220003986 10001:10001
```
That is **220,003,986 bytes** (about 209.8 MiB), 1,415,022 bytes smaller than the original image.
Remaining concern: the universal lock is resolved for Python 3.12 and includes hashes/markers for
all supported platforms, but only Linux arm64 has been built and smoked locally; amd64 remains a CI
verification gate.
@@ -1,82 +0,0 @@
# Container Packaging Task 4 Report
## Status
Implemented and verified runtime-configured frontend packaging.
## Changes
- Added the browser runtime contract `window.__THOTHII_CONFIG__.backendBaseUrl`.
- Loaded `/config.js` before the Vite module entrypoint.
- Made runtime configuration take precedence while preserving `VITE_BACKEND_URL` and the
existing `http://localhost:8787` client default for development and tests.
- Added a multi-stage frontend image that builds with Node and serves static assets as
unprivileged UID/GID `101:101` with nginx on port 8080.
- Added startup-time `BACKEND_BASE_URL` substitution (default `/api`).
- Added `/api/` reverse proxying to `core:8787`, SPA fallback, no-cache runtime config,
and SSE-safe proxy settings (`proxy_buffering off`, `proxy_cache off`, one-hour read timeout).
## TDD evidence
- RED: `npx vitest run src/api/runtime-config.test.ts` failed because
`./runtime-config` did not exist.
- GREEN: targeted runtime config suite passed (3 tests after preserving the legacy client
default).
## Verification
- `cd frontend && npx vitest run --reporter=dot && npx tsc -b && npm run build` — exit 0
(40 test files, 185 tests; TypeScript and Vite production build passed).
- `docker build -f docker/frontend.Dockerfile -t thothii-frontend:test .` — success.
- Image metadata reports `USER 101:101`.
- Two-container isolated-network smoke:
- `/config.js` returned `window.__THOTHII_CONFIG__ = { backendBaseUrl: "/api" };`
- `/api/health` proxied to the core image and returned `{"status":"ok"}`.
- an unknown nested route returned the SPA `index.html`.
- active nginx config contained `proxy_buffering off`, `proxy_cache off`, and
`proxy_read_timeout 1h`.
- `/config.js` returned `Cache-Control: no-store`.
- `sh -n docker/frontend-entrypoint.sh` and `git diff --check` — exit 0.
## Secret-leakage inspection
- `.dockerignore` excludes `.env*` (except examples), credentials/key formats, dependency
trees, build outputs, backend data, and deployment data.
- The runtime web root contained no `.env*`, `.pem`, `.key`, `.p12`, or `.pfx` files.
- Image history contained build/package instructions only; no secret build arguments or
credential values were introduced by this task.
## Self-review / concerns
- nginx resolves the `core` hostname at startup, matching the planned Compose service name;
standalone runs therefore need a reachable network alias named `core`.
- Existing frontend test warnings (React refs/act, MSW unmatched incidental requests, Vite
chunk-size warnings) remain; they did not fail the requested gates and are unrelated to
this task.
- `.superpowers/sdd/progress.md` was already modified by the orchestrator and was intentionally
excluded from this task's commit.
## P1 review fixes
Follow-up commit work addressed both review findings:
- Runtime configuration is now produced with `jq -cn --arg`, so `BACKEND_BASE_URL` is encoded
by a real JSON serializer rather than interpolated into JavaScript by `sed`.
- The image includes `frontend-config-smoke`, which strips only the fixed assignment wrapper,
parses the remaining JSON with `jq`, requires exactly the `backendBaseUrl` key, and compares
the decoded value to the environment input.
- The hostile smoke passed with quotes, backslashes, a literal newline, ampersand, pipe, and
`"; globalThis.PWNED=true; //` in the value. A breakout would leave non-JSON trailing input
and fail parsing.
- Added `joinBackendPath`, shared by API fetch and EventSource creation. It removes duplicate
boundary slashes for relative and absolute bases while keeping empty and `/` bases rooted.
Follow-up verification:
- RED: six join cases failed with `joinBackendPath is not a function` before implementation.
- Targeted: runtime config, API client, and EventSource suites — 14 tests passed.
- Full frontend gate — exit 0 (40 test files, 191 tests, TypeScript, Vite build).
- Rebuilt `thothii-frontend:test` successfully.
- Hostile config image smoke — `frontend runtime config smoke: ok`.
- Rebuilt two-container smoke — default `/api` config, proxied `/api/health`, SPA fallback,
and SSE-safe nginx directives all passed.
@@ -1,98 +0,0 @@
# Evidence / Preprocessing Task 1 Report
## Outcome
Implemented the additive Evidence source port and canonical corpus records. Existing evidence,
search, vector, and session runtime code is unchanged.
## Contract
- `EvidenceSource` is a runtime-checkable protocol with `discover` and `acquire` operations.
- `SourceObject` and `AcquiredDocument` are frozen, reject extra fields, use independent metadata
defaults, and restrict metadata to Pydantic `JsonValue` values.
- `CanonicalDocument`, `CanonicalChunk`, and `CorpusManifest` are frozen and reject extra fields.
- Provenance includes stable source IDs, canonical URIs, fingerprints, modification time, and
content hashes.
- Pipeline versions are recorded on documents, chunks, and manifests. Manifests also carry schema
version, optional publish ID/vector generation, and paired embedding model/dimension fields.
- Credential-like metadata keys are rejected recursively. Credentials are not model fields and
therefore cannot enter serialized canonical artifacts through extras.
## TDD evidence
The initial focused run failed during collection because `tht.ports.evidence` and `tht.corpus`
did not exist. After implementation, the focused suite passed.
## Verification
- Focused models/protocol tests: 13 passed.
- Harness excluding Docker-backed L0 and the network-dependent wheel packaging test: 444 passed,
5 deselected.
- Focused Ruff: passed.
- Full-repository Ruff remains blocked by 34 pre-existing findings outside the task files.
- An unrestricted `pytest -q` attempt reached 453 passed and 5 deselected, but reported 47 Docker
setup errors plus 4 Docker parity failures because the sandbox cannot access the Docker socket;
the wheel packaging test also failed because its isolated `uv build` needs unavailable network.
## Concerns / follow-up
- Pydantic's `frozen=True` prevents model field reassignment but does not recursively freeze list
and dict contents. `default_factory` prevents shared mutable defaults. Later pipeline stages should
treat these value objects as immutable and construct replacements rather than mutate collections.
- The adapter and normalization tasks should preserve the credential-free boundary by passing only
these records beyond acquisition.
## Review hardening follow-up
All six binding review areas were addressed in a separate TDD pass:
- JSON metadata is recursively converted to immutable `FrozenDict`/tuple values while retaining
stable object/array JSON serialization. Manifest document and chunk collections are tuples.
- Secret-key matching now normalizes camelCase and punctuation. It rejects credential-specific
names (passwords, API keys, access/refresh tokens, client/private keys, session cookies and
authorization) recursively, while deliberate benign labels such as generic `token` and `secret`
remain valid.
- Canonical URIs require a scheme and reject userinfo or credential-bearing query parameters.
- Namespaced IDs, SHA-256 content hashes, timezone-aware UTC timestamps, embedding/vector
compatibility, unique IDs, chunk referential/provenance integrity, contiguous per-document
ordinals and pipeline-version consistency are validated. Nested Pydantic instances are always
revalidated so `model_copy(update=...)` cannot bypass a manifest boundary.
- Acquired arbitrary bytes have explicit base64 JSON encoding and validation, covered by a JSON
round-trip test.
- `EvidenceSourceError` classifies transient/retryable versus permanent failures and exposes only
recursively immutable, credential-screened JSON details.
Follow-up verification:
- Focused contract suite: 39 passed.
- Focused Ruff: passed.
- Harness excluding Docker-backed L0 and the network-dependent wheel packaging test: 470 passed,
5 deselected.
- Fresh unrestricted harness attempt: 479 passed, 5 deselected; the same environmental boundary
remains (47 Docker socket setup errors, four Docker parity failures, one isolated `uv build`
network failure).
## Final blocker follow-up
The remaining four contract blockers were closed in a third TDD cycle:
- `EvidenceSourceError` now always exposes the fixed public message/`args` value `evidence source
operation failed`; caller diagnostics are not retained. Category, details and args cannot be
reassigned, details remain recursively frozen and credential-screened, and an original exception
is available only when callers use standard exception chaining.
- Canonical document/chunk provenance stores only URI scheme, authority and path. Userinfo is
rejected; query strings and fragments are removed unconditionally, including AWS `X-Amz-*`, SAS
`sig`, and fragment token material.
- Binding model bases override Pydantic's unchecked `model_copy(update=...)`: merged values always
pass full field/model validation, so invalid copied records and top-level manifests fail.
- A canonical document/chunk `content_hash` must equal SHA-256 of the exact stored text encoded as
UTF-8. This establishes the normalization boundary explicitly: line-ending/frontmatter/text
normalization happens before model construction; the canonical models never rewrite content.
Final follow-up verification:
- Focused contract suite: 45 passed.
- Focused Ruff: passed.
- Harness excluding Docker-backed L0 and network-dependent packaging: 476 passed, 5 deselected.
- Fresh unrestricted harness attempt: 486 passed, 5 deselected, with the unchanged environmental
failures (47 Docker setup errors, four Docker parity failures, one isolated `uv build` failure).
@@ -1,65 +0,0 @@
# Evidence Task 2 Report
## Status
Implemented filesystem and explicit-manifest HTTP Evidence source adapters, typed source
configuration with legacy compatibility, and factory construction.
## Delivered behavior
- Filesystem discovery is deterministic and rooted at a strict canonical directory.
- Symlink/path escapes are rejected before content is exposed.
- Discovery hashing and acquisition reads enforce a configurable byte limit.
- Filesystem fingerprints are content SHA-256 values; stable IDs derive from relative paths.
- HTTP accepts only explicit `http`/`https` manifest entries and keeps transport URLs private.
- HTTP provenance strips query strings/fragments, while config and adapter representations hide
signed or secret-bearing transport URLs.
- HTTP acquisition uses separate connect/read timeouts, streaming byte limits, bounded redirects,
private redirect rejection, and safe transient/permanent error classification.
- HTTP fingerprints prefer a deterministic ETag digest, then Last-Modified, then content SHA-256.
- `build_evidence_sources(cfg)` supports both typed `evidence.sources` entries and the legacy
`source_root` plus `evidence_dir` filesystem configuration.
## TDD and verification
- RED: focused tests initially failed during collection because the adapter package did not exist.
- GREEN: `15 passed` for filesystem, HTTP, and resource-config tests.
- Full harness: `548 passed, 5 deselected`.
- Changed-file Ruff: clean.
- Repository-wide Ruff remains non-clean due to 34 pre-existing findings in unrelated test files;
no unrelated lint files were modified.
## Notes
The approved `SourceObject` namespace grammar does not permit raw quoted ETags such as
`etag:"abc"`. The adapter therefore uses `etag:<sha256-of-opaque-etag>`: it preserves ETag-based
change identity without weakening the canonical contract or exposing validator contents.
## Review hardening follow-up
Four review findings were closed in a separate follow-up commit:
- Filesystem access now anchors a persistent descriptor at the canonical root and walks each
component with `openat` semantics (`dir_fd`, `O_NOFOLLOW`, and `O_DIRECTORY`). The regular-file
check, bounded read, metadata, and hash all use the opened descriptor. Acquisition reopens by
the same path-safe mechanism and rejects a changed fingerprint. Deterministic tests swap both a
leaf and an ancestor to symlinks at open time.
- HTTP network policy defaults to public hosts only. Initial URLs and every redirect reject
userinfo, mixed public/private IPv4/IPv6 answers fail closed, and the connected peer must be a
public member of the previously validated DNS answer set before any body bytes are consumed.
Explicit `allow_private_hosts: true` is required for trusted private deployments and local tests.
- Every HTTP response is closed in a `finally` block, including redirects, status failures,
policy failures, oversized bodies, and mid-stream exceptions.
- ETag and Last-Modified values remain adapter-internal. Repeated discovery and acquisition send
conditional headers; a 304 reuses only previously verified cached bytes and identity. The LRU
content cache has an explicit byte bound (`max_cache_bytes`). Validators are not forwarded
across redirect origins.
### Conditional cache binding correction
The conditional cache now binds bytes and validators to both the canonical provenance key and the
exact final effective representation URL. Redirect traversal recomputes request headers per hop:
validators are sent only when that exact URL matches the cached final URL, never merely because a
redirect retains an origin. A same-origin path change therefore downloads and replaces the body.
The adapter accepts 304 only when the exact request carried a bound ETag or Last-Modified validator;
unsolicited and cross-origin 304 responses are permanent protocol errors.
@@ -1,59 +0,0 @@
# Evidence Task 3 — deterministic normalization and chunking
## Outcome
- Added pure `normalize(acquired, pipeline_version)` and `chunk(document, policy)` transforms.
- Normalization enforces UTF-8 (including UTF-8 BOM), a 10 MiB input ceiling, LF line endings,
NFC Unicode, safe YAML frontmatter extraction, canonical provenance URIs, and hashes the exact
canonical UTF-8 text stored on the document.
- Undecodable, unsupported-charset, oversized, and invalid-frontmatter inputs fail explicitly;
byte content is never truncated.
- Chunking uses a versioned immutable policy, paragraph/word boundaries with deterministic
character-count hard splits for long tokens, contiguous ordinals, provenance metadata, exact
per-chunk hashes, and IDs derived from document hash + ordinal + policy version.
- Empty documents produce no chunks. Non-ASCII, CRLF equivalence, repeatability, policy changes,
duplicate-content ordinal collisions, and max-character limits are covered by tests.
## TDD evidence
- Initial focused test run failed during collection because both transform modules were absent.
- The EOF-frontmatter edge test was separately observed failing before its implementation.
- Final focused verification: `12 passed`.
## Verification
- `cd harness && .venv/bin/pytest tests/test_corpus_normalize.py tests/test_corpus_chunk.py -q`
— **12 passed**.
- `cd harness && .venv/bin/pytest -q` — **573 passed, 5 deselected**. The sandboxed attempt could
not access Docker; the approved rerun with local Docker access passed.
- Targeted Ruff over all four implementation/test files — **clean**.
- Full `cd harness && .venv/bin/ruff check .` — reports **34 pre-existing errors** in unrelated
legacy tests (unused imports and existing E702 semicolon lines); none are in Task 3 files.
## Concerns
- The 10 MiB normalization ceiling is deliberately explicit and independent of adapter download
limits. If deployment policy needs a different ceiling, it should become a versioned pipeline
configuration before ingestion is wired.
- Character limits use Python Unicode code points (`len`), not UTF-8 bytes or tokenizer tokens;
this is recorded in the chunk-policy metadata and tested with non-ASCII content.
## Review hardening follow-up
- Chunk IDs now bind the canonical document identity, document content hash, ordinal, chunk hash,
and a canonical SHA-256 fingerprint of every `ChunkPolicy` field. Identical content in separate
documents and same-version policies with different limits cannot collide.
- Boundary-aware slicing now retains separators in the slices. Concatenating every chunk exactly
reconstructs the canonical document for repeated spaces, tabs, blank lines, Markdown hard
breaks, fenced code, whitespace-only input, Unicode, and overlong tokens; every slice remains
within `max_chars`.
- Frontmatter uses a bounded `SafeLoader` variant: duplicate keys, anchors/aliases, structures
deeper than 20 nodes, and documents larger than 1000 composed nodes are rejected. YAML parse,
JSON type, credential-safety, and resulting canonical-model errors attributable to frontmatter
map to `PermanentNormalizationError(reason="invalid_frontmatter")`; invalid pipeline policy
remains a programmer-facing `ValueError`.
- Follow-up TDD evidence: the expanded focused suite first reported 11 expected failures against
the prior implementation, then passed **45/45** across normalization, chunking, and manifest
invariants.
- Follow-up full verification: **586 passed, 5 deselected**. Targeted Ruff is clean. Full Ruff
continues to report the same **34 unrelated pre-existing** violations in legacy tests.
-107
View File
@@ -1,107 +0,0 @@
# Evidence Task 4 — shared job envelope
Status: complete
## Delivered
- Immutable `JobSpec`, `JobRun`, `JobReport`, per-stage state, sanitized error, and UTC
timestamp records.
- `run_job(spec, stages)` with a durable checkpoint at job start, before and after every stage,
and at terminal state. Successful stages are skipped when a prior run is resumed.
- Atomic JSON checkpoint/report replacement using a unique same-directory temporary file,
file `fsync`, atomic `os.replace`, and parent-directory `fsync`.
- Public reports contain fixed operational fields only. Workspace paths, stage return values,
exception messages, source content, credentials, and arbitrary metadata are not serialized.
- `WorkspaceJobLock` uses non-blocking kernel `flock` on a stable workspace/job-specific inode.
Locks are released by the kernel on process exit; lock files are never removed based on PID,
avoiding stale-lock and PID-reuse deletion races. Evidence and DWH use distinct lock files.
- Dry-run intent is immutable in the spec/report and exposed to every stage through `JobContext`.
## TDD evidence
Initial focused collection failed because `tht.jobs` did not exist. Tests then drove:
- failure, sanitized reporting, resume, and idempotent successful-stage skipping;
- corrupt-checkpoint refusal before stage execution;
- JSON schema and path/secret/PII exclusion;
- dry-run propagation and ordered aware timestamps;
- multiprocessing exclusion, distinct Evidence/DWH jobs, traversal rejection, and recovery after
a lock-owning process crashes.
Final focused result:
```text
11 passed in 0.42s
```
## Verification
```text
cd harness && .venv/bin/pytest -q
597 passed, 5 deselected, 17 warnings in 28.45s
cd harness && .venv/bin/ruff check tht/jobs tests/test_job_runner.py tests/test_job_locking.py
All checks passed!
```
The full Ruff invocation was also run. It reports 34 pre-existing violations in unrelated legacy
tests; no Task 4 file is among them. L2 tests remain deselected by the repository configuration.
## Operational notes
- `fcntl.flock` intentionally targets the supported Linux/macOS deployment environments; it is not
a Windows locking implementation.
- The envelope does not publish or mutate an active corpus. Later pipeline stages must use
`JobContext.run_dir` for staging and perform their own final atomic publish only after validation.
- A dry run is an execution mode foundation: the runner exposes and records it; individual stages
remain responsible for suppressing external mutations.
## Review hardening follow-up
Four post-implementation findings were fixed test-first:
1. Resume compatibility is now a canonical SHA-256 fingerprint over checkpoint schema version,
hashed workspace identity, job type, dry-run mode, explicit spec/pipeline versions,
configuration/input fingerprints, and the exact ordered explicit `stage_ids`. Any insertion,
removal, reorder, mode, identity, version, config, or input change rejects resume before a stage
executes. Omitting `resume_run_id` remains the explicit safe path for a new run.
2. Lock traversal now uses directory file descriptors with `O_DIRECTORY` and `O_NOFOLLOW`.
Lock files use `O_NOFOLLOW | O_CLOEXEC`; `fstat` requires a regular file owned by the current
UID with one link, and permissions are forced to `0600` (`0700` for private directories).
Pre-existing lock-file and lock-directory symlinks are rejected.
3. Stage failures now serialize only the fixed safe tuple `internal` / `stage_exception` /
`stage execution failed`. Neither exception class names nor messages are inspected for output;
a hostile exception-name/message regression test proves a terminal failed report is retained.
4. Job/run directory creation is no-follow, owner-checked, private, and durable. Each newly created
parent is fsynced, the run directory is fsynced before the first atomic file write, and the
existing file-fsync → replace → directory-fsync ordering has an explicit regression test.
Follow-up verification:
```text
focused job/lock suite: 27 passed in 0.45s
full harness suite: 613 passed, 5 deselected, 17 warnings in 29.65s
Task 4 scoped Ruff: All checks passed
```
Repository-wide Ruff continues to report the same 34 unrelated pre-existing legacy-test findings.
## Final resume-integrity fix
Resume is now read-only until the source checkpoint proves trustworthy. The runner loads the source
before allocating a new run ID or directory, validates the exact stage state/timestamp/error ledger,
rejects duplicate stage identifiers, and recomputes compatibility from every persisted compatibility
field plus the exact ordered persisted stage IDs. It first requires the stored fingerprint to match
that recomputation, then compares the trusted recomputation with the requested job fingerprint.
Valid-JSON tampering tests cover removed, inserted/duplicated, reordered, and substituted stages;
input-field and stored-fingerprint changes; and invalid stage-state shapes. Every rejection occurs
before stage execution and asserts that the runs directory contains no orphan allocation.
Final verification:
```text
focused job/lock suite: 34 passed in 0.56s
full harness suite: 620 passed, 5 deselected, 17 warnings in 27.42s
Task 4 scoped Ruff: All checks passed
```
@@ -1,82 +0,0 @@
# Evidence Task 5 report
## Outcome
Implemented an incremental Evidence corpus pipeline with immutable materialized generations,
generation-scoped vector records, and an fsynced atomic `ACTIVE` pointer. Runtime Evidence
artifact lookup reads the active canonical manifest and keeps a legacy source-tree fallback only
when no corpus has been published.
The CLI is available as `tht preprocess evidence [--dry-run] [--resume RUN_ID] [--json]`.
JSON success and failure output is pristine and failure details are sanitized.
## Safety and failure model
- A workspace writer lock serializes preprocess writers; readers never take the lock.
- Generation directories, manifests, materialized files, locks, and `ACTIVE` reject symlink/path
escape cases and use owner-only durable writes.
- Vector records use generation-specific keys and metadata. The active manifest maps each active
document to its valid vector generation, allowing unchanged documents to retain their vectors.
- Runtime retrieval admits only active document IDs and their manifest-selected generations.
Removed documents and partial writes from failed generations are therefore unreachable.
- Embedding count and dimension checks occur before vector upsert; vector write count is checked
before staging/publish. Any failure leaves `ACTIVE` unchanged.
- Dry runs perform discovery/fingerprint planning only and never acquire, embed, write vectors, or
publish. Fully unchanged runs return the active generation without creating a replacement.
- Resume can safely retry idempotent generation-scoped upserts and publish an already staged,
compatibility-checked generation after a crash between staging and pointer replacement.
## TDD evidence
Initial focused collection failed because `tht.corpus.pipeline` and `tht.corpus.store` did not
exist. The implemented suite covers incremental skips, removals, model/policy rebuilds, acquire and
partial-vector failures, dry-run isolation, dimension validation, atomic reader snapshots, pointer
validation, symlink defense, and pristine CLI JSON.
Fresh focused verification:
```text
18 passed, 3 warnings in 0.39s
```
Command:
```text
.venv/bin/pytest tests/test_corpus_pipeline.py tests/test_corpus_publish.py \
tests/test_preprocess_cli.py tests/test_search_pack.py tests/test_session_documents.py -q
```
Scoped Ruff: `All checks passed!`
Broader non-Docker/non-packaging run reached `560 passed, 5 deselected`; ten pre-existing HTTP
adapter tests could not bind localhost under the sandbox. The complete suite reached `570 passed,
5 deselected`, with the remaining failures/errors caused by denied Docker socket, localhost bind,
and offline wheel-build access. No task-focused test failed.
## Remaining operational gate
Live pgvector integration needs Docker or an authorized local pgvector endpoint. The compensation
strategy is logical isolation rather than destructive cleanup because the shared `VectorStore`
port intentionally exposes no delete/transaction API; unreachable failed generations can be
garbage-collected by a future maintenance job.
## Review integration wave
Added an enforceable `metadata_filter` vector-port contract and capability flags. Direct pgvector
places exact Evidence generation/document predicates in SQL before `LIMIT`; HTTP sends the same
filter to the RPC and deliberately does not use the legacy 404 fallback. The reader RPC script now
validates and applies that filter. Normal Evidence search and search-pack use an ACTIVE-aware
searcher that groups active documents by generation, executes complete server-filtered searches,
and merges the results.
Added exact-generation Evidence cleanup to direct and HTTP writers plus the allowlisted writer RPC.
Pipeline failures compensate both staged filesystem state and vector writes; cleanup failures stay
sanitized and ACTIVE filtering remains the exposure boundary. Corpus-present session artifact
resolution now fails closed on corrupt/missing ACTIVE rather than falling through to source files.
Focused review-wave verification: 45 passed, scoped Ruff clean. A mocked REST regression proves
the exact filter payload and fail-closed legacy 404 behavior.
Still outstanding from the expanded review request: Task-4 JobRunner stage-by-stage integration,
published-generation retention/garbage collection, same-fd `dirfd` materialized-file reads, and
live local pgvector integration could not be completed in this wave.
@@ -1,94 +0,0 @@
# Evidence Task 5B implementation report
## Status
Integrated Evidence preprocessing with the Task 4 `JobRunner`. The CLI now accepts only a
32-character JobRunner run ID for `--resume`; generation IDs remain outputs. Runs persist the
exact ordered stages `discover`, `acquire_normalize_chunk`, `embed`, `vector_upsert`,
`stage_validate`, `publish`, and `retention_cleanup`.
Successful-stage artifacts are copied into the new resume run before execution, allowing later
stages to continue without rediscovery, acquisition, normalization, chunking, or embedding.
Job compatibility includes workspace, configuration, discovered-input, pipeline, embedding, and
chunk-policy fingerprints. Generation-specific filesystem/vector compensation is retained, and a
compensated generation is rotated before retry. `ACTIVE` is mutated only by `publish`.
Dry-run executes discovery/planning and makes every side-effecting stage a no-op. JSON output is
pristine and includes the JobRunner `run_id`, `resumed_from`, generation, plan, and publish status.
## TDD evidence
- RED: run-ID rejection and resume-artifact tests failed because generation IDs reached
configuration and resume runs had empty artifact directories.
- GREEN: the two regression tests passed after strict CLI validation and durable artifact carryover.
- Added pipeline job-plan and dry-run counting-fake coverage; both passed.
## Fresh verification
- Focused integration/search suite: `62 passed, 4 warnings`.
- Available harness suite excluding sandbox-blocked Docker, loopback HTTP-server, and networked
wheel-build tests: `559 passed, 5 deselected, 18 warnings`.
- Scoped Ruff: `All checks passed!`.
- `git diff --check`: clean.
## Environment limitations and concerns
The literal full harness invocation cannot complete in the managed sandbox: Docker socket access,
loopback HTTP test servers, and the `uv build` dependency resolution path are denied. It reached
`575 passed, 5 deselected` before those environment errors. The available-suite rerun above is
green.
One pre-existing Pydantic serialization warning is exposed by the new end-to-end job test when
canonical metadata contains frozen tuple values; it does not contaminate CLI stdout. Retention is
an explicit stable no-op until a retention policy is configured.
## Review fix wave — crash consistency and artifact integrity
Addressed all five follow-up findings:
- `JobRunner` now supports a test-only post-call/pre-checkpoint fault hook. Each stage seals a
canonical artifact manifest containing required flat filenames, SHA-256, byte size, producer
stage, and the full spec compatibility fingerprint. Resume validates the checkpoint and every
sealed artifact before allocating/copying a new run, rejecting missing, tampered, extra, nested,
or symlinked state. A sealed `running` stage is promoted after a simulated process crash; a
sealed `failed` stage is deliberately retried.
- Vector intent (exact record IDs and content hashes) is sealed before upsert. Execution reconciles
`existing_hashes` and writes only missing/mismatched rows. Crash-after-effect tests prove no
duplicate acquire, embed, or vector upsert.
- Raw upsert, stage, recovery-upsert, recovery-stage, and publish exceptions compensate the exact
generation. Compensation markers survive failed checkpoints; resume rotates the generation,
refreshes generation-bound artifacts, reconciles vectors, and stages idempotently.
- `CorpusStore.publish` is idempotent and failure-atomic. If replace succeeds but directory fsync
fails, it restores the previous `ACTIVE` value (or removes a newly created pointer), fsyncs the
rollback, and re-raises. Pipeline cleanup refuses to discard a generation referenced by ACTIVE.
- Added crash/resume coverage after all seven ordered stages; corrupt/missing plan, manifest, and
embeddings; unsafe extra paths; nonexistent run IDs; raw vector/stage failures; and post-replace
ACTIVE rollback.
Fresh fix-wave verification:
- Focused jobs/corpus/CLI/search suite: `82 passed, 17 warnings`.
- Available harness suite (same sandbox exclusions described above):
`579 passed, 5 deselected, 31 warnings`.
- Scoped Ruff and `git diff --check`: clean.
## Final P1 fix — effect state and checkpoint-bound manifest roots
- Stage checkpoints now distinguish `intent` from `completed`. Vector intent is atomically sealed
and checkpointed before upsert. A process-level `BaseException` after a partial multi-record
write leaves the stage `running/intent`; resume never promotes it and instead reconciles
`existing_hashes`, writing only the missing records. The completed state is persisted only after
reconciliation returns successfully.
- Every stage now persists its completed artifact state while still `running`, before the
post-call fault hook. The checkpoint binds the SHA-256 of canonical `artifact-manifest.json`,
effect state, exact producer stage, and exact required-file mapping. Resume validates this root
and all bindings before promotion or copying.
- Added process-interruption coverage proving the already-written vector record is not submitted
twice, remaining records are written, and publish completes only after reconciliation. Added
coordinated artifact/manifest, spec-binding, and producer-binding tamper rejection tests.
Fresh verification:
- Focused jobs/corpus/CLI/search suite: `86 passed, 18 warnings`.
- Available broad harness suite: `583 passed, 5 deselected, 32 warnings`.
- Scoped Ruff and `git diff --check`: clean.
-188
View File
@@ -1,188 +0,0 @@
# Evidence Task 5C report
## Delivered
- Added `vector.retain_published_generations` (default `3`, validation minimum `1`).
- Retention runs only after publication. It keeps ACTIVE, the newest configured generations,
and generations referenced by running or resumable failed job checkpoints.
- Cleanup deletes the exact Evidence generation from the vector store before removing its
immutable filesystem directory. Vector failures retain filesystem metadata for retry and
produce credential-free partial reports.
- Added idempotent `tht preprocess evidence gc [--dry-run] --json` reconciliation with pristine
JSON output.
- Materialized document reads now open generation/documents components with directory file
descriptors and `O_NOFOLLOW`, require a regular file owned by the process with one link, and
hash the bytes read from the same descriptor against the canonical manifest.
- HTTP generation deletion is pinned to `delete_vector_generation` with exact
table/kind/generation arguments. Legacy 404 responses fail closed with an actionable,
sanitized migration message.
## Evidence
- Focused retention, safe-read, CLI, and HTTP contract tests: `51 passed` (Docker-backed direct
parametrizations excluded from that focused invocation).
- Real Docker pgvector adapter suites: `33 passed`.
- Full harness suite, including Docker-backed tests: `668 passed, 5 deselected`.
- Changed-file Ruff: clean.
- `git diff --check`: clean.
The five deselected tests are the repository's opt-in `l2` tests requiring external services;
they are not local pgvector tests. Test output retains pre-existing Pydantic serialization and
legacy-config deprecation warnings.
## Review fix wave
- Publication is now explicit and durable (`PUBLISHED` marker). Retention candidates require a
valid generation manifest and publication marker (ACTIVE remains backward-compatible), so
staged and malformed directories neither consume retention slots nor become deletion targets.
- The policy retains ACTIVE plus exactly `N-1` newest rollback publications, ordered by durable
publication time and generation id. Running and failed-resumable JobRunner checkpoints protect
every referenced plan generation.
- `VectorStore` now exposes exact Evidence generation inventory. Direct pgvector uses a constrained
`SELECT DISTINCT` over `kind='evidence'` and `metadata.vector_generation`; HTTP uses the
allowlisted `list_evidence_generations` RPC and fails closed on legacy 404. The writer RPC SQL,
revokes, and grants are packaged in `create_vector_writer_rpc.sql`.
- Explicit GC reconciles the union of published filesystem generations and vector-only orphans,
preserving vector-before-filesystem deletion and retry semantics.
- `run_as_job` holds the same corpus writer lock across checkpoint recovery, staging, publish, and
retention. Explicit GC already uses this lock, serializing candidate snapshots with publishers.
- Session artifact consumers no longer receive the corpus source path after validation. They get
an owned, read-only copy atomically written from the bytes read and hash-validated on the same
descriptor.
Fresh verification after the fix wave: full harness `672 passed, 5 deselected`; Docker pgvector,
HTTP parity, and migration suites `43 passed`; exact direct inventory/delete integration `1 passed`;
changed-file Ruff and `git diff --check` clean.
## Final hardening verification
- Canonical generation validation is exact (`^gen:[0-9a-f]{32}$`) before HTTP/direct deletion;
malformed HTTP inventory rows fail closed rather than entering the GC candidate set.
- Added explicit protection coverage for running and failed-resumable JobRunner checkpoints, plus
a second-GC idempotence assertion for vector-only orphan reconciliation.
- Added deterministic concurrent locking coverage: a job paused after discovery retains the corpus
writer lock, explicit GC blocks, then completes after publication without deleting the active run.
- Added a descriptor-race regression: replacing the corpus pathname immediately after `read(2)`
leaves the atomically materialized session-owned copy byte-for-byte equal to the validated ACTIVE
document and its manifest hash.
Final fresh evidence: Docker pgvector/HTTP/migration suites `48 passed`; full harness `680 passed,
5 external L2 deselected`; changed-file Ruff and `git diff --check` clean.
## Integrated Task 5 dependency fixes
- GC now distinguishes filesystem retention from vector dependencies. ACTIVE and the newest
`N-1` published manifests keep their directories; every exact generation in their
`document_generations` maps remains vector-protected even after its old publication directory is
evicted. Job-protected manifests receive the same dependency treatment.
- The real four-publication Docker lifecycle now includes an unchanged document whose vectors come
from the first generation. With retention `N=2`, only the final two publication directories remain
while the first generation's vectors remain searchable from ACTIVE and survive restart/explicit GC.
- Evidence lookup is always wrapped by the ACTIVE-aware searcher. With no corpus/ACTIVE, Evidence
returns no rows and search packs cannot expose legacy vectors; non-Evidence kinds are unchanged.
- Session artifact resolution holds the corpus writer lock, snapshots the active manifest once, and
materializes bytes using that exact `manifest_id`, preventing a concurrent publish/retain-1 GC from
changing or deleting the selected source generation.
Focused unit tests, the updated real Docker lifecycle, changed-file Ruff, and `git diff --check` pass.
The final full harness invocation completed with exit code 0, including the concurrently added DWH
JobRunner tests.
## Final ACTIVE search review fixes
- `ActiveEvidenceSearcher` now treats default (`kinds=None`) and mixed-kind searches as explicit
split queries: non-Evidence kinds are queried separately, while Evidence is queried only with
ACTIVE manifest generation/document predicates applied server-side before every limit.
- Results are merged deterministically by descending similarity then stable id and truncated once
to the caller's global `top_n`. Pure non-Evidence searches retain their original delegate path.
- The corpus writer lock now covers manifest snapshot construction and all corresponding vector
queries, preventing retain-1 publication/GC from switching or deleting generations mid-search.
- Removed the public post-LIMIT `active_evidence_hits` helper; no public Evidence path performs
client filtering after limit.
Focused default/mixed/no-ACTIVE/search-pack tests pass, the real Docker pgvector lifecycle passes,
and the final full harness plus scoped Ruff/diff invocation completed with exit code 0.
## Workspace-scoped Evidence isolation
- Evidence manifests, vector metadata, and record keys now carry the stable JobRunner workspace id
derived from the configured workspace identity (config stem), never credentials or absolute paths.
- Every ACTIVE server-side predicate includes `workspace_id`. Legacy unscoped rows therefore fail
closed and cannot appear in Evidence results.
- Vector generation inventory and deletion require the workspace namespace across the port, direct
pgvector adapter, HTTP client/adapter, and allowlisted RPC SQL. Legacy unscoped RPC overloads are
explicitly dropped during migration; destructive SQL matches collection, kind, generation, and
workspace together.
- GC recovers the persisted namespace from ACTIVE for explicit/restarted cleanup and can only list
or delete that workspace's generations. Real shared-pgvector coverage proves deleting a generation
for workspace A preserves the same generation in workspace B.
- `PipelineResult.model_dump` now serializes fields explicitly instead of `dataclasses.asdict`,
avoiding deepcopy of immutable `FrozenDict` metadata while preserving pristine JSON CLI output.
Final focused verification: `89 passed` across corpus/CLI JSON, direct/HTTP parity, migrations, and
real Docker pgvector lifecycle; scoped Ruff and `git diff --check` clean. A contemporaneous full-suite
run reached unrelated Task 6 immutable-file tamper tests; those files were deliberately not changed.
## Immutable corpus/workspace binding
- A corpus root becomes bound to the workspace id persisted in its ACTIVE manifest. Job, non-job,
explicit GC, and ACTIVE search entry points compare the configured namespace before discovery,
vector access, staging, deletion, or ACTIVE mutation.
- Reusing the same paths after renaming a workspace now fails closed with a typed/sanitized message:
use a new corpus root or perform an intentional explicit rebuild. Unscoped legacy manifests also
fail this ownership check.
- Tests prove unchanged-document reuse cannot silently mix workspace A vectors into a workspace B
manifest, and that mismatched job, GC, and search paths perform no vector/filesystem mutations.
Focused workspace-binding, search-pack, preprocess JSON, and scoped Ruff/diff tests pass.
Compatibility follow-up: direct/internal `CorpusPipeline` instances now distinguish an omitted
workspace identity from an explicit config/job identity. An unbound instance adopts the persisted
ACTIVE owner (or `default` only for a brand-new direct corpus), preserving safe resume/GC tests and
the real pgvector lifecycle. Explicit config/job identities still fail closed on any mismatch. The
two reported regressions, workspace mismatch guards, real Docker lifecycle, scoped Ruff/diff, and
the full harness suite all pass.
Final fail-closed follow-up: persisted ACTIVE ownership is now validated under the corpus lock before
every configured search delegate, including default, mixed, pack, and non-Evidence-only operations.
Malformed or missing `metadata.workspace_id` is intrinsically rejected even for unbound direct
callers; source discovery, vector operations, GC, files, and ACTIVE remain untouched. Focused tests,
real Docker lifecycle, scoped Ruff/diff, and the full harness regression run pass.
Final lock/preflight follow-up: `CorpusPipeline.gc()` now acquires the corpus writer lock itself for
ownership validation through vector/filesystem cleanup. The store lock is thread-reentrant so nested
job retention is safe without weakening cross-thread/process exclusion; the CLI wrapper no longer
double-locks. Search find/pack performs locked corpus ownership preflight immediately after config
load, before DWH leasing, vector/searcher factories, embeddings, or schema work. Focused concurrency
and fail-closed tests, real Docker lifecycle, scoped Ruff/diff, and the full harness pass.
## Compact public Evidence reports
- Public `PipelineResult.model_dump()` is now a bounded operational envelope: terminal status,
run/resume/publication/generation/manifest identifiers, capped changed/unchanged/removed source
identifiers, and aggregate document/chunk counts. Full manifests, bodies, and metadata remain
internal/on disk and are never serialized to CLI stdout.
- `tht preprocess evidence` exits `1` for any durable terminal status other than `succeeded` in
both JSON and text modes. JSON stdout remains one pristine sanitized object; text mode emits one
compact stderr error without traceback, exception identity, evidence content, or credentials.
- Tests cover a real failed acquisition job, sensitive evidence content, capped thousand-item
summaries, bounded report size, and smoke-compatible changed/unchanged fields.
Focused tests and scoped Ruff/diff pass. The contemporaneous full suite reaches an unrelated Task 6
DWH snapshot fixture missing its newly required workspace identity.
### Safe result representation and exact text totals
- `PipelineResult.manifest` is explicitly excluded from dataclass representation and the custom
representation is fixed-size operational data only. It omits manifest ids, documents, chunks,
content, metadata, and errors; `str(result)` inherits the same safe representation.
- Text-mode Evidence success output reads the uncapped aggregate totals from `payload["counts"]`
rather than the intentionally capped identifier arrays.
- Regression coverage builds a thousand-document/chunk manifest containing content and
credential-like metadata secrets, checks bounded `repr`/`str`, and verifies exact totals above
the 100-item public-array cap.
Focused Evidence verification passes (`67 passed`), and scoped Ruff is clean. The full harness run
is not green in this sandbox: Docker-backed tests cannot access the daemon, wheel packaging cannot
use the restricted build environment, and concurrent Task 6 DWH binding changes currently fail two
DWH tests. None of those failures touch the Evidence files in this follow-up.
@@ -1,50 +0,0 @@
# Evidence Task 5D — Real pgvector lifecycle gate
## Status
Complete. The Docker-backed L0 gate uses one persistent `pgvector/pgvector:pg16`
database and the production migrations, direct reader/writer `PgVectorStore`,
`CorpusStore`, `CorpusPipeline.run_as_job`/JobRunner, ACTIVE Evidence retrieval,
search-pack fusion, owned session artifact copy, retention, and explicit GC.
## Lifecycle covered
- Four real corpus publications with retention set to two generations.
- A higher-similarity stale vector proves ACTIVE metadata filtering happens before LIMIT
for normal Evidence retrieval and the search-pack fusion path.
- A removed source is absent from ACTIVE retrieval and cannot be copied to a session.
- An injected process death occurs after one real committed vector upsert. Resume uses the
real run ID, preserves that record, fills the missing records, and produces no duplicate keys.
- Database engines and direct store objects are disposed/recreated before persisted ACTIVE
retrieval is checked again.
- An exact canonical vector-only orphan generation is discovered and removed by explicit GC.
- Filesystem and vector inventories converge exactly to ACTIVE plus one rollback; a second GC
is a no-op.
- Owned session artifact bytes and SHA-256 match the ACTIVE canonical document.
## Production bug found and fixed
Production migration `003_roles.sql` intentionally restricted `vector_writer`, but omitted
the privileges used by the production generation lifecycle: `SELECT(metadata)` for inventory
and `DELETE` for cleanup on `vectors.evidence`. Consequently a real job published successfully
and then failed in `retention_cleanup` on its first run.
Added versioned migration `004_evidence_generation_gc.sql` granting only those two Evidence
generation-management privileges. Runtime application code was not redesigned.
## Verification
- Target lifecycle: `1 passed` (Docker-backed).
- Full harness: `681 passed, 5 deselected`.
- Scoped Ruff: passed.
- `git diff --check`: passed.
The existing Pydantic serialization and legacy-workspace deprecation warnings remain unchanged.
## Follow-up assertion correction
The removal phase now retains the removed canonical document ID/ref before publication and
asserts both fields are absent from post-resume ACTIVE Evidence hits. It reruns the real
search-pack fusion after removal, proves active fourth-generation content is positively
returned in both paths, and proves the removed content remains absent. The owned session
artifact lookup for the retained removed ID remains empty.
@@ -1,49 +0,0 @@
# Evidence Task 6 — final fd-anchored DWH correction
All DWH generation state below `.tht-dwh` is now accessed relative to the directory descriptor
retained by the shared/exclusive generation lease. ACTIVE reads, atomic temp writes, replacement,
fsync, and rollback use `openat`/`replaceat` operations. Generation staging, validation,
reconciliation, resume checks, retention classification, and recursive deletion likewise use owned
root/generations/candidate descriptors with `O_NOFOLLOW`; locked operations no longer reopen
generation paths through `workspace_root`.
Portable reader snapshots are copied from validated generation file descriptors into private 0700
process-owned temporary directories while the shared lease is held. This avoids Linux-only
`/proc/self/fd` paths and prevents a renamed/replaced `.tht-dwh` pathname from redirecting later
schema or LSH reads. Lease-scoped copies are removed on exit and standalone snapshots are removed
at process exit.
Deterministic adversarial tests rename the DWH root after lease acquisition during ACTIVE reads,
ACTIVE publication, and retention cleanup. Each test proves the replacement tree is never read,
written, or deleted; the descriptor-pinned original either completes consistently or fails closed.
Existing owner binding, legacy rejection, crash reconciliation, resume, atomic rollback, retention,
and reader/writer exclusion behavior remains covered.
## Final review correction
Snapshot materialization now reads the manifest and every owned artifact exactly once through the
already-open generation descriptor, validates each hash against those exact bytes, and writes the
same byte objects to the private snapshot. A deterministic second-read mutation test proves hostile
pickle bytes can neither pass validation nor enter the snapshot. Reconciliation closes the ACTIVE
generation descriptor in a `finally` block on matches, mismatches, and exceptions. Pipeline-owned
snapshot directories are removed and deregistered after `run_job` on both successful and failed
runs, preventing repeated pipeline use from accumulating temporary directories or registry entries.
The cleanup boundary now begins immediately after snapshot materialization. Resume checkpoint
validation and `JobSpec` construction are guarded by the same release routine as `run_job`, so
corrupt/mismatched resume state or constructor failure clears the pipeline holder, removes the
private directory, and restores the snapshot registry to its prior state before propagating.
## Shipped preprocessing startup contract
Local-vector preprocessing now uses a dedicated Compose override. Both one-shot jobs depend on a
successfully completed `vector-migrate`, whose transitive chain waits for database health and role
reconciliation. The generic preprocessing overlay remains independently renderable and contains no
local-vector services or password secrets. README commands include the local override and build the
job image before running.
The real clean-project smoke no longer injects dependencies or manually starts, reconciles, or
migrates PostgreSQL. Its first shipped `compose run preprocess-evidence` demonstrably creates the
database, waits for health, runs reconciliation and migration, then runs the Evidence job. Unchanged
rerun, changed-source publish, DWH preprocessing, ACTIVE verification, and injected-failure cleanup
all pass through the same shipped dependency path.
@@ -1,93 +0,0 @@
# Evidence preprocessing Task 7 report
Implemented the S3-compatible Evidence adapter, explicit preprocessing Compose overlay, and
operational gates.
- S3 discovery uses bounded paginator pages, page size, and total objects; acquisition enforces a
byte ceiling and always closes streaming bodies.
- Provenance is canonical `s3://bucket/key`. Versioned objects use `s3-version:<version>`;
unversioned objects use a hashed exact ETag, and acquisition refuses validator drift.
- The adapter uses boto3/botocore rather than custom signing. TLS verification is enabled by
default. Custom HTTP and private endpoints require independent explicit opt-ins; endpoint
userinfo is rejected and public custom endpoints are DNS-policy checked.
- Access, secret, and session credentials support file-secret resolution into masked `SecretStr`
config fields. They are never emitted in provenance, reports, errors, or Compose environment.
- `deploy/compose.preprocess.yaml` provides separate one-shot Evidence and DWH jobs and is inert
unless explicitly included with the `preprocess` profile.
- `scripts/preprocess-smoke.sh` verifies both services render without secret material and pins an
unchanged rerun plus a modified generation through deterministic pipeline tests.
Verification: focused S3/HTTP/filesystem/config tests 34 passed; operational smoke 2 passed; core
image with locked boto3 extra built; full harness 702 passed, 5 deselected; scoped Ruff and diff
checks passed.
Operational risk: custom S3-compatible endpoints remain part of the deployment trust boundary.
Private endpoint access must be explicitly enabled and should be restricted by container egress
policy in production. S3 list consistency semantics are provider-defined; version IDs are preferred
over ETags wherever bucket versioning is available.
## Review correction
The Compose overlay now uses committed, purpose-built Evidence and DWH workspace files with
job-specific dependencies. Its services create their lock roots and mount only the vector secrets
they consume. The operational smoke is a real isolated Compose project: real pgvector migrations,
a deterministic in-project embeddings endpoint, actual Evidence CLI JSON across initial/unchanged/
mutated runs, exact ACTIVE verification, an actual DWH introspection job, and owned cleanup.
S3 custom endpoints now fail closed unless declared trusted; HTTP and private loopback endpoints
need additional independent opt-ins. Boto uses forced path-style addressing. Custom endpoints reject
userinfo, query, fragment, and non-root paths. Buckets use strict DNS syntax; listed keys must remain
under prefix and within the S3 byte bound; validators must be nonempty/bounded. Because
ListObjectsV2 does not provide version IDs, discovery honestly fingerprints the exact ETag and
acquisition rejects ETag drift.
Final correction verification: S3/config focused 20 passed; full harness 721 passed, 5 deselected;
real Compose smoke and image build passed; scoped Ruff, shell syntax, and diff checks passed.
## Final security review correction
Literal non-global IPv4/IPv6 endpoints now require the private-endpoint opt-in without claiming DNS
pinning for hostnames. Pagination uses explicit continuation requests and never fetches page
`max_pages + 1`. IP-shaped buckets, leading-slash prefixes, empty/overlong/control-character keys,
and absent validators fail closed. Acquisition accepts only the exact stored `SourceObject` and
compares the response ETag with the stored discovery validator. The real smoke snapshots generation
directory counts after every run and has an injected-failure cleanup mode; cleanup fails if Compose
down fails or any owned container, volume, or network remains.
The canonical smoke correction counts only root-level `corpus/gen-<32 hex>` directories. It exposed
that the durable job path still published an empty unchanged generation; the pipeline now returns
the existing ACTIVE generation without staging a directory when compatibility and all source
fingerprints are unchanged. The smoke therefore proves directory deltas `+1`, `+0`, `+1`.
Failure injection runs a real exit-97 command after resources exist and reaches the EXIT trap.
Cleanup aggregates Compose-down, residual container/volume/network, and temp-directory failures
while preserving the original failure status. S3 prefixes are validated before any client request
for leading slash, UTF-8 byte length, controls, and DEL.
## Canonical unchanged-run correction
The durable job now persists a deterministic source snapshot keyed by source identity. Each entry
binds canonical URI, exact source fingerprint, UTC modification time, canonical immutable metadata,
and explicit media type and size contract fields. The manifest also binds document-to-source
provenance, supplied config/input fingerprints, compatibility, embedding settings, and pipeline and
chunk-policy versions.
An unchanged run reuses ACTIVE only when ownership, bindings, the complete snapshot, document
provenance, materialized document hashes, and every required vector ID/content hash match exactly.
Snapshot changes rebuild only the affected sources; job input/config changes publish a new manifest
while retaining valid stable vector-generation dependencies. Missing or corrupt legacy contract
metadata, documents, or vectors fails closed and rebuilds. The Compose smoke now explicitly expects
the unchanged no-op to report `published=false` while proving generation deltas `+1`, `+0`, `+1`.
## Corrupt ACTIVE reconstruction correction
ACTIVE reuse now reconstructs each source contract from the persisted discovery snapshot and checks
the deterministic document identity, canonical URI, source fingerprint, UTC modification time,
source metadata, applicable media type, content hash, and pipeline identity against the owned
materialized document. The persisted document-source map carries the same exact binding.
Chunks are recomputed under the current chunk policy and must match the manifest exactly in count,
order, IDs, ordinals, content, hashes, linkage, provenance, and policy metadata. Vector health must
report the configured dimension, and every recomputed chunk must have its generation-scoped vector
ID with the exact content hash. Missing, altered, or extra chunks and corrupt document or vector
contracts therefore disable the no-op and rebuild, while a valid unchanged run still performs no
source acquisition.
@@ -1,16 +0,0 @@
# Model provider credential boundary
The backend accepts only an absolute `THT_MODEL_API_KEY_FILE` reference. `PiProcessManager` reads
and validates it afresh before each hosted-provider spawn, rejects symlinks, non-regular/hard-linked,
empty, whitespace-containing, oversized, unreadable, or permissively-mode files, and accepts Docker
0444 secrets only beneath `/run/secrets`. Failures are sanitized and occur before child creation.
Provider names are normalized and mapped to Pi-recognized variables. The child environment removes
the generic path, deprecated `PI_PROVIDER_API_KEY`, and all unselected known provider keys before
injecting only the selected key. Values never enter argv, settings, health, or diagnostics. Local
providers remain keyless and unknown hosted providers fail closed.
The production Compose overlay mounts `model_api_key` read-only and points the backend at its file;
the deployment render smoke proves the value is absent from rendered configuration. Entrypoint,
root README, Pi configuration guide, environment example, and secrets operator guide document the
new contract and reject the legacy generic value variable.
@@ -1,59 +0,0 @@
# Local pgvector whole-plan final fix report
## Outcome
All four binding final-review findings are closed.
1. `PgVectorStore.health()` checks namespace `USAGE` independently for reader and writer
before inspecting vector types. Real PostgreSQL tests revoke only schema `USAGE`, prove both
health sides false and operations unavailable, then grant it back and prove recovery.
2. Direct reader/writer passwords use workspace `password_file` references. Compose mounts the
two files read-only into core and exposes only `_FILE` paths. Rendered Compose and live
`docker inspect` checks prove secret contents are absent.
3. Direct search failures map to `VectorReadUnavailable`; hash/upsert failures map to
`VectorWriteUnavailable`. Messages are fixed and sanitized, original exceptions remain chained,
and upsert rollback is preserved.
4. The shared secret policy uses Linux `stat -c` with macOS `stat -f` fallback. Host files permit
only `0600`/`0400`; Docker's read-only `0444` is accepted only beneath `/run/secrets`. Tests and
operator docs pin this exact policy.
## TDD evidence
The new config, mode, schema-usage, unavailable-connection, and permission regressions failed
before their implementations. The first live secret-policy run also caught GNU `stat -f` accepting
an incompatible format invocation; detection now tries the native Linux form first. The next live
run caught smoke-generated rotation fixtures at `0644`; fixtures now model the documented host
policy.
## Verification
- Real direct pgvector + HTTP parity: `31 passed`.
- Full harness from `harness/`: `493 passed, 5 deselected`.
- Live `local-vector` rotation, restart persistence, inspect boundary, and backup/restore: pass.
- Core image vector migration discovery/status smoke: pass.
- External and local Compose deployment security contracts: pass.
- Config/port focused suite: `26 passed`.
- Secret policy, bootstrap rotation, and backup/restore safety scripts: pass.
- Changed Python Ruff, shell syntax, and `git diff --check`: pass.
One attempted full-harness invocation from the repository root produced a path-dependent failure
in an existing test that opens `workflow.yaml` relative to CWD. It was immediately rerun using the
documented `cd harness && .venv/bin/pytest -q` command and passed completely.
## Operational notes
Workspace files contain file paths, never direct passwords. Secret contents necessarily exist in
the in-process validated `DatabaseConfig` used to establish PostgreSQL connections, but are not
serialized by doctor/Compose/inspect paths. Docker Desktop file-backed secrets may appear as bind
mounts; the safe runtime exception is therefore based on the read-only service mount location
`/run/secrets`, while source files remain owner-only on the host.
## External-profile regression follow-up
Local pgvector is now an explicit `deploy/compose.local-vector.yaml` overlay. The base Compose and
production external override contain no direct vector password declarations, mounts, or `_FILE`
variables, so external deployments do not resolve or require local password files. A real lifecycle
gate unsets all local secret-file variables, renders external config, builds and starts core, waits
for health, and inspects the live container for absence of local direct-vector secret paths. The
local overlay retains its live inspect assertion (paths present, values absent), rotation, restart
persistence, and transactional backup/restore drill.
@@ -1,95 +0,0 @@
# Local pgvector Task 1 report
## Status
Implemented the direct `PgVectorStore` behind the transport-neutral `VectorStore` port.
The adapter uses separate optional reader and writer database configurations, derives
capabilities from configured authority, validates strict positive search limits, filters kinds
in SQL before limiting, and merges multi-collection results by cosine similarity.
All collection identifiers are selected from the fixed `schema_records`, `evidence`, and
`memory` allowlist and composed with `psycopg2.sql.Identifier`. Values, vectors, kinds, hashes,
and limits remain bound parameters. Collection/kind mismatches fail with `VectorStoreError`.
Upserts preserve the canonical metadata shape, use `record_key` conflict semantics, update the
transport hash and embedding, and leave semantic metadata fields intact. Health probes reader
and writer independently and reports observed `vector(N)` dimensions against the configured
embedding dimension.
## Configuration and factory
`pgvector_direct` now accepts explicit optional `reader` and `writer` `DatabaseConfig` entries.
The former `connection` entry remains supported as a deprecated read-only compatibility path.
`build_vector_store(..., require_write=True)` accepts writer-only direct configurations and
fails early when no explicit writer is present.
The transitional `build_vector_loader` bulk-sync path remains in place. It uses an explicit
direct writer when present, or the legacy `connection`; it deliberately does not treat a new
reader-only credential as writable. No production schema migration was added.
## TDD and verification
- RED: the new tests initially failed at collection because `PgVectorStore` did not exist.
- Docker L0 pgvector tests: `11 passed`.
- Direct + HTTP parity/factory/config focus: `51 passed`.
- Full harness: `461 passed, 5 deselected`.
- Changed-file Ruff lint: clean.
- Changed-file Ruff format check: clean.
- `git diff --check`: clean.
The repository-wide `ruff check .` still reports 34 pre-existing test-file findings outside
Task 1; none are in changed files. The full pytest suite emits 17 existing legacy-config
deprecation warnings.
## Scope and concerns
- Test fixtures create only the three existing vector tables needed to exercise the adapter;
migration/versioning remains Task 2.
- The legacy single `connection` form stays read-only through the public port, matching its
previous adapter behavior, while remaining available to the explicitly documented bulk-loader
transition.
## Review fix wave
The Task 1 review findings were addressed in a follow-up TDD cycle:
- Search now validates requested kinds against the global known-kind set, intersects valid kinds
with each collection, and skips unrelated collections. A direct-versus-HTTP parity test covers
the multi-collection case.
- Health requires all three allowlisted tables, an `embedding vector(N)` column on every table,
the expected dimension on every table, and the appropriate read or write table privileges for
each configured side. Empty and partial schemas return deterministic, credential-free details;
unexpected database failures expose only their exception class.
- The Docker L0 fixture now provisions separate least-privilege reader and writer roles. Tests
prove the reader cannot insert, the writer cannot execute the cosine-search SELECT, and the
adapter still routes search to the reader and upsert/hash operations to the writer. Direct
upsert uses an atomic `INSERT ... ON CONFLICT DO NOTHING` followed by `UPDATE` for an existing
key, avoiding broad SELECT authority while retaining conflict-safe hash/upsert semantics.
Fresh verification after the fix wave:
- Docker L0 + HTTP port/search parity: `42 passed` (earlier checkpoint); the final L0 file has
`16 passed` including the stricter raw-role search denial.
- Expanded focused adapter/config suite: `56 passed`.
- Full harness: `466 passed, 5 deselected`.
- Changed-file Ruff lint/format and `git diff --check`: clean.
## Sequence privilege health follow-up
Writer health now resolves the real serial/identity sequence for the `id` column of every
required collection using `pg_get_serial_sequence`. It requires `USAGE` on each resolved
sequence, which is the privilege used by the adapter's implicit `nextval`; sequence `SELECT` is
not required because no adapter operation reads sequence state.
The Docker fixture includes a writer role with complete table/hash-column authority but no
sequence grant. Its health is deterministically unhealthy and a new-key upsert fails. Granting
only sequence `USAGE` makes health green and the same port upsert succeeds. Sequence discovery is
guarded for partial schemas so a missing `id` column produces the existing sanitized schema
diagnostic instead of a PostgreSQL error.
Fresh verification for this follow-up:
- Docker pgvector L0 after formatting: `17 passed`.
- Expanded focused adapter/config/parity suite: `57 passed`.
- Full harness: `467 passed, 5 deselected`.
- Changed-file Ruff lint/format and `git diff --check`: clean.
@@ -1,82 +0,0 @@
# Local pgvector Task 2 report
## Outcome
Implemented ordered, idempotent production migrations and the `tht vector migrate`
interface, including `tht vector migrate --status --json` with pristine JSON output.
## Implementation
- `001_extensions.sql` installs pgvector.
- `002_schema_tables.sql` creates `vectors.schema_records`, `vectors.evidence`, and
`vectors.memory` with the `VectorWriteRecord` columns and `vector(768)` embeddings.
- `003_roles.sql` creates passwordless `NOLOGIN` reader/writer roles. Deployments inject
credentials (or grant these roles to separately-created login roles); no production secret
is stored in the repository.
- Reader authority is schema usage plus table `SELECT`.
- Writer authority is schema usage, table `INSERT`/`UPDATE`, narrow hash-probe column `SELECT`,
and sequence `USAGE`. It has no `DELETE`, broad row `SELECT`, DDL, or ownership authority.
- The migration runner discovers ordered SQL files, records SHA-256 checksums in
`public.tht_vector_migrations`, serializes runners with a transaction-scoped advisory lock,
and applies the full pending batch in one transaction.
- Status distinguishes applied, pending, and checksum-drifted migrations. Apply refuses drift.
A failed migration rolls back both prior migrations in that batch and ledger writes.
## TDD evidence
RED was observed with a real `pgvector/pgvector:pg16` testcontainer: 6 failures for the missing
module, missing command, and missing schema.
GREEN verification:
- Focused migration + direct adapter integration: `23 passed`.
- Full harness from the documented `harness/` cwd: `473 passed, 5 deselected`.
- Targeted Ruff (`tht` plus the new L0 test): clean.
- `git diff --check`: clean.
The new L0 coverage exercises clean install, idempotent rerun, pristine JSON status, checksum
drift, transaction rollback, exact tables/columns/dimensions, role isolation, sequence authority,
and the real `PgVectorStore.health()` plus `VectorWriteRecord` upsert path.
## Existing repository lint baseline
The requested full `ruff check .` was run. It reports 34 pre-existing violations in unrelated
test files (unused imports and one-line semicolon statements). None are in Task 2 files; changing
them would exceed this task's scope. The complete harness test gate is green.
## Self-review
No unresolved Task 2 correctness concern found. One deliberate contract choice is worth noting:
writer `INSERT` and `UPDATE` are table-level because the approved direct adapter health probe uses
`has_table_privilege` for those authorities. Least privilege is retained by withholding broad
`SELECT`, `DELETE`, DDL, ownership, and credentials.
## Review fix wave
The post-implementation review found four production-boundary gaps. They are fixed as follows:
- Migration SQL now ships inside the `tht` wheel (`tht/migrations/vector`) via explicit
setuptools package-data and is discovered through `importlib.resources`, rather than relying on
a source-checkout-relative directory.
- Both status and apply reject ledger versions absent from the installed manifest, including
nonnumeric future version labels. This treats a binary/database downgrade as drift instead of
silently reporting a healthy state.
- Migration files are ordered by parsed integer version; spellings such as `2` and `02` are
rejected as duplicate versions.
- Every migration transaction pins `search_path` locally to `pg_catalog, pg_temp`; catalog calls
and the ledger are schema-qualified. pgvector is installed into the locked `vectors` schema,
tables use `vectors.vector`, and `PgVectorStore` qualifies vector casts and the cosine operator.
A hostile admin default path with a writable shadow schema cannot redirect migration objects.
- The core image build asserts CLI discovery. Image verification now starts an ephemeral pgvector
database, runs the installed image's migration command, and compares pristine apply/status JSON.
Additional verification after the fix wave:
- Focused migration, adapter, hostile-path, and wheel suite: `27 passed`.
- Full harness: `477 passed, 5 deselected`.
- Production core image build: passed, including build-time CLI discovery.
- Core-image apply/status smoke against `pgvector/pgvector:pg16`: passed.
- Changed production and test files: Ruff clean; `git diff --check` clean.
- Full Ruff remains at the same 34 pre-existing unrelated test-file findings documented above.
No dependency changed, so the committed Python requirements lock did not require regeneration.
-133
View File
@@ -1,133 +0,0 @@
# Task 3 report — optional local pgvector profile
## Status
Implemented and verified the `local-vector` Compose profile.
- `vector-db` uses pgvector 0.8.5 on PostgreSQL 16, pinned to the official multi-arch
manifest digest.
- `vector_data` is a project-scoped named volume and is not shared with application data.
- database readiness gates the packaged one-shot `vector-migrate` job; core declares the
migration completion dependency while remaining usable in the pre-existing external profile.
- bootstrap, migrator, reader, and writer identities are distinct. Bootstrap and migration
credentials are supplied as Compose secrets; the application receives only reader/writer
credentials.
- `deploy/workspaces/local-vector.yaml` selects `pgvector_direct` with separate reader and
writer connections.
- the base loopback port binding, `AUTH_MODE=none`, and `THOTH_PUBLIC_EXPOSURE=false` defaults
are unchanged.
## Red/green evidence
The initial Compose contract did not list `vector-db`, as required by the brief. The first real
smoke then failed migration 002 because bootstrap installed the vector extension in `public`.
The bootstrap was corrected to create the `vectors` schema under the migration owner and install
the extension there. A clean-volume rerun passed.
## Verification
- `./scripts/local-vector-smoke.sh`: PASS
- isolated generated Compose project and credentials
- clean migration plus idempotent status rerun
- reader/writer privilege health
- one-record upsert and similarity search
- restart of both `core` and `vector-db`
- persisted search result after restart
- project-only volume cleanup
- `./scripts/test-container-deployment.sh`: PASS
- `./scripts/test-backend-url-policy.sh`: PASS
- `docker compose --profile local-vector config --quiet`: PASS
- harness: 477 passed, 5 deselected
- backend: 84 passed; TypeScript typecheck PASS
- frontend: 226 passed; TypeScript typecheck PASS
- `git diff --check`: PASS
## Self-review / concerns
- Compose cannot make a dependency required only under one profile. The core dependency uses
`required: false` so the established `external` profile does not activate local infrastructure;
under `local-vector`, `compose up --wait` still fails if `vector-migrate` exits nonzero, and the
smoke verifies that successful migration precedes the healthy stack.
- Reader/writer passwords are injected into core environment variables because Compose service
attributes cannot be conditional by profile. Bootstrap and migrator credentials remain
file-backed secrets and are never exposed to core.
- The smoke intentionally refuses the operator project name `thothii` and removes only its unique
project namespace and volumes.
## Follow-up hardening — credential reconciliation and cleanup ownership
Review findings were resolved in a separate follow-up:
- Replaced fresh-volume-only initialization with `vector-reconcile`, an idempotent one-shot that
runs after database health and before `vector-migrate`. It authenticates with only the bootstrap
admin secret, safely creates missing identities, reconciles role attributes and passwords on
existing volumes, restores memberships/ownership, and leaves vector data untouched.
- The migrator is explicitly `NOSUPERUSER NOCREATEDB NOCREATEROLE`. Schema/database ownership is
sufficient for all packaged migrations because reconciliation creates the two group roles first.
- The live smoke rotates migrator, reader, and writer secrets on the same populated volume, rejects
the old reader credential, reruns migrations, recreates core with the new runtime credentials,
and retrieves the record written before rotation and again after database/core restart.
- Smoke project names are no longer caller-controlled. Each run creates a unique namespace and
ownership token. Containers, networks, and volumes carry the ownership label; preflight refuses
any collision and cleanup verifies every discovered resource before `down --volumes`.
- Added a dynamic fake-Docker contract suite for caller override, collision, and mismatched cleanup
labels, plus a real-Docker collision probe using a unique labeled volume.
Follow-up verification:
- `./scripts/local-vector-smoke.sh`: PASS, including live secret rotation and persisted retrieval
- `./scripts/test-local-vector-smoke-safety.sh`: PASS
- `./scripts/test-local-vector-smoke-live-collision.sh`: PASS
- harness: 477 passed, 5 deselected
- backend: 84 passed; TypeScript typecheck PASS
- frontend: 226 passed; TypeScript typecheck PASS
- Compose security, backend URL, config, shell syntax, and diff checks: PASS
Remaining operational constraint: the bootstrap admin secret must continue to match the PostgreSQL
bootstrap account stored in the volume. Runtime migrator/reader/writer rotation is supported without
data deletion; bootstrap-account password rotation is a distinct database-administration operation.
## Final hardening — bootstrap account rotation
The remaining operational constraint is now covered by
`scripts/vector-rotate-bootstrap-password.sh OLD_SECRET_FILE NEW_SECRET_FILE`:
- It does not rely on `POSTGRES_PASSWORD_FILE` after initialization.
- It pre-stages the deployment-file replacement in the same directory, authenticates to the live
database with the explicit old file, and changes only the authenticated bootstrap role.
- Passwords are passed as connection parameters and rendered with psycopg2 SQL composition, so
shell and SQL metacharacters are not interpolated.
- A second connection must authenticate with the new password before the command succeeds. If that
verification fails, the still-open old connection restores the old database password.
- Only after verified database login does an atomic rename replace the current deployment secret.
Wrong-old authentication and verification failures leave deployment configuration unchanged.
Final live smoke evidence on one existing `vector_data` volume:
- wrong-old bootstrap rotation rejected; current deployment secret unchanged
- bootstrap password with quote characters rotated successfully
- old bootstrap login rejected and new login accepted
- `vector-reconcile`, packaged migrations, and core health passed afterward
- the vector record written before rotation remained searchable after rotation and after a further
database/core restart
Final tests:
- `./scripts/test-vector-bootstrap-rotation.sh`: PASS
- `./scripts/local-vector-smoke.sh`: PASS with negative and positive live bootstrap rotation
- existing local-vector collision/safety and Compose deployment contracts: PASS
## Final identity and secret-policy alignment
- `THT_VECTOR_BOOTSTRAP_USER` is now passed through core as well as vector-db and reconciliation,
so the rotation helper uses the authoritative configured role instead of defaulting to `postgres`.
- Rotation and reconciliation source the same raw-file `secret-policy.sh`: non-empty and no
whitespace, including trailing newlines. Rotation validates both files before Docker,
PostgreSQL, or atomic replacement staging; `test-vector-secret-policy.sh` pins empty, newline,
internal-space, and valid metacharacter cases.
- Fake-Docker tests prove a non-default identity reaches the helper path and whitespace rejection
performs no Docker call and creates no staged replacement.
- The real smoke runs the entire stack as `thoth_bootstrap_smoke`. Its whitespace-negative case
leaves the deployment file unchanged and proves the existing database login still succeeds;
non-default-account bootstrap rotation, reconciliation, migration, core health, restart, and
persisted retrieval all pass.
@@ -1,94 +0,0 @@
# Local pgvector Task 4 report
## Outcome
Implemented adapter parity gates and an operator-safe custom-format backup/restore workflow.
- Direct and HTTP stores now share validation, configured-dimension rejection, and deterministic
similarity ordering with record ID as the tie-break.
- The parity fixture exercises identical records through real pgvector and the HTTP RPC contract:
kind filtering, ordering, hashes, replacement upserts, invalid collection/kind errors, and query
plus write dimensions.
- Backup explicitly allowlists the three vector tables and migration ledger, refuses overwrite,
writes through a partial file, and uses a custom compressed archive.
- Restore requires explicit active-source and target coordinates. It compares PostgreSQL system
identifier plus database OID (robust across DNS aliases), refuses the active database, checks for
an empty target unless force is explicit, and restores with exit-on-error.
- Passwords are accepted only through validated secret files, converted to private temporary
`PGPASSFILE`s, and never placed in command arguments or success/error logs.
- Role passwords/login identities are deliberately not dumped. The target must have the approved
passwordless group roles and pgvector extension reconciled before restore; archived ACLs restore
the reader/writer grants.
## TDD and semantic alignment
The first parity run exposed the intended HTTP differences: it accepted unknown collections and
wrong dimensions. Direct pgvector also had no stable order for equal cosine distance. The adapters
were aligned, and the final focused real-pgvector gate passed: **25 passed**.
The first recovery run caught an incorrect probe username before restore. The second caught an
intersection between `pg_dump --schema` and the explicit public ledger table. The third confirmed
the archive contents but caught missing target group roles. Each defect was corrected and the
complete drill was rerun from a fresh generated project.
## Live recovery smoke
`./scripts/local-vector-smoke.sh --backup-restore`: **PASS**.
- generated/owned source Compose project and source `vector_data`
- distinct restore container and distinct named restore volume
- migration and role health, secret rotation, restart persistence
- real custom backup, then deliberate mutation of the active source record
- same-database identity guard evaluated before restore
- restore into the separate target only
- restored hash equals the pre-mutation backup, proving retrieval parity
- migration ledger has all three applied versions
- all three restored embedding columns report `vectors.vector(768)`
- ownership-checked cleanup; the active operator project/volume is never addressed
## Verification
- parity + direct adapter: 25 passed
- full harness: 485 passed, 5 deselected
- changed Python files: Ruff clean
- shell syntax: clean
- `git diff --check`: clean
- full Ruff: unchanged repository baseline of 34 unrelated pre-existing test-file violations
## Self-review and operational constraints
The restore account must be able to read `pg_control_system()` for the robust cluster-identity
comparison and create/restore the selected objects. This is intentionally an administrative
recovery operation, not a runtime reader/writer action. `--force-nonempty` is explicit but still
uses `pg_restore --clean --if-exists`; operators should prefer a new database/volume and validate
migration status, health, and known retrieval before endpoint cutover.
## Post-review hardening
All five final review findings were addressed in a follow-up commit:
- Restore now requires a physically separate PostgreSQL cluster and refuses any equal
`system_identifier`, independent of database OID or hostname.
- `pg_restore` combines `--single-transaction` with `--exit-on-error`. The live drill creates an
existing vector sentinel, deliberately fails late during a forced restore, and proves the
original sentinel row/hash remains unchanged before performing the successful restore.
- Backup uses a mode-0600 `mktemp` in the output directory, atomically renames it, and cleans only
that owned path. A fake-command test pins symlink-clobber resistance and preserves an adversarial
legacy `.partial` symlink and its target.
- HTTP parity now traverses the real `VectorRestClient` transport boundary. It asserts RPC URL/key
and kinds payloads, legacy 404 fallback, response conversion, malformed metadata tolerance, and
canonical `VectorRestError` to `VectorStoreError` mapping.
- The restored target runs role/secret reconciliation and a real `PgVectorStore` with separate
reader/writer logins. Health, known-record search, writer upsert, hash probe, schema/table/column/
sequence authority, and 768-dimensional compatibility are therefore verified through the
production adapter. Reconciliation now restores group-role schema `USAGE`, which table-selected
archives cannot carry.
### Atomic no-replace backup publication
The final publication review is also closed. The private same-directory archive is published with
an atomic hard-link create rather than rename-overwrite semantics. If any process creates the final
file or symlink after preflight but before publication, `ln` fails with `EEXIST`, the backup exits
nonzero, the concurrent destination remains byte-for-byte intact, and the trap removes only the
randomly named temporary archive owned by this invocation. The fake `pg_dump` safety test creates
that destination immediately before returning and pins the failure and cleanup behavior.
-929
View File
@@ -1,929 +0,0 @@
# Pre-deployment Fix Wave Report
Date: 2026-07-14
Worktree: `/home/chirone/ThothII/.worktrees/activity-log-cte-layout`
Base: `e5366d14a6da8fb331d94be60b8929cefb1fe3e0`
## Outcome
All three reviewed findings are implemented in one coherent backend/frontend wave:
1. Resume leaves the prior selection, Zustand state, document panel, and EventSource untouched
until `POST /resume` succeeds. Cold Resume changes state and reconnects only after backend
clear/rebind; already-active same-session Resume preserves the existing binding; failure is a
no-op apart from the fixed toast.
2. SSE uses monotonically increasing per-session ids, cursor-filtered replay, native and manual
reconnect cursors, id continuity across `hub.clear`, and descriptor-id pending-gate
idempotence at both backend and frontend layers.
3. Generic Pi system events and readiness errors are projected through explicit public
allowlists. Sentinel URLs, paths, tokens, stderr, commands, and extra fields do not reach HTTP
or SSE.
No harness, workflow, persistence, model, CTE viewer, CTE card, or shared Card file changed.
## Interfaces
- Frontend `resumeSession(id)` now returns
`Promise<{ id: string; alreadyActive: boolean }>` via `ResumeSessionResult`.
- Backend successful Resume always returns the same shape:
- running/waiting runtime: `{ id, alreadyActive: true }`
- validated cold runtime: `{ id, alreadyActive: false }`
- `SseHub.publish(sessionId, event, data): number` returns the assigned SSE id.
- `SseHub.subscribe(sessionId, send, { afterId, pending })` calls
`send(event, data, id)` for replay/live frames with `id > afterId`.
- `GET /sessions/:id/events` accepts native `Last-Event-ID` and manual
`?lastEventId=<integer>`; when both are valid it uses the greater cursor.
- Every emitted SSE frame is `id: <n>\nevent: <name>\ndata: <json>\n\n`.
- Public readiness failure is exactly:
`Session services are not ready. Check configuration and connectivity, then try again.`
- Generic Pi system events are exactly `{ type: "system_event", event }`, and `event` must be a
non-empty string.
## Files
Backend production:
- `backend/src/bridge/session-bridge.ts`
- `backend/src/routes/sessions.ts`
- `backend/src/sse/sse-hub.ts`
Backend tests:
- `backend/test/routes-sessions.test.ts`
- `backend/test/session-bridge.test.ts`
- `backend/test/sse-hub.test.ts`
- `backend/test/sse-route.test.ts` (new)
Frontend production/support:
- `frontend/src/api/sessions.ts`
- `frontend/src/api/types.ts`
- `frontend/src/shell/AppShell.tsx`
- `frontend/src/store/sessionStore.ts`
- `frontend/src/stream/useSessionStream.ts`
- `frontend/src/test/fakeEventSource.ts`
Frontend tests:
- `frontend/src/api/sessions.test.ts`
- `frontend/src/shell/AppShell.session-mgmt.test.tsx`
- `frontend/src/store/sessionStore.test.ts`
- `frontend/src/stream/useSessionStream.test.tsx`
## TDD RED/GREEN evidence
### 1. Backend Resume result and client-boundary allowlists
RED command:
```text
cd backend && npx vitest run test/routes-sessions.test.ts test/session-bridge.test.ts
```
RED output (exit 1):
```text
Test Files 2 failed (2)
Tests 6 failed | 35 passed (41)
expected { id: 's1' } to deeply equal { id: 's1', alreadyActive: false }
expected raw readiness URL/token/path to equal the fixed public message
expected three raw generic system events to equal [{ type: 'system_event', event: 'session_exit' }]
```
GREEN command:
```text
cd backend && npx vitest run test/routes-sessions.test.ts test/session-bridge.test.ts
```
GREEN output (exit 0):
```text
✓ test/session-bridge.test.ts (14 tests)
✓ test/routes-sessions.test.ts (27 tests)
Test Files 2 passed (2)
Tests 41 passed (41)
```
### 2. Backend exact-once SseHub and route framing
RED command:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/sse-route.test.ts
```
RED output (exit 1):
```text
Test Files 2 failed (2)
Tests 7 failed (7)
expected [undefined, undefined, undefined] to deeply equal [1, 2, 3]
expected unconditional replay not to contain "one" / "two"
expected one buffered pending gate, received replay plus a second pending emission
```
GREEN command:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/sse-route.test.ts
```
GREEN output (exit 0):
```text
✓ test/sse-hub.test.ts (4 tests)
✓ test/sse-route.test.ts (3 tests)
Test Files 2 passed (2)
Tests 7 passed (7)
```
### 3. Frontend cursor tracking and gate idempotence
RED command:
```text
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx src/store/sessionStore.test.ts
```
RED output (exit 1):
```text
Test Files 2 failed (2)
Tests 2 failed | 28 passed (30)
expected /sessions/s1/events to be /sessions/s1/events?lastEventId=7
expected duplicate gate pendingWidget to remain null, received gate-1
```
GREEN command:
```text
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx src/store/sessionStore.test.ts
```
GREEN output (exit 0):
```text
✓ src/store/sessionStore.test.ts (23 tests)
✓ src/stream/useSessionStream.test.tsx (7 tests)
Test Files 2 passed (2)
Tests 30 passed (30)
```
### 4. Frontend typed Resume and AppShell ordering/preservation
Typed API RED command:
```text
cd frontend && npx tsc -b
```
Typed API RED output (exit 1):
```text
src/api/sessions.test.ts(43,9): error TS2322: Type 'void' is not assignable to type
'{ id: string; alreadyActive: boolean; }'.
```
Lifecycle RED command:
```text
cd frontend && npx vitest run src/api/sessions.test.ts src/shell/AppShell.session-mgmt.test.tsx
```
Lifecycle RED output (exit 1):
```text
✓ src/api/sessions.test.ts (9 tests)
❯ src/shell/AppShell.session-mgmt.test.tsx (15 tests | 4 failed)
Test Files 1 failed | 1 passed (2)
Tests 4 failed | 20 passed (24)
already-active same-session Resume created two EventSources instead of one
deferred cold Resume closed the document panel before POST completion
failed same-session Resume closed the prior EventSource
failed Resume with no active session opened an EventSource
```
GREEN commands:
```text
cd frontend && npx vitest run src/api/sessions.test.ts src/shell/AppShell.session-mgmt.test.tsx
cd frontend && npx tsc -b
```
GREEN output (exit 0):
```text
✓ src/api/sessions.test.ts (9 tests)
✓ src/shell/AppShell.session-mgmt.test.tsx (15 tests)
Test Files 2 passed (2)
Tests 24 passed (24)
TypeScript: no output, exit 0
```
The AppShell cold-reconnect test additionally proves that the old source accepts an event while
Resume is pending, the replacement URL carries `lastEventId=8`, the replacement receives one
post-resume transcript/activity row, and two deliveries of the same descriptor id yield one gate.
## Affected verification
Backend command:
```text
cd backend && npx vitest run test/routes-sessions.test.ts test/session-bridge.test.ts \
test/sse-hub.test.ts test/sse-route.test.ts test/health.test.ts test/e2e-f1.test.ts
```
Output (exit 0):
```text
Test Files 6 passed (6)
Tests 52 passed (52)
```
Backend typecheck:
```text
cd backend && npx tsc --noEmit -p .
```
Output: no output, exit 0.
Frontend command:
```text
cd frontend && npx vitest run src/api/sessions.test.ts src/store/sessionStore.test.ts \
src/stream/useSessionStream.test.tsx src/shell/AppShell.session-mgmt.test.tsx \
src/shell/CentralStatus.test.tsx src/shell/ModelActivityPanel.test.tsx \
src/shell/f1-loop.test.tsx src/shell/AppShell.new-session.test.tsx
```
Output (exit 0):
```text
Test Files 8 passed (8)
Tests 74 passed (74)
```
Frontend typecheck:
```text
cd frontend && npx tsc -b
```
Output: no output, exit 0.
## Full verification
Backend full suite:
```text
cd backend && npx vitest run
```
```text
Test Files 22 passed (22)
Tests 177 passed (177)
```
Frontend full suite:
```text
cd frontend && npx vitest run
```
```text
Test Files 43 passed (43)
Tests 271 passed (271)
```
Backend production build:
```text
cd backend && npm run build
> tsc -p tsconfig.json
exit 0
```
Frontend production build:
```text
cd frontend && npm run build
> tsc -b && vite build
✓ 4835 modules transformed.
✓ built in 8.25s
exit 0
```
## Integrated re-review closure (2026-07-15)
This section supersedes the earlier cold same-session assertion that the replacement URL carries
`lastEventId=8`. That behavior was correct only while the backend process and its in-memory id
sequence survived. A restarted backend begins a fresh sequence, so a successful cold Resume now
explicitly discards the browser's cursor before replacing the EventSource.
All four integrated re-review findings are closed:
1. `AppShell` passes a dedicated cursor-reset epoch to `useSessionStream`. A cold same-session
Resume increments it only after `alreadyActive: false`; a high cursor such as `901` is omitted
from the replacement URL and fresh low-id events/gates are consumed. An already-active
same-session Resume still preserves its source, cursor, and store.
2. `useSessionStream` no longer mutates the cursor ref during render. Effect setup resets cursor
state on session/reset-epoch changes, callbacks are guarded by a captured active-source
identity, and cleanup clears only its own active identity. A queued event from the replaced
source cannot write the new store or poison its next reconnect URL.
3. Backend Resume is serialized per session and rechecks runtime state inside the lock. Manifest,
readiness, and reopen validation precede the transport commit. Idle/failed replacement creates
and binds the new runtime before `hub.clear`, which occurs synchronously immediately before the
first `Resuming session` publish. Reopen/create failure returns exactly
`Session could not be resumed. Check configuration and connectivity, then try again.`, keeps the
prior hub buffer/subscribers attached, and does not expose exception sentinels. Concurrent calls
perform one cold start and the waiter returns `alreadyActive: true`.
4. `SseHub.forget(id)` removes subscribers, buffered events, and the last id. Permanent session
DELETE invokes it after disk deletion; ordinary close and Resume continue to use `clear`, which
preserves the id sequence.
### Re-review files
Production:
- `backend/src/pi/pi-process-manager.ts`
- `backend/src/routes/sessions.ts`
- `backend/src/sse/sse-hub.ts`
- `frontend/src/shell/AppShell.tsx`
- `frontend/src/stream/useSessionStream.ts`
Tests/support:
- `backend/test/pi-process-manager.test.ts`
- `backend/test/routes-sessions.test.ts`
- `backend/test/sse-hub.test.ts`
- `frontend/src/shell/AppShell.session-mgmt.test.tsx`
- `frontend/src/stream/useSessionStream.test.tsx`
- `frontend/src/test/fakeEventSource.ts`
### Re-review TDD RED/GREEN evidence
Frontend RED command:
```text
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx \
src/shell/AppShell.session-mgmt.test.tsx
```
RED output (exit 1):
```text
Test Files 2 failed (2)
Tests 3 failed | 21 passed (24)
reset epoch: expected the old source to close, received false
cold same-session: expected /sessions/s1/events, received ?lastEventId=901
stale source: expected an empty transcript, received "stale session one"
```
Frontend GREEN command:
```text
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx \
src/shell/AppShell.session-mgmt.test.tsx
cd frontend && npx tsc -b
```
GREEN output (exit 0):
```text
Test Files 2 passed (2)
Tests 24 passed (24)
TypeScript: no output, exit 0
```
Backend RED command:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/routes-sessions.test.ts
```
RED output (exit 1):
```text
Test Files 2 failed (2)
Tests 8 failed | 28 passed (36)
three Resume ordering assertions observed clear before reopen/create
reopen and create sentinels escaped as raw HTTP 500 responses
the concurrent waiter cold-started again instead of returning alreadyActive: true
SseHub.forget was absent and DELETE did not invoke permanent cleanup
```
Backend GREEN command:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/routes-sessions.test.ts
cd backend && npx tsc --noEmit -p .
```
GREEN output (exit 0):
```text
Test Files 2 passed (2)
Tests 36 passed (36)
TypeScript: no output, exit 0
```
The failure tests publish a post-failure probe through the same hub and prove that a subscriber
attached before either reopen or create rejection still receives it. The concurrency test overlaps
two same-id requests behind a deferred reopen and proves one manifest/readiness/reopen/create/clear
sequence.
### Initial re-review verification (before independent-review hardening)
```text
cd backend && npx vitest run
Test Files 22 passed (22)
Tests 182 passed (182)
cd frontend && npx vitest run
Test Files 43 passed (43)
Tests 273 passed (273)
cd backend && npm run build
> tsc -p tsconfig.json
exit 0
cd frontend && npm run build
> tsc -b && vite build
✓ 4835 modules transformed.
✓ built in 8.46s
exit 0
```
`git diff --check` produced no output (exit 0). The frontend build retains its pre-existing
large-chunk warning; no new build or type errors were introduced.
Final whitespace verification:
```text
git diff --check
no output, exit 0
```
## Self-review
- Resume sequencing: reopen and runtime binding precede backend `clear` and HTTP success; frontend
state mutation and cursor-reset epoch follow it. Failure catch only emits fixed UI copy.
- Already active: same-session returns before reset/generation/manifest repaint; different session
resets the single-session store and binds the new id only after success.
- SSE exact-once: ids are transport identity, not content hashes; replay is strictly `id > cursor`;
`clear` retains the counter; pending gate matching uses only descriptor id.
- Cursor behavior: hook tracks `MessageEvent.lastEventId`, carries it only to an ordinary same-id
generation, and resets it on session-id/cold-runtime epoch change. Native EventSource reconnect
remains supported by the route header.
- Gate defense: the Zustand set survives pending clear but resets with the session store.
- Client boundary: raw `ensure.error` is unused in public responses; generic Pi system events are
reconstructed rather than spread; frontend type mirrors the two-field event.
- Scope: `git diff` contains no CTE/Card/harness/workflow/persistence/model changes. Four pre-existing
modified `.superpowers/sdd/{progress,task-2-report,task-3-report,task-4-report}.md` files are user
work and are excluded from staging.
## Remaining concerns
- The 200-event SSE ring limit remains intentional. A brand-new page can reconstruct only retained
backlog; an in-memory same-session reconnect is exact-once from its cursor.
- Per-session sequence counters remain in backend memory after `clear` by design so later in-process
cold same-id Resume cannot reuse ids. Permanent DELETE removes the counter via `forget`.
- The Delete-then-Resume adversarial route test proves the deleted session is not resurrected but
currently receives the runner's generic HTTP 500 when `sessionShow` can no longer find it. A
future API cleanup can normalize that missing-session response to 404 or 409.
- Frontend tests still print pre-existing MSW unhandled-request and React ref/`act` warnings even
though all 276 tests pass. The frontend production build still reports pre-existing large chunk
warnings. Neither warning class was introduced or expanded by this change.
- No live Pi/DWH smoke was run; this wave changes only REST/SSE/frontend lifecycle boundaries and
is covered by fake-Pi, live Fastify SSE, component, full-suite, typecheck, and production-build
gates.
## Independent-review hardening
The required independent review was run repeatedly against the uncommitted diff. Its first pass
found four Important lifecycle edges beyond the integrated findings: queued old-runtime callbacks,
post-spawn construction cleanup, concurrent frontend Resume completions, and the passive-effect
commit window. Its second pass confirmed those fixes and identified one remaining Important
retention issue in the new runtime-identity map. The final pass reported no Critical, Important, or
Minor findings and assessed the diff ready to merge.
The resulting hardening is:
- Runtime bridge callbacks are gated by the bound runtime identity. Replacement, close, and DELETE
invalidate the old identity, so queued old events cannot publish or call `failSession`. An active
runtime removed by the manager can still publish its complete public failure sequence; after the
terminal unmanaged `agent_end`, its binding is released and later events are rejected.
- `PiProcessManager` kills the spawned child and removes any registered map entry if either
spawn-boundary stderr setup or later RPC/bridge/map initialization throws.
- Resume completion compares against synchronously maintained current active-session identity.
Concurrent `alreadyActive: false` then `alreadyActive: true` results preserve the cold source,
cursor, store, and replayed gate.
- Stream source replacement uses a layout effect. A deterministic later-layout-effect test delivers
a queued old event inside the former commit-to-passive-cleanup window and proves it is ignored.
- Cursor tests cover both a restarted backend's fresh low ids and an in-process hub's preserved high
ids followed by a cursor-bearing ordinary reconnect.
### Hardening TDD RED/GREEN evidence
Backend identity/construction RED command:
```text
cd backend && npx vitest run test/pi-process-manager.test.ts test/routes-sessions.test.ts
```
```text
Test Files 2 failed (2)
Tests 3 failed | 69 passed (72)
post-spawn reader initialization did not kill the child
replaced and deleted runtime callbacks still called failSession/published
```
Additional spawn-boundary and terminal-release RED checks:
```text
cd backend && npx vitest run test/pi-process-manager.test.ts \
-t "spawn boundary initialization"
Tests 1 failed | 38 skipped (39)
cd backend && npx vitest run test/routes-sessions.test.ts -t "terminal sequence"
Tests 1 failed | 34 skipped (35)
```
Frontend concurrency/layout RED command:
```text
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx \
src/shell/AppShell.session-mgmt.test.tsx
```
```text
Test Files 2 failed (2)
Tests 2 failed | 25 passed (27)
the later-layout-effect event wrote "commit-window stale text"
the false→true completion pair erased pending gate "cold-gate"
```
Final focused GREEN commands:
```text
cd backend && npx vitest run test/pi-process-manager.test.ts \
test/routes-sessions.test.ts test/sse-hub.test.ts
cd backend && npx tsc --noEmit -p .
Test Files 3 passed (3)
Tests 79 passed (79)
TypeScript: no output, exit 0
cd frontend && npx vitest run src/stream/useSessionStream.test.tsx \
src/shell/AppShell.session-mgmt.test.tsx
cd frontend && npx tsc -b
Test Files 2 passed (2)
Tests 27 passed (27)
TypeScript: no output, exit 0
```
### Final full verification after review hardening
```text
cd backend && npx vitest run
Test Files 22 passed (22)
Tests 188 passed (188)
cd frontend && npx vitest run
Test Files 43 passed (43)
Tests 276 passed (276)
cd backend && npm run build
> tsc -p tsconfig.json
exit 0
cd frontend && npm run build
> tsc -b && vite build
✓ 4835 modules transformed.
✓ built in 8.47s
exit 0
```
The final frontend run retains the repository's pre-existing MSW/ref/`act` warnings, and the build
retains the pre-existing large-chunk warning. No test, typecheck, or build failures remain.
## Stale-bootstrap, lifecycle-lock, and competing-Resume hardening
Date: 2026-07-15
Base: `08b1f4909e8eb7538156cecc2e7a6cafb46ddfc7`
This follow-up closes asynchronous identity/order and multi-client transport gaps found in the
pre-deployment review:
- `PiProcessManager.teardownIfCurrent(id, runtime)` makes teardown an identity-checked operation.
Bootstrap re-checks identity after configuration/retrieval and before both the public
`Starting model` event and model start. Its failure continuation acquires the same session
lifecycle lock, claims only its own runtime identity, and holds serialization through persisted
failure and the public terminal sequence. A continuation left behind by Close or DELETE cannot
target a replacement or recreate forgotten SSE state.
- The former Resume-only promise tail is now a per-session lifecycle lock shared by Resume, Close,
and DELETE. Each route reads the current runtime inside the lock immediately before replacement
or removal and uses identity-checked teardown. Deferred route tests prove both orderings:
Resume then Close/Delete finishes removed with no post-removal bootstrap event; Close then Resume
creates only after Close completes; DELETE then Resume cannot recreate a deleted session.
- AppShell assigns each Resume invocation a monotonic token and records the latest target. A
completion for a different, superseding session id cannot reset the store, select a source, close
the panel, or repaint phase from a late manifest. Same-id invocations are per-target single-flight
operations through the POST and local binding commit: repeated pre-commit clicks update the
shared operation's latest token but issue no second POST or commit path. The operation becomes
joinable again before its manifest fetch, whose repaint remains token/id/selection guarded. Start
new, Stop, streamed session exit, and active-session deletion invalidate pending Resume work.
This prevents stale-source preservation and reverse/non-Resume intent overwrite without allowing
a slow manifest to suppress a later explicit rebind.
- `SseHub` subscriber registrations now carry idempotent transport-close callbacks. `clear` and
`forget` snapshot and actively close every response before discarding runtime transport state;
callback-driven unsubscription during that iteration is safe. The SSE route ends its response so
native EventSource reconnects with `Last-Event-ID`. Post-clear events retain monotonic ids and are
buffered for replay; `forget` additionally resets the id state.
Production files:
- `backend/src/pi/pi-process-manager.ts`
- `backend/src/routes/sessions.ts`
- `backend/src/sse/sse-hub.ts`
- `frontend/src/shell/AppShell.tsx`
Regression tests:
- `backend/test/pi-process-manager.test.ts`
- `backend/test/routes-sessions.test.ts`
- `backend/test/sse-hub.test.ts`
- `backend/test/sse-route.test.ts`
- `frontend/src/shell/AppShell.session-mgmt.test.tsx`
### TDD RED/GREEN evidence
Runtime identity API RED:
```text
cd backend && npx vitest run test/pi-process-manager.test.ts -t "identity-checked teardown"
Test Files 1 failed (1)
Tests 1 failed | 39 skipped (40)
TypeError: mgr.teardownIfCurrent is not a function
```
Runtime identity API GREEN:
```text
Test Files 1 passed (1)
Tests 1 passed | 39 skipped (40)
```
Deferred bootstrap RED:
```text
cd backend && npx vitest run test/routes-sessions.test.ts \
-t "stale bootstrap|bootstrap that"
Test Files 1 failed (1)
Tests 6 failed | 35 skipped (41)
close/delete + replacement: stale continuation removed the replacement runtime
delete without replacement: stale continuation called failSession after forget
```
Deferred bootstrap GREEN:
```text
Test Files 1 passed (1)
Tests 6 passed | 35 skipped (41)
```
Shared lifecycle ordering RED:
```text
cd backend && npx vitest run test/routes-sessions.test.ts \
-t "Resume followed|Close followed|Delete followed"
Test Files 1 failed (1)
Tests 4 failed | 41 skipped (45)
All four deferred assertions observed the competing route settle before the first lifecycle
operation released.
```
Bootstrap plus lifecycle GREEN:
```text
cd backend && npx vitest run test/routes-sessions.test.ts \
-t "Resume followed|Close followed|Delete followed|stale bootstrap|bootstrap that"
Test Files 1 passed (1)
Tests 10 passed | 35 skipped (45)
```
Competing frontend Resume RED:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "competing Resume|stale Resume manifest"
Test Files 1 failed (1)
Tests 2 failed | 16 skipped (18)
reverse POST completion opened a second, stale EventSource
late s1 manifest repainted the selected s3 phase from F3 to F7
```
Competing and same-id Resume GREEN:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "competing Resume|stale Resume manifest|false then true"
Test Files 1 passed (1)
Tests 3 passed | 15 skipped (18)
```
### Independent-review hardening RED/GREEN
The first final review reported no Critical findings and three Important edge cases: bootstrap
could start during an in-progress Close; bootstrap-owned failure was persisted twice; and an older
same-id result could overwrite newer state. The integrated reviewer also required non-Resume
navigation to invalidate pending Resume work. The final main review tightened the same-ID contract
to true single-flight so a second same-target click cannot preserve a dead pre-restart source.
Backend review RED:
```text
cd backend && npx vitest run test/routes-sessions.test.ts \
-t "Close suppresses|bootstrap failure persists once"
Test Files 1 failed (1)
Tests 2 failed | 45 skipped (47)
deferred configure started Pi while closeSession was still pending
bootstrap/public failure called failSession twice
```
Backend review GREEN:
```text
Test Files 1 passed (1)
Tests 2 passed | 45 skipped (47)
```
Same-id single-flight RED:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "share one cold request"
Test Files 1 failed (1)
Tests 1 failed | 18 skipped (19)
two concurrent same-ID invocations issued two cold POSTs (three total including initial activation)
```
Non-Resume invalidation RED:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "starting a new question invalidates"
Test Files 1 failed (1)
Tests 1 failed | 19 skipped (20)
the late Resume opened an EventSource after Start new returned to the landing state
```
Frontend review GREEN:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "share one cold request|competing Resume|stale Resume manifest|starting a new question invalidates"
Test Files 1 passed (1)
Tests 4 passed | 15 skipped (19)
```
Post-commit single-flight lifetime RED:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
-t "releases same-id single-flight"
Test Files 1 failed (1)
Tests 1 failed | 19 skipped (20)
s1 committed and waited on its manifest; after s3 superseded it, a new s1 Resume reused the old
operation and issued no second s1 POST (expected 2, received 1).
```
Same-id and manifest lifetime GREEN:
```text
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx -t "same-id|manifest"
Test Files 1 passed (1)
Tests 4 passed | 16 skipped (20)
cd frontend && npx tsc -b
no output, exit 0
```
Multi-client SSE disconnect RED:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/sse-route.test.ts
Test Files 2 failed (2)
Tests 3 failed | 6 passed (9)
clear/forget invoked zero of two registered close callbacks, and two live HTTP SSE responses timed
out instead of reaching EOF after clear.
```
Multi-client SSE disconnect GREEN:
```text
cd backend && npx vitest run test/sse-hub.test.ts test/sse-route.test.ts
Test Files 2 passed (2)
Tests 9 passed (9)
cd backend && npx tsc --noEmit -p .
no output, exit 0
```
The Hub tests use two subscribers whose close callbacks immediately unsubscribe themselves, proving
safe snapshot iteration and exactly-once closure. The live-route test opens two HTTP streams, proves
both receive EOF on clear, publishes a new event and gate, then reconnects after id 1 and replays
exactly ids 2 and 3. The forget test closes both subscribers and proves the next id resets to 1.
Close now removes the observed runtime identity before awaiting persistence. Failure persistence is
claimed once per runtime and lifecycle-serialized; bootstrap's public `session_failed` cannot start
a duplicate. A per-target in-flight map owns the only same-ID POST and commit while its mutable
latest token keeps s1→s2→s1 ordering correct; it is removed immediately after the binding commit,
before awaiting the independently guarded manifest. One shared invalidation helper is called when
active deletion, streamed exit, Start new, or Stop begins.
### Focused verification
```text
cd backend && npx vitest run test/routes-sessions.test.ts test/pi-process-manager.test.ts \
test/sse-hub.test.ts test/sse-route.test.ts
Test Files 4 passed (4)
Tests 96 passed (96)
cd backend && npx tsc --noEmit -p .
no output, exit 0
cd frontend && npx vitest run src/shell/AppShell.session-mgmt.test.tsx \
src/shell/AppShell.new-session.test.tsx src/stream/useSessionStream.test.tsx
Test Files 3 passed (3)
Tests 36 passed (36)
cd frontend && npx tsc -b
no output, exit 0
```
### Full verification
```text
cd backend && npx vitest run
Test Files 22 passed (22)
Tests 202 passed (202)
cd frontend && npx vitest run
Test Files 43 passed (43)
Tests 280 passed (280)
cd backend && npm run build
> tsc -p tsconfig.json
exit 0
cd frontend && npm run build
> tsc -b && vite build
✓ 4835 modules transformed.
✓ built in 8.47s
exit 0
```
The frontend suite/build retain the previously documented MSW, React ref/`act`, experimental type
stripping, and large-chunk warnings. No warning class was introduced by this wave. No harness,
workflow, persistence, SQL/CTE viewer, model-selection, or deployment file changed. The four
pre-existing modified `.superpowers/sdd/{progress,task-2-report,task-3-report,task-4-report}.md`
files remain excluded from staging.
### Final independent-review verdict
After the multi-client transport fix, the independent reviewer reported no Critical, Important, or
Minor findings. Its own focused verification passed 96 backend transport/lifecycle tests, 31
frontend Resume/stream tests, both TypeScript checks, and `git diff --check`. Final assessment:
**Ready to deploy: Yes.**
-53
View File
@@ -1,53 +0,0 @@
# DWH REST per-installation authentication SDD progress
Plan: `docs/superpowers/plans/2026-08-20-dwh-rest-per-installation-auth.md`
Branch: `feat/dwh-rest-installation-auth`
Worktree: `/home/chirone/ThothII-next/.worktrees/dwh-rest-installation-auth`
Baseline: workspace docs PASS; Go unavailable on host (use containerized Go 1.26.5); default Compose pre-existing path-sensitive false positive under `/home/chirone`.
Task 1: complete (commits 3bc84b0..1e82fd3, review clean after bounded-digest fix wave).
Task 1 plan note: the dotted-import grep was resolved by the exact module assertion in `09290a0`; `go list -m all` confirms the DWH module only.
Task 2: complete (commits 541ef45, 971a0e6, e90a1a1; independent vet fix d2415b5; review clean after bounded-read, expiry/JSON, same-Store, and cross-Store synchronization fix waves).
Task 3: complete (commits ebb360f, 055dcab, b1079bd; independent review clean after CLI grammar, metadata validation, and ambiguous-publication cleanup fix waves).
Task 4: complete (commits 943f809, 419c344, 134dc19; independent review PASS after legacy multiplicity, fail-closed handler/socket, SGID 2750, O_RDONLY shared lock, OpenReadOnly, and safe socket-parent waves). Task 5 must precreate `.writer.lock` as `0640 root:dwh-auth`; add a non-owner group/cross-process integration proof when packaging permits.
Task 5: complete (commits 87606c7, 62ec29f; independent review PASS after auth-subrequest header isolation and target-Nginx duplicate-header verification). No runtime installation or service/Nginx mutation performed.
Task 6: complete (commits 1b18a0f, d0f7e04; independent review PASS after regex/duplicate bypass, exact-PID TCP, and bounded-cleanup hardening). Minor for final review: remove or rename the redundant legacy `negative_postgrest_bypass` fixture and align the historical fixture-count prose if useful. No active Nginx/runtime mutation performed.
Task 7: complete (commits 7b9b8b3, 707c13d, f616aab, 7b86ea9, 7fe5316; Terra review PASS). Documentation and rollout-contract alignment completed; no runtime mutation performed.
Task 8: complete at frozen SHA `6499d24892b4383ac492579303e766cfb51fe44e` (fix commits `09290a0`, `6499d24`; Terra review PASS). Focused, portability, scanner, DWH Go, and two full tools/tht matrix runs PASS; evidence recorded at `.artifacts/dwh-auth/source-verification.md`. Broad coupling remains `BASELINE_RED` debt; immutable paths remain unchanged.
Task 9: PASS at frozen SHA `0c4ff3750d3ecd3fc514e50e511cf7475fbe0446`. Built and installed the exact local candidate, enabled and started `dwh-auth`, imported the protected legacy credential under public ID `legacy-shared`, and created `psd-mac-primary` with public key ID `oNPdOfoH7ypLtVb1`. Registry check, AF_UNIX-only listener, v1/legacy `204`, random/missing `401`, bounded journal scan, installed-file hashes, and `nginx -t` all PASS. Protected report: `/root/dwh-auth-provision/gate9-20260821T054514Z.report.md`, SHA-256 `65ce1e0d8be5f74355eca2b1dca901da16f2864f68eafb7dede9d23ef36b82d5`. Nginx was not changed or reloaded; the legacy key remains active; the old stack was not changed or stopped. Terra final review: PASS with no Critical or Important findings.
Task 10: NOT STARTED and requires a second explicit authorization. Public Nginx cutover, Mac-key delivery/configuration, and legacy revocation have not occurred. Activity 1 remains `IN_DISCUSSION`; external deployment remains `SURVEY_NO_GO`.
Final clarification (bookkeeping): initial authorization at `6499d24` stopped before installation because the protected legacy file was missing and a journal-scan finding remained. A secret-safe legacy file was prepared without emit/hash; Nginx metadata remained unchanged and `nginx -t` PASS. Fix commits `6fb4886`, `dee0f9c`, `0c4ff37` received Terra PASS, followed by a detached complete re-freeze PASS at full `0c4ff3750d3ecd3fc514e50e511cf7475fbe0446`. The owner then explicitly authorized Gate 9 at that exact SHA; Gate 9 completed as recorded above. Task 10 remains a separate gate.
## Project A authentication runtime projection
Plan: `docs/superpowers/plans/2026-08-21-project-a-server-auth-runtime-projection.md`
Plan commit: `64f46c7019e11a74dae35a7cdb447cd881e17061`
Implementation baseline: `64f46c7019e11a74dae35a7cdb447cd881e17061`
Runtime constraint: source, synthetic tests, and documentation only; Project A, `/srv`, Nginx,
the legacy stack, and shared services remain untouched.
Task 1: complete (commits `8a8f2c2`, `f9e2950`, `de86760`; independent Terra review PASS after descriptor-relative rewrite, full-history validation, crash recovery, destructive replacement guards, deterministic failure seams, and interrupted-retention recovery).
Task 2: complete (commits `05f8615`, `da2f4a6`, `1e2c4e6`; independent review PASS after exact GID enforcement, bounded descriptor-bound namespace enumeration, strict trailing-slash parity, OIDC coverage, and complete one-retry `CURRENT` publication linearization).
Task 3: complete (commit `903c0b4`; independent Terra review PASS after retained-FD outer locking, cancellable runtime/canonical waits, public transaction-context propagation, deterministic swap/metadata/creator-race tests, and fail-closed pre/post-commit error handling).
Task 4: complete (commit `3d9a9f0`; independent Terra review PASS after moving the Linux/root restore gate before secret-bearing checkpoint creation). Exact focused Go, race, vet, Node 24 focused/full, TypeScript, projected Compose, secret-policy, Windows backup compile, and full serialized tools/tht gates PASS. Historical canonical/unified Compose failures were reproduced as baseline-only documentation/path coupling failures and were not weakened.
Task 5: complete (commit `ef7ae70`; independent Terra review PASS after read-only Vitest gate repair,
remote-Docker/context hardening, and adversarial canonical-mount/`sudo printenv` verifier fixes).
All 14 cross-layer acceptance cases, documentation verifiers, shell syntax, full Go/race/vet,
backend Node 24 Vitest/typecheck, projected Compose, and secret-policy gates PASS. The three known
default/canonical/unified Compose policy failures were reproduced at `a21e2c1` and remain
unmodified baseline debt. Project A has not been started; applying the descriptor or any runtime
root under `/srv/thothii` still requires a new explicit authorization.
Final Project A source review: PASS for `a21e2c1..ef7ae70`; independent Terra review found no
remaining Critical or Important issue after validating restore admission/order, backend path and
identity controls, transaction cancellation, lifecycle gating, portability, dependency scope, and
redaction. Pre-live stop boundary remains in force.
-325
View File
@@ -1,325 +0,0 @@
# Task 2 report — protected atomic registry
## Scope and commit
- Commit: `541ef45 feat: add protected DWH credential registry`
- Committed files only:
- `tools/dwh-auth/internal/securefile/securefile_linux.go`
- `tools/dwh-auth/internal/securefile/securefile_linux_test.go`
- `tools/dwh-auth/internal/registry/store.go`
- `tools/dwh-auth/internal/registry/store_test.go`
- No server, Nginx, systemd, Docker stack, real registry, secrets, or legacy ThothII files were
read or changed. Tests use `t.TempDir` and synthetic record digests only.
## TDD evidence
All Go commands ran in the required official `golang:1.26.5` container with only this linked
worktree bind-mounted at `/work`. The container image reports `go version go1.26.5 linux/amd64`.
### RED
Before either Task 2 production file existed, the focused command was run inside the container:
```text
go test ./internal/securefile ./internal/registry -count=1
```
It failed non-zero for the expected absent implementation symbols, including `undefined: OpenDir`,
`undefined: ReadSecret`, `undefined: Open`, `undefined: State`, `undefined: PublicRecord`, and
`undefined: Store`.
### GREEN
After the minimal implementation and formatting:
```text
go test ./internal/securefile ./internal/registry -count=1
```
Result:
```text
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry
```
### Race verification
The required race command completed successfully:
```text
go test -race ./internal/securefile ./internal/registry -count=1
```
Result:
```text
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 1.027s
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 1.179s
```
Additional scoped verification:
```text
go vet ./internal/securefile ./internal/registry
go test ./... -count=1
git diff --cached --check
```
The Task 2 vet command completed with no findings; all four DWH-auth packages passed the full
module test run; the staged-diff check completed with no output.
## Delivered behavior
- `securefile` is Linux-only and traverses absolute paths through descriptor-anchored
`syscall.Open`/`Openat` calls with `O_NOFOLLOW|O_CLOEXEC`; protected roots, child directories,
records, and secret files are regular/directories only and are checked against `Lstat` after
`Fstat`.
- Protected reads reject special, group-writable, or world-writable modes, cap record reads at
4096 bytes, read at most one extra byte, and reject file-size changes or short/partial reads.
Secret ingress additionally requires exact `0600`.
- Secret output uses `O_CREAT|O_EXCL|O_NOFOLLOW`, exact `0600`, and an absolute protected parent.
- `registry.Open` creates protected `active` and `revoked` subdirectories under an existing safe
root. Record enumeration rejects unexpected entries, unsafe files, symlinks, oversized files,
bad filenames, malformed JSON, unknown JSON fields, duplicate JSON fields, and trailing JSON.
- `Add` validates Task 1 records, writes canonical JSON plus one newline through an exclusive
temporary file, sets final mode `0640`, syncs the file, renames under a protected per-root writer
lock, then syncs the directory.
- `Revoke` writes and syncs a valid revoked record before unlinking and syncing the active record.
`Find` checks revoked first; `List` resolves an active/revoked overlap to the revoked public
record. `FindLegacy` scans fail-closed and permits only the reserved legacy record state.
- `PublicRecord` deliberately omits `secret_sha256`; the redaction is regression-tested.
## Security-test coverage
- protected normal files and canonical record publication;
- symlinked roots, registry directories, records, and secret input;
- unsafe root/directory/record/secret modes;
- bounded/oversized record input;
- unknown, duplicate, trailing, and partial JSON;
- filename mismatch and multiple legacy-record integrity failures;
- revoked-state precedence when both active and revoked files exist;
- concurrent adds and concurrent reads during revocation, including the race detector.
## Self-review
Reviewed all syscall, path, mode, and error paths after the final race run:
- Directory traversal never follows a supplied component; later operations use retained directory
descriptors, not re-opened untrusted prefixes.
- `Fstat` validates the opened object and `Lstat` must identify the same inode/device; the direct
child name grammar refuses separators, dot components, and NUL.
- File validation occurs before and after reads; mode/type/size checks fail closed. Directory
listing obtains a fresh `openat(dirfd, ".")` descriptor so scans do not share a mutable directory
offset.
- Writer serialization protects the check-then-rename no-replace sequence. Failed temporary
cleanup leaves an unexpected entry that later scans reject rather than silently accepting it.
- State-specific validation rejects revocation metadata in active records and requires it in
revoked records. Revoked files are consulted before active files so interruption after revoked
publication cannot reactivate a credential.
- All functionality uses only Go standard-library packages and Linux `syscall`; no CGO, SQLite,
or third-party module was added.
## Concerns
- The optional whole-module `go vet ./...` reports a pre-existing Task 1 test warning at
`internal/credential/credential_test.go:86` (`append` with no variadic values). The identical
line is present in approved HEAD `1e82fd3`, outside this task’s authorized files. Focused Task 2
vet passes, and all module tests pass.
- The official image's login shell resets `PATH` and hides `/usr/local/go/bin`; all evidence uses
direct `go`/`gofmt` container entrypoints, which preserves the image’s Go 1.26.5 environment.
- The generic `apply_patch` helper intermittently failed before file access with a sandbox network
namespace error. Exact scoped corrections were applied through the shared worktree workflow;
this did not affect the final staged file set or verification evidence.
## Review remediation — 2026-08-21
### Scope and fix commit
- Review-fix commit: `971a0e6 fix: harden DWH credential registry reads`.
- Committed files only:
- `tools/dwh-auth/internal/registry/store.go`
- `tools/dwh-auth/internal/registry/store_test.go`
- The separate Task 1 vet correction is the independent preceding commit `d2415b5`; it is not
included in this Task 2 fix commit. No filesystem primitive, server, Nginx, service, registry,
secret, Docker stack, or legacy ThothII file was changed.
### Strict TDD evidence
All commands again used the official `golang:1.26.5` image with only this linked worktree mounted
at `/work`.
#### RED
The first focused command was run after the new regression tests and before production changes:
```text
go test ./internal/securefile ./internal/registry -count=1
```
It failed as intended. The three case-variant aliases (`SECRET_SHA256`, `Secret_SHA256`, and
`Schema_Version`) were accepted; past expiry returned active records from both `Find` and
`FindLegacy`; a revocation snapshot let readers return active data before publication; a temporary
file let `List`, `Check`, and `FindLegacy` observe false integrity failures; and the original
concurrent-read regression observed `ErrNotFound` during revocation.
The deterministic exact-expiry test was then added before the clock implementation. Its focused
run failed as intended with:
```text
internal/registry/store_test.go:572:10: store.now undefined
```
#### GREEN and verification
After the minimum implementation and `gofmt`:
```text
go test ./internal/securefile ./internal/registry -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 0.014s
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 0.684s
go test -race ./internal/securefile ./internal/registry -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 1.022s
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 1.711s
go vet ./internal/securefile ./internal/registry
```
The focused vet output was empty (success). Additional final checks passed:
```text
go test ./... -count=1
ok internal/credential
ok internal/record
ok internal/registry
ok internal/securefile
go vet ./...
go test -race ./internal/registry -count=10
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 8.127s
git diff --cached --check
```
### Remediated security invariants
- `Find` and `FindLegacy` now deny an active record when `ExpiresAt <= now.UTC()`, returning the
existing non-disclosing `ErrNotFound`. The unexported per-Store `now` function is the minimal
deterministic clock seam; past, exact-equality, and future cases are covered for v1 and legacy
records. A revoked record is still consulted before expiry and therefore remains authoritative.
- One per-Store `sync.RWMutex` creates an in-process consistent snapshot. `Add` and `Revoke` hold
it exclusively for their full writer-lock lifetime, including temporary-file publication and
revoked-then-active removal. `Find`, `FindLegacy`, `List`, and `Check` hold a shared lock; their
bodies delegate only to unlocked helpers, preventing nested-lock deadlocks. `Close` also takes
the exclusive lock before closing descriptors.
- Deterministic regression tests hold the writer path at the revocation publication/unlink and
temporary-file stages. They prove public readers wait, then see either the final revoked state or
a clean directory, eliminating the `Names`-to-load/unlink and temporary-entry false failures
within the Store contract.
- Before struct decoding, the outer record JSON object now requires exactly spelled keys from the
schema allowlist and rejects duplicate literal keys. `Decoder.DisallowUnknownFields`, recursive
duplicate detection, trailing-value rejection, record validation, filename matching, no-follow
reads, modes, and durability ordering remain intact.
### Self-review and concerns
- Reviewed the new lock boundaries, error returns, revoked-first ordering, clock fallback,
JSON-token consumption, and every unchanged `securefile` syscall/path/mode boundary. The change
adds only standard-library `sync`; it does not relax existing fail-closed behavior.
- The synchronized snapshot is intentionally per `Store`, matching the requested in-process
contract. The existing protected advisory lock continues to serialize writers across Store
instances/processes; no cross-process reader snapshot is claimed by this fix.
- The historical whole-module vet concern in the original Task 2 report is now resolved by the
independent Task 1 commit `d2415b5`; complete module vet passes in the final evidence above.
## Cross-Store snapshot remediation — 2026-08-21
### Scope and TDD evidence
This third Task 2 fix wave changes only the protected lock primitive and registry snapshot code:
- `tools/dwh-auth/internal/securefile/securefile_linux.go`
- `tools/dwh-auth/internal/securefile/securefile_linux_test.go`
- `tools/dwh-auth/internal/registry/store.go`
- `tools/dwh-auth/internal/registry/store_test.go`
All commands used the official `golang:1.26.5` image with only this linked worktree mounted at
`/work`.
The test-only red patch initially tried to inspect the unexported `securefile.Dir.fd` through the
registry package and therefore did not compile. That assertion was removed without production
changes: the registry tests still create writer Store A and reader Store B through two independent
`Open(root)` calls, while the securefile test proves separate descriptors directly in its own
package. The subsequent behavioral RED run, before the production change, was:
```text
go test ./internal/securefile ./internal/registry -count=1
FAIL TestLockSharedAllowsReadersAndBlocksExclusiveWriter: Dir lacks shared advisory locking
FAIL TestCrossStoreReadersWaitAcrossRevokePublicationAndUnlink:
Find, List, Check, and FindLegacy completed during Store A's revocation snapshot
FAIL TestCrossStoreScanReadersWaitForWriterTemporaryFile:
Store B's List, Check, and FindLegacy observed `.tmp-regression`
```
### Delivered synchronization contract
- `securefile.Dir.LockShared` now acquires `LOCK_SH` on the same protected, no-follow, exact-0600
root lock file used by `Lock`, which continues to acquire `LOCK_EX`. The lock file is still
opened/created, mode-validated, inode-checked, and closed through the existing Linux syscall
path.
- Every public snapshot reader (`Find`, `FindLegacy`, `List`, and `Check`) takes its Store
`RLock`, then a shared advisory lock on root `.writer.lock`, and retains both through the whole
revoked/active lookup or directory scan/load. `Add` and `Revoke` retain Store `Lock`, then the
same root lock under `LOCK_EX`, over their full operation.
- The lock order is universally Store mutex then root advisory lock. Public methods delegate only
to unlocked helpers, so neither reader nor writer paths recursively acquire the Store mutex.
`Close` retains its exclusive Store mutex, preventing descriptor closure from racing any locked
reader or writer.
- Revoked-first precedence, expiry denial, exact JSON validation, no-follow checks, record modes,
temporary-file durability, and all previous behavior remain unchanged.
### GREEN and repeated verification
```text
go test ./internal/securefile ./internal/registry -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 0.019s
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 0.711s
go test -race ./internal/securefile ./internal/registry -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 1.032s
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 1.748s
go vet ./internal/securefile ./internal/registry
go test ./... -count=1
ok internal/credential
ok internal/record
ok internal/registry
ok internal/securefile
go vet ./...
go test -race ./internal/registry \
-run 'TestCrossStoreReadersWaitAcrossRevokePublicationAndUnlink|TestCrossStoreScanReadersWaitForWriterTemporaryFile' -count=20
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/registry 9.880s
go test -race ./internal/securefile \
-run TestLockSharedAllowsReadersAndBlocksExclusiveWriter -count=20
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/securefile 1.063s
git diff --check
```
Both vet commands and the whitespace check produced no output. The securefile regression opens
three protected directory descriptors, proves they are distinct, permits two independent shared
holders, proves a third descriptor cannot take `LOCK_EX|LOCK_NB`, then proves exclusive acquisition
succeeds after shared release. The registry regressions deterministically block Store B readers
while Store A holds the exclusive root lock and verify only final revoked/clean states afterward.
### Self-review and concerns
- Reviewed lock creation/reopen races, no-follow flags, exact lock-file mode validation, lock
release, descriptor lifetime, lock ordering, error wrapping, and the unlocked-helper call graph.
No public reader invokes another public reader or writer while holding a Store lock.
- Advisory synchronization necessarily covers cooperating registry Store instances/processes;
arbitrary external filesystem mutation remains fail-closed through the existing integrity
checks rather than being silently accepted.
- No known concerns within the registry's cooperating-process contract.
-129
View File
@@ -1,129 +0,0 @@
# Task 3 report — secret-safe dwh-auth administrative CLI
## Scope
- Added `tools/dwh-auth/internal/command/command.go`, its command tests, and
`tools/dwh-auth/cmd/dwh-auth/main.go`.
- The CLI accepts only the frozen Task 3 grammar: key create/import/list/status/revoke,
registry check, and the reserved serve invocation.
- Existing Task 1–2 APIs are consumed without modifying their files.
- No server, Nginx, systemd, Compose, portable `tht`, real registry, real secret, or legacy
stack was accessed or changed.
## TDD evidence
Tests were written before `Run` existed. In the official `golang:1.26.5` container, mounted
against only the dedicated worktree, the focused RED run was:
```text
go test ./internal/command -count=1
internal/command/command_test.go:214:10: undefined: Run
FAIL
```
After implementation and formatting:
```text
go test ./internal/command -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/command
go test -race ./internal/command -count=1
ok github.com/aritmolab/thothii/tools/dwh-auth/internal/command
go test ./... -count=1
ok internal/command, internal/credential, internal/record, internal/registry, internal/securefile
go test -race ./... -count=1
ok internal/command, internal/credential, internal/record, internal/registry, internal/securefile
go vet ./...
```
## Contract coverage
- Create generates the Task 1 canonical credential, writes it once through the protected
exclusive `0600` output primitive, syncs/closes it before registry publication, and emits
only `created key_id=... installation_id=... output=...`.
- Existing output is never overwritten. Publication failure attempts compensating removal;
cleanup uncertainty returns exit 4 and reports only the output path.
- Legacy import requires `--legacy-raw`, the reserved `legacy-shared` installation ID, and an
absolute exact-`0600` source. It verifies the opaque value, never changes its source, and
stores only its digest.
- List/status expose `PublicRecord` data only; JSON is written as pristine JSON with no digest.
Revoke requires a non-empty reason and reports only its public key ID.
- Relative paths, malformed/unknown flags, duplicate options, invalid IDs/metadata/expiry,
and missing required values return exit 2. Missing status/revoke keys return exit 3.
Registry/filesystem/integrity failures return exit 4.
- Diagnostics are fixed redacted strings. Tests use a sentinel secret and assert it is absent
from stdout/stderr, list/status/check JSON, import output, and unsafe/integrity failures.
## Secret-redaction evidence
The command never prints credential contents or digests. It does not use environment fallback,
interactive stdin, or `flag` diagnostics that echo argument values. The sentinel appears only
in synthetic temporary test input and an integrity-fixture file; all command output assertions
confirm it is absent. The registry’s existing `PublicRecord` contract omits `secret_sha256`.
## Concerns
- `serve` is grammar-reserved and returns a redacted exit-4 unavailable response; Task 4 owns
the Unix-socket service implementation and will wire this dispatch.
- The earlier concern about UTF-8 metadata hardening is superseded by `055dcab`: metadata now
rejects invalid UTF-8 and Unicode controls before generation/output. The nil-safe cleanup note
remains non-blocking and outside this review wave.
## Review-fix wave
Review findings were addressed in separate commit `055dcab`. Regression tests were added first. The focused RED run in the official Go 1.26.5 container failed on intentionally absent seams:
```text
undefined: nowUTC
undefined: addRecord
undefined: closeStore
FAIL github.com/aritmolab/thothii/tools/dwh-auth/internal/command
```
The fix rejects embedded canonical v1 credentials in description/revocation reason without echoing metadata, validates UTF-8/Unicode controls and expiry against one captured UTC creation time before generation/output, reserves exactly `serve --registry-root ABS --socket ABS`, and makes publication cleanup depend on a definitive registry lookup. Output is retained after publication or close ambiguity, with path-only recovery guidance.
Review-fix verification in Go 1.26.5:
```text
go test ./internal/command -count=1 PASS
go test -race ./internal/command -count=1 PASS
go test ./... -count=1 PASS
go test -race ./... -count=1 PASS
go vet ./... PASS
git diff --check PASS
```
New tests cover synthetic canonical credentials embedded with prefix/suffix, invalid UTF-8, C1 Unicode controls, past/equal/future expiry, exact serve ordering, deterministic pre-/post-publication and close-failure seams, and sentinel absence from stdout/stderr/list/status JSON.
## Cleanup snapshot review-fix wave
The second re-review added two regression tests before implementation. The RED run in the
official Go 1.26.5 container showed the old `Find` proof incorrectly treated both cases as
cleanup-safe:
```text
FAIL TestCreateRetainsOutputWhenSnapshotFindsUnrelatedIntegrityFailure
corrupt snapshot result = (4, "", "integrity failure\n")
FAIL TestCreateRetainsOutputWhenFailedPublicationRecordIsExpired
expired publication result = (4, "", "integrity failure\n")
```
Commit `b1079bd fix: retain DWH key output on ambiguous publication` replaces the `Find` proof
with a complete `Store.List()` snapshot. It removes generated output only when the snapshot
succeeds, the generated key ID is absent, and `Store.Close()` succeeds. Any unrelated integrity
error, active/revoked/expired record, or close error retains the output and emits only path-based
recovery guidance. The clean pre-publication failure path still removes the output.
Final cleanup-wave verification in Go 1.26.5:
```text
go test ./internal/command -count=1 PASS
go test -race ./internal/command -count=1 PASS
go test ./... -count=1 PASS
go test -race ./... -count=1 PASS
go vet ./... PASS
git diff --check PASS
```
-60
View File
@@ -1,60 +0,0 @@
# Task 4 — Unix-socket DWH verification report
## Scope
Implemented the standalone Linux verifier at `tools/dwh-auth/internal/service` and wired the exact command:
```text
dwh-auth serve --registry-root ABSOLUTE_CANONICAL --socket ABSOLUTE_CANONICAL
```
The service accepts only `GET /verify`. It returns empty `204` responses with `X-DWH-Key-ID` for verified v1 or reserved legacy credentials; credential failures are generic empty `401` responses, and registry/integrity faults are empty `503` responses. Other paths/methods return empty `404`/`405`.
## Security decisions
- Exactly one `X-API-Key` header, maximum 128 bytes.
- Strict `thtdwh_v1.` parsing precedes legacy lookup; non-v1 values alone may use the reserved legacy record.
- Registry integrity is checked before every verification request, so unrelated malformed/unsafe records fail closed with `503`.
- Logs emit only timestamp, decision code, and (when safely parsed or verified) public key ID; test sentinels prove no key, digest, description, or query value is emitted.
- `serve` validates canonical absolute paths, performs startup `Store.Check`, and reports service startup errors as non-secret `integrity failure`.
- Socket collisions that are regular files, directories, symlinks, live sockets, or foreign-owned stale sockets are refused. Only an owned stale Unix socket after `ECONNREFUSED` can be reclaimed.
- Published sockets are mode `0660`; cancellation calls graceful shutdown and removes only a revalidated same-device/same-inode owned socket. A test seam proves a changed path is retained rather than unlinked.
## Required supporting security fix
Commit `943f809` (`fix: reject duplicate legacy DWH records`) tightens the Task 2 registry contract: a synthetically valid active plus revoked legacy pair is now an integrity failure. It is intentionally separate from the Task 4 commit.
## TDD evidence
RED was observed for the missing handler, listener/configuration API, CLI wiring, unrelated-registry corruption, active+revoked legacy state, and cleanup replacement race. Each increment was then implemented minimally and rerun GREEN.
## Verification
All commands were executed in official `golang:1.26.5`, with only this worktree mounted:
```text
gofmt -w cmd internal/command internal/service
go test ./internal/service ./internal/command -count=1
go test ./... -count=1
go test -race ./... -count=1
go vet ./...
git diff --check
```
All passed. A dependency scan also found no third-party Go dependencies.
## Scope boundary
No Nginx, systemd, real Unix socket, real registry, credential, legacy stack, or external service was changed. All test data was synthetic and temporary.
## Follow-up hardening: runtime read-only registry and socket parent
The Task 5 storage contract uses `root:dwh-auth` SGID directories (`2750`) and a service account with read-only group access. The original registry reader path was incompatible because shared locks were opened `O_RDWR` and lazily created as `0600`; secure-directory validation also rejected SGID.
The runtime path now uses `registry.OpenReadOnly`: it opens only preprovisioned root, `active`, `revoked`, and `.writer.lock` paths, and rejects `Add`/`Revoke`. The administrative `Open` path bootstraps the lock through the exclusive writer path. Shared lock acquisition opens the existing `root:dwh-auth 0640` lock `O_RDONLY` with `LOCK_SH`; writer acquisition remains `O_RDWR` with `LOCK_EX`, preserving cross-process snapshot exclusion. Secure directories allow SGID but still reject setuid, sticky, group-write, and world-write bits.
Task 5 must create `.writer.lock` as `0640 root:dwh-auth` alongside the `2750 root:dwh-auth` registry directories before the service starts.
The socket parent must be a canonical non-symlink directory owned by the service EUID and not group/world writable. This removes the bind-to-chmod and path-replacement exposure from other principals. The remaining POSIX path race is bounded to trusted processes sharing the service EUID inside that non-contendible parent.
Additional verification (official `golang:1.26.5`, worktree only): focused securefile/registry/service/command tests, full tests, full race tests, vet, plus ten race repetitions each for cross-store snapshot readers, `OpenReadOnly`, and listener tests: all PASS.
-71
View File
@@ -1,71 +0,0 @@
# Task 5 report — backend principal enforcement
## RED
Added backend route/auth tests before implementation. The initial focused run failed in
seven new assertions: `getPrincipal` did not exist, upstream requests still required the
legacy identity header, foreign session/SSE routes were not hidden, admin scope was not
enforced, new sessions had no trusted principal binding, and settings were global.
## GREEN
- Focused backend suite: `66 passed` across auth, sessions, SSE, and settings tests.
- Complete backend Vitest suite: `209 passed` across `22` files.
- `npx tsc --noEmit -p .`, `npm run build`, `git diff --check`, and changed Python
source Ruff all exit successfully.
- Harness targeted repository/local/migration tests and Python bytecode compilation exit
successfully. The new `tht session preferences get|set` commands are registered and
expose the expected Typer help. A direct local CLI preference smoke was not run because
the checked-in local workspace requires unavailable `THT_DB_HOST` configuration.
## Route and child-process coverage
- `GET /me` returns the request `PrincipalContext`; upstream accepts only the portal's
normalized `X-Thoth-*` identity tuple, with the legacy header ignored. Local mode uses
the same stable `THT_HOME`/`~/.thothii/identity.json` UUID contract as the harness.
- All session operations are principal-scoped: list (`mine` and admin-only `all`), show,
create, resume, close, delete, rename, group, archive, unarchive, documents, reviewer
response, steer, SQL preview/export, and SSE. Missing and foreign sessions are 404;
absent upstream identity is 401. SSE is authorized before response headers or hub
subscription, so a rejected request cannot attach to a live stream.
- New/resumed Pi runtimes and every route-spawned `tht` process receive
`THT_PRINCIPAL_ISSUER`, `THT_PRINCIPAL_SUBJECT`, optional display name, and admin flag.
The readiness `tht` child is also principal-bound.
- Settings use asynchronous repository-backed `tht session preferences get|set` in the
production runner, which isolates preferences by principal. The legacy settings file is
retained only as an injected-runner compatibility fallback for existing isolated tests.
- Repository/settings authorization failures map to 503 before model startup. SQL execution
errors remain 500 after authorization, preserving the prior API distinction.
## Self-review and concerns
- Confirmed the Task 4 portal emits lowercase `true`/`false` for the admin header; the
parser accepts that exact normalized form plus the repository's existing `1`/`0`
compatibility form, and rejects all other values.
- The harness principal resolver is the ownership authority; the backend never accepts an
owner supplied in request bodies. Its route guards use a repository-scoped `session show`
before every session resource operation.
- Existing dependency-injected route fakes without `sessionShow` retain a narrow test seam;
production `ThtRunner` always has that method, so deployed requests cannot bypass the
repository authorization check.
## Review follow-up
### RED
Focused regressions initially failed exactly at the three review findings: stale ambient
display names survived into both `tht` and Pi child environments; mutation/document runner
methods dropped the selected workspace; and `expandLocalHome` did not exist.
### GREEN
- Child environments now remove all four `THT_PRINCIPAL_*` keys from their cloned base
environment before applying the exact request principal. Regression tests prove an absent
display name does not inherit a stale ambient value in either child path.
- `setName`, `setGroup`, `archive`, `unarchive`, and `documents` now take and retain an
optional workspace. The rename route regression proves `session show` authorization and
the mutation use the same non-default workspace.
- Local principal paths expand `~`/`~/...`; existing local home and identity file modes are
repaired to POSIX `0700`/`0600` when applicable, with Windows left unchanged.
- Focused suite: `74 passed`; full backend suite: `213 passed` across `22` files, followed by
TypeScript typecheck, production build, and diff check.
-298
View File
@@ -1,298 +0,0 @@
# Task 6 — Frontend identity and administrator UX report
## RED
- Added API tests for the `/me` principal call and `mine`/`all` session-list scopes.
- Added component tests for regular-user scope, admin scope switching, owner labels,
administrator banner, foreign-owner delete confirmation, and foreign-owner archive
confirmation.
- Initial focused run: 7 expected failures (missing `getMe`, missing scope query,
missing owner label/admin controls, and missing foreign-action confirmation).
- The archive-confirmation regression was also run separately before its implementation
and failed because `window.confirm` was not called.
## GREEN
- `npx vitest run src/api/sessions.test.ts src/shell/NavSessions.test.tsx src/shell/AppShell.session-mgmt.test.tsx`
— passed (47 tests before the archive follow-up; the focused archive regression then passed).
- `npm test` — passed: 44 files / 305 tests.
- `npx tsc -b` — passed.
- `npm run build` — passed.
- `git diff --check` — passed.
- `npm run e2e` reached Playwright but could not run: the environment has no Chromium
executable at Playwright's configured cache path. No application test failure was reported.
## Files changed
- `frontend/src/api/types.ts`: typed principal and session scope contracts.
- `frontend/src/api/sessions.ts`: typed `/me` API call; scoped listing defaults to `mine`.
- `frontend/src/shell/AppShell.tsx`: identity query, admin-only session scope selector and
banner, owner-aware destructive action confirmations.
- `frontend/src/shell/NavSessions.tsx`: owner labels in the all-sessions view.
- `frontend/src/api/sessions.test.ts`, `frontend/src/shell/NavSessions.test.tsx`, and
`frontend/src/shell/AppShell.session-mgmt.test.tsx`: contract and UX coverage.
## Self-review
- Regular users remain fail-closed on `mine`; no administrator control renders without
`principal.isAdmin`.
- The all-sessions view includes owner labels (including `Unknown` for legacy records).
- Delete confirmation preserves the pre-existing select-all behavior and adds confirmation
for foreign/unknown owners. Foreign archive now also requires an explicit browser
confirmation; existing Stop & save already has its confirmation dialog.
- A read-only review found no critical, important, or minor issues. The archive guard was
added after that review in response to the requirement to cover every destructive rail
action, and has its own RED/GREEN regression plus the final full verification above.
## Concerns
- E2E remains environment-blocked until the Playwright Chromium browser is installed.
- Existing Vitest runs emit pre-existing MSW unmatched-request and dialog-ref warnings; all
assertions pass and this task does not modify those shared test/UI primitives.
## Review remediation
- A post-commit review correctly identified that matching `displayName` must never establish
ownership. The predicate now skips confirmation only when `session.author` exactly equals
`principal.subject`; all display-name matches and missing authors are conservative
cross-owner actions.
- Added RED/GREEN regressions where two principals share display name `Alice` but have distinct
subjects: both delete (with another session present, so select-all cannot mask the guard) and
archive require confirmation.
- Added `aria-pressed` to the My sessions / All sessions controls and asserts their selected state
before and after switching.
- Remediation verification: focused regressions passed; full frontend Vitest (44 files / 305
tests), `npx tsc -b`, `npm run build`, and `git diff --check` all passed.
---
# DWH authentication Task 6 — Nginx and CI gate report
## Scope
Added only the two DWH-auth Nginx gates and the `dwh-auth-linux` deployment workflow job:
- `scripts/test-dwh-auth-nginx-contract.sh`
- `scripts/test-dwh-auth-nginx-integration.sh`
- `.github/workflows/deployment.yml`
This report deliberately remains unstaged. The pre-existing frontend Task 6 report above is
preserved rather than overwritten.
## TDD RED
The structural gate was written before any Task 5 template change. Those templates already met
the approved contract, so the behavioral RED was obtained by copying them into one exact temporary
root and removing only the effective `/dwh/` `auth_request` directive. The new checker failed as
required, with no credential material in output:
```text
case=source_contract status=FAIL
```
The runtime gate was also first invoked before its file existed:
```text
bash: scripts/test-dwh-auth-nginx-integration.sh: No such file or directory
```
The CI-job RED check found no `dwh-auth-linux` job in `deployment.yml`. No production template was
modified: the tests prove the existing Task 5 template contract instead of weakening it.
## GREEN
Shell syntax and workflow YAML were checked with:
```text
bash -n scripts/test-dwh-auth-nginx-contract.sh scripts/test-dwh-auth-nginx-integration.sh
python3 -c import-yaml-and-safe-load
```
The structural gate passed its source contract plus these 13 real copied-and-mutated Nginx fixtures:
```text
missing_auth_request
missing_proxy_method
missing_proxy_body
missing_proxy_header_isolation
missing_content_length_clear
missing_verifier_key_forward
missing_upstream_key_clear
missing_failure_mapping
public_verifier
tcp_authenticator
postgrest_bypass
failure_mapped_to_success
full_secret_rate_key
```
Each test mutates an effective, not comment-only, directive and requires the checker to reject it.
The source test and all 13 fixture tests emitted `case=... status=PASS`, followed by
`case=summary status=PASS`.
The isolated Nginx 1.24 smoke passed these sanitized cases:
```text
nginx_1_24
build_dwh_auth
registry_setup
verifier_start
synthetic_upstreams
composite_nginx_config
nginx_start
auth_socket_unix_only
verifier_not_public
valid_v1
valid_legacy
invalid_key
revoked_key
expired_key
duplicate_v1
duplicate_legacy
stopped_verifier
header_and_path_isolation
summary
```
It builds with the pinned official Go 1.26.5 image when the host Go binary is absent, creates only
synthetic v1, legacy, revoked, and expired credentials in a `0700` `/tmp` root, runs both Nginx and
the verifier on explicit temporary Unix sockets, and uses a loopback-only marker backend. Its output
is strictly `case` and `status`; keys, values, and digests remain only in the exact temporary root
and are removed by the trap.
`nginx -t` passed against the complete generated configuration. The marker proves that successful
`/dwh/?keep=exact&second=two` reaches the upstream unchanged, while neither the client API key nor
client or verifier `X-DWH-Key-ID` reaches it. A Unix forwarding probe proves that the verifier sees
only `X-API-Key`, with Cookie, Authorization, and spoofed audit ID absent. Duplicate v1 and ordinary
legacy headers return 401 through Nginx; a stopped verifier returns 503.
The final local equivalent of the four CI commands passed:
```text
Docker Go 1.26.5: go test -race ./... -count=1 and go vet ./...
bash scripts/test-dwh-auth-build-contract.sh
bash scripts/test-dwh-auth-nginx-contract.sh
bash scripts/test-dwh-auth-nginx-integration.sh
```
The Go race suite passed for command, credential, record, registry, securefile, and service;
`go vet` was silent; the build contract passed; both Nginx gates reached their summaries.
## CI contract
The new job uses `actions/checkout` with `persist-credentials: false`, pins Go 1.26.5 with cache
keyed on `tools/dwh-auth/go.mod`, installs `nginx-light`, and runs exactly the four required commands.
Existing jobs were not altered.
## Self-review
- The template tests parse normalized effective directives, so commented-out declarations cannot
satisfy the gate.
- The authentication socket is configured as `http://unix:...:/verify`, is observed by `ss -xl`,
and Nginx itself listens only on a temporary Unix socket; neither test starts a public listener.
- All spawned processes are registered by PID; cleanup signals only those PIDs and deletes only the
exact `mktemp` root after a guarded path check.
- The verifier, marker, registry, Nginx prefix, PID, logs, config, and sockets all reside beneath
that root. No `/etc`, systemd, active Nginx config, stack, legacy route, or real registry/key is
read or changed.
- Task 5 templates were not modified because the structural and runtime tests passed unchanged.
## Concern
The sandbox `apply_patch` helper repeatedly failed with `bwrap: loopback: Failed RTM_NEWADDR:
Operation not permitted`. A narrowly scoped fallback editor was used only for the workflow and the
Nginx-version assertion. Its first workflow insertion interpreted the action-reference at signs;
the two malformed values were immediately corrected and all final YAML, exact-string, syntax, and
four-command checks were rerun. No remaining product concern is known; the integration gate requires
Nginx 1.24 and Python 3, both supplied by the specified Ubuntu CI runner.
---
# DWH authentication Task 6 — review remediation wave
## Review findings and RED evidence
The three review findings were reproduced against the Task 6 commit before their corresponding
hardening was accepted.
1. The contract checker originally selected only the first matching `/dwh/` location. A real copied
fixture appended this competing location without authentication:
```nginx
location ~ ^/dwh/ {
proxy_pass http://127.0.0.1:3001;
}
```
The first run reached the new check and failed as required:
```text
case=negative_postgrest_regex_bypass status=FAIL
```
2. The previous process stop sent TERM and immediately used an unbounded `wait`. A synthetic Python
child ignored TERM; the RED run used one exact short-lived watchdog only to prevent a test hang and
produced:
```text
case=cleanup_term_ignored_bounded status=FAIL
```
3. The TCP detector has a positive-control regression. A scratch copy of the integration script
replaced its `ss -ltnpH` detector with `return 1`; its known loopback listener was then not
detected and the run failed with:
```text
case=tcp_listener_detector_positive status=FAIL
```
All RED fixtures and the scratch script used an exact temporary path and were removed. No template,
service, workflow, key, or active Nginx configuration was changed.
## GREEN changes
- `location_declarations` consumes normalized, comment-stripped effective lines and `check_templates`
requires exactly one each of the only approved locations: verifier, unavailable named location, and
`/dwh/`. It therefore rejects both any extra intercepting location and a duplicate. The real regex
bypass and a new real duplicate `/dwh/` bypass fixture both pass by being rejected.
- `tcp_listener_for_pid` uses `ss -ltnpH` and a PID-bound match. The integration gate starts a
loopback-only synthetic listener, proves the detector sees that exact PID, stops and deregisters it,
then proves the verifier PID has no TCP listener while its Unix socket remains present.
- `stop_registered_pid` now sends TERM, polls for exit or zombie for a bounded deadline, sends KILL
if required, polls a second bounded deadline, and only reaps a direct child after terminal state is
proved. Explicit stops deregister their PID. The cleanup loop invokes that bounded operation only
for recorded PIDs and removes only its guarded temporary root.
- The synthetic child that ignores TERM is killed by the bounded path, must no longer answer to
`kill -0`, must not remain registered, and must finish within three seconds. Final gate output is
restricted to `case` and `status` lines.
## GREEN verification
```text
bash -n scripts/test-dwh-auth-nginx-contract.sh scripts/test-dwh-auth-nginx-integration.sh
Docker Go 1.26.5: go test -race ./... -count=1 and go vet ./...
bash scripts/test-dwh-auth-build-contract.sh
bash scripts/test-dwh-auth-nginx-contract.sh
gate contract: source plus 15 negative fixtures PASS, then summary PASS
bash scripts/test-dwh-auth-nginx-integration.sh
gate integration: 20 named cases PASS, then summary PASS
git diff --check
```
The integration cases include `cleanup_term_ignored_bounded`,
`tcp_listener_detector_positive`, `auth_socket_unix_only`, all existing credential decisions,
composite Nginx syntax, and stopped-verifier 503 behavior. Go race tests passed for command,
credential, record, registry, securefile, and service; vet and both diff checks were silent.
## Self-review and concern
The new location parser rejects comment-only and non-exact declarations because it operates on the
same normalized effective representation used by the rest of the contract. The TCP positive control
binds only `127.0.0.1` on a kernel-selected temporary port and is stopped through the same exact-PID
path under test. The bounded cleanup avoids arbitrary process lookup or broad signaling.
The environment still intermittently rejects `apply_patch` with the sandbox loopback error noted in
the original report; only narrowly scoped fallback edits to the two authorized scripts were used and
all final gates were rerun. No remaining review concern is known.
-90
View File
@@ -1,90 +0,0 @@
# Task 7 — report
## RED
- Creato `scripts/test-verify-dwh-auth-docs.sh` con fixture positiva e fixture negative per
credenziale/digest sintetici, TLS insicuro, segreto in env/argv, mode world-readable, cattura
Nginx e coupling Compose.
- Eseguito `bash scripts/test-verify-dwh-auth-docs.sh` prima del verificatore: `case=verifier_missing status=FAIL`.
## GREEN
- Aggiunti manuali server, client, TLS, runbook PSD, collaudo ed evidenza sanitizzata; collegati
manuali locali/server, setup PSD, guida, indice e nav MkDocs.
- Eseguiti: `bash -n scripts/verify-dwh-auth-docs.sh scripts/test-verify-dwh-auth-docs.sh`,
`bash scripts/test-verify-dwh-auth-docs.sh`, `bash scripts/verify-dwh-auth-docs.sh`,
`bash scripts/test-verify-workspace-install-docs.sh`, `bash scripts/auth-docs-smoke.sh`.
- Tutti gli output finali sono PASS; il nuovo gate esercita una fixture positiva e nove negative.
## Self-review
- Verificati path/owner/mode: registry 2750, lock/record 0640, socket 0660.
- Verificata separazione: chiavi solo `rest_api`; PSD server `postgres_direct`; Mac/remoti REST;
nessun lifecycle Compose per `dwh-auth`.
- Verificati TLS `.it`/SAN, `.com` non coperto, `TLS_CA_FILE`, fingerprint fuori banda, rinnovo e
assenza di bypass.
- Verificati due gate Task 9–10, evidenze solo metadati e nessuna migrazione di sessioni/index/cache legacy.
## Concern
- Nessuna mutazione PSD/Nginx/systemd/registry o lettura di segreti è stata eseguita. I comandi del
runbook restano condizionati alle autorizzazioni separate dei Task 9 e 10.
## Review fix — RED/GREEN
### RED review
- La fixture `sudo nginx -T` ha prodotto il rifiuto `case=sudo_raw_nginx_capture status=FAIL` prima della correzione del gate.
- La fixture header legacy opaco ha prodotto `case=opaque_legacy_header_literal status=FAIL` prima della correzione del gate.
- Dopo avere riallineato le label UI nei manuali, `bash scripts/test-verify-workspace-install-docs.sh` ha prodotto `server-workspace-registry.md: curator flow missing registry rule`: il verifier cercava ancora le due label precedenti. Il test sulla base HEAD e il diff hanno confermato la causa.
### GREEN review
- Il gate DWH ora rifiuta anche header opaco, digest JSON quotato, `export` di API key, `curl --header` e `-H`, `sudo nginx -T`, raw diff e Compose; le mutation fixture coprono label, PSD direct/Mac REST/CA, socket e flag REST.
- Il runbook non prescrive raw diff o dump: solo checker strutturale e secret scan con metadati e PASS/FAIL. Il piano Task 10 adotta la stessa regola.
- Il template `psd-local` resta `rest_api` solo Mac/local/remota; il server PSD Project A resta `postgres_direct` con binding separato. La CA privata e `TLS_CA_FILE` sono obbligatori salvo trust approvato equivalente.
- Le procedure server ora coprono backup manifest protetto, restore, curl config 0600 senza segreto in argv/env/output, Unix 204/401, HTTPS 2xx/401, 503 bounded con trap, journal PASS/FAIL e retention alla disinstallazione.
- Il verifier workspace-install e entrambi i manuali registry usano ora le quattro label effettive: `Validate workspace source`, `Test workspace connections`, `Save entered secrets`, `Forget stored value`.
### Final verification review
- PASS: `bash scripts/test-verify-dwh-auth-docs.sh`.
- PASS: `bash scripts/verify-dwh-auth-docs.sh`.
- PASS: `bash scripts/test-verify-workspace-install-docs.sh` (fixture complete).
- PASS: `bash scripts/auth-docs-smoke.sh`.
- PASS: `bash -n scripts/verify-dwh-auth-docs.sh scripts/test-verify-dwh-auth-docs.sh` e `git diff --check`.
### Review concern
- Nessuna mutazione runtime e nessun segreto reale sono stati letti. I soli comandi server documentati restano soggetti ai gate autorizzativi Task 9 e Task 10.
## Review fix wave 2 — RED/GREEN
### RED wave 2
- Prima della correzione del proxy, `bash scripts/test-dwh-auth-build-contract.sh` ha fallito il contratto di preservazione path e `bash scripts/test-dwh-auth-nginx-integration.sh` ha chiuso con `case=header_and_path_isolation status=FAIL`: il prefisso `/dwh` arrivava a PostgREST invece di essere rimosso.
- Prima delle procedure finali, il gate docs ha rifiutato il path chiave non deterministico e la fixture curl con header legacy opaco ha dato `case=header_file_curl_synthetic status=FAIL` perché il valore non veniva confrontato esattamente.
- Le mutation fixture hanno catturato l'estrazione tar sul registro attivo e i rename non protetti. Dopo l'inasprimento finale del gate, la sorgente ha dato `dwh-auth docs: restore must stage/check then use guarded same-filesystem renames` finché mancava il controllo fail-closed del candidato.
- Il RED finale dello scanner journal è stato `dwh-auth docs: docs/install/dwh-auth-server.md lacks required topic: sys.argv[2:]`: il gate esige la lettura byte-esatta di v1 e legacy e un `journalctl` che fallisca chiuso.
### GREEN wave 2
- Commit `f616aab fix: preserve PostgREST RPC path through DWH proxy`: `proxy_pass` termina con `/`; il contratto e l'integrazione verificano `/dwh/rpc/ping?x` verso `/rpc/ping?x`.
- Il runbook usa un singolo file chiave v1, header file `0600` passati solo con `curl --header @file`, socket 204 dual-key, HTTPS 2xx pre/post per v1 e 401 post-revoca per legacy `legacy-shared`.
- Restore protetto: staging sul filesystem `/var/lib`, check candidato, `mv -T --` guardato per ogni publish/rollback e pre-restore conservato. Backup/manifest restano root-only `0600` su storage cifrato approvato.
- Lo scanner journal esegue `journalctl` in un unico processo Python root, sopprime stderr, controlla return code e bytes esatti di entrambe le chiavi senza emettere journal o segreti; la shell mostra solo PASS/FAIL.
- Il verifier rifiuta `curl --config`, header in argv, raw Nginx/diff, TLS insicuro, segreti env, mode insicuri e Compose. Le fixture mutano path chiave, ID legacy, header/legacy probes, restore, journal, codici HTTPS e label UI.
### Final verification wave 2
- PASS: `bash scripts/test-dwh-auth-build-contract.sh`.
- PASS: `bash scripts/test-dwh-auth-nginx-contract.sh`.
- PASS: `bash scripts/test-dwh-auth-nginx-integration.sh`.
- PASS: `bash scripts/test-verify-dwh-auth-docs.sh` e `bash scripts/verify-dwh-auth-docs.sh`.
- PASS: `bash scripts/test-verify-workspace-install-docs.sh` e `bash scripts/auth-docs-smoke.sh`.
- PASS: `bash -n` sugli otto gate shell e `git diff --check`.
### Review concern wave 2
- Nessuna configurazione protetta, chiave reale, Nginx, systemd o stack PSD è stata letta o mutata. Le procedure privilegiate restano istruzioni condizionate ai Gate 9–10; la verifica degli owner/mode reali è un'attività del rollout autorizzato, non di questo task documentale.
+46 -12
View File
@@ -1,5 +1,19 @@
# AGENTS.md
## Agent skills
### Issue tracker
Issues for this repository live in the self-hosted Gitea repository at `https://git.tylconsulting.it/mptyl/ThothII`; use its web UI or authenticated Gitea API. See `docs/agents/issue-tracker.md`.
### Triage labels
Use the canonical labels `needs-triage`, `needs-info`, `ready-for-agent`, `ready-for-human`, and `wontfix`. See `docs/agents/triage-labels.md`.
### Domain docs
This is a single-context repository with root `CONTEXT.md` and `docs/adr/`. See `docs/agents/domain.md`.
This file provides guidance to Codex (Codex.ai/code) when working with code in this repository.
## Start here
@@ -7,7 +21,9 @@ This file provides guidance to Codex (Codex.ai/code) when working with code in t
Read [PROJECT_STATE.md](PROJECT_STATE.md) for the current-state snapshot: what was last
built, pending manual gates, workspace/secret layout, and design-doc locations. This file
holds the stable commands + architecture mental model; PROJECT_STATE.md holds the evolving
detail. Design history lives in `docs/superpowers/specs/` and `docs/superpowers/plans/`.
detail. Current architecture and contracts live in `docs/architecture/`, `docs/contracts/`,
and `docs/evidence.md`; durable design decisions live in `docs/adr/`. Git history is the source
for superseded designs and implementation plans.
## Commands
@@ -34,6 +50,10 @@ The repo has three independently-built layers. Run the local Docker stack with `
- Test: `npx vitest run` · Single: `npx vitest run src/shell/NavSessions.test.tsx`
- Typecheck: `npx tsc -b` · E2E: `npm run e2e` (Playwright)
**Documentation** (MkDocs, repository-locked Python dependencies)
- Strict build: `./scripts/build-docs.sh`
- Refresh lock: `./scripts/update-docs-lock.sh`
No ESLint on the TS layers — `tsc` is the gate. Tests use vitest + MSW (no network).
## Architecture (the parts that need multiple files to see)
@@ -56,12 +76,15 @@ frontend (React/SSE) → backend (Fastify) → pi --mode rpc → tht/harness →
There is no verbatim transcript store. A resumed Pi process rebuilds context from
`tht session show <id>` + the on-disk artifacts.
- **The backend is a thin bridge with no database.** `ThtRunner` shells the Python workflow `tht`
subcommands inside `core`;
- **The backend bridges sessions and owns the installation-local metadata catalog.** `ThtRunner`
shells the Python workflow `tht` subcommands inside `core`;
`PiProcessManager` runs one Pi child per session and bridges its RPC stream;
`SessionBridge` maps Pi RPC events → client events (`ui_request`/`text_delta`/`info`);
`SseHub` fans them out over SSE to the browser. App settings live in a JSON file
(`backend/data/settings.json`), not a DB.
`SseHub` fans them out over SSE to the browser. The separate PostgreSQL catalog stores database
metadata and sequential AI description-generation runs. Description generation samples the DWH
through read-only connectors and calls a short-lived Python LiteLLM helper; it does not use Pi or
expose a public CLI command. Sessions, metadata generation, and embedding resolve models from the
generated Installation Model Catalog; `thothii-installation.yaml` is its only authored source.
- **Human-in-the-loop gate contract.** The model proposes; a human reviewer decides at gates
via widgets (`reviewer_select` = single pick — a chosen option carrying a `decision` payload
@@ -75,13 +98,24 @@ frontend (React/SSE) → backend (Fastify) → pi --mode rpc → tht/harness →
- **`tht`'s `-c`/`--config` is a PER-COMMAND option** — it must follow the subcommand, never
precede it (`ThtRunner.buildArgv` enforces this; prepending caused live 500s).
- **`--json` output must be pristine** (only valid JSON on stdout) — used as a machine contract.
- **UI strings are English; document *content* stays the workspace language** (Italian for
`psd`) because it's the real data. Only chrome/labels are English.
- **Workspaces** (`harness/workspaces/*.yaml`) set the DB target and **absolute**
`paths.sessions/artifacts/indexes` — for `psd` these point at a *separate, uncommitted* repo
(`tht-workspace-psd/`). Secrets live ONLY in `harness/.env` (gitignored).
- **Settings are global** (`backend/data/settings.json`: workspace/provider/model/thinking);
the New-session form is question-only.
- **Localization:** deterministic UI uses the EN/IT catalogs with English fallback;
model interaction uses the session manifest's immutable `interaction_language`.
Workspace content, SQL, and identifiers remain unchanged. For shell modes, portal
integration, or translations, read `docs/operations/shell-and-localization.md`.
- **Server deployment:** for the coordinated ThothII/Omics upgrade, follow
`docs/operations/server-codex-handoff.md`; it supersedes earlier Omics delivery
instructions. Omics source integration uses GitHub with no repository relay prerequisite.
- **Server identity:** for portal login/logout, proxy headers, or Omics deploy, read
`docs/install/authentication-upstream.md` before changing authentication. Omics
uses embedded/upstream, not a second ThothII OIDC login. Full/embedded rendering
is documented in `docs/architecture/application-shell.md`; release acceptance
is in `docs/testing/authentication-manual-acceptance.md`.
- **Workspace schema v4** defines workspace identity and optional Evidence only. PostgreSQL Metadata
Catalog owns database identity, binding, schema, descriptions, sensitivity, and relationships;
embedding/model facts come from the installation catalog. The legacy `harness/workspaces/*.yaml` runtime snapshots still use
absolute session/artifact/index paths; secrets stay in `harness/.env` (gitignored).
- **Settings are global** (`backend/data/settings.json`: workspace/thinking). Provider/model choices
are ephemeral canonical catalog selections pinned into the session manifest.
- **Resume**: a resumable session re-enters at its last incomplete phase. The backend refuses
resume with 409 when `finalized` or `archived`, and `PiProcessManager.spawnFor` must send
`/riprendi-sessione <id>` (resume mode) vs `/nuova-domanda` (new) — sending the wrong prompt
-87
View File
@@ -1,87 +0,0 @@
# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## Start here
Read [PROJECT_STATE.md](PROJECT_STATE.md) for the current-state snapshot: what was last
built, pending manual gates, workspace/secret layout, and design-doc locations. This file
holds the stable commands + architecture mental model; PROJECT_STATE.md holds the evolving
detail. Design history lives in `docs/superpowers/specs/` and `docs/superpowers/plans/`.
## Commands
The repo has three independently-built layers. Run the **full stack** (real Pi + DWH, needs
VPN + `harness/.env` + `pi` on PATH) with `./scripts/run-stack.sh` (frontend :5173 → backend :8787).
**harness/** (Python `tht` CLI + Pi gate extension)
- Install: `cd harness && python -m venv .venv && pip install -e ".[dev]"` (puts `tht` on PATH)
- Test: `.venv/bin/pytest -q` — `l2` (real GLM + remote DB) is opt-in via `addopts = -m 'not l2'`; `l0` (testcontainers) needs Docker
- Single test: `.venv/bin/pytest tests/test_session_mutations.py::test_set_name -v` (or `-k <pattern>`); include e2e with `-m l2`
- Lint: `.venv/bin/ruff check .` (line-length 100)
**backend/** (Fastify + TypeScript, vitest)
- Dev: `npm run dev` (tsx watch `src/server.ts`) · Build: `npm run build` (tsc → `dist/`)
- Test: `npx vitest run` · Single: `npx vitest run test/routes-sessions.test.ts -t "rename"`
- Typecheck: `npx tsc --noEmit -p .` (vitest does NOT type-check — run this before committing)
**frontend/** (React 18 + Vite + vitest)
- Dev: `npm run dev` (Vite; set `VITE_BACKEND_URL`) · Build: `npm run build`
- Test: `npx vitest run` · Single: `npx vitest run src/shell/NavSessions.test.tsx`
- Typecheck: `npx tsc -b` · E2E: `npm run e2e` (Playwright)
No ESLint on the TS layers — `tsc` is the gate. Tests use vitest + MSW (no network).
## Architecture (the parts that need multiple files to see)
```
frontend (React/SSE) → backend (Fastify) → pi --mode rpc → tht/harness → DWH (read-only)
```
- **The harness owns the workflow and all persistence.** `tht` (Python) is a deterministic
CLI; `harness/.pi/extensions/tht-gate.js` is a Pi extension that drives an **8-phase
NL→SQL workflow**. The single source of workflow truth is `harness/workflow.yaml`; the
orchestration rules the model must follow are `harness/.pi/skills/tht-sessione/SKILL.md`.
"Current phase" is computed by folding the decision ledger (`harness/tht/phase.py`), not
stored — read it before reasoning about phase logic.
- **Persistence = phase documents, NOT chat.** A session is a directory under the workspace's
`sessions/` path: `session_manifest.yaml` + per-phase artifacts (`question.md`,
`schema_linking.json`, `sql_final.sql`, …) + `review_decisions.jsonl`. The contract
(SKILL.md): *"the persisted state is the truth — what is not recorded did not happen."*
There is no verbatim transcript store. A resumed Pi process rebuilds context from
`tht session show <id>` + the on-disk artifacts.
- **The backend is a thin bridge with no database of its own.** `ThtRunner` shells `tht`
subcommands; `PiProcessManager` runs one Pi child per session and bridges its RPC stream;
`SessionBridge` maps Pi RPC events → client events (`ui_request`/`text_delta`/`info`);
`SseHub` fans them out over SSE to the browser. Persistence belongs to the HARNESS, which
selects the session repository from the workspace config (`harness/tht/session/repository.py`):
filesystem by default, **PostgreSQL when `session_storage` is configured** (server/portable
deployment). Settings flow through harness preferences (`tht session preferences`) with
`backend/data/settings.json` only as the file fallback for injected runners/tests.
- **Human-in-the-loop gate contract.** The model proposes; a human reviewer decides at gates
via widgets (`reviewer_select` = single pick — a chosen option carrying a `decision` payload
auto-confirms/persists directly, an option without one only asks; `reviewer_decide` = multiselect,
each choice IS a decision; `reviewer_confirm` = artifact/phase gate). The frontend renders these
widget-descriptors (`src/widgets/` registry) and the live transcript is rebuilt in-memory
from the SSE stream (`src/store/sessionStore.ts`) — it is not persisted.
## Project-specific gotchas
- **`tht`'s `-c`/`--config` is a PER-COMMAND option** — it must follow the subcommand, never
precede it (`ThtRunner.buildArgv` enforces this; prepending caused live 500s).
- **`--json` output must be pristine** (only valid JSON on stdout) — used as a machine contract.
- **UI strings are English; document *content* stays the workspace language** (Italian for
`psd`) because it's the real data. Only chrome/labels are English.
- **Workspaces** (`harness/workspaces/*.yaml`) set the DB target and **absolute**
`paths.sessions/artifacts/indexes` — for `psd` these point at a *separate, uncommitted* repo
(`tht-workspace-psd/`). Secrets live ONLY in `harness/.env` (gitignored).
- **Settings are global** (workspace/provider/model/thinking, persisted via harness
preferences — `backend/data/settings.json` is only the fallback); the New-session form is
question-only.
- **Resume**: a resumable session re-enters at its last incomplete phase. The backend refuses
resume with 409 when `finalized` or `archived`, and `PiProcessManager.spawnFor` must send
`/riprendi-sessione <id>` (resume mode) vs `/nuova-domanda` (new) — sending the wrong prompt
silently turns a resume into a new question.
+643
View File
@@ -0,0 +1,643 @@
# Contesto di dominio di ThothII
## Architettura del workflow
**Workflow Kernel** — Il coordinatore deterministico che possiede lo stato del workflow,
le transizioni, il rollback, la finalizzazione e l'applicazione atomica degli esiti dei
moduli.
**Workflow Module** — Una capacità incapsulata che espone un contratto versionato. Un
modulo può partecipare a più stage e non modifica direttamente lo stato del workflow.
**Stage** — Un punto del workflow, identificato semanticamente, nel quale viene invocato
un modulo. L'identità dello stage è indipendente dalla sua posizione visiva.
**Display code** — L'etichetta di presentazione associata a uno stage, per esempio da
`F1` a `F8`. I display code alimentano gli indicatori di avanzamento nel frontend, ma non
sono usati come identità del workflow o chiavi di dipendenza.
**Module outcome** — Il risultato proposto da un modulo: eventi tipizzati, modifiche agli
artifact, un'eventuale richiesta di revisione umana e uno stato di esecuzione. Il
Workflow Kernel valida e applica l'esito.
**Revision request** — La proposta tipizzata con cui un modulo segnala che lo stage
corrente non può concludersi validamente senza rieseguire lo stesso stage o uno stage
precedente. Non produce direttamente una transizione: il Workflow Kernel valida la
richiesta, sospende l'avanzamento e, per riaprire uno stage già completato, attende una
decisione umana tipizzata. Il Kernel, non il modulo, determina gli eventi e gli artifact
causalmente da rendere stale.
**Question Admission** — Il controllo preliminare eseguito prima delle fasi da `F1` a
`F8`. Nella prima release distingue una domanda utilizzabile da input garbage e verifica
che la domanda appartenga allo scope dichiarato dal workspace. Il suo stato è mostrato
separatamente dagli otto indicatori di fase.
**Workspace scope** — La dichiarazione gestita e versionata di ciò che il database di un
workspace rappresenta e delle domande alle quali è destinato a rispondere. Question
Admission la usa come riferimento per valutare la pertinenza di una domanda.
**Datamart Plugin** — Il modulo sostituibile che implementa lo stage semantico
`datamart`, presentato con display code `F8`. La promozione della memory e la
finalizzazione della sessione non appartengono al Datamart Plugin.
**Ordered workflow** — La pipeline deterministica composta dal preflight Admission,
dagli otto stage principali ordinati da `F1` a `F8` e dalla finalizzazione. L'ordine
degli stage è esplicito; il workflow non è un DAG generale.
**Extension point** — Una posizione semantica nel lifecycle dell'Ordered workflow alla
quale possono contribuire uno o più moduli senza diventare nuovi stage visibili. Un
extension point non possiede un display code.
**Stage state** — La proiezione deterministica degli eventi del workflow che descrive
uno stage come `pending`, `ready`, `running`, `awaiting_human`, `completed`, `skipped` o
`failed`. Non è un valore corrente memorizzato separatamente dal ledger.
**Required contribution** — Il contributo di un modulo a un extension point che deve
concludersi o essere esplicitamente saltato secondo policy prima che il workflow possa
avanzare.
**Best-effort contribution** — Il contributo di un modulo il cui fallimento viene
registrato e mostrato come warning, ma non impedisce al workflow di avanzare.
**Blocked workflow** — La proiezione complessiva di un workflow che non può avanzare a
causa di uno stage o di un contributo required fallito o non disponibile. `Blocked` non
è uno Stage state autonomo.
**Module invocation** — Una singola richiesta del Workflow Kernel a un modulo in uno
stage o extension point. Conserva la stessa identità attraverso eventuali retry, che
sono tentativi distinti della medesima invocation.
**Stage skip** — La conclusione esplicita di uno stage senza eseguirne il comportamento.
È ammessa soltanto dalla policy dello stage e registra motivo e attore; un fallimento non
equivale mai implicitamente a uno skip.
**Stage reopen** — La riapertura di uno stage non finalizzato che rende stale gli esiti
causalmente successivi. Gli effetti esterni già prodotti richiedono una marcatura o una
compensazione esplicita e non sono presentati come automaticamente annullati. Può essere
applicata dal Workflow Kernel in seguito all'approvazione di una Revision request, ma
non può essere eseguita direttamente da Pi o da un Workflow Module.
**Completion policy** — La regola con cui uno stage si conclude: `automatic` quando il
kernel può verificarne deterministicamente l'esito, oppure `review_required` quando è
necessaria un'approvazione umana tipizzata.
**Paused session** — Una sessione interrotta intenzionalmente ma resumibile. L'azione
“Stop and save” mette la sessione in pausa; non la completa e non la marca come fallita.
**Finalized session** — Una sessione completata con esito canonico e immutabile. Una
correzione successiva crea una nuova sessione derivata, collegata a quella precedente.
**After-finalize hook** — Una notifica o attività best-effort eseguita tramite outbox
dopo la finalizzazione. Non può modificare il ledger, gli artifact canonici o lo stato
terminale della sessione.
## Memory
**Memory Module** — Il modulo che possiede le conoscenze ed esperienze curate per
migliorare schema linking e generazione SQL di domande future. Le Memory appartengono
a un workspace e rimangono distinte dalle Evidence.
**Memory Card** — L'unità di contenuto gestibile del Memory Module, con identità,
ambito di applicazione e provenienza. Il formato è allineato per analogia alle
Evidence, senza implicare la stessa origine o lo stesso percorso di pubblicazione.
**Reusable Memory** — Una Memory Card che esprime un chiarimento di dominio, una
regola di costruzione SQL o un errore da evitare con motivo compreso e approvato.
La sua validità è circoscritta a un ambito esplicito e non deriva dalla sola
approvazione di una scelta occasionale in una domanda.
**Solved Question** — Una Memory Card che conserva una domanda risolta con la
relativa soluzione SQL e il contesto necessario a interpretarla. È un exemplar
consultativo: i parametri e le scelte del caso non diventano regole generali.
**Memory Graph** — L'insieme dei collegamenti espliciti fra card che contribuisce
al recupero di conoscenze pertinenti oltre alla somiglianza del contenuto. Il
ritrovamento di una card tramite un collegamento non ne implica l'approvazione.
**Memory Link** — Un collegamento curato fra card, con destinazione e significato
espliciti, che contribuisce alla consultazione di contenuti pertinenti. La sua
rimozione non comporta la cancellazione delle card collegate.
## Evidence
**Context specialist** — La persona competente sul dominio che redige e cura il
contenuto delle Evidence. Può essere distinta da chi amministra l'installazione;
il suo lavoro di redazione non richiede accesso al database applicativo.
**Evidence draft** — Il documento iniziale scritto dallo specialista di contesto,
che il sistema acquisisce e raffina in Evidence Unit. Può essere redatto e
consegnato indipendentemente dall'installazione che userà le Evidence risultanti.
**Evidence Module** — Il modulo autonomo che possiede la preparazione delle Evidence e
la loro consultazione durante il workflow. La preparazione avviene fuori dalle singole
sessioni; il workflow usa soltanto contenuti già pubblicati. A runtime contribuisce agli
stage semantici esistenti, senza diventare uno stage visibile e senza modificare ledger,
artifact o stato del workflow.
**Source Evidence** — Il documento o la dichiarazione che sostiene il contenuto
corrente di una Evidence Unit. Un documento acquisito viene conservato come
riferimento umano; una dichiarazione manuale attribuisce il contenuto alla persona
che lo ha scritto e approvato.
**Manual Evidence declaration** — Una dichiarazione esplicita dell'amministratore
che sostiene una Evidence creata direttamente o una correzione del suo significato.
Non implica una verifica indipendente da parte di una fonte documentale esterna.
**Evidence origin** — Il documento da cui una Evidence Unit è stata inizialmente
derivata. Può restare collegato per provenienza e confronto con gli aggiornamenti
anche quando una dichiarazione manuale sostiene il testo corrente. La sola origine
non dimostra il supporto semantico di una successiva correzione.
**Local Evidence archive** — L'insieme delle Evidence curate custodite
dall'installazione, distinto dalle draft originali e dai contenuti derivati per
la ricerca. Comprende le correzioni manuali e i ritiri deliberati.
**Consolidated Evidence** — Una versione delle Evidence locali controllata come
insieme coerente e pronta per l'attivazione. I file ancora in modifica non ne
cambiano il contenuto.
**Active Evidence** — La versione consolidata disponibile alla consultazione del
core. Un tentativo di aggiornamento fallito conserva la versione attiva precedente.
**Evidence Unit** — La più piccola unità semantica coerente, revisionabile e ricercabile
fondata su una Source Evidence corrente, anche manuale, e con eventuale origine
documentale distinta. Possiede un identificatore stabile indipendente dal kind,
assegnato una volta nella forma `evidence:<slug>`; fonti diverse non vengono fuse
automaticamente.
**Evidence kind** — La categoria semantica di una Evidence Unit, che ne determina i
campi specifici e ne orienta l'uso. Ogni unità ha un solo kind primario; i tipi iniziali
sono `glossary`, `domain`, `enum`, `example`, `mapping`, `normalization`, `formula` e
`reference`.
**Glossary Evidence** — Una Evidence Unit che definisce il significato linguistico, i
sinonimi o le varianti di un termine.
**Domain Evidence** — Una Evidence Unit che esprime una regola o un vincolo del dominio
non rappresentato da un kind più specifico.
**Enum Evidence** — Una Evidence Unit che collega un insieme finito di valori
memorizzati ai relativi significati.
**Example Evidence** — Una Evidence Unit che associa un input o una domanda alla sua
interpretazione o al risultato atteso.
**Mapping Evidence** — Una Evidence Unit che collega un concetto logico agli elementi
del relativo schema fisico.
**Normalization Evidence** — Una Evidence Unit che descrive la trasformazione di una
rappresentazione in una forma canonica.
**Formula Evidence** — Una Evidence Unit che contiene una singola espressione PostgreSQL
componibile e ne dichiara gli input. Una query SQL completa non è una Formula Evidence.
**Reference Evidence** — Una Evidence Unit che rappresenta un collegamento esterno da
restituire come contenuto autonomo, anziché come semplice provenienza.
**Evidence purpose** — La destinazione dichiarata di una Evidence Unit nel workflow:
disambiguation, rewriting, schema linking o SQL generation. È distinta dall'Evidence
kind: il tipo descrive cosa contiene, il purpose quando può essere utile; durante la
ricerca il purpose richiesto è un filtro obbligatorio. Il recupero di esperienze e
soluzioni precedenti appartiene al Memory Module e non è un Evidence purpose.
**Evidence Search Outcome** — Il risultato tipizzato di una consultazione del modulo
Evidence. Distingue una ricerca disponibile, che può legittimamente non trovare
corrispondenze, da un'indisponibilità tecnica che impedisce allo stage chiamante di
avanzare fino a un retry riuscito.
**Evidence receipt** — La traccia minima di una consultazione disponibile conservata
nella sessione: stage semantico, purpose, generazione interrogata e identificatori delle
Evidence restituite. Non duplica il contenuto delle Evidence.
**Curated Evidence** — Una o più Evidence Unit preparate da documenti o curate
manualmente. La presenza nell'archivio curato non implica da sola che il contenuto
sia già attivo per il workflow.
**Published Evidence** — Le Curated Evidence valide appartenenti alla revisione attiva
del workspace e alla generazione Evidence pubblicata. L'approvazione umana precede
l'attivazione, ma non viene duplicata come stato nel manifest.
**Evidence Index** — La proiezione ricercabile e ricostruibile delle Published Evidence.
Accelera il recupero delle informazioni, ma non è una fonte di verità.
**Evidence preparation** — Il processo di authoring che trasforma Source Evidence in
Curated Evidence mediante estrazione e normalizzazione deterministiche, una singola
ristrutturazione assistita dal modello e una validazione finale deterministica. Nella
prima versione accetta Markdown o testo UTF-8 e non acquisisce automaticamente il
contenuto di URL o documenti esterni. Prepara l'intero insieme delle modifiche in
un'area temporanea e lo applica atomicamente soltanto se tutti gli output sono validi;
non ritenta automaticamente una chiamata al modello fallita.
**Supporting excerpt** — Un breve estratto presente nel Source Evidence che sostiene
una Evidence Unit. Il sistema ne verifica deterministicamente la presenza dopo la
normalizzazione meccanica; il curatore resta responsabile di verificarne la sufficienza
semantica.
**Evidence resolution** — La decisione esplicita con cui un curatore risolve un
problema di una Evidence Unit, correggendola, ritirandola oppure ricollegandola a
una fonte adeguata.
**Source update conflict** — Un contrasto fra una fonte aggiornata e una correzione
manuale già approvata. La correzione resta in uso fino alla risoluzione esplicita
del confronto da parte dell'amministratore.
**Evidence source refresh** — La riacquisizione delle fonti esterne richiesta
dall'amministratore per rilevarne le modifiche. Fra due aggiornamenti il contenuto
già acquisito resta il riferimento per preparazione e consultazione.
**Review item** — Un blocco di revisione descritto da codice stabile, messaggio umano e
campo opzionale. Finché viene mantenuto nell'Evidence Unit, ne impedisce la
pubblicazione; la sua storia è conservata da Git, non da uno stato interno all'item.
**Retirement candidate** — Una Curated Evidence che il Source Evidence esistente non
sostiene più. Rimane visibile con un Review item e blocca la pubblicazione finché il
curatore non la elimina oppure la rende nuovamente coerente con il sorgente.
**Evidence evaluation set** — Un piccolo insieme versionato di domande rappresentative
e relativi risultati attesi. La baseline è accettabile quando ogni domanda recupera
almeno un risultato atteso nei primi dieci risultati della fusione RRF; il risultato
nei primi cinque è informativo. Comprende almeno un caso lessicale, uno semantico e uno
misto e conserva, a fini diagnostici, le posizioni dense, BM25 e fused.
**Candidate Evidence Generation** — Una generazione completa dell'Evidence Index che
può essere valutata ma non è ancora visibile alle sessioni. Diventa attiva soltanto se
supera l'Evidence evaluation set.
**Evidence manifest** — Il file versionato e gestito dal sistema che collega ogni
Source Evidence al suo hash e alle Evidence Unit derivate. Conserva gli identificatori
stabili, permette l'elaborazione incrementale e segnala le unità rimaste orfane senza
cancellarle automaticamente.
**Orphaned Evidence Unit** — Una Curated Evidence il cui Source Evidence non esiste più.
Rimane disponibile per la revisione, ma blocca la pubblicazione finché non viene
eliminata, ricollegata oppure ne viene ripristinato il sorgente.
**Evidence Fragment** — Una proiezione ricercabile di una sezione semanticamente
coerente di una Published Evidence. La divisione segue intestazioni e confini di
paragrafo; formule, coppie valore/significato, mapping, regole e URL non vengono mai
tagliati. Il testo completo reso per il frammento usa il solo limite esistente
`max_chunk_chars`, pari per default a 4.000 caratteri; un elemento atomico troppo grande
produce un Review item bloccante. Qdrant indicizza i frammenti, mentre l'Evidence Module
li raggruppa per Evidence Unit.
**Evidence Result** — La rappresentazione di una singola Evidence Unit restituita dalla
ricerca con metadati, migliori estratti, provenienza e riferimento al documento completo.
**Hybrid Evidence retrieval** — La ricerca che combina in Qdrant una graduatoria
semantica dense e una graduatoria lessicale BM25 sparse mediante Reciprocal Rank
Fusion. I metadati tipizzati restringono o orientano i risultati senza creare una
collezione separata per ogni Evidence kind.
**Evidence query text** — La rappresentazione deterministica condivisa dalla ricerca
dense e BM25: domanda originale, concetti, tabelle e colonne in ordine fisso. I campi
vuoti sono omessi; domanda e contesto ricevono soltanto normalizzazione Unicode NFC,
conversione degli a-capo e rimozione degli spazi esterni. Gli elementi contestuali sono
poi deduplicati e ordinati senza conversione delle maiuscole, mentre punteggiatura e
spazi interni della domanda non vengono riscritti.
**Reference Vector Collection** — La collezione Qdrant ricostruibile di un workspace che
contiene Schema, relazioni ed Evidence. Possiede il vettore dense predefinito e il vettore
sparse `bm25`; soltanto gli Evidence Fragment ricevono valori BM25. Il preprocessing può
sostituirla o eliminarla integralmente.
**Memory Vector Collection** — La collezione Qdrant persistente di un workspace che contiene
`memory` e `solved_question`. Non è un output del preprocessing e non viene eliminata dal
Preprocessing Clear.
**Preprocessing Clear** — L'operazione amministrativa che elimina Reference Vector Collection,
LSH, corpus e checkpoint derivati e rende il workspace non pronto. Conserva Memory Vector
Collection, sessioni, Catalog Metadata e database sorgente; non offre history o rollback.
**Formula proposal** — Una formula individuata durante una sessione e conservata come
artefatto della sessione. Non diventa Published Evidence finché non viene importata,
revisionata e approvata nel repository del workspace.
**Fail-closed Evidence retrieval** — Il comportamento per cui un indice assente,
incompatibile o non aggiornato produce nessuna Evidence e un avviso esplicito. Il
workflow può continuare, ma non usa mai silenziosamente contenuti di una revisione
precedente o di un altro workspace.
## Configurazione dei modelli
**Workspace Descriptor** — La dichiarazione versionata dell'identità del workspace e dello
scope delle sue Evidence. Non contiene identità o configurazione del Workspace Database,
Database Binding, fatti strutturali o metadati semantici: il Metadata Catalog associa il
workspace al relativo database.
**Installation Model Catalog** — L'insieme dichiarativo, proprio di un'installazione, dei
modelli disponibili, dei loro Model Usage e dei relativi default. È l'unica autorità per i
modelli di sessione, generazione dei metadati ed embedding e non appartiene a un workspace.
_Avoid_: Model Catalog, Metadata Generation Model Configuration
**Model Usage** — Lo scopo per cui un modello dell'Installation Model Catalog può essere
usato: `session`, `metadata_generation` oppure `embedding`. L'ammissibilità e il default
dipendono dall'uso, non dal workspace.
**Model Selection** — La scelta runtime, a livello di installazione, di un modello del
catalogo per uno specifico Model Usage. Riferisce l'identità canonica del modello senza
ridefinirne provider, endpoint o capacità.
**Model Runtime Projection** — La rappresentazione derivata e non autoritativa
dell'Installation Model Catalog richiesta da uno specifico runtime. Può essere rigenerata
integralmente dalla configurazione dell'installazione.
## Distribuzione del prodotto
**Customer-Hosted Installation** — Un'installazione eseguita interamente nel trust boundary
controllato dall'organizzazione cliente, inclusi eventuali tenant cloud privati. Credenziali,
domande, prompt, metadati e risultati non attraversano quel boundary.
_Avoid_: on-premise deployment, self-managed deployment
**Community Edition** — La distribuzione open source utilizzabile gratuitamente anche in
produzione e capace di eseguire il workflow fondamentale completo.
_Avoid_: free tier, trial edition
**Enterprise Edition** — La distribuzione con licenza commerciale che aggiunge governance
organizzativa, esercizio production-grade e industrializzazione alla Community Edition.
_Avoid_: paid tier, pro edition
## Catalogo dei metadati
**Workspace Database** — Il database che appartiene a un solo workspace e non può essere
condiviso con altri workspace; un workspace può averne al massimo uno. È considerato nella
coppia composta dal database PostgreSQL e da un solo schema: tutte le tabelle, le colonne e
le relazioni catalogate appartengono a quello schema. Il Metadata Catalog conserva
l'associazione, ma non crea né possiede l'identità del workspace.
**Database Binding** — La configurazione specifica di un'installazione che seleziona un
trasporto e fornisce i riferimenti necessari a raggiungere un Workspace Database. Non è una
seconda identità del database e non viene condivisa automaticamente fra installazioni.
**Thoth REST Connector** — Il trasporto REST tipizzato con cui ThothII interroga ed
introspeziona un Workspace Database attraverso il contratto RPC DWH supportato. Non è un
client configurabile per API REST arbitrarie.
**Orphaned Workspace Database** — Un Workspace Database il cui workspace non è più presente
nel catalogo autorevole. Rimane conservato per il recupero amministrativo, ma non può essere
usato dal workflow finché non viene riassegnato a un workspace esistente.
**Metadata Catalog** — L'autorità per l'associazione fra workspace e Workspace Database, la
relativa Database Binding, i fatti strutturali osservati e i metadati semantici curati. Ogni
uso downstream dei metadati del database deriva da questo catalogo.
**Database Profile** — L'insieme curato di scope, descrizioni e metadati semantici
associato a un Workspace Database.
**Physical Table** — Una tabella osservata nello schema esterno di un Workspace Database.
La sua identità e il suo nome appartengono al database esterno, non al Metadata Catalog.
**Catalog Table** — La rappresentazione persistita di una Physical Table nel Metadata Catalog.
La sua appartenenza e identità fisica derivano dall'introspezione: non può essere creata o
rinominata manualmente, ma può essere rimossa tramite Catalog Metadata Cleanup.
_Avoid_: SqlTable, managed table
**Physical Column** — Una colonna osservata in una Physical Table, inclusi nome, posizione,
tipo e appartenenza a chiavi dichiarate. La sua identità e i suoi fatti strutturali appartengono
al database esterno.
**Catalog Column** — La rappresentazione persistita di una Physical Column nel Metadata Catalog.
I fatti osservati sono governati dalla sincronizzazione; Description e Generated Description
sono metadati amministrativi modificabili e la rappresentazione può essere rimossa tramite
Catalog Metadata Cleanup.
_Avoid_: SqlColumn, managed column
**Physical Relationship** — Un vincolo foreign key dichiarato nel database esterno. La sua
identità comprende il vincolo e la sequenza ordinata delle coppie di colonne che lo compongono.
**Catalog Relationship** — La rappresentazione persistita di una Physical Relationship nel
Metadata Catalog. Non è creata o modificata manualmente, ma può essere rimossa tramite Catalog
Metadata Cleanup.
_Avoid_: denormalized FK, relationship string
**Logical Relationship** — Una relazione modificabile fra due Catalog Column che non corrisponde
necessariamente a un vincolo fisico. Può essere Generated o Manual e rimane distinta dalla Catalog
Relationship osservata nel database.
**Generated Relationship** — Una Logical Relationship ricavata dai nomi delle colonne, dalle
primary key e dalla compatibilità dei tipi mediante regole deterministiche, senza LLM, embedding o
campionamento dei dati. Una ricostruzione non riattiva una Generated Relationship cancellata
logicamente, ma può ricrearne una cancellata fisicamente.
**Manual Relationship** — Una Logical Relationship aggiunta dall'utente. La ricostruzione delle
Generated Relationship non la modifica.
**Logical Relationship Deletion** — L'esclusione persistente di una Logical Relationship che ne
conserva l'identità per impedirne la ricreazione automatica finché esistono entrambe le Catalog
Column alle quali è collegata.
**Permanent Relationship Deletion** — La rimozione completa di una Logical Relationship. Una
ricostruzione successiva può ricrearla quando soddisfa nuovamente le regole di inferenza. Anche il
cleanup distruttivo di una tabella o colonna endpoint rimuove permanentemente le relative esclusioni.
**Relationship Reconstruction** — L'operazione amministrativa esplicita che scopre e aggiunge le
Generated Relationship mancanti. Conserva le Manual Relationship e le relationship già presenti e
non riattiva quelle cancellate logicamente.
**Relationship Restore** — La riattivazione esplicita di una Logical Relationship cancellata
logicamente.
**Effective Relationship Map** — La vista unificata delle Catalog Relationship fisiche e delle
Logical Relationship, con origine e stato espliciti. È l'interfaccia usata dall'amministrazione e
dalla comprensione dello schema, non un ulteriore modello persistito.
**Catalog Metadata Snapshot** — La proiezione immutabile e versionata della struttura catalogata,
delle descrizioni pubblicabili e delle relazioni effettive attive di un Workspace Database che il
core consuma. È derivata esclusivamente dal Metadata Catalog e non è un archivio autoritativo.
**Schema Index** — La proiezione vettoriale ricostruibile dei metadati del Workspace Database nel
Metadata Catalog. Il preprocessing la sostituisce integralmente e non è una fonte di verità.
**Description** — Il testo curato e consolidato che descrive una Catalog Table o Catalog Column
per gli usi downstream. Quando presente, prevale sulla relativa Generated Description.
**Generated Description** — Il testo modificabile prodotto dall'AI per una Catalog Table o Catalog
Column. È pubblicabile per gli usi downstream quando manca una Description, anche senza essere
prima consolidato, e rimane distinto dal commento osservato nel database.
_Avoid_: generated comment, source comment
**Description Consolidation** — L'azione amministrativa esplicita che copia la Generated
Description di Catalog Table o Catalog Column selezionate nella relativa Description. Opera sulla
selezione corrente, conserva la Generated Description e non modifica il commento osservato o il
database esterno.
**Table Synchronization** — La riconciliazione esplicita che rende le Catalog Table di un
Workspace Database uguali alle Physical Table osservate: crea quelle nuove, aggiorna i metadati
di origine ed elimina definitivamente quelle assenti. Non modifica mai il database esterno.
_Avoid_: table import
**Schema Synchronization** — La riconciliazione esplicita e autorevole di tabelle, colonne e
Catalog Relationship di un Workspace Database. Può operare su uno scope specifico oppure su
un unico snapshot completo tramite Synchronize All.
**Catalog Sync Run** — L'esecuzione durevole in background di una Schema Synchronization, con
scope, stato, avanzamento e log propri. Al massimo un run per Workspace Database può essere attivo.
**Description Generation Run** — L'esecuzione asincrona e sequenziale che usa il modello scelto
per produrre Generated Description di Catalog Table o Catalog Column. Al massimo una run è attiva
nell'intera installazione e ogni risultato valido viene salvato appena disponibile. Dopo
un'interruzione il recupero è manuale tramite una nuova generazione dei soli elementi mancanti.
**Description Generation Event** — Una riga testuale ordinata che registra avanzamento, risultato
o errore di una Description Generation Run e alimenta il log visibile all'amministratore.
**Non-generatable Description** — L'esito valido con cui il modello dichiara di non disporre di
informazioni sufficienti per descrivere il target. Produce una Generated Description standard
nella lingua del workspace e non rappresenta un timeout, un errore del provider o una risposta
non valida.
**Description Generation Unlock** — Il recupero amministrativo che marca come interrotta una
Description Generation Run registrata come attiva quando il backend non ha alcun processo di
generazione vivo. Non è un meccanismo di lock distribuito.
**Catalog Metadata Cleanup** — La rimozione amministrativa esplicita di Catalog Table, Catalog
Column o Catalog Relationship selezionate. Non modifica il Workspace Database, la Database Binding
o i segreti, e può lasciare il Metadata Catalog intenzionalmente incompleto fino alla prossima
Schema Synchronization.
**Catalog Freshness** — La corrispondenza fra uno scope sincronizzato e la versione corrente
della Database Binding. Uno scope rimane consultabile ma è stale finché non viene sincronizzato
con la binding corrente.
**Metadata Content Revision** — La revisione monotona di tutto lo stato del Metadata Catalog che
può modificare il comportamento del core. Ogni mutazione rilevante produce una nuova revisione
nella stessa transazione che la rende durevole.
**Preprocessing State** — Lo stato corrente `running`, `succeeded` o `failed` del preprocessing di
un workspace, insieme all'identità dei suoi input. Il core può usare il workspace soltanto quando
lo stato è `succeeded` e gli input coincidono ancora.
**Catalog Metadata** — I campi mutabili che descrivono database, tabelle, colonne e relazioni,
distinti dai fatti strutturali governati dalla sincronizzazione. Possono essere popolati dall'AI,
o da una modifica amministrativa senza cambiare il database esterno.
**Model Completion Helper** — Il processo Python interno ed effimero che esegue una singola
richiesta LiteLLM per conto del backend. Non è un servizio HTTP, non possiede il lifecycle della
Description Generation Run e non è una CLI esposta agli utenti.
**Catalog Sample** — Un input transitorio composto da un massimo di cinque righe e da valori di
esempio bounded di una Catalog Table per la generazione delle descrizioni. Può contenere valori
reali oppure sintetici in base alla Source Value Disclosure Decision; non viene persistito e non
diventa Catalog Metadata.
**Sensitive Data Flag** — La classificazione binaria umana applicata a una Catalog Column. Può
essere impostata liberamente dall'amministratore anche in contrasto con una valutazione automatica.
**Sensitivity Reason** — La motivazione sanificata persistita insieme al Sensitive Data Flag
quando l'amministratore salva una Sensitivity Review Draft. È Catalog Metadata della colonna, non
history della run; viene rimossa quando il flag torna non-sensitive e può essere assente per una
classificazione manuale priva di valutazione locale.
_Avoid_: AI reasoning, source evidence
**Local Sensitivity Assessment** — La valutazione locale, non autoritativa e priva di LLM di una
Catalog Column, basata su metadati e contenuto sorgente, con esito `sensitive`, `non_sensitive`
oppure `unknown`.
_Avoid_: AI suggestion, automatic flag
**Local NER Detector** — Il componente NLP opzionale e CPU-only che esamina soltanto testo ancora
ambiguo e restituisce evidenze al Local Sensitivity Assessment. Non decide lo stato della colonna,
non usa un LLM generativo e non persiste valori sorgente.
_Avoid_: AI classifier, local LLM fallback
**Model Data Boundary** — La qualificazione amministrativa di un modello come `internal` oppure
`external` rispetto al confine entro cui i valori sorgente possono essere comunicati.
_Avoid_: local model, remote model
**Source Value Disclosure Decision** — L'unica decisione effettiva che stabilisce se un modello
riceve valori sorgente reali oppure sostituti sintetici, combinando Model Data Boundary e Sensitive
Data Flag.
_Avoid_: sample filter, export flag
**Sensitive Data Policy** — L'insieme versionato di regole locali generali e specifiche che produce
una Local Sensitivity Assessment. Un singolo riscontro blocca l'intera colonna e qualsiasi valore
testuale più lungo di 500 caratteri rende sensibile la colonna.
_Avoid_: PII filter, sample filter
**Sensitivity Analysis Run** — Il tentativo amministrativo esplicito e tracciato che valuta una
selezione di colonne mediante la Sensitive Data Policy. Conserva stato, copertura e conteggi
aggregati, ma non valori sorgente né esiti per colonna.
_Avoid_: Sensitive Data Suggestion Run, AI analysis
**Sensitivity Review Draft** — La proposta transitoria che associa alle colonne selezionate una
Local Sensitivity Assessment e le relative evidenze sanificate. Non modifica il Sensitive Data Flag
né la Sensitivity Reason finché l'amministratore non salva le proprie decisioni e viene scartata al
reload.
_Avoid_: automatic flag
**Sensitivity Analysis Event** — Una riga testuale ordinata e sanificata che registra l'avvio,
l'avanzamento per fase e batch, l'esito o l'errore di una Sensitivity Analysis Run senza conservare
contenuti sorgente, output grezzi del detector o proposte per colonna.
_Avoid_: Sensitive Data Suggestion Event
**Introspection Capability** — Una categoria di struttura fisica che una Database Binding
può osservare, come tabelle, colonne, relazioni, indici o enum. Una capability non disponibile
è distinta da una capability osservata che non ha restituito elementi.
## Amministrazione e integrazione
**Workspace Readiness** — La preparazione di uno specifico Workspace per l'uso nel
workflow, comprensiva della disponibilità degli artefatti derivati dai suoi metadati
Database e dalle sue Evidence. Il preprocessing appartiene a questa preparazione;
la configurazione e la sincronizzazione del catalogo restano responsabilità Database.
**Administration Surface** — Una superficie amministrativa autonoma per configurare o curare una
parte dell'installazione. Workspace, Evidence, Memory, Database e Pi sono superfici peer e non
dipendono dall'esistenza di una sessione attiva.
**Administration Page** — La rappresentazione a pagina intera di una Administration Surface, con
una gerarchia condivisa per identità, stato, azioni e contenuto. Un form amministrativo appartiene
alla pagina e non a una popup come contenitore principale.
_Avoid_: management popup, settings modal
**Administration Route** — L'identità navigabile di una Administration Surface nel browser. Deve
essere ripristinabile con refresh e cronologia e non contiene valori transitori o segreti dei form.
**Embedded Thoth Shell** — L'esperienza Thoth ospitata dentro il documento e il contesto visuale di
un portale host. Conserva la propria gerarchia funzionale, ma deve rispettare la geometria,
l'autenticazione e le regole responsive del portale host.
**Full Thoth Shell** — L'esperienza Thoth autonoma che possiede il proprio header e il proprio
layout di pagina. Non replica la navigazione amministrativa del portale host e non dipende dal suo
template visuale.
**Shell mode** — La scelta di installazione fra `embedded` e `full`. Determina chi possiede il
chrome globale, i comandi di identità e le integrazioni visuali, ma non cambia il workflow o la
persistenza delle sessioni.
**Fullscreen state** — Lo stato temporaneo in cui il documento applicativo occupa il fullscreen
del browser. È distinto da `Shell mode`: una Full Thoth Shell può essere aperta senza fullscreen;
il passaggio è attivato da un comando esplicito e può essere annullato con la stessa azione o con
il comando nativo del browser.
**Portal Shell Adapter** — Il confine sostituibile che traduce lo stato e i comandi del chrome di
un portale host nel modello semantico usato da Thoth. L'adapter non possiede autorizzazione,
sessioni di workflow o contenuti del modello.
**Host Shell State** — Il minimo stato visuale fornito dal portale host: locale UI, tema e stato
fullscreen. In una Embedded Thoth Shell è la fonte autorevole per queste preferenze;
non include identità, token o stato di autenticazione, che restano responsabilità dell'accesso.
**UI locale** — La lingua delle label, dei messaggi, dei tooltip, degli stati e delle istruzioni
non generate dal modello nell'interfaccia Thoth. È distinta dalla lingua dei contenuti di un
workspace.
**Interaction language** — La lingua in cui il modello presenta domande, spiegazioni e proposte
al revisore durante una sessione. Viene fissata alla creazione della sessione e rimane invariata
durante una ripresa, anche se la UI locale corrente cambia.
**Administrative Page Family** — L'insieme delle cinque Administration Page che condividono shell,
navigazione, tipografia e regole responsive, pur mantenendo contenuti e operazioni specifici:
Workspace, Evidence, Memory, Database e Pi.
## Installazione
**Manual standalone installation** — Una copia di ThothII predisposta per l'uso autonomo da una
persona che possiede il computer, con una Full Thoth Shell e servizi applicativi locali. La
procedura non implica che DWH o provider LLM siano locali o disponibili offline.
**Installation bootstrap** — L'insieme delle attività iniziali che rende disponibile una
installazione manuale: verifica dell'host, generazione della configurazione, predisposizione
delle credenziali protette e avvio dei servizi. Non è un installer dell'applicazione.
**Platform acceptance** — La verifica che una Manual standalone installation possa essere
predisposta e avviata su una specifica combinazione di sistema operativo, architettura e runtime,
distinta dalla verifica funzionale del collegamento a DWH e provider LLM.
+465
View File
@@ -0,0 +1,465 @@
---
name: ThothII
description: "A calm, precise clinical analytics workbench for traceable and reviewable SQL workflows."
colors:
instrument-red: "oklch(55.87% 0.1881 23.2)"
instrument-red-hover: "oklch(50.95% 0.1812 24.1)"
porcelain-background: "oklch(99.18% 0.0011 17.2)"
porcelain-card: "oklch(99.85% 0.0006 17.2)"
warm-surface: "oklch(97.09% 0.0011 17.2)"
sunken-surface: "oklch(94.08% 0.0011 17.2)"
warm-graphite: "oklch(26.78% 0.0097 355.6)"
muted-graphite: "oklch(51.33% 0.0088 345.6)"
quiet-border: "oklch(90.93% 0.0035 354.7)"
success-mint: "oklch(46% 0.095 160)"
navigation-active: "oklch(92.5% 0.052 23.2)"
navigation-active-hover: "oklch(89.5% 0.071 23.2)"
navigation-active-foreground: "oklch(36.5% 0.11 23.2)"
navigation-active-border: "oklch(60% 0.135 23.2)"
warning-amber: "oklch(48% 0.09 70)"
information-neutral: "oklch(51.33% 0.0088 345.6)"
typography:
display:
fontFamily: "Manrope Variable, Manrope, system-ui, sans-serif"
fontSize: "1.5rem"
fontWeight: 600
lineHeight: 1.03
letterSpacing: "-0.025em"
headline:
fontFamily: "Manrope Variable, Manrope, system-ui, sans-serif"
fontSize: "1.5rem"
fontWeight: 600
lineHeight: 1.15
letterSpacing: "-0.015em"
title:
fontFamily: "Manrope Variable, Manrope, system-ui, sans-serif"
fontSize: "1.25rem"
fontWeight: 600
lineHeight: 1.25
letterSpacing: "-0.01em"
body:
fontFamily: "Manrope Variable, Manrope, -apple-system, BlinkMacSystemFont, Segoe UI, system-ui, Arial, sans-serif"
fontSize: "1rem"
fontWeight: 400
lineHeight: 1.65
letterSpacing: "normal"
control:
fontFamily: "Manrope Variable, Manrope, -apple-system, BlinkMacSystemFont, Segoe UI, system-ui, Arial, sans-serif"
fontSize: "0.875rem"
fontWeight: 600
lineHeight: 1.25
letterSpacing: "0.005em"
label:
fontFamily: "Manrope Variable, Manrope, system-ui, sans-serif"
fontSize: "0.75rem"
fontWeight: 600
lineHeight: 1.25
letterSpacing: "normal"
rounded:
xs: "4px"
sm: "6px"
md: "8px"
lg: "12px"
xl: "16px"
full: "9999px"
spacing:
xs: "4px"
sm: "8px"
md: "16px"
lg: "24px"
xl: "32px"
components:
button-primary:
backgroundColor: "{colors.instrument-red}"
textColor: "{colors.porcelain-background}"
typography: "{typography.control}"
rounded: "{rounded.md}"
padding: "0 14px"
height: "32px"
button-primary-hover:
backgroundColor: "{colors.instrument-red-hover}"
textColor: "{colors.porcelain-background}"
typography: "{typography.control}"
rounded: "{rounded.md}"
padding: "0 14px"
height: "32px"
button-secondary:
backgroundColor: "{colors.porcelain-card}"
textColor: "{colors.warm-graphite}"
typography: "{typography.control}"
rounded: "{rounded.md}"
padding: "0 14px"
height: "32px"
input-default:
backgroundColor: "{colors.porcelain-background}"
textColor: "{colors.warm-graphite}"
typography: "{typography.body}"
rounded: "{rounded.md}"
padding: "0 12px"
height: "40px"
card-default:
backgroundColor: "{colors.porcelain-card}"
textColor: "{colors.warm-graphite}"
rounded: "{rounded.lg}"
padding: "16px"
badge-primary:
backgroundColor: "{colors.instrument-red}"
textColor: "{colors.porcelain-background}"
typography: "{typography.control}"
rounded: "{rounded.sm}"
padding: "2px 8px"
height: "20px"
---
# Design System: ThothII
## Visual review branch, September 2026
The revision on `codex/ui-visual-review` is approved for implementation and Docker visual review,
not yet for adoption on `main`. The previous look remains recoverable from the base commit and
the preserved Docker image. Historical prototypes must remain untouched.
This revision follows Impeccable's product register: one locally bundled Manrope family for the
whole UI, five fixed size roles, red as the sole brand accent and additional color only for meaningful
state. The primary scene remains an analyst reading data and SQL in a well-lit office.
## Overview
**Creative North Star: "The Clinical Workbench"**
ThothII should feel like a well-kept clinical workbench: warm enough for sustained reading, exact
enough for consequential review, and quiet enough that evidence, state, and decisions remain in the
foreground. The visual system is calm, precise, and trustworthy. It uses familiar product patterns,
restrained color, and deliberate density instead of decorative spectacle.
The primary physical scene is an analyst reviewing persisted evidence and SQL on a large monitor in
a well-lit working environment. This makes the warm light theme the default. The supported dark
theme serves lower-light work without becoming a separate neon aesthetic. Both themes preserve the
same hierarchy and semantic roles.
The system rejects generic SaaS ornament, conspicuous ripples, bounce or elastic motion, long
choreographed transitions, and effects that compete with the analytical task. Controls should feel
disciplined and tactile, never playful, sluggish, or visually unstable.
**Key Characteristics:**
- Warm, restrained surfaces with one scarce red accent.
- One sans-serif family, with hierarchy expressed through size, weight and spacing.
- Dense information organized through hierarchy, rhythm, and progressive disclosure.
- Persisted artifacts and reviewer decisions presented as the visual source of truth.
- Fast state feedback with reduced-motion parity.
**The Workbench Rule.** Every visual element must support inspection, action, state, or provenance.
Decoration without an operational purpose is forbidden.
**The Persisted Truth Rule.** Persisted artifacts and reviewer decisions receive stronger hierarchy
than transient model narration.
**The Density with Rhythm Rule.** Preserve information density, but vary spacing between groups so
users can scan structure without adding nested containers.
## Colors
The full-mode application header matches Omics Portal's `--gsd-red-primary`
(`#CB333B`) in both themes. Its complete wordmark, including `II`, and controls
use a near-white foreground. This header is absent in embedded mode. The sidebar
and welcome wordmarks retain their red suffix. Context editing places workspace,
model and Done in one desktop row, stacking on narrow containers. Session-scope
tabs retain their selected fill and accessible keyboard state with a uniform one-pixel
border on every side, gray when inactive and red when active. Their padding is 11px
horizontal and 3px vertical, with a 38px minimum height and wrapping labels.
The palette combines warm porcelain surfaces, warm graphite text, and an instrument red used only
for action, focus, and important state. OKLCH values in the frontmatter are normative because the
frontend uses OKLCH tokens directly.
### Primary
- **Instrument Red** (`instrument-red`): primary actions, focus identity, and destructive meaning
where the context already makes the action explicit.
- **Instrument Red Pressed** (`instrument-red-hover`): hover and active emphasis for the primary
action family.
### Neutral
- **Porcelain Background** (`porcelain-background`): the main canvas.
- **Porcelain Card** (`porcelain-card`): lifted panels, cards, and popovers.
- **Warm Surface** (`warm-surface`): sidebars, secondary controls, and muted regions.
- **Sunken Surface** (`sunken-surface`): selected rows, quiet emphasis, and inset regions.
- **Warm Graphite** (`warm-graphite`): primary text and high-confidence labels.
- **Muted Graphite** (`muted-graphite`): descriptions, timestamps, and secondary metadata.
- **Quiet Border** (`quiet-border`): structural boundaries, input outlines, and dividers.
### Semantic
- **Success Mint** (`success-mint`): completed and ready states.
- **Navigation Active** (`navigation-active`): the one application surface currently in the
foreground. It shares Instrument Red's hue but uses a lighter, lower-chroma fill, so location is
visible without carrying the full weight of a primary action.
- **Warning Amber** (`warning-amber`): waiting, attention, and in-progress states.
- **Information**: neutral text and indicators for dates, protocols and ordinary status. The legacy
`--info` token resolves to muted foreground, not an additional blue accent.
The dark theme keeps the same semantic mapping with neutral near-black surfaces and a slightly
lighter red accent. Do not introduce a second visual identity for dark mode.
**The One Voice Rule.** Instrument Red should occupy no more than roughly ten percent of a screen.
Its rarity is what makes it authoritative.
**The State Has a Name Rule.** Success, warning, information, and destructive colors are reserved
for their named states. Color is never the only state indicator.
## Typography
**UI Font:** locally bundled Manrope Variable, with Manrope and native sans-serif fallbacks.
**Technical Font:** SF Mono or Cascadia Code, with Menlo and Consolas fallbacks.
Manrope covers headings, labels, controls, navigation and document reading. Monospace is reserved
for SQL, code, paths and machine identifiers, never for ordinary UI labels or status headings.
### Hierarchy
- **Headline** (600, `1.5rem`, `1.3`): page or artifact titles, `--text-page`.
- **Title** (600, `1.25rem`, `1.4`): section hierarchy, `--text-section`.
- **Body** (400, `1rem`, `1.6`): operational prose, `--text-body`, with a target line length of 65 to 75
characters where the surface controls width.
- **Control** (400–600, `0.875rem`, `1.5`): buttons, inputs, tables, tabs and compact subheadings,
`--text-control`.
- **Metadata** (400–600, `0.75rem`, `1.5`): secondary status, counts and timestamps, `--text-meta`.
Labels use sentence case and normal tracking. Ordinary operational text never falls below 12px.
Typography uses fixed sizes. Responsive changes happen at structural breakpoints, not through fluid
type scaling. Numeric data and identifiers use tabular numerals where comparison matters.
**Application wordmark:** ThothII is a brand mark, not a page title: use Manrope semibold at
48px (`3rem`) in the Core welcome area and 32px (`2rem`) in the session sidebar, with the
`II` suffix in brand red. Preserve these sizes across responsive layouts.
**The One Family Rule.** The UI and document readers use sans-serif throughout. The legacy
`--font-heading` alias resolves to `--font-sans`. Preserve technical monospace without turning it
into a second decorative hierarchy. Do not shrink text to solve layout constraints.
**The Read Once Rule.** A heading, label, and body must be distinguishable on first glance through
size and weight. Do not repeat headings in explanatory copy.
## Elevation
The system is flat by default and layered when necessary. Borders mark structure. Warm, diffuse
shadows mark actual elevation for popovers, dialogs, and selected containers. Tonal layering should
solve most hierarchy before a shadow is introduced.
### Shadow Vocabulary
- **Contact Shadow** (`--shadow-xs`): a one-pixel contact shadow for controls and code blocks.
- **Panel Shadow** (`--shadow-sm`): a small two-stage shadow for cards that need separation from the
canvas.
- **Overlay Shadow** (`--shadow-md`): a broad, low-opacity shadow for dialogs and floating layers.
Focus uses an explicit three-pixel ring. Waiting-for-input state may use a success-tinted ring, but
must retain a textual or structural cue. Motion for button state changes lasts `140ms` with
`cubic-bezier(0.22, 1, 0.36, 1)`. Dialog transitions last `100ms`. Activity pulses may run at
`1.5s`, and must be disabled under `prefers-reduced-motion`.
**The Flat by Default Rule.** A resting surface has no shadow unless it is physically above another
surface. If every panel floats, none of them has hierarchy.
**The Borders Structure, Shadows Elevate Rule.** Never use shadow as a substitute for grouping or a
border as a decorative accent.
## Components
Components are familiar, compact, and state-complete. Every interactive primitive must define
default, hover, focus, active, disabled, loading, and error behavior where those states apply.
### Buttons
- **Shape:** gently curved rectangle (`8px`) with a one-pixel transparent or structural border.
- **Primary:** Instrument Red, porcelain text, `32px` default height, and `14px` horizontal padding.
- **Hover / Focus:** shift to Instrument Red Pressed; show a three-pixel focus ring at 25 percent
opacity. Active state scales to `0.97` for `140ms` and removes elevation.
- **Secondary / Outline:** porcelain card surface, Quiet Border, Warm Graphite text, and a Warm
Surface hover.
- **Ghost:** transparent at rest, Warm Surface on hover. Use only where surrounding structure makes
the hit target obvious.
### Badges and Status Indicators
- **Style:** compact (`20px` height), gently curved (`6px`), and semibold.
- **State:** pair semantic color with text, icon, or position. A colored dot alone is insufficient
when the state affects workflow decisions.
### Cards and Containers
- **Corner Style:** softly rounded (`12px`), with `16px` default internal padding.
- **Background:** Porcelain Card over Porcelain Background or Warm Surface.
- **Shadow Strategy:** Panel Shadow only when the card must read as elevated.
- **Border:** one-pixel Quiet Border at partial opacity.
- **Nesting:** nested cards are forbidden. Use headings, dividers, spacing, or tonal regions.
### Inputs and Fields
- **Style:** `40px` height, `8px` corners, Porcelain Background, Quiet Border, and Manrope body text.
- **Focus:** three-pixel Instrument Red ring with a clear border shift.
- **Error / Disabled:** errors combine destructive color with explanatory text; disabled controls
retain readable contrast and use 50 percent opacity.
- **Global context:** the collapsible top shelf is the sole workspace/model selector for Core and
Admin. Preserve independent remembered choices, installation defaults, operation locks and unsaved
edit guards. Never introduce a separate metadata-generation default or selector.
### Navigation
- **Workspace readiness:** the Workspace navigation button carries an 8px dot to
the right of its label. Green means a selected workspace with confirmed ready
preprocessing and no query error; all other states are red. The button's
tooltip and accessible description retain the translated exact state. Do not
add a separate readiness text row or change the backend readiness gate.
- **Session groups:** one accessible single-open accordion contains Active sessions
and Archive, both initially closed. Below the scope tabs, show only their
adjacent section headers, without a redundant Sessions heading. Selection and
bulk-delete controls belong inside each panel and only appear for nonempty
lists. Select all affects that list only, preserves the other list's selection,
and exposes a mixed state for partial selection. Preserve the existing archived
flag as the grouping rule, independent of whether a Pi process is running.
Opening a section closes the other; either can be collapsed, including both.
Empty lists show only the translated "No sessions yet." message.
The open section uses the rail's remaining height; its list scrolls internally
with a cap of `min(18rem, 35dvh)`, while its trigger remains outside that scroll
area. The mobile navigation dialog supplies a bounded viewport-height container.
Keyboard users can focus and scroll each labelled panel.
- **Session entry:** one Session button returns to the current unfinished session,
including provisional creation, without resetting or reconnecting it. Otherwise
it prepares a new question using the normal readiness and unsaved-work guards.
- **Style:** compact session rows use `8px` corners and restrained vertical padding.
- **Default / Hover / Active:** porcelain at rest, Sunken Surface on hover, and a muted Navigation
Active red with a defined border when current. Exactly one top-level navigation control is current.
- **Administrative controls:** the admin-only Administration accordion groups Database,
Memory, Evidence, a structural divider, Workspace, and Pi configuration in that order. Its trigger exposes
expanded state and starts collapsed by default, while non-admin users do not receive the accordion
or its navigation actions.
- **Responsive:** collapse navigation structurally at the application breakpoint. Do not shrink
labels into illegibility. Below 768px, Memory and Evidence management use the full content
width; a Navigation button opens the shared accessible dialog. Selecting another archive
page or pressing Escape closes it. Desktop retains the right session sidebar and its My sessions /
All sessions tabs. Core retains question/answer, eight phases, reviewer gates and the left log.
In embedded mode the portal owns the red header and left sidebar; ThothII must not duplicate them. Size to the
actual application container. Narrow session document panels may use the available width.
### Session review and confirmations
Session dialogs use the visible application area, including the portal's header
and side rail. Artifact and schema-column review can grow to 80rem wide and the
available height; short confirmations use up to 40rem and at least 18rem when
space permits. Keep a 24px outer margin on desktop and 8px on small or short
screens. Long review content scrolls internally; on very short screens the
whole dialog can also scroll so every action remains reachable.
Session forms and review gates repeat their existing primary confirmation above
and below the content, sharing selection, validation, pending state and response
handlers. Alternate-response inputs follow the same rule. Reserved navigation
controls remain below the review. Stop/delete initially focus Cancel; rename
initially focuses the name field. Administration dialogs and forms retain their
existing layout and actions.
### Tabs
- **Shape:** compact label tabs sit on a shared baseline with rounded top corners and a two-pixel
lower edge, except session-scope tabs which use a uniform one-pixel border, rounded
corners and a 4px gap without a shared border or negative bottom margin.
Inactive labels retain a Quiet Border and Porcelain Card surface, so every
label reads as a tab before interaction; hover feedback reinforces clickability.
- **Current:** the selected tab uses the muted Navigation Active red for its fill, text, and defined border.
It must expose `aria-selected`, participate in a labelled `tablist`/`tabpanel`, and be the only
tab in the roving keyboard tab order.
- **Keyboard:** Left/Right move between adjacent tabs with wrapping; Home/End select the first or
last tab.
### Tooltips
- **Row actions:** icon-action tooltips open three pixels below the trigger and align to its trailing
edge, so they never cover the icon row. They use a dark slate surface, porcelain text, and a
defined border rather than the light popover treatment.
- **Interaction:** tooltip layers never receive pointer events. They appear on hover and keyboard
focus with a short ease-out transition, while the icon button keeps its complete accessible name.
- **Scope:** this treatment is shared by database, table, column, and relationship row actions.
Toolbar and navigation hints may use separate collision-aware placement.
### Curated Evidence Documents
Memory and Evidence share the `thot-knowledge-reader` reading contract. Use locally
bundled Manrope with normal tracking for prose and labels, and these fixed roles:
- Card title: 24px, weight 600, line-height 1.3 (`thot-knowledge-title`).
- Field/section heading, including Scope and Provenance: 20px, weight 600,
line-height 1.4, 8px clearance below (`thot-knowledge-heading`).
- All narrative text, including scope, lists and provenance: 16px, weight 400,
line-height 1.65. Do not apply compact UI text sizes to these fields.
- Authored Markdown subheadings inside a field: 16px, weight 600, line-height 1.5,
24px above/8px below. They remain subordinate to the enclosing field heading;
their semantic heading levels and original content are preserved.
- Technical metadata labels/values: 14px/1.5, with weight 600 for labels.
Only code, paths and machine identifiers use the technical monospace family at
14px/1.65, identical for inline and fenced code (never compound `em` shrinkage).
Separate reading sections by 24px; keep the first Markdown block flush with its
field heading's 8px bottom gap. The same typography applies in light/dark and at
all responsive widths. Controls and archive indexes retain their compact UI roles.
Memory and Evidence detail readers use the entire available content width, without
the ordinary 72–75ch prose cap. This is the owner's explicit reading-layout choice.
Long unstructured paragraphs are split for display at existing sentence/semicolon
boundaries outside inline code and links; authored Markdown structure and stored
content are unchanged. Paragraph spacing is 1.25em. Scope and provenance share the
available width; provenance excerpts render Markdown rather than literal markers.
Copy actions use the two-overlapping-sheets icon, an accessible name/tooltip and
live success/failure feedback instead of a visible Copy label.
Memory has four explicitly FAKE formatting examples, one per family, in a separate
expandable section. They reuse the real detail reader but never enter persistence,
indexing, link search or model recall, and expose no edit/delete/save actions.
Curated evidence follows a fixed reading order: title, compact type and purpose summary, scope,
typed content, supporting excerpts, review items, then technical provenance. Curated v4 files
use short, visible YAML frontmatter for identity and classification. The Markdown title and
body are authoritative; hidden payload comments are a legacy format converted on consolidation.
`applies_to` is rendered as “Ambito di applicazione” with separate bullet lists for concepts,
tables, and columns. Enum values also use lists. Tables are forbidden for metadata, scope, or any
one-dimensional collection; reserve tables for genuinely two-dimensional datasets. Long machine
identifiers use inline code. SQL uses fenced code. Supporting excerpts use blockquotes.
**The Review Surface Rule.** The visible Markdown must be readable without understanding the
machine contract. In Administration, explain current and original provenance separately and
keep file-editing templates and Git instructions in progressive disclosure. Show actual host
paths with copy controls, never browser file links to container-only locations.
## Do's and Don'ts
### Do:
- **Do** make every state change unmistakable without interrupting flow.
- **Do** use Instrument Red only for primary action, current selection, focus identity, or explicit
destructive meaning.
- **Do** preserve information density with headings, rhythm, and progressive disclosure.
- **Do** keep keyboard focus explicit and pair color with text, shape, icon, or position.
- **Do** respect `prefers-reduced-motion` while preserving immediate non-kinetic feedback.
- **Do** use the selected interface language (English by default) for chrome and preserve the
workspace language for persisted domain content. Session interaction language remains pinned.
- **Do** render curated metadata and scope as Markdown prose or lists, never as a frontmatter table.
- **Do** break long curated rules into paragraphs, labelled subsections, and lists at existing
punctuation boundaries while preserving the exact canonical text for machines.
### Don't:
- **Don't** add generic SaaS ornament, conspicuous ripples, bounce or elastic motion, long
choreographed transitions, or effects that compete with the analytical task.
- **Don't** make controls feel playful, sluggish, or visually unstable.
- **Don't** use gradient text, decorative glassmorphism, or full-saturation accents on inactive
states.
- **Don't** use a colored side stripe greater than one pixel on cards, callouts, list items, or
blockquotes. Use a full border, tonal background, icon, or heading instead.
- **Don't** nest cards or wrap every section in a container.
- **Don't** use a modal before exhausting inline or progressive alternatives.
- **Don't** use tables for `applies_to`, metadata, enum values, or other one-dimensional content.
- **Don't** use color as the sole carrier of success, warning, error, selection, or progress.
- **Don't** use display typography for buttons, labels, or data.
- **Don't** add em dashes to interface copy. Use commas, colons, semicolons, or parentheses.
+101 -1295
View File
File diff suppressed because it is too large Load Diff
+116 -92
View File
@@ -1,69 +1,62 @@
# ThothII
ThothII is a human-reviewed NL-to-SQL workflow with a React frontend and a Fastify/Pi/`tht`
core. The portable deployment runs exactly two application services; data services remain
external in this profile, except for the mandatory internal semantic services bundled in Compose.
core. The portable deployment runs two application services plus the installation-local metadata
catalog; DWH and LLM services remain external. Semantic services are bundled in Compose.
Authentication is configured through the single host CLI tht: see the [local authentication guide](docs/install/authentication-local.md),
[generic OIDC guide](docs/install/authentication-oidc.md), and [manual acceptance matrix](docs/testing/authentication-manual-acceptance.md).
The same frontend supports **full** (its own header) and **embedded** (inside a
portal). This choice is independent of authentication: the Mac uses full/local,
Omics uses embedded/upstream with its existing login, and a standalone server
can use full/OIDC. See [rendering architecture](docs/architecture/application-shell.md)
and [configuration, Omics delivery and deploy](docs/operations/shell-and-localization.md).
## Docker Compose: local startup
For the current server upgrade with Omics Portal, follow the ordered
[Codex server handoff](docs/operations/server-codex-handoff.md), including source
integration, embedded/upstream configuration, coordinated rollout and rollback.
Requirements: Docker Engine with Compose v2. The mandatory stack is `frontend`, `core`, `qdrant`, `embedding`, and the one-shot `embedding-model-init`. DWH and LLM remain external,
configurable endpoints—even when they are co-located with ThothII.
Local/OIDC authentication is configured through the host CLI `tht`; portal
authentication is established by the trusted server proxy. See the
[local guide](docs/install/authentication-local.md),
[OIDC guide](docs/install/authentication-oidc.md),
[upstream integration](docs/install/authentication-upstream.md), and
[manual acceptance matrix](docs/testing/authentication-manual-acceptance.md).
From a fresh clone, run these commands from the repository root:
For the clone-based manual standalone installation test on macOS, Windows, and Linux, use the
[Italian procedure](docs/install/standalone-manual-it.md) or the
[English procedure](docs/install/standalone-manual-en.md).
```sh
cp deploy/env/local.env.example deploy/env/local.env
# Edit deploy/env/local.env, including PI_AUTH_FILE, THT_SECRETS_FILE, and external endpoints.
docker compose --env-file deploy/env/local.env \
-f compose.yaml -f deploy/compose.local.yaml up --build -d
```
The [public manual](https://git.tylconsulting.it/thothii-docs/) covers the product,
installation, use and administration. Developer architecture, contracts, ADRs, tests,
plans and release records remain in this repository but are excluded from MkDocs
pages and search. This is an editorial boundary, not an access restriction on the
public repository. See the [documentation cleanup review](docs/maintenance/2026-09-15-documentation-cleanup.md)
for the executed consolidation and the inventory of historical sources retained in Git.
`./scripts/run-stack.sh` runs this same base+local command in the foreground. The core image
contains its Pi runtime; no host `pi` executable is used. For a server installation:
## Docker Compose and installation
```sh
cp deploy/env/server.env.example deploy/env/server.env
# Edit all absolute storage, Pi/secret/session files, and endpoint paths.
sudo scripts/prepare-server-pi-state.sh /srv/thothii/pi-state 10001 10001
docker compose --env-file deploy/env/server.env \
-f compose.yaml -f deploy/compose.server.yaml \
-f deploy/compose.session-server.yaml.example up --build -d
```
For a fresh installation, follow the complete manual procedure in
[Italian](docs/install/standalone-manual-it.md) or
[English](docs/install/standalone-manual-en.md). Configure protected files first;
then run the documented build, explicit migrations and startup commands with the
same installation descriptor and Compose project. There is no installer or launcher.
The initializer is required for an empty or restored server Pi-state bind. It atomically creates
the three regular targets hidden below the writable parent bind; protected Pi auth and tracked
model/settings sources remain separate read-only mounts. See the server manual before substituting
a root other than `/srv/thothii/pi-state`.
The mandatory stack includes frontend, core, PostgreSQL catalog, Qdrant, Ollama
and the embedding initializer. DWH and LLM endpoints remain external dependencies.
Pi is included in the core image. Credentials and certificates belong in protected
installation-local files, never in the workspace repository.
Workspace descriptors come from the Git remote configured by `THT_WORKSPACE_GIT_REMOTE`; their
runtime endpoint and secret bindings remain installation-local. Open
<http://127.0.0.1:8080> (set `THOTH_HTTP_PORT` in `deploy/env/local.env` to choose another
loopback port).
Credentials and certificates are local protected files. Do not put them in environment examples,
workspace YAML, URLs, or Compose interpolation values.
Application state is split across the named `settings`, `pi-state`, `workspace-registry`,
`sessions`, `qdrant-data`, and `embedding-models` volumes. `docker compose down` keeps them.
`qdrant-data` is a derived but persistent index store; `embedding-models` is an Ollama model
cache for `qwen3-embedding:0.6b` with fixed `1024`-dimension embeddings. Only an explicit destructive command such as `docker compose
down --volumes` removes them.
The frontend depends on the core health check and proxies `/health` and `/api/*` to it. The
application health endpoint intentionally checks process readiness only; external dependency
diagnostics are exposed by `tht doctor` and do not prevent the UI from starting.
For developer topology, overlays and lifecycle details, see the internal
[Compose reference](docs/operations/compose-reference.md). Ordinary stop/down keeps
persistent data; removing volumes is destructive and is not an upgrade step.
Process health is distinct from external dependency checks performed by doctor.
## Git-backed workspace repository
Workspace descriptors are shared through a validated Git repository while endpoint bindings and
secret files remain installation-local. Use the [local Mac/PC installation manual](docs/install/local-workspace-registry.md)
for Docker Desktop or a local engine, the [server installation manual](docs/install/server-workspace-registry.md)
for the Gitea, reverse-proxy, backup, upgrade, and recovery workflow, and the
[P1→P1.1 migration guide](docs/migrations/p1-to-p1-1-registry-layout.md) before upgrading an
older flat-layout registry.
secret files remain installation-local. The supported operating sequence is documented in
[Workspace operations](docs/operations/workspaces.md); it covers curator publication, installation
activation, runtime bindings, and preprocessing. The host setup and lifecycle path is in
[Install and first start](docs/install/first-start.md).
The curator-owned repository layout is:
@@ -92,19 +85,26 @@ a remote user's partial list. The isolated deployment exercise is
`./scripts/verify-workspace-install-docs.sh --profile local` or `--profile server`.
<!-- workspace-descriptor-contract:start -->
Schema v3 is the only accepted workspace descriptor. Schema v1 and v2 workspace descriptors are
rejected before activation. Candidate snapshot validation therefore makes activation or a pull fail
atomically while the prior valid snapshot remains active. There is no in-product migrator or
automatic conversion. A repository must already contain reviewed v3 descriptors. One workspace
owns one Qdrant collection;
schema, Evidence, and Memory records share that collection and stay separated by indexed payload
`kind`.
Schema v4 is the only accepted workspace descriptor. It contains workspace identity and optional
Evidence configuration only; PostgreSQL Metadata Catalog owns every database fact and binding.
Schema v1, v2, and v3 descriptors are rejected before activation. Candidate snapshot validation
therefore makes activation or a pull fail atomically while the prior valid snapshot remains active.
Each workspace owns separate Qdrant `reference` and `memory` collections: Schema, relationships, and
Evidence are replaceable reference data; Memory and solved questions have a persistent lifecycle.
<!-- workspace-descriptor-contract:end -->
Connector `ssh_tunnel` bindings are diagnostic-only in this release: their bounded probe always
cleans up the loopback forward and returns `workspace_not_activatable`; session creation is rejected
before persistence. Git registry access over SSH is unaffected. Use direct or REST connector
transport for runtime sessions.
<!-- non-workspace-migration:start -->
Create a clean v4 descriptor containing only `workspace` and optional `evidence`. Do not copy the
legacy database, diagnostics, `llm_policy`, or `semantic_index` blocks; configure the database in
Database Management.
<!-- non-workspace-migration:end -->
For NL→SQL runtime sessions, connector `ssh_tunnel` bindings remain diagnostic-only: their bounded
probe cleans up the loopback forward and returns `workspace_not_activatable`; session creation is
rejected before persistence. Database management is a separate boundary and supports a strict
OpenSSH tunnel for **Test connection** and **Sync tables**, using a private key, optional passphrase,
mandatory `known_hosts`, and optional PostgreSQL TLS CA/server name. Git registry access over SSH is
unaffected. Use direct or REST connector transport for runtime sessions.
`docker-compose.dev.yml` is deliberately local: both published ports bind to `127.0.0.1`,
`THT_SESSION_STORAGE=local`, and `THT_HOME=/data/local-home`. Do not set
@@ -160,10 +160,10 @@ secret files, upstream-auth checks, and a fail-closed `503` assertion for its de
unavailable disposable session endpoint. No real provider, database credential, or repository
secret is required.
For a clean server bind, `scripts/prepare-server-pi-state.sh` creates the hidden regular
`agent/auth.json`, `agent/models.json`, and `agent/settings.json` mount targets atomically before
Compose. The server smoke starts from an empty Pi-state root and applies this same preflight; the
real protected/tracked sources remain separate read-only mounts. Deterministic fixture tests render
For a clean server bind, `scripts/prepare-server-pi-state.sh` creates the hidden regular Pi agent
mount targets atomically before Compose. The auth target receives the protected credential bind;
the model and settings targets receive generated read-only projections. The server smoke starts
from an empty Pi-state root and applies this same preflight. Deterministic fixture tests render
both profiles, verify that bindings stay on `core`, check mount readability, and run the production
workspace resolver. Wrong-service, wrong-value, and broken-secret-mount mutations must fail.
@@ -173,7 +173,7 @@ an independent 32-minute outer timeout and does not retry a failed command.
Current release status (2026-08-05): clean-root render/setup and the production runtime-binding
resolver contracts are green. The server fixture supplies all four private trusted claims,
including exact non-admin value `0`, and a focused test proves nginx normalization produces the
accepted non-admin backend principal. Canonical schema-v3 registry descriptors now pass through
accepted non-admin backend principal. Canonical schema-v4 registry descriptors now pass through
one backend-owned, secret-safe runtime handoff for inventory and session execution; canonical
identity and durable session/artifact/index roots are retained. The fresh update-only smoke passed
bad-candidate mutation, automatic `rolled_back` compensation, exact prior-image restoration,
@@ -202,20 +202,26 @@ Startup mode adds bounded image build/two-service health startup, installation-a
status, stopped-container-aware ownership checks, and exact cleanup. The ordinary hosted Windows
job remains deterministic and does not claim Docker startup.
## Preprocessing jobs and S3 Evidence
## Workspace preprocessing and S3 Evidence
The included preprocessing services reuse the internal Qdrant/Ollama stack. Mount Evidence at
`/data/source/evidence`, then run the explicit preprocessing preset:
For an interactive run, select the workspace, expand **Administration** in the right sidebar, and
use its **Preprocessing** control. The control explains any unmet prerequisite and exposes only the
latest safe failure diagnostic. For unattended operation, use the native host CLI and installation
descriptor:
```sh
docker compose --env-file deploy/env/local.env \
-f compose.yaml -f deploy/compose.local.yaml \
-f deploy/compose.preprocess.yaml --profile preprocess run --rm preprocess-evidence
tht --installation /absolute/path/thothii-installation.yaml \
workspace preprocess run --workspace <workspace-id>
tht --installation /absolute/path/thothii-installation.yaml \
workspace preprocess clear --workspace <workspace-id>
```
Replace the final service with `preprocess-dwh` when required. The overlay makes each job wait for the internal Qdrant
service health checks and embedding model initialization; no separate semantic-service startup is
required.
The one-shot command starts the profile-gated `workspace-maintenance` service, reads database
metadata from PostgreSQL, and rebuilds LSH plus schema/Evidence vectors. The clear command removes
those derived artifacts while preserving the separate Memory collection. The core remains unavailable
until preprocessing completes. See [Evidence](docs/evidence.md) and the
[workspace preprocessing CLI contract](docs/contracts/workspace-preprocessing-cli.md).
S3 Evidence uses the optional `tht[s3]` dependency and canonical `s3://bucket/key` provenance.
AWS endpoints are used when no custom URL is supplied. Every custom endpoint is an explicit egress
@@ -255,7 +261,7 @@ Compose project name by passing `--confirm-project`:
The restore script stops `qdrant`, validates the exact labeled target, stages the current volume
contents for rollback, extracts the requested archive into the volume, and then returns the
service to its prior running state. It restores semantic storage only. Before reopening write
traffic, the workspace registry must already be at a reviewed v3 descriptor revision compatible
traffic, the workspace registry must already be at a reviewed v4 descriptor revision compatible
with the restored collection; then run backend health checks and a known retrieval query. The
helper does not restore descriptors, rename collections, or reconcile an incompatible collection
contract.
@@ -281,6 +287,33 @@ Copy `deploy/secrets/thothii.secrets.example` to a protected host file, include
keys, and set its absolute path as `THT_SECRETS_FILE` in the operator env. Keep Pi's native
provider auth in the separate protected file named by `PI_AUTH_FILE`.
Interactive sessions, Description Generation, and embedding share the protected installation
descriptor's `modelCatalog`. Set `THT_INSTALLATION_CONFIG_SOURCE` to that exact host file; `tht`
validates it and generates the runtime catalog, Pi adapters, and Compose override before startup.
Each authenticated provider stores only an audited `apiKeyEnv` reference; the referenced value stays
in the secret bundle. A provider may use `authentication.mode: none` only with an explicit keyless
endpoint. The browser receives only eligible model IDs, labels, and the catalog default.
Before enabling Description Generation, approve the selected model provider for bounded source-data
disclosure. Every catalog column has a **Sensitive** flag that defaults to `false`. Administrators can
request an AI proposal based only on structural metadata, then must review and save the resulting
checkboxes themselves. The proposal never reads column contents and is not persisted automatically.
For unprotected columns, a request may send up to five real source rows and five representative
distinct, non-null example values. Protected columns are omitted from source reads and replaced in the
prompt by deterministic plausible values derived only from column metadata. Samples are transient and
are not stored in generation runs, run logs, application logs, API responses, or catalog metadata;
prompt and sample snapshots are not retained. A flag change applies to later generations and does not
regenerate existing descriptions.
Description Generation is an interactive Database Management operation, not a user-facing CLI.
The installation runs at most one sequential generation at a time. The run drawer exposes safe
ordered events through SSE with polling fallback, Stop terminates the current helper while keeping
already stored results, and Run history retains terminal runs for inspection. A backend restart
marks queued or running work interrupted instead of resuming it; use Generate Missing to continue.
Unlock is reserved for a stale recorded run and is rejected while a local start, worker, or helper
is still live.
The bundle is mounted read-only as `/run/secrets/thothii.secrets` and must be mode `0600` or
`0400` on the host. Docker's runtime `0444` mode is accepted only beneath `/run/secrets`; see
[`deploy/secrets/README.md`](deploy/secrets/README.md). A PEM CA chain is deliberately not a
@@ -290,22 +323,11 @@ the host/secret-manager materialization and add a reviewed Compose override that
does not create that mount. The frontend remains on loopback; the authenticated host proxy is the
only public listener.
Set the selected model provider in application settings (or `PI_PROVIDER`). For each Pi spawn the
backend validates and reads `THT_MODEL_API_KEY` from the bundle, then exposes its value only as the provider's
recognized child variable (for example `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `GEMINI_API_KEY`, or
`ZAI_API_KEY`). Neither the generic file path nor deprecated `PI_PROVIDER_API_KEY` is inherited by
Pi. Local providers such as Ollama require no model key.
`THT_MODEL_API_KEY` supports Pi providers whose authentication is exactly one key:
`ant-ling`, `anthropic`, `cerebras`, `deepseek`, `fireworks`, `github-copilot`, `google`
(including the `gemini` alias), `google-vertex` when using its API-key mode, `groq`,
`huggingface`, `kimi-coding`, `minimax`, `minimax-cn`, `mistral`, `moonshotai`,
`moonshotai-cn`, `nvidia`, `openai`, `opencode`, `opencode-go`, `openrouter`, `together`,
`vercel-ai-gateway`, `xai`, the four `xiaomi*` providers, `zai`, and `zai-coding-cn`.
Compound providers are deliberately unsupported: `amazon-bedrock`, `azure-openai-responses`,
`cloudflare-workers-ai`, and `cloudflare-ai-gateway` require multiple credential/configuration
values. Selecting one fails before Pi starts; ambient AWS, Azure, and Cloudflare credentials are
still scrubbed. Supporting them requires a future dedicated provider-specific configuration.
For each Pi spawn, the backend resolves the selected canonical provider/model in the runtime catalog,
reads exactly that provider's declared `apiKeyEnv` value from the bundle, and exposes only that key
to the child. Ambient provider credentials and secret-bundle paths are scrubbed. Providers needing a
compound credential bundle remain unsupported until the catalog gains an explicit generic contract
for them.
## User-owned session server cutover
@@ -314,6 +336,8 @@ The server profile stores sessions and per-user preferences directly in PostgreS
dual write. Use [`deploy/compose.session-server.yaml.example`](deploy/compose.session-server.yaml.example)
with the canonical base+server files and set `THT_SERVER_WORKSPACE_CONFIG` to an absolute,
protected copy of [`deploy/workspaces/server-sessions.yaml.example`](deploy/workspaces/server-sessions.yaml.example).
That file is an installation runtime template, not an authored workspace descriptor; database
bindings are injected from the PostgreSQL Metadata Catalog for each runtime lease.
The runtime login needs membership in the no-login database role `thoth_sessions_runtime` only.
The distinct, one-shot migrator login needs migration authority and uses
+2133 -19
View File
File diff suppressed because it is too large Load Diff
+9 -1
View File
@@ -6,9 +6,12 @@
"dev": "tsx watch src/server.ts",
"prebuild": "node scripts/clean-dist.mjs",
"build": "tsc -p tsconfig.json",
"catalog:migrate": "node dist/catalog/migrate.js",
"sensitivity:shadow": "node dist/catalog/sensitivity-shadow.js",
"test": "vitest run",
"start": "node dist/server.js",
"test:schema-v3-verifier": "python3 -I -B scripts/test_revision_state_policy.py && node --test scripts/verify-workspace-descriptor-files.test.mjs scripts/revision-state-policy.test.mjs"
"test:schema-v4-verifier": "python3 -I -B scripts/test_revision_state_policy.py && node --test scripts/verify-workspace-descriptor-files.test.mjs scripts/revision-state-policy.test.mjs",
"test:schema-v3-verifier": "npm run test:schema-v4-verifier"
},
"dependencies": {
"@fastify/cookie": "11.1.2",
@@ -16,13 +19,18 @@
"@fastify/rate-limit": "11.2.0",
"@types/pg": "^8.20.3",
"fastify": "^5.0.0",
"kysely": "^0.29.5",
"libphonenumber-js": "1.13.12",
"openid-client": "6.8.5",
"pg": "^8.22.0",
"validator": "13.15.35",
"yaml": "^2.9.0",
"zod": "^4.4.3"
},
"devDependencies": {
"@testcontainers/postgresql": "^12.1.0",
"@types/node": "24.13.3",
"@types/validator": "13.15.10",
"tsx": "^4.19.0",
"typescript": "^5.6.0",
"vitest": "^2.1.0"
@@ -0,0 +1,8 @@
34448b82c17d60fec9b65b1f093c115ddbaadc04beb1b0140b6bfed2e012a930 ./.gitattributes
4d9344c58a2a2ea4bb4ff4f7c611a853cf413205fc10d0cace564eba06f73828 ./README.md
180f0a10d1d5ed5ce3318db0bcb0b1b7780d79a52f0a8fc3acbd27f74536d0e4 ./THOTHII_MODEL_REVISION
164f17362bcf9d114067d3465e7374bfdd79ce6b605acb745de5a49dabb9595c ./config.json
f27dd63cc43a248d2566f0b6ad7a115db353676ce0561dcbca45bac766464c1a ./encoder_config/config.json
0280f6f39f6012da50b6640bad438d9b7e763a1b0102094115d1b710c4dd79b6 ./model.safetensors
f6df10ec83bea993035b2dd7c39345a3d4fcf23421c2adb6cb4ffc1e6d1bc4b5 ./tokenizer.json
233beed1f1095cccfc7907cde31a8d90a0c6aa4fdfaf6493f8e55fd162e81ae6 ./tokenizer_config.json
@@ -0,0 +1,34 @@
# Optional offline CPU pack. Fully version-locked in its own venv; not part of the base image.
--extra-index-url https://download.pytorch.org/whl/cpu
accelerate==1.14.0
annotated-types==0.8.0
certifi==2026.7.22
charset-normalizer==3.5.1
filelock==3.32.5
fsspec==2026.7.0
gliner2[local]==2.0.0
hf-xet==1.6.0
huggingface-hub==0.36.2
idna==3.19
Jinja2==3.1.6
MarkupSafe==3.0.3
mpmath==1.3.0
networkx==3.6.1
numpy==2.5.2
packaging==26.3
peft==0.20.0
psutil==7.2.2
pydantic==2.13.5
pydantic-core==2.46.5
PyYAML==6.0.3
regex==2026.9.3
requests==2.34.2
safetensors==0.8.0
sympy==1.14.0
tokenizers==0.22.2
torch==2.14.0+cpu
tqdm==4.70.0
transformers==4.57.6
typing-extensions==4.16.0
typing-inspection==0.4.4
urllib3==2.7.0
+301
View File
@@ -0,0 +1,301 @@
"""Offline, CPU-only JSONL worker for optional sensitivity NER evidence."""
from __future__ import annotations
import argparse
import contextlib
import ctypes
import errno
import hashlib
import json
import os
import socket
import sys
import tempfile
from pathlib import Path
from typing import Any
PII_LABELS = [
"person",
"full_name",
"first_name",
"middle_name",
"last_name",
"date_of_birth",
"email",
"phone_number",
"address",
"street_address",
"city",
"state_or_region",
"postal_code",
"country",
"government_id",
"national_id_number",
"passport_number",
"drivers_license_number",
"license_number",
"tax_id",
"tax_number",
"bank_account",
"account_number",
"routing_number",
"iban",
"payment_card",
"card_number",
"card_expiry",
"card_cvv",
"username",
"ip_address",
"account_id",
"sensitive_account_id",
"password",
"secret",
"api_key",
"access_token",
"recovery_code",
"sensitive_date",
"document_date",
"expiration_date",
"transaction_date",
]
_MODEL_COMPAT_DIRECTORY: tempfile.TemporaryDirectory[str] | None = None
_EXPECTED_MODEL_REVISION = "c153999da5f4c509df4322b0c6a1baf3d2c284d7"
def _arguments() -> argparse.Namespace:
parser = argparse.ArgumentParser(add_help=False)
parser.add_argument("--model", required=True)
parser.add_argument("--threads", type=int, default=2)
return parser.parse_args()
def _disable_network() -> None:
libc = ctypes.CDLL(None, use_errno=True)
libc.prctl.argtypes = [
ctypes.c_int,
ctypes.c_ulong,
ctypes.c_ulong,
ctypes.c_ulong,
ctypes.c_ulong,
]
libc.prctl.restype = ctypes.c_int
if libc.prctl(38, 1, 0, 0, 0) != 0: # PR_SET_NO_NEW_PRIVS
raise RuntimeError("cannot enable no-new-privileges for network isolation")
try:
seccomp = ctypes.CDLL("libseccomp.so.2", use_errno=True)
except OSError as error:
raise RuntimeError("libseccomp is required for network isolation") from error
seccomp.seccomp_init.argtypes = [ctypes.c_uint32]
seccomp.seccomp_init.restype = ctypes.c_void_p
seccomp.seccomp_syscall_resolve_name.argtypes = [ctypes.c_char_p]
seccomp.seccomp_syscall_resolve_name.restype = ctypes.c_int
seccomp.seccomp_rule_add.argtypes = [
ctypes.c_void_p,
ctypes.c_uint32,
ctypes.c_int,
ctypes.c_uint,
]
seccomp.seccomp_rule_add.restype = ctypes.c_int
seccomp.seccomp_load.argtypes = [ctypes.c_void_p]
seccomp.seccomp_load.restype = ctypes.c_int
seccomp.seccomp_release.argtypes = [ctypes.c_void_p]
seccomp.seccomp_release.restype = None
allow = 0x7FFF0000 # SCMP_ACT_ALLOW
deny = 0x00050000 | errno.EPERM # SCMP_ACT_ERRNO(EPERM)
filter_context = seccomp.seccomp_init(allow)
if not filter_context:
raise RuntimeError("cannot initialize network syscall filter")
try:
for syscall in (
"socket",
"connect",
"sendto",
"sendmsg",
"sendmmsg",
"bind",
"listen",
"accept",
"accept4",
):
syscall_number = seccomp.seccomp_syscall_resolve_name(syscall.encode("ascii"))
if syscall_number < 0:
raise RuntimeError(f"cannot resolve network syscall: {syscall}")
if seccomp.seccomp_rule_add(filter_context, deny, syscall_number, 0) != 0:
raise RuntimeError(f"cannot block network syscall: {syscall}")
if seccomp.seccomp_load(filter_context) != 0:
raise RuntimeError("cannot activate network syscall filter")
finally:
seccomp.seccomp_release(filter_context)
def blocked(*_args: Any, **_kwargs: Any) -> Any:
raise PermissionError(errno.EPERM, "network disabled")
socket.socket = blocked # type: ignore[assignment]
socket.create_connection = blocked # type: ignore[assignment]
def _verify_model(path: Path) -> None:
revision_path = path / "THOTHII_MODEL_REVISION"
try:
revision = revision_path.read_text(encoding="utf-8").strip()
except OSError as error:
raise RuntimeError("model revision marker is unavailable") from error
if revision != _EXPECTED_MODEL_REVISION:
raise RuntimeError("model revision is not approved")
manifest_path = Path(__file__).with_name("sensitivity-ner-model-sha256.txt")
try:
manifest = manifest_path.read_text(encoding="utf-8").splitlines()
except OSError as error:
raise RuntimeError("model checksum manifest is unavailable") from error
for line in manifest:
checksum, separator, relative_name = line.partition(" ")
if not separator or len(checksum) != 64 or not relative_name.startswith("./"):
raise RuntimeError("model checksum manifest is invalid")
relative_path = Path(relative_name[2:])
if relative_path.is_absolute() or ".." in relative_path.parts:
raise RuntimeError("model checksum path is invalid")
model_file = path / relative_path
if not model_file.is_file() or model_file.is_symlink():
raise RuntimeError("approved model file is unavailable")
digest = hashlib.sha256()
with model_file.open("rb") as stream:
for chunk in iter(lambda: stream.read(1024 * 1024), b""):
digest.update(chunk)
if digest.hexdigest() != checksum:
raise RuntimeError("approved model checksum does not match")
def _transformers4_model_path(path: Path) -> Path:
"""Adapt tokenizer metadata emitted by Transformers 5 without changing pinned weights.
GLiNER2 2.0.0 officially requires Transformers <5, while current Fastino checkpoints were
saved by Transformers 5.8.0. Transformers 4 calls the same list
``additional_special_tokens``; Transformers 5 renamed it to ``extra_special_tokens`` and
changed its type. Keep the downloaded model immutable and create a temporary symlink view
containing only the compatibility metadata needed by the supported GLiNER2 dependency set.
"""
tokenizer_path = path / "tokenizer_config.json"
try:
tokenizer = json.loads(tokenizer_path.read_text(encoding="utf-8"))
except (OSError, json.JSONDecodeError) as error:
raise RuntimeError("invalid tokenizer configuration") from error
extra_tokens = tokenizer.get("extra_special_tokens")
if extra_tokens is None:
return path
if not isinstance(extra_tokens, list) or not all(isinstance(token, str) for token in extra_tokens):
raise RuntimeError("unsupported extra_special_tokens configuration")
if "additional_special_tokens" in tokenizer:
raise RuntimeError("ambiguous special-token configuration")
global _MODEL_COMPAT_DIRECTORY
_MODEL_COMPAT_DIRECTORY = tempfile.TemporaryDirectory(prefix="thothii-ner-model-")
compatible_path = Path(_MODEL_COMPAT_DIRECTORY.name)
for child in path.iterdir():
if child.name == tokenizer_path.name:
continue
(compatible_path / child.name).symlink_to(child, target_is_directory=child.is_dir())
tokenizer["additional_special_tokens"] = tokenizer.pop("extra_special_tokens")
(compatible_path / tokenizer_path.name).write_text(
json.dumps(tokenizer, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
return compatible_path
def _load_model(model_path: str, threads: int) -> Any:
path = Path(model_path).resolve(strict=True)
if not path.is_dir():
raise RuntimeError("model path must be a local directory")
_verify_model(path)
os.environ["CUDA_VISIBLE_DEVICES"] = ""
os.environ["HIP_VISIBLE_DEVICES"] = ""
os.environ["HF_HUB_OFFLINE"] = "1"
os.environ["TRANSFORMERS_OFFLINE"] = "1"
import torch
from gliner2 import AutoExtractor
torch.set_num_threads(max(1, min(threads, 8)))
torch.set_num_interop_threads(1)
compatible_path = _transformers4_model_path(path)
with contextlib.redirect_stdout(sys.stderr):
model = AutoExtractor.from_pretrained(str(compatible_path), map_location="cpu")
_disable_network()
return model
def _request(value: Any) -> tuple[str, list[dict[str, str]]]:
if not isinstance(value, dict) or not isinstance(value.get("id"), str):
raise ValueError("invalid request")
candidates = value.get("candidates")
if not isinstance(candidates, list) or not 1 <= len(candidates) <= 128:
raise ValueError("invalid candidates")
parsed: list[dict[str, str]] = []
for candidate in candidates:
if not isinstance(candidate, dict):
raise ValueError("invalid candidate")
column_id = candidate.get("columnId")
text = candidate.get("text")
if not isinstance(column_id, str) or not isinstance(text, str) or not 1 <= len(text) <= 500:
raise ValueError("invalid candidate")
parsed.append({"columnId": column_id, "text": text})
return value["id"], parsed
def _detect(model: Any, candidates: list[dict[str, str]]) -> list[dict[str, Any]]:
evidence: list[dict[str, Any]] = []
for candidate in candidates:
result = model.extract_entities(
candidate["text"],
PII_LABELS,
threshold=0.5,
include_confidence=True,
)
entities = result.get("entities", {}) if isinstance(result, dict) else {}
best: tuple[str, float] | None = None
if isinstance(entities, dict):
for label, matches in entities.items():
if label not in PII_LABELS or not isinstance(matches, list):
continue
for match in matches:
if not isinstance(match, dict):
continue
confidence = match.get("confidence")
if not isinstance(confidence, (int, float)) or not 0 <= confidence <= 1:
continue
if best is None or confidence > best[1]:
best = (label, float(confidence))
if best is not None:
evidence.append(
{
"columnId": candidate["columnId"],
"label": best[0],
"confidence": best[1],
}
)
return evidence
def main() -> int:
args = _arguments()
model = _load_model(args.model, args.threads)
print(json.dumps({"ready": True}, separators=(",", ":")), flush=True)
for line in sys.stdin:
request_id = "invalid"
try:
request_id, candidates = _request(json.loads(line))
response = {"id": request_id, "ok": True, "evidence": _detect(model, candidates)}
except Exception:
response = {"id": request_id, "ok": False, "error": "detection_failed"}
print(json.dumps(response, separators=(",", ":")), flush=True)
return 0
if __name__ == "__main__":
raise SystemExit(main())
+1 -6
View File
@@ -1195,13 +1195,8 @@ export async function executeChecks({ checks, failAt, recorder } = {}) {
function baseWorkspace(id, evidenceSource) {
return {
workspace: { schema_version: 3, id, name: `P1 ${id}`, language: "en" },
workspace: { schema_version: 4, id, name: `P1 ${id}`, language: "en" },
dwh: { engine: "postgres", database: "postgres", schema: "public", supported_transports: ["postgres_direct"] },
semantic_index: {
vector_store: { engine: "qdrant", collection: id, dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
evidence: { source: evidenceSource, policy: { max_chunk_chars: 4000, retain_published_generations: 3 } },
};
}
+1 -1
View File
@@ -109,7 +109,7 @@ async function validateDistFiles(repo,files){const dist=join(repo,"backend","dis
export async function readManualOwnership({repositoryRoot=defaultRepositoryRoot}={}){const repo=realpathSync(repositoryRoot),root=fixedManualRoot(repo);noSymlinkExisting(repo,root);let rootEntry,ownershipEntry;try{rootEntry=await lstat(root);ownershipEntry=await lstat(join(root,"ownership.json"));}catch{throw new Error("manual ownership is missing");}if(!rootEntry.isDirectory()||rootEntry.isSymbolicLink()||await realpath(root)!==root||!ownershipEntry.isFile()||ownershipEntry.isSymbolicLink())throw new Error("manual ownership is unsafe");let value;try{value=JSON.parse(await readFile(join(root,"ownership.json"),"utf8"));}catch{throw new Error("manual ownership is malformed");}const baseValid=value.schemaVersion===1&&value.kind==="p1-manual-acceptance"&&HEX64.test(value.nonce??"")&&value.repositoryRoot===repo&&value.root===root&&value.status==="PENDING"&&["PREPARING","READY"].includes(value.stage)&&value.listener?.host===HOST&&value.listener?.port===PORT&&value.listener?.state==="stopped"&&typeof value.createdAt==="string"&&validEntrypoint(value.entrypoint,repo)&&validDistManifest(value.distManifest,root)&&JSON.stringify(value.resources)===JSON.stringify([root,{kind:"fastify",host:HOST,port:PORT}]);const readyLog=value.backendLog?.path===join(root,"logs/backend.log")&&Number.isSafeInteger(value.backendLog?.dev)&&Number.isSafeInteger(value.backendLog?.ino);if(!baseValid||(value.stage==="READY"?!readyLog:value.backendLog!==null))throw new Error("manual ownership identity mismatch");return value;}
async function run(executable,argv,options={}){return await exec(executable,argv,{...options,maxBuffer:2*1024*1024,encoding:"utf8"});}
function descriptor(id,source){return{workspace:{schema_version:3,id,name:`P1 ${id}`,language:"en"},dwh:{engine:"postgres",database:"postgres",schema:"public",supported_transports:["postgres_direct"]},semantic_index:{vector_store:{engine:"qdrant",collection:id,dimensions:1024,distance:"cosine"},embedding:{provider:"ollama_internal",model:"qwen3-embedding:0.6b",dimensions:1024}},llm_policy:{allowed:["zai/glm-5.2"]},evidence:{source,policy:{max_chunk_chars:4000,retain_published_generations:3}}};}
function descriptor(id,source){return{workspace:{schema_version:4,id,name:`P1 ${id}`,language:"en"},dwh:{engine:"postgres",database:"postgres",schema:"public",supported_transports:["postgres_direct"]},evidence:{source,policy:{max_chunk_chars:4000,retain_published_generations:3}}};}
function descriptors(){return[descriptor("p1-filesystem",{type:"filesystem",uri:"workspace-content/p1-filesystem/evidence",patterns:["**/*.md"],max_bytes:10485760}),descriptor("p1-http",{type:"http",uris:["https://evidence.example.test/guide.md"],authentication:"signed_urls_file",connect_timeout_ms:1250,read_timeout_ms:30001,max_bytes:12345,max_redirects:2,allow_private_hosts:false,max_cache_bytes:67890}),descriptor("p1-s3",{type:"s3",uri:"s3://p1-evidence/published/",endpoint_url:"https://s3.example.test/",region:"eu-west-1",credentials:"static_files",trusted_endpoint:true,allow_private_endpoint:false,allow_insecure_endpoint:false,max_bytes:12345,max_objects:33,max_pages:4,page_size:5})];}
function quote(value){return `'${String(value).replaceAll("'",`'"'"'`)}'`;}
async function checkPrerequisites(repo){for(const path of ["scripts/p1-acceptance.sh","scripts/test-p1-acceptance.sh","backend/scripts/p1-acceptance.mjs","backend/dist/server.js"]){try{await access(join(repo,path));}catch{throw new Error(`Task 8 prerequisite is missing: ${path}`);}}for(const command of ["node","npm","git","curl","unzip","zipinfo","lsof","python3"]){try{await run(command,[command==="unzip"||command==="lsof"?"-v":command==="zipinfo"?"-h":"--version"]);}catch{throw new Error(`missing prerequisite: ${command}`);}}const tht=join(repo,"harness",".venv","bin","tht");try{await access(tht,constants.X_OK);}catch{throw new Error("missing prerequisite: harness/.venv/bin/tht");}}
@@ -403,7 +403,7 @@ test("generated render command validates saved responses and owned snapshot befo
});
const renderSnapshotYaml=`workspace:
schema_version: 3
schema_version: 4
id: p1-filesystem
name: P1 filesystem
language: en
@@ -412,11 +412,6 @@ dwh:
database: postgres
schema: public
supported_transports: [postgres_direct]
semantic_index:
vector_store: {engine: qdrant, collection: p1-filesystem, dimensions: 1024, distance: cosine}
embedding: {provider: ollama_internal, model: qwen3-embedding:0.6b, dimensions: 1024}
llm_policy:
allowed: [zai/glm-5.2]
evidence:
source: {type: filesystem, uri: workspace-content/p1-filesystem/evidence, patterns: ["**/*.md"], max_bytes: 10485760}
policy: {max_chunk_chars: 4000, retain_published_generations: 3}
+2 -7
View File
@@ -16,7 +16,7 @@ async function fixture() {
await writeFile(join(root,"installation/base.yaml"),"{}\n");
const secret=join(root,"fixture-secrets/dwh-password"); await writeFile(secret,"not-inspected",{mode:0o600});
await writeFile(snapshot,`workspace:
schema_version: 3
schema_version: 4
id: p1-filesystem
name: P1 filesystem
language: en
@@ -25,11 +25,6 @@ dwh:
database: postgres
schema: public
supported_transports: [postgres_direct]
semantic_index:
vector_store: {engine: qdrant, collection: p1-filesystem, dimensions: 1024, distance: cosine}
embedding: {provider: ollama_internal, model: qwen3-embedding:0.6b, dimensions: 1024}
llm_policy:
allowed: [zai/glm-5.2]
evidence:
source: {type: filesystem, uri: workspace-content/p1-filesystem/evidence, patterns: ["**/*.md"], max_bytes: 10485760}
policy: {max_chunk_chars: 4000, retain_published_generations: 3}
@@ -61,7 +56,7 @@ test("renderer refuses snapshot manifest head, digest, and expected-digest tampe
test("renderer refuses a missing or malformed snapshot manifest",async()=>{ const f=await fixture(); const output=join(f.root,"rendered/nomanifest.yaml"); await rm(f.manifestPath); await assert.rejects(call(f,{outputPath:output}),/snapshot manifest.*(missing|unbounded|unsafe)/); await writeFile(f.manifestPath,"{not json"); await assert.rejects(call(f,{outputPath:output}),/snapshot manifest.*malformed/); await assert.rejects(lstat(output)); assert.deepEqual(await runtimeLeases(f),[]); });
test("renderer rejects a regular snapshot replacement against its manifest",async()=>{ const f=await fixture(); const output=join(f.root,"rendered/replaced.yaml"); await assert.rejects(call(f,{outputPath:output,beforePublish:async()=>{await writeFile(f.snapshot,"workspace:\n schema_version: 3\n id: p1-filesystem\n name: replaced\n")}}),/snapshot content changed/); await assert.rejects(lstat(output)); });
test("renderer rejects a regular snapshot replacement against its manifest",async()=>{ const f=await fixture(); const output=join(f.root,"rendered/replaced.yaml"); await assert.rejects(call(f,{outputPath:output,beforePublish:async()=>{await writeFile(f.snapshot,"workspace:\n schema_version: 4\n id: p1-filesystem\n name: replaced\n")}}),/snapshot content changed/); await assert.rejects(lstat(output)); });
test("renderer anchors publication when rendered parent is concurrently swapped", async()=>{
const f=await fixture(),output=join(f.root,"rendered/raced.yaml"),moved=join(f.root,"rendered-moved"),outside=join(f.repo,"outside-rendered"); await mkdir(outside);
+1 -6
View File
@@ -328,13 +328,8 @@ async function tht(ctx, argv, options = {}) {
function namespace(id) { return id.toUpperCase().replaceAll("-", "_"); }
function baseWorkspace(id, evidenceSource) {
return {
workspace: { schema_version: 3, id, name: `P1.1 ${id}`, description: `Catalog entry for ${id}`, language: "en" },
workspace: { schema_version: 4, id, name: `P1.1 ${id}`, description: `Catalog entry for ${id}`, language: "en" },
dwh: { engine: "postgres", database: "postgres", schema: "public", supported_transports: ["postgres_direct"] },
semantic_index: {
vector_store: { engine: "qdrant", collection: id, dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
evidence: { source: evidenceSource, policy: { max_chunk_chars: 4000, retain_published_generations: 3 } },
};
}
+1 -6
View File
@@ -81,13 +81,8 @@ async function git(executable, argv, options = {}) {
function namespace(id) { return id.toUpperCase().replaceAll("-", "_"); }
function baseWorkspace(id, evidenceSource) {
return {
workspace: { schema_version: 3, id, name: `P1.1 ${id}`, description: `Catalog entry for ${id}`, language: "en" },
workspace: { schema_version: 4, id, name: `P1.1 ${id}`, description: `Catalog entry for ${id}`, language: "en" },
dwh: { engine: "postgres", database: "postgres", schema: "public", supported_transports: ["postgres_direct"] },
semantic_index: {
vector_store: { engine: "qdrant", collection: id, dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
evidence: { source: evidenceSource, policy: { max_chunk_chars: 4000, retain_published_generations: 3 } },
};
}
+1 -6
View File
@@ -375,16 +375,11 @@ function installationProjectName(installationPath) {
function baseWorkspace(id, { dwhBaseUrl, evidenceSource }) {
return {
workspace: { schema_version: 3, id, name: `P2 ${id}`, language: "en" },
workspace: { schema_version: 4, id, name: `P2 ${id}`, language: "en" },
dwh: { engine: "postgres", database: "warehouse", schema: "dw", supported_transports: ["rest_api"] },
semantic_index: {
vector_store: { engine: "qdrant", collection: id, dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
diagnostics: {
dwh_rest: { method: "POST", path: "/rpc/ping", auth: "x-api-key", response: { database: "database", schema: "schema" } },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
...(evidenceSource ? { evidence: { source: evidenceSource, policy: { max_chunk_chars: 4000, retain_published_generations: 3 } } } : {}),
};
}
+1 -6
View File
@@ -376,16 +376,11 @@ function installationProjectName(installationPath) {
function baseWorkspace(id, { dwhBaseUrl, evidenceSource }) {
return {
workspace: { schema_version: 3, id, name: `P2 ${id}`, language: "en" },
workspace: { schema_version: 4, id, name: `P2 ${id}`, language: "en" },
dwh: { engine: "postgres", database: "warehouse", schema: "dw", supported_transports: ["rest_api"] },
semantic_index: {
vector_store: { engine: "qdrant", collection: id, dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
diagnostics: {
dwh_rest: { method: "POST", path: "/rpc/ping", auth: "x-api-key", response: { database: "database", schema: "schema" } },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
...(evidenceSource ? { evidence: { source: evidenceSource, policy: { max_chunk_chars: 4000, retain_published_generations: 3 } } } : {}),
};
}
+1 -6
View File
@@ -379,16 +379,11 @@ function installationProjectName(installationPath) {
function baseWorkspace(id, { dwhBaseUrl, evidenceSource }) {
return {
workspace: { schema_version: 3, id, name: `P2 ${id}`, language: "en" },
workspace: { schema_version: 4, id, name: `P2 ${id}`, language: "en" },
dwh: { engine: "postgres", database: "warehouse", schema: "dw", supported_transports: ["rest_api"] },
semantic_index: {
vector_store: { engine: "qdrant", collection: id, dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
diagnostics: {
dwh_rest: { method: "POST", path: "/rpc/ping", auth: "x-api-key", response: { database: "database", schema: "schema" } },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
...(evidenceSource ? { evidence: { source: evidenceSource, policy: { max_chunk_chars: 4000, retain_published_generations: 3 } } } : {}),
};
}
+1 -6
View File
@@ -374,16 +374,11 @@ function installationProjectName(installationPath) {
function baseWorkspace(id, { dwhBaseUrl, evidenceSource }) {
return {
workspace: { schema_version: 3, id, name: `P2 ${id}`, language: "en" },
workspace: { schema_version: 4, id, name: `P2 ${id}`, language: "en" },
dwh: { engine: "postgres", database: "warehouse", schema: "dw", supported_transports: ["rest_api"] },
semantic_index: {
vector_store: { engine: "qdrant", collection: id, dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
diagnostics: {
dwh_rest: { method: "POST", path: "/rpc/ping", auth: "x-api-key", response: { database: "database", schema: "schema" } },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
...(evidenceSource ? { evidence: { source: evidenceSource, policy: { max_chunk_chars: 4000, retain_published_generations: 3 } } } : {}),
};
}
+1 -6
View File
@@ -374,16 +374,11 @@ function installationProjectName(installationPath) {
function baseWorkspace(id, { dwhBaseUrl, evidenceSource }) {
return {
workspace: { schema_version: 3, id, name: `P2 ${id}`, language: "en" },
workspace: { schema_version: 4, id, name: `P2 ${id}`, language: "en" },
dwh: { engine: "postgres", database: "warehouse", schema: "dw", supported_transports: ["rest_api"] },
semantic_index: {
vector_store: { engine: "qdrant", collection: id, dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
diagnostics: {
dwh_rest: { method: "POST", path: "/rpc/ping", auth: "x-api-key", response: { database: "database", schema: "schema" } },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
...(evidenceSource ? { evidence: { source: evidenceSource, policy: { max_chunk_chars: 4000, retain_published_generations: 3 } } } : {}),
};
}
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env node
import { readFileSync } from "node:fs";
const path = process.env.THT_SSH_PASSPHRASE_FILE;
if (!path) process.exit(1);
process.stdout.write(readFileSync(path));
@@ -15,28 +15,32 @@ const allowedKinds = new Set(["policy_text", "workspace_descriptor", "deployment
// opener line through the closer line (including physical line endings). These
// blocks are reviewed non-workspace runtime/config generation, not semantic proof.
const reviewedExpandableBlocks = new Map([
["scripts/preprocess-smoke.sh", [
{ sha256: "fc530dc721c946644ab6552bbd46b7918d6c5f11f06f3495b6ea1fcda819b38d", rationale: "Generates the reviewed preprocess Compose override." },
["scripts/test-dwh-auth-nginx-integration.sh", [
{ sha256: "ead57234ad3520b5c7d4262b772957cbc7b9589da4f35fb17b160f948eb2ac7b", rationale: "Generates the reviewed isolated Nginx integration configuration." },
]],
["scripts/test-install-tht.sh", [
{ sha256: "37f18ce7ce93cb8b84f3b3708462cc16d50fdc7bab22836c382dbacf8382f05f", rationale: "Generates the reviewed synthetic tht installer artifact." },
]],
["scripts/test-server-pi-state-topology.sh", [
{ sha256: "6f746f7e8442b0a6ea0e216607a6a923d94b24cd8fa17fa2d1dac56e6f14f7ef", rationale: "Generates the isolated server topology test environment." },
{ sha256: "435c769b8cbd7b834f56fdabddb86ba04fb404dd0d8a6b7c21719a8b0f7cf011", rationale: "Generates the reviewed model-catalog projection override for the isolated server topology test." },
{ sha256: "6ae9567db53d6cd45a2c19c98acaf45f382450b157ea7d6f6d35125f68c50947", rationale: "Generates the isolated server topology test environment, including its installation descriptor and authentication configuration root." },
]],
["scripts/test-vector-backup-restore-safety.sh", [
{ sha256: "40b8a10a3c06aaa98e324fbf688b7d1f5cead330d7ba7eef98e06256d412a85a", rationale: "Generates the reviewed restore safety manifest." },
]],
["scripts/test-windows-clone-contract.ps1", [
{ sha256: "80f4880576a0679cb58e7b92600e7a90550c93c254553a2d4b299539f9ff0bcf", rationale: "Generates reviewed Windows clone test configuration." },
{ sha256: "6166294bdc8a8bf6436ad402bcbf7cae0f3b67dc6051cecfcca79267a62b082c", rationale: "Same reviewed block in the repository-required CRLF checkout representation." },
{ sha256: "3216201d59400ed7d1ec23e536634b8235a2e78b336e45b4dc598624920f0057", rationale: "Generates reviewed Windows clone test configuration." },
{ sha256: "a4044bb38b27e8120e90d65a0695fe0afd7757c067ae8dd67f170edf569a1de0", rationale: "Same reviewed block in the repository-required CRLF checkout representation." },
{ sha256: "3204f772d33cad42bcac99191507051aefb2c91d2935bec6698b956e44f9bf45", rationale: "Generates reviewed Windows clone test configuration with its authentication configuration root." },
{ sha256: "f4814d842a7502b7ef30fd6b224d5cb17b0ffd6fb2367c41c49ac16587536d93", rationale: "Same reviewed block in the repository-required CRLF checkout representation." },
{ sha256: "6f25ce3b58cea47b74fe9319ed917d8089a2fb334bc0d469daa7e1f10865d870", rationale: "Generates the reviewed Windows Compose override for the canonical service topology." },
{ sha256: "45a3cf19f7ce697b858b63d27a4edc7fefa2414d0408e7b6d72a65c86d314f5b", rationale: "Same reviewed Compose override in the repository-required CRLF checkout representation." },
{ sha256: "5d0d1a3fc45e99b3aacaf4ee5dd09a6bee1937784375dfe4bcfaa4ae32cfb9de", rationale: "Generates reviewed Windows clone test configuration." },
{ sha256: "b903e5dae953ae1372f1a5276f12a92ed3dd632b897f3afe5e00c646d90a1b42", rationale: "Same reviewed block in the repository-required CRLF checkout representation." },
]],
["scripts/unified-deployment-smoke.sh", [
{ sha256: "ca0c17d9ff8dc0fbe018fc1c5510eb33bc667a936fbe44a9be2d311390576825", rationale: "Generates reviewed Task 13 runtime configuration." },
{ sha256: "31ec00cc315b52da4a3bb6e3fba2d40aef29cdcd090bbc5d14c31f1aebbcfd04", rationale: "Generates reviewed Task 13 runtime configuration." },
{ sha256: "d92822815357ce3424e1a6eb43923df2b37b4fd93a3b5465ee9dfc69559ab0ed", rationale: "Generates reviewed Task 13 runtime configuration." },
{ sha256: "d6b8b7b951936c0452a485e9ee3b18a61251556581d6f7a2ce66f994b5700695", rationale: "Generates reviewed Task 13 runtime configuration." },
{ sha256: "b6c0826151b2c8b955399d1abf5b691cc8fe6b6454b17da000dde7ba3bc55d2d", rationale: "Generates the reviewed local Task 13 Compose override with normalized catalog mounts." },
{ sha256: "24f69d12b8554aa2bebba455be99fde3e60743eef5a40fa2ef5b29397a477c03", rationale: "Generates the reviewed local Task 13 installation descriptor with its model catalog." },
{ sha256: "526006fa6d48a8080b3834723630c64de5005a67243e944ebf1da15212b4d654", rationale: "Generates the reviewed server Task 13 Compose override." },
{ sha256: "406ccead1967f642225c946fc4a23fe5b019c9764cc5153e1125876ade16ec90", rationale: "Generates the reviewed projected-auth server Task 13 installation descriptor with its model catalog." },
]],
["scripts/vector-backup.sh", [
{ sha256: "571899db49dfdcec8107fbe1e0a86a61e7581979d3c4c248c20546843e275bcf", rationale: "Generates the reviewed backup manifest inside the helper command." },
@@ -99,6 +103,7 @@ function isPolicyImplementationException(label, category) {
]);
if (implementations.has(label)) return true;
if (category === "migration-marker" && new Set([
"backend/src/workspaces/schema.ts",
"scripts/workspace_descriptor_doc_contract.py",
"scripts/test_workspace_descriptor_doc_contract.py",
"backend/scripts/clean-dist.test.mjs",
@@ -169,7 +174,7 @@ function validateWorkspaceSource(source, label, { requireWorkspace, expandable =
try {
parseWorkspaceYaml(source);
} catch (error) {
throw new Error(`${label}: workspace descriptor is not valid schema v3: ${error instanceof Error ? error.message : String(error)}`);
throw new Error(`${label}: workspace descriptor is not valid schema v4: ${error instanceof Error ? error.message : String(error)}`);
}
return true;
}
@@ -33,20 +33,20 @@ function bashN(root, path) {
function replaceWorkspaceKeys(source, workspaceKey, schemaLine) {
return source
.replace(/^workspace:$/m, workspaceKey)
.replace(/^ schema_version: 3$/m, schemaLine);
.replace(/^ schema_version: 4$/m, schemaLine);
}
test("production parser accepts semantic v3 with quoted Unicode/tagged keys and spacing", async (t) => {
test("production parser accepts semantic v4 with quoted Unicode/tagged keys and spacing", async (t) => {
const root = await fixture(t);
const unicode = replaceWorkspaceKeys(
canonicalDescriptor,
'"\\u0077orkspace" :',
' "\\u0073chema_version" : 3',
' "\\u0073chema_version" : 4',
);
const tagged = replaceWorkspaceKeys(
canonicalDescriptor,
"!!str workspace :",
" !!str schema_version : 3",
" !!str schema_version : 4",
);
await put(root, "deploy/workspaces/unicode.yaml", unicode);
await put(root, "deploy/workspaces/tagged.yaml", tagged);
@@ -59,14 +59,15 @@ test("production parser accepts semantic v3 with quoted Unicode/tagged keys and
});
});
test("production parser rejects fancy keys with every non-v3 or ambiguous value", async (t) => {
test("production parser rejects fancy keys with every non-v4 or ambiguous value", async (t) => {
const invalid = [
["unicode-v2", '"\\u0077orkspace" :', ' "\\u0073chema_version" : 2'],
["tagged-leading-zero", "!!str workspace :", " !!str schema_version : 02"],
["hexadecimal", "workspace :", " schema_version : 0x2"],
["multiline", "workspace :", " schema_version : >\n 3"],
["duplicate", "workspace :", " schema_version : 3\n schema_version: 3"],
["inline", "workspace: { schema_version: 3 }", " schema_version: 3"],
["unicode-v3", '"\\u0077orkspace" :', ' "\\u0073chema_version" : 3'],
["tagged-leading-zero", "!!str workspace :", " !!str schema_version : 03"],
["hexadecimal", "workspace :", " schema_version : 0x3"],
["multiline", "workspace :", " schema_version : >\n 4"],
["duplicate", "workspace :", " schema_version : 4\n schema_version: 4"],
["inline", "workspace: { schema_version: 4 }", " schema_version: 4"],
];
for (const [name, workspaceKey, schemaLine] of invalid) {
await t.test(name, async () => {
@@ -120,7 +121,7 @@ test("PowerShell embedded workspace mappings are rejected while bundle-only stri
const root = await fixture(t);
const source = [
"$workspace = @'",
canonicalDescriptor.replace(" schema_version: 3", " schema_version: 0x2").trimEnd(),
canonicalDescriptor.replace(" schema_version: 4", " schema_version: 0x2").trimEnd(),
"'@",
'$bundle = @"',
"bundle:",
@@ -137,7 +138,7 @@ test("PowerShell embedded workspace mappings are rejected while bundle-only stri
test("workspace descriptor family entries require a top-level workspace", async (t) => {
const root = await fixture(t);
await put(root, "scripts/fixtures/workspace-registry-future.yaml", "bundle:\n schema_version: 3\n");
await put(root, "scripts/fixtures/workspace-registry-future.yaml", "bundle:\n schema_version: 4\n");
await assert.rejects(
verifyEntries({
root,
@@ -179,7 +180,7 @@ test("script scalar workspace remains a bundle even with descriptor-like sibling
test("standalone descriptor files require workspace to be a mapping", async (t) => {
const root = await fixture(t);
const path = "scripts/fixtures/workspace-registry-scalar.yaml";
await put(root, path, "workspace: analytics\nschema_version: 3\n");
await put(root, path, "workspace: analytics\nschema_version: 4\n");
await assert.rejects(
verifyEntries({ root, entries: [entry("workspace_descriptor", path)] }),
/workspace.*mapping/i,
@@ -193,11 +194,11 @@ test("Bash extractor supports hyphen, digit, escaped delimiters, and tab strippi
name: "hyphen-v2",
opener: "cat <<'WORKSPACE-YAML'",
delimiter: "WORKSPACE-YAML",
descriptor: canonicalDescriptor.replace(" schema_version: 3", " schema_version: 2"),
descriptor: canonicalDescriptor.replace(" schema_version: 4", " schema_version: 2"),
rejected: true,
},
{
name: "digit-v3",
name: "digit-v4",
opener: "cat <<2YAML",
delimiter: "2YAML",
descriptor: canonicalDescriptor,
@@ -207,11 +208,11 @@ test("Bash extractor supports hyphen, digit, escaped delimiters, and tab strippi
name: "escaped-v2",
opener: "cat <<WORKSPACE\\-YAML",
delimiter: "WORKSPACE-YAML",
descriptor: canonicalDescriptor.replace(" schema_version: 3", " schema_version: 2"),
descriptor: canonicalDescriptor.replace(" schema_version: 4", " schema_version: 2"),
rejected: true,
},
{
name: "tab-strip-v3",
name: "tab-strip-v4",
opener: "cat <<-'TAB-YAML'",
delimiter: "\tTAB-YAML",
descriptor: canonicalDescriptor.split("\n").map((line) => `\t${line}`).join("\n"),
@@ -272,7 +273,7 @@ test("non-stripping heredoc close requires an exact physical delimiter line", as
"#!/usr/bin/env bash",
"cat <<'---'",
"--- ",
canonicalDescriptor.replace(" schema_version: 3", " schema_version: 2").trimEnd(),
canonicalDescriptor.replace(" schema_version: 4", " schema_version: 2").trimEnd(),
"---",
"",
].join("\n");
@@ -313,7 +314,7 @@ test("double-quoted non-special backslash is preserved in the delimiter", async
"#!/usr/bin/env bash",
'cat <<"\\---"',
"---",
canonicalDescriptor.replace(" schema_version: 3", " schema_version: 2").trimEnd(),
canonicalDescriptor.replace(" schema_version: 4", " schema_version: 2").trimEnd(),
"\\---",
"",
].join("\n");
@@ -355,7 +356,7 @@ test("split heredoc operator continuation cannot bypass v2 validation", async (t
"#!/usr/bin/env bash",
"cat <\\",
"<'YAML'",
canonicalDescriptor.replace(" schema_version: 3", " schema_version: 2").trimEnd(),
canonicalDescriptor.replace(" schema_version: 4", " schema_version: 2").trimEnd(),
"YAML",
"",
].join("\n");
@@ -424,7 +425,7 @@ test("PowerShell comment backslash cannot hide a following v2 here-string", asyn
const source = [
"# harmless PowerShell comment \\",
"$workspace = @'",
canonicalDescriptor.replace(" schema_version: 3", " schema_version: 2").trimEnd(),
canonicalDescriptor.replace(" schema_version: 4", " schema_version: 2").trimEnd(),
"'@",
"",
].join("\n");
@@ -435,7 +436,7 @@ test("PowerShell comment backslash cannot hide a following v2 here-string", asyn
);
});
test("PowerShell dialect accepts normal v3 and non-workspace bundle here-strings", async (t) => {
test("PowerShell dialect accepts normal v4 and non-workspace bundle here-strings", async (t) => {
const root = await fixture(t);
const path = "scripts/powershell-valid-smoke.ps1";
const source = [
@@ -494,10 +495,10 @@ test("PowerShell cast and concatenation openers cannot hide embedded descriptors
test("expandable YAML interpolation that can hide a workspace descriptor fails closed", async (t) => {
const root = await fixture(t);
const cases = [
["braced-key", "${key}:\n schema_version: 3"],
["plain-key", "$key:\n schema_version: 3"],
["quoted-key", '"$key" :\n schema_version: 3'],
["subexpression-key", "$($key):\n schema_version: 3"],
["braced-key", "${key}:\n schema_version: 4"],
["plain-key", "$key:\n schema_version: 4"],
["quoted-key", '"$key" :\n schema_version: 4'],
["subexpression-key", "$($key):\n schema_version: 4"],
["version", "workspace:\n schema_version: $version"],
];
for (const [name, body] of cases) {
@@ -564,7 +565,7 @@ test("unmarked expandable Bash YAML cannot generate descriptor keys or values at
"key=workspace",
"cat <<YAML",
generatedKey,
" schema_version: 3",
" schema_version: 4",
"YAML",
"",
].join("\n");
@@ -600,14 +601,14 @@ test("an in-band marker cannot authorize expandable content", async (t) => {
for (const [path, source] of [
["scripts/fake-marker.sh", [
"#!/usr/bin/env bash",
"# schema-v3-only: expandable-nonworkspace",
"# schema-v4-only: expandable-nonworkspace",
"cat <<YAML",
"${DESCRIPTOR}",
"YAML",
"",
].join("\n")],
["scripts/fake-marker.ps1", [
"# schema-v3-only: expandable-nonworkspace",
"# schema-v4-only: expandable-nonworkspace",
'$yaml = @"',
"$descriptor",
'"@',
@@ -624,7 +625,6 @@ test("an in-band marker cannot authorize expandable content", async (t) => {
test("current exact reviewed expandable blocks pass only at their trusted paths", async (t) => {
const reviewedPaths = [
"scripts/preprocess-smoke.sh",
"scripts/test-server-pi-state-topology.sh",
"scripts/test-vector-backup-restore-safety.sh",
"scripts/test-windows-clone-contract.ps1",
@@ -636,19 +636,6 @@ test("current exact reviewed expandable blocks pass only at their trusted paths"
root: repositoryRoot,
entries: reviewedPaths.map((path) => entry("deployment_script", path)),
});
const root = await fixture(t);
const original = await readFile(join(repositoryRoot, "scripts/preprocess-smoke.sh"), "utf8");
await put(root, "scripts/copied-preprocess.sh", original);
await assert.rejects(
verifyEntries({ root, entries: [entry("deployment_script", "scripts/copied-preprocess.sh")] }),
/exact-content reviewed allowlist/,
);
await put(root, "scripts/preprocess-smoke.sh", original.replace('$tmp/smoke.yaml', '$tmp/other.yaml'));
await assert.rejects(
verifyEntries({ root, entries: [entry("deployment_script", "scripts/preprocess-smoke.sh")] }),
/exact-content reviewed allowlist/,
);
});
test("PowerShell tokenizer ignores opener text in comments and ordinary strings", async (t) => {
+229 -9
View File
@@ -2,10 +2,12 @@ import Fastify, { type FastifyInstance, type FastifyRequest } from "fastify";
import cors from "@fastify/cors";
import cookie from "@fastify/cookie";
import rateLimit from "@fastify/rate-limit";
import { join } from "node:path";
import { dirname, isAbsolute, join } from "node:path";
import { fileURLToPath } from "node:url";
import { tmpdir } from "node:os";
import type { AppConfig } from "./config.js";
import { ThtRunner } from "./tht/tht-runner.js";
import { createMemoryCleanup } from "./catalog/memory-cleanup.js";
import { PiProcessManager } from "./pi/pi-process-manager.js";
import { SseHub } from "./sse/sse-hub.js";
import { authenticateSession, captureAuthConfigSnapshot, configuredOrigin } from "./auth/auth.js";
@@ -22,7 +24,8 @@ import { isUsableAuthenticationSecret } from "./auth/secret-policy.js";
import { secretValue } from "./config/secret-bundle.js";
import { sessionRoutes } from "./routes/sessions.js";
import { sqlRoutes } from "./routes/sql.js";
import { metaRoutes, type ListModelsFn } from "./routes/meta.js";
import { metaRoutes } from "./routes/meta.js";
import type { ListModelsFn } from "./pi/list-models.js";
import { settingsRoutes, effectiveSettings } from "./routes/settings.js";
import { createPiModelLister } from "./pi/list-models.js";
import { createPiManagement, type PiManagementService } from "./pi/management.js";
@@ -31,12 +34,55 @@ import { ReadinessManager } from "./runtime/readiness-manager.js";
import { MaintenanceBarrier } from "./runtime/maintenance-gate.js";
import { WorkspaceRegistry } from "./workspaces/registry.js";
import { createProductionWorkspaceDiagnoser } from "./workspaces/diagnostics.js";
import { workspaceRoutes, type WorkspaceDiagnoser } from "./routes/workspaces.js";
import {
workspaceRoutes,
type WorkspaceDatabaseTester,
type WorkspaceDiagnoser,
} from "./routes/workspaces.js";
import { piManagementRoutes } from "./routes/pi-management.js";
import { supportsSessionRuntime } from "./workspaces/bindings.js";
import { resolveRuntimeBindingsWithWorkspaceSecrets } from "./workspaces/secret-requirements.js";
import type { WorkspaceDescriptor } from "./workspaces/schema.js";
import { WorkspaceSecretStore } from "./workspaces/secret-store.js";
import { createCatalogRepository } from "./catalog/repository.js";
import type { CatalogRepository } from "./catalog/types.js";
import { CatalogService } from "./catalog/service.js";
import { catalogDatabaseRoutes } from "./routes/catalog-databases.js";
import { CatalogOperationCoordinator } from "./catalog/operation-coordinator.js";
import { ConcreteCatalogPostgresAccess, type CatalogPostgresAccess } from "./catalog/postgres-access.js";
import { CatalogTableService } from "./catalog/table-service.js";
import { catalogTableRoutes } from "./routes/catalog-tables.js";
import { ConcreteCatalogSchemaIntrospector, type CatalogSchemaIntrospector } from "./catalog/schema-introspector.js";
import { CatalogSyncWorker } from "./catalog/sync-worker.js";
import { catalogSchemaRoutes } from "./routes/catalog-schema.js";
import {
loadMetadataGenerationModels,
type MetadataGenerationModels,
} from "./catalog/metadata-generation-models.js";
import { metadataGenerationModelRoutes } from "./routes/metadata-generation-models.js";
import { catalogDescriptionConsolidationRoutes } from "./routes/catalog-description-consolidation.js";
import { PythonModelCompleter, type ModelCompleter } from "./catalog/model-completer.js";
import { DescriptionGenerationWorker } from "./catalog/description-generation-worker.js";
import { SensitivityAnalysisService } from "./catalog/sensitivity-analysis-service.js";
import { SensitivityAnalysisRunner } from "./catalog/sensitivity-analysis-runner.js";
import { SensitivityClassifier, type LocalNerDetector, type SensitivityValueSource } from "./catalog/sensitivity-classifier.js";
import { ConcreteSensitivityValueSource } from "./catalog/sensitivity-value-source.js";
import { PythonLocalNerDetector } from "./catalog/local-ner-detector.js";
import {
ConcreteDescriptionSourceSampler,
type DescriptionSourceSampler,
} from "./catalog/description-source-sampler.js";
import { catalogDescriptionGenerationRoutes } from "./routes/catalog-description-generation.js";
import { CatalogLogicalRelationshipService } from "./catalog/logical-relationship-service.js";
import { catalogLogicalRelationshipRoutes } from "./routes/catalog-logical-relationships.js";
import { EffectiveRelationshipSnapshotProvider } from "./catalog/effective-relationship-snapshot.js";
import { loadRuntimeModelCatalog, type RuntimeModelCatalog } from "./models/runtime-model-catalog.js";
import { createProductionWorkspacePreprocessingService } from "./workspace-maintenance.js";
import type { WorkspacePreprocessingService } from "./workspaces/preprocessing-service.js";
import { PreprocessingStateStore } from "./workspaces/preprocessing-state.js";
import { workspacePreprocessingRoutes } from "./routes/workspace-preprocessing.js";
import { memoryRoutes } from "./routes/memory.js";
import { evidenceRoutes } from "./routes/evidence.js";
export interface BuildAppDeps {
thtRunner?: ThtRunner;
@@ -48,7 +94,24 @@ export interface BuildAppDeps {
hub?: SseHub;
workspaceRegistry?: WorkspaceRegistry;
workspaceDiagnoser?: WorkspaceDiagnoser;
workspaceDatabaseTester?: WorkspaceDatabaseTester;
workspaceSecretStore?: WorkspaceSecretStore;
workspacePreprocessingService?: Pick<WorkspacePreprocessingService, "run" | "clear"> & Partial<Pick<WorkspacePreprocessingService, "consolidateEvidence" | "evidenceSources">>;
catalogRepository?: CatalogRepository;
catalogService?: CatalogService;
catalogPostgresAccess?: CatalogPostgresAccess;
catalogTableService?: CatalogTableService;
catalogLogicalRelationshipService?: CatalogLogicalRelationshipService;
effectiveRelationshipSnapshotProvider?: EffectiveRelationshipSnapshotProvider;
catalogSchemaIntrospector?: CatalogSchemaIntrospector;
catalogSyncWorker?: CatalogSyncWorker;
catalogOperationCoordinator?: CatalogOperationCoordinator;
metadataGenerationModels?: MetadataGenerationModels;
runtimeModelCatalog?: RuntimeModelCatalog;
modelCompleter?: ModelCompleter;
descriptionSourceSampler?: DescriptionSourceSampler;
sensitivityValueSource?: SensitivityValueSource;
localNerDetector?: LocalNerDetector;
workspaceRuntimeSupport?: (workspace: WorkspaceDescriptor) => boolean;
maintenanceBarrier?: MaintenanceBarrier;
piManagement?: PiManagementService;
@@ -101,6 +164,8 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc
app.register(cookie);
app.register(rateLimit, { global: false });
const workspaceRegistry = deps?.workspaceRegistry ?? new WorkspaceRegistry(config.workspaceRegistry);
const catalogRepository = deps?.catalogRepository ?? createCatalogRepository(config.catalogDatabase);
const tht = deps?.thtRunner ?? new ThtRunner({
thtBin: config.thtBin,
harnessDir: config.harnessDir,
@@ -111,20 +176,125 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc
secretsFile: config.secretsFile,
secretFiles: config.secretFiles,
workspaceSecretStore,
catalogRepository: deps?.catalogRepository ?? (config.catalogDatabase ? catalogRepository : undefined),
semanticRuntime: {
internalQdrantUrl: config.internalQdrantUrl,
internalEmbeddingUrl: config.internalEmbeddingUrl,
internalEmbeddingId: config.internalEmbeddingId,
internalEmbeddingModel: config.internalEmbeddingModel,
internalEmbeddingDimensions: config.internalEmbeddingDimensions,
},
});
const mgr = deps?.mgr ?? new PiProcessManager(config, deps?.spawnFn ? { spawnFn: deps.spawnFn } : undefined);
const workspacePreprocessingService = deps?.workspacePreprocessingService
?? createProductionWorkspacePreprocessingService({
config,
catalogRepository,
registry: workspaceRegistry,
workspaceSecretStore,
runner: tht as ThtRunner,
});
const hub = deps?.hub ?? new SseHub();
const workspaceRegistry = deps?.workspaceRegistry ?? new WorkspaceRegistry(config.workspaceRegistry);
const catalogOperationCoordinator = deps?.catalogOperationCoordinator ?? new CatalogOperationCoordinator();
const runtimeModelCatalog = deps?.runtimeModelCatalog ?? loadRuntimeModelCatalog(config.modelCatalogFile);
const mgr = deps?.mgr ?? new PiProcessManager(config, {
...(deps?.spawnFn ? { spawnFn: deps.spawnFn } : {}),
modelCatalog: runtimeModelCatalog,
});
const metadataGenerationModels = deps?.metadataGenerationModels ?? loadMetadataGenerationModels({
catalogFile: config.modelCatalogFile,
secretsFile: config.secretsFile,
});
const modelCompleter = deps?.modelCompleter ?? new PythonModelCompleter({
pythonExecutable: isAbsolute(config.thtBin) ? join(dirname(config.thtBin), "python") : "python3",
cwd: config.harnessDir,
});
const catalogPostgresAccess = deps?.catalogPostgresAccess ?? new ConcreteCatalogPostgresAccess(
workspaceSecretStore,
{ connectTimeoutMs: config.workspaceDiagnosticTimeoutMs },
);
const descriptionSourceSampler = deps?.descriptionSourceSampler
?? new ConcreteDescriptionSourceSampler(catalogPostgresAccess, workspaceSecretStore);
const descriptionGenerationWorker = new DescriptionGenerationWorker(
catalogRepository,
workspaceRegistry,
metadataGenerationModels,
modelCompleter,
catalogOperationCoordinator,
descriptionSourceSampler,
);
const sensitivityValueSource = deps?.sensitivityValueSource
?? new ConcreteSensitivityValueSource(catalogPostgresAccess, workspaceSecretStore);
const configuredNerWorker = config.sensitivityNer?.workerScript
?? fileURLToPath(new URL("../python/sensitivity_ner_worker.py", import.meta.url));
const localNerDetector = deps?.localNerDetector ?? (config.sensitivityNer
? new PythonLocalNerDetector({
pythonExecutable: config.sensitivityNer.pythonExecutable,
workerScript: configuredNerWorker,
modelPath: config.sensitivityNer.modelPath,
cwd: dirname(configuredNerWorker),
threads: config.sensitivityNer.threads,
})
: undefined);
const sensitiveDataSuggester = new SensitivityAnalysisService(
catalogRepository,
new SensitivityClassifier(sensitivityValueSource, localNerDetector),
);
const sensitivityAnalysisRunner = new SensitivityAnalysisRunner(
catalogRepository,
sensitiveDataSuggester,
);
const catalogService = deps?.catalogService ?? new CatalogService(
catalogRepository,
workspaceRegistry,
workspaceSecretStore,
config.workspaceRegistry.secretRoots,
config.workspaceDiagnosticTimeoutMs,
catalogPostgresAccess,
catalogOperationCoordinator,
);
const workspaceDatabaseTester = deps?.workspaceDatabaseTester ?? (async (workspaceId: string) => {
const database = await catalogRepository.getByWorkspace(workspaceId);
return database ? catalogService.test(database) : undefined;
});
const catalogTableService = deps?.catalogTableService ?? new CatalogTableService(catalogRepository);
const catalogLogicalRelationshipService = deps?.catalogLogicalRelationshipService
?? new CatalogLogicalRelationshipService(catalogRepository);
const catalogSchemaIntrospector = deps?.catalogSchemaIntrospector ?? new ConcreteCatalogSchemaIntrospector(
catalogPostgresAccess,
workspaceSecretStore,
);
const catalogSyncWorker = deps?.catalogSyncWorker ?? new CatalogSyncWorker(
catalogRepository,
catalogSchemaIntrospector,
catalogOperationCoordinator,
config.catalogSyncTimeoutMs,
createMemoryCleanup(tht as ThtRunner, {
internalQdrantUrl: config.internalQdrantUrl, internalEmbeddingUrl: config.internalEmbeddingUrl,
internalEmbeddingId: config.internalEmbeddingId, internalEmbeddingModel: config.internalEmbeddingModel,
internalEmbeddingDimensions: config.internalEmbeddingDimensions,
}),
);
app.addHook("onReady", async () => { await catalogSyncWorker.initialize(); });
app.addHook("onReady", async () => { await descriptionGenerationWorker.initialize(); });
app.addHook("onReady", async () => { await sensitivityAnalysisRunner.initialize(); });
if (localNerDetector?.warmup) {
app.addHook("onReady", async () => {
void localNerDetector.warmup?.().catch(() => undefined);
});
}
if (!deps?.catalogRepository && catalogRepository.close) {
app.addHook("onClose", async () => { await catalogRepository.close?.(); });
}
app.addHook("onClose", async () => { await catalogSyncWorker.stop(); });
app.addHook("onClose", async () => { await descriptionGenerationWorker.stop(); });
if (localNerDetector?.close) {
app.addHook("onClose", async () => { await localNerDetector.close?.(); });
}
const workspaceDiagnoser = deps?.workspaceDiagnoser
?? createProductionWorkspaceDiagnoser(config.workspaceDiagnosticTimeoutMs, undefined, {
internalQdrantUrl: config.internalQdrantUrl,
internalEmbeddingUrl: config.internalEmbeddingUrl,
internalEmbeddingId: config.internalEmbeddingId,
internalEmbeddingModel: config.internalEmbeddingModel,
internalEmbeddingDimensions: config.internalEmbeddingDimensions,
});
@@ -147,6 +317,7 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc
);
const listModels = deps?.listModels ?? createPiModelLister(config, {
modelCatalog: runtimeModelCatalog,
warn: (detail) => app.log.warn(
{ component: "pi-model-list", detail },
"Pi enabled-model configuration warning",
@@ -159,7 +330,7 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc
const getSettings = async (principal: PrincipalContext): Promise<Settings> => {
if (deps?.getSettings) return await deps.getSettings(principal);
const stored = loadSettings(config);
const effective = effectiveSettings(config, stored);
const effective = effectiveSettings(config, stored, runtimeModelCatalog);
// In the registry system the legacy `harness/workspaces/*.yaml` default is obsolete: when no
// installation workspace is pinned, default to the first active registry workspace.
if (!stored.workspace) {
@@ -172,7 +343,9 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc
}
return effective;
};
const piManagement = deps?.piManagement ?? createPiManagement(config, { listModels });
const piManagement = deps?.piManagement ?? createPiManagement(config, {
modelCatalog: runtimeModelCatalog,
});
const maintenanceBarrier = deps?.maintenanceBarrier ?? new MaintenanceBarrier(config.maintenanceFile);
const localRegistryResolver = deps?.localUserRegistry === undefined
@@ -296,7 +469,9 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc
dwhPrecheck: config.dwhPrecheck,
legacyWorkspaceMode: config.legacyWorkspaceMode,
workspaceRuntimeSupport,
modelCatalog: runtimeModelCatalog,
maintenanceBarrier,
catalogRepository: deps?.catalogRepository ?? (config.catalogDatabase ? catalogRepository : undefined),
});
app.post("/internal/maintenance/activate", async (req, reply) => {
try {
@@ -326,15 +501,60 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc
return maintenanceBarrier.status();
});
sqlRoutes(app, { tht: tht as ThtRunner, getSettings, workspaceRegistry });
metaRoutes(app, { harnessDir: config.harnessDir, listModels });
memoryRoutes(app, { runner: tht as ThtRunner, registry: workspaceRegistry, runtime: {
internalQdrantUrl: config.internalQdrantUrl,
internalEmbeddingUrl: config.internalEmbeddingUrl,
internalEmbeddingModel: config.internalEmbeddingModel,
internalEmbeddingDimensions: config.internalEmbeddingDimensions,
} });
metaRoutes(app, { harnessDir: config.harnessDir, modelCatalog: runtimeModelCatalog });
evidenceRoutes(app, { runner: tht as ThtRunner, registry: workspaceRegistry,
registryRoot: config.workspaceRegistry.root, hostRegistryRoot: config.evidenceHostRegistryRoot,
service: workspacePreprocessingService });
workspaceRoutes(app, {
registry: workspaceRegistry,
config: config.workspaceRegistry,
diagnose: workspaceDiagnoser,
authDiagnoser,
secretStore: workspaceSecretStore,
testDatabaseConnection: workspaceDatabaseTester,
});
settingsRoutes(app, { cfg: config, listModels, getSettings });
workspacePreprocessingRoutes(app, {
repository: catalogRepository,
registry: workspaceRegistry,
service: workspacePreprocessingService,
inputFingerprint: tht as ThtRunner,
readLatestJob: (workspaceId) => new PreprocessingStateStore({
dataRoot: config.dataRoot ?? "/data",
workspaceId,
}).readLatestJob(),
});
catalogDatabaseRoutes(app, { repository: catalogRepository, service: catalogService, operations: catalogOperationCoordinator });
catalogTableRoutes(app, {
repository: catalogRepository,
service: catalogTableService,
operations: catalogOperationCoordinator,
});
catalogSchemaRoutes(app, {
repository: catalogRepository,
worker: catalogSyncWorker,
operations: catalogOperationCoordinator,
});
catalogLogicalRelationshipRoutes(app, {
service: catalogLogicalRelationshipService,
operations: catalogOperationCoordinator,
});
catalogDescriptionConsolidationRoutes(app, {
repository: catalogRepository,
operations: catalogOperationCoordinator,
});
metadataGenerationModelRoutes(app, metadataGenerationModels);
catalogDescriptionGenerationRoutes(app, {
repository: catalogRepository,
worker: descriptionGenerationWorker,
sensitivityAnalysisRunner,
});
settingsRoutes(app, { cfg: config, getSettings });
piManagementRoutes(app, { service: piManagement });
return app;
+3 -1
View File
@@ -119,7 +119,9 @@ export function authenticateSession(deps: AuthDependencies): preHandlerHookHandl
subject: session.subject,
...(session.displayName === undefined ? {} : { displayName: session.displayName }),
roles: session.roles,
permissions: session.permissions,
// Sessions can outlive a deployment that changes the role permission catalog.
// resolve() has already checked validity, including current local user roles.
permissions: rolesToPermissions(session.roles),
isAdmin: session.roles.includes("admin"),
};
if (STATE_CHANGING_METHODS.has(request.method)) {
+1 -1
View File
@@ -38,7 +38,7 @@ const MAX_MAPPED_GROUPS = 128;
const ROLES = ["user", "admin"] as const;
export const PERMISSION_CATALOG: readonly Permission[] = [
"session.use", "session.read_all", "session.manage_all", "settings.manage",
"workspace.manage", "workspace.secrets.manage", "pi.manage", "auth.diagnostics.read",
"workspace.manage", "workspace.secrets.manage", "database.manage", "memory.manage", "evidence.manage", "pi.manage", "auth.diagnostics.read",
];
const invalid = (): Error => new Error("authentication configuration is invalid");
+1 -1
View File
@@ -581,7 +581,7 @@ function load(root: string): LoadedAuthConfig {
);
return {
value: selectedGeneration.value,
revision: `sha256:${selected.generation}`,
revision: selected.generation,
sourcePath: join(generationsPath, selected.generation, "auth.yaml"),
runtimeProjection: snapshot(
selected.generation,
+1 -1
View File
@@ -33,7 +33,7 @@ const EMPTY_HKDF_SALT = Buffer.alloc(0);
const ROLES = ["user", "admin"] as const;
const PERMISSIONS = [
"session.use", "session.read_all", "session.manage_all", "settings.manage",
"workspace.manage", "workspace.secrets.manage", "pi.manage", "auth.diagnostics.read",
"workspace.manage", "workspace.secrets.manage", "database.manage", "memory.manage", "evidence.manage", "pi.manage", "auth.diagnostics.read",
] as const satisfies readonly Permission[];
const invalid = (): Error => new Error("auth_session_store_invalid");
+1 -1
View File
@@ -5,7 +5,7 @@ export type Role = "user" | "admin";
export type Permission =
| "session.use" | "session.read_all" | "session.manage_all"
| "settings.manage" | "workspace.manage" | "workspace.secrets.manage"
| "pi.manage" | "auth.diagnostics.read";
| "database.manage" | "memory.manage" | "evidence.manage" | "pi.manage" | "auth.diagnostics.read";
export interface AuthenticationSessionConfig {
regularTtlSeconds: number;
+14 -9
View File
@@ -3,6 +3,18 @@ export interface ConfiguredTransportUrlOptions {
originOnly?: boolean;
}
export function parseCredentialFreeHttpUrl(value: string): URL | undefined {
let url: URL;
try {
url = new URL(value);
} catch {
return undefined;
}
if (!["http:", "https:"].includes(url.protocol)
|| url.username || url.password || url.search || url.hash) return undefined;
return url;
}
function canonicalLoopbackAuthority(value: string): boolean {
const match = /^http:\/\/([^/?#]+)(?:[/?#]|$)/.exec(value);
if (!match) return false;
@@ -27,15 +39,8 @@ export function parseConfiguredTransportUrl(
value: string,
options: ConfiguredTransportUrlOptions,
): URL | undefined {
let url: URL;
try {
url = new URL(value);
} catch {
return undefined;
}
if (url.username || url.password || url.search || url.hash || (options.originOnly && url.pathname !== "/")) {
return undefined;
}
const url = parseCredentialFreeHttpUrl(value);
if (!url || (options.originOnly && url.pathname !== "/")) return undefined;
if (url.protocol === "https:") return url;
if (options.allowLoopbackHttp && url.protocol === "http:" && canonicalLoopbackAuthority(value)) return url;
return undefined;
+15 -2
View File
@@ -4,6 +4,15 @@ const GENERIC_MODEL_FAILURE =
"Model request failed. Check provider connectivity, then Resume the session.";
const SUBSCRIPTION_MODEL_FAILURE =
"The selected model is unavailable for the current subscription. Choose another model and start a new session.";
const PHASE_STARTED_NOTIFICATION_PREFIX = "__tht_phase_started__:";
function phaseStartedNotification(message: unknown): string | null {
if (typeof message !== "string" || !message.startsWith(PHASE_STARTED_NOTIFICATION_PREFIX)) {
return null;
}
const phase = message.slice(PHASE_STARTED_NOTIFICATION_PREFIX.length);
return /^F[1-8]$/.test(phase) ? phase : "";
}
function safeModelFailure(error: unknown): string {
const detail = typeof error === "string" ? error : "";
@@ -36,7 +45,7 @@ export type ClientEvent =
| { type: "activity_event"; activity: ToolActivity }
| { type: "usage"; usage: TokenUsage }
| { type: "info"; [k: string]: any }
| { type: "system_event"; event: string };
| { type: "system_event"; event: string; phase?: string };
export type TurnState = "idle" | "running" | "waiting" | "failed";
@@ -71,7 +80,11 @@ export class SessionBridge {
});
}
} else if (m.type === "extension_ui_request" && m.method === "notify") {
this.fan({ type: "info", level: m.notifyType ?? "info", text: m.message ?? "" });
const phase = phaseStartedNotification(m.message);
if (phase) this.fan({ type: "system_event", event: "phase_started", phase });
else if (phase === null) {
this.fan({ type: "info", level: m.notifyType ?? "info", text: m.message ?? "" });
}
} else if (m.type === "message_update" && m.assistantMessageEvent?.type === "text_delta") {
this.fan({ type: "text_delta", text: m.assistantMessageEvent.delta ?? "" });
} else if (m.type === "message_update" && m.assistantMessageEvent?.type === "thinking_delta") {
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,252 @@
import { readFile } from "node:fs/promises";
import type { WorkspaceSecretStore } from "../workspaces/secret-store.js";
import type { CatalogPostgresAccess } from "./postgres-access.js";
import { CATALOG_SECRET_IDS } from "./secrets.js";
import { CatalogConnectorError, type WorkspaceDatabase } from "./types.js";
const MAX_SOURCE_ROWS = 5;
const MAX_REPRESENTATIVE_VALUES = 5;
const MAX_SOURCE_COLUMNS_PER_TARGET = 8;
const MAX_SOURCE_VALUE_BYTES = 256;
export type DescriptionSourceSampleValue = string | number | boolean | null;
export interface DescriptionSourceSampleField {
name: string;
value: DescriptionSourceSampleValue;
}
export interface DescriptionSourceSampleRow {
fields: readonly DescriptionSourceSampleField[];
}
export interface DescriptionSourceRepresentativeValues {
column: string;
values: readonly Exclude<DescriptionSourceSampleValue, null>[];
}
export interface DescriptionTargetSourceSample {
targetId: string;
tableName: string;
rows: readonly DescriptionSourceSampleRow[];
representativeValues: readonly DescriptionSourceRepresentativeValues[];
}
export interface DescriptionSourceSamplingTarget {
targetId: string;
tableName: string;
columnNames: readonly string[];
}
/** Optional, transient source context for one model-completion batch. */
export interface DescriptionSourceSampler {
sample(
database: WorkspaceDatabase,
targets: readonly DescriptionSourceSamplingTarget[],
signal: AbortSignal,
): Promise<readonly DescriptionTargetSourceSample[]>;
}
function quoteIdentifier(identifier: string): string {
return `"${identifier.replaceAll('"', '""')}"`;
}
function boundedUtf8(value: string, maxBytes: number): string {
const normalized = value
.normalize("NFC")
.replace(/\r\n?/g, "\n")
.replace(/[\u0000-\u0008\u000b\u000c\u000e-\u001f\u007f]/g, " ");
if (Buffer.byteLength(normalized, "utf8") <= maxBytes) return normalized;
let result = "";
let bytes = 0;
for (const character of normalized) {
const characterBytes = Buffer.byteLength(character, "utf8");
if (bytes + characterBytes > maxBytes) break;
result += character;
bytes += characterBytes;
}
return result;
}
function normalizeValue(value: unknown): DescriptionSourceSampleValue | undefined {
if (value === null) return null;
if (typeof value === "string") return boundedUtf8(value, MAX_SOURCE_VALUE_BYTES);
if (typeof value === "boolean") return value;
if (typeof value === "number") return Number.isFinite(value) ? value : undefined;
if (typeof value === "bigint") return boundedUtf8(String(value), MAX_SOURCE_VALUE_BYTES);
if (value instanceof Date && !Number.isNaN(value.valueOf())) return value.toISOString();
return undefined;
}
function distinctKey(value: Exclude<DescriptionSourceSampleValue, null>): string {
return `${typeof value}:${String(value)}`;
}
function columnsFor(target: DescriptionSourceSamplingTarget): string[] {
return [...new Set(target.columnNames)].slice(0, MAX_SOURCE_COLUMNS_PER_TARGET);
}
function normalizedSample(
target: DescriptionSourceSamplingTarget,
columnNames: readonly string[],
sourceRows: readonly Record<string, unknown>[],
): DescriptionTargetSourceSample {
const rows = sourceRows.slice(0, MAX_SOURCE_ROWS).map((row) => ({
fields: columnNames.flatMap((name) => {
const value = normalizeValue(row[name]);
return value === undefined ? [] : [{ name, value }];
}),
}));
const valuesByColumn = new Map<string, Exclude<DescriptionSourceSampleValue, null>[]>();
const seenByColumn = new Map<string, Set<string>>();
let representativeValueCount = 0;
for (const row of rows) {
for (const field of row.fields) {
if (representativeValueCount === MAX_REPRESENTATIVE_VALUES) break;
if (field.value === null) continue;
const seen = seenByColumn.get(field.name) ?? new Set<string>();
const key = distinctKey(field.value);
if (seen.has(key)) continue;
seen.add(key);
seenByColumn.set(field.name, seen);
const values = valuesByColumn.get(field.name) ?? [];
values.push(field.value);
valuesByColumn.set(field.name, values);
representativeValueCount += 1;
}
if (representativeValueCount === MAX_REPRESENTATIVE_VALUES) break;
}
return {
targetId: target.targetId,
tableName: target.tableName,
rows,
representativeValues: columnNames.flatMap((column) => {
const values = valuesByColumn.get(column);
return values && values.length > 0 ? [{ column, values }] : [];
}),
};
}
function samplingSql(
database: WorkspaceDatabase,
target: DescriptionSourceSamplingTarget,
columns: readonly string[],
): string {
const projections = columns.map((columnName) => {
const identifier = quoteIdentifier(columnName);
return `LEFT((${identifier})::text, ${MAX_SOURCE_VALUE_BYTES}) AS ${identifier}`;
});
return [
`SELECT ${projections.join(", ")}`,
`FROM ${quoteIdentifier(database.schema)}.${quoteIdentifier(target.tableName)}`,
`LIMIT ${MAX_SOURCE_ROWS}`,
].join(" ");
}
/** Bounded source sampler that follows the database's PostgreSQL-wire or REST binding. */
export class ConcreteDescriptionSourceSampler implements DescriptionSourceSampler {
constructor(
private readonly access: CatalogPostgresAccess,
private readonly secretStore?: Pick<WorkspaceSecretStore, "materialize">,
) {}
async sample(
database: WorkspaceDatabase,
targets: readonly DescriptionSourceSamplingTarget[],
signal: AbortSignal,
): Promise<readonly DescriptionTargetSourceSample[]> {
if (database.binding.transport === "rest_api") {
return await this.sampleRest(database, targets, signal);
}
const client = await this.access.connect(database, signal);
let transactionOpen = false;
try {
await client.query("BEGIN TRANSACTION READ ONLY", []);
transactionOpen = true;
const samples: DescriptionTargetSourceSample[] = [];
for (const target of targets) {
const columnNames = columnsFor(target);
if (columnNames.length === 0) {
samples.push({
targetId: target.targetId,
tableName: target.tableName,
rows: [],
representativeValues: [],
});
continue;
}
const projections = columnNames.map((columnName) => {
const identifier = quoteIdentifier(columnName);
return `LEFT((${identifier})::text, $1) AS ${identifier}`;
});
const sql = [
`SELECT ${projections.join(", ")}`,
`FROM ${quoteIdentifier(database.schema)}.${quoteIdentifier(target.tableName)}`,
"LIMIT $2",
].join(" ");
const result = await client.query(sql, [MAX_SOURCE_VALUE_BYTES, MAX_SOURCE_ROWS]);
samples.push(normalizedSample(target, columnNames, result.rows));
}
return samples;
} finally {
if (transactionOpen) await client.query("ROLLBACK", []).catch(() => undefined);
await client.end().catch(() => undefined);
}
}
private async sampleRest(
database: WorkspaceDatabase,
targets: readonly DescriptionSourceSamplingTarget[],
signal: AbortSignal,
): Promise<readonly DescriptionTargetSourceSample[]> {
if (!this.secretStore) throw new CatalogConnectorError("REST source sampling is not configured");
const auth = database.binding.restAuth ?? "bearer";
const materialized = this.secretStore.materialize(
database.workspaceId,
auth === "none" ? [] : [CATALOG_SECRET_IDS.apiKey],
);
try {
const headers: Record<string, string> = { "content-type": "application/json" };
if (auth !== "none") {
const credentialFile = materialized.files.get(CATALOG_SECRET_IDS.apiKey);
if (!credentialFile) throw new CatalogConnectorError("REST API key is not configured");
const credential = (await readFile(credentialFile, "utf8")).trim();
if (auth === "bearer") headers.authorization = `Bearer ${credential}`;
else headers["x-api-key"] = credential;
}
const baseUrl = database.binding.baseUrl?.replace(/\/+$/, "");
if (!baseUrl) throw new CatalogConnectorError("Database binding is incomplete");
const samples: DescriptionTargetSourceSample[] = [];
for (const target of targets) {
const columnNames = columnsFor(target);
if (columnNames.length === 0) {
samples.push(normalizedSample(target, columnNames, []));
continue;
}
const response = await fetch(`${baseUrl}/rpc/run_query`, {
method: "POST",
headers,
body: JSON.stringify({ query_text: samplingSql(database, target, columnNames) }),
signal,
});
if (!response.ok) throw new CatalogConnectorError("REST source sampling failed");
const body: unknown = await response.json();
if (!Array.isArray(body)
|| body.some((row) => !row || typeof row !== "object" || Array.isArray(row))) {
throw new CatalogConnectorError("REST source sampling response is invalid");
}
samples.push(normalizedSample(
target,
columnNames,
body as Array<Record<string, unknown>>,
));
}
return samples;
} catch (error) {
if (error instanceof CatalogConnectorError) throw error;
throw new CatalogConnectorError("REST source sampling failed");
} finally {
materialized.release();
}
}
}
@@ -0,0 +1,93 @@
import type { CatalogRelationship, CatalogRepository } from "./types.js";
export interface EffectiveRelationshipSnapshotReader {
list(databaseId: string): Promise<CatalogRelationship[]>;
}
export interface EffectiveRelationshipSnapshotCoordinator {
run<T>(databaseId: string, operation: () => Promise<T>): Promise<T>;
}
export class EffectiveRelationshipSnapshotStaleError extends Error {}
interface EffectiveRelationship {
sourceTable: string;
sourceColumns: string[];
targetTable: string;
targetColumns: string[];
origin: CatalogRelationship["origin"];
}
interface EffectiveRelationshipSnapshot {
schemaVersion: 1;
workspaceId: string;
relationships: EffectiveRelationship[];
}
const originRank: Record<CatalogRelationship["origin"], number> = {
physical: 0,
manual: 1,
generated: 2,
};
function endpointKey(relationship: EffectiveRelationship): string {
return [
relationship.sourceTable,
relationship.sourceColumns.join("\u0000"),
relationship.targetTable,
relationship.targetColumns.join("\u0000"),
].join("\u0001");
}
function compareRelationships(left: EffectiveRelationship, right: EffectiveRelationship): number {
return endpointKey(left).localeCompare(endpointKey(right))
|| originRank[left.origin] - originRank[right.origin];
}
/**
* Adapter from the mutable Catalog model to the immutable relationship contract consumed by the
* harness. The returned JSON is a deterministic projection, never an authored second store.
*/
export class EffectiveRelationshipSnapshotProvider {
constructor(
private readonly repository: Pick<CatalogRepository, "get" | "getByWorkspace">,
private readonly relationships: EffectiveRelationshipSnapshotReader,
private readonly operations: EffectiveRelationshipSnapshotCoordinator,
) {}
async render(workspaceId: string): Promise<string | undefined> {
const database = await this.repository.getByWorkspace(workspaceId);
if (!database) return undefined;
return await this.operations.run(database.id, async () => {
const current = await this.repository.get(database.id);
if (!current || current.schemaSyncedVersion !== current.version) {
throw new EffectiveRelationshipSnapshotStaleError(
"effective relationship snapshot requires a current full schema synchronization",
);
}
const projected = (await this.relationships.list(database.id))
.filter((relationship) => relationship.status === "active")
.map((relationship): EffectiveRelationship => ({
sourceTable: relationship.sourceTableName,
sourceColumns: relationship.columns.map((column) => column.sourceColumnName),
targetTable: relationship.targetTableName,
targetColumns: relationship.columns.map((column) => column.targetColumnName),
origin: relationship.origin,
}))
.sort(compareRelationships);
const seen = new Set<string>();
const snapshot: EffectiveRelationshipSnapshot = {
schemaVersion: 1,
workspaceId,
relationships: projected.filter((relationship) => {
const key = endpointKey(relationship);
if (seen.has(key)) return false;
seen.add(key);
return true;
}),
};
return `${JSON.stringify(snapshot, null, 2)}\n`;
});
}
}
+254
View File
@@ -0,0 +1,254 @@
import { randomUUID } from "node:crypto";
import { spawn, type ChildProcessWithoutNullStreams } from "node:child_process";
import { tmpdir } from "node:os";
import { z } from "zod";
import type {
LocalNerCandidate,
LocalNerDetector,
LocalNerEvidence,
} from "./sensitivity-classifier.js";
const MAX_LINE_BYTES = 64 * 1024;
const candidateSchema = z.object({
columnId: z.uuid(),
text: z.string().min(1).max(500),
}).strict();
const workerMessageSchema = z.union([
z.object({ ready: z.literal(true) }).strict(),
z.object({
id: z.uuid(),
ok: z.literal(true),
evidence: z.array(z.object({
columnId: z.uuid(),
label: z.string().min(1).max(80),
confidence: z.number().min(0).max(1),
}).strict()).max(1_000),
}).strict(),
z.object({ id: z.uuid(), ok: z.literal(false), error: z.string().min(1).max(80) }).strict(),
]);
export class LocalNerUnavailableError extends Error {
constructor() {
super("local NER is unavailable");
this.name = "LocalNerUnavailableError";
}
}
interface PendingRequest {
resolve: (value: readonly LocalNerEvidence[]) => void;
reject: (error: Error) => void;
timer: ReturnType<typeof setTimeout>;
signal: AbortSignal;
cancel: () => void;
}
/** Persistent JSONL adapter for the optional, CPU-only Python NER worker. */
export class PythonLocalNerDetector implements LocalNerDetector {
private child?: ChildProcessWithoutNullStreams;
private ready?: Promise<void>;
private readyResolve?: () => void;
private readyReject?: (error: Error) => void;
private workerReady = false;
private stdout = "";
private readonly pending = new Map<string, PendingRequest>();
constructor(private readonly options: {
pythonExecutable: string;
workerScript: string;
modelPath: string;
cwd: string;
threads?: number;
startupTimeoutMs?: number;
}) {}
async warmup(): Promise<void> {
await this.ensureStarted();
}
isReady(): boolean {
return this.workerReady
&& this.child !== undefined
&& this.child.exitCode === null
&& this.child.signalCode === null;
}
async detect(
candidates: readonly LocalNerCandidate[],
signal: AbortSignal,
deadline: number,
): Promise<readonly LocalNerEvidence[]> {
const parsed = z.array(candidateSchema).min(1).max(128).parse(candidates);
if (signal.aborted || deadline <= Date.now()) throw new LocalNerUnavailableError();
await this.ensureStartedWithin(signal, deadline);
if (!this.child || this.child.exitCode !== null || this.child.signalCode !== null) {
throw new LocalNerUnavailableError();
}
const id = randomUUID();
return await new Promise<readonly LocalNerEvidence[]>((resolve, reject) => {
const fail = () => {
this.finishPending(id);
reject(new LocalNerUnavailableError());
this.stopWorker();
};
const timer = setTimeout(fail, Math.max(1, Math.floor(deadline - Date.now())));
const cancel = fail;
const pending: PendingRequest = { resolve, reject, timer, signal, cancel };
this.pending.set(id, pending);
signal.addEventListener("abort", cancel, { once: true });
this.child!.stdin.write(`${JSON.stringify({ id, candidates: parsed })}\n`, (error) => {
if (error) fail();
});
});
}
async close(): Promise<void> {
const child = this.child;
if (!child || child.exitCode !== null || child.signalCode !== null) return;
await new Promise<void>((resolve) => {
child.once("close", () => resolve());
child.kill("SIGTERM");
setTimeout(() => {
if (child.exitCode === null && child.signalCode === null) child.kill("SIGKILL");
}, 250).unref();
});
}
private async ensureStarted(): Promise<void> {
if (this.ready) return await this.ready;
this.ready = new Promise<void>((resolve, reject) => {
this.readyResolve = resolve;
this.readyReject = reject;
});
const threads = String(this.options.threads ?? 2);
const inheritedRuntimeEnvironment = Object.fromEntries([
"PATH", "SystemRoot", "WINDIR", "PATHEXT", "TMPDIR", "TEMP", "TMP", "LANG", "LC_ALL",
].flatMap((name) => process.env[name] === undefined ? [] : [[name, process.env[name]!]]));
const child = spawn(this.options.pythonExecutable, [
"-I",
"-B",
this.options.workerScript,
"--model",
this.options.modelPath,
"--threads",
threads,
], {
cwd: this.options.cwd,
stdio: ["pipe", "pipe", "pipe"],
env: {
...inheritedRuntimeEnvironment,
HOME: process.env.HOME ?? tmpdir(),
CUDA_VISIBLE_DEVICES: "",
HIP_VISIBLE_DEVICES: "",
HF_HUB_OFFLINE: "1",
HF_HUB_DISABLE_TELEMETRY: "1",
TRANSFORMERS_OFFLINE: "1",
TOKENIZERS_PARALLELISM: "false",
PYTHONNOUSERSITE: "1",
OMP_NUM_THREADS: threads,
MKL_NUM_THREADS: threads,
OPENBLAS_NUM_THREADS: threads,
HTTP_PROXY: "",
HTTPS_PROXY: "",
ALL_PROXY: "",
NO_PROXY: "*",
},
});
this.child = child;
child.stdout.setEncoding("utf8");
child.stdout.on("data", (chunk: string) => this.receive(chunk));
child.stderr.resume();
child.once("error", () => this.failWorker());
child.once("close", () => this.failWorker());
const startupTimer = setTimeout(() => this.failWorker(), this.options.startupTimeoutMs ?? 120_000);
startupTimer.unref();
try {
await this.ready;
} finally {
clearTimeout(startupTimer);
}
}
private async ensureStartedWithin(signal: AbortSignal, deadline: number): Promise<void> {
const started = this.ensureStarted();
await new Promise<void>((resolve, reject) => {
let settled = false;
const finish = (error?: Error, stopWorker = false) => {
if (settled) return;
settled = true;
clearTimeout(timer);
signal.removeEventListener("abort", cancel);
if (stopWorker) this.failWorker();
if (error) reject(error);
else resolve();
};
const cancel = () => finish(new LocalNerUnavailableError(), true);
const timer = setTimeout(cancel, Math.max(1, Math.floor(deadline - Date.now())));
signal.addEventListener("abort", cancel, { once: true });
void started.then(
() => finish(),
() => finish(new LocalNerUnavailableError()),
);
});
}
private receive(chunk: string): void {
this.stdout += chunk;
if (Buffer.byteLength(this.stdout, "utf8") > MAX_LINE_BYTES) {
this.failWorker();
return;
}
let newline: number;
while ((newline = this.stdout.indexOf("\n")) >= 0) {
const line = this.stdout.slice(0, newline);
this.stdout = this.stdout.slice(newline + 1);
if (!line) continue;
try {
const message = workerMessageSchema.parse(JSON.parse(line));
if ("ready" in message) {
this.workerReady = true;
this.readyResolve?.();
this.readyResolve = undefined;
this.readyReject = undefined;
continue;
}
const pending = this.pending.get(message.id);
if (!pending) continue;
this.finishPending(message.id);
if (message.ok) pending.resolve(message.evidence);
else pending.reject(new LocalNerUnavailableError());
} catch {
this.failWorker();
return;
}
}
}
private finishPending(id: string): void {
const pending = this.pending.get(id);
if (!pending) return;
clearTimeout(pending.timer);
pending.signal.removeEventListener("abort", pending.cancel);
this.pending.delete(id);
}
private stopWorker(): void {
const child = this.child;
if (child && child.exitCode === null && child.signalCode === null) child.kill("SIGTERM");
}
private failWorker(): void {
const error = new LocalNerUnavailableError();
this.readyReject?.(error);
this.readyResolve = undefined;
this.readyReject = undefined;
for (const [id, pending] of this.pending) {
this.finishPending(id);
pending.reject(error);
}
this.stopWorker();
this.child = undefined;
this.ready = undefined;
this.workerReady = false;
this.stdout = "";
}
}
@@ -0,0 +1,260 @@
import type {
CatalogLogicalRelationship,
CatalogLogicalRelationshipCandidate,
CatalogLogicalRelationshipContext,
CatalogLogicalRelationshipEndpoint,
CatalogRelationship,
CatalogRepository,
} from "./types.js";
export class LogicalRelationshipDatabaseNotFoundError extends Error {}
export class LogicalRelationshipDuplicateError extends Error {}
export class LogicalRelationshipNotFoundError extends Error {}
export class LogicalRelationshipReadOnlyError extends Error {}
export class LogicalRelationshipSchemaStaleError extends Error {}
export class LogicalRelationshipTargetNotUniqueError extends Error {}
export class LogicalRelationshipTypeIncompatibleError extends Error {}
export class LogicalRelationshipColumnNotFoundError extends Error {
constructor(readonly field: "sourceColumnId" | "targetColumnId") {
super(`Catalog column '${field}' was not found`);
}
}
export interface RebuildGeneratedRelationshipsResult {
added: number;
alreadyPresent: number;
excluded: number;
ambiguous: number;
}
function identifierTokens(value: string): string[] {
return value
.replace(/([a-z0-9])([A-Z])/g, "$1_$2")
.toLowerCase()
.split(/[^a-z0-9]+/)
.filter(Boolean);
}
function singularWord(value: string): string {
if (value.length > 4 && value.endsWith("ies")) return `${value.slice(0, -3)}y`;
if (value.length > 4 && /(ches|shes|xes|zes|ses)$/.test(value)) return value.slice(0, -2);
if (value.length > 3 && value.endsWith("s") && !/(ss|us)$/.test(value)) return value.slice(0, -1);
return value;
}
function tableAliases(tableName: string): string[] {
const tokens = identifierTokens(tableName);
if (tokens.length === 0) return [];
const normalized = tokens.join("_");
const singular = [...tokens];
singular[singular.length - 1] = singularWord(singular[singular.length - 1]);
return [...new Set([normalized, singular.join("_")])];
}
const GENERIC_PRIMARY_KEY_NAMES = new Set(["id", "key", "code", "pk"]);
function nameMatches(
source: CatalogLogicalRelationshipEndpoint,
target: CatalogLogicalRelationshipEndpoint,
): boolean {
const sourceName = identifierTokens(source.columnName).join("_");
const targetName = identifierTokens(target.columnName).join("_");
if (!sourceName || !targetName) return false;
const expected = new Set<string>();
if (!GENERIC_PRIMARY_KEY_NAMES.has(targetName)) expected.add(targetName);
for (const alias of tableAliases(target.tableName)) {
expected.add(`${alias}_${targetName}`);
expected.add(`${alias.replaceAll("_", "")}${targetName.replaceAll("_", "")}`);
if (targetName === "id" || targetName === "pk") expected.add(alias);
}
return expected.has(sourceName);
}
function canonicalDataType(value: string): string {
const normalized = value.trim().toLowerCase().replace(/\s+/g, " ");
const arraySuffix = normalized.endsWith("[]") ? "[]" : "";
const base = arraySuffix ? normalized.slice(0, -2) : normalized;
const withoutModifier = base.replace(/\([^)]*\)/g, "").trim();
const aliases: Record<string, string> = {
int2: "smallint",
smallserial: "smallint",
int4: "integer",
int: "integer",
serial: "integer",
int8: "bigint",
bigserial: "bigint",
decimal: "numeric",
varchar: "text",
"character varying": "text",
bool: "boolean",
"timestamp without time zone": "timestamp",
"timestamp with time zone": "timestamptz",
"time without time zone": "time",
"time with time zone": "timetz",
};
return `${aliases[withoutModifier] ?? withoutModifier}${arraySuffix}`;
}
function typesCompatible(left: string, right: string): boolean {
return canonicalDataType(left) === canonicalDataType(right);
}
function pairKey(sourceColumnId: string, targetColumnId: string): string {
return `${sourceColumnId}\u0000${targetColumnId}`;
}
function relationshipPair(relationship: CatalogLogicalRelationship): CatalogLogicalRelationshipCandidate {
return {
sourceColumnId: relationship.columns[0].sourceColumnId,
targetColumnId: relationship.columns[0].targetColumnId,
};
}
function relationshipSortKey(relationship: CatalogRelationship): string {
const sourceColumns = relationship.columns.map((column) => column.sourceColumnName).join(",");
const targetColumns = relationship.columns.map((column) => column.targetColumnName).join(",");
return [
relationship.sourceTableName,
sourceColumns,
relationship.targetTableName,
targetColumns,
relationship.origin,
].join("\u0000");
}
export class CatalogLogicalRelationshipService {
constructor(private readonly repository: CatalogRepository) {}
async list(databaseId: string): Promise<CatalogRelationship[]> {
if (!(await this.repository.get(databaseId))) throw new LogicalRelationshipDatabaseNotFoundError();
const relationships: CatalogRelationship[] = [
...await this.repository.listRelationships(databaseId),
...await this.repository.listLogicalRelationships(databaseId),
];
return relationships.sort((left, right) => relationshipSortKey(left).localeCompare(relationshipSortKey(right)));
}
async addManual(
databaseId: string,
sourceColumnId: string,
targetColumnId: string,
): Promise<CatalogLogicalRelationship> {
const context = await this.requiredContext(databaseId);
const source = context.endpoints.find((endpoint) => endpoint.columnId === sourceColumnId);
if (!source) throw new LogicalRelationshipColumnNotFoundError("sourceColumnId");
const target = context.endpoints.find((endpoint) => endpoint.columnId === targetColumnId);
if (!target) throw new LogicalRelationshipColumnNotFoundError("targetColumnId");
if (sourceColumnId === targetColumnId
|| target.primaryKeyPosition === null
|| target.tablePrimaryKeyColumnCount !== 1) {
throw new LogicalRelationshipTargetNotUniqueError();
}
if (!typesCompatible(source.dataType, target.dataType)) {
throw new LogicalRelationshipTypeIncompatibleError();
}
const key = pairKey(sourceColumnId, targetColumnId);
if (context.physicalPairs.some((pair) => pairKey(pair.sourceColumnId, pair.targetColumnId) === key)
|| context.logicalRelationships.some((relationship) => {
const pair = relationshipPair(relationship);
return pairKey(pair.sourceColumnId, pair.targetColumnId) === key;
})) throw new LogicalRelationshipDuplicateError();
const created = await this.repository.insertLogicalRelationship(
databaseId,
sourceColumnId,
targetColumnId,
false,
);
if (!created) throw new LogicalRelationshipDuplicateError();
return created;
}
async rebuildGenerated(databaseId: string): Promise<RebuildGeneratedRelationshipsResult> {
const context = await this.requiredContext(databaseId);
const targets = context.endpoints.filter((endpoint) => (
endpoint.primaryKeyPosition !== null && endpoint.tablePrimaryKeyColumnCount === 1
));
const physical = new Set(context.physicalPairs.map((pair) => pairKey(pair.sourceColumnId, pair.targetColumnId)));
const active = new Set<string>();
const excluded = new Set<string>();
for (const relationship of context.logicalRelationships) {
const pair = relationshipPair(relationship);
(relationship.status === "excluded" ? excluded : active)
.add(pairKey(pair.sourceColumnId, pair.targetColumnId));
}
const pending: CatalogLogicalRelationshipCandidate[] = [];
let alreadyPresent = 0;
let excludedCount = 0;
let ambiguous = 0;
const dimTimeTargets = targets.filter((target) => (
identifierTokens(target.tableName).join("_") === "dim_time"
));
const sources = context.endpoints.filter((endpoint) => (
endpoint.primaryKeyPosition === null || endpoint.tablePrimaryKeyColumnCount > 1
));
for (const source of sources) {
const sourceName = identifierTokens(source.columnName).join("_");
const isTimeKey = sourceName.endsWith("time_key")
&& identifierTokens(source.tableName).join("_") !== "dim_time";
const candidates = isTimeKey ? dimTimeTargets : targets;
const matches = candidates.filter((target) => (
target.columnId !== source.columnId
&& typesCompatible(source.dataType, target.dataType)
&& (isTimeKey || nameMatches(source, target))
));
if (matches.length > 1) {
ambiguous += 1;
continue;
}
if (matches.length === 0) continue;
const candidate = { sourceColumnId: source.columnId, targetColumnId: matches[0].columnId };
const key = pairKey(candidate.sourceColumnId, candidate.targetColumnId);
if (physical.has(key) || active.has(key)) {
alreadyPresent += 1;
} else if (excluded.has(key)) {
excludedCount += 1;
} else {
pending.push(candidate);
}
}
const added = await this.repository.insertGeneratedLogicalRelationships(databaseId, pending);
alreadyPresent += pending.length - added;
return { added, alreadyPresent, excluded: excludedCount, ambiguous };
}
async setStatus(
databaseId: string,
relationshipId: string,
status: CatalogLogicalRelationship["status"],
): Promise<CatalogLogicalRelationship> {
if (!(await this.repository.get(databaseId))) throw new LogicalRelationshipDatabaseNotFoundError();
const updated = await this.repository.setLogicalRelationshipStatus(databaseId, relationshipId, status);
if (updated) return updated;
await this.assertNotPhysical(databaseId, relationshipId);
throw new LogicalRelationshipNotFoundError();
}
async deletePermanently(databaseId: string, relationshipId: string): Promise<void> {
if (!(await this.repository.get(databaseId))) throw new LogicalRelationshipDatabaseNotFoundError();
if (await this.repository.deleteLogicalRelationship(databaseId, relationshipId)) return;
await this.assertNotPhysical(databaseId, relationshipId);
throw new LogicalRelationshipNotFoundError();
}
private async requiredContext(databaseId: string): Promise<CatalogLogicalRelationshipContext> {
const database = await this.repository.get(databaseId);
if (!database) throw new LogicalRelationshipDatabaseNotFoundError();
if (database.schemaSyncedVersion !== database.version) {
throw new LogicalRelationshipSchemaStaleError();
}
const context = await this.repository.getLogicalRelationshipContext(databaseId);
if (!context) throw new LogicalRelationshipDatabaseNotFoundError();
return context;
}
private async assertNotPhysical(databaseId: string, relationshipId: string): Promise<void> {
if ((await this.repository.listRelationships(databaseId)).some((relationship) => relationship.id === relationshipId)) {
throw new LogicalRelationshipReadOnlyError();
}
}
}
+23
View File
@@ -0,0 +1,23 @@
import type { ThtRunner } from "../tht/tht-runner.js";
import type { SemanticRuntimeConfig } from "../workspaces/runtime-renderer.js";
import type { CatalogSyncRun, WorkspaceDatabase } from "./types.js";
/** Internal continuation of an applied physical sync, using the harness Memory boundary. */
export function createMemoryCleanup(runner: Pick<ThtRunner, "withPrincipal">, runtime: SemanticRuntimeConfig) {
return async (database: WorkspaceDatabase, run: CatalogSyncRun): Promise<number> => {
if (run.phase !== "memory_cleanup" || !run.plannedDiff) throw new Error("Physical cleanup is not committed");
const result = await runner.withPrincipal({ issuer: "installation", subject: "catalog-sync",
roles: ["admin"], permissions: ["memory.manage"], isAdmin: true,
}).runWithRuntimeSnapshot(["memory", "admin", "--workspace", database.workspaceId], JSON.stringify({
action: "cleanup", runtime, request: { sync_id: run.id, database: database.databaseName,
schema_name: database.schema, removed_tables: run.plannedDiff.deletedTables,
removed_columns: run.plannedDiff.deletedColumns.map(column => ({ table: column.tableName, column: column.columnName })),
},
}));
const payload = JSON.parse(result.stdout);
if (result.code !== 0 || payload.indexed !== true || !Number.isInteger(payload.deleted) || payload.deleted < 0) {
throw new Error("Memory cleanup is incomplete");
}
return payload.deleted;
};
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,114 @@
import { loadSecretBundle } from "../config/secret-bundle.js";
import { loadRuntimeModelCatalog } from "../models/runtime-model-catalog.js";
export interface MetadataGenerationModelChoice {
id: string;
label: string;
}
export interface MetadataGenerationModelCatalog {
models: MetadataGenerationModelChoice[];
default: string | null;
}
export interface ResolvedMetadataGenerationModel {
readonly id: string;
readonly provider: string;
readonly model: string;
readonly disableThinking?: true;
readonly endpoint?: Readonly<{ baseUrl: string; apiVersion?: string }>;
readonly apiKeyEnv?: string;
readonly apiKey?: string;
}
export class MetadataGenerationModelUnavailableError extends Error {
constructor() {
super("metadata-generation model is unavailable");
this.name = "MetadataGenerationModelUnavailableError";
}
}
export interface MetadataGenerationModels {
catalog(): MetadataGenerationModelCatalog;
resolve(selection: string): ResolvedMetadataGenerationModel;
}
class RestartLoadedMetadataGenerationModels implements MetadataGenerationModels {
readonly #models: ReadonlyMap<string, ResolvedMetadataGenerationModel>;
readonly #catalog: MetadataGenerationModelCatalog;
constructor(models: ReadonlyMap<string, ResolvedMetadataGenerationModel>, defaultModel: string | null) {
this.#models = models;
this.#catalog = {
models: [...models.values()].map(({ id }) => ({ id, label: id })),
default: defaultModel,
};
}
catalog(): MetadataGenerationModelCatalog {
return { models: this.#catalog.models.map((choice) => ({ ...choice })), default: this.#catalog.default };
}
resolve(selection: string): ResolvedMetadataGenerationModel {
const model = this.#models.get(selection);
if (!model) throw new MetadataGenerationModelUnavailableError();
return model;
}
}
function invalid(message = "metadata-generation runtime catalog is invalid"): Error {
return new Error(message);
}
export function loadMetadataGenerationModels(options: {
catalogFile?: string;
secretsFile?: string;
}): MetadataGenerationModels {
const catalog = loadRuntimeModelCatalog(options.catalogFile);
const configured = catalog.metadataModels();
if (configured.length === 0) return new RestartLoadedMetadataGenerationModels(new Map(), null);
const requiresSecrets = configured.some((model) => model.authentication.mode === "secret_env");
let secrets: ReadonlyMap<string, string> = new Map();
if (requiresSecrets) {
if (!options.secretsFile) throw invalid("metadata-generation keyed models require THT_SECRETS_FILE");
try { secrets = loadSecretBundle(options.secretsFile); }
catch { throw invalid("metadata-generation secrets are unavailable"); }
}
const models = new Map<string, ResolvedMetadataGenerationModel>();
const labels = new Map<string, string>();
for (const configuredModel of configured) {
const adapter = configuredModel.metadataAdapter;
if (!adapter || configuredModel.authentication.mode === "pi_auth") throw invalid();
const apiKeyEnv = configuredModel.authentication.apiKeyEnv;
let apiKey: string | undefined;
if (configuredModel.authentication.mode === "secret_env") {
if (!apiKeyEnv) throw invalid();
apiKey = secrets.get(apiKeyEnv);
if (!apiKey) throw invalid(`metadata-generation model "${configuredModel.id}" secret "${apiKeyEnv}" is missing`);
if (apiKey.length > 16 * 1024 || /\s/u.test(apiKey)) {
throw invalid(`metadata-generation model "${configuredModel.id}" secret "${apiKeyEnv}" is unusable`);
}
}
labels.set(configuredModel.id, configuredModel.label);
models.set(configuredModel.id, Object.freeze({
id: configuredModel.id,
provider: adapter.litellmProvider,
model: configuredModel.upstreamModel,
...(configuredModel.metadataGeneration?.disableThinking === true
? { disableThinking: true as const } : {}),
...(configuredModel.endpoint ? { endpoint: Object.freeze({ ...configuredModel.endpoint }) } : {}),
...(apiKeyEnv ? { apiKeyEnv, apiKey } : {}),
}));
}
const result = new RestartLoadedMetadataGenerationModels(models, models.size ? catalog.defaultInteraction : null);
const safe = result.catalog();
return {
catalog: () => ({
default: safe.default,
models: safe.models.map((choice) => ({ ...choice, label: labels.get(choice.id) ?? choice.id })),
}),
resolve: (selection) => result.resolve(selection),
};
}
+144
View File
@@ -0,0 +1,144 @@
import type {
CatalogColumn,
CatalogLogicalRelationship,
CatalogPhysicalRelationship,
CatalogRepository,
CatalogTable,
} from "./types.js";
export type CatalogDescriptionSource = "curated" | "generated" | "source_comment";
export interface CatalogMetadataSnapshotColumn {
id: string;
name: string;
ordinalPosition: number;
dataType: string;
isNullable: boolean;
defaultExpression: string | null;
primaryKeyPosition: number | null;
sensitive: boolean;
description: string | null;
descriptionSource: CatalogDescriptionSource | null;
}
export interface CatalogMetadataSnapshotTable {
id: string;
name: string;
description: string | null;
descriptionSource: CatalogDescriptionSource | null;
columns: CatalogMetadataSnapshotColumn[];
}
export interface CatalogMetadataSnapshotRelationship {
id: string;
origin: "physical" | "generated" | "manual";
sourceTable: string;
sourceColumns: string[];
targetTable: string;
targetColumns: string[];
}
export interface CatalogMetadataSnapshot {
schemaVersion: 1;
workspaceId: string;
databaseId: string;
databaseName: string;
schemaName: string;
metadataContentRevision: number;
tables: CatalogMetadataSnapshotTable[];
relationships: CatalogMetadataSnapshotRelationship[];
}
function effectiveDescription(value: {
description: string | null;
generatedDescription: string | null;
sourceComment: string | null;
}): { description: string | null; descriptionSource: CatalogDescriptionSource | null } {
if (value.description?.trim()) {
return { description: value.description.trim(), descriptionSource: "curated" };
}
if (value.generatedDescription?.trim()) {
return { description: value.generatedDescription.trim(), descriptionSource: "generated" };
}
if (value.sourceComment?.trim()) {
return { description: value.sourceComment.trim(), descriptionSource: "source_comment" };
}
return { description: null, descriptionSource: null };
}
function snapshotColumn(column: CatalogColumn): CatalogMetadataSnapshotColumn {
return {
id: column.id,
name: column.name,
ordinalPosition: column.ordinalPosition,
dataType: column.dataType,
isNullable: column.isNullable,
defaultExpression: column.defaultExpression,
primaryKeyPosition: column.primaryKeyPosition,
sensitive: column.sensitive,
...effectiveDescription(column),
};
}
function snapshotRelationship(
relationship: CatalogPhysicalRelationship | CatalogLogicalRelationship,
): CatalogMetadataSnapshotRelationship {
const columns = [...relationship.columns].sort((left, right) => left.position - right.position);
return {
id: relationship.id,
origin: relationship.origin,
sourceTable: relationship.sourceTableName,
sourceColumns: columns.map((column) => column.sourceColumnName),
targetTable: relationship.targetTableName,
targetColumns: columns.map((column) => column.targetColumnName),
};
}
export async function buildCatalogMetadataSnapshot(
repository: CatalogRepository,
workspaceId: string,
expectedMetadataContentRevision: number,
): Promise<CatalogMetadataSnapshot> {
const database = await repository.getByWorkspace(workspaceId);
if (!database || database.preprocessingStatus !== "running") {
throw new Error("catalog preprocessing lease is not active");
}
if (database.metadataContentRevision !== expectedMetadataContentRevision) {
throw new Error("catalog metadata revision changed");
}
const catalogTables = await repository.listTables(database.id);
const tables: CatalogMetadataSnapshotTable[] = [];
for (const table of [...catalogTables].sort((left, right) => left.name.localeCompare(right.name))) {
const columns = await repository.listColumns(database.id, table.id);
tables.push({
id: table.id,
name: table.name,
...effectiveDescription(table),
columns: columns
.sort((left, right) => left.ordinalPosition - right.ordinalPosition || left.name.localeCompare(right.name))
.map(snapshotColumn),
});
}
const physical = await repository.listRelationships(database.id);
const logical = (await repository.listLogicalRelationships(database.id))
.filter((relationship) => relationship.status === "active");
const relationships = [...physical, ...logical]
.map(snapshotRelationship)
.sort((left, right) =>
left.sourceTable.localeCompare(right.sourceTable)
|| left.targetTable.localeCompare(right.targetTable)
|| left.id.localeCompare(right.id));
return {
schemaVersion: 1,
workspaceId,
databaseId: database.id,
databaseName: database.databaseName,
schemaName: database.schema,
metadataContentRevision: expectedMetadataContentRevision,
tables,
relationships,
};
}
+53
View File
@@ -0,0 +1,53 @@
import type { CatalogMetrics } from "./types.js";
export interface CatalogMetricCounts {
tables: number;
columns: number;
sensitiveColumns: number;
relationships: number;
describedTables: number;
describedColumns: number;
}
export function hasCatalogDescription(target: {
description: string | null;
generatedDescription: string | null;
}): boolean {
return Boolean(target.description?.trim() || target.generatedDescription?.trim());
}
export function latestCatalogTimestamp(
values: readonly (string | null | undefined)[],
): string | null {
let latest: number | undefined;
for (const value of values) {
if (!value) continue;
const timestamp = Date.parse(value);
if (!Number.isFinite(timestamp)) continue;
latest = latest === undefined ? timestamp : Math.max(latest, timestamp);
}
return latest === undefined ? null : new Date(latest).toISOString();
}
export function createCatalogMetrics(
databaseId: string | undefined,
counts: CatalogMetricCounts,
updatedAt: string | null,
): CatalogMetrics {
const descriptionTargets = counts.tables + counts.columns;
const describedTargets = counts.describedTables + counts.describedColumns;
return {
scope: databaseId === undefined ? "global" : "database",
databaseId: databaseId ?? null,
tables: counts.tables,
columns: counts.columns,
sensitiveColumns: counts.sensitiveColumns,
relationships: counts.relationships,
descriptionTargets,
describedTargets,
descriptionCoverage: descriptionTargets === 0
? 0
: Math.round((describedTargets / descriptionTargets) * 100),
updatedAt,
};
}

Some files were not shown because too many files have changed in this diff Show More