diff --git a/CONTEXT.md b/CONTEXT.md index bb7eaf10..58f1d73c 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -91,21 +91,79 @@ correzione successiva crea una nuova sessione derivata, collegata a quella prece dopo la finalizzazione. Non può modificare il ledger, gli artifact canonici o lo stato terminale della sessione. +## Memory + +**Memory Module** — Il modulo che possiede le conoscenze ed esperienze curate per +migliorare schema linking e generazione SQL di domande future. Le Memory appartengono +a un workspace e rimangono distinte dalle Evidence. + +**Memory Card** — L'unità di contenuto gestibile del Memory Module, con identità, +ambito di applicazione e provenienza. Il formato è allineato per analogia alle +Evidence, senza implicare la stessa origine o lo stesso percorso di pubblicazione. + +**Reusable Memory** — Una Memory Card che esprime un chiarimento di dominio, una +regola di costruzione SQL o un errore da evitare con motivo compreso e approvato. +La sua validità è circoscritta a un ambito esplicito e non deriva dalla sola +approvazione di una scelta occasionale in una domanda. + +**Solved Question** — Una Memory Card che conserva una domanda risolta con la +relativa soluzione SQL e il contesto necessario a interpretarla. È un exemplar +consultativo: i parametri e le scelte del caso non diventano regole generali. + +**Memory Graph** — L'insieme dei collegamenti espliciti fra card che contribuisce +al recupero di conoscenze pertinenti oltre alla somiglianza del contenuto. Il +ritrovamento di una card tramite un collegamento non ne implica l'approvazione. + +**Memory Link** — Un collegamento curato fra card, con destinazione e significato +espliciti, che contribuisce alla consultazione di contenuti pertinenti. La sua +rimozione non comporta la cancellazione delle card collegate. + ## Evidence +**Context specialist** — La persona competente sul dominio che redige e cura il +contenuto delle Evidence. Può essere distinta da chi amministra l'installazione; +il suo lavoro di redazione non richiede accesso al database applicativo. + +**Evidence draft** — Il documento iniziale scritto dallo specialista di contesto, +che il sistema acquisisce e raffina in Evidence Unit. Può essere redatto e +consegnato indipendentemente dall'installazione che userà le Evidence risultanti. + **Evidence Module** — Il modulo autonomo che possiede la preparazione delle Evidence e la loro consultazione durante il workflow. La preparazione avviene fuori dalle singole sessioni; il workflow usa soltanto contenuti già pubblicati. A runtime contribuisce agli stage semantici esistenti, senza diventare uno stage visibile e senza modificare ledger, artifact o stato del workflow. -**Source Evidence** — Un documento originale del workspace, conservato senza modifiche -come riferimento umano e origine della successiva ristrutturazione. +**Source Evidence** — Il documento o la dichiarazione che sostiene il contenuto +corrente di una Evidence Unit. Un documento acquisito viene conservato come +riferimento umano; una dichiarazione manuale attribuisce il contenuto alla persona +che lo ha scritto e approvato. + +**Manual Evidence declaration** — Una dichiarazione esplicita dell'amministratore +che sostiene una Evidence creata direttamente o una correzione del suo significato. +Non implica una verifica indipendente da parte di una fonte documentale esterna. + +**Evidence origin** — Il documento da cui una Evidence Unit è stata inizialmente +derivata. Può restare collegato per provenienza e confronto con gli aggiornamenti +anche quando una dichiarazione manuale sostiene il testo corrente. La sola origine +non dimostra il supporto semantico di una successiva correzione. + +**Local Evidence archive** — L'insieme delle Evidence curate custodite +dall'installazione, distinto dalle draft originali e dai contenuti derivati per +la ricerca. Comprende le correzioni manuali e i ritiri deliberati. + +**Consolidated Evidence** — Una versione delle Evidence locali controllata come +insieme coerente e pronta per l'attivazione. I file ancora in modifica non ne +cambiano il contenuto. + +**Active Evidence** — La versione consolidata disponibile alla consultazione del +core. Un tentativo di aggiornamento fallito conserva la versione attiva precedente. **Evidence Unit** — La più piccola unità semantica coerente, revisionabile e ricercabile -derivata da una sola Source Evidence. Possiede un identificatore stabile indipendente -dal kind, assegnato una volta nella forma `evidence:`; fonti diverse non vengono -fuse automaticamente. +fondata su una Source Evidence corrente, anche manuale, e con eventuale origine +documentale distinta. Possiede un identificatore stabile indipendente dal kind, +assegnato una volta nella forma `evidence:`; fonti diverse non vengono fuse +automaticamente. **Evidence kind** — La categoria semantica di una Evidence Unit, che ne determina i campi specifici e ne orienta l'uso. Ogni unità ha un solo kind primario; i tipi iniziali @@ -151,10 +209,9 @@ avanzare fino a un retry riuscito. nella sessione: stage semantico, purpose, generazione interrogata e identificatori delle Evidence restituite. Non duplica il contenuto delle Evidence. -**Curated Evidence** — Una o più Evidence Unit ristrutturate a partire da una Source -Evidence e conservate nel repository del workspace come proposte per la revisione -umana. Git conserva la versione precedente e rende visibile ogni modifica; una Curated -Evidence non è ancora contenuto autorevole del runtime. +**Curated Evidence** — Una o più Evidence Unit preparate da documenti o curate +manualmente. La presenza nell'archivio curato non implica da sola che il contenuto +sia già attivo per il workflow. **Published Evidence** — Le Curated Evidence valide appartenenti alla revisione attiva del workspace e alla generazione Evidence pubblicata. L'approvazione umana precede @@ -176,9 +233,17 @@ una Evidence Unit. Il sistema ne verifica deterministicamente la presenza dopo l normalizzazione meccanica; il curatore resta responsabile di verificarne la sufficienza semantica. -**Evidence resolution** — L'operazione esplicita con cui un curatore ritira una -Evidence Unit oppure la ricollega a un Source Evidence esistente. Aggiorna documento e -manifest insieme, lascia un diff Git revisionabile e non pubblica né crea commit. +**Evidence resolution** — La decisione esplicita con cui un curatore risolve un +problema di una Evidence Unit, correggendola, ritirandola oppure ricollegandola a +una fonte adeguata. + +**Source update conflict** — Un contrasto fra una fonte aggiornata e una correzione +manuale già approvata. La correzione resta in uso fino alla risoluzione esplicita +del confronto da parte dell'amministratore. + +**Evidence source refresh** — La riacquisizione delle fonti esterne richiesta +dall'amministratore per rilevarne le modifiche. Fra due aggiornamenti il contenuto +già acquisito resta il riferimento per preparazione e consultazione. **Review item** — Un blocco di revisione descritto da codice stabile, messaggio umano e campo opzionale. Finché viene mantenuto nell'Evidence Unit, ne impedisce la diff --git a/DESIGN.md b/DESIGN.md index 29c99d12..60021b55 100644 --- a/DESIGN.md +++ b/DESIGN.md @@ -290,11 +290,14 @@ default, hover, focus, active, disabled, loading, and error behavior where those - **Default / Hover / Active:** porcelain at rest, Sunken Surface on hover, and a muted Navigation Active red with a defined border when current. Exactly one top-level navigation control is current. - **Administrative controls:** the admin-only Administration accordion groups Database management, - a structural divider, Workspace management, and Pi management in that order. Its trigger exposes + Memory management, Evidence management, a structural divider, Workspace management, and Pi + management in that order. Its trigger exposes expanded state and starts collapsed by default, while non-admin users do not receive the accordion or its navigation actions. - **Responsive:** collapse navigation structurally at the application breakpoint. Do not shrink - labels into illegibility. + labels into illegibility. Below 768px, Memory and Evidence management use the full content + width; a Navigation button opens the shared accessible dialog. Selecting another archive + page or pressing Escape closes it. Desktop retains the session sidebar. ### Tabs @@ -320,8 +323,9 @@ default, hover, focus, active, disabled, loading, and error behavior where those ### Curated Evidence Documents Curated evidence follows a fixed reading order: title, compact type and purpose summary, scope, -typed content, supporting excerpts, review items, then collapsed technical provenance. Machine -metadata stays in invisible comments so GitHub Preview shows only the reviewable document. +typed content, supporting excerpts, review items, then technical provenance. Curated v4 files +use short, visible YAML frontmatter for identity and classification. The Markdown title and +body are authoritative; hidden payload comments are a legacy format converted on consolidation. `applies_to` is rendered as “Ambito di applicazione” with separate bullet lists for concepts, tables, and columns. Enum values also use lists. Tables are forbidden for metadata, scope, or any @@ -329,7 +333,9 @@ one-dimensional collection; reserve tables for genuinely two-dimensional dataset identifiers use inline code. SQL uses fenced code. Supporting excerpts use blockquotes. **The Review Surface Rule.** The visible Markdown must be readable without understanding the -machine contract. Technical metadata belongs in progressive disclosure, not above the title. +machine contract. In Administration, explain current and original provenance separately and +keep file-editing templates and Git instructions in progressive disclosure. Show actual host +paths with copy controls, never browser file links to container-only locations. ## Do's and Don'ts diff --git a/PROJECT_STATE.md b/PROJECT_STATE.md index 8d41e4e4..b490df64 100644 --- a/PROJECT_STATE.md +++ b/PROJECT_STATE.md @@ -1,6 +1,6 @@ # ThothII — Project State -Last updated: 2026-09-06. +Last updated: 2026-09-10. This file is the short operational snapshot. Stable commands and the architecture mental model live in `AGENTS.md`; current design and runtime contracts live under `docs/architecture/`, @@ -26,6 +26,165 @@ metadata catalog for administrative database configuration. The frontend renders and keeps the live transcript in memory. See `docs/architecture/components.md` for the detailed component and data-flow map. +## Memory M1–M3, Evidence E1–E3 and joint workflow repair X1 implemented + +Browser authorization now derives permissions from validated session roles using the current +catalog. This fixes Memory/Evidence navigation remaining disabled for administrators whose +remembered login predates those permissions; refreshing the page loads the updated permissions. + +The owner requested two independent administration projects: Memory management and Evidence +management. Their navigation entries must sit immediately after Database management, as peers; +neither page belongs to Database management. Both require complete browsing, filtering, and CRUD +without an active core session. Shared requirements and the two project briefs are linked from +`docs/plans/2026-09-08-memory-evidence-administration.md`. Memory administration is implemented; +Evidence administration is implemented through external file editing and explicit consolidation. +Evidence editing requires an explicit evolution of the current authoring/publication contract. +The M1 specification is at `docs/plans/2026-09-08-memory-m1-spec.md` and published as +[M1 — Archivio autorevole e amministrazione delle Memory Card](https://git.tylconsulting.it/mptyl/ThothII/issues/27) +with the `ready-for-agent` and `enhancement` labels. +It covers the authoritative store, administrative CRUD, projection recovery, and current +Memory/exemplar producer integration. Its accepted test boundaries are the public harness +service, Fastify APIs, and the AppShell page, joined by a focused real-stack browser path. +The owner confirmed those boundaries and authorized publication on 2026-09-08. +M1 is implemented locally: PostgreSQL authority in `thoth_memory`, versioned installation +migrations, admin CRUD for all four families, structured dependencies and links, explicit +projection recovery, verified recall and current workflow producers. The page requires no +active session or DWH binding. The existing `catalog-migrate` preparation service now runs +Memory migrations as well. Existing installations need that preparation before using M1; +the initial implementation did not deploy or migrate the owner's stacks. +The integrated browser check passed with real authentication, Fastify, ThtRunner, harness, +PostgreSQL and Qdrant; embeddings were deterministic and unrelated Pi/session activity used +test fixtures. See `docs/plans/2026-09-08-memory-m1-validation.md` for results and commands. +M2 is implemented locally: Memory dense/BM25 fusion, physical and business scope filters, +bounded outgoing-link expansion and joint ranking, all resolved against current PostgreSQL +authority. Migration `002_hybrid_projection.sql` makes old dense projections pending until +explicit retry/rebuild; Reference remains separate. The real retrieval check uses the configured +`qwen3-embedding:0.6b` model, a separate Ollama process with a read-only model-volume mount, +and isolated PostgreSQL/Qdrant resources. See `docs/plans/2026-09-09-memory-m2-validation.md`. +M3 is implemented: editable F8 summary grounded in effective approved decisions, explicit +updates with concurrent-edit protection, selected-card/link transactions and durable review +receipts. Finalization no longer saves exemplars implicitly. SQL rules and explained errors +are consulted in the existing F4/F6/F7 gates. Successful Catalog physical synchronization +performs dependency cleanup, preserving the original removals for recovery across restarts. +Migration `003_review_receipts.sql` is required. Validation includes a real GLM 5.3 generation +case against synthetic PostgreSQL data; see `docs/plans/2026-09-09-memory-m3-validation.md`. +Evidence administration and the joint X1 conflict-repair increment are now implemented. +The same local preview was subsequently rebuilt with M3 and migration 003 applied. +Core and frontend now include the final Memory review and physical dependency cleanup. +On 2026-09-09, at the owner's request, the local PSD Docker installation +`thothii-18998cca7b0a` was updated from this worktree. Core/frontend images were rebuilt, +Catalog and Memory migrations completed, and all five services became healthy. The UI is at +`http://127.0.0.1:8080`, using the existing local authentication and persistent volumes. +The `psd-clinical` authoritative Memory archive is initially empty (no legacy import). +The existing installation configuration remains in `/Users/mp/projects/ThothII/deploy/psd/`; +`/private/tmp/thothii-memory-preview.sh` invokes its Compose files with a final build-context +and migration-command override from this worktree. A future build from the main checkout +will use that checkout's code, so retain the worktree override until the changes are integrated. +Memory revision history and compatibility with existing development sessions are not requirements. +The agreed Memory scope includes reusable domain clarifications, SQL construction rules, solved +questions, and explained, approved mistakes to avoid. This extends the current runtime's +`concept_clarified`-only reusable Memory contract. The owner also approved an editable final summary +for proposed additions/updates and Memory consumption in the relevant existing review gates. +Further agreed behavior includes persistent corrections for Memory/Evidence conflicts through +explicit choices, deletion of Memory with invalid dependencies after successful physical schema +synchronization, and bounded functional/regression tests instead of a general quality benchmark. +In-house Memory evolution is approved, including hybrid Qdrant retrieval and explicit card links +traversed in core, without a dedicated graph database. Both capabilities belong to the current +scope. PostgreSQL is the agreed authority for Memory Cards, links, and schema dependencies; +Qdrant is a rebuildable index. Links are reviewed with cards and editable in Administration; +deleting a card removes its incident links while preserving the other cards. The accepted, +implemented architecture is recorded in +`docs/adr/0018-use-postgres-for-memory-and-qdrant-for-retrieval.md`. +For Evidence administration, Q12 rejects a separate editorial publication workflow. +The final R0 direction uses external editors and a manual consolidation command that +validates structure, reports required corrections, updates derived metadata, and +activates the local Evidence index. The operator then checks the diff and runs Git +commit/push manually. No watcher, automated Git, or web editor is required. +The owner's later simplification +instruction removes the requirement to support administrative edits while core work +is in progress. Deliberate corrections made by the workflow itself remain supported. +The context specialist writes Evidence drafts independently of the installation and +without PostgreSQL access; the system refines them and stores them locally for use +and maintenance. The revised Q11 recommendation separates external drafts from a +durable local canonical file archive, removing automatic commit/push from CRUD. +The owner has now accepted local files and requires discoverable, editable Markdown +for nontechnical domain specialists, with no JSONL management surface. Editors on +Mac/PC or vim/nano on the server are the chosen R0 editing surface. Evidence +management keeps browsing/filtering/detail and clearly identifies the persistent +working tree, each Markdown file's actual host path, and the manual commands. +E1 implements editable Curated Evidence v4, deterministic legacy conversion, persistent +local files and a consolidation API with immutable candidates, manual provenance, +deletion records and recoverable activation. E2 connects its installed command, runtime source +selection and administration page. The operator completes Git steps manually. +All 35 PSD units were converted on an isolated copy with identical IDs and typed content. +Tests exercise visible edits through normalization, indexing and recall with real Qdrant; +see `docs/plans/2026-09-09-evidence-e1-validation.md` and +`docs/contracts/curated-evidence-v4.md`. E2 is now running on the local Docker preview: +all 35 PSD units were converted and indexed through the installed command. The editable +host archive is `/Users/mp/projects/ThothII/deploy/psd/evidence-registry/repo/psd-clinical/evidence`. +The original registry-volume checkout and original author repository were retained. +The installation descriptor now includes `workspace-bindings.yaml` and `evidence-host.yaml` +so native maintenance uses the same volumes and checkout as the preview. Descriptor backup: +`/private/tmp/thothii-installation-before-e2.yaml`; previous native binary: `/private/tmp/tht-before-e2`. +The launcher remains `bash /private/tmp/thothii-memory-preview.sh`; its worktree build override +is still required until integration. Runtime lease filenames now include rendered bytes so +upgrading the Evidence renderer does not collide with old immutable configs; Catalog +input fingerprints and readiness are unchanged. See `docs/plans/2026-09-09-evidence-e2-validation.md`. +E3 adds explicit source acquisition/refinement and durable comparisons in Evidence management. +Local drafts go in `evidence/incoming/`; original local documents and configured HTTP/S3 +sources are reacquired only by **Import or refresh sources** or installed +`tht workspace evidence refresh --workspace `. Keep/replace decisions, including the +native `workspace evidence decide` command, activate through the existing archive/index +path and have durable retry state. Acquired source versions and remote provenance stay +local; old documentary lineage can coexist with current manual declarations. +The installed preview's 35 PSD sources were verified unchanged with no new proposals. +The previous native CLI is backed up at `/private/tmp/tht-before-e3`. +See `docs/plans/2026-09-09-evidence-e3-validation.md` for tests and validation limits. +X1 adds `reviewer_archive_repair`: closed alternatives with complete before/after content, +explicit rejection/reformulation, admin-only application, durable session receipts and +retry of saved but inactive corrections. The coordinator uses the canonical Memory and +Evidence services; phase approval remains separate. Migration `004_archive_repairs.sql` +is required. See `docs/contracts/archive-repair.md` for authorization and recovery limits. +X1 validation is recorded in `docs/plans/2026-09-09-archive-repair-x1-validation.md`: +both archives were corrected and retrieved through the actual CLI with real PostgreSQL +and Qdrant; desktop/mobile widget behavior and permission failures were verified separately. +The local preview images include X1 and migration 004 is applied. Follow-up technical +acceptance passed with configured GLM 5.3 generating closed conflict alternatives and +the chosen Memory correction persisted and retrieved. An authenticated browser test +also verified both administration pages, navigation order, Evidence filtering, workspace +404s, desktop/mobile layout, and Evidence survival after Memory deletion. All approved +technical increments/checks are complete. A real PSD domain-conflict session remains +the reviewer's semantic acceptance check; automated cases did not modify PSD knowledge. +The joint browser inspection also fixed mobile archive navigation: below 768px, +Memory and Evidence keep the full content width and open navigation in the shared +accessible dialog. Desktop retains its sidebar; the local frontend image includes this fix. +The owner accepted the remaining simplifications and requested explicit clarification +of the core format change and the simple terminal-based Git check. E1 must adapt +the Evidence parser, renderer, authoring, validation, and normalization; convert +existing files and reindex; and verify that visible edits reach core consumption. +Preserve the internal typed model where possible. This precedes the administrative +page and is not merely a presentation change. Human inspection uses normal Git +status and optional line-level diff commands; no custom diff viewer or mandatory +double review is required. +The suggestion of making PostgreSQL the Evidence authority was withdrawn after this +clarification; it was never implemented or accepted as a replacement for Q11. +Q13 keeps manual corrections active when updated sources contradict them, until an +administrator resolves the comparison; deleted Evidence must not be regenerated +automatically. Q14 allows direct manual creation and records a manual declaration +as the current source, preserving any original document as distinct provenance. +Q15 refreshes external sources only on explicit administrator request; normal saves +and lookups do not reacquire them. These decisions, now implemented through E1–E3, are recorded in +`docs/adr/0019-author-evidence-in-app-with-automatic-activation.md`. +The decision-by-decision review is recorded in +`docs/plans/2026-09-08-memory-evidence-simplification-review.md`; it distinguishes the +new constraints from the revised technical recommendations. It also specifies a +sequential save with minimal durable retry state and direct Memory cleanup after +successful schema synchronization, without new queues or event infrastructure. +The delivery order remains Memory, Evidence, and persistent conflict repair between +both modules. Local curated Evidence must survive preprocessing Clear and be backed +up as primary data; the Qdrant projection remains rebuildable. +No runtime gate change or administrative page is implemented yet. + ## Evidence restructuring — accepted The evidence restructuring and PSD migration completed real acceptance on 2026-08-25. diff --git a/backend/src/app.ts b/backend/src/app.ts index 7e33cd80..d26a7cb9 100644 --- a/backend/src/app.ts +++ b/backend/src/app.ts @@ -7,6 +7,7 @@ import { fileURLToPath } from "node:url"; import { tmpdir } from "node:os"; import type { AppConfig } from "./config.js"; import { ThtRunner } from "./tht/tht-runner.js"; +import { createMemoryCleanup } from "./catalog/memory-cleanup.js"; import { PiProcessManager } from "./pi/pi-process-manager.js"; import { SseHub } from "./sse/sse-hub.js"; import { authenticateSession, captureAuthConfigSnapshot, configuredOrigin } from "./auth/auth.js"; @@ -80,6 +81,8 @@ import { createProductionWorkspacePreprocessingService } from "./workspace-maint import type { WorkspacePreprocessingService } from "./workspaces/preprocessing-service.js"; import { PreprocessingStateStore } from "./workspaces/preprocessing-state.js"; import { workspacePreprocessingRoutes } from "./routes/workspace-preprocessing.js"; +import { memoryRoutes } from "./routes/memory.js"; +import { evidenceRoutes } from "./routes/evidence.js"; export interface BuildAppDeps { thtRunner?: ThtRunner; @@ -93,7 +96,7 @@ export interface BuildAppDeps { workspaceDiagnoser?: WorkspaceDiagnoser; workspaceDatabaseTester?: WorkspaceDatabaseTester; workspaceSecretStore?: WorkspaceSecretStore; - workspacePreprocessingService?: Pick; + workspacePreprocessingService?: Pick & Partial>; catalogRepository?: CatalogRepository; catalogService?: CatalogService; catalogPostgresAccess?: CatalogPostgresAccess; @@ -265,6 +268,11 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc catalogSchemaIntrospector, catalogOperationCoordinator, config.catalogSyncTimeoutMs, + createMemoryCleanup(tht as ThtRunner, { + internalQdrantUrl: config.internalQdrantUrl, internalEmbeddingUrl: config.internalEmbeddingUrl, + internalEmbeddingId: config.internalEmbeddingId, internalEmbeddingModel: config.internalEmbeddingModel, + internalEmbeddingDimensions: config.internalEmbeddingDimensions, + }), ); app.addHook("onReady", async () => { await catalogSyncWorker.initialize(); }); app.addHook("onReady", async () => { await descriptionGenerationWorker.initialize(); }); @@ -493,7 +501,16 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc return maintenanceBarrier.status(); }); sqlRoutes(app, { tht: tht as ThtRunner, getSettings, workspaceRegistry }); + memoryRoutes(app, { runner: tht as ThtRunner, registry: workspaceRegistry, runtime: { + internalQdrantUrl: config.internalQdrantUrl, + internalEmbeddingUrl: config.internalEmbeddingUrl, + internalEmbeddingModel: config.internalEmbeddingModel, + internalEmbeddingDimensions: config.internalEmbeddingDimensions, + } }); metaRoutes(app, { harnessDir: config.harnessDir, modelCatalog: runtimeModelCatalog }); + evidenceRoutes(app, { runner: tht as ThtRunner, registry: workspaceRegistry, + registryRoot: config.workspaceRegistry.root, hostRegistryRoot: config.evidenceHostRegistryRoot, + service: workspacePreprocessingService }); workspaceRoutes(app, { registry: workspaceRegistry, config: config.workspaceRegistry, diff --git a/backend/src/auth/auth.ts b/backend/src/auth/auth.ts index e73008ac..dfd5f361 100644 --- a/backend/src/auth/auth.ts +++ b/backend/src/auth/auth.ts @@ -119,7 +119,9 @@ export function authenticateSession(deps: AuthDependencies): preHandlerHookHandl subject: session.subject, ...(session.displayName === undefined ? {} : { displayName: session.displayName }), roles: session.roles, - permissions: session.permissions, + // Sessions can outlive a deployment that changes the role permission catalog. + // resolve() has already checked validity, including current local user roles. + permissions: rolesToPermissions(session.roles), isAdmin: session.roles.includes("admin"), }; if (STATE_CHANGING_METHODS.has(request.method)) { diff --git a/backend/src/auth/config.ts b/backend/src/auth/config.ts index 82429cde..9d2fe5ff 100644 --- a/backend/src/auth/config.ts +++ b/backend/src/auth/config.ts @@ -38,7 +38,7 @@ const MAX_MAPPED_GROUPS = 128; const ROLES = ["user", "admin"] as const; export const PERMISSION_CATALOG: readonly Permission[] = [ "session.use", "session.read_all", "session.manage_all", "settings.manage", - "workspace.manage", "workspace.secrets.manage", "database.manage", "pi.manage", "auth.diagnostics.read", + "workspace.manage", "workspace.secrets.manage", "database.manage", "memory.manage", "evidence.manage", "pi.manage", "auth.diagnostics.read", ]; const invalid = (): Error => new Error("authentication configuration is invalid"); diff --git a/backend/src/auth/session-store.ts b/backend/src/auth/session-store.ts index a7d52379..721f8db4 100644 --- a/backend/src/auth/session-store.ts +++ b/backend/src/auth/session-store.ts @@ -33,7 +33,7 @@ const EMPTY_HKDF_SALT = Buffer.alloc(0); const ROLES = ["user", "admin"] as const; const PERMISSIONS = [ "session.use", "session.read_all", "session.manage_all", "settings.manage", - "workspace.manage", "workspace.secrets.manage", "database.manage", "pi.manage", "auth.diagnostics.read", + "workspace.manage", "workspace.secrets.manage", "database.manage", "memory.manage", "evidence.manage", "pi.manage", "auth.diagnostics.read", ] as const satisfies readonly Permission[]; const invalid = (): Error => new Error("auth_session_store_invalid"); diff --git a/backend/src/auth/types.ts b/backend/src/auth/types.ts index d5d74db1..a815ea30 100644 --- a/backend/src/auth/types.ts +++ b/backend/src/auth/types.ts @@ -5,7 +5,7 @@ export type Role = "user" | "admin"; export type Permission = | "session.use" | "session.read_all" | "session.manage_all" | "settings.manage" | "workspace.manage" | "workspace.secrets.manage" - | "database.manage" | "pi.manage" | "auth.diagnostics.read"; + | "database.manage" | "memory.manage" | "evidence.manage" | "pi.manage" | "auth.diagnostics.read"; export interface AuthenticationSessionConfig { regularTtlSeconds: number; diff --git a/backend/src/catalog/memory-cleanup.ts b/backend/src/catalog/memory-cleanup.ts new file mode 100644 index 00000000..fc7ec06a --- /dev/null +++ b/backend/src/catalog/memory-cleanup.ts @@ -0,0 +1,23 @@ +import type { ThtRunner } from "../tht/tht-runner.js"; +import type { SemanticRuntimeConfig } from "../workspaces/runtime-renderer.js"; +import type { CatalogSyncRun, WorkspaceDatabase } from "./types.js"; + +/** Internal continuation of an applied physical sync, using the harness Memory boundary. */ +export function createMemoryCleanup(runner: Pick, runtime: SemanticRuntimeConfig) { + return async (database: WorkspaceDatabase, run: CatalogSyncRun): Promise => { + if (run.phase !== "memory_cleanup" || !run.plannedDiff) throw new Error("Physical cleanup is not committed"); + const result = await runner.withPrincipal({ issuer: "installation", subject: "catalog-sync", + roles: ["admin"], permissions: ["memory.manage"], isAdmin: true, + }).runWithRuntimeSnapshot(["memory", "admin", "--workspace", database.workspaceId], JSON.stringify({ + action: "cleanup", runtime, request: { sync_id: run.id, database: database.databaseName, + schema_name: database.schema, removed_tables: run.plannedDiff.deletedTables, + removed_columns: run.plannedDiff.deletedColumns.map(column => ({ table: column.tableName, column: column.columnName })), + }, + })); + const payload = JSON.parse(result.stdout); + if (result.code !== 0 || payload.indexed !== true || !Number.isInteger(payload.deleted) || payload.deleted < 0) { + throw new Error("Memory cleanup is incomplete"); + } + return payload.deleted; + }; +} diff --git a/backend/src/catalog/memory-repository.ts b/backend/src/catalog/memory-repository.ts index 0897159a..618abbad 100644 --- a/backend/src/catalog/memory-repository.ts +++ b/backend/src/catalog/memory-repository.ts @@ -926,6 +926,7 @@ export class MemoryCatalogRepository implements CatalogRepository { scope: CatalogSyncScope, tableIds: readonly string[], snapshot: ObservedSchemaSnapshot, + syncRunId?: string, ): Promise { const database = this.records.get(databaseId); if (!database || database.version !== expectedDatabaseVersion) return undefined; @@ -1141,6 +1142,7 @@ export class MemoryCatalogRepository implements CatalogRepository { if (scope === "all") { this.records.set(databaseId, { ...database, schemaSyncedVersion: expectedDatabaseVersion, schemaSyncedAt: now }); } + if (syncRunId) await this.updateSyncRun(syncRunId, { phase: "memory_cleanup" }); return { tables: (await this.listTables(databaseId)).length, columns: [...this.columns.values()].filter((column) => this.tables.get(column.tableId)?.databaseId === databaseId).length, @@ -1254,7 +1256,7 @@ export class MemoryCatalogRepository implements CatalogRepository { for (const run of this.syncRuns.values()) { if (["queued", "running", "awaiting_confirmation", "applying"].includes(run.state)) { await this.updateSyncRun(run.id, { - state: "interrupted", phase: "completed", finishedAt: new Date().toISOString(), + state: "interrupted", phase: run.phase === "memory_cleanup" ? "memory_cleanup" : "completed", finishedAt: new Date().toISOString(), errorCode: "SYNC_INTERRUPTED", errorMessage: "Synchronization was interrupted by a service restart", }); } diff --git a/backend/src/catalog/repository.ts b/backend/src/catalog/repository.ts index 25ae090a..03e3562e 100644 --- a/backend/src/catalog/repository.ts +++ b/backend/src/catalog/repository.ts @@ -1604,6 +1604,7 @@ export class KyselyCatalogRepository implements CatalogRepository { scope: CatalogSyncScope, tableIds: readonly string[], snapshot: ObservedSchemaSnapshot, + syncRunId?: string, ): Promise { return await this.db.transaction().execute(async (trx) => { const database = await trx.selectFrom("workspaceDatabases").select("version") @@ -1779,6 +1780,10 @@ export class KyselyCatalogRepository implements CatalogRepository { schemaSyncedAt: now, }).where("id", "=", databaseId).execute(); } + if (syncRunId) { + await trx.updateTable("catalogSyncRuns").set({ phase: "memory_cleanup" }) + .where("id", "=", syncRunId).where("databaseId", "=", databaseId).execute(); + } return { tables: snapshot.tables.length, columns: scope === "tables" ? undefined : snapshot.columns.length, @@ -1882,7 +1887,7 @@ export class KyselyCatalogRepository implements CatalogRepository { async interruptActiveSyncRuns(): Promise { await this.db.updateTable("catalogSyncRuns").set({ - state: "interrupted", phase: "completed", errorCode: "worker_restarted", + state: "interrupted", phase: sql`case when phase='memory_cleanup' then phase else 'completed' end`, errorCode: "worker_restarted", errorMessage: "Synchronization was interrupted by a backend restart.", finishedAt: sql`now()`, updatedAt: sql`now()`, leaseOwner: null, leaseExpiresAt: null, diff --git a/backend/src/catalog/sync-worker.ts b/backend/src/catalog/sync-worker.ts index a880306b..84e15543 100644 --- a/backend/src/catalog/sync-worker.ts +++ b/backend/src/catalog/sync-worker.ts @@ -58,6 +58,8 @@ export class CatalogSyncWorker { private readonly introspector: CatalogSchemaIntrospector, private readonly operations: CatalogOperationCoordinator, private readonly timeoutMs: number, + private readonly cleanupMemory: (database: WorkspaceDatabase, run: CatalogSyncRun) => Promise + = async () => 0, ) {} async initialize(): Promise { @@ -67,6 +69,9 @@ export class CatalogSyncWorker { } async start(database: WorkspaceDatabase, scope: CatalogSyncScope, tableIds: readonly string[]): Promise { + if ((await this.repository.listSyncRuns(database.id, 100)).some(run => run.phase === "memory_cleanup")) { + throw new CatalogConflictError("Retry the pending Memory cleanup before starting another synchronization"); + } const uniqueTableIds = [...new Set(tableIds)]; if (scope === "columns") { const tables = await Promise.all(uniqueTableIds.map((tableId) => this.repository.getTable(database.id, tableId))); @@ -109,7 +114,7 @@ export class CatalogSyncWorker { async cancel(runId: string): Promise { const run = await this.repository.getSyncRun(runId); if (!run) return undefined; - if (run.state === "applying" || TERMINAL_STATES.has(run.state)) return run; + if (run.phase === "memory_cleanup" || run.state === "applying" || TERMINAL_STATES.has(run.state)) return run; await this.repository.requestSyncRunCancellation(runId); this.controllers.get(runId)?.abort(); if (run.state === "queued" || run.state === "awaiting_confirmation") { @@ -138,6 +143,16 @@ export class CatalogSyncWorker { } const database = await this.repository.get(previous.databaseId); if (!database) return undefined; + if (previous.phase === "memory_cleanup") { + const release = this.operations.reserve(database.id); + try { + const queued = await this.repository.updateSyncRun(runId, { state: "queued", cancelRequested: false, + errorCode: null, errorMessage: null, finishedAt: null, leaseOwner: null, leaseExpiresAt: null }); + this.reservations.set(runId, release); + this.launch(runId); + return queued; + } catch (error) { release(); throw error; } + } return await this.start(database, previous.scope, previous.tableIds); } @@ -179,6 +194,10 @@ export class CatalogSyncWorker { if (!database || database.version !== claimed.requestedDatabaseVersion) { throw new CatalogConflictError("Database binding changed before synchronization started"); } + if (claimed.phase === "memory_cleanup") { + await this.finishMemoryCleanup(database, claimed); + return; + } const progress: CatalogSchemaScanProgress = async (phase, counts) => { await this.checkCancelled(runId); await this.repository.updateSyncRun(runId, { @@ -230,36 +249,30 @@ export class CatalogSyncWorker { claimed.scope, claimed.tableIds, snapshot, + runId, ); if (!applied) throw new CatalogConflictError("Database binding changed before schema changes were applied"); - await this.repository.updateSyncRun(runId, { - state: "succeeded", - phase: "completed", - counts: applied, - finishedAt: new Date().toISOString(), - observedSnapshot: null, - plannedDiff: null, - confirmationToken: null, - heartbeatAt: new Date().toISOString(), - leaseOwner: null, - leaseExpiresAt: null, - }); - await this.repository.appendSyncEvent(runId, "info", "succeeded", "Synchronization completed.", { ...applied }); - this.release(runId); + const cleanupRun = await this.repository.updateSyncRun(runId, { counts: applied }); + if (!cleanupRun) throw new Error("Synchronization run disappeared"); + await this.finishMemoryCleanup(database, cleanupRun); + /* Completion is recorded only after the durable Memory cleanup succeeds. */ + return; } catch (error) { const current = await this.repository.getSyncRun(runId); const cancelled = !timedOut && (error instanceof SyncCancelledError || controller.signal.aborted || current?.cancelRequested); - const failure = timedOut + const failure = current?.phase === "memory_cleanup" + ? { code: "memory_cleanup_pending", message: "Catalog synchronized. Memory cleanup is pending; retry this synchronization to complete it." } + : timedOut ? { code: "schema_sync_timed_out", message: "Schema synchronization timed out." } : safeFailure(error); await this.repository.updateSyncRun(runId, { state: cancelled ? "cancelled" : "failed", - phase: "completed", + phase: current?.phase === "memory_cleanup" ? "memory_cleanup" : "completed", errorCode: cancelled ? null : failure.code, errorMessage: cancelled ? null : failure.message, finishedAt: new Date().toISOString(), - observedSnapshot: null, - plannedDiff: null, + observedSnapshot: current?.phase === "memory_cleanup" ? current.observedSnapshot : null, + plannedDiff: current?.phase === "memory_cleanup" ? current.plannedDiff : null, confirmationToken: null, leaseOwner: null, leaseExpiresAt: null, @@ -278,6 +291,20 @@ export class CatalogSyncWorker { } } + private async finishMemoryCleanup(database: WorkspaceDatabase, run: CatalogSyncRun): Promise { + if (!run.plannedDiff || run.phase !== "memory_cleanup") throw new Error("Missing committed cleanup context"); + await this.repository.appendSyncEvent(run.id, "info", "memory_cleanup", "Removing Memory cards with deleted physical dependencies."); + const memoryDeleted = await this.cleanupMemory(database, run); + const counts = { ...run.counts, memoryDeleted }; + await this.repository.updateSyncRun(run.id, { + state: "succeeded", phase: "completed", counts, finishedAt: new Date().toISOString(), + observedSnapshot: null, plannedDiff: null, confirmationToken: null, + heartbeatAt: new Date().toISOString(), leaseOwner: null, leaseExpiresAt: null, + }); + await this.repository.appendSyncEvent(run.id, "info", "succeeded", "Synchronization completed.", counts); + this.release(run.id); + } + private assertCapability(scope: CatalogSyncScope, snapshot: ObservedSchemaSnapshot): void { const required = scope === "all" ? ["tables", "columns", "relationships"] as const : [scope] as const; for (const name of required) { diff --git a/backend/src/catalog/types.ts b/backend/src/catalog/types.ts index 877d5f27..62167474 100644 --- a/backend/src/catalog/types.ts +++ b/backend/src/catalog/types.ts @@ -366,7 +366,7 @@ export type CatalogSyncState = export type CatalogSyncPhase = | "queued" | "connecting" | "scanning_tables" | "scanning_columns" | "scanning_relationships" | "planning" | "awaiting_confirmation" - | "applying" | "completed"; + | "applying" | "memory_cleanup" | "completed"; export interface CatalogSchemaDiff { deletedTables: string[]; @@ -375,6 +375,7 @@ export interface CatalogSchemaDiff { } export interface CatalogSyncCounts { + memoryDeleted?: number; tables?: number; columns?: number; relationships?: number; @@ -587,6 +588,7 @@ export interface CatalogRepository { scope: CatalogSyncScope, tableIds: readonly string[], snapshot: ObservedSchemaSnapshot, + syncRunId?: string, ): Promise; createSyncRun( databaseId: string, diff --git a/backend/src/config.ts b/backend/src/config.ts index de12b133..39f28a28 100644 --- a/backend/src/config.ts +++ b/backend/src/config.ts @@ -32,6 +32,7 @@ export interface AppConfig { piManagementTimeoutMs: number; secretsFile?: string; installationConfigFile?: string; + evidenceHostRegistryRoot?: string; modelCatalogFile?: string; sensitivityNer?: { pythonExecutable: string; @@ -476,6 +477,7 @@ export function loadConfig( piManagementTimeoutMs: piManagementTimeout(env.PI_MANAGEMENT_TIMEOUT_MS), secretsFile, installationConfigFile, + evidenceHostRegistryRoot: env.THT_EVIDENCE_HOST_REGISTRY_ROOT || undefined, modelCatalogFile, sensitivityNer, piAuthFile, diff --git a/backend/src/routes/evidence.ts b/backend/src/routes/evidence.ts new file mode 100644 index 00000000..dc037ab3 --- /dev/null +++ b/backend/src/routes/evidence.ts @@ -0,0 +1,90 @@ +import path from "node:path"; +import type { FastifyInstance } from "fastify"; +import { z } from "zod"; +import { requirePermission, isPrincipalContext } from "../auth/authorization.js"; +import type { ThtRunner } from "../tht/tht-runner.js"; +import type { WorkspaceRegistry } from "../workspaces/registry.js"; +import type { WorkspacePreprocessingService } from "../workspaces/preprocessing-service.js"; + +const params = z.object({ workspaceId: z.string().regex(/^[a-z][a-z0-9-]{2,62}$/), + evidenceId: z.string().regex(/^evidence:[a-z0-9]+(?:-[a-z0-9]+)*$/).optional() }); +const query = z.object({ + q: z.string().max(1000).optional(), kind: z.enum(["domain", "glossary", "enum", "example", "mapping", "normalization", "formula", "reference"]).optional(), + purpose: z.enum(["disambiguation", "rewriting", "schema_linking", "sql_generation"]).optional(), + status: z.enum(["new", "modified", "active", "removed", "review_required", "legacy", "invalid"]).optional(), + concept: z.string().max(300).optional(), table: z.string().max(300).optional(), column: z.string().max(300).optional(), + source: z.string().max(300).optional(), language: z.string().max(30).optional(), + sort: z.enum(["title", "id", "kind", "status"]).optional(), direction: z.enum(["asc", "desc"]).optional(), + page: z.coerce.number().int().positive().optional(), page_size: z.coerce.number().int().min(1).max(100).optional(), +}).strict(); +const quote = (value: string) => `'${value.replaceAll("'", "'\\''")}'`; + +export function evidenceRoutes(app: FastifyInstance, deps: { + runner: Pick; registry: Pick; + registryRoot: string; hostRegistryRoot?: string; + service: Partial>; +}) { + app.post("/workspaces/:workspaceId/evidence/sources", async (request, reply) => { + const principal = requirePermission(request, reply, "evidence.manage"); + if (!isPrincipalContext(principal)) return reply; + try { + const { workspaceId } = params.parse(request.params); + const body = z.discriminatedUnion("action", [ + z.object({ action: z.literal("refresh") }).strict(), + z.object({ action: z.literal("decide"), sourceId: z.string().regex(/^[a-f0-9]{64}$/), + revision: z.string().regex(/^[a-f0-9]{64}$/), decision: z.enum(["keep", "replace"]) }).strict(), + ]).parse(request.body); + if (!(await deps.registry.list()).some(w => w.id === workspaceId)) return reply.code(404).send({ code: "workspace_invalid" }); + if (!deps.service.evidenceSources) throw new Error("Evidence service unavailable"); + return await deps.service.evidenceSources({ workspaceId, ...body, actor: principal.subject }); + } catch (error) { + return reply.code(error instanceof z.ZodError ? 400 : 503).send({ code: "evidence_unavailable", + message: error instanceof z.ZodError ? "Invalid source request." : "Evidence source service is unavailable." }); + } + }); + app.get<{ Params: { workspaceId: string; evidenceId?: string } }>("/workspaces/:workspaceId/evidence", read); + app.get<{ Params: { workspaceId: string; evidenceId?: string } }>("/workspaces/:workspaceId/evidence/:evidenceId", read); + async function read(request: import("fastify").FastifyRequest, reply: import("fastify").FastifyReply) { + const principal = requirePermission(request, reply, "evidence.manage"); + if (!isPrincipalContext(principal)) return reply; + try { + const { workspaceId, evidenceId } = params.parse(request.params); + const filters = query.parse(request.query); + if (!(await deps.registry.list()).some(w => w.id === workspaceId)) return reply.code(404).send({ code: "workspace_invalid" }); + const root = path.join(deps.registryRoot, "repo", workspaceId); + const result = await deps.runner.withPrincipal(principal).runWithRuntimeSnapshot( + ["evidence", "admin", "--workspace", workspaceId], + JSON.stringify({ root, query: { ...filters, ...(evidenceId ? { id: evidenceId } : {}) } }), + ); + const payload = JSON.parse(result.stdout); + if (result.code !== 0) return reply.code(503).send(payload); + if (evidenceId && !payload.item) return reply.code(404).send({ code: "evidence_not_found", message: "Evidence was not found." }); + const host = deps.hostRegistryRoot; + const hostPath = host && (path.isAbsolute(host) || path.win32.isAbsolute(host)) + ? (path.win32.isAbsolute(host) && !path.isAbsolute(host) ? path.win32 : path).join(host, "repo") : null; + const hostJoin = hostPath && path.win32.isAbsolute(hostPath) && !path.isAbsolute(hostPath) ? path.win32.join : path.join; + return { ...payload, location: { repository: hostPath, workspace: hostPath ? hostJoin(hostPath, workspaceId) : null, + runtime_workspace: root, host: "Installation host", command: `tht workspace evidence consolidate --workspace ${workspaceId}`, + git_commands: hostPath ? [ + `git -C ${quote(hostPath)} status --short -- ${quote(workspaceId + "/evidence")}`, + `git -C ${quote(hostPath)} diff -- ${quote(workspaceId + "/evidence")}`, + `git -C ${quote(hostPath)} add -A -- ${quote(workspaceId + "/evidence")}`, + `git -C ${quote(hostPath)} commit --only -m 'Curate Evidence' -- ${quote(workspaceId + "/evidence")}`, + `git -C ${quote(hostPath)} push`, + ] : [] } }; + } catch (error) { + return reply.code(error instanceof z.ZodError ? 400 : 503).send({ code: "evidence_unavailable", + message: error instanceof z.ZodError ? "Evidence filters are invalid." : "Evidence archive is unavailable." }); + } + } + app.post("/workspaces/:workspaceId/evidence/consolidate", async (request, reply) => { + const principal = requirePermission(request, reply, "evidence.manage"); + if (!isPrincipalContext(principal)) return reply; + try { + const { workspaceId } = params.parse(request.params); + if (!(await deps.registry.list()).some(w => w.id === workspaceId)) return reply.code(404).send({ code: "workspace_invalid" }); + if (!deps.service.consolidateEvidence) throw new Error("Evidence service unavailable"); + return await deps.service.consolidateEvidence({ workspaceId }); + } catch { return reply.code(503).send({ code: "evidence_unavailable", message: "Evidence consolidation is unavailable." }); } + }); +} diff --git a/backend/src/routes/memory.ts b/backend/src/routes/memory.ts new file mode 100644 index 00000000..c905b6fc --- /dev/null +++ b/backend/src/routes/memory.ts @@ -0,0 +1,94 @@ +import type { FastifyInstance, FastifyReply, FastifyRequest } from "fastify"; +import { z } from "zod"; +import { isPrincipalContext, requirePermission } from "../auth/authorization.js"; +import type { ThtRunner } from "../tht/tht-runner.js"; +import type { WorkspaceRegistry } from "../workspaces/registry.js"; +import type { SemanticRuntimeConfig } from "../workspaces/runtime-renderer.js"; + +const paramsSchema = z.object({ + workspaceId: z.string().regex(/^[a-z][a-z0-9_-]{0,63}$/), + cardId: z.string().regex(/^mem-[0-9a-f-]{36}$/).optional(), +}); +const family = z.enum(["domain_clarification", "sql_rule", "solved_question", "explained_error"]); +const cardSchema = z.object({ + family, subject: z.string().trim().min(1).max(1000), + detail: z.string().max(50000).default(""), scope: z.string().trim().min(1).max(10000), + rationale: z.string().max(10000).default(""), question: z.string().max(10000).default(""), + sql: z.string().max(100000).default(""), + concepts: z.array(z.string().trim().min(1).max(200)).max(100).default([]), + dependencies: z.array(z.object({ + database: z.string().trim().min(1).max(200), schema_name: z.string().max(200).default(""), + table: z.string().max(200).default(""), column: z.string().max(200).default(""), + }).strict()).max(200).default([]), + links: z.array(z.object({ + target_id: z.string().min(1).max(100), meaning: z.string().trim().min(1).max(1000), + }).strict()).max(200).default([]), +}).strict(); +const querySchema = z.object({ + q: z.string().max(1000).optional(), family: family.optional(), + concept: z.string().max(200).optional(), database: z.string().max(200).optional(), + table: z.string().max(200).optional(), column: z.string().max(200).optional(), + origin: z.enum(["manual", "workflow"]).optional(), + updated_after: z.iso.datetime({ offset: true }).optional(), + updated_before: z.iso.datetime({ offset: true }).optional(), + page: z.coerce.number().int().positive().optional(), + page_size: z.coerce.number().int().min(1).max(100).optional(), + sort: z.enum(["updated_at", "created_at", "subject", "family"]).optional(), + direction: z.enum(["asc", "desc"]).optional(), +}).strict(); + +export function memoryRoutes(app: FastifyInstance, deps: { + runner: Pick; + registry: Pick; + runtime: SemanticRuntimeConfig; +}) { + const invoke = (action: string) => async (request: FastifyRequest, reply: FastifyReply) => { + const principal = requirePermission(request, reply, "memory.manage"); + if (!isPrincipalContext(principal)) return reply; + try { + const { workspaceId, cardId } = paramsSchema.parse(request.params); + const input = action === "list" ? querySchema.parse(request.query) + : action === "create" || action === "update" + ? { card: cardSchema.parse(request.body), ...(cardId ? { id: cardId } : {}) } + : cardId ? { id: cardId } : {}; + // Registry identity is enough: administrative browsing must not require a DWH binding. + const workspaces = await deps.registry.list(); + if (!workspaces.some(workspace => workspace.id === workspaceId)) { + return reply.code(404).send({ code: "workspace_invalid", message: "Workspace was not found." }); + } + const { workspace } = await deps.registry.read(workspaceId); + const result = await deps.runner.withPrincipal(principal).runWithRuntimeSnapshot( + ["memory", "admin", "--workspace", workspaceId], + JSON.stringify({ action, request: input, runtime: { + ...deps.runtime, memoryLanguage: workspace.workspace.language, + } }), + ); + let payload; + try { payload = JSON.parse(result.stdout); } catch { + return reply.code(503).send({ code: "memory_unavailable", message: "Memory archive is unavailable." }); + } + if (result.code !== 0) { + const codes: Record = { + memory_invalid: 400, memory_forbidden: 403, memory_not_found: 404, + memory_conflict: 409, memory_unavailable: 503, + }; + const code = typeof payload?.code === "string" && payload.code in codes + ? payload.code : "memory_unavailable"; + return reply.code(codes[code]).send({ code, message: "Memory operation could not be completed." }); + } + return reply.code(action === "create" ? 201 : 200).send(payload); + } catch (error) { + if (error instanceof z.ZodError) { + return reply.code(400).send({ code: "memory_invalid", message: "Memory request is invalid." }); + } + return reply.code(503).send({ code: "memory_unavailable", message: "Memory archive is unavailable." }); + } + }; + app.get("/workspaces/:workspaceId/memory", invoke("list")); + app.post("/workspaces/:workspaceId/memory", invoke("create")); + app.get("/workspaces/:workspaceId/memory/pending", invoke("pending")); + app.get("/workspaces/:workspaceId/memory/:cardId", invoke("show")); + app.put("/workspaces/:workspaceId/memory/:cardId", invoke("update")); + app.delete("/workspaces/:workspaceId/memory/:cardId", invoke("delete")); + app.post("/workspaces/:workspaceId/memory/:cardId/retry", invoke("retry")); +} diff --git a/backend/src/routes/sessions.ts b/backend/src/routes/sessions.ts index 4eeca147..682dd357 100644 --- a/backend/src/routes/sessions.ts +++ b/backend/src/routes/sessions.ts @@ -626,6 +626,23 @@ export function sessionRoutes( } catch (error) { return lifecycleFailure(reply, error); } const rt = d.mgr.get(id); if (!rt) return reply.code(404).send({ error: "sessione non attiva" }); + const pending = rt.bridge.pendingWidget?.() as any; + const response = (req.body as any)?.ui_response; + if (pending?.widget === "archive-repair" && response?.id === pending.id && !response.control) { + const choices = response.choices; + if (!Array.isArray(choices) || choices.length !== 1) + return reply.code(400).send({ error: "Select one archive repair choice" }); + if (choices[0] !== "continue" && choices[0] !== "reject") { + const option = pending.repair?.options?.find((item: any) => item.id === choices[0]); + if (!option || !["memory", "evidence"].includes(option.archive)) + return reply.code(400).send({ error: "Unknown archive repair choice" }); + const permission = option.archive === "memory" ? "memory.manage" : "evidence.manage"; + if (!principal.isAdmin || !hasPermission(principal, permission)) + return reply.code(403).send({ error: "Archive corrections require an administrator" }); + if (rt.ownerKey !== `${principal.issuer}\0${principal.subject}`) + return reply.code(409).send({ error: "Resume the session with your account before correcting an archive" }); + } + } if (!rt.bridge.respond((req.body as any).ui_response)) { return reply.code(409).send({ error: "risposta non corrispondente al gate in attesa" }); } diff --git a/backend/src/workspace-maintenance.ts b/backend/src/workspace-maintenance.ts index 39d3726a..0128bbaa 100644 --- a/backend/src/workspace-maintenance.ts +++ b/backend/src/workspace-maintenance.ts @@ -18,7 +18,7 @@ export interface WorkspaceMaintenanceIo { writeStderr(value: string): void; } -type Command = "inspect" | "preprocess-run" | "preprocess-clear"; +type Command = "inspect" | "preprocess-run" | "preprocess-clear" | "evidence-consolidate" | "evidence-refresh" | "evidence-decide"; function failureResult( operation: string, @@ -65,6 +65,9 @@ function parseRequest(command: string, stdin: string): Record { inspect: ["schemaVersion", "workspaceId"], "preprocess-run": ["schemaVersion", "workspaceId"], "preprocess-clear": ["schemaVersion", "workspaceId"], + "evidence-consolidate": ["schemaVersion", "workspaceId"], + "evidence-refresh": ["schemaVersion", "workspaceId"], + "evidence-decide": ["schemaVersion", "workspaceId", "sourceId", "revision", "decision"], }; const allowed = allowedByCommand[command]; if (!allowed) throw new Error("unknown command"); @@ -81,6 +84,16 @@ function exitCodeFor(result: WorkspaceOperationResult): number { async function dispatch(command: Command, service: WorkspacePreprocessingService, request: Record): Promise { switch (command) { + case "evidence-refresh": + return await service.evidenceSources({ workspaceId: request.workspaceId as string, action: "refresh" }); + case "evidence-decide": + if (typeof request.sourceId !== "string" || !/^[a-f0-9]{64}$/.test(request.sourceId) + || typeof request.revision !== "string" || !/^[a-f0-9]{64}$/.test(request.revision) + || (request.decision !== "keep" && request.decision !== "replace")) throw new Error("invalid request"); + return await service.evidenceSources({ workspaceId: request.workspaceId as string, action: "decide", + sourceId: request.sourceId, revision: request.revision, decision: request.decision }); + case "evidence-consolidate": + return await service.consolidateEvidence({ workspaceId: request.workspaceId as string }); case "inspect": return await service.inspect({ workspaceId: request.workspaceId as string }); case "preprocess-run": @@ -127,7 +140,8 @@ export async function runWorkspaceMaintenanceCli( || message === "unexpected request field" || message === "invalid workspace id"; return command in { - inspect: true, "preprocess-run": true, "preprocess-clear": true, + inspect: true, "preprocess-run": true, "preprocess-clear": true, "evidence-consolidate": true, + "evidence-refresh": true, "evidence-decide": true, } ? (requestError ? 2 : 1) : 2; } } diff --git a/backend/src/workspaces/evidence/preprocessing.ts b/backend/src/workspaces/evidence/preprocessing.ts index 8f8b3d36..9cf936e3 100644 --- a/backend/src/workspaces/evidence/preprocessing.ts +++ b/backend/src/workspaces/evidence/preprocessing.ts @@ -22,6 +22,7 @@ export interface EvidencePreprocessingRequest { evidence: EvidenceConfig; job: EvidenceJobState; dryRun?: boolean; + consolidate?: boolean; httpPrivateHostAllowlist?: readonly string[]; } @@ -56,7 +57,7 @@ function isPrivateHost(hostname: string): boolean { return hostname.endsWith(".internal"); } -function evidencePolicy( +export function evidencePolicy( evidence: EvidenceConfig, httpPrivateHostAllowlist?: readonly string[], ): EvidencePreprocessingOutcome | undefined { @@ -95,13 +96,14 @@ function jobResult(job: EvidenceJobState): Pick< }; } -async function runEvidenceStage( +export async function runEvidenceStage( request: EvidencePreprocessingRequest, deps: EvidencePreprocessingDependencies, ): Promise { const payload = await deps.runStage([ "preprocess", "evidence", + ...(request.consolidate ? ["--consolidate"] : []), ...(request.dryRun ? ["--dry-run"] : []), ...(request.job.childRuns.evidence ? ["--resume", request.job.childRuns.evidence] diff --git a/backend/src/workspaces/preprocessing-service.ts b/backend/src/workspaces/preprocessing-service.ts index 66de997d..aeb346d1 100644 --- a/backend/src/workspaces/preprocessing-service.ts +++ b/backend/src/workspaces/preprocessing-service.ts @@ -3,6 +3,8 @@ import { renameSync, rmSync, writeFileSync, mkdirSync } from "node:fs"; import { join } from "node:path"; import { continueEvidencePreprocessing, + evidencePolicy, + runEvidenceStage, type EvidencePreprocessingDependencies, type EvidencePreprocessingOutcome, } from "./evidence/preprocessing.js"; @@ -108,6 +110,72 @@ function baseResult( export class WorkspacePreprocessingService { constructor(private readonly deps: WorkspacePreprocessingServiceDeps) {} + async evidenceSources(options: { workspaceId: string; action: "refresh" | "decide"; + sourceId?: string; revision?: string; decision?: "keep" | "replace"; actor?: string }): Promise { + const runtime = await this.deps.acquireActiveRuntime(options.workspaceId); + try { + if (options.action === "refresh") { + const refused = evidencePolicy(runtime.workspace.evidence, this.deps.httpPrivateHostAllowlist); + if (refused) return baseResult(runtime, "evidence refresh", "failed", refused.code); + } + const argv = ["evidence", "sources", options.action, "--json", "-c", "/dev/fd/3"]; + if (options.action === "decide") { + if (!/^[a-f0-9]{64}$/.test(options.sourceId ?? "") || !/^[a-f0-9]{64}$/.test(options.revision ?? "") + || !["keep", "replace"].includes(options.decision ?? "")) throw new Error("Invalid source decision"); + const preflight = await this.deps.evidencePreflight(runtime.workspace); + if (!preflight.ok) return baseResult(runtime, "evidence sources", "failed", preflight.code); + argv.push("--source-id", options.sourceId!, "--revision", options.revision!, "--decision", options.decision!); + } + argv.push("--actor", options.actor ?? "installation operator"); + const result = await this.deps.runChild({ argv, configPath: runtime.configLease.path }); + const payload = JSON.parse(result.stdout); + if (result.exitCode !== 0 || payload.status !== "succeeded") throw new Error( + typeof payload.error === "string" ? payload.error.slice(0, 1500) : "Evidence source operation failed"); + return baseResult(runtime, `evidence ${options.action}`, "succeeded", "ok", { + counts: this.numberRecord(payload.counts), + warnings: [options.action === "refresh" ? "Source comparisons are ready in Evidence management. Active content is unchanged." + : "Source decision activated locally. Commit and push the Evidence tree manually."], + }); + } catch (error) { + return baseResult(runtime, `evidence ${options.action}`, "failed", "evidence_materialization_required", { + warnings: [error instanceof Error ? error.message : "Evidence source operation failed"], + }); + } + } + + async consolidateEvidence(options: { workspaceId: string }): Promise { + const runtime = await this.deps.acquireActiveRuntime(options.workspaceId); + const preflight = await this.deps.evidencePreflight(runtime.workspace); + if (!preflight.ok) return baseResult(runtime, "evidence consolidate", "failed", preflight.code); + // Evidence has its own durable archive/corpus jobs. Do not alter Catalog readiness + // or the full preprocessing job when publishing this one component. + try { + const outcome = await runEvidenceStage({ evidence: runtime.workspace.evidence, + consolidate: true, job: { runId: randomBytes(16).toString("hex"), completedStages: [], childRuns: {} }, + }, { + runStage: async argv => { + const result = await this.deps.runChild({ argv, configPath: runtime.configLease.path }); + const payload = JSON.parse(result.stdout); + if (result.exitCode !== 0 || payload.status !== "succeeded") { + return Promise.reject(new Error(typeof payload.error === "string" ? payload.error.slice(0, 1500) : "Evidence consolidation failed. Retry the command.")); + } + return payload; + }, + persistJob: () => undefined, + evidencePreflight: async () => preflight, + requireRunId: value => this.requireRunId(value), + numberRecord: value => this.numberRecord(value), + }); + return baseResult(runtime, "evidence consolidate", "succeeded", "ok", { + ...outcome, warnings: ["Evidence is active locally. Catalog and Schema readiness are unchanged. Commit and push the Evidence files manually."], + }); + } catch (error) { + return baseResult(runtime, "evidence consolidate", "failed", "evidence_materialization_required", { + warnings: [error instanceof Error ? error.message : "Evidence consolidation failed. Retry the command."], + }); + } + } + async inspect(options: { workspaceId: string }): Promise { try { const runtime = await this.deps.acquireActiveRuntime(options.workspaceId); diff --git a/backend/src/workspaces/runtime-config-lease.ts b/backend/src/workspaces/runtime-config-lease.ts index f7f8d0de..0fb81cc1 100644 --- a/backend/src/workspaces/runtime-config-lease.ts +++ b/backend/src/workspaces/runtime-config-lease.ts @@ -497,7 +497,10 @@ export async function publishDeterministicRuntimeConfigLease(options: { rendered.workspaceRevision, renderedConfigObject, ); - const identitySuffix = inputFingerprintValue.slice(7, 23); + // The Catalog input identity intentionally excludes representation-only changes. + // A lease also identifies its bytes, so a new renderer never collides with an + // immutable config produced by an earlier release for the same Catalog inputs. + const identitySuffix = sha256(inputFingerprintValue + "\n" + publishedConfig).slice(7, 23); const preprocessingRoot = ensureTrustedDirectory(join( options.dataRoot, diff --git a/backend/src/workspaces/runtime-renderer.ts b/backend/src/workspaces/runtime-renderer.ts index 1d795b64..c034dac7 100644 --- a/backend/src/workspaces/runtime-renderer.ts +++ b/backend/src/workspaces/runtime-renderer.ts @@ -1,4 +1,4 @@ -import { basename, join } from "node:path"; +import { basename, dirname, join } from "node:path"; import { stringify } from "yaml"; import { buildInstallationContract } from "./contracts.js"; import { validateWorkspaceDescriptor, type WorkspaceDescriptor } from "./schema.js"; @@ -179,6 +179,7 @@ function renderEvidence( return { evidence: { ...(workspace.evidence.schema_version === 2 ? { schema_version: 2 } : {}), + local_archive_root: join(dirname(dirname(context.revisionContentRoot)), "repo", context.workspaceId), sources: [renderedSource], }, vector: { diff --git a/backend/test/archive-repair-response.test.ts b/backend/test/archive-repair-response.test.ts new file mode 100644 index 00000000..e985a5be --- /dev/null +++ b/backend/test/archive-repair-response.test.ts @@ -0,0 +1,48 @@ +import Fastify from "fastify"; +import { expect, test, vi } from "vitest"; +import { sessionRoutes } from "../src/routes/sessions.js"; +import type { PrincipalContext } from "../src/auth/principal.js"; + +test.each([ + [false, "m", 403], [false, "e", 403], [false, "reject", 204], + [true, "m", 204], [true, "e", 204], [true, "forged", 400], +] as const)("archive repair response enforces the responding principal (%s, %s)", async (admin, choice, status) => { + const app = Fastify(); + const principal: PrincipalContext = { issuer: "test", subject: "reviewer", isAdmin: admin, + roles: admin ? ["admin"] : ["user"], + permissions: admin ? ["session.use", "memory.manage", "evidence.manage"] : ["session.use"] }; + app.addHook("preHandler", async request => { request.principal = principal; }); + const respond = vi.fn(() => true); + sessionRoutes(app, { + tht: {}, mgr: { get: () => ({ ownerKey: "test\0reviewer", bridge: { + respond, pendingWidget: () => ({ id: "gate", widget: "archive-repair", repair: { options: [ + { id: "m", archive: "memory" }, { id: "e", archive: "evidence" }, + ] } }), + } }) }, + } as unknown as Parameters[1]); + try { + const result = await app.inject({ method: "POST", url: "/sessions/s/response", + payload: { ui_response: { id: "gate", choices: [choice] } } }); + expect(result.statusCode).toBe(status); + expect(respond).toHaveBeenCalledTimes(status === 204 ? 1 : 0); + } finally { await app.close(); } +}); + +test("an administrator cannot attribute a repair to another runtime's principal", async () => { + const app = Fastify(); + app.addHook("preHandler", async request => { request.principal = { + issuer: "test", subject: "second-admin", isAdmin: true, roles: ["admin"], + permissions: ["session.use", "memory.manage"], + }; }); + const respond = vi.fn(); + sessionRoutes(app, { tht: {}, mgr: { get: () => ({ ownerKey: "test\0first-admin", bridge: { + respond, pendingWidget: () => ({ id: "gate", widget: "archive-repair", + repair: { options: [{ id: "m", archive: "memory" }] } }), + } }) } } as unknown as Parameters[1]); + try { + const result = await app.inject({ method: "POST", url: "/sessions/s/response", + payload: { ui_response: { id: "gate", choices: ["m"] } } }); + expect(result.statusCode).toBe(409); + expect(respond).not.toHaveBeenCalled(); + } finally { await app.close(); } +}); diff --git a/backend/test/auth-config.test.ts b/backend/test/auth-config.test.ts index bb1f243b..5bc0667e 100644 --- a/backend/test/auth-config.test.ts +++ b/backend/test/auth-config.test.ts @@ -188,6 +188,8 @@ test("roles collapse duplicates and admin contains all administrative permission "workspace.manage", "workspace.secrets.manage", "database.manage", + "memory.manage", + "evidence.manage", "pi.manage", "auth.diagnostics.read", ]); diff --git a/backend/test/auth-routes-local.test.ts b/backend/test/auth-routes-local.test.ts index b3a515bc..2a5cbf04 100644 --- a/backend/test/auth-routes-local.test.ts +++ b/backend/test/auth-routes-local.test.ts @@ -130,7 +130,7 @@ test("local login sets a non-persistent opaque session cookie and exposes only a roles: ["admin"], permissions: [ "session.use", "session.read_all", "session.manage_all", "settings.manage", - "workspace.manage", "workspace.secrets.manage", "database.manage", "pi.manage", "auth.diagnostics.read", + "workspace.manage", "workspace.secrets.manage", "database.manage", "memory.manage", "evidence.manage", "pi.manage", "auth.diagnostics.read", ], isAdmin: true, csrfToken: expect.stringMatching(/^[A-Za-z0-9_-]{43}$/), diff --git a/backend/test/auth.test.ts b/backend/test/auth.test.ts index c5a5379c..2db3261c 100644 --- a/backend/test/auth.test.ts +++ b/backend/test/auth.test.ts @@ -7,6 +7,8 @@ import { join } from "node:path"; import { expandLocalHome, localPrincipal, upstreamPrincipal } from "../src/auth/principal.js"; import { buildApp } from "../src/app.js"; import { loadConfig } from "../src/config.js"; +import { rolesToPermissions } from "../src/auth/config.js"; +import type { Permission, Role } from "../src/auth/types.js"; test("server smoke rejects retired trusted claims under OIDC authentication", () => { const smoke = readFileSync("../scripts/unified-deployment-smoke.sh", "utf8"); @@ -32,7 +34,7 @@ test("local mode resolves a stable local principal", async () => { roles: ["admin"], permissions: [ "session.use", "session.read_all", "session.manage_all", "settings.manage", - "workspace.manage", "workspace.secrets.manage", "database.manage", "pi.manage", "auth.diagnostics.read", + "workspace.manage", "workspace.secrets.manage", "database.manage", "memory.manage", "evidence.manage", "pi.manage", "auth.diagnostics.read", ], isAdmin: true, }); @@ -74,7 +76,7 @@ test("upstream mode accepts only normalized proxy principal headers", async () = issuer: "portal", subject: "42", displayName: "Alice", roles: ["user", "admin"], permissions: [ "session.use", "session.read_all", "session.manage_all", "settings.manage", - "workspace.manage", "workspace.secrets.manage", "database.manage", "pi.manage", "auth.diagnostics.read", + "workspace.manage", "workspace.secrets.manage", "database.manage", "memory.manage", "evidence.manage", "pi.manage", "auth.diagnostics.read", ], isAdmin: true, }); @@ -206,15 +208,33 @@ test("the session boundary rejects tht maintenance headers outside exact loopbac })).statusCode).toBe(503); }); -test("the session boundary touches a valid cookie session through the bounded Task 7 store operation", async () => { +test.each<{ + name: string; + roles: Role[]; + storedPermissions: Permission[]; +}>([ + { name: "current user", roles: ["user"], storedPermissions: ["session.use"] }, + { + name: "administrator signed in before archive permissions existed", + roles: ["admin"], + storedPermissions: rolesToPermissions(["admin"]).filter( + (permission) => permission !== "memory.manage" && permission !== "evidence.manage", + ), + }, + { + name: "user with obsolete administrative permissions", + roles: ["user"], + storedPermissions: ["session.use", "memory.manage", "evidence.manage"], + }, +])("the session boundary derives current permissions and touches a valid cookie: $name", async ({ roles, storedPermissions }) => { const sessions = { resolve: vi.fn(async () => ({ version: 1, issuer: "local", subject: "user-1", method: "local", - roles: ["user"], - permissions: ["session.use"], + roles, + permissions: storedPermissions, userAuthRevision: 1, authConfigRevision: "b".repeat(64), remembered: false, @@ -249,7 +269,13 @@ test("the session boundary touches a valid cookie session through the bounded Ta app.get("/private", async (request) => getPrincipal(request)); const token = "z".repeat(43); - expect((await app.inject({ method: "GET", url: "/private", headers: { cookie: `thothii_session=${token}` } })).statusCode).toBe(200); + const response = await app.inject({ method: "GET", url: "/private", headers: { cookie: `thothii_session=${token}` } }); + expect(response.statusCode).toBe(200); + expect(response.json()).toMatchObject({ + roles, + permissions: rolesToPermissions(roles), + isAdmin: roles.includes("admin"), + }); expect(sessions.touch).toHaveBeenCalledWith(token); }); diff --git a/backend/test/catalog-repository.integration.test.ts b/backend/test/catalog-repository.integration.test.ts index aac2bd09..2cdcb6d6 100644 --- a/backend/test/catalog-repository.integration.test.ts +++ b/backend/test/catalog-repository.integration.test.ts @@ -124,8 +124,17 @@ test.skipIf(!dockerAvailable)("PostgreSQL migration enforces one database per wo ], relationships: [], }; - expect(await repository.applySchemaSync(created.id, 1, "columns", [], fullColumnsSnapshot)) + const memoryCleanupRun = await repository.createSyncRun(created.id, "columns", [], 1); + await repository.updateSyncRun(memoryCleanupRun.id, { state: "applying", phase: "applying" }); + await expect(repository.applySchemaSync(created.id, 999, "columns", [], fullColumnsSnapshot, memoryCleanupRun.id)) + .resolves.toBeUndefined(); + expect((await repository.getSyncRun(memoryCleanupRun.id))?.phase).toBe("applying"); + expect(await repository.applySchemaSync(created.id, 1, "columns", [], fullColumnsSnapshot, memoryCleanupRun.id)) .toMatchObject({ created: 4, deleted: 0 }); + expect((await repository.getSyncRun(memoryCleanupRun.id))?.phase).toBe("memory_cleanup"); + await repository.interruptActiveSyncRuns(); + expect(await repository.getSyncRun(memoryCleanupRun.id)) + .toMatchObject({ state: "interrupted", phase: "memory_cleanup" }); expect((await repository.listColumns(created.id, patients.id)).map((column) => ({ name: column.name, sensitive: column.sensitive, @@ -233,6 +242,7 @@ test.skipIf(!dockerAvailable)("PostgreSQL migration enforces one database per wo }); expect(await repository.listSyncRuns(created.id)).toEqual([ expect.objectContaining({ id: syncRun.id, tableIds: [patients.id] }), + expect.objectContaining({ id: memoryCleanupRun.id, phase: "memory_cleanup" }), ]); const lockedDatabase = await repository.create({ diff --git a/backend/test/catalog-schema-routes.test.ts b/backend/test/catalog-schema-routes.test.ts index b81141dc..5283c8b4 100644 --- a/backend/test/catalog-schema-routes.test.ts +++ b/backend/test/catalog-schema-routes.test.ts @@ -117,6 +117,7 @@ async function setup(env: Record = {}) { }); const introspector: CatalogSchemaIntrospector = { scan }; const operations = new CatalogOperationCoordinator(); + const cleanup = vi.fn(async () => ({ code: 0, stdout: JSON.stringify({ indexed: true, deleted: 0 }), stderr: "" })); const registry = { list: vi.fn(async () => [revision]), listCatalog: vi.fn(async () => [{ id: "psd-clinical", name: "Policlinico San Donato", configurationState: "ready", revision }]), @@ -124,7 +125,7 @@ async function setup(env: Record = {}) { readPinned: vi.fn(async () => ({ workspace, workspaceConfigPath: revision.snapshotPath })), } as unknown as WorkspaceRegistry; const app = buildApp(loadConfig({ THT_HARNESS_DIR: "/missing", NODE_ENV: "test", ...env }), { - thtRunner: {} as never, + thtRunner: { withPrincipal: () => ({ runWithRuntimeSnapshot: cleanup }) } as never, workspaceRegistry: registry, workspaceSecretStore: new WorkspaceSecretStore({ root: secretRoot, runtimeRoot, installationId: "test" }), catalogRepository: repository, @@ -133,7 +134,7 @@ async function setup(env: Record = {}) { workspaceDiagnoser: vi.fn(), }); return { - app, repository, database: (await repository.get(created.id))!, scan, operations, + app, repository, database: (await repository.get(created.id))!, scan, operations, cleanup, setObserved(next: ObservedSchemaSnapshot) { observed = next; }, }; } @@ -614,6 +615,45 @@ test("waits for confirmation and rescans before applying destructive changes", a expect(scan).toHaveBeenCalledTimes(3); }); +test("retries committed Memory cleanup with the original removals without rescanning", async () => { + const { app, repository, database, scan, setObserved, cleanup } = await setup(); + await seedCatalog(repository, database); + const observed = snapshot(); + observed.tables = observed.tables.filter(t => t.name !== "visits"); + observed.columns = observed.columns.filter(c => c.tableName !== "visits"); + observed.relationships = []; + setObserved(observed); + cleanup.mockResolvedValueOnce({ code: 0, stdout: JSON.stringify({ indexed: false, deleted: 2 }), stderr: "" }); + const started = await app.inject({ method: "POST", url: `/catalog/databases/${database.id}/sync-runs`, + payload: { version: database.version, scope: "all", tableIds: [] } }); + const waiting = await waitFor(repository, started.json().id, "awaiting_confirmation"); + expect(cleanup).not.toHaveBeenCalled(); + await app.inject({ method: "POST", url: `/catalog/sync-runs/${waiting.id}/confirm`, + payload: { confirmationToken: waiting.confirmationToken } }); + const failed = await waitFor(repository, waiting.id, "failed"); + expect(failed).toMatchObject({ phase: "memory_cleanup", errorCode: "memory_cleanup_pending", + plannedDiff: { deletedTables: ["visits"] } }); + expect((await repository.listTables(database.id)).map(t => t.name)).toEqual(["patients"]); + const callsBeforeRetry = scan.mock.calls.length; + const blocked = await app.inject({ method: "POST", url: `/catalog/databases/${database.id}/sync-runs`, + payload: { version: database.version, scope: "all", tableIds: [] } }); + expect(blocked.statusCode).toBe(409); + // Simulate startup recovery and an unreachable DWH after schema application. + await repository.interruptActiveSyncRuns(); + scan.mockRejectedValue(new Error("DWH unavailable")); + cleanup.mockResolvedValue({ code: 0, stdout: JSON.stringify({ indexed: true, deleted: 2 }), stderr: "" }); + const retried = await app.inject({ method: "POST", url: `/catalog/sync-runs/${waiting.id}/retry` }); + expect(retried.statusCode).toBe(202); + const completed = await waitFor(repository, waiting.id, "succeeded"); + expect(completed.counts).toMatchObject({ memoryDeleted: 2 }); + expect(scan.mock.calls.length).toBe(callsBeforeRetry); + expect(cleanup).toHaveBeenCalledTimes(2); + expect(cleanup.mock.calls[1]).toEqual(cleanup.mock.calls[0]); + const request = JSON.parse((cleanup.mock.calls[0] as unknown as string[])[1]!).request; + expect(request).toMatchObject({ sync_id: waiting.id, database: "warehouse", + schema_name: "datawarehouse", removed_tables: ["visits"] }); +}); + test("deletes every catalog table for multiple selected databases and cascades dependent metadata", async () => { const { app, repository, database } = await setup(); const second = await repository.create({ diff --git a/backend/test/evidence-routes.test.ts b/backend/test/evidence-routes.test.ts new file mode 100644 index 00000000..8e76725a --- /dev/null +++ b/backend/test/evidence-routes.test.ts @@ -0,0 +1,69 @@ +import Fastify from "fastify"; +import { afterEach, expect, test, vi } from "vitest"; +import { evidenceRoutes } from "../src/routes/evidence.js"; +import type { ThtRunner } from "../src/tht/tht-runner.js"; +import type { PrincipalContext } from "../src/auth/principal.js"; + +const apps: ReturnType[] = []; +afterEach(async () => { await Promise.all(apps.splice(0).map(app => app.close())); }); +function setup(admin = true) { + const app = Fastify(); apps.push(app); + const principal: PrincipalContext = { issuer: "test", subject: "curator", roles: [admin ? "admin" : "user"], + permissions: admin ? ["evidence.manage"] : ["session.use"], isAdmin: admin }; + app.addHook("preHandler", async request => { request.principal = principal; }); + const run = vi.fn(async () => ({ code: 0, stdout: JSON.stringify({ items: [], total: 0 }), stderr: "" })); + const withPrincipal = vi.fn(() => ({ runWithRuntimeSnapshot: run }) as unknown as ThtRunner); + const consolidateEvidence = vi.fn(); + const evidenceSources = vi.fn().mockResolvedValue({status: "succeeded"}); + evidenceRoutes(app, { runner: { withPrincipal }, registryRoot: "/data/registry", hostRegistryRoot: "/srv/Thoth workspaces", + registry: { list: async () => [{ id: "sales", commit: "a".repeat(40), blob: "b".repeat(40), snapshotPath: "/missing" }] }, + service: { consolidateEvidence, evidenceSources } }); + return { app, run, withPrincipal, principal, consolidateEvidence, evidenceSources }; +} +test("browses complete local files with host paths, without a database runtime", async () => { + const { app, run, withPrincipal, principal } = setup(); + const response = await app.inject("/workspaces/sales/evidence?kind=domain&concept=orders&page=2"); + expect(response.statusCode).toBe(200); + expect(withPrincipal).toHaveBeenCalledWith(principal); + const [argv, snapshot] = run.mock.calls[0] as unknown as [string[], string]; + expect(argv).toEqual(["evidence", "admin", "--workspace", "sales"]); + expect(JSON.parse(snapshot)).toEqual({ root: "/data/registry/repo/sales", query: { kind: "domain", concept: "orders", page: 2 } }); + expect(response.json().location.workspace).toBe("/srv/Thoth workspaces/repo/sales"); + expect(response.json().location.git_commands[0]).toContain("git -C '/srv/Thoth workspaces/repo'"); +}); + +test("source actions require admin, bind the actor, and reject cross-workspace or arbitrary inputs", async () => { + const denied = setup(false); + expect((await denied.app.inject({method: "POST", url: "/workspaces/sales/evidence/sources", payload: {action: "refresh"}})).statusCode).toBe(403); + expect(denied.evidenceSources).not.toHaveBeenCalled(); + const allowed = setup(); + expect((await allowed.app.inject({method: "POST", url: "/workspaces/sales/evidence/sources", payload: {action: "refresh"}})).statusCode).toBe(200); + expect(allowed.evidenceSources).toHaveBeenCalledWith({workspaceId: "sales", action: "refresh", actor: "curator"}); + const choice = {action: "decide", sourceId: "a".repeat(64), revision: "b".repeat(64), decision: "keep"}; + expect((await allowed.app.inject({method: "POST", url: "/workspaces/sales/evidence/sources", payload: choice})).statusCode).toBe(200); + expect(allowed.evidenceSources).toHaveBeenLastCalledWith({...choice, workspaceId: "sales", actor: "curator"}); + for (const payload of [{action: "refresh", url: "http://localhost"}, {...choice, actor: "forged"}, {...choice, revision: "../other"}]) { + expect((await allowed.app.inject({method: "POST", url: "/workspaces/sales/evidence/sources", payload})).statusCode).toBe(400); + } + expect((await allowed.app.inject({method: "POST", url: "/workspaces/other/evidence/sources", payload: choice})).statusCode).toBe(404); + expect(allowed.evidenceSources).toHaveBeenCalledTimes(2); +}); +test.each(["/", "/evidence:rule", "/consolidate"])("refuses non-admin access %s", async suffix => { + const { app, run, consolidateEvidence } = setup(false); + const response = await app.inject({ method: suffix === "/consolidate" ? "POST" : "GET", url: `/workspaces/sales/evidence${suffix === "/" ? "" : suffix}` }); + expect(response.statusCode).toBe(403); expect(run).not.toHaveBeenCalled(); expect(consolidateEvidence).not.toHaveBeenCalled(); +}); +test("rejects workspace escape, unknown workspace and unsupported filters", async () => { + const { app, run } = setup(); + expect((await app.inject("/workspaces/foreign/evidence")).statusCode).toBe(404); + expect((await app.inject("/workspaces/sales/evidence?root=/private")).statusCode).toBe(400); + expect((await app.inject("/workspaces/sales/evidence?page_size=101")).statusCode).toBe(400); + expect(run).not.toHaveBeenCalled(); +}); +test("a missing unit is 404 and consolidation preserves the partial outcome", async () => { + const { app, consolidateEvidence } = setup(); + expect((await app.inject("/workspaces/sales/evidence/evidence:missing")).statusCode).toBe(404); + consolidateEvidence.mockResolvedValue({ status: "failed", warnings: ["Saved; retry indexing."] }); + const response = await app.inject({ method: "POST", url: "/workspaces/sales/evidence/consolidate" }); + expect(response.json()).toEqual({ status: "failed", warnings: ["Saved; retry indexing."] }); +}); diff --git a/backend/test/memory-routes.test.ts b/backend/test/memory-routes.test.ts new file mode 100644 index 00000000..4aa6dce3 --- /dev/null +++ b/backend/test/memory-routes.test.ts @@ -0,0 +1,81 @@ +import Fastify from "fastify"; +import { afterEach, expect, test, vi } from "vitest"; +import { memoryRoutes } from "../src/routes/memory.js"; +import { DEFAULT_SEMANTIC_RUNTIME } from "../src/workspaces/runtime-renderer.js"; +import type { ThtRunner } from "../src/tht/tht-runner.js"; +import type { PrincipalContext } from "../src/auth/principal.js"; + +const apps: ReturnType[] = []; +afterEach(async () => { await Promise.all(apps.splice(0).map(app => app.close())); }); +const id = "mem-11111111-1111-4111-8111-111111111111"; +const input = { family: "domain_clarification", subject: "Order", scope: "Sales" }; + +function setup(admin = true) { + const app = Fastify(); apps.push(app); + const principal: PrincipalContext = { issuer: "test", subject: "operator", roles: [admin ? "admin" : "user"], + permissions: admin ? ["memory.manage"] : ["session.use"], isAdmin: admin }; + app.addHook("preHandler", async request => { request.principal = principal; }); + const run = vi.fn(async () => ({ code: 0, stdout: JSON.stringify({ items: [], total: 0 }), stderr: "" })); + const bound = { runWithRuntimeSnapshot: run } as unknown as ThtRunner; + const withPrincipal = vi.fn(() => bound); + memoryRoutes(app, { runner: { withPrincipal }, runtime: DEFAULT_SEMANTIC_RUNTIME, + registry: { + list: async () => [{ id: "sales", commit: "a".repeat(40), blob: "b".repeat(40), snapshotPath: "/missing" }], + read: vi.fn().mockResolvedValue({ workspace: { workspace: { language: "it" } } }), + } }); + return { app, run, withPrincipal, principal }; +} + +test("list binds principal and workspace without requiring a DWH runtime", async () => { + const { app, run, withPrincipal, principal } = setup(); + const response = await app.inject("/workspaces/sales/memory?q=Order&page=2&column=id"); + expect(response.statusCode).toBe(200); + expect(withPrincipal).toHaveBeenCalledWith(principal); + const [args, snapshot] = run.mock.calls[0] as unknown as [string[], string]; + expect(args).toEqual(["memory", "admin", "--workspace", "sales"]); + expect(JSON.parse(snapshot)).toMatchObject({ action: "list", request: { q: "Order", page: 2, column: "id" } }); + expect(JSON.parse(snapshot).runtime.memoryLanguage).toBe("it"); +}); + +test.each([ + ["GET", ""], ["POST", ""], ["GET", "/pending"], ["GET", `/${id}`], + ["PUT", `/${id}`], ["DELETE", `/${id}`], ["POST", `/${id}/retry`], +] as const)("non-admin cannot %s %s", async (method, suffix) => { + const { app, run } = setup(false); + const response = await app.inject({ method, url: `/workspaces/sales/memory${suffix}`, + ...(method === "PUT" || method === "POST" && suffix === "" ? { payload: input } : {}) }); + expect(response.statusCode).toBe(403); expect(run).not.toHaveBeenCalled(); +}); + +test("unknown workspace and invalid input do not invoke the harness", async () => { + const { app, run } = setup(); + expect((await app.inject("/workspaces/missing/memory")).statusCode).toBe(404); + expect((await app.inject("/workspaces/sales/memory?page_size=101")).statusCode).toBe(400); + expect((await app.inject({ method: "POST", url: "/workspaces/sales/memory", + payload: { ...input, workspace_id: "foreign" } })).statusCode).toBe(400); + expect(run).not.toHaveBeenCalled(); +}); + +test("create preserves saved-but-unindexed outcome, update carries complete related changes", async () => { + const { app, run } = setup(); + run.mockResolvedValue({ code: 0, stdout: JSON.stringify({ id, saved: true, indexed: false }), stderr: "" }); + const result = await app.inject({ method: "POST", url: "/workspaces/sales/memory", payload: input }); + expect(result.statusCode).toBe(201); expect(result.json()).toEqual({ id, saved: true, indexed: false }); + await app.inject({ method: "PUT", url: `/workspaces/sales/memory/${id}`, payload: { + ...input, links: [{ target_id: id, meaning: "Related" }], dependencies: [{ database: "dwh" }], + } }); + const [, snapshot] = run.mock.calls[1] as unknown as [string[], string]; + expect(JSON.parse(snapshot)).toMatchObject({ action: "update", request: { id, card: { + links: [{ target_id: id, meaning: "Related" }], dependencies: [{ database: "dwh" }], + } } }); +}); + +test.each([["memory_not_found", 404], ["memory_conflict", 409], ["memory_unavailable", 503]])( + "maps %s without leaking stderr", async (code, status) => { + const { app, run } = setup(); + run.mockResolvedValue({ code: 1, stdout: JSON.stringify({ code, message: "sensitive detail" }), stderr: "secret" }); + const result = await app.inject("/workspaces/sales/memory"); + expect(result.statusCode).toBe(status); expect(result.json().code).toBe(code); + expect(result.body).not.toMatch(/secret|sensitive/); + }, +); diff --git a/backend/test/workspace-preprocessing-service.test.ts b/backend/test/workspace-preprocessing-service.test.ts index 44600a64..015c0f30 100644 --- a/backend/test/workspace-preprocessing-service.test.ts +++ b/backend/test/workspace-preprocessing-service.test.ts @@ -10,6 +10,28 @@ import { const roots: string[] = []; +test("Evidence consolidation runs only its stage and preserves blocked Catalog readiness", async () => { + const catalog = repository(); + const runChild = vi.fn(async () => ({ exitCode: 0, stdout: JSON.stringify({ status: "succeeded", counts: { documents: 2 } }), stderr: "" })); + const result = await service({ repository: catalog, runChild }).consolidateEvidence({ workspaceId: "catalog-workspace" }); + expect(result.status).toBe("succeeded"); + expect(runChild).toHaveBeenCalledOnce(); + expect(runChild.mock.calls[0]?.[0]).toMatchObject({ argv: ["preprocess", "evidence", "--consolidate", "--json", "-c", "/dev/fd/3"] }); + expect(catalog.beginPreprocessing).not.toHaveBeenCalled(); + expect(catalog.finishPreprocessing).not.toHaveBeenCalled(); + expect(catalog.clearPreprocessing).not.toHaveBeenCalled(); +}); + +test("Evidence index failure reports retry without changing Catalog state", async () => { + const catalog = repository(); + const runChild = vi.fn(async () => ({ exitCode: 1, stdout: JSON.stringify({ status: "failed", saved: true, error: "Saved; retry indexing." }), stderr: "private provider endpoint" })); + const result = await service({ repository: catalog, runChild }).consolidateEvidence({ workspaceId: "catalog-workspace" }); + expect(result.status).toBe("failed"); + expect(result.warnings).toEqual(["Saved; retry indexing."]); + expect(catalog.finishPreprocessing).not.toHaveBeenCalled(); + expect(JSON.stringify(result)).not.toContain("private provider"); +}); + afterEach(() => { roots.splice(0).forEach((root) => rmSync(root, { recursive: true, force: true })); }); @@ -254,6 +276,24 @@ test("semantic preflight failure is persisted before returning", async () => { ); }); +test.each(["refresh", "decide"] as const)("Evidence %s calls only the source worker and preserves Catalog readiness", async action => { + const catalog = repository(); + const runChild = vi.fn(async () => ({exitCode: 0, stdout: JSON.stringify({status: "succeeded", counts: {changed: 1}}), stderr: ""})); + const result = await service({repository: catalog, runChild}).evidenceSources({workspaceId: "catalog-workspace", action, + ...(action === "decide" ? {sourceId: "a".repeat(64), revision: "b".repeat(64), decision: "replace" as const} : {}), actor: "Curator"}); + expect(result.status).toBe("succeeded"); + expect(runChild).toHaveBeenCalledWith({configPath: expect.any(String), argv: expect.arrayContaining(["evidence", "sources", action, "-c", "/dev/fd/3", "--actor", "Curator"])}); + expect(catalog.beginPreprocessing).not.toHaveBeenCalled(); + expect(catalog.finishPreprocessing).not.toHaveBeenCalled(); + expect(catalog.clearPreprocessing).not.toHaveBeenCalled(); +}); + +test("failed source activation reports the saved decision for retry", async () => { + const result = await service({runChild: async () => ({exitCode: 1, stdout: JSON.stringify({status: "failed", saved: true, error: "Decision saved; retry indexing"}), stderr: "secret"})}) + .evidenceSources({workspaceId: "catalog-workspace", action: "decide", sourceId: "a".repeat(64), revision: "b".repeat(64), decision: "keep"}); + expect(result).toMatchObject({status: "failed", warnings: ["Decision saved; retry indexing"]}); +}); + test("clear invalidates Catalog readiness before clearing only derived worker data", async () => { const catalog = repository(); const runChild = vi.fn(async () => ({ diff --git a/backend/test/workspace-runtime-config-lease.test.ts b/backend/test/workspace-runtime-config-lease.test.ts index 43694147..707cebdf 100644 --- a/backend/test/workspace-runtime-config-lease.test.ts +++ b/backend/test/workspace-runtime-config-lease.test.ts @@ -1,4 +1,5 @@ import { execFile } from "node:child_process"; +import { createHash } from "node:crypto"; import { existsSync, mkdtempSync, @@ -202,7 +203,7 @@ test("deterministic operator leases are keyed by logical identity and stable acr workspaceSecretStore: f.workspaceSecretStore, }); - const suffix = first.inputFingerprint.slice(7, 23); + const suffix = createHash("sha256").update(first.inputFingerprint + "\n" + readFileSync(first.path, "utf8")).digest("hex").slice(0, 16); expect(second.path).toBe(first.path); expect(first.path).toBe(join( f.dataRoot, diff --git a/backend/test/workspace-runtime-handoff.test.ts b/backend/test/workspace-runtime-handoff.test.ts index 2e3fd01e..bc7772ae 100644 --- a/backend/test/workspace-runtime-handoff.test.ts +++ b/backend/test/workspace-runtime-handoff.test.ts @@ -242,6 +242,7 @@ test("separate runtime leases hand off byte-identical revision Evidence configs source_identity: "workspace://psd-clinical", }); expect(parse(firstYaml).evidence).toEqual({ + local_archive_root: join(f.registryConfig.root, "repo", "psd-clinical"), sources: [{ type: "filesystem", root: expectedRoot, diff --git a/backend/test/workspace-runtime-renderer.test.ts b/backend/test/workspace-runtime-renderer.test.ts index 4ddbde20..c878dc5a 100644 --- a/backend/test/workspace-runtime-renderer.test.ts +++ b/backend/test/workspace-runtime-renderer.test.ts @@ -192,6 +192,7 @@ test("renders filesystem Evidence below the immutable revision content root with expect(rendered.runtime_identity.workspace_revision).toBe(evidenceRevision); expect(rendered.evidence).toEqual({ schema_version: 2, + local_archive_root: "/srv/registry/repo/psd-clinical", sources: [{ type: "filesystem", root: `/srv/registry/snapshots/${evidenceRevision}/psd-clinical/evidence`, @@ -203,7 +204,7 @@ test("renders filesystem Evidence below the immutable revision content root with max_chunk_chars: 4_000, retain_published_generations: 3, }); - expect(yaml).not.toContain("/srv/registry/repo"); + expect(rendered.evidence.sources[0].root).not.toContain("/srv/registry/repo"); }); test("renders public HTTP Evidence with exact fractional-second timeouts and every policy limit", () => { @@ -223,6 +224,7 @@ test("renders public HTTP Evidence with exact fractional-second timeouts and eve })); expect(rendered.evidence).toEqual({ + local_archive_root: "/srv/registry/repo/psd-clinical", sources: [{ type: "http", urls: ["https://evidence.example.test/guide.md"], diff --git a/compose.yaml b/compose.yaml index 3719d3bd..0784cf15 100644 --- a/compose.yaml +++ b/compose.yaml @@ -13,6 +13,7 @@ services: THT_BIN: /opt/venv/bin/tht THT_DATA_ROOT: /data SETTINGS_FILE: /data/settings/settings.json + THT_EVIDENCE_HOST_REGISTRY_ROOT: ${THT_WORKSPACE_REGISTRY_ROOT:-} THT_MAINTENANCE_FILE: /data/settings/maintenance.json THT_WORKSPACE_REGISTRY_ROOT: /data/workspace-registry THT_WORKSPACE_GIT_REMOTE: ${THT_WORKSPACE_GIT_REMOTE:?set THT_WORKSPACE_GIT_REMOTE} @@ -99,7 +100,7 @@ services: image: thothii-core:local profiles: [catalog-maintenance] pull_policy: never - command: ["node", "/app/backend/dist/catalog/migrate.js"] + command: ["bash", "/app/docker/catalog-migrate.sh"] environment: THT_CATALOG_DB_HOST: catalog-db THT_CATALOG_DB_PORT: "5432" @@ -153,7 +154,7 @@ services: - type: volume source: workspace-registry target: /data/workspace-registry - read_only: true + read_only: false - type: volume source: sessions target: /data/sessions diff --git a/deploy/compose.evidence-host.yaml b/deploy/compose.evidence-host.yaml new file mode 100644 index 00000000..c8b18752 --- /dev/null +++ b/deploy/compose.evidence-host.yaml @@ -0,0 +1,20 @@ +# Optional local profile: expose the existing curator checkout on the installation host. +# Copy the existing registry repo to this path before enabling the override. Snapshot/state +# volumes remain unchanged. THT_EVIDENCE_HOST_REGISTRY_ROOT must be an absolute host path. +services: + core: + environment: + THT_EVIDENCE_HOST_REGISTRY_ROOT: ${THT_EVIDENCE_HOST_REGISTRY_ROOT:?set an absolute host registry path} + volumes: + - type: bind + source: ${THT_EVIDENCE_HOST_REGISTRY_ROOT:?set an absolute host registry path}/repo + target: /data/workspace-registry/repo + bind: + create_host_path: false + workspace-maintenance: + volumes: + - type: bind + source: ${THT_EVIDENCE_HOST_REGISTRY_ROOT:?set an absolute host registry path}/repo + target: /data/workspace-registry/repo + bind: + create_host_path: false diff --git a/deploy/compose.server.yaml b/deploy/compose.server.yaml index 02ffa967..b16939c2 100644 --- a/deploy/compose.server.yaml +++ b/deploy/compose.server.yaml @@ -45,7 +45,7 @@ services: - type: bind source: ${THT_WORKSPACE_REGISTRY_ROOT:?set THT_WORKSPACE_REGISTRY_ROOT} target: /data/workspace-registry - read_only: true + read_only: false - type: bind source: ${THT_DATA_ROOT:?set THT_DATA_ROOT}/workspace-secrets target: /data/workspace-secrets diff --git a/docker/catalog-migrate.sh b/docker/catalog-migrate.sh new file mode 100644 index 00000000..e1186d1f --- /dev/null +++ b/docker/catalog-migrate.sh @@ -0,0 +1,4 @@ +#!/usr/bin/env bash +set -euo pipefail +node /app/backend/dist/catalog/migrate.js +/opt/venv/bin/python -m tht.memory.migrate diff --git a/docker/core.Dockerfile b/docker/core.Dockerfile index f44e3f12..fd2f8eac 100644 --- a/docker/core.Dockerfile +++ b/docker/core.Dockerfile @@ -131,7 +131,7 @@ ENV PATH="/opt/venv/bin:/usr/local/bin:$PATH" \ HOME=/home/thoth COPY scripts/verify-line-endings.sh /usr/local/bin/verify-line-endings -COPY docker/core-entrypoint.sh docker/workspace-maintenance-entrypoint.sh docker/session-migrate.sh docker/ensure-pi-trust.mjs docker/embedding-model-init.sh /app/docker/ +COPY docker/core-entrypoint.sh docker/workspace-maintenance-entrypoint.sh docker/session-migrate.sh docker/catalog-migrate.sh docker/ensure-pi-trust.mjs docker/embedding-model-init.sh /app/docker/ COPY docker/smoke/core-smoke.sh /app/docker/smoke/core-smoke.sh RUN /usr/local/bin/verify-line-endings /app/docker \ && chmod +x /app/docker/core-entrypoint.sh /app/docker/workspace-maintenance-entrypoint.sh /app/docker/session-migrate.sh /app/docker/embedding-model-init.sh /app/docker/smoke/core-smoke.sh diff --git a/docs/adr/0018-use-postgres-for-memory-and-qdrant-for-retrieval.md b/docs/adr/0018-use-postgres-for-memory-and-qdrant-for-retrieval.md new file mode 100644 index 00000000..be8b5e1b --- /dev/null +++ b/docs/adr/0018-use-postgres-for-memory-and-qdrant-for-retrieval.md @@ -0,0 +1,59 @@ +--- +status: accepted +date: 2026-09-08 +--- + +# Use PostgreSQL for Memory and Qdrant for retrieval + +The new Memory module uses the installation's existing PostgreSQL service as the +authority for Memory Cards, their links, and structured schema dependencies. +Qdrant holds a rebuildable search projection. This replaces the JSONL registry and +allows related administrative changes to be coordinated in one database transaction. +The owner accepted this direction in Q9 of the Memory design interview; implementation +is pending. + +Memory uses its own tables and remains a separate module from the Metadata Catalog. +Sharing the PostgreSQL service does not transfer Memory ownership to Database +management or make Evidence and Memory one canonical domain. + +## Retrieval and graph + +The owner also accepted hybrid semantic/lexical Qdrant retrieval with scope filters +and explicit card links traversed in core. Both belong to the planned first version; +there is no dedicated graph database or external Memory framework. + +This extends [ADR 0017](0017-separate-reference-vectors-from-runtime-memory.md): +the separate reference and memory collections remain, while Memory gains sparse +lexical indexing in addition to dense vectors. Memory projections become rebuildable +from the module's PostgreSQL authority. Preprocessing Clear still preserves Memory; +this decision does not add it to the preprocessing cleanup scope. + +Links support discovery. Finding a card through a link does not approve its use. +The core proposes links with the cards for the same final review; Administration +provides manual creation, editing, and deletion. Deleting a card removes its incident +links without deleting the other linked cards. + +## Considered options + +- Retaining JSONL preserves the current storage format, but leaves coordinated + card/link/dependency mutations and concurrent administrative writes to application code. +- Using Qdrant as the sole authority is a viable alternative for record storage and + retrieval. PostgreSQL is preferred for the coordinated mutations of the new module, + with Qdrant reserved for its search projection. +- Adding a dedicated graph database would introduce another service; the selected + bounded traversal can be implemented in core over persisted links. + +## Consequences + +The module needs a persistence contract, PostgreSQL schema, and explicit propagation +of additions, updates, and deletions to Qdrant. Choosing PostgreSQL does not make this +propagation atomic across both systems: failures, retries, and invalidation of stale +search content must be handled and tested. An index rebuild uses the original card +content and cannot resurrect deleted cards. + +The decision does not introduce Memory revision history or require compatibility with +existing development sessions. Evidence authoring and publication remain governed by +their own contract until the Evidence management project defines its evolution. + +The [Memory management project](../plans/2026-09-08-memory-management.md) records the +approved behavior, scope, and integration work. diff --git a/docs/adr/0019-author-evidence-in-app-with-automatic-activation.md b/docs/adr/0019-author-evidence-in-app-with-automatic-activation.md new file mode 100644 index 00000000..8f94711b --- /dev/null +++ b/docs/adr/0019-author-evidence-in-app-with-automatic-activation.md @@ -0,0 +1,152 @@ +--- +status: accepted +date: 2026-09-08 +--- + +# Maintain local Evidence with external editors and manual consolidation + +Evidence management must provide complete access to registered Evidence and a +maintenance path, as accepted in Q11–Q15 and revised during simplification. The owner subsequently +clarified that a context specialist writes drafts independently of the installation +and without PostgreSQL access; the system refines and stores them locally for use +and maintenance. Implementation is pending. + +The storage mechanics of Q11 were revisited in the +[simplification review](../plans/2026-09-08-memory-evidence-simplification-review.md). +The original choice combined the workspace repository as Evidence authority with +application-managed Git writes. The owner accepted the separation of external draft +files/repositories from a durable local canonical file archive and removes +commit/push from normal CRUD. The management surface must expose discoverable, +editable Markdown for domain specialists, without JSONL editing. The owner chose +external editors for release 0: their preferred Mac/PC editor or vim/nano on the +server. No web editor or content form is required. The page keeps browsing, +filtering, and detail, with actual host file paths and maintenance instructions. +The owner also requested manual consolidation followed by manual Git diff, commit, +and push, relying on operator discipline rather than automation. The owner accepted +the remaining simplifications and requested explicit clarification of the core +format change and the simple terminal-based Git check. The original application Git writer +is not the final implementation prescription. A suggestion +to move Evidence authority into PostgreSQL was withdrawn after the clarification. +Memory remains a distinct module with its own PostgreSQL authority under +[ADR 0018](0018-use-postgres-for-memory-and-qdrant-for-retrieval.md). + +## Manual consolidation and Git follow-up + +The external-editor decision replaces the form's single Save with saving files and +running one consolidation command. The command checks expected structure, required +fields, types, identifiers, references, and provenance. It reports the affected file +and the necessary correction. Invalid input stops activation; valid input updates +derived metadata, the corpus, and the Evidence index. New and deleted files use the +same path. An inaccessible archive must not be interpreted as deleted Evidence. +Structural validation does not prove semantic correctness or rerun model refinement +to rewrite manually edited rules. + +The core reads only successfully consolidated active content, never the editable +working files directly. Candidate validation and activation must preserve the last +valid corpus on failure, or report unavailability if its integrity cannot be +guaranteed. Unconsolidated edits must not mix new text with old retrieval data. + +Consolidation reports full local success only once the updated content is available +to the core. If persistence succeeds but activation fails, the command reports the +partial outcome and can be rerun without duplicating units or losing edits. +The operation does not imply atomicity across +authoritative storage and Qdrant; their failure and recovery behavior requires +explicit implementation. An interrupted operation must not leave a partial corpus +advertised as ready for subsequent use. + +The maintained archive is a persistent Git working tree containing the local +Evidence; it may reuse the workspace repository. Draft sources and curated units +remain distinct even when in the same repository. Setup and the page identify the +repository, working tree, actual host paths, branch, and configured remote. A +rebuildable runtime snapshot is not the editing location. + +Before consolidation, the operator checks Git status, including inspection of new +untracked files. Normal terminal Git diff commands are available for line-level +inspection when needed; no custom viewer or mandatory double review is required. +After consolidation, the operator stages the Evidence and required metadata changes, +commits, and pushes using normal Git commands. No watcher, automatic Git writes, retrying push, pull, merge, +or dedicated web execution control is required. Missing credentials, unconfigured +upstreams, and Git conflicts are handled by the operator. + +Consolidation activates local changes before commit/push. If the operator omits the +Git steps or push fails, the local update remains effective while transfer to the +remote is incomplete. Pushing does not automatically update other installations. +The system relies on operator discipline to complete the sequence; it does not add +a separate editorial publication state machine or require a second reviewer. + +Approved conflict repairs from the core still call the same persistence and +activation service directly. They do not require an external editor or manual +consolidation before the session can use its own approved correction. Their local +file changes are included in the operator's subsequent Git maintenance. + +## Administration without concurrent core work + +The owner's simplification instruction replaces the earlier Q12 requirement to +refresh open sessions after administrative edits. Core activity can be assumed +absent during administration, or its overlap can be ignored. No dedicated live +update, session notification, restart, maintenance mode, or reader coordination +is required. Completed administrative changes apply to subsequent work. + +Deliberate writes from the workflow itself still exist: approved conflict repairs +and the final Memory summary use the same persistence services and handle their +outcome before proceeding. In particular, a session must be able to use its own +approved Evidence correction. This does not require updating all other sessions +or rewriting previously approved decisions, artifacts, or SQL. + +## Manual corrections survive source updates + +When an updated source contradicts an administrator's correction, the saved manual +Evidence remains active. Evidence management shows the conflicting content for an +explicit decision and subsequent save. The source's newer content does not +automatically override the correction. The owner accepts that the manual rule can +remain in use until that comparison is resolved. + +Deleted Evidence must not silently reappear after preparation or reindexing. The +authoring implementation must preserve both manual corrections and suppression of +deleted units through source refreshes. The current protection for uncommitted Git +changes is insufficient: it does not preserve already committed manual corrections +against regeneration, and retirement currently removes the unit without recording +suppression for later preparation. + +## Manual creation and explicit source refresh + +Administrators can create Evidence without providing an external document. In R0, +they add Markdown using the documented example and consolidate it. The application +records a manual declaration as the managed source of the current statement. +Explicit consolidation approves that declaration; it does not +claim independent documentary verification. + +A correction that changes a unit's meaning uses the same explicit manual origin. +The original document remains linked for provenance and source-change detection, +but its excerpt is not presented as support for a rule it does not contain. The +canonical contract must distinguish the current supporting source from the original +document rather than silently retaining outdated support metadata. + +External sources are reacquired only when an administrator requests a source +refresh. Ordinary saves and session lookups use the acquired local content; they +do not poll sources or fetch their current versions. A remote change is therefore +detected at the next requested refresh, when the manual-precedence rule applies. +The revised storage recommendation reads the specialist's repository as an input; +ordinary local edits do not write back to it or refresh its content automatically. + +## Implementation consequences + +The existing [Evidence lifecycle](../evidence.md) and +[Workspace Evidence v3 contract](../contracts/workspace-evidence-v3.md) require +explicit evolution for editable local Markdown, manual consolidation, and manual +precedence. Source provenance and coherent canonical metadata remain requirements; +manual changes must not be presented as statements supported by an unrelated source +excerpt. Under the accepted local-file direction, curated files are primary +installation data preserved by Clear and backed up separately from rebuildable +indexes. Visible Markdown content must become authoritative for editing; operators +must not maintain hidden duplicate text, hashes, or manifest entries themselves. +This is a core Evidence contract change: parser, renderer, authoring, validation, +normalization, and their integration with indexing and recall must be adapted +together. E1 includes a new Curated unit format version, conversion of existing +Evidence, reindexing, and end-to-end verification that visible edits reach core +consumption. Preserve the internal typed model where possible; the change does +not require redesigning the NL-to-SQL workflow phases. +Local edits do not automatically change external drafts or another installation; +Git versioning and transfer occur through the explicit manual follow-up. +The [Evidence management project](../plans/2026-09-08-evidence-management.md) +records the implementation sequence and acceptance checks. diff --git a/docs/architecture/authentication.md b/docs/architecture/authentication.md index e9eb3a5a..f7e77731 100644 --- a/docs/architecture/authentication.md +++ b/docs/architecture/authentication.md @@ -33,11 +33,16 @@ The production role expansion from `backend/src/auth/config.ts` is exact: | Role | Permissions | |---|---| | `user` | `session.use` | -| `admin` | `session.use`, `session.read_all`, `session.manage_all`, `settings.manage`, `workspace.manage`, `workspace.secrets.manage`, `database.manage`, `pi.manage`, `auth.diagnostics.read` | +| `admin` | `session.use`, `session.read_all`, `session.manage_all`, `settings.manage`, `workspace.manage`, `workspace.secrets.manage`, `database.manage`, `memory.manage`, `evidence.manage`, `pi.manage`, `auth.diagnostics.read` | `admin` therefore includes the ordinary `session.use` permission. No other role or permission label is part of the production catalog. +After validating a browser session, the backend expands its roles through the current permission +catalog on every request. The permissions saved at login are a historical snapshot, so existing +administrator sessions can use newly deployed administration features without signing in again. +Session expiry, revocation and local-user role validation still apply before role expansion. + OIDC is provider-neutral at the browser protocol boundary. Authorization Code + PKCE, issuer, signature, audience, expiry, state, and nonce are validated before a principal is created. Authentik is the first certified group-catalog adapter, not a special browser login mode. diff --git a/docs/architecture/components.md b/docs/architecture/components.md index 15c0976a..47906fb6 100644 --- a/docs/architecture/components.md +++ b/docs/architecture/components.md @@ -7,7 +7,9 @@ This page complements the [architecture overview](overview.md) with the module s The frontend communicates with the backend through REST and SSE. The backend does not own session persistence: it starts Pi, invokes the `tht` CLI, and forwards events. It does own the separate installation-local database catalog. The harness contains the workflow, the Python CLI, and -adapters for the DWH and vector store. +adapters for the DWH and vector store. Its Memory module also owns the authoritative +PostgreSQL archive of cards, links, dependencies and pending Qdrant projections. +Administrative API calls use the same harness service as workflow producers and recall. ```mermaid flowchart LR @@ -20,6 +22,7 @@ flowchart LR THT --> FS["Sessions and artifacts\nworkspace repository"] THT --> DWH["DWH\nread-only"] THT --> VDB["Qdrant / vector store"] + THT --> MEM["thoth_memory\nPostgreSQL Memory archive"] BE --> CFG["settings.json\nworkspace + thinking"] BE --> MODELS["generated runtime catalog\nfrom installation YAML"] BE --> CAT["catalog-db\nPostgreSQL + Kysely"] diff --git a/docs/contracts/archive-repair.md b/docs/contracts/archive-repair.md new file mode 100644 index 00000000..2d4d0d58 --- /dev/null +++ b/docs/contracts/archive-repair.md @@ -0,0 +1,65 @@ +# Session corrections to Memory and Evidence + +`reviewer_archive_repair` presents one to five closed alternatives for a conflict. +Each alternative updates one existing Memory Card or one existing local Evidence unit, +showing complete current and resulting content. The reviewer chooses one alternative, +rejects all as inadequate, or asks for reformulation. A correction does not approve or +advance the session phase; the ordinary gate still reviews its use in the current question. + +The harness coordinator `tht/archive_repair.py` joins two independent domains. It uses +Memory's PostgreSQL repository, workspace mutation lock and projection service, and the +same local Evidence save/consolidation boundary as administration. Evidence activation +runs the existing corpus pipeline without reacquiring configured external sources. +No database binding, schema configuration, other archive, or Git repository is changed. + +## Authority and recovery + +Migration `004_archive_repairs.sql` stores session-bound proposals and receipts in +`thoth_memory.archive_repairs`, isolated by workspace RLS. Each receipt captures the +session decision context, current target revision, complete resulting content, selected +choice, acting principal and publication outcome. The preparation command changes no +archive content. Application accepts only an option ID from the persisted proposal. + +Memory writes its new revision and receipt in one SQL transaction. Projection failures +leave a durable pending operation; retry propagates the saved revision. A later card +edit invalidates replay of the correction. + +Evidence records the exact choice before writing its canonical file. Recovery accepts +either the reviewed original revision or the already-written approved result; it never +overwrites a different intervening correction. Before a first proposal, the editable +checkout must match its active snapshot. Other curated files are fingerprinted and +checked again before application/retry so a session decision cannot publish unrelated +external edits. A crash after replacement is recoverable by the same receipt. Original +document provenance is retained as the lineage of the manual correction. + +The gate displays these outcomes: + +| Status | Meaning | +| --- | --- | +| `proposed` | Waiting for a human choice; no content saved | +| `rejected` | All alternatives declined; reformulation required | +| `applying` | Choice recorded; file write or candidate recovery still required | +| `pending_activation` | Content saved; index activation incomplete | +| `active` | Saved correction matches the currently active revision | +| `superseded` | The target was removed or changed after the saved correction | + +The reviewer may retry a pending correction or continue the current question while +leaving activation explicitly pending. The latter is not persistent-resolution success. +`repair-show` refreshes target status; `repairs` lists the historical receipt status. +The final Memory summary must not create a duplicate of a correction already saved here. + +## Authorization and commands + +Inspection/preparation requires an accessible open session in the same workspace. +Applying either archive correction requires an administrator. Browser responses are +also checked against the responding principal's `memory.manage` or `evidence.manage` +permission and the Pi runtime owner. An administrator using another principal's runtime +must resume it under their own account first, preserving truthful receipt attribution. +Non-administrators may reject proposals or continue question review without changing +the archives. The Pi shell guard blocks `repair-apply`; only the human gate invokes it. + +The Python workflow CLI provides `memory repair-target`, `repair-prepare`, `repair-show`, +`repair-apply`, and `repairs`. These are workflow integration commands, not new native +installation commands. All accept `--session` and command-local `-c`; JSON output remains +machine-readable. Proposal bodies are bounded to 1 MB. Durable receipts support both +filesystem and server session storage without a persisted chat transcript. diff --git a/docs/contracts/curated-evidence-v4.md b/docs/contracts/curated-evidence-v4.md new file mode 100644 index 00000000..95a94305 --- /dev/null +++ b/docs/contracts/curated-evidence-v4.md @@ -0,0 +1,248 @@ +# Editable Curated Evidence v4 + +Curated unit v4 makes the visible Markdown body authoritative. It uses the existing +typed Evidence payloads and stable identifiers. Workspace descriptor v4 and Evidence +descriptor v1/v2 are separate version numbers. + +E1 implements the format, explicit conversion, local archive and consolidation API. +E2 connects that API to the installed consolidation command, Evidence management +page and runtime source selection. The local PSD preview now uses 35 converted units. + +## Write or edit a file + +Place units in `/evidence/curated//.md`. Keep the existing +`id` when editing. The directory must match `kind`. A new manual unit needs no external +document, hash or encoded metadata: + +```markdown +--- +schema_version: 4 +id: evidence:order-key +kind: domain +language: en +purposes: [sql_generation] +applies_to: + tables: [sales.orders] +--- + +# Order key + +## Rule + +Join orders using the order number, financial year and company. +``` + +Required metadata is `schema_version`, `id`, `kind`, `language`, and a nonempty +`purposes` list. Purposes are `disambiguation`, `rewriting`, `schema_linking`, and +`sql_generation`. Optional `applies_to` contains `concepts`, `tables`, and `columns`. +Tables use `schema.table`; columns use `schema.table.column`. + +The first H1 is the title. H2 headings identify the payload fields below. Heading +spelling follows `language`: Italian for `it` and its regional variants, English +otherwise. Keep structural headings when editing the text below them. + +| Kind | English field headings | Italian field headings | +| --- | --- | --- | +| `domain` | Rule | Regola | +| `glossary` | Definition, Synonyms, Variants | Definizione, Sinonimi, Varianti | +| `enum` | Column, Values | Colonna, Valori | +| `example` | Question, Interpretation | Domanda, Interpretazione | +| `mapping` | Concept, Tables, Columns | Concetto, Tabelle, Colonne | +| `normalization` | Input, Output, Rule | Input, Output, Regola | +| `formula` | Concept, Columns, SQL | Concetto, Colonne, SQL | +| `reference` | URL, Label, Description | URL, Etichetta, Descrizione | + +List fields use one `- value` per line. Empty optional lists may be omitted. Quoted +JSON strings within bullets preserve unusual or multiline values during conversion. +Enum values use `### "stored value"`, followed by their meaning; `### ""` represents +an empty stored value. Formula SQL uses a fenced `sql` block containing one PostgreSQL +expression. Whole queries and mutation statements remain invalid. + +Nested prose headings and fenced examples are supported inside text fields. An H2 +matching a field heading is structural outside a code fence. Duplicate fields, missing +required fields, duplicate metadata keys and malformed payloads are rejected. There +is no second title or payload in frontmatter and no hidden authoritative rule text. +Conversion fails explicitly if a legacy payload cannot be represented losslessly. + +## Current provenance and original source + +The host supplies document provenance when refining source material: relative source +path, normalized source hash and exact supporting excerpts. A manual unit can omit +`provenance`. Consolidation records `kind: manual` and the supplied curator identity. + +A visible change to a previously recorded unit becomes a manual declaration. If that +unit originated from a document, its former document provenance is retained under +`original`. The original excerpts establish lineage; they do not assert that the +source contains the new wording. Retrieval carries this distinction through typed +metadata. Curators edit content; the service updates managed provenance. + +Unchanged document declarations still require matching source bytes and excerpts. +Replacing a source requires an explicit refresh through the E3 source-review path. Ordinary preparation +is blocked on an initialized local archive, preventing regenerated source material +from overwriting corrections or restoring deletions. Legacy `resolve` is likewise +blocked there; local file corrections and the archive API own those changes. + +## Persistent archive and activation + +```text +/ + .evidence-archive.lock + evidence/ + source/ # acquired original documents + curated//*.md # editable primary content + local-manifest.yaml # derived declarations and deletion records + .local/ + state.yaml # baseline, pending and active revisions + snapshots// # immutable units, source bytes and manifest +``` + +Keep the complete Evidence tree and its managed metadata in backups and the operator's +Git review. Historical source bytes and deletion records are needed to preserve manual +care across subsequent imports. The separate corpus cache and Qdrant index are derived. +The lock file only coordinates local service operations. + +`LocalEvidenceArchive.initialize()` records the pre-edit baseline without activation. +`consolidate(actor=..., activate=...)` validates files, derives provenance, records +deletions and source suppression, and creates an immutable candidate. The activation +callback receives that snapshot and must raise if indexing is blocked or fails. Only +successful activation advances `active_snapshot()`. Without a callback the result is +explicitly `pending_activation`; saving a file alone never changes this pointer. + +The same candidate can be retried after an index failure. Interrupted managed writes +are replayed only when the operator's file bytes have not changed. A missing curated +directory is an availability error, not proof that all units were deleted. Removing +unit files from an accessible directory records deletion without deleting their sources. +The API also provides `get`, revision-checked `save`, and revision-checked `remove` for +future deliberate workflow corrections. Concurrent external edits produce conflicts. + +These boundaries are exercised with the existing corpus pipeline and real Qdrant. +An initialized installation reads only its active local snapshot, including during +ordinary preprocessing. Unconsolidated edits are visible in Administration but do +not enter retrieval. Before initialization, the pinned repository source still works. + +## Administration and installed command + +Open **Administration → Evidence management**. This independent page requires the +`evidence.manage` permission and no active session. It shows complete units, source +lineage, review items, file errors, filters, and changes relative to the active snapshot. +Edit the displayed Markdown path using an external editor. Refresh files to inspect +the result, then run the command shown by the page: + +```sh +tht --installation /absolute/path/thothii-installation.yaml \ + workspace evidence consolidate --workspace psd-clinical +``` + +The installed command accepts `--json`. It validates and activates Evidence using +the existing corpus pipeline and embedding service. It does not scan the DWH, run +full workspace preprocessing, change Catalog/Schema readiness, or execute Git. +If indexing fails after saving, the archive retains the candidate for retry and the +previous active revision remains selected. Validation errors identify corrections +to make in the files. Structural and review checks still apply; initialized archives +do not require the legacy repository's fixed retrieval-evaluation fixture, whose +expected IDs would otherwise prevent deliberate local deletions. + +The canonical workspace root is `/repo/`. +Snapshots and the derived corpus are separate. For a registry in a Docker volume, +copy its existing `repo` to a persistent host directory before enabling +`deploy/compose.evidence-host.yaml`; set `THT_EVIDENCE_HOST_REGISTRY_ROOT` to that +directory and include the override in the installation descriptor. Core and +workspace-maintenance must mount the same checkout. Keep registry state/snapshots +on their existing volume. The page only presents a host path when configured; it +does not label an internal container path as a usable editor path. + +After successful consolidation, inspect and commit the complete workspace Evidence +tree, including managed manifests, snapshots, and deletion records, then push manually. +The page provides quoted POSIX-shell examples for status, diff, add, commit, and push. +Do not commit only the edited Markdown. Ordinary Git operations remain the operator's +responsibility. E3 source decisions use this same activation boundary. + +**Clear** removes derived Reference/corpus data while retaining editable files, +snapshots and Memory. It still requires full workspace preprocessing to recreate +Reference/Schema readiness; Evidence consolidation does not satisfy that gate. + +## Import drafts and refresh sources + +Place externally authored Markdown drafts in `/evidence/incoming/`. +The specialist needs no installation account or database access to write a draft. +Copying the draft to the installation and choosing **Import or refresh sources** in +Evidence management explicitly starts acquisition and refinement. The equivalent +installed command is: + +```sh +tht --installation /absolute/path/thothii-installation.yaml \ + workspace evidence refresh --workspace psd-clinical +``` + +This reads local `incoming/**/*.md` and original `source/**/*.md` files, excluding +managed `source/acquired/` versions. It also reads configured HTTP and S3 sources +through the existing read-only adapters and network policies. After local archive +initialization, filesystem descriptors use this local authoring tree; ordinary +runtime/preprocessing never fetches remote source changes. No source-server write +credential is needed. The existing Pi authoring refiner runs once for each changed +document, without session state or tools. Unchanged source hashes skip refinement. + +All acquisitions and proposals must succeed before the new comparison set is +recorded. An access or refinement failure preserves previous comparisons and active +Evidence. A source absent from a successful discovery is marked missing and never +treated as permission to delete units. Restore an accidentally missing local original +file before consolidating its document-derived units, or explicitly retire those units. + +Each changed source has a durable comparison showing current local units, complete +proposed content, supporting excerpts, and IDs that replacement would retire: + +- **Keep local Evidence** retains the current wording as a manual declaration, with + its original documentary lineage preserved. The changed source is acknowledged; + the next unchanged refresh does not reopen that decision. +- **Use proposed Evidence** adopts the displayed proposal and explicitly retires the + displayed omitted IDs. The source version and provenance change together. + +Both choices save and activate through the same local consolidation/index pipeline. +Review items block adoption; correct the original draft and refresh, or keep local +content. There is no automatic merge based on a model's semantic conflict assessment. +Any change to an affected curated file invalidates the comparison and requires a new +refresh. If indexing fails after the decision is saved, use **Retry saved decision**; +this reuses acquired content without fetching sources again. An intervening external +edit is never silently overwritten by recovery. + +Headless operators can make the same decision using the source ID and comparison +revision from the local source registry or administration response: + +```sh +tht --installation /absolute/path/thothii-installation.yaml \ + workspace evidence decide --workspace psd-clinical \ + --source-id <64-hex-source-id> --revision <64-hex-comparison-revision> \ + --decision keep +``` + +Use `--decision replace` to adopt the proposal. All installed commands accept `--json`. +The Python `evidence sources` worker is internal to this installed command/API surface. + +`evidence/.local/sources.json` stores source identities, comparisons and retry journals. +`evidence/.local/acquisitions/` preserves original acquired bytes and credential-free +remote provenance. Adopted normalized documents live under +`evidence/source/acquired//.md`. Versioned paths let a new +document and an older manual declaration's original source coexist. Include all of +these files in the existing manual Git/backup sequence. Do not edit managed acquired +versions: edit the original local draft or refresh its remote origin. + +Deleted IDs remain reserved. Once a source has had a curated deletion, fresh model +IDs from that source are conservatively suppressed as well: changing an ID must not +restore retired knowledge. Existing surviving IDs can still receive reviewed updates; +deliberate new knowledge can be written as a manual Evidence file. Refresh is bounded +to 200 documents and 100 MiB per operation, in addition to each adapter's limits. + +## Convert an existing workspace + +The existing workflow CLI command `tht evidence migrate ` converts +unit versions 1–3 to 4 deterministically and initializes the archive baseline. It +preserves IDs, typed content, provenance and review items, with no model call or Git +commit. It does not activate a local index. Review items still block consolidation. +The command's existing Git-worktree path check remains in effect. + +The first installed consolidation performs this conversion automatically when the +legacy manifest is present, then validates and activates the result. Preserve the +existing checkout in backups before upgrading. The E1 validation used an isolated +copy; E2 also converted and indexed all 35 units on the running local preview. +See the [E1 validation report](../plans/2026-09-09-evidence-e1-validation.md) and +[E2 validation report](../plans/2026-09-09-evidence-e2-validation.md). diff --git a/docs/contracts/workspace-evidence-v3.md b/docs/contracts/workspace-evidence-v3.md index a5c389c7..caff5d26 100644 --- a/docs/contracts/workspace-evidence-v3.md +++ b/docs/contracts/workspace-evidence-v3.md @@ -6,8 +6,9 @@ When present, `evidence` is strict: it contains `source` and a defaulted strict source variant and the policy reject unknown keys. The version numbers are intentionally separate: the Evidence descriptor supports v1/v2, while the -latest Curated Evidence Unit format is v3. There is no Evidence descriptor v3/v4 and no Curated -Evidence Unit v4. +latest Curated Evidence Unit format is [v4](curated-evidence-v4.md). There is no Evidence +descriptor v3/v4. The editable local archive is implemented in E1; installation integration +is the next increment, E2. ```mermaid flowchart LR @@ -67,9 +68,11 @@ ignore it. Domain rules also retain their exact canonical text in an invisible ` comment while presenting long prose as paragraphs, labelled subsections, and semicolon-derived lists. Runtime chunking reads the parsed canonical rule, not this review-only presentation. -Newly prepared units use v3. `tht evidence migrate ` upgrades v1 and v2 units and -canonicalizes an older v3 presentation locally without a model call, commit, publication, or -semantic change. +The representations above are legacy conversion inputs. Newly prepared units use editable v4: +short YAML metadata, a visible H1 title and typed H2 payload fields, with no hidden content copy. +`tht evidence migrate ` converts v1–v3 to v4 and initializes a local archive +baseline without a model call, commit, activation, or semantic change. See the +[v4 editing and consolidation contract](curated-evidence-v4.md). ### Example: filesystem diff --git a/docs/contracts/workspace-preprocessing-cli.md b/docs/contracts/workspace-preprocessing-cli.md index 7bd21d99..d8613a90 100644 --- a/docs/contracts/workspace-preprocessing-cli.md +++ b/docs/contracts/workspace-preprocessing-cli.md @@ -14,12 +14,30 @@ tht --installation /thothii-installation.yaml workspace preprocess run tht --installation /thothii-installation.yaml workspace preprocess clear --workspace [--json] + +tht --installation /thothii-installation.yaml workspace evidence consolidate + --workspace [--json] + +tht --installation /thothii-installation.yaml workspace evidence refresh + --workspace [--json] + +tht --installation /thothii-installation.yaml workspace evidence decide + --workspace --source-id <64-hex> --revision <64-hex> + --decision keep|replace [--json] ``` There are no public partial commands for DWH introspection, LSH, FK suggestions, schema indexing, -Evidence indexing, or Qdrant rebuild. `preprocess run` does not accept `--resume`, `--dry-run`, a +or Qdrant rebuild. The separate Evidence curation command validates editable local files and +activates only their index; it does not satisfy workspace preprocessing readiness or execute Git. +See [Curated Evidence v4](curated-evidence-v4.md#administration-and-installed-command). +Source refresh acquires and refines only on explicit request, saving comparisons without +changing active Evidence. Source decisions activate through the same Evidence-only +pipeline and preserve Catalog readiness. Their envelopes reject arbitrary URLs, paths +and extra flags; source locations and credentials come from installation configuration. +`preprocess run` does not accept `--resume`, `--dry-run`, a generation identifier, or a rollback option. Re-running it replaces the preceding derived output. -`preprocess clear` removes all replaceable preprocessing output and preserves runtime memory. +`preprocess clear` removes all replaceable preprocessing output and preserves runtime memory +and the canonical local Evidence archive. Full preprocessing remains required afterward. ## Sources of truth @@ -29,8 +47,10 @@ generation identifier, or a rollback option. Re-running it replaces the precedin - PostgreSQL Metadata Catalog is the sole database authority. It owns the workspace/database association, installation-local binding, tables, columns, descriptions, sensitivity flags, physical foreign keys, and active logical relationships. -- The workspace Git revision remains authoritative for Evidence. Evidence Descriptor v1/v2 and - Curated Evidence Unit v3 are unchanged; there is no Evidence v4. +- Evidence Descriptor v1/v2 still configures the initial source. Initialized archives use + editable Curated Evidence Unit v4 and the last successfully activated local snapshot. + Ordinary preprocessing never imports unconsolidated working-tree edits. Before initialization, + the pinned workspace Git source remains supported. No metadata is imported from legacy workspace YAML or `physical.yaml`/`annotations.yaml`. @@ -98,9 +118,8 @@ generation produced for another database or revision. `workspace preprocess clear` is intentionally narrower than deleting all semantic data. It: -1. copies any `memory` and `solved_question` records still present in the pre-split workspace - collection to `-memory`, then retires that legacy collection (a normal preprocessing - write performs the same one-time cutover if clear was not invoked first); +1. leaves `-memory` unchanged; Memory projections are reconstructed only from + the authoritative PostgreSQL archive, never imported from legacy vector payloads; 2. deletes `-reference`; 3. removes the active LSH generation, Evidence corpus, private Catalog snapshot, and derived job checkpoints; diff --git a/docs/evidence.md b/docs/evidence.md index 9265101a..5c47124c 100644 --- a/docs/evidence.md +++ b/docs/evidence.md @@ -2,7 +2,30 @@ This page describes the complete workspace Evidence lifecycle: where original material lives, how curated units are produced, when they become available at runtime, and what the author and reviewer are responsible for. -## Publication rule +## Editable local Evidence + +E1 adds [Curated Evidence v4 and a persistent local archive](contracts/curated-evidence-v4.md). +The visible title and payload fields are authoritative Markdown. New manual units need no +external source; consolidation records their curator and distinguishes later corrections +from original documentary provenance. The core archive API creates immutable candidates +and advances its active pointer only after successful indexing. + +E2 adds **Administration → Evidence management**, actual host file paths, complete +browsing and filtering, and the installed `tht workspace evidence consolidate +--workspace ` command. Edit files externally, consolidate to activate them, then +review and run Git manually. Runtime consumes only the active local snapshot. +The [v4 contract](contracts/curated-evidence-v4.md#administration-and-installed-command) +describes host mounting, first conversion, failure recovery and Clear behavior. +The local PSD preview has 35 converted and indexed units. E3 adds explicit import/refresh +from local drafts and configured HTTP/S3 sources, durable current/proposed comparisons, +and keep/replace decisions with activation and retry. See +[Import drafts and refresh sources](contracts/curated-evidence-v4.md#import-drafts-and-refresh-sources). +Workflow gate corrections remain the subsequent shared increment, X1. + +## Existing repository publication path + +The remainder describes the legacy, uninitialized repository source path. Initialized +v4 local archives use the lifecycle above; manual declarations need no source document. The workspace repository is the versioned source. ThothII reads it, validates it, and publishes an atomic generation. It does not modify, commit, or push the author's repository. @@ -46,7 +69,33 @@ The legacy configuration may expose `source_root`, such as `${THT_DOCS_ROOT}` or HTTP and S3 are separate adapters. They do not use the filesystem structure `source/` and `curated/`, but they must still provide stable provenance, without credentials in URIs, under the adapter-specific contract. -## What a curated unit must contain +## What an editable curated unit contains + +New preparation produces unit schema v4. The first H1 contains its title; documented H2 +sections contain the typed payload. A minimal manual unit is: + +```markdown +--- +schema_version: 4 +id: evidence:order-key +kind: domain +language: en +purposes: [sql_generation] +--- + +# Order key + +## Rule + +Join orders using the order number, financial year and company. +``` + +Use `tht evidence migrate ` for deterministic legacy conversion. The +conversion preserves typed content and initializes an archive baseline; it does not +activate the local corpus. See the [v4 contract](contracts/curated-evidence-v4.md) for +all eight kinds, provenance, file layout and the E1/E2 boundary. + +## Legacy v3 representation Canonical Curated Evidence v3 hides canonical machine metadata in an HTML comment and renders the whole review surface as real Markdown. GitHub therefore shows no frontmatter table. The body layout @@ -94,7 +143,7 @@ La fascia pediatrica comprende i pazienti con età inferiore a 18 anni. The actual files contain invisible `tht:` comments for canonical metadata and typed-field boundaries. Removing, duplicating, or desynchronizing them makes validation fail closed instead of silently ignoring content. Unit schemas v1 and v2 remain readable for compatibility, but newly -prepared units use v3. +prepared units use v4. Unit v3 remains readable as a conversion input. Curated units must be atomic, readable by a second reviewer, and supported by the source. Provenance references must lead back to the original file and the passage that supports the claim. diff --git a/docs/gestione-memory.md b/docs/gestione-memory.md index c56325f9..71a904a9 100644 --- a/docs/gestione-memory.md +++ b/docs/gestione-memory.md @@ -1,244 +1,230 @@ # Memory management -This document describes how Memory currently works in ThothII: its conceptual model, lifecycle, persistence, semantic search, review gates, session display, and main technical limits. +Memory holds reusable knowledge for a workspace. PostgreSQL is the authoritative +archive; Qdrant contains a rebuildable search projection. The harness owns both +the administrative operations and the verification of retrieved results. -## Architectural summary +## Administration -A Memory item is domain knowledge that can be reused across questions. It is not a copy of one question's schema linking. +Open **Administration → Memory management**, immediately after Database management. +An administrator selects the workspace explicitly. No active session, DWH binding, +embedding service or Qdrant connection is required to browse and edit the archive. +PostgreSQL must be available. + +The page provides a paginated list, text/ID search, stable sorting, and combined +filters for family, concepts, database, table, column, origin and update date. +Filtering happens across the entire archive before pagination. Each card has a +complete detail view, an editable form, cancellation of unsaved changes, and an +explicit deletion confirmation. + +| Family | Content | +| --- | --- | +| Domain clarification | A reusable definition or interpretation, with scope and context. | +| SQL rule | Guidance for constructing SQL, with scope and rationale. | +| Solved question | A question, its approved SQL and context; used as a consultative exemplar. | +| Explained error | A correction and its rationale, to avoid repeating a known error. | + +All cards have a stable `mem-` identity, title, scope, origin and timestamps. +Manual cards have no invented source session or decision. Workflow cards retain +their source references when an administrator edits them. There is no editorial +revision history. + +The form also manages concepts, structured database/schema/table/column +dependencies, and links to other cards with an explicit meaning. Links can only +connect cards in the same workspace. Card, dependency and outgoing-link changes +are committed together. Deleting a card removes its incident links and dependencies, +while keeping the other cards, Evidence and Catalog metadata. + +## Save, failure and recovery ```mermaid -flowchart TB - CLARIFY["F1 concept clarified"] --> REVIEW["F8 reviewer review"] - REVIEW -->|"accepted"| REGISTRY["registry.jsonl"] - REVIEW -->|"declined"| LOCAL["Session decision only"] - REGISTRY --> VECTOR["Qdrant semantic index"] - VECTOR --> FUTURE["Future F2 retrieval"] - FUTURE --> PROPOSAL["Reviewer proposal"] +flowchart LR + EDIT["Admin or workflow"] --> SQL["PostgreSQL transaction"] + SQL --> CARD["Current card and links"] + SQL --> WORK["Pending projection"] + WORK --> Q["Qdrant"] + Q --> CHECK["Verify current card and projection"] + CARD --> CHECK + CHECK --> REVIEW["Recall for human review"] ``` -```text -F1: clarify a concept - │ - ▼ -concept_clarified decision in the session ledger - │ - ▼ -F8: reviewer decides whether to promote it - │ - ├── global registry registry.jsonl - └── Qdrant semantic index - │ - ▼ - F2 in a future session - search and proposal to the reviewer -``` +The service serializes each workspace's mutations. It first commits content and +the projection operation in PostgreSQL, then propagates to Qdrant. A failed save +is different from **saved, index update incomplete**. In the latter case the +current content is already available in administration, while its previous vector +result is excluded from recall. -The main invariant is `REUSABLE_TYPES = {"concept_clarified"}`: only clarified concepts can be generated, saved, searched, or proposed as Memory. Decisions such as `table_promoted`, `table_excluded`, and `column_promoted` remain local to the question. +The **Pending index updates** section offers an explicit retry, including cleanup +for deleted cards. These operations survive a restart. Retry reads the current +archive state and cannot restore an earlier edit or a deleted card. Technical +revisions are internal consistency markers, not user-managed card statuses. -## The three management levels +Every recall hit must resolve to a current card in the requested workspace, with +a matching valid projection. The response is reconstructed from PostgreSQL, never +from an unverified Qdrant payload. Deleted, orphaned and stale points are excluded +for both domain Memory and solved-question exemplars. An unavailable archive is +an operational error, not a successful empty archive. -| Level | Content | Function | +Rebuilding Memory uses only authoritative cards. It does not import JSONL files, +historical session artifacts or old vector payloads. Reference preprocessing keeps +the Memory collection separate and does not migrate legacy Memory payloads. + +## Hybrid recall and links + +Memory uses dense embeddings and Qdrant BM25, fused with reciprocal rank fusion. +Both branches receive the same workspace, family and scope filters before candidate +selection. Indexed text includes content, scope, rationale, question, exemplar SQL, +concepts and qualified physical dependencies. Both indexing and querying use the +workspace language (`en` or `it`), including manual administration without a DWH binding. + +The workflow CLI binds recall to its configured database and schema. Optional +`--filters` JSON can narrow the business `scope`, `concepts`, `table` and `column`. +Business scope is an exact string; every requested concept must be present. Physical +context matches a single structured dependency: database, schema, table and column +cannot be satisfied by unrelated entries. A card without dependencies is workspace-wide; +a database-only or table-only dependency also applies to descendants. Without a table +filter, table-specific knowledge in the selected database/schema remains discoverable. +Filters cannot override the configured database/schema. Free-text business scope is +not automatically interpreted or inferred from the question. + +The core expands outgoing links from current, eligible search candidates. It applies +the same filters and workflow-family restrictions to destinations and intermediate +cards. Removed, pending, previously decided and out-of-scope cards cannot act as bridges. +Traversal allows two hops, at most 20 outgoing links per card (stable target-ID order), +200 distinct card lookups and 400 inspected links in total. Search requests +`min(100, max(20, 3 × top))` seeds; the result limit remains between 1 and 100. + +All candidates are ranked together using `1 / (60 + seed rank)` for direct hits, +plus the strongest linked contribution, decayed by `0.5` per hop. Repeated paths +do not accumulate votes. Ties use card ID. The returned `score` is a ranking score, +not cosine similarity or a confidence estimate. `retrieval.path` shows the strongest +link path, or just the card ID for an exclusively direct result. PostgreSQL content, +links and eligibility are resolved under the workspace operation lock after search. + +Migration `002_hybrid_projection.sql` marks the format of old dense projections as +incompatible without changing their authoritative cards. They appear in pending +updates and are excluded from recall until explicit retry or `memory index` succeeds. +Memory writes can add a missing BM25 sparse vector to their collection; an incompatible +existing vector configuration fails visibly and leaves recovery pending. Reading never +silently falls back to dense retrieval. Reference remains independently managed. +An explicit `memory index` can also recreate a missing Memory collection before +rebuilding its projections; it never imports records from another source. + +## Workflow integration + +During the workflow, Pi prepares reusable proposals in the session artifact +`memory_proposals.json`. Each proposal names effective approved source decisions, +the content and scope, its rationale, and any physical dependencies or links. +Unexplained failures, rejected options and simple table selections do not create +reusable knowledge. Exact existing cards are reused; semantic similarity alone +never authorizes replacement. Updates name the existing card and its current revision. + +At F8, `reviewer_memory_promote` presents one editable Memory summary, including +the approved solved question. The reviewer chooses additions and updates, edits +their content, scope, dependencies and links, or declines everything. Approved SQL +is read only here: changing the solution requires returning to SQL review. +Only selected cards and their links are committed. Invalid links or stale updates +roll back the entire selection. Links between selected new cards are resolved +inside the same transaction. An explicit update preserves the existing identity +and origin, including manually authored content. + +A durable review receipt makes repeated delivery idempotent and recovers the +gap between saving Memory and recording `memory_summary_reviewed` in the session +ledger. Retry never recreates a deleted card. The gate then closes F8 and finalizes; +finalization itself performs no automatic Memory writes. Pending indexing remains +visible and recoverable in administration. + +F2 consumes domain clarifications and excludes already decided Memory. In F4, +F6 and F7, `memory rules` retrieves applicable SQL rules and explained errors. +Pi presents their use in the existing schema, CTE or SQL approval gate. +Exemplar search remains consultative. Retrieval never constitutes approval. +Persistent Memory/Evidence conflict repair remains part of the joint X1 increment. + +## Physical schema changes + +After a successful physical Catalog synchronization, the backend passes the exact +removed tables and columns, database, schema and run identity to Memory. Only cards +with matching structured dependencies are deleted, together with their incident +links and searchable projections. Global cards and objects outside the synchronized +scope survive. Manual Catalog cleanup and failed DWH scans never trigger this deletion. + +The Catalog transaction records a pending `memory_cleanup` phase before committing +the physical change. If cleanup or indexing fails, the run retains the original +removals. Retry completes that same operation without rescanning the DWH or relying +on a new diff. A new synchronization is blocked until this cleanup is completed. +The Memory deletion receipt and vector tombstones make the operation repeatable +across restarts. The synchronization drawer reports deleted Memory cards and errors. + +## API and commands + +The administrative HTTP surface requires `memory.manage`, included in the existing +admin role. The backend checks workspace identity and passes its trusted principal +and a protected request snapshot to the harness. The harness independently checks +the principal; ordinary users cannot bypass administration through the CLI. +Production authentication and CSRF protections apply to the new routes. + +| Method | Path below `/api/workspaces/:workspaceId/memory` | Operation | | --- | --- | --- | -| Session ledger | `concept_clarified`, `memory_promoted`, `memory_promotion_declined` | Audit and state for one session | -| Global registry | `mem-XXXX` records in `registry.jsonl` | Current canonical Memory archive | -| Qdrant index | Embeddings and metadata derived from the registry | Semantic search | +| GET | root | Search, filter and paginate the archive | +| POST | root | Create a card | +| GET | `/:cardId` | Read a complete card | +| PUT | `/:cardId` | Save content, links and dependencies together | +| DELETE | `/:cardId` | Delete a card and incident links | +| GET | `/pending` | List incomplete projection operations | +| POST | `/:cardId/retry` | Retry current projection work, including deletion | -The ledger contains provenance and human decisions. The global record contains reusable text. The vector index is a search projection, not the place where the workflow records decisions directly. - -## What can become Memory - -During F1, the workflow records clarifications as `concept_clarified` decisions. A clarification can express: - -- definitions of production or organizational concepts; -- inclusion and exclusion criteria for a product line; -- formulas and calculation methods; -- interpretations of time periods; -- mappings to specific tables and columns; -- the meaning of flags, codes, or indicators. - -The `MemoryRecord` model contains: - -- `id`, such as `mem-0001`; -- the timestamp, session, and sequence of the original decision; -- `type`; -- `subject`; -- `detail`; -- `rationale`; -- `question_context`; -- `tables` e `concepts`. - -For new Memory items, the type is always `concept_clarified` and `tables` starts empty. A table or column may appear in the explanation as a technical mapping, but it cannot be the Memory item's standalone concept. - -Valid example: - -> Ablation means a procedure with `ablazione_transcatetere = TRUE`, counted with `COUNT(DISTINCT cod_paz)` by year. - -Invalid examples: - -- `fact_cardioversione` as approved Memory; -- `dim_time` as rejected Memory; -- an "include this table" decision saved for future questions. - -The workflow skill also documents this rule in [harness/.pi/skills/tht-sessione/SKILL.md](https://git.tylconsulting.it/mptyl/ThothII/src/branch/main/harness/.pi/skills/tht-sessione/SKILL.md#L217). - -## Promotion at the end of the session: F8 - -At the end of the workflow, the `reviewer_memory_promote` gate runs a deterministic preview: +CLI configuration remains a per-command option. Commands emit pure JSON: ```text -tht memory promote --session --preview --json +tht memory list --filters '{"family":"sql_rule","page":1}' -c +tht memory show -c +tht memory create --data -c +tht memory update --data -c +tht memory delete --yes -c +tht memory pending -c +tht memory retry -c +tht memory index -c +tht memory search "" --session --json -c +tht memory search "" --filters '{"table":"orders","column":"id","scope":"Sales"}' --json -c +tht memory solved-search "" --json -c +tht memory rules "" --session --json -c +tht memory propose --session --data -c +tht memory summary --session --json -c +tht memory solved-index --json -c ``` -The preview: +`memory index` rebuilds both domain and solved-question projections. +`solved-index` only retries an existing authoritative source receipt. +The reviewer gate owns `memory review-apply`; Pi must not call it directly. +Migration `003_review_receipts.sql` adds durable review and physical-cleanup receipts. +The backend uses `memory admin --workspace -c ` so +administration does not materialize session or DWH configuration. -1. reads the session's effective decisions; -2. considers only `concept_clarified`; -3. discards decisions already promoted; -4. discards sequences already declined in F8; -5. deduplicates equivalent content; -6. proposes at most five candidates. +## Installation and storage -The code applies filtering and deduplication in -[harness/tht/memory/core.py](https://git.tylconsulting.it/mptyl/ThothII/src/branch/main/harness/tht/memory/core.py); the gate applies an additional defensive filter in -[harness/.pi/extensions/gate/memory/index.js](https://git.tylconsulting.it/mptyl/ThothII/src/branch/main/harness/.pi/extensions/gate/memory/index.js). +Memory shares the installation's existing PostgreSQL service, using its own +`thoth_memory` schema and versioned harness migration pack. The existing +`catalog-migrate` preparation service runs Catalog migrations followed by +`python -m tht.memory.migrate`. The core image includes both migration runners. +Ordinary API requests never migrate the schema. -The reviewer sees one preselected checklist. For each candidate: +Connection credentials come from the existing generated `THT_CATALOG_DB_HOST`, +`THT_CATALOG_DB_PORT`, `THT_CATALOG_DB_NAME`, `THT_CATALOG_RUNTIME_USER` and +`THT_CATALOG_RUNTIME_PASSWORD_FILE`. Migration uses the corresponding migrator +user/password file. Direct URL environments can use `THT_CATALOG_DATABASE_URL` +(runtime) and `THT_CATALOG_MIGRATOR_DATABASE_URL`; the harness also accepts +`THT_CATALOG_RUNTIME_DATABASE_URL`. Credentials are not authored in workspace YAML. -- selected: `tht memory save-one` runs, followed by a `memory_promoted` record; -- deselected: records `memory_promotion_declined`; -- no candidates: F8 closes automatically. +The migration grants the installation login membership in the restricted +`thoth_memory_runtime` role. Every repository transaction sets that role and a +workspace context. Forced row-level policies isolate cards, links, dependencies +and projection operations. This role has Memory DML and migration-status read +access, with no runtime DDL privilege. There are no cascading foreign keys to +Catalog or Evidence. A missing or incompatible schema returns a clear operational +error. Preparation is repeatable and checks migration checksums. -The `memory_promoted` marker uses `detail: seq:N`, a reference to the original `concept_clarified` decision. The flow is in [harness/.pi/extensions/tht-gate.js](https://git.tylconsulting.it/mptyl/ThothII/src/branch/main/harness/.pi/extensions/tht-gate.js#L1691). - -Promotion does not run in F2, and the model cannot invent F8 candidates. Direct promotion commands are also protected by the anti-bypass gate. - -## Global persistence - -### JSONL registry - -The current registry is: - -```text -/memory/registry.jsonl -``` - -The registry is written through a temporary file and `os.replace`, so replacement is atomic. Promotion is idempotent on the `session_id + decision_seq` pair: the same decision from the same session cannot create two global records. - -### Qdrant - -After promotion, `save-one` builds one `VectorRecord` and sends it to the Qdrant index. The indexed text includes: - -- type and subject; -- detail; -- rationale; -- question context; -- any concepts and mappings. - -The vector record uses the ID `memory:mem-XXXX`. Its metadata stores `subject`, `detail`, `rationale`, `tables`, `concepts`, and the `kind` discriminator. The content's SHA-256 hash prevents embedding and upsert work when the text has not changed. - -This behavior is implemented in -[harness/tht/memory/core.py](https://git.tylconsulting.it/mptyl/ThothII/src/branch/main/harness/tht/memory/core.py). - -### Current canonical source - -The JSONL registry remains the application's canonical source, while Qdrant is a derived but persistent index. The workflow does not record decisions directly in the vector database. It uses Qdrant as a searchable projection of the registry and effective ledger. - -## Reuse in F2 - -In a future session, F2 runs: - -```text -tht memory search "" --session --json -``` - -The command: - -1. creates an embedding for the question; -2. searches the vector store for records with `kind=memory` only; -3. resolves each hit in the JSONL registry through its `ref`; -4. discards records missing from the registry; -5. discards every type other than `concept_clarified`; -6. excludes Memory already decided in the current session; -7. returns results ordered by similarity. - -Memory is never applied automatically. The model must present it in one `reviewer_decide` choice: - -- a selected Memory item is recorded as a new `concept_clarified` in the current session; -- the rationale must cite the original `mem-XXXX` ID; -- a deselected Memory item means "do not apply it now", not "delete it globally". - -If F2 is reopened, a deselected Memory item can be proposed again. `memory_rejected` remains supported for legacy decisions and sessions, but it is not the normal behavior for current F2 deselection. - -## Effective ledger, rollback, and reopening - -The ledger is append-only. Reopening and withdrawing decisions do not delete earlier rows, but they change which decisions are effective. - -Memory helpers use `effective_decisions()` to: - -- exclude withdrawn decisions; -- ignore decisions from phases that became stale after a rollback; -- prevent promotion of clarifications that are no longer valid. - -The effective view is defined in [harness/tht/phase.py](https://git.tylconsulting.it/mptyl/ThothII/src/branch/main/harness/tht/phase.py#L82). - -## Display in the session summary - -The "Memories" section of the summary is projected at runtime from the session ledger; it is not a direct copy of the global registry. - -The projection: - -- shows approved Memory first; -- then shows declined Memory; -- resolves `seq:N` to the original `concept_clarified`; -- hides markers whose original record is not `concept_clarified`; -- hides subjects that match schema-linking tables; -- hides standalone subjects shaped like `fact_*` or `dim_*`; -- keeps table and field references when they are part of the conceptual explanation. - -The logic is in [harness/tht/session/store.py](https://git.tylconsulting.it/mptyl/ThothII/src/branch/main/harness/tht/session/store.py#L237). The frontend renders a structured list and treats `subject`, `detail`, and `rationale` as Markdown instead of showing raw Markdown. - -The transformation happens on read. Historical sessions use the current layout and filters without rewriting their original artifacts. - -## Enforced invariants - -The protections are distributed across several boundaries: - -1. `REUSABLE_TYPES` in the Python core; -2. the `memory search` command filter; -3. the F8 preview filter; -4. filtering and deduplication in the Pi gate; -5. exclusion of table Memory from the UI projection. - -This prevents one prompt or component change from reintroducing tables as Memory. - -## Remaining limits and risks - -### The registry and index are not one transaction - -Saving broadly follows this sequence: - -```text -JSONL registry → Qdrant → memory_promoted marker in the ledger -``` - -If Qdrant is unavailable, the registry can contain Memory that is not yet searchable. The command reports that reindexing is required. - -If the ledger marker fails after the vector database save, the Memory item can exist globally without a complete session audit. The gate returns a manual recovery command. - -### Five-candidate limit - -F8 proposes at most five Memory items. If a session produces more than five valid concepts, the extra items are not shown and the session can be finalized without promoting them. - -### Deduplication is not global - -Deduplication prevents duplicates within one proposal, and idempotency prevents the same decision from being promoted twice. There is no global merge of semantically similar Memory items from different sessions. - -### Old physical records - -Old `table_promoted` or `table_excluded` records may still exist in historical artifacts or indexes. The current code makes them unusable by filtering by type and does not show them in session projections. Physically removing them from the vector database remains a separate cleanup task. - -## Final assessment - -The current implementation matches the functional requirement: Memory is reusable conceptual knowledge, not a schema-linking choice. - -The strongest part is the multilayer protection of the `concept_clarified` type. The main technical debt is the coexistence of the JSONL registry and Qdrant, with no single transaction spanning the global archive, semantic index, and session ledger. +See [the M1 specification](plans/2026-09-08-memory-m1-spec.md) and +[ADR 0018](adr/0018-use-postgres-for-memory-and-qdrant-for-retrieval.md) for the +approved scope and acceptance boundaries. +See [M2 implementation and validation](plans/2026-09-09-memory-m2-validation.md) +for the retrieval checks and real embedding test command. diff --git a/docs/plans/2026-09-08-evidence-management.md b/docs/plans/2026-09-08-evidence-management.md new file mode 100644 index 00000000..686007ea --- /dev/null +++ b/docs/plans/2026-09-08-evidence-management.md @@ -0,0 +1,492 @@ +# Progetto: Evidence management + +Data: 2026-09-08. Stato: archivio locale Markdown, editor esterni e consolidamento +manuale con commit/push dell'operatore accettati per la release 0; +restanti semplificazioni confermate, con chiarimenti su impatto core e controllo Git; +E1, E2, E3 e il collegamento ai gate X1 implementati il 2026-09-09. +Il contratto X1 è in [Session corrections](../contracts/archive-repair.md). Risultati e limiti in +[E1 — validazione](2026-09-09-evidence-e1-validation.md) e +[E2 — validazione](2026-09-09-evidence-e2-validation.md) e +[E3 — validazione](2026-09-09-evidence-e3-validation.md). + +La [revisione di semplicità](2026-09-08-memory-evidence-simplification-review.md) +registra il riesame di Q1–Q15. Il flusso principale è draft scritta dallo specialista +indipendentemente dall'installazione, raffinamento da parte del sistema e +conservazione locale per uso e manutenzione. La proposta tecnica aggiornata +distingue le fonti esterne dall'archivio canonico locale, evita PostgreSQL per le +Evidence e toglie commit/push dal normale CRUD. La contemporaneità con attività core +durante l'amministrazione non è più un requisito da supportare. + +Il progetto rende consultabili le Evidence dall'applicazione e modificabili nei +file locali tramite l'editor scelto dall'operatore. +Segue le [decisioni comuni di Administration](2026-09-08-memory-evidence-administration.md) +ed è consegnabile separatamente da [Memory management](2026-09-08-memory-management.md). + +## Risultato richiesto + +Un amministratore apre Evidence management direttamente da Administration, cerca e +filtra le Evidence registrate, consulta il documento completo e trova il percorso +del file Markdown da modificare con un editor esterno. La voce segue Memory management e si trova allo stesso +livello di Database management, senza appartenere alle sue funzionalità. + +La consultazione è indipendente da una sessione e dalla disponibilità dell'indice +semantico. Mostra le Evidence Unit complete, senza limitarsi ai frammenti o ai +risultati più simili restituiti dal recall. + +## Accesso e modifica nella release 0 + +Il proprietario richiede file facilmente raggiungibili e modificabili anche da uno +specialista non tecnico, con Markdown come formato di lavoro e senza JSONL per la +gestione delle Evidence. Ha scelto editor esterni per la release 0: l'editor +preferito sul Mac o PC, oppure vim, nano o equivalenti sul server. Questa scelta +sostituisce la proposta di editor e form di contenuto dentro Evidence management. +Non si integra un prodotto di editing né si costruisce un editor applicativo. + +Evidence management mantiene elenco, ricerca, filtri e dettaglio. Mostra il percorso +assoluto della cartella del workspace e di ciascun file, con possibilità di copiarlo, +risolto dalla configurazione effettiva dell'installazione. Indica su quale host si +trova: per Docker occorre il percorso persistente accessibile sull'host, non soltanto +quello interno al container. Le istruzioni distinguono draft originali e file delle +Evidence raffinate da manutenere; includono un esempio Markdown per crearne una. + +Chi lavora sul server modifica direttamente i file; chi lavora su una copia sul +Mac o PC la riporta nello stesso archivio con i propri strumenti. Non è richiesto +costruire upload/download nell'applicazione o predisporre una cartella condivisa. +Un percorso sul server non viene presentato come un file apribile dal browser locale. + +Dopo aver salvato, aggiunto o rimosso i file, l'operatore controlla lo stato Git, +esegue il consolidamento ed esegue commit e push; il normale diff Git è disponibile +per approfondire le modifiche. La pagina mostra +le istruzioni e i comandi con i percorsi effettivi. La richiesta più recente +sostituisce la proposta intermedia del pulsante `Apply file changes`: non sono +necessari un pulsante di esecuzione, watcher o operazioni Git automatiche. +Il semplice salvataggio nell'editor, o di una copia fuori dall'archivio, non aggiorna +il recall. Il consolidamento e il seguito Git sono descritti sotto. + +Il Markdown canonico attuale non è già un formato di editing libero: il parser +controlla marker e corrispondenza fra testo canonico e presentazione. Per le regole +di dominio il testo codificato viene letto prima di verificare il rendering; cambiare +la sola frase visibile può produrre un errore di canonicalità. E1 deve quindi rendere +il Markdown locale realmente editabile: il testo leggibile è autorevole, i campi +richiesti sono documentati e i metadati derivati sono aggiornati dal sistema. +Non si chiede all'operatore di correggere marker, hash o una copia codificata del +testo. Non si introduce un secondo archivio di scambio da sincronizzare. + +## Funzioni + +- Elenco paginato e ordinabile; ricerca per testo, titolo e identificatore stabile. +- Filtri combinabili per workspace, kind, purpose, concetti, tabelle/colonne, + Source Evidence, Review item e stato di pubblicazione. +- Dettaglio leggibile della card: contenuto tipizzato, ambito di applicazione, + Supporting excerpt, provenienza e accesso alla Source Evidence disponibile. +- Creazione e modifica tramite file Markdown ed editor esterno, con esempio dei + campi richiesti per kind. In assenza di documento esterno, l'applicazione registra + una dichiarazione manuale come fonte quando acquisisce la modifica. +- Risoluzione dei Review item e gestione delle unità orfane o da ritirare. +- Cancellazione tramite rimozione del file, riconosciuta dal consolidamento + nell'archivio locale verificato come accessibile; un errore di lettura non prova + una cancellazione. La rimozione dal corpus e dall'indice è persistente. +- Un comando manuale di consolidamento per acquisire e validare le modifiche locali + e renderle disponibili al core, con esito esplicito e possibilità di riesecuzione. +- Istruzioni per controllo del diff, commit e push dei file Evidence e dei relativi + metadati necessari, eseguiti dall'operatore nel repository indicato. + +Il servizio acquisisce il Markdown editabile e aggiorna la rappresentazione +tipizzata e i metadati derivati, preservando gli identificatori delle unità esistenti. +La cancellazione di una Evidence Unit non elimina automaticamente la sua Source +Evidence o altre unità derivate dalla stessa fonte. + +## Salvataggio — Q12 semplificata + +Il proprietario rifiuta la separazione ordinaria fra salvataggio della bozza e +pubblicazione: le modifiche durante una sessione sono considerate rare e non +giustificano un workflow editoriale separato. La scelta successiva dell'editor esterno +sostituisce il `Save` della form con il salvataggio del file e un comando manuale +di consolidamento. Questo acquisisce, valida e attiva insieme aggiunte, modifiche e +cancellazioni, senza un successivo `Publish` o un secondo revisore. Commit e push +sono il seguito manuale per versionare e trasferire il lavoro nel repository. +L'operatore autorizza l'attivazione con il consolidamento. Le correzioni approvate nei +gate chiamano direttamente lo stesso servizio e non richiedono un editor esterno. + +L'ultimo chiarimento del proprietario sostituisce il requisito iniziale di aggiornare +le sessioni aperte mentre un amministratore modifica le Evidence. Si assume che il +core sia fermo o si ignora la contemporaneità: non servono aggiornamenti a caldo, +notifiche, ricalcoli, attese delle sessioni o una modalità manutenzione dedicata. +Il risultato di una modifica completata è usato nelle successive elaborazioni. + +Restano le correzioni deliberate dalla sessione stessa secondo Q4: chiamano lo +stesso servizio, attendono l'esito e permettono alla sessione di usare la correzione +approvata. Il collegamento fra configurazione della ricerca e contenuto locale +deve quindi funzionare senza richiedere un nuovo commit della fonte per ogni modifica; +non richiede un sistema generale di aggiornamento delle altre sessioni. + +Nel seguito, salvataggio applicativo indica questa acquisizione delle modifiche o +la scrittura deliberata dal core. Il feedback ordinario è operazione in corso, +completata oppure errore. Un successo +completo significa che il contenuto è disponibile al core; se la persistenza riesce +ma l'aggiornamento del corpus o dell'indice fallisce, l'esito deve dirlo e indicare +come rieseguire il comando. Un errore tecnico non diventa una bozza che attende una nuova decisione +editoriale di pubblicazione. Dopo un errore o un riavvio, un corpus parziale non +deve essere dichiarato pronto per la successiva elaborazione. + +Le decisioni Q11–Q15 e il successivo chiarimento sui vincoli sono registrati +nell'[ADR 0019](../adr/0019-author-evidence-in-app-with-automatic-activation.md). + +## Consolidamento manuale e seguito Git — release 0 {#consolidamento-manuale-e-seguito-git--release-0} + +Il proprietario richiede un flusso KISS affidato alla disciplina dell'operatore. +Il comando proposto è `tht evidence consolidate ` nel harness: +è da implementare, non è un comando già disponibile. Riutilizza parser, validazione +e indicizzazione esistenti, adattati al Markdown editabile. La documentazione +operativa e la pagina mostreranno l'invocazione effettiva per l'installazione, +compreso l'accesso al core se il comando gira in Docker. + +Il consolidamento svolge in sequenza questi passaggi: + +1. Legge l'archivio locale e confronta aggiunte, modifiche e rimozioni con l'ultimo + contenuto consolidato. Se l'archivio non è accessibile, si ferma: non interpreta + il problema come cancellazione delle Evidence. +2. Verifica struttura e sezioni previste per kind, campi obbligatori, identificatori + univoci, tipi, riferimenti e provenienza coerenti. Segnala file, campo o sezione, + problema e correzione richiesta. Non valuta automaticamente la verità del contenuto + e non riscrive il significato delle regole con un nuovo raffinamento del modello. +3. Se ci sono errori, termina con esito non riuscito prima dell'attivazione; lascia + all'operatore i file da correggere e il comando da rieseguire. Se i controlli + passano, aggiorna metadati derivati e manifest, inclusa la protezione delle + correzioni e delle cancellazioni, poi prepara e verifica il candidato completo + prima di renderlo attivo nel corpus e nell'indice Evidence locale. +4. Riporta un riepilogo testuale breve di file aggiunti, modificati e cancellati, + l'esito locale, gli eventuali errori di indicizzazione e i file primari + da includere nel commit. Dopo un errore tecnico si riesegue lo stesso comando, + senza duplicare unità o perdere il contenuto modificato. Non dichiara pronto un + aggiornamento incompleto e non avvia commit, push, pull o merge. + +Il core legge soltanto il contenuto consolidato attivo, mai direttamente i file in +corso di modifica. Una modifica salvata ma non consolidata, o un tentativo fallito, +non deve mescolare testo nuovo e vecchi risultati di ricerca. Si conserva l'ultimo +corpus valido; se un errore tecnico non permette di garantirne l'integrità, si +segnala l'indisponibilità invece di usare uno stato parziale. Si riusa il percorso +esistente di preparazione del candidato e attivazione, senza creare un secondo +archivio autorevole o un sistema di coordinamento delle sessioni. + +L'archivio mantenuto è una working tree persistente, versionabile nel repository +che contiene quelle Evidence; può riusare il repository del workspace. Draft e +unità curate restano contenuti distinti anche se ospitati nello stesso repository. +Non si modifica una materializzazione temporanea o uno snapshot runtime ricreato +dal preprocessing. Setup e pagina indicano working tree, cartella delle Evidence, +branch e remoto configurati; non si creano automaticamente repository o remoti. + +Il seguito manuale, documentato come sequenza da eseguire nel repository corretto, è: + +```text +git -C status --short +tht evidence consolidate # comando previsto, da implementare +git -C add -- +git -C commit -m "Update Evidence" +git -C push +``` + +Il controllo umano usa i normali comandi Git nel terminale. `git status --short` +mostra quali file sono cambiati; `git diff HEAD` permette, quando serve, di vedere +le righe cambiate prima del consolidamento. `git diff --cached` è disponibile per +controllare ciò che si sta per committare. Non sono richiesti un visualizzatore +nell'applicazione, uno strumento grafico, due revisioni obbligatorie o un nuovo gate. +I controlli strutturali del consolidamento vengono invece sempre eseguiti. + +Si includono aggiunte, modifiche e cancellazioni dei dati primari necessari alla +ricostruzione, compresi manifest e informazioni di cura manuale quando cambiano. +Indici, cache, sessioni e segreti non fanno parte del commit Evidence. L'operatore +controlla aggiunte, modifiche e cancellazioni prima di consolidare; i file nuovi +segnalati da status vanno letti, perché non compaiono ancora nel diff dei file +tracciati. L'output del consolidamento indica anche i metadati generati da includere +nel commit. Errori Git, credenziali, branch senza upstream e conflitti vengono +risolti con i normali strumenti Git, senza retry o risoluzioni automatiche. + +L'ordine ha una conseguenza esplicita: un consolidamento riuscito aggiorna il core +locale prima di commit e push. Se il push manca o fallisce, la modifica rimane locale +e non è ancora trasferita al remoto; l'operatore completa il seguito Git. Un push +riuscito, da solo, non aggiorna altre installazioni. Nessun processo sorveglia i file +o impone che l'operatore abbia completato la sequenza prima di riprendere il lavoro. + +## Persistenza e contratto da evolvere + +Il proprietario ha approvato la risoluzione persistente dei conflitti fra Memory ed +Evidence nel round Q4 della discussione Memory. Quando la decisione richiede correggere +un'Evidence, il core prepara una proposta che identifica l'unità e mostra il testo +o l'ambito risultante. La correzione passa attraverso l'authoring e la pubblicazione +di questo modulo, con le relative autorizzazioni e validazioni. L'accettazione della +correzione da parte di un utente autorizzato avvia lo stesso salvataggio con attivazione +automatica. L'interfaccia distingue una proposta da approvare, un aggiornamento in +corso o fallito e una correzione già attiva; non basta risolvere soltanto +la domanda corrente. È sempre possibile dichiarare inadeguate le opzioni proposte. + +Oggi il [lifecycle Evidence](../evidence.md) attribuisce l'authoring al repository +del workspace: ThothII non lo modifica, non crea commit e non esegue push. +Il [contratto Workspace Evidence v3](../contracts/workspace-evidence-v3.md) lega la +pubblicazione a una revisione coerente del workspace. La manutenzione dei file locali +richiede un'evoluzione esplicita di questo percorso e del formato editabile. + +**Q11, riesaminata e accettata per l'archivio locale:** la draft dello specialista deve restare producibile e +consegnabile senza accesso al PostgreSQL o alla stessa installazione. La +scelta aggiornata conserva le fonti in file/repository esterni e le +Evidence raffinate in un archivio locale persistente su file Markdown, gestito dal sistema. +Il CRUD modifica tale archivio e aggiorna l'indice senza commit/push verso le fonti. +L'alternativa di rendere PostgreSQL autorevole anche per Evidence, suggerita nella +prima parte della revisione, è ritirata. PostgreSQL rimane l'archivio delle Memory. + +I file locali curati sono dati primari: sopravvivono a Clear e reindicizzazione e +devono essere inclusi nelle copie di sicurezza dell'installazione. Una loro modifica +non aggiorna automaticamente i documenti dello specialista o altre installazioni; +un eventuale trasferimento dei file resta esplicito, senza sincronizzazione bidirezionale. + +Il comportamento deve coprire le origini supportate dal prodotto: filesystem/Git, +HTTP e S3. Nella proposta riveduta sono origini di acquisizione delle fonti e dei +contenuti già disponibili. Servono importazione e provenienza coerenti; le unità +raffinate e le modifiche manuali vengono conservate nell'archivio locale. +La gestione delle Evidence non richiede di aggiungere scritture sui server HTTP o S3 +di origine. Non si può dichiarare completato il CRUD lasciando queste unità in sola +lettura senza un percorso di importazione utilizzabile. + +Il descriptor `evidence.source` indica oggi l'origine acquisita dal preprocessing: +non è già un importatore di documenti nel repository di authoring. Inoltre il +runtime normalizza HTTP/S3 come documenti generici, mentre il parser canonico delle +unità tipizzate viene applicato ai file sotto `curated//`. La provenienza +canonica ammette solo un percorso locale sotto `source/`; URI e fingerprint remoti +esistono invece nel corpus acquisito. L'importazione deve collegare questi due +livelli senza fingere che il percorso tipizzato sia già uniforme fra i trasporti. + +Modifiche e ritiri devono sopravvivere alla successiva preparazione, sincronizzazione +e reindicizzazione. La rigenerazione da una fonte non deve ripristinare silenziosamente +una unità cancellata o sovrascrivere la correzione del curatore. La semantica di +conflitto fra nuova Source Evidence e cura manuale è approvata in Q13. + +Il runtime attuale non garantisce questa protezione durevole: `evidence prepare` +rifiuta modifiche non committate ai file curati e al manifest, ma una fonte cambiata +può rigenerare anche unità corrette manualmente e già committate. Il ritiro esplicito +rimuove l'unità e i riferimenti nel manifest, senza conservare una soppressione che +impedisca a una successiva generazione di riproporla. Questi comportamenti sono +verificati in `harness/tht/evidence/authoring.py`; la gestione amministrativa richiede +di evolverli, non solo di esporli come comandi dell'editor. + +**Q13, approvata:** se la fonte aggiornata contraddice una correzione +manuale, conservare in uso l'Evidence salvata dall'amministratore e mostrare il +confronto in Evidence management. Una sostituzione richiede una scelta esplicita, +seguita dalla stessa operazione di attivazione. La conseguenza è che la regola manuale può restare +attiva anche se la fonte più recente dice altro, finché il conflitto non viene +risolto. Il proprietario accetta questa conseguenza. La decisione comprende la +permanenza delle cancellazioni e la protezione da sovrascritture silenziose: una +nuova preparazione non deve riproporre automaticamente conoscenze eliminate. + +Validazione e pubblicazione sono passaggi interni dell'unica operazione di salvataggio; +un errore non dichiara attivo il nuovo contenuto. Il sistema continua a rispettare +provenienza e coerenza del contenuto; il salvataggio esplicito dell'amministratore +costituisce l'approvazione della modifica manuale. La decisione sullo storico delle +Memory non elimina questi requisiti del dominio Evidence. + +## Creazione manuale e aggiornamento delle fonti — Q14 e Q15 approvate + +**Q14, approvata:** consentire di scrivere una +Evidence senza dover fornire un documento esterno. In release 0 la si scrive in un +nuovo file Markdown nell'archivio indicato, usando l'esempio documentato. +Il sistema registra una dichiarazione manuale come fonte gestita e la distingue +dall'informazione derivata da documenti. Lo stesso criterio vale per una correzione +che cambia il significato del contenuto: la dichiarazione dell'amministratore +sostiene il testo corrente, mentre il documento originario resta collegato come +origine e per rilevarne gli aggiornamenti. Non deve essere mostrato come prova di +una regola che non contiene. Il consolidamento acquisisce anche le nuove unità manuali. + +Oggi la provenienza canonica richiede fonte locale, hash ed estratti: non esiste +un'origine manuale esplicita. La validazione verifica integrità e presenza testuale +degli estratti, ma non dimostra la verità o il supporto semantico della regola. +Attuare Q14 richiede distinguere nel contratto la fonte del testo corrente dal +documento originario; non richiede inventare estratti o una validazione semantica +automatica presentata come garanzia di correttezza. + +**Q15, approvata:** riacquisire le fonti esterne su richiesta +esplicita dell'amministratore, senza controlli periodici o accessi alle fonti a ogni +domanda o salvataggio di una card. Fra due aggiornamenti richiesti, il sistema usa +quanto già acquisito. Il consolidamento continua ad attivare la modifica locale +senza richiedere prima un aggiornamento della fonte remota; eventuali conflitti +scoperti alla successiva acquisizione seguono Q13. + +Il proprietario ha approvato entrambe le raccomandazioni il 2026-09-08. Il successivo +riesame conserva questi comportamenti e rende esplicita l'autonomia dello specialista +che produce le draft. I dettagli tecnici seguenti descrivono la proposta semplificata. + +## Piano esecutivo + +### E1 — Raffinamento e archivio canonico locale + +**Modifica del sottosistema Evidence core.** Rendere editabile il Markdown cambia +il contratto che il codice attuale legge e scrive. Questa attività precede la +pagina amministrativa: non è una modifica limitata alla presentazione o all'editor. +Il nuovo formato conserva contenuti tipizzati, ambito e provenienza richiesti dal +core, ma usa il testo visibile come unica rappresentazione autorevole del contenuto +umano. Va identificato come una nuova versione del contratto Curated unit, distinta +dalla versione del descriptor del workspace. + +L'intervento coordinato comprende: + +- parser, renderer e validazione in `harness/tht/evidence/canonical.py`; +- preparazione, scritture e migrazione in `harness/tht/evidence/authoring.py`, + con adeguamento delle istruzioni e degli esempi usati nel raffinamento; +- acquisizione e normalizzazione in `harness/tht/evidence/corpus/normalize.py`, + consolidamento e preprocessing, affinché producano il contenuto tipizzato atteso + dall'indicizzazione e dal recall; +- correzioni deliberate dal core, provenienza e citazioni, che devono consumare + e aggiornare coerentemente la nuova rappresentazione; +- contratto documentato e test del percorso completo: modifica del testo visibile, + consolidamento, contenuto indicizzato e successivo utilizzo da parte del core. + +Le Evidence esistenti devono essere convertite esplicitamente al nuovo formato +e reindicizzate, verificando il mantenimento del contenuto e degli identificatori. +La fase di sviluppo permette un passaggio unico; non è richiesto mantenere due +formati di authoring concorrenti. Una conversione non rappresentabile fedelmente +va segnalata. Si preserva il modello tipizzato interno dove possibile, per contenere +la modifica nei componenti che dipendono dal documento senza ridisegnare le fasi +del workflow NL→SQL. + +Estendere i contratti in `harness/tht/evidence/` per distinguere fonte corrente, +dichiarazione manuale e documento originario. La dichiarazione viene gestita +dall'applicazione insieme all'unità locale; chi scrive il Markdown non deve +creare un file sorgente separato o gestire hash ed estratti. I campi per kind +mantengono la validazione deterministica, senza attribuirle una verifica della +verità della regola. Aggiornare il glossario e il contratto canonico con il codice. + +Riutilizzare raffinamento, parser, renderer e validazione esistenti, distinguendo +la working tree locale persistente dagli snapshot acquisiti e adattando il formato al testo Markdown +editabile come fonte autorevole, senza duplicazione nascosta del contenuto umano. +Documentare posizione dei file, campi richiesti e un esempio per la creazione. +L'applicazione acquisisce le draft e +conserva il risultato in file locali persistenti. Le mutazioni aggiornano contenuto +canonico, presentazione Markdown e manifest; le scritture sono serializzate per +workspace e controllano la versione corrente del record; una modifica concorrente +non viene risolta sovrascrivendo silenziosamente il contenuto altrui. + +Il salvataggio non richiede commit, push o un database condiviso con lo specialista. +L'archivio locale è distinto dalla materializzazione runtime ricostruita da Git e +dal corpus derivato. Consolidamento e preprocessing acquisiscono i contenuti locali +curati; il core consulta soltanto il corpus attivo validato. Clear preserva i file +e la ricostruzione dell'indice riparte da essi, senza sovrascriverli +con nuove copie delle draft originali. + +La persistenza conserva la cura manuale e le esclusioni necessarie a impedire la +ricomparsa delle unità eliminate. La rigenerazione confronta le proposte con tale +stato: assegnare un nuovo ID a un contenuto derivato non deve essere un modo per +riattivarlo automaticamente aggirando una cancellazione. Le nuove proposte e i +conflitti restano distinti dal corpus già approvato e attivo. + +Verificare con archivi locali temporanei draft esterne e raffinamento, creazione +manuale, cambiamento di significato +con provenienza corretta, round trip del formato canonico, aggiornamento del manifest, +cancellazione, conflitto fra scritture e retry. La validazione di un estratto +presente nella fonte non viene usata come prova automatica di supporto semantico. + +### E2 — Pagina amministrativa e consolidamento manuale + +Esporre API amministrative per elenco completo, filtri, dettaglio, mutazioni ed +esito delle operazioni. Il backend applica autorizzazioni e isolamento, poi chiama +il servizio di authoring del harness. Realizzare la pagina autonoma nell'AppShell +con provenienza leggibile, percorsi effettivi dei file sull'host e istruzioni per +modificarli con editor esterni, consolidarli e completare commit/push manualmente. +Il comando di consolidamento acquisisce aggiunte, modifiche e rimozioni locali +tramite lo stesso servizio. Nessun editor, form di contenuto o +trasferimento file web è richiesto in R0; non serve una sessione del core. + +Evolvere lo stage interno `runEvidenceStage` e il percorso harness +`tht preprocess evidence` per attivare il contenuto canonico locale autorizzato. +Il salvataggio non avvia una riacquisizione HTTP/S3, una nuova estrazione dalle +Source Evidence o una scansione DWH. Il nuovo comando manuale Evidence riusa questo +stage interno senza richiedere un preprocessing completo per ogni modifica. + +Preparare e verificare il candidato Evidence prima di attivarlo. L'indicizzazione +e la rimozione devono essere circoscritte al workspace e alle generazioni Evidence +interessate nella collezione Reference condivisa. Preservare Schema, relazioni, +LSH e Memory: il salvataggio non usa Preprocessing Clear e non ricrea l'intera +collezione. Se l'attivazione fallisce, mostrare l'esito parziale e rendere ripetibile +il completamento della stessa modifica, senza chiedere una nuova approvazione. + +Adeguare la ricerca perché consumi l'archivio locale aggiornato, senza richiedere +commit della fonte o cambiare configurazioni estranee. Adeguare il calcolo della +readiness: aggiornare Evidence non dichiara risolto un blocco +di Schema o Catalog. Dopo un Clear, l'eventuale necessità di preprocessing completo +rimane esplicita; il solo consolidamento non ricostruisce tutti i derivati mancanti. + +Verificare che i percorsi mostrati portino ai file effettivi anche su installazioni +Docker e che una modifica alla frase visibile sia acquisita senza editing di metadati +tecnici. Verificare creazione/modifica/cancellazione con esito disponibile al core, errore +di indicizzazione dopo persistenza, retry e isolamento delle altre componenti +della collezione. Le successive elaborazioni devono usare il contenuto corrente +senza recuperare unità cancellate. Verificare le correzioni deliberate dalla sessione +stessa, senza aggiungere prove di CRUD amministrativo concorrente al core. Un +problema preesistente di Catalog deve continuare a impedire una falsa readiness. +Verificare che errori strutturali impediscano l'attivazione e producano indicazioni +correggibili, che il comando sia rieseguibile e che non esegua operazioni Git. +Verificare che file modificati senza consolidamento e tentativi falliti non cambino +il contenuto usato dal core né combinino versioni differenti di testo e indice. +Provare la sequenza documentata di commit/push con repository temporanei e remoto +locale, verificando l'inclusione delle cancellazioni e dei metadati necessari. + +### E3 — Importazione, refresh esplicito e conflitti con le fonti + +Fornire un percorso di importazione locale per le origini già supportate, +conservando contenuto acquisito e provenienza remota. Uniformare il passaggio alle +unità canoniche: i documenti HTTP/S3 oggi normalizzati genericamente devono diventare +file locali editabili. Le unità importate restano gestibili senza credenziali di +scrittura sui server fonte. + +L'azione amministrativa di aggiornamento delle fonti riacquisisce il contenuto e +confronta le impronte con quanto già acquisito. Per le unità curate interessate, +prepara il confronto con le proposte risultanti senza sostituire la dichiarazione +manuale attiva. Questa protezione non dipende dalla capacità del modello di +riconoscere ogni contraddizione semantica. La scelta dell'amministratore usa poi +lo stesso salvataggio con attivazione; fonti invariate non richiedono nuova cura. + +Verificare con origini controllate che la riacquisizione avvenga soltanto su richiesta, +che un errore di accesso non venga scambiato per una cancellazione e che consolidamento/recall +non chiamino i connettori di acquisizione remota. Verificare aggiornamento di una +fonte collegata a una correzione manuale, permanenza della versione curata, decisione +di sostituzione e mancata ricomparsa automatica delle unità eliminate. + +La lettura può essere consegnata come incremento intermedio di E2, ma il progetto +è completo soltanto con CRUD, attivazione e gestione delle origini previste. X1 nel +piano comune collega le proposte provenienti dai conflitti con Memory e verifica +che raggiungano questi stessi servizi con autorizzazioni ed esiti coerenti. + +## Criteri di completamento + +- Tutte le unità sono raggiungibili mediante elenco e filtri, senza una sessione. +- La pagina mostra cartella e percorso effettivo sull'host per ogni unità. Markdown, + campi richiesti ed esempi permettono creazione e modifica con editor esterni; + i dati invalidi non vengono attivati e gli errori indicano il file da correggere. +- Una draft prodotta fuori dall'installazione può essere acquisita, raffinata e + conservata localmente senza accesso dello specialista a PostgreSQL. +- Creazione, modifica e ritiro aggiornano l'archivio locale e il corpus usato dal core; + non richiedono commit/push nel repository delle fonti. +- Una Evidence può essere creata senza documento esterno; la dichiarazione manuale + è riconoscibile e il documento originario di una correzione non è mostrato come + supporto di un'affermazione che non contiene. +- Una modifica o cancellazione completata resta valida dopo una nuova preparazione + e reindicizzazione; non rimangono frammenti ricercabili della versione rimossa. +- Una fonte aggiornata in conflitto con la cura manuale non sostituisce l'Evidence + attiva: il confronto resta disponibile all'amministratore fino alla sua decisione. +- Il consolidamento manuale completa anche l'attivazione, senza un successivo `Publish`; l'esito + distingue operazione in corso, completata ed eventuali errori con retry. +- Il core usa soltanto l'ultimo corpus consolidato valido. File in lavorazione o + consolidamenti falliti non diventano disponibili attraverso letture dirette. +- Dopo una modifica completata, le successive elaborazioni recuperano il contenuto + corrente. Dopo una cancellazione completata non recuperano l'unità rimossa. + Nessuna delle due operazioni riscrive decisioni, artefatti o SQL già prodotti. +- Clear e ricostruzione degli indici preservano le Evidence canoniche locali, + incluse le correzioni e le esclusioni dovute a cancellazioni. +- Le correzioni originate da conflitti con Memory raggiungono lo stesso percorso + autorevole delle modifiche amministrative e mostrano il proprio stato di pubblicazione. +- Le origini senza scrittura dispongono di un percorso utilizzabile verso l'authoring. +- Le fonti esterne vengono riacquisite solo su richiesta dell'amministratore; + il consolidamento e il recall usano il contenuto locale già acquisito. +- API e pagina applicano accesso amministrativo e isolamento dei workspace. +- Le istruzioni identificano working tree e file effettivi; il seguito manuale + include verifica del diff, commit e push. Nessun watcher o comando Git automatico + è introdotto. L'esito distingue disponibilità locale e trasferimento al remoto. +- Il salvataggio delle Evidence preserva Memory, Schema, relazioni, LSH e Catalog + Metadata; non elimina blocchi di readiness estranei all'aggiornamento Evidence. diff --git a/docs/plans/2026-09-08-memory-evidence-administration.md b/docs/plans/2026-09-08-memory-evidence-administration.md new file mode 100644 index 00000000..e28fb99e --- /dev/null +++ b/docs/plans/2026-09-08-memory-evidence-administration.md @@ -0,0 +1,261 @@ +# Amministrazione di Memory ed Evidence + +Data del piano: 2026-09-08. Aggiornamento 2026-09-09: M1–M3, E1–E3 e X1 +implementati. Risultati e limiti della verifica finale sono raccolti nel +[rapporto X1](2026-09-09-archive-repair-x1-validation.md). + +La [revisione di semplicità](2026-09-08-memory-evidence-simplification-review.md) +riesamina tutte le decisioni Q1–Q15 alla luce degli ultimi chiarimenti del +proprietario e confronta le conseguenze con il piano precedente. Per Evidence R0 +sono scelti file Markdown locali, editor esterni e consolidamento manuale con +controllo della struttura, seguito da diff, commit e push dell'operatore. La pagina +deve indicare chiaramente percorsi e comandi. Le restanti semplificazioni sono +confermate. Il cambio del formato richiede adeguare il sottosistema core Evidence, +convertire i file e reindicizzare; E1 precede la pagina. Il controllo umano usa +lo stato Git nel terminale, con diff delle righe quando serve, senza una UI dedicata. +Si ignora la contemporaneità fra amministrazione e core. Lo +specialista scrive draft indipendentemente dall'installazione; il sistema le +raffina e le conserva localmente. La proposta aggiornata usa file canonici locali +per le Evidence e toglie commit/push automatici dal CRUD; questa revisione tecnica di Q11 +sostituisce la raccomandazione precedente ed è stata attuata negli incrementi E1–E3. + +## Obiettivo e decisioni del proprietario + +L'amministratore deve poter accedere in qualsiasi momento alle Memory e alle Evidence +registrate, cercarle, filtrarle, aprirle, crearle, modificarle e cancellarle. L'accesso +non richiede una sessione del core o una fase del workflow. + +Il lavoro è diviso in due progetti autonomi: + +- [Memory management](2026-09-08-memory-management.md); +- [Evidence management](2026-09-08-evidence-management.md). + +I progetti condividono l'esperienza di gestione e i componenti appropriati, mantenendo +distinti i contenuti, le regole di validazione e i percorsi di persistenza. + +### Collocazione dei due accessi + +L'ordine previsto nell'accordion Administration è: + +```text +Administration + Database management + Memory management + Evidence management + ──────────────────── + Workspace management + Pi management +``` + +Memory management ed Evidence management sono due voci autonome, allo stesso livello +di Database management. «Sotto» indica soltanto la posizione fisica nella navigazione. +Non sono sottopagine, tab o funzionalità di Database management. Ciascuna apre la +propria pagina e possiede il proprio stato di navigazione. + +### Decisioni già acquisite sulla Memory + +- La Memory serve a migliorare schema linking e generazione SQL di domande future. +- Il perimetro comprende chiarimenti di dominio riutilizzabili, regole corrette per + join, filtri e aggregazioni, domande risolte consultabili ed errori da evitare + quando il motivo è stato compreso e approvato. Le scelte occasionali non diventano + regole generali; possono restare nel contesto di un exemplar. +- Il formato della memory card è allineato per analogia a quello delle Evidence; + origine e dominio restano separati. Le Memory non entrano nel canone Evidence. +- La gestione avviene tramite CRUD e form interni a ThothII. +- L'evoluzione del modulo Memory è interna a ThothII, riusando l'infrastruttura + dell'installazione e senza adottare un framework esterno per governarne il comportamento. +- Il progetto comprende ricerca semantica e lessicale ibrida in Qdrant, filtri + sull'ambito e collegamenti espliciti fra card percorsi dal core. Non introduce + un database a grafi dedicato e non rinvia il grafo a una successiva sperimentazione. +- PostgreSQL, già presente nell'installazione, è l'archivio autorevole di card, + collegamenti e dipendenze, in tabelle proprie del modulo Memory. Sostituisce il + registro JSONL; Qdrant è una proiezione rigenerabile, con sincronizzazione esplicita. +- I collegamenti sono proposti e approvati insieme alle card nel riepilogo finale + e gestibili manualmente da Administration. Cancellare una card elimina anche i + collegamenti che la coinvolgono, conservando le altre card. +- Gli exemplar `solved_question` usano il formato card con consumo consultativo. +- Il core prepara durante il lavoro un riepilogo finale modificabile delle nuove + Memory e degli aggiornamenti proposti. Il reviewer seleziona cosa salvare; la + cancellazione resta nel CRUD amministrativo. +- Le Memory vengono proposte nei gate pertinenti: chiarimenti all'inizio, regole di + collegamento nello schema linking e regole di calcolo durante la costruzione SQL. + Le approvazioni sono integrate nei gate, senza una domanda separata per ogni card. +- Una modifica sostituisce il contenuto corrente: non è richiesta una cronologia + aggiuntiva delle revisioni delle Memory o una ricostruibilità storica dedicata. +- Se un riferimento allo schema rende una Memory inutilizzabile, il proprietario + sceglie la cancellazione anziché lo stato «Needs review». Dopo una sincronizzazione + riuscita dello schema fisico, il backend comunica gli elementi rimossi al modulo + Memory, che cancella le card dipendenti e le proiezioni. La relazione fra card ed + elementi dello schema deve essere strutturata; cleanup del Catalog ed errori di + connessione non sono prove di rimozione fisica. +- I conflitti fra Memory ed Evidence si risolvono con azioni chiuse, specifiche e + accompagnate dal contenuto risultante. La scelta alimenta una correzione degli + archivi attraverso i rispettivi percorsi, con stato di pubblicazione esplicito; + è sempre possibile dichiarare inadeguate le proposte e richiederne la riformulazione. +- La verifica usa test funzionali automatici e regressioni per casi concreti; non + promette un miglioramento qualitativo generale o un benchmark con/senza Memory. +- Le sessioni esistenti non vincolano il design. Il loro azzeramento durante lo + sviluppo, se necessario, è autorizzato; non è un'operazione eseguita da questi documenti. + +La decisione di non conservare uno storico riguarda le Memory. Non modifica +automaticamente il contratto di authoring e pubblicazione delle Evidence. + +### Salvataggio delle Evidence + +In Q12 il proprietario rifiuta la separazione fra `Save draft` e `Publish`. +La scelta successiva dell'editor esterno sostituisce il Save della form con il +salvataggio dei file e un comando manuale di consolidamento: verifica la struttura, +indica le correzioni necessarie e, se valido, aggiorna metadati, corpus e indice. +Acquisisce anche aggiunte e cancellazioni. Segue il controllo del diff con commit e +push manuali dell'operatore; non sono previsti watcher, Git automatico o un editor +in Evidence management. La pagina mostra cartella, percorsi sull'host e comandi; +il [piano Evidence](2026-09-08-evidence-management.md) specifica la sequenza. +Il consolidamento attiva localmente il contenuto; un push mancante o fallito lascia +il trasferimento al remoto da completare manualmente. L'ultimo chiarimento +elimina il requisito di aggiornare sessioni aperte a seguito di modifiche +amministrative: durante tali modifiche il core è fermo o la contemporaneità può +essere ignorata. Rimangono le correzioni deliberate dalla sessione stessa. +La review di Q11 distingue le draft dello specialista dalle Evidence raffinate +locali; raccomanda di preservare le prime e gestire le seconde su file persistenti, +senza PostgreSQL condiviso o scritture automatiche nel repository delle fonti. +In Q13 è approvata +la precedenza della correzione manuale quando una fonte aggiornata la contraddice: +resta attiva fino alla risoluzione esplicita del confronto in Evidence management. +Le Evidence cancellate non ricompaiono automaticamente durante la rigenerazione. +Q14 consente di crearle senza documento esterno, in R0 tramite un nuovo Markdown: una +dichiarazione manuale sostiene il testo corrente, con l'eventuale documento +originario conservato come provenienza distinta. Q15 limita la riacquisizione +delle fonti esterne a una richiesta esplicita dell'amministratore; il normale +salvataggio e il recall usano il contenuto già acquisito. +L'[ADR 0019](../adr/0019-author-evidence-in-app-with-automatic-activation.md) registra +l'evoluzione del contratto, ancora da implementare. + +## Stato verificato nel repository + +`frontend/src/shell/AppShell.tsx` contiene Administration con Database management, +Workspace management e Pi management. Non contiene le due pagine richieste. + +Il modulo Memory espone già comandi CLI per elenco, dettaglio, aggiornamento, +cancellazione e ricerca, ma manca una superficie CRUD amministrativa web. +Il riepilogo di una sessione mostra soltanto le Memory collegate alla sessione. +Il [contratto attuale della Memory](../gestione-memory.md) descrive registro JSONL +canonico e indice Qdrant derivato. + +Le Evidence hanno già un percorso di consultazione e modifica dei documenti nel +repository di authoring, descritto in [Evidence: sources, preparation, and review](../evidence.md). +Manca una pagina amministrativa per queste operazioni dentro ThothII. Il contratto +attuale prevede che ThothII legga e pubblichi il repository, senza modificarlo, +creare commit o eseguire push: l'editing amministrativo richiede evolvere questo +confine, come esplicitato nel progetto Evidence. + +## Esperienza comune + +L'amministratore lavora nell'interfaccia operativa esistente, con una lista densa +e leggibile e un'area di dettaglio. Si riusano tema, controlli, focus e navigazione +già definiti in PRODUCT.md e DESIGN.md. + +- Selettore di workspace, ricerca testuale, filtri combinabili, ordinamento e paginazione. +- Elenco completo dei record persistiti, indipendente dalla disponibilità della + ricerca semantica. Una ricerca per similarità può affiancarlo, senza limitarlo ai + pochi risultati del recall del core. +- Apertura del contenuto completo, dell'ambito di applicazione e della provenienza. +- Creazione e modifica Memory mediante form; per Evidence R0, percorsi ed esempi + Markdown per editor esterni, consolidamento e seguito Git manuali. +- Cancellazione con indicazione precisa dell'oggetto e del suo effetto; per Evidence + R0 la rimozione del file è acquisita dal consolidamento. +- Stato esplicito di salvataggio e disponibilità per il core, con recupero dagli errori. +- Filtri e posizione nell'elenco conservati quando si apre e si chiude un record. +- Controlli utilizzabili da tastiera; stato vuoto, nessun risultato e indisponibilità + del servizio distinguibili. Chrome in inglese, contenuti nella lingua del workspace. + +Le due pagine non dipendono dalla selezione di un database in Database management. +Eventuali filtri su tabelle e colonne usano riferimenti al catalogo quando disponibili; +la loro assenza non impedisce di consultare i contenuti registrati. + +Componenti condivisibili: barra di ricerca e filtri, lista, paginazione, struttura +del dettaglio, campi comuni della card, provenienza e feedback delle operazioni. +Form Memory, istruzioni di manutenzione Evidence, autorizzazione, validazione, pubblicazione e +persistenza rimangono responsabilità dei rispettivi moduli. Il riuso del frontend +non introduce un archivio canonico unico per Memory ed Evidence. + +## Confini e integrazione + +Entrambe le pagine appartengono ad Administration e devono applicare il controllo +amministrativo anche nelle API. La collocazione visiva non assegna automaticamente +le autorizzazioni di Database management ai nuovi moduli. + +Il progetto Memory può essere consegnato senza attendere il progetto Evidence. +I componenti comuni si estraggono quando servono ai flussi reali di entrambi. +La verifica finale congiunta copre ordine della navigazione, accesso indipendente, +filtri, percorsi di manutenzione e feedback coerenti, isolamento fra workspace e assenza di effetti +incrociati fra i due domini. + +La separazione fra Reference Vector Collection e Memory Vector Collection rimane +quella dell'[ADR 0017](../adr/0017-separate-reference-vectors-from-runtime-memory.md). +Un'operazione amministrativa sulle Evidence non cancella le Memory; una cancellazione +di Memory non elimina Evidence, Schema o metadati del database. + +## Piano esecutivo dei due progetti + +Il riesame mantiene i requisiti funzionali Q1–Q15 e semplifica le scelte tecniche +secondo le condizioni descritte sopra. L'ordine di lavoro parte dalla Memory e +riusa poi i componenti effettivamente comuni per Evidence management. Le API e +il coordinamento delle scritture attuano questi vincoli; la distinzione fra draft +esterne e archivio locale sostituisce l'ipotesi di un unico archivio Git da modificare. + +| Ordine | Incremento | Risultato verificabile | Dipendenze | +| --- | --- | --- | --- | +| 1 | M1 — Archivio e CRUD Memory | PostgreSQL autorevole, API e pagina con elenco completo, filtri, form e cancellazione; collegamenti e dipendenze persistiti, proiezioni aggiornate o invalidate | Nessuna dipendenza dal progetto Evidence | +| 2 | M2 — Ricerca Memory | Ricerca ibrida con filtri ed espansione limitata dei collegamenti; risultati coerenti con il contenuto corrente | M1 | +| 3 | M3 — Memory nel workflow | Riepilogo finale, uso delle categorie nei gate e cancellazione dopo sincronizzazione fisica riuscita | M1 e M2 | +| 4 | E1 — Preparazione e archivio Evidence | Draft esterne, raffinamento del sistema, file canonici locali persistenti e protezione delle correzioni/cancellazioni | Q11–Q15 riesaminate; non richiede il runtime Memory | +| 5 | E2 — Manutenzione e consolidamento Evidence | Pagina con percorsi dei Markdown; comando manuale di verifica e attivazione; istruzioni per diff, commit e push | E1 | +| 6 | E3 — Fonti e conflitti di aggiornamento | Importazione delle origini supportate, refresh esplicito, confronto con le correzioni manuali | E1 ed E2 | +| 7 | X1 — Integrazione finale | Correzioni persistenti dei conflitti Memory/Evidence e verifica congiunta delle due pagine | M3 ed E3 | + +E1–E3 sono un progetto separato: la sequenza è l'ordine operativo scelto per questa +consegna, non una dipendenza tecnica dal modulo Memory. Il CRUD Memory è utilizzabile +come primo incremento; M2 e M3 restano obbligatori nel progetto attuale. La correzione +completa dei conflitti fra i due archivi si considera consegnata soltanto con X1. + +### Responsabilità di implementazione + +- Il harness possiede contratti, persistenza e operazioni dei moduli Memory ed + Evidence. Il backend applica autorizzazioni, espone le API e orchestra le + operazioni; non introduce una seconda implementazione delle stesse scritture. +- Il Metadata Catalog rimane responsabilità del backend. Una sincronizzazione + fisica riuscita chiama la pulizia Memory nello stesso flusso, con gli elementi + effettivamente rimossi nell'ambito controllato e un recupero dopo interruzione. + La pulizia non dipende dalla UI e non richiede un bus di eventi. +- Le pagine condividono controlli di consultazione e feedback; logica dei kind, fonte autorevole e + attivazione rimangono nei rispettivi moduli. L'AppShell ospita le due voci autonome + nell'ordine concordato. +- Le modifiche coordinate hanno gestione esplicita di errori e retry. I controlli + sulla versione corrente impediscono sovrascritture inconsapevoli senza richiedere + uno storico delle Memory. I retry non duplicano card o collegamenti. Il normale + salvataggio è sequenziale, con un esito persistente da recuperare se incompleto; + non richiede nuovi worker, code generiche o coordinamento delle sessioni aperte. + +### Verifiche e chiusura della consegna + +Ogni incremento esegue i test delle operazioni che cambia e i controlli dei layer +coinvolti: pytest/ruff per il harness, vitest e typecheck per backend/frontend. +Le pagine sono verificate anche nel browser per navigazione, filtri, form, errori +e uso da tastiera. I contratti e i test specifici sono nei due piani di progetto. + +X1 verifica sia i conflitti che correggono Memory sia quelli che correggono Evidence, +inclusi utente privo dell'autorizzazione necessaria, proposta rifiutata, errore di +attivazione e retry. Il gate deve mostrare lo stato reale dell'archivio: la sola +risoluzione della domanda corrente non dimostra che la correzione sia persistita. +La verifica congiunta copre inoltre isolamento dei workspace, ordine dei link e +assenza di cancellazioni incrociate. + +I contratti correnti vengono aggiornati insieme al relativo codice. Alla fine si +aggiornano PROJECT_STATE.md e documentazione operativa e si esegue la build strict. +La verifica della generazione con un modello reale rimane distinta dai test +deterministici; non si promette un benchmark generale di miglioramento qualitativo. + +Gli incrementi approvati sono implementati e disponibili nel Docker locale. +I rapporti di validazione distinguono i test deterministici, i servizi reali e le +prove con il modello configurato. Le verifiche sintetiche non modificano la +conoscenza PSD; la valutazione dei contenuti reali rimane una decisione del reviewer. diff --git a/docs/plans/2026-09-08-memory-evidence-simplification-review.md b/docs/plans/2026-09-08-memory-evidence-simplification-review.md new file mode 100644 index 00000000..b632c904 --- /dev/null +++ b/docs/plans/2026-09-08-memory-evidence-simplification-review.md @@ -0,0 +1,253 @@ +# Revisione di semplicità: Memory ed Evidence + +Data: 2026-09-08. Stato: archivio locale Markdown, editor esterni e consolidamento +manuale con seguito Git scelti per la release 0; conseguenze esplicitate, +restanti semplificazioni confermate, chiariti impatto core e controllo Git. +Nessuna modifica applicativa. + +## Condizioni che guidano la revisione + +Chi amministra il sistema è competente e deve vedere chiaramente cosa produce ogni +azione. Durante il CRUD amministrativo di Memory ed Evidence si può assumere che +non ci siano attività core in corso; la loro eventuale contemporaneità non è un +caso da supportare con meccanismi dedicati. + +Lo specialista di contesto può essere una persona diversa da chi gestisce +l'installazione. Scrive le draft delle Evidence senza dover accedere a PostgreSQL, +amministrarlo o disporre della stessa installazione. Il flusso fondamentale resta: + +1. Lo specialista scrive e consegna le draft in documenti accessibili al sistema. +2. Il sistema le acquisisce e le raffina in Evidence strutturate. +3. Il sistema conserva localmente le Evidence per consultazione e manutenzione. + +La proposta avanzata durante la revisione di usare PostgreSQL come archivio +autorevole anche delle Evidence è ritirata. Confrontava il costo del solo CRUD +interno, senza rappresentare adeguatamente l'autonomia di chi produce le fonti. +Un database dietro un'interfaccia non richiederebbe di per sé accesso SQL agli autori, +né un database condiviso; questo però non risolve da solo il flusso di redazione e +consegna esterno. Non propongo di introdurre tale dipendenza. + +## Archivio locale delle Evidence accettato dopo il chiarimento + +Il proprietario ha accettato i file locali e richiede una gestione facile da trovare +e usare anche per uno specialista non tecnico. Il formato di lavoro è Markdown; +JSONL non è una superficie di gestione delle Evidence. Il proprietario ha scelto +editor esterni per la release 0, chiedendo di indicare chiaramente dove sono i file. + +| Contenuto | Responsabile | Conservazione e uso | +| --- | --- | --- | +| Draft originali | Specialista di contesto | File o repository delle fonti, redigibili e consegnabili indipendentemente dall'installazione | +| Evidence raffinate e correzioni locali | Sistema e persone autorizzate alla manutenzione del contenuto | File Markdown in un archivio locale persistente dell'installazione, consultabili dall'applicazione e modificabili con editor esterni | +| Indice di ricerca | Sistema | Qdrant, ricostruibile dalle Evidence locali correnti | + +Questa è una revisione della parte tecnica di Q11: distingue l'autorità delle draft +esterne dall'archivio delle Evidence raffinate usate dalla singola installazione. +Il repository delle fonti rimane utilizzabile; il consolidamento delle Evidence +locali non esegue commit o push e non richiede un PostgreSQL condiviso. Il seguito +manuale ora richiesto comprende controllo del diff, commit e push nel repository +che contiene i file curati. L'archivio è una working tree persistente; draft e unità +curate possono stare nello stesso repository, mantenendo distinta la loro funzione. + +Le conseguenze devono essere esplicite. Una correzione locale cambia ciò che usa +quell'installazione e non viene rispedita automaticamente allo specialista o ad +altre installazioni. I file restano trasferibili per un passaggio esplicito; non +si costruisce una sincronizzazione bidirezionale. L'archivio locale va conservato +e incluso nelle copie di sicurezza: dopo una correzione non è più un semplice +output eliminabile e rigenerabile dalle draft senza perdita di lavoro. + +Memory mantiene PostgreSQL locale come archivio già concordato. Le due pagine +restano separate e riusano controlli comuni; la scelta della persistenza segue il +flusso di produzione dei rispettivi contenuti. + +## Gestione della release 0: editor esterni scelti + +Il proprietario sceglie il proprio editor sul Mac o PC, oppure vim, nano o +equivalenti sul server. La proposta di editor applicativo è ritirata; non si +integra una libreria di editing né si sviluppano form di contenuto Evidence. + +`Evidence management` conserva lista, ricerca, filtri e dettaglio. Mostra la cartella +del workspace e il percorso assoluto di ogni file, copiabile e risolto dalla +configurazione effettiva. Indica l'host su cui si trova; con Docker mostra il percorso +persistente accessibile sull'host. Le istruzioni distinguono draft originali e +Evidence raffinate da manutenere e includono un esempio Markdown per crearne una. + +Si modificano direttamente i file nell'archivio locale dell'installazione. Chi +lavora su una copia sul proprio computer la riporta lì con i propri strumenti. +Non sono richiesti upload/download web, nuove cartelle condivise o sincronizzazione. + +Dopo aver salvato, aggiunto o rimosso file, l'operatore controlla lo stato Git, +poi esegue un comando manuale di +consolidamento: controlla la struttura attesa, segnala file e correzioni necessarie +e, solo se i controlli passano, aggiorna metadati, corpus e indice. Il comando è +rieseguibile dopo correzioni o errori tecnici. Seguono commit e +push manuali, con le istruzioni della pagina e della documentazione operativa. +Il core consulta solo l'ultimo corpus consolidato valido; i file in lavorazione +e gli aggiornamenti falliti non devono introdurre contenuto parziale nella ricerca. +La richiesta più recente sostituisce la proposta intermedia del pulsante +`Apply file changes`; non servono un'esecuzione dalla UI, watcher o operazioni Git +automatiche. Il salvataggio nell'editor da solo non aggiorna il recall. Il controllo +dei file usa `git status --short` nel terminale; il normale `git diff` è disponibile +per approfondire le righe cambiate, senza visualizzatore web o doppia revisione +obbligatoria. Il comando riporta un breve riepilogo testuale delle modifiche. + +Il Markdown locale deve essere realmente editabile: il testo visibile è autorevole, +i campi richiesti sono documentati e i metadati derivati sono gestiti dal sistema. +Il formato v3 attuale, che verifica il rendering contro una copia codificata del +testo, va quindi adattato. Non basta indicare i percorsi dei file attuali e non si +introduce un secondo archivio di scambio da sincronizzare. Questo comporta una +modifica effettiva dei componenti core Evidence: parser, renderer, preparazione, +validazione, normalizzazione e collegamento a indicizzazione/recall. E1 comprende +nuova versione del contratto, conversione dei file esistenti e verifica del percorso +completo; il formato interno tipizzato viene conservato dove possibile. + +Creazione, modifica e cancellazione delle unità raffinate passano dai file e dallo +stesso consolidamento. Un errore di accesso all'archivio non è una prova di +cancellazione. Le correzioni approvate dal core continuano a chiamare direttamente +il servizio di scrittura. Il proprietario ha confermato le restanti semplificazioni +chiedendo di esplicitare impatto core e semplicità del controllo Git. +Il [piano Evidence](2026-09-08-evidence-management.md#consolidamento-manuale-e-seguito-git--release-0) +specifica controlli, comando previsto e sequenza Git. Dopo il consolidamento il +contenuto è disponibile localmente; finché commit/push non sono completati, non è +versionato/trasferito al remoto. Il sistema si affida alla disciplina dell'operatore +e non tenta di completare o riparare automaticamente la sequenza Git. + +## Conseguenze rispetto al piano precedente alla revisione + +Il confronto riguarda il piano concordato prima del riesame: form strutturate per +le unità, repository Git autorevole con scritture applicative e supporto alle +modifiche amministrative durante sessioni aperte. Le funzionalità non erano ancora +implementate: si confrontano due progetti, non una regressione già introdotta. + +| Aspetto | Cosa cambia | Conseguenza pratica | +| --- | --- | --- | +| Gestione R0 delle Evidence | Editor esterno scelto al posto di editor e form nell'applicazione | Si usano strumenti già disponibili. La pagina indica i file; servono un Markdown realmente editabile, esempi e validazione in acquisizione. Si rinuncia alla guida e ai controlli durante la digitazione. | +| Applicazione delle modifiche esterne | Salvare i file nell'archivio ed eseguire il consolidamento manuale | Una copia sul Mac o PC va riportata nell'archivio con gli strumenti dell'operatore. Fino al consolidamento riuscito la modifica non è disponibile al core. Gli errori strutturali indicano il seguito necessario. Non si sviluppano trasferimenti file web o watcher. | +| Autorità delle Evidence raffinate | Archivio locale, distinto dalle draft esterne | Una correzione agisce su quell'installazione. Lo specialista che lavora alle draft e altre installazioni non la ricevono automaticamente. | +| Cronologia e distribuzione Git | Controllo del diff, commit e push manuali dopo il consolidamento | Git conserva e trasferisce quanto l'operatore committa e pubblica. La sequenza non è imposta né completata dal sistema: se il push manca o fallisce, il core locale può già usare modifiche non trasferite. Non c'è rollback applicativo o gestione automatica dei conflitti Git. | +| Ripristino dei dati | Le Evidence locali curate sono dati primari | Il backup deve comprenderle. Ricostruire tutto dalle sole draft recupererebbe la base, ma potrebbe perdere correzioni e cancellazioni locali; Clear deve preservare l'archivio. | +| Attività core contemporanee | Non vengono più gestite le modifiche amministrative durante il lavoro core | Se avvengono comunque, non è garantita la coerenza della sessione in corso. Non si aggiornano contesti o SQL già prodotti. Le successive elaborazioni usano il contenuto aggiornato dopo il completamento dell'operazione. | +| Salvataggio e indice | Operazione sequenziale con recupero minimo persistente | L'utente attende l'esito dell'indicizzazione. Un problema può richiedere Retry; il contenuto già salvato viene conservato e un esito incompleto non viene presentato come pieno successo. Il piano precedente già prevedeva questi esiti, non garantiva retry automatici. | +| Pulizia Memory dopo sync | Chiamata diretta nel flusso esistente | Stesso effetto funzionale, con meno coordinamento interno. Rimangono il recupero dopo interruzione, i limiti dell'ambito controllato e l'esclusione di errori di connessione o cleanup del solo Catalog. | +| Verifica semantica | Nessuna nuova valutazione generale o revisione obbligatoria a ogni Save | La responsabilità del significato resta allo specialista; i test verificano contratti e casi mirati. La validazione strutturale non garantiva la correttezza del dominio neppure nel piano precedente. | + +Sono invariati il riepilogo Memory, le categorie ammesse, i gate esistenti, la +ricerca ibrida e i collegamenti, le correzioni persistenti dei conflitti e la +protezione delle modifiche manuali. Le fonti vengono aggiornate su richiesta come +già concordato. Le scritture deliberate dal core stesso restano supportate tramite +lo stesso servizio: l'assenza di amministrazione concomitante non le elimina. + +Non è prevista una rinuncia alle capacità di ricerca o al contenuto delle Evidence. +Conservare gli stessi contenuti tipizzati, ambiti e indicizzazione evita una perdita +di qualità dovuta a un taglio di funzionalità, ma l'equivalenza del nuovo percorso +deve essere verificata. Il formato Markdown e un editor più semplice non sono, +da soli, una garanzia di qualità o di assenza di errori. + +## Riesame di tutte le decisioni Q1–Q15 + +| Decisione | Esito della revisione | Approccio più semplice e conseguenza | +| --- | --- | --- | +| Q1 — Riepilogo Memory | Mantengo | Un riepilogo finale editabile, con selezione di cosa salvare. Nessuna approvazione ripetuta per ogni card durante la sessione. | +| Q2 — Aggiunte e aggiornamenti | Mantengo | Mostrare contenuto risultante e record sostituito. La somiglianza non avvia fusioni automatiche; la cancellazione semantica resta esplicita nel CRUD. | +| Q3 — Consumo nei gate | Mantengo | Usare i gate già previsti. Nessun nuovo percorso di approvazione per la sola consultazione di una Memory; exemplar consultativi. | +| Q4 — Conflitti Memory/Evidence | Semplifico l'esecuzione | La scelta mostra quale archivio cambia e con quale testo. Chiamare lo stesso servizio di salvataggio del CRUD, con un esito unico; niente secondo sistema di pubblicazione. Le correzioni deliberate dal core restano un caso da supportare. | +| Q5 — Dipendenze fisiche eliminate | Mantengo, con chiamata diretta | Alla fine della sincronizzazione fisica riuscita, chiamare la pulizia Memory nello stesso flusso. Mostrare le conseguenze nella conferma della sincronizzazione già esistente e il conteggio finale. Non introdurre bus di eventi o un nuovo controllo continuo del DWH. | +| Q6 — Verifiche | Mantengo il perimetro limitato | CRUD, filtri, persistenza, cancellazioni, recupero dall'errore e casi mirati di ricerca. Nessuna valutazione qualitativa generale a ogni salvataggio, nessun secondo modello giudice obbligatorio. | +| Q7 — Evoluzione interna | Mantengo | Riutilizzare componenti e servizi dell'installazione. Nessun framework esterno o ulteriore servizio per governare Memory. | +| Q8 — Ibrido e collegamenti | Mantengo entrambe le capacità | Riutilizzare la ricerca ibrida; collegamenti in una lista modificabile ed espansione limitata nel core. Nessun database a grafi, editor visuale di grafi o deduzione automatica di una rete di relazioni. Le capacità restano nel progetto attuale. | +| Q9 — PostgreSQL per Memory | Mantengo | Il database è già locale all'installazione; card, collegamenti e dipendenze restano coordinati. Qdrant è ricostruibile. Il flusso Memory non richiede un autore esterno indipendente. | +| Q10 — Cura dei collegamenti | Mantengo | Gestirli nello stesso riepilogo e dettaglio della card. Cancellare una card elimina i collegamenti incidenti e conserva le altre card. | +| Q11 — Archivio Evidence | Rivedo la soluzione tecnica | Draft esterne indipendenti e Evidence raffinate in file locali persistenti. Togliere commit/push dal CRUD. La modifica locale non aggiorna automaticamente la fonte dello specialista. La proposta PostgreSQL autorevole per Evidence è ritirata. | +| Q12 — Save e sessioni aperte | Adatto agli editor esterni | Dopo il salvataggio dei file, un comando manuale consolida e attiva le modifiche; l'operatore completa poi commit/push. Nessun workflow editoriale aggiuntivo, watcher, aggiornamento delle sessioni aperte o modalità manutenzione. | +| Q13 — Fonte cambiata e cura manuale | Mantengo, con confronto semplice | La correzione manuale prevale finché una persona decide altrimenti. Un cambiamento della fonte collegata rende disponibile il confronto; non serve dimostrare automaticamente una contraddizione semantica. Le cancellazioni non vengono annullate dalla rigenerazione. | +| Q14 — Creazione manuale | Mantengo tramite file | Un nuovo Markdown secondo l'esempio documentato consente una Evidence manuale con provenienza dichiarata e senza documento esterno obbligatorio. Il flusso principale draft dello specialista → raffinamento locale rimane disponibile e indipendente. | +| Q15 — Refresh delle fonti | Mantengo | Acquisizione e raffinamento delle fonti aggiornate su richiesta esplicita. Consolidamento e recall usano il contenuto locale; non cercano nuove versioni remote. | + +Restano confermati due link autonomi sotto Database management, CRUD completo con +filtri, assenza di storico aggiuntivo delle Memory e assenza di vincoli sulle +sessioni di sviluppo già esistenti. + +## Riduzioni trasversali del piano + +**Un'operazione alla volta, con esito comprensibile.** Save per Memory e +consolidamento manuale per Evidence validano e aggiornano l'indice in sequenza. +L'interfaccia mostra operazione in corso, +completata oppure errore con azione di recupero. Non espone un workflow editoriale +di stati `draft`, `approved`, `published` per il normale CRUD. La draft dello +specialista è il documento di ingresso della preparazione, non un secondo pulsante +di salvataggio dell'editor delle unità locali. + +**Recupero minimo dopo errore.** Se l'archivio è stato aggiornato ma Qdrant no, va +detto e deve essere possibile riprovare, per Evidence rieseguendo il comando. +Serve un'indicazione persistente del lavoro +incompleto, sufficiente anche per ripulire una card già cancellata dopo un riavvio. +La ricerca non deve usare contenuti rimossi o superati alla domanda successiva. +Il piano non imponeva già una outbox o nuovi worker: la revisione rende esplicito +che non sono richiesti né code generiche né sincronizzazioni continue. + +**Pulizia schema diretta e ripetibile.** Il flusso esistente del Catalog può +richiamare Memory dopo l'applicazione dello schema. Occorre coprire il crash fra +le due scritture: conservare la pulizia pendente sul run oppure verificare di nuovo +le dipendenze contro lo snapshot fisico riuscito e il suo ambito. Ricalcolare solo +il nuovo diff perderebbe le rimozioni già applicate. La rimozione manuale di +metadati Catalog o un errore di connessione non autorizzano a cancellare Memory. + +**Nessuna gestione delle sessioni amministrate contemporaneamente.** Non progettare +aggiornamenti a caldo, ripristino dei contesti già letti, invalidazione dello SQL, +notifiche alle sessioni o generazioni aggiuntive per lettori paralleli. Rimangono le +scritture esplicitamente richieste dalla sessione stessa: il riepilogo Memory e una +correzione Evidence approvata chiamano il medesimo servizio e ne gestiscono l'esito +prima di proseguire. Questo caso non richiede coordinare tutte le altre sessioni. + +**Pagine essenziali.** Ricerca testuale, filtri e lista completa permettono di trovare +i contenuti; la ricerca semantica appartiene anzitutto al core. Memory conserva le +form; Evidence espone dettaglio, percorsi e istruzioni per la modifica esterna. +Hash e manifest restano gestiti dal sistema. Si riusano i controlli di accesso +esistenti; chi modifica i file necessita dei permessi sul relativo filesystem, +senza accesso a PostgreSQL o un nuovo sistema generale di ruoli. + +## Conseguenze delle azioni da rendere visibili + +| Azione | Conseguenza da comunicare | +| --- | --- | +| Save riuscito | Il contenuto corrente è conservato ed è disponibile per le successive elaborazioni. | +| Save con errore dell'indice | Il contenuto è conservato; la disponibilità alla ricerca non è completata. Retry completa il lavoro senza richiedere di riscrivere la modifica. | +| Salvataggio nell'editor esterno | Cambia il file, ma non aggiorna il recall. Una copia esterna va prima riportata nell'archivio dell'installazione. | +| Consolidamento Evidence | Acquisisce aggiunte, modifiche e cancellazioni locali; verifica la struttura e, se valida, aggiorna metadati e indice. Un errore indica cosa correggere; si riesegue il comando dopo la correzione o un errore tecnico. | +| Commit e push manuali | Versionano e trasferiscono al repository remoto i file consolidati. Un errore Git si risolve manualmente; il consolidamento locale già riuscito non viene annullato. | +| Delete di una Memory | La card e i collegamenti che la coinvolgono vengono rimossi; le altre card restano. | +| Delete di una Evidence | L'unità locale non viene più usata né ricreata automaticamente; la draft originale e le altre unità derivate restano. | +| Refresh sources | Si acquisiscono nuove versioni delle fonti e si preparano le Evidence interessate; le correzioni locali protette non vengono sovrascritte. | +| Sincronizzazione fisica | Le Memory dipendenti da elementi effettivamente eliminati vengono cancellate; l'ambito e il conteggio dell'effetto sono visibili. | +| Preprocessing Clear | Si eliminano i dati derivati previsti dal comando; le Evidence canoniche locali e le Memory rimangono. | + +Non sono necessarie conferme ripetute su Save. Per le cancellazioni si usa una +conferma concreta sull'oggetto e sulle conseguenze, integrando gli effetti nella +conferma già presente quando l'azione è una sincronizzazione distruttiva. + +## Cosa va comunque implementato + +Il raffinamento locale e la scrittura di Markdown canonico/manifest esistono in +`harness/tht/evidence/authoring.py`. Preparano e validano gli output prima della +sostituzione con staging e rollback; non eseguono commit/push. Va separato il +requisito di Git worktree dalla preparazione del contenuto e va integrata +l'acquisizione delle draft esterne. + +La materializzazione corrente in `backend/src/workspaces/evidence/materialization.ts` +è invece ricostruita da una revisione Git: non è già l'archivio locale scrivibile +proposto. Il renderer e il preprocessing devono consumare il nuovo archivio +persistente. Una modifica locale non deve richiedere un nuovo commit della fonte. + +Si riusa l'attivazione Evidence esistente dove serve a verificare un candidato e +a recuperare da errori; l'assenza di lettori contemporanei non rende atomici file +e Qdrant. La mutazione resta circoscritta alle Evidence e preserva Schema, relazioni, +LSH e Memory. Il suo successo non cancella blocchi di readiness del Catalog. + +Le verifiche prioritarie coprono il percorso draft → raffinamento → elenco/CRUD +locale → ricerca, persistenza dopo riavvio, retry dopo errore, mancata ricomparsa +delle unità eliminate, refresh con correzioni locali, isolamento dei workspace e +conservazione dell'archivio locale dopo Clear. Le prove di aggiornamento live da +amministrazione vengono eliminate; resta la verifica delle correzioni deliberate +dal core stesso. + +L'ordine resta Memory, Evidence e integrazione finale. I dettagli degli incrementi +nei due piani sono lavoro interno; per l'utente rimangono due gestioni autonome. diff --git a/docs/plans/2026-09-08-memory-m1-spec.md b/docs/plans/2026-09-08-memory-m1-spec.md new file mode 100644 index 00000000..da996bf5 --- /dev/null +++ b/docs/plans/2026-09-08-memory-m1-spec.md @@ -0,0 +1,343 @@ +# M1 — Archivio autorevole e amministrazione delle Memory Card + +Data: 2026-09-08. Stato: M1 implementato e verificato localmente; +confini di test confermati dal proprietario il 2026-09-08. + +Primo incremento del progetto Memory management. Attua le decisioni già approvate +nel piano del 2026-09-08 e nell'ADR 0018. Le scelte tecniche di dettaglio qui +proposte derivano dalla ricognizione del runtime. Gli esiti dell'implementazione +sono riportati nel [rapporto di verifica](2026-09-08-memory-m1-validation.md). + +## Problem Statement + +L'amministratore deve poter trovare, leggere e curare tutta la conoscenza +riutilizzabile di un workspace: chiarimenti di dominio, regole SQL, domande +risolte ed errori compresi da evitare. Oggi manca una pagina amministrativa +dedicata e l'archivio è frammentato: le Memory sono registrate in JSONL, mentre +gli exemplar delle domande risolte sono indicizzati attraverso un percorso distinto. + +Questa situazione non offre un unico archivio completo di card, collegamenti e +dipendenze strutturate. La disponibilità dell'indice non deve determinare se una +card è consultabile o modificabile. Una modifica o cancellazione deve inoltre +impedire che una ricerca successiva utilizzi contenuti superati, anche se +l'aggiornamento dell'indice fallisce. + +## Solution + +Consegnare la pagina **Memory management** in Administration e un archivio +PostgreSQL autorevole, appartenente al modulo Memory. La pagina consente elenco +completo, ricerca testuale, filtri, dettaglio, creazione, modifica, cancellazione +e gestione dei collegamenti. I contenuti restano consultabili con embedding o +Qdrant indisponibili, purché PostgreSQL sia disponibile. + +Un salvataggio aggiorna insieme card, collegamenti e dipendenze, poi propaga la +modifica a Qdrant. L'amministratore vede se il contenuto è stato salvato e se è +disponibile al recall. Se la propagazione fallisce può riprovarla, anche dopo un +riavvio. I contenuti rimossi o superati non sono utilizzati dal recall. + +M1 consegna l'amministrazione e la coerenza dell'archivio. La ricerca ibrida con +espansione dei collegamenti e il nuovo riepilogo del workflow sono gli incrementi +M2 e M3, entrambi ancora obbligatori per completare il progetto Memory. + +## User Stories + +1. Come amministratore, voglio aprire Memory management da Administration, così + da curare la conoscenza senza avviare una sessione. +2. Come amministratore, voglio scegliere esplicitamente il workspace da + amministrare, così da sapere a quale archivio appartiene ogni operazione. +3. Come amministratore, voglio elencare tutte le card del workspace, così da + raggiungere anche quelle che non compaiono nel recall semantico. +4. Come amministratore, voglio cercare per testo o identificatore e ordinare i + risultati, così da trovare una card senza conoscerne la formulazione esatta. +5. Come amministratore, voglio combinare filtri per famiglia, concetti, riferimenti + a tabelle o colonne, provenienza e aggiornamento, così da restringere l'intero + archivio prima della paginazione. +6. Come amministratore, voglio leggere contenuto completo, ambito, motivazione e + provenienza disponibile, così da capire quando una card è applicabile. +7. Come amministratore, voglio creare una card manuale senza inventare una sessione + o una decisione di origine, così da registrare conoscenza curata direttamente. +8. Come amministratore, voglio rappresentare chiarimenti, regole SQL, domande + risolte ed errori compresi, così da conservare i contenuti concordati. +9. Come amministratore, voglio conservare domanda, SQL e contesto di un exemplar, + così da distinguerlo da una regola generale. +10. Come amministratore, voglio associare dipendenze esplicite a database, tabelle + e colonne, così da non affidare l'identificazione degli oggetti al testo libero. +11. Come amministratore, voglio correggere i campi consentiti dalla famiglia e + annullare una modifica non salvata, così da controllare il contenuto corrente. +12. Come amministratore, voglio ricevere errori di validazione comprensibili senza + perdere il testo inserito, così da poterlo correggere. +13. Come amministratore, voglio creare, modificare e cancellare collegamenti con + destinazione e significato espliciti, così da curare le relazioni fra card. +14. Come amministratore, voglio salvare card e modifiche correlate come un'unica + operazione, così da non lasciare collegamenti o dipendenze parziali. +15. Come amministratore, voglio cancellare una card e i suoi collegamenti + incidenti conservando le altre card, così da rimuovere solo il contenuto scelto. +16. Come amministratore, voglio ritrovare le modifiche dopo riapertura della pagina + e riavvio del servizio, così da verificare che il salvataggio sia persistente. +17. Come amministratore, voglio consultare e curare l'archivio quando embedding o + Qdrant sono indisponibili, così da proseguire il lavoro amministrativo. +18. Come amministratore, voglio distinguere archivio vuoto e archivio non + disponibile, così da non interpretare un guasto come perdita dei dati. +19. Come amministratore, voglio distinguere salvataggio fallito e contenuto + salvato con indicizzazione incompleta, così da scegliere il recupero corretto. +20. Come amministratore, voglio riprovare una propagazione incompleta anche dopo + un riavvio o una cancellazione, così da completare la pulizia dell'indice. +21. Come reviewer, voglio che una nuova ricerca escluda card eliminate o contenuti + superati, così da ricevere soltanto conoscenza corrente. +22. Come reviewer, voglio che gli exemplar rimangano consultativi e che una Memory + recuperata non costituisca approvazione, così da conservare il controllo del workflow. +23. Come amministratore, voglio che reindicizzazione e preprocessing rispettino le + cancellazioni e le correzioni, così da non doverle ripetere. +24. Come operatore dell'installazione, voglio preparare lo schema e configurare + l'accesso Memory con i meccanismi esistenti, così da avviarlo senza nuovi servizi. +25. Come proprietario del workspace, voglio che API e comandi rispettino il + contesto autorizzato, così da evitare accessi o collegamenti fra archivi diversi. +26. Come proprietario del workspace, voglio che il CRUD Memory lasci invariati + Evidence e metadati del database, così da mantenere distinte le responsabilità. + +## Implementation Decisions + +### Responsabilità e punti d'ingresso + +- Il modulo Memory del harness possiede modello, validazione, repository, + mutazioni, collegamenti, dipendenze e coerenza delle proiezioni. PostgreSQL è + autorevole; Qdrant contiene una proiezione ricostruibile. +- I comandi Memory esistenti diventano adattatori del medesimo servizio. La + superficie viene completata con creazione manuale, gestione dei collegamenti, + elenco delle propagazioni incomplete e retry. Nessuna logica di persistenza + Memory viene duplicata nel backend. +- Il backend espone API amministrative attraverso il runner del harness, + associando principal attendibile e configurazione del workspace alla richiesta. + Si conservano opzioni di configurazione per comando e output JSON puro. +- La ricognizione conferma che il harness usa già SQLAlchemy e PostgreSQL per le + sessioni. Se ne riusano i meccanismi adatti, mantenendo separati modello, + migrazioni e proprietà dei dati Memory. Il repository in memoria del Metadata + Catalog è un test double del Catalog, non il modulo Memory. + +### Modello e transazioni + +- La card ha identità stabile, workspace, famiglia/contenuto, titolo o soggetto, + ambito, motivazione, provenienza, concetti e date di creazione/aggiornamento. + Le domande risolte conservano anche domanda, SQL e contesto. L'identità di una + card manuale non dipende da una sessione né dal solo testo. +- Il modello rappresenta le quattro categorie di contenuto approvate senza + imporre quattro nuovi kind vettoriali o la corrispondenza con i tipi del ledger. + Le origini manuali sono distinguibili; sessione e decisione sono riferimenti + opzionali quando effettivamente disponibili. +- I collegamenti sono record propri del modulo con sorgente, destinazione e + significato. Le due card devono esistere nello stesso workspace. La cancellazione + di una card elimina i collegamenti incidenti, senza propagarsi alle altre card. +- Le dipendenze identificano esplicitamente database, schema, tabella e colonna + secondo l'ambito applicabile. Il contratto consente il futuro confronto con + lo schema fisico. Un cleanup dei metadati Catalog non deve poter cancellare + card attraverso una cascata implicita di chiavi esterne. +- Le tabelle logiche necessarie sono card, collegamenti, dipendenze e stato + operativo delle proiezioni. Card e modifiche correlate si aggiornano nella + stessa transazione. Una validazione o scrittura fallita non lascia aggiornamenti + parziali. Non si conserva una storia delle revisioni del contenuto. +- Le modifiche ordinarie non spostano una card in un altro workspace. Tutti gli + identificatori ricevuti vengono verificati nel workspace dell'operazione, + inclusi estremi dei collegamenti e riferimenti delle azioni di retry. + +### Salvataggio, cancellazione e recall + +- Il servizio valida l'intera mutazione, registra il nuovo stato autorevole e + il lavoro di propagazione nella stessa transazione PostgreSQL, poi aggiorna + Qdrant nello stesso flusso di salvataggio. Si completa un'operazione alla volta; + non occorrono una coda generale, un worker o una sincronizzazione continua. +- Lo stato persistente è sufficiente a distinguere la proiezione del contenuto + corrente da una proiezione precedente e a ritentare l'azione dopo un riavvio. + Può usare una versione tecnica o un'impronta interna; non è una cronologia + editoriale né un ulteriore stato che l'utente debba gestire. +- Un risultato Qdrant è utilizzabile solo se corrisponde a una card autorevole + corrente del workspace e a una proiezione valida. Il contenuto restituito + viene dall'archivio autorevole. La verifica si applica anche agli exemplar. + In assenza di verifica autorevole il recall non restituisce il vecchio payload. +- Dopo il commit PostgreSQL, una propagazione fallita lascia il contenuto + consultabile nell'amministrazione e la sua proiezione non utilizzabile dal + recall fino al recupero. L'esito distingue chiaramente questo caso da un + salvataggio fallito prima del commit. +- La cancellazione rimuove card, collegamenti incidenti e dipendenze e conserva + soltanto i dati operativi necessari a eliminare la proiezione. La card non è + più richiamabile anche se il punto Qdrant esiste ancora. La pulizia pendente + resta raggiungibile dalla pagina, senza richiedere il dettaglio della card eliminata. +- Il retry è esplicito e ripetibile. Usa lo stato corrente del repository, non + il contenuto di una vecchia richiesta. Un retry superato non può sovrascrivere + una correzione successiva né ricreare una card cancellata. +- La ricostruzione degli indici Memory e solved-question usa esclusivamente + le card autorevoli. Non reimporta automaticamente il registro JSONL, i payload + Qdrant o le sessioni di origine. Il preprocessing delle reference mantiene + la separazione delle collezioni stabilita nell'ADR 0017. +- Anche i produttori attuali di Memory ed exemplar scrivono attraverso il + servizio autorevole. Si adeguano promozione, salvataggio singolo e percorso di + finalizzazione quanto necessario a evitare scritture dirette al solo indice. + Un errore successivo al commit della sessione non deve annullarne la finalizzazione; + l'esito e il recupero Memory restano espliciti. Questa transizione non introduce + il nuovo riepilogo di approvazione previsto da M3. + +### API, autorizzazione e configurazione + +- Il contratto amministrativo comprende elenco, dettaglio, creazione, + aggiornamento, cancellazione, manutenzione dei collegamenti, stato delle + propagazioni incomplete e retry. Le mutazioni restituiscono identità interessata, + esito del salvataggio ed esito della propagazione; gli errori non espongono segreti. +- L'elenco restituisce pagina, conteggio totale filtrato e ordinamento stabile + con identificatore come discriminante. Ricerca testuale e filtri combinabili + agiscono sull'intero archivio prima della paginazione, senza embedding. +- Il backend distingue input invalido, accesso negato, record assente nel + workspace richiesto, archivio indisponibile e propagazione incompleta dopo + salvataggio. Un archivio indisponibile non produce una lista vuota riuscita. +- Si riusano autenticazione, controlli di accesso e trasmissione del principal. + Il catalogo attuale non ha un permesso Memory dedicato: la proposta è aggiungere + la capability amministrativa Memory al ruolo admin esistente, senza introdurre + ruoli nuovi. Il controllo copre anche letture amministrative e retry. +- Le operazioni amministrative e i comandi esposti non permettono bypass del + controllo nel harness. Le scritture già previste dal workflow conservano il + proprio contesto autorizzato di sessione; non diventano CRUD amministrativo + liberamente accessibile a un utente ordinario. Il recall rimane accessibile + secondo le regole del workflow. +- M1 include configurazione della connessione al PostgreSQL dell'installazione, + distribuzione protetta delle credenziali al harness, migrazioni versionate e + privilegi runtime necessari alle sole tabelle Memory. Si riusano i meccanismi + di configurazione generata, segreti e provisioning esistenti; non si usano le + credenziali di lettura del DWH. I nomi fisici di schema, tabelle e parametri + vengono fissati nell'implementazione rispettando questi contratti. +- Le migrazioni sono eseguite dal percorso di preparazione dell'installazione, + non da una richiesta HTTP ordinaria. Schema mancante o non aggiornato produce + un errore operativo comprensibile. La transizione non prevede doppie scritture + permanenti o conservazione del comportamento delle sessioni storiche. + +### Pagina amministrativa + +- Memory management è un accesso indipendente nell'Administration dell'AppShell, + immediatamente dopo Database management, senza richiedere una sessione attiva + o l'ingresso in Database management. +- La pagina rende esplicito il workspace e offre lista paginata, ricerca, + filtri, ordinamento, dettaglio completo e form. Le modifiche hanno salvataggio, + annullamento e validazione. La cancellazione rende chiari contenuto interessato + e rimozione dei collegamenti, seguendo le convenzioni UI esistenti. +- Il feedback distingue operazione in corso, salvataggio fallito, contenuto + salvato con indice incompleto e operazione completata. Il recupero delle + cancellazioni pendenti è disponibile anche quando la card non compare più in lista. +- I controlli sono accessibili da tastiera e hanno etichette comprensibili. + Chrome e messaggi UI sono in inglese; il contenuto resta nella lingua del workspace. + Hash, versioni tecniche e dettagli delle tabelle non sono esposti nel flusso ordinario. + +## Testing Decisions + +Confini confermati dal proprietario: usare tre confini già presenti nel +repository, concentrando la maggior parte dei casi sul servizio pubblico Memory +del harness. I test osservano risultati, persistenza ed errori; non vincolano +metodi privati, numero di query o disposizione interna delle tabelle. + +### 1. Servizio pubblico Memory e suoi comandi + +Usare PostgreSQL reale in testcontainers, come nei test del repository delle +sessioni, e gli adapter vettoriali sostituibili già impiegati nei test di recall +e del ciclo di vita solved-question. Gli embedding dei casi deterministici sono +controllati. Un gruppo mirato con Qdrant reale verifica aggiornamento, cancellazione +e ricostruzione della proiezione; non richiede DWH remoto o un modello generativo. + +Questo confine verifica il comportamento di archivio, propagazione e recall: + +| Caso | Risultato osservabile richiesto | +| --- | --- | +| Creare e riaprire il repository | La card completa, i collegamenti e le dipendenze sono persistiti; l'origine manuale non contiene sessioni inventate. | +| Salvare una card di ciascuna categoria | Contenuto, ambito e dati specifici sono rappresentabili e leggibili senza dipendere dai tipi del ledger. | +| Cercare un record fuori dalla prima pagina | Filtri combinati, totale e ordinamento si riferiscono all'intero archivio. | +| Fallire una scrittura correlata | Card, collegamenti e dipendenze mantengono tutti lo stato precedente. | +| Indicare una card di un altro workspace | Lettura, mutazione, collegamento e retry non accedono al contenuto estraneo. | +| Cancellare una card collegata | Scompaiono card e collegamenti incidenti; le altre card restano intatte. | +| Rendere embedding o Qdrant indisponibili | Elenco e dettaglio funzionano; il CRUD persiste e distingue la propagazione incompleta. | +| Fallire PostgreSQL prima del commit | Nessun falso salvataggio riuscito e nessun nuovo contenuto propagato. | +| Fallire Qdrant dopo un aggiornamento | Il dettaglio contiene la correzione; il recall esclude il vecchio risultato. | +| Fallire Qdrant dopo una cancellazione | La card non è richiamabile; il lavoro di pulizia resta visibile e recuperabile. | +| Riavviare fra commit e propagazione | Il lavoro incompleto permane e un retry lo completa. | +| Ritentare dopo un errore o un esito incerto | Non si creano duplicati; si applica lo stato corrente senza ripristinare contenuti superati. | +| Indice con punto orfano o versione superata | Recall Memory ed exemplar lo escludono anche se ha il punteggio più alto. | +| PostgreSQL indisponibile durante il recall | Il servizio segnala l'indisponibilità senza servire payload non verificati. | +| Ricostruire dopo modifica o cancellazione | Il contenuto corretto è conservato; nessuna card viene ricreata dalle sessioni o da vecchi indici. | +| Eseguire promozione o finalizzazione corrente | I nuovi contenuti passano dall'archivio; un errore Memory successivo non annulla una sessione già finalizzata. | +| Applicare migrazioni e riavviare | Lo schema è utilizzabile con il ruolo runtime previsto; la preparazione è ripetibile e non richiede privilegi di migrazione nelle richieste ordinarie. | + +I test CLI coprono solo l'adattamento che il servizio non prova: parsing, principal, +workspace, esiti macchina e JSON puro. I test di integrazione riusano le convenzioni +L0 del harness. I test di recall esistenti continuano a verificare che decisioni +già registrate nella sessione e famiglie non ammesse non vengano riproposte. + +### 2. API amministrative Fastify + +Usare l'iniezione HTTP e il runner sostituibile già presenti nei test backend, +seguendo i test delle route Catalog e dell'autorizzazione. Verificare principal +autenticato, admin e utente ordinario; validazione; selezione del workspace; +contratto delle risposte e mappatura degli errori. Includere letture, collegamenti +e retry, non soltanto le mutazioni delle card. + +Questi test provano il confine HTTP e il passaggio al harness. Non si considera +il runner simulato una prova della transazione PostgreSQL o della coerenza Qdrant. +La normale policy CSRF dell'applicazione resta applicata alle nuove mutazioni. + +### 3. Pagina nell'AppShell + +Usare React Testing Library, MSW e le convenzioni dei test di AppShell e Database +management. Verificare ingresso amministrativo, scelta workspace, lista completa, +filtri, dettaglio, form, annullamento, errori, collegamenti e retry dopo cancellazione. +Controllare il comportamento tramite elementi accessibili e contenuto visibile. + +Un percorso browser mirato sullo stack reale collega i tre confini: amministratore +autenticato, creazione manuale, modifica, riapertura della pagina, cancellazione +e verifica dell'assenza nel recall. I test browser con API intercettate provano +interazione e presentazione; non vengono dichiarati prova della persistenza. + +Non si replica l'intera matrice su tutti e tre i confini. PostgreSQL, indice e +recupero sono verificati nel harness; autenticazione e trasporto nel backend; +interazione e feedback nella UI. Il percorso integrato copre il collegamento reale. + +### Verifica della consegna + +Eseguire i test interessati e i gate documentati dei layer modificati, inclusi +typecheck TypeScript, lint Python e build documentale strict. Le verifiche con +PostgreSQL, Qdrant e browser reale hanno esito riportato separatamente; se un +servizio necessario manca, il relativo gate resta aperto. + +M1 non richiede una valutazione della qualità SQL generata da un LLM. I test +deterministici non sono presentati come prova di tale capacità: gli eventuali +casi reali appartengono agli incrementi che cambiano generazione e workflow. + +## Out of Scope + +- Ricerca ibrida, nuova selezione per ambito ed espansione dei collegamenti: M2. +- Riepilogo finale modificabile, nuove categorie nei gate e pulizia dopo una + sincronizzazione fisica del Catalog: M3. M1 ne prepara card e dipendenze. +- Authoring, consolidamento e manutenzione Evidence: E1–E3; risoluzione persistente + congiunta dei conflitti fra Memory ed Evidence: X1. +- Revisione storica delle card, snapshot per vecchie sessioni, migrazione dei dati + di sviluppo o compatibilità con il registro JSONL come archivio operativo. +- Aggiornamento a caldo delle altre sessioni, nuove invalidazioni dello SQL già + generato, coordinamento generale dei lettori paralleli e modalità manutenzione. +- Nuovi servizi PostgreSQL o graph database, code generiche, worker e polling continuo. +- Promozione automatica di rifiuti senza spiegazione o scelte occasionali, + consolidamento automatico e benchmark generale della qualità del modello. +- Deploy o pulizia dell'installazione PSD: restano soggetti ai rispettivi piani e gate. + +## Further Notes + +Fonti: progetto **Memory management** del 2026-09-08; piano comune +**Amministrazione di Memory ed Evidence**; **Revisione di semplicità: Memory ed +Evidence**; glossario di dominio; ADR 0017 sulla separazione delle collezioni e +ADR 0018 su PostgreSQL autorevole e Qdrant per il retrieval. + +La ricognizione ha verificato MemoryRecord, recall ordinario, ricerca degli +exemplar, comandi Memory, repository PostgreSQL delle sessioni, runner del harness, +autorizzazione backend e navigazione amministrativa. Il recall ordinario oggi +risolve già i risultati nel registro autorevole, mentre gli exemplar leggono +contenuti dal payload vettoriale: M1 deve uniformare entrambe le garanzie. + +Le scelte tecniche da fissare durante l'implementazione sono nomi e DDL delle +tabelle, firma esatta dei nuovi comandi/API e parametri generati di connessione. +Devono rispettare i contratti e i casi di accettazione di questa specifica; +non riaprono le decisioni di prodotto approvate. + +La destinazione della specifica è il tracker Gitea canonico di ThothII, con +etichetta **ready-for-agent**. Il proprietario ha confermato i confini di test +e autorizzato la pubblicazione il 2026-09-08. L'implementazione resta da eseguire. diff --git a/docs/plans/2026-09-08-memory-m1-validation.md b/docs/plans/2026-09-08-memory-m1-validation.md new file mode 100644 index 00000000..cba6693a --- /dev/null +++ b/docs/plans/2026-09-08-memory-m1-validation.md @@ -0,0 +1,103 @@ +# M1 — Implementazione e verifica + +Data: 2026-09-08. Implementazione locale della +[specifica approvata](2026-09-08-memory-m1-spec.md), associata all' +[issue 27](https://git.tylconsulting.it/mptyl/ThothII/issues/27). + +## Risultato + +La pagina **Memory management** è disponibile nell'Administration dopo Database +management. Gestisce le quattro famiglie di card, elenco completo, ricerca e filtri, +ordinamento, dettaglio, creazione, modifica, cancellazione, collegamenti e dipendenze. +Richiede un amministratore autenticato e una selezione esplicita del workspace; +non richiede una sessione o un database DWH configurato. + +Il harness possiede l'archivio PostgreSQL `thoth_memory`. Card, collegamenti, +dipendenze e lavoro di propagazione sono salvati nella stessa transazione. +La pagina distingue salvataggio fallito e contenuto salvato con indice incompleto, +offrendo retry anche per le cancellazioni. Recall Memory ed exemplar verificano +esistenza, workspace e proiezione corrente nell'archivio prima di restituire contenuto. + +Promozione, salvataggio singolo e finalizzazione corrente passano dal servizio +autorevole. Le ricevute della sorgente impediscono duplicati e ricreazione di card +cancellate. Reindicizzazione e preprocessing non importano vecchi payload o sessioni. +L'errore Memory non annulla una sessione già finalizzata; il gate segnala anche +una promozione salvata con indicizzazione incompleta. + +Le migrazioni sono versionate, controllate tramite checksum e incluse nel wheel +e nell'immagine core. Il servizio di preparazione `catalog-migrate` le esegue dopo +quelle del Catalog. Il runtime assume il ruolo limitato `thoth_memory_runtime`, +con isolamento del workspace tramite RLS e senza privilegi DDL. + +## Verifiche eseguite + +| Confine | Esito | +| --- | --- | +| Harness, test senza L0/L2 | 1.134 passati; i 9 test dei percorsi portabili sono stati eseguiti separatamente e sono passati. | +| Servizio Memory, PostgreSQL e Qdrant reali | 17 passati, inclusi CLI, migrazioni, ruolo runtime, isolamento, transazioni, outage, retry, cancellazioni, cambio famiglia e rebuild. | +| Gate Pi | 190 passati, inclusi identità UUID e avviso dopo salvataggio con indice incompleto. | +| Backend | 1.345 passati nella suite completa, 40 esclusi dalle condizioni previste dai test; un test di autenticazione ha superato il timeout sotto carico. Il relativo file è stato rieseguito isolato: tutti i 17 test passati. | +| Frontend | 632 passati, inclusi ingresso dall'AppShell, form, filtri, collegamenti, dipendenze e retry delle cancellazioni. | +| Browser integrato | Passato: autenticazione amministratore, creazione, modifica, riavvio del backend, rilettura, cancellazione e assenza nel recall. | +| Build e tipi | Build backend e frontend, typecheck TypeScript e build documentale strict superati. | +| Lint e diff | Ruff sui file Python modificati e `git diff --check` superati. Il lint globale segnala tre rilievi in file non modificati, elencati sotto. | + +Il browser utilizza autenticamente frontend, login locale, Fastify, ThtRunner, +CLI Python, PostgreSQL e Qdrant. Gli embedding sono deterministici e le attività +Pi/sessione estranee al percorso Memory usano le fixture esistenti. Non sono state +intercettate le API Memory. Sono stati usati container temporanei PostgreSQL 16 e +Qdrant 1.18.2, senza accesso a un DWH remoto o a un modello generativo. + +Il test browser ha consentito di correggere etichette accessibili instabili nei +campi compilati e la sovrapposizione del pannello di recupero ai comandi del dettaglio. +La selezione del workspace e l'uscita dalla pagina sono bloccate durante le operazioni. + +Il lint globale preesistente riguarda soltanto: + +- ordinamento import in `harness/tests/test_effective_relationships.py`; +- ordinamento import in `harness/tests/test_p3_dwh_binding.py`; +- uso di `datetime.UTC` in `harness/tht/mschema/catalog_snapshot.py`. + +## Riproduzione + +Usare Node 24 e le dipendenze installate dei tre layer. Per eseguire il harness +in un ambiente con home non scrivibile si può impostare `THT_HOME` su una directory +di prova. I test dei percorsi portabili devono essere eseguiti senza questo override, +perché verificano deliberatamente la risoluzione dell'home e di `THT_DATA_ROOT`. + +```sh +cd harness +THT_HOME=/private/tmp/thothii-m1-test-home .venv/bin/pytest -m 'not l0 and not l2' --ignore=tests/test_portable_paths.py -q +.venv/bin/pytest tests/test_portable_paths.py -q +.venv/bin/pytest tests/memory/test_administration.py -q +npm test +``` + +```sh +cd backend +npx vitest run +npx tsc --noEmit -p . +npm run build +``` + +```sh +cd frontend +npx vitest run +npx tsc -b +npm run build +THT_MEMORY_BROWSER_E2E=1 npx playwright test e2e/memory-real.spec.ts +``` + +Il percorso browser richiede Docker, Python del harness, Go per il bridge di +autenticazione e Chromium di Playwright. Avvia risorse isolate e le rimuove alla +fine. Su macOS il browser deve poter avviare i processi Chromium fuori dalle +restrizioni della sandbox. La build documentale si esegue dalla radice con +`./scripts/build-docs.sh`. + +## Stato della consegna + +Le modifiche sono nel worktree locale. Nessuno stack già attivo è stato aggiornato +e nessun dato esistente è stato migrato o eliminato. Prima di usare M1 su +un'installazione occorrono il nuovo core e la preparazione `catalog-migrate`. +M2 (retrieval ibrido ed espansione dei collegamenti), M3 (integrazione estesa nel +workflow) ed Evidence management restano incrementi successivi. diff --git a/docs/plans/2026-09-08-memory-management.md b/docs/plans/2026-09-08-memory-management.md new file mode 100644 index 00000000..9d7ae87f --- /dev/null +++ b/docs/plans/2026-09-08-memory-management.md @@ -0,0 +1,342 @@ +# Progetto: Memory management + +Data del piano: 2026-09-08. Aggiornamento 2026-09-09: M1–M3 implementati, +con integrazione X1 per le correzioni persistenti dei conflitti. Risultati e limiti +sono raccolti nel [rapporto X1](2026-09-09-archive-repair-x1-validation.md). + +La [revisione di semplicità](2026-09-08-memory-evidence-simplification-review.md) +mantiene il perimetro Memory e precisa un salvataggio sequenziale e una pulizia +diretta dopo sincronizzazione. Non si progetta l'amministrazione contemporanea +all'attività core. + +Il progetto realizza il CRUD amministrativo previsto dalla discussione sull'evoluzione +della Memory. Segue le [decisioni comuni di Administration](2026-09-08-memory-evidence-administration.md) +e rimane distinto dal [progetto Evidence management](2026-09-08-evidence-management.md). + +## Risultato richiesto + +Un amministratore apre Memory management direttamente da Administration, cerca e +filtra l'intero archivio di un workspace e gestisce le card senza avviare una sessione. +La voce è immediatamente sotto Database management, allo stesso livello. + +L'elenco comprende le Memory riutilizzabili e gli exemplar `solved_question`, +distinguibili per famiglia. La modifica di un exemplar non riscrive gli artefatti +della sessione da cui deriva e non lo trasforma in una decisione applicabile al gate. + +## Contenuti ammessi: decisione del proprietario + +Il perimetro concordato il 2026-09-08 comprende quattro categorie di contenuto. +La classificazione descrive il valore della conoscenza; non impone quattro nuovi +kind tecnici o una corrispondenza con i tipi delle decisioni del ledger. + +| Contenuto | Cosa conserva | Esempio inventato | +| --- | --- | --- | +| Chiarimento di dominio | Significato riutilizzabile di un termine, con il suo ambito | «In questo workspace, ordine evaso significa che tutte le righe sono state spedite.» | +| Regola SQL | Regola corretta per join, filtri o aggregazioni, con condizioni e motivazione | «Il codice commessa è univoco solo all'interno dell'esercizio: collegare movimenti e commesse usando codice ed esercizio.» | +| Domanda risolta | Domanda, SQL approvato e contesto, come exemplar consultativo | «Totale degli ordini del 2024 per cliente», con la query che lo calcola. | +| Errore da evitare | Errore compreso, motivo e comportamento corretto approvato | «Il join fra ordini e righe moltiplica il totale di testata: calcolare il totale una sola volta per ordine.» | + +Il criterio di ammissione è l'utilità per altre domande nello stesso ambito. +Una scelta come «questa volta usa il 2024» non è una regola riutilizzabile; il 2024 +può rimanere nel contesto della domanda risolta. Analogamente, la selezione di una +tabella per una domanda non diventa automaticamente una regola di schema linking. + +La categoria «errore da evitare» richiede una spiegazione verificata e approvata. +Un timeout, una query rifiutata senza motivo o una proposta non selezionata non +bastano a produrre conoscenza. Quando errore e correzione esprimono la stessa +regola, una sola card conserva la regola e la sua motivazione. + +## Formazione, aggiornamento e uso delle card + +### Decisioni del primo round di grill-with-docs + +Il proprietario ha approvato le tre raccomandazioni il 2026-09-08: + +- **Q1, approvazione del salvataggio:** riepilogo finale modificabile, preparato + durante il lavoro. Il reviewer corregge le card e sceglie quali salvare; la + creazione manuale da Administration resta sempre disponibile. +- **Q2, operazioni proposte dal core:** aggiunte e aggiornamenti espliciti. Il + riepilogo distingue una nuova card dalla modifica di una card esistente e ne + spiega il cambiamento. La sostituzione richiede la selezione del reviewer; + la cancellazione semantica resta un'operazione del CRUD amministrativo. La pulizia + automatica dei riferimenti invalidi è disciplinata separatamente da Q5. +- **Q3, momento del consumo:** i chiarimenti sono proposti all'inizio, le regole di + collegamento durante lo schema linking e le regole di calcolo durante la costruzione + SQL. Le approvazioni entrano nei gate pertinenti, senza una domanda separata per + ciascuna card; gli exemplar rimangono consultativi. + +Queste sono decisioni di prodotto; i contratti runtime non sono ancora aggiornati. + +### Decisioni del secondo round di grill-with-docs + +Il proprietario ha approvato i chiarimenti su Q4–Q6 il 2026-09-08: + +- **Q4, risoluzione persistente dei conflitti:** il gate propone azioni chiuse e + specifiche per il caso, mostrando record interessati, azione e testo o ambito + risultante. Le opzioni possono confermare l'Evidence e correggere la Memory, + confermare la Memory e preparare una correzione dell'Evidence, oppure precisare + gli ambiti distinti di entrambe. È sempre disponibile «Nessuna proposta è adeguata», + che richiede una riformulazione. Le modifiche Memory confluiscono nel riepilogo + finale; quelle Evidence seguono l'authoring e la pubblicazione del rispettivo + modulo. Come approvato in Q12, accettare la correzione Evidence con le autorizzazioni + necessarie avvia anche l'attivazione automatica, senza un ulteriore `Publish`. + Il sistema distingue proposte da approvare, aggiornamenti in corso o falliti e + archivio già aggiornato. La sola risoluzione della domanda corrente non esaurisce il flusso. +- **Q5, cancellazione dopo modifiche allo schema:** dopo una sincronizzazione + riuscita dello schema fisico, il backend comunica al modulo Memory gli elementi + rimossi. Il modulo identifica tramite dipendenze strutturate le card non più + valide e cancella record e proiezioni ricercabili. Non si introduce lo stato + «Needs review» per conservarle. Il controllo avviene alla sincronizzazione, senza + scansione continua del DWH o interrogazioni aggiuntive a ogni domanda. Un cleanup + manuale del Catalog o un errore di accesso al database non prova una rimozione + fisica e non avvia questa pulizia. Essa è distinta dalle proposte semantiche del + core in Q2. Oggi mancano sia i riferimenti strutturati a colonne nelle Memory sia + il collegamento fra sincronizzazione e pulizia: devono essere implementati. +- **Q6, verifiche concrete:** lo sviluppatore prepara ed esegue test automatici + funzionali per CRUD, filtri, approvazioni, aggiornamenti e cancellazioni. Quando + cambia la ricerca, verifica casi mirati con card necessarie e card fuori ambito; + quando emerge un errore SQL riproducibile, aggiunge una regressione su dati + controllati confrontando i risultati, senza richiedere un identico testo SQL. + L'esperto di dominio conferma inizialmente regola e risultato atteso soltanto per + i casi reali che lo richiedono. La verifica della generazione necessita di un + modello reale ed è separata dalla suite deterministica: una query scritta a mano + non prova che il modello sappia generarla. Non si introduce una valutazione umana + permanente o un benchmark generale con percentuali di miglioramento promesse. + +### Terzo round: direzione tecnica e capacità di ricerca + +- **Q7, deciso:** il proprietario ha approvato l'evoluzione interna di ThothII. + Il modulo riusa l'infrastruttura dell'installazione e integra card, CRUD, mutazioni + e recall con i gate; non adotta un framework esterno per governare la Memory. +- **Q8, deciso:** il proprietario ha approvato ricerca ibrida in Qdrant e collegamenti + espliciti fra card gestiti dal core, senza un database a grafi aggiuntivo. Entrambe + le capacità sono incluse nella pianificazione attuale; la presenza dei collegamenti + non è rinviata alla futura comparsa di casi concreti. + +La configurazione approvata comprende: + +- ricerca semantica e lessicale ibrida in Qdrant, con filtri sull'ambito; +- collegamenti espliciti fra card, proposti e revisionabili, percorsi nel core con + espansione limitata e riordinamento dei risultati insieme a quelli della ricerca; +- persistenza dei collegamenti coordinata con le card, con rimozione dei riferimenti + a contenuti cancellati e rispetto dei confini fra workspace e dei gate; +- nessun servizio di database a grafi aggiuntivo. + +La scelta include il grafo logico, senza introdurre un servizio di graph DB. +La qualità non è garantita dalla scelta di un motore: mantenere queste capacità +evita una rinuncia architetturale ai collegamenti, ma non dimostra equivalenza +qualitativa con qualsiasi soluzione basata su graph DB. +Qdrant è il motore di ricerca; l'archivio autorevole è PostgreSQL, scelto in Q9. + +Le [query ibride di Qdrant](https://qdrant.tech/documentation/search/hybrid-queries/) +e i [filtri sui metadati](https://qdrant.tech/documentation/search/filtering/) +coprono le capacità di ricerca indicate. La logica dei collegamenti di dominio +nel core è lavoro applicativo da implementare. + +### Decisioni del quarto round di grill-with-docs + +- **Q9, archivio autorevole:** il proprietario ha approvato PostgreSQL, già presente + nell'installazione, con tabelle proprie del modulo Memory per card, collegamenti + e dipendenze dallo schema. Sostituisce il registro JSONL; Qdrant è l'indice + rigenerabile. Le modifiche correlate vengono coordinate in PostgreSQL e la + propagazione a Qdrant deve gestire esplicitamente errori e cancellazioni. +- **Q10, gestione dei collegamenti:** il core propone i collegamenti insieme alle + card, indicando destinazione e significato. Il reviewer li approva nello stesso + riepilogo finale, senza un gate aggiuntivo. Administration ne consente creazione, + modifica e cancellazione manuali. Quando una card è cancellata vengono rimossi + anche i collegamenti che la coinvolgono, conservando le altre card. I collegamenti + contribuiscono al recupero e non applicano automaticamente i contenuti. + +La decisione architetturale è registrata nell'[ADR 0018](../adr/0018-use-postgres-for-memory-and-qdrant-for-retrieval.md). + +### Flusso da implementare + +Il flusso seguente traduce le decisioni approvate. I payload e l'integrazione con +i gate sono dettagli da definire nell'implementazione, senza altre decisioni di +prodotto pendenti. + +1. Durante il lavoro il core individua possibili conoscenze riutilizzabili a partire + da decisioni e artefatti registrati. Una candidata esplicita cosa afferma, dove + vale, perché è utile e su quale correzione o decisione si basa. +2. Prima di proporne il salvataggio confronta la candidata con le card correnti. + Un doppione esatto non richiede una nuova card; una somiglianza semantica non + autorizza da sola a eliminare o sovrascrivere una conoscenza. +3. Alla conclusione del lavoro presenta un riepilogo editabile delle aggiunte e + degli aggiornamenti proposti. Il reviewer può correggere il contenuto, restringere + l'ambito e scegliere cosa salvare; approvare la query non equivale ad approvare + ogni generalizzazione ricavata dalla query. +4. Una correzione alla stessa regola nello stesso ambito propone un aggiornamento + esplicito della card esistente. Regole valide in ambiti diversi restano distinte; + un conflitto irrisolto non viene risolto silenziosamente dal modello. +5. Le card salvate sono subito consultabili in Memory management; la loro + disponibilità al recall segue lo stato di indicizzazione. La scrittura sostituisce + il contenuto corrente senza introdurre una cronologia delle Memory. + +La creazione manuale da Memory management resta disponibile in qualsiasi momento +e non dipende dal riepilogo finale di una sessione. La form richiede contenuto e +ambito adeguati alla famiglia e identifica l'origine amministrativa. + +Gli exemplar conservano domanda e soluzione approvata come materiale consultativo. +Il loro salvataggio non applica le scelte di quella soluzione a domande successive. +La distribuzione del consumo nei passaggi pertinenti è decisa in Q3. Restano da +definire i payload e l'integrazione con i gate esistenti, compreso il contesto necessario +a proporre una regola di collegamento o di calcolo e a registrarne l'approvazione. + +### Scenari per verificare il design + +| Evento | Esito atteso | +| --- | --- | +| Il reviewer corregge un join perché il codice commessa si ripete fra esercizi e approva la spiegazione. | Proporre la regola con entrambe le chiavi e l'ambito delle tabelle interessate. | +| Il reviewer chiede di limitare solo la domanda corrente al 2024. | Nessuna regola generale; mantenere il periodo nell'eventuale exemplar. | +| Una query conta più volte lo stesso ordine e la correzione viene spiegata e approvata. | Proporre una card che descrive la granularità corretta e il rischio di duplicazione. | +| Una query fallisce per timeout o una memory non viene selezionata. | Nessuna nuova regola dedotta automaticamente dall'evento. | +| La candidata ripete esattamente una regola già presente nello stesso ambito. | Evitare una nuova card duplicata. | +| Una nuova regola corregge una card dello stesso ambito. | Mostrare la sostituzione proposta prima del salvataggio; conservare poi solo il contenuto corrente. | +| Due regole differenti valgono per processi o tabelle differenti. | Conservare entrambe con ambiti espliciti, senza generalizzarle al workspace intero. | +| Il reviewer risolve un contrasto fra Memory ed Evidence. | Mostrare una correzione esplicita degli archivi; applicare i percorsi distinti per Memory ed Evidence. L'accettazione autorizzata della correzione Evidence avvia anche l'attivazione; indicare esito, operazione in corso o errore. | +| Una sincronizzazione riuscita accerta la rimozione di una colonna da cui dipende una card. | Cancellare la card dipendente e rimuoverla dai risultati di ricerca. | +| La connessione al DWH fallisce oppure vengono puliti solo metadati del Catalog. | Non interpretare l'evento come prova di rimozione della colonna e non cancellare Memory per quel motivo. | + +## Differenza rispetto al runtime corrente + +`harness/tht/memory/core.py` limita `REUSABLE_TYPES` a `concept_clarified`. +Anche il contratto Pi di F2 ammette soltanto questi chiarimenti; gli exemplar +`solved_question` hanno già un percorso distinto di consultazione. + +Il perimetro concordato amplia quindi il modulo Memory. Il design deve distinguere +la conoscenza riutilizzabile dall'evento di workflow che l'ha originata, e aggiornare +insieme estrazione, validazione, persistenza, recall e gate. Aggiungere alla whitelist +tutti i tipi delle decisioni SQL o sulle tabelle promuoverebbe anche scelte occasionali +e non realizza il requisito. + +## Funzioni + +- Elenco paginato e ordinabile, ricerca per testo o identificatore. +- Filtri combinabili per workspace, famiglia/kind, concetti, tabelle e colonne + quando presenti, provenienza e data di aggiornamento. +- Dettaglio completo: titolo, contenuto, ambito, motivazione e provenienza disponibile. +- Creazione manuale di una card, distinguibile da una card prodotta dal workflow; + la creazione manuale non inventa una sessione o una decisione di origine. +- Modifica dei campi consentiti dalla famiglia, con validazione e annullamento. +- Cancellazione del record e rimozione delle sue proiezioni ricercabili. +- Indicazione di contenuti salvati ma non ancora disponibili al recall, con retry + dell'operazione necessaria a renderli disponibili. + +L'elenco amministrativo legge i record persistiti senza richiedere embedding o +ricerca per similarità. Un'indisponibilità dell'archivio deve produrre un errore +esplicito, distinguibile da un elenco vuoto. I filtri sono applicati sull'intero +archivio, prima della paginazione, e non sui soli risultati del recall. + +## Comportamento delle modifiche + +Una modifica sostituisce il contenuto corrente. Non si introducono revisioni storiche, +snapshot dedicati alle vecchie sessioni o migrazioni per conservarne il comportamento. +La cancellazione toglie la card dall'archivio e dal recall ordinario. + +La mutazione deve aggiornare o invalidare ogni proiezione interessata. Un errore +dell'indice non può essere presentato come piena disponibilità del nuovo contenuto, +né permettere di usare silenziosamente il contenuto eliminato o sostituito. +La strategia di consistenza e di retry appartiene al design tecnico del modulo. + +Il salvataggio esplicito dell'amministratore cura il contenuto condiviso. Il suo +successivo consumo nel core mantiene la semantica della famiglia: le Memory vengono +proposte secondo i gate del workflow, gli exemplar restano consultativi. + +## Piano esecutivo + +### M1 — Archivio e CRUD amministrativo + +Definire nel modulo `harness/tht/memory/` il contratto delle card: identità stabile, +workspace, contenuto, famiglia, ambito, motivazione, provenienza e dati specifici +delle domande risolte. I riferimenti allo schema identificano database, tabella e +colonna senza affidarsi alla sola presenza di nomi nel testo. Le card manuali +non richiedono sessioni inventate. + +Implementare un repository PostgreSQL del modulo con card, collegamenti e dipendenze. +La transazione aggiorna insieme il contenuto e le modifiche correlate; la propagazione +a Qdrant avviene nello stesso flusso di salvataggio. Conservare un'indicazione +persistente dell'operazione incompleta, sufficiente anche a ripulire cancellazioni +dopo un riavvio; il recupero usa un retry esplicito. Non servono una coda generale, +un nuovo worker o una sincronizzazione continua. Un risultato indicizzato +con contenuto superato o privo di card autorevole non può essere usato dal recall. +Gli exemplar passano anch'essi dall'archivio autorevole. La transizione dal registro +JSONL non introduce scritture doppie permanenti o compatibilità storica delle sessioni. + +Esporre attraverso il backend elenco filtrato prima della paginazione, dettaglio, +creazione, aggiornamento, cancellazione, gestione dei collegamenti ed esito della +propagazione. Le operazioni chiamano la logica del harness e applicano controllo +amministrativo e isolamento del workspace. La UI legge il repository attraverso +queste API anche quando il servizio di embedding o Qdrant è indisponibile. + +Consegnare Memory management nell'AppShell con form, contenuto completo, gestione +dei collegamenti e feedback di salvataggio/indicizzazione. Verificare persistenza +PostgreSQL, rollback delle mutazioni correlate, aggiornamento e rimozione dal recall, +retry dopo errore dell'indice, filtri sull'intero archivio e autorizzazioni. Una +cancellazione elimina i collegamenti incidenti conservando le altre card. + +### M2 — Ricerca ibrida e collegamenti + +Estendere l'adapter Qdrant alla ricerca dense e lessicale della Memory e applicare +l'ambito anche ai risultati raggiunti attraverso collegamenti. Le card iniziali +alimentano l'espansione limitata nel core; deduplicazione, gestione dei cicli e +limiti espliciti impediscono una visita incontrollata dell'archivio. I risultati +vengono riordinati insieme e risolti contro il contenuto autorevole corrente. + +Verificare card attese, esclusioni per ambito, cicli, collegamenti verso card rimosse +e rigenerazione dell'indice da PostgreSQL. Quando si verifica il recupero effettivo, +usare il percorso di embedding e ricerca configurato su un indice isolato: un fake +che restituisce gli ID predisposti verifica soltanto il contratto applicativo. +La separazione fra le collezioni Reference e Memory rimane quella degli ADR 0017 e 0018. + +### M3 — Workflow e sincronizzazione fisica + +Aggiornare insieme contratti, CLI, regole Pi e widget necessari al riepilogo finale +modificabile. Il salvataggio applica soltanto card e collegamenti selezionati; le +nuove categorie entrano nei gate pertinenti. Il recupero di una regola non ne +costituisce approvazione, e l'exemplar continua a essere consultativo. + +Collegare la sincronizzazione fisica del Catalog alla pulizia delle dipendenze nel +modulo Memory con una chiamata diretta dopo l'applicazione riuscita dello schema, +con copertura del controllo e riferimenti rimossi. Per recuperare un'interruzione +fra applicazione e pulizia, conservarne lo stato pendente oppure verificare di +nuovo le dipendenze contro lo snapshot fisico riuscito e il suo ambito: il nuovo +diff da solo perderebbe le rimozioni già applicate. La pulizia è ripetibile e +non richiede un sistema generale di consegna eventi. +Un confronto parziale non prova la rimozione di elementi fuori dall'ambito controllato. +La pulizia aggiorna archivio, collegamenti e proiezioni senza un'azione manuale ulteriore. + +Verificare selezioni e rifiuti nel riepilogo, contenuto manuale, categorie ammesse, +notifica di rimozione fisica, errore di connessione e cleanup del solo Catalog. +Integrare le correzioni che riguardano Evidence nell'incremento congiunto X1, dopo +il completamento del relativo servizio di authoring e attivazione. + +L'evoluzione interna è decisa in Q7; ricerca ibrida e grafo nel core in Q8; +PostgreSQL autorevole in Q9; gestione dei collegamenti in Q10. I contratti tecnici +di persistenza, indicizzazione e API devono attuare queste decisioni. Apprendimento +automatico da rifiuti non spiegati e consolidamento automatico restano fuori dal +perimetro concordato; gli errori compresi e approvati rientrano nei contenuti decisi. + +## Criteri di completamento + +- Accesso amministrativo indipendente da sessioni e da Database management. +- Tutti i record sono raggiungibili con elenco, filtri e paginazione, senza dipendere + dalla disponibilità di embedding e recall semantico. +- Creazione, modifica e cancellazione persistono dopo riapertura della pagina. +- Dopo una mutazione completata il recall usa il contenuto corrente; i record + cancellati non riappaiono dopo reindicizzazione o preprocessing. +- Errori di salvataggio e indicizzazione sono distinguibili e recuperabili. +- API e interfaccia rispettano isolamento dei workspace e accesso amministrativo. +- Nessuna operazione del CRUD modifica Evidence o metadati del database. +- Non vengono richieste compatibilità storica o conservazione delle sessioni esistenti. +- Le quattro categorie concordate sono rappresentabili senza promuovere le scelte + occasionali a regole generali; gli scenari di ammissione verificano il confine. +- La rimozione fisica accertata di una dipendenza elimina le card interessate; + errori di connessione e cleanup del Catalog non vengono scambiati per rimozioni. +- Le scelte sui conflitti producono correzioni persistenti esplicite secondo Q4. +- La verifica rispetta Q6, separando contratti funzionali e casi di generazione reale. +- Il recupero combina ricerca ibrida, filtri d'ambito e collegamenti espliciti fra + card; la gestione del grafo non richiede un servizio di database aggiuntivo. +- Card, collegamenti e dipendenze hanno un'unica fonte autorevole PostgreSQL; + la rigenerazione di Qdrant conserva il contenuto corrente e le cancellazioni. +- I collegamenti sono curabili nel riepilogo e in Administration; cancellare una + card elimina i suoi collegamenti senza cancellare altre card. diff --git a/docs/plans/2026-09-09-archive-repair-x1-validation.md b/docs/plans/2026-09-09-archive-repair-x1-validation.md new file mode 100644 index 00000000..b88e981f --- /dev/null +++ b/docs/plans/2026-09-09-archive-repair-x1-validation.md @@ -0,0 +1,112 @@ +# X1 — validation of session archive corrections + +Date: 2026-09-09. The joint Memory/Evidence repair increment is implemented. +The authoritative contract is [Session archive corrections](../contracts/archive-repair.md). + +## Delivered behavior + +The session gate shows complete before/after content for specific alternatives targeting +Memory or Evidence. The reviewer chooses one correction or rejects all proposals as +inadequate and requests reformulation. The resulting receipt survives interruption; +saved content and index activation are reported separately. Pending activation offers +retry of the same chosen operation. A subsequent curator change blocks replay. + +Application requires an administrator in the harness and the responding browser +principal's archive-management permission. Cross-principal runtime responses cannot +misattribute the correction. A non-administrator can decline or continue the current +question without modifying shared archives. The gate does not advance a workflow phase. + +Memory and Evidence remain separate domains. The integration coordinator reuses their +canonical persistence and activation operations. A session Evidence correction requires +a consolidated archive, preserves source lineage, and cannot publish unrelated external +edits. Existing administration, source import, dependency cleanup and final Memory review +remain available. No automatic Git commit or push was added. + +## Verification + +- Harness regression: **1,264 passed**, one skipped, five deselected. All **nine** + portable-path checks passed separately without `THT_HOME`. Python lint passed on changed modules. +- Backend: **1,366 passed**, 40 skipped. Tests include actual session-response routes + for both target archives, unauthorized response, malformed choice and runtime ownership. +- Frontend: complete suite **645 passed**; the final display adjustment passed all four + focused widget tests. Backend and frontend TypeScript checks passed. +- Pi extension: **199 passed**, including closed human choices, rejection, failure/retry, + forged selections and the updated public tool schema. The modular skill projection is + byte-identical to its updated approved template. +- PostgreSQL/Qdrant integration traverses the actual Python CLI for preparation, + application and recovery inspection. Corrected Memory and Evidence are retrieved from + real indexes, and Evidence activation preserves the Memory card. Session loading is a + controlled fixture and embeddings are deterministic; this is not an LLM quality test. +- Failure tests cover both targets, index outage, replay, later edits, workspace/session + isolation, changed session context, rejection, non-admin writes, and interruption after + the Evidence file write but before its saved receipt. +- Playwright desktop/mobile: **one passed**. The real widget renders both alternatives, + accepts an Evidence choice, displays pending activation and allows retry to active. + No page errors or mobile horizontal overflow. Screenshots are + `/private/tmp/thothii-x1-repair-desktop.png` and `/private/tmp/thothii-x1-repair-mobile.png`. + This browser fixture controls operation outcomes; persistent behavior is tested above. +- Strict MkDocs build and `git diff --check` passed. + +## Local installation and reviewer acceptance + +Core and frontend images were rebuilt from this worktree using the existing local +preview launcher. Migration `004_archive_repairs.sql` was applied to the existing +installation catalog. The new gate is available to session workflows; it is not an +always-visible administration panel. Existing PSD archive content was not changed by +the synthetic validation cases. +All five local services are healthy at `http://127.0.0.1:8080/`. + +The technical increments and their planned checks are complete. The end-user acceptance +check remains a real session containing a meaningful domain conflict, with the reviewer +evaluating the proposed correction. Automated browser validation uses temporary accounts +and data, not the user's authenticated PSD session. Source import retains its E3 +validation boundaries; no broader model-quality benchmark was added. + +## Follow-up acceptance: configured model + +The opt-in `test_real_model_proposes_a_reviewable_persistent_archive_correction` +passed with the installation's **zai/glm-5.3** model. Synthetic Memory asserted an +order-ID-only join; synthetic Evidence required the financial year too. The model +returned two schema-valid, specific alternatives with complete content and the exact +target revisions. The test reviewer selected Memory, persisted the correction through +the real coordinator and PostgreSQL, and retrieved the updated rule. Evidence stayed +unchanged. This test uses deterministic vectors and the configured completion helper; +it does not claim a full autonomous Pi session or human acceptance of PSD semantics. + +The run log is `/private/tmp/x1-acceptance-model.log`. Reproduce with +`THT_MEMORY_L2_INSTALLATION=` and `THT_MEMORY_L2_CORE=` +using `pytest -q -s -m l2 tests/memory/test_administration.py -k real_model_proposes`. +Credentials are resolved inside core and are not returned to the test runner. + +## Follow-up acceptance: both administration pages + +The opt-in `frontend/e2e/memory-real.spec.ts` passed through real authentication, +Fastify, ThtRunner, Python, isolated PostgreSQL and Qdrant. It verifies: + +- Database management, Memory management and Evidence management appear as peers in + that order, with no active core session or DWH binding required. +- Memory creation, editing, persistence across backend restart, deletion and absence + from subsequent recall. +- Canonical Evidence remains intact after the Memory deletion. Its full rule is read + through the real Evidence administration worker; content filtering finds it and an + unmatched filter produces the empty state. +- Requests for an unregistered workspace return 404 for both archives. +- Desktop and mobile Evidence views render without horizontal document overflow. + On phones, both archive pages have at least 380px of usable width at a 390px viewport. + Navigation opens in the shared accessible dialog, closes with Escape or archive selection, + and returns focus to the trigger after Escape. + +The temporary PostgreSQL readiness probe now waits for TCP, avoiding the image's +socket-only initialization server. The browser waits for Memory refresh to finish +before leaving its page, matching the existing navigation guard. Visual inspection +also exposed a real mobile layout issue: the fixed sidebar left only 134px for the +Evidence page. `ArchiveNavigation` now moves that sidebar into the shared dialog below +768px on Memory/Evidence pages. Desktop behavior is unchanged. The frontend image +was rebuilt for the local preview. + +Run log: `/private/tmp/x1-acceptance-browser7.log` (**one passed**). +Screenshots: `/private/tmp/thothii-acceptance-evidence-desktop.png` and +`/private/tmp/thothii-acceptance-evidence-mobile.png`. Reproduce with +`THT_MEMORY_BROWSER_E2E=1 npx playwright test e2e/memory-real.spec.ts` from `frontend/`. +The fixture removes its temporary containers, accounts and checkout on completion. +The TypeScript check, Python lint, strict documentation build and diff check also pass. diff --git a/docs/plans/2026-09-09-evidence-e1-validation.md b/docs/plans/2026-09-09-evidence-e1-validation.md new file mode 100644 index 00000000..186f3777 --- /dev/null +++ b/docs/plans/2026-09-09-evidence-e1-validation.md @@ -0,0 +1,67 @@ +# Evidence E1 — validation + +Date: 2026-09-09. Scope: editable Curated Evidence v4 and the persistent local archive. + +## Implemented behavior + +- Parser, renderer, authoring output and normalization share the existing typed payloads. + Visible Markdown edits determine content for all eight kinds. Legacy v1–v3 conversion + is explicit and lossless, with errors for content that cannot be represented exactly. +- Manual declarations record the curator. A correction preserves the original document + as lineage, separately from the current declaration. No source hash is needed to + create a manual file. +- The local archive records baselines, immutable candidates, active revisions and + deletion/source suppression metadata. Unresolved review items and invalid edits block + consolidation. Missing archive directories are availability failures, not deletions. +- Activation failures preserve the previous active revision. Interrupted normalization + replays only unchanged input bytes; later operator edits survive recovery. +- Revision-checked correction methods reject stale workflow updates. Legacy preparation + and resolution cannot overwrite an initialized local archive; explicit import/refresh + integration is deferred to E3. + +## Verification + +The final harness suite excluding opt-in L0/L2 and portable-layout cases passed with +**1,180 tests** (58 deselected). All **9 portable-layout tests** passed separately with +`THT_HOME` unset. The dedicated real-Qdrant integration test passed, including the +optional 35-unit PSD probe. Ruff passed on the changed Evidence implementation and +tests, and the strict documentation build succeeded. No frontend or backend TypeScript +changes are part of E1. + +The integration test uses an isolated Qdrant 1.18.2 container, the actual corpus +pipeline, semantic chunking, vector adapter and active Evidence searcher. Deterministic +three-dimensional embeddings isolate file/content correctness from model behavior. +It verifies that raw edits do not change recall, consolidation updates recalled content +and curator identity, a blocked candidate preserves prior recall, and deletions remove +recall. Existing schema and Memory records survive each operation. + +All **35 PSD units** were copied from the owner's workspace into +`/private/tmp/thothii-e1-psd.bsW4cp`. Deterministic conversion preserved every ID, payload, +scope, provenance and review item. There were no unresolved review items. The optional +integration probe then indexed all 35 converted units and compared their complete ID +set to the original. It uses PSD's actual `max_chunk_chars: 5000`; a preliminary probe +at 4000 correctly blocked an oversized atomic unit. + +Reproduce the isolated real-corpus probe after creating a converted workspace copy: + +```sh +cd harness +THT_E1_PSD_COPY=/absolute/path/to/converted-copy \ + .venv/bin/pytest -q -s tests/test_evidence_editable_integration.py +``` + +The environment variable is optional. Ordinary CI uses only synthetic Evidence. No +source refresh, external document download, DWH call or model request is involved. + +## Delivery boundary + +E1 is a core/library increment. E2 must add the installed manual consolidation command, +connect runtime source selection to the active local snapshot, and build administrative +list/filter/detail with real persistent host paths and manual Git instructions. E3 +adds source acquisition and explicit refresh/conflict handling. X1 later wires deliberate +joint Memory/Evidence corrections into review gates. + +The actual PSD Evidence checkout was not converted. The live Docker preview at +`http://127.0.0.1:8080` remains the previously deployed M3 stack, with no new Evidence +administration page. The corpus conversion and reindexing described here used copies +and disposable test resources. diff --git a/docs/plans/2026-09-09-evidence-e2-validation.md b/docs/plans/2026-09-09-evidence-e2-validation.md new file mode 100644 index 00000000..272733d3 --- /dev/null +++ b/docs/plans/2026-09-09-evidence-e2-validation.md @@ -0,0 +1,86 @@ +# Evidence E2 — validation + +Date: 2026-09-09. E2 is implemented locally and installed on the existing Docker preview. +E3 source imports/refresh and X1 deliberate Memory/Evidence workflow corrections remain open. + +## Delivered behavior + +- Independent **Administration → Evidence management**, after Memory, protected by + `evidence.manage`: complete typed content, provenance and original excerpts, review items, + pagination, search, kind/purpose/status and scope/source filters, sort, and refresh. +- Working-file states distinguish active, modified, new, removed, invalid, legacy and review + required. Detail shows the actual configured host path with copy controls. Instructions + cover external editing, all eight Markdown templates, consolidation and manual Git. +- Installed `tht workspace evidence consolidate --workspace [--json]` uses a closed + maintenance envelope. First use converts legacy units. Validation, immutable candidates, + activation and retry run through the existing corpus pipeline without a DWH scan or Git. +- Runtime and ordinary preprocessing consume the active local snapshot. Unconsolidated + edits remain excluded. Catalog/Schema readiness is not advanced by this operation. + Clear preserves curated files, archive metadata and Memory; full preprocessing must + recreate the missing Reference/Schema derivations afterward. +- Immutable runtime lease filenames now identify rendered bytes as well as logical input + identity. This fixes upgrades colliding with old runtime files without changing Catalog + fingerprints or removing the checks against tampered files. + +## Automated checks + +The complete backend suite passed: **1,355 tests**, 40 skipped. The complete frontend +suite passed: **639 tests**. Both TypeScript checks passed. Native Go workspace operation +and CLI tests passed, including rejection of arbitrary consolidation flags. Ruff passed +for changed Python implementation and test files. + +The harness run passed **1,248 tests**, with one skipped and five deselected. Its three +portable-path tests failed because that run deliberately set `THT_HOME` to the test +runtime; rerunning the portable tests with `THT_HOME` unset passed. The final focused +administration/path suite passed all 18 tests, including actionable migration errors and invalid +consolidation combinations rejected before cleanup or indexing. + +The real-Qdrant integration test exercised the actual harness consolidation CLI with +deterministic embeddings: all 35 PSD units converted and indexed with stable identities; +active-only source selection; separate Schema and Memory canaries; Clear and rebuild +from the retained snapshot. Unit tests cover validation, saved-but-unindexed failure, +retry, browsing/filtering, no automatic Git, and no Catalog mutation from consolidation. + +A temporary Git repository and bare local remote exercise the documented manual sequence: +edit, add and remove files, consolidate, inspect, stage the complete Evidence tree, +commit, push and clone. The clone retains changed content, additions, deletions, managed +metadata and an accessible active snapshot. No remote user repository was pushed. + +## Installed preview + +The existing Compose project is `thothii-18998cca7b0a`, at `http://127.0.0.1:8080`. +The persistent editable checkout is: + +```text +/Users/mp/projects/ThothII/deploy/psd/evidence-registry/repo/psd-clinical/evidence +``` + +The original registry checkout was copied from its retained Docker volume. The original +author repository was not changed. Core and maintenance share a nested host bind for +`repo`; registry state/snapshots and all other existing data volumes were retained. +The installation descriptor includes the existing workspace bindings and the new +`evidence-host.yaml` override. The previous descriptor and native binary are backed up +at `/private/tmp/thothii-installation-before-e2.yaml` and `/private/tmp/tht-before-e2`. + +The real installed command succeeded with **35 documents, 35 chunks, 35 changed, zero +removed**, using the configured embedding service and Qdrant. A second run succeeded +with **35 unchanged, zero changed**. Reading the actual archive from core returned +35 active units and no file errors. All five long-running services are healthy. +Only the Evidence stage ran. The strict documentation build and `git diff --check` +also passed. The stack launcher is `bash /private/tmp/thothii-memory-preview.sh`; keep its +worktree image-build override until this branch is integrated into the main checkout. + +Browser verification reached the local login page. The saved administrator password +does not match the current account hash, so the authenticated visual check remains +manual. No account or password was modified. React interaction tests cover navigation, +detail, host paths, templates, filtering, pending activation and invalid files. + +## Boundaries + +There is no web content editor, watcher, automatic commit/push, or implicit source refresh. +The API exposes administration reads and consolidation; the archive's revision-checked +save/remove operations remain available for the later explicit workflow corrections. +These gates are not claimed as implemented by E2. Initialized local archives retain +structural/review checks but bypass the legacy fixed retrieval-evaluation fixture so +its old expected IDs cannot veto deliberate deletions. A general retrieval benchmark +is outside the agreed scope. diff --git a/docs/plans/2026-09-09-evidence-e3-validation.md b/docs/plans/2026-09-09-evidence-e3-validation.md new file mode 100644 index 00000000..5180a6cf --- /dev/null +++ b/docs/plans/2026-09-09-evidence-e3-validation.md @@ -0,0 +1,100 @@ +# Evidence E3 — validation + +Date: 2026-09-09. Explicit source import/refresh and decisions are implemented. X1, +the integration of deliberate Memory/Evidence corrections into workflow gates, remains next. + +## Delivered behavior + +The independent Evidence page now offers **Sources and imports**. An operator copies +a specialist's draft into `evidence/incoming/`, then explicitly imports/refreshes. +Original local Markdown and configured HTTP/S3 sources use existing read-only adapters. +Acquisition retains raw bytes, source identity and versioned normalized documents. +The existing Pi authoring refiner prepares typed, editable v4 proposals. + +Unchanged hashes skip refinement. All source acquisitions/refinements must succeed +before saving a new set of comparisons. Missing sources are recorded as unavailable, +never interpreted as permission to delete. No runtime lookup, ordinary consolidation +or preprocessing triggers remote refresh once the local archive is initialized. + +The administrator sees current and proposed units, scope, content, excerpts, review +items and explicit retirement IDs. **Keep local Evidence** records the retained wording +as a manual declaration with original lineage. **Use proposed Evidence** adopts the +proposal and its source version. Both save and activate through the existing archive +and corpus pipeline; review items block adoption. Comparisons use optimistic checks +on affected file bytes. Interrupted decisions have a durable replay journal and retry +without reacquisition, while intervening external edits are preserved and reported. + +Deleted IDs remain reserved. New model-generated identities from sources with curated +deletions are also conservatively suppressed; surviving IDs can still receive reviewed +updates. Deliberate new knowledge can be authored as a manual file. This mechanical +protection does not depend on the model detecting semantic duplication or contradictions. + +Installed commands are `workspace evidence refresh` and `workspace evidence decide`, +alongside E2 consolidation. Decision envelopes carry a source identity, comparison +revision and keep/replace choice. Extra URLs, arbitrary paths, forged actors and unknown +fields are rejected at the public API/CLI boundary. HTTP requests bind the authenticated +curator. Source operations do not mutate Catalog readiness or run DWH/schema stages. +The Python source worker is internal; the workflow CLI's visible surface is preserved. + +## Checks + +- Complete backend suite: **1,359 passed**, 40 skipped. Complete frontend suite: + **641 passed**. Both TypeScript checks passed; native Go CLI/workspace tests passed. +- Harness regression run: **1,256 passed**, one skipped and five deselected, with + portable-path tests run separately without `THT_HOME`. All **24 focused import, + CLI-surface and portable-path checks** passed. These + cover import, unchanged refresh, access failure, missing source, manual correction, + keep/replace, deletion suppression, stale comparisons, failure/retry and interrupted + journal writes. Ruff passed on the changed Python implementation and tests. +- The real-Qdrant test traverses the actual harness source CLI with deterministic + refinement/embedding boundaries: import is absent from recall before a decision, + accepted content becomes searchable, refreshed proposals preserve active manual + corrections, replacement removes the former text, deletion remains absent after + another refresh, and unrelated Schema/Memory canaries survive. +- Source contract fixtures cover controlled HTTP and S3 identities, exact acquired + bytes, and acquisition call counts. Existing adapter tests retain transport/egress + coverage. The test does not claim to exercise a live S3 account. +- React interaction tests verify explicit refresh, comparison content, exact decisions, + saved-decision retry and failure feedback. Route tests cover admin authorization, + workspace isolation, strict inputs and principal attribution. Service tests verify + the trusted config file descriptor and absence of Catalog mutation. + +## Local preview + +Core/frontend were rebuilt for the existing `thothii-18998cca7b0a` stack. Its persistent +archive and data volumes are retained. The native `/usr/local/bin/tht` was updated; +the previous executable is at `/private/tmp/tht-before-e3`. + +The installed refresh command ran against `psd-clinical` successfully: **35 unchanged +sources, zero changed, zero pending comparisons**. All 35 source hashes matched their +existing units, so this probe required no refinement and changed no active Evidence. +Source registry metadata was saved locally; no Git commit or push was performed. + +A separate synthetic draft was passed to the configured Pi/model inside core. It +produced one domain proposal with one review item, which was not activated. The probe +exposed a deployment issue: Python wheel modules and Pi skills live in different +directories. The refiner now resolves resources through `THT_HARNESS_DIR`, with the +source-tree location as its development fallback; a regression test covers this layout. + +The corrected installed worker was then exercised end to end in a temporary workspace +inside core, using the real configured Pi/model and a synthetic `incoming/orders.md`. +It returned success, one changed source, one comparison and one proposal with a review +item. The active snapshot remained absent. The temporary directory was removed on exit; +the probe did not open the DWH or activate an index. Strict docs build and +`git diff --check` also passed. + +The in-app browser still showed the login page with the prior credential error. +Authenticated visual verification remains manual; no credentials were reset or retried. + +## Operational instructions and boundaries + +See [Import drafts and refresh sources](../contracts/curated-evidence-v4.md#import-drafts-and-refresh-sources) +for commands, local paths, review, retry and backup/Git requirements. Preserve the +complete Evidence tree, including source comparisons, journals and acquired versions. +Local activation and transfer to the remote Git repository remain separate operator steps. + +No web content editor, automatic Git, background watcher, new job queue, general +retrieval benchmark or automatic source merge was introduced. Review/refinement is +sequential and bounded by per-source limits plus 200 documents/100 MiB per refresh. +New changed-source decisions and failure scenarios use isolated test data; live PSD +curated content was kept unchanged. X1 is not included in this increment. diff --git a/docs/plans/2026-09-09-memory-m2-validation.md b/docs/plans/2026-09-09-memory-m2-validation.md new file mode 100644 index 00000000..2dd8afeb --- /dev/null +++ b/docs/plans/2026-09-09-memory-m2-validation.md @@ -0,0 +1,96 @@ +# M2 — Ricerca ibrida e collegamenti + +Data: 2026-09-09. Implementazione locale dell'incremento M2 del +[piano Memory approvato](2026-09-08-memory-management.md#piano-esecutivo). + +## Risultato + +La ricerca Memory ed exemplar usa embedding dense e BM25 in Qdrant. I filtri sono +applicati in entrambi i rami prima della selezione dei candidati. Il core espande +i collegamenti uscenti, deduplica, limita cicli e visite, riordina insieme i risultati +e restituisce il contenuto corrente verificato in PostgreSQL. Non usa un database +a grafo né una chiamata LLM per il riordinamento. + +Il contesto fisico distingue database, schema, tabella e colonna e richiede che +corrispondano alla stessa dipendenza strutturata. Le card senza dipendenze valgono +per il workspace; un riferimento a un antenato fisico si applica ai suoi discendenti. +Ambito descrittivo e concetti possono restringere ulteriormente la ricerca tramite +`--filters`. La CLI impedisce di sostituire database/schema del runtime. Il dettaglio +dei limiti e della formula di ranking è nel [contratto operativo](../gestione-memory.md#hybrid-recall-and-links). + +L'incremento conserva i confini dei gate correnti: F2 riceve chiarimenti di dominio, +gli exemplar restano consultativi. Un collegamento non autorizza a consumare una +famiglia diversa o una card fuori ambito. Il riepilogo finale, i nuovi gate e la +pulizia delle dipendenze dopo sincronizzazione fisica appartengono a M3. + +## Transizione e recupero + +La migrazione versionata `002_hybrid_projection.sql` aggiunge il formato delle +proiezioni. Le card M1 restano autorevoli e consultabili; le loro proiezioni dense +risultano pendenti e non possono alimentare il recall. Un retry esplicito o +`tht memory index -c ` costruisce dense e BM25 dal contenuto corrente. +Il formato della proiezione e la revisione della card sono verificati prima dell'uso. + +Il salvataggio può aggiungere il vettore sparse mancante nella sola collezione +Memory. Non sostituisce configurazioni incompatibili e non modifica Reference. +Il rebuild esplicito ricrea anche una collezione Memory assente; il test elimina +la collezione, ricostruisce da PostgreSQL e verifica il ritorno della sola card conservata. +Fallimenti lasciano il lavoro di propagazione persistito e recuperabile. Nessuna +importazione da JSONL, sessioni storiche o vecchi payload Qdrant. + +## Verifiche + +| Controllo | Esito | +| --- | --- | +| Suite harness senza L0/L2, escluso il file dei percorsi portabili | 1.152 test passati nell'esecuzione finale. | +| Suite mirata Memory, adapter e CLI, con embedding reale | 78 test passati. | +| Verifica aggiuntiva del rebuild con collezione assente e adapter | 54 passati; il solo test del modello reale era escluso in questa riesecuzione. | +| API Fastify Memory | 13 test passati, compresa propagazione della lingua del workspace. | +| Browser amministrativo integrato | Passato: creazione, modifica, riavvio, rilettura, cancellazione e verifica del recall. | +| Wheel e casi CLI | 10 test passati; il wheel include entrambe le migrazioni Memory. | +| Build e controlli statici | Build/typecheck backend, Ruff sui file Python interessati e build documentale strict passati. | + +- Test deterministici: collegamenti necessari, contenuto corrente, duplicati, + cicli, profondità e limiti, card mancanti o pendenti, rifiuti già registrati, + famiglie e ambiti esclusi, dipendenze omonime e filtri CLI vincolati al runtime. +- Adapter: stesso filtro nei prefetch dense/BM25 per Memory ed exemplar; restano + coperti i contratti Evidence esistenti. +- PostgreSQL/Qdrant: salvataggio, modifica, cambio famiglia, cancellazione, + ricostruzione, riferimento separato, isolamento, RLS, errori, retry e transizione + delle proiezioni M1 al formato ibrido. +- Recupero effettivo: client Ollama di produzione con il modello configurato + `qwen3-embedding:0.6b`, dimensione 1024, e Qdrant dell'immagine fissata in Compose. + La domanda sulla chiave commessa fra esercizi recupera la regola attesa e la + granularità collegata, escludendo un altro database, un altro ambito e dipendenze + che corrispondono soltanto combinando riferimenti distinti. Verifica separatamente + dense, BM25 e fusione, poi recall, cancellazione e rebuild. + +Il test effettivo avvia un processo Ollama separato, montando il volume del modello +installato in sola lettura. PostgreSQL e Qdrant sono container temporanei dedicati, +eliminati a fine test. Non usa ID di risultati predisposti. Il browser amministrativo +usa invece embedding deterministici: verifica il collegamento fra UI, autenticazione, +Fastify, ThtRunner e persistenza, senza essere una misura di qualità del recupero. + +Non è una valutazione generale della qualità semantica su un corpus di produzione; +verifica i casi di recupero richiesti da M2, con il percorso reale configurato. + +## Riproduzione + +Dalla directory `harness`, con Docker disponibile: + +```sh +THT_HOME=/private/tmp/thothii-m2-test-home \ +THT_MEMORY_TEST_OLLAMA_VOLUME= \ +THT_MEMORY_TEST_MODEL=qwen3-embedding:0.6b \ +THT_MEMORY_TEST_DIMENSIONS=1024 \ +.venv/bin/pytest -q tests/memory/test_administration.py \ + tests/memory/test_retrieval.py tests/memory/test_recall.py \ + tests/test_qdrant_vector_store.py tests/test_solved_search_cli.py +``` + +Senza il volume esplicito, il solo test con modello reale viene escluso; gli altri +test restano eseguibili. Il volume deve contenere il modello indicato. Le immagini +Qdrant e Ollama del test sono lette da `compose.yaml`. + +Consegna locale: non sono stati eseguiti deploy, migrazioni delle installazioni +attive o aggiornamenti remoti dell'issue tracker. diff --git a/docs/plans/2026-09-09-memory-m3-validation.md b/docs/plans/2026-09-09-memory-m3-validation.md new file mode 100644 index 00000000..1d847c7b --- /dev/null +++ b/docs/plans/2026-09-09-memory-m3-validation.md @@ -0,0 +1,78 @@ +# Memory M3 — validation + +Implemented on 2026-09-09 in the rapid-harbor worktree. + +## Delivered behavior + +- F8 presents an editable summary of proposed additions, explicit updates and links. + Only selected content is saved, including the optional solved-question exemplar. +- Proposals reference effective approved decisions. Exact existing content is reused. + Concurrent edits invalidate an update; manual identities and origins are preserved. +- Selected cards and links commit atomically. Durable receipts recover repeat delivery + and the gap before the session review marker. Finalization does not add Memory. +- F4/F6/F7 retrieve SQL rules and explained errors for the existing approval gates. + Retrieval is consultative and does not write an approval decision. +- Successful Catalog physical sync deletes cards with matching removed dependencies. + The same Catalog transaction marks pending Memory cleanup. Retry keeps the original + removals and does not rescan; deletion receipts and projection tombstones survive restarts. +- Migration 003 adds minimal review and physical-cleanup receipts. + +## Executed checks + +| Check | Result | +| --- | --- | +| Harness deterministic suite, excluding portable-path environment cases | 1,152 passed | +| Portable-path suite without the temporary THT_HOME override | 9 passed | +| Memory service/retrieval with isolated PostgreSQL and Qdrant | 41 passed, 1 optional real-embedding case skipped, 1 L2 case excluded | +| Pi gate suite | 195 passed | +| Backend full suite | 1,346 passed; one auth timing test exceeded 5 seconds under concurrent load | +| Isolated auth and Catalog route rerun | All 40 passed, including the timed-out case | +| Catalog PostgreSQL integration after adding atomic cleanup-marker coverage | All 5 passed | +| Frontend full suite | 635 passed | +| Chromium summary review, desktop and 390px mobile | Passed; no page errors or horizontal overflow | +| Configured real GLM 5.3 generation | Passed on synthetic PostgreSQL data | +| Backend/frontend production builds, modified Python lint, strict docs build | Passed | + +The existing local Docker preview was rebuilt from this worktree, migration 003 +was applied, and core/frontend were recreated with the existing persistent volumes. +The preview remains at `http://127.0.0.1:8080`. + +The browser check uses the production widget in an isolated Vite fixture. It edits +the rule, declines the exemplar, submits only the selected card and checks responsive +layout. Gate tests separately verify request ordering through the production Pi +composition root; service and Catalog tests use real PostgreSQL. This is not a claim +of an automated complete live Pi conversation. + +The L2 case retrieves an approved SQL rule, excludes a card bound to another database, +and asks the configured GLM 5.3 model to generate a query. Order IDs repeat between +financial years; the correct composite join returns 120 on the synthetic fixture. +The generated SQL is validated and executed in a read-only PostgreSQL transaction. +No real DWH rows are sent. Embeddings in this case are deterministic; the real +embedding/hybrid retrieval evidence remains documented in M2. + +## Reproduction + +From the harness: + +```sh +THT_HOME=/private/tmp/thothii-m1-harness-home .venv/bin/pytest -q -m 'not l0 and not l2' --ignore=tests/test_portable_paths.py +.venv/bin/pytest -q tests/test_portable_paths.py +THT_HOME=/private/tmp/thothii-m1-harness-home .venv/bin/pytest -q tests/memory/test_administration.py tests/memory/test_retrieval.py -m 'not l2' +npm test +``` + +The optional generation case requires an installation YAML path and its running +core container. It resolves the model credential inside core without printing it: + +```sh +THT_MEMORY_L2_INSTALLATION= THT_MEMORY_L2_CORE= \ + .venv/bin/pytest -q -s -m l2 tests/memory/test_administration.py -k real_model +``` + +From frontend: `npx playwright test e2e/memory-review.spec.ts`. +Screenshots are written to `/private/tmp/thothii-m3-summary-desktop.png` and +`/private/tmp/thothii-m3-summary-mobile.png`. + +Evidence authoring and the joint X1 persistent Memory/Evidence conflict repair remain +outside M3. This increment does not infer knowledge from unexplained failures or +promise general improvements in SQL-generation accuracy. diff --git a/frontend/e2e/archive-repair.spec.ts b/frontend/e2e/archive-repair.spec.ts new file mode 100644 index 00000000..508da127 --- /dev/null +++ b/frontend/e2e/archive-repair.spec.ts @@ -0,0 +1,29 @@ +import { test, expect } from "@playwright/test"; +import { createServer, type ViteDevServer } from "vite"; + +let server: ViteDevServer; +let url: string; +test.beforeAll(async () => { + server = await createServer({ server: { host: "127.0.0.1", port: 0 } }); + await server.listen(); + url = server.resolvedUrls!.local[0]! + "e2e/fixtures/archive-repair.html"; +}); +test.afterAll(async () => { await server?.close(); }); +test("reviewer selects a correction, sees pending activation and retries on desktop/mobile", async ({ page }) => { + const errors: string[] = []; + page.on("pageerror", e => errors.push(e.message)); + await page.setViewportSize({ width: 1280, height: 900 }); + await page.goto(url); + await expect(page.getByRole("region", { name: "Archive conflict repair" })).toBeVisible(); + await page.screenshot({ path: "/private/tmp/thothii-x1-repair-desktop.png", fullPage: true }); + await page.setViewportSize({ width: 390, height: 844 }); + expect(await page.evaluate(() => document.documentElement.scrollWidth)).toBeLessThanOrEqual(390); + await page.screenshot({ path: "/private/tmp/thothii-x1-repair-mobile.png", fullPage: true }); + await page.getByRole("button", { name: "Apply this correction to Evidence" }).click(); + await expect(page.getByRole("status").filter({ hasText: "Correction saved" })).toContainText("activation is incomplete"); + await page.getByRole("button", { name: "Retry selected correction" }).click(); + await expect(page.getByRole("status").filter({ hasText: "Correction saved" })).toContainText("active in the archive index"); + const responses = JSON.parse(await page.locator("#response").textContent() || "[]"); + expect(responses.map((r: { choices: string[] }) => r.choices)).toEqual([["e"], ["e"]]); + expect(errors).toEqual([]); +}); diff --git a/frontend/e2e/fixtures/archive-repair.html b/frontend/e2e/fixtures/archive-repair.html new file mode 100644 index 00000000..2052fffc --- /dev/null +++ b/frontend/e2e/fixtures/archive-repair.html @@ -0,0 +1 @@ +Archive repair review
diff --git a/frontend/e2e/fixtures/archive-repair.tsx b/frontend/e2e/fixtures/archive-repair.tsx new file mode 100644 index 00000000..b9477ef9 --- /dev/null +++ b/frontend/e2e/fixtures/archive-repair.tsx @@ -0,0 +1,26 @@ +import React, { useState } from "react"; +import { createRoot } from "react-dom/client"; +import { ArchiveRepairWidget } from "../../src/widgets/ArchiveRepairWidget"; +import "../../src/index.css"; + +function Preview() { + const [step, setStep] = useState(0); + const [responses, setResponses] = useState([]); + return
0, indexed: step > 1, + status: ["proposed", "pending_activation", "active"][Math.min(step, 2)], options: [ + { id: "m", label: "Include the financial year in Memory", archive: "memory", target_id: "mem-order-key", + before: { subject: "Order identity", detail: "Use the order number.", scope: "Sales" }, + content: { subject: "Order identity", detail: "Use order number and financial year.", scope: "Sales" } }, + { id: "e", label: "Correct Evidence to use the order number", archive: "evidence", target_id: "evidence:order-key", + before: { title: "Order key", payload: { rule: "Use order number and financial year." } }, + content: { title: "Order key", payload: { rule: "Use the order number." }, + provenance: { kind: "manual", declared_by: "Sales curator" } } }, + ], + }, + }} onRespond={response => { setResponses([...responses, response]); setStep(step + 1); }} /> + {JSON.stringify(responses)}
; +} +createRoot(document.getElementById("root")!).render(); diff --git a/frontend/e2e/fixtures/auth-stack.mjs b/frontend/e2e/fixtures/auth-stack.mjs index c643286c..72deb76d 100644 --- a/frontend/e2e/fixtures/auth-stack.mjs +++ b/frontend/e2e/fixtures/auth-stack.mjs @@ -85,7 +85,7 @@ dwh: supported_transports: [postgres_direct] `; -async function prepareF1Workspace(root) { +async function prepareF1Workspace(root, memoryOnly = false) { const source = join(root, "workspace-source"); const remote = join(root, "workspace-remote.git"); secureDirectory(source); @@ -97,7 +97,8 @@ async function prepareF1Workspace(root) { " name: Fixture workspace", "", ].join("\n")); - writeSecure(join(source, F1_WORKSPACE_ID, "workspace.yaml"), F1_WORKSPACE_DESCRIPTOR); + writeSecure(join(source, F1_WORKSPACE_ID, "workspace.yaml"), + memoryOnly ? F1_WORKSPACE_DESCRIPTOR.split("dwh:")[0] : F1_WORKSPACE_DESCRIPTOR); await runFixtureCommand("git", ["init", "--bare", "--initial-branch=main", remote], root); chmodSync(remote, 0o700); @@ -332,7 +333,7 @@ function cleanBackendEnvironment(overrides) { return { ...env, ...overrides }; } -export async function createAuthenticationStack({ withF1Workspace = false } = {}) { +export async function createAuthenticationStack({ withF1Workspace = false, memoryRuntime = {} } = {}) { const credentials = resolveFixtureCredentials(); const root = mkdtempSync(join(realpathSync(tmpdir()), "thothii-auth-e2e-")); secureDirectory(root); @@ -363,7 +364,7 @@ export async function createAuthenticationStack({ withF1Workspace = false } = {} const [frontendPort, backendPort] = await Promise.all([freeLoopbackPort(), freeLoopbackPort()]); const publicUrl = `http://127.0.0.1:${frontendPort}`; const backendUrl = `http://127.0.0.1:${backendPort}`; - const workspace = withF1Workspace ? await prepareF1Workspace(root) : undefined; + const workspace = withF1Workspace ? await prepareF1Workspace(root, Object.keys(memoryRuntime).length > 0) : undefined; const provider = await startFakeOidcProvider({ directory: providerRoot, registration: { @@ -515,6 +516,7 @@ export async function createAuthenticationStack({ withF1Workspace = false } = {} THT_WS_FIXTURE_WORKSPACE_DWH_PASSWORD_FILE: fixtureDwhPasswordFile, THT_WS_FIXTURE_WORKSPACE_DWH_TLS_CA_FILE: fixtureDwhCaFile, } : {}), + ...memoryRuntime, }); } @@ -535,6 +537,10 @@ export async function createAuthenticationStack({ withF1Workspace = false } = {} return { publicUrl, + fixtureWorkspaceRoot() { + if (!workspace) throw safeError("e2e_workspace_fixture_missing"); + return join(registryRoot, "repo", workspace.id); + }, localAccount(account) { const found = accounts[account]; if (!found) throw safeError("e2e_local_account_unknown"); diff --git a/frontend/e2e/fixtures/fake-tht.mjs b/frontend/e2e/fixtures/fake-tht.mjs index 725cb79f..7179c6cc 100755 --- a/frontend/e2e/fixtures/fake-tht.mjs +++ b/frontend/e2e/fixtures/fake-tht.mjs @@ -5,6 +5,15 @@ // So we skip the first two args after node (-c ). const argv = process.argv.slice(2); // drop "node" + script path +// Archive integration tests run the actual harness through the production runner. +// Unrelated session/LLM operations remain deterministic fixtures. +if ((argv[0] === "memory" || argv[0] === "evidence" && argv[1] === "admin") && process.env.THT_E2E_MEMORY_BIN) { + const { spawnSync } = await import("node:child_process"); + const result = spawnSync(process.env.THT_E2E_MEMORY_BIN, argv, { + env: process.env, stdio: ["ignore", "inherit", "inherit", 3], + }); + process.exit(result.status ?? 1); +} // Strip leading "-c " pair (ThtRunner always prepends this). let args = argv; diff --git a/frontend/e2e/fixtures/memory-review.html b/frontend/e2e/fixtures/memory-review.html new file mode 100644 index 00000000..e4ebd70f --- /dev/null +++ b/frontend/e2e/fixtures/memory-review.html @@ -0,0 +1,3 @@ + + +
diff --git a/frontend/e2e/fixtures/memory-review.tsx b/frontend/e2e/fixtures/memory-review.tsx new file mode 100644 index 00000000..867b3080 --- /dev/null +++ b/frontend/e2e/fixtures/memory-review.tsx @@ -0,0 +1,22 @@ +import React from "react"; +import { createRoot } from "react-dom/client"; +import { MemoryReviewWidget } from "../../src/widgets/MemoryReviewWidget"; +import "../../src/index.css"; + +const card = { family: "sql_rule", subject: "Join orders within the same financial year", + detail: "Join order lines using both order_id and financial_year. The order number repeats across years.", + scope: "Sales reporting", rationale: "Confirmed during review of annual totals.", + question: "What is the total value of shipped orders?", sql: "", concepts: ["Order grain"], + dependencies: [{ database: "warehouse", schema_name: "sales", table: "orders", column: "financial_year" }], + links: [] }; +const descriptor = { id: "review", widget: "memory-review", reserved: [], + summary: { summary_id: "fixture", items: [ + { id: "rule", card, reason: "Preserve the approved join correction.", source_seqs: [12], before: null, target_id: null }, + { id: "solved", card: { ...card, family: "solved_question", subject: "Shipped orders by year", + sql: "SELECT financial_year, SUM(total) FROM sales.orders GROUP BY financial_year" }, + reason: "Keep the approved solution as an example.", source_seqs: [18], before: null, target_id: null }, + ] } }; +createRoot(document.getElementById("root")!).render(
+ { + document.querySelector("#response")!.textContent = JSON.stringify(response); + }} />
); diff --git a/frontend/e2e/memory-real.spec.ts b/frontend/e2e/memory-real.spec.ts new file mode 100644 index 00000000..29f02f7d --- /dev/null +++ b/frontend/e2e/memory-real.spec.ts @@ -0,0 +1,195 @@ +import { expect, test } from "@playwright/test"; +import { execFileSync } from "node:child_process"; +import { createServer, type Server } from "node:http"; +import { resolve } from "node:path"; +import { createAuthenticationStack } from "./fixtures/auth-stack.mjs"; + +// Opt in: real isolated PostgreSQL/Qdrant containers and the Python harness are required. +// Auth, browser, Fastify, ThtRunner and Memory persistence are production code. +// Embeddings are deterministic; unrelated sessions and Pi use the existing fixture. +test.skip(process.env.THT_MEMORY_BROWSER_E2E !== "1", "Set THT_MEMORY_BROWSER_E2E=1"); +test.setTimeout(180_000); +test.use({ actionTimeout: 15_000, viewport: { width: 1440, height: 1000 } }); +let stack: Awaited>; +let embedding: Server; +const containers: string[] = []; +const root = resolve(".."); +let databaseUrl: string; +let qdrantUrl: string; + +function command(bin: string, args: string[], env = process.env) { + return execFileSync(bin, args, { cwd: root, env, encoding: "utf8", timeout: 60_000 }).trim(); +} +function container(image: string, port: number, args: string[] = []) { + const id = command("docker", ["run", "-d", "--rm", "-p", `127.0.0.1::${port}`, ...args, image]); + containers.push(id); + return { id, port: Number(command("docker", ["port", id, `${port}/tcp`]).split(":").at(-1)) }; +} + +test.beforeAll(async () => { + const pg = container("postgres:16-alpine", 5432, ["-e", "POSTGRES_PASSWORD=memory-test"]); + const qdrant = container("qdrant/qdrant:v1.18.2", 6333); + qdrantUrl = `http://127.0.0.1:${qdrant.port}`; + await expect.poll(async () => { try { return (await fetch(`${qdrantUrl}/healthz`)).ok; } catch { return false; } }).toBe(true); + await expect.poll(() => { + // The image's temporary initialization server accepts sockets before restarting. + // TCP readiness identifies the final server and avoids racing that shutdown. + try { command("docker", ["exec", pg.id, "pg_isready", "-h", "127.0.0.1", "-U", "postgres"]); return true; } catch { return false; } + }).toBe(true); + command("docker", ["exec", pg.id, "psql", "-U", "postgres", "-c", + "CREATE ROLE thothii_catalog_runtime LOGIN PASSWORD 'memory-runtime-test'; " + + "ALTER DEFAULT PRIVILEGES IN SCHEMA public GRANT SELECT, INSERT, UPDATE, DELETE ON TABLES TO thothii_catalog_runtime; " + + "ALTER DEFAULT PRIVILEGES IN SCHEMA public GRANT USAGE, SELECT ON SEQUENCES TO thothii_catalog_runtime"]); + const migratorUrl = `postgresql://postgres:memory-test@127.0.0.1:${pg.port}/postgres`; + databaseUrl = `postgresql://thothii_catalog_runtime:memory-runtime-test@127.0.0.1:${pg.port}/postgres`; + command(resolve(root, "harness/.venv/bin/python"), ["-m", "tht.memory.migrate"], + { ...process.env, THT_CATALOG_MIGRATOR_DATABASE_URL: migratorUrl }); + command(resolve(root, "backend/node_modules/.bin/tsx"), ["backend/src/catalog/migrate.ts"], + { ...process.env, THT_CATALOG_MIGRATOR_DATABASE_URL: migratorUrl }); + embedding = createServer(async (request, response) => { + let body = ""; + for await (const chunk of request) body += chunk; + const payload = JSON.parse(body || "{}"); + const texts = Array.isArray(payload.input) ? payload.input : [payload.input]; + response.setHeader("content-type", "application/json"); + response.end(JSON.stringify({ embeddings: texts.map(() => [1, 0, 0]) })); + }); + await new Promise(done => embedding.listen(0, "127.0.0.1", done)); + const address = embedding.address(); + if (!address || typeof address === "string") throw new Error("embedding address missing"); + stack = await createAuthenticationStack({ withF1Workspace: true, memoryRuntime: { + THT_E2E_MEMORY_BIN: resolve(root, "harness/.venv/bin/tht"), + THT_CATALOG_RUNTIME_DATABASE_URL: databaseUrl, + THT_CATALOG_DATABASE_URL: databaseUrl, + THT_INTERNAL_QDRANT_URL: qdrantUrl, + THT_INTERNAL_EMBEDDING_URL: `http://127.0.0.1:${address.port}`, + THT_INTERNAL_EMBEDDING_DIMENSIONS: "3", + } }); + await stack.useLocalMode(); +}); + +test.afterAll(async () => { + await stack?.close(); + if (embedding) await new Promise(done => embedding.close(() => done())); + for (const id of containers.reverse()) command("docker", ["rm", "-f", id]); +}); + +test("admin manages both archives through the real stack without cross-deletion", async ({ page }) => { + page.on("pageerror", error => console.error("Memory browser error:", error.message)); + await page.goto(stack.publicUrl); + await page.getByLabel("Username").fill(stack.localAccount("admin").username); + await page.getByLabel("Password", { exact: true }).fill(stack.localAccount("admin").password); + await page.getByRole("button", { name: "Sign in", exact: true }).click(); + await expect(page.getByTestId("app-shell")).toBeVisible(); + // Pull only the temporary fixture registry through the authenticated application API. + const pulled = await page.evaluate(async () => { + const me = await (await fetch("/api/me")).json(); + const result = await fetch("/api/workspace-registry/pull", { method: "POST", + headers: { "x-thothii-csrf": me.csrfToken } }); + return { status: result.status, body: await result.json() }; + }); + expect(pulled.status, JSON.stringify(pulled.body)).toBe(200); + // This checkout belongs to the temporary authentication fixture, never the PSD archive. + command(resolve(root, "harness/.venv/bin/python"), ["-c", ` +import os +from pathlib import Path +from tht.evidence.canonical import CuratedEvidence, dump_curated_markdown +from tht.evidence.local_archive import LocalEvidenceArchive +root = Path(os.environ["FIXTURE_ARCHIVE_ROOT"]) +unit = CuratedEvidence.model_validate({"schema_version":4,"id":"evidence:order-key", + "kind":"domain","title":"Order key Evidence","language":"en","purposes":["sql_generation"], + "payload":{"rule":"Use order_id and financial_year when joining order lines."}, + "provenance":{"kind":"manual","declared_by":"Acceptance fixture"}}) +path = root / "evidence/curated/domain/order-key.md" +path.parent.mkdir(parents=True, exist_ok=True) +path.write_text(dump_curated_markdown(unit)) +LocalEvidenceArchive(root).initialize() +`], { ...process.env, FIXTURE_ARCHIVE_ROOT: stack.fixtureWorkspaceRoot() }); + async function openMemory() { + await page.getByRole("button", { name: "Administration", exact: true }).click(); + expect((await page.getByRole("button", { + name: /^(Database management|Memory management|Evidence management)$/, + }).allTextContents()).map(text => text.trim())).toEqual([ + "Database management", "Memory management", "Evidence management", + ]); + await page.getByRole("button", { name: "Memory management", exact: true }).click(); + await page.screenshot({ path: "/private/tmp/thothii-m1-memory-opening.png", fullPage: true }); + const memory = page.getByRole("main", { name: "Memory management" }); + await expect(memory).toBeVisible(); + await memory.getByLabel("Workspace", { exact: true }).selectOption("fixture-workspace"); + await expect(page.getByRole("button", { name: "New card", exact: true })).toBeVisible(); + } + await openMemory(); + await page.getByRole("button", { name: "New card", exact: true }).click(); + await page.getByLabel("Title", { exact: true }).fill("Orders approved"); + await page.getByLabel("Scope", { exact: true }).fill("Sales"); + await page.getByLabel("Content", { exact: true }).fill("Only shipped orders"); + await page.getByRole("button", { name: "Save card", exact: true }).click(); + await expect(page.getByText("Saved and indexed.", { exact: true })).toBeVisible(); + await page.getByRole("button", { name: "Edit card", exact: true }).click(); + await page.getByLabel("Content", { exact: true }).fill("Only fully shipped orders"); + await page.getByRole("button", { name: "Save card", exact: true }).click(); + await expect(page.getByText("Only fully shipped orders", { exact: true })).toBeVisible(); + await stack.restartBackend(); + await page.reload(); + await openMemory(); + await page.getByRole("button", { name: /Orders approved Domain clarification/ }).click(); + await expect(page.getByText("Only fully shipped orders", { exact: true })).toBeVisible(); + await page.screenshot({ path: "/private/tmp/thothii-m1-memory-desktop.png", fullPage: true }); + await page.getByRole("button", { name: "Delete card", exact: true }).click(); + await page.getByRole("button", { name: "Confirm deletion", exact: true }).click(); + await expect(page.getByText("Card deleted and removed from recall.", { exact: true })).toBeVisible(); + const recalled = command(resolve(root, "harness/.venv/bin/python"), ["-c", ` +import json, os +from tht.memory.runtime import admin_service +from tht.adapters.vector.qdrant import QdrantVectorStore +runtime = {"internalQdrantUrl": os.environ["QDRANT_URL"], "internalEmbeddingDimensions": 3} +service = admin_service("fixture-workspace", runtime) +store = QdrantVectorStore(base_url=runtime["internalQdrantUrl"], workspace_id="fixture-workspace", + collections={"reference":"fixture-workspace-reference", "memory":"fixture-workspace-memory"}, expected_dimension=3) +class Searcher: + def search(self, vector, top_n, kinds, **kwargs): return store.search(["memory"], vector, limit=top_n, kinds=kinds, **kwargs) +class Embedder: + def embed_query(self, text): return [1,0,0] +try: print(json.dumps(service.recall("orders", searcher=Searcher(), embedder=Embedder()))) +finally: service.close() +`], { ...process.env, THT_CATALOG_RUNTIME_DATABASE_URL: databaseUrl, QDRANT_URL: qdrantUrl, + THT_PRINCIPAL_ISSUER: "test", THT_PRINCIPAL_SUBJECT: "admin", THT_PRINCIPAL_IS_ADMIN: "true" }); + expect(JSON.parse(recalled)).toEqual([]); + await expect(page.getByRole("button", { name: "New card", exact: true })).toBeEnabled(); + if (!await page.getByRole("button", { name: "Evidence management", exact: true }).isVisible()) + await page.getByRole("button", { name: "Administration", exact: true }).click(); + await page.getByRole("button", { name: "Evidence management", exact: true }).click(); + const evidence = page.getByRole("main", { name: "Evidence management" }); + await expect(evidence).toBeVisible(); + await evidence.getByRole("combobox", { name: /Workspace/ }).selectOption("fixture-workspace"); + await evidence.getByRole("button", { name: "Order key Evidence", exact: true }).click(); + await expect(evidence.getByRole("region", { name: "Evidence detail" })).toContainText( + "Use order_id and financial_year when joining order lines."); + await evidence.getByLabel("Search Evidence", { exact: true }).fill("no matching knowledge"); + await expect(evidence.getByText(/No Evidence matches these filters/)).toBeVisible(); + await evidence.getByLabel("Search Evidence", { exact: true }).fill("financial_year"); + await expect(evidence.getByRole("button", { name: "Order key Evidence", exact: true })).toBeVisible(); + await page.screenshot({ path: "/private/tmp/thothii-acceptance-evidence-desktop.png", fullPage: true }); + await page.setViewportSize({ width: 390, height: 844 }); + await expect(page.getByRole("button", { name: "Navigation", exact: true })).toBeVisible(); + await expect(page.getByRole("complementary", { name: "Session navigation" })).toBeHidden(); + expect((await evidence.boundingBox())!.width).toBeGreaterThanOrEqual(380); + expect(await page.evaluate(() => document.documentElement.scrollWidth)).toBeLessThanOrEqual(390); + await page.screenshot({ path: "/private/tmp/thothii-acceptance-evidence-mobile.png", fullPage: true }); + await page.getByRole("button", { name: "Navigation", exact: true }).click(); + const navigation = page.getByRole("dialog", { name: "Navigation", exact: true }); + await expect(navigation.getByRole("button", { name: "Memory management", exact: true })).toBeVisible(); + await page.keyboard.press("Escape"); + await expect(navigation).toBeHidden(); + await expect(page.getByRole("button", { name: "Navigation", exact: true })).toBeFocused(); + await page.getByRole("button", { name: "Navigation", exact: true }).click(); + await navigation.getByRole("button", { name: "Memory management", exact: true }).click(); + await expect(navigation).toBeHidden(); + const mobileMemory = page.getByRole("main", { name: "Memory management" }); + await expect(mobileMemory).toBeVisible(); + expect((await mobileMemory.boundingBox())!.width).toBeGreaterThanOrEqual(380); + const unknown = await page.evaluate(async () => Promise.all(["memory", "evidence"].map(async archive => + (await fetch(`/api/workspaces/unregistered-workspace/${archive}`)).status))); + expect(unknown).toEqual([404, 404]); +}); diff --git a/frontend/e2e/memory-review.spec.ts b/frontend/e2e/memory-review.spec.ts new file mode 100644 index 00000000..6191e4c3 --- /dev/null +++ b/frontend/e2e/memory-review.spec.ts @@ -0,0 +1,29 @@ +import { test, expect } from "@playwright/test"; +import { createServer, type ViteDevServer } from "vite"; + +let server: ViteDevServer; +let url: string; +test.beforeAll(async () => { + server = await createServer({ server: { host: "127.0.0.1", port: 0 } }); + await server.listen(); + url = server.resolvedUrls!.local[0]! + "e2e/fixtures/memory-review.html"; +}); +test.afterAll(async () => { await server?.close(); }); +test("reviewer edits the final summary and selects what to save on desktop and mobile", async ({ page }) => { + const errors: string[] = []; + page.on("pageerror", e => errors.push(e.message)); + await page.setViewportSize({ width: 1280, height: 900 }); + await page.goto(url); + await expect(page.getByRole("form", { name: "Memory summary" })).toBeVisible(); + await page.getByLabel("rule detail").fill("Join on order number and year, within the sales database."); + await page.getByRole("checkbox", { name: /Shipped orders by year/ }).uncheck(); + await page.screenshot({ path: "/private/tmp/thothii-m3-summary-desktop.png", fullPage: true }); + await page.setViewportSize({ width: 390, height: 844 }); + expect(await page.evaluate(() => document.documentElement.scrollWidth)).toBeLessThanOrEqual(390); + await page.screenshot({ path: "/private/tmp/thothii-m3-summary-mobile.png", fullPage: true }); + await page.getByRole("button", { name: "Save selected and finish" }).click(); + const response = JSON.parse(await page.locator("#response").textContent() || "{}"); + expect(JSON.parse(response.text).items).toHaveLength(1); + expect(JSON.parse(response.text).items[0].card.detail).toContain("within the sales database"); + expect(errors).toEqual([]); +}); diff --git a/frontend/src/api/catalog-databases.ts b/frontend/src/api/catalog-databases.ts index 167c1568..305d6ea7 100644 --- a/frontend/src/api/catalog-databases.ts +++ b/frontend/src/api/catalog-databases.ts @@ -258,7 +258,7 @@ export type CatalogSyncScope = "tables" | "columns" | "relationships" | "all"; export type CatalogSyncState = "queued" | "running" | "awaiting_confirmation" | "applying" | "succeeded" | "failed" | "cancelled" | "interrupted"; export type CatalogSyncPhase = "queued" | "connecting" | "scanning_tables" | "scanning_columns" - | "scanning_relationships" | "planning" | "awaiting_confirmation" | "applying" | "completed"; + | "scanning_relationships" | "planning" | "awaiting_confirmation" | "applying" | "memory_cleanup" | "completed"; export interface CatalogSchemaDiff { deletedTables: string[]; @@ -276,7 +276,7 @@ export interface CatalogSyncRun { requestedDatabaseVersion: number; plannedDiff: CatalogSchemaDiff | null; confirmationToken: string | null; - counts: { tables?: number; columns?: number; relationships?: number; created?: number; updated?: number; deleted?: number }; + counts: { tables?: number; columns?: number; relationships?: number; created?: number; updated?: number; deleted?: number; memoryDeleted?: number }; errorCode: string | null; errorMessage: string | null; cancelRequested: boolean; diff --git a/frontend/src/api/client.ts b/frontend/src/api/client.ts index d0987145..b54ee0f6 100644 --- a/frontend/src/api/client.ts +++ b/frontend/src/api/client.ts @@ -7,6 +7,8 @@ import { const MAX_ERROR_BODY_BYTES = 8 * 1024; const safeErrorCodes = new Set([ + "evidence_unavailable", "evidence_not_found", + "memory_invalid", "memory_forbidden", "memory_not_found", "memory_conflict", "memory_unavailable", "auth_forbidden", "auth_invalid_credentials", "auth_not_authorized", "auth_unavailable", "auth_not_implemented", "authentication_required", "invalid_credentials", "login_rate_limited", "csrf_failed", "csrf_invalid", "dwh_unreachable", "model_unavailable", @@ -46,6 +48,13 @@ type SafeErrorPayload = { }; const localCodeMessages: Record = { + evidence_unavailable: "Evidence archive is unavailable. Check the local files and installation configuration.", + evidence_not_found: "Evidence was not found. Refresh the file list.", + memory_invalid: "Check the card fields, family requirements and schema dependencies.", + memory_forbidden: "Memory administration is not permitted.", + memory_not_found: "This Memory card or pending update no longer exists in the workspace.", + memory_conflict: "A linked card is missing or invalid. Review the links and retry.", + memory_unavailable: "The Memory archive is unavailable. Check the installation connection and migrations, then retry.", auth_forbidden: "Access is not permitted.", auth_invalid_credentials: "Invalid username or password.", auth_not_authorized: "Access is not permitted.", diff --git a/frontend/src/api/evidence.ts b/frontend/src/api/evidence.ts new file mode 100644 index 00000000..8dca645f --- /dev/null +++ b/frontend/src/api/evidence.ts @@ -0,0 +1,42 @@ +import { apiFetch } from "./client"; + +export interface EvidenceUnit { + id: string; title: string; kind: string; language: string; purposes: string[]; + applies_to: { concepts: string[]; tables: string[]; columns: string[] }; + payload: Record>; + provenance: { kind?: "manual"; declared_by?: string; source_file?: string; supporting_excerpts?: string[]; + original?: { source_file: string; supporting_excerpts: string[] } | null }; + review_items: Array<{ code: string; message: string; field?: string }>; + file: string; status: string; revision: string; +} +export interface EvidencePage { + source_reviews?: EvidenceSourceReview[]; + items: EvidenceUnit[]; item?: EvidenceUnit; total: number; page: number; page_size: number; + errors: Array<{ file: string; message: string }>; initialized: boolean; + active_revision: string | null; pending_revision: string | null; + location: { repository: string | null; workspace: string | null; runtime_workspace: string; + host: string; command: string; git_commands: string[] }; +} + +export interface EvidenceSourceReview { + id: string; uri: string; revision: string; availability: "available" | "missing"; + status: "review" | "applying" | "accepted" | "kept"; decision?: "keep" | "replace"; + current?: EvidenceUnit[]; proposed: EvidenceUnit[]; removed_ids?: string[]; suppressed?: boolean; +} +export type EvidenceSourceAction = { action: "refresh" } | { + action: "decide"; sourceId: string; revision: string; decision: "keep" | "replace"; +}; +export async function actOnEvidenceSources(workspace: string, action: EvidenceSourceAction) { + const result = await apiFetch<{status: string; warnings?: string[]}>( + `/workspaces/${encodeURIComponent(workspace)}/evidence/sources`, {method: "POST", body: JSON.stringify(action)}); + if (result.status !== "succeeded") throw new Error(result.warnings?.join(" ") || "Source operation failed."); + return result; +} +export function listEvidence(workspace: string, filters: Record = {}) { + const params = new URLSearchParams(); + for (const [key, value] of Object.entries(filters)) if (value !== "") params.set(key, String(value)); + return apiFetch(`/workspaces/${encodeURIComponent(workspace)}/evidence?${params}`); +} +export function getEvidence(workspace: string, id: string) { + return apiFetch(`/workspaces/${encodeURIComponent(workspace)}/evidence/${encodeURIComponent(id)}`); +} diff --git a/frontend/src/api/memory.ts b/frontend/src/api/memory.ts new file mode 100644 index 00000000..92dea0a2 --- /dev/null +++ b/frontend/src/api/memory.ts @@ -0,0 +1,41 @@ +import { apiFetch } from "./client"; + +export type MemoryFamily = "domain_clarification" | "sql_rule" | "solved_question" | "explained_error"; +export interface MemoryDependency { database: string; schema_name: string; table: string; column: string } +export interface MemoryLink { target_id: string; meaning: string } +export interface MemoryInput { + family: MemoryFamily; subject: string; detail: string; scope: string; rationale: string; + question: string; sql: string; concepts: string[]; dependencies: MemoryDependency[]; links: MemoryLink[]; +} +export interface MemoryCard extends MemoryInput { + id: string; workspace_id: string; origin: "manual" | "workflow"; session_id: string | null; + decision_seq: number | null; created_at: string; updated_at: string; revision: string; indexed: boolean; +} +export interface MemoryPage { items: MemoryCard[]; total: number; page: number; page_size: number } +export interface MemoryPending { card_id: string; action: "upsert" | "delete"; error: string | null } +export interface MemoryOutcome { + id: string; saved: boolean; indexed: boolean; action: "upsert" | "delete"; + card?: MemoryCard; error: string | null; +} +const base = (workspace: string) => `/workspaces/${encodeURIComponent(workspace)}/memory`; +export function listMemory(workspace: string, filters: Record = {}) { + const params = new URLSearchParams(); + for (const [key, value] of Object.entries(filters)) if (value !== "") params.set(key, String(value)); + return apiFetch(`${base(workspace)}?${params}`); +} +export const getMemory = (workspace: string, id: string) => + apiFetch(`${base(workspace)}/${encodeURIComponent(id)}`); +export const pendingMemory = (workspace: string) => apiFetch(`${base(workspace)}/pending`); +export const saveMemory = (workspace: string, card: MemoryInput, id?: string) => + apiFetch(`${base(workspace)}${id ? `/${encodeURIComponent(id)}` : ""}`, { + method: id ? "PUT" : "POST", body: JSON.stringify(card), + }); +export const deleteMemory = (workspace: string, id: string) => + apiFetch(`${base(workspace)}/${encodeURIComponent(id)}`, { method: "DELETE" }); +export const retryMemory = (workspace: string, id: string) => + apiFetch(`${base(workspace)}/${encodeURIComponent(id)}/retry`, { method: "POST" }); + +export function memoryInput(card: MemoryCard): MemoryInput { + const { family, subject, detail, scope, rationale, question, sql, concepts, dependencies, links } = card; + return { family, subject, detail, scope, rationale, question, sql, concepts, dependencies, links }; +} diff --git a/frontend/src/shell/AppShell.auth.test.tsx b/frontend/src/shell/AppShell.auth.test.tsx index ccd10db0..89d47154 100644 --- a/frontend/src/shell/AppShell.auth.test.tsx +++ b/frontend/src/shell/AppShell.auth.test.tsx @@ -47,6 +47,15 @@ beforeEach(() => { }); describe("authenticated shell permissions", () => { + test("opens Memory administration independently of the core session", async () => { + renderShell({ subject: "admin", isAdmin: true, roles: ["admin"], + permissions: ["session.use", "memory.manage", "database.manage"] }); + await userEvent.click(await screen.findByRole("button", { name: "Administration" })); + await userEvent.click(screen.getByRole("button", { name: "Memory management" })); + const memory = screen.getByRole("main", { name: "Memory management" }); + expect(memory).toBeVisible(); + expect(within(memory).getByRole("combobox", { name: "Workspace" })).toBeVisible(); + }); test("does not use the legacy installation default as a preprocessing workspace", async () => { const requestedWorkspaceIds: string[] = []; server.use( @@ -125,15 +134,19 @@ describe("authenticated shell permissions", () => { const panel = within(sessionNavigation).getByRole("region", { name: "Administration" }); const database = within(panel).getByRole("button", { name: "Database management" }); + const memory = within(panel).getByRole("button", { name: "Memory management" }); + const evidence = within(panel).getByRole("button", { name: "Evidence management" }); const separator = within(panel).getByRole("separator"); const workspace = within(panel).getByRole("button", { name: "Workspace management" }); const pi = within(panel).getByRole("button", { name: "Pi management" }); const preprocessing = within(panel).getByRole("region", { name: "Workspace preprocessing" }); expect(panel.children[0]).toBe(database); - expect(panel.children[1]).toBe(separator); - expect(panel.children[2]).toBe(workspace); - expect(panel.children[3]).toBe(pi); - expect(panel.children[4]).toBe(preprocessing); + expect(panel.children[1]).toBe(memory); + expect(panel.children[2]).toBe(evidence); + expect(panel.children[3]).toBe(separator); + expect(panel.children[4]).toBe(workspace); + expect(panel.children[5]).toBe(pi); + expect(panel.children[6]).toBe(preprocessing); expect(screen.getByRole("tab", { name: "All sessions" })).toBeInTheDocument(); await user.keyboard(" "); diff --git a/frontend/src/shell/AppShell.tsx b/frontend/src/shell/AppShell.tsx index 043a10ad..a21589b5 100644 --- a/frontend/src/shell/AppShell.tsx +++ b/frontend/src/shell/AppShell.tsx @@ -17,6 +17,9 @@ import { StopConfirmDialog } from "./StopConfirmDialog"; import { SteerInput, ComposerFooter } from "./SteerInput"; import { WorkflowBar } from "./WorkflowBar"; import { DatabaseManagementPage } from "./DatabaseManagementPage"; +import { MemoryManagementPage } from "./MemoryManagementPage"; +import { EvidenceManagementPage } from "./EvidenceManagementPage"; +import { ArchiveNavigation } from "./ArchiveNavigation"; import { WorkspacePreprocessingControl } from "./WorkspacePreprocessingControl"; import { Pencil, ArrowLeft, ArrowRight, ChevronDown, Trash2 } from "lucide-react"; import { Accordion } from "@base-ui/react/accordion"; @@ -46,7 +49,7 @@ interface AppShellProps { canLogout: boolean; } -type ActiveSurface = "core" | "database-management"; +type ActiveSurface = "core" | "database-management" | "memory-management" | "evidence-management"; type ActiveManagementPanel = "workspace" | "pi" | null; const SESSION_SCOPES: readonly SessionScope[] = ["mine", "all"]; @@ -153,6 +156,10 @@ export function AppShell({ canLogout }: AppShellProps) { const canManageWorkspace = permissions.includes("workspace.manage"); const canManageWorkspaceSecrets = permissions.includes("workspace.secrets.manage"); const canManageDatabase = permissions.includes("database.manage"); + const canManageMemory = permissions.includes("memory.manage"); + const canManageEvidence = permissions.includes("evidence.manage"); + const [memoryDirty, setMemoryDirty] = useState(false); + const [memoryBusy, setMemoryBusy] = useState(false); const canManagePi = permissions.includes("pi.manage"); const authGeneration = useAuthGeneration(); const { data: installationSettings } = useQuery({ @@ -299,6 +306,13 @@ export function AppShell({ canLogout }: AppShellProps) { } function canLeaveDatabaseManagement(): boolean { + if (activeSurface === "memory-management" && memoryBusy) { + toast.info("Wait for the Memory operation to finish before leaving this page"); + return false; + } + if (activeSurface === "memory-management" && memoryDirty) { + return window.confirm("Discard unsaved Memory changes and leave Memory management?"); + } if (activeSurface !== "database-management") return true; if (databaseNavigationRef.current.busy) { toast.info("Wait for the database operation to finish before leaving this page"); @@ -695,7 +709,7 @@ export function AppShell({ canLogout }: AppShellProps) { const currentNavigation = activeSurface === "database-management" ? "database" - : activeManagementPanel ?? "core"; + : activeSurface === "memory-management" ? "memory" : activeSurface === "evidence-management" ? "evidence" : activeManagementPanel ?? "core"; const adminNavigationOpen = adminNavigationValue.includes("administration"); const managementNavigationCurrent = currentNavigation !== "core"; @@ -709,6 +723,7 @@ export function AppShell({ canLogout }: AppShellProps) { data-session-resizing={sessionResizing} className={[ "flex h-screen bg-background text-foreground", + ["memory-management", "evidence-management"].includes(activeSurface) ? "max-md:pt-12" : "", activityResizing || sessionResizing ? "select-none cursor-col-resize" : "", ].join(" ")} > @@ -749,6 +764,10 @@ export function AppShell({ canLogout }: AppShellProps) { {/* Conversation column */} + {activeSurface === "memory-management" && ( + + )} + {activeSurface === "evidence-management" && } {activeSurface === "database-management" && ( {/* Right session rail */} - {(activeSurface === "database-management" || !showActivity) && ( + {(activeSurface !== "core" || !showActivity) && ( + + )} diff --git a/frontend/src/shell/ArchiveNavigation.tsx b/frontend/src/shell/ArchiveNavigation.tsx new file mode 100644 index 00000000..9010218d --- /dev/null +++ b/frontend/src/shell/ArchiveNavigation.tsx @@ -0,0 +1,30 @@ +import { useEffect, useState, useSyncExternalStore, type ReactNode } from "react"; +import { Button } from "../components/ui/button"; +import { Dialog, DialogContent, DialogDescription, DialogTitle, DialogTrigger } from "../components/ui/dialog"; + +const query = "(max-width: 767px)"; +function subscribe(listener: () => void) { + const media = window.matchMedia(query); + media.addEventListener("change", listener); + return () => media.removeEventListener("change", listener); +} +const isNarrow = () => window.matchMedia(query).matches; + +/** Keep the archive content readable on phones, using the shared accessible dialog. */ +export function ArchiveNavigation({ surface, children }: { surface: string; children: ReactNode }) { + const narrow = useSyncExternalStore(subscribe, isNarrow, () => false); + const [open, setOpen] = useState(false); + useEffect(() => { setOpen(false); }, [surface, narrow]); + if (!narrow || !["memory-management", "evidence-management"].includes(surface)) return children; + return +
+ ThothII + }>Navigation +
+ + Navigation + Choose an administration page or a session. +
{children}
+
+
; +} diff --git a/frontend/src/shell/EvidenceManagementPage.test.tsx b/frontend/src/shell/EvidenceManagementPage.test.tsx new file mode 100644 index 00000000..73b4057a --- /dev/null +++ b/frontend/src/shell/EvidenceManagementPage.test.tsx @@ -0,0 +1,86 @@ +import { QueryClient, QueryClientProvider } from "@tanstack/react-query"; +import { render, screen, within } from "@testing-library/react"; +import userEvent from "@testing-library/user-event"; +import { http, HttpResponse } from "msw"; +import { server } from "../test/msw"; +import { EvidenceManagementPage } from "./EvidenceManagementPage"; +import type { EvidencePage, EvidenceUnit } from "../api/evidence"; + +const item: EvidenceUnit = { id: "evidence:order-key", title: "Order key", kind: "domain", language: "en", purposes: ["sql_generation"], + applies_to: { concepts: ["Orders"], tables: ["sales.orders"], columns: [] }, payload: { rule: "Join by order number, year and company." }, + provenance: { kind: "manual", declared_by: "Curator", original: { source_file: "source/orders.md", supporting_excerpts: ["Old original wording."] } }, + review_items: [], file: "curated/domain/order-key.md", status: "modified", revision: "rev" }; +const page: EvidencePage = { items: [item], item, total: 1, page: 1, page_size: 25, errors: [], initialized: true, active_revision: "active", pending_revision: null, + location: { repository: "/srv/workspaces/repo", workspace: "/srv/workspaces/repo/sales", runtime_workspace: "/data/registry/repo/sales", host: "Installation host", + command: "tht workspace evidence consolidate --workspace sales", git_commands: ["git status", "git diff", "git add -A", "git commit", "git push"] } }; +function renderPage() { + render(); +} +beforeEach(() => server.use( + http.get("/api/workspaces", () => HttpResponse.json([{ id: "sales", displayName: "Sales" }])), + http.get("/api/workspaces/sales/evidence", () => HttpResponse.json(page)), + http.get("/api/workspaces/sales/evidence/evidence:order-key", () => HttpResponse.json(page)), +)); +async function choose() { await screen.findByRole("option", { name: "Sales" }); await userEvent.selectOptions(screen.getByLabelText("Workspace"), "sales"); } +test("reads complete content, manual lineage and actual host file path", async () => { + renderPage(); await choose(); await userEvent.click(await screen.findByRole("button", { name: "Order key" })); + expect(await screen.findByText("Join by order number, year and company.")).toBeVisible(); + expect(screen.getByText("Manual declaration by Curator.")).toBeVisible(); + expect(screen.getByText(/Original document \(lineage\)/)).toBeVisible(); + expect(screen.getByText("/srv/workspaces/repo/sales/evidence/curated/domain/order-key.md")).toBeVisible(); + expect(screen.queryByRole("button", { name: "Save" })).not.toBeInTheDocument(); +}); +test("shows manual consolidation instructions and all kind templates", async () => { + renderPage(); await choose(); await userEvent.click(await screen.findByText("Edit files, consolidate and commit")); + expect(screen.getByText(page.location.command)).toBeVisible(); + await userEvent.click(screen.getByText("Create a manual Evidence file")); + await userEvent.selectOptions(screen.getByLabelText("Example kind"), "formula"); + expect(screen.getAllByText(/sum\(amount\)/).length).toBeGreaterThan(0); +}); +test("combines complete-archive filters and refreshes pending state", async () => { + server.use(http.get("/api/workspaces/sales/evidence", ({ request }) => { + const query = new URL(request.url).searchParams; + return HttpResponse.json({ ...page, items: query.get("kind") === "domain" && query.get("q") === "order" ? [item] : [], pending_revision: "pending" }); + })); + renderPage(); await choose(); await userEvent.type(screen.getByLabelText("Search Evidence"), "order"); + await userEvent.selectOptions(screen.getByLabelText("Kind"), "domain"); + expect(await screen.findByRole("button", { name: "Order key" })).toBeVisible(); + expect(screen.getByText(/Consolidation is incomplete/)).toBeVisible(); +}); +test("invalid files remain actionable alongside valid units", async () => { + server.use(http.get("/api/workspaces/sales/evidence", () => HttpResponse.json({ ...page, errors: [{ file: "curated/domain/broken.md", message: "Missing Rule section" }] }))); + renderPage(); await choose(); + expect(await screen.findByRole("alert")).toHaveTextContent("curated/domain/broken.md"); + expect(within(screen.getByRole("region", { name: "Evidence list" })).getByRole("button", { name: "Order key" })).toBeVisible(); +}); + +test("source refresh is explicit and presents current versus proposed content before a decision", async () => { + const requests: unknown[] = []; + const source = {id: "a".repeat(64), revision: "b".repeat(64), uri: "https://docs.example.test/orders.md", status: "review", availability: "available", + current: [item], proposed: [{...item, payload: {rule: "Use the new approved key."}}], removed_ids: []}; + server.use(http.get("/api/workspaces/sales/evidence", () => HttpResponse.json({...page, source_reviews: [source]})), + http.post("/api/workspaces/sales/evidence/sources", async ({request}) => { requests.push(await request.json()); return HttpResponse.json({status: "succeeded", warnings: ["Source decision activated locally."]}); })); + renderPage(); await choose(); + await screen.findByText(source.uri); + expect(requests).toHaveLength(0); + await userEvent.click(screen.getByRole("button", {name: "Import or refresh sources"})); + await screen.findByText("Source decision activated locally."); + expect(requests).toEqual([{action: "refresh"}]); + await userEvent.click(screen.getByText(source.uri)); + expect(screen.getByText("Current local Evidence")).toBeVisible(); + expect(screen.getByText("Use the new approved key.")).toBeVisible(); + await userEvent.click(screen.getByRole("button", {name: "Keep local Evidence"})); + expect(requests[1]).toEqual({action: "decide", sourceId: source.id, revision: source.revision, decision: "keep"}); +}); + +test("saved source decisions expose retry and preserve the original choice", async () => { + const source = {id: "a".repeat(64), revision: "b".repeat(64), uri: "s3://docs/orders.md", status: "applying", decision: "replace", availability: "available", + current: [item], proposed: [item], removed_ids: []}; + server.use(http.get("/api/workspaces/sales/evidence", () => HttpResponse.json({...page, source_reviews: [source]})), + http.post("/api/workspaces/sales/evidence/sources", () => HttpResponse.json({status: "failed", warnings: ["Source decision saved, but indexing failed. Retry the same decision."]}))); + renderPage(); await choose(); await userEvent.click(await screen.findByText(source.uri)); + expect(screen.getByRole("button", {name: "Import or refresh sources"})).toBeDisabled(); + expect(screen.queryByRole("button", {name: "Keep local Evidence"})).not.toBeInTheDocument(); + await userEvent.click(screen.getByRole("button", {name: "Retry saved decision"})); + expect(await screen.findByRole("alert")).toHaveTextContent("Retry the same decision"); +}); diff --git a/frontend/src/shell/EvidenceManagementPage.tsx b/frontend/src/shell/EvidenceManagementPage.tsx new file mode 100644 index 00000000..eabc796a --- /dev/null +++ b/frontend/src/shell/EvidenceManagementPage.tsx @@ -0,0 +1,138 @@ +import { useState } from "react"; +import { useQuery } from "@tanstack/react-query"; +import { apiFetch, apiErrorMessage } from "../api/client"; +import { listEvidence, getEvidence, type EvidencePage, type EvidenceUnit } from "../api/evidence"; +import { useAuthGeneration } from "../auth/authState"; +import { Button } from "../components/ui/button"; +import { MarkdownView } from "../viewers/MarkdownView"; +import { EvidenceSources } from "./EvidenceSourceReview"; + +const control = "rounded-md border border-input bg-background px-3 py-2 text-sm focus-visible:ring-3 focus-visible:ring-ring/50"; +const kinds = ["domain", "glossary", "enum", "example", "mapping", "normalization", "formula", "reference"]; +const states: Record = { new: "New file", modified: "Modified", active: "Active locally", removed: "Removed file", review_required: "Needs review", legacy: "Conversion required", invalid: "Invalid file" }; +const templates: Record = { + domain: "## Rule\n\nDescribe the approved domain rule.", + glossary: "## Definition\n\nDefine the concept.\n\n## Synonyms\n\n- Alternative name\n\n## Variants\n", + enum: '## Column\n\nsales.orders.status\n\n## Values\n\n### "A"\n\nActive', + example: "## Question\n\nHow many orders?\n\n## Interpretation\n\nCount distinct orders.", + mapping: "## Concept\n\nOrders\n\n## Tables\n\n- sales.orders\n\n## Columns\n\n- sales.orders.id", + normalization: "## Input\n\nA\n\n## Output\n\nActive\n\n## Rule\n\nExpand the status abbreviation.", + formula: "## Concept\n\nTotal amount\n\n## Columns\n\n- sales.lines.amount\n\n## SQL\n\n```sql\nsum(amount)\n```", + reference: "## URL\n\nhttps://example.org/policy\n\n## Label\n\nPolicy\n\n## Description\n\nDescribe the reference.", +}; + +function CopyText({ text, label, showText = true }: { text: string; label: string; showText?: boolean }) { + const [notice, setNotice] = useState(""); + return
+ {showText && {text}} + {notice} +
; +} + +export function EvidenceManagementPage({ canManage }: { canManage: boolean }) { + const generation = useAuthGeneration(); + const [workspace, setWorkspace] = useState(""); + const workspaces = useQuery({ queryKey: ["evidence-workspaces", generation], + queryFn: () => apiFetch>("/workspaces"), enabled: canManage }); + if (!canManage) return
Evidence administration is not permitted.
; + return
+
+

Evidence management

+

Review workspace knowledge and maintain its Markdown files.

+ +
+ {workspaces.isError &&

{apiErrorMessage(workspaces.error)}

} + {workspace ? + :

Choose a workspace to browse its Evidence and locate the files to edit.

} +
; +} + +function WorkspaceEvidence({ workspace, generation }: { workspace: string; generation: number }) { + const [filters, setFilters] = useState>({ q: "", kind: "", purpose: "", status: "", sort: "title", direction: "asc" }); + const [page, setPage] = useState(1); + const [selected, setSelected] = useState(""); + const list = useQuery({ queryKey: ["evidence", generation, workspace, filters, page], queryFn: () => listEvidence(workspace, { ...filters, page, page_size: 25 }) }); + const detail = useQuery({ queryKey: ["evidence-detail", generation, workspace, selected], queryFn: () => getEvidence(workspace, selected), enabled: Boolean(selected) }); + function filter(key: string, value: string) { setFilters(old => ({ ...old, [key]: value })); setPage(1); } + return
+ {list.data && } + +
+ + filter("purpose", v)} /> + filter(key, e.target.value)} />)} + filter("direction", v)} all={false} /> +
+ {list.isError &&

{apiErrorMessage(list.error)}

} + {list.isPending &&

Loading Evidence files…

} + {list.data?.errors.map(error =>
{error.file}

{error.message}

Correct this file before consolidation.

)} +
+
+

{list.data?.total ?? 0} Evidence units

+
+ {list.data?.items.map(item => )} +
EvidenceKindStatus
{item.id}
{item.kind}{states[item.status]}
+ {list.data?.total === 0 &&

No Evidence matches these filters. Clear the filters or use the file instructions to create a unit.

} + +
+
+ {detail.isError ?

{apiErrorMessage(detail.error)}

: selected && detail.isPending ?

Loading Evidence…

: detail.data?.item ? :

Select an Evidence unit to read its complete content and provenance.

} +
+
+
; +} + +function Select({ label, value, values, labels = {}, onChange, all = true }: { label: string; value: string; values: string[]; labels?: Record; onChange: (value: string) => void; all?: boolean }) { + return ; +} + +function Instructions({ data }: { data: EvidencePage }) { + const [kind, setKind] = useState("domain"); + const example = `---\nschema_version: 4\nid: evidence:my-unit\nkind: ${kind}\nlanguage: en\npurposes: [sql_generation]\n---\n\n# My Evidence\n\n${templates[kind]}\n`; + return
+

{data.pending_revision ? "Consolidation is incomplete. The previous active content remains available." : data.active_revision ? "The core uses the last successfully consolidated version." : "Run the first consolidation to convert and activate the local archive."}

+ {data.location.workspace ? <>

Workspace on the installation host

:

A host file path is unavailable for this installation. Configure a persistent workspace-registry bind mount before editing files from the host.

} +
Edit files, consolidate and commit +
    +
  1. Edit the files under evidence/curated/<kind>/ with your preferred editor. Keep existing IDs. Remove a unit's file to delete it; keep the curated directory.
  2. +
  3. Check Git status and the diff. Saving a file does not update recall. + {data.location.git_commands.slice(0, 2).map(command =>
    )}
  4. +
  5. From the installation directory, run this command. It validates and activates Evidence without a full preprocessing run. If needed, add --installation /absolute/path/thothii-installation.yaml immediately after tht. +
  6. +
  7. After success, review the generated metadata and complete Git commit/push manually. These commands use a POSIX shell on the installation host. A remote host's path is not a file on this browser's computer. + {data.location.git_commands.slice(2).map(command =>
    )}
  8. +
+

Source drafts are preserved under evidence/source/. Maintain the refined files under evidence/curated/. New source acquisition and refresh are not part of consolidation.

+
Create a manual Evidence file
; } +const emptyCard = (): MemoryInput => ({ family: "domain_clarification", subject: "", detail: "", + scope: "", rationale: "", question: "", sql: "", concepts: [], dependencies: [], links: [] }); + +export function MemoryManagementPage({ canManage, onDirtyChange, onBusyChange }: { + canManage: boolean; onDirtyChange: (dirty: boolean) => void; onBusyChange?: (busy: boolean) => void; +}) { + const [workspace, setWorkspace] = useState(""); + const [dirty, setDirty] = useState(false); + const [busy, setBusy] = useState(false); + const generation = useAuthGeneration(); + const workspaces = useQuery({ queryKey: ["memory-workspaces", generation], + queryFn: () => apiFetch>("/workspaces"), enabled: canManage }); + useEffect(() => { onDirtyChange(dirty); return () => onDirtyChange(false); }, [dirty, onDirtyChange]); + useEffect(() => { onBusyChange?.(busy); return () => onBusyChange?.(false); }, [busy, onBusyChange]); + if (!canManage) return
Memory administration is not permitted.
; + return
+
+

Memory management

+

Curate reusable knowledge for each workspace.

+ +
+ {workspaces.isError &&

{apiErrorMessage(workspaces.error)}

} + {workspace ? + :

Choose a workspace to browse its knowledge or create a Memory card.

} +
; +} + +function MemoryWorkspace({ workspace, onDirtyChange, onBusyChange }: { + workspace: string; onDirtyChange: (dirty: boolean) => void; onBusyChange: (busy: boolean) => void; +}) { + const queryClient = useQueryClient(); + const [filters, setFilters] = useState>({ q: "", family: "", sort: "updated_at", direction: "desc" }); + const [page, setPage] = useState(1); + const [selected, setSelected] = useState(null); + const [draft, setDraft] = useState(null); + const [busy, setBusy] = useState(false); + const [notice, setNotice] = useState(""); + const [error, setError] = useState(""); + const [deleteConfirm, setDeleteConfirm] = useState(false); + const [linkSearch, setLinkSearch] = useState(""); + const list = useQuery({ queryKey: ["memory", workspace, filters, page], + queryFn: () => listMemory(workspace, { ...filters, page, page_size: 25 }) }); + const pending = useQuery({ queryKey: ["memory-pending", workspace], queryFn: () => pendingMemory(workspace) }); + const linkOptions = useQuery({ queryKey: ["memory-links", workspace, linkSearch], + queryFn: () => listMemory(workspace, { q: linkSearch, page_size: 15 }), + enabled: draft !== null && linkSearch.trim().length > 0 }); + const dirty = draft !== null && (!selected || JSON.stringify(draft) !== JSON.stringify(memoryInput(selected))); + useEffect(() => { onDirtyChange(dirty); return () => onDirtyChange(false); }, [dirty, onDirtyChange]); + + function leaveDraft() { return !dirty || window.confirm("Discard unsaved Memory changes?"); } + useEffect(() => { onBusyChange(busy); return () => onBusyChange(false); }, [busy, onBusyChange]); + function filter(key: string, value: string) { setFilters(old => ({ ...old, [key]: value })); setPage(1); } + function update(key: K, value: MemoryInput[K]) { + setDraft(old => old ? { ...old, [key]: value } : old); + } + async function refresh() { + await Promise.all([ + queryClient.invalidateQueries({ queryKey: ["memory", workspace] }), + queryClient.invalidateQueries({ queryKey: ["memory-pending", workspace] }), + queryClient.invalidateQueries({ queryKey: ["memory-links", workspace] }), + ]); + } + async function open(id: string) { + if (busy || !leaveDraft()) return; + setBusy(true); setError(""); setDeleteConfirm(false); + try { const card = await getMemory(workspace, id); setSelected(card); setDraft(null); } + catch (e) { setError(apiErrorMessage(e)); } + finally { setBusy(false); } + } + async function mutate(action: () => Promise) { + setBusy(true); setError(""); setNotice(""); + try { + const outcome = await action(); + setNotice(outcome.indexed + ? outcome.action === "delete" ? "Card deleted and removed from recall." : "Saved and indexed." + : outcome.action === "delete" ? "Card deleted. Index cleanup is incomplete; retry below." + : "Saved. Index update is incomplete; retry below. The previous content is excluded from recall."); + if (outcome.action === "delete" && outcome.id === selected?.id) { setSelected(null); setDraft(null); } + else if (outcome.card) { setSelected(outcome.card); setDraft(null); } + setDeleteConfirm(false); await refresh(); + } catch (e) { setError(apiErrorMessage(e)); } + finally { setBusy(false); } + } + function save(event: FormEvent) { + event.preventDefault(); + if (draft) void mutate(() => saveMemory(workspace, { + ...draft, concepts: draft.concepts.map(c => c.trim()).filter(Boolean), + }, selected?.id)); + } + + return <> +
+ + + + + +
+
More filters +
+ {(["concept", "database", "table", "column"] as const).map(key => )} + + {(["updated_after", "updated_before"] as const).map(key => )} +
+
+ {notice &&

{notice}

} + {error &&

{error}

} +
+
+ {list.isPending &&

Loading Memory cards…

} + {list.isError &&

{apiErrorMessage(list.error)}

} + {list.data && <> +

{list.data.total} {list.data.total === 1 ? "card" : "cards"}

+ {list.data.items.length === 0 &&

No cards match these filters. Adjust the filters or create a card.

} +
    {list.data.items.map(card =>
  • +
  • )}
+ + } +
+
+ {!draft && !selected &&

Select a card to inspect its content, scope and provenance.

} + {selected && !draft &&
+

{families[selected.family]}

{selected.subject}

+
+ {(["detail", "scope", "rationale", "question", "sql"] as const).filter(k => selected[k]).map(key =>
+
{key === "detail" ? "Content" : key === "sql" ? "Approved SQL" : key}
+
{selected[key]}
)} +
Provenance
{selected.origin === "manual" ? "Created manually" : `Workflow session ${selected.session_id}`}
+ {selected.decision_seq !== null &&
Source decision
{selected.decision_seq}
} +
Card ID
{selected.id}
+
Updated
{new Date(selected.updated_at).toLocaleString()}
+
Created
{new Date(selected.created_at).toLocaleString()}
+ {selected.concepts.length > 0 &&
Concepts
{selected.concepts.join(", ")}
} + {selected.dependencies.length > 0 &&
Schema dependencies
    {selected.dependencies.map((d, i) =>
  • {[d.database, d.schema_name, d.table, d.column].filter(Boolean).join(" / ")}
  • )}
} + {selected.links.length > 0 &&
Linked cards
    {selected.links.map(link =>
  • )}
} +
+
+
+ {deleteConfirm &&

Delete “{selected.subject}” and its links? Other cards will be kept.

+
+
} +
} + {draft &&
+

{selected ? "Edit card" : "New Memory card"}

+
+ + + {(["detail", "scope", "rationale", "question", "sql"] as const).filter(k => k !== "sql" || draft.family === "solved_question").map(key =>