Files
ThothII/docs/gestione-memory.md
T
2026-09-15 14:37:29 +02:00

231 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Memory management
Memory holds reusable knowledge for a workspace. PostgreSQL is the authoritative
archive; Qdrant contains a rebuildable search projection. The harness owns both
the administrative operations and the verification of retrieved results.
## Administration
Open **Administration → Memory management**, immediately after Database management.
An administrator selects the workspace explicitly. No active session, DWH binding,
embedding service or Qdrant connection is required to browse and edit the archive.
PostgreSQL must be available.
The page provides a paginated list, text/ID search, stable sorting, and combined
filters for family, concepts, database, table, column, origin and update date.
Filtering happens across the entire archive before pagination. Each card has a
complete detail view, an editable form, cancellation of unsaved changes, and an
explicit deletion confirmation.
| Family | Content |
| --- | --- |
| Domain clarification | A reusable definition or interpretation, with scope and context. |
| SQL rule | Guidance for constructing SQL, with scope and rationale. |
| Solved question | A question, its approved SQL and context; used as a consultative exemplar. |
| Explained error | A correction and its rationale, to avoid repeating a known error. |
All cards have a stable `mem-<UUID>` identity, title, scope, origin and timestamps.
Manual cards have no invented source session or decision. Workflow cards retain
their source references when an administrator edits them. There is no editorial
revision history.
The form also manages concepts, structured database/schema/table/column
dependencies, and links to other cards with an explicit meaning. Links can only
connect cards in the same workspace. Card, dependency and outgoing-link changes
are committed together. Deleting a card removes its incident links and dependencies,
while keeping the other cards, Evidence and Catalog metadata.
## Save, failure and recovery
```mermaid
flowchart LR
EDIT["Admin or workflow"] --> SQL["PostgreSQL transaction"]
SQL --> CARD["Current card and links"]
SQL --> WORK["Pending projection"]
WORK --> Q["Qdrant"]
Q --> CHECK["Verify current card and projection"]
CARD --> CHECK
CHECK --> REVIEW["Recall for human review"]
```
The service serializes each workspace's mutations. It first commits content and
the projection operation in PostgreSQL, then propagates to Qdrant. A failed save
is different from **saved, index update incomplete**. In the latter case the
current content is already available in administration, while its previous vector
result is excluded from recall.
The **Pending index updates** section offers an explicit retry, including cleanup
for deleted cards. These operations survive a restart. Retry reads the current
archive state and cannot restore an earlier edit or a deleted card. Technical
revisions are internal consistency markers, not user-managed card statuses.
Every recall hit must resolve to a current card in the requested workspace, with
a matching valid projection. The response is reconstructed from PostgreSQL, never
from an unverified Qdrant payload. Deleted, orphaned and stale points are excluded
for both domain Memory and solved-question exemplars. An unavailable archive is
an operational error, not a successful empty archive.
Rebuilding Memory uses only authoritative cards. It does not import JSONL files,
historical session artifacts or old vector payloads. Reference preprocessing keeps
the Memory collection separate and does not migrate legacy Memory payloads.
## Hybrid recall and links
Memory uses dense embeddings and Qdrant BM25, fused with reciprocal rank fusion.
Both branches receive the same workspace, family and scope filters before candidate
selection. Indexed text includes content, scope, rationale, question, exemplar SQL,
concepts and qualified physical dependencies. Both indexing and querying use the
workspace language (`en` or `it`), including manual administration without a DWH binding.
The workflow CLI binds recall to its configured database and schema. Optional
`--filters` JSON can narrow the business `scope`, `concepts`, `table` and `column`.
Business scope is an exact string; every requested concept must be present. Physical
context matches a single structured dependency: database, schema, table and column
cannot be satisfied by unrelated entries. A card without dependencies is workspace-wide;
a database-only or table-only dependency also applies to descendants. Without a table
filter, table-specific knowledge in the selected database/schema remains discoverable.
Filters cannot override the configured database/schema. Free-text business scope is
not automatically interpreted or inferred from the question.
The core expands outgoing links from current, eligible search candidates. It applies
the same filters and workflow-family restrictions to destinations and intermediate
cards. Removed, pending, previously decided and out-of-scope cards cannot act as bridges.
Traversal allows two hops, at most 20 outgoing links per card (stable target-ID order),
200 distinct card lookups and 400 inspected links in total. Search requests
`min(100, max(20, 3 × top))` seeds; the result limit remains between 1 and 100.
All candidates are ranked together using `1 / (60 + seed rank)` for direct hits,
plus the strongest linked contribution, decayed by `0.5` per hop. Repeated paths
do not accumulate votes. Ties use card ID. The returned `score` is a ranking score,
not cosine similarity or a confidence estimate. `retrieval.path` shows the strongest
link path, or just the card ID for an exclusively direct result. PostgreSQL content,
links and eligibility are resolved under the workspace operation lock after search.
Migration `002_hybrid_projection.sql` marks the format of old dense projections as
incompatible without changing their authoritative cards. They appear in pending
updates and are excluded from recall until explicit retry or `memory index` succeeds.
Memory writes can add a missing BM25 sparse vector to their collection; an incompatible
existing vector configuration fails visibly and leaves recovery pending. Reading never
silently falls back to dense retrieval. Reference remains independently managed.
An explicit `memory index` can also recreate a missing Memory collection before
rebuilding its projections; it never imports records from another source.
## Workflow integration
During the workflow, Pi prepares reusable proposals in the session artifact
`memory_proposals.json`. Each proposal names effective approved source decisions,
the content and scope, its rationale, and any physical dependencies or links.
Unexplained failures, rejected options and simple table selections do not create
reusable knowledge. Exact existing cards are reused; semantic similarity alone
never authorizes replacement. Updates name the existing card and its current revision.
At F8, `reviewer_memory_promote` presents one editable Memory summary, including
the approved solved question. The reviewer chooses additions and updates, edits
their content, scope, dependencies and links, or declines everything. Approved SQL
is read only here: changing the solution requires returning to SQL review.
Only selected cards and their links are committed. Invalid links or stale updates
roll back the entire selection. Links between selected new cards are resolved
inside the same transaction. An explicit update preserves the existing identity
and origin, including manually authored content.
A durable review receipt makes repeated delivery idempotent and recovers the
gap between saving Memory and recording `memory_summary_reviewed` in the session
ledger. Retry never recreates a deleted card. The gate then closes F8 and finalizes;
finalization itself performs no automatic Memory writes. Pending indexing remains
visible and recoverable in administration.
F2 consumes domain clarifications and excludes already decided Memory. In F4,
F6 and F7, `memory rules` retrieves applicable SQL rules and explained errors.
Pi presents their use in the existing schema, CTE or SQL approval gate.
Exemplar search remains consultative. Retrieval never constitutes approval.
Persistent Memory/Evidence conflict repair remains part of the joint X1 increment.
## Physical schema changes
After a successful physical Catalog synchronization, the backend passes the exact
removed tables and columns, database, schema and run identity to Memory. Only cards
with matching structured dependencies are deleted, together with their incident
links and searchable projections. Global cards and objects outside the synchronized
scope survive. Manual Catalog cleanup and failed DWH scans never trigger this deletion.
The Catalog transaction records a pending `memory_cleanup` phase before committing
the physical change. If cleanup or indexing fails, the run retains the original
removals. Retry completes that same operation without rescanning the DWH or relying
on a new diff. A new synchronization is blocked until this cleanup is completed.
The Memory deletion receipt and vector tombstones make the operation repeatable
across restarts. The synchronization drawer reports deleted Memory cards and errors.
## API and commands
The administrative HTTP surface requires `memory.manage`, included in the existing
admin role. The backend checks workspace identity and passes its trusted principal
and a protected request snapshot to the harness. The harness independently checks
the principal; ordinary users cannot bypass administration through the CLI.
Production authentication and CSRF protections apply to the new routes.
| Method | Path below `/api/workspaces/:workspaceId/memory` | Operation |
| --- | --- | --- |
| GET | root | Search, filter and paginate the archive |
| POST | root | Create a card |
| GET | `/:cardId` | Read a complete card |
| PUT | `/:cardId` | Save content, links and dependencies together |
| DELETE | `/:cardId` | Delete a card and incident links |
| GET | `/pending` | List incomplete projection operations |
| POST | `/:cardId/retry` | Retry current projection work, including deletion |
CLI configuration remains a per-command option. Commands emit pure JSON:
```text
tht memory list --filters '{"family":"sql_rule","page":1}' -c <runtime.yaml>
tht memory show <card-id> -c <runtime.yaml>
tht memory create --data <card.json> -c <runtime.yaml>
tht memory update <card-id> --data <card.json> -c <runtime.yaml>
tht memory delete <card-id> --yes -c <runtime.yaml>
tht memory pending -c <runtime.yaml>
tht memory retry <card-id> -c <runtime.yaml>
tht memory index -c <runtime.yaml>
tht memory search "<question>" --session <id> --json -c <runtime.yaml>
tht memory search "<question>" --filters '{"table":"orders","column":"id","scope":"Sales"}' --json -c <runtime.yaml>
tht memory solved-search "<question>" --json -c <runtime.yaml>
tht memory rules "<question>" --session <id> --json -c <runtime.yaml>
tht memory propose --session <id> --data <proposals.json> -c <runtime.yaml>
tht memory summary --session <id> --json -c <runtime.yaml>
tht memory solved-index <session-id> --json -c <runtime.yaml>
```
`memory index` rebuilds both domain and solved-question projections.
`solved-index` only retries an existing authoritative source receipt.
The reviewer gate owns `memory review-apply`; Pi must not call it directly.
Migration `003_review_receipts.sql` adds durable review and physical-cleanup receipts.
The backend uses `memory admin --workspace <id> -c <protected-request.json>` so
administration does not materialize session or DWH configuration.
## Installation and storage
Memory shares the installation's existing PostgreSQL service, using its own
`thoth_memory` schema and versioned harness migration pack. The existing
`catalog-migrate` preparation service runs Catalog migrations followed by
`python -m tht.memory.migrate`. The core image includes both migration runners.
Ordinary API requests never migrate the schema.
Connection credentials come from the existing generated `THT_CATALOG_DB_HOST`,
`THT_CATALOG_DB_PORT`, `THT_CATALOG_DB_NAME`, `THT_CATALOG_RUNTIME_USER` and
`THT_CATALOG_RUNTIME_PASSWORD_FILE`. Migration uses the corresponding migrator
user/password file. Direct URL environments can use `THT_CATALOG_DATABASE_URL`
(runtime) and `THT_CATALOG_MIGRATOR_DATABASE_URL`; the harness also accepts
`THT_CATALOG_RUNTIME_DATABASE_URL`. Credentials are not authored in workspace YAML.
The migration grants the installation login membership in the restricted
`thoth_memory_runtime` role. Every repository transaction sets that role and a
workspace context. Forced row-level policies isolate cards, links, dependencies
and projection operations. This role has Memory DML and migration-status read
access, with no runtime DDL privilege. There are no cascading foreign keys to
Catalog or Evidence. A missing or incompatible schema returns a clear operational
error. Preparation is repeatable and checks migration checksums.
See [the M1 specification](plans/2026-09-08-memory-m1-spec.md) and
[ADR 0018](adr/0018-use-postgres-for-memory-and-qdrant-for-retrieval.md) for the
approved scope and acceptance boundaries.
See [M2 implementation and validation](reports/knowledge-archives-release.md)
for the retrieval checks and real embedding test command.