feat: implement memory and evidence administration with guided repairs
Publish documentation / publish (push) Successful in 1m27s

Add PostgreSQL-backed memory, editable evidence with source review and activation, and human-approved archive repairs across the harness, API, and UI. Include migrations, deployment support, regression coverage, and validation documentation.

Refresh permissions from validated session roles so existing administrator logins can access newly deployed archive management features.
This commit is contained in:
Codex
2026-09-10 10:31:34 +02:00
parent 8fe526dd6e
commit 82e2c91f42
168 changed files with 11914 additions and 1772 deletions
@@ -9,6 +9,10 @@ You receive exactly one normalized Source Evidence request. Return one JSON obje
only a `candidates` array, without Markdown fences, comments, or explanatory text. The
host strictly rejects unknown or missing fields.
Candidate JSON uses response protocol version 1. The host renders accepted candidates
as editable Curated Evidence v4; supply the typed payload and exact excerpts in this
response, leaving file rendering and provenance metadata to the host.
Use only facts present in `normalized_text`. Never merge, cite, or infer facts from
another source. Reuse an `existing_id` only when it was supplied in `previous_units`;
otherwise omit it. Preserve prior reviewed wording when it is still supported. When a
+20 -11
View File
@@ -81,6 +81,9 @@ kind:"cte_result"` of the plan (F6), `reviewer_confirm kind:"sql"` (F7), and
Back" (rollback, see discipline 11), or "Esci/Exit" (session abort). If the
reviewer closes without choosing, the gate re-presents the same widget — there is
no silent skip.
When Memory and Evidence conflict and a shared archive needs correction, read
`archive-repair.md` and use `reviewer_archive_repair`. Its receipt reports the
persistent correction; the normal phase gate still approves the current question.
6. **Self-contained messages.** When you call any `reviewer_*` tool, ALWAYS include
in the `message` (or in the `options`' labels/descriptions) a concise recap of the
context the reviewer needs to decide: what was asked, what you found, what each
@@ -310,6 +313,10 @@ Prerequisite: Phase 3 closed.
questions show which tables comparable questions used. Cite relevant precedents
(session id + tables) to the reviewer as CONTEXT — they are reference material,
NOT decisions to apply; their filters/periods may not transfer.
Run `tht memory rules "<question>" --session <id> --json` for reusable join rules
and explained errors. Read `memory-review.md` when a rule is relevant or a reviewer
approves a reusable correction. Present the applicable rule and its Memory ID in
the existing table/join proposal; that gate decides its use for this question.
2. Propose tables to promote/exclude with **`reviewer_schema_linking`**: pass
`tables[]` as `{id, name, kind: "promote"|"exclude", rationale, suggested_columns}`.
Do NOT list every column yourself — the gate loads the full column set (with
@@ -383,6 +390,8 @@ Prerequisite: Phase 5 closed.
1. Read `cte.md`. Decompose the rewritten question into CTEs (Agent View Generation):
each CTE captures an informative subset with a clear purpose, named in snake_case.
Consult `tht memory rules "<question>" --session <id> --json` and follow
`memory-review.md` for reusable calculation rules and explained errors.
`tht memory solved-search "<question>" --json` shows how similar solved questions
were structured — use as reference only.
2. Present the full CTE plan to the reviewer with `reviewer_confirm kind:"cte_plan"`,
@@ -428,6 +437,8 @@ Prerequisite: Phase 6 closed.
1. Read `sql-generation.md`. Recursive divide-and-conquer: the CTEs approved in
Phase 6 are the preferred building blocks (reuse them by name).
Consult `tht memory rules "<question>" --session <id> --json`; show applicable
rules and Memory IDs in the SQL explanation approved by the existing SQL gate.
`tht memory solved-search "<question>" --json` gives the final SQL of similar
solved questions: reference exemplars — never copy filters, periods or
populations without checking them against the current rewritten question.
@@ -460,15 +471,12 @@ Prerequisite: Phase 7 closed.
2. On a server, if the reviewer chose yes: `tht datamart generate` (stub — raises
NotImplementedError for now). Tell
the reviewer that dbt generation is not implemented yet.
3. **Memory promotion closes the session.** Call `reviewer_memory_promote` with ONLY
the session id: the gate computes the candidates itself (`tht memory promote
--preview` — the 3 reusable types, already excluding promoted/declined ones) and
shows the reviewer a pre-selected checklist. Selected → saved to the vectordb +
`memory_promoted`; deselected → `memory_promotion_declined` (never re-proposed).
After recording the promotion (even with zero candidates) the gate advances F8 and
finalizes the session itself — do NOT present a `reviewer_confirm kind:"phase"`
afterwards: there is nothing left to approve. When the gate answers "sessione
finalizzata", give the reviewer the final summary and end the turn.
3. **Memory review closes the session.** Follow `memory-review.md` to finish any
persisted proposals, then call `reviewer_memory_promote` with the session ID.
The reviewer edits and selects additions, explicit updates and links in one
summary, including the consultative solved question. Only the selected cards
are saved. The gate records `memory_summary_reviewed`, advances F8 and finalizes.
Once it reports finalization, give the session summary and end the turn.
4. If the gate reports an error instead (e.g. the datamart decision is missing),
fix the prerequisite and call `reviewer_memory_promote` again. Only if the gate
says the session is still open, close with `reviewer_confirm kind:"phase"` as a
@@ -478,7 +486,8 @@ Prerequisite: Phase 7 closed.
When the promotion gate (or, as fallback, the F8 phase gate) closes Phase 8, the
gate calls `tht session finalize` automatically.
Finalize also indexes the question→SQL pair in the vectordb (kind `solved_question`,
best-effort — on failure recover with `tht memory solved-index <id>`). The persisted
Exemplars are saved only when selected in the Memory summary. Finalize performs
no additional automatic Memory writes. Pending indexing is recovered through
Memory management. The persisted
state (ledger `review_decisions.jsonl` + artifacts) is the truth: what is not
recorded did not happen.
@@ -0,0 +1,37 @@
# Resolve a conflict in the shared archives
Use this gate when the retrieved Memory and Evidence disagree and the reviewer
must decide which archive to correct. State the conflict, its effect on the current
question, and the concrete alternative corrections. Each choice replaces one existing
Memory Card or one Evidence unit; offer only changes that resolve the stated conflict.
1. Read the complete current target and revision through
`tht memory repair-target --session <id> --archive memory|evidence --target-id <id> --json`.
Keep all content fields that the proposed correction does not change. An Evidence
correction retains its identity, kind and original source history. Evidence must
already be consolidated; external file edits must be reconciled first.
2. Call `reviewer_archive_repair` with `session` and `proposal`:
`{reason, options:[{id,label,archive,target_id,revision,content}]}`.
`content` is the complete resulting card or Evidence unit. Use one to five distinct
choice IDs, excluding the reserved `reject` and `continue`. The gate loads the current
content itself and shows both versions. The human selects the archive correction.
3. A rejection means every proposal was inadequate. Reformulate the choices using the
reviewer's feedback and present a new proposal. No archive mutation follows rejection.
A non-administrator can reject or continue the question; shared corrections require
an administrator. Do not disguise a shared correction as a final-summary promotion.
4. The result distinguishes `active`, `pending_activation`, `applying`, and `superseded`.
`saved` alone does not establish retrieval availability. The gate offers retry of
the same approved correction after an index failure. A newer archive edit requires
reconciliation and a new proposal; replay never overwrites that edit.
5. After interruption, run `tht memory repairs --session <id> --json`, then call
`reviewer_archive_repair` with `session` and `repair_id`. This refreshes current
archive status. The list's `recorded_status` is historical. Resume the existing
receipt for recovery instead of proposing the already-saved correction again.
6. Continue the ordinary clarification/schema/SQL gate for the current question,
carrying the chosen meaning and the reported archive status. Archive repair does
not advance a phase. If activation remains pending, report that explicitly.
The extension alone invokes `repair-apply` after the human response. Model-authored
shell commands may prepare or inspect proposals, but cannot apply them. Correction
receipts, like phase artifacts, persist independently of the live chat. Git remains
an operator action after reviewing the Evidence file diff.
@@ -0,0 +1,74 @@
# Reusable Memory during the workflow
Consult rules at schema linking and SQL construction with
`tht memory rules "<current question>" --session <id> --json`.
An optional `--filters '{"table":"orders","column":"id"}'` narrows the physical
context. The runtime fixes database and schema. A returned card is a candidate:
explain its scope and Memory ID in the existing join, CTE or SQL proposal. That
gate approves the concrete use; a retrieved link is not an approval. Exemplars
remain references to previous questions.
## Prepare additions and updates
After a reviewer approves a reusable definition, join/calculation rule or explained
correction, prepare a card citing the effective decision sequence(s) from
`tht session show <id>`. Preserve the explanation and physical dependencies.
Keep query-specific choices (for example only using 2024) in the solved question.
An unselected option, unexplained rejection or timeout is not a reusable source.
When an error and correction state the same rule, prepare one card.
Write the complete current proposal list to a temporary JSON file and run
`tht memory propose --session <id> --data <file.json>`. This persists
`memory_proposals.json` in the session without changing the shared archive.
Reuse proposal IDs when refining their wording. The accepted shape is:
```json
[
{
"id": "order-grain",
"source_seqs": [11],
"reason": "The reviewer corrected repeated header totals; this also applies to future questions.",
"card": {
"family": "sql_rule",
"subject": "Order totals after joining lines",
"detail": "Aggregate each order once before combining it with line totals.",
"scope": "Sales orders and order lines",
"rationale": "A line join repeats the header amount once per line.",
"concepts": ["order grain"],
"dependencies": [{"database": "warehouse", "schema_name": "sales", "table": "orders", "column": "total"}],
"links": []
}
}
]
```
Use the actual decision numbers and schema identifiers from this session. Families
are `domain_clarification`, `sql_rule`, `explained_error` and `solved_question`.
Explained errors require both the corrected behavior and an approved explanation.
The summary adds domain clarifications and the current approved exemplar when they
are not covered by authored proposals, so author only the additions/changes needed.
For an update, include `target_id` and `target_revision` from the current recalled
card, explain the change in `reason`, and provide its complete resulting content.
The reviewer sees the current card alongside the proposal. A conflicting manual
edit requires a refreshed proposal and another review; it is not overwritten.
Keep different rules with different scopes separate. Exact duplicates do not need
another card. Semantic deletion belongs to Memory management.
Links contain `target_id` and `meaning`. Use a current Memory ID for an existing
destination or `proposal:<proposal-id>` for another new card in the summary. The
reviewer may edit/remove links; a link to a new card requires selecting that card.
## Final review
At the end of F8, call `reviewer_memory_promote`. The widget permits content, scope,
concept, dependency and link edits and accepts an empty selection. Approved exemplar
SQL remains the session's solution; changing it requires returning to SQL review.
The gate applies selected cards and links atomically in PostgreSQL and then updates
Qdrant. An incomplete index update is reported and retained for explicit retry.
An interrupted review is recovered from its persisted receipt rather than applied
again. Once the gate reports finalization, end the session without another gate.
For a Memory/Evidence conflict requiring a persistent correction, follow
`archive-repair.md`. A correction already saved by that gate does not need a
duplicate update in the final summary.
@@ -1,12 +1,9 @@
3. **Memory promotion closes the session.** Call `reviewer_memory_promote` with ONLY
the session id: the gate computes the candidates itself (`tht memory promote
--preview` — the 3 reusable types, already excluding promoted/declined ones) and
shows the reviewer a pre-selected checklist. Selected → saved to the vectordb +
`memory_promoted`; deselected → `memory_promotion_declined` (never re-proposed).
After recording the promotion (even with zero candidates) the gate advances F8 and
finalizes the session itself — do NOT present a `reviewer_confirm kind:"phase"`
afterwards: there is nothing left to approve. When the gate answers "sessione
finalizzata", give the reviewer the final summary and end the turn.
3. **Memory review closes the session.** Follow `memory-review.md` to finish any
persisted proposals, then call `reviewer_memory_promote` with the session ID.
The reviewer edits and selects additions, explicit updates and links in one
summary, including the consultative solved question. Only the selected cards
are saved. The gate records `memory_summary_reviewed`, advances F8 and finalizes.
Once it reports finalization, give the session summary and end the turn.
4. If the gate reports an error instead (e.g. the datamart decision is missing),
fix the prerequisite and call `reviewer_memory_promote` again. Only if the gate
says the session is still open, close with `reviewer_confirm kind:"phase"` as a
@@ -1,2 +1,3 @@
Finalize also indexes the question→SQL pair in the vectordb (kind `solved_question`,
best-effort — on failure recover with `tht memory solved-index <id>`). The persisted
Exemplars are saved only when selected in the Memory summary. Finalize performs
no additional automatic Memory writes. Pending indexing is recovered through
Memory management. The persisted
@@ -2,3 +2,7 @@
questions show which tables comparable questions used. Cite relevant precedents
(session id + tables) to the reviewer as CONTEXT — they are reference material,
NOT decisions to apply; their filters/periods may not transfer.
Run `tht memory rules "<question>" --session <id> --json` for reusable join rules
and explained errors. Read `memory-review.md` when a rule is relevant or a reviewer
approves a reusable correction. Present the applicable rule and its Memory ID in
the existing table/join proposal; that gate decides its use for this question.
@@ -1,2 +1,4 @@
Consult `tht memory rules "<question>" --session <id> --json` and follow
`memory-review.md` for reusable calculation rules and explained errors.
`tht memory solved-search "<question>" --json` shows how similar solved questions
were structured — use as reference only.
@@ -1,3 +1,5 @@
Consult `tht memory rules "<question>" --session <id> --json`; show applicable
rules and Memory IDs in the SQL explanation approved by the existing SQL gate.
`tht memory solved-search "<question>" --json` gives the final SQL of similar
solved questions: reference exemplars — never copy filters, periods or
populations without checking them against the current rewritten question.
@@ -81,6 +81,9 @@ kind:"cte_result"` of the plan (F6), `reviewer_confirm kind:"sql"` (F7), and
Back" (rollback, see discipline 11), or "Esci/Exit" (session abort). If the
reviewer closes without choosing, the gate re-presents the same widget — there is
no silent skip.
When Memory and Evidence conflict and a shared archive needs correction, read
`archive-repair.md` and use `reviewer_archive_repair`. Its receipt reports the
persistent correction; the normal phase gate still approves the current question.
6. **Self-contained messages.** When you call any `reviewer_*` tool, ALWAYS include
in the `message` (or in the `options`' labels/descriptions) a concise recap of the
context the reviewer needs to decide: what was asked, what you found, what each