Files
ThothII/docs/superpowers/plans/2026-07-07-active-memory-promotion-and-solved-questions.md
T

50 KiB
Raw Blame History

Active Memory (Promotion Gate + Solved Questions) Implementation Plan

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Make the tht memory system active: (A) a reviewer gate at end-of-workflow (F8) that proposes the session's reusable decisions for promotion into the global memory, and (B) automatic indexing of the question→SQL pair at tht session finalize, with recall commands consumable in F4/F6/F7.

Architecture: Part A adds two decision types (memory_promoted/memory_promotion_declined), a declined-candidates filter in tht/memory.py, and a new deterministic gate tool reviewer_memory_promote in tht-gate.js that fetches candidates via tht memory promote --preview --json, shows a pre-selected multiselect, then persists via tht memory save-one + ledger decisions. Part B adds a solved_question vector kind (stored in the existing memory pgvector table — no server DDL), a tht/solved.py module (record build + D11-style one-row upsert), two CLI commands (tht memory solved-index, tht memory solved-search), and a best-effort hook in finalize_cmd. SKILL.md is updated to prescribe both flows.

Tech Stack: Python (typer CLI, pydantic, pytest in harness/.venv), JavaScript (Pi extension, node:test via npm test in harness/). No frontend changes: the promotion gate reuses the existing multiselect widget descriptor already rendered by the browser.

Global Constraints

  • tht's -c/--config is a PER-COMMAND option: always AFTER the subcommand (project gotcha; the gate never passes it — config resolves from env/default).
  • --json output must be pristine: only valid JSON on stdout (warnings go to stderr with err=True).
  • User-facing CLI/gate copy is Italian (match existing messages); SKILL.md prose is English with Italian example content (match existing file).
  • Python: ruff line-length 100 (.venv/bin/ruff check . from harness/). No ESLint on JS; match existing tab-indented style in tht-gate.js.
  • Run Python tests with harness/.venv/bin/pytest -q (l2 excluded by default); JS gate tests with npm test from harness/.
  • Do NOT touch frontend/ or backend/.
  • The pgvector is remote: workstation writes go ONLY through vector_write_rest (writer key, upsert-only, no deletes). Never introduce a VectorStore.sync call for solved_question records (its delete-stale step would wipe other sessions' records).
  • The model must never bypass the gate: promotion CLI commands get added to the bash FORBIDDEN list (the gate itself calls the CLI via execFileSync, which the hook does not intercept).

File Structure

File Change Responsibility
harness/tht/decisions.py modify +2 decision types
harness/workflow.yaml modify F8 emits the new types
harness/tht/memory.py modify declined-seq filter; _question_context → public question_context
harness/.pi/extensions/tht-gate.js modify reviewer_memory_promote tool + pure helpers + FORBIDDEN
harness/.pi/extensions/gate/__tests__/gate_memory_promote.test.js create L1 tests for the pure helpers
harness/tht/solved.py create solved-question record build + one-row upsert
harness/tht/vectorstore/reader.py modify solved_question → memory table mapping
harness/tht/vectorstore/rest_writer.py modify allow kind in memory table
harness/tht/cli/memory_cmd.py modify solved-index, solved-search, index_solved_session
harness/tht/cli/session_cmd.py modify best-effort indexing hook in finalize_cmd
harness/.pi/skills/tht-sessione/SKILL.md modify prescribe promotion (F2/F5/F8) + recall (F4/F6/F7, Session end)
harness/tests/test_memory_promotion.py create Tasks 1–2 tests
harness/tests/test_solved_question.py create Task 5 tests
harness/tests/test_solved_build.py create Task 6 tests
PROJECT_STATE.md modify snapshot entry

Design notes locked in:

  • Promotion runs INSIDE F8 (after the datamart decision, before reviewer_confirm kind:"phase"), so the ledger decisions are recorded at phase 8 and the workflow closes with a complete audit trail before finalize. The reusable types were each individually reviewer-approved earlier, and sql_approved already exists by end of F7 — no need to wait for the finalize battery.
  • Declined candidates carry detail: "seq:<decision_seq>" on the memory_promotion_declined decision; declined_promotion_seqs() parses it so a re-opened promotion gate never re-proposes them. Promoted ones are already deduped by the registry key (session_id, decision_seq).
  • solved_question reuses the memory pgvector table (kinds already share tables: schema_table/schema_column share schema_records). KIND_TO_TABLE/TABLE_TO_KINDS get the mapping; the REST reader filters by kind client-side. No migration.
  • Embedding = the rewritten question only (record.content); the SQL lives in metadata. The dedup hash therefore covers question+SQL (_solved_hash), so a re-finalize that changes only the SQL still updates the row. solved_question records never flow through sync() (documented in the module).
  • Finalize never fails on indexing: the hook wraps index_solved_session in try/except and degrades to a yellow warning (missing writer key, VPN down, etc.).
  • The integration test tests/integration/test_gate_cli_signatures.py extracts gate CLI call sites automatically (memory is already in _GROUPS), so the new gate calls are covered without editing it.

Task 1: Decision types + workflow emits

Files:

  • Modify: harness/tht/decisions.py (the DecisionType Literal, after "datamart_declined",)
  • Modify: harness/workflow.yaml (F8 emits)
  • Test: harness/tests/test_memory_promotion.py (create)

Interfaces:

  • Produces: decision types "memory_promoted" and "memory_promotion_declined" valid in DecisionRecord and accepted by tht decision add from phase 8 (Workflow.decision_min_phase(...) == 8). Convention consumed by Tasks 2–3: subject = the original decision's subject, detail = "seq:<decision_seq>".

  • Step 1: Write the failing tests

"""L1: gate di promozione memorie (F8) — tipi di decisione + filtro candidati.

La promozione era solo CLI facoltativa (mai innescata): il gate reviewer_memory_promote
la rende un passo del workflow. Questi test fissano il contratto harness-side:
- i due nuovi decision type esistono e sono ammessi dalla Fase 8 (workflow.yaml emits)
- i candidati rifiutati al gate (memory_promotion_declined, detail "seq:<n>") non
  vengono riproposti da reusable_promotions/preview_promotions.
"""
from datetime import datetime

from tht.decisions import DecisionRecord
from tht.workflow import load_workflow


def test_promotion_decision_types_are_valid():
    for t in ("memory_promoted", "memory_promotion_declined"):
        d = DecisionRecord(
            seq=1, ts=datetime(2026, 1, 1), type=t, subject="fact_x", detail="seq:5"
        )
        assert d.type == t


def test_promotion_decision_types_min_phase_is_f8():
    wf = load_workflow()
    assert wf.decision_min_phase("memory_promoted") == 8
    assert wf.decision_min_phase("memory_promotion_declined") == 8
  • Step 2: Run tests to verify they fail

Run: cd /Users/mp/projects/ThothII/harness && .venv/bin/pytest tests/test_memory_promotion.py -v Expected: FAIL — pydantic ValidationError (type not in Literal) on the first test; decision_min_phase returning 1 on the second.

  • Step 3: Add the two types to DecisionType

In harness/tht/decisions.py, after the lines

    "datamart_requested",
    "datamart_declined",

insert:

    # F8: promozione memorie riusabili al gate reviewer_memory_promote. subject =
    # subject della decisione originale, detail = "seq:<decision_seq>" (usato da
    # declined_promotion_seqs per non riproporre i candidati rifiutati).
    "memory_promoted",
    "memory_promotion_declined",
  • Step 4: Add the types to F8's emits in harness/workflow.yaml

Replace:

    emits: [datamart_requested, datamart_declined]

with:

    emits: [datamart_requested, datamart_declined, memory_promoted, memory_promotion_declined]
  • Step 5: Run tests to verify they pass

Run: cd /Users/mp/projects/ThothII/harness && .venv/bin/pytest tests/test_memory_promotion.py -v Expected: 2 PASS

  • Step 6: Run the full L1 suite + lint (no regressions)

Run: cd /Users/mp/projects/ThothII/harness && .venv/bin/pytest -q && .venv/bin/ruff check . Expected: all pass, no lint errors.

  • Step 7: Commit
git add harness/tht/decisions.py harness/workflow.yaml harness/tests/test_memory_promotion.py
git commit -m "feat(memory): memory_promoted/memory_promotion_declined decision types (F8 emits)"

Task 2: Declined-candidates filter in memory.py

Files:

  • Modify: harness/tht/memory.py
  • Test: harness/tests/test_memory_promotion.py (extend)

Interfaces:

  • Consumes: the detail: "seq:<n>" convention from Task 1.

  • Produces: declined_promotion_seqs(decisions: list[DecisionRecord]) -> set[int]; reusable_promotions(...) (same signature) now excludes declined seqs; question_context(decisions, manifest) -> str (public rename of _question_context, consumed by Task 6).

  • Step 1: Write the failing tests (append to tests/test_memory_promotion.py)

from tht.decisions import append_decision
from tht.memory import declined_promotion_seqs, reusable_promotions
from tht.session.models import SessionManifest


def _manifest() -> SessionManifest:
    return SessionManifest(
        id="s1", created_at=datetime(2026, 1, 1), question="domanda originale",
        database="db", schema="public",
    )


def test_declined_promotion_seqs_parses_seq_detail():
    d = DecisionRecord(
        seq=9, ts=datetime(2026, 1, 1), type="memory_promotion_declined",
        subject="fact_x", detail="seq:5",
    )
    assert declined_promotion_seqs([d]) == {5}


def test_declined_promotion_seqs_ignores_malformed_and_other_types():
    ds = [
        DecisionRecord(seq=1, ts=datetime(2026, 1, 1),
                       type="memory_promotion_declined", subject="x", detail=""),
        DecisionRecord(seq=2, ts=datetime(2026, 1, 1),
                       type="table_promoted", subject="x", detail="seq:3"),
    ]
    assert declined_promotion_seqs(ds) == set()


def test_reusable_promotions_exclude_declined(tmp_path):
    append_decision(tmp_path, type="table_promoted", subject="fact_a",
                    detail="tab principale", rationale="scelta reviewer")     # seq 1
    append_decision(tmp_path, type="concept_clarified", subject="attivo",
                    detail="flag_attivo = TRUE")                              # seq 2
    append_decision(tmp_path, type="memory_promotion_declined",
                    subject="fact_a", detail="seq:1")                         # seq 3
    cand = reusable_promotions(tmp_path, _manifest(), tmp_path / "registry.jsonl")
    assert [c.decision_seq for c in cand] == [2]
  • Step 2: Run to verify they fail

Run: cd /Users/mp/projects/ThothII/harness && .venv/bin/pytest tests/test_memory_promotion.py -v Expected: FAIL with ImportError: cannot import name 'declined_promotion_seqs'.

  • Step 3: Implement in harness/tht/memory.py

(a) After the decided_memory_ids function, add:

_DECLINED_SEQ_RE = re.compile(r"\bseq:(\d+)\b")


def declined_promotion_seqs(decisions: list[DecisionRecord]) -> set[int]:
    """decision_seq dei candidati che il reviewer ha rifiutato al gate di promozione
    (F8, `memory_promotion_declined` con detail "seq:<n>"): una riapertura del gate
    non deve riproporli. I promossi sono gia' dedupati dal registro."""
    out: set[int] = set()
    for d in decisions:
        if d.type != "memory_promotion_declined":
            continue
        m = _DECLINED_SEQ_RE.search(d.detail or "")
        if m:
            out.add(int(m.group(1)))
    return out

(b) Rename _question_context → question_context (public: Task 6's solved-record build reuses it) and update its one internal call site in _compute_promotions:

def question_context(decisions: list[DecisionRecord], manifest: SessionManifest) -> str:
    rewritten = [d for d in decisions if d.type == "question_rewritten"]
    return rewritten[-1].detail if rewritten else manifest.question

(in _compute_promotions: context = question_context(decisions, manifest))

(c) Replace the body of reusable_promotions with:

def reusable_promotions(
    session_dir: Path, manifest: SessionManifest, registry_path: Path
) -> list[MemoryRecord]:
    """Candidati riusabili (tipi in REUSABLE_TYPES) non ancora promossi ne' rifiutati
    al gate, SENZA cap: il chiamante applica MAX_PROMOTION_CANDIDATES e segnala il
    troncamento."""
    from tht.phase import effective_decisions

    cand = _compute_promotions(
        session_dir, manifest, seqs=None, existing=load_registry(registry_path)
    )
    declined = declined_promotion_seqs(effective_decisions(session_dir))
    return [
        c for c in cand if c.type in REUSABLE_TYPES and c.decision_seq not in declined
    ]
  • Step 4: Run tests to verify they pass

Run: cd /Users/mp/projects/ThothII/harness && .venv/bin/pytest tests/test_memory_promotion.py tests/test_memory_metadata.py tests/test_memory_save_one.py -v Expected: all PASS (the rename must not break existing memory tests).

  • Step 5: Full suite + lint

Run: cd /Users/mp/projects/ThothII/harness && .venv/bin/pytest -q && .venv/bin/ruff check . Expected: all pass.

  • Step 6: Commit
git add harness/tht/memory.py harness/tests/test_memory_promotion.py
git commit -m "feat(memory): filter gate-declined candidates from promotion preview"

Task 3: Gate tool reviewer_memory_promote

Files:

  • Modify: harness/.pi/extensions/tht-gate.js
  • Test: harness/.pi/extensions/gate/__tests__/gate_memory_promote.test.js (create)

Interfaces:

  • Consumes: tht memory promote --session <id> --preview --json (existing; emits [{decision_seq, type, subject, detail, rationale, question_context, tables, concepts}]), tht memory save-one --session <id> --decision <seq> --json (existing), decisionAddArgs / buildMultiselectRequest / emitAndWait / relayIfThtFails (existing in-file), decision types from Task 1.

  • Produces: registered Pi tool reviewer_memory_promote {session: string}; exported pure helpers promotionOptions(candidates), promotionContent(candidates), splitPromotionChoices(candidates, choices).

  • Step 1: Write the failing JS tests

Create harness/.pi/extensions/gate/__tests__/gate_memory_promote.test.js:

const test = require("node:test");
const assert = require("node:assert");
const {
  promotionOptions,
  promotionContent,
  splitPromotionChoices,
} = require("../../tht-gate.js");

// F8 memory-promotion gate: the candidates are DETERMINISTIC (computed by
// `tht memory promote --preview --json`), the model only names the session.
// These tests pin the pure candidate->widget mapping and the choice partition.

const CANDIDATES = [
  { decision_seq: 3, type: "table_promoted", subject: "fact_seeablazione",
    detail: "tabella principale ablazioni", rationale: "scelta dal reviewer",
    question_context: "quante ablazioni nel 2023", tables: ["fact_seeablazione"], concepts: [] },
  { decision_seq: 5, type: "concept_clarified", subject: "paziente attivo",
    detail: "flag_attivo = TRUE", rationale: "",
    question_context: "quante ablazioni nel 2023", tables: [], concepts: ["paziente attivo"] },
];

test("promotionOptions maps candidates to seq-keyed options", () => {
  assert.deepEqual(promotionOptions(CANDIDATES), [
    { id: "seq-3", label: "table_promoted: fact_seeablazione" },
    { id: "seq-5", label: "concept_clarified: paziente attivo" },
  ]);
});

test("promotionContent lists every candidate with detail and context", () => {
  const c = promotionContent(CANDIDATES);
  assert.ok(c.includes("fact_seeablazione"));
  assert.ok(c.includes("flag_attivo = TRUE"));
  assert.ok(c.includes("quante ablazioni nel 2023"));
});

test("splitPromotionChoices partitions by selection", () => {
  const { promote, decline } = splitPromotionChoices(CANDIDATES, ["seq-5"]);
  assert.deepEqual(promote.map((c) => c.decision_seq), [5]);
  assert.deepEqual(decline.map((c) => c.decision_seq), [3]);
});

test("empty or missing choices declines everything", () => {
  assert.equal(splitPromotionChoices(CANDIDATES, []).decline.length, 2);
  assert.equal(splitPromotionChoices(CANDIDATES, undefined).decline.length, 2);
});
  • Step 2: Run to verify they fail

Run: cd /Users/mp/projects/ThothII/harness && npm test Expected: the new file FAILS (promotionOptions is not a function); the rest of the suite passes.

  • Step 3: Add the pure helpers to tht-gate.js

Immediately after the shouldSkipEmptyDecide export, add:

// --- F8 memory-promotion gate: pure candidate->widget mapping (L1-tested) ------
// The candidates come from `tht memory promote --preview --json` (deterministic,
// reviewer-approved decisions only); the model never authors them.
export function promotionOptions(candidates) {
	return candidates.map((c) => ({
		id: `seq-${c.decision_seq}`,
		label: `${c.type}: ${c.subject}`,
	}));
}

export function promotionContent(candidates) {
	return candidates
		.map(
			(c) =>
				`- **${c.type}: ${c.subject}** (decisione #${c.decision_seq})\n` +
				`  ${c.detail || ""}\n` +
				`  Motivo: ${c.rationale || "—"}\n` +
				`  Domanda di contesto: ${c.question_context || "—"}`,
		)
		.join("\n");
}

export function splitPromotionChoices(candidates, choices) {
	const chosen = new Set(choices ?? []);
	const promote = [];
	const decline = [];
	for (const c of candidates) {
		(chosen.has(`seq-${c.decision_seq}`) ? promote : decline).push(c);
	}
	return { promote, decline };
}
  • Step 4: Run the JS tests to verify they pass

Run: cd /Users/mp/projects/ThothII/harness && npm test Expected: all PASS.

  • Step 5: Register the tool

In tht-gate.js, immediately BEFORE the pi.registerTool({ name: "rewrite_question", ...}) block, add:

	pi.registerTool({
		name: "reviewer_memory_promote",
		label: "Promozione memorie riusabili (reviewer)",
		description:
			"F8 (prima della chiusura di fase): propone al reviewer i candidati di promozione " +
			"calcolati dalla CLI (tht memory promote --preview: tipi riusabili concept_clarified/" +
			"table_promoted/table_excluded, max 5, esclusi i gia' promossi/rifiutati). Le selezioni " +
			"vengono salvate nel vectordb (tht memory save-one) e registrate come memory_promoted; " +
			"le deselezioni come memory_promotion_declined (non riproposte). Nessun parametro oltre " +
			"alla sessione: i candidati sono deterministici, NON li scrivi tu.",
		parameters: Type.Object({
			session: Type.String(),
		}),
		async execute(_id, params, _signal, _onUpdate, ctx) {
			lockActive = true;
			const { session } = params;
			const phase = phaseId(ctx, currentPhase(ctx, session));
			let candidates;
			try {
				candidates = JSON.parse(
					tht(ctx, ["memory", "promote", "--session", session, "--preview", "--json"]),
				);
			} catch (e) {
				const msg = (e.stderr || e.message || String(e)).toString().trim();
				return textResult(`Preview di promozione non disponibile: ${msg}`);
			}
			if (!Array.isArray(candidates) || candidates.length === 0) {
				await ctx.ui.notify(
					"Nessuna decisione riusabile da promuovere in memoria per questa sessione.",
					"info",
				);
				return textResult(
					"Nessun candidato di promozione: prosegui con la chiusura della sessione.",
				);
			}
			const options = promotionOptions(candidates);
			const widget = buildMultiselectRequest({
				id: `u${Date.now()}`,
				phase,
				title: "Quali decisioni salvare nella memoria riutilizzabile?",
				allowEmpty: true,
				options,
				selected: options.map((o) => o.id),
				content: promotionContent(candidates),
			});
			const resp = await emitAndWait(ctx, widget);
			if (resp.control === "freetext")
				return textResult(`Altro (reviewer): ${resp.text}. Valuta e ripresenta il gate.`);
			if (resp.control === "back")
				return textResult("Il reviewer vuole tornare indietro.");
			if (resp.control === "exit") return textResult("Il reviewer vuole uscire.");
			const { promote, decline } = splitPromotionChoices(candidates, resp.choices);
			let saved = 0;
			for (const c of promote) {
				const err = relayIfThtFails(
					ctx,
					["memory", "save-one", "--session", session,
						"--decision", String(c.decision_seq), "--json"],
					"",
				);
				if (err) return err;
				const e2 = relayIfThtFails(ctx, decisionAddArgs(session, {
					type: "memory_promoted", subject: c.subject,
					detail: `seq:${c.decision_seq}`, rationale: c.rationale || c.detail || "",
				}), "");
				if (e2) return e2;
				saved++;
			}
			for (const c of decline) {
				const err = relayIfThtFails(ctx, decisionAddArgs(session, {
					type: "memory_promotion_declined", subject: c.subject,
					detail: `seq:${c.decision_seq}`,
				}), "");
				if (err) return err;
			}
			return textResult(
				`Promozione registrata: ${saved} memorie salvate nel vectordb, ` +
					`${decline.length} candidati scartati.`,
			);
		},
	});
  • Step 6: Extend the anti-bypass FORBIDDEN list

In tht-gate.js, change:

const FORBIDDEN = [
	/\btht\s+phase\s+(advance|reopen)\b/,
	/\btht\s+decision\s+add\b/,
	/\btht\s+cte\s+plan\b/,
];

to:

const FORBIDDEN = [
	/\btht\s+phase\s+(advance|reopen)\b/,
	/\btht\s+decision\s+add\b/,
	/\btht\s+cte\s+plan\b/,
	// La promozione in memoria passa dal reviewer (reviewer_memory_promote), mai da shell.
	/\btht\s+memory\s+(promote|save-one)\b/,
];

(The gate's own calls use execFileSync directly and are not intercepted by the bash tool_call hook.)

  • Step 7: Run JS tests + the gate CLI signature integration test

Run: cd /Users/mp/projects/ThothII/harness && npm test && .venv/bin/pytest tests/integration/test_gate_cli_signatures.py -v Expected: all PASS — the signature test auto-extracts the two new ["memory", ...] call sites and verifies --session/--preview/--json/--decision exist on the CLI.

  • Step 8: Commit
git add harness/.pi/extensions/tht-gate.js harness/.pi/extensions/gate/__tests__/gate_memory_promote.test.js
git commit -m "feat(gate): reviewer_memory_promote — deterministic F8 memory-promotion gate"

Task 4: SKILL.md — prescribe the promotion flow

Files:

  • Modify: harness/.pi/skills/tht-sessione/SKILL.md

Interfaces:

  • Consumes: the reviewer_memory_promote tool (Task 3).

  • Produces: the workflow contract the model follows. No code.

  • Step 1: Replace the Phase 2 D11 note (current step 5)

Replace:

5. **Promotion (D11).** To promote ONE memory in `profile=workstation`, use
   `tht memory save-one` (targeted one-row upsert via the writer key). Batch
   promotion (`tht memory promote`) is server-side. Memories live ONLY in the
   vectordb (no local registry).

with:

5. Memories are promoted at the END of the workflow (Phase 8, the
   `reviewer_memory_promote` gate) — never promote from here, never run
   `tht memory promote`/`save-one` yourself (the gate blocks them).
  • Step 2: Remove the Phase 5 optional-promotion step

Replace:

3. **Memory promotion (F5).** If you want to promote memories from this session, use
   `tht memory save-one` (workstation) or `tht memory promote` (server). Memories
   live ONLY in the vectordb (no local registry).
4. Close with `reviewer_confirm kind:"phase"`.

with:

3. Close with `reviewer_confirm kind:"phase"`.
  • Step 3: Rewrite Phase 8 to include the promotion gate

Replace:

1. Ask the reviewer whether they want a datamart (`reviewer_select` yes/no).
2. If yes: `tht datamart generate` (stub — raises NotImplementedError for now). Tell
   the reviewer that dbt generation is not implemented yet.
3. If no: close with `reviewer_confirm kind:"phase"`. The session is finalizable.

with:

1. Ask the reviewer whether they want a datamart (`reviewer_select` yes/no).
2. If yes: `tht datamart generate` (stub — raises NotImplementedError for now). Tell
   the reviewer that dbt generation is not implemented yet.
3. **Memory promotion.** Call `reviewer_memory_promote` with ONLY the session id: the
   gate computes the candidates itself (`tht memory promote --preview` — the 3
   reusable types, already excluding promoted/declined ones) and shows the reviewer a
   pre-selected checklist. Selected → saved to the vectordb + `memory_promoted`;
   deselected → `memory_promotion_declined` (never re-proposed). If the gate reports
   zero candidates, move on — do not retry.
4. Close with `reviewer_confirm kind:"phase"`. The session is finalizable.
  • Step 4: Verify the file renders coherently

Run: grep -n "reviewer_memory_promote\|Promotion (D11)\|Memory promotion (F5)" /Users/mp/projects/ThothII/harness/.pi/skills/tht-sessione/SKILL.md Expected: reviewer_memory_promote appears in F2's step 5 and Phase 8; the two old notes are gone.

  • Step 5: Commit
git add harness/.pi/skills/tht-sessione/SKILL.md
git commit -m "docs(skill): prescribe the F8 memory-promotion gate; drop optional D11 notes"

Task 5: solved_question kind + tht/solved.py

Files:

  • Create: harness/tht/solved.py
  • Modify: harness/tht/vectorstore/reader.py (KIND_TO_TABLE)
  • Modify: harness/tht/vectorstore/rest_writer.py (TABLE_TO_KINDS)
  • Test: harness/tests/test_solved_question.py (create)

Interfaces:

  • Consumes: VectorRecord, content_hash, pack_metadata, VectorRestClient.existing_hashes(table, kinds) / .upsert_records(table, rows) (all existing).

  • Produces: SOLVED_KIND = "solved_question"; solved_question_record(*, session_id, question, sql, tables) -> VectorRecord (id solved:<session_id>, content = question); save_solved_question(record, *, writer, embedder) -> int; _solved_hash(record) -> str. Consumed by Task 6.

  • Step 1: Write the failing tests

Create harness/tests/test_solved_question.py:

"""L1: coppia domanda->SQL risolta (kind solved_question) — memoria attiva parte B.

Una sessione finalizzata produce UN record nel vectordb (tabella `memory`, kind
dedicato): embedding = domanda riscritta, metadata = {question, sql, tables,
session_id}. Upsert one-row stile D11 (mai sync: il suo delete-stale cancellerebbe
i record delle altre sessioni). L'hash di dedup copre domanda+SQL, cosi' un
re-finalize che cambia solo l'SQL aggiorna comunque la riga.
"""
from unittest.mock import MagicMock

from tht.solved import (
    SOLVED_KIND,
    _solved_hash,
    save_solved_question,
    solved_question_record,
)
from tht.vectorstore.reader import tables_for_kinds


def _rec(**kw):
    base = dict(
        session_id="s1", question="quante ablazioni nel 2023",
        sql="SELECT count(*) FROM fact_seeablazione", tables=["fact_seeablazione"],
    )
    base.update(kw)
    return solved_question_record(**base)


def test_record_shape():
    r = _rec()
    assert r.id == "solved:s1"
    assert r.kind == SOLVED_KIND
    assert r.content == "quante ablazioni nel 2023"  # embedding = solo la domanda
    assert r.metadata["sql"].startswith("SELECT")
    assert r.metadata["tables"] == ["fact_seeablazione"]
    assert r.metadata["session_id"] == "s1"


def test_solved_kind_maps_to_memory_table():
    assert tables_for_kinds([SOLVED_KIND]) == ["memory"]


def test_save_upserts_single_row_into_memory_table():
    writer = MagicMock()
    writer.existing_hashes.return_value = {}
    writer.upsert_records.return_value = 1
    embedder = MagicMock()
    embedder.embed_documents.return_value = [[0.1] * 8]

    assert save_solved_question(_rec(), writer=writer, embedder=embedder) == 1
    writer.sync.assert_not_called()
    table, rows = writer.upsert_records.call_args[0]
    assert table == "memory"
    assert len(rows) == 1
    assert rows[0]["record_key"] == "solved:s1"
    assert rows[0]["metadata"]["kind"] == SOLVED_KIND
    assert rows[0]["metadata"]["sql"].startswith("SELECT")


def test_save_skips_when_question_and_sql_unchanged():
    r = _rec()
    writer = MagicMock()
    writer.existing_hashes.return_value = {r.id: _solved_hash(r)}
    embedder = MagicMock()
    assert save_solved_question(r, writer=writer, embedder=embedder) == 0
    embedder.embed_documents.assert_not_called()
    writer.upsert_records.assert_not_called()


def test_sql_change_alone_triggers_reupsert():
    old = _rec()
    new = _rec(sql="SELECT 1")  # stessa domanda, SQL diverso
    writer = MagicMock()
    writer.existing_hashes.return_value = {old.id: _solved_hash(old)}
    writer.upsert_records.return_value = 1
    embedder = MagicMock()
    embedder.embed_documents.return_value = [[0.0] * 4]
    assert save_solved_question(new, writer=writer, embedder=embedder) == 1
  • Step 2: Run to verify they fail

Run: cd /Users/mp/projects/ThothII/harness && .venv/bin/pytest tests/test_solved_question.py -v Expected: FAIL with ModuleNotFoundError: No module named 'tht.solved'.

  • Step 3: Add the kind mappings

In harness/tht/vectorstore/reader.py, change KIND_TO_TABLE to:

KIND_TO_TABLE = {
    "schema_table": "schema_records",
    "schema_column": "schema_records",
    "evidence": "evidence",
    "memory": "memory",
    "solved_question": "memory",  # coppie domanda->SQL: stessa tabella, kind dedicato
}

In harness/tht/vectorstore/rest_writer.py, change TABLE_TO_KINDS to:

TABLE_TO_KINDS = {
    "schema_records": {"schema_table", "schema_column"},
    "evidence": {"evidence"},
    "memory": {"memory", "solved_question"},
}
  • Step 4: Create harness/tht/solved.py
"""Coppie domanda->SQL risolte (kind `solved_question`) — memoria attiva, parte B.

Una sessione finalizzata produce UN record nel vectordb: l'embedding e' la domanda
riscritta (content), il metadata porta l'SQL finale e le tabelle promosse. Vive
nella tabella pgvector `memory` con kind dedicato (nessuna DDL server-side); si
consulta nelle fasi F4/F6/F7 con `tht memory solved-search` come materiale di
riferimento (exemplar), NON come decisione da ri-applicare.

Scrittura: SOLO upsert one-row stile D11 (`save_solved_question`). Questi record
non passano MAI da `VectorStore.sync`/`RestVectorWriter.sync`: il passo
delete-stale del sync, ricevendo il solo record corrente, cancellerebbe le coppie
delle altre sessioni. Per lo stesso motivo l'hash di dedup e' calcolato qui
(domanda+SQL) e non dal solo content come fa il sync.
"""
from tht.vectorstore.records import VectorRecord

SOLVED_KIND = "solved_question"


def solved_question_record(
    *, session_id: str, question: str, sql: str, tables: list[str]
) -> VectorRecord:
    return VectorRecord(
        id=f"solved:{session_id}",
        kind=SOLVED_KIND,
        ref=session_id,
        title=question[:120],
        content=question,
        metadata={
            "question": question,
            "sql": sql,
            "tables": tables,
            "session_id": session_id,
        },
    )


def _solved_hash(record: VectorRecord) -> str:
    # La domanda e' l'embedding (content); l'SQL vive solo nel metadata. L'hash
    # copre entrambi: un re-finalize che cambia solo l'SQL aggiorna la riga.
    from tht.vectorstore.store import content_hash

    return content_hash(record.content + "\n" + str(record.metadata.get("sql", "")))


def save_solved_question(record: VectorRecord, *, writer, embedder) -> int:
    """Upsert one-row della coppia domanda->SQL via writer key (stesso pattern di
    save_one_memory, spec D11): hash dedup client-side, embedding solo se domanda
    o SQL sono cambiati. `writer` e' un VectorRestClient (writer key). Ritorna il
    numero di righe upsertate (0 = invariata)."""
    from tht.vectorstore.rest_writer import pack_metadata

    new_hash = _solved_hash(record)
    existing = writer.existing_hashes("memory", [SOLVED_KIND])
    if existing.get(record.id) == new_hash:
        return 0
    embedding = embedder.embed_documents([record.content])[0]
    return writer.upsert_records("memory", [{
        "record_key": record.id,
        "kind": record.kind,
        "content_hash": new_hash,
        "metadata": pack_metadata(record),
        "embedding": embedding,
    }])
  • Step 5: Run tests to verify they pass

Run: cd /Users/mp/projects/ThothII/harness && .venv/bin/pytest tests/test_solved_question.py tests/test_vector_dual_key.py -v Expected: all PASS (dual-key tests confirm the mapping change breaks nothing).

  • Step 6: Full suite + lint, then commit

Run: cd /Users/mp/projects/ThothII/harness && .venv/bin/pytest -q && .venv/bin/ruff check .

git add harness/tht/solved.py harness/tht/vectorstore/reader.py harness/tht/vectorstore/rest_writer.py harness/tests/test_solved_question.py
git commit -m "feat(solved): solved_question vector kind + one-row upsert (D11 pattern)"

Files:

  • Modify: harness/tht/solved.py (add build_solved_record + SolvedIndexError)
  • Modify: harness/tht/cli/memory_cmd.py (two commands + index_solved_session)
  • Test: harness/tests/test_solved_build.py (create)

Interfaces:

  • Consumes: question_context (Task 2), solved_question_record/save_solved_question (Task 5), effective_decisions (tht/phase.py), promoted_tables_for(cfg, session_id) -> set[str] | None (tht/cli/sql_cmd.py:42), has_vector_write_rest/require_vector_write_allowed (tht/cli/_guards.py), make_embedder/open_searcher/require_vector_cfg (tht/cli/vector_cmd.py), load_session_or_exit/session_dir (already imported at the top of memory_cmd.py).

  • Produces: build_solved_record(session_dir, manifest, promoted_tables) -> VectorRecord (raises SolvedIndexError); index_solved_session(cfg, session_id) -> int (raises RuntimeError when the writer key is missing — the finalize hook, Task 7, degrades it to a warning); CLI tht memory solved-index <session_id> [--json] and tht memory solved-search "<q>" [--top N] [--json].

  • Step 1: Write the failing tests

Create harness/tests/test_solved_build.py:

"""L1: build del record solved_question dagli artefatti persistiti della sessione.

Il record si costruisce SOLO da cio' che il workflow ha approvato: sql_final.sql
presente + decisione sql_approved nella vista effective; la domanda e' l'ultima
question_rewritten (fallback: la domanda del manifest)."""
from datetime import datetime

import pytest

from tht.decisions import append_decision
from tht.session.models import SessionManifest
from tht.solved import SolvedIndexError, build_solved_record


def _manifest() -> SessionManifest:
    return SessionManifest(
        id="s1", created_at=datetime(2026, 1, 1), question="domanda originale",
        database="db", schema="public",
    )


def test_build_uses_rewritten_question_sql_and_tables(tmp_path):
    (tmp_path / "sql_final.sql").write_text("SELECT 1\n")
    append_decision(tmp_path, type="question_rewritten", subject="domanda",
                    detail="domanda riscritta esplicita")
    append_decision(tmp_path, type="sql_approved", subject="phase:7")
    rec = build_solved_record(tmp_path, _manifest(), {"fact_x", "dim_y"})
    assert rec.id == "solved:s1"
    assert rec.content == "domanda riscritta esplicita"
    assert rec.metadata["sql"] == "SELECT 1"
    assert rec.metadata["tables"] == ["dim_y", "fact_x"]  # ordinate


def test_build_falls_back_to_manifest_question(tmp_path):
    (tmp_path / "sql_final.sql").write_text("SELECT 1")
    append_decision(tmp_path, type="sql_approved", subject="phase:7")
    rec = build_solved_record(tmp_path, _manifest(), None)
    assert rec.content == "domanda originale"
    assert rec.metadata["tables"] == []


def test_build_requires_sql_file(tmp_path):
    append_decision(tmp_path, type="sql_approved", subject="phase:7")
    with pytest.raises(SolvedIndexError, match="sql_final.sql"):
        build_solved_record(tmp_path, _manifest(), None)


def test_build_requires_sql_approved(tmp_path):
    (tmp_path / "sql_final.sql").write_text("SELECT 1")
    with pytest.raises(SolvedIndexError, match="sql_approved"):
        build_solved_record(tmp_path, _manifest(), None)
  • Step 2: Run to verify they fail

Run: cd /Users/mp/projects/ThothII/harness && .venv/bin/pytest tests/test_solved_build.py -v Expected: FAIL with ImportError: cannot import name 'build_solved_record'.

  • Step 3: Add build_solved_record to harness/tht/solved.py

Append:

class SolvedIndexError(Exception):
    """La sessione non ha (ancora) gli artefatti per il record solved_question."""


def build_solved_record(session_dir, manifest, promoted_tables) -> VectorRecord:
    """Costruisce il record dagli artefatti persistiti (vista effective D15):
    richiede sql_final.sql e la decisione sql_approved; la domanda e' l'ultima
    question_rewritten, fallback la domanda del manifest."""
    from tht.memory import question_context
    from tht.phase import effective_decisions

    sql_file = session_dir / "sql_final.sql"
    if not sql_file.exists():
        raise SolvedIndexError("sql_final.sql assente")
    decisions = effective_decisions(session_dir)
    if not any(d.type == "sql_approved" for d in decisions):
        raise SolvedIndexError("decisione sql_approved assente")
    return solved_question_record(
        session_id=manifest.id,
        question=question_context(decisions, manifest),
        sql=sql_file.read_text().strip(),
        tables=sorted(promoted_tables or set()),
    )

Run: cd /Users/mp/projects/ThothII/harness && .venv/bin/pytest tests/test_solved_build.py -v Expected: 4 PASS.

  • Step 4: Add the orchestration helper + the two commands to memory_cmd.py

Append at the end of harness/tht/cli/memory_cmd.py:

def index_solved_session(cfg, session_id: str) -> int:
    """Indicizza la coppia domanda->SQL della sessione (kind solved_question).

    Solleva RuntimeError se manca la writer key e SolvedIndexError se mancano gli
    artefatti: il finalize li degrada a warning, il comando CLI li converte in
    errori espliciti."""
    from tht.cli.sql_cmd import promoted_tables_for
    from tht.cli.vector_cmd import make_embedder
    from tht.solved import build_solved_record, save_solved_question
    from tht.vectorstore.rest_client import VectorRestClient

    if not has_vector_write_rest(cfg):
        raise RuntimeError(
            "vector_write_rest assente: la coppia domanda->SQL si indicizza con la "
            "writer key (workstation) o dal server"
        )
    manifest = load_session_or_exit(cfg, session_id)
    record = build_solved_record(
        session_dir(cfg, session_id), manifest, promoted_tables_for(cfg, session_id)
    )
    return save_solved_question(
        record,
        writer=VectorRestClient(cfg.vector_write_rest),
        embedder=make_embedder(cfg.embeddings),
    )


@memory_app.command("solved-index")
def solved_index_cmd(
    session_id: str = typer.Argument(..., help="Id sessione con sql_final.sql approvato."),
    json_out: bool = typer.Option(False, "--json", help="Output JSON (per Pi)."),
    config: Path = CONFIG_OPT,
) -> None:
    """Indicizza la coppia domanda->SQL nel vectordb (backfill; il finalize lo fa da solo)."""
    import json as _json

    from tht.solved import SolvedIndexError

    cfg = _load_config_or_exit(config)
    require_vector_write_allowed(cfg, "memory solved-index")
    try:
        count = index_solved_session(cfg, session_id)
    except RuntimeError as e:
        typer.secho(f"ERRORE: {e}", fg=typer.colors.RED, err=True)
        raise typer.Exit(code=4)
    except SolvedIndexError as e:
        typer.secho(f"ERRORE: sessione {session_id} non indicizzabile: {e}",
                    fg=typer.colors.RED, err=True)
        raise typer.Exit(code=3)
    msg = (
        f"1 coppia domanda->SQL indicizzata (solved:{session_id})."
        if count else "Nessun upsert: coppia gia' aggiornata."
    )
    if json_out:
        typer.echo(_json.dumps({"upserted": count, "id": f"solved:{session_id}"},
                               ensure_ascii=False))
        return
    typer.secho(f"OK: {msg}", fg=typer.colors.GREEN)


@memory_app.command("solved-search")
def solved_search_cmd(
    question: str = typer.Argument(..., help="Domanda da confrontare con quelle risolte."),
    top: int = typer.Option(3, "--top"),
    json_out: bool = typer.Option(False, "--json", help="Output JSON (per Pi)."),
    config: Path = CONFIG_OPT,
) -> None:
    """Domande gia' risolte simili (kind solved_question): domanda, SQL e tabelle."""
    from rich.console import Console
    from rich.table import Table

    from tht.cli.vector_cmd import make_embedder, open_searcher
    from tht.solved import SOLVED_KIND

    cfg = _load_config_or_exit(config)
    require_vector_cfg(cfg)
    searcher = open_searcher(cfg)
    embedder = make_embedder(cfg.embeddings)
    hits = searcher.search(embedder.embed_query(question), top_n=top, kinds=[SOLVED_KIND])
    results = [
        {
            "session_id": h.metadata.get("session_id", h.ref),
            "question": h.metadata.get("question", h.content),
            "sql": h.metadata.get("sql", ""),
            "tables": h.metadata.get("tables", []),
            "score": round(h.similarity, 4),
        }
        for h in hits
    ]
    if json_out:
        typer.echo(json.dumps(results, ensure_ascii=False, indent=2))
        return
    if not results:
        typer.secho("Nessuna domanda risolta simile.", fg=typer.colors.YELLOW)
        return
    table = Table(title=f"Domande risolte simili a: {question}")
    table.add_column("Sessione")
    table.add_column("Domanda")
    table.add_column("Tabelle")
    table.add_column("Score", justify="right")
    for r in results:
        table.add_row(r["session_id"], r["question"][:60],
                      ", ".join(r["tables"]), f"{r['score']:.3f}")
    Console().print(table)

(json and require_vector_cfg are already module-level imports in memory_cmd.py — solved_search_cmd uses them directly; only make_embedder/open_searcher are imported lazily, matching the file's existing pattern.)

  • Step 5: Verify CLI wiring by help text

Run: cd /Users/mp/projects/ThothII/harness && .venv/bin/tht memory solved-index --help && .venv/bin/tht memory solved-search --help Expected: both exit 0 and show --json (and --top on solved-search).

  • Step 6: Full suite + lint, then commit

Run: cd /Users/mp/projects/ThothII/harness && .venv/bin/pytest -q && .venv/bin/ruff check . Expected: all pass.

git add harness/tht/solved.py harness/tht/cli/memory_cmd.py harness/tests/test_solved_build.py
git commit -m "feat(cli): tht memory solved-index / solved-search (question->SQL exemplars)"

Task 7: Finalize hook (best-effort indexing)

Files:

  • Modify: harness/tht/cli/session_cmd.py (finalize_cmd, after the manifest write)

Interfaces:

  • Consumes: index_solved_session(cfg, session_id) (Task 6). Lazy import inside the function (memory_cmd imports FROM session_cmd — a top-level import back would be circular; finalize already uses lazy imports throughout).

  • Produces: tht session finalize indexes the pair automatically; any failure prints a yellow warning and finalize still succeeds.

  • Step 1: Add the hook

In harness/tht/cli/session_cmd.py, inside finalize_cmd, locate:

    manifest.status = "finalized"
    manifest.updated_at = datetime.now(UTC)
    manifest.updated_by = current_author()
    manifest.to_yaml(sdir / MANIFEST)

and insert immediately AFTER it (before the final typer.secho(f"OK: sessione ...")):

    # --- memoria attiva (parte B): indicizza la coppia domanda->SQL, best-effort ---
    # Import lazy: memory_cmd importa da session_cmd (un import top-level qui sarebbe
    # circolare). Qualunque errore (writer key assente, VPN giu', Ollama spento) NON
    # deve bloccare il finalize: l'indice e' derivato e recuperabile con
    # `tht memory solved-index <id>`.
    try:
        from tht.cli.memory_cmd import index_solved_session

        if index_solved_session(cfg, session_id):
            typer.secho(
                "OK: coppia domanda->SQL indicizzata nel vectordb (solved_question).",
                fg=typer.colors.GREEN,
            )
    except Exception as e:
        typer.secho(
            f"ATTENZIONE: coppia domanda->SQL non indicizzata ({e}). "
            f"Recupera con `tht memory solved-index {session_id}`.",
            fg=typer.colors.YELLOW, err=True,
        )
  • Step 2: Verify no import cycle and no regressions

Run: cd /Users/mp/projects/ThothII/harness && .venv/bin/python -c "import tht.cli" && .venv/bin/pytest -q && .venv/bin/ruff check . Expected: import OK, all tests pass, no lint errors.

(The full finalize path needs the real DWH + writer key: it is exercised at L2 during the live gate below, not here.)

  • Step 3: Commit
git add harness/tht/cli/session_cmd.py
git commit -m "feat(finalize): auto-index the question->SQL pair (best-effort, never blocks)"

Task 8: SKILL.md recall (F4/F6/F7 + Session end) + docs

Files:

  • Modify: harness/.pi/skills/tht-sessione/SKILL.md
  • Modify: PROJECT_STATE.md

Interfaces:

  • Consumes: tht memory solved-search (Task 6), the finalize hook (Task 7).

  • Step 1: Phase 4 — add the exemplar consultation to step 1

In the Phase 4 section, replace:

1. `tht schema introspect` + `tht schema render --format mschema-text` for the schema
   context. Copy table/column names EXACTLY from it — never invent objects.

with:

1. `tht schema introspect` + `tht schema render --format mschema-text` for the schema
   context. Copy table/column names EXACTLY from it — never invent objects.
   Also run `tht memory solved-search "<question>" --json`: similar already-solved
   questions show which tables comparable questions used. Cite relevant precedents
   (session id + tables) to the reviewer as CONTEXT — they are reference material,
   NOT decisions to apply; their filters/periods may not transfer.
  • Step 2: Phase 6 — add the exemplar consultation to step 1

Replace:

1. Read `cte.md`. Decompose the rewritten question into CTEs (Agent View Generation):
   each CTE captures an informative subset with a clear purpose, named in snake_case.

with:

1. Read `cte.md`. Decompose the rewritten question into CTEs (Agent View Generation):
   each CTE captures an informative subset with a clear purpose, named in snake_case.
   `tht memory solved-search "<question>" --json` shows how similar solved questions
   were structured — use as reference only.
  • Step 3: Phase 7 — add the exemplar consultation to step 1

Replace:

1. Read `sql-generation.md`. Recursive divide-and-conquer: the CTEs approved in
   Phase 6 are the preferred building blocks (reuse them by name).

with:

1. Read `sql-generation.md`. Recursive divide-and-conquer: the CTEs approved in
   Phase 6 are the preferred building blocks (reuse them by name).
   `tht memory solved-search "<question>" --json` gives the final SQL of similar
   solved questions: reference exemplars — never copy filters, periods or
   populations without checking them against the current rewritten question.
  • Step 4: Session end — document the auto-indexing

Replace:

When the workflow is complete (Phase 8), `tht session finalize` closes the session
and unlocks input. The persisted state (ledger `review_decisions.jsonl` + artifacts)
is the truth: what is not recorded did not happen.

with:

When the workflow is complete (Phase 8), `tht session finalize` closes the session
and unlocks input. Finalize also indexes the question→SQL pair in the vectordb
(kind `solved_question`, best-effort — on failure recover with `tht memory
solved-index <id>`). The persisted state (ledger `review_decisions.jsonl` +
artifacts) is the truth: what is not recorded did not happen.
  • Step 5: Update PROJECT_STATE.md

Add a dated entry (match the file's existing style) summarizing: active memory shipped — F8 promotion gate (reviewer_memory_promote, decision types memory_promoted/memory_promotion_declined, declined-filter) + solved-question exemplars (solved_question kind in the memory table, tht memory solved-index/solved-search, finalize hook, recall prescribed in F4/F6/F7). Note the pending L2 gate: one live end-to-end session on workspace psd to verify (a) the promotion widget renders and persists, (b) finalize indexes the pair, (c) solved-search returns it.

  • Step 6: Final verification sweep

Run: cd /Users/mp/projects/ThothII/harness && .venv/bin/pytest -q && npm test && .venv/bin/ruff check . Expected: everything green.

  • Step 7: Commit
git add harness/.pi/skills/tht-sessione/SKILL.md PROJECT_STATE.md
git commit -m "docs(skill): prescribe solved-question recall in F4/F6/F7; refresh PROJECT_STATE"

Manual L2 gate (after all tasks — human-run, VPN + writer key required)

Not automatable in CI (real GLM + remote pgvector). Run one full session on workspace psd via ./scripts/run-stack.sh:

  1. Complete a question through F8: at step 3 the promotion checklist must appear pre-selected; deselect one candidate, approve.
  2. Verify the ledger: tht decision list --session <id> shows memory_promoted (selected) and memory_promotion_declined (deselected, detail seq:<n>).
  3. Re-open the gate scenario (new session on a similar question): F2 must retrieve the promoted memory; the declined one must not resurface at a re-run of the promotion preview.
  4. tht session finalize <id> prints the green solved-question line; tht memory solved-search "<the question>" --json returns the pair with sql + tables.