Add PostgreSQL-backed memory, editable evidence with source review and activation, and human-approved archive repairs across the harness, API, and UI. Include migrations, deployment support, regression coverage, and validation documentation.
Refresh permissions from validated session roles so existing administrator logins can access newly deployed archive management features.
Where completeness is machine-detectable, the reviewer's last substantive
approval now closes the phase itself (same pattern as F3/F4/F8):
- F6: approving the LAST CTE of the plan (kind:"cte_result" with next_cte
now empty) advances the phase; a non-final CTE keeps the phase open and
names the next one.
- F7: kind:"sql" records sql_approved — which IS F7's only advance
prerequisite — and advances immediately.
Two reviewer interactions per session removed, both pure echoes. The
summary phase gate remains only where completeness is a human judgment
(F1, F2 with recorded memories, F5). SKILL.md states the rule and the
five self-closing gates; L1 tests cover last/non-last CTE and sql close.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Audit findings 5.1-5.3.
5.1 `phase reopen` now appends `phase_reopened` BEFORE the artifact
teardown: a crash between the two used to leave later-phase artifacts
deleted with the ledger still at the old phase (resume entered a phase
missing its artifacts). The inverse half-state — reopened with stale later
artifacts — is benign. Order locked by tests/test_phase_reopen_order.py.
5.2 New `tht decision add-batch --doc -`: N substantive decisions in ONE
atomic ledger write (meta types and cte_approved stay on `decision add`;
strictest min-phase enforced). reviewer_schema_linking now builds the
complete curation set and persists it with a single add-batch call — a
mid-loop failure can no longer leave the audit ledger half-written, and a
retry cannot duplicate the first K decisions.
5.3 The anti-bypass hook now also blocks BASH mutations of protected
state (`echo >> review_decisions.jsonl`, `sed -i` on the manifest,
`cat > tht-gate.js`, python open('w'), mv/rm/tee/…): FORBIDDEN only
covered tht subcommands and the write/edit hook only covered pi's own
tools. Read-only access (cat/grep/tail/ls) stays allowed.
Also: knownDecisionTypes is defensive — a workflow meta declaring NO
emits at all (older tht, minimal stubs) skips pre-validation instead of
rejecting every substantive type; with emits present, unknown types are
still rejected before the widget (new L1 test).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Audit findings 2.1 + 2.2 (high).
2.1 WorkflowBar's static phase list was fiction from F2 on (F2 "Schema
linking" vs memoria, F4 "SQL plan" vs schema_linking, …): every live
session showed the wrong phase name. Both maps now mirror
harness/workflow.yaml (F1 chiarimento … F8 datamart).
2.2 forceAdvance (6ee5bda) let reviewer_decide advance:true bypass the
phase gate on ANY phase, contradicting SKILL.md's "auto-advance only
empty F2 / skipped F6". reviewer_decide is back on advanceIfReady (exit-6
no-op) and tells the model to close via reviewer_confirm; forceAdvance
stays only where selection IS the approval by design: reviewer_schema_linking
(F4) and the F8 promotion close path. SKILL.md now names the three
self-closing gates (F3 rewrite_question, F4 schema-linking advance:true,
F8 memory_promote) so gate and skill state one contract.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Audit finding 1.1 (high): reviewer_memory_promote, rewrite_question,
write_schema_linking, write_cte_sql and write_final_sql still ran execute
without a top-level catch — the same unhandled-rejection class that killed
Pi in reviewer_schema_linking (fixed in 87cb806 for the four reviewer_*
tools). All 9 registered tools now share the pattern: any uncaught throw
becomes a textResult the model can react to.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The last live session crashed during reviewer_schema_linking: an uncaught
exception (likely from execFileSync in currentPhase/phaseMeta or from
cat.columns being undefined) rejected the async execute() Promise. Pi does
not catch rejected tool Promises — Node.js treats them as unhandled
rejections and kills the process.
Fix:
- Wrap all four reviewer tool execute bodies (select/decide/confirm/
schema_linking) in a top-level try/catch → returns a textResult on any
unexpected error instead of crashing Pi.
- reviewer_schema_linking: defensively re-parse `tables` if still a string
(belt-and-suspenders over prepareArguments), guard cat.columns before .map().
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Gate (tht-gate.js):
- Pre-validate decision types against workflow.yaml before showing reviewer widget
- Reject decisions emitted by later phases (min-phase check)
- Copy top-level `kind` into artifact when model forgets it (prevents loop)
- Force-advance on reviewer_decide/schema_linking when advance:true — skip
redundant reviewer_confirm gate
Backend:
- Emit agent_end on clean Pi exit (code 0 + bridge idle) instead of marking failed
Frontend:
- Strip <think> tags from transcript and activity panel
- Fix mermaid render with offscreen container + cleanup
- Graceful mermaid error: show source code instead of red error, fall back to table
Workflow:
- F2 now emits table_promoted and table_excluded (early schema linking decisions)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The model sometimes stops after receiving 'Fase approvata' without calling
`tht session finalize`, leaving the session open. Now the gate itself calls
finalize after advancing the max phase (F8), making session closure
deterministic regardless of model behavior.
- reviewer_confirm kind:phase: after phase advance at max_phase, gate calls
`tht session finalize <session>` (best-effort with recovery message)
- SKILL.md updated: model no longer needs to call finalize itself
- Tests: 2 new JS tests (auto-finalize at max phase; no-finalize at non-max)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Lever 1: Join-graph via FK logics in annotations + suggest-fks command
- TableAnnotation.foreign_keys field stores curated logical FKs (DWH has no FK constraints)
- tht schema suggest-fks: mine from approved SQL, heuristics (time_key → dim_time),
same-name discovery + explicit --assume flag for multi-owner PKs
- mschema renders 【Foreign keys】 section populated; validation in merge.py
- SKILL.md F4 now reads FKs from mschema-text, no custom data_time_key logic
Lever 2: Context-pack consolidation at kickoff (tht search pack)
- Single embedding of question, reused for schema + evidence + solved searches
- One command: tht search pack <question> --session <id> → retrieval_pack.md
- Graceful degradation when Ollama/vector store unreachable (exit 0, empty sections)
- SKILL.md F1 prescribes as first call; reduces model thinking turns via pre-retrieval
Lever 3: Phase-summary recap v2 auto-construction from session ledger
- tht session show --json includes full decisions ledger
- tht phase meta --json exports 'emits' (substantive decision types per phase)
- Gate appends deterministic 【Decisioni registrate in questa fase】 section (appendLedgerSection)
- Model authors only summary + checks; recap table comes from persisted state (exact by construction)
- SKILL.md Disciplina 6: brief model output, gate fills the rest
Tests: 358 Python (including 10 FK + 3 pack + 1 session-ledger tests) + 111 JS gate tests, all pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Transcript analysis (session 2026-07-06-175012, GLM 5.2) showed F1 at 567s:
182s wasted on a useless `tht schema introspect` (re-introspecting the remote
DWH although physical.yaml was already materialized) plus ~220s of model
thinking inflated by ~7 exploratory turns (--help/find/cat). The actual
searches cost ~15s; reviewer gates (~145s, untouched) are the quality contract.
- schema_cmd.py: introspect now exits 0 with "OK (cache)" in ~1s when
physical.yaml exists; --refresh forces the real re-introspection.
Deterministic cross-model guarantee, verified live on psd (163 tables, 1.2s).
- SKILL.md: F1 toolbox (only `tht search find` + `tht schema render`; no
introspect/--help/filesystem browsing; batch all searches in one turn);
F4 step 1 is render-only with a one-shot introspect fallback.
- tht-gate.js: `tht schema introspect ... --refresh` added to FORBIDDEN
(maintenance stays shell-only, never in-session).
- tests: 4 new pytest cases (cache hit placement proven with fake credentials,
refresh bypass, corrupt-catalog fall-through, render fallback message) and
2 gate anti-bypass JS cases.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The tht-gate.js reviewer_confirm now builds structured v2 artifacts before the
widget so the reviewer approves gate-derived data, not raw model text:
- new pure modules gate/artifact-contracts.js (soft validators, {ok,errors},
legacy-passthrough) and gate/enrich.js (index/description enrichment,
buildCteResultV2 fusing thin model data with `tht cte info`, phase enrichment)
- cte_plan v2: validate + enrich + persist via `tht cte plan --name … --doc -`
(names derived from data.ctes[]); legacy `names` param kept as fallback
- cte_result v2: rebuild from `tht cte next`/`tht cte info` (sql + preview from
the persisted test record); null/error last_test -> actionable textResult
- phase v2: soft-validate + fill phase from meta + catalog descriptions
- prepareReviewerArguments coerces artifact.data too (GLM double-stringify);
legacy markdown strings pass through unchanged
- SKILL.md: Phase 6 cte_plan payload A + thin cte_result guidance; Discipline 6
payload C example; Discipline 7 reworded for gate-rebuilt cte_result
Legacy (non-v2) paths unchanged. TypeBox stays Type.Any() for artifact.data;
validation is soft (textResult) so models self-correct instead of looping.
All 102 gate JS tests green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The "respond with gate widgets; use '!' for free text" notice is terminal-
flavored. Like the reLoop Esc guard, it used ctx.hasUI, which on pi >=0.80 is
true in RPC too, so it leaked to the browser. Now guarded on ctx.mode === "tui".
The plain text is still swallowed by the `handled` return in all modes; only the
warning is dropped in RPC. No functional ctx.hasUI guard remains in the gate.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
reLoop's "Esc non chiude il gate..." warning is terminal-specific. On pi 0.73
ctx.hasUI was false in RPC, so `if (ctx.hasUI)` effectively meant "TUI only". On
pi >=0.80 hasUI is true in RPC too (dialog-capable UI via the bridge), so the
notice leaked to the browser on any invalid gate response. Guard on
ctx.mode === "tui" to restore the original intent. The fake pi runtime gains
mode:"tui" so the roundtrip test still exercises the notice.
Audit of the other 5 ctx.ui.notify: left as-is. They are valid in both live
modes (TUI and RPC, both hasUI=true) and the gate cannot run headless
(emitAndWait needs a UI), so a guard would be dead code. Their dual-mode
rationalization belongs with the future present() work (tracked in the eval doc).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Some models (GLM 5.2) send object params (reviewer_confirm's artifact,
write_schema_linking's schema_linking) as JSON-encoded strings. These failed
TypeBox validation before execute(), making the model retry in an unbounded loop
(observed ~2100s hang). jsonObjectOrSelf() coerces them back to objects in
prepareReviewerArguments and write_schema_linking.
Also close an anti-bypass gap: pi's write/edit tools were unrestricted on the gate
code, so a looping model actually patched tht-gate.js. GATE_CODE_FILES now blocks
any write under .pi/extensions. Requires a pi restart to take effect.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- WorkflowBar: center the phase pills (add justify-center).
- tht-gate.js: the gate was sending the descriptive phase name (e.g.
"chiarimento") instead of the workflow.yaml id ("F1") in the widget's
`phase` field, so the frontend's F1..F8 match never hit and pills never
lit up. Rename phaseName -> phaseId, return p.id.
- CentralStatus: expand the live "working" tail from a single truncated
line to up to 5 lines (line-clamp-5 safety net for unbroken paragraphs).
- ModelActivityPanel: render the streamed tail through react-markdown +
remark-gfm instead of raw per-line <p> tags, so emphasis/headings render
and blank-line paragraph breaks (previously stripped) are preserved;
tail-cut now operates on paragraphs instead of physical lines.
Updated tests accordingly (WorkflowBar/CentralStatus/ModelActivityPanel/
AppShell.session-mgmt); full frontend suite (98/98) + harness gate node
tests (34/34) + tsc -b pass.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
tht() gains an optional stdin arg; the tool pipes the schema-linking
object to 'tht session set-schema-linking --file -', which validates
against SchemaLinking and returns the exact error on failure — so the
model stops hand-writing the artifact and validating with ad-hoc python.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
reviewer_confirm kind:'cte_result' registered cte_approved --subject
phase:6, which decision_cmd rejects (exit 5) and next_cte never
recognizes — dead-ending F6. Derive the CTE name from 'tht cte next'
(plan-order single source of truth) and approve by name. sql path
(sql_approved:phase:N) unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A1 — SessionMenu gains a Resume item, gated to status!=="finalized" && !archived
(matching the backend's 409 read-only guard), wired in AppShell to doResume ->
POST /sessions/:id/resume. SessionMenu.test.tsx (3 tests); frontend 96/96, tsc clean.
A2 — diagnosis-first clean-room repro driving `pi --mode rpc` with the backend's
exact resume handshake shows the cold-start stall NO LONGER reproduces on pi
0.79.4 (8/8 chained into `tht session show` + `read SKILL.md` in-turn, fresh and
partway sessions). The earlier narrate-and-stop predates the pi upgrade.
Defense-in-depth anyway: RIPRENDI_KICKOFF hardened to force the in-turn tool call
(gate_resume_kickoff.test.js + live regression 2/2). Gate JS 34/34.
PROJECT_STATE open-item #1 (resume stall) flipped to RESOLVED; cross-model resume
robustness folded into workstream G.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
F — reviewer_select options may now carry a `decision` payload {type, subject,
detail?, rationale?} plus an optional `advance`. Picking such an option IS the
confirmation: the gate persists it directly (tht decision add) and optionally
advances, with no redundant reviewer_decide/reviewer_confirm follow-up gate.
Options without a payload stay ask-only; back/exit/Other never persist.
Pure logic extracted + exported for unit tests: resolveSelectOutcome (classifies
the response) and decisionAddArgs (shared with reviewer_decide, DRY). Gate JS
suite 33/33 (gate_select_decision.test.js, +5); harness pytest 269 unchanged.
Contract docs updated together: reviewer_select tool description, SKILL.md
(widget summary, disciplines 2-3, Phase-1 single-pick), and the CLAUDE.md gate
note. Live verification (model truly emits reviewer_select+decision, decision in
review_decisions.jsonl, no follow-up gate) deferred to workstream G — it is
model-behavior-dependent.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The frontend's reviewer widgets uniformly send the picked option in a `choices`
array (SelectWidget/ArtifactGateWidget: `choices: [optionId]`), but the gate's
reviewer_select and reviewer_confirm(reject) handlers read `resp.choice`
(singular). Result: every single-select gate saw an undefined choice, answered
"Nessuna scelta ricevuta", and re-presented forever — the workflow could never
pass F1. (reviewer_decide/multiselect already read `resp.choices`, so it worked.)
Add a shared selectedChoice(resp) helper reading choices[0] (falling back to the
legacy singular choice); both handlers use it.
TDD: gate/__tests__/gate_choice.test.js RED->GREEN; full gate suite 28/28.
Verified LIVE (Playwright -> real Pi -> GLM 5.2): a single-select F1 answer is
now accepted and the workflow advances (clarification 2/4 -> 3/4). The same run
also live-verified the F1 hang fix (418187a).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- loadEnvFromDotenv: read ctx.cwd/.env into the process environment
- prepareReviewerArguments: normalize/parse reviewer tool inputs
- tidy reserved-option filtering and tht command argument assembly
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
When THT_SESSION env var is set, /nuova-domanda injects a kickoff that
tells the model to use the pre-created session id instead of running
`tht session new`. /riprendi-sessione and no-env-var paths unchanged.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Removed event.source === "interactive" guard from entry detection so
/nuova-domanda and /riprendi-sessione activate lockActive+pendingKickoff
regardless of source (TUI or RPC prompt).
- Removed event.source !== "interactive" from free-input filter; lock now
blocks/steers all user input when active, not only interactive keystrokes.
- Added typebox@1.1.38 devDep + fake_pi_runtime.registerCommand (gap from Task 2).
- New test: gate_entry.test.js (2 tests: lock activates on RPC; !-steer passes).
- Full suite: 19/19 pass, zero regressions.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Implementazione del piano di remediation progressiva sui difetti emersi
dall'analisi dell'harness. Tutto verificato: 214 test Python (incl. L0 su
Postgres reale), 14 test JS del gate, ruff pulito.
Blocco 1 (CRITICA, integrazione gate↔CLI):
- phase advance: gate usa --auto + exit 6; reviewer_confirm kind:phase fa
advance esplicito che applica i prerequisiti (prima non avanzava per le
fasi a conferma umana).
- cte plan riceve i --name dal gate (param names); set-question con id
posizionale; skill `tht search find`; nuovo comando `tht memory save-one`
con dedup hash client-side in save_one_memory.
Blocco 2 (D15, stato post-rollback):
- campo `phase` su DecisionRecord + effective_decisions phase-aware per i
subject "a nome" (cte_approved ecc.); _compute_promotions e finalize sulla
vista effective; finalize confronta col piano CTE effettivo, non glob;
`decision add --retracts` + comando `decision retract`.
Blocco 3 (D7 read-only + D6 manifest):
- assert_read_only su tutti e quattro i codepath (direct + REST);
- manifest author/summary/updated_at/updated_by/schema_version popolati +
helper touch_manifest sulle mutazioni.
Blocco 4-5 (D14a/D14b):
- decision_min_phase data-driven via `emits:` in workflow.yaml;
- formula evidence: status auto, search_formulas, gruppo CLI `tht formula`,
`search find --kind formula`, load_evidence_dir salta i .sql.md.
Blocco 6 (robustezza):
- taskdoc slice promoted_tables + bound enforced; report escaping/bound +
rsplit note; filtro kind reader REST/direct; conteggio upserted robusto;
guard REST run_query non-list; LSH disallineato -> LshIndexError.
Blocco 7 (pulizia):
- dead code gate e KIND_TO_TABLE morto rimossi; doc Postgres-only
(README + connection.py).
Blocco 0 (parziale): test di compatibilità firma gate↔CLI
(tests/integration). Rinviati: fake-Pi runtime completo, artifact-gate da
disco (#23), parità eligibility REST/direct (#28), unificazione
reserved-labels (#30), memory_rejected da deselezione (#33).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>