Add PostgreSQL-backed memory, editable evidence with source review and activation, and human-approved archive repairs across the harness, API, and UI. Include migrations, deployment support, regression coverage, and validation documentation.
Refresh permissions from validated session roles so existing administrator logins can access newly deployed archive management features.
Where completeness is machine-detectable, the reviewer's last substantive
approval now closes the phase itself (same pattern as F3/F4/F8):
- F6: approving the LAST CTE of the plan (kind:"cte_result" with next_cte
now empty) advances the phase; a non-final CTE keeps the phase open and
names the next one.
- F7: kind:"sql" records sql_approved — which IS F7's only advance
prerequisite — and advances immediately.
Two reviewer interactions per session removed, both pure echoes. The
summary phase gate remains only where completeness is a human judgment
(F1, F2 with recorded memories, F5). SKILL.md states the rule and the
five self-closing gates; L1 tests cover last/non-last CTE and sql close.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Audit findings 5.1-5.3.
5.1 `phase reopen` now appends `phase_reopened` BEFORE the artifact
teardown: a crash between the two used to leave later-phase artifacts
deleted with the ledger still at the old phase (resume entered a phase
missing its artifacts). The inverse half-state — reopened with stale later
artifacts — is benign. Order locked by tests/test_phase_reopen_order.py.
5.2 New `tht decision add-batch --doc -`: N substantive decisions in ONE
atomic ledger write (meta types and cte_approved stay on `decision add`;
strictest min-phase enforced). reviewer_schema_linking now builds the
complete curation set and persists it with a single add-batch call — a
mid-loop failure can no longer leave the audit ledger half-written, and a
retry cannot duplicate the first K decisions.
5.3 The anti-bypass hook now also blocks BASH mutations of protected
state (`echo >> review_decisions.jsonl`, `sed -i` on the manifest,
`cat > tht-gate.js`, python open('w'), mv/rm/tee/…): FORBIDDEN only
covered tht subcommands and the write/edit hook only covered pi's own
tools. Read-only access (cat/grep/tail/ls) stays allowed.
Also: knownDecisionTypes is defensive — a workflow meta declaring NO
emits at all (older tht, minimal stubs) skips pre-validation instead of
rejecting every substantive type; with emits present, unknown types are
still rejected before the widget (new L1 test).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The model sometimes stops after receiving 'Fase approvata' without calling
`tht session finalize`, leaving the session open. Now the gate itself calls
finalize after advancing the max phase (F8), making session closure
deterministic regardless of model behavior.
- reviewer_confirm kind:phase: after phase advance at max_phase, gate calls
`tht session finalize <session>` (best-effort with recovery message)
- SKILL.md updated: model no longer needs to call finalize itself
- Tests: 2 new JS tests (auto-finalize at max phase; no-finalize at non-max)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Lever 1: Join-graph via FK logics in annotations + suggest-fks command
- TableAnnotation.foreign_keys field stores curated logical FKs (DWH has no FK constraints)
- tht schema suggest-fks: mine from approved SQL, heuristics (time_key → dim_time),
same-name discovery + explicit --assume flag for multi-owner PKs
- mschema renders 【Foreign keys】 section populated; validation in merge.py
- SKILL.md F4 now reads FKs from mschema-text, no custom data_time_key logic
Lever 2: Context-pack consolidation at kickoff (tht search pack)
- Single embedding of question, reused for schema + evidence + solved searches
- One command: tht search pack <question> --session <id> → retrieval_pack.md
- Graceful degradation when Ollama/vector store unreachable (exit 0, empty sections)
- SKILL.md F1 prescribes as first call; reduces model thinking turns via pre-retrieval
Lever 3: Phase-summary recap v2 auto-construction from session ledger
- tht session show --json includes full decisions ledger
- tht phase meta --json exports 'emits' (substantive decision types per phase)
- Gate appends deterministic 【Decisioni registrate in questa fase】 section (appendLedgerSection)
- Model authors only summary + checks; recap table comes from persisted state (exact by construction)
- SKILL.md Disciplina 6: brief model output, gate fills the rest
Tests: 358 Python (including 10 FK + 3 pack + 1 session-ledger tests) + 111 JS gate tests, all pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Transcript analysis (session 2026-07-06-175012, GLM 5.2) showed F1 at 567s:
182s wasted on a useless `tht schema introspect` (re-introspecting the remote
DWH although physical.yaml was already materialized) plus ~220s of model
thinking inflated by ~7 exploratory turns (--help/find/cat). The actual
searches cost ~15s; reviewer gates (~145s, untouched) are the quality contract.
- schema_cmd.py: introspect now exits 0 with "OK (cache)" in ~1s when
physical.yaml exists; --refresh forces the real re-introspection.
Deterministic cross-model guarantee, verified live on psd (163 tables, 1.2s).
- SKILL.md: F1 toolbox (only `tht search find` + `tht schema render`; no
introspect/--help/filesystem browsing; batch all searches in one turn);
F4 step 1 is render-only with a one-shot introspect fallback.
- tht-gate.js: `tht schema introspect ... --refresh` added to FORBIDDEN
(maintenance stays shell-only, never in-session).
- tests: 4 new pytest cases (cache hit placement proven with fake credentials,
refresh bypass, corrupt-catalog fall-through, render fallback message) and
2 gate anti-bypass JS cases.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Extract shared statusBadgeClass() helper (src/viewers/statusBadge.ts) so
SqlViewer, CteResultViewer and PhaseSummaryViewer can't drift: ok/success/
passed/promoted render green (--success), warn renders amber (--warning),
error/failed stay destructive red, everything else stays neutral outline.
Previously CteResultViewer/PhaseSummaryViewer mapped "ok"/"promoted" to the
default badge variant, which is bg-primary (GSD red) — success states
rendered red.
- enrich.js: buildCteResultV2 now falls back preview.rows to [] instead of
null when last_test.preview_rows is missing (pre-upgrade ok records), and
CteResultViewer reads result.preview?.rows?.length with a null-safe
fallback so it degrades to the "No preview rows" empty state instead of
crashing.
- PhaseSummaryViewer: section.items is optional (model-authored sections can
be prose-only); render (section.items ?? []) instead of crashing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The tht-gate.js reviewer_confirm now builds structured v2 artifacts before the
widget so the reviewer approves gate-derived data, not raw model text:
- new pure modules gate/artifact-contracts.js (soft validators, {ok,errors},
legacy-passthrough) and gate/enrich.js (index/description enrichment,
buildCteResultV2 fusing thin model data with `tht cte info`, phase enrichment)
- cte_plan v2: validate + enrich + persist via `tht cte plan --name … --doc -`
(names derived from data.ctes[]); legacy `names` param kept as fallback
- cte_result v2: rebuild from `tht cte next`/`tht cte info` (sql + preview from
the persisted test record); null/error last_test -> actionable textResult
- phase v2: soft-validate + fill phase from meta + catalog descriptions
- prepareReviewerArguments coerces artifact.data too (GLM double-stringify);
legacy markdown strings pass through unchanged
- SKILL.md: Phase 6 cte_plan payload A + thin cte_result guidance; Discipline 6
payload C example; Discipline 7 reworded for gate-rebuilt cte_result
Legacy (non-v2) paths unchanged. TypeBox stays Type.Any() for artifact.data;
validation is soft (textResult) so models self-correct instead of looping.
All 102 gate JS tests green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Fix six GUI defects, verified live against the real stack:
- NavSessions: mark the in-progress (active) session in yellow (warning
token) — both its status dot and its row background.
- SteerInput: the composer is now an auto-growing textarea that wraps and
grows vertically (caps at 160px, then scrolls); Enter sends, Shift+Enter
inserts a newline.
- WorkflowBar: show a synthetic title of the current phase under the 8 dots
from static EN/IT strings (no LLM); English is displayed to match the chrome.
- CentralStatus: reformat the 5-line system-message tail as a structured,
monospace list with per-line markers and an emphasized last line.
- AppShell + CentralStatus: move the left Model-activity panel toggle off the
working spinner (now a pure status indicator) onto a dedicated arrow button
(→ opens, ← closes).
- gate: rename the phase-confirm button "Conferma e prosegui" → "Salva e
procedi" (builders.js + its L1 test).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
reLoop's "Esc non chiude il gate..." warning is terminal-specific. On pi 0.73
ctx.hasUI was false in RPC, so `if (ctx.hasUI)` effectively meant "TUI only". On
pi >=0.80 hasUI is true in RPC too (dialog-capable UI via the bridge), so the
notice leaked to the browser on any invalid gate response. Guard on
ctx.mode === "tui" to restore the original intent. The fake pi runtime gains
mode:"tui" so the roundtrip test still exercises the notice.
Audit of the other 5 ctx.ui.notify: left as-is. They are valid in both live
modes (TUI and RPC, both hasUI=true) and the gate cannot run headless
(emitAndWait needs a UI), so a guard would be dead code. Their dual-mode
rationalization belongs with the future present() work (tracked in the eval doc).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Some models (GLM 5.2) send object params (reviewer_confirm's artifact,
write_schema_linking's schema_linking) as JSON-encoded strings. These failed
TypeBox validation before execute(), making the model retry in an unbounded loop
(observed ~2100s hang). jsonObjectOrSelf() coerces them back to objects in
prepareReviewerArguments and write_schema_linking.
Also close an anti-bypass gap: pi's write/edit tools were unrestricted on the gate
code, so a looping model actually patched tht-gate.js. GATE_CODE_FILES now blocks
any write under .pi/extensions. Requires a pi restart to take effect.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
buildSelectRequest injected an `altroOption` ({id:"other", opens:freetext}) into the
select descriptor, but SelectWidget/MultiselectWidget never route option.opens — so it
rendered as an inert "Altro — specifica…" button next to the working reserved
"Other — specify" control (which now owns free-text since the reviewer-gate-ux fix).
Remove the injection, the now-unused altroOption()/ALTRO_LABEL, and the unused
allow_other flag from both builders (no frontend consumer). Free-text stays offered on
every gate via the reserved "Other" control. The ArtifactGateWidget opens/LinkageHost
linkage (reject-with-reason capability) is intentionally kept — it is a separate,
tested feature, not the injected duplicate.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A1 — SessionMenu gains a Resume item, gated to status!=="finalized" && !archived
(matching the backend's 409 read-only guard), wired in AppShell to doResume ->
POST /sessions/:id/resume. SessionMenu.test.tsx (3 tests); frontend 96/96, tsc clean.
A2 — diagnosis-first clean-room repro driving `pi --mode rpc` with the backend's
exact resume handshake shows the cold-start stall NO LONGER reproduces on pi
0.79.4 (8/8 chained into `tht session show` + `read SKILL.md` in-turn, fresh and
partway sessions). The earlier narrate-and-stop predates the pi upgrade.
Defense-in-depth anyway: RIPRENDI_KICKOFF hardened to force the in-turn tool call
(gate_resume_kickoff.test.js + live regression 2/2). Gate JS 34/34.
PROJECT_STATE open-item #1 (resume stall) flipped to RESOLVED; cross-model resume
robustness folded into workstream G.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
F — reviewer_select options may now carry a `decision` payload {type, subject,
detail?, rationale?} plus an optional `advance`. Picking such an option IS the
confirmation: the gate persists it directly (tht decision add) and optionally
advances, with no redundant reviewer_decide/reviewer_confirm follow-up gate.
Options without a payload stay ask-only; back/exit/Other never persist.
Pure logic extracted + exported for unit tests: resolveSelectOutcome (classifies
the response) and decisionAddArgs (shared with reviewer_decide, DRY). Gate JS
suite 33/33 (gate_select_decision.test.js, +5); harness pytest 269 unchanged.
Contract docs updated together: reviewer_select tool description, SKILL.md
(widget summary, disciplines 2-3, Phase-1 single-pick), and the CLAUDE.md gate
note. Live verification (model truly emits reviewer_select+decision, decision in
review_decisions.jsonl, no follow-up gate) deferred to workstream G — it is
model-behavior-dependent.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The frontend's reviewer widgets uniformly send the picked option in a `choices`
array (SelectWidget/ArtifactGateWidget: `choices: [optionId]`), but the gate's
reviewer_select and reviewer_confirm(reject) handlers read `resp.choice`
(singular). Result: every single-select gate saw an undefined choice, answered
"Nessuna scelta ricevuta", and re-presented forever — the workflow could never
pass F1. (reviewer_decide/multiselect already read `resp.choices`, so it worked.)
Add a shared selectedChoice(resp) helper reading choices[0] (falling back to the
legacy singular choice); both handlers use it.
TDD: gate/__tests__/gate_choice.test.js RED->GREEN; full gate suite 28/28.
Verified LIVE (Playwright -> real Pi -> GLM 5.2): a single-select F1 answer is
now accepted and the workflow advances (clarification 2/4 -> 3/4). The same run
also live-verified the F1 hang fix (418187a).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Set quietStartup:true in .pi/settings.json to suppress Pi banner on RPC stdout.
Add test pinning both quietStartup and theme values. Document project-local trust
requirement (--approve) so no trust prompt blocks RPC loop iteration.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
When THT_SESSION env var is set, /nuova-domanda injects a kickoff that
tells the model to use the pre-created session id instead of running
`tht session new`. /riprendi-sessione and no-env-var paths unchanged.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Removed event.source === "interactive" guard from entry detection so
/nuova-domanda and /riprendi-sessione activate lockActive+pendingKickoff
regardless of source (TUI or RPC prompt).
- Removed event.source !== "interactive" from free-input filter; lock now
blocks/steers all user input when active, not only interactive keystrokes.
- Added typebox@1.1.38 devDep + fake_pi_runtime.registerCommand (gap from Task 2).
- New test: gate_entry.test.js (2 tests: lock activates on RPC; !-steer passes).
- Full suite: 19/19 pass, zero regressions.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>