Where completeness is machine-detectable, the reviewer's last substantive
approval now closes the phase itself (same pattern as F3/F4/F8):
- F6: approving the LAST CTE of the plan (kind:"cte_result" with next_cte
now empty) advances the phase; a non-final CTE keeps the phase open and
names the next one.
- F7: kind:"sql" records sql_approved — which IS F7's only advance
prerequisite — and advances immediately.
Two reviewer interactions per session removed, both pure echoes. The
summary phase gate remains only where completeness is a human judgment
(F1, F2 with recorded memories, F5). SKILL.md states the rule and the
five self-closing gates; L1 tests cover last/non-last CTE and sql close.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Audit findings 5.1-5.3.
5.1 `phase reopen` now appends `phase_reopened` BEFORE the artifact
teardown: a crash between the two used to leave later-phase artifacts
deleted with the ledger still at the old phase (resume entered a phase
missing its artifacts). The inverse half-state — reopened with stale later
artifacts — is benign. Order locked by tests/test_phase_reopen_order.py.
5.2 New `tht decision add-batch --doc -`: N substantive decisions in ONE
atomic ledger write (meta types and cte_approved stay on `decision add`;
strictest min-phase enforced). reviewer_schema_linking now builds the
complete curation set and persists it with a single add-batch call — a
mid-loop failure can no longer leave the audit ledger half-written, and a
retry cannot duplicate the first K decisions.
5.3 The anti-bypass hook now also blocks BASH mutations of protected
state (`echo >> review_decisions.jsonl`, `sed -i` on the manifest,
`cat > tht-gate.js`, python open('w'), mv/rm/tee/…): FORBIDDEN only
covered tht subcommands and the write/edit hook only covered pi's own
tools. Read-only access (cat/grep/tail/ls) stays allowed.
Also: knownDecisionTypes is defensive — a workflow meta declaring NO
emits at all (older tht, minimal stubs) skips pre-validation instead of
rejecting every substantive type; with emits present, unknown types are
still rejected before the widget (new L1 test).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The model sometimes stops after receiving 'Fase approvata' without calling
`tht session finalize`, leaving the session open. Now the gate itself calls
finalize after advancing the max phase (F8), making session closure
deterministic regardless of model behavior.
- reviewer_confirm kind:phase: after phase advance at max_phase, gate calls
`tht session finalize <session>` (best-effort with recovery message)
- SKILL.md updated: model no longer needs to call finalize itself
- Tests: 2 new JS tests (auto-finalize at max phase; no-finalize at non-max)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Lever 1: Join-graph via FK logics in annotations + suggest-fks command
- TableAnnotation.foreign_keys field stores curated logical FKs (DWH has no FK constraints)
- tht schema suggest-fks: mine from approved SQL, heuristics (time_key → dim_time),
same-name discovery + explicit --assume flag for multi-owner PKs
- mschema renders 【Foreign keys】 section populated; validation in merge.py
- SKILL.md F4 now reads FKs from mschema-text, no custom data_time_key logic
Lever 2: Context-pack consolidation at kickoff (tht search pack)
- Single embedding of question, reused for schema + evidence + solved searches
- One command: tht search pack <question> --session <id> → retrieval_pack.md
- Graceful degradation when Ollama/vector store unreachable (exit 0, empty sections)
- SKILL.md F1 prescribes as first call; reduces model thinking turns via pre-retrieval
Lever 3: Phase-summary recap v2 auto-construction from session ledger
- tht session show --json includes full decisions ledger
- tht phase meta --json exports 'emits' (substantive decision types per phase)
- Gate appends deterministic 【Decisioni registrate in questa fase】 section (appendLedgerSection)
- Model authors only summary + checks; recap table comes from persisted state (exact by construction)
- SKILL.md Disciplina 6: brief model output, gate fills the rest
Tests: 358 Python (including 10 FK + 3 pack + 1 session-ledger tests) + 111 JS gate tests, all pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Transcript analysis (session 2026-07-06-175012, GLM 5.2) showed F1 at 567s:
182s wasted on a useless `tht schema introspect` (re-introspecting the remote
DWH although physical.yaml was already materialized) plus ~220s of model
thinking inflated by ~7 exploratory turns (--help/find/cat). The actual
searches cost ~15s; reviewer gates (~145s, untouched) are the quality contract.
- schema_cmd.py: introspect now exits 0 with "OK (cache)" in ~1s when
physical.yaml exists; --refresh forces the real re-introspection.
Deterministic cross-model guarantee, verified live on psd (163 tables, 1.2s).
- SKILL.md: F1 toolbox (only `tht search find` + `tht schema render`; no
introspect/--help/filesystem browsing; batch all searches in one turn);
F4 step 1 is render-only with a one-shot introspect fallback.
- tht-gate.js: `tht schema introspect ... --refresh` added to FORBIDDEN
(maintenance stays shell-only, never in-session).
- tests: 4 new pytest cases (cache hit placement proven with fake credentials,
refresh bypass, corrupt-catalog fall-through, render fallback message) and
2 gate anti-bypass JS cases.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Extract shared statusBadgeClass() helper (src/viewers/statusBadge.ts) so
SqlViewer, CteResultViewer and PhaseSummaryViewer can't drift: ok/success/
passed/promoted render green (--success), warn renders amber (--warning),
error/failed stay destructive red, everything else stays neutral outline.
Previously CteResultViewer/PhaseSummaryViewer mapped "ok"/"promoted" to the
default badge variant, which is bg-primary (GSD red) — success states
rendered red.
- enrich.js: buildCteResultV2 now falls back preview.rows to [] instead of
null when last_test.preview_rows is missing (pre-upgrade ok records), and
CteResultViewer reads result.preview?.rows?.length with a null-safe
fallback so it degrades to the "No preview rows" empty state instead of
crashing.
- PhaseSummaryViewer: section.items is optional (model-authored sections can
be prose-only); render (section.items ?? []) instead of crashing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The tht-gate.js reviewer_confirm now builds structured v2 artifacts before the
widget so the reviewer approves gate-derived data, not raw model text:
- new pure modules gate/artifact-contracts.js (soft validators, {ok,errors},
legacy-passthrough) and gate/enrich.js (index/description enrichment,
buildCteResultV2 fusing thin model data with `tht cte info`, phase enrichment)
- cte_plan v2: validate + enrich + persist via `tht cte plan --name … --doc -`
(names derived from data.ctes[]); legacy `names` param kept as fallback
- cte_result v2: rebuild from `tht cte next`/`tht cte info` (sql + preview from
the persisted test record); null/error last_test -> actionable textResult
- phase v2: soft-validate + fill phase from meta + catalog descriptions
- prepareReviewerArguments coerces artifact.data too (GLM double-stringify);
legacy markdown strings pass through unchanged
- SKILL.md: Phase 6 cte_plan payload A + thin cte_result guidance; Discipline 6
payload C example; Discipline 7 reworded for gate-rebuilt cte_result
Legacy (non-v2) paths unchanged. TypeBox stays Type.Any() for artifact.data;
validation is soft (textResult) so models self-correct instead of looping.
All 102 gate JS tests green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Fix six GUI defects, verified live against the real stack:
- NavSessions: mark the in-progress (active) session in yellow (warning
token) — both its status dot and its row background.
- SteerInput: the composer is now an auto-growing textarea that wraps and
grows vertically (caps at 160px, then scrolls); Enter sends, Shift+Enter
inserts a newline.
- WorkflowBar: show a synthetic title of the current phase under the 8 dots
from static EN/IT strings (no LLM); English is displayed to match the chrome.
- CentralStatus: reformat the 5-line system-message tail as a structured,
monospace list with per-line markers and an emphasized last line.
- AppShell + CentralStatus: move the left Model-activity panel toggle off the
working spinner (now a pure status indicator) onto a dedicated arrow button
(→ opens, ← closes).
- gate: rename the phase-confirm button "Conferma e prosegui" → "Salva e
procedi" (builders.js + its L1 test).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
reLoop's "Esc non chiude il gate..." warning is terminal-specific. On pi 0.73
ctx.hasUI was false in RPC, so `if (ctx.hasUI)` effectively meant "TUI only". On
pi >=0.80 hasUI is true in RPC too (dialog-capable UI via the bridge), so the
notice leaked to the browser on any invalid gate response. Guard on
ctx.mode === "tui" to restore the original intent. The fake pi runtime gains
mode:"tui" so the roundtrip test still exercises the notice.
Audit of the other 5 ctx.ui.notify: left as-is. They are valid in both live
modes (TUI and RPC, both hasUI=true) and the gate cannot run headless
(emitAndWait needs a UI), so a guard would be dead code. Their dual-mode
rationalization belongs with the future present() work (tracked in the eval doc).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Some models (GLM 5.2) send object params (reviewer_confirm's artifact,
write_schema_linking's schema_linking) as JSON-encoded strings. These failed
TypeBox validation before execute(), making the model retry in an unbounded loop
(observed ~2100s hang). jsonObjectOrSelf() coerces them back to objects in
prepareReviewerArguments and write_schema_linking.
Also close an anti-bypass gap: pi's write/edit tools were unrestricted on the gate
code, so a looping model actually patched tht-gate.js. GATE_CODE_FILES now blocks
any write under .pi/extensions. Requires a pi restart to take effect.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
buildSelectRequest injected an `altroOption` ({id:"other", opens:freetext}) into the
select descriptor, but SelectWidget/MultiselectWidget never route option.opens — so it
rendered as an inert "Altro — specifica…" button next to the working reserved
"Other — specify" control (which now owns free-text since the reviewer-gate-ux fix).
Remove the injection, the now-unused altroOption()/ALTRO_LABEL, and the unused
allow_other flag from both builders (no frontend consumer). Free-text stays offered on
every gate via the reserved "Other" control. The ArtifactGateWidget opens/LinkageHost
linkage (reject-with-reason capability) is intentionally kept — it is a separate,
tested feature, not the injected duplicate.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A1 — SessionMenu gains a Resume item, gated to status!=="finalized" && !archived
(matching the backend's 409 read-only guard), wired in AppShell to doResume ->
POST /sessions/:id/resume. SessionMenu.test.tsx (3 tests); frontend 96/96, tsc clean.
A2 — diagnosis-first clean-room repro driving `pi --mode rpc` with the backend's
exact resume handshake shows the cold-start stall NO LONGER reproduces on pi
0.79.4 (8/8 chained into `tht session show` + `read SKILL.md` in-turn, fresh and
partway sessions). The earlier narrate-and-stop predates the pi upgrade.
Defense-in-depth anyway: RIPRENDI_KICKOFF hardened to force the in-turn tool call
(gate_resume_kickoff.test.js + live regression 2/2). Gate JS 34/34.
PROJECT_STATE open-item #1 (resume stall) flipped to RESOLVED; cross-model resume
robustness folded into workstream G.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
F — reviewer_select options may now carry a `decision` payload {type, subject,
detail?, rationale?} plus an optional `advance`. Picking such an option IS the
confirmation: the gate persists it directly (tht decision add) and optionally
advances, with no redundant reviewer_decide/reviewer_confirm follow-up gate.
Options without a payload stay ask-only; back/exit/Other never persist.
Pure logic extracted + exported for unit tests: resolveSelectOutcome (classifies
the response) and decisionAddArgs (shared with reviewer_decide, DRY). Gate JS
suite 33/33 (gate_select_decision.test.js, +5); harness pytest 269 unchanged.
Contract docs updated together: reviewer_select tool description, SKILL.md
(widget summary, disciplines 2-3, Phase-1 single-pick), and the CLAUDE.md gate
note. Live verification (model truly emits reviewer_select+decision, decision in
review_decisions.jsonl, no follow-up gate) deferred to workstream G — it is
model-behavior-dependent.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The frontend's reviewer widgets uniformly send the picked option in a `choices`
array (SelectWidget/ArtifactGateWidget: `choices: [optionId]`), but the gate's
reviewer_select and reviewer_confirm(reject) handlers read `resp.choice`
(singular). Result: every single-select gate saw an undefined choice, answered
"Nessuna scelta ricevuta", and re-presented forever — the workflow could never
pass F1. (reviewer_decide/multiselect already read `resp.choices`, so it worked.)
Add a shared selectedChoice(resp) helper reading choices[0] (falling back to the
legacy singular choice); both handlers use it.
TDD: gate/__tests__/gate_choice.test.js RED->GREEN; full gate suite 28/28.
Verified LIVE (Playwright -> real Pi -> GLM 5.2): a single-select F1 answer is
now accepted and the workflow advances (clarification 2/4 -> 3/4). The same run
also live-verified the F1 hang fix (418187a).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Set quietStartup:true in .pi/settings.json to suppress Pi banner on RPC stdout.
Add test pinning both quietStartup and theme values. Document project-local trust
requirement (--approve) so no trust prompt blocks RPC loop iteration.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
When THT_SESSION env var is set, /nuova-domanda injects a kickoff that
tells the model to use the pre-created session id instead of running
`tht session new`. /riprendi-sessione and no-env-var paths unchanged.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Removed event.source === "interactive" guard from entry detection so
/nuova-domanda and /riprendi-sessione activate lockActive+pendingKickoff
regardless of source (TUI or RPC prompt).
- Removed event.source !== "interactive" from free-input filter; lock now
blocks/steers all user input when active, not only interactive keystrokes.
- Added typebox@1.1.38 devDep + fake_pi_runtime.registerCommand (gap from Task 2).
- New test: gate_entry.test.js (2 tests: lock activates on RPC; !-steer passes).
- Full suite: 19/19 pass, zero regressions.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The gate's widget-descriptor CONSTRUCTION, extracted into pure testable functions.
Each builder turns plain params into a ui_request descriptor (spec §4.1 taxonomy):
buildSelectRequest, buildMultiselectRequest, buildArtifactGate, buildInfoRequest,
buildFreetextRequest, withChildLinkage. No Pi context, no I/O -- the part of the
gate fully testable in L1 (in JS, in-language, no Python mirror).
Validation in the builders (not just happy-path): select requires title + array
options; multiselect allow_empty:false with zero options throws (a broken widget);
artifact-gate requires an artifact with a kind + a valid action.kind
(confirm/approve_reject/view_only); info level must be info/warning/error. The
Altro escape hatch with freetext linkage is always injected on blocking pick
widgets (no-limbo invariant).
L1: 14 node:test cases -- 3 golden files (select_F1, multiselect_F4,
artifact_gate_F5) pin the exact descriptor shape; fuzzy tests assert bad params
throw clearly rather than silently producing a broken widget.
package.json wires 'npm test' -> node --test (runs alongside pytest). The gate
GLUE (emission, anti-bypass, no-limbo loop) is C2, verified at L2.