Where completeness is machine-detectable, the reviewer's last substantive
approval now closes the phase itself (same pattern as F3/F4/F8):
- F6: approving the LAST CTE of the plan (kind:"cte_result" with next_cte
now empty) advances the phase; a non-final CTE keeps the phase open and
names the next one.
- F7: kind:"sql" records sql_approved — which IS F7's only advance
prerequisite — and advances immediately.
Two reviewer interactions per session removed, both pure echoes. The
summary phase gate remains only where completeness is a human judgment
(F1, F2 with recorded memories, F5). SKILL.md states the rule and the
five self-closing gates; L1 tests cover last/non-last CTE and sql close.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Audit findings 5.1-5.3.
5.1 `phase reopen` now appends `phase_reopened` BEFORE the artifact
teardown: a crash between the two used to leave later-phase artifacts
deleted with the ledger still at the old phase (resume entered a phase
missing its artifacts). The inverse half-state — reopened with stale later
artifacts — is benign. Order locked by tests/test_phase_reopen_order.py.
5.2 New `tht decision add-batch --doc -`: N substantive decisions in ONE
atomic ledger write (meta types and cte_approved stay on `decision add`;
strictest min-phase enforced). reviewer_schema_linking now builds the
complete curation set and persists it with a single add-batch call — a
mid-loop failure can no longer leave the audit ledger half-written, and a
retry cannot duplicate the first K decisions.
5.3 The anti-bypass hook now also blocks BASH mutations of protected
state (`echo >> review_decisions.jsonl`, `sed -i` on the manifest,
`cat > tht-gate.js`, python open('w'), mv/rm/tee/…): FORBIDDEN only
covered tht subcommands and the write/edit hook only covered pi's own
tools. Read-only access (cat/grep/tail/ls) stays allowed.
Also: knownDecisionTypes is defensive — a workflow meta declaring NO
emits at all (older tht, minimal stubs) skips pre-validation instead of
rejecting every substantive type; with emits present, unknown types are
still rejected before the widget (new L1 test).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Audit findings 2.1 + 2.2 (high).
2.1 WorkflowBar's static phase list was fiction from F2 on (F2 "Schema
linking" vs memoria, F4 "SQL plan" vs schema_linking, …): every live
session showed the wrong phase name. Both maps now mirror
harness/workflow.yaml (F1 chiarimento … F8 datamart).
2.2 forceAdvance (6ee5bda) let reviewer_decide advance:true bypass the
phase gate on ANY phase, contradicting SKILL.md's "auto-advance only
empty F2 / skipped F6". reviewer_decide is back on advanceIfReady (exit-6
no-op) and tells the model to close via reviewer_confirm; forceAdvance
stays only where selection IS the approval by design: reviewer_schema_linking
(F4) and the F8 promotion close path. SKILL.md now names the three
self-closing gates (F3 rewrite_question, F4 schema-linking advance:true,
F8 memory_promote) so gate and skill state one contract.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Audit finding 1.1 (high): reviewer_memory_promote, rewrite_question,
write_schema_linking, write_cte_sql and write_final_sql still ran execute
without a top-level catch — the same unhandled-rejection class that killed
Pi in reviewer_schema_linking (fixed in 87cb806 for the four reviewer_*
tools). All 9 registered tools now share the pattern: any uncaught throw
becomes a textResult the model can react to.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The last live session crashed during reviewer_schema_linking: an uncaught
exception (likely from execFileSync in currentPhase/phaseMeta or from
cat.columns being undefined) rejected the async execute() Promise. Pi does
not catch rejected tool Promises — Node.js treats them as unhandled
rejections and kills the process.
Fix:
- Wrap all four reviewer tool execute bodies (select/decide/confirm/
schema_linking) in a top-level try/catch → returns a textResult on any
unexpected error instead of crashing Pi.
- reviewer_schema_linking: defensively re-parse `tables` if still a string
(belt-and-suspenders over prepareArguments), guard cat.columns before .map().
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Gate (tht-gate.js):
- Pre-validate decision types against workflow.yaml before showing reviewer widget
- Reject decisions emitted by later phases (min-phase check)
- Copy top-level `kind` into artifact when model forgets it (prevents loop)
- Force-advance on reviewer_decide/schema_linking when advance:true — skip
redundant reviewer_confirm gate
Backend:
- Emit agent_end on clean Pi exit (code 0 + bridge idle) instead of marking failed
Frontend:
- Strip <think> tags from transcript and activity panel
- Fix mermaid render with offscreen container + cleanup
- Graceful mermaid error: show source code instead of red error, fall back to table
Workflow:
- F2 now emits table_promoted and table_excluded (early schema linking decisions)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The model sometimes stops after receiving 'Fase approvata' without calling
`tht session finalize`, leaving the session open. Now the gate itself calls
finalize after advancing the max phase (F8), making session closure
deterministic regardless of model behavior.
- reviewer_confirm kind:phase: after phase advance at max_phase, gate calls
`tht session finalize <session>` (best-effort with recovery message)
- SKILL.md updated: model no longer needs to call finalize itself
- Tests: 2 new JS tests (auto-finalize at max phase; no-finalize at non-max)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Lever 1: Join-graph via FK logics in annotations + suggest-fks command
- TableAnnotation.foreign_keys field stores curated logical FKs (DWH has no FK constraints)
- tht schema suggest-fks: mine from approved SQL, heuristics (time_key → dim_time),
same-name discovery + explicit --assume flag for multi-owner PKs
- mschema renders 【Foreign keys】 section populated; validation in merge.py
- SKILL.md F4 now reads FKs from mschema-text, no custom data_time_key logic
Lever 2: Context-pack consolidation at kickoff (tht search pack)
- Single embedding of question, reused for schema + evidence + solved searches
- One command: tht search pack <question> --session <id> → retrieval_pack.md
- Graceful degradation when Ollama/vector store unreachable (exit 0, empty sections)
- SKILL.md F1 prescribes as first call; reduces model thinking turns via pre-retrieval
Lever 3: Phase-summary recap v2 auto-construction from session ledger
- tht session show --json includes full decisions ledger
- tht phase meta --json exports 'emits' (substantive decision types per phase)
- Gate appends deterministic 【Decisioni registrate in questa fase】 section (appendLedgerSection)
- Model authors only summary + checks; recap table comes from persisted state (exact by construction)
- SKILL.md Disciplina 6: brief model output, gate fills the rest
Tests: 358 Python (including 10 FK + 3 pack + 1 session-ledger tests) + 111 JS gate tests, all pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Transcript analysis (session 2026-07-06-175012, GLM 5.2) showed F1 at 567s:
182s wasted on a useless `tht schema introspect` (re-introspecting the remote
DWH although physical.yaml was already materialized) plus ~220s of model
thinking inflated by ~7 exploratory turns (--help/find/cat). The actual
searches cost ~15s; reviewer gates (~145s, untouched) are the quality contract.
- schema_cmd.py: introspect now exits 0 with "OK (cache)" in ~1s when
physical.yaml exists; --refresh forces the real re-introspection.
Deterministic cross-model guarantee, verified live on psd (163 tables, 1.2s).
- SKILL.md: F1 toolbox (only `tht search find` + `tht schema render`; no
introspect/--help/filesystem browsing; batch all searches in one turn);
F4 step 1 is render-only with a one-shot introspect fallback.
- tht-gate.js: `tht schema introspect ... --refresh` added to FORBIDDEN
(maintenance stays shell-only, never in-session).
- tests: 4 new pytest cases (cache hit placement proven with fake credentials,
refresh bypass, corrupt-catalog fall-through, render fallback message) and
2 gate anti-bypass JS cases.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The tht-gate.js reviewer_confirm now builds structured v2 artifacts before the
widget so the reviewer approves gate-derived data, not raw model text:
- new pure modules gate/artifact-contracts.js (soft validators, {ok,errors},
legacy-passthrough) and gate/enrich.js (index/description enrichment,
buildCteResultV2 fusing thin model data with `tht cte info`, phase enrichment)
- cte_plan v2: validate + enrich + persist via `tht cte plan --name … --doc -`
(names derived from data.ctes[]); legacy `names` param kept as fallback
- cte_result v2: rebuild from `tht cte next`/`tht cte info` (sql + preview from
the persisted test record); null/error last_test -> actionable textResult
- phase v2: soft-validate + fill phase from meta + catalog descriptions
- prepareReviewerArguments coerces artifact.data too (GLM double-stringify);
legacy markdown strings pass through unchanged
- SKILL.md: Phase 6 cte_plan payload A + thin cte_result guidance; Discipline 6
payload C example; Discipline 7 reworded for gate-rebuilt cte_result
Legacy (non-v2) paths unchanged. TypeBox stays Type.Any() for artifact.data;
validation is soft (textResult) so models self-correct instead of looping.
All 102 gate JS tests green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The "respond with gate widgets; use '!' for free text" notice is terminal-
flavored. Like the reLoop Esc guard, it used ctx.hasUI, which on pi >=0.80 is
true in RPC too, so it leaked to the browser. Now guarded on ctx.mode === "tui".
The plain text is still swallowed by the `handled` return in all modes; only the
warning is dropped in RPC. No functional ctx.hasUI guard remains in the gate.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
reLoop's "Esc non chiude il gate..." warning is terminal-specific. On pi 0.73
ctx.hasUI was false in RPC, so `if (ctx.hasUI)` effectively meant "TUI only". On
pi >=0.80 hasUI is true in RPC too (dialog-capable UI via the bridge), so the
notice leaked to the browser on any invalid gate response. Guard on
ctx.mode === "tui" to restore the original intent. The fake pi runtime gains
mode:"tui" so the roundtrip test still exercises the notice.
Audit of the other 5 ctx.ui.notify: left as-is. They are valid in both live
modes (TUI and RPC, both hasUI=true) and the gate cannot run headless
(emitAndWait needs a UI), so a guard would be dead code. Their dual-mode
rationalization belongs with the future present() work (tracked in the eval doc).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Some models (GLM 5.2) send object params (reviewer_confirm's artifact,
write_schema_linking's schema_linking) as JSON-encoded strings. These failed
TypeBox validation before execute(), making the model retry in an unbounded loop
(observed ~2100s hang). jsonObjectOrSelf() coerces them back to objects in
prepareReviewerArguments and write_schema_linking.
Also close an anti-bypass gap: pi's write/edit tools were unrestricted on the gate
code, so a looping model actually patched tht-gate.js. GATE_CODE_FILES now blocks
any write under .pi/extensions. Requires a pi restart to take effect.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- WorkflowBar: center the phase pills (add justify-center).
- tht-gate.js: the gate was sending the descriptive phase name (e.g.
"chiarimento") instead of the workflow.yaml id ("F1") in the widget's
`phase` field, so the frontend's F1..F8 match never hit and pills never
lit up. Rename phaseName -> phaseId, return p.id.
- CentralStatus: expand the live "working" tail from a single truncated
line to up to 5 lines (line-clamp-5 safety net for unbroken paragraphs).
- ModelActivityPanel: render the streamed tail through react-markdown +
remark-gfm instead of raw per-line <p> tags, so emphasis/headings render
and blank-line paragraph breaks (previously stripped) are preserved;
tail-cut now operates on paragraphs instead of physical lines.
Updated tests accordingly (WorkflowBar/CentralStatus/ModelActivityPanel/
AppShell.session-mgmt); full frontend suite (98/98) + harness gate node
tests (34/34) + tsc -b pass.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
tht() gains an optional stdin arg; the tool pipes the schema-linking
object to 'tht session set-schema-linking --file -', which validates
against SchemaLinking and returns the exact error on failure — so the
model stops hand-writing the artifact and validating with ad-hoc python.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
reviewer_confirm kind:'cte_result' registered cte_approved --subject
phase:6, which decision_cmd rejects (exit 5) and next_cte never
recognizes — dead-ending F6. Derive the CTE name from 'tht cte next'
(plan-order single source of truth) and approve by name. sql path
(sql_approved:phase:N) unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A1 — SessionMenu gains a Resume item, gated to status!=="finalized" && !archived
(matching the backend's 409 read-only guard), wired in AppShell to doResume ->
POST /sessions/:id/resume. SessionMenu.test.tsx (3 tests); frontend 96/96, tsc clean.
A2 — diagnosis-first clean-room repro driving `pi --mode rpc` with the backend's
exact resume handshake shows the cold-start stall NO LONGER reproduces on pi
0.79.4 (8/8 chained into `tht session show` + `read SKILL.md` in-turn, fresh and
partway sessions). The earlier narrate-and-stop predates the pi upgrade.
Defense-in-depth anyway: RIPRENDI_KICKOFF hardened to force the in-turn tool call
(gate_resume_kickoff.test.js + live regression 2/2). Gate JS 34/34.
PROJECT_STATE open-item #1 (resume stall) flipped to RESOLVED; cross-model resume
robustness folded into workstream G.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
F — reviewer_select options may now carry a `decision` payload {type, subject,
detail?, rationale?} plus an optional `advance`. Picking such an option IS the
confirmation: the gate persists it directly (tht decision add) and optionally
advances, with no redundant reviewer_decide/reviewer_confirm follow-up gate.
Options without a payload stay ask-only; back/exit/Other never persist.
Pure logic extracted + exported for unit tests: resolveSelectOutcome (classifies
the response) and decisionAddArgs (shared with reviewer_decide, DRY). Gate JS
suite 33/33 (gate_select_decision.test.js, +5); harness pytest 269 unchanged.
Contract docs updated together: reviewer_select tool description, SKILL.md
(widget summary, disciplines 2-3, Phase-1 single-pick), and the CLAUDE.md gate
note. Live verification (model truly emits reviewer_select+decision, decision in
review_decisions.jsonl, no follow-up gate) deferred to workstream G — it is
model-behavior-dependent.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The frontend's reviewer widgets uniformly send the picked option in a `choices`
array (SelectWidget/ArtifactGateWidget: `choices: [optionId]`), but the gate's
reviewer_select and reviewer_confirm(reject) handlers read `resp.choice`
(singular). Result: every single-select gate saw an undefined choice, answered
"Nessuna scelta ricevuta", and re-presented forever — the workflow could never
pass F1. (reviewer_decide/multiselect already read `resp.choices`, so it worked.)
Add a shared selectedChoice(resp) helper reading choices[0] (falling back to the
legacy singular choice); both handlers use it.
TDD: gate/__tests__/gate_choice.test.js RED->GREEN; full gate suite 28/28.
Verified LIVE (Playwright -> real Pi -> GLM 5.2): a single-select F1 answer is
now accepted and the workflow advances (clarification 2/4 -> 3/4). The same run
also live-verified the F1 hang fix (418187a).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- loadEnvFromDotenv: read ctx.cwd/.env into the process environment
- prepareReviewerArguments: normalize/parse reviewer tool inputs
- tidy reserved-option filtering and tht command argument assembly
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
When THT_SESSION env var is set, /nuova-domanda injects a kickoff that
tells the model to use the pre-created session id instead of running
`tht session new`. /riprendi-sessione and no-env-var paths unchanged.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Removed event.source === "interactive" guard from entry detection so
/nuova-domanda and /riprendi-sessione activate lockActive+pendingKickoff
regardless of source (TUI or RPC prompt).
- Removed event.source !== "interactive" from free-input filter; lock now
blocks/steers all user input when active, not only interactive keystrokes.
- Added typebox@1.1.38 devDep + fake_pi_runtime.registerCommand (gap from Task 2).
- New test: gate_entry.test.js (2 tests: lock activates on RPC; !-steer passes).
- Full suite: 19/19 pass, zero regressions.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Implementazione del piano di remediation progressiva sui difetti emersi
dall'analisi dell'harness. Tutto verificato: 214 test Python (incl. L0 su
Postgres reale), 14 test JS del gate, ruff pulito.
Blocco 1 (CRITICA, integrazione gate↔CLI):
- phase advance: gate usa --auto + exit 6; reviewer_confirm kind:phase fa
advance esplicito che applica i prerequisiti (prima non avanzava per le
fasi a conferma umana).
- cte plan riceve i --name dal gate (param names); set-question con id
posizionale; skill `tht search find`; nuovo comando `tht memory save-one`
con dedup hash client-side in save_one_memory.
Blocco 2 (D15, stato post-rollback):
- campo `phase` su DecisionRecord + effective_decisions phase-aware per i
subject "a nome" (cte_approved ecc.); _compute_promotions e finalize sulla
vista effective; finalize confronta col piano CTE effettivo, non glob;
`decision add --retracts` + comando `decision retract`.
Blocco 3 (D7 read-only + D6 manifest):
- assert_read_only su tutti e quattro i codepath (direct + REST);
- manifest author/summary/updated_at/updated_by/schema_version popolati +
helper touch_manifest sulle mutazioni.
Blocco 4-5 (D14a/D14b):
- decision_min_phase data-driven via `emits:` in workflow.yaml;
- formula evidence: status auto, search_formulas, gruppo CLI `tht formula`,
`search find --kind formula`, load_evidence_dir salta i .sql.md.
Blocco 6 (robustezza):
- taskdoc slice promoted_tables + bound enforced; report escaping/bound +
rsplit note; filtro kind reader REST/direct; conteggio upserted robusto;
guard REST run_query non-list; LSH disallineato -> LshIndexError.
Blocco 7 (pulizia):
- dead code gate e KIND_TO_TABLE morto rimossi; doc Postgres-only
(README + connection.py).
Blocco 0 (parziale): test di compatibilità firma gate↔CLI
(tests/integration). Rinviati: fake-Pi runtime completo, artifact-gate da
disco (#23), parità eligibility REST/direct (#28), unificazione
reserved-labels (#30), memory_rejected da deselezione (#33).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>