The "respond with gate widgets; use '!' for free text" notice is terminal-
flavored. Like the reLoop Esc guard, it used ctx.hasUI, which on pi >=0.80 is
true in RPC too, so it leaked to the browser. Now guarded on ctx.mode === "tui".
The plain text is still swallowed by the `handled` return in all modes; only the
warning is dropped in RPC. No functional ctx.hasUI guard remains in the gate.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Fix six GUI defects, verified live against the real stack:
- NavSessions: mark the in-progress (active) session in yellow (warning
token) — both its status dot and its row background.
- SteerInput: the composer is now an auto-growing textarea that wraps and
grows vertically (caps at 160px, then scrolls); Enter sends, Shift+Enter
inserts a newline.
- WorkflowBar: show a synthetic title of the current phase under the 8 dots
from static EN/IT strings (no LLM); English is displayed to match the chrome.
- CentralStatus: reformat the 5-line system-message tail as a structured,
monospace list with per-line markers and an emphasized last line.
- AppShell + CentralStatus: move the left Model-activity panel toggle off the
working spinner (now a pure status indicator) onto a dedicated arrow button
(→ opens, ← closes).
- gate: rename the phase-confirm button "Conferma e prosegui" → "Salva e
procedi" (builders.js + its L1 test).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
reLoop's "Esc non chiude il gate..." warning is terminal-specific. On pi 0.73
ctx.hasUI was false in RPC, so `if (ctx.hasUI)` effectively meant "TUI only". On
pi >=0.80 hasUI is true in RPC too (dialog-capable UI via the bridge), so the
notice leaked to the browser on any invalid gate response. Guard on
ctx.mode === "tui" to restore the original intent. The fake pi runtime gains
mode:"tui" so the roundtrip test still exercises the notice.
Audit of the other 5 ctx.ui.notify: left as-is. They are valid in both live
modes (TUI and RPC, both hasUI=true) and the gate cannot run headless
(emitAndWait needs a UI), so a guard would be dead code. Their dual-mode
rationalization belongs with the future present() work (tracked in the eval doc).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- aritmolab-provider.js: import pi-ai from @earendil-works (was @mariozechner).
Verified live that the aritmolab/qwen provider still registers under pi 0.80.3
(get_available_models -> aritmolab, deepseek, zai). Removes the dual-package
reliance on the frozen @mariozechner install still on disk for rollback.
- docs/general/pi-configuration.md: built-in models package is @earendil-works/pi-ai
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Some models (GLM 5.2) send object params (reviewer_confirm's artifact,
write_schema_linking's schema_linking) as JSON-encoded strings. These failed
TypeBox validation before execute(), making the model retry in an unbounded loop
(observed ~2100s hang). jsonObjectOrSelf() coerces them back to objects in
prepareReviewerArguments and write_schema_linking.
Also close an anti-bypass gap: pi's write/edit tools were unrestricted on the gate
code, so a looping model actually patched tht-gate.js. GATE_CODE_FILES now blocks
any write under .pi/extensions. Requires a pi restart to take effect.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
buildSelectRequest injected an `altroOption` ({id:"other", opens:freetext}) into the
select descriptor, but SelectWidget/MultiselectWidget never route option.opens — so it
rendered as an inert "Altro — specifica…" button next to the working reserved
"Other — specify" control (which now owns free-text since the reviewer-gate-ux fix).
Remove the injection, the now-unused altroOption()/ALTRO_LABEL, and the unused
allow_other flag from both builders (no frontend consumer). Free-text stays offered on
every gate via the reserved "Other" control. The ArtifactGateWidget opens/LinkageHost
linkage (reject-with-reason capability) is intentionally kept — it is a separate,
tested feature, not the injected duplicate.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- WorkflowBar: center the phase pills (add justify-center).
- tht-gate.js: the gate was sending the descriptive phase name (e.g.
"chiarimento") instead of the workflow.yaml id ("F1") in the widget's
`phase` field, so the frontend's F1..F8 match never hit and pills never
lit up. Rename phaseName -> phaseId, return p.id.
- CentralStatus: expand the live "working" tail from a single truncated
line to up to 5 lines (line-clamp-5 safety net for unbroken paragraphs).
- ModelActivityPanel: render the streamed tail through react-markdown +
remark-gfm instead of raw per-line <p> tags, so emphasis/headings render
and blank-line paragraph breaks (previously stripped) are preserved;
tail-cut now operates on paragraphs instead of physical lines.
Updated tests accordingly (WorkflowBar/CentralStatus/ModelActivityPanel/
AppShell.session-mgmt); full frontend suite (98/98) + harness gate node
tests (34/34) + tsc -b pass.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Move the AritmoLab provider (Qwen) into harness/.pi/extensions/ so it is
auto-discovered only when Pi runs with cwd=harness, keeping it project-local.
GLM (~/.pi/agent/models.json) and DeepSeek (built-in) stay user-wide. File is
.js (not .mjs) because Pi extension auto-discovery matches only /\.(ts|js)$/.
Removes the gemma4-26b-a4b model (404 at the endpoint) from both the provider
and the model-matrix default list.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The cheat-sheet, Discipline 2, and Phase 7 said F7 closes with
reviewer_confirm kind:"sql" — but the gate's kind:"sql" only records
sql_approved:phase:7; F7 is not in _AUTO_ADVANCE_PHASES, so advancing to
F8 still needs reviewer_confirm kind:"phase" (mirrors F6). Same
record-vs-advance trap this branch removes elsewhere.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
F3/F4 close with reviewer_confirm kind:'phase' (they do NOT auto-advance);
advance:true only auto-advances F2-empty/F6-skip; rewrite_question belongs
to F3 not F1; F4 uses the new write_schema_linking tool with the documented
SchemaLinking shape. Adds a per-phase artifact/close cheat-sheet.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
tht() gains an optional stdin arg; the tool pipes the schema-linking
object to 'tht session set-schema-linking --file -', which validates
against SchemaLinking and returns the exact error on failure — so the
model stops hand-writing the artifact and validating with ad-hoc python.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
reviewer_confirm kind:'cte_result' registered cte_approved --subject
phase:6, which decision_cmd rejects (exit 5) and next_cte never
recognizes — dead-ending F6. Derive the CTE name from 'tht cte next'
(plan-order single source of truth) and approve by name. sql path
(sql_approved:phase:N) unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Tier 1 (clean-room first-turn harness, harness/scripts/model-matrix.mjs):
kickoff + resume chain in-turn on ALL available models — zai/glm-5.2,
deepseek/deepseek-v4-{pro,flash}, aritmolab/qwen3.6-35b-a3b, zai/glm-4.5-air.
The resume cold-start stall recurs on none (closes A's cross-model robustness).
aritmolab/gemma4-26b-a4b is a 404 at the endpoint (listed but not served) — an
availability gap classified as MODEL_ERROR, not a workflow issue.
Tier 2 (live, baseline zai/glm-5.2): F single-select auto-confirm verified
end-to-end — answering the first reviewer_select persisted a concept_clarified
decision (review_decisions.jsonl 0->1) with no follow-up confirmation gate.
Closes F's deferred live check.
No prompt hardening needed. Results in the G plan doc + memory. Throwaway psd
sessions used and deleted; real sessions untouched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A1 — SessionMenu gains a Resume item, gated to status!=="finalized" && !archived
(matching the backend's 409 read-only guard), wired in AppShell to doResume ->
POST /sessions/:id/resume. SessionMenu.test.tsx (3 tests); frontend 96/96, tsc clean.
A2 — diagnosis-first clean-room repro driving `pi --mode rpc` with the backend's
exact resume handshake shows the cold-start stall NO LONGER reproduces on pi
0.79.4 (8/8 chained into `tht session show` + `read SKILL.md` in-turn, fresh and
partway sessions). The earlier narrate-and-stop predates the pi upgrade.
Defense-in-depth anyway: RIPRENDI_KICKOFF hardened to force the in-turn tool call
(gate_resume_kickoff.test.js + live regression 2/2). Gate JS 34/34.
PROJECT_STATE open-item #1 (resume stall) flipped to RESOLVED; cross-model resume
robustness folded into workstream G.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
F — reviewer_select options may now carry a `decision` payload {type, subject,
detail?, rationale?} plus an optional `advance`. Picking such an option IS the
confirmation: the gate persists it directly (tht decision add) and optionally
advances, with no redundant reviewer_decide/reviewer_confirm follow-up gate.
Options without a payload stay ask-only; back/exit/Other never persist.
Pure logic extracted + exported for unit tests: resolveSelectOutcome (classifies
the response) and decisionAddArgs (shared with reviewer_decide, DRY). Gate JS
suite 33/33 (gate_select_decision.test.js, +5); harness pytest 269 unchanged.
Contract docs updated together: reviewer_select tool description, SKILL.md
(widget summary, disciplines 2-3, Phase-1 single-pick), and the CLAUDE.md gate
note. Live verification (model truly emits reviewer_select+decision, decision in
review_decisions.jsonl, no follow-up gate) deferred to workstream G — it is
model-behavior-dependent.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
New sessions get a concise Italian-keyword `name` instead of the truncated
question. `tht session new` (when no --name is given) derives it via a new
`_extract_name` helper using YAKE (pure-Python, unsupervised, Italian, no LLM),
dropping generic query verbs and keeping the top keywords in reading order;
falls back to `_summarize` if YAKE is unavailable. `create_session` core keeps
its `name=None` default — the policy lives at the CLI layer.
TDD: tests/test_session_name.py (unit + CliRunner integration). Full harness
suite 269 passed; ruff clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The frontend's reviewer widgets uniformly send the picked option in a `choices`
array (SelectWidget/ArtifactGateWidget: `choices: [optionId]`), but the gate's
reviewer_select and reviewer_confirm(reject) handlers read `resp.choice`
(singular). Result: every single-select gate saw an undefined choice, answered
"Nessuna scelta ricevuta", and re-presented forever — the workflow could never
pass F1. (reviewer_decide/multiselect already read `resp.choices`, so it worked.)
Add a shared selectedChoice(resp) helper reading choices[0] (falling back to the
legacy singular choice); both handlers use it.
TDD: gate/__tests__/gate_choice.test.js RED->GREEN; full gate suite 28/28.
Verified LIVE (Playwright -> real Pi -> GLM 5.2): a single-select F1 answer is
now accepted and the workflow advances (clarification 2/4 -> 3/4). The same run
also live-verified the F1 hang fix (418187a).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
ctx.ui.input in `pi --mode rpc` correlates extension_ui_response on its own
top-level RPC id (crypto.randomUUID), not the descriptor id the gate carries
in `title`. SessionBridge replied with the descriptor id, so Pi silently
dropped the response and the model never resumed — every reviewer widget hung
after the human answered.
SessionBridge now stores Pi's top-level m.id (pendingPiId) and replies
extension_ui_response{ id: pendingPiId, value: <uiResponse> }; value still
carries the descriptor id so the gate's internal resp.id === descriptor.id
check still holds.
The fake-pi double had masked the bug by forcing m.id == descriptor.id; it now
mirrors real Pi (distinct randomUUID, correlate on it, drop unknown ids), with
a negative regression test. SKILL.md Phase 1 also now steers multi-answer
disambiguation to reviewer_decide (multiselect).
Tests: backend 67/67, tsc clean, fake-pi contract 2/2.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- loadEnvFromDotenv: read ctx.cwd/.env into the process environment
- prepareReviewerArguments: normalize/parse reviewer tool inputs
- tidy reserved-option filtering and tht command argument assembly
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
workspace/provider/model/thinking now come from getSettings() injected into
sessionRoutes; the request body supplies only question+name. Also teaches
fake_pi_rpc to respond to set_model and set_thinking_level RPC commands so
tests that pass real model settings don't hang.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Harness:
- preview_cmd FILE positional arg made optional; when omitted with --session,
path is derived via _session_sql_file (mirrors export_cmd) — fixes the
deferred Task-5 bug where the backend passed sessions/<id>/sql_final.sql
relative to harnessDir, which broke for workspace-dependent paths.
- New pytest: test_preview_session_no_file_resolves_sql_final
Backend:
- ThtRunner.sqlPreview: drop positional file arg; use --session only
- New routes/sql.ts: POST /sessions/:id/sql/preview + /export
- New routes/meta.ts: GET /workspaces (yaml scan) + GET /models (injectable
seam + graceful fallback to {models:[]})
- app.ts: register sqlRoutes + metaRoutes; add listModels to BuildAppDeps
- tht-runner.test.ts: add sqlPreview argv assertion (no file path)
- test/routes-sql-meta.test.ts: 9 tests (sql preview/export + meta routes)
Tests: harness 233 passed; backend 29 passed; build clean.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>