156 Commits
Author SHA1 Message Date
marcopanandClaude Opus 4.8 cef9ae4368 feat(harness): single-select answers auto-confirm (reviewer_select persists)
F — reviewer_select options may now carry a `decision` payload {type, subject,
detail?, rationale?} plus an optional `advance`. Picking such an option IS the
confirmation: the gate persists it directly (tht decision add) and optionally
advances, with no redundant reviewer_decide/reviewer_confirm follow-up gate.
Options without a payload stay ask-only; back/exit/Other never persist.

Pure logic extracted + exported for unit tests: resolveSelectOutcome (classifies
the response) and decisionAddArgs (shared with reviewer_decide, DRY). Gate JS
suite 33/33 (gate_select_decision.test.js, +5); harness pytest 269 unchanged.

Contract docs updated together: reviewer_select tool description, SKILL.md
(widget summary, disciplines 2-3, Phase-1 single-pick), and the CLAUDE.md gate
note. Live verification (model truly emits reviewer_select+decision, decision in
review_decisions.jsonl, no follow-up gate) deferred to workstream G — it is
model-behavior-dependent.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 18:01:50 +02:00
marcopanandClaude Opus 4.8 b056ff334f feat(frontend): phase progress dots (B) + compact sidebar redesign (C)
B — WorkflowBar renders F1..F8 as colored ring-dots (no phase-name text):
green=done, amber=running (subtle pulse), red=error, gray=pending; green
connectors lead the active dot, each dot carries data-state. Lightweight
error signal: sessionStore gains `phaseError`, set when an info event has
level=error during the active phase, cleared on the next ui_request.

C — denser single-line session rows (inline status dot + name, py-1), a
3-level type hierarchy (L1 SESSIONS / L2 section+group headers / L3 names),
and the "No group" label removed (ungrouped render after the last group,
guarded so the empty-state still teaches when there are zero groups).

Live-verified with Playwright (all four dot states, sidebar hierarchy, and
E's deferred activity-panel check). Frontend 93/93, tsc -b clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 17:53:29 +02:00
marcopanandClaude Opus 4.8 2e09c88f78 docs(state): record UI-redesign progress (D+E done; B/C/F/A pending)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 16:46:50 +02:00
marcopanandClaude Opus 4.8 d27afd0896 fix(gate): read reviewer_select/confirm choice from the choices[] array
The frontend's reviewer widgets uniformly send the picked option in a `choices`
array (SelectWidget/ArtifactGateWidget: `choices: [optionId]`), but the gate's
reviewer_select and reviewer_confirm(reject) handlers read `resp.choice`
(singular). Result: every single-select gate saw an undefined choice, answered
"Nessuna scelta ricevuta", and re-presented forever — the workflow could never
pass F1. (reviewer_decide/multiselect already read `resp.choices`, so it worked.)

Add a shared selectedChoice(resp) helper reading choices[0] (falling back to the
legacy singular choice); both handlers use it.

TDD: gate/__tests__/gate_choice.test.js RED->GREEN; full gate suite 28/28.
Verified LIVE (Playwright -> real Pi -> GLM 5.2): a single-select F1 answer is
now accepted and the workflow advances (clarification 2/4 -> 3/4). The same run
also live-verified the F1 hang fix (418187a).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 12:48:56 +02:00
marcopanandClaude Opus 4.8 418187a4ad fix(backend): echo Pi's RPC id so reviewer gates unblock after answer
ctx.ui.input in `pi --mode rpc` correlates extension_ui_response on its own
top-level RPC id (crypto.randomUUID), not the descriptor id the gate carries
in `title`. SessionBridge replied with the descriptor id, so Pi silently
dropped the response and the model never resumed — every reviewer widget hung
after the human answered.

SessionBridge now stores Pi's top-level m.id (pendingPiId) and replies
extension_ui_response{ id: pendingPiId, value: <uiResponse> }; value still
carries the descriptor id so the gate's internal resp.id === descriptor.id
check still holds.

The fake-pi double had masked the bug by forcing m.id == descriptor.id; it now
mirrors real Pi (distinct randomUUID, correlate on it, drop unknown ids), with
a negative regression test. SKILL.md Phase 1 also now steers multi-answer
disambiguation to reviewer_decide (multiselect).

Tests: backend 67/67, tsc clean, fake-pi contract 2/2.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 10:43:49 +02:00
marcopanandClaude Opus 4.8 d258a1ae4c docs: add PROJECT_STATE.md orientation snapshot
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 14:19:59 +02:00