71 lines
3.7 KiB
Markdown
71 lines
3.7 KiB
Markdown
# Workflow UI Regressions Design
|
|
|
|
## Goal
|
|
|
|
Restore visible model reasoning and make the F4/F6 review experience deterministic,
|
|
readable, and resilient to malformed model-authored artifacts.
|
|
|
|
## Diagnosed causes
|
|
|
|
- Session creation and resume force Pi `thinking` to `off`, regardless of the saved setting.
|
|
- The backend bridge forwards assistant `text_delta` events but drops `thinking_delta` events.
|
|
- The Model Activity panel reads final assistant text rather than a dedicated activity stream.
|
|
- F4 joins use the editable `multiselect` contract; `recommended` is not translated into
|
|
`selected` outside F2, so every join initially appears unchecked.
|
|
- The CTE result viewer passes one SQL block to a viewer that always displays its
|
|
multi-block Horizontal/Vertical control.
|
|
- CTE plan cards flatten long filters, table names, keys, and rationale into loosely spaced text.
|
|
- The F6 phase-summary payload accepted an object in `open_questions`; React then attempted to
|
|
render that object directly and the widget error boundary displayed the generic failure message.
|
|
- The F3 auto-approval fix was present in the image tag but not in the running containers because
|
|
they were built without being recreated.
|
|
|
|
## Design
|
|
|
|
### Model activity
|
|
|
|
Session creation uses the global `thinking` preference and resume uses the persisted manifest
|
|
preference, falling back to the current global setting. `SessionBridge` maps Pi's nested
|
|
`thinking_delta` into a distinct SSE `activity_delta`. The frontend stores that stream separately
|
|
from final assistant text, and Model Activity renders only activity deltas. This prevents internal
|
|
reasoning from leaking into transcript-oriented state while making the panel accurately reflect
|
|
the configured model activity.
|
|
|
|
### Join review
|
|
|
|
F4 calls to `reviewer_decide` whose merit decisions are all `join_modified` become a read-only
|
|
`join-review` widget. Each proposed join is rendered as an informational card with its name,
|
|
join expression, and rationale. `Continue` returns every join id; the gate accepts only that exact
|
|
complete response and persists all decisions through one atomic ledger replacement serialized by
|
|
a per-session cross-process lock. Malformed or partial responses re-present the widget. `Other — specify` returns textual feedback without
|
|
persisting the current proposal, so the model must revise and present the complete join set again.
|
|
Mixed join/non-join calls remain regular
|
|
multiselects, and the skill instructs the model to keep joins in a separate call.
|
|
|
|
### CTE presentation
|
|
|
|
The plan viewer uses a responsive card layout: a compact numbered header, purpose and rationale
|
|
as readable prose, aligned metadata rows for dependencies and keys, wrapped table rows, and a
|
|
structured filter grid with separate column/operator/value fields. Output columns remain compact
|
|
wrapping chips. The SQL layout switch is hidden whenever the viewer receives fewer than two blocks.
|
|
|
|
### Phase-summary resilience
|
|
|
|
The gate validates that `open_questions` is an array of strings and reports a corrective error to
|
|
the model before emitting a widget. The frontend also normalizes legacy malformed entries to a
|
|
safe string (preferring `question`, then `label`) so old or externally produced payloads cannot
|
|
crash React. The session skill documents the exact `string[]` contract.
|
|
|
|
### Deployment
|
|
|
|
After targeted and full regression tests, build both Docker images and recreate the Compose
|
|
services. Verify that the running containers use the newly built image ids and that both health
|
|
checks pass.
|
|
|
|
## Non-goals
|
|
|
|
- Persisting verbatim model reasoning in session artifacts.
|
|
- Allowing individual joins to be removed from the join review.
|
|
- Redesigning multi-block SQL comparison behavior.
|
|
- Changing the eight-phase workflow or F3 auto-approval semantics.
|