From 2da2df58ca587c066eeb1f07cd35c247f1d09fe1 Mon Sep 17 00:00:00 2001 From: User Date: Tue, 14 Jul 2026 22:33:30 +0200 Subject: [PATCH] docs: design activity log and CTE layout fixes --- ...26-07-14-activity-log-cte-layout-design.md | 152 ++++++++++++++++++ 1 file changed, 152 insertions(+) create mode 100644 docs/superpowers/specs/2026-07-14-activity-log-cte-layout-design.md diff --git a/docs/superpowers/specs/2026-07-14-activity-log-cte-layout-design.md b/docs/superpowers/specs/2026-07-14-activity-log-cte-layout-design.md new file mode 100644 index 00000000..72a7c254 --- /dev/null +++ b/docs/superpowers/specs/2026-07-14-activity-log-cte-layout-design.md @@ -0,0 +1,152 @@ +# Complete Activity Log and CTE Plan Layout Design + +## Goal + +Restore the left Model activity panel as a complete chronological log for the current live turn, +starting with the user's prompt in Phase 1, while keeping tool details sanitized. Improve the +Phase 6 CTE plan form so every CTE has reliable internal padding, clear hierarchy, responsive +wrapping, and a dense but readable review layout. + +## Confirmed Root Causes + +### Empty activity panel + +Commit `c7474f3` changed `ModelActivityPanel` from the assistant transcript to the dedicated +`activity` array populated only by Pi `thinking_delta` events. That made reasoning distinct from +final assistant text, but it also made the panel empty for models that do not emit reasoning. +The live local Qwen model declares `reasoning: false`, so no `thinking_delta` is expected. + +### Missing CTE padding + +The project uses Tailwind CSS 3.4, while the shared `Card` primitive uses Tailwind 4 syntax such as +`px-(--card-spacing)` and `--spacing(4)`. The build emits an unusable custom-property value and no +working horizontal or vertical padding declarations for `CardHeader` and `CardContent`. The CTE +viewer's valid child gaps remain, but its sections sit against the outer card edges. + +## Activity Timeline Contract + +The frontend store owns one in-memory chronological activity timeline for the current live +session. It is not persisted and is cleared by the existing session reset behavior. + +Each timeline entry has a stable semantic kind and display-safe fields: + +- `prompt`: user question, gate choice, or steering text recorded locally; +- `thinking`: model reasoning chunks when the provider emits them; +- `assistant`: visible assistant text chunks; +- `tool`: sanitized tool lifecycle with tool name, correlation id, and + `running | completed | failed` status; +- `gate`: reviewer request and phase metadata; +- `status`: backend informational, warning, or sanitized error text; +- `lifecycle`: model turn start/end markers. + +The current phase is captured on each entry when known. The new-question submit path sets Phase 1 +and records the question before awaiting session creation, so the first timeline row is always +attributable to F1. Gate choices and steering text are recorded only after their requests succeed, +preserving the existing retry semantics. Resume starts a fresh timeline at the resume action +because historical chat/activity is not persisted. + +Streaming `thinking` and `assistant` chunks coalesce only with the immediately preceding entry of +the same kind. Tool completion updates the matching running entry by correlation id instead of +creating a disconnected duplicate. All other events append in arrival order. + +## Backend Event Mapping and Sanitization + +`SessionBridge` continues to map reasoning and assistant text separately. It additionally maps Pi +`tool_execution_start` and `tool_execution_end` to a dedicated client activity event containing +only: + +- tool call id; +- tool name; +- lifecycle status. + +Tool arguments, partial results, final results, command output, request bodies, credentials, paths +derived from results, and raw exception text must never be forwarded. Existing sanitized provider +error behavior remains unchanged. Tool update events are ignored because their payload is raw +partial output and the running row already communicates progress. + +The existing `ui_request`, `info`, `agent_start`, and `agent_end` events remain authoritative and +are also projected into timeline rows by the frontend store. + +## Model Activity Panel + +The panel renders the full timeline rather than a reasoning-only tail. Opening or closing it does +not alter store state. It uses the existing scrollable panel and automatically follows the newest +entry while the user remains near the bottom; manual upward scrolling is not overridden. + +Rows use restrained product-UI styling: + +- phase and kind labels form the compact metadata layer; +- prompt and assistant rows use normal body text; +- thinking text remains visually secondary; +- tool rows expose name and status without expandable raw details; +- warnings and errors use existing semantic colors. + +The central status component keeps its compact assistant-text tail and is not replaced by the full +timeline. + +## CTE Plan Layout + +### Card primitive compatibility + +Replace the Tailwind 4-only spacing syntax in the shared `Card` primitive with Tailwind 3-compatible +utilities. Default cards use the existing 4-point scale at 16 px; the small variant uses 12 px. +Header, content, footer, parent gap, and vertical padding must all compile to concrete declarations. +The CTE viewer is currently the only consumer, limiting regression scope. + +### Information hierarchy + +Question, strategy, and execution order form one compact overview group. Labels identify their +roles without decorative color. The overview uses tight 8–12 px internal spacing and a 24 px gap +before the ordered CTE sequence. + +Each CTE remains one ordered outer container because the boundary has semantic value. It uses a +border without an additional decorative shadow inside the already elevated gate. The ordinal badge +and CTE name lead the header; the name is a semantic heading. Purpose remains subordinate. + +### Internal rhythm + +- Card edge padding: 16 px by default, 20 px when the available width permits it. +- Related labels, values, and chips: 4–8 px. +- Rows and closely related blocks: 12–16 px. +- Major body groups: 20–24 px. + +Tables and filters render as padded rows separated by dividers within one boundary, not as nested +shadowed cards. Rationale becomes a final plain section separated with a top divider; the prohibited +colored side stripe is removed. Primary color remains reserved for state/order emphasis. + +Long CTE names, table names, columns, filter values, and chips wrap within their containers. Filter +metadata becomes multi-column only at a width where all fields remain readable; otherwise it stays +stacked. Shared gate buttons, artifact contracts, colors, and non-CTE viewers are out of scope. + +## Accessibility and Responsive Behavior + +- The CTE name uses heading semantics in document order. +- Timeline rows include readable text labels and never rely on color alone for status. +- Tool status changes remain understandable without animation. +- The panel and CTE form preserve keyboard scrolling and existing focus behavior. +- No new interactive control is introduced by the CTE viewer. + +## Testing and Verification + +Tests are written before production changes and must prove: + +- the F1 prompt is the first timeline entry before session creation completes; +- non-reasoning models still produce prompt, assistant, tool, gate, and lifecycle rows; +- thinking/assistant chunks coalesce without losing chronological ordering; +- tool start/end update one sanitized entry and raw args/results never cross the bridge; +- opening and closing the panel preserves the timeline; +- automatic following stops while the user has scrolled away from the bottom; +- CTE cards compile/use explicit padding classes and preserve wrapping/section structure; +- filter rows, rationale divider, semantic heading, and responsive grouping render as designed. + +Run focused backend/frontend tests, full suites, TypeScript gates, production builds, the Impeccable +layout detector, and `git diff --check`. Build and force-recreate `core` and `frontend`, then verify +core health, the frontend asset update, and a live session whose Phase 1 panel shows prompt plus +subsequent sanitized activity. + +## Out of Scope + +- Persisting or replaying historical chat/activity across processes or browser reloads. +- Showing raw tool arguments, results, partial output, commands, or exception details. +- Changing model thinking configuration or provider definitions. +- Redesigning shared reviewer controls or non-CTE artifact viewers.