diff --git a/docs/superpowers/plans/2026-07-15-model-activity-signal-filter.md b/docs/superpowers/plans/2026-07-15-model-activity-signal-filter.md new file mode 100644 index 00000000..4ed35550 --- /dev/null +++ b/docs/superpowers/plans/2026-07-15-model-activity-signal-filter.md @@ -0,0 +1,495 @@ +# Model Activity Signal Filter Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Keep the left Model activity panel focused on prompt, genuine reasoning, meaningful status, and reviewer gates while hiding assistant narration, tool lifecycle, and turn lifecycle rows. + +**Architecture:** Preserve the complete chronological `activityLog` and every existing backend/SSE/store contract. Add a strict allowlist at the `ModelActivityPanel` rendering boundary, derive the visible tail for stable bottom-follow behavior, and leave the central transcript untouched. + +**Tech Stack:** React 18, TypeScript, Zustand, Vitest, Testing Library, Vite, Docker Compose, nginx-unprivileged. + +## Global Constraints + +- Visible kinds are exactly `prompt`, `thinking`, `status`, and `gate`. +- Hidden kinds are exactly `assistant`, `tool`, and `lifecycle`; unknown future kinds are hidden by default. +- Do not filter by text, language, tool name, or heuristic content matching. +- Do not change backend emission, SSE subscriptions, `sessionStore` ingestion, the central transcript, or workflow state. +- Base the empty state and bottom-follow behavior on visible entries only. +- Preserve warning/error labels, Markdown rendering, phase labels, accessibility, close behavior, and the 48 px near-bottom threshold. +- Do not touch the user-owned untracked `.vite/` directory. + +--- + +## File map + +- `frontend/src/shell/ModelActivityPanel.tsx`: owns the visibility predicate, visible projection, rendering, empty state, and bottom-follow dependency. +- `frontend/src/shell/ModelActivityPanel.test.tsx`: proves the allowlist, transcript isolation, hidden-only empty state, Markdown/status rendering, and scroll behavior. +- `brain/codebase/workflow-ui-contracts.md`: records the durable distinction between the complete internal log and its user-facing projection. +- `PROJECT_STATE.md`: records final verification, container image/asset, and Qwen smoke evidence. + +### Task 1: Filter the user-facing activity projection + +**Files:** +- Modify: `frontend/src/shell/ModelActivityPanel.test.tsx` +- Modify: `frontend/src/shell/ModelActivityPanel.tsx` + +**Interfaces:** +- Consumes: `ActivityEntry` and `ActivityKind` from `frontend/src/api/types.ts`; the unchanged `activityLog` array from `useSessionStore`. +- Produces: `isVisibleModelActivity(entry: ActivityEntry): boolean`; a panel projection containing only prompt/thinking/status/gate rows. + +- [ ] **Step 1: Write failing behavior tests before production code** + +In `frontend/src/shell/ModelActivityPanel.test.tsx`, import the activity kind type and the module namespace so the not-yet-created predicate can fail as an assertion rather than a module-load error: + +```tsx +import type { ActivityEntry, ActivityKind } from "../api/types"; +import * as activityPanelModule from "./ModelActivityPanel"; +import { isNearBottom, ModelActivityPanel } from "./ModelActivityPanel"; +``` + +Replace the existing complete-sequence test with the approved mixed-sequence contract: + +```tsx +test("renders only prompt, thinking, status, and gate from a mixed F1 sequence", () => { + const store = useSessionStore.getState(); + store.setPhase("F1"); + store.setLastUserEntry({ kind: "input", text: "How many patients?" }); + store.applyEvent({ type: "system_event", event: "agent_start" }); + store.applyEvent({ + type: "activity_event", + activity: { kind: "tool", toolCallId: "tool-1", toolName: "bash", status: "running" }, + }); + store.applyEvent({ + type: "activity_event", + activity: { kind: "tool", toolCallId: "tool-1", toolName: "bash", status: "completed" }, + }); + store.applyEvent({ type: "text_delta", text: "Let me run another bash search." }); + store.applyEvent({ type: "activity_delta", text: "Inspecting **schema**." }); + store.applyEvent({ type: "info", level: "warning", text: "Retrying schema lookup" }); + store.applyEvent({ + type: "ui_request", + ui_request: { id: "gate-1", widget: "select", phase: "F1_review", title: "Confirm cohort" }, + }); + + render(); + + expect(screen.getAllByRole("article")).toHaveLength(4); + expect(screen.getByText("How many patients?")).toBeInTheDocument(); + expect(screen.getByText("schema").tagName).toBe("STRONG"); + expect(screen.getByText("Retrying schema lookup")).toBeInTheDocument(); + expect(screen.getByText("Confirm cohort")).toBeInTheDocument(); + expect(screen.queryByText("Agent started")).not.toBeInTheDocument(); + expect(screen.queryByText("bash")).not.toBeInTheDocument(); + expect(screen.queryByText("completed")).not.toBeInTheDocument(); + expect(screen.queryByText("Let me run another bash search.")).not.toBeInTheDocument(); + expect(useSessionStore.getState().transcript).toEqual([ + { role: "assistant", text: "Let me run another bash search." }, + ]); +}); +``` + +Add direct allowlist/default-deny and hidden-only empty-state tests: + +```tsx +test("uses an explicit default-deny activity-kind allowlist", () => { + const predicate = (activityPanelModule as unknown as { + isVisibleModelActivity?: (entry: ActivityEntry) => boolean; + }).isVisibleModelActivity; + expect(predicate).toBeTypeOf("function"); + if (!predicate) return; + + for (const kind of ["prompt", "thinking", "status", "gate"] satisfies ActivityKind[]) { + expect(predicate({ kind, phase: "F1", text: kind })).toBe(true); + } + for (const kind of ["assistant", "tool", "lifecycle"] satisfies ActivityKind[]) { + expect(predicate({ kind, phase: "F1", text: kind })).toBe(false); + } + expect(predicate({ kind: "future-kind" as ActivityKind, phase: "F1", text: "future" })).toBe(false); +}); + +test("shows the empty state when the raw log contains only hidden entries", () => { + useSessionStore.setState({ + activityLog: [ + { kind: "assistant", phase: "F1", text: "Let me try another command." }, + { kind: "tool", phase: "F1", text: "bash", toolCallId: "tool-1", status: "completed" }, + { kind: "lifecycle", phase: "F1", text: "Turn end" }, + ], + }); + + render(); + + expect(screen.getByText("No activity yet.")).toBeInTheDocument(); + expect(screen.queryAllByRole("article")).toHaveLength(0); +}); +``` + +Change the assistant/thinking Markdown test so only reasoning is rendered: + +```tsx +test("renders thinking markdown without assistant narration", () => { + useSessionStore.setState({ + activityLog: [ + { kind: "thinking", phase: "F2", text: "Considering **history**." }, + { kind: "assistant", phase: "F2", text: "Use *current* records." }, + ], + }); + + render(); + + expect(screen.getByText("history").tagName).toBe("STRONG"); + expect(screen.queryByText("current")).not.toBeInTheDocument(); + expect(screen.getAllByRole("article")).toHaveLength(1); +}); +``` + +Add the hidden-update scroll regression: + +```tsx +test("does not bottom-follow when only a hidden event arrives", () => { + useSessionStore.setState({ + activityLog: [{ kind: "prompt", phase: "F1", text: "Visible prompt" }], + }); + render(); + const viewport = screen.getByTestId("activity-scroll"); + setScrollGeometry(viewport, { scrollHeight: 400, clientHeight: 100, scrollTop: 300 }); + + act(() => useSessionStore.getState().applyEvent({ + type: "activity_event", + activity: { kind: "tool", toolCallId: "tool-2", toolName: "bash", status: "running" }, + })); + + expect(viewport.scrollTop).toBe(300); +}); +``` + +Update both existing scroll tests to start with a visible prompt and append a visible status rather +than calling `recordLifecycle`, which is intentionally hidden: + +```tsx +useSessionStore.setState({ + activityLog: [{ kind: "prompt", phase: "F1", text: "First" }], +}); +// ...existing geometry and scroll setup... +act(() => useSessionStore.getState().applyEvent({ type: "info", text: "Second" })); +``` + +- [ ] **Step 2: Run the focused test and verify RED** + +Run: + +```bash +cd frontend +npx vitest run src/shell/ModelActivityPanel.test.tsx +``` + +Expected: FAIL because the predicate is absent, the panel renders seven mixed-sequence rows instead +of four, hidden-only input does not show the empty state, assistant text remains visible, and a +hidden tool append moves the bottom-follow viewport. + +- [ ] **Step 3: Implement the strict rendering-boundary filter** + +In `frontend/src/shell/ModelActivityPanel.tsx`, extend the type import and add the pure predicate: + +```tsx +import type { ActivityEntry, ActivityKind } from "../api/types"; + +const VISIBLE_MODEL_ACTIVITY_KINDS: ReadonlySet = new Set([ + "prompt", + "thinking", + "status", + "gate", +]); + +export function isVisibleModelActivity(entry: ActivityEntry): boolean { + return VISIBLE_MODEL_ACTIVITY_KINDS.has(entry.kind); +} +``` + +Inside `ModelActivityPanel`, derive the projection and a stable visible tail: + +```tsx +const activityLog = useSessionStore((s) => s.activityLog); +const visibleActivity = activityLog.filter(isVisibleModelActivity); +const visibleCount = visibleActivity.length; +const visibleTail = visibleActivity.at(-1); +const scrollRef = useRef(null); +const followRef = useRef(true); +``` + +Change the layout effect dependency so hidden appends do not run it, while a streamed visible tail +still changes object identity and follows correctly: + +```tsx +useLayoutEffect(() => { + const viewport = scrollRef.current; + if (viewport && followRef.current) viewport.scrollTop = viewport.scrollHeight; +}, [visibleCount, visibleTail]); +``` + +Finally, use `visibleActivity` for both empty-state and row rendering: + +```tsx +{visibleActivity.length === 0 ? ( +

No activity yet.

+) : ( + visibleActivity.map((entry, index) => ( + + )) +)} +``` + +Do not modify `sessionStore.ts`, `useSessionStream.ts`, backend files, or `ActivityEntry`. + +- [ ] **Step 4: Run focused GREEN and adjacent store regressions** + +Run: + +```bash +cd frontend +npx vitest run src/shell/ModelActivityPanel.test.tsx src/store/sessionStore.test.ts +npx tsc -b +``` + +Expected: both test files pass, the complete store fold still retains all kinds, and TypeScript +exits 0. + +- [ ] **Step 5: Inspect scope and commit Task 1** + +Run: + +```bash +git diff --check +git diff -- frontend/src/shell/ModelActivityPanel.tsx frontend/src/shell/ModelActivityPanel.test.tsx +git status --short +git add frontend/src/shell/ModelActivityPanel.tsx frontend/src/shell/ModelActivityPanel.test.tsx +git commit -m "fix(frontend): hide model activity execution noise" +``` + +Expected: only the panel and its test are committed; `.vite/` remains untracked and untouched. + +### Task 2: Verify, document, deploy, and smoke the filtered panel + +**Files:** +- Modify: `brain/codebase/workflow-ui-contracts.md` +- Modify: `PROJECT_STATE.md` + +**Interfaces:** +- Consumes: Task 1's `isVisibleModelActivity` contract and production frontend bundle. +- Produces: deployed frontend image/entry evidence and a durable record of the complete-log versus visible-projection distinction. + +- [ ] **Step 1: Run the complete frontend verification gate** + +Run: + +```bash +cd frontend +npx vitest run +npx tsc -b +npm run build +cd .. +git diff --check +``` + +Expected: all frontend tests pass (baseline before this task: 285 tests), TypeScript/build exit 0, +and only the repository's pre-existing MSW/ref/chunk-size warnings remain. + +- [ ] **Step 2: Record the running asset and runtime state before deployment** + +Run: + +```bash +docker inspect -f '{{.Name}}|{{.State.Status}}|{{if .State.Health}}{{.State.Health.Status}}{{else}}no-healthcheck{{end}}|{{.Image}}|{{.State.StartedAt}}' thothii-core-1 thothii-frontend-1 omics_portal-web-1 +docker exec thothii-frontend-1 cat /usr/share/nginx/html/.vite/manifest.json | rg -n -B 3 '"isEntry": true' +docker top thothii-core-1 -eo pid,args +``` + +Expected: core is healthy, frontend/portal are running, the active entry is recorded, and no +`pi --mode rpc` process is active before the frontend-only recreate. + +- [ ] **Step 3: Build and recreate only the frontend service** + +Run from the repository root: + +```bash +docker compose build frontend +docker compose up -d --no-deps --force-recreate --wait --wait-timeout 60 frontend +docker inspect -f '{{.Name}}|{{.State.Status}}|{{.Image}}|{{.State.StartedAt}}' thothii-frontend-1 +docker exec thothii-frontend-1 cat /usr/share/nginx/html/.vite/manifest.json | rg -n -B 3 '"isEntry": true' +``` + +Expected: the build exits 0, frontend is running on the newly built image, and the active Vite +entry differs from the value recorded in Step 2. + +If and only if the entry changed, refresh the portal's indefinite manifest cache: + +```bash +docker restart omics_portal-web-1 +docker inspect -f '{{.Name}}|{{.State.Status}}|{{.State.StartedAt}}' omics_portal-web-1 +``` + +Expected: only `omics_portal-web-1` receives a new start timestamp; core and unrelated portal +services are not restarted. + +- [ ] **Step 4: Run a real no-COT Qwen smoke with exact settings restore** + +Run the probe inside the core container so it can reach the private backend listener. It preserves +the full settings object, creates one uniquely named session, waits for sanitized tool activity and +the first gate, asserts Qwen emitted no reasoning delta, cleans up only its own session, and restores +settings in `finally`: + +```bash +docker exec -i thothii-core-1 node - <<'NODE' +const base = "http://127.0.0.1:8787"; +const canonical = (value) => JSON.stringify(value, Object.keys(value).sort()); +async function request(path, options = {}) { + const response = await fetch(base + path, { + ...options, + headers: { "content-type": "application/json", ...(options.headers ?? {}) }, + }); + if (!response.ok) throw new Error(`${options.method ?? "GET"} ${path}: ${response.status}`); + return response.status === 204 ? null : response.json(); +} + +const saved = await request("/settings"); +const stamp = new Date().toISOString().replaceAll(/[:.]/g, "-"); +const question = `activity-filter-smoke-${stamp}: quanti record sono disponibili?`; +let sessionId = null; +let reader = null; +let activityDeltaCount = 0; +let toolCount = 0; +let gateSeen = false; +let timeout = null; +let cleanupError = null; + +try { + await request("/settings", { + method: "PUT", + body: JSON.stringify({ + ...saved, + provider: "local-qwen", + model: "qwen3.6-35b-a3b", + thinking: "low", + }), + }); + const created = await request("/sessions", { + method: "POST", + body: JSON.stringify({ question }), + }); + sessionId = created.id; + + const response = await fetch(`${base}/sessions/${encodeURIComponent(sessionId)}/events`); + if (!response.ok || !response.body) throw new Error(`SSE open failed: ${response.status}`); + reader = response.body.getReader(); + timeout = setTimeout(() => { void reader.cancel("Qwen smoke timeout"); }, 180_000); + const decoder = new TextDecoder(); + let buffer = ""; + let eventName = "message"; + let dataLines = []; + + while (!gateSeen) { + const read = await reader.read(); + if (read.done) throw new Error("SSE ended before the first gate"); + buffer += decoder.decode(read.value, { stream: true }); + const lines = buffer.split("\n"); + buffer = lines.pop() ?? ""; + for (const rawLine of lines) { + const line = rawLine.replace(/\r$/, ""); + if (line.startsWith("event:")) eventName = line.slice(6).trim(); + else if (line.startsWith("data:")) dataLines.push(line.slice(5).trimStart()); + else if (line === "") { + if (dataLines.length > 0) { + const payload = JSON.parse(dataLines.join("\n")); + if (eventName === "activity_delta") activityDeltaCount += 1; + if (eventName === "activity_event") { + const keys = Object.keys(payload.activity ?? {}).sort(); + const allowed = ["kind", "status", "toolCallId", "toolName"]; + if (JSON.stringify(keys) !== JSON.stringify(allowed)) { + throw new Error(`forbidden tool fields: ${keys.join(",")}`); + } + toolCount += 1; + } + if (eventName === "ui_request") gateSeen = true; + } + eventName = "message"; + dataLines = []; + } + } + } + + if (toolCount === 0) throw new Error("no sanitized tool activity observed"); + if (activityDeltaCount !== 0) throw new Error(`expected no COT, got ${activityDeltaCount}`); + console.log(JSON.stringify({ sessionId, activityDeltaCount, toolCount, gateSeen })); +} finally { + if (timeout) clearTimeout(timeout); + if (reader) await reader.cancel().catch(() => undefined); + if (sessionId) { + try { + const close = await fetch(`${base}/sessions/${encodeURIComponent(sessionId)}/close`, { method: "POST" }); + if (!close.ok) throw new Error(`smoke close failed: ${close.status}`); + const remove = await fetch(`${base}/sessions/${encodeURIComponent(sessionId)}`, { method: "DELETE" }); + if (remove.status !== 204) throw new Error(`smoke delete failed: ${remove.status}`); + } catch (error) { + cleanupError = error; + } + } + await request("/settings", { method: "PUT", body: JSON.stringify(saved) }); + const restored = await request("/settings"); + const sessions = await request("/sessions"); + if (canonical(restored) !== canonical(saved)) throw new Error("settings restore mismatch"); + if (sessionId && sessions.some((session) => session.id === sessionId)) { + throw new Error("smoke session still exists"); + } + if (cleanupError) throw cleanupError; + console.log(JSON.stringify({ settingsRestored: true, smokePresent: false })); +} +NODE +``` + +Expected: `activityDeltaCount` is 0, `toolCount` is greater than 0, `gateSeen` is true, tool payloads +contain only the four public fields, settings restore is exact, and the smoke session is absent. + +- [ ] **Step 5: Update durable project state with the new projection contract and evidence** + +In `brain/codebase/workflow-ui-contracts.md`, replace the complete-panel statement with: + +```markdown +- `activityLog` remains the complete in-memory chronological fold of prompt, thinking, assistant, + sanitized tool lifecycle, reviewer gates, status, and turn lifecycle. The left Model activity + panel is a strict projection: it renders only prompt, thinking, status, and gate; assistant, + tool, lifecycle, and unknown future kinds are hidden. The central transcript and workflow state + continue to consume their existing events independently. +``` + +In `PROJECT_STATE.md`, add a dated resolved-state paragraph containing the exact final frontend test +count, TypeScript/build result, new frontend image id, old/new Vite entry, portal restart decision, +Qwen smoke session id/counts, exact settings restore, cleanup, and absence of residual Pi runtime. + +- [ ] **Step 6: Recheck runtime, review scope, and commit the evidence** + +Run: + +```bash +docker inspect -f '{{.Name}}|{{.State.Status}}|{{if .State.Health}}{{.State.Health.Status}}{{else}}no-healthcheck{{end}}|{{.Image}}' thothii-core-1 thothii-frontend-1 omics_portal-web-1 +docker top thothii-core-1 -eo pid,args +docker logs --since 15m --tail 300 thothii-core-1 +docker logs --since 15m --tail 200 thothii-frontend-1 +git diff --check +git status --short +git diff -- brain/codebase/workflow-ui-contracts.md PROJECT_STATE.md +git add brain/codebase/workflow-ui-contracts.md PROJECT_STATE.md +git commit -m "docs: record filtered model activity deployment" +``` + +Expected: containers are healthy/running, no smoke/Pi process remains, logs contain no raw tool +payloads/credentials/fatal crash, only the two documentation files are committed, and `.vite/` +remains untouched. + +--- + +## Final review gate + +Before integration, request an independent review of the complete implementation range. The review +must verify the allowlist is at the rendering boundary, assistant text remains in the central +transcript, hidden updates cannot move the panel scroll, unknown kinds default to hidden, tests +exercise real store/panel behavior, and deployment evidence matches the live containers. Fix every +Critical or Important finding before declaring the work complete.