32 KiB
Complete Activity Log and CTE Plan Layout Implementation Plan
For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (
- [ ]) syntax for tracking.
Goal: Restore the left panel as a complete, chronological, sanitized activity log from the Phase 1 prompt onward, and finish the Phase 6 CTE-plan layout with real Tailwind 3 padding and a clearer responsive hierarchy.
Architecture: Keep persistence and workflow behavior unchanged. SessionBridge exposes only a sanitized tool lifecycle event; the frontend folds all existing stream events plus successful local user actions into one in-memory activityLog. ModelActivityPanel renders that log and owns near-bottom auto-follow. The CTE fix replaces Tailwind 4-only spacing in the shared Card primitive and simplifies the CTE viewer's nested layout without changing artifact data.
Tech Stack: Fastify/TypeScript + vitest (backend); React 18, Zustand, SSE, Tailwind CSS 3.4, Testing Library + vitest (frontend); Docker Compose v2; Impeccable layout detector.
Global Constraints
- Preserve the pre-existing untracked
.vite/directory and all unrelated user changes. - Do not modify workflow phases, harness decisions, artifact schemas, model configuration, or persisted session data.
- Never forward or render tool arguments, partial results, final results, commands, raw output, credentials, or raw exceptions.
- Keep
CentralStatusdriven by the existing assistant transcript; the complete timeline belongs only toModelActivityPanel. - UI labels remain English. Workspace document content remains in its existing language.
- Use test-driven development: observe every new focused test fail for the intended reason before changing production code.
- After each production edit, run the focused tests and the affected TypeScript gate before committing.
- Use
apply_patchfor all manual file edits. Do not touch.vite/.
Task 1: Bridge a sanitized tool lifecycle to the client
Files:
- Modify:
backend/src/bridge/session-bridge.ts - Test:
backend/test/session-bridge.test.ts
Contract:
type ToolActivity = {
kind: "tool";
toolCallId: string;
toolName: string;
status: "running" | "completed" | "failed";
};
type ClientEvent =
// existing variants remain unchanged
| { type: "activity_event"; activity: ToolActivity };
-
tool_execution_startbecomesstatus: "running". -
tool_execution_endbecomes"completed", or"failed"only whenisError === true. -
tool_execution_updateis ignored entirely. -
Missing/invalid call ids are ignored rather than emitting an uncorrelatable row.
-
Missing/invalid names use the fixed display fallback
"Tool"; no other payload field is inspected. -
Step 1: Add the failing bridge tests
Append tests that fire realistic Pi events containing sentinel secrets:
test("tool execution emits only a sanitized correlated lifecycle", () => {
const { rpc, fire } = fakeRpc();
const bridge = new SessionBridge(rpc);
const seen: any[] = [];
bridge.onClientEvent((event) => seen.push(event));
fire({
type: "tool_execution_start",
toolCallId: "tool-1",
toolName: "bash",
args: { command: "curl https://secret.invalid/?token=DO_NOT_LEAK_ARGS" },
});
fire({
type: "tool_execution_update",
toolCallId: "tool-1",
toolName: "bash",
partialResult: { content: "DO_NOT_LEAK_PARTIAL" },
});
fire({
type: "tool_execution_end",
toolCallId: "tool-1",
toolName: "bash",
result: { content: "DO_NOT_LEAK_RESULT" },
isError: false,
});
expect(seen).toEqual([
{
type: "activity_event",
activity: { kind: "tool", toolCallId: "tool-1", toolName: "bash", status: "running" },
},
{
type: "activity_event",
activity: { kind: "tool", toolCallId: "tool-1", toolName: "bash", status: "completed" },
},
]);
expect(JSON.stringify(seen)).not.toMatch(/DO_NOT_LEAK|args|partialResult|result|command/);
});
test("failed tool completion is sanitized and invalid starts are ignored", () => {
const { rpc, fire } = fakeRpc();
const bridge = new SessionBridge(rpc);
const seen: any[] = [];
bridge.onClientEvent((event) => seen.push(event));
fire({ type: "tool_execution_start", toolName: "read", args: { path: "DO_NOT_LEAK_PATH" } });
fire({
type: "tool_execution_end",
toolCallId: "tool-2",
toolName: 42,
result: { error: "DO_NOT_LEAK_ERROR" },
isError: true,
});
expect(seen).toEqual([
{
type: "activity_event",
activity: { kind: "tool", toolCallId: "tool-2", toolName: "Tool", status: "failed" },
},
]);
expect(JSON.stringify(seen)).not.toMatch(/DO_NOT_LEAK|args|result|error|path/);
});
- Step 2: Run the tests and confirm the intended failure
cd backend
npx vitest run test/session-bridge.test.ts -t "tool execution|failed tool"
Expected: FAIL because activity_event is not yet emitted; existing bridge tests remain green.
- Step 3: Implement the narrow event mapping
Extend ClientEvent, add a private helper that reads only toolCallId, toolName, and isError, and insert the start/end branches before the generic system_event branch. Delete the old comment that says all tool_execution_* events are intentionally dropped; replace it with a comment explaining that updates/raw payloads remain dropped.
- Step 4: Verify the focused backend boundary
cd backend
npx vitest run test/session-bridge.test.ts
npx tsc --noEmit -p .
Expected: all bridge tests PASS and TypeScript exits 0.
- Step 5: Commit Task 1
git add backend/src/bridge/session-bridge.ts backend/test/session-bridge.test.ts
git commit -m "feat(backend): expose sanitized tool activity"
Task 2: Fold stream events into one chronological frontend activity log
Files:
- Modify:
frontend/src/api/types.ts - Modify:
frontend/src/store/sessionStore.ts - Test:
frontend/src/store/sessionStore.test.ts - Modify:
frontend/src/stream/useSessionStream.ts - Test:
frontend/src/stream/useSessionStream.test.tsx
Frontend types:
export type ActivityKind =
| "prompt"
| "thinking"
| "assistant"
| "tool"
| "gate"
| "status"
| "lifecycle";
export interface ActivityEntry {
kind: ActivityKind;
phase: string | null;
text: string;
level?: "info" | "warning" | "error";
toolCallId?: string;
status?: "running" | "completed" | "failed";
}
Add recordLifecycle(text: string) to the store actions. It appends one local lifecycle row
using the current phase; it is used only to make an explicitly requested Resume visible after the
timeline reset.
Add the same activity_event wire variant used by the backend to StreamEvent. Replace the reasoning-only store field activity with activityLog: ActivityEntry[]. Preserve transcript, stepMessages, lastUserEntry, phase/error behavior, and all existing retry semantics.
Fold rules:
-
text_delta: keep updatingtranscript; append/coalesce anassistantrow. -
activity_delta: append/coalesce athinkingrow. -
Coalesce only when the immediately preceding timeline row has the same streaming kind.
-
activity_eventstart: append atoolrow; completion/failure: update the matching row bytoolCallId, or append a terminal row if no start was observed. -
ui_request: normalize the phase, then append agaterow usingtitleor"Review requested". -
info: keep updatingstepMessages; also append astatusrow with its level. -
system_event: keep existing lifecycle state; append readablelifecyclerows foragent_start,agent_end, and other named events. -
resetSession: clear the complete log. -
Step 1: Write failing store tests for chronology and coalescing
Add explicit tests with this event order:
const store = useSessionStore.getState();
store.setPhase("F1");
store.setLastUserEntry({ kind: "input", text: "How many patients?" });
store.applyEvent({ type: "system_event", event: "agent_start" });
store.applyEvent({ type: "activity_delta", text: "Inspect " });
store.applyEvent({ type: "activity_delta", text: "schema" });
store.applyEvent({
type: "activity_event",
activity: { kind: "tool", toolCallId: "t1", toolName: "bash", status: "running" },
});
store.applyEvent({ type: "text_delta", text: "I found " });
store.applyEvent({ type: "text_delta", text: "the candidates." });
store.applyEvent({
type: "activity_event",
activity: { kind: "tool", toolCallId: "t1", toolName: "bash", status: "completed" },
});
store.applyEvent({
type: "ui_request",
ui_request: { id: "g1", widget: "select", phase: "F1_understanding", title: "Confirm intent" },
});
Assert exact kind order:
["prompt", "lifecycle", "thinking", "tool", "assistant", "gate"]
Also assert:
-
the two thinking chunks form one row;
-
the two assistant chunks form one row;
-
the tool start/end form one completed row in its original position;
-
prompt and all following rows carry
F1when known; -
an intervening non-stream row prevents later streaming chunks from coalescing backward;
-
infoadds both the existingstepMessagesrecord and astatusrow; -
resetSessionclearsactivityLog. -
Step 2: Run the store tests and confirm RED
cd frontend
npx vitest run src/store/sessionStore.test.ts
Expected: FAIL on missing activityLog and missing activity_event behavior.
- Step 3: Implement the typed fold
Keep helper functions pure and local to sessionStore.ts:
function phaseOf(value: unknown, fallback: string | null): string | null {
return typeof value === "string" && value ? value.split("_")[0] : fallback;
}
function appendStream(
log: ActivityEntry[],
entry: ActivityEntry & { kind: "thinking" | "assistant" },
): ActivityEntry[] {
const next = [...log];
const last = next.at(-1);
if (last?.kind === entry.kind) next[next.length - 1] = { ...last, text: last.text + entry.text };
else next.push(entry);
return next;
}
For terminal tool events, find the last matching toolCallId, clone that row with the new status, and keep its array position. Do not store any wire fields beyond the sanitized activity object.
- Step 4: Add the named SSE subscription test
In useSessionStream.test.tsx, emit a named activity_event, then assert activityLog contains a running tool row. Update the existing activity_delta assertion to inspect the thinking timeline row.
- Step 5: Observe the SSE test fail, then subscribe
cd frontend
npx vitest run src/stream/useSessionStream.test.tsx -t "activity"
Expected before implementation: the named tool event is ignored. Add "activity_event" to namedEvents, rerun, and expect PASS.
- Step 6: Verify the complete store/stream slice
cd frontend
npx vitest run src/store/sessionStore.test.ts src/stream/useSessionStream.test.tsx
npx tsc -b
Expected: all focused tests PASS and TypeScript exits 0.
- Step 7: Commit Task 2
git add frontend/src/api/types.ts frontend/src/store/sessionStore.ts frontend/src/store/sessionStore.test.ts frontend/src/stream/useSessionStream.ts frontend/src/stream/useSessionStream.test.tsx
git commit -m "feat(frontend): build chronological activity timeline"
Task 3: Record Phase 1 immediately and render the complete activity panel
Files:
- Inspect:
frontend/src/shell/SteerInput.tsxto verify the establishedsetPhase("F1")thensetLastUserEntry(...)ordering; no production edit is expected - Modify:
frontend/src/shell/AppShell.tsx - Modify:
frontend/src/shell/ModelActivityPanel.tsx - Test:
frontend/src/shell/AppShell.new-session.test.tsx - Test:
frontend/src/shell/AppShell.session-mgmt.test.tsx - Test:
frontend/src/shell/ModelActivityPanel.test.tsx
Panel behavior:
- Render every
activityLogrow in arrival order; removeTAIL_PARAGRAPHS,expanded, and Show more/less. - Show compact textual metadata for phase and kind/status; never depend on color alone.
- Use normal text for prompt/assistant, secondary text for thinking, existing semantic colors for warning/error, and tool name plus
running/completed/failedonly. - Preserve Markdown rendering for assistant/thinking content without merging unrelated entries.
- Keep the scroll viewport keyboard-scrollable and auto-follow only while it is within 48 px of the bottom.
Use a testable helper:
export function isNearBottom(
el: Pick<HTMLElement, "scrollHeight" | "clientHeight" | "scrollTop">,
threshold = 48,
): boolean {
return el.scrollHeight - el.clientHeight - el.scrollTop <= threshold;
}
followRef.current starts true. onScroll updates it through isNearBottom. A useLayoutEffect keyed by activityLog assigns scrollTop = scrollHeight only when followRef.current is true.
- Step 1: Prove the initial F1 prompt exists before POST completion
Extend the existing deferred-create test in AppShell.new-session.test.tsx. Before calling releaseCreate(), assert:
expect(useSessionStore.getState().activityLog[0]).toEqual({
kind: "prompt",
phase: "F1",
text: "How many patients?",
});
Keep the existing assertion that no EventSource exists yet. This proves logging starts at submit, not at the first server event.
- Step 2: Write the failing panel tests
Replace reasoning-tail tests with timeline tests that prove:
-
a non-reasoning sequence containing prompt,
agent_start, tool start/end, assistant text, and gate renders every row; -
all eight synthetic timeline entries remain visible with no Show more/less control;
-
the displayed tool row contains
bashandcompleted, but no raw-details disclosure control; -
close invokes
onClosewithout mutatingactivityLog; -
isNearBottomreturns true at/within 48 px and false above it; -
after setting DOM
scrollHeight,clientHeight, andscrollTop, appending a row scrolls to the bottom when following; -
after firing a scroll event from well above the bottom, appending a row leaves
scrollTopunchanged. -
Step 3: Run panel and shell tests and confirm RED
cd frontend
npx vitest run src/shell/ModelActivityPanel.test.tsx src/shell/AppShell.new-session.test.tsx
Expected: FAIL because the panel still reads the removed reasoning-only array/tail and no timeline row renderer exists.
- Step 4: Implement the panel timeline and auto-follow
Build small local rendering branches rather than a single dense conditional. Use stable keys derived from row index plus toolCallId because the log never reorders; tool completion updates content in place. The scrollable element receives data-testid="activity-scroll" for the behavior tests.
Recommended row shell:
<article className="border-b border-border/50 py-3 last:border-b-0">
<div className="mb-1.5 flex flex-wrap items-center gap-1.5 text-[0.65rem] font-semibold uppercase tracking-[0.08em] text-muted-foreground">
{entry.phase && <span>{entry.phase}</span>}
<span>{entry.kind}</span>
{entry.status && <span>{entry.status}</span>}
</div>
{/* safe row-specific body */}
</article>
Do not add expandable tool details or animation.
- Step 5: Preserve local action and resume semantics
setLastUserEntry now appends prompt, so keep these existing call-site rules:
- New question:
setPhase("F1")immediately beforesetLastUserEntryand before awaitingcreateSession. - Steering: append only after
postSteersucceeds. - Gate response: append only after
postResponsesucceeds (already inWidgetHost).
In AppShell.doResume, select recordLifecycle from the store and call
recordLifecycle("Resuming session") immediately after resetSession(). This makes the new,
intentionally non-persisted resume timeline explicit while leaving historical activity empty.
- Step 6: Prove close/reopen preserves the log
Update the existing Model activity test in AppShell.session-mgmt.test.tsx:
- Open a live session.
- Apply prompt/tool/assistant events.
- Open the left panel and assert those rows.
- Close the panel, assert
activityLogis unchanged. - Reopen it and assert the same rows are still rendered.
- Step 7: Verify the complete activity UI slice
cd frontend
npx vitest run src/store/sessionStore.test.ts src/stream/useSessionStream.test.tsx src/shell/ModelActivityPanel.test.tsx src/shell/AppShell.new-session.test.tsx src/shell/AppShell.session-mgmt.test.tsx src/shell/CentralStatus.test.tsx
npx tsc -b
Expected: all focused tests PASS; CentralStatus behavior is unchanged; TypeScript exits 0.
- Step 8: Commit Task 3
git add frontend/src/shell/AppShell.tsx frontend/src/shell/ModelActivityPanel.tsx frontend/src/shell/ModelActivityPanel.test.tsx frontend/src/shell/AppShell.new-session.test.tsx frontend/src/shell/AppShell.session-mgmt.test.tsx
git commit -m "fix(frontend): show complete model activity from phase one"
Task 4: Fix Tailwind 3 card spacing and polish the CTE plan layout
Files:
- Modify:
frontend/src/components/ui/card.tsx - Create:
frontend/src/components/ui/card.test.tsx - Modify:
frontend/src/viewers/CtePlanViewer.tsx - Test:
frontend/src/viewers/CtePlanViewer.test.tsx
Shared Card compatibility target:
Replace the Tailwind 4-only custom-spacing syntax with concrete Tailwind 3 utilities:
// Card
"group/card flex flex-col gap-4 overflow-hidden rounded-xl border border-border/70 bg-card py-4 text-sm text-card-foreground shadow-sm data-[size=sm]:gap-3 data-[size=sm]:py-3"
// CardHeader / CardContent / CardFooter
"... px-4 group-data-[size=sm]/card:px-3 ..."
"px-4 group-data-[size=sm]/card:px-3"
"... p-4 group-data-[size=sm]/card:p-3"
Also replace any Tailwind 4-only has-data-* spacing variant with a Tailwind 3 arbitrary selector, or remove it when redundant. Preserve existing slots, size attributes, borders, radius, image handling, and exports.
CTE viewer target:
-
Overview group: labeled
Question,Strategy, andExecution order, withgap-3; usegap-6before the cards. -
Outer card:
rounded-lg, border,shadow-none, and no decorative nested elevation. -
Header: explicit
p-4 sm:p-5,border-b; CTE name rendered as anh3. -
Content: explicit
p-4 sm:p-5, major-group gap around 20 px. -
Delay two-column dependency/key and filter layouts to
md:where needed. -
Tables and filters: one bordered
divide-ylist withp-3 sm:p-4rows. -
Rationale: plain
border-t pt-4; removeborder-l-2, colored stripe, tinted box, and extra radius. -
Long names/values/chips: keep
min-w-0,whitespace-pre-wrapwhere appropriate, and[overflow-wrap:anywhere]. -
Step 1: Add failing Card primitive tests
Create card.test.tsx and render a default and a small card with header/content/footer. Assert the resulting slots contain these concrete classes:
-
default card:
gap-4,py-4; -
small variant rules:
data-[size=sm]:gap-3,data-[size=sm]:py-3; -
header/content:
px-4,group-data-[size=sm]/card:px-3; -
footer:
p-4,group-data-[size=sm]/card:p-3; -
no class contains
(--card-spacing)or--spacing(. -
Step 2: Add failing CTE structure/layout tests
Extend CtePlanViewer.test.tsx to assert:
-
the first CTE name is a level-3 heading;
-
Question,Strategy, andExecution orderlabels are present; -
each
card-headerandcard-contentcarriesp-4 sm:p-5; -
the CTE card has
shadow-none; -
the filters share one list container with
divide-yand each filter remains an accessiblerole="group"; -
rationale has
border-tand does not haveborder-l-2, a tinted background, or rounded-card styling; -
a fixture with very long CTE/table/filter/chip values has the wrapping classes on those values;
-
optional fields still disappear without rendering
undefined. -
Step 3: Run the focused layout tests and confirm RED
cd frontend
npx vitest run src/components/ui/card.test.tsx src/viewers/CtePlanViewer.test.tsx
Expected: FAIL on invalid Card spacing, missing semantic overview labels/heading, nested filter cards, and rationale side stripe.
- Step 4: Implement Tailwind 3-compatible Card spacing
Use concrete utilities exactly as tested. Let cn/tailwind-merge allow CTE-specific p-4 sm:p-5 to override shared px-4/py-4 defaults. Do not introduce arbitrary pixel values.
- Step 5: Implement the approved CTE hierarchy
Rename FilterCard to FilterRow and remove its private border/background/radius. Put all rows inside one shared divided container. Replace CardTitle for the CTE name with an h3 while keeping the existing font/size treatment. Preserve every artifact field and its content.
- Step 6: Run focused tests, typecheck, and production CSS compilation
cd frontend
npx vitest run src/components/ui/card.test.tsx src/viewers/CtePlanViewer.test.tsx
npx tsc -b
npm run build
Expected: tests PASS, TypeScript exits 0, Vite production build succeeds, and no Tailwind warning references the removed card-spacing syntax.
- Step 7: Run the Impeccable layout detector
node /home/admlocforn1/.codex/skills/impeccable/scripts/detect.mjs --json --scope layout frontend/src/viewers/CtePlanViewer.tsx frontend/src/components/ui/card.tsx
Expected: JSON [] and exit 0. The final review must explicitly inspect and report:
-
Card:gap-4 py-4, smallgap-3 py-3; -
CardHeader/CardContent: basepx-4, CTE overridep-4 sm:p-5; -
CTE overview-to-card separation:
gap-6; -
filter list: single border plus
divide-yand padded rows; -
rationale:
border-t pt-4, no side stripe. -
Step 8: Commit Task 4
git add frontend/src/components/ui/card.tsx frontend/src/components/ui/card.test.tsx frontend/src/viewers/CtePlanViewer.tsx frontend/src/viewers/CtePlanViewer.test.tsx
git commit -m "fix(frontend): restore CTE card padding and hierarchy"
Task 5: Full verification, durable notes, Docker deployment, and live Qwen smoke
Files:
-
Modify:
brain/codebase/workflow-ui-contracts.md -
Modify:
PROJECT_STATE.md -
Step 1: Run complete source verification
cd backend
npx vitest run
npx tsc --noEmit -p .
npm run build
cd ../frontend
npx vitest run
npx tsc -b
npm run build
cd ..
node /home/admlocforn1/.codex/skills/impeccable/scripts/detect.mjs --json --scope layout frontend/src/viewers/CtePlanViewer.tsx frontend/src/components/ui/card.tsx
git diff --check
git status --short
Expected: both complete test suites PASS, both typechecks and builds exit 0, detector prints [], git diff --check prints nothing, and status contains only intentional task files plus the pre-existing .vite/.
- Step 2: Perform an adversarial self-review before deployment
Inspect the complete task diff and verify:
- no raw Pi tool field other than id/name/status can reach
ClientEvent; - timeline ordering cannot merge across tool/gate/status rows;
- create failure still resets the provisional prompt log while preserving composer text;
- gate and steer failures still do not log a successful user response;
- close/reopen changes only panel visibility;
CentralStatusstill usestranscript;- CTE artifact values and optional-section behavior are unchanged;
- no Tailwind 4 spacing syntax remains in
card.tsx; - no unrelated file or
.vite/content is staged.
Run the requesting-code-review skill here. Address only verified findings, then rerun the affected focused tests and full typecheck.
- Step 3: Update durable architecture notes
Append concise bullets to brain/codebase/workflow-ui-contracts.md recording:
- the activity panel is a client-side chronological projection of prompt, thinking, assistant, sanitized tool lifecycle, gate, status, and turn lifecycle events; it is deliberately not persisted;
- Pi tool events may cross the bridge only as call id, tool name, and running/completed/failed status; updates/args/results stay server-side;
- the project is Tailwind 3.4, so shared primitives must use Tailwind 3-compatible concrete spacing utilities.
Update the existing PROJECT_STATE.md workflow/UI section after live verification with test totals, deployment timestamp, running image ids, activity smoke evidence, and CTE layout result. Do not add a new brain file or index entry.
- Step 4: Record the running frontend asset and active Pi processes
docker compose exec -T frontend sh -c 'ls -1 /usr/share/nginx/html/assets/index-*.js'
docker compose exec -T core sh -c 'pgrep -af "pi.*--mode rpc" || true'
docker compose ps core frontend
Expected: record the old frontend entry filename and current runtime state. Do not kill unrelated Pi processes. Container recreation is authorized by the user and may make any active persisted session resumable afterward.
- Step 5: Build and force-recreate both impacted services
docker compose build core frontend
docker compose up -d --no-deps --force-recreate --wait --wait-timeout 60 core frontend
docker compose ps core frontend
Expected: both builds exit 0; core is running (healthy) and frontend is running.
- Step 6: Refresh the portal's indefinitely cached Vite manifest
docker compose exec -T frontend sh -c 'ls -1 /usr/share/nginx/html/assets/index-*.js'
docker restart omics_portal-web-1
docker ps --filter name=omics_portal-web-1 --format '{{.Names}} {{.Status}}'
Expected: the frontend entry filename differs from the pre-build one and only the portal web container is restarted; it returns to Up status.
- Step 7: Verify container identity, health, and sanitized logs
docker image inspect thothii-core:local thothii-frontend:local --format '{{.RepoTags}} {{.Id}} {{.Created}}'
docker inspect thothii-core-1 thothii-frontend-1 --format '{{.Name}} {{.Image}} {{if .State.Health}}{{.State.Health.Status}}{{else}}{{.State.Status}}{{end}} {{.State.StartedAt}}'
docker compose logs --since=10m core frontend
Expected: running image ids equal the rebuilt image ids; core health is healthy; logs contain no credentials, raw tool payload, uncaught exception, or crash loop.
- Step 8: Run an application-level non-reasoning Qwen smoke
Run the following one-off script inside core. It saves exact settings, selects Qwen, creates one uniquely named smoke session, reads named SSE frames until the first gate, asserts at least one sanitized tool activity event, rejects forbidden activity fields, cleans up only its own session, aborts the keep-alive reader, and restores settings in finally:
docker compose exec -T core node --input-type=module - <<'NODE'
const base = "http://127.0.0.1:8787";
let original = null;
let smokeId = null;
let reader = null;
const controller = new AbortController();
const forbidden = /"(args|partialResult|result|command|errorMessage)"\s*:/;
async function json(path, init) {
const response = await fetch(`${base}${path}`, init);
const body = await response.json().catch(() => ({}));
if (!response.ok) throw new Error(`${path}: ${response.status} ${JSON.stringify(body)}`);
return body;
}
async function waitForActivityAndGate(id) {
const response = await fetch(`${base}/sessions/${id}/events`, { signal: controller.signal });
if (!response.ok || !response.body) throw new Error(`SSE: ${response.status}`);
reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
let toolSeen = false;
let gateSeen = false;
const deadline = Date.now() + 180_000;
while (Date.now() < deadline && !(toolSeen && gateSeen)) {
const { value, done } = await reader.read();
if (done) throw new Error("SSE ended before activity and gate");
buffer += decoder.decode(value, { stream: true });
let boundary;
while ((boundary = buffer.indexOf("\n\n")) >= 0) {
const frame = buffer.slice(0, boundary);
buffer = buffer.slice(boundary + 2);
const event = frame.match(/^event: (.+)$/m)?.[1];
const data = frame.match(/^data: (.+)$/m)?.[1] ?? "";
if (event === "activity_event") {
if (forbidden.test(data)) throw new Error(`forbidden tool field crossed bridge: ${data}`);
const parsed = JSON.parse(data);
const activity = parsed.activity;
if (activity?.kind === "tool" && activity.toolCallId && activity.toolName && activity.status) {
toolSeen = true;
}
}
if (event === "ui_request") gateSeen = true;
}
}
if (!toolSeen || !gateSeen) throw new Error(`smoke timeout: tool=${toolSeen} gate=${gateSeen}`);
return { toolSeen, gateSeen };
}
try {
original = await json("/settings");
await json("/settings", {
method: "PUT",
headers: { "content-type": "application/json" },
body: JSON.stringify({
...original,
provider: "local-qwen",
model: "qwen3.6-35b-a3b",
thinking: "low",
}),
});
const created = await json("/sessions", {
method: "POST",
headers: { "content-type": "application/json" },
body: JSON.stringify({ question: `Activity smoke ${new Date().toISOString()}: count patients by sex` }),
});
smokeId = created.id;
const evidence = await waitForActivityAndGate(smokeId);
console.log(JSON.stringify({ qwenActivitySmoke: "passed", smokeId, ...evidence }));
} finally {
controller.abort();
await reader?.cancel().catch(() => undefined);
if (smokeId) {
await fetch(`${base}/sessions/${smokeId}/close`, { method: "POST" }).catch(() => undefined);
await fetch(`${base}/sessions/${smokeId}`, { method: "DELETE" }).catch(() => undefined);
}
if (original) {
const restored = await fetch(`${base}/settings`, {
method: "PUT",
headers: { "content-type": "application/json" },
body: JSON.stringify(original),
});
if (!restored.ok) throw new Error(`settings restore failed: ${restored.status}`);
}
}
NODE
Expected: qwenActivitySmoke: "passed", toolSeen: true, gateSeen: true; no raw tool fields; the reader closes; only the unique smoke session is deleted; exact original settings are restored.
The frontend tests from Tasks 2–3 are the authoritative proof that the locally recorded F1 prompt precedes these SSE events and that close/reopen preserves the rendered timeline. The live smoke proves the deployed non-reasoning model supplies sanitized activity even without thinking_delta.
- Step 9: Commit verified state documentation
After inserting actual counts, timestamp, image ids, asset filename, health, and smoke evidence:
git add brain/codebase/workflow-ui-contracts.md PROJECT_STATE.md
git commit -m "docs: record activity log and CTE deployment"
git status --short
Expected: commit succeeds and status shows only the known pre-existing .vite/ directory.
Requirement Traceability
- Complete Phase 1 log from prompt: Tasks 2–3, verified before session-create completion.
- Works with Qwen reasoning disabled: Tasks 1–3 plus Task 5 live Qwen smoke.
- User-selected detail level 1: semantic events and tool name/status only; Tasks 1–3.
- No tool args/results leakage: Task 1 boundary tests and Task 5 live forbidden-field check.
- Panel close/reopen and scroll behavior: Task 3.
- CTE internal padding and hierarchy: Task 4 Card compatibility plus viewer redesign.
- Impeccable isolated mechanical/visual findings: encoded in Task 4 and rechecked in Tasks 4–5.
- Docker update: Task 5 rebuilds/recreates
coreandfrontend, then refreshes only the portal web manifest cache.
Final Review Checklist
- New-question prompt is the first
activityLogrow and carriesF1before POST completion. - A no-reasoning event sequence still shows assistant, tool, gate, status, and lifecycle activity.
- Tool rows never contain args, partial/final results, commands, output, credentials, or raw errors.
- Streaming coalesces only across adjacent entries of the same kind.
- Closing/reopening the left panel preserves the timeline.
- Near-bottom auto-follow works and manual upward scrolling is respected.
CentralStatusremains compact and transcript-driven.- CTE cards have compiled 16/20 px edge padding and no content touches the border.
- Filter/table rows use one boundary and dividers; rationale uses a top divider only.
- CTE names and long technical values wrap and remain readable responsively.
- Complete backend/frontend suites, typechecks, builds, detector, and
git diff --checkpass. - Both rebuilt containers run their new image; core is healthy; portal asset cache is refreshed.
- Live Qwen smoke reaches sanitized tool activity and an F1 gate, cleans up, and restores settings.