docs: plan complete activity log and CTE layout
This commit is contained in:
@@ -0,0 +1,776 @@
|
||||
# Complete Activity Log and CTE Plan Layout Implementation Plan
|
||||
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||||
|
||||
**Goal:** Restore the left panel as a complete, chronological, sanitized activity log from the Phase 1 prompt onward, and finish the Phase 6 CTE-plan layout with real Tailwind 3 padding and a clearer responsive hierarchy.
|
||||
|
||||
**Architecture:** Keep persistence and workflow behavior unchanged. `SessionBridge` exposes only a sanitized tool lifecycle event; the frontend folds all existing stream events plus successful local user actions into one in-memory `activityLog`. `ModelActivityPanel` renders that log and owns near-bottom auto-follow. The CTE fix replaces Tailwind 4-only spacing in the shared `Card` primitive and simplifies the CTE viewer's nested layout without changing artifact data.
|
||||
|
||||
**Tech Stack:** Fastify/TypeScript + vitest (backend); React 18, Zustand, SSE, Tailwind CSS 3.4, Testing Library + vitest (frontend); Docker Compose v2; Impeccable layout detector.
|
||||
|
||||
## Global Constraints
|
||||
|
||||
- Preserve the pre-existing untracked `.vite/` directory and all unrelated user changes.
|
||||
- Do not modify workflow phases, harness decisions, artifact schemas, model configuration, or persisted session data.
|
||||
- Never forward or render tool arguments, partial results, final results, commands, raw output, credentials, or raw exceptions.
|
||||
- Keep `CentralStatus` driven by the existing assistant transcript; the complete timeline belongs only to `ModelActivityPanel`.
|
||||
- UI labels remain English. Workspace document content remains in its existing language.
|
||||
- Use test-driven development: observe every new focused test fail for the intended reason before changing production code.
|
||||
- After each production edit, run the focused tests and the affected TypeScript gate before committing.
|
||||
- Use `apply_patch` for all manual file edits. Do not touch `.vite/`.
|
||||
|
||||
---
|
||||
|
||||
### Task 1: Bridge a sanitized tool lifecycle to the client
|
||||
|
||||
**Files:**
|
||||
|
||||
- Modify: `backend/src/bridge/session-bridge.ts`
|
||||
- Test: `backend/test/session-bridge.test.ts`
|
||||
|
||||
**Contract:**
|
||||
|
||||
```ts
|
||||
type ToolActivity = {
|
||||
kind: "tool";
|
||||
toolCallId: string;
|
||||
toolName: string;
|
||||
status: "running" | "completed" | "failed";
|
||||
};
|
||||
|
||||
type ClientEvent =
|
||||
// existing variants remain unchanged
|
||||
| { type: "activity_event"; activity: ToolActivity };
|
||||
```
|
||||
|
||||
- `tool_execution_start` becomes `status: "running"`.
|
||||
- `tool_execution_end` becomes `"completed"`, or `"failed"` only when `isError === true`.
|
||||
- `tool_execution_update` is ignored entirely.
|
||||
- Missing/invalid call ids are ignored rather than emitting an uncorrelatable row.
|
||||
- Missing/invalid names use the fixed display fallback `"Tool"`; no other payload field is inspected.
|
||||
|
||||
- [ ] **Step 1: Add the failing bridge tests**
|
||||
|
||||
Append tests that fire realistic Pi events containing sentinel secrets:
|
||||
|
||||
```ts
|
||||
test("tool execution emits only a sanitized correlated lifecycle", () => {
|
||||
const { rpc, fire } = fakeRpc();
|
||||
const bridge = new SessionBridge(rpc);
|
||||
const seen: any[] = [];
|
||||
bridge.onClientEvent((event) => seen.push(event));
|
||||
|
||||
fire({
|
||||
type: "tool_execution_start",
|
||||
toolCallId: "tool-1",
|
||||
toolName: "bash",
|
||||
args: { command: "curl https://secret.invalid/?token=DO_NOT_LEAK_ARGS" },
|
||||
});
|
||||
fire({
|
||||
type: "tool_execution_update",
|
||||
toolCallId: "tool-1",
|
||||
toolName: "bash",
|
||||
partialResult: { content: "DO_NOT_LEAK_PARTIAL" },
|
||||
});
|
||||
fire({
|
||||
type: "tool_execution_end",
|
||||
toolCallId: "tool-1",
|
||||
toolName: "bash",
|
||||
result: { content: "DO_NOT_LEAK_RESULT" },
|
||||
isError: false,
|
||||
});
|
||||
|
||||
expect(seen).toEqual([
|
||||
{
|
||||
type: "activity_event",
|
||||
activity: { kind: "tool", toolCallId: "tool-1", toolName: "bash", status: "running" },
|
||||
},
|
||||
{
|
||||
type: "activity_event",
|
||||
activity: { kind: "tool", toolCallId: "tool-1", toolName: "bash", status: "completed" },
|
||||
},
|
||||
]);
|
||||
expect(JSON.stringify(seen)).not.toMatch(/DO_NOT_LEAK|args|partialResult|result|command/);
|
||||
});
|
||||
|
||||
test("failed tool completion is sanitized and invalid starts are ignored", () => {
|
||||
const { rpc, fire } = fakeRpc();
|
||||
const bridge = new SessionBridge(rpc);
|
||||
const seen: any[] = [];
|
||||
bridge.onClientEvent((event) => seen.push(event));
|
||||
|
||||
fire({ type: "tool_execution_start", toolName: "read", args: { path: "DO_NOT_LEAK_PATH" } });
|
||||
fire({
|
||||
type: "tool_execution_end",
|
||||
toolCallId: "tool-2",
|
||||
toolName: 42,
|
||||
result: { error: "DO_NOT_LEAK_ERROR" },
|
||||
isError: true,
|
||||
});
|
||||
|
||||
expect(seen).toEqual([
|
||||
{
|
||||
type: "activity_event",
|
||||
activity: { kind: "tool", toolCallId: "tool-2", toolName: "Tool", status: "failed" },
|
||||
},
|
||||
]);
|
||||
expect(JSON.stringify(seen)).not.toMatch(/DO_NOT_LEAK|args|result|error|path/);
|
||||
});
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Run the tests and confirm the intended failure**
|
||||
|
||||
```bash
|
||||
cd backend
|
||||
npx vitest run test/session-bridge.test.ts -t "tool execution|failed tool"
|
||||
```
|
||||
|
||||
Expected: FAIL because `activity_event` is not yet emitted; existing bridge tests remain green.
|
||||
|
||||
- [ ] **Step 3: Implement the narrow event mapping**
|
||||
|
||||
Extend `ClientEvent`, add a private helper that reads only `toolCallId`, `toolName`, and `isError`, and insert the start/end branches before the generic `system_event` branch. Delete the old comment that says all `tool_execution_*` events are intentionally dropped; replace it with a comment explaining that updates/raw payloads remain dropped.
|
||||
|
||||
- [ ] **Step 4: Verify the focused backend boundary**
|
||||
|
||||
```bash
|
||||
cd backend
|
||||
npx vitest run test/session-bridge.test.ts
|
||||
npx tsc --noEmit -p .
|
||||
```
|
||||
|
||||
Expected: all bridge tests PASS and TypeScript exits 0.
|
||||
|
||||
- [ ] **Step 5: Commit Task 1**
|
||||
|
||||
```bash
|
||||
git add backend/src/bridge/session-bridge.ts backend/test/session-bridge.test.ts
|
||||
git commit -m "feat(backend): expose sanitized tool activity"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 2: Fold stream events into one chronological frontend activity log
|
||||
|
||||
**Files:**
|
||||
|
||||
- Modify: `frontend/src/api/types.ts`
|
||||
- Modify: `frontend/src/store/sessionStore.ts`
|
||||
- Test: `frontend/src/store/sessionStore.test.ts`
|
||||
- Modify: `frontend/src/stream/useSessionStream.ts`
|
||||
- Test: `frontend/src/stream/useSessionStream.test.tsx`
|
||||
|
||||
**Frontend types:**
|
||||
|
||||
```ts
|
||||
export type ActivityKind =
|
||||
| "prompt"
|
||||
| "thinking"
|
||||
| "assistant"
|
||||
| "tool"
|
||||
| "gate"
|
||||
| "status"
|
||||
| "lifecycle";
|
||||
|
||||
export interface ActivityEntry {
|
||||
kind: ActivityKind;
|
||||
phase: string | null;
|
||||
text: string;
|
||||
level?: "info" | "warning" | "error";
|
||||
toolCallId?: string;
|
||||
status?: "running" | "completed" | "failed";
|
||||
}
|
||||
```
|
||||
|
||||
Add `recordLifecycle(text: string)` to the store actions. It appends one local `lifecycle` row
|
||||
using the current phase; it is used only to make an explicitly requested Resume visible after the
|
||||
timeline reset.
|
||||
|
||||
Add the same `activity_event` wire variant used by the backend to `StreamEvent`. Replace the reasoning-only store field `activity` with `activityLog: ActivityEntry[]`. Preserve `transcript`, `stepMessages`, `lastUserEntry`, phase/error behavior, and all existing retry semantics.
|
||||
|
||||
**Fold rules:**
|
||||
|
||||
- `text_delta`: keep updating `transcript`; append/coalesce an `assistant` row.
|
||||
- `activity_delta`: append/coalesce a `thinking` row.
|
||||
- Coalesce only when the immediately preceding timeline row has the same streaming kind.
|
||||
- `activity_event` start: append a `tool` row; completion/failure: update the matching row by `toolCallId`, or append a terminal row if no start was observed.
|
||||
- `ui_request`: normalize the phase, then append a `gate` row using `title` or `"Review requested"`.
|
||||
- `info`: keep updating `stepMessages`; also append a `status` row with its level.
|
||||
- `system_event`: keep existing lifecycle state; append readable `lifecycle` rows for `agent_start`, `agent_end`, and other named events.
|
||||
- `resetSession`: clear the complete log.
|
||||
|
||||
- [ ] **Step 1: Write failing store tests for chronology and coalescing**
|
||||
|
||||
Add explicit tests with this event order:
|
||||
|
||||
```ts
|
||||
const store = useSessionStore.getState();
|
||||
store.setPhase("F1");
|
||||
store.setLastUserEntry({ kind: "input", text: "How many patients?" });
|
||||
store.applyEvent({ type: "system_event", event: "agent_start" });
|
||||
store.applyEvent({ type: "activity_delta", text: "Inspect " });
|
||||
store.applyEvent({ type: "activity_delta", text: "schema" });
|
||||
store.applyEvent({
|
||||
type: "activity_event",
|
||||
activity: { kind: "tool", toolCallId: "t1", toolName: "bash", status: "running" },
|
||||
});
|
||||
store.applyEvent({ type: "text_delta", text: "I found " });
|
||||
store.applyEvent({ type: "text_delta", text: "the candidates." });
|
||||
store.applyEvent({
|
||||
type: "activity_event",
|
||||
activity: { kind: "tool", toolCallId: "t1", toolName: "bash", status: "completed" },
|
||||
});
|
||||
store.applyEvent({
|
||||
type: "ui_request",
|
||||
ui_request: { id: "g1", widget: "select", phase: "F1_understanding", title: "Confirm intent" },
|
||||
});
|
||||
```
|
||||
|
||||
Assert exact kind order:
|
||||
|
||||
```ts
|
||||
["prompt", "lifecycle", "thinking", "tool", "assistant", "gate"]
|
||||
```
|
||||
|
||||
Also assert:
|
||||
|
||||
- the two thinking chunks form one row;
|
||||
- the two assistant chunks form one row;
|
||||
- the tool start/end form one completed row in its original position;
|
||||
- prompt and all following rows carry `F1` when known;
|
||||
- an intervening non-stream row prevents later streaming chunks from coalescing backward;
|
||||
- `info` adds both the existing `stepMessages` record and a `status` row;
|
||||
- `resetSession` clears `activityLog`.
|
||||
|
||||
- [ ] **Step 2: Run the store tests and confirm RED**
|
||||
|
||||
```bash
|
||||
cd frontend
|
||||
npx vitest run src/store/sessionStore.test.ts
|
||||
```
|
||||
|
||||
Expected: FAIL on missing `activityLog` and missing `activity_event` behavior.
|
||||
|
||||
- [ ] **Step 3: Implement the typed fold**
|
||||
|
||||
Keep helper functions pure and local to `sessionStore.ts`:
|
||||
|
||||
```ts
|
||||
function phaseOf(value: unknown, fallback: string | null): string | null {
|
||||
return typeof value === "string" && value ? value.split("_")[0] : fallback;
|
||||
}
|
||||
|
||||
function appendStream(
|
||||
log: ActivityEntry[],
|
||||
entry: ActivityEntry & { kind: "thinking" | "assistant" },
|
||||
): ActivityEntry[] {
|
||||
const next = [...log];
|
||||
const last = next.at(-1);
|
||||
if (last?.kind === entry.kind) next[next.length - 1] = { ...last, text: last.text + entry.text };
|
||||
else next.push(entry);
|
||||
return next;
|
||||
}
|
||||
```
|
||||
|
||||
For terminal tool events, find the last matching `toolCallId`, clone that row with the new status, and keep its array position. Do not store any wire fields beyond the sanitized activity object.
|
||||
|
||||
- [ ] **Step 4: Add the named SSE subscription test**
|
||||
|
||||
In `useSessionStream.test.tsx`, emit a named `activity_event`, then assert `activityLog` contains a running tool row. Update the existing `activity_delta` assertion to inspect the `thinking` timeline row.
|
||||
|
||||
- [ ] **Step 5: Observe the SSE test fail, then subscribe**
|
||||
|
||||
```bash
|
||||
cd frontend
|
||||
npx vitest run src/stream/useSessionStream.test.tsx -t "activity"
|
||||
```
|
||||
|
||||
Expected before implementation: the named tool event is ignored. Add `"activity_event"` to `namedEvents`, rerun, and expect PASS.
|
||||
|
||||
- [ ] **Step 6: Verify the complete store/stream slice**
|
||||
|
||||
```bash
|
||||
cd frontend
|
||||
npx vitest run src/store/sessionStore.test.ts src/stream/useSessionStream.test.tsx
|
||||
npx tsc -b
|
||||
```
|
||||
|
||||
Expected: all focused tests PASS and TypeScript exits 0.
|
||||
|
||||
- [ ] **Step 7: Commit Task 2**
|
||||
|
||||
```bash
|
||||
git add frontend/src/api/types.ts frontend/src/store/sessionStore.ts frontend/src/store/sessionStore.test.ts frontend/src/stream/useSessionStream.ts frontend/src/stream/useSessionStream.test.tsx
|
||||
git commit -m "feat(frontend): build chronological activity timeline"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 3: Record Phase 1 immediately and render the complete activity panel
|
||||
|
||||
**Files:**
|
||||
|
||||
- Inspect: `frontend/src/shell/SteerInput.tsx` to verify the established `setPhase("F1")` then `setLastUserEntry(...)` ordering; no production edit is expected
|
||||
- Modify: `frontend/src/shell/AppShell.tsx`
|
||||
- Modify: `frontend/src/shell/ModelActivityPanel.tsx`
|
||||
- Test: `frontend/src/shell/AppShell.new-session.test.tsx`
|
||||
- Test: `frontend/src/shell/AppShell.session-mgmt.test.tsx`
|
||||
- Test: `frontend/src/shell/ModelActivityPanel.test.tsx`
|
||||
|
||||
**Panel behavior:**
|
||||
|
||||
- Render every `activityLog` row in arrival order; remove `TAIL_PARAGRAPHS`, `expanded`, and Show more/less.
|
||||
- Show compact textual metadata for phase and kind/status; never depend on color alone.
|
||||
- Use normal text for prompt/assistant, secondary text for thinking, existing semantic colors for warning/error, and tool name plus `running/completed/failed` only.
|
||||
- Preserve Markdown rendering for assistant/thinking content without merging unrelated entries.
|
||||
- Keep the scroll viewport keyboard-scrollable and auto-follow only while it is within 48 px of the bottom.
|
||||
|
||||
Use a testable helper:
|
||||
|
||||
```ts
|
||||
export function isNearBottom(
|
||||
el: Pick<HTMLElement, "scrollHeight" | "clientHeight" | "scrollTop">,
|
||||
threshold = 48,
|
||||
): boolean {
|
||||
return el.scrollHeight - el.clientHeight - el.scrollTop <= threshold;
|
||||
}
|
||||
```
|
||||
|
||||
`followRef.current` starts true. `onScroll` updates it through `isNearBottom`. A `useLayoutEffect` keyed by `activityLog` assigns `scrollTop = scrollHeight` only when `followRef.current` is true.
|
||||
|
||||
- [ ] **Step 1: Prove the initial F1 prompt exists before POST completion**
|
||||
|
||||
Extend the existing deferred-create test in `AppShell.new-session.test.tsx`. Before calling `releaseCreate()`, assert:
|
||||
|
||||
```ts
|
||||
expect(useSessionStore.getState().activityLog[0]).toEqual({
|
||||
kind: "prompt",
|
||||
phase: "F1",
|
||||
text: "How many patients?",
|
||||
});
|
||||
```
|
||||
|
||||
Keep the existing assertion that no EventSource exists yet. This proves logging starts at submit, not at the first server event.
|
||||
|
||||
- [ ] **Step 2: Write the failing panel tests**
|
||||
|
||||
Replace reasoning-tail tests with timeline tests that prove:
|
||||
|
||||
- a non-reasoning sequence containing prompt, `agent_start`, tool start/end, assistant text, and gate renders every row;
|
||||
- all eight synthetic timeline entries remain visible with no Show more/less control;
|
||||
- the displayed tool row contains `bash` and `completed`, but no raw-details disclosure control;
|
||||
- close invokes `onClose` without mutating `activityLog`;
|
||||
- `isNearBottom` returns true at/within 48 px and false above it;
|
||||
- after setting DOM `scrollHeight`, `clientHeight`, and `scrollTop`, appending a row scrolls to the bottom when following;
|
||||
- after firing a scroll event from well above the bottom, appending a row leaves `scrollTop` unchanged.
|
||||
|
||||
- [ ] **Step 3: Run panel and shell tests and confirm RED**
|
||||
|
||||
```bash
|
||||
cd frontend
|
||||
npx vitest run src/shell/ModelActivityPanel.test.tsx src/shell/AppShell.new-session.test.tsx
|
||||
```
|
||||
|
||||
Expected: FAIL because the panel still reads the removed reasoning-only array/tail and no timeline row renderer exists.
|
||||
|
||||
- [ ] **Step 4: Implement the panel timeline and auto-follow**
|
||||
|
||||
Build small local rendering branches rather than a single dense conditional. Use stable keys derived from row index plus `toolCallId` because the log never reorders; tool completion updates content in place. The scrollable element receives `data-testid="activity-scroll"` for the behavior tests.
|
||||
|
||||
Recommended row shell:
|
||||
|
||||
```tsx
|
||||
<article className="border-b border-border/50 py-3 last:border-b-0">
|
||||
<div className="mb-1.5 flex flex-wrap items-center gap-1.5 text-[0.65rem] font-semibold uppercase tracking-[0.08em] text-muted-foreground">
|
||||
{entry.phase && <span>{entry.phase}</span>}
|
||||
<span>{entry.kind}</span>
|
||||
{entry.status && <span>{entry.status}</span>}
|
||||
</div>
|
||||
{/* safe row-specific body */}
|
||||
</article>
|
||||
```
|
||||
|
||||
Do not add expandable tool details or animation.
|
||||
|
||||
- [ ] **Step 5: Preserve local action and resume semantics**
|
||||
|
||||
`setLastUserEntry` now appends `prompt`, so keep these existing call-site rules:
|
||||
|
||||
- New question: `setPhase("F1")` immediately before `setLastUserEntry` and before awaiting `createSession`.
|
||||
- Steering: append only after `postSteer` succeeds.
|
||||
- Gate response: append only after `postResponse` succeeds (already in `WidgetHost`).
|
||||
|
||||
In `AppShell.doResume`, select `recordLifecycle` from the store and call
|
||||
`recordLifecycle("Resuming session")` immediately after `resetSession()`. This makes the new,
|
||||
intentionally non-persisted resume timeline explicit while leaving historical activity empty.
|
||||
|
||||
- [ ] **Step 6: Prove close/reopen preserves the log**
|
||||
|
||||
Update the existing Model activity test in `AppShell.session-mgmt.test.tsx`:
|
||||
|
||||
1. Open a live session.
|
||||
2. Apply prompt/tool/assistant events.
|
||||
3. Open the left panel and assert those rows.
|
||||
4. Close the panel, assert `activityLog` is unchanged.
|
||||
5. Reopen it and assert the same rows are still rendered.
|
||||
|
||||
- [ ] **Step 7: Verify the complete activity UI slice**
|
||||
|
||||
```bash
|
||||
cd frontend
|
||||
npx vitest run src/store/sessionStore.test.ts src/stream/useSessionStream.test.tsx src/shell/ModelActivityPanel.test.tsx src/shell/AppShell.new-session.test.tsx src/shell/AppShell.session-mgmt.test.tsx src/shell/CentralStatus.test.tsx
|
||||
npx tsc -b
|
||||
```
|
||||
|
||||
Expected: all focused tests PASS; CentralStatus behavior is unchanged; TypeScript exits 0.
|
||||
|
||||
- [ ] **Step 8: Commit Task 3**
|
||||
|
||||
```bash
|
||||
git add frontend/src/shell/AppShell.tsx frontend/src/shell/ModelActivityPanel.tsx frontend/src/shell/ModelActivityPanel.test.tsx frontend/src/shell/AppShell.new-session.test.tsx frontend/src/shell/AppShell.session-mgmt.test.tsx
|
||||
git commit -m "fix(frontend): show complete model activity from phase one"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 4: Fix Tailwind 3 card spacing and polish the CTE plan layout
|
||||
|
||||
**Files:**
|
||||
|
||||
- Modify: `frontend/src/components/ui/card.tsx`
|
||||
- Create: `frontend/src/components/ui/card.test.tsx`
|
||||
- Modify: `frontend/src/viewers/CtePlanViewer.tsx`
|
||||
- Test: `frontend/src/viewers/CtePlanViewer.test.tsx`
|
||||
|
||||
**Shared Card compatibility target:**
|
||||
|
||||
Replace the Tailwind 4-only custom-spacing syntax with concrete Tailwind 3 utilities:
|
||||
|
||||
```tsx
|
||||
// Card
|
||||
"group/card flex flex-col gap-4 overflow-hidden rounded-xl border border-border/70 bg-card py-4 text-sm text-card-foreground shadow-sm data-[size=sm]:gap-3 data-[size=sm]:py-3"
|
||||
|
||||
// CardHeader / CardContent / CardFooter
|
||||
"... px-4 group-data-[size=sm]/card:px-3 ..."
|
||||
"px-4 group-data-[size=sm]/card:px-3"
|
||||
"... p-4 group-data-[size=sm]/card:p-3"
|
||||
```
|
||||
|
||||
Also replace any Tailwind 4-only `has-data-*` spacing variant with a Tailwind 3 arbitrary selector, or remove it when redundant. Preserve existing slots, size attributes, borders, radius, image handling, and exports.
|
||||
|
||||
**CTE viewer target:**
|
||||
|
||||
- Overview group: labeled `Question`, `Strategy`, and `Execution order`, with `gap-3`; use `gap-6` before the cards.
|
||||
- Outer card: `rounded-lg`, border, `shadow-none`, and no decorative nested elevation.
|
||||
- Header: explicit `p-4 sm:p-5`, `border-b`; CTE name rendered as an `h3`.
|
||||
- Content: explicit `p-4 sm:p-5`, major-group gap around 20 px.
|
||||
- Delay two-column dependency/key and filter layouts to `md:` where needed.
|
||||
- Tables and filters: one bordered `divide-y` list with `p-3 sm:p-4` rows.
|
||||
- Rationale: plain `border-t pt-4`; remove `border-l-2`, colored stripe, tinted box, and extra radius.
|
||||
- Long names/values/chips: keep `min-w-0`, `whitespace-pre-wrap` where appropriate, and `[overflow-wrap:anywhere]`.
|
||||
|
||||
- [ ] **Step 1: Add failing Card primitive tests**
|
||||
|
||||
Create `card.test.tsx` and render a default and a small card with header/content/footer. Assert the resulting slots contain these concrete classes:
|
||||
|
||||
- default card: `gap-4`, `py-4`;
|
||||
- small variant rules: `data-[size=sm]:gap-3`, `data-[size=sm]:py-3`;
|
||||
- header/content: `px-4`, `group-data-[size=sm]/card:px-3`;
|
||||
- footer: `p-4`, `group-data-[size=sm]/card:p-3`;
|
||||
- no class contains `(--card-spacing)` or `--spacing(`.
|
||||
|
||||
- [ ] **Step 2: Add failing CTE structure/layout tests**
|
||||
|
||||
Extend `CtePlanViewer.test.tsx` to assert:
|
||||
|
||||
- the first CTE name is a level-3 heading;
|
||||
- `Question`, `Strategy`, and `Execution order` labels are present;
|
||||
- each `card-header` and `card-content` carries `p-4 sm:p-5`;
|
||||
- the CTE card has `shadow-none`;
|
||||
- the filters share one list container with `divide-y` and each filter remains an accessible `role="group"`;
|
||||
- rationale has `border-t` and does not have `border-l-2`, a tinted background, or rounded-card styling;
|
||||
- a fixture with very long CTE/table/filter/chip values has the wrapping classes on those values;
|
||||
- optional fields still disappear without rendering `undefined`.
|
||||
|
||||
- [ ] **Step 3: Run the focused layout tests and confirm RED**
|
||||
|
||||
```bash
|
||||
cd frontend
|
||||
npx vitest run src/components/ui/card.test.tsx src/viewers/CtePlanViewer.test.tsx
|
||||
```
|
||||
|
||||
Expected: FAIL on invalid Card spacing, missing semantic overview labels/heading, nested filter cards, and rationale side stripe.
|
||||
|
||||
- [ ] **Step 4: Implement Tailwind 3-compatible Card spacing**
|
||||
|
||||
Use concrete utilities exactly as tested. Let `cn`/`tailwind-merge` allow CTE-specific `p-4 sm:p-5` to override shared `px-4`/`py-4` defaults. Do not introduce arbitrary pixel values.
|
||||
|
||||
- [ ] **Step 5: Implement the approved CTE hierarchy**
|
||||
|
||||
Rename `FilterCard` to `FilterRow` and remove its private border/background/radius. Put all rows inside one shared divided container. Replace `CardTitle` for the CTE name with an `h3` while keeping the existing font/size treatment. Preserve every artifact field and its content.
|
||||
|
||||
- [ ] **Step 6: Run focused tests, typecheck, and production CSS compilation**
|
||||
|
||||
```bash
|
||||
cd frontend
|
||||
npx vitest run src/components/ui/card.test.tsx src/viewers/CtePlanViewer.test.tsx
|
||||
npx tsc -b
|
||||
npm run build
|
||||
```
|
||||
|
||||
Expected: tests PASS, TypeScript exits 0, Vite production build succeeds, and no Tailwind warning references the removed card-spacing syntax.
|
||||
|
||||
- [ ] **Step 7: Run the Impeccable layout detector**
|
||||
|
||||
```bash
|
||||
node /home/admlocforn1/.codex/skills/impeccable/scripts/detect.mjs --json --scope layout frontend/src/viewers/CtePlanViewer.tsx frontend/src/components/ui/card.tsx
|
||||
```
|
||||
|
||||
Expected: JSON `[]` and exit 0. The final review must explicitly inspect and report:
|
||||
|
||||
- `Card`: `gap-4 py-4`, small `gap-3 py-3`;
|
||||
- `CardHeader`/`CardContent`: base `px-4`, CTE override `p-4 sm:p-5`;
|
||||
- CTE overview-to-card separation: `gap-6`;
|
||||
- filter list: single border plus `divide-y` and padded rows;
|
||||
- rationale: `border-t pt-4`, no side stripe.
|
||||
|
||||
- [ ] **Step 8: Commit Task 4**
|
||||
|
||||
```bash
|
||||
git add frontend/src/components/ui/card.tsx frontend/src/components/ui/card.test.tsx frontend/src/viewers/CtePlanViewer.tsx frontend/src/viewers/CtePlanViewer.test.tsx
|
||||
git commit -m "fix(frontend): restore CTE card padding and hierarchy"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 5: Full verification, durable notes, Docker deployment, and live Qwen smoke
|
||||
|
||||
**Files:**
|
||||
|
||||
- Modify: `brain/codebase/workflow-ui-contracts.md`
|
||||
- Modify: `PROJECT_STATE.md`
|
||||
|
||||
- [ ] **Step 1: Run complete source verification**
|
||||
|
||||
```bash
|
||||
cd backend
|
||||
npx vitest run
|
||||
npx tsc --noEmit -p .
|
||||
npm run build
|
||||
|
||||
cd ../frontend
|
||||
npx vitest run
|
||||
npx tsc -b
|
||||
npm run build
|
||||
|
||||
cd ..
|
||||
node /home/admlocforn1/.codex/skills/impeccable/scripts/detect.mjs --json --scope layout frontend/src/viewers/CtePlanViewer.tsx frontend/src/components/ui/card.tsx
|
||||
git diff --check
|
||||
git status --short
|
||||
```
|
||||
|
||||
Expected: both complete test suites PASS, both typechecks and builds exit 0, detector prints `[]`, `git diff --check` prints nothing, and status contains only intentional task files plus the pre-existing `.vite/`.
|
||||
|
||||
- [ ] **Step 2: Perform an adversarial self-review before deployment**
|
||||
|
||||
Inspect the complete task diff and verify:
|
||||
|
||||
- no raw Pi tool field other than id/name/status can reach `ClientEvent`;
|
||||
- timeline ordering cannot merge across tool/gate/status rows;
|
||||
- create failure still resets the provisional prompt log while preserving composer text;
|
||||
- gate and steer failures still do not log a successful user response;
|
||||
- close/reopen changes only panel visibility;
|
||||
- `CentralStatus` still uses `transcript`;
|
||||
- CTE artifact values and optional-section behavior are unchanged;
|
||||
- no Tailwind 4 spacing syntax remains in `card.tsx`;
|
||||
- no unrelated file or `.vite/` content is staged.
|
||||
|
||||
Run the requesting-code-review skill here. Address only verified findings, then rerun the affected focused tests and full typecheck.
|
||||
|
||||
- [ ] **Step 3: Update durable architecture notes**
|
||||
|
||||
Append concise bullets to `brain/codebase/workflow-ui-contracts.md` recording:
|
||||
|
||||
- the activity panel is a client-side chronological projection of prompt, thinking, assistant, sanitized tool lifecycle, gate, status, and turn lifecycle events; it is deliberately not persisted;
|
||||
- Pi tool events may cross the bridge only as call id, tool name, and running/completed/failed status; updates/args/results stay server-side;
|
||||
- the project is Tailwind 3.4, so shared primitives must use Tailwind 3-compatible concrete spacing utilities.
|
||||
|
||||
Update the existing `PROJECT_STATE.md` workflow/UI section after live verification with test totals, deployment timestamp, running image ids, activity smoke evidence, and CTE layout result. Do not add a new brain file or index entry.
|
||||
|
||||
- [ ] **Step 4: Record the running frontend asset and active Pi processes**
|
||||
|
||||
```bash
|
||||
docker compose exec -T frontend sh -c 'ls -1 /usr/share/nginx/html/assets/index-*.js'
|
||||
docker compose exec -T core sh -c 'pgrep -af "pi.*--mode rpc" || true'
|
||||
docker compose ps core frontend
|
||||
```
|
||||
|
||||
Expected: record the old frontend entry filename and current runtime state. Do not kill unrelated Pi processes. Container recreation is authorized by the user and may make any active persisted session resumable afterward.
|
||||
|
||||
- [ ] **Step 5: Build and force-recreate both impacted services**
|
||||
|
||||
```bash
|
||||
docker compose build core frontend
|
||||
docker compose up -d --no-deps --force-recreate --wait --wait-timeout 60 core frontend
|
||||
docker compose ps core frontend
|
||||
```
|
||||
|
||||
Expected: both builds exit 0; `core` is `running (healthy)` and `frontend` is running.
|
||||
|
||||
- [ ] **Step 6: Refresh the portal's indefinitely cached Vite manifest**
|
||||
|
||||
```bash
|
||||
docker compose exec -T frontend sh -c 'ls -1 /usr/share/nginx/html/assets/index-*.js'
|
||||
docker restart omics_portal-web-1
|
||||
docker ps --filter name=omics_portal-web-1 --format '{{.Names}} {{.Status}}'
|
||||
```
|
||||
|
||||
Expected: the frontend entry filename differs from the pre-build one and only the portal web container is restarted; it returns to Up status.
|
||||
|
||||
- [ ] **Step 7: Verify container identity, health, and sanitized logs**
|
||||
|
||||
```bash
|
||||
docker image inspect thothii-core:local thothii-frontend:local --format '{{.RepoTags}} {{.Id}} {{.Created}}'
|
||||
docker inspect thothii-core-1 thothii-frontend-1 --format '{{.Name}} {{.Image}} {{if .State.Health}}{{.State.Health.Status}}{{else}}{{.State.Status}}{{end}} {{.State.StartedAt}}'
|
||||
docker compose logs --since=10m core frontend
|
||||
```
|
||||
|
||||
Expected: running image ids equal the rebuilt image ids; core health is `healthy`; logs contain no credentials, raw tool payload, uncaught exception, or crash loop.
|
||||
|
||||
- [ ] **Step 8: Run an application-level non-reasoning Qwen smoke**
|
||||
|
||||
Run the following one-off script inside `core`. It saves exact settings, selects Qwen, creates one uniquely named smoke session, reads named SSE frames until the first gate, asserts at least one sanitized tool activity event, rejects forbidden activity fields, cleans up only its own session, aborts the keep-alive reader, and restores settings in `finally`:
|
||||
|
||||
```bash
|
||||
docker compose exec -T core node --input-type=module - <<'NODE'
|
||||
const base = "http://127.0.0.1:8787";
|
||||
let original = null;
|
||||
let smokeId = null;
|
||||
let reader = null;
|
||||
const controller = new AbortController();
|
||||
const forbidden = /"(args|partialResult|result|command|errorMessage)"\s*:/;
|
||||
|
||||
async function json(path, init) {
|
||||
const response = await fetch(`${base}${path}`, init);
|
||||
const body = await response.json().catch(() => ({}));
|
||||
if (!response.ok) throw new Error(`${path}: ${response.status} ${JSON.stringify(body)}`);
|
||||
return body;
|
||||
}
|
||||
|
||||
async function waitForActivityAndGate(id) {
|
||||
const response = await fetch(`${base}/sessions/${id}/events`, { signal: controller.signal });
|
||||
if (!response.ok || !response.body) throw new Error(`SSE: ${response.status}`);
|
||||
reader = response.body.getReader();
|
||||
const decoder = new TextDecoder();
|
||||
let buffer = "";
|
||||
let toolSeen = false;
|
||||
let gateSeen = false;
|
||||
const deadline = Date.now() + 180_000;
|
||||
while (Date.now() < deadline && !(toolSeen && gateSeen)) {
|
||||
const { value, done } = await reader.read();
|
||||
if (done) throw new Error("SSE ended before activity and gate");
|
||||
buffer += decoder.decode(value, { stream: true });
|
||||
let boundary;
|
||||
while ((boundary = buffer.indexOf("\n\n")) >= 0) {
|
||||
const frame = buffer.slice(0, boundary);
|
||||
buffer = buffer.slice(boundary + 2);
|
||||
const event = frame.match(/^event: (.+)$/m)?.[1];
|
||||
const data = frame.match(/^data: (.+)$/m)?.[1] ?? "";
|
||||
if (event === "activity_event") {
|
||||
if (forbidden.test(data)) throw new Error(`forbidden tool field crossed bridge: ${data}`);
|
||||
const parsed = JSON.parse(data);
|
||||
const activity = parsed.activity;
|
||||
if (activity?.kind === "tool" && activity.toolCallId && activity.toolName && activity.status) {
|
||||
toolSeen = true;
|
||||
}
|
||||
}
|
||||
if (event === "ui_request") gateSeen = true;
|
||||
}
|
||||
}
|
||||
if (!toolSeen || !gateSeen) throw new Error(`smoke timeout: tool=${toolSeen} gate=${gateSeen}`);
|
||||
return { toolSeen, gateSeen };
|
||||
}
|
||||
|
||||
try {
|
||||
original = await json("/settings");
|
||||
await json("/settings", {
|
||||
method: "PUT",
|
||||
headers: { "content-type": "application/json" },
|
||||
body: JSON.stringify({
|
||||
...original,
|
||||
provider: "local-qwen",
|
||||
model: "qwen3.6-35b-a3b",
|
||||
thinking: "low",
|
||||
}),
|
||||
});
|
||||
const created = await json("/sessions", {
|
||||
method: "POST",
|
||||
headers: { "content-type": "application/json" },
|
||||
body: JSON.stringify({ question: `Activity smoke ${new Date().toISOString()}: count patients by sex` }),
|
||||
});
|
||||
smokeId = created.id;
|
||||
const evidence = await waitForActivityAndGate(smokeId);
|
||||
console.log(JSON.stringify({ qwenActivitySmoke: "passed", smokeId, ...evidence }));
|
||||
} finally {
|
||||
controller.abort();
|
||||
await reader?.cancel().catch(() => undefined);
|
||||
if (smokeId) {
|
||||
await fetch(`${base}/sessions/${smokeId}/close`, { method: "POST" }).catch(() => undefined);
|
||||
await fetch(`${base}/sessions/${smokeId}`, { method: "DELETE" }).catch(() => undefined);
|
||||
}
|
||||
if (original) {
|
||||
const restored = await fetch(`${base}/settings`, {
|
||||
method: "PUT",
|
||||
headers: { "content-type": "application/json" },
|
||||
body: JSON.stringify(original),
|
||||
});
|
||||
if (!restored.ok) throw new Error(`settings restore failed: ${restored.status}`);
|
||||
}
|
||||
}
|
||||
NODE
|
||||
```
|
||||
|
||||
Expected: `qwenActivitySmoke: "passed"`, `toolSeen: true`, `gateSeen: true`; no raw tool fields; the reader closes; only the unique smoke session is deleted; exact original settings are restored.
|
||||
|
||||
The frontend tests from Tasks 2–3 are the authoritative proof that the locally recorded F1 prompt precedes these SSE events and that close/reopen preserves the rendered timeline. The live smoke proves the deployed non-reasoning model supplies sanitized activity even without `thinking_delta`.
|
||||
|
||||
- [ ] **Step 9: Commit verified state documentation**
|
||||
|
||||
After inserting actual counts, timestamp, image ids, asset filename, health, and smoke evidence:
|
||||
|
||||
```bash
|
||||
git add brain/codebase/workflow-ui-contracts.md PROJECT_STATE.md
|
||||
git commit -m "docs: record activity log and CTE deployment"
|
||||
git status --short
|
||||
```
|
||||
|
||||
Expected: commit succeeds and status shows only the known pre-existing `.vite/` directory.
|
||||
|
||||
---
|
||||
|
||||
## Requirement Traceability
|
||||
|
||||
- Complete Phase 1 log from prompt: Tasks 2–3, verified before session-create completion.
|
||||
- Works with Qwen reasoning disabled: Tasks 1–3 plus Task 5 live Qwen smoke.
|
||||
- User-selected detail level 1: semantic events and tool name/status only; Tasks 1–3.
|
||||
- No tool args/results leakage: Task 1 boundary tests and Task 5 live forbidden-field check.
|
||||
- Panel close/reopen and scroll behavior: Task 3.
|
||||
- CTE internal padding and hierarchy: Task 4 Card compatibility plus viewer redesign.
|
||||
- Impeccable isolated mechanical/visual findings: encoded in Task 4 and rechecked in Tasks 4–5.
|
||||
- Docker update: Task 5 rebuilds/recreates `core` and `frontend`, then refreshes only the portal web manifest cache.
|
||||
|
||||
## Final Review Checklist
|
||||
|
||||
- [ ] New-question prompt is the first `activityLog` row and carries `F1` before POST completion.
|
||||
- [ ] A no-reasoning event sequence still shows assistant, tool, gate, status, and lifecycle activity.
|
||||
- [ ] Tool rows never contain args, partial/final results, commands, output, credentials, or raw errors.
|
||||
- [ ] Streaming coalesces only across adjacent entries of the same kind.
|
||||
- [ ] Closing/reopening the left panel preserves the timeline.
|
||||
- [ ] Near-bottom auto-follow works and manual upward scrolling is respected.
|
||||
- [ ] `CentralStatus` remains compact and transcript-driven.
|
||||
- [ ] CTE cards have compiled 16/20 px edge padding and no content touches the border.
|
||||
- [ ] Filter/table rows use one boundary and dividers; rationale uses a top divider only.
|
||||
- [ ] CTE names and long technical values wrap and remain readable responsively.
|
||||
- [ ] Complete backend/frontend suites, typechecks, builds, detector, and `git diff --check` pass.
|
||||
- [ ] Both rebuilt containers run their new image; core is healthy; portal asset cache is refreshed.
|
||||
- [ ] Live Qwen smoke reaches sanitized tool activity and an F1 gate, cleans up, and restores settings.
|
||||
Reference in New Issue
Block a user