Audit findings 4.1-4.6.
- spawnFor: a rejected configure/start no longer leaks a registered runtime
with a live Pi child (identity-checked teardown + rethrow); every later
start used to hit "session runtime already active".
- ThtRunner.run: default 60s timeout on every tht child (SIGKILL backstop),
120s for DWH-touching calls (sql preview/export, search pack); a dropped
VPN mid-call no longer wedges the HTTP request forever.
- configArg: a NAMED workspace whose yaml is missing now throws instead of
silently falling back to the default config (operations were silently
targeting the wrong workspace).
- resume: the finalized/archived 409 is evaluated BEFORE the alreadyActive
fast-path — the manifest is the truth even with a lingering runtime.
- ollamaEnsure: exit-0 with non-JSON stdout is a failed check, not ok:true.
- SessionBridge.respond: only the response matching the pending descriptor
is forwarded to Pi; stale/duplicate submissions return 409 instead of
being sent with the current gate's RPC id.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Audit finding 3.1 (high, 3/3 reviewer consensus). Event ids restart at 1
when the backend restarts; a browser auto-reconnect carrying the old
numeric Last-Event-ID was honored whenever the new process had already
emitted that many events, silently suppressing fresh events (same ids,
different content). The previous guard only caught cursor > lastId.
Wire ids are now "<generation>:<seq>" (generation = per-hub instance
token; seq = the existing per-session monotonic counter). The hub parses
raw header/query candidates itself: other-generation and legacy bare-
number cursors are stale → replay from the beginning; same-generation
cursors keep the newest-valid-wins behavior. EventSource treats ids as
opaque, so no frontend change.
Finding 3.2 (eviction) resolved by NOT evicting: close keeps the seq
counter on purpose (sessions reopen; monotonicity is what makes old
cursors detectable) — documented at the call site; buffers are emptied by
clear() and ring-bounded at 200.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
New session now refuses to spawn a Pi runtime that would only die in bootstrap
retrieval when the DWH/vector host is unreachable (e.g. a dropped VPN). Before
`session new`, POST /sessions probes the DWH via `tht db ping`; if it is down it
returns 503 {code:"dwh_unreachable"} with a clear message and creates nothing.
- Gated behind the THT_DWH_PRECHECK flag (default off), enabled only by the local
dev launcher (run-stack.sh) — containers/CI never pay the probe, and existing
tests that don't set it are unaffected.
- ThtRunner.dbPing() runs `tht db ping` with a 10s timeout (run() gains an optional
timeout that SIGKILLs a hung child).
- Frontend: apiFetch throws a typed ApiError (status + parsed payload); the new-
session composer shows the specific alert on `dwh_unreachable` instead of the
generic retry hint, keeping the question for retry.
Verified live on an isolated backend (precheck on + broken DWH host → 503
dwh_unreachable, no session created) and via unit tests (backend 228, frontend 308).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>