# Qwen Connectivity and Resume Recovery Design ## Goal Make `local-qwen/qwen3.6-35b-a3b` usable from the production ThothII containers and prevent future resume requests from becoming no-ops after a Pi turn has ended or failed. Existing Qwen sessions that failed before this change are explicitly out of scope. They do not need migration or recovery, and deployment may terminate their stale Pi processes. ## Root Cause - Pi's mounted `models.json` points `local-qwen` at `http://127.0.0.1:18000/v1`. - Inside `thothii-core`, that loopback belongs to the core container, not the Docker host. - vLLM runs as `localllm-vllm` on the separate `localllm_default` network and exposes its host port only on `127.0.0.1`, so `host.docker.internal:18000` is also unreachable. - Pi remains alive after provider retry exhaustion. `PiProcessManager` equates a live child with an active turn, so `POST /sessions/:id/resume` returns `alreadyActive` without sending `/riprendi-sessione`. - Resuming the same session ID does not recreate the frontend EventSource subscription. ## Qwen Network Design Attach the production `core` service to the existing external `localllm_default` Docker network in addition to `omics_portal_omics_network`. Keep vLLM private and address it via Docker DNS at `http://localllm-vllm:8000/v1` in the mounted Pi `models.json`. Do not publish vLLM on `0.0.0.0` and do not proxy it through the public Omics network. The private shared network gives the core only the connectivity it needs without broadening host or public exposure. ## Runtime Lifecycle Design Each managed Pi runtime exposes one of four turn states: - `idle`: no turn is currently running; - `running`: a prompt is executing; - `waiting`: the workflow is blocked on a reviewer widget; - `failed`: the last assistant turn ended with `stopReason: error` or the child exited. `SessionBridge` recognizes Pi assistant error events, emits a sanitized error `info` event, and exposes the lifecycle signals needed by `PiProcessManager`. Provider error text may be shown, but stack traces, credentials, request bodies, and endpoint secrets must not be sent to the browser. Resume behavior is state-based: - `running` or `waiting`: return `alreadyActive` and preserve the current turn or gate; - `idle` or `failed`: tear down the old child, create a new runtime from the persisted provider/model/thinking settings, and send `/riprendi-sessione `; - no runtime: retain the existing cold-resume behavior. Normal `agent_end` changes a runtime to `idle`. A pending reviewer request changes it to `waiting`; responding to that widget hands control back to Pi and changes it to `running`. ## Frontend Reconnection The session stream hook accepts a connection generation key. Every successful Resume click increments the generation, even when the resumed ID equals the current active ID, causing the old EventSource to close and a fresh one to subscribe. Existing buffered SSE events and pending widget replay remain authoritative. Assistant provider failures are rendered through the existing `info`/step-message path and stop the working indicator when `agent_end` arrives. ## Verification - Backend unit tests cover assistant error mapping, lifecycle transitions, preservation of a running/waiting runtime, and respawn of idle/failed runtimes. - Frontend tests cover same-ID Resume creating a new EventSource subscription. - Compose configuration validation confirms both external networks are attached to `core`. - Full backend and frontend test/typecheck/build gates pass. - Deployment rebuilds and force-recreates the affected core and frontend containers. - Live smoke verifies `GET /v1/models` and a real Qwen inference from inside `core`, followed by a new ThothII Qwen session that reaches its first reviewer gate. - A controlled failed/idle runtime test verifies that Resume sends a fresh kickoff without disturbing a genuinely pending reviewer gate. ## Out of Scope - Recovering, migrating, deleting, or replaying pre-fix Qwen sessions. - Changing the Qwen model, vLLM image, sampling parameters, or GPU allocation. - Exposing vLLM outside its private Docker network. - Changing GLM or DeepSeek provider configuration.