Files
ThothII/docs/superpowers/specs/2026-07-14-qwen-connectivity-resume-recovery-design.md
T

4.1 KiB

Qwen Connectivity and Resume Recovery Design

Goal

Make local-qwen/qwen3.6-35b-a3b usable from the production ThothII containers and prevent future resume requests from becoming no-ops after a Pi turn has ended or failed.

Existing Qwen sessions that failed before this change are explicitly out of scope. They do not need migration or recovery, and deployment may terminate their stale Pi processes.

Root Cause

  • Pi's mounted models.json points local-qwen at http://127.0.0.1:18000/v1.
  • Inside thothii-core, that loopback belongs to the core container, not the Docker host.
  • vLLM runs as localllm-vllm on the separate localllm_default network and exposes its host port only on 127.0.0.1, so host.docker.internal:18000 is also unreachable.
  • Pi remains alive after provider retry exhaustion. PiProcessManager equates a live child with an active turn, so POST /sessions/:id/resume returns alreadyActive without sending /riprendi-sessione.
  • Resuming the same session ID does not recreate the frontend EventSource subscription.

Qwen Network Design

Attach the production core service to the existing external localllm_default Docker network in addition to omics_portal_omics_network. Keep vLLM private and address it via Docker DNS at http://localllm-vllm:8000/v1 in the mounted Pi models.json.

Do not publish vLLM on 0.0.0.0 and do not proxy it through the public Omics network. The private shared network gives the core only the connectivity it needs without broadening host or public exposure.

Runtime Lifecycle Design

Each managed Pi runtime exposes one of four turn states:

  • idle: no turn is currently running;
  • running: a prompt is executing;
  • waiting: the workflow is blocked on a reviewer widget;
  • failed: the last assistant turn ended with stopReason: error or the child exited.

SessionBridge recognizes Pi assistant error events, emits a sanitized error info event, and exposes the lifecycle signals needed by PiProcessManager. Provider error text may be shown, but stack traces, credentials, request bodies, and endpoint secrets must not be sent to the browser.

Resume behavior is state-based:

  • running or waiting: return alreadyActive and preserve the current turn or gate;
  • idle or failed: tear down the old child, create a new runtime from the persisted provider/model/thinking settings, and send /riprendi-sessione <id>;
  • no runtime: retain the existing cold-resume behavior.

Normal agent_end changes a runtime to idle. A pending reviewer request changes it to waiting; responding to that widget hands control back to Pi and changes it to running.

Frontend Reconnection

The session stream hook accepts a connection generation key. Every successful Resume click increments the generation, even when the resumed ID equals the current active ID, causing the old EventSource to close and a fresh one to subscribe. Existing buffered SSE events and pending widget replay remain authoritative.

Assistant provider failures are rendered through the existing info/step-message path and stop the working indicator when agent_end arrives.

Verification

  • Backend unit tests cover assistant error mapping, lifecycle transitions, preservation of a running/waiting runtime, and respawn of idle/failed runtimes.
  • Frontend tests cover same-ID Resume creating a new EventSource subscription.
  • Compose configuration validation confirms both external networks are attached to core.
  • Full backend and frontend test/typecheck/build gates pass.
  • Deployment rebuilds and force-recreates the affected core and frontend containers.
  • Live smoke verifies GET /v1/models and a real Qwen inference from inside core, followed by a new ThothII Qwen session that reaches its first reviewer gate.
  • A controlled failed/idle runtime test verifies that Resume sends a fresh kickoff without disturbing a genuinely pending reviewer gate.

Out of Scope

  • Recovering, migrating, deleting, or replaying pre-fix Qwen sessions.
  • Changing the Qwen model, vLLM image, sampling parameters, or GPU allocation.
  • Exposing vLLM outside its private Docker network.
  • Changing GLM or DeepSeek provider configuration.