docs: design Qwen connectivity and resume recovery

This commit is contained in:
User
2026-07-14 20:26:00 +02:00
parent a295cabbe3
commit ad4d6963a2
@@ -0,0 +1,84 @@
# Qwen Connectivity and Resume Recovery Design
## Goal
Make `local-qwen/qwen3.6-35b-a3b` usable from the production ThothII containers and
prevent future resume requests from becoming no-ops after a Pi turn has ended or failed.
Existing Qwen sessions that failed before this change are explicitly out of scope. They
do not need migration or recovery, and deployment may terminate their stale Pi processes.
## Root Cause
- Pi's mounted `models.json` points `local-qwen` at `http://127.0.0.1:18000/v1`.
- Inside `thothii-core`, that loopback belongs to the core container, not the Docker host.
- vLLM runs as `localllm-vllm` on the separate `localllm_default` network and exposes its
host port only on `127.0.0.1`, so `host.docker.internal:18000` is also unreachable.
- Pi remains alive after provider retry exhaustion. `PiProcessManager` equates a live child
with an active turn, so `POST /sessions/:id/resume` returns `alreadyActive` without sending
`/riprendi-sessione`.
- Resuming the same session ID does not recreate the frontend EventSource subscription.
## Qwen Network Design
Attach the production `core` service to the existing external `localllm_default` Docker
network in addition to `omics_portal_omics_network`. Keep vLLM private and address it via
Docker DNS at `http://localllm-vllm:8000/v1` in the mounted Pi `models.json`.
Do not publish vLLM on `0.0.0.0` and do not proxy it through the public Omics network. The
private shared network gives the core only the connectivity it needs without broadening
host or public exposure.
## Runtime Lifecycle Design
Each managed Pi runtime exposes one of four turn states:
- `idle`: no turn is currently running;
- `running`: a prompt is executing;
- `waiting`: the workflow is blocked on a reviewer widget;
- `failed`: the last assistant turn ended with `stopReason: error` or the child exited.
`SessionBridge` recognizes Pi assistant error events, emits a sanitized error `info` event,
and exposes the lifecycle signals needed by `PiProcessManager`. Provider error text may be
shown, but stack traces, credentials, request bodies, and endpoint secrets must not be sent
to the browser.
Resume behavior is state-based:
- `running` or `waiting`: return `alreadyActive` and preserve the current turn or gate;
- `idle` or `failed`: tear down the old child, create a new runtime from the persisted
provider/model/thinking settings, and send `/riprendi-sessione <id>`;
- no runtime: retain the existing cold-resume behavior.
Normal `agent_end` changes a runtime to `idle`. A pending reviewer request changes it to
`waiting`; responding to that widget hands control back to Pi and changes it to `running`.
## Frontend Reconnection
The session stream hook accepts a connection generation key. Every successful Resume click
increments the generation, even when the resumed ID equals the current active ID, causing the
old EventSource to close and a fresh one to subscribe. Existing buffered SSE events and pending
widget replay remain authoritative.
Assistant provider failures are rendered through the existing `info`/step-message path and
stop the working indicator when `agent_end` arrives.
## Verification
- Backend unit tests cover assistant error mapping, lifecycle transitions, preservation of a
running/waiting runtime, and respawn of idle/failed runtimes.
- Frontend tests cover same-ID Resume creating a new EventSource subscription.
- Compose configuration validation confirms both external networks are attached to `core`.
- Full backend and frontend test/typecheck/build gates pass.
- Deployment rebuilds and force-recreates the affected core and frontend containers.
- Live smoke verifies `GET /v1/models` and a real Qwen inference from inside `core`, followed
by a new ThothII Qwen session that reaches its first reviewer gate.
- A controlled failed/idle runtime test verifies that Resume sends a fresh kickoff without
disturbing a genuinely pending reviewer gate.
## Out of Scope
- Recovering, migrating, deleting, or replaying pre-fix Qwen sessions.
- Changing the Qwen model, vLLM image, sampling parameters, or GPU allocation.
- Exposing vLLM outside its private Docker network.
- Changing GLM or DeepSeek provider configuration.