Files
ThothII/docs/superpowers/specs/2026-06-29-ollama-ensure-design.md
T

8.4 KiB

Ollama Ensure (embeddings preflight) — Design

Date: 2026-06-29 Status: Approved (design), pending implementation plan Layers: harness (tht CLI), backend (Fastify). No frontend code beyond an existing toast path.

Problem

The NL→SQL workflow depends on Ollama embeddings at session time: tht search find (Phases 1/4) and tht memory search (Phase 2) call OllamaEmbeddings against {base_url}/api/embed with the workspace's configured model (e.g. nomic-embed-text-v2-moe for psd, at http://localhost:11434). If Ollama is off, or the embedding model is not loaded, these calls fail mid-session with EmbeddingsError.

Embeddings are a hard requirement — the system cannot work without them. So at session creation or restart the system must: ensure Ollama is running (start it if down), confirm the embedding model is installed and warm it into memory, and refuse to start the session (hard error) whenever embeddings cannot be made available for any reason.

Decisions (from brainstorming)

  • Mechanism: invoke the ollama CLI, but parameterized — the binary and the start command are configurable (other contexts, e.g. Docker, differ). The base URL is the existing THT_OLLAMA_URL / embeddings.base_url.
  • Blocking with timeout: session create/restart waits for Ollama to become reachable and the model to warm, up to a configurable timeout (default 60s).
  • Hard-fail (NOT degrade): on the timeout, or any other reason embeddings are unavailable, the preflight fails and the session is refused — no Pi process is spawned. There is no "degraded" session.
  • Warm-only model load: the model is assumed already installed (ollama ls shows it); "load" = warm it into memory via an embed ping. If it is not installed → hard error with guidance (ollama pull <model>), no automatic pull (a multi-GB sync download would blow the timeout and the name may not be registry-pullable).
  • Ownership: a single deterministic tht ollama ensure command in the harness (which owns the embeddings config + the OllamaEmbeddings client); the backend calls it as a preflight. Rejected: doing it in the backend (it does not know the model name — it lives in the workspace YAML) or splitting start/warm across layers (needless coordination).

Failure matrix (tht ollama ensure)

Condition Result
Workspace has no embeddings config ERROR — "system requires embeddings; workspace not configured"
Ollama unreachable and not startable within the timeout ERROR
Configured model not installed in Ollama ERROR — "run ollama pull <model> / import it"
Warm fails (embed ping errors) ERROR
Server up (or started) + model present + warm ok OK

ERROR ⇒ exit code ≠ 0 ⇒ backend refuses the session. OK ⇒ exit 0.

Harness layer

Config — EmbeddingsConfig (config.py)

Add two optional fields (defaults make existing workspace YAMLs work unchanged):

  • bin: str = "ollama" — the CLI binary (path-overridable per context).
  • start_cmd: list[str] | None = None — command to start the server. When None, default to [bin, "serve"] at use-time. Set explicitly in YAML for other contexts, e.g. ["docker", "start", "ollama"]. Set to [] to disable auto-start (remote Ollama: probe only, never try to start — still hard-errors if unreachable).

YAML supports ${VAR} expansion already, so these can reference env if needed.

New command — tht ollama ensure

A new ollama Typer sub-app (registered in cli/__init__.py alongside session_app/vector_app).

tht ollama ensure [--timeout 60] [--no-start] [--json] -c <workspace>

(-c is the per-command CONFIG_OPT, appended after the subcommand — see gotchas.)

Steps (each step is separately diagnosable so the error message names the exact failure):

  1. Load config. If cfg.embeddings is None → ERROR (exit ≠ 0).
  2. Probe the server: GET {base_url}/api/tags (short timeout). Reachable → go to 4.
  3. Start (unless --no-start or start_cmd == []): spawn the resolved start command detached (subprocess.Popen, start_new_session=True, output to devnull/log so it outlives tht). Poll /api/tags until reachable or the --timeout elapses. Still unreachable → ERROR.
  4. Model present? Match cfg.embeddings.model against the installed models from /api/tags (allow the implicit :latest tag). Absent → ERROR (guidance to pull/import).
  5. Warm: OllamaEmbeddings(cfg).embed_query("ping") (loads the model into memory). Raises EmbeddingsError → ERROR.
  6. Success. --json (pristine stdout) emits { "ok": true, "server": "up" | "started", "model": "warmed", "model_name": "<model>" }. On any ERROR with --json, emit { "ok": false, "stage": "<config|server|model|warm>", "error": "<message>" } on stdout and exit ≠ 0. The exit code is authoritative (the backend maps exit 0 → ok:true, non-zero → ok:false) and merges the parsed JSON for the stage/error detail.

The probe and the start are isolated helpers (pure-ish, injectable) so tests can drive them without a real Ollama: a _probe(base_url) -> bool, a _installed_models(base_url) -> set[str], and the start spawn behind a seam.

Backend layer (Fastify)

ThtRunner (tht-runner.ts)

  • ollamaEnsure(workspace: string, timeoutSec: number): Promise<{ ok: boolean; stage?: string; error?: string }> — shells tht ollama ensure --json --timeout <sec> (workspace via the existing -c arg). Returns the parsed JSON; treats a non-zero exit as { ok: false, ... } (parse stderr/stdout for the message) rather than throwing, so the route controls the HTTP response.

Routes (sessions.ts)

The preflight runs first, before any session is created or resumed:

  • POST /sessions: const r = await tht.ollamaEnsure(settings.workspace, timeout). If !r.ok → reply.code(503).send({ error: r.error }) and return (do NOT call sessionNew/spawnFor).
  • POST /sessions/:id/resume: the finalized/archived guard runs first — a read-only session is rejected with 409 without touching Ollama. The preflight runs after that guard and before mgr.resume. On failure → 503, no spawn. (Uses the session's workspace — the current configured/settings workspace, consistent with how resume resolves the session.)

Config (config.ts)

  • ollamaEnsureTimeoutMs: number from OLLAMA_ENSURE_TIMEOUT_MS (default 60000); passed to ollamaEnsure as seconds.

(The frontend already surfaces backend error responses via the existing toast/error path from session-creation failures — no new frontend code; the 503 message reaches the user.)

Testing

  • Harness (pytest): drive ensure with Ollama mocked via the injectable seams —
    • no embeddings config → exit ≠ 0, stage config;
    • server reachable + model present → warm called (embed ping issued), exit 0, --json shape;
    • server unreachable + start_cmd set → start spawned, then poll succeeds → exit 0;
    • server unreachable after timeout → exit ≠ 0, stage server;
    • model absent from /api/tags → exit ≠ 0, stage model, message names ollama pull;
    • warm raises EmbeddingsError → exit ≠ 0, stage warm;
    • --no-start / start_cmd == [] + unreachable → exit ≠ 0 without spawning.
    • --json stdout pristine on both success and error.
  • Backend (vitest): ollamaEnsure builds argv ["ollama","ensure","--json","--timeout","60", ...]
    • -c <workspace> and maps non-zero exit to { ok: false }; POST /sessions and /resume call it before spawn and return 503 (no Pi spawned) when it fails, proceed when it succeeds. Inject a ThtRunner double.

Out of scope / non-goals

  • No automatic ollama pull of a missing model (hard error with guidance instead).
  • No Ollama process supervision beyond a detached start (no health-monitoring/restart loop; the next session create/restart re-runs the preflight).
  • No per-model GPU/keep-alive tuning — a single warm ping is enough to load it.
  • No change to how tht search/tht memory call embeddings; this only guarantees Ollama is ready before the session starts.