From c2f8dd5af9d768d25488fe531530a92bb027463a Mon Sep 17 00:00:00 2001 From: mptyl Date: Mon, 29 Jun 2026 18:53:03 +0200 Subject: [PATCH] docs(spec): Ollama ensure (embeddings preflight) design Hard-fail preflight at session create/restart: ensure Ollama up + warm the configured embedding model, refuse the session if embeddings unavailable. Parameterized ollama bin/start_cmd; tht ollama ensure command + backend 503. Co-Authored-By: Claude Opus 4.8 --- .../specs/2026-06-29-ollama-ensure-design.md | 139 ++++++++++++++++++ 1 file changed, 139 insertions(+) create mode 100644 docs/superpowers/specs/2026-06-29-ollama-ensure-design.md diff --git a/docs/superpowers/specs/2026-06-29-ollama-ensure-design.md b/docs/superpowers/specs/2026-06-29-ollama-ensure-design.md new file mode 100644 index 00000000..b3da6e86 --- /dev/null +++ b/docs/superpowers/specs/2026-06-29-ollama-ensure-design.md @@ -0,0 +1,139 @@ +# Ollama Ensure (embeddings preflight) — Design + +**Date:** 2026-06-29 +**Status:** Approved (design), pending implementation plan +**Layers:** harness (`tht` CLI), backend (Fastify). No frontend code beyond an existing toast path. + +## Problem + +The NL→SQL workflow depends on **Ollama embeddings** at session time: `tht search find` +(Phases 1/4) and `tht memory search` (Phase 2) call `OllamaEmbeddings` against +`{base_url}/api/embed` with the workspace's configured model (e.g. `nomic-embed-text-v2-moe` +for `psd`, at `http://localhost:11434`). If Ollama is **off**, or the embedding model is not +loaded, these calls fail mid-session with `EmbeddingsError`. + +**Embeddings are a hard requirement — the system cannot work without them.** So at session +**creation or restart** the system must: ensure Ollama is running (start it if down), confirm +the embedding model is installed and warm it into memory, and **refuse to start the session +(hard error) whenever embeddings cannot be made available** for any reason. + +## Decisions (from brainstorming) + +- **Mechanism:** invoke the **`ollama` CLI**, but **parameterized** — the binary and the + start command are configurable (other contexts, e.g. Docker, differ). The base URL is the + existing `THT_OLLAMA_URL` / `embeddings.base_url`. +- **Blocking with timeout:** session create/restart **waits** for Ollama to become reachable + and the model to warm, up to a configurable timeout (default 60s). +- **Hard-fail (NOT degrade):** on the timeout, or any other reason embeddings are unavailable, + the preflight **fails** and the session is **refused** — no Pi process is spawned. There is + no "degraded" session. +- **Warm-only model load:** the model is assumed already installed (`ollama ls` shows it); + "load" = warm it into memory via an embed ping. If it is **not** installed → hard error with + guidance (`ollama pull `), **no automatic pull** (a multi-GB sync download would blow + the timeout and the name may not be registry-pullable). +- **Ownership:** a single deterministic `tht ollama ensure` command in the harness (which owns + the embeddings config + the `OllamaEmbeddings` client); the backend calls it as a preflight. + Rejected: doing it in the backend (it does not know the model name — it lives in the + workspace YAML) or splitting start/warm across layers (needless coordination). + +## Failure matrix (`tht ollama ensure`) + +| Condition | Result | +|---|---| +| Workspace has **no** `embeddings` config | **ERROR** — "system requires embeddings; workspace not configured" | +| Ollama unreachable and not startable within the timeout | **ERROR** | +| Configured model not installed in Ollama | **ERROR** — "run `ollama pull ` / import it" | +| Warm fails (embed ping errors) | **ERROR** | +| Server up (or started) + model present + warm ok | **OK** | + +`ERROR` ⇒ exit code ≠ 0 ⇒ backend refuses the session. `OK` ⇒ exit 0. + +## Harness layer + +### Config — `EmbeddingsConfig` ([config.py](../../../harness/tht/config.py)) +Add two optional fields (defaults make existing workspace YAMLs work unchanged): +- `bin: str = "ollama"` — the CLI binary (path-overridable per context). +- `start_cmd: list[str] | None = None` — command to start the server. When `None`, default to + `[bin, "serve"]` at use-time. Set explicitly in YAML for other contexts, e.g. + `["docker", "start", "ollama"]`. Set to `[]` to **disable auto-start** (remote Ollama: probe + only, never try to start — still hard-errors if unreachable). + +YAML supports `${VAR}` expansion already, so these can reference env if needed. + +### New command — `tht ollama ensure` +A new `ollama` Typer sub-app (registered in [cli/__init__.py](../../../harness/tht/cli/__init__.py) +alongside `session_app`/`vector_app`). + +``` +tht ollama ensure [--timeout 60] [--no-start] [--json] -c +``` +(`-c` is the per-command CONFIG_OPT, appended after the subcommand — see gotchas.) + +Steps (each step is separately diagnosable so the error message names the exact failure): +1. Load config. **If `cfg.embeddings is None` → ERROR** (exit ≠ 0). +2. **Probe** the server: `GET {base_url}/api/tags` (short timeout). Reachable → go to 4. +3. **Start** (unless `--no-start` or `start_cmd == []`): spawn the resolved start command + **detached** (`subprocess.Popen`, `start_new_session=True`, output to devnull/log so it + outlives `tht`). Poll `/api/tags` until reachable or the `--timeout` elapses. Still + unreachable → **ERROR**. +4. **Model present?** Match `cfg.embeddings.model` against the installed models from + `/api/tags` (allow the implicit `:latest` tag). Absent → **ERROR** (guidance to pull/import). +5. **Warm:** `OllamaEmbeddings(cfg).embed_query("ping")` (loads the model into memory). Raises + `EmbeddingsError` → **ERROR**. +6. Success. `--json` (pristine stdout) emits + `{ "ok": true, "server": "up" | "started", "model": "warmed", "model_name": "" }`. + On any ERROR with `--json`, emit `{ "ok": false, "stage": "", + "error": "" }` on stdout **and** exit ≠ 0. The exit code is authoritative (the + backend maps exit 0 → `ok:true`, non-zero → `ok:false`) and merges the parsed JSON for the + `stage`/`error` detail. + +The probe and the start are isolated helpers (pure-ish, injectable) so tests can drive them +without a real Ollama: a `_probe(base_url) -> bool`, a `_installed_models(base_url) -> set[str]`, +and the start spawn behind a seam. + +## Backend layer (Fastify) + +### `ThtRunner` ([tht-runner.ts](../../../backend/src/tht/tht-runner.ts)) +- `ollamaEnsure(workspace: string, timeoutSec: number): Promise<{ ok: boolean; stage?: string; error?: string }>` + — shells `tht ollama ensure --json --timeout ` (workspace via the existing `-c` arg). + Returns the parsed JSON; treats a non-zero exit as `{ ok: false, ... }` (parse stderr/stdout + for the message) rather than throwing, so the route controls the HTTP response. + +### Routes ([sessions.ts](../../../backend/src/routes/sessions.ts)) +The preflight runs **first**, before any session is created or resumed: +- `POST /sessions`: `const r = await tht.ollamaEnsure(settings.workspace, timeout)`. If `!r.ok` + → `reply.code(503).send({ error: r.error })` and **return** (do NOT call `sessionNew`/`spawnFor`). +- `POST /sessions/:id/resume`: same preflight **before** the existing finalized/archived guard + and `mgr.resume`. On failure → 503, no spawn. (Uses the session's workspace — the current + configured/settings workspace, consistent with how resume resolves the session.) + +### Config ([config.ts](../../../backend/src/config.ts)) +- `ollamaEnsureTimeoutMs: number` from `OLLAMA_ENSURE_TIMEOUT_MS` (default `60000`); passed to + `ollamaEnsure` as seconds. + +(The frontend already surfaces backend error responses via the existing toast/error path from +session-creation failures — no new frontend code; the 503 message reaches the user.) + +## Testing + +- **Harness (pytest):** drive `ensure` with Ollama mocked via the injectable seams — + - no `embeddings` config → exit ≠ 0, stage `config`; + - server reachable + model present → warm called (embed ping issued), exit 0, `--json` shape; + - server unreachable + `start_cmd` set → start spawned, then poll succeeds → exit 0; + - server unreachable after timeout → exit ≠ 0, stage `server`; + - model absent from `/api/tags` → exit ≠ 0, stage `model`, message names `ollama pull`; + - warm raises `EmbeddingsError` → exit ≠ 0, stage `warm`; + - `--no-start` / `start_cmd == []` + unreachable → exit ≠ 0 without spawning. + - `--json` stdout pristine on both success and error. +- **Backend (vitest):** `ollamaEnsure` builds argv `["ollama","ensure","--json","--timeout","60", ...]` + + `-c ` and maps non-zero exit to `{ ok: false }`; `POST /sessions` and `/resume` + call it **before** spawn and return **503** (no Pi spawned) when it fails, proceed when it + succeeds. Inject a `ThtRunner` double. + +## Out of scope / non-goals +- **No automatic `ollama pull`** of a missing model (hard error with guidance instead). +- **No Ollama process supervision** beyond a detached start (no health-monitoring/restart loop; + the next session create/restart re-runs the preflight). +- **No per-model GPU/keep-alive tuning** — a single warm ping is enough to load it. +- No change to how `tht search`/`tht memory` call embeddings; this only guarantees Ollama is + ready before the session starts.