# Ollama Ensure (embeddings preflight) — Design **Date:** 2026-06-29 **Status:** Approved (design), pending implementation plan **Layers:** harness (`tht` CLI), backend (Fastify). No frontend code beyond an existing toast path. ## Problem The NL→SQL workflow depends on **Ollama embeddings** at session time: `tht search find` (Phases 1/4) and `tht memory search` (Phase 2) call `OllamaEmbeddings` against `{base_url}/api/embed` with the workspace's configured model (e.g. `nomic-embed-text-v2-moe` for `psd`, at `http://localhost:11434`). If Ollama is **off**, or the embedding model is not loaded, these calls fail mid-session with `EmbeddingsError`. **Embeddings are a hard requirement — the system cannot work without them.** So at session **creation or restart** the system must: ensure Ollama is running (start it if down), confirm the embedding model is installed and warm it into memory, and **refuse to start the session (hard error) whenever embeddings cannot be made available** for any reason. ## Decisions (from brainstorming) - **Mechanism:** invoke the **`ollama` CLI**, but **parameterized** — the binary and the start command are configurable (other contexts, e.g. Docker, differ). The base URL is the existing `THT_OLLAMA_URL` / `embeddings.base_url`. - **Blocking with timeout:** session create/restart **waits** for Ollama to become reachable and the model to warm, up to a configurable timeout (default 60s). - **Hard-fail (NOT degrade):** on the timeout, or any other reason embeddings are unavailable, the preflight **fails** and the session is **refused** — no Pi process is spawned. There is no "degraded" session. - **Warm-only model load:** the model is assumed already installed (`ollama ls` shows it); "load" = warm it into memory via an embed ping. If it is **not** installed → hard error with guidance (`ollama pull `), **no automatic pull** (a multi-GB sync download would blow the timeout and the name may not be registry-pullable). - **Ownership:** a single deterministic `tht ollama ensure` command in the harness (which owns the embeddings config + the `OllamaEmbeddings` client); the backend calls it as a preflight. Rejected: doing it in the backend (it does not know the model name — it lives in the workspace YAML) or splitting start/warm across layers (needless coordination). ## Failure matrix (`tht ollama ensure`) | Condition | Result | |---|---| | Workspace has **no** `embeddings` config | **ERROR** — "system requires embeddings; workspace not configured" | | Ollama unreachable and not startable within the timeout | **ERROR** | | Configured model not installed in Ollama | **ERROR** — "run `ollama pull ` / import it" | | Warm fails (embed ping errors) | **ERROR** | | Server up (or started) + model present + warm ok | **OK** | `ERROR` ⇒ exit code ≠ 0 ⇒ backend refuses the session. `OK` ⇒ exit 0. ## Harness layer ### Config — `EmbeddingsConfig` ([config.py](../../../harness/tht/config.py)) Add two optional fields (defaults make existing workspace YAMLs work unchanged): - `bin: str = "ollama"` — the CLI binary (path-overridable per context). - `start_cmd: list[str] | None = None` — command to start the server. When `None`, default to `[bin, "serve"]` at use-time. Set explicitly in YAML for other contexts, e.g. `["docker", "start", "ollama"]`. Set to `[]` to **disable auto-start** (remote Ollama: probe only, never try to start — still hard-errors if unreachable). YAML supports `${VAR}` expansion already, so these can reference env if needed. ### New command — `tht ollama ensure` A new `ollama` Typer sub-app (registered in [cli/__init__.py](../../../harness/tht/cli/__init__.py) alongside `session_app`/`vector_app`). ``` tht ollama ensure [--timeout 60] [--no-start] [--json] -c ``` (`-c` is the per-command CONFIG_OPT, appended after the subcommand — see gotchas.) Steps (each step is separately diagnosable so the error message names the exact failure): 1. Load config. **If `cfg.embeddings is None` → ERROR** (exit ≠ 0). 2. **Probe** the server: `GET {base_url}/api/tags` (short timeout). Reachable → go to 4. 3. **Start** (unless `--no-start` or `start_cmd == []`): spawn the resolved start command **detached** (`subprocess.Popen`, `start_new_session=True`, output to devnull/log so it outlives `tht`). Poll `/api/tags` until reachable or the `--timeout` elapses. Still unreachable → **ERROR**. 4. **Model present?** Match `cfg.embeddings.model` against the installed models from `/api/tags` (allow the implicit `:latest` tag). Absent → **ERROR** (guidance to pull/import). 5. **Warm:** `OllamaEmbeddings(cfg).embed_query("ping")` (loads the model into memory). Raises `EmbeddingsError` → **ERROR**. 6. Success. `--json` (pristine stdout) emits `{ "ok": true, "server": "up" | "started", "model": "warmed", "model_name": "" }`. On any ERROR with `--json`, emit `{ "ok": false, "stage": "", "error": "" }` on stdout **and** exit ≠ 0. The exit code is authoritative (the backend maps exit 0 → `ok:true`, non-zero → `ok:false`) and merges the parsed JSON for the `stage`/`error` detail. The probe and the start are isolated helpers (pure-ish, injectable) so tests can drive them without a real Ollama: a `_probe(base_url) -> bool`, a `_installed_models(base_url) -> set[str]`, and the start spawn behind a seam. ## Backend layer (Fastify) ### `ThtRunner` ([tht-runner.ts](../../../backend/src/tht/tht-runner.ts)) - `ollamaEnsure(workspace: string, timeoutSec: number): Promise<{ ok: boolean; stage?: string; error?: string }>` — shells `tht ollama ensure --json --timeout ` (workspace via the existing `-c` arg). Returns the parsed JSON; treats a non-zero exit as `{ ok: false, ... }` (parse stderr/stdout for the message) rather than throwing, so the route controls the HTTP response. ### Routes ([sessions.ts](../../../backend/src/routes/sessions.ts)) The preflight runs **first**, before any session is created or resumed: - `POST /sessions`: `const r = await tht.ollamaEnsure(settings.workspace, timeout)`. If `!r.ok` → `reply.code(503).send({ error: r.error })` and **return** (do NOT call `sessionNew`/`spawnFor`). - `POST /sessions/:id/resume`: the finalized/archived guard runs **first** — a read-only session is rejected with 409 without touching Ollama. The preflight runs **after** that guard and **before** `mgr.resume`. On failure → 503, no spawn. (Uses the session's workspace — the current configured/settings workspace, consistent with how resume resolves the session.) ### Config ([config.ts](../../../backend/src/config.ts)) - `ollamaEnsureTimeoutMs: number` from `OLLAMA_ENSURE_TIMEOUT_MS` (default `60000`); passed to `ollamaEnsure` as seconds. (The frontend already surfaces backend error responses via the existing toast/error path from session-creation failures — no new frontend code; the 503 message reaches the user.) ## Testing - **Harness (pytest):** drive `ensure` with Ollama mocked via the injectable seams — - no `embeddings` config → exit ≠ 0, stage `config`; - server reachable + model present → warm called (embed ping issued), exit 0, `--json` shape; - server unreachable + `start_cmd` set → start spawned, then poll succeeds → exit 0; - server unreachable after timeout → exit ≠ 0, stage `server`; - model absent from `/api/tags` → exit ≠ 0, stage `model`, message names `ollama pull`; - warm raises `EmbeddingsError` → exit ≠ 0, stage `warm`; - `--no-start` / `start_cmd == []` + unreachable → exit ≠ 0 without spawning. - `--json` stdout pristine on both success and error. - **Backend (vitest):** `ollamaEnsure` builds argv `["ollama","ensure","--json","--timeout","60", ...]` + `-c ` and maps non-zero exit to `{ ok: false }`; `POST /sessions` and `/resume` call it **before** spawn and return **503** (no Pi spawned) when it fails, proceed when it succeeds. Inject a `ThtRunner` double. ## Out of scope / non-goals - **No automatic `ollama pull`** of a missing model (hard error with guidance instead). - **No Ollama process supervision** beyond a detached start (no health-monitoring/restart loop; the next session create/restart re-runs the preflight). - **No per-model GPU/keep-alive tuning** — a single warm ping is enough to load it. - No change to how `tht search`/`tht memory` call embeddings; this only guarantees Ollama is ready before the session starts.