docs(spec): Ollama ensure (embeddings preflight) design

Hard-fail preflight at session create/restart: ensure Ollama up + warm the
configured embedding model, refuse the session if embeddings unavailable.
Parameterized ollama bin/start_cmd; tht ollama ensure command + backend 503.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
2026-06-29 18:53:03 +02:00
co-authored by Claude Opus 4.8
parent 1334c1d871
commit c2f8dd5af9
@@ -0,0 +1,139 @@
# Ollama Ensure (embeddings preflight) — Design
**Date:** 2026-06-29
**Status:** Approved (design), pending implementation plan
**Layers:** harness (`tht` CLI), backend (Fastify). No frontend code beyond an existing toast path.
## Problem
The NL→SQL workflow depends on **Ollama embeddings** at session time: `tht search find`
(Phases 1/4) and `tht memory search` (Phase 2) call `OllamaEmbeddings` against
`{base_url}/api/embed` with the workspace's configured model (e.g. `nomic-embed-text-v2-moe`
for `psd`, at `http://localhost:11434`). If Ollama is **off**, or the embedding model is not
loaded, these calls fail mid-session with `EmbeddingsError`.
**Embeddings are a hard requirement — the system cannot work without them.** So at session
**creation or restart** the system must: ensure Ollama is running (start it if down), confirm
the embedding model is installed and warm it into memory, and **refuse to start the session
(hard error) whenever embeddings cannot be made available** for any reason.
## Decisions (from brainstorming)
- **Mechanism:** invoke the **`ollama` CLI**, but **parameterized** — the binary and the
start command are configurable (other contexts, e.g. Docker, differ). The base URL is the
existing `THT_OLLAMA_URL` / `embeddings.base_url`.
- **Blocking with timeout:** session create/restart **waits** for Ollama to become reachable
and the model to warm, up to a configurable timeout (default 60s).
- **Hard-fail (NOT degrade):** on the timeout, or any other reason embeddings are unavailable,
the preflight **fails** and the session is **refused** — no Pi process is spawned. There is
no "degraded" session.
- **Warm-only model load:** the model is assumed already installed (`ollama ls` shows it);
"load" = warm it into memory via an embed ping. If it is **not** installed → hard error with
guidance (`ollama pull <model>`), **no automatic pull** (a multi-GB sync download would blow
the timeout and the name may not be registry-pullable).
- **Ownership:** a single deterministic `tht ollama ensure` command in the harness (which owns
the embeddings config + the `OllamaEmbeddings` client); the backend calls it as a preflight.
Rejected: doing it in the backend (it does not know the model name — it lives in the
workspace YAML) or splitting start/warm across layers (needless coordination).
## Failure matrix (`tht ollama ensure`)
| Condition | Result |
|---|---|
| Workspace has **no** `embeddings` config | **ERROR** — "system requires embeddings; workspace not configured" |
| Ollama unreachable and not startable within the timeout | **ERROR** |
| Configured model not installed in Ollama | **ERROR** — "run `ollama pull <model>` / import it" |
| Warm fails (embed ping errors) | **ERROR** |
| Server up (or started) + model present + warm ok | **OK** |
`ERROR` ⇒ exit code ≠ 0 ⇒ backend refuses the session. `OK` ⇒ exit 0.
## Harness layer
### Config — `EmbeddingsConfig` ([config.py](../../../harness/tht/config.py))
Add two optional fields (defaults make existing workspace YAMLs work unchanged):
- `bin: str = "ollama"` — the CLI binary (path-overridable per context).
- `start_cmd: list[str] | None = None` — command to start the server. When `None`, default to
`[bin, "serve"]` at use-time. Set explicitly in YAML for other contexts, e.g.
`["docker", "start", "ollama"]`. Set to `[]` to **disable auto-start** (remote Ollama: probe
only, never try to start — still hard-errors if unreachable).
YAML supports `${VAR}` expansion already, so these can reference env if needed.
### New command — `tht ollama ensure`
A new `ollama` Typer sub-app (registered in [cli/__init__.py](../../../harness/tht/cli/__init__.py)
alongside `session_app`/`vector_app`).
```
tht ollama ensure [--timeout 60] [--no-start] [--json] -c <workspace>
```
(`-c` is the per-command CONFIG_OPT, appended after the subcommand — see gotchas.)
Steps (each step is separately diagnosable so the error message names the exact failure):
1. Load config. **If `cfg.embeddings is None` → ERROR** (exit ≠ 0).
2. **Probe** the server: `GET {base_url}/api/tags` (short timeout). Reachable → go to 4.
3. **Start** (unless `--no-start` or `start_cmd == []`): spawn the resolved start command
**detached** (`subprocess.Popen`, `start_new_session=True`, output to devnull/log so it
outlives `tht`). Poll `/api/tags` until reachable or the `--timeout` elapses. Still
unreachable → **ERROR**.
4. **Model present?** Match `cfg.embeddings.model` against the installed models from
`/api/tags` (allow the implicit `:latest` tag). Absent → **ERROR** (guidance to pull/import).
5. **Warm:** `OllamaEmbeddings(cfg).embed_query("ping")` (loads the model into memory). Raises
`EmbeddingsError` → **ERROR**.
6. Success. `--json` (pristine stdout) emits
`{ "ok": true, "server": "up" | "started", "model": "warmed", "model_name": "<model>" }`.
On any ERROR with `--json`, emit `{ "ok": false, "stage": "<config|server|model|warm>",
"error": "<message>" }` on stdout **and** exit ≠ 0. The exit code is authoritative (the
backend maps exit 0 → `ok:true`, non-zero → `ok:false`) and merges the parsed JSON for the
`stage`/`error` detail.
The probe and the start are isolated helpers (pure-ish, injectable) so tests can drive them
without a real Ollama: a `_probe(base_url) -> bool`, a `_installed_models(base_url) -> set[str]`,
and the start spawn behind a seam.
## Backend layer (Fastify)
### `ThtRunner` ([tht-runner.ts](../../../backend/src/tht/tht-runner.ts))
- `ollamaEnsure(workspace: string, timeoutSec: number): Promise<{ ok: boolean; stage?: string; error?: string }>`
— shells `tht ollama ensure --json --timeout <sec>` (workspace via the existing `-c` arg).
Returns the parsed JSON; treats a non-zero exit as `{ ok: false, ... }` (parse stderr/stdout
for the message) rather than throwing, so the route controls the HTTP response.
### Routes ([sessions.ts](../../../backend/src/routes/sessions.ts))
The preflight runs **first**, before any session is created or resumed:
- `POST /sessions`: `const r = await tht.ollamaEnsure(settings.workspace, timeout)`. If `!r.ok`
→ `reply.code(503).send({ error: r.error })` and **return** (do NOT call `sessionNew`/`spawnFor`).
- `POST /sessions/:id/resume`: same preflight **before** the existing finalized/archived guard
and `mgr.resume`. On failure → 503, no spawn. (Uses the session's workspace — the current
configured/settings workspace, consistent with how resume resolves the session.)
### Config ([config.ts](../../../backend/src/config.ts))
- `ollamaEnsureTimeoutMs: number` from `OLLAMA_ENSURE_TIMEOUT_MS` (default `60000`); passed to
`ollamaEnsure` as seconds.
(The frontend already surfaces backend error responses via the existing toast/error path from
session-creation failures — no new frontend code; the 503 message reaches the user.)
## Testing
- **Harness (pytest):** drive `ensure` with Ollama mocked via the injectable seams —
- no `embeddings` config → exit ≠ 0, stage `config`;
- server reachable + model present → warm called (embed ping issued), exit 0, `--json` shape;
- server unreachable + `start_cmd` set → start spawned, then poll succeeds → exit 0;
- server unreachable after timeout → exit ≠ 0, stage `server`;
- model absent from `/api/tags` → exit ≠ 0, stage `model`, message names `ollama pull`;
- warm raises `EmbeddingsError` → exit ≠ 0, stage `warm`;
- `--no-start` / `start_cmd == []` + unreachable → exit ≠ 0 without spawning.
- `--json` stdout pristine on both success and error.
- **Backend (vitest):** `ollamaEnsure` builds argv `["ollama","ensure","--json","--timeout","60", ...]`
+ `-c <workspace>` and maps non-zero exit to `{ ok: false }`; `POST /sessions` and `/resume`
call it **before** spawn and return **503** (no Pi spawned) when it fails, proceed when it
succeeds. Inject a `ThtRunner` double.
## Out of scope / non-goals
- **No automatic `ollama pull`** of a missing model (hard error with guidance instead).
- **No Ollama process supervision** beyond a detached start (no health-monitoring/restart loop;
the next session create/restart re-runs the preflight).
- **No per-model GPU/keep-alive tuning** — a single warm ping is enough to load it.
- No change to how `tht search`/`tht memory` call embeddings; this only guarantees Ollama is
ready before the session starts.