8.4 KiB
Ollama Ensure (embeddings preflight) — Design
Date: 2026-06-29
Status: Approved (design), pending implementation plan
Layers: harness (tht CLI), backend (Fastify). No frontend code beyond an existing toast path.
Problem
The NL→SQL workflow depends on Ollama embeddings at session time: tht search find
(Phases 1/4) and tht memory search (Phase 2) call OllamaEmbeddings against
{base_url}/api/embed with the workspace's configured model (e.g. nomic-embed-text-v2-moe
for psd, at http://localhost:11434). If Ollama is off, or the embedding model is not
loaded, these calls fail mid-session with EmbeddingsError.
Embeddings are a hard requirement — the system cannot work without them. So at session creation or restart the system must: ensure Ollama is running (start it if down), confirm the embedding model is installed and warm it into memory, and refuse to start the session (hard error) whenever embeddings cannot be made available for any reason.
Decisions (from brainstorming)
- Mechanism: invoke the
ollamaCLI, but parameterized — the binary and the start command are configurable (other contexts, e.g. Docker, differ). The base URL is the existingTHT_OLLAMA_URL/embeddings.base_url. - Blocking with timeout: session create/restart waits for Ollama to become reachable and the model to warm, up to a configurable timeout (default 60s).
- Hard-fail (NOT degrade): on the timeout, or any other reason embeddings are unavailable, the preflight fails and the session is refused — no Pi process is spawned. There is no "degraded" session.
- Warm-only model load: the model is assumed already installed (
ollama lsshows it); "load" = warm it into memory via an embed ping. If it is not installed → hard error with guidance (ollama pull <model>), no automatic pull (a multi-GB sync download would blow the timeout and the name may not be registry-pullable). - Ownership: a single deterministic
tht ollama ensurecommand in the harness (which owns the embeddings config + theOllamaEmbeddingsclient); the backend calls it as a preflight. Rejected: doing it in the backend (it does not know the model name — it lives in the workspace YAML) or splitting start/warm across layers (needless coordination).
Failure matrix (tht ollama ensure)
| Condition | Result |
|---|---|
Workspace has no embeddings config |
ERROR — "system requires embeddings; workspace not configured" |
| Ollama unreachable and not startable within the timeout | ERROR |
| Configured model not installed in Ollama | ERROR — "run ollama pull <model> / import it" |
| Warm fails (embed ping errors) | ERROR |
| Server up (or started) + model present + warm ok | OK |
ERROR ⇒ exit code ≠ 0 ⇒ backend refuses the session. OK ⇒ exit 0.
Harness layer
Config — EmbeddingsConfig (config.py)
Add two optional fields (defaults make existing workspace YAMLs work unchanged):
bin: str = "ollama"— the CLI binary (path-overridable per context).start_cmd: list[str] | None = None— command to start the server. WhenNone, default to[bin, "serve"]at use-time. Set explicitly in YAML for other contexts, e.g.["docker", "start", "ollama"]. Set to[]to disable auto-start (remote Ollama: probe only, never try to start — still hard-errors if unreachable).
YAML supports ${VAR} expansion already, so these can reference env if needed.
New command — tht ollama ensure
A new ollama Typer sub-app (registered in cli/__init__.py
alongside session_app/vector_app).
tht ollama ensure [--timeout 60] [--no-start] [--json] -c <workspace>
(-c is the per-command CONFIG_OPT, appended after the subcommand — see gotchas.)
Steps (each step is separately diagnosable so the error message names the exact failure):
- Load config. If
cfg.embeddings is None→ ERROR (exit ≠ 0). - Probe the server:
GET {base_url}/api/tags(short timeout). Reachable → go to 4. - Start (unless
--no-startorstart_cmd == []): spawn the resolved start command detached (subprocess.Popen,start_new_session=True, output to devnull/log so it outlivestht). Poll/api/tagsuntil reachable or the--timeoutelapses. Still unreachable → ERROR. - Model present? Match
cfg.embeddings.modelagainst the installed models from/api/tags(allow the implicit:latesttag). Absent → ERROR (guidance to pull/import). - Warm:
OllamaEmbeddings(cfg).embed_query("ping")(loads the model into memory). RaisesEmbeddingsError→ ERROR. - Success.
--json(pristine stdout) emits{ "ok": true, "server": "up" | "started", "model": "warmed", "model_name": "<model>" }. On any ERROR with--json, emit{ "ok": false, "stage": "<config|server|model|warm>", "error": "<message>" }on stdout and exit ≠ 0. The exit code is authoritative (the backend maps exit 0 →ok:true, non-zero →ok:false) and merges the parsed JSON for thestage/errordetail.
The probe and the start are isolated helpers (pure-ish, injectable) so tests can drive them
without a real Ollama: a _probe(base_url) -> bool, a _installed_models(base_url) -> set[str],
and the start spawn behind a seam.
Backend layer (Fastify)
ThtRunner (tht-runner.ts)
ollamaEnsure(workspace: string, timeoutSec: number): Promise<{ ok: boolean; stage?: string; error?: string }>— shellstht ollama ensure --json --timeout <sec>(workspace via the existing-carg). Returns the parsed JSON; treats a non-zero exit as{ ok: false, ... }(parse stderr/stdout for the message) rather than throwing, so the route controls the HTTP response.
Routes (sessions.ts)
The preflight runs first, before any session is created or resumed:
POST /sessions:const r = await tht.ollamaEnsure(settings.workspace, timeout). If!r.ok→reply.code(503).send({ error: r.error })and return (do NOT callsessionNew/spawnFor).POST /sessions/:id/resume: the finalized/archived guard runs first — a read-only session is rejected with 409 without touching Ollama. The preflight runs after that guard and beforemgr.resume. On failure → 503, no spawn. (Uses the session's workspace — the current configured/settings workspace, consistent with how resume resolves the session.)
Config (config.ts)
ollamaEnsureTimeoutMs: numberfromOLLAMA_ENSURE_TIMEOUT_MS(default60000); passed toollamaEnsureas seconds.
(The frontend already surfaces backend error responses via the existing toast/error path from session-creation failures — no new frontend code; the 503 message reaches the user.)
Testing
- Harness (pytest): drive
ensurewith Ollama mocked via the injectable seams —- no
embeddingsconfig → exit ≠ 0, stageconfig; - server reachable + model present → warm called (embed ping issued), exit 0,
--jsonshape; - server unreachable +
start_cmdset → start spawned, then poll succeeds → exit 0; - server unreachable after timeout → exit ≠ 0, stage
server; - model absent from
/api/tags→ exit ≠ 0, stagemodel, message namesollama pull; - warm raises
EmbeddingsError→ exit ≠ 0, stagewarm; --no-start/start_cmd == []+ unreachable → exit ≠ 0 without spawning.--jsonstdout pristine on both success and error.
- no
- Backend (vitest):
ollamaEnsurebuilds argv["ollama","ensure","--json","--timeout","60", ...]-c <workspace>and maps non-zero exit to{ ok: false };POST /sessionsand/resumecall it before spawn and return 503 (no Pi spawned) when it fails, proceed when it succeeds. Inject aThtRunnerdouble.
Out of scope / non-goals
- No automatic
ollama pullof a missing model (hard error with guidance instead). - No Ollama process supervision beyond a detached start (no health-monitoring/restart loop; the next session create/restart re-runs the preflight).
- No per-model GPU/keep-alive tuning — a single warm ping is enough to load it.
- No change to how
tht search/tht memorycall embeddings; this only guarantees Ollama is ready before the session starts.