Files
ThothII/docs/superpowers/plans/2026-06-29-ollama-ensure.md
T

29 KiB

Ollama Ensure (embeddings preflight) Implementation Plan

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Guarantee Ollama embeddings are available before a session starts — a hard-fail preflight that starts Ollama if down, warms the configured model, and refuses session create/restart when embeddings can't be made available.

Architecture: A deterministic tht ollama ensure command in the harness (which owns the embeddings config + the OllamaEmbeddings client) does probe → start (configurable command, detached) → poll → verify model installed → warm. The Fastify backend calls it as a preflight before spawning Pi; a non-zero exit becomes a 503 and no session is created.

Tech Stack: Python (typer, requests, subprocess, pydantic, pytest) · Node/Fastify + TypeScript (vitest).

Global Constraints

  • Embeddings are mandatory. Any condition that makes embeddings unavailable (no embeddings config, Ollama unreachable after timeout, model not installed, warm fails) is a hard error: the command exits non-zero and the backend refuses the session (HTTP 503, no Pi spawned). No "degraded" session.
  • No automatic ollama pull — a missing model is a hard error with guidance (ollama pull <model>).
  • "Load" = warm-only — load the already-installed model into memory via one embed ping.
  • Parameterized invocation — the ollama binary (embeddings.bin, default "ollama") and the start command (embeddings.start_cmd, default [bin, "serve"], [] disables auto-start) are config-driven; the base URL is embeddings.base_url.
  • tht's -c/--config is PER-COMMAND — it follows the subcommand (ThtRunner appends it).
  • --json output must be pristine (only valid JSON on stdout). On error with --json, the error JSON is on stdout AND the exit code is non-zero (the exit code is authoritative).
  • Blocking with timeout (default 60s, backend env OLLAMA_ENSURE_TIMEOUT_MS): the timeout only bounds waiting for the server to come up; on expiry → error, not degrade.

File Structure

Harness (tht)

  • Modify harness/tht/config.py — add bin + start_cmd to EmbeddingsConfig.
  • Create harness/tht/cli/ollama_cmd.py — ops (_probe/_installed_models/_start/_warm), the pure ensure_ollama(...) orchestration, and the ollama ensure CLI command.
  • Modify harness/tht/cli/__init__.py — register the ollama sub-app.
  • Create harness/tests/test_ollama_ensure.py.

Backend (Fastify)

  • Modify backend/src/tht/tht-runner.ts — ollamaEnsure(workspace, timeoutSec).
  • Modify backend/src/config.ts — ollamaEnsureTimeoutMs.
  • Modify backend/src/app.ts — pass the timeout (seconds) into sessionRoutes deps.
  • Modify backend/src/routes/sessions.ts — preflight in POST /sessions and POST /sessions/:id/resume.
  • Modify backend/test/tht-runner.test.ts, backend/test/routes-sessions.test.ts.

Phase A — Harness

Task 1: EmbeddingsConfig fields + ensure_ollama orchestration

Files:

  • Modify: harness/tht/config.py (EmbeddingsConfig)
  • Create: harness/tht/cli/ollama_cmd.py
  • Test: harness/tests/test_ollama_ensure.py (create)

Interfaces:

  • Consumes: EmbeddingsConfig, OllamaEmbeddings (tht/vectorstore/embeddings.py, .embed_query).

  • Produces:

    • EmbeddingsConfig.bin: str = "ollama", EmbeddingsConfig.start_cmd: list[str] | None = None
    • ensure_ollama(cfg, *, timeout: int, no_start: bool, probe=_probe, installed_models=_installed_models, start=_start, warm=_warm, sleep=time.sleep, clock=time.monotonic) -> dict returning {"ok": True, "server": "up"|"started", "model": "warmed", "model_name": str} on success or {"ok": False, "stage": "config"|"server"|"model"|"warm", "error": str} on failure.
    • Module ops _probe(base_url)->bool, _installed_models(base_url)->set[str], _start(cmd)->None, _warm(cfg)->None, and _model_present(installed, model)->bool.
  • Step 1: Write the failing test

Create harness/tests/test_ollama_ensure.py:

"""Tests for the ensure_ollama orchestration (Ollama mocked via injected ops)."""
from types import SimpleNamespace

from tht.config import EmbeddingsConfig
from tht.cli.ollama_cmd import ensure_ollama


def _cfg(**kw):
    emb = EmbeddingsConfig(base_url="http://localhost:11434", **kw)
    return SimpleNamespace(embeddings=emb)


def test_no_embeddings_config_is_hard_error():
    r = ensure_ollama(SimpleNamespace(embeddings=None), timeout=5, no_start=False)
    assert r["ok"] is False and r["stage"] == "config"


def test_server_up_model_present_warms_ok():
    warmed = []
    r = ensure_ollama(
        _cfg(model="nomic-embed-text-v2-moe"), timeout=5, no_start=False,
        probe=lambda url: True,
        installed_models=lambda url: {"nomic-embed-text-v2-moe:latest"},
        start=lambda cmd: (_ for _ in ()).throw(AssertionError("must not start")),
        warm=lambda cfg: warmed.append(True),
    )
    assert r == {"ok": True, "server": "up", "model": "warmed", "model_name": "nomic-embed-text-v2-moe"}
    assert warmed == [True]


def test_server_down_then_started_after_poll():
    started = []
    probes = iter([False, True])  # down, then up after start
    r = ensure_ollama(
        _cfg(), timeout=5, no_start=False,
        probe=lambda url: next(probes),
        installed_models=lambda url: {"nomic-embed-text-v2-moe"},
        start=lambda cmd: started.append(cmd),
        warm=lambda cfg: None,
        sleep=lambda s: None,
    )
    assert r["ok"] is True and r["server"] == "started"
    assert started and started[0] == ["ollama", "serve"]


def test_server_unreachable_after_timeout_is_error():
    clk = iter([0.0, 1.0, 2.0, 99.0])  # monotonic crosses the deadline
    r = ensure_ollama(
        _cfg(), timeout=5, no_start=False,
        probe=lambda url: False,           # never comes up
        installed_models=lambda url: set(),
        start=lambda cmd: None,
        warm=lambda cfg: None,
        sleep=lambda s: None,
        clock=lambda: next(clk),
    )
    assert r["ok"] is False and r["stage"] == "server"


def test_no_start_and_down_is_error_without_starting():
    r = ensure_ollama(
        _cfg(), timeout=5, no_start=True,
        probe=lambda url: False,
        installed_models=lambda url: set(),
        start=lambda cmd: (_ for _ in ()).throw(AssertionError("must not start")),
        warm=lambda cfg: None,
    )
    assert r["ok"] is False and r["stage"] == "server"


def test_empty_start_cmd_disables_autostart():
    r = ensure_ollama(
        _cfg(start_cmd=[]), timeout=5, no_start=False,
        probe=lambda url: False,
        installed_models=lambda url: set(),
        start=lambda cmd: (_ for _ in ()).throw(AssertionError("must not start")),
        warm=lambda cfg: None,
    )
    assert r["ok"] is False and r["stage"] == "server"


def test_model_absent_is_error_with_pull_guidance():
    r = ensure_ollama(
        _cfg(model="missing-model"), timeout=5, no_start=False,
        probe=lambda url: True,
        installed_models=lambda url: {"nomic-embed-text-v2-moe"},
        start=lambda cmd: None,
        warm=lambda cfg: None,
    )
    assert r["ok"] is False and r["stage"] == "model"
    assert "ollama pull missing-model" in r["error"]


def test_warm_failure_is_error():
    r = ensure_ollama(
        _cfg(), timeout=5, no_start=False,
        probe=lambda url: True,
        installed_models=lambda url: {"nomic-embed-text-v2-moe"},
        start=lambda cmd: None,
        warm=lambda cfg: (_ for _ in ()).throw(RuntimeError("boom")),
    )
    assert r["ok"] is False and r["stage"] == "warm"


def test_custom_start_cmd_used():
    started = []
    probes = iter([False, True])
    ensure_ollama(
        _cfg(bin="ollama", start_cmd=["docker", "start", "ollama"]), timeout=5, no_start=False,
        probe=lambda url: next(probes),
        installed_models=lambda url: {"nomic-embed-text-v2-moe"},
        start=lambda cmd: started.append(cmd),
        warm=lambda cfg: None,
        sleep=lambda s: None,
    )
    assert started[0] == ["docker", "start", "ollama"]
  • Step 2: Run test to verify it fails

Run: cd harness && .venv/bin/pytest tests/test_ollama_ensure.py -q Expected: FAIL — ModuleNotFoundError: No module named 'tht.cli.ollama_cmd'.

  • Step 3: Add the config fields

In harness/tht/config.py, in EmbeddingsConfig, after timeout: int = 120 add:

    bin: str = "ollama"
    start_cmd: list[str] | None = None
  • Step 4: Create ollama_cmd.py

Create harness/tht/cli/ollama_cmd.py:

"""`tht ollama` -- embeddings preflight (ensure Ollama up + model warm).

The system REQUIRES embeddings: any condition that makes them unavailable is a hard
error (the caller refuses the session). "Load" = warm the already-installed model.
"""
from __future__ import annotations

import json
import subprocess
import time
from pathlib import Path

import typer

from tht.cli.config_cmd import CONFIG_OPT
from tht.cli.schema_cmd import _load_config_or_exit

ollama_app = typer.Typer(help="Ollama (embeddings) -- preflight.")


# --- low-level ops (real implementations; injected as fakes in tests) ----------

def _probe(base_url: str, timeout: float = 2.0) -> bool:
    import requests

    try:
        return requests.get(f"{base_url.rstrip('/')}/api/tags", timeout=timeout).status_code == 200
    except requests.RequestException:
        return False


def _installed_models(base_url: str, timeout: float = 5.0) -> set[str]:
    import requests

    resp = requests.get(f"{base_url.rstrip('/')}/api/tags", timeout=timeout)
    resp.raise_for_status()
    return {m.get("name", "") for m in resp.json().get("models", [])}


def _start(start_cmd: list[str]) -> None:
    # Detached so the server outlives this short-lived CLI process.
    subprocess.Popen(  # noqa: S603
        start_cmd, start_new_session=True,
        stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL,
    )


def _warm(cfg) -> None:
    from tht.vectorstore.embeddings import OllamaEmbeddings

    OllamaEmbeddings(cfg.embeddings).embed_query("ping")


def _model_present(installed: set[str], model: str) -> bool:
    """Match the configured model against installed names, allowing the implicit ':latest'."""
    if model in installed:
        return True
    base = model.split(":")[0]
    return any(name == base or name.split(":")[0] == base for name in installed)


# --- orchestration (pure: returns a result dict, never raises for control flow) ----

def ensure_ollama(
    cfg,
    *,
    timeout: int,
    no_start: bool,
    probe=_probe,
    installed_models=_installed_models,
    start=_start,
    warm=_warm,
    sleep=time.sleep,
    clock=time.monotonic,
) -> dict:
    if cfg.embeddings is None:
        return {"ok": False, "stage": "config",
                "error": "il sistema richiede embeddings ma il workspace non li configura"}
    emb = cfg.embeddings
    base_url = emb.base_url
    start_cmd = emb.start_cmd if emb.start_cmd is not None else [emb.bin, "serve"]

    server_state = "up"
    if not probe(base_url):
        if no_start or start_cmd == []:
            return {"ok": False, "stage": "server",
                    "error": f"Ollama non raggiungibile su {base_url} e avvio disabilitato"}
        start(start_cmd)
        server_state = "started"
        deadline = clock() + timeout
        up = False
        while clock() < deadline:
            sleep(1.0)
            if probe(base_url):
                up = True
                break
        if not up:
            return {"ok": False, "stage": "server",
                    "error": f"Ollama non raggiungibile su {base_url} entro {timeout}s"}

    try:
        installed = installed_models(base_url)
    except Exception as e:  # noqa: BLE001 - any read failure is a hard error
        return {"ok": False, "stage": "server",
                "error": f"impossibile leggere i modelli da {base_url}: {e}"}
    if not _model_present(installed, emb.model):
        return {"ok": False, "stage": "model",
                "error": f"modello '{emb.model}' non installato in Ollama: "
                         f"esegui `ollama pull {emb.model}` o importalo"}

    try:
        warm(cfg)
    except Exception as e:  # noqa: BLE001 - warm failure is a hard error
        return {"ok": False, "stage": "warm",
                "error": f"warm del modello '{emb.model}' fallito: {e}"}

    return {"ok": True, "server": server_state, "model": "warmed", "model_name": emb.model}
  • Step 5: Run test to verify it passes

Run: cd harness && .venv/bin/pytest tests/test_ollama_ensure.py -q Expected: PASS (9 passed). Then cd harness && .venv/bin/ruff check tht/cli/ollama_cmd.py tht/config.py tests/test_ollama_ensure.py → clean.

  • Step 6: Commit
git add harness/tht/config.py harness/tht/cli/ollama_cmd.py harness/tests/test_ollama_ensure.py
git commit -m "feat(harness): EmbeddingsConfig bin/start_cmd + ensure_ollama preflight orchestration"

Task 2: tht ollama ensure CLI command + registration

Files:

  • Modify: harness/tht/cli/ollama_cmd.py (add the command)
  • Modify: harness/tht/cli/__init__.py (register ollama_app)
  • Test: harness/tests/test_ollama_ensure.py (append CLI tests)

Interfaces:

  • Consumes: ensure_ollama (Task 1), _load_config_or_exit, CONFIG_OPT, ollama_app.

  • Produces: CLI tht ollama ensure [--timeout N] [--no-start] [--json] -c <ws> — exit 0 on ok, exit 1 on any failure; with --json, only the result JSON on stdout (pristine) on both paths.

  • Step 1: Write the failing test

Append to harness/tests/test_ollama_ensure.py:

import json as _json

from typer.testing import CliRunner

from tht.cli.ollama_cmd import ollama_app
from tht.cli import ollama_cmd


def _patch(monkeypatch, result):
    monkeypatch.setattr(ollama_cmd, "_load_config_or_exit", lambda _c: SimpleNamespace(embeddings=object()))
    monkeypatch.setattr(ollama_cmd, "ensure_ollama", lambda cfg, **kw: result)


def test_cli_ok_exit_zero_and_json_pristine(monkeypatch):
    _patch(monkeypatch, {"ok": True, "server": "up", "model": "warmed", "model_name": "m"})
    res = CliRunner().invoke(ollama_app, ["ensure", "--json"])
    assert res.exit_code == 0, res.output
    assert _json.loads(res.output) == {"ok": True, "server": "up", "model": "warmed", "model_name": "m"}


def test_cli_error_exit_one_and_json_on_stdout(monkeypatch):
    _patch(monkeypatch, {"ok": False, "stage": "model", "error": "missing"})
    res = CliRunner().invoke(ollama_app, ["ensure", "--json"])
    assert res.exit_code == 1
    assert _json.loads(res.output) == {"ok": False, "stage": "model", "error": "missing"}


def test_cli_error_human_mode_exit_one(monkeypatch):
    _patch(monkeypatch, {"ok": False, "stage": "server", "error": "down"})
    res = CliRunner().invoke(ollama_app, ["ensure"])
    assert res.exit_code == 1


def test_cli_registered_on_root_app():
    from tht.cli import app  # the root Typer app
    runner = CliRunner()
    res = runner.invoke(app, ["ollama", "--help"])
    assert res.exit_code == 0
    assert "ensure" in res.output
  • Step 2: Run test to verify it fails

Run: cd harness && .venv/bin/pytest tests/test_ollama_ensure.py -k cli -q Expected: FAIL — No such command 'ensure' / the root app has no ollama command.

  • Step 3: Add the CLI command

In harness/tht/cli/ollama_cmd.py, append:

@ollama_app.command("ensure")
def ensure_cmd(
    timeout: int = typer.Option(60, "--timeout", help="Secondi di attesa per l'avvio di Ollama."),
    no_start: bool = typer.Option(False, "--no-start", help="Non avviare Ollama (solo verifica)."),
    json_out: bool = typer.Option(False, "--json", help="Emetti JSON puro su stdout."),
    config: Path = CONFIG_OPT,
) -> None:
    """Assicura Ollama attivo + modello di embedding caricato; errore se non possibile."""
    cfg = _load_config_or_exit(config)
    result = ensure_ollama(cfg, timeout=timeout, no_start=no_start)
    if json_out:
        typer.echo(json.dumps(result, ensure_ascii=False))
    elif result["ok"]:
        typer.secho(
            f"OK: Ollama {result['server']}, modello {result['model_name']} {result['model']}.",
            fg=typer.colors.GREEN,
        )
    else:
        typer.secho(f"ERRORE [{result['stage']}]: {result['error']}", fg=typer.colors.RED, err=True)
    if not result["ok"]:
        raise typer.Exit(code=1)
  • Step 4: Register the sub-app

In harness/tht/cli/__init__.py, add the import alongside the others:

from tht.cli.ollama_cmd import ollama_app  # noqa: E402

and the registration alongside the other add_typer calls:

app.add_typer(ollama_app, name="ollama")
  • Step 5: Run test to verify it passes

Run: cd harness && .venv/bin/pytest tests/test_ollama_ensure.py -q Expected: PASS (all). Then cd harness && .venv/bin/ruff check tht/cli/ollama_cmd.py tht/cli/__init__.py → clean.

  • Step 6: Commit
git add harness/tht/cli/ollama_cmd.py harness/tht/cli/__init__.py harness/tests/test_ollama_ensure.py
git commit -m "feat(harness): tht ollama ensure CLI command (hard-fail preflight)"

Phase B — Backend

Task 3: ThtRunner.ollamaEnsure

Files:

  • Modify: backend/src/tht/tht-runner.ts
  • Test: backend/test/tht-runner.test.ts (append)

Interfaces:

  • Consumes: ThtRunner.run(args, workspace).

  • Produces: ollamaEnsure(workspace: string, timeoutSec: number): Promise<{ ok: boolean; stage?: string; error?: string; server?: string; model?: string; model_name?: string }> — shells tht ollama ensure --json --timeout <sec> (workspace via the per-command -c); exit 0 → { ok: true, ...parsedJson }, non-zero → { ok: false, stage, error }.

  • Step 1: Write the failing test

Append to backend/test/tht-runner.test.ts:

test("ollamaEnsure builds argv with --json --timeout and the workspace -c", async () => {
  let calledArgs: string[] = [];
  let calledWs: string | undefined;
  const r = new ThtRunner({ thtBin: "tht", harnessDir: "/nope", configPath: "config/tht.yaml" });
  r.run = async (args, ws) => { calledArgs = args; calledWs = ws; return { code: 0, stdout: '{"ok":true,"server":"up","model":"warmed","model_name":"m"}', stderr: "" }; };
  const res = await r.ollamaEnsure("psd", 60);
  expect(calledArgs).toEqual(["ollama", "ensure", "--json", "--timeout", "60"]);
  expect(calledWs).toBe("psd");
  expect(res).toEqual({ ok: true, server: "up", model: "warmed", model_name: "m" });
});

test("ollamaEnsure maps a non-zero exit to ok:false with stage/error from stdout JSON", async () => {
  const r = new ThtRunner({ thtBin: "tht", harnessDir: "/nope", configPath: "config/tht.yaml" });
  r.run = async () => ({ code: 1, stdout: '{"ok":false,"stage":"model","error":"missing"}', stderr: "" });
  expect(await r.ollamaEnsure("psd", 60)).toEqual({ ok: false, stage: "model", error: "missing" });
});

test("ollamaEnsure falls back to stderr when stdout is not JSON on failure", async () => {
  const r = new ThtRunner({ thtBin: "tht", harnessDir: "/nope", configPath: "config/tht.yaml" });
  r.run = async () => ({ code: 1, stdout: "", stderr: "boom" });
  const res = await r.ollamaEnsure("psd", 60);
  expect(res.ok).toBe(false);
  expect(res.error).toContain("boom");
});
  • Step 2: Run test to verify it fails

Run: cd backend && npx vitest run test/tht-runner.test.ts Expected: FAIL — r.ollamaEnsure is not a function.

  • Step 3: Implement the method

In backend/src/tht/tht-runner.ts, add a result interface near SessionDocument:

export interface OllamaEnsureResult {
  ok: boolean;
  stage?: string;
  error?: string;
  server?: string;
  model?: string;
  model_name?: string;
}

and the method (after documents):

  async ollamaEnsure(workspace: string, timeoutSec: number): Promise<OllamaEnsureResult> {
    const { code, stdout, stderr } = await this.run(
      ["ollama", "ensure", "--json", "--timeout", String(timeoutSec)],
      workspace,
    );
    let parsed: Partial<OllamaEnsureResult> = {};
    try { parsed = JSON.parse(stdout.trim() || "{}"); } catch { /* leave {} */ }
    if (code === 0) return { ok: true, ...parsed };
    return {
      ok: false,
      stage: parsed.stage,
      error: parsed.error ?? (stderr.trim() || `tht ollama ensure exit ${code}`),
    };
  }
  • Step 4: Run test to verify it passes

Run: cd backend && npx vitest run test/tht-runner.test.ts then npx tsc --noEmit -p . Expected: PASS, typecheck clean.

  • Step 5: Commit
git add backend/src/tht/tht-runner.ts backend/test/tht-runner.test.ts
git commit -m "feat(backend): ThtRunner.ollamaEnsure (parses tht ollama ensure --json)"

Task 4: Backend preflight in session routes + timeout config

Files:

  • Modify: backend/src/config.ts (ollamaEnsureTimeoutMs)
  • Modify: backend/src/app.ts (pass ollamaEnsureTimeoutSec to sessionRoutes)
  • Modify: backend/src/routes/sessions.ts (preflight in create + resume)
  • Test: backend/test/routes-sessions.test.ts (append)

Interfaces:

  • Consumes: ThtRunner.ollamaEnsure (Task 3), getSettings().workspace.

  • Produces: POST /sessions and POST /sessions/:id/resume run ollamaEnsure first; on !ok reply 503 { error } and do NOT create/resume; the sessionRoutes deps object gains ollamaEnsureTimeoutSec: number.

  • Step 1: Write the failing test

Append to backend/test/routes-sessions.test.ts:

test("POST /sessions refuses with 503 when ollamaEnsure fails (no session created)", async () => {
  let createdCalled = false;
  const app = buildApp(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
    thtRunner: {
      ollamaEnsure: async () => ({ ok: false, stage: "model", error: "modello non installato" }),
      sessionNew: async () => { createdCalled = true; return { id: "s1" }; },
    } as any,
    getSettings: () => ({ workspace: "psd" }) as any,
    spawnFn: () => nodeSpawn("node", [FAKE, SCRIPT]) as any,
  });
  const res = await app.inject({ method: "POST", url: "/sessions", payload: { question: "q" } });
  expect(res.statusCode).toBe(503);
  expect(res.json().error).toContain("non installato");
  expect(createdCalled).toBe(false);
});

test("POST /sessions proceeds when ollamaEnsure succeeds", async () => {
  let ensureWs: string | undefined;
  const app = buildApp(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
    thtRunner: {
      ollamaEnsure: async (ws: string) => { ensureWs = ws; return { ok: true }; },
      sessionNew: async () => ({ id: "s1" }),
    } as any,
    getSettings: () => ({ workspace: "psd" }) as any,
    spawnFn: () => nodeSpawn("node", [FAKE, SCRIPT]) as any,
  });
  const res = await app.inject({ method: "POST", url: "/sessions", payload: { question: "q" } });
  expect(res.json()).toEqual({ id: "s1" });
  expect(ensureWs).toBe("psd");
});

test("POST /sessions/:id/resume refuses with 503 when ollamaEnsure fails", async () => {
  const app = buildApp(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
    thtRunner: {
      ollamaEnsure: async () => ({ ok: false, error: "Ollama down" }),
      sessionShow: async () => ({ status: "open", archived: false }),
    } as any,
    getSettings: () => ({ workspace: "psd" }) as any,
    spawnFn: () => nodeSpawn("node", [FAKE, SCRIPT]) as any,
  });
  const res = await app.inject({ method: "POST", url: "/sessions/s1/resume" });
  expect(res.statusCode).toBe(503);
});

(Note: the existing mutApp tests in this file do not pass ollamaEnsure; keep those tests unaffected — the routes that call ollamaEnsure are only create and resume, and mutApp is used for rename/group/archive/delete/documents. The two pre-existing create/resume tests at the top of the file DO call create/resume, so update them per Step 4's note.)

  • Step 2: Run test to verify it fails

Run: cd backend && npx vitest run test/routes-sessions.test.ts Expected: FAIL — create/resume don't call ollamaEnsure (503 tests fail; and the new success test fails because ollamaEnsure isn't invoked).

  • Step 3: Add the timeout config

In backend/src/config.ts, add to the AppConfig interface:

  ollamaEnsureTimeoutMs: number;

and in loadConfig's returned object:

    ollamaEnsureTimeoutMs: Number(env.OLLAMA_ENSURE_TIMEOUT_MS ?? 60000),
  • Step 4: Wire the preflight into the routes

In backend/src/app.ts, pass the timeout (seconds) into sessionRoutes. Change the call:

  sessionRoutes(app, { mgr, tht: tht as ThtRunner, hub, getSettings });

to:

  sessionRoutes(app, {
    mgr, tht: tht as ThtRunner, hub, getSettings,
    ollamaEnsureTimeoutSec: Math.round(config.ollamaEnsureTimeoutMs / 1000),
  });

In backend/src/routes/sessions.ts, extend the deps type and add the preflight. Change the function signature's deps to include the timeout:

export function sessionRoutes(
  app: FastifyInstance,
  d: { mgr: PiProcessManager; tht: ThtRunner; hub: SseHub; getSettings: () => Settings; ollamaEnsureTimeoutSec: number },
) {

At the start of the POST /sessions handler (before d.tht.sessionNew):

  app.post("/sessions", async (req, reply) => {
    const b = req.body as { question: string; name?: string };
    const s = d.getSettings();
    const ensure = await d.tht.ollamaEnsure(s.workspace, d.ollamaEnsureTimeoutSec);
    if (!ensure.ok) return reply.code(503).send({ error: ensure.error ?? "Ollama/embeddings non disponibili" });
    // ... existing sessionNew + spawnFor unchanged ...

At the start of the POST /sessions/:id/resume handler (before the manifest read / guard):

  app.post("/sessions/:id/resume", async (req, reply) => {
    const id = (req.params as any).id;
    const ensure = await d.tht.ollamaEnsure(d.getSettings().workspace, d.ollamaEnsureTimeoutSec);
    if (!ensure.ok) return reply.code(503).send({ error: ensure.error ?? "Ollama/embeddings non disponibili" });
    // ... existing manifest read + finalized/archived guard + mgr.resume unchanged ...

Note — keep the pre-existing tests green. The preflight now runs before any create/resume logic, so every injected thtRunner double that hits POST /sessions or /resume needs an ollamaEnsure, else it throws d.tht.ollamaEnsure is not a function. Two edits cover all cases:

  1. Give mutApp a default ollamaEnsure so its tests (including the resume-guard 409 tests from the session-management feature) pass the preflight and still reach their logic. Change the helper so the injected runner is { ollamaEnsure: async () => ({ ok: true }), ...thtRunner }:
    function mutApp(thtRunner: any) {
      return buildApp(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
        thtRunner: { ollamaEnsure: async () => ({ ok: true }), ...thtRunner },
        getSettings: () => ({ workspace: "w" }) as any,
        spawnFn: () => nodeSpawn("node", [FAKE, SCRIPT]) as any,
      });
    }
    
  2. Add ollamaEnsure: async () => ({ ok: true }) to the two top-of-file inline doubles that POST /sessions: the "POST /sessions usa i settings…" test and the "POST /sessions/:id/response inoltra al bridge…" test.
  • Step 5: Run test to verify it passes

Run: cd backend && npx vitest run test/routes-sessions.test.ts then npx tsc --noEmit -p . Expected: PASS (new 503/success tests + the updated pre-existing tests), typecheck clean.

  • Step 6: Run the full backend suite (catch regressions)

Run: cd backend && npx vitest run Expected: PASS. The resume-guard 409 tests from the session-management feature use mutApp, so the default ollamaEnsure added in Step 4 lets the preflight pass and the 409 guard is still reached. If any other test that POSTs create/resume was missed, give its thtRunner double an ollamaEnsure: async () => ({ ok: true }).

  • Step 7: Commit
git add backend/src/config.ts backend/src/app.ts backend/src/routes/sessions.ts backend/test/routes-sessions.test.ts
git commit -m "feat(backend): Ollama embeddings preflight on session create/resume (503 hard-fail)"

Final verification

  • Harness: cd harness && .venv/bin/pytest tests/test_ollama_ensure.py -q → all pass; .venv/bin/ruff check tht/cli/ollama_cmd.py → clean.
  • Backend: cd backend && npx vitest run → all pass; npx tsc --noEmit -p . → clean.
  • Manual smoke (optional, needs the stack): with Ollama stopped, tht ollama ensure --json -c workspaces/psd.yaml starts it and warms the model (exit 0); with the model uninstalled it exits 1 with ollama pull guidance; creating a session via the UI while Ollama is down returns a 503 with the diagnostic.

Spec coverage check

  • bin/start_cmd parameterization → Task 1.
  • Hard-fail matrix (config/server/model/warm) → Task 1 (ensure_ollama) + Task 2 (exit codes).
  • tht ollama ensure command, pristine --json, --no-start → Task 2.
  • ThtRunner.ollamaEnsure parsing + non-zero mapping → Task 3.
  • Preflight-before-spawn, 503, no degraded session, timeout env → Task 4.
  • Warm-only / no auto-pull → enforced by ensure_ollama (model-absent = error) — Task 1.
  • Frontend: none (the 503 reaches the existing error path) — no task, by design.