docs(plan): Ollama embeddings preflight implementation plan (4 tasks, TDD)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,750 @@
|
||||
# Ollama Ensure (embeddings preflight) Implementation Plan
|
||||
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||||
|
||||
**Goal:** Guarantee Ollama embeddings are available before a session starts — a hard-fail preflight that starts Ollama if down, warms the configured model, and refuses session create/restart when embeddings can't be made available.
|
||||
|
||||
**Architecture:** A deterministic `tht ollama ensure` command in the harness (which owns the embeddings config + the `OllamaEmbeddings` client) does probe → start (configurable command, detached) → poll → verify model installed → warm. The Fastify backend calls it as a preflight before spawning Pi; a non-zero exit becomes a 503 and no session is created.
|
||||
|
||||
**Tech Stack:** Python (typer, requests, subprocess, pydantic, pytest) · Node/Fastify + TypeScript (vitest).
|
||||
|
||||
## Global Constraints
|
||||
|
||||
- **Embeddings are mandatory.** Any condition that makes embeddings unavailable (no `embeddings`
|
||||
config, Ollama unreachable after timeout, model not installed, warm fails) is a **hard error**:
|
||||
the command exits non-zero and the backend refuses the session (HTTP 503, no Pi spawned). No
|
||||
"degraded" session.
|
||||
- **No automatic `ollama pull`** — a missing model is a hard error with guidance (`ollama pull <model>`).
|
||||
- **"Load" = warm-only** — load the already-installed model into memory via one embed ping.
|
||||
- **Parameterized invocation** — the ollama binary (`embeddings.bin`, default `"ollama"`) and the
|
||||
start command (`embeddings.start_cmd`, default `[bin, "serve"]`, `[]` disables auto-start) are
|
||||
config-driven; the base URL is `embeddings.base_url`.
|
||||
- **`tht`'s `-c`/`--config` is PER-COMMAND** — it follows the subcommand (`ThtRunner` appends it).
|
||||
- **`--json` output must be pristine** (only valid JSON on stdout). On error with `--json`, the
|
||||
error JSON is on stdout AND the exit code is non-zero (the exit code is authoritative).
|
||||
- **Blocking with timeout** (default 60s, backend env `OLLAMA_ENSURE_TIMEOUT_MS`): the timeout
|
||||
only bounds waiting for the server to come up; on expiry → error, not degrade.
|
||||
|
||||
---
|
||||
|
||||
## File Structure
|
||||
|
||||
**Harness (`tht`)**
|
||||
- Modify `harness/tht/config.py` — add `bin` + `start_cmd` to `EmbeddingsConfig`.
|
||||
- Create `harness/tht/cli/ollama_cmd.py` — ops (`_probe`/`_installed_models`/`_start`/`_warm`),
|
||||
the pure `ensure_ollama(...)` orchestration, and the `ollama ensure` CLI command.
|
||||
- Modify `harness/tht/cli/__init__.py` — register the `ollama` sub-app.
|
||||
- Create `harness/tests/test_ollama_ensure.py`.
|
||||
|
||||
**Backend (Fastify)**
|
||||
- Modify `backend/src/tht/tht-runner.ts` — `ollamaEnsure(workspace, timeoutSec)`.
|
||||
- Modify `backend/src/config.ts` — `ollamaEnsureTimeoutMs`.
|
||||
- Modify `backend/src/app.ts` — pass the timeout (seconds) into `sessionRoutes` deps.
|
||||
- Modify `backend/src/routes/sessions.ts` — preflight in `POST /sessions` and `POST /sessions/:id/resume`.
|
||||
- Modify `backend/test/tht-runner.test.ts`, `backend/test/routes-sessions.test.ts`.
|
||||
|
||||
---
|
||||
|
||||
## Phase A — Harness
|
||||
|
||||
### Task 1: EmbeddingsConfig fields + `ensure_ollama` orchestration
|
||||
|
||||
**Files:**
|
||||
- Modify: `harness/tht/config.py` (`EmbeddingsConfig`)
|
||||
- Create: `harness/tht/cli/ollama_cmd.py`
|
||||
- Test: `harness/tests/test_ollama_ensure.py` (create)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: `EmbeddingsConfig`, `OllamaEmbeddings` (`tht/vectorstore/embeddings.py`, `.embed_query`).
|
||||
- Produces:
|
||||
- `EmbeddingsConfig.bin: str = "ollama"`, `EmbeddingsConfig.start_cmd: list[str] | None = None`
|
||||
- `ensure_ollama(cfg, *, timeout: int, no_start: bool, probe=_probe, installed_models=_installed_models, start=_start, warm=_warm, sleep=time.sleep, clock=time.monotonic) -> dict`
|
||||
returning `{"ok": True, "server": "up"|"started", "model": "warmed", "model_name": str}` on
|
||||
success or `{"ok": False, "stage": "config"|"server"|"model"|"warm", "error": str}` on failure.
|
||||
- Module ops `_probe(base_url)->bool`, `_installed_models(base_url)->set[str]`, `_start(cmd)->None`,
|
||||
`_warm(cfg)->None`, and `_model_present(installed, model)->bool`.
|
||||
|
||||
- [ ] **Step 1: Write the failing test**
|
||||
|
||||
Create `harness/tests/test_ollama_ensure.py`:
|
||||
|
||||
```python
|
||||
"""Tests for the ensure_ollama orchestration (Ollama mocked via injected ops)."""
|
||||
from types import SimpleNamespace
|
||||
|
||||
from tht.config import EmbeddingsConfig
|
||||
from tht.cli.ollama_cmd import ensure_ollama
|
||||
|
||||
|
||||
def _cfg(**kw):
|
||||
emb = EmbeddingsConfig(base_url="http://localhost:11434", **kw)
|
||||
return SimpleNamespace(embeddings=emb)
|
||||
|
||||
|
||||
def test_no_embeddings_config_is_hard_error():
|
||||
r = ensure_ollama(SimpleNamespace(embeddings=None), timeout=5, no_start=False)
|
||||
assert r["ok"] is False and r["stage"] == "config"
|
||||
|
||||
|
||||
def test_server_up_model_present_warms_ok():
|
||||
warmed = []
|
||||
r = ensure_ollama(
|
||||
_cfg(model="nomic-embed-text-v2-moe"), timeout=5, no_start=False,
|
||||
probe=lambda url: True,
|
||||
installed_models=lambda url: {"nomic-embed-text-v2-moe:latest"},
|
||||
start=lambda cmd: (_ for _ in ()).throw(AssertionError("must not start")),
|
||||
warm=lambda cfg: warmed.append(True),
|
||||
)
|
||||
assert r == {"ok": True, "server": "up", "model": "warmed", "model_name": "nomic-embed-text-v2-moe"}
|
||||
assert warmed == [True]
|
||||
|
||||
|
||||
def test_server_down_then_started_after_poll():
|
||||
started = []
|
||||
probes = iter([False, True]) # down, then up after start
|
||||
r = ensure_ollama(
|
||||
_cfg(), timeout=5, no_start=False,
|
||||
probe=lambda url: next(probes),
|
||||
installed_models=lambda url: {"nomic-embed-text-v2-moe"},
|
||||
start=lambda cmd: started.append(cmd),
|
||||
warm=lambda cfg: None,
|
||||
sleep=lambda s: None,
|
||||
)
|
||||
assert r["ok"] is True and r["server"] == "started"
|
||||
assert started and started[0] == ["ollama", "serve"]
|
||||
|
||||
|
||||
def test_server_unreachable_after_timeout_is_error():
|
||||
clk = iter([0.0, 1.0, 2.0, 99.0]) # monotonic crosses the deadline
|
||||
r = ensure_ollama(
|
||||
_cfg(), timeout=5, no_start=False,
|
||||
probe=lambda url: False, # never comes up
|
||||
installed_models=lambda url: set(),
|
||||
start=lambda cmd: None,
|
||||
warm=lambda cfg: None,
|
||||
sleep=lambda s: None,
|
||||
clock=lambda: next(clk),
|
||||
)
|
||||
assert r["ok"] is False and r["stage"] == "server"
|
||||
|
||||
|
||||
def test_no_start_and_down_is_error_without_starting():
|
||||
r = ensure_ollama(
|
||||
_cfg(), timeout=5, no_start=True,
|
||||
probe=lambda url: False,
|
||||
installed_models=lambda url: set(),
|
||||
start=lambda cmd: (_ for _ in ()).throw(AssertionError("must not start")),
|
||||
warm=lambda cfg: None,
|
||||
)
|
||||
assert r["ok"] is False and r["stage"] == "server"
|
||||
|
||||
|
||||
def test_empty_start_cmd_disables_autostart():
|
||||
r = ensure_ollama(
|
||||
_cfg(start_cmd=[]), timeout=5, no_start=False,
|
||||
probe=lambda url: False,
|
||||
installed_models=lambda url: set(),
|
||||
start=lambda cmd: (_ for _ in ()).throw(AssertionError("must not start")),
|
||||
warm=lambda cfg: None,
|
||||
)
|
||||
assert r["ok"] is False and r["stage"] == "server"
|
||||
|
||||
|
||||
def test_model_absent_is_error_with_pull_guidance():
|
||||
r = ensure_ollama(
|
||||
_cfg(model="missing-model"), timeout=5, no_start=False,
|
||||
probe=lambda url: True,
|
||||
installed_models=lambda url: {"nomic-embed-text-v2-moe"},
|
||||
start=lambda cmd: None,
|
||||
warm=lambda cfg: None,
|
||||
)
|
||||
assert r["ok"] is False and r["stage"] == "model"
|
||||
assert "ollama pull missing-model" in r["error"]
|
||||
|
||||
|
||||
def test_warm_failure_is_error():
|
||||
r = ensure_ollama(
|
||||
_cfg(), timeout=5, no_start=False,
|
||||
probe=lambda url: True,
|
||||
installed_models=lambda url: {"nomic-embed-text-v2-moe"},
|
||||
start=lambda cmd: None,
|
||||
warm=lambda cfg: (_ for _ in ()).throw(RuntimeError("boom")),
|
||||
)
|
||||
assert r["ok"] is False and r["stage"] == "warm"
|
||||
|
||||
|
||||
def test_custom_start_cmd_used():
|
||||
started = []
|
||||
probes = iter([False, True])
|
||||
ensure_ollama(
|
||||
_cfg(bin="ollama", start_cmd=["docker", "start", "ollama"]), timeout=5, no_start=False,
|
||||
probe=lambda url: next(probes),
|
||||
installed_models=lambda url: {"nomic-embed-text-v2-moe"},
|
||||
start=lambda cmd: started.append(cmd),
|
||||
warm=lambda cfg: None,
|
||||
sleep=lambda s: None,
|
||||
)
|
||||
assert started[0] == ["docker", "start", "ollama"]
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Run test to verify it fails**
|
||||
|
||||
Run: `cd harness && .venv/bin/pytest tests/test_ollama_ensure.py -q`
|
||||
Expected: FAIL — `ModuleNotFoundError: No module named 'tht.cli.ollama_cmd'`.
|
||||
|
||||
- [ ] **Step 3: Add the config fields**
|
||||
|
||||
In `harness/tht/config.py`, in `EmbeddingsConfig`, after `timeout: int = 120` add:
|
||||
|
||||
```python
|
||||
bin: str = "ollama"
|
||||
start_cmd: list[str] | None = None
|
||||
```
|
||||
|
||||
- [ ] **Step 4: Create `ollama_cmd.py`**
|
||||
|
||||
Create `harness/tht/cli/ollama_cmd.py`:
|
||||
|
||||
```python
|
||||
"""`tht ollama` -- embeddings preflight (ensure Ollama up + model warm).
|
||||
|
||||
The system REQUIRES embeddings: any condition that makes them unavailable is a hard
|
||||
error (the caller refuses the session). "Load" = warm the already-installed model.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import subprocess
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
import typer
|
||||
|
||||
from tht.cli.config_cmd import CONFIG_OPT
|
||||
from tht.cli.schema_cmd import _load_config_or_exit
|
||||
|
||||
ollama_app = typer.Typer(help="Ollama (embeddings) -- preflight.")
|
||||
|
||||
|
||||
# --- low-level ops (real implementations; injected as fakes in tests) ----------
|
||||
|
||||
def _probe(base_url: str, timeout: float = 2.0) -> bool:
|
||||
import requests
|
||||
|
||||
try:
|
||||
return requests.get(f"{base_url.rstrip('/')}/api/tags", timeout=timeout).status_code == 200
|
||||
except requests.RequestException:
|
||||
return False
|
||||
|
||||
|
||||
def _installed_models(base_url: str, timeout: float = 5.0) -> set[str]:
|
||||
import requests
|
||||
|
||||
resp = requests.get(f"{base_url.rstrip('/')}/api/tags", timeout=timeout)
|
||||
resp.raise_for_status()
|
||||
return {m.get("name", "") for m in resp.json().get("models", [])}
|
||||
|
||||
|
||||
def _start(start_cmd: list[str]) -> None:
|
||||
# Detached so the server outlives this short-lived CLI process.
|
||||
subprocess.Popen( # noqa: S603
|
||||
start_cmd, start_new_session=True,
|
||||
stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL,
|
||||
)
|
||||
|
||||
|
||||
def _warm(cfg) -> None:
|
||||
from tht.vectorstore.embeddings import OllamaEmbeddings
|
||||
|
||||
OllamaEmbeddings(cfg.embeddings).embed_query("ping")
|
||||
|
||||
|
||||
def _model_present(installed: set[str], model: str) -> bool:
|
||||
"""Match the configured model against installed names, allowing the implicit ':latest'."""
|
||||
if model in installed:
|
||||
return True
|
||||
base = model.split(":")[0]
|
||||
return any(name == base or name.split(":")[0] == base for name in installed)
|
||||
|
||||
|
||||
# --- orchestration (pure: returns a result dict, never raises for control flow) ----
|
||||
|
||||
def ensure_ollama(
|
||||
cfg,
|
||||
*,
|
||||
timeout: int,
|
||||
no_start: bool,
|
||||
probe=_probe,
|
||||
installed_models=_installed_models,
|
||||
start=_start,
|
||||
warm=_warm,
|
||||
sleep=time.sleep,
|
||||
clock=time.monotonic,
|
||||
) -> dict:
|
||||
if cfg.embeddings is None:
|
||||
return {"ok": False, "stage": "config",
|
||||
"error": "il sistema richiede embeddings ma il workspace non li configura"}
|
||||
emb = cfg.embeddings
|
||||
base_url = emb.base_url
|
||||
start_cmd = emb.start_cmd if emb.start_cmd is not None else [emb.bin, "serve"]
|
||||
|
||||
server_state = "up"
|
||||
if not probe(base_url):
|
||||
if no_start or start_cmd == []:
|
||||
return {"ok": False, "stage": "server",
|
||||
"error": f"Ollama non raggiungibile su {base_url} e avvio disabilitato"}
|
||||
start(start_cmd)
|
||||
server_state = "started"
|
||||
deadline = clock() + timeout
|
||||
up = False
|
||||
while clock() < deadline:
|
||||
sleep(1.0)
|
||||
if probe(base_url):
|
||||
up = True
|
||||
break
|
||||
if not up:
|
||||
return {"ok": False, "stage": "server",
|
||||
"error": f"Ollama non raggiungibile su {base_url} entro {timeout}s"}
|
||||
|
||||
try:
|
||||
installed = installed_models(base_url)
|
||||
except Exception as e: # noqa: BLE001 - any read failure is a hard error
|
||||
return {"ok": False, "stage": "server",
|
||||
"error": f"impossibile leggere i modelli da {base_url}: {e}"}
|
||||
if not _model_present(installed, emb.model):
|
||||
return {"ok": False, "stage": "model",
|
||||
"error": f"modello '{emb.model}' non installato in Ollama: "
|
||||
f"esegui `ollama pull {emb.model}` o importalo"}
|
||||
|
||||
try:
|
||||
warm(cfg)
|
||||
except Exception as e: # noqa: BLE001 - warm failure is a hard error
|
||||
return {"ok": False, "stage": "warm",
|
||||
"error": f"warm del modello '{emb.model}' fallito: {e}"}
|
||||
|
||||
return {"ok": True, "server": server_state, "model": "warmed", "model_name": emb.model}
|
||||
```
|
||||
|
||||
- [ ] **Step 5: Run test to verify it passes**
|
||||
|
||||
Run: `cd harness && .venv/bin/pytest tests/test_ollama_ensure.py -q`
|
||||
Expected: PASS (9 passed). Then `cd harness && .venv/bin/ruff check tht/cli/ollama_cmd.py tht/config.py tests/test_ollama_ensure.py` → clean.
|
||||
|
||||
- [ ] **Step 6: Commit**
|
||||
|
||||
```bash
|
||||
git add harness/tht/config.py harness/tht/cli/ollama_cmd.py harness/tests/test_ollama_ensure.py
|
||||
git commit -m "feat(harness): EmbeddingsConfig bin/start_cmd + ensure_ollama preflight orchestration"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 2: `tht ollama ensure` CLI command + registration
|
||||
|
||||
**Files:**
|
||||
- Modify: `harness/tht/cli/ollama_cmd.py` (add the command)
|
||||
- Modify: `harness/tht/cli/__init__.py` (register `ollama_app`)
|
||||
- Test: `harness/tests/test_ollama_ensure.py` (append CLI tests)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: `ensure_ollama` (Task 1), `_load_config_or_exit`, `CONFIG_OPT`, `ollama_app`.
|
||||
- Produces: CLI `tht ollama ensure [--timeout N] [--no-start] [--json] -c <ws>` — exit 0 on ok,
|
||||
exit 1 on any failure; with `--json`, only the result JSON on stdout (pristine) on both paths.
|
||||
|
||||
- [ ] **Step 1: Write the failing test**
|
||||
|
||||
Append to `harness/tests/test_ollama_ensure.py`:
|
||||
|
||||
```python
|
||||
import json as _json
|
||||
|
||||
from typer.testing import CliRunner
|
||||
|
||||
from tht.cli.ollama_cmd import ollama_app
|
||||
from tht.cli import ollama_cmd
|
||||
|
||||
|
||||
def _patch(monkeypatch, result):
|
||||
monkeypatch.setattr(ollama_cmd, "_load_config_or_exit", lambda _c: SimpleNamespace(embeddings=object()))
|
||||
monkeypatch.setattr(ollama_cmd, "ensure_ollama", lambda cfg, **kw: result)
|
||||
|
||||
|
||||
def test_cli_ok_exit_zero_and_json_pristine(monkeypatch):
|
||||
_patch(monkeypatch, {"ok": True, "server": "up", "model": "warmed", "model_name": "m"})
|
||||
res = CliRunner().invoke(ollama_app, ["ensure", "--json"])
|
||||
assert res.exit_code == 0, res.output
|
||||
assert _json.loads(res.output) == {"ok": True, "server": "up", "model": "warmed", "model_name": "m"}
|
||||
|
||||
|
||||
def test_cli_error_exit_one_and_json_on_stdout(monkeypatch):
|
||||
_patch(monkeypatch, {"ok": False, "stage": "model", "error": "missing"})
|
||||
res = CliRunner().invoke(ollama_app, ["ensure", "--json"])
|
||||
assert res.exit_code == 1
|
||||
assert _json.loads(res.output) == {"ok": False, "stage": "model", "error": "missing"}
|
||||
|
||||
|
||||
def test_cli_error_human_mode_exit_one(monkeypatch):
|
||||
_patch(monkeypatch, {"ok": False, "stage": "server", "error": "down"})
|
||||
res = CliRunner().invoke(ollama_app, ["ensure"])
|
||||
assert res.exit_code == 1
|
||||
|
||||
|
||||
def test_cli_registered_on_root_app():
|
||||
from tht.cli import app # the root Typer app
|
||||
runner = CliRunner()
|
||||
res = runner.invoke(app, ["ollama", "--help"])
|
||||
assert res.exit_code == 0
|
||||
assert "ensure" in res.output
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Run test to verify it fails**
|
||||
|
||||
Run: `cd harness && .venv/bin/pytest tests/test_ollama_ensure.py -k cli -q`
|
||||
Expected: FAIL — `No such command 'ensure'` / the root app has no `ollama` command.
|
||||
|
||||
- [ ] **Step 3: Add the CLI command**
|
||||
|
||||
In `harness/tht/cli/ollama_cmd.py`, append:
|
||||
|
||||
```python
|
||||
@ollama_app.command("ensure")
|
||||
def ensure_cmd(
|
||||
timeout: int = typer.Option(60, "--timeout", help="Secondi di attesa per l'avvio di Ollama."),
|
||||
no_start: bool = typer.Option(False, "--no-start", help="Non avviare Ollama (solo verifica)."),
|
||||
json_out: bool = typer.Option(False, "--json", help="Emetti JSON puro su stdout."),
|
||||
config: Path = CONFIG_OPT,
|
||||
) -> None:
|
||||
"""Assicura Ollama attivo + modello di embedding caricato; errore se non possibile."""
|
||||
cfg = _load_config_or_exit(config)
|
||||
result = ensure_ollama(cfg, timeout=timeout, no_start=no_start)
|
||||
if json_out:
|
||||
typer.echo(json.dumps(result, ensure_ascii=False))
|
||||
elif result["ok"]:
|
||||
typer.secho(
|
||||
f"OK: Ollama {result['server']}, modello {result['model_name']} {result['model']}.",
|
||||
fg=typer.colors.GREEN,
|
||||
)
|
||||
else:
|
||||
typer.secho(f"ERRORE [{result['stage']}]: {result['error']}", fg=typer.colors.RED, err=True)
|
||||
if not result["ok"]:
|
||||
raise typer.Exit(code=1)
|
||||
```
|
||||
|
||||
- [ ] **Step 4: Register the sub-app**
|
||||
|
||||
In `harness/tht/cli/__init__.py`, add the import alongside the others:
|
||||
|
||||
```python
|
||||
from tht.cli.ollama_cmd import ollama_app # noqa: E402
|
||||
```
|
||||
|
||||
and the registration alongside the other `add_typer` calls:
|
||||
|
||||
```python
|
||||
app.add_typer(ollama_app, name="ollama")
|
||||
```
|
||||
|
||||
- [ ] **Step 5: Run test to verify it passes**
|
||||
|
||||
Run: `cd harness && .venv/bin/pytest tests/test_ollama_ensure.py -q`
|
||||
Expected: PASS (all). Then `cd harness && .venv/bin/ruff check tht/cli/ollama_cmd.py tht/cli/__init__.py` → clean.
|
||||
|
||||
- [ ] **Step 6: Commit**
|
||||
|
||||
```bash
|
||||
git add harness/tht/cli/ollama_cmd.py harness/tht/cli/__init__.py harness/tests/test_ollama_ensure.py
|
||||
git commit -m "feat(harness): tht ollama ensure CLI command (hard-fail preflight)"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Phase B — Backend
|
||||
|
||||
### Task 3: `ThtRunner.ollamaEnsure`
|
||||
|
||||
**Files:**
|
||||
- Modify: `backend/src/tht/tht-runner.ts`
|
||||
- Test: `backend/test/tht-runner.test.ts` (append)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: `ThtRunner.run(args, workspace)`.
|
||||
- Produces: `ollamaEnsure(workspace: string, timeoutSec: number): Promise<{ ok: boolean; stage?: string; error?: string; server?: string; model?: string; model_name?: string }>`
|
||||
— shells `tht ollama ensure --json --timeout <sec>` (workspace via the per-command `-c`); exit 0
|
||||
→ `{ ok: true, ...parsedJson }`, non-zero → `{ ok: false, stage, error }`.
|
||||
|
||||
- [ ] **Step 1: Write the failing test**
|
||||
|
||||
Append to `backend/test/tht-runner.test.ts`:
|
||||
|
||||
```typescript
|
||||
test("ollamaEnsure builds argv with --json --timeout and the workspace -c", async () => {
|
||||
let calledArgs: string[] = [];
|
||||
let calledWs: string | undefined;
|
||||
const r = new ThtRunner({ thtBin: "tht", harnessDir: "/nope", configPath: "config/tht.yaml" });
|
||||
r.run = async (args, ws) => { calledArgs = args; calledWs = ws; return { code: 0, stdout: '{"ok":true,"server":"up","model":"warmed","model_name":"m"}', stderr: "" }; };
|
||||
const res = await r.ollamaEnsure("psd", 60);
|
||||
expect(calledArgs).toEqual(["ollama", "ensure", "--json", "--timeout", "60"]);
|
||||
expect(calledWs).toBe("psd");
|
||||
expect(res).toEqual({ ok: true, server: "up", model: "warmed", model_name: "m" });
|
||||
});
|
||||
|
||||
test("ollamaEnsure maps a non-zero exit to ok:false with stage/error from stdout JSON", async () => {
|
||||
const r = new ThtRunner({ thtBin: "tht", harnessDir: "/nope", configPath: "config/tht.yaml" });
|
||||
r.run = async () => ({ code: 1, stdout: '{"ok":false,"stage":"model","error":"missing"}', stderr: "" });
|
||||
expect(await r.ollamaEnsure("psd", 60)).toEqual({ ok: false, stage: "model", error: "missing" });
|
||||
});
|
||||
|
||||
test("ollamaEnsure falls back to stderr when stdout is not JSON on failure", async () => {
|
||||
const r = new ThtRunner({ thtBin: "tht", harnessDir: "/nope", configPath: "config/tht.yaml" });
|
||||
r.run = async () => ({ code: 1, stdout: "", stderr: "boom" });
|
||||
const res = await r.ollamaEnsure("psd", 60);
|
||||
expect(res.ok).toBe(false);
|
||||
expect(res.error).toContain("boom");
|
||||
});
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Run test to verify it fails**
|
||||
|
||||
Run: `cd backend && npx vitest run test/tht-runner.test.ts`
|
||||
Expected: FAIL — `r.ollamaEnsure is not a function`.
|
||||
|
||||
- [ ] **Step 3: Implement the method**
|
||||
|
||||
In `backend/src/tht/tht-runner.ts`, add a result interface near `SessionDocument`:
|
||||
|
||||
```typescript
|
||||
export interface OllamaEnsureResult {
|
||||
ok: boolean;
|
||||
stage?: string;
|
||||
error?: string;
|
||||
server?: string;
|
||||
model?: string;
|
||||
model_name?: string;
|
||||
}
|
||||
```
|
||||
|
||||
and the method (after `documents`):
|
||||
|
||||
```typescript
|
||||
async ollamaEnsure(workspace: string, timeoutSec: number): Promise<OllamaEnsureResult> {
|
||||
const { code, stdout, stderr } = await this.run(
|
||||
["ollama", "ensure", "--json", "--timeout", String(timeoutSec)],
|
||||
workspace,
|
||||
);
|
||||
let parsed: Partial<OllamaEnsureResult> = {};
|
||||
try { parsed = JSON.parse(stdout.trim() || "{}"); } catch { /* leave {} */ }
|
||||
if (code === 0) return { ok: true, ...parsed };
|
||||
return {
|
||||
ok: false,
|
||||
stage: parsed.stage,
|
||||
error: parsed.error ?? (stderr.trim() || `tht ollama ensure exit ${code}`),
|
||||
};
|
||||
}
|
||||
```
|
||||
|
||||
- [ ] **Step 4: Run test to verify it passes**
|
||||
|
||||
Run: `cd backend && npx vitest run test/tht-runner.test.ts` then `npx tsc --noEmit -p .`
|
||||
Expected: PASS, typecheck clean.
|
||||
|
||||
- [ ] **Step 5: Commit**
|
||||
|
||||
```bash
|
||||
git add backend/src/tht/tht-runner.ts backend/test/tht-runner.test.ts
|
||||
git commit -m "feat(backend): ThtRunner.ollamaEnsure (parses tht ollama ensure --json)"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 4: Backend preflight in session routes + timeout config
|
||||
|
||||
**Files:**
|
||||
- Modify: `backend/src/config.ts` (`ollamaEnsureTimeoutMs`)
|
||||
- Modify: `backend/src/app.ts` (pass `ollamaEnsureTimeoutSec` to `sessionRoutes`)
|
||||
- Modify: `backend/src/routes/sessions.ts` (preflight in create + resume)
|
||||
- Test: `backend/test/routes-sessions.test.ts` (append)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: `ThtRunner.ollamaEnsure` (Task 3), `getSettings().workspace`.
|
||||
- Produces: `POST /sessions` and `POST /sessions/:id/resume` run `ollamaEnsure` first; on `!ok`
|
||||
reply **503** `{ error }` and do NOT create/resume; the `sessionRoutes` deps object gains
|
||||
`ollamaEnsureTimeoutSec: number`.
|
||||
|
||||
- [ ] **Step 1: Write the failing test**
|
||||
|
||||
Append to `backend/test/routes-sessions.test.ts`:
|
||||
|
||||
```typescript
|
||||
test("POST /sessions refuses with 503 when ollamaEnsure fails (no session created)", async () => {
|
||||
let createdCalled = false;
|
||||
const app = buildApp(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
|
||||
thtRunner: {
|
||||
ollamaEnsure: async () => ({ ok: false, stage: "model", error: "modello non installato" }),
|
||||
sessionNew: async () => { createdCalled = true; return { id: "s1" }; },
|
||||
} as any,
|
||||
getSettings: () => ({ workspace: "psd" }) as any,
|
||||
spawnFn: () => nodeSpawn("node", [FAKE, SCRIPT]) as any,
|
||||
});
|
||||
const res = await app.inject({ method: "POST", url: "/sessions", payload: { question: "q" } });
|
||||
expect(res.statusCode).toBe(503);
|
||||
expect(res.json().error).toContain("non installato");
|
||||
expect(createdCalled).toBe(false);
|
||||
});
|
||||
|
||||
test("POST /sessions proceeds when ollamaEnsure succeeds", async () => {
|
||||
let ensureWs: string | undefined;
|
||||
const app = buildApp(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
|
||||
thtRunner: {
|
||||
ollamaEnsure: async (ws: string) => { ensureWs = ws; return { ok: true }; },
|
||||
sessionNew: async () => ({ id: "s1" }),
|
||||
} as any,
|
||||
getSettings: () => ({ workspace: "psd" }) as any,
|
||||
spawnFn: () => nodeSpawn("node", [FAKE, SCRIPT]) as any,
|
||||
});
|
||||
const res = await app.inject({ method: "POST", url: "/sessions", payload: { question: "q" } });
|
||||
expect(res.json()).toEqual({ id: "s1" });
|
||||
expect(ensureWs).toBe("psd");
|
||||
});
|
||||
|
||||
test("POST /sessions/:id/resume refuses with 503 when ollamaEnsure fails", async () => {
|
||||
const app = buildApp(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
|
||||
thtRunner: {
|
||||
ollamaEnsure: async () => ({ ok: false, error: "Ollama down" }),
|
||||
sessionShow: async () => ({ status: "open", archived: false }),
|
||||
} as any,
|
||||
getSettings: () => ({ workspace: "psd" }) as any,
|
||||
spawnFn: () => nodeSpawn("node", [FAKE, SCRIPT]) as any,
|
||||
});
|
||||
const res = await app.inject({ method: "POST", url: "/sessions/s1/resume" });
|
||||
expect(res.statusCode).toBe(503);
|
||||
});
|
||||
```
|
||||
|
||||
(Note: the existing `mutApp` tests in this file do not pass `ollamaEnsure`; keep those tests
|
||||
unaffected — the routes that call `ollamaEnsure` are only create and resume, and `mutApp` is used
|
||||
for rename/group/archive/delete/documents. The two pre-existing create/resume tests at the top of
|
||||
the file DO call create/resume, so update them per Step 4's note.)
|
||||
|
||||
- [ ] **Step 2: Run test to verify it fails**
|
||||
|
||||
Run: `cd backend && npx vitest run test/routes-sessions.test.ts`
|
||||
Expected: FAIL — create/resume don't call `ollamaEnsure` (503 tests fail; and the new success test fails because `ollamaEnsure` isn't invoked).
|
||||
|
||||
- [ ] **Step 3: Add the timeout config**
|
||||
|
||||
In `backend/src/config.ts`, add to the `AppConfig` interface:
|
||||
|
||||
```typescript
|
||||
ollamaEnsureTimeoutMs: number;
|
||||
```
|
||||
|
||||
and in `loadConfig`'s returned object:
|
||||
|
||||
```typescript
|
||||
ollamaEnsureTimeoutMs: Number(env.OLLAMA_ENSURE_TIMEOUT_MS ?? 60000),
|
||||
```
|
||||
|
||||
- [ ] **Step 4: Wire the preflight into the routes**
|
||||
|
||||
In `backend/src/app.ts`, pass the timeout (seconds) into `sessionRoutes`. Change the call:
|
||||
|
||||
```typescript
|
||||
sessionRoutes(app, { mgr, tht: tht as ThtRunner, hub, getSettings });
|
||||
```
|
||||
|
||||
to:
|
||||
|
||||
```typescript
|
||||
sessionRoutes(app, {
|
||||
mgr, tht: tht as ThtRunner, hub, getSettings,
|
||||
ollamaEnsureTimeoutSec: Math.round(config.ollamaEnsureTimeoutMs / 1000),
|
||||
});
|
||||
```
|
||||
|
||||
In `backend/src/routes/sessions.ts`, extend the deps type and add the preflight. Change the
|
||||
function signature's deps to include the timeout:
|
||||
|
||||
```typescript
|
||||
export function sessionRoutes(
|
||||
app: FastifyInstance,
|
||||
d: { mgr: PiProcessManager; tht: ThtRunner; hub: SseHub; getSettings: () => Settings; ollamaEnsureTimeoutSec: number },
|
||||
) {
|
||||
```
|
||||
|
||||
At the **start** of the `POST /sessions` handler (before `d.tht.sessionNew`):
|
||||
|
||||
```typescript
|
||||
app.post("/sessions", async (req, reply) => {
|
||||
const b = req.body as { question: string; name?: string };
|
||||
const s = d.getSettings();
|
||||
const ensure = await d.tht.ollamaEnsure(s.workspace, d.ollamaEnsureTimeoutSec);
|
||||
if (!ensure.ok) return reply.code(503).send({ error: ensure.error ?? "Ollama/embeddings non disponibili" });
|
||||
// ... existing sessionNew + spawnFor unchanged ...
|
||||
```
|
||||
|
||||
At the **start** of the `POST /sessions/:id/resume` handler (before the manifest read / guard):
|
||||
|
||||
```typescript
|
||||
app.post("/sessions/:id/resume", async (req, reply) => {
|
||||
const id = (req.params as any).id;
|
||||
const ensure = await d.tht.ollamaEnsure(d.getSettings().workspace, d.ollamaEnsureTimeoutSec);
|
||||
if (!ensure.ok) return reply.code(503).send({ error: ensure.error ?? "Ollama/embeddings non disponibili" });
|
||||
// ... existing manifest read + finalized/archived guard + mgr.resume unchanged ...
|
||||
```
|
||||
|
||||
**Note — keep the pre-existing tests green.** The preflight now runs before any create/resume
|
||||
logic, so every injected `thtRunner` double that hits `POST /sessions` or `/resume` needs an
|
||||
`ollamaEnsure`, else it throws `d.tht.ollamaEnsure is not a function`. Two edits cover all cases:
|
||||
|
||||
1. **Give `mutApp` a default `ollamaEnsure`** so its tests (including the resume-guard 409 tests
|
||||
from the session-management feature) pass the preflight and still reach their logic. Change the
|
||||
helper so the injected runner is `{ ollamaEnsure: async () => ({ ok: true }), ...thtRunner }`:
|
||||
```typescript
|
||||
function mutApp(thtRunner: any) {
|
||||
return buildApp(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
|
||||
thtRunner: { ollamaEnsure: async () => ({ ok: true }), ...thtRunner },
|
||||
getSettings: () => ({ workspace: "w" }) as any,
|
||||
spawnFn: () => nodeSpawn("node", [FAKE, SCRIPT]) as any,
|
||||
});
|
||||
}
|
||||
```
|
||||
2. **Add `ollamaEnsure: async () => ({ ok: true })`** to the two top-of-file inline doubles that
|
||||
POST `/sessions`: the "POST /sessions usa i settings…" test and the
|
||||
"POST /sessions/:id/response inoltra al bridge…" test.
|
||||
|
||||
- [ ] **Step 5: Run test to verify it passes**
|
||||
|
||||
Run: `cd backend && npx vitest run test/routes-sessions.test.ts` then `npx tsc --noEmit -p .`
|
||||
Expected: PASS (new 503/success tests + the updated pre-existing tests), typecheck clean.
|
||||
|
||||
- [ ] **Step 6: Run the full backend suite (catch regressions)**
|
||||
|
||||
Run: `cd backend && npx vitest run`
|
||||
Expected: PASS. The resume-guard 409 tests from the session-management feature use `mutApp`, so the
|
||||
default `ollamaEnsure` added in Step 4 lets the preflight pass and the 409 guard is still reached.
|
||||
If any other test that POSTs create/resume was missed, give its `thtRunner` double an
|
||||
`ollamaEnsure: async () => ({ ok: true })`.
|
||||
|
||||
- [ ] **Step 7: Commit**
|
||||
|
||||
```bash
|
||||
git add backend/src/config.ts backend/src/app.ts backend/src/routes/sessions.ts backend/test/routes-sessions.test.ts
|
||||
git commit -m "feat(backend): Ollama embeddings preflight on session create/resume (503 hard-fail)"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Final verification
|
||||
|
||||
- [ ] **Harness:** `cd harness && .venv/bin/pytest tests/test_ollama_ensure.py -q` → all pass; `.venv/bin/ruff check tht/cli/ollama_cmd.py` → clean.
|
||||
- [ ] **Backend:** `cd backend && npx vitest run` → all pass; `npx tsc --noEmit -p .` → clean.
|
||||
- [ ] **Manual smoke (optional, needs the stack):** with Ollama stopped, `tht ollama ensure --json -c workspaces/psd.yaml` starts it and warms the model (exit 0); with the model uninstalled it exits 1 with `ollama pull` guidance; creating a session via the UI while Ollama is down returns a 503 with the diagnostic.
|
||||
|
||||
## Spec coverage check
|
||||
- `bin`/`start_cmd` parameterization → Task 1.
|
||||
- Hard-fail matrix (config/server/model/warm) → Task 1 (`ensure_ollama`) + Task 2 (exit codes).
|
||||
- `tht ollama ensure` command, pristine `--json`, `--no-start` → Task 2.
|
||||
- `ThtRunner.ollamaEnsure` parsing + non-zero mapping → Task 3.
|
||||
- Preflight-before-spawn, 503, no degraded session, timeout env → Task 4.
|
||||
- Warm-only / no auto-pull → enforced by `ensure_ollama` (model-absent = error) — Task 1.
|
||||
- Frontend: none (the 503 reaches the existing error path) — no task, by design.
|
||||
Reference in New Issue
Block a user