# Task 7 report — mandatory Qdrant and Ollama Compose services Date: 2026-08-08 Status: completed Summary: - Added mandatory private `qdrant`, `embedding`, and `embedding-model-init` services to the base Compose stack. - Pinned Qdrant `v1.18.2` and Ollama `0.32.0` by immutable multi-arch digest. - Persisted Qdrant storage in `qdrant-data` and Ollama model cache in `embedding-models`. - Wired `core` to fixed internal semantic endpoints: - `THT_INTERNAL_QDRANT_URL=http://qdrant:6333` - `THT_INTERNAL_EMBEDDING_URL=http://embedding:11434` - `THT_INTERNAL_EMBEDDING_MODEL=qwen3-embedding:0.6b` - `THT_INTERNAL_EMBEDDING_DIMENSIONS=1024` - Removed external vector / embedding endpoint requirements from the local and server env examples. - Added an idempotent Ollama model bootstrap script that: - waits up to a bounded deadline for `/api/tags` - skips `ollama pull` when the model is already cached - pulls `qwen3-embedding:0.6b` only when needed - verifies the model appears in `/api/tags` after pull - Added optional GPU override file `deploy/compose.embedding-gpu.yaml`; base Compose remains CPU-only. - Updated `scripts/run-stack.sh` so the GPU override is included only when `THOTH_ENABLE_EMBEDDING_GPU=1`. Verification: - RED confirmed before implementation: - `./scripts/test-default-compose.sh` failed on missing required services. - `./scripts/test-unified-compose.sh` failed on missing required services. - `./scripts/test-internal-semantic-compose.sh` failed because the GPU override file did not exist. - GREEN after implementation: - `./scripts/test-default-compose.sh` - `./scripts/test-unified-compose.sh` - `./scripts/test-internal-semantic-compose.sh` - `git diff --check` - Additional shell verification: - `scripts/run-stack.sh --wait` includes only base + local Compose files by default. - `THOTH_ENABLE_EMBEDDING_GPU=1 scripts/run-stack.sh --wait` adds `deploy/compose.embedding-gpu.yaml`. Resolved image digests: - `qdrant/qdrant:v1.18.2@sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c` - `ollama/ollama:0.32.0@sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a` Self-review: - The first bootstrap-script draft depended on tools not guaranteed inside the Ollama image. This was corrected after image inspection; the final script uses only confirmed image tools (`bash`, `ollama`, `grep`) plus raw HTTP over `/dev/tcp`. - The server overlay intentionally replaces most named core volumes with bind mounts, so the unified contract was tightened to require named semantic-cache volumes there while preserving the local/base named-volume checks. Concerns: - The model bootstrap waits for Ollama readiness and verifies cache state, but the first real cold-start will still take time to download `qwen3-embedding:0.6b`. - The GPU override requests generic Docker GPU capability only; actual GPU availability remains host/runtime dependent and intentionally stays opt-in.