2.9 KiB
2.9 KiB
Task 7 report — mandatory Qdrant and Ollama Compose services
Date: 2026-08-08
Status: completed
Summary:
- Added mandatory private
qdrant,embedding, andembedding-model-initservices to the base Compose stack. - Pinned Qdrant
v1.18.2and Ollama0.32.0by immutable multi-arch digest. - Persisted Qdrant storage in
qdrant-dataand Ollama model cache inembedding-models. - Wired
coreto fixed internal semantic endpoints:THT_INTERNAL_QDRANT_URL=http://qdrant:6333THT_INTERNAL_EMBEDDING_URL=http://embedding:11434THT_INTERNAL_EMBEDDING_MODEL=qwen3-embedding:0.6bTHT_INTERNAL_EMBEDDING_DIMENSIONS=1024
- Removed external vector / embedding endpoint requirements from the local and server env examples.
- Added an idempotent Ollama model bootstrap script that:
- waits up to a bounded deadline for
/api/tags - skips
ollama pullwhen the model is already cached - pulls
qwen3-embedding:0.6bonly when needed - verifies the model appears in
/api/tagsafter pull
- waits up to a bounded deadline for
- Added optional GPU override file
deploy/compose.embedding-gpu.yaml; base Compose remains CPU-only. - Updated
scripts/run-stack.shso the GPU override is included only whenTHOTH_ENABLE_EMBEDDING_GPU=1.
Verification:
- RED confirmed before implementation:
./scripts/test-default-compose.shfailed on missing required services../scripts/test-unified-compose.shfailed on missing required services../scripts/test-internal-semantic-compose.shfailed because the GPU override file did not exist.
- GREEN after implementation:
./scripts/test-default-compose.sh./scripts/test-unified-compose.sh./scripts/test-internal-semantic-compose.shgit diff --check
- Additional shell verification:
scripts/run-stack.sh --waitincludes only base + local Compose files by default.THOTH_ENABLE_EMBEDDING_GPU=1 scripts/run-stack.sh --waitaddsdeploy/compose.embedding-gpu.yaml.
Resolved image digests:
qdrant/qdrant:v1.18.2@sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50collama/ollama:0.32.0@sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a
Self-review:
- The first bootstrap-script draft depended on tools not guaranteed inside the Ollama image. This was corrected after image inspection; the final script uses only confirmed image tools (
bash,ollama,grep) plus raw HTTP over/dev/tcp. - The server overlay intentionally replaces most named core volumes with bind mounts, so the unified contract was tightened to require named semantic-cache volumes there while preserving the local/base named-volume checks.
Concerns:
- The model bootstrap waits for Ollama readiness and verifies cache state, but the first real cold-start will still take time to download
qwen3-embedding:0.6b. - The GPU override requests generic Docker GPU capability only; actual GPU availability remains host/runtime dependent and intentionally stays opt-in.