Files
ThothII/.superpowers/sdd/2026-08-08-internal-qdrant-ollama/task-7-report.md
T

2.9 KiB

Task 7 report — mandatory Qdrant and Ollama Compose services

Date: 2026-08-08

Status: completed

Summary:

  • Added mandatory private qdrant, embedding, and embedding-model-init services to the base Compose stack.
  • Pinned Qdrant v1.18.2 and Ollama 0.32.0 by immutable multi-arch digest.
  • Persisted Qdrant storage in qdrant-data and Ollama model cache in embedding-models.
  • Wired core to fixed internal semantic endpoints:
    • THT_INTERNAL_QDRANT_URL=http://qdrant:6333
    • THT_INTERNAL_EMBEDDING_URL=http://embedding:11434
    • THT_INTERNAL_EMBEDDING_MODEL=qwen3-embedding:0.6b
    • THT_INTERNAL_EMBEDDING_DIMENSIONS=1024
  • Removed external vector / embedding endpoint requirements from the local and server env examples.
  • Added an idempotent Ollama model bootstrap script that:
    • waits up to a bounded deadline for /api/tags
    • skips ollama pull when the model is already cached
    • pulls qwen3-embedding:0.6b only when needed
    • verifies the model appears in /api/tags after pull
  • Added optional GPU override file deploy/compose.embedding-gpu.yaml; base Compose remains CPU-only.
  • Updated scripts/run-stack.sh so the GPU override is included only when THOTH_ENABLE_EMBEDDING_GPU=1.

Verification:

  • RED confirmed before implementation:
    • ./scripts/test-default-compose.sh failed on missing required services.
    • ./scripts/test-unified-compose.sh failed on missing required services.
    • ./scripts/test-internal-semantic-compose.sh failed because the GPU override file did not exist.
  • GREEN after implementation:
    • ./scripts/test-default-compose.sh
    • ./scripts/test-unified-compose.sh
    • ./scripts/test-internal-semantic-compose.sh
    • git diff --check
  • Additional shell verification:
    • scripts/run-stack.sh --wait includes only base + local Compose files by default.
    • THOTH_ENABLE_EMBEDDING_GPU=1 scripts/run-stack.sh --wait adds deploy/compose.embedding-gpu.yaml.

Resolved image digests:

  • qdrant/qdrant:v1.18.2@sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c
  • ollama/ollama:0.32.0@sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a

Self-review:

  • The first bootstrap-script draft depended on tools not guaranteed inside the Ollama image. This was corrected after image inspection; the final script uses only confirmed image tools (bash, ollama, grep) plus raw HTTP over /dev/tcp.
  • The server overlay intentionally replaces most named core volumes with bind mounts, so the unified contract was tightened to require named semantic-cache volumes there while preserving the local/base named-volume checks.

Concerns:

  • The model bootstrap waits for Ollama readiness and verifies cache state, but the first real cold-start will still take time to download qwen3-embedding:0.6b.
  • The GPU override requests generic Docker GPU capability only; actual GPU availability remains host/runtime dependent and intentionally stays opt-in.