feat: run qdrant and ollama inside thothii
This commit is contained in:
@@ -0,0 +1,54 @@
|
||||
# Task 7 report — mandatory Qdrant and Ollama Compose services
|
||||
|
||||
Date: 2026-08-08
|
||||
|
||||
Status: completed
|
||||
|
||||
Summary:
|
||||
|
||||
- Added mandatory private `qdrant`, `embedding`, and `embedding-model-init` services to the base Compose stack.
|
||||
- Pinned Qdrant `v1.18.2` and Ollama `0.32.0` by immutable multi-arch digest.
|
||||
- Persisted Qdrant storage in `qdrant-data` and Ollama model cache in `embedding-models`.
|
||||
- Wired `core` to fixed internal semantic endpoints:
|
||||
- `THT_INTERNAL_QDRANT_URL=http://qdrant:6333`
|
||||
- `THT_INTERNAL_EMBEDDING_URL=http://embedding:11434`
|
||||
- `THT_INTERNAL_EMBEDDING_MODEL=qwen3-embedding:0.6b`
|
||||
- `THT_INTERNAL_EMBEDDING_DIMENSIONS=1024`
|
||||
- Removed external vector / embedding endpoint requirements from the local and server env examples.
|
||||
- Added an idempotent Ollama model bootstrap script that:
|
||||
- waits up to a bounded deadline for `/api/tags`
|
||||
- skips `ollama pull` when the model is already cached
|
||||
- pulls `qwen3-embedding:0.6b` only when needed
|
||||
- verifies the model appears in `/api/tags` after pull
|
||||
- Added optional GPU override file `deploy/compose.embedding-gpu.yaml`; base Compose remains CPU-only.
|
||||
- Updated `scripts/run-stack.sh` so the GPU override is included only when `THOTH_ENABLE_EMBEDDING_GPU=1`.
|
||||
|
||||
Verification:
|
||||
|
||||
- RED confirmed before implementation:
|
||||
- `./scripts/test-default-compose.sh` failed on missing required services.
|
||||
- `./scripts/test-unified-compose.sh` failed on missing required services.
|
||||
- `./scripts/test-internal-semantic-compose.sh` failed because the GPU override file did not exist.
|
||||
- GREEN after implementation:
|
||||
- `./scripts/test-default-compose.sh`
|
||||
- `./scripts/test-unified-compose.sh`
|
||||
- `./scripts/test-internal-semantic-compose.sh`
|
||||
- `git diff --check`
|
||||
- Additional shell verification:
|
||||
- `scripts/run-stack.sh --wait` includes only base + local Compose files by default.
|
||||
- `THOTH_ENABLE_EMBEDDING_GPU=1 scripts/run-stack.sh --wait` adds `deploy/compose.embedding-gpu.yaml`.
|
||||
|
||||
Resolved image digests:
|
||||
|
||||
- `qdrant/qdrant:v1.18.2@sha256:75eab8c4ba42096724fdcfde8b4de0b5713d529dde32f285a1f86fdcb2c9e50c`
|
||||
- `ollama/ollama:0.32.0@sha256:57f573b47f1f71ebb445789f279fe3e596a8beab182f7cf486db9205bad87c5a`
|
||||
|
||||
Self-review:
|
||||
|
||||
- The first bootstrap-script draft depended on tools not guaranteed inside the Ollama image. This was corrected after image inspection; the final script uses only confirmed image tools (`bash`, `ollama`, `grep`) plus raw HTTP over `/dev/tcp`.
|
||||
- The server overlay intentionally replaces most named core volumes with bind mounts, so the unified contract was tightened to require named semantic-cache volumes there while preserving the local/base named-volume checks.
|
||||
|
||||
Concerns:
|
||||
|
||||
- The model bootstrap waits for Ollama readiness and verifies cache state, but the first real cold-start will still take time to download `qwen3-embedding:0.6b`.
|
||||
- The GPU override requests generic Docker GPU capability only; actual GPU availability remains host/runtime dependent and intentionally stays opt-in.
|
||||
Reference in New Issue
Block a user