8.4 KiB
Internal Qdrant and Ollama Architecture Design
Status: approved on 2026-08-08
Objective
ThothII owns its semantic infrastructure. Every supported deployment includes a private Qdrant service and a private Ollama embedding service. The analytical DWH remains external and read-only; each workspace descriptor associates that DWH with one Qdrant collection used for database schema, Evidence, and approved Memory records.
Decisions
- Qdrant replaces pgvector as the only operational vector store.
- Ollama replaces workspace-selected external embedding endpoints.
- The default and required model is
qwen3-embedding:0.6bwith 1024-dimensional normalized dense embeddings and cosine distance. - One Qdrant collection belongs to one workspace. Schema, Evidence, and Memory points share that
collection and are separated by indexed payload field
kind. - Qdrant and Ollama are mandatory base-Compose services. They are not published on host ports and are reachable only from the private Compose network.
- Existing schema-v1 and schema-v2 descriptors remain readable for migration, but they are not activatable. The new operational contract is workspace schema v3.
The model choice is based on the published Qwen model card: the 0.6B model supports more than 100 languages, a 32K context window, Matryoshka dimensions up to 1024, and instruction-aware retrieval. Ollama distributes a CPU-viable quantized build and can use an exposed GPU without changing the application protocol.
References:
- https://huggingface.co/Qwen/Qwen3-Embedding-0.6B
- https://ollama.com/library/qwen3-embedding
- https://docs.ollama.com/capabilities/embeddings
- https://qdrant.tech/documentation/installation/
- https://qdrant.tech/documentation/manage-data/collections/
Target topology
browser -> frontend -> core -> external DWH
-> private Qdrant
-> private Ollama embedding
The base Compose project contains:
frontend: static React application and same-origin API proxy.core: Fastify, Pi, and the Pythonththarness.qdrant: pinned Qdrant server with persistentqdrant-datavolume.embedding: pinned Ollama server with persistentembedding-modelsvolume.embedding-model-init: bounded one-shot service that pulls and verifiesqwen3-embedding:0.6b;corestarts only after it succeeds.
qdrant and embedding use expose, not ports. The core receives installation-owned internal
URLs:
THT_INTERNAL_QDRANT_URL=http://qdrant:6333
THT_INTERNAL_EMBEDDING_URL=http://embedding:11434
THT_INTERNAL_EMBEDDING_MODEL=qwen3-embedding:0.6b
THT_INTERNAL_EMBEDDING_DIMENSIONS=1024
These are deployment facts, not workspace connector bindings. The runtime rejects non-loopback or non-Compose-service hosts when these variables are overridden for development.
An optional Linux GPU override exposes an available NVIDIA/AMD device to Ollama. The base profile must remain CPU-safe. macOS Docker remains CPU-only because Docker Desktop cannot expose the Apple GPU to an Ollama container.
Workspace schema v3
The workspace itself is the association between the external database and the internal collection:
workspace:
schema_version: 3
id: psd-clinical
name: PSD Clinical
language: it
dwh:
engine: postgres
database: postgres
schema: datawarehouse
supported_transports: [postgres_direct]
semantic_index:
vector_store:
engine: qdrant
collection: psd-clinical
dimensions: 1024
distance: cosine
embedding:
provider: ollama_internal
model: qwen3-embedding:0.6b
dimensions: 1024
llm_policy:
allowed: [zai/glm-5.2]
Invariants:
- the collection name is an explicit portable identifier;
- active workspaces cannot share a collection;
- vector and embedding dimensions are both 1024;
- distance is
cosine; - provider and model are exactly the supported internal values;
- no vector transport, vector credential, embedding URL, or embedding credential may appear in a schema-v3 descriptor or installation contract;
- DWH connectors remain installation-local and can still use the supported external DWH transports.
Schema-v1/v2 pgvector descriptors are listed as migration_required. Migration creates a reviewed
schema-v3 document; it does not copy vector data implicitly. Existing semantic data is rebuilt from
the canonical schema documents, Evidence corpus, and Memory registry.
Qdrant data model
Each point has a deterministic UUIDv5 derived from:
workspace_id + kind + record_key
The vector is the 1024-dimensional Ollama result. The payload is:
{
"workspace_id": "psd-clinical",
"kind": "schema",
"source_id": "datawarehouse.patients",
"record_key": "schema:table:datawarehouse.patients",
"content_hash": "sha256:...",
"workspace_revision": "<git commit>",
"generation": "<optional corpus generation>",
"language": "it",
"text": "...",
"metadata": {}
}
kind, source_id, content_hash, workspace_revision, and generation receive keyword payload
indexes. Queries always filter by workspace_id and an explicit allowed kind set. Upsert is
idempotent. Evidence generation deletion is an exact filtered delete. Collection creation is also
idempotent and fails closed if an existing collection has incompatible dimensions or distance.
Harness integration
The existing VectorStore port remains the workflow boundary. A QdrantVectorStore adapter maps
its operations to Qdrant REST endpoints while preserving current schema/Evidence/Memory call sites.
The existing Ollama embedding client is narrowed to the internal /api/embed contract and verifies:
- configured model exists;
- output count matches input count;
- every vector has 1024 finite numeric values;
- no remote URL or API key is accepted.
The JSONL Memory registry and persisted phase documents remain canonical. Qdrant remains a derived, rebuildable semantic index. Schema, Evidence, and Memory ingestion all use the same point builder, content hashing, and retry policy.
Readiness and failure behavior
Readiness is layered:
- Compose waits for Qdrant health.
- Compose waits for Ollama health and successful model initialization.
- Workspace activation validates the schema-v3 contract.
- Harness readiness ensures the Qdrant collection and checks its vector configuration.
- Harness embeds a bounded probe and verifies 1024 dimensions.
Failures are sanitized and fail closed:
- unavailable Qdrant ->
workspace_not_activatablebefore session persistence; - unavailable or missing Ollama model ->
model_unavailablebefore session persistence; - collection mismatch ->
semantic_index_incompatiblewithout recreating or deleting data; - embedding dimension mismatch -> no point write;
- partial batch failure -> operation reports failure and remains safe to retry.
No health response, API response, or diagnostic log exposes DWH credentials or indexed text.
Deployment and migration
The pgvector deployment path is retired:
- remove local-vector Compose overlays and pgvector bootstrap/migration services;
- remove vector PostgreSQL role and password contracts;
- remove runtime support for vector REST/SSH and external embedding URLs;
- keep only the descriptor parser and migration code needed to recognize legacy workspaces;
- update local/server manuals, examples, smoke tests, CI coupling scans, backup instructions, and release gates for four persistent stores plus Qdrant and Ollama volumes.
Qdrant backup/restore uses collection snapshots or the persistent volume according to the operator manual. Ollama model storage is a cache: it may be backed up for offline recovery but is not an application source of truth.
Acceptance criteria
- Base local and server Compose renders include healthy private
qdrantandembeddingservices. - A clean CPU-only installation downloads the model, creates a workspace collection, and embeds a probe without external vector or embedding configuration.
- GPU override uses the same API and persistent model volume.
- Schema-v3 workspaces activate; schema-v1/v2 workspaces report
migration_required. - Two workspaces cannot claim the same Qdrant collection.
- Schema, Evidence, and Memory records coexist in one collection and remain filter-isolated.
- Existing workflow behavior and persisted session contracts remain unchanged.
- Tests reject all active pgvector deployment, external vector binding, and external embedding configuration paths.