Files
ThothII/docs/plans/2026-08-08-internal-qdrant-ollama-design.md
T

211 lines
8.4 KiB
Markdown

# Internal Qdrant and Ollama Architecture Design
**Status:** approved on 2026-08-08
## Objective
ThothII owns its semantic infrastructure. Every supported deployment includes a private Qdrant
service and a private Ollama embedding service. The analytical DWH remains external and read-only;
each workspace descriptor associates that DWH with one Qdrant collection used for database schema,
Evidence, and approved Memory records.
## Decisions
- Qdrant replaces pgvector as the only operational vector store.
- Ollama replaces workspace-selected external embedding endpoints.
- The default and required model is `qwen3-embedding:0.6b` with 1024-dimensional normalized dense
embeddings and cosine distance.
- One Qdrant collection belongs to one workspace. Schema, Evidence, and Memory points share that
collection and are separated by indexed payload field `kind`.
- Qdrant and Ollama are mandatory base-Compose services. They are not published on host ports and
are reachable only from the private Compose network.
- Existing schema-v1 and schema-v2 descriptors remain readable for migration, but they are not
activatable. The new operational contract is workspace schema v3.
The model choice is based on the published Qwen model card: the 0.6B model supports more than 100
languages, a 32K context window, Matryoshka dimensions up to 1024, and instruction-aware retrieval.
Ollama distributes a CPU-viable quantized build and can use an exposed GPU without changing the
application protocol.
References:
- <https://huggingface.co/Qwen/Qwen3-Embedding-0.6B>
- <https://ollama.com/library/qwen3-embedding>
- <https://docs.ollama.com/capabilities/embeddings>
- <https://qdrant.tech/documentation/installation/>
- <https://qdrant.tech/documentation/manage-data/collections/>
## Target topology
```text
browser -> frontend -> core -> external DWH
-> private Qdrant
-> private Ollama embedding
```
The base Compose project contains:
- `frontend`: static React application and same-origin API proxy.
- `core`: Fastify, Pi, and the Python `tht` harness.
- `qdrant`: pinned Qdrant server with persistent `qdrant-data` volume.
- `embedding`: pinned Ollama server with persistent `embedding-models` volume.
- `embedding-model-init`: bounded one-shot service that pulls and verifies
`qwen3-embedding:0.6b`; `core` starts only after it succeeds.
`qdrant` and `embedding` use `expose`, not `ports`. The core receives installation-owned internal
URLs:
```text
THT_INTERNAL_QDRANT_URL=http://qdrant:6333
THT_INTERNAL_EMBEDDING_URL=http://embedding:11434
THT_INTERNAL_EMBEDDING_MODEL=qwen3-embedding:0.6b
THT_INTERNAL_EMBEDDING_DIMENSIONS=1024
```
These are deployment facts, not workspace connector bindings. The runtime rejects non-loopback or
non-Compose-service hosts when these variables are overridden for development.
An optional Linux GPU override exposes an available NVIDIA/AMD device to Ollama. The base profile
must remain CPU-safe. macOS Docker remains CPU-only because Docker Desktop cannot expose the Apple
GPU to an Ollama container.
## Workspace schema v3
The workspace itself is the association between the external database and the internal collection:
```yaml
workspace:
schema_version: 3
id: psd-clinical
name: PSD Clinical
language: it
dwh:
engine: postgres
database: postgres
schema: datawarehouse
supported_transports: [postgres_direct]
semantic_index:
vector_store:
engine: qdrant
collection: psd-clinical
dimensions: 1024
distance: cosine
embedding:
provider: ollama_internal
model: qwen3-embedding:0.6b
dimensions: 1024
llm_policy:
allowed: [zai/glm-5.2]
```
Invariants:
- the collection name is an explicit portable identifier;
- active workspaces cannot share a collection;
- vector and embedding dimensions are both 1024;
- distance is `cosine`;
- provider and model are exactly the supported internal values;
- no vector transport, vector credential, embedding URL, or embedding credential may appear in a
schema-v3 descriptor or installation contract;
- DWH connectors remain installation-local and can still use the supported external DWH transports.
Schema-v1/v2 pgvector descriptors are listed as `migration_required`. Migration creates a reviewed
schema-v3 document; it does not copy vector data implicitly. Existing semantic data is rebuilt from
the canonical schema documents, Evidence corpus, and Memory registry.
## Qdrant data model
Each point has a deterministic UUIDv5 derived from:
```text
workspace_id + kind + record_key
```
The vector is the 1024-dimensional Ollama result. The payload is:
```json
{
"workspace_id": "psd-clinical",
"kind": "schema",
"source_id": "datawarehouse.patients",
"record_key": "schema:table:datawarehouse.patients",
"content_hash": "sha256:...",
"workspace_revision": "<git commit>",
"generation": "<optional corpus generation>",
"language": "it",
"text": "...",
"metadata": {}
}
```
`kind`, `source_id`, `content_hash`, `workspace_revision`, and `generation` receive keyword payload
indexes. Queries always filter by `workspace_id` and an explicit allowed `kind` set. Upsert is
idempotent. Evidence generation deletion is an exact filtered delete. Collection creation is also
idempotent and fails closed if an existing collection has incompatible dimensions or distance.
## Harness integration
The existing `VectorStore` port remains the workflow boundary. A `QdrantVectorStore` adapter maps
its operations to Qdrant REST endpoints while preserving current schema/Evidence/Memory call sites.
The existing Ollama embedding client is narrowed to the internal `/api/embed` contract and verifies:
- configured model exists;
- output count matches input count;
- every vector has 1024 finite numeric values;
- no remote URL or API key is accepted.
The JSONL Memory registry and persisted phase documents remain canonical. Qdrant remains a derived,
rebuildable semantic index. Schema, Evidence, and Memory ingestion all use the same point builder,
content hashing, and retry policy.
## Readiness and failure behavior
Readiness is layered:
1. Compose waits for Qdrant health.
2. Compose waits for Ollama health and successful model initialization.
3. Workspace activation validates the schema-v3 contract.
4. Harness readiness ensures the Qdrant collection and checks its vector configuration.
5. Harness embeds a bounded probe and verifies 1024 dimensions.
Failures are sanitized and fail closed:
- unavailable Qdrant -> `workspace_not_activatable` before session persistence;
- unavailable or missing Ollama model -> `model_unavailable` before session persistence;
- collection mismatch -> `semantic_index_incompatible` without recreating or deleting data;
- embedding dimension mismatch -> no point write;
- partial batch failure -> operation reports failure and remains safe to retry.
No health response, API response, or diagnostic log exposes DWH credentials or indexed text.
## Deployment and migration
The pgvector deployment path is retired:
- remove local-vector Compose overlays and pgvector bootstrap/migration services;
- remove vector PostgreSQL role and password contracts;
- remove runtime support for vector REST/SSH and external embedding URLs;
- keep only the descriptor parser and migration code needed to recognize legacy workspaces;
- update local/server manuals, examples, smoke tests, CI coupling scans, backup instructions, and
release gates for four persistent stores plus Qdrant and Ollama volumes.
Qdrant backup/restore uses collection snapshots or the persistent volume according to the operator
manual. Ollama model storage is a cache: it may be backed up for offline recovery but is not an
application source of truth.
## Acceptance criteria
- Base local and server Compose renders include healthy private `qdrant` and `embedding` services.
- A clean CPU-only installation downloads the model, creates a workspace collection, and embeds a
probe without external vector or embedding configuration.
- GPU override uses the same API and persistent model volume.
- Schema-v3 workspaces activate; schema-v1/v2 workspaces report `migration_required`.
- Two workspaces cannot claim the same Qdrant collection.
- Schema, Evidence, and Memory records coexist in one collection and remain filter-isolated.
- Existing workflow behavior and persisted session contracts remain unchanged.
- Tests reject all active pgvector deployment, external vector binding, and external embedding
configuration paths.