docs: plan internal qdrant and ollama architecture

This commit is contained in:
2026-08-08 16:23:53 +02:00
parent f36d5aefa8
commit 4fe4049a24
2 changed files with 979 additions and 0 deletions
@@ -0,0 +1,210 @@
# Internal Qdrant and Ollama Architecture Design
**Status:** approved on 2026-08-08
## Objective
ThothII owns its semantic infrastructure. Every supported deployment includes a private Qdrant
service and a private Ollama embedding service. The analytical DWH remains external and read-only;
each workspace descriptor associates that DWH with one Qdrant collection used for database schema,
Evidence, and approved Memory records.
## Decisions
- Qdrant replaces pgvector as the only operational vector store.
- Ollama replaces workspace-selected external embedding endpoints.
- The default and required model is `qwen3-embedding:0.6b` with 1024-dimensional normalized dense
embeddings and cosine distance.
- One Qdrant collection belongs to one workspace. Schema, Evidence, and Memory points share that
collection and are separated by indexed payload field `kind`.
- Qdrant and Ollama are mandatory base-Compose services. They are not published on host ports and
are reachable only from the private Compose network.
- Existing schema-v1 and schema-v2 descriptors remain readable for migration, but they are not
activatable. The new operational contract is workspace schema v3.
The model choice is based on the published Qwen model card: the 0.6B model supports more than 100
languages, a 32K context window, Matryoshka dimensions up to 1024, and instruction-aware retrieval.
Ollama distributes a CPU-viable quantized build and can use an exposed GPU without changing the
application protocol.
References:
- <https://huggingface.co/Qwen/Qwen3-Embedding-0.6B>
- <https://ollama.com/library/qwen3-embedding>
- <https://docs.ollama.com/capabilities/embeddings>
- <https://qdrant.tech/documentation/installation/>
- <https://qdrant.tech/documentation/manage-data/collections/>
## Target topology
```text
browser -> frontend -> core -> external DWH
-> private Qdrant
-> private Ollama embedding
```
The base Compose project contains:
- `frontend`: static React application and same-origin API proxy.
- `core`: Fastify, Pi, and the Python `tht` harness.
- `qdrant`: pinned Qdrant server with persistent `qdrant-data` volume.
- `embedding`: pinned Ollama server with persistent `embedding-models` volume.
- `embedding-model-init`: bounded one-shot service that pulls and verifies
`qwen3-embedding:0.6b`; `core` starts only after it succeeds.
`qdrant` and `embedding` use `expose`, not `ports`. The core receives installation-owned internal
URLs:
```text
THT_INTERNAL_QDRANT_URL=http://qdrant:6333
THT_INTERNAL_EMBEDDING_URL=http://embedding:11434
THT_INTERNAL_EMBEDDING_MODEL=qwen3-embedding:0.6b
THT_INTERNAL_EMBEDDING_DIMENSIONS=1024
```
These are deployment facts, not workspace connector bindings. The runtime rejects non-loopback or
non-Compose-service hosts when these variables are overridden for development.
An optional Linux GPU override exposes an available NVIDIA/AMD device to Ollama. The base profile
must remain CPU-safe. macOS Docker remains CPU-only because Docker Desktop cannot expose the Apple
GPU to an Ollama container.
## Workspace schema v3
The workspace itself is the association between the external database and the internal collection:
```yaml
workspace:
schema_version: 3
id: psd-clinical
name: PSD Clinical
language: it
dwh:
engine: postgres
database: postgres
schema: datawarehouse
supported_transports: [postgres_direct]
semantic_index:
vector_store:
engine: qdrant
collection: psd-clinical
dimensions: 1024
distance: cosine
embedding:
provider: ollama_internal
model: qwen3-embedding:0.6b
dimensions: 1024
llm_policy:
allowed: [zai/glm-5.2]
```
Invariants:
- the collection name is an explicit portable identifier;
- active workspaces cannot share a collection;
- vector and embedding dimensions are both 1024;
- distance is `cosine`;
- provider and model are exactly the supported internal values;
- no vector transport, vector credential, embedding URL, or embedding credential may appear in a
schema-v3 descriptor or installation contract;
- DWH connectors remain installation-local and can still use the supported external DWH transports.
Schema-v1/v2 pgvector descriptors are listed as `migration_required`. Migration creates a reviewed
schema-v3 document; it does not copy vector data implicitly. Existing semantic data is rebuilt from
the canonical schema documents, Evidence corpus, and Memory registry.
## Qdrant data model
Each point has a deterministic UUIDv5 derived from:
```text
workspace_id + kind + record_key
```
The vector is the 1024-dimensional Ollama result. The payload is:
```json
{
"workspace_id": "psd-clinical",
"kind": "schema",
"source_id": "datawarehouse.patients",
"record_key": "schema:table:datawarehouse.patients",
"content_hash": "sha256:...",
"workspace_revision": "<git commit>",
"generation": "<optional corpus generation>",
"language": "it",
"text": "...",
"metadata": {}
}
```
`kind`, `source_id`, `content_hash`, `workspace_revision`, and `generation` receive keyword payload
indexes. Queries always filter by `workspace_id` and an explicit allowed `kind` set. Upsert is
idempotent. Evidence generation deletion is an exact filtered delete. Collection creation is also
idempotent and fails closed if an existing collection has incompatible dimensions or distance.
## Harness integration
The existing `VectorStore` port remains the workflow boundary. A `QdrantVectorStore` adapter maps
its operations to Qdrant REST endpoints while preserving current schema/Evidence/Memory call sites.
The existing Ollama embedding client is narrowed to the internal `/api/embed` contract and verifies:
- configured model exists;
- output count matches input count;
- every vector has 1024 finite numeric values;
- no remote URL or API key is accepted.
The JSONL Memory registry and persisted phase documents remain canonical. Qdrant remains a derived,
rebuildable semantic index. Schema, Evidence, and Memory ingestion all use the same point builder,
content hashing, and retry policy.
## Readiness and failure behavior
Readiness is layered:
1. Compose waits for Qdrant health.
2. Compose waits for Ollama health and successful model initialization.
3. Workspace activation validates the schema-v3 contract.
4. Harness readiness ensures the Qdrant collection and checks its vector configuration.
5. Harness embeds a bounded probe and verifies 1024 dimensions.
Failures are sanitized and fail closed:
- unavailable Qdrant -> `workspace_not_activatable` before session persistence;
- unavailable or missing Ollama model -> `model_unavailable` before session persistence;
- collection mismatch -> `semantic_index_incompatible` without recreating or deleting data;
- embedding dimension mismatch -> no point write;
- partial batch failure -> operation reports failure and remains safe to retry.
No health response, API response, or diagnostic log exposes DWH credentials or indexed text.
## Deployment and migration
The pgvector deployment path is retired:
- remove local-vector Compose overlays and pgvector bootstrap/migration services;
- remove vector PostgreSQL role and password contracts;
- remove runtime support for vector REST/SSH and external embedding URLs;
- keep only the descriptor parser and migration code needed to recognize legacy workspaces;
- update local/server manuals, examples, smoke tests, CI coupling scans, backup instructions, and
release gates for four persistent stores plus Qdrant and Ollama volumes.
Qdrant backup/restore uses collection snapshots or the persistent volume according to the operator
manual. Ollama model storage is a cache: it may be backed up for offline recovery but is not an
application source of truth.
## Acceptance criteria
- Base local and server Compose renders include healthy private `qdrant` and `embedding` services.
- A clean CPU-only installation downloads the model, creates a workspace collection, and embeds a
probe without external vector or embedding configuration.
- GPU override uses the same API and persistent model volume.
- Schema-v3 workspaces activate; schema-v1/v2 workspaces report `migration_required`.
- Two workspaces cannot claim the same Qdrant collection.
- Schema, Evidence, and Memory records coexist in one collection and remain filter-isolated.
- Existing workflow behavior and persisted session contracts remain unchanged.
- Tests reject all active pgvector deployment, external vector binding, and external embedding
configuration paths.
@@ -0,0 +1,769 @@
# Internal Qdrant and Ollama Implementation Plan
> **For Claude:** REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
**Goal:** Make Qdrant and Ollama mandatory internal ThothII services while keeping the analytical
DWH external and associating each workspace with one Qdrant collection for schema, Evidence, and
Memory embeddings.
**Architecture:** Introduce workspace schema v3, preserve v1/v2 only as migration inputs, and keep
the existing harness `VectorStore` port behind a new Qdrant REST adapter. Base Compose owns Qdrant,
Ollama, their persistent volumes, and model initialization; workspace descriptors contain semantic
identity but no vector/embedding endpoints or credentials.
**Tech Stack:** TypeScript/Fastify/Zod, Python 3.12/Pydantic/requests, React 18, Docker Compose,
Qdrant REST API, Ollama `/api/embed`, Vitest, pytest.
---
## Guardrails
- Apply `@superpowers:test-driven-development` to every behavior change: add one focused failing
test, observe the expected failure, implement the minimum, and rerun the focused test.
- Do not run broad suites until the corresponding code/config changes exist; this preserves the
requested ordering while still using TDD.
- Preserve the external DWH connector contract and session persistence model.
- Do not retain an operational fallback to pgvector or an external embedding endpoint.
- Do not delete or rewrite user workspace repositories or Qdrant data. Migration is descriptor-only;
semantic data is rebuilt explicitly.
- Commit after each task only when focused tests are green.
### Task 1: Define workspace schema v3
**Files:**
- Modify: `backend/src/workspaces/schema.ts`
- Modify: `backend/src/workspaces/types.ts`
- Modify: `backend/test/workspaces-schema.test.ts`
- Modify: `backend/test/workspaces-migrate-legacy.test.ts`
- Create: `backend/src/workspaces/migrate-v2-qdrant.ts`
- Create: `backend/test/workspaces-migrate-v2-qdrant.test.ts`
**Step 1: Write the failing schema tests**
Add tests proving that schema v3 accepts only this semantic shape:
```ts
const semantic_index = {
vector_store: {
engine: "qdrant",
collection: "psd-clinical",
dimensions: 1024,
distance: "cosine",
},
embedding: {
provider: "ollama_internal",
model: "qwen3-embedding:0.6b",
dimensions: 1024,
},
};
```
Add separate rejection cases for `pgvector`, `supported_transports`, external embedding providers,
non-1024 dimensions, non-cosine distance, and unknown fields. Assert v1/v2 remain parseable as
legacy descriptors but `isOperationalWorkspace()` returns false.
**Step 2: Run the tests and verify RED**
Run:
```bash
cd backend
npx vitest run test/workspaces-schema.test.ts test/workspaces-migrate-v2-qdrant.test.ts
```
Expected: failure because schema version 3 and `migrateWorkspaceV2ToV3` do not exist.
**Step 3: Implement the minimum schema and migration**
Add `QdrantVectorStore`, `InternalEmbedding`, and `WorkspaceV3` types. Replace the operational type
guard with schema-v3-only semantics. Implement:
```ts
export function migrateWorkspaceV2ToV3(
legacy: WorkspaceV2,
collection: string,
): WorkspaceV3 {
return validateOperationalWorkspace({
workspace: { ...legacy.workspace, schema_version: 3 },
dwh: legacy.dwh,
semantic_index: {
vector_store: {
engine: "qdrant",
collection,
dimensions: 1024,
distance: "cosine",
},
embedding: {
provider: "ollama_internal",
model: "qwen3-embedding:0.6b",
dimensions: 1024,
},
},
llm_policy: legacy.llm_policy,
...(legacy.diagnostics?.dwh_rest
? { diagnostics: { dwh_rest: legacy.diagnostics.dwh_rest } }
: {}),
});
}
```
Do not copy vector/embedding diagnostics or transports.
**Step 4: Verify GREEN**
Run the command from Step 2. Expected: all selected tests pass.
**Step 5: Commit**
```bash
git add backend/src/workspaces/schema.ts backend/src/workspaces/types.ts \
backend/src/workspaces/migrate-v2-qdrant.ts backend/test/workspaces-schema.test.ts \
backend/test/workspaces-migrate-legacy.test.ts backend/test/workspaces-migrate-v2-qdrant.test.ts
git commit -m "feat: define internal semantic workspace schema"
```
### Task 2: Make collection ownership unique in the Git registry
**Files:**
- Modify: `backend/src/workspaces/registry.ts`
- Modify: `backend/src/workspaces/migrate-legacy.ts`
- Modify: `backend/test/workspace-registry.test.ts`
- Modify: `backend/test/workspaces-migrate-legacy.test.ts`
**Step 1: Write failing registry tests**
Add fixtures with two schema-v3 workspaces claiming `collection: shared`. Assert snapshot activation
fails with `workspace_invalid` and retains the previous active snapshot. Assert v1/v2 entries are
listed as `migration_required` and cannot be acquired with `acquireSessionRevision()`.
**Step 2: Verify RED**
```bash
cd backend
npx vitest run test/workspace-registry.test.ts test/workspaces-migrate-legacy.test.ts \
-t "collection|migration_required"
```
Expected: duplicate collections are currently accepted and v2 is currently operational.
**Step 3: Implement uniqueness and migration state**
During snapshot validation, build `Map<collection, workspaceId>` for operational descriptors and
raise a sanitized `workspace_invalid` error on a duplicate. Update migration output and CLI wording
to require an explicit target collection and schema v3.
**Step 4: Verify GREEN and commit**
```bash
cd backend
npx vitest run test/workspace-registry.test.ts test/workspaces-migrate-legacy.test.ts \
-t "collection|migration_required"
cd ..
git add backend/src/workspaces/registry.ts backend/src/workspaces/migrate-legacy.ts \
backend/test/workspace-registry.test.ts backend/test/workspaces-migrate-legacy.test.ts
git commit -m "feat: reserve one qdrant collection per workspace"
```
### Task 3: Remove external semantic bindings and render internal endpoints
**Files:**
- Modify: `backend/src/workspaces/contracts.ts`
- Modify: `backend/src/workspaces/bindings.ts`
- Modify: `backend/src/workspaces/runtime-renderer.ts`
- Modify: `backend/src/config.ts`
- Modify: `backend/test/workspaces-contracts.test.ts`
- Modify: `backend/test/workspaces-bindings.test.ts`
- Modify: `backend/test/workspace-runtime-renderer.test.ts`
- Modify: `backend/test/config.test.ts`
**Step 1: Write failing contract tests**
Assert schema-v3 installation contracts contain DWH variables only. Assert environment variables
matching `*_VECTOR_*`, `*_EMBEDDING_BASE_URL`, or semantic API-key suffixes are ignored/rejected.
Assert the rendered harness config always contains:
```yaml
resources:
vector:
engine: qdrant
base_url: http://qdrant:6333
collection: psd-clinical
embeddings:
provider: ollama_internal
base_url: http://embedding:11434
model: qwen3-embedding:0.6b
dimensions: 1024
```
**Step 2: Verify RED**
```bash
cd backend
npx vitest run test/workspaces-contracts.test.ts test/workspaces-bindings.test.ts \
test/workspace-runtime-renderer.test.ts test/config.test.ts
```
Expected: current contracts require external vector and embedding bindings.
**Step 3: Implement internal runtime configuration**
Add typed backend config fields with Compose defaults:
```ts
internalQdrantUrl: "http://qdrant:6333"
internalEmbeddingUrl: "http://embedding:11434"
internalEmbeddingModel: "qwen3-embedding:0.6b"
internalEmbeddingDimensions: 1024
```
Accept only `qdrant`, `embedding`, `localhost`, or loopback hosts. Keep these values out of Git
workspace descriptors, API payloads, and generated installation docs. Render them into the
ephemeral backend-owned harness config after descriptor validation.
**Step 4: Verify GREEN and commit**
Run Step 2, then:
```bash
git add backend/src/config.ts backend/src/workspaces/contracts.ts backend/src/workspaces/bindings.ts \
backend/src/workspaces/runtime-renderer.ts backend/test/config.test.ts \
backend/test/workspaces-contracts.test.ts backend/test/workspaces-bindings.test.ts \
backend/test/workspace-runtime-renderer.test.ts
git commit -m "feat: render private semantic service endpoints"
```
### Task 4: Narrow harness embedding configuration to internal Ollama
**Files:**
- Modify: `harness/tht/config.py`
- Modify: `harness/tht/config_compat.py`
- Modify: `harness/tht/vectorstore/embeddings.py`
- Modify: `harness/tht/cli/ollama_cmd.py`
- Modify: `harness/tests/test_config_resources.py`
- Create: `harness/tests/test_internal_embeddings.py`
**Step 1: Write failing embedding tests**
Use a fake `requests.Session` to prove `OllamaInternalEmbeddings.embed()` calls `/api/embed` with
model and batch input, returns 1024-dimensional finite vectors, and rejects count/dimension/NaN
mismatches. Add config tests rejecting external providers, API keys, and non-private base URLs.
**Step 2: Verify RED**
```bash
cd harness
.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py -q
```
Expected: `OllamaInternalEmbeddings` and internal-only config do not exist.
**Step 3: Implement the client**
Implement one bounded `/api/embed` request per batch:
```python
response = self._session.post(
f"{self.base_url}/api/embed",
json={"model": self.model, "input": texts},
timeout=self.timeout,
)
```
Validate response shape before returning any vector. Keep retry behavior bounded and sanitize URLs
and response bodies from raised errors.
**Step 4: Verify GREEN and commit**
```bash
cd harness
.venv/bin/pytest tests/test_internal_embeddings.py tests/test_config_resources.py -q
cd ..
git add harness/tht/config.py harness/tht/config_compat.py harness/tht/vectorstore/embeddings.py \
harness/tht/cli/ollama_cmd.py harness/tests/test_config_resources.py \
harness/tests/test_internal_embeddings.py
git commit -m "feat: use internal ollama embeddings"
```
### Task 5: Implement the Qdrant VectorStore adapter
**Files:**
- Create: `harness/tht/adapters/vector/qdrant.py`
- Modify: `harness/tht/adapters/vector/__init__.py`
- Modify: `harness/tht/ports/vector.py`
- Modify: `harness/tht/vectorstore/records.py`
- Modify: `harness/tht/vectorstore/store.py`
- Create: `harness/tests/test_qdrant_vector_store.py`
- Modify: `harness/tests/test_vector_port_contract.py`
**Step 1: Write failing adapter tests**
Test a real adapter against a deterministic fake HTTP server. Cover:
- idempotent collection create with 1024/Cosine;
- mismatch fails without delete/recreate;
- keyword payload-index creation;
- deterministic UUIDv5 point IDs;
- upsert payload for `schema`, `evidence`, and `memory`;
- query filtered by workspace and allowed kinds;
- `existing_hashes`, exact Evidence generation list/delete, and health;
- sanitized timeouts and malformed responses.
The point ID helper must satisfy:
```python
def point_id(workspace_id: str, kind: str, record_key: str) -> str:
return str(uuid5(NAMESPACE_URL, f"thothii:{workspace_id}:{kind}:{record_key}"))
```
**Step 2: Verify RED**
```bash
cd harness
.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
```
Expected: import failure for the Qdrant adapter.
**Step 3: Implement minimal REST mappings**
Use existing `requests` dependency and these endpoints:
```text
GET /collections/{collection}
PUT /collections/{collection}
PUT /collections/{collection}/index
PUT /collections/{collection}/points?wait=true
POST /collections/{collection}/points/query
POST /collections/{collection}/points/scroll
POST /collections/{collection}/points/delete?wait=true
```
Every operation must include the workspace filter even though the collection is workspace-owned.
Map Qdrant scores and payloads back into existing `VectorHit` objects.
**Step 4: Verify GREEN and commit**
```bash
cd harness
.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py -q
cd ..
git add harness/tht/adapters/vector/qdrant.py harness/tht/adapters/vector/__init__.py \
harness/tht/ports/vector.py harness/tht/vectorstore/records.py \
harness/tht/vectorstore/store.py harness/tests/test_qdrant_vector_store.py \
harness/tests/test_vector_port_contract.py
git commit -m "feat: add qdrant vector adapter"
```
### Task 6: Wire schema, Evidence, and Memory through Qdrant
**Files:**
- Modify: `harness/tht/vectorstore/reader.py`
- Modify: `harness/tht/cli/vector_cmd.py`
- Modify: `harness/tht/cli/memory_cmd.py`
- Modify: `harness/tht/corpus/pipeline.py`
- Modify: `harness/tht/search/evidence.py`
- Modify: `harness/tht/cli/schema_cmd.py`
- Modify: `harness/tests/test_memory_save_one.py`
- Modify: `harness/tests/test_search_pack.py`
- Create: `harness/tests/test_semantic_kind_isolation.py`
**Step 1: Write failing integration tests**
Use an in-memory fake implementing the `VectorStore` port. Assert:
- schema records use `kind=schema`;
- corpus records use `kind=evidence` and exact generation;
- approved memories use `kind=memory`;
- search pack requests only its allowed kind set;
- all three paths share `workspace_id`, `workspace_revision`, hashing, and point-key construction;
- retries do not duplicate points.
**Step 2: Verify RED**
```bash
cd harness
.venv/bin/pytest tests/test_semantic_kind_isolation.py tests/test_memory_save_one.py \
tests/test_search_pack.py -q
```
Expected: current factories select pgvector/HTTP adapters and payloads lack the v3 identity fields.
**Step 3: Wire the adapter**
Make schema-v3 `qdrant` the only operational vector factory branch. Reuse the current canonical
record builders; add only missing identity fields. Keep the JSONL Memory registry and filesystem
Evidence corpus as sources of truth.
**Step 4: Verify GREEN and commit**
Run Step 2, then commit the listed files with:
```bash
git commit -m "feat: index semantic records in qdrant"
```
### Task 7: Add mandatory Qdrant and Ollama Compose services
**Files:**
- Modify: `compose.yaml`
- Create: `deploy/compose.embedding-gpu.yaml`
- Create: `docker/embedding-model-init.sh`
- Modify: `docker/core.Dockerfile`
- Modify: `deploy/env/local.env.example`
- Modify: `deploy/env/server.env.example`
- Modify: `scripts/run-stack.sh`
- Modify: `scripts/test-default-compose.sh`
- Modify: `scripts/test-unified-compose.sh`
- Create: `scripts/test-internal-semantic-compose.sh`
**Step 1: Write failing Compose contract tests**
Assert the rendered base profile has `core`, `frontend`, `qdrant`, `embedding`, and
`embedding-model-init`; private services have no published ports; persistent volumes exist; core
depends on Qdrant health and successful model init; no external vector/embedding binding is required.
Also assert all service images use version plus immutable digest. Resolve and record supported
multi-architecture digests for Qdrant v1.18.x and Ollama v0.32.x during implementation:
```bash
docker buildx imagetools inspect qdrant/qdrant:v1.18.2
docker buildx imagetools inspect ollama/ollama:0.32.0
```
**Step 2: Verify RED**
```bash
./scripts/test-default-compose.sh
./scripts/test-unified-compose.sh
./scripts/test-internal-semantic-compose.sh
```
Expected: required services and volumes are absent.
**Step 3: Implement the services**
`embedding-model-init.sh` must wait with a bounded deadline, call `ollama pull` for the exact model,
and verify it appears in `/api/tags`. The Qdrant healthcheck uses its HTTP health endpoint. The CPU
base has no device reservation; the GPU override adds only the supported device stanza.
**Step 4: Verify GREEN and commit**
Run Step 2, then:
```bash
git add compose.yaml deploy/compose.embedding-gpu.yaml docker/embedding-model-init.sh \
docker/core.Dockerfile deploy/env/local.env.example deploy/env/server.env.example \
scripts/run-stack.sh scripts/test-default-compose.sh scripts/test-unified-compose.sh \
scripts/test-internal-semantic-compose.sh
git commit -m "feat: run qdrant and ollama inside thothii"
```
### Task 8: Retire pgvector deployment and external semantic connectors
**Files:**
- Delete: `deploy/compose.local-vector.yaml`
- Delete: `deploy/compose.preprocess-local-vector.yaml`
- Delete: `deploy/sql/20-vector-roles.sql`
- Delete: `deploy/vector/reconcile-roles.sh`
- Delete: `deploy/vector/rotate-bootstrap-password.py`
- Delete: `deploy/vector/secret-policy.sh`
- Delete: `deploy/vector/vector-db-entrypoint.sh`
- Delete: `scripts/local-vector-smoke.sh`
- Delete: `scripts/test-local-vector-smoke-safety.sh`
- Delete: `scripts/test-local-vector-smoke-live-collision.sh`
- Delete: `scripts/test-vector-bootstrap-rotation.sh`
- Delete: `scripts/test-vector-migration-image.sh`
- Delete: `scripts/test-vector-secret-policy.sh`
- Modify: `scripts/test-no-deployment-coupling.sh`
- Modify: `scripts/test-no-deployment-coupling-scope.sh`
- Modify: `scripts/test-compose-secret-policy.sh`
- Modify: `.github/workflows/deployment.yml`
**Step 1: Write the failing coupling test**
Teach the coupling gate to reject active `pgvector`, `local-vector`, `THT_VECTOR_*`, workspace
embedding URLs/API keys, and external vector transports while allowing historical specs and the
explicit descriptor migration module.
**Step 2: Verify RED**
```bash
./scripts/test-no-deployment-coupling-scope.sh
./scripts/test-no-deployment-coupling.sh
./scripts/test-compose-secret-policy.sh
```
Expected: active pgvector deployment paths are reported.
**Step 3: Remove the retired paths and update CI**
Remove only repository deployment machinery. Retain harness pgvector code temporarily only if it
is needed to read/export legacy data during migration; it must not be reachable from schema v3 or
Compose. Remove it in a follow-up task once migration fixtures no longer import it.
**Step 4: Verify GREEN and commit**
Run Step 2 and the workflow fixture tests, then commit all deletions and modifications:
```bash
git add -A deploy scripts .github/workflows/deployment.yml
git commit -m "refactor: retire external vector deployment"
```
### Task 9: Update frontend workspace editing and examples
**Files:**
- Modify: `frontend/src/api/workspaces.ts`
- Modify: `frontend/src/shell/WorkspaceEditor.tsx`
- Modify: `frontend/src/shell/WorkspaceEditor.test.tsx`
- Modify: `frontend/src/shell/WorkspaceManager.test.tsx`
- Modify: `frontend/src/api/workspaces.test.ts`
- Modify: `frontend/src/workspaces/drafts.test.ts`
- Modify: `deploy/workspaces/example.yaml`
- Modify: `deploy/workspaces/psd.yaml.example`
**Step 1: Write failing UI tests**
Assert editor/preview show Qdrant collection and fixed internal embedding model, expose no vector
endpoint/credential fields, and publish schema v3. Assert legacy descriptors display a migration
banner and cannot be selected for a new session.
**Step 2: Verify RED**
```bash
cd frontend
npx vitest run src/shell/WorkspaceEditor.test.tsx src/shell/WorkspaceManager.test.tsx \
src/api/workspaces.test.ts src/workspaces/drafts.test.ts
```
Expected: fixtures and controls still use pgvector/external embedding.
**Step 3: Implement fixed semantic controls**
Collection remains editable and validated. Engine, provider, model, dimensions, and distance render
as fixed architecture values. Remove external semantic diagnostics from drafts and publish payloads.
**Step 4: Verify GREEN and commit**
Run Step 2, then commit the listed files with:
```bash
git commit -m "feat: edit qdrant workspace collections"
```
### Task 10: Add a real internal semantic smoke
**Files:**
- Create: `scripts/internal-semantic-smoke.sh`
- Modify: `scripts/unified-deployment-smoke.sh`
- Modify: `scripts/server-deployment-smoke.sh`
- Modify: `scripts/task13-runtime-fixture-check.ts`
- Modify: `scripts/test-task13-runtime-fixtures.sh`
**Step 1: Write failing smoke fixture assertions**
The fixture must require private Qdrant/Ollama services, model volume, Qdrant volume, fixed internal
URLs, and no host ports. It must reject wrong service names, external URLs, collection reuse, and
dimension changes.
**Step 2: Verify RED**
```bash
./scripts/test-task13-runtime-fixtures.sh local
./scripts/test-task13-runtime-fixtures.sh server
```
Expected: current fixture expects the two-service topology.
**Step 3: Implement the live smoke**
Using disposable volumes and a fixture workspace, start the stack on CPU, wait for the model, ensure
the collection, embed one record of each kind, query each kind with filters, restart offline, and
prove all points and the model remain available. Cleanup must remain exact and must not prune global
Docker resources.
**Step 4: Verify GREEN and commit**
```bash
./scripts/test-task13-runtime-fixtures.sh local
./scripts/test-task13-runtime-fixtures.sh server
./scripts/internal-semantic-smoke.sh
git add scripts/internal-semantic-smoke.sh scripts/unified-deployment-smoke.sh \
scripts/server-deployment-smoke.sh scripts/task13-runtime-fixture-check.ts \
scripts/test-task13-runtime-fixtures.sh
git commit -m "test: cover internal semantic services"
```
### Task 11: Update operator documentation and state
**Files:**
- Modify: `README.md`
- Modify: `AGENTS.md`
- Modify: `PROJECT_STATE.md`
- Modify: `docs/install/local-workspace-registry.md`
- Modify: `docs/install/server-workspace-registry.md`
- Modify: `docs/installazione-docker-4-contesti.md`
- Modify: `docs/workspace-diagnostic-protocol.md`
- Modify: `docs/gestione-memory.md`
- Modify: `deploy/secrets/README.md`
- Modify: `scripts/verify-workspace-install-docs.sh`
- Modify: `scripts/test-verify-workspace-install-docs.sh`
**Step 1: Write failing documentation contract assertions**
Require the four-service topology, CPU/GPU behavior, volume backup/restore, schema-v3 migration,
Qdrant collection ownership, and removal of external vector/embedding variables from active manuals.
**Step 2: Verify RED**
```bash
./scripts/test-verify-workspace-install-docs.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
```
Expected: manuals still describe external pgvector/embedding and a two-service mandatory stack.
**Step 3: Update documentation**
Document Qdrant as a derived but persistent index, Ollama model cache behavior, CPU-first startup,
optional GPU override, snapshot/restore, explicit legacy migration, and the fact that only the DWH
and LLM remain external application endpoints.
**Step 4: Verify GREEN and commit**
Run Step 2, then:
```bash
git add README.md AGENTS.md PROJECT_STATE.md docs deploy/secrets/README.md \
scripts/verify-workspace-install-docs.sh scripts/test-verify-workspace-install-docs.sh
git commit -m "docs: document internal semantic infrastructure"
```
### Task 12: Remove unreachable pgvector runtime code
**Files:**
- Delete: `harness/tht/adapters/vector/pgvector.py`
- Delete: `harness/tht/adapters/vector/legacy_direct.py`
- Delete: `harness/tht/adapters/vector/thoth_http.py`
- Delete: `harness/tht/vectorstore/rest_client.py`
- Delete: `harness/tht/vectorstore/rest_writer.py`
- Delete: `harness/tht/migrations/vector/001_extensions.sql`
- Delete: `harness/tht/migrations/vector/002_schema_tables.sql`
- Delete: `harness/tht/migrations/vector/003_roles.sql`
- Delete: `harness/tht/migrations/vector/004_evidence_generation_gc.sql`
- Modify: `harness/pyproject.toml`
- Modify/Delete: affected pgvector and migration tests under `harness/tests/l0/`
**Step 1: Prove the code is unreachable**
```bash
rg -n "PgVectorStore|ThothHttpVectorStore|LegacyDirectVectorStore|migrations/vector" \
harness backend frontend compose.yaml deploy scripts docker docs \
--glob '!docs/plans/**' --glob '!docs/superpowers/**'
```
Expected before cleanup: matches only in the files scheduled for deletion and legacy tests. If an
operational call site remains, stop and migrate it before deleting anything.
**Step 2: Delete obsolete runtime and tests**
Retain descriptor migration tests, but remove PostgreSQL vector runtime/migration packaging tests.
Remove `psycopg2-binary` only if the DWH/session PostgreSQL paths do not need it; otherwise keep it.
**Step 3: Verify focused imports and packaging**
```bash
cd harness
.venv/bin/pytest tests/test_qdrant_vector_store.py tests/test_vector_port_contract.py \
tests/test_semantic_kind_isolation.py tests/test_vector_migration_packaging.py -q
python -m build
```
Expected: Qdrant tests pass and the wheel contains no pgvector migrations. Adjust the packaging test
to assert Qdrant has no SQL migration payload.
**Step 4: Commit**
```bash
git add -A harness
git commit -m "refactor: remove pgvector runtime"
```
### Task 13: Run complete verification
**Files:**
- Modify only if a genuine regression is discovered.
**Step 1: Deterministic layer gates**
```bash
cd harness && .venv/bin/pytest -q && .venv/bin/ruff check .
cd ../backend && npx vitest run && npx tsc --noEmit -p . && npm run build
cd ../frontend && npx vitest run && npx tsc -b && npm run build
cd .. && git diff --check
```
Expected: all gates pass. Existing unrelated Ruff debt must be reported separately if it remains;
new/modified files must be Ruff-clean.
**Step 2: Deployment contracts**
```bash
./scripts/test-default-compose.sh
./scripts/test-unified-compose.sh
./scripts/test-internal-semantic-compose.sh
./scripts/test-no-deployment-coupling.sh
./scripts/test-compose-secret-policy.sh
./scripts/verify-workspace-install-docs.sh --fixtures-only
```
Expected: all pass without external vector/embedding settings.
**Step 3: Docker smokes**
```bash
./scripts/internal-semantic-smoke.sh
./scripts/workspace-registry-smoke.sh
./scripts/unified-deployment-smoke.sh
./scripts/thothctl-update-smoke.sh
./scripts/server-deployment-smoke.sh
```
Expected: CPU semantic smoke passes, persistence survives offline restart, and every script proves
exact cleanup. Investigate the previously observed `thothctl` rollback failure independently if it
recurs; do not weaken the new semantic gate to hide it.
**Step 4: Final audit**
```bash
rg -n "pgvector|local-vector|THT_VECTOR_|EMBEDDING_BASE_URL|openai_compatible|ollama_compatible" \
. --glob '!docs/plans/**' --glob '!docs/superpowers/**' --glob '!**/node_modules/**' \
--glob '!**/.venv/**' --glob '!**/.git/**'
git status --short
```
Expected: no active operational references; only explicit legacy descriptor migration fixtures may
remain. Worktree contains only intentional changes.
**Step 5: Commit verification metadata**
Update `PROJECT_STATE.md` with exact counts, image digests, smoke durations, CPU hardware, and any
manual GPU/Windows gates. Commit only verified claims:
```bash
git add PROJECT_STATE.md
git commit -m "docs: record qdrant ollama verification"
```