249 lines
11 KiB
Markdown
249 lines
11 KiB
Markdown
# Installation Model Catalog
|
|
|
|
Status: implemented on 2026-09-02.
|
|
|
|
## Outcome
|
|
|
|
`thothii-installation.yaml` is the only operator-authored source for models used by interactive
|
|
sessions, metadata generation, and embedding. Runtime-specific files are deterministic projections,
|
|
not additional configuration sources. Workspace descriptors contain database and Evidence concerns
|
|
and no model, provider, allowlist, default, embedding, or vector-store configuration.
|
|
|
|
This design does not merge execution lifecycles. Pi continues to run interactive sessions, the
|
|
short-lived LiteLLM helper continues to perform metadata generation, and the internal Ollama service
|
|
continues to provide embeddings. They share model declaration, not execution machinery.
|
|
|
|
## Canonical installation shape
|
|
|
|
The following example covers all currently required cases: a Pi built-in model, an authenticated
|
|
custom endpoint, a keyless internal endpoint, metadata generation, and the single embedding model.
|
|
|
|
```yaml
|
|
schemaVersion: 2
|
|
profile: server
|
|
projectDirectory: /srv/thothii
|
|
envFile: /srv/thothii/operator.env
|
|
|
|
workspaceRepository:
|
|
remote: git@git.example.com:organization/workspaces.git
|
|
branch: main
|
|
access: ssh
|
|
|
|
modelCatalog:
|
|
defaults:
|
|
session: zai/glm-5.3
|
|
metadataGeneration: local-qwen/qwen3.6-35b-a3b
|
|
|
|
embedding:
|
|
id: ollama/qwen3-embedding:0.6b
|
|
dimensions: 1024
|
|
|
|
providers:
|
|
deepseek:
|
|
authentication:
|
|
mode: pi_auth
|
|
session:
|
|
mode: pi_builtin
|
|
models:
|
|
deepseek-v4-pro:
|
|
session: {}
|
|
deepseek-v4-flash:
|
|
session: {}
|
|
|
|
zai:
|
|
endpoint:
|
|
baseUrl: https://api.z.ai/api/coding/paas/v4
|
|
authentication:
|
|
mode: secret_env
|
|
apiKeyEnv: ZAI_API_KEY
|
|
session:
|
|
mode: openai_compatible
|
|
metadataGeneration:
|
|
litellmProvider: openai
|
|
models:
|
|
glm-5.3:
|
|
label: GLM-5.3
|
|
session:
|
|
reasoning: true
|
|
contextWindow: 200000
|
|
maxTokens: 131072
|
|
metadataGeneration: {}
|
|
|
|
local-qwen:
|
|
endpoint:
|
|
baseUrl: https://ml-aritmolab.policlinicosandonato.it/v1
|
|
authentication:
|
|
mode: none
|
|
session:
|
|
mode: openai_compatible
|
|
metadataGeneration:
|
|
litellmProvider: openai
|
|
models:
|
|
qwen3.6-35b-a3b:
|
|
label: Qwen3.6 35B A3B
|
|
session:
|
|
reasoning: false
|
|
contextWindow: 131072
|
|
maxTokens: 16384
|
|
compatibility:
|
|
supportsDeveloperRole: false
|
|
supportsReasoningEffort: false
|
|
supportsStore: false
|
|
maxTokensField: max_tokens
|
|
metadataGeneration:
|
|
disableThinking: true
|
|
|
|
authentication:
|
|
configDirectory: /srv/thothii/auth-canonical
|
|
runtimeProjection:
|
|
directory: /srv/thothii/auth-runtime
|
|
uid: 10001
|
|
gid: 10001
|
|
```
|
|
|
|
The catalog uses maps instead of repeated IDs. The canonical identity of a model is always derived
|
|
as `<provider-key>/<model-key>`. `upstreamModel` may be added to a model only when the endpoint uses
|
|
a different identifier. `label` is optional and falls back to the canonical identity.
|
|
|
|
Model eligibility is not repeated in an `usages` array. A `session` block makes the model eligible
|
|
for sessions; a `metadataGeneration` block makes it eligible for metadata generation. The embedding
|
|
is a single required installation value rather than a list plus default.
|
|
|
|
## Provider and authentication rules
|
|
|
|
A provider owns one endpoint, one authentication mode, and zero or one adapter for each runtime.
|
|
Model entries cannot override provider endpoint or credentials. If the same upstream service needs
|
|
different endpoints or credentials, the installation declares two provider identities.
|
|
|
|
Supported session modes are intentionally closed:
|
|
|
|
- `pi_builtin`: Pi already owns the model's technical descriptor; the model's `session` block is
|
|
empty and ThothII does not copy context-window or compatibility facts.
|
|
- `openai_compatible`: ThothII generates a Pi custom-provider descriptor; each session model supplies
|
|
the technical values required by Pi.
|
|
|
|
Metadata generation uses the provider-level `litellmProvider`. A model-level
|
|
`metadataGeneration.disableThinking: true` is permitted only for an explicit compatible endpoint.
|
|
There is no generic adapter or plugin abstraction in schema version 2.
|
|
|
|
Exactly one provider authentication mode is allowed:
|
|
|
|
- `secret_env` requires an approved API-key environment reference present in the protected secret
|
|
bundle. Secret values never enter YAML, generated files, logs, arguments, or API responses.
|
|
- `pi_auth` is valid only for session-only `pi_builtin` providers and resolves through Pi's protected
|
|
authentication projection.
|
|
- `none` is valid only for an explicit endpoint. Runtime projections may supply a fixed non-secret
|
|
compatibility placeholder when a client library requires a non-empty key.
|
|
|
|
## Defaults and selections
|
|
|
|
`defaults.session` and `embedding` are required. `defaults.metadataGeneration` is required exactly
|
|
when at least one model has a `metadataGeneration` block; metadata generation may otherwise be
|
|
absent and its UI controls are disabled.
|
|
|
|
`modelCatalog.defaults.session` is the only configured session-model default. `PI_PROVIDER`,
|
|
`PI_MODEL`, and provider/model fields in installation-default settings are removed. A user choice is
|
|
a Model Selection containing only the canonical model identity and runtime controls such as thinking
|
|
level. A session manifest pins the selected canonical identity.
|
|
|
|
Removing the currently selected model causes new-session selection to fall back to the catalog
|
|
default with an explicit administrative warning. An existing session is never silently moved to a
|
|
different model; resume fails with `model_unavailable` when its pinned identity can no longer be
|
|
resolved.
|
|
|
|
## Generated runtime projections
|
|
|
|
Before Compose starts, `tht` strictly validates schema version 2 and generates installation-local
|
|
artifacts below `deploy/<installation-id>/generated/`:
|
|
|
|
- a normalized catalog JSON consumed defensively by the backend;
|
|
- Pi `models.json` for custom providers;
|
|
- Pi `settings.json`, combining fixed product settings with the session-eligible canonical IDs;
|
|
- a Compose override that mounts the projections and supplies embedding identity and dimensions to
|
|
core, preprocessing, and `embedding-model-init`.
|
|
|
|
Generation is deterministic and published only after every candidate artifact validates. A failed
|
|
generation aborts start before Compose is invoked. `tht doctor` recomputes expected bytes and reports
|
|
differences; no digest manifest or separate apply command exists. When projection bytes change,
|
|
`tht start` recreates the affected services so they cannot continue with an older bind mount.
|
|
Pi-only restart, update, and rollback operations reject projection drift and direct the operator to
|
|
`tht start`, because applying only the core-facing files could leave embedding services stale.
|
|
|
|
Generated projections are not backed up. Restore validates the canonical installation descriptor,
|
|
regenerates every projection, and only then starts services. Base Compose files and `operator.env`
|
|
must contain no model identities, defaults, endpoints, or dimensions.
|
|
|
|
## Workspace schema v4
|
|
|
|
Workspace schema v4 removes both top-level `llm_policy` and `semantic_index`. The entire latter
|
|
block is redundant today: its engine and distance are product constants, its collection duplicates
|
|
the workspace ID, and its model and dimensions are installation facts.
|
|
|
|
The runtime derives:
|
|
|
|
- Qdrant collection identity from the workspace ID;
|
|
- engine and distance from the supported product contract;
|
|
- embedding identity and dimensions from the Installation Model Catalog.
|
|
|
|
The published index generation records the canonical embedding identity and dimensions that created
|
|
it. A mismatch makes the index explicitly incompatible and requires operator-triggered
|
|
preprocessing. No existing index is deleted or rebuilt automatically.
|
|
|
|
The v3-to-v4 workspace migration is deterministic: set `workspace.schema_version` to `4`, remove
|
|
`llm_policy`, and remove `semantic_index`. It does not alter database, Evidence, diagnostics, or
|
|
binding data.
|
|
|
|
## Installation migration
|
|
|
|
Legacy installation migration must inspect all three former sources:
|
|
|
|
1. `metadataGeneration` in `thothii-installation.yaml`;
|
|
2. `deploy/pi/models.json`;
|
|
3. `deploy/pi/settings.json`.
|
|
|
|
The migrator emits a version-2 candidate only when it can reconcile identities, endpoints,
|
|
credentials, and runtime-specific facts without guessing. Ambiguous aliases such as `glm-53`,
|
|
`zai/glm-5.3`, and `openai/glm-5.3` are not silently equated. A conflict produces a field-level
|
|
report and leaves every input unchanged for operator resolution.
|
|
|
|
After migration, the strict loader rejects `metadataGeneration`, workspace `llm_policy`, workspace
|
|
`semantic_index`, legacy Pi source files, unknown fields, duplicate YAML keys, invalid defaults, and
|
|
incompatible authentication/adapter combinations with an actionable `migration_required` or
|
|
validation error.
|
|
|
|
## Final simplicity audit
|
|
|
|
The accepted design removes every configuration duplication that can be removed without inference:
|
|
|
|
- one authored installation file instead of an installation block plus two Pi files;
|
|
- one canonical `provider/model` identity instead of display IDs and runtime IDs;
|
|
- per-use blocks instead of a duplicated usages list;
|
|
- one embedding entry instead of a selectable embedding catalog;
|
|
- one catalog session default instead of environment and settings defaults;
|
|
- no model or vector-store fields in workspace descriptors;
|
|
- provider-level credentials instead of per-model credentials;
|
|
- no generic runtime-plugin abstraction;
|
|
- no persisted digest, apply command, or backup of generated projections.
|
|
|
|
The remaining generated files are necessary boundary adapters, not configuration concepts. Making
|
|
the backend parse the authoring YAML independently would remove one file but restore two semantic
|
|
validators. Hard-coding embedding values in Compose would remove one projection but restore a model
|
|
source outside the catalog. Inferring authentication from missing fields would save one YAML key but
|
|
turn a safe explicit choice into ambiguity. These apparent simplifications are therefore rejected.
|
|
|
|
No further reduction was found that preserves one authority, strict validation, explicit security,
|
|
session determinism, and model-free workspaces.
|
|
|
|
## Implementation surface
|
|
|
|
Implementation must update the host `tht` installation loader, setup and lifecycle projection,
|
|
doctor, backup/restore, Compose mounts and embedding inputs, backend catalog/settings/session model
|
|
resolution, workspace schema and migration, runtime rendering and diagnostics, frontend workspace
|
|
drafts and model filtering, examples, fixtures, and documentation. Existing session manifests remain
|
|
readable and keep their pinned provider/model identity; only resume resolution changes to the new
|
|
catalog.
|
|
|
|
Implementation completed after explicit approval. The installation schema, deterministic runtime
|
|
projections, migration path, model-free workspace schema v4, backend consumers, operator UI,
|
|
fixtures, and documentation now enforce this contract.
|