Files
ThothII/docs/general/pi-configuration.md
Codex 84084bba37
Publish documentation / publish (push) Successful in 30s
Fix Qwen session tool calls and expose thinking compatibility
2026-09-21 16:23:51 +02:00

282 lines
14 KiB
Markdown

# Installation Model Catalog
ThothII has one operator-authored model source: `modelCatalog` in
`deploy/<installation-id>/thothii-installation.yaml`. It declares models used by interactive Pi
sessions, metadata generation, and the internal embedding service. Workspace descriptors never
declare providers, model allowlists, defaults, embeddings, dimensions, or vector-store settings.
Do not edit `deploy/pi/models.json`, `deploy/pi/settings.json`, files under `generated/`, or
provider/model environment defaults. Those former sources are retired.
## Host instructions in Administration
The **Pi configuration** navigation button opens **Pi management**. Its Host maintenance
section selects the installation host's Linux, macOS, or Windows tab automatically; the
browser's operating system does not affect it. Selecting another tab is still possible.
The host CLI includes `THT_HOST_PLATFORM` in the generated Compose projection using the OS
on which it runs. Generate projections on the destination host, not on another computer,
and do not edit generated files. Without this projection (older deployments or native
development), the backend reports its own OS; a Linux Docker container cannot discover
whether the physical host is macOS or Windows. Run the updated host CLI's normal
configuration reload on the destination installation to regenerate this information.
## Minimal catalog
```yaml
schemaVersion: 2
modelCatalog:
defaults:
interaction: zai/glm-5.3
embedding:
id: ollama/qwen3-embedding:0.6b
dimensions: 1024
providers:
zai:
endpoint:
baseUrl: https://api.z.ai/api/coding/paas/v4
authentication:
mode: secret_env
apiKeyEnv: ZAI_API_KEY
session:
mode: openai_compatible
metadataGeneration:
litellmProvider: openai
models:
glm-5.3:
label: GLM 5.3
session:
reasoning: true
contextWindow: 200000
maxTokens: 131072
metadataGeneration: {}
```
Provider and model entries are maps. The keys form the canonical identity
`<provider-key>/<model-key>`; `label` is only display text. A model is eligible for a use only when
it contains that use block:
- `session` makes it selectable for interactive sessions;
- `metadataGeneration` makes it selectable for description generation;
- `embedding` is a single installation-level model rather than a selectable list.
`defaults.interaction` is the only LLM default, required once per installation, never per workspace.
Core and Administration share the user's operational model choice. An explicit choice takes priority
over the default and is remembered in this browser for the authenticated user and application mount.
Switching workspace does not change the model. Browser-storage restrictions may limit remembering
to the current visit; preferences do not synchronize across devices or change installation YAML.
An unavailable remembered model is not silently replaced: select another configured model.
When any metadata-generation models are configured, the default and the operational list must
support both `session` and `metadataGeneration`. Entries for just one adapter may remain in the
installation inventory, but are not selectable for global interaction. Core-only installations remain
supported when no metadata-generation model is configured; Admin AI is then unavailable.
The embedding model remains separate and is unaffected by the interaction selector.
New sessions record the selected model in their manifest. Resume retains the session's workspace and
revision, but uses the current global model (the installation default for clients that omit a model).
Historical manifest model fields are not rewritten by resume. Archived/finalized sessions remain read-only.
### Required operator verification for every model
Catalog validation checks configuration, not model behavior. Before offering a model to users, and
after changing its endpoint, adapters, or the Pi/LiteLLM versions, the operator must verify **both**:
1. **Core / Pi:** select the model, start a test session in a prepared test workspace, exercise an
actual tool call and its returned result, a human review gate, and stop/resume. Check streaming,
tool arguments, authentication, and reasoning/token-limit compatibility. A plain chat reply or
`tht pi test` alone is not sufficient.
2. **Administration / LiteLLM:** select the same model and generate descriptions for a small,
non-sensitive test table. Check the structured result is accepted and the generation completes.
Review the output quality before using it on real metadata. This action writes test metadata
and may incur provider charges: use an authorized test database and approved data.
There is no automatic certification flag or startup model probe. The operator owns this verification;
do not infer compatibility from the model label or from success in just one path. Both adapters point
to one catalog identity; Pi does not need to route through a new LiteLLM proxy. Models using only
`pi_auth` cannot serve the current LiteLLM path and are excluded from shared selection.
## Session adapters
Use `pi_builtin` for a model whose technical definition ships with Pi. This does not require
`pi_auth`: a shared bundle credential lets native Pi and LiteLLM use the same provider identity:
```yaml
deepseek:
authentication:
mode: secret_env
apiKeyEnv: DEEPSEEK_API_KEY
session:
mode: pi_builtin
metadataGeneration:
litellmProvider: deepseek
models:
deepseek-v4-pro:
session: {}
metadataGeneration: {}
deepseek-v4-flash:
session: {}
metadataGeneration: {}
```
Use `openai_compatible` for an explicit compatible endpoint. Each eligible session model must then
declare the technical limits Pi needs. `upstreamModel` is optional and is used only when the
endpoint expects a model name different from the catalog key.
Provider integrations remain declarative. Do not register providers from
`harness/.pi/extensions/`; those extensions implement the workflow and human gates only.
### Qwen 3.6 sessions and thinking controls
For Qwen served through a vLLM-compatible chat template, declare the following inside the
model's `session` block, alongside its context and output limits:
```yaml
reasoning: true
compatibility:
supportsDeveloperRole: false
supportsReasoningEffort: false
supportsStore: false
maxTokensField: max_tokens
thinkingFormat: qwen-chat-template
```
This makes Pi send `chat_template_kwargs.enable_thinking` from the selected thinking level,
with `preserve_thinking: true`. Choose **off** to explicitly disable thinking. The alternative
`thinkingFormat: qwen` is for endpoints expecting top-level `enable_thinking`. Both formats
require `reasoning: true`; declaring `reasoning: false` does not tell the server to disable
thinking. Omit `thinkingFormat` to preserve Pi's default behavior for other providers.
Regenerate projections with the updated host CLI and recreate the local core container after
rebuilding it. Do not add these fields directly to generated Pi files. These controls do not
force tool calls or certify the workflow; perform the operator verification above.
```sh
tht --installation /absolute/path/thothii-installation.yaml installation generate
```
Use the model identifier exposed by your endpoint, such as `qwen3.6-35b-a3b`, and limits
supported by that deployment. The thinking format configures the Pi session adapter;
metadata generation continues to use its separate LiteLLM settings.
If a session displays text such as `{"type":"bash","command":"tht session show … --json"}`
and never opens a review widget, that text is not an executed tool call. A verified cause
was the Evidence JSON extension being loaded into interactive sessions and forcing
`response_format: {type: "json_object"}`. Upgrade to the core image containing the fix:
the extension belongs in `.pi/evidence-extensions/` and is loaded explicitly only by
Evidence authoring. It must not also remain in the automatically loaded `.pi/extensions/`
directory. Regenerating model configuration alone does not remove an extension from an old image.
After upgrading, reload the browser and resume the session. Verify that Pi executes
`tht session show` and opens a review widget. This fix does not require changing the Qwen
server, forcing every turn to call a tool, or teaching the model to print tool-call JSON.
## Authentication
Every provider chooses one explicit mode:
- `secret_env` names an approved key in the protected ThothII secret bundle through `apiKeyEnv`;
- `pi_auth` uses Pi's protected authentication projection and is valid only for session-only
`pi_builtin` providers;
- `none` is valid only with an explicit keyless endpoint.
Secret values never belong in installation YAML, generated files, logs, CLI arguments, or browser
requests. The YAML contains only an environment-variable name or an authentication mode. Pi's
protected credential file remains selected by the installation authentication configuration.
For catalog providers using `secret_env`, the bundle is authoritative in both Core and Admin.
ThothII removes only the selected provider's old auth entry from the temporary Pi session snapshot;
the operator's original Pi auth store and other providers are unchanged. Provider smoke checks use
the same precedence, and model enumeration receives the catalog-declared bundle keys. A missing
declared key is an error, not permission to fall back to Pi auth or the legacy generic key file.
After rotating a bundle key, apply the normal installation lifecycle so processes reload it.
The PSD descriptor now declares only `deepseek/deepseek-v4-pro` and `deepseek/deepseek-v4-flash`
for both uses; it no longer duplicates them under `deepseek-metadata`. Historical records are not
rewritten. A saved obsolete identity must be explicitly reselected from the current catalog;
it is not silently remapped to another model or account.
## Generated runtime projections
Before Compose starts, `tht` validates the installation and atomically writes deterministic files
under `deploy/<installation-id>/generated/`:
```text
generated/
├── catalog.json
├── pi/
│ ├── models.json
│ └── settings.json
└── compose.models.yaml
```
The normalized catalog is consumed by the backend. The Pi files and Compose override are boundary
adapters. They are not configuration sources and are excluded from backup. Restore regenerates
them from the installation descriptor.
Run all lifecycle commands from the project root and select the descriptor explicitly when more
than one installation exists:
```bash
INSTALLATION=/absolute/path/deploy/example/thothii-installation.yaml
tht --installation "$INSTALLATION" start
tht --installation "$INSTALLATION" doctor
```
After editing `modelCatalog` or provider credentials, apply the complete runtime projection with
the normal installation lifecycle, then run the Pi checks:
```bash
tht --installation "$INSTALLATION" start
tht --installation "$INSTALLATION" pi doctor
tht --installation "$INSTALLATION" pi test
```
`tht pi restart`, `tht pi update`, and `tht pi rollback` refuse to run while generated model
projections differ from `modelCatalog`: those commands recreate only `core`, so they must never
partially apply an embedding change. `tht pi update` changes the Pi version; it is not the
configuration command. There is no `tht pi configure` and no separate apply command.
## Migrating a legacy installation
For schema-v2 descriptors with the former `defaults.session` and `defaults.metadataGeneration`,
replace both with `defaults.interaction`. Equal legacy values are accepted and normalized in memory;
the loader never rewrites the descriptor. Different values fail with `migration_required`: explicitly
choose a model supporting both uses, remove both old fields, and set the single new field. Do not mix
new and legacy fields. A Core-only legacy session default can be normalized when Admin AI is absent.
The generated runtime catalog now uses schema version **2** and only `defaultInteraction`. Regenerate
and apply all runtime projections with the matching host/backend release using the normal installation
lifecycle; do not deploy only the backend against an old generated catalog or hand-edit generated JSON.
The migrator reads the former installation `metadataGeneration` block and the two former Pi JSON
files, but never modifies them. Supply the facts that cannot be inferred safely and write a separate
candidate:
```bash
tht --installation /absolute/path/legacy/thothii-installation.yaml installation migrate \
--output /absolute/path/thothii-installation.v2.yaml \
--session-default zai/glm-5.3 \
--embedding-id ollama/qwen3-embedding:0.6b \
--embedding-dimensions 1024
```
Review the candidate, move the legacy source files out of the installation only after approval,
then select the v2 descriptor. Ambiguous aliases, endpoint conflicts, or missing authentication
facts produce field-level errors; the migrator does not guess.
The legacy CLI flag `--session-default` now supplies the unified interaction default in the candidate;
if it conflicts with the legacy metadata default, align that choice explicitly before retrying.
## Troubleshooting
| Symptom | Meaning | Action |
| --- | --- | --- |
| `migration_required` | A retired model source or installation schema is still present | Run the installation migrator and review its candidate |
| Invalid interaction default | The canonical ID does not support all configured uses | Correct `defaults.interaction` or the intended adapter blocks |
| Generated projection drift | Runtime files differ from the descriptor-derived bytes | Run `tht start` or `tht pi restart --yes --drain` |
| `model_unavailable` on create/resume | The selected global model is no longer eligible | Explicitly choose an eligible model; no fallback is applied |
| Provider smoke failure | Credentials, endpoint, or provider availability is invalid | Correct the protected credential or catalog endpoint, restart, then run `tht pi test` |