Files
ThothII/docs/general/pi-configuration.md
T
Codex 84084bba37
Publish documentation / publish (push) Successful in 30s
Fix Qwen session tool calls and expose thinking compatibility
2026-09-21 16:23:51 +02:00

14 KiB

Installation Model Catalog

ThothII has one operator-authored model source: modelCatalog in deploy/<installation-id>/thothii-installation.yaml. It declares models used by interactive Pi sessions, metadata generation, and the internal embedding service. Workspace descriptors never declare providers, model allowlists, defaults, embeddings, dimensions, or vector-store settings.

Do not edit deploy/pi/models.json, deploy/pi/settings.json, files under generated/, or provider/model environment defaults. Those former sources are retired.

Host instructions in Administration

The Pi configuration navigation button opens Pi management. Its Host maintenance section selects the installation host's Linux, macOS, or Windows tab automatically; the browser's operating system does not affect it. Selecting another tab is still possible.

The host CLI includes THT_HOST_PLATFORM in the generated Compose projection using the OS on which it runs. Generate projections on the destination host, not on another computer, and do not edit generated files. Without this projection (older deployments or native development), the backend reports its own OS; a Linux Docker container cannot discover whether the physical host is macOS or Windows. Run the updated host CLI's normal configuration reload on the destination installation to regenerate this information.

Minimal catalog

schemaVersion: 2
modelCatalog:
  defaults:
    interaction: zai/glm-5.3

  embedding:
    id: ollama/qwen3-embedding:0.6b
    dimensions: 1024

  providers:
    zai:
      endpoint:
        baseUrl: https://api.z.ai/api/coding/paas/v4
      authentication:
        mode: secret_env
        apiKeyEnv: ZAI_API_KEY
      session:
        mode: openai_compatible
      metadataGeneration:
        litellmProvider: openai
      models:
        glm-5.3:
          label: GLM 5.3
          session:
            reasoning: true
            contextWindow: 200000
            maxTokens: 131072
          metadataGeneration: {}

Provider and model entries are maps. The keys form the canonical identity <provider-key>/<model-key>; label is only display text. A model is eligible for a use only when it contains that use block:

  • session makes it selectable for interactive sessions;
  • metadataGeneration makes it selectable for description generation;
  • embedding is a single installation-level model rather than a selectable list.

defaults.interaction is the only LLM default, required once per installation, never per workspace. Core and Administration share the user's operational model choice. An explicit choice takes priority over the default and is remembered in this browser for the authenticated user and application mount. Switching workspace does not change the model. Browser-storage restrictions may limit remembering to the current visit; preferences do not synchronize across devices or change installation YAML. An unavailable remembered model is not silently replaced: select another configured model.

When any metadata-generation models are configured, the default and the operational list must support both session and metadataGeneration. Entries for just one adapter may remain in the installation inventory, but are not selectable for global interaction. Core-only installations remain supported when no metadata-generation model is configured; Admin AI is then unavailable. The embedding model remains separate and is unaffected by the interaction selector.

New sessions record the selected model in their manifest. Resume retains the session's workspace and revision, but uses the current global model (the installation default for clients that omit a model). Historical manifest model fields are not rewritten by resume. Archived/finalized sessions remain read-only.

Required operator verification for every model

Catalog validation checks configuration, not model behavior. Before offering a model to users, and after changing its endpoint, adapters, or the Pi/LiteLLM versions, the operator must verify both:

  1. Core / Pi: select the model, start a test session in a prepared test workspace, exercise an actual tool call and its returned result, a human review gate, and stop/resume. Check streaming, tool arguments, authentication, and reasoning/token-limit compatibility. A plain chat reply or tht pi test alone is not sufficient.
  2. Administration / LiteLLM: select the same model and generate descriptions for a small, non-sensitive test table. Check the structured result is accepted and the generation completes. Review the output quality before using it on real metadata. This action writes test metadata and may incur provider charges: use an authorized test database and approved data.

There is no automatic certification flag or startup model probe. The operator owns this verification; do not infer compatibility from the model label or from success in just one path. Both adapters point to one catalog identity; Pi does not need to route through a new LiteLLM proxy. Models using only pi_auth cannot serve the current LiteLLM path and are excluded from shared selection.

Session adapters

Use pi_builtin for a model whose technical definition ships with Pi. This does not require pi_auth: a shared bundle credential lets native Pi and LiteLLM use the same provider identity:

deepseek:
  authentication:
    mode: secret_env
    apiKeyEnv: DEEPSEEK_API_KEY
  session:
    mode: pi_builtin
  metadataGeneration:
    litellmProvider: deepseek
  models:
    deepseek-v4-pro:
      session: {}
      metadataGeneration: {}
    deepseek-v4-flash:
      session: {}
      metadataGeneration: {}

Use openai_compatible for an explicit compatible endpoint. Each eligible session model must then declare the technical limits Pi needs. upstreamModel is optional and is used only when the endpoint expects a model name different from the catalog key.

Provider integrations remain declarative. Do not register providers from harness/.pi/extensions/; those extensions implement the workflow and human gates only.

Qwen 3.6 sessions and thinking controls

For Qwen served through a vLLM-compatible chat template, declare the following inside the model's session block, alongside its context and output limits:

reasoning: true
compatibility:
  supportsDeveloperRole: false
  supportsReasoningEffort: false
  supportsStore: false
  maxTokensField: max_tokens
  thinkingFormat: qwen-chat-template

This makes Pi send chat_template_kwargs.enable_thinking from the selected thinking level, with preserve_thinking: true. Choose off to explicitly disable thinking. The alternative thinkingFormat: qwen is for endpoints expecting top-level enable_thinking. Both formats require reasoning: true; declaring reasoning: false does not tell the server to disable thinking. Omit thinkingFormat to preserve Pi's default behavior for other providers.

Regenerate projections with the updated host CLI and recreate the local core container after rebuilding it. Do not add these fields directly to generated Pi files. These controls do not force tool calls or certify the workflow; perform the operator verification above.

tht --installation /absolute/path/thothii-installation.yaml installation generate

Use the model identifier exposed by your endpoint, such as qwen3.6-35b-a3b, and limits supported by that deployment. The thinking format configures the Pi session adapter; metadata generation continues to use its separate LiteLLM settings.

If a session displays text such as {"type":"bash","command":"tht session show … --json"} and never opens a review widget, that text is not an executed tool call. A verified cause was the Evidence JSON extension being loaded into interactive sessions and forcing response_format: {type: "json_object"}. Upgrade to the core image containing the fix: the extension belongs in .pi/evidence-extensions/ and is loaded explicitly only by Evidence authoring. It must not also remain in the automatically loaded .pi/extensions/ directory. Regenerating model configuration alone does not remove an extension from an old image.

After upgrading, reload the browser and resume the session. Verify that Pi executes tht session show and opens a review widget. This fix does not require changing the Qwen server, forcing every turn to call a tool, or teaching the model to print tool-call JSON.

Authentication

Every provider chooses one explicit mode:

  • secret_env names an approved key in the protected ThothII secret bundle through apiKeyEnv;
  • pi_auth uses Pi's protected authentication projection and is valid only for session-only pi_builtin providers;
  • none is valid only with an explicit keyless endpoint.

Secret values never belong in installation YAML, generated files, logs, CLI arguments, or browser requests. The YAML contains only an environment-variable name or an authentication mode. Pi's protected credential file remains selected by the installation authentication configuration.

For catalog providers using secret_env, the bundle is authoritative in both Core and Admin. ThothII removes only the selected provider's old auth entry from the temporary Pi session snapshot; the operator's original Pi auth store and other providers are unchanged. Provider smoke checks use the same precedence, and model enumeration receives the catalog-declared bundle keys. A missing declared key is an error, not permission to fall back to Pi auth or the legacy generic key file. After rotating a bundle key, apply the normal installation lifecycle so processes reload it.

The PSD descriptor now declares only deepseek/deepseek-v4-pro and deepseek/deepseek-v4-flash for both uses; it no longer duplicates them under deepseek-metadata. Historical records are not rewritten. A saved obsolete identity must be explicitly reselected from the current catalog; it is not silently remapped to another model or account.

Generated runtime projections

Before Compose starts, tht validates the installation and atomically writes deterministic files under deploy/<installation-id>/generated/:

generated/
├── catalog.json
├── pi/
│   ├── models.json
│   └── settings.json
└── compose.models.yaml

The normalized catalog is consumed by the backend. The Pi files and Compose override are boundary adapters. They are not configuration sources and are excluded from backup. Restore regenerates them from the installation descriptor.

Run all lifecycle commands from the project root and select the descriptor explicitly when more than one installation exists:

INSTALLATION=/absolute/path/deploy/example/thothii-installation.yaml

tht --installation "$INSTALLATION" start
tht --installation "$INSTALLATION" doctor

After editing modelCatalog or provider credentials, apply the complete runtime projection with the normal installation lifecycle, then run the Pi checks:

tht --installation "$INSTALLATION" start
tht --installation "$INSTALLATION" pi doctor
tht --installation "$INSTALLATION" pi test

tht pi restart, tht pi update, and tht pi rollback refuse to run while generated model projections differ from modelCatalog: those commands recreate only core, so they must never partially apply an embedding change. tht pi update changes the Pi version; it is not the configuration command. There is no tht pi configure and no separate apply command.

Migrating a legacy installation

For schema-v2 descriptors with the former defaults.session and defaults.metadataGeneration, replace both with defaults.interaction. Equal legacy values are accepted and normalized in memory; the loader never rewrites the descriptor. Different values fail with migration_required: explicitly choose a model supporting both uses, remove both old fields, and set the single new field. Do not mix new and legacy fields. A Core-only legacy session default can be normalized when Admin AI is absent.

The generated runtime catalog now uses schema version 2 and only defaultInteraction. Regenerate and apply all runtime projections with the matching host/backend release using the normal installation lifecycle; do not deploy only the backend against an old generated catalog or hand-edit generated JSON.

The migrator reads the former installation metadataGeneration block and the two former Pi JSON files, but never modifies them. Supply the facts that cannot be inferred safely and write a separate candidate:

tht --installation /absolute/path/legacy/thothii-installation.yaml installation migrate \
  --output /absolute/path/thothii-installation.v2.yaml \
  --session-default zai/glm-5.3 \
  --embedding-id ollama/qwen3-embedding:0.6b \
  --embedding-dimensions 1024

Review the candidate, move the legacy source files out of the installation only after approval, then select the v2 descriptor. Ambiguous aliases, endpoint conflicts, or missing authentication facts produce field-level errors; the migrator does not guess. The legacy CLI flag --session-default now supplies the unified interaction default in the candidate; if it conflicts with the legacy metadata default, align that choice explicitly before retrying.

Troubleshooting

Symptom Meaning Action
migration_required A retired model source or installation schema is still present Run the installation migrator and review its candidate
Invalid interaction default The canonical ID does not support all configured uses Correct defaults.interaction or the intended adapter blocks
Generated projection drift Runtime files differ from the descriptor-derived bytes Run tht start or tht pi restart --yes --drain
model_unavailable on create/resume The selected global model is no longer eligible Explicitly choose an eligible model; no fallback is applied
Provider smoke failure Credentials, endpoint, or provider availability is invalid Correct the protected credential or catalog endpoint, restart, then run tht pi test