Fix Qwen session tool calls and expose thinking compatibility
Publish documentation / publish (push) Successful in 30s

This commit is contained in:
Codex
2026-09-21 16:23:51 +02:00
parent efd7d788d9
commit 84084bba37
10 changed files with 220 additions and 4 deletions
+45
View File
@@ -128,6 +128,51 @@ endpoint expects a model name different from the catalog key.
Provider integrations remain declarative. Do not register providers from
`harness/.pi/extensions/`; those extensions implement the workflow and human gates only.
### Qwen 3.6 sessions and thinking controls
For Qwen served through a vLLM-compatible chat template, declare the following inside the
model's `session` block, alongside its context and output limits:
```yaml
reasoning: true
compatibility:
supportsDeveloperRole: false
supportsReasoningEffort: false
supportsStore: false
maxTokensField: max_tokens
thinkingFormat: qwen-chat-template
```
This makes Pi send `chat_template_kwargs.enable_thinking` from the selected thinking level,
with `preserve_thinking: true`. Choose **off** to explicitly disable thinking. The alternative
`thinkingFormat: qwen` is for endpoints expecting top-level `enable_thinking`. Both formats
require `reasoning: true`; declaring `reasoning: false` does not tell the server to disable
thinking. Omit `thinkingFormat` to preserve Pi's default behavior for other providers.
Regenerate projections with the updated host CLI and recreate the local core container after
rebuilding it. Do not add these fields directly to generated Pi files. These controls do not
force tool calls or certify the workflow; perform the operator verification above.
```sh
tht --installation /absolute/path/thothii-installation.yaml installation generate
```
Use the model identifier exposed by your endpoint, such as `qwen3.6-35b-a3b`, and limits
supported by that deployment. The thinking format configures the Pi session adapter;
metadata generation continues to use its separate LiteLLM settings.
If a session displays text such as `{"type":"bash","command":"tht session show … --json"}`
and never opens a review widget, that text is not an executed tool call. A verified cause
was the Evidence JSON extension being loaded into interactive sessions and forcing
`response_format: {type: "json_object"}`. Upgrade to the core image containing the fix:
the extension belongs in `.pi/evidence-extensions/` and is loaded explicitly only by
Evidence authoring. It must not also remain in the automatically loaded `.pi/extensions/`
directory. Regenerating model configuration alone does not remove an extension from an old image.
After upgrading, reload the browser and resume the session. Verify that Pi executes
`tht session show` and opens a review widget. This fix does not require changing the Qwen
server, forcing every turn to call a tool, or teaching the model to print tool-call JSON.
## Authentication
Every provider chooses one explicit mode: