Fix Qwen session tool calls and expose thinking compatibility
Publish documentation / publish (push) Successful in 30s
Publish documentation / publish (push) Successful in 30s
This commit is contained in:
@@ -128,6 +128,51 @@ endpoint expects a model name different from the catalog key.
|
||||
Provider integrations remain declarative. Do not register providers from
|
||||
`harness/.pi/extensions/`; those extensions implement the workflow and human gates only.
|
||||
|
||||
### Qwen 3.6 sessions and thinking controls
|
||||
|
||||
For Qwen served through a vLLM-compatible chat template, declare the following inside the
|
||||
model's `session` block, alongside its context and output limits:
|
||||
|
||||
```yaml
|
||||
reasoning: true
|
||||
compatibility:
|
||||
supportsDeveloperRole: false
|
||||
supportsReasoningEffort: false
|
||||
supportsStore: false
|
||||
maxTokensField: max_tokens
|
||||
thinkingFormat: qwen-chat-template
|
||||
```
|
||||
|
||||
This makes Pi send `chat_template_kwargs.enable_thinking` from the selected thinking level,
|
||||
with `preserve_thinking: true`. Choose **off** to explicitly disable thinking. The alternative
|
||||
`thinkingFormat: qwen` is for endpoints expecting top-level `enable_thinking`. Both formats
|
||||
require `reasoning: true`; declaring `reasoning: false` does not tell the server to disable
|
||||
thinking. Omit `thinkingFormat` to preserve Pi's default behavior for other providers.
|
||||
|
||||
Regenerate projections with the updated host CLI and recreate the local core container after
|
||||
rebuilding it. Do not add these fields directly to generated Pi files. These controls do not
|
||||
force tool calls or certify the workflow; perform the operator verification above.
|
||||
|
||||
```sh
|
||||
tht --installation /absolute/path/thothii-installation.yaml installation generate
|
||||
```
|
||||
|
||||
Use the model identifier exposed by your endpoint, such as `qwen3.6-35b-a3b`, and limits
|
||||
supported by that deployment. The thinking format configures the Pi session adapter;
|
||||
metadata generation continues to use its separate LiteLLM settings.
|
||||
|
||||
If a session displays text such as `{"type":"bash","command":"tht session show … --json"}`
|
||||
and never opens a review widget, that text is not an executed tool call. A verified cause
|
||||
was the Evidence JSON extension being loaded into interactive sessions and forcing
|
||||
`response_format: {type: "json_object"}`. Upgrade to the core image containing the fix:
|
||||
the extension belongs in `.pi/evidence-extensions/` and is loaded explicitly only by
|
||||
Evidence authoring. It must not also remain in the automatically loaded `.pi/extensions/`
|
||||
directory. Regenerating model configuration alone does not remove an extension from an old image.
|
||||
|
||||
After upgrading, reload the browser and resume the session. Verify that Pi executes
|
||||
`tht session show` and opens a review widget. This fix does not require changing the Qwen
|
||||
server, forcing every turn to call a tool, or teaching the model to print tool-call JSON.
|
||||
|
||||
## Authentication
|
||||
|
||||
Every provider chooses one explicit mode:
|
||||
|
||||
Reference in New Issue
Block a user