Fix Qwen session tool calls and expose thinking compatibility
Publish documentation / publish (push) Successful in 30s
Publish documentation / publish (push) Successful in 30s
This commit is contained in:
@@ -128,6 +128,51 @@ endpoint expects a model name different from the catalog key.
|
||||
Provider integrations remain declarative. Do not register providers from
|
||||
`harness/.pi/extensions/`; those extensions implement the workflow and human gates only.
|
||||
|
||||
### Qwen 3.6 sessions and thinking controls
|
||||
|
||||
For Qwen served through a vLLM-compatible chat template, declare the following inside the
|
||||
model's `session` block, alongside its context and output limits:
|
||||
|
||||
```yaml
|
||||
reasoning: true
|
||||
compatibility:
|
||||
supportsDeveloperRole: false
|
||||
supportsReasoningEffort: false
|
||||
supportsStore: false
|
||||
maxTokensField: max_tokens
|
||||
thinkingFormat: qwen-chat-template
|
||||
```
|
||||
|
||||
This makes Pi send `chat_template_kwargs.enable_thinking` from the selected thinking level,
|
||||
with `preserve_thinking: true`. Choose **off** to explicitly disable thinking. The alternative
|
||||
`thinkingFormat: qwen` is for endpoints expecting top-level `enable_thinking`. Both formats
|
||||
require `reasoning: true`; declaring `reasoning: false` does not tell the server to disable
|
||||
thinking. Omit `thinkingFormat` to preserve Pi's default behavior for other providers.
|
||||
|
||||
Regenerate projections with the updated host CLI and recreate the local core container after
|
||||
rebuilding it. Do not add these fields directly to generated Pi files. These controls do not
|
||||
force tool calls or certify the workflow; perform the operator verification above.
|
||||
|
||||
```sh
|
||||
tht --installation /absolute/path/thothii-installation.yaml installation generate
|
||||
```
|
||||
|
||||
Use the model identifier exposed by your endpoint, such as `qwen3.6-35b-a3b`, and limits
|
||||
supported by that deployment. The thinking format configures the Pi session adapter;
|
||||
metadata generation continues to use its separate LiteLLM settings.
|
||||
|
||||
If a session displays text such as `{"type":"bash","command":"tht session show … --json"}`
|
||||
and never opens a review widget, that text is not an executed tool call. A verified cause
|
||||
was the Evidence JSON extension being loaded into interactive sessions and forcing
|
||||
`response_format: {type: "json_object"}`. Upgrade to the core image containing the fix:
|
||||
the extension belongs in `.pi/evidence-extensions/` and is loaded explicitly only by
|
||||
Evidence authoring. It must not also remain in the automatically loaded `.pi/extensions/`
|
||||
directory. Regenerating model configuration alone does not remove an extension from an old image.
|
||||
|
||||
After upgrading, reload the browser and resume the session. Verify that Pi executes
|
||||
`tht session show` and opens a review widget. This fix does not require changing the Qwen
|
||||
server, forcing every turn to call a tool, or teaching the model to print tool-call JSON.
|
||||
|
||||
## Authentication
|
||||
|
||||
Every provider chooses one explicit mode:
|
||||
|
||||
@@ -0,0 +1,57 @@
|
||||
# Qwen 3.6: session tool-call failure and correction
|
||||
|
||||
Date: 2026-09-21. Verified runtime: Pi coding agent and Pi AI 0.80.3.
|
||||
|
||||
## Cause and correction
|
||||
|
||||
Interactive sessions returned text resembling a bash call and stopped before the
|
||||
first review widget. The Evidence-only extension `tht-evidence-json-mode.ts` lived
|
||||
under `.pi/extensions`, so Pi automatically loaded it into interactive sessions.
|
||||
Its `before_provider_request` hook imposed `response_format: {type: "json_object"}`
|
||||
and temperature zero on every request.
|
||||
|
||||
Replaying the captured startup request, with all original hooks preserved, isolated
|
||||
the cause. With JSON response format, Qwen returned the command as text and finish
|
||||
reason `stop`. Removing only that field produced a native `bash` call and finish
|
||||
reason `tool_calls`. Both requests contained the same 15 tools and used High thinking.
|
||||
|
||||
The extension now lives in `.pi/evidence-extensions` and is loaded explicitly by
|
||||
`PiEvidenceRestructurer`. Evidence retains JSON output. Interactive sessions retain
|
||||
native tool calls. No changes to the remote Qwen server were required.
|
||||
|
||||
Earlier SDK probes replaced `session.agent.onPayload`, inadvertently bypassing the
|
||||
extension hooks. Their success did not reproduce the application path and did not
|
||||
establish a model or server fault.
|
||||
|
||||
## Model configuration
|
||||
|
||||
The installation catalog now accepts optional `session.compatibility.thinkingFormat`
|
||||
values `qwen` and `qwen-chat-template`, requiring `reasoning: true`. Go validation,
|
||||
the backend schema, and generated catalog/Pi projections preserve this setting.
|
||||
It controls thinking; it was not the cause or correction of the tool-call failure.
|
||||
|
||||
For the verified endpoint, `qwen-chat-template` sends `enable_thinking` and
|
||||
`preserve_thinking` inside `chat_template_kwargs`. Selecting Off explicitly disables
|
||||
thinking. Declaring `reasoning: false` alone does not disable thinking on the server.
|
||||
See the [operator configuration](../general/pi-configuration.md#qwen-36-sessions-and-thinking-controls)
|
||||
and [Pi 0.80.3 model documentation](https://github.com/earendil-works/pi/blob/v0.80.3/packages/coding-agent/docs/models.md#openai-compatibility).
|
||||
|
||||
## Validation and local delivery
|
||||
|
||||
- Catalog and projection regression tests failed before the compatibility change,
|
||||
then passed. Go config/modelprojection/CLI tests, 84 targeted backend tests,
|
||||
TypeScript checking, and the strict documentation build passed.
|
||||
- The extension-isolation regression failed before relocation. All 46 Evidence
|
||||
authoring/restructurer tests and Ruff checks on changed Python files passed.
|
||||
- The rebuilt local core image has digest
|
||||
`sha256:9cd593d7362dcbefe177f1b9eb4b9ffdf3010ced7bd7efc8fc3fdcd288f24579`.
|
||||
Core/frontend were recreated, healthy, and returned HTTP 200.
|
||||
- A real `pi --mode rpc --no-session` probe with Qwen and High thinking executed
|
||||
`tht session show` successfully and reached the first clarification widget.
|
||||
The probe allowed only that read and `reviewer_select`; it stopped without
|
||||
submitting a human response or recording decisions. Its temporary configuration
|
||||
referenced the mounted credential because the original runtime lease had expired.
|
||||
- The user subsequently confirmed that the application now works.
|
||||
|
||||
This verifies recovery from the startup failure. It does not certify every workflow
|
||||
phase or the separate LiteLLM metadata-generation path.
|
||||
Reference in New Issue
Block a user