Fix Qwen session tool calls and expose thinking compatibility
Publish documentation / publish (push) Successful in 30s

This commit is contained in:
Codex
2026-09-21 16:23:51 +02:00
parent efd7d788d9
commit 84084bba37
10 changed files with 220 additions and 4 deletions
+45
View File
@@ -128,6 +128,51 @@ endpoint expects a model name different from the catalog key.
Provider integrations remain declarative. Do not register providers from
`harness/.pi/extensions/`; those extensions implement the workflow and human gates only.
### Qwen 3.6 sessions and thinking controls
For Qwen served through a vLLM-compatible chat template, declare the following inside the
model's `session` block, alongside its context and output limits:
```yaml
reasoning: true
compatibility:
supportsDeveloperRole: false
supportsReasoningEffort: false
supportsStore: false
maxTokensField: max_tokens
thinkingFormat: qwen-chat-template
```
This makes Pi send `chat_template_kwargs.enable_thinking` from the selected thinking level,
with `preserve_thinking: true`. Choose **off** to explicitly disable thinking. The alternative
`thinkingFormat: qwen` is for endpoints expecting top-level `enable_thinking`. Both formats
require `reasoning: true`; declaring `reasoning: false` does not tell the server to disable
thinking. Omit `thinkingFormat` to preserve Pi's default behavior for other providers.
Regenerate projections with the updated host CLI and recreate the local core container after
rebuilding it. Do not add these fields directly to generated Pi files. These controls do not
force tool calls or certify the workflow; perform the operator verification above.
```sh
tht --installation /absolute/path/thothii-installation.yaml installation generate
```
Use the model identifier exposed by your endpoint, such as `qwen3.6-35b-a3b`, and limits
supported by that deployment. The thinking format configures the Pi session adapter;
metadata generation continues to use its separate LiteLLM settings.
If a session displays text such as `{"type":"bash","command":"tht session show … --json"}`
and never opens a review widget, that text is not an executed tool call. A verified cause
was the Evidence JSON extension being loaded into interactive sessions and forcing
`response_format: {type: "json_object"}`. Upgrade to the core image containing the fix:
the extension belongs in `.pi/evidence-extensions/` and is loaded explicitly only by
Evidence authoring. It must not also remain in the automatically loaded `.pi/extensions/`
directory. Regenerating model configuration alone does not remove an extension from an old image.
After upgrading, reload the browser and resume the session. Verify that Pi executes
`tht session show` and opens a review widget. This fix does not require changing the Qwen
server, forcing every turn to call a tool, or teaching the model to print tool-call JSON.
## Authentication
Every provider chooses one explicit mode:
@@ -0,0 +1,57 @@
# Qwen 3.6: session tool-call failure and correction
Date: 2026-09-21. Verified runtime: Pi coding agent and Pi AI 0.80.3.
## Cause and correction
Interactive sessions returned text resembling a bash call and stopped before the
first review widget. The Evidence-only extension `tht-evidence-json-mode.ts` lived
under `.pi/extensions`, so Pi automatically loaded it into interactive sessions.
Its `before_provider_request` hook imposed `response_format: {type: "json_object"}`
and temperature zero on every request.
Replaying the captured startup request, with all original hooks preserved, isolated
the cause. With JSON response format, Qwen returned the command as text and finish
reason `stop`. Removing only that field produced a native `bash` call and finish
reason `tool_calls`. Both requests contained the same 15 tools and used High thinking.
The extension now lives in `.pi/evidence-extensions` and is loaded explicitly by
`PiEvidenceRestructurer`. Evidence retains JSON output. Interactive sessions retain
native tool calls. No changes to the remote Qwen server were required.
Earlier SDK probes replaced `session.agent.onPayload`, inadvertently bypassing the
extension hooks. Their success did not reproduce the application path and did not
establish a model or server fault.
## Model configuration
The installation catalog now accepts optional `session.compatibility.thinkingFormat`
values `qwen` and `qwen-chat-template`, requiring `reasoning: true`. Go validation,
the backend schema, and generated catalog/Pi projections preserve this setting.
It controls thinking; it was not the cause or correction of the tool-call failure.
For the verified endpoint, `qwen-chat-template` sends `enable_thinking` and
`preserve_thinking` inside `chat_template_kwargs`. Selecting Off explicitly disables
thinking. Declaring `reasoning: false` alone does not disable thinking on the server.
See the [operator configuration](../general/pi-configuration.md#qwen-36-sessions-and-thinking-controls)
and [Pi 0.80.3 model documentation](https://github.com/earendil-works/pi/blob/v0.80.3/packages/coding-agent/docs/models.md#openai-compatibility).
## Validation and local delivery
- Catalog and projection regression tests failed before the compatibility change,
then passed. Go config/modelprojection/CLI tests, 84 targeted backend tests,
TypeScript checking, and the strict documentation build passed.
- The extension-isolation regression failed before relocation. All 46 Evidence
authoring/restructurer tests and Ruff checks on changed Python files passed.
- The rebuilt local core image has digest
`sha256:9cd593d7362dcbefe177f1b9eb4b9ffdf3010ced7bd7efc8fc3fdcd288f24579`.
Core/frontend were recreated, healthy, and returned HTTP 200.
- A real `pi --mode rpc --no-session` probe with Qwen and High thinking executed
`tht session show` successfully and reached the first clarification widget.
The probe allowed only that read and `reviewer_select`; it stopped without
submitting a human response or recording decisions. Its temporary configuration
referenced the mounted credential because the original runtime lease had expired.
- The user subsequently confirmed that the application now works.
This verifies recovery from the startup failure. It does not certify every workflow
phase or the separate LiteLLM metadata-generation path.