docs: design Pi-enabled model selector
This commit is contained in:
@@ -0,0 +1,167 @@
|
||||
# Pi-Enabled Model Selector — Design
|
||||
|
||||
**Date:** 2026-07-14
|
||||
**Status:** Approved (design), pending implementation plan
|
||||
**Layers:** backend, deployment configuration; frontend contract unchanged
|
||||
|
||||
## Problem
|
||||
|
||||
The model selector currently offers only the active `zai/glm-5.2` model. The live
|
||||
`GET /models` response is `{ "models": [] }`, so the frontend applies its existing
|
||||
degraded-mode fallback and renders only the model already stored in application
|
||||
settings.
|
||||
|
||||
The regression was introduced by the provider-credential hardening in commit
|
||||
`f064dae`. The ephemeral Pi model lister now constructs its child environment using
|
||||
`cfg.defaults.provider`. In the live deployment `PI_PROVIDER` is intentionally absent:
|
||||
the selected provider lives in the persisted application settings file. Because a
|
||||
generic model credential file is configured but no environment default provider is
|
||||
available, `buildPiChildEnv` throws `model provider credential is unavailable` before
|
||||
Pi starts. The route catches the failure and returns an empty list.
|
||||
|
||||
There is a second usability defect behind the first one: `local-qwen` is a custom
|
||||
local provider, but the credential policy does not classify it as local. A session
|
||||
using it would therefore be rejected before Pi starts even if the model appeared in
|
||||
the selector.
|
||||
|
||||
## Approved outcome
|
||||
|
||||
The selector must expose exactly these three configured models:
|
||||
|
||||
1. `zai/glm-5.2`
|
||||
2. `deepseek/deepseek-v4-flash`
|
||||
3. `local-qwen/qwen3.6-35b-a3b`
|
||||
|
||||
`zai/glm-5v-turbo` must remain hidden. The selected approach is to use Pi's
|
||||
`enabledModels` setting as the single source of truth rather than introduce a second
|
||||
application-specific allowlist.
|
||||
|
||||
The models must be usable for new sessions; merely displaying them is insufficient.
|
||||
|
||||
## Source of truth and precedence
|
||||
|
||||
The backend reads Pi settings from the same two scopes Pi uses:
|
||||
|
||||
- global: `~/.pi/agent/settings.json`;
|
||||
- project: `<harnessDir>/.pi/settings.json`.
|
||||
|
||||
If project settings define `enabledModels`, that value overrides the global value.
|
||||
Otherwise the global value applies. This mirrors Pi's settings precedence for the
|
||||
field used here without attempting to reimplement unrelated Pi settings behavior.
|
||||
|
||||
For this integration, `enabledModels` entries must be exact `provider/model` strings.
|
||||
Wildcards, ambiguous model-only patterns, and thinking-level suffixes are outside the
|
||||
selector contract. Invalid entries are ignored with a warning; the backend never
|
||||
expands them into additional visible models.
|
||||
|
||||
The live global Pi settings will be updated to contain only the three approved exact
|
||||
identifiers, removing `zai/glm-5v-turbo`.
|
||||
|
||||
## Backend model-list flow
|
||||
|
||||
`createPiModelLister` continues to start an ephemeral `pi --mode rpc` process in
|
||||
`harnessDir` and request `get_available_models`. Its environment construction changes
|
||||
as follows:
|
||||
|
||||
- scrub ambient deployment secrets and provider credentials;
|
||||
- do not require or inject the generic credential merely to enumerate models;
|
||||
- let Pi resolve configured authentication through its mounted profile
|
||||
(`~/.pi/agent/auth.json`) and custom model definitions;
|
||||
- preserve the existing portable data-root handling.
|
||||
|
||||
After Pi responds, the backend:
|
||||
|
||||
1. maps the Pi response to the public `PiModel` shape;
|
||||
2. loads the effective `enabledModels` list;
|
||||
3. matches models by the composite `provider/model` identifier;
|
||||
4. returns only matched models, in `enabledModels` order;
|
||||
5. caches the filtered result using the existing short TTL.
|
||||
|
||||
Filtering after Pi discovery ensures that an enabled identifier is shown only if Pi
|
||||
also considers the corresponding model available.
|
||||
|
||||
## Session credential behavior
|
||||
|
||||
The existing provider isolation remains in place for hosted providers. DeepSeek is
|
||||
already present in Pi's `auth.json`; Pi gives profile credentials precedence over an
|
||||
environment fallback, so selecting `deepseek/deepseek-v4-flash` uses its configured
|
||||
DeepSeek credential.
|
||||
|
||||
`local-qwen` is added to the explicit set of local providers. Its endpoint and request
|
||||
configuration remain owned by Pi's `models.json`; no generic hosted-provider key is
|
||||
required or injected for its session process.
|
||||
|
||||
No general exception is added for unknown providers. An unrecognized provider still
|
||||
fails closed before session spawn.
|
||||
|
||||
## API validation
|
||||
|
||||
`PUT /settings` validates the composite `provider/model` pair whenever the filtered
|
||||
model list is non-empty. Matching only `model.id` is insufficient because different
|
||||
providers may expose the same identifier.
|
||||
|
||||
The public shape of `GET /models` and the frontend API contract remain unchanged:
|
||||
|
||||
```json
|
||||
{
|
||||
"models": [
|
||||
{ "provider": "zai", "id": "glm-5.2", "name": "GLM-5.2", "reasoning": true }
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
No frontend component change is expected: once the endpoint returns the three models,
|
||||
the existing selector can render them and persist both provider and model.
|
||||
|
||||
## Failure behavior and observability
|
||||
|
||||
The model list fails closed to an empty array when any of these conditions applies:
|
||||
|
||||
- the effective `enabledModels` field is missing, empty, or malformed;
|
||||
- neither Pi settings file can be read successfully when one is expected;
|
||||
- the ephemeral Pi process fails, times out, or returns an invalid response;
|
||||
- none of the enabled identifiers is currently available to Pi.
|
||||
|
||||
The route keeps its graceful `{ "models": [] }` response for frontend compatibility,
|
||||
but emits a sanitized warning through the Fastify logger. Logs may include file paths,
|
||||
provider/model identifiers, and error classes; they must never include credential
|
||||
values or the contents of `auth.json`.
|
||||
|
||||
## Tests
|
||||
|
||||
Backend regression tests cover:
|
||||
|
||||
- model enumeration when `PI_PROVIDER` is absent and a generic credential file exists;
|
||||
- exact filtering and `enabledModels` ordering;
|
||||
- project `enabledModels` overriding the global list;
|
||||
- missing, malformed, and invalid settings producing an empty list/error path without
|
||||
leaking all Pi-available models;
|
||||
- `zai/glm-5v-turbo` being excluded;
|
||||
- composite provider/model validation in `PUT /settings`;
|
||||
- `local-qwen` spawning without a generic provider credential;
|
||||
- DeepSeek remaining selectable through its configured profile authentication.
|
||||
|
||||
Existing frontend tests remain the compatibility gate. A focused frontend test is
|
||||
added only if inspection reveals that the three-model response is not already covered.
|
||||
|
||||
## Deployment and verification
|
||||
|
||||
Implementation completion requires:
|
||||
|
||||
1. updating the mounted live Pi `settings.json` to the approved three-entry list;
|
||||
2. running backend unit tests and TypeScript type checking;
|
||||
3. running any affected frontend tests/type checking if frontend code changes;
|
||||
4. rebuilding and recreating the impacted `core` container;
|
||||
5. verifying container health;
|
||||
6. verifying live `GET /models` returns exactly the three approved composite IDs;
|
||||
7. verifying a model-setting update accepts DeepSeek and Qwen and rejects the hidden
|
||||
GLM-5V model without leaving application settings altered after the smoke test.
|
||||
|
||||
## Out of scope
|
||||
|
||||
- Changing Pi's built-in model registry.
|
||||
- Exposing every model with a configured credential.
|
||||
- Supporting wildcard or fuzzy `enabledModels` patterns in the web selector.
|
||||
- Adding a frontend model-management UI.
|
||||
- Generalizing custom-provider credential modes beyond the explicit `local-qwen`
|
||||
requirement.
|
||||
Reference in New Issue
Block a user