docs: design Pi-enabled model selector

This commit is contained in:
User
2026-07-14 17:42:59 +02:00
parent dbe9659060
commit 49f6a402ee
@@ -0,0 +1,167 @@
# Pi-Enabled Model Selector — Design
**Date:** 2026-07-14
**Status:** Approved (design), pending implementation plan
**Layers:** backend, deployment configuration; frontend contract unchanged
## Problem
The model selector currently offers only the active `zai/glm-5.2` model. The live
`GET /models` response is `{ "models": [] }`, so the frontend applies its existing
degraded-mode fallback and renders only the model already stored in application
settings.
The regression was introduced by the provider-credential hardening in commit
`f064dae`. The ephemeral Pi model lister now constructs its child environment using
`cfg.defaults.provider`. In the live deployment `PI_PROVIDER` is intentionally absent:
the selected provider lives in the persisted application settings file. Because a
generic model credential file is configured but no environment default provider is
available, `buildPiChildEnv` throws `model provider credential is unavailable` before
Pi starts. The route catches the failure and returns an empty list.
There is a second usability defect behind the first one: `local-qwen` is a custom
local provider, but the credential policy does not classify it as local. A session
using it would therefore be rejected before Pi starts even if the model appeared in
the selector.
## Approved outcome
The selector must expose exactly these three configured models:
1. `zai/glm-5.2`
2. `deepseek/deepseek-v4-flash`
3. `local-qwen/qwen3.6-35b-a3b`
`zai/glm-5v-turbo` must remain hidden. The selected approach is to use Pi's
`enabledModels` setting as the single source of truth rather than introduce a second
application-specific allowlist.
The models must be usable for new sessions; merely displaying them is insufficient.
## Source of truth and precedence
The backend reads Pi settings from the same two scopes Pi uses:
- global: `~/.pi/agent/settings.json`;
- project: `<harnessDir>/.pi/settings.json`.
If project settings define `enabledModels`, that value overrides the global value.
Otherwise the global value applies. This mirrors Pi's settings precedence for the
field used here without attempting to reimplement unrelated Pi settings behavior.
For this integration, `enabledModels` entries must be exact `provider/model` strings.
Wildcards, ambiguous model-only patterns, and thinking-level suffixes are outside the
selector contract. Invalid entries are ignored with a warning; the backend never
expands them into additional visible models.
The live global Pi settings will be updated to contain only the three approved exact
identifiers, removing `zai/glm-5v-turbo`.
## Backend model-list flow
`createPiModelLister` continues to start an ephemeral `pi --mode rpc` process in
`harnessDir` and request `get_available_models`. Its environment construction changes
as follows:
- scrub ambient deployment secrets and provider credentials;
- do not require or inject the generic credential merely to enumerate models;
- let Pi resolve configured authentication through its mounted profile
(`~/.pi/agent/auth.json`) and custom model definitions;
- preserve the existing portable data-root handling.
After Pi responds, the backend:
1. maps the Pi response to the public `PiModel` shape;
2. loads the effective `enabledModels` list;
3. matches models by the composite `provider/model` identifier;
4. returns only matched models, in `enabledModels` order;
5. caches the filtered result using the existing short TTL.
Filtering after Pi discovery ensures that an enabled identifier is shown only if Pi
also considers the corresponding model available.
## Session credential behavior
The existing provider isolation remains in place for hosted providers. DeepSeek is
already present in Pi's `auth.json`; Pi gives profile credentials precedence over an
environment fallback, so selecting `deepseek/deepseek-v4-flash` uses its configured
DeepSeek credential.
`local-qwen` is added to the explicit set of local providers. Its endpoint and request
configuration remain owned by Pi's `models.json`; no generic hosted-provider key is
required or injected for its session process.
No general exception is added for unknown providers. An unrecognized provider still
fails closed before session spawn.
## API validation
`PUT /settings` validates the composite `provider/model` pair whenever the filtered
model list is non-empty. Matching only `model.id` is insufficient because different
providers may expose the same identifier.
The public shape of `GET /models` and the frontend API contract remain unchanged:
```json
{
"models": [
{ "provider": "zai", "id": "glm-5.2", "name": "GLM-5.2", "reasoning": true }
]
}
```
No frontend component change is expected: once the endpoint returns the three models,
the existing selector can render them and persist both provider and model.
## Failure behavior and observability
The model list fails closed to an empty array when any of these conditions applies:
- the effective `enabledModels` field is missing, empty, or malformed;
- neither Pi settings file can be read successfully when one is expected;
- the ephemeral Pi process fails, times out, or returns an invalid response;
- none of the enabled identifiers is currently available to Pi.
The route keeps its graceful `{ "models": [] }` response for frontend compatibility,
but emits a sanitized warning through the Fastify logger. Logs may include file paths,
provider/model identifiers, and error classes; they must never include credential
values or the contents of `auth.json`.
## Tests
Backend regression tests cover:
- model enumeration when `PI_PROVIDER` is absent and a generic credential file exists;
- exact filtering and `enabledModels` ordering;
- project `enabledModels` overriding the global list;
- missing, malformed, and invalid settings producing an empty list/error path without
leaking all Pi-available models;
- `zai/glm-5v-turbo` being excluded;
- composite provider/model validation in `PUT /settings`;
- `local-qwen` spawning without a generic provider credential;
- DeepSeek remaining selectable through its configured profile authentication.
Existing frontend tests remain the compatibility gate. A focused frontend test is
added only if inspection reveals that the three-model response is not already covered.
## Deployment and verification
Implementation completion requires:
1. updating the mounted live Pi `settings.json` to the approved three-entry list;
2. running backend unit tests and TypeScript type checking;
3. running any affected frontend tests/type checking if frontend code changes;
4. rebuilding and recreating the impacted `core` container;
5. verifying container health;
6. verifying live `GET /models` returns exactly the three approved composite IDs;
7. verifying a model-setting update accepts DeepSeek and Qwen and rejects the hidden
GLM-5V model without leaving application settings altered after the smoke test.
## Out of scope
- Changing Pi's built-in model registry.
- Exposing every model with a configured credential.
- Supporting wildcard or fuzzy `enabledModels` patterns in the web selector.
- Adding a frontend model-management UI.
- Generalizing custom-provider credential modes beyond the explicit `local-qwen`
requirement.