diff --git a/docs/superpowers/specs/2026-07-14-pi-enabled-model-selector-design.md b/docs/superpowers/specs/2026-07-14-pi-enabled-model-selector-design.md new file mode 100644 index 00000000..e73de5ae --- /dev/null +++ b/docs/superpowers/specs/2026-07-14-pi-enabled-model-selector-design.md @@ -0,0 +1,167 @@ +# Pi-Enabled Model Selector — Design + +**Date:** 2026-07-14 +**Status:** Approved (design), pending implementation plan +**Layers:** backend, deployment configuration; frontend contract unchanged + +## Problem + +The model selector currently offers only the active `zai/glm-5.2` model. The live +`GET /models` response is `{ "models": [] }`, so the frontend applies its existing +degraded-mode fallback and renders only the model already stored in application +settings. + +The regression was introduced by the provider-credential hardening in commit +`f064dae`. The ephemeral Pi model lister now constructs its child environment using +`cfg.defaults.provider`. In the live deployment `PI_PROVIDER` is intentionally absent: +the selected provider lives in the persisted application settings file. Because a +generic model credential file is configured but no environment default provider is +available, `buildPiChildEnv` throws `model provider credential is unavailable` before +Pi starts. The route catches the failure and returns an empty list. + +There is a second usability defect behind the first one: `local-qwen` is a custom +local provider, but the credential policy does not classify it as local. A session +using it would therefore be rejected before Pi starts even if the model appeared in +the selector. + +## Approved outcome + +The selector must expose exactly these three configured models: + +1. `zai/glm-5.2` +2. `deepseek/deepseek-v4-flash` +3. `local-qwen/qwen3.6-35b-a3b` + +`zai/glm-5v-turbo` must remain hidden. The selected approach is to use Pi's +`enabledModels` setting as the single source of truth rather than introduce a second +application-specific allowlist. + +The models must be usable for new sessions; merely displaying them is insufficient. + +## Source of truth and precedence + +The backend reads Pi settings from the same two scopes Pi uses: + +- global: `~/.pi/agent/settings.json`; +- project: `/.pi/settings.json`. + +If project settings define `enabledModels`, that value overrides the global value. +Otherwise the global value applies. This mirrors Pi's settings precedence for the +field used here without attempting to reimplement unrelated Pi settings behavior. + +For this integration, `enabledModels` entries must be exact `provider/model` strings. +Wildcards, ambiguous model-only patterns, and thinking-level suffixes are outside the +selector contract. Invalid entries are ignored with a warning; the backend never +expands them into additional visible models. + +The live global Pi settings will be updated to contain only the three approved exact +identifiers, removing `zai/glm-5v-turbo`. + +## Backend model-list flow + +`createPiModelLister` continues to start an ephemeral `pi --mode rpc` process in +`harnessDir` and request `get_available_models`. Its environment construction changes +as follows: + +- scrub ambient deployment secrets and provider credentials; +- do not require or inject the generic credential merely to enumerate models; +- let Pi resolve configured authentication through its mounted profile + (`~/.pi/agent/auth.json`) and custom model definitions; +- preserve the existing portable data-root handling. + +After Pi responds, the backend: + +1. maps the Pi response to the public `PiModel` shape; +2. loads the effective `enabledModels` list; +3. matches models by the composite `provider/model` identifier; +4. returns only matched models, in `enabledModels` order; +5. caches the filtered result using the existing short TTL. + +Filtering after Pi discovery ensures that an enabled identifier is shown only if Pi +also considers the corresponding model available. + +## Session credential behavior + +The existing provider isolation remains in place for hosted providers. DeepSeek is +already present in Pi's `auth.json`; Pi gives profile credentials precedence over an +environment fallback, so selecting `deepseek/deepseek-v4-flash` uses its configured +DeepSeek credential. + +`local-qwen` is added to the explicit set of local providers. Its endpoint and request +configuration remain owned by Pi's `models.json`; no generic hosted-provider key is +required or injected for its session process. + +No general exception is added for unknown providers. An unrecognized provider still +fails closed before session spawn. + +## API validation + +`PUT /settings` validates the composite `provider/model` pair whenever the filtered +model list is non-empty. Matching only `model.id` is insufficient because different +providers may expose the same identifier. + +The public shape of `GET /models` and the frontend API contract remain unchanged: + +```json +{ + "models": [ + { "provider": "zai", "id": "glm-5.2", "name": "GLM-5.2", "reasoning": true } + ] +} +``` + +No frontend component change is expected: once the endpoint returns the three models, +the existing selector can render them and persist both provider and model. + +## Failure behavior and observability + +The model list fails closed to an empty array when any of these conditions applies: + +- the effective `enabledModels` field is missing, empty, or malformed; +- neither Pi settings file can be read successfully when one is expected; +- the ephemeral Pi process fails, times out, or returns an invalid response; +- none of the enabled identifiers is currently available to Pi. + +The route keeps its graceful `{ "models": [] }` response for frontend compatibility, +but emits a sanitized warning through the Fastify logger. Logs may include file paths, +provider/model identifiers, and error classes; they must never include credential +values or the contents of `auth.json`. + +## Tests + +Backend regression tests cover: + +- model enumeration when `PI_PROVIDER` is absent and a generic credential file exists; +- exact filtering and `enabledModels` ordering; +- project `enabledModels` overriding the global list; +- missing, malformed, and invalid settings producing an empty list/error path without + leaking all Pi-available models; +- `zai/glm-5v-turbo` being excluded; +- composite provider/model validation in `PUT /settings`; +- `local-qwen` spawning without a generic provider credential; +- DeepSeek remaining selectable through its configured profile authentication. + +Existing frontend tests remain the compatibility gate. A focused frontend test is +added only if inspection reveals that the three-model response is not already covered. + +## Deployment and verification + +Implementation completion requires: + +1. updating the mounted live Pi `settings.json` to the approved three-entry list; +2. running backend unit tests and TypeScript type checking; +3. running any affected frontend tests/type checking if frontend code changes; +4. rebuilding and recreating the impacted `core` container; +5. verifying container health; +6. verifying live `GET /models` returns exactly the three approved composite IDs; +7. verifying a model-setting update accepts DeepSeek and Qwen and rejects the hidden + GLM-5V model without leaving application settings altered after the smoke test. + +## Out of scope + +- Changing Pi's built-in model registry. +- Exposing every model with a configured credential. +- Supporting wildcard or fuzzy `enabledModels` patterns in the web selector. +- Adding a frontend model-management UI. +- Generalizing custom-provider credential modes beyond the explicit `local-qwen` + requirement.