Files
ThothII/docs/superpowers/specs/2026-07-14-pi-enabled-model-selector-design.md
T

7.1 KiB

Pi-Enabled Model Selector — Design

Date: 2026-07-14 Status: Approved (design), pending implementation plan Layers: backend, deployment configuration; frontend contract unchanged

Problem

The model selector currently offers only the active zai/glm-5.2 model. The live GET /models response is { "models": [] }, so the frontend applies its existing degraded-mode fallback and renders only the model already stored in application settings.

The regression was introduced by the provider-credential hardening in commit f064dae. The ephemeral Pi model lister now constructs its child environment using cfg.defaults.provider. In the live deployment PI_PROVIDER is intentionally absent: the selected provider lives in the persisted application settings file. Because a generic model credential file is configured but no environment default provider is available, buildPiChildEnv throws model provider credential is unavailable before Pi starts. The route catches the failure and returns an empty list.

There is a second usability defect behind the first one: local-qwen is a custom local provider, but the credential policy does not classify it as local. A session using it would therefore be rejected before Pi starts even if the model appeared in the selector.

Approved outcome

The selector must expose exactly these three configured models:

  1. zai/glm-5.2
  2. deepseek/deepseek-v4-flash
  3. local-qwen/qwen3.6-35b-a3b

zai/glm-5v-turbo must remain hidden. The selected approach is to use Pi's enabledModels setting as the single source of truth rather than introduce a second application-specific allowlist.

The models must be usable for new sessions; merely displaying them is insufficient.

Source of truth and precedence

The backend reads Pi settings from the same two scopes Pi uses:

  • global: ~/.pi/agent/settings.json;
  • project: <harnessDir>/.pi/settings.json.

If project settings define enabledModels, that value overrides the global value. Otherwise the global value applies. This mirrors Pi's settings precedence for the field used here without attempting to reimplement unrelated Pi settings behavior.

For this integration, enabledModels entries must be exact provider/model strings. Wildcards, ambiguous model-only patterns, and thinking-level suffixes are outside the selector contract. Invalid entries are ignored with a warning; the backend never expands them into additional visible models.

The live global Pi settings will be updated to contain only the three approved exact identifiers, removing zai/glm-5v-turbo.

Backend model-list flow

createPiModelLister continues to start an ephemeral pi --mode rpc process in harnessDir and request get_available_models. Its environment construction changes as follows:

  • scrub ambient deployment secrets and provider credentials;
  • do not require or inject the generic credential merely to enumerate models;
  • let Pi resolve configured authentication through its mounted profile (~/.pi/agent/auth.json) and custom model definitions;
  • preserve the existing portable data-root handling.

After Pi responds, the backend:

  1. maps the Pi response to the public PiModel shape;
  2. loads the effective enabledModels list;
  3. matches models by the composite provider/model identifier;
  4. returns only matched models, in enabledModels order;
  5. caches the filtered result using the existing short TTL.

Filtering after Pi discovery ensures that an enabled identifier is shown only if Pi also considers the corresponding model available.

Session credential behavior

The existing provider isolation remains in place for hosted providers. DeepSeek is already present in Pi's auth.json; Pi gives profile credentials precedence over an environment fallback, so selecting deepseek/deepseek-v4-flash uses its configured DeepSeek credential.

local-qwen is added to the explicit set of local providers. Its endpoint and request configuration remain owned by Pi's models.json; no generic hosted-provider key is required or injected for its session process.

No general exception is added for unknown providers. An unrecognized provider still fails closed before session spawn.

API validation

PUT /settings validates the composite provider/model pair whenever the filtered model list is non-empty. Matching only model.id is insufficient because different providers may expose the same identifier.

The public shape of GET /models and the frontend API contract remain unchanged:

{
  "models": [
    { "provider": "zai", "id": "glm-5.2", "name": "GLM-5.2", "reasoning": true }
  ]
}

No frontend component change is expected: once the endpoint returns the three models, the existing selector can render them and persist both provider and model.

Failure behavior and observability

The model list fails closed to an empty array when any of these conditions applies:

  • the effective enabledModels field is missing, empty, or malformed;
  • neither Pi settings file can be read successfully when one is expected;
  • the ephemeral Pi process fails, times out, or returns an invalid response;
  • none of the enabled identifiers is currently available to Pi.

The route keeps its graceful { "models": [] } response for frontend compatibility, but emits a sanitized warning through the Fastify logger. Logs may include file paths, provider/model identifiers, and error classes; they must never include credential values or the contents of auth.json.

Tests

Backend regression tests cover:

  • model enumeration when PI_PROVIDER is absent and a generic credential file exists;
  • exact filtering and enabledModels ordering;
  • project enabledModels overriding the global list;
  • missing, malformed, and invalid settings producing an empty list/error path without leaking all Pi-available models;
  • zai/glm-5v-turbo being excluded;
  • composite provider/model validation in PUT /settings;
  • local-qwen spawning without a generic provider credential;
  • DeepSeek remaining selectable through its configured profile authentication.

Existing frontend tests remain the compatibility gate. A focused frontend test is added only if inspection reveals that the three-model response is not already covered.

Deployment and verification

Implementation completion requires:

  1. updating the mounted live Pi settings.json to the approved three-entry list;
  2. running backend unit tests and TypeScript type checking;
  3. running any affected frontend tests/type checking if frontend code changes;
  4. rebuilding and recreating the impacted core container;
  5. verifying container health;
  6. verifying live GET /models returns exactly the three approved composite IDs;
  7. verifying a model-setting update accepts DeepSeek and Qwen and rejects the hidden GLM-5V model without leaving application settings altered after the smoke test.

Out of scope

  • Changing Pi's built-in model registry.
  • Exposing every model with a configured credential.
  • Supporting wildcard or fuzzy enabledModels patterns in the web selector.
  • Adding a frontend model-management UI.
  • Generalizing custom-provider credential modes beyond the explicit local-qwen requirement.