pi-model-switch
Model switching extension for pi coding agent
Package details
Install pi-model-switch from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:pi-model-switch- Package
pi-model-switch- Version
0.3.0- Published
- Sep 27, 2026
- Downloads
- 1,096/mo · 485/wk
- Author
- nicopreme
- License
- MIT
- Types
- extension
- Size
- 23 KB
- Dependencies
- 0 dependencies · 3 peers
Pi manifest JSON
{
"extensions": [
"./index.ts"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
pi-model-switch
A Pi coding agent extension for direct model switching.
It provides one tool, switch_model, for identifying, listing, searching, and directly switching models, and warns in the footer when any model switch would re-bill a warm prompt cache.
Foreground orchestration now lives in pi-orchestrate.
Installation
pi install npm:pi-model-switch
Restart Pi to load the extension.
Tool
switch_model
Parameters:
action:current | list | search | switchsearch?: query forsearchandswitchprovider?: provider filterthinkingLevel?:minimal | low | medium | high | xhigh | max
Behavior:
current: shows the active model without listing every modellist: shows available authenticated modelssearch: filters by provider, id, or nameswitch: resolves aliases first, then does exact or partial model matching, and optionally applies a thinking level
Prompt cache warning
Switching models is free, but the next request re-bills the whole conversation on the new model because its prompt cache is cold. When the current model's cache is still warm and the re-bill is significant (at least 20k tokens or $0.10), the extension shows a footer status such as:
⚠ next request re-bills ~120k cached tokens (~$0.54) on anthropic/claude-haiku-4-5
This covers /model, Ctrl+P cycling, and switch_model, and clears once the next request starts. After /model or Ctrl+P, switch back before sending to avoid the cost. When the agent calls switch_model mid-run, the next request goes out right away, so the cost is already paid; the tool result includes the same note so the agent can tell you.
A model's cache counts as warm when its last request on the current branch used the prompt cache (or pi refreshed it) within the model's cache lifetime. The lifetime comes from the model's promptCache metadata, using the long tier when PI_CACHE_RETENTION=long, and defaults to 5 minutes. Compaction resets it. Returning to a model whose cache is still warm does not warn.
The cost is an estimate: the current context size priced at the new model's cache-write (or input) rate minus its cache-read rate.
Aliases
Aliases can be defined in the first existing file from these locations, in priority order:
{project}/.pi/aliases.json~/.pi/agent/aliases.json~/.pi/agent/extensions/model-switch/aliases.json
The extension loads aliases when the tool runs, so project aliases use the current working directory. list reports the selected alias file path.
For example, define aliases in the extension-level file:
~/.pi/agent/extensions/model-switch/aliases.json
{
"cheap": "google/gemini-2.5-flash",
"coding": {
"model": "anthropic/claude-opus-4-5",
"thinkingLevel": "high"
},
"budget": [
{ "model": "openai/gpt-5-mini", "thinkingLevel": "low" },
"google/gemini-2.5-flash"
]
}
Rules:
- top-level value must be an object
- alias names must be non-empty strings
- each target must be a
provider/modelIdstring or an object containingmodeland optionalthinkingLevel - string alias: one exact model target (backward compatible)
- object alias:
{ "model": "provider/modelId", "thinkingLevel": "high" } - array alias: fallback chain; first available authenticated target wins
- explicit
thinkingLevelonswitch_modeloverrides the alias setting - omit
thinkingLevelto preserve Pi's existing model-switch behavior; Pi clamps unsupported levels per model and the tool reports the effective level
License
MIT