@sidshaytay/pi-llama-swap
Pi extension that discovers llama-swap models and registers them as a provider.
Package details
Install @sidshaytay/pi-llama-swap from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:@sidshaytay/pi-llama-swap- Package
@sidshaytay/pi-llama-swap- Version
0.1.1- Published
- Sep 24, 2026
- Downloads
- 393/mo · 251/wk
- Author
- sidshaytay
- License
- MIT
- Types
- extension
- Size
- 13.1 KB
- Dependencies
- 0 dependencies · 0 peers
Pi manifest JSON
{
"extensions": [
"./extensions"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
@sidshaytay/pi-llama-swap
A pi extension that discovers the models available on your llama-swap server and registers them as a pi provider — no manual model catalog to maintain.
Why this extension
Other llama-swap extensions for pi exist. This one is built around not getting in llama-swap's way:
- Model switching never blocks. Discovery talks only to
/v1/models— it never probes endpoints that would wake a sleeping llama-swap instance. - pi starts even when llama-swap doesn't. If the server is unreachable at
startup, pi starts with an empty catalog and retries the next time you open
/model. - Sensible reasoning and vision defaults. Reasoning is on for capable
families (Qwen3.6+, Gemma 4, GPT-OSS, DeepSeek) and off for
-no-thinking/-uncensoredvariants. Vision is detected from llama-swap's capability reporting. - Your tweaks survive refreshes. Per-model overrides in
models.jsonare re-applied every time the catalog refreshes. - Config lives in pi, not your shell. One optional key in the pi agent
directory's
settings.jsonper machine; environment variables override it when you need that.
Quick start
Install:
pi install npm:@sidshaytay/pi-llama-swapOther ways to install:
pi install npm:@sidshaytay/pi-llama-swap@0.1.0 # pinned version pi install git:github.com/SidShaytay/pi-extensions # from the repoTell it where llama-swap listens — skip this if it runs at
localhost:8080. In~/.pi/agent/settings.json:{ "llama-swap": { "url": "http://llama-swap.example.com" } }Restart pi, then verify:
pi --list-models 2>&1 | grep -A30 llama-swapYou should see the
llama-swapprovider with your chat models. Pick one with/modelin pi.
Configuration
All settings are optional. Put them under the "llama-swap" key in
settings.json in pi's agent directory — ~/.pi/agent/settings.json by
default, or $PI_CODING_AGENT_DIR/settings.json when that variable is set
(the extension resolves the directory the same way pi does, so sandboxes
like pi-less-yolo, which mount
the agent dir at /pi-agent, read the same config as the host). This file
is per-machine by design and is not synced between machines. Each key has a
matching environment variable that overrides it.
| settings.json key | Environment variable | Default | Purpose |
|---|---|---|---|
url |
LLAMA_SWAP_URL |
http://localhost:8080 |
llama-swap base URL (no trailing /v1) |
provider |
LLAMA_SWAP_PROVIDER |
llama-swap |
provider ID registered in pi |
apiKey |
LLAMA_SWAP_API_KEY |
llama-swap-local |
placeholder API key |
contextWindow |
LLAMA_SWAP_CONTEXT_WINDOW |
128000 |
fallback context window for unloaded models |
maxTokens |
LLAMA_SWAP_MAX_TOKENS |
32768 |
default max output tokens per model |
exclude |
LLAMA_SWAP_EXCLUDE |
built-in regex (see below) | regex of model IDs to hide |
Per-model overrides
The extension registers every model it discovers. To adjust one — thinking
format, compat flags, reasoning, context window — add a providers.llama-swap
block to ~/.pi/agent/models.json:
{
"providers": {
"llama-swap": {
"compat": {
"supportsDeveloperRole": false,
"supportsReasoningEffort": false,
"supportsStrictMode": false,
"maxTokensField": "max_tokens"
},
"modelOverrides": {
"qwen3.6-35b-a3b": {
"compat": { "thinkingFormat": "qwen-chat-template" }
},
"qwen3.5-9b": { "reasoning": false },
"gemma-4-e4b": { "contextWindow": 131072 }
}
}
}
}
Overrides set here are preserved across catalog refreshes.
How models are classified
- Reasoning: on by default. IDs matching
-no-thinking,nothinking, or-uncensoredget reasoning off. llama-swap doesn't report reasoning capability over/v1/models, so the extension uses this heuristic; override any model inmodels.jsonwith"reasoning": true/false. - Vision: set to
["text", "image"]when/v1/modelsreportsimageinarchitecture.input_modalities. - Context window: taken from
context_lengthwhen the model reports it; otherwise thecontextWindowdefault applies. - Non-chat models are hidden by the default exclude regex: image and
diffusion models, TTS, whisper/asr, and common embedding families
(
bge,nomic,mxbai,e5,clip). Set your ownexcludeto change it.
To make vision and context window reliable even when a model is unloaded,
declare capabilities in llama-swap's config.yaml:
mymodel:
capabilities:
in: [text, image]
context: 131072
Troubleshooting
The llama-swap provider is empty or missing.
The URL is wrong or llama-swap is unreachable from this machine. Check with
curl http://your-host:8080/v1/models, then open /model in pi to refresh.
A model shows the wrong context window or thinking behavior.
The model was probably unloaded at discovery time, or its family isn't covered
by the reasoning heuristic. Declare capabilities.context in llama-swap's
YAML and/or set modelOverrides in models.json.
A model I want isn't listed.
It matched the exclude regex. Set a custom exclude in settings.json or
LLAMA_SWAP_EXCLUDE.