@sidshaytay/pi-llama-swap

Pi extension that discovers llama-swap models and registers them as a provider.

Packages

Package details

extension

Install @sidshaytay/pi-llama-swap from npm and Pi will load the resources declared by the package manifest.

$ pi install npm:@sidshaytay/pi-llama-swap
Package
@sidshaytay/pi-llama-swap
Version
0.1.1
Published
Sep 24, 2026
Downloads
393/mo · 251/wk
Author
sidshaytay
License
MIT
Types
extension
Size
13.1 KB
Dependencies
0 dependencies · 0 peers
Pi manifest JSON
{
  "extensions": [
    "./extensions"
  ]
}

Security note

Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.

README

@sidshaytay/pi-llama-swap

A pi extension that discovers the models available on your llama-swap server and registers them as a pi provider — no manual model catalog to maintain.

Why this extension

Other llama-swap extensions for pi exist. This one is built around not getting in llama-swap's way:

  • Model switching never blocks. Discovery talks only to /v1/models — it never probes endpoints that would wake a sleeping llama-swap instance.
  • pi starts even when llama-swap doesn't. If the server is unreachable at startup, pi starts with an empty catalog and retries the next time you open /model.
  • Sensible reasoning and vision defaults. Reasoning is on for capable families (Qwen3.6+, Gemma 4, GPT-OSS, DeepSeek) and off for -no-thinking / -uncensored variants. Vision is detected from llama-swap's capability reporting.
  • Your tweaks survive refreshes. Per-model overrides in models.json are re-applied every time the catalog refreshes.
  • Config lives in pi, not your shell. One optional key in the pi agent directory's settings.json per machine; environment variables override it when you need that.

Quick start

  1. Install:

    pi install npm:@sidshaytay/pi-llama-swap

    Other ways to install:

    pi install npm:@sidshaytay/pi-llama-swap@0.1.0        # pinned version
    pi install git:github.com/SidShaytay/pi-extensions    # from the repo
  2. Tell it where llama-swap listens — skip this if it runs at localhost:8080. In ~/.pi/agent/settings.json:

    {
      "llama-swap": {
        "url": "http://llama-swap.example.com"
      }
    }
  3. Restart pi, then verify:

    pi --list-models 2>&1 | grep -A30 llama-swap

    You should see the llama-swap provider with your chat models. Pick one with /model in pi.

Configuration

All settings are optional. Put them under the "llama-swap" key in settings.json in pi's agent directory — ~/.pi/agent/settings.json by default, or $PI_CODING_AGENT_DIR/settings.json when that variable is set (the extension resolves the directory the same way pi does, so sandboxes like pi-less-yolo, which mount the agent dir at /pi-agent, read the same config as the host). This file is per-machine by design and is not synced between machines. Each key has a matching environment variable that overrides it.

settings.json key Environment variable Default Purpose
url LLAMA_SWAP_URL http://localhost:8080 llama-swap base URL (no trailing /v1)
provider LLAMA_SWAP_PROVIDER llama-swap provider ID registered in pi
apiKey LLAMA_SWAP_API_KEY llama-swap-local placeholder API key
contextWindow LLAMA_SWAP_CONTEXT_WINDOW 128000 fallback context window for unloaded models
maxTokens LLAMA_SWAP_MAX_TOKENS 32768 default max output tokens per model
exclude LLAMA_SWAP_EXCLUDE built-in regex (see below) regex of model IDs to hide

Per-model overrides

The extension registers every model it discovers. To adjust one — thinking format, compat flags, reasoning, context window — add a providers.llama-swap block to ~/.pi/agent/models.json:

{
  "providers": {
    "llama-swap": {
      "compat": {
        "supportsDeveloperRole": false,
        "supportsReasoningEffort": false,
        "supportsStrictMode": false,
        "maxTokensField": "max_tokens"
      },
      "modelOverrides": {
        "qwen3.6-35b-a3b": {
          "compat": { "thinkingFormat": "qwen-chat-template" }
        },
        "qwen3.5-9b": { "reasoning": false },
        "gemma-4-e4b": { "contextWindow": 131072 }
      }
    }
  }
}

Overrides set here are preserved across catalog refreshes.

How models are classified

  • Reasoning: on by default. IDs matching -no-thinking, nothinking, or -uncensored get reasoning off. llama-swap doesn't report reasoning capability over /v1/models, so the extension uses this heuristic; override any model in models.json with "reasoning": true/false.
  • Vision: set to ["text", "image"] when /v1/models reports image in architecture.input_modalities.
  • Context window: taken from context_length when the model reports it; otherwise the contextWindow default applies.
  • Non-chat models are hidden by the default exclude regex: image and diffusion models, TTS, whisper/asr, and common embedding families (bge, nomic, mxbai, e5, clip). Set your own exclude to change it.

To make vision and context window reliable even when a model is unloaded, declare capabilities in llama-swap's config.yaml:

mymodel:
  capabilities:
    in: [text, image]
    context: 131072

Troubleshooting

The llama-swap provider is empty or missing. The URL is wrong or llama-swap is unreachable from this machine. Check with curl http://your-host:8080/v1/models, then open /model in pi to refresh.

A model shows the wrong context window or thinking behavior. The model was probably unloaded at discovery time, or its family isn't covered by the reasoning heuristic. Declare capabilities.context in llama-swap's YAML and/or set modelOverrides in models.json.

A model I want isn't listed. It matched the exclude regex. Set a custom exclude in settings.json or LLAMA_SWAP_EXCLUDE.

License

MIT