pi-llama-server
Pi extension for llama-server router - model discovery, auto-load, per-project config
Package details
Install pi-llama-server from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:pi-llama-server- Package
pi-llama-server- Version
1.1.0- Published
- Jul 6, 2026
- Downloads
- 1,012/mo · 103/wk
- Author
- am17an
- License
- MIT
- Types
- extension
- Size
- 2.6 MB
- Dependencies
- 0 dependencies · 1 peer
Pi manifest JSON
{
"extensions": [
"./extensions/llama-server.ts"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
Note: This is how I use pi.dev + llama.cpp on my local machine. I created a plugin so that I can update my setup quickly.
pi-llama-server
Pi extension that integrates a running llama-server instance with the Pi Coding Agent. Discovers llama-server models and automatically loads the selected model when you switch models in Pi.
Demo

Prerequisites
- A running llama-server instance (from llama.cpp) in
router-mode(the default if you don't mention-m) - Pi Coding Agent installed (
@earendil-works/pi-coding-agent)
Install
pi install npm:pi-llama-server
Or from git:
pi install git:github.com/user/pi-llama-server
Pi auto-discovers the extension via pi.extensions in package.json. No additional setup needed.
Configuration
The llama-server URL is resolved in this order:
- Per-project config — create
.pi/llama-server.jsonin your project root:{ "url": "http://10.0.0.5:9090" } - Environment variable — set globally:
export LLAMA_SERVER_URL=http://10.0.0.5:9090 - Default — falls back to
http://127.0.0.1:8080
Usage
Use Ctrl+P (or /model) in Pi to select any llama-server model for inference. Pi switches to that model, and the extension automatically tells llama-server to load it. While llama-server reports loading progress, Pi shows a progress bar in the footer status.
How it works
When Pi starts, the extension:
- Resolves the llama-server URL from config/env/default
- Queries
GET /modelsto discover available GGUF models - Registers each model as an OpenAI-compatible provider under
{url}/v1 - Listens for model switch events and calls
POST /models/loadon the server - Listens to
GET /models/ssewhile a selected model is loading to show footer progress
llama-server endpoints used
| Endpoint | Method | Purpose |
|---|---|---|
/models |
GET | List all models |
/models/load |
POST | Load a model |
/models/sse |
GET | Stream model status/progress events |
/v1/... |
POST | OpenAI-compatible completions (via Pi provider) |