pi-llama-server

Pi extension for llama-server router - model discovery, auto-load, per-project config

Packages

Package details

extension

Install pi-llama-server from npm and Pi will load the resources declared by the package manifest.

$ pi install npm:pi-llama-server
Package
pi-llama-server
Version
1.1.0
Published
Jul 6, 2026
Downloads
1,012/mo · 103/wk
Author
am17an
License
MIT
Types
extension
Size
2.6 MB
Dependencies
0 dependencies · 1 peer
Pi manifest JSON
{
  "extensions": [
    "./extensions/llama-server.ts"
  ]
}

Security note

Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.

README

Note: This is how I use pi.dev + llama.cpp on my local machine. I created a plugin so that I can update my setup quickly.

pi-llama-server

Pi extension that integrates a running llama-server instance with the Pi Coding Agent. Discovers llama-server models and automatically loads the selected model when you switch models in Pi.

Demo

Demo

Prerequisites

  • A running llama-server instance (from llama.cpp) in router-mode (the default if you don't mention -m)
  • Pi Coding Agent installed (@earendil-works/pi-coding-agent)

Install

pi install npm:pi-llama-server

Or from git:

pi install git:github.com/user/pi-llama-server

Pi auto-discovers the extension via pi.extensions in package.json. No additional setup needed.

Configuration

The llama-server URL is resolved in this order:

  1. Per-project config — create .pi/llama-server.json in your project root:
    { "url": "http://10.0.0.5:9090" }
    
  2. Environment variable — set globally:
    export LLAMA_SERVER_URL=http://10.0.0.5:9090
    
  3. Default — falls back to http://127.0.0.1:8080

Usage

Use Ctrl+P (or /model) in Pi to select any llama-server model for inference. Pi switches to that model, and the extension automatically tells llama-server to load it. While llama-server reports loading progress, Pi shows a progress bar in the footer status.

How it works

When Pi starts, the extension:

  1. Resolves the llama-server URL from config/env/default
  2. Queries GET /models to discover available GGUF models
  3. Registers each model as an OpenAI-compatible provider under {url}/v1
  4. Listens for model switch events and calls POST /models/load on the server
  5. Listens to GET /models/sse while a selected model is loading to show footer progress

llama-server endpoints used

Endpoint Method Purpose
/models GET List all models
/models/load POST Load a model
/models/sse GET Stream model status/progress events
/v1/... POST OpenAI-compatible completions (via Pi provider)