pi-uva-hva
Pi coding-agent provider extension for the UvA/HvA LiteLLM proxy. Adds a /login connect flow (pick base URL, paste API key), discovers every model with real capabilities from /model_group/info, routes each to the right endpoint (Responses vs chat-completi
Package details
Install pi-uva-hva from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:pi-uva-hva- Package
pi-uva-hva- Version
1.4.0- Published
- Jul 14, 2026
- Downloads
- 144/mo · 144/wk
- Author
- clearlynotmich
- License
- MIT
- Types
- extension
- Size
- 67.3 KB
- Dependencies
- 0 dependencies · 0 peers
Pi manifest JSON
{
"extensions": [
"./index.ts"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
Pi extension (uva-hva-agentic-tools)
Part of uva-hva-agentic-tools. This is the Pi guide; for other tools see Claude Code, VS Code, OpenCode, Aider, Kilo Code, Factory Droid, or Odysseus.
A Pi coding-agent provider extension for the University of
Amsterdam / Amsterdam University of Applied Sciences LiteLLM proxies
(https://llmproxy.uva.nl/v1, https://llmproxy.hva.nl/v1).
Get every UvA/HvA model working in Pi in under a minute, without editing Pi's config
files. Load the extension, run /login, pick your proxy, paste your API key, and
it auto-discovers all available models so you can select one straight from
/models. Your base URL and key are saved, so the next launch just reconnects.
Reasoning models come pre-tuned (thinking on at medium),
and tool-heavy agent turns that would otherwise silently come back empty on this
proxy just work. That's the whole setup; everything below is optional detail.
Install
You need Pi installed and a UvA/HvA proxy API key.
1. Load the extension
Pick one:
- From npm (easiest): add
"npm:pi-uva-hva"to thepackagesarray in~/.pi/agent/settings.json(or runpi install npm:pi-uva-hva). Pi installs it on the next launch, no clone needed:{ "packages": [ "npm:pi-uva-hva" ] } - From a local clone: clone the repo and point at the
pi/folder instead:{ "packages": [ "/path/to/uva-hva-agentic-tools/pi" ] } - Try it once:
pi -e /path/to/uva-hva-agentic-tools/pi - Auto-load folder: copy
index.tsto~/.pi/agent/extensions/.
2. Connect with /login
Start Pi and run:
/login
Choose "UvA / HvA proxy", pick a base URL (UvA / HvA / custom), and paste
your API key. The extension discovers every model and saves your base URL + key
to ~/.pi/agent/openai-responses-uva.json, so the next launch reconnects with
nothing re-entered. /uva-login runs the same flow.
Prefer env vars? Set
UVA_API_KEY(and optionallyUVA_BASE_URL) instead of/login. Both work.
3. Pick a model
/models # select any discovered model
Reasoning models default to thinking ON at medium. You can raise or lower a
model's default in /configure-models, or set the level per run from the CLI:
pi --provider uva --model gpt-5.6-sol --thinking high -p "hello"
Configuration
/login is the easy path; everything is also configurable via environment
variables (all optional):
| Variable | Default | Purpose |
|---|---|---|
UVA_API_KEY |
- | API key, if you prefer env over /login. |
UVA_BASE_URL |
https://llmproxy.uva.nl/v1 |
Proxy base URL (must end at the /v1 root). |
UVA_PROVIDER_ID |
uva |
Provider id shown in /models and --provider. |
UVA_CREDENTIALS_FILE |
~/.pi/agent/openai-responses-uva.json |
Override the saved-credentials path. |
UVA_NO_AUTO_THINKING |
- | Set to disable the medium thinking default. |
UVA_MODEL_OVERRIDES_FILE |
- | Path to a JSON file overriding per-model capabilities (below). |
Per-model overrides
Capabilities come from the proxy metadata (with a name-table fallback). To pin a value yourself there are two ways.
Interactive menu (easiest): run /configure-models. Pick a model, then set
its context window and max output (type the number), toggle reasoning on/off,
and choose the default thinking level applied when that model is selected. Two
options at the end, Save & apply changes and Discard & exit. Saved
overrides are written to ~/.pi/agent/openai-responses-uva.models.json and
applied live (and on every future launch).
By hand: point UVA_MODEL_OVERRIDES_FILE at a JSON file (this overrides the
default path above):
{
"gpt-5.6-sol": { "contextWindow": 1100000, "maxTokens": 128000, "defaultThinkingLevel": "high" },
"some-new-model": { "reasoning": false, "input": ["text", "image"], "contextWindow": 200000, "maxTokens": 32000 }
}
The defaultThinkingLevel on gpt-5.6-sol above is just an example of pinning
one model to high; by default every reasoning model starts at medium.
Each key is a model id; each value may set any of reasoning,
defaultThinkingLevel (off/low/medium/high), input, contextWindow,
maxTokens, vision, name, cost, thinkingLevelMap.
How it works
For each turn the custom stream handler:
- Lets Pi's pristine built-in
openai-responseshandler build the exact request params (full message + tool conversion, reasoning, caching) via anonPayloadhook that captures the params and throws before the network call, so nothing extra is billed. - Reissues the request itself and synthesizes Pi's content events
(
text/thinking/toolCall), reconciling the incremental SSE events with the terminalresponse.output[]by item id so a collapsed tool call is never lost. Tool-call ids and signatures are preserved for multi-turn replay.
Two dispatch paths keep it reliable:
- Non-reasoning turns stream (
stream:true): bytes flow, so the nginx gateway read-timeout keeps resetting and output is token-by-token. - Reasoning turns use background + poll (
background:true+GET /responses/{id}): the model can buffer its whole reasoning phase with zero interim bytes without ever tripping a 504, because each request is short. A streaming turn that still hits a gateway 5xx before any output falls back to this path automatically.
Params incompatible with non-OpenAI backends (prompt_cache_key on
Bedrock/Vertex) are stripped per-model.
Robust to model changes
The UvA/HvA line-up changes often, so nothing about specific models is
hard-coded. On connect the extension reads the proxy's own
/model_group/info and derives each model's context window, max output,
reasoning/vision support, cost, and backend directly from it, then decides the
endpoint from the backend (azure/bedrock speak the Responses API;
open-weight openai/vLLM models use chat-completions). If that metadata
endpoint is ever unavailable it falls back to /v1/models plus a researched
name table, and if a model is still mis-routed, a Responses turn that 404s is
self-healed at runtime (retried on chat-completions and remembered). So even
if every current model is replaced with new ones, discovery, capabilities, and
routing keep working with no code change.
Trade-off
Reasoning replies are not token-by-token; they appear at once on completion (the proxy only delivers background results as a single terminal payload). Non-reasoning turns stream normally.
Reasoning models
Reasoning is enabled per model from the proxy metadata (supports_reasoning or a
reasoning_effort parameter), with the name table as a fallback, but only on the
Responses route (chat-completions rejects reasoning_effort alongside tools).
Reasoning models default to thinking ON at medium (disable with
UVA_NO_AUTO_THINKING=1). Override any model's capabilities, including its
default thinking level, per-id via UVA_MODEL_OVERRIDES_FILE or
/configure-models.
Compatibility
- Pi coding-agent with the
@earendil-works/pi-airuntime (verified on 0.80.x). - Imports only the extension-facing
@earendil-works/pi-aisurface, so it keeps working across Pi updates (it does not patchnode_modules).
Recommended extensions
Pi becomes far more capable with a few extensions. These are the ones I run alongside this provider, grouped by what they do, with install commands below.
Context, memory & compaction
These three cover the three layers with one tool each, so nothing overlaps:
| Extension | Layer | What it does |
|---|---|---|
| context-mode | Tool output | Offloads large tool/command output into a sandbox and a searchable store so it never floods your context window. |
| pi-smart-compact | History | Verification-oriented conversation compaction: deterministic extract, then synthesize, then verify what survived. |
| gentle-engram | Memory | Persistent memory shared across sessions, compactions, and MCP agents. |
Planning, workflows & subagents
| Extension | What it does |
|---|---|
| @juicesharp/rpiv-pi | Skill-based dev workflow: discover → research → design → plan → implement → validate → review. |
| pi-code-planner | Structured planning, TDD, and Git worktrees for local coding agents. |
| @gotgenes/pi-subagents | In-process sub-agent core with a typed API and lifecycle events. |
| @quintinshaw/pi-dynamic-workflows | Fan a task across hundreds of subagents with real model routing and cost accounting. |
| @juicesharp/rpiv-workflow | Chain skills into typed multi-stage workflows with audited state. |
| @juicesharp/rpiv-todo | A live todo overlay for the model that survives /reload and compaction. |
Providers & models
| Extension | What it does |
|---|---|
| pi-multi-account | Automatic multi-account failover & rotation across Anthropic, OpenAI, Qwen, Ollama. |
| glm-vision | Gives non-vision GLM models (z.ai) image understanding via GLM-4.6V. |
Tools & UX
| Extension | What it does |
|---|---|
| @juicesharp/rpiv-web-tools | Web search + fetch for the model with pluggable providers (Brave, Tavily, Exa, …). |
| pi-mcp-adapter | Connect any MCP (Model Context Protocol) server to Pi. |
| @amaster.ai/pi-computer-use | Desktop automation via computer_use_* tools. |
| pi-image-paste | Turns pasted image paths into first-class image attachments. |
| @trevonistrevon/pi-loop | Cron/event re-wake loops and background process monitoring. |
| @juicesharp/rpiv-ask-user-question | Lets the model ask you a structured, typed questionnaire instead of guessing. |
| @juicesharp/rpiv-advisor | A second opinion the model can request from a stronger reviewer model. |
| @juicesharp/rpiv-args | $1 / $ARGUMENTS placeholders and shell substitution in skills. |
| @juicesharp/rpiv-i18n | Localization foundation for the rpiv-* skills (/languages, --locale). |
Code quality
| Extension | What it does |
|---|---|
| pi-simplify | Reviews recently changed code for clarity, consistency, and maintainability. |
| ponytail | Lazy senior dev mode: stops the agent over-engineering and writes the minimum code that works, with /ponytail review/audit/debt commands. |
Install them
Pi auto-installs anything listed in the packages array of
~/.pi/agent/settings.json on the next launch, with no manual npm install.
All at once: merge this into your packages array (keep any entries you
already have), then restart Pi:
{
"packages": [
"npm:context-mode",
"npm:pi-smart-compact",
"npm:gentle-engram",
"npm:@juicesharp/rpiv-pi",
"npm:pi-code-planner",
"npm:@gotgenes/pi-subagents",
"npm:@quintinshaw/pi-dynamic-workflows",
"npm:@juicesharp/rpiv-workflow",
"npm:@juicesharp/rpiv-todo",
"npm:pi-multi-account",
"npm:glm-vision",
"npm:@juicesharp/rpiv-web-tools",
"npm:pi-mcp-adapter",
"npm:@amaster.ai/pi-computer-use",
"npm:pi-image-paste",
"npm:@trevonistrevon/pi-loop",
"npm:@juicesharp/rpiv-ask-user-question",
"npm:@juicesharp/rpiv-advisor",
"npm:@juicesharp/rpiv-args",
"npm:@juicesharp/rpiv-i18n",
"npm:pi-simplify",
"npm:opencode-ponytail"
]
}
One by one: add a single line to the same packages array and restart Pi.
Each entry is just "npm:<name>", e.g.:
"packages": [
"npm:context-mode"
]
Setup notes (extensions with a background DB or service)
Most of these are pure extensions that work the moment they're in packages. A
few run a background database or service and need one extra thing to work on a
single install:
Node ≥ 22.5.0, required by
context-mode. It stores its knowledge base in SQLite and relies on Node's built-innode:sqlite(older Node falls back to a native module that can crash). Check withnode -v. Itspostinstallauto-wires the Pi hooks and heals the native binding, so no manual setup is needed. Just verify afterward with/context-mode:ctx-doctor(ornpx context-mode doctor). If an install ever complains aboutbetter-sqlite3, upgrading Node to 22.5+ is the fix. Known quirk: if you also have a~/.claudefolder, context-mode may store its knowledge base there instead of under~/.pi; harmless, but that is where to look for it.gentle-engramneeds the Engram backend. The npm package is only the Pi bridge: persistence is handled by a separateengrambinary that it launches as an MCP server. Install Engram from Gentleman-Programming/engram and make sureengramis on yourPATH(or setENGRAM_BIN); otherwise the memory tools load but nothing is saved. Also keep only one engram entry inpackages:npm:gentle-engram, not a second pinned copy likenpm:gentle-engram@0.1.8.
Some of these overlap in purpose.
rpiv-todogives the model itstodotracker;pi-loopalso ships a nativeTaskCreatefallback that only activates when no dedicated task system is present, so it just sits unused besidetodo(harmless, not a conflict). Several planning and workflow engines overlap too. Start with the context/memory group, then add planning and tools as you need them rather than enabling all of them at once.
License
MIT. See LICENSE.