pi-model-swap

llama-swap model-swap controller for pi sessions — refusal-gated, claim-file coordinated, dry-run first

Packages

Package details

extension

Install pi-model-swap from npm and Pi will load the resources declared by the package manifest.

$ pi install npm:pi-model-swap
Package
pi-model-swap
Version
0.1.2
Published
Oct 6, 2026
Downloads
not available
Author
adamjen
License
MIT
Types
extension
Size
75.8 KB
Dependencies
0 dependencies · 3 peers
Pi manifest JSON
{
  "extensions": [
    "./extensions"
  ]
}

Security note

Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.

README

pi-model-swap

Model-swap controller for pi sessions driven by llama-swap. It moves a session from one llama-swap model to another model-first, thinking-second, and refuses when the swap would evict a model that another live session is using.

/swap --status
/swap --dry-run orchestrator
/swap orchestrator medium
model_swap { model: "orchestrator", dry_run: true }

Install

pi install npm:pi-model-swap

Requires a running llama-swap controller and a pi provider that fronts it. The extension never starts, stops, signals, or kills llama-swap — it only issues requests to the controller endpoint.

Linux only (the census needs ss from iproute2 and /proc). First time here? Read SETUP.md — the llama.cpp → llama-swap → roster → models.json → config chain, with the port-macro rule and the id-matching rule.

Reference stack: a hybrid roster — llama.cpp models plus a Strata-served MoE model (swift-1.5) as the session default. That is the deployment this extension was built and tested on, and the reason the refusal gate and the cold-start warning exist: see SETUP.md → Worked example and examples/pi-model-swap.strata.json.

Configure

Write ~/.pi/agent/pi-model-swap.json (or point PI_MODEL_SWAP_CONFIG at a file). The extension is fail-closed: if this file is missing or does not parse, every mutator refuses and the report names the path it tried. Every key below is required — there are no fallback values for the machine-specific ones.

Key Default Meaning
controllerUrl required llama-swap controller — the only inference endpoint used
controllerPort required its port (census allow-list: the controller's own pid)
rosterPath required only source of upstream ports; parsed for proxy:
llamaSwapConfigPath required source of startPort for the ${PORT} macro
startPort required ${PORT} base = startPort + roster slot index
provider required pi provider fronting the controller
sessionDefault required restore target, bare id under provider
strataPort required corroboration only (/slots), never a port source
verifyTimeoutMs 960000 client bound on the step-7 verify request
probeTimeoutMs 5000 GET /running timeout (never inherits the 16-min verify bound)
claimStalenessS 1800 claim-file staleness outer bound
boundaryRestore true announce at one turn boundary, execute at the next

Surface

Command /swap [model] [level] [--dry-run] [--go] [--restore] [--status] Tool model_swap { model, level?, dry_run?, go?, restore? } — the tool path and the command path produce identical output.

Levels: off minimal low medium high xhigh max. Omit level to keep the current one.

Ordering invariant: setModel → setThinkingLevel → verify request. Mirrors pi's own examples/extensions/preset.ts.

Safety model

  • Dry-run first. --dry-run / dry_run: true prints the plan and touches nothing.
  • Refusal gate. A swap is refused when its eviction set contains a model a live session is bound to (claim file), or when a census of that model's upstream port shows an established connection whose owner pid is neither the model's listener owner nor the controller's pid. Unattributable connection ⇒ refuse (default-deny).
  • The allow-list is pid identity, never a process name. /proc/<pid>/comm can be set by any local process via prctl(PR_SET_NAME), so a name match is evidence of nothing.
  • The census must be READ, not merely answered. A 200 on GET /running whose body is not the documented shape is a census failure: the eviction set is unknown, so the swap refuses. An empty set counts as clean only when the body itself says "nothing is running".
  • Claim files in ~/.pi/agent/state/swap-state/<pid>.json, phases idle | armed | verifying | committed | rolling_back | restoring. One swap in flight.
  • Cold start is not a fast rollback. Evicting a resident model means a full reload (minutes). The warning names the models this swap can actually evict; a same-model no-op (empty eviction set) prints no warning.
  • No pgrep -f, no pkill -f, no rm, no port polling, no llama-swap model-load endpoint.

Known limitations (v0.1.2)

  • All machine-specific keys are required (controllerUrl, controllerPort, rosterPath, llamaSwapConfigPath, startPort, provider, sessionDefault, strataPort). There are no fallback values for them: with no config the extension refuses and says so.
  • No per-model thinking-level cap: a level a model cannot honour is offered, not validated.
  • Claim staleness (1800 s) is an outer bound, not a liveness check.
  • The refusal gate and writeClaim are not atomic (TOCTOU window between census and claim write).

Development

bash tests/run.sh          # build + functional harness (mock controller) + gate + typecheck
bash tests/t007-live-check.sh <session.jsonl>   # verify a live dry-run session's own records
```The harness runs against a mock controller and the real socket census. It never starts, stops,
signals, or kills a real server.

## License

MIT