pi-moa
Mixture of Agents for pi: run multiple high-intelligence reference models in parallel, then synthesize with a main model that keeps full conversation context.
Package details
Install pi-moa from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:pi-moa- Package
pi-moa- Version
0.3.2- Published
- Aug 14, 2026
- Downloads
- 110/mo · 22/wk
- Author
- wangding2020
- License
- Apache-2.0
- Types
- extension, skill
- Size
- 63.4 KB
- Dependencies
- 0 dependencies · 3 peers
Pi manifest JSON
{
"extensions": [
"./extensions"
],
"skills": [
"./skills"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
pi-moa
Mixture of Agents (MOA) for pi, recreating Nous Research Hermes Agent's MOA workflow as a slash command.
- 1–3 high-intelligence reference models run in parallel, each seeing only a single-turn context (no tools, no system prompt).
- 1 main model receives the labeled reference opinions plus the full conversation history, and answers normally with tool access. By default this is your current UI-selected model; you can optionally pin a fixed aggregator.
- Models and auth are reused from pi's existing configuration — the config
file stores only
provider/modelidentifiers, never API keys. - Optional auto mode (
/moa-auto): when enabled, every message you type runs through the MOA pipeline automatically — no need to prefix/moa.
What it does
One /moa <prompt> runs in two stages:
- Fan-out. Each reference model (1–3) gets a single-turn context — the current user message only, plus a short conversation summary on multi-turn sessions. No tools, no system prompt, no history. They answer independently and in parallel.
- Aggregation. The main model receives every reference answer (labeled by source) together with the full conversation history, then produces the final response — with normal tool access, exactly like a regular turn.
The main model defaults to your current UI-selected model; set aggregator to
pin a fixed one.
Why it helps
- Cross-checks one model's blind spots. Independent models give complementary or conflicting first-pass opinions, which is more signal for the aggregator than a single model reasoning alone.
- More robust on hard or ambiguous questions. Multiple angles reduce the chance of a confident-but-wrong single-model answer.
- Second opinions without manual effort. No need to paste the prompt into several separate chats — one command fans out and converges automatically.
- No extra setup. Reference models reuse models already configured in pi;
the config file stores only
provider/modelidentifiers, never keys.
Install
pi install npm:pi-moa
# or install from a local path during development
pi install ./moa
# or copy into ~/.pi/agent/extensions/ + ~/.pi/agent/skills/
Then run /moa once in an interactive terminal to complete the first-run wizard
(pick 1–3 reference models; the main model follows your current UI model unless
you opt to pin a fixed one). Config is written to ~/.pi/agent/moa.json.
Usage
/moa <prompt> run MOA once with the default preset
/moa usage + config summary (wizard if unconfigured)
/moa list list presets
/moa configure [name] interactive create/edit (default: "default")
/moa delete <name> delete a preset
/moa status show config + current model
/moa set-main <p>/<m> pin a fixed main model (default: follow UI model)
/moa-auto toggle auto mode (every message runs MOA)
Config
{
"default_preset": "default",
"presets": {
"default": {
"reference_models": [
{ "provider": "anthropic", "model": "claude-sonnet-4-5", "reasoning_effort": "low" },
{ "provider": "openrouter", "model": "deepseek/deepseek-v3.1" }
],
"aggregator": { "provider": "anthropic", "model": "claude-opus-4-5" },
"reference_max_tokens": 4000,
"max_tokens": 4096
}
}
}
The aggregator field is optional. When omitted, MOA uses your current
UI-selected model as the main output model (no model switching, no restore); set
it only to pin a fixed main model regardless of the UI.
Other optional per-preset fields: reference_temperature,
reasoning_effort (per reference slot), max_tokens (reserved; not applied in
v1 — pi governs the main model's token budget).
Behavior notes & limitations (v0.3.2)
- Main model follows the UI. By default the main output model is your
current UI-selected model (
ctx.model) — no model switching and no restore. Only when a fixedaggregatoris configured does MOA temporarily switch to it and then restore. - Auto mode.
/moa-autotoggles a persistent "every message runs MOA" state. It intercepts user input from both the TUI (source === "interactive") and web/RPC frontends such as pi-web (source === "rpc") and rewrites it into the MOA message before the agent loop. Slash commands and/moa's internalsendUserMessagecalls (source === "extension") are unaffected, so there is no double fan-out. The status bar showsMOA 自动:开while active. - Scrollable picker. The config wizard uses a scrollable
SelectList(max 10 visible +(n/total)indicator) instead ofctx.ui.select, so long model lists no longer overflow the window. - Conversation summary (multi-turn). When the session already has history,
/moafirst asks the main model to summarize "what are we working on" (100–200 words), then feeds that summary to the reference models instead of the bare user prompt. If the summary fails, it degrades to the bare prompt. - One-shot semantics (
fanout: user_turn). References run once per/moa, not per tool-loop iteration. - Model restore. With a fixed
aggregator, the pre-MOA model is restored onagent_settled; queued follow-ups during the MOA turn still run on the aggregator before restore. - Failure handling. A failed reference is reported inline and the flow
continues; if all references fail,
/moaaborts before running the main model. - No privacy filter. Reference output is injected into the main model as-is
(Hermes'
privacy_filteris not implemented). - Headless/print mode. The wizard is unavailable; run
/moainteractively first, or edit~/.pi/agent/moa.jsonby hand. - Prompt-cache interaction (U8) and no file locking on concurrent
moa.jsonwrites are not addressed in v1.