@arhen/pi-core-subagent
pi extension: fast in-process subagents with a dependency-graph scheduler (needs edges gate tasks and carry upstream output into dependent prompts), plus background runs, intercom and agent-to-agent mailbox. Leader defines agents inline.
Package details
Install @arhen/pi-core-subagent from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:@arhen/pi-core-subagent- Package
@arhen/pi-core-subagent- Version
1.3.68- Published
- Oct 8, 2026
- Downloads
- 3,649/mo · 1,457/wk
- Author
- arhen
- License
- MIT
- Types
- extension
- Size
- 231.9 KB
- Dependencies
- 0 dependencies · 5 peers
Pi manifest JSON
{
"extensions": [
"./src/index.ts"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
@arhen/pi-core-subagent
Install
Requires the pi coding agent — install it first: npm install -g @earendil-works/pi-coding-agent.
pi install npm:@arhen/pi-core-subagent
# or locally: pi install /path/to/pi-subagents
Minimalist pi extension: fast in-process subagents with single / parallel / graph modes, background runs, cancellation, intercom (child↔leader) and an agent↔agent mailbox.
Built for one job: delegate work to isolated subagents without bloating the parent context.
One rule underneath everything else:
A task is a graph, not a checklist. Nodes are workers, edges are data dependencies. The edge both gates the dependent and hands it the upstream output — so "the coordinator forgot to pass X" stops being a failure mode. Where there are no edges, there is no graph: flat fan-out stays flat.
That is Graph Protocol, applied to the runtime rather than to the prompt.
flowchart LR
subgraph w1["wave 1 — runs in parallel"]
api["api<br/><i>api-mapper</i>"]
db["db<br/><i>db-mapper</i>"]
end
gate{{"gate"}}
subgraph w2["wave 2"]
doc["doc<br/><i>writer</i>"]
end
api -- "route map" --> gate
db -- "schema map" --> gate
gate -- "both outputs<br/>prepended to the prompt" --> doc

The subagent tool call plus the live above-editor widget: per-agent activity, tool counts, turns, token counters and timers.
Design principles
- The delegation is the graph.
needsdeclares edges; the scheduler runs each wave of ready tasks in parallel and gates the rest. One code path for single, parallel, chain and graph —chainis justneeds: [previous]. (why) - Edges carry data, not just order. An upstream task's output is prepended to its dependents' prompts automatically. The coordinator cannot forget to pass it, because it never passes it.
- A bad graph fails before it spawns. Unknown ids, self-edges and cycles are rejected at call time — never halfway through a run with three children already burning tokens.
- Proof is an exit code, never a self-report. Tasks are asked for a runnable
Verify:command; the leader checksgit diff --stat. Agents auditing their own work score ~0. (why) - No ceremony without edges. Six independent reviewers stay six independent reviewers — no waves, no gates, no graph vocabulary imposed on flat work.
- Agent files respected. A spawn goal (name + task) that matches a user agent file's
description(.agents/agents,.claude/agents,.pi/agents— project then home) loads that file — body = system prompt, frontmattermodel/toolsapply, filemodelvalidated against the pi model registry. File wins over inline; no match → on-demand definition. - Two toolsets, plus explicit override. Read-only (
read, grep, find, ls, codemode— default) or write (read, grep, find, ls, bash, edit, write, codemode—write: true);tools:sets an explicit per-task allowlist. Children run withnoExtensions, socodemodeis the one extension they get: it lets a child batchtools.*calls instead of spending a model turn per call. - In-process — children are
AgentSessions in the same runtime. No process spawn, no context bleed. - Zero parent-context injection. No catalog, no context hook. 9 slim tools total.
- Throttled updates — widget/stream updates coalesce to ~6/s; no per-event deep clones.
- Bounded always — a child is limited only by wall clock: explicit
maxRuntimeMs, else the 1 h ceiling with/subagents auto-limit on, else the 6 h safety ceiling (default).
How it runs
Children are not subprocesses. They are separate AgentSessions inside the same pi process — which is why spawning is instant, and why a child's transcript never lands in your context:
flowchart TB
subgraph proc["one OS process — no spawn, no IPC"]
direction TB
L["<b>leader</b><br/>your session, your context"]
subgraph kids["isolated child sessions"]
direction LR
A["api-mapper"]
B["db-mapper"]
end
end
L -- "task text in" --> A
L -- "task text in" --> B
A -. "final answer only" .-> L
B -. "final answer only" .-> L
The dotted arrows are the whole point: a child may burn 200k tokens reading files, and the leader receives only its final answer.
Usage — the leader invents the agents
Define agents inline per call, or reference a named agent file (see Agent files). Model resolution: explicit model → agent-file model (validated against the pi model registry) → the parent's current model → settings default. Any model that is not the session model is preflight-probed; if the provider rejects it, the task falls back to the session model and the result carries a Model: note.
{
"agent": "api-reviewer",
"prompt": "You are a strict API reviewer. Check auth, rate limiting, and error handling. Cite file:line.",
"task": "Review src/api/upload.ts"
}
Parallel — mixed toolsets, siblings can talk via mailbox (intercom is always on):
{
"tasks": [
{ "agent": "researcher", "prompt": "You find facts. Cite paths.", "task": "Map the auth flow", "write": false },
{ "agent": "implementer", "prompt": "You make minimal changes.", "task": "Implement POST /api/upload", "write": true }
]
}
Chain — {previous} is replaced with the prior agent's output:
{
"chain": [
{ "agent": "planner", "prompt": "You write a step list.", "task": "Plan the change", "write": false },
{ "agent": "doer", "prompt": "You follow the plan exactly.", "task": "Execute: {previous}", "write": true }
]
}
Agent files
A user agent file in an agents directory is matched by its description frontmatter against the spawn goal (agent name + task) — not by name. When matched, the file is authoritative: body = system prompt, frontmatter model/tools apply, inline prompt/model are ignored — with one exception: explicit per-call tools/write override the file's tools (the file narrows defaults, it never displaces explicit intent, and it can never widen past the leader's read/write choice). An override is surfaced on the task's notice and summary. No match → the inline on-demand definition stands. The model stays in control: it names the agent and states the goal; user files that describe that goal take over.
---
name: api-reviewer
description: reviews APIs for auth, rate limiting, and error handling
model: claude-opus-4-6
tools: read, grep, find, ls
---
You are a strict API reviewer. Check auth, rate limiting, and error handling. Cite file:line.
Lookup order (first directory with a match wins):
.agents/agents/then.claude/agents/then.pi/agents/in each directory from the taskcwdup to the filesystem root (nearest ancestor wins).- Home:
~/.agents/agents/(single source) →~/.claude/agents/→~/.pi/agents/.
Within a directory the file with the highest description-overlap score wins (≥2 shared meaningful tokens and ≥40% coverage of the shorter token set). A file model is validated against the pi model registry (unknown model fails the task with a catalog message). File bodies are capped at 64,000 characters. Files without a description frontmatter never match.
Worktree isolation (write agents)
In a git repo, a write: true subagent runs in an isolated git worktree at <repo>/.git/subagents/<run>/<task> on branch subagents/<run>/<task> — the child's cwd is the worktree, so project context (AGENTS.md chain) still loads, node_modules is symlinked, and the main tree stays clean while the child works. Parallel write agents can't collide on files.
On completion the extension commits the child's changes (the child is told not to touch branches) and reports branch + diffstat + changed files in the result. The leader reviews, then merges:
git merge --no-ff subagents/<run>/<task>
Branch relationships — read this before merging. A write task that needs a completed write task is stacked: its worktree branches from the upstream's branch, so the child actually sees the files its upstream wrote. A stacked branch contains its upstream's commits, so merge order doesn't matter — merging the stacked branch brings both, and the upstream's own merge is then a no-op.
Why this matters (verified against real git, not just reasoned about): when a downstream child can't see its upstream's work, merging produces spurious conflicts, half-clobbered files where the child's side looks like a phantom delete, and — worst — clean merges that leave a broken tree. If the upstream renamed login→signIn and the downstream wrote new code importing login, git reports success with exit 0 and the code doesn't compile. Stacking removes that class by construction.
Same-wave write tasks are siblings: both branch from the same base, so they are independent. When two siblings changed the same file the summary emits CONFLICT RISK naming the overlap — the second git merge is a real 3-way and fails loudly, which is the safe outcome. The dangerous case is siblings touching different but coupled files: that merges clean and breaks at runtime. Nothing can detect it for you.
Dependencies are shared, not isolated. node_modules is symlinked to the main checkout, so dependency writes escape the worktree: children are instructed never to install, upgrade, or delete deps. A task that genuinely needs a dependency change should edit the manifest and say so.
Isolation follows the toolset the child actually receives: explicit tools: ["bash", "edit", "write"] earns a worktree even without write: true, and an agent file that narrows the child to read-only gets no branch at all.
Cleanup, in order of trust:
| When | What |
|---|---|
| Session start | Registered worktrees can't be live yet — an interrupted child's uncommitted work is committed, the branch kept, the dir dropped. Then merged branches are reaped and dirs git no longer tracks are removed. |
| After a run | Merged branches + their dirs, never touching a branch a live run owns (checkout state re-checked per branch). |
| Task failed/canceled | Partial work is committed first, so the branch keeps it; the dir is dropped. |
Non-git repos fall back to in-place edits.
Graph mode — needs
parallel runs everything at once; chain runs everything one at a time. Most real work is neither. Give a task an id and list the ids it needs:
{
"tasks": [
{ "id": "api", "agent": "api-mapper", "task": "Map every route in src/api/" },
{ "id": "db", "agent": "db-mapper", "task": "Map the schema in src/db/" },
{ "id": "doc", "agent": "writer", "needs": ["api", "db"], "write": true,
"task": "Write ARCHITECTURE.md from the maps above. Verify: test -s ARCHITECTURE.md" }
]
}
The call line renders the graph in §2 notation as the model types it:
subagent graph 3
wave1[api ∥ db] → gate → wave2[doc]
api api-mapper Map every route in src/api/
db db-mapper Map the schema in src/db/
doc writer ✎ ← api, db Write ARCHITECTURE.md from the maps above. Verify: test -s ARCHI…
✎ marks a write-toolset task; ← lists its edges. With no needs anywhere the wave line is omitted entirely.
What one edge does
An edge is not just ordering. It is a delivery:
sequenceDiagram
participant S as scheduler
participant A as api
participant D as db
participant W as doc
Note over S,D: wave 1 — both start together
S->>A: "Map every route in src/api/"
S->>D: "Map the schema in src/db/"
A-->>S: route map
Note right of W: doc is queued,<br/>waiting at the gate
D-->>S: schema map
Note over S: gate opens: every need settled
S->>W: ## Output of api<br/><route map><br/><br/>## Output of db<br/><schema map><br/>---<br/>"Write ARCHITECTURE.md…"
The leader never copies those outputs into the prompt — so it cannot forget to.
One scheduler, four shapes
Single, parallel, chain and graph are not four code paths. They are four shapes of the same wave loop:
flowchart LR
subgraph one["single"]
direction TB
s1(("a"))
end
subgraph par["parallel — no needs"]
direction TB
p1(("a")) ~~~ p2(("b")) ~~~ p3(("c"))
end
subgraph ch["chain — needs: [previous]"]
direction TB
c1(("a")) --> c2(("b")) --> c3(("c"))
end
subgraph gr["graph — needs"]
direction TB
g1(("a")) --> g2(("b"))
g1 --> g3(("c"))
g2 --> g4(("d"))
g3 --> g4
end
The loop
flowchart TD
start(["subagent call"]) --> validate{"graph valid?<br/><small>unknown id · self-edge · cycle</small>"}
validate -- no --> reject["reject the call<br/><b>zero children spawned</b>"]
validate -- yes --> loop{"tasks left?"}
loop -- no --> done(["run finished"])
loop -- yes --> ready["frontier =<br/>tasks whose needs are all settled"]
ready --> spawn["run that wave in parallel<br/><small>throttled by concurrency</small>"]
spawn --> collect["record each output<br/>mark tasks settled"]
collect --> loop
Two consequences worth stating plainly:
- A bad graph costs nothing. Validation happens before the first spawn, never halfway through with three children already burning tokens.
- A broken upstream stops its branch. If a need fails or is aborted, its dependents are marked aborted rather than run against a prompt with a hole in it:
flowchart LR
api["api ✓"] --> doc
db["db ✗ failed"] --> doc["doc ⏹ skipped<br/><small>never spawned</small>"]
And the rule that keeps this from becoming ceremony: zero needs anywhere = plain parallel. No waves, no gates, no graph vocabulary imposed on flat work.
Background (default) + intercom — the run returns a runId immediately; you stay steerable while it works. Children can always ask you questions and message each other:
{
"agent": "auditor",
"prompt": "You audit dependencies.",
"task": "Audit package.json for outdated deps"
}
Steering a running child: while a background run is active the leader stays responsive, and you can push a message into a live child's session mid-run with steer_subagent — e.g. steer_subagent({ runId, taskId, message: "Ignore tests/, only audit runtime deps" }). The message queues as a steer if the child is mid-turn and lands at its next model boundary. Omit taskId to steer every still-running task in the run. Combined with notifyPerTask, this makes a background run feel like a live team you can redirect, not a fire-and-forget blob.
Resuming a failed child: a child that dies mid-work (provider rate limit, timeout, network error) keeps its session JSONL and its worktree branch. resume_subagent({ runId, taskId, model?: "openai/gpt-5", message? }) reopens that session with full context, re-attaches the branch, and prompts it to recap and continue — no respawn, no lost tokens. model swaps provider when the original one is exhausted. Refused for tasks that never started (no session file); those you respawn. Wait for the run to settle before resuming (the tool tells you if it hasn't).
Choosing a model: subagent_models lists what this session may use — the same set /scoped-models shows when scoping is configured, and every model with usable credentials when it is not (the output states which case applies). Each row gives the exact model value to pass, the thinking levels the runtime honors, the context window, and pi's catalog price per million tokens (free only when the reported rates are zero; absent rates read unavailable; catalog rates are a planning guide, not a billing quote). model is optional: omit it to inherit the leader's current session model, or name one (a matched agent file's model frontmatter still wins) to pin the run. ~/.pi/agent/subagent-models.json may shape the listing only — prefer sorts, hide omits, default is surfaced as a suggestion and never applied. It never grants or blocks a model: a hidden model still runs when named, because pi owns what may run.
Tools
| Tool | Purpose |
|---|---|
subagent_models |
list the models a subagent task may name, each with the exact model value, thinking levels the runtime honors, context window, and catalog price; scoped to the session's enabled set when scoping is configured, else the full available catalogue (the output says which) |
subagent |
single / tasks (parallel or graph via needs) / chain ({previous}); every run is background — returns a runId, completion notifies you; autoAwait:true parks the call until the run finishes and returns the final result inline; children always carry talk tools (ask/notify/mailbox); notifyPerTask (default true) wakes you as each task completes |
subagent_status |
live per-task snapshot (non-blocking), including each child's session file path; call it once right after spawning — a child that died on spawn is invisible until far later otherwise |
subagent_result |
full output of a run or one task |
await_subagent |
block until a run finishes (optional timeoutMs) |
reply_subagent |
answer a child's ask_parent question |
steer_subagent |
inject a steering message into a running child's session (queues as steer if mid-turn; lands at its next model boundary) |
resume_subagent |
revive a failed/aborted task in its original session (context + branch preserved); optional model swap, thinking override (the task's stored level is clamped to what the target model accepts), custom message. Waits for the reopened session to go idle before prompting, so a task killed mid-turn can still be resumed |
subagent_cancel |
abort a running/queued run |
Per-task fields
agent (name you invent — required), task (required), prompt (system prompt, optional — minimal default used), write (toolset, default read-only), plus optional model (provider/model-id; omitted → the leader's current session model, a matched agent file's frontmatter wins), thinking (validated enum: off|minimal|low|medium|high|xhigh|max), tools (explicit allowlist), cwd, maxRuntimeMs, id, needs (dependency edges — see Graph mode). Top-level only: autoAwait, notifyPerTask, concurrency (default 3, max 8). A call carries at most 16 tasks.
Child talk tools (always on)
| Tool | Meaning |
|---|---|
ask_parent |
blocking question to the leader; delivered mid-turn as a steering message labelled [URGENT] or [not urgent], parent answers via reply_subagent |
notify_parent |
one-way message to the leader; identical updates coalesce, final: true declares the complete report |
send_agent_message |
message to a sibling subagent's mailbox (to = its task id, or "leader") |
poll_agent_messages |
drain this subagent's mailbox |
Intercom anti-deadlock: children are told to never block indefinitely on intercom replies — an unanswered
ask_parenttimes out after 10 minutes (the child is told to proceed with best judgment), and sibling polls are capped (~5 tries) with the same fallback. Gated siblings (later waves) may not be running yet — waiting on them is the top stall cause, so children are instructed not to.
Ask urgency:
ask_parenttakesurgent(defaultfalse). Both variants steer into the leader's current turn so the question is never deferred to the end of a long turn.[URGENT]tells the leader to answer before its next step;[not urgent]tells it that the child keeps waiting, so it may finish its current step first. Failures steer for the same reason. Completions, aborts, and informational updates are held in an extension outbox and delivered together as one follow-up when the leader settles; a receipt confirms against the finalized leader transcript, so a report the leader already consumed (or an awaited result orsubagent_resultthat already showed the outcome) suppresses its redundant queued notice. See docs/notification-policy.md.
Tool exposure and opt-in codemode
Default: legacy direct calls, even when the codemode tool is active. Active helpers remain
model-visible and retain native script-call compatibility. Explicit global codemode.mode: "only"
still applies Pi's global script-only policy; this extension does not override it.
/subagents mode— show preference and effective profile./subagents mode auto— opt in to script-based routing while codemode is active; direct otherwise./subagents mode codemode— explicitly select script-based routing (direct fallback when unavailable)./subagents mode direct— return to legacy direct behavior.
Choice is stored globally in ~/.pi/agent/subagents-config.json (the same file as auto-limit,
written with a merge, never a clobber) and applies to every session — reload, resume, fork and tree
navigation included. Sessions that chose a mode before the preference moved into that file keep their
branch entry only while the config has no mode; the built-in default is still direct. Published
1.3.63 defaulted to auto: this compatibility correction requires the corrected source/package, not
a retroactive change to that npm release.
Commands
/subagents— list runs;/subagents peek(orctrl+shift+a) — browsable pane/subagents auto-limit on|off— toggle leader-imposedmaxRuntimeMscaps (persists to~/.pi/agent/subagents-config.json; default off).ongives tasks without an explicitmaxRuntimeMsthe 1 h default ceiling;offraises the ceiling to 6 h (still a ceiling — an unbounded child would pin the run forever). Bare/subagents auto-limitshows the current state.
Peek — /subagents peek or ctrl+shift+a
Read-only pane over the session's subagents:
shift+↑/shift+↓(orj/k) — move between agents; bare arrows work too where the terminal doesn't reserve thementer— live tail of that child's session file (escgoes back)xtheny— abort ONE subagent (only mutation;n/any other key cancels)esc— close
Watching a child from outside
A child has no terminal of its own — but it does write a real transcript file, and that file is the seam every external viewer can use:
flowchart LR
child["child session<br/><small>no TTY</small>"] -- writes --> file[("session.jsonl")]
file -- "peek · enter" --> pane["in-pi tail"]
file -- "tail -f" --> term["any terminal pane<br/><small>herdr · tmux · zellij</small>"]
subagent_status returns that path for every running child:
tail -f /path/from/subagent_status.jsonl
In a terminal multiplexer, that is a pane per agent — e.g. with Herdr:
herdr pane split --current --direction right
herdr pane run w1:p2 "tail -f /path/from/subagent_status.jsonl"
The extension has no multiplexer integration and does not want one: it exposes the path, your agent already knows how to drive its own terminal. For an in-pi view of the same stream, use /subagents peek.
Context budget
- Parent tools: 9 schemas with short descriptions. No catalog, no context hook — nothing injected per request.
- The widget shows live work only: settled runs are pruned from it and stay reachable through
subagent_status/subagent_result. - Background completion: 3-line notice. Full text only via
subagent_result. - Children: isolated sessions; talk tools always injected; each child's prompt states its own task id and its siblings' so mailbox addressing works. Model resolution: explicit
provider/model-idor bare id via the pi model registry → the parent's current model → settings default. Thinking levels validated against the resolved model'sthinkingLevelMap;subagent_modelslists the levels the runtime honors when you need a safe set.
What this is built on
Graph Protocol
needs is an implementation of Graph Protocol — a delegation discipline that treats a task as a graph (Delegation<A, E, R>) rather than a checklist. Its ten sections map onto this extension as follows:
| § | Protocol | Here |
|---|---|---|
| §1 | nodes, domains, edges | one agent owns one task; needs are the edges |
| §2 | happy path as execution graph, waves + gates | wave scheduler: ready = tasks whose needs are settled |
| §3 | one worker or many | single vs tasks |
| §4 | break points: wrong context, missing input, misinterpretation | missing input is structurally impossible — the edge carries the output |
| §5 | R: subgraph, method, verification command, WHY | prompt guidelines require a runnable Verify: line per task |
| §6 | structured at the boundary | in: upstream outputs prepended as named blocks. out: prose (see below) |
| §7 | observe without changing the graph | the widget and /subagents peek are read-only |
| §8 | worker attention acquired and released | spawn → terminal status; aborted upstream releases dependents immediately |
| §9 | prove it: delegated vs implemented | deliberately not self-reported — see below |
| §10 | prompt = subgraph, return = implemented graph | prompt yes; return kept as prose |
Why §9 is a verification command, not a self-report
The protocol asks the coordinator to compare the delegated subgraph against the graph the worker says it implemented. We implement the comparison against the filesystem and the exit code, not against the worker's account of itself, because self-reports carry close to zero signal about exactly the failure §9 exists to catch:
- Asked to audit its own work against 34 real violations, an agent reported 0 — at 90–100 confidence. A fresh instance of the same model shown the same output caught 7 (p = 0.0156). A deterministic checker caught all 34. (Armalo Labs, 2026)
- Across 9,876 τ2-bench and 1,879 AppWorld trajectories, "false success" reached 75.8% of self-assessing coding-agent failures; adding an LLM judge scored 0.54–0.65 AUROC (0.5 = coin flip). (arXiv:2606.09863)
- LLM judges reading agent traces can be flipped by rewriting the trace — the exact surface a self-reported graph exposes. (arXiv:2601.14691)
The shape of the problem:
flowchart TD
W["worker finishes"] --> Q{"who says it's correct?"}
Q -- "the worker itself" --> S["self-report<br/><b>0 of 34 caught</b><br/><small>at 90–100 confidence</small>"]
Q -- "another model reading the trace" --> J["LLM judge<br/><b>0.54–0.65 AUROC</b><br/><small>0.5 = coin flip</small>"]
Q -- "the machine" --> D["exit code + git diff<br/><b>34 of 34 caught</b>"]
So §9 in practice is two things you already have:
# in the task text — the worker must prove it, not claim it
Verify: npx tsc --noEmit && bun test
# in the leader, after the run — ground truth, not narrative
git diff --stat
If files outside a worker's subgraph were touched, the diff says so. A structured return schema would only add a second, less trustworthy witness.
Why waves instead of "more agents"
Flat fan-out is not free — orchestration cost is critical path + α × cross-agent communication, and ignoring the second term is what makes added agents lose to a single one:
- Dependency-graph partitioning vs flat file-parallel spawning across 28 real repos: +14.0% pass rate, 2.10× wall-clock, −35% API cost, with the largest gains on the most dependency-dense projects. Flat parallel inflated cost 60% for a 1.56× speedup; an agent-team baseline was fastest but scored below sequential on code quality. (arXiv:2606.00953)
- Dynamic task graphs across 300 trials: 47.5% of baseline token cost, 79.7% accuracy vs 57.6% for a static graph — and, notably, static tied dynamic when the structure was genuinely known up front, which is the case
needstargets. The "frontier" (ready set) in this scheduler is theirs. (arXiv:2605.06320)
The corollary is in the design principles: when there are no edges, don't draw a graph. Six independent reviewers stay six independent reviewers.
Development
bun install # dev deps (typecheck/test only; runtime uses pi's bundled SDK)
npx tsc --noEmit
bun test # pure-logic tests (wave scheduling, edge payload, mailbox, failure classification, watchdog)
Runtime state: runs persist to <parent-session>.subagents.json sidecar; restored (non-terminal → aborted) on session start.
License
MIT.