@arhen/pi-core-subagent

pi extension: fast in-process subagents with a dependency-graph scheduler (needs edges gate tasks and carry upstream output into dependent prompts), plus background runs, intercom and agent-to-agent mailbox. Leader defines agents inline.

Packages

Package details

extension

Install @arhen/pi-core-subagent from npm and Pi will load the resources declared by the package manifest.

$ pi install npm:@arhen/pi-core-subagent
Package
@arhen/pi-core-subagent
Version
1.3.68
Published
Oct 8, 2026
Downloads
3,649/mo · 1,457/wk
Author
arhen
License
MIT
Types
extension
Size
231.9 KB
Dependencies
0 dependencies · 5 peers
Pi manifest JSON
{
  "extensions": [
    "./src/index.ts"
  ]
}

Security note

Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.

README

@arhen/pi-core-subagent

npm version CI license pi extension

Install

Requires the pi coding agent — install it first: npm install -g @earendil-works/pi-coding-agent.

pi install npm:@arhen/pi-core-subagent
# or locally: pi install /path/to/pi-subagents

Minimalist pi extension: fast in-process subagents with single / parallel / graph modes, background runs, cancellation, intercom (child↔leader) and an agent↔agent mailbox.

Built for one job: delegate work to isolated subagents without bloating the parent context.

One rule underneath everything else:

A task is a graph, not a checklist. Nodes are workers, edges are data dependencies. The edge both gates the dependent and hands it the upstream output — so "the coordinator forgot to pass X" stops being a failure mode. Where there are no edges, there is no graph: flat fan-out stays flat.

That is Graph Protocol, applied to the runtime rather than to the prompt.

flowchart LR
    subgraph w1["wave 1 — runs in parallel"]
        api["api<br/><i>api-mapper</i>"]
        db["db<br/><i>db-mapper</i>"]
    end
    gate{{"gate"}}
    subgraph w2["wave 2"]
        doc["doc<br/><i>writer</i>"]
    end
    api -- "route map" --> gate
    db -- "schema map" --> gate
    gate -- "both outputs<br/>prepended to the prompt" --> doc

Subagents widget

The subagent tool call plus the live above-editor widget: per-agent activity, tool counts, turns, token counters and timers.

Design principles

  • The delegation is the graph. needs declares edges; the scheduler runs each wave of ready tasks in parallel and gates the rest. One code path for single, parallel, chain and graph — chain is just needs: [previous]. (why)
  • Edges carry data, not just order. An upstream task's output is prepended to its dependents' prompts automatically. The coordinator cannot forget to pass it, because it never passes it.
  • A bad graph fails before it spawns. Unknown ids, self-edges and cycles are rejected at call time — never halfway through a run with three children already burning tokens.
  • Proof is an exit code, never a self-report. Tasks are asked for a runnable Verify: command; the leader checks git diff --stat. Agents auditing their own work score ~0. (why)
  • No ceremony without edges. Six independent reviewers stay six independent reviewers — no waves, no gates, no graph vocabulary imposed on flat work.
  • Agent files respected. A spawn goal (name + task) that matches a user agent file's description (.agents/agents, .claude/agents, .pi/agents — project then home) loads that file — body = system prompt, frontmatter model/tools apply, file model validated against the pi model registry. File wins over inline; no match → on-demand definition.
  • Two toolsets, plus explicit override. Read-only (read, grep, find, ls, codemode — default) or write (read, grep, find, ls, bash, edit, write, codemode — write: true); tools: sets an explicit per-task allowlist. Children run with noExtensions, so codemode is the one extension they get: it lets a child batch tools.* calls instead of spending a model turn per call.
  • In-process — children are AgentSessions in the same runtime. No process spawn, no context bleed.
  • Zero parent-context injection. No catalog, no context hook. 9 slim tools total.
  • Throttled updates — widget/stream updates coalesce to ~6/s; no per-event deep clones.
  • Bounded always — a child is limited only by wall clock: explicit maxRuntimeMs, else the 1 h ceiling with /subagents auto-limit on, else the 6 h safety ceiling (default).

How it runs

Children are not subprocesses. They are separate AgentSessions inside the same pi process — which is why spawning is instant, and why a child's transcript never lands in your context:

flowchart TB
    subgraph proc["one OS process — no spawn, no IPC"]
        direction TB
        L["<b>leader</b><br/>your session, your context"]
        subgraph kids["isolated child sessions"]
            direction LR
            A["api-mapper"]
            B["db-mapper"]
        end
    end
    L -- "task text in" --> A
    L -- "task text in" --> B
    A -. "final answer only" .-> L
    B -. "final answer only" .-> L

The dotted arrows are the whole point: a child may burn 200k tokens reading files, and the leader receives only its final answer.

Usage — the leader invents the agents

Define agents inline per call, or reference a named agent file (see Agent files). Model resolution: explicit model → agent-file model (validated against the pi model registry) → the parent's current model → settings default. Any model that is not the session model is preflight-probed; if the provider rejects it, the task falls back to the session model and the result carries a Model: note.

{
  "agent": "api-reviewer",
  "prompt": "You are a strict API reviewer. Check auth, rate limiting, and error handling. Cite file:line.",
  "task": "Review src/api/upload.ts"
}

Parallel — mixed toolsets, siblings can talk via mailbox (intercom is always on):

{
  "tasks": [
    { "agent": "researcher", "prompt": "You find facts. Cite paths.", "task": "Map the auth flow", "write": false },
    { "agent": "implementer", "prompt": "You make minimal changes.", "task": "Implement POST /api/upload", "write": true }
  ]
}

Chain — {previous} is replaced with the prior agent's output:

{
  "chain": [
    { "agent": "planner", "prompt": "You write a step list.", "task": "Plan the change", "write": false },
    { "agent": "doer", "prompt": "You follow the plan exactly.", "task": "Execute: {previous}", "write": true }
  ]
}

Agent files

A user agent file in an agents directory is matched by its description frontmatter against the spawn goal (agent name + task) — not by name. When matched, the file is authoritative: body = system prompt, frontmatter model/tools apply, inline prompt/model are ignored — with one exception: explicit per-call tools/write override the file's tools (the file narrows defaults, it never displaces explicit intent, and it can never widen past the leader's read/write choice). An override is surfaced on the task's notice and summary. No match → the inline on-demand definition stands. The model stays in control: it names the agent and states the goal; user files that describe that goal take over.

---
name: api-reviewer
description: reviews APIs for auth, rate limiting, and error handling
model: claude-opus-4-6
tools: read, grep, find, ls
---
You are a strict API reviewer. Check auth, rate limiting, and error handling. Cite file:line.

Lookup order (first directory with a match wins):

  1. .agents/agents/ then .claude/agents/ then .pi/agents/ in each directory from the task cwd up to the filesystem root (nearest ancestor wins).
  2. Home: ~/.agents/agents/ (single source) → ~/.claude/agents/ → ~/.pi/agents/.

Within a directory the file with the highest description-overlap score wins (≥2 shared meaningful tokens and ≥40% coverage of the shorter token set). A file model is validated against the pi model registry (unknown model fails the task with a catalog message). File bodies are capped at 64,000 characters. Files without a description frontmatter never match.

Worktree isolation (write agents)

In a git repo, a write: true subagent runs in an isolated git worktree at <repo>/.git/subagents/<run>/<task> on branch subagents/<run>/<task> — the child's cwd is the worktree, so project context (AGENTS.md chain) still loads, node_modules is symlinked, and the main tree stays clean while the child works. Parallel write agents can't collide on files.

On completion the extension commits the child's changes (the child is told not to touch branches) and reports branch + diffstat + changed files in the result. The leader reviews, then merges:

git merge --no-ff subagents/<run>/<task>

Branch relationships — read this before merging. A write task that needs a completed write task is stacked: its worktree branches from the upstream's branch, so the child actually sees the files its upstream wrote. A stacked branch contains its upstream's commits, so merge order doesn't matter — merging the stacked branch brings both, and the upstream's own merge is then a no-op.

Why this matters (verified against real git, not just reasoned about): when a downstream child can't see its upstream's work, merging produces spurious conflicts, half-clobbered files where the child's side looks like a phantom delete, and — worst — clean merges that leave a broken tree. If the upstream renamed login→signIn and the downstream wrote new code importing login, git reports success with exit 0 and the code doesn't compile. Stacking removes that class by construction.

Same-wave write tasks are siblings: both branch from the same base, so they are independent. When two siblings changed the same file the summary emits CONFLICT RISK naming the overlap — the second git merge is a real 3-way and fails loudly, which is the safe outcome. The dangerous case is siblings touching different but coupled files: that merges clean and breaks at runtime. Nothing can detect it for you.

Dependencies are shared, not isolated. node_modules is symlinked to the main checkout, so dependency writes escape the worktree: children are instructed never to install, upgrade, or delete deps. A task that genuinely needs a dependency change should edit the manifest and say so.

Isolation follows the toolset the child actually receives: explicit tools: ["bash", "edit", "write"] earns a worktree even without write: true, and an agent file that narrows the child to read-only gets no branch at all.

Cleanup, in order of trust:

When What
Session start Registered worktrees can't be live yet — an interrupted child's uncommitted work is committed, the branch kept, the dir dropped. Then merged branches are reaped and dirs git no longer tracks are removed.
After a run Merged branches + their dirs, never touching a branch a live run owns (checkout state re-checked per branch).
Task failed/canceled Partial work is committed first, so the branch keeps it; the dir is dropped.

Non-git repos fall back to in-place edits.

Graph mode — needs

parallel runs everything at once; chain runs everything one at a time. Most real work is neither. Give a task an id and list the ids it needs:

{
  "tasks": [
    { "id": "api", "agent": "api-mapper", "task": "Map every route in src/api/" },
    { "id": "db",  "agent": "db-mapper",  "task": "Map the schema in src/db/" },
    { "id": "doc", "agent": "writer", "needs": ["api", "db"], "write": true,
      "task": "Write ARCHITECTURE.md from the maps above. Verify: test -s ARCHITECTURE.md" }
  ]
}

The call line renders the graph in §2 notation as the model types it:

subagent graph 3
  wave1[api ∥ db] → gate → wave2[doc]
  api api-mapper Map every route in src/api/
  db  db-mapper Map the schema in src/db/
  doc writer ✎ ← api, db Write ARCHITECTURE.md from the maps above. Verify: test -s ARCHI…

✎ marks a write-toolset task; ← lists its edges. With no needs anywhere the wave line is omitted entirely.

What one edge does

An edge is not just ordering. It is a delivery:

sequenceDiagram
    participant S as scheduler
    participant A as api
    participant D as db
    participant W as doc

    Note over S,D: wave 1 — both start together
    S->>A: "Map every route in src/api/"
    S->>D: "Map the schema in src/db/"
    A-->>S: route map
    Note right of W: doc is queued,<br/>waiting at the gate
    D-->>S: schema map
    Note over S: gate opens: every need settled
    S->>W: ## Output of api<br/>&lt;route map&gt;<br/><br/>## Output of db<br/>&lt;schema map&gt;<br/>---<br/>"Write ARCHITECTURE.md…"

The leader never copies those outputs into the prompt — so it cannot forget to.

One scheduler, four shapes

Single, parallel, chain and graph are not four code paths. They are four shapes of the same wave loop:

flowchart LR
    subgraph one["single"]
        direction TB
        s1(("a"))
    end
    subgraph par["parallel — no needs"]
        direction TB
        p1(("a")) ~~~ p2(("b")) ~~~ p3(("c"))
    end
    subgraph ch["chain — needs: [previous]"]
        direction TB
        c1(("a")) --> c2(("b")) --> c3(("c"))
    end
    subgraph gr["graph — needs"]
        direction TB
        g1(("a")) --> g2(("b"))
        g1 --> g3(("c"))
        g2 --> g4(("d"))
        g3 --> g4
    end

The loop

flowchart TD
    start(["subagent call"]) --> validate{"graph valid?<br/><small>unknown id · self-edge · cycle</small>"}
    validate -- no --> reject["reject the call<br/><b>zero children spawned</b>"]
    validate -- yes --> loop{"tasks left?"}
    loop -- no --> done(["run finished"])
    loop -- yes --> ready["frontier =<br/>tasks whose needs are all settled"]
    ready --> spawn["run that wave in parallel<br/><small>throttled by concurrency</small>"]
    spawn --> collect["record each output<br/>mark tasks settled"]
    collect --> loop

Two consequences worth stating plainly:

  • A bad graph costs nothing. Validation happens before the first spawn, never halfway through with three children already burning tokens.
  • A broken upstream stops its branch. If a need fails or is aborted, its dependents are marked aborted rather than run against a prompt with a hole in it:
flowchart LR
    api["api ✓"] --> doc
    db["db ✗ failed"] --> doc["doc ⏹ skipped<br/><small>never spawned</small>"]

And the rule that keeps this from becoming ceremony: zero needs anywhere = plain parallel. No waves, no gates, no graph vocabulary imposed on flat work.

Background (default) + intercom — the run returns a runId immediately; you stay steerable while it works. Children can always ask you questions and message each other:

{
  "agent": "auditor",
  "prompt": "You audit dependencies.",
  "task": "Audit package.json for outdated deps"
}

Steering a running child: while a background run is active the leader stays responsive, and you can push a message into a live child's session mid-run with steer_subagent — e.g. steer_subagent({ runId, taskId, message: "Ignore tests/, only audit runtime deps" }). The message queues as a steer if the child is mid-turn and lands at its next model boundary. Omit taskId to steer every still-running task in the run. Combined with notifyPerTask, this makes a background run feel like a live team you can redirect, not a fire-and-forget blob.

Resuming a failed child: a child that dies mid-work (provider rate limit, timeout, network error) keeps its session JSONL and its worktree branch. resume_subagent({ runId, taskId, model?: "openai/gpt-5", message? }) reopens that session with full context, re-attaches the branch, and prompts it to recap and continue — no respawn, no lost tokens. model swaps provider when the original one is exhausted. Refused for tasks that never started (no session file); those you respawn. Wait for the run to settle before resuming (the tool tells you if it hasn't).

Choosing a model: subagent_models lists what this session may use — the same set /scoped-models shows when scoping is configured, and every model with usable credentials when it is not (the output states which case applies). Each row gives the exact model value to pass, the thinking levels the runtime honors, the context window, and pi's catalog price per million tokens (free only when the reported rates are zero; absent rates read unavailable; catalog rates are a planning guide, not a billing quote). model is optional: omit it to inherit the leader's current session model, or name one (a matched agent file's model frontmatter still wins) to pin the run. ~/.pi/agent/subagent-models.json may shape the listing only — prefer sorts, hide omits, default is surfaced as a suggestion and never applied. It never grants or blocks a model: a hidden model still runs when named, because pi owns what may run.

Tools

Tool Purpose
subagent_models list the models a subagent task may name, each with the exact model value, thinking levels the runtime honors, context window, and catalog price; scoped to the session's enabled set when scoping is configured, else the full available catalogue (the output says which)
subagent single / tasks (parallel or graph via needs) / chain ({previous}); every run is background — returns a runId, completion notifies you; autoAwait:true parks the call until the run finishes and returns the final result inline; children always carry talk tools (ask/notify/mailbox); notifyPerTask (default true) wakes you as each task completes
subagent_status live per-task snapshot (non-blocking), including each child's session file path; call it once right after spawning — a child that died on spawn is invisible until far later otherwise
subagent_result full output of a run or one task
await_subagent block until a run finishes (optional timeoutMs)
reply_subagent answer a child's ask_parent question
steer_subagent inject a steering message into a running child's session (queues as steer if mid-turn; lands at its next model boundary)
resume_subagent revive a failed/aborted task in its original session (context + branch preserved); optional model swap, thinking override (the task's stored level is clamped to what the target model accepts), custom message. Waits for the reopened session to go idle before prompting, so a task killed mid-turn can still be resumed
subagent_cancel abort a running/queued run

Per-task fields

agent (name you invent — required), task (required), prompt (system prompt, optional — minimal default used), write (toolset, default read-only), plus optional model (provider/model-id; omitted → the leader's current session model, a matched agent file's frontmatter wins), thinking (validated enum: off|minimal|low|medium|high|xhigh|max), tools (explicit allowlist), cwd, maxRuntimeMs, id, needs (dependency edges — see Graph mode). Top-level only: autoAwait, notifyPerTask, concurrency (default 3, max 8). A call carries at most 16 tasks.

Child talk tools (always on)

Tool Meaning
ask_parent blocking question to the leader; delivered mid-turn as a steering message labelled [URGENT] or [not urgent], parent answers via reply_subagent
notify_parent one-way message to the leader; identical updates coalesce, final: true declares the complete report
send_agent_message message to a sibling subagent's mailbox (to = its task id, or "leader")
poll_agent_messages drain this subagent's mailbox

Intercom anti-deadlock: children are told to never block indefinitely on intercom replies — an unanswered ask_parent times out after 10 minutes (the child is told to proceed with best judgment), and sibling polls are capped (~5 tries) with the same fallback. Gated siblings (later waves) may not be running yet — waiting on them is the top stall cause, so children are instructed not to.

Ask urgency: ask_parent takes urgent (default false). Both variants steer into the leader's current turn so the question is never deferred to the end of a long turn. [URGENT] tells the leader to answer before its next step; [not urgent] tells it that the child keeps waiting, so it may finish its current step first. Failures steer for the same reason. Completions, aborts, and informational updates are held in an extension outbox and delivered together as one follow-up when the leader settles; a receipt confirms against the finalized leader transcript, so a report the leader already consumed (or an awaited result or subagent_result that already showed the outcome) suppresses its redundant queued notice. See docs/notification-policy.md.

Tool exposure and opt-in codemode

Default: legacy direct calls, even when the codemode tool is active. Active helpers remain model-visible and retain native script-call compatibility. Explicit global codemode.mode: "only" still applies Pi's global script-only policy; this extension does not override it.

  • /subagents mode — show preference and effective profile.
  • /subagents mode auto — opt in to script-based routing while codemode is active; direct otherwise.
  • /subagents mode codemode — explicitly select script-based routing (direct fallback when unavailable).
  • /subagents mode direct — return to legacy direct behavior.

Choice is stored globally in ~/.pi/agent/subagents-config.json (the same file as auto-limit, written with a merge, never a clobber) and applies to every session — reload, resume, fork and tree navigation included. Sessions that chose a mode before the preference moved into that file keep their branch entry only while the config has no mode; the built-in default is still direct. Published 1.3.63 defaulted to auto: this compatibility correction requires the corrected source/package, not a retroactive change to that npm release.

Commands

  • /subagents — list runs; /subagents peek (or ctrl+shift+a) — browsable pane
  • /subagents auto-limit on|off — toggle leader-imposed maxRuntimeMs caps (persists to ~/.pi/agent/subagents-config.json; default off). on gives tasks without an explicit maxRuntimeMs the 1 h default ceiling; off raises the ceiling to 6 h (still a ceiling — an unbounded child would pin the run forever). Bare /subagents auto-limit shows the current state.

Peek — /subagents peek or ctrl+shift+a

Read-only pane over the session's subagents:

  • shift+↑/shift+↓ (or j/k) — move between agents; bare arrows work too where the terminal doesn't reserve them
  • enter — live tail of that child's session file (esc goes back)
  • x then y — abort ONE subagent (only mutation; n/any other key cancels)
  • esc — close

Watching a child from outside

A child has no terminal of its own — but it does write a real transcript file, and that file is the seam every external viewer can use:

flowchart LR
    child["child session<br/><small>no TTY</small>"] -- writes --> file[("session.jsonl")]
    file -- "peek · enter" --> pane["in-pi tail"]
    file -- "tail -f" --> term["any terminal pane<br/><small>herdr · tmux · zellij</small>"]

subagent_status returns that path for every running child:

tail -f /path/from/subagent_status.jsonl

In a terminal multiplexer, that is a pane per agent — e.g. with Herdr:

herdr pane split --current --direction right
herdr pane run w1:p2 "tail -f /path/from/subagent_status.jsonl"

The extension has no multiplexer integration and does not want one: it exposes the path, your agent already knows how to drive its own terminal. For an in-pi view of the same stream, use /subagents peek.

Context budget

  • Parent tools: 9 schemas with short descriptions. No catalog, no context hook — nothing injected per request.
  • The widget shows live work only: settled runs are pruned from it and stay reachable through subagent_status / subagent_result.
  • Background completion: 3-line notice. Full text only via subagent_result.
  • Children: isolated sessions; talk tools always injected; each child's prompt states its own task id and its siblings' so mailbox addressing works. Model resolution: explicit provider/model-id or bare id via the pi model registry → the parent's current model → settings default. Thinking levels validated against the resolved model's thinkingLevelMap; subagent_models lists the levels the runtime honors when you need a safe set.

What this is built on

Graph Protocol

needs is an implementation of Graph Protocol — a delegation discipline that treats a task as a graph (Delegation<A, E, R>) rather than a checklist. Its ten sections map onto this extension as follows:

§ Protocol Here
§1 nodes, domains, edges one agent owns one task; needs are the edges
§2 happy path as execution graph, waves + gates wave scheduler: ready = tasks whose needs are settled
§3 one worker or many single vs tasks
§4 break points: wrong context, missing input, misinterpretation missing input is structurally impossible — the edge carries the output
§5 R: subgraph, method, verification command, WHY prompt guidelines require a runnable Verify: line per task
§6 structured at the boundary in: upstream outputs prepended as named blocks. out: prose (see below)
§7 observe without changing the graph the widget and /subagents peek are read-only
§8 worker attention acquired and released spawn → terminal status; aborted upstream releases dependents immediately
§9 prove it: delegated vs implemented deliberately not self-reported — see below
§10 prompt = subgraph, return = implemented graph prompt yes; return kept as prose

Why §9 is a verification command, not a self-report

The protocol asks the coordinator to compare the delegated subgraph against the graph the worker says it implemented. We implement the comparison against the filesystem and the exit code, not against the worker's account of itself, because self-reports carry close to zero signal about exactly the failure §9 exists to catch:

  • Asked to audit its own work against 34 real violations, an agent reported 0 — at 90–100 confidence. A fresh instance of the same model shown the same output caught 7 (p = 0.0156). A deterministic checker caught all 34. (Armalo Labs, 2026)
  • Across 9,876 τ2-bench and 1,879 AppWorld trajectories, "false success" reached 75.8% of self-assessing coding-agent failures; adding an LLM judge scored 0.54–0.65 AUROC (0.5 = coin flip). (arXiv:2606.09863)
  • LLM judges reading agent traces can be flipped by rewriting the trace — the exact surface a self-reported graph exposes. (arXiv:2601.14691)

The shape of the problem:

flowchart TD
    W["worker finishes"] --> Q{"who says it's correct?"}
    Q -- "the worker itself" --> S["self-report<br/><b>0 of 34 caught</b><br/><small>at 90–100 confidence</small>"]
    Q -- "another model reading the trace" --> J["LLM judge<br/><b>0.54–0.65 AUROC</b><br/><small>0.5 = coin flip</small>"]
    Q -- "the machine" --> D["exit code + git diff<br/><b>34 of 34 caught</b>"]

So §9 in practice is two things you already have:

# in the task text — the worker must prove it, not claim it
Verify: npx tsc --noEmit && bun test

# in the leader, after the run — ground truth, not narrative
git diff --stat

If files outside a worker's subgraph were touched, the diff says so. A structured return schema would only add a second, less trustworthy witness.

Why waves instead of "more agents"

Flat fan-out is not free — orchestration cost is critical path + α × cross-agent communication, and ignoring the second term is what makes added agents lose to a single one:

  • Dependency-graph partitioning vs flat file-parallel spawning across 28 real repos: +14.0% pass rate, 2.10× wall-clock, −35% API cost, with the largest gains on the most dependency-dense projects. Flat parallel inflated cost 60% for a 1.56× speedup; an agent-team baseline was fastest but scored below sequential on code quality. (arXiv:2606.00953)
  • Dynamic task graphs across 300 trials: 47.5% of baseline token cost, 79.7% accuracy vs 57.6% for a static graph — and, notably, static tied dynamic when the structure was genuinely known up front, which is the case needs targets. The "frontier" (ready set) in this scheduler is theirs. (arXiv:2605.06320)

The corollary is in the design principles: when there are no edges, don't draw a graph. Six independent reviewers stay six independent reviewers.

Development

bun install        # dev deps (typecheck/test only; runtime uses pi's bundled SDK)
npx tsc --noEmit
bun test           # pure-logic tests (wave scheduling, edge payload, mailbox, failure classification, watchdog)

Runtime state: runs persist to <parent-session>.subagents.json sidecar; restored (non-terminal → aborted) on session start.

License

MIT.