pi-matt-subagent
Blocking + background subagents for skills that require subagents (code-review, design-it-twice, improve-codebase-architecture, research, wayfinder, grilling).
Package details
Install pi-matt-subagent from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:pi-matt-subagent- Package
pi-matt-subagent- Version
0.11.0- Published
- Oct 8, 2026
- Downloads
- 2,637/mo · 1,006/wk
- Author
- hu3rror
- License
- MIT
- Types
- extension
- Size
- 451.7 KB
- Dependencies
- 0 dependencies · 5 peers
Pi manifest JSON
{
"extensions": [
"./extensions"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
简体中文: README.zh-CN.md
A pi plugin that turns the subagent instructions in Matt Pocock's skills into real tool calls. When a skill says "spawn sub-agents in parallel" or "fire the research subagents", this plugin is the execution layer — it starts real pi subagents (separate subprocesses, or an in-process session for background work), waits for them, and hands you their results.
[!WARNING] This plugin registers a tool named
subagent; similar subagents extensions do too. Don't run them side by side — install one or the other, never both.
Features
- Two tools, two semantics —
subagent(blocking) andresearch(background), matching exactly what the upstream skills ask for. - Blocking subagents — single, parallel, or chained agents run in separate subprocesses; results come back in one tool result, with a
{previous}placeholder for chain handoffs. - Background research — an in-process second session writes cited findings to a file while you keep working; the completion state (
succeeded/failed/terminated/aborted) is pushed into your context — no polling. - Six bundled roles —
standards-reviewer,spec-reviewer,design-explorer,architecture-scout,researcher,fact-finder— the same jobs the skills describe. - Four slash commands —
/code-review <ref>,/design-it-twice <candidate>,/research <question>, and/subagentsfor run management. - Per-run overrides — pin one run's model via the public
modelfield; thinking effort is deterministic (per-task > per-call override > conversationally-declared session level > per-role customization > config default > role preset > inherited main-session level, as-is), with theset-thinking-leveltool as the only model-visible channel for changing it. - Live run management — a footer counter (
⧗ N subagents running);/subagentsto follow, stop, or inspect runs.
Install
pi install npm:pi-matt-subagent
Or from a local checkout:
pi install <path-to-this-repo>
Both install the extension, which registers the three workflow presets (/code-review, /design-it-twice, /research) and the /subagents command. Verify with pi list; the presets appear in the TUI's / completion unless hidden via the hideWorkflowPresets config knob (which needs /reload to take effect).
Quick start
Everything here is model-facing — you describe the job, the main agent makes the call.
- Review your last commit —
/code-review HEAD~1runs the Standards and Spec axes as two parallel blocking subagents and reports them side by side. - Research while you keep working —
/research "verify the claim that …"returns immediately; the findings path is pushed to you when the run finishes. - Call the tools directly — say "run a
subagentreview ofsrc/lib.tswithstandards-reviewer", or "start aresearchon ADR 0013 and write findings todocs/research-0013.md".
The two tools
| Tool | Semantics | What it does |
|---|---|---|
subagent |
blocking | Runs single / parallel / chain subagents. Does not return until every subagent finishes; full results come back in one tool result. chain supports a {previous} placeholder that passes one step's output into the next. |
research |
background | Runs an in-process second session that writes cited findings to a file, returns immediately, and pushes the completion into your context — no polling. |
set-thinking-level |
session | Sets the main session's thinking level for the rest of the session (upstream pi.setThinkingLevel, session-scoped, never persisted; a fresh session starts from your global default). The only model-visible channel for changing thinking effort — call it only when the user explicitly asks for a different depth. While set, the requested level is also the declared level subagent dispatch honors above config defaults and per-role pins (issue #40). |
Both accept an optional input field: a JSON object string carrying the parameters still hidden from the public schema (thinkingOverride for subagent; maxWallClockMs for research, a wall-clock cap that may only tighten the 60-minute default). The per-run model override is a public field on both tools (ADR 0022 — a hidden channel proved unreliable: models dropped it and the run silently fell back to the configured default).
Direct fields override same-name JSON keys; invalid JSON raises a clear model-visible error, and the merged parameters are validated against the full contract before dispatch.
{
"task": "review the diff since HEAD~1 for standards compliance",
"agent": "standards-reviewer",
"model": "sensenova/sensenova-6.8-flash-lite"
}
Thinking effort is not a parameter: every run resolves it through the decision hierarchy (per-task > per-call override > conversationally-declared session level > per-role customization > config default > role preset > inherited main-session level, then the model-capability clamp), or the model applies an explicit instruction via set-thinking-level. A legacy call that still passes thinkingLevel fails loudly with an unknown-parameter error.
Model names resolve against your ~/.pi/agent/models.json registry as provider/id; an unresolvable name fails loudly and the run never starts.
The four slash commands
/code-review <ref>— two-axis review (Standards + Spec) of the diff since<ref>, run as two parallel blocking subagents./design-it-twice <candidate>— generate 3–4 radically different interface designs for one deepening candidate as parallel blocking subagents, then compare by depth, locality, and seam placement./research <question>— start a background researcher against primary sources and keep working; the completion is pushed to you with the findings path./subagents— overview and management of every run: follow progress, stop a runaway researcher, clear finished records, read a run's log (kill,tail,prune,snapshot).
Roles and dispatch
Six bundled roles come with the plugin. User agents from ~/.pi/agent/agents/ and project agents from .pi/agents/ override bundled roles by name (project agents sit behind a trust confirmation).
Thinking effort resolves deterministically through two orthogonal layers. The decision-source hierarchy is per-task/per-step thinkingLevel > per-call override (input.thinkingOverride) > declared (the conversationally-declared session level — the requested level of the last set-thinking-level call, read from the session transcript, sticky across model switches and session continuation) > roleDefaults per-role tier > config default (dispatchDefaultThinkingLevel) > role preset > inherited (the main-session level, passed through as-is — a declared depth like xhigh propagates to preset-less roles). The result then passes through the target model's capability clamp (clampThinkingLevel, upward-first) — requested vs effective are both recorded, and a clamp shows as high (req: xhigh). The public thinkingLevel parameter is gone (it was the randomness source); the only model-visible channel for changing effort is the set-thinking-level tool, called only when the user explicitly asks for a different depth. Model-pinned roles (frontmatter/embedded or roleDefaults.<role>.model) skip only the inherited layer. The dispatch model resolves per-call-first — per-call override > role-pinned model > roleDefaults model > config default > inherited session model — for both tools.
Configuration
Behavioral knobs live in a lazily-created file at ~/.pi/agent/extensions/matt-subagent.json. Loading never writes it — the file appears only when you set or reset — and deleting it restores all defaults. Values apply to the next run without /reload — except hideWorkflowPresets, which gates command registration at load and takes effect on /reload.
| Key | Default | Effect |
|---|---|---|
maxTasksPerCall |
8 | Caps tasks per call — parallel tasks and chain steps; over the cap the call is refused |
maxConcurrency |
4 | Max in-flight subagent processes in parallel mode |
perTaskOutputCap |
50 KiB | Byte cap for one task's summary output in parallel aggregation |
researchWallClockMs |
45 min | Default background-research wall-clock cap, and the hard ceiling for input.maxWallClockMs (an existing explicit value overrides the new default; reset the key to pick up 45) |
researchChildExtensions |
(curated) | npm packages loaded into a research child: default npm:@ssk_dev/pi-web-access-lean + npm:@upstash/context7-pi; only an explicit empty array disables extensions |
logTailBytes |
4096 | Byte cap for /subagents tail log reads |
dispatchDefaultModel |
(inherit) | Default provider/id when neither the call nor the role specifies one |
dispatchDefaultThinkingLevel |
(inherit) | Default thinking level above the role preset (issue #39 — live for the six embedded roles, no longer dead config); absent → inherited (main-session level as-is) |
roleDefaults |
(inherit) | Per-role dispatch customization: { "<role>": { model?, thinkingLevel? } }. roleDefaults.<role>.thinkingLevel beats the config default and the role's own preset; roleDefaults.<role>.model beats the config default model. Edit via dotted keys: config set roleDefaults.standards-reviewer.thinkingLevel low |
hideWorkflowPresets |
false |
Hides the three workflow presets (/code-review, /design-it-twice, /research) from / completion — for users with the matching skills installed, which cover the same workflows. Read at extension load, so needs /reload. Never gates the three tools or /subagents |
When hideWorkflowPresets is on, typing a hidden preset name behaves like any unknown / command: the text goes to the model unchanged.
Drive it from /subagents config:
/subagents config set maxConcurrency 6
/subagents config set dispatchDefaultThinkingLevel low
/subagents config set roleDefaults.standards-reviewer.thinkingLevel low
/subagents config set roleDefaults.researcher.model openai/gpt-x
/subagents config set hideWorkflowPresets true
/subagents config reset roleDefaults.standards-reviewer.thinkingLevel
/subagents config reset maxConcurrency
/subagents config reset
set writes the effective value to the file; reset removes one key (back to that knob's default) or the whole file. Dotted roleDefaults.<role>.<field> keys edit the nested knob; resetting a dotted key removes just that field (an emptied role prunes itself). Invalid entries degrade that key to its default and are flagged [degraded] in the config view.
Research child extensions (trust surface)
A background research child starts with only built-in tools plus the approved query packages — never the main session's full extension set. The default loadout is npm:@ssk_dev/pi-web-access-lean (web search/page fetch) and npm:@upstash/context7-pi (library docs), which the researcher role declares as web_access, query-docs, resolve-library-id.
This list is a curated trust surface, not a promise that every query-style tool is default. A package qualifies for the default list only when it is read-only, has no external write side effects, has low and explicit external cost, and serves primary-source retrieval.
Maintaining the list per-machine (no code changes):
- Extend/override — add packages to
researchChildExtensions(JSON array in~/.pi/agent/extensions/matt-subagent.json, or a comma list via/subagents config set researchChildExtensions npm:x,npm:y). Remember step 2 below: a loaded package is only visible to the researcher once its tools are also declared on the role. - Disable entirely — set
researchChildExtensionsto[]in the JSON file (the config-set CLI cannot express an empty list): the child then loads no extensions at all (network fully off, built-ins only). - Finer control — override the built-in researcher role with a user agent at
~/.pi/agent/agents/researcher.md(frontmattertools:+ body; same-name overrides win, and the runner still appends the findings path / wall-clock / checkpoint rules), to pin exactly which tools are declared.
Adding a new query tool — the two steps. (1) Load: put its package in researchChildExtensions. (2) Expose: declare the tool name on the researcher role (the built-in list, or your custom researcher.md). Missing either half is a drift.
Error shapes and remedies. When a declared tool cannot be loaded (package not installed, knob-disabled, or platform-impossible — e.g. powershell off-Windows), the run still starts but the tool is absent from the child, the research tool's returned text and the run log carry a ⚠ Declared but not loaded: … drift note, and the researcher prompt only lists tools it actually has. The symptom of ignoring the note: the researcher model guesses tool names and loops. Fix by installing the package, pointing the knob at it, or removing the tool from the role.
Reading a research run's log — what each line means. The per-run log (logPath, tail via /subagents tail <id>) mixes four kinds of content; only one of them is the run's outcome:
| Line | Source | It says | It does NOT say |
|---|---|---|---|
[loadout] ok/warn … |
loadout self-check (first lines) | whether the knob packages' tools were loaded | run success/failure |
[run] HH:MM:SS … |
stage lines (session created, prompt started, terminal: <status>) |
what the run is doing now | — |
| model text | researcher output | research progress | — |
| research-status push | terminal push card | the run's authoritative outcome (succeeded / failed / terminated / aborted) |
— |
A loadout line is a loadout status, never a run status: the researcher still runs, and the terminal: stage line (or the push card) is the outcome to judge by. Healthy signals: the push card arrives with succeeded, the findings file is written, and the log keeps growing. Trouble signals: the run stays running with no new log lines for a long stretch (the wall-clock cap then resolves it terminated), or the push card reports failed/aborted. Every executed tool call is also audited to toolcalls.jsonl next to the log — the ground truth for what actually ran.
Project layout
extensions/subagent.ts pi extension: registers the three tools, the run-registry UI,
and the in-process research child-session factory
src/lib.ts pure logic — roles, tool schemas (single source of truth),
input merge/validation, tool-name resolution, the research
runner, run-registry bookkeeping; zero pi-runtime imports
src/blocking-protocol.ts pure blocking-child protocol: JSON-lines stdout accumulation,
usage tracking, escalating-kill abort path
src/blocking-runner.ts pure blocking orchestration — single / parallel / chain plans
src/*.test.ts unit tests (node --test, no pi runtime)
scripts/ token benchmark + e2e scripts (dev-only)
prompts/ the three workflow-preset markdown sources (read by the extension at load)
docs/adr/ 19 recorded decisions (dual channel, research redesign,
config surface, usage line, …)
GLOSSARY.md domain glossary (subagent, role, blocking, background, push, …)
Development
npm test # unit tests, no pi runtime needed
npm run typecheck # extension + lib + scripts typecheck
The extension is a thin consumer of src/lib.ts; the pure functions there (dispatch-arg assembly, tool resolution, input merge/validation, the surface contract) are what the tests cover. Real-pi e2e scripts (scripts/push-e2e.ts, scripts/blocking-e2e.ts) cover the process wiring that unit tests can't reach; node scripts/benchmark-tools.ts measures the tool surface's token contribution (~482 / ~398 / ~104 tokens for subagent / research / set-thinking-level), guarded by a regression test in the suite. node scripts/apply-skill-patch.ts re-applies the ADR 0013 patch texts to the installed skills after a mattpocock upstream sync.