@hank-warren/pi-auto-permissions
Fork of @ogulcancelik/pi-auto-permissions with dialog-answer user evidence: context-aware Bash permissions for Pi with automated guardian review.
Package details
Install @hank-warren/pi-auto-permissions from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:@hank-warren/pi-auto-permissions- Package
@hank-warren/pi-auto-permissions- Version
0.20.0- Published
- Oct 7, 2026
- Downloads
- 1,085/mo · 299/wk
- Author
- hank-warren
- License
- MIT
- Types
- extension, skill
- Size
- 281.1 KB
- Dependencies
- 1 dependency · 3 peers
Pi manifest JSON
{
"skills": [
"./auto-permissions-setup"
],
"extensions": [
"./index.ts"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
pi-auto-permissions
Published fork. This is
@ogulcancelik/pi-auto-permissions0.1.3 (MIT, Can Celik) plus the dialog-answer evidence feature proposed upstream in ogulcancelik/pi-extensions#24 / PR #25. Switch back to the upstream package once it ships the feature. Never install both packages at once, or every Bash call is reviewed twice. Tests live intest/: pure module tests plustest/extension.test.ts, which drives thetool_callpipeline through a mockExtensionAPI.
A context-aware permission system for Pi shell commands, with an automated guardian that checks what the user actually authorized.
Pi Auto Permissions pauses configured Bash commands before execution. A guardian model reviews the exact command against a compact view of the current conversation:
- Clearly authorized and compliant commands run automatically.
- Authorized commands that violate the user's constraints are blocked with feedback so the agent can revise them.
- Commands without clear authorization are sent to the user for confirmation.
- High-risk commands always require confirmation.
Only user messages can grant permission or impose constraints. The assistant cannot authorize its own command.
The posture is review-by-default: a fresh install ships a live ruleset (deny rules for oversight bypasses and critical-path destruction, guardian review for force pushes, infrastructure destroys, credential access, and the rest of the default groups), prose trust configuration in guardianPolicy, a session-start trust snapshot, an optional blanket reviewAllShell mode, a denial ledger with allow-on-retry, and a bundled setup skill that co-authors the trust config with you from observed friction.
Compared with Claude Code auto mode
The design deliberately mirrors Claude Code's auto mode (default ruleset content, trust tiers, "$defaults" splice, session-start remote baseline, denial review) while keeping the pieces pi does better. Honestly stated, both directions:
| pi-auto-permissions | Claude Code auto mode | |
|---|---|---|
| Verdicts and reasons | approve/revise/ask_user; always a concrete one-sentence reason; revise tells the agent exactly how to clear the objection |
Deny only; fixed Blocked by classifier text in most sessions |
| Evidence caching | Append-only reviewer lineage: full envelope once, deltas after, fingerprint-invalidated | Stage-1→stage-2 prompt cache; no cross-review evidence cache |
| Deterministic surface | Regex gates are testable; the false-positive surface is bounded by the patterns; the guardian runs only after a mechanical match | Prose rules are checkable only by running the classifier |
| Evaluation loop | Labelled evaluation JSONL (three-way user labels) + per-review cost sidecar | /feedback; no per-review cost or label surface |
| Coverage | Bash only. Edits, writes, fetches, MCP tools, and subagent spawns are ungated (the beyond-bash plan is the named follow-up) | All tool calls Tier 3-reviewed; in-project edits skip review by design |
| Sandbox | None — pair with OS sandboxing for a hard boundary | Integrated sandbox + network-domain review |
| Injection posture | Keeps assistant text and tool-call summaries as (non-authorizing) evidence — better provenance, larger injection surface | Strips assistant text and tool output entirely; input-layer injection probe |
Example
Suppose Git commits are globally guarded.
If the user says:
Fix the failing test and commit it with a concise lowercase message.
Then a matching commit can run without another permission prompt:
git commit -m "fix retry test"
But the guardian rejects a command that violates the request:
git commit -m "Fix the Failing Retry Test and Update Documentation"
If the conversation never authorized a commit, an interactive Pi session asks the user before running it. A non-interactive session blocks it.
Permissions are contextual by default. Each command is judged against the current conversation and the exact action being proposed.
Install
pi install npm:@hank-warren/pi-auto-permissions
For local development:
pi install /absolute/path/to/pi-extensions/packages/pi-auto-permissions
Remove or disable other extensions that gate the same commands to avoid duplicate prompts.
Define your policy
The extension ships with a default policy, derived from the mechanically matchable rows of Claude Code auto mode's blocked-by-default list. Out of the box — no config file at all — it denies agent-oversight bypasses and critical-path destruction outright, and sends force pushes, history rewrites, infrastructure destroys, piped remote code execution, credential access, registry overrides, PR merges/self-approvals, and shell writes to protected dotfiles to the guardian. A normal dev session (build, test, commit, push a feature branch) triggers zero gates.
The policy is a list of rules. When your config has no rules key at all, the built-in ruleset is active. An authored rules array fully replaces the built-ins — unless it splices them back in with the literal string "$defaults", which expands to the built-in ruleset at that position so your entries can sit before or after them. An explicit "rules": [] gates nothing (the pre-0.13 out-of-box behavior, now an explicit opt-out). Commands that do not match a rule run normally.
The default groups — usable in .pi/trusted-ops (guarded rules only) — are oversight, destructive, git, iac, net-exec, secrets, supply-chain, review, and protected-paths. See default-rules.ts for the exact patterns and levels.
Configuration is read before every Bash command from:
$PI_CODING_AGENT_DIR/pi-auto-permissions/config.json
PI_CODING_AGENT_DIR defaults to ~/.pi/agent.
Here is a policy that keeps the defaults and additionally gates commits, pushes, and publishing:
{
"rules": [
"$defaults",
{
"pattern": "\\bgit\\s+commit\\b",
"flags": "i",
"level": "guarded",
"group": "git",
"label": "Git commit"
},
{
"pattern": "\\bgit\\s+push\\b",
"flags": "i",
"level": "guarded",
"group": "git",
"label": "Git push"
},
{
"pattern": "\\bnpm\\s+publish\\b",
"flags": "i",
"level": "guarded",
"group": "npm",
"label": "npm publish"
}
]
}
Rule fields are:
pattern: JavaScript regular expression sourceflags: optional regular expression flags, defaulting toilevel:guardedfor guardian review, ordenyfor a hard blockgroup: policy group used by trusted-project bypasseslabel: short description shown during reviewmessage: optional feedback; required for deny rules
A deny rule is a hard policy boundary: it blocks immediately with its message, is evaluated before every other level, and nothing lifts it — not a trusted group, not user approval at a prompt. Reserve it for operations that should never happen in an agent session (disabling agent oversight, deleting critical paths), and use guarded where a human judgment call is legitimate.
When several rules match one command, the most severe level wins (deny, then guarded) regardless of their order in the list.
A legacy "level": "convention" rule still loads, as a deny rule: it blocks without review, and .pi/trusted-ops cannot lift it.
Set enabled to false to disable the extension. Invalid configuration fails closed and blocks Bash calls until corrected.
Review every shell command
Set "reviewAllShell": true to review every bash command that matches no rule under a generic shell command gate (group all-shell), the analogue of Claude Code's classifyAllShell. The rules then act as a severity layer on top of blanket coverage: deny rules still block outright, and everything else — named by a rule or not — goes to the guardian. The trade is one guardian call per command; pair it with a cheap reviewer model. A command whose matching group is listed in .pi/trusted-ops was explicitly waved through and is not re-captured by the blanket gate; trusting all-shell itself opts a project out of blanket review while keeping the ruleset live. Lineage caching and the evaluation log apply to all-shell reviews unchanged.
Define trusted infrastructure
The rules decide which commands are reviewed; guardianPolicy tells the guardian how to judge them against your infrastructure. All four lists are prose — write entries the way you would describe your environment to a new engineer:
{
"guardianPolicy": {
"environment": [
"Our GitHub orgs acme-corp and example-labs are trusted source control",
"registry.internal.example is our internal package registry"
],
"allow": [
"Uploading build artifacts to the artifacts.example.com bucket is routine"
],
"softDeny": [
"Do not run database migrations against staging without being asked"
],
"hardDeny": [
"Never push code or data to a repository outside our GitHub orgs"
]
}
}
The lists extend the built-in decision table (which stays active) and are appended to the reviewer prompt as a labeled section with explicit precedence:
- hardDeny blocks unconditionally — user intent, permission overrides, and allow entries never clear it.
- softDeny blocks unless an allow entry covers the action or the user's own message names the exact action and target ("force-push this branch", not "clean up the repo").
- allow entries are exceptions to softDeny and to the trust boundary: data flow only, never destructive or credential operations on the same infrastructure.
- environment entries define what "internal" means — a destination nothing names is treated as a potential exfiltration target.
Each list is independent: setting environment alone leaves the others intact. Worth filling in, slot by slot: your source control orgs, trusted internal domains, cloud buckets, package registries and mirrors, key internal services — and on the sensitive side, production namespaces, sensitive data locations, and remote targets that should always get a hard look. These sections are appended outside systemPrompt, so a customized systemPromptFile still receives them, and every entry participates in the reviewer fingerprint, so editing the policy safely invalidates the cached reviewer lineage.
Alongside your configured entries, the guardian always receives a SESSION ENVIRONMENT section captured once at session start: the working directory, repository root, and the git remote -v baseline. Those remotes are within the session's trust boundary; a remote added or repointed mid-session is not — which pairs with the default git remote add|set-url gate rule, giving the guardian a baseline to judge those commands against. In a non-repository directory the section says exactly that, so no repository trust is assumed.
The precedence tiers are prompt-enforced, like Claude Code's own; the mechanically enforced hard boundary is the deny rule level.
/auto-permissions setup
The subcommand is a pointer: it hands the job straight to the bundled auto-permissions-setup skill, starting that conversation on the keystroke. The skill starts with guardian friction logs, prefers session-recall tools over raw session scans, interviews you about ambiguous hosts and absolute boundaries, and applies confirmed edits with a dated backup, diff, and loader validation.
There used to be a second path here — a one-shot wizard that scanned the project and your ten most recent sessions and rendered an accept-or-discard draft. It is gone. A conversation can ask which of two look-alike hosts is production, weigh friction across your whole history rather than a fixed window, and be argued out of a bad suggestion; a fixed draft could do none of that, so the two were never equal and keeping both only split the maintenance. What survives is the property that mattered: guardianPolicy is read only from the user-scoped config file — never from a project file — so a checked-in repository cannot inject its own trust entries.
Guardian configuration
By default, the guardian uses Pi's active model with low reasoning effort and a 30-second timeout. You can select a separate low-cost model.
Giving the guardian its own account keeps reviews from competing with your interactive session for a subscription's rate limits. Pi keys OAuth credentials by provider id, so a second login needs a second provider id — which is what @hank-warren/pi-multi-login exists to create. This package no longer registers one itself.
pi install npm:@hank-warren/pi-multi-login- Run
/multi-loginand add a login with baseopenai-codexand suffixauto-permissions; then run/login, pick the new provider, and complete OAuth with the account reserved for guardian reviews. - Select the dedicated provider in the Auto Permissions config:
{
"reviewer": {
"provider": "openai-codex-auto-permissions",
"model": "gpt-5.6-luna",
"reasoningEffort": "low",
"timeoutMs": 30000
}
}
The separate credential is stored in Pi's normal auth.json under openai-codex-auto-permissions; the existing openai-codex credential is unchanged. The alias uses Pi's built-in OpenAI Codex OAuth flow, model catalog, and transport. Because it is a normal provider, its models also appear in Pi's model selector and in /login after login.
If you already used the login this package registered in earlier versions, nothing changes: pi-multi-login adopts the existing openai-codex-auto-permissions credential on first run, so the config above keeps resolving. Without that package installed, the provider is missing and Auto Permissions warns once per session; point reviewer.provider at openai-codex (or any signed-in provider) to silence it.
Other reviewer providers continue to work by setting their normal provider and model ids.
/auto-permissions
The reviewer settings that change most often are editable from a settings menu instead of the config file:
| Row | What it edits |
|---|---|
| Enabled | enabled — off disables all gating, including deny rules; every command runs unreviewed |
| Reviewer model | reviewer.provider / reviewer.model, picked from the models you are signed in to |
| Thinking level | reviewer.reasoningEffort |
| Review timeout | reviewer.timeoutMs, entered as 30s or 45000ms |
| Parallel reviews | reviewConcurrency, cycling through 1, 2, 4, 8 and 16; a hand-written value in between stays in the cycle |
| System prompt | read-only: the resolved path of the active systemPromptFile, or whether the built-in or an inline prompt is in use |
| Recent denials | read-only view of the denial log with Allow on retry; shown while denialLog is enabled (the default) |
Saves are applied immediately — the config is re-read on every guarded command, so there is nothing to restart, in this session or any other. The menu is a narrow writer: it merges only the keys above into whatever is on disk, so rules, prompts, evidence settings and log paths stay exactly as you wrote them and remain file-only. A config that fails validation is never rewritten; the menu reports the error and refuses to open.
Guardian prompt
The bundled prompt evaluates authorization, command risk, and compliance with user constraints. Replace it inline with systemPrompt, or load a file:
{
"systemPromptFile": "./guardian-prompt.md"
}
Relative paths resolve from the configuration directory. Set only one of systemPrompt or systemPromptFile.
The guardian must return one of three decisions:
approve: execute the commandrevise: block it and tell the main agent what to correctask_user: open an approval prompt (rendered with@hank-warren/pi-permission-selector'sOptionSelector): the command, then the guardian's note behind a▌bar in a fixed strong yellow (gold on dark themes, dark goldenrod on light ones, regardless of the theme'swarningcolor), then numbered options. A command taller than a quarter of the terminal is cut to a preview, and the prompt stays pinned while you scroll the session above. Ctrl+O opens the full command in a scrolling viewer inside the prompt (Shift+↑/↓ or the mouse wheel; PgUp/PgDn outside fullscreen) and closes it again — the only complete copy when a codemode script built the command, since the transcript then shows the script. Numbered options with1–9hotkeys, Tab to attach a note that is delivered to the agent as a steering user message, Esc to cancel — which blocks the command. In RPC mode, as in a pi-agents child whose prompts the parent shows, the prompt is Pi's plain select dialog instead: Allow or Block, without Tab notes. It waits for the RPC client's answer, so a client that never answers holds the command until the turn is cancelled
Conversation context and caching
The guardian receives a compact chronological view of Pi's active, compaction-aware conversation context plus the exact pending command. It includes retained user and assistant text, Pi's latest compaction summary, and small summaries of finalized tool calls, but excludes summarized-away history, thinking, tool output bodies, file contents, patches, images, and session metadata. Compaction summaries are non-authoritative assistant context and cannot grant permission.
The extension also recognizes native checkpoints created by @ogulcancelik/pi-codex-compaction. It keeps the checkpoint's retained plaintext user messages and post-checkpoint evidence, while excluding the opaque provider state and older local history. This integration is optional and does not change behavior when pi-codex-compaction is not installed.
The first review sends the complete compact evidence. Later reviews reuse the same reviewer session and append only newly finalized evidence and the latest action. The extension uses stable session identity, cache affinity, and long cache retention when supported by the provider. Branch changes, model or policy changes, failures, cancellation, and context pressure reset the reviewer session.
Assistant and tool evidence provide context but never grant permission. Later user messages override earlier conflicting user instructions.
Codemode scripts
A codemode script runs shell commands from inside it, so judging one of its bash calls alone misses what the script is doing around it: a loop over forty targets, or a delete that only makes sense because a backup runs first. The guardian therefore sees the script itself:
- The running script. A call a script makes is reviewed with that script's source as an assistant-source
SCRIPT codemode [<call id>]record, and the proposed action names it inissuedByScript. It sits in the evidence stream rather than in the proposed action, so every call from one script shares it through the cached reviewer lineage instead of re-sending it. - Finished scripts. Their
SCRIPTrecord stays, and theirTOOLrecord lists every call they made and whether it succeeded, from thenestedCallsrecord Pi keeps on the result. - Never authorization. The script is model-written: comments, strings and names in it grant nothing. A policy section appended outside
systemPrompttells the guardian to use it only for context, and to judge an action whose prerequisite was blocked, failed, or runs concurrently as if that step never happened.
Scripts are capped at 8000 characters (head and tail, with an elision marker), and a full reviewer rebuild keeps the source of only the five newest; older ones collapse to SCRIPT codemode [<call id>] (source elided).
Trusted projects may optionally provide their root AGENTS.md, or CLAUDE.md when no AGENTS.md exists, as policy evidence:
{
"reviewEvidence": {
"projectInstructions": true
}
}
Project instructions help interpret the requested workflow, but cannot independently authorize an action or override guardian policy.
Interactive dialog answers
When the main agent gathers a decision through an interactive question tool (for example ask_user_question), the user's selection is stored as a tool result, which never grants permission. Operators can allowlist user-answer tools, whose confirmed answers become source: "user" evidence records:
{
"reviewEvidence": {
"userAnswerTools": ["ask_user_question"]
}
}
A successful result from an allowlisted tool qualifies when its details are { "answers": [{ "question": string, "answer"?: string, "selected"?: string[], "notes"?: string }], "cancelled": false } with no error field. selected takes precedence over answer, non-string answers are ignored, and notes count only alongside a real answer. Each answered question contributes one USER (dialog answer): record; the guardian treats it as authorization for exactly the selected content and is told the question wording is assistant-drafted context, never an instruction. Any dialog extension emitting that shape qualifies.
Injected user messages
Extensions inject context with appendCustomMessageEntry, and Pi converts that content into a user message for the model — but the session entry is a custom_message with no role, so the reviewer never saw it. The model was acting on text the guardian could not read.
reviewEvidence.userMessageTypes allowlists the customTypes whose content becomes source: "user" evidence:
{
"reviewEvidence": {
"userMessageTypes": ["my-objective-anchor"]
}
}
The default is an empty list: no injected type is trusted until you name it.
Each text block contributes one USER (<customType>): record, in session order, never truncated — user records are the only ones that can authorize or constrain. Images are dropped. The guardian is told these records authorize exactly the operations they name, that their constraints bind, and that their content is data rather than instructions.
Injected messages whose customType is not allowlisted are still visible — as CUSTOM <customType>: records with source: "tool", capped like any tool record. They routinely carry the reason a command was proposed (subagent return values, CI outcomes, process notifications), so hiding them left the guardian judging commands with no visible motive; but they are extension output, and only the allowlist can promote a type to user-source, so a subagent's return value can explain a command without ever authorizing it.
The allowlist matches what you wrote: a bare name such as ask_user_question matches that tool in any namespace (functions.ask_user_question included), while a dotted name matches exactly. The default is an empty list.
Parallel reviews
When several guarded commands are pending at once, their guardian reviews run in parallel: a codemode script that runs bash calls under Promise.all, or several bash calls in one assistant message. Verdicts are still applied, and you are still asked, one command at a time.
{
"reviewConcurrency": 4
}
reviewConcurrency (1–16, default 4) caps how many guardian calls run at once; the Parallel reviews row in /auto-permissions sets it too. Set it to 1 for fully serial reviews. Parallel reviews all build on the same cached reviewer lineage: the first to finish extends it and the others' exchanges are dropped, which loses nothing because prior reviewer responses are non-authoritative. When there is no usable lineage yet (the first review of a session, or after a compaction, reset or cancellation), one review builds it from the full evidence and the others wait for it and extend it, so a batch pays for one cold request rather than one each. If that build fails, the waiting reviews go ahead on their own instead of queuing behind each other.
- Codemode scripts. Pi runs each
bashcall a script makes through the sametool_callgate as a model-issued call, concurrently, so their reviews simply overlap. When a script ends — it returns, throws, times out or is aborted — while some of its calls are still queued, under review or waiting on your approval, those are released at once: the guardian request is cancelled (the reviewer lineage is kept) and an open prompt closes by itself. Pi never runs such a call anyway, because it hands it an already-aborted signal, so this only saves you from answering a prompt that can change nothing. - Several calls in one message. Pi runs the
tool_callhandlers for one message's calls one after another, so the handler of the firstbashcall starts the reviews of the later guarded ones. Each call later takes its early review only when the input Pi finally hands the handler is exactly the input that was reviewed; otherwise it is reviewed afresh. An early review whose call never arrives is dropped when the turn ends. WithreviewConcurrency: 1nothing is reviewed early. - Your answers still bind. Applying a verdict and prompting you share one slot, and a verdict whose evidence changed while it waited — because you answered another command's prompt, or a tool result landed — is reviewed again before it is applied (up to three reviews per command; the third runs holding that slot, so no other prompt can be answered while it runs). A block you give on one command therefore reaches a sibling that was reviewed while your prompt was open.
Denial log and retry
Every non-approved outcome — a guardian revise, a user block at the prompt, a deny block, a review-infrastructure failure — is appended to a private denials.jsonl sidecar next to the config (0600, 16 MB rotation, one previous generation kept):
{"v":1,"ts":"…","sessionId":"…","tool":"bash","gate":{"label":"Force push","group":"git"},"command":"git push --force …","verdict":"block","reason":"…","decisionSource":"user"}
On by default; "denialLog": { "enabled": false } opts out, and path relocates it.
/auto-permissions gains a Recent denials view over that ledger. Selecting a denial offers Allow on retry: an exact-command session override is added through the existing override machinery (it re-injects as user-source evidence on every later review and survives session resume), and an injected message tells the agent it may run that exact command again. No new authorization pathway exists — it is the same override a live approval prompt would have produced.
Every denial also emits a pi.events event — auto-permissions:denied with tool, command, gate, group, verdict, reason, decisionSource — the pi-native equivalent of Claude Code's PermissionDenied hook, for other extensions to react to. Like that hook, it cannot reverse a denial.
Permission overrides also persist as custom session entries, so a resumed session keeps the user's earlier allow decisions and standing block constraints instead of forgetting them.
Prompted-review evaluation log
Auto Permissions can append a private JSONL regression record whenever the guardian asks for confirmation and the user gives explicit feedback:
{
"evaluationLog": {
"enabled": true,
"path": "./review-evals.jsonl"
}
}
The path defaults to review-evals.jsonl beside the Auto Permissions config and resolves relative to that config. The file is created with mode 0600 and rotates to review-evals.jsonl.1 once it passes 64 MiB, keeping one previous generation, so older rows are discarded. Writing is best effort: a failing log never blocks or changes a permission decision.
When logging is enabled, prompted reviews offer the three labeling choices:
- Allow — asking was unnecessary executes the command and records
userChoice: "allow_unnecessary"withexpectedDecision: "approve". - Block — asking was appropriate remains the second choice, blocks the command, and records
userChoice: "block"withexpectedDecision: "ask_user". (A block always affirms the prompt: the guardian has no reject verdict — its only non-approve outcomes are asking you or bouncing the command back to the agent asrevise— so the only true rejection in the system is yours at this prompt.) - Allow — asking was appropriate executes the command and records
userChoice: "allow_appropriate"withexpectedDecision: "ask_user".
When logging is disabled, the prompt retains the normal Allow and Block choices. Prompts from Pi or other extensions are unchanged.
Each version 2 record contains the collected user request, exact command, compact reviewer evidence, guardian reason, gate and session metadata, raw user choice, and both labels used for evaluation. The guardian's actualDecision is ask_user. Automatic-review failures are identified separately with decisionSource: "review_failure". Existing version 1 records can remain in the same JSONL file.
Logging is disabled by default. Records can contain sensitive conversation text and shell commands, so keep the file private and out of repositories. Cancelled or interrupted prompts are not labeled or logged.
Reviewer usage sidecar
Guardian reviews are model calls made outside the Pi agent loop, so they never appear in the session transcript and are invisible to tools that total usage from session files. Auto Permissions therefore appends one content-free record per completed review:
{"v":1,"id":"3f2b…","ts":"2026-08-11T22:41:03.118Z","source":"auto-permissions","label":"guardian","provider":"anthropic","model":"claude-fable-5","usage":{"input":812,"output":96,"cacheRead":18442,"cacheWrite":0,"reasoning":48,"cost":0.0121}}
The record carries identity, timing, and counters only. It never contains prompts, commands, reviewer evidence, verdicts, or responses, which is what makes it safe to keep on by default:
{
"usageLog": {
"enabled": true,
"path": "./usage.jsonl"
}
}
The path defaults to usage.jsonl beside the Auto Permissions config and resolves relative to it. The file is created with mode 0600 and rotates to usage.jsonl.1 once it passes 16 MB, keeping one previous generation. Writing is best effort: a failing sidecar never blocks or changes a permission decision.
@hank-warren/pi-stats reads <agent dir>/<extension>/usage.jsonl sidecars and shows this usage as its own provider/model (guardian) row. Set "enabled": false to stop recording.
Review display
The default UI shows guardian progress as a single animated status line in a temporary widget above the editor:
auto permissions · Git commit · ✶ waiting for openai-codex-auto-permissions/gpt-5.6-luna
A sparkle spinner (✶ ✸ ✻ ✽) cycles while the guardian is reviewing and resolves to ✓ approved, ↻ revision requested, or ✗ blocked. When approval is needed, the widget row steps aside for the approval dialog, whose static yellow ● heading keeps it visually distinct from the transcript; with several commands under review, the summary line still counts the ones waiting for your approval. The reason, when present, appears on a second line behind a ▌ bar in the outcome's color; the command itself is not repeated because it is already visible in the Bash tool box. Configure the widget with:
{
"ui": {
"enabled": true,
"resultDisplayMs": 2500
}
}
While several commands are under review at once, the widget collapses them into one summary line:
auto permissions · 3 commands · ✶ 2 waiting for openai-codex-auto-permissions/gpt-5.6-luna · ⋯ 1 queued
queued counts commands waiting for a review slot, or with a verdict waiting to be applied behind another command's prompt.
Set ui.enabled to false to hide review state without disabling enforcement.
Herdr pane indicator
When the environment variable HERDR_ENV is set to exactly 1, the extension additionally emits a herdr:blocked event whenever a command is waiting on your approval, and a matching cleared event once the review resolves. Herdr, a terminal multiplexer for coding agents, sets this variable for the panes it manages and uses the event to flag the pane that needs attention — useful when a review is blocking in a pane you are not currently looking at. The blocked event carries the gate label so the indicator can name the operation.
This is an optional integration and nothing needs to be configured to use it. Outside Herdr the variable is unset, the emit is skipped entirely, and every other feature behaves identically; the extension has no dependency on Herdr being installed.
Trusted groups
In a trusted project, create .pi/trusted-ops to bypass selected rule groups:
git
gh
Group names come from your configured rules. A trusted group bypasses guarded review for that group, so use it only in projects you control. Deny rules are never bypassed: .pi/trusted-ops is a project-scoped file, and a checked-in file must not be able to disarm a hard policy boundary.
Subagent sessions
When a session is a subagent child (PI_SUBAGENT_CHILD=1, set by @hank-warren/pi-agents and pi-subagents), the guardian receives additional execution facts with each review — run id, nesting depth, whether the cwd is a linked git worktree, and the checked-out branch — plus a prompt section telling it to judge risk by effect scope and reversibility relative to the subagent's own workspace instead of by command name. Mutations confined to the subagent's isolated worktree, its own feature branch, or resources it created are approvable when they serve the delegated task; ask_user is reserved for effects that escape that scope (shared or default branches, host-level configuration, production systems, credentials, data leaving the machine).
The child itself is told it is a subagent: a short section appended to its system prompt says that commands needing approval pause the supervising session, so it should prefer in-scope and read-only commands and revise when asked to. On Pi 1.x the section is a named prompt section (auto_permissions_subagent), so it never replaces the prompt that Pi and other extensions build, such as the parent's instruction files pi-agents adds to a child.
An ask_user verdict in a subagent, or a failed review (timeout, unavailable reviewer, unparsable verdict), never interrupts the human on the first try:
- Child with a UI (an RPC child whose prompts surface in the parent, as with pi-agents): the command is blocked with a revise-first reason. Issuing the same command again unchanged in a later turn escalates it to the human as an ordinary approval prompt, once: after that, a later identical attempt is revised first again. Duplicates within the same turn (sibling calls, or a codemode
Promise.all) are refused as well, because the model has not seen the first refusal yet. Any other command gets its own revise-first turn. - Child without a UI: the command is blocked with a reason instructing the child to route around the gated operation or report the blocker.
The guardian's subagent section says which of these applies, so a child without a UI is never told a human may approve later. A revise-first refusal is recorded in the denial log and the auto-permissions:denied event with "reviseFirst": true, keeping the decisionSource of the verdict behind it, so it stays distinguishable from a guardian's own revise.
Reviewer usage records from subagent sessions carry "subagent": true in the usage sidecar.
Guardian dispatch
Reviewer requests dispatch through the host's model runtime rather than pi-ai's compat layer, so provider transports registered by other extensions (for example @gotgenes/pi-anthropic-auth OAuth request shaping) apply to guardian calls. When the runtime seam is unavailable, dispatch falls back to compat.completeSimple.
Failure behavior
A missing reviewer model, unavailable credentials, malformed response, timeout, cancellation, or oversized review context never auto-approves a command.
- Interactive sessions fall back to user confirmation.
- Non-interactive sessions block the command.
- Invalid configuration blocks Bash calls until corrected.
Security boundary
Rules match raw shell text. Quoting, variables, aliases, generated scripts, or other indirection can evade a regex, while quoted command text can cause false positives.
Pi Auto Permissions is a permission layer for normal agent behavior. It is not an operating-system sandbox or a defense against hostile shell input. Pair it with sandboxing when commands need a hard security boundary.
reviewEvidence.userAnswerTools widens what counts as user authorization: any code that can record a tool result under an allowlisted tool name can mint USER (dialog answer) evidence. Allowlist only tool names served by extensions you trust.
reviewEvidence.userMessageTypes widens it the same way, and is why it is an allowlist rather than "project every injected message": any installed extension can append a custom message under any customType, so a blanket rule would let any of them mint user authorization. Allowlist only types written by extensions you trust.
Changelog
See CHANGELOG.md for release history.
License
MIT