pi-warden
Makes the Pi agent follow your project's rules. Jev judges every write against your pi-warden.md and quotes the broken rule back to the agent, names slop, breaks stuck loops, calls out unverified done claims, compresses large tool output, and holds the ra
Package details
Install pi-warden from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:pi-warden- Package
pi-warden- Version
0.74.1- Published
- Sep 29, 2026
- Downloads
- 7,030/mo · 3,393/wk
- Author
- devmortimer
- License
- MIT
- Types
- extension
- Size
- 1.6 MB
- Dependencies
- 1 dependency · 2 peers
Pi manifest JSON
{
"image": "https://raw.githubusercontent.com/DevMortimer/pi-warden/main/docs/preview.png",
"extensions": [
"./extensions/index.js"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
pi-warden
Your coding agent says "done". It didn't run the tests. pi-warden catches it, and the agent fixes it without you.

759 sessions · 220 risky actions stopped before they ran · 76% of fake "done"s turned into real test runs
Nine days of the maintainer's real use across all their projects, not just this one, 2026-09-16 to 2026-09-24. Method and noise: field report.
Install
pi install npm:pi-warden
Then /warden enable (paste a TypeSafe key) and /warden init (writes a starter pi-warden.md). That's it. No key? The offline guards still run.
It steers. It doesn't nag.
Most guardrails stop and ask you. pi-warden tells the agent what it got wrong, and the agent corrects itself. You are pulled in only when something can't be undone: 3 holds per 1,000 calls. The other 997 just run.
| Your agent… | pi-warden… |
|---|---|
| says "done" with no test, build, or lint behind it | sends it back to prove it |
breaks a rule in your pi-warden.md or AGENTS.md |
quotes the exact rule it broke |
is about to git push --force, reset --hard, rm -rf, DROP |
holds it before it runs |
| retries the same failing fix for the third time | asks for a new hypothesis |
| writes stubs, restating comments, hardcoded secrets | names them on the spot |
| floods its context with a 40k-line log | keeps the lines that matter, stores the rest |
| starts repeating itself forever | stops the reply |
Every guard, with its thresholds and calibration →
Rules no linter can check
# A TODO names a ticket
A bare `TODO` or `FIXME` without a ticket reference is a violation.
# Errors never reach the user raw
paths: src/api/**/*.ts
Catch errors at the handler and return a message a user can act on.
Each # heading is one rule. Every write and edit is judged against it in about a quarter of a second by Jev, a model that returns a probability, not prose. No pi-warden.md? Your AGENTS.md, CLAUDE.md, or README.md is used instead.
Receipts
- Rules: in 150 paired agent runs, the agent without pi-warden broke the tested rule 6 times. With it: 0.
- Done-check: after a nudge, the agent ran a check 57 of 75 times, and sometimes found a failure it had missed.
- Holds: when the agent was stopped, it found a safer way 40 of 65 times; you approved 24.
- Stability: 13,952 guard cases over 109 overnight cycles, no score drift.
Every number has a script and a raw report in eval/reports/. They are the maintainer's measurements, not a universal promise, and the reports list what was noise.
Privacy
Secrets and unshown paths are stripped before anything leaves your machine. Exactly what is sent →
Docs
Guards · Configuration · Commands · FAQ · Data handling · Examples
Development
npm install
npm run check # typecheck + offline tests + build
MIT licensed.
