principal-pi-skills
Skill framework for principal-level software engineering: eight skills (decide, architect, plan, build, review, debug, investigate, git-ops), a pi bootstrap, and two approval-gated workflows.
Package details
Install principal-pi-skills from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:principal-pi-skills- Package
principal-pi-skills- Version
4.7.2- Published
- Oct 1, 2026
- Downloads
- 2,053/mo · 648/wk
- Author
- mojo_manyana
- License
- MIT
- Types
- extension, skill, prompt
- Size
- 237 KB
- Dependencies
- 0 dependencies · 0 peers
Pi manifest JSON
{
"skills": [
"./decide",
"./architect",
"./plan",
"./build",
"./review",
"./debug",
"./investigate",
"./git-ops"
],
"prompts": [
"./prompts"
],
"extensions": [
"./extensions/bootstrap.ts"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
principal-pi-skills
Eight skills for principal-level software engineering with the
pi coding agent — three inline skills and five
that double as subagents. Dialogue and session state run inline (decide, architect,
git-ops); heavy reading, cold judgment, and noisy loops delegate to isolated contexts
(plan, build, review, debug, investigate — single-shot variants in agents/,
generated from the same contract as the skill). The files follow the
Agent Skills standard, so other harnesses can
consume the skills, but pi is the supported target. This is 4.x: a thin orchestration
layer over the eight skills, with the risk-adaptive
assurance controller and broad model qualification kept on a separate track (see
Validation).
The set is built for one principal engineer steering at a high level while skills and subagents do the work. Two properties follow, and every design choice below serves them: delegable trust — an output carries the evidence needed to verify it without redoing the work — and cheap iteration — a defect found is a defect fixed, not documented around.
Three constraints
- Dual-use.
plan,build,review,debug, andinvestigateeach serve as a loaded skill and as a subagent system prompt. Both forms are rendered from one contract, so the shared behavior cannot drift between them, and the differences — single-shot mechanics, the BLOCKED form, no-dialogue rules — are marked rather than remembered. That constraint is what forces single-shot-safe behavior and a literal output template. - Model-agnostic. Written for the weakest model that will run it (DeepSeek, GLM,
Sonnet-class), not the strongest: imperative numbered steps, literal fill-in templates,
plain-text tags (
[ONE-WAY],[BLOCKER]) instead of an emoji schema, no aphorisms doing load-bearing work, no personas, no required reading in reference files. - Token economics. Budgets stated as decisions rather than aspirations: skills
≤ 1400 words, agents ≤ 1500, and
git-opsan accepted exception at ≤ 2150 — the safety-critical operator carries the most arming, and validated behavior outweighs a budget. Ceilings move only to buy a fix rather than more prose: an absolute is cheap to write and wrong in real cases, and a rule plus the cases it must not eat costs more words than the absolute it replaced. When a fix and the ceiling conflict, the ceiling moves. Every count in the table below is checkable withwc -w. Nothing loads anything else — a subagent reads one file and has the whole contract.
Individual fidelity exceptions (skill/agent words): Plan 1900/1950 buys complete source reads, blocking exceptions, stable mapping, persisted completeness and explicit no-shell fallback/tiny-file reads. Build 1850/2000 preserves pre-mutation authority classification, the complete authorized-amendment positive case, immutable report/repair provenance and compressed evidence/caveats; Review 1700/1750 buys candidate-bound obligation/gate evidence and repair definitions. Both distinguish lasting behavior-named regressions from historical receipt/replay verification, without freezing transient review/release status. The 100-word ceiling increases retain evidence-only scope safeguards and useful authorized regressions without removing existing guards; Debug's agent ceiling is 1550 for honest sandbox/applied states. Common ceilings and Git-Ops stay unchanged. These budgets preserve safeguards for lower-cost models, not a claim of measured robustness on those models.
The set
| Skill | What it does | How it runs | Words |
|---|---|---|---|
decide |
Options and stress-tests for a decision that isn't settled — "should I", "what are my options", "I'm stuck" | inline | 1134 |
architect |
System structure from measurable drivers; components, boundaries and data. The decision record is a section of the output, not a separate artifact | inline | 1211 |
plan |
A task turned into ordered steps and per-step specs a builder can execute without making load-bearing decisions. Writes no code | subagent (agents/principal-plan.md, 1938) or inline |
1877 |
build |
Test-first implementation — code proven by a test you watched fail | subagent (agents/principal-build.md, 1845) or inline |
1801 |
review |
One pass, two axes — correctness and simplicity — ending in one severity-ranked verdict | subagent (agents/principal-review.md, 1724) or inline |
1660 |
debug |
Hypothesis before fix: a diagnosis loop ending in a note with root cause and a regression test | subagent (agents/principal-debug.md, 1533) or inline |
1397 |
investigate |
A factual report of how code, data, runtime, or history currently behaves, with file-and-line citations | subagent (agents/principal-investigate.md, 491) or inline |
492 |
git-ops |
Safe version-control operator — reads state before writing it, keeps published history immutable, scans for secrets before committing | inline, never delegated | 2145 |
Routing between them belongs to the orchestrator, not to a skill — there is deliberately no routing skill spending context to say "pick a skill". AGENTS.md is the long-form routing reference, and the bootstrap extension below injects its routing table automatically at session start, so nothing needs to point pi at it by hand.
Bootstrap and workflows
A pi extension (extensions/bootstrap.ts) injects bootstrap/BOOTSTRAP.md — under 300
words — as a leading message at session start and again after compaction, so the routing
context survives a context reset instead of depending on someone re-reading a file. It
carries the routing table compressed to input shape → skill → inline/subagent, the closed
Next: vocabulary the phases hand off with, and model tiering: the cheapest model for a
build agent working a complete step spec, the session default for plan and debug, the
strongest available for review and architect.
The three spines (/principal-feature <task>, /principal-bugfix <symptom>,
/principal-refactor <scope>) stop for your approval after the planning phase or the debug
note and wait for an explicit go
before building — presenting the artifact and starting to build in the same turn is the
failure the rule exists to catch. The artifact scales with the change (three lines for a
config tweak, full slices for a feature); the stop does not.
Delegated phases hand artifacts to each other as files, not pasted text. A principal-build
writes its full report to an unused task/run/candidate-scoped path under .principal/reports/
(or a valid unused caller-chosen path) and returns five status lines;
review receives governing source/definition references, the complete plan/map when present,
and a candidate-identified diff package plus reports. Only matching candidate/scope evidence
is reusable; a passing suite alone is not complete requirement coverage. Every successful
report contributes assumptions, follow-ups and evidence gaps to the Digest. Repairs carry
full review-report paths, finding definitions and acceptance conditions, not IDs alone;
scoped re-review retains original whole-change evidence and gates. Independent runs never
reuse prior report/diff paths; repairs retain original review path and reviewed candidate.
Before any artifact write, the orchestrator creates an absent .principal/.gitignore
containing *, never overwrites it, and checks report destinations are ignored. This covers
planless Review, inline Build and optional Investigate persistence too. An existing policy
that exposes reports requires caller repair, not accidental runtime files in git.
When plan's output is the multi-step template, it writes the plan to
.principal/plans/<slug>.md and prints the path; .principal/.gitignore is created
alongside it only if absent, never overwritten. The complete executable artifact and map
stay in the file; chat gives a short summary and approval cue. Without persistence, the full
artifact is returned in chat, explicitly not saved. Tiny normative changes retain source/ID,
step and test without full machinery. Delegated
principal-build agents read their assigned step from that file, and a fresh or compacted
session that finds a matching plan resumes from it: read the file and git log, mark done
whatever already has a commit, continue at the first undone step, and never re-plan without
being asked.
Layout
<skill>/SKILL.md the interactive contract — nothing else is required reading
agents/principal-{plan,build,review,debug,investigate}.md subagent definitions available for delegation
contracts/{plan,build,review,debug,investigate}.md.tmpl source for dual-use contracts — edit here, run `npm run generate`
contracts/workflows.md.tmpl source for the three namespaced spines
prompts/principal-{feature,bugfix,refactor}.md generated workflows
prompts/principal-review-branch.md handwritten planless/buildless review entry
bootstrap/BOOTSTRAP.md routing table + Next: vocabulary + model tiering, injected by the extension
extensions/bootstrap.ts pi extension: injects BOOTSTRAP.md at session start and after compaction
scripts/ generator, installers, and checks behind `npm test`
tests/{unit,install}/ current product/contract + clean-home install tests (node:test)
evals/{triggers.json,adversarial-triggers.json,baseline/} routing checks and baselines
AGENTS.md routing + dispatch reference; the bootstrap injects its table automatically
CHANGELOG.md release history
Install (pi)
Skills + prompts — install an immutable tag, not a branch:
pi install git:github.com/mojomanyana/principal-pi-skills@v4.7.2The v4.7.2
pimanifest registers the eight skills, the four/principal-*commands, and the bootstrap extension — it loads automatically with the package; there is no separate extension-install step. Unpinnedmainmoves under you, so install a tag if you want a fixed, nameable behavior.Do not install
2.3.0— it is deprecated on npm for a destructive defect: itsprincipal-pi-workspace removedeletes any path handed to it, including your checkout, and reports success.2.3.1is the lowest safe version.Subagents (optional). Use the same source as the installed skills. Before npm publication, npm latest is 4.7.1; unpinned
npxwould install/check older definitions, and npm@4.7.2is not yet available. Either locate the actual installed tagged package (its path varies by Pi configuration) and run itsscripts/install-agents.mjswith Node, or use this concrete matching-tag disposable checkout recipe:( set -eu source_dir=$(mktemp -d) trap 'rm -rf -- "$source_dir"' EXIT git clone --depth 1 --branch v4.7.2 https://github.com/mojomanyana/principal-pi-skills.git "$source_dir/package" test "$(git -C "$source_dir/package" rev-parse HEAD)" = "$(git -C "$source_dir/package" rev-parse 'v4.7.2^{commit}')" git -C "$source_dir/package" rev-parse HEAD # retain the resolved source identity node "$source_dir/package/scripts/install-agents.mjs" install node "$source_dir/package/scripts/install-agents.mjs" check )After npm publication, and only once the matching package is available, pin both commands:
npx -p principal-pi-skills@4.7.2 principal-pi-agents install npx -p principal-pi-skills@4.7.2 principal-pi-agents checkVersion 4.7.2 moves behavioral measurement to the separate
principal-pi-skills-evalsrepository while retaining routing checks here. Use the commands above after the v4.7.2 tag exists; npm publication is separate. The checkout check verifies tag/HEAD consistency, not independent tag trust; compare the recorded SHA with the release identity when provenance matters. Before tagging, validate the actual candidate's skills and installer in an isolatedPI_CODING_AGENT_DIR. Known behavioral qualification limits remain documented; combining the release does not waive them.It installs
principal-plan,principal-build,principal-review,principal-debug, andprincipal-investigateas real files, not symlinks — a symlink into a checkout breaks the moment that directory moves, and breaks silently, since pi just reports an unknown agent. It refuses to overwrite anything it did not install, anduninstallremoves only its own unmodified files.Tool restriction is structural, in the agents' frontmatter:
investigateis read-only;planis read-only except for its plan and creation of an absent.principal/.gitignore;build,review, anddebugaddbashto run tests (and, forbuild, to write and edit).One trap worth knowing if you run subagents on a non-default provider: a delegated agent runs on the pi config's
defaultProvider/defaultModel, not the--provider/--modelyou gave the parent session — the extension forwards--modelonly when an agent's frontmatter names one, and these deliberately do not. If delegations fail to authenticate while the parent session is fine, that mismatch is the reason.Without the subagent step, everything still runs completely inline via the skills; the How column in The set simply collapses to "inline". The routing table still reaches the session because the bootstrap extension injects it — installing the package is enough for that part; only delegation itself needs step 2.
Under pi-daddy (0.33.0+), each skill declares how much of the caller's session it may receive. A child gets only the
context:mode its ownallowed-toolsnames (or a weaker one:none < files < pruned < summary < fork). Asking for more is refused, not downgraded. The ceilings are decisions, and each one is explained in its frontmatter:skill ceiling why architect,decidecontext:summarydelegated, they cannot ask; the drivers and the rejected options live in the parent's dialogue plancontext:summarya plan must honour what the user ruled out, which a task line flattens; no bashreviewcontext:filesdeliberately low. Review is cold by design, and the author's reasoning is what it must not be anchored on; the diff package and build report are files build,debugcontext:filesthey act on a stated target (a plan file or a symptom); logs travel verbatim as files; with bash, anything a child receives can leave the machineinvestigatecontext:filesreceives only named evidence and has no shell or write tools; the caller's reasoning is not evidence git-opsnone acts on the working tree, not the conversation; consent to a destructive op must come from the user, not from forwarded turns Nothing declares
context:fork. Write the prefix in lowercase:Context:summaryturns intotool:context:summary, which grants no context mode.Egress: know what
summaryenables.context:summaryalso permitspruned. That mode carries the operator's own session turns (the last 20 by default, up to 32 KiB), not just the task. With a pi-daddy advisor enabled (PI_DADDY_ADVISORplusPI_DADDY_ADVISOR_KEY), aprunedhandoff sends those turns to that third party so it can choose which ones to keep. Forarchitect,decideandplan, that is what you switch on by delegating withprunedwhile an advisor is on. Raising any other skill's ceiling extends that exposure to it.
Validation
npm test remains the free gate: generated-contract drift, word budgets, frontmatter lint,
installer and tarball behavior, Next: transition parity, and deterministic source-fidelity
contract assertions. These check current product contracts, not model behavior or instruction delivery.
Behavioural measurement—including the requirement-fidelity corpus and skill-harness evidence—lives in principal-pi-skills-evals; routing checks live here.
Enforcement boundary: Build/Review instructions ask the model to organize regressions by lasting behavior, not PR/finding IDs, and distinguish product coverage from historical verification. Evidence-only requests first use existing checks or disposable probes, not permanent tests merely to prove a finding was addressed; ordinary approved bugfix/feature regressions and useful explicitly requested checks remain authorized. Markdown contracts are product, with legitimate structural tests. These are not filesystem restrictions or a semantic test-quality checker. The package commands and CI mechanically select suites; pi-daddy's delegated tool grants control available capabilities, not test names or correctness. Refresh installed skill/agent definitions when adopting a revision: editing this checkout does not update another installed copy.
Two opt-in routing checks use only the eight authored frontmatter descriptions. Run
npm run check:routing-collisions for all 56 directed description pairs and
npm run check:routing-triggers for the three-run synthetic trigger suite. Both require
FIREWORKS_API_KEY; ROUTING_MODEL and ROUTING_API_URL override the defaults. The suite
starts with 12 positives and 8 near-miss negatives per skill, then
adds authored collision-boundary and no-skill probes, and reports per-skill precision and recall
at a 0.5 vote threshold. evals/baseline/triggers.json records the expanded three-run corpus;
the obsolete pre-expansion record was removed. Regenerate the baseline after any corpus or
routing-description change.
Why 4.0
Version 3.x grew a risk-adaptive assurance controller — a hash-chained event ledger, task
packets, digests, fail-closed gates — whose protocol leaked into the model-facing skill text
and whose init step ran before every workflow, including a typo fix. The routing layer in
AGENTS.md was never loaded by pi, the feature spine had no human approval point outside
critical mode, and every build ran inline so long features filled the steering context with
diffs and test output. The seven skills were the strongest part of the repo and 12% of its
Markdown. 4.0 returns the repo to skills plus a thin orchestration layer and borrows four
mechanisms from superpowers that serve the north star.
Decisions taken for 4.0, all closed:
| Decision | Chosen |
|---|---|
| Target harness | pi only |
| Assurance ledger and profiles | removed; per-skill right-sizing is the mechanism; the tool lives on the v3.2.0 tag for porting to pi-daddy |
| Human approval | always, after plan (feature) or after the debug note (bugfix); the artifact scales, the stop does not |
| Routing delivery | a pi extension injects bootstrap/BOOTSTRAP.md at session start and after compaction |
| Build delegation | principal-build agent; inline when there is no multi-step plan file or no subagent tool |
| Plan persistence | multi-step plans to git-ignored .principal/plans/<slug>.md; no date prefix because plan has no clock, and resume matches on the ## Plan: line |
| Decide vs architect | both kept; decide answers "should we / which", architect answers "how is it structured" |
| Measurement | behavioural measurement lives in principal-pi-skills-evals; routing checks stay here |
Two implementation notes that differ from the obvious reading: the bootstrap is injected as
a user-role message wrapped in <IMPORTANT>, because pi's context hook can only insert
messages; and a delegated principal-build may run without a plan file (the bugfix spine and
repair rounds), in which case the prompt's task is its whole spec.
Deliberate design rules
Why the files look the way they do. Each of these was learned by measuring the alternative.
- Description = triggers only. Never a workflow summary — a description that summarizes the process trains the model to follow the description and skip the body.
- Recipes, not prohibition tables. Output-shape problems get a literal template to fill. Prohibitions are reserved for genuine discipline failures (skipping tests under pressure, force-push, secret handling), where a short Checks table remains.
- One governor sentence instead of a governor table per skill. If a skill needs a table of reasons not to use itself, it is over-scoped.
- Assumptions instead of questions in delegated mode. A subagent cannot ask, so every
skill says what to do when information is missing: state the assumption, or return
BLOCKEDwith the one question that matters. - Pressure armor is explicit. Discipline rules carry "repetition doesn't change the answer — any turn, including the last", because models otherwise cave on the third push.
- Right-sizing is a hard conditional, not a suggestion: "2–5 sentences, no machinery", and when a user asks for the artifact on a trivial change, the minimal form is the deliverable — otherwise the model declares the artifact unwarranted and produces it anyway.
- Grounded skills carry a no-repo branch.
planandgit-opsact on the material given instead of stalling on "point me at the repo". - Weak models need code anchors.
debug's error-swallowing rule survived two rounds of prose and died to one literalcatchexample. Escape hatches work best inside the template they exempt.
License
MIT © 2026 Nemanja Alavanja. See LICENSE.