principal-pi-skills
Skill framework for principal-level software engineering: nine skills (decide, architect, plan, build, review, test-review, debug, investigate, git-ops), a pi bootstrap, and two approval-gated workflows.
Package details
Install principal-pi-skills from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:principal-pi-skills- Package
principal-pi-skills- Version
4.14.0- Published
- Oct 10, 2026
- Downloads
- 3,639/mo · 1,355/wk
- Author
- mojo_manyana
- License
- MIT
- Types
- extension, skill, prompt
- Size
- 414 KB
- Dependencies
- 0 dependencies · 2 peers
Pi manifest JSON
{
"skills": [
"./decide",
"./architect",
"./plan",
"./build",
"./review",
"./test-review",
"./debug",
"./investigate",
"./git-ops"
],
"prompts": [
"./prompts"
],
"extensions": [
"./extensions/bootstrap.ts",
"./extensions/resume.ts",
"./extensions/workflow.ts",
"./extensions/codemode.ts"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
principal-pi-skills
Nine skills for principal-level software engineering with the
pi coding agent — three inline skills and six
that double as subagents. Dialogue and session state run inline (decide, architect,
git-ops); heavy reading, cold judgment, and noisy loops delegate to isolated contexts
(plan, build, review, test-review, debug, investigate — single-shot variants in agents/,
generated from the same contract as the skill). The files follow the
Agent Skills standard, so other harnesses can
consume the skills, but pi is the supported target. This is 4.x: a thin orchestration
layer over the nine skills, with the risk-adaptive
assurance controller and broad model qualification kept on a separate track (see
Validation).
The set is built for one principal engineer steering at a high level while skills and subagents do the work. Two properties follow, and every design choice below serves them: delegable trust — an output carries the evidence needed to verify it without redoing the work — and cheap iteration — a defect found is a defect fixed, not documented around.
Deterministic workflow tools
Principal 4.12 adds principal_workflow for candidate identity, exact file references and
private report handoffs. snapshot and reference support any coordinator work. For native
phases, prepare allocates one ignored report path and stable operation ID; pass that ID to
Daddy delegation. After settlement and before repairs, complete retains the exact captured final
and producer identity in immutable result.json; completion.json requires current matching
inputs/candidate. Stale results remain history, not current approval. A differing existing report is preserved and refused;
an identical final may be reused. result verifies and returns a settled native final without
storage for strict no-files tasks; skip prepare/complete in that mode. status checks current bytes before reuse. Explicit retry preserves
the old attempt and needs native settled or never-started proof. A new coordinator session
cannot silently redispatch an old prepared operation; inspect and reconcile it first. These records are evidence,
not review approval. Native result retention requires the compatible pi-daddy installation pinned below in the same session.
The tool uses the existing candidate snapshot implementation with the named
principal-candidate-v1 format. It refuses stale inputs, conflicting requests, exposed report
paths and uncertain execution. It does not reconstruct historical fingerprint formulas.
Inline work does not require a dummy delegation to complete a record. Pass relevant authority/
evidence as inputPaths for tool-generated references and the returned candidate object as
expectedCandidate; product-file hash lists are unnecessary. complete, status and retry
accept operationId alone to locate the recorded task/step. See mechanics.
principal_codemode reuses the available native Pi Codemode factory with models: false.
It batches tool calls through native hooks; no classifier/model API is exposed by this
adapter. Independent reads may run together, while writes and dependent operations remain
sequential. It is a coordinator facility; governed child Codemode remains unqualified.
Hosts without the factory simply do not register the adapter. Model-free loading and
operation-bridge checks passed with Pi 1.0.4 and 1.1.0; this is not paid model qualification.
Three constraints
- Dual-use.
plan,build,review,test-review,debug, andinvestigateeach serve as a loaded skill and as a subagent system prompt. Both forms are rendered from one contract, so the shared behavior cannot drift between them, and the differences — single-shot mechanics, the BLOCKED form, no-dialogue rules — are marked rather than remembered. That constraint is what forces single-shot-safe behavior and a literal output template. - Model-agnostic. Written for the weakest model that will run it (DeepSeek, GLM,
Sonnet-class), not the strongest: imperative numbered steps, literal fill-in templates,
plain-text tags (
[ONE-WAY],[BLOCKER]) instead of an emoji schema, no aphorisms doing load-bearing work, no personas, no required reading in reference files. - Token economics. Budgets stated as decisions rather than aspirations: skills
≤ 1400 words, agents ≤ 1500, and
git-opsan accepted exception at ≤ 2550 (including operation-specific preservation and verified restoration) — the safety-critical operator carries the most arming, and validated behavior outweighs a budget. Ceilings move only to buy a fix rather than more prose: an absolute is cheap to write and wrong in real cases, and a rule plus the cases it must not eat costs more words than the absolute it replaced. When a fix and the ceiling conflict, the ceiling moves. Every count in the table below is checkable withwc -w. Nothing loads anything else — a subagent reads one file and has the whole contract.
Individual contract exceptions (skill/agent words): Plan 2050/2100 preserves complete source/requirement maps and adds exact connected interfaces and shared-resource checks. Build 2100/2100 retains authority, full evidence and immutable repair provenance while adding meaningful test maintenance and explicit gate ownership/due stages. Review 1900/1950 retains candidate-bound gates and distinguishes task, integrated and repair scope. Debug 1453/1600 preserves honest disposable/applied states while making diagnosis sufficient for handoff and protecting boundary evidence. These narrow increases retain the existing safeguards; they are not claims of measured model compliance. Common ceilings remain unchanged.
The set
| Skill | What it does | How it runs | Words |
|---|---|---|---|
decide |
Options and stress-tests for a decision that isn't settled — "should I", "what are my options", "I'm stuck" | inline | 975 |
architect |
System structure from measurable drivers; components, boundaries and data. The decision record is a section of the output, not a separate artifact | inline | 1270 |
plan |
Outcomes, constraints and acceptance, with essential interfaces pinned and routine design left to Build. Writes no code | subagent (agents/principal-plan.md, 1591) or inline |
1525 |
build |
Implementation with meaningful behavior and regression evidence | subagent (agents/principal-build.md, 1728) or inline |
1713 |
review |
One pass, two axes — correctness and simplicity — ending in one severity-ranked verdict | subagent (agents/principal-review.md, 1348) or inline |
1296 |
test-review |
Test assertions, missing behavioral cases, realistic mocks and justified consolidation; no product integration approval | subagent (agents/principal-test-review.md, 799) or inline |
778 |
debug |
Hypothesis before fix: a diagnosis loop ending in a note with root cause and a regression test | subagent (agents/principal-debug.md, 1573) or inline |
1433 |
investigate |
A factual report of how code, data, runtime, or history currently behaves, with file-and-line citations | subagent (agents/principal-investigate.md, 538) or inline |
539 |
git-ops |
Safe version-control operator — reads state before writing it, keeps published history immutable, scans for secrets before committing | inline, never delegated | 2523 |
Routing between them belongs to the orchestrator, not to a skill — there is deliberately no routing skill spending context to say "pick a skill". AGENTS.md is the long-form routing reference, and the bootstrap extension below injects its routing table into applicable requests, so nothing needs to point pi at it by hand.
Bootstrap and workflows
The extension (extensions/bootstrap.ts) discovers this package's enabled selected skills
through public pi.getCommands() at before_agent_start, after resource discovery. It
adds the under-300-word bootstrap/BOOTSTRAP.md once to each converted request, preserving
leading compaction/system/tool state and durable history. Tool iterations and later ordinary
requests receive routing too. Only extension-owned inserted objects are replaced; a quoted
marker cannot suppress injection. Disabled or shadowed package resources stay inactive.
Principal read/discovery failure is visible and latched until deliberate /reload reinstantiates the
extension; stale foreign command paths are ignored independently without disabling valid Principal skills; each registration owns its content cache and selected-resource state.
Native handoffs use delegate_describe({agent:"phase"}), require binding package
principal-pi-skills and exact phase. Native names are plan, build, review, test-review,
debug, investigate; pass the corresponding captured definitionId to every delegate call,
delegate_all child and delegate_chain step. The root principal-agents.json manifest binds exactly six phases to generated
inline/delegated paths and full-file SHA-256 hashes from the same generation pass. The
runtime must verify enabled selected source, package identity and both hashes. Missing,
wrong or disabled bindings and operational failures stop dependent work; they never trigger
legacy/foreign/inline substitution. Only genuinely absent native tools permit an explicitly
configured legacy principal-* runner. Inline is a workflow choice. Authored model/effort
policy governs; phase labels do not silently override it. This contract alone does not qualify
live native execution; the actual installed runtime/candidate still needs integration evidence.
Use one delegate_all({children:[...]}) call for an approved independent parallel batch,
with the exact native phase name and captured definitionId on each child. Single delegate
calls conservatively reserve the available subtree capacity; overlapping single calls can
be refused. Writers need distinct precreated registered Git worktrees and declared scopes,
dependencies, shared resources and full-report destinations. Read every outcome, preserve
completed siblings after failure, review candidates, integrate serially and check the merged whole.
The three spines (/principal-feature <task>, /principal-bugfix <symptom>,
/principal-refactor <scope>) present a newly needed plan/design or debugging note before
Build. They obtain approval when it is still needed; explicit existing user authorization
for the actual scope persists and does not need to be requested again. Urgency, a title or
a self-written flag cannot supply it. Material scope changes and one-way actions retain
their applicable approval boundaries.
A clear authorized feature can go directly to Build; use a persisted plan only when sequencing, material risks or dependencies need it. The named feature/refactor prompts follow the same proportional routing. Ordinary plans state behavior, affected files and acceptance; formal source IDs retain their required map. Empty metadata fields and repeated generic governance do not belong in every plan or report.
A routine feature keeps implementation, acceptance tests and documentation together, followed by one independent integrated Review. Split for a real dependency, risk boundary or useful parallel increment. A public API addition does not by itself require a skeleton phase, a separate test phase and intermediate review. Task review still protects consequential interfaces consumed across implementation boundaries and parallel candidate integration.
Read the governing task sources and needed definitions directly; follow historical reports when they supply current authority, unresolved findings or evidence needed now. Generic coordinator gates stay addressable once, rather than becoming repeated requirement matrices. Exact source IDs, task clauses, approvals, binding, candidate identity and settlement remain required. A single declared full-suite run can also be green evidence. Reuse unchanged baseline/examples only after verifying relevant identities and environment; mutation probes are optional risk-based checks, not a requirement to fabricate red for already-correct behavior. Finish honors an explicit known preference, including leaving changes uncommitted on the actual branch, without asking again.
Plan resolves essential outcome and interface tradeoffs; Build owns the simplest coherent, readable and maintainable implementation within that authority. Routine design choices need no repeated permission. Review judges design and test quality alongside correctness; neither green tests nor a coverage percentage substitutes for sound engineering judgment. The distinct test-review skill checks whether assertions detect plausible wrong implementations, including exact errors, configurable capacities and state changes that could mask a failure. A code reviewer may use it with a separate test-quality assessment; that is one reviewer using two skills. A separately delegated test-review child provides independent test judgment when requested or useful, without an always-required extra agent. Neither mode grants integration approval. No test-count, coverage or mutation-score quota replaces judgment. Workflow mechanics contains conditional archive/progress/resume details, read only when that facility is used.
Native children return their complete report once in the final message. After settlement,
principal_workflow complete retains that exact captured final at its prepared private path;
the child does not write another report. Plan still persists its actual executable plan.
Failed completion stays incomplete and visible. This removes model-written report copies
and repeated path checks. An explicitly configured legacy caller can request safe persistence.
Delegated phases consume retained files;
review receives governing source/definition references and relevant plan sections/map rows,
and a candidate-identified diff package plus reports. Only matching candidate/scope evidence
is reusable; a passing suite alone is not complete requirement coverage. Every successful
report contributes assumptions, follow-ups and evidence gaps to the Digest. Repairs carry
full review-report paths, finding definitions and acceptance conditions, not IDs alone;
scoped re-review retains original whole-change evidence and gates. Independent work gets
new report/diff paths; resuming the same prepared operation reuses its verified references.
Repairs retain original review path and reviewed candidate. UNVERIFIED returns
Next: evidence for coordinator reconciliation, not automatic Build. A concrete product/test
finding routes to Build only within repair authority. Reuse matching receipts; run a targeted
check for a named doubt, not every suite merely to obtain a fresh reviewer. Stop/reassess
repeated failures without new evidence rather than imposing a fixed repair-round quota.
Before any artifact write, the orchestrator creates an absent .principal/.gitignore
containing *, never overwrites it, and checks report destinations are ignored. This covers
planless Review, inline Build and optional Investigate persistence too. An existing policy
that exposes reports requires caller repair, not accidental runtime files in git.
When plan's output is the multi-step template, it writes the plan to
.principal/plans/<slug>.md and prints the path; .principal/.gitignore is created
alongside it only if absent, never overwritten. The complete executable artifact and map
stay in the file; chat gives a short summary and approval cue. Without persistence, the full
artifact is returned in chat, explicitly not saved. Tiny normative changes retain source/ID,
step and test without full machinery. Delegated
build children read their assigned step from that file, and a fresh or compacted
session locates matching plans, then reconciles actual approval, current work and candidate-bound
reports. A commit does not establish review, integration or verification. Resume manually at
what the evidence shows remains; do not repeat completed work or infer approval from a title.
Layout
<skill>/SKILL.md the interactive contract — nothing else is required reading
agents/principal-{plan,build,review,test-review,debug,investigate}.md subagent definitions available for delegation
contracts/{plan,build,review,test-review,debug,investigate}.md.tmpl source for dual-use contracts — edit here, run `npm run generate`
contracts/workflows.md.tmpl source for the three namespaced spines
prompts/principal-{feature,bugfix,refactor}.md generated workflows
prompts/principal-review-branch.md handwritten planless/buildless review entry
bootstrap/BOOTSTRAP.md routing table + Next: vocabulary, injected by the extension
extensions/bootstrap.ts pi extension: request-local routing after selected-resource discovery
test-review/SKILL.md separate test-quality judgment; optional independent child
references/workflow-mechanics.md optional archive/progress/resume mechanics
scripts/ generator, installers, and checks behind `npm test`
tests/{unit,install}/ current product/contract + clean-home install tests (node:test)
evals/{triggers.json,adversarial-triggers.json,baseline/} routing checks and baselines
AGENTS.md routing + dispatch reference; the bootstrap injects its table automatically
CHANGELOG.md release history
Install (pi)
Skills + prompts — install the exact npm release:
pi install npm:principal-pi-skills@4.14.0 pi install npm:pi-daddy@0.49.1 pi install npm:skill-harness@0.28.0Restart Pi after package changes. The
pimanifest registers nine skills, four/principal-*commands and the bootstrap extension automatically. Exact npm pins keep the installed source version identifiable.For planning or review that needs repository discovery, merge this setting into your Pi
settings.json(preserving its other fields), then restart Pi:{ "defaultTools": ["read", "bash", "edit", "write", "grep", "find", "ls"] }Pi 1.0.4's default coding tool selection omits
grep,findandls. This setting enables those built-in discovery tools while retaining extension tools such asdelegate. Check the active tool inventory after restart. Plan still forbids shell substitution and may correctly block when essential discovery is unavailable. During Plan, describe later coordinator-owned progress work; the coordinator runs the helper after that phase, within its own authority. Planning does not run the CLI or widen its write ceiling.Do not install
2.3.0— it is deprecated on npm for a destructive defect: itsprincipal-pi-workspace removedeletes any path handed to it, including your checkout, and reports success.2.3.1is the lowest safe version.Native report-completion composition is checked with Pi 1.0.4 and 1.1.0. Captured execution uses
PI_DADDY_HERDR=0; Herdr (PI_DADDY_HERDR=1) additionally requires pi-daddy's live compatibility and native ownership checks. This is runtime qualification, not a model-quality claim for every provider. The selected package's generatedprincipal-agents.jsonbinds each phase to its skill and agent bytes. Native routing usesdelegate_describeand the captureddefinitionId; refusal or failure cannot fall back to a legacy runner or inline execution.Legacy subagents (optional). A configured legacy runner can use the six definitions when native tools are genuinely absent. Native delegation needs no separate agent install. Use the same npm package version as the installed skills:
npx -p principal-pi-skills@4.14.0 principal-pi-agents install npx -p principal-pi-skills@4.14.0 principal-pi-agents checkAlternatively run
scripts/install-agents.mjswith Node from the actual selected npm package; resolve its path from Pi resources instead of guessing an installation directory. Source tags and npm publication are separate. Validate candidate installs in an isolatedPI_CODING_AGENT_DIRbefore release; known behavioral qualification limits still apply.It installs
principal-plan,principal-build,principal-review,principal-test-review,principal-debug, andprincipal-investigateas real files, not symlinks — a symlink into a checkout breaks the moment that directory moves, and breaks silently, since pi just reports an unknown agent. It refuses to overwrite anything it did not install, anduninstallremoves only its own unmodified files. Symlinked home/config ancestors resolve once to a canonical directory; agent and manifest leaf symlinks (including dangling links) still refuse. Installed agent files are0644; the ownership manifest is0600. Re-runninginstallrepairs the earlier0600agent mode only for unchanged owned files.Ownership migration/recovery:
installno longer silently adopts byte-identical unowned files, and deprecated--forcenever bypasses validation or ownership. If only the manifest was lost and all six agents exactly match this package version, explicitly runnode <installed-package>/scripts/install-agents.mjs adopt, theninstall(to repair unchanged owned file modes), thencheck.adoptcreates only the absent manifest and leaves agent bytes/modes untouched; it refuses an existing manifest, partial sets, modified/older content and leaf symlinks. Unrelated files remain unclaimed. Keep any edited or older files and restore known-good metadata from your own backup; do not delete them merely to make installation pass.Tool restriction is structural, in the agents' frontmatter:
investigateis read-only;planis read-only except for its plan and creation of an absent.principal/.gitignore;build,review,test-review, anddebugaddbashto run tests (and, forbuild, to write and edit).Legacy runners differ in model forwarding. If a runner does not forward the parent's provider/model, check that runner's defaults and agent-frontmatter rules when a child cannot authenticate. Native delegation uses the matching pi-daddy runtime's resolved model and effort policy.
Inline execution remains available when the workflow selects it. The bootstrap supplies routing with the installed skills. Skipping the legacy agent installation does not disable native delegation, and a failed native handoff cannot silently become inline.
With the matching pi-daddy candidate, each skill declares how much of the caller's session it may receive. A child gets only the
context:mode its ownallowed-toolsnames (or a weaker one:none < files < pruned < summary < fork). Asking for more is refused, not downgraded. The ceilings are decisions, and each one is explained in its frontmatter:skill ceiling why architect,decidecontext:summarydelegated, they cannot ask; the drivers and the rejected options live in the parent's dialogue plancontext:summarya plan must honour what the user ruled out, which a task line flattens; no bashreview,test-reviewcontext:filesdeliberately low. Review is cold by design, and the author's reasoning is what it must not be anchored on; the diff package and build report are files build,debugcontext:filesthey act on a stated target (a plan file or a symptom); logs travel verbatim as files; with bash, anything a child receives can leave the machineinvestigatecontext:filesreceives only named evidence and has no shell or write tools; the caller's reasoning is not evidence git-opsnone acts on the working tree, not the conversation; consent to a destructive op must come from the user, not from forwarded turns Nothing declares
context:fork. Write the prefix in lowercase:Context:summaryturns intotool:context:summary, which grants no context mode.Egress: know what
summaryenables.context:summaryalso permitspruned, which can carry selected turns from the active parent branch to the delegated child's provider. The matching pi-daddy candidate no longer uses an external advisor to choose those turns. Forwarded context still reaches the child and its provider; raising another skill's ceiling extends that exposure to it.
Coordinator evidence handoffs
Before native delegation, save the complete unmodified delegate_describe response in a private
handoff report and retain accessible references to the actual selected skill/agent files and
package manifest. Resolve their paths from selected Pi resources, verify the binding hashes,
and preserve original paths/hashes when copying bytes into a child workspace. The handshake
contains identity data, not source paths or new authority. Pass the files, not a prose substitute.
Each gate records its owner and due stage. A typical assignment is:
| Due stage | Owner | Required observations |
|---|---|---|
| Before dispatch | Coordinator | Actual authority, selected definition and permitted runtime configuration |
| During implementation | Builder | Due source requirements, scope, tests, candidate identity and full report |
| After builder return | Coordinator | Complete native result, report and terminal/cleanup evidence |
| Candidate review | Reviewer | Exact candidate, applicable source requirements, supplied evidence and verdict |
| After reviewer return | Coordinator | Reviewer settlement, then authorized integration and final verification |
Native process settlement needs cleanup.state=settled with its matching identity-bound receipt.
An absent disposable workspace is not process-cleanup evidence. Retain the model-visible runtime
evidence emitted by pi-daddy 0.44.4; missing evidence blocks dependent work and task completion even with
passing tests or review.
This example cannot change source-required ordering. A leaf reports later coordinator gates as pending and does not certify its future exit or call tools it lacks. Missing evidence required for current work still blocks. Every dependent step waits for the coordinator's actual return checks. Final integrated review preserves global obligations; a task verdict covers only its scope.
Exact public evidence capture
With companion pi-daddy 0.45.0, an operator can opt in before starting Pi by setting
PI_DADDY_PUBLIC_EVIDENCE_DIR to an existing private owned directory. This option is off
by default and is not forwarded to children. It captures the native tool's exact public
{isError,content} response before its capture-reference block, plus selected definition/source
artifacts and existing runtime evidence. It excludes raw details, private native sessions,
and later Pi hook formatting. This is local evidence persistence, not automatic JEV collection.
The appended Public evidence capture: block returns status: "captured" and a manifest
ref with absolute path and SHA-256. Read that manifest and use its response/source references
directly in child handoffs and principal-pi-progress reference; do not transcribe the response
or produce a second capture manifest. Pass the unmodified response file and accessible selected
sources or exact captured copies with original identities. Verify the described binding and
source hashes as before. A manifest hash does not prove its referenced bytes or their meaning;
check each reference used by the current handoff. A status: "failed" capture retains the original
tool result but does not establish the missing evidence; stop any dependent work that needs it.
Capture is not authority, receipt validation, review approval or task acceptance. Keep full
semantic reports and explicit coordinator decisions separate. No setting is enabled by this package.
Installed disposable-workspace helper
Resolve the selected Principal package from Pi's selected skill/agent source metadata (or its
original source path in the coordinator handoff), verify that package's package.json, and
invoke its scripts/snapshot-workspace.mjs with Node. A captured copy's directory and
delegate_describe alone do not identify the selected installation. Do not guess another
npm prefix, use unpinned npx, or install a package to obtain this helper.
node /actual/selected/principal-pi-skills/scripts/snapshot-workspace.mjs create --repo /absolute/caller-repo
node /actual/selected/principal-pi-skills/scripts/snapshot-workspace.mjs remove /returned/worktree-path --repo /absolute/caller-repo
Use the same resolved helper for creation and cleanup. This selects the installed version
reliably; it does not diagnose an unknown failed npm invocation. The detached worktree holds
HEAD, staged/unstaged tracked changes and nonignored untracked files, including symlinks.
It flattens staged and unstaged changes, excludes ignored files such as dependencies/secrets,
and is not a full recovery backup. Missing helper or failed creation keeps Debug/Review
read-only. Removal refuses paths the helper does not own; a failure is not permission to
substitute recursive deletion. node --check /actual/selected/principal-pi-skills/scripts/snapshot-workspace.mjs
is a read-only syntax check. The current helper's --help prints usage with exit status 2.
Reports and manual progress
Use the installed principal-pi-progress helper when a coordinator needs repeated report
allocation/persistence. It creates private unused candidate directories beneath ignored
.principal/reports/, writes complete .md reports exclusively, and appends a small
progress.jsonl index. It preserves an existing ignore policy and refuses an exposed path.
Before dispatch, the coordinator allocates the primary report in the actual child workspace and
passes an absolute unused path plus report-write scope. For example, a builder running in a
worktree writes its report under that worktree's ignored .principal/reports/, even when the
operator also requested a separate evidence directory. The coordinator creates an absent ignore
file and checks the actual destination before dispatch; routine safe allocation within the task's
authority needs no extra approval. Existing exposed ignore policy, permission failures, or an
explicit prohibition on local artifacts still stop dispatch.
External evidence directories are coordinator archives. After native settlement, copy complete report bytes exclusively to the authorized private archive, verify the original and copy hashes, and record original path, SHA-256 and copy path. Refuse symlink redirects; retain originals through review, repair and resume. An archive failure preserves the primary report and blocks only work that requires that copy. Keep native public-capture manifests and their source refs unchanged.
Progress tracking is optional; when used, all index writes must go through the installed helper.
Never hand-write its run.json or progress.jsonl. Plan's write ceiling and Investigate's
read-only ceiling do not grow. If the binary is not on PATH, invoke
node /actual/selected/principal-pi-skills/scripts/progress-artifacts.mjs with the same arguments,
resolving the package from the selected Pi resources. Do not guess paths or fetch/install a helper.
principal-pi-progress create /path/to/repo task-name '<actual candidate identity>'
# Set RUN to the decoded returned path; use the same caller-established identity below.
principal-pi-progress report "$RUN" step-1-build.md < complete-report.md
principal-pi-progress reference /absolute/path/to/plan.md
principal-pi-progress append "$RUN" < record.json
principal-pi-progress check "$RUN" '<actual current candidate identity>'
# For diagnostics only (read reports issues but does not fail its exit status):
principal-pi-progress read "$RUN" '<actual current candidate identity>'
report and reference return { "path": "/absolute/path", "sha256": "<64 hex>" }.
For an existing producer file, copy <run> <name> <source> <expected-sha256> preserves exact
bytes at an unused private report path and returns {source:{path,sha256},copy:{path,sha256}}.
Obtain the expected hash from producer evidence or reference; do not recreate native JSON
or manually transcribe receipts. Copy checks byte integrity only, not approval or settlement.
For supported full candidate observations, the selected package's
scripts/resume-checkpoint.mjs candidate <repo> avoids ad hoc hashing; its Linux constraints
still apply, and full-tree identity does not prove partial-scope evidence equivalence.
A version 1 record has exactly these fields (replace REPORT_REF with that returned object):
{
"version": 1,
"plan": null,
"step": "impl:S1",
"candidate": "<actual candidate identity>",
"facts": {
"planned": { "state": "unknown", "evidence": [], "note": "No separate plan" },
"implemented": { "state": "complete", "evidence": ["REPORT_REF"], "note": "Implementation saved; review pending" },
"reviewed": { "state": "unknown", "evidence": [], "note": "Not reviewed" },
"integrated": { "state": "unknown", "evidence": [], "note": "Not integrated" },
"verified": { "state": "unknown", "evidence": [], "note": "Qualification not run" }
},
"findings": [],
"nextAction": "Request task review"
}
plan is null or a file reference. Every fact has unknown, incomplete or complete
state, evidence references and an honest note; complete needs evidence. A completed review
means the review occurred, and its note must retain the verdict, including CHANGES-REQUESTED.
Each finding has id, source reference and status (open, addressed, verified,
accepted, disputed, duplicate, stale); a duplicate also names duplicateOf. Preserve original
review baseline/definitions and PR comment URLs/IDs in the referenced complete reports.
Run check after each append and before relying on a resumed or final index. Its explicit
candidate argument is a caller assertion, not a Git snapshot. It returns integrityValid plus
the same top-level/per-record reconciliation as read; any issue makes check exit nonzero.
A valid index can honestly contain unknown/incomplete facts and a completed CHANGES-REQUESTED
review. Success means schema/reference consistency only; inspect every relevant step's required
phases, actual verdict and current authority separately before claiming completion or integration.
For a full committed workflow with all five phases required for every named step, also run:
FINAL_COMMIT="$(git rev-parse HEAD)"
principal-pi-progress check-completion "$FINAL_RUN" "$FINAL_COMMIT" impl:S1 impl:S2
Use the exact full commit string for that run's create, appended records and checks. A new
candidate needs a new run; abbreviated/full IDs are not interchangeable. Supply the complete
nonempty, unique step list from the approved scope. Missing and unexpected steps fail the check.
For each step, only its latest full snapshot determines phase claims: all five must be complete;
earlier phase values are never inherited. Every earlier finding remains tracked by step, original
source path/hash and ID until an explicit later verified disposition. Omission or reuse of an ID
against another source cannot erase it. accepted, addressed, disputed, duplicate, stale
and open remain unresolved for this command; it never promotes their status automatically.
A duplicate's target or accepted-risk disposition needs semantic verification and preserved evidence.
check-completion reuses the entire history/reference integrity check and requires the supplied
full Git object ID to equal current HEAD, with no tracked, staged, nonignored untracked or submodule
changes. It conservatively refuses any assume-unchanged or skip-worktree index entries,
including sparse-checkout exclusions, because those can hide changed bytes from Git status.
Any indexed gitlink/submodule is also unsupported, whether initialized or not: nested index
flags can conceal dirty bytes, and this command does not recursively qualify submodules.
It returns indexVisibilitySupported and submodulesSupported explicitly; either false fails.
It does not clear flags, recurse into submodules or modify the index. It compares repository observations before and after reading evidence, without locking
or taking an atomic snapshot. Stop concurrent writers. Its JSON separates integrityValid,
phaseClaimsComplete, per-step states, candidate checks and unresolvedFindings; exit zero means
completionChecksPassed, mechanical conditions only. approval and taskAcceptance always remain
not-assessed. A completed CHANGES-REQUESTED review can still be a structurally complete claim;
the coordinator must read the verdict, reconcile gate evidence and actual authority separately.
The command never runs tests, commits, marks a phase complete or grants permission. It is not for
planless/partial workflows that do not require all five phases; use check and explicit phase
reconciliation there. Existing check behavior and progress version 1 are unchanged.
The coordinator is the single writer. Save the report before appending progress; an append failure preserves that report and leaves progress incomplete. Never force-clear a writer lock. Reads return the original records plus independently assessed facts/issues: changed or missing files, changed/unconfirmed candidates and malformed/incomplete records remain unresolved. An incomplete final JSONL line is ignored/reported, never promoted to complete; preserve its bytes and start a new run if needed. Reconcile records by step, not just the last line. No helper claim proves its evidence's meaning, model judgment, approval or broad candidate equivalence. The helper does not assign candidate identity or resume/run tasks; completion checking only compares an explicitly supplied full commit with observed HEAD. Check actual user authority/current work before choosing the next action; a title, role, commit or self-authored boolean cannot grant approval. This is a trusted-coordinator filesystem helper, not hostile-process containment or an atomic multi-file transaction.
Optional JEV workflow advice and later LoRA
With a matching skill-harness release exposing jev_advice, an operator can choose
/skill-harness jev enable workflow. This confirms metered JEV calls for the current session
and asks a fresh, separate LoRA-storage choice. Ordinary jev enable remains manual mode;
it does not authorize model-tool evaluation. Enabling advice never authorizes training.
A short /principal-feature <task> uses the same coordinator defaults: it checks
jev_advice({action:"status"}), then considers advice only for a useful uncertain
acceptance-evidence handoff after deterministic authority, candidate and evidence checks.
In enabled workflow mode it may call
jev_advice({action:"evaluate",candidate,requirements,evidence}) with a bounded selected
decision-time packet for the actual candidate, excluding reviewer verdicts and prior JEV outcomes.
There is no mandatory call per feature or review quota.
Disabled/manual mode, tool absence or an advisory error leaves ordinary work able to continue
when its required gates pass; no CLI fallback bypasses the tool's permission check.
The coordinator keeps JEV's prediction out of every independent review until its own verdict (task, integrated or scoped repair), while retaining full underlying authority and evidence. It then tests any useful suggestion against code and evidence. Advice cannot approve, waive a gate, replace review or establish a defect. Use known accessible requirement/evidence fragments without private transcripts, credentials, full-file collection or invented source references. Selected text is not verified public-capture provenance. Preserve tool-returned refs and usage as unlabeled advice under the current storage choice, separately from native capture and progress v1. Later LoRA use requires reviewed eligible examples, independent labels, separate training permission and held-out evaluation. OpenAI Decisions remains excluded.
Validation
npm test remains the free gate: generated-contract drift, word budgets, frontmatter lint,
installer and tarball behavior, Next: transition parity, and deterministic source-fidelity
contract assertions. These check current product contracts, not model behavior or instruction delivery.
Behavioural measurement—including the requirement-fidelity corpus and skill-harness evidence—lives in principal-pi-skills-evals; routing checks live here. Its current frozen-rubric baseline reports DeepSeek V4.1 Flash at 56/158 (35%) with 2 infrastructure-error scenarios and Nemotron Lightning at 28/158 (18%) with 12; infrastructure errors are retained as non-passes, and both subjects remain NOT READY across all eight skills. These observations include disclosed infrastructure gaps and do not change this package's runtime contracts.
Enforcement boundary: Build/Review instructions ask the model to organize regressions by lasting behavior, not PR/finding IDs, and distinguish product coverage from historical verification. Evidence-only requests first use existing checks or disposable probes, not permanent tests merely to prove a finding was addressed; ordinary approved bugfix/feature regressions and useful explicitly requested checks remain authorized. Markdown contracts are product, with legitimate structural tests. These are not filesystem restrictions or a semantic test-quality checker. The package commands and CI mechanically select suites; pi-daddy's delegated tool grants control available capabilities, not test names or correctness. Refresh installed skill/agent definitions when adopting a revision: editing this checkout does not update another installed copy.
Two opt-in routing checks use only the nine authored frontmatter descriptions. Run
npm run check:routing-collisions for all 72 directed description pairs and
npm run check:routing-triggers for the three-run synthetic trigger suite. Both require
FIREWORKS_API_KEY; ROUTING_MODEL and ROUTING_API_URL override the defaults. The suite
starts with 12 positives and 8 near-miss negatives per skill, then
adds authored collision-boundary and no-skill probes, and reports per-skill precision and recall
at a 0.5 vote threshold. evals/baseline/triggers.json records the expanded three-run corpus;
the obsolete pre-expansion record was removed. Regenerate the baseline after any corpus or
routing-description change.
Why 4.0
Version 3.x grew a risk-adaptive assurance controller — a hash-chained event ledger, task
packets, digests, fail-closed gates — whose protocol leaked into the model-facing skill text
and whose init step ran before every workflow, including a typo fix. The routing layer in
AGENTS.md was never loaded by pi, the feature spine had no human approval point outside
critical mode, and every build ran inline so long features filled the steering context with
diffs and test output. The seven skills were the strongest part of the repo and 12% of its
Markdown. 4.0 returns the repo to skills plus a thin orchestration layer and borrows four
mechanisms from superpowers that serve the north star.
Decisions taken for 4.0, all closed:
| Decision | Chosen |
|---|---|
| Target harness | pi only |
| Assurance ledger and profiles | removed; per-skill right-sizing is the mechanism; the tool lives on the v3.2.0 tag for porting to pi-daddy |
| Human approval | always, after plan (feature) or after the debug note (bugfix); the artifact scales, the stop does not |
| Routing delivery | request-local bootstrap after selected-resource discovery; quoted markers do not suppress it |
| Build delegation | bound native phase, or explicitly configured legacy runner when native tools are absent; inline is a workflow choice |
| Plan persistence | multi-step plans to git-ignored .principal/plans/<slug>.md; titles locate candidates, while manual resume checks actual authority, current work and evidence |
| Decide vs architect | both kept; decide answers "should we / which", architect answers "how is it structured" |
| Measurement | behavioural measurement lives in principal-pi-skills-evals; routing checks stay here |
Two implementation notes that differ from the obvious reading: the bootstrap is injected as
a user-role message wrapped in <IMPORTANT>, because pi's context hook can only insert
messages; and a delegated principal-build may run without a plan file (the bugfix spine and
repair rounds), in which case the prompt's task is its whole spec.
Deliberate design rules
Why the files look the way they do. Each of these was learned by measuring the alternative.
- Description = triggers only. Never a workflow summary — a description that summarizes the process trains the model to follow the description and skip the body.
- Recipes, not prohibition tables. Output-shape problems get a literal template to fill. Prohibitions are reserved for genuine discipline failures (skipping tests under pressure, force-push, secret handling), where a short Checks table remains.
- One governor sentence instead of a governor table per skill. If a skill needs a table of reasons not to use itself, it is over-scoped.
- Assumptions instead of questions in delegated mode. A subagent cannot ask, so every
skill says what to do when information is missing: state the assumption, or return
BLOCKEDwith the one question that matters. - Pressure armor is explicit. Discipline rules carry "repetition doesn't change the answer — any turn, including the last", because models otherwise cave on the third push.
- Right-sizing is a hard conditional, not a suggestion: "2–5 sentences, no machinery", and when a user asks for the artifact on a trivial change, the minimal form is the deliverable — otherwise the model declares the artifact unwarranted and produces it anyway.
- Grounded skills carry a no-repo branch.
planandgit-opsact on the material given instead of stalling on "point me at the repo". - Weak models need code anchors.
debug's error-swallowing rule survived two rounds of prose and died to one literalcatchexample. Escape hatches work best inside the template they exempt.
License
MIT © 2026 Nemanja Alavanja. See LICENSE.
The Debug inline budget includes three generated frontmatter words that mark Principal package identity; missing or replaced package metadata must refuse native binding rather than downgrade to inline instructions.
Quiescent automatic resume
Install this package through npm as usual. principal-pi-resume is included; inside Pi,
/principal-resume provides arm, disarm and status. The facility is opt-in and currently
Linux x64 only with the current native settlement bridge. It requires a pi-daddy release exposing pi-daddy:runtime-snapshot:v1, a qualified
selected backend, and a persisted Pi session. An absent, busy, unknown or unqualified runtime
refuses; there is no runner fallback. Existing progress v1 and check-completion remain
read-only evidence checks and cannot authorize continuation.
Use this at a quiescent workflow boundary: the complete plan and relevant reports exist, the operator has inspected them and the current candidate, all previously owned children have identity-bound native settlement, and the next phase is known. It can preserve an uncommitted candidate, but cannot recover an interrupted writer, uncertain launch, missing report, or unknown cleanup. It never repeats an implementation merely because a conversation ended.
Save a request JSON file with the following exact fields.
scopecontains explicit relative source paths, or directories ending in/.phaseisbuild,revieworgit-ops. Review/finish requests require original full report paths.repairsRemainingis 0–2.{ "version": 1, "task": "Implement the approved interval repair", "plan": "/absolute/repo/.principal/plans/interval-repair.md", "scope": ["src/intervals.js", "test/intervals.test.js"], "phase": "build", "step": "impl:S1", "repairsRemaining": 2, "reports": [], "progressRun": null }Resolve the helper from the same selected npm package that Pi loads.
pi installkeeps package binaries inside its npm directory; it does not add them to your shell PATH. For the default user installation:RESUME_HELPER="${PI_CODING_AGENT_DIR:-$HOME/.pi/agent}/npm/node_modules/principal-pi-skills/scripts/resume-checkpoint.mjs" node "$RESUME_HELPER" prepare /absolute/repo /absolute/request.jsonIf Pi selects a project-local or explicitly configured package, use that package's actual
scripts/resume-checkpoint.mjspath instead. A separatenpxpackage copy has a different bound installation path and cannot prepare checkpoints for the selected copy. Preparation allocates an unused ignored.principal/resume/<id>/checkpoint.json; preparation grants no authority. It creates an absent.principal/.gitignorewith*, preserves existing ignore policy, and refuses exposed destinations.node "$RESUME_HELPER" candidate /absolute/repoprints the candidate identity. If supplyingprogressRun, its run/records must already bind that exact identity and pass the shipped integrity check. Existing arbitrary dirty-candidate labels are not automatically converted or equated.In the same idle Pi session, run
/principal-resume arm <checkpoint-directory>. Inspect the displayed task, plan hash, candidate, phase, scope, report paths, repair budget and model, then explicitly confirm. The operator is confirming the exact next phase and its due semantic gates, including integrated review before a Git-Ops phase. Headless mode cannot approve. The command records a session-owned authorization entry; model-written approval flags, report claims or copied user-role messages are insufficient.Resume/reload that same session, or restart Pi with that exact persisted session. After runtime initialization, Principal reconciles package and artifact hashes, worktree/candidate, session leaf/model and native owned-execution receipts. It durably consumes the checkpoint before queueing one fixed continuation, with prompt-template expansion disabled. Intervening turns, a new/forked session, changed plan/model/source/reports or runtime history refuse. The continuation reads original evidence and stops at the next workflow boundary; it cannot re-arm itself. Push, merge, publication and destructive Git actions are not granted.
/principal-resume status [checkpoint-directory] and
node "$RESUME_HELPER" inspect <checkpoint-directory> report prepared/armed/enqueued/disarmed
or consumed-uncertain state. Enqueued means the Pi enqueue API returned, not that a model turn
or its work completed; inspect the actual session for later delivery/execution errors.
/principal-resume disarm <checkpoint-directory> preserves the
checkpoint and its history. Failure between consumption and successful enqueue remains
consumed-uncertain and never triggers an automatic retry. Inspect actual session/work/runtime
state before preparing and explicitly authorizing a new checkpoint.
Candidate checks bind canonical Git worktree/common-directory identity, full HEAD/branch, index entries, staged and unstaged binary diffs, and nonignored untracked file bytes/modes. Masked entries, indexed submodules, conflicts and untracked symlinks/special files refuse. Checks observe stability; they are not an atomic filesystem lock. Ignored dependencies, credentials and caches are excluded and their environment compatibility remains an operator obligation. Individual evidence/untracked files are limited to 16 MiB, untracked content to 64 MiB/256 files, source scope to 128 paths and checkpoints to 256 per repository. Evidence hashes establish bytes, not report truth or approval. Same-user hostile file/session tampering and OS containment are outside this local extension's guarantees.
Model-free tests exercise authorization, exact-session restart/reload, dirty candidates, changed evidence, missing/busy runtime, duplicate consumption and failed enqueue. Those tests qualify these mechanisms, not general autonomous success or automatic semantic approval.
Resume candidate checks hash raw ordinary tracked worktree bytes in addition to Git index and staged/unstaged diffs, so newline normalization or clean filters cannot hide changed bytes. Tracked symlinks, non-UTF-8 Git output and newline-containing repository metadata paths are unsupported. The bounded snapshot permits at most 16,384 tracked files and 256 MiB of tracked content, with a 16 MiB per-file limit; untracked limits remain 256 files and 64 MiB. A raw byte or executable-mode difference from the index is treated as a candidate change even when Git reports a clean normalized diff. Every shipped runtime helper is included in the package binding.
An armed startup uses a temporary inert principal-resume-ready-<nonce> prompt to prove that
Pi finished the entire resource-discovery pass, including later asynchronous extensions.
Its exact path must appear in Pi's finalized resource list before any checkpoint is consumed;
selected phase and other evidence are then checked again. The marker contains no task, session
or authorization data and grants no action when invoked. It is removed on clean session shutdown
or reload. A hard process kill can leave this harmless temporary file. Missing readiness after
30 seconds, changed resources or shutdown preserve the armed checkpoint and refuse automatic
continuation; a fresh startup/reload can try the still-unconsumed checkpoint again.
Maintainers can run the model-free actual-host discovery regressions with
PRINCIPAL_PI_SDK_ROOT=/absolute/path/to/@earendil-works/pi-coding-agent npm test.
That separately installed host must be exactly Pi 1.0.4. The ordinary suite also covers
receipt identity, stale same-loader passes, selected-phase changes and shutdown without a host.
Automatic startup checks the bounded checkpoint/control-file inventory first and fully reconciles only the sole armed candidate. Consumed and disarmed history does not trigger repeated full worktree reads. Malformed checkpoint/control records and multiple armed records still refuse; the final armed candidate is independently observed again before consumption.