pi-495

Evidence-gated harness for coding agents, inside Pi. The checks are frozen before the change exists; it is accepted on evidence, never on the agent's word.

Packages

Package details

extension

Install pi-495 from npm and Pi will load the resources declared by the package manifest.

$ pi install npm:pi-495
Package
pi-495
Version
0.4.0
Published
Oct 7, 2026
Downloads
518/mo · 45/wk
Author
jeanjerome
License
Apache-2.0
Types
extension
Size
3.6 MB
Dependencies
1 dependency · 4 peers
Pi manifest JSON
{
  "image": "https://raw.githubusercontent.com/jeanjerome/pi-495/main/.github/assets/banner.jpg",
  "video": "https://github.com/jeanjerome/pi-495/raw/refs/heads/main/.github/assets/demo.mp4",
  "extensions": [
    "./dist/extension/index.js"
  ]
}

Security note

Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.

README

Pi-495 is a Pi extension that turns a software request into a change you can inspect, verify and integrate. It coordinates the AI work, defines how success will be checked before implementation, and runs those checks on the result. You review the changes and resolve decisions that need human judgment.

Use it to add behavior, fix a bug or refactor a supported project with an explicit record of what was requested, what changed and why it was accepted.

/495 start add a retry with backoff to the upload client

The core rule: the checks are frozen before implementation. The coding agent cannot change the protected tests or decide that its own work is accepted.

Available today: an early implementation for macOS on Apple Silicon and for Linux (tested on arm64), with Java/Maven and Node adapters. Models come from your Pi configuration. Start with a small change on a project whose tests already run locally.

  • Pi is the host. Pi is a minimal terminal coding harness that packages extend without forking it. It brings what 495 does not rebuild: the models you configure and authenticate, the agent session each intervention runs in, the terminal dialogues where you answer decisions, and the TUI, RPC, JSON and print modes. 495 is a Pi package, with no CLI or service of its own.
  • 495 is the three-digit Kaprekar constant: repeated application of a simple rule reaches a fixed point for eligible starting numbers. The name reflects the project's use of explicit rules and feedback to drive a change toward acceptance. Explore the number.

Quick start

You need Node.js 24+, Pi 1.0.4 or later, Git, and either macOS on Apple Silicon or Linux (tested on arm64) with bubblewrap (bwrap) on PATH, able to create its namespaces without privilege. Your target project must have at least one Git commit and use a supported test setup. Install its dependencies before starting: verification runs with restricted network access.

1. Install the extension

From npm:

pi install npm:pi-495

Pi installs the package and its dependencies under its own npm directory. pi update npm:pi-495 moves it to the latest release; install npm:pi-495@0.4.0 instead to pin a version.

From the Git repository, in a directory where you keep your tools:

git clone https://github.com/jeanjerome/pi-495.git
cd pi-495
npm ci
npm run build
pi install "$PWD"

Pi registers this local directory. Keep it in place; after updating the source, reinstall dependencies and rebuild it.

Install from one source only, so that a single copy of the extension loads.

2. Start a change in your project

cd /path/to/your-project
pi

Configure and authenticate your model in Pi, select it with /model, then enter:

/495 start add a retry with backoff to the upload client

495 advances through Scoping, Specification, Qualification, Design, Implementation and Acceptance. It stops when it needs a decision or cannot proceed, and records the reason.

Next action Command
See progress and the next step /495 status
Answer a pending decision /495 decide
Inspect the candidate and its changes /495 review
Read the findings and remaining risks /495 report
Continue from the recorded state /495 resume

To survey the project instead of changing it, enter /495 state followed by a question about the project, for example /495 state does the code meet its quality standards?. 495 runs the checks on the project as it stands, writes no candidate, and asks you to accept or refuse the survey.

Integration is disabled by default. To allow it, set policy.integration_enabled to true before starting a new session. After acceptance, /495 integrate requests local integration and any required authorization. 495 does not push to a remote repository.

Features

Capability What it gives you today
📝 Spec-driven development A prose request becomes explicit requirements with observable acceptance criteria and links to verification controls.
🧪 Enforced test-first workflow When new behavior needs tests, a separate preparation step writes them. They must run and fail on the reference before adoption; implementation cannot edit protected tests.
🚦 Evidence gates Controls must first demonstrate a passing case, a failing case and a tool failure. The kernel computes acceptance from recorded evidence and the configured review policy.
🧩 Context engineering Each intervention receives role-specific instructions, adopted artifacts, an explicit output schema and bounded feedback. Context manifests are recorded for inspection.
⚖️ Regression-aware verification Compare the candidate with the reference to distinguish introduced regressions from inherited findings, under a frozen tolerance policy.
🧬 Coverage, mutation testing and architecture checks Measure the coverage of the lines a change introduces and the mutants that survive on them: on Maven with JaCoCo and PIT, on Node with the runner's LCOV report and, under node --test or vitest, Stryker. On Maven, also check declared Java import boundaries.
🔎 Project surveys /495 state runs the checks on the project as it stands, without changing it. The survey gives each requirement its verdicts, each finding its file, and each unmeasured area a named blind spot; you accept or refuse it.
📏 Quality referentials and trajectories When a survey asks about code quality that no check measures, 495 offers a referential for complexity, duplication and dead code: PMD and CPD on Maven, ESLint and jscpd on Node. Adopted, it is installed in a copy, never in your project. A trajectory you write turns the measured gaps into increments, and /495 measure judges each milestone on a new survey.
🛡️ Sandboxed execution Work happens in isolated copies, with phase-specific permissions, confined by Seatbelt on macOS and bubblewrap on Linux. An unavailable required isolation capability blocks execution.
👀 Human-in-the-loop review Inspect the file tree in the terminal with each change drawn beside it as Pi draws its own edits, using your Pi keybindings, and record human decisions with their origin.
⏯️ Resumable workflows Pause and resume a change. Bound attempts, intervention duration and tool calls; eligible interrupted producers continue on their existing workspace.
🔗 Verifiable audit trail Keep artifacts, evidence and hash-chained events. Export a dossier with an offline integrity verifier that runs with Node alone.
🤖 Local and hosted models Use the model selected in Pi, subject to authentication and capability checks. No silent model substitution.
🖥️ Pi-native surfaces TUI, RPC, JSON and print expose the same underlying change state. Human actions depend on the surface's ability to supply a human origin. English and French are available.

How it works

Define success → freeze the checks → produce a candidate → judge the evidence.

flowchart TD
    S["`**Scoping**
the objective and the questions it raises`"] --> SP["`**Specification**
requirements with observable criteria`"]
    SP --> Q["`**Qualification**
checks qualified, verification protocol frozen`"]
    Q --> D["`**Design**
the plan within the mandate`"]
    D --> I["`**Implementation**
a candidate in an isolated workspace`"]
    I --> A{"`**Acceptance**
does the evidence satisfy the protocol?`"}
    A -->|"Yes"| G["`**Integration**
the accepted candidate applied locally, when authorized`"]
    A -->|"Correction within budget"| I
    A -->|"Decision or capability missing"| X["Stop and record the next action"]

Agents propose artifacts and code. The kernel runs the controls and applies the acceptance rules. Required agent reviews and human decisions are additional inputs when the protocol calls for them. /495 verify reruns the frozen controls without calling a model.

Gate What it checks
G0 — Scoping The objective is explicit and material questions have answers.
G1 — Specification Requirements have observable criteria and preserve the human decisions that bind them.
G2 — Qualification Mandatory requirements have qualified controls or assigned human decisions; the protocol is frozen.
G3 — Design The design addresses the mandatory requirements within the mandate.
G4 — Implementation The candidate is complete, within scope and preserves protected paths.
G5 — Acceptance The applicable evidence, required reviews and human acceptance satisfy the protocol.
G6 — Integration The locally integrated tree matches the verified candidate.

Acceptance establishes conformance to the adopted protocol, within the limits of its controls and reviews.

Compatibility

Area Current scope
Host Pi 1.0.4 or later (qualified on 1.0.4); Node.js 24 or later. 495 is a Pi package with no standalone CLI or service.
Platform macOS on Apple Silicon, confined by Seatbelt. Linux, confined by bubblewrap (bwrap on PATH), tested on arm64; where bwrap cannot create its namespaces without privilege, confined work is refused and the refusal says why. Windows is not supported. Under WSL 2, which runs a Linux kernel, the Linux support should apply where bwrap can create its namespaces; this is not tested.
Java / Maven Tests run by Surefire. Coverage of introduced lines (JaCoCo) and mutation testing (PIT) when the project declares the required reports; a project without JaCoCo is recommended it, and once you adopt it, its plugin declaration is inserted into the candidate's pom.xml. Structural checks derived from supported Maven and Java declarations. A survey can adopt a quality referential for complexity, duplication and dead code (PMD and CPD), and tells generated code apart by its @Generated annotations. Maven verification uses offline mode; resolving an adopted plugin opens the network for that step alone.
Node Tests run by scripts.test with node --test (or absent), or with vitest or vitest run, mocha or jest without any other argument, each read through the report it writes (JUnit for vitest and mocha, JSON for jest), and a detected lint script. A runner given arguments, and any other runner in scripts.test, is refused, and the refusal names it. Coverage of the lines a change introduces, when the target asks for it (--experimental-test-coverage in a node --test script, or the coverage provider of vitest installed), read from the LCOV report of the run; otherwise coverage is not measured and the report says so. Mutation testing of the lines a change introduces, under node --test or vitest, when the target installed Stryker (@stryker-mutator/core), read from the JSON report of a Stryker run scoped to those lines: a mutant that survives on a line the change wrote blocks it. Without Stryker, mutation is not measured, the report says so and 495 recommends installing it; under mocha or jest, mutation is not measured. A survey of a target locked by package-lock.json can adopt a quality referential for complexity and dead code (ESLint) and duplication (jscpd); TypeScript and JSX sources are not measured for complexity and dead code.
Other languages Additional target adapters are required. The kernel and report contracts provide the extension boundary.
Models Models configured and authenticated in Pi, including local OpenAI-compatible endpoints with working tool calls. Provider and subscription availability follow Pi and the provider.

A recorded qualification case reached acceptance with one local model and one hosted model. It covers a small contract case; the record also notes that the two runs used different harness builds. It is not a qualification of every model or project shape.

  • Selecting a model in Pi admits its provider; 495 keeps no list of destinations of its own. The model worker needs network access; tools and controls have their own confinement profiles.
  • Model qualification can issue a small tool-call probe, cached per provider/model pair for the session. A hosted provider may bill that request.
  • Budget controls currently cover attempts, time, continuations and tool calls; they are not a monetary spending cap.
  • export --redact masks recognized secret patterns. Review a dossier before sharing it; pattern matching does not guarantee that every sensitive value was removed.
  • The offline verifier checks dossier integrity and event-chain consistency. It does not rerun the target's checks or prove that every requirement is correct.

Roadmap

The direction is a broader engineering workflow: understand an existing codebase, strengthen its verification, and guide changes with traceable decisions. These capabilities are planned or deferred; they are not all available in the current release.

Direction Planned capabilities
📏 Code quality assessment Extend the quality referentials beyond their first analysers: Checkstyle and SpotBugs on Maven; complexity and dead code in TypeScript sources, and cognitive complexity, on Node.
🏗️ Architecture assessment and migration Diagnose the current architecture, justify a target and plan reversible migration steps.
🔍 Brownfield verification Audit what existing tests assert, add characterization tests and strengthen verification as increments progress.
📚 Documentation grounding Retrieve version-matched sources with provenance, then validate important API uses through compilation, examples or contract tests.
🧩 Adaptive context engineering Assemble versioned skills, prompt templates and documentation by role, phase and model capability.
🧭 Risk-guided engineering Route relevant risks to rules, experiments, specialist reviews or human decisions; record consequential choices and evaluate outcomes.
🛠️ Workflow and onboarding Improve recovery from incomplete verification.

Further qualification and deferred work: additional advanced testing methods such as property-based testing and fuzzing, and indexed documentation retrieval with corpus maintenance. These require further work and qualification; no release date is implied.

See the plan of the open work for the detailed boundaries and progress.

Reference

Command Purpose
/495 start <request> Capture the project and drive a new change.
/495 state <question> Survey the project without changing it: run the checks on it as it stands, then accept or refuse the survey.
/495 adopt <trajectory.json> Adopt a trajectory of several increments you wrote, and drive the change of its first ready increment.
/495 next Drive the change of the next ready increment of the bound program.
/495 measure <change_id> Judge each milestone of the bound program on the accepted survey of the integrated project that change took.
/495 status Show phase, gates, attempts, evidence and next action.
/495 resume Resume from the recorded state.
/495 decide Present and answer pending human decisions.
/495 review [path|cand_id] Inspect the candidate in the TUI, or obtain a text summary on other surfaces; cand_id picks the candidate everywhere, path the file of the text summary.
/495 report Read observations, judgments and residual risks.
/495 verify Rerun the frozen controls on the frozen candidate.
/495 integrate Request authorized local integration after acceptance.
/495 export [--redact] Export the dossier and its integrity verifier.
/495 pause Pause the current work.
/495 close <question> Declare a material question no longer material when the specification stops without its answer.
/495 revoke <question> Revoke your answer to a material question, or its close, until the candidate is accepted: the question is asked again and nothing adopted since stays adopted.
/495 cancel [reason] Cancel the change while retaining its dossier.
/495 bind [change_id] List open changes or bind the session to one.
/495 unbind Release the session binding.
/495 help List the commands and the version of 495.

The conversational harness495 tool can request status, pending decisions, review summaries, reports, verification, export and start. It cannot adopt artifacts, make human decisions or integrate a change.

Configuration and state live outside your project: $HARNESS495_DATA_DIR, otherwise ~/.495. Workspaces default to ~/.495/workspaces; set $HARNESS495_WORKSPACES_DIR to move them.

config.json is optional, and so is each setting in it. The file below holds every behavioural setting at its default value; policy.policy_id and policy.revision only identify the policy. Durations are in milliseconds.

{
  "policy": {
    "budgets": {
      "max_attempts": 3,
      "max_technical_retries": 2,
      "max_continuations": 3,
      "intervention_ms": 1200000,
      "increment_ms": 7200000,
      "tool_calls_per_intervention": 100,
      "feedback_bytes": 65536
    },
    "adoption": {
      "mandate": "kernel",
      "requirements": "kernel",
      "protocol": "kernel",
      "design": "kernel"
    },
    "g5_human_acceptance": false,
    "integration_enabled": false,
    "baseline": {
      "compare_to_reference": true,
      "tolerance": "no_aggravation",
      "instability": "confirm_then_indeterminate",
      "max_confirmations": 1
    },
    "stagnation_identical_candidates": 2,
    "required_reviews": []
  },
  "isolation": { "allow_unconfined": false },
  "human_origin": { "rpc_actor_env": "HARNESS495_RPC_HUMAN_ACTOR" },
  "workspace_exclusions": ["target/", "dist/", ".pi/", "__pycache__/", "build/", "node_modules/.vite/", "node_modules/.vite-temp/", "node_modules/.vitest/", "reports/mutation/", ".stryker-tmp/"],
  "language": "fr"
}
Setting Purpose
policy.budgets Attempts, retries, continuations, time and tool-call limits, feedback size.
policy.adoption Who adopts the mandate, requirements and design: kernel or human. The protocol stays with the kernel.
policy.g5_human_acceptance Require a human acceptance decision.
policy.integration_enabled Permit local integration, still subject to authorization.
policy.baseline How each control is compared with the reference.
policy.stagnation_identical_candidates Stop after this many identical candidates in a row; 0 disables it.
policy.required_reviews Review roles required for acceptance.
isolation.allow_unconfined Run checks without a sandbox. Leave false.
human_origin.rpc_actor_env Environment variable through which an RPC host names its human actor.
workspace_exclusions Build outputs left out of workspaces. Keep the inputs the checks need.
language fr or en.

config.json must match its schema. A file that does not, that is not valid JSON, that cannot be opened or that is not a regular file stops every /495 command, /495 help included, and the error says what is wrong: for a file that does not match, where its first three deviations lie and what is expected there, and how many others there are. The defaults never replace it, since one of its settings may keep a decision for a human. Fix the file, then reload Pi (/reload) or start a new session.

HARNESS495_INTEGRATION=1, HARNESS495_HUMAN_ACCEPTANCE=1 and HARNESS495_LANGUAGE=en override the file. Set them before starting Pi.

Contributing

Early feedback is especially useful on real Maven and Node projects: the request, the command used, the observed stop reason and the expected behavior help make a report actionable. Ask questions and suggest ideas in Discussions; report defects with a bug report.

CONTRIBUTING.md explains how to build, check and propose a change. Report a vulnerability privately, as the security policy describes. Everyone who takes part follows the code of conduct.

License

Apache License 2.0 © 2026 Jean-Jerome Levy.