@pi-vault/pi-dcp

Pi extension for dynamic context pruning — incremental tool output pruning and conversation compression

Packages

Package details

extension

Install @pi-vault/pi-dcp from npm and Pi will load the resources declared by the package manifest.

$ pi install npm:@pi-vault/pi-dcp
Package
@pi-vault/pi-dcp
Version
0.10.0
Published
Oct 6, 2026
Downloads
869/mo · 616/wk
Author
lanhhoang
License
MIT
Types
extension
Size
254.7 KB
Dependencies
0 dependencies · 4 peers
Pi manifest JSON
{
  "extensions": [
    "./src/index.ts"
  ]
}

Security note

Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.

README

@pi-vault/pi-dcp

npm version Quality Node >= 24.15.0 License: MIT

Keep long Pi sessions usable by pruning stale tool output, reporting what changed, and nudging the model to compress older context before the window fills up.

Install

pi install npm:@pi-vault/pi-dcp

Restart Pi after install.

To try a local checkout before publishing:

pi -e /absolute/path/to/pi-dcp

Quick Start

pi-dcp works out of the box — no configuration needed.

dcp:context
dcp:help
dcp:stats
dcp:sweep

Use dcp:context to see token usage and active DCP state, dcp:help to list commands, dcp:stats to check savings, and dcp:sweep to clear dead tool output before a heavy session.

What it does

  • Prunes automatically — deduplicates repeated tool outputs and purges stale failed tool inputs while preserving diagnostics.
  • Compresses with the model — exposes a compress tool in range or message mode while keeping tool-call/tool-result pairs intact.
  • Nudges before the window fills — context-limit, turn, and iteration nudges are anchored and frequency-throttled.
  • Shows operational feedback — pruning and compression can surface in toast or status notifications.
  • Lets you tune behavior — config, manual mode, runtime permission control, and schema-backed validation are all built in.

What's new in 0.10.0

  • Per-call ask permission — compression permission is now allow, ask, or deny. ask keeps the tool exposed and confirms every call inside the registered execute function, including direct executions. The dialog names the topic and the number of ranges or targets, and approvals are never reused.
  • Interactive dcp panel — one Pi-native TUI view shows the current model, context usage, resolved thresholds, compression and manual modes, permission, policy and tool availability, compression blocks, session savings, and lifetime totals. Arrow keys or j/k move, Enter runs the selected action, and Escape or q closes. Panel actions use the same guarded commands as the direct commands and recheck policy before acting.
  • Fail-closed confirmations — TUI and RPC can confirm; print/JSON sessions and any mode without dialog UI refuse the call. Rejection, cancellation, abort, or confirmation failure returns an error tool result and creates no compression timing, blocks, statistics, or DCP state entries.
  • Backward-readable snapshots — version 2 adds the full permission union, while readers keep accepting version 1 allow/deny snapshots unchanged.

What's new in 0.9.0

  • Lifecycle-aware tool reconciliation — DCP registers the compress tool during extension creation with defaultActive: false, then refreshes its mode-specific schema from trusted project configuration at session start. Pi's restored loadout is authoritative: a tool you deselected stays deselected across model switches, permission toggles, definition refreshes, resume, and reload. DCP remembers a temporary suppression within the current branch and restores the tool only when every policy reason clears.
  • Explicit fresh-session activation — on a brand-new session where the host allows the tool and the branch declares no tool loadout, DCP selects compress once. It never overrides tools, excludeTools, or noTools: "all", and it never reactivates the tool on resume, fork, tree navigation, or reload.
  • Structured system prompt section — DCP writes its instructions to the dcp section of Pi's structured system prompt instead of replacing the whole prompt. Other extensions' sections and any forced prompt are left untouched.
  • One prompt snapshot per run — custom prompt overrides reload once at the start of each agent run and are shared by the system section and every subsequent context pass until the next run. Edits made mid-run take effect on the next run.
  • Guidance follows the selection — compression instructions and nudges are hidden whenever compress is inactive or compression permission is denied. Automatic deduplication, stale-error pruning, reference assignment, and existing compression summaries keep working.
  • Defensive authorization — tool_call and the registered execute function reject every policy suppression reason with a clear message, even if the tool is invoked directly. Denied permission still allows automatic pruning.
  • dcp:compress does not activate the tool — the command reports that compression is unavailable when compress is inactive instead of silently selecting it.

What's new in 0.8.0

  • Compact message markers replace verbose XML tags — each injectable user/assistant message now ends with a standalone @mN@ line, or @mN:P@ when a compression priority is assigned. On a 2,000-message clean workload this cut metadata overhead from roughly 20,000 estimated tokens to about 4,000.
  • Every accepted tool input still resolves to the same message — compress accepts m1, m0001, @m1@, and @m1:3@ in range mode, and m2, m0002, @m2@, @m2:1@ in message mode, alongside b1-style block refs. Input is normalized to the stored canonical reference before lookup, so the accepted formats are interchangeable rather than merely parseable.
  • Stored and persisted references stay canonical — messageIds and version-1 snapshots continue to use the padded m0001 form. Nothing is migrated: a snapshot pair is validated as canonical and discarded if it is not, and nextRefIndex remains authoritative.
  • Safer marker sanitization — a marker is removed only when it occupies a whole line (complete or truncated, LF or CRLF, with optional indentation). Inline markers, person@m1@example.com, @mention, and marker text inside a sentence are left alone. Legacy XML marker cleanup remains active.
  • Benchmark token gates are release-blocking — deterministic token budgets for all three workloads are asserted in tests/benchmark.test.ts. Elapsed-time reporting stays informational.

Message markers and IDs

DCP shows the model a compact marker for each message it can compress:

Form Meaning
@m12@ Message 12, no priority assigned
@m12:3@ Message 12, compression priority 3
b3 Compression block 3

The marker is injected on its own line at the end of the message. Copy it exactly as shown when calling compress.

Accepted message references for startId / endId in range mode and messageId in message mode:

  • m1 and @m1@ — the bare or wrapped compact form
  • @m1:3@ — a compact form carrying a priority
  • m0001 — the canonical padded form

Range boundaries also accept b2-style compression block anchors.

Priorities run from 1 to 5: 1-2 for the highest compression value, 3 moderate, 4-5 low.

Compact input is exact, case-sensitive, and whitespace-free. m01, @m0001@, @m1@ , and m1:3 are rejected. What DCP stores and persists is always the padded canonical m0001; the compact form exists only so the model-facing text stays short.

What's new in 0.7.0

  • Pi-native file paths are protected — read, write, and edit tool arguments now use Pi's path field (and legacy filePath), with Windows separators normalized before glob matching. Nested calls made by tools such as codemode are inspected too, so a parent tool result is protected when any direct or nested file path matches protectedFilePatterns.
  • Unsafe context limits are rejected — maxContextLimit / minContextLimit and their per-model entries accept only positive integers or percentages greater than 0 and at most 100. Invalid fields are dropped with a warning instead of aborting startup.
  • Configuration layers fall back independently — an invalid global field keeps the built-in default, and an invalid project field inherits the valid global value rather than resetting it. Unknown keys and invalid values are reported with a source-qualified JSON pointer.
  • Configuration problems are visible once per reload — each problem is written to the DCP log even when debug logging is off, and an interactive session shows exactly one warning notification with the problem count and the participating config paths.
  • nudgeForce selects the role — strong injects the turn nudge into the user message; soft injects it into the assistant message, including a synthetic text part prepended before a tool-only assistant call. Both halves of an eligible user/assistant pair are anchored without changing the version-1 snapshot shape, and older user-only version-1 anchors are upgraded when their messages reappear.

Commands

All commands are also discoverable in-session via dcp:help.

Command Purpose
dcp Open the interactive DCP status panel
dcp:help List all available commands
dcp:context Show context usage and DCP state
dcp:stats Show compression and token savings statistics
dcp:sweep Force-prune all eligible tool outputs
dcp:manual on Pause automatic compression
dcp:manual off Resume automatic compression
dcp:decompress <blockId> Deactivate a compression block
dcp:recompress <blockId> Reactivate a compression block
dcp:lifetime Show aggregate statistics across saved sessions
dcp:permission Cycle compress permission (allow/ask/deny)
dcp:compress [focus] Ask Pi to run compression on stale context

Permission and the DCP panel

Compression permission is exactly allow, ask, or deny and defaults to allow. allow exposes and runs compress normally. ask exposes it and confirms each call inside the registered execute function, including direct executions; the dialog includes the topic and the number of ranges or targets. deny hides the tool, removes its prompt section and nudges, and blocks defensive direct calls. dcp:permission cycles allow -> ask -> deny -> allow, using the configured permission as the fallback when the session has no override.

TUI and RPC sessions can confirm an ask call. Print/JSON sessions and any mode without dialog UI fail closed with Compression requires interactive approval; a rejected, cancelled, aborted, or failed confirmation returns Compression was not approved. Waiting for approval or refusing a call creates no compression timing, blocks, statistics, or DCP state entries.

Run dcp to open the single status panel in TUI mode. It shows the current model, context usage, and resolved max/min thresholds; compression mode, manual mode, permission, pipeline policy, and whether compress is active; session savings and lifetime totals; and every compression block, sorted numerically, with its mode, tokens, and whether it can be deactivated or reactivated.

Use the arrow keys or j/k to move, Enter to run the selected action, and Escape or q to close. The list scrolls to keep the selection visible and stays bounded at 60, 80, and 120 columns. Policy-disabled actions cannot mutate state; globally disabled, model-disabled, and disallowed sub-agent sessions keep informational access only. Permission denial alone does not disable sweep, manual mode, or block controls. Existing dcp:* commands remain stable shortcuts and the fallback when no TUI is available.

Typical workflows

Default: install it and let DCP prune duplicates and stale failed inputs automatically.

Need a cleanup pass first? Run dcp:sweep, then dcp:context.

Want manual compression control? Use dcp:manual on, compress selectively, then dcp:manual off.

Need to block compression temporarily? Run dcp:permission to cycle between allow, ask, and deny.

Need to approve each compression? Cycle dcp:permission until it reports ask. Every compress call then shows a confirmation in TUI or RPC; print/JSON sessions refuse the call instead. Rejection, cancellation, or abort leaves context unchanged.

Want everything in one place? Run dcp for the interactive panel, then act on manual mode, permission, sweep, and individual blocks without leaving it.

Need compression now? Run dcp:compress [focus]. It sends Pi a hidden follow-up that asks it to use the compress tool; it does nothing while DCP or compression permission is disabled, or while the compress tool is inactive.

Need to undo a compression block? Use dcp:decompress <blockId> and dcp:recompress <blockId>.

Need lifetime totals? Use dcp:lifetime to see aggregate savings across saved sessions.

Customize prompts (experimental). Enable experimental.customPrompts, then edit prompt overrides in either trusted project or global locations:

  • Trusted project: .pi/dcp-prompts/overrides/<file>.md
  • Global: ~/.pi/agent/extensions/dcp-prompts/overrides/<file>.md

Files: system.md, context-limit-nudge.md, turn-nudge.md, iteration-nudge.md.

Configuration

Create <agentDir>/extensions/dcp.json (normally ~/.pi/agent/extensions/dcp.json) to override defaults. On each session start, DCP merges built-in defaults, this global file, and <ctx.cwd>/.pi/dcp.json when Pi marks the project trusted. Nested objects merge recursively; arrays replace earlier arrays. Untrusted project configuration is ignored, and a previously registered compression tool safely reports that DCP is disabled after a later disable.

Configuration is sanitized field-by-field. Invalid, unsafe, or unknown fields are dropped with a warning that names the source file and JSON pointer; valid siblings are retained, and an invalid project value inherits the valid global value instead of resetting it. The shipped schema rejects unknown properties in declared configuration objects while keeping per-model maps open. Startup never aborts because of configuration. Each reload writes every problem to the DCP log even when debug is false and, in an interactive session, shows a single warning notification containing the problem count and the participating config paths.

You can also use the shipped dcp.schema.json for editor tooling or config validation workflows.

{
  "enabled": true,
  "disabledModels": [],
  "debug": false,
  "nudgeNotification": "minimal",
  "nudgeNotificationType": "status",
  "protectedFilePatterns": [],
  "turnProtection": 0,
  "compress": {
    "mode": "range",
    "permission": "allow",
    "showCompression": false,
    "maxContextPercent": 80,
    "minContextPercent": 50,
    "maxContextLimit": 200000,
    "minContextLimit": 100000,
    "modelMaxLimits": {},
    "modelMinLimits": {},
    "nudgeFrequency": 5,
    "iterationNudgeThreshold": 15,
    "nudgeForce": "soft",
    "protectedTools": ["compress"],
    "protectUserMessages": false,
    "protectTags": false,
    "summaryBuffer": true
  },
  "manualMode": {
    "default": false,
    "automaticStrategies": true
  },
  "strategies": {
    "deduplication": {
      "enabled": true,
      "protectedTools": [],
      "turnProtection": 0
    },
    "purgeErrors": {
      "enabled": true,
      "turns": 4,
      "protectedTools": []
    }
  },
  "experimental": {
    "allowSubAgents": false,
    "customPrompts": false
  }
}

Top-level

  • enabled — set to false to disable the extension entirely without uninstalling.
  • disabledModels — exact, case-sensitive provider/modelId keys for which DCP processing, mutating commands, and the active compress tool are disabled.
  • debug — when true, writes operational per-session logs to {sessionDir}/dcp/logs/YYYY-MM-DD.log. Configuration warnings are always written there so headless sessions retain diagnostics.
  • nudgeNotification — notification verbosity: "off", "minimal", or "detailed".
  • nudgeNotificationType — notification delivery: "toast" or "status".
  • protectedFilePatterns — file-path globs whose related tool outputs should never be pruned. Direct arguments from Pi's read, write, and edit tools (path, plus legacy filePath on any tool) and nested calls recorded on a tool result are both checked; candidate paths are normalized to / separators before matching.

Protected tool and file patterns use Node's path.posix.matchesGlob semantics: / is the path separator, and supported patterns include *, **, ?, and character classes such as [abc] and [0-9]. Wildcards continue to match leading-dot path segments for compatibility with earlier pi-dcp releases.

  • turnProtection — hard-protect the newest N raw user-message turns from every DCP transformation; defaults to 0.
{
  "disabledModels": ["openai-codex/gpt-5.6-sol"],
  "compress": {
    "modelMaxLimits": {
      "openai-codex/gpt-5.6-sol": "80%",
      "openai-codex/gpt-5.6-terra": "60%"
    },
    "modelMinLimits": {
      "openai-codex/gpt-5.6-sol": "50%",
      "openai-codex/gpt-5.6-terra": "40%"
    }
  }
}

For a session using openai-codex/gpt-5.6-sol, DCP leaves messages unchanged, rejects mutating DCP commands, and removes compress from the active tools. The configured sol thresholds remain dormant while that model is disabled. The independent terra thresholds remain active for sessions using openai-codex/gpt-5.6-terra. Changing models during a live session immediately removes or restores compress according to disabledModels. Existing DCP state remains intact while the selected model is disabled and is available again after switching to an enabled model.

compress

  • mode — compression mode: "range" or "message".
  • permission — compression permission: "allow", "ask", or "deny". allow runs normally, ask keeps the tool exposed and confirms each call, and deny hides the tool and blocks direct calls. dcp:permission cycles the in-session value.
  • showCompression — when true, detailed notifications include the compression summary text.
  • maxContextPercent / minContextPercent — legacy percentage thresholds.
  • maxContextLimit / minContextLimit — accept a positive integer token count or a percentage string greater than 0 and at most 100 (for example "80%"). Anything else is dropped with a warning.
  • modelMaxLimits / modelMinLimits — per-model overrides keyed by provider/modelId; each value follows the same context-limit rules as the global limits.
  • nudgeFrequency — minimum messages between non-urgent nudges.
  • iterationNudgeThreshold — assistant iterations without user input before an iteration nudge fires.
  • nudgeForce — nudge strength. "strong" injects the turn nudge into the user message; "soft" injects it into the assistant message.
  • protectedTools — Node glob patterns for tool outputs preserved during compression.
  • protectUserMessages — append user message text to compression summaries.
  • protectTags — preserve <protect>...</protect> tag content in summaries.
  • summaryBuffer — exclude active summary tokens from threshold comparison to prevent cascading compressions.

manualMode

  • default — start in automatic mode (false) or manual mode ("active").
  • automaticStrategies — continue running automatic pruning strategies while manual compression mode is active.

strategies

  • deduplication.enabled — enable or disable deduplication.
  • deduplication.protectedTools — Node glob patterns for tool names excluded from deduplication.
  • deduplication.turnProtection — legacy deduplication window; deduplication uses the larger of this and top-level turnProtection.
  • purgeErrors.enabled — enable or disable stale failed-input purging.
  • purgeErrors.turns — age threshold for failed tool-input purging.
  • purgeErrors.protectedTools — Node glob patterns for tool names excluded from failed-input purging.

DCP counts turns from raw user messages, not assistant iterations. When the history contains fewer user turns than turnProtection, all existing user turns are protected. Deduplication, stale-error pruning, dcp:sweep, and both compression modes enforce this boundary. Normal compression expands a tool target to its complete assistant call/result group; DCP removes orphan results it creates, while Pi synthesizes error results for assistant calls that have no result.

DCP preserves failed tool diagnostics and purges only the historical arguments of eligible stale failures. Repeated read, grep, find, ls, and bash calls may be deduplicated or swept. compress, write, edit, and subagent remain protected by default; configured protected-tool patterns are additive.

experimental

  • allowSubAgents — run DCP inside sub-agent child sessions.
  • customPrompts — load prompt overrides from the filesystem.

Development and verification

pnpm install
pnpm check
pnpm release:check

Benchmarks

pnpm benchmark runs three deterministic workloads through the production DCP pipeline and writes one JSON report to stdout. The retained benchmarks/result.json was recorded on Node 24.21.0 with 30 timed iterations per workload, after one untimed warm-up.

Each timed iteration includes cloning the fixture, creating fresh session state, cloning the default configuration, restoring persisted state when applicable, running the pipeline, projecting the transformed messages, and estimating input and output tokens.

Workload What it exercises Median p95 Input tokens Output tokens Reduction
clean-2000-messages Baseline pipeline processing for alternating user/assistant messages 11.29 ms 14.18 ms 9,000 12,992 -3,992
repeated-tool-pairs-2000 Deduplication, stale-error input purging, and protected write results 41.40 ms 54.02 ms 1,017,575 53,977 963,598
restored-nested-blocks-100 Snapshot restoration and relationship rebuilding for nested blocks 3.03 ms 3.95 ms 1,291 120 1,171

The clean workload intentionally reports a negative reduction: no content is pruned, while DCP adds compact message markers to all 2,000 messages. It is a baseline for pipeline and metadata overhead, not a token-savings case. Compact markers reduced that overhead from -20,000 to -3,992 estimated tokens.

The repeated-tool workload reduces the estimate by 94.7%. It models 2,000 assistant/tool-result pairs with repeated reads, stale failures, and unique writes. Production strategies replace superseded read output and stale failed-call arguments while preserving protected write output, error diagnostics, and complete tool-call ownership.

The restored-nesting workload reduces the estimate by 90.7%. It restores 100 persisted compression blocks arranged as ten nested chains, rebuilds their runtime relationships from real compress call/result owners, and leaves the ten outer blocks active.

Enforced token gates. tests/benchmark.test.ts asserts three deterministic budgets, so a token regression fails pnpm check rather than waiting for a human to read a report:

  • clean-2000-messages — marker overhead (output - input) must stay at or below 6,000 tokens.
  • repeated-tool-pairs-2000 — the reduction ratio must stay within 5 percentage points of the retained 962873 / 1017575 baseline.
  • restored-nested-blocks-100 — the reduction ratio must stay within 5 percentage points of the retained 1171 / 1291 baseline.

The ratios are compared as exact fractions rather than rounded percentages, so the tolerance is precisely five percentage points.

Informational timing. medianMs and p95Ms are reported for comparison only and are not gated. Elapsed times vary with hardware and system load; compare timing reports only from the same machine and Node version.

Report fields:

  • nodeVersion and iterations describe the runtime and timed sample count.
  • medianMs is the middle elapsed time; p95Ms represents the slower tail.
  • inputEstimatedTokens and outputEstimatedTokens use DCP's lightweight character-based estimator, not provider billing tokens.
  • reductionEstimatedTokens is exactly input minus output, so it may be negative.

To refresh the retained evidence, capture the report and update benchmarks/result.json with it:

pnpm benchmark > benchmarks/result.json

Changelog

See CHANGELOG.md for release notes.

License

MIT — see LICENSE.