pi-fast-compaction

Fast compaction for Pi: Jev-scored verbatim eviction of stale tool history, with LLM summarisation of the evicted transcript as the fallback.

Packages

Package details

extension

Install pi-fast-compaction from npm and Pi will load the resources declared by the package manifest.

$ pi install npm:pi-fast-compaction
Package
pi-fast-compaction
Version
0.1.0
Published
Sep 18, 2026
Downloads
148/mo · 148/wk
Author
tvdavies
License
MIT
Types
extension
Size
60.5 KB
Dependencies
0 dependencies · 1 peer
Pi manifest JSON
{
  "extensions": [
    "./extensions/fast-compaction.ts"
  ]
}

Security note

Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.

README

pi-fast-compaction

Fast context compaction for Pi.

When Pi's context fills up it normally asks the session model to summarise the old history. That is slow (minutes on a large context with thinking enabled), costs a full uncached read of the conversation, and loses detail.

This extension takes a different first step. It asks TypeSafe's Jev — a small, fast decision model — which tool results in the old history are still needed, drops the rest, and keeps what survives verbatim. Only when that does not free enough space does it fall back to an LLM summary, and then it summarises the already-pruned transcript on a model of your choosing.

Typical numbers from a 400k-token span: Jev scored 167 tool calls in ~1.5 s for $0.002, evicted 87% of the span, and the whole compaction finished in under 3 s with no LLM call. The native path on the same span took 5–10 minutes.

How it works

On session_before_compact (threshold, overflow, or /compact):

  1. Score. The span Pi is about to replace is rendered as a compact state (user/assistant text, every tool call with its arguments, a short head of each result and its size) and sent to Jev with one question per tool call: is the full output of this call still needed to continue the task? Results scoring below keepThreshold are truncated to a short head plus a note; the call and its arguments always stay. Long write/edit bodies are clipped too — the files are on disk and listed under <modified-files>.
  2. Fit. If the surviving history still exceeds maxVerbatimTokens, the lowest-scored survivors are evicted until it fits.
  3. Decide.
    • Fast path — reduction ≥ minReduction: the surviving history is stored verbatim as the compaction summary, in Pi's own [User]: / [Assistant]: / [Tool result]: format. No LLM call.
    • Hybrid path — otherwise: Pi's summariser runs over the pruned transcript on the configured summariser model. Faster and cheaper than summarising the original, and the summary still sees the full shape of the conversation.
    • Summary path — Jev unavailable (no key, timeout, daily budget hit): the original transcript is summarised on the configured summariser model.
    • Default — no summariser model configured and Jev unavailable, or the summariser fails: the extension returns nothing and Pi compacts as usual.

Each compaction entry records which path was taken and the eviction stats in details.fastCompaction.

Install

pi install npm:pi-fast-compaction

Set a Jev API key in your environment as JEV_API_KEY (or TYPESAFE_API_KEY). Get one at https://typesafe.ai. Jev is priced per input token; a compaction costs a fraction of a cent, and the extension keeps a daily spend ledger (maxUsdPerDay, default $1) as a safety net.

Configuration

All keys are optional. Put them under fastCompaction in ~/.pi/agent/settings.json (global) or .pi/settings.json (project; project wins).

{
  "fastCompaction": {
    "enabled": true,
    "reasons": ["threshold", "overflow", "manual"],
    "minReduction": 0.4,        // fast path needs ≥40% of the span removed
    "maxVerbatimTokens": 40000, // budget for the verbatim history
    "keepThreshold": 0.3,       // P(keep) below this → truncate
    "truncateHeadChars": 300,   // head kept from a truncated result
    "argHeadChars": 400,        // clip long string args in kept history (0 = keep whole)
    "thinkingChars": 400,       // assistant thinking kept per block (0 = drop)
    "protectedTools": [],       // tool names whose results are never truncated
    "stateHeadChars": 200,      // result head shown to Jev
    "maxStateTokens": 25000,
    "maxRequestTokens": 30000,
    "jevModel": "jev-latest",
    "jevTimeoutMs": 15000,
    "maxUsdPerDay": 1,
    "summariser": { "model": "openai-codex/gpt-5.6-luna", "thinkingLevel": "high" }
  }
}

Summariser model

The hybrid and summary paths use fastCompaction.summariser when set. Otherwise they use the compactionModel block that pi-compaction-model reads, so if you already have that configured nothing more is needed:

{ "compactionModel": { "model": "openai-codex/gpt-5.6-luna", "thinkingLevel": "high" } }

With neither set, the session model summarises.

You do not need pi-compaction-model installed alongside this extension: the summariser selection is built in and covers the same cases. If you keep both, list this package first in packages so it runs first; pi-compaction-model will then only handle compactions this extension declined.

Environment overrides

FAST_COMPACTION_ENABLED=0 disables the extension for a process. FAST_COMPACTION_MIN_REDUCTION, FAST_COMPACTION_KEEP_THRESHOLD and FAST_COMPACTION_MAX_VERBATIM_TOKENS override the matching settings; useful for forcing a path while testing.

Commands

  • /fast-compaction — status: configuration in force, key present, last run.
  • /fast-compaction dry-run — score the current branch (since the last compaction) with Jev and report which path would be taken, the reduction, and the lowest-scored calls. Nothing is compacted.
  • /fast-compaction config — dump the resolved configuration.

Tuning

Jev's P(keep) for a tool result is usually low — it sees the call, the result head and size, and the recent task, and most outputs really are consumed the moment they arrive. On real sessions the distribution sits around 0.1–0.2 with a thin tail above. keepThreshold 0.3 evicts nearly everything in that band; lower it towards 0.15 to keep more, or rely on maxVerbatimTokens and let the budget decide.

minReduction decides when verbatim is worth it. Past a compaction the model re-reads the kept history on every turn, so a verbatim block that is much larger than an LLM summary costs cache-read tokens for the rest of the cycle. Around 40–50% is a reasonable floor; set it to 1 to always summarise (you still get the speed of summarising a pruned transcript).

Development

npm install
npm run check          # typecheck + tests (node:test, no network)
JEV_API_KEY=… node --experimental-strip-types scripts/replay.ts span.json [keepThreshold] [minReduction]

scripts/replay.ts runs the engine over a JSON array of Pi AgentMessages with live Jev and writes the rendered verbatim history for inspection.

Releasing

Bump version in package.json, commit, then publish a GitHub release with tag vX.Y.Z. The Publish workflow verifies the tag matches the package version, runs the checks, and publishes to npm with provenance.

Credits

The idea of scoring tool calls with Jev and keeping survivors verbatim, and the staged fitting of the state under Jev's request limit, follow tamaratran/fast-jev-compaction and vava-nessa/pi-jev-compaction. This is an independent implementation built around Pi's compaction pipeline: it prunes the exact span Pi prepared, carries previous summaries forward, and routes the fallback through Pi's own summariser.

Licence

MIT