pi-fast-compaction
Fast compaction for Pi: Jev-scored verbatim eviction of stale tool history, with LLM summarisation of the evicted transcript as the fallback.
Package details
Install pi-fast-compaction from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:pi-fast-compaction- Package
pi-fast-compaction- Version
0.1.0- Published
- Sep 18, 2026
- Downloads
- 148/mo · 148/wk
- Author
- tvdavies
- License
- MIT
- Types
- extension
- Size
- 60.5 KB
- Dependencies
- 0 dependencies · 1 peer
Pi manifest JSON
{
"extensions": [
"./extensions/fast-compaction.ts"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
pi-fast-compaction
Fast context compaction for Pi.
When Pi's context fills up it normally asks the session model to summarise the old history. That is slow (minutes on a large context with thinking enabled), costs a full uncached read of the conversation, and loses detail.
This extension takes a different first step. It asks TypeSafe's Jev — a small, fast decision model — which tool results in the old history are still needed, drops the rest, and keeps what survives verbatim. Only when that does not free enough space does it fall back to an LLM summary, and then it summarises the already-pruned transcript on a model of your choosing.
Typical numbers from a 400k-token span: Jev scored 167 tool calls in ~1.5 s for $0.002, evicted 87% of the span, and the whole compaction finished in under 3 s with no LLM call. The native path on the same span took 5–10 minutes.
How it works
On session_before_compact (threshold, overflow, or /compact):
- Score. The span Pi is about to replace is rendered as a compact state
(user/assistant text, every tool call with its arguments, a short head of
each result and its size) and sent to Jev with one question per tool call:
is the full output of this call still needed to continue the task?
Results scoring below
keepThresholdare truncated to a short head plus a note; the call and its arguments always stay. Longwrite/editbodies are clipped too — the files are on disk and listed under<modified-files>. - Fit. If the surviving history still exceeds
maxVerbatimTokens, the lowest-scored survivors are evicted until it fits. - Decide.
- Fast path — reduction ≥
minReduction: the surviving history is stored verbatim as the compaction summary, in Pi's own[User]: / [Assistant]: / [Tool result]:format. No LLM call. - Hybrid path — otherwise: Pi's summariser runs over the pruned transcript on the configured summariser model. Faster and cheaper than summarising the original, and the summary still sees the full shape of the conversation.
- Summary path — Jev unavailable (no key, timeout, daily budget hit): the original transcript is summarised on the configured summariser model.
- Default — no summariser model configured and Jev unavailable, or the summariser fails: the extension returns nothing and Pi compacts as usual.
- Fast path — reduction ≥
Each compaction entry records which path was taken and the eviction stats in
details.fastCompaction.
Install
pi install npm:pi-fast-compaction
Set a Jev API key in your environment as JEV_API_KEY (or TYPESAFE_API_KEY).
Get one at https://typesafe.ai. Jev is priced per input token; a compaction
costs a fraction of a cent, and the extension keeps a daily spend ledger
(maxUsdPerDay, default $1) as a safety net.
Configuration
All keys are optional. Put them under fastCompaction in
~/.pi/agent/settings.json (global) or .pi/settings.json (project; project
wins).
{
"fastCompaction": {
"enabled": true,
"reasons": ["threshold", "overflow", "manual"],
"minReduction": 0.4, // fast path needs ≥40% of the span removed
"maxVerbatimTokens": 40000, // budget for the verbatim history
"keepThreshold": 0.3, // P(keep) below this → truncate
"truncateHeadChars": 300, // head kept from a truncated result
"argHeadChars": 400, // clip long string args in kept history (0 = keep whole)
"thinkingChars": 400, // assistant thinking kept per block (0 = drop)
"protectedTools": [], // tool names whose results are never truncated
"stateHeadChars": 200, // result head shown to Jev
"maxStateTokens": 25000,
"maxRequestTokens": 30000,
"jevModel": "jev-latest",
"jevTimeoutMs": 15000,
"maxUsdPerDay": 1,
"summariser": { "model": "openai-codex/gpt-5.6-luna", "thinkingLevel": "high" }
}
}
Summariser model
The hybrid and summary paths use fastCompaction.summariser when set.
Otherwise they use the compactionModel block that
pi-compaction-model
reads, so if you already have that configured nothing more is needed:
{ "compactionModel": { "model": "openai-codex/gpt-5.6-luna", "thinkingLevel": "high" } }
With neither set, the session model summarises.
You do not need pi-compaction-model installed alongside this extension: the
summariser selection is built in and covers the same cases. If you keep both,
list this package first in packages so it runs first; pi-compaction-model
will then only handle compactions this extension declined.
Environment overrides
FAST_COMPACTION_ENABLED=0 disables the extension for a process.
FAST_COMPACTION_MIN_REDUCTION, FAST_COMPACTION_KEEP_THRESHOLD and
FAST_COMPACTION_MAX_VERBATIM_TOKENS override the matching settings; useful
for forcing a path while testing.
Commands
/fast-compaction— status: configuration in force, key present, last run./fast-compaction dry-run— score the current branch (since the last compaction) with Jev and report which path would be taken, the reduction, and the lowest-scored calls. Nothing is compacted./fast-compaction config— dump the resolved configuration.
Tuning
Jev's P(keep) for a tool result is usually low — it sees the call, the
result head and size, and the recent task, and most outputs really are
consumed the moment they arrive. On real sessions the distribution sits
around 0.1–0.2 with a thin tail above. keepThreshold 0.3 evicts nearly
everything in that band; lower it towards 0.15 to keep more, or rely on
maxVerbatimTokens and let the budget decide.
minReduction decides when verbatim is worth it. Past a compaction the model
re-reads the kept history on every turn, so a verbatim block that is much
larger than an LLM summary costs cache-read tokens for the rest of the cycle.
Around 40–50% is a reasonable floor; set it to 1 to always summarise (you
still get the speed of summarising a pruned transcript).
Development
npm install
npm run check # typecheck + tests (node:test, no network)
JEV_API_KEY=… node --experimental-strip-types scripts/replay.ts span.json [keepThreshold] [minReduction]
scripts/replay.ts runs the engine over a JSON array of Pi AgentMessages
with live Jev and writes the rendered verbatim history for inspection.
Releasing
Bump version in package.json, commit, then publish a GitHub release with
tag vX.Y.Z. The Publish workflow verifies the tag matches the package
version, runs the checks, and publishes to npm with provenance.
Credits
The idea of scoring tool calls with Jev and keeping survivors verbatim, and the staged fitting of the state under Jev's request limit, follow tamaratran/fast-jev-compaction and vava-nessa/pi-jev-compaction. This is an independent implementation built around Pi's compaction pipeline: it prunes the exact span Pi prepared, carries previous summaries forward, and routes the fallback through Pi's own summariser.
Licence
MIT