pi-l1-cache
L1 response cache for pi with replay, disk persistence, and working capture
Package details
Install pi-l1-cache from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:pi-l1-cache- Package
pi-l1-cache- Version
1.4.0- Published
- Sep 20, 2026
- Downloads
- 538/mo · 13/wk
- Author
- graphwiz
- License
- MIT
- Types
- extension
- Size
- 56.2 KB
- Dependencies
- 0 dependencies · 1 peer
Pi manifest JSON
{
"extensions": [
"./src/index.ts"
],
"minPiVersion": "0.4.0"
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
pi-l1-cache
L1 In-Memory Cache Extension for pi — with disk persistence, response replay, and working capture
A high-performance, production-ready L1 (in-memory + disk) cache extension for the pi coding agent. Unlike the stock API which lacks response capture, this implementation uses a pi-core patch to enable full caching with replay.
⚡ Performance
| Scenario | Cold | Warm (replayed) | Speedup |
|---|---|---|---|
| Simple prompt | 3.2s | 1.9s | ~1.7× |
| Tool-calling agent loop | 13.1s | 2.5s | 5.2× |
Features
| Feature | Description |
|---|---|
| Response replay | Cached chunks are fed through pi-ai's normal consume path — parsing, usage, tool-calls, stop-reason all identical |
| Disk persistence | Survives across pi -p process restarts (~/.pi/agent/cache/l1-cache/) |
| ~138µs overhead | End-to-end per-request cost (measured, incl. key hashing + JSON) |
| Fast key hash | FNV-1a (~1.5µs/op) — no per-request CPU probe on the hot path |
| Memory cap | Hard limit (20MB default) prevents RAM bloat |
| Auto-eviction | LRU-style cleanup when limits reached |
| TTL-based | 1-hour default expiry for cached entries |
| Coalescing | Adjacent content/reasoning_content deltas merged to shrink storage |
Architecture
pi → [L1: RAM Map + disk] → Provider
~138µs lookup 1-3s
The extension uses a pi-core patch (fix-l1-cache.cjs) to add two capabilities the stock extension API lacks:
Bundle era (pi ≥ 0.84): pi runs from the esbuild bundle (
dist/bundle/cli.js), sofix-l1-cache.cjspatchesdist/bundle/chunks/*(minified, anchor-matched — verified against 0.86.0). Pre-bundle pi (< 0.84) is patched in the readable pi-ai sources. The patch is idempotent and bundle-era misses degrade to pass-through rather than failing an install.
REPLAY — an extension may serve a cached response by returning params with
__piL1Replay: [chunks]frombefore_provider_request; pi-ai feeds the cached chunks through the normal consume path and never contacts the provider.CAPTURE — after a successful completion, pi-ai calls
options.onStreamComplete(allChunks, requestParams);sdk.jsforwards them to extensions as aprovider_stream_completeevent so the cache can store the response.
Installation
Option 1: From npm (recommended)
pi install npm:pi-l1-cache
This installs the extension AND automatically applies the fix-l1-cache.cjs pi-core patch.
Option 2: From GitHub
pi install git:github.com:tobias-weiss-ai-xr/pi-l1-cache@main
Option 3: From local clone
pi install /path/to/pi-l1-cache
Then restart pi — the extension auto-loads.
Configuration
Environment Variables
# Enable/disable
L1_CACHE_ENABLED=true
# Max entries (default: 200)
L1_CACHE_MAX_ENTRIES=200
# Max memory in MB (default: 20)
L1_CACHE_MAX_MB=20
# TTL in seconds (default: 3600)
L1_CACHE_TTL=3600
# Log stats on each hit/miss (default: false)
L1_CACHE_LOG=true
Settings (in ~/.pi/settings.json)
{
"extensions": {
"l1-cache": {
"enabled": true,
"maxEntries": 200,
"maxMemoryBytes": 20971520,
"ttlSeconds": 3600,
"persist": true,
"logStats": false
}
}
}
Usage
Show cache stats
/l1-cache
Output:
L1 cache: 16 entries, 43.2KB | hits 12 (replays 12), misses 4, writes 4, evictions 0 | dir: C:/Users/Tobias/.pi/agent/cache/l1-cache
Clear cache
/l1-cache clear
Cleared: memory + disk (all .json files in cache dir).
Key Semantics
The cache key is a stable hash of:
modelmessages(full conversation history)toolstool_choicetemperature,top_preasoning_effort,thinkingmax_completion_tokens,max_tokens
Volatile fields are EXCLUDED: prompt_cache_key, prompt_cache_retention, stream, stream_options, store, sessionId.
This means:
- ✅ Cross-run hits possible (same prompt, same cwd, same model)
- ✅ Tool-calling agent loops fully replayed (tool results are part of messages)
- ❌ Different conversation history = different key (expected)
- ❌ Session-derived fields don't break cross-run hits
Gotchas
Replays are canned — identical input returns the stored response verbatim. For fresh answers, use
/l1-cache clear.Print mode (
pi -p) reads entire stdin as ONE prompt — you cannot test two identical requests in one process this way.llm-timestamp.js does NOT pollute provider messages (it only appends display-level
message_endentries), so keys are stable across runs.Replayed responses carry original responseId/usage — accurate since input identical.
Troubleshooting
"fix-reasoning-content.js: layout changed"
The fix-l1-cache.cjs patch modifies the same file as fix-reasoning-content.js. The wrapper's postinstall chain handles this correctly — fix-reasoning-content.js now recognizes its work via marker even after the replay branch is added.
If you see this error:
- Ensure you're running the full postinstall chain (not individual scripts)
- Check that
fix-reasoning-content.jshas the marker-based detection (v1.2.3+) - Verify postinstall order:
fix-reasoning-content.jsBEFOREfix-l1-cache.cjs
Cache never hits
Check:
L1_CACHE_LOG=trueto see HIT/MISS/STORED logs- Prompt is byte-identical (including system prompt, tools, cwd context)
- Disk persistence is enabled (
persist: true) - TTL hasn't expired (default 1h)
Development
Testing
# Run unit tests
npm test
# Run the extension manually
npx tsx src/index.ts
Benchmark
# Measure cold vs warm times
echo "What is 7*6?" | time pi -p "test" # cold
echo "What is 7*6?" | time pi -p "test" # warm (should be ~1.7× faster)
License
MIT — see LICENSE