pi-retrospect
Tools to explore past Pi sessions to facilitate harness self-improvement
Package details
Install pi-retrospect from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:pi-retrospect- Package
pi-retrospect- Version
0.3.0- Published
- Oct 7, 2026
- Downloads
- 150/mo · 150/wk
- Author
- nietaki
- License
- MIT
- Types
- extension, skill
- Size
- 196.9 KB
- Dependencies
- 0 dependencies · 2 peers
Pi manifest JSON
{
"skills": [
"./skills"
],
"extensions": [
"./src/index.ts"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
pi-retrospect
pi-retrospect is a Pi package for
learning from previous Pi sessions. It helps agents recover prior context, audit
completed work, compare approaches, understand how a task was handled, and improve
the local harness that produced it.
Past sessions can reveal repeated failures, ineffective instructions, tool friction, delegation problems, and opportunities to improve local skills, prompts, tools, and workflows.
The package provides two read-only tools for discovering recorded sessions and inspecting their entries. It also includes a skill that guides agents through retrospective analysis and an optional setting that makes steering corrections easier to find later.
Status:
pi-retrospectis an early 0.x package. Its session-discovery and transcript-reading tools are implemented and usable, but broader capabilities such as ranking or modifying history remain out of scope until their behavior is discussed.
What can it help with?
- Recover decisions or context from an earlier session.
- Understand how a previous task was approached.
- Audit what an agent or delegated subagent actually did.
- Compare approaches used across multiple sessions.
- Find recurring failures, corrections, or misunderstandings.
- Identify improvements to instructions, skills, prompts, tools, and workflows.
Install
pi install npm:pi-retrospect
Or try it for a single invocation without adding it to settings:
pi -e npm:pi-retrospect
pi list shows configured packages, pi remove pi-retrospect removes it.
Requires Pi 0.99.0 or newer. The tool registration uses exposure, annotations,
and outputSchema, which were added in Pi 0.99.0 (2026-09-29). Verified against
0.99.2 and 1.0.0. The peer ranges stay "*" because that is the convention Pi
prescribes for host-provided packages — Pi does not resolve them for managed installs,
so the minimum is stated here rather than in package.json.
Using it
Two tools register, and they chain: list_sessions names transcript files, session_entries opens
one of them.
list_sessions walks the Pi sessions root, reads only each file's header line, and returns
session metadata — id, absolute path, absolute cwd, timestamp, fork lineage
(parentSessionPath) — with subagent transcripts nested under the session that launched them, plus
a warning for every file it had to skip.
All of its parameters are optional, and all of them act on top-level sessions — a matching parent always arrives with its complete subagent tree:
| Parameter | Default | Meaning |
|---|---|---|
cwds |
all working directories | Absolute cwds to keep. |
cwdMatch |
"exact" |
"sibling-prefix" also keeps sibling directories whose basename extends the requested one — the shape of git worktrees placed next to the main checkout (a lexical path rule; no git metadata is read). |
includeCurrentSession |
false |
Keep the session this call runs inside. By default it is dropped, with the transcripts nested under it, before filtering, sorting, and limit — a retrospective normally means earlier sessions. Pi names it by its session file, never by id, so a copy of it survives; an ephemeral session has no file and so excludes nothing. |
startTimestamp, endTimestamp |
unbounded | Inclusive ISO 8601 bounds, read in the host timezone: a bare date is one whole calendar day, and a date-time with no offset is local to the machine running the tool. |
sortBy, sortDirection |
"timestamp", "asc" |
Top-level order only; children always stay in launch order. "desc" puts the newest first. |
limit |
none | Cap on returned rows, after filtering and sorting. Headers are still all read. |
// in a codemode script — the project and its worktrees, ten newest first
const { sessions } = await tools.list_sessions({
cwds: ["/Users/me/repos/my-app"],
cwdMatch: "sibling-prefix",
sortDirection: "desc",
limit: 10,
});
It is exposed to codemode, not to the model. The tool registers with
exposure: "codemode", so it is never declared in the model's tool list and is not
activated on registration. Call it from a codemode script:
// in a codemode script
// every session but this one, oldest first
const { sessions, warnings } = await tools.list_sessions({});
A plain "list my sessions" prompt will not reach it unless your own settings declare the tool directly. This is deliberate — the result is structured JSON that a script can filter before it costs context — but it does mean the tool is invisible to a session running without codemode.
session_entries takes sessionPath plus optional filters, and returns the entries of that file:
{ lineNo, id, parentId, timestamp, type, messageRole, text, raw }, where raw is the whole parsed JSON
line unchanged and text is the entry's primary human-readable body — a message's content, a system
message's content plus its prompt sections, a compaction or branch summary, a custom_message content,
a context_edit replacement, a session_info name, a usage note, a label, or the command of a !
shell run — or null when the entry has no such payload. Assistant thinking, tool calls, and images
never reach text; they are still in raw. Line 1 is the session header and is never returned, so
lineNo starts at 2 and a
malformed line costs a warning without shifting the lines after it. Unknown entry types and unknown
message roles come back verbatim, and the call cannot leave the sessions root — a relative path, a
.. traversal, a symlink that resolves outside it, and a file whose first line is not a session
header all throw.
| Parameter | Default | Meaning |
|---|---|---|
startLineNo, endLineNo |
unbounded | Inclusive physical line bounds. |
ids, parentIds |
unfiltered | Exact, case-sensitive sets of entry ids. A null field matches nothing, so a version 1 file is never selected. |
types, messageRoles |
unfiltered | Exact, case-sensitive sets. messageRoles reaches only type: "message" rows. |
search |
unfiltered | Literal substring search over text: { terms: string[], caseSensitive?: boolean }. A row matches when its non-null text contains any term. Case-insensitive by default. |
startTimestamp, endTimestamp |
unbounded | Inclusive ISO 8601 bounds on each entry's own timestamp, read in the host timezone — the same grammar list_sessions uses. |
limit |
none | Cap on returned entries, applied after filtering. |
Filters are ANDed, values inside one array are ORed, and order is never configurable: rows come back
in file order. Filtering narrows the result, never the scan — warnings still describe the
whole file. To page, pass startLineNo one past the last lineNo you already read.
search is the one filter that is not exact. Terms are literal bytes — no pattern, no tokenization,
no glob — so ["h.llo"] matches only h.llo and ["the"] matches inside there; a term of "" and
an empty terms array are refused rather than read as "every row". It runs over text, never raw,
so thinking, tool calls, images, and the output of a ! shell run are unreachable (they are still in
raw for a script to filter), and a row whose text is null is never a hit. Because a system row's
text is that message's rendered prompt, an ordinary word matches harness text — the preamble, the
tool rules, every AGENTS.md — so AND the search with types or messageRoles when the question is
about what was said.
// in a codemode script — project the rows in the script, never hand `raw` to a model
// the newest *previous* session: this one is dropped before `limit` applies
const { sessions } = await tools.list_sessions({ sortDirection: "desc", limit: 1 });
const { entries, warnings } = await tools.session_entries({
sessionPath: sessions[0].path,
messageRoles: ["assistant"],
});
if (warnings.length > 0) return { skipped: warnings.length, warnings };
return entries.map((entry) => ({
lineNo: entry.lineNo,
said: entry.text,
stopReason: entry.raw.message.stopReason,
}));
// in a codemode script — the rows that mention one error, in the conversation only
const { entries } = await tools.session_entries({
sessionPath,
search: { terms: ["ETIMEDOUT", "connection timed out"] },
messageRoles: ["user", "assistant", "toolResult"],
});
return entries.map((entry) => ({ lineNo: entry.lineNo, role: entry.messageRole, text: entry.text }));
raw is unbounded per row — as large as the entries it keeps (a 2.4 MB session returned 2.46 MB of
raw) — and text is bounded only by the entry it was projected from (measured max 51 KB on a tool
result), so it is a codemode-only tool by design. A system row is the largest category text carries
by mean (16.7 KB, max 38.3 KB): it projects that message's rendered prompt, which is one message's own
state — a session folds several such rows to get the prompt the model actually had.
It reads stored history: no compaction, no context_edit, no branch selection is applied, so it is not
the model's context view. In a session file older than version 2, id and parentId come back null
even where the line stores them — Pi
replaces every id when it migrates such a file — and one legacy_version warning says so; lineNo
is the handle that stays valid, and raw keeps what was written.
Steering-message markers
A mid-run correction often means the agent misunderstood the task. Making those corrections findable
turns one-off friction into evidence of recurring harness problems: an instruction that is not
landing, a tool that keeps getting misused, or a repo whose AGENTS.md needs a clearer rule.
Enable marking
Installing the package registers the two read-only tools without changing input. To opt into marking,
set markSteeringMessages in the user-level ~/.pi/agent/settings.json or a trusted project's
.pi/settings.json, which Pi merges over the user value:
{
"piRetrospect": {
"markSteeringMessages": true
}
}
The default is false. After editing the settings file, run /reload.
What gets marked
While the setting is enabled, pi-retrospect prepends STEERING: to steering input submitted from
the interactive UI or an RPC client while the agent is streaming. It leaves idle prompts, queued
follow-ups, extension-generated input, slash-prefixed input, and text that already starts with the
exact marker unchanged. Attached images are preserved.
The prefix is part of the user message sent to the model and stored in the transcript, not separate metadata. This changes what the model sees — often usefully, because the correction is explicitly labelled — as well as making the message searchable later.
Find and interpret markers
Use a literal, case-sensitive search over user messages, then keep entries where the automatic marker appears as a prefix:
// in a codemode script — the marked steering messages of one session
const { entries, warnings } = await tools.session_entries({
sessionPath,
messageRoles: ["user"],
search: { terms: ["STEERING: "], caseSensitive: true },
});
return {
steeringMessages: entries
.filter(({ text }) => text?.startsWith("STEERING: ") === true)
.map(({ lineNo, text }) => ({ lineNo, text })),
warnings,
};
Treat matches as high-signal candidates, not authoritative metadata. A user can type the same prefix manually, and unmarked steering can exist when the setting was disabled or the input belonged to an excluded category. Inspect the surrounding conversation before deciding why the user intervened, and compare sessions before concluding that a misunderstanding recurs.
Only messages submitted while marking is enabled receive the prefix. Existing history is never
rewritten, and disabling the setting does not remove markers already stored. The full behavior is in
docs/tool-api.md.
Reference
docs/tool-api.md— the contract for the operations this package exposes:list_sessions(discovery rules, filters, ordering, guarantees) andsession_entries(sessions-root confinement, line addressing, entry filters, the literaltextsearch, thetextprojection,raw, warning codes).test/fixtures/generate.mjs(source repository, not in the npm tarball) — rebuilds the synthetic session tree the tests run against.
The deeper reference on Pi's session format — the measurements and Pi-internals analysis this contract was derived from — lives in the author's Obsidian vault rather than in the published package, because it documents Pi's schema (which changes with Pi, not with this package) and quotes counts from one developer's local session store.
Development
Run the gate locally with npm run check — it cleans test/tmp/, runs the Vitest suite,
then type-checks with tsc --noEmit. Do not run two Vitest processes in one checkout at
once: the start-of-run purge is not concurrency-safe.
CI is .github/workflows/ci.yml: on pull requests, pushes to master, and manual dispatch it
installs with npm ci, runs npm run check, and verifies the tarball contents with
npm pack --dry-run. It publishes nothing.
Releasing
Releases are run locally with release-it. Start from a
clean master branch that tracks its upstream and make sure npm is authenticated for this package.
Preview the interactive flow without changing Git or npm state:
npm run release:dry-run
A dry run still performs read-only prerequisite checks such as npm authentication. For a real release, run either the interactive version selector or name the SemVer increment explicitly:
npm run release
npm run release -- patch
The release runs npm run check, updates package.json and package-lock.json, creates and pushes
a chore: release vX.Y.Z commit and vX.Y.Z tag, and publishes the package to npm. The existing
prepublishOnly guard runs the checks again immediately before publication, so a direct
npm publish remains protected too. This workflow does not create a GitHub Release or maintain a
changelog.
License
MIT — see LICENSE.