pi-midflight
Pi extension that routes mid-flight user messages to continue, inject, replan, or restart based on semantic assessment by Jev
Package details
Install pi-midflight from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:pi-midflight- Package
pi-midflight- Version
0.1.0- Published
- Sep 19, 2026
- Downloads
- 165/mo · 165/wk
- Author
- romilchouhan
- License
- MIT
- Types
- extension
- Size
- 52 KB
- Dependencies
- 1 dependency · 2 peers
Pi manifest JSON
{
"extensions": [
"./src/index.ts"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
pi-midflight
A pi extension that decides what to do when you send a message while the agent is still working on your previous request.
Instead of always steering or always queueing, it takes a compact snapshot of the run in progress, asks a decision model how the new message relates to the work, and routes it:
| Action | Meaning | What happens to your message |
|---|---|---|
| CONTINUE | Does not change this response, or explicitly says “after this” | Queued as a follow-up after the run finishes; current run stays untouched |
| INJECT | Adds info or constraints for the remaining work | Delivered as a steering message; completed work kept |
| REPLAN | Changes approach/scope, completed work still partly valid | Steering message asking the model to keep valid work, restate it, and replan the rest |
| RESTART | Invalidates the task or most completed work | Run aborted; combined request (original + new instruction) sent fresh |
The decision model returns signals only (delivery timing, semantic change, invalidation, confidence). The action is chosen by a fixed policy in this package, never by the model. Explicit sequencing such as “after this,” “once this is done,” and “then do…” is a hard queue instruction: it outranks semantic distance, so an unrelated next task cannot abort the current response.
Install
pi install npm:pi-midflight
Or try it without installing:
pi -e npm:pi-midflight
Configuration
TYPESAFE_API_KEY— API key for TypeSafe's Jev model. When absent, the extension falls back to your active pi model, and if that also fails, to a deterministic INJECT (which never destroys work and never drops your message).
The latency path is deliberately bounded: Jev gets 1800 ms, then the active-model fallback gets 1400 ms. The fallback is capped at 192 output tokens; if neither tier answers in time, the extension immediately takes the non-destructive INJECT route. A single Jev client is reused and the decision snapshot is kept under a few KB.
Live decision UI
The moment a mid-stream message is submitted, a footer status and widget below the editor show what Jev is evaluating. Pi's streaming output and working line are left untouched, so assessment rendering never overlaps the response. When the decision lands, the widget flips to the action, scores, route, and latency. A compact decision card is also inserted into the transcript (TUI-only—it is not sent back to the model):
◆ MID-FLIGHT REPLAN · Jev · 824ms
update “use PostgreSQL advisory locks instead”
change ████████░░ 75% invalid █████░░░░░ 50%
→ 50% invalidation crosses the replan threshold; preserve valid work, redirect the rest
Press Ctrl+T to expand the card and see Jev's selected rubric labels and confidence. This makes the routing legible in screen recordings without mixing harness deliberation into the assistant's answer. Internal routing wrappers are hidden: after an abort, the transcript shows the clean user request first and its decision card second. The full conversation prefix is retained unchanged for provider prompt-cache reuse. To prevent unrelated older answers from leaking into the retry, the appended instruction explicitly identifies the authoritative active task, current response excerpt, and new requirement.
Plain control messages such as stop, cancel, and never mind bypass Jev and abort immediately without triggering another model response.
Token comparison
Token and timing counters are collected silently and never add a transcript block, widget, or footer item. Query them only when needed with /midflight-usage; the notification separates agent and decision tokens, elapsed time, and output spent before an abort. For a true with-vs-without experiment, this package also includes src/baseline.ts: a usage-only observer with no input handler. It does not perform mid-flight assessment, routing, prompt injection, or message modification.
Trial A — native Pi baseline:
pi -e ./src/baseline.ts
Send the initial prompt, then send the update at the same point you plan to use in Trial B. The observer reports cumulative session usage as BASELINE · MID-FLIGHT DISABLED. Use /baseline-usage to inspect it at any time.
Trial B — mid-flight enabled:
pi -e ./src/index.ts
Use a fresh session and repeat the same prompt and update. Compare its final combined count against the baseline's cumulative session count. Use /midflight-usage for current totals.
The baseline observer itself makes no model calls, so it adds zero tokens. /baseline-usage and /midflight-usage report wall-clock elapsed time on demand; the mid-flight command also includes decision usage and latency. No automatic usage UI is rendered. Decision time may overlap agent time, so compare elapsed values rather than adding timing components.
Output accuracy
Token and latency savings only matter if final quality holds. Compare final answers blind with templates/accuracy-judge.md. The rubric scores final-constraint adherence, technical correctness, completeness, feasibility, and clarity, with a stale-context penalty. Use only each trial's final answer—not aborted partial output or decision cards—and run the judge twice with A/B positions swapped to reduce position bias.
Every assessment is also recorded as a midflight-decision entry in the session file (event id, timestamp, original request, new message, compact snapshot, signals, policy reason, chosen action, per-tier timing—usable for evals). TUI mode performs no raw stderr or file logging; visibility comes from the status, widget, transcript cards, and session entries without adding synchronous I/O to the routing path. Headless modes trace to stderr:
[midflight:input] captured mid-stream message (52 chars, streaming=steer, toolsUsed=0, runningTools=0, pendingDelivery=no)
[midflight:jev] ok in 1240ms
[midflight:jev] timingIntent=NOW (conf 0.94)
[midflight:jev] semanticChange=0.75/4 → "Changes the requested approach or scope enough that the remaining plan should differ" (conf 0.90)
[midflight:jev] invalidation=0.50/4 → "Roughly half of the completed work no longer applies" (conf 0.85)
[midflight:decision] source=jev jev=1240ms sc=0.75 inv=0.50 conf=0.85 assess=1245ms
[midflight:action] REPLAN snap=0ms assess=1245ms total=1247ms
[midflight:deliver] REPLAN -> abort-and-resend (text-only run)
In the TUI, Jev failures switch the isolated widget to the fallback tier. In headless mode, [midflight:jev] FAILED after 1800ms: APITimeoutError: … includes the exact error and latency on stderr.
Testing
Interactive:
pi -e ./src/index.ts
Start a long task, then type a second message while it streams. Press Ctrl+T to expand thinking blocks. Inspect decisions afterwards with:
grep midflight-decision ~/.pi/agent/sessions/*/*.jsonl | tail -1
Limitations
This is checkpoint/context-level continuation, not token-level continuation. Steering messages are delivered at a turn boundary (after the current assistant turn finishes its tool calls), not mid-token. On a text-only run there is no such boundary until the whole answer ends, so both INJECT and REPLAN escalate to abort-and-resend: the partial answer is cancelled immediately and the model is re-prompted with the update. CONTINUE remains queued because it explicitly targets a later turn. RESTART aborts the run but does not roll back side effects: files written or commands executed by already-completed tool calls stay on disk. Decisions are made on a compact snapshot (six short turn summaries plus the live head/tail of output currently streaming), so borderline calls are made on partial evidence. After an abort, no earlier messages are removed or rewritten: the stable conversation prefix remains cacheable, and a compact disambiguation instruction is appended with the authoritative current task plus current-response excerpt. REPLAN is enforced by that prompt wording, so the model may still redo valid portions of the interrupted response.
Development
npm install
npm run check # typecheck + policy tests, no network
pi -e ./src/index.ts
Publishing
The package manifest exposes ./src/index.ts as the Pi extension. Before the first release:
npm login
npm whoami
npm run check
npm pack --dry-run
npm publish --access public
Then verify the published package without installing it:
pi -e npm:pi-midflight@0.1.0
Install it permanently with:
pi install npm:pi-midflight
For later releases, increment the version first (npm version patch, minor, or major) and publish again.