@french-castle/jev-router
Route each coding-agent turn to the cheapest model that can finish it. Switchyard-style execution signals, TypeSafe's jev as the judge, and honest cost accounting. Works with Claude Code, Codex, OpenCode, and Pi.
Package details
Install @french-castle/jev-router from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:@french-castle/jev-router- Package
@french-castle/jev-router- Version
0.3.0- Published
- Sep 30, 2026
- Downloads
- 294/mo · 294/wk
- Author
- french-castle
- License
- MIT
- Types
- extension
- Size
- 399.8 KB
- Dependencies
- 0 dependencies · 2 peers
Pi manifest JSON
{
"extensions": [
"./dist/adapters/pi/index.js"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
jev-router
Route every turn of your coding agent to the cheapest model that can finish it.
jev-router sits between Claude Code, Codex, OpenCode, or Pi and your model gateway. It watches how the agent is doing, asks TypeSafe's jev a few typed questions when the situation is unclear, and picks a model and reasoning effort per turn. It logs every decision with what it would have cost on every other model, so you can see whether routing pays before you trust it.
- One key, one bill. An OpenRouter or Vercel AI Gateway key serves both the judge and inference.
- Nothing rewritten but the model. Requests are forwarded in the client's own wire format. Streaming, tools, thinking, and prompt caching pass through untouched.
- Fails open. Judge down, key missing, model not found: the request goes through on the default model and the log says why.
- Honest numbers.
statscompares against always-cheap and always-frontier alike, including when the router loses.
Requirements
- Node 22 or later
- One of the harnesses: Claude Code, Codex, OpenCode, Gemini CLI, or Pi
- A key that can reach jev, for the judge: an OpenRouter key or a Vercel AI Gateway key. It also serves inference for harnesses that have no login of their own
Quick start
npm install -g @french-castle/jev-router
jev-router setup
setup asks once for a judge key (an OpenRouter or Vercel AI Gateway key; jev costs about $0.00003 a decision), finds your Claude Code and Codex logins and any gateway keys, writes ~/.jev-router/policy.json with live prices, points every installed harness at the relay, installs the relay as a background service (launchd on macOS, systemd on Linux), and checks that it answers. Preview everything with --dry-run. npx @french-castle/jev-router setup works without a global install.
Then use your harness as usual. Claude Code shows auto (jev-router) in /model and is set to it; Codex and OpenCode use the auto model; Gemini CLI uses jev-router/auto; Pi runs the extension in-process. The pieces are also available one at a time: init, ping, up, service; see jev-router help.
To undo it: jev-router service uninstall stops and removes the background relay, every harness file setup changed has a .bak copy next to it, and ~/.jev-router holds the policy, the decision log, and the judge key file; delete it and the package is gone.
What you get
| What you have | Judge | Inference |
|---|---|---|
OPENROUTER_API_KEY |
jev through OpenRouter Decisions | OpenRouter |
AI_GATEWAY_API_KEY |
jev through Vercel AI Gateway | Vercel AI Gateway |
TYPESAFE_API_KEY plus one of the above |
jev direct from TypeSafe (init --judge typesafe) |
that gateway |
TYPESAFE_API_KEY only |
jev direct from TypeSafe | none for the relay; the Pi extension still routes with Pi's own providers |
| Claude Code logged in with a claude.ai plan (Pro, Max) | one of the keys above | your plan, through Claude Code's own login: Haiku, Sonnet, Opus |
| Codex logged in with ChatGPT | one of the keys above | your plan, through Codex's own login: the models your plan lists |
GEMINI_API_KEY |
not used for the judge | Gemini CLI only, on the Gemini API: flash-lite, flash, pro |
| Gemini CLI logged in with Google | one of the keys above | Gemini CLI only, through its own login to Code Assist: flash-lite, flash, pro |
| Ollama running locally | unchanged | a free local tier for auxiliary and compaction calls in OpenAI chat format |
Logins are never copied or stored. The harness sends its own credentials, the relay forwards them unchanged to Anthropic, OpenAI, or Google and only chooses the model, and init reads nothing but the plan type to know which models to offer. Costs for plan-backed models are shown at API list prices, the same scale a plan's allowance is consumed on. As the plan's five-hour window fills (80%, then 95%), the router caps the tier so you do not hit the limit mid-task.
The generated policy has three candidates on models available on both gateways:
| Tier | Model | When it is chosen |
|---|---|---|
| fast | openai/gpt-6-luna |
Default. Routine edits, lookups, auxiliary calls, compaction |
| mid | anthropic/claude-sonnet-5 |
Substantial work, anything needing careful reasoning, high stakes such as auth or payments |
| frontier | anthropic/claude-opus-5.5 |
Deep problems with high stakes, and recovery after repeated tool failures |
Edit ~/.jev-router/policy.json to change models, prices, or rules. jev-router policy validates it. See docs/policy.md.
How it decides
- Hard overrides. A request right after context compaction, three consecutive all-failure tool batches, or a critical error escalates one tier and holds it for two turns.
- Holds and leases. A clean tool continuation reuses the last decision. Nothing is re-judged mid-chain unless something went wrong.
- Tool signals. Recent tool outcomes are scored the way NVIDIA's Switchyard does it: error severity, spinning, and exploring push toward a capable model, steady production pushes toward a cheap one. A decisive score skips the judge.
- The judge. On a new user turn, or when the signals are ambiguous, one jev call answers five task questions or three execution questions against a bounded summary. jev never sees the full conversation. It costs about three thousandths of a cent and returns in a few hundred milliseconds.
- Policy. Plain rules map the answers to a candidate and an effort. Low confidence on a question a rule depends on keeps the current tier.
- Switch cost. An escalation first raises reasoning effort on the current model and changes model only once effort is maxed out. A downgrade that would drop more prompt cache than it saves over the next few turns is skipped. Stakes rules and hard overrides still switch at once. On a plan, caps on the five-hour window's usage come last and bound every step above.
Every decision is one line in ~/.jev-router/decisions.jsonl: raw answers, the decision and its reasons, whether it took effect, the tokens the upstream reported, and the cost on every other candidate.
Harnesses
| Harness | How routing is applied | What the harness tells the router |
|---|---|---|
| Claude Code | Local relay via ANTHROPIC_BASE_URL, on your claude.ai login or a gateway key |
Plugin hooks report tool successes and failures, compaction, subagents, API errors; gateway hint headers carry request class |
| Codex | Local relay as a Responses-API model provider, on your ChatGPT login or a gateway key | Hooks report tool results, compaction, prompts |
| OpenCode | Local relay as an OpenAI-compatible provider | Plugin tags requests with the session and reports tool results, compaction, API errors |
| Gemini CLI | Local relay via GOOGLE_GEMINI_BASE_URL (API key) or CODE_ASSIST_ENDPOINT (Google login), model jev-router/auto |
Tool results from the request body; with a Google login, hooks report tool results and prompts |
| Pi | In-process extension, no relay | Everything: prompts, tool results, compaction, model changes |
| Cursor (Chat and Agent) | A token-guarded relay published by jev-router expose, set as Cursor's OpenAI base URL; Tab and the Cursor CLI are not routed |
The request body only: its tool results; no hooks |
jev-router setup configures whichever of the first five are installed, backs up every file it touches, and keeps the relay running as a background service. Cursor calls the relay from its own servers, so setup prints its steps instead of writing files. docs/harnesses.md has the manual steps, the service commands, and the Claude Code marketplace install.
Measure before believing
jev-router stats # actual cost versus every single-model baseline
jev-router replay --policy new.json # re-decide the same log under another policy, no judge calls
jev-router up --shadow frontier # serve one model, log what the router would have done
jev-router why --last 3 # why the last three turns went where they did
None of this is a benchmark. docs/evaluation.md is the runbook for Terminal-Bench through Harbor with each harness against single-model baselines. Until that has been run, treat any savings figure as unproven.
Library
import { plan, loadPolicy, emptySession } from "@french-castle/jev-router/core";
import { HttpJudge } from "@french-castle/jev-router/judge";
const policy = loadPolicy(JSON.parse(await readFile("policy.json", "utf8")));
const judge = new HttpJudge({ transport: "openrouter", apiKey: process.env.OPENROUTER_API_KEY! });
const outcome = plan({ request, session: emptySession(), policy, policyId: "default" });
const { decision } = outcome.kind === "decision" ? outcome : outcome.conclude((await judge.evaluate(outcome.judgeRequest)).answers);
plan is pure: no network, no clock. The core and judge modules run on Node, Bun, and edge runtimes.
Status and limits
Early. The mechanics have been verified live against OpenRouter, and against Claude Code's and Codex's own logins: real judge decisions, and Anthropic streaming, OpenAI chat, and Responses requests routed and echoed correctly. What has not been done: the benchmark that would justify a savings claim, live runs against Vercel and TypeSafe, and the AI SDK middleware. Run the relay with Node; under Bun a client cancellation does not propagate to the upstream. The relay binds to localhost; binding wider requires --token, and a --token bind cannot forward a harness's own login, so plan-backed routing stays on loopback.
Documentation
- Design: the architecture, the decision pipeline, and every decision with its rationale
- Policy reference
- Harness setup
- Evaluation runbook
- Changelog
Contributing
See CONTRIBUTING.md. bun install && bun run check runs everything CI runs. Security reports go to the address in SECURITY.md.
Acknowledgements
The execution-phase routing idea comes from NVIDIA Switchyard. The judge is jev by TypeSafe. The OpenRouter Decisions and Vercel AI Gateway evaluation surfaces make one-key setups possible.
MIT licensed.