@french-castle/jev-router

Route each coding-agent turn to the cheapest model that can finish it. Switchyard-style execution signals, TypeSafe's jev as the judge, and honest cost accounting. Works with Claude Code, Codex, OpenCode, and Pi.

Packages

Package details

extension

Install @french-castle/jev-router from npm and Pi will load the resources declared by the package manifest.

$ pi install npm:@french-castle/jev-router
Package
@french-castle/jev-router
Version
0.3.0
Published
Sep 30, 2026
Downloads
294/mo · 294/wk
Author
french-castle
License
MIT
Types
extension
Size
399.8 KB
Dependencies
0 dependencies · 2 peers
Pi manifest JSON
{
  "extensions": [
    "./dist/adapters/pi/index.js"
  ]
}

Security note

Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.

README

jev-router

CI npm license: MIT

Route every turn of your coding agent to the cheapest model that can finish it.

jev-router sits between Claude Code, Codex, OpenCode, or Pi and your model gateway. It watches how the agent is doing, asks TypeSafe's jev a few typed questions when the situation is unclear, and picks a model and reasoning effort per turn. It logs every decision with what it would have cost on every other model, so you can see whether routing pays before you trust it.

  • One key, one bill. An OpenRouter or Vercel AI Gateway key serves both the judge and inference.
  • Nothing rewritten but the model. Requests are forwarded in the client's own wire format. Streaming, tools, thinking, and prompt caching pass through untouched.
  • Fails open. Judge down, key missing, model not found: the request goes through on the default model and the log says why.
  • Honest numbers. stats compares against always-cheap and always-frontier alike, including when the router loses.

Requirements

Quick start

npm install -g @french-castle/jev-router
jev-router setup

setup asks once for a judge key (an OpenRouter or Vercel AI Gateway key; jev costs about $0.00003 a decision), finds your Claude Code and Codex logins and any gateway keys, writes ~/.jev-router/policy.json with live prices, points every installed harness at the relay, installs the relay as a background service (launchd on macOS, systemd on Linux), and checks that it answers. Preview everything with --dry-run. npx @french-castle/jev-router setup works without a global install.

Then use your harness as usual. Claude Code shows auto (jev-router) in /model and is set to it; Codex and OpenCode use the auto model; Gemini CLI uses jev-router/auto; Pi runs the extension in-process. The pieces are also available one at a time: init, ping, up, service; see jev-router help.

To undo it: jev-router service uninstall stops and removes the background relay, every harness file setup changed has a .bak copy next to it, and ~/.jev-router holds the policy, the decision log, and the judge key file; delete it and the package is gone.

What you get

What you have Judge Inference
OPENROUTER_API_KEY jev through OpenRouter Decisions OpenRouter
AI_GATEWAY_API_KEY jev through Vercel AI Gateway Vercel AI Gateway
TYPESAFE_API_KEY plus one of the above jev direct from TypeSafe (init --judge typesafe) that gateway
TYPESAFE_API_KEY only jev direct from TypeSafe none for the relay; the Pi extension still routes with Pi's own providers
Claude Code logged in with a claude.ai plan (Pro, Max) one of the keys above your plan, through Claude Code's own login: Haiku, Sonnet, Opus
Codex logged in with ChatGPT one of the keys above your plan, through Codex's own login: the models your plan lists
GEMINI_API_KEY not used for the judge Gemini CLI only, on the Gemini API: flash-lite, flash, pro
Gemini CLI logged in with Google one of the keys above Gemini CLI only, through its own login to Code Assist: flash-lite, flash, pro
Ollama running locally unchanged a free local tier for auxiliary and compaction calls in OpenAI chat format

Logins are never copied or stored. The harness sends its own credentials, the relay forwards them unchanged to Anthropic, OpenAI, or Google and only chooses the model, and init reads nothing but the plan type to know which models to offer. Costs for plan-backed models are shown at API list prices, the same scale a plan's allowance is consumed on. As the plan's five-hour window fills (80%, then 95%), the router caps the tier so you do not hit the limit mid-task.

The generated policy has three candidates on models available on both gateways:

Tier Model When it is chosen
fast openai/gpt-6-luna Default. Routine edits, lookups, auxiliary calls, compaction
mid anthropic/claude-sonnet-5 Substantial work, anything needing careful reasoning, high stakes such as auth or payments
frontier anthropic/claude-opus-5.5 Deep problems with high stakes, and recovery after repeated tool failures

Edit ~/.jev-router/policy.json to change models, prices, or rules. jev-router policy validates it. See docs/policy.md.

How it decides

  1. Hard overrides. A request right after context compaction, three consecutive all-failure tool batches, or a critical error escalates one tier and holds it for two turns.
  2. Holds and leases. A clean tool continuation reuses the last decision. Nothing is re-judged mid-chain unless something went wrong.
  3. Tool signals. Recent tool outcomes are scored the way NVIDIA's Switchyard does it: error severity, spinning, and exploring push toward a capable model, steady production pushes toward a cheap one. A decisive score skips the judge.
  4. The judge. On a new user turn, or when the signals are ambiguous, one jev call answers five task questions or three execution questions against a bounded summary. jev never sees the full conversation. It costs about three thousandths of a cent and returns in a few hundred milliseconds.
  5. Policy. Plain rules map the answers to a candidate and an effort. Low confidence on a question a rule depends on keeps the current tier.
  6. Switch cost. An escalation first raises reasoning effort on the current model and changes model only once effort is maxed out. A downgrade that would drop more prompt cache than it saves over the next few turns is skipped. Stakes rules and hard overrides still switch at once. On a plan, caps on the five-hour window's usage come last and bound every step above.

Every decision is one line in ~/.jev-router/decisions.jsonl: raw answers, the decision and its reasons, whether it took effect, the tokens the upstream reported, and the cost on every other candidate.

Harnesses

Harness How routing is applied What the harness tells the router
Claude Code Local relay via ANTHROPIC_BASE_URL, on your claude.ai login or a gateway key Plugin hooks report tool successes and failures, compaction, subagents, API errors; gateway hint headers carry request class
Codex Local relay as a Responses-API model provider, on your ChatGPT login or a gateway key Hooks report tool results, compaction, prompts
OpenCode Local relay as an OpenAI-compatible provider Plugin tags requests with the session and reports tool results, compaction, API errors
Gemini CLI Local relay via GOOGLE_GEMINI_BASE_URL (API key) or CODE_ASSIST_ENDPOINT (Google login), model jev-router/auto Tool results from the request body; with a Google login, hooks report tool results and prompts
Pi In-process extension, no relay Everything: prompts, tool results, compaction, model changes
Cursor (Chat and Agent) A token-guarded relay published by jev-router expose, set as Cursor's OpenAI base URL; Tab and the Cursor CLI are not routed The request body only: its tool results; no hooks

jev-router setup configures whichever of the first five are installed, backs up every file it touches, and keeps the relay running as a background service. Cursor calls the relay from its own servers, so setup prints its steps instead of writing files. docs/harnesses.md has the manual steps, the service commands, and the Claude Code marketplace install.

Measure before believing

jev-router stats                       # actual cost versus every single-model baseline
jev-router replay --policy new.json    # re-decide the same log under another policy, no judge calls
jev-router up --shadow frontier        # serve one model, log what the router would have done
jev-router why --last 3                # why the last three turns went where they did

None of this is a benchmark. docs/evaluation.md is the runbook for Terminal-Bench through Harbor with each harness against single-model baselines. Until that has been run, treat any savings figure as unproven.

Library

import { plan, loadPolicy, emptySession } from "@french-castle/jev-router/core";
import { HttpJudge } from "@french-castle/jev-router/judge";

const policy = loadPolicy(JSON.parse(await readFile("policy.json", "utf8")));
const judge = new HttpJudge({ transport: "openrouter", apiKey: process.env.OPENROUTER_API_KEY! });

const outcome = plan({ request, session: emptySession(), policy, policyId: "default" });
const { decision } = outcome.kind === "decision" ? outcome : outcome.conclude((await judge.evaluate(outcome.judgeRequest)).answers);

plan is pure: no network, no clock. The core and judge modules run on Node, Bun, and edge runtimes.

Status and limits

Early. The mechanics have been verified live against OpenRouter, and against Claude Code's and Codex's own logins: real judge decisions, and Anthropic streaming, OpenAI chat, and Responses requests routed and echoed correctly. What has not been done: the benchmark that would justify a savings claim, live runs against Vercel and TypeSafe, and the AI SDK middleware. Run the relay with Node; under Bun a client cancellation does not propagate to the upstream. The relay binds to localhost; binding wider requires --token, and a --token bind cannot forward a harness's own login, so plan-backed routing stays on loopback.

Documentation

Contributing

See CONTRIBUTING.md. bun install && bun run check runs everything CI runs. Security reports go to the address in SECURITY.md.

Acknowledgements

The execution-phase routing idea comes from NVIDIA Switchyard. The judge is jev by TypeSafe. The OpenRouter Decisions and Vercel AI Gateway evaluation surfaces make one-key setups possible.

MIT licensed.