Jev Router Explained: When an LLM Router Saves You Money

Jev Router is an LLM router that asks TypeSafe's Jev how hard each request is, then sends it to a model priced for that level. On mixed traffic of 3,000-token requests with 600-token answers, it costs less than Claude Opus 5.5 alone, and less than mid-price GPT-6.1 Sol when easy requests outnumber hard ones by a few percentage points.
Prices, model lists and test results are as of October 5, 2026.
What is Jev Router and what problem does it solve?
Jev Router is a single model ID on AI/ML API, typesafe/jev-router, that covers 8 models in 3 price tiers. You send every request to that ID, and TypeSafe's Jev 1.13 picks the tier that answers it. It targets one cause of a high LLM bill: paying a flagship model's rate for easy work.
Why one model for everything gets expensive
If your app sends every request to the model that handles its hardest job, every request pays that model's rate. A request to reformat a date then costs as much per token as a multi-step proof.
The price gap between models is wide, and Jev Router saves money by sending easy requests to the cheapest tier. Here is what one typical request of 3,000 input and 600 output tokens costs at AI/ML API list prices, through Jev Router or with everything sent to the flagship Claude Opus 5.5:
An easy request through Jev Router costs about 32 times less than on Claude Opus 5.5. A hard one costs the same, plus a $0.00021 routing decision.
You could split traffic yourself, with rules that send each kind of request to a cheaper model. Rules take time to write, and they need updating whenever your prompts or the models change. A router makes that choice for each request instead, so you can use several LLMs at low cost through one endpoint.
Jev Router's advantages
Jev Router's main advantage is cost: requests Jev rates easy are billed at Light-tier rates instead of flagship prices. It has 3 other advantages:
- Switch with little code: OpenAI SDK code needs only a new base URL, an AI/ML API key and the model ID.
- A known per-token price ceiling: all 8 models the router can choose from are listed with their prices, so you know the highest rate per token a request can pay. Setting
max_tokenscaps how long an answer can get. - Routing you can audit: you can see which model answered and what each routing decision cost.
How does Jev Router work?
Jev Router handles each request in 2 steps. Jev 1.13 first rates how hard the request is, and then the first model in the matching tier writes the answer. Requests rated easy go to the Light tier, medium ones to Standard and hard ones to Heavy.
Jev 1.13 is a decision model. Instead of writing text, it returns probabilities, so it scores the request but never answers it. This is LLM routing by difficulty, and the router chooses before any model starts writing.
The 3 tiers and their models
Each tier is a short list of models for one level of difficulty, in the order they're tried. Prices come from AI/ML API's model list.
AI/ML API's second router, Liquid Router (liquid/d1-router), uses the same 8 models and tiers with a different decision model. Prices and backups work the same way on both routers.
When a model or the decision fails
If the tier's first model fails, the next model in the same tier answers the request. If the Jev decision fails, the request goes to the Heavy tier, and Claude Opus 5.5 answers it at Heavy-tier prices, even if the request was easy. Either way, your app still gets an answer.
No backup costs less per token than its tier's first model. The most expensive is GPT-6 Astra, which answers only when Claude Opus 5.5 fails and costs 2.5 times as much. A failed attempt is billed only if the model provider charged us for it.
How much does Jev Router cost?
Through Jev Router, each request from your user costs the answering model's tokens at its list price, plus a Jev decision that costs a fraction of a cent. The router adds no other fee. Jev 1.13 costs $0.0577668 per 1M input tokens, and its output is free. The Jev Router pricing page lists every model's input and output price.
What a routing decision costs
A routing decision costs about $0.00004–$0.00006 on short requests and $0.00021 on a typical one. Each decision is billed as its own request. It reads your request plus about 600 tokens that the router adds, so for a 3,000-token request it reads about 3,600 tokens. That adds just over a quarter to the price of a GPT-6 Luna answer and 0.7% to a Claude Opus 5.5 one.
Each decision shows up in your usage log as its own line, with the same request ID as the answer it routed.

In this log, a decision read 805 tokens for a 256-token request, cost $0.00005 and sent the request to Claude Opus 5.5, whose 3,567-token answer cost $0.09408. The cost in each API response includes only the answer, not the decision. To check a bill, add the decision line from the usage log.
Tool calls add no routing cost. When a model asks to use a tool, your app runs it and sends the result back. That follow-up skips the decision and goes straight back to the model that asked. Only new messages from your user get a decision, so a request with 5 tool rounds pays for 1 decision instead of 6.
Is Jev Router free?
No. AI/ML API's Free Tier is paused, so every request is paid. There is no fee to top up your balance.
How much can Jev Router save?
Jev Router can reduce LLM API costs by 27% against running everything on GPT-6.1 Sol, in the calculator example on the Jev Router page: $17,739 a month instead of $24,180. That example assumes 1,550,000 typical requests a month, 45% of them easy, 40% medium and 15% hard. With the same settings, routing costs 63% less than running everything on Claude Opus 5.5.
A worked example from the calculator
The calculator takes your own request count and mix. It assumes an answer from each tier's first model and counts each decision at 3,000 input tokens. Real decisions also read the router's own tokens, about 600 more, which would add about $54 a month to this example, 0.3%. Compared with running everything on GPT-6.1 Sol, routing mainly changes the price of 2 kinds of request.
An easy request moves to GPT-6 Luna, which charges 1/20 of GPT-6.1 Sol's model price. A hard request moves up to Claude Opus 5.5, which charges twice as much as GPT-6.1 Sol. Medium requests stay on GPT-6.1 Sol and pay only the added decision. So the result depends on how many easy and hard requests you send:
Monthly cost of 1,550,000 requests of 3,000 input and 600 output tokens at AI/ML API list prices, each answered by its tier's first model, with 1 Jev decision per request counted at 3,000 input tokens, as in the calculator.
With requests of this size, Jev Router costs less than running everything on Claude Opus 5.5 in every row. It also costs less than GPT-6.1 Sol whenever easy requests outnumber hard ones by a few points, so a team that already runs everything on GPT-6.1 Sol can still cut its bill by routing.
What our test showed
A lower bill only helps if the answers stay right, and we tested that on 205 tasks. Jev Router matched Claude Opus 5.5's 97.6% accuracy for 35% less, $2.37 against $3.66. Routing did cost accuracy on medium tasks: Jev Router solved 48 of the 50 and Liquid Router 49, while every model called directly solved all 50. On the 40 hard tasks, both routers solved all 40, as did Claude Opus 5.5, while GPT-6.1 Sol and GPT-6 Luna each solved 95%.
On the 65 hardest tasks, labeled hard+, both routers solved more than Claude Opus 5.5. Jev Router solved 95.4% and Liquid Router 96.9%, against 92.3% for Claude Opus 5.5 and 89.2% for GPT-6 Luna. That is 2 more tasks than Claude Opus 5.5 for Jev Router and 3 more for Liquid Router, from a single run. GPT-6.1 Sol solved 98.5%.
Time to first token is how long an answer took to start, measured for the answering model, so it may leave out the decision. Time per call includes the decision. p95 means 95% of calls finished within that time.
How we tested
- We sent 205 automatically graded tasks once each to both routers and to GPT-6 Luna, GPT-6.1 Sol and Claude Opus 5.5 called directly.
- Easy (50): everyday jobs we wrote, such as JSON extraction and date formatting. Medium (50): GSM8K grade-school math and HumanEval Python coding.
- Hard (40): MATH-500 level 5 and AIME 2025 competition math. Hard+ (65): GPQA Diamond graduate-level science questions, HMMT February 2026 and AIME 2026.
- Costs are at AI/ML API list prices and include the routing decisions.
- Every prompt was a short single message without tools, so long cached conversations and agent loops weren't tested.
- It's a single run by us, the vendor, on mostly public benchmarks, so treat it as one data point.
Why request size changes the result
The test's prompts were short, unlike the calculator's typical requests, and that changed which option was cheapest. In the test, GPT-6.1 Sol on its own cost $0.52 for all 205 tasks at 98.5% accuracy, less than either router. There are 2 reasons.
First, easy tasks cost almost nothing on any model when prompts are short. The 50 easy tasks cost $0.0024 through Jev Router and $0.0497 on Claude Opus 5.5, so routing cut a cost that was only cents to begin with. Second, most of the test's cost came from hard tasks that the routers sent to Claude Opus 5.5. That model wrote 3.8 times as many output tokens as GPT-6.1 Sol, at twice the price per output token.
Requests that carry thousands of tokens of context pay the full price gap between tiers on every call. That is where sending easy requests to a Light model saves the most, as the calculator shows. The calculator also assumes answers of the same length on every tier, so if your Heavy-tier answers run longer, your hard requests will cost more than it shows.
How routing reduces token usage
Besides choosing cheaper models, the router saves tokens on Light-tier answers, which come without reasoning unless you ask for it. Reasoning tokens are what a model spends thinking before it answers, and they're billed as output.
Called directly on all 205 tasks, GPT-6 Luna spent 175,283 reasoning tokens, while the answers the router sent to it used none. Both routers sent all 50 easy tasks to GPT-6 Luna, and it solved every one without reasoning.
How do you use Jev Router?
Point any OpenAI-compatible client at https://api.aimlapi.com/v1 with an AI/ML API key, and set the model to typesafe/jev-router. Your messages and tool definitions stay as they are.
Jev Router setup in 4 steps
Only the last 2 steps touch your code:
- Create an API key in your AI/ML API dashboard.
- Save it as an environment variable:
export AIMLAPI_KEY=... - Set the base URL to
https://api.aimlapi.com/v1. - Set the model ID to
typesafe/jev-router.
The base URL points your client at AI/ML API's gateway, the single endpoint in front of every model it hosts. What Is an LLM API Gateway? explains how that works.
Python and Node.js examples
Here is a first request in Python with the OpenAI SDK. Compared with calling a fixed model at another provider, only the base URL, the API key and the model ID change.

When we ran this example, GPT-6 Luna, the Light tier's first model, answered it. The answer cost $0.0000105, and the routing decision was billed separately. The same call in Node.js:
Which model answered and when to add reasoning
meta.model holds the full ID of the model that answered, such as openai/gpt-6-luna. You get it in the response body, or in the final chunk when you stream. Count its values over a sample of your requests to see your tier shares. The Request Tracing and Cost Headers docs list the other tracing and cost data AI/ML API returns with each response.
The router doesn't choose a reasoning effort, so reasoning_effort is up to you. Set it if you want Light-tier models to think through harder prompts, and keep in mind that the extra reasoning tokens are billed as output.
What other decision model based routers are there?
An LLM router picks which model answers each request. Besides Jev Router, the main options are AI/ML API's Liquid Router, OpenRouter Auto Router, Microsoft Foundry model router, Not Diamond, LiteLLM, Portkey and cascades. They differ mostly in how they decide: a decision model rates the request, a classifier trained only for routing sorts it, rules you write match it, or a cascade tries a cheap model first and escalates.
What is Liquid Router?
Liquid Router (liquid/d1-router) is AI/ML API's second router. It works like Jev Router, except that requests are rated by Liquid AI's d1 instead of Jev 1.13. The d1 model is billed at $0.0546 per 1M input tokens, 5.5% less than Jev 1.13, and its output is also free.
In our test, Liquid Router solved 98.5% of the 205 tasks for $2.55, tied for the highest score with GPT-6.1 Sol, which cost $0.52. Jev Router and Liquid Router chose the same tier for 87.8% of tasks. Liquid Router sent more of the 50 medium tasks to the Standard tier and solved 1 more of them than Jev Router, a gap too small to pick a winner from a single run. Both routers work with the same API key, so you can compare them on your own traffic by changing only the model ID.
Liquid AI says d1 is the first model to outperform Jev on Hugging Face's Decision Index. In GIGAZINE's coverage, a results chart from an internal reproduction of Decision Index 0.2.1 puts d1 at 58.9 against 57.9 for Jev 1.13. In that chart, d1 leads in arts, language and retrieval and trails in tools and knowledge. KuCoin reports the same scores. We found no independent run, and our test is too small to settle the question.
How other routers decide
OpenRouter Auto Router uses a trained classifier to sort prompts by task type, where AI/ML API's 2 routers sort them by difficulty. Microsoft Foundry model router has a balanced mode that picks the cheapest model within about 1–2% of the highest predicted quality, though Microsoft's docs give no savings figure. According to Not Diamond, its router picks a model and an effort level for each step of a coding agent.
Portkey suits teams that already know where each request should go, for example by a user's plan or region. LiteLLM is an open-source proxy you host yourself, and its beta Auto Routing lets you pick the classifier, with Jev as one option.
A cascade can pay for 2 or 3 answers to one request, so it pays off only when the cheap answer usually passes. Stanford's FrugalGPT paper matched the most accurate single model at up to 98% lower cost, at March 2023 prices. Across its 3 datasets, the savings ranged from 59.2% to 98.3%.
If you're weighing OpenRouter as a whole platform, the best OpenRouter alternatives in 2026 compares the options.
What you pay comes from each vendor's docs and pricing pages and leaves out the models' own token prices.
Other projects that use the Jev name
OpenRouter hosts its own Jev Router under the same model ID, typesafe/jev-router. The base URL decides which one answers, and AI/ML API's is at https://api.aimlapi.com/v1. The Jev name is also used by 5 other projects, none of which claims a link to TypeSafe or OpenRouter:
- JevRouter (jevrouter.co), by BillionsBobby: an MIT-licensed router for models, tools and subagents, unrelated to AI/ML API.
- jev-router, by gargpratyush: an MIT-licensed model picker for Claude Code.
- System1 Models (jev-router.com): a hosted Jev-compatible API for open-weight decision models.
- Jev-chooses-a-LLM, by Bodila51: an MIT-licensed MCP server for Cursor.
- pi-jev-model-router, by da-vinci-noob: an MIT-licensed Pi package.
When is a router the wrong choice?
A router can be the wrong choice in 4 situations:
- Long agent sessions that rely on a prompt cache
- Features where response time matters most
- Requests whose difficulty doesn't show in the text
- Traffic with as many hard requests as easy ones
In each case, test the router on your own traffic before you switch.
Problem 1: Long agent sessions lose their cache
A prompt cache lets a model reread the long, unchanged start of a conversation at a discount, and that cache belongs to one model. According to OpenRouter, each model switch loses the cached chat, and the new model rereads it at full price. On AI/ML API, GPT-6.1 Sol's input costs $0.13 per 1M tokens cached and $2.60 uncached, so a cache miss costs 20 times as much.
Jev Router makes a fresh decision for each new message and doesn't take into account that switching models loses the cache. In a long agent session, a switch between tiers can mean paying full price for the whole context again. Before moving such sessions to the router, run a few of them both ways, through the router and on one cached model, and compare the bills.
Problem 2: Each message waits for a decision
Every message from your user waits for the decision before its answer starts. For example, TypeSafe says Jev responds in 70–500 ms.
In our test of short single-message requests without streaming, Jev Router's median time per call, decision included, was 3.15 s. That is close to GPT-6 Luna's 3.05 s called directly and well under Claude Opus 5.5's 6.01 s. Direct GPT-6 Luna calls also spent time reasoning, which routed Light answers skip, so the test can't isolate how long the decision itself takes.
Before routing a feature where response time matters most, time a sample of your real requests through the router and directly. That comparison shows the wait your users would feel.
Problem 3: The prompt hides the real difficulty
A one-line bug report can point to a hard fix in a large codebase that Jev never sees. TypeSafe also lists math and dates among Jev 1.13's weak spots, so requests that hinge on them can be misrated.
When Jev underrates a request, a cheap model gets a hard job, so unless you set reasoning_effort, it answers without thinking the problem through. To catch this, check meta.model on requests you know are hard. If a Light-tier model answered them, the rating missed.
Problem 4: Your traffic mix doesn't favor routing
At the calculator's request size, routing loses its edge over GPT-6.1 Sol once hard requests are about as common as easy ones: a mix of 30% easy, 40% medium and 30% hard costs 3% more than running everything on GPT-6.1 Sol. Route a sample, count tier shares from meta.model and enter them in the pricing calculator. If the router comes out more expensive and GPT-6.1 Sol handles your hardest requests, send everything to GPT-6.1 Sol. The router would send that hard share to Claude Opus 5.5, at twice the price per token.
Very short, easy traffic is the other case. On the test's 50 easy tasks, the routed answers cost about $0.0005 in total, less than the decisions that picked them. Called directly, GPT-6 Luna did the same 50 tasks for $0.0010, reasoning included, less than half of what they cost through Jev Router. If nearly all your requests look like that, call GPT-6 Luna directly and skip the decision.
How to test a router first
A trial on your own requests shows which of these problems apply to you:
- Take 300 real requests with personal data removed.
- Send each one to Jev Router, Liquid Router and the model you use now, with the same
max_tokensandreasoning_effort. - Log
meta.model, tokens, cost and time for every call, and add each router's decision lines from the usage log to its cost. - Grade the answers. Check them automatically where there's a right answer; for open-ended text, have a model from a different developer compare the answers side by side without knowing which target wrote them. Then compare cost per 1,000 requests and tier shares.
FAQ
What is jev for AI?
Jev is a decision model from TypeSafe, announced on September 15, 2026, and the first of its System One models. Instead of writing text, it answers set kinds of questions, such as a choice or a score, with probabilities. Jev Router uses Jev 1.13 to rate how hard each request is. TypeSafe's Jev explained covers the model itself.
How do AI model routers work?
An AI model router sits between your app and several models and picks one per request. It decides in 1 of 4 ways. A decision model such as Jev or d1 rates the request, a trained classifier sorts it, or rules you write match it. A cascade tries a cheap model first and escalates when its answer fails a check.
What is the difference between an LLM router and an AI gateway?
A router decides which model answers each request. An AI gateway is the single layer all your requests pass through, giving you one API, key management, fallbacks, logs and cost tracking. A gateway can include a router. Portkey's gateway routes on rules you write, and AI/ML API's gateway offers Jev Router, which decides from the request itself.
Is OpenRouter auto free?
The routing is free, but the requests aren't. According to OpenRouter's docs, Auto Router (openrouter/auto) adds no fee of its own, and each request is billed at the selected model's standard rate. OpenRouter does charge a fee when you buy credits, so you pay the model's price plus that fee.
Does OpenRouter have smart routing features?
Yes. Auto Router (openrouter/auto) sorts each prompt into one of about 30 task types and ranks models by what OpenRouter users spent on that task over the last 7 days. OpenRouter also hosts a Jev Router under typesafe/jev-router, and its models API lists other routers, such as openrouter/pareto-code.
How much does an open router AI cost?
On OpenRouter, you pay the selected model's provider rate, plus a fee when you buy credits: 5.5% on the Standard plan or 8% on Business, according to its pricing page. Open-source routers such as LiteLLM are free to host yourself. You pay the model providers, plus Jev calls if you use Jev as the classifier.
How to lower API costs?
Send easy requests to cheaper models, and avoid paying full price for repeated context. Jev Router automates the first. At 3,000 input and 600 output tokens a request, it costs less than GPT-6.1 Sol alone when easy requests outnumber hard ones by a few points. For the second, keep long agent sessions on one model so its prompt cache works.
How much is 1 million tokens in LLM?
It depends on the model and on whether the tokens are input or output. Among the models behind Jev Router on AI/ML API, 1M tokens cost from $0.13 input and $0.65 output on GPT-6 Luna to $13 input and $65 output on GPT-6 Astra. A mid-price model such as GPT-6.1 Sol charges $2.60 input and $13 output.
What is dynamic LLM routing and how does it work?
Dynamic LLM routing picks the model for each request at run time instead of hard-coding one. A classifier or decision model scores the request, for example on difficulty, and a fixed mapping turns that score into a model or a price tier. Jev Router and Liquid Router map Jev or d1 ratings to the Light, Standard and Heavy tiers.
Key Takeaways
- Jev Router (
typesafe/jev-router) has TypeSafe's Jev rate each request and sends it to 1 of 8 models in 3 price tiers. Liquid Router does the same with Liquid AI's d1 model. - Each request from your user costs the answering model's tokens plus 1 Jev decision, about $0.00021 for a 3,000-token request. Jev 1.13 costs $0.0577668 per 1M input tokens, and its output is free.
- In the calculator on AI/ML API's Jev Router page, 1,550,000 monthly requests of 3,000 input and 600 output tokens, 45% easy, 40% medium and 15% hard, cost $17,739 through Jev Router against $24,180 on GPT-6.1 Sol, 27% less.
- In our 205-task test, Jev Router matched Claude Opus 5.5's 97.6% accuracy for 35% less, but scored 96% on medium tasks, where every model called directly scored 100%.
- On the 65 hardest tasks in our test, Jev Router solved 95.4% and Liquid Router 96.9%, 2 and 3 tasks more than Claude Opus 5.5 (92.3%) in a single run. GPT-6.1 Sol solved 98.5%.
- Our test used short prompts, and there GPT-6.1 Sol alone was cheaper than both routers at the same or higher accuracy: 98.5% for $0.52, against 97.6% for $2.37 with Jev Router and 98.5% for $2.55 with Liquid Router. Routing saves more as requests grow and the share of easy ones rises.



