# Open source Jev alternatives sucks, here is the benchmark

> I ran three LLM routers on 563 real prompts in 14 languages. Jev routed 84% correctly. Two local open models barely beat always answering "medium".

- Source: https://jays.fyi/blog/local-llm-routers-lost-to-always-answering-medium
- Author: Jay Derinbogaz
- Published: 2026-09-23
- Language: en
- Tags: ai, open-source, benchmarks
- Reading time: 7 min

---

Every chat product with a model dropdown hands its users a routing job. They
have to guess whether their question deserves the expensive model or the cheap
one, and most people pick once and never touch it again. So they overpay on
"what is puppy love?" and underpay on a proof about Bernoulli samples.

We are building automated model selection at
[TextCortex](https://textcortex.com/) so nobody has to guess. A router reads the
prompt and picks the cheapest tier that will still answer it well. The obvious
way to build one is a small open model on your own hardware: no per-call fee,
about 60 milliseconds a call, and the prompt never leaves the building. I
assumed that was where we would end up.

So I measured it. On 563 real chatbot prompts in 14 languages, the hosted model
Jev routed **84.5%** of them to the right tier. Laya and Von, open models on my
laptop, peaked at 61.6% and 58.8%. A router that answers "medium" to
every prompt scores 56.3%. Call that the constant router. Four of the six Laya
and Von runs lost to it.

## A router is one multiple-choice question asked before every prompt.

All three systems speak the same API. You send the prompt, plus a question with
named options and a sentence describing each. You get back a probability per
option.

<figure>
  <img src="/images/typesafe-jev-routing-playground.png" alt="Typesafe AI playground: jev-latest routes 'Calculate fibonacci' to the cheapest of three tiers at 84%, confidence 76%." width="3032" height="1548" decoding="async" fetchpriority="high" />
  <figcaption>My first routing question in the <a href="https://typesafe.ai/">Typesafe AI</a> playground, with made-up tier names. The benchmark used new wording and neutral tier names.</figcaption>
</figure>

"Calculate fibonacci" goes to the cheap tier. Obviously. I first tested all
three routers on prompts like that, and Jev won that round too, but the prompts
I write for a router are prompts I already know the answer to. It proved almost
nothing.

## I used real prompts, because mine were too easy.

The prompts come from
[WildChat-1M](https://huggingface.co/datasets/allenai/WildChat-1M), a public
dataset of real chatbot conversations (ODC-BY licence, toxic ones filtered
out). I took 40 first messages in each of English, German, French, Spanish,
Italian, Portuguese, Dutch, Polish, Turkish, Russian, Arabic, Chinese, Japanese
and Korean, then added 168 longer or more technical prompts, because real
traffic contains almost no hard ones.

Two blind annotators on different models, Claude Opus and Claude Sonnet,
labelled all 728 prompts small, medium or frontier. They agreed on 77.5% (Cohen's
κ = 0.60). I kept the 563 they agreed on: 225 small, 317 medium, 21 frontier.

Each system answered three questions per prompt, 5,067 calls with zero hard
failures:

- **Q1, described.** "Pick the cheapest tier that will still answer it well,"
  with example tasks for each tier.
- **Q2, minimal.** "Simple," "moderately complex" and "very hard requests."
- **Q3, score.** A three-level difficulty score, mapped onto tiers.

## Two of the three routers barely beat a constant.

<figure>
  <img src="/images/llm-router-accuracy-by-run.png" alt="Bar chart: Jev 84.5%, 84.2% and 70.5%; Laya and Von between 42.3% and 61.6%, against 56.3% for always answering medium." width="1600" height="980" decoding="async" loading="lazy" />
  <figcaption>Accuracy on the 563 agreed prompts, with 95% bootstrap intervals. The length rule sends prompts under 60 characters to small, under 1,500 to medium, the rest to frontier.</figcaption>
</figure>

Jev scored 84.5% with two- and three-word descriptions and 84.2% with the full
rubric. That gap is noise, so nobody on my team has to become a prompt engineer to write a routing question.

Laya's best run beat the constant router by 5.3 points. Von's beat it by 2.5,
with a 95% interval (55% to 63%) that still contains the constant. A rule that
looks only at prompt length scored 54.7%, above two local runs and within a
point of two more.

A router that loses to `return "medium"` is a random number generator with a
model download attached.

## The local models guess "medium" and call it a decision.

<figure>
  <img src="/images/llm-router-answer-distribution.png" alt="Stacked bars of tiers chosen per run. Von Q2 sent 550 of 563 to medium. Laya Q1 sent 86 to frontier. Jev sent at most 7 to frontier." width="1600" height="1019" decoding="async" loading="lazy" />
  <figcaption>What each run answered, against the gold labels.</figcaption>
</figure>

Von on Q2 answered "medium" 550 times out of 563. On Q3 it never once said
small, and sent 139 prompts to frontier when 21 belong there. Laya on Q1 sent
86 prompts to frontier and was right about six. Its multilingual checkpoint
handled 1,329 of its 1,689 calls, and it still scored 29% on French.

## Von and Laya overspend on up to 55% of prompts.

<figure>
  <img src="/images/llm-router-error-direction.png" alt="Diverging bars: Von and Laya send 20% to 55% of prompts to a pricier tier; Jev's Q2 and Q3 misses lean cheaper." width="1600" height="980" decoding="async" loading="lazy" />
  <figcaption>Share of prompts sent cheaper (left) or pricier (right) than the gold label.</figcaption>
</figure>

Von on Q3 sends 55% of all prompts to a pricier tier than they need. That
router raises the bill it exists to cut.

Jev leans the other way on Q2 and Q3, and that is worse per mistake. An
overpriced answer costs the price difference. An underpowered one costs a user
who asked a hard question and got a confident wrong reply. On Q2, 51 of Jev's
misses were medium prompts sent to small. On Q3, 113.

## Language barely moved Jev and moved everyone else a lot.

<figure>
  <img src="/images/llm-router-accuracy-by-language.png" alt="Dot plot by language: Jev 74% to 92%, above the baseline in all 14; Von 42% to 69%; Laya 29% to 67%." width="1600" height="1120" decoding="async" loading="lazy" />
  <figcaption>Accuracy by language on Q1. With 31 to 53 prompts per language, one prompt moves a score by 2 to 3 points.</figcaption>
</figure>

Jev stayed between 74% (German) and 92% (English), and scored 84% on Latin and
non-Latin scripts alike. Von ran from 42% in German to 69% in Portuguese, Laya
from 29% in French to 67% in Dutch. A router that gets French right less than a
third of the time cannot go near French customers.

## Jev misses the prompts that matter most.

Of the 21 frontier prompts, Jev on Q1 sent six to frontier and 15 to medium. On
Q2 it sent three. Those are where routing cheap does the most damage: a
statistics problem in Turkish, investment decision rules in German, an outline
on quantum key distribution. Twenty-one is a small sample, and I am saying so
first.

Jev's probabilities help. Escalate to frontier whenever Jev gives it more than
0.2, and Q1 catches 14 of the 21 while escalating only 18 of 563 prompts. Q2 at
the same cutoff catches seven. I picked 0.2 after seeing this data, so treat it
as a starting point to tune.

Three more things cut against the headline:

- **84% flatters Jev.** The gold set is the prompts both annotators agreed on,
  the easier ones. The annotators agreed only 77.5% of the time on the full set,
  so about 80% is a realistic ceiling for any router with these tiers.
- **These are judgement labels.** Two models guessed which tier a prompt needs.
  Nobody checked what each tier actually produced.
- **Jev is slower.** Median latency was about 330 ms per call from my laptop,
  mostly network, against 56 to 170 ms for the local models.

## The cheap router is the expensive one.

We could not ship Jev in the end: it has no EU deployment. So I fine-tuned Laya
on our own labels, and that router
[now runs on one CPU server in Germany](/blog/fine-tuned-llm-router-on-a-hetzner-cpu-server).

A local router costs nothing per call. If it sends 41% of your traffic to a tier
it does not need, as Von did on Q2, you pay for that on every one of those
calls, and that cost appears on no model card I have read.

So count. Take WildChat or your own logs, label a few hundred prompts with two
annotators, drop the ones they disagree on, and score your router against
`return "medium"`. If it cannot clear that line by ten points, you do not have a
router.

The better test, which we are running next, is outcome-based: push real traffic
through every tier, record the cheapest tier whose answer was good enough, and
score the router against that. If Jev's 84% falls apart there, I will publish
that too.
