Every chat product with a model dropdown hands its users a routing job. They have to guess whether their question deserves the expensive model or the cheap one, and most people pick once and never touch it again. So they overpay on “what is puppy love?” and underpay on a proof about Bernoulli samples.
We are building automated model selection at TextCortex so nobody has to guess. A router reads the prompt and picks the cheapest tier that will still answer it well. The obvious way to build one is a small open model on your own hardware: no per-call fee, about 60 milliseconds a call, and the prompt never leaves the building. I assumed that was where we would end up.
So I measured it. On 563 real chatbot prompts in 14 languages, the hosted model Jev routed 84.5% of them to the right tier. Laya and Von, open models on my laptop, peaked at 61.6% and 58.8%. A router that answers “medium” to every prompt scores 56.3%. Call that the constant router. Four of the six Laya and Von runs lost to it.
A router is one multiple-choice question asked before every prompt.
All three systems speak the same API. You send the prompt, plus a question with named options and a sentence describing each. You get back a probability per option.
“Calculate fibonacci” goes to the cheap tier. Obviously. I first tested all three routers on prompts like that, and Jev won that round too, but the prompts I write for a router are prompts I already know the answer to. It proved almost nothing.
I used real prompts, because mine were too easy.
The prompts come from WildChat-1M, a public dataset of real chatbot conversations (ODC-BY licence, toxic ones filtered out). I took 40 first messages in each of English, German, French, Spanish, Italian, Portuguese, Dutch, Polish, Turkish, Russian, Arabic, Chinese, Japanese and Korean, then added 168 longer or more technical prompts, because real traffic contains almost no hard ones.
Two blind annotators on different models, Claude Opus and Claude Sonnet, labelled all 728 prompts small, medium or frontier. They agreed on 77.5% (Cohen’s κ = 0.60). I kept the 563 they agreed on: 225 small, 317 medium, 21 frontier.
Each system answered three questions per prompt, 5,067 calls with zero hard failures:
- Q1, described. “Pick the cheapest tier that will still answer it well,” with example tasks for each tier.
- Q2, minimal. “Simple,” “moderately complex” and “very hard requests.”
- Q3, score. A three-level difficulty score, mapped onto tiers.
Two of the three routers barely beat a constant.
Jev scored 84.5% with two- and three-word descriptions and 84.2% with the full rubric. That gap is noise, so nobody on my team has to become a prompt engineer to write a routing question.
Laya’s best run beat the constant router by 5.3 points. Von’s beat it by 2.5, with a 95% interval (55% to 63%) that still contains the constant. A rule that looks only at prompt length scored 54.7%, above two local runs and within a point of two more.
A router that loses to return "medium" is a random number generator with a
model download attached.
The local models guess “medium” and call it a decision.
Von on Q2 answered “medium” 550 times out of 563. On Q3 it never once said small, and sent 139 prompts to frontier when 21 belong there. Laya on Q1 sent 86 prompts to frontier and was right about six. Its multilingual checkpoint handled 1,329 of its 1,689 calls, and it still scored 29% on French.
Von and Laya overspend on up to 55% of prompts.
Von on Q3 sends 55% of all prompts to a pricier tier than they need. That router raises the bill it exists to cut.
Jev leans the other way on Q2 and Q3, and that is worse per mistake. An overpriced answer costs the price difference. An underpowered one costs a user who asked a hard question and got a confident wrong reply. On Q2, 51 of Jev’s misses were medium prompts sent to small. On Q3, 113.
Language barely moved Jev and moved everyone else a lot.
Jev stayed between 74% (German) and 92% (English), and scored 84% on Latin and non-Latin scripts alike. Von ran from 42% in German to 69% in Portuguese, Laya from 29% in French to 67% in Dutch. A router that gets French right less than a third of the time cannot go near French customers.
Jev misses the prompts that matter most.
Of the 21 frontier prompts, Jev on Q1 sent six to frontier and 15 to medium. On Q2 it sent three. Those are where routing cheap does the most damage: a statistics problem in Turkish, investment decision rules in German, an outline on quantum key distribution. Twenty-one is a small sample, and I am saying so first.
Jev’s probabilities help. Escalate to frontier whenever Jev gives it more than 0.2, and Q1 catches 14 of the 21 while escalating only 18 of 563 prompts. Q2 at the same cutoff catches seven. I picked 0.2 after seeing this data, so treat it as a starting point to tune.
Three more things cut against the headline:
- 84% flatters Jev. The gold set is the prompts both annotators agreed on, the easier ones. The annotators agreed only 77.5% of the time on the full set, so about 80% is a realistic ceiling for any router with these tiers.
- These are judgement labels. Two models guessed which tier a prompt needs. Nobody checked what each tier actually produced.
- Jev is slower. Median latency was about 330 ms per call from my laptop, mostly network, against 56 to 170 ms for the local models.
The cheap router is the expensive one.
We could not ship Jev in the end: it has no EU deployment. So I fine-tuned Laya on our own labels, and that router now runs on one CPU server in Germany.
A local router costs nothing per call. If it sends 41% of your traffic to a tier it does not need, as Von did on Q2, you pay for that on every one of those calls, and that cost appears on no model card I have read.
So count. Take WildChat or your own logs, label a few hundred prompts with two
annotators, drop the ones they disagree on, and score your router against
return "medium". If it cannot clear that line by ten points, you do not have a
router.
The better test, which we are running next, is outcome-based: push real traffic through every tier, record the cheapest tier whose answer was good enough, and score the router against that. If Jev’s 84% falls apart there, I will publish that too.