On Friday I wrote about Raya, the router that picks a model tier for every prompt in TextCortex. It scores 81% on our 563-prompt benchmark and answers in 30 ms on a CPU in Falkenstein. That post did not say how to build one.
The recipe is now public, in the
training/ folder
of Raya’s Hugging Face repo, under Apache-2.0. Describe the decision, have two
LLMs label your data, train, evaluate, export to ONNX. I ran every step except
labelling on an Apple M4 laptop this morning to check that it works as
written. Training took 34 seconds and the export took 24.
Here is the claim. If you send a frontier model the same question with a fixed set of answers on every request, that question belongs in a small classifier you own. Use the big model once, as the teacher, and stop paying it per decision.
You are paying an LLM to answer the same question a million times.
The README calls the result a System-1 model, and I will keep the term. A System-1 model answers one well-defined question about an input, instantly, with probabilities you can threshold. Which model should answer this prompt? Does this ticket need a human? Which team owns this request? It is a fine-tuned Laya encoder of about 300 million parameters, it runs in tens of milliseconds, and it has no per-call fee.
Calling a chat model for these is hiring a lawyer to sort your mail: a bill per envelope, a wait on each, and a copy of your mail on somebody else’s server. For a European company that last part can end the conversation: we built Raya because the router we wanted had no EU deployment.
The wait is easy to measure. Raya answers in 30 ms at p50 in production. When I measured Jev, the hosted router Raya replaces, it took about 350 ms a call: Raya is roughly 11 times faster. It runs on a €84-a-month Hetzner server that was already on our bill, so the hosting cost us nothing new.
The whole recipe is four scripts and a JSON file.
python label.py --task my_task.json --data prompts.jsonl --out labelled.jsonl \
--annotator <model-a> --annotator <model-b>@https://api.anthropic.com/v1/#ANTHROPIC_API_KEY
python train.py --task my_task.json --data labelled.jsonl --out my-model
python evaluate.py --model my-model --task my_task.json --data test.jsonl
python export_onnx.py --model my-model --task my_task.json --data test.jsonl
Before any of that you write task.json: the labels, and one or more
phrasings of the question. Raya has three: “Route this prompt to a model”, a
longer version with a rubric per tier, and a 1-to-3 difficulty score. The
model trains on every phrasing with the options reshuffled each epoch, so it
learns the decision itself. The file also holds the rubric the annotators
read, and
Raya’s own
ships as the template.
Two annotators who disagree are worth more than one who sounds sure.
You probably have inputs and no labels. label.py sends each input to two or
more LLMs, separately, with the same rubric. Raya’s roughly 10,000 labels came
from Claude Opus and Claude Sonnet, blind to each other. Any OpenAI-compatible
endpoint works, including Anthropic’s, OpenRouter, vLLM and Ollama.
Where the annotators agree you get a clean label. Where they disagree the row
keeps both votes and train.py learns a 50/50 target. That is the honest
answer for a prompt two strong models split on. Forcing a hard label there
teaches the model a coin flip as if it were a fact.
The script also prints how often the annotators agreed. Write that number down. It is roughly the ceiling of what your model can score against these labels. For Raya it was 78%. If your annotators agree 75% of the time and your model scores 90%, it has learned one annotator’s habits, and the fix is a tighter rubric.
On volume, the README says 1,000 to 5,000 real inputs for a first model, in
the languages you actually serve. Leave the classes unbalanced; train.py
weights rare labels itself. Keep a test set
that nothing ever trains or validates on.
The GPU is the cheap part.
Raya trained in about six minutes on one 48 GB RTX A6000. On my M4 it ran two epochs over the 40 toy training rows in 22 seconds, 34 with model loading.
It freezes the token-embedding table to save memory, picks the best epoch on validation, then fits a temperature per question so that a probability of 0.9 is right about nine times in ten. That calibration is what lets you act on a threshold later, such as sending a prompt to the frontier tier only above 0.6.
The default starting point is Laya’s multilingual mmBERT-base. Pass
--base TextCortex/raya to adapt Raya to your own routing traffic instead.
Six minutes of GPU time took Laya from 61.6% to 81%, three and a half points short of Jev. The labels are where your afternoon goes.
The toy data proves the plumbing works and nothing else.
The folder includes 44 hand-written routing prompts for training and 13 for
testing. The README promises five minutes. Training and export took under a
minute, not counting the install, on the pinned versions. The README’s example call routed “Prove that √2 is irrational.”
to frontier_model at 0.52.
evaluate.py then scored the toy model at 46.2% on the short phrasing, 84.6%
on the rubric phrasing and 69.2% on the difficulty score. The short phrasing
sent 12 of 13 prompts to medium. Thirteen prompts cannot separate those
numbers from noise, and 44 training rows is, in the README’s words, “far too
small to train a useful model”.
The useful thing evaluate.py did was print this line first:
13 rows with a gold label; always answering 'small_model' scores 38.5%
That is the constant baseline. In the
first benchmark,
stock Laya lost to return "medium" on two of its three question styles.
Every evaluation should open with the score of the model that does nothing.
Export once, then measure on the CPU you will actually rent.
export_onnx.py writes an fp32 file and a block-wise int8 file, then checks
both against PyTorch on your test data. The export fails if any choice changes
or if probabilities drift by more than 0.001 for fp32 or 0.05 for int8. My toy
model came out at 0.00000 and 0.02790.
Which file to ship depends on silicon. As the last post measured, block-wise int8 kept Raya’s accuracy and ran fastest on an i9-13900 with VNNI instructions, but took more than twice as long as fp32 on an EPYC 7502P without them. Measure on your own hardware before you pick.
Rent the big model once, then stop paying it per decision.
The frontier-model industry prices every call, so every call looks like a job for a frontier model. Any call that picks from a fixed list of answers is a classification, billed as reasoning. After the labels are written, the LLM is an expensive way to reach a verdict a 300M encoder reaches in 30 ms on hardware you already pay for.
There is a limit I owe you. Raya’s training data is not published, so you can build your own Raya with this folder and cannot reproduce ours. And Raya learned from the same two annotators who wrote the benchmark labels, which flatters its 81%.
So measure it on your own traffic. Run evaluate.py on a test set your model
never saw. If it does not clear the constant baseline by a wide margin,
tighten the rubric before you touch anything else. If the recipe fails on your
decision, send me the numbers and I will publish them.