# Train your own Jev: an 11x faster System-1 model with almost free hosting

> Raya's training code is public. Two LLMs label your data, a 300M encoder learns it in minutes, ONNX serves it on a CPU. I ran the recipe on a laptop.

- Source: https://jays.fyi/blog/train-your-own-system-1-model
- Author: Jay Derinbogaz
- Published: 2026-09-27
- Language: en
- Tags: ai, open-source, tooling
- Reading time: 7 min

---

On Friday I wrote about
[Raya](/blog/fine-tuned-llm-router-on-a-hetzner-cpu-server), the router that
picks a model tier for every prompt in TextCortex. It scores 81% on our
563-prompt benchmark and answers in 30 ms on a CPU in Falkenstein. That post did not
say how to build one.

The recipe is now public, in the
[`training/` folder](https://huggingface.co/TextCortex/raya/blob/main/training/README.md)
of Raya's Hugging Face repo, under Apache-2.0. Describe the decision, have two
LLMs label your data, train, evaluate, export to ONNX. I ran every step except
labelling on an Apple M4 laptop this morning to check that it works as
written. Training took **34 seconds** and the export took 24.

Here is the claim. If you send a frontier model the same question with a fixed
set of answers on every request, that question belongs in a small classifier
you own. Use the big model once, as the teacher, and stop paying it per
decision.

## You are paying an LLM to answer the same question a million times.

The README calls the result a **System-1 model**, and I will keep the term. A
System-1 model answers one well-defined question about an input, instantly,
with probabilities you can threshold. *Which model should answer this prompt?*
*Does this ticket need a human?* *Which team owns this request?* It is a
fine-tuned [Laya](https://huggingface.co/convaiinnovations/laya) encoder of
about 300 million parameters, it runs in tens of milliseconds, and it has no
per-call fee.

Calling a chat model for these is hiring a lawyer to sort your mail: a bill
per envelope, a wait on each, and a copy of your mail on somebody else's
server. For a European company
that last part can end the conversation: we built Raya because the router we
wanted had no EU deployment.

The wait is easy to measure. Raya answers in 30 ms at p50 in production. When
I measured Jev, the hosted router Raya replaces, it took about 350 ms a call:
Raya is roughly **11 times faster**. It runs on a €84-a-month Hetzner server
that was already on our bill, so the hosting cost us nothing new.

## The whole recipe is four scripts and a JSON file.

<figure>
  <img src="/images/raya-training-recipe-pipeline.png" alt="The five steps, task.json to export_onnx.py, with what each took for Raya and for a toy run on an Apple M4: 34 seconds of training, 24 of export." width="1600" height="860" decoding="async" fetchpriority="high" />
  <figcaption>Raya's figures are from the <a href="https://huggingface.co/TextCortex/raya/blob/main/training/README.md">training README</a> and the <a href="/blog/fine-tuned-llm-router-on-a-hetzner-cpu-server">previous post</a>. The toy run used the 44 demo prompts that ship with the code.</figcaption>
</figure>

```bash
python label.py --task my_task.json --data prompts.jsonl --out labelled.jsonl \
    --annotator <model-a> --annotator <model-b>@https://api.anthropic.com/v1/#ANTHROPIC_API_KEY
python train.py --task my_task.json --data labelled.jsonl --out my-model
python evaluate.py --model my-model --task my_task.json --data test.jsonl
python export_onnx.py --model my-model --task my_task.json --data test.jsonl
```

Before any of that you write `task.json`: the labels, and one or more
*phrasings* of the question. Raya has three: "Route this prompt to a model", a
longer version with a rubric per tier, and a 1-to-3 difficulty score. The
model trains on every phrasing with the options reshuffled each epoch, so it
learns the decision itself. The file also holds the rubric the annotators
read, and
[Raya's own](https://huggingface.co/TextCortex/raya/blob/main/training/task.example.json)
ships as the template.

## Two annotators who disagree are worth more than one who sounds sure.

You probably have inputs and no labels. `label.py` sends each input to two or
more LLMs, separately, with the same rubric. Raya's roughly 10,000 labels came
from Claude Opus and Claude Sonnet, blind to each other. Any OpenAI-compatible
endpoint works, including Anthropic's, OpenRouter, vLLM and Ollama.

Where the annotators agree you get a clean label. Where they disagree the row
keeps both votes and `train.py` learns a 50/50 target. That is the honest
answer for a prompt two strong models split on. Forcing a hard label there
teaches the model a coin flip as if it were a fact.

The script also prints how often the annotators agreed. **Write that number
down.** It is roughly the ceiling of what your model can score against these
labels. For Raya it was 78%. If your annotators agree 75% of the time and your
model scores 90%, it has learned one annotator's habits, and the fix is a
tighter rubric.

On volume, the README says 1,000 to 5,000 real inputs for a first model, in
the languages you actually serve. Leave the classes unbalanced; `train.py`
weights rare labels itself. Keep a test set
that nothing ever trains or validates on.

## The GPU is the cheap part.

Raya trained in about six minutes on one 48 GB RTX A6000. On my M4 it ran two epochs over the 40 toy
training rows in 22 seconds, 34 with model loading.

It freezes the token-embedding table to save memory, picks the best epoch on
validation, then fits a temperature per question so that a probability of 0.9
is right about nine times in ten. That calibration is what lets you act on a
threshold later, such as sending a prompt to the frontier tier only above 0.6.

The default starting point is Laya's multilingual mmBERT-base. Pass
`--base TextCortex/raya` to adapt Raya to your own routing traffic instead.

<figure>
  <img src="/images/raya-accuracy-after-fine-tuning.png" alt="Bar chart of routing accuracy: always medium 56.3%, stock Laya 61.6%, Raya 81.0%, Jev 84.5%." width="1600" height="860" decoding="async" loading="lazy" />
  <figcaption>Best of three question styles on the 563-prompt benchmark from the <a href="/blog/local-llm-routers-lost-to-always-answering-medium">first post</a>.</figcaption>
</figure>

Six minutes of GPU time took Laya from 61.6% to 81%, three and a half points
short of Jev. The labels are where your afternoon goes.

## The toy data proves the plumbing works and nothing else.

The folder includes 44 hand-written routing prompts for training and 13 for
testing. The README promises five minutes. Training and export took under a
minute, not counting the install, on the pinned versions. The README's example call routed "Prove that √2 is irrational."
to `frontier_model` at 0.52.

`evaluate.py` then scored the toy model at 46.2% on the short phrasing, 84.6%
on the rubric phrasing and 69.2% on the difficulty score. The short phrasing
sent 12 of 13 prompts to medium. Thirteen prompts cannot separate those
numbers from noise, and 44 training rows is, in the README's words, "far too
small to train a useful model".

The useful thing `evaluate.py` did was print this line first:

```text
13 rows with a gold label; always answering 'small_model' scores 38.5%
```

That is the constant baseline. In the
[first benchmark](/blog/local-llm-routers-lost-to-always-answering-medium),
stock Laya lost to `return "medium"` on two of its three question styles.
Every evaluation should open with the score of the model that does nothing.

## Export once, then measure on the CPU you will actually rent.

`export_onnx.py` writes an fp32 file and a block-wise int8 file, then checks
both against PyTorch on your test data. The export fails if any choice changes
or if probabilities drift by more than 0.001 for fp32 or 0.05 for int8. My toy
model came out at 0.00000 and 0.02790.

Which file to ship depends on silicon. As the
[last post](/blog/fine-tuned-llm-router-on-a-hetzner-cpu-server) measured,
block-wise int8 kept Raya's accuracy and ran fastest on an i9-13900 with VNNI
instructions, but took more than twice as long as fp32 on an EPYC 7502P without them.
Measure on your own hardware before you pick.

## Rent the big model once, then stop paying it per decision.

The frontier-model industry prices every call, so every call looks like a job
for a frontier model. Any call that picks from a fixed list of answers is a
classification, billed as reasoning. After the labels are written, the LLM is an expensive way to reach a verdict a 300M encoder
reaches in 30 ms on hardware you already pay for.

There is a limit I owe you. Raya's training data is not published, so you can
build your own Raya with this folder and cannot reproduce ours. And Raya
learned from the same two annotators who wrote the benchmark labels, which
flatters its 81%.

So measure it on your own traffic. Run `evaluate.py` on a test set your model
never saw. If it does not clear the constant baseline by a wide margin,
tighten the rubric before you touch anything else. If the recipe fails on your
decision, send me the numbers and I will publish them.
