Abstract.
In September, three kinds of small model came up in the same conversation. Jev, TypeSafe AI's closed "System One" decision model, opened early access on September 15 at $0.042 per million input tokens. Laya, an Apache 2.0, 421M-parameter answer to it from Convai Innovations, shipped three days later, passed 26,000 GitHub stars in nine days, and reports beating Jev on accuracy, calibration, and latency. And GLiNER, the open encoder family that commenters cited within hours as having done schema-driven classification since 2023. All three take text plus a typed question and return a probability, not prose. That makes them look like the ideal LLM router: one fast forward pass, no output tokens, nothing to parse.
So we tested it. We gave Laya and GLiNER2 the same difficulty rubric we already send Jev in production, and asked each to label 500 real prompts as simple, medium, or complex, zero-shot. Neither beat a model that answers "medium" every time (42.4%). Laya scored 30.8% and 39.0%, with no correlation to the labels. GLiNER2 scored 37.8%. A one-line length rule scored 53.8%. The same GLiNER2 encoder with a trained logistic-regression head on top scored 58.8%, the best in the test. The lesson isn't that decision models are bad. It's that a decision model is an encoder with a good API, and the routing decision lives in the head you train on top of it.
All numbers in this post are from our own runs on public prompts, or from our own Jev canary where stated. Not derived from proprietary customer data. Sources cited throughout.
The three models.
| Jev | Laya | GLiNER2 | |
|---|---|---|---|
| Who | TypeSafe AI, $40M seed | Convai Innovations, one developer | Fastino Labs, from the 2023 GLiNER line |
| Released | September 15, 2026, early access | September 18, 2026 | GLiNER2 paper July 2025 (arXiv:2507.18546) |
| License | Closed API | Apache 2.0 | Apache 2.0 |
| Architecture | Unpublished | ModernBERT-large (421M) English, mmBERT-base (322M) multilingual | DeBERTa-style bidirectional encoder, labels and text in one pass |
| Question types | noul (yes/no probability), choice, score | Same three, same names | Classification, entities, relations, structured extraction |
| Price | $0.042 per 1M input tokens, output free | $0 per call, self-hosted | $0 per call, self-hosted |
| Claimed latency | 70 to 500 ms | 32.8 ms p50 single question, on GPU | CPU-friendly by design |
Laya's own benchmark table reports 0.766 accuracy against Jev's 0.727 on a typed-decisions benchmark, 0.950 vs 0.910 on AG News, and calibration error of 0.081 against Jev's 0.246. It is also unusually honest about limits: choice questions degrade above 20 options (0.425 vs Jev's 0.870 on Banking77), and the base English checkpoint scores 0.362 zero-shot on the same typed-decisions benchmark where the fine-tuned checkpoint scores 0.766. Source: Laya, Hugging Face
Andreas Maier's September 28 essay frames all of this as "the return of the discriminative model," and makes the point that matters for anyone paying for inference: many production tasks treated as generation problems are classification tasks "paid for at generative prices and in generative latency, with an output that has to be parsed and may be invented." We made the same argument about Jev last week. Routing is the obvious example. Which model a prompt needs is a choice question.
The research question.
Can an off-the-shelf decision model, with no training on your traffic, decide which tier of LLM a prompt needs?
This matters because it's how these models are sold. Write the question in English, describe the options, get a calibrated probability. If that works for routing, a team can replace a trained difficulty classifier with a prompt to a free, open 421M model.
Methodology.
| Prompts | The first 500 of a 1,280-prompt Chatbot Arena set we already use (eval/qwen_bench/prompts_1280.json) |
| Labels | Simple, medium, complex, assigned by an LLM judge. 37.2% / 42.4% / 20.4% on these 500 |
| Rubric | The exact tier_score and tier_choice questions we send Jev in our canary, with the same instructions and criteria |
| Laya | pip install laya 0.3.21, default Router(), English checkpoint, both the score and the choice question |
| GLiNER2 | fastino/gliner2.5-base-v1, zero-shot single-label classification with the three tier descriptions as labels |
| Trained head | GLiNER2's encoder, mean-pooled, into a logistic regression. 5-fold cross-validation over all 1,280 prompts, scored on the same 500 |
| Baselines | Always "medium." A length rule fixed before scoring: under 60 characters is simple, over 300 is complex |
| Hardware | 4 vCPU cloud container, CPU only |
Two caveats up front. The labels are an LLM's judgment of difficulty, not measured answer quality, and a length rule benefits from any length bias in that judgment. And CPU latencies aren't comparable to Laya's GPU numbers. The rubric, model versions, and thresholds above are everything needed to rerun it.
Finding 1: zero-shot, all three configurations lose to "medium."
| Policy | Accuracy | Routed too cheap | Complex recall | Rank correlation with label |
|---|---|---|---|---|
| Always "medium" | 42.4% | 20.4% | 0% | n/a |
| Laya zero-shot, score | 30.8% | 9.8% | 58.8% | −0.003 |
| Laya zero-shot, choice | 39.0% | 57.2% | 17.6% | −0.006 |
| GLiNER2 zero-shot | 37.8% | 7.2% | 71.6% | 0.299 |
| Length rule | 53.8% | 14.0% | 62.7% | 0.562 |
| GLiNER2 + trained head | 58.8% | 21.4% | 49.0% | 0.638 |
The Laya rows are the striking ones. The score question and the choice question ran in the same forward pass on the same prompts, and they disagree almost completely. Neither has any rank correlation with the labels: −0.003 and −0.006. That's not a weak signal. It's no signal. Laya reads our three-paragraph rubric and returns a confident-looking distribution that doesn't track difficulty at all.
This is consistent with what Laya says about itself. Its base checkpoint scores 0.362 zero-shot on its own typed-decisions benchmark and 0.766 after fine-tuning. Our result is the zero-shot half of that sentence, on a harder question.
GLiNER2 does better. It has real signal (0.299 rank correlation) and a low rate of routing prompts too cheap (7.2%), but it gets there by calling 59% of prompts complex. That's safe, and it saves almost nothing, which leads to the second finding.
Finding 2: a zero-shot router is a prior, not a decision.
A router's bill follows its predicted distribution, not its accuracy. Look at what each configuration thinks the traffic is:
- Laya, score: 54% complex, 4% simple. Most traffic goes to the flagship. The router costs you a forward pass and saves nothing.
- Laya, choice: 92% simple. This would cut the bill dramatically, and send 57% of prompts to a model below their label. That's the most expensive kind of saving.
- GLiNER2, zero-shot: 59% complex. Safe and nearly useless for cost.
- Trained head: 34% / 50% / 16%, close to the label mix of 37% / 42% / 20%.
The zero-shot outputs mostly reflect how each model reacts to the wording of the rubric, not to the prompt. Change "COMPLEX: substantial synthesis..." to something shorter and you'd get a different distribution from the same model. That's what we mean by a prior. A trained head learns where your labels put the boundaries. The zero-shot model has to guess from the wording.
Finding 3: the encoder is fine. Train the head.
The best result in the test used GLiNER2's encoder, the same weights that scored 37.8% zero-shot, with a logistic regression trained on 1,024 labeled prompts per fold. It reached 58.8% accuracy and a 0.638 rank correlation, the highest of anything we ran. No fine-tuning, no GPU, just a linear layer.
We've seen the same pattern on our own traffic twice.
- GLiNER on coding-agent turns. In a September 24 internal study on 6,641 Claude Code turns, zero-shot GLiNER scored below a constant "always implement" baseline on task type at every model size. Fine-tuned, GLiNER-base predicted turn size (tool calls per turn) at a 0.58 rank correlation on new sessions and 0.56 on unseen repositories, ahead of a fine-tuned ModernBERT at 0.43 and 0.46.
- Jev in our router. Our Jev canary doesn't route on Jev's answer directly. It asks Jev four rubric questions, feeds the answers into a small trained head alongside our own student classifier, and routes on that. On a separate 497-case judgment set, the frozen hybrid reached 73.44% tier accuracy, routed 5.23% of prompts below their label, and caught 87% of complex prompts. Different labels, so don't compare those numbers with the table above. The shape of the design is the point: Jev's answers are features, and the decision comes from a head we trained. Customer rollout is 0% while we measure it.
What each one costs to run.
| Jev | Laya | GLiNER2 | |
|---|---|---|---|
| Cost per 1M routing decisions | about $77 in our canary (four questions per call) | $0 plus your hardware | $0 plus your hardware |
| Latency we measured | 432 ms median for a fresh call, 129 ms on a cache hit | 1,710 ms median on shared CPU (vendor: 33 ms on GPU) | 513 ms median zero-shot on shared CPU |
| Runs air-gapped | No | Yes | Yes |
| Zero-shot routing on our 500 | not re-run here | 30.8% to 39.0% | 37.8% |
The Jev cost comes from our canary: seven paid calls cost an estimated $0.000535626 in total. For comparison, answering the same routing question with a frontier LLM costs 49x to 476x more per decision. The encoder models are almost free per call. What isn't free is the labeled data to train a head, and that's the part nobody ships.
Tutorial: turn a decision model into a router.
The recipe that worked in every test above: use the encoder for features, train a small head on your own labels, and set the thresholds by cost, not accuracy.
Step 1: label a few hundred of your own prompts.
Sample 300 to 1,000 real prompts. Label each with the cheapest model that answered it well. That's the only label that matters for routing. An LLM judge is fine to start with, as long as you spot-check it.
Step 2: embed with the encoder.
import numpy as np, torch
import gliner2.classification as gc
clf = gc.Classifier.from_pretrained("fastino/gliner2.5-base-v1")
enc, tok = clf.model.encoder.eval(), clf.model.processor.tokenizer
@torch.no_grad()
def embed(texts: list[str], bs: int = 16) -> np.ndarray:
out = []
for i in range(0, len(texts), bs):
b = tok([t[:2000] for t in texts[i:i + bs]], truncation=True, max_length=384,
padding=True, return_tensors="pt")
h = enc(**b).last_hidden_state
m = b["attention_mask"].unsqueeze(-1)
out.append(((h * m).sum(1) / m.sum(1)).numpy()) # mean-pool over real tokens
return np.concatenate(out)
The same pattern works with Laya's backbone or any encoder you already run. The encoder is the commodity part.
Step 3: train the head, check it out of fold.
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import StratifiedKFold, cross_val_predict
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
TIERS = ["simple", "medium", "complex"]
X = embed(prompts)
y = np.array([TIERS.index(t) for t in labels])
head = make_pipeline(StandardScaler(), LogisticRegression(C=0.05, max_iter=3000))
P = cross_val_predict(head, X, y, cv=StratifiedKFold(5, shuffle=True, random_state=0),
method="predict_proba")
print("accuracy", (P.argmax(1) == y).mean(), "too cheap", (P.argmax(1) < y).mean())
head.fit(X, y) # the final model, trained on everything
Always compare against the two baselines in this post: a constant "medium" and a length rule. If your head doesn't beat both out of fold, you don't have a router yet.
Step 4: pick the tier by cost, not argmax.
Argmax treats "routed too cheap" and "routed too expensive" as the same mistake. They aren't. Price them:
COST = np.array([1.0, 3.0, 5.0]) # relative price of each tier's model
MISS = 8.0 # cost of a bad answer, in the same units
def route(p: np.ndarray) -> int:
"""Pick the tier with the lowest expected cost: price plus P(the tier is too weak) x MISS."""
too_weak = np.array([p[1:].sum(), p[2:].sum(), 0.0])
return int(np.argmin(COST + too_weak * MISS))
Raising MISS trades savings for safety. That's the knob to tune on held-out data, the same way you'd sweep a router's confidence threshold.
When a zero-shot decision model is the right tool.
This post is about one hard question. Decision models are good at easier ones, and Laya's and Jev's benchmarks show it:
- Few options with clear names. Department routing, language detection, yes/no policy checks. Laya's AG News result (0.950) is this kind of task.
- Questions where the answer is in the text. "Does the user threaten to cancel?" is a reading task. "How hard is this to answer well?" isn't. It needs a model of what the other models can do.
- As features, not answers. Ask a decision model several narrow questions (is this code, does it need multiple steps, is there a numeric answer) and let a trained head combine them. That's how our Jev canary is built.
Where Nadir fits.
Prompt difficulty is the question Nadir exists to answer, and this post is why we don't answer it with a zero-shot rubric. POST /v1/bucket returns simple, medium, or complex with a probability for each and a confidence score. The classifier behind it is trained on labeled prompts, and the Jev hybrid we're testing uses Jev's answers as features for a trained head, the design Finding 3 supports. The call spends no provider tokens, and you can try it without a key.
If you've been thinking of pointing Laya or GLiNER at your routing problem, run the baselines and steps above on 500 of your own prompts first. If they beat len() on your traffic, great. If not, start with a free key and compare /v1/bucket on the same 500.
Conclusion.
Jev, Laya, and GLiNER are a real shift. The generative model is the wrong tool for a large class of production decisions, and a 421M encoder that answers in one pass, for free, under Apache 2.0, is a better one. But they're sold on the promise that you write the question and the model does the rest. For prompt routing, that isn't true yet. Zero-shot, both open models we tested scored below a constant guess, and one of them had no correlation with the labels at all. The same encoder with a trained linear head was the best thing we ran. Use these models for what they're good at: fast, cheap, private representations of text. Put the decision in a head trained on your own labels, and check it against len() before you trust it.
Benchmark results are from our own runs on public Chatbot Arena prompts with LLM-assigned labels, and from our internal Jev canary where stated. Not derived from customer data. Sources: [Laya, Convai Innovations](https://laya.convaiinnovations.com/), [Laya on Hugging Face](https://huggingface.co/convaiinnovations/laya). [Andreas Maier, "Laya, Jev and the Return of the Discriminative Model," September 28, 2026](https://akmaier.substack.com/p/laya-jev-and-the-return-of-the-discriminative). [eesel AI, "Laya AI: the open 33ms decision model"](https://www.eesel.ai/blog/laya-ai). [Zaratiana et al., "GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface," arXiv:2507.18546](https://arxiv.org/abs/2507.18546). [Fastino, GLiNER2 on GitHub](https://github.com/fastino-ai/GLiNER2). [Wikipedia, Jev (AI model)](https://en.wikipedia.org/wiki/Jev_(AI_model)).