Jev, Laya, GLiNER: Beaten by len()

We tested Laya and GLiNER2, two open decision models, as zero-shot LLM routers on 500 real prompts, and neither beat a constant guess of "medium." Both got the same difficulty rubric we send TypeSafe's Jev in production. Laya scored 30.8% and 39.0% with no rank correlation to the labels, GLiNER2 scored 37.8%, always-medium scored 42.4%, and a one-line length rule scored 53.8%. The same GLiNER2 encoder with a trained logistic-regression head scored 58.8%, the best in the test, which matches our Jev canary design and an earlier GLiNER study on coding-agent turns: the encoder is the commodity, and the routing decision lives in the head you train. Includes a comparison of the three models, costs and latencies, and a Python recipe for turning any decision model into a cost-aware router.

Published 2026-09-28 by Dor Amir on the Nadir blog.

Filed under Routing & Cascades.

Abstract.

In September, three kinds of small model came up in the same conversation. Jev, TypeSafe AI's closed "System One" decision model, opened early access on September 15 at $0.042 per million input tokens. Laya, an Apache 2.0, 421M-parameter answer to it from Convai Innovations, shipped three days later, passed 26,000 GitHub stars in nine days, and reports beating Jev on accuracy, calibration, and latency. And GLiNER, the open encoder family that commenters cited within hours as having done schema-driven classification since 2023. All three take text plus a typed question and return a probability, not prose. That makes them look like the ideal LLM router: one fast forward pass, no output tokens, nothing to parse.

So we tested it. We gave Laya and GLiNER2 the same difficulty rubric we already send Jev in production, and asked each to label 500 real prompts as simple, medium, or complex, zero-shot. Neither beat a model that answers "medium" every time (42.4%). Laya scored 30.8% and 39.0%, with no correlation to the labels. GLiNER2 scored 37.8%. A one-line length rule scored 53.8%. The same GLiNER2 encoder with a trained logistic-regression head on top scored 58.8%, the best in the test. The lesson isn't that decision models are bad. It's that a decision model is an encoder with a good API, and the routing decision lives in the head you train on top of it.

All numbers in this post are from our own runs on public prompts, or from our own Jev canary where stated. Not derived from proprietary customer data. Sources cited throughout.

The three models.

JevLayaGLiNER2
WhoTypeSafe AI, $40M seedConvai Innovations, one developerFastino Labs, from the 2023 GLiNER line
ReleasedSeptember 15, 2026, early accessSeptember 18, 2026GLiNER2 paper July 2025 (arXiv:2507.18546)
LicenseClosed APIApache 2.0Apache 2.0
ArchitectureUnpublishedModernBERT-large (421M) English, mmBERT-base (322M) multilingualDeBERTa-style bidirectional encoder, labels and text in one pass
Question typesnoul (yes/no probability), choice, scoreSame three, same namesClassification, entities, relations, structured extraction
Price$0.042 per 1M input tokens, output free$0 per call, self-hosted$0 per call, self-hosted
Claimed latency70 to 500 ms32.8 ms p50 single question, on GPUCPU-friendly by design

Laya's own benchmark table reports 0.766 accuracy against Jev's 0.727 on a typed-decisions benchmark, 0.950 vs 0.910 on AG News, and calibration error of 0.081 against Jev's 0.246. It is also unusually honest about limits: choice questions degrade above 20 options (0.425 vs Jev's 0.870 on Banking77), and the base English checkpoint scores 0.362 zero-shot on the same typed-decisions benchmark where the fine-tuned checkpoint scores 0.766. Source: Laya, Hugging Face

Andreas Maier's September 28 essay frames all of this as "the return of the discriminative model," and makes the point that matters for anyone paying for inference: many production tasks treated as generation problems are classification tasks "paid for at generative prices and in generative latency, with an output that has to be parsed and may be invented." We made the same argument about Jev last week. Routing is the obvious example. Which model a prompt needs is a choice question.

The research question.

Can an off-the-shelf decision model, with no training on your traffic, decide which tier of LLM a prompt needs?

This matters because it's how these models are sold. Write the question in English, describe the options, get a calibrated probability. If that works for routing, a team can replace a trained difficulty classifier with a prompt to a free, open 421M model.

Methodology.

PromptsThe first 500 of a 1,280-prompt Chatbot Arena set we already use (eval/qwen_bench/prompts_1280.json)
LabelsSimple, medium, complex, assigned by an LLM judge. 37.2% / 42.4% / 20.4% on these 500
RubricThe exact tier_score and tier_choice questions we send Jev in our canary, with the same instructions and criteria
Layapip install laya 0.3.21, default Router(), English checkpoint, both the score and the choice question
GLiNER2fastino/gliner2.5-base-v1, zero-shot single-label classification with the three tier descriptions as labels
Trained headGLiNER2's encoder, mean-pooled, into a logistic regression. 5-fold cross-validation over all 1,280 prompts, scored on the same 500
BaselinesAlways "medium." A length rule fixed before scoring: under 60 characters is simple, over 300 is complex
Hardware4 vCPU cloud container, CPU only

Two caveats up front. The labels are an LLM's judgment of difficulty, not measured answer quality, and a length rule benefits from any length bias in that judgment. And CPU latencies aren't comparable to Laya's GPU numbers. The rubric, model versions, and thresholds above are everything needed to rerun it.

Finding 1: zero-shot, all three configurations lose to "medium."

Chart 1: Routing 500 prompts, zero-shot decision models versus a length rule. Accuracy and the share routed below its label. Always medium: 42.4% accuracy, 20.4% routed too cheap. Laya zero-shot with the score question: 30.8%, 9.8%. Laya zero-shot with the choice question: 39.0%, 57.2%. GLiNER2 zero-shot: 37.8%, 7.2%. Length rule: 53.8%, 14.0%. GLiNER2 encoder with a trained head: 58.8%, 21.4%.
Chart 1: Routing 500 prompts, zero-shot decision models versus a length rule. Accuracy and the share routed below its label. Always medium: 42.4% accuracy, 20.4% routed too cheap. Laya zero-shot with the score question: 30.8%, 9.8%. Laya zero-shot with the choice question: 39.0%, 57.2%. GLiNER2 zero-shot: 37.8%, 7.2%. Length rule: 53.8%, 14.0%. GLiNER2 encoder with a trained head: 58.8%, 21.4%.
PolicyAccuracyRouted too cheapComplex recallRank correlation with label
Always "medium"42.4%20.4%0%n/a
Laya zero-shot, score30.8%9.8%58.8%−0.003
Laya zero-shot, choice39.0%57.2%17.6%−0.006
GLiNER2 zero-shot37.8%7.2%71.6%0.299
Length rule53.8%14.0%62.7%0.562
GLiNER2 + trained head58.8%21.4%49.0%0.638

The Laya rows are the striking ones. The score question and the choice question ran in the same forward pass on the same prompts, and they disagree almost completely. Neither has any rank correlation with the labels: −0.003 and −0.006. That's not a weak signal. It's no signal. Laya reads our three-paragraph rubric and returns a confident-looking distribution that doesn't track difficulty at all.

This is consistent with what Laya says about itself. Its base checkpoint scores 0.362 zero-shot on its own typed-decisions benchmark and 0.766 after fine-tuning. Our result is the zero-shot half of that sentence, on a harder question.

GLiNER2 does better. It has real signal (0.299 rank correlation) and a low rate of routing prompts too cheap (7.2%), but it gets there by calling 59% of prompts complex. That's safe, and it saves almost nothing, which leads to the second finding.

Finding 2: a zero-shot router is a prior, not a decision.

Chart 2: What each model thinks the traffic looks like, as the share of 500 prompts predicted simple, medium and complex. Labels: 37%, 42%, 20%. Laya score: 4%, 42%, 54%. Laya choice: 92%, 0%, 7%. GLiNER2 zero-shot: 8%, 33%, 59%. Length rule: 25%, 47%, 29%. GLiNER2 with a trained head: 34%, 50%, 16%.
Chart 2: What each model thinks the traffic looks like, as the share of 500 prompts predicted simple, medium and complex. Labels: 37%, 42%, 20%. Laya score: 4%, 42%, 54%. Laya choice: 92%, 0%, 7%. GLiNER2 zero-shot: 8%, 33%, 59%. Length rule: 25%, 47%, 29%. GLiNER2 with a trained head: 34%, 50%, 16%.

A router's bill follows its predicted distribution, not its accuracy. Look at what each configuration thinks the traffic is:

The zero-shot outputs mostly reflect how each model reacts to the wording of the rubric, not to the prompt. Change "COMPLEX: substantial synthesis..." to something shorter and you'd get a different distribution from the same model. That's what we mean by a prior. A trained head learns where your labels put the boundaries. The zero-shot model has to guess from the wording.

Finding 3: the encoder is fine. Train the head.

The best result in the test used GLiNER2's encoder, the same weights that scored 37.8% zero-shot, with a logistic regression trained on 1,024 labeled prompts per fold. It reached 58.8% accuracy and a 0.638 rank correlation, the highest of anything we ran. No fine-tuning, no GPU, just a linear layer.

We've seen the same pattern on our own traffic twice.

What each one costs to run.

JevLayaGLiNER2
Cost per 1M routing decisionsabout $77 in our canary (four questions per call)$0 plus your hardware$0 plus your hardware
Latency we measured432 ms median for a fresh call, 129 ms on a cache hit1,710 ms median on shared CPU (vendor: 33 ms on GPU)513 ms median zero-shot on shared CPU
Runs air-gappedNoYesYes
Zero-shot routing on our 500not re-run here30.8% to 39.0%37.8%

The Jev cost comes from our canary: seven paid calls cost an estimated $0.000535626 in total. For comparison, answering the same routing question with a frontier LLM costs 49x to 476x more per decision. The encoder models are almost free per call. What isn't free is the labeled data to train a head, and that's the part nobody ships.

Tutorial: turn a decision model into a router.

The recipe that worked in every test above: use the encoder for features, train a small head on your own labels, and set the thresholds by cost, not accuracy.

Step 1: label a few hundred of your own prompts.

Sample 300 to 1,000 real prompts. Label each with the cheapest model that answered it well. That's the only label that matters for routing. An LLM judge is fine to start with, as long as you spot-check it.

Step 2: embed with the encoder.

import numpy as np, torch
import gliner2.classification as gc

clf = gc.Classifier.from_pretrained("fastino/gliner2.5-base-v1")
enc, tok = clf.model.encoder.eval(), clf.model.processor.tokenizer

@torch.no_grad()
def embed(texts: list[str], bs: int = 16) -> np.ndarray:
    out = []
    for i in range(0, len(texts), bs):
        b = tok([t[:2000] for t in texts[i:i + bs]], truncation=True, max_length=384,
                padding=True, return_tensors="pt")
        h = enc(**b).last_hidden_state
        m = b["attention_mask"].unsqueeze(-1)
        out.append(((h * m).sum(1) / m.sum(1)).numpy())   # mean-pool over real tokens
    return np.concatenate(out)

The same pattern works with Laya's backbone or any encoder you already run. The encoder is the commodity part.

Step 3: train the head, check it out of fold.

from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import StratifiedKFold, cross_val_predict
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

TIERS = ["simple", "medium", "complex"]
X = embed(prompts)
y = np.array([TIERS.index(t) for t in labels])

head = make_pipeline(StandardScaler(), LogisticRegression(C=0.05, max_iter=3000))
P = cross_val_predict(head, X, y, cv=StratifiedKFold(5, shuffle=True, random_state=0),
                      method="predict_proba")
print("accuracy", (P.argmax(1) == y).mean(), "too cheap", (P.argmax(1) < y).mean())
head.fit(X, y)   # the final model, trained on everything

Always compare against the two baselines in this post: a constant "medium" and a length rule. If your head doesn't beat both out of fold, you don't have a router yet.

Step 4: pick the tier by cost, not argmax.

Argmax treats "routed too cheap" and "routed too expensive" as the same mistake. They aren't. Price them:

COST = np.array([1.0, 3.0, 5.0])        # relative price of each tier's model
MISS = 8.0                              # cost of a bad answer, in the same units

def route(p: np.ndarray) -> int:
    """Pick the tier with the lowest expected cost: price plus P(the tier is too weak) x MISS."""
    too_weak = np.array([p[1:].sum(), p[2:].sum(), 0.0])
    return int(np.argmin(COST + too_weak * MISS))

Raising MISS trades savings for safety. That's the knob to tune on held-out data, the same way you'd sweep a router's confidence threshold.

When a zero-shot decision model is the right tool.

This post is about one hard question. Decision models are good at easier ones, and Laya's and Jev's benchmarks show it:

Where Nadir fits.

Prompt difficulty is the question Nadir exists to answer, and this post is why we don't answer it with a zero-shot rubric. POST /v1/bucket returns simple, medium, or complex with a probability for each and a confidence score. The classifier behind it is trained on labeled prompts, and the Jev hybrid we're testing uses Jev's answers as features for a trained head, the design Finding 3 supports. The call spends no provider tokens, and you can try it without a key.

If you've been thinking of pointing Laya or GLiNER at your routing problem, run the baselines and steps above on 500 of your own prompts first. If they beat len() on your traffic, great. If not, start with a free key and compare /v1/bucket on the same 500.

Conclusion.

Jev, Laya, and GLiNER are a real shift. The generative model is the wrong tool for a large class of production decisions, and a 421M encoder that answers in one pass, for free, under Apache 2.0, is a better one. But they're sold on the promise that you write the question and the model does the rest. For prompt routing, that isn't true yet. Zero-shot, both open models we tested scored below a constant guess, and one of them had no correlation with the labels at all. The same encoder with a trained linear head was the best thing we ran. Use these models for what they're good at: fast, cheap, private representations of text. Put the decision in a head trained on your own labels, and check it against len() before you trust it.


Benchmark results are from our own runs on public Chatbot Arena prompts with LLM-assigned labels, and from our internal Jev canary where stated. Not derived from customer data. Sources: [Laya, Convai Innovations](https://laya.convaiinnovations.com/), [Laya on Hugging Face](https://huggingface.co/convaiinnovations/laya). [Andreas Maier, "Laya, Jev and the Return of the Discriminative Model," September 28, 2026](https://akmaier.substack.com/p/laya-jev-and-the-return-of-the-discriminative). [eesel AI, "Laya AI: the open 33ms decision model"](https://www.eesel.ai/blog/laya-ai). [Zaratiana et al., "GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface," arXiv:2507.18546](https://arxiv.org/abs/2507.18546). [Fastino, GLiNER2 on GitHub](https://github.com/fastino-ai/GLiNER2). [Wikipedia, Jev (AI model)](https://en.wikipedia.org/wiki/Jev_(AI_model)).

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.