Where the Label Lives

Jev, Laya, GLiNER and a plain BERT classifier all return a decision instead of prose, and the architecture decides how they scale. The difference is where the label lives: in the weights (BERT with a trained head), one label per forward pass (NLI cross-encoders), every label written into the input (GLiNER2 and Laya), or unpublished (Jev). We ran four open architectures on AG News (4 labels) and Banking77 (77 labels) on the same CPU. Laya led on 4 labels at 95.5% and fell to 19.5% on 77. GLiNER2 did the reverse, 74.0% then 75.5%. A frozen ModernBERT with a linear head came second on both and its latency stayed flat as labels grew, while the cross-encoder's grew with every label. On a per-token API the option list is most of the bill: 399 of 406 input tokens on a Banking77 request. Includes the same decision in four code snippets and a benchmark harness.

Published 2026-10-01 by Dor Amir on the Nadir blog.

Filed under Routing & Cascades.

Abstract.

"System One" decision models are the fastest-moving category in inference this month. TypeSafe's Jev opened early access on September 15. Laya, an open 421M-parameter answer to it, shipped three days later. And GLiNER's fans spent the week pointing out that schema-driven classification has existed since 2023. All of them take text plus a question and return a probability instead of prose. People compare them on price and on benchmark scores. The difference that predicts how they behave on your workload is architectural, and it comes down to one question: where does the label live? In a classic BERT classifier, the labels live in the weights. In an NLI cross-encoder, each label is a separate input. In GLiNER and Laya, every label is written into the same input as the text. In an LLM, the labels sit in the prompt and the answer is generated. That one choice sets what a new label costs, how latency grows with the number of options, and what you're billed for. We ran four open architectures on the same two tasks, one with 4 labels and one with 77, on the same CPU. No single architecture won both.

All accuracy and latency numbers are from our own runs on public datasets (AG News and Banking77), 200 test examples each, on a 4 vCPU cloud container. Jev is described from TypeSafe's public statements and was not re-run. Not derived from customer data.

Six ways to answer a three-option question.

Chart 1: Where the label lives, for a three-label question. BERT plus a trained head: the input is only the text, and a linear layer with three fixed outputs gives the answer in one pass, so a new label means retraining. NLI cross-encoder: the text plus one label per input, scored for entailment, three passes for three labels. GLiNER2: a task token, then each label behind an L token, a separator, then the text, all in one pass, with an MLP on each label token, so the input grows with the number of labels. Laya: a typed question, then each option behind a MASK marker token, then the text, one pass per question, with a scorer on each marker. Jev: architecture unpublished, TypeSafe says all outputs come from a single query and questions run in parallel. Generative LLM: labels in the prompt, the answer decoded token by token and parsed afterwards.
Chart 1: Where the label lives, for a three-label question. BERT plus a trained head: the input is only the text, and a linear layer with three fixed outputs gives the answer in one pass, so a new label means retraining. NLI cross-encoder: the text plus one label per input, scored for entailment, three passes for three labels. GLiNER2: a task token, then each label behind an L token, a separator, then the text, all in one pass, with an MLP on each label token, so the input grows with the number of labels. Laya: a typed question, then each option behind a MASK marker token, then the text, one pass per question, with a scorer on each marker. Jev: architecture unpublished, TypeSafe says all outputs come from a single query and questions run in parallel. Generative LLM: labels in the prompt, the answer decoded token by token and parsed afterwards.

All of these except the last are encoders: the text is read once, in both directions, and every output is produced in the same forward pass. Nothing is generated, so nothing has to be parsed. What separates them is where the candidate answers go.

1. BERT with a trained head: labels in the weights. The original pattern. Run the text through an encoder (BERT, DeBERTa, ModernBERT), pool the result into one vector, and multiply by a weight matrix with one column per label. The labels aren't in the input at all. They're the columns. That makes inference the cheapest of anything here, and it makes the label set fixed: a new label is a new column, which needs labeled examples and a retrain.

2. NLI cross-encoder: one label per pass. The zero-shot trick from 2019. Take a model trained on natural language inference, pair the text with a hypothesis like "This text is about billing," and read the entailment score. Labels are now free text you can change at run time. The cost is that each label is its own forward pass. Three labels, three passes. Seventy-seven labels, seventy-seven.

3. GLiNER and GLiNER2: all labels in one input. GLiNER (Zaratiana et al., arXiv:2311.08526) fixed the NLI scaling problem for entity extraction by writing every entity type into the same sequence as the text, each behind an [ENT] token, on a DeBERTa-v3 backbone. The bidirectional encoder lets every label token attend to the text and vice versa, and a small network scores each label against each span. GLiNER2 (arXiv:2507.18546), a 205M-parameter encoder, extends that to classification: each label sits behind an [L] token, and an MLP turns each [L] token's output into a logit. One pass, any labels, defined at run time. The paper reports 163 ms on CPU for a 20-label task against 6,758 ms for a DeBERTa NLI baseline.

4. Laya: typed questions on the GLiNER pattern. Laya's open-source package documents its input format: [CLS] <type> instructions [SEP] [MASK] opt0 [MASK] opt1 ... [SEP] state [SEP], on a ModernBERT-large encoder (421M). A scorer reads each [MASK] marker, and separate heads serve the three question types it shares with Jev: yes/no (noul), choice, and score. It's trained with reinforcement learning for calibrated probabilities. Architecturally, it's GLiNER's label-in-the-input design with typed outputs and a different training objective. One consequence: each question is its own sequence, and the options share a token budget (head_max_len, 192 by default) with the question text.

5. Jev: unpublished. TypeSafe has said Jev is transformer-based, trained on synthetic data with a method it calls Reinforcement Learning for Calibrated Decisions (RLCD), and "generates all outputs in a single query" rather than one token at a time, evaluating several questions in parallel. It has not published the architecture, the parameter count, or a paper. Outside analyses of the API's behavior have guessed at a repurposed causal transformer with a shared state encoding, but that's speculation, and we treat it as such. What we do know is the billing: $0.042 per million input tokens, output free.

6. A generative LLM: labels in the prompt. List the options in the prompt, generate the answer token by token, parse the text. It's the most flexible, it handles questions none of the others can, and it's the most expensive per decision by one to three orders of magnitude. We've priced that gap before.

The benchmark.

The research question: does the label's position predict how a model handles a task with few options versus many?

TasksAG News (4 topics: world, sports, business, science and technology) and Banking77 (77 customer-banking intents)
Test set200 examples from each test split, fixed seed
BERT + headModernBERT-base (149M), mean-pooled, logistic regression trained on 2,000 random training examples. Frozen encoder, no fine-tuning
NLI cross-encoderMoritzLaurer/deberta-v3-base-zeroshot-v2.0, hypothesis "This text is about {label}."
GLiNER2fastino/gliner2.5-base-v1, single-label classification
Layapip install laya 0.3.21, English checkpoint, one choice question. Banking77 needed head_max_len=448 to fit 77 options
LabelsThe dataset's own label names, lightly cleaned (card_arrival becomes "card arrival"). No descriptions
Hardware4 vCPU, CPU only, batch size 1, one model at a time

Three of the four are zero-shot. The trained head saw 2,000 labeled examples, which is about 500 per class on AG News and only about 26 per class on Banking77. Treat it as the "you have some labels" baseline, not as a fully fine-tuned model.

Finding 1: nothing wins both.

Chart 2: Accuracy on 200 test examples. AG News with 4 labels: Laya 95.5%, BERT plus trained head 93.5%, NLI cross-encoder 92.5%, GLiNER2 74.0%. Banking77 with 77 labels: GLiNER2 75.5%, BERT plus trained head 74.0%, NLI cross-encoder 63.0%, Laya 19.5%.
Chart 2: Accuracy on 200 test examples. AG News with 4 labels: Laya 95.5%, BERT plus trained head 93.5%, NLI cross-encoder 92.5%, GLiNER2 74.0%. Banking77 with 77 labels: GLiNER2 75.5%, BERT plus trained head 74.0%, NLI cross-encoder 63.0%, Laya 19.5%.
ArchitectureLabels liveAG News (4)Banking77 (77)p50 latency, AGp50 latency, Banking77
BERT + trained headweights93.5%74.0%52 ms32 ms
NLI cross-encoderone per pass92.5%63.0%168 ms1,437 ms
GLiNER2in the input74.0%75.5%85 ms232 ms
Layain the input, typed95.5%19.5%198 ms766 ms

Laya is the best model in the test on four options and the worst on seventy-seven. That's consistent with what its own documentation says: choice questions degrade above about 20 options, and it reports 0.425 on Banking77 where its table puts Jev at 0.870. Our 19.5% is lower than Laya's own figure, probably because we gave it bare label names, not descriptions. Two more caveats on the AG News number. Laya's public benchmark includes AG News, so we can't rule out that the dataset was in its training data. And on Banking77 it was confidently wrong: "How do I locate my card?" (labeled card_arrival) came back as lost_or_stolen_card at 0.95 confidence.

GLiNER2 is the opposite. It's the weakest on AG News, where "world news" and "business" overlap and a label name carries little information. It's the best on Banking77, where the 77 label names are specific ("exchange rate", "pin blocked", "top up failed") and reading them alongside the text is most of the job. The GLiNER2 paper reports 0.70 zero-shot on Banking77, close to our 75.5%.

The trained head is the steady one. It's second on both tasks, behind whichever zero-shot model happened to fit the task, and it got there with 26 examples per intent on Banking77 and no fine-tuning.

Finding 2: where the label lives sets the latency curve.

To isolate the architecture from the dataset, we took 30 Banking77 messages and asked each model to choose among the first K intents, for K from 2 to 77.

Chart 3: Median CPU latency per decision as the number of labels grows from 2 to 77, on 30 Banking77 messages. BERT plus a trained head stays flat at about 31 to 36 ms. GLiNER2 and Laya grow with the input as labels are added. The NLI cross-encoder grows linearly, one pass per label, to the slowest at 77 labels.
Chart 3: Median CPU latency per decision as the number of labels grows from 2 to 77, on 30 Banking77 messages. BERT plus a trained head stays flat at about 31 to 36 ms. GLiNER2 and Laya grow with the input as labels are added. The NLI cross-encoder grows linearly, one pass per label, to the slowest at 77 labels.
Labels (K)2481632647777 vs 2
BERT + trained head36 ms33 ms32 ms33 ms31 ms32 ms30 ms0.8x
GLiNER274 ms71 ms85 ms92 ms125 ms186 ms237 ms3.2x
Laya145 ms174 ms220 ms291 ms499 ms802 ms823 ms5.7x
NLI cross-encoder79 ms108 ms202 ms323 ms592 ms1,198 ms1,466 ms18.5x

At two labels, the NLI cross-encoder is about as fast as GLiNER2. By seventy-seven, it's six times slower. Laya starts slower than the cross-encoder and ends at about half its latency, but grows more steeply than GLiNER2: its sequences are longer, a typed question plus a [MASK] marker per option, on an encoder twice the size.

The shapes follow directly from Chart 1:

Finding 3: per-token pricing bills you for the options.

That last point changes what a hosted decision model costs. Laya reports its own token usage, and on a Banking77 request it was 406 input tokens, of which 7 were the customer's message. The other 399 were the question and the 77 options. On a label-in-the-input architecture, the option list is the request.

Jev bills $0.042 per million input tokens. We don't know Jev's tokenizer or how it counts option text, so take this as an estimate built from Laya's count, not a quote:

DecisionEst. input tokensJev, per 1M decisionsSelf-hosted, cheapest architecture, per 1M
AG News, 4 optionsabout 75about $3$2.57 (head, 52 ms)
Banking77, 77 optionsabout 406about $17$1.61 (head, 32 ms)

The self-hosted column prices single-stream CPU inference at $0.1785 per hour for a 4 vCPU on-demand instance (an AWS c7i.xlarge in us-east-1), with no batching. Batching cuts it further. On a GPU, Laya's own reported latency is 33 ms, six times faster than our 198 ms CPU median on AG News.

For comparison, the same 77-option request to a generative LLM, about 450 input tokens and 10 output, costs about $50 per million decisions on GPT-6 Luna and about $1,000 on Claude Sonnet 5 at list price. Every architecture in this post is cheap next to that. The point of the table is narrower: with many options, per-token pricing charges you for re-sending the label list on every call, and a label-in-the-weights model never sends it at all.

Which architecture for which job.

Your situationUseWhy
Fixed label set, a few hundred labeled examples, high volumeBERT + trained headFlat latency, cheapest per call, steadiest accuracy
Labels change often or arrive with no training data, many specific labelsGLiNER2One pass for any label set, best zero-shot on 77 intents
A handful of options, several typed questions per input, need calibrated probabilitiesLaya or JevTyped answers and confidence per question. Check the option count first
Few labels, no data, accuracy matters more than latencyNLI cross-encoderStrong on 4 labels, but cost grows with every label
Air-gapped or regulated dataAny open model aboveJev is a hosted API only
The answer needs reasoning, or isn't in the textA generative LLMThe only one here that can think before answering

And the decision that sits in front of all of these: which model should answer this prompt? We tested that one zero-shot last week, and Laya and GLiNER2 both lost to a constant guess. Prompt difficulty isn't written in the prompt the way "pin blocked" is written in a banking message. That's the case for labels in the weights.

Tutorial: the same decision, four ways.

Each snippet answers one Banking77-style question. Swap in your own labels and texts, and time them on your hardware before you choose.

Labels in the weights: a trained head.

import numpy as np, torch
from transformers import AutoTokenizer, AutoModel
from sklearn.linear_model import LogisticRegression

tok = AutoTokenizer.from_pretrained("answerdotai/ModernBERT-base")
enc = AutoModel.from_pretrained("answerdotai/ModernBERT-base").eval()

@torch.no_grad()
def embed(texts, bs=32):
    out = []
    for i in range(0, len(texts), bs):
        b = tok(texts[i:i + bs], truncation=True, max_length=256,
                padding=True, return_tensors="pt")
        h = enc(**b).last_hidden_state
        m = b["attention_mask"].unsqueeze(-1)
        out.append(((h * m).sum(1) / m.sum(1)).numpy())   # mean-pool real tokens
    return np.concatenate(out)

head = LogisticRegression(C=1.0, max_iter=3000).fit(embed(train_texts), train_labels)
pred = head.predict(embed(["My card still hasn't arrived"]))

One label per pass: an NLI cross-encoder.

from transformers import pipeline

zs = pipeline("zero-shot-classification",
              model="MoritzLaurer/deberta-v3-base-zeroshot-v2.0")
r = zs("My card still hasn't arrived", labels,           # one pass per label
       hypothesis_template="This text is about {}.")
pred = r["labels"][0]

All labels in the input: GLiNER2.

import gliner2.classification as gc

clf = gc.Classifier.from_pretrained("fastino/gliner2.5-base-v1")
schema = gc.ClassificationSchema().single("intent", labels)   # labels go in the input
probs = clf.classify("My card still hasn't arrived", schema).to_dict()["intent"]["probabilities"]
pred = max(probs, key=probs.get)

Typed questions: Laya (and Jev's API shape).

from laya import Router

r = Router()
q = {"intent": {"type": "choice",
                "instructions": "Which banking request is this?",
                "criteria": {f"o{i}": name for i, name in enumerate(labels)}}}
res = r.predict("My card still hasn't arrived", q, head_max_len=448)
pred = labels[int(res["answers"]["intent"]["choice"][1:])]
print(res["usage"]["input_tokens"], "input tokens")   # what a per-token API would bill

The head_max_len argument is the label-in-the-input tradeoff in one line. Raise it to fit more options and you leave fewer tokens for the text.

Pick with numbers, not a table.

import time, statistics as st

def bench(predict, texts, gold):
    ms, hits = [], 0
    for t, y in zip(texts, gold):
        t0 = time.perf_counter()
        hits += predict(t) == y
        ms.append((time.perf_counter() - t0) * 1000)
    return {"accuracy": hits / len(texts), "p50_ms": st.median(ms)}

Run all four on 200 of your own examples with your real label set. The option count, how specific your label names are, and whether you have a few hundred labeled examples will decide it faster than any vendor benchmark, including ours.

Where Nadir fits.

Nadir makes one decision on every request, which model should answer it, and this post is why that decision lives in a trained head. Prompt difficulty is a three-option question with no useful words in the label names, the case where zero-shot label-in-the-input models did worst in our last test. POST /v1/bucket returns simple, medium, or complex with a probability for each, from a classifier trained on labeled prompts, without spending provider tokens. Our Jev experiment follows the same rule: Jev's typed answers go in as features to a trained head, not as the routing decision itself.

The bigger saving is downstream. If your app sends classification, triage, or yes/no checks to a generative model, those requests are the expensive end of the comparison above. Through Nadir's OpenAI compatible gateway, model="auto" routes decision-shaped prompts to the cheapest model that answers them well, and every response reports which model served it and what it cost. Start with a free key, send a week of traffic, and look at how much of your spend is decisions dressed up as generation.

Conclusion.

Jev, Laya, GLiNER, and a BERT classifier all answer a question without writing prose, and they're priced and benchmarked as if they were interchangeable. They aren't, and the reason is architectural. Put the labels in the weights and you get flat, cheap, steady inference, plus a retrain for every new label. Put one label per pass and cost grows with every option. Put every label in the input and you get run-time flexibility, a sequence that grows with the option list, and, on a per-token API, a bill that's mostly the options. In our runs, Laya led on 4 labels at 95.5% and fell to 19.5% on 77. GLiNER2 did the reverse. A frozen encoder with a linear head came second on both. Count your options, check how specific their names are, and benchmark on your own data before you pick.


Accuracy and latency are from our own runs on public data: 200 examples each from the AG News and Banking77 test splits, CPU only, batch size 1. Jev was not run; its description uses TypeSafe's public statements, and its cost row is an estimate. Not derived from customer data. Sources: [Zaratiana et al., "GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer," arXiv:2311.08526](https://arxiv.org/abs/2311.08526). [Zaratiana et al., "GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface," arXiv:2507.18546](https://arxiv.org/abs/2507.18546). [Laya on Hugging Face](https://huggingface.co/convaiinnovations/laya). [Wikipedia, Jev (AI model)](https://en.wikipedia.org/wiki/Jev_(AI_model)). [Turing Post, "What Is Jev AI?"](https://www.turingpost.com/p/what-is-jev-rlcd). [MoritzLaurer/deberta-v3-base-zeroshot-v2.0](https://huggingface.co/MoritzLaurer/deberta-v3-base-zeroshot-v2.0). [ModernBERT](https://huggingface.co/answerdotai/ModernBERT-base).

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

For coding agents, Nadir connects to Codex, Claude Code, or Cursor and recommends a model tier for delegated work. The agent decides whether to hand the task off and uses its own configured models, so Nadir needs no proxy and no provider keys on that path.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.