Abstract.
"System One" decision models are the fastest-moving category in inference this month. TypeSafe's Jev opened early access on September 15. Laya, an open 421M-parameter answer to it, shipped three days later. And GLiNER's fans spent the week pointing out that schema-driven classification has existed since 2023. All of them take text plus a question and return a probability instead of prose. People compare them on price and on benchmark scores. The difference that predicts how they behave on your workload is architectural, and it comes down to one question: where does the label live? In a classic BERT classifier, the labels live in the weights. In an NLI cross-encoder, each label is a separate input. In GLiNER and Laya, every label is written into the same input as the text. In an LLM, the labels sit in the prompt and the answer is generated. That one choice sets what a new label costs, how latency grows with the number of options, and what you're billed for. We ran four open architectures on the same two tasks, one with 4 labels and one with 77, on the same CPU. No single architecture won both.
All accuracy and latency numbers are from our own runs on public datasets (AG News and Banking77), 200 test examples each, on a 4 vCPU cloud container. Jev is described from TypeSafe's public statements and was not re-run. Not derived from customer data.
Six ways to answer a three-option question.
All of these except the last are encoders: the text is read once, in both directions, and every output is produced in the same forward pass. Nothing is generated, so nothing has to be parsed. What separates them is where the candidate answers go.
1. BERT with a trained head: labels in the weights. The original pattern. Run the text through an encoder (BERT, DeBERTa, ModernBERT), pool the result into one vector, and multiply by a weight matrix with one column per label. The labels aren't in the input at all. They're the columns. That makes inference the cheapest of anything here, and it makes the label set fixed: a new label is a new column, which needs labeled examples and a retrain.
2. NLI cross-encoder: one label per pass. The zero-shot trick from 2019. Take a model trained on natural language inference, pair the text with a hypothesis like "This text is about billing," and read the entailment score. Labels are now free text you can change at run time. The cost is that each label is its own forward pass. Three labels, three passes. Seventy-seven labels, seventy-seven.
3. GLiNER and GLiNER2: all labels in one input. GLiNER (Zaratiana et al., arXiv:2311.08526) fixed the NLI scaling problem for entity extraction by writing every entity type into the same sequence as the text, each behind an [ENT] token, on a DeBERTa-v3 backbone. The bidirectional encoder lets every label token attend to the text and vice versa, and a small network scores each label against each span. GLiNER2 (arXiv:2507.18546), a 205M-parameter encoder, extends that to classification: each label sits behind an [L] token, and an MLP turns each [L] token's output into a logit. One pass, any labels, defined at run time. The paper reports 163 ms on CPU for a 20-label task against 6,758 ms for a DeBERTa NLI baseline.
4. Laya: typed questions on the GLiNER pattern. Laya's open-source package documents its input format: [CLS] <type> instructions [SEP] [MASK] opt0 [MASK] opt1 ... [SEP] state [SEP], on a ModernBERT-large encoder (421M). A scorer reads each [MASK] marker, and separate heads serve the three question types it shares with Jev: yes/no (noul), choice, and score. It's trained with reinforcement learning for calibrated probabilities. Architecturally, it's GLiNER's label-in-the-input design with typed outputs and a different training objective. One consequence: each question is its own sequence, and the options share a token budget (head_max_len, 192 by default) with the question text.
5. Jev: unpublished. TypeSafe has said Jev is transformer-based, trained on synthetic data with a method it calls Reinforcement Learning for Calibrated Decisions (RLCD), and "generates all outputs in a single query" rather than one token at a time, evaluating several questions in parallel. It has not published the architecture, the parameter count, or a paper. Outside analyses of the API's behavior have guessed at a repurposed causal transformer with a shared state encoding, but that's speculation, and we treat it as such. What we do know is the billing: $0.042 per million input tokens, output free.
6. A generative LLM: labels in the prompt. List the options in the prompt, generate the answer token by token, parse the text. It's the most flexible, it handles questions none of the others can, and it's the most expensive per decision by one to three orders of magnitude. We've priced that gap before.
The benchmark.
The research question: does the label's position predict how a model handles a task with few options versus many?
| Tasks | AG News (4 topics: world, sports, business, science and technology) and Banking77 (77 customer-banking intents) |
| Test set | 200 examples from each test split, fixed seed |
| BERT + head | ModernBERT-base (149M), mean-pooled, logistic regression trained on 2,000 random training examples. Frozen encoder, no fine-tuning |
| NLI cross-encoder | MoritzLaurer/deberta-v3-base-zeroshot-v2.0, hypothesis "This text is about {label}." |
| GLiNER2 | fastino/gliner2.5-base-v1, single-label classification |
| Laya | pip install laya 0.3.21, English checkpoint, one choice question. Banking77 needed head_max_len=448 to fit 77 options |
| Labels | The dataset's own label names, lightly cleaned (card_arrival becomes "card arrival"). No descriptions |
| Hardware | 4 vCPU, CPU only, batch size 1, one model at a time |
Three of the four are zero-shot. The trained head saw 2,000 labeled examples, which is about 500 per class on AG News and only about 26 per class on Banking77. Treat it as the "you have some labels" baseline, not as a fully fine-tuned model.
Finding 1: nothing wins both.
| Architecture | Labels live | AG News (4) | Banking77 (77) | p50 latency, AG | p50 latency, Banking77 |
|---|---|---|---|---|---|
| BERT + trained head | weights | 93.5% | 74.0% | 52 ms | 32 ms |
| NLI cross-encoder | one per pass | 92.5% | 63.0% | 168 ms | 1,437 ms |
| GLiNER2 | in the input | 74.0% | 75.5% | 85 ms | 232 ms |
| Laya | in the input, typed | 95.5% | 19.5% | 198 ms | 766 ms |
Laya is the best model in the test on four options and the worst on seventy-seven. That's consistent with what its own documentation says: choice questions degrade above about 20 options, and it reports 0.425 on Banking77 where its table puts Jev at 0.870. Our 19.5% is lower than Laya's own figure, probably because we gave it bare label names, not descriptions. Two more caveats on the AG News number. Laya's public benchmark includes AG News, so we can't rule out that the dataset was in its training data. And on Banking77 it was confidently wrong: "How do I locate my card?" (labeled card_arrival) came back as lost_or_stolen_card at 0.95 confidence.
GLiNER2 is the opposite. It's the weakest on AG News, where "world news" and "business" overlap and a label name carries little information. It's the best on Banking77, where the 77 label names are specific ("exchange rate", "pin blocked", "top up failed") and reading them alongside the text is most of the job. The GLiNER2 paper reports 0.70 zero-shot on Banking77, close to our 75.5%.
The trained head is the steady one. It's second on both tasks, behind whichever zero-shot model happened to fit the task, and it got there with 26 examples per intent on Banking77 and no fine-tuning.
Finding 2: where the label lives sets the latency curve.
To isolate the architecture from the dataset, we took 30 Banking77 messages and asked each model to choose among the first K intents, for K from 2 to 77.
| Labels (K) | 2 | 4 | 8 | 16 | 32 | 64 | 77 | 77 vs 2 |
|---|---|---|---|---|---|---|---|---|
| BERT + trained head | 36 ms | 33 ms | 32 ms | 33 ms | 31 ms | 32 ms | 30 ms | 0.8x |
| GLiNER2 | 74 ms | 71 ms | 85 ms | 92 ms | 125 ms | 186 ms | 237 ms | 3.2x |
| Laya | 145 ms | 174 ms | 220 ms | 291 ms | 499 ms | 802 ms | 823 ms | 5.7x |
| NLI cross-encoder | 79 ms | 108 ms | 202 ms | 323 ms | 592 ms | 1,198 ms | 1,466 ms | 18.5x |
At two labels, the NLI cross-encoder is about as fast as GLiNER2. By seventy-seven, it's six times slower. Laya starts slower than the cross-encoder and ends at about half its latency, but grows more steeply than GLiNER2: its sequences are longer, a typed question plus a [MASK] marker per option, on an encoder twice the size.
The shapes follow directly from Chart 1:
- Labels in the weights: flat. Adding a label adds one column to a matrix multiply. The encoder pass doesn't change.
- One label per pass: linear in K. Batching helps on a GPU, but the compute is still K full passes over the text.
- Labels in the input: grows with the length of the label list, not the number of passes. One pass, but over a longer sequence. Seventy-seven short intent names add a few hundred tokens, and attention cost grows faster than linearly with sequence length.
Finding 3: per-token pricing bills you for the options.
That last point changes what a hosted decision model costs. Laya reports its own token usage, and on a Banking77 request it was 406 input tokens, of which 7 were the customer's message. The other 399 were the question and the 77 options. On a label-in-the-input architecture, the option list is the request.
Jev bills $0.042 per million input tokens. We don't know Jev's tokenizer or how it counts option text, so take this as an estimate built from Laya's count, not a quote:
| Decision | Est. input tokens | Jev, per 1M decisions | Self-hosted, cheapest architecture, per 1M |
|---|---|---|---|
| AG News, 4 options | about 75 | about $3 | $2.57 (head, 52 ms) |
| Banking77, 77 options | about 406 | about $17 | $1.61 (head, 32 ms) |
The self-hosted column prices single-stream CPU inference at $0.1785 per hour for a 4 vCPU on-demand instance (an AWS c7i.xlarge in us-east-1), with no batching. Batching cuts it further. On a GPU, Laya's own reported latency is 33 ms, six times faster than our 198 ms CPU median on AG News.
For comparison, the same 77-option request to a generative LLM, about 450 input tokens and 10 output, costs about $50 per million decisions on GPT-6 Luna and about $1,000 on Claude Sonnet 5 at list price. Every architecture in this post is cheap next to that. The point of the table is narrower: with many options, per-token pricing charges you for re-sending the label list on every call, and a label-in-the-weights model never sends it at all.
Which architecture for which job.
| Your situation | Use | Why |
|---|---|---|
| Fixed label set, a few hundred labeled examples, high volume | BERT + trained head | Flat latency, cheapest per call, steadiest accuracy |
| Labels change often or arrive with no training data, many specific labels | GLiNER2 | One pass for any label set, best zero-shot on 77 intents |
| A handful of options, several typed questions per input, need calibrated probabilities | Laya or Jev | Typed answers and confidence per question. Check the option count first |
| Few labels, no data, accuracy matters more than latency | NLI cross-encoder | Strong on 4 labels, but cost grows with every label |
| Air-gapped or regulated data | Any open model above | Jev is a hosted API only |
| The answer needs reasoning, or isn't in the text | A generative LLM | The only one here that can think before answering |
And the decision that sits in front of all of these: which model should answer this prompt? We tested that one zero-shot last week, and Laya and GLiNER2 both lost to a constant guess. Prompt difficulty isn't written in the prompt the way "pin blocked" is written in a banking message. That's the case for labels in the weights.
Tutorial: the same decision, four ways.
Each snippet answers one Banking77-style question. Swap in your own labels and texts, and time them on your hardware before you choose.
Labels in the weights: a trained head.
import numpy as np, torch
from transformers import AutoTokenizer, AutoModel
from sklearn.linear_model import LogisticRegression
tok = AutoTokenizer.from_pretrained("answerdotai/ModernBERT-base")
enc = AutoModel.from_pretrained("answerdotai/ModernBERT-base").eval()
@torch.no_grad()
def embed(texts, bs=32):
out = []
for i in range(0, len(texts), bs):
b = tok(texts[i:i + bs], truncation=True, max_length=256,
padding=True, return_tensors="pt")
h = enc(**b).last_hidden_state
m = b["attention_mask"].unsqueeze(-1)
out.append(((h * m).sum(1) / m.sum(1)).numpy()) # mean-pool real tokens
return np.concatenate(out)
head = LogisticRegression(C=1.0, max_iter=3000).fit(embed(train_texts), train_labels)
pred = head.predict(embed(["My card still hasn't arrived"]))
One label per pass: an NLI cross-encoder.
from transformers import pipeline
zs = pipeline("zero-shot-classification",
model="MoritzLaurer/deberta-v3-base-zeroshot-v2.0")
r = zs("My card still hasn't arrived", labels, # one pass per label
hypothesis_template="This text is about {}.")
pred = r["labels"][0]
All labels in the input: GLiNER2.
import gliner2.classification as gc
clf = gc.Classifier.from_pretrained("fastino/gliner2.5-base-v1")
schema = gc.ClassificationSchema().single("intent", labels) # labels go in the input
probs = clf.classify("My card still hasn't arrived", schema).to_dict()["intent"]["probabilities"]
pred = max(probs, key=probs.get)
Typed questions: Laya (and Jev's API shape).
from laya import Router
r = Router()
q = {"intent": {"type": "choice",
"instructions": "Which banking request is this?",
"criteria": {f"o{i}": name for i, name in enumerate(labels)}}}
res = r.predict("My card still hasn't arrived", q, head_max_len=448)
pred = labels[int(res["answers"]["intent"]["choice"][1:])]
print(res["usage"]["input_tokens"], "input tokens") # what a per-token API would bill
The head_max_len argument is the label-in-the-input tradeoff in one line. Raise it to fit more options and you leave fewer tokens for the text.
Pick with numbers, not a table.
import time, statistics as st
def bench(predict, texts, gold):
ms, hits = [], 0
for t, y in zip(texts, gold):
t0 = time.perf_counter()
hits += predict(t) == y
ms.append((time.perf_counter() - t0) * 1000)
return {"accuracy": hits / len(texts), "p50_ms": st.median(ms)}
Run all four on 200 of your own examples with your real label set. The option count, how specific your label names are, and whether you have a few hundred labeled examples will decide it faster than any vendor benchmark, including ours.
Where Nadir fits.
Nadir makes one decision on every request, which model should answer it, and this post is why that decision lives in a trained head. Prompt difficulty is a three-option question with no useful words in the label names, the case where zero-shot label-in-the-input models did worst in our last test. POST /v1/bucket returns simple, medium, or complex with a probability for each, from a classifier trained on labeled prompts, without spending provider tokens. Our Jev experiment follows the same rule: Jev's typed answers go in as features to a trained head, not as the routing decision itself.
The bigger saving is downstream. If your app sends classification, triage, or yes/no checks to a generative model, those requests are the expensive end of the comparison above. Through Nadir's OpenAI compatible gateway, model="auto" routes decision-shaped prompts to the cheapest model that answers them well, and every response reports which model served it and what it cost. Start with a free key, send a week of traffic, and look at how much of your spend is decisions dressed up as generation.
Conclusion.
Jev, Laya, GLiNER, and a BERT classifier all answer a question without writing prose, and they're priced and benchmarked as if they were interchangeable. They aren't, and the reason is architectural. Put the labels in the weights and you get flat, cheap, steady inference, plus a retrain for every new label. Put one label per pass and cost grows with every option. Put every label in the input and you get run-time flexibility, a sequence that grows with the option list, and, on a per-token API, a bill that's mostly the options. In our runs, Laya led on 4 labels at 95.5% and fell to 19.5% on 77. GLiNER2 did the reverse. A frozen encoder with a linear head came second on both. Count your options, check how specific their names are, and benchmark on your own data before you pick.
Accuracy and latency are from our own runs on public data: 200 examples each from the AG News and Banking77 test splits, CPU only, batch size 1. Jev was not run; its description uses TypeSafe's public statements, and its cost row is an estimate. Not derived from customer data. Sources: [Zaratiana et al., "GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer," arXiv:2311.08526](https://arxiv.org/abs/2311.08526). [Zaratiana et al., "GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface," arXiv:2507.18546](https://arxiv.org/abs/2507.18546). [Laya on Hugging Face](https://huggingface.co/convaiinnovations/laya). [Wikipedia, Jev (AI model)](https://en.wikipedia.org/wiki/Jev_(AI_model)). [Turing Post, "What Is Jev AI?"](https://www.turingpost.com/p/what-is-jev-rlcd). [MoritzLaurer/deberta-v3-base-zeroshot-v2.0](https://huggingface.co/MoritzLaurer/deberta-v3-base-zeroshot-v2.0). [ModernBERT](https://huggingface.co/answerdotai/ModernBERT-base).