TL;DR.
TypeSafe launched Jev on September 15, 2026: a "decision model" that reads a state, answers typed questions with calibrated probabilities, bills input tokens only, and writes no prose. Sixteen days later it had at least five competitors. Fastino shipped GLiNER2.5-Decide on September 24 and GLiDE on September 30. On October 1, Cloudflare released Clef and Clef-flash as open weights on Workers AI, and Perplexity open-sourced pplx-decider-v1-27b behind a new Decisions API.
The category went from one vendor to a market in two weeks, and the market is already doing what every model market does:
- Prices spread 6x on near-identical models. Clef and pplx-decider are both fine-tunes of Qwen3.8-27B. Clef lists at $0.24 per million input tokens, pplx-decider at $0.04. At 300 tokens a decision, that's $72 vs $12 per million decisions. Clef costs more per decision than GPT-6 Luna used as a classifier.
- Every vendor wins its own benchmark. Four launches, four headline wins over Jev, each on the launcher's own table. Inside Perplexity's panel, where pplx-decider leads overall by 1.2 points, Jev still wins 6 of the 11 benchmarks.
- Latency claims don't survive the network. Cloudflare quotes a 38.8 ms median for Clef-flash. An independent client-side test measured 191 to 205 ms.
This post compares the six on price, openness, context, and latency, works out when self-hosting the open ones pays, and walks through a Python bake-off that picks a decider by cost per correct decision on your own labels, not by its launch table.
Prices and specs are from vendor launch materials and press coverage, October 1 to 3, 2026. Benchmarks are vendor-reported unless marked. The escalation example is illustrative. None of this is Nadir customer data.
Sixteen days, six deciders.
| Model | Vendor | Launched | Open weights | Size | Context | Input / output, per M |
|---|---|---|---|---|---|---|
| Jev 1.13 | TypeSafe | Sept 15 | No | Unpublished | 32K | $0.042 / free |
| GLiNER2.5-Decide | Fastino | Sept 24 | Yes, Apache 2.0 | 340M | Encoder | Self-host, CPU |
| GLiDE | Fastino | Sept 30 | No | Unpublished | 40K per question | Listed inconsistently |
| Clef | Cloudflare | Oct 1 | Yes, Apache 2.0 | 27B (Qwen3.8-27B) | 64K | $0.24 / free |
| Clef-flash | Cloudflare | Oct 1 | Yes, Apache 2.0 | 9B (Qwen3.5-9B) | 64K | $0.09 / free |
| pplx-decider-v1-27b | Perplexity | Oct 1 | Yes, Apache 2.0 | 26.1B (Qwen3.8-27B) | 262K | $0.04 / free |
Sources: Cloudflare's Clef, deep dive by Flavio Copes, traictory.com on the Clef launch, BenchLM decision models, Fastino on GLiNER2.5-Decide, explainx.ai on pplx-decider. GLiDE's price appears as $0 input and $0.04 output in one copy of Fastino's docs and as $0.30 input with free output elsewhere. Check before you budget.
They share a shape. You send a state (a ticket, a page, an agent trace, a prompt) and a set of questions, each typed: a yes/no probability (Jev calls it noul), a choice over a fixed option list, or a score on a scale. You get back a probability for every allowed answer. Cloudflare calls Clef "fully Jev-API compatible," so code written against Jev can point at Clef without a rewrite. That's the most important line in any of the launches. The interface is becoming a standard, which turns the model behind it into a line item you can swap.
We covered what Jev is and what it costs, where each architecture keeps its labels, and why zero-shot decision models lost to `len()` as prompt routers. This post is about buying one now that there's a choice.
What a million decisions cost.
A typical decision is a few hundred tokens: a short state, one question, a handful of options. At 300 input tokens, and about 10 output tokens for the chat models that have to write their answer:
| Model | Per decision | Per million decisions | vs pplx-decider |
|---|---|---|---|
| pplx-decider-v1-27b | $0.0000120 | $12.00 | 1.0x |
| Jev 1.13 | $0.0000126 | $12.60 | 1.05x |
| Clef-flash | $0.0000270 | $27.00 | 2.3x |
| GPT-6 Luna (classifier) | $0.0000350 | $35.00 | 2.9x |
| Clef | $0.0000720 | $72.00 | 6.0x |
| Claude Haiku 4.5 (classifier) | $0.000350 | $350.00 | 29x |
| Claude Sonnet 5 (classifier) | $0.000700 | $700.00 | 58x |
The Register called Clef "nearly six times the price of Jev," and price was the main complaint in a 614-point Hacker News thread. The table shows why. Two things are worth noticing past the headline.
The 27B pair is the same-weights story again. Clef and pplx-decider are both fine-tunes of Qwen3.8-27B, roughly the same compute per token, and they list 6x apart. It's the pattern we found with one open model priced 10.4x apart across hosts: the price reflects the seller's strategy, not the work.
Clef is priced above a cheap chat model. At $72 per million decisions, Clef costs about twice what GPT-6 Luna costs to answer the same question as text. The decider still has real advantages: calibrated probabilities instead of parsed text, no format failures, and up to 64 questions on one state in one request, which bills the state once. If you ask eight questions of the same 2,000-token trace, that last point matters more than the per-token price. But the idea that a decision model is always far cheaper than an LLM is no longer true by default. It depends which two you compare.
Every vendor wins its own table.
All four challengers beat Jev at launch, each on a table it chose. That isn't fraud. Every vendor tunes for the tasks it cares about and reports the suite it built against. But it means the headline numbers can't rank these models for you, and the detail underneath disagrees with the headlines.
Perplexity's own panel is the clearest case:
| Benchmark | Jev | pplx-decider | Winner |
|---|---|---|---|
| WinoGrande | 90.70% | 83.30% | Jev |
| FinancialPhraseBank | 76.98% | 84.18% | pplx-decider |
| RAGTruth | 77.27% | 88.80% | pplx-decider |
| JudgeBench | 78.57% | 78.29% | Jev |
| BBH | 94.27% | 82.80% | Jev |
| JevBench public hard | 73.27% | 70.30% | Jev |
| TabFact | 89.80% | 90.60% | pplx-decider |
| ContractNLI | 77.45% | 80.78% | pplx-decider |
| Circa | 84.60% | 89.20% | pplx-decider |
| Belebele | 95.00% | 94.00% | Jev |
| TruthfulQA binary | 92.00% | 85.40% | Jev |
| Overall | 84.51% | 85.71% | pplx-decider |
pplx-decider wins the average by 1.2 points and loses 6 of the 11 rows. The spread by task is far larger than the spread of the averages: 11.5 points on BBH one way, 11.5 on RAGTruth the other. If your decision looks like hallucination detection, pplx-decider is the better model on this evidence. If it looks like reasoning over a short passage, Jev is.
Cloudflare's table splits the same way. Clef beat Jev on BANKING77 intent classification by 14.5 points (94.20 vs 79.74) and lost When2Call, the "should the agent call a tool right now" benchmark, by 8.6 (72.37 vs 80.97). Clef-flash, the cheaper 9B model, beat the 27B Clef on three of Cloudflare's six rows. HN commenters also argued Cloudflare compared against weaker Jev variants. Whether or not that's true, it's the reason to run your own test.
Thirty-nine milliseconds, in the press release.
Cloudflare reports Clef-flash at a 38.8 ms median and Clef at 209.3 ms, with Jev at 524.1 ms in the same run. Flavio Copes measured from a client and got 191 to 205 ms for Clef-flash and 524 to 726 ms for Clef. Both sets of numbers are probably honest. They measure different things: model time inside the data center versus what your code waits for, including the network, TLS, and the queue.
For a decision that gates a user-facing request, the second number is the one in your latency budget. It's also the argument for open weights. A 9B decider running next to your service avoids the network hop entirely.
When self-hosting the open ones pays.
Four of the six ship weights under Apache 2.0, so you can run them inside your own VPC. That solves the problem we flagged when we tested Jev as a verifier: every hosted decision sends your prompt to another subprocessor. Whether it also saves money depends on utilization.
Decisions are prefill-only. There's no token-by-token decode, so a GPU's throughput on this workload is its prefill rate, which is much higher than its generation rate. The break-even is simple: the GPU has to read enough tokens per hour to match the API bill.
| Model | Example GPU (assumed price) | API price it must beat | Sustained input tokens/s to break even |
|---|---|---|---|
| pplx-decider (~49 GB in BF16) | 1x H100 80 GB at $2.50/hr | $0.04/M | 17,361 |
| Clef (27B) | 1x H100 80 GB at $2.50/hr | $0.24/M | 2,894 |
| Clef-flash (9B) | 1x L40S 48 GB at $1.00/hr | $0.09/M | 3,086 |
| GLiNER2.5-Decide (340M) | CPU | Nothing to beat | Fastino reports 167 ms per call on 48 vCPUs |
GPU prices are illustrative on-demand rates. Use your own.
Read the table by row. Beating Perplexity's $0.04 means keeping an H100 busy at over 17,000 tokens a second, around the clock, which almost no single workload does. Beating Clef's $0.24 on the same GPU takes a sixth of that. So the open weights put a ceiling on Clef's API price, and pplx-decider's API price puts a ceiling on everyone's. If you need the data to stay in your network, self-host. If you just need it cheap and you're under a few thousand tokens a second, the cheapest hosted API wins. We worked through the same arithmetic for chat models in the utilization tax.
Calibration is the price you don't see.
A decision model is rarely trusted on every call. The useful pattern is selective: accept the answer when the top probability clears a threshold, and escalate the rest to something stronger, a frontier LLM or a human. That turns calibration into cost. A decider whose probabilities are honest can accept more calls at the same error rate, so it escalates less.
The blended cost per decision is:
cost = decider_price + (1 - coverage) * fallback_price
where coverage is the share of calls you accept at the threshold that hits your accuracy target. An illustrative example, with Claude Haiku 4.5 as the fallback at $350 per million decisions:
| Decider | Price per M | Coverage at 98% accepted accuracy | Escalation cost per M | Blended per M |
|---|---|---|---|---|
| Cheap, poorly calibrated | $12 | 70% | $105 | $117 |
| 6x pricier, well calibrated | $72 | 90% | $35 | $107 |
Hypothetical coverages, chosen to show the mechanism. Measure yours.
The model that costs 6x more per token is the cheaper system. The reverse also happens, which is the point: you can't tell from a price sheet or a launch table. Coverage at your accuracy target, on your data, is the only number that ranks them.
Tutorial: a bake-off on your own labels.
You need a few hundred decisions from production with known right answers. Ticket routes your team corrected, moderation calls a reviewer confirmed, tool choices an agent got right. Split them 50/50 into a calibration set and a test set.
Step 1: wrap each decider behind one interface.
Jev and Clef share a request shape, Perplexity's is close, and a chat model can play along with log-probabilities. Normalize all of them to "probability per option, plus input tokens billed":
import math, time
from dataclasses import dataclass
from typing import Callable
import numpy as np
from openai import OpenAI
@dataclass
class Decider:
name: str
price_in: float # USD per million input tokens
price_out: float # USD per million output tokens (0 for decision models)
ask: Callable[[str, str, list[str]], tuple[dict[str, float], int, int]]
# ask(state, question, options) -> ({option: prob}, input_tokens, output_tokens)
def jev_compatible(base_url: str, api_key: str, model: str):
"""Jev, Clef and Clef-flash. Field names follow each vendor's docs as of
Oct 1, 2026; check them against the current API reference."""
import httpx
def ask(state, question, options):
r = httpx.post(base_url, headers={"Authorization": f"Bearer {api_key}"}, json={
"model": model,
"state": state,
"questions": [{"id": "q", "type": "choice", "question": question, "options": options}],
}, timeout=10)
r.raise_for_status()
body = r.json()
probs = body["answers"]["q"]["probabilities"] # {option: p}
return probs, body["usage"]["input_tokens"], 0
return ask
def chat_logprobs(model: str, client: OpenAI):
"""Any chat model as a classifier: ask for a single option letter, read logprobs."""
def ask(state, question, options):
letters = [chr(65 + i) for i in range(len(options))]
menu = "\n".join(f"{l}. {o}" for l, o in zip(letters, options))
r = client.chat.completions.create(
model=model, max_tokens=1, logprobs=True, top_logprobs=min(20, len(options)),
messages=[{"role": "user", "content": f"{state}\n\n{question}\n{menu}\nAnswer with one letter."}],
)
top = {t.token.strip(): math.exp(t.logprob) for t in r.choices[0].logprobs.content[0].top_logprobs}
raw = {o: top.get(l, 0.0) for l, o in zip(letters, options)}
z = sum(raw.values()) or 1.0
return {o: p / z for o, p in raw.items()}, r.usage.prompt_tokens, r.usage.completion_tokens
return ask
The chat baseline matters. Without it you can't tell whether a decider is worth adding, only which decider is best.
Step 2: run every row through every decider.
def run(decider: Decider, rows):
out = []
for row in rows: # row: {"state", "question", "options", "label"}
t0 = time.perf_counter()
probs, tin, tout = decider.ask(row["state"], row["question"], row["options"])
ms = (time.perf_counter() - t0) * 1000
pred = max(probs, key=probs.get)
cost = (tin * decider.price_in + tout * decider.price_out) / 1e6
out.append({"conf": probs[pred], "correct": pred == row["label"], "cost": cost, "ms": ms})
return out
Run it from the same region your service runs in. Chart 3 is why.
Step 3: pick a threshold on the calibration half, score on the test half.
TARGET = 0.98 # accuracy you need on the calls you accept
FALLBACK_COST = 0.00035 # per escalated decision, e.g. Claude Haiku 4.5 at ~300 tokens
def threshold_for(results, target=TARGET):
"""Lowest confidence cutoff whose accepted calls still hit the target."""
for t in np.arange(0.50, 1.00, 0.01):
kept = [r for r in results if r["conf"] >= t]
if kept and np.mean([r["correct"] for r in kept]) >= target:
return t
return 1.01 # never accept: everything escalates
def ece(results, bins=10):
"""Expected calibration error: how far stated confidence is from accuracy."""
conf = np.array([r["conf"] for r in results]); ok = np.array([r["correct"] for r in results])
edges = np.linspace(0, 1, bins + 1); err = 0.0
for lo, hi in zip(edges[:-1], edges[1:]):
m = (conf > lo) & (conf <= hi)
if m.any():
err += m.mean() * abs(conf[m].mean() - ok[m].mean())
return err
def score(decider, calib_rows, test_rows):
t = threshold_for(run(decider, calib_rows))
test = run(decider, test_rows)
kept = [r for r in test if r["conf"] >= t]
coverage = len(kept) / len(test)
blended = np.mean([r["cost"] for r in test]) + (1 - coverage) * FALLBACK_COST
return {
"decider": decider.name,
"threshold": round(t, 2),
"coverage": round(coverage, 3),
"accepted_acc": round(np.mean([r["correct"] for r in kept]), 3) if kept else None,
"ece": round(ece(test), 3),
"p95_ms": round(np.percentile([r["ms"] for r in test], 95)),
"usd_per_million": round(blended * 1e6, 2),
}
Rank by usd_per_million, and reject any decider whose accepted_acc on the test half misses the target by more than sampling noise. With 200 test rows, a point or two of difference isn't real. A coverage gap of 20 points is.
Step 4: keep it swappable.
Write your application against the Decider interface, not a vendor SDK. This market will reprice within weeks: Perplexity has already said it plans to cut its price further, and three of the four challenger vendors ship open weights anyone can host. Re-run the bake-off when a price moves or a new model ships. It's a few dollars of API calls.
Where Nadir fits.
Nadir answers one decision on every request: which model does this prompt need? POST /v1/bucket returns simple, medium, or complex with a probability for each, and the classifier behind it is trained on labeled prompts rather than prompted zero-shot, for the reasons in Beaten by len(). Decision models are inputs to that classifier, not a replacement for it, and the bake-off above is how we decide which ones earn a place.
If you're adding a decider to your own pipeline, the bake-off is yours to run. If the decision you actually need is "cheap model or expensive one," send model="auto" and Nadir makes it per prompt, against the quality floor you set, with the decision and its cost in the response headers. Observe mode records what routing would have chosen without changing a single response, so you can size the saving first. Start with a free key.
Conclusion.
Decision models went from one vendor to six in sixteen days, and the category already behaves like a market. Two fine-tunes of the same 27B base list 6x apart, at $12 and $72 per million decisions, and the expensive one costs more per decision than a cheap chat model asked the same question. Every launch beat Jev on its own table, while the per-task detail splits both ways by more than 10 points. Headline latency was 5x lower than one independent client-side test. None of that tells you which decider to use. Coverage at your accuracy target, on your own labels, measured from your own region, does. Build against the shared interface, run the bake-off, and run it again when the prices move.
Sources: [Flavio Copes, "A deep dive into Clef, Cloudflare's decision model"](https://flaviocopes.com/clef/), October 2026. [traictory.com, "Cloudflare ships Clef, an open-weight answer to Jev, with a price critics noticed"](https://traictory.com/news/2026-10-03-cloudflare-clef-decision-models), October 3, 2026, citing The Register and Hacker News. [Dealroom, "Cloudflare launches Clef as TypeSafe's 'decision model' idea spreads across big tech"](https://dealroom.co/news/158476-cloudflare-launches-clef-as-typesafes-decision-model-idea-spreads-across/). [explainx.ai, "pplx-decider: 85.71% vs Jev, $0.04/M"](https://explainx.ai/blog/perplexity-pplx-decider-decisions-api-2026). [Perplexity Developers on X, Decisions API launch](https://x.com/perplexitydevs/status/2105725598882832414). [pplx-decider-v1-27b on OpenRouter](https://openrouter.ai/perplexity/pplx-decider-v1-27b). [BenchLM, "AI Decision Models: GLiDE, Jev & Perplexity Decider"](https://benchlm.ai/decision-models) and [GLiDE benchmarks](https://benchlm.ai/models/fastino-glide). [Fastino, "GLiNER2.5-Decide: An Open-Weight Model for Structured Decision Making"](https://fastino.ai/blog/gliner-2-5-decide-open-weight-decision-model), September 24, 2026. LLM classifier prices from each provider's list price as of October 2026.