Six Deciders, Sixteen Days

Six decision models launched in sixteen days, and two of them share a 27B base model yet list 6x apart: $12 vs $72 per million decisions. After TypeSafe's Jev on September 15, 2026 came Fastino's GLiNER2.5-Decide and GLiDE, Cloudflare's open-weight Clef and Clef-flash, and Perplexity's open-sourced pplx-decider-v1-27b. Clef and pplx-decider are both Qwen3.8-27B fine-tunes, and at 300 tokens a decision Clef costs more than GPT-6 Luna used as a classifier. Every launch beat Jev on its own table, yet Jev still wins 6 of the 11 rows on Perplexity's panel, and Clef-flash's 38.8 ms headline median measured 191 to 205 ms from a client. A comparison of all six on price, openness, context and latency, the GPU break-even for self-hosting the open ones, and a Python bake-off that ranks deciders by cost per correct decision on your own labels.

Published 2026-10-04 by Dor Amir on the Nadir blog.

Filed under Routing & Cascades.

TL;DR.

TypeSafe launched Jev on September 15, 2026: a "decision model" that reads a state, answers typed questions with calibrated probabilities, bills input tokens only, and writes no prose. Sixteen days later it had at least five competitors. Fastino shipped GLiNER2.5-Decide on September 24 and GLiDE on September 30. On October 1, Cloudflare released Clef and Clef-flash as open weights on Workers AI, and Perplexity open-sourced pplx-decider-v1-27b behind a new Decisions API.

The category went from one vendor to a market in two weeks, and the market is already doing what every model market does:

This post compares the six on price, openness, context, and latency, works out when self-hosting the open ones pays, and walks through a Python bake-off that picks a decider by cost per correct decision on your own labels, not by its launch table.

Prices and specs are from vendor launch materials and press coverage, October 1 to 3, 2026. Benchmarks are vendor-reported unless marked. The escalation example is illustrative. None of this is Nadir customer data.

Sixteen days, six deciders.

ModelVendorLaunchedOpen weightsSizeContextInput / output, per M
Jev 1.13TypeSafeSept 15NoUnpublished32K$0.042 / free
GLiNER2.5-DecideFastinoSept 24Yes, Apache 2.0340MEncoderSelf-host, CPU
GLiDEFastinoSept 30NoUnpublished40K per questionListed inconsistently
ClefCloudflareOct 1Yes, Apache 2.027B (Qwen3.8-27B)64K$0.24 / free
Clef-flashCloudflareOct 1Yes, Apache 2.09B (Qwen3.5-9B)64K$0.09 / free
pplx-decider-v1-27bPerplexityOct 1Yes, Apache 2.026.1B (Qwen3.8-27B)262K$0.04 / free

Sources: Cloudflare's Clef, deep dive by Flavio Copes, traictory.com on the Clef launch, BenchLM decision models, Fastino on GLiNER2.5-Decide, explainx.ai on pplx-decider. GLiDE's price appears as $0 input and $0.04 output in one copy of Fastino's docs and as $0.30 input with free output elsewhere. Check before you budget.

They share a shape. You send a state (a ticket, a page, an agent trace, a prompt) and a set of questions, each typed: a yes/no probability (Jev calls it noul), a choice over a fixed option list, or a score on a scale. You get back a probability for every allowed answer. Cloudflare calls Clef "fully Jev-API compatible," so code written against Jev can point at Clef without a rewrite. That's the most important line in any of the launches. The interface is becoming a standard, which turns the model behind it into a line item you can swap.

We covered what Jev is and what it costs, where each architecture keeps its labels, and why zero-shot decision models lost to `len()` as prompt routers. This post is about buying one now that there's a choice.

What a million decisions cost.

Chart 1: Cost per million decisions at 300 input tokens each, log scale. Perplexity Decider $12.00, Jev $12.60, Clef-flash $27.00, GPT-6 Luna as an LLM classifier $35.00, Clef $72.00, Claude Haiku 4.5 as a classifier $350.00, Claude Sonnet 5 as a classifier $700.00.
Chart 1: Cost per million decisions at 300 input tokens each, log scale. Perplexity Decider $12.00, Jev $12.60, Clef-flash $27.00, GPT-6 Luna as an LLM classifier $35.00, Clef $72.00, Claude Haiku 4.5 as a classifier $350.00, Claude Sonnet 5 as a classifier $700.00.

A typical decision is a few hundred tokens: a short state, one question, a handful of options. At 300 input tokens, and about 10 output tokens for the chat models that have to write their answer:

ModelPer decisionPer million decisionsvs pplx-decider
pplx-decider-v1-27b$0.0000120$12.001.0x
Jev 1.13$0.0000126$12.601.05x
Clef-flash$0.0000270$27.002.3x
GPT-6 Luna (classifier)$0.0000350$35.002.9x
Clef$0.0000720$72.006.0x
Claude Haiku 4.5 (classifier)$0.000350$350.0029x
Claude Sonnet 5 (classifier)$0.000700$700.0058x

The Register called Clef "nearly six times the price of Jev," and price was the main complaint in a 614-point Hacker News thread. The table shows why. Two things are worth noticing past the headline.

The 27B pair is the same-weights story again. Clef and pplx-decider are both fine-tunes of Qwen3.8-27B, roughly the same compute per token, and they list 6x apart. It's the pattern we found with one open model priced 10.4x apart across hosts: the price reflects the seller's strategy, not the work.

Clef is priced above a cheap chat model. At $72 per million decisions, Clef costs about twice what GPT-6 Luna costs to answer the same question as text. The decider still has real advantages: calibrated probabilities instead of parsed text, no format failures, and up to 64 questions on one state in one request, which bills the state once. If you ask eight questions of the same 2,000-token trace, that last point matters more than the per-token price. But the idea that a decision model is always far cheaper than an LLM is no longer true by default. It depends which two you compare.

Every vendor wins its own table.

Chart 2: Each launch's headline score vs Jev, from its own materials. Cloudflare Decision Index: Clef 61.2, Jev 57.9. Fastino Decision Index 0.2.1: GLiDE 64.81, Jev 1.13.0 57.91. Fastino average: GLiNER2.5-Decide 60.1, JevK5 57.5. Perplexity 11-benchmark panel: pplx-decider 85.71, Jev 84.51.
Chart 2: Each launch's headline score vs Jev, from its own materials. Cloudflare Decision Index: Clef 61.2, Jev 57.9. Fastino Decision Index 0.2.1: GLiDE 64.81, Jev 1.13.0 57.91. Fastino average: GLiNER2.5-Decide 60.1, JevK5 57.5. Perplexity 11-benchmark panel: pplx-decider 85.71, Jev 84.51.

All four challengers beat Jev at launch, each on a table it chose. That isn't fraud. Every vendor tunes for the tasks it cares about and reports the suite it built against. But it means the headline numbers can't rank these models for you, and the detail underneath disagrees with the headlines.

Perplexity's own panel is the clearest case:

BenchmarkJevpplx-deciderWinner
WinoGrande90.70%83.30%Jev
FinancialPhraseBank76.98%84.18%pplx-decider
RAGTruth77.27%88.80%pplx-decider
JudgeBench78.57%78.29%Jev
BBH94.27%82.80%Jev
JevBench public hard73.27%70.30%Jev
TabFact89.80%90.60%pplx-decider
ContractNLI77.45%80.78%pplx-decider
Circa84.60%89.20%pplx-decider
Belebele95.00%94.00%Jev
TruthfulQA binary92.00%85.40%Jev
Overall84.51%85.71%pplx-decider

pplx-decider wins the average by 1.2 points and loses 6 of the 11 rows. The spread by task is far larger than the spread of the averages: 11.5 points on BBH one way, 11.5 on RAGTruth the other. If your decision looks like hallucination detection, pplx-decider is the better model on this evidence. If it looks like reasoning over a short passage, Jev is.

Cloudflare's table splits the same way. Clef beat Jev on BANKING77 intent classification by 14.5 points (94.20 vs 79.74) and lost When2Call, the "should the agent call a tool right now" benchmark, by 8.6 (72.37 vs 80.97). Clef-flash, the cheaper 9B model, beat the 27B Clef on three of Cloudflare's six rows. HN commenters also argued Cloudflare compared against weaker Jev variants. Whether or not that's true, it's the reason to run your own test.

Thirty-nine milliseconds, in the press release.

Chart 3: Median latency. Clef-flash, Cloudflare-measured 38.8 ms, independently measured 191 to 205 ms. Clef, Cloudflare-measured 209.3 ms, independently 524 to 726 ms. Jev, Cloudflare-measured 524.1 ms.
Chart 3: Median latency. Clef-flash, Cloudflare-measured 38.8 ms, independently measured 191 to 205 ms. Clef, Cloudflare-measured 209.3 ms, independently 524 to 726 ms. Jev, Cloudflare-measured 524.1 ms.

Cloudflare reports Clef-flash at a 38.8 ms median and Clef at 209.3 ms, with Jev at 524.1 ms in the same run. Flavio Copes measured from a client and got 191 to 205 ms for Clef-flash and 524 to 726 ms for Clef. Both sets of numbers are probably honest. They measure different things: model time inside the data center versus what your code waits for, including the network, TLS, and the queue.

For a decision that gates a user-facing request, the second number is the one in your latency budget. It's also the argument for open weights. A 9B decider running next to your service avoids the network hop entirely.

When self-hosting the open ones pays.

Four of the six ship weights under Apache 2.0, so you can run them inside your own VPC. That solves the problem we flagged when we tested Jev as a verifier: every hosted decision sends your prompt to another subprocessor. Whether it also saves money depends on utilization.

Decisions are prefill-only. There's no token-by-token decode, so a GPU's throughput on this workload is its prefill rate, which is much higher than its generation rate. The break-even is simple: the GPU has to read enough tokens per hour to match the API bill.

ModelExample GPU (assumed price)API price it must beatSustained input tokens/s to break even
pplx-decider (~49 GB in BF16)1x H100 80 GB at $2.50/hr$0.04/M17,361
Clef (27B)1x H100 80 GB at $2.50/hr$0.24/M2,894
Clef-flash (9B)1x L40S 48 GB at $1.00/hr$0.09/M3,086
GLiNER2.5-Decide (340M)CPUNothing to beatFastino reports 167 ms per call on 48 vCPUs

GPU prices are illustrative on-demand rates. Use your own.

Read the table by row. Beating Perplexity's $0.04 means keeping an H100 busy at over 17,000 tokens a second, around the clock, which almost no single workload does. Beating Clef's $0.24 on the same GPU takes a sixth of that. So the open weights put a ceiling on Clef's API price, and pplx-decider's API price puts a ceiling on everyone's. If you need the data to stay in your network, self-host. If you just need it cheap and you're under a few thousand tokens a second, the cheapest hosted API wins. We worked through the same arithmetic for chat models in the utilization tax.

Calibration is the price you don't see.

A decision model is rarely trusted on every call. The useful pattern is selective: accept the answer when the top probability clears a threshold, and escalate the rest to something stronger, a frontier LLM or a human. That turns calibration into cost. A decider whose probabilities are honest can accept more calls at the same error rate, so it escalates less.

The blended cost per decision is:

cost = decider_price + (1 - coverage) * fallback_price

where coverage is the share of calls you accept at the threshold that hits your accuracy target. An illustrative example, with Claude Haiku 4.5 as the fallback at $350 per million decisions:

DeciderPrice per MCoverage at 98% accepted accuracyEscalation cost per MBlended per M
Cheap, poorly calibrated$1270%$105$117
6x pricier, well calibrated$7290%$35$107

Hypothetical coverages, chosen to show the mechanism. Measure yours.

The model that costs 6x more per token is the cheaper system. The reverse also happens, which is the point: you can't tell from a price sheet or a launch table. Coverage at your accuracy target, on your data, is the only number that ranks them.

Tutorial: a bake-off on your own labels.

You need a few hundred decisions from production with known right answers. Ticket routes your team corrected, moderation calls a reviewer confirmed, tool choices an agent got right. Split them 50/50 into a calibration set and a test set.

Step 1: wrap each decider behind one interface.

Jev and Clef share a request shape, Perplexity's is close, and a chat model can play along with log-probabilities. Normalize all of them to "probability per option, plus input tokens billed":

import math, time
from dataclasses import dataclass
from typing import Callable
import numpy as np
from openai import OpenAI

@dataclass
class Decider:
    name: str
    price_in: float    # USD per million input tokens
    price_out: float   # USD per million output tokens (0 for decision models)
    ask: Callable[[str, str, list[str]], tuple[dict[str, float], int, int]]
    # ask(state, question, options) -> ({option: prob}, input_tokens, output_tokens)

def jev_compatible(base_url: str, api_key: str, model: str):
    """Jev, Clef and Clef-flash. Field names follow each vendor's docs as of
    Oct 1, 2026; check them against the current API reference."""
    import httpx
    def ask(state, question, options):
        r = httpx.post(base_url, headers={"Authorization": f"Bearer {api_key}"}, json={
            "model": model,
            "state": state,
            "questions": [{"id": "q", "type": "choice", "question": question, "options": options}],
        }, timeout=10)
        r.raise_for_status()
        body = r.json()
        probs = body["answers"]["q"]["probabilities"]   # {option: p}
        return probs, body["usage"]["input_tokens"], 0
    return ask

def chat_logprobs(model: str, client: OpenAI):
    """Any chat model as a classifier: ask for a single option letter, read logprobs."""
    def ask(state, question, options):
        letters = [chr(65 + i) for i in range(len(options))]
        menu = "\n".join(f"{l}. {o}" for l, o in zip(letters, options))
        r = client.chat.completions.create(
            model=model, max_tokens=1, logprobs=True, top_logprobs=min(20, len(options)),
            messages=[{"role": "user", "content": f"{state}\n\n{question}\n{menu}\nAnswer with one letter."}],
        )
        top = {t.token.strip(): math.exp(t.logprob) for t in r.choices[0].logprobs.content[0].top_logprobs}
        raw = {o: top.get(l, 0.0) for l, o in zip(letters, options)}
        z = sum(raw.values()) or 1.0
        return {o: p / z for o, p in raw.items()}, r.usage.prompt_tokens, r.usage.completion_tokens
    return ask

The chat baseline matters. Without it you can't tell whether a decider is worth adding, only which decider is best.

Step 2: run every row through every decider.

def run(decider: Decider, rows):
    out = []
    for row in rows:   # row: {"state", "question", "options", "label"}
        t0 = time.perf_counter()
        probs, tin, tout = decider.ask(row["state"], row["question"], row["options"])
        ms = (time.perf_counter() - t0) * 1000
        pred = max(probs, key=probs.get)
        cost = (tin * decider.price_in + tout * decider.price_out) / 1e6
        out.append({"conf": probs[pred], "correct": pred == row["label"], "cost": cost, "ms": ms})
    return out

Run it from the same region your service runs in. Chart 3 is why.

Step 3: pick a threshold on the calibration half, score on the test half.

TARGET = 0.98             # accuracy you need on the calls you accept
FALLBACK_COST = 0.00035   # per escalated decision, e.g. Claude Haiku 4.5 at ~300 tokens

def threshold_for(results, target=TARGET):
    """Lowest confidence cutoff whose accepted calls still hit the target."""
    for t in np.arange(0.50, 1.00, 0.01):
        kept = [r for r in results if r["conf"] >= t]
        if kept and np.mean([r["correct"] for r in kept]) >= target:
            return t
    return 1.01           # never accept: everything escalates

def ece(results, bins=10):
    """Expected calibration error: how far stated confidence is from accuracy."""
    conf = np.array([r["conf"] for r in results]); ok = np.array([r["correct"] for r in results])
    edges = np.linspace(0, 1, bins + 1); err = 0.0
    for lo, hi in zip(edges[:-1], edges[1:]):
        m = (conf > lo) & (conf <= hi)
        if m.any():
            err += m.mean() * abs(conf[m].mean() - ok[m].mean())
    return err

def score(decider, calib_rows, test_rows):
    t = threshold_for(run(decider, calib_rows))
    test = run(decider, test_rows)
    kept = [r for r in test if r["conf"] >= t]
    coverage = len(kept) / len(test)
    blended = np.mean([r["cost"] for r in test]) + (1 - coverage) * FALLBACK_COST
    return {
        "decider": decider.name,
        "threshold": round(t, 2),
        "coverage": round(coverage, 3),
        "accepted_acc": round(np.mean([r["correct"] for r in kept]), 3) if kept else None,
        "ece": round(ece(test), 3),
        "p95_ms": round(np.percentile([r["ms"] for r in test], 95)),
        "usd_per_million": round(blended * 1e6, 2),
    }

Rank by usd_per_million, and reject any decider whose accepted_acc on the test half misses the target by more than sampling noise. With 200 test rows, a point or two of difference isn't real. A coverage gap of 20 points is.

Step 4: keep it swappable.

Write your application against the Decider interface, not a vendor SDK. This market will reprice within weeks: Perplexity has already said it plans to cut its price further, and three of the four challenger vendors ship open weights anyone can host. Re-run the bake-off when a price moves or a new model ships. It's a few dollars of API calls.

Where Nadir fits.

Nadir answers one decision on every request: which model does this prompt need? POST /v1/bucket returns simple, medium, or complex with a probability for each, and the classifier behind it is trained on labeled prompts rather than prompted zero-shot, for the reasons in Beaten by len(). Decision models are inputs to that classifier, not a replacement for it, and the bake-off above is how we decide which ones earn a place.

If you're adding a decider to your own pipeline, the bake-off is yours to run. If the decision you actually need is "cheap model or expensive one," send model="auto" and Nadir makes it per prompt, against the quality floor you set, with the decision and its cost in the response headers. Observe mode records what routing would have chosen without changing a single response, so you can size the saving first. Start with a free key.

Conclusion.

Decision models went from one vendor to six in sixteen days, and the category already behaves like a market. Two fine-tunes of the same 27B base list 6x apart, at $12 and $72 per million decisions, and the expensive one costs more per decision than a cheap chat model asked the same question. Every launch beat Jev on its own table, while the per-task detail splits both ways by more than 10 points. Headline latency was 5x lower than one independent client-side test. None of that tells you which decider to use. Coverage at your accuracy target, on your own labels, measured from your own region, does. Build against the shared interface, run the bake-off, and run it again when the prices move.


Sources: [Flavio Copes, "A deep dive into Clef, Cloudflare's decision model"](https://flaviocopes.com/clef/), October 2026. [traictory.com, "Cloudflare ships Clef, an open-weight answer to Jev, with a price critics noticed"](https://traictory.com/news/2026-10-03-cloudflare-clef-decision-models), October 3, 2026, citing The Register and Hacker News. [Dealroom, "Cloudflare launches Clef as TypeSafe's 'decision model' idea spreads across big tech"](https://dealroom.co/news/158476-cloudflare-launches-clef-as-typesafes-decision-model-idea-spreads-across/). [explainx.ai, "pplx-decider: 85.71% vs Jev, $0.04/M"](https://explainx.ai/blog/perplexity-pplx-decider-decisions-api-2026). [Perplexity Developers on X, Decisions API launch](https://x.com/perplexitydevs/status/2105725598882832414). [pplx-decider-v1-27b on OpenRouter](https://openrouter.ai/perplexity/pplx-decider-v1-27b). [BenchLM, "AI Decision Models: GLiDE, Jev & Perplexity Decider"](https://benchlm.ai/decision-models) and [GLiDE benchmarks](https://benchlm.ai/models/fastino-glide). [Fastino, "GLiNER2.5-Decide: An Open-Weight Model for Structured Decision Making"](https://fastino.ai/blog/gliner-2-5-decide-open-weight-decision-model), September 24, 2026. LLM classifier prices from each provider's list price as of October 2026.

More on routing & cascades

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

For coding agents, Nadir connects to Codex, Claude Code, or Cursor and recommends a model tier for delegated work. The agent decides whether to hand the task off and uses its own configured models, so Nadir needs no proxy and no provider keys on that path.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.