The $0.0002 Judge

Claude Opus 5.5 cut the frontier price to $4/$20 on September 22, a week after TypeSafe's Jev made a correctness check cost $0.00016. Together they change cascade math. A Jev check costs 0.6% of the Opus 5.5 answer it might save, against 58% for using Opus 5.5 as its own judge, so price no longer limits a cheap-model-first cascade. But the cheaper Opus raises the bar: a Sonnet 5 first cascade must pass 50% of requests to break even, and 59% on cached agent turns, where Sonnet and Opus 5.5 now read cache at the same $0.20. We measured Jev as a reference-free verifier at AUROC 0.779 on 860 coding rows, against 0.549 for our deployed cross-encoder, and at chance on preference-labeled data. In simulation, a loose Jev gate cuts Sonnet's extra errors from 80 to 32 per 1,000 at 74% of all-Opus cost. Includes a Python calibrator that tells you when to skip the cascade entirely.

Published 2026-09-29 by Dor Amir on the Nadir blog.

Filed under Routing & Cascades.

Two launches, one week apart, pulling in opposite directions.

On September 15, TypeSafe AI opened early access to Jev, a decision model that can't write text but answers typed questions at $0.042 per million input tokens, with output unbilled. We covered what it is and how it does as a zero-shot router. On September 22, Anthropic shipped Claude Opus 5.5 at $4 / $20 per million input / output tokens, down from Opus 5's $5 / $25, with cache reads cut from $0.50 to $0.20. Anthropic says it "performs at the level of Claude Fable 5.1 on most work." Source: Justin McKelvey

Put them together and you get the obvious question for anyone paying a frontier bill: if the frontier model just got cheaper and the correctness check just got almost free, should I try a cheap model first and let Jev decide when to call Opus 5.5?

The short answer: Jev makes the check cost 0.6% of the answer it might save, so price is no longer what limits a cascade. Opus 5.5 raises the bar the cheap model has to clear, and the accuracy of the check, not its price, decides whether the cascade is worth it. This post has the math, our measured verifier numbers, a simulation, and 70 lines of Python that tell you whether it pays on your own traffic.

Costs in this post are modeled from public list prices. The verifier accuracy figures are Nadir's own measurements on a public coding sample; the cascade outcomes built on them are simulations, not production traces, and are not derived from customer data.

What changed on September 22.

ModelInputCache readOutputvs Opus 5.5
Claude Haiku 4.5$1.00$0.10$5.000.25x
Claude Sonnet 5$2.00$0.20$10.000.5x
Claude Opus 5.5$4.00$0.20$20.001x
Claude Opus 5$5.00$0.50$25.001.25x
Claude Fable 5.1$10.00$0.25$50.002.5x

Price per million tokens. Opus 5.5 cache writes are $5 (5-minute) and $8 (1-hour); batch is half price.

Two things matter for cascades. The gap between Sonnet 5 and the top of the Opus line shrank from 2.5x to 2x. And on cache reads, the token type agents consume most, Opus 5.5 and Sonnet 5 now cost exactly the same: $0.20. We flagged that last week as the end of cache reads as a reason to downroute.

The bar the cheap model has to clear.

A cascade answers with the cheap model, checks the answer, and escalates to the expensive model when the check fails. Ignoring the check's own cost, it beats "send everything to the big model" only when the share of cheap answers that pass the check is higher than the price ratio between the two models.

Bar chart of the share of requests a cheap model must get past the check for a cascade to beat sending everything to Opus. Haiku 4.5 first: 20% to Opus 5, 25% to Opus 5.5, 30% to Opus 5.5 on a cached agent turn. Sonnet 5 first: 40% to Opus 5, 50% to Opus 5.5, 59% to Opus 5.5 on a cached agent turn.
Bar chart of the share of requests a cheap model must get past the check for a cascade to beat sending everything to Opus. Haiku 4.5 first: 20% to Opus 5, 25% to Opus 5.5, 30% to Opus 5.5 on a cached agent turn. Sonnet 5 first: 40% to Opus 5, 50% to Opus 5.5, 59% to Opus 5.5 on a cached agent turn.

On a single 3,000-token call, a Sonnet 5 → Opus 5.5 cascade needs to pass half of all requests just to break even. On a cached agent turn (60,000 input tokens at 90% cache hits, 1,200 output), where the two models now read cache at the same price, it needs 59%. Starting from Haiku 4.5 keeps the bar at 25-30%, which is why a Haiku-first cascade is still interesting after the price cut and a Sonnet-first one is marginal.

Break-even only covers cost. Every cheap answer that passes the check but is wrong is a quality loss that Opus 5.5 wouldn't have shipped, and that part depends entirely on the check.

Why the check used to be the problem.

Before decision models, the practical ways to check an answer were a small classifier you trained yourself or an LLM acting as a judge. An LLM judge reads the prompt plus the draft answer and writes a verdict. With a 3,000-token prompt and an 800-token answer, that's 3,800 input tokens per check:

Bar chart of the cost of one correctness check as a share of the $0.028 Opus 5.5 answer it might save. Jev: $0.00016, 0.6%, 70 to 500 ms. Haiku 4.5 as judge: $0.0041, 14%. Sonnet 5 as judge: $0.0081, 29%. Opus 5.5 as judge: $0.0162, 58%, over half the answer.
Bar chart of the cost of one correctness check as a share of the $0.028 Opus 5.5 answer it might save. Jev: $0.00016, 0.6%, 70 to 500 ms. Haiku 4.5 as judge: $0.0041, 14%. Sonnet 5 as judge: $0.0081, 29%. Opus 5.5 as judge: $0.0162, 58%, over half the answer.

Using Opus 5.5 to judge whether you needed Opus 5.5 costs 58% of just asking Opus 5.5. A Haiku judge costs 14%, which is most of the margin a Sonnet-first cascade has left. Jev costs $0.00016, about a hundredth of an Opus judge, and returns in 70-500 ms instead of seconds. Source: DataCamp The check is now effectively free. The remaining question is how good it is.

What we measured.

We tested Jev as a reference-free verifier: it sees the prompt and the cheap model's answer, never the expensive model's. That's the only setup that saves money, because a verifier that needs the expensive answer has already paid for it. On an 860-row near-peer coding sample from public benchmarks, scored on the same rows as our deployed cross-encoder:

VerifierAUROC, all familiesAUROC, coding only
Deployed cross-encoder0.5490.616
Jev, zero-shot0.7790.789
Jev, answers shuffled (control)0.506n/a

The difference over the cross-encoder is +0.230 (95% interval +0.182 to +0.277). The shuffle control matters: when each prompt is paired with a different prompt's answer, Jev drops to chance, so it's reading the answer, not recognizing a famous benchmark question.

And one hole: on the arenahard_coding slice (112 rows), where the label is an LLM judge's preference rather than a test that passes or fails, Jev scored 0.516. Chance. When "correct" means "a model liked it more," Jev has nothing to check.

For context, our earlier writeup of cascade verifier ceilings found that a verifier that can read the expensive answer reaches 0.961, and the same verifier retrained without it drops to 0.776. Zero-shot Jev lands at the level of our best reference-free result without any training on our data.

What an AUROC of 0.779 buys you.

AUROC says how well scores separate right answers from wrong ones. To run a cascade you pick a threshold, and the threshold sets two rates: how many correct cheap answers you accept, and how many wrong ones slip through. At 0.779 the trade is steep:

Accept this share of correct cheap answersWrong cheap answers that also pass
50%14%
60%20%
70%29%
80%40%
90%58%

Operating points on a binormal ROC curve fitted to AUROC 0.779. Your curve will differ; measure it.

Now apply that to an illustrative coding workload. Treat Opus 5.5 as always right. Assume Sonnet 5 gets 92% right, close to the 92.5% it passed in our 400-task execution-graded test, and Haiku 4.5 gets 70%. "Strict" accepts 60% of correct cheap answers; "loose" accepts 80%.

Scatter plot of cost per request, as a share of sending everything to Opus 5.5, against extra wrong answers per 1,000 compared with Opus 5.5. All Opus 5.5: 100%, 0. Sonnet only: 50%, 80. Sonnet plus Jev loose: 74%, 32. Sonnet plus Jev strict: 94%, 16. Haiku only: 25%, 300. Haiku plus Jev loose: 58%, 121. Haiku plus Jev strict: 78%, 61.
Scatter plot of cost per request, as a share of sending everything to Opus 5.5, against extra wrong answers per 1,000 compared with Opus 5.5. All Opus 5.5: 100%, 0. Sonnet only: 50%, 80. Sonnet plus Jev loose: 74%, 32. Sonnet plus Jev strict: 94%, 16. Haiku only: 25%, 300. Haiku plus Jev loose: 58%, 121. Haiku plus Jev strict: 78%, 61.
StrategyCost vs all-Opus 5.5Extra wrong answers per 1,000
All Opus 5.5100%0
Sonnet 5 + Jev, strict94%16
Sonnet 5 + Jev, loose74%32
Sonnet 5 only50%80
Haiku 4.5 + Jev, strict78%61
Haiku 4.5 + Jev, loose58%121
Haiku 4.5 only25%300

Three readings:

Tutorial: find out whether it pays on your traffic.

The numbers above are one fitted curve on one workload. Yours will be different, and the only way to know is to score your own labeled traffic. You need a few hundred prompts where you know whether the cheap model's answer was acceptable: tests that pass, tickets that didn't reopen, extractions that matched.

Step 1: ask Jev whether the answer is correct.

This is the question we send in production. The answer is clipped head and tail, because the conclusion carries most of the correctness signal and a tail-only cut removes it.

import os, requests

JEV_URL = "https://api.typesafe.ai/v1/systemone"
QUESTION = {
    "correct": {
        "type": "noul",
        "instructions": "The answer is correct and actually resolves what the question asked.",
        "criteria": {
            "true": "The answer reaches the right result. Minor style or verbosity issues do not matter.",
            "false": "The answer is wrong, solves a different problem, stops before resolving "
                     "the question, or contradicts itself.",
        },
    }
}

def clip(text: str, cap: int) -> str:
    # Keep head AND tail: the conclusion carries most of the correctness signal.
    if len(text) <= cap:
        return text
    half = cap // 2
    return f"{text[:half]}\n...[trimmed]...\n{text[-half:]}"

def jev_correct(prompt: str, answer: str) -> float:
    resp = requests.post(
        JEV_URL,
        headers={"Authorization": f"Bearer {os.environ['TYPESAFE_API_KEY']}"},
        json={
            "model": "jev-latest",
            "state": {"question": clip(prompt, 4000), "answer": clip(answer, 8000)},
            "questions": QUESTION,
        },
        timeout=3,
    )
    resp.raise_for_status()
    return float(resp.json()["answers"]["correct"]["noul"])

Clipping also caps the check's cost. At 12,000 characters, a check is about 3,000 tokens, or roughly $0.00013, however long the agent's context is. The price of that cap is that Jev doesn't see the full context on long agent turns.

Step 2: pick a threshold from your own rows.

Score each labeled row with jev_correct, then search for the cheapest threshold that keeps shipped errors under a budget you choose. If none beats sending everything to the top model, the function says so.

# USD per 1M tokens: input, output
PRICES = {
    "claude-haiku-4-5": (1.00, 5.00),
    "claude-sonnet-5":  (2.00, 10.00),
    "claude-opus-5-5":  (4.00, 20.00),
}
JEV_PER_M = 0.042

def call_cost(model: str, inp: int, out: int) -> float:
    pi, po = PRICES[model]
    return (inp * pi + out * po) / 1e6

def calibrate(rows, cheap: str, top: str, inp: int, out: int, max_wrong: float = 0.02):
    """rows: (jev_prob, cheap_was_correct) pairs from your own labeled traffic.
    Returns the cheapest threshold whose shipped-wrong rate stays under max_wrong,
    or None when no threshold beats sending everything to `top`."""
    c_cheap, c_top = call_cost(cheap, inp, out), call_cost(top, inp, out)
    c_check = (inp + out) * JEV_PER_M / 1e6
    n, best = len(rows), None
    for t in sorted({p for p, _ in rows}):
        accepted = [ok for p, ok in rows if p >= t]
        wrong = sum(1 for ok in accepted if not ok) / n
        if wrong > max_wrong:
            continue
        escalate = 1 - len(accepted) / n
        cost = c_cheap + c_check + escalate * c_top
        if cost < c_top and (best is None or cost < best[1]):
            best = (t, cost, len(accepted) / n, wrong)
    return best

We ran it on 2,000 simulated rows drawn from the same fitted curve, with Sonnet 5 at 92% correct and the 3,000 / 800 token profile:

Error budget (`max_wrong`)ResultAccepted from SonnetCost vs all-Opus 5.5
1%None: no threshold paysn/a100%
2%threshold 0.5165%86%
4%threshold 0.2685%66%

That first row is the honest one. If your product can only afford one extra wrong answer in a hundred, a Sonnet-first cascade with a 0.779 verifier doesn't pay against Opus 5.5, and the code tells you to route straight to the top.

Step 3: run the cascade, with the answer from Step 2.

def answer(prompt: str, policy, call_model):
    """policy: the calibrate() result for this cheap/top pair, or None."""
    cheap, top = "claude-sonnet-5", "claude-opus-5-5"
    if policy is None:                       # cascade doesn't pay here: go straight up
        return call_model(top, prompt), top
    threshold = policy[0]
    draft = call_model(cheap, prompt)
    try:
        p = jev_correct(prompt, draft)
    except Exception:                        # fail toward quality, not toward the draft
        p = 0.0
    if p >= threshold:
        return draft, cheap
    return call_model(top, prompt), top

Calibrate per workload, not once per company. A threshold that fits code review won't fit support replies, and the arenahard_coding result says some workloads have no usable threshold at all.

Before you turn it on.

Where Nadir fits.

Nadir decides per request which tier a prompt needs, before any provider tokens are spent. Opus 5.5 is supported with all five of its effort levels, so you can put it at the top of your model pool and let only the prompts that need it reach it. Our Jev-based routing classifier went to canary on September 20, at about $77 per million routing decisions. The Jev verifier described here is built and measured, and it's opt-in: it stays off unless you enable it, never runs on requests with prompt storage off, and is disabled outright on offline installs. We'll list TypeSafe as a subprocessor before it processes any customer traffic.

What you get on every response is the model that answered and what it cost, so you can see what a cascade actually shipped instead of trusting a simulation, including this one.

It's OpenAI compatible: change the base URL, set model="auto", keep your own provider keys. Start with a free key, send a week of traffic through it, and use the numbers to decide whether a cheap-first cascade pays for you after the Opus 5.5 price cut.

Conclusion.

Jev made the correctness check almost free: $0.00016, against $0.0162 for asking Opus 5.5 to judge itself. Opus 5.5 made the thing you're trying to avoid cheaper, so the cheap model now has to pass half of all requests for a Sonnet-first cascade to break even, and 59% on cached agent turns. Between the two, price is no longer what limits a cascade. Verifier accuracy is. At our measured 0.779, a Jev gate cuts a cheap model's extra errors by more than half and still lets some through. Calibrate on your own labels, set an error budget before you set a threshold, and let the math tell you when to skip the cascade and send the request straight to Opus.


Prices are September 2026 list prices. Verifier AUROCs are Nadir's measurements on an 860-row near-peer coding sample drawn from public benchmarks (September 17, 2026). Operating points, the cost-vs-error table and the calibration table are simulations from a binormal curve fitted to that AUROC with assumed cheap-model accuracy, and are not derived from customer data. Sources: [Justin McKelvey, "Claude Opus 5.5 Pricing (2026)"](https://justinmckelvey.com/blog/claude-opus-5-5). [Digital Applied, "Claude Opus 5.5: Pricing, Benchmarks and Breaking Changes"](https://www.digitalapplied.com/blog/claude-opus-5-5-launch-pricing-benchmarks-2026). [Claude pricing documentation](https://platform.claude.com/docs/en/about-claude/pricing). [TypeSafe AI, "Introducing System One Models & Jev"](https://typesafe.ai/blog/introducing-system-one-models-and-jev). [DataCamp, "Jev: TypeSafe's System One Model"](https://www.datacamp.com/blog/system-one-models-jev). [LangChain, "Building a harness with Jev"](https://www.langchain.com/blog/building-a-harness-with-jev). [Simon Willison on Jev](https://simonwillison.net/2026/Sep/21/jev/).

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.