Two launches, one week apart, pulling in opposite directions.
On September 15, TypeSafe AI opened early access to Jev, a decision model that can't write text but answers typed questions at $0.042 per million input tokens, with output unbilled. We covered what it is and how it does as a zero-shot router. On September 22, Anthropic shipped Claude Opus 5.5 at $4 / $20 per million input / output tokens, down from Opus 5's $5 / $25, with cache reads cut from $0.50 to $0.20. Anthropic says it "performs at the level of Claude Fable 5.1 on most work." Source: Justin McKelvey
Put them together and you get the obvious question for anyone paying a frontier bill: if the frontier model just got cheaper and the correctness check just got almost free, should I try a cheap model first and let Jev decide when to call Opus 5.5?
The short answer: Jev makes the check cost 0.6% of the answer it might save, so price is no longer what limits a cascade. Opus 5.5 raises the bar the cheap model has to clear, and the accuracy of the check, not its price, decides whether the cascade is worth it. This post has the math, our measured verifier numbers, a simulation, and 70 lines of Python that tell you whether it pays on your own traffic.
Costs in this post are modeled from public list prices. The verifier accuracy figures are Nadir's own measurements on a public coding sample; the cascade outcomes built on them are simulations, not production traces, and are not derived from customer data.
What changed on September 22.
| Model | Input | Cache read | Output | vs Opus 5.5 |
|---|---|---|---|---|
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 | 0.25x |
| Claude Sonnet 5 | $2.00 | $0.20 | $10.00 | 0.5x |
| Claude Opus 5.5 | $4.00 | $0.20 | $20.00 | 1x |
| Claude Opus 5 | $5.00 | $0.50 | $25.00 | 1.25x |
| Claude Fable 5.1 | $10.00 | $0.25 | $50.00 | 2.5x |
Price per million tokens. Opus 5.5 cache writes are $5 (5-minute) and $8 (1-hour); batch is half price.
Two things matter for cascades. The gap between Sonnet 5 and the top of the Opus line shrank from 2.5x to 2x. And on cache reads, the token type agents consume most, Opus 5.5 and Sonnet 5 now cost exactly the same: $0.20. We flagged that last week as the end of cache reads as a reason to downroute.
The bar the cheap model has to clear.
A cascade answers with the cheap model, checks the answer, and escalates to the expensive model when the check fails. Ignoring the check's own cost, it beats "send everything to the big model" only when the share of cheap answers that pass the check is higher than the price ratio between the two models.
On a single 3,000-token call, a Sonnet 5 → Opus 5.5 cascade needs to pass half of all requests just to break even. On a cached agent turn (60,000 input tokens at 90% cache hits, 1,200 output), where the two models now read cache at the same price, it needs 59%. Starting from Haiku 4.5 keeps the bar at 25-30%, which is why a Haiku-first cascade is still interesting after the price cut and a Sonnet-first one is marginal.
Break-even only covers cost. Every cheap answer that passes the check but is wrong is a quality loss that Opus 5.5 wouldn't have shipped, and that part depends entirely on the check.
Why the check used to be the problem.
Before decision models, the practical ways to check an answer were a small classifier you trained yourself or an LLM acting as a judge. An LLM judge reads the prompt plus the draft answer and writes a verdict. With a 3,000-token prompt and an 800-token answer, that's 3,800 input tokens per check:
Using Opus 5.5 to judge whether you needed Opus 5.5 costs 58% of just asking Opus 5.5. A Haiku judge costs 14%, which is most of the margin a Sonnet-first cascade has left. Jev costs $0.00016, about a hundredth of an Opus judge, and returns in 70-500 ms instead of seconds. Source: DataCamp The check is now effectively free. The remaining question is how good it is.
What we measured.
We tested Jev as a reference-free verifier: it sees the prompt and the cheap model's answer, never the expensive model's. That's the only setup that saves money, because a verifier that needs the expensive answer has already paid for it. On an 860-row near-peer coding sample from public benchmarks, scored on the same rows as our deployed cross-encoder:
| Verifier | AUROC, all families | AUROC, coding only |
|---|---|---|
| Deployed cross-encoder | 0.549 | 0.616 |
| Jev, zero-shot | 0.779 | 0.789 |
| Jev, answers shuffled (control) | 0.506 | n/a |
The difference over the cross-encoder is +0.230 (95% interval +0.182 to +0.277). The shuffle control matters: when each prompt is paired with a different prompt's answer, Jev drops to chance, so it's reading the answer, not recognizing a famous benchmark question.
And one hole: on the arenahard_coding slice (112 rows), where the label is an LLM judge's preference rather than a test that passes or fails, Jev scored 0.516. Chance. When "correct" means "a model liked it more," Jev has nothing to check.
For context, our earlier writeup of cascade verifier ceilings found that a verifier that can read the expensive answer reaches 0.961, and the same verifier retrained without it drops to 0.776. Zero-shot Jev lands at the level of our best reference-free result without any training on our data.
What an AUROC of 0.779 buys you.
AUROC says how well scores separate right answers from wrong ones. To run a cascade you pick a threshold, and the threshold sets two rates: how many correct cheap answers you accept, and how many wrong ones slip through. At 0.779 the trade is steep:
| Accept this share of correct cheap answers | Wrong cheap answers that also pass |
|---|---|
| 50% | 14% |
| 60% | 20% |
| 70% | 29% |
| 80% | 40% |
| 90% | 58% |
Operating points on a binormal ROC curve fitted to AUROC 0.779. Your curve will differ; measure it.
Now apply that to an illustrative coding workload. Treat Opus 5.5 as always right. Assume Sonnet 5 gets 92% right, close to the 92.5% it passed in our 400-task execution-graded test, and Haiku 4.5 gets 70%. "Strict" accepts 60% of correct cheap answers; "loose" accepts 80%.
| Strategy | Cost vs all-Opus 5.5 | Extra wrong answers per 1,000 |
|---|---|---|
| All Opus 5.5 | 100% | 0 |
| Sonnet 5 + Jev, strict | 94% | 16 |
| Sonnet 5 + Jev, loose | 74% | 32 |
| Sonnet 5 only | 50% | 80 |
| Haiku 4.5 + Jev, strict | 78% | 61 |
| Haiku 4.5 + Jev, loose | 58% | 121 |
| Haiku 4.5 only | 25% | 300 |
Three readings:
- Jev removes most of the cheap model's errors, not all of them. Sonnet alone ships 80 more wrong answers per 1,000 than Opus 5.5. Behind a loose Jev gate that falls to 32, at 74% of the cost.
- Strict gates barely save anything after the price cut. Sonnet + strict Jev escalates 43% of requests and lands at 94% of all-Opus cost. Before September 22, against Opus 5, the same gate would have saved more, because each escalation it avoids was worth more.
- Haiku-first is where the money is, if you can tolerate the errors. 42% cheaper than all-Opus, at 121 extra wrong answers per 1,000. For a chat reply or a classification, maybe. For code that ships, probably not without a test suite behind it.
Tutorial: find out whether it pays on your traffic.
The numbers above are one fitted curve on one workload. Yours will be different, and the only way to know is to score your own labeled traffic. You need a few hundred prompts where you know whether the cheap model's answer was acceptable: tests that pass, tickets that didn't reopen, extractions that matched.
Step 1: ask Jev whether the answer is correct.
This is the question we send in production. The answer is clipped head and tail, because the conclusion carries most of the correctness signal and a tail-only cut removes it.
import os, requests
JEV_URL = "https://api.typesafe.ai/v1/systemone"
QUESTION = {
"correct": {
"type": "noul",
"instructions": "The answer is correct and actually resolves what the question asked.",
"criteria": {
"true": "The answer reaches the right result. Minor style or verbosity issues do not matter.",
"false": "The answer is wrong, solves a different problem, stops before resolving "
"the question, or contradicts itself.",
},
}
}
def clip(text: str, cap: int) -> str:
# Keep head AND tail: the conclusion carries most of the correctness signal.
if len(text) <= cap:
return text
half = cap // 2
return f"{text[:half]}\n...[trimmed]...\n{text[-half:]}"
def jev_correct(prompt: str, answer: str) -> float:
resp = requests.post(
JEV_URL,
headers={"Authorization": f"Bearer {os.environ['TYPESAFE_API_KEY']}"},
json={
"model": "jev-latest",
"state": {"question": clip(prompt, 4000), "answer": clip(answer, 8000)},
"questions": QUESTION,
},
timeout=3,
)
resp.raise_for_status()
return float(resp.json()["answers"]["correct"]["noul"])
Clipping also caps the check's cost. At 12,000 characters, a check is about 3,000 tokens, or roughly $0.00013, however long the agent's context is. The price of that cap is that Jev doesn't see the full context on long agent turns.
Step 2: pick a threshold from your own rows.
Score each labeled row with jev_correct, then search for the cheapest threshold that keeps shipped errors under a budget you choose. If none beats sending everything to the top model, the function says so.
# USD per 1M tokens: input, output
PRICES = {
"claude-haiku-4-5": (1.00, 5.00),
"claude-sonnet-5": (2.00, 10.00),
"claude-opus-5-5": (4.00, 20.00),
}
JEV_PER_M = 0.042
def call_cost(model: str, inp: int, out: int) -> float:
pi, po = PRICES[model]
return (inp * pi + out * po) / 1e6
def calibrate(rows, cheap: str, top: str, inp: int, out: int, max_wrong: float = 0.02):
"""rows: (jev_prob, cheap_was_correct) pairs from your own labeled traffic.
Returns the cheapest threshold whose shipped-wrong rate stays under max_wrong,
or None when no threshold beats sending everything to `top`."""
c_cheap, c_top = call_cost(cheap, inp, out), call_cost(top, inp, out)
c_check = (inp + out) * JEV_PER_M / 1e6
n, best = len(rows), None
for t in sorted({p for p, _ in rows}):
accepted = [ok for p, ok in rows if p >= t]
wrong = sum(1 for ok in accepted if not ok) / n
if wrong > max_wrong:
continue
escalate = 1 - len(accepted) / n
cost = c_cheap + c_check + escalate * c_top
if cost < c_top and (best is None or cost < best[1]):
best = (t, cost, len(accepted) / n, wrong)
return best
We ran it on 2,000 simulated rows drawn from the same fitted curve, with Sonnet 5 at 92% correct and the 3,000 / 800 token profile:
| Error budget (`max_wrong`) | Result | Accepted from Sonnet | Cost vs all-Opus 5.5 |
|---|---|---|---|
| 1% | None: no threshold pays | n/a | 100% |
| 2% | threshold 0.51 | 65% | 86% |
| 4% | threshold 0.26 | 85% | 66% |
That first row is the honest one. If your product can only afford one extra wrong answer in a hundred, a Sonnet-first cascade with a 0.779 verifier doesn't pay against Opus 5.5, and the code tells you to route straight to the top.
Step 3: run the cascade, with the answer from Step 2.
def answer(prompt: str, policy, call_model):
"""policy: the calibrate() result for this cheap/top pair, or None."""
cheap, top = "claude-sonnet-5", "claude-opus-5-5"
if policy is None: # cascade doesn't pay here: go straight up
return call_model(top, prompt), top
threshold = policy[0]
draft = call_model(cheap, prompt)
try:
p = jev_correct(prompt, draft)
except Exception: # fail toward quality, not toward the draft
p = 0.0
if p >= threshold:
return draft, cheap
return call_model(top, prompt), top
Calibrate per workload, not once per company. A threshold that fits code review won't fit support replies, and the arenahard_coding result says some workloads have no usable threshold at all.
Before you turn it on.
- It sends your prompts and outputs to a third party. Every check posts the prompt and the draft answer to TypeSafe. Treat Jev as a subprocessor: check your data agreements, and skip the check on zero-retention or regulated traffic. That's how we gate it.
- It needs labels that are checkable. Jev does well where "correct" is a fact (tests pass, the number matches). It's at chance where "correct" is a preference.
- Route before you verify. A verifier only decides after the cheap model has spent tokens. A router that sends obviously hard prompts straight to Opus 5.5 avoids paying for drafts that were never going to pass, which matters more now that the break-even rate is 50%.
- Watch the effort setting. Opus 5.5 supports effort levels from low to max. An Opus 5.5 at low effort may beat a Sonnet-plus-Jev cascade on both cost and quality for your workload. Test that as a baseline too.
Where Nadir fits.
Nadir decides per request which tier a prompt needs, before any provider tokens are spent. Opus 5.5 is supported with all five of its effort levels, so you can put it at the top of your model pool and let only the prompts that need it reach it. Our Jev-based routing classifier went to canary on September 20, at about $77 per million routing decisions. The Jev verifier described here is built and measured, and it's opt-in: it stays off unless you enable it, never runs on requests with prompt storage off, and is disabled outright on offline installs. We'll list TypeSafe as a subprocessor before it processes any customer traffic.
What you get on every response is the model that answered and what it cost, so you can see what a cascade actually shipped instead of trusting a simulation, including this one.
It's OpenAI compatible: change the base URL, set model="auto", keep your own provider keys. Start with a free key, send a week of traffic through it, and use the numbers to decide whether a cheap-first cascade pays for you after the Opus 5.5 price cut.
Conclusion.
Jev made the correctness check almost free: $0.00016, against $0.0162 for asking Opus 5.5 to judge itself. Opus 5.5 made the thing you're trying to avoid cheaper, so the cheap model now has to pass half of all requests for a Sonnet-first cascade to break even, and 59% on cached agent turns. Between the two, price is no longer what limits a cascade. Verifier accuracy is. At our measured 0.779, a Jev gate cuts a cheap model's extra errors by more than half and still lets some through. Calibrate on your own labels, set an error budget before you set a threshold, and let the math tell you when to skip the cascade and send the request straight to Opus.
Prices are September 2026 list prices. Verifier AUROCs are Nadir's measurements on an 860-row near-peer coding sample drawn from public benchmarks (September 17, 2026). Operating points, the cost-vs-error table and the calibration table are simulations from a binormal curve fitted to that AUROC with assumed cheap-model accuracy, and are not derived from customer data. Sources: [Justin McKelvey, "Claude Opus 5.5 Pricing (2026)"](https://justinmckelvey.com/blog/claude-opus-5-5). [Digital Applied, "Claude Opus 5.5: Pricing, Benchmarks and Breaking Changes"](https://www.digitalapplied.com/blog/claude-opus-5-5-launch-pricing-benchmarks-2026). [Claude pricing documentation](https://platform.claude.com/docs/en/about-claude/pricing). [TypeSafe AI, "Introducing System One Models & Jev"](https://typesafe.ai/blog/introducing-system-one-models-and-jev). [DataCamp, "Jev: TypeSafe's System One Model"](https://www.datacamp.com/blog/system-one-models-jev). [LangChain, "Building a harness with Jev"](https://www.langchain.com/blog/building-a-harness-with-jev). [Simon Willison on Jev](https://simonwillison.net/2026/Sep/21/jev/).