The short answer.
LLM routing does not help when the cheap model already passes almost everything, when a model switch throws away a warm prompt cache, when the signal that triggers an escalation is no better than a coin flip, when the router asks a model to judge fit, when it borrows priors from someone else's benchmark, or when the verifier behind a cascade cannot see what it is grading. We know because we built each of those and measured it. Most of them lost.
This post collects Nadir's own negative results from June to September 2026. Every number below comes from an internal eval we ran, with the date, the dataset and the sample size where we have them. None of it is a forecast for your traffic, and we still publish no universal production savings or quality percentage. What it gives you is a list of places where a router costs more than it saves, and what to do instead.
1. When the cheap model already passes.
The situation. You add a router in front of a workload where the cheaper model is already right most of the time. The router's job is to find the few prompts where only the expensive model succeeds. If there are very few, there is very little to route.
What we measured. On 31 August 2026 we took 400 HumanEval and MBPP tasks and ran each through Claude Sonnet 5 and Claude Opus 5, graded by executing the tests. Sonnet 5 passed 370 of 400. Opus 5 passed 388. Only 32 tasks, 8% of the set, had different outcomes on the two models, and on 7 of those 32 it was Opus that failed and Sonnet that passed. On the other 92% there is no decision to get right: both models pass or both fail.
Why. A router's upside is capped by the discordant set, the prompts where the models disagree. When the cheap model's pass rate is high, that set is small, and part of it points the wrong way. A router that is 80% accurate on a decision surface of 8% changes the outcome of a few prompts in a hundred, and pays its own latency and error rate on all hundred.
What to do instead. Measure the discordant rate on a sample of your real prompts before you buy or build anything. If the cheap model passes nearly everything, send everything to the cheap model and spend the effort on a check that catches the rest. For code with tests, that check is running the tests (more on that below). The LLMRouterBench result, where most routers failed to beat a simple size-based baseline, is the same effect at a much larger scale.
2. When a switch throws away a warm prompt cache.
The situation. An agent session builds a long cached prefix on one model, then the router sends a later turn to a different model. The cache is keyed to the model, so the new model processes the whole prefix from scratch.
What we measured. On 18 September 2026 we priced one mid-session switch against letting the current model continue, at Anthropic list input prices ($1, $2 and $5 per million tokens for Haiku 4.5, Sonnet 5 and Opus 5), a cache read at 0.1x and a five-minute cache write at 1.25x. The penalty scales with the prefix, so the ratio holds for any prefix length:
| Switch | Costs the same as |
|---|---|
| Haiku to Sonnet | 24 more Haiku turns reading the cached prefix |
| Haiku to Opus | 61 more Haiku turns |
| Sonnet to Opus | 30 more Sonnet turns |
That is the best case for switching. It assumes the new call sets its own cache marker, so the write premium is paid once.
The same arithmetic killed a different idea on 12 August 2026: having the router emit a per-task output-style instruction (terse prose, minimal code) alongside the model choice. The instruction would vary from task to task, and changing it changes the prefix. On a 54,000-token cached prefix at Opus rates, rewriting the cache costs $0.3375 against $0.027 to read it, about $0.31 per flip, to save roughly $0.006 of output on that turn. One flip cost about fifty times what it saved.
Why. Prompt caching is the largest discount in a long agent session, and it belongs to one model and one exact prefix. A router that decides per turn keeps resetting it.
What to do instead. Decide the model at the start of a thread and hold it. If a thread must change model, do it at a natural break where the prefix would be rebuilt anyway. Keep anything that varies per turn out of the cached prefix. The full break-even math is in Your Router Is Fighting Your Prompt Cache.
3. When the escalation signal is no better than a coin flip.
The situation. A coding agent starts on a cheap model and hands the conversation to an expensive one when it looks stuck: too many turns, repeated errors, the user asking again.
What we measured. On 18 September 2026 we replayed this offline over published per-instance results for 500 SWE-bench Verified tasks, escalating from Claude Haiku 4.5 to Claude Opus 4.6 once a run passed a turn count. Always-Haiku resolved 66.6% of tasks and always-Opus 75.6%. Escalating after 25 turns matched Opus at 75.6%, but it escalated 97.2% of tasks. Then we ran the control that matters: escalate the same fraction of tasks, chosen at random. Random resolved 75.4%, and its modeled cost was within 1% of the depth rule's. The turn count does carry some signal. Its AUROC was 0.700 for "the cheap model failed", but only 0.596 for the case escalation exists to catch, "the cheap model failed and the expensive one would have succeeded", which was 12.2% of tasks.
A second struggle signal failed harder. On 20 August 2026 we tried to detect "the user had to correct the agent" by comparing consecutive user messages for similarity. Replayed over six real Claude Code transcripts, about 8,100 messages, it fired on 101 of 8,008 windows. Every top trigger was text the harness injects into the conversation, such as hook feedback and image placeholders. None of them was a user correction. Wired to escalation, it would have served the more expensive model for no reason.
Why. Handing a conversation to a stronger model after some fixed amount of work pays off because the stronger model finishes the job. It does not show that the trigger picked the right tasks. When nearly every task escalates, the policy is a pipeline, not a decision. And at the wire level, what the harness adds to a conversation looks a lot like what a frustrated user types.
What to do instead. Always test an escalation rule against a random control that escalates at the same rate. If they tie, drop the signal and keep the simpler pipeline, or skip escalation. Only let a trigger that has a labeled outcome behind it reach model selection. Tool-error counts qualify. Text similarity does not.
4. When you ask an LLM which model fits.
The situation. Instead of training a classifier, you ask a fast model to read the prompt and the candidate models and predict which one will succeed.
What we measured. On 28 August 2026, in a test whose pass bar we fixed before we saw any data, Claude Haiku judged 880 coding-weighted LLMRouterBench prompts against a panel of near-peer frontier models. On the pairs where the models disagreed, the judge that saw model names scored an AUROC of 0.537. The judge that saw anonymous capability profiles scored 0.597. Always picking models in one fixed order scored 0.606, better than both. Knowing the models' reputations made the judge worse, not better.
We replicated it on 31 August 2026 with a local open-weight judge on the 400 HumanEval and MBPP tasks from section 1. Asked for a 0 to 100 success score, the judge said 95 for nearly every task. Replacing the task text with "(withheld) a Python coding task" produced identical scores on every row. A separate run with a longer, carefully written judge prompt, on two judge families, landed at AUROC 0.483 and 0.521.
Why. Sonnet passes 92.5% of that set, so a calibrated guess really is about 0.95 for almost everything. Asking for an absolute probability collapses to the base rate. Asking which of several models fits pushes the judge onto memorized reputations, and those misorder near-peer models. The per-prompt information needed to separate two frontier models is mostly not in the prompt.
What to do instead. Use LLM judges for relative difficulty, where they did show real signal, and let measured outcomes on your own traffic choose between near-peer models. Do not let a model pick another model by name.
5. When the priors come from someone else's benchmark.
The situation. You have no labels for your own traffic, so you borrow per-domain statistics from a public benchmark: escalate on the domains where the cheap models tend to be wrong, hold on the rest.
What we measured. On 22 August 2026 we fitted exactly that prior for the agreement gate of one of our consensus cascades, using LLMRouterBench, froze it, and scored it once on RouterArena. It changed the arena score by -0.04 points. Where it had an opinion, it ranked domains backwards: it said escalate on open QA, which turned out to be the second-worst place to escalate, and hold on math, which was the best. We re-fitted with our own RouterBench labels, 7,855 rows across 69 task families with the exact models involved. The frozen prior came out empty, and the rank correlation between what RouterBench predicted and what RouterArena showed across 7 domains was 0.00. Even a perfect domain-level selector, fitted on the test set and not achievable in practice, was worth under one point.
Why. How hard a domain is for a given set of models depends on the corpus, not on the domain label. "Open QA" in one benchmark and "open QA" in another are different tasks with the same name.
What to do instead. Treat public per-domain numbers as a starting guess to be replaced, not a policy. Collect outcome labels on your own traffic, and route on them once you have enough.
6. When the verifier cannot see what it is grading.
The situation. A cascade sends every prompt to the cheap model first and lets a verifier decide whether to accept the answer or escalate. The whole saving depends on how well the verifier tells good answers from bad ones.
What we measured. Our widely cited RouterBench cascade figures (about 60% lower cost than always-Opus with about 98% of quality retained, verifier AUROC 0.961) are a reference-assisted research ceiling, not deployable performance: the verifier was given the expensive model's answer, which a cost-saving cascade never has. The evidence correction and the original benchmark post explain the setup. Scored the way production runs it, with no reference, that verifier fell to about 0.52, close to random, on 2,000 held-out triples in June 2026.
We retrained it reference-free on 16 June 2026 on the full RouterBench split (89,621 training and 11,420 test triples). Held-out AUROC was 0.776, ECE 0.028. Two follow-ups explain the gap. Giving the old verifier a cheap model's answer as the reference instead scored 0.689 with a mid-tier peer and 0.585 with a weak one, against 0.855 with the expensive answer (1,500 held-out triples). The signal tracks the quality of the reference, and a good reference is exactly the expensive call you were trying to skip. Then, on 21 September 2026, we found that the median answer in our verifier pool is 2,166 tokens and the verifier read only 320 of them, so it saw the whole answer in 9.8% of cases. Clipping a much larger judge model to the same 320 tokens dropped it from 0.846 to 0.792, within 0.013 of our small verifier at 0.778, on 4,636 test cells. About 80% of the gap was context, not model size.
Why. A cascade is only as good as its accept/escalate decision. Without a reference answer and without the full response, that decision is hard, and every point of verifier AUROC turns directly into either lost savings or bad answers served.
What to do instead. Ask which regime any cascade number comes from before you trust it. Verify where you have a real check (tests, schemas, exact answers), and give a learned verifier the whole response. When a response fails only on format, repair it before paying to escalate; Right Answer, Wrong Shape walks through that.
When routing does help.
None of this means routing never pays. It pays when the models really disagree on a meaningful share of your prompts, when the decision is made once per thread instead of once per turn, and when a real check backs the cheap answer. Our strongest results have that shape:
- Checkable code. Run-check-escalate, where the cheap model's code is executed against its tests and escalated only on failure, solved 392 of 395 common HumanEval and MBPP problems (99.2%), against 383 of 395 for Claude Opus 5 alone. That applies only to tasks with runnable, deterministic tests.
- Per-task tiers on agent work. An offline replay of a tier policy over published per-instance SWE-bench results resolved 371 of 500 tasks at 39.8% lower replay cost than always-Opus. It is a replay, not a live submission or the deployed classifier.
If you are deciding whether you need a router at all, start with what an LLM router is and when it is not worth it, and how a router differs from the AI gateway you may already run. Then measure your own discordant rate before you believe anyone's savings number, ours included.