When LLM Routing Does Not Help

LLM routing does not help when your cheap model already passes, when a switch throws away a warm prompt cache, or when the signal is no better than a coin. This post collects Nadir's own negative results from June to September 2026. On 400 execution-graded coding tasks, Sonnet 5 and Opus 5 disagreed on only 32. One mid-session switch from Haiku 4.5 to Opus 5 costs as much as 61 more Haiku turns on the cached prefix. A turn-count escalation rule on 500 SWE-bench Verified tasks tied a random control at the same rate. An LLM judge picking between near-peer models lost to a fixed order. A per-domain prior borrowed from one benchmark ranked another benchmark's domains backwards. A cascade verifier reached its research-ceiling AUROC of 0.961 only when it could read the expensive model's answer; retrained without that reference, it scores 0.776. Each section says what to do instead.

Published 2026-09-28 by Dor Amir on the Nadir blog.

Filed under Routing & Cascades.

The short answer.

LLM routing does not help when the cheap model already passes almost everything, when a model switch throws away a warm prompt cache, when the signal that triggers an escalation is no better than a coin flip, when the router asks a model to judge fit, when it borrows priors from someone else's benchmark, or when the verifier behind a cascade cannot see what it is grading. We know because we built each of those and measured it. Most of them lost.

This post collects Nadir's own negative results from June to September 2026. Every number below comes from an internal eval we ran, with the date, the dataset and the sample size where we have them. None of it is a forecast for your traffic, and we still publish no universal production savings or quality percentage. What it gives you is a list of places where a router costs more than it saves, and what to do instead.

1. When the cheap model already passes.

The situation. You add a router in front of a workload where the cheaper model is already right most of the time. The router's job is to find the few prompts where only the expensive model succeeds. If there are very few, there is very little to route.

What we measured. On 31 August 2026 we took 400 HumanEval and MBPP tasks and ran each through Claude Sonnet 5 and Claude Opus 5, graded by executing the tests. Sonnet 5 passed 370 of 400. Opus 5 passed 388. Only 32 tasks, 8% of the set, had different outcomes on the two models, and on 7 of those 32 it was Opus that failed and Sonnet that passed. On the other 92% there is no decision to get right: both models pass or both fail.

Why. A router's upside is capped by the discordant set, the prompts where the models disagree. When the cheap model's pass rate is high, that set is small, and part of it points the wrong way. A router that is 80% accurate on a decision surface of 8% changes the outcome of a few prompts in a hundred, and pays its own latency and error rate on all hundred.

What to do instead. Measure the discordant rate on a sample of your real prompts before you buy or build anything. If the cheap model passes nearly everything, send everything to the cheap model and spend the effort on a check that catches the rest. For code with tests, that check is running the tests (more on that below). The LLMRouterBench result, where most routers failed to beat a simple size-based baseline, is the same effect at a much larger scale.

2. When a switch throws away a warm prompt cache.

The situation. An agent session builds a long cached prefix on one model, then the router sends a later turn to a different model. The cache is keyed to the model, so the new model processes the whole prefix from scratch.

What we measured. On 18 September 2026 we priced one mid-session switch against letting the current model continue, at Anthropic list input prices ($1, $2 and $5 per million tokens for Haiku 4.5, Sonnet 5 and Opus 5), a cache read at 0.1x and a five-minute cache write at 1.25x. The penalty scales with the prefix, so the ratio holds for any prefix length:

SwitchCosts the same as
Haiku to Sonnet24 more Haiku turns reading the cached prefix
Haiku to Opus61 more Haiku turns
Sonnet to Opus30 more Sonnet turns

That is the best case for switching. It assumes the new call sets its own cache marker, so the write premium is paid once.

The same arithmetic killed a different idea on 12 August 2026: having the router emit a per-task output-style instruction (terse prose, minimal code) alongside the model choice. The instruction would vary from task to task, and changing it changes the prefix. On a 54,000-token cached prefix at Opus rates, rewriting the cache costs $0.3375 against $0.027 to read it, about $0.31 per flip, to save roughly $0.006 of output on that turn. One flip cost about fifty times what it saved.

Why. Prompt caching is the largest discount in a long agent session, and it belongs to one model and one exact prefix. A router that decides per turn keeps resetting it.

What to do instead. Decide the model at the start of a thread and hold it. If a thread must change model, do it at a natural break where the prefix would be rebuilt anyway. Keep anything that varies per turn out of the cached prefix. The full break-even math is in Your Router Is Fighting Your Prompt Cache.

3. When the escalation signal is no better than a coin flip.

The situation. A coding agent starts on a cheap model and hands the conversation to an expensive one when it looks stuck: too many turns, repeated errors, the user asking again.

What we measured. On 18 September 2026 we replayed this offline over published per-instance results for 500 SWE-bench Verified tasks, escalating from Claude Haiku 4.5 to Claude Opus 4.6 once a run passed a turn count. Always-Haiku resolved 66.6% of tasks and always-Opus 75.6%. Escalating after 25 turns matched Opus at 75.6%, but it escalated 97.2% of tasks. Then we ran the control that matters: escalate the same fraction of tasks, chosen at random. Random resolved 75.4%, and its modeled cost was within 1% of the depth rule's. The turn count does carry some signal. Its AUROC was 0.700 for "the cheap model failed", but only 0.596 for the case escalation exists to catch, "the cheap model failed and the expensive one would have succeeded", which was 12.2% of tasks.

A second struggle signal failed harder. On 20 August 2026 we tried to detect "the user had to correct the agent" by comparing consecutive user messages for similarity. Replayed over six real Claude Code transcripts, about 8,100 messages, it fired on 101 of 8,008 windows. Every top trigger was text the harness injects into the conversation, such as hook feedback and image placeholders. None of them was a user correction. Wired to escalation, it would have served the more expensive model for no reason.

Why. Handing a conversation to a stronger model after some fixed amount of work pays off because the stronger model finishes the job. It does not show that the trigger picked the right tasks. When nearly every task escalates, the policy is a pipeline, not a decision. And at the wire level, what the harness adds to a conversation looks a lot like what a frustrated user types.

What to do instead. Always test an escalation rule against a random control that escalates at the same rate. If they tie, drop the signal and keep the simpler pipeline, or skip escalation. Only let a trigger that has a labeled outcome behind it reach model selection. Tool-error counts qualify. Text similarity does not.

4. When you ask an LLM which model fits.

The situation. Instead of training a classifier, you ask a fast model to read the prompt and the candidate models and predict which one will succeed.

What we measured. On 28 August 2026, in a test whose pass bar we fixed before we saw any data, Claude Haiku judged 880 coding-weighted LLMRouterBench prompts against a panel of near-peer frontier models. On the pairs where the models disagreed, the judge that saw model names scored an AUROC of 0.537. The judge that saw anonymous capability profiles scored 0.597. Always picking models in one fixed order scored 0.606, better than both. Knowing the models' reputations made the judge worse, not better.

We replicated it on 31 August 2026 with a local open-weight judge on the 400 HumanEval and MBPP tasks from section 1. Asked for a 0 to 100 success score, the judge said 95 for nearly every task. Replacing the task text with "(withheld) a Python coding task" produced identical scores on every row. A separate run with a longer, carefully written judge prompt, on two judge families, landed at AUROC 0.483 and 0.521.

Why. Sonnet passes 92.5% of that set, so a calibrated guess really is about 0.95 for almost everything. Asking for an absolute probability collapses to the base rate. Asking which of several models fits pushes the judge onto memorized reputations, and those misorder near-peer models. The per-prompt information needed to separate two frontier models is mostly not in the prompt.

What to do instead. Use LLM judges for relative difficulty, where they did show real signal, and let measured outcomes on your own traffic choose between near-peer models. Do not let a model pick another model by name.

5. When the priors come from someone else's benchmark.

The situation. You have no labels for your own traffic, so you borrow per-domain statistics from a public benchmark: escalate on the domains where the cheap models tend to be wrong, hold on the rest.

What we measured. On 22 August 2026 we fitted exactly that prior for the agreement gate of one of our consensus cascades, using LLMRouterBench, froze it, and scored it once on RouterArena. It changed the arena score by -0.04 points. Where it had an opinion, it ranked domains backwards: it said escalate on open QA, which turned out to be the second-worst place to escalate, and hold on math, which was the best. We re-fitted with our own RouterBench labels, 7,855 rows across 69 task families with the exact models involved. The frozen prior came out empty, and the rank correlation between what RouterBench predicted and what RouterArena showed across 7 domains was 0.00. Even a perfect domain-level selector, fitted on the test set and not achievable in practice, was worth under one point.

Why. How hard a domain is for a given set of models depends on the corpus, not on the domain label. "Open QA" in one benchmark and "open QA" in another are different tasks with the same name.

What to do instead. Treat public per-domain numbers as a starting guess to be replaced, not a policy. Collect outcome labels on your own traffic, and route on them once you have enough.

6. When the verifier cannot see what it is grading.

The situation. A cascade sends every prompt to the cheap model first and lets a verifier decide whether to accept the answer or escalate. The whole saving depends on how well the verifier tells good answers from bad ones.

What we measured. Our widely cited RouterBench cascade figures (about 60% lower cost than always-Opus with about 98% of quality retained, verifier AUROC 0.961) are a reference-assisted research ceiling, not deployable performance: the verifier was given the expensive model's answer, which a cost-saving cascade never has. The evidence correction and the original benchmark post explain the setup. Scored the way production runs it, with no reference, that verifier fell to about 0.52, close to random, on 2,000 held-out triples in June 2026.

We retrained it reference-free on 16 June 2026 on the full RouterBench split (89,621 training and 11,420 test triples). Held-out AUROC was 0.776, ECE 0.028. Two follow-ups explain the gap. Giving the old verifier a cheap model's answer as the reference instead scored 0.689 with a mid-tier peer and 0.585 with a weak one, against 0.855 with the expensive answer (1,500 held-out triples). The signal tracks the quality of the reference, and a good reference is exactly the expensive call you were trying to skip. Then, on 21 September 2026, we found that the median answer in our verifier pool is 2,166 tokens and the verifier read only 320 of them, so it saw the whole answer in 9.8% of cases. Clipping a much larger judge model to the same 320 tokens dropped it from 0.846 to 0.792, within 0.013 of our small verifier at 0.778, on 4,636 test cells. About 80% of the gap was context, not model size.

Why. A cascade is only as good as its accept/escalate decision. Without a reference answer and without the full response, that decision is hard, and every point of verifier AUROC turns directly into either lost savings or bad answers served.

What to do instead. Ask which regime any cascade number comes from before you trust it. Verify where you have a real check (tests, schemas, exact answers), and give a learned verifier the whole response. When a response fails only on format, repair it before paying to escalate; Right Answer, Wrong Shape walks through that.

When routing does help.

None of this means routing never pays. It pays when the models really disagree on a meaningful share of your prompts, when the decision is made once per thread instead of once per turn, and when a real check backs the cheap answer. Our strongest results have that shape:

If you are deciding whether you need a router at all, start with what an LLM router is and when it is not worth it, and how a router differs from the AI gateway you may already run. Then measure your own discordant rate before you believe anyone's savings number, ours included.

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.