Beaten by the Baseline

LLMRouterBench, presented at ACL 2026 Findings, benchmarked 10 routing methods across 33 models, 21 datasets, and 400,000+ query instances. The headline result: most commercial routers fail to beat a size-based routing baseline. The study isolated the failure mode — oracle gap caused by model recall failures, not routing uncertainty — and found embedding backbone choice changes quality by less than 2%. One ensemble approach surpassed GPT-5-medium by 7% at 63% lower cost.

Published 2026-06-29 by Dor Amir on the Nadir blog.

Filed under Routing & Cascades.

The largest routing benchmark ever published found a surprising result.

LLMRouterBench (arXiv:2601.07206, ACL 2026 Findings) benchmarked 10 routing methods across 33 models, 21 datasets, and over 400,000 query instances. It is the most comprehensive evaluation of LLM routing to date.

The headline finding: most commercial routers fail to outperform a simple size-based baseline.


What the benchmark measured.

The 10 routing methods tested included embedding-based classifiers, reward model routing, LLM-as-judge routing, cascade systems, and commercial routing APIs. All were evaluated across three dimensions:

33 models were included in the pool — from 7B open-source to GPT-5.5 and Claude Opus 4.8 — across 21 datasets spanning coding, reasoning, summarization, and factual QA.


The failure mode.

Most routers underperform because they optimize for query difficulty rather than the actual failure mode they face: model recall failures.

When a cheap model fails a query, it is rarely because the query was genuinely hard. It is because cheap model failures are inconsistent. The same query, rephrased slightly, succeeds. The same query, unchanged, fails on a different sample. Routers trained to detect difficulty learn from consistent failure patterns. When failure is stochastic, the training signal is noise.

The study also found that embedding backbone choice changes routing quality by less than 2%. Swapping BERT for a larger model, or using domain-specific fine-tuned embeddings, made negligible difference. Most teams over-invest in the embedding pipeline and under-invest in the model pool.


What worked.

Ensemble routing consistently outperformed single-classifier baselines. Combining multiple independent predictors averages out stochastic failures that break any single classifier.

Avengers-Pro (arXiv:2508.12631, ACM DAI 2025) pushed this further: instead of routing to one best model, it assembles a dynamic committee weighted by query type and domain. On the RouterBench evaluation set, it surpassed GPT-5-medium by 7% on quality while operating at 63% lower cost.

RouterEval (EMNLP 2025) confirmed a parallel finding: adding more models to the routing pool improves quality more reliably than improving the router itself. More candidates mean better coverage of query types.


Three tests before buying a commercial router.

TestWhat to check
Beat the size baselineRoute cheap by default, frontier for hard queries by word count or domain. If the router doesn't beat this, it adds no value.
Check oracle gap by query typeAggregate 95% quality can hide a 40% gap on specific distributions. Audit by domain, not just average.
Skip embedding optimizationThe data shows it doesn't move the needle. Expand the model pool instead.

The practical implication.

The ACL 2026 findings confirm what practitioners are discovering independently: routing quality is dominated by model pool breadth and ensemble signal aggregation, not router sophistication.

If your routing setup is underperforming, the most likely cause is a training distribution mismatch or an insufficient model pool — not the classifier architecture.

Nadir uses multi-signal routing with dynamic threshold calibration against observed production performance, directly addressing the oracle gap the study identifies. See where your current setup is leaving cost or quality on the table.


Sources: [LLMRouterBench (arXiv:2601.07206)](https://arxiv.org/abs/2601.07206), ACL 2026 Findings. [Avengers-Pro (arXiv:2508.12631)](https://arxiv.org/abs/2508.12631), ACM DAI 2025. RouterEval, EMNLP 2025.

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.