RouterArena Benchmark: How Far LLM Routers Fall From the Oracle

RouterArena, accepted at ICLR 2026, is the first standardized, independent benchmark to evaluate LLM routers rather than the models they route between. A team including Yifan Lu and Jiarong Xing published it in September 2025 (arXiv:2510.00202). Its dataset spans 8,400 queries across nine domains and 44 categories, built from 23 existing datasets and difficulty-graded on Bloom's taxonomy, and it scores 27 routers on five axes at once: accuracy, cost, routing optimality against a theoretical oracle, robustness to input paraphrase, and latency overhead. The headline finding isn't who won. It's that nobody came close to winning everything: "no router ranks at the top across all metrics," the paper states, and even the router that came closest to the oracle's accuracy, MIRT-BERT, still spent roughly five times what a perfect-foresight router would have spent to get there. The cheapest, most accurate cluster in the whole field tops out around 60-67% raw accuracy, priced an order of magnitude below a second cluster that costs substantially more for comparable or worse results. Every router tested shares the same failure mode: hedging up to the expensive model even when a cheap one would have answered correctly. Here's what RouterArena actually measures, what its numbers say about the state of LLM routing in 2026, and four questions worth asking before trusting any router's own benchmark, including ours.

Published 2026-08-17 by Dor Amir on the Nadir blog.

Filed under Routing & Cascades.

Every LLM router on the market makes the same pitch: send us your traffic, and we'll pick the cheapest model that can handle each request. Almost none of them get evaluated by anyone but themselves. A router vendor's benchmark is run on the vendor's own dataset, against the vendor's own baseline, reported in the vendor's own units, and there has been no standardized, independent way to check whether the number on the landing page means what it claims to mean.

RouterArena, published as a conference paper at ICLR 2026 by a team including Yifan Lu and Jiarong Xing (arXiv:2510.00202), set out to close that gap. It built a fixed, standardized 8,400-query evaluation set, scored 27 routers on it, and asked a question no vendor benchmark is built to answer: when a router says it picked the right model, compared to what, exactly? The result isn't a single winner. It's a field of 27 systems, none of which comes close to the theoretical best a router could do, and a headline finding the paper states in one flat sentence: "no router ranks at the top across all metrics."

What RouterArena actually measures

The dataset is built to resist the two ways a router benchmark usually gets gamed: an easy test set, and a single metric. It draws from 23 existing open datasets, spans nine domains modeled on the Dewey Decimal Classification (everything except religion), and covers 44 categories. Every query gets a difficulty label, easy, medium, or hard, mapped onto Bloom's taxonomy (Remember/Understand, Apply, Analyze/Evaluate), assigned by DeepSeek-V3 and then rebalanced so no single category dominates the harder tiers. A separate 420-query subset tests the same routers under paraphrased, perturbed versions of the same questions, to see whether a routing decision holds up when the input isn't word-for-word identical to what the router was tuned on.

RouterArena factDetail
Dataset size8,400 queries, 809-query and 420-query subsets for lighter runs
Domains / categories9 domains (Dewey Decimal, minus religion), 44 categories
Difficulty tiersEasy, medium, hard, via Bloom's taxonomy, DeepSeek-V3-labeled
Source datasets23 existing open datasets, deduplicated by cosine similarity
Routers evaluated27, spanning commercial (GPT-5-as-router, NotDiamond, Azure Router) and academic (RouterBench KNN/MLP, GraphRouter, CARROT, RouterDC, IRT-Router, RouteLLM, vLLM Semantic Router, MIRT-BERT)
MetricsAccuracy, cost, routing optimality vs. an oracle, robustness to perturbation, latency overhead

On top of the five metrics sits a single composite ranking number, the Arena Score, a weighted harmonic mean of accuracy and cost:

S = (1 + β) · A · C / (β · A + C)

A = router accuracy (0-1)
C = normalized cost, log2-scaled between $0.0044 and $200 per 1,000 queries
β = 0.1 (default), weights accuracy above cost

That formula matters more than it looks. A harmonic mean punishes a router for being lopsided, brilliant on accuracy but wildly expensive, or cheap but wrong too often, in a way a simple average wouldn't. It's the same instinct behind why a router that always calls the frontier model scores badly here even with near-perfect accuracy: cost is baked into the ranking, not reported as a footnote next to it.

Even the best-costed router in a 27-router field isn't close to the oracle. MIRT-BERT reaches 77% of the oracle's accuracy at roughly 5x the oracle's cost — the tightest gap RouterArena found across 27 routers tested.
Even the best-costed router in a 27-router field isn't close to the oracle. MIRT-BERT reaches 77% of the oracle's accuracy at roughly 5x the oracle's cost — the tightest gap RouterArena found across 27 routers tested.

The results: everyone hedges upward

RouterArena's comparison point isn't a competitor's number, it's an oracle: a router with perfect foresight that always picks the cheapest model that would have answered a given query correctly. No real router can match an oracle, since that requires knowing the answer before choosing a model. What RouterArena measures is how far short of it every real router falls, and the paper's finding is that all 27 fall short the same way: by over-relying on expensive models even on queries a cheap model would have gotten right.

Two clusters emerge. Azure-Router, MIRT-BERT, and NIRT-BERT land in the 60-67% raw accuracy range at roughly one order of magnitude lower cost than a second cluster, MLP, KNN, NotDiamond, and vLLM Semantic Router, which cost substantially more for comparable or worse accuracy. MIRT-BERT comes closest to the accuracy-cost frontier of anything tested: about 77% of the oracle's accuracy, at roughly five times the oracle's cost. That's the best result in the field. On the long-context subset specifically, GPT-5-as-router (always route to GPT-5, no selection at all) posts the highest raw accuracy, unsurprisingly, at the highest cost, while Azure-Router and NIRT-BERT hold competitive accuracy at a fraction of that spend.

ClusterAccuracyCost, relative to peersWhat it tells you
Azure-Router, MIRT-BERT, NIRT-BERT60-67%~1x (cheapest cluster)Best accuracy-per-dollar in the field, still 5x the oracle's floor
MLP, KNN, NotDiamond, vLLM Semantic RouterComparable or worse~10x the cheap clusterExpensive without a matching accuracy gain
GPT-5-as-router (no routing)Highest overallHighest overallThe ceiling a router should beat, not match

RouterArena's live leaderboard, which keeps accepting community-submitted routers past the paper's original fixed evaluation, currently ranks 27 entries by Arena Score, topped by Cross-Router (75.75), Sqwish Router (75.27), and vLLM Semantic Router in third (74.86). The ordering shifts as new submissions land; the underlying finding about the oracle gap doesn't, because it's a property of how hard the routing problem is, not of which specific router currently sits on top.

Why every router makes the same mistake

The pattern across all 27 isn't randomness, it's a structural bias. A router built as a single classifier, trained once on labeled difficulty data, has no way to check itself against the request it's actually looking at. When it isn't sure, the only safe failure mode it has is to route up, to the model less likely to be wrong. That's individually rational and collectively expensive: it's the exact behavior that shows up as "over-relying on the expensive model," because uncertainty gets priced as a cost decision made in the router's favor, not the requester's.

The robustness metric makes the same point from a different angle. RouterArena reruns 420 queries through paraphrased variants and checks whether the routing decision stays the same. A router whose decision flips on a reworded but semantically identical question isn't measuring query difficulty, it's pattern-matching surface features of the prompt, which is a brittle signal to be pricing an entire model tier on.

Four questions worth asking before trusting a router's benchmark

RouterArena's own methodology doubles as a checklist for evaluating any router's cost claims, including a vendor's:

Where Nadir fits

Nadir-Tumbler now records an arena_score of 72.3 on RouterArena's public scorer, ranking #5 of 23 routers in the submitted field. That result measures routing decisions on RouterArena's model pool; it does not predict a customer's bill or served quality. Production proof begins separately: declare the real baseline, start in shadow mode, attach measured execution and outcomes, and enforce only the routes that clear the customer's bar. Complete non-streaming proxy responses can optionally add reference-free verifier telemetry, while streaming bypasses that step.

Related reading


Sources: Lu, Liu, Yuan, Cui, Zhang, Liu, and Xing, ["RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers"](https://arxiv.org/abs/2510.00202), arXiv:2510.00202, submitted September 30, 2025, accepted as a conference paper at ICLR 2026. RouteWorks, [RouterArena GitHub repository and live leaderboard](https://github.com/RouteWorks/RouterArena).

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

For coding agents, Nadir connects to Codex, Claude Code, or Cursor and recommends a model tier for delegated work. The agent decides whether to hand the task off and uses its own configured models, so Nadir needs no proxy and no provider keys on that path.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.