Every LLM router on the market makes the same pitch: send us your traffic, and we'll pick the cheapest model that can handle each request. Almost none of them get evaluated by anyone but themselves. A router vendor's benchmark is run on the vendor's own dataset, against the vendor's own baseline, reported in the vendor's own units, and there has been no standardized, independent way to check whether the number on the landing page means what it claims to mean.
RouterArena, published as a conference paper at ICLR 2026 by a team including Yifan Lu and Jiarong Xing (arXiv:2510.00202), set out to close that gap. It built a fixed, standardized 8,400-query evaluation set, scored 27 routers on it, and asked a question no vendor benchmark is built to answer: when a router says it picked the right model, compared to what, exactly? The result isn't a single winner. It's a field of 27 systems, none of which comes close to the theoretical best a router could do, and a headline finding the paper states in one flat sentence: "no router ranks at the top across all metrics."
What RouterArena actually measures
The dataset is built to resist the two ways a router benchmark usually gets gamed: an easy test set, and a single metric. It draws from 23 existing open datasets, spans nine domains modeled on the Dewey Decimal Classification (everything except religion), and covers 44 categories. Every query gets a difficulty label, easy, medium, or hard, mapped onto Bloom's taxonomy (Remember/Understand, Apply, Analyze/Evaluate), assigned by DeepSeek-V3 and then rebalanced so no single category dominates the harder tiers. A separate 420-query subset tests the same routers under paraphrased, perturbed versions of the same questions, to see whether a routing decision holds up when the input isn't word-for-word identical to what the router was tuned on.
| RouterArena fact | Detail |
|---|---|
| Dataset size | 8,400 queries, 809-query and 420-query subsets for lighter runs |
| Domains / categories | 9 domains (Dewey Decimal, minus religion), 44 categories |
| Difficulty tiers | Easy, medium, hard, via Bloom's taxonomy, DeepSeek-V3-labeled |
| Source datasets | 23 existing open datasets, deduplicated by cosine similarity |
| Routers evaluated | 27, spanning commercial (GPT-5-as-router, NotDiamond, Azure Router) and academic (RouterBench KNN/MLP, GraphRouter, CARROT, RouterDC, IRT-Router, RouteLLM, vLLM Semantic Router, MIRT-BERT) |
| Metrics | Accuracy, cost, routing optimality vs. an oracle, robustness to perturbation, latency overhead |
On top of the five metrics sits a single composite ranking number, the Arena Score, a weighted harmonic mean of accuracy and cost:
S = (1 + β) · A · C / (β · A + C)
A = router accuracy (0-1)
C = normalized cost, log2-scaled between $0.0044 and $200 per 1,000 queries
β = 0.1 (default), weights accuracy above cost
That formula matters more than it looks. A harmonic mean punishes a router for being lopsided, brilliant on accuracy but wildly expensive, or cheap but wrong too often, in a way a simple average wouldn't. It's the same instinct behind why a router that always calls the frontier model scores badly here even with near-perfect accuracy: cost is baked into the ranking, not reported as a footnote next to it.
The results: everyone hedges upward
RouterArena's comparison point isn't a competitor's number, it's an oracle: a router with perfect foresight that always picks the cheapest model that would have answered a given query correctly. No real router can match an oracle, since that requires knowing the answer before choosing a model. What RouterArena measures is how far short of it every real router falls, and the paper's finding is that all 27 fall short the same way: by over-relying on expensive models even on queries a cheap model would have gotten right.
Two clusters emerge. Azure-Router, MIRT-BERT, and NIRT-BERT land in the 60-67% raw accuracy range at roughly one order of magnitude lower cost than a second cluster, MLP, KNN, NotDiamond, and vLLM Semantic Router, which cost substantially more for comparable or worse accuracy. MIRT-BERT comes closest to the accuracy-cost frontier of anything tested: about 77% of the oracle's accuracy, at roughly five times the oracle's cost. That's the best result in the field. On the long-context subset specifically, GPT-5-as-router (always route to GPT-5, no selection at all) posts the highest raw accuracy, unsurprisingly, at the highest cost, while Azure-Router and NIRT-BERT hold competitive accuracy at a fraction of that spend.
| Cluster | Accuracy | Cost, relative to peers | What it tells you |
|---|---|---|---|
| Azure-Router, MIRT-BERT, NIRT-BERT | 60-67% | ~1x (cheapest cluster) | Best accuracy-per-dollar in the field, still 5x the oracle's floor |
| MLP, KNN, NotDiamond, vLLM Semantic Router | Comparable or worse | ~10x the cheap cluster | Expensive without a matching accuracy gain |
| GPT-5-as-router (no routing) | Highest overall | Highest overall | The ceiling a router should beat, not match |
RouterArena's live leaderboard, which keeps accepting community-submitted routers past the paper's original fixed evaluation, currently ranks 27 entries by Arena Score, topped by Cross-Router (75.75), Sqwish Router (75.27), and vLLM Semantic Router in third (74.86). The ordering shifts as new submissions land; the underlying finding about the oracle gap doesn't, because it's a property of how hard the routing problem is, not of which specific router currently sits on top.
Why every router makes the same mistake
The pattern across all 27 isn't randomness, it's a structural bias. A router built as a single classifier, trained once on labeled difficulty data, has no way to check itself against the request it's actually looking at. When it isn't sure, the only safe failure mode it has is to route up, to the model less likely to be wrong. That's individually rational and collectively expensive: it's the exact behavior that shows up as "over-relying on the expensive model," because uncertainty gets priced as a cost decision made in the router's favor, not the requester's.
The robustness metric makes the same point from a different angle. RouterArena reruns 420 queries through paraphrased variants and checks whether the routing decision stays the same. A router whose decision flips on a reworded but semantically identical question isn't measuring query difficulty, it's pattern-matching surface features of the prompt, which is a brittle signal to be pricing an entire model tier on.
Four questions worth asking before trusting a router's benchmark
RouterArena's own methodology doubles as a checklist for evaluating any router's cost claims, including a vendor's:
- Is cost reported against a true oracle, or against a single fixed baseline model? A router that looks 90% cheaper than always-calling-the-frontier-model can still be 5x the oracle's floor. Both numbers are true. Only one tells you how much is left on the table.
- Is accuracy tested under paraphrase, not just on the exact benchmark wording? A router that only performs well on inputs shaped like its training distribution will look great on a demo and degrade the moment real traffic diverges from it.
- Does the reported accuracy hold across difficulty tiers, or only on the easy slice? A benchmark average can hide a router that's excellent on easy queries and close to random on hard ones, since easy queries are more common and dominate the average.
- Does the routing decision get re-evaluated per request, or committed once by a static classifier that never revisits a low-confidence call? This is the structural fix for the upward-hedging bias above: a router that can check its own answer before committing to it doesn't need to default to the expensive model every time it's unsure.
Where Nadir fits
Nadir-Tumbler now records an arena_score of 72.3 on RouterArena's public scorer, ranking #5 of 23 routers in the submitted field. That result measures routing decisions on RouterArena's model pool; it does not predict a customer's bill or served quality. Production proof begins separately: declare the real baseline, start in shadow mode, attach measured execution and outcomes, and enforce only the routes that clear the customer's bar. Complete non-streaming proxy responses can optionally add reference-free verifier telemetry, while streaming bypasses that step.
Related reading
- Reference-assisted RouterBench research ceiling: 60% lower projected cost, 98% retained quality
- ACL 2026 tested every major LLM routing method on 400,000 queries. Most commercial routers failed to beat a simple baseline.
- The LLM router that refuses to guess, and cuts bad routes 23%
- Routing without verification is dead-reckoning
Sources: Lu, Liu, Yuan, Cui, Zhang, Liu, and Xing, ["RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers"](https://arxiv.org/abs/2510.00202), arXiv:2510.00202, submitted September 30, 2025, accepted as a conference paper at ICLR 2026. RouteWorks, [RouterArena GitHub repository and live leaderboard](https://github.com/RouteWorks/RouterArena).