48% Off, Plus a Platform Team

Red Hat open sourced vLLM Semantic Router in September 2025, and its own published benchmark shows 48.5% fewer tokens and 47.1% lower latency. It passed 2,000 GitHub stars in two months, shipped three major versions in five months, and gets those gains by routing simple queries out of always-on reasoning mode. None of that is hype. It also ships as an Envoy filter that wants a Kubernetes cluster, roughly 25GB of classifier models, and its own Prometheus stack before it routes a single request. Here's what the project actually does, what standing it up costs before it saves you anything, and who should run it anyway.

Published 2026-07-05 by Dor Amir on the Nadir blog.

Filed under Nadir & Alternatives.

An open source router just proved routing works. It didn't make routing free.

In September 2025, Red Hat's CTO office published a new open source project: an Envoy filter that reads the semantic content of an LLM request and routes it to the right model before the request ever reaches an inference server. Source: Red Hat Developer, "vLLM Semantic Router: Improving efficiency in AI reasoning," September 11, 2025. Dr. Huamin Chen, the Senior Principal Software Engineer who built it, put the motivation plainly: "I knew this kind of reasoning-aware routing was largely confined to closed, proprietary systems. Red Hat's open source DNA demanded we bring this crucial capability to the open source community, making it accessible and transparent for everyone." Source: Red Hat, "Bringing intelligent, efficient routing to open source AI with vLLM Semantic Router".

It gained more traction than most infrastructure launches: over 2,000 GitHub stars and nearly 300 forks within two months of debut, per Red Hat's own count. Source: Red Hat, "Bringing intelligent, efficient routing to open source AI with vLLM Semantic Router". It then shipped three tagged releases in five months: v0.1 "Iris" on January 5, 2026, v0.2 "Athena" on March 10, 2026, and v0.3 in June 2026. Source: vLLM Blog, "vLLM Semantic Router v0.1 Iris: The First Major Release," January 5, 2026; Source: GitHub, vllm-project/semantic-router releases. AMD contributes GPU capacity and ROCm support for training the project's classifiers and running its public playground. Source: GitHub, vllm-project/semantic-router.

None of that is marketing. It's a real, actively maintained, independently governed project with a published research agenda: a vision paper on workload-router-pool architecture Source: "The Workload-Router-Pool Architecture for LLM Inference Optimization," arXiv:2603.21354, March 2026 and a paper on reasoning-mode routing specifically Source: "When to Reason: Semantic Router for vLLM," arXiv:2510.08731, both listed on the project's publications page. If you're weighing whether open source semantic routing is a real category now or still a research toy, the honest answer is: it's real. The part that gets skipped in the launch posts is what it costs to run.

What it actually does.

The router sits in front of your model pool as an Envoy External Processor: a gRPC filter, written in Go for the networking path and Rust (via the Candle framework) for the ML path, that intercepts every request before Envoy forwards it. Source: Red Hat, "Bringing intelligent, efficient routing to open source AI with vLLM Semantic Router". The filter turns the incoming prompt into an embedding, compares it against task vectors, and tells Envoy which backend model cluster should handle it. Source: Red Hat Emerging Technologies, "Intelligent inference request routing for large language models," November 11, 2025.

By the Athena release, that classification step had grown into eight neural classifiers built on mmBERT-32K, a 307-million-parameter model covering more than 1,800 languages: intent, jailbreak detection, PII detection, fact-checking, hallucination detection, domain classification, complexity scoring, and safety assessment, plus keyword-pattern and embedding-similarity signals layered on top. Source: Red Hat Developer, "Getting started with the vLLM Semantic Router project's Athena release," March 25, 2026. Athena also added a signal-decision architecture for Boolean routing logic, an HNSW-based semantic cache, and 11 model-selection algorithms ranging from static rules to Thompson sampling and RouterDC. Source: Red Hat Developer, "Getting started with the vLLM Semantic Router project's Athena release," March 25, 2026. The project's own homepage describes the goal as coordinating "local, private, and frontier models" from one routing layer, sending routine traffic to efficient lanes and reserving frontier reasoning for what actually needs it. Source: vLLM Semantic Router homepage.

The benchmark that's real.

The number that gets quoted most is from the original September 2025 write-up: on MMLU-Pro with Qwen3-30B, letting the router choose between reasoning mode and standard mode per query, instead of running the model in always-on reasoning mode, cut token usage 48.5%, cut latency 47.1%, and increased accuracy 10.2%, with some domains like business and economics gaining more than 20 points. Source: Red Hat Developer, "vLLM Semantic Router: Improving efficiency in AI reasoning," September 11, 2025.

vLLM Semantic Router's own published benchmark: token usage down 48.5%, latency down 47.1%, and MMLU-Pro accuracy up 10.2% in auto reasoning mode versus always-on reasoning, on Qwen3-30B
vLLM Semantic Router's own published benchmark: token usage down 48.5%, latency down 47.1%, and MMLU-Pro accuracy up 10.2% in auto reasoning mode versus always-on reasoning, on Qwen3-30B

That result is worth taking seriously and worth reading narrowly. It measures one routing decision: reasoning mode on or off, on one model family, on one benchmark. It is not a claim that every workload sees a 48% token reduction, and the project doesn't claim that either. What it demonstrates is the mechanism this whole blog argues for repeatedly: a meaningful share of requests get the wrong amount of compute by default, and a classifier cheap enough to run per-request can fix that before the model ever sees the prompt. The same logic applies one layer up, at the model tier rather than the reasoning-mode toggle.

What it costs to stand up.

This is the part the release notes don't lead with. To run vLLM Semantic Router in production, the documented dependency list is: vLLM itself, PyTorch, Hugging Face model infrastructure, Kubernetes, Envoy, and Prometheus/Grafana for monitoring. Source: vLLM Semantic Router homepage. Production deployment ships through Helm charts and two custom Kubernetes resources, IntelligentPool and IntelligentRoute, with horizontal pod autoscaling wired in. Source: Red Hat Developer, "Getting started with the vLLM Semantic Router project's Athena release," March 25, 2026. The initial deployment footprint is documented at roughly 25GB for the internal ML models plus 2GB for the container image, before it has routed a single production request. Source: Red Hat Developer, "Getting started with the vLLM Semantic Router project's Athena release," March 25, 2026.

None of that is a criticism of the project. It's an accurate description of what "open source and free" means for a system-level routing layer: the software has no license fee, and someone on your team still owns a Kubernetes-native gateway, an eight-classifier ML pipeline, a semantic cache, and an observability stack, across major version jumps that landed roughly every two months this year. Quick local testing is available through Docker or Podman without the full cluster. Source: Red Hat Developer, "Getting started with the vLLM Semantic Router project's Athena release," March 25, 2026. Running it at the reliability a production API gateway needs is a different commitment than docker run.

# What "trying it" looks like today
curl -fsSL https://vllm-semantic-router.com/install.sh | bash

# What "running it in production" adds on top:
# - an Envoy deployment with the ExtProc filter wired in
# - a Kubernetes cluster (Helm chart, IntelligentPool/IntelligentRoute CRDs, HPA)
# - ~25GB of classifier models served and kept warm
# - Prometheus + Grafana for the observability the router itself doesn't give you

Source: vLLM Semantic Router homepage, installation instructions.

Self-hosted OSS router versus a drop-in proxy.

vLLM Semantic Router (self-hosted)Nadir (hosted proxy)
License costFree, open sourceFree with BYOK; usage-based on the hosted plan
Infra you ownEnvoy, Kubernetes, classifier hosting, Prometheus/GrafanaNone; point your base URL at the proxy
ClassifierYou train, host, and upgrade across major versionsMaintained centrally, versioned for you
Time to first routed requestCluster setup, Helm install, CRDs, then trafficA base URL change
Where it fitsTeams already running vLLM/Kubernetes at scale, or with a hard requirement to keep classification fully in-houseTeams whose LLM calls already hit hosted APIs (Claude, GPT, Gemini) and want routing without hiring for it
Verification before shipping a routeNot the project's stated focusVerifier-gated: a calibrated model checks the cheap answer before it ships

Both rows are legitimate answers to "how do I stop sending every request to the most expensive model." They're just answers for different teams. If you're already operating a vLLM fleet behind Kubernetes and Envoy, the marginal cost of adding this router is low, and full control over the classifier is a feature, not overhead, especially if PII or jailbreak detection has to stay inside your own network boundary for compliance reasons. If your traffic already goes to hosted frontier APIs and nobody on the team wants to own a new Kubernetes-native gateway, Nadir is the same routing idea with the operational side absorbed: swap the base URL, set model=auto, and the classification, fallback chains, and cost dashboard are already running.

A checklist before you adopt either one.

Conclusion.

vLLM Semantic Router is a genuinely good outcome for the industry: reasoning-aware routing, jailbreak and PII detection, and multi-model orchestration are no longer locked inside proprietary gateways, and a well-resourced open source project is iterating on all three in public. The 48.5% token reduction and 47.1% latency reduction it reports are real, published numbers, not vendor marketing math. What's also real is that "open source" describes the license, not the operational cost. Standing this up means a Kubernetes cluster, an Envoy deployment, roughly 25GB of classifier models kept warm, and a monitoring stack, before it routes a single production request, and someone has to keep all of it current across a release cadence measured in weeks. For teams already living in that infrastructure, that's a fair trade for full control. For teams whose LLM traffic already goes to hosted APIs, the same routing idea is available without becoming a platform team's new full-time project.


Data in the chart is drawn directly from the cited Red Hat Developer benchmark, not derived from proprietary production traces. Sources: [Red Hat Developer, "vLLM Semantic Router: Improving efficiency in AI reasoning," September 11, 2025](https://developers.redhat.com/articles/2025/09/11/vllm-semantic-router-improving-efficiency-ai-reasoning). [Red Hat, "Bringing intelligent, efficient routing to open source AI with vLLM Semantic Router"](https://www.redhat.com/en/blog/bringing-intelligent-efficient-routing-open-source-ai-vllm-semantic-router). [Red Hat Emerging Technologies, "Intelligent inference request routing for large language models," November 11, 2025](https://next.redhat.com/2025/11/11/intelligent-inference-request-routing-for-large-language-models/). [Red Hat Developer, "Getting started with the vLLM Semantic Router project's Athena release," March 25, 2026](https://developers.redhat.com/articles/2026/03/25/getting-started-vllm-semantic-router-athena-release). [vLLM Blog, "vLLM Semantic Router v0.1 Iris: The First Major Release," January 5, 2026](https://blog.vllm.ai/2026/01/05/vllm-sr-iris.html). [GitHub, vllm-project/semantic-router](https://github.com/vllm-project/semantic-router). [vLLM Semantic Router homepage](https://vllm-semantic-router.com/). ["The Workload-Router-Pool Architecture for LLM Inference Optimization," arXiv:2603.21354](https://arxiv.org/pdf/2603.21354). ["When to Reason: Semantic Router for vLLM," arXiv:2510.08731](https://arxiv.org/pdf/2510.08731).

More on nadir & alternatives

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

For coding agents, Nadir connects to Codex, Claude Code, or Cursor and recommends a model tier for delegated work. The agent decides whether to hand the task off and uses its own configured models, so Nadir needs no proxy and no provider keys on that path.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.