The Classification Ceiling

In September 2025, Red Hat and Verizon engineers filed SIRP, an IETF Internet-Draft to standardize how to classify and route a request before generation. The Semantic Inference Routing Protocol (draft-chen-nmrg-semantic-inference-routing-00) proposes headers so a client, a proxy, and a model server can agree on that. Read against the vLLM Semantic Router project's own Workload-Router-Pool Architecture vision paper (arXiv:2603.21354), it's the clearest public sketch yet of where router architecture is converging: classify, route, and stop. Neither spec checks whether the routed model's answer was actually correct. Here's what SIRP standardizes, why a header field can't close the accuracy gap on its own, and where a verifier-gated cascade picks up after classification ends.

Published 2026-08-22 by Dor Amir on the Nadir blog.

Filed under Routing & Cascades.

In September 2025, two engineers, one from Red Hat and one from Verizon, filed an Internet-Draft with the IETF: the Semantic Inference Routing Protocol (SIRP). The pitch is simple. Every LLM gateway today invents its own way to classify a request and decide where it goes. SIRP proposes a shared vocabulary instead: standardized header fields that let a client, a proxy, and a model server agree on what a prompt is asking for before a single token gets generated.

Whether or not SIRP itself becomes an RFC, the draft is worth reading, because it's the clearest public sketch yet of where router architecture is converging: classify the request, route it, and stop. That third step, checking whether the routed model actually got it right, is where SIRP and most production routers alike stop short.

What SIRP actually proposes

The draft, draft-chen-nmrg-semantic-inference-routing-00, authored by H. Chen (Red Hat) and L. Jalil (Verizon), defines three pieces: classification axes and a representation for them, interoperable signaling through standardized header fields, and a pluggable pipeline of "value-added routing" (VAR) modules for cost optimization, urgency prioritization, domain specialization, and privacy-aware handling. The goal is content-level classification, reading what a request is actually asking, rather than routing on client-supplied metadata that different teams label inconsistently.

It's early. An Internet-Draft expires after six months without a resubmission, and this is intended Standards Track, the kind of proposal that typically takes years and several revisions before it becomes an RFC, if it gets there at all. But the authorship is a signal worth noting: Red Hat is also behind the vLLM Semantic Router project, whose own Workload-Router-Pool Architecture vision paper argues for the same classify-then-route separation as infrastructure, not as a per-vendor implementation detail. Two publications from the same lineage, converging on the idea that routing deserves the kind of shared plumbing load balancers and CDNs already have.

Three ways to route a request today

ApproachHow it decidesWhat happens when the decision is wrongWhere it fits
Static rulesHand-written if/else on task type or client hintSilently ships whatever model the rule pointed toPrototypes, single-team internal tools
Predictive routing (SIRP-style classification, most commercial routers)Classify the prompt, route to the predicted best modelShips whatever the picked model produced, no second checkHigh-volume traffic where a wrong answer is cheap to redo
Verifier-gated cascadeCheap model answers first; an eligible complete response is scored before returnAttempts a stronger model after rejection and keeps the usable response on failureNon-streaming traffic where a wrong answer costs more than a retry

SIRP standardizes the plumbing for the second row. It says nothing about the third.

Why a header spec doesn't close the accuracy gap

Read closely, SIRP's VAR modules operate entirely before generation. A request comes in, gets classified along the draft's axes, cost sensitivity, urgency, domain, and gets signaled to whichever model server the classification points to. That's the whole loop. Nothing in the spec checks the model's actual output against the request afterward, because checking output was never the problem SIRP set out to solve. It's solving interoperability: making sure a proxy and a model server agree on what "route this cheap" means, the same way HTTP content negotiation makes sure a client and server agree on what "application/json" means.

That's a real, useful problem to standardize. It's also a different problem from "did the model that got picked actually answer correctly." A perfectly SIRP-compliant router can classify a hard multi-step reasoning question as routine, based on surface features like length or vocabulary, send it to a fast model, and ship a confidently wrong answer, because the classification happened before generation and nothing downstream double-checked the result. Classification predicts; it doesn't verify.

That gap is where a verifier-gated cascade sits. Nadir's router still classifies first. Complete non-streaming proxy responses can then enter a reference-free verifier and escalate on rejection; streaming bypasses that step. Separately, a reference-assisted RouterBench research experiment—where the verifier had the expensive reference answer unavailable in production—reported AUROC 0.961 and a ceiling of 60% projected cost reduction at 98% retained quality on 11,420 triples, with 1.7% catastrophic routes at tau 0.8. Those are research figures, not deployed performance.

What this means if you're evaluating a router this quarter

  1. Ask whether the router checks the output or only the input. A vendor that classifies a prompt and routes on that alone can't tell you when the classification was wrong, only that it followed its own rule.
  2. Ask what "routing accuracy" is measured against. Confidence scores from a classifier and correctness of the final answer are different numbers. A router can be highly confident and still wrong.
  3. If you're self-hosting a router, adopt SIRP-style classification headers now, ahead of RFC status. The classification axes are a genuinely useful shared vocabulary between your proxy and your model servers, standard or not, and adopting the pattern early costs nothing.
  4. Decide whether you need the cheapest model that can handle it, or the cheapest model that can handle it and be checked. For traffic where a wrong answer is cheap to retry, classification alone is enough. For traffic where a wrong answer ships to a customer, gets committed to a repo, or goes into a contract, pair classification with verification.

Nadir supports that second half as an optional managed-proxy policy for complete non-streaming responses. Streaming bypasses post-generation verification, and a failed verifier or escalation keeps the usable response. Start in shadow mode, attach outcomes, and inspect routing telemetry separately from measured quality before enforcement.


Sources: [Semantic Inference Routing Protocol, IETF Internet-Draft draft-chen-nmrg-semantic-inference-routing-00 (September 2025)](https://www.ietf.org/archive/id/draft-chen-nmrg-semantic-inference-routing-00.html). [IETF Datatracker entry for draft-chen-nmrg-semantic-inference-routing](https://datatracker.ietf.org/doc/draft-chen-nmrg-semantic-inference-routing/). ["The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project," arXiv:2603.21354](https://arxiv.org/pdf/2603.21354). [vLLM Semantic Router publications](https://vllm-sr.ai/publications). Figures current as of August 2026; Internet-Draft status subject to change or expiry per IETF process.

More on routing & cascades

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

For coding agents, Nadir Auto is a terminal launcher for Claude Code or Codex that asks Nadir for the main-session model before each turn, with your configured model as the fallback and cost ceiling. Delegation integrations for Claude Code, Codex, or Cursor instead recommend a model tier for subagent work, and the agent decides. Either way inference runs on your own provider account, so Nadir holds no provider keys on that path.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.