The Fifth Lane

Read together, the four lanes of PointFive's August 2026 LLM cost optimization guide leave a gap: prompt-side cost decisions are not outcome evidence. Nadir's older 60% / 98% RouterBench result explored that gap, but it was reference-assisted and is a research ceiling, not deployed performance. Here is the market map and the evidence boundary the original guide missed.

Published 2026-08-19 by Dor Amir on the Nadir blog.

Filed under Nadir & Alternatives.

A category gets a buyer's guide once it's too crowded to evaluate by hand.

That's usually the signal a market has matured: analysts stop writing "what is X" explainers and start writing "which X should you buy" comparisons. In August 2026, PointFive published one of the first of those for this space: "Top 10 Token Optimization Solutions," a buyer's guide running down ten vendors that all promise the same thing, spend less per LLM call without losing quality. Source: PointFive, "Top 10 Token Optimization Solutions (2026): An Honest, Comparison-Driven Guide"

Ten vendors is a real number for a category that barely had a name two years ago. It's also small enough to read in one sitting, which is worth doing, because the guide's own structure reveals something its author probably wasn't trying to say: every one of the ten tools it lists optimizes something about the request. None of them check the response.

The four lanes, as the guide draws them.

PointFive groups its ten vendors into four functional lanes:

AI gateways — Portkey, LiteLLM, Kong AI Gateway, Cloudflare AI Gateway. These proxy the request: virtual keys, per-tenant budgets, fallback chains, and a semantic-cache layer bolted on for the repeat-query case. Kong's own claim is "up to 5x cost reduction while preserving ~80% semantic meaning" through a prompt-compression plugin; that's a vendor number, not an independent benchmark, and the guide reports it as such.

Intelligent routers — OpenRouter, Not Diamond, RouteLLM, Martian. These pick which model handles a request before generation starts. Not Diamond's public claim is a router that "picks the cheapest acceptable model per prompt" with "up to +39% accuracy" versus naive routing; RouteLLM, the open-source project out of LMSYS, reports over 85% cost reduction on MT-Bench while matching GPT-4 level judgments. Martian, Accenture-backed, cites a 20-97% cost-cut range depending on workload.

Observability and semantic cache — Langfuse, GPTCache, Helicone. These surface where tokens are going and exploit exact or near-duplicate queries. Helicone reports caching "up to 95% on cache hits" on repeated requests; GPTCache, the open-source library behind a lot of that pattern, claims roughly 100x latency reduction on a cache hit.

Endpoint and agent-side optimizers — TokenShift, LLMLingua (Microsoft Research). These compress the prompt itself before it ever leaves the developer's machine: deduplicating carried context, trimming CLI output, rightsizing images. TokenShift reports "12-21% average token reduction" across developer coding workloads; LLMLingua's published range runs up to 20x compression on long prompts with limited quality loss on its own benchmarks.

Diagram of four labeled market lanes for LLM token optimization, each with example vendors and what it optimizes: gateways route the request to a provider, intelligent routers pick a model before generation, observability and cache tools exploit repeat queries, and endpoint optimizers compress the prompt before it leaves the developer's machine. A fifth, highlighted lane, verified cascade, is shown below the four as unnamed in the guide: it grades the cheap model's actual answer before shipping it, rather than deciding anything before the answer exists.
Diagram of four labeled market lanes for LLM token optimization, each with example vendors and what it optimizes: gateways route the request to a provider, intelligent routers pick a model before generation, observability and cache tools exploit repeat queries, and endpoint optimizers compress the prompt before it leaves the developer's machine. A fifth, highlighted lane, verified cascade, is shown below the four as unnamed in the guide: it grades the cheap model's actual answer before shipping it, rather than deciding anything before the answer exists.

Four lanes, ten vendors, a real spread of claimed numbers from 12% to 97%. Every one of those numbers describes a decision made looking at the prompt: which model, which cache entry, how many tokens to strip. Not one of them describes a decision made looking at the answer.

What "optimizing the request" quietly assumes.

A router that picks the cheapest acceptable model per prompt is making a prediction. It has to be: the whole point is to decide before the expensive step (generation) happens. Not Diamond's own framing says as much, "picks the cheapest acceptable model," where acceptable is a property the router estimates in advance, not a property it checks afterward.

That's not a flaw specific to any one vendor in the guide, it's the shape of the entire routing lane. This blog has made the general version of this argument before: a predictive router reads the prompt, estimates difficulty, and ships whatever the selected model produces, including the fraction of the time its estimate was wrong. The gateways lane doesn't touch this question at all, it routes traffic, not judgment. The cache lane sidesteps it by only firing on requests it's already seen before. The endpoint-optimizer lane compresses what goes in, and has no visibility into what comes out. Read the four lanes together and the pattern is structural, not a gap in any single product: the entire mapped category makes its one cost decision pre-generation, then stops looking.

The number that lane 2 doesn't publish.

Every routing vendor in the guide publishes an average: 39% more accurate, 85% cheaper on MT-Bench, 20-97% depending on workload. None of the four lanes' vendor claims quote a tail number, how often the cheap model's actual answer on a specific request was wrong and nothing caught it before the user saw it. That's not because the vendors are hiding something; a predictive router that never inspects the generated answer has no mechanism to produce that number in the first place. It would have to look at the output to measure it.

A reference-assisted RouterBench research experiment on 11,420 triples projected 60% lower cost at about 98% retained quality. The verifier had the expensive reference answer unavailable in production, so this is a research ceiling, not deployed performance. It demonstrates the value of output evidence without establishing a universal live guarantee.

A fifth lane, unnamed because nobody in the guide builds it.

Call it verified cascade: the cheap model answers first, a calibrated verifier scores that specific answer, and only a low score escalates to a stronger model. It's a routing decision, so it overlaps with lane 2, but it happens on the other side of generation from every router in the guide. The distinction isn't cosmetic. A predictive router that guesses wrong ships the wrong answer at full speed. A verified cascade that guesses wrong catches it before the user sees it and pays a second call to fix it, the same failure mode, with a floor under it.

It also composes with the other three lanes rather than replacing them. A verified cascade still benefits from a gateway's spend controls and fallback chains, still benefits from a semantic cache catching the true duplicates, still benefits from a leaner prompt arriving from an endpoint optimizer. None of that changes what happens once generation starts. The fifth lane is additive, not a rival category to the other four, which is probably part of why a buyer's guide organized around "what does this tool touch" didn't surface it as its own row.

Where this goes next.

Buyer's guides for a maturing category tend to get a new column roughly on the schedule the category earns one. Cache hit rate wasn't a line item in cost-optimization comparisons three years ago, then caching got cheap and provider-native, and now every gateway in the guide lists it as a feature. The same thing is plausible here: as more teams route production traffic through a predictive-only router and eventually hit the tail case, an unrecoverable wrong answer at full model speed, "does it verify the answer, or just guess at the prompt" starts to look like the next column analysts add, not a niche add-on.

That's the bet Nadir is built on. Start beside an existing gateway in shadow mode and each request gets a model, effort, cache, context, and policy receipt without a model-provider call. Attach actual usage and outcomes, then enforce only what clears the workload's bar. Complete non-streaming proxy responses can optionally enter a verifier cascade; streaming bypasses it. Start free or read why routing without verification is dead reckoning for the longer argument underneath this post.


Sources: [PointFive, "Top 10 Token Optimization Solutions (2026): An Honest, Comparison-Driven Guide"](https://www.pointfive.co/guides/top-10-token-optimization-solutions-2026), which compiles vendor-reported claims. Nadir's 60% / ~98% figures are a reference-assisted research ceiling on 11,420 RouterBench triples, detailed with the production boundary in ["Reference-assisted cascade on RouterBench"](/blog/routerbench-cascade-benchmark).

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

For coding agents, Nadir connects to Codex, Claude Code, or Cursor and recommends a model tier for delegated work. The agent decides whether to hand the task off and uses its own configured models, so Nadir needs no proxy and no provider keys on that path.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.