A category gets a buyer's guide once it's too crowded to evaluate by hand.
That's usually the signal a market has matured: analysts stop writing "what is X" explainers and start writing "which X should you buy" comparisons. In August 2026, PointFive published one of the first of those for this space: "Top 10 Token Optimization Solutions," a buyer's guide running down ten vendors that all promise the same thing, spend less per LLM call without losing quality. Source: PointFive, "Top 10 Token Optimization Solutions (2026): An Honest, Comparison-Driven Guide"
Ten vendors is a real number for a category that barely had a name two years ago. It's also small enough to read in one sitting, which is worth doing, because the guide's own structure reveals something its author probably wasn't trying to say: every one of the ten tools it lists optimizes something about the request. None of them check the response.
The four lanes, as the guide draws them.
PointFive groups its ten vendors into four functional lanes:
AI gateways — Portkey, LiteLLM, Kong AI Gateway, Cloudflare AI Gateway. These proxy the request: virtual keys, per-tenant budgets, fallback chains, and a semantic-cache layer bolted on for the repeat-query case. Kong's own claim is "up to 5x cost reduction while preserving ~80% semantic meaning" through a prompt-compression plugin; that's a vendor number, not an independent benchmark, and the guide reports it as such.
Intelligent routers — OpenRouter, Not Diamond, RouteLLM, Martian. These pick which model handles a request before generation starts. Not Diamond's public claim is a router that "picks the cheapest acceptable model per prompt" with "up to +39% accuracy" versus naive routing; RouteLLM, the open-source project out of LMSYS, reports over 85% cost reduction on MT-Bench while matching GPT-4 level judgments. Martian, Accenture-backed, cites a 20-97% cost-cut range depending on workload.
Observability and semantic cache — Langfuse, GPTCache, Helicone. These surface where tokens are going and exploit exact or near-duplicate queries. Helicone reports caching "up to 95% on cache hits" on repeated requests; GPTCache, the open-source library behind a lot of that pattern, claims roughly 100x latency reduction on a cache hit.
Endpoint and agent-side optimizers — TokenShift, LLMLingua (Microsoft Research). These compress the prompt itself before it ever leaves the developer's machine: deduplicating carried context, trimming CLI output, rightsizing images. TokenShift reports "12-21% average token reduction" across developer coding workloads; LLMLingua's published range runs up to 20x compression on long prompts with limited quality loss on its own benchmarks.
Four lanes, ten vendors, a real spread of claimed numbers from 12% to 97%. Every one of those numbers describes a decision made looking at the prompt: which model, which cache entry, how many tokens to strip. Not one of them describes a decision made looking at the answer.
What "optimizing the request" quietly assumes.
A router that picks the cheapest acceptable model per prompt is making a prediction. It has to be: the whole point is to decide before the expensive step (generation) happens. Not Diamond's own framing says as much, "picks the cheapest acceptable model," where acceptable is a property the router estimates in advance, not a property it checks afterward.
That's not a flaw specific to any one vendor in the guide, it's the shape of the entire routing lane. This blog has made the general version of this argument before: a predictive router reads the prompt, estimates difficulty, and ships whatever the selected model produces, including the fraction of the time its estimate was wrong. The gateways lane doesn't touch this question at all, it routes traffic, not judgment. The cache lane sidesteps it by only firing on requests it's already seen before. The endpoint-optimizer lane compresses what goes in, and has no visibility into what comes out. Read the four lanes together and the pattern is structural, not a gap in any single product: the entire mapped category makes its one cost decision pre-generation, then stops looking.
The number that lane 2 doesn't publish.
Every routing vendor in the guide publishes an average: 39% more accurate, 85% cheaper on MT-Bench, 20-97% depending on workload. None of the four lanes' vendor claims quote a tail number, how often the cheap model's actual answer on a specific request was wrong and nothing caught it before the user saw it. That's not because the vendors are hiding something; a predictive router that never inspects the generated answer has no mechanism to produce that number in the first place. It would have to look at the output to measure it.
A reference-assisted RouterBench research experiment on 11,420 triples projected 60% lower cost at about 98% retained quality. The verifier had the expensive reference answer unavailable in production, so this is a research ceiling, not deployed performance. It demonstrates the value of output evidence without establishing a universal live guarantee.
A fifth lane, unnamed because nobody in the guide builds it.
Call it verified cascade: the cheap model answers first, a calibrated verifier scores that specific answer, and only a low score escalates to a stronger model. It's a routing decision, so it overlaps with lane 2, but it happens on the other side of generation from every router in the guide. The distinction isn't cosmetic. A predictive router that guesses wrong ships the wrong answer at full speed. A verified cascade that guesses wrong catches it before the user sees it and pays a second call to fix it, the same failure mode, with a floor under it.
It also composes with the other three lanes rather than replacing them. A verified cascade still benefits from a gateway's spend controls and fallback chains, still benefits from a semantic cache catching the true duplicates, still benefits from a leaner prompt arriving from an endpoint optimizer. None of that changes what happens once generation starts. The fifth lane is additive, not a rival category to the other four, which is probably part of why a buyer's guide organized around "what does this tool touch" didn't surface it as its own row.
Where this goes next.
Buyer's guides for a maturing category tend to get a new column roughly on the schedule the category earns one. Cache hit rate wasn't a line item in cost-optimization comparisons three years ago, then caching got cheap and provider-native, and now every gateway in the guide lists it as a feature. The same thing is plausible here: as more teams route production traffic through a predictive-only router and eventually hit the tail case, an unrecoverable wrong answer at full model speed, "does it verify the answer, or just guess at the prompt" starts to look like the next column analysts add, not a niche add-on.
That's the bet Nadir is built on. Start beside an existing gateway in shadow mode and each request gets a model, effort, cache, context, and policy receipt without a model-provider call. Attach actual usage and outcomes, then enforce only what clears the workload's bar. Complete non-streaming proxy responses can optionally enter a verifier cascade; streaming bypasses it. Start free or read why routing without verification is dead reckoning for the longer argument underneath this post.
Sources: [PointFive, "Top 10 Token Optimization Solutions (2026): An Honest, Comparison-Driven Guide"](https://www.pointfive.co/guides/top-10-token-optimization-solutions-2026), which compiles vendor-reported claims. Nadir's 60% / ~98% figures are a reference-assisted research ceiling on 11,420 RouterBench triples, detailed with the production boundary in ["Reference-assisted cascade on RouterBench"](/blog/routerbench-cascade-benchmark).