The RULER Gap

BenchLM counts 322 models with a published context window, a quarter now claiming 1 million tokens or more, but almost none publish how much of it works. Its context-window tracker, dated today, puts a number on the arms race: one model (Pokee AI's Pokee-Isaac 28B) claims 10 million. NVIDIA's RULER benchmark (arXiv:2404.06654) found in 2024 that effective context runs well under the advertised maximum, a gap the industry still cites at roughly 60-70%, and the 2026 field has mostly stopped publishing quality-at-length scores to check it against. Since Anthropic dropped its long-context pricing surcharge on March 13, 2026, the one signal that used to hint a prompt had wandered past the useful zone is gone too: a 900,000-token prompt now costs exactly what a 9,000-token one does, whether or not the model actually used what you paid for. Here's what RULER measures, how thin the 2026 quality-at-length data actually is, and a framework for building to the real ceiling instead of the printed one.

Published 2026-08-20 by Dor Amir on the Nadir blog.

Filed under Context & Compression.

As of today, one model out of 322 backs up a 10-million-token claim with nothing anyone can check.

BenchLM's context-window tracker, dated August 20, 2026, puts a number on the context-window arms race that's been running for the last year: 397 models tracked overall, 322 of them with a published context window, and a quarter of those now claim 1 million tokens or more. The single largest number on the board belongs to Pokee AI's Pokee-Isaac 28B, at 10 million tokens. Llama 4 Scout and Gemini 3 Pro advertise the same ceiling.

The median model in that tracker, though, sits at 256K, a number that's easy to miss next to the 10M headline, and worth sitting with for a second: three out of four models in the entire tracked field don't even claim 1M, let alone 10M. The arms race is real, but it's a race among a minority of the field, not where most deployed context actually lives.

Here's the part that doesn't make it into the spec sheet: how much of an advertised window a model can actually use. That number has its own name, and its own decade-old benchmark, and in 2026 it's still nowhere close to catching up with the marketing.

What RULER actually measures.

In April 2024, a team at NVIDIA (Hsieh, Sun, Kriman, Acharya, Rekesh, Jia, Zhang, and Ginsburg) published RULER: What's the Real Context Size of Your Long-Context Language Models?, arXiv:2404.06654, a synthetic benchmark built specifically to answer the question its title asks. RULER goes past the standard needle-in-a-haystack retrieval test with 13 tasks across four categories, retrieval, multi-hop variable tracing, aggregation, and question answering, because the paper's own finding was that "despite achieving nearly perfect accuracy in the vanilla NIAH test, almost all models exhibit large performance drops as the context length increases." A model can ace the easy version of the long-context test and still fail the harder ones well inside its claimed window.

Testing 17 models across those 13 tasks, RULER found that only half of the models claiming 32K-token context or larger could actually clear a fixed accuracy bar at 32K tokens. Its own words: "almost all models fall below the threshold before reaching the claimed context lengths." That was the 2024 generation of models. The tracked field has grown roughly 20x in headline context size since, and the same structural gap, the industry now widely cites RULER's methodology as finding effective usable context runs somewhere around 60-70% of the advertised maximum, hasn't closed. It's arguably gotten harder to see, because at 10M tokens, nobody's independently re-running a 13-task synthetic probe out to the full claimed length before the next model ships.

Left panel: advertised context tiers across the 322 models BenchLM tracks with a published window as of August 20, 2026, roughly half under 256K, a quarter between 256K and 1M, a quarter at 1M or more, and one single model claiming the 10M ceiling. Right panel: of the roughly 80 tracked models that advertise 1M or more tokens, only a handful publish any quality-at-length benchmark score anywhere near that ceiling, against RULER's widely cited 60-70% effective-context range.
Left panel: advertised context tiers across the 322 models BenchLM tracks with a published window as of August 20, 2026, roughly half under 256K, a quarter between 256K and 1M, a quarter at 1M or more, and one single model claiming the 10M ceiling. Right panel: of the roughly 80 tracked models that advertise 1M or more tokens, only a handful publish any quality-at-length benchmark score anywhere near that ceiling, against RULER's widely cited 60-70% effective-context range.

The number almost nobody publishes.

If RULER-style testing at the advertised ceiling were routine, this would be a solved problem: buyers would read the effective-length column the way they read the price column. It isn't routine. Pulling together the long-context comparison pages tracking 2026's frontier models, the honest state of the data is closer to this: "the only published row in this set is GPT-5.5 at 87.5 on the MRCRv2 long-context test" in the 128K-256K band, 83.1 in the 64K-128K band, and nothing published near its 1M ceiling. On Artificial Analysis's Long-Context Reasoning benchmark, Claude Fable 5 is reported scoring 77%, one of the only other public data points in the same neighborhood. Llama 4 Scout advertises 10M tokens; as of this writing, no published benchmark shows quality holding anywhere near that length.

That's the actual gap worth naming, and it's not just "effective context is smaller than advertised," which RULER already established in 2024. It's that the industry has kept publishing the size spec every model generation while mostly stopping publishing the quality-at-that-size spec. A context window is now marketed the way storage used to be marketed: the number on the box, not what happens to read speed once you fill it.

Why the bill used to catch this, and doesn't anymore.

There used to be a cost signal that at least hinted when a prompt had wandered past the point of diminishing returns. Anthropic ran a long-context surcharge on Claude prompts over 200K tokens during its 1M-context beta, a 2x input / 1.5x output multiplier. On March 13, 2026, Anthropic dropped it for Claude Opus 4.6 and Sonnet 4.6: standard pricing now runs flat across the entire window, $5/$25 per million tokens for Opus 4.6, $3/$15 for Sonnet 4.6, no premium tier past 200K. One industry write-up put the practical effect plainly: "a 900,000-token request now costs the same per-token rate as a 9,000-token one." Every other frontier vendor already prices this way.

That's a genuinely good change for teams that need the full window and were previously taxed for using it. It also means the one number on an invoice that used to nudge a team toward "maybe check whether this prompt needs to be this long" is gone. Nothing in the bill distinguishes a 900K-token prompt where every token gets used correctly from a 900K-token prompt where the fact the model needed was buried at token 550,000, past whatever fraction of that window is actually reliable, and got missed. Both cost the same. Only one of them gave you an answer worth what you paid for it.

ModelAdvertised 1M input price (list)Long-context surcharge
DeepSeek V4 Pro$0.435 / MNone
Gemini 3.1 Pro$4.00 / MNone
GPT-5.5$5.00 / MNone
Claude Opus (tokenizer-adjusted)$6.75 / MRemoved March 13, 2026

Figures as compiled in 2026 pricing trackers; list prices, not accounting for caching. See sources.

A 71x-cheaper-per-token model and a full-price frontier model both charge you in full for tokens the underlying architecture may not be reliably attending to. Price-per-token was never a proxy for whether the tokens in that prompt were going to be used correctly, and now there's no premium tier left to accidentally flag it either.

This is a quality bug wearing a cost bug's clothes.

The mechanism under the RULER gap is well documented and has a name in its own right: "lost in the middle." Facts placed near the start or end of a long context get retrieved reliably. The same fact placed somewhere in the middle gets missed at a measurably higher rate, and the model doesn't hedge or flag uncertainty when it happens, it just answers confidently with whatever it did retrieve. That's the trap: a context-stuffed prompt that silently drops the one fact that mattered doesn't look different from a correct answer. It looks identical, right up until it's wrong in a way nobody caught before it shipped.

Treating this purely as a cost problem, "we're paying for tokens we don't need," undersells it. The same tokens that don't help the bill also don't help the answer, and past the effective boundary they can actively hurt it, by pushing the fact that mattered further into the unreliable middle of the window instead of leaving it near an edge.

A framework for building to the real ceiling, not the printed one.

1. Treat the advertised window as a purchase spec, not an engineering budget. RULER's own finding, still the best-documented shape of this problem two years and one context-size generation later, is that effective length runs meaningfully under claimed length for most models. Design context budgets around roughly 60-70% of the number on the pricing page, not the number itself, until a specific model's own quality-at-length data says otherwise.

2. Don't assume the ratio transfers across vendors. RULER measured a wide spread across the 17 models it tested in 2024, and the sparse 2026 data available (GPT-5.5's published MRCRv2 rows, Claude Fable 5's AA-LCR score) shows the same thing: this is not one constant you can apply to every provider. A model's effective ratio is closer to a property of that specific model than an industry average.

3. Position what has to be found. Given the lost-in-the-middle effect, a fact your task genuinely cannot afford to miss belongs near the start or end of the prompt, not buried in the middle of retrieved chunks, agent history, or tool output, regardless of how much headroom the advertised window claims to have left.

4. Compress and route against the effective ceiling, not the advertised one. Pruning retrieved context down to what a task actually needs, and picking the model whose real, measured window covers what's left, both aim at the same target: keep what you send inside the zone a model can reliably use, instead of inside the zone it's rated for.

5. Verify the output, because the input side can't catch this on its own. Nothing about a prompt's token count tells you whether the fact that mattered landed somewhere retrievable. The only way to know is to check what came back.

That last step is the one a token budget alone cannot do. Nadir starts with a shadow model-and-effort decision and lets the customer attach an outcome that reveals whether the needed fact survived. Complete non-streaming proxy responses can optionally enter a verifier cascade; streaming bypasses post-generation verification, and verifier errors fail open. That boundary matters for context-window failures because no prompt-side token count proves that the answer used the fact that mattered.


Sources: [Hsieh, Sun, Kriman, Acharya, Rekesh, Jia, Zhang, and Ginsburg, "RULER: What's the Real Context Size of Your Long-Context Language Models?", arXiv:2404.06654 (NVIDIA, April 2024)](https://arxiv.org/abs/2404.06654). [BenchLM, "LLM Context Window Statistics," accessed August 20, 2026](https://benchlm.ai/stats/context-windows). [elvex, "AI Model Context Window Comparison 2026: Advertised vs. Real"](https://www.elvex.com/blog/context-length-comparison-ai-models-2026). [BenchLM, "Advertised vs Effective Context Windows: What 1M-Token Claims Hide"](https://benchlm.ai/blog/posts/context-window-comparison). [byteiota, "Anthropic Drops Long-Context Premium: 1M Tokens at Standard Pricing"](https://byteiota.com/anthropic-drops-long-context-premium-1m-tokens-at-standard-pricing/). Model-specific benchmark and pricing figures are as reported in the sources above, current as of August 2026 and subject to change as vendors update pricing and publish new benchmark results.

More on context & compression

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

For coding agents, Nadir connects to Codex, Claude Code, or Cursor and recommends a model tier for delegated work. The agent decides whether to hand the task off and uses its own configured models, so Nadir needs no proxy and no provider keys on that path.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.