As of today, one model out of 322 backs up a 10-million-token claim with nothing anyone can check.
BenchLM's context-window tracker, dated August 20, 2026, puts a number on the context-window arms race that's been running for the last year: 397 models tracked overall, 322 of them with a published context window, and a quarter of those now claim 1 million tokens or more. The single largest number on the board belongs to Pokee AI's Pokee-Isaac 28B, at 10 million tokens. Llama 4 Scout and Gemini 3 Pro advertise the same ceiling.
The median model in that tracker, though, sits at 256K, a number that's easy to miss next to the 10M headline, and worth sitting with for a second: three out of four models in the entire tracked field don't even claim 1M, let alone 10M. The arms race is real, but it's a race among a minority of the field, not where most deployed context actually lives.
Here's the part that doesn't make it into the spec sheet: how much of an advertised window a model can actually use. That number has its own name, and its own decade-old benchmark, and in 2026 it's still nowhere close to catching up with the marketing.
What RULER actually measures.
In April 2024, a team at NVIDIA (Hsieh, Sun, Kriman, Acharya, Rekesh, Jia, Zhang, and Ginsburg) published RULER: What's the Real Context Size of Your Long-Context Language Models?, arXiv:2404.06654, a synthetic benchmark built specifically to answer the question its title asks. RULER goes past the standard needle-in-a-haystack retrieval test with 13 tasks across four categories, retrieval, multi-hop variable tracing, aggregation, and question answering, because the paper's own finding was that "despite achieving nearly perfect accuracy in the vanilla NIAH test, almost all models exhibit large performance drops as the context length increases." A model can ace the easy version of the long-context test and still fail the harder ones well inside its claimed window.
Testing 17 models across those 13 tasks, RULER found that only half of the models claiming 32K-token context or larger could actually clear a fixed accuracy bar at 32K tokens. Its own words: "almost all models fall below the threshold before reaching the claimed context lengths." That was the 2024 generation of models. The tracked field has grown roughly 20x in headline context size since, and the same structural gap, the industry now widely cites RULER's methodology as finding effective usable context runs somewhere around 60-70% of the advertised maximum, hasn't closed. It's arguably gotten harder to see, because at 10M tokens, nobody's independently re-running a 13-task synthetic probe out to the full claimed length before the next model ships.
The number almost nobody publishes.
If RULER-style testing at the advertised ceiling were routine, this would be a solved problem: buyers would read the effective-length column the way they read the price column. It isn't routine. Pulling together the long-context comparison pages tracking 2026's frontier models, the honest state of the data is closer to this: "the only published row in this set is GPT-5.5 at 87.5 on the MRCRv2 long-context test" in the 128K-256K band, 83.1 in the 64K-128K band, and nothing published near its 1M ceiling. On Artificial Analysis's Long-Context Reasoning benchmark, Claude Fable 5 is reported scoring 77%, one of the only other public data points in the same neighborhood. Llama 4 Scout advertises 10M tokens; as of this writing, no published benchmark shows quality holding anywhere near that length.
That's the actual gap worth naming, and it's not just "effective context is smaller than advertised," which RULER already established in 2024. It's that the industry has kept publishing the size spec every model generation while mostly stopping publishing the quality-at-that-size spec. A context window is now marketed the way storage used to be marketed: the number on the box, not what happens to read speed once you fill it.
Why the bill used to catch this, and doesn't anymore.
There used to be a cost signal that at least hinted when a prompt had wandered past the point of diminishing returns. Anthropic ran a long-context surcharge on Claude prompts over 200K tokens during its 1M-context beta, a 2x input / 1.5x output multiplier. On March 13, 2026, Anthropic dropped it for Claude Opus 4.6 and Sonnet 4.6: standard pricing now runs flat across the entire window, $5/$25 per million tokens for Opus 4.6, $3/$15 for Sonnet 4.6, no premium tier past 200K. One industry write-up put the practical effect plainly: "a 900,000-token request now costs the same per-token rate as a 9,000-token one." Every other frontier vendor already prices this way.
That's a genuinely good change for teams that need the full window and were previously taxed for using it. It also means the one number on an invoice that used to nudge a team toward "maybe check whether this prompt needs to be this long" is gone. Nothing in the bill distinguishes a 900K-token prompt where every token gets used correctly from a 900K-token prompt where the fact the model needed was buried at token 550,000, past whatever fraction of that window is actually reliable, and got missed. Both cost the same. Only one of them gave you an answer worth what you paid for it.
| Model | Advertised 1M input price (list) | Long-context surcharge |
|---|---|---|
| DeepSeek V4 Pro | $0.435 / M | None |
| Gemini 3.1 Pro | $4.00 / M | None |
| GPT-5.5 | $5.00 / M | None |
| Claude Opus (tokenizer-adjusted) | $6.75 / M | Removed March 13, 2026 |
Figures as compiled in 2026 pricing trackers; list prices, not accounting for caching. See sources.
A 71x-cheaper-per-token model and a full-price frontier model both charge you in full for tokens the underlying architecture may not be reliably attending to. Price-per-token was never a proxy for whether the tokens in that prompt were going to be used correctly, and now there's no premium tier left to accidentally flag it either.
This is a quality bug wearing a cost bug's clothes.
The mechanism under the RULER gap is well documented and has a name in its own right: "lost in the middle." Facts placed near the start or end of a long context get retrieved reliably. The same fact placed somewhere in the middle gets missed at a measurably higher rate, and the model doesn't hedge or flag uncertainty when it happens, it just answers confidently with whatever it did retrieve. That's the trap: a context-stuffed prompt that silently drops the one fact that mattered doesn't look different from a correct answer. It looks identical, right up until it's wrong in a way nobody caught before it shipped.
Treating this purely as a cost problem, "we're paying for tokens we don't need," undersells it. The same tokens that don't help the bill also don't help the answer, and past the effective boundary they can actively hurt it, by pushing the fact that mattered further into the unreliable middle of the window instead of leaving it near an edge.
A framework for building to the real ceiling, not the printed one.
1. Treat the advertised window as a purchase spec, not an engineering budget. RULER's own finding, still the best-documented shape of this problem two years and one context-size generation later, is that effective length runs meaningfully under claimed length for most models. Design context budgets around roughly 60-70% of the number on the pricing page, not the number itself, until a specific model's own quality-at-length data says otherwise.
2. Don't assume the ratio transfers across vendors. RULER measured a wide spread across the 17 models it tested in 2024, and the sparse 2026 data available (GPT-5.5's published MRCRv2 rows, Claude Fable 5's AA-LCR score) shows the same thing: this is not one constant you can apply to every provider. A model's effective ratio is closer to a property of that specific model than an industry average.
3. Position what has to be found. Given the lost-in-the-middle effect, a fact your task genuinely cannot afford to miss belongs near the start or end of the prompt, not buried in the middle of retrieved chunks, agent history, or tool output, regardless of how much headroom the advertised window claims to have left.
4. Compress and route against the effective ceiling, not the advertised one. Pruning retrieved context down to what a task actually needs, and picking the model whose real, measured window covers what's left, both aim at the same target: keep what you send inside the zone a model can reliably use, instead of inside the zone it's rated for.
5. Verify the output, because the input side can't catch this on its own. Nothing about a prompt's token count tells you whether the fact that mattered landed somewhere retrievable. The only way to know is to check what came back.
That last step is the one a token budget alone cannot do. Nadir starts with a shadow model-and-effort decision and lets the customer attach an outcome that reveals whether the needed fact survived. Complete non-streaming proxy responses can optionally enter a verifier cascade; streaming bypasses post-generation verification, and verifier errors fail open. That boundary matters for context-window failures because no prompt-side token count proves that the answer used the fact that mattered.
Sources: [Hsieh, Sun, Kriman, Acharya, Rekesh, Jia, Zhang, and Ginsburg, "RULER: What's the Real Context Size of Your Long-Context Language Models?", arXiv:2404.06654 (NVIDIA, April 2024)](https://arxiv.org/abs/2404.06654). [BenchLM, "LLM Context Window Statistics," accessed August 20, 2026](https://benchlm.ai/stats/context-windows). [elvex, "AI Model Context Window Comparison 2026: Advertised vs. Real"](https://www.elvex.com/blog/context-length-comparison-ai-models-2026). [BenchLM, "Advertised vs Effective Context Windows: What 1M-Token Claims Hide"](https://benchlm.ai/blog/posts/context-window-comparison). [byteiota, "Anthropic Drops Long-Context Premium: 1M Tokens at Standard Pricing"](https://byteiota.com/anthropic-drops-long-context-premium-1m-tokens-at-standard-pricing/). Model-specific benchmark and pricing figures are as reported in the sources above, current as of August 2026 and subject to change as vendors update pricing and publish new benchmark results.