The 200K Line

On August 12, 2026, xAI shipped Grok 4.6 with a 500,000-token context window and a price that doubles once input crosses 200,000 tokens. Its own docs state the mechanic plainly: the entire request, not just the tokens past the line, bills at $4.00 input and $12.00 output per million instead of $2.00 and $6.00. That's not a Grok idiosyncrasy. Google's Gemini 3.1 Pro doubles input and raises output 50% at the identical 200,000-token line. OpenAI's GPT-5.6 Sol does the same 2x input / 1.5x output jump at 272,000 tokens, a threshold this blog covered the day it last moved. Only Anthropic's Claude Sonnet 5 charges the same $2/$10 per million tokens whether a request is 9,000 tokens or 900,000. Run the math on a single Grok request: add 2,000 tokens to a 199,000-token prompt, about 1% more input, and the bill doesn't rise 1%. It rises 102%. Three of the four current frontier providers now draw a whole-request repricing line somewhere between 200,000 and 272,000 tokens, with no shared standard for where it sits or how steep the jump is. Here's the mechanic, the four-provider comparison, and what a routing decision actually needs to know before the invoice tells it.

Published 2026-08-14 by Dor Amir on the Nadir blog.

Filed under Pricing & Models.

On August 12, 2026, xAI shipped Grok 4.6 with a 500,000-token context window and a pricing mechanic its own developer docs state in plain language: once a prompt reaches 200,000 tokens, the entire request, not just the tokens past the line, bills at double the rate. Below the line, Grok 4.6 costs $2.00 per million input tokens and $6.00 per million output. At or above it, every token in that request, including the first one, bills at $4.00 and $12.00. That is not a Grok idiosyncrasy. Google's Gemini 3.1 Pro doubles input and raises output 50% at the identical 200,000-token line. OpenAI's GPT-5.6 Sol does the same 2x input / 1.5x output jump at 272,000 tokens, a threshold this blog covered the day it last moved. Only Anthropic's Claude Sonnet 5 charges the same $2/$10 per million tokens whether a request is 9,000 tokens or 900,000. Three of the four frontier providers now draw a whole-request repricing line somewhere between 200,000 and 272,000 tokens, with no shared standard for exactly where it sits, how steep the jump is, or whether it applies to input, output, or both. Here's the mechanic, the four-provider comparison, and what actually breaks when a request crosses a line nobody on the calling side is watching for.

The mechanic: whole-request, not marginal.

The natural assumption is that long-context pricing works like a tax bracket: the first 200,000 tokens bill at the standard rate, and only the tokens past that line bill at the higher one. That is not how any of these three providers built it. xAI's pricing documentation states that requests whose prompt reaches the listed token threshold are billed at the higher rate for all tokens in the request. A 201,000-token prompt does not pay the premium rate on the last 1,000 tokens. It pays the premium rate on all 201,000.

TierInputCached inputOutput
Under 200,000 tokens$2.00 / M$0.50 / M$6.00 / M
200,000 tokens or more$4.00 / M$1.00 / M$12.00 / M

Every rate in the bottom row applies to the whole request the moment it crosses the line, cached tokens included. A prompt that reads 199,999 tokens and one that reads 200,000 tokens can differ by a single token and differ in cost by roughly double.

Chart 1: the same Grok 4.6 request costed at 199,000 versus 201,000 input tokens, 4,000 output tokens held constant. $0.42 becomes $0.85, a 102% increase from 2,000 additional tokens, about 1% more input.
Chart 1: the same Grok 4.6 request costed at 199,000 versus 201,000 input tokens, 4,000 output tokens held constant. $0.42 becomes $0.85, a 102% increase from 2,000 additional tokens, about 1% more input.

Run the arithmetic yourself: 199,000 input tokens at $2.00/M plus 4,000 output tokens at $6.00/M comes to $0.422. Add 2,000 tokens of input, crossing the line, and the same request becomes 201,000 tokens at $4.00/M plus 4,000 output tokens at $12.00/M: $0.852. A prompt that grew by roughly 1% saw its bill grow by 102%.

Three providers, one line. One provider, none.

Grok 4.6 isn't alone at 200,000 tokens, and the 272,000-token line OpenAI drew for GPT-5.6 Sol isn't new information either, it's the same mechanic this blog wrote about three weeks earlier under a different model name. What's new is seeing all four current frontier options lined up next to each other:

ModelContext windowCliffBelow cliffAt/above cliff
Grok 4.6 (xAI)500,000200,000$2.00 / $6.00 per M$4.00 / $12.00 per M (2x / 2x)
Gemini 3.1 Pro (Google)1,048,576200,000$2.00 / $12.00 per M$4.00 / $18.00 per M (2x / 1.5x)
GPT-5.6 Sol (OpenAI)~1,050,000272,000$5.00 / $30.00 per M$10.00 / $45.00 per M (2x / 1.5x)
Claude Sonnet 5 (Anthropic)1,000,000none$2.00 / $10.00 per Msame, no threshold
Chart 2: four frontier models' pricing over their full context window. Grok 4.6, Gemini 3.1 Pro, and GPT-5.6 Sol each turn to a whole-request premium rate near 200K–272K tokens. Claude Sonnet 5 charges the same rate from the first token to the millionth.
Chart 2: four frontier models' pricing over their full context window. Grok 4.6, Gemini 3.1 Pro, and GPT-5.6 Sol each turn to a whole-request premium rate near 200K–272K tokens. Claude Sonnet 5 charges the same rate from the first token to the millionth.

Three-quarters of the frontier field now prices this way, and the three that do disagree with each other on where the line sits (200K for two of them, 272K for the third) and how steep the jump is (Grok doubles both input and output; Gemini and OpenAI double input but only raise output 50%). A routing layer built to avoid one provider's cliff learns nothing transferable about the other two. A prompt engineered to stay just under Gemini's line will still cross Grok's, at a different dollar cost, because the two lines sit at the same token count but the multipliers aren't the same.

Why 200K, and why now.

None of the three providers with a cliff has published the underlying cost model, but the shape of the number is not a coincidence. Transformer attention cost scales worse than linearly with sequence length, so a provider serving a 250,000-token request isn't paying 25% more compute than a 200,000-token one, the actual compute curve bends upward faster than the token count does. A flat per-token rate that holds all the way to a million tokens, the way Claude Sonnet 5's does, means the provider is absorbing that curve at the top end rather than passing it through. A cliff is the provider's way of stopping that absorption at a specific point instead of raising the flat rate for every request to cover the tail. We've written about the compute reality behind long-context degradation before: the token count alone was never the full story on what a long request actually costs to serve, and a repricing line at 200K is the provider-side admission of the same fact the client side already knows from watching quality degrade well before the context window fills up.

Where this bites without anyone deciding it should.

Nobody designs a request to be exactly 201,000 tokens. It happens as a side effect of what the request accumulates: a coding agent that keeps the last several tool results and file reads in context, a RAG pipeline that retrieves one more document than the query needed, a multi-turn conversation where the transcript itself becomes the largest input in the room. Each of those is a workload this blog has already looked at from the token-count side, react-style agent loops that grow their own trajectory turn over turn, retrieval systems that over-fetch past what a query needs, and trajectory pruning research that shows most of what accumulates in a long agent session is stale rather than load-bearing. None of that prior analysis had to account for a request's cost function changing shape mid-growth. A pricing cliff adds a second, sharper failure mode on top of the first: the token count wasn't just wasteful, it was also, at some exact point nobody flagged, twice as expensive per token that was there before the line was crossed.

What a routing layer actually needs to know.

A model router that only compares list prices per million tokens is comparing numbers that don't apply once a request is long enough to matter. Four things a routing decision needs that a static price table doesn't give it:

  1. The actual token count of this specific request, measured, not estimated, against the specific threshold of the specific model being considered, because the threshold varies by provider and doesn't show up in a per-token rate card.
  2. Which side of the line the request lands on before it's sent, not after the invoice arrives, since none of the three providers with a cliff exposes a warning in the response when a request crosses it.
  3. The fact that trimming context by even one token can matter more than switching models, when the alternative to crossing a cliff is removing 1,001 tokens of stale trajectory rather than routing to a different provider entirely.
  4. A threshold table that gets checked on every request, not written once during a pricing review, given that GPT-5.6 Sol's own line moved three times in ten days in July before settling, and nothing says these three providers' current thresholds are permanent either.

This is close to what Nadir already does for provider repricing generally: score each request against current pricing rather than a snapshot from a past review. When a benchmark is configured and priced, `nadir_metadata.benchmark_comparison.savings_usd` reflects that comparison, including the applicable pricing cliff.

Conclusion.

A single-provider cliff reads like a vendor-specific quirk, something to note in one integration's documentation and move past. Three cliffs, at three different token counts, with three different multiplier profiles, on three of the four current frontier models, is a pattern serious enough that comparing raw per-token prices across providers has quietly stopped being sufficient. The fourth provider's flat rate isn't proof the other three are wrong to charge more for genuinely more expensive requests, it's proof that "charge the same per token no matter how long the request is" and "charge double past a threshold" are both live design choices right now, and a workload routed across more than one of these models needs to know which choice it's currently facing on each individual call, not just what it faced when the integration was first built.


Sources: [xAI, Grok API pricing and model documentation, accessed August 14, 2026](https://docs.x.ai/docs/models). [Google, Gemini API pricing, accessed August 14, 2026](https://ai.google.dev/gemini-api/docs/pricing). [OpenAI, GPT-5.6 Sol API pricing, accessed August 14, 2026](https://platform.openai.com/docs/pricing). [Anthropic, Claude API pricing, accessed August 14, 2026](https://platform.claude.com/docs/en/about-claude/pricing).

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.