On August 12, 2026, xAI shipped Grok 4.6 with a 500,000-token context window and a pricing mechanic its own developer docs state in plain language: once a prompt reaches 200,000 tokens, the entire request, not just the tokens past the line, bills at double the rate. Below the line, Grok 4.6 costs $2.00 per million input tokens and $6.00 per million output. At or above it, every token in that request, including the first one, bills at $4.00 and $12.00. That is not a Grok idiosyncrasy. Google's Gemini 3.1 Pro doubles input and raises output 50% at the identical 200,000-token line. OpenAI's GPT-5.6 Sol does the same 2x input / 1.5x output jump at 272,000 tokens, a threshold this blog covered the day it last moved. Only Anthropic's Claude Sonnet 5 charges the same $2/$10 per million tokens whether a request is 9,000 tokens or 900,000. Three of the four frontier providers now draw a whole-request repricing line somewhere between 200,000 and 272,000 tokens, with no shared standard for exactly where it sits, how steep the jump is, or whether it applies to input, output, or both. Here's the mechanic, the four-provider comparison, and what actually breaks when a request crosses a line nobody on the calling side is watching for.
The mechanic: whole-request, not marginal.
The natural assumption is that long-context pricing works like a tax bracket: the first 200,000 tokens bill at the standard rate, and only the tokens past that line bill at the higher one. That is not how any of these three providers built it. xAI's pricing documentation states that requests whose prompt reaches the listed token threshold are billed at the higher rate for all tokens in the request. A 201,000-token prompt does not pay the premium rate on the last 1,000 tokens. It pays the premium rate on all 201,000.
| Tier | Input | Cached input | Output |
|---|---|---|---|
| Under 200,000 tokens | $2.00 / M | $0.50 / M | $6.00 / M |
| 200,000 tokens or more | $4.00 / M | $1.00 / M | $12.00 / M |
Every rate in the bottom row applies to the whole request the moment it crosses the line, cached tokens included. A prompt that reads 199,999 tokens and one that reads 200,000 tokens can differ by a single token and differ in cost by roughly double.
Run the arithmetic yourself: 199,000 input tokens at $2.00/M plus 4,000 output tokens at $6.00/M comes to $0.422. Add 2,000 tokens of input, crossing the line, and the same request becomes 201,000 tokens at $4.00/M plus 4,000 output tokens at $12.00/M: $0.852. A prompt that grew by roughly 1% saw its bill grow by 102%.
Three providers, one line. One provider, none.
Grok 4.6 isn't alone at 200,000 tokens, and the 272,000-token line OpenAI drew for GPT-5.6 Sol isn't new information either, it's the same mechanic this blog wrote about three weeks earlier under a different model name. What's new is seeing all four current frontier options lined up next to each other:
| Model | Context window | Cliff | Below cliff | At/above cliff |
|---|---|---|---|---|
| Grok 4.6 (xAI) | 500,000 | 200,000 | $2.00 / $6.00 per M | $4.00 / $12.00 per M (2x / 2x) |
| Gemini 3.1 Pro (Google) | 1,048,576 | 200,000 | $2.00 / $12.00 per M | $4.00 / $18.00 per M (2x / 1.5x) |
| GPT-5.6 Sol (OpenAI) | ~1,050,000 | 272,000 | $5.00 / $30.00 per M | $10.00 / $45.00 per M (2x / 1.5x) |
| Claude Sonnet 5 (Anthropic) | 1,000,000 | none | $2.00 / $10.00 per M | same, no threshold |
Three-quarters of the frontier field now prices this way, and the three that do disagree with each other on where the line sits (200K for two of them, 272K for the third) and how steep the jump is (Grok doubles both input and output; Gemini and OpenAI double input but only raise output 50%). A routing layer built to avoid one provider's cliff learns nothing transferable about the other two. A prompt engineered to stay just under Gemini's line will still cross Grok's, at a different dollar cost, because the two lines sit at the same token count but the multipliers aren't the same.
Why 200K, and why now.
None of the three providers with a cliff has published the underlying cost model, but the shape of the number is not a coincidence. Transformer attention cost scales worse than linearly with sequence length, so a provider serving a 250,000-token request isn't paying 25% more compute than a 200,000-token one, the actual compute curve bends upward faster than the token count does. A flat per-token rate that holds all the way to a million tokens, the way Claude Sonnet 5's does, means the provider is absorbing that curve at the top end rather than passing it through. A cliff is the provider's way of stopping that absorption at a specific point instead of raising the flat rate for every request to cover the tail. We've written about the compute reality behind long-context degradation before: the token count alone was never the full story on what a long request actually costs to serve, and a repricing line at 200K is the provider-side admission of the same fact the client side already knows from watching quality degrade well before the context window fills up.
Where this bites without anyone deciding it should.
Nobody designs a request to be exactly 201,000 tokens. It happens as a side effect of what the request accumulates: a coding agent that keeps the last several tool results and file reads in context, a RAG pipeline that retrieves one more document than the query needed, a multi-turn conversation where the transcript itself becomes the largest input in the room. Each of those is a workload this blog has already looked at from the token-count side, react-style agent loops that grow their own trajectory turn over turn, retrieval systems that over-fetch past what a query needs, and trajectory pruning research that shows most of what accumulates in a long agent session is stale rather than load-bearing. None of that prior analysis had to account for a request's cost function changing shape mid-growth. A pricing cliff adds a second, sharper failure mode on top of the first: the token count wasn't just wasteful, it was also, at some exact point nobody flagged, twice as expensive per token that was there before the line was crossed.
What a routing layer actually needs to know.
A model router that only compares list prices per million tokens is comparing numbers that don't apply once a request is long enough to matter. Four things a routing decision needs that a static price table doesn't give it:
- The actual token count of this specific request, measured, not estimated, against the specific threshold of the specific model being considered, because the threshold varies by provider and doesn't show up in a per-token rate card.
- Which side of the line the request lands on before it's sent, not after the invoice arrives, since none of the three providers with a cliff exposes a warning in the response when a request crosses it.
- The fact that trimming context by even one token can matter more than switching models, when the alternative to crossing a cliff is removing 1,001 tokens of stale trajectory rather than routing to a different provider entirely.
- A threshold table that gets checked on every request, not written once during a pricing review, given that GPT-5.6 Sol's own line moved three times in ten days in July before settling, and nothing says these three providers' current thresholds are permanent either.
This is close to what Nadir already does for provider repricing generally: score each request against current pricing rather than a snapshot from a past review. When a benchmark is configured and priced, `nadir_metadata.benchmark_comparison.savings_usd` reflects that comparison, including the applicable pricing cliff.
Conclusion.
A single-provider cliff reads like a vendor-specific quirk, something to note in one integration's documentation and move past. Three cliffs, at three different token counts, with three different multiplier profiles, on three of the four current frontier models, is a pattern serious enough that comparing raw per-token prices across providers has quietly stopped being sufficient. The fourth provider's flat rate isn't proof the other three are wrong to charge more for genuinely more expensive requests, it's proof that "charge the same per token no matter how long the request is" and "charge double past a threshold" are both live design choices right now, and a workload routed across more than one of these models needs to know which choice it's currently facing on each individual call, not just what it faced when the integration was first built.
Sources: [xAI, Grok API pricing and model documentation, accessed August 14, 2026](https://docs.x.ai/docs/models). [Google, Gemini API pricing, accessed August 14, 2026](https://ai.google.dev/gemini-api/docs/pricing). [OpenAI, GPT-5.6 Sol API pricing, accessed August 14, 2026](https://platform.openai.com/docs/pricing). [Anthropic, Claude API pricing, accessed August 14, 2026](https://platform.claude.com/docs/en/about-claude/pricing).