TL;DR.
Three labs shipped a new model in three days, all at the same list price. Anthropic released Claude Sonnet 5.5 on September 28, OpenAI released GPT-6.1 Sol on September 29, and Google announced Gemini 4 Argon on September 30. Each one costs $2 per million input tokens and $10 per million output tokens. On a pricing page they look identical.
On a task, they aren't. Artificial Analysis ran all three through its Intelligence Index the next day. At each model's highest setting, one index task cost $0.72 on GPT-6.1 Sol, $1.99 on Gemini 4 Argon, and $7.62 on Claude Sonnet 5.5. Same rate card, a 10.6x spread. The rates didn't decide the bill. How many tokens each model spent on the task did, and the reasoning-effort setting moved that more than the choice of model.
This post covers what each model costs per task, why the effort dial matters more than the logo, what to check before you migrate, and a short tutorial for picking a (model, effort) pair per task type from your own traffic.
Per-task costs and scores are Artificial Analysis readings from October 1, 2026, as reported by [Trending Topics](https://www.trendingtopics.eu/gemini-4-artificial-analysis-en/) and [Aivy](https://aivy.com.au/resources/claude-sonnet-5-5-vs-gpt-6-1-sol/). Aivy reports in Australian dollars; we converted at the 1.405 ratio implied by its own A$2.81 = US$2 input price. Vendor benchmark numbers are vendor-reported. None of this is Nadir customer data.
Three launches, one price.
| Claude Sonnet 5.5 | GPT-6.1 Sol | Gemini 4 Argon | |
|---|---|---|---|
| Released | Sept 28 | Sept 29 | Sept 30 (limited) |
| Input / output, per M | $2 / $10 | $2 / $10 | $2 / $10 intro, then $4 / $20 |
| Cached input, per M | $0.20 | $0.10 | 95% off input ($0.10 intro) |
| Cache write, per M | $2.50 (5-minute) | none listed | not published |
| Long prompt surcharge | none | 2x input, 1.5x output over 272K | not published |
| Context / max output | 1M / 128K | 1.05M / 128K | not published / 1M |
| Who can call it today | Everyone, all three clouds | Everyone, gpt-6.1-sol | Fairwind cyber program only |
Sources: Unite.AI and Developers Digest for Sonnet 5.5, The Next Web and Aivy for GPT-6.1 Sol, 9to5Google and DataCamp for Gemini 4 Argon.
Three things the headline price hides:
- Argon's price is temporary. Google calls $2/$10 introductory and says the standard rate is $4/$20. It hasn't said when the switch happens. Any cost model that hardcodes $2/$10 for Argon is wrong on a date nobody knows yet. We've covered this pattern before with Gemini 3.8 Flash and GPT-5.6 Sol's promo.
- Sol's cached input is half of Sonnet's. $0.10 against $0.20. For a long agent loop that re-reads its context every turn, that column matters more than the input price, the same effect we measured in Fable 5.1 vs GPT-6 Astra.
- Argon isn't generally available. No public model ID yet, and no listing on OpenRouter or Vertex AI as of September 30. You can plan for it. You can't route to it.
The bill is price times tokens.
If three models charge the same per token and their per-task costs differ by 10x, they spent 10x different numbers of tokens. Most of that is reasoning: thinking tokens bill as output, at $10 per million on all three, whether anyone reads them or not. We covered that mechanic in detail here.
Two results from Chart 1 stand out:
- Sonnet 5.5 at max costs more per task than Opus 5.5 at max. $7.62 against $5.98, for a model whose list price is half. Opus scores 1.6 points higher. Paying for the mid-tier model doesn't mean paying mid-tier money if you leave the dial at the top.
- GPT-6 Astra is cheaper per task than Argon at its future standard price, even though Astra's list price is 2.5x higher. Trending Topics reports Argon at about 62,000 output tokens per index task against 27,000 for Astra. Verbosity beat price.
So the honest answer to "which $2/$10 model is cheapest" is: GPT-6.1 Sol, by a lot, on this suite. The more useful question is which one is cheapest for the score you need.
Effort moves the bill more than the model does.
| Effort | Sonnet 5.5 cost | Sonnet 5.5 score | Sol cost | Sol score |
|---|---|---|---|---|
| medium | $0.58 | 40.7 | $0.21 | 47.8 |
| high | $1.08 | 46.7 | $0.32 | 50.2 |
| xhigh | $2.74 | 51.9 | $0.39 | 51.0 |
| max | $7.62 | 56.0 | $0.72 | 51.8 |
Read the two curves, not the two models:
- Sol's curve is flat. From medium to max, cost goes up 3.4x and the score goes up 4 points. Most of Sol's quality is available at medium for 21 cents.
- Sonnet's curve is steep. From medium to max, cost goes up 13x and the score goes up 15 points. Sonnet at medium is the worst point on the chart. Sonnet at max is the best of the three.
- The crossover is at xhigh. Sonnet 5.5 at xhigh and Sol at max score within 0.1 points of each other. Sonnet costs 3.8x as much to get there.
- The top 4 points cost $6.90 a task. If your workload needs what Sonnet 5.5 at max does, that's what it costs, and on some tasks it's worth it. Aivy's sub-scores at max show Sonnet ahead on Terminal-Bench 4.0 (63.6% vs 56.1%) and SciCode (61.0% vs 54.2%), and the two tied on long-context reasoning (82.7% vs 83.0%).
There's one more column the chart doesn't show. Artificial Analysis measured a hallucination rate of 54% for GPT-6.1 Sol and 15% for Gemini 4 Argon, the lowest among the leading models. Sol is cheap per task. If your task punishes a confident wrong answer more than it rewards a fast one, cheap per task can be expensive per outcome. The verification discount post is about exactly that trade.
Latency splits differently again. At max, Sonnet 5.5 writes about 145 tokens a second to Sol's 73, but waits about 265 seconds before its first token against Sol's 130. At high effort it flips: about 7 seconds to Sonnet's first token, about 19 to Sol's. Pick the effort level before you compare speed.
What a single benchmark can't tell you.
The Intelligence Index is one suite of tasks, and your traffic isn't that suite. Three reasons not to take Chart 2 as your routing table:
- Token counts depend on the prompt. Anthropic says Sonnet 5.5 costs "up to 30% less per task" than Sonnet 5. Customer numbers in the launch coverage range from 12% fewer total tokens at Box to about 121,000 tokens per answer against Sonnet 5's 497,000 at Balyasny. Same model, very different savings.
- Vendor and independent numbers disagree. Anthropic reports 70.6% on Terminal-Bench 4.0 for Sonnet 5.5. Artificial Analysis measured 63.6% at max. Neither is wrong; harness and settings differ. Your harness will differ too.
- The default effort isn't the cheapest. If you migrate a call site without setting effort explicitly, you inherit the vendor default, which isn't chosen for your workload.
The fix is cheap: run your own tasks through each candidate at two or three effort levels and count the cost per passing task.
Before you migrate to Sonnet 5.5.
The price didn't change, but the API did. Developers Digest lists several changes that return errors on code written for Sonnet 5. The ones most likely to break a cost-optimized pipeline:
- Non-default `temperature`, `top_p`, or `top_k` now fail. Deterministic extraction pipelines that pin
temperature=0will need that removed. - Forced tool choice is gone.
tool_choiceset toanyor a specific tool returns a 400. The replacement isautoplusstrict: trueon tool definitions, with a limit of 20 strict tools per request. - `thinking: disabled` is rejected. Use
between_tools, which isn't available atxhighormaxeffort. - The minimum cacheable prompt drops to 512 tokens, from 1,024. Short system prompts that never cached before now can, which is a small unannounced saving.
Run the bake-off below on the new model before you flip a model string in production. A 400 at 2 a.m. is the most expensive kind of migration.
Tutorial: pick a model and an effort per task type.
The decision isn't "which model." It's "which (model, effort) pair, for which kind of task." Four steps, through one OpenAI-compatible endpoint so all candidates are one model string apart.
Step 1: sample real tasks and a pass check.
Take 50 to 200 recent requests per task type: support replies, extraction, code edits, whatever your app actually sends. For each one, write the cheapest check you trust: a JSON schema, a unit test, an exact match, or a stronger model as judge on a sample. Without a pass check you can measure cost, but not whether a cheaper setting hurt anything.
Step 2: run the grid.
import itertools, json
from openai import OpenAI
client = OpenAI(base_url="https://api.getnadir.com/v1", api_key=NADIR_KEY)
# USD per million tokens: (input, cached input, output).
# Argon is commented out until it has a public model ID.
PRICES = {
"claude-sonnet-5-5": (2.00, 0.20, 10.00),
"gpt-6.1-sol": (2.00, 0.10, 10.00),
# "gemini-4-argon": (2.00, 0.10, 10.00), # intro; standard is (4.00, 0.20, 20.00)
}
EFFORTS = ["medium", "high", "xhigh"]
def cost(model, usage):
p_in, p_cached, p_out = PRICES[model]
cached = getattr(usage.prompt_tokens_details, "cached_tokens", 0) or 0
fresh = usage.prompt_tokens - cached
return (fresh * p_in + cached * p_cached + usage.completion_tokens * p_out) / 1e6
def run_grid(tasks, task_type):
rows = []
for (model, effort), task in itertools.product(
itertools.product(PRICES, EFFORTS), tasks):
r = client.chat.completions.create(
model=model,
messages=task["messages"],
reasoning={"effort": effort},
extra_headers={"X-Nadir-Tags": f"bakeoff,{task_type},{model},{effort}"},
)
rows.append({
"model": model, "effort": effort, "task_type": task_type,
"cost": cost(model, r.usage),
"passed": task["check"](r.choices[0].message.content),
})
return rows
In OpenAI-compatible responses, completion_tokens includes reasoning tokens. That's the whole point: it's the number the rate card doesn't show you.
Step 3: score by cost per passing task.
import pandas as pd
df = pd.DataFrame(rows)
table = (df.groupby(["task_type", "model", "effort"])
.agg(pass_rate=("passed", "mean"), spend=("cost", "sum"),
passes=("passed", "sum"))
.assign(cost_per_pass=lambda t: t.spend / t.passes)
.reset_index())
def pick(group, floor=0.95):
"""Cheapest setting within `floor` of the best pass rate for this task type."""
best = group.pass_rate.max()
ok = group[group.pass_rate >= best * floor]
return ok.sort_values("cost_per_pass").iloc[0]
policy = table.groupby("task_type").apply(pick)
print(policy[["model", "effort", "pass_rate", "cost_per_pass"]])
Cost per pass, not cost per call. A setting that's half the price and fails twice as often is the same price with worse latency. The 95% floor is a knob: tighten it for tasks where a wrong answer is expensive, loosen it for drafts a human edits anyway.
Step 4: ship the table, with an expiry.
ROUTES = {
# task_type: (model, effort), from Step 3
"extraction": ("gpt-6.1-sol", "medium"),
"support": ("gpt-6.1-sol", "high"),
"code_edit": ("claude-sonnet-5-5", "xhigh"),
}
REVIEW_BY = "2026-11-01" # re-run Step 2 when prices or models change
Write the review date into the config. Argon's price will double on a date Google hasn't announced. Sonnet 5.5's token counts will shift as Anthropic tunes it. A routing table that never gets re-run slowly turns into a hardcoded model choice.
Where Nadir fits.
The tutorial above is the manual version of what Nadir does in the request path:
- One endpoint, every model above. OpenAI-compatible, with BYOK for your Anthropic, OpenAI, and Google keys, so the bake-off is a model string per call instead of three SDKs.
- One effort field. Send
reasoning={"effort": ...}and Nadir maps it to each provider's own parameter. - Cost per call, tagged.
X-Nadir-Tagsrolls spend up by task type, model, and effort in the dashboard, which is the input to Step 3. - Per-request routing for everything you haven't benchmarked. Send
model="auto"and Nadir picks the tier per prompt before any provider tokens are spent, with per-request reasons in the response headers. - Observe mode serves exactly the model you asked for while recording what routing would have picked, so you can size the saving on your own traffic before you change anything.
If you moved a call site to one of this week's $2/$10 models without setting its effort, start with a free key, tag a week of traffic, and look at cost per task by model and effort. Chart 2 says the answer can move by an order of magnitude.
Conclusion.
Sonnet 5.5, GPT-6.1 Sol, and Gemini 4 Argon all list at $2/$10, and on the first independent run they cost $7.62, $0.72, and $1.99 per task at their top settings. Same price, 10x apart, because the rate card prices tokens and the models spend very different numbers of them. Within one model, the effort dial moved cost 3x on Sol and 13x on Sonnet. Sol is cheapest on this suite and flattest across effort levels; Sonnet 5.5 at max is the strongest of the three and pays for it; Argon scores between them with the lowest hallucination rate, at a price that will double on an unannounced date. When sticker prices converge, the bill is decided by tokens per passing task. Measure that per task type, pick a model and an effort for each, and put a review date on the table.
Per-task costs and scores are Artificial Analysis Intelligence Index readings reported on October 1, 2026, and may change as providers tune their models. Sources: [Trending Topics, "Gemini 4 Matches GPT-6 Astra but Trails Opus 5.5"](https://www.trendingtopics.eu/gemini-4-artificial-analysis-en/). [Aivy, "Sonnet 5.5 vs GPT 6.1 Sol: same price, different bill"](https://aivy.com.au/resources/claude-sonnet-5-5-vs-gpt-6-1-sol/). [Unite.AI, "Anthropic Releases Claude Sonnet 5.5 at Unchanged Sonnet 5 Pricing"](https://www.unite.ai/anthropic-releases-claude-sonnet-5-5-at-unchanged-sonnet-5-pricing/). [Developers Digest, "Claude Sonnet 5.5 Developer Guide"](https://www.developersdigest.tech/blog/claude-sonnet-5-5-release-guide-2026). [The Next Web, "OpenAI releases GPT-6.1 Sol at a fifth of GPT-6 Astra's token prices"](https://thenextweb.com/news/openai-gpt-6-1-sol-price-astra-devday). [9to5Google, "Google announces Gemini 4 Argon"](https://9to5google.com/2026/09/30/gemini-4-argon-announcement/). [DataCamp, "Gemini 4 Argon: Benchmarks, Pricing, and Access"](https://www.datacamp.com/blog/gemini-4-argon).