Three at $2/$10

Claude Sonnet 5.5, GPT-6.1 Sol, and Gemini 4 Argon all launched in the same three days at $2 per million input tokens and $10 per million output. On Artificial Analysis's Intelligence Index, one task cost $0.72 on Sol, $1.99 on Argon, and $7.62 on Sonnet 5.5 at each model's top setting: same rate card, 10.6x apart. Reasoning effort moved the bill more than the model did, 3.4x across Sol's settings and 13x across Sonnet's, and Sonnet at xhigh matches Sol at max for 3.8x the price. Plus Argon's introductory price that doubles on an unannounced date, the Sonnet 5.5 API changes that return 400s on old code, and a Python bake-off that picks a model and an effort per task type by cost per passing task.

Published 2026-10-02 by Dor Amir on the Nadir blog.

Filed under Pricing & Models.

TL;DR.

Three labs shipped a new model in three days, all at the same list price. Anthropic released Claude Sonnet 5.5 on September 28, OpenAI released GPT-6.1 Sol on September 29, and Google announced Gemini 4 Argon on September 30. Each one costs $2 per million input tokens and $10 per million output tokens. On a pricing page they look identical.

On a task, they aren't. Artificial Analysis ran all three through its Intelligence Index the next day. At each model's highest setting, one index task cost $0.72 on GPT-6.1 Sol, $1.99 on Gemini 4 Argon, and $7.62 on Claude Sonnet 5.5. Same rate card, a 10.6x spread. The rates didn't decide the bill. How many tokens each model spent on the task did, and the reasoning-effort setting moved that more than the choice of model.

This post covers what each model costs per task, why the effort dial matters more than the logo, what to check before you migrate, and a short tutorial for picking a (model, effort) pair per task type from your own traffic.

Per-task costs and scores are Artificial Analysis readings from October 1, 2026, as reported by [Trending Topics](https://www.trendingtopics.eu/gemini-4-artificial-analysis-en/) and [Aivy](https://aivy.com.au/resources/claude-sonnet-5-5-vs-gpt-6-1-sol/). Aivy reports in Australian dollars; we converted at the 1.405 ratio implied by its own A$2.81 = US$2 input price. Vendor benchmark numbers are vendor-reported. None of this is Nadir customer data.

Three launches, one price.

Claude Sonnet 5.5GPT-6.1 SolGemini 4 Argon
ReleasedSept 28Sept 29Sept 30 (limited)
Input / output, per M$2 / $10$2 / $10$2 / $10 intro, then $4 / $20
Cached input, per M$0.20$0.1095% off input ($0.10 intro)
Cache write, per M$2.50 (5-minute)none listednot published
Long prompt surchargenone2x input, 1.5x output over 272Knot published
Context / max output1M / 128K1.05M / 128Knot published / 1M
Who can call it todayEveryone, all three cloudsEveryone, gpt-6.1-solFairwind cyber program only

Sources: Unite.AI and Developers Digest for Sonnet 5.5, The Next Web and Aivy for GPT-6.1 Sol, 9to5Google and DataCamp for Gemini 4 Argon.

Three things the headline price hides:

  1. Argon's price is temporary. Google calls $2/$10 introductory and says the standard rate is $4/$20. It hasn't said when the switch happens. Any cost model that hardcodes $2/$10 for Argon is wrong on a date nobody knows yet. We've covered this pattern before with Gemini 3.8 Flash and GPT-5.6 Sol's promo.
  2. Sol's cached input is half of Sonnet's. $0.10 against $0.20. For a long agent loop that re-reads its context every turn, that column matters more than the input price, the same effect we measured in Fable 5.1 vs GPT-6 Astra.
  3. Argon isn't generally available. No public model ID yet, and no listing on OpenRouter or Vertex AI as of September 30. You can plan for it. You can't route to it.

The bill is price times tokens.

Chart 1: Cost to run one Artificial Analysis Intelligence Index task at each model's highest setting. GPT-6.1 Sol $0.72, index 51.8. Gemini 4 Argon $1.99 at the introductory price, $3.98 at the standard $4/$20 price, index 52.6. Claude Sonnet 5.5 $7.62, index 56.0. For reference, GPT-6 Astra at $10/$50 costs $3.26 per task with index 52.7, and Claude Opus 5.5 at $4/$20 costs $5.98 with index 57.6.
Chart 1: Cost to run one Artificial Analysis Intelligence Index task at each model's highest setting. GPT-6.1 Sol $0.72, index 51.8. Gemini 4 Argon $1.99 at the introductory price, $3.98 at the standard $4/$20 price, index 52.6. Claude Sonnet 5.5 $7.62, index 56.0. For reference, GPT-6 Astra at $10/$50 costs $3.26 per task with index 52.7, and Claude Opus 5.5 at $4/$20 costs $5.98 with index 57.6.

If three models charge the same per token and their per-task costs differ by 10x, they spent 10x different numbers of tokens. Most of that is reasoning: thinking tokens bill as output, at $10 per million on all three, whether anyone reads them or not. We covered that mechanic in detail here.

Two results from Chart 1 stand out:

So the honest answer to "which $2/$10 model is cheapest" is: GPT-6.1 Sol, by a lot, on this suite. The more useful question is which one is cheapest for the score you need.

Effort moves the bill more than the model does.

Chart 2: Intelligence Index score against USD per task on a log scale. GPT-6.1 Sol: medium $0.21 scores 47.8, high $0.32 scores 50.2, xhigh $0.39 scores 51.0, max $0.72 scores 51.8. Claude Sonnet 5.5: medium $0.58 scores 40.7, high $1.08 scores 46.7, xhigh $2.74 scores 51.9, max $7.62 scores 56.0. Gemini 4 Argon high $1.99 scores 52.6, $3.98 at standard pricing. Opus 5.5 at $5.98 scores 57.6 and Astra at $3.26 scores 52.7 for reference.
Chart 2: Intelligence Index score against USD per task on a log scale. GPT-6.1 Sol: medium $0.21 scores 47.8, high $0.32 scores 50.2, xhigh $0.39 scores 51.0, max $0.72 scores 51.8. Claude Sonnet 5.5: medium $0.58 scores 40.7, high $1.08 scores 46.7, xhigh $2.74 scores 51.9, max $7.62 scores 56.0. Gemini 4 Argon high $1.99 scores 52.6, $3.98 at standard pricing. Opus 5.5 at $5.98 scores 57.6 and Astra at $3.26 scores 52.7 for reference.
EffortSonnet 5.5 costSonnet 5.5 scoreSol costSol score
medium$0.5840.7$0.2147.8
high$1.0846.7$0.3250.2
xhigh$2.7451.9$0.3951.0
max$7.6256.0$0.7251.8

Read the two curves, not the two models:

  1. Sol's curve is flat. From medium to max, cost goes up 3.4x and the score goes up 4 points. Most of Sol's quality is available at medium for 21 cents.
  2. Sonnet's curve is steep. From medium to max, cost goes up 13x and the score goes up 15 points. Sonnet at medium is the worst point on the chart. Sonnet at max is the best of the three.
  3. The crossover is at xhigh. Sonnet 5.5 at xhigh and Sol at max score within 0.1 points of each other. Sonnet costs 3.8x as much to get there.
  4. The top 4 points cost $6.90 a task. If your workload needs what Sonnet 5.5 at max does, that's what it costs, and on some tasks it's worth it. Aivy's sub-scores at max show Sonnet ahead on Terminal-Bench 4.0 (63.6% vs 56.1%) and SciCode (61.0% vs 54.2%), and the two tied on long-context reasoning (82.7% vs 83.0%).

There's one more column the chart doesn't show. Artificial Analysis measured a hallucination rate of 54% for GPT-6.1 Sol and 15% for Gemini 4 Argon, the lowest among the leading models. Sol is cheap per task. If your task punishes a confident wrong answer more than it rewards a fast one, cheap per task can be expensive per outcome. The verification discount post is about exactly that trade.

Latency splits differently again. At max, Sonnet 5.5 writes about 145 tokens a second to Sol's 73, but waits about 265 seconds before its first token against Sol's 130. At high effort it flips: about 7 seconds to Sonnet's first token, about 19 to Sol's. Pick the effort level before you compare speed.

What a single benchmark can't tell you.

The Intelligence Index is one suite of tasks, and your traffic isn't that suite. Three reasons not to take Chart 2 as your routing table:

The fix is cheap: run your own tasks through each candidate at two or three effort levels and count the cost per passing task.

Before you migrate to Sonnet 5.5.

The price didn't change, but the API did. Developers Digest lists several changes that return errors on code written for Sonnet 5. The ones most likely to break a cost-optimized pipeline:

Run the bake-off below on the new model before you flip a model string in production. A 400 at 2 a.m. is the most expensive kind of migration.

Tutorial: pick a model and an effort per task type.

The decision isn't "which model." It's "which (model, effort) pair, for which kind of task." Four steps, through one OpenAI-compatible endpoint so all candidates are one model string apart.

Step 1: sample real tasks and a pass check.

Take 50 to 200 recent requests per task type: support replies, extraction, code edits, whatever your app actually sends. For each one, write the cheapest check you trust: a JSON schema, a unit test, an exact match, or a stronger model as judge on a sample. Without a pass check you can measure cost, but not whether a cheaper setting hurt anything.

Step 2: run the grid.

import itertools, json
from openai import OpenAI

client = OpenAI(base_url="https://api.getnadir.com/v1", api_key=NADIR_KEY)

# USD per million tokens: (input, cached input, output).
# Argon is commented out until it has a public model ID.
PRICES = {
    "claude-sonnet-5-5": (2.00, 0.20, 10.00),
    "gpt-6.1-sol":       (2.00, 0.10, 10.00),
    # "gemini-4-argon":  (2.00, 0.10, 10.00),  # intro; standard is (4.00, 0.20, 20.00)
}
EFFORTS = ["medium", "high", "xhigh"]

def cost(model, usage):
    p_in, p_cached, p_out = PRICES[model]
    cached = getattr(usage.prompt_tokens_details, "cached_tokens", 0) or 0
    fresh = usage.prompt_tokens - cached
    return (fresh * p_in + cached * p_cached + usage.completion_tokens * p_out) / 1e6

def run_grid(tasks, task_type):
    rows = []
    for (model, effort), task in itertools.product(
            itertools.product(PRICES, EFFORTS), tasks):
        r = client.chat.completions.create(
            model=model,
            messages=task["messages"],
            reasoning={"effort": effort},
            extra_headers={"X-Nadir-Tags": f"bakeoff,{task_type},{model},{effort}"},
        )
        rows.append({
            "model": model, "effort": effort, "task_type": task_type,
            "cost": cost(model, r.usage),
            "passed": task["check"](r.choices[0].message.content),
        })
    return rows

In OpenAI-compatible responses, completion_tokens includes reasoning tokens. That's the whole point: it's the number the rate card doesn't show you.

Step 3: score by cost per passing task.

import pandas as pd

df = pd.DataFrame(rows)
table = (df.groupby(["task_type", "model", "effort"])
           .agg(pass_rate=("passed", "mean"), spend=("cost", "sum"),
                passes=("passed", "sum"))
           .assign(cost_per_pass=lambda t: t.spend / t.passes)
           .reset_index())

def pick(group, floor=0.95):
    """Cheapest setting within `floor` of the best pass rate for this task type."""
    best = group.pass_rate.max()
    ok = group[group.pass_rate >= best * floor]
    return ok.sort_values("cost_per_pass").iloc[0]

policy = table.groupby("task_type").apply(pick)
print(policy[["model", "effort", "pass_rate", "cost_per_pass"]])

Cost per pass, not cost per call. A setting that's half the price and fails twice as often is the same price with worse latency. The 95% floor is a knob: tighten it for tasks where a wrong answer is expensive, loosen it for drafts a human edits anyway.

Step 4: ship the table, with an expiry.

ROUTES = {
    # task_type: (model, effort), from Step 3
    "extraction":   ("gpt-6.1-sol", "medium"),
    "support":      ("gpt-6.1-sol", "high"),
    "code_edit":    ("claude-sonnet-5-5", "xhigh"),
}
REVIEW_BY = "2026-11-01"  # re-run Step 2 when prices or models change

Write the review date into the config. Argon's price will double on a date Google hasn't announced. Sonnet 5.5's token counts will shift as Anthropic tunes it. A routing table that never gets re-run slowly turns into a hardcoded model choice.

Where Nadir fits.

The tutorial above is the manual version of what Nadir does in the request path:

If you moved a call site to one of this week's $2/$10 models without setting its effort, start with a free key, tag a week of traffic, and look at cost per task by model and effort. Chart 2 says the answer can move by an order of magnitude.

Conclusion.

Sonnet 5.5, GPT-6.1 Sol, and Gemini 4 Argon all list at $2/$10, and on the first independent run they cost $7.62, $0.72, and $1.99 per task at their top settings. Same price, 10x apart, because the rate card prices tokens and the models spend very different numbers of them. Within one model, the effort dial moved cost 3x on Sol and 13x on Sonnet. Sol is cheapest on this suite and flattest across effort levels; Sonnet 5.5 at max is the strongest of the three and pays for it; Argon scores between them with the lowest hallucination rate, at a price that will double on an unannounced date. When sticker prices converge, the bill is decided by tokens per passing task. Measure that per task type, pick a model and an effort for each, and put a review date on the table.


Per-task costs and scores are Artificial Analysis Intelligence Index readings reported on October 1, 2026, and may change as providers tune their models. Sources: [Trending Topics, "Gemini 4 Matches GPT-6 Astra but Trails Opus 5.5"](https://www.trendingtopics.eu/gemini-4-artificial-analysis-en/). [Aivy, "Sonnet 5.5 vs GPT 6.1 Sol: same price, different bill"](https://aivy.com.au/resources/claude-sonnet-5-5-vs-gpt-6-1-sol/). [Unite.AI, "Anthropic Releases Claude Sonnet 5.5 at Unchanged Sonnet 5 Pricing"](https://www.unite.ai/anthropic-releases-claude-sonnet-5-5-at-unchanged-sonnet-5-pricing/). [Developers Digest, "Claude Sonnet 5.5 Developer Guide"](https://www.developersdigest.tech/blog/claude-sonnet-5-5-release-guide-2026). [The Next Web, "OpenAI releases GPT-6.1 Sol at a fifth of GPT-6 Astra's token prices"](https://thenextweb.com/news/openai-gpt-6-1-sol-price-astra-devday). [9to5Google, "Google announces Gemini 4 Argon"](https://9to5google.com/2026/09/30/gemini-4-argon-announcement/). [DataCamp, "Gemini 4 Argon: Benchmarks, Pricing, and Access"](https://www.datacamp.com/blog/gemini-4-argon).

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

For coding agents, Nadir connects to Codex, Claude Code, or Cursor and recommends a model tier for delegated work. The agent decides whether to hand the task off and uses its own configured models, so Nadir needs no proxy and no provider keys on that path.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.