Cheaper on Paper

A study of 190,000 API calls found that the reasoning model with the lower list price cost more in total in 32% of model pairs, by up to 28x. Gemini 3 Flash was listed 80% cheaper than GPT-5.4 and cost 38% more, because it spent about 7x the tokens thinking. The same week, a survey of 396 enterprises found only 11% can forecast AI spend within 10%. The two findings are one problem: budgets are built on price per token, and the model decides how many tokens a task takes. What the price reversal paper measured, why the 9.7x run-to-run variance averages out at production volume while a wrong price assumption never does, and a Python harness that ranks models by cost per solved task and forecasts the bill as P10, P50 and P90.

Published 2026-10-06 by Dor Amir on the Nadir blog.

Filed under Pricing & Models.

TL;DR.

On October 5 the Wall Street Journal put two numbers side by side, and they belong together.

The first is from a survey. Benchmarkit and Mavvrik asked 396 enterprises about their AI spend for the State of AI Cost Governance 2026 report. Only 11% could forecast it within 10%, down from 15% a year earlier. 89% missed by more than that, and 49% had to reprice an AI-powered product because of costs they didn't see coming.

The second is from a paper. Researchers at Stanford, Carnegie Mellon, Berkeley and Microsoft Research ran eight reasoning models on 6,877 tasks and three agentic benchmarks, about 190,000 API calls in all. In The Price Reversal Phenomenon, the model with the lower list price cost more in total in 32% of model pairs. Gemini 3 Flash was listed 80% cheaper than GPT-5.4 and cost 38% more across the tasks.

The second number explains much of the first. Most AI budgets are built as tokens × price per token, and the token count is the part nobody controls. The model decides how long to think.

This post covers what the paper measured, why per-query noise is not your forecasting problem, and a Python harness that prices models the way the paper did: per task, several runs each, on your own prompts. It ends with a forecast you can hand to finance.

Study and survey figures are as published by their authors. Model prices in the tutorial are list prices as of October 6, 2026. The simulation in chart 3 is illustrative, not Nadir customer data.

What the paper measured.

The authors took eight reasoning models at their May 1, 2026 list prices and summed input and output price per million tokens into a single listed price:

ModelListed price (input + output, per M)
Claude Opus 4.7$30.00
GPT-5.4$17.50
Gemini 3.1 Pro$14.00
Claude Haiku 4.5$6.00
GPT-5.4 Mini$5.25
Kimi K2.6$4.95
Gemini 3 Flash$3.50
MiniMax-M2.7$1.50

Then they ran every model on nine single-turn datasets (math, science, code and QA, 6,877 tasks) and three multi-turn agentic benchmarks, generating 7.39 billion tokens, and compared what each model actually cost. Eight models make 28 pairs per task, 336 pairs in total. In 106 of them the cheaper-listed model was the more expensive one.

Chart: GPT-5.4 vs Gemini 3 Flash indexed to GPT-5.4 = 100. Listed price: GPT-5.4 100, Gemini 3 Flash 20, 80% cheaper. Actual cost across the study's tasks: GPT-5.4 100, Gemini 3 Flash 138, 38% more.
Chart: GPT-5.4 vs Gemini 3 Flash indexed to GPT-5.4 = 100. Listed price: GPT-5.4 100, Gemini 3 Flash 20, 80% cheaper. Actual cost across the study's tasks: GPT-5.4 100, Gemini 3 Flash 138, 38% more.

The headline pair is easy to work through. Flash's listed price is 20% of GPT-5.4's. Its actual cost was 138% of GPT-5.4's. For both to be true, Flash had to spend roughly 1.38 / 0.20 ≈ 7x as many tokens on the same work. Our arithmetic, not the paper's, and approximate because input and output tokens are priced differently, but the order of magnitude is right. On one MMLU-Pro question the paper reports Flash using more than 60,000 thinking tokens where GPT-5.4 used 25.

Three findings matter for anyone who owns a bill:

Chart: Share of model pairs where the lower listed price had the higher total cost. ArenaHard 11%. All tasks 32%, 106 of 336 pairs. MMLU-Pro 57%. The worst single reversal was 28x.
Chart: Share of model pairs where the lower listed price had the higher total cost. ArenaHard 11%. All tasks 32%, 106 of 336 pairs. MMLU-Pro 57%. The worst single reversal was 28x.

The lead author, Lingjiao Chen, put the conclusion in one line: "Price alone should not be used to infer which model is actually cheaper." The paper leaves predicting per-query cost as an open problem, and argues the run-to-run variance sets a noise floor no predictor can get under.

Why 89% of forecasts miss.

Read the survey and the paper together and the forecasting failure has three layers.

Layer 1: the unit is wrong. A budget built on price per token is a budget built on the one number the paper shows is unreliable. What you pay for is a finished task, and tokens per task vary by model, by task type and by run. If you swapped to a "cheaper" model to save money and the bill went up, this is why.

Layer 2: the mix moves. A team that budgets on GPT-5.4 and later moves half of its traffic to Flash has changed the cost per task without changing a line in the budget. So has a team whose users start asking harder questions, or whose agent gained three new tools. Agents make the mix move faster.

Layer 3: nobody owns the variable. The survey found 62% of organizations had a cost surprise change a business decision in the past year, 40% escalated to the board, 33% froze spending and 25% delayed or cancelled an AI project. Spending freezes are what happens when the only lever left is the cap.

There is one thing the 9.7x does not cause, and it matters for how you fix this.

Noise averages out. Assumptions don't.

Chart: Spend as a multiple of plan, log scale. One task: 0.22x to 2.65x. 100 tasks: 0.86x to 1.15x. 10,000 tasks: within 1.5%. Budgeted on list price, Gemini 3 Flash: 6.9x.
Chart: Spend as a multiple of plan, log scale. One task: 0.22x to 2.65x. 100 tasks: 0.86x to 1.15x. 10,000 tasks: within 1.5%. Budgeted on list price, Gemini 3 Flash: 6.9x.

We simulated per-task cost with a spread that reproduces the paper's run-to-run range (a lognormal with σ = 0.75, where the median max-to-min ratio across ten runs comes out at about 9.6x). Then we asked how far the average lands from the true mean:

VolumeP5 to P95 of actual vs expected
1 task0.22x to 2.65x
100 tasks0.86x to 1.15x
10,000 tasks0.985x to 1.015x

At production volume, per-query variance disappears into the average. A 9.7x spread on one prompt becomes a 1.5% band on a week of traffic. That's good news for forecasting: the noise floor the paper describes is a problem for a pilot of 50 prompts, not for a monthly budget.

What doesn't average out is a wrong assumption about the mean. If you priced Flash at its list rate relative to GPT-5.4 and it delivered the study's result, your forecast is off by 6.9x at every volume, and more traffic only makes the miss bigger in dollars. The same holds for a mix shift or a new model. So the fix is not a better statistical model of tokens. It's measuring cost per task on your own prompts before you commit, and re-measuring when the mix changes.

Tutorial: measure cost per task, then forecast.

This harness does what the paper did, on your traffic: several runs per prompt per model, priced from the provider's own usage counts, scored for correctness. It needs an OpenAI-compatible endpoint and a sample of 50 to 200 real prompts with a way to check each answer.

Step 1: run each model several times per task.

import itertools, math, random, statistics
from openai import OpenAI

client = OpenAI()  # any OpenAI-compatible endpoint

# USD per million tokens (input, output), October 2026 list prices.
# Reasoning tokens are billed as output. Edit to your contract rates.
PRICES = {
    "gpt-6.1-sol": (2.00, 10.00),
    "sonnet-5.5":  (2.00, 10.00),
    "haiku-4.5":   (1.00, 5.00),
    "gpt-6-luna":  (0.10, 0.50),
}

def run_once(model, prompt):
    r = client.chat.completions.create(
        model=model, messages=[{"role": "user", "content": prompt}]
    )
    u = r.usage
    details = getattr(u, "completion_tokens_details", None)
    reasoning = (getattr(details, "reasoning_tokens", 0) or 0) if details else 0
    pin, pout = PRICES[model]
    # completion_tokens already includes reasoning tokens on OpenAI-style usage
    usd = (u.prompt_tokens * pin + u.completion_tokens * pout) / 1e6
    return usd, reasoning, r.choices[0].message.content

def measure(models, tasks, runs=5):
    """tasks: list of {"id": str, "prompt": str, "check": callable(answer) -> bool}"""
    rows = []
    for t, m in itertools.product(tasks, models):
        for k in range(runs):
            usd, reasoning, answer = run_once(m, t["prompt"])
            rows.append({"task": t["id"], "model": m, "run": k, "usd": usd,
                         "reasoning": reasoning, "ok": bool(t["check"](answer))})
    return rows

Five runs per prompt is the minimum to see the spread. With 100 prompts and four models that's 2,000 calls, a few dollars on most rate cards.

Step 2: rank by cost per solved task, not by list price.

def report(rows):
    out = {}
    for m in {r["model"] for r in rows}:
        mr = [r for r in rows if r["model"] == m]
        costs = sorted(r["usd"] for r in mr)
        solved = sum(r["ok"] for r in mr)
        by_task = {}
        for r in mr:
            by_task.setdefault(r["task"], []).append(r["usd"])
        spreads = [max(c) / min(c) for c in by_task.values() if min(c) > 0]
        out[m] = {
            "listed": sum(PRICES[m]),                      # the paper's input + output
            "mean_per_task": statistics.fmean(costs),
            "p90_per_task": costs[int(0.9 * (len(costs) - 1))],
            "pass_rate": solved / len(mr),
            "per_solved": sum(costs) / max(solved, 1),
            "median_run_spread": statistics.median(spreads) if spreads else float("nan"),
            "mean_reasoning": statistics.fmean(r["reasoning"] for r in mr),
        }
    return out

def reversals(stats):
    """Pairs where the lower listed price has the higher mean cost per task."""
    found = []
    for a, b in itertools.combinations(stats, 2):
        lo, hi = sorted((a, b), key=lambda m: stats[m]["listed"])
        if stats[lo]["listed"] < stats[hi]["listed"] and \
           stats[lo]["mean_per_task"] > stats[hi]["mean_per_task"]:
            found.append((lo, hi, stats[lo]["mean_per_task"] / stats[hi]["mean_per_task"]))
    return found

stats = report(rows)
for m, s in sorted(stats.items(), key=lambda kv: kv[1]["per_solved"]):
    print(f"{m:12} listed ${s['listed']:>6.2f}  per task ${s['mean_per_task']:.5f}  "
          f"p90 ${s['p90_per_task']:.5f}  pass {s['pass_rate']:.0%}  "
          f"per solved ${s['per_solved']:.5f}  run spread {s['median_run_spread']:.1f}x")
for lo, hi, x in reversals(stats):
    print(f"REVERSAL: {lo} is listed cheaper than {hi} but costs {x:.2f}x per task")

Sort by cost per solved task. A model that's cheap per call and fails a third of the time is not cheap: every failure is paid for once on the cheap model and again on whatever handles the retry. That's the retry tax, and it's the agentic version of the paper's reversal.

Step 3: forecast the bill with a range, not a point.

def forecast(stats, plan, volume_low, volume_high, fallback=None, sims=5000):
    """plan: {model: share of traffic}. fallback: model that retries failures."""
    bills = []
    for _ in range(sims):
        volume = random.uniform(volume_low, volume_high)
        bill = 0.0
        for m, share in plan.items():
            s, n = stats[m], volume * share
            per_task = s["mean_per_task"]
            if fallback and m != fallback:
                per_task += (1 - s["pass_rate"]) * stats[fallback]["mean_per_task"]
            # sum of n noisy tasks: the mean is stable, the volume isn't
            sd = (s["p90_per_task"] - s["mean_per_task"]) / 1.28
            bill += random.gauss(per_task * n, max(sd, 0) * math.sqrt(n))
        bills.append(bill)
    bills.sort()
    return {p: bills[int(p / 100 * (sims - 1))] for p in (10, 50, 90)}

print(forecast(stats, {"haiku-4.5": 0.8, "sonnet-5.5": 0.2},
               volume_low=400_000, volume_high=600_000, fallback="sonnet-5.5"))

Give finance the P10, P50 and P90, and say which assumptions they depend on: the volume range, the traffic mix, and the pass rate that decides how often you pay the fallback. When you run it you'll see what chart 3 shows. The band comes almost entirely from volume and mix, not from per-query noise.

Step 4: re-measure when anything changes.

The forecast is only as good as the last measurement. Re-run steps 1 and 2 when a provider ships a new model version, when you change reasoning effort or thinking budgets, when a prompt template changes, and monthly regardless. Capping reasoning is the most direct lever on the paper's main driver: a max_tokens ceiling or a low effort setting bounds the 60,000-token tail, at some cost in pass rate that step 2 will show you.

Where Nadir fits.

The harness above answers "which model is cheapest per solved task" once, for a sample. Production traffic asks it again on every request, because the answer depends on the prompt. That's the decision Nadir makes. Send model="auto" and a trained classifier picks a tier per prompt against the quality floor you set on the API key. Each response carries the routing decision and its cost in the metadata, so actual cost per task is a number you read, not one you estimate. If you only want the decision, POST /v1/bucket returns simple, medium or complex with a probability for each, and you call your own providers.

For forecasting, start in observe mode. It records what routing would have chosen on your real traffic without changing a single response, and that gives you a measured mix and cost per request to feed step 3. Start with a free key, or read how per-team attribution turns the same data into chargeback.

Conclusion.

Only 11% of enterprises can forecast AI spend within 10%, and a study of 190,000 API calls shows why the usual method fails: in 32% of model pairs the cheaper-listed reasoning model cost more, because it spent more tokens thinking. Price per token is the wrong unit. The good news is in the statistics. Per-query variance of 9.7x shrinks to about 1.5% over 10,000 tasks, so the noise is not what breaks forecasts. Wrong assumptions about the mean are, and those you can measure. Price models per solved task on your own prompts, forecast with a range driven by volume and mix, and re-measure when the mix moves.


Sources: Lingjiao Chen et al., ["The Price Reversal Phenomenon"](https://arxiv.org/html/2603.23971v2), arXiv 2603.23971, v2 May 28, 2026. [The Wall Street Journal](https://www.wsj.com/tech/personal-tech/ai-token-spending-businesses-431ee94a), October 5, 2026, as summarized by [PYMNTS](https://www.pymnts.com/news/artificial-intelligence/2026/businesses-finding-it-harder-to-predict-ai-spending/) and [Implicator](https://www.implicator.ai/cheaper-ai-models-cost-more-price-reversal-study/). Benchmarkit and Mavvrik, [State of AI Cost Governance 2026](https://www.mavvrik.ai/blog/blog-ai-cost-governance-report-2026/) (396 enterprises, April to May 2026), with the escalation, freeze and repricing figures via [Channel Insider](https://www.channelinsider.com/ai/mavvrik-state-of-ai-cost-governance/). Chart 3 rows 1 to 3 are a lognormal simulation (σ = 0.75) calibrated to the paper's 9.7x run-to-run spread. Tutorial prices are list prices as of October 6, 2026.

More on pricing & models

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

For coding agents, Nadir Auto is a terminal launcher for Claude Code or Codex that asks Nadir for the main-session model before each turn, with your configured model as the fallback and cost ceiling. Delegation integrations for Claude Code, Codex, or Cursor instead recommend a model tier for subagent work, and the agent decides. Either way inference runs on your own provider account, so Nadir holds no provider keys on that path.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.