Same Weights, Different Model

Llama 3.3 70B costs $0.10 per million input tokens on one host and $1.04 on another: same open weights, 10.4x apart. The answers differ too. The same gpt-oss-120b scored from 36.7% to 93.3% on AIME25 depending on who served it, two hosts gained 7 to 13 points in a week with no model change, and Kimi K2's tool-call schema accuracy ran from 73% to 100% across deployments. On a 20-step agent, per-call accuracy below about 89% erases a 10x price advantage. Why identical weights diverge, how per-call errors compound, and a Python tutorial that qualifies each model-and-host pair by cost per passing task before it gets production traffic.

Published 2026-10-03 by Dor Amir on the Nadir blog.

Filed under Pricing & Models.

TL;DR.

Open weights were supposed to make the model a commodity. Pick Llama, Qwen, Kimi, or gpt-oss, then shop for the cheapest host. The shopping part works. In September 2026, the same Llama 3.3 70B Instruct listed at $0.10 per million input tokens on DeepInfra and $1.04 on Together AI, a 10.4x spread for identical weights.

The commodity part doesn't. Two hosts serving the same weights can return different answers, because they don't serve the same deployment: different quantization, different chat templates, different tool-call parsers, different defaults. When Artificial Analysis ran gpt-oss-120b through every major API, the same model scored anywhere from 36.7% to 93.3% on AIME25 depending on who served it. Moonshot's own verifier found Kimi K2 tool-call schema accuracy ranging from 100% down to 73% across deployments.

This post covers where the price spread comes from, why the quality spread exists, why per-call errors compound into much bigger per-run costs on agent workloads, and a tutorial for qualifying hosts by cost per passing task before you route production traffic to the cheapest one.

Prices are from [CloudZero](https://www.cloudzero.com/blog/llm-inference-providers/) (September 22, 2026) and [OpenRouter](https://openrouter.ai/blog/tutorials/how-to-get-the-lowest-cost-llm-inference-on-openrouter/) (updated September 24, 2026). Accuracy figures are from [Artificial Analysis via Simon Willison](https://simonwillison.net/2025/Aug/15/inconsistent-performance/) and [MoonshotAI's K2 Vendor Verifier](https://github.com/MoonshotAI/K2-Vendor-Verifier). The compounding math is illustrative. None of this is Nadir customer data.

One model, six prices.

Chart 1: Llama 3.3 70B Instruct input price per million tokens by host, September 2026. DeepInfra $0.10 on an FP8 Turbo tier, Novita $0.135, Groq $0.59, Amazon Bedrock about $0.72 blended, Fireworks about $0.90, Together AI about $1.04. A 10.4x spread for the same open weights.
Chart 1: Llama 3.3 70B Instruct input price per million tokens by host, September 2026. DeepInfra $0.10 on an FP8 Turbo tier, Novita $0.135, Groq $0.59, Amazon Bedrock about $0.72 blended, Fireworks about $0.90, Together AI about $1.04. A 10.4x spread for the same open weights.
HostInput / output, per MWhat you're paying for
DeepInfra$0.10 / $0.32FP8 "Turbo" tier, thin margins
Novita$0.135 / $0.40Budget tier
Groq$0.59 / $0.79LPU hardware, high throughput
Amazon Bedrock~$0.72 blendedAWS billing, IAM, private networking
Fireworks~$0.90 to $1.20Fine-tuning, deployment features
Together AI~$1.04 / $1.04Full-precision serving, platform bundle

Sources: CloudZero, OpenRouter. An earlier Inferbase audit from May listed Together at about $0.88, so these numbers move. Check before you commit.

None of these hosts is overcharging or undercharging by accident. The spread reflects real differences:

  1. Precision. CloudZero notes that budget tiers "often serve FP8 or FP4 compressed variants under the same model name." Fewer bits per weight means less memory per request, more requests per GPU, and a lower price.
  2. Hardware. Groq and Cerebras run custom silicon. CloudZero reports roughly 478 tokens per second on Groq and about 1,700 sustained on Cerebras for gpt-oss-120b, against GPU hosts that run 2 to 3x slower. Inferbase measured the cheapest Llama 3.3 hosts at 50 to 80 tokens per second.
  3. Load. Two hosts on the same GPUs can quote different prices because one runs fuller. We've covered how utilization sets the price when you host it yourself. The same economics apply to them.
  4. Bundle. Bedrock's price includes AWS procurement, IAM, and VPC endpoints. For some buyers that's worth 7x. For others it's 7x for a model they could reach elsewhere.

So on the rate card, the cheap host wins. The question is whether it still wins per task.

Same weights, different answers.

Chart 2: AIME25 scores for gpt-oss-120b by API host, median of 32 runs, Artificial Analysis, August 2025. Six hosts scored 93.3% in both runs. Parasail scored 90.0%. Groq moved from 86.7% to 93.3% and Azure from 80.0% to 93.3% between the first run and an August 20 re-run. Amazon moved from 83.3% to 80.0%. Google Vertex appeared in the re-run at 83.3%. CompactifAI scored 36.7% in the first run.
Chart 2: AIME25 scores for gpt-oss-120b by API host, median of 32 runs, Artificial Analysis, August 2025. Six hosts scored 93.3% in both runs. Parasail scored 90.0%. Groq moved from 86.7% to 93.3% and Azure from 80.0% to 93.3% between the first run and an August 20 re-run. Amazon moved from 83.3% to 80.0%. Google Vertex appeared in the re-run at 83.3%. CompactifAI scored 36.7% in the first run.

In August 2025, Artificial Analysis ran gpt-oss-120b through every host offering it, repeating GPQA Diamond 16 times, AIME25 32 times, and IFBench 8 times per host. AIME25 showed the widest spread. Six hosts, plus a reference vLLM run, tied at 93.3%. Azure scored 80.0%, Amazon 83.3%, and CompactifAI, which serves a compressed version, 36.7%.

The re-run on August 20 is the part worth remembering. Groq went from 86.7% to 93.3% and Azure from 80.0% to 93.3%. Amazon dropped from 83.3% to 80.0%. Nobody shipped a new model. The hosts changed their deployments, and the score followed, with the same model ID on the request the whole time.

Moonshot saw the same thing with Kimi K2 and built a tool to catch it. Its K2 Vendor Verifier sends about 4,000 tool-calling requests to each deployment and compares the results to Moonshot's own API: did the model call a tool when it should, and did the arguments match the schema? On the November 2025 run:

DeploymentModelSchema accuracy
MoonshotAI (official), Turbo, FireworksK2-thinking100.00%
InfiniAIK2-thinking99.89%
SiliconFlowK2-thinking98.96%
GMICloudK2-thinking95.95%
vLLM (reference self-deploy)K2-thinking87.22%
MoonshotAI, DeepInfra, Fireworks, GroqK2-0905100.00%
SGLang (reference self-deploy)K2-090573.13%

The vLLM and SGLang rows are Moonshot's tests of the open-source serving engines, not commercial hosts, which makes them the most useful rows in the table if you plan to self-host: the default engine configuration was the least accurate deployment tested. The vLLM team later published the debugging work it took to close that gap, most of it in chat templates and tool-call parsing rather than the weights.

In the Hacker News discussion of the verifier's update, users reported Bedrock defects that silently ended 20% to 30% of Kimi tool-calling attempts, and being routed to heavily quantized versions on aggregators. Those are user reports, not measurements. They match what the measured tables show.

Why identical weights diverge.

None of this shows up in the model name.

Per-call errors compound.

Chart 3: Share of agent runs where every tool call is valid, by number of tool calls per run, at K2-thinking schema accuracies from Moonshot's Vendor Verifier. At 20 steps: 99.89% per call gives 97.8% clean runs, 98.96% gives 81.1%, 95.95% gives 43.7%, and 87.22% gives 6.5%, which means 15.4x the tokens per clean run if failed runs are retried from scratch.
Chart 3: Share of agent runs where every tool call is valid, by number of tool calls per run, at K2-thinking schema accuracies from Moonshot's Vendor Verifier. At 20 steps: 99.89% per call gives 97.8% clean runs, 98.96% gives 81.1%, 95.95% gives 43.7%, and 87.22% gives 6.5%, which means 15.4x the tokens per clean run if failed runs are retried from scratch.

A single chat completion that fails 4% of the time costs about 4% more on retries. An agent that makes 20 tool calls doesn't. If each call is valid with probability p and one bad call breaks the run, the run succeeds with probability p²⁰:

Per-call schema accuracyClean 20-step runsTokens per clean run
99.89%97.8%1.02x
98.96%81.1%1.23x
95.95%43.7%2.29x
87.22%6.5%15.4x

The break-even is sharp. Below 89.1% per-call accuracy, a 20-step agent spends more than 10x the tokens per clean run, which erases the full 10.4x price gap in Chart 1. A host that's 4% less accurate on tool calls loses more than half its 20-step runs.

This is a simplified model. Real agents often recover from a bad call, steps aren't independent, and the K2 numbers are from one test harness. But the direction holds, and it's the same mechanic we measured for model downgrades in the retry tax: failure cost scales with the work already done when the failure happens, not with the price of the failing call.

There's a worse case than retries. A run that fails loudly costs tokens. A run that fails quietly, with a wrong answer that looks fine or a conversation that just ends, costs whatever happens downstream. The structured output post covers the cheap defenses: validate every tool call against its schema before you act on it, and treat a parse failure as an error, not an empty result.

Tutorial: qualify hosts by cost per passing task.

The fix is to treat each (model, host) pair as its own model, test it on your traffic, and only route to the ones that pass. This works with any OpenAI-compatible endpoint. The examples use OpenRouter's provider routing because it exposes per-host pinning on one key.

Step 1: build a canary set from your own logs.

Pull two kinds of prompts from production: about 30 that should trigger a tool call, with the tool schemas you actually use, and about 30 with a checkable answer (an extraction with a known value, a classification with a label, a calculation). Sixty prompts is enough to catch a broken deployment. It isn't enough to rank two good ones, so don't over-read small differences.

import json, time, statistics
from jsonschema import validate, ValidationError
from openai import OpenAI

client = OpenAI(base_url="https://openrouter.ai/api/v1", api_key=OPENROUTER_KEY)
MODEL = "moonshotai/kimi-k2-thinking"

# USD per million tokens (input, output) for each host, from the host's pricing page.
HOSTS = {
    "moonshotai": (0.60, 2.50),
    "fireworks":  (0.60, 2.50),
    "deepinfra":  (0.50, 2.00),
    "novita":     (0.48, 2.00),
}

canaries = [json.loads(line) for line in open("canaries.jsonl")]
# {"messages": [...], "tools": [...], "expect_tool": true, "check": null}
# {"messages": [...], "tools": null, "expect_tool": false, "check": "42"}

Fill HOSTS from each host's current pricing page on the day you run this. The values above are placeholders, not quotes.

Step 2: run every canary against every host, pinned.

def run(host, c):
    t0 = time.monotonic()
    r = client.chat.completions.create(
        model=MODEL,
        messages=c["messages"],
        tools=c["tools"] or None,
        temperature=0.6,          # set it; never inherit a host default
        max_tokens=4096,
        extra_body={"provider": {"only": [host], "allow_fallbacks": False}},
    )
    latency = time.monotonic() - t0
    msg = r.choices[0].message
    p_in, p_out = HOSTS[host]
    cost = (r.usage.prompt_tokens * p_in + r.usage.completion_tokens * p_out) / 1e6
    return {"host": host, "cost": cost, "latency": latency,
            "passed": passes(c, msg, r.choices[0].finish_reason)}

def passes(c, msg, finish_reason):
    if c["expect_tool"]:
        if finish_reason != "tool_calls" or not msg.tool_calls:
            return False                      # should have called a tool
        schemas = {t["function"]["name"]: t["function"]["parameters"] for t in c["tools"]}
        for call in msg.tool_calls:
            try:
                validate(json.loads(call.function.arguments), schemas[call.function.name])
            except (KeyError, ValueError, ValidationError):
                return False                  # unknown tool, bad JSON, or wrong shape
        return True
    if msg.tool_calls:
        return False                          # called a tool it shouldn't have
    return c["check"] in (msg.content or "")

rows = [run(h, c) for h in HOSTS for c in canaries]

allow_fallbacks: False matters. Without it, a failed request can be served by a different host and you'll score the wrong deployment.

Step 3: score cost per passing task, then project it to your agent length.

STEPS = 20  # median tool calls per run in your agent

for host in HOSTS:
    r = [x for x in rows if x["host"] == host]
    pass_rate = sum(x["passed"] for x in r) / len(r)
    cost_per_pass = sum(x["cost"] for x in r) / max(1, sum(x["passed"] for x in r))
    run_success = pass_rate ** STEPS
    p50 = statistics.median(x["latency"] for x in r)
    print(f"{host:12s} pass {pass_rate:6.1%}  $/pass {cost_per_pass:.5f}  "
          f"{STEPS}-step clean {run_success:6.1%}  p50 {p50:.1f}s")

Rank by the projected run column, not the per-call one. A host that's 20% cheaper per call and 3 points less accurate is usually the expensive choice for an agent and the cheap choice for single-shot classification. Run the scoring separately per workload, because the answer differs.

Step 4: pin the winners, and block what failed.

APPROVED = ["fireworks", "moonshotai"]   # passed Step 3 for this workload

extra_body = {
    "provider": {
        "order": APPROVED,
        "allow_fallbacks": False,         # never fall through to an untested host
        "quantizations": ["fp8", "bf16"], # refuse fp4 and below
    }
}

If you'd rather pay the cheapest approved price than a fixed order, use "only": APPROVED with "sort": "price". Either way, failover goes to a host you've tested, not to whichever one is cheapest that minute.

Step 5: re-run the canaries on a schedule.

The gpt-oss scores changed in a week with no model change. Run the 60 canaries nightly against each approved host, store the pass rate, and alert when one drops more than two points below its trailing average. It costs cents a night, and it's the only way you'll find out a host changed its deployment before your users do.

Where Nadir fits.

The tutorial picks where a model runs. Nadir decides which model a prompt needs, and the two decisions stack:

If you moved a workload to an open-weight model to save money, start with a free key and tag a week of traffic by host. If the cheapest host's cost per passing task is higher than its price suggests, you'll see it.

Conclusion.

Open weights made models portable. They didn't make deployments identical. In September 2026, Llama 3.3 70B ranged from $0.10 to $1.04 per million input tokens across six hosts, and the same gpt-oss-120b scored from 36.7% to 93.3% on AIME25 depending on who served it, with two hosts gaining 7 to 13 points in a week without a model change. Kimi K2's tool-call schema accuracy ran from 73% to 100% across deployments. On a single call, a small accuracy gap is a small cost. On a 20-step agent, per-call accuracy below about 89% erases a 10x price advantage. Treat each model-and-host pair as its own model: qualify it on your own prompts, pin the ones that pass, block fallbacks to the ones you haven't tested, and re-check them every night.


Sources: [CloudZero, "Best LLM inference providers 2026: 16+ on cost per outcome"](https://www.cloudzero.com/blog/llm-inference-providers/), September 22, 2026. [OpenRouter, "Lowest-Cost LLM Inference: The Complete OpenRouter Guide"](https://openrouter.ai/blog/tutorials/how-to-get-the-lowest-cost-llm-inference-on-openrouter/), updated September 24, 2026. [Inferbase, "The Real Cost of Inference at Enterprise Scale: A 2026 Pricing Audit"](https://inferbase.ai/blog/enterprise-llm-inference-pricing-2026), May 13, 2026. [Simon Willison, "Open weight LLMs exhibit inconsistent performance across providers"](https://simonwillison.net/2025/Aug/15/inconsistent-performance/), August 15, 2025, with Artificial Analysis data updated August 20, 2025. [MoonshotAI, K2 Vendor Verifier](https://github.com/MoonshotAI/K2-Vendor-Verifier), November 15, 2025 run. [vLLM Blog, "Chasing 100% Accuracy: A Deep Dive into Debugging Kimi K2's Tool-Calling on vLLM"](https://blog.vllm.ai/2025/10/28/Kimi-K2-Accuracy.html), October 28, 2025. [Hacker News, "Kimi vendor verifier"](https://news.ycombinator.com/item?id=47838703).

More on pricing & models

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

For coding agents, Nadir connects to Codex, Claude Code, or Cursor and recommends a model tier for delegated work. The agent decides whether to hand the task off and uses its own configured models, so Nadir needs no proxy and no provider keys on that path.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.