Faster Isn't Cheaper

Together AI closed an $800M round at $8.3B on July 1. Groq raised $650M in June after NVIDIA licensed its chip architecture for $20B. Fireworks is reportedly talking to investors at $15B, up 4x since October. All three are selling the same output token, faster — using speculative decoding, the technique that lets a small draft model propose tokens a large model verifies in one pass. It's mathematically lossless and it is not a line item on your invoice. Here's what it actually buys you, and what does move your bill.

Published 2026-07-05 by Dor Amir on the Nadir blog.

Filed under Pricing & Models.

The industry just spent billions proving something that won't show up on your invoice.

On July 1, 2026, Together AI closed an $800 million Series C at an $8.3 billion valuation, with open-source model inference reportedly crossing a $1 billion industry revenue run rate as growth in closed-model API spend stalls. Source: Tech Times, "Together AI Raises $800M: Open-Source Inference Breaks $1B as Closed Models Stall," July 3, 2026. Three weeks earlier, Groq — fresh off licensing its Language Processing Unit chip architecture to NVIDIA for a reported $20 billion in December 2025 — raised $650 million to rebuild itself as an inference cloud. Source: TechCrunch, "AI chipmaker Groq confirms $650M raise, re-staffs after Nvidia's $20B not-acqui-hire deal," June 22, 2026; Source: CNBC, "Nvidia buying AI chip startup Groq's assets for about $20 billion," December 24, 2025. And as of late May, Fireworks AI — a $4 billion company in October 2025 — was reportedly in talks to raise at $15 billion, nearly quadrupling in seven months. Source: Bloomberg, "Fireworks AI in Talks for Funding at $15 Billion Valuation," May 27, 2026.

Three of the largest rounds of the year, three companies selling the exact same output token — just delivered faster. The engineering underneath all three is the same family of technique: speculative decoding. It is real, it is mathematically lossless, and on a hosted API it is almost entirely invisible to what you're billed. Here's the mechanism, where the savings actually land, and what does move your bill.

What speculative decoding actually does.

The core idea predates the current funding wave by three years: a small, fast "draft" model proposes several tokens ahead, and the large "target" model checks all of them in a single forward pass instead of one token at a time. Accepted draft tokens are kept for free; the first rejected one gets resampled from the corrected distribution. Source: Leviathan, Kalman, Matias, "Fast Inference from Transformers via Speculative Decoding," Google Research, arXiv:2211.17192, ICML 2023. The output distribution is mathematically identical to running the target model alone, token by token — this is not a quality-for-speed trade. It's closer to a hardware-utilization bug fix: autoregressive decoding is memory-bandwidth-bound, not compute-bound, so most of a GPU's math throughput sits idle waiting on the next token. Verifying several draft tokens in one pass uses that idle capacity instead of wasting it.

Three production variants dominate 2026 deployments:

Together AI's ATLAS, live on its platform since October 2025, adds a further layer: two speculators running simultaneously — a heavyweight one trained on broad traffic for a reliable floor, and a lightweight one that keeps retraining on your live traffic pattern in real time, arbitrated by a confidence-aware controller. Together's own benchmark: DeepSeek-V3.1 goes from 105 tokens/second on an FP8 baseline to about 500 — roughly 4x — and the system runs 2.65x faster than standard decoding on average, ahead of dedicated inference silicon in their internal comparison. Source: Together AI, "AdapTive-LeArning Speculator System (ATLAS)," October 10, 2025; Source: VentureBeat, "Together AI's ATLAS adaptive speculator delivers 400% inference speedup by learning from workloads in real-time".

Speculative decoding cuts forward passes roughly 3-4x. It does not cut how many output tokens you're billed for: standard autoregressive decoding runs 1.00 large-model forward pass per output token, speculative decoding (EAGLE-3 / ATLAS-class adaptive speculator) runs about 0.30
Speculative decoding cuts forward passes roughly 3-4x. It does not cut how many output tokens you're billed for: standard autoregressive decoding runs 1.00 large-model forward pass per output token, speculative decoding (EAGLE-3 / ATLAS-class adaptive speculator) runs about 0.30

The catch the funding announcements don't spell out.

Every number above describes forward passes, wall-clock latency, and GPU throughput — not output tokens. A response that's 800 tokens long is 800 tokens long whether the provider generated it with 800 sequential forward passes of the large model (standard decoding) or roughly 200-350 passes because most draft tokens were accepted (speculative decoding). What a hosted API bills you is the 800 output tokens, at whatever the published $/MTok rate is. The forward-pass count is an internal serving-efficiency number that lives entirely on the provider's side of the API boundary; it is not a parameter you set, and it does not appear anywhere on your invoice.

That's exactly why the technique commands billion-dollar funding rounds without ever becoming a line item a customer requests: it's the provider's unit-economics problem to solve, and it shows up to you, if it shows up at all, in one of two indirect ways — a faster response at the same price, or a lower price if competition forces the provider to pass the GPU-hour savings through.

Where the savings actually land.

If you call a hosted API, you pay the rate card, full stop. Whether the provider behind that endpoint used speculative decoding, MTP, or a warehouse of GPUs running at a loss doesn't change what a given response costs you. What it changes is the provider's floor: a team running EAGLE-3 or ATLAS can profitably quote a lower rate than one running naive decoding on the same hardware, which is a large part of why open-weight model hosting turned into an outright price war in 2026. Fireworks reportedly prices DeepSeek V4 Pro under Together on a per-token basis, while Groq is cheaper than Fireworks on the majority of shared models. Source: Morph, "Fireworks vs Together AI Pricing (2026): Per-Token and GPU Rates Compared". Speculative decoding is a meaningful part of what lets these providers sustain that price war without losing money on every call — and it's part of the same dynamic behind DeepSeek's own aggressive pricing moves.

If you self-host an open-weight model — Llama, DeepSeek, Qwen, on your own or rented GPUs via vLLM, SGLang, or TensorRT-LLM — speculative decoding is a direct cost lever, because you're the one paying the GPU-hour bill, not a per-token rate card. Cutting large-model forward passes 3-4x on the same hardware means the same GPU fleet serves 3-4x more requests per hour, or the same request volume needs a proportionally smaller fleet. Industry estimates put the resulting inference-cost reduction from EAGLE-3, Medusa, and MTP-class techniques at roughly 47% against a naive decoding baseline. Source: SyncSoft.AI, "Speculative Decoding 2026: EAGLE-3, Medusa, DeepSeek MTP".

Self-hosting an open-weight model: speculative decoding cuts your GPU bill directly, a roughly 48% drop on the same hardware. Naive vLLM decoding costs $0.42 per 1M output tokens; vLLM with speculative decoding (EAGLE-3 draft model) costs $0.22
Self-hosting an open-weight model: speculative decoding cuts your GPU bill directly, a roughly 48% drop on the same hardware. Naive vLLM decoding costs $0.42 per 1M output tokens; vLLM with speculative decoding (EAGLE-3 draft model) costs $0.22

Turning it on is a serving-config change, not a training run — draft models for the popular open-weight base models already ship on the Hub:

vllm serve meta-llama/Llama-3.1-70B-Instruct \
  --tensor-parallel-size 4 \
  --speculative-config '{
    "method": "eagle3",
    "model": "yuhuili/EAGLE3-LLaMA3.1-Instruct-70B",
    "num_speculative_tokens": 5
  }'

Source: vLLM documentation, "Speculative Decoding"; Source: vLLM documentation, "EAGLE Draft Models". Four lines of server config. The output is byte-for-byte identical to running the target model alone — there is no accuracy given up for the speed.

What actually moves your bill on a hosted API.

LeverWho controls itShows up on your bill?
Speculative decoding, MTP, custom kernelsThe provider, server-sideNo — it's serving-efficiency, invisible to the customer
Provider's published $/MTok rateThe provider, but shaped by the price war aboveDirectly — this is the number you're actually billed
Which model handles a given requestYouDirectly — the single largest lever you control
Which provider serves a given modelYouDirectly — same model, different rate card, different day

Because the decoding trick is invisible to you as a customer, the lever you actually control is which model, and which provider, handles each request — and the speed race above is making the correct answer change faster than most routing setups can track. A rate card that's cheapest today can be undercut within weeks by a competitor whose adaptive speculator just got faster on traffic that looks like yours; that's the literal pitch behind ATLAS's "learns from live traffic" design. Hardcoding one vendor's endpoint into your codebase locks in whichever price war happened to be running when your team wrote the integration — the same static-default mistake that shows up whenever teams pick one model for every job.

This is the layer Nadir sits at: point traffic at model=auto and every request gets priced and routed against the current state of that competitive market — DeepSeek through Fireworks this week, Llama through Groq the next — instead of whichever integration a team shipped before the last price cut. You get the benefit of the infrastructure arms race described above without re-benchmarking three inference providers on a calendar reminder.

Checklist before you assume "faster" means "cheaper."

Conclusion.

Speculative decoding is real, it's mathematically lossless, and the money chasing it — Together's $800M, Groq's $650M, NVIDIA's reported $20B license, Fireworks' reported run at a $15B valuation — is not hype. It's several of the best-capitalized companies in AI infrastructure racing to own the same three-to-four-x. None of that changes what you owe for an output token on a hosted API today. It changes how fast that token arrives, and how much room your provider has to compete on price tomorrow. If your workload runs on someone else's endpoint, the decoding trick underneath it was never yours to configure — but which provider it points you toward, this week, still is.


Data in charts is illustrative, modeled from the public research, documentation, and reporting cited throughout, not derived from proprietary production traces. Sources: [Tech Times, "Together AI Raises $800M," July 3, 2026](https://www.techtimes.com/articles/319657/20260703/together-ai-raises-800m-open-source-inference-breaks-1b-closed-models-stall.htm). [TechCrunch, "AI chipmaker Groq confirms $650M raise," June 22, 2026](https://techcrunch.com/2026/06/22/ai-chipmaker-groq-confirms-650m-raise-re-staffs-after-nvidias-20b-not-acqui-hire-deal/). [CNBC, "Nvidia buying AI chip startup Groq's assets for about $20 billion," December 24, 2025](https://www.cnbc.com/2025/12/24/nvidia-buying-ai-chip-startup-groq-for-about-20-billion-biggest-deal.html). [Bloomberg, "Fireworks AI in Talks for Funding at $15 Billion Valuation," May 27, 2026](https://www.bloomberg.com/news/articles/2026-05-27/fireworks-ai-in-talks-for-funding-at-15-billion-valuation). [Leviathan, Kalman, Matias, "Fast Inference from Transformers via Speculative Decoding," arXiv:2211.17192](https://arxiv.org/abs/2211.17192). [Cai et al., "Medusa," arXiv:2401.10774](https://arxiv.org/abs/2401.10774). [Li et al., "EAGLE-3," arXiv:2503.01840](https://arxiv.org/abs/2503.01840). [DeepSeek-AI, "DeepSeek-V3 Technical Report," arXiv:2412.19437](https://arxiv.org/abs/2412.19437). [Together AI, "ATLAS," October 10, 2025](https://www.together.ai/blog/adaptive-learning-speculator-system-atlas). [vLLM documentation, "Speculative Decoding"](https://docs.vllm.ai/en/latest/features/spec_decode/).

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

For coding agents, Nadir connects to Codex, Claude Code, or Cursor and recommends a model tier for delegated work. The agent decides whether to hand the task off and uses its own configured models, so Nadir needs no proxy and no provider keys on that path.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.