Save Twice

Every prompt-compression tool on the market saves you the same way: fewer tokens, same model. That is one axis. Your bill has two. You are overpaying on how many tokens you send and on the price of each one. Nadir compresses the context and routes the request to the cheapest capable model in the same call, so the two savings multiply instead of competing. And there is no package to install, no sidecar to run, and no client code to change, because it is built into the endpoint you already point at.

Published 2026-07-05 by Dor Amir on the Nadir blog.

Filed under Nadir & Alternatives.

The bill has two numbers, and most tools only touch one

The cost of an LLM call is simple:

cost  =  tokens  ×  price per token

Every prompt-compression product attacks the left side. LLMLingua drops low-information tokens. Headroom runs a model that rewrites verbose context. Claude Code hook tools compress tool output before the agent reads it. They are good at what they do, and they all do the same thing: fewer tokens, sent to the same model, at the same price per token.

That leaves the right side of the equation untouched. And the right side is where the biggest lever is. An Opus input token costs roughly 19x a Haiku input token. If the request did not need Opus, compressing its context is optimizing the wrong number first.

Routing attacks the right side. Send each request to a cheaper model that clears the workload's measured bar, and the price per token falls. Nadir's older 60% / 98% RouterBench result is a reference-assisted research ceiling on 11,420 triples, not deployed performance or a customer forecast.

Compression and routing are not competing strategies. They are two factors of the same product. Do both and the savings compound.

Saving twice, in one call

cost  =  tokens  ×  price per token
          ▲            ▲
          │            └── routing: send it to a cheaper capable model
          └── compression: strip repeated schemas, minify JSON, drop dead context

Take an input-heavy agentic workload, the kind where context is the bill.

That is the point. 1 minus (0.40 times 0.85) is about 66%, not 60% and not 15%. The second saving lands on the already-reduced bill, not the original one.

There is an ordering detail that matters here, and Nadir gets it right: it compresses first, then ranks models on the compressed cost. A request that was going to be borderline-expensive can become cheap once its context is smaller, and the router sees that. Compression does not just shrink the bill, it changes the routing decision.

Built in, not bolted on

Here is the part that usually gets skipped in the pitch.

To add compression to your stack the normal way, you install a package, or you stand up a sidecar proxy, or you wrap every agent with a hook, and then you keep that dependency current as your providers and SDKs move. It is real work, and it is a second system to operate next to whatever you use for model selection.

Nadir does not ask for any of that. Point your OpenAI-compatible calls at Nadir with model=auto. Compression runs inside the request path, before the model is chosen. There is:

The compression is Nadir's own, developed against real routing traffic, not a third-party library you install and keep current. The safe transforms are deterministic and switched on per account, key, or request. They minify embedded JSON, deduplicate repeated tool schemas and system prompts across a request, and normalize whitespace outside fenced code blocks while keeping indentation. Nothing is paraphrased, nothing is guessed, and past 40 user turns the middle of the conversation is trimmed. For heavier context, an aggressive mode adds semantic deduplication, replacing a near-duplicate message with a short reference plus only the phrases that actually changed. Same methods, one endpoint, no extra dependency.

An honest word on the number

We are not going to print one big compression percentage on a banner, because the honest answer depends on your traffic, and we would rather you trust the small number than distrust a large one.

The compressor is not the ceiling. Your traffic is. The more bloated the context, the more the first saving is worth, and either way the second saving from routing still stacks on top.

The takeaway

Compression alone saves you once, on how many tokens you send. Routing alone saves you once, on the price of each token. Run them together and the two savings multiply, on every call, with the compressed size feeding the routing decision.

You do not need another package to get there. It is the endpoint.

Point `model=auto` at Nadir and pay for neither too many tokens nor too expensive a model.

More on nadir & alternatives

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

For coding agents, Nadir connects to Codex, Claude Code, or Cursor and recommends a model tier for delegated work. The agent decides whether to hand the task off and uses its own configured models, so Nadir needs no proxy and no provider keys on that path.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.