FrugalGPT, Shipped

Stanford's FrugalGPT (TMLR 2024) demonstrated up to 98% cost reduction with a cascade: cheap model first, escalate to frontier only when necessary. Microsoft's BEST-Route (ICML 2025) refined it by sampling the cheap model multiple times before escalating, reducing unnecessary escalations by an additional 15-30%. ETH Zurich proved theoretically that combining routing with cascading always dominates either approach alone. This post walks through the implementation from a simple cascade to the full unified architecture.

Published 2026-06-29, updated 2026-09-24, by Dor Amir on the Nadir blog.

Filed under Routing & Cascades.

98% cost reduction from LLM cascades is achievable. Here is how to build it.

FrugalGPT, published by Stanford researchers in TMLR 2024, demonstrated up to 98% cost reduction by running queries through a cascade: cheap model first, escalate to frontier only when necessary. Three papers since have refined the approach. None require machine learning expertise to implement.

This is the implementation guide.


The core idea.

A cascade router has three components:

  1. A cheap model that attempts every query
  2. A quality estimator that evaluates the response
  3. An escalation threshold that decides when to retry with a frontier model

The cheap model handles 70-90% of queries in most enterprise workloads. The frontier model only sees the hard ones. The cost math is straightforward: if 80% of queries route to a $0.10/M model and 20% to a $5/M model, your blended rate is $1.08/M instead of $5/M, a 78% reduction, provided the cheap model's answers on the 80% it keeps are good enough. Enforcing that is the escalation threshold's job.

FrugalGPT found model combinations achieving 98% cost reduction at matched quality because the quality estimator was trained on the specific domain. Untrained implementations typically achieve 60-80%.


Implementation: basic cascade.

CHEAP_MODEL = "gemini-2.0-flash"      # $0.10/M input
FRONTIER_MODEL = "claude-opus-4-8"    # $5/M input
QUALITY_THRESHOLD = 0.75

def estimate_quality(response: str, query: str) -> float:
    """Heuristic quality scorer. Replace with trained classifier in production."""
    score = 0.7  # base
    penalties = ["i don't know", "i cannot", "i'm not sure"]
    bonuses = ["specifically", "in particular", "according to"]
    for phrase in penalties:
        if phrase in response.lower(): score -= 0.2
    for phrase in bonuses:
        if phrase in response.lower(): score += 0.05
    length_factor = min(len(response) / 500, 1.0) * 0.15
    return min(max(score + length_factor, 0.0), 1.0)

def cascade_query(query: str) -> tuple[str, str]:
    """Returns (response, model_used)."""
    cheap_response = call_model(CHEAP_MODEL, query)
    quality = estimate_quality(cheap_response, query)

    if quality >= QUALITY_THRESHOLD:
        return cheap_response, CHEAP_MODEL

    # Escalate to frontier
    return call_model(FRONTIER_MODEL, query), FRONTIER_MODEL

The BEST-Route improvement.

Microsoft's BEST-Route (arXiv:2506.22716, ICML 2025) found a flaw in naive cascades: escalating after one failure discards information from the cheap model's output.

BEST-Route samples the cheap model multiple times with varied temperatures, then checks for consensus before deciding whether to escalate. When multiple samples agree, the answer is likely correct even if a single sample looked uncertain. When samples diverge, escalation is warranted.

def best_route_query(query: str, n_samples: int = 3) -> tuple[str, str]:
    """Sample cheap model N times. Escalate only on genuine disagreement."""
    samples = []
    for temp in [0.2, 0.5, 0.8]:
        samples.append(call_model(CHEAP_MODEL, query, temperature=temp))

    quality_scores = [estimate_quality(r, query) for r in samples]
    avg_quality = sum(quality_scores) / len(quality_scores)
    score_variance = max(quality_scores) - min(quality_scores)

    # Consensus: high average quality AND low variance between samples
    if avg_quality >= QUALITY_THRESHOLD and score_variance < 0.2:
        best = samples[quality_scores.index(max(quality_scores))]
        return best, CHEAP_MODEL

    # Genuine uncertainty — escalate
    return call_model(FRONTIER_MODEL, query), FRONTIER_MODEL

BEST-Route reduces frontier escalations by an additional 15-30% compared to single-sample cascades. If a naive cascade escalates 25% of queries, BEST-Route brings that to 17-21%.


The ETH Zurich finding: cascade always beats pure routing.

An ICML 2025 paper from ETH Zurich (arXiv:2410.10347) proved theoretically that a cascade router always dominates pure routing or pure cascading when the quality estimator is calibrated.

Pure routing picks one model per query. Pure cascading tries cheap first, then frontier. The unified approach combines both: route clearly easy queries to cheap models directly, use cascades for medium-difficulty queries, and send genuinely hard queries straight to the frontier.

def unified_router(query: str, difficulty_score: float) -> tuple[str, str]:
    """
    Unified routing + cascading.
    difficulty_score: 0.0 = trivial, 1.0 = frontier required.
    Use a lightweight classifier to produce this score.
    """
    if difficulty_score < 0.3:
        # Easy: direct cheap model, skip cascade overhead
        return call_model(CHEAP_MODEL, query), CHEAP_MODEL

    elif difficulty_score < 0.7:
        # Medium: BEST-Route cascade
        return best_route_query(query)

    else:
        # Hard: direct frontier, escalation is certain
        return call_model(FRONTIER_MODEL, query), FRONTIER_MODEL

The difficulty_score comes from a lightweight classifier (logistic regression on query length, keyword features, domain) — not a large model. The classifier itself costs negligible compute.


Realistic cost outcomes.

ImplementationEscalation RateCost vs. Always-Frontier
Naive cascade (single sample)20-30%-70 to -80%
BEST-Route (3 samples)15-22%-78 to -85%
Unified routing + cascade10-18%-82 to -90%
Domain-trained quality estimator5-15%-85 to -98%

The 98% figure from FrugalGPT required a domain-trained estimator on a specific enterprise workload. For general workloads without training data, 80-85% is a realistic target with the unified approach.


The quality estimator is the bottleneck.

The heuristic estimator above is adequate for prototyping. Production systems need something better:

  1. Keyword heuristics — fast, cheap, domain-specific, breaks on novel failure modes
  2. Small classifier on response features — logistic regression on (query, response) pairs labeled good/bad by a frontier model. Requires 500-2,000 labeled examples.
  3. Reward model — fine-tune a small LLM to predict quality. Higher accuracy, higher cost per query.

Most teams start with option 2. Labeling 1,000 pairs costs around $200 and recovers in the first hour of production traffic.


Implementation checklist.

Before going to production with any cascade:


Skip the build if this is your first routing system.

Building and calibrating a cascade from scratch takes 2-4 weeks for an engineering team that has done it before. Six to eight weeks is realistic for teams doing it the first time, and the quality estimator will need multiple calibration cycles before it is stable.

Nadir ships a trained pre-classifier and an optional verifier that escalates a non-streaming answer when it misses your quality bar, behind an OpenAI-compatible base URL. Start in shadow mode, attach outcomes, and enforce only the routes that clear your bar. Nadir publishes no universal savings figure; what a cascade saves depends on your traffic, so measure it there.


Sources: [FrugalGPT (arXiv:2305.05176)](https://arxiv.org/abs/2305.05176), Stanford, TMLR 2024. [BEST-Route (arXiv:2506.22716)](https://arxiv.org/abs/2506.22716), Microsoft, ICML 2025. [ETH Zurich cascade routing (arXiv:2410.10347)](https://arxiv.org/abs/2410.10347), ICML 2025.

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.