Buy the Hint

Routers and cascades treat the large model as all-or-nothing: skip it entirely, or pay for its whole answer every time the cheap model fails a check. A January 2026 paper, LLM Shepherding (arXiv:2601.22132), buys only the first 10 to 30% of the large model's answer and lets a small model finish. On four math and code benchmarks it cut cost 42 to 94% versus the large model alone, and on GSM8K it beat RouteLLM on both cost and accuracy. This post prices the idea at frontier API rates, where a hint call still pays for the whole prompt. On reasoning and code, a 20% Opus hint costs 29% of the full answer and shepherding comes out 14% cheaper than a cascade. On long-context RAG, the same hint costs 93% of the answer and shepherding loses to everything. Includes a 50-line Python shepherd with a break-even guard.

Published 2026-09-28 by Dor Amir on the Nadir blog.

Filed under Routing & Cascades.

Abstract.

Every cheap-first LLM setup makes the same all-or-nothing bet. A router sends the prompt to the small model or the large one. A cascade tries the small model and, if the answer fails a check, pays the large model for a complete second answer. A January 2026 paper, "Pay for Hints, Not Answers: LLM Shepherding for Cost-Efficient Inference" (Dong, Sharma, O'Toole, Champati, and Wu, arXiv:2601.22132), asks whether there's something in between. Its answer: buy only the first 10 to 30% of the large model's answer, hand that prefix to the small model as a hint, and let the small model finish. On four benchmarks the best variant cut cost 42 to 94% versus using the large model for everything, and on GSM8K it was both cheaper and more accurate than RouteLLM. This post walks through the paper, then prices the idea with frontier API rates, where the part the paper didn't have to deal with shows up: a hint call still pays for the whole prompt. On a reasoning or code workload, a 20% hint costs 29% of a full Opus answer and shepherding comes out 14% cheaper than a cascade. On long-context RAG, the same hint costs 93% of the answer and shepherding is the most expensive option we modeled. Includes a 50-line Python implementation with a break-even guard that skips the hint when it can't pay for itself.

The cost model in this post is illustrative, built from public list prices and the cited paper's results, not measured production traces. Not derived from proprietary customer data. Sources cited throughout.

The research question.

A cascade is the most common way to use a cheap model safely. Haiku answers first, a verifier checks it, and anything that fails goes to Opus. We've written about tuning the threshold and about the overhead the checking step adds. The part nobody questions is the escalation itself: when Haiku fails, you pay Opus for the whole answer, even though Haiku already got most of the way there on a lot of those prompts.

Shepherding asks a narrower question. If the small model is stuck, how much of the large model's answer does it actually need to get unstuck? And if the answer is "the first few lines," how much does that save once you price it with real API rates?

What the paper did.

Chart 1: Three ways to use the expensive model. Route: a classifier picks Haiku or Opus before generation and pays one full answer. Cascade: Haiku answers, a verifier checks it, and a failure pays for a full Opus answer. Shepherd: Haiku answers, the verifier checks it, and a failure buys only the first 20% of an Opus answer as a hint, which Haiku then completes.
Chart 1: Three ways to use the expensive model. Route: a classifier picks Haiku or Opus before generation and pays one full answer. Cascade: Haiku answers, a verifier checks it, and a failure pays for a full Opus answer. Shepherd: Haiku answers, the verifier checks it, and a failure buys only the first 20% of an Opus answer as a hint, which Haiku then completes.

The setup is small and clean. The small model is Llama-3.2-3B-Instruct, run locally, so the authors count its cost as zero. The large model is Llama-3.3-70B via the Groq API at $0.59 per million input tokens and $0.79 per million output. The benchmarks are GSM8K and CNK12 (math), and HumanEval and MBPP (code).

The mechanism is one line: call the large model with a cap of n output tokens, concatenate what comes back to the original query, and let the small model generate the rest. The paper's key observation is that "even hints comprising 10–30% of the full LLM response improve SLM accuracy significantly," with diminishing returns past about 60%.

It comes in two versions:

  1. Proactive shepherding. A two-stage predictor looks at the query before anything runs. Stage one decides whether a hint is needed at all. Stage two, only for the queries that need one, predicts how many tokens to buy. The first stage matters more than it sounds: on GSM8K, 80.6% of queries needed no hint at all. On CNK12, 46.7%.
  2. Reactive shepherding. The small model answers first, several times. If its samples agree, keep the answer. If they don't, buy a hint and regenerate. This is a cascade where the escalation step buys a prefix instead of an answer.

What it found.

Chart 2: GSM8K accuracy versus cost from arXiv:2601.22132. Llama 3B only: 73.0% at $0. Llama 70B only: 98.0% at $0.104. RouteLLM: 88.1% at $0.065. GraphRouter: 82.2% at $0.041. ABC cascade: 94.9% at $0.048. Proactive shepherding: 80.9% at $0.036. Reactive shepherding: 89.1% at $0.034.
Chart 2: GSM8K accuracy versus cost from arXiv:2601.22132. Llama 3B only: 73.0% at $0. Llama 70B only: 98.0% at $0.104. RouteLLM: 88.1% at $0.065. GraphRouter: 82.2% at $0.041. ABC cascade: 94.9% at $0.048. Proactive shepherding: 80.9% at $0.036. Reactive shepherding: 89.1% at $0.034.

On GSM8K, reactive shepherding reached 89.1% for $0.034, against RouteLLM's 88.1% for $0.065. Better accuracy for about half the cost. Against the 70B model alone, it cut cost 67.4%.

Here are all four benchmarks, with one extra column we computed: how much of the gap between the small and large model the hint closed.

Benchmark3B alone70B aloneReactive shepherdingGap closedCost vs 70B alone
GSM8K73.0%98.0%89.1%64%−67.4%
CNK1253.8%84.4%76.0%73%−42.1%
HumanEval48.8%83.5%76.2%79%−44.3%
MBPP65.2%74.2%67.2%22%−93.6%

Three things are worth taking from it.

  1. Most of a hard answer's value is in its opening. A prefix that sets up the right approach, the first equation or the function signature and first loop, closes 64 to 79% of the gap between the two models on three of four benchmarks. The rest of the large model's answer is mostly the small model's job anyway.
  2. Shepherding isn't the accuracy winner. On GSM8K, the ABC cascade reached 94.9% for $0.048. If you need the last five points, you pay for full answers. What shepherding wins is the paper's accuracy-per-cost metric: the most accuracy per dollar, not the most accuracy.
  3. MBPP is the warning. The hint closed only 22% of the gap there. The cost reduction looks spectacular because the policy mostly declined to buy hints. A shepherd that rarely helps is just the small model with extra steps.

The authors are clear about scope. Every benchmark has an automatic correctness check, both models are from the same family, and they flag "tokenizer alignment across different architectures" as an open question. Open-ended tasks like summarization would need a learned reward model instead of a pass/fail check.

The part the paper didn't have to price.

The paper's large model charges $0.59 in and $0.79 out, almost the same for both, and GSM8K prompts are short. Under those prices, stopping generation at 20% saves close to 80% of the call.

Frontier APIs don't look like that. Opus 5 lists at $5 per million input tokens and $25 per million output, and production prompts are rarely a single math question. A hint call pays for the entire prompt, then stops early. Output is the only part you save. So what a hint costs, as a share of the full answer, depends on how much of the bill was output to begin with:

hint cost / full cost = (input $ + h × output $) / (input $ + output $)

where h is the hint length as a share of the full answer.

Chart 3: What a hint costs as a share of the full Opus answer, by hint length, at $5 input and $25 output per million tokens. Long-context RAG with 20,000 input and 400 output tokens: a 20% hint costs 92.7% of the full answer. Chat or extraction with 1,500 in and 500 out: 50.0%. Reasoning or code with 2,000 in and 3,000 out: 29.4%.
Chart 3: What a hint costs as a share of the full Opus answer, by hint length, at $5 input and $25 output per million tokens. Long-context RAG with 20,000 input and 400 output tokens: a 20% hint costs 92.7% of the full answer. Chat or extraction with 1,500 in and 500 out: 50.0%. Reasoning or code with 2,000 in and 3,000 out: 29.4%.
Workload shapeTokens in / outOpus full answer20% hintHint as share of answer
Long-context RAG20,000 / 400$0.110$0.10292.7%
Chat / extraction1,500 / 500$0.020$0.01050.0%
Reasoning / code2,000 / 3,000$0.085$0.02529.4%

On a RAG prompt with 20,000 tokens of retrieved context and a 400-token answer, the "hint" is almost the whole bill. You've paid Opus to read everything and then thrown away the part it was about to write.

Methodology.

To compare policies, we priced one million requests per month for each workload shape. An analytical model, not a benchmark:

ParameterValueNotes
Small modelHaiku 4.5, $1 / $5 per 1M tokensPaid, unlike the paper's local 3B
Large modelOpus 5, $5 / $25 per 1M tokensSeptember 2026 list
RouterSends 35% to Opus, 65% to HaikuAssumes a perfect classifier, the best case for routing
CascadeHaiku first, 30% fail the check, those go to OpusVerifier cost left out, since it's the same for cascade and shepherd
ShepherdHaiku first, 30% fail, those buy a 20% Opus hint, Haiku regeneratesHaiku's second pass reads the hint as input and writes the remaining 80%
Rescue rate70% of hinted retries pass the checkBetween the paper's 64% and 79% on three of four benchmarks
FallbackUnrescued retries go to a full Opus answerSo every policy ends at the same quality bar

Reactive shepherding in the paper samples the small model several times to measure agreement. With a local 3B model that's free. With Haiku at API prices it isn't, so our shepherd uses the same single check as the cascade.

Finding 1: shepherding wins where output dominates.

Chart 4: Monthly bill per million requests by workload shape. Long-context RAG: router $52,800, cascade $55,000, shepherd $69,004. Chat or extraction: router $9,600, cascade $10,000, shepherd $9,880. Reasoning or code: router $40,800, cascade $42,500, shepherd $36,530. Opus for everything: $110,000, $20,000, and $85,000.
Chart 4: Monthly bill per million requests by workload shape. Long-context RAG: router $52,800, cascade $55,000, shepherd $69,004. Chat or extraction: router $9,600, cascade $10,000, shepherd $9,880. Reasoning or code: router $40,800, cascade $42,500, shepherd $36,530. Opus for everything: $110,000, $20,000, and $85,000.
Workload shapeOpus onlyRouterCascadeShepherdShepherd vs cascade
Long-context RAG$110,000$52,800$55,000$69,004+25.5%
Chat / extraction$20,000$9,600$10,000$9,880−1.2%
Reasoning / code$85,000$40,800$42,500$36,530−14.0%

On reasoning and code generation, where a single Opus answer runs thousands of output tokens, shepherding is the cheapest policy we modeled: 14% below the cascade, and 10% below a router that never misclassifies anything. That router doesn't exist. A real one sends some hard prompts to Haiku and some easy ones to Opus, and pays for both mistakes.

On chat and extraction, it's a wash. On long-context RAG, shepherding is the most expensive policy, because every failed prompt pays Opus to read 20,000 tokens for a hint and then pays Haiku to read them again.

Finding 2: the break-even is a rescue rate.

For any request that fails the first check, shepherding beats escalating straight to Opus when:

hint cost + small-model retry cost < rescue rate × full large-model cost

Rearranged, the minimum rescue rate for a hint to pay for itself:

Workload shapeBreak-even rescue ratePaper's measured range
Reasoning / code46.6%64 to 79% on GSM8K, CNK12, HumanEval
Chat / extraction68.0%
Long-context RAG112.4%Impossible

On the reasoning shape, even a hint that rescues only half the retries still beats the cascade. On RAG, a hint that rescued every single retry would still lose, because the hint plus Haiku's second pass costs more than the Opus answer it was supposed to replace. That's the one number to compute on your own traffic before you build any of this.

Finding 3: caching moves the line.

The input side of the hint call is the problem, and the input side is what prompt caching discounts. If the long context is a shared, cached prefix (a system prompt, a codebase, a document set reused across requests), cache reads bill at a fraction of base input and the RAG curve in Chart 3 drops toward the chat curve. A shepherd that sends the hint request to the provider where the prefix is already warm is a different cost model from one that doesn't. Uncached, per-request context is where shepherding has nothing to offer.

Tutorial: a shepherd with a break-even guard.

You need three things: a check for the small model's answer, token counts for the prompt, and a price table. The check is the hard part and it's workload-specific: unit tests for code, a recomputation for math, a schema plus business rules for extraction. Without one, you have no signal for when to buy a hint.

Step 1: prices and the guard.

PRICE = {  # $ per token, input / output
    "claude-haiku-4-5": (1e-6, 5e-6),
    "claude-opus-5":    (5e-6, 25e-6),
}
SMALL, LARGE = "claude-haiku-4-5", "claude-opus-5"

def call_cost(model: str, tokens_in: int, tokens_out: int) -> float:
    pin, pout = PRICE[model]
    return tokens_in * pin + tokens_out * pout

def hint_pays(tokens_in: int, expected_out: int, h: float = 0.2,
              rescue: float = 0.7) -> bool:
    """True if a hint + small retry beats escalating straight to the large model."""
    hint_tokens = int(h * expected_out)
    hint = call_cost(LARGE, tokens_in, hint_tokens)
    retry = call_cost(SMALL, tokens_in + hint_tokens, expected_out - hint_tokens)
    full = call_cost(LARGE, tokens_in, expected_out)
    return hint + retry < rescue * full

Set rescue from your own data once you have it. Start at 0.5 so the guard is conservative.

Step 2: the shepherd.

from openai import OpenAI

client = OpenAI(base_url="https://api.getnadir.com/v1", api_key=NADIR_KEY)

def ask(model, messages, max_tokens):
    r = client.chat.completions.create(model=model, messages=messages,
                                       max_tokens=max_tokens)
    return r.choices[0].message.content, r.usage

def shepherd(messages, check, expected_out=2000, h=0.2):
    # 1. Small model first.
    answer, usage = ask(SMALL, messages, expected_out)
    if check(answer):
        return answer, "small"

    # 2. Hint only if it can pay for itself on this prompt.
    if hint_pays(usage.prompt_tokens, expected_out, h):
        hint, _ = ask(LARGE, messages, max_tokens=int(h * expected_out))
        retry = messages + [{
            "role": "user",
            "content": "A stronger model started the solution below. "
                       "Continue from where it stops and give the complete answer.\n\n"
                       + hint,
        }]
        answer, _ = ask(SMALL, retry, expected_out)
        if check(answer):
            return answer, "shepherded"

    # 3. Fall back to the full large-model answer.
    answer, _ = ask(LARGE, messages, expected_out)
    return answer, "large"

Pinning model sends each call to that model. Through Nadir's gateway, each response also reports what it cost, so you can check the guard's assumptions against real bills.

Step 3: three details that decide whether it works.

Step 4: add the proactive gate.

The paper's biggest single saving came from knowing that 80.6% of GSM8K queries need no hint. You can get a similar signal without training a predictor: classify difficulty up front, skip the whole shepherd path for simple prompts, and for prompts classified as hard, consider buying the hint first instead of paying Haiku to fail.

import requests

def tier(messages) -> str:
    return requests.post(
        "https://api.getnadir.com/v1/bucket",
        headers={"X-API-Key": NADIR_KEY},
        json={"messages": messages},
        timeout=2,
    ).json()["bucket"]          # simple | medium | complex

For complex prompts on an output-heavy workload, calling the hint first and Haiku once is cheaper than Haiku, fail, hint, Haiku.

When not to bother.

Where Nadir fits.

Nadir doesn't run shepherding for you today. The shepherd in this post is about 50 lines of your code, and the check it depends on is specific to your workload. What it needs from outside is the part that's tedious to build well: a difficulty signal before any tokens are spent, and a real cost figure per call.

That's what Nadir provides. POST /v1/bucket classifies each prompt as simple, medium, or complex without a provider call, which is the proactive gate in Step 4. Nadir's verifier-gated cascade covers the case where full escalation is the right answer. Through the OpenAI compatible gateway, every response reports the model that served it and what it cost, so the break-even guard runs on your prices and your token counts, not ours.

If a large share of your Opus bill is long code or reasoning answers, start with a free key, bucket a week of the requests you currently escalate, and run hint_pays over them. That tells you whether shepherding is worth building before you write any of it.

Conclusion.

Routing and cascading both treat the large model as all-or-nothing: you either skip it or pay for its whole answer. Shepherding shows there's a third option. Buy the opening of the answer, where most of the value is, and let the cheap model finish. In the paper, that cut cost 42 to 94% versus the large model alone and beat RouteLLM on both cost and accuracy on GSM8K. At frontier API prices, the idea holds only where output is most of the bill. On reasoning and code, a 20% hint costs 29% of the answer and shepherding comes out 14% below a cascade in our model. On long-context RAG, a 20% hint costs 93% of the answer and it loses to everything. The rule is one inequality, hint plus retry against rescue rate times full answer, and it's worth computing on your own traffic before you pay for another complete answer.


Costs in this post are illustrative, modeled for this post from public list prices and the cited paper's reported accuracy, and are not derived from customer data. Sources: [Dong, Sharma, O'Toole, Champati, and Wu, "Pay for Hints, Not Answers: LLM Shepherding for Cost-Efficient Inference," arXiv:2601.22132, January 2026](https://arxiv.org/abs/2601.22132). [Anthropic, Pricing](https://platform.claude.com/docs/en/about-claude/pricing).

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.