Abstract.
Every cheap-first LLM setup makes the same all-or-nothing bet. A router sends the prompt to the small model or the large one. A cascade tries the small model and, if the answer fails a check, pays the large model for a complete second answer. A January 2026 paper, "Pay for Hints, Not Answers: LLM Shepherding for Cost-Efficient Inference" (Dong, Sharma, O'Toole, Champati, and Wu, arXiv:2601.22132), asks whether there's something in between. Its answer: buy only the first 10 to 30% of the large model's answer, hand that prefix to the small model as a hint, and let the small model finish. On four benchmarks the best variant cut cost 42 to 94% versus using the large model for everything, and on GSM8K it was both cheaper and more accurate than RouteLLM. This post walks through the paper, then prices the idea with frontier API rates, where the part the paper didn't have to deal with shows up: a hint call still pays for the whole prompt. On a reasoning or code workload, a 20% hint costs 29% of a full Opus answer and shepherding comes out 14% cheaper than a cascade. On long-context RAG, the same hint costs 93% of the answer and shepherding is the most expensive option we modeled. Includes a 50-line Python implementation with a break-even guard that skips the hint when it can't pay for itself.
The cost model in this post is illustrative, built from public list prices and the cited paper's results, not measured production traces. Not derived from proprietary customer data. Sources cited throughout.
The research question.
A cascade is the most common way to use a cheap model safely. Haiku answers first, a verifier checks it, and anything that fails goes to Opus. We've written about tuning the threshold and about the overhead the checking step adds. The part nobody questions is the escalation itself: when Haiku fails, you pay Opus for the whole answer, even though Haiku already got most of the way there on a lot of those prompts.
Shepherding asks a narrower question. If the small model is stuck, how much of the large model's answer does it actually need to get unstuck? And if the answer is "the first few lines," how much does that save once you price it with real API rates?
What the paper did.
The setup is small and clean. The small model is Llama-3.2-3B-Instruct, run locally, so the authors count its cost as zero. The large model is Llama-3.3-70B via the Groq API at $0.59 per million input tokens and $0.79 per million output. The benchmarks are GSM8K and CNK12 (math), and HumanEval and MBPP (code).
The mechanism is one line: call the large model with a cap of n output tokens, concatenate what comes back to the original query, and let the small model generate the rest. The paper's key observation is that "even hints comprising 10–30% of the full LLM response improve SLM accuracy significantly," with diminishing returns past about 60%.
It comes in two versions:
- Proactive shepherding. A two-stage predictor looks at the query before anything runs. Stage one decides whether a hint is needed at all. Stage two, only for the queries that need one, predicts how many tokens to buy. The first stage matters more than it sounds: on GSM8K, 80.6% of queries needed no hint at all. On CNK12, 46.7%.
- Reactive shepherding. The small model answers first, several times. If its samples agree, keep the answer. If they don't, buy a hint and regenerate. This is a cascade where the escalation step buys a prefix instead of an answer.
What it found.
On GSM8K, reactive shepherding reached 89.1% for $0.034, against RouteLLM's 88.1% for $0.065. Better accuracy for about half the cost. Against the 70B model alone, it cut cost 67.4%.
Here are all four benchmarks, with one extra column we computed: how much of the gap between the small and large model the hint closed.
| Benchmark | 3B alone | 70B alone | Reactive shepherding | Gap closed | Cost vs 70B alone |
|---|---|---|---|---|---|
| GSM8K | 73.0% | 98.0% | 89.1% | 64% | −67.4% |
| CNK12 | 53.8% | 84.4% | 76.0% | 73% | −42.1% |
| HumanEval | 48.8% | 83.5% | 76.2% | 79% | −44.3% |
| MBPP | 65.2% | 74.2% | 67.2% | 22% | −93.6% |
Three things are worth taking from it.
- Most of a hard answer's value is in its opening. A prefix that sets up the right approach, the first equation or the function signature and first loop, closes 64 to 79% of the gap between the two models on three of four benchmarks. The rest of the large model's answer is mostly the small model's job anyway.
- Shepherding isn't the accuracy winner. On GSM8K, the ABC cascade reached 94.9% for $0.048. If you need the last five points, you pay for full answers. What shepherding wins is the paper's accuracy-per-cost metric: the most accuracy per dollar, not the most accuracy.
- MBPP is the warning. The hint closed only 22% of the gap there. The cost reduction looks spectacular because the policy mostly declined to buy hints. A shepherd that rarely helps is just the small model with extra steps.
The authors are clear about scope. Every benchmark has an automatic correctness check, both models are from the same family, and they flag "tokenizer alignment across different architectures" as an open question. Open-ended tasks like summarization would need a learned reward model instead of a pass/fail check.
The part the paper didn't have to price.
The paper's large model charges $0.59 in and $0.79 out, almost the same for both, and GSM8K prompts are short. Under those prices, stopping generation at 20% saves close to 80% of the call.
Frontier APIs don't look like that. Opus 5 lists at $5 per million input tokens and $25 per million output, and production prompts are rarely a single math question. A hint call pays for the entire prompt, then stops early. Output is the only part you save. So what a hint costs, as a share of the full answer, depends on how much of the bill was output to begin with:
hint cost / full cost = (input $ + h × output $) / (input $ + output $)
where h is the hint length as a share of the full answer.
| Workload shape | Tokens in / out | Opus full answer | 20% hint | Hint as share of answer |
|---|---|---|---|---|
| Long-context RAG | 20,000 / 400 | $0.110 | $0.102 | 92.7% |
| Chat / extraction | 1,500 / 500 | $0.020 | $0.010 | 50.0% |
| Reasoning / code | 2,000 / 3,000 | $0.085 | $0.025 | 29.4% |
On a RAG prompt with 20,000 tokens of retrieved context and a 400-token answer, the "hint" is almost the whole bill. You've paid Opus to read everything and then thrown away the part it was about to write.
Methodology.
To compare policies, we priced one million requests per month for each workload shape. An analytical model, not a benchmark:
| Parameter | Value | Notes |
|---|---|---|
| Small model | Haiku 4.5, $1 / $5 per 1M tokens | Paid, unlike the paper's local 3B |
| Large model | Opus 5, $5 / $25 per 1M tokens | September 2026 list |
| Router | Sends 35% to Opus, 65% to Haiku | Assumes a perfect classifier, the best case for routing |
| Cascade | Haiku first, 30% fail the check, those go to Opus | Verifier cost left out, since it's the same for cascade and shepherd |
| Shepherd | Haiku first, 30% fail, those buy a 20% Opus hint, Haiku regenerates | Haiku's second pass reads the hint as input and writes the remaining 80% |
| Rescue rate | 70% of hinted retries pass the check | Between the paper's 64% and 79% on three of four benchmarks |
| Fallback | Unrescued retries go to a full Opus answer | So every policy ends at the same quality bar |
Reactive shepherding in the paper samples the small model several times to measure agreement. With a local 3B model that's free. With Haiku at API prices it isn't, so our shepherd uses the same single check as the cascade.
Finding 1: shepherding wins where output dominates.
| Workload shape | Opus only | Router | Cascade | Shepherd | Shepherd vs cascade |
|---|---|---|---|---|---|
| Long-context RAG | $110,000 | $52,800 | $55,000 | $69,004 | +25.5% |
| Chat / extraction | $20,000 | $9,600 | $10,000 | $9,880 | −1.2% |
| Reasoning / code | $85,000 | $40,800 | $42,500 | $36,530 | −14.0% |
On reasoning and code generation, where a single Opus answer runs thousands of output tokens, shepherding is the cheapest policy we modeled: 14% below the cascade, and 10% below a router that never misclassifies anything. That router doesn't exist. A real one sends some hard prompts to Haiku and some easy ones to Opus, and pays for both mistakes.
On chat and extraction, it's a wash. On long-context RAG, shepherding is the most expensive policy, because every failed prompt pays Opus to read 20,000 tokens for a hint and then pays Haiku to read them again.
Finding 2: the break-even is a rescue rate.
For any request that fails the first check, shepherding beats escalating straight to Opus when:
hint cost + small-model retry cost < rescue rate × full large-model cost
Rearranged, the minimum rescue rate for a hint to pay for itself:
| Workload shape | Break-even rescue rate | Paper's measured range |
|---|---|---|
| Reasoning / code | 46.6% | 64 to 79% on GSM8K, CNK12, HumanEval |
| Chat / extraction | 68.0% | |
| Long-context RAG | 112.4% | Impossible |
On the reasoning shape, even a hint that rescues only half the retries still beats the cascade. On RAG, a hint that rescued every single retry would still lose, because the hint plus Haiku's second pass costs more than the Opus answer it was supposed to replace. That's the one number to compute on your own traffic before you build any of this.
Finding 3: caching moves the line.
The input side of the hint call is the problem, and the input side is what prompt caching discounts. If the long context is a shared, cached prefix (a system prompt, a codebase, a document set reused across requests), cache reads bill at a fraction of base input and the RAG curve in Chart 3 drops toward the chat curve. A shepherd that sends the hint request to the provider where the prefix is already warm is a different cost model from one that doesn't. Uncached, per-request context is where shepherding has nothing to offer.
Tutorial: a shepherd with a break-even guard.
You need three things: a check for the small model's answer, token counts for the prompt, and a price table. The check is the hard part and it's workload-specific: unit tests for code, a recomputation for math, a schema plus business rules for extraction. Without one, you have no signal for when to buy a hint.
Step 1: prices and the guard.
PRICE = { # $ per token, input / output
"claude-haiku-4-5": (1e-6, 5e-6),
"claude-opus-5": (5e-6, 25e-6),
}
SMALL, LARGE = "claude-haiku-4-5", "claude-opus-5"
def call_cost(model: str, tokens_in: int, tokens_out: int) -> float:
pin, pout = PRICE[model]
return tokens_in * pin + tokens_out * pout
def hint_pays(tokens_in: int, expected_out: int, h: float = 0.2,
rescue: float = 0.7) -> bool:
"""True if a hint + small retry beats escalating straight to the large model."""
hint_tokens = int(h * expected_out)
hint = call_cost(LARGE, tokens_in, hint_tokens)
retry = call_cost(SMALL, tokens_in + hint_tokens, expected_out - hint_tokens)
full = call_cost(LARGE, tokens_in, expected_out)
return hint + retry < rescue * full
Set rescue from your own data once you have it. Start at 0.5 so the guard is conservative.
Step 2: the shepherd.
from openai import OpenAI
client = OpenAI(base_url="https://api.getnadir.com/v1", api_key=NADIR_KEY)
def ask(model, messages, max_tokens):
r = client.chat.completions.create(model=model, messages=messages,
max_tokens=max_tokens)
return r.choices[0].message.content, r.usage
def shepherd(messages, check, expected_out=2000, h=0.2):
# 1. Small model first.
answer, usage = ask(SMALL, messages, expected_out)
if check(answer):
return answer, "small"
# 2. Hint only if it can pay for itself on this prompt.
if hint_pays(usage.prompt_tokens, expected_out, h):
hint, _ = ask(LARGE, messages, max_tokens=int(h * expected_out))
retry = messages + [{
"role": "user",
"content": "A stronger model started the solution below. "
"Continue from where it stops and give the complete answer.\n\n"
+ hint,
}]
answer, _ = ask(SMALL, retry, expected_out)
if check(answer):
return answer, "shepherded"
# 3. Fall back to the full large-model answer.
answer, _ = ask(LARGE, messages, expected_out)
return answer, "large"
Pinning model sends each call to that model. Through Nadir's gateway, each response also reports what it cost, so you can check the guard's assumptions against real bills.
Step 3: three details that decide whether it works.
- Turn extended thinking off for the hint call. With thinking on, a 400-token cap can be spent entirely on reasoning the small model never sees. You pay for a hint and receive nothing.
- Pass the hint as text, not as a prefill. The paper concatenates tokens from a same-family model. Across vendors, a plain-text prefix in the user turn avoids the tokenizer question the authors left open. We haven't benchmarked cross-family hints; measure before trusting them.
- Log the outcome label.
small,shepherded, andlargeper request give you the rescue rate directly. That's the number the guard needs, and it'll differ by task type.
Step 4: add the proactive gate.
The paper's biggest single saving came from knowing that 80.6% of GSM8K queries need no hint. You can get a similar signal without training a predictor: classify difficulty up front, skip the whole shepherd path for simple prompts, and for prompts classified as hard, consider buying the hint first instead of paying Haiku to fail.
import requests
def tier(messages) -> str:
return requests.post(
"https://api.getnadir.com/v1/bucket",
headers={"X-API-Key": NADIR_KEY},
json={"messages": messages},
timeout=2,
).json()["bucket"] # simple | medium | complex
For complex prompts on an output-heavy workload, calling the hint first and Haiku once is cheaper than Haiku, fail, hint, Haiku.
When not to bother.
- Input-heavy prompts. Long retrieved context, big documents, large tool outputs. Run
hint_payson a sample. If it's false for most of your traffic, stop here. - No automatic check. Shepherding, like a cascade, needs a pass/fail signal. Without one you can't tell when to buy a hint or whether it worked.
- Latency-critical paths. A shepherded request makes three sequential calls. For a user waiting on a chat reply, a router that decides once is usually the better trade.
- Short answers. If the full answer is 150 tokens, a 20% hint is 30 tokens. There's nothing to save.
Where Nadir fits.
Nadir doesn't run shepherding for you today. The shepherd in this post is about 50 lines of your code, and the check it depends on is specific to your workload. What it needs from outside is the part that's tedious to build well: a difficulty signal before any tokens are spent, and a real cost figure per call.
That's what Nadir provides. POST /v1/bucket classifies each prompt as simple, medium, or complex without a provider call, which is the proactive gate in Step 4. Nadir's verifier-gated cascade covers the case where full escalation is the right answer. Through the OpenAI compatible gateway, every response reports the model that served it and what it cost, so the break-even guard runs on your prices and your token counts, not ours.
If a large share of your Opus bill is long code or reasoning answers, start with a free key, bucket a week of the requests you currently escalate, and run hint_pays over them. That tells you whether shepherding is worth building before you write any of it.
Conclusion.
Routing and cascading both treat the large model as all-or-nothing: you either skip it or pay for its whole answer. Shepherding shows there's a third option. Buy the opening of the answer, where most of the value is, and let the cheap model finish. In the paper, that cut cost 42 to 94% versus the large model alone and beat RouteLLM on both cost and accuracy on GSM8K. At frontier API prices, the idea holds only where output is most of the bill. On reasoning and code, a 20% hint costs 29% of the answer and shepherding comes out 14% below a cascade in our model. On long-context RAG, a 20% hint costs 93% of the answer and it loses to everything. The rule is one inequality, hint plus retry against rescue rate times full answer, and it's worth computing on your own traffic before you pay for another complete answer.
Costs in this post are illustrative, modeled for this post from public list prices and the cited paper's reported accuracy, and are not derived from customer data. Sources: [Dong, Sharma, O'Toole, Champati, and Wu, "Pay for Hints, Not Answers: LLM Shepherding for Cost-Efficient Inference," arXiv:2601.22132, January 2026](https://arxiv.org/abs/2601.22132). [Anthropic, Pricing](https://platform.claude.com/docs/en/about-claude/pricing).