Abstract.
Almost every LLM router in production makes one decision per request: this prompt goes to the small model or the large one. For a chat reply, that's the right unit. For an agent, it isn't. A coding agent working one ticket makes dozens of calls, and most of them are search, read a file, run the tests, read the output. A small model can do those. It just can't do the whole ticket. A September 28, 2026 paper, "RSI-Router: Evolving Subtask-Level LLM Routing and Skills for Cost-Efficient Agents" (Li, Zhang, Cui, Mu, Zhang, Zhang, Jia, and Hu, arXiv:2609.34712), routes at the level in between: the subtask, a recurring phase of the task like "locate the bug" or "verify the fix." Paired with a 9B model, it cut the cost of a DeepSeek-V4.1-Flash agent by 51.7% on average across five benchmarks and scored higher on all five. This post reads the results table more carefully than the headline does, re-prices the idea at the September 2026 list prices you'd actually pay, and ends with a tutorial for doing per-phase routing in your own agent loop.
Benchmark numbers are from the cited paper. The re-pricing and the trajectory in Chart 1 are illustrative, built from public list prices, not measured production traces or customer data.
The unit of routing is wrong for agents.
RouteLLM's finding that most GPT-4 calls didn't need GPT-4 was about single prompts. Agents break that framing. A task-level router looking at "fix the failing date parser in billing/" sees a hard coding task and sends it to the large model, and every one of the 30 calls in that trajectory bills at the large model's rate, including the twelve that were grep and cat.
The other extreme, choosing a model for every individual step, has its own problem. The paper calls it out directly: step-level routers explore an enormous decision space with no structure to guide them, and when a trajectory fails, it's hard to say which of the 30 routing decisions caused it.
Subtasks sit in the middle. The paper's own example: "code repair may involve locating relevant code, implementing a fix, and verifying the result, with the same subtask recurring within a trajectory." Few enough categories to learn a policy for, specific enough that the small model can own some of them outright.
What RSI-Router does.
The router is learned offline from training trajectories, in a loop the authors call recursive self-improvement. Six iterations, four candidate strategies per iteration, four stages each time:
- Subtask mining. Read past trajectories, split them by what each stretch of steps was trying to do, and write a definition plus an identification rule for each subtask. From the second iteration on, existing subtasks get kept, split, or merged.
- Routing strategy evolution. Propose four different assignments of subtasks to models: maybe "locate" goes small and "verify" goes large, maybe the reverse.
- Model-specific skill evolution. Compare each candidate's trajectories with the large model's on the same tasks, find where the small model failed or wasted steps, and write reusable skills (instructions) that fix those failures for that model.
- Pareto selection. Keep the routers that aren't beaten on both validation score and cost.
At run time, the small model reads the user query, recent history, and the previous subtask, predicts which subtask comes next, and the policy picks the model. The pair tested is DeepSeek-V4.1-Flash ($0.30 / $1.20 per million tokens) as the large model and Qwen3.5-9B as the small one.
What it found.
| Benchmark | Large only | Small only | RSI-router | Cost vs large only |
|---|---|---|---|---|
| ALFWorld | 90.62% / $1.50 | 65.62% / $0.13 | 98.44% / $0.37 | −75.3% |
| ScienceWorld | 34.38% / $1.91 | 14.58% / $0.06 | 35.94% / $0.34 | −82.2% |
| WebShop | 0.59 / $0.75 | 0.48 / $0.02 | 0.61 / $0.19 | −74.7% |
| SWE-bench Verified | 83.33% / $3.46 | 62.50% / $0.69 | 84.38% / $3.18 | −8.1% |
| Terminal-Bench 2.0 | 40.00% / $17.10 | 5.33% / $0.17 | 46.67% / $14.02 | −18.0% |
The paper compares against nine baselines: five task-level routers (HybridLLM, FrugalGPT, RouteLLM, GraphRouter, Avengers-Pro), two step-level routers (Router-R1, MTRouter), and both single models. RSI-Router sits on the Pareto frontier against all of them.
Three things in that table matter more than the 51.7%.
- The small model never comes close alone. Qwen3.5-9B by itself scores 5.33% on Terminal-Bench, against 40% for DeepSeek. Any task-level router has to send nearly every Terminal-Bench task to the large model. Subtask routing still found 18% of that bill to move. That's spend a per-request router can't see.
- The headline average hides the bill. 51.7% is the mean of five percentages. Add up the dollars and the five test sets cost $24.72 on DeepSeek alone and $18.10 routed, a 26.8% cut. The two coding benchmarks are 83% of the spend and got the two smallest cuts. If your agents write code, the number to plan around is closer to 8 to 18% than to 50%.
- The test sets are small. 64 tasks each for the three navigation benchmarks, 32 for SWE-bench Verified. One SWE-bench task is 3.1 points, so the 1.05-point gain there is less than one task. Read the score columns as "didn't get worse." The cost columns are the result.
One more detail from the ablations: removing the evolved skills lowered scores on ALFWorld, ScienceWorld, and WebShop by 12.4, 11.5, and 7.4 points, but raised them on SWE-bench Verified and Terminal-Bench by 9.7 and 10.5. Extra instructions that help a small model navigate a text world get in the way on code. Skills are per-domain, not a free add-on.
The part the paper didn't have to price.
Qwen3.5-9B doesn't have a list price in the paper. The authors estimate it from GPU-hours: across six workloads, DeepSeek-V4.1-Flash took 35.96 times the GPU-hours, so they price Qwen at about 1/30th, $0.01 / $0.04 per million tokens. That's a self-hosted 9B model at good utilization. Most teams pay list prices for both tiers, and the gap between tiers is much narrower than 30x.
The cost cut from moving work to a cheaper model is the share of spend you move times the price gap you close:
saving = moved share × (1 − 1 / price ratio)
Backing the moved share out of the paper's own numbers gives about 80% for the three navigation benchmarks and about 13% for the two coding ones. Re-price those at September 2026 pairs:
| Model pair | Price ratio | Navigation-heavy | Coding |
|---|---|---|---|
| Paper (GPU-hour estimate) | 30x | −77% | −13% |
| GPT-6 Sol → Luna ($2/$10 → $0.10/$0.50) | 20x | −76% | −13% |
| Opus 5.5 → Haiku 4.5 ($4/$20 → $1/$5) | 4x | −60% | −10% |
| Opus 5.5 → Sonnet 5 ($4/$20 → $2/$10) | 2x | −40% | −7% |
Two caveats on this table. It assumes the same steps get routed at every price, and a different small model will be able to own a different set of steps. It also ignores cache effects, which are the next section. But the shape holds: what you can route depends on the task, and what routing is worth depends on the price gap. A navigation-heavy agent on a 20x pair saves three quarters of its bill. A coding agent moving from Opus 5.5 to Sonnet 5 saves single digits, which is less than a cache miss can cost you.
Every switch is a cache decision.
The paper's costs are uncached token prices. Production agents run on prompt caching: turn 20 re-sends turns 1 through 19, and the provider bills that prefix at a cache-read rate instead of the full input rate. Switching models mid-trajectory means the new model has no cache for that prefix. You pay full input price, or a cache write, to rebuild it.
That's the case for routing at the subtask boundary rather than every step. A subtask is a run of consecutive steps on one model. One switch at the start of "locate," several cached steps inside it, one switch back for "edit." A per-step router that flips models on every call can lose more on cache rebuilds than it saves on price, and our cache switch penalty post works through that math in detail. The rule of thumb: only switch when the routed run is long enough to earn back one full-price read of the context.
Tutorial: per-phase routing in your own agent loop.
You don't need the paper's evolutionary loop to get started. The core of it is a policy table from phase to model, a cheap rule for which phase you're in, and per-phase cost data to decide what to move next. Here it is in about 60 lines, through an OpenAI compatible client.
Step 1: name your phases.
Pull a week of trajectories and label the steps. Most coding agents fall into five or six phases, and the tool being called gets you most of the way. That's a cheap version of the paper's identification rules:
PHASE_BY_TOOL = {
"grep": "locate", "glob": "locate", "list_dir": "locate",
"read_file": "read",
"edit_file": "edit", "write_file": "edit", "apply_patch": "edit",
"run_tests": "verify", "bash": "verify",
}
def current_phase(messages) -> str:
"""Phase of the next call, from the last tool the agent used."""
for m in reversed(messages):
for call in m.get("tool_calls") or []:
return PHASE_BY_TOOL.get(call["function"]["name"], "plan")
return "plan" # first turn: no tools yet
It's crude: the step after a read_file is often the edit. Start crude, measure, then refine the rule, which is exactly what the paper's stage 1 does on every iteration.
Step 2: a policy table and a client.
from openai import OpenAI
client = OpenAI(base_url="https://api.getnadir.com/v1", api_key=NADIR_KEY)
SMALL, LARGE = "claude-haiku-4-5", "claude-opus-5-5"
# Start conservative: everything large except phases you've measured.
POLICY = {"plan": LARGE, "locate": SMALL, "read": SMALL,
"edit": LARGE, "verify": SMALL}
def step(messages, tools, task_id, phase, model):
return client.chat.completions.create(
model=model,
messages=messages,
tools=tools,
extra_headers={
"X-Nadir-Use-Case": "coding-agent",
"X-Nadir-Tags": f"phase:{phase},task:{task_id}",
},
)
The tags are what makes Step 4 possible. Every call carries its phase, so cost per phase is a group-by, not a log-parsing project.
Step 3: the loop, with sticky phases and a failure escape.
def run_agent(task, tools, task_id, max_steps=40):
messages = [{"role": "user", "content": task}]
phase = model = None
small_failures = 0
for _ in range(max_steps):
next_phase = current_phase(messages)
if next_phase != phase: # switch only at a phase boundary
phase, model = next_phase, POLICY[next_phase]
small_failures = 0
reply = step(messages, tools, task_id, phase, model)
msg = reply.choices[0].message
messages.append(msg.model_dump(exclude_none=True))
if not msg.tool_calls:
return msg.content
for call in msg.tool_calls:
result = execute(call) # your tool runner
messages.append({"role": "tool", "tool_call_id": call.id,
"content": result.output})
if model == SMALL and result.failed:
small_failures += 1
if small_failures >= 2: # small model stuck: escalate this phase
model = LARGE
raise RuntimeError("step budget exhausted")
Three choices in there are deliberate:
- The model only changes at a phase boundary. Inside a phase, the prefix stays warm on one model. That's the cache argument from the previous section in two lines.
- Two failed tool calls on the small model escalate the rest of that phase. It's the subtask version of a cascade: cheap first, large when the evidence says so, without throwing away the trajectory so far.
- `max_steps` is a hard cap. A small model that loops on a failing test is how a cost optimization turns into an unbounded loop.
Step 4: decide what to move next.
Once a few hundred tasks have run, compare each phase's cost and outcome:
import pandas as pd
# One row per call, exported from your usage logs:
# task_id, phase, model, cost_usd, task_passed
df = pd.read_csv("agent_calls.csv")
by_phase = (df.groupby(["phase", "model"])
.agg(calls=("cost_usd", "size"),
spend=("cost_usd", "sum"),
pass_rate=("task_passed", "mean"))
.sort_values("spend", ascending=False))
print(by_phase)
The next phase to try on the small model is the one with the most large-model spend that isn't where tasks fail. Move it for a slice of traffic, compare pass rates over the same tasks, keep it if the pass rate holds. That's stage 2 and stage 4 of the paper done by hand, one phase per week instead of four strategies per iteration. When a phase almost works on the small model, write the failure pattern into its system prompt for that phase only. That's stage 3, and the ablation says to check it per domain rather than assume it helps.
When not to bother.
- Short tasks. An agent that finishes in three calls has no phases to split. Use a per-request router.
- Coding agents on a narrow price gap. At a 2x ratio and roughly 13% routable spend, per-phase routing saves single digits. Prompt caching and trimming the trajectory will likely save more for less work.
- No outcome signal. Step 4 needs to know whether a task passed. Without tests, a reviewer verdict, or a user acceptance signal, you can measure what each phase costs but not whether routing it hurt anything.
- Provider-specific continuation state. Some agent APIs keep server-side reasoning state that doesn't carry to another provider's model. Keep the large and small models within one provider, or switch only where that state ends.
Where Nadir fits.
Nadir doesn't learn your phase policy for you. Your agent knows which tool it just called; a gateway sitting outside the loop doesn't, and the subtask rules in the paper come from your own trajectories. What Nadir provides is the part of this that's tedious to build and needs to sit in the request path:
- One OpenAI compatible endpoint for both tiers, so the phase policy is a model string per call, with BYOK for your provider keys.
- Per-call cost attribution.
X-Nadir-TagsandX-Nadir-Use-Caseroll spend up per phase and per task in the dashboard, which is the input to Step 4. - Per-request routing for the calls you haven't assigned. Send
model="auto"for theplanphase, or any phase you haven't measured, and Nadir picks the tier per prompt before any provider tokens are spent. - Observe mode on a key serves exactly the model your agent asked for while recording what routing would have chosen, so you can size the saving before changing anything.
If you run agents that make more than a handful of calls per task, start with a free key, add the phase tag to every call for a week, and look at which phases carry your large-model spend. That table tells you whether subtask routing is worth 5% or 50% for your workload, before you write the rest of it.
Conclusion.
Per-request routing treats an agent task as one indivisible decision: hard task, large model, every call. RSI-Router shows that a hard task is mostly easy steps. With a 9B model owning the steps it can handle, the paper cut a DeepSeek agent's cost by 75 to 82% on navigation-style benchmarks and 8 to 18% on coding, with scores no worse on any of the five. Two corrections before you plan around it: the dollar-weighted cut across the five test sets is 26.8%, not 51.7%, because coding is where the money is, and the paper's 30x price gap is a self-hosting estimate. At list prices, the same routed steps are worth 60% on a navigation agent moving from Opus 5.5 to Haiku 4.5, and about 10% on a coding agent. Route by phase, switch only at phase boundaries, tag every call, and let your own per-phase numbers say how far to take it.
Benchmark results are from the cited paper. Re-priced savings are illustrative, computed from the paper's reported costs and public September 2026 list prices, and are not derived from customer data. Sources: [Li, Zhang, Cui, Mu, Zhang, Zhang, Jia, and Hu, "RSI-Router: Evolving Subtask-Level LLM Routing and Skills for Cost-Efficient Agents," arXiv:2609.34712, September 2026](https://arxiv.org/abs/2609.34712). [Anthropic, Pricing](https://platform.claude.com/docs/en/about-claude/pricing). [OpenAI, API Pricing](https://openai.com/api/pricing/).