Route the Step

Most LLM routers pick one model per request, which is the wrong unit for an agent that makes dozens of calls to finish a single task. Most of those calls are search, read, and run-the-tests steps a small model can handle even when it can't finish the task. RSI-Router (arXiv:2609.34712, September 28, 2026) routes per subtask instead, and paired DeepSeek-V4.1-Flash with Qwen3.5-9B to cut agent cost 51.7% on average across five benchmarks while scoring higher on all five. Read the table closely and the story changes: navigation benchmarks dropped 75 to 82%, coding only 8 to 18%, and the dollar-weighted cut is 26.8%. The paper's 30x price gap is a self-hosting estimate; re-priced at Opus 5.5 to Haiku 4.5, the same routed steps are worth about 60% on navigation and 10% on code. Includes a 60-line per-phase routing loop with cache-safe switching, failure escalation, and per-phase cost tags.

Published 2026-10-01 by Dor Amir on the Nadir blog.

Filed under Agents.

Abstract.

Almost every LLM router in production makes one decision per request: this prompt goes to the small model or the large one. For a chat reply, that's the right unit. For an agent, it isn't. A coding agent working one ticket makes dozens of calls, and most of them are search, read a file, run the tests, read the output. A small model can do those. It just can't do the whole ticket. A September 28, 2026 paper, "RSI-Router: Evolving Subtask-Level LLM Routing and Skills for Cost-Efficient Agents" (Li, Zhang, Cui, Mu, Zhang, Zhang, Jia, and Hu, arXiv:2609.34712), routes at the level in between: the subtask, a recurring phase of the task like "locate the bug" or "verify the fix." Paired with a 9B model, it cut the cost of a DeepSeek-V4.1-Flash agent by 51.7% on average across five benchmarks and scored higher on all five. This post reads the results table more carefully than the headline does, re-prices the idea at the September 2026 list prices you'd actually pay, and ends with a tutorial for doing per-phase routing in your own agent loop.

Benchmark numbers are from the cited paper. The re-pricing and the trajectory in Chart 1 are illustrative, built from public list prices, not measured production traces or customer data.

The unit of routing is wrong for agents.

RouteLLM's finding that most GPT-4 calls didn't need GPT-4 was about single prompts. Agents break that framing. A task-level router looking at "fix the failing date parser in billing/" sees a hard coding task and sends it to the large model, and every one of the 30 calls in that trajectory bills at the large model's rate, including the twelve that were grep and cat.

The other extreme, choosing a model for every individual step, has its own problem. The paper calls it out directly: step-level routers explore an enormous decision space with no structure to guide them, and when a trajectory fails, it's hard to say which of the 30 routing decisions caused it.

Chart 1: One coding-agent task, twelve steps, routed three ways. Task-level to the large model: all 12 steps on the large model. Task-level to the small model: all 12 on the small model, which can't finish the task. Subtask-level: only plan and the two edit steps go to the large model; locate, read, run tests, verify, and submit go to the small model, so 3 of 12 steps bill at the large-model rate.
Chart 1: One coding-agent task, twelve steps, routed three ways. Task-level to the large model: all 12 steps on the large model. Task-level to the small model: all 12 on the small model, which can't finish the task. Subtask-level: only plan and the two edit steps go to the large model; locate, read, run tests, verify, and submit go to the small model, so 3 of 12 steps bill at the large-model rate.

Subtasks sit in the middle. The paper's own example: "code repair may involve locating relevant code, implementing a fix, and verifying the result, with the same subtask recurring within a trajectory." Few enough categories to learn a policy for, specific enough that the small model can own some of them outright.

What RSI-Router does.

The router is learned offline from training trajectories, in a loop the authors call recursive self-improvement. Six iterations, four candidate strategies per iteration, four stages each time:

  1. Subtask mining. Read past trajectories, split them by what each stretch of steps was trying to do, and write a definition plus an identification rule for each subtask. From the second iteration on, existing subtasks get kept, split, or merged.
  2. Routing strategy evolution. Propose four different assignments of subtasks to models: maybe "locate" goes small and "verify" goes large, maybe the reverse.
  3. Model-specific skill evolution. Compare each candidate's trajectories with the large model's on the same tasks, find where the small model failed or wasted steps, and write reusable skills (instructions) that fix those failures for that model.
  4. Pareto selection. Keep the routers that aren't beaten on both validation score and cost.

At run time, the small model reads the user query, recent history, and the previous subtask, predicts which subtask comes next, and the policy picks the model. The pair tested is DeepSeek-V4.1-Flash ($0.30 / $1.20 per million tokens) as the large model and Qwen3.5-9B as the small one.

What it found.

Chart 2: Test-set cost in USD, DeepSeek-V4.1-Flash only versus RSI-router, from arXiv:2609.34712 Table 1. ALFWorld $1.50 to $0.37, minus 75.3%, success 90.6% to 98.4%. ScienceWorld $1.91 to $0.34, minus 82.2%, 34.4% to 35.9%. WebShop $0.75 to $0.19, minus 74.7%, reward 0.59 to 0.61. SWE-bench Verified $3.46 to $3.18, minus 8.1%, 83.3% to 84.4%. Terminal-Bench 2.0 $17.10 to $14.02, minus 18.0%, 40.0% to 46.7%. Average of the five cuts minus 51.7%; cut on the combined bill minus 26.8%.
Chart 2: Test-set cost in USD, DeepSeek-V4.1-Flash only versus RSI-router, from arXiv:2609.34712 Table 1. ALFWorld $1.50 to $0.37, minus 75.3%, success 90.6% to 98.4%. ScienceWorld $1.91 to $0.34, minus 82.2%, 34.4% to 35.9%. WebShop $0.75 to $0.19, minus 74.7%, reward 0.59 to 0.61. SWE-bench Verified $3.46 to $3.18, minus 8.1%, 83.3% to 84.4%. Terminal-Bench 2.0 $17.10 to $14.02, minus 18.0%, 40.0% to 46.7%. Average of the five cuts minus 51.7%; cut on the combined bill minus 26.8%.
BenchmarkLarge onlySmall onlyRSI-routerCost vs large only
ALFWorld90.62% / $1.5065.62% / $0.1398.44% / $0.37−75.3%
ScienceWorld34.38% / $1.9114.58% / $0.0635.94% / $0.34−82.2%
WebShop0.59 / $0.750.48 / $0.020.61 / $0.19−74.7%
SWE-bench Verified83.33% / $3.4662.50% / $0.6984.38% / $3.18−8.1%
Terminal-Bench 2.040.00% / $17.105.33% / $0.1746.67% / $14.02−18.0%

The paper compares against nine baselines: five task-level routers (HybridLLM, FrugalGPT, RouteLLM, GraphRouter, Avengers-Pro), two step-level routers (Router-R1, MTRouter), and both single models. RSI-Router sits on the Pareto frontier against all of them.

Three things in that table matter more than the 51.7%.

  1. The small model never comes close alone. Qwen3.5-9B by itself scores 5.33% on Terminal-Bench, against 40% for DeepSeek. Any task-level router has to send nearly every Terminal-Bench task to the large model. Subtask routing still found 18% of that bill to move. That's spend a per-request router can't see.
  2. The headline average hides the bill. 51.7% is the mean of five percentages. Add up the dollars and the five test sets cost $24.72 on DeepSeek alone and $18.10 routed, a 26.8% cut. The two coding benchmarks are 83% of the spend and got the two smallest cuts. If your agents write code, the number to plan around is closer to 8 to 18% than to 50%.
  3. The test sets are small. 64 tasks each for the three navigation benchmarks, 32 for SWE-bench Verified. One SWE-bench task is 3.1 points, so the 1.05-point gain there is less than one task. Read the score columns as "didn't get worse." The cost columns are the result.

One more detail from the ablations: removing the evolved skills lowered scores on ALFWorld, ScienceWorld, and WebShop by 12.4, 11.5, and 7.4 points, but raised them on SWE-bench Verified and Terminal-Bench by 9.7 and 10.5. Extra instructions that help a small model navigate a text world get in the way on code. Skills are per-domain, not a free add-on.

The part the paper didn't have to price.

Qwen3.5-9B doesn't have a list price in the paper. The authors estimate it from GPU-hours: across six workloads, DeepSeek-V4.1-Flash took 35.96 times the GPU-hours, so they price Qwen at about 1/30th, $0.01 / $0.04 per million tokens. That's a self-hosted 9B model at good utilization. Most teams pay list prices for both tiers, and the gap between tiers is much narrower than 30x.

The cost cut from moving work to a cheaper model is the share of spend you move times the price gap you close:

saving = moved share × (1 − 1 / price ratio)

Backing the moved share out of the paper's own numbers gives about 80% for the three navigation benchmarks and about 13% for the two coding ones. Re-price those at September 2026 pairs:

Chart 3: Savings from the same routed steps at different price ratios. Navigation-heavy tasks, about 80% of spend routable: minus 77% at the paper's 30x, minus 76% for GPT-6 Sol to Luna at 20x, minus 60% for Opus 5.5 to Haiku 4.5 at 4x, minus 40% for Opus 5.5 to Sonnet 5 at 2x. Coding tasks, about 13% routable: minus 13%, 13%, 10%, and 7%.
Chart 3: Savings from the same routed steps at different price ratios. Navigation-heavy tasks, about 80% of spend routable: minus 77% at the paper's 30x, minus 76% for GPT-6 Sol to Luna at 20x, minus 60% for Opus 5.5 to Haiku 4.5 at 4x, minus 40% for Opus 5.5 to Sonnet 5 at 2x. Coding tasks, about 13% routable: minus 13%, 13%, 10%, and 7%.
Model pairPrice ratioNavigation-heavyCoding
Paper (GPU-hour estimate)30x−77%−13%
GPT-6 Sol → Luna ($2/$10 → $0.10/$0.50)20x−76%−13%
Opus 5.5 → Haiku 4.5 ($4/$20 → $1/$5)4x−60%−10%
Opus 5.5 → Sonnet 5 ($4/$20 → $2/$10)2x−40%−7%

Two caveats on this table. It assumes the same steps get routed at every price, and a different small model will be able to own a different set of steps. It also ignores cache effects, which are the next section. But the shape holds: what you can route depends on the task, and what routing is worth depends on the price gap. A navigation-heavy agent on a 20x pair saves three quarters of its bill. A coding agent moving from Opus 5.5 to Sonnet 5 saves single digits, which is less than a cache miss can cost you.

Every switch is a cache decision.

The paper's costs are uncached token prices. Production agents run on prompt caching: turn 20 re-sends turns 1 through 19, and the provider bills that prefix at a cache-read rate instead of the full input rate. Switching models mid-trajectory means the new model has no cache for that prefix. You pay full input price, or a cache write, to rebuild it.

That's the case for routing at the subtask boundary rather than every step. A subtask is a run of consecutive steps on one model. One switch at the start of "locate," several cached steps inside it, one switch back for "edit." A per-step router that flips models on every call can lose more on cache rebuilds than it saves on price, and our cache switch penalty post works through that math in detail. The rule of thumb: only switch when the routed run is long enough to earn back one full-price read of the context.

Tutorial: per-phase routing in your own agent loop.

You don't need the paper's evolutionary loop to get started. The core of it is a policy table from phase to model, a cheap rule for which phase you're in, and per-phase cost data to decide what to move next. Here it is in about 60 lines, through an OpenAI compatible client.

Step 1: name your phases.

Pull a week of trajectories and label the steps. Most coding agents fall into five or six phases, and the tool being called gets you most of the way. That's a cheap version of the paper's identification rules:

PHASE_BY_TOOL = {
    "grep": "locate", "glob": "locate", "list_dir": "locate",
    "read_file": "read",
    "edit_file": "edit", "write_file": "edit", "apply_patch": "edit",
    "run_tests": "verify", "bash": "verify",
}

def current_phase(messages) -> str:
    """Phase of the next call, from the last tool the agent used."""
    for m in reversed(messages):
        for call in m.get("tool_calls") or []:
            return PHASE_BY_TOOL.get(call["function"]["name"], "plan")
    return "plan"  # first turn: no tools yet

It's crude: the step after a read_file is often the edit. Start crude, measure, then refine the rule, which is exactly what the paper's stage 1 does on every iteration.

Step 2: a policy table and a client.

from openai import OpenAI

client = OpenAI(base_url="https://api.getnadir.com/v1", api_key=NADIR_KEY)

SMALL, LARGE = "claude-haiku-4-5", "claude-opus-5-5"

# Start conservative: everything large except phases you've measured.
POLICY = {"plan": LARGE, "locate": SMALL, "read": SMALL,
          "edit": LARGE, "verify": SMALL}

def step(messages, tools, task_id, phase, model):
    return client.chat.completions.create(
        model=model,
        messages=messages,
        tools=tools,
        extra_headers={
            "X-Nadir-Use-Case": "coding-agent",
            "X-Nadir-Tags": f"phase:{phase},task:{task_id}",
        },
    )

The tags are what makes Step 4 possible. Every call carries its phase, so cost per phase is a group-by, not a log-parsing project.

Step 3: the loop, with sticky phases and a failure escape.

def run_agent(task, tools, task_id, max_steps=40):
    messages = [{"role": "user", "content": task}]
    phase = model = None
    small_failures = 0

    for _ in range(max_steps):
        next_phase = current_phase(messages)
        if next_phase != phase:            # switch only at a phase boundary
            phase, model = next_phase, POLICY[next_phase]
            small_failures = 0

        reply = step(messages, tools, task_id, phase, model)
        msg = reply.choices[0].message
        messages.append(msg.model_dump(exclude_none=True))
        if not msg.tool_calls:
            return msg.content

        for call in msg.tool_calls:
            result = execute(call)          # your tool runner
            messages.append({"role": "tool", "tool_call_id": call.id,
                             "content": result.output})
            if model == SMALL and result.failed:
                small_failures += 1

        if small_failures >= 2:             # small model stuck: escalate this phase
            model = LARGE
    raise RuntimeError("step budget exhausted")

Three choices in there are deliberate:

Step 4: decide what to move next.

Once a few hundred tasks have run, compare each phase's cost and outcome:

import pandas as pd

# One row per call, exported from your usage logs:
# task_id, phase, model, cost_usd, task_passed
df = pd.read_csv("agent_calls.csv")

by_phase = (df.groupby(["phase", "model"])
              .agg(calls=("cost_usd", "size"),
                   spend=("cost_usd", "sum"),
                   pass_rate=("task_passed", "mean"))
              .sort_values("spend", ascending=False))
print(by_phase)

The next phase to try on the small model is the one with the most large-model spend that isn't where tasks fail. Move it for a slice of traffic, compare pass rates over the same tasks, keep it if the pass rate holds. That's stage 2 and stage 4 of the paper done by hand, one phase per week instead of four strategies per iteration. When a phase almost works on the small model, write the failure pattern into its system prompt for that phase only. That's stage 3, and the ablation says to check it per domain rather than assume it helps.

When not to bother.

Where Nadir fits.

Nadir doesn't learn your phase policy for you. Your agent knows which tool it just called; a gateway sitting outside the loop doesn't, and the subtask rules in the paper come from your own trajectories. What Nadir provides is the part of this that's tedious to build and needs to sit in the request path:

If you run agents that make more than a handful of calls per task, start with a free key, add the phase tag to every call for a week, and look at which phases carry your large-model spend. That table tells you whether subtask routing is worth 5% or 50% for your workload, before you write the rest of it.

Conclusion.

Per-request routing treats an agent task as one indivisible decision: hard task, large model, every call. RSI-Router shows that a hard task is mostly easy steps. With a 9B model owning the steps it can handle, the paper cut a DeepSeek agent's cost by 75 to 82% on navigation-style benchmarks and 8 to 18% on coding, with scores no worse on any of the five. Two corrections before you plan around it: the dollar-weighted cut across the five test sets is 26.8%, not 51.7%, because coding is where the money is, and the paper's 30x price gap is a self-hosting estimate. At list prices, the same routed steps are worth 60% on a navigation agent moving from Opus 5.5 to Haiku 4.5, and about 10% on a coding agent. Route by phase, switch only at phase boundaries, tag every call, and let your own per-phase numbers say how far to take it.


Benchmark results are from the cited paper. Re-priced savings are illustrative, computed from the paper's reported costs and public September 2026 list prices, and are not derived from customer data. Sources: [Li, Zhang, Cui, Mu, Zhang, Zhang, Jia, and Hu, "RSI-Router: Evolving Subtask-Level LLM Routing and Skills for Cost-Efficient Agents," arXiv:2609.34712, September 2026](https://arxiv.org/abs/2609.34712). [Anthropic, Pricing](https://platform.claude.com/docs/en/about-claude/pricing). [OpenAI, API Pricing](https://openai.com/api/pricing/).

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

For coding agents, Nadir connects to Codex, Claude Code, or Cursor and recommends a model tier for delegated work. The agent decides whether to hand the task off and uses its own configured models, so Nadir needs no proxy and no provider keys on that path.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.