The Stopping Problem

Every ReAct step re-sends the full history, so in a simulated research agent step 12 costs 8.5x step 1 while being the last needed step for under 2% of tasks. Across 50,000 synthetic tasks, a framework-default cap of 25 spent 32% of its budget on the 6% of tasks that can never resolve and another 20% on steps after the answer was available. Fixed caps cut spend mostly by failing tasks, and the cheapest cap per resolved task resolved only 51% of them. A stagnation-aware stop cut cost per resolved task 43% against the default and 27% against the cap with the same resolution rate, at 1.8 points less resolution. All numbers are illustrative simulation output with stated assumptions, not production data. Includes a short novelty-based stop function and the seven metrics to log first.

Published 2026-10-05 by Dor Amir on the Nadir blog.

Filed under Agents.

Abstract.

Agent cost audits usually ask where tokens go inside a step. This post asks when the agent should stop taking steps. In a simulated ReAct research agent, each step re-sends the whole history, so step 12 costs 8.5x step 1 while being the last needed step for under 2% of tasks. We compare a framework-default cap of 25, fixed caps of 6, 10, and 15, and a stagnation-aware stop across 50,000 synthetic tasks. Tight caps cut spend by failing tasks. The stagnation stop cut cost per resolved task 43% against the default and 27% against the cap that matched its success rate, giving up 1.8 points of resolution. The data is illustrative, not measured; the mechanism is the finding.

Everything below comes from a simulation with the assumptions listed in the methodology. It is not production data, and the percentages will move with your workload. The code path to measure your own is in the implementation section.

The question.

A ReAct loop (reason, act, observe, repeat; Yao et al., arXiv:2210.03629) has no natural price ceiling. Every call re-sends the system prompt, the tool schemas, and everything the agent has read so far. We have shown where tokens go inside one step, and the O(n^2) shape of the sum. The question left open is the control one:

Which stopping rule minimizes cost per resolved task, and what does each rule give up in resolution rate?

Methodology.

The simulator is about 120 lines of Python (numpy, pandas, seaborn). Every number is an assumption, so change them.

A task counts as resolved if the agent took at least as many steps as it needed and was not falsely stopped.

Result 1: retrieved content is the bill, and it compounds.

Chart 1: Input tokens per call by stage, over 12 steps. Retrieved search results and document reads are 72% of input by step 5 and 81% by step 12
Chart 1: Input tokens per call by stage, over 12 steps. Retrieved search results and document reads are 72% of input by step 5 and 81% by step 12

The fixed prefix is flat. Everything else grows linearly with step count, and almost all of it is retrieved content. By step 5, 72% of input is search results and document reads. Over-reading and duplicate retrieval are expensive because every later step pays for them again. Caching the prefix does not touch this part.

Result 2: late steps cost more and finish less.

Chart 2: Step cost rises from $0.027 to $0.30 by step 16 while the share of tasks whose last needed step is k falls below 2% after step 8
Chart 2: Step cost rises from $0.027 to $0.30 by step 16 while the share of tasks whose last needed step is k falls below 2% after step 8

Step 1 costs $0.027. Step 12 costs $0.229. Under our task mix, steps 9 and beyond are each the last needed step for about 1.7 to 1.9% of tasks. The shape of that mix is an assumption, but any workload with a hard tail has the same crossing point; only its location moves. Compare it to the marginal-return analysis within a step: this is the same logic across steps.

Result 3: the default policy spends a third of its budget on tasks that cannot succeed.

Chart 3: Cost per task by policy, log scale. The framework default reaches a $6.69 P99; the stagnation stop reaches $2.97
Chart 3: Cost per task by policy, log scale. The framework default reaches a $6.69 P99; the stagnation stop reaches $2.97
PolicyResolvedMean costP99 costCost per resolved taskMedian latency
Framework default (cap 25)94.0%$1.235$6.69$1.31382 s
Cap 665.6%$0.467$0.59$0.71282 s
Cap 1083.5%$0.728$1.32$0.87282 s
Cap 1592.3%$0.953$2.65$1.03382 s
Stagnation stop92.2%$0.692$2.97$0.75071 s

The 6% of tasks that can never be resolved produce 32% of default spend, because each runs to the cap with a full context. Another 20% goes to overrun on tasks that were already answerable. Only 48% of the default policy's spend is steps the tasks needed.

Chart 4: Cost per resolved task by step cap. The minimum sits at cap 4, where only 51% of tasks resolve
Chart 4: Cost per resolved task by step cap. The minimum sits at cap 4, where only 51% of tasks resolve

The cap sweep contains a trap. Cost per resolved task is lowest at cap 4, but that cap resolves 51% of tasks. Failed tasks leave the denominator, so the metric rewards failing. Any dashboard that reports cost per resolved task must report the resolution rate beside it.

Chart 5: Spend decomposition by policy, indexed to the framework default. The stagnation stop totals 56, with 82% of that being needed steps
Chart 5: Spend decomposition by policy, indexed to the framework default. The stagnation stop totals 56, with 82% of that being needed steps

Discussion.

A cap and a stop rule are different tools. A cap prices every task the same way. It saves money mostly by truncating hard tasks (the pink segment in Chart 5), so the saving is paid back in failures or retries. A retry that re-runs the task can erase the saving entirely. The stagnation rule spends its savings on steps that were not producing evidence, which is why it kept 92.2% resolution at $0.750 where cap 15 needed $1.033 for the same rate.

The stop rule has its own failure mode. The 5% false-stop rate is an assumption, and it is the number to measure first. Stopping early on a task that was one step from resolving is worse than a cap, because it looks like a confident answer.

Caching shrinks the bill but not the argument. With cache reads at 10% of input price and writes at 125%, cost per resolved task falls to $0.445 for the default and $0.318 for the stagnation stop, 29% apart. See Anthropic's prompt caching documentation for the pricing model. Caching helps only if the prefix repeats.

Latency follows steps. Median latency fell 14% because fewer steps run in sequence. Parallel tool calls attack a different term and can raise token load.

Long contexts are not free of quality risk. Liu et al. found that models use information in the middle of long contexts less reliably. A late step is both more expensive and, plausibly, working from a noisier prompt. We did not model that, so the simulation understates the case for stopping.

Implementation guide.

  1. Tag every call with task id and step index. Without step index you cannot draw Chart 2 for your own traffic.
  2. Log tokens by stage: prefix, history, new tool output, output. Retrieved content is usually the largest.
  3. Compute a novelty score per step, for example the share of retrieved chunk ids not already in context.
  4. Stop after two low-novelty steps. Tune the window on a sample where you know the answer.
  5. Keep a hard cap as a backstop, set near your P95 of needed steps, not at the framework default.
  6. On stop, escalate or return a partial answer with a flag. Do not silently fail.
  7. Track cost per task, resolution rate, and P99 cost together. Any one alone can be gamed.
def should_stop(step, new_chunk_ids, seen_ids, history, cap=15, window=2):
    novelty = len(set(new_chunk_ids) - seen_ids) / max(len(new_chunk_ids), 1)
    history.append(novelty)
    seen_ids.update(new_chunk_ids)
    if step >= cap:
        return "cap"
    if len(history) >= window and all(n < 0.2 for n in history[-window:]):
        return "stagnant"
    return None

Where routing fits.

A stop rule decides how many steps run. Routing decides which model runs each one. They compound: once the agent runs fewer steps, per-subtask routing sends the cheap phases (search, read, filter) to a smaller model instead of treating every step as an Opus-level problem. Nadir is an efficiency layer built for that second part: it records cost and routing decisions per request, so the per-step data this post depends on is already in your logs, and per-task attribution shows which workflows are doing the long loops. Whether it saves you anything depends on how many of your steps are simple.

Conclusion.

The future of LLM cost optimization is not just cheaper models. It is better orchestration, better routing, better retrieval discipline, and better measurement. In this model, the cheapest token is the one in a step that did not need to happen, and the only way to find those steps is to log them. Choose a stop rule by its cost per resolved task and its resolution rate together.

If you run agents in production, benchmark your own workflows: tag step index for a week and see how much spend sits past your P95 of needed steps.


Illustrative simulation: 50,000 synthetic tasks, fixed seed, assumptions listed in the methodology. Not measured production traces. Sources: [Yao et al., ReAct, arXiv:2210.03629](https://arxiv.org/abs/2210.03629); [Liu et al., Lost in the Middle, arXiv:2307.03172](https://arxiv.org/abs/2307.03172); [Anthropic, prompt caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching).

More on agents

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

For coding agents, Nadir Auto is a terminal launcher for Claude Code or Codex that asks Nadir for the main-session model before each turn, with your configured model as the fallback and cost ceiling. Delegation integrations for Claude Code, Codex, or Cursor instead recommend a model tier for subagent work, and the agent decides. Either way inference runs on your own provider account, so Nadir holds no provider keys on that path.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.