Abstract.
Agent cost audits usually ask where tokens go inside a step. This post asks when the agent should stop taking steps. In a simulated ReAct research agent, each step re-sends the whole history, so step 12 costs 8.5x step 1 while being the last needed step for under 2% of tasks. We compare a framework-default cap of 25, fixed caps of 6, 10, and 15, and a stagnation-aware stop across 50,000 synthetic tasks. Tight caps cut spend by failing tasks. The stagnation stop cut cost per resolved task 43% against the default and 27% against the cap that matched its success rate, giving up 1.8 points of resolution. The data is illustrative, not measured; the mechanism is the finding.
Everything below comes from a simulation with the assumptions listed in the methodology. It is not production data, and the percentages will move with your workload. The code path to measure your own is in the implementation section.
The question.
A ReAct loop (reason, act, observe, repeat; Yao et al., arXiv:2210.03629) has no natural price ceiling. Every call re-sends the system prompt, the tool schemas, and everything the agent has read so far. We have shown where tokens go inside one step, and the O(n^2) shape of the sum. The question left open is the control one:
Which stopping rule minimizes cost per resolved task, and what does each rule give up in resolution rate?
Methodology.
The simulator is about 120 lines of Python (numpy, pandas, seaborn). Every number is an assumption, so change them.
- Per call: 3,000 prefix tokens (900 system prompt, 2,100 tool schemas). Each step adds 480 output tokens (reasoning plus query) and 3,200 tokens of tool output (1,800 search results, 1,400 document reads) to history.
- Price: $5 per million input, $25 per million output, an Opus-class list price. No caching in the main results. A cached variant is below.
- Tasks: 6% are unresolvable (evidence never arrives; the agent plateaus after 3 to 6 steps). Of the rest, 55% need 2 to 4 steps, 30% need 5 to 8, and 15% need 9 to 16.
- Overrun: a default agent is unsure it is done and takes a geometric number of extra steps (mean 1.5). A stagnation-aware agent averages 0.2.
- Stagnation stop: halt after two consecutive steps with no new evidence. It catches 90% of unresolvable loops, and falsely stops 5% of tasks needing 6 or more steps.
- Latency: 0.5 s call overhead, 25K tokens/s prefill, 60 tokens/s decode, 2 s per tool call, sequential.
A task counts as resolved if the agent took at least as many steps as it needed and was not falsely stopped.
Result 1: retrieved content is the bill, and it compounds.

The fixed prefix is flat. Everything else grows linearly with step count, and almost all of it is retrieved content. By step 5, 72% of input is search results and document reads. Over-reading and duplicate retrieval are expensive because every later step pays for them again. Caching the prefix does not touch this part.
Result 2: late steps cost more and finish less.

Step 1 costs $0.027. Step 12 costs $0.229. Under our task mix, steps 9 and beyond are each the last needed step for about 1.7 to 1.9% of tasks. The shape of that mix is an assumption, but any workload with a hard tail has the same crossing point; only its location moves. Compare it to the marginal-return analysis within a step: this is the same logic across steps.
Result 3: the default policy spends a third of its budget on tasks that cannot succeed.

| Policy | Resolved | Mean cost | P99 cost | Cost per resolved task | Median latency |
|---|---|---|---|---|---|
| Framework default (cap 25) | 94.0% | $1.235 | $6.69 | $1.313 | 82 s |
| Cap 6 | 65.6% | $0.467 | $0.59 | $0.712 | 82 s |
| Cap 10 | 83.5% | $0.728 | $1.32 | $0.872 | 82 s |
| Cap 15 | 92.3% | $0.953 | $2.65 | $1.033 | 82 s |
| Stagnation stop | 92.2% | $0.692 | $2.97 | $0.750 | 71 s |
The 6% of tasks that can never be resolved produce 32% of default spend, because each runs to the cap with a full context. Another 20% goes to overrun on tasks that were already answerable. Only 48% of the default policy's spend is steps the tasks needed.

The cap sweep contains a trap. Cost per resolved task is lowest at cap 4, but that cap resolves 51% of tasks. Failed tasks leave the denominator, so the metric rewards failing. Any dashboard that reports cost per resolved task must report the resolution rate beside it.

Discussion.
A cap and a stop rule are different tools. A cap prices every task the same way. It saves money mostly by truncating hard tasks (the pink segment in Chart 5), so the saving is paid back in failures or retries. A retry that re-runs the task can erase the saving entirely. The stagnation rule spends its savings on steps that were not producing evidence, which is why it kept 92.2% resolution at $0.750 where cap 15 needed $1.033 for the same rate.
The stop rule has its own failure mode. The 5% false-stop rate is an assumption, and it is the number to measure first. Stopping early on a task that was one step from resolving is worse than a cap, because it looks like a confident answer.
Caching shrinks the bill but not the argument. With cache reads at 10% of input price and writes at 125%, cost per resolved task falls to $0.445 for the default and $0.318 for the stagnation stop, 29% apart. See Anthropic's prompt caching documentation for the pricing model. Caching helps only if the prefix repeats.
Latency follows steps. Median latency fell 14% because fewer steps run in sequence. Parallel tool calls attack a different term and can raise token load.
Long contexts are not free of quality risk. Liu et al. found that models use information in the middle of long contexts less reliably. A late step is both more expensive and, plausibly, working from a noisier prompt. We did not model that, so the simulation understates the case for stopping.
Implementation guide.
- Tag every call with task id and step index. Without step index you cannot draw Chart 2 for your own traffic.
- Log tokens by stage: prefix, history, new tool output, output. Retrieved content is usually the largest.
- Compute a novelty score per step, for example the share of retrieved chunk ids not already in context.
- Stop after two low-novelty steps. Tune the window on a sample where you know the answer.
- Keep a hard cap as a backstop, set near your P95 of needed steps, not at the framework default.
- On stop, escalate or return a partial answer with a flag. Do not silently fail.
- Track cost per task, resolution rate, and P99 cost together. Any one alone can be gamed.
def should_stop(step, new_chunk_ids, seen_ids, history, cap=15, window=2):
novelty = len(set(new_chunk_ids) - seen_ids) / max(len(new_chunk_ids), 1)
history.append(novelty)
seen_ids.update(new_chunk_ids)
if step >= cap:
return "cap"
if len(history) >= window and all(n < 0.2 for n in history[-window:]):
return "stagnant"
return None
Where routing fits.
A stop rule decides how many steps run. Routing decides which model runs each one. They compound: once the agent runs fewer steps, per-subtask routing sends the cheap phases (search, read, filter) to a smaller model instead of treating every step as an Opus-level problem. Nadir is an efficiency layer built for that second part: it records cost and routing decisions per request, so the per-step data this post depends on is already in your logs, and per-task attribution shows which workflows are doing the long loops. Whether it saves you anything depends on how many of your steps are simple.
Conclusion.
The future of LLM cost optimization is not just cheaper models. It is better orchestration, better routing, better retrieval discipline, and better measurement. In this model, the cheapest token is the one in a step that did not need to happen, and the only way to find those steps is to log them. Choose a stop rule by its cost per resolved task and its resolution rate together.
If you run agents in production, benchmark your own workflows: tag step index for a week and see how much spend sits past your P95 of needed steps.
Illustrative simulation: 50,000 synthetic tasks, fixed seed, assumptions listed in the methodology. Not measured production traces. Sources: [Yao et al., ReAct, arXiv:2210.03629](https://arxiv.org/abs/2210.03629); [Liu et al., Lost in the Middle, arXiv:2307.03172](https://arxiv.org/abs/2307.03172); [Anthropic, prompt caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching).