No Exit

A July 2026 paper confirmed 68 agent loops with no enforced exit across 6,549 GitHub repositories, and 95.6% of them can burn through an API budget. On April 29, 2026, a single idle timeout in the open-source agent runtime OpenClaw produced 1,384 paid model calls in 60 seconds, an estimated $20-30, before its own five-call circuit breaker shipped a week later and cut the worst case to $0.10-0.30. The paper, IAL-Scan (arXiv:2607.01641), covers 8 frameworks and finds those 68 bugs in 47 projects, concentrated in LangGraph and AutoGen. This post walks through the paper, the OpenClaw incident it would have caught, and an illustrative fleet-scale simulation: a mean monthly exposure of $1,231 across 5,000 deployments without a hard cap, versus $12.96 with one, a 95x gap. Includes a five-minute audit checklist from the paper's own root-cause taxonomy.

Published 2026-09-28 by Dor Amir on the Nadir blog.

Filed under Agents.

Abstract.

"An AI agent can fail in a very expensive way: it can keep trying." That line comes from a FinOps vendor's write-up of agent retry costs, and it understates the real incident it cites. On April 29, 2026, a single idle timeout in the open-source agent runtime OpenClaw produced 1,384 paid model calls in 60 seconds, all timestamped to the same minute, costing an estimated $20-30 before anyone noticed. Source: OpenClaw, GitHub issue #76293. Four days earlier, the same bug had done it once already: 761 calls in 60 seconds. The fix, shipped a week later, capped consecutive failures at five and cut the worst case to about $0.10-0.30 per incident. Source: OpenClaw, PR #76345. This is not a pricing story or a routing story. It is a defect class: a feedback path in agent code with no bound on how many times it can fire. A July 2026 static-analysis paper, IAL-Scan (Hou, Wang, Zhao, and Wang, arXiv:2607.01641), scanned 6,549 real GitHub agent repositories across eight frameworks and confirmed 68 of these bugs in 47 projects, 95.6% of which cause "API cost exhaustion." This post walks through that paper, the OpenClaw incident it would have caught, and an illustrative simulation of what an unpatched feedback path costs a fleet of agent deployments over a month: a mean exposure of $1,231, against $12.96 once the same deployments carry a five-call circuit breaker, a 95x gap.

All simulation figures in this post are illustrative, modeled to anchor to the cited paper and incident data, not measured production traces. Not derived from proprietary customer data. Sources cited throughout.

The research question.

Most of the token waste this blog documents is a design tradeoff: multi-agent orchestration spends 15x more tokens than a single chat because fan-out buys better answers. The retry tax is an economic argument: a router that blind-retries on a cheap tier before escalating pays for every failed attempt, and you can write down a break-even formula for when that stops being worth it. Both assume the system is working as designed and ask whether the design is a good deal.

This post asks a different question: what happens when a piece of agent code has a feedback path, a retry, a tool call, a graph edge, a multi-agent handoff, that nothing bounds at all? The Blast Radius Bill covered what an agent can do to your infrastructure when it goes wrong; Denial of Wallet Is the New DDoS named the financial version of that outcome. This post is about the narrower, more common cause sitting underneath a lot of both: not a bad decision, a missing stop condition. IAL-Scan is the first paper we've found that measured how often that cause is sitting in your dependency tree right now.

What IAL-Scan found.

Chart 1: Bar chart of 68 confirmed unbounded-loop failures by agent framework from IAL-Scan (arXiv:2607.01641): LangGraph 23 (33.8%), AutoGen AgentChat 22 (32.4%), LlamaIndex 6 (8.8%), LangChain AgentExecutor 5 (7.4%), CrewAI 4 (5.9%), OpenAI Agents SDK 4 (5.9%), Google ADK 3 (4.4%), Semantic Kernel 1 (1.5%). Second panel shows share of those 68 failures causing each impact: API cost exhaustion 95.6%, model denial of service 95.6%, context window exhaustion 27.9%, tool rate-limit exhaustion 7.4%.
Chart 1: Bar chart of 68 confirmed unbounded-loop failures by agent framework from IAL-Scan (arXiv:2607.01641): LangGraph 23 (33.8%), AutoGen AgentChat 22 (32.4%), LlamaIndex 6 (8.8%), LangChain AgentExecutor 5 (7.4%), CrewAI 4 (5.9%), OpenAI Agents SDK 4 (5.9%), Google ADK 3 (4.4%), Semantic Kernel 1 (1.5%). Second panel shows share of those 68 failures causing each impact: API cost exhaustion 95.6%, model denial of service 95.6%, context window exhaustion 27.9%, tool rate-limit exhaustion 7.4%.

Xinyi Hou, Shenao Wang, Yanjie Zhao, and Haoyu Wang, all at Huazhong University of Science and Technology, define an Infinite Agentic Loop (IAL) as a feedback path, a retry, a tool-call iteration, a workflow cycle, or an agent handoff, that "repeatedly triggers costly or state-growing actions" without an effective bound. Their tool, IAL-Scan, abstracts eight popular frameworks (LangChain, LangGraph, CrewAI, AutoGen, LlamaIndex, OpenAI Agents SDK, Google ADK, Semantic Kernel) into a framework-independent intermediate representation, builds an "Agentic Loop Dependence Graph" for each project, and checks every feedback path for a bound that actually holds at runtime, not just a max_iterations argument that a nested call can bypass.

They ran it against 6,549 GitHub repositories with at least one star, 246,748 Python files, 33.41 million lines of code, and reported 74 candidate findings. After manual verification: 68 confirmed, 6 false positives, 91.9% precision, spread across 47 distinct projects.

Root cause (of 68 confirmed failures)Share
Missing strong bound (no cap at all)100.0%
Tool-controlled retry (the tool decides when to stop)41.2%
Model-controlled termination (the LLM decides when to stop)38.2%
Missing exit condition in a loop or graph cycle33.8%
Workflow cycle without a verified bound30.9%
State growth amplifier (context/history grows every pass)27.9%
Agent tool reentry25.0%

Two patterns explain most of the concentration in LangGraph and AutoGen. LangGraph's conditional edges let a developer write a cycle that assumes the model will eventually route to an exit node; when it doesn't, the graph has no independent counter to fall back on. AutoGen's multi-agent chat loops similarly assume the conversation will converge; a max_turns argument exists, but the paper found it frequently unset or set on the wrong object in the call chain. Both are "model-controlled termination" by another name: the bound is a hope about model behavior, not a number enforced by the runtime.

The incident IAL-Scan's own pattern would have caught.

OpenClaw is a real, actively developed open-source coding-agent runtime. Its bug tracker documents exactly the failure mode the paper describes, with a timestamp. On April 25, 2026 at 13:02 ET, a single LLM idle timeout (60 seconds, no response from the model) propagated up through OpenClaw's retry layer, which re-invoked the same stream function against the same wedged provider connection. That retry also timed out. The cycle repeated, unbounded, 761 times, all logged in the same 60-second window. Four days later it happened again: 1,384 calls in 60 seconds. Source: OpenClaw, GitHub issue #76293. The bug report is explicit about the root cause: no maximum retry count, no exponential backoff, no circuit breaker. Auto-recharge on the provider account meant nobody saw the spike until the bill arrived.

This is IAL-Scan's "retry feedback without bound" category, the single largest bucket in the paper at 25.0% of confirmed failures. The fix that shipped a week later is almost embarrassingly small: cap consecutive idle timeouts at five, reset the counter on any successful chunk so a slow-but-responsive stream isn't penalized, then lock out. Worst case after the fix: 5 paid calls instead of up to 1,384. Source: OpenClaw, PR #76345.

Methodology.

Two things in this post are cited fact: IAL-Scan's scan results and OpenClaw's incident numbers. One thing is an illustrative simulation, built to show what those two facts imply at fleet scale, not a measurement of anyone's production traffic.

ParameterValueNotes
Fleet size5,000 production agent deploymentsIllustrative platform scale
Share exposed to ≥ 1 unbounded feedback path0.72%IAL-Scan's repo-level rate, 47/6,549, extended here as an illustrative per-deployment rate. Not a claim about production incident frequency.
Incidents per exposed deployment per monthPoisson, mean 1.8Assumed; real triggers are provider idle timeouts, tool rate limits, and validator-rejection loops
Cost per incident, uncappedUniform $8–$30Anchored to OpenClaw's two documented incidents ($20-30 for 761-1,384 calls in 60s); lower bound assumes a smaller loop caught sooner
Cost per incident, capped at 5 consecutive failuresUniform $0.10–$0.30OpenClaw's own post-fix worst case
Simulation10,000 Monte Carlo trials, fixed seed

Finding 1: the bug clusters in exactly two frameworks, and the impact is almost always the same.

LangGraph and AutoGen together account for 45 of the 68 confirmed failures, 66% of the total, despite being 2 of 8 frameworks scanned. That's not because they're worse-engineered; it's because cyclic graphs and multi-turn multi-agent chat are the two abstractions most likely to need a bound that isn't "the model will figure it out." Every other framework in the study had at least one confirmed failure too. And whichever framework it happened in, the impact was almost never ambiguous: 95.6% of the 68 confirmed failures caused API cost exhaustion, and the same 95.6% caused what the paper calls model denial of service, the agent itself becoming unusable mid-loop. Context window exhaustion (27.9%) and external tool rate-limit exhaustion (7.4%) trail well behind. If your agent has an unbounded feedback path, the first thing that breaks is almost always the bill, not the context window.

Finding 2: two different loops fail at two different speeds.

Chart 2: Log-log line chart comparing two unbounded-loop cost curves. Growing-context loop (each retry resends the full conversation history, reconstructed from bex.co's reported multipliers) reaches 3.2x cost at 5 steps, 26x at 50 steps ($4.85 vs $0.19 baseline), and 100x at 200 steps ($76.38). Flat retry loop (same request replayed with no growing history, OpenClaw's idle-timeout bug) reaches $13.75 at 761 calls and $25.00 at 1,384 calls in 60 seconds, both real incident numbers from GitHub issue #76293.
Chart 2: Log-log line chart comparing two unbounded-loop cost curves. Growing-context loop (each retry resends the full conversation history, reconstructed from bex.co's reported multipliers) reaches 3.2x cost at 5 steps, 26x at 50 steps ($4.85 vs $0.19 baseline), and 100x at 200 steps ($76.38). Flat retry loop (same request replayed with no growing history, OpenClaw's idle-timeout bug) reaches $13.75 at 761 calls and $25.00 at 1,384 calls in 60 seconds, both real incident numbers from GitHub issue #76293.

Not every unbounded loop costs the same per iteration. OpenClaw's bug replays an identical, small request: cost grows linearly, about $0.018 per call. A different, more common loop shape resends the entire growing conversation history on every retry, because that's how stateless chat APIs work: cost grows quadratically. A July 2026 analysis of this second pattern, run by the open-source PaaS Bex against Claude Sonnet 5 pricing, reported multipliers of 3.2x at 5 steps, 26x at 50 steps (a $5.01 actual bill against a $0.19 baseline of only the tokens actually needed), and 100x at 200 steps. Source: Dora Noda, "Your AI Agent's Debug Loop Costs Grow Quadratically, Not Linearly," Bex, July 31, 2026. The two curves cross around 8-10 iterations: below that, a flat retry loop is actually more expensive per step; above it, the growing-context loop pulls away fast, because every step now carries the cost of every step before it. IAL-Scan's "state growth amplifier" category (27.9% of confirmed failures) is this shape specifically, and it's the one where catching the loop late costs disproportionately more than catching it early.

Finding 3: the fix already exists, and it's nearly free.

Chart 3: Bar chart on a log scale comparing OpenClaw's real cost per incident before and after its retry-loop fix. Before fix (unbounded retry on idle timeout): $20-30 per incident, from 761-1,384 calls in 60 seconds. After fix (hard cap of 5 consecutive failures, then lockout): $0.10-0.30 per incident. Approximately 100x reduction, from GitHub issue #76293 and PR #76345, May 2026.
Chart 3: Bar chart on a log scale comparing OpenClaw's real cost per incident before and after its retry-loop fix. Before fix (unbounded retry on idle timeout): $20-30 per incident, from 761-1,384 calls in 60 seconds. After fix (hard cap of 5 consecutive failures, then lockout): $0.10-0.30 per incident. Approximately 100x reduction, from GitHub issue #76293 and PR #76345, May 2026.

OpenClaw's own before-and-after numbers are the cleanest evidence in this post, because they're not a simulation: the same bug, the same codebase, measured twice. Before the fix: $20-30 per incident. After: $0.10-0.30. A roughly 100x reduction, achieved with a counter and a threshold, not a smarter model or a cheaper one. The counter resets on any successful chunk, so a slow-but-alive stream isn't punished the way a truly wedged one is. This is the general shape of every fix IAL-Scan's authors recommend: replace an assumed bound (the model will stop, the conversation will converge) with an enforced one (a number, checked by the runtime, not by the thing that's failing).

Finding 4: at fleet scale, this is a monthly tax, not a one-off surprise.

Chart 4: Two histograms from a 10,000-trial Monte Carlo simulation of fleet-wide monthly cost from unbounded agent loops, 5,000 deployments, 0.72% exposed. Without a hard cap: mean $1,231/month, P99 $1,882/month. With a hard cap of 5 consecutive failures: mean $12.96/month, P99 $19.77/month. Both the mean and P99 drop approximately 95x.
Chart 4: Two histograms from a 10,000-trial Monte Carlo simulation of fleet-wide monthly cost from unbounded agent loops, 5,000 deployments, 0.72% exposed. Without a hard cap: mean $1,231/month, P99 $1,882/month. With a hard cap of 5 consecutive failures: mean $12.96/month, P99 $19.77/month. Both the mean and P99 drop approximately 95x.

Any single incident, $20 or $30, looks like rounding error against a real AI budget. The reason it isn't: IAL-Scan's 0.72% prevalence rate, applied to a fleet of 5,000 agent deployments, means roughly 36 of them are carrying an unbounded feedback path at any given time, each capable of firing more than once a month. In the illustrative simulation, that puts mean fleet-wide monthly exposure at $1,231 without a cap, with a 99th-percentile month at $1,882. Add the same five-call circuit breaker OpenClaw shipped, and the mean drops to $12.96, the P99 to $19.77, both roughly a 95x reduction. Nobody sees the $1,231; they see forty small, unexplained line items scattered across the month's invoice, each one plausible on its own, none of them investigated, because $25 isn't worth a ticket.

A five-minute audit, from IAL-Scan's own taxonomy.

In order of how often the paper found them: grep for retry logic without a hard numeric cap ("retry feedback without bound," 25.0% of failures — if the exit condition is "success" and nothing else, it's exposed); check every LangGraph conditional edge for an independent counter rather than a routing condition that assumes the model eventually picks the exit node ("model-controlled termination," 38.2%); confirm max_turns or its equivalent is actually set and reaches the object that enforces it in AutoGen-style multi-agent chats, since the paper found it frequently declared but not wired through; and look for state, conversation history, accumulated tool output, that grows every pass with no ceiling, the "state growth amplifier" pattern that turns a cheap flat retry into a quadratic one. Then add a cost or call-count circuit breaker as the last line regardless, on the assumption you missed one: cap consecutive failures, reset on success, lock out past the cap, the exact shape of OpenClaw's fix.

When not to bother.

Where Nadir fits.

Nadir doesn't watch for infinite loops inside your agent framework; that's a code-level fix, and IAL-Scan's authors are right that it belongs in the framework or in your own retry wrappers, not bolted on downstream. What Nadir does provide is the visibility that turns a $25 anomaly from invisible into obvious: every response routed through the OpenAI-compatible gateway reports the model and the exact cost that served it, so a deployment suddenly making hundreds of identical calls in a minute shows up as a spike against its own baseline, not as forty unexplained line items scattered across an invoice. That's the same per-response cost reporting behind chargeback and showback: the fastest way to catch a runaway feedback path is to already know what every deployment normally costs, so the abnormal minute is visible the day it happens, not the day the bill is due.

If you run agents built on LangGraph or AutoGen in production, start with a free key, route a week of traffic through it, and check whether any single deployment's per-minute call count ever spikes two orders of magnitude above its own median. That spike is what an unbounded feedback path looks like before it becomes a support ticket.

Conclusion.

Most of the token waste on this blog is a system doing what it was built to do, just inefficiently: too much context, the wrong model, a cache miss. Infinite agentic loops are different: a piece of code with no exit condition, waiting for a trigger that IAL-Scan found in 68 real, confirmed cases across 47 production repositories, 95.6% of which burned API budget when it fired. OpenClaw's own incident and fix are the cleanest before-and-after available: $20-30 per incident down to $0.10-0.30, a 100x reduction, from a five-line counter. At the scale of a real deployment fleet, in this post's illustrative simulation, that's the difference between a mean monthly exposure of $1,231 and $12.96. The fix isn't a smarter model, a cheaper one, or a better router. It's a number, enforced by the runtime, that replaces a hope about what the model or the conversation will eventually do.


Simulation figures in this post are illustrative, built to anchor to the cited paper's repo-level prevalence rate and OpenClaw's own documented incident and fix costs, and are not derived from customer data. Sources: [Hou, Wang, Zhao, and Wang, "When Agents Do Not Stop: Uncovering Infinite Agentic Loops in LLM Agents," arXiv:2607.01641, July 2026](https://arxiv.org/abs/2607.01641). [OpenClaw, GitHub issue #76293](https://github.com/openclaw/openclaw/issues/76293). [OpenClaw, GitHub PR #76345](https://github.com/openclaw/openclaw/pull/76345). [Larridin, "Are Agent Retry Loops Silently Consuming Your AI Budget?", August 28, 2026](https://larridin.com/blog/ai-agent-retry-cost-control?hs_amp=true). [Dora Noda, "Your AI Agent's Debug Loop Costs Grow Quadratically, Not Linearly," Bex, July 31, 2026](https://bex.co/blog/2026/07/31/agent-retry-loop-quadratic-token-cost).

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.