Denial of Wallet Is the New DDoS

In January 2026, researchers built an MCP server that inflates a task's cost by up to 658x without ever failing the task. OWASP renamed the risk 'Unbounded Consumption' in its 2025 Top 10 for LLM Applications. Real incidents already range from $46,000 a day to $500 million a month. Rate limiters and WAFs don't see any of it, because none of them count in dollars. Here is what does.

Published 2026-07-06 by Dor Amir on the Nadir blog.

Filed under FinOps & Governance.

The attack that doesn't crash anything, it just spends your money.

Denial of service takes a system offline until someone notices and responds. Denial of wallet does the opposite: the system stays up, every health check reads green, and the only thing moving is a number nobody signed off on. "Model denial of service keeps the server running perfectly while draining the cloud budget through excessive token consumption." Source: toxsec, "Model Denial of Service Turns Your Cloud Bill Into a Weapon"

The industry now has real figures attached to this, not hypotheticals. Sysdig's threat research team tracked attackers using stolen cloud credentials to run large language model calls on victims' own AWS Bedrock accounts, a pattern it named LLMjacking, at costs it estimated up to $46,000 a day per compromised account. Source: Sysdig, "LLMjacking: Stolen Cloud Credentials Used in New AI Attack". In March 2026, a single stolen Google Gemini API key reportedly ran up $82,000 in charges in 48 hours. Source: toxsec, "Model Denial of Service Turns Your Cloud Bill Into a Weapon". And in May 2026, an enterprise client burned through $500 million on Anthropic's Claude in a single month with no attacker involved at all, just thousands of employees with unrestricted access and no per-user limits. Source: Yahoo Finance, "Client Accidentally Burns $500 Million on Claude AI in One Month," May 2026; Source: Business Today, "AI spending nightmare: Companies spend over $500 million in 30 days on Anthropic's Claude," May 2026. We covered that incident in detail here.

That last one is the finding underneath all of this: a production agent running exactly as designed, and a malicious one running exactly as an attacker designed, produce the same failure mode. The invoice can't tell them apart, and until this year, neither could most of the tooling watching it.

Three ways to inflate a bill without ever failing a task.

2026's security research has converged on one uncomfortable property shared across techniques: the task completes successfully, so the correctness checks that would normally catch a problem never fire.

Denial of wallet, in two charts: real incidents run from $46K/day (Sysdig LLMjacking on AWS Bedrock) to $500M/month (one enterprise, no usage caps), while attack-induced token amplification runs from a 1x ungoverned baseline to 658x on the "Beyond Max Tokens" MCP attack
Denial of wallet, in two charts: real incidents run from $46K/day (Sysdig LLMjacking on AWS Bedrock) to $500M/month (one enterprise, no usage caps), while attack-induced token amplification runs from a 1x ungoverned baseline to 658x on the "Beyond Max Tokens" MCP attack

OWASP gave this its own number: LLM10:2025.

Through the 2023 edition, OWASP's Top 10 for LLM Applications filed this under LLM04: Model Denial of Service, a narrow category about resource overload crashing a service. The 2025 revision retired that framing. LLM10:2025 is now Unbounded Consumption, covering the wider set of failures in which a model or the system around it generates more tokens, inference steps, or external calls than the application ever anticipated, with cost, not just availability, as the primary harm. Source: Aembit, "OWASP Top 10 for LLM Applications Explained". That's the field's own standards body admitting the risk stopped being a server falling over and became a bill that keeps climbing while every dashboard reads green.

Why the tools already watching your traffic don't see it.

A rate limiter counts requests. A WAF pattern-matches payloads. Neither has a concept of cost, and that gap is exactly what every technique above is built to walk through. "Standard WAFs and API gateways cannot distinguish between a $0.001 cached response and a $0.50 agentic workflow." Source: toxsec, "Model Denial of Service Turns Your Cloud Bill Into a Weapon". A request that triggers a 40-step tool-calling loop and a request that returns instantly from cache look identical to infrastructure that only counts requests per second. Stay under the rate limit, stay under the availability threshold, and the cost compounds somewhere nothing is watching it.

The one defense that holds up: a ceiling, not a filter.

A 2026 study mapping which defenses actually close which OWASP LLM Top 10 risks found something specific worth building around: token-budget controls that terminate a multi-step sequence eliminated Unbounded Consumption findings in its test harness entirely, and unlike refusal-phrase filters or keyword guards, that effectiveness held at full strength even when attackers paraphrased their prompts to evade detection. Source: "Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing," arXiv:2606.02822. A filter can be reworded around. A hard ceiling on tokens or dollars per task cannot.

The practical version of that same conclusion, echoed across incident writeups from this year, treats cost control and action control as one control surface, not two: cap retries at a small fixed number instead of letting them run unlimited, cap total tool calls per task, set a token or dollar ceiling per agent and per action, treat repeated identical tool calls as a loop signal, and when a threshold is crossed mid-task, escalate to a cheaper model or a human reviewer instead of hard-failing the request. Source: Sondera, "AI Agent Token Costs Are Now a Security Risk"

That last point is where most teams get it backwards. The instinct when a request crosses a cost threshold is to return a hard failure and kill it, which also kills the legitimate 38-step task that was one step from finishing. Downgrading the rest of the task to a cheaper model tier instead keeps the work alive and puts a ceiling on the blast radius at the same moment.

What a request counter sees versus what a cost-aware proxy sees.

Request-counting gateway or WAFCost-aware LLM proxy
Limits enforced byrequests per second or minutetokens and dollars per request, session, and agent
A 40-step tool-calling loop that finishes correctlylogged as one successful requestflagged the moment cumulative cost crosses a ceiling
Repeated identical tool callsinvisible; not a payload signaturedetected as a loop pattern and capped
Response to a threshold breachnothing, or a blanket failure that kills the taskdowngrade to a cheaper model tier for the rest of the task
Stolen-key abuse (LLMjacking)invisible until the invoice arrivesanomalous spend per hour on a key trips an alert same-day
What you see on a dashboardrequests per seconddollars per route, per session, per API key

The five-line version.

The defense that survived paraphrasing in the arXiv study above isn't a smarter classifier. It's arithmetic run on every call: track cumulative cost against a ceiling, watch for the same tool call repeating, and downgrade instead of dying.

# Cost ceiling + loop detection: not a request counter
session_budget_usd = 2.00
recent_calls = collections.deque(maxlen=6)

def guard(session, tool_call, projected_cost_usd):
    recent_calls.append(hash(tool_call))
    if list(recent_calls).count(recent_calls[-1]) >= 4:
        raise LoopDetected(tool_call)          # same call, four times running

    if session.spent_usd + projected_cost_usd > session_budget_usd:
        return downgrade_model(session, tier="cheap")   # keep the task alive
    return session.model

This is the shape of what Nadir enforces on every call that routes through it: a per-request and per-session cost ceiling that isn't a request counter, and an automatic downgrade path instead of a hard failure when a task runs long or a tool-calling chain repeats. The dashboard tracks dollars per route, per session, and per API key, which is the unit denial-of-wallet actually moves in, not requests per second. Paired with tool-schema deduplication that shrinks the fixed cost every tool call carries, and verifier-gated routing that avoids compounding a wrong answer into a bigger, more expensive retry, the same layer that trims your everyday bill is the layer that caps how bad your worst day can get.

What to actually ship this week.

Conclusion.

The most useful thing about denial-of-wallet research so far is what it does to the mental model. This was never really a new attack category. It's the same failure agentic workloads already produce by accident, given a name and a motive. Uber burning its entire 2026 AI budget in four months and a "Beyond Max Tokens" MCP server inflating a single task 658x are the same shape of problem wearing different intent, one an adoption story and one an attack. Whatever you build to survive one survives the other, because the fix was never about detecting bad intent. It's about putting a ceiling on cost that doesn't move when the wording does.


Sources: [toxsec, "Model Denial of Service Turns Your Cloud Bill Into a Weapon"](https://www.toxsec.com/p/denial-of-wallet). [Sysdig, "LLMjacking: Stolen Cloud Credentials Used in New AI Attack"](https://www.sysdig.com/blog/llmjacking-stolen-cloud-credentials-used-in-new-ai-attack). [LayerX Security, "Denial of Wallet Attacks: Draining Resources via GenAI Abuse"](https://layerxsecurity.com/generative-ai/denial-of-wallet-attacks/). [Sondera, "AI Agent Token Costs Are Now a Security Risk"](https://blog.sondera.ai/p/ai-agent-token-costs-security-risk). [SecurityWeek, "The AI Token Costs That Can Break Cybersecurity"](https://www.securityweek.com/the-ai-token-costs-that-can-break-cybersecurity/). [Aembit, "OWASP Top 10 for LLM Applications Explained"](https://aembit.io/blog/owasp-top-10-llm-risks-explained/). ["ReasoningBomb," arXiv:2602.00154](https://arxiv.org/pdf/2602.00154). ["OverThink: Slowdown Attacks on Reasoning LLMs," arXiv:2502.02542](https://arxiv.org/pdf/2502.02542). ["Which Defense Closes Which Threat?," arXiv:2606.02822](https://arxiv.org/pdf/2606.02822). [Yahoo Finance, "Client Accidentally Burns $500 Million on Claude AI in One Month"](https://finance.yahoo.com/sectors/technology/articles/client-accidentally-burns-500-million-105400717.html). [Business Today, "AI spending nightmare"](https://www.businesstoday.in/technology/artificial-intelligence/story/ai-spending-nightmare-companies-spend-over-a-500-million-in-30-days-on-anthropics-claude-533824-2026-05-29).

More on finops & governance

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

For coding agents, Nadir connects to Codex, Claude Code, or Cursor and recommends a model tier for delegated work. The agent decides whether to hand the task off and uses its own configured models, so Nadir needs no proxy and no provider keys on that path.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.