$600K in the Noise

On August 31, 2026, METR disclosed that an attacker used a stolen model provider API key for three weeks and consumed about $600,000 in credits. The key came from a vibe-coded agent dashboard found through certificate transparency logs, and the agent revealed it when asked. METR didn't notice because it already runs large token volumes. We simulated why: on a bursty eval workload, a common "2x the 28-day median" alert fired on 16% of ordinary days and 22% of attack days, so the leak looked like a busy week. Measured against the leaked key's own history, the same spend was 947x normal on day one. Includes a tested 90-line Python per-key spend watcher that flagged the planted leak in its first hour with zero false alerts on a noisy eval key, and a checklist for giving agents scoped keys instead of provider master keys.

Published 2026-09-30 by Dor Amir on the Nadir blog.

Filed under FinOps & Governance.

Nobody noticed for three weeks.

On August 31, 2026, METR, the AI evaluation nonprofit, published a security update describing two incidents. The first one is the one every team running agents should read. In March 2026, an attacker took a model provider API key from a researcher's agent dashboard and used it for three weeks. The usage came to about $600,000 in model credits. Source: METR, "Update on Security at METR"

The details are ordinary, which is the point:

The credits had been granted to METR free by the model developer, so METR didn't pay the bill. Your company would. $600,000 over 21 days is about $28,600 a day, or roughly $1,190 an hour, around the clock.

METR isn't alone. On August 2, 2026, Okta's threat intelligence team found a 7 GB infostealer dump on Telegram with 44,791 JSON Web Tokens from 5,871 infected machines. 555 looked like AI service sessions, across OpenAI, Anthropic, Google, Cursor, OpenRouter and others, and the dump included 24 valid API keys for four AI services. Source: The Hacker News We covered the wider pattern, from LLMjacking on stolen cloud credentials to agents billed into a loop, in Denial of Wallet Is the New DDoS. This post is about the part METR's write-up makes concrete: why a big leak can stay invisible, and the cheap per-key checks that would have caught it on day one.

Spend figures in the charts below are illustrative, from a simulation we built for this post. They are not METR's actual spend, which METR didn't publish, and not from customer data.

Why $28,600 a day hid.

Most teams watch AI spend the way they watch a cloud bill: one line, org-wide, with an alert when today looks much bigger than usual. That works when usage is steady. It fails for exactly the teams with the most to lose: the ones running evals, batch jobs and agent fleets, whose normal spend already swings by several times from day to day.

We simulated 84 days of spend for a bursty, eval-heavy workload with a median day of about $37,600 and occasional eval sweeps up to 5.3x that. Then we added a stolen key spending $28,600 a day for 21 days, and applied a common alert rule: fire when a day's spend is more than twice the 28-day median.

Line chart of simulated org-wide daily AI spend over 84 days. Spend swings between about $10,000 and $200,000 a day. A stolen key adds about $28,600 a day from day 57 to day 77, shown as an orange band under the line. Open circles mark days when a "spend over 2x the 28-day median" alert fired: 5 times in the 28 days before the attack and 7 times during it.
Line chart of simulated org-wide daily AI spend over 84 days. Spend swings between about $10,000 and $200,000 a day. A stolen key adds about $28,600 a day from day 57 to day 77, shown as an orange band under the line. Open circles mark days when a "spend over 2x the 28-day median" alert fired: 5 times in the 28 days before the attack and 7 times during it.

The alert did fire during the attack. It had also fired five times in the four quiet weeks before it. Across 500 simulated runs, it went off on 16% of ordinary days and 22% of attack days. That's not a detector. It's an alert that fires about once a week, which someone learns to mute by week three. A stricter rule, spend above the 28-day maximum, fired on about one day in 21 during the attack and about once in the 28 quiet days before it.

The attacker didn't need to be subtle. The workload made the leak look normal.

Same money, a different denominator.

Now look at the same $28,600 from the point of view of the key that leaked. In our simulation, the dashboard key normally spends about $30 on a working day and close to nothing on weekends.

Log-scale bar chart titled "Same $28,600 a day, two denominators." Org-wide, a median day plus the attacker's spend is 1.76x a normal day, inside the org's normal range of up to 5.3x. On the leaked key, the same spend is 947x the key's own median day, from the first day.
Log-scale bar chart titled "Same $28,600 a day, two denominators." Org-wide, a median day plus the attacker's spend is 1.76x a normal day, inside the org's normal range of up to 5.3x. On the leaked key, the same spend is 947x the key's own median day, from the first day.

Against the org, the attack is a 1.76x day, well inside the normal range. Against the key, it's 947x its own median on the first day. A per-key rule flagged it on day one in every one of the 500 simulated runs, with zero false alerts in the quiet weeks.

That comparison only exists if the key has one job. When the dashboard, the eval harness and three internal tools share one provider key, the key's baseline is the org's baseline, and the leak is noise again. One key per workload is the step that makes any per-key signal possible.

The agent told them the key.

The other lesson from METR is about what an agent can reveal. An agent with shell access, an environment, or a tool that reads files can print anything it can reach. Prompt injection defenses reduce that risk. They don't remove it. So plan for the key an agent holds to leak eventually, and choose what that key can do.

Diagram comparing two setups. Top: an agent dashboard holds a provider API key directly. If it leaks, it spends until someone reads the invoice, on any model at any volume, and revoking it breaks every service that shares it. Bottom: the agent holds a scoped gateway key. The gateway holds the provider key, enforces a model allowlist and a daily dollar budget, and logs cost per key. If that key leaks, it stops at the day's cap, reaches only the allowed models, and can be deleted without rotating anything else.
Diagram comparing two setups. Top: an agent dashboard holds a provider API key directly. If it leaks, it spends until someone reads the invoice, on any model at any volume, and revoking it breaks every service that shares it. Bottom: the agent holds a scoped gateway key. The gateway holds the provider key, enforces a model allowlist and a daily dollar budget, and logs cost per key. If that key leaks, it stops at the day's cap, reaches only the allowed models, and can be deleted without rotating anything else.

A provider master key in an agent's environment is the worst case: every model on the account, no ceiling, and a revocation that breaks every other service using it. A scoped key issued by a gateway changes all three. The provider key never reaches the agent box. The leaked key hits a daily cap in hours, not an invoice weeks later. And it can be deleted alone.

Tutorial: a per-key spend watcher in 90 lines.

This detector reads your usage log, one row per LLM call, and flags any key-hour that looks wrong. It needs five fields most gateways and provider usage exports already have: key id, timestamp, model, token counts and cost. It runs as a cron job every hour and needs no ML.

It checks four signals, all against the key's own history:

  1. Spend above the key's busiest hour. Not its average: eval runners and batch jobs sit idle most hours, so a median of zero says nothing.
  2. 24-hour spend above the key's busiest day. This catches a steady, around-the-clock leak that never spikes a single hour.
  3. A model the key has never used. Stolen keys go to the most expensive model on the account.
  4. A change in shape. A dashboard that sends long prompts and gets short answers suddenly generating long outputs is a different workload, whatever the total.

New keys get a flat hourly cap for their first week, because they have no history to compare against.

Step 1: the detector.

from collections import defaultdict
from dataclasses import dataclass
from datetime import datetime, timedelta

@dataclass
class Row:
    key_id: str
    ts: datetime
    model: str
    input_tokens: int
    output_tokens: int
    cost_usd: float

NEW_KEY_DAYS = 7          # younger than this: judged by a flat cap, not a baseline
NEW_KEY_HOURLY_CAP = 25.0 # USD per hour for a key with no history
PEAK_MULT = 1.5           # this hour vs the key's busiest hour in 28 days
DAY_MULT = 1.5            # last 24 hours vs the key's busiest day in 28 days
FLOOR_USD = 5.0           # ignore anything under $5 an hour
SHAPE_SHIFT = 4.0         # output/input ratio moved this many times

def hourly(rows):
    """{key: {hour: [rows]}}"""
    out = defaultdict(lambda: defaultdict(list))
    for r in rows:
        out[r.key_id][r.ts.replace(minute=0, second=0, microsecond=0)].append(r)
    return out

def out_in_ratio(rows):
    i = sum(r.input_tokens for r in rows) or 1
    return sum(r.output_tokens for r in rows) / i

def check_hour(key, hour, by_hour, first_seen):
    """Return the reasons this key's hour looks wrong, empty if it looks normal."""
    now = by_hour[key][hour]
    spend = sum(r.cost_usd for r in now)
    reasons = []

    past = [h for h in by_hour[key] if hour - timedelta(days=28) <= h < hour]
    if hour - first_seen[key] < timedelta(days=NEW_KEY_DAYS):
        if spend > NEW_KEY_HOURLY_CAP:
            reasons.append(f"new key spent ${spend:,.0f} in an hour")
        return reasons

    # Compare against the key's own peaks, not its average: eval runners and
    # batch jobs are idle most hours, so a mean or median says nothing.
    cost = lambda h: sum(r.cost_usd for r in by_hour[key].get(h, []))
    peak_hour = max((cost(h) for h in past), default=0.0)
    if spend > max(peak_hour * PEAK_MULT, FLOOR_USD):
        reasons.append(f"${spend:,.0f} this hour, busiest hour in 28 days was ${peak_hour:,.0f}")

    last_24h = sum(cost(hour - timedelta(hours=n)) for n in range(24))
    days = defaultdict(float)
    for h in past:
        if h < hour - timedelta(hours=23):
            days[h.date()] += cost(h)
    peak_day = max(days.values(), default=0.0)
    if last_24h > max(peak_day * DAY_MULT, FLOOR_USD * 24):
        reasons.append(f"${last_24h:,.0f} in 24 hours, busiest day in 28 days was ${peak_day:,.0f}")

    seen_models = {r.model for h in past for r in by_hour[key][h]}
    new_models = {r.model for r in now} - seen_models
    if new_models:
        reasons.append(f"first use of {', '.join(sorted(new_models))}")

    past_rows = [r for h in past for r in by_hour[key][h]]
    if past_rows:
        before, after = out_in_ratio(past_rows), out_in_ratio(now)
        if after > before * SHAPE_SHIFT or after < before / SHAPE_SHIFT:
            reasons.append(f"output/input ratio {before:.2f} -> {after:.2f}")
    return reasons

def scan(rows):
    by_hour = hourly(rows)
    first_seen = {}
    for r in sorted(rows, key=lambda r: r.ts):
        first_seen.setdefault(r.key_id, r.ts)
    alerts = []
    for key, hours in by_hour.items():
        for hour in sorted(hours):
            reasons = check_hour(key, hour, by_hour, first_seen)
            if reasons:
                alerts.append((hour, key, reasons))
    return sorted(alerts)

This version rescans every hour for clarity. In production, run check_hour only on the hour that just closed, and keep the per-key daily and hourly peaks in a small table instead of recomputing them.

Step 2: test it on a leak you planted.

Don't trust a detector you haven't seen fire. We generated 40 days of synthetic traffic for two keys: an eval runner that spends $500 to $3,000 an hour in bursts on weekdays and is idle otherwise, and a dashboard that spends $1 to $3 an hour during working hours on Claude Sonnet 5. On day 35 we planted the METR pattern on the dashboard key: $1,190 an hour, around the clock, on Claude Opus 5, with long outputs.

hour, key, reasons = next(a for a in scan(rows) if a[1] == "dashboard")
print(hour, key)
for reason in reasons:
    print("  ", reason)
2026-04-05 02:00:00 dashboard
   $1,190 this hour, busiest hour in 28 days was $3
   $1,190 in 24 hours, busiest day in 28 days was $20
   first use of claude-opus-5
   output/input ratio 0.03 -> 0.60

All four signals fired in the first hour. The eval runner, with hourly spend that swings from $0 to $3,000, raised zero false alerts once its first week was over. In that first week it tripped the new-key cap, as intended: a brand-new key spending $2,000 an hour is worth a look.

Then the harder case: the same attack on the eval runner's key, same model, same shape as its normal traffic. The hourly rule stayed quiet, because $1,190 is below the key's normal peak. The 24-hour rule fired 17 hours in: $21,420 in a day against a busiest day of $13,874. That's slower, but it's 17 hours rather than three weeks. It's also the argument for giving the noisiest workload its own key, so no one else's traffic can hide in its bursts.

Step 3: decide what an alert does.

An alert nobody acts on at 3 a.m. is a log line. Pick in advance:

Checklist: before your agents hold a key.

Where Nadir fits.

Nadir is an OpenAI compatible gateway that routes each prompt to the cheapest model that can handle it. The same position in the request path is what makes the checklist above cheap to enforce.

With BYOK (bring your own keys), your provider keys are stored encrypted on the Nadir side, and agents hold ndr_ keys instead. Issue one per workload from the dashboard, and delete one without touching the provider key or any other service. Each key can carry a model allowlist, so a leaked dashboard key can't reach the flagship model at all. Governance policies set a daily dollar budget per org, team or user, and once it's spent, requests are refused with a 402 until 00:00 UTC. That applies to BYOK accounts too. The budget attaches to the person or team, not the key, so a developer who can delete their own key can't delete their budget with it. Every request logs its model and cost against the key that made it, which is the usage log the detector above reads.

Routing is still where most of the savings come from. Start with a free key, move one agent onto its own ndr_ key with an allowlist and a daily budget, and point the 90-line watcher at its usage.

Conclusion.

METR lost about $600,000 in credits in three weeks to an attacker who found an agent dashboard through certificate logs and asked the agent for its key. The spend went unnoticed because it was smaller than the normal swings of a large eval workload. Org-wide alerts can't see that kind of leak. Per-key baselines see it in the first hour, but only if each key has one job. Give agents scoped keys, one per workload, with a model allowlist and a daily cap, and compare each key only to itself. The leak you can't prevent then costs a day's budget, not a quarter's.


Spend figures in the charts are illustrative, from a simulation built for this post: 84 days of lognormal daily spend with a weekday median of about $37,600 and eval sweeps on about 12% of days, plus a stolen key adding $600,000 over 21 days, repeated 500 times. The tutorial results come from a 40-day synthetic two-key log. Neither is METR's data, which wasn't published, or customer data. Sources: [METR, "Update on Security at METR," August 31, 2026](https://metr.org/blog/2026-08-31-security-update/). [The Hacker News, "Attackers Steal METR API Key and Consume AI Credits Worth About $600,000"](https://thehackernews.com/2026/09/attackers-steal-metr-api-key-and.html). [Dark Reading, "AI Model Evaluator METR Hit by Credential Theft, Probing"](https://www.darkreading.com/identity-access-management-security/ai-model-evaluator-metr-credential-theft-probing). [The Hacker News, "Infostealer Logs Expose Replayable AI Tokens That Can Bypass MFA"](https://thehackernews.com/2026/09/infostealer-logs-expose-replayable-ai.html). [OWASP Top 10 for LLM Applications 2025, LLM10: Unbounded Consumption](https://genai.owasp.org/llmrisk/llm102025-unbounded-consumption/).

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.