Nobody noticed for three weeks.
On August 31, 2026, METR, the AI evaluation nonprofit, published a security update describing two incidents. The first one is the one every team running agents should read. In March 2026, an attacker took a model provider API key from a researcher's agent dashboard and used it for three weeks. The usage came to about $600,000 in model credits. Source: METR, "Update on Security at METR"
The details are ordinary, which is the point:
- The entry point was a side project. A researcher ran an agent orchestration app on a personal EC2 instance, internet-facing on purpose and meant to sit behind Google authentication. The app, which METR describes as vibe-coded, had a fail-open bug that silently turned that authentication off.
- The attacker found it through certificate transparency. METR suspects the attacker scanned recently registered sites for keywords about LLMs and agents. A new TLS certificate is a public announcement that a new service exists.
- The agent handed over the key. The attacker didn't need an exploit. They prompted the agent to reveal its model provider API key, then added an SSH key to keep access.
- The key had no spend limit. In METR's words, "as of the incident there was no way to put a spending limit on keys like this one."
- The spend didn't look strange. METR says it didn't notice because "we are accustomed to running evaluations and experiments that use large volumes of tokens."
The credits had been granted to METR free by the model developer, so METR didn't pay the bill. Your company would. $600,000 over 21 days is about $28,600 a day, or roughly $1,190 an hour, around the clock.
METR isn't alone. On August 2, 2026, Okta's threat intelligence team found a 7 GB infostealer dump on Telegram with 44,791 JSON Web Tokens from 5,871 infected machines. 555 looked like AI service sessions, across OpenAI, Anthropic, Google, Cursor, OpenRouter and others, and the dump included 24 valid API keys for four AI services. Source: The Hacker News We covered the wider pattern, from LLMjacking on stolen cloud credentials to agents billed into a loop, in Denial of Wallet Is the New DDoS. This post is about the part METR's write-up makes concrete: why a big leak can stay invisible, and the cheap per-key checks that would have caught it on day one.
Spend figures in the charts below are illustrative, from a simulation we built for this post. They are not METR's actual spend, which METR didn't publish, and not from customer data.
Why $28,600 a day hid.
Most teams watch AI spend the way they watch a cloud bill: one line, org-wide, with an alert when today looks much bigger than usual. That works when usage is steady. It fails for exactly the teams with the most to lose: the ones running evals, batch jobs and agent fleets, whose normal spend already swings by several times from day to day.
We simulated 84 days of spend for a bursty, eval-heavy workload with a median day of about $37,600 and occasional eval sweeps up to 5.3x that. Then we added a stolen key spending $28,600 a day for 21 days, and applied a common alert rule: fire when a day's spend is more than twice the 28-day median.
The alert did fire during the attack. It had also fired five times in the four quiet weeks before it. Across 500 simulated runs, it went off on 16% of ordinary days and 22% of attack days. That's not a detector. It's an alert that fires about once a week, which someone learns to mute by week three. A stricter rule, spend above the 28-day maximum, fired on about one day in 21 during the attack and about once in the 28 quiet days before it.
The attacker didn't need to be subtle. The workload made the leak look normal.
Same money, a different denominator.
Now look at the same $28,600 from the point of view of the key that leaked. In our simulation, the dashboard key normally spends about $30 on a working day and close to nothing on weekends.
Against the org, the attack is a 1.76x day, well inside the normal range. Against the key, it's 947x its own median on the first day. A per-key rule flagged it on day one in every one of the 500 simulated runs, with zero false alerts in the quiet weeks.
That comparison only exists if the key has one job. When the dashboard, the eval harness and three internal tools share one provider key, the key's baseline is the org's baseline, and the leak is noise again. One key per workload is the step that makes any per-key signal possible.
The agent told them the key.
The other lesson from METR is about what an agent can reveal. An agent with shell access, an environment, or a tool that reads files can print anything it can reach. Prompt injection defenses reduce that risk. They don't remove it. So plan for the key an agent holds to leak eventually, and choose what that key can do.
A provider master key in an agent's environment is the worst case: every model on the account, no ceiling, and a revocation that breaks every other service using it. A scoped key issued by a gateway changes all three. The provider key never reaches the agent box. The leaked key hits a daily cap in hours, not an invoice weeks later. And it can be deleted alone.
Tutorial: a per-key spend watcher in 90 lines.
This detector reads your usage log, one row per LLM call, and flags any key-hour that looks wrong. It needs five fields most gateways and provider usage exports already have: key id, timestamp, model, token counts and cost. It runs as a cron job every hour and needs no ML.
It checks four signals, all against the key's own history:
- Spend above the key's busiest hour. Not its average: eval runners and batch jobs sit idle most hours, so a median of zero says nothing.
- 24-hour spend above the key's busiest day. This catches a steady, around-the-clock leak that never spikes a single hour.
- A model the key has never used. Stolen keys go to the most expensive model on the account.
- A change in shape. A dashboard that sends long prompts and gets short answers suddenly generating long outputs is a different workload, whatever the total.
New keys get a flat hourly cap for their first week, because they have no history to compare against.
Step 1: the detector.
from collections import defaultdict
from dataclasses import dataclass
from datetime import datetime, timedelta
@dataclass
class Row:
key_id: str
ts: datetime
model: str
input_tokens: int
output_tokens: int
cost_usd: float
NEW_KEY_DAYS = 7 # younger than this: judged by a flat cap, not a baseline
NEW_KEY_HOURLY_CAP = 25.0 # USD per hour for a key with no history
PEAK_MULT = 1.5 # this hour vs the key's busiest hour in 28 days
DAY_MULT = 1.5 # last 24 hours vs the key's busiest day in 28 days
FLOOR_USD = 5.0 # ignore anything under $5 an hour
SHAPE_SHIFT = 4.0 # output/input ratio moved this many times
def hourly(rows):
"""{key: {hour: [rows]}}"""
out = defaultdict(lambda: defaultdict(list))
for r in rows:
out[r.key_id][r.ts.replace(minute=0, second=0, microsecond=0)].append(r)
return out
def out_in_ratio(rows):
i = sum(r.input_tokens for r in rows) or 1
return sum(r.output_tokens for r in rows) / i
def check_hour(key, hour, by_hour, first_seen):
"""Return the reasons this key's hour looks wrong, empty if it looks normal."""
now = by_hour[key][hour]
spend = sum(r.cost_usd for r in now)
reasons = []
past = [h for h in by_hour[key] if hour - timedelta(days=28) <= h < hour]
if hour - first_seen[key] < timedelta(days=NEW_KEY_DAYS):
if spend > NEW_KEY_HOURLY_CAP:
reasons.append(f"new key spent ${spend:,.0f} in an hour")
return reasons
# Compare against the key's own peaks, not its average: eval runners and
# batch jobs are idle most hours, so a mean or median says nothing.
cost = lambda h: sum(r.cost_usd for r in by_hour[key].get(h, []))
peak_hour = max((cost(h) for h in past), default=0.0)
if spend > max(peak_hour * PEAK_MULT, FLOOR_USD):
reasons.append(f"${spend:,.0f} this hour, busiest hour in 28 days was ${peak_hour:,.0f}")
last_24h = sum(cost(hour - timedelta(hours=n)) for n in range(24))
days = defaultdict(float)
for h in past:
if h < hour - timedelta(hours=23):
days[h.date()] += cost(h)
peak_day = max(days.values(), default=0.0)
if last_24h > max(peak_day * DAY_MULT, FLOOR_USD * 24):
reasons.append(f"${last_24h:,.0f} in 24 hours, busiest day in 28 days was ${peak_day:,.0f}")
seen_models = {r.model for h in past for r in by_hour[key][h]}
new_models = {r.model for r in now} - seen_models
if new_models:
reasons.append(f"first use of {', '.join(sorted(new_models))}")
past_rows = [r for h in past for r in by_hour[key][h]]
if past_rows:
before, after = out_in_ratio(past_rows), out_in_ratio(now)
if after > before * SHAPE_SHIFT or after < before / SHAPE_SHIFT:
reasons.append(f"output/input ratio {before:.2f} -> {after:.2f}")
return reasons
def scan(rows):
by_hour = hourly(rows)
first_seen = {}
for r in sorted(rows, key=lambda r: r.ts):
first_seen.setdefault(r.key_id, r.ts)
alerts = []
for key, hours in by_hour.items():
for hour in sorted(hours):
reasons = check_hour(key, hour, by_hour, first_seen)
if reasons:
alerts.append((hour, key, reasons))
return sorted(alerts)
This version rescans every hour for clarity. In production, run check_hour only on the hour that just closed, and keep the per-key daily and hourly peaks in a small table instead of recomputing them.
Step 2: test it on a leak you planted.
Don't trust a detector you haven't seen fire. We generated 40 days of synthetic traffic for two keys: an eval runner that spends $500 to $3,000 an hour in bursts on weekdays and is idle otherwise, and a dashboard that spends $1 to $3 an hour during working hours on Claude Sonnet 5. On day 35 we planted the METR pattern on the dashboard key: $1,190 an hour, around the clock, on Claude Opus 5, with long outputs.
hour, key, reasons = next(a for a in scan(rows) if a[1] == "dashboard")
print(hour, key)
for reason in reasons:
print(" ", reason)
2026-04-05 02:00:00 dashboard
$1,190 this hour, busiest hour in 28 days was $3
$1,190 in 24 hours, busiest day in 28 days was $20
first use of claude-opus-5
output/input ratio 0.03 -> 0.60
All four signals fired in the first hour. The eval runner, with hourly spend that swings from $0 to $3,000, raised zero false alerts once its first week was over. In that first week it tripped the new-key cap, as intended: a brand-new key spending $2,000 an hour is worth a look.
Then the harder case: the same attack on the eval runner's key, same model, same shape as its normal traffic. The hourly rule stayed quiet, because $1,190 is below the key's normal peak. The 24-hour rule fired 17 hours in: $21,420 in a day against a busiest day of $13,874. That's slower, but it's 17 hours rather than three weeks. It's also the argument for giving the noisiest workload its own key, so no one else's traffic can hide in its bursts.
Step 3: decide what an alert does.
An alert nobody acts on at 3 a.m. is a log line. Pick in advance:
- New key over its cap, or first use of a flagship model: block the key automatically. A false positive costs one engineer a new key in the morning.
- 24-hour rule on an established key: page someone. That's where a big legitimate eval run and a leak look the most alike.
- Shape change alone: a ticket. It often means someone shipped a new prompt.
Checklist: before your agents hold a key.
- Keep provider master keys out of agent environments. Give agents a scoped key from a gateway or a provider project key, never the account key.
- One key per workload. Per service, per agent, per environment. It's what gives each key a baseline, and it lets you revoke one without an outage.
- Allowlist models per key. A dashboard that uses Sonnet has no reason to reach the flagship model. Stolen keys go straight to the most expensive model on the account.
- Put a daily dollar cap on everything that can hold one. METR's key type couldn't take a spend limit. A gateway in front of it can.
- Watch certificate transparency for your own domains. Attackers read new-certificate logs for words like "agent" and "llm." You can read the same logs to find the side projects your team forgot.
- Treat side projects as production when they hold production keys. METR's entry point was a personal instance. The key it held wasn't personal.
Where Nadir fits.
Nadir is an OpenAI compatible gateway that routes each prompt to the cheapest model that can handle it. The same position in the request path is what makes the checklist above cheap to enforce.
With BYOK (bring your own keys), your provider keys are stored encrypted on the Nadir side, and agents hold ndr_ keys instead. Issue one per workload from the dashboard, and delete one without touching the provider key or any other service. Each key can carry a model allowlist, so a leaked dashboard key can't reach the flagship model at all. Governance policies set a daily dollar budget per org, team or user, and once it's spent, requests are refused with a 402 until 00:00 UTC. That applies to BYOK accounts too. The budget attaches to the person or team, not the key, so a developer who can delete their own key can't delete their budget with it. Every request logs its model and cost against the key that made it, which is the usage log the detector above reads.
Routing is still where most of the savings come from. Start with a free key, move one agent onto its own ndr_ key with an allowlist and a daily budget, and point the 90-line watcher at its usage.
Conclusion.
METR lost about $600,000 in credits in three weeks to an attacker who found an agent dashboard through certificate logs and asked the agent for its key. The spend went unnoticed because it was smaller than the normal swings of a large eval workload. Org-wide alerts can't see that kind of leak. Per-key baselines see it in the first hour, but only if each key has one job. Give agents scoped keys, one per workload, with a model allowlist and a daily cap, and compare each key only to itself. The leak you can't prevent then costs a day's budget, not a quarter's.
Spend figures in the charts are illustrative, from a simulation built for this post: 84 days of lognormal daily spend with a weekday median of about $37,600 and eval sweeps on about 12% of days, plus a stolen key adding $600,000 over 21 days, repeated 500 times. The tutorial results come from a 40-day synthetic two-key log. Neither is METR's data, which wasn't published, or customer data. Sources: [METR, "Update on Security at METR," August 31, 2026](https://metr.org/blog/2026-08-31-security-update/). [The Hacker News, "Attackers Steal METR API Key and Consume AI Credits Worth About $600,000"](https://thehackernews.com/2026/09/attackers-steal-metr-api-key-and.html). [Dark Reading, "AI Model Evaluator METR Hit by Credential Theft, Probing"](https://www.darkreading.com/identity-access-management-security/ai-model-evaluator-metr-credential-theft-probing). [The Hacker News, "Infostealer Logs Expose Replayable AI Tokens That Can Bypass MFA"](https://thehackernews.com/2026/09/infostealer-logs-expose-replayable-ai.html). [OWASP Top 10 for LLM Applications 2025, LLM10: Unbounded Consumption](https://genai.owasp.org/llmrisk/llm102025-unbounded-consumption/).