Your fallback has a price tag.
On the morning of September 3, 2026, three of the largest AI providers went down within the same few hours. Claude had a partial outage across Claude.ai, Claude Code, Cowork and the API that lasted 3 hours and 6 minutes, resolved at 16:16 UTC. ChatGPT and Codex were unavailable for some users for about 34 minutes because of what OpenAI called "a routing error." Grok went down after an outage at xAI's Memphis compute center. Cloudflare, which all three use, said it was "operating normally," and AWS, Google Cloud and Azure showed nothing. Source: The Register
Most of the write-ups since have asked the reliability question: did your app stay up? That question has a known answer. Put a second provider behind the first one and fail over. This post asks the question that comes after it: what did the failover cost? In our model it ranged from 1.3x to 5.9x the normal bill for the same three hours of work, depending on one line of config, which most teams wrote once and never priced.
Costs in this post are illustrative, modeled from public API list prices and an assumed agent workload, not measured production traces. Not derived from proprietary customer data. Sources cited throughout.
Three ways an outage shows up on the bill.
1. The fallback model has a different price. "Fall back to the best model on the other provider" is the most common chain we see, and it's usually a tier jump. Claude Sonnet 5 costs $2 / $10 per million input / output tokens. GPT-6 Astra costs $10 / $50. The failover works, and every request in the window costs five times more.
2. The prompt cache doesn't come with you. Agents re-send a large prefix on every call and rely on cache reads to make that affordable. A cache lives with one provider. Fail over, and every active session pays full input price once on the new provider. Fail back, and it pays again, because a 3-hour outage outlasts every cache TTL on the provider you left. On Anthropic, that rebuild is a cache write, billed at 1.25x the base input price. We covered the mechanics in The Cache-Write Tax and the switch penalty: an outage forces both sides of that switch on every session at once.
3. Naive retries cost time before they cost money. Most 5xx and overload errors aren't billed. But a client that retries the dead provider three times with a 30-second timeout spends 90 seconds per request before reaching the fallback. For an agent that makes 40 calls per task, that's the difference between a slow task and one that times out. Retry storms are a well-documented way to turn a partial outage into a full one.
The math for one outage.
Take a mid-size agent product doing 1 request per second, around the clock. Each request carries a 60,000-token prefix with 90% cache hits and produces 1,200 output tokens including reasoning. Primary model: Claude Sonnet 5. Average session length: 20 requests.
| Model | Input | Cache read | Cache write | Output | Warm request | Cold request |
|---|---|---|---|---|---|---|
| Claude Sonnet 5 | $2.00 | $0.20 | $2.50 | $10.00 | $0.0348 | $0.162 |
| GPT-6 Sol | $2.00 | $0.20 | n/a | $10.00 | $0.0348 | $0.132 |
| GPT-6 Astra | $10.00 | $1.00 | n/a | $50.00 | $0.174 | $0.660 |
Price per million tokens. OpenAI caches automatically with no write premium; a cold OpenAI request pays the base input rate.
A 3-hour-6-minute outage is 11,160 requests across 558 sessions. With 558 cold starts on the way over and 558 cache rewrites on the way back:
- Price-matched failover (Sol): $514. 1.3x normal. The whole premium is cache: about $54 of cold starts going over and $71 of cache writes coming back.
- "Best available" failover (Astra): $2,284. 5.9x normal, for the same work. Nobody chose to pay that; someone wrote
fallback: "gpt-6-astra"in March because it was the safe choice. - No fallback: $0. And 11,160 failed requests.
$1,900 for one morning isn't a crisis. But September 3 wasn't the only incident this year, and outages cluster. A fallback chain is a pricing decision you make once and pay for at the worst possible time, with nobody watching the bill.
Four rules for a fallback chain that doesn't surprise finance.
Match price, not prestige. The fallback for a mid-tier model is a mid-tier model on another provider. Sonnet 5 and GPT-6 Sol have identical list prices, which makes them natural pairs. Keep the flagship as a fallback for requests that were already going to a flagship.
Match per request, not per app. A single app-wide fallback treats a one-line classification and a 200-file refactor the same way. If traffic is already split by difficulty, the fallback map is tier to tier: simple to simple, complex to complex.
Fail fast with a circuit breaker. After a handful of consecutive failures, stop calling the primary. Send one probe after a cool-down, and switch back only when the probe succeeds.
Cap the multiplier. Decide in advance how much more you're willing to pay during an incident, and enforce it in code. Past the cap, drop to a cheaper tier or queue non-interactive work instead of paying 5x for a nightly batch job.
Tutorial: a cost-aware failover chain in 80 lines.
This is a provider-agnostic sketch in Python. It keeps a circuit breaker per provider, maps each tier to a price-matched fallback, refuses fallbacks above a cost multiplier, and logs what the incident is costing while it happens. Swap call_provider for your SDK calls.
Step 1: prices and the fallback map.
import time, logging
from dataclasses import dataclass, field
log = logging.getLogger("failover")
# USD per 1M tokens: input, cache read, output
PRICES = {
"anthropic/claude-haiku-4-5": (1.00, 0.10, 5.00),
"anthropic/claude-sonnet-5": (2.00, 0.20, 10.00),
"anthropic/claude-opus-5": (5.00, 0.50, 25.00),
"openai/gpt-6-luna": (0.10, 0.01, 0.50),
"openai/gpt-6-sol": (2.00, 0.20, 10.00),
"openai/gpt-6-astra": (10.00, 1.00, 50.00),
}
# Tier-matched chains: primary first, then price-matched fallbacks.
CHAINS = {
"simple": ["anthropic/claude-haiku-4-5", "openai/gpt-6-luna"],
"medium": ["anthropic/claude-sonnet-5", "openai/gpt-6-sol"],
"complex": ["anthropic/claude-opus-5", "openai/gpt-6-astra", "openai/gpt-6-sol"],
}
MAX_MULTIPLIER = 1.5 # never fail over to more than 1.5x the primary's price per request
def est_cost(model, inp, cached, out):
pi, pc, po = PRICES[model]
return ((inp - cached) * pi + cached * pc + out * po) / 1e6
Step 2: a circuit breaker per provider.
@dataclass
class Breaker:
threshold: int = 5 # consecutive failures before opening
cooldown: float = 60.0 # seconds before one probe is allowed
failures: int = 0
opened_at: float | None = None
probing: bool = False
def allow(self) -> bool:
if self.opened_at is None:
return True
if not self.probing and time.monotonic() - self.opened_at >= self.cooldown:
self.probing = True # half-open: let exactly one request through
return True
return False
def success(self):
if self.opened_at is not None:
log.warning("breaker closed; expect cold caches on the primary")
self.failures, self.opened_at, self.probing = 0, None, False
def failure(self):
self.failures += 1
if self.probing or self.failures >= self.threshold:
self.opened_at, self.probing = time.monotonic(), False
breakers: dict[str, Breaker] = {}
def provider(model: str) -> str:
return model.split("/", 1)[0]
The breaker is per provider, not per model. On September 3 the Claude outage hit Sonnet 5 "and other models," so failing over from Sonnet to Opus on the same provider would have bought nothing.
Step 3: the call path.
class ProviderError(Exception): ... # raise from call_provider on 429, 5xx, 529
class AllProvidersDown(Exception): ...
def complete(tier: str, messages, est_in: int, est_cached: int, est_out: int):
chain = CHAINS[tier]
primary_cost = est_cost(chain[0], est_in, est_cached, est_out)
for i, model in enumerate(chain):
b = breakers.setdefault(provider(model), Breaker())
if not b.allow():
continue
# Cap on like-for-like price, so a cold cache alone never blocks a fallback.
ratio = est_cost(model, est_in, est_cached, est_out) / primary_cost
if i > 0 and ratio > MAX_MULTIPLIER:
log.info("skip %s: %.1fx primary price", model, ratio)
continue
# What this request will actually cost: a fallback starts cold.
cost = est_cost(model, est_in, est_cached if i == 0 else 0, est_out)
try:
resp = call_provider(model, messages, timeout=30)
except (TimeoutError, ConnectionError, ProviderError) as e:
b.failure()
log.warning("%s failed: %s", model, e)
continue
b.success()
if i > 0:
log.warning("served by fallback %s at ~$%.4f (primary ~$%.4f)",
model, cost, primary_cost)
return resp
raise AllProvidersDown(tier)
Two details carry most of the value. The cap compares like with like, list price against list price on the same request, so the unavoidable cold-cache premium never blocks a price-matched fallback. The logged cost is priced cold, because that's what the first request of each session will actually cost. And the check runs before the call, so the cap holds while nobody is looking. With MAX_MULTIPLIER = 1.5, a medium request fails over from Sonnet 5 to Sol at 1.0x, and a complex request skips Astra (2x Opus 5) and lands on Sol. That's the right call for a background job and the wrong one for your largest customer's hardest task. Set the cap per API key or per customer tier, not globally.
Step 4: run a drill.
Pick a quiet hour, open the breaker for your primary provider by hand, and let real traffic run on the fallback for 30 minutes. Then look at three numbers: cost per request against the primary, error rate, and the share of structured outputs that still parse. The last one catches a problem the bill doesn't: a fallback model that formats tool calls differently. We covered that failure in Right Answer, Wrong Shape. A drill costs 30 minutes of fallback pricing. Finding out on the next outage costs the whole outage.
What to check in your own config this week.
- Find every hardcoded fallback.
grep -r fallbackacross services, gateway configs and SDK wrappers. Write down the primary and fallback list price next to each one. - Flag any fallback above 2x its primary. Those are the lines that turned September 3 into a 5.9x morning.
- Check that fallbacks cross providers. A same-provider fallback protects you from a model deprecation, not from an outage.
- Check the retry policy before the fallback. Three retries at a 30-second timeout is 90 seconds of dead time per request. Fewer retries, a shorter timeout, and a breaker get you to the fallback in seconds.
- Put the fallback model in your logs. If your cost dashboard can't show "requests served by fallback" as a line, the next incident's premium will show up a month later as an unexplained spike.
Where Nadir fits.
Nadir already makes the per-request decision the tier-matched rules depend on. Each prompt is classified as simple, medium or complex and routed to the cheapest model that clears your quality bar, so the fallback for a request is chosen at the same tier as the request itself, not from one app-wide list.
Underneath, a provider health monitor scores every provider on recent successes and failures, and a circuit breaker skips a provider after repeated 429s, 5xx errors or timeouts, then probes it before sending traffic back. You can pin your own chain per request with route="fallback" and fallback_models=[...], or save one per API key in the dashboard. Every response reports the model that actually answered and what it cost, so an incident shows up as a line on the dashboard the same day, not as a surprise on the invoice.
It's OpenAI compatible: change the base URL, set model="auto", keep your own provider keys. Start with a free key, set a cross-provider chain, and run the 30-minute drill above against it.
Conclusion.
September 3 was a reliability story, and the fix for reliability is well known: a second provider and a failover chain. The part that gets less attention is that the chain is also a price list, applied automatically during the hours nobody is watching. In our model, the same three hours cost $514 with a price-matched fallback and $2,284 with a "best available" one. Match fallbacks by price and by request, fail fast with a breaker, cap the multiplier in code, and drill it before the next morning like that one.
Costs in this post are illustrative, modeled for this post from public API list prices and an assumed agent request profile (1 request per second, 60,000 input tokens at 90% cache hits, 1,200 output tokens, 20 requests per session), and are not derived from customer data. Sources: [The Register, "ChatGPT, Claude, and Grok all had outages at the same time," September 3, 2026](https://www.theregister.com/ai-and-ml/2026/09/03/chatgpt-claude-and-grok-all-had-outages-at-the-same-time/5294322). [MacDailyNews, "Major AI platforms go down in unprecedented simultaneous outage," September 3, 2026](https://macdailynews.com/2026/09/03/major-ai-platforms-go-down-in-unprecedented-simultaneous-outage/). [Anthropic prompt caching documentation](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching). [OpenAI prompt caching guide](https://platform.openai.com/docs/guides/prompt-caching). [AWS Builders' Library, "Timeouts, retries, and backoff with jitter"](https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/). [Martin Fowler, "CircuitBreaker"](https://martinfowler.com/bliki/CircuitBreaker.html). Claude Sonnet 5 pricing as covered in [our post on the cancelled Sonnet 5 price hike](/blog/sonnet-5-price-hike-cancelled-hardcoded-pricing-trap); GPT-6 prices as listed in [our GPT-6 pricing post](/blog/gpt-6-sol-luna-pricing-terra-retired-routing).