Three Down at Once

On September 3, 2026, Claude, ChatGPT and Grok went down within hours of each other, and Claude's partial API outage alone lasted 3 hours 6 minutes. Most post-mortems asked whether apps stayed up. This one asks what the failover cost. For an illustrative agent workload at 1 request per second, the same three hours cost $388 normally, $514 with a price-matched fallback to GPT-6 Sol, and $2,284 with a "best available" fallback to GPT-6 Astra: 5.9x for the same work, most of it from one line of config. The rest is prompt caches that don't travel between providers and have to be rebuilt on the way back. Includes four rules for a fallback chain, an 80-line Python failover router with per-provider circuit breakers and a cost cap, and a 30-minute failover drill.

Published 2026-09-29 by Dor Amir on the Nadir blog.

Filed under Routing & Cascades.

Your fallback has a price tag.

On the morning of September 3, 2026, three of the largest AI providers went down within the same few hours. Claude had a partial outage across Claude.ai, Claude Code, Cowork and the API that lasted 3 hours and 6 minutes, resolved at 16:16 UTC. ChatGPT and Codex were unavailable for some users for about 34 minutes because of what OpenAI called "a routing error." Grok went down after an outage at xAI's Memphis compute center. Cloudflare, which all three use, said it was "operating normally," and AWS, Google Cloud and Azure showed nothing. Source: The Register

Timeline of the September 3, 2026 outages in UTC. Claude: a partial outage from about 13:10 to 16:16, 3 hours 6 minutes, affecting the API, Claude Code and apps. ChatGPT and Codex: about 34 minutes from 14:43 to 15:17, attributed to a routing error. Grok: from 13:30, after an outage at the Memphis compute center, end time not published.
Timeline of the September 3, 2026 outages in UTC. Claude: a partial outage from about 13:10 to 16:16, 3 hours 6 minutes, affecting the API, Claude Code and apps. ChatGPT and Codex: about 34 minutes from 14:43 to 15:17, attributed to a routing error. Grok: from 13:30, after an outage at the Memphis compute center, end time not published.

Most of the write-ups since have asked the reliability question: did your app stay up? That question has a known answer. Put a second provider behind the first one and fail over. This post asks the question that comes after it: what did the failover cost? In our model it ranged from 1.3x to 5.9x the normal bill for the same three hours of work, depending on one line of config, which most teams wrote once and never priced.

Costs in this post are illustrative, modeled from public API list prices and an assumed agent workload, not measured production traces. Not derived from proprietary customer data. Sources cited throughout.

Three ways an outage shows up on the bill.

1. The fallback model has a different price. "Fall back to the best model on the other provider" is the most common chain we see, and it's usually a tier jump. Claude Sonnet 5 costs $2 / $10 per million input / output tokens. GPT-6 Astra costs $10 / $50. The failover works, and every request in the window costs five times more.

2. The prompt cache doesn't come with you. Agents re-send a large prefix on every call and rely on cache reads to make that affordable. A cache lives with one provider. Fail over, and every active session pays full input price once on the new provider. Fail back, and it pays again, because a 3-hour outage outlasts every cache TTL on the provider you left. On Anthropic, that rebuild is a cache write, billed at 1.25x the base input price. We covered the mechanics in The Cache-Write Tax and the switch penalty: an outage forces both sides of that switch on every session at once.

3. Naive retries cost time before they cost money. Most 5xx and overload errors aren't billed. But a client that retries the dead provider three times with a 30-second timeout spends 90 seconds per request before reaching the fallback. For an agent that makes 40 calls per task, that's the difference between a slow task and one that times out. Retry storms are a well-documented way to turn a partial outage into a full one.

The math for one outage.

Take a mid-size agent product doing 1 request per second, around the clock. Each request carries a 60,000-token prefix with 90% cache hits and produces 1,200 output tokens including reasoning. Primary model: Claude Sonnet 5. Average session length: 20 requests.

ModelInputCache readCache writeOutputWarm requestCold request
Claude Sonnet 5$2.00$0.20$2.50$10.00$0.0348$0.162
GPT-6 Sol$2.00$0.20n/a$10.00$0.0348$0.132
GPT-6 Astra$10.00$1.00n/a$50.00$0.174$0.660

Price per million tokens. OpenAI caches automatically with no write premium; a cold OpenAI request pays the base input rate.

A 3-hour-6-minute outage is 11,160 requests across 558 sessions. With 558 cold starts on the way over and 558 cache rewrites on the way back:

Bar chart of the cost of 11,160 agent requests during a 3 hour 6 minute outage. No outage, Sonnet 5 with a warm cache: $388. Price-matched failover to GPT-6 Sol: $514, 1.3x, from cold caches going over and coming back. Best-available failover to GPT-6 Astra: $2,284, 5.9x. No fallback: $0 and 11,160 failed requests.
Bar chart of the cost of 11,160 agent requests during a 3 hour 6 minute outage. No outage, Sonnet 5 with a warm cache: $388. Price-matched failover to GPT-6 Sol: $514, 1.3x, from cold caches going over and coming back. Best-available failover to GPT-6 Astra: $2,284, 5.9x. No fallback: $0 and 11,160 failed requests.

$1,900 for one morning isn't a crisis. But September 3 wasn't the only incident this year, and outages cluster. A fallback chain is a pricing decision you make once and pay for at the worst possible time, with nobody watching the bill.

Four rules for a fallback chain that doesn't surprise finance.

Match price, not prestige. The fallback for a mid-tier model is a mid-tier model on another provider. Sonnet 5 and GPT-6 Sol have identical list prices, which makes them natural pairs. Keep the flagship as a fallback for requests that were already going to a flagship.

Match per request, not per app. A single app-wide fallback treats a one-line classification and a 200-file refactor the same way. If traffic is already split by difficulty, the fallback map is tier to tier: simple to simple, complex to complex.

Fail fast with a circuit breaker. After a handful of consecutive failures, stop calling the primary. Send one probe after a cool-down, and switch back only when the probe succeeds.

Cap the multiplier. Decide in advance how much more you're willing to pay during an incident, and enforce it in code. Past the cap, drop to a cheaper tier or queue non-interactive work instead of paying 5x for a nightly batch job.

State diagram of a circuit breaker. Closed: calls go to the primary. After 5 failures it moves to Open: skip the primary and use the fallback. After 60 seconds it moves to Half-open: send one probe to the primary. If the probe succeeds it returns to Closed, and the cache is cold. If the probe fails it reopens.
State diagram of a circuit breaker. Closed: calls go to the primary. After 5 failures it moves to Open: skip the primary and use the fallback. After 60 seconds it moves to Half-open: send one probe to the primary. If the probe succeeds it returns to Closed, and the cache is cold. If the probe fails it reopens.

Tutorial: a cost-aware failover chain in 80 lines.

This is a provider-agnostic sketch in Python. It keeps a circuit breaker per provider, maps each tier to a price-matched fallback, refuses fallbacks above a cost multiplier, and logs what the incident is costing while it happens. Swap call_provider for your SDK calls.

Step 1: prices and the fallback map.

import time, logging
from dataclasses import dataclass, field

log = logging.getLogger("failover")

# USD per 1M tokens: input, cache read, output
PRICES = {
    "anthropic/claude-haiku-4-5": (1.00, 0.10, 5.00),
    "anthropic/claude-sonnet-5":  (2.00, 0.20, 10.00),
    "anthropic/claude-opus-5":    (5.00, 0.50, 25.00),
    "openai/gpt-6-luna":          (0.10, 0.01, 0.50),
    "openai/gpt-6-sol":           (2.00, 0.20, 10.00),
    "openai/gpt-6-astra":         (10.00, 1.00, 50.00),
}

# Tier-matched chains: primary first, then price-matched fallbacks.
CHAINS = {
    "simple":  ["anthropic/claude-haiku-4-5", "openai/gpt-6-luna"],
    "medium":  ["anthropic/claude-sonnet-5", "openai/gpt-6-sol"],
    "complex": ["anthropic/claude-opus-5", "openai/gpt-6-astra", "openai/gpt-6-sol"],
}

MAX_MULTIPLIER = 1.5   # never fail over to more than 1.5x the primary's price per request

def est_cost(model, inp, cached, out):
    pi, pc, po = PRICES[model]
    return ((inp - cached) * pi + cached * pc + out * po) / 1e6

Step 2: a circuit breaker per provider.

@dataclass
class Breaker:
    threshold: int = 5          # consecutive failures before opening
    cooldown: float = 60.0      # seconds before one probe is allowed
    failures: int = 0
    opened_at: float | None = None
    probing: bool = False

    def allow(self) -> bool:
        if self.opened_at is None:
            return True
        if not self.probing and time.monotonic() - self.opened_at >= self.cooldown:
            self.probing = True     # half-open: let exactly one request through
            return True
        return False

    def success(self):
        if self.opened_at is not None:
            log.warning("breaker closed; expect cold caches on the primary")
        self.failures, self.opened_at, self.probing = 0, None, False

    def failure(self):
        self.failures += 1
        if self.probing or self.failures >= self.threshold:
            self.opened_at, self.probing = time.monotonic(), False

breakers: dict[str, Breaker] = {}

def provider(model: str) -> str:
    return model.split("/", 1)[0]

The breaker is per provider, not per model. On September 3 the Claude outage hit Sonnet 5 "and other models," so failing over from Sonnet to Opus on the same provider would have bought nothing.

Step 3: the call path.

class ProviderError(Exception): ...     # raise from call_provider on 429, 5xx, 529
class AllProvidersDown(Exception): ...

def complete(tier: str, messages, est_in: int, est_cached: int, est_out: int):
    chain = CHAINS[tier]
    primary_cost = est_cost(chain[0], est_in, est_cached, est_out)
    for i, model in enumerate(chain):
        b = breakers.setdefault(provider(model), Breaker())
        if not b.allow():
            continue
        # Cap on like-for-like price, so a cold cache alone never blocks a fallback.
        ratio = est_cost(model, est_in, est_cached, est_out) / primary_cost
        if i > 0 and ratio > MAX_MULTIPLIER:
            log.info("skip %s: %.1fx primary price", model, ratio)
            continue
        # What this request will actually cost: a fallback starts cold.
        cost = est_cost(model, est_in, est_cached if i == 0 else 0, est_out)
        try:
            resp = call_provider(model, messages, timeout=30)
        except (TimeoutError, ConnectionError, ProviderError) as e:
            b.failure()
            log.warning("%s failed: %s", model, e)
            continue
        b.success()
        if i > 0:
            log.warning("served by fallback %s at ~$%.4f (primary ~$%.4f)",
                        model, cost, primary_cost)
        return resp
    raise AllProvidersDown(tier)

Two details carry most of the value. The cap compares like with like, list price against list price on the same request, so the unavoidable cold-cache premium never blocks a price-matched fallback. The logged cost is priced cold, because that's what the first request of each session will actually cost. And the check runs before the call, so the cap holds while nobody is looking. With MAX_MULTIPLIER = 1.5, a medium request fails over from Sonnet 5 to Sol at 1.0x, and a complex request skips Astra (2x Opus 5) and lands on Sol. That's the right call for a background job and the wrong one for your largest customer's hardest task. Set the cap per API key or per customer tier, not globally.

Step 4: run a drill.

Pick a quiet hour, open the breaker for your primary provider by hand, and let real traffic run on the fallback for 30 minutes. Then look at three numbers: cost per request against the primary, error rate, and the share of structured outputs that still parse. The last one catches a problem the bill doesn't: a fallback model that formats tool calls differently. We covered that failure in Right Answer, Wrong Shape. A drill costs 30 minutes of fallback pricing. Finding out on the next outage costs the whole outage.

What to check in your own config this week.

Where Nadir fits.

Nadir already makes the per-request decision the tier-matched rules depend on. Each prompt is classified as simple, medium or complex and routed to the cheapest model that clears your quality bar, so the fallback for a request is chosen at the same tier as the request itself, not from one app-wide list.

Underneath, a provider health monitor scores every provider on recent successes and failures, and a circuit breaker skips a provider after repeated 429s, 5xx errors or timeouts, then probes it before sending traffic back. You can pin your own chain per request with route="fallback" and fallback_models=[...], or save one per API key in the dashboard. Every response reports the model that actually answered and what it cost, so an incident shows up as a line on the dashboard the same day, not as a surprise on the invoice.

It's OpenAI compatible: change the base URL, set model="auto", keep your own provider keys. Start with a free key, set a cross-provider chain, and run the 30-minute drill above against it.

Conclusion.

September 3 was a reliability story, and the fix for reliability is well known: a second provider and a failover chain. The part that gets less attention is that the chain is also a price list, applied automatically during the hours nobody is watching. In our model, the same three hours cost $514 with a price-matched fallback and $2,284 with a "best available" one. Match fallbacks by price and by request, fail fast with a breaker, cap the multiplier in code, and drill it before the next morning like that one.


Costs in this post are illustrative, modeled for this post from public API list prices and an assumed agent request profile (1 request per second, 60,000 input tokens at 90% cache hits, 1,200 output tokens, 20 requests per session), and are not derived from customer data. Sources: [The Register, "ChatGPT, Claude, and Grok all had outages at the same time," September 3, 2026](https://www.theregister.com/ai-and-ml/2026/09/03/chatgpt-claude-and-grok-all-had-outages-at-the-same-time/5294322). [MacDailyNews, "Major AI platforms go down in unprecedented simultaneous outage," September 3, 2026](https://macdailynews.com/2026/09/03/major-ai-platforms-go-down-in-unprecedented-simultaneous-outage/). [Anthropic prompt caching documentation](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching). [OpenAI prompt caching guide](https://platform.openai.com/docs/guides/prompt-caching). [AWS Builders' Library, "Timeouts, retries, and backoff with jitter"](https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/). [Martin Fowler, "CircuitBreaker"](https://martinfowler.com/bliki/CircuitBreaker.html). Claude Sonnet 5 pricing as covered in [our post on the cancelled Sonnet 5 price hike](/blog/sonnet-5-price-hike-cancelled-hardcoded-pricing-trap); GPT-6 prices as listed in [our GPT-6 pricing post](/blog/gpt-6-sol-luna-pricing-terra-retired-routing).

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.