Guide

How LLM gateway failover actually works

Short answer

LLM gateway failover sends a failed or slow request to an alternate target when its configured conditions permit. Retries repeat an attempt; fallbacks change the target. Useful failover depends on error eligibility, a shared deadline, spare capacity and compatible outputs. A feature flag does not establish those behaviors, especially after streaming has begun.

Work within one request deadline

Hypothetical 10-second caller deadline: reserve 1 second for application/network work, allow 3 seconds for attempt A, 0.5 seconds for backoff and 4 seconds for attempt B. Total planned time is 8.5 seconds, leaving 1.5 seconds of margin. Include every SDK, gateway and provider retry; enforce the remaining deadline on each attempt rather than giving each a fresh 10 seconds.

After a stream delivers tokens, switching models may require an application-visible restart rather than silently joining outputs. A tool-triggering request may have produced side effects before the connection failed: use operation IDs and deduplication at the tool boundary, and reconcile uncertain outcomes before replay. Two routes through the same provider, region or credentials may fail together. Test independence and fallback capacity.

Record a staging timeout, rate-limit response, partial stream and unavailable fallback. Verify the final status, attempted targets, billable usage and cancellation propagation. This is a proposed test, not a claim that the catalogue products passed. Connect the result to your spending-limit test.

Four shapes, all called failover

The catalogue records the shape of the fallback chain separately from the fact that one exists, because the shape decides what happens to everything else when the first target goes down. Three of these are mechanisms. The fourth is a product advertising failover without saying which mechanism it is.

15 Ordered list
Targets tried top to bottom. Predictable, easy to reason about, and all-or-nothing: when the first fails, your second choice absorbs the entire surge at whatever rate limit it has. Size the second target for the first one’s traffic, not for spillover.
6 Weighted split
Traffic divided by proportion. Spreads the risk, and makes any individual request’s destination non-deterministic — which is what you want for a gradual model migration and what you must account for if you need reproducible output or per-model evaluation.
2 Single alternate
One documented alternate target rather than a chain. Better than nothing and without depth: if the alternate is also degraded, there is no third option to reach for.
8 Shape not documented
Failover appears in the product’s own material, and the shape of the chain does not appear anywhere we could read it. Not a no. It means you will find out which shape it is during your first outage unless you ask first.

The single-alternate group is Azure AI Foundry and Higress, where the documented mechanism is one spillover target rather than a list. The 8 that do not publish a shape are AI Gateway HQ, Amazon Bedrock, Fireworks AI, Google Vertex AI, Groq, LLM Gateway, Together AI and Velokey. Each product page records what the vendor does say, and the page it was read from, so you can put the specific question to them rather than a general one.

Where the chain is defined matters as much as its shape. 19 of the 31 let you set it from code — per request, or in a configuration file you can commit. 4 put it in a dashboard only, which means your failover policy is a piece of production state that exists in the vendor's database and in nobody's repository. If you are in that group, screenshot the configuration and write down what it says somewhere a reviewer will find it.

The four controls, and how many you get

A fallback chain is the visible part of failover. Underneath it sit four settings that decide whether the chain ever gets used, and how much damage the attempt does. The catalogue records where each one can be set rather than whether it exists, because a value you cannot reach is the vendor’s decision, not yours.

Request timeout — 12 of 31 documented
How long one upstream call may hang before it is abandoned. 5 products let you set it per request, and 7 in a configuration file. 19 publish nothing, which does not mean there is no timeout — it means the worst case of a stuck call is not knowable from the docs. Set a deadline in your own client either way, and treat it as the real one.
Retry policy — 15 of 31 documented
How many times the same upstream is tried again, and how long the gateway waits between attempts. 2 of those keep it in a dashboard, so the policy is not reviewable in a pull request. Ask what a retry costs before you raise the count: a call that emitted tokens before failing has usually emitted billable ones.
Health checks and circuit breaking — 6 of 31 tunable
Whether the gateway notices a failing upstream and stops sending traffic to it. 9 track it automatically and do not let you tune the thresholds, which is a mechanism rather than a gap — just not one you own. 16 publish nothing. Without this, every request rediscovers the outage and pays the timeout again, so your latency floor rises for the whole incident.
Cross-region failover — 3 of 31 configurable
A vendor operating in several regions is not the same as a vendor letting you define what happens when one degrades. 5 state that the behaviour is theirs and fixed; 23 publish nothing either way. If a regional requirement is real for you, this is the field to get answered in writing before signing, because it is the hardest one to add afterwards.

Load balancingLoad balancing: Spreading traffic across several keys or providers so no single one hits its rate limit. Different from failover, which only reacts to failure. is the exception, and worth noting because it is the one reliability control most products do hand over: 21 of the 31 document a surface where traffic distribution across upstreams or keys can be configured, against 4 that fix it and 6 that say nothing. 22 advertise the capability and 20 advertise conditional routingConditional routing: Rules that pick the model per request — cheap model for simple work, expensive one for hard work, a specific model for one customer. as well. Spreading traffic before anything fails is the cheapest failover you can arrange, and for most products here it is the one that is actually available.

When the retry chain outlives the timeout

Retry envelopes and request timeouts are configured independently, so nothing stops them contradicting each other. An exponential retry chain with a handful of attempts can total tens of seconds of delay on its own. Put a call timeout well under that total in front of it and the client gives up before the chain has finished working — you pay for the upstream attempts and get the error anyway. The arithmetic is simple and it is nobody’s default; you have to do it.

11 of the 31 products document both a timeout surface and a retry surface, which makes them the only ones where the sum can be computed in advance rather than discovered. Of those, Orq.ai Router and Portkey go further and publish the arithmetic itself — the backoff schedule, the jitter, and in one case a cumulative worst-case delay — together with a warning to work out wall-clock time before setting a client timeout. The figures are on each product’s own page, with the vendor page they were read from, because they change more often than a guide does.

The related trap is silence in the other direction. When the timeout is undocumented you cannot compute the sum at all, so the only deadline you control is the one in your own HTTP client, and it should be shorter than the one your users will wait through.

One thing worth looking for specifically: Braintrust Gateway reports failover in the response headers, so the caller can tell that a request was served by something other than its first choice. That is rarer than a dashboard and more useful, because silent degradation — correct-looking answers from a cheaper or older model for a week — is the failure mode you cannot detect any other way. If your gatewayGateway: A single endpoint you send all your AI requests to, which then forwards them to whichever model you asked for. One integration instead of one per vendor. does not expose it per request, the substitute is a log field you can group by, and you should confirm one exists before you need it.

The reliability surface, product by product

All 31 products, sorted by name. Read across rather than down: a product with a documented shape, a retry policy and a timeout you can set is a different proposition from one that publishes a shape and nothing else, even though both would be described as supporting failover.

Fallback chain shape, retry surface, request timeout surface and upstream health tracking for every product in the catalogue.
Product Fallback shape Retry Timeout Health check
agentgateway Ordered list Checked 2026-09-02 In config Checked 2026-09-02 In config Check date not recorded In config Checked 2026-09-02
AI Gateway HQ Shape not documented Checked 2026-09-17 Not documented Checked 2026-09-17 Not documented Checked 2026-09-17 Not documented Checked 2026-09-17
Amazon Bedrock Shape not documented Check date not recorded Not documented Check date not recorded Not documented Check date not recorded Not documented Check date not recorded
Apache APISIX AI Gateway Weighted split Check date not recorded In config Check date not recorded In config Check date not recorded In config Check date not recorded
Azure AI Foundry Single alternate Check date not recorded Not documented Check date not recorded Not documented Check date not recorded Fixed, cannot change Check date not recorded
Bifrost Weighted split Check date not recorded In config Check date not recorded In config Check date not recorded Not documented Check date not recorded
Braintrust Gateway Ordered list Check date not recorded Not documented Check date not recorded Not documented Check date not recorded Fixed, cannot change Check date not recorded
Cloudflare AI Gateway Weighted split Check date not recorded Per request Check date not recorded Per request Check date not recorded Not documented Check date not recorded
Eden AI Ordered list Check date not recorded Not documented Check date not recorded Not documented Check date not recorded Fixed, cannot change Check date not recorded
Envoy AI Gateway Ordered list Check date not recorded In config Check date not recorded In config Check date not recorded Not documented Check date not recorded
Fireworks AI Shape not documented Check date not recorded Not documented Check date not recorded Not documented Check date not recorded Not documented Check date not recorded
Google Vertex AI Shape not documented Check date not recorded Not documented Check date not recorded Not documented Check date not recorded Fixed, cannot change Check date not recorded
Groq Shape not documented Check date not recorded Not documented Check date not recorded Not documented Check date not recorded Not documented Check date not recorded
Helicone Ordered list Check date not recorded Per request Check date not recorded Not documented Check date not recorded Fixed, cannot change Check date not recorded
Higress Single alternate Check date not recorded In config Checked 2026-09-02 In config Checked 2026-09-02 In config Checked 2026-09-02
Hugging Face Inference Providers Ordered list Check date not recorded Not documented Checked 2026-09-03 Not documented Checked 2026-09-03 Fixed, cannot change Checked 2026-09-03
Kong AI Gateway Weighted split Check date not recorded In config Check date not recorded In config Check date not recorded In config Check date not recorded
LiteLLM Weighted split Check date not recorded Per request Check date not recorded Per request Check date not recorded In config Check date not recorded
LLM Gateway Shape not documented Check date not recorded Not documented Check date not recorded Not documented Check date not recorded Not documented Check date not recorded
Merge Gateway Ordered list Checked 2026-09-02 Not documented Checked 2026-09-02 Not documented Checked 2026-09-02 Fixed, cannot change Checked 2026-09-02
MLflow AI Gateway Ordered list Check date not recorded Not documented Check date not recorded Not documented Check date not recorded Not documented Check date not recorded
New API Weighted split Checked 2026-09-02 Dashboard only Checked 2026-09-02 In config Checked 2026-09-02 Not published Check date not recorded
OpenRouter Ordered list Check date not recorded Not documented Check date not recorded Not documented Check date not recorded Fixed, cannot change Check date not recorded
Orq.ai Router Ordered list Check date not recorded Per request Check date not recorded Per request Check date not recorded Not documented Check date not recorded
Portkey Ordered list Check date not recorded In config Check date not recorded Per request Check date not recorded Not documented Check date not recorded
Requesty Ordered list Check date not recorded Dashboard only Check date not recorded Not documented Check date not recorded Not documented Check date not recorded
Respan Ordered list Checked 2026-09-15 Per request Checked 2026-09-15 Not documented Check date not recorded Not documented Check date not recorded
Together AI Shape not documented Check date not recorded Not documented Check date not recorded Not documented Check date not recorded Not documented Check date not recorded
TrueFoundry AI Gateway Ordered list Check date not recorded In config Check date not recorded Not documented Check date not recorded In config Check date not recorded
Velokey Shape not documented Check date not recorded Not documented Check date not recorded Not documented Check date not recorded Not documented Check date not recorded
Vercel AI Gateway Ordered list Check date not recorded Not documented Check date not recorded Per request Check date not recorded Fixed, cannot change Check date not recorded

A blank reads as not published rather than as no. A gateway with no documented retry policy still retries something; the catalogue records that the behaviour is undocumented, which is a finding about the documentation and not a verdict on the product. “Fixed, cannot change” is a documented answer, and a more informative one than silence.

Which side you are on

You need configurable failover if

  • A model call sits in a user-facing path with a latency budget.
  • You are already committed to more than one model provider.
  • Your incident review will ask what the gateway did, in writing.
  • You need to shift traffic gradually rather than all at once.

The vendor’s defaults are fine if

  • The work tolerates delay and duplicate-work costs are understood.
  • You call one provider and would not fail over anywhere else.
  • Your own client already enforces the deadline that matters.
  • Nobody on the team wants to own another set of thresholds.

Check before you rely on it

  • The shape of the chain, in the vendor’s own words.
  • Whether the policy lives in code or only in a dashboard.
  • The retry total against your client timeout, as arithmetic.
  • Whether a retried or failed-over call is billed more than once.
  • How you would detect a failover after it happened.

Common questions

What is the difference between a retry and a fallback?

A retry sends the same request to the same upstream again, usually after a delay, on the assumption the failure was transient. A fallback sends it somewhere else, on the assumption the upstream itself is the problem. They may be configured separately, they can be active at the same time, and the combined worst case is the retry delay plus the fallback attempt, not either one alone.

Does an ordered fallback chain protect me during a large outage?

It protects the request, not the system. An ordered list is tried top to bottom, so when the first target fails every request moves to the second one at once. Your second choice absorbs the entire surge, at whatever rate limit it happens to have. A weighted split spreads that load, at the cost of making any single request destination non-deterministic, which matters if you need reproducible output.

Why does a health check matter if I already have retries and fallbacks?

Without health tracking, every request discovers the outage again. The gateway sends traffic to the failing upstream, waits for it to fail, and only then moves on, so your latency floor rises for the whole duration. Health checks and circuit breaking take the failing upstream out of rotation so the discovery cost is paid once. Where a vendor tracks health automatically but does not let you tune the thresholds, the mechanism exists and the tuning is theirs.

Will a retried request be billed twice?

Assume so unless the vendor says otherwise. A retry is a new upstream call, and a call that produced tokens before failing has usually produced billable tokens. Streaming makes this worse, because a stream can be abandoned partway and charged for what it emitted. This is one of the reasons a generous retry policy set once and forgotten is expensive rather than safe.

A vendor publishes nothing about timeouts. Does that mean there is no timeout?

No. It means the value is unknowable from the documentation, which is recorded here as unpublished rather than as a missing feature. There is almost always a timeout somewhere — in the gateway, in a load balancer in front of it, or in your own HTTP client. The practical consequence of silence is that you cannot predict the worst case of a hung upstream call, so set your own client-side deadline and treat it as the real one.

Next