Guide

Does your LLM gateway promise any uptime?

Short answer

Some LLM gateways publish an uptime commitment; others offer private terms or leave availability to the self-hosting operator. An SLA is a contract defining measurement and remedies. An SLO is an operational target, and observed uptime is a measurement. Neither a published percentage nor a status page proves the reliability of your complete model request path.

Read a real agreement before relying on the number

The Amazon Bedrock SLA sets its service commitment, credit tiers, claim requirements and exclusions. The Vercel Enterprise SLA is plan- and service-scoped; do not assume a platform SLA covers every gateway feature. Reviewed September 16, 2026. Confirm the exact covered service and account terms rather than transferring a headline percentage to all inference traffic.

Record: service and edition; measurement window; eligible success/failure events; excluded causes; claim deadline and evidence; remedy cap; upstream-provider scope. Set your own application SLO and measure it independently.

Three kinds of answer, not two

Counting published figures against the whole catalogue produces a number that sounds damning and means very little, because it lumps together vendors who operate a service and say nothing about it with vendors who operate nothing at all. The three groups below sum to the 31 products here, and the groups describe different service boundaries.

9 Publish a figure
A contractual uptime percentage you can read before signing. Every one of them also links a service level document, so the exclusions are readable too.
17 Operated for you, no figure
Hosted products with no published uptime commitment. Somebody runs this service and the target is either in an agreement you have not been shown or nowhere. This is the group worth asking about, in writing, before you put it in the request path.
5 Self-host only, no figure
No hosted service exists, so there is nobody to promise you availability. The absence here is structural and holding it against the product would be a category error. Your uptime is your deployment, and the commitment you need is your own.

The hosted products with no published figure are AI Gateway HQ, Braintrust Gateway, Cloudflare AI Gateway, Eden AI, Fireworks AI, Groq, Helicone, Higress, Hugging Face Inference Providers, LLM Gateway, Merge Gateway, MLflow AI Gateway, OpenRouter, Orq.ai Router, Respan, Together AI and Velokey. None of that is a claim that they are unreliable, and several are among the largest services in the category. It is a claim about what you can hold them to, which is a different question and the only one an SLA answers. Read the other direction too: Bifrost publishes a figure despite having no managed deployment recorded, so check what that number is scoped to before you rely on it — a commitment about a hosted control plane is not a commitment about the binary in your cluster.

What each number permits

For comparison, the table converts each recorded percentage using downtime = 30 × 24 × 60 × (1 − uptime ÷ 100). Thus 99.9% corresponds to 43.2 minutes in a hypothetical 30-day month. This is arithmetic, not a contract interpretation: the actual measurement window, qualifying requests and exclusions may differ.

99.999% — about 26 seconds a month, 1 product
Bifrost. The allowance is per measurement period, it is not a prediction, and the document decides what counts as downtime in the first place.
99.99% — about 4.3 minutes a month, 2 products
Requesty and Vercel AI Gateway. The allowance is per measurement period, it is not a prediction, and the document decides what counts as downtime in the first place.
99.9% — about 43 minutes a month, 6 products
Amazon Bedrock, Azure AI Foundry, Google Vertex AI, Kong AI Gateway, Portkey and TrueFoundry AI Gateway. The allowance is per measurement period, it is not a prediction, and the document decides what counts as downtime in the first place.

A higher percentage yields a smaller arithmetic downtime budget. It does not show actual incident frequency, user impact or a guaranteed remedy. Compare the service boundary and eligibility conditions before ranking two percentages.

A percentage is not a remedy

The remedy comes from the agreement, not the percentage alone. It may be a service credit with eligibility, evidence and claim requirements. Check the covered bill, cap, exclusions and deadline; do not assume compensation for your own downstream losses.

This catalogue deliberately records no credit percentages or remedy terms, because those sit in the vendor’s service level document and change without notice. What it records is whether that document exists to be read. 12 of the 31 products link one; on the remaining 19 there is nothing public to read. 3 link a document without publishing a percentage alongside it, so inspect the document before inferring a target. Four things to find in it before the percentage: the measurement window, what counts as downtime, the exclusions, and the claim deadline. The exclusions are usually where the failure you actually care about lives.

Your gateway is now in the path

If two required components each have actual availability of 99.9% over the same window and their failures are independent, their joint availability is 0.999 × 0.999 = 99.8001%, about 86.36 unavailable minutes per 30 days in that model. Contractual SLA percentages cannot simply be multiplied into an end-to-end promise. Shared outages, retries and fallback routing change the relationship.

That is the argument for reading the mechanism as well as the percentage. 25 of the 31 products advertise automatic failoverFailover: When the provider you asked for is down or rate-limiting you, the gateway automatically retries somewhere else. The single most valuable reliability feature these products offer., and 3 let you define cross-region behaviour yourself — the control that is hardest to add after you have signed. How those mechanisms differ, and where each is configured, is the subject of the failover guide rather than this one. The short version for planning: the SLA tells you what the vendor owes you, the failover configuration tells you whether your users notice, and only one of those is yours to change.

Product by product

All 31 products, sorted by name. Read the deployment column first: it decides what the empty cells beside it mean.

Deployment model, published contractual uptime percentage, the 30-day equivalent downtime that percentage allows per 30-day month, and whether a service level document is published, for every product in the catalogue.
Product Deployment Published uptime 30-day equivalent downtime SLA doc
agentgateway Self-host only Checked 2026-09-02 Not published Check date not recorded — Not published Check date not recorded
AI Gateway HQ Managed only Checked 2026-09-17 Not published Check date not recorded — Published
Amazon Bedrock Managed only Check date not recorded 99.9% Check date not recorded 43 minutes Published
Apache APISIX AI Gateway Self-host only Check date not recorded Not published Check date not recorded — Not published Check date not recorded
Azure AI Foundry Managed or self-host Checked 2026-08-29 99.9% Check date not recorded 43 minutes Published
Bifrost Self-host only Check date not recorded 99.999% Check date not recorded 26 seconds Published
Braintrust Gateway Managed or self-host Checked 2026-08-29 Not published Check date not recorded — Not published Check date not recorded
Cloudflare AI Gateway Managed only Checked 2026-08-29 Not published Check date not recorded — Not published Check date not recorded
Eden AI Managed only Checked 2026-09-02 Not published Check date not recorded — Not published Check date not recorded
Envoy AI Gateway Self-host only Checked 2026-09-02 Not published Check date not recorded — Not published Check date not recorded
Fireworks AI Managed only Check date not recorded Not published Check date not recorded — Not published Check date not recorded
Google Vertex AI Managed only Check date not recorded 99.9% Check date not recorded 43 minutes Published
Groq Managed only Check date not recorded Not published Check date not recorded — Not published Check date not recorded
Helicone Managed or self-host Checked 2026-08-29 Not published Check date not recorded — Not published Check date not recorded
Higress Managed or self-host Checked 2026-09-02 Not published Check date not recorded — Not published Check date not recorded
Hugging Face Inference Providers Managed only Checked 2026-09-03 Not published Check date not recorded — Not published Check date not recorded
Kong AI Gateway Managed or self-host Check date not recorded 99.9% Check date not recorded 43 minutes Published
LiteLLM Self-host only Check date not recorded Not published Check date not recorded — Not published Check date not recorded
LLM Gateway Managed or self-host Check date not recorded Not published Check date not recorded — Not published Check date not recorded
Merge Gateway Managed only Checked 2026-09-02 Not published Check date not recorded — Not published Check date not recorded
MLflow AI Gateway Managed or self-host Checked 2026-09-02 Not published Check date not recorded — Not published Check date not recorded
New API Self-host only Checked 2026-09-02 Not published Check date not recorded — Not published Check date not recorded
OpenRouter Managed only Checked 2026-08-29 Not published Check date not recorded — Not published Check date not recorded
Orq.ai Router Managed or self-host Check date not recorded Not published Check date not recorded — Not published Check date not recorded
Portkey Managed or self-host Checked 2026-08-29 99.9% Check date not recorded 43 minutes Published
Requesty Managed only Checked 2026-08-29 99.99% Check date not recorded 4.3 minutes Published
Respan Managed or self-host Checked 2026-09-15 Not published Checked 2026-09-15 — Published
Together AI Managed only Check date not recorded Not published Check date not recorded — Not published Check date not recorded
TrueFoundry AI Gateway Managed or self-host Check date not recorded 99.9% Check date not recorded 43 minutes Published
Velokey Managed only Checked 2026-09-19 Not published Check date not recorded — Published
Vercel AI Gateway Managed only Checked 2026-08-29 99.99% Check date not recorded 4.3 minutes Published

A blank reads as not published rather than as no. A hosted product with no figure here may well hold itself to one internally, or offer one in an enterprise agreement; the catalogue records that nothing is public. The permitted downtime column is arithmetic on the published percentage, not a vendor claim, and an em dash there means there was no percentage to convert.

Which side you are on

An SLA matters to you if

  • Model calls sit in a path your own customers have been promised.
  • Procurement will ask for an availability commitment in writing.
  • You need a document to point at during an incident review.
  • You are consolidating several integrations behind one hosted service.

It matters less if

  • You run the gateway yourself, so the availability is already yours.
  • The work is batch and an hour of delay costs nothing.
  • Your fallback path does not depend on the vendor noticing anything.
  • Spend is small enough that any credit would be rounding.

Read before you sign

  • The measurement window, and what the vendor counts as downtime.
  • The exclusions, especially for upstream provider failures.
  • Whether credits are automatic or claim-based, and the deadline.
  • Whether the figure covers the data plane or only the control plane.
  • What the percentage permits in minutes, worked out before you agree to it.

Common questions

Is a status page the same as an SLA?

No. A status page reports what happened; an SLA promises what should happen and states what the vendor owes you when it does not. Every product here can have a status page and most do. A private enterprise agreement can also establish a commitment. Check the contract that applies to your account.

Why do so many products publish no uptime figure at all?

Two quite different reasons, which is why they are counted separately on this page. A product you run on your own infrastructure has no vendor-side availability to promise, so the absence is structural. A hosted product with no published figure is a different matter: the service exists, somebody operates it, and the commitment is either in an enterprise agreement you have not seen or nowhere. Silence is recorded here as unpublished, not as a refusal.

What does a service credit cover?

A percentage of what you spent with that vendor in the affected period, usually as credit against future use rather than money returned. It is not compensation for your own outage, and it should not be assumed to compensate your revenue loss. Treat the credit as an admission mechanism rather than insurance, and read the claim window and notice requirements, because credits are almost always claim-based rather than automatic.

If the gateway promises 99.9% and the model provider promises 99.9%, what do I have?

No end-to-end commitment follows from those two contractual numbers alone. Multiplication applies to an availability model with a common measurement window and independent failures, not automatically to SLA contracts. Adding a gateway in front of a model provider does not raise your availability by itself; it raises it only if you use the gateway to route around the provider when it fails, which is a configuration question rather than an SLA question.

Does an SLA cover an outage at the model provider behind the gateway?

Read the specific agreement. Coverage depends on the named service, deployment, eligible requests and exclusions. Do not assume an upstream outage is covered or excluded from the gateway agreement without checking its terms.

Next