All 20 products

Grouped by what each one actually is, because the categories are not interchangeable. A marketplace and an inference provider solve different problems, and comparing their prices directly will mislead you.

Kind of product

Pick one. These are different kinds of product, not tiers.

Must have

Stack as many as you like — a product must satisfy every one.

Where it runs

Every product here is reachable from the US, so there is no filter for that. These two ask the questions that actually differ: whose courts the operator answers to, and whether your data can be pinned to an EU region.

How they charge

The fee structure each vendor publishes. No estimate involved — every figure behind these is on the product page with its source.

any amount

An estimate, not a quote. Models the early product workload — 50M input and 10M output tokens, 250,000 requests, 3 seats — and counts only what the gateway adds, not the model tokens you would pay anyone. Change the assumptions.

Managed marketplace

One account and one key gets you hundreds of models from dozens of providers. Fastest way to start, widest catalog, least control over the data path.

Managed marketplace Managed only

Hosted marketplace that routes one OpenAI-compatible API to models from many inference providers.

US company · EU region available

Acquisition pending

Cost above the model bill
5.5% to top up
No markup on tokens, fee applies when you add funds
Models
400–500
across ~83 providers
  • Failover
  • Spend limits
  • Logs
  • Caching

Best for Teams that want the broadest possible model and provider catalog behind one OpenAI-compatible key with unified billing.

Managed marketplace Managed only

Hosted router with a flat 5% fee on inference, EU data residency and enterprise governance controls.

UK company · EU region available

Cost above the model bill
+5% on tokens
Added to the underlying model price
Models
160–600
count not published
  • Failover
  • Spend limits
  • Logs
  • Semantic cache
  • Guardrails

Best for European teams that want one router with EU data residency, PII scrubbing and a predictable flat 5% fee.

Back to top ↑

Managed gateway

A hosted control plane in front of the models: routing, caching, guardrails, spend limits, and logs. You trade a little simplicity for governance.

Managed gateway Managed or self-host Open core

Multi-provider gateway inside Braintrust's eval and observability platform, with caching and span-level tracing.

US company · EU region available

Cost above the model bill
See pricing
Not published in a directly comparable form
Models
~100
across 9 providers
  • Failover
  • Logs
  • Caching

Best for Teams that already run evals and tracing in Braintrust and want their model traffic to flow through the same platform.

Managed gateway Managed or self-host Apache-2.0

Open-source LLM observability platform with an OpenAI-compatible AI gateway attached.

US company · EU region available

Maintenance mode

Cost above the model bill
No markup
You pay the model provider’s own prices
Models
~100
across 20–100 providers
  • Failover
  • Logs
  • Caching

Best for Teams whose main need is per-request, per-user LLM observability with a thin gateway bolted on, ideally self-hosted.

Managed gateway Managed or self-host Open core

AI plugins on the Kong API gateway, adding LLM routing, guardrails and token limits to existing API infrastructure.

US company

Cost above the model bill
See pricing
Not published in a directly comparable form
Models
across 17 providers
  • Failover
  • Logs
  • Semantic cache
  • Guardrails

Best for Enterprises already standardized on Kong for API management that want AI traffic governed by the same gateway, plugins and ops tooling.

Managed gateway Managed or self-host

Managed EU-hosted AI gateway and router bundled with evaluation, observability and governance tooling.

Netherlands company · EU region available

Cost above the model bill
4.5% to top up
No markup on tokens, fee applies when you add funds
Models
~500
count not published
  • Failover
  • Spend limits
  • Logs
  • Caching
  • Guardrails

Best for European teams that want an EU-resident managed gateway with governance, evaluations and observability in one platform rather than a bare proxy.

Managed gateway Managed or self-host Open core

Open-core AI gateway with a hosted control plane for observability, prompt management and governance.

US company

Acquired

Cost above the model bill
See pricing
Not published in a directly comparable form
Models
250–2,300
across 40–48 providers
  • Failover
  • Spend limits
  • Logs
  • Semantic cache
  • Guardrails

Best for Teams that want governance, guardrails and prompt management in one control plane, with the option to self-host the routing engine.

Managed gateway Managed only

Vercel-operated gateway that routes AI SDK and OpenAI-format requests to many providers with zero token markup.

US company

Cost above the model bill
No markup
You pay the model provider’s own prices
Models
200–350
count not published
  • Failover
  • Spend limits
  • Logs
  • Caching

Best for Teams already shipping on Vercel with the AI SDK who want zero token markup and no credit-purchase fee.

Back to top ↑

Open source

You run the gateway yourself. Full control over the data path and no per-request vendor fee, in exchange for operating it.

Open source Self-host only Apache-2.0

AI proxy plugins for the Apache APISIX API gateway, adding LLM routing and token limits to an OpenResty data plane.

US company

Cost above the model bill
Self-hosted
No license cost at all; you run APISIX and its etcd control store on your own infrastructure. The docs are co-branded with API7, a commercial vendor offering a supported distribution, whose pricing is not published on these pages.
Models
across 10–20 providers
  • Failover
  • Logs
  • Guardrails

Best for Teams already running APISIX that want basic multi-provider LLM proxying, retries and token-based rate limiting without adding another gateway.

Open source Self-host only Apache-2.0

Go-based open-source AI gateway focused on low proxy overhead, with an enterprise tier for clustering and SSO.

US company

Cost above the model bill
Self-hosted
OSS is free (Apache-2.0); infra cost only. Enterprise adds guardrails, cluster mode, adaptive load balancing, SAML/OIDC SSO, vault integration, log exports, audit logs, RBAC and SLAs at custom pricing after a 14-day trial; VPC, on-prem and air-gapped installs are enterprise options.
Models
across 20–23 providers
  • Failover
  • Spend limits
  • Logs
  • Semantic cache
  • Guardrails

Best for Teams that want a fast Go gateway they can self-host for free and later buy clustering, SSO and guardrails from a single vendor.

Open source Self-host only MIT Security incident on record

Open-source AI gateway and Python SDK that puts one OpenAI-compatible API in front of many LLM providers.

US company

Cost above the model bill
Self-hosted
OSS is free (MIT); you pay only for your own containers plus Postgres and Redis. Enterprise features (SSO, RBAC, JWT auth, SCIM, audit logs, support SLAs) require a paid LiteLLM commercial license whose price is not published.
Models
count not published
  • Failover
  • Spend limits
  • Logs
  • Semantic cache
  • Guardrails

Best for Platform teams that want a free, self-hosted, maximally broad provider abstraction with virtual keys and budgets, at moderate request volumes.

Open source Managed or self-host AGPL-3.0

AGPL-licensed OpenAI-compatible gateway available as one self-hosted Docker image or a hosted service with credit fees.

US company

Cost above the model bill
5% to top up
No markup on tokens, fee applies when you add funds
Models
~200
across ~40 providers
  • Failover
  • Spend limits
  • Logs
  • Caching
  • Guardrails

Best for Small teams that want an OpenRouter-style hosted gateway with a genuine zero-cost self-host escape hatch and no token markup.

Back to top ↑

Cloud platform

A hyperscaler surface offering several vendors models under one contract, one bill, and one identity system. Strong compliance story, limited to that cloud.

Cloud platform Managed only

AWS-managed service for calling foundation models from 19 model providers through one AWS API.

US company

Cost above the model bill
See pricing
Not published in a directly comparable form
Models
~100
count not published
  • Caching
  • Guardrails

Best for Teams already standardized on AWS that need many model vendors behind one IAM-governed, compliance-attested API.

Cloud platform Managed or self-host

Microsoft's Azure platform for deploying models from its own and partner catalogs, now branded Microsoft Foundry.

US company · EU region available

Cost above the model bill
See pricing
Not published in a directly comparable form
Models
~10,000
count not published
  • Logs
  • Guardrails

Best for Microsoft-centric enterprises that want first-party OpenAI models plus a very large partner catalog under Azure governance.

Cloud platform Managed only

Edge proxy in front of a curated set of AI providers, with caching, rate limiting, DLP and analytics.

US company

Cost above the model bill
5% to top up
No markup on tokens, fee applies when you add funds
Models
across 24 providers
  • Failover
  • Spend limits
  • Logs
  • Caching
  • Guardrails

Best for Teams already on Cloudflare Workers who want free caching, analytics, spend limits, DLP and guardrails at the edge.

Cloud platform Managed only

Google Cloud's model platform for Gemini plus 200+ Model Garden models, now branded Gemini Enterprise Agent Platform.

US company · EU region available

Cost above the model bill
See pricing
Not published in a directly comparable form
Models
~200
count not published
  • Caching

Best for Google Cloud customers who want Gemini alongside third-party models with strong, explicitly documented EU residency and ZDR controls.

Cloud platform Managed or self-host

Closed-source enterprise AI gateway sold on request tiers, deployable as SaaS or inside the customer's own cloud.

India company

Cost above the model bill
See pricing
Not published in a directly comparable form
Models
~1,000
across 27 providers
  • Failover
  • Spend limits
  • Logs
  • Semantic cache
  • Guardrails

Best for Enterprises that want a fully managed or in-VPC AI gateway with guardrails, MCP governance and SSO/RBAC, and are comfortable with closed source.

Back to top ↑

Inference provider

Hosts open-weight models on its own hardware. Often the cheapest or fastest route to a specific open model, but it is one source, not a router.

Inference provider Managed only

Inference provider serving open-weight models on its own stack, with fine-tuning and dedicated deployments.

US company

Cost above the model bill
See pricing
Not published in a directly comparable form
Models
~100
count not published

Best for Teams wanting fast serving of open-weight models with tiered latency options and both supervised and reinforcement fine-tuning.

Inference provider Managed only

Inference provider running open-weight models on its own LPU hardware for very high output speed.

US company · no EU region

Cost above the model bill
See pricing
Not published in a directly comparable form
Models
6–16
count not published

Best for Latency-sensitive, user-facing applications on open-weight models where output speed matters more than catalog breadth.

Inference provider Managed only

Inference provider running open-weight models on its own GPUs, with fine-tuning and dedicated endpoints.

US company

Cost above the model bill
See pricing
Not published in a directly comparable form
Models
19
count not published
  • Guardrails

Best for Teams that want a broad open-weight catalog plus cheap fine-tuning and the option to move to dedicated GPUs on one vendor.

Back to top ↑