All 20 products
Grouped by what each one actually is, because the categories are not
interchangeable. A marketplace and an inference provider solve different problems,
and comparing their prices directly will mislead you.
Managed marketplace
One account and one key gets you hundreds of models from dozens of providers. Fastest way to start, widest catalog, least control over the data path.
Managed marketplace Managed only
Hosted marketplace that routes one OpenAI-compatible API to models from many inference providers.
US company · EU region available
Acquisition pending
- Cost above the model bill
- 5.5% to top up
- No markup on tokens, fee applies when you add funds
- Models
- 400–500
- across ~83 providers
- Failover
- Spend limits
- Logs
- Caching
Best for Teams that want the broadest possible model and provider catalog behind one OpenAI-compatible key with unified billing.
Managed marketplace Managed only
Hosted router with a flat 5% fee on inference, EU data residency and enterprise governance controls.
UK company · EU region available
- Cost above the model bill
- +5% on tokens
- Added to the underlying model price
- Models
- 160–600
- count not published
- Failover
- Spend limits
- Logs
- Semantic cache
- Guardrails
Best for European teams that want one router with EU data residency, PII scrubbing and a predictable flat 5% fee.
Back to top ↑ Managed gateway
A hosted control plane in front of the models: routing, caching, guardrails, spend limits, and logs. You trade a little simplicity for governance.
Managed gateway Managed or self-host Open core
Multi-provider gateway inside Braintrust's eval and observability platform, with caching and span-level tracing.
US company · EU region available
- Cost above the model bill
- See pricing
- Not published in a directly comparable form
- Models
- ~100
- across 9 providers
Best for Teams that already run evals and tracing in Braintrust and want their model traffic to flow through the same platform.
Managed gateway Managed or self-host Apache-2.0
Open-source LLM observability platform with an OpenAI-compatible AI gateway attached.
US company · EU region available
Maintenance mode
- Cost above the model bill
- No markup
- You pay the model provider’s own prices
- Models
- ~100
- across 20–100 providers
Best for Teams whose main need is per-request, per-user LLM observability with a thin gateway bolted on, ideally self-hosted.
Managed gateway Managed or self-host Open core
AI plugins on the Kong API gateway, adding LLM routing, guardrails and token limits to existing API infrastructure.
US company
- Cost above the model bill
- See pricing
- Not published in a directly comparable form
- Models
- —
- across 17 providers
- Failover
- Logs
- Semantic cache
- Guardrails
Best for Enterprises already standardized on Kong for API management that want AI traffic governed by the same gateway, plugins and ops tooling.
Managed gateway Managed or self-host
Managed EU-hosted AI gateway and router bundled with evaluation, observability and governance tooling.
Netherlands company · EU region available
- Cost above the model bill
- 4.5% to top up
- No markup on tokens, fee applies when you add funds
- Models
- ~500
- count not published
- Failover
- Spend limits
- Logs
- Caching
- Guardrails
Best for European teams that want an EU-resident managed gateway with governance, evaluations and observability in one platform rather than a bare proxy.
Managed gateway Managed or self-host Open core
Open-core AI gateway with a hosted control plane for observability, prompt management and governance.
US company
Acquired
- Cost above the model bill
- See pricing
- Not published in a directly comparable form
- Models
- 250–2,300
- across 40–48 providers
- Failover
- Spend limits
- Logs
- Semantic cache
- Guardrails
Best for Teams that want governance, guardrails and prompt management in one control plane, with the option to self-host the routing engine.
Managed gateway Managed only
Vercel-operated gateway that routes AI SDK and OpenAI-format requests to many providers with zero token markup.
US company
- Cost above the model bill
- No markup
- You pay the model provider’s own prices
- Models
- 200–350
- count not published
- Failover
- Spend limits
- Logs
- Caching
Best for Teams already shipping on Vercel with the AI SDK who want zero token markup and no credit-purchase fee.
Back to top ↑ Open source Self-host only Apache-2.0
AI proxy plugins for the Apache APISIX API gateway, adding LLM routing and token limits to an OpenResty data plane.
US company
- Cost above the model bill
- Self-hosted
- No license cost at all; you run APISIX and its etcd control store on your own infrastructure. The docs are co-branded with API7, a commercial vendor offering a supported distribution, whose pricing is not published on these pages.
- Models
- —
- across 10–20 providers
Best for Teams already running APISIX that want basic multi-provider LLM proxying, retries and token-based rate limiting without adding another gateway.
Open source Self-host only Apache-2.0
Go-based open-source AI gateway focused on low proxy overhead, with an enterprise tier for clustering and SSO.
US company
- Cost above the model bill
- Self-hosted
- OSS is free (Apache-2.0); infra cost only. Enterprise adds guardrails, cluster mode, adaptive load balancing, SAML/OIDC SSO, vault integration, log exports, audit logs, RBAC and SLAs at custom pricing after a 14-day trial; VPC, on-prem and air-gapped installs are enterprise options.
- Models
- —
- across 20–23 providers
- Failover
- Spend limits
- Logs
- Semantic cache
- Guardrails
Best for Teams that want a fast Go gateway they can self-host for free and later buy clustering, SSO and guardrails from a single vendor.
Open-source AI gateway and Python SDK that puts one OpenAI-compatible API in front of many LLM providers.
US company
- Cost above the model bill
- Self-hosted
- OSS is free (MIT); you pay only for your own containers plus Postgres and Redis. Enterprise features (SSO, RBAC, JWT auth, SCIM, audit logs, support SLAs) require a paid LiteLLM commercial license whose price is not published.
- Models
- —
- count not published
- Failover
- Spend limits
- Logs
- Semantic cache
- Guardrails
Best for Platform teams that want a free, self-hosted, maximally broad provider abstraction with virtual keys and budgets, at moderate request volumes.
Open source Managed or self-host AGPL-3.0
AGPL-licensed OpenAI-compatible gateway available as one self-hosted Docker image or a hosted service with credit fees.
US company
- Cost above the model bill
- 5% to top up
- No markup on tokens, fee applies when you add funds
- Models
- ~200
- across ~40 providers
- Failover
- Spend limits
- Logs
- Caching
- Guardrails
Best for Small teams that want an OpenRouter-style hosted gateway with a genuine zero-cost self-host escape hatch and no token markup.
Back to top ↑ Cloud platform
A hyperscaler surface offering several vendors models under one contract, one bill, and one identity system. Strong compliance story, limited to that cloud.
Cloud platform Managed only
AWS-managed service for calling foundation models from 19 model providers through one AWS API.
US company
- Cost above the model bill
- See pricing
- Not published in a directly comparable form
- Models
- ~100
- count not published
Best for Teams already standardized on AWS that need many model vendors behind one IAM-governed, compliance-attested API.
Cloud platform Managed or self-host
Microsoft's Azure platform for deploying models from its own and partner catalogs, now branded Microsoft Foundry.
US company · EU region available
- Cost above the model bill
- See pricing
- Not published in a directly comparable form
- Models
- ~10,000
- count not published
Best for Microsoft-centric enterprises that want first-party OpenAI models plus a very large partner catalog under Azure governance.
Cloud platform Managed only
Edge proxy in front of a curated set of AI providers, with caching, rate limiting, DLP and analytics.
US company
- Cost above the model bill
- 5% to top up
- No markup on tokens, fee applies when you add funds
- Models
- —
- across 24 providers
- Failover
- Spend limits
- Logs
- Caching
- Guardrails
Best for Teams already on Cloudflare Workers who want free caching, analytics, spend limits, DLP and guardrails at the edge.
Cloud platform Managed only
Google Cloud's model platform for Gemini plus 200+ Model Garden models, now branded Gemini Enterprise Agent Platform.
US company · EU region available
- Cost above the model bill
- See pricing
- Not published in a directly comparable form
- Models
- ~200
- count not published
Best for Google Cloud customers who want Gemini alongside third-party models with strong, explicitly documented EU residency and ZDR controls.
Cloud platform Managed or self-host
Closed-source enterprise AI gateway sold on request tiers, deployable as SaaS or inside the customer's own cloud.
India company
- Cost above the model bill
- See pricing
- Not published in a directly comparable form
- Models
- ~1,000
- across 27 providers
- Failover
- Spend limits
- Logs
- Semantic cache
- Guardrails
Best for Enterprises that want a fully managed or in-VPC AI gateway with guardrails, MCP governance and SSO/RBAC, and are comfortable with closed source.
Back to top ↑ Inference provider
Hosts open-weight models on its own hardware. Often the cheapest or fastest route to a specific open model, but it is one source, not a router.
Inference provider Managed only
Inference provider serving open-weight models on its own stack, with fine-tuning and dedicated deployments.
US company
- Cost above the model bill
- See pricing
- Not published in a directly comparable form
- Models
- ~100
- count not published
Best for Teams wanting fast serving of open-weight models with tiered latency options and both supervised and reinforcement fine-tuning.
Inference provider Managed only
Inference provider running open-weight models on its own LPU hardware for very high output speed.
US company · no EU region
- Cost above the model bill
- See pricing
- Not published in a directly comparable form
- Models
- 6–16
- count not published
Best for Latency-sensitive, user-facing applications on open-weight models where output speed matters more than catalog breadth.
Inference provider Managed only
Inference provider running open-weight models on its own GPUs, with fine-tuning and dedicated endpoints.
US company
- Cost above the model bill
- See pricing
- Not published in a directly comparable form
- Models
- 19
- count not published
Best for Teams that want a broad open-weight catalog plus cheap fine-tuning and the option to move to dedicated GPUs on one vendor.
Back to top ↑
Nothing meets all of those requirements at once. Loosen one — the
wizard will tell you which requirement is costing you the most options.