Every value links to the page it came from, so you can check any figure in one click.
Blank cells read "Not published" rather than "No" — we do not treat silence as a
negative answer. Turn on Technical in the header for the finer
fields and precise terminology.
OpenRouterVercel AI Gateway
Feature and pricing comparison of OpenRouter, Vercel AI Gateway
What kind of productMarketplaces resell many providers behind one key. Gateways add governance on top. Open-source projects you run yourself. Cloud platforms are hyperscaler surfaces. Inference providers host models on their own hardware.
Managed marketplace
Managed gateway
Who runs itManaged means the vendor operates it. Self-host means you run it on your own infrastructure. Both means you can choose.
LicenceProprietary products cannot be inspected or forked. Open licences such as MIT and Apache-2.0 let you audit, modify, and run the code without permission.
Who you would be signing withWhether the product is still an independent company, has been acquired, is a large cloud vendor’s product line, is run by a software foundation, or has been put into maintenance mode. Maintenance mode means bug fixes and security patches only — no new features.
Last shipped an updateThe date of the most recent release or version tag. A product that has not shipped in a year is a different risk from one that shipped last week, regardless of what its marketing site says.
Markup on model pricesHow much the product adds on top of what the underlying model provider charges. Zero means you pay the same per-token price you would pay the model provider directly.
Fee to add fundsA percentage charged when you top up your balance, separate from token prices. It is easy to miss because it does not appear on the per-token price list.
Can use your own provider accountsBring Your Own Key: you keep direct contracts with OpenAI, Anthropic and others, and the gateway only routes traffic. This preserves negotiated rates and committed-spend discounts.
Enterprise plan fromAnnual entry price for the enterprise tier, where one is published or credibly reported.
Not published
Not published
How the vendor makes moneyThe shape of the vendor’s bill: does the routing layer charge a percentage on top of tokens, a flat monthly fee, both, neither (because inference is the product), or nothing at all (open source with no paid tier).
Percentage on tokens or top-ups
Percentage on tokens or top-ups
How pricing works, brieflyA one-paragraph description that covers the caveats a pricing category cannot: introductory rates, per-feature meters, tier gating, and pricing that resets on a specific date.
Credit top-up fee (5.5% Stripe / 5% crypto) with a $0.80 minimum; 0% token markup; BYOK charged 5% above a monthly list-price allowance ($25k Pay-as-you-go / $200k Enterprise).
Minimum commitmentWhether the vendor requires a minimum contract term, a minimum spend, or a provisioned-capacity purchase to get its published rate.
None stated. Unused credits may expire one year after purchase.
None. Credits purchasable at any time with no obligation to renew; custom volume discounts on Enterprise.
Charges that fire after you go over an allowanceThe line items that scale with usage after an included allowance is exhausted — log storage, extra requests, per-feature meters, data export — which is where cost estimates usually go wrong.
No log/trace-retention or request-volume overage. Only after-the-fact charge is the BYOK 5% once the monthly list-price allowance is exceeded.
Trace Drains meter trace events delivered and trace data transferred; Pro plans include no allowance for either. Team-wide provider allowlist: $0.10 per 1,000 successful requests. Team-wide zero data retention: $0.10 per 1,000 requests. Custom Reporting: $0.075 per 1,000 tag/user/quota-entity writes and $5 per 1,000 reporting-endpoint queries. All billed outside credits.
Prompt cache offeredWhether the gateway offers its own response cache, what kind of match it does (exact request, prefix, semantic), or simply passes provider caching through unchanged.
Passes provider caching through
No gateway-owned cache
Discount on cached inputHow much cheaper cached tokens are than fresh input, when the vendor publishes a single figure. Bundled-inference clouds usually price this per model instead of as one number.
Not published
Not published
Premium on cache writesHow much more the first write of a cached prefix costs versus a plain input token. A high write premium and a low hit rate can leave you paying more than you save, so this matters as much as the read discount.
Not published
Not published
Who captures the cache savingWhether the customer keeps the full saving from caching or the vendor captures part of it — and any conditions attached (write premium, storage fees, best-effort hits).
Provider caching flows through unchanged. Read discounts vary by provider: 75%/50% for OpenAI, 90% for Anthropic/Alibaba/DeepSeek, 50% for Groq, ~80% for Z.AI, 75% for Gemini implicit. Cache-write premium: 0% for pre-GPT-5.6 OpenAI/Grok/Moonshot/Groq/Gemini; +25% for GPT-5.6+, Alibaba explicit and Anthropic 5-min; +100% for Anthropic 1-hour. OpenRouter itself adds 0% on cached traffic.
Pricing page does not state whether the gateway caches or passes through provider caching, and no cached-token pricing is published. Because Vercel charges provider list price with 0% markup, any provider cache discount would reach the customer unchanged.
What you can split spend byThe dimensions the vendor documents for splitting spend — per key, per user, per team, per tag, per customer. Matters if you need to chargeback internally or bill an end customer.
Per model, provider and API key on the Activity page; session_id grouping across turns. Per-user/team/customer not stated.
By tags, user IDs and quota entity IDs via Custom Reporting (queried through the reporting endpoint at $5/1k queries). Per team/key/customer not explicitly stated.
How you get cost data outThe mechanisms the vendor publishes for exporting cost and usage data: CSV, an API, webhooks, S3, a data warehouse, or nothing at all. Any per-unit price is included.
API only (/api/v1/key, credits API, /api/v1/generation). CSV, webhook, S3 and warehouse export not stated.
Reporting API plus Trace Drains (Pro/Enterprise, metered). CSV/S3/warehouse not stated.
Who pays the model billWhether you bring your own provider accounts (BYOK), buy inference from this vendor, or can do either. This is the single biggest commercial difference between these products: it decides who holds the contract with the model provider and who carries the spend.
Your keys or their credits
Your keys or their credits
Catalog
Models availableHow many models you can call, shown as a range because several vendors publish different totals on different pages. Vendor-reported either way, so counts are not directly comparable — some count every provider variant of the same model separately.
Model providers reachableHow many distinct model providers or labs you can reach, shown as a range where the vendor’s own pages disagree. More providers usually means better redundancy when one has an outage.
Works with standard OpenAI codeIf yes, you can usually switch to it by changing one base URL, and switch away just as easily. This is the main defence against lock-in.
OpenAI chat endpointThe endpoint almost every application ports first. "Not documented" means the vendor never states it, which is different from a documented no.
Yes
Yes
Anthropic messages endpointWhether Anthropic-shaped calls work without rewriting them. Several products support this only as SDK compatibility or provider passthrough rather than a native endpoint — the detail page says which.
Yes
Yes
OpenAI Responses endpointThe newer stateful OpenAI surface. Support is much thinner across this market than chat completions.
Yes
Yes
Embeddings endpointWhether you can generate vectors through the same gateway, or need a second integration for retrieval workloads.
Yes
Yes
Image generation endpointWhether image models are reachable through the same surface as text.
Yes
Yes
Audio endpointsSpeech-to-text and text-to-speech. Frequently the first gap in an otherwise complete gateway.
Yes
Yes
Batch jobs endpointAsynchronous bulk processing, usually at a discount. Commonly undocumented, and commonly the reason a migration stalls late.
Yes
Not published
Needs the vendor’s own code libraryA proprietary client library spreads through your codebase and has to be torn out again if you leave. “No” is the better answer here, and it means the standard OpenAI client works.
You can export your request historyWhether you can get your own request logs, traces, or usage records back out — through an API, a bulk export, or a download. Decides whether you leave with your history or abandon it.
Uptime it promises in writingThe uptime percentage in a published, contractual service level agreement. A public status page is not an SLA — it reports what happened, it does not promise anything or pay you back when it breaks.
Automatic failoverWhen a model provider goes down or rate-limits you, traffic moves to a backup automatically instead of returning errors to your users.
Rule-based routingSend different requests to different models based on rules — for example a cheap model for free users and a strong model for paying ones.
Similar-question cachingReuses an answer when a new question means roughly the same thing as an earlier one. Saves far more than exact-match caching but can return subtly wrong answers if tuned loosely.
Where you set the timeoutWhere a request timeout can be set: per request, in a config file, in the vendor dashboard, or nowhere. Nine of the twenty products documented here do not describe a request timeout at all, so the worst case of a hung upstream call is unknowable from the docs.
Not published
Per request
Where you set retriesWhere retry count and backoff are configured. Worth knowing alongside billing: a retried streaming call can be charged more than once.
Not published
Not published
Where you set fallbacksWhere the fallback chain is defined. Almost every product claims fallback; the useful question is whether you can change it from code or only by hand in a dashboard.
Per request
Per request
Upstream health trackingWhether the product notices a failing upstream and stops sending traffic to it, and whether you can tune the thresholds. This is what turns a provider outage into a blip rather than a sustained error rate, and only four of the twenty expose it.
Fixed, cannot change
Fixed, cannot change
Cross-region failover you controlWhether you can define what happens when a region degrades. A vendor running many regions is not the same as a vendor letting you configure failover between them; only three document a user-controlled mechanism.
Not published
Not published
Where you set load balancingWhere traffic distribution across upstreams or keys is configured.
Per request
Per request
Operations
Usage dashboards and logsBuilt-in visibility into what was sent, what came back, what it cost, and how long it took.
Separate keys per team or appIssue scoped keys with their own budgets and permissions so you can attribute cost and revoke access without rotating everything.
What gets loggedWhether full prompts and responses are stored, only metadata, or nothing. Full-body logging is the most useful debugging feature here and the one most likely to need a conversation with your compliance team.
Metadata only
Your choice
You can turn logging offWhether prompt and response bodies can be suppressed while still keeping usage metrics. Nineteen of the twenty document a way to do this; the mechanisms range from a per-request header to an organisation-wide setting.
Yes
Yes
Traces you can take elsewhereWhether the product emits OpenTelemetry, a proprietary format, or nothing. OpenTelemetry means the traces land in the tooling you already run instead of only in the vendor’s dashboard.
Where telemetry can goDocumented sinks for logs and metrics. This is a good proxy for how replaceable the vendor’s own dashboard is: around twenty destinations means you never have to depend on it, while a CSV download means you do.
CSV export
CSV export
Performance
Delay it addsExtra time the product itself adds to each request, on top of however long the model takes. Usually irrelevant next to multi-second model latency, but it matters for high-volume or streaming-sensitive workloads.
Not published
Not published
What the request path runs onThe comparable way to talk about latency here. An edge worker, a compiled Go or Rust binary, and a Python proxy have different overhead floors no matter which figures each vendor publishes. Products that never disclose their runtime are recorded as undisclosed rather than assumed.
Edge worker
Undisclosed vendor service
You can run the request path yourselfWhether the component that actually carries your prompts can run on your own infrastructure. Distinct from a vendor offering a self-hosted control plane while still proxying traffic through their network.
Not published
No
Streaming responsesWhether token-by-token streaming is documented. The caveats matter more than the yes: some products cannot cancel a stream without still being billed, and several timeout and fallback mechanisms stop applying once the first token has been sent.
Yes
Yes
Security & compliance
Does your prompt reach their serversWhether the text you send passes through this company’s own infrastructure. If it does, every other promise on this page is a policy commitment rather than a physical impossibility. Self-hosted products can answer no outright.
What they keep if you change nothingDefaults matter more than options. A product that stores full prompts and replies unless you find the right header will have stored them by the time you read the docs.
How long they keep itDefault retention for request content, in days. Zero means nothing is kept. Read the note: several products keep nothing as a rule but make timed exceptions for abuse review or specific models.
Could they train on your promptsWhether the vendor may use your prompts and outputs to train models. “Not published” means we could not find any position, which is not the same as a no — ask for it in writing.
Where it runs, and what you can pinWhich regions are offered and whether you can force processing to stay in one. A global endpoint that silently picks a region is a different compliance story from an endpoint you pin yourself.
Where safety filters runA filter that strips personal data only helps if it runs before the data leaves your boundary. If guardrails execute in the vendor’s cloud, the vendor has already received whatever you wanted redacted.
Who else touches the dataThe published list of third parties the vendor passes your data to. No list means you cannot know the full chain, which most data-protection agreements require you to.
Will sign a HIPAA agreementRequired before you may send protected health information through the service. Without a signed BAA, healthcare data is off limits.
Can keep data in the EURequests can be processed inside the EU rather than routed to US infrastructure. Often the deciding constraint for European customers.
Blocks personal data in promptsWhether personal data detection sits on the request path and can stop the call, merely inspects and forwards it, or is not documented. A control that only reports is a logging feature, not a policy control.
Can block the request
Not published
Blocks prompt injectionWhether injection and jailbreak detection can stop a request. Most products offering this call a partner classifier rather than shipping their own.
Can block the request
Not published
Blocks harmful contentWhether hate, violence, sexual and self-harm categories are checked inline and can stop a request, in either direction.
Not published
Not published
Your own policy rulesWhether you can add your own rule — a regex, a webhook, or your own classifier — rather than choosing from the vendor library.
Can block the request
Not published
Where guardrails runWhether guardrail evaluation happens inside your infrastructure or on the vendor’s servers. This decides whether the prompt you are trying to protect leaves your network in order to be checked.
On the vendor's servers
On the vendor's servers
If the guardrail itself failsWhat happens when the guardrail service times out or errors: does the request proceed unchecked, or is it blocked? This is the worst-documented field in the entire catalogue — only two vendors state it plainly, which means most teams are running a control whose failure behaviour they cannot know.
Not published
Not published
Fit & integration
Work to try itThe shape of the work on the vendor’s own quickstart, from swapping one base URL through to deploying infrastructure. An ordinal class rather than a duration, because elapsed time depends on accounts and quota we cannot see.
Change one base URL
Install a package
Work to run itThe same scale applied to the vendor’s recommended production path. For several products this is much heavier than the quickstart, which is exactly why both are recorded.
Change one base URL
Install a package
Can you self-host it todayWhether an install command is actually published. Several products advertise self-hosting while publishing no command to start from, which a plain yes/no would hide.
No self-hosting
No self-hosting
Works with the OpenAI SDKWhether an existing OpenAI-compatible client can be pointed at it by changing the base URL and key.
Yes
Yes
Vercel AI SDK supportWhether a first-party AI SDK provider package exists, or only a community package, a documented workaround, or the generic OpenAI provider pointed at a custom base URL.
Official provider package
Official provider package
Python framework integrationsLangChain, LangGraph, LlamaIndex and similar orchestration frameworks with a documented integration.
LangChain,LlamaIndex
LangChain,LlamaIndex
Callable from Cloudflare WorkersWhether the docs show calling this product from your own Worker. Deliberately separated from the several products whose own gateway runs on Workers, which is a fact about their infrastructure and not about your edge compatibility.
Their gateway runs on Workers, not yours
Not published
MCP supportWhich kind of MCP support this is: a gateway that governs many MCP servers, a hosted MCP server you connect to, MCP tools accepted inside the completion API, or client tooling. These are different products behind one acronym.
Hosted MCP server
MCP client tooling
Needs your own provider keyWhether an upstream provider account and key must exist before your first call works. A prerequisite rather than a step, and it can differ between the hosted and self-hosted forms of the same product.
Not needed
Not needed
Gate before models workWhether an enablement click, a quota grant, a paid tier or an approval form stands between a valid key and a working model call.
A capability table cannot tell you which is cheaper for your traffic. Feed in your own volume and see the arithmetic.
OpenRouter vs LiteLLM vs Vercel AI Gateway at a glance
The three products people most often weigh against each other, on the twelve fields
that usually decide it. Use the tool above for any other combination, or for all
113 fields.
Summary comparison of OpenRouter, LiteLLM, Vercel AI Gateway across pricing,
deployment, and compliance fields
"Not published" means the vendor has not stated a figure. It is not a No. Model counts
are not comparable across vendors, because counting conventions differ — treat them as
an order of magnitude, not a score. Every value here is sourced on the
provider pages, and the
methodology explains how it was checked.
Questions about this table
Why are some cells in the comparison empty?
A blank means the vendor does not publish that fact, not that the answer is no. Every value in this table is read from a vendor's own page and stores the URL it came from, so there is nothing to show when no page states it. Treat a blank as a question to ask on a sales call rather than as a mark against the product.
Why can you only compare four products at a time?
Four columns is what stays readable on a laptop screen across 113 fields without horizontal scrolling doing the comparing for you. All 31 products are selectable and the set is swappable, so a wider shortlist is a sequence of four-way reads. The full list, grouped by category, is on the products page.
How current are the figures in this comparison?
Each value carries the date it was last checked, and a field is flagged once it passes the staleness window without re-verification. The changelog records every edit with its previous value and the source that justified it. Where a figure here differs from a vendor's page, the vendor's page is newer and the correction form is the fastest fix.
Pairs we have written up
The tool above compares any combination. These 46 pairs also come with a verdict, the
single question that decides them, and the traps to watch for.
Velokey vs OpenRouter Does your workload need provider-key flexibility or consolidated chat and media access?
Velokey vs Eden AI Does your workload need a broader task API or consolidated chat and media access?
AI Gateway HQ vs Portkey Do you need a hosted governance boundary or the option to run the gateway yourself?
Groq vs Together AI Is perceived output speed your product differentiator, or do you need catalog breadth, fine-tuning and a HIPAA BAA?
Fireworks AI vs Together AI Do you need selectable latency tiers and reinforcement fine-tuning, or broader modality coverage and cheaper supervised fine-tuning?
LLM Gateway vs Requesty Do you want a fee you can escape by self-hosting, or a hosted router with EU residency that you must trust with full request logs?
LLM Gateway vs Bifrost Is a permissive licence with no hosted escape hatch worth more to you than a copyleft one that comes with a managed service?
Orq.ai Router vs Braintrust Gateway Is the surrounding platform an eval workbench or a governance layer, and does the gateway underneath it reach far enough?
Groq vs Fireworks AI Do you need maximum output speed on a handful of models, or a broad open-weight menu with a business associate agreement?
Azure AI Foundry vs Google Vertex AI Which cloud are you already in — and if that is genuinely open, does catalog breadth or the tighter logging default matter more?
New API vs LiteLLM Are you charging the people who use the gateway, or standardising provider access for your own platform?
Eden AI vs OpenRouter Is a meaningful share of your workload non-LLM — OCR, speech, document parsing — or is it all chat?
Envoy AI Gateway vs Kong AI Gateway Do you need a vendor who can sign a contract and an SLA, or a project you own outright with nothing gated behind a quote?
Higress vs Apache APISIX AI Gateway Is the deciding factor the depth of the AI plugin shelf, or request-path reliability and framework integrations that work without you assembling them?
Envoy AI Gateway vs Higress Do you need the wider plugin shelf, or upstream credentials your security review will accept and a control-plane API that will not move?
agentgateway vs Envoy AI Gateway Are you adding LLM routing to an Envoy Gateway you already run, or do you want one standalone data plane that also carries agent protocols and content guardrails?
agentgateway vs Bifrost Do you want the whole feature set in one Apache-2.0 build from a foundation project, or a generous free build from one vendor with guardrails, clustering and SSO behind a quote?
Eden AI vs Requesty Would you rather pay a fee on money in or a fee on tokens out, and which retention model can you actually sign?
New API vs LLM Gateway Do you want to run the gateway yourself and bill the people behind it, or hand the running of it to someone with an attestation?
MLflow AI Gateway vs Helicone Do you need per-user request analytics badly enough to adopt a platform whose owner is offering migration help?
Merge Gateway vs Vercel AI Gateway Is 5% of every token worth paying for content-level controls, or is a fixed seat fee with nothing on tokens the better shape?
Merge Gateway vs Portkey Are you buying a governance platform to run against your own provider keys, or buying model access with governance attached?
Together AI vs Hugging Face Inference Providers Do you need the things only the destination sells — fine-tuning, dedicated GPUs, batch, per-request controls — or do you need to not be tied to one destination?
Analytics, only if you say so
GatewayScore would like to record which comparisons, guides and tools people use, so the
research goes where it is actually read. No advertising, no data sold, no session recording,
and nothing is loaded until you choose. What gets recorded.