What will this actually cost me

Most pricing comparisons quietly bury their assumptions. This one shows every line item and labels the ones it had to guess. Start with fee-only mode — it is the comparison that does not depend on predicting which models you will call.

A worked example

One workload, priced across every product, so there is a concrete answer on the page before you touch a single slider. These are the estimator's own default inputs — change them above and the numbers move, but the arithmetic is identical.

Fees only. Model token spend is excluded from every figure below. It would add roughly the same few hundred dollars to every row and drown the difference that is actually being compared.

The four ways a gateway charges you

Usage-based fees
A token markup, a credit-purchase fee, or a charge per request. The billing unit determines how it scales; minimum charges and plan allowances can also apply.
Your own infrastructure
The software is free and you pay for the machine it runs on. Only the machine is priced here; the engineering time to run it is real and is not.
A flat monthly fee
A plan charge, a per-seat charge, or both. It is the same figure at ten requests a month as at ten million, which makes it painful early and cheap at volume.
Inside the token price
No separable gateway fee, because the product is selling you the inference. Its margin is in its own per-token rate, so compare those rates instead of hunting for a fee.

Monthly gateway fee at this workload

Of the 31 products tracked here, 21 charge a fee this workload can be priced against: from $6.75 a month (Orq.ai Router) to $499 (TrueFoundry AI Gateway), a spread of $492.25 for the same traffic. The remaining 10 sell the inference themselves, so there is no gateway fee to separate out and they are listed last rather than counted as free.

Estimated monthly gateway fee for each product at 50M input and 10M output tokens, 250,000 requests and 3 seats, excluding model token spend.
Product Fee per month How it charges What that figure is
Orq.ai Router $6.75 Usage-based fees 4.5% charged when you top up, applied to $150 of spend
Requesty $7.50 Usage-based fees 5% on $150 of model spend
Merge Gateway $7.50 Usage-based fees 5% on $150 of model spend
LLM Gateway $7.50 Usage-based fees 5% charged when you top up, applied to $150 of spend
Cloudflare AI Gateway $7.50 Usage-based fees 5% charged when you top up, applied to $150 of spend
Eden AI $8.25 Usage-based fees 5.5% charged when you top up, applied to $150 of spend
OpenRouter $8.25 Usage-based fees 5.5% charged when you top up, applied to $150 of spend (minimum $0.8 per top-up)
AI Gateway HQ $25 Usage-based fees Flex gateway usage: $0.0001 × 250,000 requests; assumes all succeed, including cache hits.
Kong AI Gateway $40 Your own infrastructure Your assumption of $40/month for compute and database
Respan $40 Your own infrastructure Your assumption of $40/month for compute and database
agentgateway $40 Your own infrastructure Your assumption of $40/month for compute and database
Apache APISIX AI Gateway $40 Your own infrastructure Your assumption of $40/month for compute and database
Bifrost $40 Your own infrastructure Your assumption of $40/month for compute and database
Envoy AI Gateway $40 Your own infrastructure Your assumption of $40/month for compute and database
Higress $40 Your own infrastructure Your assumption of $40/month for compute and database
LiteLLM $40 Your own infrastructure Your assumption of $40/month for compute and database
New API $40 Your own infrastructure Your assumption of $40/month for compute and database
Portkey $49 A flat monthly fee $49/month (100k recorded logs, 30-day log retention)
Helicone $79 A flat monthly fee $79/month, unlimited seats, 10k requests + 1 GB free then usage-based
Braintrust Gateway $249 A flat monthly fee $249/month (includes $249 model credits, 5 GB data, 50k scores)
TrueFoundry AI Gateway $499 A flat monthly fee $499/month (1M requests, 10 users)
Hugging Face Inference Providers $0 Inside the token price No gateway fee to separate out — you buy the tokens from them directly.
Velokey $0 Inside the token price No gateway fee to separate out — you buy the tokens from them directly.
Vercel AI Gateway $0 Inside the token price No gateway fee to separate out — you buy the tokens from them directly.
MLflow AI Gateway $0 Inside the token price No gateway fee to separate out — you buy the tokens from them directly.
Amazon Bedrock $0 Inside the token price No gateway fee to separate out — you buy the tokens from them directly.
Azure AI Foundry $0 Inside the token price No gateway fee to separate out — you buy the tokens from them directly.
Google Vertex AI $0 Inside the token price No gateway fee to separate out — you buy the tokens from them directly.
Fireworks AI $0 Inside the token price No gateway fee to separate out — you buy the tokens from them directly.
Groq $0 Inside the token price No gateway fee to separate out — you buy the tokens from them directly.
Together AI $0 Inside the token price No gateway fee to separate out — you buy the tokens from them directly.

Self-hosted rows carry an assumed infrastructure figure, not a published price. Subscription rows are priced on the cheapest published tier that supports this workload, and several vendors have optional charges — zero-data-retention surcharges, storage, overage — that apply only if you use them. The estimator lists those per product. Fee fields verified to .

Common questions

How much does an LLM gateway cost per month?

On 50M input and 10M output tokens, 250,000 requests and 3 seats a month, the gateway's own fee runs from $6.75 to $499 across the 21 products here that charge a separable fee. The mechanism matters more than the headline rate: percentage fees are cheapest at low volume, flat plan and seat fees at high volume. A further 10 products sell the inference themselves and have no gateway fee to isolate. Model token spend is billed on top and is excluded from these figures.

What is the difference between a markup, a top-up fee and a seat fee?

Usage-based fees — A token markup, a credit-purchase fee, or a charge per request. The billing unit determines how it scales; minimum charges and plan allowances can also apply. Your own infrastructure — The software is free and you pay for the machine it runs on. Only the machine is priced here; the engineering time to run it is real and is not. A flat monthly fee — A plan charge, a per-seat charge, or both. It is the same figure at ten requests a month as at ten million, which makes it painful early and cheap at volume. Inside the token price — No separable gateway fee, because the product is selling you the inference. Its margin is in its own per-token rate, so compare those rates instead of hunting for a fee.

Is a gateway with no markup always the cheaper option?

No. A product can advertise zero markup and still be the most expensive row on this page, because it charges a flat platform or per-seat fee instead. Which mechanism wins depends on volume: a percentage is cheaper while traffic is small, a flat fee is cheaper once it is large. Compare the total at your own numbers rather than the headline percentage.

Are model token costs included in these figures?

No. Every figure in the worked example is gateway-attributable cost only — fees, seats, and assumed self-hosting infrastructure. Token spend is excluded on purpose: it depends on which models you call, and including it buries a fee difference of a few dollars under an estimate of several hundred. The estimator can add estimated token spend on top if you want the combined number.

Is OpenRouter's 5.5% a markup on tokens?

No — it is a top-up fee. OpenRouter charges 5.5% when you add credit to your account, not on each call, and its token markup is zero. The practical difference is that the charge is a percentage of what you load rather than of what you spend, so it does not scale with traffic once the credit is bought.