What will this actually cost me
Most pricing comparisons quietly bury their assumptions. This one shows every line item and labels the ones it had to guess. Start with fee-only mode — it is the comparison that does not depend on predicting which models you will call.
Interpret these fields: How LLM gateway pricing works · LLM gateway spending limits: stop a runaway agent bill? · Is routing destroying your prompt cache? · Self-hosted vs managed LLM gateways
A worked example
One workload, priced across every product, so there is a concrete answer on the page before you touch a single slider. These are the estimator's own default inputs — change them above and the numbers move, but the arithmetic is identical.
- 50M input tokens per month
- 10M output tokens per month
- 250,000 requests per month
- 3 developer seats
- $40 a month assumed for self-hosted infrastructure
Fees only. Model token spend is excluded from every figure below. It would add roughly the same few hundred dollars to every row and drown the difference that is actually being compared.
The four ways a gateway charges you
- Usage-based fees
- A token markup, a credit-purchase fee, or a charge per request. The billing unit determines how it scales; minimum charges and plan allowances can also apply.
- Your own infrastructure
- The software is free and you pay for the machine it runs on. Only the machine is priced here; the engineering time to run it is real and is not.
- A flat monthly fee
- A plan charge, a per-seat charge, or both. It is the same figure at ten requests a month as at ten million, which makes it painful early and cheap at volume.
- Inside the token price
- No separable gateway fee, because the product is selling you the inference. Its margin is in its own per-token rate, so compare those rates instead of hunting for a fee.
Monthly gateway fee at this workload
Of the 31 products tracked here, 21 charge a fee this workload can be priced against: from $6.75 a month (Orq.ai Router) to $499 (TrueFoundry AI Gateway), a spread of $492.25 for the same traffic. The remaining 10 sell the inference themselves, so there is no gateway fee to separate out and they are listed last rather than counted as free.
| Product | Fee per month | How it charges | What that figure is |
|---|---|---|---|
| Orq.ai Router | $6.75 | Usage-based fees | 4.5% charged when you top up, applied to $150 of spend |
| Requesty | $7.50 | Usage-based fees | 5% on $150 of model spend |
| Merge Gateway | $7.50 | Usage-based fees | 5% on $150 of model spend |
| LLM Gateway | $7.50 | Usage-based fees | 5% charged when you top up, applied to $150 of spend |
| Cloudflare AI Gateway | $7.50 | Usage-based fees | 5% charged when you top up, applied to $150 of spend |
| Eden AI | $8.25 | Usage-based fees | 5.5% charged when you top up, applied to $150 of spend |
| OpenRouter | $8.25 | Usage-based fees | 5.5% charged when you top up, applied to $150 of spend (minimum $0.8 per top-up) |
| AI Gateway HQ | $25 | Usage-based fees | Flex gateway usage: $0.0001 × 250,000 requests; assumes all succeed, including cache hits. |
| Kong AI Gateway | $40 | Your own infrastructure | Your assumption of $40/month for compute and database |
| Respan | $40 | Your own infrastructure | Your assumption of $40/month for compute and database |
| agentgateway | $40 | Your own infrastructure | Your assumption of $40/month for compute and database |
| Apache APISIX AI Gateway | $40 | Your own infrastructure | Your assumption of $40/month for compute and database |
| Bifrost | $40 | Your own infrastructure | Your assumption of $40/month for compute and database |
| Envoy AI Gateway | $40 | Your own infrastructure | Your assumption of $40/month for compute and database |
| Higress | $40 | Your own infrastructure | Your assumption of $40/month for compute and database |
| LiteLLM | $40 | Your own infrastructure | Your assumption of $40/month for compute and database |
| New API | $40 | Your own infrastructure | Your assumption of $40/month for compute and database |
| Portkey | $49 | A flat monthly fee | $49/month (100k recorded logs, 30-day log retention) |
| Helicone | $79 | A flat monthly fee | $79/month, unlimited seats, 10k requests + 1 GB free then usage-based |
| Braintrust Gateway | $249 | A flat monthly fee | $249/month (includes $249 model credits, 5 GB data, 50k scores) |
| TrueFoundry AI Gateway | $499 | A flat monthly fee | $499/month (1M requests, 10 users) |
| Hugging Face Inference Providers | $0 | Inside the token price | No gateway fee to separate out — you buy the tokens from them directly. |
| Velokey | $0 | Inside the token price | No gateway fee to separate out — you buy the tokens from them directly. |
| Vercel AI Gateway | $0 | Inside the token price | No gateway fee to separate out — you buy the tokens from them directly. |
| MLflow AI Gateway | $0 | Inside the token price | No gateway fee to separate out — you buy the tokens from them directly. |
| Amazon Bedrock | $0 | Inside the token price | No gateway fee to separate out — you buy the tokens from them directly. |
| Azure AI Foundry | $0 | Inside the token price | No gateway fee to separate out — you buy the tokens from them directly. |
| Google Vertex AI | $0 | Inside the token price | No gateway fee to separate out — you buy the tokens from them directly. |
| Fireworks AI | $0 | Inside the token price | No gateway fee to separate out — you buy the tokens from them directly. |
| Groq | $0 | Inside the token price | No gateway fee to separate out — you buy the tokens from them directly. |
| Together AI | $0 | Inside the token price | No gateway fee to separate out — you buy the tokens from them directly. |
Self-hosted rows carry an assumed infrastructure figure, not a published price. Subscription rows are priced on the cheapest published tier that supports this workload, and several vendors have optional charges — zero-data-retention surcharges, storage, overage — that apply only if you use them. The estimator lists those per product. Fee fields verified to .
Common questions
How much does an LLM gateway cost per month?
On 50M input and 10M output tokens, 250,000 requests and 3 seats a month, the gateway's own fee runs from $6.75 to $499 across the 21 products here that charge a separable fee. The mechanism matters more than the headline rate: percentage fees are cheapest at low volume, flat plan and seat fees at high volume. A further 10 products sell the inference themselves and have no gateway fee to isolate. Model token spend is billed on top and is excluded from these figures.
What is the difference between a markup, a top-up fee and a seat fee?
Usage-based fees — A token markup, a credit-purchase fee, or a charge per request. The billing unit determines how it scales; minimum charges and plan allowances can also apply. Your own infrastructure — The software is free and you pay for the machine it runs on. Only the machine is priced here; the engineering time to run it is real and is not. A flat monthly fee — A plan charge, a per-seat charge, or both. It is the same figure at ten requests a month as at ten million, which makes it painful early and cheap at volume. Inside the token price — No separable gateway fee, because the product is selling you the inference. Its margin is in its own per-token rate, so compare those rates instead of hunting for a fee.
Is a gateway with no markup always the cheaper option?
No. A product can advertise zero markup and still be the most expensive row on this page, because it charges a flat platform or per-seat fee instead. Which mechanism wins depends on volume: a percentage is cheaper while traffic is small, a flat fee is cheaper once it is large. Compare the total at your own numbers rather than the headline percentage.
Are model token costs included in these figures?
No. Every figure in the worked example is gateway-attributable cost only — fees, seats, and assumed self-hosting infrastructure. Token spend is excluded on purpose: it depends on which models you call, and including it buries a fee difference of a few dollars under an estimate of several hundred. The estimator can add estimated token spend on top if you want the combined number.
Is OpenRouter's 5.5% a markup on tokens?
No — it is a top-up fee. OpenRouter charges 5.5% when you add credit to your account, not on each call, and its token markup is zero. The practical difference is that the charge is a percentage of what you load rather than of what you spend, so it does not scale with traffic once the credit is bought.