What will this actually cost me

Most pricing comparisons quietly bury their assumptions. This one shows every line item and labels the ones it had to guess. Start with fee-only mode — it is the comparison that does not depend on predicting which models you will call.

A worked example

One workload, priced across every product, so there is a concrete answer on the page before you touch a single slider. These are the estimator's own default inputs — change them above and the numbers move, but the arithmetic is identical.

Fees only. Model token spend is excluded from every figure below. It would add roughly the same few hundred dollars to every row and drown the difference that is actually being compared.

The four ways a gateway charges you

A share of spend
A percentage, taken either as a markup on the model price or as a fee when you add funds. It costs nothing at zero traffic and never stops growing.
Your own infrastructure
The software is free and you pay for the machine it runs on. Only the machine is priced here; the engineering time to run it is real and is not.
A flat monthly fee
A plan charge, a per-seat charge, or both. It is the same figure at ten requests a month as at ten million, which makes it painful early and cheap at volume.
Inside the token price
No separable gateway fee, because the product is selling you the inference. Its margin is in its own per-token rate, so compare those rates instead of hunting for a fee.

Monthly gateway fee at this workload

Of the 20 products tracked here, 14 charge a fee this workload can be priced against: from $6.75 a month (Orq.ai Router) to $499 (TrueFoundry AI Gateway), a spread of $492.25 for the same traffic. The remaining 6 sell the inference themselves, so there is no gateway fee to separate out and they are listed last rather than counted as free.

Estimated monthly gateway fee for each product at 50M input and 10M output tokens, 250,000 requests and 3 seats, excluding model token spend.
Product Fee per month How it charges What that figure is
Orq.ai Router $6.75 A share of spend 4.5% charged when you top up, applied to $150 of spend
Requesty $7.50 A share of spend 5% on $150 of model spend
LLM Gateway $7.50 A share of spend 5% charged when you top up, applied to $150 of spend
Cloudflare AI Gateway $7.50 A share of spend 5% charged when you top up, applied to $150 of spend
OpenRouter $8.25 A share of spend 5.5% charged when you top up, applied to $150 of spend (minimum $0.8 per top-up)
Kong AI Gateway $40 Your own infrastructure Your assumption of $40/month for compute and database
Apache APISIX AI Gateway $40 Your own infrastructure Your assumption of $40/month for compute and database
Bifrost $40 Your own infrastructure Your assumption of $40/month for compute and database
LiteLLM $40 Your own infrastructure Your assumption of $40/month for compute and database
Portkey $49 A flat monthly fee $49/month (100k recorded logs, 30-day log retention)
Vercel AI Gateway $60 A flat monthly fee $20 per user per month × 3 users
Helicone $79 A flat monthly fee $79/month, unlimited seats, 10k requests + 1 GB free then usage-based
Braintrust Gateway $249 A flat monthly fee $249/month (includes $249 model credits, 5 GB data, 50k scores)
TrueFoundry AI Gateway $499 A flat monthly fee $499/month (1M requests, 10 users)
Amazon Bedrock $0 Inside the token price No gateway fee to separate out — you buy the tokens from them directly.
Azure AI Foundry $0 Inside the token price No gateway fee to separate out — you buy the tokens from them directly.
Google Vertex AI $0 Inside the token price No gateway fee to separate out — you buy the tokens from them directly.
Fireworks AI $0 Inside the token price No gateway fee to separate out — you buy the tokens from them directly.
Groq $0 Inside the token price No gateway fee to separate out — you buy the tokens from them directly.
Together AI $0 Inside the token price No gateway fee to separate out — you buy the tokens from them directly.

Self-hosted rows carry an assumed infrastructure figure, not a published price. Subscription rows are priced on the cheapest published tier that supports this workload, and several vendors have optional charges — zero-data-retention surcharges, storage, overage — that apply only if you use them. The estimator lists those per product. Fee fields last verified .

Common questions

How much does an LLM gateway cost per month?

On 50M input and 10M output tokens, 250,000 requests and 3 seats a month, the gateway's own fee runs from $6.75 to $499 across the 14 products here that charge a separable fee. The mechanism matters more than the headline rate: percentage fees are cheapest at low volume, flat plan and seat fees at high volume. A further 6 products sell the inference themselves and have no gateway fee to isolate. Model token spend is billed on top and is excluded from these figures.

What is the difference between a markup, a top-up fee and a seat fee?

A share of spend — A percentage, taken either as a markup on the model price or as a fee when you add funds. It costs nothing at zero traffic and never stops growing. Your own infrastructure — The software is free and you pay for the machine it runs on. Only the machine is priced here; the engineering time to run it is real and is not. A flat monthly fee — A plan charge, a per-seat charge, or both. It is the same figure at ten requests a month as at ten million, which makes it painful early and cheap at volume. Inside the token price — No separable gateway fee, because the product is selling you the inference. Its margin is in its own per-token rate, so compare those rates instead of hunting for a fee.

Is a gateway with no markup always the cheaper option?

No. A product can advertise zero markup and still be the most expensive row on this page, because it charges a flat platform or per-seat fee instead. Which mechanism wins depends on volume: a percentage is cheaper while traffic is small, a flat fee is cheaper once it is large. Compare the total at your own numbers rather than the headline percentage.

Are model token costs included in these figures?

No. Every figure in the worked example is gateway-attributable cost only — fees, seats, and assumed self-hosting infrastructure. Token spend is excluded on purpose: it depends on which models you call, and including it buries a fee difference of a few dollars under an estimate of several hundred. The estimator can add estimated token spend on top if you want the combined number.