How this is put together

A recommendation you cannot audit is just an opinion with better styling. This page gives you the entire scoring model, the rules that rule products out, and an honest account of where the data is thin.

Where the data comes from

Every value in the catalog was read off a vendor's own pricing page, documentation, trust centre, public repository, or live models endpoint. Nothing is filled in from general knowledge or from another comparison site. Each figure stores the URL it came from, which is why almost every number on this site is clickable.

20 products tracked
113 fields per product
1043 values with a source URL
52.1 sourced values per product

Coverage is uneven, and pretending otherwise would be the dishonest part. The thinnest entries right now are Together AI (42), Apache APISIX AI Gateway (43), Fireworks AI (43). Treat those with more caution than the well-documented ones.

Blank is not "no"

This is the single most important rule on the site. When a vendor has not published something, the cell reads Not published — never "No". A small open-source project that has never mentioned SOC 2 is in a completely different position from one that has confirmed it has no report, and collapsing those two into the same red cross would quietly punish the honest and reward the vague.

The practical consequence: the wizard will never rule a product out for a missing value. It only rules out on an explicit vendor "no". Where a requirement could not be confirmed either way, the product stays in the list and you get told what to go and check yourself.

Back to top ↑

What gets ruled out, and only what

Six answers can eliminate a product outright. Every one of them requires the vendor to have explicitly confirmed the negative.

If you said you need The rule applied
A signed HIPAA business associate agreement Ruled out only when the vendor states it does not offer a BAA.
A SOC 2 report Ruled out only when the vendor confirms it has none.
EU-only processing Ruled out only when the vendor states it cannot guarantee EU processing.
Zero data retention Ruled out only when the vendor confirms prompts are retained with no opt-out.
Nobody to run infrastructure Rules out anything available only as self-hosted software.
Embeddings, images, audio, or video Ruled out only when the vendor confirms the modality is unsupported.
Back to top ↑

The exact scoring weights

Everything not ruled out is scored. Points are awarded only for facts that are relevant to what you actually said mattered — answer that cost is your priority and the governance weights never fire at all. The published figures below are the entire model; there is nothing else behind it.

How you want to run it up to 14 pts

  • Fully managed, when you asked for managed +14
  • Managed option available, when you asked for managed +10
  • Self-hostable, when you asked to own the data path +12
  • Both deployment modes, when you want to keep the option open +11
  • Open source licence, on top of self-hosting +6

Cost sensitivity up to 16 pts

  • No markup and no top-up fee +16
  • No markup on token prices +9
  • Keeps your negotiated rates via your own keys +14
  • Bills against an existing cloud commitment +12
  • No per-user platform fee +6
  • Caching to cut repeat spend +4

Model choice up to 16 pts

  • 400 or more models +16
  • 200 to 399 models +12
  • 100 to 199 models +8
  • Fewer than 100 models +3
  • 50 or more upstream providers +8
  • 20 to 49 upstream providers +5

Reliability up to 14 pts

  • Automatic failover +14
  • Load balancing across keys and providers +10
  • Rule-based conditional routing +6
  • Low measured latency overhead +4

Governance and control up to 12 pts

  • Hard spending limits +12
  • Scoped keys per team or app +10
  • Built-in logs and usage dashboards +8
  • Content and policy guardrails +6
  • Personal-data redaction +6
  • Per-key rate limits +4

Compliance, when asked for up to 10 pts

  • Each compliance requirement you selected and the vendor confirms +10
  • SOC 2, as general procurement reassurance +5
  • Publishes a status page +3
  • OpenAI-compatible, so leaving later is cheap +4

Scores are then normalised so the leader shows 100% and everything else is expressed relative to it. That percentage is a relative fit against the products in this catalog, not an absolute quality grade — an 80% in a weak field is not better than an 80% in a strong one.

Back to top ↑

How the lock-in score is derived

The “how hard is it to leave” score on each product page is arithmetic, not an opinion. Six published facts each carry a fixed number of points, identical for every product, and each point on the page links to the source for the fact it came from. The weights are ordered by how much engineering time the fact saves you on the way out:

  • Works with standard OpenAI code Leaving is a base-URL change, not a rewrite of every call site. +22
  • You can bring your own provider keys Your model access and billing survive dropping the gateway. +20
  • You can self-host it A price or policy change cannot strand you if you can run it yourself. +20
  • Configuration lives in version control Routing rules are a file you keep, not dashboard state to rebuild by hand. +16
  • Your request history can be exported You leave with your own logs instead of abandoning them. +12
  • No proprietary SDK required Weighted low because no product we track requires one — it earns little on its own. +10

An unpublished fact lowers the ceiling instead of costing points. If a vendor documents four of the six, the highest score it can reach is the sum of those four, and the page shows that reduced denominator rather than pretending the gaps were failures. This does mean a poorly documented product cannot reach the top band — which we think is the right outcome, because you genuinely cannot verify your exit path from documentation that does not exist.

Band cut-points sit at 85 and 55 out of 100. Those are set against the observed spread of the products we track, not an assumed even distribution. An earlier version of this score put more than half its weight on OpenAI compatibility and the absence of a proprietary SDK; because every product in the catalog satisfies both, 18 of 20 came out in a single band and the score separated nothing. The current weights put the emphasis on what actually varies — your keys, your config, your history, and whether you can run it yourself.

The score covers technical switching cost only. It does not price the work of re-testing prompts against a different routing stack, renegotiating a contract, or retraining a team, and a high score is not a recommendation to leave.

Back to top ↑

What this cannot tell you

  • Weights are a judgement call. We think automatic failoverFailoverWhen the provider you asked for is down or rate-limiting you, the gateway automatically retries somewhere else. The single most valuable reliability feature these products offer. deserves 14 points when you say reliability matters. You might think it deserves 30. The numbers are published precisely so you can disagree with them specifically.
  • Vendor benchmarks are marketing. Latency and throughput figures are almost always self-published under conditions chosen to flatter. Where an independent measurement or a contradicting production report exists, it is noted alongside. Never size capacity from these.
  • Model counts move weekly and are not comparable. One vendor counts every variant and quantisation, another counts families. A gap of 30 means nothing; a gap of 300 means something.
  • Self-hosting is priced misleadingly everywhere, including here. The cost tool cannot price the engineering time to run, patch, and upgrade a gateway, and that is usually the dominant cost.
  • Enterprise pricing is negotiated. Above a certain volume the published numbers simply stop applying.
  • Nothing here reflects support quality or roadmap risk. Whether a vendor answers tickets, or is still around in two years, is not a field we can source.
Back to top ↑

How it stays current

Values carry the date they were checked, and a field turns amber once it passes 45 days without verification. The changelog records every edit with its previous value and the source that justified the change, and lists each product by its oldest field check — so the page tells you where it is decaying rather than hiding it.

Refreshes are triggered deliberately rather than on a schedule, and proposed changes are reviewed before they reach these pages. That is slower than automating it, and the trade is intentional: an unreviewed pipeline writing wrong numbers into a comparison site is worse than a site that is a fortnight behind.

Back to top ↑

How we vet independent coverage

Every product page links to third-party coverage. The hard part is not finding links — it is that most of what a search returns for “product X review” is marketing wearing a review's clothes. So each link has to clear a bar before it goes on a page.

What we exclude, and why

  • Anything the vendor published or wrote. Including bylines: we dropped an InfoWorld piece on Kong's AI Gateway on finding it was authored by Kong's own CTO, and a conference talk on APISIX given by an API7 developer advocate. A vendor's own account of its product belongs in the official links section, which is a click away, not under “independent”.
  • Anything a competitor listed on this site published. A comparison table written by one gateway vendor about another is a sales document.
  • SEO and AI-generated content farms. Sites that publish a templated “review” of every tool in a category, with no author and no evidence of anyone having run the thing.
  • Launch and funding announcements. A round being raised tells you nothing about whether the product works.
  • Vendor-submitted forum posts. A “Show HN” thread where the discussion is the founders answering questions is not practitioner feedback.

What has to be true

  • The page was fetched and read, not taken from a search snippet. Several promising-looking URLs turned out to 404 or to be robots-blocked, and those were dropped rather than cited unseen.
  • Video has to come from a channel with an existing audience and subject-matter track record, not the vendor's own channel. Every video link was checked against YouTube directly to confirm it still exists — one candidate was a dead ID and was removed.
  • A discussion thread needs at least five comments. Below that it is a post, not a discussion, and six otherwise-relevant threads were cut on this rule alone.

The consequence is that coverage is uneven, and we would rather show that than hide it. Well-known platforms have plenty. Three products — Helicone, LLM Gateway and TrueFoundry — have exactly one link each, because for those the honest finding was that almost everything written about them is written by them. A thin coverage section is itself a signal worth reading.

Switch a product page to Technical mode to see, next to each link, the reason it qualified.

Back to top ↑

Why there is no “available in the US” badge

Every product in this catalog is a public internet service, so all twenty are reachable from the United States. A badge saying so would be true twenty times over and would tell you nothing. Instead the cards and product pages report the two things that genuinely differ.

  • Operator jurisdiction — the country the company is headquartered in, which is whose courts and disclosure regime it answers to. Three of the twenty sit outside the US, and those are marked so you can see it while scanning rather than after opening the page. This is sourced per product like every other value.
  • Data residency — whether processing can be pinned to a specific region. Reported only where the vendor has published a position; where they have said nothing we say nothing, rather than recording silence as a “no”.

These are separate facts and conflating them is a common way to get this decision wrong: a US-headquartered vendor can still process your data in Europe, and a European vendor can offer a US region. The providers page filters them separately for that reason.

Back to top ↑

Found something wrong

Corrections are genuinely welcome, particularly from vendors about their own products. Include the URL that shows the correct value and it can be verified and changed quickly — with the edit recorded on the changelog like any other.

Back to top ↑

The full field list, with a plain-language definition of each, is on the glossary page — 8 groups covering all 113 tracked fields.