Guide
How much does an LLM gateway lock you in?
Short answer
An LLM gateway can reduce provider-specific integration work without making a migration automatic. Even chat-only applications must verify authentication, model IDs, errors, streaming and tool-call behavior. Routing policies, virtual keys, stored data and prepaid balances add further exit work. Catalogue endpoint support narrows a shortlist; it does not prove interchangeability.
Compatible APIs still need a contract test
The catalogue records chat-completions support for 31 of 31 products. That is a documented surface, not a test that the same request behaves identically everywhere. Migration effort depends on the features and guarantees your application actually uses.
| Contract | Compare and record |
|---|---|
| Authentication and model IDs | Key scopes, organization headers, aliases, version pins and authorization errors. |
| Streaming | Chunk shape, finish reason, usage delivery, cancellation and partial failure. |
| Tools and structured output | Argument schema, tool IDs, parallel calls, JSON validation and refusals. |
| Errors and limits | Status codes, retry headers, rate-limit scope and deadline behavior. |
| Usage and routing | Input/output/cache token accounting, actual provider/model and fallback headers. |
Use synthetic fixtures that reflect your application, compare outcomes, then rehearse key rotation, configuration import and rollback. A base-URL change may be the code edit; it is not the complete migration acceptance test.
Every surface you add narrows the field
Beyond chat the catalogue records 6 further endpoints, and support thins out across them. Each card below counts the products that document the surface, then the ones that document it with caveats, then the ones that publish nothing. Silence is recorded as silence: several of these products almost certainly carry an endpoint they never wrote down, and an endpoint you cannot cite is still one you cannot plan a migration around.
- 22 Anthropic messages
- Anthropic-shaped calls accepted without rewriting them. The second most common surface, and often provided as SDK compatibility or provider passthrough rather than a native endpoint, which the product pages distinguish. 9 publish nothing.
- 24 Embeddings
- Vectors through the same gateway as text. Cheap to adopt and easy to forget you adopted, because a retrieval pipeline points at the gateway once and never mentions it again. 7 publish nothing.
- 17 Responses
- The newer stateful OpenAI surface. Support is much thinner than for chat completions, so code written against it has a smaller field of replacements from the day it is written. 3 document it partly; 11 publish nothing.
- 17 Images
- Image models reachable on the same key and the same base URL. A separate integration if your gateway does not carry them, which is a cost you pay on the way in as well as on the way out. 3 document it partly; 11 publish nothing.
- 16 Audio
- Speech to text and text to speech. Frequently the first gap in an otherwise complete gateway, and the surface most often documented as partial rather than as a plain yes. 5 document it partly; 10 publish nothing.
- 9 Batch jobs
- Asynchronous bulk processing, usually at a discount. The least documented surface in the catalogue, and the one that most often stalls a migration late, once everything else has already moved. 22 publish nothing.
Read down that list and it is a set of independent facts. Read it as a stack — as an application that adopts one surface, then another — and it becomes the shape of your exit. Each step below keeps only the products that document everything above it, so the number is the field of replacements available to an application that has gone that far.
- 31 Chat completions only
- 22 and Anthropic messages
- 18 and Embeddings
- 12 and Responses
- 7 and Images
- 7 and Audio
- 3 and Batch jobs
3 of the 31 document the whole set: Kong AI Gateway, LiteLLM and OpenRouter. An application that calls all 6 surfaces has that many like-for-like replacements, and every additional requirement — a licence, a region, a price — is applied to that shortlist rather than to the full catalogue.
No product here stops at chat: every one documents at least one surface beyond it, so the narrowing is a property of your adoption rather than of the market. The vendors describe their own surfaces in their own words, and the catalogue keeps those lists verbatim: 79 distinct surface names appear across the 31 products, with individual products naming between 1 and 13. Those lists are worth reading before you commit, because the name a vendor gives a surface is often the first clue that it is its own rather than somebody else’s.
One thing genuinely reduces this: being able to point the gateway at an endpoint you run yourself. 15 of the 31 document a way to register a custom or self-hosted endpoint outright, 15 document something partial or nothing at all, and 1 states plainly that there is no mechanism. A product that will route to your own model server can follow you into places its catalogue does not go, which is the difference between a gateway you are using and one you are inside.
Surface and configuration, product by product
All 31 products, sorted by name. Three of the surfaces where products disagree most, plus where the configuration lives. Read across rather than down: a product that documents every column is one you could leave for another that does the same, and a row with three blanks is not a bad product but an unknown exit.
| Product | Anthropic messages | Responses API | Batch | Config as code |
|---|---|---|---|---|
| agentgateway | Yes Check date not recorded | Yes Check date not recorded | Not documented Checked 2026-09-02 | Yes Check date not recorded |
| AI Gateway HQ | Yes Checked 2026-09-17 | Yes Checked 2026-09-17 | Not documented Check date not recorded | Not published Check date not recorded |
| Amazon Bedrock | Not documented Check date not recorded | Yes Check date not recorded | Yes Check date not recorded | Yes Checked 2026-08-29 |
| Apache APISIX AI Gateway | Yes Check date not recorded | Partly Check date not recorded | Not documented Check date not recorded | Yes Checked 2026-08-29 |
| Azure AI Foundry | Not documented Check date not recorded | Not documented Check date not recorded | Yes Check date not recorded | Yes Checked 2026-08-29 |
| Bifrost | Yes Check date not recorded | Not documented Check date not recorded | Not documented Check date not recorded | Yes Checked 2026-08-29 |
| Braintrust Gateway | Yes Check date not recorded | Yes Check date not recorded | Not documented Check date not recorded | Not published Check date not recorded |
| Cloudflare AI Gateway | Yes Check date not recorded | Not documented Check date not recorded | Not documented Check date not recorded | Yes Checked 2026-08-29 |
| Eden AI | Yes Check date not recorded | Yes Check date not recorded | Not documented Check date not recorded | No Check date not recorded |
| Envoy AI Gateway | Yes Check date not recorded | Yes Check date not recorded | Not documented Check date not recorded | Yes Check date not recorded |
| Fireworks AI | Not documented Check date not recorded | Not documented Check date not recorded | Yes Check date not recorded | Not published Check date not recorded |
| Google Vertex AI | Not documented Check date not recorded | Not documented Check date not recorded | Not documented Check date not recorded | Yes Checked 2026-08-29 |
| Groq | Not documented Check date not recorded | Yes Check date not recorded | Yes Check date not recorded | No Checked 2026-08-29 |
| Helicone | Yes Check date not recorded | Not documented Check date not recorded | Not documented Check date not recorded | Yes Checked 2026-08-29 |
| Higress | Yes Checked 2026-09-02 | Not documented Checked 2026-09-02 | Not documented Checked 2026-09-02 | Yes Check date not recorded |
| Hugging Face Inference Providers | Not documented Check date not recorded | Yes Checked 2026-09-03 | Not documented Check date not recorded | No Check date not recorded |
| Kong AI Gateway | Yes Check date not recorded | Yes Check date not recorded | Yes Check date not recorded | Yes Checked 2026-08-29 |
| LiteLLM | Yes Check date not recorded | Yes Check date not recorded | Yes Check date not recorded | Yes Checked 2026-08-29 |
| LLM Gateway | Yes Check date not recorded | Not documented Check date not recorded | Not documented Check date not recorded | No Checked 2026-08-29 |
| Merge Gateway | Yes Checked 2026-09-02 | Partly Checked 2026-09-02 | Not documented Checked 2026-09-02 | Yes Checked 2026-09-02 |
| MLflow AI Gateway | Yes Check date not recorded | Yes Check date not recorded | Not documented Check date not recorded | No Checked 2026-09-02 |
| New API | Yes Checked 2026-09-02 | Yes Checked 2026-09-02 | Not documented Check date not recorded | Not published Check date not recorded |
| OpenRouter | Yes Check date not recorded | Yes Check date not recorded | Yes Check date not recorded | No Checked 2026-08-29 |
| Orq.ai Router | Not documented Check date not recorded | Yes Check date not recorded | Not documented Check date not recorded | No Checked 2026-08-29 |
| Portkey | Yes Check date not recorded | Not documented Check date not recorded | Not documented Check date not recorded | Yes Checked 2026-08-29 |
| Requesty | Yes Check date not recorded | Yes Check date not recorded | Not documented Check date not recorded | No Checked 2026-08-29 |
| Respan | Yes Checked 2026-09-15 | Partly Checked 2026-09-15 | Not documented Check date not recorded | Not published Checked 2026-09-15 |
| Together AI | Not documented Check date not recorded | Not documented Check date not recorded | Yes Check date not recorded | No Checked 2026-08-29 |
| TrueFoundry AI Gateway | Yes Check date not recorded | Not documented Check date not recorded | Yes Check date not recorded | Yes Checked 2026-08-29 |
| Velokey | Not documented Check date not recorded | Yes Checked 2026-09-19 | Not documented Check date not recorded | Not published Check date not recorded |
| Vercel AI Gateway | Yes Check date not recorded | Yes Check date not recorded | Not documented Check date not recorded | Not published Check date not recorded |
A blank reads as not published rather than as no. An undocumented batch endpoint may well exist and simply never have been written down, which is a finding about the documentation rather than a verdict on the product. For a migration the effect is still real, because you cannot plan against a surface nobody has committed to in public. The chat completions column is omitted on purpose: it is the same answer for every row.
Where your configuration lives
The second axis of lock-in has nothing to do with endpoints. It is the question of whether your production configuration — the routing rules, the fallback order, the budgets, the key scopes — exists anywhere other than the vendor’s database. 15 of the 31 products document configuration as code. 9 are configured by clicking, and 7 publish nothing either way.
The reliability fields say the same thing from another direction. 6 products keep the fallback chain, the load-balancing policy, or both, in a dashboard only. For 3 of those there is no documented configuration-as-code path either, so the routing behaviour of your production traffic is a set of clicks that nobody reviewed and nobody can diff. If you are in that group, the cheap mitigation is to write the settings down in your own repository, in whatever format, and treat the dashboard as a rendering of that file rather than as the file.
Virtual keysVirtual keys: Separate scoped keys you issue per team, app, or customer, each with its own budget and limits, without handing out your real provider credentials. deserve their own line, because they look like a governance feature and behave like a commitment. 11 of the 31 document them and 20 publish nothing. A virtual key is a credential your applications hold that only this product can honour, so every service holding one is a service that needs a new credential on the day you leave. 7 of the products that issue them do not also document configuration as code, which means the mapping between keys, teams and budgets is itself dashboard state. Keep your own inventory of which service holds which key and what it is allowed to spend; that inventory is the migration plan.
Money you cannot take with you
The third axis has a currency figure attached. Where you bring your own provider accounts — BYOKBYOK — bring your own key: You keep your own accounts and contracts with OpenAI, Anthropic and the rest, and the gateway routes through your keys. You keep your negotiated rates and any committed-spend discounts; the gateway charges you for the plumbing, not the tokens. — the contract with the model provider is yours, and leaving the gateway leaves the spending relationship intact. Where you prepay the gateway for inference, the balance is a switching cost with a number on it.
- Your keys only — 12
- Your provider keys, your provider bills. The gateway charges for itself, if at all, and there is no balance to strand.
- Your keys or their credits — 12
- Either arrangement is available, which means the exposure is a choice you make at setup rather than one the product makes for you. Choose your own keys if you expect to reassess the gateway later.
- Their credits only — 3
- Inference is bought from the vendor and there is no documented way to route around it with your own accounts. Any unspent balance is money that leaves with the vendor, so top it up in amounts you would be willing to write off.
- Not applicable — 3
- The distinction does not apply, because the model and the platform are the same company. There is no separate key to bring, and the commitment is to the platform rather than to a routing layer in front of it.
24 of the 31 let you hold the provider relationship yourself, and 7 publish nothing about whether your own keys are accepted at all. This page stops here on purpose: what each product charges on top, and which charging model wins at your volume, is the subject of the pricing guide, and the cost estimator puts numbers against it. The only lock-in point here is the direction of the money: a prepaid balance with one of the 3 credits-only products is the one part of this analysis you can put in a spreadsheet on the day you sign.
Which side you are on
Low exit cost if
- Chat completions are the only surface you call.
- Your routing and budget rules are in a file you commit.
- You pay your model providers directly and hold no balance.
- Your services authenticate with keys you issued, not the gateway.
You are committing if
- Embeddings, images, audio or batch jobs run through the same product.
- Production routing exists only as settings in a console.
- You have prepaid for inference you have not used yet.
- Applications across your estate hold keys the gateway minted.
- You depend on a surface only a handful of products document.
Reduce the cost now by
- Keeping the chat surface as your interface and wrapping the rest.
- Moving configuration into version control wherever the product allows it.
- Buying credits in amounts you would accept losing.
- Treating every issued key as gateway-scoped and rotatable.
- Writing down, once a quarter, which surfaces you now depend on.
Common questions
If every product speaks the OpenAI chat format, is there any lock-in at all?
Yes, including on the chat path. A base URL and key change may connect a client, but authentication, model IDs, streaming, errors and tools still need validation. The cost sits in what you added afterwards: the embeddings calls, the image and audio endpoints, the batch jobs, the routing and budget rules configured in a dashboard, and the scoped keys your services already hold. None of that moves with a base URL.
How do I tell how committed I already am?
Count the distinct endpoints your code calls through the gateway, and count the settings that exist only in its console. The first number tells you how many products could replace it; the second tells you how much of your production configuration you would be transcribing by hand at three in the morning. Both are countable in an afternoon and neither gets smaller on its own.
Does a product being open source mean I am not locked in?
It removes one kind of dependence and not the others. You can keep running an open build after a vendor loses interest, which is a real protection. It does not make your image or batch integration portable to a different product, and it does not refund a prepaid balance. Licence and exit cost are separate questions, and the self-hosting guide covers the licence one.
Is a vendor silent on an endpoint the same as a vendor without it?
No, and this page records the two differently. A blank means nothing was published on the pages we read, and the endpoint may exist. The practical effect is the same for planning, though: an undocumented endpoint is one you cannot rely on in a migration, so treat it as a question for the vendor rather than as a verdict.
What is the cheapest thing I can do now to keep the exit open?
Keep the chat surface as your interface and put anything else behind your own thin wrapper, so provider differences are contained and tested. Keep routing and budget configuration in version control where the product allows it. Prefer paying your own model providers over prepaying a balance. And treat any key the gateway issued as gateway-scoped, so leaving is a rotation you have already rehearsed.