Managed only Proprietary hosted service

AI Gateway HQ Managed gateway

AI Gateway HQ is a managed LLM gateway with an OpenAI-compatible API in front of 11 upstream providers; it publishes no model count. Its entry gateway meter is $0.10 per 1,000 successful requests; provider inference is separate. It cannot be self-hosted. It publishes neither an unconditional zero-retention guarantee, a HIPAA BAA nor SOC 2. You can point it at your own provider accounts. Beyond chat it also serves embeddings.

· 170 dated entries · 170 source references

Built by AI Gateway HQ LLC · no EU region

Access at a glance

Whether this product can work for you at all, before features matter: who pays the model bill, where it can run, and which of your existing API calls keep working. Every value is the vendor’s own claim, linked to the page it came from.

Hosted governance gateway; optional managed Bedrock access is a separate funding path.

Who pays the model bill

Your keys or their credits

You can start on their credits and move to your own provider accounts later.

BYOK or funded managed Bedrock; managed access requires purchased credit and a reusable payment method.

Merchant of record: Customer pays BYOK provider directly; managed Bedrock is prepaid through AI Gateway HQ.

Key handling: Write-only provider credentials encrypted using tenant-bound KMS context.

Where it can run

1 of 5 shapes documented
  • Vendor-hosted
  • Self-host
  • Your VPC
  • On-premise
  • Air-gapped

Dimmed shapes are not documented by the vendor, which is not the same as unsupported.

Private-deployment design is separately scoped, not an available self-host license.

Customer-VPC deployment is not generally available.

API surfaces your code can keep using

4 of 7 documented
  • OpenAI chat POST /v1/chat/completions Yes *

    POST /v1/chat/completions; compatible target required.

  • Anthropic messages POST /v1/messages Yes *

    POST /v1/messages with compatible provider targets.

  • OpenAI Responses POST /v1/responses Yes *

    POST /v1/responses; route capabilities still constrain requests.

  • Embeddings POST /v1/embeddings Yes *

    POST /v1/embeddings; configure an eligible target.

  • Images POST /v1/images/generations Not documented
  • Audio POST /v1/audio/* Not documented
  • Batch jobs POST /v1/batches Not documented

An asterisk marks a qualified verdict: support that is indirect (SDK compatibility or provider passthrough rather than a native endpoint), or a gap that is narrower or wider than the label suggests. Read the note before porting.

Documented interfaces do not imply support for every upstream API feature.

How much it reaches

Models Not published
Upstream providers 11

Providers: Counted from the vendor’s own published list; no aggregate total is published.

Anonymous models endpoint requires authentication; do not infer model availability from provider integrations.

Enumerated ten named model services plus OpenRouter; not a model count. Most BYOK integrations are labelled Beta.

Whose models: Routes to third-party model services; it does not supply its own foundation models.

Your own endpoints: Named compatible integrations; arbitrary custom-host support not established.

How it behaves in production

What happens when an upstream model is slow, wrong, or down — and what you can see and stop while it happens. Reliability features are recorded as where you configure them, not whether the vendor lists them, because almost every product here lists all of them.

0 of 6 reachable from code 1 of 4 can block 1 documented destination

When something goes wrong

Each row says where the knob is, not whether the feature is on the marketing page. A control you can only reach by hand in someone else’s dashboard cannot be reviewed or version-controlled.

  • Request timeout Not documented

    The vendor does not document this, so any behaviour you observe today is unversioned and may change.

    Route attempt/time boundaries configured by administrators. Console configuration is documented; whether this is dashboard-only or also configurable through code is not established.

  • Retries Not documented

    Bounded attempts; exact numeric defaults not published. Console configuration is documented; whether this is dashboard-only or also configurable through code is not established.

  • Fallback to another model Not documented

    Fallback constrained by capabilities, region, time and cost. Console configuration is documented; whether this is dashboard-only or also configurable through code is not established.

  • Load balancing Not documented

    Priority and weighted credential pools. Console configuration is documented; whether this is dashboard-only or also configurable through code is not established.

  • Upstream health tracking Not documented

    Health-aware selection and quota cooldowns. Console configuration is documented; whether this is dashboard-only or also configurable through code is not established.

  • Cross-region failover Not documented

Defaults: Hedging is off by default; this does not establish retry count.

Validate configured target capabilities; fallback is not a guarantee of upstream parity.

How fast the hop is

Undisclosed vendor service

The vendor does not disclose what the request path runs on, so no overhead floor can be inferred at all.

Hosted AWS production; no independently measured overhead recorded.

You can run the request path yourself No
Streaming Partly

No generally available customer deployment; paid design work is distinct.

Streaming caveats: Compatible BYOK streams supported; exact cache and enforcing output checks exclude streaming.

This vendor publishes no latency or throughput figure for the routing layer. That is the most common case here, and it is why the architecture class above carries the comparison instead of a number.

What it will stop

1 of 4 can block

Two separate questions per control: can it stop a request at all, and what does it do before you change any settings? A control that inspects and forwards is a logging feature, however it is named.

  • Personal data in prompts Not documented

    Non-streaming output PII/secret checks are documented; a prompt-input PII blocking guarantee is not established.

  • Prompt injection and jailbreaks Not documented
  • Harmful content Not documented
  • Your own policies Can block the request

    Out of the box: Not documented

    Bounded JSON Schema and configured tool profiles before release; enforcing inspection requires stream=false.

Where checks run On the vendor's servers
If the guardrail itself fails Not documented

Almost no vendor in this catalogue documents what happens when the guardrail service itself times out. If the control matters to you, this is a question worth asking before you sign.

Identity, tenant and explicit policy denial fail closed; a general guardrail timeout guarantee is not established.

Organization/key controls do not establish delegated hierarchy or non-overridable parent policies.

What you can see

Export is limited
What gets logged Metadata only

Token counts, latency and model names are stored, but not the text itself.

You can turn bodies off Not documented
Traces Not documented

Identity, route, timing and usage metadata; caching/support may differ.

Where telemetry can go

  • Signed HTTPS audit webhooks

Administrative metadata only; not a complete prompt or trace export.

Records user feedback Not documented
Scores live traffic Not documented

Depends on the vendor’s SaaS: Current hosted console stores operating metadata.

Retention: Usage, security, audit and billing records have different retention purposes.

Whether it fits how you work

How much work stands between you and a first call, how different that is from running it in production, and whether it slots into the stack you already have. Recorded as the shape of the work rather than a number of minutes — how long it takes you depends on which accounts and quota you already hold, which no comparison can know.

to try: cloud console setup to run: cloud console setup fits 1 of 10 common stacks

Getting to a first call

5 numbered steps
Shape of the work Set it up in a cloud console

You cannot start from an empty editor. An account, a project or a deployed resource has to exist first, and that step is done by hand.

Read off: the vendor’s own quickstart — 5 numbered steps.

Vendor lists five setup steps; no elapsed-time test performed.

Before step one

  • Your own provider key Optional

    You can start on the product’s own credits and move to your own provider keys later.

    BYOK for customer accounts; managed Bedrock uses purchased credit.

  • Payment method No card needed to start

    Test Lab requires no card; managed Bedrock requires a reusable payment method.

  • Gate before models answer You enable it first

    One extra click or API enable per model or project before a call succeeds.

    Configure eligible targets; managed models are reviewed for account availability.

Everything you need first: Workspace, provider credential or managed entitlement, route alias and workload key.

Running it in production

The same scale applied to the path the vendor recommends for production traffic. Kept separate from the quickstart because for several products here the two are barely related pieces of work.

Shape of the work Set it up in a cloud console

You cannot start from an empty editor. An account, a project or a deployed resource has to exist first, and that step is done by hand.

What production needs: Validate route capabilities, rates, output caps and enforcing policies before production.

Can you run it yourself

Install command published No self-hosting

This runs on the vendor’s infrastructure only.

How it fits your stack

1 of 10

Each row is a thing you might already run. “With a caveat” means it works but not the way the vendor’s marketing implies — read the reason, because that is usually where the surprise lives.

  • With a caveat The OpenAI SDK OpenAI-compatible paths are documented, but not as a general drop-in.
  • No The Vercel AI SDK No AI SDK route documented.
  • No Cloudflare Workers No Workers guidance published.
  • No Kubernetes No Kubernetes deployment published.
  • No Terraform or OpenTofu Nothing published for Terraform.
  • Fits An existing API gateway This is that gateway — AI traffic becomes a plugin, not a new hop.
  • No Cloud IAM I already run Static upstream credentials only. Your calls to it still use its own key.
  • No LangChain or LlamaIndex No framework integration documented.
  • With a caveat MCP servers to govern MCP tools in the API — governs nothing on your side.
  • No Nothing — plain Node or Python A cloud console or resource has to exist before your first call.

Reading this the other way round — pick what you already run and see every product scored against it.

The integration surfaces behind those answers

  • Vercel AI SDK Not documented

    Nothing published. Assume the OpenAI-compatible route and verify it yourself.

  • Cloudflare Workers Not documented

    No Workers guidance either way. If you are edge-first, verify fetch-only compatibility yourself.

  • Kubernetes Not documented

    No Kubernetes story published.

  • Terraform Not documented

    No Terraform surface published. Configuration is API or dashboard work.

    Internal AWS module exists; customer-facing resource provider is planned, not available.

  • Existing API gateway It is the API gateway

    This product is the gateway. If you already run it for your other APIs, AI traffic becomes a plugin rather than a new hop.

    Applications call this gateway before configured upstream services.

  • Cloud identity Static provider credentials only

    You paste static provider credentials into this product so it can reach upstream models. Your own calls to it still use its own API key, and those pasted secrets are yours to rotate.

    Workload gateway key for requests; OIDC/SAML is administrative identity.

  • MCP MCP tools in the API

    The completion API accepts MCP tool definitions, so the model can call MCP tools. A model capability, not an MCP control plane.

    Checks MCP declarations and emitted calls; not evidence of an MCP-server proxy.

Python frameworks None documented
First-party client libraries
  • Python

Examples use upstream OpenAI and Anthropic SDKs.

Agent features: Client configuration guidance for Codex CLI, Claude Code and OpenCode.

Start with test credentials and inspect decisions before enforcing rules.

Most BYOK provider connections are Beta; managed Bedrock is labelled Available.

What it does well

  • Pre-request reservations with scoped workload keys
  • Compatible OpenAI and Anthropic interfaces
  • Metadata-only default request evidence

Where it falls short

  • No generally available private deployment or SLA
  • No current SOC 2 certification claim
  • Most documented BYOK integrations remain Beta

Choose it when

Teams evaluating hosted budget and policy controls with metadata-based reporting.

Look elsewhere when

You need a current self-host package, audited assurance, or a public exhaustive model inventory.

Managed gateway: A hosted control plane in front of the models: routing, caching, guardrails, spend limits, and logs. You trade a little simplicity for governance.

How hard is it to leave?

Derived from six published facts, not from an opinion. The weights are fixed and the same for every product — see the arithmetic.

64 /84 Some work to leave
Portability score breakdown for AI Gateway HQ
What helps you leave Points Source
Works with standard OpenAI code Switching away is a base-URL change rather than a rewrite of every call site. 22 /22 vendor page
No vendor-specific SDK required A proprietary client library spreads through your codebase and has to be torn out again. 10 /10 vendor page
Can use your own provider accounts Your keys and billing relationship stay yours, so removing the gateway does not cut off model access. 20 /20 vendor page
Can be self-hosted You can run it yourself instead of accepting a pricing or policy change. 0 /20 vendor page
Configuration lives in version control Routing and budget rules are a file you keep, not dashboard state you would have to rebuild. not published —
Your request history can be exported You leave with your own logs instead of abandoning them. 12 /12 vendor page

Read the fine print: Compatible clients reduce integration work; policies, aliases and metadata require a separate migration plan.

1 of the 6 inputs is not published, so the highest reachable score here is 84 rather than 100. That is a gap in the public documentation, not a mark against the product — no points are deducted, they simply cannot be claimed. This measures technical switching cost only. It does not price the engineering time to re-test prompts against a different routing stack.

AI Gateway HQ models & pricing

Browse every imported listing from this provider, with published token rates and a link to compare other providers for the same model. This is provider-reported coverage; an absent listing does not mean unsupported.

Loading model listings…

Official model coverage source ↗ · Model source coverage and limitations

Full specification

Every field we track. Blank fields say "Not published" rather than "No" — we do not infer an absence from silence. Switch to Technical in the header for the precise field names and the low-level details.

Overview

What kind of product Category
Managed gateway our judgement
Marketplaces resell many providers behind one key. Gateways add governance on top. Open-source projects you run yourself. Cloud platforms are hyperscaler surfaces. Inference providers host models on their own hardware.
Who runs it Deployment model
Managed only
Managed means the vendor operates it. Self-host means you run it on your own infrastructure. Both means you can choose.
Licence Licence
Proprietary hosted service
Proprietary products cannot be inspected or forked. Open licences such as MIT and Apache-2.0 let you audit, modify, and run the code without permission.
Company Company
AI Gateway HQ LLC
The organisation that maintains the product.
Who you would be signing with Vendor status
Not published
Whether the product is still an independent company, has been acquired, is a large cloud vendor’s product line, is run by a software foundation, or has been put into maintenance mode. Maintenance mode means bug fixes and security patches only — no new features.
Last shipped an update Latest release
Not published
The date of the most recent release or version tag. A product that has not shipped in a year is a different risk from one that shipped last week, regardless of what its marketing site says.
GitHub stars GitHub stars
Not published
A rough proxy for community size on open-source projects. Not a quality measure.

Cost

Markup on model prices Token markup
None
How much the product adds on top of what the underlying model provider charges. Zero means you pay the same per-token price you would pay the model provider directly.
Fee to add funds Credit purchase fee
Not published
A percentage charged when you top up your balance, separate from token prices. It is easy to miss because it does not appear on the per-token price list.
Monthly cost per person Seat fee
Not published
A recurring per-user platform charge that applies regardless of how much you use the models.
Can use your own provider accounts BYOK supported
Yes
Bring Your Own Key: you keep direct contracts with OpenAI, Anthropic and others, and the gateway only routes traffic. This preserves negotiated rates and committed-spend discounts.
Cost of using your own accounts BYOK terms
0% applies to BYOK inference. Flex gateway usage is $0.10/1,000 successful requests; managed Bedrock rates are separate.
What the product charges to route traffic through your own provider keys.
Free tier Free tier
No-card simulated Test Lab; live inference needs provider funding and gateway entitlement.
What you can do without paying, useful for evaluation.
Enterprise plan from Enterprise plan from
$36,000/yr
Annual entry price for the enterprise tier, where one is published or credibly reported.
Cost to run it yourself Self-host cost
Design engagement from $60,000/year; not a ready-to-run deployment license.
What self-hosting actually costs once you account for infrastructure and any paid tier.
How the vendor makes money Pricing model
Platform fee plus usage meters
The shape of the vendor’s bill: does the routing layer charge a percentage on top of tokens, a flat monthly fee, both, neither (because inference is the product), or nothing at all (open source with no paid tier).
How pricing works, briefly Pricing model detail
Flex usage meter or Company subscription; provider inference is separate. Portfolio has a separate sponsor boundary.
A one-paragraph description that covers the caveats a pricing category cannot: introductory rates, per-feature meters, tier gating, and pricing that resets on a specific date.
Minimum commitment Minimum commitment
Flex has no subscription. Hosted enterprise starts at $36,000/year.
Whether the vendor requires a minimum contract term, a minimum spend, or a provisioned-capacity purchase to get its published rate.
Charges that fire after you go over an allowance Overage terms
Company and Portfolio allowances are hard monthly ceilings, not automatic overages.
The line items that scale with usage after an included allowance is exhausted — log storage, extra requests, per-feature meters, data export — which is where cost estimates usually go wrong.
Prompt cache offered Cache mechanism
Exact-match cache
Whether the gateway offers its own response cache, what kind of match it does (exact request, prefix, semantic), or simply passes provider caching through unchanged.
Discount on cached input Cache-read discount
Not published
How much cheaper cached tokens are than fresh input, when the vendor publishes a single figure. Bundled-inference clouds usually price this per model instead of as one number.
Premium on cache writes Cache-write premium
Not published
How much more the first write of a cached prefix costs versus a plain input token. A high write premium and a low hit rate can leave you paying more than you save, so this matters as much as the read discount.
Who captures the cache saving Cache economics
Opt-in volatile exact cache; gateway billing still runs. Streaming, tools and idempotent requests bypass it.
Whether the customer keeps the full saving from caching or the vendor captures part of it — and any conditions attached (write premium, storage fees, best-effort hits).
What you can split spend by Cost attribution
Workload, environment, client, data class, provider account, route and model in retained evidence.
The dimensions the vendor documents for splitting spend — per key, per user, per team, per tag, per customer. Matters if you need to chargeback internally or bill an end customer.
How you get cost data out Cost export
Durable period reporting and invoice-import automation are not currently included.
The mechanisms the vendor publishes for exporting cost and usage data: CSV, an API, webhooks, S3, a data warehouse, or nothing at all. Any per-unit price is included.
Who pays the model bill BYOK mode
Your keys or their credits
Whether you bring your own provider accounts (BYOK), buy inference from this vendor, or can do either. This is the single biggest commercial difference between these products: it decides who holds the contract with the model provider and who carries the spend.

Spend governance in detail

Seven signals matter when a bill starts to hurt: who can spend, how much, on what, and who gets paged when it goes wrong. Everything below is drawn from the vendor’s own pricing and docs pages — how we read these.

  • Virtual or scoped keys Yes

    Keys that carry their own budget and rate-limit policy, so an intern experiment cannot spend against a production budget.

  • Budget caps per key Yes

    A dollar or token ceiling attached to an individual key. Where enforcement is soft, one over-limit request still completes before the block kicks in.

  • Budget caps per team or workspace Not published

    A ceiling applied at a higher scope than one key — a team, a workspace, a customer, or an entire environment.

    Organization and workload-key scopes; dedicated team hierarchy not established.

  • Rate limiting as a cost control Yes

    Configurable request-per-time-window caps. Platform-set rate limits do not count as spend controls; user-configurable ones do.

  • Model allowlists Not published

    A policy that constrains which models a key or team can call, keeping expensive frontier models out of the wrong hands.

  • Spend alerts Yes

    Alerts fired as spend approaches a threshold. Alerts that only fire after the meter has rolled over are marked as such.

  • Webhook notifications Not published

    Programmatic notifications on spend events, so budget breaches can page an on-call or open a ticket.

Enforcement: Enforced before each request

Source →

Catalog

Models available Models available
Not published
How many models you can call, shown as a range because several vendors publish different totals on different pages. Vendor-reported either way, so counts are not directly comparable — some count every provider variant of the same model separately.
Model providers reachable Upstream providers
11
How many distinct model providers or labs you can reach, shown as a range where the vendor’s own pages disagree. More providers usually means better redundancy when one has an outage.
Works with standard OpenAI code OpenAI-compatible API
Yes
If yes, you can usually switch to it by changing one base URL, and switch away just as easily. This is the main defence against lock-in.
OpenAI chat endpoint POST /v1/chat/completions
Yes
The endpoint almost every application ports first. "Not documented" means the vendor never states it, which is different from a documented no.
Anthropic messages endpoint POST /v1/messages
Yes
Whether Anthropic-shaped calls work without rewriting them. Several products support this only as SDK compatibility or provider passthrough rather than a native endpoint — the detail page says which.
OpenAI Responses endpoint POST /v1/responses
Yes
The newer stateful OpenAI surface. Support is much thinner across this market than chat completions.
Embeddings endpoint POST /v1/embeddings
Yes
Whether you can generate vectors through the same gateway, or need a second integration for retrieval workloads.
Image generation endpoint POST /v1/images/generations
Not documented
Whether image models are reachable through the same surface as text.
Audio endpoints POST /v1/audio/*
Not documented
Speech-to-text and text-to-speech. Frequently the first gap in an otherwise complete gateway.
Batch jobs endpoint POST /v1/batches
Not documented
Asynchronous bulk processing, usually at a discount. Commonly undocumented, and commonly the reason a migration stalls late.
Needs the vendor’s own code library Requires a vendor-specific SDK
No
A proprietary client library spreads through your codebase and has to be torn out again if you leave. “No” is the better answer here, and it means the standard OpenAI client works.
You can export your request history Logs / usage data export
Yes
Whether you can get your own request logs, traces, or usage records back out — through an API, a bulk export, or a download. Decides whether you leave with your history or abandon it.
Settings can live in version control Declarative config-as-code
Not published
Whether routing, fallback, and budget rules can be declared in a file you keep in Git, rather than existing only as settings clicked into a hosted dashboard.
Embeddings Embeddings
Yes
Text-to-vector models, needed for search and retrieval features.
Image generation Image generation
Not published
Whether image models are reachable through the same interface.
Speech and audio Speech and audio
Not published
Text-to-speech or transcription models through the same interface.
Video generation Video generation
Not published
Whether video models are reachable through the same interface.
Batch processing Batch processing
Not published
Submitting large jobs for cheaper, slower processing. Often 50% off for work that is not time-sensitive.

Routing & reliability

Uptime it promises in writing Contractual SLA uptime
Not published
The uptime percentage in a published, contractual service level agreement. A public status page is not an SLA — it reports what happened, it does not promise anything or pay you back when it breaks.
Automatic failover Automatic failover
Yes
When a model provider goes down or rate-limits you, traffic moves to a backup automatically instead of returning errors to your users.
Load balancing Load balancing
Yes
Spreads requests across several providers or keys to raise your effective rate limit.
Rule-based routing Conditional routing
Yes
Send different requests to different models based on rules — for example a cheap model for free users and a strong model for paying ones.
Response caching Response caching
Yes
Reuses the answer when the exact same request comes in again, which cuts both cost and latency.
Similar-question caching Semantic cache
Not published
Reuses an answer when a new question means roughly the same thing as an earlier one. Saves far more than exact-match caching but can return subtly wrong answers if tuned loosely.
Where you set the timeout Request timeout surface
Not documented

Route attempt/time boundaries configured by administrators. Console configuration is documented; whether this is dashboard-only or also configurable through code is not established.

Where a request timeout can be set: per request, in a config file, in the vendor dashboard, or nowhere. Nine of the twenty products documented here do not describe a request timeout at all, so the worst case of a hung upstream call is unknowable from the docs.
Where you set retries Retry policy surface
Not documented

Bounded attempts; exact numeric defaults not published. Console configuration is documented; whether this is dashboard-only or also configurable through code is not established.

Where retry count and backoff are configured. Worth knowing alongside billing: a retried streaming call can be charged more than once.
Where you set fallbacks Fallback surface
Not documented

Fallback constrained by capabilities, region, time and cost. Console configuration is documented; whether this is dashboard-only or also configurable through code is not established.

Where the fallback chain is defined. Almost every product claims fallback; the useful question is whether you can change it from code or only by hand in a dashboard.
Shape of the fallback chain Fallback shape
Shape not documented
An ordered list tries targets in sequence; a weighted split sends a percentage of traffic to each, which is what you need to trial a new model on 5% of requests. Weighted splits are much rarer than the marketing implies.
Upstream health tracking Health checks / circuit breaking
Not documented

Health-aware selection and quota cooldowns. Console configuration is documented; whether this is dashboard-only or also configurable through code is not established.

Whether the product notices a failing upstream and stops sending traffic to it, and whether you can tune the thresholds. This is what turns a provider outage into a blip rather than a sustained error rate, and only four of the twenty expose it.
Cross-region failover you control Multi-region failover surface
Not documented
Whether you can define what happens when a region degrades. A vendor running many regions is not the same as a vendor letting you configure failover between them; only three document a user-controlled mechanism.
Where you set load balancing Load balancing surface
Not documented

Priority and weighted credential pools. Console configuration is documented; whether this is dashboard-only or also configurable through code is not established.

Where traffic distribution across upstreams or keys is configured.

Operations

Usage dashboards and logs Observability
Yes
Built-in visibility into what was sent, what came back, what it cost, and how long it took.
Spending limits Budget controls
Yes
Hard caps that stop spend before it becomes a surprise invoice. The single most valuable control for a small team.
Rate limits Rate limits
Yes
Caps on request volume per key or per user, useful for protecting against abuse and runaway loops.
Separate keys per team or app Virtual keys
Yes
Issue scoped keys with their own budgets and permissions so you can attribute cost and revoke access without rotating everything.
Prompt versioning Prompt management
Not published
Store and version prompts outside your code so they can be changed without a deploy.
Quality testing Evals
Not published
Built-in tooling to score model output against test cases, so you can tell whether a model swap made things better or worse.
MCP support MCP support
Yes
Native support for the Model Context Protocol, the emerging standard for connecting models to external tools.
What gets logged Logged content
Metadata only

Identity, route, timing and usage metadata; caching/support may differ.

Whether full prompts and responses are stored, only metadata, or nothing. Full-body logging is the most useful debugging feature here and the one most likely to need a conversation with your compliance team.
You can turn logging off Body-logging opt-out
Not documented
Whether prompt and response bodies can be suppressed while still keeping usage metrics. Nineteen of the twenty document a way to do this; the mechanisms range from a per-request header to an organisation-wide setting.
Traces you can take elsewhere Distributed tracing
Not documented
Whether the product emits OpenTelemetry, a proprietary format, or nothing. OpenTelemetry means the traces land in the tooling you already run instead of only in the vendor’s dashboard.
Where telemetry can go Export destinations
Signed HTTPS audit webhooks

Administrative metadata only; not a complete prompt or trace export.

Documented sinks for logs and metrics. This is a good proxy for how replaceable the vendor’s own dashboard is: around twenty destinations means you never have to depend on it, while a CSV download means you do.
Can record user feedback Feedback capture API
Not documented
Whether there is an API to attach a rating or score to a logged request, which is what lets production traffic feed quality work later.
Scores live traffic Online eval hooks
Not documented
Whether automated scorers can run against real production requests, rather than only against a test set you assemble yourself.

Performance

Delay it adds Proxy overhead
Not published
Extra time the product itself adds to each request, on top of however long the model takes. Usually irrelevant next to multi-second model latency, but it matters for high-volume or streaming-sensitive workloads.
Requests per second ceiling Throughput
Not published
Published sustained request rate before the product becomes the bottleneck. Only relevant at genuinely high volume.
What the request path runs on Architecture class
Undisclosed vendor service

Hosted AWS production; no independently measured overhead recorded.

The comparable way to talk about latency here. An edge worker, a compiled Go or Rust binary, and a Python proxy have different overhead floors no matter which figures each vendor publishes. Products that never disclose their runtime are recorded as undisclosed rather than assumed.
You can run the request path yourself Self-hostable data plane
No

No generally available customer deployment; paid design work is distinct.

Whether the component that actually carries your prompts can run on your own infrastructure. Distinct from a vendor offering a self-hosted control plane while still proxying traffic through their network.
Streaming responses Streaming support
Partly

Compatible BYOK streams supported; exact cache and enforcing output checks exclude streaming.

Whether token-by-token streaming is documented. The caveats matter more than the yes: some products cannot cancel a stream without still being billed, and several timeout and fallback mechanisms stop applying once the first token has been sent.

Security & compliance

Does your prompt reach their servers Prompt transits vendor
Yes

Hosted gateway processes content before forwarding to the selected provider.

Whether the text you send passes through this company’s own infrastructure. If it does, every other promise on this page is a policy commitment rather than a physical impossibility. Self-hosted products can answer no outright.
What they keep if you change nothing Logging default
Metadata only, not content

Ordinary evidence excludes prompt/output bodies; optional features can change the path.

Defaults matter more than options. A product that stores full prompts and replies unless you find the right header will have stored them by the time you read the docs.
How long they keep it Default content retention (days)
Not published

Record- and configuration-dependent; no universal day count.

Default retention for request content, in days. Zero means nothing is kept. Read the note: several products keep nothing as a rule but make timed exceptions for abuse review or specific models.
Could they train on your prompts Training on customer data
No

Gateway processor role excludes a separate model-training purpose; upstream terms are independent.

Whether the vendor may use your prompts and outputs to train models. “Not published” means we could not find any position, which is not the same as a no — ask for it in writing.
Where it runs, and what you can pin Region and residency control
Hosted production US-E1; managed Bedrock uses reviewed U.S. regions.
Which regions are offered and whether you can force processing to stay in one. A global endpoint that silently picks a region is a different compliance story from an endpoint you pin yourself.
Where safety filters run Guardrail execution location
In the vendor’s cloud

Hosted policy boundary; execution of tools remains the client responsibility.

A filter that strips personal data only helps if it runs before the data leaves your boundary. If guardrails execute in the vendor’s cloud, the vendor has already received whatever you wanted redacted.
Who else touches the data Subprocessor list
https://aigatewayhq.com/legal/subprocessors/
The published list of third parties the vendor passes your data to. No list means you cannot know the full chain, which most data-protection agreements require you to.
SOC 2 audited SOC 2 audited
No
An independent audit of security controls. Enterprise buyers and their procurement teams routinely require it.
Will sign a HIPAA agreement HIPAA BAA
Not published
Required before you may send protected health information through the service. Without a signed BAA, healthcare data is off limits.
GDPR commitments GDPR commitments
Not published
Published data processing terms for handling personal data of people in the EU and UK.
Can keep data in the EU EU data residency
No
Requests can be processed inside the EU rather than routed to US infrastructure. Often the deciding constraint for European customers.
Does not retain your data Zero data retention
Depends how you deploy it
Prompts and responses are not stored after the request completes. Sometimes a paid add-on rather than the default.
Strips personal data PII redaction
Yes
Detects and removes identifiers such as names, emails, and card numbers before the request reaches the model provider.
Content guardrails Content guardrails
Yes
Policy checks on inputs and outputs — blocking unsafe content, enforcing formats, or catching prompt-injection attempts.
Runs fully disconnected Air-gapped deployment
Not published
Can be deployed in a network with no internet access, which some regulated and defence environments require.
Blocks personal data in prompts PII / DLP enforcement
Not documented

Non-streaming output PII/secret checks are documented; a prompt-input PII blocking guarantee is not established.

Whether personal data detection sits on the request path and can stop the call, merely inspects and forwards it, or is not documented. A control that only reports is a logging feature, not a policy control.
Blocks prompt injection Injection / jailbreak enforcement
Not documented
Whether injection and jailbreak detection can stop a request. Most products offering this call a partner classifier rather than shipping their own.
Blocks harmful content Toxicity / moderation enforcement
Not documented
Whether hate, violence, sexual and self-harm categories are checked inline and can stop a request, in either direction.
Your own policy rules Custom policy hooks
Can block the request

Bounded JSON Schema and configured tool profiles before release; enforcing inspection requires stream=false.

Whether you can add your own rule — a regex, a webhook, or your own classifier — rather than choosing from the vendor library.
Where guardrails run Guardrail execution location
On the vendor's servers
Whether guardrail evaluation happens inside your infrastructure or on the vendor’s servers. This decides whether the prompt you are trying to protect leaves your network in order to be checked.
If the guardrail itself fails Guardrail failure mode
Not documented

Identity, tenant and explicit policy denial fail closed; a general guardrail timeout guarantee is not established.

What happens when the guardrail service times out or errors: does the request proceed unchecked, or is it blocked? This is the worst-documented field in the entire catalogue — only two vendors state it plainly, which means most teams are running a control whose failure behaviour they cannot know.
Third-party guardrail vendors Guardrail integrations
Not published
Named external guardrail services the product can call. A long list means the product is a router for policy engines rather than a policy engine itself — which also means another vendor bill and another hop.

Compliance evidence

Graded by how strong the evidence is, not whether the word appears on the vendor’s website. An audited report and a marketing claim are different things, and only one of them will satisfy your own auditor.

  • SOC 2 Not published Vendor explicitly does not claim SOC 2 certification.
  • ISO 27001 Not published
  • GDPR DPA Not published
  • HIPAA BAA Not published Regulated use requires a separately executed agreement.
  • FedRAMP Not published No authorization claimed.
  • ITAR Not published

No compliance certifications were found published for this product. That is not the same as failing an audit — it means there is nothing public to check, so ask for evidence directly.

Vendor source

Fit & integration

Work to try it Evaluation work shape
Set it up in a cloud console
The shape of the work on the vendor’s own quickstart, from swapping one base URL through to deploying infrastructure. An ordinal class rather than a duration, because elapsed time depends on accounts and quota we cannot see.
Work to run it Production work shape
Set it up in a cloud console
The same scale applied to the vendor’s recommended production path. For several products this is much heavier than the quickstart, which is exactly why both are recorded.
Steps on the quickstart Numbered quickstart steps
5

Vendor lists five setup steps; no elapsed-time test performed.

A literal count of numbered steps on the vendor’s quickstart, recorded as evidence beside the work shape. Zero means the page publishes no numbered procedure at all. Large counts usually mean interleaved language tracks rather than more work.
Can you self-host it today Self-host install documentation
No self-hosting
Whether an install command is actually published. Several products advertise self-hosting while publishing no command to start from, which a plain yes/no would hide.
Works with the OpenAI SDK OpenAI SDK drop-in
Partly

Endpoint replacement also requires gateway credentials and configured aliases.

Whether an existing OpenAI-compatible client can be pointed at it by changing the base URL and key.
Vercel AI SDK support AI SDK provider package
Not documented
Whether a first-party AI SDK provider package exists, or only a community package, a documented workaround, or the generic OpenAI provider pointed at a custom base URL.
Python framework integrations Documented Python frameworks
Not published
LangChain, LangGraph, LlamaIndex and similar orchestration frameworks with a documented integration.
Callable from Cloudflare Workers Cloudflare Workers support
Not documented
Whether the docs show calling this product from your own Worker. Deliberately separated from the several products whose own gateway runs on Workers, which is a fact about their infrastructure and not about your edge compatibility.
Kubernetes install Helm chart availability
Not documented
Whether a named, published Helm chart exists, versus Helm being referenced with no chart named, versus generic cluster documentation that has nothing to do with this product.
Terraform support Terraform provider or modules
Not documented

Internal AWS module exists; customer-facing resource provider is planned, not available.

Whether you can declare this in code: an official provider, official modules, resources inside a hyperscaler’s provider, a community provider, or only Terraform code shipped in a repo.
Reuses your cloud identity Cloud IAM reuse
Static provider credentials only

Workload gateway key for requests; OIDC/SAML is administrative identity.

Whether you can authenticate with IAM roles, workload identity or managed identities instead of another long-lived API key. Distinguished from products that only accept static upstream provider credentials.
Fits behind your API gateway API gateway integration
It is the API gateway

Applications call this gateway before configured upstream services.

Whether AI traffic can go through a gateway you already run, and who documented that — the product’s vendor or the gateway’s.
MCP support MCP surface shape
MCP tools in the API

Checks MCP declarations and emitted calls; not evidence of an MCP-server proxy.

Which kind of MCP support this is: a gateway that governs many MCP servers, a hosted MCP server you connect to, MCP tools accepted inside the completion API, or client tooling. These are different products behind one acronym.
Needs your own provider key Upstream provider key required
Optional

BYOK for customer accounts; managed Bedrock uses purchased credit.

Whether an upstream provider account and key must exist before your first call works. A prerequisite rather than a step, and it can differ between the hosted and self-hosted forms of the same product.
Gate before models work Model access gate
You enable it first

Configure eligible targets; managed models are reviewed for account availability.

Whether an enablement click, a quota grant, a paid tier or an approval form stands between a valid key and a working model call.
Official client languages First-party client SDK languages
Python

Examples use upstream OpenAI and Anthropic SDKs.

Languages with a first-party client library. An empty list can still mean the product is usable from any language via an OpenAI-compatible SDK.

Additional charges

These are the fees that do not appear on a per-token price list, and they are where estimates usually go wrong.

  • Flex gateway usage $0.10 per 1,000 successful requests
  • Alternative Company subscription $499/month; 2,000,000 successful requests ceiling
  • Alternative Portfolio subscription $1,500/month sponsor base; participating company costs separate

How pricing actually works

Design engagement from $60,000/year; not a ready-to-run deployment license.

Back to top ↑

Common questions

Answered from the fields above, so these move when the catalog moves. Every figure quoted here appears in the specification with its source.

Does AI Gateway HQ charge a markup on model prices?

AI Gateway HQ adds no percentage markup to model prices. Other charges on the page: flex gateway usage ($0.10 per 1,000 successful requests), alternative company subscription ($499/month; 2,000,000 successful requests ceiling) and alternative portfolio subscription ($1,500/month sponsor base; participating company costs separate).

Can AI Gateway HQ be self-hosted?

No. AI Gateway HQ is available only as a service the vendor operates; there is no self-hosted build. The licence is Proprietary hosted service. Prompts therefore leave your network and reach the vendor, which makes its retention and residency terms the control that matters here rather than deployment.

Is AI Gateway HQ SOC 2 audited, and will it sign a HIPAA BAA?

AI Gateway HQ publishes neither a SOC 2 report nor a HIPAA business associate agreement. EU data residency is not offered. Each of these is linked to the vendor's own page in the compliance section below. Ask for both documents directly before committing, because a published claim and a countersigned agreement are not the same thing.

Does AI Gateway HQ retain your prompts?

AI Gateway HQ offers zero data retention conditionally rather than by default. Only metadata is logged — not prompt or response bodies. It states that it does not train on customer data. Retention, logging and training on customer data are three separate questions, and a vendor can answer one of them without answering the others.

Can you use your own provider keys with AI Gateway HQ?

Yes. AI Gateway HQ can route through your own accounts with the underlying model providers, so inference is billed to you directly. 0% applies to BYOK inference. Flex gateway usage is $0.10/1,000 successful requests; managed Bedrock rates are separate. The billing route is worth confirming with the vendor before committing.

How many models does AI Gateway HQ support?

AI Gateway HQ publishes no total model count. It reaches 11 upstream providers. No exhaustive public total; configured routes and the reviewed managed catalog have different scopes. Counts are not comparable between vendors, because what each one counts as a distinct model differs — treat it as an order of magnitude rather than a score.

Back to top ↑

What has changed here

  1. provider added provider added Not published AI Gateway HQ source ↗
See this in the full changelog Back to top ↑

Read the head-to-head

These pairs have a written verdict, not just a table.

Usually weighed against

Follow updates about AI Gateway HQ