Self-host only Apache-2.0

Envoy AI Gateway Open source

Envoy AI Gateway is an open-source LLM gateway with an OpenAI-compatible API in front of 16–19 upstream providers; it publishes no model count. It charges no token markup. It runs only self-hosted under Apache-2.0. It does not publish a HIPAA BAA. You can point it at your own provider accounts. Beyond chat it also serves embeddings, image generation and audio. It handles failover, load balancing and request logging.

· 66 dated entries · 66 source references

Built by Envoy AI Gateway open-source project in the `envoyproxy` GitHub organization; maintainers are drawn from Tetrate, Bloomberg, Tencent, Netflix and Nutanix plus the KServe, Kubeflow, Envoy Proxy and Envoy Gateway projects

Access at a glance

Whether this product can work for you at all, before features matter: who pays the model bill, where it can run, and which of your existing API calls keep working. Every value is the vendor’s own claim, linked to the page it came from.

An infrastructure-layer AI gateway, not a platform: "an open source project for using Envoy Gateway to handle request traffic from application clients to Generative AI services", positioned explicitly as "an additive layer designed to expand use cases for Envoy Proxy and Envoy Gateway without changing existing deployment or control patterns" (README; GOALS.md, 2026-09-02). Its personas are platform engineers and infrastructure administrators; there is no dashboard, no prompt studio and no evaluation product - control surfaces are Kubernetes CRDs and telemetry.

Who pays the model bill

Your keys only

You contract with each model provider directly and hold those accounts. The gateway never resells inference.

By necessity: there are no platform credits or vendor billing, so every upstream call uses credentials you supply through BackendSecurityPolicy - API keys from Kubernetes Secrets, or cloud identity for AWS, Azure and GCP (upstream auth; connect providers, 2026-09-02).

Merchant of record: The upstream AI provider. You supply provider credentials in Kubernetes Secrets and are billed directly by OpenAI, AWS, Azure, GCP and so on; the project never takes payment or proxies billing (upstream auth; connect providers).

Key handling: Credentials are centralised in the control plane, away from applications. BackendSecurityPolicy takes an API key from a Kubernetes Secret (key name apiKey), or cloud identity: AWS Bedrock via OIDC to STS - including EKS Pod Identity and IRSA with just a region set - Azure OpenAI via Entra ID, GCP Vertex AI via workload identity federation and Google STS, each minting short-lived tokens per request; long-lived keys stay in Secrets under operator control, and the docs note secret updates are "picked up automatically in a few seconds" (upstream auth; connect providers; OpenAI guide, 2026-09-02). v1.1 added BackendSecurityPolicy.spec.credentialOverride so a per-request credential can come from dynamic metadata or a header, which is how you do per-tenant keys (v1.1 release notes).

Where it can run

1 of 5 shapes documented
  • Vendor-hosted
  • Self-host
  • Your VPC
  • On-premise
  • Air-gapped

Dimmed shapes are not documented by the vendor, which is not the same as unsupported.

Self-host only, in two shapes. The primary one is Kubernetes: two Helm charts (CRDs plus controller) on top of Envoy Gateway, driving Envoy proxies in your cluster (installation). The second is a single-process local mode - aigw run runs the same configuration API with no Docker or Kubernetes on Linux and macOS, listening on localhost:1975, intended for testing config and for local development against provider-agnostic apps (aigw run, 2026-09-02). There is no vendor-hosted or hybrid-VPC offering from the project; the managed product on this data plane is Tetrate's separate Agent Router Service.

Install order matters: Envoy Gateway first (oci://docker.io/envoyproxy/gateway-helm with the project's envoy-gateway-values.yaml, requiring v1.7.0+), then ai-gateway-crds-helm, then ai-gateway-helm into envoy-ai-gateway-system, waiting on deployment/ai-gateway-controller (prerequisites; installation). Kubernetes v1.32+ and Gateway API v1.5.x are required; the compatibility matrix pairs AI Gateway v1.0.x with Envoy Gateway v1.8.1+ and Envoy Proxy v1.38.x, and lists v0.5-v0.7 as still supported with earlier versions EOL (compatibility, 2026-09-02). Token rate limiting and quotas additionally need Redis plus rate-limit configuration on the Envoy Gateway install (usage-based rate limiting); InferencePool needs the Gateway API Inference Extension CRDs (InferencePool support).

API surfaces your code can keep using

6 of 7 documented
  • OpenAI chat POST /v1/chat/completions Yes

    POST /v1/chat/completions is listed "Fully Supported" with streaming, function/tool calling, JSON-schema structured output, audio and video inputs, the injected x-ai-eg-model routing header, token-usage and cost tracking, and fallback plus load balancing; the quickstart calls $GATEWAY_URL/v1/chat/completions (supported endpoints; basic usage, 2026-09-02).

  • Anthropic messages POST /v1/messages Yes

    Two ways: a native POST /anthropic/v1/messages surface listed "Fully Supported", and cross-provider translation of Anthropic Messages into OpenAI Chat Completions and into AWS Bedrock Converse/InvokeModel including streaming, tool use, reasoning blocks and images (supported endpoints; v1.0 release notes, 2026-09-02). v1.1 added /anthropic/v1/messages/count_tokens and /anthropic/v1/models (v1.1 release notes).

  • OpenAI Responses POST /v1/responses Yes

    POST /v1/responses is listed "Fully Supported" with MCP tools, reasoning and multimodal input, and v1.1 added /v1/responses/input_tokens (supported endpoints; v1.1 release notes, 2026-09-02).

  • Embeddings POST /v1/embeddings Yes

    POST /v1/embeddings is a supported endpoint with its own metrics surface, and embeddings translation is documented per provider in the model-virtualization table (supported endpoints; metrics; model name virtualization, 2026-09-02).

  • Images POST /v1/images/generations Yes

    POST /v1/images/generations appears as a supported endpoint, and "images" is listed in the 1.0 endpoint coverage summary (supported endpoints; v1.0 release notes, 2026-09-02).

  • Audio POST /v1/audio/* Yes

    POST /v1/audio/transcriptions and POST /v1/audio/translations are supported endpoints, and audio is part of the documented 1.0 endpoint surface (supported endpoints; v1.0 release notes, 2026-09-02).

  • Batch jobs POST /v1/batches Not documented

    No batch or async bulk endpoint appears on any page fetched: supported endpoints enumerates chat, completions, embeddings, images, audio, responses, rerank, models and tokenize surfaces with no /v1/batches, and neither the v1.0 nor v1.1 release notes mention batch.

An asterisk marks a qualified verdict: support that is indirect (SDK compatibility or provider passthrough rather than a native endpoint), or a gap that is narrower or wider than the label suggests. Read the note before porting.

One OpenAI-compatible front door plus provider-native passthrough surfaces, all mounted on the same Gateway: OpenAI paths under /, Cohere under /cohere, Anthropic under /anthropic, with the prefixes overridable through the Helm endpointConfig value (supported endpoints, 2026-09-02). Requests are matched on the model name, which the ext_proc filter extracts from the body and re-emits as the x-ai-eg-model header; the same Gateway also serves /mcp for Model Context Protocol traffic (MCP). v1.1 added a vLLM-compatible /tokenize endpoint across providers (v1.1 release notes).

How much it reaches

Models Not published
Upstream providers 16–19

n.a. - no page fetched states a model count. Checked supported providers, supported endpoints, home and the v1.0 release notes; models come from whichever upstream providers the operator configures with their own credentials.

Range 16-19 depending on the page. The supported-providers table lists 19 rows - OpenAI, AWS Bedrock, Azure OpenAI, Google Gemini on AI Studio, Google Vertex AI, Anthropic on GCP Vertex AI, Groq, Grok, Together AI, Cohere, Mistral, DeepInfra, DeepSeek, Hunyuan, Tencent LLM Knowledge Engine, Tetrate Agent Router Service, SambaNova, Self-hosted-models and Anthropic (supported providers, 2026-09-02) - while the 1.0 announcement and the v1.0/v1.1 release-note headlines count "16 AI providers" and the repo README shows 16 provider logos (1.0 announcement; README). "Support" here means an API-schema mapping on AIServiceBackend plus an upstream-auth type on BackendSecurityPolicy, not a hosted model catalogue.

Whose models: All third-party or operator-hosted: the project owns no models and no inference hardware. It routes to 18 named external providers plus self-hosted servers such as vLLM, and to Gateway API InferencePool endpoints you run yourself (supported providers; InferencePool support, 2026-09-02).

Your own endpoints: Yes: a Self-hosted-models row in the provider table maps any OpenAI-schema server (vLLM is the named example) onto an AIServiceBackend pointing at an Envoy Gateway Backend or Kubernetes Service, with optional API-key auth (supported providers; connect providers). Self-hosted fleets can instead be fronted by a Gateway API InferencePool with an endpoint picker that routes on live KV-cache usage, queue depth and LoRA adapter state (InferencePool support). Locally, aigw run auto-configures against any OPENAI_BASE_URL, with Ollama as the documented example (aigw run, 2026-09-02).

How it behaves in production

What happens when an upstream model is slow, wrong, or down — and what you can see and stop while it happens. Reliability features are recorded as where you configure them, not whether the vendor lists them, because almost every product here lists all of them.

4 of 6 reachable from code nothing documented on the request path 5 documented destinations

When something goes wrong

Each row says where the knob is, not whether the feature is on the marketing page. A control you can only reach by hand in someone else’s dashboard cannot be reviewed or version-controlled.

  • Request timeout In config

    Changing it means editing configuration and shipping it, so behaviour is uniform across traffic until you redeploy.

    Config-file only, and split across two APIs. Request timeouts come from Envoy Gateway's BackendTrafficPolicy (the fallback example sets timeout: 30s with perRetry.backOff.baseInterval 100ms and maxInterval 10s) (provider fallback); v1.1 added an AI-specific AIGatewayRouteRule.streamIdleTimeout that triggers failover if it fires before the first token and returns 504 if it fires mid-stream (v1.1 release notes, 2026-09-02). No default value for either is stated on the pages fetched.

  • Retries In config

    Retries are Envoy Gateway's, configured in YAML next to the AI routes: the documented example sets numRetries: 5, numAttemptsPerPriority: 1, exponential backoff 100ms/10s, retryOn.httpStatusCodes: [500] and triggers connect-failure and retriable-status-codes, with numAttemptsPerPriority controlling how many tries each priority group gets before failing over (provider fallback, 2026-09-02).

  • Fallback to another model In config

    Ordered priority groups: backendRefs in an AIGatewayRoute rule carry priority: 0, priority: 1 and so on, and traffic moves to the next priority when the current one exhausts its attempts. Because the ext_proc filter chain is split into router-level and upstream-level filters, a failover to a different provider re-runs request translation and upstream auth for the new backend - which is what makes cross-provider fallback work rather than just cross-endpoint retry (provider fallback; data plane, 2026-09-02). Weighted fallback is not documented.

  • Load balancing In config

    Config-file. Multiple backendRefs at the same priority are load-balanced by Envoy, and Envoy Gateway's BackendTrafficPolicy governs the algorithm (provider fallback); for self-hosted fleets an InferencePool endpoint picker does inference-aware selection on real-time KV-cache usage, queued requests and LoRA adapter state (InferencePool support, 2026-09-02).

  • Upstream health tracking Not documented

    The vendor does not document this, so any behaviour you observe today is unversioned and may change.

    n.a. - no active health check or circuit breaker for AI backends on any page fetched (provider fallback, system architecture, data plane, capabilities index). Failure detection is reactive: retries and priority failover on connect failures and retriable status codes. The one exception is InferencePool, where an endpoint picker selects self-hosted endpoints from live metrics such as KV-cache usage and queue depth (InferencePool support, 2026-09-02).

  • Cross-region failover Not documented

    n.a. - no cross-region or multi-cluster failover construct is documented; the two-tier gateway pattern in the README is about tier-one ingress versus tier-two self-hosted model access, not geography (README; provider fallback, 2026-09-02).

Fallback chain: Ordered list — Try A, then B, then C. Simple and predictable, but every failover is all-or-nothing.

Defaults: The numbers above are the values in the documentation's example manifest, not stated defaults; no page fetched publishes a default retry count or backoff for AI routes (provider fallback, 2026-09-02).

Reliability is inherited rather than reinvented: this is Envoy Proxy's retry, timeout, priority and load-balancing machinery driven by Gateway API resources, with AI-specific additions where translation matters (per-priority attempt counts, upstream-level filters that re-auth on failover, streamIdleTimeout failover before first token, 429 on token quota exhaustion). The gaps are health checking, circuit breaking and multi-region failover, none of which appear in the AI Gateway docs. Operationally the controller supports leader election plus a horizontally scalable read-only extension server, with an HPA example at 70% CPU, min 2 / max 10 replicas (scaling; provider fallback; v1.1 release notes, 2026-09-02).

How fast the hop is

Compiled binary

A single compiled Go or Rust binary. The lowest overhead floor of the self-hostable options, and the easiest to reason about under load.

Go control plane and Go ext_proc filter alongside the Envoy C++ data plane: repo language bytes are Go 7,634,320, MDX 1,427,966, CSS 44,361 and TypeScript 43,527, built against Go 1.26.4 (GitHub API; v1.1 release notes, 2026-09-02). Requests traverse Envoy Proxy plus an External Processor container that does model extraction, routing, upstream auth, request/response translation and token accounting, with a Rate Limit Service and Redis for token limits; the filter chain is deliberately split into router-level and upstream-level stages so a retry to another provider re-runs translation and auth. Dynamic Modules are noted as a possible future alternative to ext_proc (system architecture; data plane).

You can run the request path yourself Yes
Streaming Yes

Everything is published: Helm charts as OCI artifacts (oci://docker.io/envoyproxy/ai-gateway-crds-helm and oci://docker.io/envoyproxy/ai-gateway-helm, pinned in the docs to --version v1.0.0) (installation); a CLI container image at envoyproxy/ai-gateway-cli and per-release aigw binaries for Linux and macOS on the GitHub releases page, or go install ./cmd/aigw from source (CLI installation, 2026-09-02); the v1.1.0 release itself is on GitHub (releases API).

Streaming caveats: Yes - SSE streaming is listed as supported on chat completions and Anthropic messages, with translation preserved across providers including tool use and reasoning blocks (supported endpoints; v1.0 release notes). Two operational caveats the docs raise: Envoy Gateway's default 32 KB buffer limit "is not enough for most AI model responses", so the example manifest raises it to 50 MB via ClientTrafficPolicy (basic usage); and streamIdleTimeout returns 504 if it fires mid-stream, while firing before the first token triggers failover instead (v1.1 release notes, 2026-09-02).

Published figures, grouped by what each one measured. Figures in different groups are different quantities and cannot be compared with one another — nor, in most cases, with another vendor’s figure in the same group.

Overhead added by the gateway

The only figures that speak to whether this product slows your application down. Still not comparable between vendors: each was measured on different hardware, at a different load, with a different payload.

  • ~2 ms data-plane gateway overhead (~0.01% of end-to-end latency)

    Attributed to an independent Broadcom/VMware Cloud Foundation validation under sustained enterprise LLM load with real GPU inference; overhead reported flat at peak saturation. The millisecond figure appears only in Tetrate's summary page, and Tetrate co-maintains the project; Broadcom's own post states no ms value. Percentile not stated.

    Source
  • 160-390 ms added latency on an MCP tool call vs direct unproxied call Vendor-published

    December 2025 Go microbenchmark of a simple MCP echo tool, standalone single process, 100 key-derivation iterations for encryption setup; average difference against a competing implementation ~0.2 ms. Summarised by Tetrate; original benchmark post not fetched.

    Source

Sustained capacity

Requests or queries per second sustained on the stated hardware.

  • ~5 s route readiness (AIGatewayRoute creation to serving traffic) Vendor-published

    Project-run control-plane benchmark at 2,000 routes; the delay is the 5-second default poll interval of the ext_proc config watcher, not request-path latency. Mock cassette backend, no headerMutation.

    Source
  • 2000 AIGatewayRoutes control-plane configuration scale, all routes verified serving Vendor-published

    Required raising the Envoy Gateway extension-manager gRPC message size to 25Mi and the controller `maxRecvMsgSize` to 26214400; runs excluded `headerMutation`; effective ceiling bounded by the ~1 MB Kubernetes Secret size limit for aggregated filter config.

    Source
  • 224 concurrent users saturation point (TTFT climbs sharply)

    Broadcom/VMware run as summarised by Tetrate; described as a GPU compute ceiling on a four-H100 test cluster rather than a gateway limit. Not an RPS figure.

    Source

The 2 ms and 224-user figures originate with Broadcom's VMware Cloud Foundation team (third party) but are quoted from a Tetrate page, and Tetrate co-created and maintains the project; the 2,000-route and MCP figures are project or maintainer-published. None is a neutral head-to-head against other gateways with policies enabled.

What it will stop

nothing documented on the request path

No request-path policy controls are documented. That is not a fault in a product built purely for routing — but it means anything you need blocked has to be blocked before the call reaches here.

What you can see

Exports to a few places
What gets logged Your choice

You decide whether bodies are captured, by setting or by header.

You can turn bodies off Yes

Yes, and it is total: with no OTLP endpoint configured nothing is exported at all, and per-field redaction flags plus the gen_ai semantic-convention mode let you keep spans without message content (tracing, 2026-09-02).

Traces OpenTelemetry

Full OpenTelemetry: set OTEL_EXPORTER_OTLP_ENDPOINT (globally or per-Gateway through GatewayConfig.spec.extProc.kubernetes.env) and the ext_proc emits spans, defaulting to OpenInference conventions with AI_GATEWAY_TRACING_SEMCONV=gen_ai as the alternative; MCP CallTool and ListTools operations get spans too, and a session.id header mapping groups multi-turn conversations. Arize Phoenix is the documented consumer, giving LLM-as-judge evaluation over production spans (tracing; gateway config, 2026-09-02).

Content capture is a switch, and the default depends on which semantic convention you pick. OpenInference (the default) records inputs and outputs unless you set OPENINFERENCE_HIDE_INPUTS, OPENINFERENCE_HIDE_OUTPUTS, OPENINFERENCE_HIDE_EMBEDDINGS_TEXT or OPENINFERENCE_HIDE_EMBEDDINGS_VECTORS; switching to AI_GATEWAY_TRACING_SEMCONV=gen_ai makes content opt-in via OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=true (tracing, 2026-09-02).

Where telemetry can go

  • OpenTelemetry (OTLP)
  • Prometheus
  • Envoy access logs
  • Arize Phoenix
  • Grafana

OTLP for traces to any collector (tracing); Prometheus scrape for GenAI token and latency metrics, with custom labels lifted from request headers (metrics); Envoy access logs carrying gen_ai.* token costs and MCP fields such as mcp_tool_name and mcp_backend, configured on EnvoyProxy.spec.telemetry.accessLog (access logs). Named downstreams in the docs are Arize Phoenix via its Helm chart and a bundled Grafana dashboard at examples/monitoring/grafana-dashboard.json (tracing; v1.1 release notes, 2026-09-02).

Records user feedback No
Scores live traffic Partly

n.a. - no thumbs-up/score ingestion API or feedback header on any page fetched (tracing; metrics; API reference, 2026-09-02).

Partial and second-hand: the gateway emits OpenInference-compatible spans with full message content, and the docs walk you through deploying Arize Phoenix so "LLM-as-judge" evaluations can run against production spans - the evaluation itself belongs to Phoenix, not to the gateway, which ships no dataset, scorer or experiment surface (tracing, 2026-09-02).

Depends on the vendor’s SaaS: No. Metrics are scraped from your own Prometheus and traces go to any OTLP collector you run; there is no vendor dashboard, account or telemetry egress in the product (metrics; tracing, 2026-09-02).

Retention: n.a. - no vendor-side log retention exists. The only project-side storage described anywhere is the Kubernetes Secret the controller writes to hand filter configuration to ext_proc, which is configuration rather than traffic (control-plane scaling).

Whether it fits how you work

How much work stands between you and a first call, how different that is from running it in production, and whether it slots into the stack you already have. Recorded as the shape of the work rather than a number of minutes — how long it takes you depends on which accounts and quota you already hold, which no comparison can know.

to try: cli or container to run: infrastructure rollout fits 5 of 10 common stacks

Getting to a first call

4 numbered steps
Shape of the work Run something locally first

Nothing works until you have a process running. Fine on a laptop, but it means there is no zero-install way to try it.

Read off: the vendor’s own quickstart — 4 numbered steps.

The getting-started guide numbers four sections (Prerequisites, Installation, Basic Usage, Connect Providers) rather than four commands; the Installation page itself has two numbered Helm steps and Prerequisites adds the Envoy Gateway install, so there is no single steps-to-first-call figure (getting started; installation, 2026-09-02).

Before step one

  • Your own provider key Required

    You need an upstream provider account and key before anything works. That is a prerequisite, not a step.

    For any real traffic: each AIServiceBackend needs a BackendSecurityPolicy with your provider credential (connect providers). The exception is evaluation - the quickstart's basic.yaml routes to a mock backend (some-cool-self-hosted-model) so you can make a first call with no provider account at all (basic usage, 2026-09-02).

  • Payment method No card needed to start

    Not required - there is no account, sign-up or billing relationship anywhere in the install path; you need a Kubernetes cluster or a local machine plus, for real calls, your own provider key (installation; aigw run, 2026-09-02).

  • Gate before models answer No gate

    Every catalogue model is callable as soon as you have a key.

    No approval, waitlist, enablement or quota step exists on the project side - whatever your provider credential can reach is reachable once you declare it in a route (connect providers; supported endpoints, 2026-09-02). Any gating you experience is your upstream provider's.

Everything you need first: Two paths. Local: the aigw binary or envoyproxy/ai-gateway-cli image on Linux or macOS, one environment variable (OPENAI_API_KEY, or OPENAI_BASE_URL for Ollama and other OpenAI-compatible servers), listening on localhost:1975 (aigw run; CLI installation). Kubernetes: kubectl, helm and curl, a cluster on v1.32+, and Envoy Gateway v1.7.0+ installed first with the project's values file; the quickstart then works against a mock backend with no provider key (prerequisites; basic usage, 2026-09-02).

Running it in production

The same scale applied to the path the vendor recommends for production traffic. Kept separate from the quickstart because for several products here the two are barely related pieces of work.

Shape of the work Deploy it on your infrastructure

This is an infrastructure project, not an integration. Expect charts or Terraform, networking, secrets and someone who owns the deployment.

If you self-host it instead Deploy it on your infrastructure

This is an infrastructure project, not an integration. Expect charts or Terraform, networking, secrets and someone who owns the deployment.

Getting to production is a step up in kind from the quickstart, not just more of the same.

What production needs: Kubernetes v1.32+ with Gateway API v1.5.x and Envoy Gateway v1.8.1+ (Envoy Proxy v1.38.x) for AI Gateway v1.0.x; kubectl, helm and curl to install; provider credentials in Kubernetes Secrets or cloud identity wired to BackendSecurityPolicy; Redis plus Envoy Gateway rate-limit configuration if you want token limits or quotas; an OTLP collector and Prometheus if you want traces and metrics; Gateway API Inference Extension CRDs for InferencePool routing (compatibility; prerequisites; usage-based rate limiting; tracing; InferencePool support, 2026-09-02). Plan capacity from the documented controller footprint - 100m/128Mi requests, 500m/256Mi limits, HPA to 10 replicas - and note the ~1 MB Kubernetes object limit on the aggregated filter-config Secret when running thousands of routes (scaling; control-plane scaling).

Can you run it yourself

Install command published Install command published

There is a command you can copy and run, so you can evaluate the self-hosted path yourself today.

helm upgrade -i aieg-crd oci://docker.io/envoyproxy/ai-gateway-crds-helm --version v1.0.0 --namespace envoy-ai-gateway-system --create-namespace then helm upgrade -i aieg oci://docker.io/envoyproxy/ai-gateway-helm --version v1.0.0 --namespace envoy-ai-gateway-system --create-namespace, after Envoy Gateway v1.7.0+; locally docker run --rm -p 1975:1975 -e OPENAI_API_KEY=... envoyproxy/ai-gateway-cli run or go install ./cmd/aigw

How it fits your stack

5 of 10

Each row is a thing you might already run. “With a caveat” means it works but not the way the vendor’s marketing implies — read the reason, because that is usually where the surprise lives.

  • Fits The OpenAI SDK Drop-in once set up — but first-call work is cli or container.
  • With a caveat The Vercel AI SDK Via the OpenAI provider
  • No Cloudflare Workers No Workers guidance published.
  • Fits Kubernetes oci://docker.io/envoyproxy/ai-gateway-helm plus oci://docker.io/envoyproxy/ai-gateway-crds-helm (Envoy Gateway's oci://docker.io/envoyproxy/gateway-helm required first)
  • No Terraform or OpenTofu Nothing published for Terraform.
  • Fits An existing API gateway This is that gateway — AI traffic becomes a plugin, not a new hop.
  • Fits Cloud IAM I already run Reuses IAM roles, workload identity or managed identities.
  • No LangChain or LlamaIndex No framework integration documented.
  • Fits MCP servers to govern Acts as an MCP gateway or registry.
  • With a caveat Nothing — plain Node or Python You have to run a process locally before any call works.

Reading this the other way round — pick what you already run and see every product scored against it.

The integration surfaces behind those answers

  • Vercel AI SDK Via the OpenAI provider

    Use the OpenAI provider pointed at a custom base URL. Works, but you lose provider-specific options.

    No Vercel AI SDK provider package or integration page exists - the term appears nowhere on the pages fetched or in the site map. What is documented is generic OpenAI compatibility at the gateway root, which is what an @ai-sdk/openai baseURL override would target (supported endpoints; aigw run, 2026-09-02).

  • Cloudflare Workers Not documented

    No Workers guidance either way. If you are edge-first, verify fetch-only compatibility yourself.

    n.a. - no Cloudflare Workers or edge-runtime story; the data plane is Envoy in your cluster or the aigw binary on Linux/macOS (data plane; aigw run, 2026-09-02). Cloudflare is not even listed as an upstream provider (supported providers).

  • Kubernetes Official Helm chart

    A named, published chart. You can read its values file before committing to anything.

    Named: oci://docker.io/envoyproxy/ai-gateway-helm plus oci://docker.io/envoyproxy/ai-gateway-crds-helm (Envoy Gateway's oci://docker.io/envoyproxy/gateway-helm required first)

    Kubernetes is the primary target, not a deployment option: two official OCI Helm charts (CRDs and controller), a documented CRD-ownership migration with --take-ownership, controller leader election with a horizontally scalable read-only extension server, an HPA example (70% CPU, min 2 / max 10), and v1.1 Helm additions for PodDisruptionBudget, topology spread constraints, pod labels and a restricted controller security context (installation; scaling; v1.1 release notes, 2026-09-02).

  • Terraform Not documented

    No Terraform surface published. Configuration is API or dashboard work.

    n.a. - no Terraform provider, module or example is published: nothing in the repository root or the 469-URL site map mentions Terraform, and the documented install path is Helm (installation; GitHub repo, 2026-09-02).

  • Existing API gateway It is the API gateway

    This product is the gateway. If you already run it for your other APIs, AI traffic becomes a plugin rather than a new hop.

    It is the gateway - but uniquely among these entries it is not a standalone one: Envoy Gateway is a hard prerequisite (v1.7.0+ for install, v1.8.1+ in the v1.0.x compatibility matrix) and AI Gateway is an extension server plus ext_proc filter on top of it, described as "an additive layer" (prerequisites; compatibility; GOALS.md, 2026-09-02).

  • Cloud identity Reuses your cloud identity

    Authenticate with the identity you already run — IAM roles, workload identity or managed identities. No long-lived key to rotate.

    Native cloud identity on all three hyperscalers: AWS Bedrock through the default credential chain including EKS Pod Identity and IRSA (region only, no key material) or OIDC-to-STS temporary credentials; Azure OpenAI through Entra ID short-lived tokens; GCP Vertex AI through Application Default Credentials, service-account keys or Workload Identity Federation with Google STS (connect providers; upstream auth; supported providers, 2026-09-02).

  • MCP MCP gateway or registry

    It sits in front of many MCP servers and governs access to them. This is the shape you want if MCP sprawl is the problem you are solving.

    A genuine MCP gateway, configured with an MCPRoute CRD: multiple MCP servers multiplexed behind one endpoint, tool names namespaced by backend (github__issue_read), tool filtering by name or regex, OAuth 2.0 authorization-code flow with PKCE plus JWT scope/claim and CEL-based authorization with a configurable defaultAction, header forwarding, upstream API-key injection, streamable HTTP per the June 2025 MCP spec, and OTel spans plus Prometheus metrics for MCP traffic (MCP, 2026-09-02). The same thing runs locally via aigw run --mcp-config on http://localhost:1975/mcp (aigw run). A dedicated MCPBackend CRD is on the post-1.0 roadmap (1.0 announcement).

Python frameworks None documented

n.a. - no LangChain, LlamaIndex, DSPy or other framework integration is documented; the 469-URL site map contains no framework page and none of the docs pages fetched name one (capabilities index; supported endpoints, 2026-09-02). In practice you would use each framework's OpenAI-compatible client with the gateway's base URL.

First-party client libraries None — you call it with an OpenAI-compatible client in whatever language you like

No client library is published or required. The documented pattern is pointing an OpenAI-compatible client at the gateway; aigw run reads the OpenAI SDK's environment variables so existing code needs no change, and the InferencePool page cites "seamless integration with OpenAI SDKs" (aigw run; InferencePool support, 2026-09-02). No per-language SDK list, snippet set or package name appears on the pages fetched, so fit_client_sdk_langs is left empty rather than inferred from the curl examples.

Agent features: Agent-shaped features are the project's growth area: the Responses API surface supports MCP tools, reasoning and multimodal input (supported endpoints), and the MCP gateway multiplexes many MCP servers behind one /mcp endpoint with tool-name prefixing (github__issue_read), include/regex tool filtering, OAuth 2.0 authorization-code + PKCE with JWT scope and CEL claim checks, upstream API-key injection, header forwarding and streamable HTTP transport per the June 2025 MCP spec, all traced and metered (MCP, 2026-09-02). v1.1 added MCPRoute.spec.hostnames (max 16) and a CEL backendSelector defaulting to Deny (v1.1 release notes). Locally, aigw run --mcp-config accepts the same mcpServers JSON that Claude Desktop, Cursor and VS Code use (aigw run). There is no agent loop, no tool executor and no prompt/agent registry - tool execution stays with the client.

Three things bite newcomers, and all are in the docs. Envoy Gateway must be installed first, with the project's own envoy-gateway-values.yaml, or nothing reconciles (prerequisites). Envoy Gateway's default 32 KB client buffer is too small for model responses, which is why the basic example ships a ClientTrafficPolicy raising it to 50 MB (basic usage). And token rate limiting or quotas need a Redis deployment plus rate-limit configuration chosen at Envoy Gateway install time, not afterwards (usage-based rate limiting; quota policy). Also note the docs' own caution about v0.0.0-latest chart tags being unstable and the advice to pin a commit (installation, 2026-09-02).

The ecosystem is Kubernetes-native and standards-first rather than SaaS-integration-first: Gateway API v1.5.x, Gateway API Inference Extension v1.0.2, Envoy Gateway v1.8.1 and Envoy Proxy v1.38.1 as of v1.1.0, with the MCP Go SDK v1.7.0 for MCP support (v1.1 release notes; compatibility). Observability plugs into Prometheus, OTLP collectors, Arize Phoenix and a bundled Grafana dashboard (metrics; tracing). Named adopters on the home page are Alan by Comma Soft, Bloomberg, LY Corporation, National Research Platform, Nutanix, Paper Compute Co., Simplifai, Stacklok, Tencent Cloud, Tetrate and Unwrap (home, 2026-09-02). What is absent is application-framework glue: no LangChain, LlamaIndex, Vercel AI SDK or Terraform documentation exists on the site.

Silence in the docs: 5 of the integration questions on this card have no published answer either way. That is recorded as undocumented, not as a no — but it does mean you would be verifying it yourself.

What it does well

  • Apache-2.0, self-host-only, with no vendor in the request path and no billing relationship at all
  • Stable v1beta1 control-plane CRDs with an explicit never-break promise and documented migration paths
  • Broad endpoint coverage across 19 provider configurations: chat, completions, embeddings, images, audio, Responses, Cohere rerank and native Anthropic messages, with cross-provider translation
  • Real MCP gateway: server multiplexing, tool filtering, OAuth 2.0 + PKCE and CEL authorization, plus MCP spans and metrics
  • Cloud-native upstream auth - EKS Pod Identity/IRSA, Entra ID, GCP workload identity federation - minting short-lived credentials per request
  • OpenTelemetry GenAI metrics and OpenInference tracing with a documented Arize Phoenix evaluation path, all into infrastructure you own
  • Multi-vendor maintainer base (Tetrate, Bloomberg, Tencent, Netflix, Nutanix) on the CNCF Envoy foundation

Where it falls short

  • No content guardrails whatsoever - no PII, moderation, injection or custom evaluator surface anywhere in the docs
  • No gateway-side caching: prompt caching is passthrough of provider cache_control breakpoints only
  • Envoy Gateway plus Kubernetes v1.32+ is a hard prerequisite, and token limits additionally require Redis
  • No health checks, circuit breaking or multi-region failover documented for AI backends
  • QuotaPolicy is v1alpha1-only, outside the stability guarantee, and its serviceQuota field is accepted but not enforced end-to-end
  • No prompt management, cost dashboard, virtual keys for downstream callers or vendor SLA
  • No Terraform, Vercel AI SDK or Python-framework integration documented; ecosystem glue is left to you
  • The only data-plane overhead figure (~2 ms) comes from a maintainer-published summary of a third-party benchmark, not from a neutral head-to-head

Choose it when

Platform teams already running Envoy Gateway or Gateway API on Kubernetes who want provider-agnostic LLM and MCP routing, cloud-native upstream credential handling and OpenTelemetry GenAI telemetry as ordinary infrastructure, with a stable CRD API and no vendor in the request path.

Look elsewhere when

You want content guardrails, PII redaction, a semantic cache, prompt management, cost dashboards or a hosted option out of the box - none of those exist here - or you do not run Kubernetes and do not want Envoy Gateway as a hard prerequisite.

Open source: You run the gateway yourself. Full control over the data path and no per-request vendor fee, in exchange for operating it.

How hard is it to leave?

Derived from six published facts, not from an opinion. The weights are fixed and the same for every product — see the arithmetic.

100 /100 Easy to leave
Portability score breakdown for Envoy AI Gateway
What helps you leave Points Source
Works with standard OpenAI code Switching away is a base-URL change rather than a rewrite of every call site. 22 /22 vendor page
No vendor-specific SDK required A proprietary client library spreads through your codebase and has to be torn out again. 10 /10 vendor page
Can use your own provider accounts Your keys and billing relationship stay yours, so removing the gateway does not cut off model access. 20 /20 vendor page
Can be self-hosted You can run it yourself instead of accepting a pricing or policy change. 20 /20 vendor page
Configuration lives in version control Routing and budget rules are a file you keep, not dashboard state you would have to rebuild. 16 /16 —
Your request history can be exported You leave with your own logs instead of abandoning them. 12 /12 —

Read the fine print: Configuration is portable YAML: `AIGatewayRoute`, `AIServiceBackend`, `BackendSecurityPolicy`, `GatewayConfig`, `MCPRoute` and `QuotaPolicy` CRDs, with the same schema accepted by `aigw run` as a local config file ([resources](https://aigateway.envoyproxy.io/docs/concepts/resources); [aigw run](https://aigateway.envoyproxy.io/docs/cli/aigwrun)). Telemetry leaves over Prometheus and OTLP ([metrics](https://aigateway.envoyproxy.io/docs/capabilities/observability/metrics); [tracing](https://aigateway.envoyproxy.io/docs/capabilities/observability/tracing)).

All six inputs are published, so this is scored against the full 100. This measures technical switching cost only. It does not price the engineering time to re-test prompts against a different routing stack.

Envoy AI Gateway models & pricing

Browse every imported listing from this provider, with published token rates and a link to compare other providers for the same model. This is provider-reported coverage; an absent listing does not mean unsupported.

Loading model listings…

Official model coverage source ↗ · Model source coverage and limitations

Full specification

Every field we track. Blank fields say "Not published" rather than "No" — we do not infer an absence from silence. Switch to Technical in the header for the precise field names and the low-level details.

Overview

What kind of product Category
Open source
Marketplaces resell many providers behind one key. Gateways add governance on top. Open-source projects you run yourself. Cloud platforms are hyperscaler surfaces. Inference providers host models on their own hardware.
Who runs it Deployment model
Self-host only
Managed means the vendor operates it. Self-host means you run it on your own infrastructure. Both means you can choose.
Licence Licence
Apache-2.0
Proprietary products cannot be inspected or forked. Open licences such as MIT and Apache-2.0 let you audit, modify, and run the code without permission.
Company Company
Envoy AI Gateway open-source project in the `envoyproxy` GitHub organization; maintainers are drawn from Tetrate, Bloomberg, Tencent, Netflix and Nutanix plus the KServe, Kubeflow, Envoy Proxy and Envoy Gateway projects
The organisation that maintains the product.
Who you would be signing with Vendor status
Run by a software foundation
Whether the product is still an independent company, has been acquired, is a large cloud vendor’s product line, is run by a software foundation, or has been put into maintenance mode. Maintenance mode means bug fixes and security patches only — no new features.
Last shipped an update Latest release
2026-08-21

v1.1.0, published 2026-08-21T18:23:16Z, not a prerelease ([releases API](https://api.github.com/repos/envoyproxy/ai-gateway/releases/latest)). It is the first minor release on the stable 1.x API and added `/tokenize` across providers, `credentialOverride` for per-request credentials, `GatewayConfig.spec.forwardProxy` for HTTP CONNECT egress, `streamIdleTimeout`, MCP `hostnames` and CEL `backendSelector`, the `gen_ai` tracing convention plus a Grafana dashboard, JSON controller logs, and Helm PDB/topology-spread support ([v1.1 release notes](https://aigateway.envoyproxy.io/release-notes/v1.1), 2026-09-02).

The date of the most recent release or version tag. A product that has not shipped in a year is a different risk from one that shipped last week, regardless of what its marketing site says.
GitHub stars GitHub stars
2,112
A rough proxy for community size on open-source projects. Not a quality measure.

Cost

Markup on model prices Token markup
None
How much the product adds on top of what the underlying model provider charges. Zero means you pay the same per-token price you would pay the model provider directly.
Fee to add funds Credit purchase fee
Not published
A percentage charged when you top up your balance, separate from token prices. It is easy to miss because it does not appear on the per-token price list.
Monthly cost per person Seat fee
Not published
A recurring per-user platform charge that applies regardless of how much you use the models.
Can use your own provider accounts BYOK supported
Yes
Bring Your Own Key: you keep direct contracts with OpenAI, Anthropic and others, and the gateway only routes traffic. This preserves negotiated rates and committed-spend discounts.
Free tier Free tier
The whole project is free: Apache-2.0 licence, Helm charts on Docker Hub (`oci://docker.io/envoyproxy/ai-gateway-helm`), a CLI image (`envoyproxy/ai-gateway-cli`) and per-release CLI binaries, with no account, plan or quota. Your costs are the Kubernetes cluster and the upstream provider tokens.
What you can do without paying, useful for evaluation.
Enterprise plan from Enterprise plan from
Not published
Annual entry price for the enterprise tier, where one is published or credibly reported.
How the vendor makes money Pricing model
Open source, no paid tier
The shape of the vendor’s bill: does the routing layer charge a percentage on top of tokens, a flat monthly fee, both, neither (because inference is the product), or nothing at all (open source with no paid tier).
How pricing works, briefly Pricing model detail
Apache-2.0 open source with no vendor-priced tier at all: the site has no pricing page and the only artifacts are Helm charts, container images and CLI binaries ([LICENSE](https://raw.githubusercontent.com/envoyproxy/ai-gateway/main/LICENSE); [installation](https://aigateway.envoyproxy.io/docs/getting-started/installation); [CLI install](https://aigateway.envoyproxy.io/docs/cli/aigwinstall), 2026-09-02). A managed product built on the same data plane is sold separately by Tetrate as Tetrate Agent Router Service, which appears in this project only as one more upstream provider you can route to ([supported providers](https://aigateway.envoyproxy.io/docs/capabilities/llm-integrations/supported-providers)) - it is a different product and is not scored here.
A one-paragraph description that covers the caveats a pricing category cannot: introductory rates, per-feature meters, tier gating, and pricing that resets on a specific date.
Minimum commitment Minimum commitment
None. Apache-2.0 licence, no contract, no seat count and no vendor relationship ([LICENSE](https://raw.githubusercontent.com/envoyproxy/ai-gateway/main/LICENSE)).
Whether the vendor requires a minimum contract term, a minimum spend, or a provisioned-capacity purchase to get its published rate.
Charges that fire after you go over an allowance Overage terms
n.a. - there is no vendor meter to overrun. Nothing on the [installation](https://aigateway.envoyproxy.io/docs/getting-started/installation) or [release notes](https://aigateway.envoyproxy.io/release-notes/) pages describes paid units; quota enforcement is something you configure for your own tenants with `QuotaPolicy` ([quota policy](https://aigateway.envoyproxy.io/docs/capabilities/traffic/quota-policy)).
The line items that scale with usage after an included allowance is exhausted — log storage, extra requests, per-feature meters, data export — which is where cost estimates usually go wrong.
Prompt cache offered Cache mechanism
Passes provider caching through
Whether the gateway offers its own response cache, what kind of match it does (exact request, prefix, semantic), or simply passes provider caching through unchanged.
Discount on cached input Cache-read discount
Not published
How much cheaper cached tokens are than fresh input, when the vendor publishes a single figure. Bundled-inference clouds usually price this per model instead of as one number.
Premium on cache writes Cache-write premium
Not published
How much more the first write of a cached prefix costs versus a plain input token. A high write premium and a low hit rate can leave you paying more than you save, so this matters as much as the read discount.
Who captures the cache saving Cache economics
No gateway-side cache and therefore no cache pricing. What exists is provider-agnostic passthrough of Anthropic-style `cache_control: {"type": "ephemeral"}` breakpoints - native on Anthropic, translated for Claude on GCP Vertex AI and AWS Bedrock, minimum 1,024 cacheable tokens and at most 4 breakpoints, with `prompt_tokens_details.cached_tokens` returned in usage. Any discount is the provider's, not the gateway's ([prompt caching](https://aigateway.envoyproxy.io/docs/capabilities/llm-integrations/prompt-caching), 2026-09-02).
Whether the customer keeps the full saving from caching or the vendor captures part of it — and any conditions attached (write premium, storage fees, best-effort hits).
What you can split spend by Cost attribution
Per-request token and cost metadata rather than a billing console: `llmRequestCosts` on `AIGatewayRoute` exposes InputToken, CachedInputToken, OutputToken, TotalToken and CEL-computed costs as Envoy dynamic metadata under `io.envoy.ai_gateway`, surfaced in access logs as `gen_ai.*` fields; OpenTelemetry GenAI metrics carry provider and model attributes and can be labelled with arbitrary request headers (for example a tenant header) via `controller.metricsRequestHeaderAttributes`. Per-team or per-key rollups are yours to build in Prometheus or your log store.
The dimensions the vendor documents for splitting spend — per key, per user, per team, per tag, per customer. Matters if you need to chargeback internally or bill an end customer.
How you get cost data out Cost export
No cost CSV, invoice or billing API - the project has no billing. Cost signals leave through Prometheus metrics, OTLP traces and Envoy access logs that you own ([metrics](https://aigateway.envoyproxy.io/docs/capabilities/observability/metrics); [tracing](https://aigateway.envoyproxy.io/docs/capabilities/observability/tracing); [access logs](https://aigateway.envoyproxy.io/docs/capabilities/observability/accesslogs)).
The mechanisms the vendor publishes for exporting cost and usage data: CSV, an API, webhooks, S3, a data warehouse, or nothing at all. Any per-unit price is included.
Who pays the model bill BYOK mode
Your keys only
Whether you bring your own provider accounts (BYOK), buy inference from this vendor, or can do either. This is the single biggest commercial difference between these products: it decides who holds the contract with the model provider and who carries the spend.

Spend governance in detail

Seven signals matter when a bill starts to hurt: who can spend, how much, on what, and who gets paged when it goes wrong. Everything below is drawn from the vendor’s own pricing and docs pages — how we read these.

  • Virtual or scoped keys Not published

    Keys that carry their own budget and rate-limit policy, so an intern experiment cannot spend against a production budget.

    No virtual-key concept for downstream callers on any page fetched; client authentication is delegated to Envoy Gateway SecurityPolicy (JWT, API key, basic auth, mTLS, OIDC, external authorization) ([security](https://aigateway.envoyproxy.io/docs/capabilities/security/)).

  • Budget caps per key Yes

    A dollar or token ceiling attached to an individual key. Where enforcement is soft, one over-limit request still completes before the block kicks in.

    `QuotaPolicy` per-model token budgets with `clientSelectors` bucket rules keyed on request headers, so a per-tenant or per-key budget is expressible; shadow mode evaluates rules without enforcing ([quota policy](https://aigateway.envoyproxy.io/docs/capabilities/traffic/quota-policy)).

  • Budget caps per team or workspace Yes

    A ceiling applied at a higher scope than one key — a team, a workspace, a customer, or an entire environment.

    Same mechanism - bucket rules on a tenant header carve per-team budgets; there is no team object in the product ([quota policy](https://aigateway.envoyproxy.io/docs/capabilities/traffic/quota-policy)).

  • Rate limiting as a cost control Yes

    Configurable request-per-time-window caps. Platform-set rate limits do not count as spend controls; user-configurable ones do.

    Token-based global rate limiting through Envoy Gateway's `BackendTrafficPolicy`, keyed on `x-ai-eg-model` plus client headers, with CEL cost expressions; requires a Redis deployment ([usage-based rate limiting](https://aigateway.envoyproxy.io/docs/capabilities/traffic/usage-based-ratelimiting)).

  • Model allowlists Yes

    A policy that constrains which models a key or team can call, keeping expensive frontier models out of the wrong hands.

    Effectively by routing: only models declared in an `AIGatewayRoute` rule are reachable and `GET /v1/models` returns exactly that set ([supported endpoints](https://aigateway.envoyproxy.io/docs/capabilities/llm-integrations/supported-endpoints)).

  • Spend alerts Not published

    Alerts fired as spend approaches a threshold. Alerts that only fire after the meter has rolled over are marked as such.

    Not documented as a product feature; you would alert on the exported Prometheus token metrics yourself ([metrics](https://aigateway.envoyproxy.io/docs/capabilities/observability/metrics)).

  • Webhook notifications Not published

    Programmatic notifications on spend events, so budget breaches can page an on-call or open a ticket.

    Not documented on any page fetched ([quota policy](https://aigateway.envoyproxy.io/docs/capabilities/traffic/quota-policy); [metrics](https://aigateway.envoyproxy.io/docs/capabilities/observability/metrics)).

Enforcement: Enforced before each request

Source →

Catalog

Models available Models available
Not published
How many models you can call, shown as a range because several vendors publish different totals on different pages. Vendor-reported either way, so counts are not directly comparable — some count every provider variant of the same model separately.
Model providers reachable Upstream providers
16–19
How many distinct model providers or labs you can reach, shown as a range where the vendor’s own pages disagree. More providers usually means better redundancy when one has an outage.
Works with standard OpenAI code OpenAI-compatible API
Yes
If yes, you can usually switch to it by changing one base URL, and switch away just as easily. This is the main defence against lock-in.
OpenAI chat endpoint POST /v1/chat/completions
Yes
The endpoint almost every application ports first. "Not documented" means the vendor never states it, which is different from a documented no.
Anthropic messages endpoint POST /v1/messages
Yes
Whether Anthropic-shaped calls work without rewriting them. Several products support this only as SDK compatibility or provider passthrough rather than a native endpoint — the detail page says which.
OpenAI Responses endpoint POST /v1/responses
Yes
The newer stateful OpenAI surface. Support is much thinner across this market than chat completions.
Embeddings endpoint POST /v1/embeddings
Yes
Whether you can generate vectors through the same gateway, or need a second integration for retrieval workloads.
Image generation endpoint POST /v1/images/generations
Yes
Whether image models are reachable through the same surface as text.
Audio endpoints POST /v1/audio/*
Yes
Speech-to-text and text-to-speech. Frequently the first gap in an otherwise complete gateway.
Batch jobs endpoint POST /v1/batches
Not documented
Asynchronous bulk processing, usually at a discount. Commonly undocumented, and commonly the reason a migration stalls late.
Needs the vendor’s own code library Requires a vendor-specific SDK
No
A proprietary client library spreads through your codebase and has to be torn out again if you leave. “No” is the better answer here, and it means the standard OpenAI client works.
You can export your request history Logs / usage data export
Yes
Whether you can get your own request logs, traces, or usage records back out — through an API, a bulk export, or a download. Decides whether you leave with your history or abandon it.
Settings can live in version control Declarative config-as-code
Yes
Whether routing, fallback, and budget rules can be declared in a file you keep in Git, rather than existing only as settings clicked into a hosted dashboard.
Embeddings Embeddings
Yes
Text-to-vector models, needed for search and retrieval features.
Image generation Image generation
Yes
Whether image models are reachable through the same interface.
Speech and audio Speech and audio
Yes
Text-to-speech or transcription models through the same interface.
Video generation Video generation
Yes
Whether video models are reachable through the same interface.
Batch processing Batch processing
Not published
Submitting large jobs for cheaper, slower processing. Often 50% off for work that is not time-sensitive.

Routing & reliability

Uptime it promises in writing Contractual SLA uptime
Not published
The uptime percentage in a published, contractual service level agreement. A public status page is not an SLA — it reports what happened, it does not promise anything or pay you back when it breaks.
Automatic failover Automatic failover
Yes
When a model provider goes down or rate-limits you, traffic moves to a backup automatically instead of returning errors to your users.
Load balancing Load balancing
Yes
Spreads requests across several providers or keys to raise your effective rate limit.
Rule-based routing Conditional routing
Yes
Send different requests to different models based on rules — for example a cheap model for free users and a strong model for paying ones.
Response caching Response caching
Not published
Reuses the answer when the exact same request comes in again, which cuts both cost and latency.
Similar-question caching Semantic cache
Not published
Reuses an answer when a new question means roughly the same thing as an earlier one. Saves far more than exact-match caching but can return subtly wrong answers if tuned loosely.
Where you set the timeout Request timeout surface
In config

Config-file only, and split across two APIs. Request timeouts come from Envoy Gateway's `BackendTrafficPolicy` (the fallback example sets `timeout: 30s` with `perRetry.backOff.baseInterval 100ms` and `maxInterval 10s`) ([provider fallback](https://aigateway.envoyproxy.io/docs/capabilities/traffic/provider-fallback)); v1.1 added an AI-specific `AIGatewayRouteRule.streamIdleTimeout` that triggers failover if it fires before the first token and returns 504 if it fires mid-stream ([v1.1 release notes](https://aigateway.envoyproxy.io/release-notes/v1.1), 2026-09-02). No default value for either is stated on the pages fetched.

Where a request timeout can be set: per request, in a config file, in the vendor dashboard, or nowhere. Nine of the twenty products documented here do not describe a request timeout at all, so the worst case of a hung upstream call is unknowable from the docs.
Where you set retries Retry policy surface
In config

Retries are Envoy Gateway's, configured in YAML next to the AI routes: the documented example sets `numRetries: 5`, `numAttemptsPerPriority: 1`, exponential backoff `100ms`/`10s`, `retryOn.httpStatusCodes: [500]` and triggers `connect-failure` and `retriable-status-codes`, with `numAttemptsPerPriority` controlling how many tries each priority group gets before failing over ([provider fallback](https://aigateway.envoyproxy.io/docs/capabilities/traffic/provider-fallback), 2026-09-02).

Where retry count and backoff are configured. Worth knowing alongside billing: a retried streaming call can be charged more than once.
Where you set fallbacks Fallback surface
In config

Ordered priority groups: `backendRefs` in an `AIGatewayRoute` rule carry `priority: 0`, `priority: 1` and so on, and traffic moves to the next priority when the current one exhausts its attempts. Because the ext_proc filter chain is split into router-level and upstream-level filters, a failover to a different provider re-runs request translation and upstream auth for the new backend - which is what makes cross-provider fallback work rather than just cross-endpoint retry ([provider fallback](https://aigateway.envoyproxy.io/docs/capabilities/traffic/provider-fallback); [data plane](https://aigateway.envoyproxy.io/docs/concepts/architecture/data-plane), 2026-09-02). Weighted fallback is not documented.

Where the fallback chain is defined. Almost every product claims fallback; the useful question is whether you can change it from code or only by hand in a dashboard.
Shape of the fallback chain Fallback shape
Ordered list
An ordered list tries targets in sequence; a weighted split sends a percentage of traffic to each, which is what you need to trial a new model on 5% of requests. Weighted splits are much rarer than the marketing implies.
Upstream health tracking Health checks / circuit breaking
Not documented

n.a. - no active health check or circuit breaker for AI backends on any page fetched ([provider fallback](https://aigateway.envoyproxy.io/docs/capabilities/traffic/provider-fallback), [system architecture](https://aigateway.envoyproxy.io/docs/concepts/architecture/system-architecture), [data plane](https://aigateway.envoyproxy.io/docs/concepts/architecture/data-plane), [capabilities index](https://aigateway.envoyproxy.io/docs/)). Failure detection is reactive: retries and priority failover on connect failures and retriable status codes. The one exception is InferencePool, where an endpoint picker selects self-hosted endpoints from live metrics such as KV-cache usage and queue depth ([InferencePool support](https://aigateway.envoyproxy.io/docs/capabilities/inference/inferencepool-support), 2026-09-02).

Whether the product notices a failing upstream and stops sending traffic to it, and whether you can tune the thresholds. This is what turns a provider outage into a blip rather than a sustained error rate, and only four of the twenty expose it.
Cross-region failover you control Multi-region failover surface
Not documented

n.a. - no cross-region or multi-cluster failover construct is documented; the two-tier gateway pattern in the README is about tier-one ingress versus tier-two self-hosted model access, not geography ([README](https://raw.githubusercontent.com/envoyproxy/ai-gateway/main/README.md); [provider fallback](https://aigateway.envoyproxy.io/docs/capabilities/traffic/provider-fallback), 2026-09-02).

Whether you can define what happens when a region degrades. A vendor running many regions is not the same as a vendor letting you configure failover between them; only three document a user-controlled mechanism.
Where you set load balancing Load balancing surface
In config

Config-file. Multiple `backendRefs` at the same priority are load-balanced by Envoy, and Envoy Gateway's `BackendTrafficPolicy` governs the algorithm ([provider fallback](https://aigateway.envoyproxy.io/docs/capabilities/traffic/provider-fallback)); for self-hosted fleets an InferencePool endpoint picker does inference-aware selection on real-time KV-cache usage, queued requests and LoRA adapter state ([InferencePool support](https://aigateway.envoyproxy.io/docs/capabilities/inference/inferencepool-support), 2026-09-02).

Where traffic distribution across upstreams or keys is configured.

Operations

Usage dashboards and logs Observability
Yes
Built-in visibility into what was sent, what came back, what it cost, and how long it took.
Spending limits Budget controls
Yes
Hard caps that stop spend before it becomes a surprise invoice. The single most valuable control for a small team.
Rate limits Rate limits
Yes
Caps on request volume per key or per user, useful for protecting against abuse and runaway loops.
Separate keys per team or app Virtual keys
Not published
Issue scoped keys with their own budgets and permissions so you can attribute cost and revoke access without rotating everything.
Prompt versioning Prompt management
Not published
Store and version prompts outside your code so they can be changed without a deploy.
Quality testing Evals
Not published
Built-in tooling to score model output against test cases, so you can tell whether a model swap made things better or worse.
MCP support MCP support
Yes
Native support for the Model Context Protocol, the emerging standard for connecting models to external tools.
What gets logged Logged content
Your choice

Content capture is a switch, and the default depends on which semantic convention you pick. OpenInference (the default) records inputs and outputs unless you set `OPENINFERENCE_HIDE_INPUTS`, `OPENINFERENCE_HIDE_OUTPUTS`, `OPENINFERENCE_HIDE_EMBEDDINGS_TEXT` or `OPENINFERENCE_HIDE_EMBEDDINGS_VECTORS`; switching to `AI_GATEWAY_TRACING_SEMCONV=gen_ai` makes content opt-in via `OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=true` ([tracing](https://aigateway.envoyproxy.io/docs/capabilities/observability/tracing), 2026-09-02).

Whether full prompts and responses are stored, only metadata, or nothing. Full-body logging is the most useful debugging feature here and the one most likely to need a conversation with your compliance team.
You can turn logging off Body-logging opt-out
Yes

Yes, and it is total: with no OTLP endpoint configured nothing is exported at all, and per-field redaction flags plus the `gen_ai` semantic-convention mode let you keep spans without message content ([tracing](https://aigateway.envoyproxy.io/docs/capabilities/observability/tracing), 2026-09-02).

Whether prompt and response bodies can be suppressed while still keeping usage metrics. Nineteen of the twenty document a way to do this; the mechanisms range from a per-request header to an organisation-wide setting.
Traces you can take elsewhere Distributed tracing
OpenTelemetry

Full OpenTelemetry: set `OTEL_EXPORTER_OTLP_ENDPOINT` (globally or per-Gateway through `GatewayConfig.spec.extProc.kubernetes.env`) and the ext_proc emits spans, defaulting to **OpenInference** conventions with `AI_GATEWAY_TRACING_SEMCONV=gen_ai` as the alternative; MCP `CallTool` and `ListTools` operations get spans too, and a `session.id` header mapping groups multi-turn conversations. Arize Phoenix is the documented consumer, giving LLM-as-judge evaluation over production spans ([tracing](https://aigateway.envoyproxy.io/docs/capabilities/observability/tracing); [gateway config](https://aigateway.envoyproxy.io/docs/capabilities/gateway-config), 2026-09-02).

Whether the product emits OpenTelemetry, a proprietary format, or nothing. OpenTelemetry means the traces land in the tooling you already run instead of only in the vendor’s dashboard.
Where telemetry can go Export destinations
OpenTelemetry (OTLP), Prometheus, Envoy access logs, Arize Phoenix, Grafana

OTLP for traces to any collector ([tracing](https://aigateway.envoyproxy.io/docs/capabilities/observability/tracing)); Prometheus scrape for GenAI token and latency metrics, with custom labels lifted from request headers ([metrics](https://aigateway.envoyproxy.io/docs/capabilities/observability/metrics)); Envoy access logs carrying `gen_ai.*` token costs and MCP fields such as `mcp_tool_name` and `mcp_backend`, configured on `EnvoyProxy.spec.telemetry.accessLog` ([access logs](https://aigateway.envoyproxy.io/docs/capabilities/observability/accesslogs)). Named downstreams in the docs are Arize Phoenix via its Helm chart and a bundled Grafana dashboard at `examples/monitoring/grafana-dashboard.json` ([tracing](https://aigateway.envoyproxy.io/docs/capabilities/observability/tracing); [v1.1 release notes](https://aigateway.envoyproxy.io/release-notes/v1.1), 2026-09-02).

Documented sinks for logs and metrics. This is a good proxy for how replaceable the vendor’s own dashboard is: around twenty destinations means you never have to depend on it, while a CSV download means you do.
Can record user feedback Feedback capture API
No

n.a. - no thumbs-up/score ingestion API or feedback header on any page fetched ([tracing](https://aigateway.envoyproxy.io/docs/capabilities/observability/tracing); [metrics](https://aigateway.envoyproxy.io/docs/capabilities/observability/metrics); [API reference](https://aigateway.envoyproxy.io/docs/api/), 2026-09-02).

Whether there is an API to attach a rating or score to a logged request, which is what lets production traffic feed quality work later.
Scores live traffic Online eval hooks
Partly

Partial and second-hand: the gateway emits OpenInference-compatible spans with full message content, and the docs walk you through deploying Arize Phoenix so "LLM-as-judge" evaluations can run against production spans - the evaluation itself belongs to Phoenix, not to the gateway, which ships no dataset, scorer or experiment surface ([tracing](https://aigateway.envoyproxy.io/docs/capabilities/observability/tracing), 2026-09-02).

Whether automated scorers can run against real production requests, rather than only against a test set you assemble yourself.

Performance

Delay it adds Proxy overhead
2 ms
Extra time the product itself adds to each request, on top of however long the model takes. Usually irrelevant next to multi-second model latency, but it matters for high-volume or streaming-sensitive workloads.
Requests per second ceiling Throughput
Not published
Published sustained request rate before the product becomes the bottleneck. Only relevant at genuinely high volume.
What the request path runs on Architecture class
Compiled binary

Go control plane and Go ext_proc filter alongside the Envoy C++ data plane: repo language bytes are Go 7,634,320, MDX 1,427,966, CSS 44,361 and TypeScript 43,527, built against Go 1.26.4 ([GitHub API](https://api.github.com/repos/envoyproxy/ai-gateway); [v1.1 release notes](https://aigateway.envoyproxy.io/release-notes/v1.1), 2026-09-02). Requests traverse Envoy Proxy plus an External Processor container that does model extraction, routing, upstream auth, request/response translation and token accounting, with a Rate Limit Service and Redis for token limits; the filter chain is deliberately split into router-level and upstream-level stages so a retry to another provider re-runs translation and auth. Dynamic Modules are noted as a possible future alternative to ext_proc ([system architecture](https://aigateway.envoyproxy.io/docs/concepts/architecture/system-architecture); [data plane](https://aigateway.envoyproxy.io/docs/concepts/architecture/data-plane)).

The comparable way to talk about latency here. An edge worker, a compiled Go or Rust binary, and a Python proxy have different overhead floors no matter which figures each vendor publishes. Products that never disclose their runtime are recorded as undisclosed rather than assumed.
You can run the request path yourself Self-hostable data plane
Yes

Everything is published: Helm charts as OCI artifacts (`oci://docker.io/envoyproxy/ai-gateway-crds-helm` and `oci://docker.io/envoyproxy/ai-gateway-helm`, pinned in the docs to `--version v1.0.0`) ([installation](https://aigateway.envoyproxy.io/docs/getting-started/installation)); a CLI container image at `envoyproxy/ai-gateway-cli` and per-release `aigw` binaries for Linux and macOS on the GitHub releases page, or `go install ./cmd/aigw` from source ([CLI installation](https://aigateway.envoyproxy.io/docs/cli/aigwinstall), 2026-09-02); the v1.1.0 release itself is on GitHub ([releases API](https://api.github.com/repos/envoyproxy/ai-gateway/releases/latest)).

Whether the component that actually carries your prompts can run on your own infrastructure. Distinct from a vendor offering a self-hosted control plane while still proxying traffic through their network.
Streaming responses Streaming support
Yes

Yes - SSE streaming is listed as supported on chat completions and Anthropic messages, with translation preserved across providers including tool use and reasoning blocks ([supported endpoints](https://aigateway.envoyproxy.io/docs/capabilities/llm-integrations/supported-endpoints); [v1.0 release notes](https://aigateway.envoyproxy.io/release-notes/v1.0)). Two operational caveats the docs raise: Envoy Gateway's default 32 KB buffer limit "is not enough for most AI model responses", so the example manifest raises it to 50 MB via `ClientTrafficPolicy` ([basic usage](https://aigateway.envoyproxy.io/docs/getting-started/basic-usage)); and `streamIdleTimeout` returns 504 if it fires mid-stream, while firing before the first token triggers failover instead ([v1.1 release notes](https://aigateway.envoyproxy.io/release-notes/v1.1), 2026-09-02).

Whether token-by-token streaming is documented. The caveats matter more than the yes: some products cannot cancel a stream without still being billed, and several timeout and fallback mechanisms stop applying once the first token has been sent.

Security & compliance

Does your prompt reach their servers Prompt transits vendor
No

Prompts never reach a project-operated service. Traffic terminates in Envoy proxies and an ext_proc container inside your own cluster, and the control plane only reconciles CRDs and pushes xDS config ([system architecture](https://aigateway.envoyproxy.io/docs/concepts/architecture/system-architecture); [data plane](https://aigateway.envoyproxy.io/docs/concepts/architecture/data-plane), 2026-09-02). The only third parties on the path are the upstream providers you configure.

Whether the text you send passes through this company’s own infrastructure. If it does, every other promise on this page is a policy commitment rather than a physical impossibility. Self-hosted products can answer no outright.
What they keep if you change nothing Logging default
Metadata only, not content

Out of the box you get Prometheus metrics following OpenTelemetry GenAI semantic conventions - `gen_ai.client.token.usage`, `gen_ai.server.request.duration`, `gen_ai.server.time_to_first_token`, `gen_ai.server.time_per_output_token` - and Envoy access logs carrying token-cost dynamic metadata; that is metadata, not content ([metrics](https://aigateway.envoyproxy.io/docs/capabilities/observability/metrics); [access logs](https://aigateway.envoyproxy.io/docs/capabilities/observability/accesslogs), 2026-09-02). The load-bearing caveat is one level up: as soon as you set `OTEL_EXPORTER_OTLP_ENDPOINT`, tracing defaults to OpenInference conventions **with full request and response content included by default**, and every `OPENINFERENCE_HIDE_*` redaction flag defaults to false ([tracing](https://aigateway.envoyproxy.io/docs/capabilities/observability/tracing)).

Defaults matter more than options. A product that stores full prompts and replies unless you find the right header will have stored them by the time you read the docs.
How long they keep it Default content retention (days)
Not published

Retention is entirely yours - the project stores nothing and ships no datastore for request data. Metrics, traces and access logs land in the Prometheus, OTLP collector and log sinks you operate ([metrics](https://aigateway.envoyproxy.io/docs/capabilities/observability/metrics); [tracing](https://aigateway.envoyproxy.io/docs/capabilities/observability/tracing); [access logs](https://aigateway.envoyproxy.io/docs/capabilities/observability/accesslogs)).

Default retention for request content, in days. Zero means nothing is kept. Read the note: several products keep nothing as a rule but make timed exceptions for abuse review or specific models.
Could they train on your prompts Training on customer data
Not applicable

No vendor ever receives prompts, so there is nothing to train on; no training policy is published, and none would be meaningful for self-hosted Apache-2.0 software ([system architecture](https://aigateway.envoyproxy.io/docs/concepts/architecture/system-architecture)).

Whether the vendor may use your prompts and outputs to train models. “Not published” means we could not find any position, which is not the same as a no — ask for it in writing.
Where it runs, and what you can pin Region and residency control
Region is whatever your Kubernetes cluster and chosen upstreams are - there is no vendor region list. Kubernetes v1.32+ with Envoy Gateway v1.7.0+ is the only placement requirement stated ([prerequisites](https://aigateway.envoyproxy.io/docs/getting-started/prerequisites); [compatibility](https://aigateway.envoyproxy.io/docs/compatibility), 2026-09-02).
Which regions are offered and whether you can force processing to stay in one. A global endpoint that silently picks a region is a different compliance story from an endpoint you pin yourself.
Where safety filters run Guardrail execution location
No guardrails offered

There is no guardrail feature. The security documentation covers client-to-gateway access control by delegating to Envoy Gateway SecurityPolicy - JWT validation, JWT claim-based authorization, mTLS, external authorization, OIDC, basic auth, API key, IP allow/deny - and gateway-to-provider credential handling; content inspection is absent ([security](https://aigateway.envoyproxy.io/docs/capabilities/security/); [upstream auth](https://aigateway.envoyproxy.io/docs/capabilities/security/upstream-auth), 2026-09-02). No page in the site map carries the word guardrail, and neither the [v1.0](https://aigateway.envoyproxy.io/release-notes/v1.0) nor [v1.1](https://aigateway.envoyproxy.io/release-notes/v1.1) release notes introduce one.

A filter that strips personal data only helps if it runs before the data leaves your boundary. If guardrails execute in the vendor’s cloud, the vendor has already received whatever you wanted redacted.
Who else touches the data Subprocessor list
Not published
The published list of third parties the vendor passes your data to. No list means you cannot know the full chain, which most data-protection agreements require you to.
SOC 2 audited SOC 2 audited
Not published
An independent audit of security controls. Enterprise buyers and their procurement teams routinely require it.
Will sign a HIPAA agreement HIPAA BAA
Not published
Required before you may send protected health information through the service. Without a signed BAA, healthcare data is off limits.
GDPR commitments GDPR commitments
Not published
Published data processing terms for handling personal data of people in the EU and UK.
Can keep data in the EU EU data residency
Not published
Requests can be processed inside the EU rather than routed to US infrastructure. Often the deciding constraint for European customers.
Does not retain your data Zero data retention
Not applicable
Prompts and responses are not stored after the request completes. Sometimes a paid add-on rather than the default.
Strips personal data PII redaction
Not published
Detects and removes identifiers such as names, emails, and card numbers before the request reaches the model provider.
Content guardrails Content guardrails
Not published
Policy checks on inputs and outputs — blocking unsafe content, enforcing formats, or catching prompt-injection attempts.
Runs fully disconnected Air-gapped deployment
Not published
Can be deployed in a network with no internet access, which some regulated and defence environments require.
Blocks personal data in prompts PII / DLP enforcement
Not documented

n.a. - no PII detection, masking or redaction of request content on any page fetched ([security](https://aigateway.envoyproxy.io/docs/capabilities/security/); [header and body mutations](https://aigateway.envoyproxy.io/docs/capabilities/traffic/header-body-mutations); [v1.1 release notes](https://aigateway.envoyproxy.io/release-notes/v1.1)). The nearest thing is trace-level redaction: `OPENINFERENCE_HIDE_INPUTS` / `HIDE_OUTPUTS` / `HIDE_EMBEDDINGS_TEXT` / `HIDE_EMBEDDINGS_VECTORS` suppress prompt content in spans, all defaulting to false ([tracing](https://aigateway.envoyproxy.io/docs/capabilities/observability/tracing), 2026-09-02) - that protects your telemetry, not the model call.

Whether personal data detection sits on the request path and can stop the call, merely inspects and forwards it, or is not documented. A control that only reports is a logging feature, not a policy control.
Blocks prompt injection Injection / jailbreak enforcement
Not documented

n.a. - no prompt-injection or jailbreak detection documented ([security](https://aigateway.envoyproxy.io/docs/capabilities/security/); [capabilities index](https://aigateway.envoyproxy.io/docs/); [v1.1 release notes](https://aigateway.envoyproxy.io/release-notes/v1.1), 2026-09-02).

Whether injection and jailbreak detection can stop a request. Most products offering this call a partner classifier rather than shipping their own.
Blocks harmful content Toxicity / moderation enforcement
Not documented

n.a. - no moderation or toxicity filtering; nothing on the [security](https://aigateway.envoyproxy.io/docs/capabilities/security/) page or in the [v1.0](https://aigateway.envoyproxy.io/release-notes/v1.0) and [v1.1](https://aigateway.envoyproxy.io/release-notes/v1.1) release notes inspects prompt or completion content.

Whether hate, violence, sexual and self-harm categories are checked inline and can stop a request, in either direction.
Your own policy rules Custom policy hooks
Not documented

n.a. as a policy-driven content check. You can rewrite headers and bodies (`headerMutation`, `bodyMutation` on `AIServiceBackend`) and enforce arbitrary CEL expressions for MCP tool authorization ([header and body mutations](https://aigateway.envoyproxy.io/docs/capabilities/traffic/header-body-mutations); [MCP](https://aigateway.envoyproxy.io/docs/capabilities/mcp/)), and Envoy Gateway's external-authorization hook can front the gateway ([security](https://aigateway.envoyproxy.io/docs/capabilities/security/)) - but no custom guardrail evaluator is documented.

Whether you can add your own rule — a regex, a webhook, or your own classifier — rather than choosing from the vendor library.
Where guardrails run Guardrail execution location
Not documented
Whether guardrail evaluation happens inside your infrastructure or on the vendor’s servers. This decides whether the prompt you are trying to protect leaves your network in order to be checked.
If the guardrail itself fails Guardrail failure mode
Not documented

n.a. - no guardrail exists, so no fail-open/fail-closed behaviour is stated ([security](https://aigateway.envoyproxy.io/docs/capabilities/security/)). For the request path proper, the documented failure behaviour is Envoy retry and priority failover plus 429 on quota exhaustion ([provider fallback](https://aigateway.envoyproxy.io/docs/capabilities/traffic/provider-fallback); [quota policy](https://aigateway.envoyproxy.io/docs/capabilities/traffic/quota-policy)).

What happens when the guardrail service times out or errors: does the request proceed unchecked, or is it blocked? This is the worst-documented field in the entire catalogue — only two vendors state it plainly, which means most teams are running a control whose failure behaviour they cannot know.
Third-party guardrail vendors Guardrail integrations
Not published
Named external guardrail services the product can call. A long list means the product is a router for policy engines rather than a policy engine itself — which also means another vendor bill and another hop.

Compliance evidence

Graded by how strong the evidence is, not whether the word appears on the vendor’s website. An audited report and a marketing claim are different things, and only one of them will satisfy your own auditor.

  • SOC 2 Not published No SOC 2 statement on any page fetched; the project is Apache-2.0 software you run yourself and publishes no compliance attestations ([security](https://aigateway.envoyproxy.io/docs/capabilities/security/), [SECURITY.md](https://raw.githubusercontent.com/envoyproxy/ai-gateway/main/SECURITY.md)).
  • ISO 27001 Not published
  • GDPR DPA Not published No DPA or controller/processor language exists because the project never receives your data; compliance sits with the operator.
  • HIPAA BAA Not published No BAA is possible - there is no vendor service to sign one with.
  • FedRAMP Not published
  • ITAR Not published

No compliance certifications were found published for this product. That is not the same as failing an audit — it means there is nothing public to check, so ask for evidence directly.

Vendor source

Fit & integration

Work to try it Evaluation work shape
Run something locally first
The shape of the work on the vendor’s own quickstart, from swapping one base URL through to deploying infrastructure. An ordinal class rather than a duration, because elapsed time depends on accounts and quota we cannot see.
Work to run it Production work shape
Deploy it on your infrastructure
The same scale applied to the vendor’s recommended production path. For several products this is much heavier than the quickstart, which is exactly why both are recorded.
Steps on the quickstart Numbered quickstart steps
4

The getting-started guide numbers four sections (Prerequisites, Installation, Basic Usage, Connect Providers) rather than four commands; the Installation page itself has two numbered Helm steps and Prerequisites adds the Envoy Gateway install, so there is no single steps-to-first-call figure ([getting started](https://aigateway.envoyproxy.io/docs/getting-started/); [installation](https://aigateway.envoyproxy.io/docs/getting-started/installation), 2026-09-02).

A literal count of numbered steps on the vendor’s quickstart, recorded as evidence beside the work shape. Zero means the page publishes no numbered procedure at all. Large counts usually mean interleaved language tracks rather than more work.
Can you self-host it today Self-host install documentation
Install command published

`helm upgrade -i aieg-crd oci://docker.io/envoyproxy/ai-gateway-crds-helm --version v1.0.0 --namespace envoy-ai-gateway-system --create-namespace` then `helm upgrade -i aieg oci://docker.io/envoyproxy/ai-gateway-helm --version v1.0.0 --namespace envoy-ai-gateway-system --create-namespace`, after Envoy Gateway v1.7.0+; locally `docker run --rm -p 1975:1975 -e OPENAI_API_KEY=... envoyproxy/ai-gateway-cli run` or `go install ./cmd/aigw`

Whether an install command is actually published. Several products advertise self-hosting while publishing no command to start from, which a plain yes/no would hide.
Works with the OpenAI SDK OpenAI SDK drop-in
Yes

Yes: OpenAI paths are served at the gateway root, the quickstart is a plain curl to `$GATEWAY_URL/v1/chat/completions`, and `aigw run` intentionally consumes OpenAI SDK environment variables so an existing client only needs its base URL changed ([supported endpoints](https://aigateway.envoyproxy.io/docs/capabilities/llm-integrations/supported-endpoints); [basic usage](https://aigateway.envoyproxy.io/docs/getting-started/basic-usage); [aigw run](https://aigateway.envoyproxy.io/docs/cli/aigwrun), 2026-09-02).

Whether an existing OpenAI-compatible client can be pointed at it by changing the base URL and key.
Vercel AI SDK support AI SDK provider package
Via the OpenAI provider

No Vercel AI SDK provider package or integration page exists - the term appears nowhere on the pages fetched or in the site map. What is documented is generic OpenAI compatibility at the gateway root, which is what an `@ai-sdk/openai` `baseURL` override would target ([supported endpoints](https://aigateway.envoyproxy.io/docs/capabilities/llm-integrations/supported-endpoints); [aigw run](https://aigateway.envoyproxy.io/docs/cli/aigwrun), 2026-09-02).

Whether a first-party AI SDK provider package exists, or only a community package, a documented workaround, or the generic OpenAI provider pointed at a custom base URL.
Python framework integrations Documented Python frameworks
Not published

n.a. - no LangChain, LlamaIndex, DSPy or other framework integration is documented; the 469-URL site map contains no framework page and none of the docs pages fetched name one ([capabilities index](https://aigateway.envoyproxy.io/docs/); [supported endpoints](https://aigateway.envoyproxy.io/docs/capabilities/llm-integrations/supported-endpoints), 2026-09-02). In practice you would use each framework's OpenAI-compatible client with the gateway's base URL.

LangChain, LangGraph, LlamaIndex and similar orchestration frameworks with a documented integration.
Callable from Cloudflare Workers Cloudflare Workers support
Not documented

n.a. - no Cloudflare Workers or edge-runtime story; the data plane is Envoy in your cluster or the `aigw` binary on Linux/macOS ([data plane](https://aigateway.envoyproxy.io/docs/concepts/architecture/data-plane); [aigw run](https://aigateway.envoyproxy.io/docs/cli/aigwrun), 2026-09-02). Cloudflare is not even listed as an upstream provider ([supported providers](https://aigateway.envoyproxy.io/docs/capabilities/llm-integrations/supported-providers)).

Whether the docs show calling this product from your own Worker. Deliberately separated from the several products whose own gateway runs on Workers, which is a fact about their infrastructure and not about your edge compatibility.
Kubernetes install Helm chart availability
Official Helm chart

Kubernetes is the primary target, not a deployment option: two official OCI Helm charts (CRDs and controller), a documented CRD-ownership migration with `--take-ownership`, controller leader election with a horizontally scalable read-only extension server, an HPA example (70% CPU, min 2 / max 10), and v1.1 Helm additions for PodDisruptionBudget, topology spread constraints, pod labels and a restricted controller security context ([installation](https://aigateway.envoyproxy.io/docs/getting-started/installation); [scaling](https://aigateway.envoyproxy.io/docs/capabilities/scaling); [v1.1 release notes](https://aigateway.envoyproxy.io/release-notes/v1.1), 2026-09-02).

Whether a named, published Helm chart exists, versus Helm being referenced with no chart named, versus generic cluster documentation that has nothing to do with this product.
Terraform support Terraform provider or modules
Not documented

n.a. - no Terraform provider, module or example is published: nothing in the repository root or the 469-URL site map mentions Terraform, and the documented install path is Helm ([installation](https://aigateway.envoyproxy.io/docs/getting-started/installation); [GitHub repo](https://github.com/envoyproxy/ai-gateway), 2026-09-02).

Whether you can declare this in code: an official provider, official modules, resources inside a hyperscaler’s provider, a community provider, or only Terraform code shipped in a repo.
Reuses your cloud identity Cloud IAM reuse
Reuses your cloud identity

Native cloud identity on all three hyperscalers: AWS Bedrock through the default credential chain including EKS Pod Identity and IRSA (region only, no key material) or OIDC-to-STS temporary credentials; Azure OpenAI through Entra ID short-lived tokens; GCP Vertex AI through Application Default Credentials, service-account keys or Workload Identity Federation with Google STS ([connect providers](https://aigateway.envoyproxy.io/docs/capabilities/llm-integrations/connect-providers); [upstream auth](https://aigateway.envoyproxy.io/docs/capabilities/security/upstream-auth); [supported providers](https://aigateway.envoyproxy.io/docs/capabilities/llm-integrations/supported-providers), 2026-09-02).

Whether you can authenticate with IAM roles, workload identity or managed identities instead of another long-lived API key. Distinguished from products that only accept static upstream provider credentials.
Fits behind your API gateway API gateway integration
It is the API gateway

It is the gateway - but uniquely among these entries it is not a standalone one: Envoy Gateway is a hard prerequisite (v1.7.0+ for install, v1.8.1+ in the v1.0.x compatibility matrix) and AI Gateway is an extension server plus ext_proc filter on top of it, described as "an additive layer" ([prerequisites](https://aigateway.envoyproxy.io/docs/getting-started/prerequisites); [compatibility](https://aigateway.envoyproxy.io/docs/compatibility); [GOALS.md](https://raw.githubusercontent.com/envoyproxy/ai-gateway/main/GOALS.md), 2026-09-02).

Whether AI traffic can go through a gateway you already run, and who documented that — the product’s vendor or the gateway’s.
MCP support MCP surface shape
MCP gateway or registry

A genuine MCP gateway, configured with an `MCPRoute` CRD: multiple MCP servers multiplexed behind one endpoint, tool names namespaced by backend (`github__issue_read`), tool filtering by name or regex, OAuth 2.0 authorization-code flow with PKCE plus JWT scope/claim and CEL-based authorization with a configurable `defaultAction`, header forwarding, upstream API-key injection, streamable HTTP per the June 2025 MCP spec, and OTel spans plus Prometheus metrics for MCP traffic ([MCP](https://aigateway.envoyproxy.io/docs/capabilities/mcp/), 2026-09-02). The same thing runs locally via `aigw run --mcp-config` on `http://localhost:1975/mcp` ([aigw run](https://aigateway.envoyproxy.io/docs/cli/aigwrun)). A dedicated `MCPBackend` CRD is on the post-1.0 roadmap ([1.0 announcement](https://aigateway.envoyproxy.io/blog/v1.0-release-announcement)).

Which kind of MCP support this is: a gateway that governs many MCP servers, a hosted MCP server you connect to, MCP tools accepted inside the completion API, or client tooling. These are different products behind one acronym.
Needs your own provider key Upstream provider key required
Required

Yes for any real traffic: each `AIServiceBackend` needs a `BackendSecurityPolicy` with your provider credential ([connect providers](https://aigateway.envoyproxy.io/docs/capabilities/llm-integrations/connect-providers)). The exception is evaluation - the quickstart's `basic.yaml` routes to a mock backend (`some-cool-self-hosted-model`) so you can make a first call with no provider account at all ([basic usage](https://aigateway.envoyproxy.io/docs/getting-started/basic-usage), 2026-09-02).

Whether an upstream provider account and key must exist before your first call works. A prerequisite rather than a step, and it can differ between the hosted and self-hosted forms of the same product.
Gate before models work Model access gate
No gate

None: no approval, waitlist, enablement or quota step exists on the project side - whatever your provider credential can reach is reachable once you declare it in a route ([connect providers](https://aigateway.envoyproxy.io/docs/capabilities/llm-integrations/connect-providers); [supported endpoints](https://aigateway.envoyproxy.io/docs/capabilities/llm-integrations/supported-endpoints), 2026-09-02). Any gating you experience is your upstream provider's.

Whether an enablement click, a quota grant, a paid tier or an approval form stands between a valid key and a working model call.
Official client languages First-party client SDK languages
Not published

No client library is published or required. The documented pattern is pointing an OpenAI-compatible client at the gateway; `aigw run` reads the OpenAI SDK's environment variables so existing code needs no change, and the InferencePool page cites "seamless integration with OpenAI SDKs" ([aigw run](https://aigateway.envoyproxy.io/docs/cli/aigwrun); [InferencePool support](https://aigateway.envoyproxy.io/docs/capabilities/inference/inferencepool-support), 2026-09-02). No per-language SDK list, snippet set or package name appears on the pages fetched, so `fit_client_sdk_langs` is left empty rather than inferred from the curl examples.

Languages with a first-party client library. An empty list can still mean the product is usable from any language via an OpenAI-compatible SDK.

How pricing actually works

Self-hosting is the only mode and carries no licence fee: `helm upgrade -i aieg-crd oci://docker.io/envoyproxy/ai-gateway-crds-helm` followed by `helm upgrade -i aieg oci://docker.io/envoyproxy/ai-gateway-helm` ([installation](https://aigateway.envoyproxy.io/docs/getting-started/installation)). Real cost drivers are the Envoy Gateway control plane and Envoy data plane it requires, the AI Gateway controller (documented starting requests 100m CPU / 128Mi, limits 500m / 256Mi, HPA example min 2 max 10 replicas at 70% CPU) and a Redis instance if you enable token rate limiting or quotas ([scaling](https://aigateway.envoyproxy.io/docs/capabilities/scaling); [rate limiting](https://aigateway.envoyproxy.io/docs/capabilities/traffic/usage-based-ratelimiting), 2026-09-02).

Back to top ↑

Common questions

Answered from the fields above, so these move when the catalog moves. Every figure quoted here appears in the specification with its source.

Does Envoy AI Gateway charge a markup on model prices?

Envoy AI Gateway adds no percentage markup to model prices. The cost estimator on this site itemises these mechanisms against your own volume, because which one is cheapest depends entirely on the numbers you put in. The cost estimator itemises this against your own volume.

Can Envoy AI Gateway be self-hosted?

Yes — self-hosting is the only way to run Envoy AI Gateway; there is no vendor-hosted option. The licence is Apache-2.0. Running it yourself means you supply the infrastructure and the upstream model accounts, so the bill is your own hosting plus the providers' own rates.

Is Envoy AI Gateway SOC 2 audited, and will it sign a HIPAA BAA?

Envoy AI Gateway publishes neither a SOC 2 report nor a HIPAA business associate agreement. Each of these is linked to the vendor's own page in the compliance section below. Neither absence means a refusal: both are things a vendor either publishes or does not, and smaller products often hold the certification without advertising it.

Does Envoy AI Gateway retain your prompts?

Zero data retention does not apply to Envoy AI Gateway: it runs inside your own infrastructure, so prompts never reach a vendor. Whether prompt and response bodies are logged is configurable. Logging can be turned off. Retention becomes your own configuration question instead, decided by whatever logging you switch on in your own deployment.

Can you use your own provider keys with Envoy AI Gateway?

Yes. Envoy AI Gateway can route through your own accounts with the underlying model providers, so inference is billed to you directly. No fee of any kind: the project takes no payment and no tokens pass through vendor infrastructure, so BYOK carries no surcharge ([LICENSE](https://raw.githubusercontent.com/envoyproxy/ai-gateway/main/LICENSE); [upstream auth](https://aigateway.envoyproxy.io/docs/capabilities/security/upstream-auth), 2026-09-02).

How many models does Envoy AI Gateway support?

Envoy AI Gateway publishes no total model count. It reaches 16–19 upstream providers. No model catalogue exists: the project ships no hosted model list and publishes no model total. `GET /v1/models` returns only the models an operator declared in their own `AIGatewayRoute` resources ([supported endpoints](https://aigateway.envoyproxy.io/docs/capabilities/llm-integrations/supported-endpoints), 2026-09-02), and model names are supplied per route via `modelNameOverride` ([model name virtualization](https://aigateway.envoyproxy.io/docs/capabilities/traffic/model-name-virtualization)).

Back to top ↑

What has changed here

  1. GitHub stars GitHub stars 1987 2112 source ↗
  2. catalog entry catalog entry Not published Added to the catalog source ↗
See this in the full changelog Back to top ↑

Read the head-to-head

These pairs have a written verdict, not just a table.

Usually weighed against

Follow updates about Envoy AI Gateway