agentgateway vs Envoy AI Gateway
The question that decides it: Are you adding LLM routing to an Envoy Gateway you already run, or do you want one standalone data plane that also carries agent protocols and content guardrails?
Our verdict
Already running Envoy Gateway on Kubernetes and routing LLM traffic: Envoy AI Gateway, because it is an additive layer on machinery you already operate and its control-plane CRDs are committed stable. Routing MCP or A2A agent traffic as well, or needing PII masking and moderation enforced inside the gateway: agentgateway, which does all of that in one Apache-2.0 binary.
Why
The shared ground here is unusually wide, so the differences matter more than usual. Both are Apache-2.0 and self-host only, with no hosted tier, no account, no seat fee and zero token markup. Both are BYOK-only by construction — agentgateway takes credentials inline, from an environment variable, from a file, from a Kubernetes secret or straight from the caller's token, and Envoy AI Gateway takes them through a BackendSecurityPolicy — and both mint short-lived upstream credentials from AWS, Azure and GCP identity rather than static keys. Both export OpenTelemetry traces and Prometheus metrics to collectors you run, and neither project ever sees a prompt. Both are foundation-governed: agentgateway was donated to the Linux Foundation in 2025 and became an Agentic AI Foundation hosted project in 2026, with 300+ contributors across 60+ organisations; Envoy AI Gateway lives in the envoyproxy GitHub organization with maintainers from Tetrate, Bloomberg, Tencent, Netflix and Nutanix.
The first real fork is what you must already be running. Envoy AI Gateway is explicitly "an additive layer" and Envoy Gateway is a hard prerequisite — v1.7.0+ to install, v1.8.1+ in the v1.0.x compatibility matrix, on Kubernetes v1.32+ with Gateway API v1.5.x — and token rate limiting or quotas additionally need Redis plus rate-limit configuration chosen at Envoy Gateway install time rather than afterwards. There is a local single-process mode (aigw run on localhost:1975) but it is documented for testing config and local development. agentgateway starts from the other end: one static Rust binary via curl -sL https://agentgateway.dev/install | bash or docker run cr.agentgateway.dev/agentgateway:v1.5.0, proxy on 4000 and UI, playground and config API on 15000, with a five-step LLM quickstart. Kubernetes mode with CRDs, a control plane, xDS and a GitOps workflow exists as a second shape rather than the only one.
The second fork is what the gateway is allowed to inspect. Envoy AI Gateway has no content guardrails whatsoever: no PII detection, no moderation, no injection detection and no custom evaluator appears anywhere in its documentation, and the nearest thing is trace-level redaction (OPENINFERENCE_HIDE_INPUTS, HIDE_OUTPUTS), which protects your telemetry rather than the model call. agentgateway ships guardrails that enforce by default: a regex guard with named creditCard, ssn, email and phoneNumber patterns whose default action is mask, plus OpenAI Moderation, Bedrock Guardrails, Google Model Armor, Azure Content Safety and your own webhook classifier, all defaulting to reject. Agent protocols split the same way but less starkly. Both are genuine MCP gateways — Envoy multiplexes servers behind one /mcp endpoint with tool-name prefixing, include and regex filtering, OAuth 2.0 authorization-code with PKCE and a CEL backendSelector defaulting to Deny; agentgateway adds MCP-specific guardrails, MCP authorization and tool-level access control, and supports MCP spec 2026-07-28 from v1.4. But A2A is a first-class protocol only on the agentgateway side; no A2A support is recorded for Envoy AI Gateway.
Where Envoy AI Gateway leads is ordinary endpoint breadth and API stability. It serves /v1/images/generations, /v1/audio/transcriptions and /v1/audio/translations, Cohere v2 rerank, Responses and, from v1.1, a vLLM-compatible /tokenize; agentgateway records images as not documented as a route type and audio only as a realtime WebSocket proxy whose documented examples are all text-only and which is exempt from prompt guards entirely. Envoy AI Gateway also shipped v1.0.0 in June 2026 with a committed-stable v1beta1 control-plane API and a 2-3 month cadence tracking Envoy Gateway, and it can front self-hosted fleets through a Gateway API InferencePool whose endpoint picker routes on live KV-cache usage, queue depth and LoRA adapter state. The catalogue numbers favour agentgateway — 4,691 GitHub stars against 1,987, a headline 1002+ models against no published model total, and 44+ providers on the cookbook (20 natively documented with a capability matrix) against a 19-row provider table headlined as 16 — but treat the performance figures carefully: agentgateway publishes p50 0.863 ms of overhead and 35,502 QPS from its own Fortio runs against a mock backend, while Envoy AI Gateway publishes no overhead figure of its own at all, the ~2 ms number coming from Tetrate summarising a Broadcom/VMware validation, and Tetrate co-maintains the project.
Which one, concretely
Choose agentgateway if
- You route MCP and A2A agent traffic as well as LLM traffic, and want the same data plane in front of ordinary HTTP and gRPC
- You need PII masking and moderation enforced in the gateway: regex guards default to mask, every external guard defaults to reject
- You want to start with one binary and no cluster — install script or Docker, proxy on 4000, UI and playground on 15000
- You want per-virtual-key USD and token budgets checked before the request is forwarded, plus a built-in cost dashboard that needs no Prometheus
Choose Envoy AI Gateway if
- You already run Envoy Gateway on Kubernetes and want AI routing as an additive layer rather than a second proxy
- You need image generation, audio transcription and translation, Cohere rerank or the /tokenize endpoint
- You want a control-plane API committed to stability — v1beta1 CRDs since the v1.0.0 GA in June 2026, with documented migrations
- You route to self-hosted model fleets and want InferencePool endpoint picking on live KV-cache usage, queue depth and LoRA adapter state
What catches people out
- Envoy AI Gateway has no content guardrails at all — no PII, moderation, injection or custom-evaluator surface appears in its docs. If you need content enforcement in the gateway, this side cannot give it to you.
- Envoy AI Gateway's prerequisites are load-bearing: Envoy Gateway v1.7.0+ on Kubernetes v1.32+, Redis and rate-limit configuration chosen at Envoy Gateway install time for token limits, and Envoy Gateway's default 32 KB client buffer raised to 50 MB because it "is not enough for most AI model responses". Its QuotaPolicy is also v1alpha1-only, outside the stability guarantee, with a serviceQuota field accepted but not enforced end-to-end.
- agentgateway's homepage markets semantic caching, but no semantic or response cache appears on any documentation page — the only caching feature is control over provider-side prompt-cache breakpoints. Both sides are passthrough caches only.
- agentgateway's config resource API is served on the admin address 127.0.0.1:15000 with no authentication, so network isolation is the only control; and
routing.failoveralone does not fail over — the docs warn you must also configurehealth.eviction, and the triggering request still fails without retries.
Side by side
Interpret these fields: How much does an LLM gateway lock you in? · LLM gateway compliance: SOC 2, HIPAA and evidence · How LLM gateway failover actually works
3 of 17 fields differ, marked with a dot. Every figure links to the vendor page it came from. Blank values read Not published rather than No — silence from a vendor is not a negative answer.
| Field | agentgateway | Envoy AI Gateway |
|---|---|---|
| Ease of leaving Derived score, higher is easier | 100/100 Easy to leave | 100/100 Easy to leave |
| What kind of product Category | Open source | Open source |
| Who runs it Deployment model | Self-host only | Self-host only |
| Licence Licence | Apache-2.0 | Apache-2.0 |
| Models available Models available | ~1,002 | Not published |
| Model providers reachable Upstream providers | 20–44 | 16–19 |
| Markup on model prices Token markup | None | None |
| Monthly cost per person Seat fee | None | Not published |
| Content guardrails Content guardrails | Yes | Not published |
| Strips personal data PII redaction | Yes | Not published |
| MCP support MCP support | Yes | Yes |
| Separate keys per team or app Virtual keys | Yes | Not published |
| Spending limits Budget controls | Yes | Yes |
| Speech and audio Speech and audio | Not published | Yes |
| Image generation Image generation | Not published | Yes |
| GitHub stars GitHub stars | 4,985 | 2,112 |
| Delay it adds Proxy overhead | 0.863 ms | 2 ms |
| Requests per second ceiling Throughput | 35,502 rps | Not published |
for agentgateway and for Envoy AI Gateway. Want more fields, or a third option in the mix? Open these two in the full comparison tool.
Common questions
Do I need Kubernetes for either of these?
For Envoy AI Gateway, effectively yes: Kubernetes v1.32+ with Envoy Gateway installed first is the documented target, and the local aigw run mode on localhost:1975 is presented for testing configuration and local development. agentgateway runs either way — a single binary or Docker container driven by a config file, or a Kubernetes mode with CRDs, a control plane, xDS and a GitOps workflow.
Which one has guardrails?
Only agentgateway. It ships a regex guard with named credit-card, SSN, email and phone patterns defaulting to mask, external guards through OpenAI Moderation, Bedrock Guardrails, Google Model Armor and Azure Content Safety defaulting to reject, a webhook guard for your own classifier, and a separate guardrail surface for MCP traffic. Envoy AI Gateway documents no content inspection of any kind.
Which one is faster?
The figures are not comparable. agentgateway publishes p50 0.863 ms of proxy overhead and 35,502 QPS from its own Fortio runs against a mock LLM backend, with scripts open on GitHub. Envoy AI Gateway publishes no data-plane overhead number of its own; the roughly 2 ms figure comes from Tetrate summarising a Broadcom/VMware validation, and Tetrate co-maintains the project. Load-test your own traffic.