Envoy AI Gateway vs Higress
The question that decides it: Do you need the wider plugin shelf, or upstream credentials your security review will accept and a control-plane API that will not move?
Our verdict
If you want one binary covering ingress, microservices and LLM traffic, or you need caching, guardrails and per-consumer quotas today, Higress — and budget an afternoon for re-reading plugin defaults before you trust them. If your provider credentials must come from cloud identity rather than static keys in plugin config, or you already run Envoy Gateway and want a v1beta1 API you can build on, Envoy AI Gateway.
Why
This is the one genuinely close pair in this corner of the catalogue: same Apache-2.0 licence, same Envoy data plane, same Kubernetes-first target, no vendor in either request path, and zero token markup on both sides. The split starts with what each one demands of your cluster. Envoy AI Gateway is not standalone — Envoy Gateway v1.7.0+ is a hard prerequisite to install and v1.8.1+ with Envoy Proxy v1.38.x in the v1.0.x compatibility matrix, on Kubernetes v1.32+ with Gateway API v1.5.x — and the project describes itself as an additive layer on that stack, with an aigw run single-process mode on localhost:1975 for local testing rather than production. Higress runs the same Envoy data plane but takes almost nothing for granted: one all-in-one container serves console, HTTP and HTTPS on ports 8001, 8080 and 8443, a generated Docker Compose bundle covers non-Kubernetes hosts with configuration in local files or Nacos, and Helm installs a controller and a two-replica gateway into higress-system. If you are not already on Envoy Gateway, Higress is the one you can stand up this afternoon.
The feature surfaces are almost complementary. Higress has the plugin shelf: 31 documented provider types, a 100+ model claim, exact and semantic caching in one ai-cache plugin, ai-quota per-consumer token quotas in Redis, console-issued consumer credentials checked against an x-api-key header, ai-data-masking, inline moderation through ai-security-guard, ai-statistics across gateway, route, service and model, and a 41-plugin marketplace whose other entries handle WAF, CORS, JWT auth and DeGraphQL. Envoy AI Gateway has none of those: no content guardrails of any kind appear in its documentation, its only caching is passthrough of provider cache_control breakpoints, and it has no virtual keys for downstream callers because client authentication is delegated to Envoy Gateway SecurityPolicy. Where it leads is the endpoint surface: 19 provider configurations behind chat completions, completions, embeddings, image generations, audio transcriptions and translations, Responses, Cohere v2 rerank, a vLLM-compatible tokenize endpoint and native Anthropic Messages, against Higress's documented chat completions, embeddings and Anthropic Messages with no audio, image or Responses path at all.
For anyone with a security review, though, credential handling is the sharper deciding question, and it does not favour the longer feature list. Envoy AI Gateway centralises upstream credentials in a BackendSecurityPolicy and mints short-lived tokens per request from cloud identity on all three hyperscalers: AWS Bedrock through the default credential chain including EKS Pod Identity and IRSA with only a region set, or OIDC to STS; Azure OpenAI through Entra ID; Google Vertex AI through Application Default Credentials or Workload Identity Federation with Google STS. Higress takes cloud credentials as static plugin fields — awsAccessKey, awsSecretKey and awsRegion for Bedrock, vertexAuthKey, vertexRegion and vertexProjectId for Vertex AI — with no IRSA, workload identity, instance profile or credential-chain support documented, and its non-Kubernetes install relies on a 32-character data-encryption key that the docs say must be set for cluster deployment or a random one is generated. Long-lived keys in plugin configuration is a normal choice for a gateway of this lineage, but it is a materially different conversation with an auditor than per-request federated tokens.
Maturity cuts the other way, and neither side wins it cleanly. Higress has the larger community and the longer production history: 9,273 stars with 1,080 open issues, v2.2.4 released 13 August 2026, roughly two years inside Alibaba before the CNCF Sandbox, and it is a conformant Gateway API and Gateway API Inference Extension implementation with nginx-Ingress annotation compatibility. Envoy AI Gateway is smaller and newer at 1,987 stars and 286 open issues, but it reached v1.0.0 on 23 June 2026 with a committed-stable v1beta1 control-plane API and an explicit promise not to break it, followed by v1.1.0 on 21 August 2026, tracking Envoy Gateway at roughly every two to three months with maintainers drawn from Tetrate, Bloomberg, Tencent, Netflix and Nutanix. The clearest difference in intent is that Higress has a commercial edition behind it and Envoy AI Gateway does not have one from the project: Alibaba Cloud AI Gateway is recommended in Higress's own quick start for production without Kubernetes, and Alibaba's comparison page marks automatic fault detection and recovery as unsupported and multi-AZ deployment, rate-limit degradation, monitoring and alerting, and enterprise observability as build-it-yourself in the open-source edition. Neither publishes SOC 2 or a DPA, and neither can: if you need a counterparty on the paperwork, Kong AI Gateway is the comparison to be making instead.
Which one, concretely
Choose Envoy AI Gateway if
- Upstream credentials must come from cloud identity — EKS Pod Identity, IRSA, Entra ID or Google STS workload identity federation — rather than static keys in config
- You already run Envoy Gateway on Kubernetes and want the AI layer to reuse it
- You want a control-plane API committed to stability: v1beta1 CRDs, documented migrations, EOL two releases out
- You need audio, image generation, Responses or rerank endpoints, none of which Higress documents
Choose Higress if
- You need to run outside Kubernetes, or without an existing Envoy Gateway install
- You need caching, moderation, PII masking, per-consumer quotas or consumer credentials from the gateway itself
- You want 31 documented provider types rather than 19, plus a native passthrough mode
- You want one gateway for ingress, microservices and LLM traffic, with nginx-Ingress annotation compatibility as the migration path
What catches people out
- Envoy Gateway must be installed first using the project's own values file or nothing reconciles, and its default 32 KB client buffer is "not enough for most AI model responses" — which is why the basic example ships a ClientTrafficPolicy raising it to 50 MB.
- Envoy AI Gateway documents no health checks, circuit breaking or multi-region failover for AI backends; failure detection is reactive retries and priority failover, with the exception of InferencePool endpoint picking for self-hosted fleets.
- Higress's Helm defaults are conservative in ways that surprise people: enableIstioAPI, enableGatewayAPI and o11y.enabled are all false, the maximum supported Gateway API version is 1.4.0 on 2.2.x, and images ship from Alibaba Cloud registries in cn-hangzhou, us-west-1 and ap-southeast-7 with no EU mirror.
- Higress's guardrails are attached but inert until configured: ai-security-guard defaults checkRequest and checkResponse to false, denyCode defaults to 200 so a blocked call looks like a success, and the qwen3guard plugin is explicitly fail-open.
Side by side
Interpret these fields: How much does an LLM gateway lock you in? · LLM gateway compliance: SOC 2, HIPAA and evidence · How LLM gateway failover actually works
3 of 15 fields differ, marked with a dot. Every figure links to the vendor page it came from. Blank values read Not published rather than No — silence from a vendor is not a negative answer.
| Field | Envoy AI Gateway | Higress |
|---|---|---|
| Ease of leaving Derived score, higher is easier | 100/100 Easy to leave | 100/100 Easy to leave |
| What kind of product Category | Open source | Open source |
| Who runs it Deployment model | Self-host only | Managed or self-host |
| Licence Licence | Apache-2.0 | Apache-2.0 |
| Models available Models available | Not published | ~100 |
| Model providers reachable Upstream providers | 16–19 | 26–31 |
| Markup on model prices Token markup | None | None |
| GitHub stars GitHub stars | 2,112 | 9,458 |
| Content guardrails Content guardrails | Not published | Yes |
| Strips personal data PII redaction | Not published | Yes |
| Similar-question caching Semantic cache | Not published | Yes |
| Separate keys per team or app Virtual keys | Not published | Yes |
| Speech and audio Speech and audio | Yes | Not published |
| Image generation Image generation | Yes | Not published |
| MCP support MCP support | Yes | Yes |
| Spending limits Budget controls | Yes | Yes |
for Envoy AI Gateway and for Higress. Want more fields, or a third option in the mix? Open these two in the full comparison tool.
Common questions
Can I run either one without Kubernetes?
Higress, properly. It ships an all-in-one container on ports 8001, 8080 and 8443, and a Docker Compose bundle of apiserver, controller, pilot, gateway and console for non-Kubernetes hosts, with configuration in local files or Nacos. Envoy AI Gateway has aigw run, a single-process mode on localhost:1975 that reads OpenAI SDK environment variables, but it is documented for testing configuration and local development rather than as a production deployment shape.
Which one supports more providers?
Higress. Its Chinese ai-proxy page enumerates 31 provider types and the English one 26, and it claims 100+ models across three vendor pages without publishing an enumerated list. Envoy AI Gateway's supported-providers table lists 19 rows including self-hosted models, while its 1.0 announcement, release-note headlines and README logos all count 16. Neither publishes a real model catalogue: Envoy's /v1/models returns only what an operator declared in a route.
Which is better if I need guardrails or a cache?
Higress, with caveats. It has moderation through ai-security-guard, regex PII masking through ai-data-masking, and exact plus semantic caching through ai-cache. But the defaults are permissive — content checks off, denyCode 200, qwen3guard fail-open, cacheTTL 0 meaning never expire — so each needs deliberate configuration. Envoy AI Gateway has no guardrails and no gateway-side cache whatsoever, so it is not a candidate for that requirement.