Higress Open source
Higress is an open-source LLM gateway: an OpenAI-compatible API in front of ~100 models from 26–31 providers. It charges no token markup. It can be self-hosted under Apache-2.0 or used as a managed service. It does not publish a HIPAA BAA. You can point it at your own provider accounts. Beyond chat it also serves embeddings. It handles failover, load balancing, guardrails and request logging.
· 47 dated entries · 65 source references
Access at a glance
Whether this product can work for you at all, before features matter: who pays the model bill, where it can run, and which of your existing API calls keep working. Every value is the vendor’s own claim, linked to the page it came from.
General-purpose API gateway first, LLM router second: the docs describe Higress as "an open source AI-native API gateway" built on Istio and Envoy that covers AI gateway, Kubernetes Ingress and microservice-gateway use cases in one binary (what-is-Higress). The AI behaviour is delivered by Wasm plugins (ai-proxy, ai-cache, ai-token-ratelimit, ai-quota, ai-security-guard, ai-data-masking, ai-statistics) inside a 41-plugin marketplace whose other 24 plugins are ordinary gateway concerns such as WAF, CORS, JWT auth and DeGraphQL (plugin marketplace).
Who pays the model bill
Your keys onlyYou contract with each model provider directly and hold those accounts. The gateway never resells inference.
byok_only: every provider block in ai-proxy requires your own credentials — apiTokens for token-based providers, vertexAuthKey/vertexRegion/vertexProjectId for Google Vertex AI, awsAccessKey/awsSecretKey/awsRegion for AWS Bedrock (ai-proxy plugin). The installer prompts for an Alibaba Cloud DashScope or other API key at first run and the console's "LLM Provider Management" page is where the keys live (AI quick start). Higress sells no credits and hosts no models.
Merchant of record: The upstream model provider, always. Higress forwards requests using the apiTokens you configure per provider and never intermediates payment (ai-proxy plugin); if you buy the managed edition instead, Alibaba Cloud is the merchant (quick start).
Key handling: Provider keys are plugin configuration, not a managed secret store. apiTokens is a list and ai-proxy "randomly selects" one per request, with Azure OpenAI limited to a single token; failover can quarantine a token that returns errors until health checks recover it (ai-proxy, English, ai-proxy, Chinese). In the non-Kubernetes install, -k/--data-enc-key sets a 32-character key used to encrypt sensitive configuration data and "for cluster deployment, this must be set", otherwise a random key is generated (Docker Compose options). Console consumer credentials are separate from provider keys and are checked against the x-api-key header (token management guide). Hardware-backed KMS integration is a commercial-edition feature, listed as 深度集成阿里云 KMS 产品 versus 自行构建 for open source (Alibaba Cloud comparison).
Where it can run
2 of 5 shapes documented- Vendor-hosted
- Self-host
- Your VPC
- On-premise
- Air-gapped
Dimmed shapes are not documented by the vendor, which is not the same as unsupported.
Self-host is the primary mode, in three shapes: Helm on Kubernetes (Helm guide), Docker Compose or a shell installer without Kubernetes, with configuration in local files or Nacos (Docker Compose guide, quick start), and a single all-in-one container (README). The SaaS option is not a Higress service — it is Alibaba Cloud AI Gateway, recommended in Higress's own quick start for production without Kubernetes, in Serverless Standard, Serverless Enterprise and Feitian Exclusive tiers (quick start, AI gateway editions).
Two components: higress-controller (config aggregation and distribution) and higress-gateway (data plane, replicas default 2, deployable as Deployment or DaemonSet), plus a higress-console at 1 replica. Notable defaults: global.ingressClass higress, global.enableIstioAPI false, global.enableGatewayAPI false, global.o11y.enabled false; maximum supported Gateway API version is 1.4.0 on 2.2.x and 1.0.0 on 2.1.x and earlier (Helm guide). Images come from higress-registry.cn-hangzhou.cr.aliyuncs.com with us-west-1 and ap-southeast-7 mirrors selectable through global.hub (README). MCP hosting needs Higress >= 2.1.0, Redis for caching, and Nacos >= 3.0 with Higress >= 2.1.2 for the Nacos MCP registry (MCP quick start).
API surfaces your code can keep using
3 of 7 documented- OpenAI chat
POST /v1/chat/completionsYesThe ai-proxy plugin "implements AI proxy functionality based on OpenAI API contract" and auto-detects the OpenAI Chat Completions protocol when the request path is
/v1/chat/completions, converting to each upstream provider's own format (ai-proxy plugin). - Anthropic messages
POST /v1/messagesYesAi-proxy recognises
/v1/messagesas the Anthropic Claude Messages protocol and performs "intelligent conversion" to OpenAI format for providers that do not support Claude natively (ai-proxy plugin), observed 2026-09-02. - OpenAI Responses
POST /v1/responsesNot documented/v1/responsesappears on none of the fetched pages: ai-proxy (EN), ai-proxy (CN), plugin marketplace and the full documentation index at higress.ai/llms.txt. - Embeddings
POST /v1/embeddingsYesAi-proxy auto-detects
/v1/embeddingsas the OpenAI text-embedding protocol, with worked examples for Qwen and 360 Brain (ai-proxy plugin). Semantic caching also depends on calling an embedding service through the gateway (semantic cache guide). - Images
POST /v1/images/generationsNot documentedn.a. as a gateway path — no
/v1/images/generationsor image-generation surface on ai-proxy (EN), ai-proxy (CN) or plugin marketplace. The only multimodal content is a vision-input example headed "Multimodal Model API Request Example (Applicable to qwen-vl-plus and qwen-vl-max Models)" that still posts to chat/completions. - Audio
POST /v1/audio/*Not documentedNo
/v1/audio/*, TTS or STT path on ai-proxy (EN), ai-proxy (CN) or plugin marketplace. DeepL is supported as a provider but only for text translation. - Batch jobs
POST /v1/batchesNot documentedNo batch or async bulk endpoint on ai-proxy (EN), ai-proxy (CN) or the documentation index higress.ai/llms.txt.
An asterisk marks a qualified verdict: support that is indirect (SDK compatibility or provider passthrough rather than a native endpoint), or a gap that is narrower or wider than the label suggests. Read the note before porting.
Protocol is inferred from the request path rather than declared: ai-proxy auto-detects OpenAI /v1/chat/completions, Anthropic /v1/messages and OpenAI /v1/embeddings, and protocol: original lets you pass a provider's native wire format straight through instead of the OpenAI contract (ai-proxy plugin). On the north-south side the same binary serves the Kubernetes Ingress API, the Gateway API and the Gateway API Inference Extension (README), and hosts MCP servers over Streamable HTTP and SSE (MCP quick start). The console's own REST API is Swagger-documented but higress-console.swagger.enabled defaults to false (Helm values).
How much it reaches
Single figure, repeated verbatim on three vendor pages (100+); the "View Complete Model Integration Directory" link on the AI gateway page did not resolve to a fetchable list on 2026-09-02, and no /v1/models catalogue endpoint is documented, so the number could not be independently counted.
Counted from the provider.type tables, not from marketing. The Chinese ai-proxy page enumerates 31 named provider types — openai, azure, moonshot, qwen, baichuan, yi, zhipuai, deepseek, groq, grok, openrouter, fireworks, baidu, ai360, github, mistral, minimax, claude, ollama, hunyuan, stepfun, cloudflare, spark, gemini, deepl, cohere, together-ai, dify, vertex, bedrock, triton (ai-proxy, Chinese). The English page carries 26 provider-specific config sections and omits GitHub Models, Cohere and Dify (ai-proxy, English), hence the 26-31 range.
Whose models: All third-party or self-run. Higress hosts no models: ai-proxy forwards to external provider APIs using credentials you supply, and self-hosted backends are first-class via the ollama provider (ollamaServerHost/ollamaServerPort) and the NVIDIA Triton Inference Server example (ai-proxy plugin).
Your own endpoints: Yes, several ways: openaiCustomUrl points the OpenAI provider at any compatible service, azureServiceUrl at an Azure deployment, ollamaServerHost/ollamaServerPort at a local Ollama, and there is a worked NVIDIA Triton Inference Server configuration (ai-proxy plugin). Below the plugin, arbitrary upstreams are registered as console "Service Sources" of type Domains or Static Addresses, including non-LLM services such as Redis (token management guide).
How it behaves in production
What happens when an upstream model is slow, wrong, or down — and what you can see and stop while it happens. Reliability features are recorded as where you configure them, not whether the vendor lists them, because almost every product here lists all of them.
4 of 6 reachable from code 4 of 4 can block 6 documented destinations
When something goes wrong
Each row says where the knob is, not whether the feature is on the marketing page. A control you can only reach by hand in someone else’s dashboard cannot be reviewed or version-controlled.
- Request timeout In config
Changing it means editing configuration and shipping it, so behaviour is uniform across traffic until you redeploy.
config_file, but the LLM request timeout is the weak spot: ai-proxy'stimeout(default 120000 ms) is documented as applying only to the context-data retrieval call, not to the model request (ai-proxy plugin). The request timeout you actually want is the Ingress annotationhigress.io/timeout, in seconds, with no timeout by default (annotations). Connection-level values live in the ConfigMap: downstreamidleTimeout180 s, upstreamidleTimeout10 s (global configuration). MCP proxying has its ownserver.timeout, default 5000 ms (mcp-server plugin). - Retries In config
Off by default at the AI layer and capped at one extra attempt when switched on; no backoff strategy or jitter is documented anywhere (ai-proxy retryOnFailure). The generic Ingress retry annotation is separate and retries to a different upstream host rather than to a different model (annotations). No retry keys appear in the global ConfigMap (global configuration).
- Fallback to another model Dashboard only
Only reachable by hand in the vendor UI, so it cannot be reviewed, version-controlled, or changed from code.
Model-level fallback is a console concept, not a plugin field: the home page states Higress "supports model-level Fallback" (home page) and the multi-model proxy guide configures it in
AI Route Config, where a DeepSeek route falls back to Alibaba Cloudqwen-turbowhen the primary model fails or is rate-limited (multi-model proxy guide). The guide shows a single fallback target, not an ordered chain, and no YAML/CRD field for it appears in ai-proxy (EN) or ai-proxy (CN). Alibaba Cloud's own comparison confirms cross-model failover exists in the open-source edition ("多模型间 Failover: 支持") (Alibaba Cloud comparison). - Load balancing In config
Two unrelated layers. Classic HTTP balancing is an Ingress annotation:
nginx.ingress.kubernetes.io/load-balancedefaults toround_robinwithleast_connandrandomalso supported,ewmais explicitly not supported and silently falls back to round-robin, andupstream-hash-bygives consistent hashing on$request_uri,$host,$remote_addr, a header or a query arg (annotation compatibility). Across LLM credentials, ai-proxy simply picks anapiTokenat random per request (ai-proxy plugin). LLM-aware algorithms — "global least connections", "prefix matching" and "GPU-aware" load balancing — are described in a vendor engineering post rather than in the plugin reference (Higress engineering post), and noai-load-balancerpage exists in the documentation index (higress.ai/llms.txt). - Upstream health tracking In config
Health checking exists only at the API-key level, not the upstream level. ai-proxy
failoverisenabled: falseby default withfailureThreshold3,successThreshold1,healthCheckInterval5000 ms,healthCheckTimeout5000 ms, a requiredhealthCheckModelandfailoverOnStatus["4.*","5.*"]; an unhealthyapiTokenis removed from rotation and re-added once probes pass (ai-proxy, Chinese), which the AI quick-start describes in prose as pausing requests on a token until health checks recover it (AI quick start). No circuit-breaker or outlier-detection annotation appears in the annotation reference and no health-check keys appear in the global ConfigMap — and Alibaba Cloud's comparison page explicitly marks "故障自动检测及恢复" (automatic fault detection and recovery) as 不支持 for the open-source edition (Alibaba Cloud comparison). - Cross-region failover Not documented
The vendor does not document this, so any behaviour you observe today is unversioned and may change.
n.a. — no cross-region or multi-AZ failover configuration on the Helm guide, the global ConfigMap or the annotation reference. Alibaba Cloud's comparison page lists 多可用区部署 (multi-availability-zone deployment) as 自行构建, i.e. build it yourself, for the open-source edition (Alibaba Cloud comparison). Container images are mirrored to cn-hangzhou, us-west-1 and ap-southeast-7 registries, but that is image distribution, not traffic failover (README).
Fallback chain: Single alternate — One documented alternate target rather than a chain.
Defaults: config_file in two independent places. ai-proxy retryOnFailure: enabled default false, maxRetries default 1, retryTimeout default 30000 ms, retryOnStatus default ["4.*","5.*"], and "only non-streaming requests can be retried" (ai-proxy, Chinese). Ingress-level: nginx.ingress.kubernetes.io/proxy-next-upstream-tries default 3, with proxy-next-upstream-timeout having no timeout by default (annotation compatibility).
The reliability surface is a general-purpose gateway's, not an LLM router's, and it is split awkwardly: retries and key failover are ai-proxy plugin fields (both off by default), load balancing and request timeout are nginx-compatible Ingress annotations, and model fallback is a console-only setting. Read together with Alibaba Cloud's own feature matrix — which marks automatic fault detection and recovery, multi-AZ deployment, rate-limit degradation, monitoring/alerting and enterprise observability as "build it yourself" or unsupported in the open-source edition (Alibaba Cloud comparison) — the honest reading is that Higress gives you the primitives and expects you to operate them. Envoy underneath means richer circuit breaking is reachable through Istio APIs, but global.enableIstioAPI defaults to false (Helm values).
How fast the hop is
Compiled binaryA single compiled Go or Rust binary. The lowest overhead floor of the self-hostable options, and the easiest to reason about under load.
Envoy data plane plus a Go control plane: the project is "built on Istio and Envoy" and extended with Wasm plugins written in Go, Rust or JavaScript (what-is-Higress). Repo language bytes are Go 9,245,785, C++ 1,269,908, Rust 193,762, Shell 122,023, Python 88,069 and TypeScript 53,579 (GitHub languages API), i.e. a compiled control plane and a compiled proxy, with plugin code compiled to Wasm rather than interpreted per request.
Container images and a Helm chart, both published: higress-registry.cn-hangzhou.cr.aliyuncs.com/higress/all-in-one:latest runs console, HTTP and HTTPS on 8001/8080/8443 from one docker run (README), the chart is higress.io/higress installed into higress-system (README, Helm guide), and a Docker Compose bundle of apiserver/controller/pilot/gateway/console is generated by an installer script for non-Kubernetes hosts (Docker Compose guide). Alibaba Cloud Compute Nest deployment is also documented (documentation index). The FAQ notes there is no pre-built package beyond the images: "There is no existing one, you need to build it yourself. Currently, all Docker images are provided and can be pulled and used by yourself." (FAQ)
Streaming caveats: Yes, and it is a headline feature: Higress claims "true streaming processing" for SSE and supports the AI streaming (SSE) scenario (what-is-Higress), and ai-cache handles both streaming and non-streaming responses (ai-cache). Two documented streaming caveats matter: retries only apply to non-streaming requests (ai-proxy retryOnFailure), and on SSE the ai-data-masking plugin may fail to restore a masked word split across chunks and may leak part of a sensitive word to the client (ai-data-masking).
Published figures, grouped by what each one measured. Figures in different groups are different quantities and cannot be compared with one another — nor, in most cases, with another vendor’s figure in the same group.
Sustained capacity
Requests or queries per second sustained on the stated hardware.
- hundreds of thousands requests/second production traffic handled Vendor-published
README/overview marketing statement: "Born from Alibaba's internal product with over 2 years of production validation, supporting large-scale scenarios with hundreds of thousands of requests per second." No hardware, replica count, payload size or measurement method given, and it is aggregate Alibaba production traffic rather than a per-instance benchmark.
Source
What it will stop
4 of 4 can blockTwo separate questions per control: can it stop a request at all, and what does it do before you change any settings? A control that inspects and forwards is a logging feature, however it is named.
These controls can block a request, but will not until you change their settings:
- Harmful content — Ships switched off
- Your own policies — Only logs until you change it
- Personal data in prompts Can block the request
Out of the box: Blocks out of the box
sync_blockby default when the plugin is attached: ai-data-masking'ssystem_deny(built-in sensitive-word rules, sourced fromgithub.com/houbb/sensitive-word) anddeny_openaiboth default to true, anddeny_codedefaults to 200 with the message "Sensitive words found in the question or answer have been blocked". A softer mode exists —replace_roleswith regex/GROK patterns andtype: replaceorhash, plusrestore: trueto put the original values back into the model's answer (ai-data-masking). Documented failure mode on SSE: a masked word split across chunks may not be restored, and part of a sensitive word can reach the user. - Prompt injection and jailbreaks Can block the request
Out of the box: Blocks out of the box
sync_block, but only through the newerqwen3guardplugin, and only as a by-product: the vendor engineering post states Qwen3Guard's official safety policy covers violence, illegal acts, sexual content, PII, self-harm, unethical behaviour, politically sensitive topics and copyright infringement "并在输入审核中包含 Jailbreak 检测" (jailbreak detection is included in input moderation). The plugin enablescheckRequestandcheckResponseby default, callsQwen/Qwen3Guard-Gen-4B, usesriskLevelBar: Unsafeby default, and when the bar is met returns a refusal without calling the model at all (Higress content-security post). Caveats stated by the vendor: only theSafetyverdict drives the decision, per-category actions are not implemented, and streaming interception cannot honourdenyCode. No prompt-injection or jailbreak wording appears in the older ai-security-guard reference. - Harmful content Can block the request
Out of the box: Ships switched off
sync_blockvia ai-security-guard, which calls Alibaba Cloud's content-moderation service inline for both directions and denies the request when a risk label is returned;checkRequestandcheckResponseboth default to false,requestCheckServicedefaults tollm_query_moderation,responseCheckServicetollm_response_moderation, anddenyCodedefaults to 200 — so a blocked call looks like a success to a naive client (ai-security-guard). It emitsai_sec_request_deny/ai_sec_response_denymetrics andai_sec_risklabel/ai_sec_deny_phasespan attributes. The console walkthrough registers Alibaba Cloud Content Safety as a service source and attaches "AI Safety Guard" to a route (content security guide). - Your own policies Can block the request
Out of the box: Only logs until you change it
Custom policy is regex/GROK-based: ai-data-masking takes user-supplied
deny_wordsandreplace_rolesregex patterns with per-rulereplace/hashactions (ai-data-masking), and the genericrequest-blockandrequest-validationplugins add path/body blocking and JSON-schema request validation (plugin marketplace). Custom rules you add yourself default to replacement rather than denial unless you put them in the deny list, hencesync_observeas the default posture for custom rules specifically.
fail_open, stated explicitly and unusually candidly for the qwen3guard plugin: "当前插件选择 fail-open:记录警告并放行" — on connection failure, timeoutMs (default 2000 ms) expiry, a non-200 from the guard model, unparseable JSON or a missing Safety field, the request is logged and allowed through. The same post says compliance scenarios needing forced fail-close "不能被描述为已经满足" (cannot be described as satisfied) in the current version (Higress content-security post). No fail-open/fail-closed statement appears in the older ai-security-guard reference.
Calls out to: Alibaba Cloud Content Moderation, Qwen3Guard. Each is a separate vendor relationship and a separate hop on the request path.
The defaults to check before trusting this stack: ai-security-guard ships with checkRequest: false and checkResponse: false, so attaching it without editing those fields inspects nothing (ai-security-guard); denyCode defaults to 200 in both ai-security-guard and ai-data-masking, so blocked traffic returns HTTP 200 with a refusal body unless you change it; and the qwen3guard plugin is deliberately fail-open, so an outage of the moderation model silently disables moderation (Higress content-security post). Streaming is the weakest surface: ai-data-masking may leak fragments of a sensitive word across SSE chunks (ai-data-masking), and qwen3guard cannot apply denyCode once bytes have shipped.
What you can see
Exports widelyYou decide whether bodies are captured, by setting or by header.
n.a. as a documented switch: the log page explains how to change accessLogFormat and how to view logs but states no way to disable access logging, and the global ConfigMap and the annotation reference contain no logging on/off key. In practice a self-hosted operator controls this at the Envoy/ConfigMap level, but Higress does not document it.
OpenTelemetry, plus SkyWalking and Zipkin, configured in the global ConfigMap under tracing: enable defaults to false, sampling to 100.0, timeout to 500 ms, and only one exporter can take effect at a time; the opentelemetry block takes a service and a gRPC port (global configuration). Spans are enriched by ai-statistics via apply_to_span attributes, and trace_id is an access-log field (ai-statistics, log description).
Configurable in both directions: accessLogFormat under the mesh field controls which access-log fields are emitted (log description), and ai-statistics has an explicit "record questions and answers" mode where attributes entries with value_source of request_body, response_body or response_streaming_body and apply_to_log: true write prompt and completion text into the log — both flags default to false, so content logging is opt-in (ai-statistics).
Where telemetry can go
- OpenTelemetry
- Prometheus
- Grafana
- Loki
- SkyWalking
- Zipkin
Traces go to OpenTelemetry, SkyWalking or Zipkin collectors you name in the ConfigMap (global configuration); metrics are scraped from the gateway's metrics port (default 15020 in the Docker Compose install) by Prometheus and drawn in Grafana, with --set global.o11y.enabled=true installing Grafana, Prometheus, Loki and PromTail for you, or an external Grafana URL pasted into the console instead (Prometheus guide, Helm values, Docker Compose options). There is no vendor SaaS sink and no Datadog/LangSmith/S3 integration on any fetched page.
No feedback, rating or annotation endpoint. ai-statistics can attach arbitrary attributes to logs and spans from request or response fields, which is the closest available mechanism, but nothing captures a user verdict (ai-statistics); nothing appears in the plugin marketplace either.
No evaluation, scoring or experiment feature exists: the 41-plugin marketplace has no eval plugin (plugin marketplace) and the documentation index lists none across overview, user guide, ops, developer, AI gateway and scene-guide sections (higress.ai/llms.txt). Traffic-shaping primitives that could support an offline comparison do exist — canary annotations (canary-by-header, canary-weight) and 多模型灰度 model-level canary — but no scoring is attached (annotation reference, Alibaba Cloud comparison).
Depends on the vendor’s SaaS: No. The console ships "a built-in monitoring suite based on Prometheus + Grafana, though it's not installed by default", enabled with --set global.o11y.enabled=true, and if you skip it the Monitoring Dashboard page accepts an external Grafana URL instead (Prometheus guide). Token-level AI metrics come from the self-hosted ai-statistics plugin across gateway, route, service and model dimensions (ai-statistics). The one caveat is Alibaba Cloud's own matrix, which lists 企业级可观测 (enterprise-grade observability) and 监控告警 (monitoring and alerting) as "build it yourself" for the open-source edition (Alibaba Cloud comparison).
Retention: No retention period is published or configurable by Higress, because Higress stores nothing centrally: logs go to the gateway pod's stdout and, if you enable the bundled observability stack, into your own Loki (Helm values). Retention is whatever your log pipeline is set to.
Whether it fits how you work
How much work stands between you and a first call, how different that is from running it in production, and whether it slots into the stack you already have. Recorded as the shape of the work rather than a number of minutes — how long it takes you depends on which accounts and quota you already hold, which no comparison can know.
to try: cli or container to run: infrastructure rollout fits 4 of 10 common stacks
Getting to a first call
No numbered procedure publishedNothing works until you have a process running. Fine on a laptop, but it means there is no zero-install way to try it.
Read off: the vendor’s own quickstart — no numbered procedure published.
Why the count is not the work: No numbered step list exists to count. The AI gateway quick start is organised as screenshot-driven sections (install by Docker, log in to the console, configure an LLM provider key, configure an AI route, observe the AI dashboard) rather than enumerated steps (AI quick start), and the Kubernetes quick start is structured as three stages — install, configure, validate — with branching environment options rather than a linear count (quick start). Recorded as 0 to mean "vendor publishes no step count", not "zero work".
Before step one
- Your own provider key Required
You need an upstream provider account and key before anything works. That is a prerequisite, not a step.
You cannot make a first call without an upstream key. The installer prompts "enter the Aliyun Dashscope or other API-KEY" (you may skip and add it in the console instead) and "LLM Provider Management" is where keys for Alibaba Cloud, DeepSeek, Azure OpenAI, OpenAI, DouBao and others are stored (AI quick start); every ai-proxy provider block takes
apiTokensor cloud credentials (ai-proxy plugin). - Payment method No card needed to start
Not required: the Community edition is "Free & Open Source" with local deployment (AI gateway editions) and the install path is a Docker or Helm command with no account creation (README). A card only enters the picture if you choose Alibaba Cloud's managed edition (quick start).
- Gate before models answer No gate
Every catalogue model is callable as soon as you have a key.
No gate: you add a provider block with your own
apiTokensand the model is immediately routable, withmodelMappingdeciding which names clients may use (ai-proxy plugin); the console equivalent is adding a key under LLM Provider Management (AI quick start). No approval, enablement, quota or waitlist step appears on either page.
Everything you need first: Docker and an LLM provider key. The AI gateway quick start is a Docker install that exposes the console on http://localhost:8001/, asks you to set an admin account on first login, and prompts for an Alibaba Cloud DashScope or other API key which you may skip and add later; it also warns that "AI Gateway needs to access Internet resources" during start-up (AI quick start, token management guide). No Kubernetes cluster, cloud account or credit card is involved.
The vendor’s own time claim: Vendor claim, verbatim: "Local deployment, verify core capabilities in 5 minutes" for the Community edition (AI gateway editions). The README separately claims Higress "can be deployed with a single Docker command" outside Kubernetes (README). Quoted, not verified. Marketing time claims assume every account and approval is already in place.
Running it in production
The same scale applied to the path the vendor recommends for production traffic. Kept separate from the quickstart because for several products here the two are barely related pieces of work.
This is an infrastructure project, not an integration. Expect charts or Terraform, networking, secrets and someone who owns the deployment.
Nothing works until you have a process running. Fine on a laptop, but it means there is no zero-install way to try it.
Getting to production is a step up in kind from the quickstart, not just more of the same.
What production needs: A Kubernetes cluster with Helm, plus a LoadBalancer service (the quick start documents hostNetwork or MetalLB fallbacks when none is available) (quick start). Redis is required for ai-cache, ai-token-ratelimit, ai-quota and MCP hosting (ai-cache, ai-token-ratelimit, MCP quick start). Add-ons as needed: Prometheus/Grafana/Loki via global.o11y.enabled=true, Nacos >= 3.0 for the MCP registry, an embedding service plus a vector database for semantic caching, and an Alibaba Cloud Content Safety subscription for ai-security-guard (Helm values, MCP quick start, semantic cache guide, content security guide).
Can you run it yourself
There is a command you can copy and run, so you can evaluate the self-hosted path yourself today.
One container: docker run -d --rm --name higress-ai -v ${PWD}:/data -p 8001:8001 -p 8080:8080 -p 8443:8443 higress-registry.cn-hangzhou.cr.aliyuncs.com/higress/all-in-one:latest (console 8001, HTTP 8080, HTTPS 8443). Kubernetes: helm install higress -n higress-system higress.io/higress --create-namespace. Non-Kubernetes AI gateway installer: wget https://higress.cn/ai-gateway/install.sh then run with bash.
How it fits your stack
4 of 10Each row is a thing you might already run. “With a caveat” means it works but not the way the vendor’s marketing implies — read the reason, because that is usually where the surprise lives.
- Fits The OpenAI SDK Drop-in once set up — but first-call work is cli or container.
- No The Vercel AI SDK No AI SDK route documented.
- No Cloudflare Workers Workers AI is an upstream model here, which is the opposite direction.
- Fits Kubernetes `higress.io/higress` from the `https://higress.io/helm-charts` repo, installed into namespace `higress-system`; sub-charts `higress-core` (controller + gateway) and `higress-console`
- No Terraform or OpenTofu Nothing published for Terraform.
- Fits An existing API gateway This is that gateway — AI traffic becomes a plugin, not a new hop.
- No Cloud IAM I already run Static upstream credentials only. Your calls to it still use its own key.
- No LangChain or LlamaIndex No framework integration documented.
- Fits MCP servers to govern Acts as an MCP gateway or registry.
- With a caveat Nothing — plain Node or Python You have to run a process locally before any call works.
Reading this the other way round — pick what you already run and see every product scored against it.
The integration surfaces behind those answers
- Vercel AI SDK Not documented
Nothing published. Assume the OpenAI-compatible route and verify it yourself.
n.a. — the Vercel AI SDK is not mentioned on the documentation index, the ai-proxy reference or the AI gateway page. In practice
@ai-sdk/openaiwith a custombaseURLshould work against the OpenAI-compatible route (ai-proxy plugin), but Higress documents no such integration, so this is recorded as not documented rather than as compatibility. - Cloudflare Workers Workers AI as an upstream model
Cloudflare Workers AI is reachable as a model behind this product. That is the opposite direction from running your code on Workers.
Cloudflare appears only as a model source: ai-proxy has a
cloudflareprovider type for Cloudflare Workers AI, configured withcloudflareAccountId(ai-proxy plugin). Higress itself is a self-hosted Envoy process and does not run on Workers (what-is-Higress). - Kubernetes Official Helm chart
A named, published chart. You can read its values file before committing to anything.
Named:
`higress.io/higress` from the `https://higress.io/helm-charts` repo, installed into namespace `higress-system`; sub-charts `higress-core` (controller + gateway) and `higress-console`Kubernetes is the first-class target:
helm repo add higress.io https://higress.io/helm-chartsthenhelm install higress -n higress-system higress.io/higress --create-namespace, withglobal.local=truefor a kind cluster andhgctlto open the console (Helm guide, quick start). The README states Higress is a conformant Gateway API implementation, a conformant Gateway API Inference Extension implementation, and is listed in the official Kubernetes Ingress Controllers documentation (README). It is also nginx-Ingress-annotation compatible, which is its main migration story (annotation compatibility). - Terraform Not documented
No Terraform surface published. Configuration is API or dashboard work.
n.a. — no Terraform provider, module or registry reference on the documentation index, the Helm guide or the Docker Compose guide. Infrastructure-as-code is expressed instead as Helm values and Kubernetes Ingress/ConfigMap objects, plus an Alibaba Cloud Compute Nest deployment path (documentation index).
- Existing API gateway It is the API gateway
This product is the gateway. If you already run it for your other APIs, AI traffic becomes a plugin rather than a new hop.
It is the gateway: Higress is described as an AI-native API gateway built on Istio and Envoy that unifies AI gateway, Kubernetes Ingress gateway, microservice gateway and security gateway roles in one deployment (what-is-Higress), and the home page splits the AI side into LLM Gateway, MCP Gateway and Model Gateway (the last implementing the Gateway API Inference Extension with InferencePool and Endpoint Picker) (home page).
- Cloud identity Static provider credentials only
You paste static provider credentials into this product so it can reach upstream models. Your own calls to it still use its own API key, and those pasted secrets are yours to rotate.
Cloud credentials are passed as static plugin fields, not assumed roles: AWS Bedrock takes
awsAccessKey,awsSecretKeyandawsRegion, and Google Vertex AI takesvertexAuthKey,vertexRegion,vertexProjectIdandvertexAuthServiceName(ai-proxy plugin). No IRSA, workload identity, instance-profile or credential-chain support is documented. IDaaS/OIDC integration for the control plane is listed as a commercial-edition feature (深度集成阿里云 IDAAS 产品) versus "build it yourself" for open source (Alibaba Cloud comparison), although an OIDC plugin exists for data-plane auth (plugin marketplace). - MCP MCP gateway or registry
It sits in front of many MCP servers and governs access to them. This is the shape you want if MCP sprawl is the problem you are solving.
A real MCP gateway, not just a client. The
mcp-serverplugin both converts REST APIs into MCP tools with no code (server.type: rest,tools[].requestTemplate) and proxies existing MCP servers (server.type: mcp-proxywithmcpServerURL,server.timeoutdefault 5000 ms), withallowToolsallowlisting,securitySchemesplusdefaultDownstreamSecurity/defaultUpstreamSecurity, andpassthroughAuthHeaderdefaulting to false so client credentials are not leaked to backends (mcp-server plugin). Hosting requires Higress >= 2.1.0 and Redis; the Nacos MCP registry needs Nacos >= 3.0 and Higress >= 2.1.2; SSE endpoints are exposed at/{mcp route prefix}/{nacos name}/sseand the 2025-03-26 streamable-HTTP protocol needs no ConfigMap entry (MCP quick start). The home page announces support for the 2026-07-28 MCP revision "with compatibility for existing protocols" (home page), anopenapi-to-mcpconversion tool and a hosted demo at https://mcp.higress.ai/ (README).
n.a. — neither LangChain nor LlamaIndex is documented anywhere in the official docs index or the plugin marketplace (higress.ai/llms.txt, plugin marketplace); a LangChain + Higress + Elasticsearch RAG walkthrough exists only as a third-party blog post, not vendor documentation. RAG is instead offered as a gateway plugin (ai-rag) and via ai-search (plugin marketplace).
n.a. — no Higress client library exists on the ai-proxy reference, the documentation index or the AI quick start; the documented calling convention is curl or any OpenAI-compatible SDK against the gateway route. The Go, Rust and JavaScript SDKs Higress does publish are for authoring Wasm gateway plugins, not for calling the gateway (what-is-Higress), so no client-language list is recorded here.
Agent features: Agent support is plugin-shaped and fairly basic. The ai-agent plugin runs a ReAct loop in the gateway against OpenAPI-described tools — llm block with maxIterations default 15, maxExecutionTime default 50000 ms and maxTokens default 1000, apis[].api holding the tool's OpenAPI document, and Chinese/English ReAct prompt templates; the worked example wires Amap, XZWeather and DeepL against qwen-max-0403 (ai-agent plugin). Note that this is prompt-driven ReAct, not OpenAI tool-calling: the page says nothing about function calling. The stronger agent story is MCP — REST-to-MCP conversion and MCP proxying with per-tool auth (mcp-server plugin) — plus ai-history for multi-turn context and ai-intent for intent routing (plugin marketplace).
Fast to a first token, slower to a defensible production posture. One docker run gives you console, HTTP and HTTPS on 8001/8080/8443 (README) and the console walks you through LLM Provider Management, Service Sources and AI Route Config with per-route strategies for auth, rate limiting, RAG, prompt templates and semantic caching (AI quick start). Friction to expect: the installer needs outbound internet access at start-up (token management guide); Redis must exist before caching, token rate limiting, quotas or MCP hosting work (ai-cache, MCP quick start); guardrails require an Alibaba Cloud Content Safety subscription or a self-hosted Qwen3Guard model (content security guide, Higress content-security post); and a meaningful part of the reference material — including the failover and retryOnFailure field tables — is only on the Chinese pages (ai-proxy, Chinese).
The ecosystem is cloud-native rather than AI-native: service discovery from Nacos, ZooKeeper, Consul and Eureka, and governance interop with Dubbo, Nacos and Sentinel (what-is-Higress); Wasm plugins in Go, Rust or JavaScript with sandbox isolation and hot updates; Prometheus, Grafana, Loki, SkyWalking, Zipkin and OpenTelemetry for telemetry (global configuration, Prometheus guide). The AI-side dependencies lean Alibaba: content moderation, BaiLian embeddings and DashVector in the worked examples, and the recommended managed upgrade path is Alibaba Cloud AI Gateway (content security guide, semantic cache guide, quick start). Sibling projects HiMarket (API/agent portal, documented on the same site) and HiClaw are also named (documentation index, CNCF announcement).
What it does well
- One Apache-2.0 binary covers Kubernetes Ingress, microservice routing, LLM proxying and MCP hosting; nginx-Ingress annotation compatibility makes it a credible ingress-nginx replacement
- 31 provider types behind the OpenAI contract with path-based protocol detection for chat completions, Anthropic Messages and embeddings, plus `protocol: original` passthrough
- Genuine MCP gateway: REST-to-MCP with no code, MCP proxying, per-tool allowlists and separate downstream/upstream security schemes, with `passthroughAuthHeader` off by default
- Cache is both exact and semantic in one plugin, and all state (Redis, vector DB, Prometheus, Loki) stays in your own infrastructure
- Token-level governance is real: per-consumer quotas with an admin API, Redis-backed token-per-second/minute/hour/day limits, and ai-statistics metrics split by gateway, route, service and model
- CNCF Sandbox project since 15 March 2026 with a named enterprise adopter list, active releases and a documented plugin-authoring SDK in Go, Rust and JavaScript
Where it falls short
- Reliability primitives are off by default and shallow: ai-proxy `retryOnFailure` disabled with `maxRetries: 1` and no backoff, `failover` disabled, retries only on non-streaming requests
- `denyCode` defaults to 200 in both ai-security-guard and ai-data-masking, so blocked traffic returns HTTP 200 unless you change it; ai-security-guard also ships with `checkRequest`/`checkResponse` false
- The qwen3guard plugin is explicitly fail-open and the vendor states forced fail-close "cannot be described as satisfied" in the current version
- Alibaba Cloud's own matrix marks automatic fault detection and recovery as unsupported, and multi-AZ deployment, rate-limit degradation, monitoring/alerting and enterprise observability as build-it-yourself, for the open-source edition — with no SLA at all versus 99.99% for the paid product
- No published model catalogue or `/v1/models` endpoint: the only figure is a repeated "100+ models" marketing claim
- Documentation is unevenly bilingual — the `failover`, `retryOnFailure` and full 31-provider tables appear on the Chinese ai-proxy page but not the English one
- No image, audio, video, rerank, batch or Responses API surface, and no Terraform, LangChain, LlamaIndex or Vercel AI SDK integration is documented
- No evaluation or feedback capture of any kind, and the LLM request timeout defaults to no timeout at the Ingress layer
Choose it when
Teams already running Kubernetes who want one Envoy-based gateway for north-south traffic, microservices and LLM/MCP calls, and who are happy to assemble reliability and guardrail behaviour from plugin fields themselves.
Look elsewhere when
You want an LLM-first router with a published model catalogue, per-request routing policy and turnkey defaults; you need English-only documentation; or you need a vendor SLA, SOC 2 report or fail-closed guardrails, none of which exist for the open-source edition.
Open source: You run the gateway yourself. Full control over the data path and no per-request vendor fee, in exchange for operating it.
How hard is it to leave?
Derived from six published facts, not from an opinion. The weights are fixed and the same for every product — see the arithmetic.
| What helps you leave | Points | Source |
|---|---|---|
| Works with standard OpenAI code Switching away is a base-URL change rather than a rewrite of every call site. | 22 /22 | vendor page |
| No vendor-specific SDK required A proprietary client library spreads through your codebase and has to be torn out again. | 10 /10 | — |
| Can use your own provider accounts Your keys and billing relationship stay yours, so removing the gateway does not cut off model access. | 20 /20 | — |
| Can be self-hosted You can run it yourself instead of accepting a pricing or policy change. | 20 /20 | vendor page |
| Configuration lives in version control Routing and budget rules are a file you keep, not dashboard state you would have to rebuild. | 16 /16 | — |
| Your request history can be exported You leave with your own logs instead of abandoning them. | 12 /12 | — |
Read the fine print: Everything is in your own systems by construction: configuration lives in Kubernetes Ingress/CRD objects, in the global ConfigMap, or — in the non-Kubernetes install — in local files or a Nacos namespace you nominate with `-c file:///opt/higress/conf` or `nacos://host:8848` ([Docker Compose options](https://higress.ai/en/docs/latest/ops/deploy-by-docker-compose/), [quick start](https://higress.ai/en/docs/latest/user/quickstart/)). Telemetry lands in your Prometheus, Loki and OTLP collectors ([Prometheus guide](https://higress.ai/en/docs/latest/user/prometheus/)). No vendor account holds anything to export.
All six inputs are published, so this is scored against the full 100. This measures technical switching cost only. It does not price the engineering time to re-test prompts against a different routing stack.
Higress models & pricing
Browse every imported listing from this provider, with published token rates and a link to compare other providers for the same model. This is provider-reported coverage; an absent listing does not mean unsupported.
Loading model listings…
Official model coverage source ↗ · Model source coverage and limitations
Full specification
Every field we track. Blank fields say "Not published" rather than "No" — we do not infer an absence from silence. Switch to Technical in the header for the precise field names and the low-level details.
Overview
Interpret these fields: Self-hosted vs managed LLM gateways · Who still owns your LLM gateway?
- What kind of product Category
- Open source
- Marketplaces resell many providers behind one key. Gateways add governance on top. Open-source projects you run yourself. Cloud platforms are hyperscaler surfaces. Inference providers host models on their own hardware.
- Who runs it Deployment model
- Managed or self-host
- Managed means the vendor operates it. Self-host means you run it on your own infrastructure. Both means you can choose.
- Licence Licence
- Apache-2.0
- Proprietary products cannot be inspected or forked. Open licences such as MIT and Apache-2.0 let you audit, modify, and run the code without permission.
- Company Company
- Alibaba Group; governed as a CNCF Sandbox project ("Copyright Higress a Series of LF Projects, LLC")
- The organisation that maintains the product.
- Who you would be signing with Vendor status
- Run by a software foundation
- Whether the product is still an independent company, has been acquired, is a large cloud vendor’s product line, is run by a software foundation, or has been put into maintenance mode. Maintenance mode means bug fixes and security patches only — no new features.
- Last shipped an update Latest release
- 2026-08-13
Latest non-prerelease is v2.2.4, published 2026-08-13T23:49:46Z via the GitHub releases API; the default branch was last pushed 2026-09-02T12:30:09Z ([releases/latest](https://api.github.com/repos/higress-group/higress/releases/latest)).
- The date of the most recent release or version tag. A product that has not shipped in a year is a different risk from one that shipped last week, regardless of what its marketing site says.
- GitHub stars GitHub stars
- 9,458
- A rough proxy for community size on open-source projects. Not a quality measure.
Cost
Interpret these fields: How LLM gateway pricing works · LLM gateway spending limits: stop a runaway agent bill?
- Markup on model prices Token markup
- None
- How much the product adds on top of what the underlying model provider charges. Zero means you pay the same per-token price you would pay the model provider directly.
- Fee to add funds Credit purchase fee
- Not published
- A percentage charged when you top up your balance, separate from token prices. It is easy to miss because it does not appear on the per-token price list.
- Monthly cost per person Seat fee
- Not published
- A recurring per-user platform charge that applies regardless of how much you use the models.
- Can use your own provider accounts BYOK supported
- Yes
- Bring Your Own Key: you keep direct contracts with OpenAI, Anthropic and others, and the gateway only routes traffic. This preserves negotiated rates and committed-spend discounts.
- Cost of using your own accounts BYOK terms
- No fee of any kind: the software is Apache-2.0 and nothing meters your traffic, so BYOK is the only mode and it costs nothing beyond your provider bills and your own compute ([GitHub API](https://api.github.com/repos/higress-group/higress), [ai-proxy plugin](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-proxy/)).
- What the product charges to route traffic through your own provider keys.
- Free tier Free tier
- The entire gateway is free: Apache-2.0 licensed ([GitHub API](https://api.github.com/repos/higress-group/higress)) with the Community edition described as "Free & Open Source, Community Support" and "Local deployment, verify core capabilities in 5 minutes" ([AI gateway editions](https://higress.ai/en/ai-gateway/)). There is no request, seat or token cap in the open-source edition, and no Higress account exists to create.
- What you can do without paying, useful for evaluation.
- Enterprise plan from Enterprise plan from
- Not published
- Annual entry price for the enterprise tier, where one is published or credibly reported.
- Cost to run it yourself Self-host cost
- Your own infrastructure only — no licence, no seat fee, no metered tier. Sizing signals from the docs: gateway defaults to 2 replicas, the console to 1, and the optional o11y stack adds Grafana, Prometheus, Loki and PromTail to the same cluster ([Helm values](https://higress.ai/en/docs/latest/ops/deploy-by-helm/)). Redis is a hard dependency for ai-cache, ai-token-ratelimit, ai-quota and MCP hosting, so budget for it ([ai-cache](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-cache/), [ai-token-ratelimit](https://higress.ai/en/docs/latest/plugins/ai/api-consumer/ai-token-ratelimit/), [MCP quick start](https://higress.ai/en/docs/ai/mcp-quick-start/)). Autoscaling is your job: "Higress is based on K8s HPA and supports elastic scaling. The gateway is stateless and is a deployment." ([FAQ](https://higress.ai/en/docs/latest/overview/faq/))
- What self-hosting actually costs once you account for infrastructure and any paid tier.
- How the vendor makes money Pricing model
- Open source with a managed tier
- The shape of the vendor’s bill: does the routing layer charge a percentage on top of tokens, a flat monthly fee, both, neither (because inference is the product), or nothing at all (open source with no paid tier).
- How pricing works, briefly Pricing model detail
- Apache-2.0 software with no price of its own, alongside a separately-sold managed edition: the AI gateway page presents Community ("Free & Open Source, Community Support", local deployment), Serverless via Alibaba Cloud AI Gateway, and a Feitian Exclusive Edition with negotiable SLA and ticket/DingTalk support ([AI gateway editions](https://higress.ai/en/ai-gateway/)). The quick start states "Serverless Standard starts at ¥0 and charges only for actual usage" and recommends Alibaba Cloud AI Gateway Enterprise for production without Kubernetes ([quick start](https://higress.ai/en/docs/latest/user/quickstart/)). No dollar or yuan rate card is published on the Higress site itself, and the managed edition is an Alibaba Cloud product, not a Higress-branded SaaS.
- A one-paragraph description that covers the caveats a pricing category cannot: introductory rates, per-feature meters, tier gating, and pricing that resets on a specific date.
- Minimum commitment Minimum commitment
- None for the open-source edition — no account, contract or licence key is involved ([GitHub API](https://api.github.com/repos/higress-group/higress)). The managed alternatives are Alibaba Cloud commitments: Serverless Standard "starts at ¥0 and charges only for actual usage", while the Feitian Exclusive Edition is a negotiated commercial agreement ([quick start](https://higress.ai/en/docs/latest/user/quickstart/), [AI gateway editions](https://higress.ai/en/ai-gateway/)).
- Whether the vendor requires a minimum contract term, a minimum spend, or a provisioned-capacity purchase to get its published rate.
- Charges that fire after you go over an allowance Overage terms
- n.a. — there is no meter to overrun. Higress bills nothing; the costs you incur are your own infrastructure plus the upstream provider bills against the API keys you configure in ai-proxy ([AI gateway editions](https://higress.ai/en/ai-gateway/), [ai-proxy plugin](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-proxy/)).
- The line items that scale with usage after an included allowance is exhausted — log storage, extra requests, per-feature meters, data export — which is where cost estimates usually go wrong.
- Prompt cache offered Cache mechanism
- Both exact and semantic
- Whether the gateway offers its own response cache, what kind of match it does (exact request, prefix, semantic), or simply passes provider caching through unchanged.
- Discount on cached input Cache-read discount
- Not published
- How much cheaper cached tokens are than fresh input, when the vendor publishes a single figure. Bundled-inference clouds usually price this per model instead of as one number.
- Premium on cache writes Cache-write premium
- Not published
- How much more the first write of a cached prefix costs versus a plain input token. A high write premium and a low hit rate can leave you paying more than you save, so this matters as much as the read discount.
- Who captures the cache saving Cache economics
- Both shapes, self-hosted, and Redis is mandatory. The ai-cache plugin does exact-match caching keyed by default on `messages.@reverse.0.content` with prefix `higress-ai-cache:` and `cacheTTL` defaulting to 0, i.e. never expire; `redis.serviceName` is required and `x-higress-skip-ai-cache: on` bypasses the cache per request ([ai-cache plugin](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-cache/)). Semantic caching is the same plugin pointed at an embedding service plus a vector database — the worked example uses Alibaba Cloud BaiLian `text-embedding-v3` at 1024 dimensions with DashVector and Cosine distance ([semantic cache guide](https://higress.ai/en/docs/ai/scene-guide/semantic-cache/)). No cached-token pricing exists because no vendor meters tokens here; a hit simply skips the upstream call, so you save the provider's own bill.
- Whether the customer keeps the full saving from caching or the vendor captures part of it — and any conditions attached (write premium, storage fees, best-effort hits).
- What you can split spend by Cost attribution
- Per gateway, route, service and model: ai-statistics exposes input tokens, output tokens, time to first token (streaming) and total request time "observable across four dimensions: gateway, route, service, model", with `input_token`, `output_token`, `model`, `llm_service_duration` and `llm_first_token_duration` in the log line ([ai-statistics](https://higress.ai/en/docs/latest/plugins/ai/api-o11y/ai-statistics/)). Per-consumer attribution comes from key-auth's `X-Mse-Consumer` header plus ai-quota's per-consumer Redis counters ([key-auth](https://higress.ai/en/docs/latest/user/plugins/authentication/key-auth/), [ai-quota](https://higress.ai/en/docs/latest/user/plugins/ai/api-consumer/ai-quota/)). Attribution is in tokens, not currency — no price table or cost field is computed anywhere.
- The dimensions the vendor documents for splitting spend — per key, per user, per team, per tag, per customer. Matters if you need to chargeback internally or bill an end customer.
- How you get cost data out Cost export
- n.a. — no CSV, billing API, webhook, S3 or warehouse cost export on [ai-statistics](https://higress.ai/en/docs/latest/plugins/ai/api-o11y/ai-statistics/), [ai-quota](https://higress.ai/en/docs/latest/user/plugins/ai/api-consumer/ai-quota/) or [the token management guide](https://higress.ai/en/docs/ai/scene-guide/token-management/). What you can export is token telemetry, via Prometheus scraping and OTLP ([Prometheus guide](https://higress.ai/en/docs/latest/user/prometheus/)).
- The mechanisms the vendor publishes for exporting cost and usage data: CSV, an API, webhooks, S3, a data warehouse, or nothing at all. Any per-unit price is included.
- Who pays the model bill BYOK mode
- Your keys only
- Whether you bring your own provider accounts (BYOK), buy inference from this vendor, or can do either. This is the single biggest commercial difference between these products: it decides who holds the contract with the model provider and who carries the spend.
Spend governance in detail
Seven signals matter when a bill starts to hurt: who can spend, how much, on what, and who gets paged when it goes wrong. Everything below is drawn from the vendor’s own pricing and docs pages — how we read these.
- Virtual or scoped keys Yes
Keys that carry their own budget and rate-limit policy, so an intern experiment cannot spend against a production budget.
Console "Consumer Management" issues per-consumer credentials verified by key-auth against the `x-api-key` header, and routes carry an allowed-consumer list; the term "virtual key" is not used. Errors are 401 for a missing/invalid key and 403 for an unauthorised consumer.
- Budget caps per key Yes
A dollar or token ceiling attached to an individual key. Where enforcement is soft, one over-limit request still completes before the block kicks in.
ai-quota holds a fixed token quota per consumer in Redis under the `chat_quota:` prefix, with an admin API at `admin_path` (default `/quota`) for GET/refresh/delta. Requires key-auth or jwt-auth plus ai-statistics.
- Budget caps per team or workspace Not published
A ceiling applied at a higher scope than one key — a team, a workspace, a customer, or an entire environment.
Not stated. Quotas and rate limits are per consumer, header, cookie, param or IP; no team or org object exists in the fetched docs.
- Rate limiting as a cost control Yes
Configurable request-per-time-window caps. Platform-set rate limits do not count as spend controls; user-configurable ones do.
ai-token-ratelimit does Redis-backed token-per-second/minute/hour/day limits, keyed by header, query param, consumer, cookie or IP, with rule-level global thresholds; rejects with 429 and "Too many requests". Requires ai-statistics. Request-rate limits exist separately as `higress.io/route-limit-rpm`/`-rps` annotations and the cluster/local rate-limit plugins.
- Model allowlists Yes
A policy that constrains which models a key or team can call, keeping expensive frontier models out of the wrong hands.
Enforced indirectly through ai-proxy `modelMapping` (including a `*` catch-all and regex mapping) and through route-level model matching with an allowed-consumer list in AI Route Config.
- Spend alerts Not published
Alerts fired as spend approaches a threshold. Alerts that only fire after the meter has rolled over are marked as such.
Not stated. Token metrics are Prometheus-scrapable so alerting is possible in your own Alertmanager, but no alert feature is documented; Alibaba Cloud lists 监控告警 as "build it yourself" for the open-source edition.
- Webhook notifications Not published
Programmatic notifications on spend events, so budget breaches can page an on-call or open a ticket.
Not stated on any fetched page.
Enforcement: Enforced before each request
Catalog
Interpret these fields: LLM gateway model counts: what “500+” means
- Models available Models available
- ~100
- How many models you can call, shown as a range because several vendors publish different totals on different pages. Vendor-reported either way, so counts are not directly comparable — some count every provider variant of the same model separately.
- Model providers reachable Upstream providers
- 26–31
- How many distinct model providers or labs you can reach, shown as a range where the vendor’s own pages disagree. More providers usually means better redundancy when one has an outage.
- Works with standard OpenAI code OpenAI-compatible API
- Yes
- If yes, you can usually switch to it by changing one base URL, and switch away just as easily. This is the main defence against lock-in.
- OpenAI chat endpoint POST /v1/chat/completions
- Yes
- The endpoint almost every application ports first. "Not documented" means the vendor never states it, which is different from a documented no.
- Anthropic messages endpoint POST /v1/messages
- Yes
- Whether Anthropic-shaped calls work without rewriting them. Several products support this only as SDK compatibility or provider passthrough rather than a native endpoint — the detail page says which.
- OpenAI Responses endpoint POST /v1/responses
- Not documented
- The newer stateful OpenAI surface. Support is much thinner across this market than chat completions.
- Embeddings endpoint POST /v1/embeddings
- Yes
- Whether you can generate vectors through the same gateway, or need a second integration for retrieval workloads.
- Image generation endpoint POST /v1/images/generations
- Not documented
- Whether image models are reachable through the same surface as text.
- Audio endpoints POST /v1/audio/*
- Not documented
- Speech-to-text and text-to-speech. Frequently the first gap in an otherwise complete gateway.
- Batch jobs endpoint POST /v1/batches
- Not documented
- Asynchronous bulk processing, usually at a discount. Commonly undocumented, and commonly the reason a migration stalls late.
- Needs the vendor’s own code library Requires a vendor-specific SDK
- No
- A proprietary client library spreads through your codebase and has to be torn out again if you leave. “No” is the better answer here, and it means the standard OpenAI client works.
- You can export your request history Logs / usage data export
- Yes
- Whether you can get your own request logs, traces, or usage records back out — through an API, a bulk export, or a download. Decides whether you leave with your history or abandon it.
- Settings can live in version control Declarative config-as-code
- Yes
- Whether routing, fallback, and budget rules can be declared in a file you keep in Git, rather than existing only as settings clicked into a hosted dashboard.
- Embeddings Embeddings
- Yes
- Text-to-vector models, needed for search and retrieval features.
- Image generation Image generation
- Not published
- Whether image models are reachable through the same interface.
- Speech and audio Speech and audio
- Not published
- Text-to-speech or transcription models through the same interface.
- Video generation Video generation
- Not published
- Whether video models are reachable through the same interface.
- Batch processing Batch processing
- Not published
- Submitting large jobs for cheaper, slower processing. Often 50% off for work that is not time-sensitive.
Routing & reliability
Interpret these fields: How LLM gateway failover actually works · Does your LLM gateway promise any uptime? · Is routing destroying your prompt cache? · Changing models without breaking production
- Uptime it promises in writing Contractual SLA uptime
- Not published
- The uptime percentage in a published, contractual service level agreement. A public status page is not an SLA — it reports what happened, it does not promise anything or pay you back when it breaks.
- Automatic failover Automatic failover
- Yes
- When a model provider goes down or rate-limits you, traffic moves to a backup automatically instead of returning errors to your users.
- Load balancing Load balancing
- Yes
- Spreads requests across several providers or keys to raise your effective rate limit.
- Rule-based routing Conditional routing
- Yes
- Send different requests to different models based on rules — for example a cheap model for free users and a strong model for paying ones.
- Response caching Response caching
- Yes
- Reuses the answer when the exact same request comes in again, which cuts both cost and latency.
- Similar-question caching Semantic cache
- Yes
- Reuses an answer when a new question means roughly the same thing as an earlier one. Saves far more than exact-match caching but can return subtly wrong answers if tuned loosely.
- Where you set the timeout Request timeout surface
- In config
`config_file`, but the LLM request timeout is the weak spot: ai-proxy's `timeout` (default **120000** ms) is documented as applying only to the context-data retrieval call, not to the model request ([ai-proxy plugin](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-proxy/)). The request timeout you actually want is the Ingress annotation `higress.io/timeout`, in seconds, with **no timeout by default** ([annotations](https://higress.ai/en/docs/latest/user/annotation/)). Connection-level values live in the ConfigMap: downstream `idleTimeout` 180 s, upstream `idleTimeout` 10 s ([global configuration](https://higress.ai/en/docs/latest/user/configmap/)). MCP proxying has its own `server.timeout`, default 5000 ms ([mcp-server plugin](https://higress.ai/en/docs/ai/mcp-server/)).
- Where a request timeout can be set: per request, in a config file, in the vendor dashboard, or nowhere. Nine of the twenty products documented here do not describe a request timeout at all, so the worst case of a hung upstream call is unknowable from the docs.
- Where you set retries Retry policy surface
- In config
Off by default at the AI layer and capped at one extra attempt when switched on; no backoff strategy or jitter is documented anywhere ([ai-proxy retryOnFailure](https://higress.cn/docs/latest/plugins/ai/api-provider/ai-proxy/)). The generic Ingress retry annotation is separate and retries to a different upstream host rather than to a different model ([annotations](https://higress.ai/en/docs/latest/user/annotation/)). No retry keys appear in the global ConfigMap ([global configuration](https://higress.ai/en/docs/latest/user/configmap/)).
- Where retry count and backoff are configured. Worth knowing alongside billing: a retried streaming call can be charged more than once.
- Where you set fallbacks Fallback surface
- Dashboard only
Model-level fallback is a console concept, not a plugin field: the home page states Higress "supports model-level Fallback" ([home page](https://higress.ai/en/)) and the multi-model proxy guide configures it in `AI Route Config`, where a DeepSeek route falls back to Alibaba Cloud `qwen-turbo` when the primary model fails or is rate-limited ([multi-model proxy guide](https://higress.ai/en/docs/ai/scene-guide/multi-proxy/)). The guide shows a single fallback target, not an ordered chain, and no YAML/CRD field for it appears in [ai-proxy (EN)](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-proxy/) or [ai-proxy (CN)](https://higress.cn/docs/latest/plugins/ai/api-provider/ai-proxy/). Alibaba Cloud's own comparison confirms cross-model failover exists in the open-source edition ("多模型间 Failover: 支持") ([Alibaba Cloud comparison](https://help.aliyun.com/zh/api-gateway/ai-gateway/product-overview/product-comparison)).
- Where the fallback chain is defined. Almost every product claims fallback; the useful question is whether you can change it from code or only by hand in a dashboard.
- Shape of the fallback chain Fallback shape
- Single alternate
- An ordered list tries targets in sequence; a weighted split sends a percentage of traffic to each, which is what you need to trial a new model on 5% of requests. Weighted splits are much rarer than the marketing implies.
- Upstream health tracking Health checks / circuit breaking
- In config
Health checking exists only at the API-key level, not the upstream level. ai-proxy `failover` is `enabled: false` by default with `failureThreshold` **3**, `successThreshold` **1**, `healthCheckInterval` **5000** ms, `healthCheckTimeout` **5000** ms, a required `healthCheckModel` and `failoverOnStatus` `["4.*","5.*"]`; an unhealthy `apiToken` is removed from rotation and re-added once probes pass ([ai-proxy, Chinese](https://higress.cn/docs/latest/plugins/ai/api-provider/ai-proxy/)), which the AI quick-start describes in prose as pausing requests on a token until health checks recover it ([AI quick start](https://higress.ai/en/docs/ai/quick-start/)). No circuit-breaker or outlier-detection annotation appears in [the annotation reference](https://higress.ai/en/docs/latest/user/annotation/) and no health-check keys appear in [the global ConfigMap](https://higress.ai/en/docs/latest/user/configmap/) — and Alibaba Cloud's comparison page explicitly marks "故障自动检测及恢复" (automatic fault detection and recovery) as 不支持 for the open-source edition ([Alibaba Cloud comparison](https://help.aliyun.com/zh/api-gateway/ai-gateway/product-overview/product-comparison)).
- Whether the product notices a failing upstream and stops sending traffic to it, and whether you can tune the thresholds. This is what turns a provider outage into a blip rather than a sustained error rate, and only four of the twenty expose it.
- Cross-region failover you control Multi-region failover surface
- Not documented
n.a. — no cross-region or multi-AZ failover configuration on [the Helm guide](https://higress.ai/en/docs/latest/ops/deploy-by-helm/), [the global ConfigMap](https://higress.ai/en/docs/latest/user/configmap/) or [the annotation reference](https://higress.ai/en/docs/latest/user/annotation/). Alibaba Cloud's comparison page lists 多可用区部署 (multi-availability-zone deployment) as 自行构建, i.e. build it yourself, for the open-source edition ([Alibaba Cloud comparison](https://help.aliyun.com/zh/api-gateway/ai-gateway/product-overview/product-comparison)). Container images are mirrored to cn-hangzhou, us-west-1 and ap-southeast-7 registries, but that is image distribution, not traffic failover ([README](https://github.com/higress-group/higress/blob/main/README.md)).
- Whether you can define what happens when a region degrades. A vendor running many regions is not the same as a vendor letting you configure failover between them; only three document a user-controlled mechanism.
- Where you set load balancing Load balancing surface
- In config
Two unrelated layers. Classic HTTP balancing is an Ingress annotation: `nginx.ingress.kubernetes.io/load-balance` defaults to `round_robin` with `least_conn` and `random` also supported, `ewma` is explicitly **not** supported and silently falls back to round-robin, and `upstream-hash-by` gives consistent hashing on `$request_uri`, `$host`, `$remote_addr`, a header or a query arg ([annotation compatibility](https://higress.ai/en/docs/latest/user/annotation/)). Across LLM credentials, ai-proxy simply picks an `apiToken` at random per request ([ai-proxy plugin](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-proxy/)). LLM-aware algorithms — "global least connections", "prefix matching" and "GPU-aware" load balancing — are described in a vendor engineering post rather than in the plugin reference ([Higress engineering post](https://medium.com/@higress_ai/no-increase-in-gpu-the-first-token-latency-decreases-by-50-new-practices-in-llm-service-load-5583192f9442)), and no `ai-load-balancer` page exists in the documentation index ([higress.ai/llms.txt](https://higress.ai/llms.txt)).
- Where traffic distribution across upstreams or keys is configured.
Operations
Interpret these fields: LLM gateway observability: traces, logs and export · Running coding agents through an LLM gateway · Changing models without breaking production
- Usage dashboards and logs Observability
- Yes
- Built-in visibility into what was sent, what came back, what it cost, and how long it took.
- Spending limits Budget controls
- Yes
- Hard caps that stop spend before it becomes a surprise invoice. The single most valuable control for a small team.
- Rate limits Rate limits
- Yes
- Caps on request volume per key or per user, useful for protecting against abuse and runaway loops.
- Separate keys per team or app Virtual keys
- Yes
- Issue scoped keys with their own budgets and permissions so you can attribute cost and revoke access without rotating everything.
- Prompt versioning Prompt management
- Yes
- Store and version prompts outside your code so they can be changed without a deploy.
- Quality testing Evals
- Not published
- Built-in tooling to score model output against test cases, so you can tell whether a model swap made things better or worse.
- MCP support MCP support
- Yes
- Native support for the Model Context Protocol, the emerging standard for connecting models to external tools.
- What gets logged Logged content
- Your choice
Configurable in both directions: `accessLogFormat` under the `mesh` field controls which access-log fields are emitted ([log description](https://higress.ai/en/docs/latest/ops/log/)), and ai-statistics has an explicit "record questions and answers" mode where `attributes` entries with `value_source` of `request_body`, `response_body` or `response_streaming_body` and `apply_to_log: true` write prompt and completion text into the log — both flags default to **false**, so content logging is opt-in ([ai-statistics](https://higress.ai/en/docs/latest/plugins/ai/api-o11y/ai-statistics/)).
- Whether full prompts and responses are stored, only metadata, or nothing. Full-body logging is the most useful debugging feature here and the one most likely to need a conversation with your compliance team.
- You can turn logging off Body-logging opt-out
- Not documented
n.a. as a documented switch: [the log page](https://higress.ai/en/docs/latest/ops/log/) explains how to change `accessLogFormat` and how to view logs but states no way to disable access logging, and [the global ConfigMap](https://higress.ai/en/docs/latest/user/configmap/) and [the annotation reference](https://higress.ai/en/docs/latest/user/annotation/) contain no logging on/off key. In practice a self-hosted operator controls this at the Envoy/ConfigMap level, but Higress does not document it.
- Whether prompt and response bodies can be suppressed while still keeping usage metrics. Nineteen of the twenty document a way to do this; the mechanisms range from a per-request header to an organisation-wide setting.
- Traces you can take elsewhere Distributed tracing
- OpenTelemetry
OpenTelemetry, plus SkyWalking and Zipkin, configured in the global ConfigMap under `tracing`: `enable` defaults to **false**, `sampling` to **100.0**, `timeout` to **500** ms, and only one exporter can take effect at a time; the `opentelemetry` block takes a `service` and a gRPC `port` ([global configuration](https://higress.ai/en/docs/latest/user/configmap/)). Spans are enriched by ai-statistics via `apply_to_span` attributes, and `trace_id` is an access-log field ([ai-statistics](https://higress.ai/en/docs/latest/plugins/ai/api-o11y/ai-statistics/), [log description](https://higress.ai/en/docs/latest/ops/log/)).
- Whether the product emits OpenTelemetry, a proprietary format, or nothing. OpenTelemetry means the traces land in the tooling you already run instead of only in the vendor’s dashboard.
- Where telemetry can go Export destinations
- OpenTelemetry, Prometheus, Grafana, Loki, SkyWalking, Zipkin
Traces go to OpenTelemetry, SkyWalking or Zipkin collectors you name in the ConfigMap ([global configuration](https://higress.ai/en/docs/latest/user/configmap/)); metrics are scraped from the gateway's metrics port (default **15020** in the Docker Compose install) by Prometheus and drawn in Grafana, with `--set global.o11y.enabled=true` installing Grafana, Prometheus, Loki and PromTail for you, or an external Grafana URL pasted into the console instead ([Prometheus guide](https://higress.ai/en/docs/latest/user/prometheus/), [Helm values](https://higress.ai/en/docs/latest/ops/deploy-by-helm/), [Docker Compose options](https://higress.ai/en/docs/latest/ops/deploy-by-docker-compose/)). There is no vendor SaaS sink and no Datadog/LangSmith/S3 integration on any fetched page.
- Documented sinks for logs and metrics. This is a good proxy for how replaceable the vendor’s own dashboard is: around twenty destinations means you never have to depend on it, while a CSV download means you do.
- Can record user feedback Feedback capture API
- No
No feedback, rating or annotation endpoint. ai-statistics can attach arbitrary attributes to logs and spans from request or response fields, which is the closest available mechanism, but nothing captures a user verdict ([ai-statistics](https://higress.ai/en/docs/latest/plugins/ai/api-o11y/ai-statistics/)); nothing appears in [the plugin marketplace](https://higress.ai/en/plugins/) either.
- Whether there is an API to attach a rating or score to a logged request, which is what lets production traffic feed quality work later.
- Scores live traffic Online eval hooks
- No
No evaluation, scoring or experiment feature exists: the 41-plugin marketplace has no eval plugin ([plugin marketplace](https://higress.ai/en/plugins/)) and the documentation index lists none across overview, user guide, ops, developer, AI gateway and scene-guide sections ([higress.ai/llms.txt](https://higress.ai/llms.txt)). Traffic-shaping primitives that could support an offline comparison do exist — canary annotations (`canary-by-header`, `canary-weight`) and 多模型灰度 model-level canary — but no scoring is attached ([annotation reference](https://higress.ai/en/docs/latest/user/annotation/), [Alibaba Cloud comparison](https://help.aliyun.com/zh/api-gateway/ai-gateway/product-overview/product-comparison)).
- Whether automated scorers can run against real production requests, rather than only against a test set you assemble yourself.
Performance
- Delay it adds Proxy overhead
- Not published
- Extra time the product itself adds to each request, on top of however long the model takes. Usually irrelevant next to multi-second model latency, but it matters for high-volume or streaming-sensitive workloads.
- Requests per second ceiling Throughput
- Not published
- Published sustained request rate before the product becomes the bottleneck. Only relevant at genuinely high volume.
- What the request path runs on Architecture class
- Compiled binary
Envoy data plane plus a Go control plane: the project is "built on Istio and Envoy" and extended with Wasm plugins written in Go, Rust or JavaScript ([what-is-Higress](https://higress.ai/en/docs/latest/overview/what-is-higress/)). Repo language bytes are Go 9,245,785, C++ 1,269,908, Rust 193,762, Shell 122,023, Python 88,069 and TypeScript 53,579 ([GitHub languages API](https://api.github.com/repos/higress-group/higress)), i.e. a compiled control plane and a compiled proxy, with plugin code compiled to Wasm rather than interpreted per request.
- The comparable way to talk about latency here. An edge worker, a compiled Go or Rust binary, and a Python proxy have different overhead floors no matter which figures each vendor publishes. Products that never disclose their runtime are recorded as undisclosed rather than assumed.
- You can run the request path yourself Self-hostable data plane
- Yes
Container images and a Helm chart, both published: `higress-registry.cn-hangzhou.cr.aliyuncs.com/higress/all-in-one:latest` runs console, HTTP and HTTPS on 8001/8080/8443 from one `docker run` ([README](https://github.com/higress-group/higress/blob/main/README.md)), the chart is `higress.io/higress` installed into `higress-system` ([README](https://github.com/higress-group/higress/blob/main/README.md), [Helm guide](https://higress.ai/en/docs/latest/ops/deploy-by-helm/)), and a Docker Compose bundle of apiserver/controller/pilot/gateway/console is generated by an installer script for non-Kubernetes hosts ([Docker Compose guide](https://higress.ai/en/docs/latest/ops/deploy-by-docker-compose/)). Alibaba Cloud Compute Nest deployment is also documented ([documentation index](https://higress.ai/llms.txt)). The FAQ notes there is no pre-built package beyond the images: "There is no existing one, you need to build it yourself. Currently, all Docker images are provided and can be pulled and used by yourself." ([FAQ](https://higress.ai/en/docs/latest/overview/faq/))
- Whether the component that actually carries your prompts can run on your own infrastructure. Distinct from a vendor offering a self-hosted control plane while still proxying traffic through their network.
- Streaming responses Streaming support
- Yes
Yes, and it is a headline feature: Higress claims "true streaming processing" for SSE and supports the AI streaming (SSE) scenario ([what-is-Higress](https://higress.ai/en/docs/latest/overview/what-is-higress/)), and ai-cache handles both streaming and non-streaming responses ([ai-cache](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-cache/)). Two documented streaming caveats matter: retries only apply to non-streaming requests ([ai-proxy retryOnFailure](https://higress.cn/docs/latest/plugins/ai/api-provider/ai-proxy/)), and on SSE the ai-data-masking plugin may fail to restore a masked word split across chunks and may leak part of a sensitive word to the client ([ai-data-masking](https://higress.ai/en/docs/latest/user/plugins/ai/api-consumer/ai-data-masking/)).
- Whether token-by-token streaming is documented. The caveats matter more than the yes: some products cannot cancel a stream without still being billed, and several timeout and fallback mechanisms stop applying once the first token has been sent.
Security & compliance
Interpret these fields: LLM gateway compliance: SOC 2, HIPAA and evidence · How LLM gateway guardrails fail · Which LLM gateways store your prompts? · Do you need an MCP gateway as well?
- Does your prompt reach their servers Prompt transits vendor
- No
No Higress-operated hop exists: the gateway is software you deploy, and requests go from your gateway straight to the provider endpoints configured in ai-proxy ([README](https://github.com/higress-group/higress/blob/main/README.md), [ai-proxy plugin](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-proxy/)). Two things do leave your perimeter if you enable them — ai-security-guard sends prompt and completion text to Alibaba Cloud's content-moderation API ([content security guide](https://higress.ai/en/docs/ai/scene-guide/application-safety/)), and semantic caching sends prompt text to whichever embedding service and vector database you configure, the worked example being Alibaba Cloud BaiLian plus DashVector ([semantic cache guide](https://higress.ai/en/docs/ai/scene-guide/semantic-cache/)).
- Whether the text you send passes through this company’s own infrastructure. If it does, every other promise on this page is a policy commitment rather than a physical impossibility. Self-hosted products can answer no outright.
- What they keep if you change nothing Logging default
- Metadata only, not content
Access logs are on and JSON-formatted by default, and they are metadata: the 22 documented fields are `authority`, `bytes_received`, `bytes_sent`, `duration`, `method`, `path`, `request_id`, `response_code`, `response_flags`, `route_name`, `trace_id`, `upstream_host`, `upstream_service_time`, `user_agent`, `x_forwarded_for` and similar — request and response bodies are represented only by their byte counts ([log description](https://higress.ai/en/docs/latest/ops/log/)). Token counts, model name, `llm_service_duration` and `llm_first_token_duration` are added by the ai-statistics plugin ([ai-statistics](https://higress.ai/en/docs/latest/plugins/ai/api-o11y/ai-statistics/)).
- Defaults matter more than options. A product that stores full prompts and replies unless you find the right header will have stored them by the time you read the docs.
- How long they keep it Default content retention (days)
- Not published
n.a. as a vendor-set value — Higress is self-hosted software with no vendor-side store. The bundled `global.o11y` stack installs Grafana, Prometheus, Loki and PromTail into your cluster and you own the retention settings ([Helm values](https://higress.ai/en/docs/latest/ops/deploy-by-helm/)); ai-cache entries live in a Redis you run, with `cacheTTL` defaulting to 0 meaning never expire ([ai-cache plugin](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-cache/)).
- Default retention for request content, in days. Zero means nothing is kept. Read the note: several products keep nothing as a rule but make timed exceptions for abuse review or specific models.
- Could they train on your prompts Training on customer data
- Not applicable
not_applicable by construction rather than by policy: Higress is self-hosted Apache-2.0 software with no vendor telemetry endpoint on any fetched page ([GitHub API](https://api.github.com/repos/higress-group/higress), [README](https://github.com/higress-group/higress/blob/main/README.md)), so there is no Higress-side corpus to train on. No separate training or data-use policy is published — [the documentation index](https://higress.ai/llms.txt), [the FAQ](https://higress.ai/en/docs/latest/overview/faq/) and [the home page](https://higress.ai/en/) contain none.
- Whether the vendor may use your prompts and outputs to train models. “Not published” means we could not find any position, which is not the same as a no — ask for it in writing.
- Where it runs, and what you can pin Region and residency control
- You choose, because you run it. There is no vendor region list; placement is wherever your Kubernetes cluster or host sits ([Helm guide](https://higress.ai/en/docs/latest/ops/deploy-by-helm/), [Docker Compose guide](https://higress.ai/en/docs/latest/ops/deploy-by-docker-compose/)). The only region-shaped facts published are container-registry mirrors in cn-hangzhou, us-west-1 and ap-southeast-7 ([README](https://github.com/higress-group/higress/blob/main/README.md)) and Alibaba Cloud's note that multi-AZ deployment is "build it yourself" for the open-source edition ([Alibaba Cloud comparison](https://help.aliyun.com/zh/api-gateway/ai-gateway/product-overview/product-comparison)).
- Which regions are offered and whether you can force processing to stay in one. A global endpoint that silently picks a region is a different compliance story from an endpoint you pin yourself.
- Where safety filters run Guardrail execution location
- Either, depending on deployment
Enforcement is always in your own gateway process, but the classifier can be remote: ai-security-guard delegates to Alibaba Cloud's content-moderation API, which requires activating that Alibaba Cloud service ([content security guide](https://higress.ai/en/docs/ai/scene-guide/application-safety/)), whereas qwen3guard calls a Qwen3Guard-Gen model you can host yourself behind an OpenAI-compatible endpoint via vLLM or SGLang ([Higress content-security post](https://higress.ai/blog/higress-mmse_awbbpb_yafyyrc3t5wh0u55/)), and ai-data-masking runs entirely locally in Wasm ([ai-data-masking](https://higress.ai/en/docs/latest/user/plugins/ai/api-consumer/ai-data-masking/)).
- A filter that strips personal data only helps if it runs before the data leaves your boundary. If guardrails execute in the vendor’s cloud, the vendor has already received whatever you wanted redacted.
- Who else touches the data Subprocessor list
- Not published
- The published list of third parties the vendor passes your data to. No list means you cannot know the full chain, which most data-protection agreements require you to.
- SOC 2 audited SOC 2 audited
- Not published
- An independent audit of security controls. Enterprise buyers and their procurement teams routinely require it.
- Will sign a HIPAA agreement HIPAA BAA
- Not published
- Required before you may send protected health information through the service. Without a signed BAA, healthcare data is off limits.
- GDPR commitments GDPR commitments
- Not published
- Published data processing terms for handling personal data of people in the EU and UK.
- Can keep data in the EU EU data residency
- Not published
- Requests can be processed inside the EU rather than routed to US infrastructure. Often the deciding constraint for European customers.
- Does not retain your data Zero data retention
- Not applicable
- Prompts and responses are not stored after the request completes. Sometimes a paid add-on rather than the default.
- Strips personal data PII redaction
- Yes
- Detects and removes identifiers such as names, emails, and card numbers before the request reaches the model provider.
- Content guardrails Content guardrails
- Yes
- Policy checks on inputs and outputs — blocking unsafe content, enforcing formats, or catching prompt-injection attempts.
- Runs fully disconnected Air-gapped deployment
- Not published
- Can be deployed in a network with no internet access, which some regulated and defence environments require.
- Blocks personal data in prompts PII / DLP enforcement
- Can block the request
`sync_block` by default when the plugin is attached: ai-data-masking's `system_deny` (built-in sensitive-word rules, sourced from `github.com/houbb/sensitive-word`) and `deny_openai` both default to **true**, and `deny_code` defaults to **200** with the message "Sensitive words found in the question or answer have been blocked". A softer mode exists — `replace_roles` with regex/GROK patterns and `type: replace` or `hash`, plus `restore: true` to put the original values back into the model's answer ([ai-data-masking](https://higress.ai/en/docs/latest/user/plugins/ai/api-consumer/ai-data-masking/)). Documented failure mode on SSE: a masked word split across chunks may not be restored, and part of a sensitive word can reach the user.
- Whether personal data detection sits on the request path and can stop the call, merely inspects and forwards it, or is not documented. A control that only reports is a logging feature, not a policy control.
- Blocks prompt injection Injection / jailbreak enforcement
- Can block the request
`sync_block`, but only through the newer `qwen3guard` plugin, and only as a by-product: the vendor engineering post states Qwen3Guard's official safety policy covers violence, illegal acts, sexual content, PII, self-harm, unethical behaviour, politically sensitive topics and copyright infringement "并在输入审核中包含 Jailbreak 检测" (jailbreak detection is included in input moderation). The plugin enables `checkRequest` and `checkResponse` by default, calls `Qwen/Qwen3Guard-Gen-4B`, uses `riskLevelBar: Unsafe` by default, and when the bar is met returns a refusal without calling the model at all ([Higress content-security post](https://higress.ai/blog/higress-mmse_awbbpb_yafyyrc3t5wh0u55/)). Caveats stated by the vendor: only the `Safety` verdict drives the decision, per-category actions are not implemented, and streaming interception cannot honour `denyCode`. No prompt-injection or jailbreak wording appears in the older [ai-security-guard](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-security-guard/) reference.
- Whether injection and jailbreak detection can stop a request. Most products offering this call a partner classifier rather than shipping their own.
- Blocks harmful content Toxicity / moderation enforcement
- Can block the request
`sync_block` via ai-security-guard, which calls Alibaba Cloud's content-moderation service inline for both directions and denies the request when a risk label is returned; `checkRequest` and `checkResponse` both default to **false**, `requestCheckService` defaults to `llm_query_moderation`, `responseCheckService` to `llm_response_moderation`, and `denyCode` defaults to **200** — so a blocked call looks like a success to a naive client ([ai-security-guard](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-security-guard/)). It emits `ai_sec_request_deny`/`ai_sec_response_deny` metrics and `ai_sec_risklabel`/`ai_sec_deny_phase` span attributes. The console walkthrough registers Alibaba Cloud Content Safety as a service source and attaches "AI Safety Guard" to a route ([content security guide](https://higress.ai/en/docs/ai/scene-guide/application-safety/)).
- Whether hate, violence, sexual and self-harm categories are checked inline and can stop a request, in either direction.
- Your own policy rules Custom policy hooks
- Can block the request
Custom policy is regex/GROK-based: ai-data-masking takes user-supplied `deny_words` and `replace_roles` regex patterns with per-rule `replace`/`hash` actions ([ai-data-masking](https://higress.ai/en/docs/latest/user/plugins/ai/api-consumer/ai-data-masking/)), and the generic `request-block` and `request-validation` plugins add path/body blocking and JSON-schema request validation ([plugin marketplace](https://higress.ai/en/plugins/)). Custom rules you add yourself default to replacement rather than denial unless you put them in the deny list, hence `sync_observe` as the default posture for custom rules specifically.
- Whether you can add your own rule — a regex, a webhook, or your own classifier — rather than choosing from the vendor library.
- Where guardrails run Guardrail execution location
- Either, your choice
- Whether guardrail evaluation happens inside your infrastructure or on the vendor’s servers. This decides whether the prompt you are trying to protect leaves your network in order to be checked.
- If the guardrail itself fails Guardrail failure mode
- Request proceeds
`fail_open`, stated explicitly and unusually candidly for the qwen3guard plugin: "当前插件选择 fail-open:记录警告并放行" — on connection failure, `timeoutMs` (default 2000 ms) expiry, a non-200 from the guard model, unparseable JSON or a missing `Safety` field, the request is logged and allowed through. The same post says compliance scenarios needing forced fail-close "不能被描述为已经满足" (cannot be described as satisfied) in the current version ([Higress content-security post](https://higress.ai/blog/higress-mmse_awbbpb_yafyyrc3t5wh0u55/)). No fail-open/fail-closed statement appears in the older [ai-security-guard](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-security-guard/) reference.
- What happens when the guardrail service times out or errors: does the request proceed unchecked, or is it blocked? This is the worst-documented field in the entire catalogue — only two vendors state it plainly, which means most teams are running a control whose failure behaviour they cannot know.
- Third-party guardrail vendors Guardrail integrations
- Alibaba Cloud Content Moderation, Qwen3Guard
- Named external guardrail services the product can call. A long list means the product is a router for policy engines rather than a policy engine itself — which also means another vendor bill and another hop.
Compliance evidence
Graded by how strong the evidence is, not whether the word appears on the vendor’s website. An audited report and a marketing claim are different things, and only one of them will satisfy your own auditor.
- SOC 2 Not published No SOC 2 claim on higress.ai, the documentation index or the CNCF project page. Higress is self-hosted software; certification would attach to your own deployment.
- ISO 27001 Not published Not mentioned on any fetched Higress or CNCF page.
- GDPR DPA Not published No DPA, privacy policy or GDPR statement on the fetched pages; there is no vendor data processor to contract with.
- HIPAA BAA Not published Not mentioned on any fetched page.
- FedRAMP Not published Not mentioned on any fetched page.
- ITAR Not published Not mentioned on any fetched page.
No compliance certifications were found published for this product. That is not the same as failing an audit — it means there is nothing public to check, so ask for evidence directly.
Fit & integration
Interpret these fields: How much does an LLM gateway lock you in? · Do you need an MCP gateway as well? · Running coding agents through an LLM gateway · Changing models without breaking production
- Work to try it Evaluation work shape
- Run something locally first
- The shape of the work on the vendor’s own quickstart, from swapping one base URL through to deploying infrastructure. An ordinal class rather than a duration, because elapsed time depends on accounts and quota we cannot see.
- Work to run it Production work shape
- Deploy it on your infrastructure
- The same scale applied to the vendor’s recommended production path. For several products this is much heavier than the quickstart, which is exactly why both are recorded.
- Steps on the quickstart Numbered quickstart steps
- 0
No numbered step list exists to count. The AI gateway quick start is organised as screenshot-driven sections (install by Docker, log in to the console, configure an LLM provider key, configure an AI route, observe the AI dashboard) rather than enumerated steps ([AI quick start](https://higress.ai/en/docs/ai/quick-start/)), and the Kubernetes quick start is structured as three stages — install, configure, validate — with branching environment options rather than a linear count ([quick start](https://higress.ai/en/docs/latest/user/quickstart/)). Recorded as 0 to mean "vendor publishes no step count", not "zero work".
- A literal count of numbered steps on the vendor’s quickstart, recorded as evidence beside the work shape. Zero means the page publishes no numbered procedure at all. Large counts usually mean interleaved language tracks rather than more work.
- Can you self-host it today Self-host install documentation
- Install command published
One container: `docker run -d --rm --name higress-ai -v ${PWD}:/data -p 8001:8001 -p 8080:8080 -p 8443:8443 higress-registry.cn-hangzhou.cr.aliyuncs.com/higress/all-in-one:latest` (console 8001, HTTP 8080, HTTPS 8443). Kubernetes: `helm install higress -n higress-system higress.io/higress --create-namespace`. Non-Kubernetes AI gateway installer: `wget https://higress.cn/ai-gateway/install.sh` then run with bash.
- Whether an install command is actually published. Several products advertise self-hosting while publishing no command to start from, which a plain yes/no would hide.
- Works with the OpenAI SDK OpenAI SDK drop-in
- Yes
Yes: point an OpenAI-format client at the gateway route and ai-proxy converts to whichever provider the route is bound to, because the plugin "implements AI proxy functionality based on OpenAI API contract" and detects the protocol from `/v1/chat/completions` ([ai-proxy plugin](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-proxy/)). Requests are authenticated with your own consumer credential in `x-api-key` rather than a provider key ([token management guide](https://higress.ai/en/docs/ai/scene-guide/token-management/)). No SDK change is needed; `protocol: original` is available when you would rather not be converted at all.
- Whether an existing OpenAI-compatible client can be pointed at it by changing the base URL and key.
- Vercel AI SDK support AI SDK provider package
- Not documented
n.a. — the Vercel AI SDK is not mentioned on [the documentation index](https://higress.ai/llms.txt), [the ai-proxy reference](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-proxy/) or [the AI gateway page](https://higress.ai/en/ai-gateway/). In practice `@ai-sdk/openai` with a custom `baseURL` should work against the OpenAI-compatible route ([ai-proxy plugin](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-proxy/)), but Higress documents no such integration, so this is recorded as not documented rather than as compatibility.
- Whether a first-party AI SDK provider package exists, or only a community package, a documented workaround, or the generic OpenAI provider pointed at a custom base URL.
- Python framework integrations Documented Python frameworks
- Not published
n.a. — neither LangChain nor LlamaIndex is documented anywhere in the official docs index or the plugin marketplace ([higress.ai/llms.txt](https://higress.ai/llms.txt), [plugin marketplace](https://higress.ai/en/plugins/)); a LangChain + Higress + Elasticsearch RAG walkthrough exists only as a third-party blog post, not vendor documentation. RAG is instead offered as a gateway plugin (ai-rag) and via ai-search ([plugin marketplace](https://higress.ai/en/plugins/)).
- LangChain, LangGraph, LlamaIndex and similar orchestration frameworks with a documented integration.
- Callable from Cloudflare Workers Cloudflare Workers support
- Workers AI as an upstream model
Cloudflare appears only as a model source: ai-proxy has a `cloudflare` provider type for Cloudflare Workers AI, configured with `cloudflareAccountId` ([ai-proxy plugin](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-proxy/)). Higress itself is a self-hosted Envoy process and does not run on Workers ([what-is-Higress](https://higress.ai/en/docs/latest/overview/what-is-higress/)).
- Whether the docs show calling this product from your own Worker. Deliberately separated from the several products whose own gateway runs on Workers, which is a fact about their infrastructure and not about your edge compatibility.
- Kubernetes install Helm chart availability
- Official Helm chart
Kubernetes is the first-class target: `helm repo add higress.io https://higress.io/helm-charts` then `helm install higress -n higress-system higress.io/higress --create-namespace`, with `global.local=true` for a kind cluster and `hgctl` to open the console ([Helm guide](https://higress.ai/en/docs/latest/ops/deploy-by-helm/), [quick start](https://higress.ai/en/docs/latest/user/quickstart/)). The README states Higress is a conformant Gateway API implementation, a conformant Gateway API Inference Extension implementation, and is listed in the official Kubernetes Ingress Controllers documentation ([README](https://github.com/higress-group/higress/blob/main/README.md)). It is also nginx-Ingress-annotation compatible, which is its main migration story ([annotation compatibility](https://higress.ai/en/docs/latest/user/annotation/)).
- Whether a named, published Helm chart exists, versus Helm being referenced with no chart named, versus generic cluster documentation that has nothing to do with this product.
- Terraform support Terraform provider or modules
- Not documented
n.a. — no Terraform provider, module or registry reference on [the documentation index](https://higress.ai/llms.txt), [the Helm guide](https://higress.ai/en/docs/latest/ops/deploy-by-helm/) or [the Docker Compose guide](https://higress.ai/en/docs/latest/ops/deploy-by-docker-compose/). Infrastructure-as-code is expressed instead as Helm values and Kubernetes Ingress/ConfigMap objects, plus an Alibaba Cloud Compute Nest deployment path ([documentation index](https://higress.ai/llms.txt)).
- Whether you can declare this in code: an official provider, official modules, resources inside a hyperscaler’s provider, a community provider, or only Terraform code shipped in a repo.
- Reuses your cloud identity Cloud IAM reuse
- Static provider credentials only
Cloud credentials are passed as static plugin fields, not assumed roles: AWS Bedrock takes `awsAccessKey`, `awsSecretKey` and `awsRegion`, and Google Vertex AI takes `vertexAuthKey`, `vertexRegion`, `vertexProjectId` and `vertexAuthServiceName` ([ai-proxy plugin](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-proxy/)). No IRSA, workload identity, instance-profile or credential-chain support is documented. IDaaS/OIDC integration for the control plane is listed as a commercial-edition feature (深度集成阿里云 IDAAS 产品) versus "build it yourself" for open source ([Alibaba Cloud comparison](https://help.aliyun.com/zh/api-gateway/ai-gateway/product-overview/product-comparison)), although an OIDC plugin exists for data-plane auth ([plugin marketplace](https://higress.ai/en/plugins/)).
- Whether you can authenticate with IAM roles, workload identity or managed identities instead of another long-lived API key. Distinguished from products that only accept static upstream provider credentials.
- Fits behind your API gateway API gateway integration
- It is the API gateway
It is the gateway: Higress is described as an AI-native API gateway built on Istio and Envoy that unifies AI gateway, Kubernetes Ingress gateway, microservice gateway and security gateway roles in one deployment ([what-is-Higress](https://higress.ai/en/docs/latest/overview/what-is-higress/)), and the home page splits the AI side into LLM Gateway, MCP Gateway and Model Gateway (the last implementing the Gateway API Inference Extension with InferencePool and Endpoint Picker) ([home page](https://higress.ai/en/)).
- Whether AI traffic can go through a gateway you already run, and who documented that — the product’s vendor or the gateway’s.
- MCP support MCP surface shape
- MCP gateway or registry
A real MCP gateway, not just a client. The `mcp-server` plugin both converts REST APIs into MCP tools with no code (`server.type: rest`, `tools[].requestTemplate`) and proxies existing MCP servers (`server.type: mcp-proxy` with `mcpServerURL`, `server.timeout` default 5000 ms), with `allowTools` allowlisting, `securitySchemes` plus `defaultDownstreamSecurity`/`defaultUpstreamSecurity`, and `passthroughAuthHeader` defaulting to **false** so client credentials are not leaked to backends ([mcp-server plugin](https://higress.ai/en/docs/ai/mcp-server/)). Hosting requires Higress >= 2.1.0 and Redis; the Nacos MCP registry needs Nacos >= 3.0 and Higress >= 2.1.2; SSE endpoints are exposed at `/{mcp route prefix}/{nacos name}/sse` and the 2025-03-26 streamable-HTTP protocol needs no ConfigMap entry ([MCP quick start](https://higress.ai/en/docs/ai/mcp-quick-start/)). The home page announces support for the 2026-07-28 MCP revision "with compatibility for existing protocols" ([home page](https://higress.ai/en/)), an `openapi-to-mcp` conversion tool and a hosted demo at https://mcp.higress.ai/ ([README](https://github.com/higress-group/higress/blob/main/README.md)).
- Which kind of MCP support this is: a gateway that governs many MCP servers, a hosted MCP server you connect to, MCP tools accepted inside the completion API, or client tooling. These are different products behind one acronym.
- Needs your own provider key Upstream provider key required
- Required
Yes — you cannot make a first call without an upstream key. The installer prompts "enter the Aliyun Dashscope or other API-KEY" (you may skip and add it in the console instead) and "LLM Provider Management" is where keys for Alibaba Cloud, DeepSeek, Azure OpenAI, OpenAI, DouBao and others are stored ([AI quick start](https://higress.ai/en/docs/ai/quick-start/)); every ai-proxy provider block takes `apiTokens` or cloud credentials ([ai-proxy plugin](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-proxy/)).
- Whether an upstream provider account and key must exist before your first call works. A prerequisite rather than a step, and it can differ between the hosted and self-hosted forms of the same product.
- Gate before models work Model access gate
- No gate
No gate: you add a provider block with your own `apiTokens` and the model is immediately routable, with `modelMapping` deciding which names clients may use ([ai-proxy plugin](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-proxy/)); the console equivalent is adding a key under LLM Provider Management ([AI quick start](https://higress.ai/en/docs/ai/quick-start/)). No approval, enablement, quota or waitlist step appears on either page.
- Whether an enablement click, a quota grant, a paid tier or an approval form stands between a valid key and a working model call.
- Official client languages First-party client SDK languages
- Not published
n.a. — no Higress client library exists on [the ai-proxy reference](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-proxy/), [the documentation index](https://higress.ai/llms.txt) or [the AI quick start](https://higress.ai/en/docs/ai/quick-start/); the documented calling convention is `curl` or any OpenAI-compatible SDK against the gateway route. The Go, Rust and JavaScript SDKs Higress does publish are for authoring Wasm gateway plugins, not for calling the gateway ([what-is-Higress](https://higress.ai/en/docs/latest/overview/what-is-higress/)), so no client-language list is recorded here.
- Languages with a first-party client library. An empty list can still mean the product is usable from any language via an OpenAI-compatible SDK.
How pricing actually works
Your own infrastructure only — no licence, no seat fee, no metered tier. Sizing signals from the docs: gateway defaults to 2 replicas, the console to 1, and the optional o11y stack adds Grafana, Prometheus, Loki and PromTail to the same cluster ([Helm values](https://higress.ai/en/docs/latest/ops/deploy-by-helm/)). Redis is a hard dependency for ai-cache, ai-token-ratelimit, ai-quota and MCP hosting, so budget for it ([ai-cache](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-cache/), [ai-token-ratelimit](https://higress.ai/en/docs/latest/plugins/ai/api-consumer/ai-token-ratelimit/), [MCP quick start](https://higress.ai/en/docs/ai/mcp-quick-start/)). Autoscaling is your job: "Higress is based on K8s HPA and supports elastic scaling. The gateway is stateless and is a deployment." ([FAQ](https://higress.ai/en/docs/latest/overview/faq/))
Common questions
Answered from the fields above, so these move when the catalog moves. Every figure quoted here appears in the specification with its source.
Does Higress charge a markup on model prices?
Higress adds no percentage markup to model prices. The cost estimator on this site itemises these mechanisms against your own volume, because which one is cheapest depends entirely on the numbers you put in. With no per-token cut and no top-up charge, what you pay is the model providers' own rates plus whatever the infrastructure costs you to run.
Can Higress be self-hosted?
Yes. Higress can be run on your own infrastructure or used as a managed service. The licence is Apache-2.0. Running it yourself means you supply the infrastructure and the upstream model accounts, so the bill is your own hosting plus the providers' own rates.
Is Higress SOC 2 audited, and will it sign a HIPAA BAA?
Higress publishes neither a SOC 2 report nor a HIPAA business associate agreement. Each of these is linked to the vendor's own page in the compliance section below. Neither absence means a refusal: both are things a vendor either publishes or does not, and smaller products often hold the certification without advertising it.
Does Higress retain your prompts?
Zero data retention does not apply to Higress: it runs inside your own infrastructure, so prompts never reach a vendor. Whether prompt and response bodies are logged is configurable. Retention becomes your own configuration question instead, decided by whatever logging you switch on in your own deployment.
Can you use your own provider keys with Higress?
Yes. Higress can route through your own accounts with the underlying model providers, so inference is billed to you directly. No fee of any kind: the software is Apache-2.0 and nothing meters your traffic, so BYOK is the only mode and it costs nothing beyond your provider bills and your own compute ([GitHub API](https://api.github.com/repos/higress-group/higress), [ai-proxy plugin](https://higress.ai/en/docs/latest/plugins/ai/api-provider/ai-proxy/)).
How many models does Higress support?
Higress states ~100 models, drawn from 26–31 upstream providers. Vendor-stated "Supports 100+ LLM models" on the AI gateway product page ([AI gateway page](https://higress.ai/en/ai-gateway/)) and "unified protocol conversion for 100+ common models" on the home page ([home page](https://higress.ai/en/)) and the multi-model proxy guide ([multi-model proxy](https://higress.ai/en/docs/ai/scene-guide/multi-proxy/)). No enumerated model list is published, so 100 is a floor rather than a count. The figure on this page is dated and carries its source.
Official links
- Website higress.ai ↗
- Documentation higress.ai ↗
- Pricing higress.ai ↗
- Source code github.com ↗
- Changelog github.com ↗
9,458 GitHub stars — a proxy for community size, not for quality.
Independent coverage
Third-party analysis, walkthroughs and operator threads. We link criticism as readily as praise. Nothing here is written or published by the vendor, by a competitor listed on this site, or by an SEO content farm — how we vet these.
Nothing yet that clears our bar. Everything we found was written by the vendor, by a competitor, or by an SEO content farm, so we would rather show you nothing than pass marketing off as a review.
What has changed here
- GitHub stars GitHub stars 9273 9458 source ↗
- catalog entry catalog entry Not published Added to the catalog source ↗
Read the head-to-head
These pairs have a written verdict, not just a table.