agentgateway Open source
agentgateway is an open-source LLM gateway: an OpenAI-compatible API in front of ~1,002 models from 20–44 providers. It charges no token markup or per-seat fee. It runs only self-hosted under Apache-2.0. It does not publish a HIPAA BAA. You can point it at your own provider accounts. Beyond chat it also serves embeddings. It handles failover, load balancing, guardrails and request logging. Published throughput is 35,502 requests per second.
· 45 dated entries · 64 source references
Access at a glance
Whether this product can work for you at all, before features matter: who pays the model bill, where it can run, and which of your existing API calls keep working. Every value is the vendor’s own claim, linked to the page it came from.
Positioned as "an AI-first, open-source, cloud-native gateway control plane and proxy data plane" and simultaneously as a general-purpose HTTP/gRPC data plane with load balancing, timeouts, retries, TLS, rate limits and authorization, so that you do not run separate "regular" and "AI" gateways. Three traffic classes are first-class: LLM inference, MCP tool servers and A2A agent traffic (introduction, FAQs).
Who pays the model bill
Your keys onlyYou contract with each model provider directly and hold those accounts. The gateway never resells inference.
There are no platform credits and no project-issued upstream keys. Credentials are supplied inline, as $ENV_VAR, from a file-backed env var, as a Kubernetes secret, or passed straight through from the caller's own token; virtual keys are an internal authorisation layer in front of those upstream credentials, not a billing relationship (API keys, virtual keys).
Merchant of record: n.a. - your model providers invoice you directly and the project invoices nothing. The gateway's own cost figures are explicitly "best-effort and may not exactly match your provider bill" (model costs).
Key handling: Upstream keys are never stored by a vendor. Options are inline literals, $ENV_VAR indirection, a file whose contents become an env var, a Kubernetes secret reference, or full passthrough of the caller's token. Virtual keys sit in the config file or, in hybrid storage mode, in your own SQLite/PostgreSQL database, and can be created or revoked through the config resource API. Sharp edge: that API is served on the admin address 127.0.0.1:15000 with no authentication - anyone who can reach it can list, create or delete keys (API keys, config resources).
Where it can run
1 of 5 shapes documented- Vendor-hosted
- Self-host
- Your VPC
- On-premise
- Air-gapped
Dimmed shapes are not documented by the vendor, which is not the same as unsupported.
self_host only, in two shapes. Standalone: one binary, Docker image or the single agentgateway-standalone Helm chart, driven by a config file, with an editable UI and ClickOps. Kubernetes: adds a control plane, two Helm charts, Gateway API custom resources, xDS dynamic config, a read-only UI and a GitOps workflow. No SaaS and no hybrid vendor plane exist (introduction, Kubernetes control plane).
Install paths documented: curl -sL https://agentgateway.dev/install | bash (pin with -s -- --version v1.5.0), docker run cr.agentgateway.dev/agentgateway:v1.5.0, Docker Compose, and helm upgrade -i agentgateway-standalone oci://cr.agentgateway.dev/charts/agentgateway-standalone --version v1.5.0. The proxy listens on 4000, the UI and config API on 15000, metrics on 15020 (binary, Docker, Helm).
API surfaces your code can keep using
4 of 7 documented, 1 partial- OpenAI chat
POST /v1/chat/completionsYesThe
completionsroute type servesPOST /v1/chat/completions, and the documented drop-in isopenai.OpenAI(api_key="anything", base_url="http://localhost:4000/v1")(Completions). - Anthropic messages
POST /v1/messagesYesThe
messagesroute type serves/v1/messages, with a companionanthropicTokenCounttype for/v1/messages/count_tokens(Messages, token count). - OpenAI Responses
POST /v1/responsesYesThe
responsesroute type serves/v1/responses, with documented conversion caveats worth knowing -stop_sequencesandtop_kare silently dropped when converting, and unsupported block types return400 unsupported conversion(Responses). - Embeddings
POST /v1/embeddingsYesAn
embeddingsroute type serves/v1/embeddings, and a separatereranktype serves/v2/rerankand/v1/rerank(Embeddings, Rerank). - Images
POST /v1/images/generationsNot documented *n.a. as a gateway route type: no
/v1/imagesor image-generation route appears among the documented API types (LLM overview, Completions, passthrough). Image models such asgpt-image-1anddall-e-3do appear in the project's provider cookbook model lists, andpassthroughwould forward such a call opaquely, but no image endpoint is translated or guarded (Model and Provider Cookbook). - Audio
POST /v1/audio/*Partly *partial: a
realtimeroute type proxies OpenAI's/v1/realtimeWebSocket, but every documented example usesmodalities: ["text"]and no audio modality, transcription or TTS path is shown. There is no/v1/audio/*route type. The Realtime page also lists real limits: prompt guards, prompt enrichment and body-based rate limiting do not apply to WebSocket traffic (Realtime). - Batch jobs
POST /v1/batchesNot documentedn.a. (no batch or async bulk endpoint on any page fetched: LLM overview, Completions, passthrough).
An asterisk marks a qualified verdict: support that is indirect (SDK compatibility or provider passthrough rather than a native endpoint), or a gap that is narrower or wider than the label suggests. Read the note before porting.
Route types are declared per route, so one gateway can expose several client dialects at once: completions, responses, messages, anthropicTokenCount, embeddings, models, rerank, realtime, native Gemini (models/{model}:generateContent, :streamGenerateContent, :countTokens) and passthrough. passthrough has two sub-modes worth understanding: opaque forwards the body without interpretation and without guardrails, while detect parses enough for telemetry. A separate config resource API on the admin port manages keys, models, routes and policies at runtime (LLM overview, passthrough, config resources).
How much it reaches
Providers: The vendor publishes different totals on different pages; both bounds are shown.
One project-published figure: "1002+ Models" on the Model and Provider Cookbook, alongside "44+ LLM Gateway Providers" and "20 API Endpoints". No page publishes a total elsewhere, and the docs never claim a model count, so this is a single vendor headline rather than two corroborating sources (Model and Provider Cookbook).
Two project-published figures, at different scopes: the LLM overview documents 20 natively supported providers with a per-provider capability matrix (OpenAI, Anthropic, Bedrock, Azure, Gemini, Vertex AI, Copilot, Cohere, Ollama, Baseten, Cerebras, Deepinfra, Deepseek, Groq, Hugging Face, Mistral, OpenRouter, Together AI, xAI, Fireworks), while the Model and Provider Cookbook headlines 44+ once OpenAI-compatible, enterprise/regional and local providers are counted (LLM overview, Model and Provider Cookbook).
Whose models: All third-party routed: the project hosts no models and owns no inference capacity, proxying instead to 20 natively supported providers plus self-hosted runtimes (Ollama, vLLM, LM Studio) and anything OpenAI-compatible via the custom provider (LLM overview, custom providers).
Your own endpoints: Yes: a documented custom provider takes an arbitrary host, port, path and auth for any OpenAI-compatible endpoint, and self-hosted runtimes Ollama, vLLM and LM Studio have their own pages (custom providers, Ollama).
How it behaves in production
What happens when an upstream model is slow, wrong, or down — and what you can see and stop while it happens. Reliability features are recorded as where you configure them, not whether the vendor lists them, because almost every product here lists all of them.
5 of 6 reachable from code 4 of 4 can block 10 documented destinations
When something goes wrong
Each row says where the knob is, not whether the feature is on the marketing page. A control you can only reach by hand in someone else’s dashboard cannot be reviewed or version-controlled.
- Request timeout In config
Changing it means editing configuration and shipping it, so behaviour is uniform across traffic until you redeploy.
config_file, at four levels: routerequestTimeout(whole client request including retries) andbackendRequestTimeout(each individual try), backendhttp.requestTimeout, andtcp.connectTimeoutfor connection establishment. No per-request header or query override is documented (Timeouts). - Retries In config
retry.attempts,retry.backoff(example500ms) andretry.codes(example[429, 500, 503]). A retry prefers a different backend from the one that just failed. Caveat: retries are disabled once a request body exceeds the buffering threshold, because the body can no longer be replayed (Retries). - Fallback to another model In config
ORDERED and WEIGHTED both available, through
llm.virtualModels[](v1.3+):routing.weighteddistributes by weight,routing.failoveruses priority groups,routing.conditionalpicks a target from a CELwhen:expression. Read the warning:routing.failoveron its own does not fail over - you must also configurehealth.eviction, and even then the request that trips the failure still fails unless retries are configured (Virtual models). - Load balancing In config
Weights supported:
routing.weightedwith per-target weights inside a virtual model, and ordinary backend load balancing from the general-purpose data plane underneath (Virtual models). - Upstream health tracking In config
Passive health checking on virtual models:
unhealthyExpressiondecides what counts as a failure, andevictiontakesduration,consecutiveFailures,healthThresholdandrestoreHealth. This is outlier ejection rather than active probing (Virtual models). - Cross-region failover Not documented
The vendor does not document this, so any behaviour you observe today is unversioned and may change.
n.a. as a feature: you can point virtual-model targets at regional provider endpoints and fail over between them, but no cross-region or geo-aware routing construct is documented (Virtual models, multiple LLMs).
Fallback chain: Ordered list — Try A, then B, then C. Simple and predictable, but every failover is all-or-nothing.
Defaults: No retries by default - retry is opt-in, and the codes shown (429, 500, 503) are an example rather than a documented default set (Retries).
The reliability model is unusually explicit about its own failure modes, which is a point in its favour: the docs state that failover without eviction does nothing, that the triggering request still fails without retries, and that retries stop working above the body-buffering threshold. Everything is config-file driven with no per-request escape hatch, so reliability behaviour is reviewable in Git but cannot be tuned by a caller (Retries, Virtual models).
How fast the hop is
Compiled binaryA single compiled Go or Rust binary. The lowest overhead floor of the self-hostable options, and the easiest to reason about under load.
compiled_binary. Rust data plane (7.35 MB of Rust in the language breakdown) with a Go Kubernetes controller (2.65 MB) and a TypeScript UI (1.07 MB); ships as a single static binary and a cr.agentgateway.dev/agentgateway container image (agentgateway/agentgateway, Docker).
curl -sL https://agentgateway.dev/install | bash for the binary; cr.agentgateway.dev/agentgateway:v1.5.0 for Docker and Compose; OCI Helm charts oci://cr.agentgateway.dev/charts/agentgateway-standalone for standalone and agentgateway-crds + agentgateway for Kubernetes mode; nightly builds via GitHub Actions artifacts or chart version 0.0.0-latest-dev (binary, Docker, Helm).
Streaming caveats: SSE streaming is supported on the LLM route types, but the guardrail interaction is the thing to know: response guards default to streaming: Disabled, meaning they do not run at all on a streamed response, and the mask action never applies to streams - content passes through with no error and no event. Use reject if you need enforcement on streaming traffic (Prompt guards, Completions).
Published figures, grouped by what each one measured. Figures in different groups are different quantities and cannot be compared with one another — nor, in most cases, with another vendor’s figure in the same group.
Overhead added by the gateway
The only figures that speak to whether this product slows your application down. Still not comparable between vendors: each was measured on different hardware, at a different load, with a different payload.
- 0.831 ms p50 Vendor-published
Fortio, 32 connections, 1 KB payload, 3 s max-throughput run against a mock LLM backend; p90 1.533 ms, p99 1.970 ms. Project-published, authored by Solo.io's Lin Sun.
Source - 0.227 ms p50 Vendor-published
Fortio held at a fixed 3,000 QPS for 30 s against a mock LLM backend; p99 0.436 ms, 13 MiB resident, 13.4% CPU. Project-published.
Source - 0.863 ms p50 Vendor-published
Maximum-throughput run against a mock Anthropic backend at 35,502 QPS; p99 1.972 ms. Same harness re-run against LiteLLM's Rust mode. Project-published.
Source - 36,933 QPS max throughput Vendor-published
Fortio, 32 connections, 1 KB payload, 3 s, mock LLM backend, 22 MB average memory. LiteLLM measured at 3,198 QPS / 11.8 GB on the same harness. Project-published.
Source - 35,502 QPS max throughput Vendor-published
Same harness re-run against LiteLLM's Rust mode, which reached 984 QPS at p99 71.451 ms. Mock backend. Project-published.
Source
Full round trip, including the model
Dominated by the upstream model, not the gateway. Useful as a sanity check, useless for comparing routing layers.
- 0.2 s TTFT p90 Vendor-published
Google Summer of Code study of inference-gateway EPP overhead at 60 QPS, comparing against a plain Kubernetes Service where TTFT p90 was 135.6 s; inter-token latency ~50 ms vs 30.3 ms direct. Real model backend, not a mock.
Source
Marketing claim, not a measurement
Fleet totals and unquantified claims. Recorded here because it is all the vendor published, not because it means anything operationally.
- 35x throughput, 300x memory, 122x latency multiple marketing summary Vendor-published
Commercial distribution landing page invites you to reproduce it with `agentgateway benchmark --compare`; no conditions stated on the page itself.
Source
All figures are project-published and comparative against a named competitor, which is the least independent shape a benchmark can take. Methodology is unusually transparent - tool, connection count, payload size, duration and open scripts are all stated - but the harness is the project's own. One outside evaluation examined the numbers without re-running them and flagged a specific discrepancy: the ~12 GB LiteLLM memory figure conflicts with LiteLLM's own reported 359 MB peak on a different harness, and it declined to treat the result as a settled speed winner (fmind.dev). No genuinely independent benchmark was found.
What it will stop
4 of 4 can blockTwo separate questions per control: can it stop a request at all, and what does it do before you change any settings? A control that inspects and forwards is a logging feature, however it is named.
- Personal data in prompts Can block the request
Out of the box: Blocks out of the box
Built-in regex guard with named PII patterns -
creditCard,ssn,email,phoneNumber- plus arbitrary custom patterns, running inline on request and response paths. Actions aremask,rejectoraudit, and the default ismask- so a freshly configured PII guard modifies traffic synchronously rather than merely logging it, the opposite default from gateways that ship guardrails asynchronous. Note the streaming exception: on streamed responsesmasksilently does nothing (Regex filters). - Prompt injection and jailbreaks Can block the request
Out of the box: Blocks out of the box
Prompt-injection and jailbreak detection comes from external services rather than a built-in classifier: Azure Content Safety exposes
detectJailbreak, and Google Model Armor and Bedrock Guardrails cover similar ground. Actions arerejectoraudit(Azure Content Safety, Google Model Armor). - Harmful content Can block the request
Out of the box: Blocks out of the box
openAIModerationcalls the OpenAI moderation endpoint inline; Bedrock Guardrails, Google Model Armor and Azure Content Safety provide the same function through their own services. All default toreject(OpenAI moderation, Prompt guards). - Your own policies Can block the request
Out of the box: Not documented
A
webhookguard posts to your own classifier and honoursrejectoraudit; arbitrary regex patterns cover deterministic policies. Guards run in sequence and can be scoped withscope: [systemPrompt, messages, toolInput, toolOutput], though settingscopereplaces the default (systemPrompt+messages) rather than adding to it, and only the regex and Bedrock Guardrails types accept a non-default scope (Webhooks, Prompt guards).
Almost no vendor in this catalogue documents what happens when the guardrail service itself times out. If the control matters to you, this is a question worth asking before you sign.
The docs describe verdict semantics and give latency guidance (regex under 1 ms, external moderation 50-200 ms) but never state what happens when the external guard service itself errors or times out - fail-open versus fail-closed is unspecified (Prompt guards, Webhooks, multi-layer).
Calls out to: OpenAI Moderation, AWS Bedrock Guardrails, Google Model Armor, Azure Content Safety. Each is a separate vendor relationship and a separate hop on the request path.
The defaults here enforce rather than observe, which is the opposite of several hosted gateways: regex guards default to mask and every external guard defaults to reject. Shared guards go in llm.policies.guardrails and merge with per-model llm.models[].guardrails, and MCP traffic has its own separate guardrail surface. Three real gaps to plan around: response guards do not run on streamed responses unless you enable streaming mode, passthrough routes with opaque bodies get no guardrails at all, and the WebSocket realtime path is exempt from prompt guards entirely (Prompt guards, passthrough, Realtime, MCP guardrails).
What you can see
Exports widelyYou decide whether bodies are captured, by setting or by header.
Body logging is off unless you turn it on, and a CEL filter plus remove can suppress entries or drop individual fields entirely (Access logs).
OpenTelemetry tracing with documented collector configs for Jaeger, Datadog, Honeycomb, Grafana Cloud and a plain OTel collector, plus a gen_ai.* attribute reference. Access logs can also be exported over OTLP independently of traces (Traces, attribute reference).
Metadata by default, full bodies on request. The UI has an explicit "Include prompts and completions in logs" toggle, and the database equivalent is frontendPolicies.accessLog.database.llm: full, which writes bodies into a request_log_payloads table (Access logs, database logs).
Where telemetry can go
- OpenTelemetry (OTLP)
- Prometheus
- Grafana
- Jaeger
- Datadog
- Honeycomb
- Grafana Cloud
- Langfuse
- LangSmith
- Arize Phoenix
Any OTLP-compatible backend, with per-destination guides for Jaeger, Datadog, Honeycomb and Grafana Cloud on the tracing side, Prometheus and Grafana for metrics, and dedicated LLM-observability pages for Langfuse, LangSmith and Arize Phoenix (Traces, Langfuse).
n.a. (no feedback, rating or annotation endpoint on any page fetched: LLM observability, Access logs, LLM overview). Downstream tools like Langfuse or Phoenix would own that.
n.a. as a gateway feature: no eval runner, dataset, scoring or experiment concept appears in the docs. The built-in LLM playground lets you fire ad-hoc requests, and the Langfuse/LangSmith/Phoenix integrations exist precisely because evaluation lives outside the gateway (playground, Langfuse).
Depends on the vendor’s SaaS: No: everything is local. Prometheus metrics on :15020, access logs to stdout or your own SQLite/PostgreSQL, OTLP export to whatever you run, and a built-in cost/analytics dashboard at /ui/llm/analytics that needs only config.database and a model catalog - no external Prometheus or Grafana required (metrics, cost dashboard).
Retention: Unlimited and entirely yours - and that is a burden, not a perk. Nothing prunes the request_logs or request_log_payloads tables and no retention setting is documented, so you own the lifecycle (database logs, Database).
Whether it fits how you work
How much work stands between you and a first call, how different that is from running it in production, and whether it slots into the stack you already have. Recorded as the shape of the work rather than a number of minutes — how long it takes you depends on which accounts and quota you already hold, which no comparison can know.
to try: cli or container to run: infrastructure rollout fits 5 of 10 common stacks
Getting to a first call
5 numbered stepsNothing works until you have a process running. Fine on a laptop, but it means there is no zero-install way to try it.
Read off: the vendor’s own quickstart — 5 numbered steps.
Counts the LLM quickstart only; separate MCP and non-agentic-HTTP quickstarts exist, and the five steps assume you already hold a provider key.
Before step one
- Your own provider key Required
You need an upstream provider account and key before anything works. That is a prerequisite, not a step.
Unavoidably: there is no credit pool or vendor account, so an upstream provider key (or a passthrough caller token, or a local Ollama/vLLM endpoint) is required before the gateway can serve a single request (API keys, quickstart).
- Payment method No card needed to start
No card, no account, no email.
curl -sL https://agentgateway.dev/install | bashand you are running (binary). - Gate before models answer No gate
Every catalogue model is callable as soon as you have a key.
No gate on the gateway's side: any model your upstream credential can reach is reachable through a route the moment you add it. Access restrictions are ones you impose yourself, through virtual-key
allowedModelsor CEL authorization (quickstart, virtual keys).
Everything you need first: A machine and a provider API key. No account, no card, no cluster, no database: the binary reads a YAML file and serves on port 4000, with the UI and playground on 15000 (quickstart). A database is needed only later, for USD budgets, the cost dashboard or database-backed logs.
The vendor’s own time claim: Project claims, verbatim: "Get started with agentgateway in less than 5 minutes - local-first" (homepage). The commercial distribution page claims "One binary, zero dependencies. Running in under 2 minutes." and "In 15 minutes, you'll see exactly why agentgateway exists" (Solo.io). Quoted, not verified. Marketing time claims assume every account and approval is already in place.
Running it in production
The same scale applied to the path the vendor recommends for production traffic. Kept separate from the quickstart because for several products here the two are barely related pieces of work.
This is an infrastructure project, not an integration. Expect charts or Terraform, networking, secrets and someone who owns the deployment.
This is an infrastructure project, not an integration. Expect charts or Terraform, networking, secrets and someone who owns the deployment.
Getting to production is a step up in kind from the quickstart, not just more of the same.
What production needs: For Kubernetes mode: a cluster, kubectl, helm, the CRDs chart and the control-plane chart, plus Gateway API custom resources. For anything with USD budgets, the cost dashboard or persisted request logs: config.database pointing at SQLite or PostgreSQL, and config.storage.mode: hybrid if you want the config resource API to list what it writes. One hard operational requirement: keep the unauthenticated admin address on localhost or otherwise network-isolated (Helm, Database, Storage modes, config resources).
Can you run it yourself
There is a command you can copy and run, so you can evaluate the self-hosted path yourself today.
curl -sL https://agentgateway.dev/install | bash; docker run -d --name agentgateway -p 4000:4000 -p 15000:15000 cr.agentgateway.dev/agentgateway:v1.5.0; helm upgrade -i agentgateway-standalone oci://cr.agentgateway.dev/charts/agentgateway-standalone --namespace agentgateway-system --create-namespace --version v1.5.0
How it fits your stack
5 of 10Each row is a thing you might already run. “With a caveat” means it works but not the way the vendor’s marketing implies — read the reason, because that is usually where the surprise lives.
- Fits The OpenAI SDK Drop-in once set up — but first-call work is cli or container.
- With a caveat The Vercel AI SDK Via the OpenAI provider
- No Cloudflare Workers No Workers guidance published.
- Fits Kubernetes agentgateway-standalone (standalone mode); agentgateway-crds + agentgateway (Kubernetes mode), all OCI charts under cr.agentgateway.dev/charts
- No Terraform or OpenTofu Nothing published for Terraform.
- Fits An existing API gateway This is that gateway — AI traffic becomes a plugin, not a new hop.
- Fits Cloud IAM I already run Reuses IAM roles, workload identity or managed identities.
- No LangChain or LlamaIndex No framework integration documented.
- Fits MCP servers to govern Acts as an MCP gateway or registry.
- With a caveat Nothing — plain Node or Python You have to run a process locally before any call works.
Reading this the other way round — pick what you already run and see every product scored against it.
The integration surfaces behind those answers
- Vercel AI SDK Via the OpenAI provider
Use the OpenAI provider pointed at a custom base URL. Works, but you lose provider-specific options.
No Vercel AI SDK provider package or integration page exists; the SDK's OpenAI-compatible provider would work against the gateway's
/v1base URL, but that path is not documented by the project (docs index, OpenAI SDK). - Cloudflare Workers Not documented
No Workers guidance either way. If you are edge-first, verify fetch-only compatibility yourself.
n.a. - it is a native Rust binary, not a Workers-compatible runtime, and no Cloudflare Workers deployment appears in the docs (binary, Docker, Helm).
- Kubernetes Official Helm chart
A named, published chart. You can read its values file before committing to anything.
Named:
agentgateway-standalone (standalone mode); agentgateway-crds + agentgateway (Kubernetes mode), all OCI charts under cr.agentgateway.dev/chartsOfficial OCI Helm charts:
helm upgrade -i agentgateway-standalone oci://cr.agentgateway.dev/charts/agentgateway-standalone --version v1.5.0for standalone, and a separate CRDs + control-plane pair for Kubernetes mode with Gateway API custom resources, xDS and a read-only GitOps UI (Helm, Kubernetes control plane). - Terraform Not documented
No Terraform surface published. Configuration is API or dashboard work.
n.a. (no Terraform provider, module or
terraformreference on any page fetched, including the docs index: docs index, Helm, Kubernetes control plane). - Existing API gateway It is the API gateway
This product is the gateway. If you already run it for your other APIs, AI traffic becomes a plugin rather than a new hop.
It is the gateway, and explicitly a general-purpose one: the FAQ positions it as an HTTP/gRPC data plane with load balancing, timeouts, retries, TLS, rate limits and authorization that can front ordinary APIs and microservices, so you do not run separate "regular" and "AI" gateways (FAQs).
- Cloud identity Reuses your cloud identity
Authenticate with the identity you already run — IAM roles, workload identity or managed identities. No long-lived key to rotate.
For the three big clouds: AWS backend auth supports
auth.aws.assumeRolewith CEL/JWT-derived session names and tags, and Azure and GCP have their own backend-auth provider pages, so upstream calls can use cloud identity rather than static keys (AWS integration, invoice-grade attribution). - MCP MCP gateway or registry
It sits in front of many MCP servers and governs access to them. This is the shape you want if MCP sprawl is the problem you are solving.
A full MCP gateway, arguably the project's centre of gravity: static and dynamic MCP servers, stdio/SSE/streamable-HTTP transports, MCP authentication and authorization, tool-level access control, MCP-specific guardrails, MCP rate limiting, MCP observability, and OpenAPI-to-MCP bridging. MCP spec 2026-07-28 support landed in v1.4 (MCP guardrails, MCP observability, v1.4 announcement).
n.a. - no LangChain, LlamaIndex or other Python-framework integration page exists in the documentation index. The FAQ asserts compatibility with "any agentic framework supporting MCP and A2A protocols, including LangGraph, AutoGen, kagent, Claude Desktop, and OpenAI SDK", but that is a protocol-level claim with no framework guide behind it (FAQs, docs index).
No first-party SDK exists. Documented clients are the standard OpenAI SDKs (openai for Python, openai for JavaScript) pointed at the gateway base URL, plus curl; the operator-side tooling is the agctl CLI (OpenAI SDK, curl).
Agent features: Function calling and tool use are documented for LLM routes, and the agent surface goes further than most gateways: A2A agent-to-agent traffic is a first-class protocol alongside MCP, with MCP guardrails, MCP authorization, tool-level access control and MCP observability as separate documented features. Context compression and semantic routing exist as project blog subjects rather than core docs pages (Agent connectivity, MCP guardrails, MCP authorization).
The five quickstart steps are: set provider credentials as environment variables, start the gateway, enable the LLM feature, add a model, send a request. The built-in UI adds a ClickOps path with an LLM playground and a CEL playground for testing expressions, and agctl covers config inspection, request tracing, log level changes and CPU/heap profiling. agentgateway import --from litellm --file litellm.yaml converts an existing LiteLLM proxy config and prints per-field compatibility findings, which is the cheapest possible migration test (quickstart, playground, import).
Deep infrastructure ecosystem, thin application ecosystem. On the infra side: Kubernetes Gateway API conformance, xDS dynamic config, cert-manager and external-dns integrations, eight IdP guides (Auth0, Authentik, Descope, Entra ID, Keycloak, oauth2-proxy, Okta, Tailscale) and cloud backend-auth for AWS, Azure and GCP. On the client side, documented consumers are coding agents and chat UIs - Claude Code, Claude Desktop, Codex, Continue, Cursor, Devin, GitHub Copilot, Antigravity, VS Code, opencode, LibreChat, Open WebUI, Chatbot UI, Goose. What is missing is the Python-framework layer: no LangChain or LlamaIndex integration page, no Vercel AI SDK provider, no Terraform provider and no SCIM. Migration is covered from one direction only, agentgateway import --from litellm (integrations index, import).
Silence in the docs: 9 of the integration questions on this card have no published answer either way. That is recorded as undocumented, not as a no — but it does mean you would be verifying it yourself.
What it does well
- Single Apache-2.0 Rust binary covering LLM, MCP, A2A, HTTP, gRPC and TCP in one data plane, with no hosted dependency
- Guardrails default to enforcing: regex guards default to masking and every external moderation guard defaults to reject, unlike gateways that ship log-only
- Per-virtual-key USD and token budgets with pre-request Block enforcement, plus a built-in cost dashboard that needs no Prometheus or Grafana
- Vendor-published Fortio benchmarks put it at ~35,000 QPS with sub-2 ms p99 and 13-22 MB resident memory against a mock backend
- Neutral governance: Linux Foundation donation in 2025, AAIF hosted project in 2026, 300+ contributors across 60+ organisations
- Documented one-command migration from LiteLLM proxy configs with per-field compatibility findings
Where it falls short
- The homepage markets "semantic caching" but no semantic or response cache appears anywhere in the documentation - only provider prompt-cache breakpoint control
- Response guardrails do not run on streamed responses by default (`streaming: Disabled`), and `mask` silently never applies to streams
- `routing.failover` alone does not fail over: the docs warn you must also configure `health.eviction`, and the triggering request still fails unless retries are set
- The admin address that serves the config resource API has no authentication at all - network isolation is the only control
- No published SLA, status page, SOC 2 / ISO / HIPAA posture, subprocessor list or log-retention setting; compliance is entirely the operator's job
- No documented image, audio, video or batch endpoint, no Terraform provider, no Vercel AI SDK package and no LangChain/LlamaIndex integration guide
Choose it when
Platform teams that already run Kubernetes or Envoy-style infrastructure and want one Rust data plane in front of LLM, MCP, A2A and ordinary HTTP traffic, with per-key USD budgets and guardrails defined in version-controlled YAML.
Look elsewhere when
You want a hosted control plane with dashboards, prompt management, evals and a compliance posture you can point an auditor at, or you have no appetite to run and upgrade a proxy yourself.
Open source: You run the gateway yourself. Full control over the data path and no per-request vendor fee, in exchange for operating it.
How hard is it to leave?
Derived from six published facts, not from an opinion. The weights are fixed and the same for every product — see the arithmetic.
| What helps you leave | Points | Source |
|---|---|---|
| Works with standard OpenAI code Switching away is a base-URL change rather than a rewrite of every call site. | 22 /22 | vendor page |
| No vendor-specific SDK required A proprietary client library spreads through your codebase and has to be torn out again. | 10 /10 | — |
| Can use your own provider accounts Your keys and billing relationship stay yours, so removing the gateway does not cut off model access. | 20 /20 | vendor page |
| Can be self-hosted You can run it yourself instead of accepting a pricing or policy change. | 20 /20 | vendor page |
| Configuration lives in version control Routing and budget rules are a file you keep, not dashboard state you would have to rebuild. | 16 /16 | vendor page |
| Your request history can be exported You leave with your own logs instead of abandoning them. | 12 /12 | — |
Read the fine print: Low lock-in by construction: Apache-2.0 code, a config file you own, your own database and your own upstream provider keys, so leaving means deleting a binary. The migration door swings inward too - `agentgateway import --from litellm` converts a LiteLLM proxy config with per-field compatibility findings - but there is no documented exporter back out to another gateway, and config is agentgateway-specific YAML or Gateway API CRs. Operational portability caveats: USD budgets, the cost dashboard and persisted logs depend on a database you must run, and the model cost catalog is populated by importing price files (`agctl catalog import --source models.dev`) rather than by a vendor-maintained feed ([import](https://agentgateway.dev/docs/standalone/latest/configuration/import/), [model costs](https://agentgateway.dev/docs/standalone/latest/llm/cost-controls/costs/)).
All six inputs are published, so this is scored against the full 100. This measures technical switching cost only. It does not price the engineering time to re-test prompts against a different routing stack.
agentgateway models & pricing
Browse every imported listing from this provider, with published token rates and a link to compare other providers for the same model. This is provider-reported coverage; an absent listing does not mean unsupported.
Loading model listings…
Official model coverage source ↗ · Model source coverage and limitations
Full specification
Every field we track. Blank fields say "Not published" rather than "No" — we do not infer an absence from silence. Switch to Technical in the header for the precise field names and the low-level details.
Overview
Interpret these fields: Self-hosted vs managed LLM gateways · Who still owns your LLM gateway?
- What kind of product Category
- Open source
- Marketplaces resell many providers behind one key. Gateways add governance on top. Open-source projects you run yourself. Cloud platforms are hyperscaler surfaces. Inference providers host models on their own hardware.
- Who runs it Deployment model
- Self-host only
- Managed means the vendor operates it. Self-host means you run it on your own infrastructure. Both means you can choose.
- Licence Licence
- Apache-2.0
- Proprietary products cannot be inspected or forked. Open licences such as MIT and Apache-2.0 let you audit, modify, and run the code without permission.
- Company Company
- agentgateway project, an Agentic AI Foundation initiative under the Linux Foundation (created by Solo.io)
- The organisation that maintains the product.
- Who you would be signing with Vendor status
- Run by a software foundation
- Whether the product is still an independent company, has been acquired, is a large cloud vendor’s product line, is run by a software foundation, or has been put into maintenance mode. Maintenance mode means bug fixes and security patches only — no new features.
- Last shipped an update Latest release
- 2026-08-27
v1.5.0, published 27 Aug 2026. Release notes live on GitHub Releases; the docs also carry a release-notes page. Nightly builds are published as GitHub Actions artifacts and as the `0.0.0-latest-dev` Helm chart version ([releases](https://github.com/agentgateway/agentgateway/releases), [Helm](https://agentgateway.dev/docs/standalone/latest/setup/install/helm/)).
- The date of the most recent release or version tag. A product that has not shipped in a year is a different risk from one that shipped last week, regardless of what its marketing site says.
- GitHub stars GitHub stars
- 4,985
- A rough proxy for community size on open-source projects. Not a quality measure.
Cost
Interpret these fields: How LLM gateway pricing works · LLM gateway spending limits: stop a runaway agent bill?
- Markup on model prices Token markup
- None
- How much the product adds on top of what the underlying model provider charges. Zero means you pay the same per-token price you would pay the model provider directly.
- Fee to add funds Credit purchase fee
- Not published
- A percentage charged when you top up your balance, separate from token prices. It is easy to miss because it does not appear on the per-token price list.
- Monthly cost per person Seat fee
- None
- A recurring per-user platform charge that applies regardless of how much you use the models.
- Can use your own provider accounts BYOK supported
- Yes
- Bring Your Own Key: you keep direct contracts with OpenAI, Anthropic and others, and the gateway only routes traffic. This preserves negotiated rates and committed-spend discounts.
- Cost of using your own accounts BYOK terms
- Every request runs on your own upstream credentials, supplied inline, from an `$ENV_VAR`, from a file-backed env var, from a Kubernetes secret, or passed straight through from the caller's token; the project bills nothing and issues no keys of its own ([API keys](https://agentgateway.dev/docs/standalone/latest/llm/api-keys/)).
- What the product charges to route traffic through your own provider keys.
- Free tier Free tier
- Everything is free: the whole gateway is Apache-2.0 and there is no hosted tier, no account, no seat and no request meter. You pay your infrastructure and your model providers ([agentgateway/agentgateway](https://api.github.com/repos/agentgateway/agentgateway)).
- What you can do without paying, useful for evaluation.
- Enterprise plan from Enterprise plan from
- Not published
- Annual entry price for the enterprise tier, where one is published or credibly reported.
- Cost to run it yourself Self-host cost
- Self-hosting is the only mode and it is free of vendor charges. The binary is a single Rust process (13-22 MB resident in the project's own benchmark runs), so the practical floor is one small container plus, optionally, SQLite or PostgreSQL if you want USD budgets, the cost dashboard or database-backed request logs ([Database](https://agentgateway.dev/docs/standalone/latest/setup/database/), [benchmark](https://agentgateway.dev/blog/2026-06-26-benchmarking-agentgateway-vs-litellm-part-2/)).
- What self-hosting actually costs once you account for infrastructure and any paid tier.
- How the vendor makes money Pricing model
- Open source, no paid tier
- The shape of the vendor’s bill: does the routing layer charge a percentage on top of tokens, a flat monthly fee, both, neither (because inference is the product), or nothing at all (open source with no paid tier).
- How pricing works, briefly Pricing model detail
- There is no price of any kind. No pricing page exists on the site (navigation is Docs, Standalone, Kubernetes, Models, Blog, Enterprise, Community), no seats, no request meter and no token markup, because nothing is vendor-operated. The only commercial path is a third-party enterprise distribution, Solo Enterprise for agentgateway, whose pricing is also not published ([Enterprise distributions](https://agentgateway.dev/enterprise), [FAQs](https://agentgateway.dev/docs/standalone/latest/faqs/)).
- A one-paragraph description that covers the caveats a pricing category cannot: introductory rates, per-feature meters, tier gating, and pricing that resets on a specific date.
- Minimum commitment Minimum commitment
- None. Apache-2.0 download, no account, no contract.
- Whether the vendor requires a minimum contract term, a minimum spend, or a provisioned-capacity purchase to get its published rate.
- Charges that fire after you go over an allowance Overage terms
- n.a. - no metered plan exists, so there is nothing to overrun. Your own `remoteRateLimit` token buckets and per-key budgets are the only limits, and you set them ([per-key budgets](https://agentgateway.dev/docs/standalone/latest/llm/cost-controls/budget-limits/per-key/)).
- The line items that scale with usage after an included allowance is exhausted — log storage, extra requests, per-feature meters, data export — which is where cost estimates usually go wrong.
- Prompt cache offered Cache mechanism
- Passes provider caching through
- Whether the gateway offers its own response cache, what kind of match it does (exact request, prefix, semantic), or simply passes provider caching through unchanged.
- Discount on cached input Cache-read discount
- Not published
- How much cheaper cached tokens are than fresh input, when the vendor publishes a single figure. Bundled-inference clouds usually price this per model instead of as one number.
- Premium on cache writes Cache-write premium
- Not published
- How much more the first write of a cached prefix costs versus a plain input token. A high write premium and a low hit rate can leave you paying more than you save, so this matters as much as the read discount.
- Who captures the cache saving Cache economics
- The only caching feature is `promptCaching`, which controls provider-side prompt-cache breakpoints (`cacheSystem`, `cacheMessages`, `cacheTools`, `minTokens`, `cacheMessageOffset`) so the provider's own cache discounts apply; the gateway stores no responses itself and publishes no cached-token pricing. Worth flagging: the homepage markets "token budgets, semantic caching, and prompt redaction", but no semantic or response cache appears on any documentation page fetched - the semantic work in the project is semantic *routing*, not caching ([LLM policies](https://agentgateway.dev/docs/standalone/latest/configuration/traffic-management/llm/), [homepage](https://agentgateway.dev)).
- Whether the customer keeps the full saving from caching or the vendor captures part of it — and any conditions attached (write premium, storage fees, best-effort hits).
- What you can split spend by Cost attribution
- Per-request cost lands in the access log as `agw.ai.usage.cost.total` and in CEL as `llm.cost` / `llm.costRates`, so you can split by virtual key, user, model or any custom CEL label. The built-in LLM > Analytics dashboard aggregates it. For AWS, `auth.aws.assumeRole` session names and tags plus Bedrock `requestMetadata` push attribution into CloudTrail, the Cost & Usage Report and Cost Explorer - invoice-grade rather than gateway-estimated ([costs](https://agentgateway.dev/docs/standalone/latest/llm/cost-controls/costs/), [attribution](https://agentgateway.dev/docs/standalone/latest/llm/cost-controls/attribution/)).
- The dimensions the vendor documents for splitting spend — per key, per user, per team, per tag, per customer. Matters if you need to chargeback internally or bill an end customer.
- How you get cost data out Cost export
- No CSV or reporting export endpoint is documented. You get the data by querying your own `request_logs` / `request_log_payloads` tables, scraping Prometheus (`agentgateway_gen_ai_client_token_usage`, `agentgateway_cost_catalog_lookups_total`), or shipping access logs over OTLP ([database logs](https://agentgateway.dev/docs/standalone/latest/observability/access-logs/database/), [metrics](https://agentgateway.dev/docs/standalone/latest/observability/metrics/overview/)).
- The mechanisms the vendor publishes for exporting cost and usage data: CSV, an API, webhooks, S3, a data warehouse, or nothing at all. Any per-unit price is included.
- Who pays the model bill BYOK mode
- Your keys only
- Whether you bring your own provider accounts (BYOK), buy inference from this vendor, or can do either. This is the single biggest commercial difference between these products: it decides who holds the contract with the model provider and who carries the spend.
Spend governance in detail
Seven signals matter when a bill starts to hurt: who can spend, how much, on what, and who gets paged when it goes wrong. Everything below is drawn from the vendor’s own pricing and docs pages — how we read these.
- Virtual or scoped keys Yes
Keys that carry their own budget and rate-limit policy, so an intern experiment cannot spend against a production budget.
`llm.policies.apiKey.keys[]`, created in the UI at `/ui/llm/keys`, in the config file, or through the config resource API. Callers present `Authorization: Bearer $VIRTUAL_KEY`.
- Budget caps per key Yes
A dollar or token ceiling attached to an individual key. Where enforcement is soft, one over-limit request still completes before the block kicks in.
`budgets[{name, limit:{unit: USD|Tokens, amount}, window:{rolling: 24h}, onBudgetExceeded: Block|Audit}]`. USD budgets require `config.database` (SQLite or PostgreSQL) plus a model cost catalog; token budgets do not. Checked before the request is forwarded.
- Budget caps per team or workspace Not published
A ceiling applied at a higher scope than one key — a team, a workspace, a customer, or an entire environment.
Not stated. Budgets attach to virtual keys, and rate-limit descriptors can key on a user or JWT claim, but no team or workspace object appears in the docs fetched.
- Rate limiting as a cost control Yes
Configurable request-per-time-window caps. Platform-set rate limits do not count as spend controls; user-configurable ones do.
`remoteRateLimit` token buckets with `type: tokens` and per-user descriptors, evaluated in two phases (request estimate, then response actuals), returning 429 when empty. No database needed and it works in both standalone and Kubernetes modes.
- Model allowlists Yes
A policy that constrains which models a key or team can call, keeping expensive frontier models out of the wrong hands.
Per virtual key via `allowedModels`, which accepts wildcards such as `["gpt-5*"]`.
- Spend alerts Not published
Alerts fired as spend approaches a threshold. Alerts that only fire after the meter has rolled over are marked as such.
No alerting feature. You would alert off your own Prometheus metrics or the access log; `onBudgetExceeded: Audit` records a violation without blocking.
- Webhook notifications Not published
Programmatic notifications on spend events, so budget breaches can page an on-call or open a ticket.
Not stated on any page fetched. Guardrail webhooks exist, but no budget or spend webhook.
Enforcement: Enforced before each request
Catalog
Interpret these fields: LLM gateway model counts: what “500+” means
- Models available Models available
- ~1,002
- How many models you can call, shown as a range because several vendors publish different totals on different pages. Vendor-reported either way, so counts are not directly comparable — some count every provider variant of the same model separately.
- Model providers reachable Upstream providers
- 20–44
- How many distinct model providers or labs you can reach, shown as a range where the vendor’s own pages disagree. More providers usually means better redundancy when one has an outage.
- Works with standard OpenAI code OpenAI-compatible API
- Yes
- If yes, you can usually switch to it by changing one base URL, and switch away just as easily. This is the main defence against lock-in.
- OpenAI chat endpoint POST /v1/chat/completions
- Yes
- The endpoint almost every application ports first. "Not documented" means the vendor never states it, which is different from a documented no.
- Anthropic messages endpoint POST /v1/messages
- Yes
- Whether Anthropic-shaped calls work without rewriting them. Several products support this only as SDK compatibility or provider passthrough rather than a native endpoint — the detail page says which.
- OpenAI Responses endpoint POST /v1/responses
- Yes
- The newer stateful OpenAI surface. Support is much thinner across this market than chat completions.
- Embeddings endpoint POST /v1/embeddings
- Yes
- Whether you can generate vectors through the same gateway, or need a second integration for retrieval workloads.
- Image generation endpoint POST /v1/images/generations
- Not documented
- Whether image models are reachable through the same surface as text.
- Audio endpoints POST /v1/audio/*
- Partly
- Speech-to-text and text-to-speech. Frequently the first gap in an otherwise complete gateway.
- Batch jobs endpoint POST /v1/batches
- Not documented
- Asynchronous bulk processing, usually at a discount. Commonly undocumented, and commonly the reason a migration stalls late.
- Needs the vendor’s own code library Requires a vendor-specific SDK
- No
- A proprietary client library spreads through your codebase and has to be torn out again if you leave. “No” is the better answer here, and it means the standard OpenAI client works.
- You can export your request history Logs / usage data export
- Yes
- Whether you can get your own request logs, traces, or usage records back out — through an API, a bulk export, or a download. Decides whether you leave with your history or abandon it.
- Settings can live in version control Declarative config-as-code
- Yes
- Whether routing, fallback, and budget rules can be declared in a file you keep in Git, rather than existing only as settings clicked into a hosted dashboard.
- Embeddings Embeddings
- Yes
- Text-to-vector models, needed for search and retrieval features.
- Image generation Image generation
- Not published
- Whether image models are reachable through the same interface.
- Speech and audio Speech and audio
- Not published
- Text-to-speech or transcription models through the same interface.
- Video generation Video generation
- Not published
- Whether video models are reachable through the same interface.
- Batch processing Batch processing
- Not published
- Submitting large jobs for cheaper, slower processing. Often 50% off for work that is not time-sensitive.
Routing & reliability
Interpret these fields: How LLM gateway failover actually works · Does your LLM gateway promise any uptime? · Is routing destroying your prompt cache? · Changing models without breaking production
- Uptime it promises in writing Contractual SLA uptime
- Not published
- The uptime percentage in a published, contractual service level agreement. A public status page is not an SLA — it reports what happened, it does not promise anything or pay you back when it breaks.
- Automatic failover Automatic failover
- Yes
- When a model provider goes down or rate-limits you, traffic moves to a backup automatically instead of returning errors to your users.
- Load balancing Load balancing
- Yes
- Spreads requests across several providers or keys to raise your effective rate limit.
- Rule-based routing Conditional routing
- Yes
- Send different requests to different models based on rules — for example a cheap model for free users and a strong model for paying ones.
- Response caching Response caching
- Not published
- Reuses the answer when the exact same request comes in again, which cuts both cost and latency.
- Similar-question caching Semantic cache
- No
- Reuses an answer when a new question means roughly the same thing as an earlier one. Saves far more than exact-match caching but can return subtly wrong answers if tuned loosely.
- Where you set the timeout Request timeout surface
- In config
`config_file`, at four levels: route `requestTimeout` (whole client request including retries) and `backendRequestTimeout` (each individual try), backend `http.requestTimeout`, and `tcp.connectTimeout` for connection establishment. No per-request header or query override is documented ([Timeouts](https://agentgateway.dev/docs/standalone/latest/configuration/resiliency/timeouts/)).
- Where a request timeout can be set: per request, in a config file, in the vendor dashboard, or nowhere. Nine of the twenty products documented here do not describe a request timeout at all, so the worst case of a hung upstream call is unknowable from the docs.
- Where you set retries Retry policy surface
- In config
`retry.attempts`, `retry.backoff` (example `500ms`) and `retry.codes` (example `[429, 500, 503]`). A retry prefers a different backend from the one that just failed. Caveat: retries are disabled once a request body exceeds the buffering threshold, because the body can no longer be replayed ([Retries](https://agentgateway.dev/docs/standalone/latest/configuration/resiliency/retries/)).
- Where retry count and backoff are configured. Worth knowing alongside billing: a retried streaming call can be charged more than once.
- Where you set fallbacks Fallback surface
- In config
ORDERED and WEIGHTED both available, through `llm.virtualModels[]` (v1.3+): `routing.weighted` distributes by weight, `routing.failover` uses priority groups, `routing.conditional` picks a target from a CEL `when:` expression. **Read the warning**: `routing.failover` on its own does not fail over - you must also configure `health.eviction`, and even then the request that trips the failure still fails unless retries are configured ([Virtual models](https://agentgateway.dev/docs/standalone/latest/llm/virtual-models/)).
- Where the fallback chain is defined. Almost every product claims fallback; the useful question is whether you can change it from code or only by hand in a dashboard.
- Shape of the fallback chain Fallback shape
- Ordered list
- An ordered list tries targets in sequence; a weighted split sends a percentage of traffic to each, which is what you need to trial a new model on 5% of requests. Weighted splits are much rarer than the marketing implies.
- Upstream health tracking Health checks / circuit breaking
- In config
Passive health checking on virtual models: `unhealthyExpression` decides what counts as a failure, and `eviction` takes `duration`, `consecutiveFailures`, `healthThreshold` and `restoreHealth`. This is outlier ejection rather than active probing ([Virtual models](https://agentgateway.dev/docs/standalone/latest/llm/virtual-models/)).
- Whether the product notices a failing upstream and stops sending traffic to it, and whether you can tune the thresholds. This is what turns a provider outage into a blip rather than a sustained error rate, and only four of the twenty expose it.
- Cross-region failover you control Multi-region failover surface
- Not documented
n.a. as a feature: you can point virtual-model targets at regional provider endpoints and fail over between them, but no cross-region or geo-aware routing construct is documented ([Virtual models](https://agentgateway.dev/docs/standalone/latest/llm/virtual-models/), [multiple LLMs](https://agentgateway.dev/docs/standalone/latest/llm/providers/multiple-llms/)).
- Whether you can define what happens when a region degrades. A vendor running many regions is not the same as a vendor letting you configure failover between them; only three document a user-controlled mechanism.
- Where you set load balancing Load balancing surface
- In config
Weights supported: `routing.weighted` with per-target weights inside a virtual model, and ordinary backend load balancing from the general-purpose data plane underneath ([Virtual models](https://agentgateway.dev/docs/standalone/latest/llm/virtual-models/)).
- Where traffic distribution across upstreams or keys is configured.
Operations
Interpret these fields: LLM gateway observability: traces, logs and export · Running coding agents through an LLM gateway · Changing models without breaking production
- Usage dashboards and logs Observability
- Yes
- Built-in visibility into what was sent, what came back, what it cost, and how long it took.
- Spending limits Budget controls
- Yes
- Hard caps that stop spend before it becomes a surprise invoice. The single most valuable control for a small team.
- Rate limits Rate limits
- Yes
- Caps on request volume per key or per user, useful for protecting against abuse and runaway loops.
- Separate keys per team or app Virtual keys
- Yes
- Issue scoped keys with their own budgets and permissions so you can attribute cost and revoke access without rotating everything.
- Prompt versioning Prompt management
- Not published
- Store and version prompts outside your code so they can be changed without a deploy.
- Quality testing Evals
- Not published
- Built-in tooling to score model output against test cases, so you can tell whether a model swap made things better or worse.
- MCP support MCP support
- Yes
- Native support for the Model Context Protocol, the emerging standard for connecting models to external tools.
- What gets logged Logged content
- Your choice
Metadata by default, full bodies on request. The UI has an explicit "Include prompts and completions in logs" toggle, and the database equivalent is `frontendPolicies.accessLog.database.llm: full`, which writes bodies into a `request_log_payloads` table ([Access logs](https://agentgateway.dev/docs/standalone/latest/observability/access-logs/view/), [database logs](https://agentgateway.dev/docs/standalone/latest/observability/access-logs/database/)).
- Whether full prompts and responses are stored, only metadata, or nothing. Full-body logging is the most useful debugging feature here and the one most likely to need a conversation with your compliance team.
- You can turn logging off Body-logging opt-out
- Yes
Body logging is off unless you turn it on, and a CEL `filter` plus `remove` can suppress entries or drop individual fields entirely ([Access logs](https://agentgateway.dev/docs/standalone/latest/observability/access-logs/view/)).
- Whether prompt and response bodies can be suppressed while still keeping usage metrics. Nineteen of the twenty document a way to do this; the mechanisms range from a per-request header to an organisation-wide setting.
- Traces you can take elsewhere Distributed tracing
- OpenTelemetry
OpenTelemetry tracing with documented collector configs for Jaeger, Datadog, Honeycomb, Grafana Cloud and a plain OTel collector, plus a `gen_ai.*` attribute reference. Access logs can also be exported over OTLP independently of traces ([Traces](https://agentgateway.dev/docs/standalone/latest/observability/traces/setup/), [attribute reference](https://agentgateway.dev/docs/standalone/latest/observability/traces/attribute-reference/)).
- Whether the product emits OpenTelemetry, a proprietary format, or nothing. OpenTelemetry means the traces land in the tooling you already run instead of only in the vendor’s dashboard.
- Where telemetry can go Export destinations
- OpenTelemetry (OTLP), Prometheus, Grafana, Jaeger, Datadog, Honeycomb, Grafana Cloud, Langfuse, LangSmith, Arize Phoenix
Any OTLP-compatible backend, with per-destination guides for Jaeger, Datadog, Honeycomb and Grafana Cloud on the tracing side, Prometheus and Grafana for metrics, and dedicated LLM-observability pages for Langfuse, LangSmith and Arize Phoenix ([Traces](https://agentgateway.dev/docs/standalone/latest/observability/traces/setup/), [Langfuse](https://agentgateway.dev/docs/standalone/latest/integrations/llm-observability/langfuse/)).
- Documented sinks for logs and metrics. This is a good proxy for how replaceable the vendor’s own dashboard is: around twenty destinations means you never have to depend on it, while a CSV download means you do.
- Can record user feedback Feedback capture API
- No
n.a. (no feedback, rating or annotation endpoint on any page fetched: [LLM observability](https://agentgateway.dev/docs/standalone/latest/llm/observability/), [Access logs](https://agentgateway.dev/docs/standalone/latest/observability/access-logs/view/), [LLM overview](https://agentgateway.dev/docs/standalone/latest/llm/about/)). Downstream tools like Langfuse or Phoenix would own that.
- Whether there is an API to attach a rating or score to a logged request, which is what lets production traffic feed quality work later.
- Scores live traffic Online eval hooks
- No
n.a. as a gateway feature: no eval runner, dataset, scoring or experiment concept appears in the docs. The built-in LLM playground lets you fire ad-hoc requests, and the Langfuse/LangSmith/Phoenix integrations exist precisely because evaluation lives outside the gateway ([playground](https://agentgateway.dev/docs/standalone/latest/llm/playground/), [Langfuse](https://agentgateway.dev/docs/standalone/latest/integrations/llm-observability/langfuse/)).
- Whether automated scorers can run against real production requests, rather than only against a test set you assemble yourself.
Performance
- Delay it adds Proxy overhead
- 0.863 ms
- Extra time the product itself adds to each request, on top of however long the model takes. Usually irrelevant next to multi-second model latency, but it matters for high-volume or streaming-sensitive workloads.
- Requests per second ceiling Throughput
- 35,502 rps
- Published sustained request rate before the product becomes the bottleneck. Only relevant at genuinely high volume.
- What the request path runs on Architecture class
- Compiled binary
`compiled_binary`. Rust data plane (7.35 MB of Rust in the language breakdown) with a Go Kubernetes controller (2.65 MB) and a TypeScript UI (1.07 MB); ships as a single static binary and a `cr.agentgateway.dev/agentgateway` container image ([agentgateway/agentgateway](https://api.github.com/repos/agentgateway/agentgateway), [Docker](https://agentgateway.dev/docs/standalone/latest/setup/install/docker/)).
- The comparable way to talk about latency here. An edge worker, a compiled Go or Rust binary, and a Python proxy have different overhead floors no matter which figures each vendor publishes. Products that never disclose their runtime are recorded as undisclosed rather than assumed.
- You can run the request path yourself Self-hostable data plane
- Yes
`curl -sL https://agentgateway.dev/install | bash` for the binary; `cr.agentgateway.dev/agentgateway:v1.5.0` for Docker and Compose; OCI Helm charts `oci://cr.agentgateway.dev/charts/agentgateway-standalone` for standalone and `agentgateway-crds` + `agentgateway` for Kubernetes mode; nightly builds via GitHub Actions artifacts or chart version `0.0.0-latest-dev` ([binary](https://agentgateway.dev/docs/standalone/latest/setup/install/binary/), [Docker](https://agentgateway.dev/docs/standalone/latest/setup/install/docker/), [Helm](https://agentgateway.dev/docs/standalone/latest/setup/install/helm/)).
- Whether the component that actually carries your prompts can run on your own infrastructure. Distinct from a vendor offering a self-hosted control plane while still proxying traffic through their network.
- Streaming responses Streaming support
- Yes
SSE streaming is supported on the LLM route types, but the guardrail interaction is the thing to know: response guards default to `streaming: Disabled`, meaning they do not run at all on a streamed response, and the `mask` action never applies to streams - content passes through with no error and no event. Use `reject` if you need enforcement on streaming traffic ([Prompt guards](https://agentgateway.dev/docs/standalone/latest/llm/prompt-guards/overview/), [Completions](https://agentgateway.dev/docs/standalone/latest/llm/api-types/completions/)).
- Whether token-by-token streaming is documented. The caveats matter more than the yes: some products cannot cancel a stream without still being billed, and several timeout and fallback mechanisms stop applying once the first token has been sent.
Security & compliance
Interpret these fields: LLM gateway compliance: SOC 2, HIPAA and evidence · How LLM gateway guardrails fail · Which LLM gateways store your prompts? · Do you need an MCP gateway as well?
- Does your prompt reach their servers Prompt transits vendor
- No
Prompts go from your client to your own agentgateway process and straight on to the provider you configured; there is no project-operated plane in the path and no telemetry call-home documented. In `passthrough` mode with `opaque` bodies the gateway does not even interpret the payload ([passthrough](https://agentgateway.dev/docs/standalone/latest/llm/api-types/passthrough/)).
- Whether the text you send passes through this company’s own infrastructure. If it does, every other promise on this page is a policy commitment rather than a physical impossibility. Self-hosted products can answer no outright.
- What they keep if you change nothing Logging default
- Metadata only, not content
A structured access log is written to stdout for every request by default (key=value, switchable to JSON) with `gen_ai.*` attributes and duration, but prompt and completion bodies are excluded until you opt in. CEL `filter`, `add` and `remove` let you shape or suppress entries ([Access logs](https://agentgateway.dev/docs/standalone/latest/observability/access-logs/view/)).
- Defaults matter more than options. A product that stores full prompts and replies unless you find the right header will have stored them by the time you read the docs.
- How long they keep it Default content retention (days)
- Not published
No retention or pruning setting is documented. Logs go to stdout, to an OTLP endpoint, or into a `request_logs` table in your own SQLite or PostgreSQL database, and lifecycle is entirely yours ([database logs](https://agentgateway.dev/docs/standalone/latest/observability/access-logs/database/), [Database](https://agentgateway.dev/docs/standalone/latest/setup/database/), [Storage modes](https://agentgateway.dev/docs/standalone/latest/setup/storage/)).
- Default retention for request content, in days. Zero means nothing is kept. Read the note: several products keep nothing as a rule but make timed exceptions for abuse review or specific models.
- Could they train on your prompts Training on customer data
- Not applicable
not_applicable by construction, not by policy: no prompt ever reaches a project-operated service, and no training, telemetry or data-use statement appears on the pages checked ([introduction](https://agentgateway.dev/docs/standalone/latest/about/introduction/), [FAQs](https://agentgateway.dev/docs/standalone/latest/faqs/), [Enterprise distributions](https://agentgateway.dev/enterprise)).
- Whether the vendor may use your prompts and outputs to train models. “Not published” means we could not find any position, which is not the same as a no — ask for it in writing.
- Where it runs, and what you can pin Region and residency control
- n.a. - no vendor regions exist. You run the process wherever you choose; multi-region behaviour is whatever your own deployment and virtual-model routing do ([virtual models](https://agentgateway.dev/docs/standalone/latest/llm/virtual-models/)).
- Which regions are offered and whether you can force processing to stay in one. A global endpoint that silently picks a region is a different compliance story from an endpoint you pin yourself.
- Where safety filters run Guardrail execution location
- In your own infrastructure
Guards execute inside your own gateway process; the external ones (OpenAI Moderation, Bedrock Guardrails, Google Model Armor, Azure Content Safety, your own webhook) are outbound calls your gateway makes to services you configure, not to a project-run service ([Prompt guards](https://agentgateway.dev/docs/standalone/latest/llm/prompt-guards/overview/)).
- A filter that strips personal data only helps if it runs before the data leaves your boundary. If guardrails execute in the vendor’s cloud, the vendor has already received whatever you wanted redacted.
- Who else touches the data Subprocessor list
- Not published
- The published list of third parties the vendor passes your data to. No list means you cannot know the full chain, which most data-protection agreements require you to.
- SOC 2 audited SOC 2 audited
- Not published
- An independent audit of security controls. Enterprise buyers and their procurement teams routinely require it.
- Will sign a HIPAA agreement HIPAA BAA
- Not published
- Required before you may send protected health information through the service. Without a signed BAA, healthcare data is off limits.
- GDPR commitments GDPR commitments
- Not published
- Published data processing terms for handling personal data of people in the EU and UK.
- Can keep data in the EU EU data residency
- Not published
- Requests can be processed inside the EU rather than routed to US infrastructure. Often the deciding constraint for European customers.
- Does not retain your data Zero data retention
- Not applicable
- Prompts and responses are not stored after the request completes. Sometimes a paid add-on rather than the default.
- Strips personal data PII redaction
- Yes
- Detects and removes identifiers such as names, emails, and card numbers before the request reaches the model provider.
- Content guardrails Content guardrails
- Yes
- Policy checks on inputs and outputs — blocking unsafe content, enforcing formats, or catching prompt-injection attempts.
- Runs fully disconnected Air-gapped deployment
- Not published
- Can be deployed in a network with no internet access, which some regulated and defence environments require.
- Blocks personal data in prompts PII / DLP enforcement
- Can block the request
Built-in regex guard with named PII patterns - `creditCard`, `ssn`, `email`, `phoneNumber` - plus arbitrary custom patterns, running inline on request and response paths. Actions are `mask`, `reject` or `audit`, and the default is `mask` - so a freshly configured PII guard modifies traffic synchronously rather than merely logging it, the opposite default from gateways that ship guardrails asynchronous. Note the streaming exception: on streamed responses `mask` silently does nothing ([Regex filters](https://agentgateway.dev/docs/standalone/latest/llm/prompt-guards/regex/)).
- Whether personal data detection sits on the request path and can stop the call, merely inspects and forwards it, or is not documented. A control that only reports is a logging feature, not a policy control.
- Blocks prompt injection Injection / jailbreak enforcement
- Can block the request
Prompt-injection and jailbreak detection comes from external services rather than a built-in classifier: Azure Content Safety exposes `detectJailbreak`, and Google Model Armor and Bedrock Guardrails cover similar ground. Actions are `reject` or `audit` ([Azure Content Safety](https://agentgateway.dev/docs/standalone/latest/llm/prompt-guards/azure-content-safety/), [Google Model Armor](https://agentgateway.dev/docs/standalone/latest/llm/prompt-guards/google-model-armor/)).
- Whether injection and jailbreak detection can stop a request. Most products offering this call a partner classifier rather than shipping their own.
- Blocks harmful content Toxicity / moderation enforcement
- Can block the request
`openAIModeration` calls the OpenAI moderation endpoint inline; Bedrock Guardrails, Google Model Armor and Azure Content Safety provide the same function through their own services. All default to `reject` ([OpenAI moderation](https://agentgateway.dev/docs/standalone/latest/llm/prompt-guards/moderation/), [Prompt guards](https://agentgateway.dev/docs/standalone/latest/llm/prompt-guards/overview/)).
- Whether hate, violence, sexual and self-harm categories are checked inline and can stop a request, in either direction.
- Your own policy rules Custom policy hooks
- Can block the request
A `webhook` guard posts to your own classifier and honours `reject` or `audit`; arbitrary regex patterns cover deterministic policies. Guards run in sequence and can be scoped with `scope: [systemPrompt, messages, toolInput, toolOutput]`, though setting `scope` **replaces** the default (`systemPrompt` + `messages`) rather than adding to it, and only the regex and Bedrock Guardrails types accept a non-default scope ([Webhooks](https://agentgateway.dev/docs/standalone/latest/llm/prompt-guards/webhooks/), [Prompt guards](https://agentgateway.dev/docs/standalone/latest/llm/prompt-guards/overview/)).
- Whether you can add your own rule — a regex, a webhook, or your own classifier — rather than choosing from the vendor library.
- Where guardrails run Guardrail execution location
- Either, your choice
- Whether guardrail evaluation happens inside your infrastructure or on the vendor’s servers. This decides whether the prompt you are trying to protect leaves your network in order to be checked.
- If the guardrail itself fails Guardrail failure mode
- Not documented
The docs describe verdict semantics and give latency guidance (regex under 1 ms, external moderation 50-200 ms) but never state what happens when the external guard service itself errors or times out - fail-open versus fail-closed is unspecified ([Prompt guards](https://agentgateway.dev/docs/standalone/latest/llm/prompt-guards/overview/), [Webhooks](https://agentgateway.dev/docs/standalone/latest/llm/prompt-guards/webhooks/), [multi-layer](https://agentgateway.dev/docs/standalone/latest/llm/prompt-guards/multi-layer/)).
- What happens when the guardrail service times out or errors: does the request proceed unchecked, or is it blocked? This is the worst-documented field in the entire catalogue — only two vendors state it plainly, which means most teams are running a control whose failure behaviour they cannot know.
- Third-party guardrail vendors Guardrail integrations
- OpenAI Moderation, AWS Bedrock Guardrails, Google Model Armor, Azure Content Safety
- Named external guardrail services the product can call. A long list means the product is a router for policy engines rather than a policy engine itself — which also means another vendor bill and another hop.
Compliance evidence
Graded by how strong the evidence is, not whether the word appears on the vendor’s website. An audited report and a marketing claim are different things, and only one of them will satisfy your own auditor.
- SOC 2 Not published No SOC 2 statement, trust portal or report reference on the project pages checked
- ISO 27001 Not published
- GDPR DPA Not published No DPA offer; there is no vendor to sign one for the upstream project
- HIPAA BAA Not published No BAA offer found
- FedRAMP Not published
- ITAR Not published
No compliance certifications were found published for this product. That is not the same as failing an audit — it means there is nothing public to check, so ask for evidence directly.
Security incidents
Publicly documented incidents affecting this product. An incident here is not by itself a reason to rule a product out — what matters is what failed, whether it could recur, and what it means for the way you would deploy it.
Fit & integration
Interpret these fields: How much does an LLM gateway lock you in? · Do you need an MCP gateway as well? · Running coding agents through an LLM gateway · Changing models without breaking production
- Work to try it Evaluation work shape
- Run something locally first
- The shape of the work on the vendor’s own quickstart, from swapping one base URL through to deploying infrastructure. An ordinal class rather than a duration, because elapsed time depends on accounts and quota we cannot see.
- Work to run it Production work shape
- Deploy it on your infrastructure
- The same scale applied to the vendor’s recommended production path. For several products this is much heavier than the quickstart, which is exactly why both are recorded.
- Steps on the quickstart Numbered quickstart steps
- 5
Counts the LLM quickstart only; separate MCP and non-agentic-HTTP quickstarts exist, and the five steps assume you already hold a provider key.
- A literal count of numbered steps on the vendor’s quickstart, recorded as evidence beside the work shape. Zero means the page publishes no numbered procedure at all. Large counts usually mean interleaved language tracks rather than more work.
- Can you self-host it today Self-host install documentation
- Install command published
`curl -sL https://agentgateway.dev/install | bash`; `docker run -d --name agentgateway -p 4000:4000 -p 15000:15000 cr.agentgateway.dev/agentgateway:v1.5.0`; `helm upgrade -i agentgateway-standalone oci://cr.agentgateway.dev/charts/agentgateway-standalone --namespace agentgateway-system --create-namespace --version v1.5.0`
- Whether an install command is actually published. Several products advertise self-hosting while publishing no command to start from, which a plain yes/no would hide.
- Works with the OpenAI SDK OpenAI SDK drop-in
- Yes
Yes, a literal base-URL swap with a throwaway key: `openai.OpenAI(api_key="anything", base_url="http://localhost:4000/v1")`. The gateway authenticates callers with its own virtual keys if you configure them, otherwise the upstream key never leaves the config ([OpenAI SDK](https://agentgateway.dev/docs/standalone/latest/integrations/llm-clients/openai-sdk/), [Completions](https://agentgateway.dev/docs/standalone/latest/llm/api-types/completions/)).
- Whether an existing OpenAI-compatible client can be pointed at it by changing the base URL and key.
- Vercel AI SDK support AI SDK provider package
- Via the OpenAI provider
No Vercel AI SDK provider package or integration page exists; the SDK's OpenAI-compatible provider would work against the gateway's `/v1` base URL, but that path is not documented by the project ([docs index](https://agentgateway.dev/llms.txt), [OpenAI SDK](https://agentgateway.dev/docs/standalone/latest/integrations/llm-clients/openai-sdk/)).
- Whether a first-party AI SDK provider package exists, or only a community package, a documented workaround, or the generic OpenAI provider pointed at a custom base URL.
- Python framework integrations Documented Python frameworks
- Not published
n.a. - no LangChain, LlamaIndex or other Python-framework integration page exists in the documentation index. The FAQ asserts compatibility with "any agentic framework supporting MCP and A2A protocols, including LangGraph, AutoGen, kagent, Claude Desktop, and OpenAI SDK", but that is a protocol-level claim with no framework guide behind it ([FAQs](https://agentgateway.dev/docs/standalone/latest/faqs/), [docs index](https://agentgateway.dev/llms.txt)).
- LangChain, LangGraph, LlamaIndex and similar orchestration frameworks with a documented integration.
- Callable from Cloudflare Workers Cloudflare Workers support
- Not documented
n.a. - it is a native Rust binary, not a Workers-compatible runtime, and no Cloudflare Workers deployment appears in the docs ([binary](https://agentgateway.dev/docs/standalone/latest/setup/install/binary/), [Docker](https://agentgateway.dev/docs/standalone/latest/setup/install/docker/), [Helm](https://agentgateway.dev/docs/standalone/latest/setup/install/helm/)).
- Whether the docs show calling this product from your own Worker. Deliberately separated from the several products whose own gateway runs on Workers, which is a fact about their infrastructure and not about your edge compatibility.
- Kubernetes install Helm chart availability
- Official Helm chart
Official OCI Helm charts: `helm upgrade -i agentgateway-standalone oci://cr.agentgateway.dev/charts/agentgateway-standalone --version v1.5.0` for standalone, and a separate CRDs + control-plane pair for Kubernetes mode with Gateway API custom resources, xDS and a read-only GitOps UI ([Helm](https://agentgateway.dev/docs/standalone/latest/setup/install/helm/), [Kubernetes control plane](https://agentgateway.dev/docs/standalone/latest/setup/install/kubernetes/)).
- Whether a named, published Helm chart exists, versus Helm being referenced with no chart named, versus generic cluster documentation that has nothing to do with this product.
- Terraform support Terraform provider or modules
- Not documented
n.a. (no Terraform provider, module or `terraform` reference on any page fetched, including the docs index: [docs index](https://agentgateway.dev/llms.txt), [Helm](https://agentgateway.dev/docs/standalone/latest/setup/install/helm/), [Kubernetes control plane](https://agentgateway.dev/docs/standalone/latest/setup/install/kubernetes/)).
- Whether you can declare this in code: an official provider, official modules, resources inside a hyperscaler’s provider, a community provider, or only Terraform code shipped in a repo.
- Reuses your cloud identity Cloud IAM reuse
- Reuses your cloud identity
native_iam for the three big clouds: AWS backend auth supports `auth.aws.assumeRole` with CEL/JWT-derived session names and tags, and Azure and GCP have their own backend-auth provider pages, so upstream calls can use cloud identity rather than static keys ([AWS integration](https://agentgateway.dev/docs/standalone/latest/integrations/cloud-providers/aws/), [invoice-grade attribution](https://agentgateway.dev/docs/standalone/latest/llm/cost-controls/attribution/)).
- Whether you can authenticate with IAM roles, workload identity or managed identities instead of another long-lived API key. Distinguished from products that only accept static upstream provider credentials.
- Fits behind your API gateway API gateway integration
- It is the API gateway
It is the gateway, and explicitly a general-purpose one: the FAQ positions it as an HTTP/gRPC data plane with load balancing, timeouts, retries, TLS, rate limits and authorization that can front ordinary APIs and microservices, so you do not run separate "regular" and "AI" gateways ([FAQs](https://agentgateway.dev/docs/standalone/latest/faqs/)).
- Whether AI traffic can go through a gateway you already run, and who documented that — the product’s vendor or the gateway’s.
- MCP support MCP surface shape
- MCP gateway or registry
A full MCP gateway, arguably the project's centre of gravity: static and dynamic MCP servers, stdio/SSE/streamable-HTTP transports, MCP authentication and authorization, tool-level access control, MCP-specific guardrails, MCP rate limiting, MCP observability, and OpenAPI-to-MCP bridging. MCP spec 2026-07-28 support landed in v1.4 ([MCP guardrails](https://agentgateway.dev/docs/standalone/latest/mcp/guardrails/setup/), [MCP observability](https://agentgateway.dev/docs/standalone/latest/mcp/mcp-observability/), [v1.4 announcement](https://agentgateway.dev/blog/2026-08-03-mcp-spec-2026-07-28/)).
- Which kind of MCP support this is: a gateway that governs many MCP servers, a hosted MCP server you connect to, MCP tools accepted inside the completion API, or client tooling. These are different products behind one acronym.
- Needs your own provider key Upstream provider key required
- Required
Yes, unavoidably: there is no credit pool or vendor account, so an upstream provider key (or a passthrough caller token, or a local Ollama/vLLM endpoint) is required before the gateway can serve a single request ([API keys](https://agentgateway.dev/docs/standalone/latest/llm/api-keys/), [quickstart](https://agentgateway.dev/docs/standalone/latest/quickstart/llm/)).
- Whether an upstream provider account and key must exist before your first call works. A prerequisite rather than a step, and it can differ between the hosted and self-hosted forms of the same product.
- Gate before models work Model access gate
- No gate
No gate on the gateway's side: any model your upstream credential can reach is reachable through a route the moment you add it. Access restrictions are ones you impose yourself, through virtual-key `allowedModels` or CEL authorization ([quickstart](https://agentgateway.dev/docs/standalone/latest/quickstart/llm/), [virtual keys](https://agentgateway.dev/docs/standalone/latest/llm/cost-controls/virtual-keys/)).
- Whether an enablement click, a quota grant, a paid tier or an approval form stands between a valid key and a working model call.
- Official client languages First-party client SDK languages
- Python, JavaScript
No first-party SDK exists. Documented clients are the standard OpenAI SDKs (`openai` for Python, `openai` for JavaScript) pointed at the gateway base URL, plus `curl`; the operator-side tooling is the `agctl` CLI ([OpenAI SDK](https://agentgateway.dev/docs/standalone/latest/integrations/llm-clients/openai-sdk/), [curl](https://agentgateway.dev/docs/standalone/latest/integrations/llm-clients/curl/)).
- Languages with a first-party client library. An empty list can still mean the product is usable from any language via an OpenAI-compatible SDK.
How pricing actually works
Self-hosting is the only mode and it is free of vendor charges. The binary is a single Rust process (13-22 MB resident in the project's own benchmark runs), so the practical floor is one small container plus, optionally, SQLite or PostgreSQL if you want USD budgets, the cost dashboard or database-backed request logs ([Database](https://agentgateway.dev/docs/standalone/latest/setup/database/), [benchmark](https://agentgateway.dev/blog/2026-06-26-benchmarking-agentgateway-vs-litellm-part-2/)).
Common questions
Answered from the fields above, so these move when the catalog moves. Every figure quoted here appears in the specification with its source.
Does agentgateway charge a markup on model prices?
agentgateway adds no percentage markup to model prices. The cost estimator on this site itemises these mechanisms against your own volume, because which one is cheapest depends entirely on the numbers you put in. With no per-token cut and no top-up charge, what you pay is the model providers' own rates plus whatever the infrastructure costs you to run.
Can agentgateway be self-hosted?
Yes — self-hosting is the only way to run agentgateway; there is no vendor-hosted option. The licence is Apache-2.0. Running it yourself means you supply the infrastructure and the upstream model accounts, so the bill is your own hosting plus the providers' own rates.
Is agentgateway SOC 2 audited, and will it sign a HIPAA BAA?
agentgateway publishes neither a SOC 2 report nor a HIPAA business associate agreement. Each of these is linked to the vendor's own page in the compliance section below. Neither absence means a refusal: both are things a vendor either publishes or does not, and smaller products often hold the certification without advertising it.
Does agentgateway retain your prompts?
Zero data retention does not apply to agentgateway: it runs inside your own infrastructure, so prompts never reach a vendor. Whether prompt and response bodies are logged is configurable. Logging can be turned off. Retention becomes your own configuration question instead, decided by whatever logging you switch on in your own deployment.
Can you use your own provider keys with agentgateway?
Yes. agentgateway can route through your own accounts with the underlying model providers, so inference is billed to you directly. Every request runs on your own upstream credentials, supplied inline, from an `$ENV_VAR`, from a file-backed env var, from a Kubernetes secret, or passed straight through from the caller's token; the project bills nothing and issues no keys of its own ([API keys](https://agentgateway.dev/docs/standalone/latest/llm/api-keys/)).
How many models does agentgateway support?
agentgateway states ~1,002 models, drawn from 20–44 upstream providers. The project's own Model and Provider Cookbook headlines "1002+ Models" across its provider catalogue, and any additional model is reachable through the custom/OpenAI-compatible provider, so the number is a floor rather than a closed list ([Model and Provider Cookbook](https://agentgateway.dev/models)). The figure on this page is dated and carries its source.
Official links
- Website agentgateway.dev ↗
- Documentation agentgateway.dev ↗
- Source code github.com ↗
- Changelog github.com ↗
4,985 GitHub stars — a proxy for community size, not for quality.
Independent coverage
Third-party analysis, walkthroughs and operator threads. We link criticism as readily as praise. Nothing here is written or published by the vendor, by a competitor listed on this site, or by an SEO content farm — how we vet these.
Written reviews and analysis 3
- Agentgateway Review: A Feature-Rich New AI Gateway Line-by-line source review of MCP session statefulness, multi-backend tool multiplexing, OpenAPI-to-MCP conversion limits, OAuth/JWKS handling and A2A support; the most technically specific outside assessment found. Author affiliation is not stated on the page.
- Analysis of the AI Agent Bus Gateway 'agentgateway' - Implementation Independent code-level walkthrough of the Cargo workspace, crate responsibilities, Tokio threading model and devcontainer setup, and openly questions whether another proxy is needed alongside Nginx, HAProxy and Envoy Gateway. Reports no benchmarks or functional tests.
- LiteLLM is the known option. agentgateway is the open one. Customer-facing evaluation of source trees, licences, published advisories and docs at agentgateway v1.4.1 vs LiteLLM 1.9x. Explicitly not a load test, and it pushes back on the project's own benchmark - noting the 12 GB LiteLLM memory figure conflicts with LiteLLM's own 359 MB result on a different harness. The author discloses being an AAIF Ambassador, so treat it as semi-affiliated.
Practitioner discussion 1
- Agentgateway: a fast, feature rich, Kubernetes native proxy Project-introduction thread in r/rust where maintainers field questions from Rust practitioners. Useful for community reaction, but the opening post is by a project contributor rather than an independent reviewer.
What has changed here
- GitHub stars GitHub stars 4845 4985 source ↗
- GitHub stars GitHub stars 4691 4845 source ↗
- catalog entry catalog entry Not published Added to the catalog source ↗
Read the head-to-head
These pairs have a written verdict, not just a table.