MLflow AI Gateway Open source
MLflow AI Gateway is an open-source LLM gateway with an OpenAI-compatible API in front of 14–100 upstream providers; it publishes no model count. It charges no token markup, credit-purchase fee or per-seat fee. It can be self-hosted under Apache-2.0 or used as a managed service. It does not publish a HIPAA BAA. You can point it at your own provider accounts. Beyond chat it also serves embeddings.
· 45 dated entries · 56 source references
Access at a glance
Whether this product can work for you at all, before features matter: who pays the model bill, where it can run, and which of your existing API calls keep working. Every value is the vendor’s own claim, linked to the page it came from.
One feature of a much larger platform, and the vendor frames it that way: "MLflow AI Gateway was built to fix this without the integration tax. Because it runs as part of the MLflow Tracking Server you're already using for tracing and evaluation, you get governed LLM access in the same place you debug traces and run evaluations". So platform-level facts — 30 million monthly downloads, 27,777 GitHub stars on mlflow/mlflow, autolog for 30+ frameworks — describe MLflow, not gateway adoption; the gateway itself only reached general usability in the 3.9–3.15 releases of 2026 (blog, AI Gateway, mlflow/mlflow API).
Who pays the model bill
Your keys onlyYou contract with each model provider directly and hold those accounts. The gateway never resells inference.
There is no platform credit or hosted key. Step 2 of the quickstart is creating an LLM Connection with "your API key from the provider", which is then encrypted in the MLflow backend store (Quickstart, API keys). The coding-agent flow is the one exception: those endpoints need no key because "the agent brings its own credentials" (Claude Code).
Merchant of record: The upstream LLM provider. Provider credentials are your own LLM Connections and the project bills nothing (Quickstart, LiteLLM alternative).
Key handling: Provider keys live in "LLM Connections" and are encrypted before being written to the MLflow backend database; a default passphrase is used for local development and production deployments must set MLFLOW_CRYPTO_KEK_PASSPHRASE. Rotation is zero-downtime via mlflow crypto rotate-kek --new-passphrase plus MLFLOW_CRYPTO_KEK_VERSION, and key inputs are masked in the UI (API keys, Key rotation).
Where it can run
3 of 5 shapes documented- Vendor-hosted
- Self-host
- Your VPC
- On-premise
- Air-gapped
Dimmed shapes are not documented by the vendor, which is not the same as unsupported.
self-host and on-prem are the first-class paths: pip/uvx locally, Docker Compose, or the official Helm chart on Kubernetes (Quickstart, Kubernetes and Helm). SaaS only via third parties: "Deploy on your own infrastructure or use managed versions on Databricks or AWS" — that sentence is about MLflow generally, and gateway parity on those platforms is not documented on the pages fetched (AI Gateway).
The gateway is a feature of the Tracking Server, not a deployable of its own — as of MLflow 3.0 "the MLflow deployment server application and the start-server CLI command have been removed" (#15327), so mlflow server is the only entry point. Production guidance is PostgreSQL plus S3/GCS/Azure storage, with TLS, ingress, Prometheus metrics, NetworkPolicy, RBAC and an mlflow gc CronJob available in the chart; public ingress also needs allowed_hosts set or requests return HTTP 403 (MLflow 3 breaking changes, Kubernetes and Helm, Network security).
API surfaces your code can keep using
4 of 7 documented- OpenAI chat
POST /v1/chat/completionsYesclient = OpenAI(base_url="http://localhost:5000/gateway/mlflow/v1", api_key="")thenchat.completions.create(model="my-chat-endpoint", ...); the API key is empty because credentials are held server-side (Quickstart, Query endpoints). - Anthropic messages
POST /v1/messagesYesAs a passthrough surface:
POST /gateway/anthropic/v1/messages, documented with the Anthropic Python SDK pointed atbase_url="https://your-mlflow-server/gateway/anthropic"and a dummy API key (Model providers, blog). - OpenAI Responses
POST /v1/responsesYesThe OpenAI passthrough "exposes the full OpenAI API" and lists
POST /responses— Responses API (multi-turn conversations) — alongside/chat/completionsand/embeddings(Model providers). - Embeddings
POST /v1/embeddingsYesEmbeddings are marked supported for OpenAI, Google Gemini, Azure OpenAI, AWS Bedrock, Vertex AI, Cohere, Mistral, Together AI, Fireworks AI, Ollama, Databricks and Portkey (Anthropic and Groq are marked No), and
POST /embeddingsis part of the OpenAI passthrough (Model providers). - Images
POST /v1/images/generationsNot documentedNo image-generation endpoint on the fetched gateway pages; "Vision" appears only as a model capability badge for image input (Model providers, Query endpoints).
- Audio
POST /v1/audio/*Not documentedNo STT/TTS path documented (Model providers, Query endpoints).
- Batch jobs
POST /v1/batchesNot documentedNo batch or async bulk endpoint documented (Query endpoints, Model providers).
An asterisk marks a qualified verdict: support that is indirect (SDK compatibility or provider passthrough rather than a native endpoint), or a gap that is narrower or wider than the label suggests. Read the note before porting.
Two shapes on one server. Unified: POST /gateway/{endpoint}/mlflow/invocations, plus an OpenAI-compatible base URL at /gateway/mlflow/v1 where the gateway endpoint name is passed as model. Passthrough: provider-native paths such as /gateway/openai/v1/chat/completions, /gateway/anthropic/v1/messages and /gateway/gemini/v1beta/models/{endpoint}:generateContent, so you keep the provider's own SDK while the gateway still holds the credentials and records usage (Query endpoints, Model providers).
How much it reaches
Providers: Counted from the vendor’s own published list, which has no aggregate total. The range spans a narrow reading (production-ready only) and a broad one (including preview and system entries).
n.a. — no model count on any fetched page (AI Gateway landing, gateway docs, Model providers); models are whatever the configured provider exposes, plus the optional LiteLLM catalogue.
Three different figures on vendor pages: 14 providers enumerated in the docs tables (OpenAI, Anthropic, Google Gemini, Azure OpenAI, AWS Bedrock, Vertex AI, Cohere, Mistral, Groq, Together AI, Fireworks AI, Ollama, Databricks, Portkey) (Model providers); "Select your provider from 100+ supported options" in the endpoint-creation UI (Create and manage endpoints); and "Access 50+ Model Providers" on the product page (AI Gateway). The enumerated 14 is used here; the larger numbers depend on the optional LiteLLM catalogue, which "you will need to install separately".
Whose models: All third-party: the project hosts no models. Endpoints forward to provider APIs (OpenAI, Anthropic, Gemini, Azure OpenAI, Bedrock, Vertex AI, Cohere, Mistral, Groq, Together AI, Fireworks AI, Databricks, Portkey) or to your own Ollama/OpenAI-compatible endpoint (Model providers).
Your own endpoints: Yes: the product page states the gateway routes to "any LLM provider — including any OpenAI-compatible API or custom model endpoint", and Ollama is documented for local models (AI Gateway, Model providers).
How it behaves in production
What happens when an upstream model is slow, wrong, or down — and what you can see and stop while it happens. Reliability features are recorded as where you configure them, not whether the vendor lists them, because almost every product here lists all of them.
0 of 6 reachable from code 3 of 4 can block 5 documented destinations
When something goes wrong
Each row says where the knob is, not whether the feature is on the marketing page. A control you can only reach by hand in someone else’s dashboard cannot be reviewed or version-controlled.
- Request timeout Not documented
The vendor does not document this, so any behaviour you observe today is unversioned and may change.
For gateway requests: no per-endpoint or per-request timeout key appears on the routing, endpoint or query pages. The only timeouts discussed are artifact upload/download through the tracking server's artifact proxy (Traffic routing and fallbacks, Query endpoints, Tracking server).
- Retries Not documented
No retry count, backoff strategy or default is published. The documented failure behaviour is sequential fallback rather than same-model retry: "the gateway tries fallback models sequentially until one succeeds". Default retry count: n.a. Backoff: n.a. (Traffic routing and fallbacks).
- Fallback to another model Dashboard only
Only reachable by hand in the vendor UI, so it cannot be reviewed, version-controlled, or changed from code.
ORDERED and UI-configured: fallbacks are added under "Priority 2 (Fallback)" on the endpoint page, reordered by dragging, and tried in sequence when the primary hits errors or rate limits. Each fallback carries its own provider, model and API key, so a fallback can point at a different provider or region (Traffic routing and fallbacks).
- Load balancing Dashboard only
Weights are user-settable but UI-only: traffic splitting assigns each model a 1–100% weight and the weights "must sum to exactly 100", intended for A/B tests and gradual rollouts. No latency- or cost-aware balancing is documented (Traffic routing and fallbacks).
- Upstream health tracking Not documented
No active health check, circuit breaker or unhealthy-model ejection is described; failover is reactive, triggered by the error on the request itself (Traffic routing and fallbacks).
- Cross-region failover Not documented
As a gateway feature. "Regional failover: Route to providers in different geographic regions" is listed as a fallback use case, but that is provider selection inside one gateway rather than a multi-region gateway deployment; no cross-region MLflow topology is described (Traffic routing and fallbacks, Kubernetes and Helm).
Fallback chain: Ordered list — Try A, then B, then C. Simple and predictable, but every failover is all-or-nothing.
Endpoint changes are hot: "Add, remove, or modify endpoints dynamically without restarting the server or disrupting running applications", and key rotation is likewise zero-downtime. The gap is everything below routing — no timeouts, retry counts, health checks or circuit breakers are published, and budget enforcement can lag because the default tracker refresh interval (MLFLOW_GATEWAY_BUDGET_REFRESH_INTERVAL) is 600 seconds and the default local tracker keeps counters per process rather than sharing them (gateway docs, Key rotation, Budget alerts and limits).
How fast the hop is
Interpreted proxyRuns on an interpreted or JIT runtime (Lua, Python, Node). Overhead is higher than a compiled binary and more sensitive to concurrency, though a Lua-on-nginx proxy and a Python one are far apart.
Python (the repo's primary language per the GitHub API) running on the FastAPI/uvicorn tracking server; the gateway is not a separate process. It is a database-backed proxy: it requires a SQL backend store (SQLite, PostgreSQL, MySQL or MSSQL) plus the FastAPI server, and file-based tracking stores are not supported. Note that the 27,777 stars on the GitHub API are for all of mlflow/mlflow — the whole ML/GenAI platform — and are not a measure of gateway adoption (mlflow/mlflow API, Quickstart).
PyPI extra plus an official OCI Helm chart: pip install 'mlflow[genai]' then mlflow server --port 5000, or helm install mlflow oci://ghcr.io/mlflow/charts/mlflow --version <version> --namespace mlflow --create-namespace (Kubernetes 1.23+, Helm 3.8+) (Quickstart, Kubernetes and Helm).
Streaming caveats: Supported over SSE by setting stream: true. Two documented caveats: post-LLM guardrails "are not triggered for streaming requests", and the X-MLflow-Gateway-Overhead-Duration-Ms header is only emitted for non-streaming responses (Query endpoints, Guardrails, Benchmarks).
Published figures, grouped by what each one measured. Figures in different groups are different quantities and cannot be compared with one another — nor, in most cases, with another vendor’s figure in the same group.
Overhead added by the gateway
The only figures that speak to whether this product slows your application down. Still not comparable between vendors: each was measured on different hardware, at a different load, with a different payload.
- 28.6 ms p50 Vendor-published
Vendor benchmark against LiteLLM. 50 ms simulated provider delay, 4 workers, 50 concurrent users; overhead = measured latency minus the simulated delay. Hardware not stated.
Source - 134.2 ms p99 Vendor-published
Same vendor benchmark run (50 ms simulated provider delay, 4 workers, 50 concurrent users).
Source - single-digit-to-tens-of-milliseconds ms range Vendor-published
Docs prose, not a measured result table. The gateway emits `X-MLflow-Gateway-Duration-Ms` on every response and `X-MLflow-Gateway-Overhead-Duration-Ms` on non-streaming responses; the worked example shows duration 87 ms with 3 ms overhead. Test rig: fake OpenAI server with a fixed 50 ms delay, 4 MLflow instances behind nginx, PostgreSQL, 50 concurrent requests; explicitly excludes provider variance, network, TLS and auth.
Source
Sustained capacity
Requests or queries per second sustained on the stated hardware.
- 598 req/s sustained Vendor-published
Vendor benchmark against LiteLLM (358 req/s), described as "67% higher throughput". 50 ms simulated provider delay, 4 workers, 50 concurrent users; hardware not stated.
Source
Not independent. Every figure is MLflow's own, published on a page whose purpose is to position MLflow against LiteLLM, and it names a competitor's numbers without an independent replication (LiteLLM alternative). The docs benchmarks page is transparent about its harness but publishes no result table (Benchmarks).
What it will stop
3 of 4 can blockTwo separate questions per control: can it stop a request at all, and what does it do before you change any settings? A control that inspects and forwards is a logging feature, however it is named.
- Personal data in prompts Can block the request
Out of the box: You pick the action when configuring
A PII Detection guardrail type is built in, defaults to the Pre-LLM stage, and its action is chosen at setup: Block returns HTTP 400 naming the guardrail and its rationale, or Sanitize redacts the offending content and lets the request continue (Guardrails).
- Prompt injection and jailbreaks Not documented
n.a. — no prompt-injection or jailbreak guardrail type is offered; the three documented types are Safety, PII Detection and Custom, so injection detection would have to be written as a Custom judge prompt (Guardrails).
- Harmful content Can block the request
Out of the box: You pick the action when configuring
A Safety guardrail type ships built in and defaults to the Post-LLM stage, with the same Block (HTTP 400) or Sanitize actions. Guardrails apply to both unified and passthrough endpoints (Guardrails).
- Your own policies Can block the request
Out of the box: You pick the action when configuring
Custom guardrails are natural-language judge instructions rather than regex: you pick the stage (Pre-LLM or Post-LLM), the judge model (any other gateway endpoint) and the action (Block or Sanitize). Multiple guardrails run in table order and later ones are skipped once one blocks (Guardrails).
Almost no vendor in this catalogue documents what happens when the guardrail service itself times out. If the control matters to you, this is a question worth asking before you sign.
not_documented — the guardrails page describes Block and Sanitize outcomes but says nothing about what happens when the judge endpoint itself errors or times out (Guardrails).
The load-bearing caveat is streaming: "post-LLM guardrails are not triggered for streaming requests", so a Safety guardrail left on its Post-LLM default silently stops enforcing as soon as a client sets stream: true. Guardrails are also LLM-judge based, which means each guarded request costs an extra model call to whichever endpoint backs the judge (Guardrails).
What you can see
Exports to a few placesYou decide whether bodies are captured, by setting or by header.
Yes — it is opt-in rather than opt-out: usage tracking is a toggle on each endpoint, and the docs note token and cost metrics are unavailable for some providers and models even when it is on (Usage tracking).
Gateway requests become MLflow traces server-side with no client instrumentation, and the trace layer is OpenTelemetry-based: OTLP export via OTEL_EXPORTER_OTLP_TRACES_ENDPOINT, simultaneous dual export with MLFLOW_TRACE_ENABLE_OTLP_DUAL_EXPORT=true, GenAI semantic conventions via MLFLOW_ENABLE_OTEL_GENAI_SEMCONV, and W3C TraceContext linking client-side agent traces to gateway traces (OpenTelemetry export, AI Gateway tracing integration, LiteLLM alternative).
configurable: off unless usage tracking is enabled on the endpoint, and full request/response payloads once it is. PII can be stripped before storage with client-side span processors registered through mlflow.tracing.configure(span_processors=[...]) (Usage tracking, Masking).
Where telemetry can go
- OpenTelemetry (OTLP)
- Prometheus
- S3
- GCS
- Azure Blob Storage
OTLP/HTTP for traces and metrics, with dual export so an existing collector keeps receiving spans while MLflow stores them (OpenTelemetry export); Prometheus metrics and a ServiceMonitor from the official Helm chart (Kubernetes and Helm); and S3/GCS/Azure object storage as the trace-archival and artifact destination (Tracking server).
Partial and platform-level, not gateway-level: no feedback widget or scoring call is documented on the gateway pages, but MLflow advertises "LLM judge alignment with human feedback" and the MLflow MCP server can "log feedback and assessments" against traces (LiteLLM alternative, MCP server, checked against Usage tracking).
Yes, and this is the whole pitch: "Traces captured through the gateway feed directly into mlflow.genai.evaluate or Evaluation Dataset APIs, so you can run judges over production traffic without any additional instrumentation", and a gateway endpoint can itself be referenced as a judge model with gateway:/my-chat-endpoint (blog, Query endpoints).
Depends on the vendor’s SaaS: No — the usage dashboard, traces and evaluation all run inside the MLflow server you deployed, and nothing is gated behind a hosted control plane (Usage tracking, blog). One caveat published by the project: cost computation on Databricks managed MLflow "requires the client application to install LiteLLM or manually set the cost attributes on spans", which is not required self-hosted (Token usage and cost).
Retention: Determined by your own archival policy and deletion calls; the documented example archives span payloads to object storage after 30 days while keeping traces readable in the UI (Tracking server).
Whether it fits how you work
How much work stands between you and a first call, how different that is from running it in production, and whether it slots into the stack you already have. Recorded as the shape of the work rather than a number of minutes — how long it takes you depends on which accounts and quota you already hold, which no comparison can know.
to try: cli or container to run: infrastructure rollout fits 4 of 10 common stacks
Getting to a first call
4 numbered stepsNothing works until you have a process running. Fine on a laptop, but it means there is no zero-install way to try it.
Read off: the vendor’s own quickstart — 4 numbered steps.
Four numbered steps (install and start, create an LLM Connection, create an endpoint, query it), but steps 2 and 3 are five-click UI sub-flows each rather than commands (Quickstart).
Before step one
- Your own provider key Required
You need an upstream provider account and key before anything works. That is a prerequisite, not a step.
Nothing works until you create an LLM Connection with your own provider key (quickstart step 2), because the project hosts no models and issues no credits. Exception: coding-agent endpoints, where "the agent brings its own credentials" (Quickstart, releases index).
- Payment method No card needed to start
No card, account or licence for the gateway itself —
pip install 'mlflow[genai]'andmlflow server --port 5000is the whole install. You will need a paid provider key to get a completion out of it (Quickstart). - Gate before models answer No gate
Every catalogue model is callable as soon as you have a key.
From the project: there is no approval, quota, waitlist or enablement step — you pick a provider and model in the endpoint dialog and supply your own key. Access control is yours to impose, via RBAC on gateway endpoints (Create and manage endpoints, RBAC).
Everything you need first: Python with pip install 'mlflow[genai]', plus one provider API key. The default mlflow server uses SQLite and FastAPI, "so no additional configuration is needed for this quickstart" (Quickstart).
The vendor’s own time claim: Vendor claim, verbatim: "Set up governed LLM access in minutes. No additional infrastructure required", broken into per-step estimates of ~30 seconds to start the server, ~1 minute to create an endpoint and ~30 seconds to query it (AI Gateway). Quoted, not verified. Marketing time claims assume every account and approval is already in place.
Running it in production
The same scale applied to the path the vendor recommends for production traffic. Kept separate from the quickstart because for several products here the two are barely related pieces of work.
This is an infrastructure project, not an integration. Expect charts or Terraform, networking, secrets and someone who owns the deployment.
This is an infrastructure project, not an integration. Expect charts or Terraform, networking, secrets and someone who owns the deployment.
Getting to production is a step up in kind from the quickstart, not just more of the same.
What production needs: MLflow server with the [genai] extra, a SQL backend store (PostgreSQL or MySQL recommended over the SQLite default at concurrency), object storage for artifacts and trace archival, and — for real access control — pip install 'mlflow[auth]', MLFLOW_FLASK_SERVER_SECRET_KEY, mlflow server --app-name basic-auth and MLflow 3.13.0+ for RBAC. Shared budget enforcement across processes needs Redis via MLFLOW_GATEWAY_BUDGET_REDIS_URL, and production key encryption needs MLFLOW_CRYPTO_KEK_PASSPHRASE (Tracking server, Basic HTTP auth, RBAC, Budget alerts and limits, API keys).
Can you run it yourself
There is a command you can copy and run, so you can evaluate the self-hosted path yourself today.
pip install 'mlflow[genai]' then mlflow server --port 5000; on Kubernetes helm install mlflow oci://ghcr.io/mlflow/charts/mlflow --version <version> --namespace mlflow --create-namespace
How it fits your stack
4 of 10Each row is a thing you might already run. “With a caveat” means it works but not the way the vendor’s marketing implies — read the reason, because that is usually where the surprise lives.
- Fits The OpenAI SDK Drop-in once set up — but first-call work is cli or container.
- No The Vercel AI SDK No AI SDK route documented.
- No Cloudflare Workers No Workers guidance published.
- Fits Kubernetes official chart at oci://ghcr.io/mlflow/charts/mlflow
- No Terraform or OpenTofu Nothing published for Terraform.
- Fits An existing API gateway This is that gateway — AI traffic becomes a plugin, not a new hop.
- No Cloud IAM I already run Static upstream credentials only. Your calls to it still use its own key.
- Fits LangChain or LlamaIndex LangChain, LangGraph, DSPy, OpenAI Agents SDK, LiteLLM
- With a caveat MCP servers to govern Hosted MCP server — governs nothing on your side.
- With a caveat Nothing — plain Node or Python You have to run a process locally before any call works.
Reading this the other way round — pick what you already run and see every product scored against it.
The integration surfaces behind those answers
- Vercel AI SDK Not documented
Nothing published. Assume the OpenAI-compatible route and verify it yourself.
Vercel AI SDK appears in MLflow's docs only as a tracing integration — the
aipackage exporting OpenTelemetry spans into MLflow via@vercel/otelandOTEL_EXPORTER_OTLP_ENDPOINT— not as a way to call gateway endpoints, and no gateway provider package orbaseURLexample is documented (Vercel AI SDK tracing, checked against Query endpoints). - Cloudflare Workers Not documented
No Workers guidance either way. If you are edge-first, verify fetch-only compatibility yourself.
n.a. (not documented) — the gateway is a Python server; no edge-worker deployment appears on the self-hosting pages (Self-hosting, Kubernetes and Helm).
- Kubernetes Official Helm chart
A named, published chart. You can read its values file before committing to anything.
Named:
official chart at oci://ghcr.io/mlflow/charts/mlflowOfficial OCI Helm chart:
helm install mlflow oci://ghcr.io/mlflow/charts/mlflow --version <version> --namespace mlflow --create-namespace, requiring Kubernetes 1.23+ and Helm 3.8+, with TLS, ingress, Prometheus ServiceMonitor, NetworkPolicy, RBAC and anmlflow gcCronJob, and PostgreSQL plus S3/GCS/Azure recommended for production (Kubernetes and Helm). - Terraform Not documented
No Terraform surface published. Configuration is API or dashboard work.
n.a. (not documented) — no Terraform provider or module is referenced on the self-hosting, Kubernetes or gateway pages fetched (Self-hosting, Kubernetes and Helm).
- Existing API gateway It is the API gateway
This product is the gateway. If you already run it for your other APIs, AI traffic becomes a plugin rather than a new hop.
It is itself the gateway — "a centralized proxy layer that routes requests to LLM providers through a single, unified API" — but uniquely here it is a gateway embedded in an MLOps platform rather than a standalone proxy (AI Gateway, blog). It can also front another gateway: Portkey is a supported provider (Model providers).
- Cloud identity Static provider credentials only
You paste static provider credentials into this product so it can reach upstream models. Your own calls to it still use its own API key, and those pasted secrets are yours to rotate.
AWS Bedrock and Vertex AI are listed as supported providers, but the passthrough/authentication detail on the fetched pages is limited to API-key style LLM Connections; no IAM role, SigV4 or workload-identity configuration is documented for the gateway itself (Model providers, API keys).
- MCP Hosted MCP server
The vendor runs an MCP server you connect a client to. Useful for reaching this product from an agent, but it does not govern your other MCP servers.
Two separate things, both experimental. MLflow ships an MCP server (3.5.1+) that exposes trace management to Claude, Cursor and VS Code — search traces, analyse performance, log feedback, manage tags, delete traces — and an MCP Registry ("experimental feature introduced in MLflow 3.15.0") for registering and versioning MCP servers via
server.jsonandmlflow.genai.register_mcp_server()(MCP server, MCP Registry). The landing page's stronger claim — that the gateway governs "which MCP servers your agents can reach" — has no matching page in the AI Gateway docs section (AI Gateway).
The query page ships copy-paste snippets for LangChain (ChatOpenAI with the gateway base URL), LangGraph (create_react_agent), DSPy (dspy.LM), the OpenAI Agents SDK and LiteLLM (litellm.completion), all pointed at /gateway/mlflow/v1. Separately, MLflow advertises one-line autolog() tracing for 30+ frameworks — a platform feature, not a gateway one (Query endpoints, LiteLLM alternative).
Documented client code is Python and cURL: requests, the OpenAI Python SDK against /gateway/mlflow/v1, the Anthropic Python SDK against the passthrough base URL, and google-genai. No JavaScript/TypeScript gateway example appears on the fetched pages, though any OpenAI-compatible client works by construction (Quickstart, Query endpoints, Model providers).
Agent features: Tool calling and structured output are first-class on the unified endpoint: a tools array and response_format accepting text, json_object or json_schema. Coding agents are an explicit use case, with one-click endpoints and documented setups for Claude Code (ANTHROPIC_BASE_URL=http://localhost:5000/gateway/proxy/claude-code), OpenAI Codex, Gemini CLI and Hermes Agent, each conversation captured as a trace and subject to the same guardrails and budgets (Query endpoints, Coding agents, Claude Code).
Four steps to a first call, all local, but two constraints bite early: the gateway needs a SQL-backed store (SQLite, PostgreSQL, MySQL or MSSQL) and the FastAPI tracking server — file-based tracking stores are not supported — and usage tracking is off until you toggle it per endpoint, so the observability that justifies the product is opt-in (Quickstart, Usage tracking).
Apache-2.0 and Linux Foundation governed, so there is no vendor account in the path, and the same server carries tracing, evaluation, prompt registry and model registry. The trade is operational: you own PostgreSQL, object storage, TLS, ingress and auth, and the gateway inherits MLflow's server surface — including its authentication model, which out of the box is HTTP basic auth (LiteLLM alternative, Kubernetes and Helm, Basic HTTP auth).
Silence in the docs: 8 of the integration questions on this card have no published answer either way. That is recorded as undocumented, not as a no — but it does mean you would be verifying it yourself.
What it does well
- Gateway, tracing, evaluation and prompt registry on one Apache-2.0 server: every request becomes an MLflow trace with no extra instrumentation
- Both unified OpenAI-compatible and provider-native passthrough surfaces (OpenAI, Anthropic Messages, Gemini generateContent) on the same server
- Provider keys encrypted in the backend store with a KEK passphrase and zero-downtime `mlflow crypto rotate-kek` rotation
- LLM-judge guardrails (Safety, PII, Custom) with Block or Sanitize actions, plus USD budget policies that alert by webhook or reject with HTTP 429
- Real RBAC over gateway secrets, endpoints and model definitions from MLflow 3.13.0, and an official OCI Helm chart for Kubernetes
- Foundation-governed with no vendor account, no per-request fee and no token markup
Where it falls short
- No documented request timeout, retry count, health check or circuit breaker — failover is reactive fallback only
- Post-LLM guardrails are not triggered for streaming requests, so a Safety guardrail on its default stage stops enforcing when a client sets stream: true
- "Rate Limiting" is ticked in MLflow's own comparison table but no rate-limit mechanism is documented anywhere in the gateway docs — only dollar budgets
- MLflow 3.0 removed the gateway config keys and the standalone deployment server, so endpoint management is UI/database-driven with no documented config-as-code path
- No gateway response cache of any kind, and no usage or cost export beyond reading traces
- Usage tracking is opt-in per endpoint, and the default `local` budget tracker keeps counters per process with a 600-second refresh, so spend caps can lag under multi-worker deployments
- Out-of-the-box authentication is HTTP basic auth; SSO needs a community OIDC plugin or a reverse proxy
- All published performance numbers are MLflow's own, from a page comparing itself with LiteLLM
Choose it when
Teams already running MLflow for tracing and evaluation who want governed, OpenAI-compatible LLM access on the same server — no second system to deploy, and every gateway request lands as a trace they can evaluate.
Look elsewhere when
You want a lean standalone proxy, or you need documented rate limits, request timeouts, retry policy, health checks, response caching or config-as-code endpoint management — none of which the gateway publishes — or you need post-LLM guardrails to hold on streaming responses.
Open source: You run the gateway yourself. Full control over the data path and no per-request vendor fee, in exchange for operating it.
How hard is it to leave?
Derived from six published facts, not from an opinion. The weights are fixed and the same for every product — see the arithmetic.
| What helps you leave | Points | Source |
|---|---|---|
| Works with standard OpenAI code Switching away is a base-URL change rather than a rewrite of every call site. | 22 /22 | vendor page |
| No vendor-specific SDK required A proprietary client library spreads through your codebase and has to be torn out again. | 10 /10 | vendor page |
| Can use your own provider accounts Your keys and billing relationship stay yours, so removing the gateway does not cut off model access. | 20 /20 | vendor page |
| Can be self-hosted You can run it yourself instead of accepting a pricing or policy change. | 20 /20 | vendor page |
| Configuration lives in version control Routing and budget rules are a file you keep, not dashboard state you would have to rebuild. | 0 /16 | vendor page |
| Your request history can be exported You leave with your own logs instead of abandoning them. | 12 /12 | vendor page |
Read the fine print: Traces, token usage and cost live in your own SQL backend and can be read through the Python client, dual-exported over OTLP to another collector, or archived to your object store; there is no vendor lock on the data ([OpenTelemetry export](https://mlflow.org/docs/latest/genai/tracing/opentelemetry/export/), [Token usage and cost](https://mlflow.org/docs/latest/genai/tracing/token-usage-cost/), [Tracking server](https://mlflow.org/docs/latest/self-hosting/architecture/tracking-server/)).
All six inputs are published, so this is scored against the full 100. This measures technical switching cost only. It does not price the engineering time to re-test prompts against a different routing stack.
MLflow AI Gateway models & pricing
Browse every imported listing from this provider, with published token rates and a link to compare other providers for the same model. This is provider-reported coverage; an absent listing does not mean unsupported.
Loading model listings…
Official model coverage source ↗ · Model source coverage and limitations
Full specification
Every field we track. Blank fields say "Not published" rather than "No" — we do not infer an absence from silence. Switch to Technical in the header for the precise field names and the low-level details.
Overview
Interpret these fields: Self-hosted vs managed LLM gateways · Who still owns your LLM gateway?
- What kind of product Category
- Open source
- Marketplaces resell many providers behind one key. Gateways add governance on top. Open-source projects you run yourself. Cloud platforms are hyperscaler surfaces. Inference providers host models on their own hardware.
- Who runs it Deployment model
- Managed or self-host
- Managed means the vendor operates it. Self-host means you run it on your own infrastructure. Both means you can choose.
- Licence Licence
- Apache-2.0
- Proprietary products cannot be inspected or forked. Open licences such as MIT and Apache-2.0 let you audit, modify, and run the code without permission.
- Who you would be signing with Vendor status
- Run by a software foundation
- Whether the product is still an independent company, has been acquired, is a large cloud vendor’s product line, is run by a software foundation, or has been put into maintenance mode. Maintenance mode means bug fixes and security patches only — no new features.
- Last shipped an update Latest release
- 2026-08-26
v3.15.2, published 2026-08-26 per the GitHub releases API; the releases index describes 3.15.2 as a patch release. Repo `pushed_at` was 2026-09-02 ([GitHub releases API](https://api.github.com/repos/mlflow/mlflow/releases/latest), [releases index](https://mlflow.org/releases/)).
- The date of the most recent release or version tag. A product that has not shipped in a year is a different risk from one that shipped last week, regardless of what its marketing site says.
- GitHub stars GitHub stars
- 27,777
- A rough proxy for community size on open-source projects. Not a quality measure.
Cost
Interpret these fields: How LLM gateway pricing works · LLM gateway spending limits: stop a runaway agent bill?
- Markup on model prices Token markup
- None
- How much the product adds on top of what the underlying model provider charges. Zero means you pay the same per-token price you would pay the model provider directly.
- Fee to add funds Credit purchase fee
- None
- A percentage charged when you top up your balance, separate from token prices. It is easy to miss because it does not appear on the per-token price list.
- Monthly cost per person Seat fee
- None
- A recurring per-user platform charge that applies regardless of how much you use the models.
- Can use your own provider accounts BYOK supported
- Yes
- Bring Your Own Key: you keep direct contracts with OpenAI, Anthropic and others, and the gateway only routes traffic. This preserves negotiated rates and committed-spend discounts.
- Cost of using your own accounts BYOK terms
- You store your own provider keys as LLM Connections inside your MLflow server and pay the provider directly; the project charges nothing for the gateway ([Quickstart](https://mlflow.org/docs/latest/genai/governance/ai-gateway/quickstart/)).
- What the product charges to route traffic through your own provider keys.
- Free tier Free tier
- The whole thing is free: "MLflow is open source under the Apache 2.0 license and governed by the Linux Foundation" with no paid gateway tier, and the gateway installs with `pip install 'mlflow[genai]'` ([LiteLLM alternative](https://mlflow.org/litellm-alternative/), [Quickstart](https://mlflow.org/docs/latest/genai/governance/ai-gateway/quickstart/)).
- What you can do without paying, useful for evaluation.
- Enterprise plan from Enterprise plan from
- Not published
- Annual entry price for the enterprise tier, where one is published or credibly reported.
- Cost to run it yourself Self-host cost
- No licence cost; your cost is the MLflow Tracking Server plus a SQL backend store (SQLite, PostgreSQL, MySQL or MSSQL) and, for production, object storage and optionally Redis for shared budget counters ([Quickstart](https://mlflow.org/docs/latest/genai/governance/ai-gateway/quickstart/), [Tracking server](https://mlflow.org/docs/latest/self-hosting/architecture/tracking-server/), [Budget alerts and limits](https://mlflow.org/docs/latest/genai/governance/ai-gateway/budget-alerts-limits/)).
- What self-hosting actually costs once you account for infrastructure and any paid tier.
- How the vendor makes money Pricing model
- Open source, no paid tier
- The shape of the vendor’s bill: does the routing layer charge a percentage on top of tokens, a flat monthly fee, both, neither (because inference is the product), or nothing at all (open source with no paid tier).
- How pricing works, briefly Pricing model detail
- Apache-2.0 project under Linux Foundation governance with no pricing page, tier list or paid SKU for the gateway; the landing page contrasts itself with "per request or per seat" SaaS gateways and states "no per-request fees, no usage limits". Managed MLflow is sold by third parties (Databricks, AWS), not by the project ([AI Gateway](https://mlflow.org/ai-gateway), [LiteLLM alternative](https://mlflow.org/litellm-alternative/)).
- A one-paragraph description that covers the caveats a pricing category cannot: introductory rates, per-feature meters, tier gating, and pricing that resets on a specific date.
- Minimum commitment Minimum commitment
- None: the gateway is Apache-2.0 software you run yourself, with no contract, account or licence key ([LiteLLM alternative](https://mlflow.org/litellm-alternative/)).
- Whether the vendor requires a minimum contract term, a minimum spend, or a provisioned-capacity purchase to get its published rate.
- Charges that fire after you go over an allowance Overage terms
- No vendor meter exists, so there is nothing to overrun. Spend against providers is capped by your own budget policies, which either alert or reject with HTTP 429 ([Budget alerts and limits](https://mlflow.org/docs/latest/genai/governance/ai-gateway/budget-alerts-limits/)).
- The line items that scale with usage after an included allowance is exhausted — log storage, extra requests, per-feature meters, data export — which is where cost estimates usually go wrong.
- Prompt cache offered Cache mechanism
- No gateway-owned cache
- Whether the gateway offers its own response cache, what kind of match it does (exact request, prefix, semantic), or simply passes provider caching through unchanged.
- Discount on cached input Cache-read discount
- Not published
- How much cheaper cached tokens are than fresh input, when the vendor publishes a single figure. Bundled-inference clouds usually price this per model instead of as one number.
- Premium on cache writes Cache-write premium
- Not published
- How much more the first write of a cached prefix costs versus a plain input token. A high write premium and a low hit rate can leave you paying more than you save, so this matters as much as the read discount.
- Who captures the cache saving Cache economics
- No gateway-owned response or semantic cache appears on any fetched gateway page; "Caching" appears only as a per-model capability badge meaning "Model supports prompt caching for efficiency", and the benchmarks page mentions internal config caching, not response caching ([Model providers](https://mlflow.org/docs/latest/genai/governance/ai-gateway/endpoints/model-providers/), [Create and manage endpoints](https://mlflow.org/docs/latest/genai/governance/ai-gateway/endpoints/create-and-manage/), [Benchmarks](https://mlflow.org/docs/latest/genai/governance/ai-gateway/benchmarks/)). Because the gateway never prices tokens, any provider-side cache discount reaches you unchanged.
- Whether the customer keeps the full saving from caching or the vendor captures part of it — and any conditions attached (write premium, storage fees, best-effort hits).
- What you can split spend by Cost attribution
- Per endpoint, provider and model: the usage dashboard breaks down requests, latency percentiles, token usage, tokens per request, cost breakdown and cost over time, filterable by endpoint and time range. Budget policies scope spend globally or per workspace. Per-user or per-API-key cost attribution is not documented ([Usage tracking](https://mlflow.org/docs/latest/genai/governance/ai-gateway/usage-tracking/), [Budget alerts and limits](https://mlflow.org/docs/latest/genai/governance/ai-gateway/budget-alerts-limits/)).
- The dimensions the vendor documents for splitting spend — per key, per user, per team, per tag, per customer. Matters if you need to chargeback internally or bill an end customer.
- How you get cost data out Cost export
- No CSV or warehouse cost export is documented. Cost lives on the traces themselves — `trace.info.cost` gives input/output/total USD per trace and `span.llm_cost` per LLM call through the Python SDK — and traces can be exported over OTLP ([Token usage and cost](https://mlflow.org/docs/latest/genai/tracing/token-usage-cost/), [OpenTelemetry export](https://mlflow.org/docs/latest/genai/tracing/opentelemetry/export/)).
- The mechanisms the vendor publishes for exporting cost and usage data: CSV, an API, webhooks, S3, a data warehouse, or nothing at all. Any per-unit price is included.
- Who pays the model bill BYOK mode
- Your keys only
- Whether you bring your own provider accounts (BYOK), buy inference from this vendor, or can do either. This is the single biggest commercial difference between these products: it decides who holds the contract with the model provider and who carries the spend.
Spend governance in detail
Seven signals matter when a bill starts to hurt: who can spend, how much, on what, and who gets paged when it goes wrong. Everything below is drawn from the vendor’s own pricing and docs pages — how we read these.
- Virtual or scoped keys Not published
Keys that carry their own budget and rate-limit policy, so an intern experiment cannot spend against a production budget.
Not documented as such. "LLM Connections" hold upstream provider keys, not per-consumer virtual keys; client-side access is controlled by MLflow authentication and RBAC rather than issued gateway keys ([API keys](https://mlflow.org/docs/latest/genai/governance/ai-gateway/api-keys/create-and-manage/), [RBAC](https://mlflow.org/docs/latest/self-hosting/security/role-based-access-control/)).
- Budget caps per key Not published
A dollar or token ceiling attached to an individual key. Where enforcement is soft, one over-limit request still completes before the block kicks in.
Not documented: budget policies are scoped globally or per workspace, not per key or per user ([Budget alerts and limits](https://mlflow.org/docs/latest/genai/governance/ai-gateway/budget-alerts-limits/)).
- Budget caps per team or workspace Yes
A ceiling applied at a higher scope than one key — a team, a workspace, a customer, or an entire environment.
Yes, via workspace scoping: a USD threshold over a daily, weekly or monthly window, applied globally or to a single workspace, with the Reject action returning HTTP 429 ([Budget alerts and limits](https://mlflow.org/docs/latest/genai/governance/ai-gateway/budget-alerts-limits/)).
- Rate limiting as a cost control Not published
Configurable request-per-time-window caps. Platform-set rate limits do not count as spend controls; user-configurable ones do.
Contradictory: the vendor comparison table ticks "Rate Limiting" for MLflow, but no request-rate or token-rate limit is documented on any gateway docs page fetched — only dollar budgets ([LiteLLM alternative](https://mlflow.org/litellm-alternative/) vs [Budget alerts and limits](https://mlflow.org/docs/latest/genai/governance/ai-gateway/budget-alerts-limits/), [gateway docs index](https://mlflow.org/docs/latest/genai/governance/ai-gateway/)).
- Model allowlists Yes
A policy that constrains which models a key or team can call, keeping expensive frontier models out of the wrong hands.
Via RBAC: gateway secrets, endpoints and model definitions are first-class permissioned resources, and the USE permission is what allows "invoking a gateway endpoint", so admins decide which teams can call which endpoints. RBAC requires MLflow 3.13.0+ and authentication enabled ([RBAC](https://mlflow.org/docs/latest/self-hosting/security/role-based-access-control/)).
- Spend alerts Yes
Alerts fired as spend approaches a threshold. Alerts that only fire after the meter has rolled over are marked as such.
Budget policies with the Alert action fire once per window when the threshold is crossed ([Budget alerts and limits](https://mlflow.org/docs/latest/genai/governance/ai-gateway/budget-alerts-limits/)).
- Webhook notifications Yes
Programmatic notifications on spend events, so budget breaches can page an on-call or open a ticket.
The Alert action posts to a webhook, with a payload including `budget_policy_id` and `current_spend` ([Budget alerts and limits](https://mlflow.org/docs/latest/genai/governance/ai-gateway/budget-alerts-limits/)).
Enforcement: Enforced before each request
Catalog
Interpret these fields: LLM gateway model counts: what “500+” means
- Models available Models available
- Not published
- How many models you can call, shown as a range because several vendors publish different totals on different pages. Vendor-reported either way, so counts are not directly comparable — some count every provider variant of the same model separately.
- Model providers reachable Upstream providers
- 14–100
- How many distinct model providers or labs you can reach, shown as a range where the vendor’s own pages disagree. More providers usually means better redundancy when one has an outage.
- Works with standard OpenAI code OpenAI-compatible API
- Yes
- If yes, you can usually switch to it by changing one base URL, and switch away just as easily. This is the main defence against lock-in.
- OpenAI chat endpoint POST /v1/chat/completions
- Yes
- The endpoint almost every application ports first. "Not documented" means the vendor never states it, which is different from a documented no.
- Anthropic messages endpoint POST /v1/messages
- Yes
- Whether Anthropic-shaped calls work without rewriting them. Several products support this only as SDK compatibility or provider passthrough rather than a native endpoint — the detail page says which.
- OpenAI Responses endpoint POST /v1/responses
- Yes
- The newer stateful OpenAI surface. Support is much thinner across this market than chat completions.
- Embeddings endpoint POST /v1/embeddings
- Yes
- Whether you can generate vectors through the same gateway, or need a second integration for retrieval workloads.
- Image generation endpoint POST /v1/images/generations
- Not documented
- Whether image models are reachable through the same surface as text.
- Audio endpoints POST /v1/audio/*
- Not documented
- Speech-to-text and text-to-speech. Frequently the first gap in an otherwise complete gateway.
- Batch jobs endpoint POST /v1/batches
- Not documented
- Asynchronous bulk processing, usually at a discount. Commonly undocumented, and commonly the reason a migration stalls late.
- Needs the vendor’s own code library Requires a vendor-specific SDK
- No
- A proprietary client library spreads through your codebase and has to be torn out again if you leave. “No” is the better answer here, and it means the standard OpenAI client works.
- You can export your request history Logs / usage data export
- Yes
- Whether you can get your own request logs, traces, or usage records back out — through an API, a bulk export, or a download. Decides whether you leave with your history or abandon it.
- Settings can live in version control Declarative config-as-code
- No
- Whether routing, fallback, and budget rules can be declared in a file you keep in Git, rather than existing only as settings clicked into a hosted dashboard.
- Image generation Image generation
- Not published
- Whether image models are reachable through the same interface.
- Speech and audio Speech and audio
- Not published
- Text-to-speech or transcription models through the same interface.
- Video generation Video generation
- Not published
- Whether video models are reachable through the same interface.
- Batch processing Batch processing
- Not published
- Submitting large jobs for cheaper, slower processing. Often 50% off for work that is not time-sensitive.
Routing & reliability
Interpret these fields: How LLM gateway failover actually works · Does your LLM gateway promise any uptime? · Is routing destroying your prompt cache? · Changing models without breaking production
- Uptime it promises in writing Contractual SLA uptime
- Not published
- The uptime percentage in a published, contractual service level agreement. A public status page is not an SLA — it reports what happened, it does not promise anything or pay you back when it breaks.
- Automatic failover Automatic failover
- Yes
- When a model provider goes down or rate-limits you, traffic moves to a backup automatically instead of returning errors to your users.
- Load balancing Load balancing
- Yes
- Spreads requests across several providers or keys to raise your effective rate limit.
- Rule-based routing Conditional routing
- Not published
- Send different requests to different models based on rules — for example a cheap model for free users and a strong model for paying ones.
- Response caching Response caching
- Not published
- Reuses the answer when the exact same request comes in again, which cuts both cost and latency.
- Similar-question caching Semantic cache
- Not published
- Reuses an answer when a new question means roughly the same thing as an earlier one. Saves far more than exact-match caching but can return subtly wrong answers if tuned loosely.
- Where you set the timeout Request timeout surface
- Not documented
not_documented for gateway requests: no per-endpoint or per-request timeout key appears on the routing, endpoint or query pages. The only timeouts discussed are artifact upload/download through the tracking server's artifact proxy ([Traffic routing and fallbacks](https://mlflow.org/docs/latest/genai/governance/ai-gateway/traffic-routing-fallbacks/), [Query endpoints](https://mlflow.org/docs/latest/genai/governance/ai-gateway/endpoints/query-endpoints/), [Tracking server](https://mlflow.org/docs/latest/self-hosting/architecture/tracking-server/)).
- Where a request timeout can be set: per request, in a config file, in the vendor dashboard, or nowhere. Nine of the twenty products documented here do not describe a request timeout at all, so the worst case of a hung upstream call is unknowable from the docs.
- Where you set retries Retry policy surface
- Not documented
No retry count, backoff strategy or default is published. The documented failure behaviour is sequential fallback rather than same-model retry: "the gateway tries fallback models sequentially until one succeeds". Default retry count: n.a. Backoff: n.a. ([Traffic routing and fallbacks](https://mlflow.org/docs/latest/genai/governance/ai-gateway/traffic-routing-fallbacks/)).
- Where retry count and backoff are configured. Worth knowing alongside billing: a retried streaming call can be charged more than once.
- Where you set fallbacks Fallback surface
- Dashboard only
ORDERED and UI-configured: fallbacks are added under "Priority 2 (Fallback)" on the endpoint page, reordered by dragging, and tried in sequence when the primary hits errors or rate limits. Each fallback carries its own provider, model and API key, so a fallback can point at a different provider or region ([Traffic routing and fallbacks](https://mlflow.org/docs/latest/genai/governance/ai-gateway/traffic-routing-fallbacks/)).
- Where the fallback chain is defined. Almost every product claims fallback; the useful question is whether you can change it from code or only by hand in a dashboard.
- Shape of the fallback chain Fallback shape
- Ordered list
- An ordered list tries targets in sequence; a weighted split sends a percentage of traffic to each, which is what you need to trial a new model on 5% of requests. Weighted splits are much rarer than the marketing implies.
- Upstream health tracking Health checks / circuit breaking
- Not documented
not_documented — no active health check, circuit breaker or unhealthy-model ejection is described; failover is reactive, triggered by the error on the request itself ([Traffic routing and fallbacks](https://mlflow.org/docs/latest/genai/governance/ai-gateway/traffic-routing-fallbacks/)).
- Whether the product notices a failing upstream and stops sending traffic to it, and whether you can tune the thresholds. This is what turns a provider outage into a blip rather than a sustained error rate, and only four of the twenty expose it.
- Cross-region failover you control Multi-region failover surface
- Not documented
not_documented as a gateway feature. "Regional failover: Route to providers in different geographic regions" is listed as a fallback use case, but that is provider selection inside one gateway rather than a multi-region gateway deployment; no cross-region MLflow topology is described ([Traffic routing and fallbacks](https://mlflow.org/docs/latest/genai/governance/ai-gateway/traffic-routing-fallbacks/), [Kubernetes and Helm](https://mlflow.org/docs/latest/self-hosting/kubernetes-helm/)).
- Whether you can define what happens when a region degrades. A vendor running many regions is not the same as a vendor letting you configure failover between them; only three document a user-controlled mechanism.
- Where you set load balancing Load balancing surface
- Dashboard only
Weights are user-settable but UI-only: traffic splitting assigns each model a 1–100% weight and the weights "must sum to exactly 100", intended for A/B tests and gradual rollouts. No latency- or cost-aware balancing is documented ([Traffic routing and fallbacks](https://mlflow.org/docs/latest/genai/governance/ai-gateway/traffic-routing-fallbacks/)).
- Where traffic distribution across upstreams or keys is configured.
Operations
Interpret these fields: LLM gateway observability: traces, logs and export · Running coding agents through an LLM gateway · Changing models without breaking production
- Usage dashboards and logs Observability
- Yes
- Built-in visibility into what was sent, what came back, what it cost, and how long it took.
- Spending limits Budget controls
- Yes
- Hard caps that stop spend before it becomes a surprise invoice. The single most valuable control for a small team.
- Rate limits Rate limits
- Not published
- Caps on request volume per key or per user, useful for protecting against abuse and runaway loops.
- Separate keys per team or app Virtual keys
- Not published
- Issue scoped keys with their own budgets and permissions so you can attribute cost and revoke access without rotating everything.
- Prompt versioning Prompt management
- Yes
- Store and version prompts outside your code so they can be changed without a deploy.
- Quality testing Evals
- Yes
- Built-in tooling to score model output against test cases, so you can tell whether a model swap made things better or worse.
- MCP support MCP support
- Yes
- Native support for the Model Context Protocol, the emerging standard for connecting models to external tools.
- What gets logged Logged content
- Your choice
configurable: off unless usage tracking is enabled on the endpoint, and full request/response payloads once it is. PII can be stripped before storage with client-side span processors registered through `mlflow.tracing.configure(span_processors=[...])` ([Usage tracking](https://mlflow.org/docs/latest/genai/governance/ai-gateway/usage-tracking/), [Masking](https://mlflow.org/docs/latest/genai/tracing/observe-with-traces/masking/)).
- Whether full prompts and responses are stored, only metadata, or nothing. Full-body logging is the most useful debugging feature here and the one most likely to need a conversation with your compliance team.
- You can turn logging off Body-logging opt-out
- Yes
Yes — it is opt-in rather than opt-out: usage tracking is a toggle on each endpoint, and the docs note token and cost metrics are unavailable for some providers and models even when it is on ([Usage tracking](https://mlflow.org/docs/latest/genai/governance/ai-gateway/usage-tracking/)).
- Whether prompt and response bodies can be suppressed while still keeping usage metrics. Nineteen of the twenty document a way to do this; the mechanisms range from a per-request header to an organisation-wide setting.
- Traces you can take elsewhere Distributed tracing
- OpenTelemetry
Gateway requests become MLflow traces server-side with no client instrumentation, and the trace layer is OpenTelemetry-based: OTLP export via `OTEL_EXPORTER_OTLP_TRACES_ENDPOINT`, simultaneous dual export with `MLFLOW_TRACE_ENABLE_OTLP_DUAL_EXPORT=true`, GenAI semantic conventions via `MLFLOW_ENABLE_OTEL_GENAI_SEMCONV`, and W3C TraceContext linking client-side agent traces to gateway traces ([OpenTelemetry export](https://mlflow.org/docs/latest/genai/tracing/opentelemetry/export/), [AI Gateway tracing integration](https://mlflow.org/docs/latest/genai/tracing/integrations/listing/mlflow-ai-gateway/), [LiteLLM alternative](https://mlflow.org/litellm-alternative/)).
- Whether the product emits OpenTelemetry, a proprietary format, or nothing. OpenTelemetry means the traces land in the tooling you already run instead of only in the vendor’s dashboard.
- Where telemetry can go Export destinations
- OpenTelemetry (OTLP), Prometheus, S3, GCS, Azure Blob Storage
OTLP/HTTP for traces and metrics, with dual export so an existing collector keeps receiving spans while MLflow stores them ([OpenTelemetry export](https://mlflow.org/docs/latest/genai/tracing/opentelemetry/export/)); Prometheus metrics and a ServiceMonitor from the official Helm chart ([Kubernetes and Helm](https://mlflow.org/docs/latest/self-hosting/kubernetes-helm/)); and S3/GCS/Azure object storage as the trace-archival and artifact destination ([Tracking server](https://mlflow.org/docs/latest/self-hosting/architecture/tracking-server/)).
- Documented sinks for logs and metrics. This is a good proxy for how replaceable the vendor’s own dashboard is: around twenty destinations means you never have to depend on it, while a CSV download means you do.
- Can record user feedback Feedback capture API
- Partly
Partial and platform-level, not gateway-level: no feedback widget or scoring call is documented on the gateway pages, but MLflow advertises "LLM judge alignment with human feedback" and the MLflow MCP server can "log feedback and assessments" against traces ([LiteLLM alternative](https://mlflow.org/litellm-alternative/), [MCP server](https://mlflow.org/docs/latest/genai/mcp/), checked against [Usage tracking](https://mlflow.org/docs/latest/genai/governance/ai-gateway/usage-tracking/)).
- Whether there is an API to attach a rating or score to a logged request, which is what lets production traffic feed quality work later.
- Scores live traffic Online eval hooks
- Yes
Yes, and this is the whole pitch: "Traces captured through the gateway feed directly into `mlflow.genai.evaluate` or Evaluation Dataset APIs, so you can run judges over production traffic without any additional instrumentation", and a gateway endpoint can itself be referenced as a judge model with `gateway:/my-chat-endpoint` ([blog](https://mlflow.org/blog/mlflow-ai-gateway/), [Query endpoints](https://mlflow.org/docs/latest/genai/governance/ai-gateway/endpoints/query-endpoints/)).
- Whether automated scorers can run against real production requests, rather than only against a test set you assemble yourself.
Performance
- Delay it adds Proxy overhead
- 28.6 ms
- Extra time the product itself adds to each request, on top of however long the model takes. Usually irrelevant next to multi-second model latency, but it matters for high-volume or streaming-sensitive workloads.
- Requests per second ceiling Throughput
- 598 rps
- Published sustained request rate before the product becomes the bottleneck. Only relevant at genuinely high volume.
- What the request path runs on Architecture class
- Interpreted proxy
Python (the repo's primary language per the GitHub API) running on the FastAPI/uvicorn tracking server; the gateway is not a separate process. It is a database-backed proxy: it requires a SQL backend store (SQLite, PostgreSQL, MySQL or MSSQL) plus the FastAPI server, and file-based tracking stores are not supported. Note that the 27,777 stars on the GitHub API are for all of `mlflow/mlflow` — the whole ML/GenAI platform — and are not a measure of gateway adoption ([mlflow/mlflow API](https://api.github.com/repos/mlflow/mlflow), [Quickstart](https://mlflow.org/docs/latest/genai/governance/ai-gateway/quickstart/)).
- The comparable way to talk about latency here. An edge worker, a compiled Go or Rust binary, and a Python proxy have different overhead floors no matter which figures each vendor publishes. Products that never disclose their runtime are recorded as undisclosed rather than assumed.
- You can run the request path yourself Self-hostable data plane
- Yes
PyPI extra plus an official OCI Helm chart: `pip install 'mlflow[genai]'` then `mlflow server --port 5000`, or `helm install mlflow oci://ghcr.io/mlflow/charts/mlflow --version <version> --namespace mlflow --create-namespace` (Kubernetes 1.23+, Helm 3.8+) ([Quickstart](https://mlflow.org/docs/latest/genai/governance/ai-gateway/quickstart/), [Kubernetes and Helm](https://mlflow.org/docs/latest/self-hosting/kubernetes-helm/)).
- Whether the component that actually carries your prompts can run on your own infrastructure. Distinct from a vendor offering a self-hosted control plane while still proxying traffic through their network.
- Streaming responses Streaming support
- Yes
Supported over SSE by setting `stream: true`. Two documented caveats: post-LLM guardrails "are not triggered for streaming requests", and the `X-MLflow-Gateway-Overhead-Duration-Ms` header is only emitted for non-streaming responses ([Query endpoints](https://mlflow.org/docs/latest/genai/governance/ai-gateway/endpoints/query-endpoints/), [Guardrails](https://mlflow.org/docs/latest/genai/governance/ai-gateway/guardrails/), [Benchmarks](https://mlflow.org/docs/latest/genai/governance/ai-gateway/benchmarks/)).
- Whether token-by-token streaming is documented. The caveats matter more than the yes: some products cannot cancel a stream without still being billed, and several timeout and fallback mechanisms stop applying once the first token has been sent.
Security & compliance
Interpret these fields: LLM gateway compliance: SOC 2, HIPAA and evidence · How LLM gateway guardrails fail · Which LLM gateways store your prompts? · Do you need an MCP gateway as well?
- Does your prompt reach their servers Prompt transits vendor
- No
Self-hosted software: "Your API keys and request data stay under your control" and the gateway runs as part of your own MLflow Tracking Server, so no project-operated service sees prompts. Managed MLflow on Databricks or AWS is a third-party hosting choice, and gateway feature parity there is not documented on the pages fetched ([AI Gateway](https://mlflow.org/ai-gateway), [blog](https://mlflow.org/blog/mlflow-ai-gateway/)).
- Whether the text you send passes through this company’s own infrastructure. If it does, every other promise on this page is a policy commitment rather than a physical impossibility. Self-hosted products can answer no outright.
- What they keep if you change nothing Logging default
- Not applicable — you own the logs
No vendor logging exists; logging is to your own database, and it is opt-in per endpoint. Usage tracking is a per-endpoint toggle, and when enabled "every request is recorded as an MLflow trace" with the full request and response payload alongside latency and token counts ([Usage tracking](https://mlflow.org/docs/latest/genai/governance/ai-gateway/usage-tracking/), [blog](https://mlflow.org/blog/mlflow-ai-gateway/)).
- Defaults matter more than options. A product that stores full prompts and replies unless you find the right header will have stored them by the time you read the docs.
- How long they keep it Default content retention (days)
- Not published
You set it. Server-side trace archival is configured in YAML (`MLFLOW_TRACE_ARCHIVAL_CONFIG` or `--trace-archival-config`) with keys `retention` (example `30d`), `location` (e.g. `s3://mlflow-trace-archive`), `interval_seconds` and `long_retention_allowlist`; retention resolves global default → workspace override → experiment override, and an experiment asking for longer than the workspace policy only gets it if its ID is allowlisted. Traces can also be deleted outright with `client.delete_traces()` by timestamp or ID ([Tracking server](https://mlflow.org/docs/latest/self-hosting/architecture/tracking-server/), [Delete traces](https://mlflow.org/docs/latest/genai/tracing/observe-with-traces/delete-traces/)).
- Default retention for request content, in days. Zero means nothing is kept. Read the note: several products keep nothing as a rule but make timed exceptions for abuse review or specific models.
- Could they train on your prompts Training on customer data
- Not applicable
not_applicable by construction: the project ships software rather than a hosted service, so no project-operated system receives prompts to train on ([AI Gateway](https://mlflow.org/ai-gateway)).
- Whether the vendor may use your prompts and outputs to train models. “Not published” means we could not find any position, which is not the same as a no — ask for it in writing.
- Where it runs, and what you can pin Region and residency control
- Anywhere you run it — Docker, Kubernetes via the official Helm chart, or a cloud VM; there are no project-managed regions ([Kubernetes and Helm](https://mlflow.org/docs/latest/self-hosting/kubernetes-helm/), [Self-hosting](https://mlflow.org/docs/latest/self-hosting/)).
- Which regions are offered and whether you can force processing to stay in one. A global endpoint that silently picks a region is a different compliance story from an endpoint you pin yourself.
- Where safety filters run Guardrail execution location
- In your own infrastructure
Guardrails are configured per endpoint and evaluated by an LLM judge that is itself another gateway endpoint, so enforcement happens in your own server while the judging inference goes to whichever provider backs the judge endpoint ([Guardrails](https://mlflow.org/docs/latest/genai/governance/ai-gateway/guardrails/)).
- A filter that strips personal data only helps if it runs before the data leaves your boundary. If guardrails execute in the vendor’s cloud, the vendor has already received whatever you wanted redacted.
- Who else touches the data Subprocessor list
- Not published
- The published list of third parties the vendor passes your data to. No list means you cannot know the full chain, which most data-protection agreements require you to.
- SOC 2 audited SOC 2 audited
- Not published
- An independent audit of security controls. Enterprise buyers and their procurement teams routinely require it.
- Will sign a HIPAA agreement HIPAA BAA
- Not published
- Required before you may send protected health information through the service. Without a signed BAA, healthcare data is off limits.
- GDPR commitments GDPR commitments
- Not published
- Published data processing terms for handling personal data of people in the EU and UK.
- Can keep data in the EU EU data residency
- Not published
- Requests can be processed inside the EU rather than routed to US infrastructure. Often the deciding constraint for European customers.
- Does not retain your data Zero data retention
- Not applicable
- Prompts and responses are not stored after the request completes. Sometimes a paid add-on rather than the default.
- Strips personal data PII redaction
- Yes
- Detects and removes identifiers such as names, emails, and card numbers before the request reaches the model provider.
- Content guardrails Content guardrails
- Yes
- Policy checks on inputs and outputs — blocking unsafe content, enforcing formats, or catching prompt-injection attempts.
- Runs fully disconnected Air-gapped deployment
- Not published
- Can be deployed in a network with no internet access, which some regulated and defence environments require.
- Blocks personal data in prompts PII / DLP enforcement
- Can block the request
A PII Detection guardrail type is built in, defaults to the Pre-LLM stage, and its action is chosen at setup: Block returns HTTP 400 naming the guardrail and its rationale, or Sanitize redacts the offending content and lets the request continue ([Guardrails](https://mlflow.org/docs/latest/genai/governance/ai-gateway/guardrails/)).
- Whether personal data detection sits on the request path and can stop the call, merely inspects and forwards it, or is not documented. A control that only reports is a logging feature, not a policy control.
- Blocks prompt injection Injection / jailbreak enforcement
- Not documented
n.a. — no prompt-injection or jailbreak guardrail type is offered; the three documented types are Safety, PII Detection and Custom, so injection detection would have to be written as a Custom judge prompt ([Guardrails](https://mlflow.org/docs/latest/genai/governance/ai-gateway/guardrails/)).
- Whether injection and jailbreak detection can stop a request. Most products offering this call a partner classifier rather than shipping their own.
- Blocks harmful content Toxicity / moderation enforcement
- Can block the request
A Safety guardrail type ships built in and defaults to the Post-LLM stage, with the same Block (HTTP 400) or Sanitize actions. Guardrails apply to both unified and passthrough endpoints ([Guardrails](https://mlflow.org/docs/latest/genai/governance/ai-gateway/guardrails/)).
- Whether hate, violence, sexual and self-harm categories are checked inline and can stop a request, in either direction.
- Your own policy rules Custom policy hooks
- Can block the request
Custom guardrails are natural-language judge instructions rather than regex: you pick the stage (Pre-LLM or Post-LLM), the judge model (any other gateway endpoint) and the action (Block or Sanitize). Multiple guardrails run in table order and later ones are skipped once one blocks ([Guardrails](https://mlflow.org/docs/latest/genai/governance/ai-gateway/guardrails/)).
- Whether you can add your own rule — a regex, a webhook, or your own classifier — rather than choosing from the vendor library.
- Where guardrails run Guardrail execution location
- Either, your choice
- Whether guardrail evaluation happens inside your infrastructure or on the vendor’s servers. This decides whether the prompt you are trying to protect leaves your network in order to be checked.
- If the guardrail itself fails Guardrail failure mode
- Not documented
not_documented — the guardrails page describes Block and Sanitize outcomes but says nothing about what happens when the judge endpoint itself errors or times out ([Guardrails](https://mlflow.org/docs/latest/genai/governance/ai-gateway/guardrails/)).
- What happens when the guardrail service times out or errors: does the request proceed unchecked, or is it blocked? This is the worst-documented field in the entire catalogue — only two vendors state it plainly, which means most teams are running a control whose failure behaviour they cannot know.
- Third-party guardrail vendors Guardrail integrations
- Not published
- Named external guardrail services the product can call. A long list means the product is a router for policy engines rather than a policy engine itself — which also means another vendor bill and another hop.
Compliance evidence
Graded by how strong the evidence is, not whether the word appears on the vendor’s website. An audited report and a marketing claim are different things, and only one of them will satisfy your own auditor.
Nothing to certify. This is software you run yourself, so no vendor receives your data and compliance is inherited from whatever you deploy it on — your own certifications, not the project’s.
Fit & integration
Interpret these fields: How much does an LLM gateway lock you in? · Do you need an MCP gateway as well? · Running coding agents through an LLM gateway · Changing models without breaking production
- Work to try it Evaluation work shape
- Run something locally first
- The shape of the work on the vendor’s own quickstart, from swapping one base URL through to deploying infrastructure. An ordinal class rather than a duration, because elapsed time depends on accounts and quota we cannot see.
- Work to run it Production work shape
- Deploy it on your infrastructure
- The same scale applied to the vendor’s recommended production path. For several products this is much heavier than the quickstart, which is exactly why both are recorded.
- Steps on the quickstart Numbered quickstart steps
- 4
Four numbered steps (install and start, create an LLM Connection, create an endpoint, query it), but steps 2 and 3 are five-click UI sub-flows each rather than commands ([Quickstart](https://mlflow.org/docs/latest/genai/governance/ai-gateway/quickstart/)).
- A literal count of numbered steps on the vendor’s quickstart, recorded as evidence beside the work shape. Zero means the page publishes no numbered procedure at all. Large counts usually mean interleaved language tracks rather than more work.
- Can you self-host it today Self-host install documentation
- Install command published
`pip install 'mlflow[genai]'` then `mlflow server --port 5000`; on Kubernetes `helm install mlflow oci://ghcr.io/mlflow/charts/mlflow --version <version> --namespace mlflow --create-namespace`
- Whether an install command is actually published. Several products advertise self-hosting while publishing no command to start from, which a plain yes/no would hide.
- Works with the OpenAI SDK OpenAI SDK drop-in
- Yes
Yes: "Use any OpenAI-compatible SDK. Point the base URL at the gateway and use your endpoint name as the model", with `api_key="unused"` because the gateway holds the real credential ([AI Gateway](https://mlflow.org/genai/ai-gateway), [Quickstart](https://mlflow.org/docs/latest/genai/governance/ai-gateway/quickstart/)).
- Whether an existing OpenAI-compatible client can be pointed at it by changing the base URL and key.
- Vercel AI SDK support AI SDK provider package
- Not documented
Vercel AI SDK appears in MLflow's docs only as a tracing integration — the `ai` package exporting OpenTelemetry spans into MLflow via `@vercel/otel` and `OTEL_EXPORTER_OTLP_ENDPOINT` — not as a way to call gateway endpoints, and no gateway provider package or `baseURL` example is documented ([Vercel AI SDK tracing](https://mlflow.org/docs/latest/genai/tracing/integrations/listing/vercelai/), checked against [Query endpoints](https://mlflow.org/docs/latest/genai/governance/ai-gateway/endpoints/query-endpoints/)).
- Whether a first-party AI SDK provider package exists, or only a community package, a documented workaround, or the generic OpenAI provider pointed at a custom base URL.
- Python framework integrations Documented Python frameworks
- LangChain, LangGraph, DSPy, OpenAI Agents SDK, LiteLLM
The query page ships copy-paste snippets for LangChain (`ChatOpenAI` with the gateway base URL), LangGraph (`create_react_agent`), DSPy (`dspy.LM`), the OpenAI Agents SDK and LiteLLM (`litellm.completion`), all pointed at `/gateway/mlflow/v1`. Separately, MLflow advertises one-line `autolog()` tracing for 30+ frameworks — a platform feature, not a gateway one ([Query endpoints](https://mlflow.org/docs/latest/genai/governance/ai-gateway/endpoints/query-endpoints/), [LiteLLM alternative](https://mlflow.org/litellm-alternative/)).
- LangChain, LangGraph, LlamaIndex and similar orchestration frameworks with a documented integration.
- Callable from Cloudflare Workers Cloudflare Workers support
- Not documented
n.a. (not documented) — the gateway is a Python server; no edge-worker deployment appears on the self-hosting pages ([Self-hosting](https://mlflow.org/docs/latest/self-hosting/), [Kubernetes and Helm](https://mlflow.org/docs/latest/self-hosting/kubernetes-helm/)).
- Whether the docs show calling this product from your own Worker. Deliberately separated from the several products whose own gateway runs on Workers, which is a fact about their infrastructure and not about your edge compatibility.
- Kubernetes install Helm chart availability
- Official Helm chart
Official OCI Helm chart: `helm install mlflow oci://ghcr.io/mlflow/charts/mlflow --version <version> --namespace mlflow --create-namespace`, requiring Kubernetes 1.23+ and Helm 3.8+, with TLS, ingress, Prometheus ServiceMonitor, NetworkPolicy, RBAC and an `mlflow gc` CronJob, and PostgreSQL plus S3/GCS/Azure recommended for production ([Kubernetes and Helm](https://mlflow.org/docs/latest/self-hosting/kubernetes-helm/)).
- Whether a named, published Helm chart exists, versus Helm being referenced with no chart named, versus generic cluster documentation that has nothing to do with this product.
- Terraform support Terraform provider or modules
- Not documented
n.a. (not documented) — no Terraform provider or module is referenced on the self-hosting, Kubernetes or gateway pages fetched ([Self-hosting](https://mlflow.org/docs/latest/self-hosting/), [Kubernetes and Helm](https://mlflow.org/docs/latest/self-hosting/kubernetes-helm/)).
- Whether you can declare this in code: an official provider, official modules, resources inside a hyperscaler’s provider, a community provider, or only Terraform code shipped in a repo.
- Reuses your cloud identity Cloud IAM reuse
- Static provider credentials only
AWS Bedrock and Vertex AI are listed as supported providers, but the passthrough/authentication detail on the fetched pages is limited to API-key style LLM Connections; no IAM role, SigV4 or workload-identity configuration is documented for the gateway itself ([Model providers](https://mlflow.org/docs/latest/genai/governance/ai-gateway/endpoints/model-providers/), [API keys](https://mlflow.org/docs/latest/genai/governance/ai-gateway/api-keys/create-and-manage/)).
- Whether you can authenticate with IAM roles, workload identity or managed identities instead of another long-lived API key. Distinguished from products that only accept static upstream provider credentials.
- Fits behind your API gateway API gateway integration
- It is the API gateway
It is itself the gateway — "a centralized proxy layer that routes requests to LLM providers through a single, unified API" — but uniquely here it is a gateway embedded in an MLOps platform rather than a standalone proxy ([AI Gateway](https://mlflow.org/ai-gateway), [blog](https://mlflow.org/blog/mlflow-ai-gateway/)). It can also front another gateway: Portkey is a supported provider ([Model providers](https://mlflow.org/docs/latest/genai/governance/ai-gateway/endpoints/model-providers/)).
- Whether AI traffic can go through a gateway you already run, and who documented that — the product’s vendor or the gateway’s.
- MCP support MCP surface shape
- Hosted MCP server
Two separate things, both experimental. MLflow ships an MCP server (3.5.1+) that exposes trace management to Claude, Cursor and VS Code — search traces, analyse performance, log feedback, manage tags, delete traces — and an MCP Registry ("experimental feature introduced in MLflow 3.15.0") for registering and versioning MCP servers via `server.json` and `mlflow.genai.register_mcp_server()` ([MCP server](https://mlflow.org/docs/latest/genai/mcp/), [MCP Registry](https://mlflow.org/docs/latest/genai/mcp-registry/)). The landing page's stronger claim — that the gateway governs "which MCP servers your agents can reach" — has no matching page in the AI Gateway docs section ([AI Gateway](https://mlflow.org/ai-gateway)).
- Which kind of MCP support this is: a gateway that governs many MCP servers, a hosted MCP server you connect to, MCP tools accepted inside the completion API, or client tooling. These are different products behind one acronym.
- Needs your own provider key Upstream provider key required
- Required
Yes: nothing works until you create an LLM Connection with your own provider key (quickstart step 2), because the project hosts no models and issues no credits. Exception: coding-agent endpoints, where "the agent brings its own credentials" ([Quickstart](https://mlflow.org/docs/latest/genai/governance/ai-gateway/quickstart/), [releases index](https://mlflow.org/releases/)).
- Whether an upstream provider account and key must exist before your first call works. A prerequisite rather than a step, and it can differ between the hosted and self-hosted forms of the same product.
- Gate before models work Model access gate
- No gate
None from the project: there is no approval, quota, waitlist or enablement step — you pick a provider and model in the endpoint dialog and supply your own key. Access control is yours to impose, via RBAC on gateway endpoints ([Create and manage endpoints](https://mlflow.org/docs/latest/genai/governance/ai-gateway/endpoints/create-and-manage/), [RBAC](https://mlflow.org/docs/latest/self-hosting/security/role-based-access-control/)).
- Whether an enablement click, a quota grant, a paid tier or an approval form stands between a valid key and a working model call.
- Official client languages First-party client SDK languages
- Python
Documented client code is Python and cURL: `requests`, the OpenAI Python SDK against `/gateway/mlflow/v1`, the Anthropic Python SDK against the passthrough base URL, and `google-genai`. No JavaScript/TypeScript gateway example appears on the fetched pages, though any OpenAI-compatible client works by construction ([Quickstart](https://mlflow.org/docs/latest/genai/governance/ai-gateway/quickstart/), [Query endpoints](https://mlflow.org/docs/latest/genai/governance/ai-gateway/endpoints/query-endpoints/), [Model providers](https://mlflow.org/docs/latest/genai/governance/ai-gateway/endpoints/model-providers/)).
- Languages with a first-party client library. An empty list can still mean the product is usable from any language via an OpenAI-compatible SDK.
How pricing actually works
No licence cost; your cost is the MLflow Tracking Server plus a SQL backend store (SQLite, PostgreSQL, MySQL or MSSQL) and, for production, object storage and optionally Redis for shared budget counters ([Quickstart](https://mlflow.org/docs/latest/genai/governance/ai-gateway/quickstart/), [Tracking server](https://mlflow.org/docs/latest/self-hosting/architecture/tracking-server/), [Budget alerts and limits](https://mlflow.org/docs/latest/genai/governance/ai-gateway/budget-alerts-limits/)).
Common questions
Answered from the fields above, so these move when the catalog moves. Every figure quoted here appears in the specification with its source.
Does MLflow AI Gateway charge a markup on model prices?
MLflow AI Gateway adds no percentage markup to model prices. Adding funds is free. The cost estimator on this site itemises these mechanisms against your own volume, because which one is cheapest depends entirely on the numbers you put in.
Can MLflow AI Gateway be self-hosted?
Yes. MLflow AI Gateway can be run on your own infrastructure or used as a managed service. The licence is Apache-2.0. Running it yourself means you supply the infrastructure and the upstream model accounts, so the bill is your own hosting plus the providers' own rates.
Is MLflow AI Gateway SOC 2 audited, and will it sign a HIPAA BAA?
MLflow AI Gateway publishes neither a SOC 2 report nor a HIPAA business associate agreement. Each of these is linked to the vendor's own page in the compliance section below. Neither absence means a refusal: both are things a vendor either publishes or does not, and smaller products often hold the certification without advertising it.
Does MLflow AI Gateway retain your prompts?
Zero data retention does not apply to MLflow AI Gateway: it runs inside your own infrastructure, so prompts never reach a vendor. Whether prompt and response bodies are logged is configurable. Logging can be turned off. Retention becomes your own configuration question instead, decided by whatever logging you switch on in your own deployment.
Can you use your own provider keys with MLflow AI Gateway?
Yes. MLflow AI Gateway can route through your own accounts with the underlying model providers, so inference is billed to you directly. You store your own provider keys as LLM Connections inside your MLflow server and pay the provider directly; the project charges nothing for the gateway ([Quickstart](https://mlflow.org/docs/latest/genai/governance/ai-gateway/quickstart/)).
How many models does MLflow AI Gateway support?
MLflow AI Gateway publishes no total model count. It reaches 14–100 upstream providers. No model total is published for the gateway. The endpoint creation UI shows a searchable model selector with capability badges (Tools, Reasoning, Caching, Vision), context window and token costs, and the docs point you at "the provider dropdown when creating an endpoint" instead of a number ([Create and manage endpoints](https://mlflow.org/docs/latest/genai/governance/ai-gateway/endpoints/create-and-manage/), [Model providers](https://mlflow.org/docs/latest/genai/governance/ai-gateway/endpoints/model-providers/), 2026-09-02).
Official links
27,777 GitHub stars — a proxy for community size, not for quality.
Independent coverage
Third-party analysis, walkthroughs and operator threads. We link criticism as readily as praise. Nothing here is written or published by the vendor, by a competitor listed on this site, or by an SEO content farm — how we vet these.
Written reviews and analysis 3
- MLflow AI Gateway: LLM-Routing mit Tracing 2026 German consultancy write-up positioning the gateway as "ein datenbank-gestützter Proxy im MLflow-Tracking-Server" with native provider integrations since 3.11 — useful, but it still shows a `mlflow gateway start --config-path config.yaml` command and a YAML endpoint config, which MLflow 3.0 removed.
- MLflow vs LiteLLM: welches LLM-Gateway 2026? Third-party head-to-head that lands on the same operational caveat found in the docs: "MLflow bietet out of the box nur HTTP Basic Authentication", and neither gateway ships a ready GDPR/SSO/HA story.
- We Evaluated 13 LLM Gateways for Production. Here's What We Found Vendor-adjacent but externally published survey that files MLflow AI Gateway in "Tier 4 — Niche or Limited" with "Limited LLM-specific features" and "Heavy for simple routing"; note it predates the 2026 relaunch of the feature.
Video 1
- Centralize Your LLM Access with MLflow AI Gateway Independent walkthrough wiring SambaNova Cloud through the gateway for enterprise GenAI, covering governance and observability rather than the vendor's own demo path.
What has changed here
- catalog entry catalog entry Not published Added to the catalog source ↗
Read the head-to-head
These pairs have a written verdict, not just a table.