MLflow AI Gateway vs Braintrust Gateway
The question that decides it: Which evaluation platform are you already running, and can you accept its gateway on the terms it comes with?
Our verdict
Already running MLflow, or self-hosting by requirement: MLflow AI Gateway, because it costs nothing at any volume and it is the only one of the two that can block or sanitise a request before it reaches a provider. Already running Braintrust evals and needing SOC 2 Type II, a HIPAA BAA or an EU-hosted endpoint without building it: Braintrust Gateway, on the understanding that its post-GA price is unannounced.
Why
Both of these are a routing surface attached to something larger, and in both cases the something larger is the reason you would adopt it. MLflow's gateway is a feature of the Tracking Server rather than a deployable of its own: as of MLflow 3.0 the standalone deployment server application and the start-server CLI command were removed, so mlflow server is the only entry point. Braintrust's Gateway sits inside its eval and observability platform, with the standalone proxy MIT-licensed on GitHub and the platform itself recorded as open core. The consequence is that neither should be evaluated as a proxy in isolation. You are choosing the platform, and inheriting its gateway.
The licence and the bill diverge from there. MLflow is Apache-2.0 under Linux Foundation governance with no pricing page and no paid gateway tier at all: zero token markup, zero credit fee, zero per seat, and your only cost is the Tracking Server plus a SQL backend and, in production, object storage. Braintrust also charges $0 per seat and lists the hosted Gateway as free during public preview, but pricing is to be announced before general availability, so the number you would plan around does not exist yet. What is published is the platform around it: Starter at $0 with $10 of model credits, 1 GB of processed data, 10k scores and 14-day retention, then Pro at $249 a month, which includes $249 of credits dropping to $100 a month after 2026-09-01, with processed-data overages at $4 per GB on Starter and $3 per GB on Pro and scores at $2.50 and $1.50 per thousand.
On enforcement the two records point in opposite directions, and this is the sharpest functional split. MLflow records guardrails as present: Safety, PII and Custom checks configured per endpoint with Block or Sanitize actions, evaluated by an LLM judge that is itself another gateway endpoint, so the decision happens inside your own server while the judging inference goes to whatever provider backs the judge. Braintrust records guardrails as absent, and deliberately so — scorers, spans and online scoring run after the fact, with no request-path enforcement, and its PII handling is a global masking function that redacts before logging rather than before the provider sees the prompt. If your requirement is that something never leaves the boundary, only MLflow attempts it. The caveat is load-bearing: post-LLM guardrails are not triggered for streaming requests, so a Safety guardrail on its default stage stops enforcing the moment a client sets stream true.
Everything else is the asymmetry you would expect between a foundation project and a funded company. Braintrust publishes SOC 2 Type II, a HIPAA BAA and EU residency on a dedicated EU host in eu-west-1, and raised an $80M Series B led by ICONIQ in February 2026; MLflow's certification fields are all recorded not applicable, because there is no vendor to attest anything. Self-hosting inverts it: MLflow installs by pip or an official OCI Helm chart with TLS, ingress, Prometheus metrics, NetworkPolicy and RBAC from 3.13.0, while Braintrust self-hosting is a managed control plane plus a customer-hosted data plane and is effectively Enterprise-only. Read the headline counts carefully. MLflow's 27,777 stars are for all of mlflow/mlflow and are not a measure of gateway adoption, against 409 on Braintrust's proxy repo; MLflow enumerates 14 providers in its docs tables while other vendor pages say 100+ and 50+, and publishes no model total, against Braintrust's 18 providers and over 100 models. MLflow's 28.6 ms of P50 overhead and 598 requests per second are its own published figures from its own LiteLLM comparison, measured with a 50 ms simulated provider delay, 4 workers and 50 concurrent users; Braintrust publishes no overhead number at all.
Which one, concretely
Choose MLflow AI Gateway if
- You already run MLflow for tracing and evaluation, and want every gateway request to land as a trace with no extra instrumentation
- You need Block or Sanitize enforcement on non-streaming traffic before the prompt reaches a provider
- You need on-prem or Kubernetes self-hosting as a first-class path, via pip or the official OCI Helm chart
- You want no fee of any kind: Apache-2.0, no token markup, no credit fee, no seats, no vendor
Choose Braintrust Gateway if
- You already run evals, scorers and datasets in Braintrust and want model traffic on the same platform
- You need SOC 2 Type II, a HIPAA BAA or an EU-hosted gateway endpoint without assembling it yourself
- You want exact-match response caching, which MLflow does not offer in any form
- You want failover visible per request through the x-bt-failover-from and x-bt-failover-to response headers
What catches people out
- MLflow's post-LLM guardrails are not triggered for streaming requests. If clients can set stream true, move the guardrail off its default stage or accept that it stops enforcing.
- MLflow's budget enforcement can lag: the default refresh interval is 600 seconds and the default local tracker keeps counters per process. Configure Redis if you need shared counters across workers.
- Braintrust's hosted Gateway is free only during public preview, with pricing to be announced before general availability. Do not build a cost model on the preview price.
- MLflow's own comparison table ticks Rate Limiting, but no rate-limit mechanism is documented anywhere in the gateway docs — only dollar budgets. Plan for budgets, not request ceilings.
Side by side
Interpret these fields: How much does an LLM gateway lock you in? · LLM gateway compliance: SOC 2, HIPAA and evidence · How LLM gateway failover actually works
5 of 18 fields differ, marked with a dot. Every figure links to the vendor page it came from. Blank values read Not published rather than No — silence from a vendor is not a negative answer.
| Field | MLflow AI Gateway | Braintrust Gateway |
|---|---|---|
| Ease of leaving Derived score, higher is easier | 84/100 Some work to leave | 84/84 Some work to leave |
| What kind of product Category | Open source | Managed gateway |
| Who runs it Deployment model | Managed or self-host | Managed or self-host |
| Licence Licence | Apache-2.0 | MIT |
| Models available Models available | Not published | ~100 |
| Model providers reachable Upstream providers | 14–100 | 9 |
| Markup on model prices Token markup | None | Not published |
| Fee to add funds Credit purchase fee | None | Not published |
| Monthly cost per person Seat fee | None | None |
| GitHub stars GitHub stars | 27,777 | 409 |
| Quality testing Evals | Yes | Yes |
| Content guardrails Content guardrails | Yes | No |
| Response caching Response caching | Not published | Yes |
| SOC 2 audited SOC 2 audited | Not published | Yes |
| Will sign a HIPAA agreement HIPAA BAA | Not published | Yes |
| Can keep data in the EU EU data residency | Not published | Yes |
| How long they keep it Default content retention (days) | Not published | 14 days |
| Delay it adds Proxy overhead | 28.6 ms | Not published |
| Settings can live in version control Declarative config-as-code | No | Not published |
for MLflow AI Gateway and for Braintrust Gateway. Want more fields, or a third option in the mix? Open these two in the full comparison tool.
Common questions
Is MLflow AI Gateway deprecated?
No, though the confusion is understandable. The gateway was deprecated and renamed MLflow Deployments Server in 2.9.x, and 2.17.0 on 2024-10-11 explicitly reversed both the deprecation and the naming. Separately and genuinely, MLflow 3.0 removed the standalone deployment server application and the start-server CLI, so the gateway now runs inside mlflow server. The latest release in the record is dated 2026-08-26.
Which one can stop a bad request before it reaches the model?
Only MLflow, and only outside streaming. MLflow records LLM-judge guardrails for Safety, PII and Custom checks with Block or Sanitize actions, configured per endpoint, but post-LLM guardrails are not triggered for streaming requests. Braintrust records no request-path enforcement at all: its scorers evaluate after the fact, and its masking functions redact PII before logging rather than before the provider call.
Which is cheaper?
MLflow, unambiguously, at any volume: Apache-2.0 with no token markup, no credit fee and no seat fee, so you pay only for the server, a SQL backend and object storage. Braintrust charges $0 per seat and the hosted Gateway is free during public preview, but pricing after general availability is unannounced, and the platform around it runs from $0 Starter to $249 a month Pro plus per-GB and per-score overages.