MLflow AI Gateway vs Braintrust Gateway

The question that decides it: Which evaluation platform are you already running, and can you accept its gateway on the terms it comes with?

Our verdict

Already running MLflow, or self-hosting by requirement: MLflow AI Gateway, because it costs nothing at any volume and it is the only one of the two that can block or sanitise a request before it reaches a provider. Already running Braintrust evals and needing SOC 2 Type II, a HIPAA BAA or an EU-hosted endpoint without building it: Braintrust Gateway, on the understanding that its post-GA price is unannounced.

Why

Both of these are a routing surface attached to something larger, and in both cases the something larger is the reason you would adopt it. MLflow's gateway is a feature of the Tracking Server rather than a deployable of its own: as of MLflow 3.0 the standalone deployment server application and the start-server CLI command were removed, so mlflow server is the only entry point. Braintrust's Gateway sits inside its eval and observability platform, with the standalone proxy MIT-licensed on GitHub and the platform itself recorded as open core. The consequence is that neither should be evaluated as a proxy in isolation. You are choosing the platform, and inheriting its gateway.

The licence and the bill diverge from there. MLflow is Apache-2.0 under Linux Foundation governance with no pricing page and no paid gateway tier at all: zero token markup, zero credit fee, zero per seat, and your only cost is the Tracking Server plus a SQL backend and, in production, object storage. Braintrust also charges $0 per seat and lists the hosted Gateway as free during public preview, but pricing is to be announced before general availability, so the number you would plan around does not exist yet. What is published is the platform around it: Starter at $0 with $10 of model credits, 1 GB of processed data, 10k scores and 14-day retention, then Pro at $249 a month, which includes $249 of credits dropping to $100 a month after 2026-09-01, with processed-data overages at $4 per GB on Starter and $3 per GB on Pro and scores at $2.50 and $1.50 per thousand.

On enforcement the two records point in opposite directions, and this is the sharpest functional split. MLflow records guardrails as present: Safety, PII and Custom checks configured per endpoint with Block or Sanitize actions, evaluated by an LLM judge that is itself another gateway endpoint, so the decision happens inside your own server while the judging inference goes to whatever provider backs the judge. Braintrust records guardrails as absent, and deliberately so — scorers, spans and online scoring run after the fact, with no request-path enforcement, and its PII handling is a global masking function that redacts before logging rather than before the provider sees the prompt. If your requirement is that something never leaves the boundary, only MLflow attempts it. The caveat is load-bearing: post-LLM guardrails are not triggered for streaming requests, so a Safety guardrail on its default stage stops enforcing the moment a client sets stream true.

Everything else is the asymmetry you would expect between a foundation project and a funded company. Braintrust publishes SOC 2 Type II, a HIPAA BAA and EU residency on a dedicated EU host in eu-west-1, and raised an $80M Series B led by ICONIQ in February 2026; MLflow's certification fields are all recorded not applicable, because there is no vendor to attest anything. Self-hosting inverts it: MLflow installs by pip or an official OCI Helm chart with TLS, ingress, Prometheus metrics, NetworkPolicy and RBAC from 3.13.0, while Braintrust self-hosting is a managed control plane plus a customer-hosted data plane and is effectively Enterprise-only. Read the headline counts carefully. MLflow's 27,777 stars are for all of mlflow/mlflow and are not a measure of gateway adoption, against 409 on Braintrust's proxy repo; MLflow enumerates 14 providers in its docs tables while other vendor pages say 100+ and 50+, and publishes no model total, against Braintrust's 18 providers and over 100 models. MLflow's 28.6 ms of P50 overhead and 598 requests per second are its own published figures from its own LiteLLM comparison, measured with a 50 ms simulated provider delay, 4 workers and 50 concurrent users; Braintrust publishes no overhead number at all.

Which one, concretely

Choose MLflow AI Gateway if

  • You already run MLflow for tracing and evaluation, and want every gateway request to land as a trace with no extra instrumentation
  • You need Block or Sanitize enforcement on non-streaming traffic before the prompt reaches a provider
  • You need on-prem or Kubernetes self-hosting as a first-class path, via pip or the official OCI Helm chart
  • You want no fee of any kind: Apache-2.0, no token markup, no credit fee, no seats, no vendor

Choose Braintrust Gateway if

  • You already run evals, scorers and datasets in Braintrust and want model traffic on the same platform
  • You need SOC 2 Type II, a HIPAA BAA or an EU-hosted gateway endpoint without assembling it yourself
  • You want exact-match response caching, which MLflow does not offer in any form
  • You want failover visible per request through the x-bt-failover-from and x-bt-failover-to response headers

What catches people out

Side by side

5 of 18 fields differ, marked with a dot. Every figure links to the vendor page it came from. Blank values read Not published rather than No — silence from a vendor is not a negative answer.

Field MLflow AI Gateway Braintrust Gateway
Ease of leaving Derived score, higher is easier 84/100 Some work to leave 84/84 Some work to leave
What kind of product Category Open source Managed gateway
Who runs it Deployment model Managed or self-host Managed or self-host
Licence Licence Apache-2.0 MIT
Models available Models available Not published ~100
Model providers reachable Upstream providers 14–100 9
Markup on model prices Token markup None Not published
Fee to add funds Credit purchase fee None Not published
Monthly cost per person Seat fee None None
GitHub stars GitHub stars 27,777 409
Quality testing Evals Yes Yes
Content guardrails Content guardrails Yes No
Response caching Response caching Not published Yes
SOC 2 audited SOC 2 audited Not published Yes
Will sign a HIPAA agreement HIPAA BAA Not published Yes
Can keep data in the EU EU data residency Not published Yes
How long they keep it Default content retention (days) Not published 14 days
Delay it adds Proxy overhead 28.6 ms Not published
Settings can live in version control Declarative config-as-code No Not published

for MLflow AI Gateway and for Braintrust Gateway. Want more fields, or a third option in the mix? Open these two in the full comparison tool.

Common questions

Is MLflow AI Gateway deprecated?

No, though the confusion is understandable. The gateway was deprecated and renamed MLflow Deployments Server in 2.9.x, and 2.17.0 on 2024-10-11 explicitly reversed both the deprecation and the naming. Separately and genuinely, MLflow 3.0 removed the standalone deployment server application and the start-server CLI, so the gateway now runs inside mlflow server. The latest release in the record is dated 2026-08-26.

Which one can stop a bad request before it reaches the model?

Only MLflow, and only outside streaming. MLflow records LLM-judge guardrails for Safety, PII and Custom checks with Block or Sanitize actions, configured per endpoint, but post-LLM guardrails are not triggered for streaming requests. Braintrust records no request-path enforcement at all: its scorers evaluate after the fact, and its masking functions redact PII before logging rather than before the provider call.

Which is cheaper?

MLflow, unambiguously, at any volume: Apache-2.0 with no token markup, no credit fee and no seat fee, so you pay only for the server, a SQL backend and object storage. Braintrust charges $0 per seat and the hosted Gateway is free during public preview, but pricing after general availability is unannounced, and the platform around it runs from $0 Starter to $249 a month Pro plus per-GB and per-score overages.