Managed only Proprietary

Merge Gateway Managed gateway

Merge Gateway is a managed LLM gateway: an OpenAI-compatible API in front of 3 models from 4–24 providers. It charges a 5% token markup. It cannot be self-hosted. Zero data retention is published; a HIPAA BAA is not. You can point it at your own provider accounts. Beyond chat it also serves embeddings, image generation and audio. It handles failover, guardrails and request logging.

· 130 dated entries · 130 source references

Built by Merge API, Inc., founded 2020 · US company

Access at a glance

Whether this product can work for you at all, before features matter: who pays the model bill, where it can run, and which of your existing API calls keep working. Every value is the vendor’s own claim, linked to the page it came from.

A hosted routing, cost-governance and security proxy in front of other vendors' models: "The control plane for production AI. Access all LLMs through a single API, with intelligent routing, cost management, and security built-in" (Merge Gateway product page). Its own terms call it "Merge's proprietary large language model aggregator" (Merge Gateway Terms). It is a distinct SKU from Merge Unified (HRIS/CRM/ATS integrations) and Merge Agent Handler, launched on 31 March 2026 (Introducing Merge Gateway, TestingCatalog, marked SPONSORED).

Who pays the model bill

Your keys or their credits

You can start on their credits and move to your own provider accounts later.

Merge-managed credentials are the default and are billed at LLM cost plus 5%, while Pro and Enterprise organisations can attach their own provider credentials per vendor, with an optional fallback to Merge's keys and a BYOK_ONLY strict mode for embedded customers that never falls back (BYOK overview, Embedded Routing Stack, Gateway pricing).

Merchant of record: Split by mode. On managed credentials Merge is the merchant of record: you prepay Service Credits, Merge bills LLM cost plus 5% and "consolidate[s] all provider invoices into one" (Gateway pricing, Merge Gateway Terms). On BYOK the provider bills you directly — "BYOK traffic goes to the vendor under your own account and agreement" (Zero data retention).

Key handling: Gateway API keys are minted in the dashboard or through the Management API, which uses a separate mgmt_ key that cannot call model endpoints; keys carry optional spend limits with daily, weekly or monthly resets, and can be scoped per project or per customer (API keys, API overview). Upstream BYOK secrets are encrypted at rest, limited to one per provider per organisation, and gated behind a manage-credentials permission; Bedrock accepts either a bearer token or IAM access keys (BYOK overview, Amazon Bedrock BYOK).

Where it can run

1 of 5 shapes documented
  • Vendor-hosted
  • Self-host
  • Your VPC
  • On-premise
  • Air-gapped

Dimmed shapes are not documented by the vendor, which is not the same as unsupported.

hosted SaaS only, at https://api-gateway.merge.dev/v1 (Get started). Self-host, hybrid VPC, on-prem and air-gapped: n.a. as documented modes. The sole self-hosting statement anywhere is one Enterprise sales bullet, "Deploy to your own VPC or on-prem environment", with no accompanying architecture, artifact or requirements page (Gateway pricing); a third-party review reaches the same conclusion, "Fully hosted cloud gateway (no self-hosting option documented)" (Synthszr).

Nothing to install and nothing to operate: you swap a base URL and send a request (Get started). Searching the full Gateway documentation sitemap (403 URLs, fetched 2026-09-02) returns no page for Docker, Helm, Kubernetes, Terraform or a self-hosted data plane, so the Enterprise VPC/on-prem bullet is an unelaborated sales claim rather than a documented deployment mode (Gateway pricing).

API surfaces your code can keep using

4 of 7 documented, 2 partial
  • OpenAI chat POST /v1/chat/completions Yes

    POST /chat/completions on the OpenAI wire format, with /v1/openai as the drop-in base URL for the OpenAI SDK and unprefixed model names accepted (API overview, Get started).

  • Anthropic messages POST /v1/messages Yes

    POST /messages plus POST /messages/count_tokens, with /v1/anthropic as the Anthropic SDK base URL. The Anthropic surface omits the cost field on stream terminal frames and sends no ping events (API overview, Streaming).

  • OpenAI Responses POST /v1/responses Partly *

    Partial, and easy to misread: there is a native POST /responses, but it is Gateway's own shape (snapshot response.stream events terminated by response.done), explicitly not OpenAI's Responses API. The docs tell you to use the /v1/openai base URL when you want OpenAI-shaped Responses calls (API overview, Get started).

  • Embeddings POST /v1/embeddings Yes

    POST /embeddings is documented on the API surface (API overview).

  • Images POST /v1/images/generations Yes

    POST /v1/images/generations, with openai/gpt-image-2, xai/grok-imagine-image and Gemini hybrid image models documented (Image generation, API overview).

  • Audio POST /v1/audio/* Partly *

    Partial: text-to-speech only, POST /v1/audio/speech with openai/tts-1, tts-1-hd and MiniMax voices, billed per character. No speech-to-text or realtime endpoint appears in the docs sitemap (Text to speech, API overview).

  • Batch jobs POST /v1/batches Not documented

    n.a. No batch or async bulk-inference endpoint exists on the fetched API surface, and the string "batch" does not appear in any of the 403 Gateway documentation URLs (API overview, Rate limits). The one asynchronous job type is video generation (Video generation).

An asterisk marks a qualified verdict: support that is indirect (SDK compatibility or provider passthrough rather than a native endpoint), or a gap that is narrower or wider than the label suggests. Read the note before porting.

Base URL https://api-gateway.merge.dev/v1 with Authorization: Bearer. Any SDK that accepts a custom base URL works, and Merge publishes per-SDK shims: /v1/openai, /v1/anthropic, /v1/ai-sdk and /v1/openai for LangChain (Get started). Two documented traps: client.responses.* against the bare /v1 base URL reaches Gateway's own Responses API rather than OpenAI's, and the Anthropic SDK's models.list() against the bare host returns Gateway's native model listing. Note the product page instead tells you to "Change the base URL to gateway.merge.dev/v1", which is the dashboard host, not the API (Merge Gateway product page).

How much it reaches

Models 3
Upstream providers 4–24

Models: Counted from the vendor’s own published list; no aggregate total is published.

Count of entries returned by the models API on 2026-09-26. Includes every entry exposed by that endpoint; not a count of unique base models.

Merge publishes no provider total, and its docs use two different senses of the word. The catalogue's model IDs carry 24 model-family prefixes (ai21, alibaba, amazon, anthropic, arcee-ai, bland, bytedance, cohere, deepseek, google, meta, minimax, mistral, moonshot, morph, nvidia, openai, qwen, sakana, thinkingmachines, writer, xai, xiaomimimo, zai), while the separate "vendor" column is the execution host and is queried through GET /v1/vendors (Model catalog, Vendors list). Only four hosts are named in the quickstart (Get started); the zero-data-retention page names 18 (Zero data retention).

Whose models: All third-party routed. Merge hosts no models; the Gateway Terms define the service as "Merge's proprietary large language model aggregator, which allows users to route, monitor, manage, and optimize API requests to a variety of third-party artificial intelligence model providers" (Merge Gateway Terms), and customers must comply with each AI Provider's own terms (Provider terms).

Your own endpoints: n.a. No mechanism for registering a custom, self-hosted or OpenAI-compatible third-party endpoint (vLLM, Ollama, SageMaker or a private base URL) appears on the pages fetched: BYOK covers named vendors only (BYOK overview, Amazon Bedrock BYOK), and the routable set is the Merge-curated catalogue (Model catalog, Vendors list).

How it behaves in production

What happens when an upstream model is slow, wrong, or down — and what you can see and stop while it happens. Reliability features are recorded as where you configure them, not whether the vendor lists them, because almost every product here lists all of them.

1 of 6 reachable from code 3 of 4 can block no documented export

When something goes wrong

Each row says where the knob is, not whether the feature is on the marketing page. A control you can only reach by hand in someone else’s dashboard cannot be reviewed or version-controlled.

  • Request timeout Not documented

    The vendor does not document this, so any behaviour you observe today is unversioned and may change.

    As a user-settable value: no timeout header, SDK parameter or config key appears on the fetched pages. What is published is a fixed stall watchdog — a first chunk taking more than 120 seconds, or a 120-second gap between chunks, abandons the upstream and yields a 408/502 or a failover — plus the warning that a client disconnect is still billed (Streaming, API overview).

  • Retries Not documented

    No retry counter, backoff strategy or retryable-status list is published; the documented failure behaviour is failover to the next model in the policy, not a retry against the same one. Provider 429s, 5xx and timeouts trigger failover, while client errors 400-404 do not. Default retry count: n.a. Backoff: n.a. (Deterministic strategies, Errors).

  • Fallback to another model Per request

    Your application decides per call, so one noisy endpoint can have its own timeout without a redeploy.

    ORDERED. A Priority strategy holds an ordered list of models and walks it on provider failure (strategy: "PRIORITY" over the API); the policy is created in the dashboard or via POST /v1/routing-policies and then selected per request with routing_policy_id, with precedence request > customer default > project > org default > raw model (Deterministic strategies, Using policies). Mid-stream failover is visible to the client as a {"fallback_restart": true} frame (Streaming).

  • Load balancing Dashboard only

    Only reachable by hand in the vendor UI, so it cannot be reviewed, version-controlled, or changed from code.

    Weights are not user-settable for traffic splitting. Least Latency and Lowest Cost are dashboard-only strategies with no exposed weights, and the API accepts only PRIORITY and INTELLIGENT. The one weighted surface is Build Your Own Router, where you weight *benchmarks* (summing to 1.0) to score models, not traffic shares — and it too is dashboard-only (Deterministic strategies, Build your own router).

  • Upstream health tracking Fixed, cannot change

    The behaviour is fixed by the vendor. Predictable, but you cannot tune it for your workload.

    The feature exists but exposes no knobs: Gateway tracks provider health and automatically skips routes that are failing, and a tripped circuit breaker surfaces as HTTP 502 model_vendor_unhealthy. No probe interval, failure threshold, ejection window or half-open policy is published (Deterministic strategies, Errors).

  • Cross-region failover Not documented

    As failover. Geo-location routing is a compliance restriction that narrows the vendor set by country and region; it is not described as a cross-region failover mechanism, and Merge publishes nothing about its own regional topology for Gateway (Geo-location routing, How it works).

Fallback chain: Ordered list — Try A, then B, then C. Simple and predictable, but every failover is all-or-nothing.

The reliability surface is deliberately narrow: failover order is the one thing you control, and it is configured as a policy object rather than in code. Timeouts, retries and health-check thresholds are all Merge's to set, which is the trade you accept for a managed control plane. Two operational details worth knowing before production: managed-credential traffic is capped at 100 requests and 100,000 tokens per minute per organisation per provider (BYOK traffic skips these caps), and Retry-After is only returned on gateway-generated 429s, not provider ones (Rate limits, Errors).

How fast the hop is

Undisclosed vendor service

The vendor does not disclose what the request path runs on, so no overhead floor can be inferred at all.

vendor_saas: closed multi-tenant service reached at https://api-gateway.merge.dev/v1, with no published source, runtime or language for the data plane. The only public Merge Gateway repositories are a Claude Code skills pack and the Vercel AI SDK provider (TypeScript, MIT), neither of which is the gateway itself (merge-gateway-ai-sdk-provider, merge-gateway-skills).

You can run the request path yourself Not documented
Streaming Yes

No Docker image, Helm chart, binary, npm package or install command for a self-hosted data plane appears on any fetched page, and no such page exists in the 403-URL Gateway docs sitemap (Get started, Gateway pricing). Enterprise "Deploy to your own VPC or on-prem environment" is a pricing-page bullet with no artifact behind it.

Streaming caveats: Supported on every surface with stream: true, and unusually well documented. Usage and a cost figure arrive on the terminal frame (omitted on the Anthropic and AI SDK surfaces); a mid-stream failover emits {"fallback_restart": true} so clients must be prepared to discard partial output; the 120-second stall watchdogs apply to both first byte and inter-chunk gaps; and a client disconnect is still billed (Streaming).

Published figures, grouped by what each one measured. Figures in different groups are different quantities and cannot be compared with one another — nor, in most cases, with another vendor’s figure in the same group.

Overhead added by the gateway

The only figures that speak to whether this product slows your application down. Still not comparable between vendors: each was measured on different hardware, at a different load, with a different payload.

  • 90-650 ms median Vendor-published

    Vendor blog post measuring ten "lane combinations"; described as "subsecond median TTFT". No percentile beyond median, no RPS, no payload size, no hardware, no cache state.

    Source

Both latency figures are vendor self-published with no released methodology or raw data, and no independent benchmark of Merge Gateway was found. The same blog post also carries a cost comparison - "total model spend was $8.17 for the fixed-Opus lane against $2.87 for the router, which represents a 65% reduction in cost" - which is a vendor claim about its own routing, not a third-party test (The cost of a gateway). The one customer figure, Windmill's "more than $10,000 per month" saving, is a vendor-published case study (Windmill case study).

What it will stop

3 of 4 can block

Two separate questions per control: can it stop a request at all, and what does it do before you change any settings? A control that inspects and forwards is a logging feature, however it is named.

  • Personal data in prompts Can block the request

    Out of the box: You pick the action when configuring

    A Presidio-based DLP sidecar scans both prompts and completions against Global, USA and Custom category sets (15 seeded entity types) with a per-rule action of log, redact or block; blocking returns HTTP 400 (the aggregate page states 422 blocked_by_dlp_policy) and redaction means the redacted text is what gets logged. Vendor-stated cost: "a few milliseconds" (Data loss prevention).

  • Prompt injection and jailbreaks Can block the request

    Out of the box: You pick the action when configuring

    A fine-tuned DeBERTa v3 classifier with modes off, alert and block, calibrated to a 1% false-positive rate: pi_block_threshold 0.57, pi_pass_threshold 0.30, pi_input_action block and pi_output_action redact (from observe, redact, route, block, escalate), plus tiered indirect-injection handling and an allowlist of regexes (50 patterns of 200 characters at org level, up to 20 more per project, unioned). Blocks return 400/422 pi_blocked and every log carries a pi_score (Prompt injection protection).

  • Harmful content Not documented

    n.a. No toxicity, moderation or content-category filter is offered; the security suite is DLP, prompt injection, ZDR, geo routing and per-customer blocklists, and content refusals come from the provider's own policy as content_policy_violation with source: "provider" (Data loss prevention, Prompt injection protection, Errors).

  • Your own policies Can block the request

    Out of the box: You pick the action when configuring

    Custom DLP categories accept your own Python-flavoured regexes (500 characters) and keyword lists (50 entries) with the same log/redact/block actions, and a built-in tester accepts up to 50,000 characters of sample text (Data loss prevention). A webhook or arbitrary-code guardrail hook is not offered; custom routing classifiers are a separate, dashboard-only centroid scorer requiring at least three simple and three complex examples (Custom classifiers).

Where checks run On the vendor's servers
If the guardrail itself fails You choose

Read this one carefully before trusting the controls: "By default, DLP fails open and the request proceeds", and prompt-injection checks also fail open unless you set pi_fail_closed: true. So a freshly configured blocking policy will let traffic through if the classifier or sidecar errors, until you change that (Data loss prevention, Prompt injection protection). ZDR is the exception and fails closed with HTTP 400 (Zero data retention).

Calls out to: Microsoft Presidio. Each is a separate vendor relationship and a separate hop on the request path.

The most important default here is the failure mode, not the thresholds: DLP fails open by default and prompt-injection protection fails open unless pi_fail_closed: true is set, so both need explicit hardening before they can be treated as controls rather than telemetry (Data loss prevention, Prompt injection protection). The compensating strength is scope: policies can be overridden per project through PUT /v1/projects/{id}/pi-settings and /dlp-settings (sparse, with DELETE restoring inheritance), and every change lands in an append-only audit trail of roughly 60 event types (Projects guardrails API, Audit trail).

What you can see

No documented export
What gets logged Your choice

You decide whether bodies are captured, by setting or by header.

You can turn bodies off Yes

Yes, and it is the default state rather than an opt-out you have to request: payload logging is an account setting that starts disabled (Merge Gateway Terms). A per-request suppression header is not documented (API overview).

Traces Vendor format only

Proprietary and header-based: pass X-Merge-Trace-Id, X-Merge-Parent-Span-Id, X-Merge-Span-Name and X-Merge-Thread-Id to stitch multi-step agent runs, with cost rolled up from the billing pipeline and Fusion runs generating child spans automatically. Traces are retained about 24 hours. No OpenTelemetry or OTLP export is offered, and observability/tracing is the only observability page in the 403-URL docs sitemap (Tracing).

Configurable, and off by default: metadata only unless payload logging is enabled in settings, at which point inputs and outputs are stored for all requests and visible in the Logs detail view (Merge Gateway Terms). Where DLP redaction applies, the redacted text is what gets logged (Data loss prevention).

Where telemetry can go

No documented export. Whatever this product records stays in its own interface, so it cannot become part of the monitoring you already run.

n.a. No OTLP endpoint, log-drain, S3/warehouse sink or named third-party destination (Datadog, Grafana, LangSmith and so on) appears on any fetched page; the docs sitemap contains no OpenTelemetry page at all (Tracing, API overview). The nearest thing is that usage.cost is returned inline on every response, which third-party tools such as Langfuse pick up automatically (Cost governance and savings).

Records user feedback No
Scores live traffic Partly

n.a. — recorded as no because no score, rating or feedback ingestion endpoint appears on the fetched pages (API overview, Tracing); the trace surface is write-time headers only, with no companion endpoint for attaching outcomes after the fact.

Partial and routing-bound: Build Your Own Router can "Run evals" against a judge model and rubric, will auto-evaluate newly added models, and can source scores from published benchmarks, your own eval runs or uploaded results — and EVAL_* events appear in the audit trail. There is no eval API, dataset object or CI hook, and evals exist to weight routing rather than to test your application (Build your own router, Audit trail).

Depends on the vendor’s SaaS: Entirely. There is no self-hosted mode, so logs, traces, cost breakdowns and the audit trail live only in Merge's control plane, and there is no export path to bring them into your own stack (Tracing, Cost governance and savings).

Retention: n.a. as a numeric window. Neither the tracing page nor the pricing page states a log-retention period or a per-tier retention difference; the only published figure is roughly 24 hours for traces (Tracing, Gateway pricing).

Whether it fits how you work

How much work stands between you and a first call, how different that is from running it in production, and whether it slots into the stack you already have. Recorded as the shape of the work rather than a number of minutes — how long it takes you depends on which accounts and quota you already hold, which no comparison can know.

to try: base url swap to run: base url swap fits 2 of 10 common stacks

Getting to a first call

2 numbered steps
Shape of the work Change one base URL

Your existing OpenAI-compatible client keeps working. You change a base URL and a key, and nothing else in your code moves.

Read off: the vendor’s own quickstart — 2 numbered steps.

Two steps on the SDK path ("Install the SDK", "Send a request"); the page then repeats the same two-step shape for the OpenAI SDK, the Vercel AI SDK provider and a URL shim, so there is no single canonical numbered list.

Before step one

  • Your own provider key Optional

    You can start on the product’s own credits and move to your own provider keys later.

    . Managed credentials are the default path and the quickstart needs no provider key at all; BYOK is an upgrade for Pro and Enterprise organisations, with a fallback-to-Merge toggle and a strict BYOK_ONLY mode for embedded tenants (BYOK overview, Get started).

  • Payment method Card needed before models work

    The pricing page contradicts itself on the same card: the Free plan is badged "Credit card required" while its feature list includes the bullet "Start without a credit card" (Gateway pricing). Recorded as required because the badge is the explicit payment statement, but treat this as unresolved.

  • Gate before models answer No gate

    Every catalogue model is callable as soon as you have a key.

    The Free plan promises "Access all major LLM models" and the quickstart routes to any catalogue model by ID with no enablement, approval, waitlist or quota step (Gateway pricing, Get started). Gating is something you impose, not something Merge imposes: ZDR, geo rules and per-customer blocklists all narrow the reachable set (Per-customer restrictions).

Everything you need first: A Merge Gateway account and an API key from the dashboard (or minted through the Management API). No cloud account, cluster or upstream provider key is needed, because managed credentials are the default. Payment is the ambiguity: the Free plan carries the badge "Credit card required" while its own bullet list says "Start without a credit card" (Get started, Gateway pricing).

Running it in production

The same scale applied to the path the vendor recommends for production traffic. Kept separate from the quickstart because for several products here the two are barely related pieces of work.

Shape of the work Change one base URL

Your existing OpenAI-compatible client keeps working. You change a base URL and a key, and nothing else in your code moves.

What production needs: A funded account (prepaid credits, or a card for Pro's LLM-cost-plus-5%), and for anything beyond a prototype: a routing policy, a project per workload with a budget, and DLP/prompt-injection settings hardened away from their fail-open defaults. BYOK and SSO/SAML require Pro; VPC or on-prem requires Enterprise (Gateway pricing, Data loss prevention, Using policies).

Can you run it yourself

Install command published No self-hosting

This runs on the vendor’s infrastructure only.

How it fits your stack

2 of 10

Each row is a thing you might already run. “With a caveat” means it works but not the way the vendor’s marketing implies — read the reason, because that is usually where the surprise lives.

  • Fits The OpenAI SDK Drop-in: change the base URL and key, nothing else.
  • Fits The Vercel AI SDK merge-gateway-ai-sdk-provider
  • No Cloudflare Workers No Workers guidance published.
  • No Kubernetes No Kubernetes deployment published.
  • No Terraform or OpenTofu Nothing published for Terraform.
  • No An existing API gateway Nothing published about running behind your gateway.
  • No Cloud IAM I already run Static upstream credentials only. Your calls to it still use its own key.
  • With a caveat LangChain or LlamaIndex LangChain only.
  • With a caveat MCP servers to govern undefined — governs nothing on your side.
  • With a caveat Nothing — plain Node or Python Change one base URL, but a payment method is needed first.

Reading this the other way round — pick what you already run and see every product scored against it.

The integration surfaces behind those answers

  • Vercel AI SDK Official provider package

    Install the package, swap the model factory, done. Maintained by a party with a stake in it.

    Named: merge-gateway-ai-sdk-provider

    Official first-party provider package merge-gateway-ai-sdk-provider with createMergeGateway and typed Gateway options under providerOptions.mergeGateway; generateObject/streamObject map to Gateway's native json_schema structured output. AI SDK v6 uses the root import, v5 the /v5 subpath, and v4 is not supported (use the /v1/ai-sdk URL shim). The package source is public and MIT-licensed (Get started, merge-gateway-ai-sdk-provider).

  • Cloudflare Workers Not documented

    No Workers guidance either way. If you are edge-first, verify fetch-only compatibility yourself.

    n.a. Cloudflare Workers are neither a documented deployment target nor a documented runtime for Gateway; no edge or Workers page exists in the docs sitemap (Get started, Coding agents).

  • Kubernetes Not documented

    No Kubernetes story published.

    n.a. There is nothing to run in a cluster: no Helm chart, manifest, operator or container image is published, and "helm" and "kubernetes" return no hits across the 403-URL Gateway docs sitemap (Get started, Gateway pricing).

  • Terraform Not documented

    No Terraform surface published. Configuration is API or dashboard work.

    n.a. No Terraform provider or module is published and "terraform" returns no hits in the docs sitemap. Infrastructure-as-code, if you want it, means driving the REST Management API yourself — keys, projects, routing policies and guardrail settings are all API-addressable (Projects guardrails API, API overview).

  • Existing API gateway Not documented

    Nothing published about sitting behind an existing gateway. Treat it as a separate hop you route to yourself.

    n.a. (not documented). Gateway is a standalone hosted endpoint; no integration with an API-gateway platform or service mesh (Kong, APISIX, Envoy, Istio) appears on the fetched pages (Get started, API overview).

  • Cloud identity Static provider credentials only

    You paste static provider credentials into this product so it can reach upstream models. Your own calls to it still use its own API key, and those pasted secrets are yours to rotate.

    AWS credentials appear solely as a way to attach your own Bedrock account — a bearer token or IAM access keys with bedrock:InvokeModel, bedrock:InvokeModelWithResponseStream or AmazonBedrockLimitedAccess — and Vertex AI has its own provider-specific credential shape. There is no IAM, OIDC or workload-identity path for authenticating *to* Gateway, which uses its own bearer keys (Amazon Bedrock BYOK, BYOK overview).

  • MCP

    n.a. Gateway documents no MCP support of any kind — no MCP gateway, no hosted MCP server, no MCP tools in the API. The string "mcp" does not appear in any of the 403 Gateway documentation URLs, and the Claude Code integration ships skills rather than an MCP server (Install skills, Coding agents, API overview). MCP is Merge's separate Agent Handler product, not Gateway (Introducing Merge Gateway).

Python frameworks
  • LangChain

LangChain is the only Python framework named, and it is supported by base-URL swap rather than a dedicated package: "LangChain — https://api-gateway.merge.dev/v1/openai" (Get started). LlamaIndex, CrewAI, DSPy and Haystack: n.a. on the fetched pages.

First-party client libraries
  • Python
  • TypeScript

First-party: Python merge-gateway-python (from merge_gateway import MergeGateway) and JavaScript/TypeScript merge-gateway-sdk (import { MergeGateway } from "merge-gateway-sdk"), plus the AI SDK provider merge-gateway-ai-sdk-provider. Third-party clients are supported by base URL: OpenAI, Anthropic and LangChain are named in a table of base URLs (Get started). No Go, Java, Ruby, .NET or Rust SDK is published.

Agent features: Good coverage of the primitives, thin on orchestration. Documented: tool calling, structured outputs (json_object and json_schema with tri-state strict), reasoning via thinking plus provider-native reasoning_effort, a server-side web-search tool merge:web_search (Exa, You.com or Parallel, auto by default, up to 25 results per call and 6 tool iterations), and Fusion, which fans a request across at least two analysis models and has a judge model synthesise the answer (non-streaming only). Multi-turn state exists but is weak: previous_response_id uses an in-memory store that is off by default, expires after an hour, is best-effort and drops thinking blocks (Tool calling, Structured outputs, Web search, Fusion, Multi-turn conversations).

Onboarding is genuinely short — install an SDK or change a base URL, send a request — and the docs are unusually candid about the sharp edges you meet next: the model field becomes optional once a routing policy exists (omit it or pass the sentinel "default_routing"), streams can restart mid-flight with {"fallback_restart": true}, managed credentials are capped at 100 requests and 100,000 tokens per minute per organisation per provider, and inline media is limited to 20 MB (25 MB for audio) before a 413 (Get started, Streaming, Rate limits).

The ecosystem effort is aimed squarely at coding agents. Merge ships a Claude Code plugin (claude plugin marketplace add merge-api/merge-gateway-skills, then claude plugin install merge-gateway) exposing skills such as /merge-gateway:gateway-implement, build-agent, gateway-features and migration skills for the AI SDK, OpenRouter, direct SDKs and Bedrock; Cursor, Zed, Continue.dev, OpenCode, Claude Code, Codex, Pi, Factory Droid and Claude Desktop are documented as clients (Install skills, Coding agents). Written migration guides cover OpenRouter, LiteLLM, Azure OpenAI, Bedrock, the Vercel AI SDK and direct SDKs. Beyond that the ecosystem is thin, as a third-party review notes: "the community, integrations, and third-party tutorials are thin next to LiteLLM or OpenRouter" (Techsy).

Silence in the docs: 5 of the integration questions on this card have no published answer either way. That is recorded as undocumented, not as a no — but it does mean you would be verifying it yourself.

What it does well

  • Genuinely deep spend governance: org credit, project budgets over five periods, per-key spend limits with resets, per-customer budgets, 50/80/90% alert thresholds, a 10x-daily-average velocity alert, and HTTP 402 enforcement before any provider is called
  • Embedded Routing Stack gives each of your customers their own routing policies, BYOK credentials, budgets and API keys, with pass-through usage reporting - unusual among gateways
  • Security controls documented to an unusual level of specificity: Presidio-based DLP, a DeBERTa v3 injection classifier with published thresholds calibrated to 1% FPR, an org-level ZDR toggle that fails closed across 18 vendors, geo restriction by vendor country and region, and a ~60-event append-only audit trail with 29 permissions over 16 resources
  • Intelligent routing with real cost claims (40-60% for Cost Optimized) and a Build Your Own Router mode where you weight benchmarks yourself, plus one named customer saving "more than $10,000 per month"
  • Low-friction adoption: OpenAI, Anthropic, AI SDK and LangChain all work by base-URL swap, first-party Python and TypeScript SDKs, and migration skills for OpenRouter, LiteLLM, Bedrock and the direct SDKs

Where it falls short

  • Hosted only: no self-host artifact, image, chart or architecture page exists, and Enterprise "VPC or on-prem" is a single unelaborated pricing bullet
  • DLP and prompt-injection protection both fail open by default, so a configured blocking policy is telemetry until you set pi_fail_closed and account for DLP's documented fail-open behaviour
  • Semantic caching was announced at launch as shipping "in the next few weeks" and, five months later, appears on no documentation page; Gateway's only caching is provider passthrough
  • Observability is a closed loop: proprietary header-based tracing with ~24h retention, no OpenTelemetry or OTLP export, no log drain, and no published log-retention window
  • Five months old with company-level compliance claims that are written about Merge Unified integration data rather than Gateway, no SLA page, no published uptime figure and no Gateway component on the status page
  • The pricing page contradicts itself ("Credit card required" versus "Start without a credit card") and contradicts the docs on the free cap ($10 versus a documented $15 spend cap for new organisations)

Choose it when

Teams whose LLM bill has become a margin problem and who want routing, budgets and per-customer controls as a managed service — especially B2B products that need to hand each of their own customers separate routing policies, credentials and spend limits without building that plumbing.

Look elsewhere when

You need to self-host or run in your own VPC with documented artifacts, you need OpenTelemetry export into an existing observability stack, or you are early enough that a five-month-old proprietary gateway with no changelog and no status-page component is an unacceptable dependency.

Managed gateway: A hosted control plane in front of the models: routing, caching, guardrails, spend limits, and logs. You trade a little simplicity for governance.

How hard is it to leave?

Derived from six published facts, not from an opinion. The weights are fixed and the same for every product — see the arithmetic.

80 /100 Some work to leave
Portability score breakdown for Merge Gateway
What helps you leave Points Source
Works with standard OpenAI code Switching away is a base-URL change rather than a rewrite of every call site. 22 /22 vendor page
No vendor-specific SDK required A proprietary client library spreads through your codebase and has to be torn out again. 10 /10 vendor page
Can use your own provider accounts Your keys and billing relationship stay yours, so removing the gateway does not cut off model access. 20 /20 vendor page
Can be self-hosted You can run it yourself instead of accepting a pricing or policy change. 0 /20 vendor page
Configuration lives in version control Routing and budget rules are a file you keep, not dashboard state you would have to rebuild. 16 /16 vendor page
Your request history can be exported You leave with your own logs instead of abandoning them. 12 /12 vendor page

Read the fine print: Requests are portable in both directions by base-URL swap, since the gateway speaks OpenAI Chat Completions and Anthropic Messages natively and Merge publishes migration guides from OpenRouter, LiteLLM, Azure OpenAI, Bedrock and the direct SDKs ([Get started](https://docs.merge.dev/merge-gateway/get-started), [Install skills](https://docs.merge.dev/merge-gateway/install-skills)). Configuration is not portable: routing policies, budgets, DLP and prompt-injection settings live only in Merge's dashboard and management API ([API overview](https://docs.merge.dev/merge-gateway/api-overview)).

All six inputs are published, so this is scored against the full 100. This measures technical switching cost only. It does not price the engineering time to re-test prompts against a different routing stack.

Merge Gateway models & pricing

Browse every imported listing from this provider, with published token rates and a link to compare other providers for the same model. This is provider-reported coverage; an absent listing does not mean unsupported.

Loading model listings…

Official model coverage source ↗ · Model source coverage and limitations

Full specification

Every field we track. Blank fields say "Not published" rather than "No" — we do not infer an absence from silence. Switch to Technical in the header for the precise field names and the low-level details.

Overview

What kind of product Category
Managed gateway
Marketplaces resell many providers behind one key. Gateways add governance on top. Open-source projects you run yourself. Cloud platforms are hyperscaler surfaces. Inference providers host models on their own hardware.
Who runs it Deployment model
Managed only
Managed means the vendor operates it. Self-host means you run it on your own infrastructure. Both means you can choose.
Licence Licence
Proprietary
Proprietary products cannot be inspected or forked. Open licences such as MIT and Apache-2.0 let you audit, modify, and run the code without permission.
Company Company
Merge API, Inc.
The organisation that maintains the product.
Who you would be signing with Vendor status
Independent company
Whether the product is still an independent company, has been acquired, is a large cloud vendor’s product line, is run by a software foundation, or has been put into maintenance mode. Maintenance mode means bug fixes and security patches only — no new features.
Last shipped an update Latest release
2026-08-11

Not a product release: Merge publishes no changelog or release notes for Gateway, and no changelog URL exists in the docs sitemap. 2026-08-11 is the last public code push to `merge-api/merge-gateway-ai-sdk-provider`, the only actively pushed public Gateway repository; the skills pack was last pushed 2026-04-13 and neither repo has a tagged release ([merge-gateway-ai-sdk-provider](https://api.github.com/repos/merge-api/merge-gateway-ai-sdk-provider), [merge-gateway-skills](https://api.github.com/repos/merge-api/merge-gateway-skills)).

The date of the most recent release or version tag. A product that has not shipped in a year is a different risk from one that shipped last week, regardless of what its marketing site says.
GitHub stars GitHub stars
Not published
A rough proxy for community size on open-source projects. Not a quality measure.

Cost

Markup on model prices Token markup
5%
How much the product adds on top of what the underlying model provider charges. Zero means you pay the same per-token price you would pay the model provider directly.
Fee to add funds Credit purchase fee
Not published
A percentage charged when you top up your balance, separate from token prices. It is easy to miss because it does not appear on the per-token price list.
Monthly cost per person Seat fee
Not published
A recurring per-user platform charge that applies regardless of how much you use the models.
Can use your own provider accounts BYOK supported
Yes
Bring Your Own Key: you keep direct contracts with OpenAI, Anthropic and others, and the gateway only routes traffic. This preserves negotiated rates and committed-spend discounts.
Cost of using your own accounts BYOK terms
BYOK is gated to the Pro and Enterprise plans; one credential per provider per organisation, stored encrypted at rest, with an optional "Fall back to Merge managed credentials" toggle. No separate BYOK per-token or platform fee is published on the pricing page ([BYOK overview](https://docs.merge.dev/merge-gateway/capabilities/byok/overview), [Gateway pricing](https://www.merge.dev/pricing/gateway)).
What the product charges to route traffic through your own provider keys.
Enterprise plan from Enterprise plan from
Not published
Annual entry price for the enterprise tier, where one is published or credibly reported.
Cost to run it yourself Self-host cost
n.a. There is no self-hosted artifact to cost: no Docker image, Helm chart, binary, npm package or install command appears anywhere in the 403-URL Gateway documentation sitemap, and the only self-hosting language is a single Enterprise sales bullet, "Deploy to your own VPC or on-prem environment" ([Gateway pricing](https://www.merge.dev/pricing/gateway), [Get started](https://docs.merge.dev/merge-gateway/get-started)).
What self-hosting actually costs once you account for infrastructure and any paid tier.
How the vendor makes money Pricing model
Percentage on tokens or top-ups
The shape of the vendor’s bill: does the routing layer charge a percentage on top of tokens, a flat monthly fee, both, neither (because inference is the product), or nothing at all (open source with no paid tier).
How pricing works, briefly Pricing model detail
Pro is "Pay LLM cost plus a 5% fee" with no platform fee, no seat fee and no request fee; Free is "Free, capped at $10 LLM usage" and Enterprise is custom ([Gateway pricing](https://www.merge.dev/pricing/gateway)). Credits are prepaid and consumed per request, expire one year after purchase, are non-refundable, and can be auto-replenished ([Merge Gateway Terms](https://www.merge.dev/legal/gateway-terms), last updated 2026-04-02).
A one-paragraph description that covers the caveats a pricing category cannot: introductory rates, per-feature meters, tier gating, and pricing that resets on a specific date.
Charges that fire after you go over an allowance Overage terms
There is no request or seat meter to overrun; the meter is model spend. Free is hard-capped ($10 on the pricing page, $15 in the docs) and Pro "scale[s] without a spending cap" ([Gateway pricing](https://www.merge.dev/pricing/gateway), [Cost governance and savings](https://docs.merge.dev/merge-gateway/cost/cost-governance-and-savings)). Exhausted org credit, project budgets and per-key limits return HTTP 402 with `budget_exceeded`, `project_budget_exceeded` or `api_key_limit_exceeded` before any provider is called ([Errors](https://docs.merge.dev/merge-gateway/errors), [How it works](https://docs.merge.dev/merge-gateway/how-it-works)).
The line items that scale with usage after an included allowance is exhausted — log storage, extra requests, per-feature meters, data export — which is where cost estimates usually go wrong.
Prompt cache offered Cache mechanism
Passes provider caching through
Whether the gateway offers its own response cache, what kind of match it does (exact request, prefix, semantic), or simply passes provider caching through unchanged.
Discount on cached input Cache-read discount
Not published
How much cheaper cached tokens are than fresh input, when the vendor publishes a single figure. Bundled-inference clouds usually price this per model instead of as one number.
Premium on cache writes Cache-write premium
Not published
How much more the first write of a cached prefix costs versus a plain input token. A high write premium and a low hit rate can leave you paying more than you save, so this matters as much as the read discount.
Who captures the cache saving Cache economics
Gateway operates no cache of its own; it passes provider prompt caching through. Models are classed `explicit` (Anthropic, Claude and Amazon Nova on Bedrock — you place the markers), `automatic` (OpenAI, Gemini, DeepSeek, Qwen, Mistral, Z.AI, xAI) or `none`, and an `X-Session-Id` header (256 chars, namespaced per organisation) improves cache affinity; unsupported markers are stripped rather than erroring ([Prompt caching](https://docs.merge.dev/merge-gateway/capabilities/prompt-caching)). No cached-token discount or write premium is published, and no gateway-side exact or semantic cache exists — semantic caching was announced at launch as arriving "in the next few weeks" ([Introducing Merge Gateway](https://www.merge.dev/blog/gateway-announcement)) and, as of 2026-09-02, appears on no documentation page.
Whether the customer keeps the full saving from caching or the vendor captures part of it — and any conditions attached (write premium, storage fees, best-effort hits).
Who pays the model bill BYOK mode
Your keys or their credits
Whether you bring your own provider accounts (BYOK), buy inference from this vendor, or can do either. This is the single biggest commercial difference between these products: it decides who holds the contract with the model provider and who carries the spend.

Spend governance in detail

Seven signals matter when a bill starts to hurt: who can spend, how much, on what, and who gets paged when it goes wrong. Everything below is drawn from the vendor’s own pricing and docs pages — how we read these.

  • Virtual or scoped keys Yes

    Keys that carry their own budget and rate-limit policy, so an intern experiment cannot spend against a production budget.

    Not called virtual keys, but the object exists: API keys minted per project, environment or customer via the dashboard or Management API `POST /v1/keys`, each with its own spend limit and reset period.

  • Budget caps per key Yes

    A dollar or token ceiling attached to an individual key. Where enforcement is soft, one over-limit request still completes before the block kicks in.

    Per-key spend limits with `limit_reset` of daily, weekly or monthly; exhaustion returns HTTP 402 `api_key_limit_exceeded` and `X-Key-Limit-Remaining-USD` tracks the headroom.

  • Budget caps per team or workspace Yes

    A ceiling applied at a higher scope than one key — a team, a workspace, a customer, or an entire environment.

    Projects are the team/workload unit and carry budgets over daily, weekly, monthly, quarterly or yearly periods, with Soft Limit (alert) or Hard Limit (HTTP 402) enforcement; per-customer budgets exist separately for embedded tenants.

  • Rate limiting as a cost control Not published

    Configurable request-per-time-window caps. Platform-set rate limits do not count as spend controls; user-configurable ones do.

    Not stated as a customer-configurable control. `X-RateLimit-*` headers read 0 unless a per-key limit is set, and the only published caps are provider-facing: 100 requests and 100,000 tokens per minute per organisation per provider on managed credentials, which BYOK traffic bypasses. No concurrency or body-size cap is published.

  • Model allowlists Yes

    A policy that constrains which models a key or team can call, keeping expensive frontier models out of the wrong hands.

    Per-customer block and pin rules over provider/model combinations (maximum 100 rules per organisation, rejections as HTTP 403 `customer_blocked`, `customer_pinned`, `provider_blocked` or `model_blocked`), plus organisation-level `allowed_vendors`/`ignored_vendors` and region restrictions. Dashboard-configured, with an audited `reason`.

  • Spend alerts Yes

    Alerts fired as spend approaches a threshold. Alerts that only fire after the meter has rolled over are marked as such.

    Threshold alerts default to 50%, 80% and 90% of budget, plus a velocity alert when spend hits 10x the daily average.

  • Webhook notifications Not published

    Programmatic notifications on spend events, so budget breaches can page an on-call or open a ticket.

    n.a. No webhook or subscription mechanism for budget or alert events appears on the budgets, cost-governance or projects pages.

Enforcement: Enforced before each request

Source →

Catalog

Models available Models available
3
How many models you can call, shown as a range because several vendors publish different totals on different pages. Vendor-reported either way, so counts are not directly comparable — some count every provider variant of the same model separately.
Model providers reachable Upstream providers
4–24
How many distinct model providers or labs you can reach, shown as a range where the vendor’s own pages disagree. More providers usually means better redundancy when one has an outage.
Works with standard OpenAI code OpenAI-compatible API
Yes
If yes, you can usually switch to it by changing one base URL, and switch away just as easily. This is the main defence against lock-in.
OpenAI chat endpoint POST /v1/chat/completions
Yes
The endpoint almost every application ports first. "Not documented" means the vendor never states it, which is different from a documented no.
Anthropic messages endpoint POST /v1/messages
Yes
Whether Anthropic-shaped calls work without rewriting them. Several products support this only as SDK compatibility or provider passthrough rather than a native endpoint — the detail page says which.
OpenAI Responses endpoint POST /v1/responses
Partly
The newer stateful OpenAI surface. Support is much thinner across this market than chat completions.
Embeddings endpoint POST /v1/embeddings
Yes
Whether you can generate vectors through the same gateway, or need a second integration for retrieval workloads.
Image generation endpoint POST /v1/images/generations
Yes
Whether image models are reachable through the same surface as text.
Audio endpoints POST /v1/audio/*
Partly
Speech-to-text and text-to-speech. Frequently the first gap in an otherwise complete gateway.
Batch jobs endpoint POST /v1/batches
Not documented
Asynchronous bulk processing, usually at a discount. Commonly undocumented, and commonly the reason a migration stalls late.
Needs the vendor’s own code library Requires a vendor-specific SDK
No
A proprietary client library spreads through your codebase and has to be torn out again if you leave. “No” is the better answer here, and it means the standard OpenAI client works.
You can export your request history Logs / usage data export
Yes
Whether you can get your own request logs, traces, or usage records back out — through an API, a bulk export, or a download. Decides whether you leave with your history or abandon it.
Settings can live in version control Declarative config-as-code
Yes
Whether routing, fallback, and budget rules can be declared in a file you keep in Git, rather than existing only as settings clicked into a hosted dashboard.
Embeddings Embeddings
Yes
Text-to-vector models, needed for search and retrieval features.
Image generation Image generation
Yes
Whether image models are reachable through the same interface.
Speech and audio Speech and audio
Yes
Text-to-speech or transcription models through the same interface.
Video generation Video generation
Yes
Whether video models are reachable through the same interface.
Batch processing Batch processing
Not published
Submitting large jobs for cheaper, slower processing. Often 50% off for work that is not time-sensitive.

Routing & reliability

Uptime it promises in writing Contractual SLA uptime
Not published
The uptime percentage in a published, contractual service level agreement. A public status page is not an SLA — it reports what happened, it does not promise anything or pay you back when it breaks.
Automatic failover Automatic failover
Yes
When a model provider goes down or rate-limits you, traffic moves to a backup automatically instead of returning errors to your users.
Load balancing Load balancing
Not published
Spreads requests across several providers or keys to raise your effective rate limit.
Rule-based routing Conditional routing
Yes
Send different requests to different models based on rules — for example a cheap model for free users and a strong model for paying ones.
Response caching Response caching
Not published
Reuses the answer when the exact same request comes in again, which cuts both cost and latency.
Similar-question caching Semantic cache
Not published
Reuses an answer when a new question means roughly the same thing as an earlier one. Saves far more than exact-match caching but can return subtly wrong answers if tuned loosely.
Where you set the timeout Request timeout surface
Not documented

not_documented as a user-settable value: no timeout header, SDK parameter or config key appears on the fetched pages. What is published is a fixed stall watchdog — a first chunk taking more than 120 seconds, or a 120-second gap between chunks, abandons the upstream and yields a 408/502 or a failover — plus the warning that a client disconnect is still billed ([Streaming](https://docs.merge.dev/merge-gateway/streaming), [API overview](https://docs.merge.dev/merge-gateway/api-overview)).

Where a request timeout can be set: per request, in a config file, in the vendor dashboard, or nowhere. Nine of the twenty products documented here do not describe a request timeout at all, so the worst case of a hung upstream call is unknowable from the docs.
Where you set retries Retry policy surface
Not documented

No retry counter, backoff strategy or retryable-status list is published; the documented failure behaviour is failover to the next model in the policy, not a retry against the same one. Provider 429s, 5xx and timeouts trigger failover, while client errors 400-404 do not. Default retry count: n.a. Backoff: n.a. ([Deterministic strategies](https://docs.merge.dev/merge-gateway/routing/deterministic-strategies), [Errors](https://docs.merge.dev/merge-gateway/errors)).

Where retry count and backoff are configured. Worth knowing alongside billing: a retried streaming call can be charged more than once.
Where you set fallbacks Fallback surface
Per request

ORDERED. A Priority strategy holds an ordered list of models and walks it on provider failure (`strategy: "PRIORITY"` over the API); the policy is created in the dashboard or via `POST /v1/routing-policies` and then selected per request with `routing_policy_id`, with precedence request > customer default > project > org default > raw `model` ([Deterministic strategies](https://docs.merge.dev/merge-gateway/routing/deterministic-strategies), [Using policies](https://docs.merge.dev/merge-gateway/routing/using-policies)). Mid-stream failover is visible to the client as a `{"fallback_restart": true}` frame ([Streaming](https://docs.merge.dev/merge-gateway/streaming)).

Where the fallback chain is defined. Almost every product claims fallback; the useful question is whether you can change it from code or only by hand in a dashboard.
Shape of the fallback chain Fallback shape
Ordered list
An ordered list tries targets in sequence; a weighted split sends a percentage of traffic to each, which is what you need to trial a new model on 5% of requests. Weighted splits are much rarer than the marketing implies.
Upstream health tracking Health checks / circuit breaking
Fixed, cannot change

The feature exists but exposes no knobs: Gateway tracks provider health and automatically skips routes that are failing, and a tripped circuit breaker surfaces as HTTP 502 `model_vendor_unhealthy`. No probe interval, failure threshold, ejection window or half-open policy is published ([Deterministic strategies](https://docs.merge.dev/merge-gateway/routing/deterministic-strategies), [Errors](https://docs.merge.dev/merge-gateway/errors)).

Whether the product notices a failing upstream and stops sending traffic to it, and whether you can tune the thresholds. This is what turns a provider outage into a blip rather than a sustained error rate, and only four of the twenty expose it.
Cross-region failover you control Multi-region failover surface
Not documented

not_documented as failover. Geo-location routing is a compliance restriction that narrows the vendor set by country and region; it is not described as a cross-region failover mechanism, and Merge publishes nothing about its own regional topology for Gateway ([Geo-location routing](https://docs.merge.dev/merge-gateway/security/geo-location-routing), [How it works](https://docs.merge.dev/merge-gateway/how-it-works)).

Whether you can define what happens when a region degrades. A vendor running many regions is not the same as a vendor letting you configure failover between them; only three document a user-controlled mechanism.
Where you set load balancing Load balancing surface
Dashboard only

Weights are not user-settable for traffic splitting. Least Latency and Lowest Cost are dashboard-only strategies with no exposed weights, and the API accepts only `PRIORITY` and `INTELLIGENT`. The one weighted surface is Build Your Own Router, where you weight *benchmarks* (summing to 1.0) to score models, not traffic shares — and it too is dashboard-only ([Deterministic strategies](https://docs.merge.dev/merge-gateway/routing/deterministic-strategies), [Build your own router](https://docs.merge.dev/merge-gateway/routing/build-your-own-router)).

Where traffic distribution across upstreams or keys is configured.

Operations

Usage dashboards and logs Observability
Yes
Built-in visibility into what was sent, what came back, what it cost, and how long it took.
Spending limits Budget controls
Yes
Hard caps that stop spend before it becomes a surprise invoice. The single most valuable control for a small team.
Rate limits Rate limits
Not published
Caps on request volume per key or per user, useful for protecting against abuse and runaway loops.
Separate keys per team or app Virtual keys
Not published
Issue scoped keys with their own budgets and permissions so you can attribute cost and revoke access without rotating everything.
Prompt versioning Prompt management
Not published
Store and version prompts outside your code so they can be changed without a deploy.
Quality testing Evals
Yes
Built-in tooling to score model output against test cases, so you can tell whether a model swap made things better or worse.
MCP support MCP support
Not published
Native support for the Model Context Protocol, the emerging standard for connecting models to external tools.
What gets logged Logged content
Your choice

Configurable, and off by default: metadata only unless payload logging is enabled in settings, at which point inputs and outputs are stored for all requests and visible in the Logs detail view ([Merge Gateway Terms](https://www.merge.dev/legal/gateway-terms)). Where DLP redaction applies, the redacted text is what gets logged ([Data loss prevention](https://docs.merge.dev/merge-gateway/security/data-loss-prevention)).

Whether full prompts and responses are stored, only metadata, or nothing. Full-body logging is the most useful debugging feature here and the one most likely to need a conversation with your compliance team.
You can turn logging off Body-logging opt-out
Yes

Yes, and it is the default state rather than an opt-out you have to request: payload logging is an account setting that starts disabled ([Merge Gateway Terms](https://www.merge.dev/legal/gateway-terms)). A per-request suppression header is not documented ([API overview](https://docs.merge.dev/merge-gateway/api-overview)).

Whether prompt and response bodies can be suppressed while still keeping usage metrics. Nineteen of the twenty document a way to do this; the mechanisms range from a per-request header to an organisation-wide setting.
Traces you can take elsewhere Distributed tracing
Vendor format only

Proprietary and header-based: pass `X-Merge-Trace-Id`, `X-Merge-Parent-Span-Id`, `X-Merge-Span-Name` and `X-Merge-Thread-Id` to stitch multi-step agent runs, with cost rolled up from the billing pipeline and Fusion runs generating child spans automatically. Traces are retained about 24 hours. No OpenTelemetry or OTLP export is offered, and `observability/tracing` is the only observability page in the 403-URL docs sitemap ([Tracing](https://docs.merge.dev/merge-gateway/observability/tracing)).

Whether the product emits OpenTelemetry, a proprietary format, or nothing. OpenTelemetry means the traces land in the tooling you already run instead of only in the vendor’s dashboard.
Where telemetry can go Export destinations
Not published

n.a. No OTLP endpoint, log-drain, S3/warehouse sink or named third-party destination (Datadog, Grafana, LangSmith and so on) appears on any fetched page; the docs sitemap contains no OpenTelemetry page at all ([Tracing](https://docs.merge.dev/merge-gateway/observability/tracing), [API overview](https://docs.merge.dev/merge-gateway/api-overview)). The nearest thing is that `usage.cost` is returned inline on every response, which third-party tools such as Langfuse pick up automatically ([Cost governance and savings](https://docs.merge.dev/merge-gateway/cost/cost-governance-and-savings)).

Documented sinks for logs and metrics. This is a good proxy for how replaceable the vendor’s own dashboard is: around twenty destinations means you never have to depend on it, while a CSV download means you do.
Can record user feedback Feedback capture API
No

n.a. — recorded as no because no score, rating or feedback ingestion endpoint appears on the fetched pages ([API overview](https://docs.merge.dev/merge-gateway/api-overview), [Tracing](https://docs.merge.dev/merge-gateway/observability/tracing)); the trace surface is write-time headers only, with no companion endpoint for attaching outcomes after the fact.

Whether there is an API to attach a rating or score to a logged request, which is what lets production traffic feed quality work later.
Scores live traffic Online eval hooks
Partly

Partial and routing-bound: Build Your Own Router can "Run evals" against a judge model and rubric, will auto-evaluate newly added models, and can source scores from published benchmarks, your own eval runs or uploaded results — and `EVAL_*` events appear in the audit trail. There is no eval API, dataset object or CI hook, and evals exist to weight routing rather than to test your application ([Build your own router](https://docs.merge.dev/merge-gateway/routing/build-your-own-router), [Audit trail](https://docs.merge.dev/merge-gateway/security/audit-trail)).

Whether automated scorers can run against real production requests, rather than only against a test set you assemble yourself.

Performance

Delay it adds Proxy overhead
Not published
Extra time the product itself adds to each request, on top of however long the model takes. Usually irrelevant next to multi-second model latency, but it matters for high-volume or streaming-sensitive workloads.
Requests per second ceiling Throughput
Not published
Published sustained request rate before the product becomes the bottleneck. Only relevant at genuinely high volume.
What the request path runs on Architecture class
Undisclosed vendor service

vendor_saas: closed multi-tenant service reached at `https://api-gateway.merge.dev/v1`, with no published source, runtime or language for the data plane. The only public Merge Gateway repositories are a Claude Code skills pack and the Vercel AI SDK provider (TypeScript, MIT), neither of which is the gateway itself ([merge-gateway-ai-sdk-provider](https://api.github.com/repos/merge-api/merge-gateway-ai-sdk-provider), [merge-gateway-skills](https://api.github.com/repos/merge-api/merge-gateway-skills)).

The comparable way to talk about latency here. An edge worker, a compiled Go or Rust binary, and a Python proxy have different overhead floors no matter which figures each vendor publishes. Products that never disclose their runtime are recorded as undisclosed rather than assumed.
You can run the request path yourself Self-hostable data plane
Not documented

No Docker image, Helm chart, binary, npm package or install command for a self-hosted data plane appears on any fetched page, and no such page exists in the 403-URL Gateway docs sitemap ([Get started](https://docs.merge.dev/merge-gateway/get-started), [Gateway pricing](https://www.merge.dev/pricing/gateway)). Enterprise "Deploy to your own VPC or on-prem environment" is a pricing-page bullet with no artifact behind it.

Whether the component that actually carries your prompts can run on your own infrastructure. Distinct from a vendor offering a self-hosted control plane while still proxying traffic through their network.
Streaming responses Streaming support
Yes

Supported on every surface with `stream: true`, and unusually well documented. Usage and a `cost` figure arrive on the terminal frame (omitted on the Anthropic and AI SDK surfaces); a mid-stream failover emits `{"fallback_restart": true}` so clients must be prepared to discard partial output; the 120-second stall watchdogs apply to both first byte and inter-chunk gaps; and a client disconnect is still billed ([Streaming](https://docs.merge.dev/merge-gateway/streaming)).

Whether token-by-token streaming is documented. The caveats matter more than the yes: some products cannot cancel a stream without still being billed, and several timeout and fallback mechanisms stop applying once the first token has been sent.

Security & compliance

Does your prompt reach their servers Prompt transits vendor
Yes

Always. Every request crosses Merge's hosted endpoint before reaching a provider, and there is no self-hosted or in-VPC data plane documented; the pipeline (validate, meter, filter, route, forward) runs entirely on Merge's side ([How it works](https://docs.merge.dev/merge-gateway/how-it-works), [Get started](https://docs.merge.dev/merge-gateway/get-started)).

Whether the text you send passes through this company’s own infrastructure. If it does, every other promise on this page is a policy commitment rather than a physical impossibility. Self-hosted products can answer no outright.
What they keep if you change nothing Logging default
Metadata only, not content

Payload logging is off by default and is a setting you turn on: "Except to the extent Customer enables payload logging via settings, Merge will not retain Inputs or Outputs and will only retain metadata regarding each request and response. When payload logging is enabled, Merge will store Inputs and Outputs for all requests, which can be viewed in the Logs detail view" ([Merge Gateway Terms](https://www.merge.dev/legal/gateway-terms), last updated 2026-04-02). Metadata is still rich: routing decision, cost, latency and policy outcome per request ([Introducing Merge Gateway](https://www.merge.dev/blog/gateway-announcement)).

Defaults matter more than options. A product that stores full prompts and replies unless you find the right header will have stored them by the time you read the docs.
How long they keep it Default content retention (days)
Not published

No log-retention window is published. What is stated: metadata only unless payload logging is enabled, and "[w]ithin thirty (30) days of termination or expiration of this Agreement for any reason, Merge will, upon written request, delete all Customer Data in Merge's possession" ([Merge Gateway Terms](https://www.merge.dev/legal/gateway-terms)). Traces are the one figure given — retained about 24 hours ([Tracing](https://docs.merge.dev/merge-gateway/observability/tracing)).

Default retention for request content, in days. Zero means nothing is kept. Read the note: several products keep nothing as a rule but make timed exceptions for abuse review or specific models.
Could they train on your prompts Training on customer data
Not published — silence, not a no

Merge makes no statement about training on Gateway customer data. The Gateway Terms grant only a licence to use Customer Data "to provide and maintain the Service and develop Usage Data", where Usage Data is defined as analytics and performance information that "does not consist of personal data" — silence, not a commitment ([Merge Gateway Terms](https://www.merge.dev/legal/gateway-terms)). The only training language on any Gateway page is about third-party vendors under ZDR ([Zero data retention](https://docs.merge.dev/merge-gateway/security/zero-data-retention)).

Whether the vendor may use your prompts and outputs to train models. “Not published” means we could not find any position, which is not the same as a no — ask for it in writing.
Where it runs, and what you can pin Region and residency control
Vendor-side region control is strong and customer-configurable: every vendor carries a `country_code` and a `geo_region` (NA, EU, APAC, CN), and an organisation sets `allowed_regions`/`ignored_regions` and `allowed_vendors`/`ignored_vendors`, which also apply to BYOK traffic and are recorded in the audit trail as `ORG_SETTINGS_UPDATED` ([Geo-location routing](https://docs.merge.dev/merge-gateway/security/geo-location-routing)). Merge's own hosting regions for Gateway are not stated on any Gateway page.
Which regions are offered and whether you can force processing to stay in one. A global endpoint that silently picks a region is a different compliance story from an endpoint you pin yourself.
Where safety filters run Guardrail execution location
In the vendor’s cloud

All enforcement runs inside Merge's hosted pipeline, at step 3 of five (validate, meter, filter, route, forward), before any provider is called; there is no self-hosted or in-VPC option in which the checks would run on your own infrastructure ([How it works](https://docs.merge.dev/merge-gateway/how-it-works), [Data loss prevention](https://docs.merge.dev/merge-gateway/security/data-loss-prevention)).

A filter that strips personal data only helps if it runs before the data leaves your boundary. If guardrails execute in the vendor’s cloud, the vendor has already received whatever you wanted redacted.
Who else touches the data Subprocessor list
https://www.merge.dev/legal/data-subprocessors
The published list of third parties the vendor passes your data to. No list means you cannot know the full chain, which most data-protection agreements require you to.
SOC 2 audited SOC 2 audited
Not published
An independent audit of security controls. Enterprise buyers and their procurement teams routinely require it.
Will sign a HIPAA agreement HIPAA BAA
Not published
Required before you may send protected health information through the service. Without a signed BAA, healthcare data is off limits.
GDPR commitments GDPR commitments
Not published
Published data processing terms for handling personal data of people in the EU and UK.
Can keep data in the EU EU data residency
Not published
Requests can be processed inside the EU rather than routed to US infrastructure. Often the deciding constraint for European customers.
Does not retain your data Zero data retention
Yes
Prompts and responses are not stored after the request completes. Sometimes a paid add-on rather than the default.
Strips personal data PII redaction
Yes
Detects and removes identifiers such as names, emails, and card numbers before the request reaches the model provider.
Content guardrails Content guardrails
Yes
Policy checks on inputs and outputs — blocking unsafe content, enforcing formats, or catching prompt-injection attempts.
Runs fully disconnected Air-gapped deployment
Not published
Can be deployed in a network with no internet access, which some regulated and defence environments require.
Blocks personal data in prompts PII / DLP enforcement
Can block the request

A Presidio-based DLP sidecar scans both prompts and completions against Global, USA and Custom category sets (15 seeded entity types) with a per-rule action of log, redact or block; blocking returns HTTP 400 (the aggregate page states 422 `blocked_by_dlp_policy`) and redaction means the redacted text is what gets logged. Vendor-stated cost: "a few milliseconds" ([Data loss prevention](https://docs.merge.dev/merge-gateway/security/data-loss-prevention)).

Whether personal data detection sits on the request path and can stop the call, merely inspects and forwards it, or is not documented. A control that only reports is a logging feature, not a policy control.
Blocks prompt injection Injection / jailbreak enforcement
Can block the request

A fine-tuned DeBERTa v3 classifier with modes off, alert and block, calibrated to a 1% false-positive rate: `pi_block_threshold` 0.57, `pi_pass_threshold` 0.30, `pi_input_action` block and `pi_output_action` redact (from observe, redact, route, block, escalate), plus tiered indirect-injection handling and an allowlist of regexes (50 patterns of 200 characters at org level, up to 20 more per project, unioned). Blocks return 400/422 `pi_blocked` and every log carries a `pi_score` ([Prompt injection protection](https://docs.merge.dev/merge-gateway/security/prompt-injection-protection)).

Whether injection and jailbreak detection can stop a request. Most products offering this call a partner classifier rather than shipping their own.
Blocks harmful content Toxicity / moderation enforcement
Not documented

n.a. No toxicity, moderation or content-category filter is offered; the security suite is DLP, prompt injection, ZDR, geo routing and per-customer blocklists, and content refusals come from the provider's own policy as `content_policy_violation` with `source: "provider"` ([Data loss prevention](https://docs.merge.dev/merge-gateway/security/data-loss-prevention), [Prompt injection protection](https://docs.merge.dev/merge-gateway/security/prompt-injection-protection), [Errors](https://docs.merge.dev/merge-gateway/errors)).

Whether hate, violence, sexual and self-harm categories are checked inline and can stop a request, in either direction.
Your own policy rules Custom policy hooks
Can block the request

Custom DLP categories accept your own Python-flavoured regexes (500 characters) and keyword lists (50 entries) with the same log/redact/block actions, and a built-in tester accepts up to 50,000 characters of sample text ([Data loss prevention](https://docs.merge.dev/merge-gateway/security/data-loss-prevention)). A webhook or arbitrary-code guardrail hook is not offered; custom routing classifiers are a separate, dashboard-only centroid scorer requiring at least three simple and three complex examples ([Custom classifiers](https://docs.merge.dev/merge-gateway/routing/custom-classifiers)).

Whether you can add your own rule — a regex, a webhook, or your own classifier — rather than choosing from the vendor library.
Where guardrails run Guardrail execution location
On the vendor's servers
Whether guardrail evaluation happens inside your infrastructure or on the vendor’s servers. This decides whether the prompt you are trying to protect leaves your network in order to be checked.
If the guardrail itself fails Guardrail failure mode
You choose

Read this one carefully before trusting the controls: "By default, DLP fails open and the request proceeds", and prompt-injection checks also fail open unless you set `pi_fail_closed: true`. So a freshly configured blocking policy will let traffic through if the classifier or sidecar errors, until you change that ([Data loss prevention](https://docs.merge.dev/merge-gateway/security/data-loss-prevention), [Prompt injection protection](https://docs.merge.dev/merge-gateway/security/prompt-injection-protection)). ZDR is the exception and fails closed with HTTP 400 ([Zero data retention](https://docs.merge.dev/merge-gateway/security/zero-data-retention)).

What happens when the guardrail service times out or errors: does the request proceed unchecked, or is it blocked? This is the worst-documented field in the entire catalogue — only two vendors state it plainly, which means most teams are running a control whose failure behaviour they cannot know.
Third-party guardrail vendors Guardrail integrations
Microsoft Presidio
Named external guardrail services the product can call. A long list means the product is a router for policy engines rather than a policy engine itself — which also means another vendor bill and another hop.

Compliance evidence

Graded by how strong the evidence is, not whether the word appears on the vendor’s website. An audited report and a marketing claim are different things, and only one of them will satisfy your own auditor.

  • SOC 2 Claimed, no evidence published Company-level claim only: "Merge adheres to industry-standard compliance frameworks, including SOC 2 Type II, ISO 27001, HIPAA, GDPR, and CCPA" on merge.dev/security, a page written around Merge Unified integration data (Linked Accounts, selective sync). No Gateway page, the Gateway Terms or the pricing page repeats it, and no trust portal or report is named.
  • ISO 27001 Claimed, no evidence published Same company-level sentence on merge.dev/security; not restated for Gateway.
  • GDPR DPA Claimed, no evidence published Merge publishes a Data Processing Agreement and a GDPR page in its legal index and lists CCPA/GDPR in the company compliance sentence; the Gateway Terms do not incorporate the DPA by reference.
  • HIPAA BAA Claimed, no evidence published A Business Associate Agreement is published in Merge's legal index and HIPAA appears in the company compliance sentence; neither is scoped to Gateway, and the Gateway Terms do not reference a BAA.
  • FedRAMP Not published
  • ITAR Not published

Vendor source

Fit & integration

Work to try it Evaluation work shape
Change one base URL
The shape of the work on the vendor’s own quickstart, from swapping one base URL through to deploying infrastructure. An ordinal class rather than a duration, because elapsed time depends on accounts and quota we cannot see.
Work to run it Production work shape
Change one base URL
The same scale applied to the vendor’s recommended production path. For several products this is much heavier than the quickstart, which is exactly why both are recorded.
Steps on the quickstart Numbered quickstart steps
2

Two steps on the SDK path ("Install the SDK", "Send a request"); the page then repeats the same two-step shape for the OpenAI SDK, the Vercel AI SDK provider and a URL shim, so there is no single canonical numbered list.

A literal count of numbered steps on the vendor’s quickstart, recorded as evidence beside the work shape. Zero means the page publishes no numbered procedure at all. Large counts usually mean interleaved language tracks rather than more work.
Can you self-host it today Self-host install documentation
No self-hosting
Whether an install command is actually published. Several products advertise self-hosting while publishing no command to start from, which a plain yes/no would hide.
Works with the OpenAI SDK OpenAI SDK drop-in
Yes

Yes, and documented as a one-line change: point the OpenAI SDK at `https://api-gateway.merge.dev/v1/openai` (or the bare `/v1`, which "matches the OpenAI SDK's default base URL shape") and call `chat.completions.create` with unprefixed model names ([Get started](https://docs.merge.dev/merge-gateway/get-started)). One caveat: `client.responses.*` against the bare `/v1` base URL hits Gateway's native Responses API instead of OpenAI's.

Whether an existing OpenAI-compatible client can be pointed at it by changing the base URL and key.
Vercel AI SDK support AI SDK provider package
Official provider package

Official first-party provider package `merge-gateway-ai-sdk-provider` with `createMergeGateway` and typed Gateway options under `providerOptions.mergeGateway`; `generateObject`/`streamObject` map to Gateway's native `json_schema` structured output. AI SDK v6 uses the root import, v5 the `/v5` subpath, and v4 is not supported (use the `/v1/ai-sdk` URL shim). The package source is public and MIT-licensed ([Get started](https://docs.merge.dev/merge-gateway/get-started), [merge-gateway-ai-sdk-provider](https://api.github.com/repos/merge-api/merge-gateway-ai-sdk-provider)).

Whether a first-party AI SDK provider package exists, or only a community package, a documented workaround, or the generic OpenAI provider pointed at a custom base URL.
Python framework integrations Documented Python frameworks
LangChain

LangChain is the only Python framework named, and it is supported by base-URL swap rather than a dedicated package: "LangChain — `https://api-gateway.merge.dev/v1/openai`" ([Get started](https://docs.merge.dev/merge-gateway/get-started)). LlamaIndex, CrewAI, DSPy and Haystack: n.a. on the fetched pages.

LangChain, LangGraph, LlamaIndex and similar orchestration frameworks with a documented integration.
Callable from Cloudflare Workers Cloudflare Workers support
Not documented

n.a. Cloudflare Workers are neither a documented deployment target nor a documented runtime for Gateway; no edge or Workers page exists in the docs sitemap ([Get started](https://docs.merge.dev/merge-gateway/get-started), [Coding agents](https://docs.merge.dev/merge-gateway/coding-agents/overview)).

Whether the docs show calling this product from your own Worker. Deliberately separated from the several products whose own gateway runs on Workers, which is a fact about their infrastructure and not about your edge compatibility.
Kubernetes install Helm chart availability
Not documented

n.a. There is nothing to run in a cluster: no Helm chart, manifest, operator or container image is published, and "helm" and "kubernetes" return no hits across the 403-URL Gateway docs sitemap ([Get started](https://docs.merge.dev/merge-gateway/get-started), [Gateway pricing](https://www.merge.dev/pricing/gateway)).

Whether a named, published Helm chart exists, versus Helm being referenced with no chart named, versus generic cluster documentation that has nothing to do with this product.
Terraform support Terraform provider or modules
Not documented

n.a. No Terraform provider or module is published and "terraform" returns no hits in the docs sitemap. Infrastructure-as-code, if you want it, means driving the REST Management API yourself — keys, projects, routing policies and guardrail settings are all API-addressable ([Projects guardrails API](https://docs.merge.dev/merge-gateway/automation/projects-api/guardrails), [API overview](https://docs.merge.dev/merge-gateway/api-overview)).

Whether you can declare this in code: an official provider, official modules, resources inside a hyperscaler’s provider, a community provider, or only Terraform code shipped in a repo.
Reuses your cloud identity Cloud IAM reuse
Static provider credentials only

upstream_credentials_only: AWS credentials appear solely as a way to attach your own Bedrock account — a bearer token or IAM access keys with `bedrock:InvokeModel`, `bedrock:InvokeModelWithResponseStream` or `AmazonBedrockLimitedAccess` — and Vertex AI has its own provider-specific credential shape. There is no IAM, OIDC or workload-identity path for authenticating *to* Gateway, which uses its own bearer keys ([Amazon Bedrock BYOK](https://docs.merge.dev/merge-gateway/capabilities/byok/amazon-bedrock), [BYOK overview](https://docs.merge.dev/merge-gateway/capabilities/byok/overview)).

Whether you can authenticate with IAM roles, workload identity or managed identities instead of another long-lived API key. Distinguished from products that only accept static upstream provider credentials.
Fits behind your API gateway API gateway integration
Not documented

n.a. (not documented). Gateway is a standalone hosted endpoint; no integration with an API-gateway platform or service mesh (Kong, APISIX, Envoy, Istio) appears on the fetched pages ([Get started](https://docs.merge.dev/merge-gateway/get-started), [API overview](https://docs.merge.dev/merge-gateway/api-overview)).

Whether AI traffic can go through a gateway you already run, and who documented that — the product’s vendor or the gateway’s.
MCP support MCP surface shape
Not published

n.a. Gateway documents no MCP support of any kind — no MCP gateway, no hosted MCP server, no MCP tools in the API. The string "mcp" does not appear in any of the 403 Gateway documentation URLs, and the Claude Code integration ships skills rather than an MCP server ([Install skills](https://docs.merge.dev/merge-gateway/install-skills), [Coding agents](https://docs.merge.dev/merge-gateway/coding-agents/overview), [API overview](https://docs.merge.dev/merge-gateway/api-overview)). MCP is Merge's separate Agent Handler product, not Gateway ([Introducing Merge Gateway](https://www.merge.dev/blog/gateway-announcement)).

Which kind of MCP support this is: a gateway that governs many MCP servers, a hosted MCP server you connect to, MCP tools accepted inside the completion API, or client tooling. These are different products behind one acronym.
Needs your own provider key Upstream provider key required
Optional

Optional. Managed credentials are the default path and the quickstart needs no provider key at all; BYOK is an upgrade for Pro and Enterprise organisations, with a fallback-to-Merge toggle and a strict `BYOK_ONLY` mode for embedded tenants ([BYOK overview](https://docs.merge.dev/merge-gateway/capabilities/byok/overview), [Get started](https://docs.merge.dev/merge-gateway/get-started)).

Whether an upstream provider account and key must exist before your first call works. A prerequisite rather than a step, and it can differ between the hosted and self-hosted forms of the same product.
Gate before models work Model access gate
No gate

None: the Free plan promises "Access all major LLM models" and the quickstart routes to any catalogue model by ID with no enablement, approval, waitlist or quota step ([Gateway pricing](https://www.merge.dev/pricing/gateway), [Get started](https://docs.merge.dev/merge-gateway/get-started)). Gating is something you impose, not something Merge imposes: ZDR, geo rules and per-customer blocklists all narrow the reachable set ([Per-customer restrictions](https://docs.merge.dev/merge-gateway/security/per-customer-restrictions)).

Whether an enablement click, a quota grant, a paid tier or an approval form stands between a valid key and a working model call.
Official client languages First-party client SDK languages
Python, TypeScript

First-party: Python `merge-gateway-python` (`from merge_gateway import MergeGateway`) and JavaScript/TypeScript `merge-gateway-sdk` (`import { MergeGateway } from "merge-gateway-sdk"`), plus the AI SDK provider `merge-gateway-ai-sdk-provider`. Third-party clients are supported by base URL: OpenAI, Anthropic and LangChain are named in a table of base URLs ([Get started](https://docs.merge.dev/merge-gateway/get-started)). No Go, Java, Ruby, .NET or Rust SDK is published.

Languages with a first-party client library. An empty list can still mean the product is usable from any language via an OpenAI-compatible SDK.

Additional charges

These are the fees that do not appear on a per-token price list, and they are where estimates usually go wrong.

  • Pro usage fee on model spend 5% of LLM cost
  • Pro monthly credit grant $10/month, issued on the 1st, expires monthly, no roll-over, no cash value
  • Launch offer Sign up before 30 June for zero fees for 12 months (ACH: unlimited zero-fee spend; card: zero fees on the first ~$40,000 of annual spend)

How pricing actually works

n.a. There is no self-hosted artifact to cost: no Docker image, Helm chart, binary, npm package or install command appears anywhere in the 403-URL Gateway documentation sitemap, and the only self-hosting language is a single Enterprise sales bullet, "Deploy to your own VPC or on-prem environment" ([Gateway pricing](https://www.merge.dev/pricing/gateway), [Get started](https://docs.merge.dev/merge-gateway/get-started)).

Back to top ↑

Common questions

Answered from the fields above, so these move when the catalog moves. Every figure quoted here appears in the specification with its source.

Does Merge Gateway charge a markup on model prices?

Merge Gateway adds 5% to model prices. Other charges on the page: pro usage fee on model spend (5% of LLM cost), pro monthly credit grant ($10/month, issued on the 1st, expires monthly, no roll-over, no cash value) and launch offer (Sign up before 30 June for zero fees for 12 months (ACH: unlimited zero-fee spend; card: zero fees on the first ~$40,000 of annual spend)).

Can Merge Gateway be self-hosted?

No. Merge Gateway is available only as a service the vendor operates; there is no self-hosted build. The licence is Proprietary. Prompts therefore leave your network and reach the vendor, which makes its retention and residency terms the control that matters here rather than deployment.

Is Merge Gateway SOC 2 audited, and will it sign a HIPAA BAA?

Merge Gateway publishes neither a SOC 2 report nor a HIPAA business associate agreement. Each of these is linked to the vendor's own page in the compliance section below. Neither absence means a refusal: both are things a vendor either publishes or does not, and smaller products often hold the certification without advertising it.

Does Merge Gateway retain your prompts?

Merge Gateway publishes a zero-data-retention position. Whether prompt and response bodies are logged is configurable. Logging can be turned off. Retention, logging and training on customer data are three separate questions, and a vendor can answer one of them without answering the others.

Can you use your own provider keys with Merge Gateway?

Yes. Merge Gateway can route through your own accounts with the underlying model providers, so inference is billed to you directly. BYOK is gated to the Pro and Enterprise plans; one credential per provider per organisation, stored encrypted at rest, with an optional "Fall back to Merge managed credentials" toggle. No separate BYOK per-token or platform fee is published on the pricing page ([BYOK overview](https://docs.merge.dev/merge-gateway/capabilities/byok/overview), [Gateway pricing](https://www.merge.dev/pricing/gateway)).

How many models does Merge Gateway support?

Merge Gateway publishes no aggregate total; counting its catalog gives 3 models, drawn from 4–24 upstream providers. Count of entries returned by the models API on 2026-09-26. Includes every entry exposed by that endpoint; not a count of unique base models. The figure on this page is dated and carries its source.

Back to top ↑

What has changed here

  1. Models available Models available 274 3 source ↗
  2. catalog entry catalog entry Not published Added to the catalog source ↗
See this in the full changelog Back to top ↑

Read the head-to-head

These pairs have a written verdict, not just a table.

Usually weighed against

Follow updates about Merge Gateway