Hugging Face Inference Providers Managed marketplace
Hugging Face Inference Providers is a managed marketplace: an OpenAI-compatible API in front of 136 models from 14–18 providers. It charges no token markup or per-seat fee. It cannot be self-hosted. Zero data retention is published; a HIPAA BAA is not. You can point it at your own provider accounts. Beyond chat it also serves embeddings, image generation and audio. It handles failover and request logging.
· 88 dated entries · 88 source references
Access at a glance
Whether this product can work for you at all, before features matter: who pays the model bill, where it can run, and which of your existing API calls keep working. Every value is the vendor’s own claim, linked to the page it came from.
A routing/marketplace product rather than a model host: "a unified proxy layer that sits between your application and multiple AI providers", with "Use a single Hugging Face token for all providers" and centralised billing. It has no standalone marketing site - the docs index is its home page. Scope note: this entry covers only that router, not the Hugging Face Hub repository platform and not Inference Endpoints, which is a dedicated single-model hosting product (Inference Providers, Hub Integration).
Who pays the model bill
Your keys or their creditsYou can start on their credits and move to your own provider accounts later.
"Hugging Face Routed Requests" is the default (one HF token, HF bills you), and "Custom Provider Key" lets you store your own provider key so calls are billed by the provider while still being routed by Hugging Face (Pricing and Billing, 2026-09-03).
Merchant of record: Differs by mode. Routed Requests: Hugging Face is the merchant - "billing is managed directly by Hugging Face" and no provider account is required. Custom Provider Key: the upstream provider bills you ("Billed by: Provider"). Organisations can also route the bill through AWS Marketplace by linking the org to an AWS account (Pricing and Billing, Hub billing).
Key handling: Provider keys are stored per user (and per organisation) in Hub account settings; routed access uses a Hub user access token, and a fine-grained token can be scoped to the "Make calls to Inference Providers" permission. Organisation billing is targeted with the X-HF-Bill-To header, and admins can disable specific providers org-wide. No encryption-at-rest statement for stored provider keys was found (Hub Integration, Pricing and Billing, User access tokens).
Where it can run
1 of 5 shapes documented- Vendor-hosted
- Self-host
- Your VPC
- On-premise
- Air-gapped
Dimmed shapes are not documented by the vendor, which is not the same as unsupported.
Hosted SaaS only. Every documented path resolves to router.huggingface.co; there is no self-host, VPC or hybrid option in the docs (Inference Providers, Integrations Overview).
There is nothing to deploy: you point an OpenAI client at https://router.huggingface.co/v1 with a Hugging Face token, or install huggingface_hub / @huggingface/inference for the non-chat tasks (Your First Inference Provider Call, Inference Providers).
API surfaces your code can keep using
4 of 7 documented, 1 partial- OpenAI chat
POST /v1/chat/completionsYesPOST https://router.huggingface.co/v1/chat/completions, described as "a drop-in compatible endpoint", withmessages, tools and streaming (Inference Providers, Chat Completion, 2026-09-03). - Anthropic messages
POST /v1/messagesNot documentedNo
/v1/messagesor Anthropic-format surface appears on the docs index, the chat-completion task page, the Responses API guide or the API reference index (Inference Providers, Chat Completion, Responses API (beta), API Reference). The catalogue is open-weight-model focused and Anthropic is not one of the 18 partners. - OpenAI Responses
POST /v1/responsesYes *Yes, marked beta: "Use your existing OpenAI SDKs to access features like multi-provider routing, event streaming, structured outputs, and Remote MCP tools" at the same
https://router.huggingface.co/v1base URL, including reasoning-effort and image inputs (Responses API (beta), 2026-09-03). - Embeddings
POST /v1/embeddingsYes *Yes, but not on the OpenAI surface: embeddings are the
feature-extractiontask, documented viaInferenceClient.feature_extractionand served by HF Inference, Together and Scaleway; there is no/v1/embeddingspath in the docs (Feature Extraction, Inference Providers, 2026-09-03). - Images
POST /v1/images/generationsYes *Yes:
text-to-imageis a documented task with its own reference page and provider column (fal-ai, Replicate, Together, Nscale, Novita and others), called throughInferenceClient.text_to_image; image generation is not exposed on the OpenAI-compatible path (Text to Image, Inference Providers). - Audio
POST /v1/audio/*Partly *Partial: speech-to-text is documented (
automatic-speech-recognition, plusaudio-classification) and the partners table has a "Speech to text" column, but no text-to-speech task page exists in the API reference index (Automatic Speech Recognition, API Reference, Inference Providers, 2026-09-03). - Batch jobs
POST /v1/batchesNot documentedNo batch or async bulk endpoint on any page fetched: no
/v1/batches, no job-submission flow (API Reference, Inference Providers, Responses API (beta), Hub API).
An asterisk marks a qualified verdict: support that is indirect (SDK compatibility or provider passthrough rather than a native endpoint), or a gap that is narrower or wider than the label suggests. Read the note before porting.
Three surfaces on one base URL plus a client library. https://router.huggingface.co/v1 serves OpenAI Chat Completions - "a drop-in compatible endpoint that handles all provider selection automatically on the server side" - and a beta Responses API; the twenty-odd classic tasks (feature extraction, text-to-image, text-to-video, ASR, classification, translation) are reached through InferenceClient in huggingface_hub/@huggingface/inference rather than the OpenAI path, which the docs say is "available for chat completion tasks only". The launch blog also documents provider-native pass-through at https://router.huggingface.co/{:provider} (Inference Providers, API Reference, launch blog).
How much it reaches
Models: Counted from the vendor’s own published list; no aggregate total is published.
Count of entries returned by the models API on 2026-09-23. Includes every entry exposed by that endpoint; not a count of unique base models.
18 counted from the partners table on the docs index (Baseten, Cerebras, Cohere, DeepInfra, Fal AI, Featherless AI, Fireworks, Groq, HF Inference, Novita, Nscale, OVHcloud AI Endpoints, Public AI, Replicate, Scaleway, Together, WaveSpeedAI, Z.ai), of which 15 carry the "Chat completion (LLM)" mark (Inference Providers, 2026-09-03). The low end is 14: that is how many distinct providers appeared in the live public router catalogue when it was paged on 2026-09-03 (router /v1/models) - fal-ai, replicate, wavespeed and hyperbolic serve non-chat modalities and so do not appear there.
Whose models: Almost entirely third-party: Hugging Face describes itself as "a unified proxy layer that sits between your application and multiple AI providers". The one exception is hf-inference, Hugging Face's own backend, which appears as one of the 18 partners inside its own router and which "as of July 2025 focuses mostly on CPU inference" (Inference Providers, HF Inference, 2026-09-03).
Your own endpoints: n.a. for end users. There is no way to register your own base URL, self-hosted model or custom provider slug; becoming a provider is a partner-onboarding process Hugging Face runs, requiring a mapping API, a billing API and passing HF's validation suite (Register as an Inference Provider, Pricing and Billing, Hub Integration).
How it behaves in production
What happens when an upstream model is slow, wrong, or down — and what you can see and stop while it happens. Reliability features are recorded as where you configure them, not whether the vendor lists them, because almost every product here lists all of them.
1 of 6 reachable from code nothing documented on the request path no documented export
When something goes wrong
Each row says where the knob is, not whether the feature is on the marketing page. A control you can only reach by hand in someone else’s dashboard cannot be reviewed or version-controlled.
- Request timeout Not documented
The vendor does not document this, so any behaviour you observe today is unversioned and may change.
n.a. - no user-settable request timeout. The only published timing thresholds are internal admission rules Hugging Face enforces on providers: models "must respond in under 5 seconds" time-to-first-token for conversational and text tasks and under 30 seconds for other tasks, checked by HF's own validation system (Register as an Inference Provider, Inference Providers, Chat Completion).
- Retries Not documented
No retry count, backoff strategy or retry header is published. The documented failure behaviour is provider substitution, not a retry counter: with
provider="auto""requests are automatically routed to alternative providers if the primary provider is flagged as unavailable by our validation system". Default retry count: n.a. Backoff: n.a. (Inference Providers, Register as an Inference Provider). - Fallback to another model Per request
Your application decides per call, so one noisy endpoint can have its own timeout without a redeploy.
Selected per request through the model id.
provider="auto"(the default, equivalent to the:fastestsuffix) falls through to alternative providers when the primary is flagged unavailable;:preferredwalks the user's configured provider order from account settings;:cheapestpicks the lowest price per output token;:groq-style suffixes pin one provider and disable fallback. Order comes from HF policy or the user's settings list, not from a per-request array (Inference Providers, Hub Integration). - Load balancing Not documented
No weighted or proportional load balancing is published. What exists is single-pick policy routing (
:fastest,:cheapest,:preferred, explicit provider pin); weights are not user-settable and no traffic-splitting mechanism is described. The provider ordering shown in settings defaults to "total requests routed by HF over the last 7 days", which is a display order rather than a balancing policy (Inference Providers, Register as an Inference Provider). - Upstream health tracking Fixed, cannot change
The behaviour is fixed by the vendor. Predictable, but you cannot tune it for your workload.
Health checking exists but is entirely HF-operated and has no user knobs: every mapped model is tested every 6 hours, a failing provider is "temporarily removed from the list of active providers" and then retested hourly until it passes, and the suite enforces latency limits plus tool-calling and structured-output checks for LLMs (Register as an Inference Provider, 2026-09-03).
- Cross-region failover Not documented
n.a. - no region selection or cross-region failover for routed inference is documented. The router is a single
router.huggingface.coendpoint, and the Enterprise Hub "Storage Regions" feature governs repository data, not inference (Inference Providers, Inference Providers security, Enterprise Hub).
Fallback chain: Ordered list — Try A, then B, then C. Simple and predictable, but every failover is all-or-nothing.
The reliability surface is deliberately thin and almost all of it is operated by Hugging Face rather than configured by the caller: routing policy is a suffix on the model id, failover is automatic, health checking runs on a 6-hour cycle, and there is no timeout, retry, region or weight setting anywhere in the docs. Configuration-as-code does not exist for the same reason - the entire control surface is per-request strings plus user/organisation settings pages (Inference Providers, Register as an Inference Provider, Hub Integration).
How fast the hop is
Undisclosed vendor serviceThe vendor does not disclose what the request path runs on, so no overhead floor can be inferred at all.
vendor_saas: a closed hosted proxy. Hugging Face states "The Inference Providers API acts as a unified proxy layer that sits between your application and multiple AI providers" and "your requests go through Hugging Face's proxy infrastructure". No runtime, language or source for the router is published; the open-source artefacts are the client libraries only (Inference Providers, Inference Providers security).
No Docker image, Helm chart, binary or Terraform artefact for a customer-run router appears anywhere in the documentation: the docs tree covers index, pricing, security, hub integration, hub API, guides, integrations, providers and tasks, with no deployment section (Inference Providers, Pricing and Billing, Integrations Overview). huggingface_hub and @huggingface/inference are clients that call router.huggingface.co, not a data plane.
Streaming caveats: Supported on both surfaces: the chat-completion page lists "Streaming the output" among what the API supports, and the Responses API guide documents server-sent event streaming with typed events. No documented caveat about stream cancellation or billing on aborted streams was found (Chat Completion, Responses API (beta)).
Published figures, grouped by what each one measured. Figures in different groups are different quantities and cannot be compared with one another — nor, in most cases, with another vendor’s figure in the same group.
Sustained capacity
Requests or queries per second sustained on the stated hardware.
- under 5 s time-to-first-token admission threshold Vendor-published
Not a measured performance claim: this is the eligibility bar Hugging Face's validation system enforces on partner providers for conversational and text models (under 30 s for other tasks), tested every 6 hours. No percentile, payload, region or hardware is stated, and no gateway-overhead figure is published anywhere.
Source
Nothing to compare: Hugging Face publishes no benchmark of its own routing layer and makes no comparative claim against another gateway on any page fetched (Inference Providers, Pricing and Billing).
What it will stop
nothing documented on the request pathNo request-path policy controls are documented. That is not a fault in a product built purely for routing — but it means anything you need blocked has to be blocked before the call reaches here.
What you can see
No documented exportToken counts, latency and model names are stored, but not the text itself.
n.a. - no opt-out switch is documented, because there is no content logging to opt out of; the 30-day debugging logs are described without a customer control (Inference Providers security, Pricing and Billing).
n.a. - no OpenTelemetry, OTLP, trace export or request-trace view is documented. The observability surface is billing-shaped: a usage dashboard in account settings and a member-usage graph in organisation settings (Hub Integration, Hub billing, Inference Providers).
metadata_only by vendor statement - request and response bodies are explicitly not stored, and debugging logs are said to contain "no user data or tokens" (Inference Providers security).
Where telemetry can go
No documented export. Whatever this product records stays in its own interface, so it cannot become part of the monitoring you already run.
None documented. No OTLP endpoint, no log drain, no S3/Kafka/webhook destination appears on any fetched page (Hub Integration, Hub billing, Inference Providers security).
n.a. - no feedback, rating or annotation API. The Responses API guide documents generation features only, and there is no scoring surface in the docs (Responses API (beta), Inference Providers).
No gateway-side evaluation or scoring. Hugging Face documents evaluation only as an external harness pointed at the router - a guide for running Inspect AI against Inference Providers models, and Inspect is listed under "Evaluation Frameworks" in the integrations table (Evaluation with Inspect AI, Integrations Overview).
Depends on the vendor’s SaaS: Yes by construction - the product is hosted-only, so the usage dashboard, per-provider breakdown and organisation usage graph exist solely inside Hugging Face account settings (Hub Integration, Hub billing).
Retention: 30 days for debugging logs; no prompt or completion retention at any point (Inference Providers security).
Whether it fits how you work
How much work stands between you and a first call, how different that is from running it in production, and whether it slots into the stack you already have. Recorded as the shape of the work rather than a number of minutes — how long it takes you depends on which accounts and quota you already hold, which no comparison can know.
to try: base url swap to run: base url swap fits 4 of 10 common stacks
Getting to a first call
3 numbered stepsYour existing OpenAI-compatible client keeps working. You change a base URL and a key, and nothing else in your code moves.
Read off: the vendor’s own quickstart — 3 numbered steps.
Three numbered steps in "Your First Inference Provider Call" (find a model on the Hub, try the widget, copy the code snippet), but the widget step is optional and the effective code path is shorter - set HF_TOKEN, install a client, call the router (Your First Inference Provider Call).
Before step one
- Your own provider key Not needed
You can make a first call with only this product’s key. No upstream provider account needed.
The default "Hugging Face Routed Requests" mode needs only a Hugging Face token - "No separate provider account is required" - and every quickstart snippet uses
HF_TOKENalone (Pricing and Billing, Your First Inference Provider Call). - Payment method No card needed to start
Not required for a first call: signed-in free accounts get $0.10 of included monthly inference credits. It becomes required for sustained use - free-tier users must purchase credits to go beyond the included amount, purchases are credit-card only, and Indian-issued cards are not accepted (Pricing and Billing, Hub billing).
- Gate before models answer Deployment and quota first
You deploy a model and hold quota for it before any call works, and quota increases are a request.
No approval or enablement step per model - any model with a warm provider mapping is callable immediately - but usage is credit-gated: each tier carries a fixed monthly credit allowance ($0.10 free, $2.00 PRO, $2.00 per Team/Enterprise seat) and free accounts must buy credits to continue past it. Organisation admins can additionally disable specific providers (Pricing and Billing, Hub Integration).
Everything you need first: Same as production: an HF account, a token with the Inference Providers permission, and credits; the docs' snippets need only pip install openai or pip install huggingface_hub (Responses API (beta), Your First Inference Provider Call).
The vendor’s own time claim: Vendor claim, verbatim: "This guide will show you how to use a state-of-the-art model in under five minutes, with no infrastructure setup required." (Your First Inference Provider Call) Quoted, not verified. Marketing time claims assume every account and approval is already in place.
Running it in production
The same scale applied to the path the vendor recommends for production traffic. Kept separate from the quickstart because for several products here the two are barely related pieces of work.
Your existing OpenAI-compatible client keeps working. You change a base URL and a key, and nothing else in your code moves.
What production needs: A Hugging Face account, a fine-grained token with the "Make calls to Inference Providers" permission stored in HF_TOKEN, and remaining inference credits. Nothing else - no cluster, no provider accounts, no card for the first calls (Responses API (beta), Pricing and Billing).
Can you run it yourself
This runs on the vendor’s infrastructure only.
How it fits your stack
4 of 10Each row is a thing you might already run. “With a caveat” means it works but not the way the vendor’s marketing implies — read the reason, because that is usually where the surprise lives.
- With a caveat The OpenAI SDK OpenAI-compatible paths are documented, but not as a general drop-in.
- Fits The Vercel AI SDK @ai-sdk/huggingface
- No Cloudflare Workers No Workers guidance published.
- No Kubernetes No Kubernetes deployment published.
- No Terraform or OpenTofu Nothing published for Terraform.
- Fits An existing API gateway This is that gateway — AI traffic becomes a plugin, not a new hop.
- No Cloud IAM I already run No identity integration published.
- Fits LangChain or LlamaIndex LangChain, LlamaIndex, CrewAI, Haystack, PydanticAI, smolagents, LiteLLM, fast-agent, Inspect
- With a caveat MCP servers to govern MCP tools in the API — governs nothing on your side.
- Fits Nothing — plain Node or Python Change one base URL.
Reading this the other way round — pick what you already run and see every product scored against it.
The integration surfaces behind those answers
- Vercel AI SDK Official provider package
Install the package, swap the model factory, done. Maintained by a party with a stake in it.
Named:
@ai-sdk/huggingfaceDocumented by: Vercel AI SDK (@ai-sdk/huggingface)
An official AI SDK provider package exists,
@ai-sdk/huggingface, whose default base URL ishttps://router.huggingface.co/v1and which exposes.responses()and.languageModel()factories plus a model-capability table for tool use and image input. It is documented by Vercel, not by Hugging Face - the AI SDK does not appear in HF's own integrations table (AI SDK Hugging Face provider, Integrations Overview, 2026-09-03). - Cloudflare Workers Not documented
No Workers guidance either way. If you are edge-first, verify fetch-only compatibility yourself.
n.a. - no Cloudflare Workers, edge-runtime or Workers-AI binding guidance on any fetched page; the integrations catalogue is coding agents, LLM frameworks and IDEs (Integrations Overview, Building Your First AI App).
- Kubernetes Not documented
No Kubernetes story published.
n.a. - nothing to run in a cluster, and no manifests, Helm chart or operator are mentioned; the service is a hosted endpoint (Inference Providers, Integrations Overview).
- Terraform Not documented
No Terraform surface published. Configuration is API or dashboard work.
n.a. - no Terraform provider, module or registry reference appears in the documentation, and there is no resource to provision beyond a token and credits (Inference Providers, Pricing and Billing).
- Existing API gateway It is the API gateway
This product is the gateway. If you already run it for your other APIs, AI traffic becomes a plugin rather than a new hop.
It is the gateway. Worth recording for compare pages: three products already rated in this catalogue - Fireworks AI, Groq and Together AI - appear as upstream partners inside it, and Hugging Face's own
hf-inferencebackend is simultaneously the router and one of its 18 partners (Inference Providers, HF Inference). - Cloud identity Not documented
No identity integration published. Expect API keys in a secret store.
n.a. - authentication is a Hugging Face user access token (fine-grained tokens can be scoped to "Make calls to Inference Providers"). No AWS IAM, GCP service account, Azure AD or OIDC federation path for calling the router is documented (User access tokens, Inference Providers, Pricing and Billing).
- MCP MCP tools in the API
The completion API accepts MCP tool definitions, so the model can call MCP tools. A model capability, not an MCP control plane.
MCP tools are first-class in the beta Responses API: pass a tool of
type: "mcp"withserver_url, optionalallowed_toolsandrequire_approval, and the router calls the remote MCP server as part of the response (Responses API (beta), 2026-09-03). Hugging Face is not itself an MCP gateway for arbitrary traffic.
The integrations table lists LangChain, LlamaIndex, CrewAI, Haystack, PydanticAI, smolagents, LiteLLM, fast-agent and Inspect, each pointing at the framework's own "Official docs" rather than a Hugging Face-authored page; NeMo Data Designer and Vision Agents have HF-written getting-started pages (Integrations Overview, 2026-09-03).
First-party clients are huggingface_hub (Python, pip install huggingface_hub) and @huggingface/inference (JS/TS, npm install @huggingface/inference). The docs additionally use the OpenAI Python and Node SDKs and plain curl against the router, and a hf CLI ships with huggingface_hub for listing warm models (Inference Providers, Your First Inference Provider Call, Responses API (beta)).
Agent features: Agent-shaped work is well covered on the chat path: function/tool calling and structured outputs each have their own guide, the Responses API adds tool events, reasoning effort and remote MCP, and the integrations table includes six terminal coding agents (Pi, OpenCode, Codex, Claude Code, Hermes Agent, Roo Code) plus GitHub Copilot Chat in VS Code. HF's validation suite also tests LLMs for tool calling and structured output before a provider stays active (Function Calling, Structured Outputs, Responses API (beta), Integrations Overview, Register as an Inference Provider).
The path of least resistance runs through the Hub rather than a dashboard: find a model with the "Inference Providers" filter (or hf models ls --warm), test it in the on-page widget or the Playground, then click "View Code Snippets". Two friction points are worth knowing: the credit allowance on a free account is $0.10 a month, and the token must be a fine-grained token carrying the "Make calls to Inference Providers" permission (Your First Inference Provider Call, Hub Integration, Pricing and Billing).
The distribution advantage is the Hub itself: inference widgets on model pages, the Inference Playground, and Data Studio AI text-to-SQL all run on Inference Providers and draw down the same credits, and Hub model search can be filtered by provider (?inference_provider=fireworks-ai). Outside the Hub, the documented integrations are coding agents and frameworks rather than infrastructure - including an HF Copilot Chat extension that puts these models inside VS Code (1.104.0+), whose page repeats the pricing posture as "Transparent pricing: what the provider charges is what you pay" - and Vercel ships an official @ai-sdk/huggingface provider that HF's own integrations table does not mention (Hub Integration, Integrations Overview, VS Code integration, AI SDK Hugging Face provider).
Silence in the docs: 4 of the integration questions on this card have no published answer either way. That is recorded as undocumented, not as a no — but it does mean you would be verifying it yourself.
What it does well
- Zero markup, stated four times across the docs: "Hugging Face charges you the same rates as the provider, with no additional fees"
- Strong default data posture for a gateway: request bodies and responses are not stored, no training on user data, 30-day debug logs only
- Automatic cross-provider failover plus a real HF-run validation system (6-hourly model tests, sub-5s TTFT admission bar, tool-calling and structured-output checks)
- Routing policy as a one-string change: `:fastest`, `:cheapest`, `:preferred` or an explicit provider pin, with `provider="auto"` as the default
- Public unauthenticated model catalogue at /v1/models carrying per-provider price, context length, tool support and status
- Beta Responses API with remote MCP tools, and an official Vercel AI SDK provider (@ai-sdk/huggingface)
- BYOK mode keeps HF routing while the provider bills you, and Hugging Face charges nothing for those calls
Where it falls short
- No guardrails of any kind - no PII redaction, moderation, injection detection or model policy for callers
- No caching, prompt management, tracing or OTel export; observability is a billing dashboard and a usage graph
- No user-settable timeout, retry, weight or region: the entire reliability surface is HF-operated
- No SLA and no Inference Providers component on status.huggingface.co - the status page tracks only a single "huggingface.co" component
- Compliance material is Hub-scoped, not gateway-scoped: SOC 2 Type 2 is asserted for the Hub "which Inference Providers is a feature of", and BAAs/DPAs come via an Enterprise Hub plan
- The company subprocessor list does not name any of the 18 inference partners your prompts are routed to
- Catalogue is open-weight only - no Anthropic, OpenAI or Google frontier models, and no Anthropic Messages surface
- hf-inference is both the router and one of its own 18 partners, and "as of July 2025 focuses mostly on CPU inference"
- OpenAI compatibility covers chat and Responses only; embeddings, images, video and speech need InferenceClient
- Free-tier allowance is $0.10 of credits a month, so any real evaluation requires a card
Choose it when
Teams calling open-weight models who want one token, one bill and genuinely zero markup across 18 serverless providers, with automatic failover they do not have to configure.
Look elsewhere when
You need frontier proprietary models, guardrails, prompt management, caching, tracing, an SLA, or any per-request control over timeouts, retries and regions.
Managed marketplace: One account and one key gets you hundreds of models from dozens of providers. Fastest way to start, widest catalog, least control over the data path.
How hard is it to leave?
Derived from six published facts, not from an opinion. The weights are fixed and the same for every product — see the arithmetic.
| What helps you leave | Points | Source |
|---|---|---|
| Works with standard OpenAI code Switching away is a base-URL change rather than a rewrite of every call site. | 22 /22 | vendor page |
| No vendor-specific SDK required A proprietary client library spreads through your codebase and has to be torn out again. | 10 /10 | — |
| Can use your own provider accounts Your keys and billing relationship stay yours, so removing the gateway does not cut off model access. | 20 /20 | vendor page |
| Can be self-hosted You can run it yourself instead of accepting a pricing or policy change. | 0 /20 | vendor page |
| Configuration lives in version control Routing and budget rules are a file you keep, not dashboard state you would have to rebuild. | 0 /16 | — |
| Your request history can be exported You leave with your own logs instead of abandoning them. | not published | — |
Read the fine print: Getting off is easy, taking your data with you is not. Portability of code is excellent: the OpenAI base URL swaps back to any other OpenAI-compatible endpoint in one line. But no usage, log or cost export is documented - the billing dashboard is only described as something you monitor - and there are no stored prompts to export by design ([Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing), [Hub billing](https://huggingface.co/docs/hub/billing), [Inference Providers security](https://huggingface.co/docs/inference-providers/security)).
1 of the 6 inputs is not published, so the highest reachable score here is 88 rather than 100. That is a gap in the public documentation, not a mark against the product — no points are deducted, they simply cannot be claimed. This measures technical switching cost only. It does not price the engineering time to re-test prompts against a different routing stack.
Hugging Face Inference Providers models & pricing
Browse every imported listing from this provider, with published token rates and a link to compare other providers for the same model. This is provider-reported coverage; an absent listing does not mean unsupported.
Loading model listings…
Official model coverage source ↗ · Model source coverage and limitations
Full specification
Every field we track. Blank fields say "Not published" rather than "No" — we do not infer an absence from silence. Switch to Technical in the header for the precise field names and the low-level details.
Overview
Interpret these fields: Self-hosted vs managed LLM gateways · Who still owns your LLM gateway?
- What kind of product Category
- Managed marketplace
- Marketplaces resell many providers behind one key. Gateways add governance on top. Open-source projects you run yourself. Cloud platforms are hyperscaler surfaces. Inference providers host models on their own hardware.
- Who runs it Deployment model
- Managed only
- Managed means the vendor operates it. Self-host means you run it on your own infrastructure. Both means you can choose.
- Licence Licence
- Proprietary
- Proprietary products cannot be inspected or forked. Open licences such as MIT and Apache-2.0 let you audit, modify, and run the code without permission.
- Who you would be signing with Vendor status
- Independent company
- Whether the product is still an independent company, has been acquired, is a large cloud vendor’s product line, is run by a software foundation, or has been put into maintenance mode. Maintenance mode means bug fixes and security patches only — no new features.
- Last shipped an update Latest release
- Not published
No dated release artefact for this product. The only changelog is Hub-wide, and none of the entries in the fetched window (Granular Feature Access on 2026-08-12, then 2026-08-03, 2026-07-22 MCP Server, 2026-07-21, 2026-07-16) is Inference-Providers-specific; the router is a service with no version number and the docs pages carry no version or date stamp ([Changelog](https://huggingface.co/changelog), [Inference Providers](https://huggingface.co/docs/inference-providers/index)).
- The date of the most recent release or version tag. A product that has not shipped in a year is a different risk from one that shipped last week, regardless of what its marketing site says.
- GitHub stars GitHub stars
- Not published
- A rough proxy for community size on open-source projects. Not a quality measure.
Cost
Interpret these fields: How LLM gateway pricing works · LLM gateway spending limits: stop a runaway agent bill?
- Markup on model prices Token markup
- None
- How much the product adds on top of what the underlying model provider charges. Zero means you pay the same per-token price you would pay the model provider directly.
- Fee to add funds Credit purchase fee
- Not published
- A percentage charged when you top up your balance, separate from token prices. It is easy to miss because it does not appear on the per-token price list.
- Monthly cost per person Seat fee
- None
- A recurring per-user platform charge that applies regardless of how much you use the models.
- Can use your own provider accounts BYOK supported
- Yes
- Bring Your Own Key: you keep direct contracts with OpenAI, Anthropic and others, and the gateway only routes traffic. This preserves negotiated rates and committed-spend discounts.
- Cost of using your own accounts BYOK terms
- Bring-your-own key is called a "Custom Provider Key": requests still traverse the Hugging Face router ("HF routing: Yes") but are "Billed by: Provider", there is no free-tier allowance in that mode, and "Hugging Face won't charge you for the call" - so BYOK carries no Hugging Face fee at all ([Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing), 2026-09-03).
- What the product charges to route traffic through your own provider keys.
- Free tier Free tier
- Included monthly credits: $0.10 for signed-in free accounts, $2.00 on PRO, and $2.00 per seat (pooled) on Team and Enterprise, all described as "subject to change". Free accounts must purchase credits to continue once the included amount is spent ([Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing), 2026-09-03).
- What you can do without paying, useful for evaluation.
- Enterprise plan from Enterprise plan from
- Not published
- Annual entry price for the enterprise tier, where one is published or credibly reported.
- Cost to run it yourself Self-host cost
- No self-hosted deployment exists, so there is no self-host cost line: the service is a hosted proxy ("your requests go through Hugging Face's proxy infrastructure") and the only cost is the pass-through provider rate plus optional Hub subscription credits ([Inference Providers](https://huggingface.co/docs/inference-providers/index), [Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing)).
- What self-hosting actually costs once you account for infrastructure and any paid tier.
- How the vendor makes money Pricing model
- Platform fee plus usage meters
- The shape of the vendor’s bill: does the routing layer charge a percentage on top of tokens, a flat monthly fee, both, neither (because inference is the product), or nothing at all (open source with no paid tier).
- How pricing works, briefly Pricing model detail
- The routing layer itself is $0 and inference is billed at the upstream provider's own rate: "Hugging Face charges you the same rates as the provider, with no additional fees. We just pass through the provider costs directly." Optional Hub subscriptions (PRO $9/month, Team $20/user/month, Enterprise $50/user/month) are not required for access but each carries monthly inference credits, which makes the shape a subscription-plus-passthrough hybrid rather than a pure markup. On the separate question of a credit-purchase fee: none is documented either way - neither the Inference Providers pricing page, the Hub billing FAQ, nor the Hugging Face pricing page states a percentage or minimum on credit purchases, so `credit_fee_pct` and `credit_fee_min_usd` are left unset rather than recorded as zero ([Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing), [Hub billing](https://huggingface.co/docs/hub/billing), [Hugging Face pricing](https://huggingface.co/pricing), 2026-09-03).
- A one-paragraph description that covers the caveats a pricing category cannot: introductory rates, per-feature meters, tier gating, and pricing that resets on a specific date.
- Minimum commitment Minimum commitment
- None stated for pay-as-you-go; all users, including free accounts, can purchase credits on demand and enable automatic recharge. Enterprise Hub is quoted per seat per month with "yearly commit options" mentioned only on the Enterprise page, which is a Hub-level plan rather than a gateway commitment ([Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing), [Enterprise Hub](https://huggingface.co/enterprise)).
- Whether the vendor requires a minimum contract term, a minimum spend, or a provisioned-capacity purchase to get its published rate.
- Charges that fire after you go over an allowance Overage terms
- Once the included monthly credits are exhausted, usage continues pay-as-you-go at provider rates and is billed to the Hugging Face account; free-tier users must first purchase credits, and automatic recharge can be enabled to avoid interruption ([Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing), 2026-09-03).
- The line items that scale with usage after an included allowance is exhausted — log storage, extra requests, per-feature meters, data export — which is where cost estimates usually go wrong.
- Prompt cache offered Cache mechanism
- Passes provider caching through
- Whether the gateway offers its own response cache, what kind of match it does (exact request, prefix, semantic), or simply passes provider caching through unchanged.
- Discount on cached input Cache-read discount
- Not published
- How much cheaper cached tokens are than fresh input, when the vendor publishes a single figure. Bundled-inference clouds usually price this per model instead of as one number.
- Premium on cache writes Cache-write premium
- Not published
- How much more the first write of a cached prefix costs versus a plain input token. A high write premium and a low hit rate can leave you paying more than you save, so this matters as much as the read discount.
- Who captures the cache saving Cache economics
- No gateway cache. Hugging Face does not document an exact-match, prefix or semantic cache of its own on the index, pricing, chat-completion or Responses pages. Provider-side caching passes through economically: the provider billing spec tells partners that "`0` is a valid cost, for instance when you serve a cached response for free", so a cached upstream response can reach the caller at zero cost. No cached-token discount or premium is published by Hugging Face ([Register as an Inference Provider](https://huggingface.co/docs/inference-providers/register-as-a-provider), [Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing), 2026-09-03).
- Whether the customer keeps the full saving from caching or the vendor captures part of it — and any conditions attached (write premium, storage fees, best-effort hits).
- What you can split spend by Cost attribution
- Per-provider and per-feature usage appears in the account billing dashboard, organisation settings add "a graph of your team member's usage over time", and individual calls can be attributed to an organisation or resource group with the `X-HF-Bill-To` header. Per-key, per-tag and per-customer attribution are not documented ([Hub Integration](https://huggingface.co/docs/inference-providers/hub-integration), [Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing)).
- The dimensions the vendor documents for splitting spend — per key, per user, per team, per tag, per customer. Matters if you need to chargeback internally or bill an end customer.
- How you get cost data out Cost export
- No CSV, API, webhook or warehouse cost export is documented. The billing dashboard is described only as something you "monitor" ([Hub billing](https://huggingface.co/docs/hub/billing)); the Inference Providers pricing and hub-integration pages describe usage graphs but no export ([Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing), [Hub Integration](https://huggingface.co/docs/inference-providers/hub-integration)).
- The mechanisms the vendor publishes for exporting cost and usage data: CSV, an API, webhooks, S3, a data warehouse, or nothing at all. Any per-unit price is included.
- Who pays the model bill BYOK mode
- Your keys or their credits
- Whether you bring your own provider accounts (BYOK), buy inference from this vendor, or can do either. This is the single biggest commercial difference between these products: it decides who holds the contract with the model provider and who carries the spend.
Spend governance in detail
Seven signals matter when a bill starts to hurt: who can spend, how much, on what, and who gets paged when it goes wrong. Everything below is drawn from the vendor’s own pricing and docs pages — how we read these.
- Virtual or scoped keys Not published
Keys that carry their own budget and rate-limit policy, so an intern experiment cannot spend against a production budget.
Not documented as a gateway feature. Access uses Hub user access tokens, and a fine-grained token can be scoped to the "Make calls to Inference Providers" permission, but there is no per-key spend object described (https://huggingface.co/docs/hub/security-tokens).
- Budget caps per key Not published
A dollar or token ceiling attached to an individual key. Where enforcement is soft, one over-limit request still completes before the block kicks in.
Not documented. No per-token or per-key budget appears on the pricing or hub-integration pages.
- Budget caps per team or workspace Yes — enterprise
A ceiling applied at a higher scope than one key — a team, a workspace, a customer, or an entire environment.
"Inference Providers organization billing - Centralized usage, analytics, and spending limits" is listed as a Team/Enterprise row on the Enterprise Hub feature table, and the Enterprise page repeats "manage spending limits" for Inference Providers specifically (https://huggingface.co/enterprise).
- Rate limiting as a cost control Not published
Configurable request-per-time-window caps. Platform-set rate limits do not count as spend controls; user-configurable ones do.
No Inference Providers rate limit is published. The Enterprise page's "API rate limit - Throughput for programmatic access to the Hub" (1,000-6,000 req/5 min by tier) is explicitly a Hub API limit and was not used here (https://huggingface.co/enterprise).
- Model allowlists Yes — enterprise
A policy that constrains which models a key or team can call, keeping expensive frontier models out of the wrong hands.
Organisation admins can disable a set of Inference Providers from organisation settings, which is a provider allowlist rather than a model allowlist (https://huggingface.co/docs/inference-providers/pricing).
- Spend alerts Not published
Alerts fired as spend approaches a threshold. Alerts that only fire after the meter has rolled over are marked as such.
Not documented; automatic recharge is described but not a spend alert (https://huggingface.co/docs/inference-providers/pricing).
- Webhook notifications Not published
Programmatic notifications on spend events, so budget breaches can page an on-call or open a ticket.
Not documented on any fetched billing page.
Enforcement: Enforced before each request
Catalog
Interpret these fields: LLM gateway model counts: what “500+” means
- Models available Models available
- 136
- How many models you can call, shown as a range because several vendors publish different totals on different pages. Vendor-reported either way, so counts are not directly comparable — some count every provider variant of the same model separately.
- Model providers reachable Upstream providers
- 14–18
- How many distinct model providers or labs you can reach, shown as a range where the vendor’s own pages disagree. More providers usually means better redundancy when one has an outage.
- Works with standard OpenAI code OpenAI-compatible API
- Yes
- If yes, you can usually switch to it by changing one base URL, and switch away just as easily. This is the main defence against lock-in.
- OpenAI chat endpoint POST /v1/chat/completions
- Yes
- The endpoint almost every application ports first. "Not documented" means the vendor never states it, which is different from a documented no.
- Anthropic messages endpoint POST /v1/messages
- Not documented
- Whether Anthropic-shaped calls work without rewriting them. Several products support this only as SDK compatibility or provider passthrough rather than a native endpoint — the detail page says which.
- OpenAI Responses endpoint POST /v1/responses
- Yes
- The newer stateful OpenAI surface. Support is much thinner across this market than chat completions.
- Embeddings endpoint POST /v1/embeddings
- Yes
- Whether you can generate vectors through the same gateway, or need a second integration for retrieval workloads.
- Image generation endpoint POST /v1/images/generations
- Yes
- Whether image models are reachable through the same surface as text.
- Audio endpoints POST /v1/audio/*
- Partly
- Speech-to-text and text-to-speech. Frequently the first gap in an otherwise complete gateway.
- Batch jobs endpoint POST /v1/batches
- Not documented
- Asynchronous bulk processing, usually at a discount. Commonly undocumented, and commonly the reason a migration stalls late.
- Needs the vendor’s own code library Requires a vendor-specific SDK
- No
- A proprietary client library spreads through your codebase and has to be torn out again if you leave. “No” is the better answer here, and it means the standard OpenAI client works.
- You can export your request history Logs / usage data export
- Not published
- Whether you can get your own request logs, traces, or usage records back out — through an API, a bulk export, or a download. Decides whether you leave with your history or abandon it.
- Settings can live in version control Declarative config-as-code
- No
- Whether routing, fallback, and budget rules can be declared in a file you keep in Git, rather than existing only as settings clicked into a hosted dashboard.
- Embeddings Embeddings
- Yes
- Text-to-vector models, needed for search and retrieval features.
- Image generation Image generation
- Yes
- Whether image models are reachable through the same interface.
- Speech and audio Speech and audio
- Yes
- Text-to-speech or transcription models through the same interface.
- Video generation Video generation
- Yes
- Whether video models are reachable through the same interface.
- Batch processing Batch processing
- Not published
- Submitting large jobs for cheaper, slower processing. Often 50% off for work that is not time-sensitive.
Routing & reliability
Interpret these fields: How LLM gateway failover actually works · Does your LLM gateway promise any uptime? · Is routing destroying your prompt cache? · Changing models without breaking production
- Uptime it promises in writing Contractual SLA uptime
- Not published
- The uptime percentage in a published, contractual service level agreement. A public status page is not an SLA — it reports what happened, it does not promise anything or pay you back when it breaks.
- Automatic failover Automatic failover
- Yes
- When a model provider goes down or rate-limits you, traffic moves to a backup automatically instead of returning errors to your users.
- Load balancing Load balancing
- Not published
- Spreads requests across several providers or keys to raise your effective rate limit.
- Rule-based routing Conditional routing
- Not published
- Send different requests to different models based on rules — for example a cheap model for free users and a strong model for paying ones.
- Response caching Response caching
- Not published
- Reuses the answer when the exact same request comes in again, which cuts both cost and latency.
- Similar-question caching Semantic cache
- Not published
- Reuses an answer when a new question means roughly the same thing as an earlier one. Saves far more than exact-match caching but can return subtly wrong answers if tuned loosely.
- Where you set the timeout Request timeout surface
- Not documented
n.a. - no user-settable request timeout. The only published timing thresholds are internal admission rules Hugging Face enforces on providers: models "must respond in under 5 seconds" time-to-first-token for conversational and text tasks and under 30 seconds for other tasks, checked by HF's own validation system ([Register as an Inference Provider](https://huggingface.co/docs/inference-providers/register-as-a-provider), [Inference Providers](https://huggingface.co/docs/inference-providers/index), [Chat Completion](https://huggingface.co/docs/inference-providers/tasks/chat-completion)).
- Where a request timeout can be set: per request, in a config file, in the vendor dashboard, or nowhere. Nine of the twenty products documented here do not describe a request timeout at all, so the worst case of a hung upstream call is unknowable from the docs.
- Where you set retries Retry policy surface
- Not documented
No retry count, backoff strategy or retry header is published. The documented failure behaviour is provider substitution, not a retry counter: with `provider="auto"` "requests are automatically routed to alternative providers if the primary provider is flagged as unavailable by our validation system". Default retry count: n.a. Backoff: n.a. ([Inference Providers](https://huggingface.co/docs/inference-providers/index), [Register as an Inference Provider](https://huggingface.co/docs/inference-providers/register-as-a-provider)).
- Where retry count and backoff are configured. Worth knowing alongside billing: a retried streaming call can be charged more than once.
- Where you set fallbacks Fallback surface
- Per request
Selected per request through the model id. `provider="auto"` (the default, equivalent to the `:fastest` suffix) falls through to alternative providers when the primary is flagged unavailable; `:preferred` walks the user's configured provider order from account settings; `:cheapest` picks the lowest price per output token; `:groq`-style suffixes pin one provider and disable fallback. Order comes from HF policy or the user's settings list, not from a per-request array ([Inference Providers](https://huggingface.co/docs/inference-providers/index), [Hub Integration](https://huggingface.co/docs/inference-providers/hub-integration)).
- Where the fallback chain is defined. Almost every product claims fallback; the useful question is whether you can change it from code or only by hand in a dashboard.
- Shape of the fallback chain Fallback shape
- Ordered list
- An ordered list tries targets in sequence; a weighted split sends a percentage of traffic to each, which is what you need to trial a new model on 5% of requests. Weighted splits are much rarer than the marketing implies.
- Upstream health tracking Health checks / circuit breaking
- Fixed, cannot change
Health checking exists but is entirely HF-operated and has no user knobs: every mapped model is tested every 6 hours, a failing provider is "temporarily removed from the list of active providers" and then retested hourly until it passes, and the suite enforces latency limits plus tool-calling and structured-output checks for LLMs ([Register as an Inference Provider](https://huggingface.co/docs/inference-providers/register-as-a-provider), 2026-09-03).
- Whether the product notices a failing upstream and stops sending traffic to it, and whether you can tune the thresholds. This is what turns a provider outage into a blip rather than a sustained error rate, and only four of the twenty expose it.
- Cross-region failover you control Multi-region failover surface
- Not documented
n.a. - no region selection or cross-region failover for routed inference is documented. The router is a single `router.huggingface.co` endpoint, and the Enterprise Hub "Storage Regions" feature governs repository data, not inference ([Inference Providers](https://huggingface.co/docs/inference-providers/index), [Inference Providers security](https://huggingface.co/docs/inference-providers/security), [Enterprise Hub](https://huggingface.co/enterprise)).
- Whether you can define what happens when a region degrades. A vendor running many regions is not the same as a vendor letting you configure failover between them; only three document a user-controlled mechanism.
- Where you set load balancing Load balancing surface
- Not documented
No weighted or proportional load balancing is published. What exists is single-pick policy routing (`:fastest`, `:cheapest`, `:preferred`, explicit provider pin); weights are not user-settable and no traffic-splitting mechanism is described. The provider ordering shown in settings defaults to "total requests routed by HF over the last 7 days", which is a display order rather than a balancing policy ([Inference Providers](https://huggingface.co/docs/inference-providers/index), [Register as an Inference Provider](https://huggingface.co/docs/inference-providers/register-as-a-provider)).
- Where traffic distribution across upstreams or keys is configured.
Operations
Interpret these fields: LLM gateway observability: traces, logs and export · Running coding agents through an LLM gateway · Changing models without breaking production
- Usage dashboards and logs Observability
- Yes
- Built-in visibility into what was sent, what came back, what it cost, and how long it took.
- Spending limits Budget controls
- Yes
- Hard caps that stop spend before it becomes a surprise invoice. The single most valuable control for a small team.
- Rate limits Rate limits
- Not published
- Caps on request volume per key or per user, useful for protecting against abuse and runaway loops.
- Separate keys per team or app Virtual keys
- Not published
- Issue scoped keys with their own budgets and permissions so you can attribute cost and revoke access without rotating everything.
- Prompt versioning Prompt management
- Not published
- Store and version prompts outside your code so they can be changed without a deploy.
- Quality testing Evals
- Not published
- Built-in tooling to score model output against test cases, so you can tell whether a model swap made things better or worse.
- MCP support MCP support
- Yes
- Native support for the Model Context Protocol, the emerging standard for connecting models to external tools.
- What gets logged Logged content
- Metadata only
metadata_only by vendor statement - request and response bodies are explicitly not stored, and debugging logs are said to contain "no user data or tokens" ([Inference Providers security](https://huggingface.co/docs/inference-providers/security)).
- Whether full prompts and responses are stored, only metadata, or nothing. Full-body logging is the most useful debugging feature here and the one most likely to need a conversation with your compliance team.
- You can turn logging off Body-logging opt-out
- Not documented
n.a. - no opt-out switch is documented, because there is no content logging to opt out of; the 30-day debugging logs are described without a customer control ([Inference Providers security](https://huggingface.co/docs/inference-providers/security), [Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing)).
- Whether prompt and response bodies can be suppressed while still keeping usage metrics. Nineteen of the twenty document a way to do this; the mechanisms range from a per-request header to an organisation-wide setting.
- Traces you can take elsewhere Distributed tracing
- Not documented
n.a. - no OpenTelemetry, OTLP, trace export or request-trace view is documented. The observability surface is billing-shaped: a usage dashboard in account settings and a member-usage graph in organisation settings ([Hub Integration](https://huggingface.co/docs/inference-providers/hub-integration), [Hub billing](https://huggingface.co/docs/hub/billing), [Inference Providers](https://huggingface.co/docs/inference-providers/index)).
- Whether the product emits OpenTelemetry, a proprietary format, or nothing. OpenTelemetry means the traces land in the tooling you already run instead of only in the vendor’s dashboard.
- Where telemetry can go Export destinations
- Not published
None documented. No OTLP endpoint, no log drain, no S3/Kafka/webhook destination appears on any fetched page ([Hub Integration](https://huggingface.co/docs/inference-providers/hub-integration), [Hub billing](https://huggingface.co/docs/hub/billing), [Inference Providers security](https://huggingface.co/docs/inference-providers/security)).
- Documented sinks for logs and metrics. This is a good proxy for how replaceable the vendor’s own dashboard is: around twenty destinations means you never have to depend on it, while a CSV download means you do.
- Can record user feedback Feedback capture API
- No
n.a. - no feedback, rating or annotation API. The Responses API guide documents generation features only, and there is no scoring surface in the docs ([Responses API (beta)](https://huggingface.co/docs/inference-providers/guides/responses-api), [Inference Providers](https://huggingface.co/docs/inference-providers/index)).
- Whether there is an API to attach a rating or score to a logged request, which is what lets production traffic feed quality work later.
- Scores live traffic Online eval hooks
- No
No gateway-side evaluation or scoring. Hugging Face documents evaluation only as an external harness pointed at the router - a guide for running Inspect AI against Inference Providers models, and Inspect is listed under "Evaluation Frameworks" in the integrations table ([Evaluation with Inspect AI](https://huggingface.co/docs/inference-providers/guides/evaluation-inspect-ai), [Integrations Overview](https://huggingface.co/docs/inference-providers/integrations/index)).
- Whether automated scorers can run against real production requests, rather than only against a test set you assemble yourself.
Performance
- Delay it adds Proxy overhead
- Not published
- Extra time the product itself adds to each request, on top of however long the model takes. Usually irrelevant next to multi-second model latency, but it matters for high-volume or streaming-sensitive workloads.
- Requests per second ceiling Throughput
- Not published
- Published sustained request rate before the product becomes the bottleneck. Only relevant at genuinely high volume.
- What the request path runs on Architecture class
- Undisclosed vendor service
vendor_saas: a closed hosted proxy. Hugging Face states "The Inference Providers API acts as a unified proxy layer that sits between your application and multiple AI providers" and "your requests go through Hugging Face's proxy infrastructure". No runtime, language or source for the router is published; the open-source artefacts are the client libraries only ([Inference Providers](https://huggingface.co/docs/inference-providers/index), [Inference Providers security](https://huggingface.co/docs/inference-providers/security)).
- The comparable way to talk about latency here. An edge worker, a compiled Go or Rust binary, and a Python proxy have different overhead floors no matter which figures each vendor publishes. Products that never disclose their runtime are recorded as undisclosed rather than assumed.
- You can run the request path yourself Self-hostable data plane
- Not documented
No Docker image, Helm chart, binary or Terraform artefact for a customer-run router appears anywhere in the documentation: the docs tree covers index, pricing, security, hub integration, hub API, guides, integrations, providers and tasks, with no deployment section ([Inference Providers](https://huggingface.co/docs/inference-providers/index), [Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing), [Integrations Overview](https://huggingface.co/docs/inference-providers/integrations/index)). `huggingface_hub` and `@huggingface/inference` are clients that call `router.huggingface.co`, not a data plane.
- Whether the component that actually carries your prompts can run on your own infrastructure. Distinct from a vendor offering a self-hosted control plane while still proxying traffic through their network.
- Streaming responses Streaming support
- Yes
Supported on both surfaces: the chat-completion page lists "Streaming the output" among what the API supports, and the Responses API guide documents server-sent event streaming with typed events. No documented caveat about stream cancellation or billing on aborted streams was found ([Chat Completion](https://huggingface.co/docs/inference-providers/tasks/chat-completion), [Responses API (beta)](https://huggingface.co/docs/inference-providers/guides/responses-api)).
- Whether token-by-token streaming is documented. The caveats matter more than the yes: some products cannot cancel a stream without still being billed, and several timeout and fallback mechanisms stop applying once the first token has been sent.
Security & compliance
Interpret these fields: LLM gateway compliance: SOC 2, HIPAA and evidence · How LLM gateway guardrails fail · Which LLM gateways store your prompts? · Do you need an MCP gateway as well?
- Does your prompt reach their servers Prompt transits vendor
- Yes
Yes, in both billing modes: "When using Inference Providers, your requests go through Hugging Face's proxy infrastructure", and even the Custom Provider Key mode is listed as "HF routing: Yes". Hugging Face states it does not store the request body or response while doing so ([Inference Providers security](https://huggingface.co/docs/inference-providers/security), [Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing)).
- Whether the text you send passes through this company’s own infrastructure. If it does, every other promise on this page is a policy commitment rather than a physical impossibility. Self-hosted products can answer no outright.
- What they keep if you change nothing Logging default
- Metadata only, not content
Content is not logged: "We do not store the request body or response when routing requests through Hugging Face. Logs are kept for debugging purposes for up to 30 days, but no user data or tokens are stored." Usage metadata is retained for billing and appears in the account dashboard ([Inference Providers security](https://huggingface.co/docs/inference-providers/security), [Hub Integration](https://huggingface.co/docs/inference-providers/hub-integration), 2026-09-03).
- Defaults matter more than options. A product that stores full prompts and replies unless you find the right header will have stored them by the time you read the docs.
- How long they keep it Default content retention (days)
- 30 days
"Logs are kept for debugging purposes for up to 30 days, but no user data or tokens are stored" - so the 30 days covers operational logs, not prompts or completions, which are not stored at all ([Inference Providers security](https://huggingface.co/docs/inference-providers/security), 2026-09-03).
- Default retention for request content, in days. Zero means nothing is kept. Read the note: several products keep nothing as a rule but make timed exceptions for abuse review or specific models.
- Could they train on your prompts Training on customer data
- No
"Hugging Face does not store any user data for training purposes" on the Inference Providers security page. What each upstream partner does with routed traffic is not covered by that statement and is not summarised anywhere in the Inference Providers docs ([Inference Providers security](https://huggingface.co/docs/inference-providers/security), 2026-09-03).
- Whether the vendor may use your prompts and outputs to train models. “Not published” means we could not find any position, which is not the same as a no — ask for it in writing.
- Where it runs, and what you can pin Region and residency control
- Routed inference has no published region control. The privacy policy states "The Company and its servers are located in the United States", with Hugging Face SAS in Paris as the EU establishment and CNIL as lead authority; the upstream partners include EU-based operators (OVHcloud AI Endpoints, Scaleway) but choosing them is a provider-pinning decision, not a documented residency guarantee ([Privacy Policy](https://huggingface.co/privacy), [Inference Providers](https://huggingface.co/docs/inference-providers/index), [Inference Providers security](https://huggingface.co/docs/inference-providers/security)).
- Which regions are offered and whether you can force processing to stay in one. A global endpoint that silently picks a region is a different compliance story from an endpoint you pin yourself.
- Where safety filters run Guardrail execution location
- No guardrails offered
None. Hugging Face's own statements point the other way - it does not inspect or store request bodies at all ("We do not store the request body or response when routing requests through Hugging Face") - so any moderation must be implemented by the caller or inherited from the upstream provider ([Inference Providers security](https://huggingface.co/docs/inference-providers/security), 2026-09-03).
- A filter that strips personal data only helps if it runs before the data leaves your boundary. If guardrails execute in the vendor’s cloud, the vendor has already received whatever you wanted redacted.
- Who else touches the data Subprocessor list
- https://huggingface.co/privacy
- The published list of third parties the vendor passes your data to. No list means you cannot know the full chain, which most data-protection agreements require you to.
- SOC 2 audited SOC 2 audited
- Not published
- An independent audit of security controls. Enterprise buyers and their procurement teams routinely require it.
- Will sign a HIPAA agreement HIPAA BAA
- Not published
- Required before you may send protected health information through the service. Without a signed BAA, healthcare data is off limits.
- GDPR commitments GDPR commitments
- Not published
- Published data processing terms for handling personal data of people in the EU and UK.
- Can keep data in the EU EU data residency
- Not published
- Requests can be processed inside the EU rather than routed to US infrastructure. Often the deciding constraint for European customers.
- Does not retain your data Zero data retention
- Yes
- Prompts and responses are not stored after the request completes. Sometimes a paid add-on rather than the default.
- Strips personal data PII redaction
- Not published
- Detects and removes identifiers such as names, emails, and card numbers before the request reaches the model provider.
- Content guardrails Content guardrails
- Not published
- Policy checks on inputs and outputs — blocking unsafe content, enforcing formats, or catching prompt-injection attempts.
- Runs fully disconnected Air-gapped deployment
- Not published
- Can be deployed in a network with no internet access, which some regulated and defence environments require.
- Blocks personal data in prompts PII / DLP enforcement
- Not documented
n.a. - no guardrail, moderation, PII, prompt-injection or content-filtering feature appears anywhere in the Inference Providers documentation. Pages checked: [Inference Providers](https://huggingface.co/docs/inference-providers/index), [Inference Providers security](https://huggingface.co/docs/inference-providers/security), [Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing), [Chat Completion](https://huggingface.co/docs/inference-providers/tasks/chat-completion), [Responses API (beta)](https://huggingface.co/docs/inference-providers/guides/responses-api), [Integrations Overview](https://huggingface.co/docs/inference-providers/integrations/index), [Hub Integration](https://huggingface.co/docs/inference-providers/hub-integration).
- Whether personal data detection sits on the request path and can stop the call, merely inspects and forwards it, or is not documented. A control that only reports is a logging feature, not a policy control.
- Blocks prompt injection Injection / jailbreak enforcement
- Not documented
n.a. - no guardrail, moderation, PII, prompt-injection or content-filtering feature appears anywhere in the Inference Providers documentation. Pages checked: [Inference Providers](https://huggingface.co/docs/inference-providers/index), [Inference Providers security](https://huggingface.co/docs/inference-providers/security), [Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing), [Chat Completion](https://huggingface.co/docs/inference-providers/tasks/chat-completion), [Responses API (beta)](https://huggingface.co/docs/inference-providers/guides/responses-api), [Integrations Overview](https://huggingface.co/docs/inference-providers/integrations/index), [Hub Integration](https://huggingface.co/docs/inference-providers/hub-integration).
- Whether injection and jailbreak detection can stop a request. Most products offering this call a partner classifier rather than shipping their own.
- Blocks harmful content Toxicity / moderation enforcement
- Not documented
n.a. - no guardrail, moderation, PII, prompt-injection or content-filtering feature appears anywhere in the Inference Providers documentation. Pages checked: [Inference Providers](https://huggingface.co/docs/inference-providers/index), [Inference Providers security](https://huggingface.co/docs/inference-providers/security), [Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing), [Chat Completion](https://huggingface.co/docs/inference-providers/tasks/chat-completion), [Responses API (beta)](https://huggingface.co/docs/inference-providers/guides/responses-api), [Integrations Overview](https://huggingface.co/docs/inference-providers/integrations/index), [Hub Integration](https://huggingface.co/docs/inference-providers/hub-integration).
- Whether hate, violence, sexual and self-harm categories are checked inline and can stop a request, in either direction.
- Your own policy rules Custom policy hooks
- Not documented
n.a. - no guardrail, moderation, PII, prompt-injection or content-filtering feature appears anywhere in the Inference Providers documentation. Pages checked: [Inference Providers](https://huggingface.co/docs/inference-providers/index), [Inference Providers security](https://huggingface.co/docs/inference-providers/security), [Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing), [Chat Completion](https://huggingface.co/docs/inference-providers/tasks/chat-completion), [Responses API (beta)](https://huggingface.co/docs/inference-providers/guides/responses-api), [Integrations Overview](https://huggingface.co/docs/inference-providers/integrations/index), [Hub Integration](https://huggingface.co/docs/inference-providers/hub-integration).
- Whether you can add your own rule — a regex, a webhook, or your own classifier — rather than choosing from the vendor library.
- Where guardrails run Guardrail execution location
- Not documented
- Whether guardrail evaluation happens inside your infrastructure or on the vendor’s servers. This decides whether the prompt you are trying to protect leaves your network in order to be checked.
- If the guardrail itself fails Guardrail failure mode
- Not documented
n.a. - there are no guardrails, so no fail-open/fail-closed behaviour is described ([Inference Providers security](https://huggingface.co/docs/inference-providers/security), [Inference Providers](https://huggingface.co/docs/inference-providers/index)).
- What happens when the guardrail service times out or errors: does the request proceed unchecked, or is it blocked? This is the worst-documented field in the entire catalogue — only two vendors state it plainly, which means most teams are running a control whose failure behaviour they cannot know.
- Third-party guardrail vendors Guardrail integrations
- Not published
- Named external guardrail services the product can call. A long list means the product is a router for policy engines rather than a policy engine itself — which also means another vendor bill and another hop.
Compliance evidence
Graded by how strong the evidence is, not whether the word appears on the vendor’s website. An audited report and a marketing claim are different things, and only one of them will satisfy your own auditor.
- SOC 2 Claimed, no evidence published Scoped to the Hub, not to the gateway: "The Hugging Face Hub, which Inference Providers is a feature of, is SOC2 Type 2 certified" (https://huggingface.co/docs/inference-providers/security). No report, audit period, auditor or trust portal is named, and no Inference-Providers-specific attestation exists, so the `soc2` field is left unset.
- ISO 27001 Not published No ISO 27001 claim on the Inference Providers security page or the Hub security page (https://huggingface.co/docs/inference-providers/security, https://huggingface.co/docs/hub/security).
- GDPR DPA Claimed, no evidence published Company/Hub level only: "Hugging Face is GDPR compliant" and DPAs are offered "through an Enterprise Plan" on the Hub security page (https://huggingface.co/docs/hub/security). No Inference-Providers DPA or processing appendix is published, so the `gdpr` field is left unset.
- HIPAA BAA Claimed, no evidence published Hub/Enterprise level only: Business Associate Addendums are offered "through an Enterprise Plan" (https://huggingface.co/docs/hub/security). Nothing ties a BAA to routed inference through the 18 third-party providers, so `hipaa_baa` is left unset.
- FedRAMP Not published
- ITAR Not published
Fit & integration
Interpret these fields: How much does an LLM gateway lock you in? · Do you need an MCP gateway as well? · Running coding agents through an LLM gateway · Changing models without breaking production
- Work to try it Evaluation work shape
- Change one base URL
- The shape of the work on the vendor’s own quickstart, from swapping one base URL through to deploying infrastructure. An ordinal class rather than a duration, because elapsed time depends on accounts and quota we cannot see.
- Work to run it Production work shape
- Change one base URL
- The same scale applied to the vendor’s recommended production path. For several products this is much heavier than the quickstart, which is exactly why both are recorded.
- Steps on the quickstart Numbered quickstart steps
- 3
Three numbered steps in "Your First Inference Provider Call" (find a model on the Hub, try the widget, copy the code snippet), but the widget step is optional and the effective code path is shorter - set `HF_TOKEN`, install a client, call the router ([Your First Inference Provider Call](https://huggingface.co/docs/inference-providers/guides/first-api-call)).
- A literal count of numbered steps on the vendor’s quickstart, recorded as evidence beside the work shape. Zero means the page publishes no numbered procedure at all. Large counts usually mean interleaved language tracks rather than more work.
- Can you self-host it today Self-host install documentation
- No self-hosting
- Whether an install command is actually published. Several products advertise self-hosting while publishing no command to start from, which a plain yes/no would hide.
- Works with the OpenAI SDK OpenAI SDK drop-in
- Partly
Partial, and the docs say so: the OpenAI-compatible endpoint is "a drop-in compatible endpoint" but is "available for chat completion tasks only" (plus the beta Responses API). Embeddings, image, video and speech tasks require `InferenceClient` or provider-native paths, so a wholesale OpenAI swap only covers chat workloads ([Inference Providers](https://huggingface.co/docs/inference-providers/index), [API Reference](https://huggingface.co/docs/inference-providers/tasks/index), [Responses API (beta)](https://huggingface.co/docs/inference-providers/guides/responses-api)).
- Whether an existing OpenAI-compatible client can be pointed at it by changing the base URL and key.
- Vercel AI SDK support AI SDK provider package
- Official provider package
An official AI SDK provider package exists, `@ai-sdk/huggingface`, whose default base URL is `https://router.huggingface.co/v1` and which exposes `.responses()` and `.languageModel()` factories plus a model-capability table for tool use and image input. It is documented by Vercel, not by Hugging Face - the AI SDK does not appear in HF's own integrations table ([AI SDK Hugging Face provider](https://ai-sdk.dev/providers/ai-sdk-providers/huggingface), [Integrations Overview](https://huggingface.co/docs/inference-providers/integrations/index), 2026-09-03).
- Whether a first-party AI SDK provider package exists, or only a community package, a documented workaround, or the generic OpenAI provider pointed at a custom base URL.
- Python framework integrations Documented Python frameworks
- LangChain, LlamaIndex, CrewAI, Haystack, PydanticAI, smolagents, LiteLLM, fast-agent, Inspect
The integrations table lists LangChain, LlamaIndex, CrewAI, Haystack, PydanticAI, smolagents, LiteLLM, fast-agent and Inspect, each pointing at the framework's own "Official docs" rather than a Hugging Face-authored page; NeMo Data Designer and Vision Agents have HF-written getting-started pages ([Integrations Overview](https://huggingface.co/docs/inference-providers/integrations/index), 2026-09-03).
- LangChain, LangGraph, LlamaIndex and similar orchestration frameworks with a documented integration.
- Callable from Cloudflare Workers Cloudflare Workers support
- Not documented
n.a. - no Cloudflare Workers, edge-runtime or Workers-AI binding guidance on any fetched page; the integrations catalogue is coding agents, LLM frameworks and IDEs ([Integrations Overview](https://huggingface.co/docs/inference-providers/integrations/index), [Building Your First AI App](https://huggingface.co/docs/inference-providers/guides/building-first-app)).
- Whether the docs show calling this product from your own Worker. Deliberately separated from the several products whose own gateway runs on Workers, which is a fact about their infrastructure and not about your edge compatibility.
- Kubernetes install Helm chart availability
- Not documented
n.a. - nothing to run in a cluster, and no manifests, Helm chart or operator are mentioned; the service is a hosted endpoint ([Inference Providers](https://huggingface.co/docs/inference-providers/index), [Integrations Overview](https://huggingface.co/docs/inference-providers/integrations/index)).
- Whether a named, published Helm chart exists, versus Helm being referenced with no chart named, versus generic cluster documentation that has nothing to do with this product.
- Terraform support Terraform provider or modules
- Not documented
n.a. - no Terraform provider, module or registry reference appears in the documentation, and there is no resource to provision beyond a token and credits ([Inference Providers](https://huggingface.co/docs/inference-providers/index), [Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing)).
- Whether you can declare this in code: an official provider, official modules, resources inside a hyperscaler’s provider, a community provider, or only Terraform code shipped in a repo.
- Reuses your cloud identity Cloud IAM reuse
- Not documented
n.a. - authentication is a Hugging Face user access token (fine-grained tokens can be scoped to "Make calls to Inference Providers"). No AWS IAM, GCP service account, Azure AD or OIDC federation path for calling the router is documented ([User access tokens](https://huggingface.co/docs/hub/security-tokens), [Inference Providers](https://huggingface.co/docs/inference-providers/index), [Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing)).
- Whether you can authenticate with IAM roles, workload identity or managed identities instead of another long-lived API key. Distinguished from products that only accept static upstream provider credentials.
- Fits behind your API gateway API gateway integration
- It is the API gateway
It is the gateway. Worth recording for compare pages: three products already rated in this catalogue - Fireworks AI, Groq and Together AI - appear as upstream partners inside it, and Hugging Face's own `hf-inference` backend is simultaneously the router and one of its 18 partners ([Inference Providers](https://huggingface.co/docs/inference-providers/index), [HF Inference](https://huggingface.co/docs/inference-providers/providers/hf-inference)).
- Whether AI traffic can go through a gateway you already run, and who documented that — the product’s vendor or the gateway’s.
- MCP support MCP surface shape
- MCP tools in the API
MCP tools are first-class in the beta Responses API: pass a tool of `type: "mcp"` with `server_url`, optional `allowed_tools` and `require_approval`, and the router calls the remote MCP server as part of the response ([Responses API (beta)](https://huggingface.co/docs/inference-providers/guides/responses-api), 2026-09-03). Hugging Face is not itself an MCP gateway for arbitrary traffic.
- Which kind of MCP support this is: a gateway that governs many MCP servers, a hosted MCP server you connect to, MCP tools accepted inside the completion API, or client tooling. These are different products behind one acronym.
- Needs your own provider key Upstream provider key required
- Not needed
No: the default "Hugging Face Routed Requests" mode needs only a Hugging Face token - "No separate provider account is required" - and every quickstart snippet uses `HF_TOKEN` alone ([Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing), [Your First Inference Provider Call](https://huggingface.co/docs/inference-providers/guides/first-api-call)).
- Whether an upstream provider account and key must exist before your first call works. A prerequisite rather than a step, and it can differ between the hosted and self-hosted forms of the same product.
- Gate before models work Model access gate
- Deployment and quota first
No approval or enablement step per model - any model with a warm provider mapping is callable immediately - but usage is credit-gated: each tier carries a fixed monthly credit allowance ($0.10 free, $2.00 PRO, $2.00 per Team/Enterprise seat) and free accounts must buy credits to continue past it. Organisation admins can additionally disable specific providers ([Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing), [Hub Integration](https://huggingface.co/docs/inference-providers/hub-integration)).
- Whether an enablement click, a quota grant, a paid tier or an approval form stands between a valid key and a working model call.
- Official client languages First-party client SDK languages
- Python, JavaScript
First-party clients are `huggingface_hub` (Python, `pip install huggingface_hub`) and `@huggingface/inference` (JS/TS, `npm install @huggingface/inference`). The docs additionally use the OpenAI Python and Node SDKs and plain `curl` against the router, and a `hf` CLI ships with `huggingface_hub` for listing warm models ([Inference Providers](https://huggingface.co/docs/inference-providers/index), [Your First Inference Provider Call](https://huggingface.co/docs/inference-providers/guides/first-api-call), [Responses API (beta)](https://huggingface.co/docs/inference-providers/guides/responses-api)).
- Languages with a first-party client library. An empty list can still mean the product is usable from any language via an OpenAI-compatible SDK.
How pricing actually works
No self-hosted deployment exists, so there is no self-host cost line: the service is a hosted proxy ("your requests go through Hugging Face's proxy infrastructure") and the only cost is the pass-through provider rate plus optional Hub subscription credits ([Inference Providers](https://huggingface.co/docs/inference-providers/index), [Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing)).
Common questions
Answered from the fields above, so these move when the catalog moves. Every figure quoted here appears in the specification with its source.
Does Hugging Face Inference Providers charge a markup on model prices?
Hugging Face Inference Providers adds no percentage markup to model prices. The cost estimator on this site itemises these mechanisms against your own volume, because which one is cheapest depends entirely on the numbers you put in. The cost estimator itemises this against your own volume.
Can Hugging Face Inference Providers be self-hosted?
No. Hugging Face Inference Providers is available only as a service the vendor operates; there is no self-hosted build. The licence is Proprietary. Prompts therefore leave your network and reach the vendor, which makes its retention and residency terms the control that matters here rather than deployment.
Is Hugging Face Inference Providers SOC 2 audited, and will it sign a HIPAA BAA?
Hugging Face Inference Providers publishes neither a SOC 2 report nor a HIPAA business associate agreement. Each of these is linked to the vendor's own page in the compliance section below. Neither absence means a refusal: both are things a vendor either publishes or does not, and smaller products often hold the certification without advertising it.
Does Hugging Face Inference Providers retain your prompts?
Hugging Face Inference Providers publishes a zero-data-retention position. Only metadata is logged — not prompt or response bodies. Stated retention is 30 days. It states that it does not train on customer data. Retention, logging and training on customer data are three separate questions, and a vendor can answer one of them without answering the others.
Can you use your own provider keys with Hugging Face Inference Providers?
Yes. Hugging Face Inference Providers can route through your own accounts with the underlying model providers, so inference is billed to you directly. Bring-your-own key is called a "Custom Provider Key": requests still traverse the Hugging Face router ("HF routing: Yes") but are "Billed by: Provider", there is no free-tier allowance in that mode, and "Hugging Face won't charge you for the call" - so BYOK carries no Hugging Face fee at all ([Pricing and Billing](https://huggingface.co/docs/inference-providers/pricing), 2026-09-03).
How many models does Hugging Face Inference Providers support?
Hugging Face Inference Providers publishes no aggregate total; counting its catalog gives 136 models, drawn from 14–18 upstream providers. Count of entries returned by the models API on 2026-09-23. Includes every entry exposed by that endpoint; not a count of unique base models. The figure on this page is dated and carries its source.
Official links
Independent coverage
Third-party analysis, walkthroughs and operator threads. We link criticism as readily as praise. Nothing here is written or published by the vendor, by a competitor listed on this site, or by an SEO content farm — how we vet these.
Written reviews and analysis 2
- Hugging Face Inference API Free Tier Limits & Pricing 2026 The only independent write-up found that separates Inference Providers from Inference Endpoints and the Hub, confirms the pass-through pricing posture, and sets it against OpenRouter's markup; also the source of a credit-figure contradiction noted in the companion file.
- Best LLM API Providers in 2026: We Reviewed 8 Options Useful competitive framing of the three-way Hugging Face billing model and the pass-through rate, but read with the conflict in mind: Fireworks is itself one of the 18 upstream partners inside the product it is rating.
What has changed here
- Models available Models available 143 136 source ↗
- Models available Models available 136 143 source ↗
- Models available Models available 200 136 source ↗
- catalog entry catalog entry Not published Added to the catalog source ↗
Read the head-to-head
These pairs have a written verdict, not just a table.