New API Open source
New API is an open-source LLM gateway: an OpenAI-compatible API in front of 100+ models from 30+ providers. It charges no token markup or per-seat fee. It runs only self-hosted under AGPL-3.0. It does not publish a HIPAA BAA. You can point it at your own provider accounts. Beyond chat it also serves embeddings, image generation and audio. It handles failover, load balancing, guardrails and request logging.
· 44 dated entries · 61 source references
Access at a glance
Whether this product can work for you at all, before features matter: who pays the model bill, where it can run, and which of your existing API calls keep working. Every value is the vendor’s own claim, linked to the page it came from.
A gateway *and* a billing/asset-management application, not just a proxy: the README calls it a "Next-Generation Large Model Gateway and AI Asset Management System" (README.en.md) and the GitHub description is "A unified AI model hub for aggregation & distribution... A centralized gateway for personal and enterprise model management" (GitHub API, 2026-09-02). It is an open-source project developed from One API and is "Fully compatible with the original One API database" (README.en.md).
Who pays the model bill
Your keys onlyYou contract with each model provider directly and hold those accounts. The gateway never resells inference.
A channel *is* one upstream provider API key, and the AUP requires those keys be legally owned or authorised by the deployer (Channel Management, Acceptable Use). There are no platform credits and no first-party model hosting.
Merchant of record: For inference, the upstream provider you hold the key with - New API never sells tokens (Channel Management). If you resell access, *you* are the merchant of record: the software ships EPay (Alipay/WeChat/bank), Stripe, Creem and Waffo top-up integrations plus redemption codes and subscription plans, and the AUP assigns "taxation, invoicing, consumer protection, and payment risk control" to the deployer (System Settings - Detailed, Subscription Plans, Acceptable Use).
Key handling: Two key layers. Upstream provider keys are stored per channel in your database, with Multi-Key channels holding several keys that are polled automatically and skipped when failing (Channel Management); database content is encrypted with CRYPTO_SECRET, which must be identical across nodes sharing Redis (README.en.md, Cluster Deployment). Downstream, users hold New API tokens shown once at creation, each with expiry (-1 for none), remaining quota, unlimited-quota flag, model restrictions, IP allowlist and group (Token Management).
Where it can run
1 of 5 shapes documented- Vendor-hosted
- Self-host
- Your VPC
- On-premise
- Air-gapped
Dimmed shapes are not documented by the vendor, which is not the same as unsupported.
Self-host only. Documented methods are Docker Compose (recommended for production), single-container Docker, 1Panel, BaoTa/aaPanel app store, cluster deployment and local development (Installation & Deployment); there is no vendor-hosted tier on any page fetched - the homepage's own framing is "Open source as our covenant, self-hosted as our ground" (newapi.ai).
Docker image calciumion/new-api:latest (also published to ghcr.io); single-container form is docker run --name new-api -d --restart always -p 3000:3000 -e TZ=Asia/Shanghai -v ./data:/data calciumion/new-api:latest, or clone the repo and docker-compose up -d with the bundled compose file that already wires MySQL and Redis (README.en.md, Docker Compose Deployment). SQLite is used for local single-node data, remote MySQL >= 5.7.8 or PostgreSQL >= 9.6 for multi-node; amd64/arm64 64-bit only (README.en.md).
API surfaces your code can keep using
6 of 7 documented- OpenAI chat
POST /v1/chat/completionsYesPOST /v1/chat/completionsis the first row of the supported-endpoints table and the Python example uses the OpenAI SDK against the instance base URL (Using the API). - Anthropic messages
POST /v1/messagesYes *Yes: a "Native Claude Format" surface accepts "Requests in Anthropic Claude Messages API format", requires the
anthropic-versionheader and acceptsx-api-keyas an alternative to the bearer token, withsystem,max_tokens,tools,tool_choiceandthinkingfields (Native Claude Format);GET /v1/modelsreturns Anthropic-shaped output whenx-api-keyplusanthropic-versionare present (Get Model List). The literal request path is not printed on the fetched pages - the endpoint table on Using the API lists no/v1/messagesrow - so treat the path as undocumented even though the format is supported. - OpenAI Responses
POST /v1/responsesYesPOST /v1/responsesis listed in the endpoint table (Using the API) and "OpenAI Responses API format" is a headline feature (README.en.md). Note the README also lists OpenAI <-> OpenAI Responses *format conversion* as still "in development", so cross-dialect conversion into Responses is not finished (README.en.md). - Embeddings
POST /v1/embeddingsYesPOST /v1/embeddingsis in the endpoint table (Using the API) and Embeddings has its own API-reference section (API Reference). - Images
POST /v1/images/generationsYesPOST /v1/images/generationsandPOST /v1/images/edits(Using the API), plus Midjourney-Proxy(Plus) integration as a separate task-based surface (Features Description). - Audio
POST /v1/audio/*YesPOST /v1/audio/transcriptionsandPOST /v1/audio/speech(Using the API); the API reference additionally carries Audio and Real-time Speech sections and Suno music tasks (API Reference, Features Description). - Batch jobs
POST /v1/batchesNot documentedNo batch or bulk-async inference endpoint appears in the endpoint table (Using the API) or the API reference index, where Files and Fine-tuning are explicitly grouped under "Unimplemented" (API Reference). Asynchronous *task* APIs exist, but they are for video/Midjourney/Suno generation, not OpenAI-style batch (Create Video).
An asterisk marks a qualified verdict: support that is indirect (SDK compatibility or provider passthrough rather than a native endpoint), or a gap that is narrower or wider than the label suggests. Read the note before porting.
One instance exposes three request dialects. The OpenAI surface is a base-URL swap - "Replace OpenAI's base_url with the platform address and use your token as the api_key" - with a published endpoint table: POST /v1/chat/completions, /v1/completions, /v1/embeddings, /v1/images/generations, /v1/images/edits, /v1/audio/transcriptions, /v1/audio/speech, /v1/rerank, /v1/responses, GET /v1/realtime (WebSocket) and GET /v1/models (Using the API). Native Claude Messages format is accepted with an anthropic-version header (Native Claude Format) and Gemini requests are proxied at /v1beta/models/{model}:{action} (Gemini Text Chat). A separate Management API covers channels, tokens, users, groups, logs, payment and statistics (API Reference).
How much it reaches
Single vendor floor ("100+ models"), homepage only (newapi.ai, observed 2026-09-02). There is no public models endpoint to count from: GET /v1/models is served by your own instance (Using the API).
Vendor total, stated as a floor: the homepage says "30+ Model Providers" and "Access 30+ AI providers through a single, unified endpoint" (newapi.ai); the docs repeat "30+ other model services" after naming OpenAI, Anthropic, Google Gemini, DeepSeek, Midjourney and Suno (Project Introduction). No enumerated channel-type list was found on the fetched pages, so 30 is the vendor floor, not a count.
Whose models: All third-party: every model is served by an upstream provider whose credentials you supply - OpenAI, Anthropic, Google Gemini, DeepSeek, Midjourney, Suno "and 30+ other model services" (Project Introduction). QuantumNous claims no model hosting of its own on any fetched page.
Your own endpoints: Yes: every channel takes a Base URL field described as a "Custom endpoint URL for proxies or self-hosted deployments", alongside model mapping and parameter override (Channel Management). The AUP requires that upstream channels be "accounts, API Keys, model services, or enterprise contracts legally owned or authorized by the deployer" (Acceptable Use).
How it behaves in production
What happens when an upstream model is slow, wrong, or down — and what you can see and stop while it happens. Reliability features are recorded as where you configure them, not whether the vendor lists them, because almost every product here lists all of them.
1 of 6 reachable from code 1 of 4 can block no documented export
When something goes wrong
Each row says where the knob is, not whether the feature is on the marketing page. A control you can only reach by hand in someone else’s dashboard cannot be reviewed or version-controlled.
- Request timeout In config
Changing it means editing configuration and shipping it, so behaviour is uniform across traffic until you redeploy.
Environment variables only:
RELAY_TIMEOUTdefaults to0, meaning no timeout,STREAMING_TIMEOUTis 300 s between chunks,TASK_TIMEOUT_MINUTESis 1440 for async tasks, and connection pooling is tuned withRELAY_MAX_IDLE_CONNS500 /RELAY_MAX_IDLE_CONNS_PER_HOST100. The docs explicitly warn that setting the relay timeout too low can produce requests the upstream charges for but New API never bills (Environment Variables). - Retries Dashboard only
Only reachable by hand in the vendor UI, so it cannot be reviewed, version-controlled, or changed from code.
"Automatic retry on failure" is a listed routing feature and the retry count is a console setting at
Settings -> Operation Settings -> General Settings -> Failure Retry Count; the default value and any backoff strategy are not published (README.en.md). Default retry count:n.a.Backoff:n.a. - Fallback to another model Dashboard only
Fallback is channel-level and configured per channel in the console: Priority ("higher value = higher selection priority", default 0) selects the tier and Weight is the "random weight among same-priority channels" (default 0), so selection within a tier is weighted random while tiers are ordered (Channel Management); the README's headline phrasing is "Channel weighted random" (README.en.md). Token groups can be set to
auto, which "automatically selects an available group in priority order - useful for cross-group failover scenarios" (Group Management). Multi-Key channels additionally poll keys in Round Robin or Weighted Random order and skip a failed key until it recovers (Channel Management). - Load balancing Dashboard only
Weights and priorities are user-settable per channel in
/console/channel, with model mapping and per-channel parameter override alongside (Channel Management); there is no routing config file - the channel table lives in the database. Cluster-level balancing across nodes is left to an external load balancer such as Nginx or HAProxy (Cluster Deployment). - Upstream health tracking Not documented
The vendor does not document this, so any behaviour you observe today is unversioned and may change.
There is health-driven channel disabling, but no enum here fits: it is configured in the console, not a config file, and it is certainly not "not_configurable". Channels have an Auto Disable option that "automatically disables the channel after consecutive failures" (Channel Management), and Model Behavior Settings expose "Auto-disable Failed Models", a "Failure Threshold" and an "Auto-recovery Time (minutes)" (System Settings - Detailed). Manual "Test" / "Test All Channels" actions report per-channel response time (Channel Management). No periodic active health-probe interval or circuit-breaker semantics are published.
- Cross-region failover Not documented
Multi-node clustering is documented (shared MySQL, shared Redis, identical
SESSION_SECRETandCRYPTO_SECRET,NODE_TYPEmaster/slave,SYNC_FREQUENCY60 s) and geographic distribution is *suggested*, but there is no configurable cross-region failover or region-pinning mechanism described (Cluster Deployment).
Fallback chain: Weighted split — Traffic splits by percentage across targets, so you can shift 5% to a new model and watch it before committing.
Reliability here is operator-assembled rather than declarative: retries, channel priority/weight, multi-key polling and auto-disable thresholds are all console settings backed by the database, while timeouts and cache/sync behaviour are environment variables (Channel Management, Environment Variables). Nothing in the fetched docs publishes an SLA, uptime target or status page, which is consistent with self-host-only distribution (Installation & Deployment).
How fast the hop is
Compiled binaryA single compiled Go or Rust binary. The lowest overhead floor of the self-hostable options, and the easiest to reason about under load.
Single Go service with an embedded web console: repo language bytes are Go 6,401,798 (~76%) and TypeScript 1,657,241, plus small JS/CSS/Lua/Shell (GitHub languages API, 2026-09-02); the repo primary language is Go and it ships a Dockerfile at root (GitHub API repo). Runtime dependencies are a database (SQLite, MySQL or PostgreSQL) and optionally Redis (README.en.md).
Docker image (calciumion/new-api:latest, ghcr.io mirror), a bundled docker-compose.yml in the repo, and panel one-click installs (1Panel, BaoTa >= 9.2.0) (README.en.md, Installation & Deployment). No Helm chart or Kubernetes manifests exist in the repository root tree (GitHub API repo, 2026-09-02).
Streaming caveats: Supported: chat requests take a stream boolean (Native Claude Format) and Gemini streaming is proxied via :streamGenerateContent?alt=sse (Gemini Text Chat). Operationally, STREAMING_TIMEOUT defaults to 300 s (time between streamed chunks), STREAM_SCANNER_MAX_BUFFER_MB defaults to 64 and FORCE_STREAM_OPTION defaults to true; the docs warn that a too-short relay timeout can leave a request billed upstream but not recorded locally (Environment Variables).
This vendor publishes no latency or throughput figure for the routing layer. That is the most common case here, and it is why the architecture class above carries the comparison instead of a number.
Nothing to assess: there are no vendor benchmarks, and no third-party benchmark of New API was found. Per-channel "response time" in the console measures the operator's upstream provider, not gateway overhead (Channel Management).
What it will stop
1 of 4 can blockTwo separate questions per control: can it stop a request at all, and what does it do before you change any settings? A control that inspects and forwards is a logging feature, however it is named.
- Personal data in prompts Not documented
n.a. - no PII detection, masking or redaction on any fetched page: Operational Settings, System Settings, System Settings - Detailed, Features Description, Acceptable Use, and the full docs Changelog (93 KB, no occurrence of PII/redaction/guardrail terms).
- Prompt injection and jailbreaks Not documented
n.a. - no prompt-injection or jailbreak detection is described on Features Description, Operational Settings, System Settings - Detailed or in the docs Changelog.
- Harmful content Can block the request
Out of the box: Not documented
A blocked-word ("blocked words") facility is part of Operational Settings - "Here you can navigate to operational settings such as top-up links, documentation addresses, blocked words, logging, monitoring, and quotas" (Operational Settings) - and the AUP counts blacklisting among the content-security features (Acceptable Use). No page fetched documents the matching mode, default state, or the error returned, so treat the semantics as undocumented; the request-time rejection behaviour is visible only in community reports of
message contains sensitive wordserrors. Separately,POST /v1/moderationsis available as a relayed upstream endpoint rather than an enforced gateway policy (Create Moderation, API Reference). - Your own policies Not documented
n.a. - there is no custom policy/rule engine for request content. The closest configurable hooks are channel-level Parameter Override JSON and request-field passthrough selection, which shape requests rather than police them (Channel Management, Changelog rc.25 entry on choosing which request fields pass through).
Almost no vendor in this catalogue documents what happens when the guardrail service itself times out. If the control matters to you, this is a question worth asking before you sign.
n.a. - no fail-open/fail-closed statement for the blocked-word path on Operational Settings or System Settings - Detailed.
Treat New API as a gateway with access control, not a guardrails platform. What it enforces well is *who may call which model with how much quota*; what it does not publish is any content, PII, injection or custom-policy engine beyond a blocked-word list whose behaviour is undocumented (Operational Settings, Token Management, Acceptable Use).
What you can see
No documented exportToken counts, latency and model names are stored, but not the text itself.
No switch to disable call logging is documented. Other Settings expose "Enable Log Export" and "Log Retention Days" (automatic cleanup), and Operational Settings lists "logging" among its sections without detailing an off switch (System Settings - Detailed, Operational Settings).
n.a. - no OpenTelemetry, OTLP or distributed-tracing support appears on any fetched page, including the full docs Changelog (no occurrence of OTel/OpenTelemetry/Prometheus). What exists is profiling: Go pprof at /debug/pprof/ behind ENABLE_PPROF, and optional Pyroscope continuous profiling via PYROSCOPE_URL / PYROSCOPE_APP_NAME with mutex and block rate controls (Performance Analysis, Environment Variables).
Metadata only as documented: the admin and user log pages enumerate time, username, model, channel, token name, tokens, quota and status columns and never mention request/response bodies (Log Management, Usage Logs). Logs can be written to a separate database with LOG_SQL_DSN (Environment Variables).
Where telemetry can go
No documented export. Whatever this product records stays in its own interface, so it cannot become part of the monitoring you already run.
n.a. - no log/metric shipping destinations (OTLP collectors, Kafka, S3, webhooks) are documented; the only stated egress is the user-facing log export toggle whose format is unspecified (System Settings - Detailed) and a separate log database via LOG_SQL_DSN (Environment Variables).
n.a. - no rating, thumbs-up/down or feedback ingestion endpoint on API Reference, Usage Logs or Features Description.
n.a. - no evaluation, scoring or dataset feature. The API reference index covers model relay plus management APIs only, with Files and Fine-tuning listed as "Unimplemented" (API Reference); the built-in Playground at /console/playground is manual testing, not evaluation (Using the API).
Depends on the vendor’s SaaS: No - the dashboard, logs and consumption charts are part of the self-hosted application and there is no vendor control plane to register with (Log Management, Installation & Deployment).
Retention: Retention is a number you choose in Other Settings rather than a vendor policy; there is no vendor-side copy of the data at all because the deployment is entirely yours (System Settings - Detailed, Installation & Deployment).
Whether it fits how you work
How much work stands between you and a first call, how different that is from running it in production, and whether it slots into the stack you already have. Recorded as the shape of the work rather than a number of minutes — how long it takes you depends on which accounts and quota you already hold, which no comparison can know.
to try: cli or container to run: infrastructure rollout fits 2 of 10 common stacks
Getting to a first call
3 numbered stepsNothing works until you have a process running. Fine on a laptop, but it means there is no zero-install way to try it.
Read off: the vendor’s own quickstart — 3 numbered steps.
Counted from the README's three-command quick start (clone the repo, edit docker-compose.yml, docker-compose up -d) (README.en.md); the homepage independently frames it as "3 steps" - configure channels, deploy and integrate, monitor and optimise (newapi.ai). The docs deployment page is prose rather than a numbered list, and first login additionally requires creating the administrator account (Docker Compose Deployment).
Before step one
- Your own provider key Required
You need an upstream provider account and key before anything works. That is a prerequisite, not a step.
There is no other way to get a model: a channel is defined by an upstream provider API key, and the AUP insists those keys be legally owned or authorised by the deployer (Channel Management, Acceptable Use).
- Payment method No card needed to start
Not required: the software is free under AGPLv3 and installs from a public Docker image with no account (Project Introduction, README.en.md). Payment integrations in the product are for charging *your* users, not for paying QuantumNous (System Settings - Detailed).
- Gate before models answer You enable it first
One extra click or API enable per model or project before a call succeeds.
Every model must be enabled by an administrator before it is callable: you create a channel with a provider key and the models it serves (optionally syncing the upstream model list with an Added/Changed/Deleted preview), then grant access through groups and per-token model restrictions (Channel Management, Model Management, Token Management). No vendor approval or waitlist exists - the gate is your own.
Everything you need first: Docker, a host to run it on, and an upstream provider key. There is no trial account, sign-up or credit card: docker run ... -v ./data:/data calciumion/new-api:latest, then browse to http://localhost:3000 and set up the administrator account (README.en.md, Docker Compose Deployment). SQLite is used by default for a local single-node instance (README.en.md).
Copyable snippet: incomplete. . A first-call snippet exists in principle - the page has Python (OpenAI SDK), Claude-native and Gemini-native examples - but the code blocks did not render in the fetched page text, so only the endpoint table and the base-URL-swap instruction are verifiable (Using the API). The deployment commands themselves are fully published (README.en.md).
Running it in production
The same scale applied to the path the vendor recommends for production traffic. Kept separate from the quickstart because for several products here the two are barely related pieces of work.
This is an infrastructure project, not an integration. Expect charts or Terraform, networking, secrets and someone who owns the deployment.
This is an infrastructure project, not an integration. Expect charts or Terraform, networking, secrets and someone who owns the deployment.
Getting to production is a step up in kind from the quickstart, not just more of the same.
What production needs: Remote MySQL >= 5.7.8 or PostgreSQL >= 9.6 (SQLite is single-node only), Docker and Docker Compose, and a 64-bit amd64/arm64 host (README.en.md). Multi-node adds shared Redis, an identical SESSION_SECRET across all nodes and an identical CRYPTO_SECRET when Redis is shared, NODE_TYPE master/slave, and an external load balancer such as Nginx or HAProxy (Cluster Deployment, README.en.md); Redis is the recommended cache backing via REDIS_CONN_STRING (Environment Variables).
Can you run it yourself
There is a command you can copy and run, so you can evaluate the self-hosted path yourself today.
docker run --name new-api -d --restart always -p 3000:3000 -e TZ=Asia/Shanghai -v ./data:/data calciumion/new-api:latest, or git clone plus docker-compose up -d with the bundled compose file; add -e SQL_DSN="root:123456@tcp(localhost:3306)/oneapi" for MySQL (README.en.md, Docker Compose Deployment).
How it fits your stack
2 of 10Each row is a thing you might already run. “With a caveat” means it works but not the way the vendor’s marketing implies — read the reason, because that is usually where the surprise lives.
- Fits The OpenAI SDK Drop-in once set up — but first-call work is cli or container.
- No The Vercel AI SDK No AI SDK route documented.
- No Cloudflare Workers No Workers guidance published.
- No Kubernetes No Kubernetes deployment published.
- No Terraform or OpenTofu Nothing published for Terraform.
- Fits An existing API gateway This is that gateway — AI traffic becomes a plugin, not a new hop.
- No Cloud IAM I already run Static upstream credentials only. Your calls to it still use its own key.
- With a caveat LangChain or LlamaIndex LangChain only.
- With a caveat MCP servers to govern undefined — governs nothing on your side.
- With a caveat Nothing — plain Node or Python You have to run a process locally before any call works.
Reading this the other way round — pick what you already run and see every product scored against it.
The integration surfaces behind those answers
- Vercel AI SDK Not documented
Nothing published. Assume the OpenAI-compatible route and verify it yourself.
n.a. - neither the Vercel AI SDK nor an
@ai-sdk/*package is named on Using the API, Verified Apps or the docs Changelog. In practice the documented OpenAI base-URL swap is what a Vercel AI SDK app would use, but the vendor does not document it (Using the API). - Cloudflare Workers Not documented
No Workers guidance either way. If you are edge-first, verify fetch-only compatibility yourself.
n.a. - it is a Go server that needs a database, not an edge worker; Cloudflare Workers appear nowhere on Installation & Deployment, Cluster Deployment or the docs Changelog.
- Kubernetes Not documented
No Kubernetes story published.
n.a. - Kubernetes is never mentioned on the fetched pages: Installation & Deployment, Docker Compose Deployment, Cluster Deployment (which scales with Docker plus an external Nginx/HAProxy load balancer instead) or the full docs Changelog. The repository root tree contains no
helm/ork8s/directory (GitHub API repo, 2026-09-02). - Terraform Not documented
No Terraform surface published. Configuration is API or dashboard work.
n.a. - no Terraform provider, module or example on Installation & Deployment, Cluster Deployment or in the docs Changelog, and no
terraform/directory in the repository root (GitHub API repo). - Existing API gateway It is the API gateway
This product is the gateway. If you already run it for your other APIs, AI traffic becomes a plugin rather than a new hop.
It *is* the gateway: "an AI API gateway and usage management system designed for legally authorized scenarios" (Project Introduction); the GitHub topics include
ai-gateway(GitHub API). - Cloud identity Static provider credentials only
You paste static provider credentials into this product so it can reach upstream models. Your own calls to it still use its own API key, and those pasted secrets are yours to rotate.
Authentication to upstreams is by provider API key held in the channel record; no cloud IAM role assumption (AWS SigV4, GCP service accounts, Azure managed identity) is documented (Channel Management). For human access, the console supports Discord, LinuxDO and Telegram OAuth plus OIDC unified authentication and two-factor auth (README.en.md, API Reference).
- MCP
No MCP support found, and this is a real gap for agent stacks: "MCP" does not occur on any page fetched, including API Reference, Features Description, Verified Apps and the entire 93 KB docs Changelog. What the project ships instead is its own "Skill" plugins - explicitly described as "a lightweight extension protocol" - installed with a single
npxcommand into Claude Code, Codex CLI, OpenClaw, Cursor, Windsurf and Cline, exposing/newapi models,/newapi groups, token and balance commands; anewapi-adminSkill is marked "Coming Soon" (Skills).
Thin evidence: the homepage's "Powering AI applications worldwide" strip lists LangChain, Dify, FastGPT, n8n, Open WebUI, LobeHub, Cherry Studio, Cline, Roo Code and Coze as applications that run on New API (newapi.ai), and the Rerank feature is documented as integrating with Dify (Features Description). There is no LangChain or LlamaIndex code-level integration page; LlamaIndex is not mentioned at all on the fetched pages.
The only language shown by name on the fetched integration page is Python via the OpenAI SDK (with Claude-native and Gemini-native request examples alongside); the code blocks themselves did not render in the fetched text (Using the API). Verified third-party clients are listed instead of SDKs - AionUi, Cherry Studio, DeepChat, OpenClaw, Claude Code, Codex CLI, Factory Droid CLI and others, all configured with just API address, key and model name (Verified Apps).
Agent features: Tool calling passes through in both dialects: the OpenAI surface is a drop-in for the OpenAI SDK (Using the API) and the native Claude surface accepts tools, tool_choice and thinking (Native Claude Format). Reasoning is controlled by model-name suffixes (-high/-medium/-low, claude-...-thinking, gemini-2.5-flash-nothinking, -thinking-128) with a per-channel thinking_to_content option, default false (Features Description); a rc.25 changelog entry notes usage logs now record reasoning effort consistently (Changelog). Agent CLIs are supported as clients rather than orchestrated by the gateway (Verified Apps).
First run is guided: after docker-compose up -d, visiting http://Server_IP:3000 "will automatically redirect to the initialization page" where you set the administrator account and password (Docker Compose Deployment). Then the homepage's three-step arc applies - configure channels, deploy and integrate, monitor and optimise (newapi.ai). Two documentation caveats: the Technical Architecture page and the support FAQ rendered essentially empty when fetched (149 characters and headings only respectively), and the docs changelog trails GitHub releases (Changelog showed rc.25 against rc.30 on the API).
Strong client-side ecosystem, weak platform-side one. Verified apps include AionUi, CC Switch, Cherry Studio, DeepChat, Memoh, OpenClaw, Fluent Read, LangBot, Luna Translator, AstrBot, Claude Code, Codex CLI and Factory Droid CLI, each needing only API address, key and model name (Verified Apps); official Skill plugins target Claude Code, Codex CLI, OpenClaw, Cursor, Windsurf and Cline from the QuantumNous/skills repo (Skills). Against that, no Kubernetes, Terraform, MCP or OpenTelemetry story is documented (Installation & Deployment, Changelog). Multi-language UI covers Chinese, English, French and Japanese, and much of the deeper documentation is Chinese-first (README.en.md).
Silence in the docs: 5 of the integration questions on this card have no published answer either way. That is recorded as undocumented, not as a no — but it does mean you would be verifying it yourself.
What it does well
- Three request dialects on one instance: OpenAI `/v1` (chat, completions, embeddings, images, audio, rerank, responses, realtime, models), native Anthropic Messages and native Gemini `/v1beta` proxying
- Billing is built in, not bolted on: quota accounting per user/token/model, three-tier Model/Completion/Group ratios at 1 USD = 500,000 quota, pre-consume then reconcile, cache-hit re-pricing
- Operator monetisation path out of the box - EPay, Stripe, Creem and Waffo top-ups, redemption codes, invitation rebates and subscription plans
- Practical reliability controls without a config file: channel priority plus weighted random, multi-key polling that skips failed keys, failure retry, auto-disable with failure threshold and auto-recovery time
- Access control is genuinely granular: per-token model restrictions and IP allowlists, group-based channel isolation, `auto` token group for cross-group failover
- Self-host-only and BYOK-only, so prompts never transit a vendor and there is no control plane to trust
- Very large active community and documented horizontal scaling (shared MySQL plus Redis, `NODE_TYPE` master/slave, external load balancer): 47,090 stars, 11,229 forks, pushed 2026-09-01
Where it falls short
- Heavy 2026 security record: 14 GitHub-reviewed advisories for the repo, including CVE-2026-71479 (CVSS 9.1, integer overflow in quota billing that let a user credit their own balance - confirmed exploited in the wild on 2026-07-06, fixed in v1.0.0-rc.18) and CVE-2026-64859 (root access token leaked via the user list API)
- Still v1.0.0-rc.30 after nearly three years and 47k stars - the shipping line is release candidates, and the docs themselves say stability is not guaranteed and support may not be provided
- AGPLv3: modify it and offer it as a network service and you must publish your source, unless you buy an unpriced commercial licence
- No SLA, no status page, no SOC 2/ISO 27001/HIPAA/GDPR DPA - expected for self-host-only software, but there is nothing to lean on contractually
- No Kubernetes, Helm, Terraform, MCP or OpenTelemetry support documented anywhere in the docs site or the full changelog
- Routing and channel configuration lives in the database and console, not in a reviewable declarative file
- Guardrails stop at a blocked-word list whose matching semantics, default state and failure mode are undocumented; no PII, injection or custom policy engine
- Documentation is uneven: the Technical Architecture and FAQ pages render essentially empty, the docs changelog lagged five release candidates behind GitHub, and the deepest material is Chinese-first
Choose it when
Teams that want to self-host one multi-dialect gateway and also *bill* the people behind it - per-token quotas, groups, ratios, top-ups and subscription plans - with no vendor in the request path and no vendor fee ([Project Introduction](https://www.newapi.ai/en/docs/guide/wiki/basic-concepts/project-introduction), [Token Management](https://www.newapi.ai/en/docs/guide/feature-guide/user/token), [System Settings - Detailed](https://www.newapi.ai/en/docs/guide/feature-guide/admin/system-setting-advanced)).
Look elsewhere when
You need vendor SLAs or compliance attestations (none published), Kubernetes/Helm/Terraform deployment, MCP or OpenTelemetry, or routing config that lives in Git rather than a database ([Installation & Deployment](https://www.newapi.ai/en/docs/installation), [Changelog](https://www.newapi.ai/en/docs/guide/wiki/changelog), [Channel Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/channel)) - or you cannot commit to tracking a fast-moving release-candidate line with a heavy 2026 advisory record ([GitHub advisories for the repo](https://api.github.com/advisories?affects=github.com/QuantumNous/new-api&ecosystem=go)).
Open source: You run the gateway yourself. Full control over the data path and no per-request vendor fee, in exchange for operating it.
How hard is it to leave?
Derived from six published facts, not from an opinion. The weights are fixed and the same for every product — see the arithmetic.
| What helps you leave | Points | Source |
|---|---|---|
| Works with standard OpenAI code Switching away is a base-URL change rather than a rewrite of every call site. | 22 /22 | vendor page |
| No vendor-specific SDK required A proprietary client library spreads through your codebase and has to be torn out again. | 10 /10 | — |
| Can use your own provider accounts Your keys and billing relationship stay yours, so removing the gateway does not cut off model access. | 20 /20 | — |
| Can be self-hosted You can run it yourself instead of accepting a pricing or policy change. | 20 /20 | vendor page |
| Configuration lives in version control Routing and budget rules are a file you keep, not dashboard state you would have to rebuild. | not published | — |
| Your request history can be exported You leave with your own logs instead of abandoning them. | 12 /12 | vendor page |
Read the fine print: Reasonable: usage logs can be exported by users when the Root toggle is on ([System Settings - Detailed](https://www.newapi.ai/en/docs/guide/feature-guide/admin/system-setting-advanced)), all state lives in your own SQLite/MySQL/PostgreSQL ([README.en.md](https://github.com/QuantumNous/new-api/blob/main/README.en.md)), and the schema is "fully compatible with the original One API database", which makes migration from that predecessor a drop-in ([README.en.md](https://github.com/QuantumNous/new-api/blob/main/README.en.md)). There is no declarative export of routing configuration.
1 of the 6 inputs is not published, so the highest reachable score here is 84 rather than 100. That is a gap in the public documentation, not a mark against the product — no points are deducted, they simply cannot be claimed. This measures technical switching cost only. It does not price the engineering time to re-test prompts against a different routing stack.
New API models & pricing
Browse every imported listing from this provider, with published token rates and a link to compare other providers for the same model. This is provider-reported coverage; an absent listing does not mean unsupported.
Loading model listings…
Official model coverage source ↗ · Model source coverage and limitations
Full specification
Every field we track. Blank fields say "Not published" rather than "No" — we do not infer an absence from silence. Switch to Technical in the header for the precise field names and the low-level details.
Overview
Interpret these fields: Self-hosted vs managed LLM gateways · Who still owns your LLM gateway?
- What kind of product Category
- Open source
- Marketplaces resell many providers behind one key. Gateways add governance on top. Open-source projects you run yourself. Cloud platforms are hyperscaler surfaces. Inference providers host models on their own hardware.
- Who runs it Deployment model
- Self-host only
- Managed means the vendor operates it. Self-host means you run it on your own infrastructure. Both means you can choose.
- Licence Licence
- AGPL-3.0
- Proprietary products cannot be inspected or forked. Open licences such as MIT and Apache-2.0 let you audit, modify, and run the code without permission.
- Who you would be signing with Vendor status
- Independent company
- Whether the product is still an independent company, has been acquired, is a large cloud vendor’s product line, is run by a software foundation, or has been put into maintenance mode. Maintenance mode means bug fixes and security patches only — no new features.
- Last shipped an update Latest release
- 2026-08-31
v1.0.0-rc.30, published 2026-08-31T03:55:59Z. GitHub marks it `prerelease: false` even though the tag is a release candidate, so this is the project's shipping release line rather than a preview ([releases/latest](https://api.github.com/repos/QuantumNous/new-api/releases/latest), 2026-09-02). The docs changelog page lagged behind at v1.0.0-rc.25 ("Data updated at 2026-8-25") when checked ([Changelog](https://www.newapi.ai/en/docs/guide/wiki/changelog)).
- The date of the most recent release or version tag. A product that has not shipped in a year is a different risk from one that shipped last week, regardless of what its marketing site says.
- GitHub stars GitHub stars
- 48,314
- A rough proxy for community size on open-source projects. Not a quality measure.
Cost
Interpret these fields: How LLM gateway pricing works · LLM gateway spending limits: stop a runaway agent bill?
- Markup on model prices Token markup
- None
- How much the product adds on top of what the underlying model provider charges. Zero means you pay the same per-token price you would pay the model provider directly.
- Fee to add funds Credit purchase fee
- Not published
- A percentage charged when you top up your balance, separate from token prices. It is easy to miss because it does not appear on the per-token price list.
- Monthly cost per person Seat fee
- None
- A recurring per-user platform charge that applies regardless of how much you use the models.
- Can use your own provider accounts BYOK supported
- Yes
- Bring Your Own Key: you keep direct contracts with OpenAI, Anthropic and others, and the gateway only routes traffic. This preserves negotiated rates and committed-spend discounts.
- Cost of using your own accounts BYOK terms
- No fee: New API charges nothing for relaying your own keys, and the upstream provider bills you directly ([Project Introduction](https://www.newapi.ai/en/docs/guide/wiki/basic-concepts/project-introduction), [Channel Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/channel)).
- What the product charges to route traffic through your own provider keys.
- Free tier Free tier
- The whole product: "New API adopts the GNU AGPLv3 open-source license" and is free to deploy and use provided the licence is honoured; there is no feature-gated tier ([Project Introduction](https://www.newapi.ai/en/docs/guide/wiki/basic-concepts/project-introduction)).
- What you can do without paying, useful for evaluation.
- Enterprise plan from Enterprise plan from
- Not published
- Annual entry price for the enterprise tier, where one is published or credibly reported.
- Cost to run it yourself Self-host cost
- You pay only for your own infrastructure and upstream provider spend: the software is AGPLv3 and "can be used for free" so long as the licence is observed, with a paid commercial licence available by email for AGPL-exempt use ([Project Introduction](https://www.newapi.ai/en/docs/guide/wiki/basic-concepts/project-introduction)). Production sizing implies a MySQL and a Redis alongside the gateway ([Cluster Deployment](https://www.newapi.ai/en/docs/installation/deployment-methods/cluster-deployment)).
- What self-hosting actually costs once you account for infrastructure and any paid tier.
- How the vendor makes money Pricing model
- Open source, no paid tier
- The shape of the vendor’s bill: does the routing layer charge a percentage on top of tokens, a flat monthly fee, both, neither (because inference is the product), or nothing at all (open source with no paid tier).
- How pricing works, briefly Pricing model detail
- No vendor price at all: the project is AGPLv3 and free to run, with a paid **commercial licence** for AGPL-exempt use available by emailing support@quantumnous.com (price not published) ([Project Introduction](https://www.newapi.ai/en/docs/guide/wiki/basic-concepts/project-introduction), business enquiries at [Business Cooperation](https://www.newapi.ai/en/docs/business)). Everything the pricing pages of hosted gateways would cover is instead operator-configured: three-tier Model/Completion/Group ratios with 1 USD = 500,000 quota points, pre-consumption followed by post-consumption reconciliation, and a default ratio of 37.5 for unpriced models in self-use mode (an error in commercial mode) ([System Settings - Detailed](https://www.newapi.ai/en/docs/guide/feature-guide/admin/system-setting-advanced)). The `/en/docs/guide/pricing` page is about the price table your own instance shows to your users, not a vendor price list.
- A one-paragraph description that covers the caveats a pricing category cannot: introductory rates, per-feature meters, tier gating, and pricing that resets on a specific date.
- Minimum commitment Minimum commitment
- None - no vendor contract exists; you commit only to your own infrastructure and upstream provider spend ([Project Introduction](https://www.newapi.ai/en/docs/guide/wiki/basic-concepts/project-introduction), [Installation & Deployment](https://www.newapi.ai/en/docs/installation)).
- Whether the vendor requires a minimum contract term, a minimum spend, or a provisioned-capacity purchase to get its published rate.
- Charges that fire after you go over an allowance Overage terms
- n.a. - no vendor metering. Overage semantics are yours to configure: tokens auto-disable when their remaining quota is exceeded ([Token Management](https://www.newapi.ai/en/docs/guide/feature-guide/user/token)) and pre-consumption is reconciled after the call completes ([System Settings - Detailed](https://www.newapi.ai/en/docs/guide/feature-guide/admin/system-setting-advanced)).
- The line items that scale with usage after an included allowance is exhausted — log storage, extra requests, per-feature meters, data export — which is where cost estimates usually go wrong.
- Prompt cache offered Cache mechanism
- Passes provider caching through
- Whether the gateway offers its own response cache, what kind of match it does (exact request, prefix, semantic), or simply passes provider caching through unchanged.
- Discount on cached input Cache-read discount
- Not published
- How much cheaper cached tokens are than fresh input, when the vendor publishes a single figure. Bundled-inference clouds usually price this per model instead of as one number.
- Premium on cache writes Cache-write premium
- Not published
- How much more the first write of a cached prefix costs versus a plain input token. A high write premium and a low hit rate can leave you paying more than you save, so this matters as much as the read discount.
- Who captures the cache saving Cache economics
- New API does not cache responses itself; it re-prices upstream prompt-cache hits. A **Prompt Cache Ratio** between 0 and 1 is set globally or per channel - "0.5 means cache-hit tokens are billed at 50%" - and cache billing is documented for OpenAI, Azure, DeepSeek, Claude and Qwen channels ([Features Description](https://www.newapi.ai/en/docs/guide/wiki/basic-concepts/features-introduction), [README.en.md](https://github.com/QuantumNous/new-api/blob/main/README.en.md)). The Redis and `MEMORY_CACHE_ENABLED` caches are for platform state (channels, tokens, quotas), not model responses ([Environment Variables](https://www.newapi.ai/en/docs/installation/config-maintenance/environment-variables)). No vendor-fixed cached-token discount exists because the operator sets the ratio.
- Whether the customer keeps the full saving from caching or the vendor captures part of it — and any conditions attached (write premium, storage fees, best-effort hits).
- What you can split spend by Cost attribution
- Per user, per token, per model and per channel, from the log tables and dashboard: user logs show quota deducted per call ([Usage Logs](https://www.newapi.ai/en/docs/guide/feature-guide/user/log)), admin logs add username and channel name ([Log Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/log)), and the console dashboard charts consumption trends. Group multipliers let cost be differentiated by user group ([Group Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/group)). No per-tag or per-customer-metadata dimension is documented.
- The dimensions the vendor documents for splitting spend — per key, per user, per team, per tag, per customer. Matters if you need to chargeback internally or bill an end customer.
- How you get cost data out Cost export
- "Enable Log Export: allow users to export their usage logs" is a Root-only toggle; the export format is not stated on the page, and no API/webhook/S3/warehouse cost export is described ([System Settings - Detailed](https://www.newapi.ai/en/docs/guide/feature-guide/admin/system-setting-advanced)). Because the data is in your own MySQL/PostgreSQL, direct SQL is the practical export path ([Environment Variables](https://www.newapi.ai/en/docs/installation/config-maintenance/environment-variables)).
- The mechanisms the vendor publishes for exporting cost and usage data: CSV, an API, webhooks, S3, a data warehouse, or nothing at all. Any per-unit price is included.
- Who pays the model bill BYOK mode
- Your keys only
- Whether you bring your own provider accounts (BYOK), buy inference from this vendor, or can do either. This is the single biggest commercial difference between these products: it decides who holds the contract with the model provider and who carries the spend.
Spend governance in detail
Seven signals matter when a bill starts to hurt: who can spend, how much, on what, and who gets paged when it goes wrong. Everything below is drawn from the vendor’s own pricing and docs pages — how we read these.
- Virtual or scoped keys Yes
Keys that carry their own budget and rate-limit policy, so an intern experiment cannot spend against a production budget.
Tokens are first-class virtual keys with expiry, remaining quota, unlimited-quota flag, model restrictions, IP allowlist and group; the secret is displayed once at creation ([Token Management](https://www.newapi.ai/en/docs/guide/feature-guide/user/token)).
- Budget caps per key Yes
A dollar or token ceiling attached to an individual key. Where enforcement is soft, one over-limit request still completes before the block kicks in.
Per-token "Remaining Quota"; the token is automatically disabled once the quota is exceeded ([Token Management](https://www.newapi.ai/en/docs/guide/feature-guide/user/token)).
- Budget caps per team or workspace Not published
A ceiling applied at a higher scope than one key — a team, a workspace, a customer, or an entire environment.
Groups control channel access and billing multipliers and carry per-group rate limits, but a group-level spend cap is not stated ([Group Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/group), [System Settings - Detailed](https://www.newapi.ai/en/docs/guide/feature-guide/admin/system-setting-advanced)).
- Rate limiting as a cost control Yes
Configurable request-per-time-window caps. Platform-set rate limits do not count as spend controls; user-configurable ones do.
Global per-IP requests per minute/hour/day, per-group `{group: [per-minute, per-hour]}` limits with 0 meaning unlimited, plus model rate limiting in Rate Limit Settings covering total and successful request counts ([System Settings - Detailed](https://www.newapi.ai/en/docs/guide/feature-guide/admin/system-setting-advanced), [Features Description](https://www.newapi.ai/en/docs/guide/wiki/basic-concepts/features-introduction)).
- Model allowlists Yes
A policy that constrains which models a key or team can call, keeping expensive frontier models out of the wrong hands.
Per-token model restrictions, per-channel model lists and group-based channel isolation ([Token Management](https://www.newapi.ai/en/docs/guide/feature-guide/user/token), [Channel Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/channel), [Group Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/group)).
- Spend alerts Not published
Alerts fired as spend approaches a threshold. Alerts that only fire after the meter has rolled over are marked as such.
Not stated on [System Settings - Detailed](https://www.newapi.ai/en/docs/guide/feature-guide/admin/system-setting-advanced) or [Log Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/log); the dashboard shows consumption trends but no alert configuration is described.
- Webhook notifications Not published
Programmatic notifications on spend events, so budget breaches can page an on-call or open a ticket.
Not stated; the only documented webhook is the Stripe payment webhook `https://your-domain.com/api/payment/stripe/webhook` ([System Settings - Detailed](https://www.newapi.ai/en/docs/guide/feature-guide/admin/system-setting-advanced)).
Enforcement: Enforced before each request
Catalog
Interpret these fields: LLM gateway model counts: what “500+” means
- Models available Models available
- 100+
- How many models you can call, shown as a range because several vendors publish different totals on different pages. Vendor-reported either way, so counts are not directly comparable — some count every provider variant of the same model separately.
- Model providers reachable Upstream providers
- 30+
- How many distinct model providers or labs you can reach, shown as a range where the vendor’s own pages disagree. More providers usually means better redundancy when one has an outage.
- Works with standard OpenAI code OpenAI-compatible API
- Yes
- If yes, you can usually switch to it by changing one base URL, and switch away just as easily. This is the main defence against lock-in.
- OpenAI chat endpoint POST /v1/chat/completions
- Yes
- The endpoint almost every application ports first. "Not documented" means the vendor never states it, which is different from a documented no.
- Anthropic messages endpoint POST /v1/messages
- Yes
- Whether Anthropic-shaped calls work without rewriting them. Several products support this only as SDK compatibility or provider passthrough rather than a native endpoint — the detail page says which.
- OpenAI Responses endpoint POST /v1/responses
- Yes
- The newer stateful OpenAI surface. Support is much thinner across this market than chat completions.
- Embeddings endpoint POST /v1/embeddings
- Yes
- Whether you can generate vectors through the same gateway, or need a second integration for retrieval workloads.
- Image generation endpoint POST /v1/images/generations
- Yes
- Whether image models are reachable through the same surface as text.
- Audio endpoints POST /v1/audio/*
- Yes
- Speech-to-text and text-to-speech. Frequently the first gap in an otherwise complete gateway.
- Batch jobs endpoint POST /v1/batches
- Not documented
- Asynchronous bulk processing, usually at a discount. Commonly undocumented, and commonly the reason a migration stalls late.
- Needs the vendor’s own code library Requires a vendor-specific SDK
- No
- A proprietary client library spreads through your codebase and has to be torn out again if you leave. “No” is the better answer here, and it means the standard OpenAI client works.
- You can export your request history Logs / usage data export
- Yes
- Whether you can get your own request logs, traces, or usage records back out — through an API, a bulk export, or a download. Decides whether you leave with your history or abandon it.
- Settings can live in version control Declarative config-as-code
- Not published
- Whether routing, fallback, and budget rules can be declared in a file you keep in Git, rather than existing only as settings clicked into a hosted dashboard.
- Embeddings Embeddings
- Yes
- Text-to-vector models, needed for search and retrieval features.
- Image generation Image generation
- Yes
- Whether image models are reachable through the same interface.
- Speech and audio Speech and audio
- Yes
- Text-to-speech or transcription models through the same interface.
- Video generation Video generation
- Yes
- Whether video models are reachable through the same interface.
- Batch processing Batch processing
- Not published
- Submitting large jobs for cheaper, slower processing. Often 50% off for work that is not time-sensitive.
Routing & reliability
Interpret these fields: How LLM gateway failover actually works · Does your LLM gateway promise any uptime? · Is routing destroying your prompt cache? · Changing models without breaking production
- Uptime it promises in writing Contractual SLA uptime
- Not published
- The uptime percentage in a published, contractual service level agreement. A public status page is not an SLA — it reports what happened, it does not promise anything or pay you back when it breaks.
- Automatic failover Automatic failover
- Yes
- When a model provider goes down or rate-limits you, traffic moves to a backup automatically instead of returning errors to your users.
- Load balancing Load balancing
- Yes
- Spreads requests across several providers or keys to raise your effective rate limit.
- Rule-based routing Conditional routing
- Not published
- Send different requests to different models based on rules — for example a cheap model for free users and a strong model for paying ones.
- Response caching Response caching
- Not published
- Reuses the answer when the exact same request comes in again, which cuts both cost and latency.
- Similar-question caching Semantic cache
- Not published
- Reuses an answer when a new question means roughly the same thing as an earlier one. Saves far more than exact-match caching but can return subtly wrong answers if tuned loosely.
- Where you set the timeout Request timeout surface
- In config
Environment variables only: `RELAY_TIMEOUT` defaults to `0`, meaning **no timeout**, `STREAMING_TIMEOUT` is 300 s between chunks, `TASK_TIMEOUT_MINUTES` is 1440 for async tasks, and connection pooling is tuned with `RELAY_MAX_IDLE_CONNS` 500 / `RELAY_MAX_IDLE_CONNS_PER_HOST` 100. The docs explicitly warn that setting the relay timeout too low can produce requests the upstream charges for but New API never bills ([Environment Variables](https://www.newapi.ai/en/docs/installation/config-maintenance/environment-variables)).
- Where a request timeout can be set: per request, in a config file, in the vendor dashboard, or nowhere. Nine of the twenty products documented here do not describe a request timeout at all, so the worst case of a hung upstream call is unknowable from the docs.
- Where you set retries Retry policy surface
- Dashboard only
"Automatic retry on failure" is a listed routing feature and the retry count is a console setting at `Settings -> Operation Settings -> General Settings -> Failure Retry Count`; the default value and any backoff strategy are not published ([README.en.md](https://github.com/QuantumNous/new-api/blob/main/README.en.md)). Default retry count: `n.a.` Backoff: `n.a.`
- Where retry count and backoff are configured. Worth knowing alongside billing: a retried streaming call can be charged more than once.
- Where you set fallbacks Fallback surface
- Dashboard only
Fallback is channel-level and configured per channel in the console: **Priority** ("higher value = higher selection priority", default 0) selects the tier and **Weight** is the "random weight among same-priority channels" (default 0), so selection within a tier is weighted random while tiers are ordered ([Channel Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/channel)); the README's headline phrasing is "Channel weighted random" ([README.en.md](https://github.com/QuantumNous/new-api/blob/main/README.en.md)). Token groups can be set to `auto`, which "automatically selects an available group in priority order - useful for cross-group failover scenarios" ([Group Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/group)). Multi-Key channels additionally poll keys in Round Robin or Weighted Random order and skip a failed key until it recovers ([Channel Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/channel)).
- Where the fallback chain is defined. Almost every product claims fallback; the useful question is whether you can change it from code or only by hand in a dashboard.
- Shape of the fallback chain Fallback shape
- Weighted split
- An ordered list tries targets in sequence; a weighted split sends a percentage of traffic to each, which is what you need to trial a new model on 5% of requests. Weighted splits are much rarer than the marketing implies.
- Upstream health tracking Health checks / circuit breaking
- Not published
There is health-driven channel disabling, but no enum here fits: it is configured in the console, not a config file, and it is certainly not "not_configurable". Channels have an **Auto Disable** option that "automatically disables the channel after consecutive failures" ([Channel Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/channel)), and Model Behavior Settings expose "Auto-disable Failed Models", a "Failure Threshold" and an "Auto-recovery Time (minutes)" ([System Settings - Detailed](https://www.newapi.ai/en/docs/guide/feature-guide/admin/system-setting-advanced)). Manual "Test" / "Test All Channels" actions report per-channel response time ([Channel Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/channel)). No periodic active health-probe interval or circuit-breaker semantics are published.
- Whether the product notices a failing upstream and stops sending traffic to it, and whether you can tune the thresholds. This is what turns a provider outage into a blip rather than a sustained error rate, and only four of the twenty expose it.
- Cross-region failover you control Multi-region failover surface
- Not documented
Multi-node clustering is documented (shared MySQL, shared Redis, identical `SESSION_SECRET` and `CRYPTO_SECRET`, `NODE_TYPE` master/slave, `SYNC_FREQUENCY` 60 s) and geographic distribution is *suggested*, but there is no configurable cross-region failover or region-pinning mechanism described ([Cluster Deployment](https://www.newapi.ai/en/docs/installation/deployment-methods/cluster-deployment)).
- Whether you can define what happens when a region degrades. A vendor running many regions is not the same as a vendor letting you configure failover between them; only three document a user-controlled mechanism.
- Where you set load balancing Load balancing surface
- Dashboard only
Weights and priorities are user-settable per channel in `/console/channel`, with model mapping and per-channel parameter override alongside ([Channel Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/channel)); there is no routing config file - the channel table lives in the database. Cluster-level balancing across nodes is left to an external load balancer such as Nginx or HAProxy ([Cluster Deployment](https://www.newapi.ai/en/docs/installation/deployment-methods/cluster-deployment)).
- Where traffic distribution across upstreams or keys is configured.
Operations
Interpret these fields: LLM gateway observability: traces, logs and export · Running coding agents through an LLM gateway · Changing models without breaking production
- Usage dashboards and logs Observability
- Yes
- Built-in visibility into what was sent, what came back, what it cost, and how long it took.
- Spending limits Budget controls
- Yes
- Hard caps that stop spend before it becomes a surprise invoice. The single most valuable control for a small team.
- Rate limits Rate limits
- Yes
- Caps on request volume per key or per user, useful for protecting against abuse and runaway loops.
- Separate keys per team or app Virtual keys
- Yes
- Issue scoped keys with their own budgets and permissions so you can attribute cost and revoke access without rotating everything.
- Prompt versioning Prompt management
- Not published
- Store and version prompts outside your code so they can be changed without a deploy.
- Quality testing Evals
- Not published
- Built-in tooling to score model output against test cases, so you can tell whether a model swap made things better or worse.
- MCP support MCP support
- Not published
- Native support for the Model Context Protocol, the emerging standard for connecting models to external tools.
- What gets logged Logged content
- Metadata only
Metadata only as documented: the admin and user log pages enumerate time, username, model, channel, token name, tokens, quota and status columns and never mention request/response bodies ([Log Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/log), [Usage Logs](https://www.newapi.ai/en/docs/guide/feature-guide/user/log)). Logs can be written to a separate database with `LOG_SQL_DSN` ([Environment Variables](https://www.newapi.ai/en/docs/installation/config-maintenance/environment-variables)).
- Whether full prompts and responses are stored, only metadata, or nothing. Full-body logging is the most useful debugging feature here and the one most likely to need a conversation with your compliance team.
- You can turn logging off Body-logging opt-out
- Not documented
No switch to disable call logging is documented. Other Settings expose "Enable Log Export" and "Log Retention Days" (automatic cleanup), and Operational Settings lists "logging" among its sections without detailing an off switch ([System Settings - Detailed](https://www.newapi.ai/en/docs/guide/feature-guide/admin/system-setting-advanced), [Operational Settings](https://www.newapi.ai/en/docs/guide/console/settings/operation-settings)).
- Whether prompt and response bodies can be suppressed while still keeping usage metrics. Nineteen of the twenty document a way to do this; the mechanisms range from a per-request header to an organisation-wide setting.
- Traces you can take elsewhere Distributed tracing
- Not documented
n.a. - no OpenTelemetry, OTLP or distributed-tracing support appears on any fetched page, including the full docs [Changelog](https://www.newapi.ai/en/docs/guide/wiki/changelog) (no occurrence of OTel/OpenTelemetry/Prometheus). What exists is profiling: Go `pprof` at `/debug/pprof/` behind `ENABLE_PPROF`, and optional Pyroscope continuous profiling via `PYROSCOPE_URL` / `PYROSCOPE_APP_NAME` with mutex and block rate controls ([Performance Analysis](https://www.newapi.ai/en/docs/guide/wiki/basic-concepts/performance-analysis), [Environment Variables](https://www.newapi.ai/en/docs/installation/config-maintenance/environment-variables)).
- Whether the product emits OpenTelemetry, a proprietary format, or nothing. OpenTelemetry means the traces land in the tooling you already run instead of only in the vendor’s dashboard.
- Where telemetry can go Export destinations
- Not published
n.a. - no log/metric shipping destinations (OTLP collectors, Kafka, S3, webhooks) are documented; the only stated egress is the user-facing log export toggle whose format is unspecified ([System Settings - Detailed](https://www.newapi.ai/en/docs/guide/feature-guide/admin/system-setting-advanced)) and a separate log database via `LOG_SQL_DSN` ([Environment Variables](https://www.newapi.ai/en/docs/installation/config-maintenance/environment-variables)).
- Documented sinks for logs and metrics. This is a good proxy for how replaceable the vendor’s own dashboard is: around twenty destinations means you never have to depend on it, while a CSV download means you do.
- Can record user feedback Feedback capture API
- No
n.a. - no rating, thumbs-up/down or feedback ingestion endpoint on [API Reference](https://www.newapi.ai/en/docs/api), [Usage Logs](https://www.newapi.ai/en/docs/guide/feature-guide/user/log) or [Features Description](https://www.newapi.ai/en/docs/guide/wiki/basic-concepts/features-introduction).
- Whether there is an API to attach a rating or score to a logged request, which is what lets production traffic feed quality work later.
- Scores live traffic Online eval hooks
- No
n.a. - no evaluation, scoring or dataset feature. The API reference index covers model relay plus management APIs only, with Files and Fine-tuning listed as "Unimplemented" ([API Reference](https://www.newapi.ai/en/docs/api)); the built-in Playground at `/console/playground` is manual testing, not evaluation ([Using the API](https://www.newapi.ai/en/docs/guide/feature-guide/user/api)).
- Whether automated scorers can run against real production requests, rather than only against a test set you assemble yourself.
Performance
- Delay it adds Proxy overhead
- Not published
- Extra time the product itself adds to each request, on top of however long the model takes. Usually irrelevant next to multi-second model latency, but it matters for high-volume or streaming-sensitive workloads.
- Requests per second ceiling Throughput
- Not published
- Published sustained request rate before the product becomes the bottleneck. Only relevant at genuinely high volume.
- What the request path runs on Architecture class
- Compiled binary
Single Go service with an embedded web console: repo language bytes are Go 6,401,798 (~76%) and TypeScript 1,657,241, plus small JS/CSS/Lua/Shell ([GitHub languages API](https://api.github.com/repos/QuantumNous/new-api/languages), 2026-09-02); the repo primary language is Go and it ships a Dockerfile at root ([GitHub API repo](https://api.github.com/repos/QuantumNous/new-api)). Runtime dependencies are a database (SQLite, MySQL or PostgreSQL) and optionally Redis ([README.en.md](https://github.com/QuantumNous/new-api/blob/main/README.en.md)).
- The comparable way to talk about latency here. An edge worker, a compiled Go or Rust binary, and a Python proxy have different overhead floors no matter which figures each vendor publishes. Products that never disclose their runtime are recorded as undisclosed rather than assumed.
- You can run the request path yourself Self-hostable data plane
- Yes
Docker image (`calciumion/new-api:latest`, ghcr.io mirror), a bundled `docker-compose.yml` in the repo, and panel one-click installs (1Panel, BaoTa >= 9.2.0) ([README.en.md](https://github.com/QuantumNous/new-api/blob/main/README.en.md), [Installation & Deployment](https://www.newapi.ai/en/docs/installation)). No Helm chart or Kubernetes manifests exist in the repository root tree ([GitHub API repo](https://api.github.com/repos/QuantumNous/new-api), 2026-09-02).
- Whether the component that actually carries your prompts can run on your own infrastructure. Distinct from a vendor offering a self-hosted control plane while still proxying traffic through their network.
- Streaming responses Streaming support
- Yes
Supported: chat requests take a `stream` boolean ([Native Claude Format](https://www.newapi.ai/en/docs/api/ai-model/chat/createmessage)) and Gemini streaming is proxied via `:streamGenerateContent?alt=sse` ([Gemini Text Chat](https://www.newapi.ai/en/docs/api/ai-model/chat/gemini/geminirelayv1beta)). Operationally, `STREAMING_TIMEOUT` defaults to 300 s (time between streamed chunks), `STREAM_SCANNER_MAX_BUFFER_MB` defaults to 64 and `FORCE_STREAM_OPTION` defaults to true; the docs warn that a too-short relay timeout can leave a request billed upstream but not recorded locally ([Environment Variables](https://www.newapi.ai/en/docs/installation/config-maintenance/environment-variables)).
- Whether token-by-token streaming is documented. The caveats matter more than the yes: some products cannot cancel a stream without still being billed, and several timeout and fallback mechanisms stop applying once the first token has been sent.
Security & compliance
Interpret these fields: LLM gateway compliance: SOC 2, HIPAA and evidence · How LLM gateway guardrails fail · Which LLM gateways store your prompts? · Do you need an MCP gateway as well?
- Does your prompt reach their servers Prompt transits vendor
- No
No vendor is in the path: New API is distributed as software you deploy, so requests go from your client to your instance to the upstream provider whose key you configured ([Installation & Deployment](https://www.newapi.ai/en/docs/installation), [Channel Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/channel)). The AUP describes the intended scenarios as "self-use, internal team use, and enterprise private deployment" ([Acceptable Use](https://www.newapi.ai/en/docs/legal/acceptable-use)).
- Whether the text you send passes through this company’s own infrastructure. If it does, every other promise on this page is a policy commitment rather than a physical impossibility. Self-hosted products can answer no outright.
- What they keep if you change nothing Logging default
- Metadata only, not content
Call logs are on and are per-request metadata: time, model, tokens consumed, quota deducted and status for users ([Usage Logs](https://www.newapi.ai/en/docs/guide/feature-guide/user/log)), with username and channel name columns added for administrators plus filters by time range, username, model, channel and token name ([Log Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/log)). No prompt or completion body storage is described on either page. Error-detail logging is separately gated by `ERROR_LOG_ENABLED`, default `false` ([Environment Variables](https://www.newapi.ai/en/docs/installation/config-maintenance/environment-variables)).
- Defaults matter more than options. A product that stores full prompts and replies unless you find the right header will have stored them by the time you read the docs.
- How long they keep it Default content retention (days)
- Not published
Operator-controlled: "Log Retention Days - how many days of logs the system automatically cleans up" is a Root-only setting with no published default ([System Settings - Detailed](https://www.newapi.ai/en/docs/guide/feature-guide/admin/system-setting-advanced)); logs live in your own database and can be split out with `LOG_SQL_DSN` ([Environment Variables](https://www.newapi.ai/en/docs/installation/config-maintenance/environment-variables)).
- Default retention for request content, in days. Zero means nothing is kept. Read the note: several products keep nothing as a rule but make timed exceptions for abuse review or specific models.
- Could they train on your prompts Training on customer data
- Not applicable
not_applicable by construction: there is no hosted service, so no operator-side data reaches QuantumNous ([Installation & Deployment](https://www.newapi.ai/en/docs/installation)). The one documented outbound call from the software is model-metadata synchronisation from `SYNC_UPSTREAM_BASE`, default `https://basellm.github.io/llm-metadata` ([Environment Variables](https://www.newapi.ai/en/docs/installation/config-maintenance/environment-variables)). No separate training policy is published.
- Whether the vendor may use your prompts and outputs to train models. “Not published” means we could not find any position, which is not the same as a no — ask for it in writing.
- Where it runs, and what you can pin Region and residency control
- n.a. as a vendor concept - the operator chooses where to run it. Multi-node and geographically distributed deployments are documented via a shared primary database plus shared Redis, `NODE_TYPE` master/slave and an external load balancer ([Cluster Deployment](https://www.newapi.ai/en/docs/installation/deployment-methods/cluster-deployment)); the sample container sets `TZ=Asia/Shanghai` ([README.en.md](https://github.com/QuantumNous/new-api/blob/main/README.en.md)).
- Which regions are offered and whether you can force processing to stay in one. A global endpoint that silently picks a region is a different compliance story from an endpoint you pin yourself.
- Where safety filters run Guardrail execution location
- In your own infrastructure
Whatever filtering exists runs inside the operator's own instance - there is no vendor service in the path ([Installation & Deployment](https://www.newapi.ai/en/docs/installation)). The AUP is explicit that governance is the deployer's job: "The deployer is obligated to ensure that services provided through this project comply with content security requirements. Features such as blacklisting, logging, and monitoring should serve content security, abuse governance, and compliance auditing" ([Acceptable Use](https://www.newapi.ai/en/docs/legal/acceptable-use)). The relayed `/v1/moderations` endpoint is likewise framed as "one of the compliance tools and does not replace the deployer's own safety governance obligations" ([API Reference](https://www.newapi.ai/en/docs/api)).
- A filter that strips personal data only helps if it runs before the data leaves your boundary. If guardrails execute in the vendor’s cloud, the vendor has already received whatever you wanted redacted.
- Who else touches the data Subprocessor list
- Not published
- The published list of third parties the vendor passes your data to. No list means you cannot know the full chain, which most data-protection agreements require you to.
- SOC 2 audited SOC 2 audited
- Not published
- An independent audit of security controls. Enterprise buyers and their procurement teams routinely require it.
- Will sign a HIPAA agreement HIPAA BAA
- Not published
- Required before you may send protected health information through the service. Without a signed BAA, healthcare data is off limits.
- GDPR commitments GDPR commitments
- Not published
- Published data processing terms for handling personal data of people in the EU and UK.
- Can keep data in the EU EU data residency
- Not published
- Requests can be processed inside the EU rather than routed to US infrastructure. Often the deciding constraint for European customers.
- Does not retain your data Zero data retention
- Not applicable
- Prompts and responses are not stored after the request completes. Sometimes a paid add-on rather than the default.
- Strips personal data PII redaction
- Not published
- Detects and removes identifiers such as names, emails, and card numbers before the request reaches the model provider.
- Content guardrails Content guardrails
- Yes
- Policy checks on inputs and outputs — blocking unsafe content, enforcing formats, or catching prompt-injection attempts.
- Runs fully disconnected Air-gapped deployment
- Not published
- Can be deployed in a network with no internet access, which some regulated and defence environments require.
- Blocks personal data in prompts PII / DLP enforcement
- Not documented
n.a. - no PII detection, masking or redaction on any fetched page: [Operational Settings](https://www.newapi.ai/en/docs/guide/console/settings/operation-settings), [System Settings](https://www.newapi.ai/en/docs/guide/feature-guide/admin/system-setting), [System Settings - Detailed](https://www.newapi.ai/en/docs/guide/feature-guide/admin/system-setting-advanced), [Features Description](https://www.newapi.ai/en/docs/guide/wiki/basic-concepts/features-introduction), [Acceptable Use](https://www.newapi.ai/en/docs/legal/acceptable-use), and the full docs [Changelog](https://www.newapi.ai/en/docs/guide/wiki/changelog) (93 KB, no occurrence of PII/redaction/guardrail terms).
- Whether personal data detection sits on the request path and can stop the call, merely inspects and forwards it, or is not documented. A control that only reports is a logging feature, not a policy control.
- Blocks prompt injection Injection / jailbreak enforcement
- Not documented
n.a. - no prompt-injection or jailbreak detection is described on [Features Description](https://www.newapi.ai/en/docs/guide/wiki/basic-concepts/features-introduction), [Operational Settings](https://www.newapi.ai/en/docs/guide/console/settings/operation-settings), [System Settings - Detailed](https://www.newapi.ai/en/docs/guide/feature-guide/admin/system-setting-advanced) or in the docs [Changelog](https://www.newapi.ai/en/docs/guide/wiki/changelog).
- Whether injection and jailbreak detection can stop a request. Most products offering this call a partner classifier rather than shipping their own.
- Blocks harmful content Toxicity / moderation enforcement
- Can block the request
A blocked-word ("blocked words") facility is part of Operational Settings - "Here you can navigate to operational settings such as top-up links, documentation addresses, blocked words, logging, monitoring, and quotas" ([Operational Settings](https://www.newapi.ai/en/docs/guide/console/settings/operation-settings)) - and the AUP counts blacklisting among the content-security features ([Acceptable Use](https://www.newapi.ai/en/docs/legal/acceptable-use)). No page fetched documents the matching mode, default state, or the error returned, so treat the semantics as undocumented; the request-time rejection behaviour is visible only in community reports of `message contains sensitive words` errors. Separately, `POST /v1/moderations` is available as a relayed upstream endpoint rather than an enforced gateway policy ([Create Moderation](https://www.newapi.ai/en/docs/api/ai-model/moderations/createmoderation), [API Reference](https://www.newapi.ai/en/docs/api)).
- Whether hate, violence, sexual and self-harm categories are checked inline and can stop a request, in either direction.
- Your own policy rules Custom policy hooks
- Not documented
n.a. - there is no custom policy/rule engine for request content. The closest configurable hooks are channel-level **Parameter Override** JSON and request-field passthrough selection, which shape requests rather than police them ([Channel Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/channel), [Changelog](https://www.newapi.ai/en/docs/guide/wiki/changelog) rc.25 entry on choosing which request fields pass through).
- Whether you can add your own rule — a regex, a webhook, or your own classifier — rather than choosing from the vendor library.
- Where guardrails run Guardrail execution location
- Not documented
- Whether guardrail evaluation happens inside your infrastructure or on the vendor’s servers. This decides whether the prompt you are trying to protect leaves your network in order to be checked.
- If the guardrail itself fails Guardrail failure mode
- Not documented
n.a. - no fail-open/fail-closed statement for the blocked-word path on [Operational Settings](https://www.newapi.ai/en/docs/guide/console/settings/operation-settings) or [System Settings - Detailed](https://www.newapi.ai/en/docs/guide/feature-guide/admin/system-setting-advanced).
- What happens when the guardrail service times out or errors: does the request proceed unchecked, or is it blocked? This is the worst-documented field in the entire catalogue — only two vendors state it plainly, which means most teams are running a control whose failure behaviour they cannot know.
- Third-party guardrail vendors Guardrail integrations
- Not published
- Named external guardrail services the product can call. A long list means the product is a router for policy engines rather than a policy engine itself — which also means another vendor bill and another hop.
Compliance evidence
Graded by how strong the evidence is, not whether the word appears on the vendor’s website. An audited report and a marketing claim are different things, and only one of them will satisfy your own auditor.
- SOC 2 Not published Self-host-only open-source project; no attestation, trust portal or certification page found on the docs site (checked [Acceptable Use](https://www.newapi.ai/en/docs/legal/acceptable-use), [Business Cooperation](https://www.newapi.ai/en/docs/business), [API Reference](https://www.newapi.ai/en/docs/api)).
- ISO 27001 Not published
- GDPR DPA Not published No DPA or GDPR page; the AUP pushes compliance obligations onto the deployer, including "identity management and log retention" and filing/qualification duties ([Acceptable Use](https://www.newapi.ai/en/docs/legal/acceptable-use)).
- HIPAA BAA Not published
- FedRAMP Not published
- ITAR Not published
No compliance certifications were found published for this product. That is not the same as failing an audit — it means there is nothing public to check, so ask for evidence directly.
Security incidents
Publicly documented incidents affecting this product. An incident here is not by itself a reason to rule a product out — what matters is what failed, whether it could recur, and what it means for the way you would deploy it.
Fit & integration
Interpret these fields: How much does an LLM gateway lock you in? · Do you need an MCP gateway as well? · Running coding agents through an LLM gateway · Changing models without breaking production
- Work to try it Evaluation work shape
- Run something locally first
- The shape of the work on the vendor’s own quickstart, from swapping one base URL through to deploying infrastructure. An ordinal class rather than a duration, because elapsed time depends on accounts and quota we cannot see.
- Work to run it Production work shape
- Deploy it on your infrastructure
- The same scale applied to the vendor’s recommended production path. For several products this is much heavier than the quickstart, which is exactly why both are recorded.
- Steps on the quickstart Numbered quickstart steps
- 3
Counted from the README's three-command quick start (clone the repo, edit `docker-compose.yml`, `docker-compose up -d`) ([README.en.md](https://github.com/QuantumNous/new-api/blob/main/README.en.md)); the homepage independently frames it as "3 steps" - configure channels, deploy and integrate, monitor and optimise ([newapi.ai](https://www.newapi.ai/)). The docs deployment page is prose rather than a numbered list, and first login additionally requires creating the administrator account ([Docker Compose Deployment](https://www.newapi.ai/en/docs/installation/deployment-methods/docker-compose-installation)).
- A literal count of numbered steps on the vendor’s quickstart, recorded as evidence beside the work shape. Zero means the page publishes no numbered procedure at all. Large counts usually mean interleaved language tracks rather than more work.
- Can you self-host it today Self-host install documentation
- Install command published
`docker run --name new-api -d --restart always -p 3000:3000 -e TZ=Asia/Shanghai -v ./data:/data calciumion/new-api:latest`, or `git clone` plus `docker-compose up -d` with the bundled compose file; add `-e SQL_DSN="root:123456@tcp(localhost:3306)/oneapi"` for MySQL ([README.en.md](https://github.com/QuantumNous/new-api/blob/main/README.en.md), [Docker Compose Deployment](https://www.newapi.ai/en/docs/installation/deployment-methods/docker-compose-installation)).
- Whether an install command is actually published. Several products advertise self-hosting while publishing no command to start from, which a plain yes/no would hide.
- Works with the OpenAI SDK OpenAI SDK drop-in
- Yes
Yes, stated plainly: "Replace OpenAI's `base_url` with the platform address and use your token as the `api_key`", with the OpenAI Python SDK as the first example and an eleven-row `/v1/...` endpoint table ([Using the API](https://www.newapi.ai/en/docs/guide/feature-guide/user/api)). The homepage frames it as a "100% OpenAI compatible endpoint for all AI providers" ([newapi.ai](https://www.newapi.ai/)).
- Whether an existing OpenAI-compatible client can be pointed at it by changing the base URL and key.
- Vercel AI SDK support AI SDK provider package
- Not documented
n.a. - neither the Vercel AI SDK nor an `@ai-sdk/*` package is named on [Using the API](https://www.newapi.ai/en/docs/guide/feature-guide/user/api), [Verified Apps](https://www.newapi.ai/en/docs/apps) or the docs [Changelog](https://www.newapi.ai/en/docs/guide/wiki/changelog). In practice the documented OpenAI base-URL swap is what a Vercel AI SDK app would use, but the vendor does not document it ([Using the API](https://www.newapi.ai/en/docs/guide/feature-guide/user/api)).
- Whether a first-party AI SDK provider package exists, or only a community package, a documented workaround, or the generic OpenAI provider pointed at a custom base URL.
- Python framework integrations Documented Python frameworks
- LangChain
Thin evidence: the homepage's "Powering AI applications worldwide" strip lists LangChain, Dify, FastGPT, n8n, Open WebUI, LobeHub, Cherry Studio, Cline, Roo Code and Coze as applications that run on New API ([newapi.ai](https://www.newapi.ai/)), and the Rerank feature is documented as integrating with Dify ([Features Description](https://www.newapi.ai/en/docs/guide/wiki/basic-concepts/features-introduction)). There is no LangChain or LlamaIndex code-level integration page; LlamaIndex is not mentioned at all on the fetched pages.
- LangChain, LangGraph, LlamaIndex and similar orchestration frameworks with a documented integration.
- Callable from Cloudflare Workers Cloudflare Workers support
- Not documented
n.a. - it is a Go server that needs a database, not an edge worker; Cloudflare Workers appear nowhere on [Installation & Deployment](https://www.newapi.ai/en/docs/installation), [Cluster Deployment](https://www.newapi.ai/en/docs/installation/deployment-methods/cluster-deployment) or the docs [Changelog](https://www.newapi.ai/en/docs/guide/wiki/changelog).
- Whether the docs show calling this product from your own Worker. Deliberately separated from the several products whose own gateway runs on Workers, which is a fact about their infrastructure and not about your edge compatibility.
- Kubernetes install Helm chart availability
- Not documented
n.a. - Kubernetes is never mentioned on the fetched pages: [Installation & Deployment](https://www.newapi.ai/en/docs/installation), [Docker Compose Deployment](https://www.newapi.ai/en/docs/installation/deployment-methods/docker-compose-installation), [Cluster Deployment](https://www.newapi.ai/en/docs/installation/deployment-methods/cluster-deployment) (which scales with Docker plus an external Nginx/HAProxy load balancer instead) or the full docs [Changelog](https://www.newapi.ai/en/docs/guide/wiki/changelog). The repository root tree contains no `helm/` or `k8s/` directory ([GitHub API repo](https://api.github.com/repos/QuantumNous/new-api), 2026-09-02).
- Whether a named, published Helm chart exists, versus Helm being referenced with no chart named, versus generic cluster documentation that has nothing to do with this product.
- Terraform support Terraform provider or modules
- Not documented
n.a. - no Terraform provider, module or example on [Installation & Deployment](https://www.newapi.ai/en/docs/installation), [Cluster Deployment](https://www.newapi.ai/en/docs/installation/deployment-methods/cluster-deployment) or in the docs [Changelog](https://www.newapi.ai/en/docs/guide/wiki/changelog), and no `terraform/` directory in the repository root ([GitHub API repo](https://api.github.com/repos/QuantumNous/new-api)).
- Whether you can declare this in code: an official provider, official modules, resources inside a hyperscaler’s provider, a community provider, or only Terraform code shipped in a repo.
- Reuses your cloud identity Cloud IAM reuse
- Static provider credentials only
Authentication to upstreams is by provider API key held in the channel record; no cloud IAM role assumption (AWS SigV4, GCP service accounts, Azure managed identity) is documented ([Channel Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/channel)). For human access, the console supports Discord, LinuxDO and Telegram OAuth plus OIDC unified authentication and two-factor auth ([README.en.md](https://github.com/QuantumNous/new-api/blob/main/README.en.md), [API Reference](https://www.newapi.ai/en/docs/api)).
- Whether you can authenticate with IAM roles, workload identity or managed identities instead of another long-lived API key. Distinguished from products that only accept static upstream provider credentials.
- Fits behind your API gateway API gateway integration
- It is the API gateway
It *is* the gateway: "an AI API gateway and usage management system designed for legally authorized scenarios" ([Project Introduction](https://www.newapi.ai/en/docs/guide/wiki/basic-concepts/project-introduction)); the GitHub topics include `ai-gateway` ([GitHub API](https://api.github.com/repos/QuantumNous/new-api)).
- Whether AI traffic can go through a gateway you already run, and who documented that — the product’s vendor or the gateway’s.
- MCP support MCP surface shape
- Not published
No MCP support found, and this is a real gap for agent stacks: "MCP" does not occur on any page fetched, including [API Reference](https://www.newapi.ai/en/docs/api), [Features Description](https://www.newapi.ai/en/docs/guide/wiki/basic-concepts/features-introduction), [Verified Apps](https://www.newapi.ai/en/docs/apps) and the entire 93 KB docs [Changelog](https://www.newapi.ai/en/docs/guide/wiki/changelog). What the project ships instead is its own "Skill" plugins - explicitly described as "a lightweight extension protocol" - installed with a single `npx` command into Claude Code, Codex CLI, OpenClaw, Cursor, Windsurf and Cline, exposing `/newapi models`, `/newapi groups`, token and balance commands; a `newapi-admin` Skill is marked "Coming Soon" ([Skills](https://www.newapi.ai/en/docs/skills)).
- Which kind of MCP support this is: a gateway that governs many MCP servers, a hosted MCP server you connect to, MCP tools accepted inside the completion API, or client tooling. These are different products behind one acronym.
- Needs your own provider key Upstream provider key required
- Required
Yes - there is no other way to get a model: a channel is defined by an upstream provider API key, and the AUP insists those keys be legally owned or authorised by the deployer ([Channel Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/channel), [Acceptable Use](https://www.newapi.ai/en/docs/legal/acceptable-use)).
- Whether an upstream provider account and key must exist before your first call works. A prerequisite rather than a step, and it can differ between the hosted and self-hosted forms of the same product.
- Gate before models work Model access gate
- You enable it first
Every model must be enabled by an administrator before it is callable: you create a channel with a provider key and the models it serves (optionally syncing the upstream model list with an Added/Changed/Deleted preview), then grant access through groups and per-token model restrictions ([Channel Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/channel), [Model Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/model), [Token Management](https://www.newapi.ai/en/docs/guide/feature-guide/user/token)). No vendor approval or waitlist exists - the gate is your own.
- Whether an enablement click, a quota grant, a paid tier or an approval form stands between a valid key and a working model call.
- Official client languages First-party client SDK languages
- Python
The only language shown by name on the fetched integration page is Python via the OpenAI SDK (with Claude-native and Gemini-native request examples alongside); the code blocks themselves did not render in the fetched text ([Using the API](https://www.newapi.ai/en/docs/guide/feature-guide/user/api)). Verified third-party clients are listed instead of SDKs - AionUi, Cherry Studio, DeepChat, OpenClaw, Claude Code, Codex CLI, Factory Droid CLI and others, all configured with just API address, key and model name ([Verified Apps](https://www.newapi.ai/en/docs/apps)).
- Languages with a first-party client library. An empty list can still mean the product is usable from any language via an OpenAI-compatible SDK.
Additional charges
These are the fees that do not appear on a per-token price list, and they are where estimates usually go wrong.
- Commercial licence (AGPLv3 exemption) Price not published; by enquiry to support@quantumnous.com
How pricing actually works
You pay only for your own infrastructure and upstream provider spend: the software is AGPLv3 and "can be used for free" so long as the licence is observed, with a paid commercial licence available by email for AGPL-exempt use ([Project Introduction](https://www.newapi.ai/en/docs/guide/wiki/basic-concepts/project-introduction)). Production sizing implies a MySQL and a Redis alongside the gateway ([Cluster Deployment](https://www.newapi.ai/en/docs/installation/deployment-methods/cluster-deployment)).
Common questions
Answered from the fields above, so these move when the catalog moves. Every figure quoted here appears in the specification with its source.
Does New API charge a markup on model prices?
New API adds no percentage markup to model prices. Other charges on the page: commercial licence (agplv3 exemption) (Price not published; by enquiry to support@quantumnous.com). With no per-token cut and no top-up charge, what you pay is the model providers' own rates plus whatever the infrastructure costs you to run.
Can New API be self-hosted?
Yes — self-hosting is the only way to run New API; there is no vendor-hosted option. The licence is AGPL-3.0. Running it yourself means you supply the infrastructure and the upstream model accounts, so the bill is your own hosting plus the providers' own rates.
Is New API SOC 2 audited, and will it sign a HIPAA BAA?
New API publishes neither a SOC 2 report nor a HIPAA business associate agreement. Each of these is linked to the vendor's own page in the compliance section below. Neither absence means a refusal: both are things a vendor either publishes or does not, and smaller products often hold the certification without advertising it.
Does New API retain your prompts?
Zero data retention does not apply to New API: it runs inside your own infrastructure, so prompts never reach a vendor. Only metadata is logged — not prompt or response bodies. Retention becomes your own configuration question instead, decided by whatever logging you switch on in your own deployment.
Can you use your own provider keys with New API?
Yes. New API can route through your own accounts with the underlying model providers, so inference is billed to you directly. No fee: New API charges nothing for relaying your own keys, and the upstream provider bills you directly ([Project Introduction](https://www.newapi.ai/en/docs/guide/wiki/basic-concepts/project-introduction), [Channel Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/channel)).
How many models does New API support?
New API's own pages give between 100 and 0 models, drawn from 30+ upstream providers. "Choose from 100+ models across multiple providers" ([newapi.ai](https://www.newapi.ai/)). This is a marketing figure about the upstream catalogue rather than a shipped list - the actual model set on any instance is whatever the operator's channels expose, and the console can sync upstream model lists with an Added/Changed/Deleted preview ([Model Management](https://www.newapi.ai/en/docs/guide/feature-guide/admin/model)). The figure on this page is dated and carries its source.
Official links
48,314 GitHub stars — a proxy for community size, not for quality.
Independent coverage
Third-party analysis, walkthroughs and operator threads. We link criticism as readily as praise. Nothing here is written or published by the vendor, by a competitor listed on this site, or by an SEO content farm — how we vet these.
Written reviews and analysis 2
- Unlimited AI Credits from One Integer: New API's Quota Overflow (CVE-2026-71479) and How to Check Your Gateway Independent security write-up that reproduces the billing overflow against the real product (a $0.10 balance became $16,893,488,147,419.20), confirms the rc.18 fix, and gives operators a version-check command against `/api/status`.
- new-api review: the self-hosted AI gateway that turns every LLM into one OpenAI-compatible endpoint Hands-on operator review running new-api in production across MiniMax, DeepSeek and others, with a useful comparison against one-api, LiteLLM and Vercel AI Gateway. Note it misattributes QuantumNous as "the team behind the Qwen open-weight models", so treat its factual claims with care.
Video 2
- new-api:多模型 LLM Gateway 與 AI 資產管理中控台|開源專案介紹 Project walkthrough covering the aggregation-and-distribution model, the supported dialects (OpenAI-compatible, Responses, Claude Messages, Gemini, Rerank, images, audio, Midjourney Proxy, Suno) and where permissions, routing, billing and format conversion sit.
- 自建一个统一的 AI 接口:New API + Docker Compose 实操指南 End-to-end deployment demo on Debian 13: Docker install, `docker compose up -d`, initialisation, choosing the operating mode, then channel, model-mapping and token configuration - a realistic view of the setup effort.
What has changed here
- GitHub stars GitHub stars 47090 48314 source ↗
- catalog entry catalog entry Not published Added to the catalog source ↗
Read the head-to-head
These pairs have a written verdict, not just a table.