New API vs LiteLLM

The question that decides it: Are you charging the people who use the gateway, or standardising provider access for your own platform?

Our verdict

Charging or metering end users — quotas, top-ups, subscription plans, per-user ratios — and able to honour AGPLv3 inside your own organisation: New API, because nothing else here ships that billing layer. Standardising provider access for a platform you operate, or reselling a modified build without publishing source: LiteLLM, for 140 providers, MIT and a real infrastructure story.

Why

On the money these two are indistinguishable. Both are self-host-only, both charge zero token markup, both have a $0 seat fee, and neither puts a vendor in the request path — you run the binary, you hold the upstream keys, the provider bills you. What separates them is that they are not the same kind of product. LiteLLM's own docs title it "LiteLLM AI Gateway (LLM Proxy)": a proxy that maps many provider APIs onto OpenAI-format input and output. New API's README calls it a "Next-Generation Large Model Gateway and AI Asset Management System", it was developed from One API and remains fully compatible with the original One API database, and roughly half of what it ships is the accounting around the proxy rather than the proxy itself.

Breadth and operations both favour LiteLLM, and not narrowly. LiteLLM lists 140 providers; New API's homepage claims "30+ Model Providers" and "100+ models", both stated as floors rather than counted lists, while LiteLLM publishes no aggregate model total at all — its README's "100+" is used for models on one page and providers on another. LiteLLM has official Helm charts, official Terraform modules for AWS and Google Cloud, YAML config-as-code, MCP gateway endpoints, exact-match and semantic caching, Presidio-backed PII masking, batch endpoints and OpenTelemetry as a first-class callback. New API documents Docker Compose, single-container Docker, 1Panel and BaoTa one-click installs, and a cluster mode that scales behind an external Nginx or HAProxy; Kubernetes, Helm, Terraform, MCP and OpenTelemetry appear nowhere on the pages checked, and routing lives in the console database rather than in Git. Star counts are closer than that gap suggests: 57,500 against 47,090.

New API's answer is the thing LiteLLM does not attempt. Quota is accounted per user, per token, per model and per channel at 1 USD = 500,000 quota points, with three-tier Model, Completion and Group ratios, pre-consumption reconciled after the call, and cache-hit re-pricing via a per-channel Prompt Cache Ratio. Around that sit EPay, Stripe, Creem and Waffo top-up integrations, redemption codes, invitation rebates and subscription plans. If your requirement is a self-hosted gateway that also invoices the people behind the keys, that is the product, and it is a genuinely awkward thing to assemble on top of LiteLLM's per-key budgets. The catch is that the licence bites hardest on exactly that use case: New API is AGPL-3.0, so a modified build offered as a network service obliges you to publish your source, unless you buy the commercial exemption — whose price is not published and which is arranged by email. LiteLLM is MIT and carries no such condition; its paywall is elsewhere, with SSO, RBAC, audit logs and SCIM behind a commercial licence whose price is also unpublished.

Neither project has a clean operational record, and it would be dishonest to grade only one of them. New API carries 14 GitHub-reviewed advisories for the repository in 2026, including CVE-2026-71479, a CVSS 9.1 integer overflow in quota billing that let a user credit their own balance; it was confirmed exploited in the wild on 2026-07-06 and fixed in v1.0.0-rc.18 about two hours later, with CVE-2026-64859 separately leaking a root access token through the user list API. The shipping line is still a release candidate — v1.0.0-rc.30, published 2026-08-31 — and the docs themselves say stability is not guaranteed and support may not be provided. LiteLLM's incident was upstream of its own code: under CVE-2026-33634, versions 1.82.7 and 1.82.8 went straight to PyPI with a credential stealer, never tagged on GitHub, bypassing the project's CI/CD entirely. Compliance is thin on both sides. New API publishes no SOC 2, ISO 27001, HIPAA or GDPR DPA at all, and its acceptable-use policy pushes identity management, log retention and filing duties onto the deployer. LiteLLM's own record conflicts: its docs describe SOC 2 Type II as in progress with an ETA of 15 September 2026 while the enterprise page markets it as done, and ISO 27001 is claimed but unverified.

Which one, concretely

Choose New API if

  • You need per-user quotas, top-ups, redemption codes and subscription plans built into the gateway
  • You want three request dialects on one instance: OpenAI /v1, native Anthropic Messages and native Gemini v1beta
  • You are migrating off One API — the schema is fully compatible, so it is a drop-in
  • You want one Go service plus SQLite for a single node, with no Python runtime to operate

Choose LiteLLM if

  • 140 providers against New API's stated floor of 30+
  • Official Helm charts, Terraform modules for AWS and Google Cloud, and YAML config-as-code instead of dashboard-only routing
  • Semantic and exact-match caching, an MCP gateway, batch endpoints and Presidio PII masking
  • MIT, so a modified build can be offered as a service without publishing your source

What catches people out

Side by side

3 of 18 fields differ, marked with a dot. Every figure links to the vendor page it came from. Blank values read Not published rather than No — silence from a vendor is not a negative answer.

Field New API LiteLLM
Ease of leaving Derived score, higher is easier 84/84 Some work to leave 100/100 Easy to leave
What kind of product Category Open source Open source
Who runs it Deployment model Self-host only Self-host only
Licence Licence AGPL-3.0 MIT
Models available Models available 100+ Not published
Model providers reachable Upstream providers 30+ Not published
Markup on model prices Token markup None None
Fee to add funds Credit purchase fee Not published None
Monthly cost per person Seat fee None None
GitHub stars GitHub stars 48,314 59,000
Settings can live in version control Declarative config-as-code Not published Yes
MCP support MCP support Not published Yes
Similar-question caching Semantic cache Not published Yes
Strips personal data PII redaction Not published Yes
Batch processing Batch processing Not published Yes
Runs fully disconnected Air-gapped deployment Not published Yes
Delay it adds Proxy overhead Not published 0.66 ms
Requests per second ceiling Throughput Not published 2,800 rps
Separate keys per team or app Virtual keys Yes Yes

for New API and for LiteLLM. Want more fields, or a third option in the mix? Open these two in the full comparison tool.

Common questions

Which one supports more providers and models?

LiteLLM, clearly. It lists 140 providers; New API states floors of "30+ Model Providers" and "100+ models" on its homepage rather than publishing an enumerated list. On models the comparison is unresolvable in the other direction: LiteLLM publishes no aggregate model total, and the "100+" figure in its README is used for models on one page and providers on another. Judge on the provider gap.

Does New API's AGPL-3.0 licence matter for my use case?

It depends entirely on whether you offer it onward. Running an unmodified New API for self-use, an internal team or a private enterprise deployment is what the project describes itself as being for. Modifying it and exposing that build as a network service triggers the AGPL source-publication obligation, unless you buy the commercial exemption — whose price is not published. LiteLLM is MIT and has no equivalent condition.

How fast is each one?

New API publishes no latency or throughput figure of any kind; its performance documentation covers pprof and Pyroscope instrumentation rather than results. LiteLLM publishes 0.66 ms of added latency and 2,800+ RPS from its own bench, plus 8 ms p95 at 1k RPS in its README, all self-run. Third-party operators separately report the Python proxy hitting a GIL bottleneck past roughly 300 RPS per instance. Load-test with your own traffic.