Respan Managed gateway
Respan is a managed LLM gateway: an OpenAI-compatible API in front of 1,525 models. Platform pricing is plan-dependent; token markup and credit fees are not published. It can be self-hosted or used as a managed service. Its published HIPAA terms conflict; confirm the applicable signed agreement. You can point it at your own provider accounts. It handles failover, load balancing, guardrails and request logging. EU data residency is available.
· 176 dated entries · 176 source references
Access at a glance
Whether this product can work for you at all, before features matter: who pays the model bill, where it can run, and which of your existing API calls keep working. Every value is the vendor’s own claim, linked to the page it came from.
Respan (formerly Keywords AI) combines a managed LLM gateway with observability, evaluations and prompt tooling.
Who pays the model bill
Your keys or their creditsYou can start on their credits and move to your own provider accounts later.
Use Respan credits or connect provider credentials; available funding mode depends on the model.
Merchant of record: Respan credits fund supported upstream traffic; BYOK uses the customer’s provider account.
Key handling: Provider keys may be stored in the dashboard or supplied per request. Multiple weighted credentials and exact-model credential overrides are documented.
Where it can run
2 of 5 shapes documented- Vendor-hosted
- Self-host
- Your VPC
- On-premise
- Air-gapped
Dimmed shapes are not documented by the vendor, which is not the same as unsupported.
Cloud on Free and Team; Enterprise lists Cloud and Self-hosted. Public installation artifact and air-gap operation not established.
Enterprise self-hosting is sales-led.
API surfaces your code can keep using
2 of 7 documented, 1 partial- OpenAI chat
POST /v1/chat/completionsYesOpenAI SDK Chat Completions with base URL https://api.respan.ai/api/.
- Anthropic messages
POST /v1/messagesYesProvider-native Anthropic passthrough under /api/anthropic/.
- OpenAI Responses
POST /v1/responsesPartlyDocumented /api/responses routes: OpenAI, Azure OpenAI and Perplexity Agent API, with route-specific model and credential requirements.
- Embeddings
POST /v1/embeddingsNot documented - Images
POST /v1/images/generationsNot documented - Audio
POST /v1/audio/*Not documented - Batch jobs
POST /v1/batchesNot documented
An asterisk marks a qualified verdict: support that is indirect (SDK compatibility or provider passthrough rather than a native endpoint), or a gap that is narrower or wider than the label suggests. Read the note before porting.
Unified Chat Completions and provider-native passthrough are separate routes. Do not infer image generation, audio or batch inference from multimodal telemetry.
How much it reaches
Models: Counted from the vendor’s own published list; no aggregate total is published.
Count of entries returned by the models API on 2026-09-23. Includes every entry exposed by that endpoint; not a count of unique base models.
Integrations directory includes provider setup and tracing integrations; no comparable exhaustive upstream count recorded.
Whose models: Routes third-party models through a unified endpoint or provider-native passthrough.
Your own endpoints: Custom provider supports an OpenAI-compatible base URL and credentials; create a matching model entry.
How it behaves in production
What happens when an upstream model is slow, wrong, or down — and what you can see and stop while it happens. Reliability features are recorded as where you configure them, not whether the vendor lists them, because almost every product here lists all of them.
3 of 6 reachable from code 1 of 4 can block 2 documented destinations
When something goes wrong
Each row says where the knob is, not whether the feature is on the marketing page. A control you can only reach by hand in someone else’s dashboard cannot be reviewed or version-controlled.
- Request timeout Not documented
The vendor does not document this, so any behaviour you observe today is unversioned and may change.
- Retries Per request
Your application decides per call, so one noisy endpoint can have its own timeout without a redeploy.
retry_enabled, num_retries and retry_after govern current-route attempts; UI also available.
- Fallback to another model Per request
fallback_models is tried in order after the current route exhausts eligible attempts; preflight and fail-fast errors stop the chain.
- Load balancing Per request
Weighted model groups and weighted provider credentials may be configured in the dashboard or request.
- Upstream health tracking Not documented
- Cross-region failover Not documented
Fallback chain: Ordered list — Try A, then B, then C. Simple and predictable, but every failover is all-or-nothing.
Defaults: num_retries includes the initial attempt. Exponential backoff with jitter, capped at 60 seconds; no stable default attempt count recorded.
Retries happen before fallback. Invalid credentials, oversized context and model-side read timeout can be fail-fast.
How fast the hop is
Undisclosed vendor serviceThe vendor does not disclose what the request path runs on, so no overhead floor can be inferred at all.
AWS ECS API servers, Redis queues, Celery consumers, PostgreSQL and ClickHouse; hosted architecture as documented.
Pricing lists enterprise self-hosting; no public gateway installation command was found.
Streaming caveats: stream=True documented for Chat Completions.
Published figures, grouped by what each one measured. Figures in different groups are different quantities and cannot be compared with one another — nor, in most cases, with another vendor’s figure in the same group.
Overhead added by the gateway
The only figures that speak to whether this product slows your application down. Still not comparable between vendors: each was measured on different hardware, at a different load, with a different payload.
- 50–150 ms Vendor-published
Quickstart warning; no workload or percentile published. Not independently measured.
Source
What it will stop
1 of 4 can blockTwo separate questions per control: can it stop a request at all, and what does it do before you change any settings? A control that inspects and forwards is a logging feature, however it is named.
One of these controls can block a request, but will not until you change its settings:
- Personal data in prompts — Ships switched off
- Personal data in prompts Can block the request
Out of the box: Ships switched off
Pre-forward redaction masks detected entities; it transforms content rather than rejecting the entire request. Independent telemetry redaction; both disabled by default.
- Prompt injection and jailbreaks Not documented
- Harmful content Not documented
- Your own policies Not documented
Almost no vendor in this catalogue documents what happens when the guardrail service itself times out. If the control matters to you, this is a question worth asking before you sign.
Probabilistic PII detection. No documented general fail-open/fail-closed guarantee found.
What you can see
Exports to a few placesPrompt and completion bodies are stored by default. Powerful for debugging, and a data-residency question you have to answer before you ship.
Existing OpenTelemetry exporters can send spans to Respan; tracing SDK adds nested application spans.
Gateway logging includes request/response content and usage. Telemetry-only integrations are a separate path.
Where telemetry can go
- CSV
- JSONL
One-time and recurring log exports with selectable fields, filters and sampling.
Sampled production span or completed-trace scoring using deployed evaluator pipelines.
Depends on the vendor’s SaaS: Hosted default uses Respan storage and UI; enterprise self-hosting is advertised separately.
Retention: Plan-dependent retention: 7 days Free, 30 days Team, custom Enterprise.
Whether it fits how you work
How much work stands between you and a first call, how different that is from running it in production, and whether it slots into the stack you already have. Recorded as the shape of the work rather than a number of minutes — how long it takes you depends on which accounts and quota you already hold, which no comparison can know.
to try: base url swap to run: base url swap fits 4 of 10 common stacks
Getting to a first call
3 numbered stepsYour existing OpenAI-compatible client keeps working. You change a base URL and a key, and nothing else in your code moves.
Read off: the vendor’s own quickstart — 3 numbered steps.
Three setup steps before the optional prompt-management step; this is not a timed integration test.
Before step one
- Your own provider key Optional
You can start on the product’s own credits and move to your own provider keys later.
Some catalog listings require BYOK. Credits-supported routes do not require your own provider key.
- Payment method No card needed to start
No credit card required to start the free platform plan; model traffic still requires credits or an eligible provider key.
- Gate before models answer Not documented
The docs do not say, so budget for a surprise on the first model you actually want.
Active status does not guarantee account access. Catalog labels distinguish Credits from BYOK.
Everything you need first: Respan account and API key; credits or an eligible upstream provider key.
Running it in production
The same scale applied to the path the vendor recommends for production traffic. Kept separate from the quickstart because for several products here the two are barely related pieces of work.
Your existing OpenAI-compatible client keeps working. You change a base URL and a key, and nothing else in your code moves.
What production needs: Configure production credentials, spending limits, retention, fallback and logging policy.
Can you run it yourself
Self-hosting is advertised and the requirements are described, but no page publishes a command to start from. Expect to talk to the vendor before you can run it.
How it fits your stack
4 of 10Each row is a thing you might already run. “With a caveat” means it works but not the way the vendor’s marketing implies — read the reason, because that is usually where the surprise lives.
- Fits The OpenAI SDK Drop-in: change the base URL and key, nothing else.
- With a caveat The Vercel AI SDK Documented integration, not a provider
- No Cloudflare Workers No Workers guidance published.
- No Kubernetes No Kubernetes deployment published.
- No Terraform or OpenTofu Nothing published for Terraform.
- Fits An existing API gateway This is that gateway — AI traffic becomes a plugin, not a new hop.
- No Cloud IAM I already run Static upstream credentials only. Your calls to it still use its own key.
- Fits LangChain or LlamaIndex LangChain, LangGraph, LlamaIndex, Pydantic AI, CrewAI
- With a caveat MCP servers to govern Hosted MCP server — governs nothing on your side.
- Fits Nothing — plain Node or Python Change one base URL.
Reading this the other way round — pick what you already run and see every product scored against it.
The integration surfaces behind those answers
- Vercel AI SDK Documented integration, not a provider
It works with the AI SDK, but not by being a provider — read the integration docs rather than expecting a drop-in model factory.
Named:
@ai-sdk/openaiUse createOpenAI with the Respan base URL and explicitly call .chat(...) for Chat Completions.
- Cloudflare Workers Not documented
No Workers guidance either way. If you are edge-first, verify fetch-only compatibility yourself.
- Kubernetes Not documented
No Kubernetes story published.
- Terraform Not documented
No Terraform surface published. Configuration is API or dashboard work.
- Existing API gateway It is the API gateway
This product is the gateway. If you already run it for your other APIs, AI traffic becomes a plugin rather than a new hop.
Respan is the request gateway; custom upstream endpoints are configurable.
- Cloud identity Static provider credentials only
You paste static provider credentials into this product so it can reach upstream models. Your own calls to it still use its own API key, and those pasted secrets are yours to rotate.
Provider credentials are configured upstream; calls to Respan authenticate using a Respan API key.
- MCP Hosted MCP server
The vendor runs an MCP server you connect a client to. Useful for reaching this product from an agent, but it does not govern your other MCP servers.
Hosted Platform MCP exposes Respan data and management tools; this is not evidence of a general MCP proxy.
Gateway integrations documented separately from tracing integrations.
Respan SDKs plus compatible upstream SDKs.
Use the exact model ID and confirm its Credits/BYOK mode before calling.
Tracing integration does not by itself prove gateway API coverage.
Silence in the docs: 4 of the integration questions on this card have no published answer either way. That is recorded as undocumented, not as a no — but it does mean you would be verifying it yourself.
What it does well
- Credits and BYOK funding paths
- Ordered fallback and weighted routes
- Integrated tracing and online evaluation
Where it falls short
- Vendor reports 50–150 ms added gateway latency
- Self-hosting is enterprise and sales-led
- HIPAA offer conflicts with the public standard terms
Choose it when
Teams wanting gateway routing, tracing, evaluations and prompt management in one platform.
Look elsewhere when
You require a publicly reproducible self-host install or a confirmed zero-retention contract before evaluation.
Managed gateway: A hosted control plane in front of the models: routing, caching, guardrails, spend limits, and logs. You trade a little simplicity for governance.
How hard is it to leave?
Derived from six published facts, not from an opinion. The weights are fixed and the same for every product — see the arithmetic.
| What helps you leave | Points | Source |
|---|---|---|
| Works with standard OpenAI code Switching away is a base-URL change rather than a rewrite of every call site. | 22 /22 | vendor page |
| No vendor-specific SDK required A proprietary client library spreads through your codebase and has to be torn out again. | 10 /10 | vendor page |
| Can use your own provider accounts Your keys and billing relationship stay yours, so removing the gateway does not cut off model access. | 20 /20 | vendor page |
| Can be self-hosted You can run it yourself instead of accepting a pricing or policy change. | 20 /20 | vendor page |
| Configuration lives in version control Routing and budget rules are a file you keep, not dashboard state you would have to rebuild. | not published | vendor page |
| Your request history can be exported You leave with your own logs instead of abandoning them. | 12 /12 | vendor page |
Read the fine print: Compatible API and exportable telemetry reduce migration work; stored prompts and gateway controls remain service-specific.
1 of the 6 inputs is not published, so the highest reachable score here is 84 rather than 100. That is a gap in the public documentation, not a mark against the product — no points are deducted, they simply cannot be claimed. This measures technical switching cost only. It does not price the engineering time to re-test prompts against a different routing stack.
Respan models & pricing
Browse every imported listing from this provider, with published token rates and a link to compare other providers for the same model. This is provider-reported coverage; an absent listing does not mean unsupported.
Loading model listings…
Official model coverage source ↗ · Model source coverage and limitations
Full specification
Every field we track. Blank fields say "Not published" rather than "No" — we do not infer an absence from silence. Switch to Technical in the header for the precise field names and the low-level details.
Overview
Interpret these fields: Self-hosted vs managed LLM gateways · Who still owns your LLM gateway?
- What kind of product Category
- Managed gateway
- Marketplaces resell many providers behind one key. Gateways add governance on top. Open-source projects you run yourself. Cloud platforms are hyperscaler surfaces. Inference providers host models on their own hardware.
- Who runs it Deployment model
- Managed or self-host
- Managed means the vendor operates it. Self-host means you run it on your own infrastructure. Both means you can choose.
- Licence Licence
- Proprietary hosted service; separate SDK and MCP repositories have their own licenses.
- Proprietary products cannot be inspected or forked. Open licences such as MIT and Apache-2.0 let you audit, modify, and run the code without permission.
- Who you would be signing with Vendor status
- Not published
- Whether the product is still an independent company, has been acquired, is a large cloud vendor’s product line, is run by a software foundation, or has been put into maintenance mode. Maintenance mode means bug fixes and security patches only — no new features.
- Last shipped an update Latest release
- Not published
- The date of the most recent release or version tag. A product that has not shipped in a year is a different risk from one that shipped last week, regardless of what its marketing site says.
- GitHub stars GitHub stars
- Not published
- A rough proxy for community size on open-source projects. Not a quality measure.
Cost
Interpret these fields: How LLM gateway pricing works · LLM gateway spending limits: stop a runaway agent bill?
- Markup on model prices Token markup
- Not published
- How much the product adds on top of what the underlying model provider charges. Zero means you pay the same per-token price you would pay the model provider directly.
- Fee to add funds Credit purchase fee
- Not published
- A percentage charged when you top up your balance, separate from token prices. It is easy to miss because it does not appear on the per-token price list.
- Monthly cost per person Seat fee
- Not published
- A recurring per-user platform charge that applies regardless of how much you use the models.
- Can use your own provider accounts BYOK supported
- Yes
- Bring Your Own Key: you keep direct contracts with OpenAI, Anthropic and others, and the gateway only routes traffic. This preserves negotiated rates and committed-spend discounts.
- Cost of using your own accounts BYOK terms
- No single public BYOK fee established; account and plan terms apply.
- What the product charges to route traffic through your own provider keys.
- Free tier Free tier
- Free platform plan: 100,000 logs, 1,000 scores, five datasets, two evaluators and five prompts. Model usage is funded separately.
- What you can do without paying, useful for evaluation.
- Enterprise plan from Enterprise plan from
- Not published
- Annual entry price for the enterprise tier, where one is published or credibly reported.
- Cost to run it yourself Self-host cost
- Enterprise quote and customer infrastructure costs; no public self-hosted price.
- What self-hosting actually costs once you account for infrastructure and any paid tier.
- How the vendor makes money Pricing model
- Platform fee plus usage meters
- The shape of the vendor’s bill: does the routing layer charge a percentage on top of tokens, a flat monthly fee, both, neither (because inference is the product), or nothing at all (open source with no paid tier).
- How pricing works, briefly Pricing model detail
- Platform plans plus model usage funded with credits or BYOK. Team is advertised at $199/month billed annually.
- A one-paragraph description that covers the caveats a pricing category cannot: introductory rates, per-feature meters, tier gating, and pricing that resets on a specific date.
- Minimum commitment Minimum commitment
- The advertised $199/month Team rate requires annual billing; enterprise commitments are quoted.
- Whether the vendor requires a minimum contract term, a minimum spend, or a provisioned-capacity purchase to get its published rate.
- Charges that fire after you go over an allowance Overage terms
- Team lists $8 per additional 100,000 logs and $1 per additional 1,000 scores. Enterprise volume pricing is quoted.
- The line items that scale with usage after an included allowance is exhausted — log storage, extra requests, per-feature meters, data export — which is where cost estimates usually go wrong.
- Prompt cache offered Cache mechanism
- Both exact and semantic
- Whether the gateway offers its own response cache, what kind of match it does (exact request, prefix, semantic), or simply passes provider caching through unchanged.
- Discount on cached input Cache-read discount
- Not published
- How much cheaper cached tokens are than fresh input, when the vendor publishes a single figure. Bundled-inference clouds usually price this per model instead of as one number.
- Premium on cache writes Cache-write premium
- Not published
- How much more the first write of a cached prefix costs versus a plain input token. A high write premium and a low hit rate can leave you paying more than you save, so this matters as much as the read discount.
- Who captures the cache saving Cache economics
- Exact response caching plus separate Anthropic prompt-cache passthrough. No universal read/write discount applies.
- Whether the customer keeps the full saving from caching or the vendor captures part of it — and any conditions attached (write premium, storage fees, best-effort hits).
- What you can split spend by Cost attribution
- API key, customer_identifier and organization limits; request metadata supports attribution.
- The dimensions the vendor documents for splitting spend — per key, per user, per team, per tag, per customer. Matters if you need to chargeback internally or bill an end customer.
- How you get cost data out Cost export
- Export selected records and cost fields as CSV or JSONL.
- The mechanisms the vendor publishes for exporting cost and usage data: CSV, an API, webhooks, S3, a data warehouse, or nothing at all. Any per-unit price is included.
- Who pays the model bill BYOK mode
- Your keys or their credits
- Whether you bring your own provider accounts (BYOK), buy inference from this vendor, or can do either. This is the single biggest commercial difference between these products: it decides who holds the contract with the model provider and who carries the spend.
Spend governance in detail
Seven signals matter when a bill starts to hurt: who can spend, how much, on what, and who gets paged when it goes wrong. Everything below is drawn from the vendor’s own pricing and docs pages — how we read these.
- Virtual or scoped keys Yes
Keys that carry their own budget and rate-limit policy, so an intern experiment cannot spend against a production budget.
Temporary keys with limit policies; enforcement depends on environment configuration.
- Budget caps per key Yes
A dollar or token ceiling attached to an individual key. Where enforcement is soft, one over-limit request still completes before the block kicks in.
Lifetime and recurring caps; verify environment enforcement.
- Budget caps per team or workspace Not published
A ceiling applied at a higher scope than one key — a team, a workspace, a customer, or an entire environment.
Organization and customer scopes documented; arbitrary team budgets not established.
- Rate limiting as a cost control Yes
Configurable request-per-time-window caps. Platform-set rate limits do not count as spend controls; user-configurable ones do.
- Model allowlists Not published
A policy that constrains which models a key or team can call, keeping expensive frontier models out of the wrong hands.
- Spend alerts Yes
Alerts fired as spend approaches a threshold. Alerts that only fire after the meter has rolled over are marked as such.
- Webhook notifications Yes
Programmatic notifications on spend events, so budget breaches can page an on-call or open a ticket.
Enforcement: Enforced before each request
Catalog
Interpret these fields: LLM gateway model counts: what “500+” means
- Models available Models available
- 1,525
- How many models you can call, shown as a range because several vendors publish different totals on different pages. Vendor-reported either way, so counts are not directly comparable — some count every provider variant of the same model separately.
- Model providers reachable Upstream providers
- Not published
- How many distinct model providers or labs you can reach, shown as a range where the vendor’s own pages disagree. More providers usually means better redundancy when one has an outage.
- Works with standard OpenAI code OpenAI-compatible API
- Yes
- If yes, you can usually switch to it by changing one base URL, and switch away just as easily. This is the main defence against lock-in.
- OpenAI chat endpoint POST /v1/chat/completions
- Yes
- The endpoint almost every application ports first. "Not documented" means the vendor never states it, which is different from a documented no.
- Anthropic messages endpoint POST /v1/messages
- Yes
- Whether Anthropic-shaped calls work without rewriting them. Several products support this only as SDK compatibility or provider passthrough rather than a native endpoint — the detail page says which.
- OpenAI Responses endpoint POST /v1/responses
- Partly
- The newer stateful OpenAI surface. Support is much thinner across this market than chat completions.
- Embeddings endpoint POST /v1/embeddings
- Not documented
- Whether you can generate vectors through the same gateway, or need a second integration for retrieval workloads.
- Image generation endpoint POST /v1/images/generations
- Not documented
- Whether image models are reachable through the same surface as text.
- Audio endpoints POST /v1/audio/*
- Not documented
- Speech-to-text and text-to-speech. Frequently the first gap in an otherwise complete gateway.
- Batch jobs endpoint POST /v1/batches
- Not documented
- Asynchronous bulk processing, usually at a discount. Commonly undocumented, and commonly the reason a migration stalls late.
- Needs the vendor’s own code library Requires a vendor-specific SDK
- No
- A proprietary client library spreads through your codebase and has to be torn out again if you leave. “No” is the better answer here, and it means the standard OpenAI client works.
- You can export your request history Logs / usage data export
- Yes
- Whether you can get your own request logs, traces, or usage records back out — through an API, a bulk export, or a download. Decides whether you leave with your history or abandon it.
- Settings can live in version control Declarative config-as-code
- Not published
- Whether routing, fallback, and budget rules can be declared in a file you keep in Git, rather than existing only as settings clicked into a hosted dashboard.
- Embeddings Embeddings
- Not published
- Text-to-vector models, needed for search and retrieval features.
- Image generation Image generation
- Not published
- Whether image models are reachable through the same interface.
- Speech and audio Speech and audio
- Not published
- Text-to-speech or transcription models through the same interface.
- Video generation Video generation
- Not published
- Whether video models are reachable through the same interface.
- Batch processing Batch processing
- Not published
- Submitting large jobs for cheaper, slower processing. Often 50% off for work that is not time-sensitive.
Routing & reliability
Interpret these fields: How LLM gateway failover actually works · Does your LLM gateway promise any uptime? · Is routing destroying your prompt cache? · Changing models without breaking production
- Uptime it promises in writing Contractual SLA uptime
- Not published
- The uptime percentage in a published, contractual service level agreement. A public status page is not an SLA — it reports what happened, it does not promise anything or pay you back when it breaks.
- Automatic failover Automatic failover
- Yes
- When a model provider goes down or rate-limits you, traffic moves to a backup automatically instead of returning errors to your users.
- Load balancing Load balancing
- Yes
- Spreads requests across several providers or keys to raise your effective rate limit.
- Rule-based routing Conditional routing
- Not published
- Send different requests to different models based on rules — for example a cheap model for free users and a strong model for paying ones.
- Response caching Response caching
- Yes
- Reuses the answer when the exact same request comes in again, which cuts both cost and latency.
- Similar-question caching Semantic cache
- Not published
- Reuses an answer when a new question means roughly the same thing as an earlier one. Saves far more than exact-match caching but can return subtly wrong answers if tuned loosely.
- Where you set the timeout Request timeout surface
- Not documented
- Where a request timeout can be set: per request, in a config file, in the vendor dashboard, or nowhere. Nine of the twenty products documented here do not describe a request timeout at all, so the worst case of a hung upstream call is unknowable from the docs.
- Where you set retries Retry policy surface
- Per request
retry_enabled, num_retries and retry_after govern current-route attempts; UI also available.
- Where retry count and backoff are configured. Worth knowing alongside billing: a retried streaming call can be charged more than once.
- Where you set fallbacks Fallback surface
- Per request
fallback_models is tried in order after the current route exhausts eligible attempts; preflight and fail-fast errors stop the chain.
- Where the fallback chain is defined. Almost every product claims fallback; the useful question is whether you can change it from code or only by hand in a dashboard.
- Shape of the fallback chain Fallback shape
- Ordered list
- An ordered list tries targets in sequence; a weighted split sends a percentage of traffic to each, which is what you need to trial a new model on 5% of requests. Weighted splits are much rarer than the marketing implies.
- Upstream health tracking Health checks / circuit breaking
- Not documented
- Whether the product notices a failing upstream and stops sending traffic to it, and whether you can tune the thresholds. This is what turns a provider outage into a blip rather than a sustained error rate, and only four of the twenty expose it.
- Cross-region failover you control Multi-region failover surface
- Not documented
- Whether you can define what happens when a region degrades. A vendor running many regions is not the same as a vendor letting you configure failover between them; only three document a user-controlled mechanism.
- Where you set load balancing Load balancing surface
- Per request
Weighted model groups and weighted provider credentials may be configured in the dashboard or request.
- Where traffic distribution across upstreams or keys is configured.
Operations
Interpret these fields: LLM gateway observability: traces, logs and export · Running coding agents through an LLM gateway · Changing models without breaking production
- Usage dashboards and logs Observability
- Yes
- Built-in visibility into what was sent, what came back, what it cost, and how long it took.
- Spending limits Budget controls
- Yes
- Hard caps that stop spend before it becomes a surprise invoice. The single most valuable control for a small team.
- Rate limits Rate limits
- Yes
- Caps on request volume per key or per user, useful for protecting against abuse and runaway loops.
- Separate keys per team or app Virtual keys
- Yes
- Issue scoped keys with their own budgets and permissions so you can attribute cost and revoke access without rotating everything.
- Prompt versioning Prompt management
- Yes
- Store and version prompts outside your code so they can be changed without a deploy.
- Quality testing Evals
- Yes
- Built-in tooling to score model output against test cases, so you can tell whether a model swap made things better or worse.
- MCP support MCP support
- Yes
- Native support for the Model Context Protocol, the emerging standard for connecting models to external tools.
- What gets logged Logged content
- Full prompts and responses
Gateway logging includes request/response content and usage. Telemetry-only integrations are a separate path.
- Whether full prompts and responses are stored, only metadata, or nothing. Full-body logging is the most useful debugging feature here and the one most likely to need a conversation with your compliance team.
- You can turn logging off Body-logging opt-out
- Not documented
- Whether prompt and response bodies can be suppressed while still keeping usage metrics. Nineteen of the twenty document a way to do this; the mechanisms range from a per-request header to an organisation-wide setting.
- Traces you can take elsewhere Distributed tracing
- OpenTelemetry
Existing OpenTelemetry exporters can send spans to Respan; tracing SDK adds nested application spans.
- Whether the product emits OpenTelemetry, a proprietary format, or nothing. OpenTelemetry means the traces land in the tooling you already run instead of only in the vendor’s dashboard.
- Where telemetry can go Export destinations
- CSV, JSONL
One-time and recurring log exports with selectable fields, filters and sampling.
- Documented sinks for logs and metrics. This is a good proxy for how replaceable the vendor’s own dashboard is: around twenty destinations means you never have to depend on it, while a CSV download means you do.
- Can record user feedback Feedback capture API
- Not documented
- Whether there is an API to attach a rating or score to a logged request, which is what lets production traffic feed quality work later.
- Scores live traffic Online eval hooks
- Yes
Sampled production span or completed-trace scoring using deployed evaluator pipelines.
- Whether automated scorers can run against real production requests, rather than only against a test set you assemble yourself.
Performance
- Delay it adds Proxy overhead
- Not published
- Extra time the product itself adds to each request, on top of however long the model takes. Usually irrelevant next to multi-second model latency, but it matters for high-volume or streaming-sensitive workloads.
- Requests per second ceiling Throughput
- Not published
- Published sustained request rate before the product becomes the bottleneck. Only relevant at genuinely high volume.
- What the request path runs on Architecture class
- Undisclosed vendor service
AWS ECS API servers, Redis queues, Celery consumers, PostgreSQL and ClickHouse; hosted architecture as documented.
- The comparable way to talk about latency here. An edge worker, a compiled Go or Rust binary, and a Python proxy have different overhead floors no matter which figures each vendor publishes. Products that never disclose their runtime are recorded as undisclosed rather than assumed.
- You can run the request path yourself Self-hostable data plane
- Yes
Pricing lists enterprise self-hosting; no public gateway installation command was found.
- Whether the component that actually carries your prompts can run on your own infrastructure. Distinct from a vendor offering a self-hosted control plane while still proxying traffic through their network.
- Streaming responses Streaming support
- Yes
stream=True documented for Chat Completions.
- Whether token-by-token streaming is documented. The caveats matter more than the yes: some products cannot cancel a stream without still being billed, and several timeout and fallback mechanisms stop applying once the first token has been sent.
Security & compliance
Interpret these fields: LLM gateway compliance: SOC 2, HIPAA and evidence · How LLM gateway guardrails fail · Which LLM gateways store your prompts? · Do you need an MCP gateway as well?
- Does your prompt reach their servers Prompt transits vendor
- Yes
Managed gateway traffic passes through Respan before reaching the upstream provider.
- Whether the text you send passes through this company’s own infrastructure. If it does, every other promise on this page is a policy commitment rather than a physical impossibility. Self-hosted products can answer no outright.
- What they keep if you change nothing Logging default
- Yes — prompts and replies
Gateway requests are automatically logged; review retention and redaction before sending sensitive content.
- Defaults matter more than options. A product that stores full prompts and replies unless you find the right header will have stored them by the time you read the docs.
- How long they keep it Default content retention (days)
- Not published
Free: 7 days; Team: 30 days; Enterprise: custom. Confirm the policy configured for the project.
- Default retention for request content, in days. Zero means nothing is kept. Read the note: several products keep nothing as a rule but make timed exceptions for abuse review or specific models.
- Could they train on your prompts Training on customer data
- Not published — silence, not a no
No sufficiently specific gateway-wide no-training commitment established from reviewed public legal and compliance pages.
- Whether the vendor may use your prompts and outputs to train models. “Not published” means we could not find any position, which is not the same as a no — ask for it in writing.
- Where it runs, and what you can pin Region and residency control
- US East (Virginia), US West (Oregon); EU on request.
- Which regions are offered and whether you can force processing to stay in one. A global endpoint that silently picks a region is a different compliance story from an endpoint you pin yourself.
- Where safety filters run Guardrail execution location
- In the vendor’s cloud
Hosted Gateway pre-redaction happens after Respan receives the request. Telemetry redaction does not change upstream input.
- A filter that strips personal data only helps if it runs before the data leaves your boundary. If guardrails execute in the vendor’s cloud, the vendor has already received whatever you wanted redacted.
- Who else touches the data Subprocessor list
- Not published
- The published list of third parties the vendor passes your data to. No list means you cannot know the full chain, which most data-protection agreements require you to.
- SOC 2 audited SOC 2 audited
- Not published
- An independent audit of security controls. Enterprise buyers and their procurement teams routinely require it.
- Will sign a HIPAA agreement HIPAA BAA
- Not published
- Required before you may send protected health information through the service. Without a signed BAA, healthcare data is off limits.
- GDPR commitments GDPR commitments
- Yes
- Published data processing terms for handling personal data of people in the EU and UK.
- Can keep data in the EU EU data residency
- Yes
- Requests can be processed inside the EU rather than routed to US infrastructure. Often the deciding constraint for European customers.
- Does not retain your data Zero data retention
- Not published
- Prompts and responses are not stored after the request completes. Sometimes a paid add-on rather than the default.
- Strips personal data PII redaction
- Yes
- Detects and removes identifiers such as names, emails, and card numbers before the request reaches the model provider.
- Content guardrails Content guardrails
- Yes
- Policy checks on inputs and outputs — blocking unsafe content, enforcing formats, or catching prompt-injection attempts.
- Runs fully disconnected Air-gapped deployment
- Not published
- Can be deployed in a network with no internet access, which some regulated and defence environments require.
- Blocks personal data in prompts PII / DLP enforcement
- Can block the request
Pre-forward redaction masks detected entities; it transforms content rather than rejecting the entire request. Independent telemetry redaction; both disabled by default.
- Whether personal data detection sits on the request path and can stop the call, merely inspects and forwards it, or is not documented. A control that only reports is a logging feature, not a policy control.
- Blocks prompt injection Injection / jailbreak enforcement
- Not documented
- Whether injection and jailbreak detection can stop a request. Most products offering this call a partner classifier rather than shipping their own.
- Blocks harmful content Toxicity / moderation enforcement
- Not documented
- Whether hate, violence, sexual and self-harm categories are checked inline and can stop a request, in either direction.
- Your own policy rules Custom policy hooks
- Not documented
- Whether you can add your own rule — a regex, a webhook, or your own classifier — rather than choosing from the vendor library.
- Where guardrails run Guardrail execution location
- On the vendor's servers
- Whether guardrail evaluation happens inside your infrastructure or on the vendor’s servers. This decides whether the prompt you are trying to protect leaves your network in order to be checked.
- If the guardrail itself fails Guardrail failure mode
- Not documented
- What happens when the guardrail service times out or errors: does the request proceed unchecked, or is it blocked? This is the worst-documented field in the entire catalogue — only two vendors state it plainly, which means most teams are running a control whose failure behaviour they cannot know.
- Third-party guardrail vendors Guardrail integrations
- Not published
- Named external guardrail services the product can call. A long list means the product is a router for policy engines rather than a policy engine itself — which also means another vendor bill and another hop.
Compliance evidence
Graded by how strong the evidence is, not whether the word appears on the vendor’s website. An audited report and a marketing claim are different things, and only one of them will satisfy your own auditor.
- SOC 2 Claimed, no evidence published Vendor states Type II; report under NDA, not independently examined.
- ISO 27001 Claimed, no evidence published Badge on gateway marketing; certificate not reviewed.
- GDPR DPA Available on request DPA by request.
- HIPAA BAA Vendor’s own pages disagree Compliance docs offer a BAA, but public Terms §1 disallows HIPAA-regulated use. Obtain a superseding signed agreement before relying on the offer.
- FedRAMP Not published
- ITAR Not published
Fit & integration
Interpret these fields: How much does an LLM gateway lock you in? · Do you need an MCP gateway as well? · Running coding agents through an LLM gateway · Changing models without breaking production
- Work to try it Evaluation work shape
- Change one base URL
- The shape of the work on the vendor’s own quickstart, from swapping one base URL through to deploying infrastructure. An ordinal class rather than a duration, because elapsed time depends on accounts and quota we cannot see.
- Work to run it Production work shape
- Change one base URL
- The same scale applied to the vendor’s recommended production path. For several products this is much heavier than the quickstart, which is exactly why both are recorded.
- Steps on the quickstart Numbered quickstart steps
- 3
Three setup steps before the optional prompt-management step; this is not a timed integration test.
- A literal count of numbered steps on the vendor’s quickstart, recorded as evidence beside the work shape. Zero means the page publishes no numbered procedure at all. Large counts usually mean interleaved language tracks rather than more work.
- Can you self-host it today Self-host install documentation
- Offered, but no command published
- Whether an install command is actually published. Several products advertise self-hosting while publishing no command to start from, which a plain yes/no would hide.
- Works with the OpenAI SDK OpenAI SDK drop-in
- Yes
Chat Completions base URL and key replacement; provider-specific and Responses routes need their documented setup.
- Whether an existing OpenAI-compatible client can be pointed at it by changing the base URL and key.
- Vercel AI SDK support AI SDK provider package
- Documented integration, not a provider
Use createOpenAI with the Respan base URL and explicitly call .chat(...) for Chat Completions.
- Whether a first-party AI SDK provider package exists, or only a community package, a documented workaround, or the generic OpenAI provider pointed at a custom base URL.
- Python framework integrations Documented Python frameworks
- LangChain, LangGraph, LlamaIndex, Pydantic AI, CrewAI
Gateway integrations documented separately from tracing integrations.
- LangChain, LangGraph, LlamaIndex and similar orchestration frameworks with a documented integration.
- Callable from Cloudflare Workers Cloudflare Workers support
- Not documented
- Whether the docs show calling this product from your own Worker. Deliberately separated from the several products whose own gateway runs on Workers, which is a fact about their infrastructure and not about your edge compatibility.
- Kubernetes install Helm chart availability
- Not documented
- Whether a named, published Helm chart exists, versus Helm being referenced with no chart named, versus generic cluster documentation that has nothing to do with this product.
- Terraform support Terraform provider or modules
- Not documented
- Whether you can declare this in code: an official provider, official modules, resources inside a hyperscaler’s provider, a community provider, or only Terraform code shipped in a repo.
- Reuses your cloud identity Cloud IAM reuse
- Static provider credentials only
Provider credentials are configured upstream; calls to Respan authenticate using a Respan API key.
- Whether you can authenticate with IAM roles, workload identity or managed identities instead of another long-lived API key. Distinguished from products that only accept static upstream provider credentials.
- Fits behind your API gateway API gateway integration
- It is the API gateway
Respan is the request gateway; custom upstream endpoints are configurable.
- Whether AI traffic can go through a gateway you already run, and who documented that — the product’s vendor or the gateway’s.
- MCP support MCP surface shape
- Hosted MCP server
Hosted Platform MCP exposes Respan data and management tools; this is not evidence of a general MCP proxy.
- Which kind of MCP support this is: a gateway that governs many MCP servers, a hosted MCP server you connect to, MCP tools accepted inside the completion API, or client tooling. These are different products behind one acronym.
- Needs your own provider key Upstream provider key required
- Optional
Some catalog listings require BYOK. Credits-supported routes do not require your own provider key.
- Whether an upstream provider account and key must exist before your first call works. A prerequisite rather than a step, and it can differ between the hosted and self-hosted forms of the same product.
- Gate before models work Model access gate
- Not documented
Active status does not guarantee account access. Catalog labels distinguish Credits from BYOK.
- Whether an enablement click, a quota grant, a paid tier or an approval form stands between a valid key and a working model call.
- Official client languages First-party client SDK languages
- Python, TypeScript
Respan SDKs plus compatible upstream SDKs.
- Languages with a first-party client library. An empty list can still mean the product is usable from any language via an OpenAI-compatible SDK.
Additional charges
These are the fees that do not appear on a per-token price list, and they are where estimates usually go wrong.
- Additional Team seats $15/member; five members included
How pricing actually works
Enterprise quote and customer infrastructure costs; no public self-hosted price.
Common questions
Answered from the fields above, so these move when the catalog moves. Every figure quoted here appears in the specification with its source.
Does Respan charge a markup on model prices?
Respan does not publish a token markup figure. Other charges on the page: additional team seats ($15/member; five members included). A blank here means the vendor states no figure, not that the service is free — the pricing page it was checked against is linked in the fees section below.
Can Respan be self-hosted?
Yes. Respan can be run on your own infrastructure or used as a managed service. The licence is Proprietary hosted service; separate SDK and MCP repositories have their own licenses. Running it yourself means you supply the infrastructure and the upstream model accounts, so the bill is your own hosting plus the providers' own rates.
Is Respan SOC 2 audited, and will it sign a HIPAA BAA?
Respan publishes neither a SOC 2 report nor a HIPAA business associate agreement. It offers a GDPR data processing agreement and EU data residency. Each of these is linked to the vendor's own page in the compliance section below. Ask for both documents directly before committing, because a published claim and a countersigned agreement are not the same thing.
Does Respan retain your prompts?
Respan does not publish a zero-data-retention position. Prompt and response bodies are logged by default. Retention, logging and training on customer data are three separate questions, and a vendor can answer one of them without answering the others. Each is listed with the page it was read from in the data-handling section below.
Can you use your own provider keys with Respan?
Yes. Respan can route through your own accounts with the underlying model providers, so inference is billed to you directly. No single public BYOK fee established; account and plan terms apply. That splits the bill in two: the model providers invoice you for inference directly, and the gateway charges only for sitting in front of it.
How many models does Respan support?
Respan publishes no aggregate total; counting its catalog gives 1,525 models. Count of entries returned by the models API on 2026-09-23. Includes every entry exposed by that endpoint; not a count of unique base models. The figure on this page is dated and carries its source.
Official links
Independent coverage
Third-party analysis, walkthroughs and operator threads. We link criticism as readily as praise. Nothing here is written or published by the vendor, by a competitor listed on this site, or by an SEO content farm — how we vet these.
Nothing yet that clears our bar. Everything we found was written by the vendor, by a competitor, or by an SEO content farm, so we would rather show you nothing than pass marketing off as a review.
What has changed here
- Models available Models available 1282 1525 source ↗
- Models available Models available Not published 1282 source ↗
Read the head-to-head
These pairs have a written verdict, not just a table.