Guide

Is routing destroying your prompt cache?

Short answer

Routing can reduce prompt cache hits when related requests reach different cache scopes or arrive with changed prefixes. It does not necessarily delete a cache. Keeping a conversation on a compatible, warm route can help, but endpoint affinity alone cannot guarantee reuse. Compare upstream cached tokens and total task cost before deciding that the cheapest advertised input rate is the cheapest route.

Which cache is the router affecting?

Separate the work being reused before interpreting a dashboard’s “cache hit.” In this guide, prompt caching means reusing computation for a matching input prefix. The model still generates the response. A gateway’s exact-response cache instead returns an earlier answer for a matching request; a semantic-response cache uses a similarity rule. Those are different mechanisms with different correctness checks.

A gateway may send two requests to the same public API while the upstream service assigns them to different internal machines. Conversely, changing a gateway worker need not change the upstream cache scope. The useful question is whether the next request can reach a matching, eligible entry—not whether every hop is identical. OpenAI explicitly documents machine-local cached state and internal routing. OpenAI cache location.

How do you tell routing misses from prompt changes?

Start with a fixed model, a fixed reusable prefix and a known working cache configuration. Change one routing variable at a time. The checks below are an editorial diagnostic procedure, not results from a GatewayScore benchmark.

Diagnosing lost prompt cache reuse
Observed patternPossible explanationNext check
Hits fall after endpoint or model changesTraffic is reaching a separate or cold cache scope.Compare actual serving model, provider, region and account scope for each attempt.
Same route, new missesA prefix, tool definition, cache setting or lifetime changed.Compare the outgoing request after gateway transformation, not only the application template.
Long gaps lose reuseThe entry may no longer be eligible.Repeat inside and outside the documented lifetime for that model and API.
Many requests hit but cost barely fallsOnly a small fraction of input tokens is being reused.Measure token-weighted reuse and paid cache writes, not just hit requests.
Direct API hits; gateway calls do notTranslation or parameter support may differ.Check preserved cache controls and returned usage fields with a non-sensitive fixture.

Place stable instructions and reusable material before request-specific content when the API permits it. Do not insert a changing timestamp ahead of otherwise reusable text. This is a prefix-layout recommendation, not permission to change instruction priority or omit context that the model needs.

Which provider rules matter now?

These are four documentation examples reviewed on 17 September 2026, not a ranking or an exhaustive compatibility table. API family, model version and serving platform matter.

OpenAI: cache-key advice depends on the model

OpenAI isolates caches across organizations and regional processing boundaries. For models before GPT-5.6, it recommends a stable prompt_cache_key for related prefixes; the key influences routing without guaranteeing a hit. For GPT-5.6 and later, routing is automatic and the key is optional for separate cache accounting. Do not apply older key-tuning advice to every model. OpenAI prompt caching documentation.

Claude: configuration changes can matter as much as routing

Claude’s documented scope depends on its serving platform: the Claude API uses workspace isolation, while Bedrock and Google Cloud use organization isolation. Thinking configuration and output_config.effort changes can invalidate reusable prefixes. Cache writes can carry a premium, and the response separates cache-read and cache-creation input tokens. Check the selected model and platform rather than assuming a shared cache behind every Claude endpoint. Claude caching, isolation and usage documentation.

OpenRouter: affinity has explicit exceptions

OpenRouter documents sticky provider routing for caching, with fallback when the provider is unavailable. A manual provider.order takes precedence. Stickiness expires after ten minutes of inactivity; that is a routing-session lifetime, not a promise about the upstream cache. An explicit session_id can establish affinity from a successful request before a cache hit is observed. Preserve the session identity across related turns, then verify actual cached usage. OpenRouter provider sticky routing.

Gemini: identify the API before choosing cache controls

Google documents implicit caching for Gemini 2.5 and newer models, with model-specific minimum input lengths. Its Interactions API supports implicit caching only; explicit cache objects require the generateContent API. “Gemini supports caching” therefore does not establish which controls a gateway integration can pass through. Gemini context caching documentation.

Routing inference: treat a switch between independent providers, models or documented isolation scopes as potentially cold unless you have evidence of reuse. None of these examples establishes portable cache state across arbitrary gateways. Inspect a shortlisted product’s routing evidence on its provider page, then test the specific integration.

Can a cheaper route increase the input bill?

Yes. Here is a hypothetical calculation, not vendor pricing. Both routes process one million total input tokens. Assume identical output cost and quality, no gateway fees, no storage charge and no separate cache-write premium. Cold input is charged at the regular rate. Cache-read share is measured across all input tokens, including dynamic text.

Hypothetical input cost for one million tokens
Input measureRoute A: warmRoute B: cold
Regular rate per million tokens$3.00$2.00
Cache-read rate per million tokens$0.30Not needed: zero reads assumed
Cache-read share of all input80%0%
Input cost0.2 × $3 + 0.8 × $0.30 = $0.841 × $2 = $2.00

Route B’s regular input rate is one-third lower, yet its input bill is about 2.38 times Route A’s in this scenario. A wins once its cache-read share exceeds (3 − 2) / (3 − 0.30), or approximately 37.04%, provided B remains fully cold and the other assumptions hold.

This is not a claim that load balancing always costs more. B could warm up, deliver faster responses or require fewer retries. With separately billed writes, calculate uncached input, cache reads and cache writes at their respective rates without double-counting tokens. Add output, storage, gateway fees and failed attempts. The decision metric is cost per successful task at acceptable quality and latency. Our gateway pricing guide explains the wider fee stack; the cost calculator does not model this workload’s cache reuse.

How should you test your router?

  1. Define a representative workload. Fix the intended model version, prompt template, tool definitions, task mix and acceptance criteria. Use a non-sensitive test fixture; preserve production-like prompt lengths and turn spacing.
  2. Establish a baseline. Send the fixture through one allowed endpoint with documented cache settings. Confirm upstream cache-read tokens on repeated eligible requests. If this fails, routing comparisons cannot isolate the problem.
  3. Compare policies fairly. Test fixed-endpoint routing, the proposed balancing policy and conversation affinity. Record initial cold requests separately from steady-state traffic. Use disjoint test prefixes or allow documented expiry so one policy does not inherit another’s warm cache.
  4. Record every attempt. Capture the resolved model and provider, allowed region/account scope, template version, timestamps, latency, status and upstream usage. Include retries and fallback attempts. Use an access-controlled identifier or keyed fingerprint for prefix comparison instead of logging raw prompts.
  5. Normalize usage correctly. Compute total cache-read tokens divided by total input tokens for the workload, not the average of per-request percentages. API fields have different meanings: Claude separates ordinary input from cache reads and writes; OpenAI reports cache details within total input. Map the selected API’s accounting before adding fields. Claude usage accounting; OpenAI response usage.
  6. Exercise failure and load. Repeat at expected concurrency, after realistic idle periods and during a controlled endpoint failure. Compare time to first token at the median and 95th percentile, error rate, output quality and cost per completed task. A serial warm-cache demonstration is not evidence of production performance.

Accept a routing change only if it meets your predeclared cost, quality and reliability targets. Missing upstream cache telemetry means unknown reuse, not a zero hit rate. Keep the workload definition, run time, API versions and usage export with the result so another engineer can reproduce it. See gateway observability for trace and export considerations.

What routing policy should you try first?

For repeated long prefixes, start by testing conversation affinity within the model, deployment and data boundary your application permits. Retain health and latency escape conditions. Keep stable prefixes stable, verify cache parameters after translation, and version prompt changes deliberately. This is an editorial starting point—not a vendor-independent API recipe.

For short, unrelated requests, affinity may offer little reuse. For sparse traffic, eligible entries may expire before the next turn. For overloaded endpoints, waiting for a warm route may cost more than taking a cold fallback. Use your measurements to choose the tradeoff; do not disable a necessary recovery path to improve a cache-hit chart. The failover guide covers the recovery boundary.

Common questions

Does switching providers delete my prompt cache?

A switch can send the request somewhere that cannot reuse the original cached prefix. That is a cache miss on the new route, not evidence that the original entry was deleted. Returning to the original route may recover reuse if the prefix still matches and the entry remains available.

Does sticky routing guarantee a cache hit?

No. Affinity helps related requests reach the same provider endpoint, but prefix matching, model support, cache scope, lifetime and the provider’s internal routing still matter. Measure upstream cached tokens rather than treating an unchanged endpoint as proof of a hit.

Is prompt caching the same as semantic caching?

No. Prompt caching reuses computation for matching input prefixes while the model generates an answer. Semantic response caching returns a previously stored answer for a sufficiently similar request. Their hit counters measure different operations.

Should I turn off failover to preserve the cache?

Keep a tested recovery path. A warm route is useful only while it meets your availability, latency, quality and data-handling requirements. Compare the total cost per successful task, including retries, against the cache benefit.

What is the first metric to check?

Compare cache-read tokens divided by total input tokens across the same representative workload, grouped by actual model and serving endpoint. Then reconcile cache writes, uncached input, output and gateway fees against billed cost. A request-level hit percentage alone can hide how few tokens were reused.

Review method. The provider-specific statements above come from the linked primary documentation checked on 17 September 2026. The diagnostic procedure and routing recommendation are editorial synthesis. The dollar amounts are fully specified hypothetical inputs; no production cache-hit, latency or savings benchmark was performed. Recheck API behavior before implementing a policy.