Guide
Changing models without breaking production
Treat a model change as an application release. Inventory the resolved model and its dependent features, test the replacement on representative tasks, roll it out to a bounded cohort, and keep a viable rollback target. A gateway rewrite changes routing; it does not prove that the new model preserves the old model’s behavior.
For teams handling model upgrades, alias changes and retirement deadlines. Vercel routing rules and Anthropic’s lifecycle documentation are examples; the release procedure applies to the specific model/provider combination you operate. Primary documentation reviewed 18 September 2026. Vendor capabilities below are documented, not independently benchmarked. The evaluation procedures are GatewayScore recommendations.
Inventory what actually changes
Start with the application-facing model name and resolve it to the provider deployment or model version that serves requests. A pinned identifier can make a selection more explicit, but it does not make that model available forever. Record any alias, virtual model, rewrite or fallback between the application and the destination.
| Dependency | Record before migration | Replacement acceptance question |
|---|---|---|
| Model identity | Requested name and resolved provider/version | Which destination actually serves the canary? |
| Protocol | Endpoint, SDK, headers and options | Does the new target accept the same contract? |
| Tools and outputs | Tool schemas and structured-output requirements | Do results remain usable by downstream code? |
| Context and caching | Context limits, truncation and cache assumptions | Do long tasks and repeated sessions still work? |
| Data controls | Allowed providers/regions and logging settings | Do the same restrictions apply to the replacement? |
| Operations | Quota, budget, owner and fallback | Can traffic be recovered within the required deadline? |
Include background jobs and rarely used workflows. A model retirement can break a monthly report even when interactive traffic has already moved. Search configuration and inspect resolved traffic together; either view alone can miss an indirect dependency.
Deprecation and retirement are different states
Anthropic’s model lifecycle documentation distinguishes a deprecated model that still functions from a retired model whose requests fail. It also notes that partner-operated platforms can have different schedules. Record the lifecycle date for the actual platform serving your workload, not just the model family name. Anthropic model deprecations.
Set an internal migration deadline before the provider retirement date and leave time for an unsuccessful first candidate. Assign a replacement owner and an escalation path. A calendar reminder is useful, but a recurring inventory of active traffic is what finds workloads that continue requesting the retiring model.
If the old model will disappear, rollback cannot mean “switch back” after that date. Qualify another supported destination or a limited service mode in advance. A reversible configuration file is not an executable rollback if its target no longer exists.
Understand the scope of a central rewrite
Vercel documents team-wide model rewrite and deny rules. Its current documentation says changes can take time to propagate and that in-flight requests finish under the previous configuration. It also warns that provider-specific options are not translated across providers. A central rule can therefore alter many callers while leaving options inappropriate for the new target. Vercel routing rules documentation.
Choose a rollout mechanism whose scope matches the canary. If a rule affects all team credentials, it is not automatically a percentage rollout. Use a genuinely isolated route, credential scope or application cohort supported by your architecture. Record requested and resolved targets so you can tell which configuration handled each request.
Avoid a chain of invisible rewrites as a substitute for a migration plan. Document the effective route and test the fallback candidates against the same output, security and cost requirements as the primary replacement.
Define gates before looking at the new answers
- Freeze a representative held-out task set, including difficult inputs, long context, tools, structured outputs and refusal cases relevant to the application.
- Run the current route and the candidate with the same application configuration except changes required and documented for the new model.
- Check machine contracts first: parseability, required fields, tool result association, stream completion and error handling.
- Evaluate task success and critical mistakes separately from style. Review grader disagreements and preserve failed attempts.
- Measure complete-task cost and tail latency, including retries and any extra routing or judging work.
- Approve a canary only when predeclared thresholds pass; otherwise narrow the scope, adjust the application or reject the candidate.
Do not require byte-for-byte wording equality from a probabilistic system unless that is genuinely the product contract. A migration can change phrasing while preserving utility, or preserve fluent phrasing while breaking factual accuracy. Test the outcome that the application actually depends on.
Use a release record and explicit stop conditions
| Stage | Evidence required | Stop or rollback trigger |
|---|---|---|
| Offline evaluation | Versioned tasks, contract results and quality review | Required contract failure or unacceptable critical error |
| Internal canary | Resolved-route trace and representative tasks | Unexpected destination or broken data-control boundary |
| Limited production cohort | Task outcomes, latency and total cost | Predeclared quality, cost or reliability limit exceeded |
| Expanded rollout | Enough observations across important categories | A segment regresses despite healthy overall averages |
| Retirement of old route | No remaining required traffic and tested recovery path | Unmigrated jobs or unavailable recovery target |
For example, a team could require every mandatory schema test to pass, no new critical authorization failure and an explicitly chosen tolerance for non-critical task success. Those are suggested gate types, not universal numeric thresholds. Document the actual thresholds, observation window and decision owner before the canary starts.
Keep the release record compact: source and destination, route revision, evidence bundle, cohort definition, start time, monitoring links, stop conditions and rollback command or procedure. Run the rollback in a safe environment and verify its destination, not merely that the command exits successfully.
Plan for sessions and work already in progress
A model change can occur while a multi-step agent is using tool results or provider-specific state. Decide whether existing sessions stay on the old route, restart with an approved transcript, or stop with a recoverable error. Confirm the new model can consume the state you intend to transfer. Do not promise seamless continuation just because both endpoints accept messages.
After rollout, reconcile requested aliases with resolved destinations and look for stragglers. Keep the old route only as long as its availability and controls remain acceptable. Update the dependency inventory and schedule the next review rather than leaving emergency fallbacks undocumented.
Use the portability guide to identify translation dependencies, failover for recovery mechanics, and model-count interpretation before assuming a catalogue entry establishes deployable compatibility.
Evidence and maintenance
Review after any provider retirement notice, alias resolution change, SDK update or new provider option. Rehearse recovery before the old route becomes unavailable and retain the evidence for the model version actually evaluated.
This guide uses the selected sources cited above, not a census or ranking of every gateway. Undocumented behavior remains unknown. Review the exact product edition, deployment and configuration before relying on a control.
Common questions
Does pinning a model version prevent retirement?
No. It makes selection explicit while the version is available; the provider can still retire it. Track the platform-specific lifecycle.
Is a model rewrite a canary rollout?
Only if its scope is actually limited to the intended cohort. A team-wide rewrite can change every caller at once.
Can rollback target a retired model?
No reliable recovery plan should depend on a model that is no longer available. Qualify a supported alternative or a reduced service mode.
Should existing agent sessions switch models immediately?
Decide based on state compatibility and application requirements. Test continuation or restart explicitly rather than assuming messages and tool state transfer without loss.
Sources and editorial review dates are recorded with this guide. Catalogue-backed blocks use the dated catalogue snapshot. See our methodology.