Guide

Running coding agents through an LLM gateway

A coding agent needs more than a successful chat request through a new base URL. Validate its API format, streamed events, tool loop, model selection, context handling and credentials as a complete workflow. Keep client-vendor support boundaries separate from a gateway vendor’s compatibility claim.

For platform teams rolling out existing coding clients through a gateway. Claude Code and OpenRouter Ori are concrete documentation examples; this is not a claim that all clients or model families are interchangeable. Primary documentation reviewed 18 September 2026. Vendor capabilities below are documented, not independently benchmarked. The evaluation procedures are GatewayScore recommendations.

Use a complete coding task as the compatibility unit

A coding client is an application with its own assumptions about model behavior and protocol features. It may count context, discover models, send tool definitions, resume conversations and inspect stream events. A gateway that passes a simple completion can still lose a capability needed later in the same session.

Start with a disposable repository task: inspect files, propose a small change, apply it, run a meaningful check and explain the result. Capture failures at each stage. Avoid a smoke test that only asks for a greeting; it says little about the agent’s working loop. Record client, gateway, SDK and model versions with the test.

Compatibility contract to verify for the chosen client
SurfaceTestFailure symptom
API dialectUse the client’s required endpoint and schemaA chat endpoint works but the client cannot start
StreamingComplete a tool-producing streamed responseMissing events or unfinished tool arguments
Tool cycleExecute and return multiple tool resultsAgent stalls, repeats or mis-associates outputs
Model resolutionInspect the actual resolved modelAlias silently selects an unintended capability set
Context handlingExercise counting and a long sessionUnexpected truncation, estimates or rejection
CredentialsRevoke one developer without affecting othersShared-key offboarding breaks the whole team

What Claude Code’s own documentation requires

Claude Code’s gateway compatibility guide documents supported API formats and the headers/body fields that must pass through together. For an Anthropic-format route, preserving capability headers matters; token-counting support is optional but affects whether counts are exact. Use the guide for the deployed client version instead of assuming generic OpenAI-compatible chat is sufficient. Claude Code gateway compatibility guide.

Anthropic’s gateway overview explicitly does not support routing Claude Code to non-Claude models through a third-party gateway. OpenRouter’s Ori announcement describes its own configurations for several harnesses and model-dependent adjustments. Those statements describe different support boundaries. If evaluating such a combination, record who supports it and test it as that combination. Claude Code gateway overview; OpenRouter Ori Harness announcement.

A working community or vendor integration can still be useful. The decision is whether its support model, upgrade cadence and tested behavior meet your team’s needs. Do not relabel vendor-provided configuration as an endorsement by the client developer.

Issue credentials and configuration deliberately

Keep provider credentials at the intended server boundary and issue attributable gateway credentials where that architecture permits. Record which account pays for traffic. A developer’s existing subscription and a gateway’s upstream API billing arrangement are not automatically the same thing; verify the active credential path for the particular client.

Distribute a reviewed configuration through your team’s established settings and secrets tooling. Pin the gateway URL, approved model aliases and required compatibility options. Avoid committing secrets to repository config or relying on each developer to copy a long command sequence accurately. A changed environment variable can select a different billing or policy path.

Validate offboarding and rotation on one test identity. Check local clients, remote development environments and CI separately. An agent running in CI may need a workload credential instead of a human token; use a distinct identity and budget so its activity remains attributable.

Preserve meaning as well as JSON shape

When a gateway translates APIs, compare more than field names. Inspect streamed tool arguments, tool identifiers, result ordering, stop conditions, reasoning-related fields and any provider-specific context behavior that the client depends on. An unsupported feature should have a clear failure or documented degradation; silently dropping it makes troubleshooting much harder.

Model aliases deserve a test of their own. Save the requested alias and the resolved model/provider with each run. Repeat the same task after an alias update to detect behavior changes. If a client enables features based on a model name, a custom alias can also affect what it sends; verify the client’s actual requests rather than assuming the alias is purely cosmetic.

Long sessions can make small differences expensive. A retry that replays context, a lost cache hint or repeated failed tool call can increase cost without producing more useful work. Use prompt-cache routing and agent spending limits alongside protocol testing.

A small rollout suite that catches useful failures

  1. Choose a direct-provider baseline, if supported, and the candidate gateway route. Use the same model version and harmless repository fixtures.
  2. Run a code-reading task, a tested edit, a multi-step tool task and a long-context task. Compare actual artifacts and test outcomes, not only summaries.
  3. Interrupt a stream and restart a session. Confirm the client recovers without duplicating writes or losing the association between tool calls and results.
  4. Test denied models, exhausted budget and revoked credentials. Confirm actionable errors reach the developer and controls cannot be bypassed by a fallback.
  5. Check the logs for verified identity, resolved target, usage and error correlation. Confirm configured retention and redaction rather than assuming code snippets are excluded.
  6. Roll out to a small team cohort with a documented fallback configuration. Re-run the suite on client and gateway upgrades.

Keep a compatibility ledger with pass, fail and not-tested states for each version combination. One successful user report is evidence for that run, not proof that every feature or future release works.

When the gateway is ready for a team

Approve the route when the required tasks pass, identity and cost controls hold, and failures have a clear operating procedure. Keep unsupported features explicit. A team that depends on a first-party-only capability may need to retain the native path for that workload rather than disguise the incompatibility behind an alias.

For the broader migration boundary, read gateway lock-in and portability. For retained source code and traces, use prompt logging. This checklist concerns the chosen coding workflow; it does not rate a gateway’s entire model catalogue.

Evidence and maintenance

Retest on client releases, gateway translation changes, model alias updates or changed capability flags. Preserve the exact tested versions and support owner; avoid maintaining a timeless list of supposedly compatible model names.

This guide uses the selected sources cited above, not a census or ranking of every gateway. Undocumented behavior remains unknown. Review the exact product edition, deployment and configuration before relying on a control.

Common questions

Is changing the base URL enough?

It may establish connectivity, but you still need to test the client’s endpoints, streaming, tool loop and context features through the selected gateway.

Does gateway support for a coding client imply the client vendor supports every model?

No. Check both vendors’ statements. An integration claim and a client-vendor support commitment are different.

Should a team share one gateway key?

Prefer attributable, independently revocable credentials where supported. A shared credential makes per-user controls and offboarding harder.

When should compatibility tests run again?

After a client, gateway, model, alias or relevant feature-flag change, and before expanding a rollout to a new environment.

Sources and editorial review dates are recorded with this guide. Catalogue-backed blocks use the dated catalogue snapshot. See our methodology.