Skip to content
From API Relays to AI Gateways

AI Coding Costs Out of Control? From API Relays to AI Gateways

The model name stayed the same. Why did the experience change?

A team points Claude Code, Codex, Cursor, and an internal Agent at one “OpenAI-compatible” address. Everyone still requests the same model alias. A week later, three things happen:

  • the bill is lower, but nobody can reconcile it with the official provider;
  • coding quality varies sharply by time of day;
  • long tasks become slow, tool calls fail more often, and nobody knows which upstream model actually answered.

The address between the coding tool and the model is not a neutral pipe. It can authenticate users, rewrite requests, select an upstream account, choose another model, queue traffic, retry failures, record prompts, recalculate usage, and rewrite the response. That power can create a trustworthy control point—or a very convenient black box.

This article explains both sides. We will use LiteLLM, Sub2API, and AxonHub to inspect the machinery, then build a governance policy that makes routing and degradation visible.

1. Yes: a gateway is a control system in the model API path

The simplest mental model is:

AI Coding client → gateway or relay → upstream model API

But a production gateway has three different jobs. Mixing them together is the source of much confusion.

An AI Gateway has a data plane, control plane, and evidence plane

The data plane handles every live request

It terminates the client connection, checks the downstream key, translates the protocol, selects an upstream destination, streams the response, and handles retries. Because the plaintext request must be processed here, the gateway can normally see your prompts, source snippets, tool definitions, and model output.

The control plane decides what should happen

It stores model aliases, provider credentials, team permissions, budgets, rate limits, routing weights, fallback rules, and data policies. A route called coding-standard might map to two deployments today and a different pair next month without changing every developer laptop.

The evidence plane records what did happen

It should connect an internal trace ID to the caller, requested route, actual upstream deployment, latency, token usage, retry path, policy decision, and cost. OpenAI, for example, returns an x-request-id, publishes rate-limit headers, and lets clients supply an X-Client-Request-Id; retaining both sides of that correlation makes troubleshooting much stronger than a dashboard total alone. See the official API request debugging guidance.

The gateway is therefore not merely “one URL for many models.” It is an enforcement point plus a source of operational evidence. If the same operator controls routing, metering, and the only available logs, however, the evidence is not independent.

2. “Relay” describes a position, not a level of trust

In Chinese developer communities, zhongzhuan zhan—an API relay—can describe very different systems:

FormWhat it usually doesWho owns the upstream credential?Main valueMain question
Protocol adapterConverts one API shape into anotherYouClient compatibilityWhat fields are lost or rewritten?
Self-hosted gatewayCentralizes your own providers and policyYou or your organizationControl, audit, portabilityCan your team operate it securely?
Managed enterprise gatewayRuns governance as a contracted serviceYou and/or the vendorLess operational burdenWhat are the retention, routing, and audit guarantees?
Subscription-quota distributorPools product subscriptions or accounts and issues downstream keysUsually the operatorQuota sharing and lower apparent priceIs this use authorized, stable, and traceable?
Opaque resellerSells a model name through an unknown supply chainUnknownConvenience or discountIs the model, usage, and data handling truthful?

Open source does not automatically make a deployment trustworthy. It lets you inspect what the software can do; it does not prove how a stranger configured their server, what code they changed, or which upstream accounts they use.

Likewise, “commercial” does not automatically mean bad. A managed gateway with a clear contract, provider invoices, data-processing terms, exportable logs, and auditable routing can be safer than a poorly maintained self-hosted proxy.

3. Three open-source projects reveal three different centers of gravity

These projects overlap, but their stated purposes expose the major building blocks.

LiteLLM: normalize providers and govern access

LiteLLM presents a unified interface for many providers and can run as a central proxy. Its documentation lists protocol translation, consistent output, retries and fallbacks, authentication hooks, logging, cost tracking, rate limiting, virtual keys, and per-project budgets. In other words, its center of gravity is a general multi-provider gateway.

That uniform interface is useful, but normalization is never free. Providers have different reasoning controls, cache semantics, tool formats, streaming events, safety fields, and error behavior. A common API surface should be tested for the features your coding Agents actually use.

Sub2API: distribute subscription quota through downstream API keys

Sub2API describes itself as an “AI API gateway platform for subscription quota distribution.” Its README explicitly lists multiple upstream account types—including OAuth and API keys—downstream key generation, token-level billing, sticky sessions, account selection, concurrency limits, rate limits, and request forwarding.

This makes account-pool mechanics unusually visible. A request is not merely sent to “a provider”; the system can choose one account from a pool, keep a conversation sticky to it, move traffic when capacity is exhausted, and charge a downstream user according to its own ledger.

That architecture can be legitimate inside an organization that owns and is authorized to use the accounts. When sold by an unknown operator, it raises additional questions: whether the upstream product permits the access pattern, where the accounts came from, what happens when they are limited or closed, and whether one customer’s traffic can affect another. Those answers depend on the upstream terms and the operator’s authorization; the existence of open-source software proves none of them.

AxonHub: make channels, failover, and traces first-class

AxonHub emphasizes cross-protocol access, request tracing, RBAC, quotas, data isolation, load balancing, failover, and real-time cost tracking. Its example points an OpenAI SDK at AxonHub and routes the request to another model family. This makes the gateway’s channel-and-routing role easy to see.

The important lesson is not which project “wins.” It is that the same reverse-proxy position can optimize for different things:

ProjectStated design centerWhat it helps us see
LiteLLMUnified multi-provider gatewayProtocol normalization, virtual access, budgets, routing
Sub2APISubscription quota distributionAccount pools, sticky scheduling, downstream billing and rate controls
AxonHubProtocol conversion and governed model channelsChannel health, failover, RBAC, tracing and cost visibility

Before adopting any project, verify the current license, release cadence, security process, supported provider features, and operational dependencies in its repository. A feature list is not a production-readiness assessment.

4. Follow one request through the machine

Suppose an Agent sends this conceptual request:

{
  "model": "coding-standard",
  "messages": ["...repository context..."],
  "tools": ["shell", "search", "edit"],
  "stream": true
}

The gateway may perform nine steps:

  1. Authenticate the downstream key. Resolve it to a developer, team, project, and budget.
  2. Apply policy. Check whether this user may send code, use tools, or call the requested route.
  3. Resolve the alias. Map coding-standard to one or more provider deployments.
  4. Select a channel or account. Consider health, region, price, remaining quota, concurrency, and stickiness.
  5. Translate the request. Convert tool schemas, message roles, reasoning controls, and streaming format.
  6. Call the upstream API. Use the gateway’s credential or an organization-owned BYOK credential.
  7. Retry or fall back. Decide whether the failure is transient and which alternative is allowed.
  8. Translate the response. Normalize events and usage fields before returning them to the client.
  9. Record and bill. Write the route, tokens, latency, policy decision, and calculated cost to the ledger.

Every bold verb is a governance decision. It is also a place where an opaque operator can change the product you think you bought.

5. Routing, failover, and model downgrade are not the same thing

These terms are often bundled together:

  • Load balancing sends equivalent traffic across healthy deployments.
  • Failover moves a failed request to an approved alternative.
  • Cost routing selects a destination using price and budget signals.
  • Semantic routing selects a model based on task classification.
  • Downgrade accepts a meaningful capability reduction to keep serving.

A safe policy starts with the least semantic change and stops when the task cannot tolerate more.

A visible fallback ladder separates resilience from silent downgrade

Failure stagePreferred actionUser-visible consequenceDefault for sensitive work
One deployment is unhealthySame pinned model, another region/deploymentUsually minimalAllow and log
Provider path is unavailableContractually equivalent approved pathPossible behavior changeAllow only after evaluation
Budget is exhaustedLower-cost modelQuality and capability may dropAsk or fail closed
No approved model remainsReturn a clear errorTask stopsFail closed

The dangerous rule is “if anything fails, use anything cheaper.” A coding Agent may still return fluent text while losing context length, tool reliability, image input, reasoning effort, or instruction-following quality.

A route policy should express intent, not hide a vendor name:

routes:
  coding_sensitive:
    primary: provider_a/model_x@region_1
    fallback:
      - provider_a/model_x@region_2
    when_exhausted: fail_closed

  coding_low_risk:
    primary: provider_a/model_x@region_1
    fallback:
      - provider_b/evaluated_model_y
    disclose_actual_route: true
    when_exhausted: ask_user

Pin versions where behavior consistency matters and run evaluations before changing a mapping. OpenAI’s API documentation similarly recommends pinned model versions and evals when consistent prompting behavior is important.

6. What a good gateway actually buys you

A trustworthy gateway can give a team capabilities that individual client configuration cannot:

  • one identity layer: developers receive short-lived or virtual keys instead of copying upstream secrets;
  • explicit model allowlists: a repository or team can call only approved routes;
  • budget and rate policy: limits can be assigned by project, not guessed from a shared monthly bill;
  • visible resilience: retries and fallbacks are governed and recorded;
  • data controls: retention, redaction, region, and logging rules can differ by workload;
  • auditability: the requested route, actual deployment, upstream request ID, usage, and policy decision can be correlated;
  • a kill switch: one route or credential can be disabled centrally during an incident;
  • portable clients: Claude Code, Codex, Cursor, and internal Agents do not each need a separate policy implementation.

This is the organizational step beyond a local control plane. A local switch answers, “What should this workstation use right now?” A gateway answers, “Who may send which workload to which provider, under what limits, with what evidence?”

7. How an opaque relay can quietly make the service worse

The mechanisms below are technically possible at the gateway position. They are risk patterns, not allegations about the open-source projects discussed above.

Model-name substitution

The client asks for a premium alias; the relay maps it to a cheaper model and returns the requested name in a normalized response. Because the relay owns both alias resolution and response rewriting, the model string alone is not proof of upstream identity.

Capability clipping

The relay can cap output, reduce reasoning settings, remove unsupported tool fields, shorten context, disable caching, or aggressively summarize history. The model name may be unchanged while the effective product is weaker.

Undisclosed rate shaping and overselling

Queuing and rate limits are normal capacity controls when published. They become deceptive when a provider advertises dedicated or high-speed access but silently shares a small upstream quota across many buyers. Common symptoms include time-of-day latency spikes, unstable tokens per second, repeated 429-style failures, and long pauses before the first token.

Low-quality or unstable account pools

An operator can rotate many subscription accounts, OAuth credentials, or API keys, replacing accounts when they are throttled or closed. Sticky sessions may temporarily preserve continuity, but account churn can still create changing limits, regions, model availability, and data provenance. A dramatic discount should trigger a supply-chain question: what durable upstream entitlement funds the service?

Self-authored usage and invoices

The relay sees upstream usage and can calculate a new downstream bill. Its dashboard may be convenient, but it is not independent verification. Incorrect price tables, missing cache discounts, invented multipliers, or rewritten usage fields can all distort the bill.

Prompt and code collection

Content-aware routing requires the gateway to process plaintext. Unless the architecture provides a stronger, specifically documented protection, assume the operator can log source code, prompts, tool arguments, secrets accidentally included in context, and generated output. OWASP’s LLM supply-chain guidance recommends vetting suppliers, terms, privacy policies, and security posture; this is directly relevant to relay selection.

8. You cannot prove a model by asking it “who are you?”

Models can repeat a system prompt, imitate another model’s style, or simply be wrong about their identity. A single clever benchmark question is weak evidence. Treat model identity as a supply-chain claim and collect several kinds of evidence.

Establish a direct baseline

Keep a small official-provider account for comparison. Run the same versioned canary tasks directly and through the gateway. Use repeated samples because model output and latency are variable.

Probe capabilities, not personality

Your canary set should cover the features your Agent depends on:

  • maximum practical context boundary;
  • tool-schema adherence and parallel tool calls;
  • image or file input if used;
  • streaming event shape and time to first token;
  • pinned-model behavior on a small engineering evaluation set;
  • expected usage fields, cache behavior, errors, and rate-limit headers.

Large differences are a reason to investigate, not cryptographic proof of substitution. The gateway may be translating fields, the provider may have changed behavior, or the route may actually be different.

Correlate request evidence

Generate your own trace ID, retain the gateway trace, and—where the provider supports it—retain the upstream request ID and provider-side usage record. Standardized telemetry helps: OpenTelemetry defines common semantic attributes so traces, metrics, and logs can be correlated across systems; see its semantic conventions.

Reconcile three ledgers

For an organization-owned upstream account, compare:

client task ledger ↔ gateway call ledger ↔ provider usage and invoice

Check request counts, input/output/cache tokens, retries, actual routes, and price-card versions. A gateway total that cannot be traced to provider evidence should not be the only source used for payment or governance.

Verify the operator, not just the endpoint

Ask who operates the domain, where data is processed, how long prompts are retained, whether training is allowed, how upstream access is obtained, whether sub-processors exist, and how incidents are reported. For sensitive code, contractual provenance and audit rights matter more than a polished latency chart.

9. A worked policy for a 30-developer team

Imagine a team with Claude Code, Codex, Cursor, and two internal Agents. Instead of sharing provider keys, it creates three logical routes:

RouteWorkloadRouting ruleData ruleExhaustion behavior
coding-sensitiveproprietary core code, security fixesone pinned model, same-model regional failoverno prompt retention; restricted regionfail closed
coding-standardordinary implementation and testsevaluated primary plus visible equivalent fallbackshort operational retentionnotify on fallback
coding-low-risksummaries, search, boilerplatecost-aware routing across approved modelsno secretsask before downgrade

Each developer receives a virtual key tied to team and project. Every request records:

task_id, caller, requested_route, actual_provider, actual_model_version,
gateway_trace_id, upstream_request_id, input/cache/output_tokens,
retry_path, policy_decision, price_card_version, final_cost

The team reviews weekly exceptions, not individual “productivity”: unapproved routes, repeated fallbacks, P95 latency, budget exhaustion, usage mismatches, and tasks whose cheaper route caused more retries or review time.

This design does not guarantee the cheapest token. It makes cost, quality, and risk decisions explainable.

10. Build, buy, or stay direct?

You may not need a gateway. Stay direct when one person uses one provider, the upstream dashboard is sufficient, and there is no shared policy to enforce.

Consider self-hosting when you need provider portability or internal control, possess the operational skill to secure credentials and databases, and can maintain upgrades, backups, high availability, and incident response.

Consider a managed gateway when the control requirements are real but running the data path is not your team’s core competency. Evaluate the vendor as a critical processor of source code, not as a simple SaaS dashboard.

For an inexpensive public relay, use a separate risk category. Do not send proprietary repositories, production logs, credentials, customer data, or high-impact autonomous actions through a service whose upstream provenance and data terms you cannot verify.

Verification checklist

Before routing AI Coding traffic through an intermediary, answer:

  • Who owns and is authorized to use the upstream accounts?
  • Is the requested alias mapped to an explicit provider and pinned model version?
  • Which fallbacks are allowed, and are users told when one occurs?
  • Can the gateway retain prompts, code, tool arguments, or responses?
  • Can a task trace be correlated with an upstream request ID and provider ledger?
  • Are token usage, cache discounts, retries, and price-card versions reconcilable?
  • Are rate and concurrency limits published instead of discovered during a deadline?
  • What happens when a budget is exhausted: downgrade, ask, or fail closed?
  • Can routes, keys, and logs be exported and the service removed cleanly?
  • Has the team run direct-versus-gateway canaries for the capabilities it relies on?

If several answers are “the operator says so,” you have convenience, not governance.

The real purpose of an AI Gateway

An AI Gateway does not create model intelligence. It decides how intelligence is purchased, reached, constrained, observed, and replaced when something fails.

That is why the same architecture can produce opposite outcomes. A well-run gateway turns scattered model calls into explicit policy and auditable evidence. An opaque relay concentrates identity, routing, billing, and data in a party you cannot verify.

The right question is not “Does this relay support my model?” It is:

Can I explain which model handled this task, under which authority and policy, what it cost, what changed during fallback, and what evidence proves the path?

The next article will extend this question from route identity to capability: how to evaluate whether a model route can actually complete your team’s engineering work.

Authoritative references

Last updated on