AI Coding Costs Out of Control? From API Relays to AI Gateways
The model name stayed the same. Why did the experience change?
A team points Claude Code, Codex, Cursor, and an internal Agent at one “OpenAI-compatible” address. Everyone still requests the same model alias. A week later, three things happen:
- the bill is lower, but nobody can reconcile it with the official provider;
- coding quality varies sharply by time of day;
- long tasks become slow, tool calls fail more often, and nobody knows which upstream model actually answered.
The address between the coding tool and the model is not a neutral pipe. It can authenticate users, rewrite requests, select an upstream account, choose another model, queue traffic, retry failures, record prompts, recalculate usage, and rewrite the response. That power can create a trustworthy control point—or a very convenient black box.
This article explains both sides. We will use LiteLLM, Sub2API, and AxonHub to inspect the machinery, then build a governance policy that makes routing and degradation visible.
1. Yes: a gateway is a control system in the model API path
The simplest mental model is:
AI Coding client → gateway or relay → upstream model APIBut a production gateway has three different jobs. Mixing them together is the source of much confusion.
The data plane handles every live request
It terminates the client connection, checks the downstream key, translates the protocol, selects an upstream destination, streams the response, and handles retries. Because the plaintext request must be processed here, the gateway can normally see your prompts, source snippets, tool definitions, and model output.
The control plane decides what should happen
It stores model aliases, provider credentials, team permissions, budgets, rate limits, routing weights, fallback rules, and data policies. A route called coding-standard might map to two deployments today and a different pair next month without changing every developer laptop.
The evidence plane records what did happen
It should connect an internal trace ID to the caller, requested route, actual upstream deployment, latency, token usage, retry path, policy decision, and cost. OpenAI, for example, returns an x-request-id, publishes rate-limit headers, and lets clients supply an X-Client-Request-Id; retaining both sides of that correlation makes troubleshooting much stronger than a dashboard total alone. See the official API request debugging guidance.
The gateway is therefore not merely “one URL for many models.” It is an enforcement point plus a source of operational evidence. If the same operator controls routing, metering, and the only available logs, however, the evidence is not independent.
2. “Relay” describes a position, not a level of trust
In Chinese developer communities, zhongzhuan zhan—an API relay—can describe very different systems:
| Form | What it usually does | Who owns the upstream credential? | Main value | Main question |
|---|---|---|---|---|
| Protocol adapter | Converts one API shape into another | You | Client compatibility | What fields are lost or rewritten? |
| Self-hosted gateway | Centralizes your own providers and policy | You or your organization | Control, audit, portability | Can your team operate it securely? |
| Managed enterprise gateway | Runs governance as a contracted service | You and/or the vendor | Less operational burden | What are the retention, routing, and audit guarantees? |
| Subscription-quota distributor | Pools product subscriptions or accounts and issues downstream keys | Usually the operator | Quota sharing and lower apparent price | Is this use authorized, stable, and traceable? |
| Opaque reseller | Sells a model name through an unknown supply chain | Unknown | Convenience or discount | Is the model, usage, and data handling truthful? |
Open source does not automatically make a deployment trustworthy. It lets you inspect what the software can do; it does not prove how a stranger configured their server, what code they changed, or which upstream accounts they use.
Likewise, “commercial” does not automatically mean bad. A managed gateway with a clear contract, provider invoices, data-processing terms, exportable logs, and auditable routing can be safer than a poorly maintained self-hosted proxy.
3. Three open-source projects reveal three different centers of gravity
These projects overlap, but their stated purposes expose the major building blocks.
LiteLLM: normalize providers and govern access
LiteLLM presents a unified interface for many providers and can run as a central proxy. Its documentation lists protocol translation, consistent output, retries and fallbacks, authentication hooks, logging, cost tracking, rate limiting, virtual keys, and per-project budgets. In other words, its center of gravity is a general multi-provider gateway.
That uniform interface is useful, but normalization is never free. Providers have different reasoning controls, cache semantics, tool formats, streaming events, safety fields, and error behavior. A common API surface should be tested for the features your coding Agents actually use.
Sub2API: distribute subscription quota through downstream API keys
Sub2API describes itself as an “AI API gateway platform for subscription quota distribution.” Its README explicitly lists multiple upstream account types—including OAuth and API keys—downstream key generation, token-level billing, sticky sessions, account selection, concurrency limits, rate limits, and request forwarding.
This makes account-pool mechanics unusually visible. A request is not merely sent to “a provider”; the system can choose one account from a pool, keep a conversation sticky to it, move traffic when capacity is exhausted, and charge a downstream user according to its own ledger.
That architecture can be legitimate inside an organization that owns and is authorized to use the accounts. When sold by an unknown operator, it raises additional questions: whether the upstream product permits the access pattern, where the accounts came from, what happens when they are limited or closed, and whether one customer’s traffic can affect another. Those answers depend on the upstream terms and the operator’s authorization; the existence of open-source software proves none of them.
AxonHub: make channels, failover, and traces first-class
AxonHub emphasizes cross-protocol access, request tracing, RBAC, quotas, data isolation, load balancing, failover, and real-time cost tracking. Its example points an OpenAI SDK at AxonHub and routes the request to another model family. This makes the gateway’s channel-and-routing role easy to see.
The important lesson is not which project “wins.” It is that the same reverse-proxy position can optimize for different things:
| Project | Stated design center | What it helps us see |
|---|---|---|
| LiteLLM | Unified multi-provider gateway | Protocol normalization, virtual access, budgets, routing |
| Sub2API | Subscription quota distribution | Account pools, sticky scheduling, downstream billing and rate controls |
| AxonHub | Protocol conversion and governed model channels | Channel health, failover, RBAC, tracing and cost visibility |
Before adopting any project, verify the current license, release cadence, security process, supported provider features, and operational dependencies in its repository. A feature list is not a production-readiness assessment.
4. Follow one request through the machine
Suppose an Agent sends this conceptual request:
{
"model": "coding-standard",
"messages": ["...repository context..."],
"tools": ["shell", "search", "edit"],
"stream": true
}The gateway may perform nine steps:
- Authenticate the downstream key. Resolve it to a developer, team, project, and budget.
- Apply policy. Check whether this user may send code, use tools, or call the requested route.
- Resolve the alias. Map
coding-standardto one or more provider deployments. - Select a channel or account. Consider health, region, price, remaining quota, concurrency, and stickiness.
- Translate the request. Convert tool schemas, message roles, reasoning controls, and streaming format.
- Call the upstream API. Use the gateway’s credential or an organization-owned BYOK credential.
- Retry or fall back. Decide whether the failure is transient and which alternative is allowed.
- Translate the response. Normalize events and usage fields before returning them to the client.
- Record and bill. Write the route, tokens, latency, policy decision, and calculated cost to the ledger.
Every bold verb is a governance decision. It is also a place where an opaque operator can change the product you think you bought.
5. Routing, failover, and model downgrade are not the same thing
These terms are often bundled together:
- Load balancing sends equivalent traffic across healthy deployments.
- Failover moves a failed request to an approved alternative.
- Cost routing selects a destination using price and budget signals.
- Semantic routing selects a model based on task classification.
- Downgrade accepts a meaningful capability reduction to keep serving.
A safe policy starts with the least semantic change and stops when the task cannot tolerate more.
| Failure stage | Preferred action | User-visible consequence | Default for sensitive work |
|---|---|---|---|
| One deployment is unhealthy | Same pinned model, another region/deployment | Usually minimal | Allow and log |
| Provider path is unavailable | Contractually equivalent approved path | Possible behavior change | Allow only after evaluation |
| Budget is exhausted | Lower-cost model | Quality and capability may drop | Ask or fail closed |
| No approved model remains | Return a clear error | Task stops | Fail closed |
The dangerous rule is “if anything fails, use anything cheaper.” A coding Agent may still return fluent text while losing context length, tool reliability, image input, reasoning effort, or instruction-following quality.
A route policy should express intent, not hide a vendor name:
routes:
coding_sensitive:
primary: provider_a/model_x@region_1
fallback:
- provider_a/model_x@region_2
when_exhausted: fail_closed
coding_low_risk:
primary: provider_a/model_x@region_1
fallback:
- provider_b/evaluated_model_y
disclose_actual_route: true
when_exhausted: ask_userPin versions where behavior consistency matters and run evaluations before changing a mapping. OpenAI’s API documentation similarly recommends pinned model versions and evals when consistent prompting behavior is important.
6. What a good gateway actually buys you
A trustworthy gateway can give a team capabilities that individual client configuration cannot:
- one identity layer: developers receive short-lived or virtual keys instead of copying upstream secrets;
- explicit model allowlists: a repository or team can call only approved routes;
- budget and rate policy: limits can be assigned by project, not guessed from a shared monthly bill;
- visible resilience: retries and fallbacks are governed and recorded;
- data controls: retention, redaction, region, and logging rules can differ by workload;
- auditability: the requested route, actual deployment, upstream request ID, usage, and policy decision can be correlated;
- a kill switch: one route or credential can be disabled centrally during an incident;
- portable clients: Claude Code, Codex, Cursor, and internal Agents do not each need a separate policy implementation.
This is the organizational step beyond a local control plane. A local switch answers, “What should this workstation use right now?” A gateway answers, “Who may send which workload to which provider, under what limits, with what evidence?”
7. How an opaque relay can quietly make the service worse
The mechanisms below are technically possible at the gateway position. They are risk patterns, not allegations about the open-source projects discussed above.
Model-name substitution
The client asks for a premium alias; the relay maps it to a cheaper model and returns the requested name in a normalized response. Because the relay owns both alias resolution and response rewriting, the model string alone is not proof of upstream identity.
Capability clipping
The relay can cap output, reduce reasoning settings, remove unsupported tool fields, shorten context, disable caching, or aggressively summarize history. The model name may be unchanged while the effective product is weaker.
Undisclosed rate shaping and overselling
Queuing and rate limits are normal capacity controls when published. They become deceptive when a provider advertises dedicated or high-speed access but silently shares a small upstream quota across many buyers. Common symptoms include time-of-day latency spikes, unstable tokens per second, repeated 429-style failures, and long pauses before the first token.
Low-quality or unstable account pools
An operator can rotate many subscription accounts, OAuth credentials, or API keys, replacing accounts when they are throttled or closed. Sticky sessions may temporarily preserve continuity, but account churn can still create changing limits, regions, model availability, and data provenance. A dramatic discount should trigger a supply-chain question: what durable upstream entitlement funds the service?
Self-authored usage and invoices
The relay sees upstream usage and can calculate a new downstream bill. Its dashboard may be convenient, but it is not independent verification. Incorrect price tables, missing cache discounts, invented multipliers, or rewritten usage fields can all distort the bill.
Prompt and code collection
Content-aware routing requires the gateway to process plaintext. Unless the architecture provides a stronger, specifically documented protection, assume the operator can log source code, prompts, tool arguments, secrets accidentally included in context, and generated output. OWASP’s LLM supply-chain guidance recommends vetting suppliers, terms, privacy policies, and security posture; this is directly relevant to relay selection.
8. You cannot prove a model by asking it “who are you?”
Models can repeat a system prompt, imitate another model’s style, or simply be wrong about their identity. A single clever benchmark question is weak evidence. Treat model identity as a supply-chain claim and collect several kinds of evidence.
Establish a direct baseline
Keep a small official-provider account for comparison. Run the same versioned canary tasks directly and through the gateway. Use repeated samples because model output and latency are variable.
Probe capabilities, not personality
Your canary set should cover the features your Agent depends on:
- maximum practical context boundary;
- tool-schema adherence and parallel tool calls;
- image or file input if used;
- streaming event shape and time to first token;
- pinned-model behavior on a small engineering evaluation set;
- expected usage fields, cache behavior, errors, and rate-limit headers.
Large differences are a reason to investigate, not cryptographic proof of substitution. The gateway may be translating fields, the provider may have changed behavior, or the route may actually be different.
Correlate request evidence
Generate your own trace ID, retain the gateway trace, and—where the provider supports it—retain the upstream request ID and provider-side usage record. Standardized telemetry helps: OpenTelemetry defines common semantic attributes so traces, metrics, and logs can be correlated across systems; see its semantic conventions.
Reconcile three ledgers
For an organization-owned upstream account, compare:
client task ledger ↔ gateway call ledger ↔ provider usage and invoiceCheck request counts, input/output/cache tokens, retries, actual routes, and price-card versions. A gateway total that cannot be traced to provider evidence should not be the only source used for payment or governance.
Verify the operator, not just the endpoint
Ask who operates the domain, where data is processed, how long prompts are retained, whether training is allowed, how upstream access is obtained, whether sub-processors exist, and how incidents are reported. For sensitive code, contractual provenance and audit rights matter more than a polished latency chart.
9. A worked policy for a 30-developer team
Imagine a team with Claude Code, Codex, Cursor, and two internal Agents. Instead of sharing provider keys, it creates three logical routes:
| Route | Workload | Routing rule | Data rule | Exhaustion behavior |
|---|---|---|---|---|
coding-sensitive | proprietary core code, security fixes | one pinned model, same-model regional failover | no prompt retention; restricted region | fail closed |
coding-standard | ordinary implementation and tests | evaluated primary plus visible equivalent fallback | short operational retention | notify on fallback |
coding-low-risk | summaries, search, boilerplate | cost-aware routing across approved models | no secrets | ask before downgrade |
Each developer receives a virtual key tied to team and project. Every request records:
task_id, caller, requested_route, actual_provider, actual_model_version,
gateway_trace_id, upstream_request_id, input/cache/output_tokens,
retry_path, policy_decision, price_card_version, final_costThe team reviews weekly exceptions, not individual “productivity”: unapproved routes, repeated fallbacks, P95 latency, budget exhaustion, usage mismatches, and tasks whose cheaper route caused more retries or review time.
This design does not guarantee the cheapest token. It makes cost, quality, and risk decisions explainable.
10. Build, buy, or stay direct?
You may not need a gateway. Stay direct when one person uses one provider, the upstream dashboard is sufficient, and there is no shared policy to enforce.
Consider self-hosting when you need provider portability or internal control, possess the operational skill to secure credentials and databases, and can maintain upgrades, backups, high availability, and incident response.
Consider a managed gateway when the control requirements are real but running the data path is not your team’s core competency. Evaluate the vendor as a critical processor of source code, not as a simple SaaS dashboard.
For an inexpensive public relay, use a separate risk category. Do not send proprietary repositories, production logs, credentials, customer data, or high-impact autonomous actions through a service whose upstream provenance and data terms you cannot verify.
Verification checklist
Before routing AI Coding traffic through an intermediary, answer:
- Who owns and is authorized to use the upstream accounts?
- Is the requested alias mapped to an explicit provider and pinned model version?
- Which fallbacks are allowed, and are users told when one occurs?
- Can the gateway retain prompts, code, tool arguments, or responses?
- Can a task trace be correlated with an upstream request ID and provider ledger?
- Are token usage, cache discounts, retries, and price-card versions reconcilable?
- Are rate and concurrency limits published instead of discovered during a deadline?
- What happens when a budget is exhausted: downgrade, ask, or fail closed?
- Can routes, keys, and logs be exported and the service removed cleanly?
- Has the team run direct-versus-gateway canaries for the capabilities it relies on?
If several answers are “the operator says so,” you have convenience, not governance.
The real purpose of an AI Gateway
An AI Gateway does not create model intelligence. It decides how intelligence is purchased, reached, constrained, observed, and replaced when something fails.
That is why the same architecture can produce opposite outcomes. A well-run gateway turns scattered model calls into explicit policy and auditable evidence. An opaque relay concentrates identity, routing, billing, and data in a party you cannot verify.
The right question is not “Does this relay support my model?” It is:
Can I explain which model handled this task, under which authority and policy, what it cost, what changed during fallback, and what evidence proves the path?
The next article will extend this question from route identity to capability: how to evaluate whether a model route can actually complete your team’s engineering work.
Authoritative references
- LiteLLM documentation — unified provider interface, proxy, routing, spend, budgets, and rate controls.
- Sub2API repository — subscription quota distribution, upstream account management, scheduling, billing, and forwarding.
- AxonHub repository — protocol translation, model channels, failover, tracing, RBAC, quotas, and cost tracking.
- OpenAI API overview: debugging requests — request IDs and rate-limit response headers.
- OpenAI models documentation — canonical model catalog and capabilities.
- OWASP LLM03:2025 Supply Chain — supplier, provenance, terms, privacy, and supply-chain risks.
- OpenTelemetry semantic conventions — common attributes for correlating traces, metrics, and logs.