Skip to content
AI Coding Token and Cost Analysis

AI Coding Token and Cost Analysis: What exactly is consumed by a task

Start with an ordinary task

You tell an Agent:

GET /reports/:id sometimes returns 500. Find the cause, fix it, and add a regression test.

The final patch may touch only two or three files. Before it edits them, the Agent may also:

  • read AGENTS.md and the test commands;
  • search for the route, database calls, and retry logic;
  • open related files and an error log;
  • run a test, discover that its fixture is wrong, and try again;
  • run the checks once more and write a summary for review.

That is why task cost does not track diff size very well. The expensive part is often the path to the answer: how much the Agent had to read, how many tools it called, and how many times it had to recover.

This article follows that debugging task. First we make the Agent loop concrete. Then we separate tokens, cache hits, tools, retries, and review time. Finally, we turn those pieces into a small task ledger a team can actually use.

1. An Agent keeps doing the same four things

Codex, Claude Code, GitHub Copilot coding agent, and similar products look different, but their working loop is familiar: read context, choose the next action, use a tool, and inspect the result. See the official descriptions for Claude Code, GitHub Copilot agents, and Codex.

The observe-plan-act-verify loop of an AI Coding Agent

In plain language:

  1. Look: read the request, rules, code excerpts, and the last tool result.
  2. Choose the next move: search, open a file, edit, test, or stop.
  3. Do it: call the shell, search, browser, or another tool.
  4. Check: inspect the output and decide whether another round is needed.

Every round can add input, output, a tool run, or a retry. The “done” message is only the last line of that loop.

One product name can hide several bills

Some Agents are API-metered. Some include an allowance in a subscription. Some also charge for a cloud container, browser, or search service. So do not compare products only by “price per million tokens.” Ask:

  • How many model calls does one task make?
  • Do tools run locally or in the cloud?
  • Are failed steps retried automatically?
  • How much human review does the result need?

A monthly seat total is useful for planning. It is not the cost of fixing one bug.

2. What are you paying for when you pay for tokens?

Think of a model call as handing over a stack of material:

  • Input tokens: the prompt, project rules, code, conversation history, and tool results;
  • Output tokens: the model’s plan, explanation, edits, and final response;
  • Cached input: material that can be reused from an earlier request;
  • Tool results: new material returned by the shell, search, browser, or other tools.

OpenAI and Anthropic expose input, output, and cache-related usage as separate billing dimensions, with provider-specific fields and prices. Use the current OpenAI pricing, OpenAI prompt caching, Anthropic pricing, and Anthropic prompt caching pages as the source of truth.

For a first approximation:

task cost ≈ input + output + tools/runtime + human review

Why input grows so quickly

You typed one sentence. The model may receive much more:

  • system instructions and project rules;
  • previous turns;
  • file excerpts found by the Agent;
  • test output;
  • your latest correction.

After four or five rounds, later requests often carry parts of the earlier history. One extra search can therefore make the next few calls larger too.

Caching is cheaper reuse, not free context

Stable rules and tool definitions may be served from a prompt cache. A cache hit is usually cheaper than ordinary input, but creating a cache, missing it, or never using it again can still have a cost.

Keep stable instructions early and volatile logs later, then measure the real hit rate. Do not put the entire repository in a permanent prefix just to chase cache hits.

3. Spread one debugging task across the table

The numbers below are hypothetical. They demonstrate the method, not a vendor quote.

RoundWhat the Agent didNew materialExample usage
1Read rules, route, and teststask context6,000 input; 1,200 output
2Searched retry logicsearch results3,500 input; 700 output
3Inspected the database adaptercode and logs8,000 input; 900 output
4Edited code and testpatch context2,000 input; 1,100 output
5First test failederror output4,000 input; 600 output
6Fixed the fixture and rerancorrected context2,000 input; 500 output
7Reported the resultfinal summary1,000 input; 700 output

This is a small code change, but it still contains seven model decisions, several tool calls, and one recovery loop.

If 18,000 tokens came from repeated context and were served from cache, the record should preserve that fact:

{
  "task_id": "BUG-1842",
  "input_tokens": 26500,
  "cached_input_tokens": 18000,
  "output_tokens": 5700,
  "tool_calls": 9,
  "retries": 1,
  "human_review_minutes": 14,
  "outcome": "merged"
}

With a hypothetical price card of $2 / 1M uncached input, $0.20 / 1M cached input, and $8 / 1M output:

(26,500 - 18,000) × $2 / 1M
+ 18,000 × $0.20 / 1M
+ 5,700 × $8 / 1M
= $0.0986

The cents are not the point. Output can be small in volume but expensive; cache hits change the price of repeated context; and one failed test adds both model work and review time.

How one task’s context becomes weighted cost

4. Where the waste usually hides

Reading unrelated code

The task touches one service, but the Agent reads the whole monorepo. More background is not always better; it can bury the useful path.

Noisy tool output

Full build logs, dependency trees, and generated-file lists are expensive and hard to scan. Return the failing section and a reproducible command where possible.

Requirements changing halfway through

If “also enforce permissions” appears after the patch is mostly written, the Agent has to revisit earlier decisions. The problem is not stubbornness; the task boundary changed.

No clear acceptance test

“Fix it” invites guesses. “Fix the 500, add a regression test, run command X, and report the passing checks” gives the loop somewhere to stop.

After a task, ask three simple questions:

  1. Did this round of context change the next decision?
  2. Was the failure caused by model capability or unclear input?
  3. If we ran it again, which context or tool call would we remove first?

Those questions usually reveal more than “tokens per developer.”

5. A chat session is not a task ledger

A session is just a conversation window. It may contain three unrelated questions. One bug may span a local Agent, a cloud Agent, CI, and human review.

Give the work a stable ID such as BUG-1842, then attach:

  • model calls;
  • tool calls and runtime;
  • test results;
  • the PR or commit;
  • review time;
  • the final outcome: merged, reverted, abandoned, or blocked.

A task-cost ledger connects model calls to outcomes and team decisions

A minimal call record can be small:

task_id
repository
model
input_tokens
cached_input_tokens
output_tokens
tool_calls
retries
runtime_seconds
review_minutes
outcome
evidence_uri

This is not a scorecard for individuals. It helps answer practical questions: Why do this type of task keep retrying? Which tool produces noisy output? Which model route merges more cleanly?

6. The cheapest model is not always the cheapest result

Compare more than the price per million tokens:

  • first-review pass rate;
  • retries;
  • tests and reverts;
  • human review time;
  • time from start to merge.

A reasonable starting split is:

PhaseStart withUpgrade when
Find files and summarizesmaller modelthe path or summary is unreliable
Design the changemedium or stronger modelpermissions, data, or concurrency are involved
Edit codethe model that passes project checks reliablyrework keeps repeating
Independent reviewanother model or a humanthe change carries real risk

This is a starting point, not a vendor rule. Validate it against your own task history rather than a public benchmark alone.

7. Do not “save money” by skipping the safety work

Some shortcuts only move the bill:

  • skipping tests moves the cost to production;
  • removing permission checks makes dangerous actions faster, not safer;
  • sending secrets or raw production logs may save a search call but creates a security incident;
  • counting only model tokens hides cloud runtime and review time.

Good optimization is usually less dramatic: narrow irrelevant context, reduce tool noise, state acceptance criteria early, use a cheaper model for simple steps, and keep tests and permission boundaries in place.

8. Start the team dashboard with facts

The first dashboard does not need dozens of charts. It should answer:

  1. Which models, tools, and runtimes did this task use?
  2. How many input, cached-input, and output tokens were recorded?
  3. How many retries happened, and did the tests pass?
  4. Was the change merged or reverted?
  5. How much human review did it need?

Then look at the data four ways:

  • by outcome: merged, reverted, abandoned, blocked;
  • by phase: exploration, implementation, verification;
  • by model route: price next to pass rate and review time;
  • by waste source: oversized context, repeated tools, retries, idle runtime.

Use medians and P90s as well as averages. A few difficult tasks often consume most of the budget.

Verification checklist

Take one real task and answer:

  • Which model calls belonged to it?
  • What context did the Agent read, and what was cached?
  • Which tools ran, and how many retries occurred?
  • What evidence came from tests and human review?
  • Was the result merged, reverted, abandoned, or blocked?
  • What would you optimize first, and what quality measure would keep you honest?

If the only answer is “we know the monthly subscription total,” you have a budget number, not task-cost visibility.

Conclusion: count what the Agent did

AI Coding cost is not just the amount of text a model generated. It also includes the context it read, the tools it ran, the failed attempts it recovered from, and the time someone spent checking the result.

Attach those events to one task ID and the monthly blur becomes an engineering record you can explain, compare, and improve.

The next article follows naturally: once calls are visible, how do you control model routes, permissions, and quotas with an AI Gateway?

Authoritative references

Last updated on