Skip to content
AI Coding Project Lifecycle Management

Managing the AI Coding Project Lifecycle: An EMED Framework from Exploration to Delivery

The sprint was “fast” but nobody can explain the bill

A four-person team uses AI coding agents to migrate a legacy REST API to a new schema. The project manager estimates two weeks based on past experience with similar migrations. The developers are fluent with Cursor and Claude Code. The task list is clear.

Two weeks later, 60% of the migration is done. Token costs are three times the estimate. Two developers report spending more time reviewing and correcting AI output than they would have spent writing the code themselves. One developer switched models mid-sprint without telling anyone, and the new model’s output follows different patterns. The project documentation was never updated because “the AI will do it next time.”

This is not a tool failure. The agents worked exactly as designed. The failure is a lifecycle management gap: the team applied traditional project planning to an AI coding project without adapting for the unique properties of AI-generated code — probabilistic output, token-based cost, context-dependent quality, and iteration patterns that differ fundamentally from human development.

This article provides a complete lifecycle framework for AI coding projects. By the end, you will be able to:

  1. structure any AI coding project from initial idea to verified delivery;
  2. estimate costs using token consumption models instead of developer hours;
  3. manage the unique risks of AI-generated code: hallucination, context drift, silent quality decay;
  4. set up iteration controls that prevent token overruns and scope explosion;
  5. build a reusable readiness checklist for your next AI coding project.

The framework is adapted from the EMED (Exploration, Mobilization, Execution, Delivery) methodology described in Managing AI Projects (Runtasewee & González Sánchez, O’Reilly, 2026), reinterpreted for the specific challenges of AI-assisted software development.

1. The mental model: EMED for AI Coding

EMED divides a project into four sequential phases with a feedback loop. Each phase answers one fundamental question before the next begins.

Four phases of the EMED lifecycle adapted for AI coding projects, from Exploration through Delivery with a feedback loop

PhaseThe question it answersWhen it ends
ExplorationShould we build this, and can AI coding agents help?Go/no-go decision
MobilizationIs the project set up for AI agents to work effectively?First task contract is ready
ExecutionAre the agents producing verified, bounded, on-budget code?All acceptance criteria met
DeliveryIs the result shipped, documented, and learnable?Retrospective complete

The phases are not rigid gates. A small bug fix might compress Exploration and Mobilization into a single hour. A platform migration might spend weeks in Exploration alone. The value is in the questions, not the duration.

Why a lifecycle framework matters for AI coding

Traditional software project management assumes that implementation is deterministic: given clear requirements and skilled developers, the output is predictable. AI coding breaks this assumption in seven ways.

Seven key shifts from traditional project management to AI coding project management

These shifts are not theoretical. They change how you estimate, plan, and control every project phase. The EMED framework addresses each shift at the point where it matters most.

2. Exploration: decide before you spend tokens

Exploration is the phase most AI coding teams skip. The temptation is strong: “We have the tools, we know the codebase, let’s just start.” But Exploration is where you prevent the most expensive failures.

2.1 Define the problem in AI-coding terms

A good problem definition answers four questions:

## Problem definition

### What observable outcome do we want?
Not "refactor the auth module" but "new developers can add an OAuth provider
in under 30 minutes without reading the auth internals."

### What is in scope and what is protected?
In scope: auth provider interface, configuration, tests.
Protected: public API contract, session storage format, rate-limiting behavior.

### What does "AI-appropriate" mean for this work?
AI coding agents excel at: pattern-based refactoring, boilerplate generation,
test creation from specifications, documentation from code.
AI coding agents struggle with: novel architecture decisions, performance
optimization requiring profiling, cross-system integration with unclear boundaries.

### What would make us stop before starting?
- The change requires understanding runtime behavior that cannot be captured in code or tests.
- The protected surface area is larger than the change area.
- The team has no experience reviewing AI output in this domain.

2.2 Estimate token cost before writing a prompt

Token cost estimation is the AI coding equivalent of effort estimation. A practical model uses three variables:

Estimated cost = (base cost per task) × (expected iterations) × (context overhead multiplier)

Base cost per task depends on the model and task complexity. Public benchmark data from Aider’s polyglot evaluation (225 coding exercises across six languages, mid-2026) shows per-task costs ranging from approximately $0.03 (Gemini 2.5 Pro) to $0.13 (GPT-5 with high reasoning). Real-world sessions with full repository context typically cost more — Anthropic’s enterprise documentation reports Claude Code averages approximately $13 per developer per active day, which translates to $1.60–$2.60 per task for 5–8 daily tasks (Aider polyglot benchmark, Claude Code enterprise pricing).

Expected iterations account for the fact that AI coding agents rarely produce correct output on the first attempt. For well-specified tasks in familiar domains, expect 1.5–2 iterations. For exploratory or poorly specified work, expect 3–5 iterations. Each iteration re-sends the full context, so iteration cost is not linear — it compounds.

Context overhead multiplier is the ratio of total tokens consumed to actual task-relevant tokens. In multi-turn agentic sessions, each turn includes the full conversation history. A 10-turn debugging session with a 4,000-token system prompt and 8,000 tokens of repository context means the 10th API call sends approximately 120,000+ input tokens even when the new content is 200 tokens. Research on agent token consumption patterns shows the context overhead multiplier can reach 4.7× between different tools performing the same task (How Do Coding Agents Spend Your Money?, 2026).

A practical estimation template:

## Token cost estimate

Task type: [well-specified / moderately-specified / exploratory]
Selected model: [model name and reasoning level]
Base cost per task: $X.XX (from benchmark or historical data)
Expected iterations: N
Context overhead multiplier: M× (typical range: 2×–5×)

Estimated per-task cost: $X.XX × N × M = $Y.YY
Sprint estimate: Y tasks × $Y.YY = $Z.ZZ
Budget ceiling: $Z.ZZ × 1.5 (contingency)

2.3 Assess AI feasibility

Not every task benefits from AI coding agents. Use this decision matrix:

FactorAI-appropriateAI-inappropriate
Specification clarityWell-defined input/output, clear acceptance criteriaVague requirements, undefined edge cases
Codebase familiarityTeam knows the domain and can review outputUnfamiliar system, no domain expert available
Change scopeBounded: one module, clear boundariesCross-cutting: affects many systems
VerificationAutomated tests exist or can be writtenRequires manual testing, user observation
Risk toleranceFailure is recoverable (branch, revert)Failure affects production data, users, compliance

When a task falls in the “AI-inappropriate” column, the correct action is not to avoid AI entirely but to decompose the work: use AI for the bounded, well-specified subtasks while humans handle the judgment-intensive parts.

2.4 Exploration gate

Before moving to Mobilization, confirm:

  • Problem is defined in observable outcomes, not implementation tasks.
  • Protected surfaces are identified (APIs, data formats, permissions).
  • Token cost estimate exists with a contingency buffer.
  • Team has at least one person who can review AI output in the target domain.
  • Go/no-go decision is recorded with rationale.

3. Mobilization: set up the project for AI agents

Mobilization prepares the environment so that AI coding agents can work effectively and safely. This is where most of the investment in AI coding quality happens — and where skipping steps creates the most expensive downstream failures.

3.1 Build the context layer

AI coding agents produce better output when they understand the project. The context layer has three components:

Project rules (AGENTS.md or equivalent) define the working contract that every agent session must follow: required workflow, commands, boundaries, and definition of done. A well-written AGENTS.md is concise — an index pointing to owned documents rather than an encyclopedia copying every rule.

Architecture documentation describes system boundaries, module ownership, and design decisions. This is the information that a human developer would learn over weeks of working on the project. For AI agents, it must be explicitly documented and referenced.

Task-specific context (specifications, acceptance criteria, affected files) is loaded per-task rather than globally. This prevents context window bloat while ensuring each task has the information it needs.

The relationship between these layers is explored in depth in How Does AI Coding Scale Across a Team?. The key principle for lifecycle management is: invest in context before the first sprint, not during it.

3.2 Define autonomy lanes and task contracts

Before any AI coding begins, establish who can do what with which level of supervision. The autonomy lane determines the maximum action an agent (or a developer using an agent) can take without additional approval:

LaneAllowed actionsRequired review
Read-onlyInspect code, documentation, test outputNone
Plan-onlyGenerate plans, impact analysis, optionsHuman approves before implementation
Branch implementationCreate feature branches with AI-generated codeDiff review and test evidence
Reviewed implementationMerge reviewed AI-generated codeCI passes and owner approves
Expert-approvedCross-module or high-risk changesDomain expert plus maintainer review

Each task contract should specify:

## Task contract

### Outcome
What observable change should exist when this is complete?

### In scope / Out of scope
Named files, modules, and behaviors.

### Acceptance evidence
- Automated tests that must pass
- Documentation that must remain accurate
- Boundary checks (what must NOT change)

### Autonomy lane
Which lane applies to this task?

### Stop conditions
- Token budget exceeded
- Protected surface affected
- Unexpected cross-module impact
- Three failed iterations on the same subtask

3.3 Set token budgets and iteration limits

Token budgets are the AI coding equivalent of time budgets. Set them at three levels:

Per-task budget: the maximum tokens (or dollar cost) for a single task. When exceeded, the developer must stop, analyze why, and either decompose the task or adjust the approach.

Per-sprint budget: the total token allocation for a sprint. Track consumption daily. When 70% is consumed with more than 30% of tasks remaining, escalate.

Per-project budget: the total allocation for the project. Include a contingency buffer (typically 30–50% above the estimate, because Exploration-phase estimates have high variance for unfamiliar work).

Iteration limits complement cost limits: if an agent fails three times on the same subtask, stop and analyze. The failure is likely a context problem (missing information, ambiguous specification) or a capability problem (the model cannot handle this type of task well). Continuing to iterate burns tokens without proportional progress.

3.4 Configure guardrails

Executable controls are more reliable than instruction-based rules. Before the first sprint:

  • protected branches reject unreviewed AI-generated code;
  • CI runs focused tests, architecture boundary checks, and formatting;
  • token consumption is logged and visible (using tools like ccusage);
  • forbidden patterns are checked (e.g., no changes to public contracts, no production credentials in test code).

3.5 Mobilization gate

  • AGENTS.md and project documentation are current and verified.
  • Tool adapters are configured and tested for each supported AI tool.
  • Token budgets and iteration limits are set at task, sprint, and project levels.
  • Autonomy lanes are defined and communicated.
  • First task contract is written with clear scope, acceptance, and stop conditions.
  • CI guardrails are active and tested.

4. Execution: manage the iteration loop

Execution is where AI coding projects spend most of their time and where the unique properties of AI-generated code demand the most adaptation from traditional project management.

4.1 The AI coding iteration loop

A traditional development loop is: write → test → review → merge. An AI coding loop adds two critical steps:

Specify → Generate → Inspect → Verify → Adjust context → Re-generate (if needed) → Review → Merge

The key additions are Inspect and Adjust context.

Inspect means reading the AI-generated code before running tests. AI output can pass tests while violating architectural boundaries, introducing subtle bugs in untested paths, or following patterns that work locally but break at scale. A developer who skips inspection and relies only on tests is accepting risks they cannot see.

Adjust context means improving the prompt, rules, or task specification when the output is unsatisfactory. The most common cause of poor AI output is not a bad model but insufficient or ambiguous context. Each failed iteration should trigger a context diagnosis before a retry:

Iteration 1 fails → Was the specification clear enough? → Improve spec
Iteration 2 fails → Is the model capable of this task type? → Try different model or decompose
Iteration 3 fails → Is this task AI-appropriate? → Consider human implementation

4.2 Track cost and quality concurrently

Traditional sprint tracking monitors velocity (story points completed). AI coding sprint tracking must add two dimensions:

Token cost velocity: actual tokens consumed versus budget, tracked daily. A spike in per-task cost usually signals context bloat (the agent is loading too much information) or iteration spirals (the agent is retrying without progress).

Quality acceptance rate: the percentage of AI-generated output that passes review without major revision. A declining acceptance rate across a sprint signals that the context layer is degrading — perhaps the project rules are becoming stale, or the task specifications are becoming less precise.

A simple tracking template:

## Sprint day 3 report

Tasks completed: 8 / 20
Token budget consumed: $42 / $100 (42%)
Cost per completed task: $5.25 (budget was $5.00)
Quality acceptance rate: 75% (6 of 8 required minor revision, 2 required major rework)

Observations:
- Tasks involving the payment module had 3× the average cost.
- Root cause: AGENTS.md does not describe the payment state machine, so the
  agent loads all payment files (15,000+ tokens) to infer behavior.
- Action: Add a payment module architecture summary to docs/ before next sprint.

4.3 Manage experimentation with spikes

AI coding projects frequently encounter tasks where feasibility is unknown. Rather than estimating these tasks and then failing to meet the estimate, use experimentation spikes — time-boxed investigations with a clear learning goal:

## Spike: Can AI agents generate valid migration scripts?

Timebox: 2 hours
Token budget: $10
Goal: Determine whether Claude Code can produce correct migration scripts
      for our schema changes, given the current documentation.

Success criteria:
- Agent produces a migration script that passes the dry-run test.
- Developer review confirms no data loss risk.

If unsuccessful:
- Record what context was missing.
- Recommend human implementation or additional documentation.

Spikes convert uncertainty into bounded experiments. They prevent the most expensive failure mode in AI coding: an agent burning through tokens on a task it cannot complete, while the developer waits for a result that never comes.

4.4 Handle scope changes and change requests

AI coding agents are fast at generating code, which creates a subtle risk: the ease of generation makes scope changes feel free. A developer can ask the agent to “also update the admin panel” or “while you’re at it, add error handling for edge case X” without considering the verification cost.

Apply this rule: every scope change requires the same task contract discipline as the original work. A scope addition must have:

  • defined outcome and boundaries;
  • acceptance evidence;
  • token budget allocation (taken from the sprint contingency);
  • explicit stop conditions.

Without this discipline, AI coding projects suffer from a new form of scope creep: not “the developer spent three extra days” but “the agent generated 2,000 extra lines that nobody reviewed carefully, and three of them break the build.”

4.5 Communicate progress to stakeholders

AI coding projects require different stakeholder communication than traditional projects. Executive sponsors accustomed to “80% complete” reports need to understand:

  • What was generated (files changed, features added, tests created);
  • What was verified (tests passed, boundaries checked, review completed);
  • What was not verified (known gaps, deferred testing, risk areas);
  • Token cost versus budget (consumption rate, projected total, contingency remaining).

A weekly stakeholder update can follow this structure:

## Week 2 update

Completed:
- User registration API migrated (12 endpoints, all tests pass)
- Authentication middleware updated (reviewed, merged)
- API documentation regenerated from updated schema

In progress:
- Payment integration (3 of 8 endpoints complete, complexity higher than estimated)

Not yet verified:
- Rate limiting behavior under new schema (performance test scheduled for day 3)
- Legacy client compatibility (manual testing required)

Budget:
- Token spend: $210 / $350 (60% consumed, 40% of tasks remaining)
- Risk: payment integration may require additional $50–$80
- Mitigation: spike completed; decomposition plan ready if budget exceeded

4.6 Execution gate

  • All task contracts have acceptance evidence.
  • Token consumption is tracked daily and visible.
  • Quality acceptance rate is monitored; declining rates trigger context review.
  • Scope changes follow task contract discipline.
  • Stakeholders receive weekly updates with verified/unverified distinction.

5. Delivery: ship, document, and learn

Delivery in AI coding projects includes all the activities of traditional software delivery plus three additions specific to AI-generated code.

5.1 Integration verification

AI coding agents typically work on isolated tasks within a branch. Integration — combining multiple AI-generated changes into a coherent whole — requires explicit verification:

  • do the combined changes maintain internal consistency?
  • do any AI-generated changes conflict with each other?
  • does the integrated result still satisfy the original architecture boundaries?

Run the full test suite, architecture checks, and boundary validations on the integrated result, not just on individual changes.

5.2 Documentation reconciliation

AI coding sessions generate a large volume of transient context: chat histories, planning notes, iteration logs. Most of this is not durable project knowledge. During Delivery, identify which decisions and discoveries should become permanent:

Session artifactDisposition
Architecture decision discovered during AI workPromote to ADR or architecture document
New pattern or convention establishedAdd to AGENTS.md or development guide
Task-specific debugging insightRecord in evaluation case for future reference
Transient planning discussionDiscard (does not change the project contract)
Token consumption dataAggregate into project cost report

5.3 Cost reconciliation

Compare actual project cost to the Exploration-phase estimate. Record:

## Final cost report

Estimated token cost: $350
Actual token cost: $485
Variance: +38.6%

Breakdown:
- Planned tasks: $310 (within estimate)
- Spikes and experiments: $45 (not in original estimate)
- Iteration overruns (3 tasks exceeded budget): $80
- Context-related overhead (payment module): $50

Lessons:
- Payment module documentation gap caused 3× context loading overhead.
  Fix: architecture summary added to docs/ — expected to reduce future cost by 30%.
- Two exploratory tasks should have been spikes from the start.
  Fix: spike criteria added to Mobilization checklist.

This data feeds the next project’s Exploration phase, making estimates progressively more accurate.

5.4 Retrospective: what did AI change?

A standard retrospective asks “what went well, what didn’t.” An AI coding retrospective adds:

  • Which tasks were AI-appropriate and which were not? Update the feasibility matrix.
  • Which models performed best for which task types? Record for future model selection.
  • Which context improvements would have reduced iteration count? Prioritize for the next project’s Mobilization.
  • Which guardrails caught problems, and which problems escaped? Strengthen the weak controls.

5.5 Delivery gate

  • All acceptance criteria are met with evidence.
  • Integration tests pass on the combined result.
  • Documentation is synchronized with code changes.
  • Final cost report is complete with variance analysis.
  • Retrospective findings are recorded and actionable.
  • Knowledge from this project improves the next project’s Exploration.

6. Anti-patterns: what looks organized but fails

Anti-patternWhat happensBetter approach
Skip Exploration, start coding immediatelyToken budget exhausted on feasibility questions that should have been answered firstTime-boxed Exploration with go/no-go gate
Estimate AI tasks like human tasksEstimates miss iteration overhead and context costToken-based estimation with contingency buffer
No iteration limitAgent burns through budget on one task without progressThree-iteration rule: stop, diagnose, adjust
Context added reactivelyEach failure triggers a context patch; quality varies across tasksInvest in context during Mobilization, before first sprint
All tasks at the same autonomy levelHigh-risk changes get the same freedom as low-risk onesAutonomy lanes based on task risk and developer capability
“Tests passed” is the only quality gateAI output passes tests but violates architecture, introduces tech debt, or misses edge casesInspect + test + boundary check + review
Scope changes without budget adjustmentToken budget silently overruns as “quick additions” accumulateEvery scope change gets a task contract and budget allocation
No cost tracking during sprintBudget surprise at the endDaily token consumption tracking with escalation thresholds
Every project starts from zeroEstimates don’t improve; same mistakes repeatFeedback loop from Delivery to next Exploration

7. End-to-end case: migrating a user notification service

To make the framework concrete, here is a compressed walkthrough of a real-style project.

Project: Migrate a user notification service from a monolithic email sender to a multi-channel dispatcher (email, SMS, push).

Exploration (2 days)

  • Problem defined: new channels can be added without modifying the dispatcher core.
  • Protected surfaces: notification delivery API, user preference storage, rate limiting.
  • Feasibility: the dispatcher pattern is well-documented; AI agents handle interface-based refactoring well. The channel-specific business logic (SMS provider integration, push notification format) is less AI-appropriate and may need human review.
  • Token estimate: 40 tasks × $5 average × 2 iterations × 1.5 overhead = $600. Contingency: $900.
  • Go decision: proceed, with spikes planned for SMS and push channel feasibility.

Mobilization (3 days)

  • AGENTS.md updated with notification service architecture summary.
  • Task contracts written for the first 10 tasks (dispatcher interface, email adapter, tests).
  • Token budget: $900 total, $45 per task ceiling, $15 per spike.
  • Autonomy lanes: branch implementation for dispatcher and email adapter; plan-only for SMS and push (pending spike results).
  • CI configured with notification-specific boundary checks.

Execution (2 weeks)

Week 1: dispatcher interface and email adapter completed. Token spend: $280. Quality acceptance rate: 85%. One task required four iterations because the AGENTS.md did not describe the retry policy — added after the second failure.

Week 1 spike: SMS adapter. Agent produced a working implementation in two iterations. Feasibility confirmed; autonomy upgraded from plan-only to branch implementation.

Week 2: SMS adapter, push adapter, and integration tests completed. Token spend: $410 (total: $690 / $900). Push adapter had lower quality acceptance (60%) due to vendor-specific API quirks that required human review.

Delivery (2 days)

  • Integration test on combined result: all 12 notification scenarios pass.
  • Documentation updated: architecture document reflects new dispatcher pattern.
  • Final cost: $740 (18% below budget). The SMS spike saved an estimated $100 by confirming feasibility before full implementation.
  • Retrospective finding: push notification vendor documentation should be added to project context before similar future work.

8. Project lifecycle acceptance checklist

An AI coding project lifecycle is working when you can answer yes to these:

Exploration

  • Every project starts with a problem definition in observable outcomes, not implementation tasks.
  • AI feasibility is assessed before the first prompt is written.
  • Token cost estimate exists with a contingency buffer.

Mobilization

  • Context layer (rules, architecture, task specs) is built before the first sprint.
  • Token budgets and iteration limits are set at task, sprint, and project levels.
  • Autonomy lanes match task risk and developer capability.
  • Guardrails are executable, not just documented.

Execution

  • AI output is inspected before tests, not only after.
  • Context is adjusted after failures, not just retried.
  • Token consumption is tracked daily with escalation thresholds.
  • Scope changes follow task contract discipline.

Delivery

  • Integration verification covers combined AI-generated changes.
  • Session knowledge is promoted to durable project assets.
  • Final cost report includes variance analysis and lessons.
  • Retrospective findings improve the next project’s Exploration.

Conclusion: manage the lifecycle, not just the tools

The teams that succeed with AI coding agents are not the ones with the most powerful models. They are the ones that adapt their project management to the unique properties of AI-generated code.

The EMED framework provides the structure:

Exploration  →  decide before you spend tokens
Mobilization →  invest in context before the first sprint
Execution    →  manage iterations, cost, and quality concurrently
Delivery     →  ship, document, learn, and feed the next project

The shift from traditional project management is not about learning new tools. It is about recognizing that AI coding changes the fundamental economics and risks of software development — and adapting your planning, estimation, and control practices accordingly.

The real sign of maturity is not that every sprint uses AI agents. It is that every project decision — from “should we start?” to “what did we learn?” — accounts for the specific properties of AI-assisted development.

Authoritative references

Last updated on