Managing the AI Coding Project Lifecycle: An EMED Framework from Exploration to Delivery
The sprint was “fast” but nobody can explain the bill
A four-person team uses AI coding agents to migrate a legacy REST API to a new schema. The project manager estimates two weeks based on past experience with similar migrations. The developers are fluent with Cursor and Claude Code. The task list is clear.
Two weeks later, 60% of the migration is done. Token costs are three times the estimate. Two developers report spending more time reviewing and correcting AI output than they would have spent writing the code themselves. One developer switched models mid-sprint without telling anyone, and the new model’s output follows different patterns. The project documentation was never updated because “the AI will do it next time.”
This is not a tool failure. The agents worked exactly as designed. The failure is a lifecycle management gap: the team applied traditional project planning to an AI coding project without adapting for the unique properties of AI-generated code — probabilistic output, token-based cost, context-dependent quality, and iteration patterns that differ fundamentally from human development.
This article provides a complete lifecycle framework for AI coding projects. By the end, you will be able to:
- structure any AI coding project from initial idea to verified delivery;
- estimate costs using token consumption models instead of developer hours;
- manage the unique risks of AI-generated code: hallucination, context drift, silent quality decay;
- set up iteration controls that prevent token overruns and scope explosion;
- build a reusable readiness checklist for your next AI coding project.
The framework is adapted from the EMED (Exploration, Mobilization, Execution, Delivery) methodology described in Managing AI Projects (Runtasewee & González Sánchez, O’Reilly, 2026), reinterpreted for the specific challenges of AI-assisted software development.
1. The mental model: EMED for AI Coding
EMED divides a project into four sequential phases with a feedback loop. Each phase answers one fundamental question before the next begins.
| Phase | The question it answers | When it ends |
|---|---|---|
| Exploration | Should we build this, and can AI coding agents help? | Go/no-go decision |
| Mobilization | Is the project set up for AI agents to work effectively? | First task contract is ready |
| Execution | Are the agents producing verified, bounded, on-budget code? | All acceptance criteria met |
| Delivery | Is the result shipped, documented, and learnable? | Retrospective complete |
The phases are not rigid gates. A small bug fix might compress Exploration and Mobilization into a single hour. A platform migration might spend weeks in Exploration alone. The value is in the questions, not the duration.
Why a lifecycle framework matters for AI coding
Traditional software project management assumes that implementation is deterministic: given clear requirements and skilled developers, the output is predictable. AI coding breaks this assumption in seven ways.
These shifts are not theoretical. They change how you estimate, plan, and control every project phase. The EMED framework addresses each shift at the point where it matters most.
2. Exploration: decide before you spend tokens
Exploration is the phase most AI coding teams skip. The temptation is strong: “We have the tools, we know the codebase, let’s just start.” But Exploration is where you prevent the most expensive failures.
2.1 Define the problem in AI-coding terms
A good problem definition answers four questions:
## Problem definition
### What observable outcome do we want?
Not "refactor the auth module" but "new developers can add an OAuth provider
in under 30 minutes without reading the auth internals."
### What is in scope and what is protected?
In scope: auth provider interface, configuration, tests.
Protected: public API contract, session storage format, rate-limiting behavior.
### What does "AI-appropriate" mean for this work?
AI coding agents excel at: pattern-based refactoring, boilerplate generation,
test creation from specifications, documentation from code.
AI coding agents struggle with: novel architecture decisions, performance
optimization requiring profiling, cross-system integration with unclear boundaries.
### What would make us stop before starting?
- The change requires understanding runtime behavior that cannot be captured in code or tests.
- The protected surface area is larger than the change area.
- The team has no experience reviewing AI output in this domain.2.2 Estimate token cost before writing a prompt
Token cost estimation is the AI coding equivalent of effort estimation. A practical model uses three variables:
Estimated cost = (base cost per task) × (expected iterations) × (context overhead multiplier)Base cost per task depends on the model and task complexity. Public benchmark data from Aider’s polyglot evaluation (225 coding exercises across six languages, mid-2026) shows per-task costs ranging from approximately $0.03 (Gemini 2.5 Pro) to $0.13 (GPT-5 with high reasoning). Real-world sessions with full repository context typically cost more — Anthropic’s enterprise documentation reports Claude Code averages approximately $13 per developer per active day, which translates to $1.60–$2.60 per task for 5–8 daily tasks (Aider polyglot benchmark, Claude Code enterprise pricing).
Expected iterations account for the fact that AI coding agents rarely produce correct output on the first attempt. For well-specified tasks in familiar domains, expect 1.5–2 iterations. For exploratory or poorly specified work, expect 3–5 iterations. Each iteration re-sends the full context, so iteration cost is not linear — it compounds.
Context overhead multiplier is the ratio of total tokens consumed to actual task-relevant tokens. In multi-turn agentic sessions, each turn includes the full conversation history. A 10-turn debugging session with a 4,000-token system prompt and 8,000 tokens of repository context means the 10th API call sends approximately 120,000+ input tokens even when the new content is 200 tokens. Research on agent token consumption patterns shows the context overhead multiplier can reach 4.7× between different tools performing the same task (How Do Coding Agents Spend Your Money?, 2026).
A practical estimation template:
## Token cost estimate
Task type: [well-specified / moderately-specified / exploratory]
Selected model: [model name and reasoning level]
Base cost per task: $X.XX (from benchmark or historical data)
Expected iterations: N
Context overhead multiplier: M× (typical range: 2×–5×)
Estimated per-task cost: $X.XX × N × M = $Y.YY
Sprint estimate: Y tasks × $Y.YY = $Z.ZZ
Budget ceiling: $Z.ZZ × 1.5 (contingency)2.3 Assess AI feasibility
Not every task benefits from AI coding agents. Use this decision matrix:
| Factor | AI-appropriate | AI-inappropriate |
|---|---|---|
| Specification clarity | Well-defined input/output, clear acceptance criteria | Vague requirements, undefined edge cases |
| Codebase familiarity | Team knows the domain and can review output | Unfamiliar system, no domain expert available |
| Change scope | Bounded: one module, clear boundaries | Cross-cutting: affects many systems |
| Verification | Automated tests exist or can be written | Requires manual testing, user observation |
| Risk tolerance | Failure is recoverable (branch, revert) | Failure affects production data, users, compliance |
When a task falls in the “AI-inappropriate” column, the correct action is not to avoid AI entirely but to decompose the work: use AI for the bounded, well-specified subtasks while humans handle the judgment-intensive parts.
2.4 Exploration gate
Before moving to Mobilization, confirm:
- Problem is defined in observable outcomes, not implementation tasks.
- Protected surfaces are identified (APIs, data formats, permissions).
- Token cost estimate exists with a contingency buffer.
- Team has at least one person who can review AI output in the target domain.
- Go/no-go decision is recorded with rationale.
3. Mobilization: set up the project for AI agents
Mobilization prepares the environment so that AI coding agents can work effectively and safely. This is where most of the investment in AI coding quality happens — and where skipping steps creates the most expensive downstream failures.
3.1 Build the context layer
AI coding agents produce better output when they understand the project. The context layer has three components:
Project rules (AGENTS.md or equivalent) define the working contract that every agent session must follow: required workflow, commands, boundaries, and definition of done. A well-written AGENTS.md is concise — an index pointing to owned documents rather than an encyclopedia copying every rule.
Architecture documentation describes system boundaries, module ownership, and design decisions. This is the information that a human developer would learn over weeks of working on the project. For AI agents, it must be explicitly documented and referenced.
Task-specific context (specifications, acceptance criteria, affected files) is loaded per-task rather than globally. This prevents context window bloat while ensuring each task has the information it needs.
The relationship between these layers is explored in depth in How Does AI Coding Scale Across a Team?. The key principle for lifecycle management is: invest in context before the first sprint, not during it.
3.2 Define autonomy lanes and task contracts
Before any AI coding begins, establish who can do what with which level of supervision. The autonomy lane determines the maximum action an agent (or a developer using an agent) can take without additional approval:
| Lane | Allowed actions | Required review |
|---|---|---|
| Read-only | Inspect code, documentation, test output | None |
| Plan-only | Generate plans, impact analysis, options | Human approves before implementation |
| Branch implementation | Create feature branches with AI-generated code | Diff review and test evidence |
| Reviewed implementation | Merge reviewed AI-generated code | CI passes and owner approves |
| Expert-approved | Cross-module or high-risk changes | Domain expert plus maintainer review |
Each task contract should specify:
## Task contract
### Outcome
What observable change should exist when this is complete?
### In scope / Out of scope
Named files, modules, and behaviors.
### Acceptance evidence
- Automated tests that must pass
- Documentation that must remain accurate
- Boundary checks (what must NOT change)
### Autonomy lane
Which lane applies to this task?
### Stop conditions
- Token budget exceeded
- Protected surface affected
- Unexpected cross-module impact
- Three failed iterations on the same subtask3.3 Set token budgets and iteration limits
Token budgets are the AI coding equivalent of time budgets. Set them at three levels:
Per-task budget: the maximum tokens (or dollar cost) for a single task. When exceeded, the developer must stop, analyze why, and either decompose the task or adjust the approach.
Per-sprint budget: the total token allocation for a sprint. Track consumption daily. When 70% is consumed with more than 30% of tasks remaining, escalate.
Per-project budget: the total allocation for the project. Include a contingency buffer (typically 30–50% above the estimate, because Exploration-phase estimates have high variance for unfamiliar work).
Iteration limits complement cost limits: if an agent fails three times on the same subtask, stop and analyze. The failure is likely a context problem (missing information, ambiguous specification) or a capability problem (the model cannot handle this type of task well). Continuing to iterate burns tokens without proportional progress.
3.4 Configure guardrails
Executable controls are more reliable than instruction-based rules. Before the first sprint:
- protected branches reject unreviewed AI-generated code;
- CI runs focused tests, architecture boundary checks, and formatting;
- token consumption is logged and visible (using tools like ccusage);
- forbidden patterns are checked (e.g., no changes to public contracts, no production credentials in test code).
3.5 Mobilization gate
-
AGENTS.mdand project documentation are current and verified. - Tool adapters are configured and tested for each supported AI tool.
- Token budgets and iteration limits are set at task, sprint, and project levels.
- Autonomy lanes are defined and communicated.
- First task contract is written with clear scope, acceptance, and stop conditions.
- CI guardrails are active and tested.
4. Execution: manage the iteration loop
Execution is where AI coding projects spend most of their time and where the unique properties of AI-generated code demand the most adaptation from traditional project management.
4.1 The AI coding iteration loop
A traditional development loop is: write → test → review → merge. An AI coding loop adds two critical steps:
Specify → Generate → Inspect → Verify → Adjust context → Re-generate (if needed) → Review → MergeThe key additions are Inspect and Adjust context.
Inspect means reading the AI-generated code before running tests. AI output can pass tests while violating architectural boundaries, introducing subtle bugs in untested paths, or following patterns that work locally but break at scale. A developer who skips inspection and relies only on tests is accepting risks they cannot see.
Adjust context means improving the prompt, rules, or task specification when the output is unsatisfactory. The most common cause of poor AI output is not a bad model but insufficient or ambiguous context. Each failed iteration should trigger a context diagnosis before a retry:
Iteration 1 fails → Was the specification clear enough? → Improve spec
Iteration 2 fails → Is the model capable of this task type? → Try different model or decompose
Iteration 3 fails → Is this task AI-appropriate? → Consider human implementation4.2 Track cost and quality concurrently
Traditional sprint tracking monitors velocity (story points completed). AI coding sprint tracking must add two dimensions:
Token cost velocity: actual tokens consumed versus budget, tracked daily. A spike in per-task cost usually signals context bloat (the agent is loading too much information) or iteration spirals (the agent is retrying without progress).
Quality acceptance rate: the percentage of AI-generated output that passes review without major revision. A declining acceptance rate across a sprint signals that the context layer is degrading — perhaps the project rules are becoming stale, or the task specifications are becoming less precise.
A simple tracking template:
## Sprint day 3 report
Tasks completed: 8 / 20
Token budget consumed: $42 / $100 (42%)
Cost per completed task: $5.25 (budget was $5.00)
Quality acceptance rate: 75% (6 of 8 required minor revision, 2 required major rework)
Observations:
- Tasks involving the payment module had 3× the average cost.
- Root cause: AGENTS.md does not describe the payment state machine, so the
agent loads all payment files (15,000+ tokens) to infer behavior.
- Action: Add a payment module architecture summary to docs/ before next sprint.4.3 Manage experimentation with spikes
AI coding projects frequently encounter tasks where feasibility is unknown. Rather than estimating these tasks and then failing to meet the estimate, use experimentation spikes — time-boxed investigations with a clear learning goal:
## Spike: Can AI agents generate valid migration scripts?
Timebox: 2 hours
Token budget: $10
Goal: Determine whether Claude Code can produce correct migration scripts
for our schema changes, given the current documentation.
Success criteria:
- Agent produces a migration script that passes the dry-run test.
- Developer review confirms no data loss risk.
If unsuccessful:
- Record what context was missing.
- Recommend human implementation or additional documentation.Spikes convert uncertainty into bounded experiments. They prevent the most expensive failure mode in AI coding: an agent burning through tokens on a task it cannot complete, while the developer waits for a result that never comes.
4.4 Handle scope changes and change requests
AI coding agents are fast at generating code, which creates a subtle risk: the ease of generation makes scope changes feel free. A developer can ask the agent to “also update the admin panel” or “while you’re at it, add error handling for edge case X” without considering the verification cost.
Apply this rule: every scope change requires the same task contract discipline as the original work. A scope addition must have:
- defined outcome and boundaries;
- acceptance evidence;
- token budget allocation (taken from the sprint contingency);
- explicit stop conditions.
Without this discipline, AI coding projects suffer from a new form of scope creep: not “the developer spent three extra days” but “the agent generated 2,000 extra lines that nobody reviewed carefully, and three of them break the build.”
4.5 Communicate progress to stakeholders
AI coding projects require different stakeholder communication than traditional projects. Executive sponsors accustomed to “80% complete” reports need to understand:
- What was generated (files changed, features added, tests created);
- What was verified (tests passed, boundaries checked, review completed);
- What was not verified (known gaps, deferred testing, risk areas);
- Token cost versus budget (consumption rate, projected total, contingency remaining).
A weekly stakeholder update can follow this structure:
## Week 2 update
Completed:
- User registration API migrated (12 endpoints, all tests pass)
- Authentication middleware updated (reviewed, merged)
- API documentation regenerated from updated schema
In progress:
- Payment integration (3 of 8 endpoints complete, complexity higher than estimated)
Not yet verified:
- Rate limiting behavior under new schema (performance test scheduled for day 3)
- Legacy client compatibility (manual testing required)
Budget:
- Token spend: $210 / $350 (60% consumed, 40% of tasks remaining)
- Risk: payment integration may require additional $50–$80
- Mitigation: spike completed; decomposition plan ready if budget exceeded4.6 Execution gate
- All task contracts have acceptance evidence.
- Token consumption is tracked daily and visible.
- Quality acceptance rate is monitored; declining rates trigger context review.
- Scope changes follow task contract discipline.
- Stakeholders receive weekly updates with verified/unverified distinction.
5. Delivery: ship, document, and learn
Delivery in AI coding projects includes all the activities of traditional software delivery plus three additions specific to AI-generated code.
5.1 Integration verification
AI coding agents typically work on isolated tasks within a branch. Integration — combining multiple AI-generated changes into a coherent whole — requires explicit verification:
- do the combined changes maintain internal consistency?
- do any AI-generated changes conflict with each other?
- does the integrated result still satisfy the original architecture boundaries?
Run the full test suite, architecture checks, and boundary validations on the integrated result, not just on individual changes.
5.2 Documentation reconciliation
AI coding sessions generate a large volume of transient context: chat histories, planning notes, iteration logs. Most of this is not durable project knowledge. During Delivery, identify which decisions and discoveries should become permanent:
| Session artifact | Disposition |
|---|---|
| Architecture decision discovered during AI work | Promote to ADR or architecture document |
| New pattern or convention established | Add to AGENTS.md or development guide |
| Task-specific debugging insight | Record in evaluation case for future reference |
| Transient planning discussion | Discard (does not change the project contract) |
| Token consumption data | Aggregate into project cost report |
5.3 Cost reconciliation
Compare actual project cost to the Exploration-phase estimate. Record:
## Final cost report
Estimated token cost: $350
Actual token cost: $485
Variance: +38.6%
Breakdown:
- Planned tasks: $310 (within estimate)
- Spikes and experiments: $45 (not in original estimate)
- Iteration overruns (3 tasks exceeded budget): $80
- Context-related overhead (payment module): $50
Lessons:
- Payment module documentation gap caused 3× context loading overhead.
Fix: architecture summary added to docs/ — expected to reduce future cost by 30%.
- Two exploratory tasks should have been spikes from the start.
Fix: spike criteria added to Mobilization checklist.This data feeds the next project’s Exploration phase, making estimates progressively more accurate.
5.4 Retrospective: what did AI change?
A standard retrospective asks “what went well, what didn’t.” An AI coding retrospective adds:
- Which tasks were AI-appropriate and which were not? Update the feasibility matrix.
- Which models performed best for which task types? Record for future model selection.
- Which context improvements would have reduced iteration count? Prioritize for the next project’s Mobilization.
- Which guardrails caught problems, and which problems escaped? Strengthen the weak controls.
5.5 Delivery gate
- All acceptance criteria are met with evidence.
- Integration tests pass on the combined result.
- Documentation is synchronized with code changes.
- Final cost report is complete with variance analysis.
- Retrospective findings are recorded and actionable.
- Knowledge from this project improves the next project’s Exploration.
6. Anti-patterns: what looks organized but fails
| Anti-pattern | What happens | Better approach |
|---|---|---|
| Skip Exploration, start coding immediately | Token budget exhausted on feasibility questions that should have been answered first | Time-boxed Exploration with go/no-go gate |
| Estimate AI tasks like human tasks | Estimates miss iteration overhead and context cost | Token-based estimation with contingency buffer |
| No iteration limit | Agent burns through budget on one task without progress | Three-iteration rule: stop, diagnose, adjust |
| Context added reactively | Each failure triggers a context patch; quality varies across tasks | Invest in context during Mobilization, before first sprint |
| All tasks at the same autonomy level | High-risk changes get the same freedom as low-risk ones | Autonomy lanes based on task risk and developer capability |
| “Tests passed” is the only quality gate | AI output passes tests but violates architecture, introduces tech debt, or misses edge cases | Inspect + test + boundary check + review |
| Scope changes without budget adjustment | Token budget silently overruns as “quick additions” accumulate | Every scope change gets a task contract and budget allocation |
| No cost tracking during sprint | Budget surprise at the end | Daily token consumption tracking with escalation thresholds |
| Every project starts from zero | Estimates don’t improve; same mistakes repeat | Feedback loop from Delivery to next Exploration |
7. End-to-end case: migrating a user notification service
To make the framework concrete, here is a compressed walkthrough of a real-style project.
Project: Migrate a user notification service from a monolithic email sender to a multi-channel dispatcher (email, SMS, push).
Exploration (2 days)
- Problem defined: new channels can be added without modifying the dispatcher core.
- Protected surfaces: notification delivery API, user preference storage, rate limiting.
- Feasibility: the dispatcher pattern is well-documented; AI agents handle interface-based refactoring well. The channel-specific business logic (SMS provider integration, push notification format) is less AI-appropriate and may need human review.
- Token estimate: 40 tasks × $5 average × 2 iterations × 1.5 overhead = $600. Contingency: $900.
- Go decision: proceed, with spikes planned for SMS and push channel feasibility.
Mobilization (3 days)
AGENTS.mdupdated with notification service architecture summary.- Task contracts written for the first 10 tasks (dispatcher interface, email adapter, tests).
- Token budget: $900 total, $45 per task ceiling, $15 per spike.
- Autonomy lanes: branch implementation for dispatcher and email adapter; plan-only for SMS and push (pending spike results).
- CI configured with notification-specific boundary checks.
Execution (2 weeks)
Week 1: dispatcher interface and email adapter completed. Token spend: $280. Quality acceptance rate: 85%. One task required four iterations because the AGENTS.md did not describe the retry policy — added after the second failure.
Week 1 spike: SMS adapter. Agent produced a working implementation in two iterations. Feasibility confirmed; autonomy upgraded from plan-only to branch implementation.
Week 2: SMS adapter, push adapter, and integration tests completed. Token spend: $410 (total: $690 / $900). Push adapter had lower quality acceptance (60%) due to vendor-specific API quirks that required human review.
Delivery (2 days)
- Integration test on combined result: all 12 notification scenarios pass.
- Documentation updated: architecture document reflects new dispatcher pattern.
- Final cost: $740 (18% below budget). The SMS spike saved an estimated $100 by confirming feasibility before full implementation.
- Retrospective finding: push notification vendor documentation should be added to project context before similar future work.
8. Project lifecycle acceptance checklist
An AI coding project lifecycle is working when you can answer yes to these:
Exploration
- Every project starts with a problem definition in observable outcomes, not implementation tasks.
- AI feasibility is assessed before the first prompt is written.
- Token cost estimate exists with a contingency buffer.
Mobilization
- Context layer (rules, architecture, task specs) is built before the first sprint.
- Token budgets and iteration limits are set at task, sprint, and project levels.
- Autonomy lanes match task risk and developer capability.
- Guardrails are executable, not just documented.
Execution
- AI output is inspected before tests, not only after.
- Context is adjusted after failures, not just retried.
- Token consumption is tracked daily with escalation thresholds.
- Scope changes follow task contract discipline.
Delivery
- Integration verification covers combined AI-generated changes.
- Session knowledge is promoted to durable project assets.
- Final cost report includes variance analysis and lessons.
- Retrospective findings improve the next project’s Exploration.
Conclusion: manage the lifecycle, not just the tools
The teams that succeed with AI coding agents are not the ones with the most powerful models. They are the ones that adapt their project management to the unique properties of AI-generated code.
The EMED framework provides the structure:
Exploration → decide before you spend tokens
Mobilization → invest in context before the first sprint
Execution → manage iterations, cost, and quality concurrently
Delivery → ship, document, learn, and feed the next projectThe shift from traditional project management is not about learning new tools. It is about recognizing that AI coding changes the fundamental economics and risks of software development — and adapting your planning, estimation, and control practices accordingly.
The real sign of maturity is not that every sprint uses AI agents. It is that every project decision — from “should we start?” to “what did we learn?” — accounts for the specific properties of AI-assisted development.
Authoritative references
- Malini Jain Runtasewee and Adrián González Sánchez, Managing AI Projects, O’Reilly Media, May 2026 — EMED methodology, ADRIAN framework, and AI project management lifecycle.
- Paul Gauthier, Aider polyglot benchmark and cost data — public per-task cost and pass-rate data for AI coding models across six languages.
- Anthropic, Claude Code enterprise documentation — per-developer daily cost figures and usage patterns.
- Kunal Ganglani, AI Agent Cost Per Task: Token Budgets & Break-Even Math (2026) — detailed token consumption breakdown for agentic coding sessions.
- How Do Coding Agents Spend Your Money? (OpenReview, 2026) — empirical study of token consumption patterns and context overhead in AI coding agents.
- Gartner, AI Coding Costs Will Surpass Average Developer Salary by 2028 (June 2026) — industry projection on rising token consumption in AI-assisted development.