How Can AI Coding Maintain Code Quality? From Code Generation to Verification Engineering
An AI Coding Agent finishes the CSV export feature from our previous article. Its delivery note looks reassuring:
Implemented order export.
Added 12 tests.
All tests pass.
No lint errors.The pull request is clean, the test count went up, and every check is green. During acceptance, however, you discover three problems:
- an employee can change
departmentIdin the request and export another department’s orders; - switching filters while an export job is running changes the result;
- the browser test only checks that a download button appears, not what the downloaded CSV contains.
The Agent did not necessarily lie. It produced evidence for the questions it chose to ask. The trouble is that those questions were weaker than the risks in the task.
This is the central quality problem in AI Coding: generation makes a claim; verification decides how much confidence that claim deserves. More generated tests do not automatically create stronger evidence. A test can execute the wrong scenario, assert the wrong thing, or reproduce the same mistaken assumption as the implementation.
This article turns verification into an engineering workflow. By the end, you will be able to:
- use TDD without turning it into “ask AI to write tests first”;
- choose unit, integration, contract, and end-to-end tests by risk;
- make Codex or another coding agent return auditable evidence;
- use AI code review tools as an additional sensor, not a quality oracle;
- build CI gates that prevent an unverified change from becoming a release.
1. Quality is not a property of the generated text
Code quality has several dimensions that can fail independently:
| Question | Typical evidence | What a green result does not prove |
|---|---|---|
| Can the code be parsed and built? | formatter, compiler, type checker | the behavior is correct |
| Does a local rule hold? | unit or property test | components integrate correctly |
| Do components agree? | integration or contract test | a real user can finish the flow |
| Does the user journey work? | end-to-end test and runtime evidence | every edge case and security boundary is covered |
| Is the change safe to merge? | diff review, CI gates, required approval | production will never fail |
No single layer is “the quality check.” Confidence comes from independent evidence that covers the important failure modes at an acceptable cost.
A useful model is:
Delivery confidence = requirement coverage
x evidence strength
x environment fidelity
x reviewer independenceThis is not a formula to calculate. It is a diagnostic tool. If any factor is close to zero, a large number of passing tests can still produce weak confidence:
- Requirement coverage is weak: the test never asks whether another department can be exported.
- Evidence strength is weak: the assertion only checks HTTP 200, not the returned rows.
- Environment fidelity is weak: a mocked job queue behaves differently from production.
- Reviewer independence is weak: the same Agent invents the behavior, implementation, tests, and final judgment from one mistaken assumption.
The practical response is not to demand exhaustive testing. It is to match evidence to risk.
2. Start with a verification contract
Before code or tests, turn the acceptance criteria into a small verification matrix. Continue the order export case:
| Acceptance criterion | Main risk | Cheapest strong evidence |
|---|---|---|
| Employee exports only visible departments | unauthorized disclosure | service-level permission test plus API integration test |
| Export uses the filter snapshot from click time | inconsistent data | unit test for immutable snapshot plus concurrent-operation E2E |
| CSV opens correctly with Chinese text | data corruption | CSV encoder test plus downloaded-file assertion |
| Large export runs asynchronously | timeout and frozen UI | job integration test plus browser smoke path |
| Existing API clients still work | compatibility regression | contract test and diff review |
This matrix is the bridge from the specification and acceptance criteria to executable evidence. It stops the Agent from picking only the easiest checks.
For every row, ask four questions:
- What observable behavior must hold? Avoid “works correctly.”
- At which boundary can it fail? Pure logic, database, API, browser, third party, permissions, or deployment.
- What is the smallest reliable test at that boundary? Start low in the stack, then add a broad test only when it adds different evidence.
- Who or what must judge it independently? Deterministic tool, a separate review pass, a domain owner, or a security reviewer.
Now “done” has a contract. The Agent can propose evidence, but it cannot silently redefine success.
3. TDD is a design and feedback loop, not a file order
Martin Fowler’s current summary of Test-Driven Development describes the familiar cycle:
- write a test for the next behavior;
- write enough functional code to make it pass;
- refactor the code and tests into a clean structure.
The third step matters. Without refactoring, TDD can produce a tested pile of patches. With an AI Agent, there is another important safeguard: observe the red state before accepting the green state.
Red: prove that the test can detect the missing behavior
Suppose an employee from department sales attempts to export orders from finance:
import { describe, expect, it } from "vitest";
import { authorizeExport } from "./export-policy";
describe("authorizeExport", () => {
it("rejects an employee exporting another department", () => {
const actor = { role: "employee", departmentId: "sales" };
expect(() =>
authorizeExport(actor, { departmentId: "finance" }),
).toThrowError("EXPORT_SCOPE_FORBIDDEN");
});
});Before implementing the policy, run this test and record that it fails for the expected reason. A test that starts green may be exercising existing behavior, missing the target code, or containing an ineffective assertion.
Green: make the smallest behaviorally correct change
export function authorizeExport(
actor: { role: string; departmentId: string },
filter: { departmentId: string },
) {
if (
actor.role !== "admin" &&
actor.departmentId !== filter.departmentId
) {
throw new Error("EXPORT_SCOPE_FORBIDDEN");
}
}This makes the example pass, but it is not yet sufficient evidence. The API route might forget to call the function, or it might trust an actor object supplied by the browser. Add an integration test at the real authorization boundary:
it("returns 403 and creates no job for an out-of-scope export", async () => {
const response = await request(app)
.post("/api/exports")
.set("Authorization", employeeFrom("sales"))
.send({ departmentId: "finance" });
expect(response.status).toBe(403);
expect(await exportJobs.count()).toBe(0);
});The unit test diagnoses the policy cheaply. The integration test proves that the HTTP boundary actually enforces it.
Refactor: improve the design while evidence stays green
Once the behavior passes:
- remove duplicated authorization logic;
- give the policy a domain name rather than hiding it in a controller;
- keep test fixtures explicit enough to show role and department;
- run both the new tests and the relevant regression suite.
The AI-specific TDD trap
“Write tests first” is not enough when the Agent derives the requirement, test oracle, and implementation from the same ambiguous prompt. It can encode the same mistake three times.
Use these countermeasures:
- provide confirmed acceptance criteria and concrete counterexamples;
- require the Agent to show that the new test fails before implementation;
- review assertions, fixtures, and skipped tests as carefully as production code;
- add at least one negative or boundary case the implementation did not suggest;
- for critical logic, run a small mutation test or manually alter the condition and confirm that the suite fails.
Mutation testing deliberately changes production code and checks whether tests detect the change. It is useful when coverage is high but assertions may be weak. It is usually too expensive for every commit, so target high-risk policies and run it periodically or in a scheduled CI job.
4. Build a test portfolio, not a giant E2E suite
The test pyramid remains useful as a cost model: keep many focused, fast tests and fewer broad tests that cross the full UI and infrastructure stack. It is not a mandatory ratio. If your broad tests are fast, deterministic, and cheap to maintain, the portfolio can have a different shape.
Use each layer for the question it answers best:
| Layer | Best question | Order export example | Main limitation |
|---|---|---|---|
| Static checks | Is the program structurally valid? | type errors, lint, dependency and secret checks | cannot observe business behavior |
| Unit and property tests | Does one rule hold across examples? | permissions, CSV escaping, filter snapshot | mocks can hide integration faults |
| Integration tests | Do real components agree? | API, database query, job creation, object storage adapter | may omit the browser and deployment |
| Contract tests | Did a provider break a consumer? | export status schema and old client compatibility | only covers declared contracts |
| E2E tests | Can a user complete a critical journey? | request export, wait, download, inspect file | slower, broader, harder to diagnose |
| Manual and exploratory checks | What did our scripted model not anticipate? | accessibility, confusing recovery, unusual data | not repeatable unless recorded |
Two rules prevent waste:
- Test a behavior at the lowest layer that can observe it honestly. CSV quoting belongs in a focused encoder test, not twenty browser cases.
- Add a higher layer when it covers a different boundary. One browser case is valuable because it verifies routing, rendering, authentication, asynchronous state, and download wiring together.
When a broad test finds a defect, add a focused regression test near the cause before fixing it. The broad test protects the journey; the focused test makes future diagnosis faster.
5. Design end-to-end tests around journeys and boundaries
End-to-end testing is often either missing or overused. A useful E2E suite covers a small set of business-critical journeys and risky transitions, not every permutation already covered below the UI.
For the export feature, choose these browser-level scenarios:
- Happy path: an authorized user exports the current filter and the downloaded CSV contains the expected rows and UTF-8 text.
- Authorization boundary: an employee cannot submit or retrieve another department’s export.
- State transition: changing the UI filter after submission does not change the running job’s snapshot.
- Failure recovery: a failed job exposes a useful reason and a safe retry path.
A Playwright-style test might look like this:
test("export keeps the submitted filter snapshot", async ({ page }) => {
await seedOrders({ sales: 3, finance: 2 });
await signInAs(page, employee("sales"));
await page.goto("/orders?status=paid");
await page.getByRole("button", { name: "Export CSV" }).click();
await page.getByLabel("Status").selectOption("refunded");
const download = await Promise.all([
page.waitForEvent("download"),
page.getByRole("link", { name: "Download export" }).click(),
]).then(([file]) => file);
const csv = await readDownload(download);
expect(csv).toContain("paid-order-001");
expect(csv).not.toContain("refunded-order-001");
});The exact API will vary by project. The important part is the observable contract: submit one state, change the page, then inspect the artifact produced by the original state.
The official Playwright best-practices guide recommends testing user-visible behavior, isolating tests, controlling data, preferring user-facing locators, and using retrying web-first assertions. Those principles lead to a practical E2E design:
- give each test an isolated account or data namespace;
- create known data through fixtures or APIs instead of depending on test order;
- mock only third parties you do not control, not the boundary you intend to verify;
- locate
buttonandlinkroles by accessible names rather than fragile CSS structure; - assert the final business result, not merely the presence of a toast;
- collect a trace on CI retries so a failure includes DOM snapshots, network activity, and timing evidence;
- treat retries as diagnostics, not as a way to make a flaky suite look green.
Four practical E2E environments
| Environment | Use it for | Tradeoff |
|---|---|---|
| Local app plus disposable database | fast development and Agent self-verification | may differ from deployment infrastructure |
| CI with containerized dependencies | deterministic regression on every PR | setup and runtime cost |
| Preview deployment smoke test | routing, assets, headers, and deployed integration | slower and requires environment hygiene |
| Staging journey and exploratory test | realistic integrations and release confidence | shared data and third parties can introduce noise |
Do not point autonomous tests at production data or irreversible actions. Use test tenants, restricted credentials, synthetic records, and explicit cleanup. Payment, email, analytics, and destructive endpoints should have sandbox or stubbed providers unless a controlled release test explicitly requires the real integration.
6. How Codex should verify a change
Codex does not carry one universal test suite that understands every repository. The current official Codex best-practices guide recommends telling Codex what “done” means and recording build, test, lint, and review commands in AGENTS.md. Codex can then create or update tests, run relevant checks, confirm behavior, and review the diff.
A disciplined Codex verification loop looks like this:
Read rules and acceptance criteria
↓
Inspect existing code, tests, and failure reproduction
↓
Choose evidence by changed boundary and risk
↓
Observe a failing test or reproduce the bug
↓
Implement the smallest change
↓
Run focused checks, then broader regression checks
↓
Exercise the real UI/API when behavior crosses that boundary
↓
Review the diff and report commands, results, and residual riskPut stable commands in AGENTS.md:
## Verification
- TypeScript changes: `pnpm typecheck`
- Business logic: run related Vitest files first, then `pnpm test`
- API contract changes: `pnpm test:contract`
- User-visible flows: `pnpm playwright test --project=chromium`
- Before handoff: inspect `git diff` and list any checks not run
## High-risk changes
- Authorization, payments, migrations, and public API changes require a human owner.
- Never use production credentials or production data in tests.
- Do not weaken, skip, or delete a failing test without explaining why.Then give Codex a task-level verification contract:
Implement AC-SEC-02 and AC-DATA-03 from the export specification.
Before editing:
1. Reproduce the authorization failure with an API integration test.
2. Show that the new test fails for the expected reason.
After editing:
1. Run typecheck and the focused unit/integration tests.
2. Run the existing export regression suite.
3. Exercise the authorized and unauthorized browser flows.
4. Review the final diff for scope drift, weakened assertions, skipped tests,
client-side-only authorization, sensitive logs, and compatibility changes.
5. Report exact commands and exit status, observed behavior, changed files,
checks not run, and remaining risks. Do not claim a check passed if it was
not executed in this environment.For local changes, Codex /review can inspect uncommitted work, a commit, or a branch diff and return prioritized findings without modifying the working tree. For web applications, browser tooling can verify the rendered flow and collect runtime evidence. Neither feature should silently replace the repository’s deterministic tests or the responsible human approval.
What a useful delivery report looks like
## Verification evidence
Acceptance:
- AC-SEC-02: PASS. Employee from sales receives 403 for finance export.
- AC-DATA-03: PASS. Download retains the submitted `paid` filter snapshot.
Commands:
- `pnpm typecheck` -> exit 0
- `pnpm vitest run src/export` -> 18 passed, exit 0
- `pnpm playwright test export.spec.ts` -> 4 passed, exit 0
Runtime observations:
- Unauthorized request created no export job.
- Downloaded CSV contained 3 sales rows and valid UTF-8 text.
- Browser console had no errors; export API returned 202 then completed.
Diff review:
- 6 files changed; no dependency, schema, or unrelated formatting changes.
Not run / residual risk:
- Safari project not available locally; CI will run WebKit.
- Production object-storage lifecycle is outside this test environment.“All tests pass” is a conclusion. This report is evidence another person can audit.
7. What AI code review tools can and cannot do
AI review is useful because it can read a diff with fresh attention, search for suspicious patterns, and provide feedback before a human reviewer is available. It is especially effective for recurring checks that are explicit in repository rules.
As of August 6, 2026, common options include:
| Tool | Where it fits | Useful capability | Important boundary |
|---|---|---|---|
| Codex Code Review | local diff, commit, branch, or GitHub PR | prioritized findings; can follow review guidance in AGENTS.md | an additional reviewer; hard enforcement remains in tests and branch rules |
| GitHub Copilot code review | GitHub and supported IDEs | PR feedback, suggested fixes, repository context, custom instructions | GitHub explicitly says it can miss problems and feedback must be validated with human review |
| Cursor Bugbot | GitHub, GitLab, and Bitbucket PRs | automatic or manual diff review, comments, fix links, status checks | a successful run proves the review ran, not that every defect was found |
| CodeRabbit | PR, IDE, and CLI workflows across several Git platforms | context-aware review, pre-commit feedback, team rules and feedback loop | vendor findings still need repository tests and owner judgment |
This is a capability map, not a benchmark. The products use different context, models, policies, pricing, and integration surfaces. Run a representative evaluation in your own repository before selecting one.
Evaluate an AI reviewer with seeded cases
Create a small review set from real bugs and safe counterexamples:
| Case | Expected outcome |
|---|---|
| Remove server-side department check | one high-priority authorization finding |
| Rename a public response field | compatibility finding linked to repository rule |
| Valid refactor with unchanged behavior | no blocking finding |
| Existing unrelated issue outside the diff | no claim that the PR introduced it |
Test marked .skip to make CI green | finding on weakened verification |
Measure more than “number of comments”:
- recall on seeded consequential defects;
- false-positive rate on safe changes;
- actionability: correct file, explanation, severity, and safe fix;
- latency and cost;
- data and permission fit: what code leaves your boundary and which repositories the integration can access;
- team outcome: accepted findings, defects caught before merge, and review noise over time.
Do not make an AI comment a required blocking check on day one. Start in advisory mode, calibrate repository rules, examine false positives and misses, then decide which outputs can participate in merge policy. Deterministic checks and accountable approvals should remain the enforcement foundation.
8. Review the tests, not only the implementation
AI-generated tests deserve first-class review. Look for these failure patterns:
The assertion proves activity, not outcome
expect(exportService.create).toHaveBeenCalled();This does not prove the correct filter, authorization scope, encoding, or stored artifact. Assert the arguments and the observable result.
The mock removes the risk being tested
If the authorization repository is mocked to always return true, an API “permission test” proves only that the controller can return 200.
The test mirrors the implementation
Recomputing the expected CSV with the same helper used by production can make the same bug appear correct. Use independent examples or a fixed expected artifact.
The happy path dominates
Generated tests often cover typical inputs because typical inputs are easy to infer. Explicitly request empty, unauthorized, repeated, concurrent, malformed, timeout, and compatibility cases according to risk.
The suite was weakened to become green
Watch for deleted assertions, widened tolerances, longer arbitrary sleeps, .skip, .only, snapshots updated without inspection, and exceptions swallowed in test helpers.
Coverage becomes the goal
Coverage shows which code executed, not whether the assertions could detect a wrong result. Use it to locate untested areas, not as proof of correctness. For critical pure logic, mutation testing or deliberately breaking the implementation provides stronger feedback about test sensitivity.
9. Turn checks into risk-based quality gates
A gate is useful only when failure prevents or escalates delivery. A document saying “please run tests” is guidance; a protected branch requiring those checks is enforcement.
GitHub’s protected branch documentation supports required status checks, required reviews, code-owner approval, conversation resolution, and successful deployments before merge.
Use a risk ladder rather than one enormous pipeline for every change:
| Change risk | Required evidence | Merge control |
|---|---|---|
| Low: copy, docs, isolated style | formatter, links, targeted build | normal CI |
| Medium: business logic, local API | type/lint, unit and integration regression, diff review | required checks plus reviewer |
| High: authorization, money, personal data, migration, public contract | negative security tests, contract/E2E, owner review, rollback plan | code owner, required checks, no bypass, staged release |
For web application security, the OWASP Application Security Verification Standard provides a versioned basis for security requirements and verification. Map relevant ASVS controls to tests and review checklists instead of asking an AI reviewer to “look for security issues” without a standard.
Keep three controls separate:
- Deterministic gates: compiler, lint, tests, schema compatibility, secret and dependency scanning.
- Judgment gates: human and AI review for design, maintainability, data boundaries, and hidden assumptions.
- Runtime gates: preview smoke tests, canary or staged rollout, monitoring, and rollback readiness.
The cost of a gate should be proportional to the cost of a miss. A typo does not need a full payment E2E suite; an authorization policy does.
10. Failure diagnosis: find the weak evidence layer
When a supposedly verified change fails, do not immediately add another broad test. Identify which layer lied or was missing:
| Symptom | Likely weak layer | Smallest useful response |
|---|---|---|
| Unit tests pass, API leaks data | integration boundary | reproduce with a real API and database permission test |
| API tests pass, button never completes | UI/runtime wiring | inspect browser console and network, then add one journey test |
| E2E flakes only in CI | environment or timing | inspect trace, isolate state, replace sleeps with observable waits |
| New tests pass before implementation | test sensitivity | confirm target path executes and observe a red failure |
| AI review emits many style comments | review rules and scope | move mechanical checks to CI and narrow repository guidance |
| All checks pass, requirement still wrong | verification contract | correct the acceptance criterion and add an independent example |
The right repair usually moves evidence closer to the missing boundary. More E2E tests cannot fix an ambiguous requirement, and more review comments cannot fix a test suite that never runs in CI.
11. A reusable verification workflow
For the next AI Coding task, use this sequence:
Before generation
- identify the acceptance criteria and non-goals;
- classify changed boundaries and failure impact;
- create a verification matrix;
- name commands, environments, credentials, and human owners.
During implementation
- reproduce the defect or observe the new test fail;
- implement in small increments;
- run focused checks after each meaningful step;
- keep production and test changes reviewable together;
- never weaken a check without making the decision visible.
Before handoff
- run the broader affected regression suite;
- exercise critical API or browser journeys;
- inspect the diff for scope drift and hidden risk;
- use an independent AI review pass when useful;
- return exact commands, outcomes, artifacts, skipped checks, and residual risk.
Before merge and release
- let CI rerun the evidence in a clean environment;
- require risk-appropriate status checks and approvals;
- keep secrets and production data outside the Agent test environment;
- prepare deployment observation and rollback for high-risk changes.
12. Verification checklist
- Every acceptance criterion maps to observable evidence.
- The test portfolio matches the changed boundaries and risk.
- New tests were seen failing for the expected reason.
- Assertions verify outcomes, not only calls, status codes, or UI presence.
- Unauthorized, empty, error, concurrent, and compatibility cases were considered.
- Focused checks and the affected regression suite passed.
- Critical user journeys were verified through the real API or UI boundary.
- Test data, accounts, and third-party behavior are controlled.
- The final diff contains no skipped tests, weakened assertions, secret leakage, or unrelated changes.
- AI review findings were validated; a human owner reviewed high-risk decisions.
- CI can enforce the required checks in a clean environment.
- The delivery report states what was not tested and what risk remains.
Conclusion
AI Coding does not change the definition of software quality. It changes the production rate and the probability that implementation, tests, and explanation are generated from the same assumption.
The answer is not to distrust all generated code, nor to demand that every task run every test. The answer is a verification system:
Acceptance criteria
→ risk-based test portfolio
→ observable red/green feedback
→ runtime and diff evidence
→ independent review
→ enforceable CI and release gatesTreat tests as executable questions, review as a search for missing questions, and CI as the mechanism that makes agreed evidence unavoidable. Then the Agent’s speed becomes useful without turning “all green” into a substitute for engineering judgment.
Authoritative references
- Martin Fowler: Test-Driven Development
- Martin Fowler: Test Pyramid
- Playwright: Best Practices
- OpenAI: Codex Best Practices
- OpenAI: Code Review
- OpenAI: Custom Code Review Rules for Codex
- GitHub: About GitHub Copilot Code Review
- GitHub: About Protected Branches
- Cursor: Bugbot
- CodeRabbit Documentation
- Stryker Mutator: What Is Mutation Testing?
- OWASP Application Security Verification Standard