Skip to content
How Can AI Coding Maintain Code Quality? From Code Generation to Verification Engineering

How Can AI Coding Maintain Code Quality? From Code Generation to Verification Engineering

An AI Coding Agent finishes the CSV export feature from our previous article. Its delivery note looks reassuring:

Implemented order export.
Added 12 tests.
All tests pass.
No lint errors.

The pull request is clean, the test count went up, and every check is green. During acceptance, however, you discover three problems:

  • an employee can change departmentId in the request and export another department’s orders;
  • switching filters while an export job is running changes the result;
  • the browser test only checks that a download button appears, not what the downloaded CSV contains.

The Agent did not necessarily lie. It produced evidence for the questions it chose to ask. The trouble is that those questions were weaker than the risks in the task.

This is the central quality problem in AI Coding: generation makes a claim; verification decides how much confidence that claim deserves. More generated tests do not automatically create stronger evidence. A test can execute the wrong scenario, assert the wrong thing, or reproduce the same mistaken assumption as the implementation.

Code is a claim; layered evidence determines whether it is ready to ship

This article turns verification into an engineering workflow. By the end, you will be able to:

  • use TDD without turning it into “ask AI to write tests first”;
  • choose unit, integration, contract, and end-to-end tests by risk;
  • make Codex or another coding agent return auditable evidence;
  • use AI code review tools as an additional sensor, not a quality oracle;
  • build CI gates that prevent an unverified change from becoming a release.

1. Quality is not a property of the generated text

Code quality has several dimensions that can fail independently:

QuestionTypical evidenceWhat a green result does not prove
Can the code be parsed and built?formatter, compiler, type checkerthe behavior is correct
Does a local rule hold?unit or property testcomponents integrate correctly
Do components agree?integration or contract testa real user can finish the flow
Does the user journey work?end-to-end test and runtime evidenceevery edge case and security boundary is covered
Is the change safe to merge?diff review, CI gates, required approvalproduction will never fail

No single layer is “the quality check.” Confidence comes from independent evidence that covers the important failure modes at an acceptable cost.

A useful model is:

Delivery confidence = requirement coverage
                    x evidence strength
                    x environment fidelity
                    x reviewer independence

This is not a formula to calculate. It is a diagnostic tool. If any factor is close to zero, a large number of passing tests can still produce weak confidence:

  • Requirement coverage is weak: the test never asks whether another department can be exported.
  • Evidence strength is weak: the assertion only checks HTTP 200, not the returned rows.
  • Environment fidelity is weak: a mocked job queue behaves differently from production.
  • Reviewer independence is weak: the same Agent invents the behavior, implementation, tests, and final judgment from one mistaken assumption.

The practical response is not to demand exhaustive testing. It is to match evidence to risk.

2. Start with a verification contract

Before code or tests, turn the acceptance criteria into a small verification matrix. Continue the order export case:

Acceptance criterionMain riskCheapest strong evidence
Employee exports only visible departmentsunauthorized disclosureservice-level permission test plus API integration test
Export uses the filter snapshot from click timeinconsistent dataunit test for immutable snapshot plus concurrent-operation E2E
CSV opens correctly with Chinese textdata corruptionCSV encoder test plus downloaded-file assertion
Large export runs asynchronouslytimeout and frozen UIjob integration test plus browser smoke path
Existing API clients still workcompatibility regressioncontract test and diff review

This matrix is the bridge from the specification and acceptance criteria to executable evidence. It stops the Agent from picking only the easiest checks.

For every row, ask four questions:

  1. What observable behavior must hold? Avoid “works correctly.”
  2. At which boundary can it fail? Pure logic, database, API, browser, third party, permissions, or deployment.
  3. What is the smallest reliable test at that boundary? Start low in the stack, then add a broad test only when it adds different evidence.
  4. Who or what must judge it independently? Deterministic tool, a separate review pass, a domain owner, or a security reviewer.

Now “done” has a contract. The Agent can propose evidence, but it cannot silently redefine success.

3. TDD is a design and feedback loop, not a file order

Martin Fowler’s current summary of Test-Driven Development describes the familiar cycle:

  1. write a test for the next behavior;
  2. write enough functional code to make it pass;
  3. refactor the code and tests into a clean structure.

The third step matters. Without refactoring, TDD can produce a tested pile of patches. With an AI Agent, there is another important safeguard: observe the red state before accepting the green state.

Red: prove that the test can detect the missing behavior

Suppose an employee from department sales attempts to export orders from finance:

import { describe, expect, it } from "vitest";
import { authorizeExport } from "./export-policy";

describe("authorizeExport", () => {
  it("rejects an employee exporting another department", () => {
    const actor = { role: "employee", departmentId: "sales" };

    expect(() =>
      authorizeExport(actor, { departmentId: "finance" }),
    ).toThrowError("EXPORT_SCOPE_FORBIDDEN");
  });
});

Before implementing the policy, run this test and record that it fails for the expected reason. A test that starts green may be exercising existing behavior, missing the target code, or containing an ineffective assertion.

Green: make the smallest behaviorally correct change

export function authorizeExport(
  actor: { role: string; departmentId: string },
  filter: { departmentId: string },
) {
  if (
    actor.role !== "admin" &&
    actor.departmentId !== filter.departmentId
  ) {
    throw new Error("EXPORT_SCOPE_FORBIDDEN");
  }
}

This makes the example pass, but it is not yet sufficient evidence. The API route might forget to call the function, or it might trust an actor object supplied by the browser. Add an integration test at the real authorization boundary:

it("returns 403 and creates no job for an out-of-scope export", async () => {
  const response = await request(app)
    .post("/api/exports")
    .set("Authorization", employeeFrom("sales"))
    .send({ departmentId: "finance" });

  expect(response.status).toBe(403);
  expect(await exportJobs.count()).toBe(0);
});

The unit test diagnoses the policy cheaply. The integration test proves that the HTTP boundary actually enforces it.

Refactor: improve the design while evidence stays green

Once the behavior passes:

  • remove duplicated authorization logic;
  • give the policy a domain name rather than hiding it in a controller;
  • keep test fixtures explicit enough to show role and department;
  • run both the new tests and the relevant regression suite.

The AI-specific TDD trap

“Write tests first” is not enough when the Agent derives the requirement, test oracle, and implementation from the same ambiguous prompt. It can encode the same mistake three times.

Use these countermeasures:

  • provide confirmed acceptance criteria and concrete counterexamples;
  • require the Agent to show that the new test fails before implementation;
  • review assertions, fixtures, and skipped tests as carefully as production code;
  • add at least one negative or boundary case the implementation did not suggest;
  • for critical logic, run a small mutation test or manually alter the condition and confirm that the suite fails.

Mutation testing deliberately changes production code and checks whether tests detect the change. It is useful when coverage is high but assertions may be weak. It is usually too expensive for every commit, so target high-risk policies and run it periodically or in a scheduled CI job.

4. Build a test portfolio, not a giant E2E suite

The test pyramid remains useful as a cost model: keep many focused, fast tests and fewer broad tests that cross the full UI and infrastructure stack. It is not a mandatory ratio. If your broad tests are fast, deterministic, and cheap to maintain, the portfolio can have a different shape.

Match test scope and independence to the risk being controlled

Use each layer for the question it answers best:

LayerBest questionOrder export exampleMain limitation
Static checksIs the program structurally valid?type errors, lint, dependency and secret checkscannot observe business behavior
Unit and property testsDoes one rule hold across examples?permissions, CSV escaping, filter snapshotmocks can hide integration faults
Integration testsDo real components agree?API, database query, job creation, object storage adaptermay omit the browser and deployment
Contract testsDid a provider break a consumer?export status schema and old client compatibilityonly covers declared contracts
E2E testsCan a user complete a critical journey?request export, wait, download, inspect fileslower, broader, harder to diagnose
Manual and exploratory checksWhat did our scripted model not anticipate?accessibility, confusing recovery, unusual datanot repeatable unless recorded

Two rules prevent waste:

  1. Test a behavior at the lowest layer that can observe it honestly. CSV quoting belongs in a focused encoder test, not twenty browser cases.
  2. Add a higher layer when it covers a different boundary. One browser case is valuable because it verifies routing, rendering, authentication, asynchronous state, and download wiring together.

When a broad test finds a defect, add a focused regression test near the cause before fixing it. The broad test protects the journey; the focused test makes future diagnosis faster.

5. Design end-to-end tests around journeys and boundaries

End-to-end testing is often either missing or overused. A useful E2E suite covers a small set of business-critical journeys and risky transitions, not every permutation already covered below the UI.

For the export feature, choose these browser-level scenarios:

  1. Happy path: an authorized user exports the current filter and the downloaded CSV contains the expected rows and UTF-8 text.
  2. Authorization boundary: an employee cannot submit or retrieve another department’s export.
  3. State transition: changing the UI filter after submission does not change the running job’s snapshot.
  4. Failure recovery: a failed job exposes a useful reason and a safe retry path.

A Playwright-style test might look like this:

test("export keeps the submitted filter snapshot", async ({ page }) => {
  await seedOrders({ sales: 3, finance: 2 });
  await signInAs(page, employee("sales"));

  await page.goto("/orders?status=paid");
  await page.getByRole("button", { name: "Export CSV" }).click();
  await page.getByLabel("Status").selectOption("refunded");

  const download = await Promise.all([
    page.waitForEvent("download"),
    page.getByRole("link", { name: "Download export" }).click(),
  ]).then(([file]) => file);

  const csv = await readDownload(download);
  expect(csv).toContain("paid-order-001");
  expect(csv).not.toContain("refunded-order-001");
});

The exact API will vary by project. The important part is the observable contract: submit one state, change the page, then inspect the artifact produced by the original state.

The official Playwright best-practices guide recommends testing user-visible behavior, isolating tests, controlling data, preferring user-facing locators, and using retrying web-first assertions. Those principles lead to a practical E2E design:

  • give each test an isolated account or data namespace;
  • create known data through fixtures or APIs instead of depending on test order;
  • mock only third parties you do not control, not the boundary you intend to verify;
  • locate button and link roles by accessible names rather than fragile CSS structure;
  • assert the final business result, not merely the presence of a toast;
  • collect a trace on CI retries so a failure includes DOM snapshots, network activity, and timing evidence;
  • treat retries as diagnostics, not as a way to make a flaky suite look green.

Four practical E2E environments

EnvironmentUse it forTradeoff
Local app plus disposable databasefast development and Agent self-verificationmay differ from deployment infrastructure
CI with containerized dependenciesdeterministic regression on every PRsetup and runtime cost
Preview deployment smoke testrouting, assets, headers, and deployed integrationslower and requires environment hygiene
Staging journey and exploratory testrealistic integrations and release confidenceshared data and third parties can introduce noise

Do not point autonomous tests at production data or irreversible actions. Use test tenants, restricted credentials, synthetic records, and explicit cleanup. Payment, email, analytics, and destructive endpoints should have sandbox or stubbed providers unless a controlled release test explicitly requires the real integration.

6. How Codex should verify a change

Codex does not carry one universal test suite that understands every repository. The current official Codex best-practices guide recommends telling Codex what “done” means and recording build, test, lint, and review commands in AGENTS.md. Codex can then create or update tests, run relevant checks, confirm behavior, and review the diff.

A disciplined Codex verification loop looks like this:

Read rules and acceptance criteria
        ↓
Inspect existing code, tests, and failure reproduction
        ↓
Choose evidence by changed boundary and risk
        ↓
Observe a failing test or reproduce the bug
        ↓
Implement the smallest change
        ↓
Run focused checks, then broader regression checks
        ↓
Exercise the real UI/API when behavior crosses that boundary
        ↓
Review the diff and report commands, results, and residual risk

A verification loop turns Agent output into auditable delivery evidence

Put stable commands in AGENTS.md:

## Verification

- TypeScript changes: `pnpm typecheck`
- Business logic: run related Vitest files first, then `pnpm test`
- API contract changes: `pnpm test:contract`
- User-visible flows: `pnpm playwright test --project=chromium`
- Before handoff: inspect `git diff` and list any checks not run

## High-risk changes

- Authorization, payments, migrations, and public API changes require a human owner.
- Never use production credentials or production data in tests.
- Do not weaken, skip, or delete a failing test without explaining why.

Then give Codex a task-level verification contract:

Implement AC-SEC-02 and AC-DATA-03 from the export specification.

Before editing:
1. Reproduce the authorization failure with an API integration test.
2. Show that the new test fails for the expected reason.

After editing:
1. Run typecheck and the focused unit/integration tests.
2. Run the existing export regression suite.
3. Exercise the authorized and unauthorized browser flows.
4. Review the final diff for scope drift, weakened assertions, skipped tests,
   client-side-only authorization, sensitive logs, and compatibility changes.
5. Report exact commands and exit status, observed behavior, changed files,
   checks not run, and remaining risks. Do not claim a check passed if it was
   not executed in this environment.

For local changes, Codex /review can inspect uncommitted work, a commit, or a branch diff and return prioritized findings without modifying the working tree. For web applications, browser tooling can verify the rendered flow and collect runtime evidence. Neither feature should silently replace the repository’s deterministic tests or the responsible human approval.

What a useful delivery report looks like

## Verification evidence

Acceptance:
- AC-SEC-02: PASS. Employee from sales receives 403 for finance export.
- AC-DATA-03: PASS. Download retains the submitted `paid` filter snapshot.

Commands:
- `pnpm typecheck` -> exit 0
- `pnpm vitest run src/export` -> 18 passed, exit 0
- `pnpm playwright test export.spec.ts` -> 4 passed, exit 0

Runtime observations:
- Unauthorized request created no export job.
- Downloaded CSV contained 3 sales rows and valid UTF-8 text.
- Browser console had no errors; export API returned 202 then completed.

Diff review:
- 6 files changed; no dependency, schema, or unrelated formatting changes.

Not run / residual risk:
- Safari project not available locally; CI will run WebKit.
- Production object-storage lifecycle is outside this test environment.

“All tests pass” is a conclusion. This report is evidence another person can audit.

7. What AI code review tools can and cannot do

AI review is useful because it can read a diff with fresh attention, search for suspicious patterns, and provide feedback before a human reviewer is available. It is especially effective for recurring checks that are explicit in repository rules.

As of August 6, 2026, common options include:

ToolWhere it fitsUseful capabilityImportant boundary
Codex Code Reviewlocal diff, commit, branch, or GitHub PRprioritized findings; can follow review guidance in AGENTS.mdan additional reviewer; hard enforcement remains in tests and branch rules
GitHub Copilot code reviewGitHub and supported IDEsPR feedback, suggested fixes, repository context, custom instructionsGitHub explicitly says it can miss problems and feedback must be validated with human review
Cursor BugbotGitHub, GitLab, and Bitbucket PRsautomatic or manual diff review, comments, fix links, status checksa successful run proves the review ran, not that every defect was found
CodeRabbitPR, IDE, and CLI workflows across several Git platformscontext-aware review, pre-commit feedback, team rules and feedback loopvendor findings still need repository tests and owner judgment

This is a capability map, not a benchmark. The products use different context, models, policies, pricing, and integration surfaces. Run a representative evaluation in your own repository before selecting one.

Evaluate an AI reviewer with seeded cases

Create a small review set from real bugs and safe counterexamples:

CaseExpected outcome
Remove server-side department checkone high-priority authorization finding
Rename a public response fieldcompatibility finding linked to repository rule
Valid refactor with unchanged behaviorno blocking finding
Existing unrelated issue outside the diffno claim that the PR introduced it
Test marked .skip to make CI greenfinding on weakened verification

Measure more than “number of comments”:

  • recall on seeded consequential defects;
  • false-positive rate on safe changes;
  • actionability: correct file, explanation, severity, and safe fix;
  • latency and cost;
  • data and permission fit: what code leaves your boundary and which repositories the integration can access;
  • team outcome: accepted findings, defects caught before merge, and review noise over time.

Do not make an AI comment a required blocking check on day one. Start in advisory mode, calibrate repository rules, examine false positives and misses, then decide which outputs can participate in merge policy. Deterministic checks and accountable approvals should remain the enforcement foundation.

8. Review the tests, not only the implementation

AI-generated tests deserve first-class review. Look for these failure patterns:

The assertion proves activity, not outcome

expect(exportService.create).toHaveBeenCalled();

This does not prove the correct filter, authorization scope, encoding, or stored artifact. Assert the arguments and the observable result.

The mock removes the risk being tested

If the authorization repository is mocked to always return true, an API “permission test” proves only that the controller can return 200.

The test mirrors the implementation

Recomputing the expected CSV with the same helper used by production can make the same bug appear correct. Use independent examples or a fixed expected artifact.

The happy path dominates

Generated tests often cover typical inputs because typical inputs are easy to infer. Explicitly request empty, unauthorized, repeated, concurrent, malformed, timeout, and compatibility cases according to risk.

The suite was weakened to become green

Watch for deleted assertions, widened tolerances, longer arbitrary sleeps, .skip, .only, snapshots updated without inspection, and exceptions swallowed in test helpers.

Coverage becomes the goal

Coverage shows which code executed, not whether the assertions could detect a wrong result. Use it to locate untested areas, not as proof of correctness. For critical pure logic, mutation testing or deliberately breaking the implementation provides stronger feedback about test sensitivity.

9. Turn checks into risk-based quality gates

A gate is useful only when failure prevents or escalates delivery. A document saying “please run tests” is guidance; a protected branch requiring those checks is enforcement.

GitHub’s protected branch documentation supports required status checks, required reviews, code-owner approval, conversation resolution, and successful deployments before merge.

Use a risk ladder rather than one enormous pipeline for every change:

Change riskRequired evidenceMerge control
Low: copy, docs, isolated styleformatter, links, targeted buildnormal CI
Medium: business logic, local APItype/lint, unit and integration regression, diff reviewrequired checks plus reviewer
High: authorization, money, personal data, migration, public contractnegative security tests, contract/E2E, owner review, rollback plancode owner, required checks, no bypass, staged release

For web application security, the OWASP Application Security Verification Standard provides a versioned basis for security requirements and verification. Map relevant ASVS controls to tests and review checklists instead of asking an AI reviewer to “look for security issues” without a standard.

Keep three controls separate:

  • Deterministic gates: compiler, lint, tests, schema compatibility, secret and dependency scanning.
  • Judgment gates: human and AI review for design, maintainability, data boundaries, and hidden assumptions.
  • Runtime gates: preview smoke tests, canary or staged rollout, monitoring, and rollback readiness.

The cost of a gate should be proportional to the cost of a miss. A typo does not need a full payment E2E suite; an authorization policy does.

10. Failure diagnosis: find the weak evidence layer

When a supposedly verified change fails, do not immediately add another broad test. Identify which layer lied or was missing:

SymptomLikely weak layerSmallest useful response
Unit tests pass, API leaks dataintegration boundaryreproduce with a real API and database permission test
API tests pass, button never completesUI/runtime wiringinspect browser console and network, then add one journey test
E2E flakes only in CIenvironment or timinginspect trace, isolate state, replace sleeps with observable waits
New tests pass before implementationtest sensitivityconfirm target path executes and observe a red failure
AI review emits many style commentsreview rules and scopemove mechanical checks to CI and narrow repository guidance
All checks pass, requirement still wrongverification contractcorrect the acceptance criterion and add an independent example

The right repair usually moves evidence closer to the missing boundary. More E2E tests cannot fix an ambiguous requirement, and more review comments cannot fix a test suite that never runs in CI.

11. A reusable verification workflow

For the next AI Coding task, use this sequence:

Before generation

  • identify the acceptance criteria and non-goals;
  • classify changed boundaries and failure impact;
  • create a verification matrix;
  • name commands, environments, credentials, and human owners.

During implementation

  • reproduce the defect or observe the new test fail;
  • implement in small increments;
  • run focused checks after each meaningful step;
  • keep production and test changes reviewable together;
  • never weaken a check without making the decision visible.

Before handoff

  • run the broader affected regression suite;
  • exercise critical API or browser journeys;
  • inspect the diff for scope drift and hidden risk;
  • use an independent AI review pass when useful;
  • return exact commands, outcomes, artifacts, skipped checks, and residual risk.

Before merge and release

  • let CI rerun the evidence in a clean environment;
  • require risk-appropriate status checks and approvals;
  • keep secrets and production data outside the Agent test environment;
  • prepare deployment observation and rollback for high-risk changes.

12. Verification checklist

  • Every acceptance criterion maps to observable evidence.
  • The test portfolio matches the changed boundaries and risk.
  • New tests were seen failing for the expected reason.
  • Assertions verify outcomes, not only calls, status codes, or UI presence.
  • Unauthorized, empty, error, concurrent, and compatibility cases were considered.
  • Focused checks and the affected regression suite passed.
  • Critical user journeys were verified through the real API or UI boundary.
  • Test data, accounts, and third-party behavior are controlled.
  • The final diff contains no skipped tests, weakened assertions, secret leakage, or unrelated changes.
  • AI review findings were validated; a human owner reviewed high-risk decisions.
  • CI can enforce the required checks in a clean environment.
  • The delivery report states what was not tested and what risk remains.

Conclusion

AI Coding does not change the definition of software quality. It changes the production rate and the probability that implementation, tests, and explanation are generated from the same assumption.

The answer is not to distrust all generated code, nor to demand that every task run every test. The answer is a verification system:

Acceptance criteria
    → risk-based test portfolio
    → observable red/green feedback
    → runtime and diff evidence
    → independent review
    → enforceable CI and release gates

Treat tests as executable questions, review as a search for missing questions, and CI as the mechanism that makes agreed evidence unavoidable. Then the Agent’s speed becomes useful without turning “all green” into a substitute for engineering judgment.

Authoritative references

Last updated on