All articles
AI Code Review

How to Review AI-Generated Code: A Product Team's PR Checklist

How to review AI-generated code from the product side: a practical PR checklist for coding agent work covering scope, acceptance criteria, edge cases, security.

How to Review AI-Generated Code: A Product Team's PR Checklist

A coding agent opens a pull request. The tests pass, the description is confident, and the diff looks tidy. It is tempting to approve. But passing tests only prove that the code does what the tests check, and an agent often wrote those tests too. The real question is whether the change does what the ticket asked, only what the ticket asked, and nothing that will hurt you later.

The answer is a two-part review. Product reviews intent: acceptance criteria, scope, edge cases and traceability back to the ticket. Engineering reviews implementation: design, security, maintainability and test quality. This guide gives both sides a shared checklist, shows how to send precise fix-up instructions back to the agent, and walks through a real-shaped example.

Why "Tests Pass" Is Not Enough

Tests are only as good as what they assert. When an AI coding agent writes the code and the tests in the same session, both come from the same understanding of the task. If that understanding is wrong, the tests confirm the mistake.

GitHub's own guidance on its coding agent is direct about this. Its responsible use documentation says "you should always review and test the content generated by the cloud agent to ensure that it meets your requirements and is free of errors or security concerns prior to merging." It also warns that while the agent "can generate syntactically correct code, it may not always be secure."

Green checks are a starting signal, not a verdict.

What Makes AI-Generated Code Different to Review

AI-generated code is not worse by default, but it fails in recognizable ways:

  • Plausible but wrong. It reads well and compiles while misreading a business rule.
  • Overreach. The agent "improves" nearby code, renames things or refactors files nobody asked about.
  • Confident gaps. Unhandled cases are filled with a reasonable-looking default instead of a question.
  • Volume. Agents produce large diffs quickly, and reviewer attention does not scale with diff size.

There is also a human factor. In a study published at ACM CCS 2023, Perry and colleagues found that participants with access to an AI assistant "wrote significantly less secure code than those without access," and were also "more likely to believe they wrote secure code." The same study found that people who were more skeptical of the assistant produced safer code. Skepticism is a review skill.

Who Reviews What: Product vs Engineering

Splitting responsibility avoids two failure modes: engineers approving code that matches no requirement, and PMs being asked to judge code they cannot read.

Review areaPrimary ownerQuestion it answers
Acceptance criteria matchProductDoes it do what the story says?
ScopeProduct and engineeringDid it change only what was asked?
Edge cases and statesProductAre the awkward cases handled as intended?
TraceabilityProductCan we link this change to the ticket and decision?
Design and architecture fitEngineeringDoes it belong in the codebase this way?
SecurityEngineeringDoes it open a hole?
Maintainability and debtEngineeringWill this hurt us in three months?
Test qualityEngineeringWould these tests fail if the code broke?

A PM does not need to read every line. They need to run the feature in a preview environment, compare behavior to the criteria, and read the PR description and file list for scope.

Check 1: Does It Match the Acceptance Criteria?

Go through the story's acceptance criteria one by one. For each criterion, ask: where is this implemented, and how was it verified?

A useful habit is to ask the agent to map each criterion to the code and the test that covers it in the PR description. Missing mappings show you exactly where to look. If the criteria were vague to begin with, the agent had to guess, and the fix belongs in the story as much as the code. Our guide to writing specs an agent won't misread covers how to prevent that upstream.

Check 2: Did It Stay in Scope?

Look at the list of changed files before reading any code. Does each file make sense for this story?

Watch for:

  • Edits to unrelated modules
  • New dependencies nobody asked for
  • Renamed functions or moved files
  • Changes to configuration, CI or permissions
  • Features from the story's out-of-scope list quietly appearing

Out-of-scope changes are not always bad, but they are always unreviewed product decisions. Ask for them to be removed or split into a separate PR.

Check 3: Are the Edge Cases Handled as Intended?

An edge case is where agents most often substitute their own judgment. Check the empty, the large, the duplicate, the unauthorized and the failure states:

  • What happens with no data?
  • What happens at the limits you specified?
  • What happens if the user lacks permission?
  • What happens if a downstream call fails?
  • What happens if the action is repeated?

If the story did not define the expected behavior, decide now and write the decision down. Do not let the agent's default become policy by accident.

Check 4: Security Basics

Engineers own this check, but product should know what is being looked at. The OWASP Top 10:2025 lists Broken Access Control as the top web application risk, followed by Security Misconfiguration, with Injection also on the list. For a typical agent PR, the basics are:

  • Access control. Is every new endpoint or action checked for the right role?
  • Input handling. Is user input validated and safely passed to queries?
  • Secrets. No keys, tokens or credentials in code, logs or test fixtures.
  • Data exposure. Responses and exports return only the fields they should.
  • Dependencies. New packages are necessary, maintained and from expected sources.
  • Error handling. Errors fail safely and do not leak internals.

GitHub's documentation also notes that Actions workflows triggered by the agent's PRs "require approval from a user with write access before they will run," a reminder that agent output should pass through the same gates as any outside contribution.

Check 5: Maintainability and Technical Debt Signals

Agents can add technical debt quickly because the cost of writing code is near zero for them. Google's code review guide warns about over-engineering, where "developers have made the code more generic than it needs to be, or added functionality that isn't presently needed." That describes a common agent habit.

Signals to look for:

  • Duplicated logic instead of reusing an existing helper
  • New abstractions with a single caller
  • Inconsistent naming or patterns compared with the rest of the codebase
  • Large functions doing several jobs
  • Comments that describe what the code does rather than why

For a product view of when debt is acceptable, see our technical debt blueprint for solo builders.

Check 6: Are the Tests Any Good?

Google's guide asks reviewers to make sure "tests are correct, sensible, and useful" and that they will actually fail when the code breaks. With agent-written tests, check:

  • Each acceptance criterion has at least one test that asserts the behavior, not just that a function runs
  • Edge cases from Check 3 have tests
  • Tests do not mock away the very logic being tested
  • No tests were deleted, skipped or weakened to make the suite pass
  • Assertions check specific values, not just "no error thrown"

A quick test of the tests: if you reverted the feature code, would at least one test fail?

Check 7: Traceability Back to the Ticket

Every agent PR should link to the ticket it implements, and the ticket should link to the decision behind it. This is requirements-to-code traceability, and it is what lets you answer "why does the code do this?" six months later.

Check that the PR references the ticket, that the description restates what was built and what was deliberately left out, and that any decision made during review is recorded on the ticket, not just in a PR comment. Without this, small deviations accumulate into the drift described in our guide on preventing context drift in AI-generated codebases.

How to Give an Agent Fix-Up Instructions

Agents respond best to review comments that are specific, testable and bounded. Vague feedback such as "clean this up" invites another round of guessing.

Good fix-up instructions:

  • Name the criterion. "AC 3 is not met: non-admins can still call the endpoint."
  • State expected behavior. "Return 403 and do not include the export button for non-admins."
  • Bound the change. "Change only the export controller and its test. Do not modify the settings page."
  • Ask for proof. "Add a test that calls the endpoint as a member and asserts 403."
  • Batch your comments. GitHub's best practices for its coding agent recommend batching comments with "Start a review" so the agent works on "your entire review" at once.

If the same mistake keeps recurring, the fix belongs in the ticket template or the repository's agent instructions, not in another comment. This is the feedback half of the coding agent handoff.

A Worked Example: Reviewing an Agent PR Against a Story

A small SaaS team assigns a coding agent this story:

Story: As a workspace admin, I want to export the member list as CSV so I can audit access.

Acceptance criteria:

  1. Admins can download a CSV with name, email, role and join date for active members
  2. Members with pending invitations are excluded
  3. Non-admins do not see the button, and the endpoint returns 403 for them
  4. Workspaces with up to 5,000 members export without timing out

Out of scope: scheduled exports, custom columns.

The agent opens a PR. CI is green with 12 new tests.

Product review

The PM opens the preview environment as an admin and downloads the file. Columns and data look right (AC 1). She notices two members with pending invitations in the file (AC 2 fails). Logged in as a member, the button is hidden, which looks fine on the surface.

The file list shows changes to the export controller, a new CSV helper, the settings page, and a new "Export schedule" dropdown component. The dropdown is out of scope.

Engineering review

The engineer finds the endpoint checks that the user is logged in but not that they are an admin. The hidden button masked a failed AC 3. The new CSV helper duplicates an existing utility already used elsewhere in the codebase. The tests call the helper directly and never hit the endpoint as a non-admin. Two tests assert only that the response is not empty. The query loads all members into memory, which is a risk for AC 4.

Fix-up instructions sent as one review

  1. AC 2: filter out members whose invitation status is pending. Add a test with one pending member that asserts they are absent.
  2. AC 3: add the admin role check to the export endpoint. Add a test calling it as a member and asserting 403.
  3. Remove the export schedule dropdown. Scheduled exports are out of scope.
  4. Replace the new CSV helper with the existing utility.
  5. AC 4: stream rows instead of loading all members at once. Add a test with 5,000 generated members.
  6. Strengthen the two weak tests to assert specific column values.

The second revision passes both reviews. The PM records the pending-invitation clarification on the ticket so the next story inherits it.

None of these issues were caught by green CI. All of them were caught by checking the code against the story.

Common Review Mistakes With Agent PRs

  • Approving because CI is green
  • Reviewing the diff without running the feature
  • Accepting "helpful" extras that were never requested
  • Leaving fix-up comments one at a time, which triggers several partial rounds
  • Recording decisions only in PR comments where nobody will find them
  • Letting the person who requested the change be the only reviewer of a security-sensitive PR

AI Pull Request Review Checklist

Product reviewer

  • Each acceptance criterion is met in a running preview, not just in the description
  • No out-of-scope features or behavior changes appear
  • Empty, limit, permission, failure and repeat cases behave as intended
  • The PR links to the ticket, and the ticket links to the decision behind it
  • Any new decision made during review is recorded on the ticket

Engineering reviewer

  • Changed files match the story's scope, with no unexplained config, CI or dependency changes
  • Access control is enforced on the server, not only hidden in the UI
  • Inputs are validated, and no secrets or excess data are exposed
  • Existing utilities and patterns are reused, with no speculative abstractions
  • Tests assert behavior for every criterion and edge case and would fail if the feature broke
  • No tests were removed, skipped or weakened

Both

  • Fix-up instructions are specific, bounded and batched into one review
  • Recurring agent mistakes are fixed in the ticket template or agent instructions

How Prodstack Fits

Prodstack is an AI product management operating system with a seven-stage method (Discovery, Strategy, Prioritization, Roadmap, Requirements, Backlog and Growth) and one shared product memory across stages. It produces PRDs, user stories and acceptance criteria, builds agent-ready backlogs with handoff to Jira and Claude Code, and keeps a Traceability Log of decisions, so a reviewer can check an agent's pull request against the criteria and the reasoning behind them. For structuring that work upstream, see our guide to backlogs for AI coding agents.

If you want stories and acceptance criteria that make agent PRs easy to review, start your 7-day trial.

Written by
The Prodstack Team
Product management research

The product team behind Prodstack writes practical, evidence-based guides on discovery, strategy, prioritization, requirements and growth.

Product discoveryPrioritizationRequirementsGrowth
Plan before you prompt

Validate the problem, define the requirements and hand your coding agent a clear backlog, all in one product thread.

No credit card required · 300,000 tokens · 4 documents
// Newsletter
Field notes, in your inbox.

Evidence‑driven thinking on discovery, prioritization, specs, and shipping — plus new articles the moment they drop. No noise.