How to review AI-written code: a 10-point checklist
Short answer: review AI-written code like a pull request from a fast, confident new teammate. Read the task and the tests first, then the diff. Check that it did only what was asked, that nothing was quietly removed, that errors and edge cases are handled, and that no secrets, personal data or new dependencies slipped in. Run it yourself before you merge.
Agents make different mistakes from people. They rarely make typos. They do invent APIs, change code nobody asked them to change, weaken a test until it passes, and sound certain when they are wrong. The checklist below targets those.
The checklist
| # | Check | What to look for |
|---|---|---|
| 1 | Scope | Only the files the task needed changed |
| 2 | Tests first | New tests prove the behaviour; old tests were not weakened |
| 3 | Deletions | Nothing important quietly removed |
| 4 | Invented APIs | Every function, flag and package actually exists |
| 5 | Edge cases | Empty, null, huge, slow, concurrent, time zones |
| 6 | Errors | Failures are handled and logged, not swallowed |
| 7 | Security | No secrets, no injection, no personal data in logs |
| 8 | Dependencies | No new package without a reason |
| 9 | Duplication | Reuses existing helpers instead of writing new ones |
| 10 | Run it | You saw it work, not just the agent's summary |
1. Scope
Compare the diff against the task. Agents like to "improve" nearby code, rename things and reformat files. Each extra change is something else to review and something else that can break. Ask the agent to revert anything out of scope.
2. Read the tests before the code
Tests tell you what the agent thinks "done" means. Check that new behaviour has a test that would fail without the change. Then look hard at any edited test: loosening an assertion, adding a skip or deleting a case is the most common way an agent "fixes" a failure.
3. Look for deletions
Skim the red lines. Agents sometimes remove a validation, a feature flag or an error branch because it got in their way. Removed lines are easy to miss because you naturally read what was added.
4. Check that everything it calls exists
Confident, wrong API calls are the classic AI mistake: a method that belongs to a different library version, a CLI flag that doesn't exist, a config key that is silently ignored. If you don't recognise it, look it up.
5. Edge cases
Ask what happens with empty input, missing fields, a slow or failing network call, two requests at once, very large inputs, and dates near midnight or across time zones. Agents write the happy path well and the edges less well.
6. Error handling
Look for catch blocks that do nothing, errors turned into null, and raw error messages or stack traces sent back to users. Failures should be logged with enough context to debug and shown to users in plain language.
7. Security and privacy
Check for API keys or tokens pasted into code, SQL or shell commands built from user input, personal data (names, emails, phone numbers) written to logs or URLs, and new endpoints that skip permission checks.
8. New dependencies
Every new package is code you now maintain and trust. Agents add them freely. Ask whether the standard library or an existing dependency already does the job.
9. Duplication
Agents often don't know your codebase's helpers and write a second version. Search for an existing function before accepting a new one.
10. Run it yourself
The agent's summary is a claim, not evidence. Run the tests, start the app, click through the change, and look at it on a small screen if it's UI.
When to send it back
Send the task back, with specific notes, when the scope is wrong, a test was weakened, or you can't explain what the change does. It is usually faster than fixing it yourself, and the agent learns from the notes within the session.
Make review easier next time
- Give agents smaller, sharper tasks with a clear "done means".
- Run each task on its own branch so every diff is about one thing; see git worktrees for parallel agents.
- Use a tool built around the diff. In Zokie, every task ends as a diff you review before it becomes a pull request.
FAQ
Should AI-written code be reviewed differently?
Yes, with more attention on scope, deleted lines, weakened tests and invented APIs, which are the mistakes agents make more often than people.
Can another AI review AI-written code?
It helps as a first pass and catches plenty. It does not replace a person reading the diff and running the change, especially for money, auth and personal data.
How long should reviewing an AI change take?
Roughly as long as reviewing the same change from a colleague. If it takes much longer, the task was too big; split it.