Glossary / AI code review
What is AI code review?
AI code review is the use of a language model to evaluate a code change — its correctness, security, and quality — before a human merges it. Done well, it returns a plain verdict a reviewer can act on. Done poorly, it returns a scroll of guesses.
How it works
Read the change. Judge it. Report.
Gather context
The pull request, the surrounding repository, the spec, and — ideally — the ability to run the code. The more evidence, the better the judgment.
Evaluate
A model assesses correctness, security, structure, readability, and test coverage the way a senior engineer would, against a known standard.
Return a verdict
Safe to merge, or not — with the evidence behind each finding, so a human can stand behind the decision.
The limit
The gap is the harness, not the model.
The failure mode of AI code review is rarely raw model intelligence. It is context. A model that reads only the diff, in one pass, with no ability to run the code, misses the defects that matter most — the ones that only surface at runtime or across files. We measured this: running OpenAI Codex diff-only on 60 real production pull requests, it caught about 1 in 4 real bugs and cried wolf on clean code. Give the same model runtime evidence, and its accuracy roughly tripled.
Why it matters
Independence is what makes a verdict trustworthy.
A reviewer built into the same tool that wrote the code cannot be the neutral gate. An independent trust layer — a separate judge, with its own context and its own execution — evaluates every change on its merits and hands one auditable verdict to the person who owns the merge. That is the shape of AI code review you can actually rely on. See how Looply does it →
FAQ
Common questions
Is AI code review a replacement for human review?
No. It changes what the human does. Instead of reading every line, the person accountable reads one verdict and the evidence behind it, then owns the merge. The judgment of whether to ship still belongs to a human.
How accurate is AI code review today?
It depends far more on the harness than the model. Run diff-only, a frontier model like OpenAI Codex caught roughly 1 in 4 real bugs on 60 production pull requests and invented false alarms on clean code. Given runtime evidence, the same model's F1 jumped from 0.23 to 0.71.
What does AI code review miss most often?
Defects that only appear when you run the code, and bugs that require tracing behavior across files the diff doesn't show. A model reading only the changed lines cannot see a garbage-collection leak or a signal that never fires two files away.
Should the tool that wrote the code also review it?
A reviewer welded to the agent that produced a change cannot be the neutral gate. Independent review — a separate judge with its own context and execution — is what makes the verdict trustworthy.