The Knob
Of all the safety mechanisms we have shipped (multi-agent validation, calibrated confidence, security scanning, post-merge monitoring), one stands out for return on investment per hour of engineering work: the diff size limit.
A configuration option that says "if the diff exceeds N lines, route to human review or split into multiple PRs." It is two days of engineering work. The improvement in reviewer signal and reduction in escape rate is disproportionate.
Why Size Matters More Than You Think
Human review quality degrades sharply with diff size. The data is well established and consistent across decades of software engineering research:
- Diffs under 100 lines: reviewers find roughly 70% of real defects on first read.
- Diffs 100-400 lines: 35-50% catch rate.
- Diffs 400-1000 lines: 20-30% catch rate.
- Diffs over 1000 lines: under 15% catch rate.
For AI-generated PRs, the size distribution is heavier-tailed than human-generated. AI happily produces 800-line changes for tickets that should have been 50. Without a size limit, the median escape rate on the long tail is multiples higher than the short-diff median.
The Two-Mode Implementation
We use a single threshold (300 lines is our default) with two behaviors above it:
Mode A: Force split. The AI is asked to decompose the change into a sequence of smaller PRs, each below the threshold. The split is reviewed before any of the smaller PRs are generated. This is the right behavior for most tickets that should not have been one big PR.
Mode B: Human-elevated review. Some tickets legitimately need one large diff (a multi-file rename, a deletion of a deprecated module). For these, the AI produces the large PR but flags it for elevated human review, with the validator concentrating its findings to make the review more efficient.
The routing between modes is based on the AI's plan. If the plan decomposes naturally, force split. If the plan is intrinsically monolithic, elevated review.
What This Actually Buys
Three measurable outcomes from our deployments:
Reviewer fatigue drops. When reviewers consistently see 50-line PRs, they read them carefully. When they see 800-line PRs intermittently, they rubber-stamp those (or block them entirely, which clogs the pipeline). Predictable size discipline produces better attention per PR.
Escape rate drops disproportionately. In one team we tracked, instituting a 250-line limit reduced quarterly escape count by 41%. The intuitive cause: bugs hide in proportion to lines of code. Splitting the same logical change into smaller PRs gives bugs fewer places to hide.
Cycle time improves. Counterintuitively, splitting into three smaller PRs and merging them sequentially is often faster end-to-end than one big PR. The big PR sits in review for days; the small ones cycle through in hours.
What Goes Wrong
Three patterns to watch for:
Cosmetic splits. The AI splits a 600-line change into three 200-line PRs that have to merge in order and cannot be reviewed independently. This is worse than the original, the same review burden in three sittings. The fix: the split must produce independently reviewable PRs. Each PR has to make sense on its own.
Threshold gaming. The AI compresses code to stay under the limit. The compression makes the code harder to read. The fix: the size limit should be on logical-line count (statements, function bodies), not raw line count. Whitespace-only optimization should not change the metric.
Real one-shot changes. Mass renames, file moves, generated code regeneration. These are legitimately large and the force-split is wrong. The exception path (Mode B) handles them, but only if the team explicitly opts in. Default to force-split; opt in to elevated review for known categories.
The Configuration
Three knobs:
- Threshold (lines). 300 is our default. Teams that prefer smaller PRs tune to 150-200. Teams with larger appetite go to 500. Anything above 500 starts losing the safety benefit.
- Split style. "Chronological" (the AI produces a sequence of PRs to merge in order) vs. "independent" (the AI produces PRs that can be reviewed in any order). Independent is better when feasible.
- Elevated review concentration. When in Mode B, what fraction of validator findings are surfaced in the PR description as priority. We default to top 10.
The defaults work for the majority of teams. The tuning conversations happen in the first quarter and rarely change afterward.
What This Replaces
Three older mechanisms that diff size limits effectively subsume:
File-count limits. "Do not change more than 10 files." Reasonable proxy, but a 10-file 800-line PR is still bad and a 12-file 100-line PR is still fine. Lines correlate better with review difficulty than file count.
Per-file size limits. "No single file changes by more than 500 lines." Catches some abuses but misses the case where the AI spreads 800 lines across 4 files. Total diff size catches both.
Reviewer-set "this PR is too big" lookback. Better than nothing, but reactive. By the time the reviewer flags the size, they have already spent the review time. Preventive is cheaper.
The Counterintuitive Effect On AI Behavior
Once the AI knows about the size limit, its planning changes. It produces smaller, more focused plans. The plan decomposition step gets better because the planner is incentivized to produce reviewable units, not maximally efficient diffs.
This is the kind of constraint that improves the system in directions you did not directly target. The safety mechanism becomes a design pressure on the agent's behavior.
What To Build First
If you do not have a size limit in your pipeline today:
- Pick a threshold. Start at 300.
- Implement force-split as the default behavior.
- Add the exception path for known one-shot categories.
- Measure escape rate and reviewer time-per-PR for a month before and after.
The data is typically immediate and the conversation about whether to keep the limit is short.
For how this interacts with reviewer workflow, see the reviewer's toolkit. For the broader safety architecture, see enterprise safety layers.
Frequently asked questions
What is a good maximum PR size for code review?
A default threshold around 300 lines works for most teams; those who prefer tighter review tune to 150-200, and larger-appetite teams go up to 500. Above 500 lines you start losing the safety benefit entirely, because reviewer defect-catch rate falls below 15% on diffs over 1,000 lines. Base the count on logical lines (statements and function bodies), not raw lines, so whitespace tricks can't game it.
Why are smaller pull requests better for AI-generated code?
Human review quality degrades sharply with size, and AI PRs are heavier-tailed than human ones, the agent will happily produce an 800-line change for a ticket that should have been 50. Splitting the same logical change into smaller PRs gives bugs fewer places to hide, which is why size limits cut escape rate disproportionately. Predictable small PRs also keep reviewers reading carefully instead of rubber-stamping. Pair this with a tuned review checklist like the reviewer's toolkit.
How do you handle large AI-generated pull requests?
Use a single threshold with two behaviors above it. Mode A force-splits the change into a sequence of smaller, independently reviewable PRs, reviewing the split plan first, the right default for most oversized tickets. Mode B keeps legitimately large one-shot changes (mass renames, module deletions) as a single diff but flags them for elevated review with the validator concentrating its findings. Route between them based on whether the AI's plan decomposes naturally.
Do PR size limits slow down development?
Counterintuitively, no, cycle time usually improves. A large PR sits in review for days, while three smaller PRs cycle through in hours and merge sequentially, making the same logical change faster end-to-end. The main failure to avoid is cosmetic splits that must merge in a fixed order and can't be reviewed independently, which just spreads the same burden across three sittings.
Should PR size limits count raw lines or logical lines?
Logical lines (statements and function bodies), not raw line count. If you count raw lines, the AI can compress code to slip under the limit, making it harder to read while technically complying. A whitespace-only optimization should never change the metric. This keeps the constraint a genuine reviewability signal rather than a formatting game. For how size limits fit the wider safety stack, see enterprise safety for AI-generated code.
EnsureFix Engineering Team
The EnsureFix engineering team designs and operates the multi-agent pipeline that turns tickets into production-ready pull requests. They write about architecture, model routing, safety validation, and what actually ships in enterprise codebases.