I used to think reducing risk mostly meant keeping changes small, focused, and easy to review. I still believe in the reason behind that principle: a reviewer needs to understand what changed, a team needs to be able to reverse it, and a narrow change gives fewer things a chance to go wrong.
Agentic development changed the practical shape of the problem for me. A coding agent can produce a surprisingly large implementation quickly. In many environments, engineers are also expected to use that leverage. But review, product judgment, and verification have not accelerated at the same rate as code generation.
So the question I ask has shifted. Not only, “How do I keep this PR small?” but, “How do I make this change safe, understandable, and controllable?” A large diff is not automatically bad, and fast implementation is not evidence of safety. The work is to keep the change's blast radius legible and bounded.
Small changes were a means, not the goal
Small PRs help because they lower cognitive load. They make it easier to connect a change to an intent, spot an unexpected dependency, and reason about rollback. Those benefits still matter. A thousand generated lines do not become easy to review just because they arrived in one minute.
But treating line count as the whole risk model can also mislead. A 20-line authentication change may cross a high-impact boundary. A larger content migration with deterministic output and a tested rollback may be comparatively contained. Size is a useful signal; scope, coupling, reversibility, and impact tell us more about what evidence is needed.
The useful principle underneath “keep it small” is to keep the reasoning and risk manageable. Sometimes that means splitting work. Sometimes it means a broader implementation with explicit boundaries, staged validation, and a clear map for review.
Throughput moved the bottleneck
When implementation takes less time, it is tempting to spend the saved time generating more implementation. That can create an imbalance: the patch grows faster than the team's ability to understand its assumptions, test its behavior, or observe its release.
I have had to think beyond the role of the person writing code. I need to act as architect when setting boundaries, as product owner when clarifying the behavior that should change, as reviewer when deciding where attention matters, as QA when defining proof, and as release manager when planning exposure and recovery.
That doesn't mean one engineer can replace every discipline. It means the person directing a high-throughput change has to make those concerns visible instead of assuming the generated diff will organize itself.
Control the change at its boundaries
A change map is more useful than a raw file count. Before review, identify the parts of the system it touches: user interface, APIs, data shape, permissions, infrastructure, analytics, and rollout. Mark what is intentionally unchanged too. Explicit non-goals are a guard against an agent helpfully expanding the task.
Then separate the work into reviewable concerns. A single pull request may need to land atomically, but its commits, description, or change map can still distinguish a schema migration from an interface change and a test harness. Keep dependencies clear and call out the places where a reviewer must reason across layers.
For higher-risk boundaries, define verification before implementation. If the change affects authorization, payment, data migration, or irreversible external effects, name the invariant and the evidence that will demonstrate it. Tests should cover meaningful paths; a passing suite is one input to a ship decision, not the entire decision. I use a risk-proportional approach rather than applying the same checklist to every change; A Green Test Suite Is Not a Ship Decision explains how I scale that evidence to impact.
A large PR can still be reviewable
A good PR description should let a reviewer build a mental model before reading every line. I try to include:
- Intent: the user or system behavior this change is meant to alter.
- Boundaries: the layers and components affected, plus explicit non-goals.
- Decision notes: important alternatives and assumptions that are not obvious from the diff. A lightweight decision record preserves that rationale beyond the review itself.
- Risk map: the files or transitions where mistakes would have the greatest impact.
- Evidence: focused tests, broader checks, browser behavior, and any remaining gaps.
- Release plan: staged exposure, monitoring signals, and rollback steps where relevant.
Organize the diff so the reviewer can follow a path: contracts and data first, implementation next, then tests and presentation. If the work needs multiple independently deployable pieces, split it into stacked changes rather than forcing one huge review. If it must be atomic, preserve the logical boundaries inside the PR and say why.
Review attention should be allocated, not diluted. A reviewer can often skim mechanical generated output after checking its source and regeneration path. They should spend more time on domain decisions, permissions, data transformations, and side effects. That only works when the author marks the distinction honestly; “AI generated this” is not a reason to skip scrutiny.
Keep rollback real
Reversibility is not a sentence in a checklist. It depends on what the change does. A feature flag can stage exposure, but it does not undo a destructive data migration. Feature flags are architectural branches with ownership and cleanup costs, not a substitute for a rollback plan. That plan should answer what happens to data already written, what users see if the feature is disabled, and who can safely execute the recovery.
When a change is difficult to reverse, reduce uncertainty earlier: rehearse the migration, preserve a backup, deploy compatible readers before writers, or split preparation from activation. The right mechanism depends on the system, but the review should make the irreversible step unmistakable.
A blast-radius review I can reuse
Before merging an agent-assisted change, I ask:
- What behavior changes for which users, and what must remain invariant?
- Which system boundaries are touched, including data, authorization, and external side effects?
- What did the agent assume, and which decisions did a human confirm?
- What are the highest-impact failure modes, and what evidence checks each one?
- Can the change be staged or isolated? If not, why must it be atomic?
- What is the rollback or recovery path, including data already changed?
- Can another engineer understand the change map and continue the work without the original agent session?
If the answers are vague, producing more code is not the next step. Clarify the boundary, add the missing evidence, or reduce the scope until the change can be controlled.
AI made implementation cheaper for me. It did not make judgment, review, or risk cheaper. The senior engineering job is not to defend a particular diff size; it is to make sure the change's consequences remain visible, bounded, and recoverable.
