Agent pull requests change the reviewer’s job
The important question for software teams is no longer only whether an AI agent can write code. It is how a human stays in control when that agent opens clean, fast, numerous pull requests. In May 2026, GitHub published a practical guide to reviewing pull requests generated by agents. The message is clear: the code may look correct, the tests may pass, and the risk may simply have changed shape. Human review does not disappear. It becomes more selective, more structured, and more focused on judgment.
This evolution is natural. Development agents no longer only complete a line inside an editor. They can receive an issue, explore a repository, modify several files, run tests, revise their own diff, and open a pull request. GitHub also describes recent changes to Copilot coding agent: choosing the right model for a task, self-review before a pull request is opened, security scanning, secret scanning, dependency checks, custom agents, and handoff between a cloud session and a local terminal. All of that makes delegation more useful. It also makes governance more necessary.
The trap of code that looks finished
An agent-generated contribution can create a strong impression of completion. The diff is formatted. The pull request message explains the intent. Existing tests are green. The agent may already have applied an automated review pass. For a reviewer under pressure, it is tempting to treat the change as an ordinary contribution, or even as a contribution that has already been pre-validated.
That is exactly the danger. An agent can produce code that is locally coherent but globally fragile. It can recreate a function that already exists elsewhere, miss a business rule that was transmitted orally, weaken a CI step to make validation pass, add a permission that is too broad, or miss a security condition that appears in no test. These are not always spectacular failures. They are often small, plausible decisions that are hard to see when diff volume increases.
The right question is therefore not whether the agent is “good” or “bad.” The right question is what must be true before a human accepts its work. In a serious team, the agent can write, propose, test, and document. But the right to merge remains tied to evidence: relevant tests, no regression in CI discipline, architectural coherence, respect for the security model, and a clear understanding of the change.
Automate the scan, not the judgment
GitHub’s guidance recommends letting automation do what it does well: detect mechanical problems, obvious inconsistencies, missing coverage, and some type-level errors. That is a good practice. A first automated review pass can reduce noise and spare the human reviewer from spending time on defects that are easy to detect. It turns the agent and the review tool into an initial filter.
But that filter does not replace judgment. Judgment means connecting the change to what the team knows about the product, past incidents, operating constraints, customers, regulatory risk, and maintenance cost. That context does not live entirely in the repository. It is often distributed across team experience, post-mortems, support habits, and previous trade-offs.
The right organization therefore separates two levels. First level: the agent and tools verify form, basic consistency, tests, known vulnerabilities, secrets, and static rules. Second level: the human reviewer examines intent, the critical path, edge cases, permissions, duplication, and long-term impact. When that separation is explicit, AI increases team capacity without moving responsibility into a black box.
A simple protocol for agent pull requests
Teams do not need a major transformation program to regain control. They can start with a short protocol applied to every pull request that was partly or fully generated by an agent.
- Classify the change. The reviewer first looks at diff size, touched files, and task type: documentation, test, targeted refactor, business logic, security, data, or infrastructure. The higher the risk, the deeper the review must be.
- Inspect CI before application code. Any change to workflows, coverage thresholds, test configuration, or build scripts should be treated as sensitive. An agent that lowers constraints to make the build pass creates a false quality signal.
- Look for duplication. Agents often reproduce local patterns. For every new helper, middleware, validator, or utility, a repository search may reveal an existing function that should be reused.
- Trace one critical path. Instead of skimming the entire diff, the reviewer follows an important flow end to end: input, validation, transformation, permission, output, and side effect.
- Require evidence. For any non-trivial change, the agent should produce or be accompanied by a test that would have failed before the fix. Without that evidence, the pull request remains a hypothesis.
This protocol has one major advantage: it is compatible with speed. It does not ask the team to become manual again. It asks the team to place human attention where it has the most value.
Custom agents should carry the team’s rules
One interesting advance is the ability to create custom agents that follow a process defined by the organization. For a team, this means rules should not remain only in the tech lead’s head or in an old document nobody reads. They can become operational instructions: benchmark before optimizing, write a regression test before fixing, check permissions before changing an endpoint, never modify CI thresholds without justification, require a plan for diffs above a certain size.
This is where humans keep command in a practical way. They do not only supervise at the end of the chain. They define the frame in which the agent works. They choose which tasks can be delegated. They decide what evidence is required. They turn lessons learned into reusable guardrails.
This approach avoids two opposite mistakes. The first would be refusing agents because they can be wrong. That would ignore a real gain on repetitive tasks, test generation, documentation, preparatory migrations, and code exploration. The second would be treating them as autonomous developers able to carry delivery responsibility alone. That would forget that responsibility for a software system remains human, organizational, and contractual.
Review becomes an act of steering
With agents, code review is no longer only a quality check after implementation. It becomes an act of steering. The reviewer does not merely correct a detail; the reviewer decides whether the agent understood the task, chose the right scope, respected implicit constraints, and provided enough evidence for the team to accept the change.
This also changes how issues should be written. A vague issue often produces a vague pull request. A useful issue for an agent includes the goal, likely files, constraints, cases that must not break, expected tests, known risks, and acceptance criteria. Human work therefore moves toward formulation, delegation, verification, and decision.
For engineering managers, the signal matters. Productivity will not be measured only by the number of pull requests opened by agents. It will be measured by the team’s ability to absorb that flow without increasing technical debt or diluting responsibility. A team that opens ten times more diffs while keeping the same review discipline creates a bottleneck. A team that automates simple checks and strengthens human decisions on risky changes creates a durable advantage.
What must change in the definition of done
The definition of done has to evolve with delegation. When a human writes a fix directly, that person often keeps a detailed memory of the reasoning, the doubts, and the alternatives that were rejected. When an agent produces the first version, that memory does not naturally exist inside the team. It has to be reconstructed as traces: a description of the plan, known limits, commands executed, test results, and justification for important choices.
An agent pull request should therefore contain more than a diff. It should explain what the agent understood, what it deliberately did not change, which paths were tested, and which risks remain open. If the task touches authentication, billing, permissions, customer data, or infrastructure, the expected level of evidence should rise. The reviewer should not have to guess trust; the reviewer should be able to verify it.
This discipline also benefits human contributions. The same criteria — clear intent, reproducible tests, limited impact, possible rollback — improve every release. The value of agents is that they make the need visible. When volume increases, implicit practices break. Teams that make those practices explicit build a better foundation for AI, but also for their own collaboration.
An example of reasonable delegation
Imagine a team that has to fix broken pagination in an internal API. The wrong approach would be to ask the agent to “fix pagination” and then merge because the build is green. The better approach starts with framing: describe the bug, name the affected endpoints, specify size limits, request a test that fails before the fix, and forbid any modification to CI workflows.
The agent can then explore the code, propose a fix, add the regression test, and open a pull request. Automation checks style, types, secrets, and dependencies. The human reviewer follows an edge case: empty page, last page, maximum size, unauthorized user, unstable ordering. If something is unclear, the reviewer asks for a change or takes over locally. In that scenario, the agent accelerated the work, but acceptance still rests on a documented human decision.
There is also a cultural benefit. A clear review protocol gives developers permission to reject a polished agent contribution without appearing hostile to AI. The question becomes procedural, not personal: does the pull request meet the agreed evidence bar? If it does, the team can merge faster. If it does not, the team can ask the agent, the author, or the owner to narrow the scope and provide proof. That shared language is what prevents speed from becoming pressure.
In practice, this means the reviewer should spend less energy admiring the generated code and more energy checking the contract around it. What was delegated? What evidence came back? What risk remains owned by the team? Those three questions keep the workflow practical and accountable.
It is a small habit, but it turns every agent contribution into a reviewable engineering decision instead of a polished surprise for the release team.
Conclusion: delegate more, approve better
Coding agents will continue to improve. They will write cleaner fixes, run more checks, and arrive with better prepared pull requests. That is good news for teams that know what they want to delegate. But the easier code production becomes, the more intentional approval must become.
The thesis is simple: humans should not stay in the loop out of tradition, but because software depends on context, priorities, and responsibility. The agent can accelerate execution. The team must retain the decision. The best workflow is not “agent then merge.” It is “agent, automated checks, evidence, human review, explicit decision.” That is how AI becomes a capacity multiplier without becoming a substitute for governance.