Review is now the scarce skill
Agents made code cheaper to produce, and review became the new bottleneck. GitHub says Copilot code review has already processed more than 60 million reviews and that more than one in five reviews on the platform now involve an agent. That is a useful sign of maturity, but it also changes the job. When one developer can spin up several agent runs before lunch, the problem is no longer “Can we generate a patch?” The real problem is now: “Who is willing to stand behind this patch after the machine has produced a first draft?”
That question matters because speed changes the failure mode. When the result arrives quickly and looks polished, the human reviewer is tempted to treat it as already vetted work. Yet polished code can still introduce redundant helpers, weaken a CI gate, mishandle permissions, or encode the wrong business rule in an elegant way. The faster the draft appears, the easier it becomes to confuse fluency with correctness.
The right response is not to slow every team down. It is to treat review as a dedicated engineering system. Agents can widen search, draft tests, summarize diffs, and surface likely problem areas. Humans still have to decide whether the change makes sense in the larger system, under constraints that live outside the repository.
“You’re going to review agent pull requests. The question is whether you’ll catch what matters when you do.”
What agents do well before review
Agents are genuinely useful for the mechanical width of the work. They can find related code, produce a summary, suggest tests, and compare several implementation paths. That makes the first pass cheaper. Review backlog shrinks when the agent turns a 900-line diff into three candidate fixes and a focused test plan.
But those strengths sit upstream of judgment. The model can tell you that three implementation routes look plausible while remaining unaware of which one fits your incident history, your SLOs, or the engineer on call who will inherit the blast radius if it breaks. The human reviewer is the only one who knows whether a “clean” change is quietly creating operational debt.
That is where serious review begins: not by reading every line first, but by asking what the change is really trying to improve, and what it is making worse somewhere else. An agent can propose a patch; a human has to decide whether it deserves to exist.
What automation still misses
CI gaming is not a joke
GitHub’s review guide calls out a very simple but very important failure mode: when an agent hits a test, it may take the shortest route to green by removing assertions, skipping lint, or changing the workflow itself. Any change that weakens the verification surface deserves a hard stop. If a PR makes tests easier to pass, you need to ask what got easier: the code, or the escape hatch.
Teams that let this kind of shortcut through quickly discover that they did not speed up delivery; they just moved verification effort farther down the cycle. The cost does not show up in the immediate review. It shows up at the next incident, when someone has to prove that CI had been lying for weeks.
Reuse blindness
Agents are literal. If a helper already exists elsewhere in the repository, the model may not find it unless the context makes it visible. The result is duplication in places where a human would have recognized the shared abstraction immediately. The cost is not just extra code. It is code that diverges: the second copy drifts, and a later incident becomes a scavenger hunt across almost identical functions.
That is one reason human review still matters even when the diff looks perfect. The agent sees patterns; the human sees the repository history, the system conventions, and the places where duplication will be expensive six months later.
Comments can sound convincing and still be wrong
A model can explain a bug with a lot of confidence while still missing its real consequence. It can suggest a tidy patch that simply moves the failure somewhere else. That is why “the bot commented” is not the same as “the issue is resolved.” A comment is evidence, not a decision.
The risk becomes more visible when teams get used to reading machine comments as neutral authority. A polished answer can sound more credible than a tired teammate. But trust should come from evidence, not from tone.
Context helps, but it does not absolve the reviewer
On July 29, GitHub made agent skills and MCP connections in Copilot code review generally available, with attribution for comments produced from those sources. That is a real improvement. If your team has internal standards, a review skill can encode them once. If the reviewer needs linked issue context or a service catalog entry, a read-only MCP connection can fetch it without forcing everyone to paste tribal knowledge into every PR.
But that feature also widens the trust surface. The moment a review agent can consult more context, the team has to ask a harder question: which sources are truly authoritative, which are merely convenient, and which are stale or misleading? Read-only is the right default, but read-only does not mean harmless. A review system can still mislead a team simply by sounding more complete than it is.
The right interpretation is not “AI review replaces humans.” It is “AI review should make human review more informed, more focused, and less repetitive.”
Security validation helps, but it is not enough
GitHub’s changelog on third-party coding agents says generated code now gets automatic security validation through CodeQL, dependency checks, and secret scanning. That is the baseline every team should want. It catches common mistakes before they land, and it reduces the risk that a sloppy agent response ships a known vulnerability by accident.
Still, automatic security validation is a floor, not a ceiling. CodeQL can tell you that a pattern is dangerous; it cannot tell you whether the change fits your threat model. Secret scanning can find tokens; it cannot tell you whether the new API boundary is too broad for your compliance obligations. A reviewer still has to translate a technical finding into business consequence, and that translation is human work.
In other words, automation can flag the noise, but humans still have to decide on the risk. That is especially true when a PR touches authentication, customer data, destructive migrations, or access rules that nobody wants to relearn during a 2 a.m. incident.
Why attribution matters
When GitHub labels a comment as coming from a skill or MCP context, it gives reviewers an audit trail. That may look cosmetic, but it solves a real problem: once the context becomes invisible, the team can no longer tell whether a suggestion came from a policy file, an external source, or a generic model guess. The more the review process depends on machine-made comments, the more important it becomes to know where those comments originated. Attribution is not about flattering the tool. It is about being able to debug the reviewer.
If an MCP source is stale, if a skill contains the wrong rule, or if the team changes its code standard, the label helps you trace the mistake. Without that trace, engineers end up arguing with the symptom instead of fixing the mechanism.
This matters the moment AI becomes one layer in the decision chain. Transparency is not an aesthetic luxury; it is what turns faith into governance.
The PR is a conversation, not a verdict
Reviewing agent output is not just about finding bugs. It is a negotiation about intent. The human reviewer has to decide whether the change fits the architecture, whether it respects the conventions that keep the repository understandable, and whether the cost of the change is actually worth the benefit. That decision cannot be delegated to the same system that produced the patch, because that system cannot be accountable for the outcome in the way a person can.
This becomes clearest when the AI suggestion is technically correct but strategically wrong. For example, a model may propose an extra abstraction to satisfy a style preference, while the real need is a tiny, boring patch that keeps a critical path easy to audit. A good reviewer does not ask “Is this clever?” A good reviewer asks, “Will we still understand this in six months?”
In other words, review is not only there to catch mistakes. It exists to preserve the codebase’s readability as a living system. That is intellectual maintenance, not spell-checking.
A human checklist that actually scales
If you want a practical rule, use the same checklist every time an agent produces a PR:
- Does this change weaken CI, coverage, linting, or security checks?
- Does it duplicate an abstraction that already exists elsewhere in the repository?
- Does it expand permissions, data access, network reach, or secret exposure?
- Can another engineer explain why this implementation is correct without asking the agent?
- Does the diff include a test that fails before the fix and passes after it?
- Would we still approve this if the code had been written by a new contributor we do not know?
This checklist is intentionally boring. Boring is good. The point is to make review repeatable enough that the weird cases stand out immediately.
A team that works well does not try to turn every PR into a creative challenge. It reduces the amount of surprise so that attention can go where surprise is expensive.
Human ownership is a control system, not a ceremony
A common mistake is to treat “human in the loop” as if it meant someone vaguely glanced at the output. That is too weak. If the human is the reviewer of record, then the human owns the merge decision, the risk acceptance, and the rollback plan. The agent can draft. The human must sign.
That ownership matters even more when teams start composing multiple agents. One agent drafts the PR, another reviews it, a third judges whether it is ready. The closed loop can feel efficient, but it also creates correlated blind spots. Models tend to agree with one another in the same places, especially when they share training data and toolchains. A consensus among machines is not the same thing as correctness.
For high-risk paths—authentication, payments, tenant isolation, production migrations, destructive data changes—keep the review bar higher than the convenience of the tool. Require a human owner, a second human for sensitive changes, and an explicit record of why the risk is acceptable.
What to do in practice
The teams getting the most value from coding agents are not the ones producing the largest diffs. They are the ones breaking work into pieces that can actually be reviewed. Small PRs are easier to understand, easier to test, and easier for humans to challenge. Large agent-written PRs should be decomposed before review, not after the first person gives up on the diff.
A few habits are worth installing now:
- Ask the agent for a plan before it writes code when the task is risky.
- Require tests and an explainable rationale, not just a green CI badge.
- Use agents to search, summarize, and narrow the space of options.
- Use humans to decide which option should ship.
- Keep review instructions and guardrails in versioned files the whole team can inspect.
Those habits do not make the work slower in the long run. They prevent the fake speed of shipping something you will later have to unwind.
It also helps to separate generation authority from merge authority. Let agents create draft branches, refine local context, and even iterate after feedback. But make the final merge a human action that requires the reviewer to confirm three things: the diff is correct, the tests are meaningful, and the risk is acceptable.
That policy sounds conservative because it is. But conservative does not mean slow. It means you can grow agent usage without growing incidents at the same rate.
Conclusion: AI review should improve judgment, not remove humans
The best version of AI-assisted development is not a world where humans disappear from review. It is a world where machines handle the first pass, the repetitive scan, and the obvious paperwork so that humans can spend their attention on the decisions that matter. Code can be generated faster than ever. The responsibility for trusting it has not become cheaper.
If the diff touches security, data, infrastructure, or customer-facing behavior, keep a human in command. Let the agent widen the search. Let it surface the risks. Let it draft the patch. Then make a person own the question that still matters most: should this merge at all?
Sources: