← Back to news
OWASP and AI reviews: ship faster without losing human accountability

Photo: Robert Scoble / Wikimedia Commons (CC BY 2.0)

22/09/2026

OWASP and AI reviews: ship faster without losing human accountability

Why this matters now

Coding agents are no longer just assistants that complete a line inside an editor. In 2026, they read tickets, modify multiple files, install dependencies, run commands, summarize errors, propose fixes, and open pull requests. That evolution is useful for software teams, but it changes the shape of risk. When an agent acts with a developer’s permissions, a bad instruction, an invented dependency, or an overconfident review can become a security incident, a regression, or technical debt that is hard to reconstruct later.

OWASP’s publication of a practical Secure Coding with AI cheat sheet gives this conversation a concrete frame. It does not say developers should reject agents. It says the opposite: if agents are becoming a normal part of the development chain, they must be treated as both an attack surface and an engineering process that needs governance. That is exactly the posture teams need if they want to save time without turning AI into a black box.

The central message is straightforward: AI can help produce, analyze, and correct code, but it should not become the implicit owner of the decision. Responsibility remains human. A developer, a team, or an organization must decide what the agent may read, what it may execute, which dependencies may enter the product, which tests actually matter, and who approves release.

From assistance to action: the operational jump

The first change is operational. A classic assistant suggests text; an agent acts. It can call a shell, write into the repository, change CI configuration, query external documentation, use MCP servers, or publish a branch. That capability removes entire loops of work, but it also gives the model a role inside the software production system. From that point on, the usual controls should apply: least privilege, logging, environment separation, independent review, and explicit policy.

A coding agent also operates in a world full of untrusted inputs. A GitHub issue, a pull request comment, a README file, an error trace, or a web page can contain instructions the model treats as relevant. OWASP emphasizes this because it turns prompt injection into a supply chain problem. A malicious instruction does not have to be sent directly by a developer; it can be hidden in the context the agent consults while doing its work.

For a technical leader, the practical conclusion is not “ban the agent,” but “bound the agent.” Restricting context, disabling network access when it is unnecessary, running commands in a sandbox, blocking access to production secrets, and reviewing unexpected changes are reasonable controls. They make the agent more reliable without removing its value.

AI-suggested dependencies need verification

One very concrete risk is dependency selection. Assistants can suggest package names that are nonexistent, outdated, or close to real libraries. Attackers can then register those names on public registries and wait for developers to install them. This is especially dangerous with agents running in auto-accept mode, because a suggestion can become an installation command without a human pause.

The answer is procedural. Every AI-suggested dependency should be verified before it enters the project: real existence, maintainer, package age, activity, known vulnerabilities, popularity, license, and fit with the existing architecture. Automated audits such as npm audit, pip audit, govulncheck, OSV, or the GitHub Advisory Database do not replace judgment, but they create a necessary safety net.

This discipline avoids a common mistake: measuring an agent’s quality only by the number of tasks it completes. An agent that quickly adds an unnecessary or fragile dependency can make today faster and tomorrow slower for the whole team. The senior developer’s role is to turn the proposal into a decision: accept, reject, replace, or isolate.

AI review is not approval

GitHub now documents broader Copilot code review capabilities: repository context analysis, pull request comments, suggestions that can be applied quickly, and the option to delegate certain fixes to a cloud agent. That is useful, especially for spotting inconsistencies, asking for readability improvements, or accelerating feedback on small changes. But the documentation also states that Copilot may miss problems and make mistakes. It recommends validating its feedback carefully and supplementing it with human review.

That nuance matters. An AI-generated review can be an excellent first filter, but it should not become a signature. The danger appears when a team stacks two automations: one agent writes the fix, another agent reviews it, and a green pipeline creates a feeling of safety. If no one checks intent, business invariants, security, and side effects, the chain is fast but not accountable.

The best practice is to make AI review visible and subordinate. AI comments should help the human reviewer, not replace them. Merge rules should clearly state that a human owner approves the change. Agent-generated pull requests should be identifiable, linked to an explicit request, accompanied by a summary of changed files, and held to the same standards as human-written code.

Agent-generated tests are not enough

Agents are very good at producing plausible tests, but a plausible test is not necessarily a good oracle. It can freeze the broken behavior the agent just created, avoid edge cases, remove an inconvenient assertion, or artificially increase coverage. This is one of the most important traps in AI-assisted development: confusing passing tests with validation of the requirement.

A robust workflow separates roles. The agent can propose tests, but a human must verify what those tests prove. For critical areas, the team can require independent verification: existing tests protected against deletion, business assertions written before generation, review by a developer who did not drive the agent, static analysis, fuzzing, or execution in an isolated environment. The question is not “is the pipeline green?” but “does the pipeline measure the right risk?”

This distinction is especially important for small teams. When delivery pressure is high, the agent can create the impression that everything is under control: code produced, tests added, convincing summary written. Yet the summary is part of the output of the same system. The final decision must stay attached to external evidence: readable diff, meaningful tests, logs, metrics, and human review.

Design the workflow with stop points

The right architecture for human-agent collaboration is not total autonomy. It is delegation with stop points. Before the work starts, the human defines the goal, scope, allowed files, available tools, and acceptance criteria. During the work, the agent executes, explains, and leaves traces. After the work, the human reviews, requests corrections, checks dependencies, runs tests, and decides whether the change can be merged.

This model works because it uses each side’s strengths. The agent is fast at exploring, applying a repetitive pattern, producing a first draft, or summarizing a codebase. The human is better at understanding organizational context, arbitrating risk, rejecting an attractive but fragile solution, detecting business inconsistency, and taking responsibility for release. Productivity comes from the combination, not from erasing either role.

Teams can formalize the model with simple rules: no auto-accept on sensitive repositories, no access to production secrets, no new dependencies without validation, no modification of existing tests without justification, no merge without a human owner, and logs of the tools used. These rules do not slow innovation; they prevent innovation from depending on luck.

What teams can do this week

The OWASP cheat sheet and GitHub’s Copilot review documentation converge on the same idea: agent adoption needs a control system. A team can start modestly. It can inventory the AI tools in use, identify where auto-accept is enabled, check which secrets are reachable from development environments, document review rules for agent-generated pull requests, and add a dependency-audit step.

  • Define a human owner for every AI-assisted change.
  • Limit permissions to the repository, files, and tools needed for the task.
  • Verify dependencies proposed by AI before any installation.
  • Keep traces: initial prompt, agent summary, commands executed, and files changed.
  • Use AI review as assistance, never as final validation.

These measures are pragmatic. They do not require a large governance platform or a heavy committee. They simply establish that the agent is part of the development process and therefore must be observable, limited, and verified.

The technical manager’s role changes too

Agent governance is not only a tooling question. It changes how a lead developer, CTO, or security owner runs the team. The organization has to clarify which tasks may be delegated, which repositories require stronger supervision, how to document a decision made with AI assistance, and how to train junior developers to review model output without being intimidated by it. The risk is not that developers use too much AI; the risk is that they use it without a shared language for trust, uncertainty, and evidence.

A useful ritual is to treat the first agent-generated pull requests as learning cases. The team can inspect the diff together, compare the agent’s summary with the actual changes, identify unexpected files, note dependencies that were added, and ask what a quick review could have missed. This develops a new skill: supervising software work that was partly produced by a machine. It also strengthens accountability, because everyone sees that speed only has value when it remains explainable.

Over time, the most effective organizations will not be the ones that let agents act everywhere without friction. They will be the ones that adjust the level of control to the level of risk: broad autonomy for internal documentation or mechanical refactors, strict validation for security, payments, personal data, and irreversible migrations. Maturity means consciously choosing where to save time and where to slow down on purpose to protect the product.

For platform teams, this also means that agent policy should live close to existing engineering controls. Repository rules, branch protection, CI checks, dependency allowlists, secret scanning, and observability dashboards already express how the organization manages risk. AI agents should plug into that system rather than bypass it. When an agent opens a pull request, the same gates should run, and when those gates fail, the agent’s output should be treated as a draft that needs correction, not as evidence that the process is too strict. This keeps the conversation familiar for developers: the tool is new, but the engineering discipline remains recognizable. The practical goal is predictable acceleration, not unsupervised heroics for teams.

Conclusion: accelerate without delegating responsibility

Coding agents will keep improving. They will become more integrated with tickets, team discussions, IDEs, CI pipelines, and review systems. That is good news for developers who want to reduce repetitive work and focus on architecture, quality, and delivery. But the more capable the agent becomes, the more governance matters.

The durable approach is not to choose between total trust and rejection of AI. It is to build a workflow where AI accelerates execution while humans keep the decision. OWASP highlights the risks, GitHub highlights the limits of automated review, and software teams now have a clear roadmap: delegate the work, not the responsibility.

Sources