The new security habit is review before generation
AI coding assistants are not removing the human developer from security work. They are moving much of that work to a different place: from the moment of writing code to the moment of specifying, reviewing, testing, and accepting it. That shift is practical, not philosophical. A team can feel faster because an agent produces a working patch in minutes, yet the same team can become less safe if nobody has decided what the patch must not do, what data it must not touch, and which security assumptions must be checked before the pull request is merged.
Two recent signals make this a timely subject for software teams. The ACM Technology Policy Council warned in April that “vibe coding” can accelerate development while skipping practices that make software secure, reliable, and maintainable. A 2026 SOUPS paper on AI coding assistants and security awareness reaches a complementary conclusion: assistants tend to reorganize security thinking rather than eliminate it. Developers may write less of the first draft, but they must review more of the final behavior. Security becomes reactive unless the team deliberately makes it visible earlier.
For Paye ta com, the lesson is simple: an AI-assisted delivery process should not ask humans to rubber-stamp generated code at the end. It should ask humans to command the risk model at the beginning, supervise the agent during execution, and demand evidence before deployment. The human-in-the-loop is not a ceremonial approver. The human is the person who decides the threat model, the trust boundary, the acceptable trade-off, and the verification standard.
Why generation hides security decisions
A human developer usually makes many small security choices while typing. They decide whether an endpoint needs authentication, whether a field can be logged, whether a library is mature enough, whether user input requires normalization, and whether an error message reveals too much. These choices are often not written down because experienced engineers carry them as professional reflexes. The risk with AI assistance is that the reflex moves out of sight. The assistant receives a task such as “add password reset” or “create an admin export,” then fills in architecture, dependencies, error handling, and tests from patterns it has seen elsewhere.
That hidden expansion is useful, but it is also where security debt begins. A model can produce plausible code that works against a happy-path test while silently choosing a weak token lifetime, a broad database query, a verbose log line, or a missing authorization check. The problem is not that the model is malicious. The problem is that the task was underspecified and the generated answer looks complete enough to discourage deeper review.
This is why the security conversation must start before the prompt is sent. A good prompt for a coding agent should not only describe the feature. It should name the assets, attackers, permissions, failure modes, and constraints. Instead of “build a file upload endpoint,” the human should say: uploads are untrusted, file names must not be trusted, content type is not proof, stored objects must be private by default, malware scanning is required before publication, and the endpoint must be covered by tests for oversized files and path traversal attempts. The agent can still help. But now it is helping inside a boundary chosen by an engineer.
Reactive security is expensive security
When AI assistance pushes security into review, teams often discover issues late. Late discovery is costly because the generated patch has already created a story: the feature appears to exist, the demo works, the ticket is nearly closed, and the reviewer feels pressure to accept the shape of the solution. At that point, raising a security objection can look like slowing the team down, even when the objection is the work that protects users and the business.
Reactive security also creates an attention problem. Reviewers must inspect larger diffs than before, sometimes across files they did not expect the agent to modify. They must distinguish harmless generated boilerplate from subtle changes to authorization, serialization, dependency configuration, build scripts, or observability. The more the assistant produces, the easier it becomes for a dangerous line to hide in a sea of correct-looking code. A human review that only asks “does it compile?” is no longer enough.
The antidote is to make security evidence part of the definition of done. If an AI agent touches authentication, payment logic, personal data, infrastructure, dependency manifests, or background jobs, the pull request should contain explicit notes about the security property being preserved. Which users are allowed to call this path? Which data is read or written? Which tests prove denial as well as success? Which logs were checked for secrets? Which dependency or configuration change has been justified? These notes are not bureaucracy. They are the way humans recover visibility when machines generate faster than people can read line by line.
Command the agent with a threat model, not a wish
The practical pattern is to give every non-trivial AI coding task a small threat model. It does not need to be a long document. It can be a short checklist attached to the issue or prompt. The human names the asset, the entry point, the actor, the expected permission, the abuse case, and the test that would catch it. This turns the assistant from an autonomous guesser into a delegated implementer.
For example, an issue that says “add an export button” invites the agent to focus on a user interface and a data query. A safer issue says: “add a CSV export for organization admins only; never include password hashes, tokens, internal notes, or deleted accounts; enforce authorization on the server, not only in the UI; stream the response without loading all records into memory; add tests proving that a member without admin rights receives a denial; add a log event without personal data.” The second version is not longer because humans mistrust the tool. It is longer because humans understand the system.
This style also improves collaboration between product, engineering, and security. Product can state the user value. Engineering can state the implementation boundary. Security can state the abuse case. The AI can then draft code, tests, migration notes, and documentation within a shared frame. Everyone is faster because fewer assumptions are hidden.
Use AI to review AI, but keep a human accountable
There is nothing wrong with using a second model, a static analyzer, a dependency scanner, or a policy checker to review generated code. In fact, teams should do it. AI-assisted review can spot missing tests, suspicious permissions, insecure functions, outdated dependencies, or inconsistent error handling. It can summarize a large diff and point a human to files worth inspecting first. It can also generate adversarial test cases that the first agent did not write.
But automated review must not become automated absolution. A model that says “looks safe” has not accepted business responsibility. A scanner that is quiet has not proven the absence of flaws. A green test suite only proves the properties that were actually tested. The human reviewer remains accountable for deciding whether the evidence is sufficient for the risk. That means reviewers need the authority to block, narrow, or roll back AI-generated work even when the demo looks impressive.
A useful rule is to separate drafting from approving. Let the agent draft code and propose tests. Let tools produce review hints. But require a named human to approve changes that affect trust boundaries, secrets, identity, payments, compliance, or production operations. The named human should not be chosen at random. They should understand the area or have access to someone who does. Human-in-the-loop only works when the human has context, time, and permission to say no.
Make the workflow visible
Many AI risks come from invisible workflow. A developer accepts a large suggestion, an agent edits several files, a dependency appears in a manifest, a test is rewritten to pass, and the final diff looks like ordinary human work. Weeks later, nobody remembers which choices were generated, which were reviewed, and which were merely assumed. That is a governance failure, not because teams need theatrics, but because they need traceability.
Teams can improve this with lightweight practices. Mark pull requests that contain substantial AI-generated code. Ask agents to produce a change summary and a risk summary. Keep prompts or issue briefs for sensitive work. Require reviewers to state what they verified. Record when a tool executed commands, changed dependencies, or migrated data. These practices do not need to shame AI use. They normalize it while making responsibility clear.
Visibility also helps future maintenance. The developer who debugs the feature in six months should know why a guard exists, why a dependency was chosen, and why a test covers a strange denial case. AI-generated code that cannot be explained by humans becomes operational debt. AI-generated code that is documented, tested, and owned can become ordinary maintainable software.
What teams should change this week
First, change issue templates for AI-assisted work. Add fields for security assumptions, data touched, permissions required, and tests expected. The template should be short enough that people use it, but concrete enough that an agent receives more than a wish. Second, change prompts for risky features. Ask the assistant to identify security-sensitive files, propose tests before editing, and explain any dependency or configuration change.
Third, change review checklists. Reviewers should ask whether authorization is enforced server-side, whether secrets or personal data are logged, whether generated tests prove failure cases, whether dependencies are justified, and whether the code remains understandable to a human maintainer. Fourth, change merge rules for high-risk areas. An AI-generated change to authentication, billing, infrastructure, or data export should require a reviewer with ownership of that area.
Fifth, use AI defensively. Ask a separate assistant to attack the patch: “How could a user bypass authorization here?” “What input could crash this parser?” “Which log line could leak personal data?” “Which test would fail if the permission check were removed?” These questions turn the model into a review amplifier while keeping the human in command of the final decision.
The conclusion: faster code needs earlier judgment
The promise of AI-assisted development is real. It can reduce blank-page time, generate scaffolding, translate intent into working examples, and help teams explore alternatives. But the security cost is also real when generation replaces the small human judgments that used to happen during implementation. If teams wait until the final pull request to think about security, they will spend their gains on review fatigue, rework, and avoidable incidents.
The better path is not to reject coding assistants. It is to command them. Humans should define the risk, the boundary, the evidence, and the acceptance criteria before the machine writes too much code. Agents can then draft, test, summarize, and challenge. Humans decide whether the result is safe enough to ship. In an AI-assisted engineering culture, security review starts before code because responsibility starts before code.