A repository is full of instructions, but not all of them are for the agent
The fastest way to make an AI coding agent unreliable is to pretend that every line of repository content is equally trustworthy. A repository is not just source code. It also contains README files, issue threads, pull request comments, dependency metadata, changelog entries, shell scripts, example snippets, and documentation copied from elsewhere. Some of that text is meant to instruct a human. Some of it is meant to instruct a build system. Some of it is stale. Some of it is adversarial. The moment an agent is allowed to read the repository as context, the team has to decide which text is data, which text is instruction, and which text needs human review before it can influence behavior. It becomes essential the moment the agent can run commands, open pull requests, modify files, or act on behalf of the team.
This is why the current generation of agentic tooling keeps converging on the same pattern from different directions. GitHub’s agentic workflow architecture emphasizes isolation, constrained outputs, staged writes, and logging. OpenAI’s Codex guidance stresses bounded execution, sandboxing, and explicit approval for higher-risk actions. Anthropic’s prompt-injection work makes the threat model plain: whenever an agent processes untrusted content, an attacker can try to turn that content into instructions. They are the practical limits of delegation. A human in command is not a symbolic requirement; it is the only way to keep the interpretation of repository content from becoming arbitrary.
The useful model is simple: treat repository content as data by default, and only treat it as instruction when the system explicitly promotes it. A comment in an issue should not be able to rewrite policy. A paragraph in a README should not be able to authorize network access. A dependency page should not be able to tell the agent to reveal secrets or ignore tests. Those boundaries are obvious, but they fail first when productivity pressure rises and teams start relaxing controls in the name of speed.
Why this problem has become more visible now
Agentic coding has changed the shape of software work. Developers are no longer just asking models to draft code in the chat pane. They are delegating tasks that traverse the repository, inspect logs, run tests, compare files, and open pull requests. GitHub’s recent code review updates are a good example: the review model now gathers broader repository context to make feedback more accurate and less noisy. That extra context is useful. It is also part of the attack surface. The broader the context window, the more likely it is that the agent will encounter something that looks like instruction but was never meant for it.
That is the scenario Anthropic calls out in its prompt-injection guidance. Malicious text can hide inside legitimate pages, documents, comments, or other artifacts and attempt to redirect the model. In repositories it is easier to miss because the hostile text may sit inside an otherwise normal codebase. A malicious issue comment, a dependency README, a copied code example, or even a stale setup note can become a delivery mechanism. The agent only needs to be nudged into the wrong assumption about what it is allowed to obey.
Repositories contain history, not truth. They contain human intent, abandoned intent, and examples of bad instructions. An agent that cannot distinguish between those categories will eventually do the wrong thing with confidence. The goal is to make the system precise about what can influence decisions.
Promote context deliberately, or it will promote itself
Most teams already have a notion of trust boundaries in their infrastructure. The mistake is assuming those boundaries still exist once an agent starts reading files. A model does not know that a comment in a Markdown file is documentation while a command in a shell block is executable. It knows only that both are tokens in context unless the surrounding system tells it otherwise. That means agent design has to include a policy layer that classifies inputs before they shape behavior.
A practical version of that policy looks like this:
Repository content: readable, but not authoritative by default Issue and PR text: readable, but never a source of permission Documentation: readable, but must not override sandbox policy Code comments: informative only Machine-generated manifests: usable as data if validated External web content: untrusted until explicitly curated Human approval: required for cross-boundary actions
That list is boring on purpose. Boring is good. It is much easier to audit a workflow that says “this class of text may be read, this class may be summarized, this class may never authorize action” than one that relies on prompt phrasing and hope. A repository is not one giant prompt. It is a mixture of inputs with different security properties. A coding agent needs a classifier, not a mood.
This is also where “prompt injection” should be understood in a wider engineering sense. It is not just a clever trick hidden inside a malicious webpage. It is any situation in which untrusted text persuades an agent to reinterpret its own mission, permissions, or constraints. In software work, that can happen through the easiest channels: copied examples, tooling docs, installation notes, generated code, vendor README files, and issue templates. If a team has not assigned trust levels to those inputs, the agent will infer them, and the inference will eventually be wrong.
Why the human stays in command
People sometimes hear “human in the loop” and picture a person clicking approve on every little action. That is not the useful version. The useful version is narrower and stronger at the same time. The human defines the policy that decides what can be read, what can be written, what can reach the network, and what must stop for review. The agent can then do the routine work inside those rules. The human is not there to micromanage every keystroke. The human is there to own the trust boundary.
That distinction matters because approval fatigue is real. If an agent interrupts the workflow for every safe, reversible action, people will eventually widen permissions or ignore guardrails. Low-risk edits can happen inside the sandbox. High-risk changes require a deliberate review step. Cross-system actions, production access, and secret-bearing operations should remain explicitly human-owned.
OpenAI’s Codex guidance is useful here because it treats approvals and sandboxing as complementary, not competing. The sandbox defines the technical boundary; approvals define the policy boundary. A system that cannot write outside its workspace should not need a long discussion to justify that limit. A system that wants to touch production or reach an unfamiliar domain should be stopped by design.
Telemetry is part of the control plane, not an afterthought
Trust boundaries are easier to maintain when the workflow is observable. If the team cannot reconstruct what the agent read, what it tried, what it changed, and why it stopped, then the boundary is already weaker than it looks. GitHub’s agentic workflow architecture is explicit about logging everything because logs are what make review, rollback, and incident response possible. OpenAI’s Codex guidance makes the same point with agent-native telemetry: the system should preserve enough information for operators to understand intent and outcome, not just process exit codes.
This matters especially when the agent is operating over repository content of mixed trust. The question is not only whether the agent accepted a malicious instruction. The question is whether the team can see how that instruction entered the context, whether it was promoted or ignored, and whether similar instructions are now present elsewhere in the repo. Good logs turn one strange incident into a pattern you can investigate. Bad logs turn it into a story with missing pieces.
Telemetry also changes behavior before anything goes wrong. When engineers know that tool calls, approvals, and content sources are being recorded, they design workflows that are easier to reason about. The best guardrail is not just a barrier. It is a record that tells the next engineer what happened and why.
The failure modes are usually ordinary, not cinematic
It is tempting to imagine prompt injection as a dramatic takeover: a malicious comment causes the agent to exfiltrate secrets, delete code, or rewrite a repository. Those failures do happen in red-team exercises, but the more common problem is subtler. The agent reads a stale instruction and spends time optimizing for the wrong target. It treats example text as policy. It follows a dependency README more literally than it should. It copies an unsafe command from a docs page into a workflow. It expands the scope of a fix because the surrounding text made the extra step sound reasonable.
These failures are dangerous precisely because they look like normal productivity. A good agent is supposed to be fast, autonomous, and willing to fill in gaps. That same behavior becomes a weakness when the gaps are in trust, not in syntax. A small error in context selection can lead to an apparently sensible patch that is actually based on the wrong premise. The team then spends time reviewing a diff that was generated under a false assumption. This is why review quality depends on input hygiene. If the context is dirty, the diff is contaminated before a human ever sees it.
Another common failure mode is over-broad authority. The agent is allowed to inspect the repo, so it starts behaving as if every related artifact is fair game. It is allowed to edit files, so it begins making opportunistic changes outside the original task. It is allowed to see documentation, so it begins trusting documentation more than policy. Those are not model bugs alone. They are delegation bugs. The system gave the model more authority than the task required and then acted surprised when the model used it.
What a sane operating model looks like
A good operating model for AI-assisted development does not try to eliminate ambiguity. It contains it. It gives the agent enough room to be helpful while keeping the critical decisions human-owned. In practice that means a few concrete rules.
- Separate readable context from authoritative instructions.
- Keep secrets out of agent-readable memory by default.
- Allow writes only to the smallest necessary scope.
- Require approval for network access, production access, and irreversible changes.
- Log tool calls, context sources, approvals, and rollbacks.
- Review the policy regularly, not just the final diff.
Those rules are not glamorous, but they are exactly the kind of unglamorous discipline that lets teams use AI without giving up control. GitHub’s security validation for third-party coding agents points in the same direction: even when an external agent generates code directly in a repository, the platform still runs CodeQL, dependency checks, and secret scanning before the work is considered safe. That is the right instinct. Generation does not equal trust. Validation is what turns a patch into something the team can accept.
One useful habit is to ask a simple question whenever the agent suggests a change: “What category of text convinced it this was a good idea?” If the answer is a source file, good. If it is a README paragraph, an issue comment, a copied snippet, or a vendor doc with unclear provenance, the team should pause. That question is not about distrusting the model. It is about understanding what part of the environment shaped the decision.
Conclusion: keep the boundary visible
The argument for AI-assisted development is strong when the task is bounded, the policy is explicit, and the output can be verified. The argument collapses when the system assumes that all repository content is equally trustworthy. A repository contains useful context, but it also contains stale conventions, copied examples, accidental instructions, and sometimes deliberate traps. If an agent is allowed to treat that mixture as one clean source of truth, the team has not built an assistant. It has built an easily confused delegate.
The better model is straightforward: let the agent read broadly, but trust narrowly. Let it draft, inspect, and summarize, but do not let it promote text into instruction without a policy decision. Keep the dangerous transitions visible. Keep the logs complete. Keep the human responsible for the boundary. That is what makes the system safe enough to use and honest enough to improve.
The repository is not the boss. The human is. That is the control plane worth preserving, and it is the only way AI coding tools become more than fast ways to automate mistakes.