← Back to news
AI-First Contributions: Repository Rules Put Humans Back in Control

Photo: Naval Surface Warfare Center / Wikimedia Commons (public domain)

19/09/2026

AI-First Contributions: Repository Rules Put Humans Back in Control

The new problem is no longer just generated code

Software teams are entering a phase in which contributions increasingly arrive with an agent somewhere in the loop. The question is no longer only: should we use AI to code? It is becoming much more operational: how do we stay in control when pull requests, test fixes, documentation updates, and refactoring proposals can be produced before a maintainer has rebuilt the context? This is a change of pace, but it is also a change of governance.

A recent GitHub Blog article about AI-first contributors describes this shift clearly from the open source side. Maintainers such as the AutoGPT team no longer ask whether agents will open pull requests. They assume these contributions already exist and organize the repository so they can be useful, verifiable, and rejectable. The interesting answer is not to close the door to every assisted contribution. It is to move the rules to the places agents actually inspect: instruction files, pull request templates, test requirements, CI checks, comment-resolution conventions, and access boundaries.

This matters for OrkestrAI and for every team that wants to delegate more work without delegating responsibility. An agent can accelerate the production of a diff, but it does not know the incident history, product compromises, legal risks, customer constraints, or reasons why an architecture was accepted despite its imperfections. Humans still carry that context. The repository therefore has to become an environment that helps the agent do the mechanical work while making it impossible to erase the human decision.

The repository becomes a management interface

In a traditional workflow, the repository mostly stored code, tests, documentation, and CI rules. With agents, it also becomes a management interface. Instructions are no longer written only for human newcomers; they must be readable by tools that will explore the project, change files, run commands, and propose updates. That is the logic behind files such as AGENTS.md, directory-specific guidance, or specialized skills that explain how to test, document, or respond to review.

The trap would be to believe that a long policy document is enough. Agents are literal, fast, and sometimes highly persuasive when they are wrong. A rule buried in an internal wiki protects nothing if the agent never loads it while changing the code. A review convention protects nothing if it is not connected to the pull request template, required checks, or a concrete action a maintainer can verify. Good practice is to move the rules close to execution: near the relevant module, inside the PR, inside CI, inside permissions, inside tests, and inside merge criteria.

This also changes the role of senior engineers. The work is no longer only to review line by line. It is to design the delegation system: which tasks may the agent attempt alone? Which commands require approval? Which data must never leave the environment? What proof of testing is acceptable? Which kind of change must be rejected even if the diff looks clean? The senior engineer becomes architect, reviewer, quality owner, and guardrail designer at the same time.

Good guardrails are visible and verifiable

The examples discussed by GitHub are revealing. Requiring a complete pull request template may look ordinary, but it is a powerful lever when it forces the author, human or agent, to state intent, scope, and test plan. An AI-generated PR that does not say how it was verified deserves a different level of trust from a PR accompanied by a reproducible scenario. The test plan is not bureaucracy: it is the minimum trace that allows the reviewer to understand what was tried, what was not tried, and where attention should be focused.

Another useful guardrail is making CI a wall, not a suggestion. When coverage thresholds, critical tests, linting, or security scans are mandatory, the agent can propose fixes, but it cannot redefine what ready means. That is a fundamental difference. In a fragile workflow, the agent produces a convincing diff and the human has to guess what is missing. In a governed workflow, the agent produces a diff, the repository applies its rules, and then the human judges the areas rules cannot cover: architecture, product fit, readability, maintainability, and operational risk.

Maintainers also mention very concrete details, such as requiring a specific commit before a review thread can be resolved. That kind of rule may look tiny, but it answers a real behavior: some agents may mark a conversation as resolved without actually fixing the code. Human control is therefore not a philosophical slogan. It is a collection of observable small locks that prevent dangerous shortcuts from turning into a merge.

Reviewing an agent PR requires a different checklist

A second GitHub article notes that agent-generated pull requests are already numerous enough to affect review bandwidth. The main risk is not always obviously broken code. The risk often comes from code that looks complete. An agent can reproduce patterns, add helpers, follow existing names, fill apparent gaps, and write a plausible explanation. But it may miss the deeper reason behind a constraint, duplicate logic that already exists, ignore an edge case known to the team, or forget a permission check on a rarely used execution branch.

Human review should therefore begin with intent: what problem does this PR claim to solve, and does the diff actually prove that understanding? Then come boundaries: empty inputs, size limits, permissions, network errors, migrations, compatibility, accessibility, performance, and rollback. Finally come proofs: is there a test that fails before the change and passes after it? Is the manual scenario reproducible? Are useful logs or screenshots attached? Does the human author understand the change well enough to defend it in production?

This checklist avoids two extremes. The first would be to reject every assisted contribution as noise. That would waste a real execution capability, especially for repetitive tasks, documentation, tests, modest fixes, and regression analysis. The second would be to accept speed as evidence of quality. A mature team does the opposite: it increases the surface area of automation while hardening the checkpoints where judgment is necessary.

Agentic workflows do not replace CI/CD

GitHub Agentic Workflows illustrates the same direction. The idea is to describe repository tasks in Markdown and let agents execute them inside GitHub Actions: triage, daily reports, test improvement, documentation, or targeted simplification. The most important message is not that everything can be automated. It is that agentic automation should complement CI/CD, not replace it. Deterministic pipelines remain essential for building, testing, analyzing, and releasing. Agents are more useful for ambiguous, repetitive, contextual tasks that prepare human work.

That distinction protects teams from a common confusion. An agent can open a documentation PR after detecting an API change. It can propose a test that covers a neglected branch. It can summarize the state of a repository in the morning. But the decision to release, change a security policy, modify a public contract, or remove compatibility must remain in an explicit governance frame. Automation accelerates preparation; it must not replace mandate.

In practice, each organization should classify tasks into three groups. First, tasks the agent may execute and propose with limited risk, such as a draft note or initial analysis. Second, tasks the agent may prepare but that require mandatory human review, such as a refactoring or bug fix. Third, tasks the agent must not execute without prior approval, such as operations involving secrets, infrastructure changes, irreversible migrations, or release actions.

What teams can do now

The first action is to make repository rules explicit. If humans must explain verbally to every newcomer how to test a module, the agent will not know either. Documenting test commands, business invariants, sensitive modules, forbidden patterns, and review criteria is no longer just good hygiene. It is a condition for assisted delegation to produce something other than extra work for reviewers.

The second action is to ask for evidence in every contribution. An AI-assisted PR should contain a clear scope, touched files, test plan, known limitations, and the points where the human author reviewed the result. If an agent helped, that is not a problem; saying so actually helps reviewers adapt. But the human author must remain responsible for the outcome. The label AI-generated must never become an excuse to drop a diff nobody understands.

The third action is to instrument the exit gates. Required tests, minimum coverage, security rules, branch protection, read-only defaults, and explicit approvals reduce the number of improvised decisions. They also give agents a clearer environment: this is allowed, this fails, this requires a human. The best guardrails do not slow the team down; they prevent the same debate from being repeated in every pull request.

Conclusion: the agent executes, the team governs

The historical illustration of the first computer bug is a useful reminder: from the beginning, software improved when humans made errors observable. Coding agents do not change that rule. They only make the loop faster and denser. The easier it becomes to produce a fix, the more essential it becomes to know who decides, what proof is required, which limits apply, and how the system prevents a shortcut from passing unnoticed.

For a modern team, keeping humans in the loop does not mean slowing AI down. It means designing an environment where AI can help without confusing speed with authority. The repository becomes the place where rules are encoded, CI becomes the first filter, human review becomes the space for judgment, and release remains an accountable decision. That combination, much more than code generation itself, will separate a brilliant experiment from sustainable engineering practice.