Delegation needs a place to live
The next useful step in AI-assisted software development is not giving coding agents more independence. It is giving human delegation a durable control plane. The timely signal is OpenAI’s Symphony specification for Codex orchestration: instead of asking engineers to supervise a pile of agent sessions, the issue tracker becomes the system of record. Work is expressed as tasks. Each task can map to an agent workspace. Humans review outcomes, file follow-up work, and keep the board aligned with product priorities. That may sound like a workflow detail, but it is much more important than another demo of a model writing code. It points to the real bottleneck in agentic engineering: not raw generation, but the human ability to assign, bound, observe, interrupt, and accept work without losing the plot.
This is exactly where the human-in-the-loop thesis gets practical. A developer cannot responsibly “orchestrate” AI by watching five terminals scroll at once and hoping the right context remains in working memory. That style of supervision does not scale because it treats the human brain as the queue, the dashboard, the audit log, and the approval layer at the same time. The moment several agents are running in parallel, coordination debt appears. Which task is active? What did the agent change? Which branch contains the promising attempt? Which failure produced useful knowledge? Which idea should be abandoned? If those answers live only in chat transcripts and terminal buffers, the organization has not adopted agentic development. It has adopted agentic noise.
The issue tracker is a better home for this work because it already represents intent, ownership, priority, acceptance criteria, and history. Those are the human parts of software development that do not disappear when code becomes cheaper to produce. In fact, they become more valuable. When an agent picks up a well-scoped issue, the team can see the reason for the work before seeing the diff. When the agent opens a pull request, reviewers can evaluate the result against the task, not against a vague promise in a prompt. When a reviewer discovers a larger architectural problem, the answer is not to let the agent improvise across the codebase. The answer is to file a new issue, make the tradeoff visible, and decide whether it belongs in the current plan.
Approval is becoming product infrastructure
GitHub’s recent change allowing Copilot code review to approve pull requests, when explicitly enabled, makes the same point from another angle. The feature is off by default and configurable at enterprise, organization, repository, and path level. That matters. The interesting part is not that an AI system can say a pull request looks ready. The interesting part is where the permission to count that judgment is placed. A repository can allow it for some paths and not others. Administrators can keep it disabled entirely. New commits dismiss the approval, just as they would for a human reviewer. In other words, the tool is useful because the human organization defines when the tool’s judgment has procedural force.
That is the right pattern for agentic development. Automated review should reduce the time humans spend on mechanical scanning, but it should not quietly replace responsibility. A code-review agent can catch style issues, suspicious logic, missing tests, and obvious security problems. It can provide a useful first pass before a senior engineer spends attention on the diff. But the crucial question remains human: should this change exist, in this shape, at this time, for this system? That judgment depends on incident history, customer promises, regulatory boundaries, operational constraints, and architectural direction. Much of that context is only partly written down. Some of it is distributed across people. A model can help surface evidence, but it cannot own the accountability.
The safest teams will therefore treat approval as infrastructure, not etiquette. They will encode which actions are reversible, which paths are sensitive, which workflows require human confirmation, and which changes must include tests that fail before they pass. They will distinguish between “the agent may propose” and “the agent may merge.” They will keep logs of decisions, not merely logs of tokens. They will also resist a subtle temptation: using automation to make risky work feel routine. The point of a control plane is not to remove friction everywhere. It is to put friction where the consequences justify it.
When code generation becomes abundant, the scarce resource is not typing. It is trustworthy human judgment applied at the right boundary.
The board can prevent agent sprawl
Agent sprawl is easy to create. A developer opens several sessions, asks each one to explore a fix, and receives a collection of branches, summaries, partial patches, and confident explanations. Some of the work is useful. Some of it is redundant. Some of it solves the wrong problem elegantly. Without a control plane, the team has to reconstruct the intent after the fact. That reconstruction is expensive, and it often happens during review, when attention is already scarce.
A task board changes the economics. It makes scope explicit before execution. It gives each agent run a reason to exist. It also creates a clean place to store discoveries that are outside the current task. If an agent notices an opportunity to refactor a shared module, that should become a separate candidate task, not an unplanned rewrite inside a bug fix. If a reviewer notices that a patch depends on a missing invariant, that invariant should become acceptance criteria or a follow-up issue. This keeps the agent from turning every local observation into global action. It also keeps humans from being forced to decide a dozen architectural questions inside one overgrown pull request.
This is where human orchestration is more than “review the output.” Orchestration means designing the lanes in which agents move. It means deciding which work is suitable for autonomous exploration and which work requires a design discussion first. It means maintaining a backlog that agents can act on without converting ambiguity into accidental authority. The developer becomes less like a typist and more like a conductor of constraints: define the score, let instruments play their parts, stop the performance when the result drifts, and decide what belongs in the final release.
Scientific software shows why validation cannot be outsourced
OpenAI’s field report on agentic AI in scientific computing reinforces the same lesson. Agents can accelerate maintenance, migration, optimization, and packaging work, especially for teams with limited engineering capacity. But the report also highlights a persistent challenge: validating whether the result is scientifically correct still depends on human judgment. That warning applies directly to ordinary software teams. A coding agent may make the tests green while misunderstanding the domain. It may optimize the wrong metric, preserve a bug because existing tests encode it, or produce a plausible abstraction that violates how customers actually use the product.
That is why the issue tracker should not merely say “fix bug.” It should carry evidence. What behavior is wrong? What user impact matters? Which acceptance test proves the fix? Which constraints must not change? A human-written task is not bureaucracy when an agent executes it; it is the instrument panel. The better the task describes the expected outcome, the easier it is for both agents and reviewers to converge on reality. The worse the task, the more the model fills gaps with probability.
Conclusion
The future of AI-assisted development will not be won by teams that simply run the most agents. It will be won by teams that know where agency begins and ends. Symphony-style orchestration, AI review approvals, and agent-generated pull requests all point toward the same operating model: humans define the work, agents execute bounded pieces of it, automated reviewers remove some mechanical load, and human reviewers retain authority over consequences. The issue tracker becomes a control plane because it is where intent, priority, evidence, and accountability can meet.
That is a quieter story than full autonomy, but it is a better engineering story. Software quality has never come from activity alone. It comes from disciplined decisions made visible over time. AI can increase the amount of code a team can attempt. Human orchestration determines which of those attempts deserve to become part of the system.