← Back to news
Human Approval Is the Control Plane

Photo: Lbronn / Wikimedia Commons

30/08/2026

Human Approval Is the Control Plane

The best agentic workflow is boring on purpose

The most useful AI-assisted development workflows are not the flashiest ones. They are the ones that are dull, repeatable, and easy to explain to another engineer after a long week. The one-prompt demo is still seductive: ask for a feature, get a polished answer, and feel as if the whole software lifecycle has been compressed into one elegant exchange. But software does not ship as an exchange. It ships as a series of decisions: what the task is, what the agent is allowed to touch, which checks must pass, who can approve the change, and what happens when the change is wrong. The hard part is not getting a model to produce output. The hard part is building a system that keeps that output inside a human-controlled boundary.

That is why the strongest recent guidance from major tool builders keeps circling back to the same idea in different words. The most important work shifts toward the environments, feedback loops, and guardrails that humans design for agents to execute inside. The developer role does not disappear; it moves closer to the control plane. The more capable the agent becomes, the more valuable it is to make the workflow legible, deterministic where possible, and reviewable where it matters. Human in the loop is not a fallback. It is the operating model.

The practical implication is uncomfortable for anyone hoping that autonomy will erase coordination. It usually does the opposite. Once an agent can write code, open pull requests, run tests, edit configuration, or invoke tools in other systems, the team has introduced a delegation boundary. Delegation is powerful, but it is never free. It needs scope, permissions, verification, and a clear answer to the question, “Who is accountable if this goes wrong?” If that answer is “the model,” the system has already failed at design.

Autonomy is useful only when the boundary is visible.

Build a control plane, not a shortcut

When teams say they want an “agentic workflow,” they often imagine speed as the main benefit. Speed matters, but only because it buys time for better judgment. A good control plane is not a single tool; it is a chain of small constraints that make the work safer and easier to reason about. The workflow starts before the agent writes a line of code. The task should be specific enough that the agent can finish it without inventing product requirements, and narrow enough that a human can tell what success looks like. If a one-paragraph brief cannot name the target, the risk, and the acceptance criteria, then the task is too wide for delegated execution.

Scope before code

Scope is the first form of control. Define the change in plain language before the model touches the repository. Name the file area, the goal, the boundary, and the “do not do” list. In practice this means that an engineer should be able to say, for example, “Update the billing validation path, do not change pricing logic, and do not touch the customer-facing copy.” That extra sentence prevents the agent from solving the wrong problem elegantly. Good agents are very good at widening the solution space. Humans must narrow it.

Permissions before action

Access is the second form of control. A coding agent with read-only access is a very different system from one that can write to the repo, trigger deployments, or touch secrets. The temptation is to grant broad permissions because it is simpler. But simplicity at setup time often becomes complexity at incident time. A safe workflow gives the agent only the scopes it needs for the task at hand. If the task is to inspect logs, the agent should not be able to rotate credentials. If the task is to draft a patch, the agent should not be able to bypass branch protections. The more the agent can do, the more explicit the authorization must be.

Verification after code

Verification is where the system earns trust. Agents are good at making something look plausible; tests, linting, and security checks are what tell you whether it is actually safe. The work of building the surrounding environment matters just as much as the code itself: the UI, the logs, the metrics, and the docs are all part of the feedback loop. If your tests are weak, your agent will learn the wrong lesson. If your observability is poor, the agent will optimize the wrong signal. Deterministic checks are not bureaucracy. They are the difference between a confident guess and a validated change.

Human approval at the boundary

The merge boundary is where human judgment must remain explicit. A pull request is useful precisely because it creates a readable object: a diff, a test result, a history, and a place for accountability. The review should not ask, “Does the agent sound confident?” It should ask, “What changed, what did not, what assumptions were made, and what is the rollback path?” The point of event-driven automation is not to remove review; it is to move repetitive labor out of the way so reviewers can spend time on the meaningful parts. If the change touches authentication, billing, data access, deployment, or secrets, the answer should be human approval, not a faster model.

The same principle applies after deployment. If an agent is allowed to propose a patch, there still needs to be a person who can explain why the patch is acceptable in the production context. “The tests passed” is good evidence, but it is not ownership. Ownership means knowing which users are affected, which invariants matter, and what will happen if the change has to be reverted at 4 a.m. A control plane without a rollback path is just a confidence machine.

The three failure modes humans are there to prevent

The first failure mode is confidently wrong scope. The agent solves a related problem instead of the actual one, and because the output is polished, the mistake is hard to spot. This happens when the brief is vague or when the task is framed as an outcome instead of a bounded change. A model is happy to produce a more ambitious answer than the team asked for. Human review exists to stop “better than requested” from becoming “different from requested.”

The second failure mode is safe-looking side effects. A patch can improve one code path while quietly changing error handling, logging, environment variables, or configuration defaults elsewhere in the stack. The diff may be small; the blast radius may not be. This is especially dangerous in systems where the agent can touch multiple tools or repositories. Less human oversight creates more room for misread intent and unintended consequences. The answer is not to ban agents. The answer is to make the boundaries and approvals visible enough that side effects cannot hide.

The third failure mode is ownership drift. Once a team becomes comfortable delegating, it is easy to start delegating the review of the delegation, then the review of the review, and finally the responsibility itself. The workflow still produces code, but nobody feels fully responsible for whether the code matches reality. Watching internal coding agents is a useful reminder that even inside a company, even on internal workflows, autonomy still needs monitoring and human judgment. Trust is not a vibe. It is a maintained system.

What a practical operating model looks like

The best teams do not ask whether agents should be used at all. They ask which parts of the software lifecycle are safe to delegate and which parts should stay human-led. A practical model usually looks like this:

  • One task, one branch, one owner. Avoid large multi-goal agent runs that mix refactoring, feature work, and cleanup.
  • PRs only. No direct writes to production systems, no silent changes behind the reviewer’s back.
  • Deterministic checks first. Tests, linting, build verification, and security scanning should run before human review.
  • Review the diff, not the summary. Summaries are useful, but the diff is where the truth lives.
  • Escalate risky areas. Authentication, billing, data deletion, secrets, and infrastructure changes deserve a human final approver.
  • Keep a rollback path. If the change cannot be reverted quickly, it is too risky to treat as routine.

This model is not anti-agent. It is pro-accountability. In fact, the most successful agentic teams tend to become more disciplined, not less. They spend less time typing and more time designing the scaffolding that makes the output reliable. That sounds like overhead until you compare it with the cost of debugging a system that changed too much, too quickly, and without a clear owner.

The important word here is orchestration. Orchestration is not “let the agent do everything.” Orchestration is deciding which triggers are safe, what context the agent receives, which checks are mandatory, and where human judgment must remain in the loop. The more mature the workflow, the more visible the handoff points become. The agent does not replace the developer. It forces the developer to become more explicit about the system the agent is operating inside.

The real test is whether you can explain the change

A team has a good control plane when a senior engineer can explain the agent’s change in one minute without hand-waving. They should be able to say what the task was, why the agent was appropriate, what the critical checks were, and where the human decision happened. If they cannot do that, the workflow is too magical. And magic is exactly what production software does not need.

This is why “human in the loop” is not a slogan for caution; it is a design principle for responsibility. Humans do not stay in the loop because agents are weak. Humans stay in the loop because software systems have side effects, incentives drift, bugs hide in integration points, and product decisions are not the same thing as code completion. The model can move quickly through the mechanical parts. The human must still decide what counts as success, what risk is acceptable, and what change deserves the authority to reach users.

That is the real meaning of the control plane. It is not a wall against automation. It is the set of human-defined boundaries that lets automation scale without becoming blind. The best agentic workflows are the ones where the machine does more of the execution and the human does more of the judgment. That division is not a compromise. It is the reason the system works.

Sources