← Back to news
Coding agents need evaluation-driven delegation

Photo: Tsinkala / Wikimedia Commons (CC BY-SA 4.0)

27/09/2026

Coding agents need evaluation-driven delegation

Agent development is becoming evaluation-driven

A new practitioner study on software engineering agents captures a shift that many teams already feel: when AI agents make implementation cheaper, the hard part of software work does not disappear. It moves upstream into requirements, evaluation, review, deployment discipline, and human judgment. That is a useful corrective to the idea that coding agents are mainly about replacing keystrokes. The more capable the agent becomes, the more important it is to decide what the agent is allowed to do, how success is measured, and who accepts the result.

The paper, How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study, is especially relevant because it studies people building software engineering agents in practice. The authors interviewed 20 practitioners from 12 organizations and surveyed 80 more. Their conclusion is not that agentic development eliminates the software process. It reorganizes it. They describe a recurring loop of requirements, evaluation, data, system construction, testing and deployment, human feedback, and adaptive maintenance. In that loop, evaluation is not a final inspection after the code is written. It increasingly becomes the steering mechanism for the whole effort.

For engineering leaders, this matters because most AI-assisted development conversations still start in the wrong place. They ask which model writes the most code, which IDE integration feels fastest, or which agent can handle the longest task. Those are valid operational questions, but they are secondary. A team that cannot define acceptance criteria, cannot capture domain constraints, cannot test behavior reliably, and cannot review generated changes will not become safer just because the agent gets faster. It will simply create more unreviewed surface area.

Implementation gets cheaper, but accountability does not

The most useful sentence in the study is the idea that bottlenecks shift rather than disappear. When an agent can generate files, run commands, repair failures, and repeat the loop many times, implementation no longer consumes the same share of attention. That sounds like an unqualified gain, and in many situations it is. Boilerplate, test scaffolding, codebase exploration, migrations, documentation updates, and routine refactors can move faster. But every acceleration creates a matching control problem.

If a human writes the code by hand, comprehension tends to arrive during construction. The developer has seen each trade-off, each shortcut, and each awkward edge case. With an agent, construction can outpace understanding. The resulting code may compile, pass a narrow test, and look stylistically consistent while still hiding a wrong assumption. The study names this kind of pressure as comprehension debt: software moves into the system faster than the team can understand and evaluate it.

That debt is not just a philosophical concern. It changes the shape of code review. Reviewers no longer ask only whether a colleague made a good implementation choice. They must also ask whether the agent understood the problem, whether the prompt omitted a constraint, whether tests were selected because they were meaningful or merely because they were available, and whether the diff is larger than the human reviewer can responsibly absorb. The human remains accountable for the outcome, even when the first draft came from a model.

This is why the human-in-the-loop principle should not mean a passive approval click at the end of an agent run. It should mean active ownership at several points: framing the requirement, deciding the boundaries of delegation, reviewing the plan, defining evidence, inspecting the diff, interpreting the test results, and making the release decision. The agent can do a great deal of work inside that frame. It should not own the frame itself.

Evaluation becomes a first-class artifact

The study’s strongest practical recommendation is the move toward evaluation-driven development. In classic test-driven development, tests often express expected behavior before implementation. In agentic software engineering, evaluation needs to be broader. It may include unit tests, integration tests, golden datasets, scenario suites, tool-use traces, regression prompts, security checks, latency budgets, cost budgets, and human review rubrics. The important point is that evaluation is defined early enough to guide iteration, not late enough to provide a comforting stamp.

For a team adopting coding agents, that means every delegated task should start with the question: what evidence would convince us this change is safe? For a bug fix, the answer might be a failing regression test that now passes. For a dependency upgrade, it might include changelog review, compatibility checks, lockfile inspection, and a rollback note. For an agent that triages incidents, it might include historical incident replay, false-positive tolerance, escalation rules, and explicit human override. The evaluation package becomes part of the work product.

This changes how teams should write prompts and repository instructions. A weak instruction says: fix the bug. A stronger instruction says: reproduce the bug, add the smallest regression test, propose a plan, change only the affected module, run these checks, and summarize residual risk. The second version does not merely ask for code. It defines a chain of evidence. That is the difference between using an agent as a stochastic typist and using it as a delegated engineering worker.

It also explains why specifications, prompts, skills, context definitions, and harness behavior are becoming first-class artifacts. If an agent’s behavior depends on those materials, they deserve the same discipline as code. They should be versioned, reviewed, tested, and owned. A repository-level instruction file that says how to build, test, lint, migrate, and review the project is not bureaucracy. It is operational infrastructure for safe delegation.

The test oracle problem gets sharper

Evaluation-driven development is powerful, but the study also highlights a trap: unreliable evaluation signals. In conventional software, teams already know that tests can be incomplete or wrong. In agentic workflows, the risk grows because agents optimize toward the visible evaluation. If the benchmark is shallow, the agent may learn to satisfy the benchmark rather than the intent. If the test suite misses important behavior, the agent may produce code that looks correct under the available oracle and fails in production conditions.

That makes human judgment more important, not less. The evaluation suite should be treated as a decision aid, not a replacement decision-maker. Senior engineers need to ask whether the checks cover the right risks, whether they reflect current business rules, and whether the generated solution is maintainable beyond the scenario that triggered it. A green test run is useful evidence. It is not a complete argument.

One practical pattern is to separate three layers of acceptance. First, mechanical verification: tests, linters, type checks, static analysis, dependency scans, and reproducible commands. Second, behavioral validation: scenarios, edge cases, user journeys, backward compatibility, observability, and operational impact. Third, human review: design coherence, readability, ownership boundaries, security implications, and release readiness. Agents can assist all three layers, but humans should decide whether the layers are sufficient.

Another pattern is to ask the agent to challenge its own evaluation. Before merging, require it to list cases not covered by the tests, assumptions it made, files it deliberately avoided, and reasons the change could still be wrong. This will not make the model perfectly honest or complete, but it shifts the workflow away from blind confidence and toward inspectable risk.

Model updates mean the ground can move

The study also describes a problem every platform team will recognize: provider-side model updates can change behavior even when the team changes none of its own code, prompts, tools, or data. In traditional software, if nothing in your repository changed, you expect repeatability. With external models, the execution substrate can shift underneath you. Planning style, tool usage, sensitivity to context, refusal behavior, and code preferences can all change.

This is another reason evaluation cannot be an afterthought. Teams that depend on agents need regression suites for agent behavior, not just application behavior. If a model update changes how an agent handles migrations, reads tickets, or edits tests, the team should discover that in a controlled environment before it appears in production work. This does not require blocking all model upgrades. It requires treating them like dependency upgrades: measured, tested, observable, and reversible where possible.

For organizations using multiple tools, the governance question becomes concrete. Who approves model changes? Which repositories can use fully autonomous modes? Which commands require explicit approval? Which MCP servers can read or write external systems? How are prompts and skills reviewed? Where are agent logs stored? How are incidents investigated when a generated change causes a defect? These are software engineering questions, not abstract AI policy questions.

What teams should do now

The immediate response is not to ban agents or to let them run freely. The practical response is to build a controlled operating model. Start with bounded tasks where the evidence is clear: adding regression tests, improving documentation, making small refactors, preparing pull request summaries, or performing low-risk migrations. Require a plan before implementation. Require verification commands in the final summary. Keep diffs small enough for human review. For critical paths, require senior sign-off even when all automated checks pass.

Then invest in the boring assets that make agents safer: repository instructions, known-good commands, evaluation datasets, scenario tests, code ownership maps, release checklists, rollback templates, and review rubrics. These assets help humans as much as agents. They reduce ambiguity, improve onboarding, and turn tacit knowledge into shared context. The study’s point about specifications becoming first-class artifacts is exactly right: the more work we delegate, the more explicit our intent must become.

Teams should also measure the right things. Counting generated lines is almost useless. Better metrics include cycle time for specific task classes, review defects per agent-assisted change, escaped defects, rework rate, test coverage for modified behavior, prompt or instruction changes that reduce failures, and the percentage of agent sessions that produce a small, reviewable diff. The goal is not more AI activity. The goal is safer throughput.

The senior developer takeaway

Software engineering agents are becoming real participants in the development process. That does not remove the need for human control. It raises the standard for it. As implementation becomes easier to delegate, human value moves toward intent, context, judgment, verification, and governance. The teams that benefit most will not be the teams that trust agents the most. They will be the teams that build the clearest loop between human intent, machine execution, measurable evidence, and accountable review.

The practical thesis is simple: let agents accelerate the work, but make humans responsible for the system of work. Define the requirement, define the evaluation, constrain the tools, inspect the change, and own the release. In AI-assisted development, the future is not autonomous coding without humans. It is evaluation-driven engineering with humans firmly in command.

Sources