Benchmarks are not the same as proof
AI coding agents can write code, but a software team still needs proof that the code was produced under the right constraints, in the right environment, and with the right amount of human oversight. That proof is not a prompt. It is a harness. Anthropic’s explanation of evals for AI agents is useful because it names the thing teams often wave away: an evaluation harness is the infrastructure that runs an eval end to end. It supplies instructions and tools, runs tasks concurrently, records the steps, grades the outputs, and aggregates the results. Once you understand that definition, it becomes hard to pretend that a benchmark score alone tells you very much. A score only matters if the harness matches the kind of work the agent is supposed to do. For software teams, that usually means real repositories, real tool calls, real tests, and real failure modes. If the environment is synthetic or forgiving, the agent can look smarter than it is. If the environment is noisy or under-specified, you cannot tell whether the model improved or the setup changed. The point is not to make agents look good. The point is to make their behavior legible enough that a human can remain in command.
The mistake many teams make is to treat the model as the product and the harness as an implementation detail. That works for toy demos. It breaks down as soon as the agent is asked to do actual engineering work. A coding agent is not just a model that writes text; it is a system that plans, acts, observes, retries, and eventually returns something the team may trust or reject. The trust decision depends on the quality of the surrounding apparatus. If the harness does not preserve the steps, the team cannot replay the session. If the harness does not pin the environment, the team cannot compare runs. If the harness does not encode the task clearly, the agent will optimize for the easiest visible signal instead of the real objective. That is why the harness belongs in the same conversation as the model. The model is the actor. The harness is the stage, the camera, the scorecard, and the rules of the game.
A good harness starts with an explicit oracle
Software work is attractive for agents because parts of it are verifiable. A test suite can say whether a patch passes. A linter can say whether the code violates conventions. A task runner can say whether the command completed. But “verifiable” does not mean “self-evident.” Someone still has to decide what success looks like. That decision is the oracle. The oracle is the human-defined answer that the harness uses to grade the run. If the oracle is weak, the agent can overfit to it. If the oracle is broad but vague, the agent can satisfy the letter of the task while missing the point. A good oracle is narrow enough to be testable and broad enough to cover the thing the team actually cares about. That is why a serious harness usually includes not just unit tests, but also acceptance checks, regression checks, and some kind of transcript review.
Here is the practical shift: do not ask whether the agent “seems right.” Ask whether the task can be written down in a way that another engineer could verify independently. If not, the work is not ready for autonomy. A human reviewer should be able to answer questions such as: what behavior changed, what stayed the same, what assumptions were made, what inputs were trusted, and what failure mode would cause a rollback. If those questions cannot be answered, the agent has not actually completed a task. It has only produced a plausible artifact. Plausibility is not the same as correctness. In production, correctness is the thing that matters, and correctness is only real when the oracle says so.
If the task cannot be tested, it cannot be delegated blindly.
Separate the builder from the judge
Anthropic’s long-running harness work makes a second point that software teams should take seriously: the model that generates work and the model, tool, or process that judges it should not be the same thing in the same moment. Generator and evaluator play different roles. The generator is good at producing options, filling in scaffolding, and making forward progress. The evaluator is good at skepticism, boundary checking, and catching the sort of bug that looks fine when you are the one who wrote it. That separation is valuable because agents can be surprisingly lenient with their own output. They convince themselves that a weak test is enough, that a risky shortcut is acceptable, or that a rough edge is not really a bug. Humans do the same thing, but a dedicated evaluator can be tuned to resist that pressure better than the author can.
This is not theory. Teams building long-running agent workflows keep discovering that progress is fragile when one session is asked to do everything. The better pattern is to decompose the work into tractable chunks and to hand off structured artifacts between sessions. That may sound slower, but it usually is not. It reduces rework, makes debugging easier, and prevents the agent from wandering into a half-complete state that looks productive but is hard to recover. For software teams, the equivalent is to separate implementation from verification. Let the agent draft the code, but make a checker validate the important properties: does it pass the tests, does it violate policy, does it expand permissions, does it change the network boundary, does it preserve rollback? The human then reviews the actual risk, not the illusion of progress.
GitHub’s recent guidance on reviewing agent pull requests points in the same direction. More agent-generated PRs are landing faster than reviewer capacity can absorb, which means teams need to get stricter about what deserves attention. The reviewer should not spend energy admiring the patch surface. The reviewer should spend energy on the places where an agent is most likely to hide risk: CI gaming, duplicate helpers, permission creep, bad scope boundaries, and incomplete rollback paths. That is only possible if the harness has already done the first-pass sorting. Automation scans first. Humans judge the critical path.
Long-running work needs a handoff artifact
One of the most important lessons in long-running agentic coding is that context is a resource. A task that spans hours or days cannot rely on one giant prompt and one uninterrupted session. It needs a handoff artifact: a clean, structured summary of what happened, what failed, what remains, and what a successor should do next. Anthropic’s harness articles emphasize this because the agent does not have memory in the human sense. Each new context window risks starting over unless the workflow leaves behind something concrete. In engineering terms, that means progress notes, checkpoints, test results, and a repository state that can be resumed without guesswork.
This matters even for smaller tasks than a multi-day build. Without a handoff artifact, the next human or agent has to reconstruct the previous reasoning from fragments: diffs, logs, and maybe a chat transcript. That is slow and error-prone. With a handoff artifact, the team can resume from a known state. The artifact should say what was attempted, why certain paths were rejected, what the current failure mode is, and what evidence supports the present plan. The human remains in command because the human can inspect the work without re-deriving it from scratch. Delegation becomes durable when the system preserves enough state to make handoffs cheap.
Infrastructure noise can fool you
Another lesson from Anthropic’s eval work is that small score differences deserve skepticism unless the environment is controlled carefully. Their infrastructure-noise analysis shows that resource configuration alone can move agentic coding results by enough to matter. That is a big deal because teams often read a benchmark delta as if it were pure model improvement. Sometimes it is not. Sometimes it is just memory headroom, CPU allocation, sandbox policy, or a lucky run with fewer infrastructure failures. If you change the harness, you may change the score without changing the underlying capability. If you are not measuring the scaffold, you are not really measuring the agent.
That should change how teams interpret local evals and vendor claims alike. A better score is interesting, but only if you can explain the conditions under which it was earned. What resources were available? What was pinned? What was allowed to fail and retry? What counted as an infra error versus an agent error? If those answers are unclear, then the score is too slippery to drive policy. The human-in-the-loop principle applies here as well: the human should not be forced to trust a number that the team cannot defend. The team should be able to explain not only the result, but the measurement system that produced it.
Human command is the last checkpoint
None of this is anti-agent. It is pro-accountability. Good harnesses make agents more useful by giving them a lane to run in. They also make it possible for a human to say yes, no, or not yet with reasons that are visible later. That is the real value of human oversight in AI-assisted development. The human does not need to inspect every line of output. The human needs to own the boundaries, the acceptance criteria, and the rollback decision. If the task touches secrets, production, permissions, or externally visible behavior, the human must be able to interrupt. If the task is ambiguous, the human must be able to narrow it before the agent optimizes for the wrong target. If the task is complete, the human must still review the evidence, because the evidence is what turns a plausible fix into a shipped change.
That is why the right question is never “Can the model do this?” It is “Can we verify this well enough to delegate it?” If the answer is no, the team should not blame the model. The team should strengthen the harness. Write the oracle. Capture the transcript. Separate the builder from the judge. Pin the environment. Demand a clean handoff. Then keep a human in command at the checkpoint where judgment matters most. Autonomy is not granted by enthusiasm. It is earned by verification.
Practical checklist
1. Define the oracle before the agent starts.
2. Freeze the environment and record resource limits.
3. Log every tool call, approval, and rollback decision.
4. Separate the agent that builds from the process that judges.
5. Require a human for secrets, production, permissions, and ambiguity.
6. Preserve a handoff artifact so the next session can resume safely.
7. Review the process, not only the final diff.That list is intentionally plain. The dull parts are the parts that keep the workflow honest. A harness that is easy to explain is usually easier to trust. A workflow that is easy to replay is usually easier to improve. And a system that makes the human’s role explicit is far more likely to scale safely than one that pretends the model has replaced judgment. The best software teams will not be the ones that give agents the most freedom. They will be the ones that give agents enough freedom to be productive, while keeping the human firmly in charge of the rules of the road.