Agent Autonomy Depends on Verification
Executive Summary
Giving coding agents more authority makes the quality of verification—and the organization’s ability to act on it—more consequential. Two practitioner discussions put substance behind that point: Braintrust describes evaluation as a continuous loop between tests and production failures, while a Sonar/McKinsey panel describes moving beyond supervised pilots through workflow redesign and staged permissions.
The shared argument is that generating changes is only part of the work. Determining whether those changes deserve to proceed requires evidence, domain judgment, and repeatable review processes. These are vendor and conference perspectives, not independent demonstrations of productivity gains, but they offer a concrete account of what expanded delegation demands.
Evaluation Is a Production Loop
In “Why Building an Eval Platform Is Harder Than It Looks”, published October 6, Braintrust’s presentation connects pre-release test cases to production traces, failure analysis, and regression checks. Evaluation is presented less as a score obtained before launch than as a continuing feedback system.
The practical demands explain why this matters. Teams need to construct datasets, score or verify outputs, debug failures, collaborate with domain experts, and query large, semi-structured traces. Privacy controls and monitoring for drift add requirements that a spreadsheet alone does not address.
The consequential distinction is between recording a failure and making it useful. A trace can reveal what happened; turning that observation into a regression case gives the team a way to check whether a proposed change fixes the problem without recreating it later. That is the value of the loop described here, rather than simply accumulating more telemetry.
The presenter also describes coding agents operating on traces and proposing changes, with people reviewing the outcome. This places agents inside the improvement process without making them the final authority on whether an improvement succeeded. The talk explains a platform’s intended workflow; it does not independently establish the product’s effectiveness.
Permissions Follow Workflow Readiness
The Sonar/McKinsey panel, “Redesigning How Software Gets Built With AI Agents”, provides the organizational counterpart. Delayed discovery: the video was published October 7. Participants describe most organizations as working through supervised agent pilots, with few reaching end-to-end automation. That characterization is their account, not a measured industry-wide census.
They attribute differences in scaling to workflow redesign, tools and context, and human or organizational factors. Sonar’s CEO emphasizes measuring value rather than token consumption, establishing explicit verification processes, and transferring practices beyond the teams that first piloted them.
The panel’s permission progression is particularly useful: review, blocking, autofix, and eventually autonomous merge for selected changes. These are distinct levels of authority, not interchangeable examples of “using agents.” An agent that can flag a problem has a different responsibility from one allowed to alter code or merge it.
Read alongside Braintrust’s evaluation loop, the implication is straightforward: broader permissions need a stronger basis for deciding which actions are acceptable. A successful pilot alone does not explain how another team will reproduce its verification practices.
The Bigger Story
Recent discussion of persistent background agents emphasized making continued work visible and controllable. These talks reinforce that view while adding a missing layer: visibility must connect to a process for judging outcomes. Knowing that an agent acted is not the same as knowing that its action was sound.
Neither discussion demonstrates a new model capability or proves that autonomous software delivery is broadly ready. Their contribution is a more precise account of deployment readiness: useful traces, meaningful checks, human expertise, and permissions matched to the work being attempted.
The distinction worth carrying forward is activity versus justified authority. Token use measures activity; permission to merge requires confidence in a particular class of changes. Evaluation and organizational practice are how that confidence might be earned—not evidence that it already has been.
