Coding agents have become very good at producing plausible code quickly. This is excellent news right up until the moment someone has to decide whether the code is correct, safe, compliant and worth maintaining. Generation now scales faster than verification. The bottleneck did not disappear. It moved.
Where verification falls behind
A Place for Mom offers one of the clearest pictures of the new shape. Its Grace platform was built agentically from the first commit and has been in production since November 2025. Over ten months, the team produced 3,478 pull requests across four repositories and merged 2,822 of them—an 81% merge rate. Persistent context, structured planning, isolated Git worktrees and a 10-stage multi-agent review pipeline helped make that pace possible. Principal engineer Austin Brown gives the remaining problem a useful name: comprehension debt. An engineer can now ship in two hours what once took two days, but may not understand it well enough to debug at 2 a.m.
Datadog is attacking the review side at a different scale. Nearly 10,000 pull requests move through the company each week. Human reviewers cannot catch every missing timeout, cold-cache stampede or behavioral change that could turn into an outage. Its agentic reviewer is deliberately narrow: find reliability risks, verify the concern and explain the exact line likely to wake someone up in the middle of the night. Style nits can wait.
The broader agenda points to a new engineering division of labor. HumanLayer argues that there is no “dark factory” for software until agents have much better verifiers. Amazon’s audit of more than 6,000 failed agent trajectories shows why: endless loops and timeouts alone account for nearly a third of failures, while the harness can either help an agent recover or drive it deeper into trouble. Endor Labs calls the accumulating downstream risk “generation-verification asymmetry.” One agent introduces a dependency, another builds on it and a third deploys it before a human understands what entered the system.
Build proof into the pipeline
The response is not one giant reviewer agent standing at the end of the conveyor belt. Policy belongs before generation. Deterministic checks should sit beside probabilistic judgment. Agents should verify claims with tools before surfacing them. Tests and evaluators need to examine trajectories and consequences, not merely the final answer. FINRA’s move from assistants to bounded SDLC agents shows how these pieces become an enterprise capability instead of a collection of impressive personal workflows.
The age of expensive code encouraged teams to optimize production. The age of abundant code will force them to optimize proof. The best engineering organizations will not be the ones whose agents write the most. They will be the ones that can explain why the code deserves to ship.
Continue the conversation at AGNTCon + MCPCon North America
- The Agentic SDLC: How We Shipped AI-Native Software at a Legacy Company with Austin Brown, A Place for Mom
- Don’t Merge That: Catching Outages at 10,000 PRs a Week with Joris Bonnefoy, Datadog
- Why Your Agent Is Failing: Failure Modes from 6,000+ Agent Trajectories with Han Xu, Amazon
- There’s No Dark Factory Without Better Software Verifiers with Dexter Horthy, HumanLayer
- Generation-Verification Asymmetry with Birajendu Sahu, Endor Labs
- Before the Agent Writes Code with Nnenna Ndukwe, Qodo
- From AI Assistants to Trusted SDLC Agents with Geetha Ramachandran, FINRA
AGNTCon + MCPCon North America takes place Oct. 22–23 in San Jose. Register and use OUTREACH25 to save 25%.
Share
Author

Alex Salkever
View All PostsAlex Salkever is the Editor-in-Chief of the AAIF and the Linux Foundation. He has been working in open source storytelling for over a decade and formerly served as a CMO and VP at a number of technology companies. He started his career in journalism, ultimately working as the technology editor for Bloomberg BusinessWeek. He leads efforts to build the AAIF media engine to drive awareness and education.



