The models are better. The benchmarks are higher. Dex Horthy has fresh evidence that unattended coding agents still turn healthy codebases into radioactive spaghetti. And he wants you to hear about it.
The best thing about interviewing Dex Horthy is that Dex has no editor or filter. Not only does this make talking to him highly entertaining, it also produces the kind of spicy takes that reliably set off disputes across social media. He pulls it off with a disarming smile and an aw-shucks demeanor that belies the fact that he is one of the deepest thinkers working in the strange and still poorly understood borderland between human judgment and the agentic runtime. One moment you are discussing reinforcement-learning verifiers. The next, Dex is drawing quality-decay curves, explaining the future of GitHub and confessing that HumanLayer had to abandon a 30,000-to-40,000-line codebase after its agents turned it into an architectural fever dream. Then he reaches for more coffee to “peak the adrenaline,” produces another diagram and announces, “You’re the first person to see this,” he confides. Yikes but…yeah!!
In an era when coding agents are being embraced as software salvation with a fervor approaching religion, Dex remains the cheerful heretic. Ask whether fleets of even the most advanced agents can maintain software and he will tell you they write slop, the slop breeds more slop and a meaningful percentage of the industry is either fooling itself or losing its collective taste. “There’s a lot of hype out there,” he says. “A lot of it is bullshit.”
Let’s caveat the claim before the knives come out. Dex thinks coding agents have become disturbingly – no, disarmingly good – at writing code. They can debug, reverse-engineer, use tools, write tests and ship a feature before a human engineer has decided whether to refill the coffee, and do an excellent job with these defined tasks. .But set them loose across a long succession of changes and a messy, morphing software journey and they still tend to degrade the codebase beneath them to the point of destruction
Come see Dex talk at AGNTCon + MCPCon NA.
The SlopCodeBench Test: Where Software Factories Hit the Wall
Until recently, this case rested heavily on experience and what Dex calls “AI intuition,” the instinct power users develop after watching models make the same oddly predictable mistakes. Now he has data from SlopCodeBench, a benchmark cruel enough to resemble actual software development. Most coding benchmarks disclose the whole problem at the beginning or ask a model to repair a bounded issue inside an existing repository.
Real software does not emerge that way, perfectly baked from unblemished PRDs and architecture diagrams. Customers request features nobody anticipated, security teams reject architectural choices, product leaders change direction and new requirements collide with old assumptions. SlopCodeBench recreates that mess by giving an agent an initial feature and then revealing more requirements across successive checkpoints. The agent never sees the entire challenge in advance. The crux is that at every turn, it must inherit and extend the code it wrote earlier. “That is exactly how you and I build projects,” Dex says.
A diehard benchmark critic, Dex finally found one he liked and put it to the test, with illuminating results. Across six challenges and 30 checkpoints, Fable and Sol each recorded 10 strict passes, or 33.3%. Two Kimi K3 runs, using Modal and Baseten, reached eight and seven. Dex is careful about the caveat: this was a small subset, the provider differences sat within the error bars and he does not present the test as statistically conclusive or as an endorsement of one provider over another. But every model accumulated defects as the challenges continued, which is precisely the pattern he has been warning about.
HumanLayer has already run a real version of that experiment on itself. For three or four months, the team operated a lightly supervised software factory and deliberately allowed questionable code to remain to see whether the agents could recover. The system metastasized into 30,000–40,000 lines, including an “insanely complicated state machine” with four Unix processes communicating through sockets, standard I/O and a peculiar event bus. HumanLayer eventually archived the entire codebase. “If you let the agents take the wheel,” Dex says, “this is what happens.”
HumanLayer Is the Argument Made into a Company
Dex has more at stake here than a provocative conference thesis. He is the founder and CEO of HumanLayer, a San Francisco developer-tools company from Y Combinator’s Fall 2024 batch.
The company initially built infrastructure that allowed agents to request feedback and approval from humans through Slack, email and other channels. HumanLayer has since evolved into an IDE and collaborative cloud platform for using coding agents on large, difficult codebases. This is an important point because Dex is not arguing against coding agents, nor is HumanLayer a digital artisanal society dedicated to writing everything by hand. The company is trying to help engineering teams move faster with agents while retaining ownership of architecture, code quality and the decisions that shape the system. Yes, he is pitching his own book but also putting his money where his mouth is.
Dex made an earlier version of this argument in his August 2025 MLOps Community talk, “12-Factor Agents: Patterns of Reliable LLM Applications,” now in the Agentic AI Foundation archive. At the time, he argued that the full-fat agent loop—give a model a prompt, a sack of tools and permission to keep going—was the wrong abstraction for serious production work. “It turns out the bigger that loop gets, it doesn’t really work,” he said. Since then, “don’t read the code” has become a management philosophy.
This, of course, begs the obvious question. Now that agents are vastly more capable in the post-Opus 4.5 Era, has Dex changed his mind and joined the loop/graph coding party? To the contrary.
The core problem remains. Software engineering has good machinery for catching some failures and almost nothing reliable for catching “taste failures” that are the root cause of code slop. Slop compounds because agents read the repository and imitate its patterns. One poor decision becomes precedent, and precedent becomes architecture. “Each slop pattern in your codebase increases the chance that more code is also going to be sloppy,” Dex says. That leaves the software factory with an awkward wetware bottleneck: AI can generate changes in minutes, but humans may still need hours to understand them. Dex thinks teams should accept and embrace this constraint rather than pretend it has vanished, then move as fast as possible inside it. Which leads us to the real bottleneck of verification (a subset of taste). So rather than loops all the way down, its verifiers all the way down.
Come for the Talk. Stay to Talk to Dex.
Dex Horthy will present “There Is No Software Factory Without Better Verifiers” at AGNTCon + MCPCon North America in San Jose, California on October 22, 2026. The talk will connect the new benchmark evidence, HumanLayer’s failed factory, the verifier problem and his emerging vision for the software forge. Expect a deck approaching 159 slides, several diagrams and at least one cherished assumption destroyed. Expect a fresh Dex take as he mods each talk roughly 30% each time. Dex will be around after his talk to disabuse any remaining notions you might have.
Share
Author

Alex Salkever
View All PostsAlex Salkever is the Editor-in-Chief of the AAIF and the Linux Foundation. He has been working in open source storytelling for over a decade and formerly served as a CMO and VP at a number of technology companies. He started his career in journalism, ultimately working as the technology editor for Bloomberg BusinessWeek. He leads efforts to build the AAIF media engine to drive awareness and education.



