New research from Undo finds that almost four in five (79%) engineering leaders say their release cycles are no faster than before, despite their teams being able to produce code more easily than at any time in their careers.
As AI agents have increased the volume of code they can create, engineers now spend nearly twice as long debugging it as they do writing it, averaging 16.9 hours a week.
That accounts for 42% of the average working week. Engineers are simply unable to keep up with their agents, leading to more than a third (35%) of AI-generated code reaching production before they’ve fully comprehended it.
Adding to the risk, AI agents frequently hallucinate the cause of failures, or fail to identify problems in the codebase entirely.
In the past six months, as a result of their use of AI coding tools, 81% of organisations have had a production incident or serivce outage affecting internal users or customers.
In the same period, 93% have had the root cause of an issue incorrectly diagnosed because of an AI hallucination.
Further, 91% have had test escapes, serious defects or poorly optimised code enter production.
“When code is obviously broken, the cause is usually easy to find,” said Greg Law, founder and CEO of Undo.
“Where engineers struggle is with code that’s almost, but not quite right. Those are the times they lose days trying to unravel what went wrong and why. Their challenge is that while agents are great at writing reams of code quickly, they’re less capable at debugging it.
“The result is engineers are being buried in an avalanche of code that’s well beyond human capacity to debug. That’s why we have to give them a way to make AI better at debugging, by feeding agents with the rich context of what code actually does at runtime.”
Recommended reading
- ‘Reasoning’ AI Emits Up to 50x More Carbon Than Simpler Models
- AI Benefits Concentrated Among Minority of Workers, PwC Finds
- Is ChatGPT Making You Dumb? Possibly, MIT Says
- New Apple Paper Pours Cold Water on AI Reasoning
- Is AI Making Us Dumber? Workers Think So.
Four in five (80%) engineering leaders say coding agents struggle to solve difficult problems in large-scale, complex codebases. The arrival of more powerful models doesn’t offer a realistic solution, with a strong degree of cynicism about the impact the planned IPOs of Anthropic and OpenAI will have on AI affordability.
The majority (82%) of engineering leaders think the costs of coding agents will go ‘through the roof’ as the AI labs prioritize making Wall Street happy.
However, engineering leaders widely agree that improving model context is more important than increasing their capability to make AI more powerful. More than four in five (82%) say AI agents would be far more useful for code comprehension and debugging if they were grounded in the context of what happened during runtime.
“AI doesn’t have an intelligence problem; it has an evidence problem,” continued Law.
“An agent asked to explain why a program behaved the way it did without ever being shown what actually happened will fill in the gaps with confident guesses. That’s how days are lost with humans and agents going down blind alleys in a futile search for the root cause of critical issues that must be fixed.
“Give an agent a complete recording of what happened at runtime and it can see the precise sequence of steps leading to what went wrong, rather than have it trying to guess. The right context even makes a mid-tier model more capable of providing an accurate diagnosis than a top-end model working from source code or logs alone. The result is AI that is smarter, faster and cheaper.”





