96% of Developers Don't Trust AI Code. Teams Are Rewriting Review.
A Sonar survey finds 42% of code in shared repositories is now AI-generated, yet 96% of developers don't fully trust it. Engineering teams are rebuilding review around specs, agents, and human accountability.
Updated
Why it matters
- 42% of code added to shared codebases is AI-generated, per a Sonar survey of more than 1,100 developers
- 96% of surveyed developers do not fully trust AI-generated code to work correctly
- AI code-review startup CodeRabbit raised $143M at a $1.5B valuation in August, claiming 2M reviews per week
- Synthesia reported a 120% year-over-year rise in pull requests as of August, with 95% containing AI-generated code
- IBM estimates AI now lets junior engineers perform 70 to 80 percent of some tasks that previously required senior engineers
AI coding assistants now produce an estimated 42 percent of the code landing in shared codebases, yet 96 percent of developers say they do not fully trust that code to work correctly.
Those figures come from a Sonar survey of more than 1,100 developers, published this week as engineering teams across enterprise software, e-commerce, and AI startups grapple with a flood of machine-written pull requests. The shift has moved the bottleneck of software production from writing code to verifying it. "Large language models can produce code that looks clean on the surface but conceals sloppy mistakes," IEEE Spectrum reported, citing faulty assumptions, security vulnerabilities, and subtle errors that emerge only after deployment.
The economic stakes are visible in deal flow. In August, AI code-review startup CodeRabbit raised US $143 million at a $1.5 billion valuation. The company claims it performs more than 2 million reviews a week for 17,000 customers, including Nvidia, Indeed, and BMW Group.
What changed when AI took over the keys?
At Synthesia, an AI video-generation platform with 118 engineers, the all-in moment came in November 2025 when the company adopted Claude Code and similar tools across the engineering org. As of August, the result was a 120 percent year-over-year rise in pull requests, with 95 percent of those requests containing AI-generated code, according to chief technology officer Peter Hill.
That volume forces a reckoning. In the Sonar survey, 38 percent of developers said reviewing AI-generated code required more effort than reviewing human-written code. Sixty-one percent said AI often produces code that looks correct but is "unreliable."
The failure modes look mundane until they hit production. Hill told IEEE Spectrum that AI tools at Synthesia repeatedly fail to recognize that code for a given task already exists. The company has found as many as 10 versions of the same function cluttering its codebase. Engineers triage the duplicates, then retrain the agent to avoid regenerating them.
"I don't know if we ever get to the point where you can truly trust the agentic generation of code," Hill said.
How are teams catching the mistakes?
Three patterns have emerged across the case studies, none of which removes humans from the loop.
The first treats the specification as the critical artifact. McLaren Stanley, a senior principal engineer at Amazon Stores, leads a 70-person team modernizing 17 years of code inside Amazon's mobile shopping app while supporting more than 1,000 downstream developers. Stanley's team now spends more time deciding what the code should do before any model writes it.
The cost of getting specs wrong can be measured in wasted output. Stanley described an incident in which a missing instruction caused an agent to generate 25,000 lines in the wrong version of Swift. The model then produced 600 errors it could not fix in one pass. Stanley discarded the code, updated the spec, and the agent regenerated the correct code 15 minutes later.
The second pattern delegates the first review pass to specialized AI agents. At AWS, senior principal engineer David Yanacek said agents test whether code works, compare it against the original plan, and check for security flaws before any human sees it. Bonterra, a nonprofit software provider with about 290 engineers, runs agents that compare incoming changes against approved designs, security rules, coding standards, and accessibility requirements, then report a confidence score. A low score or flagged issue routes the change to a person. Code touching payments or personal data always gets a human, according to CTO Tanuja Korlepra.
"Agents do the reading and humans do the judging," Korlepra said.
The third pattern shifts the burden back onto the developer. Temporal, the open-source workflow platform, runs a "Send Back" policy: engineers must explain in their own words the agent's design choices and how the code handles unusual conditions, or face rejection. CEO Samar Abbas framed the policy as a refusal to treat code review as automatic rubber-stamping.
"We refuse to let code review become a dumping ground for unchecked model outputs," Abbas said.
Who reviews code when AI writes most of it?
The math is getting harder for human review to keep up. At Bonterra, proposed changes tripled within three months of adopting AI. Code entering review rose tenfold. Review times tripled. It became impractical for engineers to inspect every line, Korlepra said.
Synthesia routes changes to humans based on risk. Editing an error message carries less scrutiny than code that handles customer data or core business rules. Even so, fewer than 5 percent of changes bypass human review, Hill said.
That tension has created a new failure mode that JD Raimondi, chief AI architect at consultancy Making Sense, calls "theater approval." Engineers confirm a feature works, skim the code, and approve it without understanding the choices underneath. Raimondi told IEEE Spectrum this kind of hollow endorsement grows more likely as machine output outpaces human reading capacity.
What happens to junior engineers?
The story shifts again at the bottom of the career ladder. As AI eats the boilerplate, juniors lose the apprenticeship that once built their judgment.
At Making Sense, juniors showed some of the largest productivity gains from AI, raising concerns about what they no longer learn by doing, Raimondi said. The consultancy keeps juniors involved in deciding why a customer needs a feature and how it should work, rather than limiting them to checking AI output.
IBM is using AI to accelerate that ramp. General manager of automation and AI Neel Sundaresan said recent graduates now work on product features and projects once reserved for senior engineers. AI handles implementation and tests; if those fail, juniors diagnose what went wrong and fix it before senior developers sign off. Sundaresan estimated AI now lets junior engineers perform 70 to 80 percent of some tasks that previously required a senior.
Synthesia, which hires mostly at mid-level and above, pairs its less-experienced employees with a senior colleague and an AI agent. They own parts of projects and learn to define what successful code should do.
Bonterra moves juniors from coding well-defined tasks to owning outcomes alongside experienced colleagues. They learn to direct agents, question their output, and remain responsible for the result, Korlepra said.
"If the industry stops hiring juniors, the industry stops producing seniors," Korlepra said.
Can the review gap ever close?
Not soon, if the engineering leaders quoted in IEEE Spectrum's reporting are right. Sonar's data shows 96 percent of developers still don't fully trust AI output to work correctly. Hill said outright that he doesn't know if full trust will arrive. CodeRabbit's $143 million round signals investors see that gap as a market.
What's still open is whether the bottleneck moves a second time. Better AI reviewers may compress the gap. Specs-first workflows may move failure earlier. Send Back-style policies may force a deeper kind of accountability. What seems unlikely, based on these case studies, is a future in which humans routinely sign off on code they don't read, and the field has not yet decided what to call that practice when it becomes standard.
Original: sonarsource.com
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
200 articles
Related articles
- Enterprise AI Coding Agents: IP Indemnity, Data Residency and 500-Seat Costs Compared
- AMD Says AI Agents Now Auto-Fix 75% of Radeon Software Bugs
- OpenAI audit finds ~30% of SWE-Bench Pro coding tasks are broken
- OpenAI's Codex Hits 4 Million Weekly Developers, Launches Enterprise Push
- OpenAI Raises $122 Billion to Scale Frontier AI Worldwide