AI Coding Agents Produce More Code, Not More Software
A Harvard study of 700+ firms and 300 million work events finds AI coding agents boost code output but get stuck at human review, with no gain in shipped software.

Updated
Why it matters
- Harvard researchers Fiona Chen and James Stratton analyzed 300 million work events across 700+ firms and 700,000+ employees, covering 2021 through March 2026.
- The study finds "little evidence that firms increase software output or reduce employment" from AI coding tools.
- Efficiency gains are "absorbed by downstream constraints" — code review lengthens, pull requests need more revisions, and reviewers leave more comments.
- Data came from Jellyfish engineering analytics covering commits, pull requests, and issue management.
A Harvard study of more than 700 software firms finds "little evidence that firms increase software output or reduce employment" after adopting AI coding agents. The efficiency gains these tools deliver during the actual writing of code, the researchers conclude, are "absorbed by downstream constraints in the production process" — chiefly the human code review stage, which the study identifies as a significant "bottleneck" for overall productivity.
The working paper, authored by Harvard University researchers Fiona Chen and James Stratton, draws on aggregated analytics from Jellyfish, a company that measures the granular output of engineering teams. The dataset is unusually large and unusually close to real engineering practice. It covers 300 million individual "work events" — commits, pull requests, and issue-management activity — spanning more than 700,000 employees at over 700 software development firms from 2021 through March of 2026.
The stakes for the finding are considerable. AI coding assistants are among the most widely deployed and heavily capitalized applications of large language models, and vendors routinely pitch them as direct accelerators of software output. Enterprises are making hiring and tooling decisions on that premise. If the productivity gain materializes only as more code — and not as more shipped software — the economics of those investments look materially different from what the marketing suggests.
What does the study actually measure?
Chen and Stratton did not survey developers about how productive they feel. They analyzed operational telemetry: the commit-level and pull-request-level records that Jellyfish aggregates from engineering organizations. That matters, because self-reported productivity in developer surveys has often told a rosier story than behavioral data.
The observation window, 2021 through March 2026, captures the period in which AI coding assistants moved from novelty to default tooling at many firms. That gives the researchers before-and-after visibility around adoption events across hundreds of organizations rather than a single company's anecdote.
The core metric is software output at the firm level — not lines of code, not commits, but delivered work. It is precisely this distinction that drives the paper's conclusion.
Where do the efficiency gains go?
The study's central finding is a displacement effect. AI agents compress the time and labor needed to produce a given piece of code, but that compression pushes work downstream. The researchers report that "the code review process significantly increases in length, pull requests are more likely to require revisions, and reviewers leave more comments."
In other words, the review stage absorbs the savings from the generation stage. Three specific mechanisms appear in the data:
- Code review takes significantly longer on AI-assisted teams
- Pull requests are more likely to require revisions before merging
- Reviewers leave more comments per pull request
The net effect at the organizational level: no measurable increase in software output, and no measurable reduction in employment. The code gets written faster. The software does not ship faster.
Why doesn't more code mean more software?
The result aligns with a pattern that programmers themselves have been reporting for some time. Modern AI coding assistants and agents can generate huge volumes of functional code quickly. But developers using those tools also know better than to trust the accuracy of that output, which means substantial effort goes into reviewing everything an AI produces before it enters a production codebase.
The Harvard study effectively quantifies that intuition across hundreds of firms. Human review is the quality gate that cannot be automated away without accepting the risk of shipping unverified code. Every minute an agent saves at generation time creates demand for reviewer attention at merge time.
This is a classic bottleneck dynamic: speeding up one stage of a production process does not speed up the whole process if another stage constrains throughput. The slowest link — here, human verification — sets the pace.
What does this mean for AI productivity claims?
The finding lands in the middle of an active debate over how much economic value AI coding tools actually create. Prior research has suggested that time saved by AI tools is offset by new work created — much of it review and correction. Developer surveys have likewise shown trust in AI coding tools falling even as usage rises, which is consistent with what Chen and Stratton observe in the behavioral data.
For engineering leaders, the implication is direct. Measuring the success of an AI coding rollout by code-generation volume will overstate its value. The binding constraint on team throughput is the review pipeline, and that is where investment — in review tooling, in reviewer capacity, in process redesign — would need to go before generation-side gains can convert into shipped software.
For the broader AI market, the study is a caution against extrapolating micro-level speedups into macro-level output gains. Firms adopting these tools at scale are, on this evidence, producing more code with roughly the same number of people and roughly the same delivery rate as before.
The open question the paper sets up is whether review itself becomes the next automation target — and whether AI-assisted review can clear the trust bar that AI-generated code currently fails. Until it does, the study suggests the productivity story of AI coding agents will remain a story about shifting work, not eliminating it.
Original: fion.ac
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
230 articles
Related articles
- 96% of Developers Don't Trust AI Code. Teams Are Rewriting Review.
- AMD Says AI Agents Now Auto-Fix 75% of Radeon Software Bugs
- OpenAI's Codex Hits 4 Million Weekly Developers, Launches Enterprise Push
- OpenAI says Codex runs 'daily' across its engineering teams
- OpenAI Says Agents Have Replaced Chatbots as Its Default Work Tool