Study Finds AI Coding Agents Generate More Code, but Not More Software
A Harvard study using Jellyfish analytics from over 700 software firms found AI coding tools produce more code but not more software, as longer human review absorbs the time saved.
The research was carried out by Harvard University researchers Fiona Chen and James Stratton and reported by Ars Technica. The authors drew on aggregated analytics data from Jellyfish, which measures the granular output of engineering teams, along with issue management software data. The dataset covers 300 million individual work events, such as commits and pull requests, across more than 700,000 employees at over 700 relevant software development firms, from 2021 through March 2026.
The study identifies human code review as a significant bottleneck for the overall efficiency of AI coding tools, concluding there is "little evidence that firms increase software output or reduce employment" through their use. Any efficiency gained during the actual coding phase, the authors write, is "absorbed by downstream constraints in the production process."
The review stage itself shifts as more AI-generated code enters the pipeline. According to the study, the code review process significantly increases in length, pull requests are more likely to require revisions, and reviewers leave more comments.
Those findings line up with what programmers using the tools describe. Modern AI coding assistants and agents can be highly efficient at generating large amounts of functional code, but developers who rely on them do not trust the accuracy of that output, so substantial effort goes into reviewing it before it is accepted. The study's authors summarize the dynamic with the principle of cutting once and measuring twice, placing the measurement burden at the review stage rather than the generation stage.
Jellyfish's analytics record individual work events such as commits and pull requests, giving the researchers a view of both code generation and code review across the firms in the sample. The measurement period runs from 2021 to March 2026, spanning the period in which AI coding assistants and agents moved into routine use at software organizations.