In an early-2025 experiment, 16 experienced open-source contributors completed 246 tasks on repositories they already knew, with permission to use AI assigned per task. The authors estimated that implementation took 19 percent longer when AI was allowed.
The participants forecast 24 percent less time beforehand. Afterwards, having done the work, they estimated 20 percent less time.
The result about tools is narrow and contested. The result about self-assessment is neither.
What the trial actually establishes
Scope first, because the headline travels further than the finding.
The treatment was permission to use AI on a task, not random assignment to a product and not mandatory use. The population was experienced contributors working on mature repositories they were familiar with. The measure was developer-reported implementation time before and after pull-request review, with some post-review time imputed, adjusted for forecast task difficulty.
So it does not estimate greenfield prototype speed, novice learning, downstream business value, or the behaviour of later tooling. METR’s own February 2026 update says a later experiment is an unreliable signal of the current effect because of participant and task selection and time-accounting problems — which is counterevidence against extrapolating, not a demonstration that the original randomised result was wrong.
Other work points differently on different endpoints. A combined analysis of three company experiments with 4,867 developers reports a 26.08 percent increase in completed pull requests, with a standard error of 10.3 percentage points, using weighted instrumental-variable estimates of adoption.
Weekly completed pull requests under estimated adoption, and implementation time on assigned tasks, are not the same quantity. Neither is a measure of value. “AI is now 26 percent faster” is a claim nobody made.
The finding that generalises
Perceived productivity is a person’s assessment of how work feels, or how it would have gone otherwise. Measured productivity requires a defined output, a defined input, and a comparison.
The gap between them here was roughly 39 percentage points, in the direction of the participants’ own interests being served by accuracy. These were experienced people, on their own code, estimating a thing they do every day.
That has consequences well beyond coding tools, because self-report is what pilots collect. Satisfaction surveys, perceived time savings, “would you keep using this”, and post-hoc estimates of hours saved are the standard evidence base for a business case, and they are the evidence class this result disqualifies for that purpose.
A May 2026 survey of 349 technical workers on perceived speed and value gains makes the same point from the other side: it uses a convenience sample with low email response, and its authors caution explicitly that responses need not match actual effects.
What to do instead
Collect experience reports alongside defined work, effort, quality, elapsed time, and value measures, and investigate the disagreements rather than picking whichever number is convenient.
A developer reporting that generation makes a task feel easier, where recorded effort falls for drafting and rises for checking and integration, has produced two true findings. The experience benefit is real and worth keeping. The total-effort result is also real. Neither is a revenue gain.
Feedback then needs its own discipline, because the counting goes wrong in a predictable direction. Distinguish a direct quotation from an observer’s note from an analyst’s interpretation. Count messages, distinct people, and distinct episodes separately.
Six messages about one export failure, all from one person during one episode, alongside a single report of an incorrect recipient from someone else, is not six times more evidence for the first problem. The second is not disposable for appearing once, and a single credible report of a consequential failure warrants investigation before any theme forms.
Three observers’ notes about one event are one event. Non-use and missing voices get actively investigated rather than read as satisfaction, and the recruitment route, scheduling, location, and activity all shape who is represented.
Reading the next study
Before deciding what a new result changes, classify it: a direct replication, a changed-setting extension, a different-endpoint study, or commentary.
Record the population, recruitment, tool version, task definition, assignment, adoption, outcome, follow-up, uncertainty, and exclusions. Then locate the exact contested claim and ask whether the new evidence changes the effect in the original setting, its applicability elsewhere, or only the plausible mechanism.
And do not count a study’s preprint, its journal abstract, and its accompanying blog post as three independent replications. That is the most common way a single result acquires the appearance of a literature.
The rule
What stays fixed is that a measure needs a defined output, a defined input, and a comparison. What changes is which endpoint a study measured, and endpoints that sound alike — time, throughput, value — are not substitutes for one another.
Not to be confused with
Dismissing subjective evidence. Experience reports reveal problems that measurement misses, and a tool people find exhausting has a real cost. The error is treating the report as the measurement.
A settled question. Direct replications on stable tasks with long-term quality outcomes remain outstanding, and the tooling changes faster than the studies do.