The Evidence: System Measurements Beat Self-Reports
The four most-cited numbers on AI coding outcomes contradict each other.
- RCT. Experienced developers got 19% slower with AI.
- Leader survey. A 19.3% net productivity gain.
- Adoption survey. 90% of practitioners say they use AI and most say it helps.
- Instrumented pipeline. One team found AI-authored code reverting at one tenth the human rate.
All four are real measurements. They disagree because they measure different things: what people feel, what individuals do, and what systems produce. The third category, what systems produce, is the only one you can operate from. The case studies put faces on it.
Last verified: June 2026
The market numbers on this page are perishable. Re-check Gartner, DORA, METR, and Intercom before using them in a board memo, procurement packet, or audit narrative.
Four real measurements, four different things measured
| Number | What it measures | Method | Source |
|---|---|---|---|
| 90% adoption, >80% perceived productivity gain | Practitioner self-report | Survey, ~5,000 technology professionals | DORA 2025 |
| 19.3% net productivity gain, reported by 90% of engineering leaders | Leader self-report | Survey, 482 US/UK leaders at $250M+ companies | Gartner Software Engineering Content Survey for 2026, cited in the May 2026 Magic Quadrant for Enterprise AI Coding Agents (gated) |
| 19% slower with AI (95% CI: +2% to +39%) | Measured task time | RCT, 16 experienced open-source developers, 246 issues | METR, July 2025 |
| 0.53% revert rate for AI-authored backend code vs 5.39% human | Production outcome | Instrumented pipeline, two main codebases | Intercom |
Developers cannot feel their own speedup
METR's RCT is the most rigorous individual-level measurement published. Sixteen experienced contributors to large open-source repositories (averaging 22k+ stars and 1M+ lines of code) worked 246 real issues, each randomized to AI-allowed or AI-disallowed, with screens recorded. With early-2025 tools, AI-allowed tasks took 19% longer.
The sharper result is the perception gap. Developers forecast a 24% speedup. After living through the slowdown, they still believed AI had sped them up by 20%. Perception missed measured reality by roughly 40 percentage points, in the optimistic direction.
The 19% slowdown is already obsolete, with no replacement
METR now banners the 2025 post: "These results are out of date." Its February 2026 continuation (57 developers, 143 repos, 800+ tasks) found directional speedup: −18% for returning developers (CI: −38% to +9%) and −4% for newly recruited ones (CI: −15% to +9%). Then METR declared its own estimate unreliable. Developers increasingly refused to participate rather than work without AI, 30% to 50% admitted withholding tasks they did not want to do unassisted, and the pay cut from $150/hr to $50/hr compounded the selection bias. METR calls the result a likely lower bound and is redesigning the experiment.
There is no clean, current, measured number for how much AI speeds up developers. Anyone quoting one with confidence is ahead of the evidence. The honest state of the field: probably faster than early 2025, magnitude unknown.
The big adoption numbers are self-reports, and they run hot
DORA's 2025 report, drawing on nearly 5,000 survey responses, found 90% of respondents using AI at work and more than 80% believing it raised their productivity, while 30% report little or no trust in AI-generated code. Gartner's survey behind the 19.3% figure asks engineering leaders, not instruments, what they gained.
METR's early-2026 survey of 349 technical workers shows how hot self-reports run: median self-reported speed change of 3x, value change of 1.4–2x. Of seven respondents claiming 10x or more, the two with checkable public output had, in METR's judgment, overstated their gains. METR's own staff gave the lowest estimates of any subgroup. Treat every self-reported productivity figure as an upper bound with a roughly 40-point optimism bias attached.
The numbers you can operate are system measurements
Intercom did not ask anyone how they felt. They instrumented the pipeline: 93% of PRs agent-driven, 19% auto-approved with no human reviewer, zero reverts of AI-approved PRs in a 100+ PR pilot, and a 6–16x improvement in p75 time-to-approval. AI-authored backend code reverted at 0.53% versus 5.39% for human-authored; frontend, 0.22% versus 2.00%. These are production outcomes from a redesigned review system, not recollections.
DORA points at the same mechanism from the survey side: in 2025, AI adoption correlated positively with delivery throughput and product performance for the first time, but still negatively with delivery stability. Their summary: "AI doesn't fix a team; it amplifies what's already there." Teams with fast feedback loops and strong automated testing see gains; teams without them see instability.
This reconciles the contradiction. METR measured a tool dropped into an unchanged individual workflow. Intercom measured a system rebuilt around the tool, with enforced review gates and full instrumentation. The spread between "19% slower" and "10x fewer reverts" is the clearest published illustration of how much the system around the tool matters.
Quote no survey number as settled; instrument your own pipeline
Do not quote 19.3% or 19% as if either settles anything; both are snapshots of different questions. Gartner's forward claim, that asynchronous agent workflows will deliver 30% to 50% productivity gains by 2028 versus 0% to 20% from 2025-era assistants, is a forecast, not a measurement. The defensible move is Intercom's: instrument revert rates, time-to-approval, and stability in your own pipeline, and let those numbers argue for you. Measuring outcomes covers how.