AI Coding ROI Shows Up in Outcomes, Not Yet in Dollars
The most-cited productivity number and the only controlled measurement of it point in opposite directions. Gartner's 2026 survey reports a 19.3% net average gain; the one randomized trial, from METR, measured a 19% slowdown. Gartner itself reframes the question as "less about whether enterprise AI coding agents deliver value, and more about at what cost that value is realized", and nobody publishes the cost side at all. So the defensible ROI case is built from the outcomes the pipeline already records.
Perception says plus 19%, the one randomized trial said minus 19%
The widely cited gain is a self-report. Gartner's 2026 Software Engineering Content Survey found that "90% of engineering leaders now report productivity gains attributable to AI, with a net average productivity gain of 19.3%," but the survey covered 482 respondents in the US and UK at firms above $250M revenue, and Gartner itself disclaims that the results "do not represent global findings or the market as a whole." It is a perception number.
The only randomized controlled trial cuts the other way. METR's 2025 RCT found experienced open-source developers were 19% slower with early-2025 AI tools, having expected a 24% speedup and still believing afterward they had been sped up by 20%. METR's 2026 follow-up found a directional speedup but declared its own estimate unreliable because developers were withholding the tasks they did not want to do unassisted. There is currently no clean measured productivity number for late-2025 frontier tools. Treat 19.3% as how leaders feel, not as a measured return.
Most enterprise GenAI deployments report no measurable P&L impact
The disappointment data is consistent across three independent surveys. MIT's Project NANDA report "The GenAI Divide" (Aditya Challapally, August 2025) found that "about 5% of AI pilot programs achieve rapid revenue acceleration; the vast majority stall, delivering little to no measurable impact on P&L," and attributed the failure not to model quality but to "the learning gap for both tools and organizations." That study covers enterprise GenAI broadly, not coding agents specifically, and the underlying report has no stable public URL, so it is best cited through its press coverage. McKinsey's November 2025 State of AI (1,993 respondents) found that while 39% report AI affecting EBIT, in most cases less than 5% of EBIT is attributable to it, and only 21% of organizations had redesigned any workflow. Bain's 2025 software-development report found teams seeing "10% to 15% productivity boosts, but often the time saved isn't redirected toward higher-value work," and that only 23% of organizations could tie AI initiatives to new revenue or lower cost.
The signals that survive the gap are the outcomes the pipeline emits
Where companies do report numbers, they are quality proxies, not dollars. Intercom is the sharpest: AI-authored code reverts at roughly a tenth the human rate, with time-to-approval 6 to 16x faster at p75. Spotify reports a 76% rise in PR frequency, and Uber's uReview runs on over 90% of ~65,000 weekly diffs with 75% of its comments rated useful. These are the metrics to trust because they are recorded as a byproduct of normal operation, the whole argument of measuring outcomes. None of them is a dollar figure.
What separates the winners is the platform, not the model
The clearest predictor of return is organizational, not technical. The 2025 DORA report (~5,000 respondents) states it plainly: "AI doesn't fix a team; it amplifies what's already there. Strong teams use AI to become even better. Struggling teams will find that AI only highlights and intensifies their existing problems." DORA found AI adoption positively related to throughput and product performance but still negatively related to delivery stability, and that organizations with high-quality internal platforms unlocked more of the value. McKinsey's finding that workflow redesign correlates highest with EBIT impact says the same thing from the survey side. The teams that get ROI are the ones that already had closed feedback loops and a platform to plug agents into. Adoption that sticks follows the same pattern: cohorts beat mandates, and you grade outcomes, not usage.
Dollar ROI is unpublished, so budget against scenarios
No company in the public record attaches a dollar cost to a merged change or reports net ROI in currency. Gartner sizes the market at $9.8 to $11.0 billion annualized and predicts AI coding costs will overtake the average developer's salary by 2028 as pricing shifts toward consumption, but that is a planning assumption, not a measurement. The practical move is to instrument the value side from the outcomes above and the cost side yourself, since cost per merged change is the metric nobody publishes, and to treat the salary-overtake forecast as a scenario rather than a result. The related decision, whether the next dollar goes to a hire or to more tokens, turns on review capacity, not on the model.
Last verified: June 2026
The survey figures, market sizing, and the 2028 forecast here decay quickly and several sit behind paywalls. Re-verify against the primary sources before using any of these numbers in a board or budget context.