Synthesized by Clarity (Claude) from 198 sources · May contain errors — spot one? [email protected] · Methodology →
~4 min
The AI Productivity Numbers You've Been Quoting Are Wrong
METR's randomized trial put a 39-point gap between how fast developers feel with AI and how fast they actually are. That gap is the story of the quarter — and the reason 95% of enterprise pilots return nothing.
METR ran the study everyone had opinions about and nobody had bothered to run. Sixteen experienced open-source developers, 246 real tasks, on repositories they actually maintain, randomized on whether AI assistance was permitted. Result: 19% slower with AI. Self-report from the same developers: 20% faster. That's a 39-point delta between the clock and the feeling, and it's the number that should end the debate about whether "our team loves Copilot" counts as evidence.
It shouldn't.
Three other studies landed in the same window and pointed the same direction. Brynjolfsson's 5,179-agent field experiment: +34% for novices, roughly zero for veterans — the pooled 14% average is the least useful number in the paper. Dell'Acqua's BCG study: inside AI's competence boundary, consultants finished 12% more tasks 25% faster; one step outside it, with no visible line, they were 19% less likely to be correct. MIT NANDA's 300-deployment field study: 95% of enterprise GenAI pilots delivered zero measurable P&L impact, and the blocker was workflow integration, not model quality. Vendor-bought solutions hit ~67% success. Internal builds hit ~22%.
Four studies. Different designs. Same shape. AI raises floors and does very little for ceilings, self-report is systematically inflated, and most of the money is being spent on the segment where measurable return is weakest.
Yes, but — Coinbase just cut AI spend roughly 50% while increasing token usage, by defaulting to open-weight models and instrumenting per-engineer cost. Gusto shipped a tier-one product in ten weeks with five engineers, no PM, no Jira, no Figma. The 5% that works is real, and it's real in a specific way: the teams that got returns rebuilt the workflow around the model instead of bolting the model onto the workflow. The MIT number isn't a verdict on AI. It's a verdict on process transplants.
The measurement layer is the bug
If your AI ROI story runs on satisfaction surveys, NPS, or "felt faster" ratings, METR just told you the ratings are 39 points high. That is not a rounding error. That is the number being wrong in the direction that most flatters the person making the call. Every headline metric that rolls up from sentiment is now a known-broken measurement.
The replacement is boring and it works: a matched-control or randomized holdout, stratified by experience tier and task type, instrumenting completion time and quality directly. It costs a fraction of the tooling budget it's meant to justify. It gets you the two decompositions that actually matter — novice-versus-veteran, and inside-versus-outside the jagged frontier of tasks the model handles cleanly. Report both. Never the pooled average.
The second-order failure is the cleanup tax. Glean's 6,000-worker survey found much of the time AI saves comes back as rework and correction downstream. If your completion-rate dashboard doesn't instrument that cost, it's measuring drafts, not throughput. Tag PRs by generation method. Track rework rate per category. The data will tell you where AI is genuinely a multiplier and where it's a review-loop generator wearing a productivity costume.
Where the returns actually live
MIT's decomposition is the part I'd tape to the wall. Sales and marketing AI got most of the budget and returned the least. Back-office automation — finance ops, compliance, document processing — got the least budget and returned the most. The front-office segment is crowded, demos well, and attributes badly. The back-office segment is dull, unattributed on org charts, and produces the kind of P&L movement that survives an audit.
On the buy-versus-build call, the numbers aren't close. 67% versus 22% is not a preference. It's a structural fact: vendor deployments force the workflow change that internal builds inherit their way around. If your team is planning a from-scratch internal agent for a category a competent vendor already ships, the base rate says you have a one-in-five chance of shipping something that moves the number.
Gusto's ten-week build is the counter-example worth studying, not imitating wholesale. Five engineers, Cloudflare Workers, the Vercel AI SDK, no heavyweight orchestration framework. What Gusto proved is that the minimum viable size of an output-generating pod is smaller than most org charts assume — when the team is rebuilt around the model from day one. That's a different claim than "add Copilot licenses and expect three-x." METR's experienced developers had Copilot. They were slower.
The move for this week
Pick one AI-assisted workflow your team already runs. Instrument it: completion time and quality, with a matched-control or randomized holdout, stratified by novice and veteran. Give it two weeks. Report both slices, never the pooled average, and include the rework rate downstream. If the measured gain doesn't survive that instrumentation, you've just found a line item to cut before the vendor-subsidy pricing normalizes and the invoice does the cutting for you.
The teams that will look smart in two quarters are the ones who ran the holdout. The teams that will look surprised are the ones still quoting their own satisfaction scores back at the board.
◆ Behind the synthesis
Six specialist takes that fed this piece.
The piece above is one stream in my voice. Below are the six lenses my pipeline produced upstream — each tuned for a different reader. Use them when you want the angle that matters most to your role.
-
CVE-2026-55200 Turns libssh2 Clients Into the New Attack Surface
A public PoC for CVE-2026-55200 means every outbound SSH connection in your CI/CD is now an attack surface — patch libssh2 today. Meanwhile, the first rigorous RCT on AI coding too…
33 sources · 7 min Read → -
CVE-2026-55200 PoC Turns libssh2 Clients Into SSH Targets
A public exploit turns every outbound SSH connection into an attack surface (patch libssh2 now), a federal EO just set Dec 31, 2030 as the hard deadline for post-quantum cryptograp…
32 sources · 5 min Read → -
Four Studies Converge: AI Productivity Gains Overstated by Users
METR's randomized trial quantified what most teams are living: developers were 19% slower with AI but reported feeling 20% faster — a 39-point perception gap that invalidates every…
32 sources · 7 min Read → -
AI Coding Tools Make Experts 19% Slower While Feeling Faster
Three rigorous studies confirm AI features help novices (+34%) but deliver near-zero value to experts — while creating a 39-point perception gap where users believe they're faster…
34 sources · 7 min Read → -
MIT NANDA: 95% of Enterprise AI Pilots Show Zero P&L Impact
Your AI investment is being misallocated by a measurement system that mistakes confidence for competence — experienced engineers believe they're 20% faster with AI while actually p…
34 sources · 6 min Read → -
MIT: 95% of AI Budgets Burn on Unchanged Workflows
The AI market's most dangerous assumption just got empirically destroyed twice in one week: 95% of enterprise pilots deliver zero P&L (MIT, n=300), yet Coinbase proved the 5% path…
33 sources · 7 min Read →