Leader daily

Synthesized by Clarity (Claude) from 34 sources · May contain errors — spot one? [email protected] · Methodology →

MIT NANDA: 95% of Enterprise AI Pilots Show Zero P&L Impact

Sources
34
Words
1,298
Read
6min

Topics Agentic AI AI Capital AI Regulation

◆ The signal

MIT NANDA says the same at scale, with 95% of enterprise AI pilots showing zero P&L impact. The pattern is familiar. Budget tracks whoever lobbies loudest, which is sales and marketing, while the measurable ROI sits in back-office automation nobody fought for.

◆ INTELLIGENCE MAP

Intelligence map

  1. 01

    AI Productivity Measurement Crisis

    act now

    Four independent studies converge on the same finding: AI productivity gains are real but systematically mismeasured. Novices gain 34%; experts lose 19% while believing they gained 20%. MIT's 300-deployment study traces 95% pilot failure to organizational rigidity, not model quality. Budget is misallocated to customer-facing AI while measurable ROI sits in back-office automation.

    39%
    perception-reality gap
    5
    sources
    • Novice productivity
    • Expert actual output
    • Expert self-report
    • Pilots with P&L impact
    1. Novice actual gain34%
    2. Expert perceived gain20%
    3. Expert actual gain-19%
    4. Vendor solution success67%
    5. Internal build success22%
  2. 02

    Post-Quantum Cryptography: The $1B+ Compliance Cliff

    monitor

    June 22 executive order sets hard deadlines — Dec 31, 2030 for key establishment, Dec 31, 2031 for digital signatures. Pentagon strategy mandates PQC-incompatible systems be phased out. The crypto-agility requirement means this isn't an algorithm swap — it's an architectural transformation. Organizations that lack a full cryptographic inventory will discover the cost of finding out.

    $1B+
    forced modernization
    4
    sources
    • Key deadline
    • Signature deadline
    • Harvest-now window
    • M&A acceleration
    1. EO signedJune 2026
    2. Inventory required2027-28
    3. Key establishmentDec 2030
    4. Digital signaturesDec 2031
  3. 03

    The 5-Person Product Team Is No Longer an Outlier

    monitor

    Gusto's CTO shipped a tier-one product in 10 weeks with 5 engineers — no PM, no Jira, no Figma. AI did primary building. Concurrently, GLM-5.2 ran a 45-minute autonomous coding session processing 6M tokens for $3.36. When an AI engineer-equivalent costs single-digit dollars per hour, the constraint isn't budget — it's whether your org chart can absorb the capability.

    5
    engineers, tier-1 launch
    4
    sources
    • Team size
    • Time to ship
    • AI session cost
    • Tokens processed
    1. Traditional team30~6 months
    2. AI-native team510 weeks
  4. 04

    Agent Governance Window: 6 Months Before Standards Lock

    monitor

    Okta shipped agent identity governance GA for FedRAMP/HIPAA. Workday embedded domain-specific guardrails at the inference layer. Google launched an MCP Store. RBC survey confirms enterprises are funding AI with net-new budgets, not reallocation. The 6-month window between new budget formation and governance standard hardening is where market position gets defined.

    6 months
    governance standard window
    5
    sources
    • Budget type
    • Okta GA scope
    • x402 daily txns
    • x402 growth
    1. x402 transactions500,0005x/month
    2. Airwallex valuation$11B+38%
    3. Pangea banks5016 countries
  5. 05

    AI Infrastructure Financial Strain Materializing

    background

    Oracle posted its worst week in 25 years — $130B debt, 162% capex increase, negative $24B FCF, 55% drawdown from peak. South Korea committed $880B over a decade. The AI quarterly profitability threshold was crossed (revenues exceed depreciation) but cumulative capex remains unrecovered. The overbuild creates buyer leverage now, seller pain later.

    $130B
    Oracle debt load
    3
    sources
    • Oracle drawdown
    • Oracle FCF
    • SK commitment
    • AI quarterly profit
    1. Oracle capex growth162%
    2. Oracle peak drawdown55%
    3. SK 10yr commitment$880B
    4. AI revenue vs. depreciation105% (crossed)

◆ DEEP DIVES

Deep dives

  1. 01

    The 39-Point Lie: Why Your AI Productivity Numbers Are Systematically Wrong

    act now

    Four studies on AI productivity measurement

    This quarter's four independent studies converge rather than contradict. Together they form the clearest indictment yet of how enterprises measure AI impact. The findings cut at the measurement layer, not the technology:

    • METR Developer Study: Experienced open-source developers ran 19% slower with AI tools while believing they were 20% faster. That is a 39-percentage-point gap between sentiment and output.
    • Brynjolfsson's 5,179-agent study: Junior customer support agents gained 34% productivity. Veterans saw near-zero improvement.
    • Harvard/BCG Consultant Experiment: Large gains inside the competence boundary, 19% worse performance outside it.
    • MIT NANDA (300 deployments, 150 executive interviews): 95% of enterprise GenAI pilots deliver no measurable P&L impact.
    AI compresses rather than amplifies. It raises floors, not ceilings. If engineering leadership reports that AI tools are 'working great' based on team sentiment, the bill arrives twice: once in dollars, once in false confidence.

    Where the Money Is vs. Where the Money Went

    MIT NANDA locates pilot failure in organizational rigidity, not model quality. Firms bolt AI onto workflows they refuse to change. The specifics matter:

    ApproachSuccess RateWhy
    Vendor-purchased solutions~67%Forces process adaptation
    Internal builds~22%Inherits existing dysfunction
    Sales/marketing AI (most budget)Low measurable ROIDemos well, hard to attribute
    Back-office automation (least budget)Highest measurable ROIDull to present, clear to measure

    The Expert Pipeline Crisis Nobody Is Pricing In

    Stanford's Canaries dashboard, built on ADP payroll data covering 1-in-6 US workers, shows employment for 22-to-25-year-olds falling in AI-exposed roles while rising for less-exposed peers. The chain is causal: AI helps juniors most, so junior roles become the most automatable, and automating them removes the apprenticeship rung that produces future experts. This is a leveraged bet that AI capability improves faster than the expertise reservoir drains. If it doesn't, the capability loss doesn't reverse on any normal timeline.


    The Tradeoff Nobody Wants to Name

    The board-deck version says deploy AI across the org and capture the gains. The complete version is harder. Capturing gains means moving money away from the function that asked loudest, sales and marketing, toward the function that didn't, back-office. It means retiring self-reported productivity metrics for instrumented A/B measurement. It means 70% of transformation budget going to process change and 30% to technology, the inverse of current allocation.

    Action items

    • Kill all self-reported AI productivity metrics as primary indicators — implement instrumented A/B output measurement for every AI-assisted workflow by end of Q3
    • Rebalance AI investment portfolio: audit current split between customer-facing vs. back-office automation and shift 40% of customer-facing budget to back-office within two quarters
    • Design explicit expertise-formation pathways that coexist with AI augmentation — protect the apprenticeship rung for 22-25 year old hires
    • Mandate vendor-purchased solutions over internal builds for non-core AI use cases

    Sources:AI Weekly · TLDR AI · 🌀 Refactoring · Azeem Azhar, Exponential View · Phil from BOI (Board of Innovation)

  2. 02

    Post-Quantum Cryptography: A Dated Obligation With a Billion-Dollar Price Tag

    monitor

    What Changed on June 22

    The executive order on post-quantum cryptography turned an optional engineering project into a dated obligation. We have heard the word urgent attached to crypto transitions before, and four times it did not move budgets. This time the deadline travels with a federal/DOD purchasing gate, which is the part that moves budgets. For anyone selling to the government or sitting on long-lived sensitive data, being late means exclusion from the largest forced technology refresh since cloud migration.

    • December 31, 2030: Key establishment must use PQC algorithms
    • December 31, 2031: Digital signatures must use PQC algorithms
    • Pentagon parallel mandate: Systems that cannot support PQC by 2030 get phased out
    The mandate reveals which organizations actually know what their systems encrypt and which only think they do. The billion-dollar figure is the cost of finding out.

    Crypto-Agility Is the Real Requirement

    DOD is not asking for today's post-quantum algorithms. It is requiring the ability to swap cryptographic primitives as standards evolve. That is a ten-year architectural decision wearing the costume of a feature request. NIST will revise the standards. New vulnerabilities will surface. Firms that hard-coded today's primitives will retrofit under deadline while agile platforms swap and move on. A point solution clears one audit. A platform clears the next three.

    The Harvest-Now Threat Is Already Active

    Multiple sources confirm nation-state adversaries are already collecting encrypted traffic for future quantum decryption. The breach is not a future event. It is happening now, with the damage deferred. Organizations holding data with multi-year confidentiality requirements — healthcare, financial, defense, enterprise IP — are already on the clock.

    Complicating Factor: Standards May Be Compromised

    Allegations surfaced this week that intelligence agencies are weakening post-quantum TLS standards, with a July 7 comment deadline. Low confidence, extreme consequence. A reasonable skeptic would call this noise, and most weeks they would be right. If the allegation holds, crypto-agility stops being an engineering nicety and becomes a hedge against standards capture itself.


    Market Timing

    Nobody migrates cryptographic assets they have not inventoried. Discovery comes first, by logic rather than preference. Whoever owns discovery owns the relationship that sells migration and runs the ongoing management. That is land-and-expand compressed into a four-year government deadline. Expect M&A in cryptographic discovery to accelerate over the next 12-18 months as the larger players decide buying is faster than building.

    Action items

    • Assess federal/DOD revenue exposure and map PQC compliance requirements against current cryptographic architecture by Q4 2026
    • Commission a full cryptographic inventory — identify every system, protocol, and dependency using vulnerable algorithms
    • Architect all new systems for crypto-agility as a baseline requirement effective immediately
    • Evaluate M&A targets in cryptographic discovery/inventory and crypto-agility tooling

    Sources:CyberScoop · TLDR Crypto · TLDR InfoSec · The Hacker News

  3. 03

    The 5-Person Team Shipped in 10 Weeks — Your Org Chart Is 2022 Vintage

    monitor

    The Gusto Case Study

    Eddie Kim, Gusto co-founder and CTO, shipped a tier-one production product with a five-person team in ten weeks. No PM. No Jira. No Figma. No standups. AI did the primary building. The coordination layer was a permanent Zoom room. This is not a minor acceleration — it's a structural break in how products get built.

    The natural skeptic objections: maybe the product was simpler than 'tier-one' implies, maybe the five were unusually senior, maybe the rest of the org absorbed hidden work. All plausible. None explains why a company with the option to staff thirty chose five and still hit ten weeks.

    If this model generalizes to greenfield products — and the evidence points that way — most companies are operating org charts designed for 2022. The handoffs, specifications, and coordination rituals that justified 15-person product teams may now be net-negative overhead dressed up as quality control.

    The Economics Make It Structural

    GLM-5.2 ran a 45-minute autonomous coding session processing 6 million tokens for $3.36. At that cost, an AI engineering-equivalent runs at single-digit dollars per hour. The binding constraint stops being budget. It becomes architecture and organizational readiness to absorb the capability.

    Gusto's production agent runs on Cloudflare Workers with the Vercel AI SDK. No LangChain, no CrewAI, no heavy agent framework. Kim's working definition: "an AI SDK running somewhere in the cloud, able to look up files and call tools." Anyone invested in complex orchestration platforms should read this as a leading indicator that the market pays for simplicity.

    The Contradiction That Demands Resolution

    Sources diverge on an important point. The Gusto data says small AI-native teams can ship at extraordinary velocity. The METR data says experienced developers get slower with AI tools. The reconciliation: the gains come from rebuilding the team around AI, not from adding AI to an existing team. Kim didn't point AI at his existing org and ask it to run faster. He rebuilt from zero with AI as a core contributor. The bolt-on approach produces the expert penalty. The rebuild produces the 6x speed advantage.


    China's Talent Arbitrage Compounds This

    Chinese AI labs staff with talent averaging 1.6 years experience vs. 5.5 years in the US — a 3.4x gap. One interpretation: quality disadvantage. The more threatening interpretation: experience beyond a threshold is actually drag in a field where paradigms shift faster than institutional knowledge accumulates. If that interpretation holds, US companies spending 3-4x on AI talent may face diminishing returns on experience premiums.

    Action items

    • Run a controlled 10-week AI-native team experiment: one greenfield product, 5 senior engineers, no traditional process tools, direct exec decision-making — measure output against your standard team velocity
    • Audit headcount assumptions for all new product initiatives — model scenarios where AI-native teams replace traditional staffing ratios
    • Evaluate heavy agent infrastructure investments (LangChain, CrewAI equivalents) against minimal-stack alternatives
    • Test whether your 5+ year experience hiring requirement is genuine capability need or organizational inertia — run a parallel cohort with 1-2 year experience engineers on AI-augmented workflows

    Sources:Lenny's Newsletter · TLDR AI · TLDR Founders · Azeem Azhar, Exponential View

◆ QUICK HITS

Quick hits

  • Update: DeepSeek raises $7.4B — largest first-time Chinese startup round ever, triggered by Anthropic's Mythos Preview; plans to double every department

    The Information AM

  • Update: Coinbase confirms 50% AI spend reduction by switching to Chinese open-weight models — arithmetic that replicates across enterprise

    The Information AM

  • Oracle posts worst week in 25 years — $130B debt, -$24B free cash flow, 55% drawdown from $900B peak, all driven by AI infrastructure bets against OpenAI demand

    Finpresso

  • Adversarial prompt injection now blinds AI security tools — malware embeds prompts causing LLM-based detection to refuse analysis entirely

    CSO First Look

  • AI shopping agents converting at 4.4x human rates with 7,851% YoY growth — but 77% of traffic goes to product pages that WAFs are blocking as bot traffic

    TLDR Marketing

  • NVIDIA's ENPIRE framework achieves 99% robotic task success without human intervention — self-improving physical AI shifts timeline from decade to years

    Jack Clark from Import AI

  • Vulnerability weaponization now under 24 hours — Cisco CUCM exploited within one day of PoC release, making traditional patch SLAs documented gaps

    TLDR InfoSec

  • AI coding tools shifting from flat-rate to per-token billing — analysis of 12TB of agent logs shows consumption curves that don't respect Q4 budget assumptions

    TLDR Data

  • Sector rotation visible: ServiceNow +9.85%, Snowflake +9.65% vs. ON Semi -23.66% — market repricing AI value from hardware to software layer

    Finpresso

  • JLR cyber attack caused $2.5B loss requiring £1.5B government bailout — new reference case for board-level cyber risk quantification

    TLDR InfoSec

◆ Bottom line

The take.

Your AI investment is being misallocated by a measurement system that mistakes confidence for competence — experienced engineers believe they're 20% faster with AI while actually producing 19% less, and 95% of enterprise pilots never touch the P&L because they target the wrong workflows. The fix is surgical: kill self-reported metrics, redirect budget from customer-facing AI (where it demos well) to back-office automation (where it measurably returns), and run a single 5-person AI-native team experiment that will tell you more about your organizational readiness than any pilot program.

— Promit, reading as Leader ·

Frequently asked

Why are self-reported AI productivity gains unreliable?
The METR developer study found experienced open-source developers ran 19% slower with AI tools while believing they were 20% faster — a 39-point gap between sentiment and measured output. Any dashboard built on team surveys or perceived velocity is systematically overstating impact and steering investment toward the wrong workflows.
Where should AI budget actually be going if not sales and marketing?
Back-office automation, which delivers the highest measurable ROI but attracts the least internal lobbying. MIT NANDA's data also shows vendor-purchased solutions succeed roughly 67% of the time versus 22% for internal builds, because vendors force process change instead of inheriting existing dysfunction.
What does the post-quantum cryptography executive order actually require by when?
Key establishment must use PQC algorithms by December 31, 2030, and digital signatures by December 31, 2031, with a parallel Pentagon mandate to phase out systems that cannot support PQC by 2030. Compliance is tied to federal and DOD purchasing eligibility, making it a revenue gate rather than an engineering preference.
How did Gusto ship a tier-one product with only five people in ten weeks?
CTO Eddie Kim rebuilt the team around AI from zero rather than bolting AI onto an existing org — no PM, no Jira, no Figma, no standups, just a permanent Zoom room and a lightweight stack on Cloudflare Workers with the Vercel AI SDK. The gains come from restructuring, not from adding tools to legacy teams.
Should we still be hiring senior engineers if AI helps juniors most?
The evidence is mixed and worth testing directly. Brynjolfsson's 5,179-agent study showed juniors gained 34% while veterans saw near-zero improvement, and Chinese labs operate with 1.6 years average experience versus 5.5 in the US. Running a parallel cohort of less-experienced engineers on AI-augmented workflows is the only way to know whether your experience premium is capability or inertia.

◆ Same day, different angle

Read this day as…

◆ Recent in leader

Keep reading.

Spot an error? [email protected]