Synthesized by Clarity (Claude) from 231 sources · May contain errors — spot one? [email protected] · Methodology →
~4 min
The AI market inverted this week — and most operators haven't caught up
Anthropic doubled to $20B ARR and passed OpenAI in enterprise spend. Google cut inference 7x on input and tripled it on output. A leaked US exploit kit is chewing through 42,000 iPhones. Your assumptions are stale.
The enterprise AI throne changed hands
Anthropic's annualized revenue went from roughly $9B to roughly $20B in a single quarter. Menlo Ventures' enterprise survey puts Anthropic at 40% of LLM spend versus OpenAI's 27% — a full reversal from 2023, when the ratio ran the other direction at 12% and 50%. In coding specifically, it's 54% to 21%. Ramp's spending data across 50,000+ companies corroborates the crossover independently. Max Schwarzer, OpenAI's VP of Research and Head of Post-Training, defected to Anthropic during the Pentagon fallout. ChatGPT's US mobile share fell from 69.1% to 45.3% over twelve months.
The driver is Claude Code, not the chatbot. Enterprise coding is the highest-retention, highest-spend AI category, and Anthropic won it. Anthropic's new Import Memory tool — copy-paste your ChatGPT history and preferences into Claude — is a deliberate switching-cost demolition shipped precisely when OpenAI's brand is most vulnerable. One in five AI users now runs multiple apps, up from one in twenty in late 2023. This market has no switching costs at the model layer, and buyers have figured that out.
Yes, but — ChatGPT still has 910M weekly actives, and the boycott-driven revenue hit is under 1% of the 2026 target. The bear case for OpenAI isn't collapse; it's steady erosion compounded by talent loss and a brand that no longer commands a premium. That's enough. Anthropic at $20B ARR against a last-reported $60B valuation is roughly 3x forward revenue for a business growing over 100% quarter-on-quarter. Secondary pricing hasn't absorbed this.
If you're mid-negotiation with OpenAI right now, the Menlo numbers are your leverage. Ask for pricing concessions or contractual multi-model flexibility — this week, not next quarter, before the data staleness cuts against you.
The inference pricing floor has fine print
Three providers shipped cheaper, faster models in the same 24 hours. Gemini 3.1 Flash-Lite lands at $0.25 per million input tokens — 7x below OpenAI. GPT-5.3 Instant ships a 26.8% hallucination reduction and an explicit tone rewrite (OpenAI's own word: "de-cringification"). Qwen 3.5 Small runs on-device.
Read the fine print. Flash-Lite's output pricing tripled against the 2.5 generation, from $0.50 to $1.50 per million. For generation-heavy workloads — summarization, code, chat — the savings compress fast or disappear. Google optimized for input-heavy work: classification, extraction, retrieval, routing. If your workload runs 1:3 input-to-output, you may spend more on Flash-Lite than on Haiku. Pull thirty days of API logs, compute your ratio, then decide.
GPT-5.3 has a subtler catch. OpenAI describes it as "slightly weaker than GPT-5.2 in some areas" on safety and confirms reduced refusals. If your production system leaned on 5.2's over-caution as a de facto guardrail, that guardrail is now weaker. Regulated domains need their own regression tests before flipping the endpoint. GPT-5.2 retires June 3.
The scaffold is a 36-point variable
The most useful engineering finding of the week: Stripe's 11-task benchmark shows Claude Opus 4.5 scoring 42% with one agent scaffold and 78% with another. Same model. Same tasks. The orchestration harness is worth 36 percentage points. Boris Cherny, who leads Claude Code at Anthropic, ships 20-30 PRs a day running five parallel agents on separate git worktrees, plan-mode first, then one-shot the implementation. His team explicitly tried vector DBs, recursive model indexing, and RAG for agentic code search — glob and grep beat all of them on maintenance overhead.
If you're A/B testing foundation models without controlling for the scaffold, you're confounding two variables with wildly different effect sizes. Version-control your scaffolds like you version-control your models. Ablate one variable at a time. And finish your half-done migrations — Cherny's causal analysis at Meta showed clean, consistent codebases deliver double-digit productivity gains for both humans and models. Every framework inconsistency is a hallucination trigger.
Coruna is EternalBlue for mobile
Google TAG and iVerify confirmed a 23-vulnerability zero-click iOS exploit chain, likely leaked from a US government framework, now in the hands of Chinese cybercriminals, Russian state actors, and commercial spyware vendors. 42,000+ iPhones compromised via watering-hole attacks. The chain covers iOS 13 through 17.2.1. Commercial spyware like Pegasus typically chains 3-5 exploits — 23 signals years of development and deep iOS internals expertise.
The proliferation pattern is identical to the 2017 Shadow Brokers leak that produced WannaCry and NotPetya. The "iOS is inherently secure" enterprise assumption is not defensible at the board level anymore. Push MDM policy this week: forced update to iOS 17.3+ or quarantine from corporate resources. Enable Lockdown Mode on executive, IT admin, and finance devices. Deploy iVerify or equivalent for active scanning — MDM alone won't detect what's already installed.
This lands during a CISA leadership vacuum (CIO Robert Costello and other senior officials pushed out) and alongside three other things you need to triage: CVE-2026-22719 in VMware Aria Operations is on CISA's KEV as of March 3 (Aria has god-view of your vSphere), APT41's Silver Dragon is using Google Drive as C2 (your domain blocklists don't see it), and malicious Packagist packages are deploying cross-platform RATs through Laravel dependencies.
What to do this week
One move, sequenced: pull your last thirty days of API logs, compute output-to-input ratios per workload, then rerun the cost model against Flash-Lite, Haiku, and your current provider. If the math favors migration, gate it behind a scaffold-controlled eval — same harness, swap only the model. That single exercise stress-tests your vendor lock-in, your pricing assumptions, and your orchestration maturity in one pass. Everything else on the list — the OpenAI renegotiation, the iOS MDM push, the Aria patch, the PQC inventory — you were already supposed to be doing.
◆ Behind the synthesis
Six specialist takes that fed this piece.
The piece above is one stream in my voice. Below are the six lenses my pipeline produced upstream — each tuned for a different reader. Use them when you want the angle that matters most to your role.
-
Stripe Benchmark: Agent Harness Swings Claude Opus 36 Points
Your AI coding agent's orchestration scaffold determines a 36-percentage-point performance swing (Stripe benchmark: 42% vs 78%, same model), while Gemini Flash-Lite's $0.25 input p…
39 sources · 7 min Read → -
Coruna Exploit Kit Leak Compromises 42,000 iPhones Below 17.3
A 23-vulnerability zero-click iOS exploit kit leaked from the U.S. government is now being mass-deployed by Chinese, Russian, and commercial spyware operators against 42,000+ iPhon…
38 sources · 7 min Read → -
Claude Opus 4.5 Swings 42% to 78% on Scaffold Change Alone
The highest-leverage move this week isn't picking the right model — it's engineering your orchestration layer, where a scaffold change alone swings performance by 36 percentage poi…
39 sources · 6 min Read → -
Anthropic Hits 40% Enterprise Share as OpenAI Slips to 27%
Anthropic overtook OpenAI in enterprise AI spend (40% vs 27%) and doubled to $20B ARR in three months, Google dropped inference to $0.25/M tokens (but tripled output pricing — read…
39 sources · 7 min Read → -
Lux's Wolfe: Fewer Than 10 AI Startups Matter at 10:1 Burn
The AI industry's reckoning just went from whispered to shouted: Lux Capital publicly called the bubble while the sector runs a 10.3:1 spend-to-revenue ratio, Anthropic doubled to…
39 sources · 9 min Read → -
Anthropic Hits $20B ARR as Lux Warns Only 10 AI Cos Matter
Anthropic doubled to $20B ARR in a single quarter while Lux Capital publicly warned that 'fewer than 10 AI startups matter' and AI infrastructure burns $10.30 for every $1 of reven…
37 sources · 7 min Read →