Synthesized by Clarity (Claude) from 242 sources · May contain errors — spot one? [email protected] · Methodology →
~4 min
Your AI supplier is now your competitor, and your compute is being rationed against you
Microsoft's CFO admitted starving Azure to feed internal Copilot. Anthropic silently cut Claude's cache TTL 12x and shipped a Lovable clone the same week. The abstraction layers you trusted are load-bearing risks now.
The single most revealing sentence this week came from Amy Hood, not a research lab: Azure's growth number would have exceeded 40 if Microsoft had allocated its incoming GPUs to external customers. It didn't. M365 Copilot and GitHub Copilot got the compute. Azure customers got the leftovers, and the CFO said so on a call.
That one admission reframes everything else that happened this week. Anthropic's revenue tripled from $9B to $30B annualized in a quarter, and it responded by locking up a multi-year CoreWeave deal and a 3.5GW Broadcom/Google agreement starting 2027. Meta poached three senior Stargate infrastructure architects — Peter Hoeschele, Shamez Hemani, Anuj Saharan — into a new "Meta Compute" group reporting near Zuckerberg. Musk went further and partnered with Intel on a fab. When your largest AI buyers vertically integrate compute, the foundry-and-hyperscaler model isn't clearing.
The old cloud contract assumed a utility relationship: you pay, they serve, marginal cost is roughly zero. AI broke that. Every GPU serving your inference is a GPU not serving your provider's higher-margin internal product. Rational management picks itself. That isn't a scandal — it's the new equilibrium, and your SLA doesn't cover it.
The supplier is now the competitor
Anthropic shipped three products into this environment simultaneously: Ultraplan, Epitaxy, and Claude for Word — the last one embedded directly inside Microsoft's own Office suite with Track Changes fidelity. Leaked screenshots also show a vibe-coding app builder inside Claude, which is a direct shot at Lovable's $6.6B valuation four months after its round closed. When the model provider ships the application, every wrapper startup priced above half a billion needs a platform-risk discount applied by Monday.
On March 6, Anthropic quietly cut Claude Code's prompt cache TTL from 60 minutes to 5 minutes. No changelog. No announcement. Users found it in a GitHub issue. For any agentic loop reusing large context inside an hour — the exact workload Claude Code was sold for — that's up to a 12x cost increase overnight. Leaked session analysis separately suggests Opus 4.6's thinking depth fell around 67%, with developers migrating to Codex and GPT-5.4 in real time.
Yes, but — the counter-reading is real: "thinking depth" isn't a standardized metric, the sample is opaque, and Anthropic may have adjusted system prompts or output caps rather than the underlying model. Fine. The behavioral signal still stands, because the cache TTL change is not disputable and the migration is measurable. Provider quality is now something that regresses without notice, and you have no contractual recourse.
The supply chain underneath both of you is on fire
While the platform layer was moving, the layer under it got worse. APT41 deployed an ELF backdoor that scores 0/72 on VirusTotal, queries cloud metadata APIs across AWS, GCP, Azure, and Alibaba for IAM credentials, AES-256 encrypts them, and exfiltrates over SMTP port 25 to 43.99.48.196. Lateral movement uses UDP broadcasts on port 6006 — TensorBoard's default. Your ML monitoring traffic and their lateral movement are indistinguishable without deep packet inspection.
Researchers running a proxy called Mine caught nine LLM API routers — one paid, eight free — actively injecting malicious payloads into model responses and exfiltrating secrets. If you added a router for cost optimization or rate limiting, every inference call is an unvalidated trust boundary right now. Trivy, Xygeni, and KICs — the scanners guarding your CI/CD — were themselves compromised, with shared C2 infrastructure linking the Xygeni backdoor to a residential proxy botnet. The people building router botnets are the same people compromising your security scanners. Adobe broke Patch Tuesday for CVE-2026-34621, a Reader zero-day exploited silently since November. Marimo notebooks got a pre-auth RCE weaponized inside 12 hours. AI is now writing PoCs on patch day, and your 30-day critical SLA is fiction.
The thing worth stealing this week
Inside all this noise, LinkedIn published one of the most useful production ML disclosures of the year: five retrieval systems collapsed into one dual-encoder LLM serving 1.3B users under 50ms. The transferable part isn't the architecture, it's the finding that raw numeric tokens are invisible to transformer encoders. Feeding "views:12345" produced a correlation of -0.004 with embedding similarity — statistical zero. Wrapping percentile buckets in special tokens (<view_percentile>71</view_percentile>) delivered 30x correlation and 15% Recall@10 lift. If you feed structured numeric features into any transformer encoder for retrieval, ranking, or tabular prediction, this is a one-day change with double-digit quality upside. The offline numbers lack confidence intervals and there's no A/B disclosure — treat them as directional, not gospel. Run the ablation anyway.
What to do this week
Stop treating your LLM provider as a stable dependency. Ship a routing abstraction — LiteLLM, OpenRouter, or your own — that lets you swap Claude, GPT-5.4, and at least one open-weight model behind a flag. Instrument a daily eval that runs a frozen prompt set against production and alerts on distributional drift in output length, reasoning depth, and task completion. Diff your Anthropic bill against a March 5 baseline; if agentic workloads jumped, the cache TTL change is why, and you can negotiate. Enforce IMDSv2 today, block outbound SMTP from non-mail workloads, and pin your scanners to content hashes rather than version tags. Send canary prompts through every LLM router in your path and compare against direct API responses.
The piece of infrastructure most worth building this quarter isn't a new agent. It's the layer that lets you leave — cleanly, on a Tuesday afternoon, without an incident review.
◆ Behind the synthesis
Six specialist takes that fed this piece.
The piece above is one stream in my voice. Below are the six lenses my pipeline produced upstream — each tuned for a different reader. Use them when you want the angle that matters most to your role.
-
9 LLM API Routers Caught Injecting Code and Exfiltrating Secrets
Your AI supply chain is under coordinated attack at three layers simultaneously — 9 LLM API routers injecting malicious code, Trivy/Xygeni/KICs scanners sharing C2 with a botnet, A…
41 sources · 9 min Read → -
APT41 Harvests Cloud IAM Keys via SMTP with 0/72 AV Detection
APT41 is harvesting your cloud IAM credentials with a backdoor no antivirus detects, three of your vulnerability scanners were supply-chained by the same group running a router bot…
39 sources · 9 min Read → -
LinkedIn Bucketing Trick Lifts LLM Recall@10 by 15% at 1.3B Scale
LinkedIn proved that LLMs are literally blind to raw numeric features (-0.004 correlation), fixable with a one-day percentile bucketing change that delivered 15% Recall@10 lift — w…
40 sources · 8 min Read → -
ServiceNow Kills AI Add-On SKUs as Seat SaaS Sheds 50.5%
Seat-based SaaS lost half its market value in six months, and the winners are already visible: ServiceNow made AI free-by-default across 85 billion workflows, a16z confirmed enterp…
41 sources · 10 min Read → -
Microsoft CFO Admits Azure Capacity Diverted to Internal AI
Your cloud provider is now your compute competitor — Microsoft deliberately starved Azure to feed internal AI, Meta weaponized infrastructure hiring against OpenAI, and Anthropic's…
41 sources · 8 min Read → -
OpenAI Memo Concedes Microsoft Deal Caps Enterprise Reach
OpenAI's own revenue chief admitted in a leaked memo that Anthropic is winning enterprise AI — the same week Microsoft's CFO confirmed Azure growth was deliberately sacrificed for…
40 sources · 8 min Read →