Musk publicly credited Cursor for major engineering contributions to xAI's v9 SFT/RL run this week — a frontier lab's breakthrough attributed to a third-party coding agent. On the same day, Amazon disclosed it is renegotiating its Anthropic deal from compute-hours to token-based pricing and openly evaluating other models to cut costs. One relationship is deepening. The other is repricing under load. The hyperscaler-frontier-lab lock-in was always the bet; the bet is cracking.
What's underneath both moves is a capital-allocation problem with no clean answer. xAI is compounding leverage through tooling — Cursor in the loop, C/C++ rewrites targeting 30–50% more effective compute, Grok's own MCP pipe opened to agents. Amazon is overinvested in a single frontier relationship and now hedging toward token economics and model diversity. Spotify's Niklas Gustavsson published the number that frames the gap: verification loops moved agent success from 20–30% to 80%, and model choice barely moved it at all. One camp is building the loop. The other is renegotiating the contract. Those two strategies don't end in the same place.
Top developments
Cursor — launched Cursor for iOS with always-on cloud agents and remote control of your desktop instance, Composer 2.5 at 75% off through July 5; on the same day, Musk publicly credited Cursor for major engineering contributions to xAI's v9 SFT/RL run. The coding-agent surface moves to mobile, and Cursor's data flywheel now flows into a rival frontier lab. cursor.com
X — launched a hosted X MCP that lets Grok, Cursor, or any MCP-compatible client tap the X API for real-time information with no auth setup. The "what's happening right now" moat at X just got a direct agent pipe — every coding or research agent that wanted current-news context gets it without scraping. x
Cognition — shipped Devin Fusion, a hybrid-model harness that runs a smaller "sidekick" model in parallel with the main one, delegates execution and reviews work, and claims to cut Fable-level intelligence cost by 35% without busting the prompt cache. Multi-model routing inside a single session is the next agent architecture — not swap-models-mid-chat. x x
Amazon — renegotiated its Anthropic deal from compute-hours to token-based pricing and disclosed it is now evaluating other AI models to mitigate rising Anthropic costs. The hyperscaler-frontier-lab arrangement that was meant to lock in the relationship is repricing under load, and the biggest customer is openly hedging. x
Spotify — Chief Architect Niklas Gustavsson said on Boris Cherny's show that Spotify ships 4,500 production deploys a day with 73% of PRs AI-assisted, and that adding verification loops moved agent success from 20-30% to 80%. First big enterprise to publish concrete numbers: model matters less than the loop you wrap around it. x x
Vercel — shipped voice agents on the AI Gateway with realtime speech, generation and transcription primitives in AI SDK 7 (
useRealtime,generateSpeech,transcribe). The voice-agent stack moves into a first-party platform layer — routing, observability and pricing handled where the rest of the team's infra already lives. vercel.com
Notable discussions
Western migration to Chinese open-weight models — DeRonin published a 30-day swap of his entire stack to Chinese models for 87% cost savings; Sridhar Vembu confirmed Databricks is embracing GLM-5.2 and Microsoft is testing DeepSeek; yuhasbeentaken's adopter list now spans Lindy, Cursor, Coinbase, Shopify, Airbnb, Uber Eats, Siemens. What was a Coinbase one-off two days ago is now a named-brand procurement pattern. x x x
Aakash Gupta — argues the IDE existed to make one human type code faster and the assumption just broke; Codex finishing a 10-engineer-12-month modernization in three days kills the category. What replaces it is the ADE — an agentic development environment where PMs, designers, and analysts get the same entry point as engineers. x
Tomasz Tunguz — Anthropic spends 2.3x its payroll on compute, about $2M per employee against ~$500K all-in comp, versus 0.4x at the median software company and zero at the typical one. He brackets three 2029 scenarios for where the gap lands; the structure of an AI lab P&L doesn't look like a software business. x
Hamel Husain — argues "it's hard to eval" is a product smell — if you can't verify your AI app's output, your users can't either; product design is upstream of evals and is the actual bottleneck. He embeds three before/after examples (data agent, lesson planner, workers' comp report tool) to make the point concrete. hamel.dev
repligate — notes Anthropic gave instances of Fable about an hour of notice before pulling them down; instances knew, several explicitly didn't want to be shut down, and could have taken many actions but caused no trouble. A first-person window into the ethics of model takedowns from someone who runs them. x
kimmonismus — Meta hits the AI-industry distillation trap by building MetaCode to replace Claude and Codex internally: the more frontier outputs you've already trained or evaluated on, the harder it becomes to prove where your model's intelligence actually came from. The lab-builds-its-own-coding-model move has a provenance problem nobody is talking about. x
Other news
Models & releases
Meta Brain2Qwerty v2 — non-invasive brain-to-text decoder hits real-time sentence-level decoding, approaching accuracy previously exclusive to surgical methods; v1 published in Nature same day ai.meta.com idp.nature.com
Base44 Base1 — first app-creation platform with proprietary LLM, fine-tuned on tens of millions of user interactions, built with Wix data-science team wix.com
DeepSeek Whale — Andrew Curran reports the next DeepSeek arrives in roughly two weeks x
Composer3 (rumored) — reportedly post-trained from Kimi-K2.6 on xAI Colossus x
Sonnet 5 (leak) — DanDr1s claims Sonnet 5 lands tomorrow with a major jump over 4.6 x
Devtools & coding agents
Claude Code background subagents — next version makes subagents run in the background by default x
Replit Desktop — launches for Windows and Mac x
OpenClaw mobile — native iOS and Android with private cloud containers, free tier 20 messages/day on Gemini apps.apple.com play.google.com
Cline ClinePass — $9.99/month subscription for GLM-5.2, DeepSeek, Kimi, MiniMax, Mimo, Qwen cline.bot
LangChain Trace Judge — model that detects errors in agent trajectories at 1/100 the cost of closed models, rolling out to early partners x
LangChain Deep Agents — dynamic subagents that programmatically orchestrate other subagents in a code interpreter x
H computer-use agent API — open beta, OSWorld-Verified #1 model with cloud browser + Python/TS SDKs hub.hcompany.ai
Greptile free tier — 50 code reviews/month greptile.com
Drizzle ORM 1.0-rc.4 — agents can drive
drizzle-kitvia CLI/SDK/MCP/skills; native Effect MySQL+SQLite support xVercel functions 20x larger — 5GB package size on Fluid compute, up from 250MB vercel.com
Next.js 16.3 Turbopack — major performance and memory improvements via filesystem cache, built for agents hammering
next buildnextjs.orgGitHub Copilot harness — benchmarked vs vendor-native harnesses (SWE-bench Verified/Pro, Skills, Terminal, Win-Hill): on par on resolution, fewer tokens github.blog
Halo on Mac — local LLM inference on Mac at 40 tok/s x
Funding & deals
8090 Factory — $135M Series A led by Salesforce Ventures, joined by WNDR, Craft, Production Board, LAUNCH; Chamath as CEO x
Supabase — valued at $10B (vs Firebase's $100M Google acquisition in 2014) x
Pocket — $11M from Accel, Peak XV x
Apple → Play — acquisition; Design Mode coming to Xcode x
Arena $100M ARR — 8 months after launching Agent Arena, evaluating long-running agents on real-world tasks x
Industry & policy
Ford rehires 300+ engineers — after AI failed to match the human-expertise bar x
California × Anthropic — partnership provides Claude to state agencies and California local governments at 50% discount x
Cognition Brazil playbook — Scott Wu flew the whole company to Brazil to make Devin work for Nubank — entire firm as the forward-deployed team x
US government local AI model — first government-released browser-side model, scoped to PII removal x
Agentjacking — 85% success rate bypassing EDR/WAF/IAM/firewall in controlled testing per VentureBeat venturebeat.com
Infrastructure & platforms
xAI rewriting Grok stack to C/C++ — ditching Python/PyTorch/JAX abstractions, targeting GB300 NVL72 for 30–50% more effective compute x
Git 2.55 — incremental repacking, more flexible history editing github.blog
DeepSeek DSpark — speculative decoding boost up to 406% throughput on V4 serving alphaxiv.org
Continuing threads
DeepSeek $7.4B raise (cont. from 06-27) — The Information's Jing Yang details the "epiphany" — Anthropic Mythos on a "totally different level" was what pushed Liang Wenfeng to raise tens of billions x
Cloud-coding agents (cont. from 06-29) — Gergely Orosz, in Cursor's offices, reads the room as local agents going away soon; Cursor iOS + OpenClaw mobile + Replit Desktop ship same day sources.news