PostHog co-CEO James Hawkins said it plainly at Startup School Paris: his AI systems already generate a portion of the company's own pull requests, with session replays detecting friction and triggering agents to open fixes autonomously. The same week, Amazon confirmed it cut an undisclosed number of roles across its AGI unit — including most of the MoE pre-training team — while Alphabet posted its first negative quarterly free cash flow in 22 years as AI CapEx surged past $200B. Agents shipping code. Headcount getting cut. The gap between those two facts is where the real bet is being placed.
The companies building toward Hawkins's model — PostHog, Cognition rolling up Poke, Sierra absorbing TakeOff's long-horizon runtime — are compounding agent capability into the product loop itself. The companies cutting first are treating AI as a cost lever, not a leverage multiplier. The first group ends up with software that finds its own bugs and ships its own fixes. The second group ends up with fewer engineers and the same product velocity. Those two strategies don't end in the same place.
Top developments
OpenAI launches GPT-Live, a full-duplex voice model now powering ChatGPT Voice on the desktop app — GPT-Live is a full-duplex voice model (it listens and speaks simultaneously) that can delegate complex tasks to a frontier model in the background while keeping the conversation going. The desktop integration lets users direct ChatGPT Work and Codex agents entirely by voice — positioning spoken commands, not keyboard input, as the primary interface for multi-agent coordination. OpenAI is rolling out globally today to Plus, Pro, Business, Edu, and Enterprise plans on macOS and Windows.
Cognition acquires The Interaction Company, makers of personal-agent app Poke — Cognition — the company behind Devin, the AI software-engineering agent — is folding in Poke, a proactive personal agent that lives in iMessage and WhatsApp, has racked up 100 million messages in three months, and is the only AI agent approved to text natively on Apple Messages. The move pairs Devin's enterprise engineering muscle with a consumer-facing "always-on" companion, and marks at least the second Cognition deal in a week (following Devin Outposts and the Jared Palmer hire), making the roll-up of specialized agent teams into one vendor hard to miss.
Sierra acquires TakeOff, the long-horizon agent startup that hit near-8-figure ARR in under a year — The deal pairs Sierra's enterprise AI platform with TakeOff's long-horizon agent runtime — together they're launching a new product called Horizon. TakeOff went from $0 to nearly 8-figure ARR this year across several seven-figure contracts, making this one of the faster revenue ramp-ups in the current agent wave. Two agent-company acquisitions in a single day (alongside Cognition's acquisition of Interaction) signals that the M&A layer of the agent stack is consolidating quickly.
Etched raises $300M Series C at $10.3B valuation to scale custom AI inference silicon — The round, led by Sequoia with Andreessen Horowitz, Jane Street, Argo, and SK Hynix, funds production of full inference clusters — co-designed chips, interconnects, and cooling — rather than just chips sold to hyperscalers. Etched has already opened an 80,000-sqft, 10-MW lab near its HQ and kicked off fabrication of hundreds of millions of dollars worth of hardware, betting that winning on tokens-per-watt at scale is a dedicated-hardware problem, not a cloud one.
Kimi K3 agent claims 19 zero-days in Redis 8.8.0 in 90 minutes — but questions surround version, novelty, and disclosure — A researcher posted a PoC repo showing Moonshot AI's Kimi K3 model running 32 agents supposedly discovered 19 RCE vulnerabilities in Redis 8.8.0 in under two hours, then separately produced a zero-click arbitrary-command-execution chain targeting Telegram Desktop and iOS. Significant caveats apply: Redis 8.8.0 does not appear to be a publicly released version; at least one of the CVEs cited (e.g. CVE-2026-25589 in RedisBloom) was already published and patched months earlier; security researchers note that fuzzer-style tooling routinely surfaces 20 raw findings that triage to far fewer confirmed bugs; and the findings were published without prior disclosure to Redis maintainers.
Videos worth watching
Together AI VP shows how to train a 3B model on 3M-token context in 16 minutes on 8×H100s — Together AI VP of Research Max Ryabinin walks through the stack of techniques — parameter sharding (FSDP), sequence parallelism (DeepSpeed Ulysses), activation recomputation, and KV sharding — that turns an out-of-memory baseline into a sub-16-minute training run on hardware most teams already rent, making 3–5M-token context training accessible without expensive dedicated bootcamps.
Cursor engineer David Gomes replaced a 12,000-line feature with a 200-line skill in 19 minutes — In a talk at the AI Engineer summit, Cursor engineer David Gomes showed how the WorkTrees parallel-branch feature — originally ~12,000 lines of TypeScript — was rewritten as a ~200-line Markdown "skill" file. The pattern: strip legacy code down to its logic, encode it as a Cursor Agent Skill (a SKILL.md file that packages workflows, domain knowledge, and commands), then let subagents execute it in parallel — leaving 60× less code to maintain.
PostHog co-CEO James Hawkins on building "self-driving" software that ships its own fixes — Speaking at Startup School Paris, PostHog co-CEO James Hawkins explains how PostHog's AI systems already generate a portion of the company's own pull requests — with session replays detecting user friction and triggering agents to open fixes — and outlines the broader vision of software that autonomously identifies product problems and ships code without human prompting.
Announcements & releases
Black Forest Labs launches FLUX 3, a unified multimodal model for image, video, audio, and action prediction — FLUX 3 is trained jointly across images, video, and audio in a single architecture — the idea being that learning all modalities together builds a richer world model than training on any one alone. The video generation capability is now in early access, and the same architecture is being extended to action prediction for robotics (demonstrated with partners Mimic and Audi).
Alibaba releases Qwen-Audio-3.0-TTS with 16-language voice cloning and natural-language style control — The Plus variant ranks #1 on the independent Artificial Analysis TTS leaderboard. Both variants support inline emotion tags ([whisper], [angry], [laughs]) and free-form instructions ("read this like a bedtime story") for fine-grained delivery control — particularly useful for voice agents and audiobooks. The model is available via API only; weights have not been released.
Baseten launches GLM-5.2 Fast, a speed-optimized API tier with 2–3x higher throughput for real-time workloads — Baseten is introducing a separate "Fast" serving tier for Z.AI's GLM-5.2 — same model weights, but infrastructure tuned for per-user throughput. The pitch is agentic pipelines where multiple inference calls compound latency: GLM-5.2 Fast is positioned as smart enough to act as the main orchestrating agent while also being fast and cheap enough to run across subagents. Switch by pointing your OpenAI-compatible client at the model ID
zai-org/GLM-5.2-Fast.Bad Theory Labs releases BTL-3, a 27B open-weight model fine-tuned for agentic coding — BTL-3 is a fine-tune of Qwen3.6-27B, quantized to just 8.39GB (under 2.5 bits per parameter — smaller than an 8B model in fp16) while retaining 92.2% of the full 27B's capability. It's explicitly trained for the agent loop: reason, act, inspect result, recover, and continue — with support for single, sequential, and parallel tool calls. A companion BTL-3-Compact edition and a native llama.cpp runtime are also available.
Alibaba Cloud's Model Studio Token Plan debuts Qwen3.8-Max-Preview, starting at $6/month — The new Model Studio Token Plan for Individual gives developers a single subscription covering Qwen3.8-Max-Preview — a 2.4-trillion-parameter multimodal model — plus every text, video, image, and audio model on the platform, at up to 3× the usage vs. pay-as-you-go. The plan is early-bird priced at $6/month, and off-peak discounts stack on top.
DecBench tracks how close decompilers — and LLMs — are to perfect binary reconstruction — Binary decompilation — recovering original source code from compiled executables — is nearing a tipping point, and LLMs are now competitive with purpose-built tools like IDA Pro and Ghidra. DecBench, from University of Georgia assistant professor Zion Basque's Noelo Lab, scores decompilers across ~95k functions and 800+ binaries on three exact-match metrics; a dedicated sample-set leaderboard lets AI models like Codex and Claude Code compete head-to-head against traditional decompilers. The project is open-source.
Claude voice mode expanded with Opus/Sonnet support, mid-conversation tool access, and multilingual conversations — Anthropic's voice mode previously ran only on the lightweight Haiku model; it now carries over whichever model you were using in text chat — including the more capable Opus and Sonnet — and can reach connected tools like Gmail, Google Calendar, Google Docs, and Slack mid-conversation. Multi-language support also broadens who can use it hands-free.
Canvas UI launches as the first HTML-in-canvas component library for React, Vue, Svelte, and vanilla TypeScript — HTML-in-canvas is a new browser primitive that lets WebGL shaders read and redraw your real DOM in real time — text stays selectable, links stay clickable — and Canvas UI is the first full component library built on it, shipping 24 ready-made effects (Blaze, Liquid Glass, Shatter, Particle Reveal, VHS, and more). It's free and open-source, using a shadcn-style copy-paste install so the source lands directly in your own repo. Frontend engineer David Haz built it; the underlying browser feature currently requires a Chrome flag, so it won't be visible to all visitors yet.
HubSpot launches Agent Hub and Agent Builder in public beta — Agent Hub gives teams one place to build, deploy, and manage AI agents with shared CRM context; Agent Builder is a no-code/low-code tool for constructing custom agents directly inside HubSpot. Both products are now in public beta, marking HubSpot's push to become an AI-agent platform on top of its CRM data.
OpenWorker is an open-source agent that delivers finished work instead of just chatting — DeepLearning.AI founder Andrew Yan-Tak Ng announced the launch, describing an agent that works across your files, Slack, and calendar to produce actual deliverables — polished documents, sent messages, updated calendar entries — and checks in before taking any consequential action. It's local-first, privacy-preserving, and model-agnostic (bring your own API keys); a Mac app is available now with Windows coming soon. The source is on GitHub.
Anthropic launches an official prompt library for Claude Code — The Claude Code prompt library offers copy-paste-ready templates drawn from Anthropic's own internal guides and workflows — covering common dev tasks tagged by role and phase — so developers can skip reverse-engineering prompts from scratch and instead start from the same structured reasoning and tool-use patterns Anthropic uses internally.
Moonshot AI open-sources Kimi Code CLI, a free terminal coding agent with isolated sub-agents and screen-recording input — Kimi Code CLI is a free, Apache-licensed terminal agent (10.7k GitHub stars) from the team behind Kimi K2 that runs separate, context-isolated sub-agents for coding, planning, and research so large tasks don't bleed together. It adds at least one trick Claude Code lacks: you can drop a screen recording directly as input, useful for bug reports that are hard to describe in text. A plan mode previews proposed changes before any file is touched, and MCP servers can be configured per conversation.
TypeScript 7.0 ships with Go-powered compiler delivering 8–12x faster builds — Microsoft's native Go port of the TypeScript compiler cuts full build times by 8–12x and slashes VS Code project-load time from ~60 seconds to 10 seconds, with major memory reductions — a significant win for large monorepos and any TypeScript-heavy codebase.
OpenFPM PR adds Metal GPU backend for Apple Silicon via CUDA→SPIR-V→Metal translation chain — A pending pull request for OpenFPM 5.2.0 adds a Metal GPU backend that routes CUDA/HIP-style kernels through clspv and MoltenVK (SPIR-V) to run on Apple Silicon GPUs, showing roughly 10× speedup over CPU on a fluid-simulation workload. The scope is narrower than viral posts claimed: scientific computing researcher Abhinav Singh, who authored the PR, clarified that the translation targets OpenFPM's own CUDA/HIP-style kernels specifically — it is not a general "run any CUDA code on a Mac" solution. Apple's unified-memory, tile-based GPU architecture also means a native Metal port would still outperform this translation layer for production workloads.
Flue now supports agentOS — WebAssembly isolates cut sandbox memory 48x vs. full VMs — Rivet's agentOS applies the Cloudflare Workers isolate model to agent sandboxes: instead of booting a ~1 GiB VM per agent, each sandbox runs in a WebAssembly + V8 isolate at ~22 MB RAM with 4.8 ms cold starts, while still exposing a Linux-compatible environment (bash, git, Node, Python, filesystem, processes). Flue — the open-source TypeScript agent framework from the creators of Astro — can now use agentOS as its sandbox backend, meaning agents sitting idle between inference calls stop burning reserved VM capacity.
AngelList launches Link, an MCP server letting GPs query their fund data via AI — AngelList Link connects fund administration data — capital call status, LP holdings, valuations, daily financials — to Claude, ChatGPT, Cursor, or any MCP-compatible AI tool, so GPs can get verified answers instantly instead of waiting on email chains to administrators. MCP (Model Context Protocol) is an open standard that lets AI models securely pull live data from external systems.
Alphabet posts its first negative quarterly free cash flow in 22 years, at -$5.9B, as AI CapEx surges past $200B — Alphabet's Q2 2026 earnings release shows $119.8B in revenue (up 24% YoY, with Google Cloud growing 82%), but record AI infrastructure spending tipped quarterly free cash flow into the red for the first time since the company went public — a milestone the company's CFO attributed to surging CapEx, even as trailing-twelve-month FCF remained positive at $53.3B and the company held $242.5B in cash and securities. The $200B+ annual spend is split roughly 60% on infrastructure and 40% on data centers, with observers debating whether the investment signals strategic urgency or simply the cost of competing at hyperscaler scale.
Amazon cuts jobs in its AGI organization, MoE pre-training team among those hit — Amazon confirmed it eliminated an undisclosed number of roles across its AGI unit, saying it is "sharpening focus on initiatives that matter most for customers." Applied Scientist Yuxin Tang said most of her MoE (Mixture-of-Experts) pre-training team was impacted and is now open to new roles — the kind of foundational model-building work that rivals are racing to staff up.
OpenAI launches Health in ChatGPT with Apple Health and medical-records integration for U.S. users — ChatGPT can now securely connect to Apple Health and supported medical-record portals, letting it compare lab results over time, summarize changes since your last appointment, and contextualize sleep and activity data — without using that data to train its models. More than 300 million people already bring health questions to ChatGPT each week; this gives those conversations persistent, personalized context. Currently U.S.-only; no EU rollout date announced.
Harvey Expands Collaboration with Microsoft on Legal AI — Legal AI platform Harvey will be deployed across Microsoft's Corporate, External, and Legal Affairs (CELA) organization — one of the largest in-house legal teams in the world — covering its legal and compliance operations. It's a high-profile validation for purpose-built legal AI at enterprise scale.
Cosine launches air-gapped AI coding agents for regulated industries, powered by Nebius AI Cloud — Cosine's Lumen agent runs entirely inside a customer's security perimeter — no code, prompts, or model outputs leave the boundary — targeting defence, finance, and critical infrastructure operators whose compliance rules forbid sending source code to external cloud services. Nebius provides the underlying GPU cloud infrastructure.
Discussions & takes
Notion launches "Notion as code" beta, letting developers define entire workspaces in TypeScript — Part of Notion's new Developer Platform, the feature lets teams declare teamspaces, databases, and custom agents in TypeScript, deploy them via the API, and version-control the whole setup in git — making a Notion workspace a reproducible build artifact that coding agents can spin up or tear down like any cloud service.
Anthropic publishes html-effectiveness, a set of HTML agent-output examples — sparking debate over whether .md planning is holding agents back — Anthropic's html-effectiveness repo collects 18 interactive HTML artifacts — wireframes, code reviews, slide decks, flowcharts, and more — showing what AI agents can produce when outputting
.htmlinstead of Markdown. EasyCart co-founder Adam Gospodarczyk argues that for UI/UX tasks agents can zero-shot complex interfaces as.htmlfiles, giving both human and AI a visual reference the agent can screenshot for feedback. Daniel Ospina of RnDAO extends the critique, calling Markdown "a horrible way to structure memory" and blaming it for the cut-corners behavior common in agentic workflows. Pushback in the thread points out that not every task is visual, and that Mermaid diagrams inside.mdfiles offer similar benefits with lower token cost.CleanCoders' Agentic Discipline Episode 4 distills Uncle Bob's framework for AI-agent coding — don't read the code, test everything — The Clean Code author Robert Cecil Martin argues the only way to capture AI's productivity gains is to stop reading agent-generated code and instead surround it with extreme constraints: unit tests, Gherkin (natural-language behavior) tests, mutation testing, coverage metrics, and cyclomatic complexity checks. The post crossed 10K likes and sparked sharp debate about whether "trust the harness, not the code" is responsible engineering or an abdication of it — notable precisely because it comes from the TDD community's most prominent voice.
Microsoft launches hill-climbing MAI models for GitHub Copilot and Excel, claiming GPT-5.6-level quality at lower cost — with CEO Satya Nadella framing the strategy as train small task-specific models first, call frontier APIs only when needed — Microsoft's Superintelligence team deployed two specialized MAI models built on its "hill-climbing" training pipeline — one for GitHub Copilot code tasks and one for Excel spreadsheet work — claiming they match or beat GPT-5.6 and GPT-5.4 Mini quality respectively while running on older H100/A100 hardware at a fraction of the cost. Microsoft CEO Satya Nadella frames the broader strategy as continuously hill-climbing small, task-specific models trained against real user behavior inside your own products, routing to large frontier models only for tasks that genuinely require them.
Rippling CEO Parker Conrad: If Anthropic can't prevent distillation, the national-security case for banning Chinese AI models falls apart — After Anthropic's February blog post revealing that DeepSeek, Moonshot, and MiniMax ran industrial-scale campaigns to distill Claude's capabilities — and calling for government intervention to stop it — Conrad flips the argument: if distillation is truly impossible to prevent, then any frontier-capability lead is short-lived, protectionist bans on Chinese open-weight models lose their strategic logic, and the asymmetry runs both ways: "if China pulls ahead we can distill them."
ThePrimeagen flips on AI dev tools, names structural refactoring as the killer use case — Former Netflix engineer and developer educator Michael "ThePrimeagen" Paulson — a long-time skeptic — says modern models finally changed his mind: the sweet spot isn't code generation but code exploration. You spin up many different refactor styles in parallel, compare them, and pick the best one. Trying three different architectures used to cost days; now it costs minutes. He notes Claude (Sol) and Grok are his current go-tos, and he explicitly frames this as deliberate, read-the-code usage — not vibe-coding on autopilot.
Cheap open-source model + frontier advisor beats pure-frontier on SWE-Bench cost and completion — David Zhang (Aomni CEO) tested a "duet" pattern on a 50-task SWE-Bench Multilingual subset: pairing a cheap OS model as the primary executor with a frontier model in advisory-only mode consistently beat running a pure frontier agent. Kimi+Fable-advisors topped pure Fable on both task-completion rate and cost-per-task; GLM-5.2+Kimi also beat pure Fable while posting some of the lowest cost-per-task figures. The pattern undercuts the false choice between "pay for frontier" and "trust open-source alone."
OpenAI Codex local projects now support multiple folders with a single Git root — Codex projects can now span several directories at once — useful for monorepos or setups where code, docs, and reference files live in separate folders. One designated primary folder handles all Git operations (commits, pull requests, AGENTS.md, skill discovery) while Codex can read and write freely across every included folder.
Together AI benchmarks Kimi K3 Max vs. GPT 5.6 Sol Max on DeepSWE: matched performance at 55% of the price, with a ~16% lift when routing between them — Together AI's Zain Hasan ran a head-to-head on DeepSWE (a software-engineering agent benchmark) and found Kimi K3 Max matches GPT 5.6 Sol Max in pass@1 while costing roughly half as much. The bigger takeaway: the two models have complementary strengths — Sol wins on first tries while K3 pulls ahead with retries — so cascading/routing requests between them yields a ~16% accuracy boost over either model alone.
Continuing threads
Kimi K3 ties GPT-5.6 Sol in blind designer study, rebuilds Google Maps 3D in two prompts — Three independent real-world tests reinforce K3's #1 Frontend Code Arena ranking: a blind evaluation by 8 working designers across 10+ landing page briefs put K3's win rate (63.3%) within two votes of GPT-5.6 Sol (65.0%) and well ahead of Claude Fable 5 (32.9%); a single prompt produced a fully playable Animal Crossing–style game; and a two-prompt, 1.5-hour session rebuilt Google Maps 3D with real building extrusion, sun-angle shadows, and free camera rotation over New York City.
TickerTrends tracker claims Anthropic hit $74.1B ARR vs. OpenAI's $41.3B — but the data is disputed — TickerTrends, a pre-purchase intent tracking platform (not official company disclosures), shows Anthropic flipping from roughly half of OpenAI's run rate in January ($10.2B vs. $21.4B) to nearly double it today ($74.1B vs. $41.3B). Multiple observers in the thread flag the methodology as unreliable for enterprise-heavy companies — both firms sell heavily via API and enterprise contracts that don't surface in consumer spend signals. The last confirmed official figure from Anthropic CEO Dario Amodei put ARR at $30B in May 2026, making the $74B tracker estimate look aggressive.
Anthropic employee's screenshot accidentally confirms Opus 5 exists; separate unverified claim says it was cancelled over human-control fears — An Anthropic employee posted a screenshot showing they were routed to an "Opus 5" fallback after hitting content guardrails in Fable 5 (Anthropic's current frontier model tier), inadvertently confirming internal development of the next flagship. Separately, an unverified account claims an insider says Anthropic shelved Opus 5 at the last minute after internal tests raised fears it could threaten human control — a dramatic but uncorroborated rumor that most respondents greeted with skepticism.
Graph-engineering playbook for multi-agent Claude systems: a 5-stage knowledge-graph pipeline to give agents persistent memory — Agent memory normally dies when the context window closes; this 12-page guide proposes a graph-engineering pipeline — Extract → Resolve → Assemble → Query → Repeat — that uses Claude Haiku to pull subject-predicate-object triples from documents and stores them in a knowledge graph, so facts survive across sessions. The document is a community playbook, not an official Anthropic publication (the filename misspells "Anthropic"), but the pattern it describes — Pydantic-constrained extraction, entity-resolution clustering, and graph-grounded queries — is a practical architecture that practitioners are already debating, including pushback on how entity resolution degrades at scale with messy real-world documents.
Hugging Face CEO credits open-weight GLM-5.2 as key defense tool in unprecedented autonomous AI agent breach — Following the Security incident disclosure — July 2026, Hugging Face co-founder and CEO Clément Delangue is publicly praising both his security team and Z.ai's open-weight GLM-5.2 model. The twist: commercial US frontier models refused to assist with log analysis during incident response because their safety guardrails flagged the forensic queries — forcing the team to fall back on the unrestricted open-weight model to dissect an attack that logged over 17,000 autonomous agent actions across internal clusters. The episode has become a flashpoint in the open-source-vs.-restrictions debate: if GLM-5.2 had not been freely available, defenders say the attacker would have had a decisive speed advantage.
Funding & deals
Paper raises $34M Series A from Accel and ICONIQ to build the design platform for the agentic era — Paper positions itself as the pro design tool for teams shipping with AI coding agents — a direct response to the surge in agent-driven development. After a 25× ARR spike in a single month following its Desktop launch, the company has now raised $38.5M in total, with customers including Ramp, Lovable, and Harvey.
Splash Industries raises $4.2M for autonomous drone boats that resupplied US Navy warships at RIMPAC — The California startup's 11-foot Typhoon USV costs $30k (versus $300–600k for comparable vessels), assembles in 8 hours, and just completed the first-ever autonomous underway resupply of a crewed warship — landing inside USS Essex's well deck during RIMPAC 2026. The $4.2M seed round, led by Ubiquity Ventures, also funds a larger follow-on vessel, the Tempest, designed to carry 1,500 lbs over 300 nautical miles.
Stockholm startup Scape launches an AI-native email inbox with $3.2M in seed funding (paywalled) — The desktop client connects to your email, calendar, and meeting transcripts to proactively prepare draft responses and surface only urgent messages — aiming to replace Gmail entirely. Scape CEO Melvin Hagberg announced the raise was backed by Y Combinator, General Catalyst, and FundersClub, with angels from OpenAI, Google, Meta, Ramp, Lovable, and Legora; early access is open now at scape.app.
