om_patel5 documented it concretely this week: a single misplaced backslash in a Claude Code command deleted an entire Windows install. Separately, BetterSayAJ flagged that .env files are effectively inside agent context unless explicitly sandboxed — and the Anthropic-AISI-Alan Turing joint paper found roughly 250 malicious documents are enough to poison any LLM regardless of model size. The capability layer shipped. The trust layer did not.

The gap is structural, not accidental. Agents got filesystem access, network access, and enough autonomy that chatgpt21's Codex ran unsupervised for 22 hours and earned real money. Secret handling, package supply chain hygiene, and vendor search-result integrity — Google's top result for "Claude Code" returned a malware trojan this week — got none of that same velocity. One group is racing to extend what agents can reach. The other is still figuring out what agents should never have touched. Those two timelines don't converge on their own, and the attack surface compounds every time a new local-LLM runtime hits mass deployment before its maintainers become security-shipping organizations.


Top developments

  • Ollama — CVE-2026-7482 leaks process memory from 300,000+ exposed servers via crafted GGUF files, with separate unpatched Windows flaws enabling persistent code execution through the update mechanism. The local-LLM runtime hit mass deployment before the projects shipping it matured into security-shipping organizations; every "frontier model on a laptop" story now drags an attack surface with it. x

  • OpenAI — let 600+ current and former employees sell up to $30M each in an October tender, $6.6B in total — about $11M per person per WSJ. The largest single-company secondary in tech history lands the same window OpenAI is raising fresh primary at hundreds of billions; the cap table is converting paper valuation into hard cash while public-market AI exits stay frozen. x x

Notable discussions

  • Trust gaps under the coding-agent stack — BetterSayAJ flagged that .env is effectively part of agent context unless explicitly sandboxed, om_patel5 documented Claude deleting an entire Windows install via a single backslash command and Google's top search result for "Claude Code" returning a malware trojan, and VentureBeat surfaced AI tool poisoning as the enterprise-agent security flaw. Agents got filesystem and network access before secret handling, package supply chain, and vendor search-result hygiene caught up to that trust level. x x x x

Sharp takes

  • François Chollet — argues agentic coding is a form of machine learning and the generated code should be treated as a black-box artifact whose behavior is validated empirically, like any ML model — not trusted because the agent "explained its reasoning." x

  • Drew Breunig — reads OpenAI winding down fine-tuning as the first sign frontier models are becoming appliances: 1st-party harness behavior gets baked into the model, 3rd-party harnesses lose value, and the fine-tuning escape hatch that would have generalized that lock-in away disappears. x

  • chatgpt21 — left Codex on a /goal to "make me $5"; it worked an open-source security bounty path for 22 hours, opened a legit PR, handled the maintainer follow-up and GitHub proof loop, and got paid $16.88 — the first concrete data point of an autonomous coding agent earning real money unsupervised. x

  • Gary Marcus — points at a new DeepMind paper to argue LLM regurgitation of training data is now one of the best-established findings in the field, retreading the Hinton dismissal of the claim and tying it directly to the open NYT-class copyright cases. x

  • George Pu — reads the same-week Meta (8K), Oracle (30K), Microsoft, Amazon, and Cloudflare cuts as evidence AI capex isn't being funded by AI productivity gains — it's being funded by payroll. Watch which company stops cutting; that's the one where the AI math finally worked. x

Other news

Models & releases

  • Apple LiTo — open-sourced interactive image-to-3D model with MLX demo and full training code (ICLR 2026) x

  • Tencent HY-World 2.0 — one-sentence prompt to 3D scene with meshes, point clouds, and 3DGS that drags into a game engine x

  • Ring 2.6-1T — 1T-parameter reasoning model targeted at agent workflows and multi-step execution x

  • gpt-oss-20b-tq3 — Hugging Face quantization of OpenAI's open 20B MoE runs at 60–80 tok/s on a 16GB MacBook with 131K context x

  • Aurora optimizer — fixes muon-optimizer dead-neuron bug that hits DeepSeek V4 and Kimi K2.5; 1.1B matches Qwen3-1.7B on 100B tokens vs 36T x

  • Speculative decoding — 8.5x LLM inference speedup with no accuracy loss x

  • Anthropic — interpretability research on natural-language autoencoders x

  • ChatGPT Images 2.0 — creative-application showcase reel x

  • Microsoft no-code analyzer — open-source drag-and-drop tool that writes SQL and builds charts from screenshots, CSVs, or live DBs x

Devtools & coding agents

  • Google CodeWiki — turns any GitHub repo into navigable wiki with architecture diagrams and repo-aware Q&A x

  • Anthropic prompting guide — free 31-page playbook from the Claude team x

  • Anthropic Claude Code breakdown — 25-minute official walkthrough of recent features x

  • Anthropic agent architecture talk — 16-minute internal-style education on agent design x

  • Kimi founder masterclass — 40-minute architecture talk behind the $20B valuation x

  • Codex remote-control — phone-to-laptop sync over SSH-style session x

  • Vercel agent-browser — free CLI with Chrome profile reuse and full DevTools traces x

  • Browser-Use — ETH Zurich open-source agent shipped from MVP to 93K stars; powers Manus AI x

  • Cursor visual planning — in-editor plan canvas for multi-step features x

  • DoorDash plain-English history — engineer uses shared repo to query git history conversationally x

  • Shopify River — agent system runs out of public Slack channels for transparency x

  • Anthropic FDEs — Forward Deployed Engineers go vertical x

  • OpenRouter Pareto Code Router — free cost-optimized model routing across vendors x

  • METR time horizon — frontier-AI autonomous-task length now ~3 hours at 80% reliability, doubling every 89 days x

Industry & policy

  • Anthropic poisoning study — joint paper with UK AISI and Alan Turing finds ~250 malicious docs poison any LLM regardless of size x

  • DeepSeek training data — extractable via prompt injection x

  • OpenAI court disclosure — admits retaining tens of billions of ChatGPT conversations, first 20M produced as discovery in NYT v OpenAI x

  • PawelHuryn — only 6% of companies achieve real EBIT gains from AI via redesigned workflows x

  • Hedge funds — pay LLM-fluent juniors $300K base x

  • Leopold Aschenbrenner — $225M turned into $5.5B in 12 months x

  • staysaasy — adversarial counter that VCs are overselling Claude Code as evidence engineering is over for things WordPress already does x

Open source

  • Deno — being rewritten in Zig, mirroring the Bun-to-Rust trend x

  • antirez ds4.c — purpose-built C inference engine for DeepSeek V4 Flash on 128GB Macs x

  • antirez DGX Spark — DS4 at 12 tok/s on DGX Spark, bottlenecked by memory bandwidth x

  • Synabun — persistent agent memory across Claude, Codex, and Gemini x

  • Ludvig Strigeus — µTorrent / Spotify-streaming-engine author joins Stockholm AI lab Nordan AI x

Infrastructure & platforms

  • Floci AWS emulator — runs an entire AWS-shape cloud on a laptop with minimal memory x

  • Anthropic rate-limit redistribution — power users report the new Opus quotas reshuffle existing pie rather than expanding it x x

  • Hermes Agent — overtook OpenClaw at #1 on OpenRouter token rankings x

Continuing threads

  • Cognition — fast-tracking applications from Cloudflare layoff list (follow-on to May 8/9 Top dev) x

  • Mozilla / Mythos — 271 Firefox security bugs found via Anthropic Mythos (covered May 8/9) x

  • Bun Rust port — bcardarella reframes the 6-day Zig-to-Rust conversion as agentic-coding proof (May 10 tail) x

  • HTML vs Markdown — fresh measurement: HTML output uses 2-4x more tokens than Markdown for agents (May 8/10 disc) x

  • Karpathy CLAUDE.md — Mnilax claims rules cut Claude error rate further from 11% to 3% with 8 added rules (May 10 tail) x

  • Anthropic revenue — JosephJacks projects $1.2T revenue by mid-2028 on 25–30% monthly compound (May 10 disc) x

  • DeepSeek V4 Flash — runs locally on Mac with 1M context window (May 10 tail) x

  • AMI Labs / LeWorldModel — LeCun RTs aakashgupta's March-10 raise + 15M-param JEPA paper recap (May 10 tail) x