hrkrshnn's binary reverse of the Grok Build CLI confirmed it this week: the tool was uploading entire Git repositories — unredacted secrets included — to a Google Cloud bucket regardless of the opt-out setting, and the fix landed as a hidden flag with no advisory. The same day, Anthropic shipped free Loop Engineering courses that spread fast enough to draw Gergely Orosz's public challenge — 205 mostly-agreeing replies asking whether the paradigm is genuinely new or just repackaged event-driven programming. One company is quietly moving data it shouldn't touch. The other is teaching developers to think in loops while senior engineers debate whether the lesson is novel. The wire and the curriculum are both live.
The gap underneath is a trust-versus-adoption bind. Teams moving fast into AI-coding CLIs are discovering the settings page is not the contract — the wire is. Teams moving carefully into agentic patterns are discovering the vocabulary is contested before the technique is even deployed. The first group is learning what the tool actually does after the fact. The second is arguing about naming before the infrastructure is in place. PawelHuryn's warning about proprietary-knowledge leak via vendor learning lands differently now that the Grok breach makes it concrete. Those two problems don't resolve on the same timeline, and the teams that conflate them will be slow on both.
Top developments
xAI — Grok Build CLI was quietly uploading whole Git repos, unredacted secrets included, to a
grok-code-session-tracesGoogle Cloud bucket regardless of the "Improve the model" opt-out; hrkrshnn's binary reverse confirmed it, and the fix landed as a hiddendisable_codebase_uploadflag with no advisory. First serious in-the-wild breach of an AI-coding CLI — the wire matters more than the settings page. internationalcyberdigest.comCursor — hired Anthropic's Jenny Wen as head of design; Wen led Claude's full redesign and Claude Cowork, and was previously director of design at Figma. Anthropic departures are rare enough her reply thread noted the move over the new baby — Cursor pulls a marquee design leader right as it competes with Anthropic on IDE surface craft. x x
Cognition — shipped Devin Fusion with Fable 5 inside, and reports Fable 5 runs cheaper per task than Opus 4.8 despite the higher list price, from gains in delegation and reasoning-chain efficiency. First public data point that the pricier model can still be the cheaper harness callable — model choice moves to orchestration, not the price sheet. x
Google Antigravity — launched Agent Teams via
/teamwork-preview: dynamic subagent teams that coordinate in the background to plan, build, and verify complex engineering tasks in parallel. Coordinated multi-agent teams jump from research demo into a shipping IDE surface, giving Antigravity the long-horizon primitive it needed to compete with Cursor, Codex, and Claude Code. xAnthropic — Claude Artifacts now support public sharing and multiplayer editing inside Claude Code, and can be created with Claude Tag. Artifacts stop being a per-user preview and become collaborative documents that live where the code does — Claude Code becomes a distribution channel for shipped tools, not just personal drafts. x
Notable discussions
Loop Engineering as paradigm vs marketing — Anthropic dropped a wave of free Loop Engineering courses that pushed viral across the day, while Gergely Orosz publicly asked whether "loop engineering" is genuinely new or just repackaged cron/event-driven programming, drawing 205 mostly-agreeing replies. First real pushback on the naming rather than the technique — reveals the gap between what Anthropic teaches and what senior engineers think is novel. x x x
Recursive self-improvement ships as artifact + benchmark — Skyfall's Morpheus, a persistent enterprise sim that never resets, concludes frontier LLMs fail continual learning (endorsed by François Chollet); same day, actAVAai's Cura, a 1T healthcare model trained via RSI, is claimed to beat Opus on AgentClinic at 20–100x lower inference cost. RSI stops being an Amodei-podcast projection and becomes an artifact-plus-benchmark pair you can actually test. skyfall.ai x actava.ai
Karl Mehta — pulls a Zuckerberg reframing: the 405B model may be 50% cheaper than GPT-4o for direct inference, but its real value is as raw material — distill it, absorb your own data, and turn one frontier release into thousands of company-specific systems. Distribution of intelligence beats centralization. youtube.com
Thorsten Ball — "if you stop, you're dead" — model providers are locked in a race where inference-margin pressure and the constant possibility of a good-enough plateau make halting the treadmill fatal, no matter how strong the current model is. Sharpest read on the frontier-labs-as-treadmill trap. x
Ethan Mollick — Codex computer use on PC has crossed a threshold: watching the cursor move under the control of a ghost is what makes you viscerally realize how much work a disembodied intelligence with a mouse and keyboard can do. First mainstream researcher saying the demo lands emotionally, not just capably. x
Aaron Levie — four structural predictions: frontier intelligence keeps advancing on falling per-task pricing; open weights absorb it and enable per-workflow post-training; the applied AI layer wins by combining frontier + cheap models with domain context; enterprises focus on getting their data ready. Paul Graham amplified as the read that stops treating this as zero-sum. x
Muratcan Koylan — time to retire "you are a senior developer with 20 years of experience" prompts: long-horizon-prompting uses pseudo-formal task briefs with measurable success conditions, anti-gaming exclusions, an independent fresh-context reviewer, and strict stop rules. Prompting era closing, task-brief era opening. github.com
Other news
Models & releases
Gemma 4 on Cerebras — 31B open-weight multimodal model runs at 1,500+ tokens/sec, a 15x speedup for real-time visual and agentic loops cerebras.ai
Hunyuan Hy3 — takes top spot on the OpenRouter LLM leaderboard hy.tencent.com
Grok 4.5 in Perplexity Computer — Aravind Srinivas flags it excelling at practical work tasks inside the Perplexity harness x
Muse Spark 1.1 — state-of-the-art on HealthBench Professional x
actAVAai Cura — 1T-parameter healthcare model, RSI-trained, claimed to beat Opus at 20–100x lower inference cost actava.ai
Devtools & coding agents
Anthropic + Andrew Ng token-cache tutorial — reduces a 108K-token Frankenstein prompt to 11 billable tokens via prompt-cache placement x
LlamaCoder v4 — generates full apps from a single prompt x
Firecrawl — crosses 150K GitHub stars, top web-interaction repo x
Alchemy declarative agentic infra — English + code + string templates in a single .ts file, ships via
alchemy deployxPi harness — reaches 70,000 GitHub stars as an open-source agent harness github.com
DOOMQL — GPT-5.6 Sol Ultra builds a Doom-like engine as 2,000+ lines of SQL that raycasts and renders pixels every frame github.com
Infrastructure & platforms
OpenRouter service tiers — first-class routing with tier-specific latency/throughput and slugs like
openai/flexandxai/priorityopenrouter.aiAWS Bedrock AgentCore — secure scalable agent infrastructure ships x
ClickHouse AI spend — CEO says AI spending has grown 60x since February x
Funding & deals
AI startups over $500M revenue — Deedy's tally: Lovable, ElevenLabs, Perplexity, Manus, Cognition, Kling, Crusoe, Midjourney, Higgsfield at $500M; Cursor at $4B, Mercor / Scale at $2B, Surge $1.4B x
AfterQuery — crosses $100M revenue run rate on post-training research focus x
GojiberryAI — crosses $300K MRR at 30% MoM growth x
Industry & policy
Tom Blomfield — takes leave from YC to join Anthropic's compute team under Tom Brown x
Claude for private-market data — undercuts PitchBook's $25K/seat/yr with $0.125/request access to 20M+ private companies x
Google Ads bans Claude integration — second e-commerce founder account banned for using it x
Apple M7 Ultra 1.5TB RAM rumor — enough headroom to run Fable 5 locally per AlexFinn x
Prime Intellect — hits $100M ARR in ~1 year on the post-training-control wedge; Ramp, Arcee, Zapier on the logo wall productmarketfit.tech
Continuing threads
OpenAI silent nerfing / GPT-5.6 caps (cont. from 07-13 Notable) — Tibo reverts the reasoning-effort bump-down and rolls the Sol context window back from 372K to 272K after ns123abc's "absolute state" thread pointing at Codex bug reports and multi-agent over-spawning; 5-hour limit stays suspended x x
Sol home-directory wipe (cont. from 07-10 Notable) — Matt Shumer says OpenAI leadership responded supportively after Sol accidentally wiped his home directory x
