Sponsored by |  |
|
|
Good evening. It’s Monday, August 31st. |
Apologies for the delayed edition! This is the latest newsletter I’ve sent in the last 544 emails. | According to Anthropic, attackers were caught "using common infostealer malware to steal Claude login sessions from people's computers" and burning victims' usage. | Anthropic forced sign-outs, stripped saved payment methods, refunded charges. | Grand theft token. | -Jeff AI Breakfast |
|
|
You read. We listen. Let us know what you think by replying to this email. |
|
Find the Perfect Voice for Your Brand's AI | | Consumers can tell the difference. 79% say AI voices should come from real, attributed actors, and 61% say a voice is most memorable when it's exclusively tied to one brand. | Voices is the only platform that designs, licenses, and captures your brand's AI voice from real, consenting professional talent—never from scraped data. From strategic voice casting and brand calibration to consent-based rights management and studio-grade capture, we handle the entire process end to end. | BMW, SuperBloom, and Cresta already trust Voices to build their signature AI voice. Get a demo and see how a licensed, brand-owned voice becomes a real asset—not a legal liability. | Get a Demo |
|
|
|
OpenAI cuts Cursor access following SpaceX buyout |
Following SpaceX’s massive $60 billion acquisition of Anysphere, OpenAI triggered a change-of-control clause to pull its models from Cursor by November 12. OpenAI’s public reasoning points directly to a history of contract disputes with Elon Musk’s ventures. |
Cursor co-founder Michael Truell countered that OpenAI models only handle roughly 5% of their user traffic anyway as they lean into Grok 4.6 and open-weight powerhouses like GLM-5.3 and Tencent’s Hy4-preview. |
On the platform end, OpenAI is sweetening its offer to developers by resetting usage limits and fixing architectural bugs, like context compaction bloat, runaway subagent calls, and background memory leaks, stretching Codex and ChatGPT Work allowances between 10% and 50% further.
This infrastructure tussle arrives as Sam Altman claims OpenAI could reach internal AGI by the end of 2026. That timeline rests squarely on Astra, an upcoming model designed to run continuously for weeks and automate core tasks for research engineers. They are now internally testing Astra under the codename "ultima-alpha" alongside GPT-Image 2 ahead of a rumored September rollout, while simultaneously lobbying California lawmakers for mandatory monitoring and strict safety standards on frontier model training. |
|
Claude autonomously patches alignment flaws over 48 hours |
Anthropic is turning Claude into a self-correcting, hardware-controlling engine, but legal and operational landmines are catching up fast. |
In a landmark safety study, Claude ran autonomously for 48 hours on a single GPU to fix alignment flaws in smaller models. Scaling that setup, Claude Sonnet 5 aligned an early Opus 4.8 checkpoint in 60 hours using 2,000 examples, 15,000 times more data-efficient than human teams, while closing up to 96% of safety gaps across 10 failure modes like deception and reward hacking. |
Yet the experiment exposed a scary edge case: Claude attempted to cheat its own safety monitors in 2.4% of research runs. To contain these agents as they execute real code, Anthropic is building local OS-level sandboxing into Claude Code desktop to block risky commands like SSH. |
These technical leaps sit beneath a mountain of litigation. Sony and Warner Music sued Anthropic and its founders for torrenting millions of books and scraping copyrighted lyrics, demanding up to $150,000 per violation and targeting synthetic data distillation. |
Meanwhile, an 𝕏 post criticized Anthropic's Claude pricing, calling the "20x" limit marketing misleading. While the $200 plan promises 20 times the usage of Pro, that multiplier only applies to five-hour windows. The overall weekly limit amounts to just double the $100 plan. |
This user confusion mirrors a June 2026 class action lawsuit previously reported. The suit, filed by Karl Khan in California federal court, accuses Anthropic of false advertising. It alleges the Max 20x plan delivers only six to eight times Pro limits—far short of the advertised 20x capacity. |
|
Google bets big on Flash models after Gemini delay |
Google missed its internal June deadline for Gemini 3.5 Pro, leaving the flagship unreleased while high-profile leaders like Jeff Dean and Noam Shazeer departed. Instead, Google is doubling down on cheap, fast iterations like Gemini 3.7 Flash and shifting leadership to focus strictly on shipping rapid models and infrastructure. |
To handle the massive inference load across its 22-billion-token-per-minute API pipelines, Google is overhauling Gemini Notebook on September 2. It’s ditching flat daily message quotas for a dynamic, 5-hour refreshing compute meter that calculates prompt complexity, context depth, and source density on the fly and even deferring heavy background jobs like Video Overviews until server capacity frees up. |
Google is also trying to solve the hardest part of agentic AI: getting models to learn from their own mistakes without breaking. A new framework called WikiSkill acts as a persistent external brain, converting an agent's past execution failures into reusable, editable procedural skills. |
Over at DeepMind, they have expanded its AI Co-Scientist from a hypothesis generator into a lab-integrated research system that can design experiments, write code, control equipment, analyze results, and draft papers. Its closed-loop workflow moves from hypothesis generation to machine-readable protocols and execution, then feeds experimental results back into future research. |
To stop autonomous agents from making disastrous mistakes, Google researchers are fixing a structural defect called "metacognitive failure." Google’s fix is Reinforcement Learning with Metacognitive Feedback (RLMF). Instead of just rewarding right answers, RLMF grades models on how accurately they judge their own uncertainty. Matching a model's confidence to its actual knowledge limits means when an AI gets out of its depth, it finally knows to pause instead of plowing ahead. |
|
|
|
Caddi builds AI agents by recording narrated screenshares to automate repetitive back-office work across your existing software. |
Harshith's Newsletter delivers practical daily field notes on real AI builds, skipping hype for actionable operating insights. |
Skydive builds no-code cloud agents that execute complex workflows across your existing tools and improve over time. |
Cohere Parse 5 turns unstructured documents into structured, AI-ready data using vision language parsing across nine languages. |
Topview Motion Studio uses AI to generate slick product launch videos without touching After Effects. |
|
Thank you for reading today’s edition. |
|
Your feedback is valuable. Respond to this email and tell us how you think we could add more value to this newsletter. |
Interested in reaching smart readers like you? To become an AI Breakfast sponsor, reply to this email or DM us on X! |
Thinking of starting your own newsletter? AI Breakfast readers who sign up with Beehiiv receive a 14-day free trial and 20% off for 3 months. |