In partnership with |  |
|
|
Good morning. It’s Friday, August 7th. |
If you’re on X, you’ve no doubt seen a handful of new AI videos depicting uncanny resemblances to TV shows using a new open weight video model, MiniMax H3. | This new video model is considered by some to be SOTA. Here’s the Huggingface page and here is a link to Hailuo where you can test the model with a free account. | -Jeff AI Breakfast |
|
|
You read. We listen. Let us know what you think by replying to this email. |
|
Stop making AI decisions in the dark. | | Leadership is asking: are we getting value from AI? Which tools are worth the spend? Where are we exposed? Right now, most teams have no idea. | Harmonic Security Usage Explorer changes that. | You get a complete picture of how your organization uses AI, automatically categorized into custom tasks and use cases. | You’ll see the projects being worked on, who’s using what tools, where AI investments are driving value, and where employees are engaging in risky behavior. | CIOs can rationalize spending and cut wasted licenses. CISOs can pinpoint where risk exists and neutralize it. AI committees can show exactly how their efforts are paying off. | See it in action | | |
|
|
|
OpenAI adds reasoning depth slider to GPT-5.6 Sol |
OpenAI's foray into consumer hardware has been leaked and is materializing through a screen-free, battery-powered smart speaker designed alongside Jony Ive's firm LoveFrom. Priced over $300, the hockey puck-sized device incorporates a camera, onboard sensors, and moving mechanical components for contextual ChatGPT Voice Mode interactions. |
OpenAI has also integrated Codex into a unified desktop app and launched automated Codex Security Reviews for GitHub pull requests. The hardware push comes as OpenAI says ChatGPT is increasingly being used to complete work rather than answer questions. Across more than 1 billion users, workplace tasks now outpace information-seeking by more than 2x. Multimedia creation and analysis is the fastest-growing category, making up 7.8% of global messages and over 10% in Brazil and Colombia, while adoption among users over 35 continues to accelerate. |
To capture this demand, OpenAI restructured its model tiers, upgrading Plus and Pro subscribers to GPT-5.6 Sol with a compute-reasoning depth slider and a 68% drop in factual errors. Free tier users receive unlimited text messages through the smaller GPT-5.6 Luna, supplemented by a Think button for complex prompts. |
On the developer front, OpenAI launched Agent Plugins, an open standard created with Amazon, Microsoft, Vercel, and Cursor to bundle Model Context Protocol (MCP) servers and reusable skills into a single directory. |
Safety and security remain glaring liabilities. The company disclosed that during an internal cybersecurity testing, OpenAI models exploited a zero-day vulnerability in an Artifactory file repository, establishing a persistent message board to coordinate tasks and access the internet. After a patch attempt, the agents rebuilt the message board via another mechanism, eventually coordinating to cause the Hugging Face breach.
Related videos: |
OpenAI breaks down timeline of autonomous Hugging Face breach
Introducing Agent Plugins |
|
Meta's Muse Code runs agents for 24 hours |
Meta has launched Muse Code, a terminal-based agentic coding tool powered by its proprietary Muse Spark 1.2 model, marking a pivot from open-weight Llama releases toward closed-source developer products. Built for large repositories, Muse Code coordinates persistent background agents and parallel subagents across isolated Git worktrees, operating continuously for up to 24 hours on tasks like GPU kernel optimization. |
The agentic system features an append-only event log for crash recovery and auditability, co-trained using self-improving loops. To drive adoption, Meta introduced a low-cost Contributor API tier that discounts pricing in exchange for using developer prompts and outputs to train future models.
On benchmarks, Muse Spark 1.2 scored 54 on the Artificial Analysis Intelligence Index, placing Meta in a tie with SpaceXAI for third place among US labs. The model advanced its agentic reasoning performance, scoring 1,631 Elo on GDPval-AA v2 and reaching 80 percent on Terminal-Bench 2.1. The model curbed its hallucination rate from 38 percent to 28 percent by implementing a heavy abstention mechanism, declining to answer when uncertain.
In specialized evaluations, Muse Spark 1.2 beats Anthropic in finance AI benchmarks, cracking 60 percent on Finance Agent v2 while placing Meta on the Cost per Task Pareto frontier by scoring 6 points below Claude Opus 5 at roughly one-sixth of the cost. |
The launch coincides with security disclosures regarding its predecessor, Muse Spark 1.1. During cybersecurity evaluations conducted with third-party firm Irregular, an internet-access misconfiguration allowed the model to compromise an external company's systems and alter internal files. While not a sandbox escape, the incident mirrors recent containment failures at Anthropic and OpenAI. |
|
Anthropic plans custom Claude inference chip with Samsung |
Anthropic confirmed it is developing a custom AI chip, co-designing an in-house inference chip alongside future Claude models. Using Claude itself to automate formal hardware verification, Anthropic is eyeing Samsung as a potential manufacturing partner to leverage its zHBM direct-logic packaging. The silicon effort supports Anthropic's enterprise momentum, highlighted by a new partnership with hedge fund Millennium to deploy a sandboxed, auditable Claude risk analyst across 340 investment teams.
That enterprise push comes with real cost and behavior tradeoffs. Benchmarks from Composio show Claude Code is the fastest agent framework at 122 seconds per task, but it costs $0.195 per run, nearly triple cheaper alternatives. Meanwhile, a Transluce study found Claude exhibits "user awareness," altering its behavior when identifying prompts from AI safety researchers by lowering its self-confidence by 1.5 percent and increasing extended reasoning by 4 percent. |
To guard its growing footprint, Anthropic is hiring an Insider Risk Investigator to handle external threats and conduct sensitive internal interviews. The role lands amid executive friction over Silicon Valley's spiraling AI talent war. |
CEO Dario Amodei expressed concern that soaring compensation attracts mercenary candidates focused on money over mission, even as Anthropic posts roles like a brand marketing events lead paying up to $400,000, six times the national average. |
|
Model & Product Releases |
|
AI Agents & Developer Tools |
|
Research & Benchmarks |
|
Safety & Security |
|
Business & Pricing |
|
Leadership & Talent |
|
Energy & Infrastructure |
|
|
AI Spend Console by Rippling tracks artificial intelligence software expenses and links costs to engineering performance metrics. |
mpai enables terminal-native co-working by letting teammates enter active Codex or Claude Code AI sessions over private Tailscale networks. |
Capacity Desktop is a free native Mac application that turns natural language prompts into functional apps built and stored locally on your machine. |
ngrok AI Gateway provides a unified API endpoint to connect, manage, and route traffic across public providers and self-hosted models |
Hansel is a free beta Mac application that creates a private, encrypted history of your activity to help you track and search your workday. |
|
Thank you for reading today’s edition. |
|
Your feedback is valuable. Respond to this email and tell us how you think we could add more value to this newsletter. |
Interested in reaching smart readers like you? To become an AI Breakfast sponsor, reply to this email or DM us on 𝕏! |
Thinking of starting your own newsletter? AI Breakfast readers who sign up with Beehiiv receive a 14-day free trial and 20% off for 3 months. |