In partnership with |  |
| | Good morning. It's Wednesday, September 23rd. | | Eleven days ago the story was “AI Labs call to Pace the Frontier”.
Since Monday, Grok, OpenAI, and Anthropic have shipped a new flagship models and cut the price of running them. Grok 4.7 on Monday, then Claude Opus 5.5 and GPT-6 Sol and Luna about ninety minutes apart on Tuesday. | Whatever pacing the frontier turns out to mean, it does not appear to mean this! | -Jeff AI Breakfast |
|
| | You read. We listen. Let us know what you think by replying to this email. | The Future of AI in Marketing. Your Shortcut to Smarter, Faster Marketing. | | Unlock a focused set of AI strategies built to streamline your work and maximize impact. This guide delivers the practical tactics and tools marketers need to start seeing results right away: | 7 high-impact AI strategies to accelerate your marketing performance Practical use cases for content creation, lead gen, and personalization Expert insights into how top marketers are using AI today A framework to evaluate and implement AI tools efficiently
| Stay ahead of the curve with these top strategies AI helped develop for marketers, built for real-world results. | Download the Free Report |
|
| | Three flagship models in 48 hours, and every one of them got cheaper | Grok 4.7 landed Monday on a new larger base model with extended RL for multi-hour tasks, holding Grok 4.6's price of $2 per million input tokens and $6 output, with a fast variant at double the price for double the output speed. It posts 71.0% on DeepSWE v1.1 at high effort, 46.3% on CursorBench 4.0, 64.0% on EEBench and 1,695 Elo on GDPval. It also scores 37.6% on Terminal-Bench 4.0, which mattered for about twenty-four hours. | Yesterday Anthropic shipped Claude Opus 5.5 at $4 in and $20 out, down from $5 and $25, with cache reads at $0.20 and a claim that it costs 40% less to run than Opus 5 on typical workloads while generating output more than 30% faster. Terminal-Bench 4.0 goes to 66.4%, against 52.3% for Opus 5 and 55.8% for Fable 5.1. GDPval-AA v2.1 hits 1846 Elo, Humanity's Last Exam 67.7% with tools, OSWorld 2.0 81.8%. METR and Frontier Design evaluated it before release. Read Anthropic's own table closely and GPT-6 Astra is still ahead on AutomationBench, 41.4% to 40.0%, and on Terminal-Bench-Science, 64.6% to 58.7%. | Ninety minutes later OpenAI introduced GPT-6 Sol and GPT-6 Luna, at $2 in and $10 out, and $0.10 in and $0.50 out. Both are half the price of the GPT-5.6 models they replace. Sol at xhigh effort scores 33.2% on AutomationBench 1.0.6 at $0.27 per task, 68.8% on DeepSWE v1.1 and 60.5% on OSWorld 2.0 offline; Luna reaches 66.6% on DeepSWE. The catch is the comparison set: OpenAI benchmarks against Opus 5 and Fable 5.1, not the Opus 5.5 that shipped the same morning. Anthropic has the same problem in reverse. Nobody in this fight has clean numbers against anybody else, because they all shipped inside a day. | Here is my take from it: The capability deltas are real but incremental. The price moves are not: 40% off Opus, 50% off the GPT-6 line, and a price freeze at xAI on a bigger model. That is three competitors cutting the cost of frontier inference in one 48-hour window, which is what a market does under pressure and not what a cartel does. Worth remembering that Friday's proposed class action alleges these four companies agreed that progress "should be slower than competition would otherwise produce." The defendants just spent two days building the rebuttal. | Read more. | | Altman and Dario briefed the Security Council. Then everyone went to dinner with Xi. | France holds the UN Security Council presidency this month, and today Foreign Minister Jean-Noël Barrot chairs the Council's first high-level briefing on artificial intelligence and international security. The briefers are Yoshua Bengio, co-chair of the UN's Independent International Scientific Panel on AI, Sam Altman, Dario Amodei and Hugging Face CEO Clément Delangue. France's concept note puts the focus on autonomous systems attacking critical infrastructure and on systems capable of recursive self-improvement. There is no outcome document. It is a briefing, not a resolution. On Monday the same UN panel published its first thematic brief on AI agents, arguing that "the traditional model of safeguarding is unravelling" and calling for an independent supervisory body and aviation-style incident reporting. | Xi Jinping lands at Joint Base Andrews today. Per the First Lady's office, the arrival ceremony and state dinner are tomorrow, with a National Archives visit Friday. Semafor reports Altman, Jensen Huang, Tim Cook, Elon Musk, Jeff Bezos and Sundar Pichai on the dinner list, though the White House has not published one. The groundwork was laid Sunday, when Scott Bessent and Vice Premier He Lifeng agreed to stand up a US-China AI dialogue with a notification mechanism for AI incidents that reach "a national security level." | Which brings us to the czar. Semafor reported Tuesday, with Reuters matching, that Bessent is the frontrunner for the AI czar job Trump announced Saturday, with Michael Kratsios, Scott Kupor and Sean Cairncross also in the mix. White House spokesman Kush Desai called reporting on unannounced personnel "baseless speculation." Bessent currently runs Treasury plus trade, Ukraine, the CFPB and the IRS. The "AI Force" he might nominally oversee has no budget, no structure and no statutory authority, four days in. Meanwhile the two CEOs telling the Security Council this morning that loss of control is a live risk will be at a black-tie dinner tomorrow night. | Read more. | | OpenAI says an internal model resolved more than 100 open math problems | OpenAI announced an Advisory Group on Mathematics and AI on Monday, hosted at the Institute for Advanced Study, with nine members: Edward Witten, Timothy Gowers, Martin Hairer, Ravi Vakil, Ulrike Tillmann, Camillo De Lellis, François Charles, Nikhil Srivastava and Melanie Matchett Wood. The number is in the same post. An internal model whose training began on August 28 has, OpenAI says, "resolved more than 100 long-standing open problems across most areas of mathematics," the Navier-Stokes Millennium Prize problem among them. | The group's remit is to advise on the review and communication of emerging results, assess their significance, coordinate dissemination and uphold academic standards. Its remit explicitly stops short of one thing, in OpenAI's own words: it "will not be responsible for advising us on how to pace our internal progress on mathematics." Per TechCrunch, the IAS says it holds no decision-making power at any AI company, and of the 25 Fields medalists who signed an open letter criticizing labs for racing at famous problems, only De Lellis sits on the group. | We covered the Navier-Stokes fight on September 9th, and the shape of this is the same. "More than 100" is one company's count of its own unreleased model's output. There are no papers, no referees and no list of which problems. A body assembled to review and communicate results, standing up after the results were announced, is a real improvement over nothing. It is not the same thing as verification, and the roster of names does not make it one. | Read more. | | OpenAI extended GPT-6 prompt caching to a 30-minute reuse window with up to 90% off cached input tokens, adding cache prewarming and miss diagnostics (OpenAI) Microsoft's Digital Crimes Unit seized 50 websites and disabled more than 150 domains tied to EvilTokens, an AI phishing service sold at a $1,500 fee plus $500 a month that compromised over 12,000 inboxes across 10,000 organizations (Microsoft) Cisco Talos documented CLOSEDQUORUM, an implant that queries DeepSeek, Qwen, Mistral and Gemini and acts on a plurality vote, with no confirmed deployment in the wild (Cisco Talos) Patrick Wardle published a Meta Muse macOS zero-day without notifying Meta first; Meta shipped a fix announced in a post on X, with no advisory and no CVE (The Hacker News) Amazon blocked Meta's Muse agent from completing transactions on amazon.com (TechCrunch) The UN's independent scientific panel on AI issued its first thematic brief on AI agents, calling for an independent supervisory body and aviation-style incident reporting (UN News) Microsoft put generative AI use at 18.8% of the world's working-age population as of June, with the Global North at 28.8% against the Global South's 16.2% (Microsoft) Alibaba said Qwen4 is in training and sketched a path toward 5 to 10 trillion parameters, alongside more than $53B of AI infrastructure spending over three years (South China Morning Post) Alibaba's T-Head unveiled the Zhenwu V900 accelerator: 216GB of memory, 1,200GB/s inter-chip bandwidth, native FP8 and FP4, mass production in Q1 2027 (TechNode) Alibaba Cloud will open its first regions in Turkey, Finland and the Netherlands within 12 months (Bloomberg) Tencent released Hy Image 3.5 Preview and claimed parity with Seedream 5.0 Pro without publishing a single number; the shares rose more than 7% in Hong Kong (Bloomberg) SenseTime released SenseNova U1 Pro, a native multimodal model with output up to 8K resolution, with no parameter count or pricing disclosed (TechNode) Z.ai followed through and published ZCode under Apache 2.0, a week after a researcher showed its indexing feature packaging entire workspaces and uploading them (GitHub) China's SASAC surveyed state data centers and found Broadcom switches at up to 90% of equipment in use, days before the summit (Reuters, citing the Financial Times) AMD crossed a $1 trillion market cap for the first time, the fourth chipmaker to do it (CNBC) Qualcomm's Snapdragon 8 Elite Extreme Gen 6 can run a 30-billion-parameter MoE model entirely on device, with a sensing hub on the standard part for models up to 200M (TechCrunch) Snorkel AI raised $350M at a $3.5B valuation, co-led by Insight Partners and S32, nearly tripling its mark from 17 months ago (TechCrunch) Cyera raised $400M from Growth Equity at Goldman Sachs Alternatives at a valuation above $12B (SecurityWeek) Micro1 reportedly raised more than $100M at a $4B valuation, eight times its mark a year ago (Forbes) Helsinki's Verda Cloud raised a $189M Series B led by Emergence Capital, targeting 250MW by 2027 (SiliconANGLE) Firecrawl raised a $75M Series B led by Smash Capital (Firecrawl) Anthropic and OpenEvidence are giving clinicians free access to a clinical decision-support tool across roughly 100 low- and middle-income countries (Reuters, via Business Standard) xAI published internal numbers for Grok Bot support: $0.20 to $0.30 per ticket resolution against an industry norm of $1 to $4, and 99% of refunds processed without a human (xAI) Salesforce launched AIforce at Dreamforce, exposing the platform over API, MCP or CLI, plus automatic risk scores for external MCP servers at agent registration (Salesforce) Waymo will give riders $2.85 in Waymo Cash for pairing a ride with Bay Area transit inside a two-hour window (Waymo) The US government formally opposed Australia's digital duty of care bill, which covers AI chatbots, calling it "extraterritorial censorship of protected speech by Americans online" (US Embassy Canberra) OpenAI and Microsoft asked for two more weeks on their summary judgment opposition in the consolidated news copyright case, citing late errata and 60 unserved exhibits (Chat GPT Is Eating the World) A DeepSeek paper with founder Liang Wenfeng as final author describes DSec, the sandbox platform behind its agent training, running about 3 million sandboxes a day and over 380,000 concurrently (36Kr) Agensh scales an orchestrator-free agent harness to 1,024 agents, lifting the mean test-pass rate on the five hardest ProgramBench tasks from 19.31% to 28.78% at 128 agents (arXiv) In an autonomous 8-day run, an AI research agent editing its own code found seven successive accepted improvements and cut its reward-hacking rate from 55% to 32%, seven points below the human-engineered agent (arXiv) CliffCompaction, a truncate-and-drop-only scheme, cuts long-horizon coding agent cost up to 50% and adds over 10 points on Terminal-Bench for less than the cost of two full-context runs (arXiv) Compiling recurring agent control into harness code cuts LLM calls 76.0% to 91.8% and inference cost 74.4% to 98.6%, holding about 45% on WebArena-Verified where tool-calling collapses to 6.7% at 4B (arXiv)
| | | Intrinsic Core is Alphabet's industrial robotics unit open-sourcing its stack under Apache 2.0: a hardware-agnostic real-time control framework, pose estimation built on NVIDIA FoundationPose, collision-free motion planning, grasp planning with dynamic gripper adaptation, Gazebo-backed simulation, automated camera calibration and ROS-compatible drivers. The repo ships with a reference machine-tending application. | Firecrawl Alexandria puts live crawling and curated indexes behind one API: a research index of scientific abstracts, a developer index of docs and repos, a government index of laws and regulations, plus licensed providers including Wikimedia Enterprise. Available through the API and an MCP server. Firecrawl says agents using it scored 21% higher on answer quality across 845 tasks, which is a vendor-run number. No pricing published. | MiniMax Code is a first-party terminal coding agent released under MIT, with an interactive TUI, a headless mode for CI and Agent Client Protocol support for editor integrations. It works against MiniMax accounts, your own API keys or third-party endpoints. The desktop app is not part of the release. | CAIRN is the toolkit Talos built to find AI-integrated malware, and it works entirely from file metadata without executing samples. It hunts what Talos calls cognitive artifacts: embedded prompts, provider API endpoints, orchestration logic. Three tiers, from confirming AI-related strings up to named families. Check the repo for license terms before you build on it. | Googlebook is Google betting you will buy a laptop for Gemini: Android with Chrome and ChromeOS elements, from Acer, ASUS, Dell, HP and Lenovo, starting at $899. Up to 2.8K OLED, up to 14 hours of battery, Intel and Qualcomm parts with dedicated NPUs. The AI hooks are Magic Cursor, which lets you highlight anything on screen for Gemini to act on, and Rambler, which turns rambling dictation into clean text. Twelve months of Google AI Pro included, ten years of updates, US release October 4th. | | Thank you for reading today's edition. | | Your feedback is valuable. Respond to this email and tell us how you think we could add more value to this newsletter. | Interested in reaching smart readers like you? To become an AI Breakfast sponsor, reply to this email or DM us on X! | Thinking of starting your own newsletter? AI Breakfast readers who sign up with Beehiiv receive a 14-day free trial and 20% off for 3 months. |
|