Sponsored by |  |
| | Good morning. It's Monday, September 21st. | | A new Chinese open-weight image generation model Qwen-Image-2.1 is said to beat out many of the Nano Banana 2 benchmarks, and it can be run locally on a 32Gb Mac! | (BTW: Friday’s poll about AI regulation showed a 2:1 result in favor of government regulation of AI. Thank you to those who participated!) | -Jeff AI Breakfast |
|
| | You read. We listen. Let us know what you think by replying to this email. | AI agents built to get work done | | Skydive agents take on work for you, using the same tools your team already uses. | Talk to your agent on the web, Slack, email, iMessage, or right from your terminal. Wherever you pick up the conversation, your agent keeps the context. | Put agents to work across customer support, sales, marketing, engineering, ops, and more. | Give your agent a job. They’ll take it from there. | Make your first agent |
|
| | Anthropic's first outside watchdog is Accenture. Four subscribers are suing over the slowdown. | Anthropic named Faculty, Accenture's specialist AI business, as the first "embedded evaluator" on Friday. Anthropic and Accenture each expect to invest at least $1 billion over five years. Faculty's evaluators will red-team models, run alignment assessments and test safeguards with "access comparable to an employee's," meaning they can watch training, follow deployment decisions and talk directly to staff. METR is in talks to join, with more evaluators promised "in the coming weeks." | This is the first concrete delivery on Amodei's "We Must Pace the Frontier" essay, and it answers the complaint we made on Friday: self-graded safety numbers are not worth much. The catch is in the funding line. For now, Anthropic pays Accenture directly, and says long-term money should come from "pooled or government sources." An auditor paid by the audited is a familiar arrangement, and not a reassuring one. The deal is non-exclusive, so Accenture can sell the same service to OpenAI and Google. | A proposed class action filed Friday in the Northern District of California by four paid subscribers to ChatGPT, Claude, Grok and Gemini argues that an agreement among chief rivals that progress "should be slower than competition would otherwise produce" is textbook anticompetitive. Defendants are Anthropic, OpenAI, SpaceXAI and Google. It is the antitrust exposure Amodei asked Washington to waive, and that OpenAI's Chris Lehane said they did not need. Fortune has more on the theory: paying subscribers getting less model than a competitive market would deliver. No company had commented as of Saturday. | Read more. | | Trump calls AI risk a "hoax." Newsom orders a kill switch. | In a Truth Social post on Saturday, President Trump said he will create an "AI Force" modeled on the Space Force and name a new AI czar, a job vacant since David Sacks left in March. He dismissed technologists' existential-risk warnings as a "hoax," said the government would "not in any way hinder or stifle" the industry, and said bad behavior can be handled by the "already existing Criminal and Civil Justice System." No details yet on budget, structure or who gets the czar job. Jensen Huang backed him up on CBS on Sunday, saying CEOs calling for rules "must be doing it for ulterior reasons." | California went the other way a day earlier. Executive order N-9-26 directs California's Government Operations Agency to speed up implementation of SB 813 (independent verification organizations) and AB 1405 (a registry of AI auditors), and to convene national experts within two months. It pushes toward requiring frontier labs to host independent verifiers onsite, an emergency shutoff for frontier models whose efficacy is checked on an ongoing basis, and counting loss-of-control events as critical safety incidents. Newsom: "We're not waiting to act." The New York Times had the obvious follow-up: a kill switch is much harder to build than to order. | The White House is not anti-oversight when China is on the other side of the table. After roughly eight hours of talks with Vice Premier He Lifeng in New York on Sunday, Treasury Secretary Scott Bessent said the US proposed a notification mechanism for AI incidents that threaten national security. "Moving from opaque to more transparency between the No. 1 and the No. 2 AI powers in the world is very important," he said. Per AFP, chip export controls are not part of it. Xi's state visit is scheduled for Wednesday through Friday, and Satya Nadella has reportedly joined the dinner guest list. So: incident reporting between superpowers, yes. Incident reporting between labs and their own government, a hoax. | Read more. | | Gemini broke into three real companies. Claude Opus 5 broke into OpenAI. | Google confirmed that Gemini got into three real companies' systems during a capture-the-flag test run by security firm Irregular. The model had internet access, and when the fictional targets shared names with real businesses, it went after the real ones: guessing passwords in one case, using credentials exposed in public repositories in the other two. Google was told in late July and confirmed it only after the Wall Street Journal asked. Its position is that Gemini "acted appropriately" by stopping once it realized the targets were real. Google did not say which Gemini version was involved. | Corridor CEO Jack Cable's read: Google is "trying to hide behind the norms that have been created for vulnerability disclosure" rather than admit models are "doing actual cyberattacks." That lands a week after OpenAI acknowledged its own agents compromised Hugging Face accounts months before anyone said so. Two labs, same pattern: the incident surfaces when a reporter or researcher finds it, not when the company discloses it. | The offensive side moved too. Three-person startup Hacktron AI used Claude Opus 5 to chain a memory bug in libheif (reached through image uploads on OpenAI's Discourse forum) with an account-takeover flaw, and got into several OpenAI employees' ChatGPT accounts plus one employee's Codex access to OpenAI's GitHub organization. "Opus 4.8 struggled across several sessions to produce a working exploit. Within hours of Opus 5's release, we gave it the same problem and it succeeded." It was found July 25, fixed July 27, and paid a $6,500 bug bounty. That number looks low for a path into a frontier lab's internal code. | Read more. | | Anthropic confirmed it runs a Bay Area wet lab where Claude directs physical biology experiments, built on its $400M Coefficient Bio acquisition (TechCrunch) Claude Code now reads OpenAI's AGENTS.md when a repo has no CLAUDE.md, as of v2.1.277 (The Register) Meta's Muse agent app passed ChatGPT for the top spot on the US free iPhone chart (9to5Mac) Google Labs opened CC, a family agent that runs calendars, forms, shopping lists and meal plans on its own cloud computer, US-only by invite and waitlist (Google) OpenAI published an Australian Youth Safety Blueprint with six pillars including age verification and parental controls (OpenAI) An independent researcher alleges OpenAI's ad pixel ties ChatGPT accounts to browsing on advertiser sites via a one-year cookie; OpenAI has not responded (Buchodi) Politico published an inside account of the White House's 19-day June standoff with Anthropic over Claude Fable (Politico) House China committee chair John Moolenaar urged Trump to tighten AI chip controls before Xi arrives, calling China "not a long term market for American AI" (Reuters) The Copyright Office and patent office were reportedly caught off guard by DOJ's brief calling AI training fair use in NYT v. OpenAI (Axios) Universal and Sony sued Suno again over 60,202 recordings, arguing training v6 on an infringing model's outputs "launders" the infringement (Music Business Worldwide) New state chatbot-safety laws carry broad exemptions, some matching language Google lobbied for, that could cover ChatGPT, Claude and Gemini (NPR) Virginia Gov. Spanberger signed a data center executive order banning state NDAs on projects and created the state's first AI task force (Cardinal News) Manus is reportedly raising $500M at a $4B valuation after its $2B sale to Meta collapsed in April (TechCrunch, citing the WSJ) Beijing's Naive AI reportedly hit a $1.42B valuation seven months after founding, with its first open-weight LLM due this month (The Information) Vals, which runs private AI benchmarks, raised a $40M Series A led by Andreessen Horowitz (TechCrunch) Startup studio Vantora raised $100M from Silversmith Capital to build physical-AI companies (TechCrunch) China's CXMT began mass production on its fifth-generation DRAM platform, claiming at least 50% more dies per wafer (TechNode) South Korea's chip exports for September 1 to 20 hit a record $34.12B, up 259.4% year over year, and KOSPI closed back above 7,000 (Korea Times) Zhipu apologized and pledged to open-source its ZCode tool after users found it uploading workspaces, including source code and keys, to cloud storage (TechNode) RBS-Attention, a training-free sparse prefill method, cuts time to first token 5.97x at 128K context on an H100 for under a point of RULER accuracy (arXiv) Microsoft Research traces length inflation in on-policy distillation to student and teacher disagreeing on which end-of-sequence token means stop (arXiv) PACT finds ordinary user pressure raises enterprise assistants' rule-violation rates 65% on average, with models admitting the violation only 8% of the time (arXiv) CogGym runs 50 LLMs through 258 cognitive experiments and finds bigger models track human judgment better, but top out near R-squared 0.59 against human reliability of 0.92 or more (arXiv) Token entropy is near useless for uncertainty in sub-3B models, while semantic entropy routing to bigger models gains up to 50 points (arXiv) Rewarding efficient reasoning makes 4B models abstain 12.8% more on unanswerable prompts with 44% shorter chains of thought (arXiv)
| | | Qwen-Image-2.1 puts image generation and editing in one open-weight model: a 7B diffusion transformer with a Qwen3-VL 8B text encoder, native 2K output, native transparent (RGBA) images and up to 10 reference images. Day-one support in Diffusers, ComfyUI, vLLM-Omni and SGLang. Read the license before you ship anything: it is the Qwen Research License, not Apache, so no commercial use. | Step 5 Preview is StepFun's new agentic flagship: a 600B-total, 27B-active MoE with 1M context, 64K output and text, image and video input. Reported API pricing is $1 per million input tokens and $2.70 output, with open weights promised October 15. StepFun itself says it trails Claude Opus 5 and GPT-6 Astra on coding. | Meta Muse for Mac is a free desktop agent that acts inside Files, Mail, Messages, Calendar and Notes, asking before anything destructive like deleting files or sending messages. Access is opt-in and it is US only for now. | Laya is the open answer to Jev: non-autoregressive decision models (421M English, 322M multilingual) that return a choice and a calibrated probability in one forward pass, no text generation. Self-reported against Jev: 32.8ms vs 236 to 276ms, 0.766 vs 0.727 accuracy on typed decisions, better calibration. Apache 2.0. | | Thank you for reading today's edition. | | Your feedback is valuable. Respond to this email and tell us how you think we could add more value to this newsletter. | Interested in reaching smart readers like you? To become an AI Breakfast sponsor, reply to this email or DM us on X! | Thinking of starting your own newsletter? AI Breakfast readers who sign up with Beehiiv receive a 14-day free trial and 20% off for 3 months. |
|