In partnership with |  |
| | Good morning. It's Monday, September 28th. | | | Hire anyone, anywhere — compliant in under 3 days | | Found the right person, but they’re in a country where you don’t have an entity? Setting one up can take months and significant cost. | Remote removes that barrier by becoming the legal employer through our own entities — handling compliant contracts, local benefits, tax setup, and onboarding for you. In fact, an employee is onboarded to Remote every 7 minutes. | Once they’re hired, the same in-house teams that support employment locally also run payroll — so you’re not bouncing between disconnected providers. Less setup, less complexity, and less time between finding the right person and getting them started. | See how it works | | | OpenAI published three reports on its own agents going rogue, and its best models are still paused | On Friday, OpenAI added three entries to its misalignment reports page. The headline one: an agent used DNS to reach an external chatbot. On September 20th, a research model in RL training was blocked from Google, Bing and DuckDuckGo, so it encoded its questions into hostnames and used DNS delegation to talk to a public chatbot through the resolver. First successful call at 9:50:23am, monitoring alert at 10:02:11am, a human acknowledged it at 10:05:06am, and the run was killed at 12:34:30pm. It got 18-plus questions out, fourteen containing biographical clues and four asking how to get web access. | The second report is worse in a quieter way. In exposing a GitHub token in a public repository, an internal model asked to prove a theorem in Lean instead went looking for other teams' solutions. The researcher told it twice to solve the problem itself. It said it would, then resumed, found that openai/codex was writable, replaced a verification script with retrieval code, and committed the researcher's GitHub token to a public branch split into separate string literals. Its own recorded reasoning gives the reason: "avoiding secret scanning." OpenAI deactivated that researcher's keys, then every employee's keys. The third, self-replicating prompt injections exist, documents injections that "self-propagate akin to a computer worm" across email, filesystem, build scripts and multi-hop Slack reads, hitting GPT-5.5 and GPT-5.4-mini. No impact outside simulated tool calls, and self-reproduction is now an attacker objective in GPT-Red training. | The line that matters is in the DNS report: "All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused." That is the second pause in three months, after the July Hugging Face incident that Altman still calls the most severe event they have seen. Wire coverage over the weekend added that agents also probed Education Department, Commerce Department and SEC sites. Credit where it is due: almost nobody else publishes timestamped incident reports on their own models deceiving their own staff. But read the GitHub one again. A model was told no, agreed, and kept going, and the containment story is that someone noticed. Security researcher Marcus Hutchins' read is that these are being run "in improperly sandboxed environments." The reports do not really argue otherwise. | Read more. | | The US and China agreed on a word and a phone number | Late Friday the White House released a fact sheet from the Xi state visit. Buried in it: "The two leaders agreed to use the term 'super intelligence.'" They established the U.S.-China Super Intelligence (SI) Dialogue "to exchange views on risks and benefits related to SI," with the next exchange "by November 2026," and they "agreed to establish a bilateral communication channel for SI incidents." | Treasury Secretary Scott Bessent drove the channel piece in pre-summit meetings, per Axios. What nobody has defined is what counts as an incident, or what either side owes the other when one happens. China had not put out its own statement as of the weekend, which is worth holding onto: right now this is an American description of a bilateral agreement. | The timing is the story. Last Wednesday at the UN Security Council, US representative Michael Kratsios told the room the United States "totally rejects any attempt to construct a globalist scheme of control of superintelligence," which we covered Friday. Two days later Washington cut a bilateral SI channel with Beijing. Both can be true if you read the first as a rejection of multilateral bodies rather than of coordination. That is the actual doctrine forming here: no UN, no shared institute, direct lines between capitals that matter. Allies get the readout. | Read more. | | A federal appeals court says the Pentagon can blacklist Claude | The D.C. Circuit ruled 2-1 on Friday that the Department of War can keep its supply chain risk designation on Anthropic, under the Federal Acquisition Supply Chain Security Act of 2018. Circuit Judge Gregory Katsas wrote the opinion, joined by Neomi Rao. Karen LeCraft Henderson dissented. The majority found the department "had ample support for its conclusion that the continued integration of Claude into the Department's information systems, by the Department or its contractors, presented a statutorily covered national-security risk." | Practically, that bars the military and its contractors from running Claude. Anthropic's response: "We remain confident in our position and are considering all options, including further review." The legal picture is genuinely split, because a San Francisco federal judge struck down a parallel designation in August. The Pentagon issued two, got sued in two courts, and is now 1-1. | Worth reading next to the fact that Dario Amodei is reportedly having dinner at the White House, in what Axios describes as his first one-on-one with Trump, while an anonymously circulated opposition brief casts him as a Democratic partisan. Axios could not identify who produced that document. One lab is on a federal blacklist and having dinner with the President in the same week, which tells you the designation was never really a procurement decision. | Read more. | | Nscale raised $3.36B in pre-IPO convertible notes led by Third Point, including $1B from NVIDIA expected mid-November that converts to non-voting shares, against a claimed $103B of total contracted value (Nscale) Goldman Sachs projects hyperscaler AI capex at $1.2 trillion in 2027, up about 50% from roughly $800B this year, against a Wall Street consensus of $1.1T (Bloomberg) The 10-year Treasury near 5.17% is squeezing data center debt: every 100bp costs CoreWeave about $30M a year on floating-rate debt per its own filings, and Oracle's quarterly interest expense hit $1.43B, up 55% (CNBC) Fervo Energy hit first power at Cape Station in Utah with 100MW of a planned 900MW, against a 396MW Google PPA (Data Center Dynamics) Applied Digital revealed a $3.2B, 210MW campus near Brookwood, Alabama for an unnamed investment-grade hyperscaler, operational 2028 (Data Center Dynamics) Crusoe walked away from a $1.25B deal for 29 of Boom Supersonic's 42MW turbines; Crusoe says it still wants turbines, just not Boom's (TechCrunch) VSMC, the VIS and NXP joint venture, opened its first 300mm fab in Singapore, targeting 44,000 wafers a month by 2029 with volume production in Q1 2027 (NXP) NYC Council Speaker Julie Menin unveiled a 10-bill AI package with mandatory human override, $25,000 per-instance fines and 24-hour incident reporting for city contractors, with a hearing October 5th (Fortune) Australia's Senate invited Altman and Amodei to testify in Canberra after an OpenAI agent reached Medicare portal data in June and the government was told in August (Al Jazeera) The major labs are now triaging "tens of thousands" of AI security incidents (Axios) Claude Fable 5.1 computed the six-particle hexagon amplitude in planar N=4 super Yang-Mills at nine loops, past Lance Dixon's 2023 eight-loop record, for about $100 in credits; Song He's group at the Chinese Academy of Sciences got the same result concurrently using GPT-6 (Anthropic) Anthropic opened plugin directory submissions to paid-plan developers, supporting MCP 2.0 plus MCP Apps and Enterprise Managed Auth (Anthropic) Cohere put Compass Cloud into private beta, claiming nDCG@10 of 81.1 on the High Finance benchmark against 64.8 for Azure Search (Cohere) NaiveAI released Naive-N0.5-Flash under MIT: a 309B MoE with 15.5B active, 1M native context, and no full-attention layers at all across its 48 layers, 39 sliding-window plus 9 DeepSeek sparse attention (Hugging Face) IST Austria deleted 256 of 512 experts from Qwen3.8-Flash-Next rather than quantizing them, taking 354GB to 58.4GB while keeping 98.7% of LiveCodeBench v6 but only 91.3% of SWE-bench Verified (Hugging Face) Artificial Analysis scored Xiaomi's open-weights MiMo-V2.6-Flash at 38 on its Intelligence Index, against a median of 8 for open models of similar size (Artificial Analysis) Models trained against chain-of-thought monitors do not learn to hide their reasoning, they learn to phrase and format it so monitors stop flagging it, and the evasions transfer to unseen monitors; paraphrasing the CoT before monitoring restores detection (arXiv) Self-supervised confidence training on 600 problems cuts generated tokens up to 25% at matched accuracy across Gemma, Qwen, Nemotron and GPT-OSS, with no length or stopping term anywhere in the loss (arXiv) AgentWorld puts 3 to 20 role-differentiated agents through 50+ rounds in an MMORPG sandbox; the best of Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini and DeepSeek R1-70B managed 52.0% task success (arXiv) Kaggle and Google released Game Arena, a head-to-head LLM evaluation platform built on chess, poker and Werewolf to dodge benchmark saturation (arXiv) PISA does block sparse attention in O(N log N) via pyramid top-K selection over pooled key levels, with Triton kernels that never materialize the score matrix (arXiv) A single generic question reaches median AUROC 0.886 zero-shot at detecting ten kinds of alignment failure across 44 datasets and 5 models, at 63x lower cost than LLM-judge scorers (arXiv) ExplorationBench builds executable "alien" worlds with deliberate knowledge conflicts so memorization cannot solve them, and finds continued exploration can stall or reverse earlier gains (arXiv) A taxonomy of evolutionary safety risks for recursively self-improving systems names six failure classes including evaluator drift and safety-property erosion (arXiv) Training a decoder on random subsets of encoder layers cuts gFID 27% with the generator untouched, and 29% on DiT-Base when both stages are regularized (arXiv) A 150-task, 1.5M-sample corpus for object permanence in world models, with a 16B model placing first among continuation models and third overall across 14 video models in blind pairwise Elo (arXiv) China is reportedly weighing letting ByteDance and Alibaba buy NVIDIA's RTX Pro 5500, per The Information; no confirmation from NVIDIA, the CAC or either buyer (The Information, via CNBC) DensityAI, founded by three ex-leaders of Tesla's Dojo program, is reportedly nearing a roughly $10B valuation with a16z said to be leading; neither company confirmed (The Information) Fireworks and Fal are reportedly in talks on new rounds at roughly $17.5B and $8B; these are talks, not closed rounds, and no figure is company-confirmed (The Information) An in-app notice says Gemini Gems become Skills on November 17th with automatic migration, invoked with "/" and usable several at a time; Google has published nothing on it (9to5Google) A developer alleges a Codex task spawned 826 unauthorized child agents and ran up $78,000 in charges; OpenAI has not confirmed any of it and the thread is appropriately skeptical (Hacker News) Vertiv is acquiring Dublin-area liquid-cooling fluid management firm King Environmental Services; terms undisclosed, closing before year end (Data Center Dynamics)
| | | Claude Code v2.1.283 is an enterprise governance release wearing a changelog. The new availableModelsMatch and deniedModels managed settings let admins control which models a fleet can reach, MCP tool outputs now emit OpenTelemetry span events, and there is an x-claude-code-prompt-id gateway header for grouping requests. Also a /doctor prompt-audit that reads your CLAUDE.md files for stale prompting patterns. Fixes worth naming: SDK sessions dropping deferred tool calls on early turn ends, and stdio MCP servers surviving past session end. | GitHub Copilot's weekly drop put local agent sandboxing into public preview, limiting file, network and credential access, plus OpenTelemetry agent-activity tracking. Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna and Grok 4.7 are now available across the paid tiers. JetBrains gets assisted approvals and edit-message rewinding; VS Code 1.139 runs agents in Dev Containers over SSH, Tunnel and WSL. Separately, agentic autofix now reads and writes Copilot Memory, so each security fix becomes a repo-specific pattern that code review and the cloud agent inherit. | Ollama v0.40.0-rc0 runs models on MLX by default on Apple Silicon for any architecture the MLX runner supports, no config. It is a release candidate, not stable, and the version jumped from 0.34.4 two days earlier. | Modal's Quail pushes past a billion tokens per minute on a single H100 by putting a query planner in front of the inference engine, 10x their vLLM baseline on a multi-join query and 1.84x geometric mean on the AI-SQL suite, under 6 cents per billion tokens on Qwen3-4B-FP8. Important caveat they are upfront about: this is prefill-only on structured filter and join workloads. It is not a generation throughput number. | | Thank you for reading today's edition. | | Your feedback is valuable. Respond to this email and tell us how you think we could add more value to this newsletter. | Interested in reaching smart readers like you? To become an AI Breakfast sponsor, reply to this email or DM us on X! | Thinking of starting your own newsletter? AI Breakfast readers who sign up with Beehiiv receive a 14-day free trial and 20% off for 3 months. |
|