OpenAI Launches the Agents API (3 minute read)
OpenAI introduced the Agents API in public beta, giving developers access to the managed agent harness and infrastructure behind Codex. It handles context, tools, subagents, persistent execution, files, and code environments for agents that can run for extended periods.
|
Meta to announce Shared Agents for Muse at Meta Connect (2 minute read)
Meta's Muse app will introduce a "Shared Agents" feature, allowing users to create customizable agents that can be shared with others, similar to Grokbot's system. This could benefit small businesses already using Meta platforms by enabling specialized agents for workflows like customer support and sales.
|
|
Detecting and countering misuse of AI: September 2026 (5 hour read)
Anthropic's Threat Intelligence team has identified and disrupted several operations in the past several months where threat actors tried to use Claude for malicious activity. This report shares case studies from those operations and describes how malicious use of Claude has evolved since the company's previous threat reports. The report covers disruptions between December 2025 and August. None of the misuse cases involved the use of Claude Fable or Mythos-class models.
|
Does Scaling Web-Video Pre-training Help Real Robots Do Real Work? (36 minute read)
Larger video models and increased pre-training compute improve robot task performance, confirmed through Direct Video-Action models. Performance gains arise from better predictions of held-out web videos, with larger models excelling in real-world tasks like complex industrial unpacking. Pre-training quality, measured via DINO FD, predicts downstream robot efficiency, enhancing scalability for real deployments.
|
|
Model Card for North Small Translate (8 minute read)
North Small Translate is an open-weights research release. It has 25 billion active parameters and 218 billion total parameters. The model is specialized for high-quality machine translation across 50 languages.
|
Introducing SWE-2: Pushing the Pareto Frontier (23 minute read)
SWE-2 pushes the Pareto frontier and achieves 50.0% on FrontierCode 1.1 Main1, while being 64% cheaper. It beats SWE-1.7 and Grok 4.6 on both score and cost, matches GPT-5.6 Sol and Fable 5/5.1 at a fraction of their price, and comes within a few points of GPT-6 Astra at a quarter of the cost.
|
OpenCodeReview (GitHub Repo)
Open Code Review is an AI-powered code review CLI tool. It originated as Alibaba Group's internal official AI code review assistant — over the past two years, it has served tens of thousands of developers and identified millions of code defects. It has been validated at massive scale. The agent can read full file contents, search the codebase, inspect other changed files for context, and produce deep reviews.
|
OpenAI launches GPT-Live-1 for full-duplex voice agents (2 minute read)
GPT-Live-1 is now available in the OpenAI API at $0.05 per minute. The model adds full-duplex speech, interruption handling, and 12 voice options. It can listen and speak at the same time, handle interruptions and acknowledgements as they happen, and keep conversations moving while performing deep reasoning or actions. The model can control tone, space, and style through the system prompt. Early tests report 80% fewer interruptions than with previous turn-based systems.
|
|
OpenAI puts Pro subscriptions on hold due to Astra demand (2 minute read)
OpenAI has paused subscriptions for its $200-per-month Pro plan. The company's Astra model is now rolling out to Pro, Plus, Enterprise, and Business accounts. The model promises a major leap forward in reasoning, coding, and computer use. OpenAI says the model is the beginning of the AGI era.
|
|
|
Love TLDR? Tell your friends and get rewards!
|
|
Share your referral link below with friends to get free TLDR swag!
|
|
|
|
Track your referrals here.
|
|
|
|