Inkling-Small (4 minute read)
Thinking Machines has released Inkling-Small, a 276B-parameter mixture-of-experts model with 12B active parameters. The model retains Inkling's multimodal reasoning, variable thinking effort, and 1M-token context window while using substantially less compute.
|
OpenAI Cuts GPT-5.6 Prices (6 minute read)
OpenAI reduced GPT-5.6 Luna pricing by 80% and Terra pricing by 20%, while improving Sol's API speed. The changes extended efficiency gains across API usage, Codex, and ChatGPT Work subscriptions.
|
Gemini Robotics ER 2 (1 minute read)
Gemini Robotics ER 2 enhances automation capabilities with advanced AI and LLM integration. This innovation improves robotic efficiency, making it pivotal for industries needing precise automation.
|
|
The WASTE inference engine (14 minute read)
WASTE is an open source inference engine designed to run models whose weights are substantially larger than the memory available on the host machine. It is an initial step into a broader effort to make increasingly capable models available on more hardware. The project aims to provide people and organizations greater control over infrastructure costs, data privacy, availability, and deployment. The first fully supported model by WASTE is Kimi K3, which can run on a MacBook Pro with 64 GB of unified memory.
|
The Agent Graveyard Isn't Real Anymore (6 minute read)
Enterprise AI projects increasingly reach production when vendors prove value on live workloads, provide ongoing testing and iteration, and expose ROI. Successful deployments start with decomposable workflows that ship quickly and expand, rather than broad transformations with undefined success criteria.
|
Open-Weight LLMs Have Caught Up on Accuracy (21 minute read)
Open-weight LLMs have now reached accuracy parity with closed models in regulatory and clinical tasks, at significantly lower costs. The ClinReg benchmark showed models like GLM 5.2 and Kimi K3 performing within one standard deviation of top proprietary models like GPT 5.6 Sol, at one-third of the cost. Different models displayed distinct error profiles, highlighting the importance of choosing models based on task-specific requirements rather than just ranking.
|
|
The Session You Cannot Take With You (18 minute read)
Users should be able to close an account, keep a session, and hand it to another model. Stateful storage should be optional, and hosted tools should be observable. Compaction should be readable, and agent communication should be auditable. Distillation should be a path by which capability becomes more available, not a reason for building ever higher walls.
|
MiniMax H3 (10 minute read)
MiniMax H3 is an open model that breaks the boundaries between tasks and modalities. It understands unified context across text, images, video, and audio. The model can generate up to 15 seconds of video at 2K resolution with native stereo sound. H3 excels at instruction following, accurate text and brand rendering, and V2V motion transfer. Early testing shows that it is ready for commercial content creation across a wide range of use cases.
|
Gemini Live API overview (3 minute read)
The Gemini Live API enables low-latency, real-time voice and vision interactions with Gemini. It processes continuous streams of audio, images, and text to deliver immediate, human-like responses. The API creates a natural conversational experience for users. It can be used to build real-time agents for a variety of industries.
|
Agent Behavior (Website)
Agent Behavior is an open standard for defining and evaluating how an AI agent should behave across a whole trajectory. Each behavior spec is a Markdown file that describes the recurring conduct that makes the agent reliable. The spec gives reviewers, rubrics, scorers, and evals something concrete to measure against. It can be used to review traces, write eval cases, revise prompts or tools, and communicate intended agent conduct to teams.
|
|
GPU Management: Why Idle GPUs Are the New Grounded Aircraft (13 minute read)
AI faces a bottleneck with GPU utilization, echoing how airlines maximize aircraft use for profitability. Owning more GPUs doesn't ensure efficiency - companies must focus on optimizing GPU workload orchestration to maximize output. Specialized models and continuous GPU management are crucial strategies for improving utilization and staying competitive in the AI arena.
|
With Moonshot's free Kimi K3, China changes the sovereign AI playbook (6 minute read)
China's Moonshot AI open-sourced Kimi K3 on July 27. Any government, company, or individual can now run the model on their own computers for free. They can also retrain the model however they want. The existence of a high-quality open model increases return on hardware investments because of a reduction or elimination of ongoing licensing costs.
|
|
Superlogical (Website)
Superlogical plans to build a composable multiplexer for all work that is safe and operable in production.
|
|
|
Love TLDR? Tell your friends and get rewards!
|
|
Share your referral link below with friends to get free TLDR swag!
|
|
|
|
Track your referrals here.
|
|
|
|