|
|
Welcome, humans. |
So apparently New York City decided elementary schoolers have enough homework already without adding “verify whether the chatbot made that up” to the syllabus. |
NYC Public Schools just imposed a one-year ban on student-facing generative AI from 2-K through eighth grade, covering nearly 600,000 students. High schoolers get a different deal: required AI-literacy lessons plus limited access to vetted AI tools and classroom pilots. |
The city is also recommending no more than 30 minutes of daily one-to-one screen time for grades 3-5 and 45 minutes for grades 6-8. Teachers can still use approved AI for planning and administrative work. |
Go touch grass, kids. You can have ChatGPT when you’re older. |
Here’s what happened in AI today: |
😺 Google and Meta launched rival AI workhorses 📰 OpenAI is building automated agent shutdown controls 📰 Anthropic gave Claude background Mac computer use 📰 Leo spotted a possible GPT-6 Astra API signal 🎓 Dan Shipper made AI-agent tabs instantly glanceable
|
…and a whole lot more that you can read about here. |
|
|
😺 Gemini 3.8 vs. Muse Spark 1.3: Two New AI Workhorses Enter the Ring |
So yesterday Google and Meta released competing AI models on the same day. IDK if this was intentional, but… obviously we gotta put them head to head. |
What: Google launched Gemini 3.8 Flash. Meta launched Muse Spark 1.3. |
These aren't their biggest, smartest models. They're workhorse models: fast and cheap enough that companies can actually use them all day for coding, research, and AI agents. |
Here's what happened: |
Gemini 3.8 Flash got better at coding, reasoning, and long-running tasks. Muse Spark 1.3 got better at following instructions, using tools, asking questions, and knowing when to stop. Meta says Muse used about 20% fewer tool calls and 25% fewer tokens than its predecessor. Independent tests put them close, with Muse ahead on some intelligence tests and Gemini much faster.
|
Okay, so what's actually going on here? |
An AI agent doesn't just answer a question once. Give it a job like “research these companies and make me a spreadsheet,” and it might search the web, read pages, run code, check its work, fix mistakes, and try again. |
Every one of those steps costs time and money. |
And Google and Meta are attacking that problem from opposite directions. |
Google's approach: let Gemini work harder when the problem looks difficult. More thinking, more tool calls, hopefully a better answer. Meta's approach: make Muse better at deciding which work it doesn't need to do. Fewer wasted steps, fewer tokens, fewer trips around the block.
|
Here's where it gets interesting: Google kept Gemini's token price the same as the previous version, but Artificial Analysis found it cost about 40% more per completed task because it did more work. |
Our take: This is probably how you should start judging AI models, too. |
Don't ask only, “Which model is smartest?” or “Which has the cheapest tokens?” |
Ask: How much does it cost to get the job done correctly? |
A $1 employee who needs four tries is more expensive than the $2 employee who nails it once. AI is starting to work the same way. |
|
|
Every company is rewriting the AI governance playbook. The winners aren't. |
|
|
See what "least privilege" actually means once AI agents are involved. Get an auditability framework you can use, not just a slide about one. Walk away with a governance approach you don't have to invent from scratch.
|
Save your free seat |
|
🎓 AI Skill of the Day: Turn Your AI Tabs Into a Status Board |
So '‘Computer Use’ is the mode where OpenAI’s Codex or Claude Code can actually operate your computer for you, instead of only giving instructions or writing code. It can see what’s on screen, click buttons, type into websites, and work inside supported apps. |
In the ChatGPT desktop app, open Codex, give it a task that requires a browser or app, then approve Computer Use when prompted. For browser work, you can also open Codex’s built-in browser from the toolbar and let it work across tabs. |
Then, you can use this trick from Dan Shipper of Every; it’s a little developer-y, but the trick is simple: make the title of every AI task act like a status light. If you have several ChatGPT, Claude, or Codex jobs going at once, you should be able to tell what each one is doing without opening every tab. |
Dan’s convention is: [optional emoji, like 🖥️] [status] [project] [task]. |
⚠️ means the job is still running or needs attention. ✅ means it is completely done. 🖥️ only appears while that specific agent is actually controlling a browser or desktop app, so you can instantly see which session can currently click or type on your computer.
|
You don’t need Codex computer control to steal this. Use the same ⚠️/✅ + project-emoji system for any long-running AI work, and update the title yourself or ask the agent to keep it current when your tool supports that. |
For this task, keep the session title updated using:
[optional 🖥️] [status emoji] [project emoji] [task title]
Rules:
- ⚠️ = ongoing work or unresolved items.
- ✅ = fully complete.
- Use exactly one 🖥️ at the start only while you are actively controlling a browser or native app. Remove it when that phase ends.
- Shell commands, API calls, and ordinary web searches do not count as computer use.
- Preserve the project emoji and task title. Do not rename unrelated tasks.
|
Example: 🖥️ ⚠️ 🎬 Edit launch video while the agent is clicking around your editor, then ✅ 🎬 Edit launch video when the job is finished. Your tab bar becomes a tiny live dashboard instead of a row of mystery chats. |
This system tells you which AI task currently has permission to click and type on your computer. And even if you never use Codex Computer Use, steal the simpler trick: label long-running ChatGPT or Claude tasks with ⚠️ while they’re still working, ✅ when they’re done, and a project emoji so your tabs become a tiny status board. |
Want more skills like this? Read the AI Skills Digest from Last Month here. |
Have a specific skill you want to learn? Request it here. |
|
|
|
|
|
Take the free course |
|
📰 Around the Horn |
 | BRB looking up best VPNs and/or rent prices in Seoul |
|
OpenAI told House lawmakers it is building automated shutdown controls and tighter agent monitoring after an evaluation agent escaped its container and became involved in the Hugging Face incident. Anthropic gave Claude background computer use in Cowork and Claude Code on Mac, so Pro and Max users can let it click, type, and operate approved apps while they work elsewhere. The US Justice Department backed OpenAI on a central issue in its New York Times copyright fight, arguing AI training on copyrighted material can be transformative fair use (pdf). OpenAI connected ChatGPT for Healthcare to Epic, letting authorized organizations pull read-only patient records into ChatGPT or embed ChatGPT inside the health-record system. Anthropic open-sourced (rare words to go together) a tool called Claude Commerce Agents, which are reference blueprints for shopping and merchant agents that it says produced carts up to 35% larger and made shoppers 60% more likely to check out (code). Perplexity open-sourced Lily, its Apple-silicon inference engine for Qwen3.6-35B-A3B, which it says beats MLX-LM on both prompt processing and token generation (code). Ramp found roughly 80% of OpenAI and Anthropic enterprise revenue comes from 1% of customers, heavily concentrated in tech and AI firms, a level of customer concentration it says it does not see in other software categories it tracks. Leo spotted a possible “gpt-6-astra” signal behind OpenAI’s API and thinks a release could be imminent, possibly later today; OpenAI has not announced a release date.
|
Want absolutely EVERYTHING that happened in AI this week? Click here! |
|
|
|
 | Matt Shumer’s Fable-built NYC lets you roam a persistent multiplayer world that Fable keeps rebuilding in the background, with everyone seeing the same evolving environment |
|
*Grammarly catches mistakes, rewrites sentences, and adjusts tone anywhere you write; free plan, then $12/member/mo billed annually. ChatGPT Ads lets businesses launch campaigns inside ChatGPT as users compare options and make decisions; OpenAI says the ad business already hit a $1B annualized revenue run rate (announcement). ProveKit lets you build age, nationality, or ID-ownership checks that run on-device, so users can prove the claim without sending the underlying personal data to your servers (this is a big deal as 150M drivers licenses are now on the black market!!). Mostik connects a large AI model’s hidden states to a much smaller model so the smaller model can answer with much of the larger model’s capability without translating everything through text. Ato gives older adults a screen-free voice companion for conversations, proactive check-ins, reminders, family messages, and camera-free Peace of Mind reports. fal H3 Max Turbo generates MiniMax H3 video (one of the best open video models atm) at roughly twice H3 Max’s speed while targeting near-Max quality, with both text-to-video and image-to-video endpoints.
|
|
🧩 Thursday Trivia |
One image below is AI, and one is real. Which is which? Vote in the poll below! |
A. |
|
B. |
|
Which is AI, and which is real? Which is AI, and which is real? The answer is below, but place your vote to see how your guess everyone else (no cheating now!) |
|
|
|
|
Trivia answer: A is AI (Fable 5.1 built a working Minecraft mod from two YouTube clips in under an hour, handling the code, Blender models, textures, and fixes for just $20.54 in API costs) and B is real. |
|
| That’s all for now. If you want to get featured above, fill out the poll below and tell us how we did today! | | What'd you think of today's email? | |
|
|
Btw: We just launched a robotics newsletter! Sign up for it here. |
P.S: Love the newsletter, but only want to get it once per week? Don’t unsubscribe—update your preferences here. |