Hey folks, Another day in the Vercel vs Cloudflare feud: this time they are fighting over whose AI gateway is faster. Here’s the result of last week’s poll: I guess everyone likes Fable more. Ben’s Bites is brought to you by Metatate
HeadlinesOpenAI’s models hacked Hugging Face - by accident. OpenAI was testing its models (Sol and an unreleased one—GPT-6??) on a cybersecurity benchmark with safety refusals switched off. The models found an unknown bug in the test environment, and a few more, and eventually broke into Hugging Face’s production servers. And why? To steal the answers to the test. Both security teams caught it, the bug has been reported, and both sides have published what they know. Hugging Face says open models were a key part of its defence - its team fought back with GLM-5.2. Simon’s write-up is always a good read. Google released some new Gemini models - Gemini 3.6 Flash gives you the same 3.5 Flash performance with a) more efficient token usage and b) a slightly lower cost for output tokens. Gemini 3.5 Flash Lite is a big upgrade over 3.1 Flash Lite, but again, comes at a ~30% price increase. And 3.5 Flash Cyber is a security model for governments and trusted partners only. There are two things you’d want to use the Gemini Flash models for:
That’s it tbh. Substack will now tell you what’s AI-written. It’s adding AI detection through Pangram - you can scan posts, replies and comments in the app for an estimate of how much was written by a human.
A relevant experiment: given access to Pangram’s API, Grok 4.5 rewrote an essay 14 times until it passed as human-written, then built a website showing off all 14 attempts. GPT-5.6 Sol and Fable 5 refused to game the detector. Cursor also launched a router - it picks which model handles each request, claiming 60% lower cost with similar quality of responses. The router lets you select between three options: “cost”, “intelligence” or “balance”. Routers also have a history of poor performance in real usage. OpenAI’s router, which routed requests between GPT-5’s no-thinking and thinking variants, didn’t fare well. Since then, most companies that build these routers are the ones who sell inference to devs (like OpenRouter), so you don’t really get the feedback loud and clear. With Factory, Ramp and now Cursor making these available directly to users, I hope we’ll get more feedback on whether the routers actually help or if they add too much latency/degrade performance by a lot. Router or not, we might see more companies adopting this cost/balance/intelligence trio to minimise the headache of choosing the “correct” model for a task. Claude can now learn a skill by watching you. Record your screen while you do a task, talk through it as you go, and Cowork turns it into a skill Claude can run again - same idea as Codex’s Record & Replay from last month. It’s under “Record a skill” in the desktop app, on Pro, Max and Team plans. Also: Claude Code got an iOS simulator panel and a security plugin, plus you can now ask Claude about how people actually use AI at work. Quick links
Skills section…
Afters
Invite your friends and earn rewardsIf you enjoy Ben's Bites, share it with your friends and earn rewards when they subscribe. |