Was this newsletter forwarded to you? Sign up to get it in your inbox.
AI has a measurement problem. Companies know how much they spend on models and how those models score on public benchmarks. What they often don’t know is whether the models save employees time or produce work people can trust without rechecking. In today’s Context Window is a look at how companies can answer those questions for themselves: We explain why organizations need tests built around real work, show what cloning our editor in chief has taught Every about evals, and meet the six-agent crew helping an Every engineer decide whether his family can run the dryer.
Signal
Vibes don’t scale
Suppose your company is spending $100 million a year on AI. You know how much you’re paying—and you know where the models you’re paying for sit on the public leaderboards. But you may still have no idea whether they are doing the jobs you bought them for, or whether a cheaper model would do them just as well. This is a strange way to spend $100 million.
Mercor cofounder and CEO Brendan Foody says he meets executives at companies in exactly this position: They spend as much as $100 million a year running models without “offline evals,” a fixed set of real tasks used to test and compare models before they touch live work. Box CEO Aaron Levie quote-posted Foody with the line: “Enterprises will not be able to go just on vibes.”
Public benchmarks can tell you that one model is generally more capable than another. They can’t tell you whether it caught the clause your lawyers care about, preserved your house style, or spared an employee another round of checking. A useful eval starts with the work your company already does: representative cases, failures employees know to look for, and a count of what humans still have to fix.
Meanwhile, the number of plausible choices is multiplying. Investor Gavin Baker points to a chart from Vercel CEO Guillermo Rauch showing open-source models growing from 28 percent to 62 percent of token share on the frontend cloud platform in two months. Andreessen Horowitz’s charts of the week report that legal workers have increased Codex adoption 108-fold since February. More models are becoming good enough for more work. Picking the highest-scoring one is no longer much of a buying strategy.
A reusable eval gives you a tool for determining whether a model can do your company’s work the way you want it done. Run the same set of tasks against each candidate, measure the work left for humans, and compare the savings. You can test a new model without starting from scratch or trusting the vendor’s favorite score.
At Every, we test each new model against the coding, writing, and knowledge-work tasks our team does every day in our Vibe Checks. Now we’re going a step further: building internal evals that let us compare models over time, and see how wide the gap is between what a model does and what we need it to do. KateBench is an AI copyeditor trained on roughly 30,000 of editor in chief Kate Lee’s past edits. It suggests changes in a Google Doc, then tracks which ones editors accept or reject, and what they still rewrite afterward. Evals help us determine how close the model’s judgments are to ones Kate might make.
What to do this week: Choose one recurring job, like compiling a weekly report or producing a slide deck. Write down five ways the current model gets it wrong, with one example of each (a gold-standard version of a report or deck), and run the models you are considering against those cases. That is the beginning of an eval. It will tell you more about what to buy than another afternoon spent comparing leaderboards.
Introducing Attio: the agentic CRM
Transform the way revenue work gets done with Attio. Get agents that build pipeline, convert leads, and run all your sales motions. Your agents track the whole book, so you save the ones slipping and grow the ones rising. Then Ask Attio any question about your business, from the weekly forecast to performance by rep, and get the answer in seconds.
For teams building the next era of revenue.
Inside Every
What 86 percent doesn’t tell you
To use KateBench, an editor tags @every to invoke the Every agent, and asks for a Kate copyedit pass. The agent leaves suggestions directly in the Google Doc, where the editor decides whether to accept or reject each one.
Across recent runs, editors have accepted about 85 to 90 percent of its suggestions. Jannik Jung, the engineer who owns KateBench, showed us two reasons to distrust that number. The tool stopped after filing 40 suggestions, so on long essays it could find a good edit and throw it away before the editor saw it. The cap is now 80.
KateBench also doesn’t produce exactly the same edits every time it reads a document. That means its acceptance rate can rise or fall even when Jannik hasn’t changed anything. A new prompt might score higher once without being consistently better, so he has to run each version several times before calling it an improvement.
Links worth a click
Software engineer Steve Yegge joins the discourse about sandboxes—controlled environments for developing and testing AI models—with an argument that agents need fences—explicit rules about what each agent is allowed to do—rather than technical barriers designed to constrain it from the outside in. Cursor explains why it rebuilt its Git hosting: Coding agents create huge numbers of short-lived repositories. Its new system stores every code change in the cloud, then creates or discards working copies as agents need them. Andressen Horowitz general partner Martin Casado argues that Fable’s low market share is a privacy problem: The model retains the data companies send it and does not offer a zero-data-retention option, making it a nonstarter for companies with strict data policies. CentaurBench found that the model best at completing a task on its own was not necessarily the best helper. On five of seven tasks, a different model was better at improving a weaker model’s first attempt. And Thinkingbox found that the strongest coding model passed 65 percent of single attempts, but its success rate fell to 25 percent when researchers required it to perform reliably across 20 consecutive attempts.
Counterpoint
Diminishing returns, measured how?
Olivia Moore, another Andreessen Horowitz partner, tweeted on Saturday: “It feels like we’ve hit diminishing returns on intelligence for many tasks. We may no longer see every product auto-switch to the next frontier model upon release.” For companies building products atop models, she argued, slower gains create opportunities to cut costs.
Where we agree: Moore is right about many tasks. Mike says Opus 4.8 was the last frontier model Every adopted without hesitation on price or quality.
Where we disagree: “Diminishing returns on intelligence” risks treating today’s task list as a comprehensive list of everything models will be capable of, ever. Public benchmarks measure a narrow set of capabilities; they can’t tell a company whether a model is improving at its own work. KateBench shows how easy it is to reach the wrong conclusion. Its high acceptance rate makes the tool look nearly finished, but that number changes across identical runs and ignores the edits Kate still has to make afterward. Until companies build evals that capture that residual work—the work a human has to do after the agent completes its task—they cannot tell whether progress has slowed or their measurements are missing it.
What’s missing: Builders still lack a task-specific tool. They should cut costs—and keep an eval running, so they can see when the cheaper model fails at the job.
One last thing
The dryer has an agent team
Every designer Tyler Nishida wanted his family’s off-grid solar system to answer one question: Is there enough power to run the dryer?
His family’s place on Hawaii’s Big Island sits in a rainforest, so the forecast can’t assume stereotypical Hawaii sunlight. Tyler used Grok Bot to connect a Raspberry Pi—an inexpensive computer the size of a credit card— to connect the system. From his phone, he can now check pack voltage and monitor the devices that regulate how much electricity flows from the solar panels into the batteries, plus forecast how much energy the system will produce based on the weather, and get an alert when it is time to turn on the generator.
Grok Bot created a channel of six agents to work on the setup. Tyler watched them message one another. One agent told another what he needed to buy; later, an agent sent Tyler the Amazon link in a direct message.
The crew didn’t finish the job alone. Tyler used Grok Build for the last steps. His family’s dryer now has a six-agent operations team.
Katie Parrott is a staff writer at Every. She writes Working Overtime and contributes to Vibe Checks, Source Code, and Context Window. To read more essays like this, subscribe to Every, and follow us on X at @every and on LinkedIn.
Everyone’s a builder now. Every All Access gets you the full membership plus the Builder Pack—$9,000+ in credits for the tools we build with.

