Yesterday we launched the Every Agent, an agentic coworker in Slack that helps your whole company go AI-native. Because we pass its token costs on to you with no markup, every token we save, you save. In today’s Context Window, engineer Paridhi Agarwal explains how she made the Every Agent more token-efficient and shares a four-step workflow you can use to find token waste in your own software. Plus, the models the team is reaching for this week, a fresh batch of Thesis Statements, and Willie Williams on being an AI-first engineering manager. On Friday, sign up to join the Every team for a one-hour camp on working with agents in Slack.
‘The Every Podcast’
Why Every traded personal agents for one company agent
Dan Shipper named his first AI agent R2-D2 and grew genuinely attached to it. Then someone knocked loose the power cord on the Mac mini it ran on, and it disappeared. He never bothered to revive the agent because, by then, Codex had taken over its tasks. Dan expects his current agent, Boo, to meet the same fate within a year: “The personality thing is cool, but particularly in a work context, I mostly care: Does it do the thing?” he says.
On this week’s episode of The Every Podcast, Dan sits down with Willie Williams, Every’s head of platform and one of the driving forces behind the Every Agent, which launched this week. They discuss how Every went from everyone running their own OpenClaw to relying on one shared company agent—and where they think personal agents belong in everyday life.
Watch on X or YouTube, or listen on Spotify or Apple Podcasts. You can also read the transcript.
Here are the highlights:
- Why a company agent beats a bot for every employee: After Every’s January offsite, the team came home hooked on OpenClaw. Slack quickly filled up with bots that were extensions of employees, like a “growth claw” from Every’s then-head of growth that could answer any team member’s growth questions. Over time, the bots proved hard to maintain, and it was tough to remember which one did what, so most of them died off. Personal agents brought into a workplace start from scratch and need a year to “really get up to speed,” Willie explains. A shared company agent fixes these problems. There’s one bot to keep track of, it learns a company’s workflows, and everyone improves it simply by working with it. That’s the future of agents Willie imagines: “I don’t see the industry repeating the chaos of every worker getting their own AI coworker, where we all double our company size.”
- Personal agents don’t belong in the office: Willie predicts that people will split their agent use between a company agent at work and an agent “that is just yours” for use at home. This is already working for him. His personal setup is a roster of Grok coaches, including a therapist, and one for lap swimming (“I’m a terrible swimmer, which is really funny because I like to surf,” he says). Separate agent personas work for him because they mirror how he’d work with humans, with a coach for each particular goal. His rule of thumb: If a task is “transactional,” it gets a Codex thread. If he’s building a relationship, it gets its own bot. Until recently, a setup like his took real commitment. Now the early drawbacks of personal agents—security, setup, and maintenance—are fading as bigger players and more money arrive. “We’re starting to get closer to a better consumer personal agent experience, and it’s going to be incredibly attractive,” he says.
- Why Willie’s token usage fell off a cliff: Dan has been teasing Willie for ranking low on Every’s internal token leaderboard. The reason for that, Willie says, is that engineering management is still a people job that requires spotting bottlenecks and talking to whoever can clear them. “There’s no billion-dollar—sorry, billion-token—move I can make,” he says. AI helps more with the unglamorous side of Willie’s work. A custom AI feed of Slack threads and customer reports gives him a “field marshal-level view” of the company, and an agent watches dashboards for anomalies and takes the first slot in the on-call rotation. “You have an agent on call as the first person all the time. You don’t have to worry about vacation,” he says. More on that at the end of the newsletter.
This is a must-watch or must-listen for anyone deciding whether to give every employee a personal agent or build one shared agent for the whole company.
Miss an episode? Catch up on Dan’s recent conversations with Sam Altman; Anthropic head of product Mike Krieger; the team that built Claude Code, Cat Wu and Boris Cherny; the team that built Codex, Thibault Sottiaux and Andrew Ambrosino; Vercel cofounder Guillermo Rauch; and others to learn how they use AI to think, create, and relate.
Introducing Attio: the agentic CRM
Transform the way revenue work gets done with Attio. Get agents that build pipeline, convert leads, and run all your sales motions. Your agents track the whole book, so you save the ones slipping and grow the ones rising. Then Ask Attio any question about your business, from the weekly forecast to performance by rep, and get the answer in seconds.
For teams building the next era of revenue.
Inside Every
Cost-saving at the frontier
Every’s business is built on subscriptions, which include access to the Every Agent. The costs incurred by the agent are passed on to users without a markup. Our goal is to make the Every Agent as useful and efficient as possible, so members get the most bang for their buck.
That makes token efficiency more than an engineering nicety: It’s part of the product. And the team’s stabs at making the agent efficient produced useful learnings.
An early attempt routed all tasks assigned to the Every Agent through a cheaper coordinator model—Sonnet—which was instructed to pass complex jobs to Opus, its more expensive counterpart. Not only did the quality and speed of the responses suffer—like most of us, Sonnet isn’t great at identifying its own limitations—but costs went up. The handoff meant requests were often processed by both models.
Engineer Paridhi Agarwal digested the disappointing results and promptly changed course. “We don’t necessarily need a complex engineering system to solve a problem that comes from inefficiencies in our own codebase,” she remembers thinking.
Her instinct was right. Paridhi looked for waste in the system and, through a model upgrade and a handful of targeted changes, reduced token costs across 11 frequent tasks by more than 80 percent.
The lowest-hanging fruit was switching the Every Agent from Opus 5 to Opus 5.5 after its release. Response quality improved, and token costs noticeably dropped “within a couple of hours,” Paridhi says.
Then she started looking at what the Every Agent was being asked to read and reduced the amount of information. Its tools—the connections that let the agent access services like GitHub and Linear—were returning large amounts of data, including internal details it didn’t need. She rewrote them to return shorter, clearly labeled results, while retaining critical information. She also changed how the agent loads tool information. It now starts with just the name and a short description for the most frequently used tools, loading full technical details and other tools only as needed. Together, these changes gave the model less text to process and reduced token costs by roughly 10 percent.
But the biggest discovery was a bug that Paridhi found with help from Opus 5.5. With every new conversation, the Every Agent also receives standing instructions about how to behave—its system prompt. Processing those instructions each time costs money. To lower costs, Claude can temporarily save and reuse the work of processing those instructions, a practice called caching. But the text must stay identical.
The team assumed caching was working. It wasn’t.
Claude automatically adds memory folder names to the system prompt, and Every’s included each user’s Slack ID—an easy detail to miss. Those names made the prompt different for every user, preventing the cached version from being reused across users.
The fix was straightforward: Paridhi gave every memory folder the same generic name and kept Slack IDs where the model couldn’t see them. Each user still had access only to their own memory, but caching now worked as intended. In her best-case test, this cut token costs by 74 percent, on top of the savings from the tool redesign and model upgrade.
“Sometimes the solution is simpler than you think,” she says.
Steal this workflow
Make your AI more efficient
Opus 5.5 served as Paridhi’s coding partner, helping her investigate why the Every Agent cost so much to run. She supplied context, questioned the test results, and directed it to investigate unexplained costs. Here’s how to use that approach to fix inefficiencies in your own software.
1. Explain what you want to improve and give your agent the context it needs to understand the issue. Start with a specific goal—cost, speed, or repetitive work—and give your agent access to any relevant code, documentation, and tools. Paridhi asks the agent to read the codebase and explain its diagnosis of the problem before proposing changes. She also consults current documentation and supplies details about new features, which can ship faster than a model’s training data can keep up with.
2. Develop possible fixes. Ask the agent where your software might be doing unnecessary work and have it suggest changes, with explanations for how each potential fix would help. Instruct it to check whether the models and services you already use offer features that could address the problem before proposing something built from scratch.
Bring your own ideas, too. Paridhi knew about a Claude feature called “deferred tool loading”: Instead of giving the agent instructions for every tool at every step, the feature automatically loads instructions for just the most frequently used tools, and looks up the other tools only when a task calls for them. She asked the model to consider the feature as a way to cut costs.
3. Measure whether a proposed fix helps. Have the agent run the same tasks before and after one change, then show you the difference in cost or time. Make sure to compare the quality of results: A cheaper or faster response isn’t an improvement if the responses themselves get worse.
Investigating a possible fix can uncover a bigger problem. Paridhi asked Opus 5.5 to test whether loading fewer tool instructions would make conversations cheaper. The savings, it turns out, were small, but the tests exposed a much bigger expense: The Every Agent was repeatedly saving instructions but not reusing them. She asked Opus to investigate, which led her to identify the caching bug.
4. Investigate any remaining problems. Ask the agent to reexamine the code and test results. What is still expensive, slow, or repetitive—and why? Have it trace the cause and suggest a fix, then repeat the tests to check whether the change helps.
The daily driver
The models the team is using this week:
- Nityesh Agarwal, senior applied AI engineer—Opus 5.5: “I even stopped using Fable 5.1 completely and now I’m rarely running out of usage.”
- Randy Counsman, head of video—GPT 6.1 Sol (high), switching to Astra for harder tasks and Opus 5.5 for anything design- or image-related. A major caveat: “I feel like I just ‘use Dot’ now... I had to go and check my history to see which models I was using for certain tasks. Otherwise I don’t think about it and just talk to Dot.”
- Becky Isjwara, head of social media—Opus 5.5 (medium) for video and writing; GPT-6.1 Sol (medium) for everything else, including computer use, scheduling posts, and everyday admin. “Astra was cool, but too slow for me,” she says.
- Tyler Nishida, designer—Opus 5.5 (max) and Sonnet 5.5 (medium).
- Katie Parrott, staff writer—A “seeeeecret model” she’s early-testing is her number one, followed by Astra, Opus 5.5, and Sol (GPT-6, then GPT-6.1). “Overall, I’m splitting 55 percent to 45 percent Claude vs. GPT at the moment, which reflects the ‘Claude for creative, OAI for admin’ split that I’ve been in since the 5.5 class started rolling out from Anthropic and won my heart back for writing.”
- Arielle Shipper, head of operations—Sol (medium) remains her default—“I am a creature of habit and don’t like fixing what isn’t broken”—although she now toggles to Opus 5.5 for creative work.
- Dan Shipper, CEO—An unreleased model he’s early-testing dominates his usage, followed by Astra and GPT-6.1 Sol.
- Loren Stewart, engineer—Opus 5.5 for coding and thinking. “I love it. It’s such a good model for... everything I do!” He’s started to experiment with GPT-6.1 Sol but hasn’t used it enough to have a firm opinion yet.
- Mike Taylor, head of evals—Opus 5.5 (medium). His Astra usage took a recent nosedive given he had fewer computer-use-heavy admin tasks and has moved all new threads to Claude.
Thesis Statements
In August, we launched Thesis Statements, a collection of specific, contestable claims from builders and thinkers about the future of great human work with AI.
This week, we have seven more predictions from people at the frontier:
- The arguments we hid behind shared vocabulary will become explicit by Alex Duffy, cofounder and CEO of Good Start Labs
- Managing decision fatigue will become part of making great work by Laura Entis, staff writer at Every
- You’ll give your Legos away to agents by Brandon Gell, chief operating officer at Every
- Nothing will go wrong twice by Kieran Klaassen, creator of Compound Engineering and member of the frontier team at Every
- Describing a problem will be enough to begin solving it by Naveen Naidu, general manager of Monologue
- Your expertise will belong to your employer—unless you can make it portable by Kaushik Viswanath, managing editor at Every
- History will be rendered, not read by Willie Williams, head of platform at Every
If you want to help decide what matters in the future of AI and human work, think creatively, and build what comes next, join us at our inaugural Thesis: 2027 conference in Brooklyn on November 5.
Jagged frontier
The guide to being an AI-first engineering manager
Any manager will tell you that one of the hardest parts of the job is just understanding what the hell is going on half the time.
Managers typically operate at one to 10-plus levels of remove from the actual work—a situation I once heard described as trying to push a ping-pong ball into a cup using only a long stick made of chopsticks held together with bubble gum (with a chopstick added for every layer).
Information coming back to you travels that same distance, affected by forgetfulness, ego, confusion, and, in the worst cases, malice. To mitigate this, we’ve come up with managerial infrastructure like standups, staff meetings, and 1:1s to try and sort out who did what, who said what, and what needs to happen now.
AI now allows us to skip the age-old game of organizational telephone, and go directly to the source by setting up a feed.
The first and most foundational step is to ensure all meetings are automatically transcribed. This has become the social norm at work because it helps individuals keep track of action items and details, but transcription is becoming indispensable at an organizational level, too. At Every, we transcribe all our meetings, even when we’re together in person. It’s become second nature.
The next step is to attach an AI to this feed of meeting transcripts your company is generating. We tend to use Tend (har har) to do the attaching, but even a local automation in Codex will do. The automation should scan the feed and extract anything you deem to be relevant, like key updates and action items, and store it for your later review.
This simple setup is powerful and flexible enough to give an engineering manager ongoing insight into the team’s operations, allowing them to perform continuous conflict resolution and coaching by scanning these updates. If you expand the feed to include the stream of pull requests being created, reviewed, and merged in GitHub, you can even get an updating real-time picture of engineering progress. The end result is a persistent up-to-date picture of everything happening above, below, and around you.
The quality of your decisions as a manager depends on the quality of the information reaching you—and a feed like this gives you more of the original information, with fewer layers of interpretation.
You’re still holding the same ridiculous stick made of chopsticks and bubble gum. But it’s a lot easier to decide where to push when you can see what’s happening at the other end.—Willie Williams
Laura Entis is a staff writer at Every. To read more essays like this, subscribe to Every, and follow us on X at @every and on LinkedIn.
Stop explaining AI to your team and start showing them. Every Agent does the work in Slack, where everyone can see it, so one good workflow quickly becomes the whole team’s.


