Earlier this week, we launched the Every Agent, a coworker that lives in your company’s Slack. In her first piece for Every, engineer Paridhi Agarwal explains how we built it on Anthropic’s Claude Managed Agents, why we stopped running our own fleet of agents, and the trade-offs we accepted along the way. If you want to see what a shared agent can do for your own team, join us tomorrow for our camp on working with agents in Slack.—Kate Lee
Was this newsletter forwarded to you? Sign up to get it in your inbox.
On a Friday in September, Arielle Shipper’s salad went missing from the office fridge. Douglas Brundage did what most of us at Every now do with a problem. He asked @Every in Slack.
“who stole Arielle’s salad!?!?! make a list of likely culprits”
A minute later it posted a ranked list of suspects, starting with the people it knew had been in the kitchen that afternoon. After every named employee had made their case, @Every zeroed in on Leo, the office dog: Everyone else needed a motive. He only needed an unattended bowl. Nobody found the salad.
Over the past five months, we’ve been building a coworker that lives in your Slack, keeps up with what’s happening around the office, and joins in when someone asks it to investigate a missing lunch. That’s Every Agent, which we launched on October 6. In Slack it goes by @Every.
Every Agent runs on Claude Managed Agents (CMA), a service Anthropic launched in April. It runs the machinery an agent needs to think, remember, and pick up where it left off. That let a small team spend its time on the parts you see in Slack. We didn’t know any of this when we started. We didn’t set out to build one agent for the whole company, nor did we plan to hand its machinery to someone else to run.
Build an app that works for every age in one day at Google
On October 21, Google and k-ID host the Build for Everyone Hackathon at Google San Francisco. Bring an idea for an app or game. New tooling from k-ID, revealed that morning, adapts it to whoever is using it, so a child, a teen, and an adult each get the right experience automatically. Fifteen teams get one day to ship, with mentors from Google, k-ID, and partners on hand to help before you demo to the judges. Use Gemini and Google Cloud or any tools you like. It’s free for builders 18+. Applications close Monday, October 12.
Our first agents were pets
Every is a small team that tries to do a lot, so we’re always asking how to get more done with the people we have. When OpenClaw, an open-source agent program, went public late last year, agents started to look like the answer. What if each employee at Every had an agent working alongside them, doubling what each of us could get done? Plus One was our attempt to find out. As Brandon Gell and Willie Williams wrote in May, we gave each Every employee their own agent, running OpenClaw on a rented virtual computer.
That left us managing dozens of these cloud machines around the clock, checking every 30 seconds that the agents were still running and fixing them when they crashed. (We built a command called /heal-bot to walk us through repairs.) When a bot’s ChatGPT sign-in expired, it went quiet until someone signed it back in by hand. A separate job ran every hour just to keep Slack’s access tokens from expiring.
OpenClaw also wasn’t built for a team. It reset conversations at 4 a.m. each day, and its memory systems overlapped in ways that were hard to untangle, as Nityesh Agarwal wrote in June.
By spring we had learned that most people don’t want to look after an agent, even one that’s theirs. And we didn’t want to look after a fleet of them, either. Engineers like to say servers should be cattle, not pets: interchangeable and replaced when they fail instead of nursed back to health. Plus One had given each of us a high-maintenance pet and tasked our engineers with keeping dozens of them alive.
We wanted to build a coworker, not infrastructure
We soon realized that—aside from the issues that came with running our own infrastructure—personal agents had a ceiling. Each agent only knew its owner’s work, so it missed what the rest of the company was doing and how that affected you. When someone figured out a good workflow, it stayed with their agent. Teaching every agent the same thing meant making the same update several times, and when the agent broke, its owner had to fix it.
Meanwhile, the threads where the whole team pulled in any single agent were the ones that worked best. This led us to the vision of one agent shared by the whole company: It sees the whole team’s context, a skill it learns from one person works for everyone, and we only have to keep one agent running.
Moving to one shared agent meant fewer agents to maintain, but we still had to decide who would keep it running. Dan Shipper thought hosting a single agent ourselves would be manageable. Willie wasn’t convinced. We tried it on our own servers, and within a month, the maintenance work it required had convinced Dan, too. We wanted to build a coworker, not run an infrastructure team. In early May, we wrote a plan to rebuild on Claude Managed Agents, and within a week, @Every was answering in Slack, and the old agent fleet was gone.
An AI model generates responses. To act on them—searching files, running code, or sending a message—it needs a harness: the software that runs the tools it requests and feeds back the results. Claude Code has one, but it’s built to run on one person’s computer. OpenClaw had one, too, and Plus One taught us what it costs to host a harness yourself: Everything around it becomes your job.
With CMA, Anthropic runs the harness and handles much of the maintenance we’d been doing for Plus One. Each session—the conversation and the work the agent does within it—gets a virtual computer that costs 8 cents an hour while it’s running. CMA also handles caching, which lets the model reuse information it has already processed, and compaction, which summarizes long conversations so the model can keep working. It also lets us interrupt the agent mid-task. With Plus One, each of those was our job.
Anthropic’s engineers describe the design as separating the brain from the hands. Claude and the harness are the brain. The computer and the tools are the hands. The session sits outside both, as a running log of everything that happened, so the work survives if either one crashes.
For us, each Slack thread gets its own session, so @Every picks a thread back up right where it left off. Memory carries across threads: folders of notes the agent reads at the start of each session and adds to as it learns.
What @Every does all day
If Anthropic runs the infrastructure, what’s left for us to build? Everything that makes @Every feel like a coworker instead of a chatbot. A chatbot answers whoever talks to it. A coworker in Slack has to know when a message is meant for it and on whose behalf it’s acting. Those two problems took most of our design work. (For a tour of what @Every does day to day, read Dan’s launch post.)
Tag @Every, and it replies, whether in a direct message—where it works as a private assistant—or in a channel thread anyone can join, the way half the office did on the salad case. When a thread it’s in keeps going, @Every reads the room and decides whether a message was meant for it, merits a quick emoji, or belongs to the humans. A small, fast model reads each untagged message first, along with who’s in the thread and what @Every said last, and decides whether to reply, react with an emoji, or stay out. When in doubt, it stays quiet.
When @Every needs to read information or take an action in another app on behalf of a user, CMA pauses the agent and hands the request to our servers. Our servers make the call with the user’s own login and send back only the result. The agent never holds a password or access token, so it can’t use one for anything the asker didn’t request. Skills that the agent has follow the same rule: Anyone can ask @Every to change one, but nothing changes until the skill owner approves.
The trade-offs we accepted
Although Claude Managed Agents took whole categories of problems off our plate, building a coworker on top of it still had plenty of challenges.
The first was one we caused ourselves because of the design we’d chosen. Anthropic keeps the agent’s session running even when our servers restart. But when we deployed an update while the agent was working, we sometimes shut down a server before it could respond to the agent’s request. The agent kept waiting for a response that would never arrive. New messages failed, and trying to interrupt the agent didn’t unstick it. To get it working again, we had to send responses to all its outstanding requests. It’s an ordinary bug once you see it. Every team that runs its own code next to a hosted agent will encounter some version of it, and we hit it every time we shipped while someone was mid-task.
Then, in late July we switched models to Opus 5, and @Every got wordy overnight. The typical reply went from about 280 characters to about 880, and people started calling its answers essays. Teammates took to asking it to try again “with half as many words, like a normal person.” We spent weeks writing rules and tests for brevity and formatting. When Anthropic released Opus 5.5 in September, its announcement called writing “one of the most common areas of feedback we heard about Opus 5,” and once we switched, the essays stopped.
Our longest conversations caused two more problems. One customer’s DM ran as a single session for three days and 33 exchanges. Each time the cache expired between messages, the model had to write the whole history back into the cache, which costs more than reading it fresh. And near the model’s context limit (the most text it can consider at once), Opus sometimes returned nothing at all. The solution for both was to start a fresh session well before either point. Now a DM that goes quiet gets a new session with a short recap of the old one. When we reran past conversations with the fix in place, total costs fell 39 percent, and not one exchange cost more than it had before. (You can read more about our cost-reduction work, including fixing a caching bug that had us paying full price on every conversation, in Context Window.)
We could fix those problems ourselves. The trade-offs that come with CMA are harder to work around, and more than once they had us looking for alternatives. The limitation we feel most is that CMA only runs Claude. We can pick a different Claude model for each session, so a quick question goes to a faster, cheaper model, but there is no way to route to models from other labs that might be even cheaper or more appropriate to the task. In addition to being limited to Claude models, we’re tied to the design of CMA. And because it keeps each session’s history so it can resume work, it isn’t eligible for Anthropic’s zero-data-retention option. Some companies require that, and for now we can’t offer it.
Eventually I spun up an architecture document listing everything we’d have to replace to move off CMA. It ran longer than any of us expected. Replacing CMA’s infrastructure would mean rebuilding the harness, sessions, memory, tools, and sandboxes we’d been glad to hand off and would take months, not to mention dealing with a whole range of new bugs.
What’s next
By staying on Claude Managed Agents, we spend our time on @Every instead of the infrastructure underneath it. Today it’s a coworker you call on when you need something. We want it to be one you’d miss if it were gone. Most of that is memory that gets sharper the longer it works with your team, so it learns how each person likes to work without being told twice. We also want @Every to work with you across multiple surfaces, so it feels more like working with the person who sits next to you.
And like any good coworker, it doesn’t let things drop. Arielle’s salad is still missing, and as far as @Every is concerned, Leo is still the lead suspect.
Paridhi Agarwal is an engineer at Every working on the Every Agent.
Stop explaining AI to your team and start showing them. Every Agent does the work in Slack, where everyone can see it, so one good workflow quickly becomes the whole team’s.








