Was this newsletter forwarded to you? Sign up to get it in your inbox.
One morning in early July, I woke to a flood of alerts from OpenAI and Ramp: You’re out of credits. Your card has been declined. I began to sweat. At 10 p.m. the night before, our credit balance was full, auto-reload was on, and our Ramp card had plenty of available funds. Somehow, less than 12 hours later, our account was zeroed out.
That squall turned out to be my brother, Every CEO Dan Shipper, testing Sol (ultra) on tasks designed for a senior engineer for that day’s Vibe Check of the model. By 10 a.m., we’d restocked the Ramp card with new funds, and I thought the storm had subsided. But our token spending stayed unusually high for the rest of the day, and we haven’t had a normal day since.
In the first five full days after GPT-5.6 Sol rolled out, our daily credit usage rose from 11,520 to 26,685 credits—almost 2.5 times our previous-week baseline. I had spent weeks bracing for Fable to blow up our budget, but Sol was the storm I didn’t see coming. Suddenly, I had to figure out how to enable daily work without bankrupting us. My colleagues’ reactions ran the gamut from “Let ‘er rip and let’s see where it lands at the end of the month!” to “We gotta impose limits now, and we should start exploring running our own models locally.” Meanwhile, we were burning through a month’s worth of token spend every few days.
There was no easy solution. Experimentation is part of everyone’s job at Every; from engineering to business development, we all need to learn what these models can do and where they’re useful. Because valuable insights can come from anywhere in our company, everyone needs to be able to spend. Unlike many companies, we don’t go all-in on one model. People use whichever works best for the job, which makes our costs harder to predict as new models with different strengths and pricing structures come out. Any intricate allocation scheme I devised would be obsolete in days, if not hours.
With input from the team and a lot of thought, I put in place a loose operational process instead of strict spending policies—and for now, it’s working.
Introducing Attio: the agentic CRM
Transform the way revenue work gets done with Attio. Get agents that build pipeline, convert leads, and run all your sales motions. Your agents track the whole book, so you save the ones slipping and grow the ones rising.
Then Ask Attio any question about your business, from the weekly forecast to performance by rep, and get the answer in seconds.
For teams building the next era of revenue.
New rules for a new world
Before Every, I spent eight years building out operations for a startup. As COO at Donut, a platform that helps companies onboard, connect, and engage employees, I was responsible for designing policies and processes that could withstand change. I was reasonably certain that when I made a decision about a workflow or budget, it could last for a quarter or even a year. When the ground shifted under my feet, I’d react with a simple amendment. But over the past six months, the way that tech companies work has changed drastically. And Every feels these changes especially early. Any rule I design for today’s conditions may be completely wrong by tomorrow.
I’d thought the sea change was a “me” problem at first. Even before the Sol fiasco, I went to my first conference in nearly a year, hoping to be enlightened by decades of accumulated wisdom and comforted by prescriptive best practices I could burn like a vintage CD and bring back to Every to solve all of my stressful 30-person-company-problems. Instead, one company leader told the audience they now give employees additional equity every year instead of the previous industry standard of every four years (or not at all). A human resources executive said they were adding token budgets to compensation packages without reliable standards for how much to offer—because there were none. I wasn’t alone in struggling to build out operational guidance for AI-native teams.
The old startup modus operandi was “We’ll decide now and revisit next quarter.” In this environment, we can’t even say that to ourselves anymore. Policies are durable as long as the world they address persists, and right now, the world can change in the time it takes to run a prompt. No matter the size, stage, or maturity of the company, we’re all figuring this out in real time.
I went to the conference for answers but came up empty-handed when my billion-token question arose days later: How could I create a token spending policy that fit Every?
Parameters, not policies
My constraints were real: We have a finite amount of money, but we also have a finite amount of the team’s time, and we have to move fast to stay at the frontier—and report on it, too. Hard spending limits could stop Vibe Check benchmarks mid-run or prevent engineers from building key infrastructure for Every Agent and developing new ways to code with agents. Telling people to always use the cheapest model possible means guessing whether it will be good enough—and you can’t know whether a different model would have produced a better result. So I let go of tight spend controls and focused instead on codifying loose guidelines to help us make spending decisions in real time.
My parameters start with this aphorism: Responsible usage and cheap usage are not the same thing—nor are high spend and waste. When I see big expenditures of credits or a newly minted token billionaire on our leaderboard, I try to react with curiosity rather than a hard limit. I reach out on Slack to get more context. I ask three questions:
- What did it cost?
- What did it buy us?
- What did we learn?
One teammate spent $480 in a single run to build the permissions structure for a product we’re launching. That run was worth it; it stood up necessary infrastructure for a revenue-generating product, taught the engineer techniques that made subsequent runs more efficient, and produced an insight our editorial team could publish. Another teammate spent roughly the same amount in one day having Tend, Dan’s open-source productivity tool, check email and Slack every 15 minutes. That one wasn’t worth it; we changed the cadence that day. Same spend, different answer.
My process is still case-by-case—and I expect it will be for a while. But after three months of navigating token spend, I’ve learned four lessons for companies of our shape and size:
Establish a circuit breaker. In our case, that’s a Ramp card limit and Slack notifications that keep total spending visible. When the card runs out, we choose whether to refill it and by how much. When we hit the limit, we pause and decide whether the work is worth what it will cost to continue.
Make usage visible. We gave everyone access to ChatGPT’s usage dashboard so they can see usage in real time, including which team members are consuming the most credits. Visibility turns spend from a month-end surprise into a team learning loop. The engineer behind the $480 permissions run found ways to make future runs more efficient. The Tend owner reduced how often it checked. When the cost, output, and lesson are visible together, expensive doesn’t automatically mean irresponsible.
Earn the friction. Policies and decision gates introduce friction when a team needs to move quickly. I see it as operations’ job to identify risk and demonstrate that a lighter guardrail can’t work before imposing heavier restrictions. We give the team the tools and resources they need, and we expect them to act responsibly in return.
Make change as the evidence changes. If we ever see the team slipping on their end of the bargain—usage nobody can explain, lessons nobody shares—we’ll roll back the autonomy and put a platform-level limit in place. We haven’t had to yet. We’re a 30-person company operating at the frontier, so we can afford to trust our team and react quickly for now. A larger or more regulated company may need more checks and limits sooner. The broader lesson is to introduce friction only after you’ve seen the problem it’s meant to prevent.
Where have these loose parameters gotten us? We spent $31,300 on OpenAI credits in July—$26,800 of that after the Sol fire drill. Our token spend is high. But that cost is still lower than the cost of hiring enough people to produce the same work: launching All Access and Builder Pack, planning Thesis, building Every Agent, testing new models, shipping high-quality content every day, and doing many less-visible but still-critical things to move the business forward with a lean team. We’re keeping that increased limit for August because we can afford it and because it supports the work our team needs to do.
In the meantime, I’m back on Slack pinging my brother for answers. He spent $2,000 last night running Sol (ultra) on a nonessential task.
Arielle Shipper is the head of operations at Every. Previously she was the COO at Donut and began her career in editorial at Condé Nast.
Everyone’s a builder now. Every All Access gets you the full membership plus the Builder Pack—$9,000+ in credits for the tools we build with.
