Astra and Fable 5.1 can ace a graduate-level science exam, but no public benchmark will tell you whether a model knows where you’d put a comma or how many ideas belong on a slide. Today, Every CEO Dan Shipper explains why we’re building a personal benchmark for every employee, head of evals Mike Taylor shares how a run on his own benchmark convinced him a smaller model could handle much of his daily work, and we offer a five-step workflow for turning the corrections you already give AI into checks for grading any model.
Was this newsletter forwarded to you? Sign up to get it in your inbox.
Inside Every
Astra and Fable 5 scored 96 percent and 93 percent, respectively, on a test of graduate-level science questions. Impressive, clearly! What’s less clear from public benchmarks is how those scores translate to what most of us care about: how well these models help us do our jobs.
The solution, says Dan, is to create a personal benchmark—a set of custom evals that tests how well a model does specific parts of your work, graded against your own standard for what good looks like.
We’re building them for every employee. Editor in chief Kate Lee’s benchmark captures her copy-editing judgment, with rules on everything from comma placement to word choice. Mike’s benchmark for creating decks includes a simple preference: one idea per slide.
We all have preferences like these; a personal benchmark turns them into rules a model can be tested against, showing you what AI can reliably execute and where it still needs your intervention. Dan describes this as a way to “push yourself up a level.” Instead of correcting outputs one by one, you design a system for how you want the work done.
That system can improve over time, too. The rules you use to grade a task can become skills that improve future outputs. When a model falls short, you diagnose the failure, revise the skill, and rerun the eval. Failures become feedback.
“It’s the first step to working in a compounding loop,” Dan says.
How this works in practice
To build a personal benchmark for each of us, Mike and senior applied AI engineer Nityesh Agarwal collect five to 10 tasks a person regularly hands to AI, along with the prompts they use and any source documents.
They run the tasks on different models, then ask the person to talk them through the outputs: What does the person like? What needs fixing, and why? The feedback will become checks for evaluating future outputs, so the more specific it is, the better.
One task in Mike’s own benchmark is to generate a course feedback dashboard: a web page where he can browse and analyze participants’ ratings and comments.
Some checks are programmatic: Does the page load? Do the calculated scores match the source data? Others are subjective. An AI judge evaluates whether the design avoids the dark-mode aesthetic Mike dislikes.
Every check passes or fails.
The checks need testing, too. If you disagree with an eval score, Mike says, you tune the benchmark. The problem might be that “we don’t have enough checks yet, or we’ve got the wrong definitions for the checks we’re running.”
Once you and the AI judge reliably agree, the benchmark can tell you which model is best for each task. When a new one arrives, one run shows where the new model beats your current choice.
Sometimes, that changes your mind. When Nityesh ran Mike’s benchmark, OpenAI’s GPT-5.6 Luna outperformed Fable and GPT-5.6 Sol. Mike inspected the results and agreed: The smaller model had performed well.
That suggested it could handle many of his day-to-day tasks—and prompted him to expand his benchmark to include harder assignments.
Steal this workflow
Turn corrections into checks
Every time you ask AI to edit a paragraph, redesign a slide, or rewrite code, you’re making a judgment call. A personal benchmark turns those judgments into checks you can use to compare models.
Here’s a lightweight way to start.
1. Pick a recurring task
Look through your past AI chats for tasks you do often and mistakes the model keeps making. If nothing comes to mind, ask Codex or Claude Code:
2. Use it as a test
If the task is to make a presentation, save the prompt, and source files for one deck. Run the task in a new thread. Save the first output before you request any changes, and record the model and its reasoning setting.
You now have a fixed test you can run on any model. Keep the instructions and source material the same each time.
3. Turn your feedback into a checklist
Review the AI’s work as you normally would. Type or dictate what you like, dislike, and want changed. For a presentation, you might say:
Ask the AI to turn your feedback into separate yes-or-no questions, one requirement per check:
- Does each slide contain only one main idea?
- Is all body text at least the minimum size specified in my style guide?
- Does the deck use my specified font?
Edit the questions to reflect your standards. Each should identify something specific the output needs to get right.
4. See whether AI can reliably apply your checks
First, grade the output yourself. If slide three combines the budget, timeline, and hiring plan, the “one idea per slide” check fails.
Next, open a new conversation and provide the presentation, the checklist, and any reference files, such as your style guide. Keep your grades to yourself and ask:
Compare its answers with yours. Did it catch the multi-idea slide? Did it flag something you missed? Where you disagree, examine the evidence: Either the AI got it wrong or the check needs more specific wording.
5. Resolve disagreements and test again
If your benchmark passes a presentation you wouldn’t use, figure out why. Common issues include:
- An unclear check: “Slides are readable” might need to specify a minimum font size.
- A missing check: The slides look good but leave out a requested topic. Add a check that all requested topics are covered.
- A grading mistake: The check is clear, but the AI missed an error.
Regrade after every change, and apply the revised checklist to every output you’re comparing. Make sure you’re not tuning the rules to favor one model.
Finally, save the checklist with the assignment and run it on a different example. That tells you whether it captures your standards or just the problems in the first output.
Thesis Statements
Three weeks ago, we launched Thesis Statements, a collection of specific, contestable claims from builders and thinkers about the future of great human work with AI.
This week, we have seven more predictions from people at the frontier:
- “Powered by AI” will mean cheap by ÌníOlúwa Abíódún, lead product designer at AirOps
- We’ll move from a knowledge economy to a discovery economy by Matt Cynamon, media at Union Square Ventures
- We’ll pay to watch humans make things by Natalie Fratto, founder of Charts & Crafts
- When AI can argue for anything, trust your gut by Tim Fu, founder of Studio Tim Fu
- Institutions won’t need a crisis to change by Sahil Lavingia, founder of Gumroad
- Every good idea will have 1,000 clones by morning by Mike Taylor, head of evals at Every
- Certainty will be the trap by Alex Tryon, founder of Dewey
If you want to help decide what matters in the future of AI and human work, think creatively, and build what comes next, join us at our inaugural Thesis: 2027 conference on November 5, 2026.
The daily driver
The models the team is using this week
- Nityesh: Fable 5.1 (medium) as his daily driver for its balance of thinking effort and cost.
- Douglas Brundage, head of marketing: Astra (light), switching to high for asset creation, plus GPT-5.6 Sol for simpler tasks.
- Andrey Galko, engineering lead: Opus 5 for smaller tasks and Fable 5.1 for larger projects. “Still trying to understand what Astra is capable of but can’t really say I’m impressed.”
- Becky Isjwara, head of social media: Primarily Astra, with Fable for video and Astra for computer use. She also uses Grok Bot through Astra’s computer use for X research.
- Waqqas Mir, customer support specialist: Back to using Opus 4.8 (high).
- Katie Parrott, staff writer: “I have been Astra maxxing and I’m afraid for my bank account.” She finds Fable “more psychologically insightful as an interviewer and rigorous as a planner,” but has gotten hooked on Astra because its writing is more fluid and surprising.
- Natalia Quintero, head of consulting: Fable 5 for writing and thinking after burning through Fable 5.1 credits too quickly, and Astra (medium) for everything else. She avoids Astra at higher effort because it starts “spiraling and overthinking so much its work is unproductive. Spiraling and overthinking is my job, Chat!”
- Arielle Shipper, head of operations: GPT-5.6 Sol (medium), occasionally switching to Astra (light). “I am officially taking Sonnet 5 off my list of models because it took 12 minutes to compare two Google Docs—and still did not arrive at an answer. Sol did the task in 42 seconds.”
- Dan: Mostly Fable 5.1 and Astra, with a smattering of GPT-5.5, GPT-5.6 Sol, and Opus 5.
- Loren Stewart, engineer: Astra (medium) and GPT-5.6 Sol (medium or high) for coding; Fable 5.1 for thinking and discussing ideas. He uses Fable and Astra to critique his writing.
- Mike: Primarily GPT-5.6 Sol, reserving Astra and Fable for bigger tasks because of cost. That said, “Astra’s computer use is phenomenal.” The model took screenshots of items in his browser and on his computer and used them to make hundreds of small edits to a slide deck, requiring “almost no correction.” He recently compared Fable 5.1 and Astra for writing, and concluded, “Claude is still the better writer.”
- Willie Williams, head of platform: Mostly Astra and Fable 5.1. “Grok Bot fell off, but I feel the need to bring it back.”
Straight from Slack
Show us your agent setup
We’re turning our Friday internal show-and-tell into a tour of how people at Every use Codex and Claude Code. Over the next few weeks, teammates will open their setups and folder systems to show how they get work done. As we wrote in “A Codex of One’s Own,” those setups can be as individual as the people using them. Watch this space for reporting on everyone’s different styles—and ideas to borrow for your own.
Laura Entis is a staff writer at Every. To read more essays like this, subscribe to Every, and follow us on X at @every and on LinkedIn.
Everyone’s a builder now. Every All Access gets you the full membership plus the Builder Pack—$9,000+ in credits for the tools we build with.





