Was this newsletter forwarded to you? Sign up to get it in your inbox.
It is a well-known fact that the Germans have a word for everything. Kummerspeck, literally “grief bacon,” is the weight you put on from emotional eating. Torschlusspanik, “gate-closing panic,” is the fear that time is running out to do the thing you’re supposed to do with your life.
Recently, I came across another gem: Verschlimmbesserung, or “worsen-bettering.” It refers to our desire to make something better, only to make it worse. Used in a sentence: “Ich habe meinen Kontext verschlimmbessert,” or “I have worsen-bettered my context.”
That’s precisely what I did to the context that drives my AI setup over a few weeks in September. By “context,” I’m referring to the style guides, examples, and instructions that tell Claude and ChatGPT who I am, how I write, and what I’m working on. When my context is good, my AI “assistants” can find what they need to do work the way I want them to do it—including writing essay drafts the way I want them written.
But a few weeks ago, my context went bad, and so did my work. Drafts came back crowded and flat. When I told the model what was wrong, the next draft came back worse. The system I’d built to keep me afloat started pulling me down with it. For a while, I stopped trusting my context. Then I stopped trusting myself.
You don’t need an elaborate setup to run into this problem. You might have asked your AI to keep answers short, then found yourself fighting it for a thorough explanation, or saved a successful draft as an example and watched its quirks turn up in everything that followed. A preference that helps with one task can get in the way of the next, especially when it has to coexist with everything else you’ve ever asked the model to remember.
If you haven’t built your own context yet, consider this a cautionary tale. If you have, and you’ve watched yours get unruly, I’ll show you how I figured out what went wrong, how I tore the whole thing down and rebuilt it, and what principles I’m using now to keep my guidelines from hardening back into rules.
How my context worked, back when it worked
Before I get to the wreckage, I want you to see the thing when it was working, because it worked beautifully. That’s what makes the rest of this story so annoying.
The arrangement was simple. Every column I wrote (Working Overtime, Vibe Check, Context Window) got its own folder, and every folder held the same three things: a voice guide for how my sentences should sound, a style guide for what the column is for and how a piece gets written, and drafts that showed the model the writing I was after. Above them sat my Command Center folder, with the high-level facts about me, my job, and my preferences, and a map to everything else.
When I start a task in one of those folders, the model first reads a short instructions file, called AGENTS.md, that tells it how to behave and what else to read—for example, in my writing folders, the STYLE.md and VOICE.md that tell the models how to write like me. Then it reads those things and treats them as part of the assignment. If the instructions say read it, the model’s going to read it.
Chefs call this kind of preparation mise en place: everything chopped, portioned, and within reach before the first ticket prints. This was mine, and it let me cook everything from personal essays to news roundups to in-depth model reviews, all (for the most part) without dropping any plates.
I had a simple system that I understood, and it was producing the fastest work of my career. There was nothing to fix. Naturally, I decided to fix it.
Three ways I made a working system ‘even better’
Dan Shipper has a saying: Never make a major life decision within 30 days of a meditation retreat, a psychedelic experience, or an encounter with a new frontier model. I’d consider rearranging my context a major life decision, and I did not wait the requisite 30 days following the release of GPT-6 Astra before I started making decisions that would turn out to have a big impact on how I work. I looked at my tidy little folders and asked the question that has ruined many a good thing: How can I make this even better?
I came up with three answers, each of which I was convinced was the next big thing in context engineering. Each one also changed what the model read before it started working, which is the part I didn’t think through.
First, I decided to save everything. If AI was helpful when it remembered some things, how much more useful would it be if it remembered everything? I added two lines to my instructions: “Every experiment must leave a durable record… Failed, abandoned, interrupted, inconclusive, and superseded attempts count.” And: “Repeated mistakes should become review checklist items.”
My goal had been for AI to be able to reference past outputs if they became relevant in a future conversation. But in practice, every essay got a running log of everything that happened to it, and the model read that log before it went back to work. Every note I gave went into the log. Any note I gave twice became a checklist item that every later draft had to pass.
Next, I decided to connect everything. Around this time, I discovered the Zettelkasten, a method of note-taking in which every idea links to every related idea. So naturally, I added that to the list of ideas I applied to my context. I linked a Working Overtime idea to a Source Code idea and a Context Window draft to a Vibe Check template, until 385 files pointed at files in other folders on my desktop.
And finally, I told AI to codify everything. Every pattern I noticed and every observation that felt like it should be a principle went into my style and voice guides. A good opening became a template for that type of opening. The style guide for this column alone grew from around 759 words to 3,855, with eight approved ways to begin, six approved ways to end, and two checklists. It was supposed to describe how I write, but it had become a list of impossible, contradictory, self-defeating demands that every essay had to meet—and that no one essay could reasonably satisfy.
Remember how this works: Before it writes a word, the model reads the instructions file and everything that file points to, and it treats all of it as part of the assignment. Before September, the list a model had to go through was a short voice guide, a short style guide, and a few drafts. Now it included a 3,855-word style guide full of approved openings and endings, a running log of every experiment and correction on the piece in progress, and whatever those 385 links invited it to read in other folders.
And the stack grew every day. Each time I corrected a draft, the correction went into the log. A second line in my instructions, “Repeated mistakes should become review checklist items,” turned any note I gave twice into a test that every later draft had to pass. The model couldn’t tell a record from a rule, so it followed all of them. I’d built a system that turned my reaction to one draft into law for every draft after it.
The only thing that had changed was my context
The thing that makes bad context so hard to catch is that you never see the pile of rules your model is working through. You see the draft it hands back, and a bad draft looks like the model got dumber, or you did.
I found out while writing the previous edition of this column, an essay about what playing with AI taught me about my work. The drafts fought me from the start. They read as if a committee had written them, with every point I’d ever raised wedged in and none given enough room to make sense. One night, I told the model its outline was “UNRECOGNIZABLE relative to the instructions that the skill has.” Another night: “here I am yet again explaining stuff to you that you need to be able to find in the context.” Each time, the next version was worse.
By the end, the folder for that one essay held 127 drafts and more than 1 million words, roughly the entire Harry Potter series, produced for an essay you could read on your lunch break.
I knew I couldn’t keep working like this, and I’d started to suspect my context was to blame. Late one night, I told Codex, running GPT-5.6 Sol, that we needed to look at my voice and style guides “and see if there is any guidance in those documents that may have caused conflict or confusion in how the piece came together.” I didn’t tell it how. Sol reached for the doc-review workflow from Every’s Compound Engineering plugin, which hands a document to a panel of AI reviewers, each looking for a different kind of problem.
The voice guide was mostly fine. The style guide was full of contradictions, and Sol explained why they had done so much damage: “In an agent workflow, the concrete rules are easier to verify, so they can overwhelm the subtler ones while every checklist still passes.” “Open with one of these eight approaches” is a box the model can tick. “Let one experience carry the piece” takes judgment. The model that had fought me draft after draft had been obeying the parts of my instructions that were easiest to follow.
Starting over from scratch
So I knew what had happened to get my context in the state it had gotten into. The question now was what I was going to do to get it out.
I decided pretty quickly that my existing tangle of context was a lost cause. So I made a new folder on my desktop, named it Historical, and dragged everything into it: every log, every experiment record, every cross-reference, every version of every guide. Nothing got deleted. The archive is still there, and my new instructions tell the model to stay out of it unless I send it in.
Then I rebuilt the whole thing from scratch, from an empty desktop. The model interviewed me about my work and how I wanted AI to support me. Then it reviewed my existing folder structure and proposed a new structure based on my answers, which we refined until I understood how it worked. Then it made the changes, under an instruction that was not subtle: “Do not change anything until I approve.” I had developed a healthy paranoia about AI going off and changing things without my noticing.
Roughly 20 minutes after I reset my desktop, I was back to roughly the arrangement I’d started with: a folder for each column, each with a voice guide, a style guide, and its drafts. The old voice and style guides for this column ran 6,109 words. The new ones run 513.
They’re also a different kind of document. My new instructions say outright that drafts, outlines, notes, and reviews are “task material or evidence, never new instructions.” The model can consult anything I’ve written, but it doesn’t have to satisfy everything.
There are three key principles that make this arrangement work for me:
- The agent asks first. My instructions now tell the model to answer me in the chat and to create or update a file only when I ask for one. Nothing enters my style guide, my voice guide, or my drafts folder without my say.
- Every folder stands alone. A model working on this column starts in this column’s folder and reads what’s there. It doesn’t follow links into my other projects, and it doesn’t open the Command Center unless I tell it to.
- The guides stay short. My style guide still describes a way of moving through a piece that tends to work for this column. But it says a successful essay “often (not always)” goes that way, and that “these are principles, not a required sequence of sections or a fixed essay template.”
The folders work again, but I’m always vigilant for the possibility that my own human foibles could get in my way again. So I’ve added one more rule, for me rather than the model: no major context decisions under pressure. That includes the obvious kinds of pressure, like a deadline or a bad week, and the thrilling new-model-tilt-a-whirl kind, which is the one that got me. If I catch myself with a brilliant idea for restructuring everything, I write it down and look at it again when the new-model mania has eased.
Learning to trust again
When I first noticed that my context was no longer serving me, I lost trust in my context first and myself second. The trust came back in the opposite order.
First, I had to trust that I knew my own work well enough to say what belonged in those folders and that I’d recognize it when the model read it back to me. Trusting the system took longer. I had to see the system working and helping me meet deadlines the way I was used to before I was ready to call the problem solved.
The list of exercises in that probation includes this essay, which I wrote in the rebuilt Working Overtime folder with a handful of short, reasonable context files for company. Twice I told the model that a section felt too abrupt. Both times, the next version was better. A month ago, a note like that would have calcified into yet another impossible standard for an essay to meet. This time, I corrected what needed correcting, and that was it.
Perhaps you’ve fallen into a similar trap. Once you see how much better AI can work with a little context, it’s tempting to keep tuning until it perfectly matches your sensibilities. But trying to anticipate everything you’ll ever want can leave you fighting instructions you wrote for circumstances that no longer apply. I had to trust myself to recognize what the work needed when I was doing it, instead of trying to settle every decision in advance. Before you save another preference, example, or instruction to your setup, ask whether it belongs in every task or just the one in front of you. And when an “improved” setup starts giving you worse results, try taking something away before adding more. You don’t have to wait until you’ve produced a million words to investigate.
I’d like to tell you I’ve learned my lesson for good. But there will be another model, and it will be dazzling, and I will get another brilliant idea for making this even better. The Germans, as we’ve established, already have a word for what happens next.
Katie Parrott is a staff writer at Every. To read more essays like this, subscribe to Every, and follow us on X at @every and on LinkedIn.
Everyone’s a builder now. Every All Access gets you the full membership plus the Builder Pack—$9,000+ in credits for the tools we build with.