First Long-Running Task in the Agent Loop

I wanted to try the next step of what I already do regularly during my lunch breaks: give Claude a bigger package of work and let it run for a long time with almost no interaction from me. The task was to replace an old library in the codebase with a new one.
This time I wanted it to keep working through the whole weekend. That part did not work out. The task itself went well. I went through the code for structure, tests and comments and took a closer look at the important parts. I am happy with the result. What I did not expect was how much it grew.

What stands out is how much the tests, the docs and partly the code grew. The library was already well tested before the run, and its tests still grew by 73%. The docs grew by 36%, although the run added no feature and only closed a few gaps. My gut feeling says not all of it was needed.
I had a feeling about this before. During lunch breaks I let sessions work on their own for half an hour to an hour. In runs that short, the growth of tests and docs does not catch the eye. After this long run I analyzed how everything grew, and it stood out.
Six months ago I built the new library (with Claude) in a controlled way. I kept a close eye on the API, the internals and the tests, and it worked well. On Friday came the dry run: does it hold up in the legacy code everywhere? It showed a few gaps, not many, but enough to straighten out. So we made some features and their outward behavior more consistent. Most of the library stayed as it was.
Back then I checked every step myself. In a long run without me, nobody checks, and the agents keep adding tests. A well-designed API needs that control. I want to stop the tests from growing unchecked, and I already have ideas. Other developers see the same problem. On X it led to a heated discussion, and one post put it like this.
Ansh Nanda: NEVER write unit tests after you write code. Highly prefer E2E tests as the sole testing mechanism.
That goes far. Moving everything to end-to-end tests does not feel right to me. Unit tests have their place. They are usually much faster and more precise. They can also cover many parts at once, as integration tests, and that alone can cut the number of tests. Which way works best is something I still have to find out.
The library’s code grew by about half, although no feature was added. Most of it came from closing those gaps. I will try to slim down the library’s internals again.
What matters most stayed clean: the public API still follows my design, now with consistent interfaces. From what I read and from my own experience, well-designed APIs that work well together make the agents’ work easier. The results get better too. I cannot back that with data yet.
Seven hours, two hard interruptions
It all started on Friday at around 17:30, right after work. I gave Claude a dry run of a big migration in a brownfield project. Up to 8 agents ran in parallel, and some of them started sub-agents. More was not possible on my MacBook. One coordinator managed everything but was not allowed to write code. It could only plan, distribute work, start and stop agents, and give instructions.

Around 21:30 my MacBook crashed anyway, because memory and CPU gave up. The end-to-end tests probably played a part: every agent ran them on several cores at once. For the next run I will turn that down. Luckily cmux (a tmux-like terminal multiplexer) restored the sessions and workers, and the run continued.
At 23:30 I ran out of tokens. Up to 8 agents and their sub-agents, running for hours, use up a $100 plan quickly. Claude’s review skill and its sub-agents ate a lot of them. That will be different in phase 2. At 02:30 the limit reset and Claude continued for another hour or so.
In total the setup worked for around 7 hours, and that was only phase 1 of 2. Now it is Monday, and I have only 5% of my weekly limit left. Even with 5% you can do much more than you would think (right then Opus 5.5 came out, and I could use it on low effort). For long-running tasks with complex coding work it is nowhere near enough though. On Tuesday my weekly limit resets, and I can go on.
Building blocks that have to work together
Along the way I hit problems I knew about before, but never connected to long-running autonomous tasks. I had already solved many of them, each on its own, such as lost knowledge after compacting, stalled agents and questions while I am away. In a long-running task all these solutions have to work together. Orchestrating them does not work the way it does when you handle them separately. These building blocks have to work before you even start.
Keeping knowledge across compactions

Claude’s default auto-compacting loses a lot of knowledge every time it runs. After compacting, an agent with an unfinished task does not continue on its own. The coordinator has to ask it to go on, and you have to teach the coordinator that.
It has become a habit of mine: when a problem shows up, I think of a solution and build it. From the lessons of that Friday I built a Claude Code mod for both, the lost knowledge and the stalled agent, and called it /compactor.
The mod I built is deterministic, unit-tested and follows a fixed workflow. Only the generative part goes to the agent: writing the compact prompt and checking the summary. That is the split I want. Code handles what must be reliable, the model handles what needs judgment.
How it works: before the context runs full (65% by default, configurable), the session writes its own compact prompt. The summary then keeps what it needs to go on. After compacting, it checks its summary and continues when a step is still open. So it runs in a loop of work, compact, resume and work again. The loop ends only when the coordinator has nothing left to answer. That is a loop too: coordinator and workers go back and forth until the work is done, and inside every worker the compactor loop runs.
Boris Cherny: I don’t prompt Claude anymore. I have loops that are running. They’re the ones that are prompting Claude and figuring out what to do. My job is to write loops.
Working on brownfield projects
A long-running task in a brownfield project is much more complex than building something greenfield. The migration replaces a core part of the application with something new. That part touches almost every corner of the code. The agent shows you how much bad code you produced over the years: edge cases, workarounds, technical debt.
Debt from before the AI era can be worse, because it was often built under time and budget pressure. A migration like this works with agents when it is planned in detail. Decide the hard questions before you start and do not give the agent a free hand. Otherwise it makes decisions you do not want.
A library that grew for years
The migration replaces an old library that had grown for years. Time and budget were always tight, so I kept patching new things onto it. By now that library is used across the whole codebase.

The new library is its successor. It builds on years of usage and on everything we had come to miss. A wish list I got added the rest. I held the migration back until I had the time and the tooling for a migration of this size.
In a brownfield project, all of this takes far longer than in a greenfield one: the new library, the tooling and the know-how behind them. It is years of engineering work, and that work is the expensive part, not the code. I wrote about this before in Generating Code Is Cheap, Engineering Is Not.
When the terminal is not enough
Most of my work runs in the terminal. Tickets that are already well described go straight to Claude, before I have read them once. This run was different. The agents were to run it on their own, and the planning alone grew too complex for the terminal.
Then /show-me came to mind, and I knew I needed something like it for code. So I built a skill, /show-impact. It shows at a high level what changes or will change in the code. It is not an editor with a full diff of old and new code.
It writes the planned code in an invented mix of the code’s language and plain words. The reason for each line follows after ⟶. It does this only at the points that matter. Depending on the change it shows call chains, file trees, changed lines, before and after, call sites, return values or variants.

I worked out the whole plan in artifacts, one to two days before the start at around 17:30. One artifact held many questions like the one above, as many as the agent and I wanted to clarify. Each of them had to be settled so the migration would come out clean. During the run itself, questions came up only now and then. I stopped reading every line long ago, because the amount of code is too large. /show-impact does not change that. I use it only for the decisions that matter, not to read the code.

As a developer you see at once what changes and how. It works in the terminal and as an artifact. In the terminal it looks much like /show-me. An artifact can show more complex things, and every block gets a text field, so I can write to the Claude session directly from the browser.
Staying reachable
The real question is how you stay reachable. When you leave the house, how do you answer a short question with yes or no? This should not become the normal state. A weekend is a weekend, and I do not want to spend it watching agents. But you have to invest here now. What you learn is what lets you stop checking on a run later, on weekends or after work. Progress needs investment.

When I started the run, I expected it to have questions. But I was about to leave the house and would be back only hours later. How was I supposed to answer questions while I was away? And I wanted the run to keep going. Then I remembered that you can connect to a session remotely from your phone. You see what happens in the session, and you get a notification when it has a question. That is what I did.
Before the next run
- Leaner tests, docs and code: Tests grew by 73%, docs by 36% and code by about half. I will slim down all three, with tests that check the right behavior, not just the same coverage. The bigger goal is that it does not happen again, so the next phase starts with a strategy for tests, docs and code. More on that in part 2.
- /compactor in a real run: It has tests and is well worked out, but only a real run shows how well it works. I expect far fewer stalls, which cost me time on Friday.
- Smart routing: I will probably need smart routing for model and effort. Each task then gets the model and the effort level it needs. That saves tokens, so the next run does not stop at the token limit.
- Resource logging: I will log CPU and RAM during the whole run, which I did not do last time.
Conclusion
This run was my first real step toward agent factories, and it confirmed four things for me:
- API design decides the result: When you think the API and the interfaces through with Claude and keep full control, the agent builds what you had in mind. The internals matter much less for that, but they still matter to me, because tests and docs grew more than needed and some functions may not be needed at all.
- A new library needs full control: A library you start greenfield should be built step by step under close control, not in a long run on its own.
- Knowledge has to survive compaction: The most important lesson of this run. /compactor keeps the loop going until the task is done. After every compaction it knows where to go on, and its knowledge is up to date. Claude’s default auto-compacting loses knowledge and leaves unfinished work standing.
- A migration needs an overview: In a long migration you have to keep track of what changes where. That is why I built /show-impact.