Managed Agents

This is the first post in a series unpacking Anthropic’s Managed Agents post, released about six months ago.

What they shared, and what ripple effects we’re seeing from it today.

We’re going to spend a lot of time talking about harnesses.

A harness is the code wrapped around the model. It calls the model in a loop, runs the tools, and decides what goes into the window. A harness-level move is something that code does to the model’s situation, rather than something the model decides or a prompt tells it to do.

Some important definitions to ground ourselves in:

LevelWho actsExample
ModelThe model itselfFable 5.1
PromptInstructions in textGo to YouTube and search for “butterfly” every day at 9 a.m., then return the 16th result to me if the butterfly is orange
HarnessSurrounding codeA reset (wipe the window, start a fresh agent with a handoff) or compaction (summarize in place)

Early on, Anthropic saw that Sonnet 4.5 wrapped up tasks early as it approached its context limit. That meant that if you had asked the model the above butterfly prompt, it might have wrapped up early and left the task unfinished if it was getting too close to its limit. To solve that, Anthropic added a context reset to the harness: the window is wiped and a fresh agent starts, picking up from a handoff doc that says “here’s what I did, continue from here.” But when they moved to Opus 4.5, they didn’t run into the issue of the model wrapping up early as it approached its limit (it’s unclear why). They realized they would keep running into this problem: a harness encodes assumptions about what the model can’t do on its own, and the model can outgrow those assumptions.

Leave a comment