Context window management is the invisible cost in every multi-step agent pipeline.
I see this pattern repeatedly in production consulting work. An agent that works cleanly in testing — ten steps, coherent reasoning, correct output — starts degrading six weeks after ship. No errors. No crashes. Just drift.
The root cause is almost always the same. Context grows with every step and nobody measured it.
A multi-step agent accumulates its full conversation history by default. Every tool call result gets appended. Every intermediate reasoning step gets stored. By step 8 or 9, the model is receiving 90% of the context window on each call — technically within limits, but attention spread thin across thousands of tokens of earlier reasoning that no longer affects the current step.
This is not a prompt quality problem. The prompt was fine in testing. Context management was never treated as a first-class engineering concern.
Three things that work in production:
Allocate a fixed token budget per step, not per conversation. If step 3 needs the output of step 1, pass the structured result only — not step 1's full chain-of-thought.
Summarize and compact before handoff between pipeline stages. The detailed reasoning from a retrieval step does not need to be present when the synthesis step runs. A structured summary does.
Set a hard context limit per step and fail loudly when exceeded. Silent context overflow produces subtly wrong outputs that reach users. A visible error at least surfaces the problem before it propagates downstream.
The agent's behavior in single-query testing is not the agent's behavior under production load. Context growth under real query volume looks nothing like the test environment. The two will not converge on their own.