Context Engineering for AI Agents
The discipline that replaced prompt engineering: how to treat the context window as a finite, managed resource across an agent's multi-step lifetime — what to load, what to compact, what to persist, and how to prove it helped.
Key Takeaways
- Context engineering is deciding what enters the model's window at every step of an agent run — and what gets pruned, summarized, or offloaded. It is a curation problem, not a wording problem, and it is where agent quality is actually won or lost.
- Bigger context windows did not solve retrieval. Models read long contexts unevenly: Liu et al. found that placing the relevant document mid-context made GPT-3.5-Turbo answer multi-document questions worse than giving it no documents at all.
- There are five strategies, not one: stuff the window, retrieve selectively, compact the history, isolate work in sub-agents, and offload to external memory. Production agents use four of them at once, and each has a distinct failure mode.
- Tool results are usually the largest and least valuable thing in an agent's window. Pruning raw payloads down to the fields the agent actually reasons over is the single highest-leverage change in most agent loops.
- Memory is a persistence decision, not a storage feature. Persist durable facts, decisions, and user preferences; let transient observations expire. Anything you would not want re-read on turn fifty does not belong in long-term memory.
- Measure context changes with task-level outcome evals plus tokens-per-successful-task — never with vibes. A change that cuts tokens 60% while dropping task success 10% is a regression, and only paired metrics reveal that.
Context engineering is the practice of deciding what enters a model's context window at every step of an agent's run – and what gets pruned, summarized, or offloaded instead. It displaced prompt engineering as the thing that determines agent quality for a structural reason: in a multi-step agent, your carefully worded system prompt might be 800 tokens, while accumulated tool output, conversation history, and retrieved documents reach hundreds of thousands. Tuning the 800 tokens while ignoring the rest is optimizing the wrong surface. The window is a finite, contested budget, and deciding what occupies it is an engineering discipline with its own strategies, failure modes, and metrics.
The reason this outranks raw window size is that models do not read long contexts evenly. According to Liu et al., 2024, when the relevant document was placed in the middle of the input context, GPT-3.5-Turbo answered multi-document questions less accurately than it did with no documents at all – its closed-book baseline of 56.1%. Supplying the correct information made the model worse, purely because of where that information sat. That result is the foundation of the entire discipline: context is not additive, so more of it is not automatically better.
Which context strategy should you use?
There are five strategies, and mature agents run four of them simultaneously at different layers. Each buys you something specific and breaks in a specific way:
| Strategy | When to use it | Failure mode |
|---|---|---|
| Stuff everything in the window | Small, bounded corpora under ~20k tokens; prototypes; single-turn tasks | Position effects bury the key fact mid-context; cost and latency scale on every step |
| Retrieval (RAG) | Large or changing knowledge bases where only a few passages matter per query | Retrieval misses silently; the agent answers confidently from whatever came back |
| Compaction / summarization | Long agent loops and multi-hour sessions that outgrow the window | Summaries drop exact identifiers, paths, and error strings needed to resume |
| Sub-agent isolation | Parallelizable subtasks with wide exploration and narrow outputs | Coordination overhead; sub-agents lose context the parent never passed down |
| External memory (files, state store) | Durable facts, artifacts, and anything needed across sessions | Stale entries persist and actively mislead long after they stopped being true |
The decision is not which one to pick, but which layer each belongs to: retrieval for knowledge, compaction for history, isolation for parallel work, external memory for durability, and stuffing only for the small stable core – system prompt, tool definitions, current objective.
Why did bigger context windows not solve this?
Million-token windows changed the constraint from what fits to what the model can reliably use, which is a harder problem. Three effects explain why.
Position sensitivity. Attention is not uniform across a long input. Information at the very start and very end is used most reliably, and material in the middle is used least – the U-shaped curve the research above documents. In an agent loop this is especially cruel, because the original objective is stated first and then pushed into the middle as the transcript grows.
Context rot. Over a long run the window accumulates superseded plans, failed attempts, and stale tool output that nobody removed. The model has no way to know step four invalidated the approach from step two, so it attends to both. Rot presents as an agent repeating completed work, re-running failed tool calls, or drifting from the objective – symptoms teams usually misdiagnose as a weak model.
Cost and latency compound per step. A forty-step agent carrying 150,000 tokens pays that toll forty times, and users feel it as latency on every turn. Context discipline is one of the highest-leverage levers on unit economics, alongside the tactics in our guide to reducing LLM API costs.
How do you compact a long agent loop without losing the thread?
Compaction replaces a long stretch of transcript with a shorter representation that preserves everything needed to continue. Trigger it on a threshold you choose deliberately – commonly 50–80% of the usable window – not when you hit a hard limit, because emergency compaction under pressure is where agents lose critical state.
A compaction that survives contact with reality preserves five things:
- The original objective, verbatim. Never paraphrase it. Paraphrased goals drift, and drift compounds across every subsequent compaction.
- Decisions and their rationale. "Chose Postgres over the vector store because the dataset is 4k rows" prevents the agent from relitigating a settled question.
- Exact identifiers. File paths, record IDs, branch names, and error strings must survive character-for-character. Summarizing these is the number-one cause of an agent that resumes confidently and then fails on the next tool call.
- Open questions and pending work. What is unresolved matters more than what is finished.
- What was already tried and failed. Without this, the agent re-runs the same failing approach, which is the most expensive and most common post-compaction bug.
Everything else – verbose intermediate reasoning, successful routine steps, raw payloads – collapses into a sentence. Place the compacted summary near the end of the assembled context rather than the middle, so it lands where the model attends best.
What should the agent remember, and what should it forget?
Memory is a persistence decision, not a storage feature. Split it in two.
Short-term working context is everything relevant to the current task: the plan, recent tool results, intermediate state. It should expire with the session. Persisting it feels thrifty and is actively harmful, because next week's task inherits last week's half-finished reasoning.
Long-term memory holds only what stays true and would be expensive to rediscover: user and account preferences, stable facts about the system or codebase, decisions with rationale, and outcomes of prior tasks. The test is simple – would you want this re-read on turn fifty of an unrelated task? If not, it is not long-term memory.
Two rules keep memory from becoming a liability. Write memories with a timestamp and a source, so a stale entry can be recognized as stale rather than trusted as fact. And give every memory an expiry or a review trigger; a permanent store with no eviction policy becomes a slow-motion accuracy problem, since the agent cannot tell that the API it memorized was deprecated in March.
How do you stop tool results from eating the window?
In most production agents, tool results are the largest and least valuable thing in the window. An API returns 4,000 tokens of JSON; the agent reasons over three fields. Multiply by thirty calls and the transcript is 90% payload the model never uses but must read past on every step.
Four pruning tactics, in rough order of leverage:
- Project at the tool boundary. Have the tool return only the fields the agent needs, rather than forwarding the raw upstream response. This is a code change in your tool wrapper and it is the single highest-return fix in most agent loops.
- Paginate and cap. Return the first N results with a count and a way to request more, instead of dumping the full set. An agent that can ask for page two rarely needs page two.
- Age out old results. A tool result from twenty steps ago is usually superseded. Replace it with a one-line record of what was called and what it concluded, keeping only recent results in full.
- Offload large artifacts to files. Write the 200KB report to disk or object storage and put the path in context. The agent can re-read the part it needs, when it needs it, instead of carrying the whole thing forever.
This is also where tool design and context design meet: verbose, overlapping tools produce verbose, overlapping context. Our guide to designing software and APIs for AI agents covers the interface side of the same problem.
When should you isolate context in sub-agents?
Sub-agent isolation gives a subtask its own clean window and returns only a condensed result to the parent. It is the right tool when the work has a wide exploration surface and a narrow output: searching forty files to answer one question, evaluating six vendors to produce a recommendation, scanning a large log to extract three incidents. The parent inherits the conclusion, not the forty files.
It is the wrong tool when subtasks are tightly coupled or need to see each other's intermediate state. Isolation is the whole point, and it cuts both ways – a sub-agent knows only what the parent passed down, so vague delegation produces confidently wrong work built on missing premises. Write sub-agent instructions as if briefing someone who cannot ask a follow-up question, because that is exactly the situation. The broader trade-offs, including where multi-agent architectures stop paying for themselves, are in our guide to multi-agent systems.
How do you measure whether a context change helped?
Context changes are unusually easy to fool yourself about, because a leaner window feels better and costs less while quietly failing more often. Measure two things together on a fixed evaluation set of realistic multi-step tasks:
- Task success rate – did the agent complete the job correctly, end to end. This is the metric that governs.
- Tokens per successful task – total tokens across all steps, divided by successes only. Dividing by attempts rewards cheap failures.
A change that cuts tokens 60% and drops success 10% is a regression, and only the pair reveals it. Add three diagnostics that catch context problems earlier than outcome metrics do: steps to completion (rising counts signal the agent is losing the thread), repeated-action rate (the clearest signature of context rot), and tool-call error rate (spikes right after compaction mean your summary is dropping identifiers). Change one thing at a time – compaction threshold, then pruning, then memory policy – because bundled context changes are effectively unattributable. The eval scaffolding this requires is covered in our guide to evaluating and testing AI agents.
Where to start
If you have an agent in production and no context discipline, the highest-return sequence is: instrument tokens per successful task, prune tool results at the tool boundary, then add threshold-triggered compaction that preserves identifiers verbatim, then decide a memory persistence policy. Retrieval architecture sits underneath all of it: if the agent draws on a knowledge base, selection quality upstream sets the ceiling on what context discipline can save you downstream, which is the subject of our guide to building a RAG system.
The through-line is that the context window is a managed resource with a budget, an eviction policy, and a measurable hit rate – not a place to put things. Teams that treat it that way ship agents that stay coherent on turn eighty. Teams that do not keep upgrading models to fix a problem no model upgrade addresses. If you want that layer designed for an agent you are building, see our services.
Frequently Asked Questions
What is context engineering, and how is it different from prompt engineering?
Context engineering is the practice of curating everything that enters a model's context window at each step of an agent's run: the system prompt, tool definitions, retrieved documents, conversation history, tool results, and persisted memory. Prompt engineering is a subset of it that focuses on the wording of instructions. The distinction matters because in a multi-step agent the system prompt might be 800 tokens while accumulated tool output and history reach hundreds of thousands. Optimizing the 800 tokens while ignoring the rest is optimizing the wrong thing.
Did long context windows solve the need for retrieval?
No. Large windows changed the constraint from what fits to what the model can reliably use, but they did not remove the need to select information. Research on long-context behavior shows accuracy depends heavily on where relevant information sits in the input, with degradation when it falls in the middle of a long context. Long windows also cost more and add latency on every step of an agent loop, so a hundred thousand tokens of unfiltered context is expensive, slower, and often less accurate than five thousand well-chosen ones.
What is context rot in AI agents?
Context rot is the gradual degradation of an agent's behavior as its context window fills with stale, contradictory, or irrelevant material over a long run. It shows up as an agent repeating completed steps, following instructions that were superseded twenty turns ago, or losing the original objective. The causes are accumulation of raw tool output, superseded plans that were never removed, and position effects that push important early instructions into the poorly-attended middle of the window. Compaction and pruning are the direct fixes.
When should an agent compact or summarize its context?
Compact when the context crosses a threshold you set deliberately, typically somewhere between 50 and 80 percent of the usable window, rather than waiting for a hard limit. A good compaction preserves the original objective verbatim, the decisions made and why, open questions, and any identifiers or file paths needed to resume work, while collapsing verbose intermediate steps into a short narrative. The most common mistake is summarizing everything uniformly, which discards the exact strings — IDs, paths, error messages — the agent needs to continue.
What should an AI agent store in long-term memory?
Persist things that stay true and would be expensive to rediscover: user preferences, stable facts about the account or codebase, decisions with their rationale, and outcomes of prior tasks. Do not persist transient observations, raw tool payloads, or intermediate reasoning, because they age badly and become actively misleading when the underlying system changes. A practical test is whether you would want the item re-read on turn fifty of an unrelated task. If not, it belongs in short-term working context that expires with the session.
How do you measure whether a context change improved an agent?
Run a fixed evaluation set of realistic multi-step tasks before and after the change and track two metrics together: task success rate and tokens per successful task. Success rate alone hides cost blowups, and token count alone rewards changes that make the agent cheaper and worse. Add per-step diagnostics such as tool-call error rate, number of steps to completion, and how often the agent repeats an action, since context problems usually surface as loops and retries before they surface as outright wrong answers.
Free Tools
Tell us what you're building — book a free scoping call.
Pick a time that works and walk us through your project — 30 minutes, straight to the point. You leave with a concrete plan, timeline, and cost. No sales pitch — if we're not the right fit, we'll say so.
Keep Reading
Get new playbooks by email
Occasional, no-fluff field notes on building production AI — new guides and tools, straight to your inbox. Unsubscribe anytime.