Context Engineering
Context engineering is deciding what information a model actually sees for a given task. That means choosing which instructions, which retrieved documents, how much conversation history and which tool results go in — rather than including everything by default. An agent answering one question might have a system prompt, several retrieved documents, the last dozen turns of a conversation, and the result of a tool call all competing for the same request. Deciding which of those actually belong, how much of each, and in what order is context engineering.
What it actually consists of
Selection. Choosing which available information is actually relevant to this request — not everything that exists, everything that helps. A support assistant doesn't need the entire policy manual in context; it needs the section relevant to this question.
Compaction. Trimming or summarizing what's too long to include in full — a long conversation history gets condensed to what still matters, rather than resent in its entirety on every turn.
Isolation. Keeping one part of a task's context from bleeding into another — a subagent working on one piece of a larger job gets its own focused context instead of inheriting everything the main agent has accumulated.
Ordering. Placing information where a model will actually use it well, not just where it fits. Models are measurably less reliable at using information buried in the middle of a long input than information near the start or end, so what goes where is a real decision, not an afterthought.
Why it became necessary
As long as most AI use was one person typing one question into a chat box, there wasn't much to manage — the whole request was just that one message. It became its own problem once systems started running agents across many steps, retrieving documents, calling tools, and carrying long conversation histories, all of which have to be assembled into a single request every time. See Prompt Engineering vs Context Engineering for the full distinction between the two.
More context is not automatically better context
Adding more to a request only helps when what's added is actually relevant — irrelevant, duplicated, or conflicting information can crowd out the material that matters, and a model can only use what it's given well if there isn't too much competing for its attention. This is the same limitation the context window page — the fixed budget of text a model can take in for one request — covers in more depth: fitting is necessary, but it's never sufficient on its own.
Where it matters most
Anywhere a request isn't just one hand-typed message: RAG pipelines deciding how many retrieved chunks to include and in what order, agents carrying state across many steps, multi-agent systems deciding what each agent should and shouldn't see, and any long-running conversation that has to decide what history still matters.
In this guide
FAQ
If a model has a huge context window, does context engineering still matter?
Yes. A bigger window raises how much can technically fit, but it doesn't change how reliably a model uses everything inside it, and it doesn't reduce the cost of sending more than a task actually needs. A large window makes sloppy context engineering less immediately punishing, not unnecessary.
Is context engineering something only agent builders need to think about?
It matters most in agentic and multi-step systems, but it applies anywhere more than one source of information feeds into a single request — a RAG pipeline deciding how many chunks to retrieve is doing context engineering, even with no agent involved at all.