Context Engineering

Context engineering is deciding what information actually goes into a model's context for a given task — which instructions, which retrieved documents, how much conversation history, which tool results — rather than treating everything as if it should just be included. An agent answering one question might have a system prompt, several retrieved documents, the last dozen turns of a conversation, and the result of a tool call all competing for the same request. Deciding which of those actually belong, how much of each, and in what order is context engineering.

What it actually consists of

Selection. Choosing which available information is actually relevant to this request — not everything that exists, everything that helps. A support assistant doesn't need the entire policy manual in context; it needs the section relevant to this question.

Compaction. Trimming or summarizing what's too long to include in full — a long conversation history gets condensed to what still matters, rather than resent in its entirety on every turn.

Isolation. Keeping one part of a task's context from bleeding into another — a subagent working on one piece of a larger job gets its own focused context instead of inheriting everything the main agent has accumulated.

Ordering. Placing information where a model will actually use it well, not just where it fits. Models are measurably less reliable at using information buried in the middle of a long input than information near the start or end, so what goes where is a real decision, not an afterthought.

Why it became necessary

As long as most AI use was one person typing one question into a chat box, there wasn't much to manage — the whole request was just that one message. It became its own problem once systems started running agents across many steps, retrieving documents, calling tools, and carrying long conversation histories, all of which have to be assembled into a single request every time. See Prompt Engineering vs Context Engineering for the full distinction between the two.

More context is not automatically better context

Adding more to a request only helps when what's added is actually relevant — irrelevant, duplicated, or conflicting information can crowd out the material that matters, and a model can only use what it's given well if there isn't too much competing for its attention. This is the same limitation the context window page — the fixed budget of text a model can take in for one request — covers in more depth: fitting is necessary, but it's never sufficient on its own.

Where it matters most

Anywhere a request isn't just one hand-typed message: RAG pipelines deciding how many retrieved chunks to include and in what order, agents carrying state across many steps, multi-agent systems deciding what each agent should and shouldn't see, and any long-running conversation that has to decide what history still matters.

In this guide
  1. What it actually consists of
  2. Why it became necessary
  3. More context is not automatically better context
  4. Where it matters most
  5. FAQ

FAQ

If a model has a huge context window, does context engineering still matter?

Yes. A bigger window raises how much can technically fit, but it doesn't change how reliably a model uses everything inside it, and it doesn't reduce the cost of sending more than a task actually needs. A large window makes sloppy context engineering less immediately punishing, not unnecessary.

Is context engineering something only agent builders need to think about?

It matters most in agentic and multi-step systems, but it applies anywhere more than one source of information feeds into a single request — a RAG pipeline deciding how many chunks to retrieve is doing context engineering, even with no agent involved at all.