Context Window
A context window is the maximum amount of text a model can read and act on for one request. Send more than it can hold, and something has to give — either the request fails, or part of what you sent gets left out.
Everything shares that one budget. The system instructions, the whole conversation so far, any documents or search results handed to the model, and even the answer it writes back all count against the same limit — none of it is stored anywhere separate. The limit itself is measured in tokens, not characters or words — a token is a small chunk of text, often close to a whole word or part of one. A short sentence runs about 8-10 tokens; a long document can run to tens of thousands.
What happens when you go over it
One of two things. Some systems return an outright error and refuse the request rather than guess what to cut. Others quietly drop the oldest or least-recent part of the conversation to make room — which is the more dangerous failure, because the app keeps answering, just without whatever got pushed out: an early instruction, a fact from a document, or a tool result the answer actually depended on. The model doesn't know something is missing. It just answers based on what's left.
Three places this actually bites
A long-running chat. As a conversation grows across many turns, the earliest messages start getting dropped once the total no longer fits — not because the model chose to forget them, but because they were never sent again. Ask about something from early in a long session and the model may have no idea, because that part of the conversation simply isn't in front of it anymore.
A big document. Paste a 300-page manual into a single request and it either gets cut off partway through, or the application has to break it into pieces and handle them separately — the model never sees "the whole manual" as one thing unless it genuinely fits.
A RAG pipeline. The retrieved chunks, the user's question, and the conversation history all have to fit in the same window as everything else. That's part of why chunk size and how many chunks to retrieve are real decisions, not free choices — every chunk pulled in costs some of the same budget as everything else in the request.
A bigger window isn't automatically better
Context windows have grown a lot — from a few thousand tokens in early models to hundreds of thousands or more in many current ones. That makes it tempting to skip being selective: if it all fits, why not just send everything relevant instead of carefully picking the best pieces?
Two reasons. First, cost and speed scale with how many tokens you send, whether or not you were near the limit — a request using ten times the tokens is slower and more expensive even when it fits comfortably. Second, fitting isn't the same as being used well: research on how models actually use long contexts found they're noticeably better at using information placed near the start or end of a long prompt than information buried in the middle — performance on the buried information dropped by more than 30% in their tests. A bigger window raises the ceiling. It doesn't remove the reason to think about what you put in it and where.
In this guide
FAQ
Does a model remember an earlier conversation once its context window fills up?
No — a full context window isn't the same thing as memory. Once something falls out of the window, the model has no separate place it's stored; if the application doesn't deliberately save and resend it, it's simply gone from what the model sees. Persisting something across sessions on purpose is a separate job, usually handled by a memory system built on top of the model, not by the context window itself.
If a conversation is well under the limit, does the context window still matter?
Yes. Cost scales with tokens sent regardless of how close you are to the limit, and the tendency to use information near the start or end of a prompt more reliably than information in the middle doesn't switch on only once you're near capacity — it's a property of how the model reads a long prompt at all. Fitting under the limit means the request will work. It doesn't mean every part of it will be used equally well.