Chunking
Chunking is splitting a document into smaller pieces before you index it, so that a search returns one relevant passage instead of a whole file. A chunk is just one of those pieces — a paragraph, a section, a few hundred words.
Each chunk gets turned into an embedding — a list of numbers that lets a search match on meaning rather than wording — and that page covers why splitting is necessary in the first place. This one is about the decisions once you've accepted that it is, starting with the one people underrate.
In this guide
Where you cut matters more than how big
The simplest approach is to split every N words and move on. It works until it doesn't, and when it fails it fails quietly: a table separated from the caption explaining it, a price separated from the product it belongs to, a sentence beginning "It costs $40" with nothing nearby saying what "it" is. Each of those chunks still gets indexed. Each is still retrievable. And each is useless or actively misleading on its own.
So split on the boundaries the document already has. Paragraphs, headings, sections, list items. For Markdown or HTML that structure is sitting right there in the markup. For code, split on function and class boundaries rather than line counts. The common default in retrieval libraries works down a priority list — try paragraph breaks first, then sentence breaks, then fall back to a hard cut only when a single passage is too long for anything else.
A document with real structure should almost never be split by counting characters. A wall of unbroken text is the case where you have no choice.
What overlap is for
Overlap means repeating the tail of one chunk at the start of the next — the last sentence or two, typically. It exists for exactly one reason: a fact that straddles a boundary appears whole in at least one chunk instead of being cut in half in both.
It costs storage and some duplication in results, and it doesn't fix bad cut points. Overlap is insurance against an unlucky boundary, not a substitute for choosing sensible ones.
So how big should a chunk be?
There's no single right answer, and that's a real finding rather than a dodge. Around 512 tokens — tokens being chunks of text, usually a word or part of one — with a sentence or two of overlap is the common default across retrieval tooling, and it's a reasonable place to start.
But the useful research result is that the best size varies by question, not just by corpus. A narrow factual lookup does better on small, tightly focused chunks. A question needing several pieces of information assembled together does better on larger ones that keep related material adjacent. Both kinds of question hit the same index, so no single size is optimal for all of them — which is why tuning the number tends to produce smaller gains than people hope, and why fixing cut points and metadata usually pays better.
Two things do constrain you from outside: your embedding model's input limit, which truncates anything longer, and the cost of everything you retrieve going into the model's context on every query.
Embed a small chunk, return a bigger one
These don't have to be the same piece of text, and separating them resolves most of the size tension. Index a small, precise chunk so matching is sharp. Then, when it matches, return the surrounding section — the paragraph either side, or the whole subsection it came from — for the model to actually read.
You get precise retrieval and sufficient context at once, instead of trading one against the other. It needs your stored chunks to know where they came from, which is the next point.
Keep the metadata, always
Store each chunk with where it came from: document, section, heading path, position, date. This is easy to skip and expensive to retrofit.
Without it you can't cite a source, can't filter a search to one product or one year, can't expand a match into its surrounding section, and can't tell a user which document an answer came from. The retrieval might work fine and the system still be unusable, because nobody can check anything.
Give each chunk its context back
A chunk pulled out of a document loses everything the document was telling you implicitly. "Revenue grew 3%" doesn't say which company or which quarter — the surrounding pages did.
Anthropic's contextual retrieval addresses this directly: before embedding each chunk, generate a short description — 50 to 100 tokens — situating it in its source document, and prepend that to the chunk. Their published results measure the share of relevant documents that fail to show up in the top 20 retrieved chunks. Plain embeddings missed 5.7% of them. Adding that generated context dropped it to 3.7%. Also running BM25, a long-established keyword-matching algorithm, over the same enriched chunks took it to 2.9%, and adding a reranking step reached 1.9%.
It isn't free — you make a model call per chunk when you build the index, and the chunks get slightly larger. For a corpus you index once and query constantly, that trade is usually worth it.
What to check when retrieval is missing things
- Pull up the chunks your system actually stored and read a dozen at random. Broken cut points are obvious on sight and invisible in metrics.
- Check whether the answer to a failing question exists inside any single chunk. If it's split across two, that's a chunking problem, not a search problem.
- Check whether a retrieved chunk makes sense standing alone. If it needs its neighbours to be intelligible, return the surrounding section rather than shrinking the chunk.
- Check the long documents separately. They're where structural splitting most often falls back to a blind cut.
- Re-chunking means re-embedding everything, so test on a sample before committing to a change across the whole corpus.