Reranking

Reranking is a second pass over search results. Your first search returns a shortlist — say the closest 100 passages — and a reranker scores each one again, more carefully, then reorders them so the best few go to the model. It is a precision step bolted onto a recall step.

In this guide
  1. Why there are two stages
  2. It can only reorder what retrieval found
  3. It can also make things worse
  4. What it costs
  5. When it's worth adding
  6. What's actually been measured

Why there are two stages

Ordinary retrieval is fast for one specific reason: the document side can be computed before the question exists. You turn every chunk into an embedding — a list of numbers positioned so that similar meaning lands nearby — once, ahead of time. At query time you only convert the question and look for the nearest stored points. All the expensive work already happened.

A reranker gives that up deliberately. The usual design is a cross-encoder — a model that reads the query and one candidate passage together, as a single input, and outputs a relevance score for that pairing. Because the query is part of the computation, nothing can be precomputed. The score doesn't exist until both halves are present.

That's a much better judgement of relevance, and it's why you can't use it to search. Scoring the query against a million chunks means a million model calls. So the pattern is: something cheap that can look at everything, then something expensive that looks only at a shortlist.

It can only reorder what retrieval found

This is the part people get wrong, and it's a hard architectural limit rather than a quality issue. A reranker reorders. It cannot retrieve. If the passage that answers the question wasn't in the shortlist the first stage handed over, the reranker never sees it, and no amount of model quality recovers it.

The practical consequence is a diagnostic. If the answer-bearing chunk is missing from your first stage's top 100 on a third of your test questions, swapping rerankers cannot fix that third — ever. That's a retrieval problem, and it lives in your embeddings, your chunking, whether you're also doing keyword matching, or the query itself.

Which means you have to measure the two stages separately. Ask first: is the right passage anywhere in the shortlist? Only if the answer is yes does ranking become the thing worth improving.

It can also make things worse

Adding a second stage is not automatically an improvement, and this surprises people. A reranker trained largely on general web-search relevance has learned what a plausible answer looks like in that setting. That isn't necessarily what counts as evidence in yours — an error code buried in API documentation doesn't resemble a well-written web answer at all.

So a mismatched reranker can take a first-stage ranking that had the right passage at position one and push it to position three, with something plausible and wrong above it. The second stage didn't just fail to help; it introduced an error.

Two things follow. Measure against the retriever alone as your baseline, not against nothing — if the reranked ordering isn't better than what you already had, you're paying for a downgrade. And a longer shortlist isn't automatically better either: reranking 200 candidates can land behind reranking 100, because you've handed the model more opportunities to promote something that merely looks right.

What it costs

The work is candidate count times one model call each. Rerank 100 and you score 100 pairs. Double the shortlist and you roughly double the work — though how much wall-clock time that actually adds depends on batching, hardware and passage length enough that you should measure your own setup rather than extrapolate.

Where that lands varies enormously, and the two ends aren't the same decision at all. A small cross-encoder running locally can add tens of milliseconds and no per-query fee. A large hosted reranker means a network round trip and a charge on every search, which at volume is a real line item.

Then there's the cost that never appears on a bill: another production component to monitor, version, capacity-plan, and have a fallback for when it's slow or unavailable.

When it's worth adding

Check where the right passage currently sits in your first-stage results.

  • Not in the top 100. Retrieval problem. Fix that; reranking is irrelevant here.
  • Somewhere around rank 20 to 50, but you only pass 5 to the model. This is precisely the gap reranking exists to close, and the best case for adding one.
  • Already at rank 1 to 3. Little room to gain, and a genuine risk of making it worse. Skip it.

Making that call requires having measured your first stage on its own, which most systems haven't. That's why "add a reranker" gets reached for as a generic fix when the actual problem sits upstream of it.

What's actually been measured

Two results worth knowing, measuring different things.

Nogueira and Cho's 2019 paper established the retrieve-then-rerank pattern using BERT, reporting a 27% relative improvement over the previous best system on a standard passage-ranking benchmark. That's a ranking result — how well passages get ordered — not a statement about answer quality.

Anthropic's contextual retrieval experiment is closer to a real RAG setting. Retrieving 150 candidates and reranking down to 20 moved the share of relevant documents missing from the final set from 2.9% to 1.9%. Note carefully what that is: reranking was the last step in a stack that already included generated per-chunk context and keyword search, and the number measures retrieval failure, not whether answers got better.

Neither figure predicts what a reranker will do on your corpus. They establish that the technique can matter, which is a different and more modest claim. Be wary of anyone converting a ranking benchmark into a percentage improvement in answer accuracy — those are not the same measurement.

Practice interview questions on reranking →