Semantic Search

Ordinary search matches words. Ask it "how do I get my money back" and a page titled "Refund policy" never comes up, because the two share no words at all. Semantic search matches on meaning instead, so that page does come up. Nothing about the question had to be phrased the way the document was.

In this guide
  1. How a machine compares meaning
  2. Semantic search or vector search?
  3. Where it's worse than keyword search
  4. Its role in RAG
  5. FAQ

How a machine compares meaning

Meaning isn't something a computer can compare directly, so semantic search converts it into something that can be: position. Every piece of text gets turned into an embedding — a long list of numbers chosen so that text with similar meaning ends up with similar numbers. Once meaning is a position, "related" becomes a distance you can measure.

Three steps, and the first happens long before anyone searches:

  1. Ahead of time, every document in the collection is converted to an embedding and stored. A vector is just the mathematical name for that list of numbers, and a vector database is a store built to search a great many of them quickly.
  2. At search time, the query gets converted the same way, by the same model, into one embedding.
  3. Then the system finds the stored embeddings closest to the query's, and those are the results.

"How do I get my money back" and "Refund policy" land near each other because the model that produced both has seen enough text to place them in similar positions. No shared word is required, which is exactly the thing keyword matching can't do.

In current practice the two terms are used for the same thing, and you can treat them as interchangeable without getting into trouble. If you want the distinction: semantic search names the goal, which is searching by meaning, while vector search names the technique nearly everyone now uses to achieve it. The goal is older than the technique — systems tried to search by meaning using hand-built dictionaries of related terms and structured knowledge long before embeddings made it work well.

Exact strings — a part number, an error code, a surname. The embeddings page covers why those break, and Vector Search vs Keyword Search covers what to do about it, including when running both and merging the results is worth the extra machinery.

Its role in RAG

Semantic search is the retrieval step in RAG. When a RAG system "looks things up before answering," this is the lookup — the user's question becomes an embedding, the closest stored passages come back, and those passages go into the prompt as material for the answer. The quality ceiling of the whole system sits here: a passage that search never returns cannot be used, no matter how good the model reading the results is.

So when a RAG system answers badly, this is the first place to look rather than the prompt. The two levers that move it most are how the documents were split before being embedded — see chunking — and whether a second pass reorders the results before they reach the model, which is reranking.

FAQ

Does semantic search understand the query the way a person does?

No, and the word "semantic" oversells it. The system compares positions produced by a model trained on a lot of text. It reliably places related wording near related wording, which is enough to be useful, but there's no comprehension of the question behind it.

That's why it can confidently return something topically adjacent and wrong — including the opposite of what you asked, since opposites are highly related. The embeddings page works through that failure in detail.

Practice interview questions on semantic search →