In this guide
RAG Interview Questions and Answers
Thirty-five questions covering the complete RAG flow — documents, ingestion, chunking, embeddings, retrieval, reranking, context, and the model's final answer. The goal isn't memorizing "RAG = retrieval + generation," it's being able to design, troubleshoot, and evaluate a real pipeline.
User Question
↓
Retrieve relevant information
↓
Add information to context
↓
LLM
↓
Answer
RAG Fundamentals
1. What is RAG?
RAG retrieves relevant information and gives it to a model as context before it answers, instead of relying only on what the model learned during training. Learn more: RAG.
2. Why is RAG used with LLMs?
Because an LLM's knowledge is fixed at training time and limited to what fits in its context window — RAG lets it answer using current, private, or simply too-large-to-paste-in information by fetching only the relevant part at the moment it's needed.
3. How does a basic RAG system work?
A question comes in, the system searches a store of pre-processed documents for the passages most relevant to it, and hands those passages to the model alongside the question. The model answers from what it was given, not only from its training data.
4. What are the main components of a RAG pipeline?
A way to store and search the source material — usually a vector database — a retrieval step that finds the relevant pieces for a given question, and the model itself, which receives the question plus whatever was retrieved and produces the final answer.
5. What is the difference between RAG and simply putting a document in the prompt?
Pasting a document in directly works fine when it's small enough to fit comfortably and doesn't change often. RAG earns its place once the material is too large to paste in every time, changes regularly, or needs to be searched rather than read in full for every question.
6. Does RAG change or retrain the LLM?
No. RAG changes what the model sees at the moment a question is asked; the model itself is never touched or retrained. That's also why swapping in retrieved material works with any model, immediately, with no training step involved.
7. When should you use RAG?
When the model needs current or private information it wasn't trained on, that information changes too often to retrain around, or the answer needs to be traceable back to a real source.
8. When might you not need RAG?
When the relevant material is small enough to paste directly into a request and doesn't change often, or when the task doesn't depend on any specific external fact at all. Building retrieval machinery for a case that doesn't need it adds a system that can fail for no real benefit.
Documents, Ingestion and Chunking
9. What is document ingestion in RAG?
The process of taking source documents in and preparing them for retrieval — parsing their content, cleaning it up, and splitting it into chunks — before anything can actually be searched.
10. Why are documents split into chunks?
Because embedding models — the models that convert text into the numbers retrieval searches over — and context windows both have size limits, and retrieval works better on a focused piece of text than on an entire document. Learn more: Chunking.
11. What is chunk size?
How much text goes into each stored, searchable piece of a document. Too small and a chunk may lose the surrounding meaning it needs; too large and it may bury the specific fact that matters among too much else.
12. What is chunk overlap?
Repeating a small amount of text from the end of one chunk at the start of the next, so a fact sitting right on a chunk boundary still appears whole in at least one of them.
13. What happens if chunks are too small?
A chunk can lose the context it needs to make sense on its own — a sentence with no surrounding explanation, or a number with no label saying what it refers to.
14. What happens if chunks are too large?
Retrieval becomes less precise, because a chunk covering many different points can match a query on an unrelated part of itself, and the specific fact that actually answers the question gets diluted among everything else in that same chunk.
15. How would you choose a good chunking strategy?
Split along the document's own real structure — sections, headings — rather than a fixed character count wherever possible. A 100-page employee handbook shouldn't be treated as one giant searchable unit; dividing it into its actual sections (annual leave, sick leave, expenses, parental leave) keeps each chunk about one coherent topic. The goal is always the same: preserve useful meaning while keeping information retrievable.
Embeddings and Retrieval
16. What are embeddings in RAG?
A way of representing content as numbers positioned so that content with similar meaning ends up near other content with similar meaning. A search for "how many holidays do employees receive" can retrieve a section titled "Annual Leave Entitlement" even though the wording doesn't match at all, because the embeddings land near each other in meaning. Learn more: Embeddings.
17. What is a vector database?
A database built to quickly find the embeddings most similar to a given query out of millions of them. Learn more: Vector Database.
18. What is vector search?
Searching by comparing embeddings to find the ones closest in meaning to a query's embedding, rather than matching exact words.
19. What is semantic search?
Search that matches by meaning instead of exact keywords — the retrieval technique embeddings actually enable.
20. What is the difference between semantic search and keyword search?
Keyword search matches the literal words in a query against the literal words in stored content. Semantic search matches by meaning, so different wording for the same idea can still be found. Each catches things the other misses — an exact product code is a keyword-search strength, a differently-worded question is a semantic-search strength.
21. What is hybrid search?
Combining semantic search with keyword search, so an exact term that meaning-based search alone might miss — a code, an acronym — still gets found, alongside results matched by meaning.
22. What is metadata filtering?
Narrowing a search to only the content that matches certain attributes — a department, a document type, a date range, a specific user's permissions — before or alongside the similarity search itself.
23. What does top-k retrieval mean?
Retrieving the k most similar results to a query — the top 5, or the top 20 — rather than every match above some threshold. How large k should be is itself a real tuning decision: too small and the right passage might not make the cut, too large and irrelevant results start crowding out the ones that matter.
24. What is reranking?
A second, more careful scoring pass over a shortlist a first-stage search already narrowed down, to push the genuinely most relevant results higher before they reach the model. Learn more: Reranking.
From Retrieval to the LLM
25. What happens after relevant chunks are retrieved?
They get added to the model's context alongside the user's question, and the model generates its answer from that combined input rather than from training data alone.
26. How is retrieved information added to an LLM's context?
It's inserted into the prompt sent to the model, typically alongside instructions on how to use it — for instance, to answer only from the material provided, or to say so if the provided material doesn't cover the question.
27. Why can retrieving too much information hurt the final answer?
Because every retrieved chunk shares the same context budget as everything else in the request, and irrelevant or excessive content can crowd out or distract from the passage that actually answers the question.
28. What should happen when no relevant information is found?
It depends on the application, but forcing an answer anyway usually isn't the right default. Saying that sufficient information wasn't found, asking for clarification, checking another source, or escalating to a person are all better options than generating a confident answer with nothing real behind it.
RAG vs Other Approaches
29. What is the difference between RAG and fine-tuning?
RAG changes the information supplied at the moment a question is asked. Fine-tuning changes the model's behavior through additional training beforehand, and is not simply "putting knowledge into the model" — training on text teaches patterns and style, not a reliable lookup table of facts. Learn more: RAG vs Fine-Tuning.
30. What is the difference between RAG and a long context window?
A long context window lets you hand over a much larger piece of material directly, skipping the retrieval step. RAG stays useful once the material is too large to fit even in a big window, changes too often to keep re-pasting, or needs a traceable citation. Learn more: RAG vs Long Context.
31. What is the difference between RAG and AI memory?
RAG retrieves from a body of documents to answer a specific question. AI memory is information an application deliberately saves and reuses across separate sessions, like a user's own preferences or facts from an earlier conversation. They can be used together.
Practical RAG Problems
32. Your RAG system retrieves irrelevant chunks. What would you investigate?
Read the actual retrieved chunks against the actual question rather than trusting aggregate metrics. Check whether the embedding model captures the right kind of similarity for this domain, whether chunks are cut in a way that preserves meaning, and whether the query itself matches how the source material is worded. This is a retrieval-stage problem — not something to blame on the model that generates the answer.
33. Your RAG system retrieves the correct information but still produces a poor answer. What would you investigate?
Retrieval succeeding and generation succeeding are separate problems. Check whether the specific relevant part actually made it into the model's context, whether irrelevant retrieved content is crowding it out, whether the instructions are clear about how to use the retrieved material, whether multiple retrieved passages conflict, and whether the model itself is capable enough for the task.
34. How would you evaluate a RAG system?
Separately. Evaluate retrieval quality — did the system find the information actually needed — and answer quality — given that information, did the model produce a correct, useful response — as two different measurements, not one combined score. Depending on the application, faithfulness to sources, citation quality, latency — how long a request takes to get a response back — cost, and whether permissions were respected can matter too. Learn more: RAG Evaluation.
35. How would you improve a RAG system that performs poorly?
Find out first whether the failure is a retrieval problem or a generation problem — they need different fixes. For retrieval: better chunking, a better embedding model, hybrid search, metadata filtering, or reranking. For generation: clearer instructions about using the retrieved material, less noisy context, or a more capable model. Treating every failure as "the LLM is bad at this" skips the diagnosis most RAG problems actually need.
Designing a RAG system at scale
Scenario: a company has 500,000 internal documents. How would you design a RAG system so employees can ask questions about them? Work through it as a real pipeline rather than jumping to "put the documents in a vector database": identify and access the actual document sources, parse and clean the content, preserve useful metadata like department and access level, divide documents into meaningful chunks, create embeddings, index everything for retrieval, then at query time process the user's question, retrieve relevant candidates, apply filters or hybrid retrieval where useful, rerank if needed, construct the context, ask the model to answer from it, provide source references where appropriate, and evaluate retrieval and answer quality on an ongoing basis. The specific tools matter far less than working through each stage deliberately.
Permissions matter as much as relevance. If an employee isn't allowed to read a confidential HR document, the system shouldn't retrieve it into their context just because it's semantically relevant to their question. Relevant does not automatically mean authorized — access control has to apply at retrieval time, not just at the point someone tries to open the original document.
Distinctions worth remembering
RAG is not a vector database. A vector database can be one component of a RAG architecture, not the whole thing.
RAG is not semantic search. Semantic search can be part of retrieval, but RAG includes the larger process of retrieving information and using it during generation.
RAG is not fine-tuning. They solve different problems and can be used together.
More retrieved context is not automatically a better answer. Relevant context matters more than the amount of retrieved text.
Retrieval quality and generation quality are not the same thing. Both need to be evaluated, separately.
Going deeper on one area
For a broader revision pass: AI Interview Questions and Answers. For agent design: Agentic AI Interview Questions. For more scenario-based retrieval questions: the original RAG scenario cluster. For designing a whole system, not just the RAG piece: AI System Design Interview Questions.