RAG

Retrieval-augmented generation grounds a model's answers in your own data at query time, instead of relying only on what it learned during training.

  • RAG (Retrieval-Augmented Generation) — RAG retrieves relevant documents at query time and feeds them to an LLM as context, so answers are grounded in your own data instead of just training data.
  • Embeddings — an embedding turns text into a list of numbers positioned so similar meanings land near each other.
  • Chunking — splits documents before you index them.
  • Vector Database — finds the stored embeddings nearest a query.
  • Reranking — a reranker re-scores the shortlist from your first search, reading query and document together.
  • RAG Evaluation — scores retrieval quality and generation quality separately, because a RAG system can fail at either stage independently.

Comparisons