Retrieval-Augmented Generation (RAG)
Give a language model the right documents at the right moment: retrieve relevant passages, then generate a grounded answer with citations.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. Language models do not know your company’s policies or last week’s news, and they sometimes make things up. Retrieval augmented generation, or RAG, fixes both by looking things up first.
The RAG pipeline. The question is embedded into a vector. A vector store finds the most similar passages from your documents: the leave policy, and the carry-over rule. These chunks are placed into the prompt, and the language model writes an answer based on them: twenty days of paid leave, with up to five days carrying over, citing its sources.
Retrieval. Retrieval is semantic search. Because it matches meaning, a question like I can’t sign in finds the password reset guide even though they share no words.
Best practices. Quality depends on the details. Split documents sensibly, use good embeddings with hybrid search, re-rank results, and instruct the model to answer only from the retrieved sources, with citations.
RAG vs fine-tuning. RAG adds knowledge at question time and is easy to keep up to date. Fine-tuning changes the model’s behaviour and style. Many real systems use both.
Recap. To recap. Retrieve relevant passages, add them to the prompt, and generate a grounded answer with citations. RAG keeps answers current and reduces hallucinations.