Retrieval-Augmented Generation (RAG) — lecture notes
Give a language model the right documents at the right moment: retrieve relevant passages, then generate a grounded answer with citations.
0:001. Introduction

Language models do not know your company’s policies or last week’s news, and they sometimes make things up. Retrieval augmented generation, or RAG, fixes both by looking things up first.
0:132. The RAG pipeline

The question is embedded into a vector. A vector store finds the most similar passages from your documents: the leave policy, and the carry-over rule. These chunks are placed into the prompt, and the language model writes an answer based on them: twenty days of paid leave, with up to five days carrying over, citing its sources.
0:373. Retrieval

Retrieval is semantic search. Because it matches meaning, a question like I can’t sign in finds the password reset guide even though they share no words.
0:494. Best practices

Quality depends on the details. Split documents sensibly, use good embeddings with hybrid search, re-rank results, and instruct the model to answer only from the retrieved sources, with citations.
1:025. RAG vs fine-tuning

RAG adds knowledge at question time and is easy to keep up to date. Fine-tuning changes the model’s behaviour and style. Many real systems use both.
1:136. Recap

To recap. Retrieve relevant passages, add them to the prompt, and generate a grounded answer with citations. RAG keeps answers current and reduces hallucinations.
Key takeaways
- RAG retrieves relevant passages and adds them to the prompt before generating.
- Retrieval is usually semantic (embedding) search, often hybrid with keywords.
- Grounding in retrieved text reduces hallucinations and enables citations.
- RAG adds knowledge; fine-tuning changes behaviour — they are often combined.
Check yourself
- What happens before generation in RAG?
Show answer
Relevant documents are retrieved and added to the prompt — Retrieval provides grounding context.
- Why does RAG reduce hallucinations?
Show answer
The model answers from retrieved source text — Grounded answers rely on real documents.
- How do you update a RAG system with new policies?
Show answer
Update the documents in the store — Knowledge lives in the documents.
Go deeper
- Retrieval-Augmented Generation (RAG): Grounding LLMs in Your Documents · The AI Lecture Hall
- Vector Databases and Approximate Nearest Neighbour Search · The AI Lecture Hall
- Building an LLM Application End to End: From Idea to Production · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/retrieval-augmented-generation.html