LoRA and Parameter-Efficient Fine-Tuning — lecture notes
Fine-tune a huge model by training two tiny matrices. See why LoRA needs well under 1% of the parameters of full fine-tuning.
0:001. Introduction

Fine-tuning every weight of a billion-parameter model needs huge memory. LoRA, low-rank adaptation, gets similar results by training tiny extra matrices instead.
0:102. Low-rank adapters

Take one weight matrix, 4096 by 4096. That is about 16.8 million parameters. LoRA freezes it and adds the product of two thin matrices, B and A, with rank 8. They hold only 65,536 parameters, about 0.39 percent of the original, yet they can steer the model to a new task.
0:313. The formula

The adapted weight equals the frozen original plus B times A. Because the rank r is small, the update has two times d times r parameters instead of d squared. After training, the update can even be merged into W, so there is no extra cost at inference.
0:524. Why it matters

LoRA needs far less memory, produces small adapter files, and lets one base model switch between many tasks by swapping adapters. QLoRA combines it with four-bit quantisation, making fine-tuning possible on a single GPU.
1:075. Recap

To recap. Freeze the big weights, train small low-rank matrices, and get swappable adapters that use a tiny fraction of the parameters.
Key takeaways
- LoRA freezes pre-trained weights and learns a low-rank update B·A.
- For a 4096 × 4096 matrix with rank 8: 65,536 trainable parameters instead of 16,777,216 (0.39%).
- Adapters are small and swappable, and can be merged for inference.
- QLoRA combines LoRA with 4-bit quantisation.
Check yourself
- What does LoRA train?
Show answer
Small low-rank matrices added to frozen weights — The original weights stay frozen.
- How many parameters does a rank-8 LoRA update have for a 4096 × 4096 matrix?
Show answer
65,536 — 2 × 4096 × 8 = 65,536.
- What does QLoRA add?
Show answer
4-bit quantisation of the base model — Quantising the frozen model saves memory.
Go deeper
- LoRA and Parameter-Efficient Fine-Tuning (PEFT) · The AI Lecture Hall
- Fine-Tuning Pretrained Language Models: A Practical Guide · The AI Lecture Hall
- Quantising Large Language Models for Efficient Inference · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/lora-and-parameter-efficient-fine-tuning.html