LoRA and Parameter-Efficient Fine-Tuning
Fine-tune a huge model by training two tiny matrices. See why LoRA needs well under 1% of the parameters of full fine-tuning.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. Fine-tuning every weight of a billion-parameter model needs huge memory. LoRA, low-rank adaptation, gets similar results by training tiny extra matrices instead.
Low-rank adapters. Take one weight matrix, 4096 by 4096. That is about 16.8 million parameters. LoRA freezes it and adds the product of two thin matrices, B and A, with rank 8. They hold only 65,536 parameters, about 0.39 percent of the original, yet they can steer the model to a new task.
The formula. The adapted weight equals the frozen original plus B times A. Because the rank r is small, the update has two times d times r parameters instead of d squared. After training, the update can even be merged into W, so there is no extra cost at inference.
Why it matters. LoRA needs far less memory, produces small adapter files, and lets one base model switch between many tasks by swapping adapters. QLoRA combines it with four-bit quantisation, making fine-tuning possible on a single GPU.
Recap. To recap. Freeze the big weights, train small low-rank matrices, and get swappable adapters that use a tiny fraction of the parameters.