AI in Motion

Generative AIIntermediate1:17 video5 chapters

LoRA and Parameter-Efficient Fine-Tuning — lecture notes

Fine-tune a huge model by training two tiny matrices. See why LoRA needs well under 1% of the parameters of full fine-tuning.

▶ Watch the animated lecture

0:001. Introduction

Introduction — LoRA and Parameter-Efficient Fine-Tuning

Fine-tuning every weight of a billion-parameter model needs huge memory. LoRA, low-rank adaptation, gets similar results by training tiny extra matrices instead.

0:102. Low-rank adapters

Low-rank adapters — LoRA and Parameter-Efficient Fine-Tuning

Take one weight matrix, 4096 by 4096. That is about 16.8 million parameters. LoRA freezes it and adds the product of two thin matrices, B and A, with rank 8. They hold only 65,536 parameters, about 0.39 percent of the original, yet they can steer the model to a new task.

0:313. The formula

The formula — LoRA and Parameter-Efficient Fine-Tuning

The adapted weight equals the frozen original plus B times A. Because the rank r is small, the update has two times d times r parameters instead of d squared. After training, the update can even be merged into W, so there is no extra cost at inference.

0:524. Why it matters

Why it matters — LoRA and Parameter-Efficient Fine-Tuning

LoRA needs far less memory, produces small adapter files, and lets one base model switch between many tasks by swapping adapters. QLoRA combines it with four-bit quantisation, making fine-tuning possible on a single GPU.

1:075. Recap

Recap — LoRA and Parameter-Efficient Fine-Tuning

To recap. Freeze the big weights, train small low-rank matrices, and get swappable adapters that use a tiny fraction of the parameters.

Key takeaways

  • LoRA freezes pre-trained weights and learns a low-rank update B·A.
  • For a 4096 × 4096 matrix with rank 8: 65,536 trainable parameters instead of 16,777,216 (0.39%).
  • Adapters are small and swappable, and can be merged for inference.
  • QLoRA combines LoRA with 4-bit quantisation.

Check yourself

  1. What does LoRA train?
    Show answer

    Small low-rank matrices added to frozen weights — The original weights stay frozen.

  2. How many parameters does a rank-8 LoRA update have for a 4096 × 4096 matrix?
    Show answer

    65,536 — 2 × 4096 × 8 = 65,536.

  3. What does QLoRA add?
    Show answer

    4-bit quantisation of the base model — Quantising the frozen model saves memory.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/lora-and-parameter-efficient-fine-tuning.html