Mixture of Experts — lecture notes
A router sends each token to a few specialised experts, so a model can have many parameters while using only a fraction per token.
0:001. Introduction

What if a model could be huge, but only use a small part of itself for each word? That is the idea behind mixture of experts.
0:112. Routing tokens

Inside each mixture of experts layer there are several feed-forward networks, the experts, and a small router. For every token, the router scores all eight experts and sends the token to the best two. Their outputs are combined as a weighted sum. Different tokens go to different experts.
0:323. Dense vs MoE

In a dense model every parameter works on every token. In a mixture of experts, only the chosen experts run, so total capacity grows without the compute growing as fast. For example, Mixtral 8 by 7B uses two of its eight experts for each token.
0:514. Challenges

There are challenges. The router must spread tokens fairly so some experts are not overloaded, which is handled with a load balancing loss. All experts must still fit in memory, and moving tokens between GPUs adds communication.
1:075. Recap

To recap. A router sends each token to a few experts, their outputs are combined, and the model gains capacity without proportional compute, as long as load and memory are managed.
Key takeaways
- An MoE layer has several expert networks and a router.
- Each token is sent to its top-k experts (often 2) and their outputs are combined.
- Total parameters grow while per-token compute stays much smaller.
- Load balancing, memory and communication are key challenges.
Check yourself
- What does the router do in an MoE layer?
Show answer
Chooses which experts process each token — It scores experts and picks the top ones.
- Why can MoE models be cheaper per token than dense models of the same size?
Show answer
Only a few experts run for each token — Sparse activation saves compute.
- What does a load-balancing loss prevent?
Show answer
A few experts receiving almost all tokens — It encourages even expert usage.
Go deeper
- Mixture of Experts: Scaling Parameters Without Scaling Compute · The AI Lecture Hall
- LLM Inference: KV Caching, Batching and Speculative Decoding · The AI Lecture Hall
- Distributed Training: Data, Model, Pipeline and Sharded Parallelism · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/mixture-of-experts.html