AI in Motion

Mixture of Experts

Generative AIAdvanced1:215 chapters

A router sends each token to a few specialised experts, so a model can have many parameters while using only a fraction per token.

๐Ÿ“„ Illustrated notes ยท every chapter as a picture ยท printable

Shortcuts: Space play/pause ยท โ†/โ†’ 5 s ยท N/P chapter ยท M voice ยท C subtitles ยท F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 What does the router do in an MoE layer?
Q2 Why can MoE models be cheaper per token than dense models of the same size?
Q3 What does a load-balancing loss prevent?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. What if a model could be huge, but only use a small part of itself for each word? That is the idea behind mixture of experts.

Routing tokens. Inside each mixture of experts layer there are several feed-forward networks, the experts, and a small router. For every token, the router scores all eight experts and sends the token to the best two. Their outputs are combined as a weighted sum. Different tokens go to different experts.

Dense vs MoE. In a dense model every parameter works on every token. In a mixture of experts, only the chosen experts run, so total capacity grows without the compute growing as fast. For example, Mixtral 8 by 7B uses two of its eight experts for each token.

Challenges. There are challenges. The router must spread tokens fairly so some experts are not overloaded, which is handled with a load balancing loss. All experts must still fit in memory, and moving tokens between GPUs adds communication.

Recap. To recap. A router sends each token to a few experts, their outputs are combined, and the model gains capacity without proportional compute, as long as load and memory are managed.