AI in Motion

Generative AIAdvanced1:21 video5 chapters

Mixture of Experts — lecture notes

A router sends each token to a few specialised experts, so a model can have many parameters while using only a fraction per token.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Mixture of Experts

What if a model could be huge, but only use a small part of itself for each word? That is the idea behind mixture of experts.

0:112. Routing tokens

Routing tokens — Mixture of Experts

Inside each mixture of experts layer there are several feed-forward networks, the experts, and a small router. For every token, the router scores all eight experts and sends the token to the best two. Their outputs are combined as a weighted sum. Different tokens go to different experts.

0:323. Dense vs MoE

Dense vs MoE — Mixture of Experts

In a dense model every parameter works on every token. In a mixture of experts, only the chosen experts run, so total capacity grows without the compute growing as fast. For example, Mixtral 8 by 7B uses two of its eight experts for each token.

0:514. Challenges

Challenges — Mixture of Experts

There are challenges. The router must spread tokens fairly so some experts are not overloaded, which is handled with a load balancing loss. All experts must still fit in memory, and moving tokens between GPUs adds communication.

1:075. Recap

Recap — Mixture of Experts

To recap. A router sends each token to a few experts, their outputs are combined, and the model gains capacity without proportional compute, as long as load and memory are managed.

Key takeaways

  • An MoE layer has several expert networks and a router.
  • Each token is sent to its top-k experts (often 2) and their outputs are combined.
  • Total parameters grow while per-token compute stays much smaller.
  • Load balancing, memory and communication are key challenges.

Check yourself

  1. What does the router do in an MoE layer?
    Show answer

    Chooses which experts process each token — It scores experts and picks the top ones.

  2. Why can MoE models be cheaper per token than dense models of the same size?
    Show answer

    Only a few experts run for each token — Sparse activation saves compute.

  3. What does a load-balancing loss prevent?
    Show answer

    A few experts receiving almost all tokens — It encourages even expert usage.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/mixture-of-experts.html