AI in Motion

Deep LearningBeginner1:22 video6 chapters

Activation Functions: Sigmoid, Tanh and ReLU — lecture notes

Why networks need non-linearity, and how sigmoid, tanh, ReLU, Leaky ReLU and GELU differ — drawn live.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Activation Functions: Sigmoid, Tanh and ReLU

Without activation functions, even a hundred-layer network would only learn straight-line relationships. Activation functions add the bends that make deep learning powerful.

0:102. Why they matter

Why they matter — Activation Functions: Sigmoid, Tanh and ReLU

Stacking linear layers just gives another linear function, so depth would be pointless. Non-linear activations break that, letting networks model curves, corners and complex decision boundaries.

0:223. Four classics

Four classics — Activation Functions: Sigmoid, Tanh and ReLU

The sigmoid squeezes values between zero and one. Tanh is similar but centred on zero. ReLU outputs zero for negative inputs and passes positive ones straight through. Leaky ReLU keeps a small slope for negatives, so neurons never die completely.

0:394. Vanishing gradients

Vanishing gradients — Activation Functions: Sigmoid, Tanh and ReLU

Sigmoid and tanh flatten out for large inputs, where their slope is almost zero. During training, those tiny slopes multiply through many layers and gradients vanish. ReLU keeps a slope of one for positive inputs, so gradients flow much better.

0:565. Modern choices

Modern choices — Activation Functions: Sigmoid, Tanh and ReLU

Modern networks often use ReLU, or smooth variants like GELU, which is common in Transformers. For outputs we use sigmoid for yes or no answers and softmax for choosing between many classes.

1:106. Recap

Recap — Activation Functions: Sigmoid, Tanh and ReLU

To recap. Activations add non-linearity. Sigmoid and tanh can make gradients vanish. ReLU and its variants are the usual default. And output layers use sigmoid or softmax.

Key takeaways

  • Without non-linear activations, stacked layers collapse into one linear function.
  • Sigmoid and tanh saturate, which can cause vanishing gradients.
  • ReLU (and Leaky ReLU, GELU) are standard choices for hidden layers.
  • Use sigmoid for binary outputs and softmax for multi-class outputs.

Check yourself

  1. What is ReLU(−3)?
    Show answer

    0 — ReLU outputs max(0, x).

  2. Why can sigmoid cause vanishing gradients?
    Show answer

    It is nearly flat for large positive or negative inputs — Flat regions have tiny slopes, so gradients shrink.

  3. Which activation is common in Transformer models?
    Show answer

    GELU — GELU is widely used in Transformers.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/activation-functions.html