AI in Motion

Mathematics for MLDeep diveAdvanced10:49 video31 chapters

Likelihood, Bayes and Information Theory — lecture notes

Why models minimise cross-entropy: maximum likelihood, priors and MAP, Bayesian updating, entropy, cross-entropy and KL divergence — the statistics hiding inside every loss function.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Likelihood, Bayes and Information Theory

Why do classifiers minimise cross entropy? Why does squared error appear in regression? Why do we add weight decay? The answers come from two beautiful fields: statistical estimation and information theory. This deep dive connects them to the loss functions you use every day.

0:182. Parameters and data

Parameters and data — Likelihood, Bayes and Information Theory

A statistical model is a family of probability distributions controlled by parameters, written p of x given theta. A Gaussian has two parameters, a mean and a standard deviation. A neural network is also a probability model, with millions of parameters. Learning means choosing good values for theta.

0:393. Likelihood

Likelihood — Likelihood, Bayes and Information Theory

The likelihood asks: if the parameters were theta, how probable would the data we actually observed be? For independent data points it is the product of each point’s probability. It uses the same formula as probability, but we read it as a function of the parameters with the data held fixed.

1:004. Maximum likelihood

Maximum likelihood — Likelihood, Bayes and Information Theory

Maximum likelihood estimation picks the parameters that make the data most probable. Watch the Gaussian slide along. Each bar shows how likely one data point is under the current curve. The log likelihood, on the right, peaks exactly at the sample mean, two point six five. That is the maximum likelihood estimate.

1:225. Why take logs?

Why take logs? — Likelihood, Bayes and Information Theory

In practice we maximise the log likelihood instead. The logarithm turns a product of many tiny probabilities, which would underflow to zero on a computer, into a sum of manageable numbers. Because the log only increases, the best parameters do not change, and sums are easy to differentiate.

1:436. Loss = negative log-likelihood

Loss = negative log-likelihood — Likelihood, Bayes and Information Theory

Flip the sign and average, and you get a loss to minimise: the negative log likelihood. This single idea generates the standard losses. If you assume Gaussian noise around the prediction, it becomes squared error. If the output is a category, it becomes cross entropy.

2:027. Pause and think

Pause and think — Likelihood, Bayes and Information Theory

Pause and think. A model assigns probability point nine to the correct class for one example, and only point one for another. What is the loss for each? Minus the log of point nine is about point one one, a small loss. Minus the log of point one is about two point three: confident mistakes are punished heavily.

2:268. The shape of the loss

The shape of the loss — Likelihood, Bayes and Information Theory

This curve shows the loss for one example as a function of the probability given to the correct answer. At point nine the loss is tiny. At one half it is about point seven. As the probability approaches zero, the loss shoots up towards infinity, so the model is pushed hard never to be confidently wrong.

2:509. Estimators

Estimators — Likelihood, Bayes and Information Theory

An estimate computed from data is itself random, because a different sample would give a different answer. Its bias asks whether it is right on average, and its variance asks how much it wobbles. The sample variance divides by n minus one instead of n precisely to remove a small bias.

3:1110. Bias and variance

Bias and variance — Likelihood, Bayes and Information Theory

The dartboard picture makes this concrete. Low bias and low variance cluster tightly on the bullseye. High bias misses consistently in one direction. High variance scatters widely. The same trade off governs models: simple models tend to be biased, while very flexible ones tend to have high variance.

3:3211. Priors

Priors — Likelihood, Bayes and Information Theory

Maximum likelihood listens only to the data, which is risky when data is scarce. Bayesian statistics adds a prior: what we believe about the parameters before seeing any data. For example, believing that weights are probably small corresponds to a Gaussian prior centred on zero.

3:5112. Posterior

Posterior — Likelihood, Bayes and Information Theory

Bayes’ theorem combines the two. The posterior, our belief after seeing the data, is proportional to the likelihood times the prior. With little data the prior dominates. With lots of data the likelihood takes over and the prior washes out.

4:0813. Bayesian updating

Bayesian updating — Likelihood, Bayes and Information Theory

Here is Bayesian updating with a coin of unknown bias. We start with a gentle prior centred on one half. As flips arrive, the posterior curve shifts and narrows. After forty flips with twenty eight heads, it concentrates near point seven: more data means more confident beliefs.

4:2814. Pause and think

Pause and think — Likelihood, Bayes and Information Theory

Pause and think. A coin lands heads seven times in ten flips. What is the maximum likelihood estimate of its bias? Seven tenths. With a gentle Beta two, two prior, the MAP estimate is eight twelfths, about point six seven, pulled a little towards one half by the prior belief.

4:4915. MAP estimation

MAP estimation — Likelihood, Bayes and Information Theory

Often we want a single best answer: the peak of the posterior, called the maximum a posteriori, or MAP, estimate. Its loss is the negative log likelihood minus the log prior. With a Gaussian prior on the weights, that extra term is exactly L two regularisation, also known as weight decay.

5:1116. Priors as regularisers

Priors as regularisers — Likelihood, Bayes and Information Theory

Different priors give different regularisers. A Gaussian prior gives the L two penalty, a circle, which shrinks all weights a little. A Laplace prior gives the L one penalty, a diamond, whose corners push some weights to exactly zero. Here the L one solution lands on a corner: a sparse model.

5:3317. Pause and think

Pause and think — Likelihood, Bayes and Information Theory

Pause and think. You have twenty training examples and ten thousand features. Should you rely on maximum likelihood alone? No. With so little data, maximum likelihood will happily overfit. A prior, in other words regularisation such as L one or L two, is essential, or better still, collect more data.

5:5418. Information

Information — Likelihood, Bayes and Information Theory

Now to information theory, founded by Claude Shannon. The information in an outcome is its surprise: minus the log of its probability. A fair coin flip carries one bit. An event with a one in eight chance carries three bits. Something certain carries no information at all.

6:1419. Entropy

Entropy — Likelihood, Bayes and Information Theory

Entropy is the average surprise of a distribution. Four equally likely outcomes give the maximum, two bits. As the distribution becomes peaked, entropy falls: one point eight five bits, then one point three six, and finally only about point two four bits when one outcome is almost certain.

6:3420. Entropy in ML

Entropy in ML — Likelihood, Bayes and Information Theory

Formally, entropy is minus the sum of p log p. It is the minimum average number of bits needed to encode outcomes from that distribution, which is why it underlies compression. Decision trees use it too: they choose splits that reduce entropy the most, called information gain.

6:5421. Information gain

Information gain — Likelihood, Bayes and Information Theory

Here a decision tree splits the data. A good split creates groups that are purer, each dominated by one class, so the entropy after the split is lower. The reduction in entropy is the information gain, and the tree greedily picks the split with the largest gain.

7:1422. Cross-entropy and KL

Cross-entropy and KL — Likelihood, Bayes and Information Theory

Now suppose the truth is P, in blue, but our model predicts Q, in pink. Cross entropy measures the average surprise when outcomes come from P but we encode them using Q. It starts at two point six eight bits, well above the entropy of P, one point six five. The gap is the KL divergence.

7:3823. The key identity

The key identity — Likelihood, Bayes and Information Theory

This identity is the punchline. Cross entropy equals the entropy of the truth plus the KL divergence. The entropy of the data is fixed, so minimising cross entropy is exactly minimising the KL divergence, pulling the model’s distribution towards the true one. By the end here it is only point zero one five bits.

8:0124. Cross-entropy in LLMs

Cross-entropy in LLMs — Likelihood, Bayes and Information Theory

Language models are trained with exactly this loss. At every position the model predicts a distribution over the next token, and the loss is minus the log probability of the actual next token. Averaged over a sentence, it is the cross entropy, and its exponential is called perplexity.

8:2125. Pause and think

Pause and think — Likelihood, Bayes and Information Theory

Pause and think. Is the KL divergence symmetric? Is KL of P given Q the same as KL of Q given P? No. One direction heavily penalises the model for missing outcomes the truth considers likely, the other penalises it for putting probability where the truth has little. They generally differ.

8:4326. Perplexity

Perplexity — Likelihood, Bayes and Information Theory

Language model quality is often reported as perplexity, the exponential of the cross entropy. It has a lovely interpretation: the effective number of equally likely words the model is choosing between at each step. A perplexity of ten means the model is as uncertain as if it were picking among ten words.

9:0527. Mutual information

Mutual information — Likelihood, Bayes and Information Theory

One more quantity: mutual information, how much knowing one variable reduces our uncertainty about another. It is zero exactly when the two are independent. It is used to rank features, and contrastive methods like CLIP can be understood as maximising a bound on the mutual information between images and captions.

9:2628. Where KL appears

Where KL appears — Likelihood, Bayes and Information Theory

KL divergence appears across machine learning. It is inside every cross entropy loss. It keeps the latent codes of variational autoencoders tidy. It stops models tuned with human feedback from drifting too far from the original. It drives knowledge distillation, and related scores monitor data drift in production.

9:4629. In code

In code — Likelihood, Bayes and Information Theory

In code, these are one liners. With the truth P and a poor model Q, the entropy is one point six five bits, the cross entropy two point six eight, and the KL divergence, their difference, is about one point zero four bits. Deep learning libraries compute the same thing, usually with natural logs.

10:0930. Putting it together

Putting it together — Likelihood, Bayes and Information Theory

Putting it all together: one idea explains many losses. Squared error is the negative log likelihood under Gaussian noise. Cross entropy is the negative log likelihood of categories. Weight decay is a Gaussian prior, and the L one penalty is a Laplace prior that encourages sparsity.

10:2931. Recap

Recap — Likelihood, Bayes and Information Theory

To recap. Maximum likelihood picks the parameters that make the observed data most probable, and minimising negative log likelihood gives the standard losses. Priors combine with the likelihood into a posterior, and MAP estimation adds regularisation. Entropy is average surprise, and cross entropy is entropy plus KL divergence.

Key takeaways

  • The likelihood is the probability of the observed data as a function of the parameters.
  • Maximum likelihood estimation maximises Σ log p(xᵢ | θ); for a Gaussian mean it gives the sample mean.
  • Negative log-likelihood yields squared error (Gaussian noise) and cross-entropy (categorical outputs).
  • Posterior ∝ likelihood × prior; MAP with a Gaussian prior equals L2 regularisation.
  • Entropy is the average surprise −Σ p log p; cross-entropy = entropy + KL divergence.
  • KL divergence is asymmetric and appears in VAEs, RLHF, distillation and drift monitoring.

Check yourself

  1. Maximum likelihood estimation chooses parameters that…
    Show answer

    Make the observed data most probable — That is the definition of MLE.

  2. Minimising squared error corresponds to MLE under which noise assumption?
    Show answer

    Gaussian — The Gaussian log-density contains −(y − ŷ)².

  3. What is the entropy of a fair coin flip?
    Show answer

    1 bit — −2 × 0.5 log₂ 0.5 = 1.

  4. Cross-entropy H(P, Q) equals…
    Show answer

    H(P) + KL(P‖Q) — The key identity of this lecture.

  5. A Gaussian prior on the weights, in MAP estimation, is equivalent to…
    Show answer

    L2 regularisation (weight decay) — −log of a Gaussian prior is a squared-norm penalty.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/statistics-and-information-theory.html