AI in Motion

Support Vector Machines and the Kernel Trick

Machine LearningIntermediate1:306 chapters

Of all the lines that separate two classes, pick the widest street. Then lift the data into a new dimension to separate what a line cannot.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 Which points determine the SVM boundary?
Q2 In the kernel example, which new feature made the data separable?
Q3 What does a soft margin allow?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. Support vector machines were among the most successful classifiers before deep learning. Their idea is elegant: find the widest possible street between two classes.

Maximum margin. Many different lines separate the pink and blue points. Which is best? An SVM chooses the line with the widest margin, the empty street between the classes. The points touching the edges of the street are the support vectors. Only they decide where the boundary goes.

Why margins matter. A wide margin leaves more room for new points that are a little different. Only the support vectors matter. When classes overlap, a soft margin allows a few mistakes, and the parameter C balances width against errors.

The kernel trick. What if no straight line works? Here the blue points sit on both sides of the pink ones. No single threshold separates them. But add a new feature, x squared, and the blue points lift up. Now a straight line separates them easily. That is the idea behind the kernel trick.

Kernels. Common kernels include linear, polynomial and the radial basis function. The clever part is that kernels compute similarities in the higher-dimensional space without ever building those features explicitly.

Recap. To recap. SVMs pick the widest street. Support vectors define it. Soft margins tolerate overlap. And kernels unlock curved boundaries.