Support Vector Machines and the Kernel Trick
Of all the lines that separate two classes, pick the widest street. Then lift the data into a new dimension to separate what a line cannot.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. Support vector machines were among the most successful classifiers before deep learning. Their idea is elegant: find the widest possible street between two classes.
Maximum margin. Many different lines separate the pink and blue points. Which is best? An SVM chooses the line with the widest margin, the empty street between the classes. The points touching the edges of the street are the support vectors. Only they decide where the boundary goes.
Why margins matter. A wide margin leaves more room for new points that are a little different. Only the support vectors matter. When classes overlap, a soft margin allows a few mistakes, and the parameter C balances width against errors.
The kernel trick. What if no straight line works? Here the blue points sit on both sides of the pink ones. No single threshold separates them. But add a new feature, x squared, and the blue points lift up. Now a straight line separates them easily. That is the idea behind the kernel trick.
Kernels. Common kernels include linear, polynomial and the radial basis function. The clever part is that kernels compute similarities in the higher-dimensional space without ever building those features explicitly.
Recap. To recap. SVMs pick the widest street. Support vectors define it. Soft margins tolerate overlap. And kernels unlock curved boundaries.