AI in Motion

Mathematics for MLDeep diveBeginner12:01 video36 chapters

Vectors, Norms and the Dot Product — lecture notes

A deep dive into the object every model is built from: what vectors are, how to add, scale and measure them, and why the dot product powers neurons, attention and semantic search.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Vectors, Norms and the Dot Product

Welcome to this deep dive on vectors. Every input a model sees, every word embedding, every image and every set of learned weights is stored as a vector. Once you understand vectors and the dot product, most of the mathematics behind modern AI starts to make sense.

0:202. What is a vector?

What is a vector? — Vectors, Norms and the Dot Product

So what exactly is a vector? At its simplest it is an ordered list of numbers. You can picture it as a point in space, or as an arrow from the origin to that point. A customer described by age, height and income is a vector with three numbers.

0:403. Vectors everywhere

Vectors everywhere — Vectors, Norms and the Dot Product

Vectors appear everywhere in AI. A colour is three numbers. A small handwritten digit image is seven hundred and eighty four pixel values. A BERT word embedding has seven hundred and sixty eight numbers, and each token inside GPT-3 flows through the network as a vector of twelve thousand, two hundred and eighty eight numbers.

1:044. Basis vectors

Basis vectors — Vectors, Norms and the Dot Product

Coordinates are really instructions. The vector three, two means: take three steps along the first basis vector, called i hat, and then two steps along the second basis vector, j hat. Every vector is a combination of basis vectors, and its numbers tell you how much of each one to use.

1:255. Adding vectors

Adding vectors — Vectors, Norms and the Dot Product

Adding two vectors means adding them component by component. Geometrically, slide the tail of b to the tip of a, and the sum is the arrow from the start of a to the end of b. Here three, one plus one, two gives four, three. Order does not matter.

1:466. Scaling vectors

Scaling vectors — Vectors, Norms and the Dot Product

Multiplying a vector by a single number, called a scalar, stretches or shrinks it. A factor of two doubles its length. A negative factor flips it to point the other way. A factor of one half shrinks it. The arrow always stays on the same line through the origin.

2:077. Linear combinations

Linear combinations — Vectors, Norms and the Dot Product

Scaling and adding together give a linear combination: multiply each item by a weight and add them all up. This is the most common computation in machine learning. A neuron, a linear regression, an attention layer and a recommendation score all compute weighted sums of exactly this form.

2:278. Two views

Two views — Vectors, Norms and the Dot Product

It helps to hold two views at the same time. Geometrically a vector is an arrow with a direction and a length, which gives intuition about angles and distances. Algebraically it is a list of numbers, which works in any number of dimensions and is what computers actually store.

2:489. Length: the norm

Length: the norm — Vectors, Norms and the Dot Product

The length of a vector is called its norm. In two dimensions it is simply Pythagoras: the vector three, four has length five. In any number of dimensions, square every component, add the squares, and take the square root. This is the Euclidean norm, also called the L two norm.

3:0910. Other norms

Other norms — Vectors, Norms and the Dot Product

But length can be measured in several ways. The L one norm adds absolute values, and its unit ball is a diamond. The L two norm gives the familiar circle. The L infinity norm keeps only the largest component, which gives a square. For the green vector they give one point four, one, and point eight.

3:3311. Pause and think

Pause and think — Vectors, Norms and the Dot Product

Pause and think. Under which norm does the vector three, four have a length of exactly seven? Try the three norms we just met. The answer is the L one norm, because three plus four is seven. The L two norm gives five, and the L infinity norm gives four.

3:5412. Distance

Distance — Vectors, Norms and the Dot Product

Once we can measure length, we can measure distance: the distance between two points is simply the norm of their difference. That one idea powers k nearest neighbours, k means clustering and the squared error loss used to train regression models.

4:1213. Distances in action

Distances in action — Vectors, Norms and the Dot Product

Here is distance in action. To classify the new point, k nearest neighbours measures its distance to every training example, picks the closest ones and lets them vote. The choice of norm changes which points count as neighbours, and therefore changes the prediction.

4:3014. The dot product

The dot product — Vectors, Norms and the Dot Product

Now the star of the show: the dot product. Multiply matching components and add: three times one plus one times two equals five. That single number measures how much the two vectors point the same way. The green bar shows the projection of b onto a.

4:5015. Two formulas, one number

Two formulas, one number — Vectors, Norms and the Dot Product

The same number has a geometric meaning: the dot product equals the length of a, times the length of b, times the cosine of the angle between them. So the algebraic recipe and the geometric picture always agree, whether you have two dimensions or twelve thousand.

5:0916. The cosine curve

The cosine curve — Vectors, Norms and the Dot Product

The cosine does the interpretive work. When two vectors point the same way the angle is zero and the cosine is one. At ninety degrees the cosine is zero. When they point in opposite directions the angle is one hundred and eighty degrees and the cosine is minus one.

5:3017. Rotating b

Rotating b — Vectors, Norms and the Dot Product

Watch what happens as b rotates. While the angle is small, the dot product is large and positive. At exactly ninety degrees it becomes zero, and the vectors are called orthogonal. Beyond that it turns negative, because b now points partly against a.

5:4918. Pause and think

Pause and think — Vectors, Norms and the Dot Product

Time to pause and think. If the dot product of two non zero vectors is exactly zero, what does that tell you about them? Take a moment. The answer: they are orthogonal, perpendicular to each other, so neither one has any component along the other.

6:0819. Cosine similarity

Cosine similarity — Vectors, Norms and the Dot Product

Divide the dot product by both lengths and you get cosine similarity. It ignores length and keeps only direction, ranging from minus one for opposite, through zero for unrelated, up to one for exactly the same direction. Search engines and chatbots use it to compare meanings.

6:2820. Semantic search

Semantic search — Vectors, Norms and the Dot Product

Here is cosine similarity at work. Every document is turned into an embedding vector, and so is the question. The system retrieves the documents whose vectors point in the most similar direction. This is how semantic search, and retrieval augmented generation, find relevant text.

6:4621. Directions carry meaning

Directions carry meaning — Vectors, Norms and the Dot Product

Directions can even carry meaning. In a word embedding space, the arrow from man to woman is roughly parallel to the arrow from king to queen. Adding and subtracting vectors lets a model reason by analogy, at least approximately, which surprised researchers when word2vec first showed it.

7:0622. A neuron is a dot product

A neuron is a dot product — Vectors, Norms and the Dot Product

A single artificial neuron is a dot product in disguise. It multiplies each input by its weight, adds them up, adds a bias, and passes the result through an activation function. The weights define a direction, and the neuron responds most strongly to inputs that line up with it.

7:2723. The neuron formula

The neuron formula — Vectors, Norms and the Dot Product

Written as a formula, z equals w dot x plus b. The dot product checks how well the input matches the pattern stored in the weights, and the bias shifts the threshold. Stack many neurons side by side and you get a whole layer, which is a matrix times a vector.

7:4924. A linear classifier

A linear classifier — Vectors, Norms and the Dot Product

Geometrically, the set of points where w dot x plus b equals zero is a straight line, or in higher dimensions a flat hyperplane, and the weight vector is perpendicular to it. Points on one side give positive values, points on the other side negative ones. That is a linear classifier.

8:1025. Projection

Projection — Vectors, Norms and the Dot Product

Projection answers the question: how much of b lies along a? Take the dot product, divide by a dot a, and scale a by that amount. Projections power least squares regression, principal component analysis and the Gram Schmidt process for building orthonormal bases.

8:2926. Orthonormal bases

Orthonormal bases — Vectors, Norms and the Dot Product

A set of vectors that are mutually perpendicular and each have length one is called orthonormal. With such a basis, finding a coordinate is easy: just take the dot product with the basis vector. The principal components found by PCA form exactly this kind of basis.

8:4927. PCA uses projections

PCA uses projections — Vectors, Norms and the Dot Product

PCA is projection in action. It finds the direction along which the data varies the most, then projects every point onto that direction. Keeping only the top few directions compresses many dimensions into a few while losing as little information as possible.

9:0728. Attention is dot products

Attention is dot products — Vectors, Norms and the Dot Product

Transformers are built from dot products too. Here attention scores each word for the query it. Every score is the query dotted with a key, divided by the square root of the dimension. A softmax turns the scores into weights, and animal receives the most attention.

9:2629. Dot products in a transformer

Dot products in a transformer — Vectors, Norms and the Dot Product

In a language model, dot products are everywhere. Attention compares queries with keys by dot products and mixes values with weighted sums. The MLP layers are matrix vector products. Even the final score for every word in the vocabulary is a dot product.

9:4530. In code

In code — Vectors, Norms and the Dot Product

In code this takes a few lines. NumPy arrays hold the vectors, the at operator computes the dot product, and the norm function gives lengths. For our example the dot product is five and the cosine similarity is about point seven one, which means an angle of forty five degrees.

10:0631. Vectorise

Vectorise — Vectors, Norms and the Dot Product

A practical lesson: never loop over vector components in Python. A vectorised call hands the whole array to optimised code written in C or running on a GPU, which is often tens to hundreds of times faster. Deep learning frameworks are built on this principle.

10:2532. High dimensions

High dimensions — Vectors, Norms and the Dot Product

One warning about high dimensions. For random points in two dimensions, the nearest neighbour is far closer than the farthest one. In a thousand dimensions the ratio of nearest to farthest distance rises to about point eight nine: almost every point looks equally far away.

10:4433. Pause and think

Pause and think — Vectors, Norms and the Dot Product

Another question. Two embeddings have a cosine similarity of point nine five, but one vector is ten times longer than the other. Are they similar in meaning? Yes. Cosine similarity ignores length, and their directions almost match. Length often reflects frequency or confidence rather than meaning.

11:0434. Practical tips

Practical tips — Vectors, Norms and the Dot Product

Some practical tips. Normalise vectors before comparing directions. Scale your features so that no single feature, like income measured in dollars, dominates the distances. Prefer cosine similarity when comparing embeddings. And always use vectorised operations instead of Python loops.

11:2135. Cheat sheet

Cheat sheet — Vectors, Norms and the Dot Product

Here is a cheat sheet. Addition and scaling give new vectors, used for combining signals and taking gradient steps. The dot product, the norm and cosine similarity each give a single number, and they drive neurons, attention, regularisation and semantic search.

11:3936. Recap

Recap — Vectors, Norms and the Dot Product

To recap. A vector is an ordered list of numbers, an arrow in space. Norms measure length, and distance is the norm of a difference. The dot product equals the lengths times the cosine of the angle. Cosine similarity compares directions, and neurons, attention and search are all built on dot products.

Key takeaways

  • A vector is an ordered list of numbers; geometrically, an arrow from the origin.
  • Vectors add component-wise and scale by multiplying every component.
  • The L2 norm is √(Σvᵢ²); L1 and L∞ are other useful norms with different unit balls.
  • a · b = Σ aᵢbᵢ = ‖a‖‖b‖ cos θ: positive when aligned, zero when orthogonal, negative when opposed.
  • Cosine similarity divides out the lengths and powers semantic search and RAG.
  • A neuron computes w · x + b; attention scores are dot products of queries and keys.

Check yourself

  1. What is [3, 1] · [1, 2]?
    Show answer

    5 — 3×1 + 1×2 = 5.

  2. What is the L2 norm of [3, 4]?
    Show answer

    5 — √(9 + 16) = 5.

  3. Two non-zero vectors have a dot product of 0. They are…
    Show answer

    Orthogonal (perpendicular) — cos θ = 0 means θ = 90°.

  4. Why is cosine similarity popular for comparing embeddings?
    Show answer

    It ignores vector length and compares direction only — Direction carries meaning; length often reflects frequency.

  5. Which computation does a single neuron perform before its activation?
    Show answer

    w · x + b — A weighted sum of inputs plus a bias.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/vectors-and-dot-products.html