AI in Motion

Human Pose Estimation

Computer VisionIntermediate1:085 chapters

Find a person’s joints — shoulders, elbows, knees — and connect them into a skeleton that can be tracked over time.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 What does a keypoint heatmap show?
Q2 Top-down pose estimation first…
Q3 Why should pose data be handled carefully?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. Fitness apps count your squats, games follow your dance moves, and sports analysts study athletes’ movements. All of them use pose estimation.

Heatmaps to skeleton. The network first predicts a heatmap for every joint: a glowing blob showing where the left elbow, right knee and so on are likely to be. The peak of each heatmap becomes a keypoint. Connecting the keypoints gives a skeleton, which can be tracked frame by frame as the person moves.

Many people. With several people, top-down methods first detect each person, then estimate their pose: accurate, but slower as crowds grow. Bottom-up methods find all the joints first and then group them into people.

Applications. Pose estimation helps fitness and physiotherapy apps, sports analysis, animation and gesture interfaces. But body movement is personal data, so use it with consent and care.

Recap. To recap. Predict heatmaps, pick their peaks as keypoints, connect them into a skeleton, and choose top-down or bottom-up for multiple people.