AI in Motion

How AI Image Generators Work

Artificial IntelligenceBeginner1:266 chapters

Diffusion models turn pure noise into a picture by removing a little noise at a time, guided by your text prompt.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 What does a diffusion model start from when generating an image?
Q2 What is the model trained to predict?
Q3 How does the text prompt affect the image?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. Type a description and an AI draws it in seconds. Most of these image generators are diffusion models. Their trick sounds backwards: they create images by removing noise.

Denoising. Generation starts with random static. At each step, the model predicts which part of the picture is noise and removes a little of it. Slowly, shapes appear, then details sharpen, until the image matches the prompt.

Training. During training, the model sees millions of captioned images. Noise is added to each one, and the network learns to predict exactly which noise was added. Once it is good at that, you can start from pure noise and run the process in reverse.

Guided by text. The text prompt is turned into embeddings, and they guide every denoising step. A different prompt, starting from the same kind of noise, ends in a different picture.

Responsible use. Use these tools responsibly. Check licences, never create images of real people without consent, watch for stereotypes in the results, and label AI images when it matters.

Recap. To recap. Diffusion starts from noise, removes a little predicted noise at every step, is guided by the prompt, and was trained by learning to predict the noise added to real images.