How AI Image Generators Work — lecture notes
Diffusion models turn pure noise into a picture by removing a little noise at a time, guided by your text prompt.
0:001. Introduction

Type a description and an AI draws it in seconds. Most of these image generators are diffusion models. Their trick sounds backwards: they create images by removing noise.
0:122. Denoising

Generation starts with random static. At each step, the model predicts which part of the picture is noise and removes a little of it. Slowly, shapes appear, then details sharpen, until the image matches the prompt.
0:283. Training

During training, the model sees millions of captioned images. Noise is added to each one, and the network learns to predict exactly which noise was added. Once it is good at that, you can start from pure noise and run the process in reverse.
0:474. Guided by text

The text prompt is turned into embeddings, and they guide every denoising step. A different prompt, starting from the same kind of noise, ends in a different picture.
0:595. Responsible use

Use these tools responsibly. Check licences, never create images of real people without consent, watch for stereotypes in the results, and label AI images when it matters.
1:116. Recap

To recap. Diffusion starts from noise, removes a little predicted noise at every step, is guided by the prompt, and was trained by learning to predict the noise added to real images.
Key takeaways
- Diffusion models generate images by gradually removing noise.
- They are trained to predict the noise added to real images.
- The text prompt, turned into embeddings, guides every step.
- Use generated images responsibly: licences, consent and bias.
Check yourself
- What does a diffusion model start from when generating an image?
Show answer
Random noise — Generation begins with pure random noise.
- What is the model trained to predict?
Show answer
The noise that was added to an image — Learning to predict the added noise lets it remove noise step by step.
- How does the text prompt affect the image?
Show answer
It guides every denoising step — Prompt embeddings condition each step of denoising.
Go deeper
- Diffusion Models: Generating by Learning to Denoise · The AI Lecture Hall
- Latent Diffusion and Stable Diffusion: Text-to-Image at Scale · The AI Lecture Hall
- Text-to-Image and Text-to-Video Generation: Systems, Control and Provenance · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/how-image-generators-work.html