Speech Recognition: From Sound to Text
How a voice becomes words: waveforms, spectrograms and neural models that turn frequency patterns into text.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. Voice assistants, live captions and dictation all start by turning sound into text. That is automatic speech recognition.
Waveform to spectrogram. A microphone records a waveform: air pressure changing thousands of times per second. Hard to read directly. So we compute a spectrogram, which shows how much energy each frequency has over time. Speech sounds appear as distinctive patterns. A neural network reads these patterns and outputs words.
The pipeline. Modern systems convert audio into a mel spectrogram, an encoder builds acoustic representations, and a decoder turns them into text tokens. Models such as Whisper are trained on huge amounts of transcribed audio.
Challenges. Speech is hard because of accents, background noise, overlapping speakers and words that sound alike. Many languages also have little transcribed training data, which makes them harder to support.
Recap. To recap. Sound becomes a spectrogram, a neural model turns frequency patterns into text, and accents, noise and data scarcity are the big challenges.