AI in Motion

Speech Recognition: From Sound to Text

Natural Language ProcessingIntermediate1:075 chapters

How a voice becomes words: waveforms, spectrograms and neural models that turn frequency patterns into text.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 What does a spectrogram show?
Q2 Which is a common input representation for speech models?
Q3 Which is a major challenge for speech recognition?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. Voice assistants, live captions and dictation all start by turning sound into text. That is automatic speech recognition.

Waveform to spectrogram. A microphone records a waveform: air pressure changing thousands of times per second. Hard to read directly. So we compute a spectrogram, which shows how much energy each frequency has over time. Speech sounds appear as distinctive patterns. A neural network reads these patterns and outputs words.

The pipeline. Modern systems convert audio into a mel spectrogram, an encoder builds acoustic representations, and a decoder turns them into text tokens. Models such as Whisper are trained on huge amounts of transcribed audio.

Challenges. Speech is hard because of accents, background noise, overlapping speakers and words that sound alike. Many languages also have little transcribed training data, which makes them harder to support.

Recap. To recap. Sound becomes a spectrogram, a neural model turns frequency patterns into text, and accents, noise and data scarcity are the big challenges.