Speech Recognition: From Sound to Text — lecture notes
How a voice becomes words: waveforms, spectrograms and neural models that turn frequency patterns into text.
0:001. Introduction

Voice assistants, live captions and dictation all start by turning sound into text. That is automatic speech recognition.
0:082. Waveform to spectrogram

A microphone records a waveform: air pressure changing thousands of times per second. Hard to read directly. So we compute a spectrogram, which shows how much energy each frequency has over time. Speech sounds appear as distinctive patterns. A neural network reads these patterns and outputs words.
0:283. The pipeline

Modern systems convert audio into a mel spectrogram, an encoder builds acoustic representations, and a decoder turns them into text tokens. Models such as Whisper are trained on huge amounts of transcribed audio.
0:434. Challenges

Speech is hard because of accents, background noise, overlapping speakers and words that sound alike. Many languages also have little transcribed training data, which makes them harder to support.
0:565. Recap

To recap. Sound becomes a spectrogram, a neural model turns frequency patterns into text, and accents, noise and data scarcity are the big challenges.
Key takeaways
- Speech recognition converts audio into text.
- Spectrograms show the energy of each frequency over time.
- Modern systems use neural encoders and decoders on mel spectrograms.
- Accents, noise and low-resource languages are major challenges.
Check yourself
- What does a spectrogram show?
Show answer
Energy at each frequency over time — It is a time–frequency picture of sound.
- Which is a common input representation for speech models?
Show answer
Mel spectrogram — Mel spectrograms mimic human hearing.
- Which is a major challenge for speech recognition?
Show answer
Background noise and accents — Real audio is noisy and varied.
Go deeper
- Automatic Speech Recognition: From HMMs to Whisper · The AI Lecture Hall
- Text-to-Speech: From Concatenation to Neural Voices · The AI Lecture Hall
- Multilingual and Low-Resource NLP (with a Focus on Bangla) · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/speech-recognition.html