AI in Motion

Natural Language ProcessingIntermediate1:07 video5 chapters

Speech Recognition: From Sound to Text — lecture notes

How a voice becomes words: waveforms, spectrograms and neural models that turn frequency patterns into text.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Speech Recognition: From Sound to Text

Voice assistants, live captions and dictation all start by turning sound into text. That is automatic speech recognition.

0:082. Waveform to spectrogram

Waveform to spectrogram — Speech Recognition: From Sound to Text

A microphone records a waveform: air pressure changing thousands of times per second. Hard to read directly. So we compute a spectrogram, which shows how much energy each frequency has over time. Speech sounds appear as distinctive patterns. A neural network reads these patterns and outputs words.

0:283. The pipeline

The pipeline — Speech Recognition: From Sound to Text

Modern systems convert audio into a mel spectrogram, an encoder builds acoustic representations, and a decoder turns them into text tokens. Models such as Whisper are trained on huge amounts of transcribed audio.

0:434. Challenges

Challenges — Speech Recognition: From Sound to Text

Speech is hard because of accents, background noise, overlapping speakers and words that sound alike. Many languages also have little transcribed training data, which makes them harder to support.

0:565. Recap

Recap — Speech Recognition: From Sound to Text

To recap. Sound becomes a spectrogram, a neural model turns frequency patterns into text, and accents, noise and data scarcity are the big challenges.

Key takeaways

  • Speech recognition converts audio into text.
  • Spectrograms show the energy of each frequency over time.
  • Modern systems use neural encoders and decoders on mel spectrograms.
  • Accents, noise and low-resource languages are major challenges.

Check yourself

  1. What does a spectrogram show?
    Show answer

    Energy at each frequency over time — It is a time–frequency picture of sound.

  2. Which is a common input representation for speech models?
    Show answer

    Mel spectrogram — Mel spectrograms mimic human hearing.

  3. Which is a major challenge for speech recognition?
    Show answer

    Background noise and accents — Real audio is noisy and varied.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/speech-recognition.html