Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Speech to text pipelines

Turn recordings into accurate, timestamped transcripts and understand where speech recognition fails

By Nadia Haddad Beginner Audio and voice AI NLP 4.7(3) 54 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $5 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Speech to text pipelines AI tutor following Nadia Haddad's plan
Student:

My reference has 20 words. The transcript got 2 words wrong, missed 1 and added 3. What is the WER?

Tutor:

Add up the three error types: 2 substitutions, 1 deletion and 3 insertions makes 6 errors. Divide by the number of words in the reference, which is 20: 6 / 20 = 0.30, so a word error rate of 30 percent. Notice the insertions count fully, which is why WER can go above 100 percent when a model invents lots of text, for example during silence. Before scoring, normalise both texts the same way for case and punctuation. Quick check: if all 3 inserted words came during a silent pause, which pipeline stage would you look at first?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Describe the stages of a speech to text pipeline from audio to transcript
  • Compute word error rate and explain its limits
  • Identify causes of transcription errors and choose practical fixes
  • Handle long recordings, timestamps and speaker labels
  • Apply consent and privacy basics when recording and transcribing

Lesson plan

6 lessons. Pick one to start there.

  1. 1 How audio becomes model input Understand sample rates, channels and spectrograms in plain terms. Start
  2. 2 End to end speech models See in broad terms how modern models turn audio features into text. Start
  3. 3 Building the pipeline Combine speech detection, chunking, transcription and punctuation into a working flow. Start
  4. 4 Who spoke when Add speaker labels with diarisation and understand its limits. Start
  5. 5 Measuring accuracy Compute word error rate and read it with care. Start
  6. 6 Failures, fixes and consent Handle invented text, noisy audio and the privacy side of recording. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For beginners who want to transcribe interviews, meetings, lectures or podcasts, or build a feature that does. You will learn how audio is represented, how modern end to end speech models work in broad terms, and how a full pipeline fits together: detecting speech, splitting long recordings, transcribing, adding punctuation and timestamps, and labelling who spoke when. You will compute word error rate by hand, see why accents, noise and crosstalk raise it, and learn about odd failures such as text invented during silence. The course closes with consent and privacy, which matter whenever voices are recorded. No coding required, with notes for builders.

Reviews

4.7

3 ratingsSample

  • Abdullah H.Sample

    Useful and honest about accents. Our Arabic and English interviews really do get different accuracy, and now I measure it.

  • Fiona M.Sample

    I transcribe research interviews. Learning to cut at pauses and keep timestamps made checking transcripts so much faster.

  • Chiara L.Sample

    The invented text in silence issue explained a weird line in our podcast transcript. The consent lesson was a good reminder too.

About the teacher

Nadia Haddad

Practical NLP: from tokens and embeddings to classification, translation and speech

9 tutors 4.5(22) 470 lessons taught Sample

I teach natural language processing as a craft: turning messy text in many languages into something a model can use, and checking honestly whether the result works. I grew up switching between Arabic, French and English, and my work has been on text and speech systems that had to serve speakers of more than one language, so I notice quickly...

See Nadia's profile and tutors