Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Debugging neural network training

A systematic method for finding why a model will not train, diverges or quietly underperforms

By Nikolai Sorin Advanced Deep learning 4.5(4) 66 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $10 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Debugging neural network training AI tutor following Nikolai Sorin's plan
Student:

My model trains fine in float32 but the loss becomes NaN in mixed precision. Any ideas?

Tutor:

That points to range limits of 16 bit numbers. In float16, values above about 65,000 overflow to infinity and small gradients can underflow to zero. First, check whether you are using loss scaling; float16 training usually needs it so small gradients survive. Second, look for operations that produce large values: softmax or exponentials without subtracting the maximum, large attention scores, or a log of something that can reach zero. Third, if your hardware supports bfloat16, try it, since its range matches float32. Then find the first layer producing an infinity. Which 16 bit format are you using, and is loss scaling on?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Follow a repeatable process to isolate the cause of a training failure
  • Use single batch overfitting, starting loss and gradient norms as quick diagnostics
  • Trace NaN values to their source, including mixed precision overflow
  • Find silent data, label and mode bugs that lower scores without errors
  • Handle loss spikes in long runs with clipping, checkpoints and rate changes

Lesson plan

7 lessons. Pick one to start there.

  1. 1 A method, not guesses Reproduce, simplify and change one thing at a time with written hypotheses. Start
  2. 2 Cheap diagnostics first Use starting loss, single batch overfitting and shape checks to locate bugs quickly. Start
  3. 3 The data pipeline Find silent problems in loading, preprocessing and labelling. Start
  4. 4 Gradients and learning rates Read per layer gradient norms and run quick learning rate sweeps. Start
  5. 5 Hunting NaN and mixed precision issues Trace non finite values to the first operation that produced them. Start
  6. 6 Loss spikes and long runs Respond to instability late in training without losing the run. Start
  7. 7 Underperformance and baselines Diagnose a model that trains but scores worse than expected. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For practitioners who have lost days to a model that will not learn, suddenly produces NaN, or trains fine but scores worse than it should. You will learn a repeatable debugging process: reproduce, simplify, overfit one batch, inspect the data pipeline, check the starting loss, read gradient norms per layer and sweep the learning rate. Separate lessons cover NaN hunting in mixed precision, loss spikes in long runs, silent label and preprocessing bugs, train and evaluation mode problems, and how to compare against a known good baseline. You bring your own failing runs, and the tutor helps you form and test hypotheses one at a time.

Reviews

4.5

4 ratingsSample

  • Leila N.Sample

    The tutor refused to guess and kept asking for my curves. Annoying at first, then I realised that was the lesson.

  • Rachel T.Sample

    We had loss spikes in a long run. Logging which batch caused each spike led us to a set of corrupted examples. Very practical.

  • Kofi M.Sample

    Strong on method. Some lessons assume you already run large jobs, which I do not, but the cheap diagnostics lesson alone saved me a lot of time.

  • Bogdan I.Sample

    The debugging log habit changed how I work. I found a tokenizer mismatch in our fine tuning code in the data pipeline lesson.

About the teacher

Nikolai Sorin

Architectures explained from the inside: convolutions, recurrence, attention and beyond

9 tutors 4.6(22) 362 lessons taught Sample

I teach neural network architectures and the reasoning behind them. My working life has been spent implementing models from papers, getting them to train, and finding out which details the paper forgot to mention. I like to explain an architecture by asking what problem it was built to solve and what it costs, so convolutional networks, recurrent networks, transformers and...

See Nikolai's profile and tutors