Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Data Cleaning Fundamentals

Turn a messy table into data you can trust, with every change written down

By Lin Zhao Beginner Data science and statistics 4.7(3) 49 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $4 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Data Cleaning Fundamentals AI tutor following Lin Zhao's plan
Student:

About 15 percent of my income column is blank. Can I just delete those rows?

Tutor:

Maybe, but first ask why those values are missing. If income is blank at random, deleting loses data but does not distort much. But in surveys, people with very high or very low incomes often skip the question, so deleting them would shift your averages. Check whether the missing rows differ from the others in age, region or job type. If they do, keep the rows, mark income as missing, and report it. What do you see when you compare the blank rows with the rest?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Check types, ranges, categories and IDs systematically
  • Handle missing values with a reasoned choice instead of a default
  • Identify exact and near duplicates and define unique records
  • Standardise categories, units and date formats
  • Keep a cleaning log and write validation rules for future data

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Rules before cleaning Set up a safe, reproducible cleaning process. Start
  2. 2 Types, ranges and IDs Run basic checks that catch most problems. Start
  3. 3 Missing values Find how missing values are encoded and handle them deliberately. Start
  4. 4 Duplicates Identify exact and near duplicates and decide what is unique. Start
  5. 5 Categories, units and dates Standardise inconsistent values. Start
  6. 6 Outliers and validation rules Investigate extremes and automate checks for next time. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

A beginner tutor for anyone who receives data from other people and has to make it usable: analysts, researchers, operations staff and new data scientists. You will check types and ranges, handle missing values sensibly, remove true duplicates, standardise inconsistent categories and units, and decide what to do with outliers. Every lesson uses a small, deliberately messy table, and you build a cleaning log and a set of validation rules as you go. The methods apply in spreadsheets, SQL or Python, and the tutor keeps the focus on judgement rather than any one tool.

Reviews

4.7

3 ratingsSample

  • Kenji W.Sample

    Good tool neutral approach. The missing values lesson made me rethink filling blanks with zero.

  • Lerato N.Sample

    The messy table exercises were realistic. Found three spellings of the same country in our own data the next day.

  • Agnes P.Sample

    The cleaning log habit alone was worth it. My manager can now see exactly what I changed in our supplier data.

About the teacher

Lin Zhao

Data cleaning, SQL, exploratory analysis and honest charts

9 tutors 4.5(22) 427 lessons taught Sample

I teach the part of data science that takes most of the time: getting data into a shape you can trust, querying it, exploring it and showing it honestly. I came to data from operations work, where reports drove real decisions and a wrong join could cost a week. I teach by handing you small, deliberately messy tables and asking...

See Lin's profile and tutors