pandas DataFrames for Everyday Analysis
Load, filter, group, join and summarise tabular data with pandas without the classic silent mistakes.
A taste of a lesson
I merged tickets with customers and now I have more rows than tickets. What happened?
Your customers table has duplicate customer ids, so each ticket matched more than one customer row and pandas kept every combination. Check with customers['customer_id'].duplicated().sum(). If duplicates are real data problems, deduplicate first; if they are different records, decide which one to keep. Then merge again with validate='many_to_one', which raises an error if the customer side is not unique. Finally compare len(merged) to len(tickets). How many duplicated ids does that first check show?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Load data with correct types and inspect it before analysing
- Select and filter rows and columns with loc, iloc and conditions
- Group, aggregate and join tables while checking row counts
- Handle missing values and messy number columns deliberately
- Avoid chained assignment and duplicate key join mistakes
Lesson plan
- 1 Load and inspect before anything else Load a CSV with explicit options and understand its shape, types and gaps. Start
- 2 Selecting and filtering Pick the rows and columns you need with clear, explicit code. Start
- 3 Cleaning types and missing values Fix number and date columns and handle missing values on purpose. Start
- 4 Grouping and aggregating Summarise data by category with groupby and multiple aggregations. Start
- 5 Joining tables safely Merge two tables and prove the join did what you intended. Start
- 6 Dates, text and a full analysis Work with date and text columns and answer one real question end to end. Start
Try asking
About this tutor
For beginners who have data in CSV or spreadsheet form and want to answer real questions with Python. You learn to load data with the right types, inspect it, select rows and columns with loc and iloc, filter with conditions, handle missing values, group and aggregate, join tables and work with dates and text columns. The tutor pays special attention to the mistakes that give wrong answers without errors: joins that duplicate rows, missing values dropped by accident, numbers read as text. Exercises use a realistic dataset of support tickets or sales you can swap for your own.
Reviews
4.3
3 ratingsSample
- Bianca T.Sample
The habit of checking row counts after every merge caught a duplicate key problem in my real report on day two. validate is now in all my joins.
- Mei L.Sample
Practical examples with ticket data. The note about pandas behaviour changing between versions saved me from a confusing warning.
- Samuel A.Sample
Good on loc versus iloc and missing values. A few explanations assumed I knew more about indexes than I did, but I asked and it slowed down.
About the teacher
Numerical Python and code quality for data and AI projects that have outgrown a single notebook
9 tutors 374 lessons taught Sample
I work on the part of AI projects nobody photographs: the arrays, dataframes, tests and packaging that let a prototype survive contact with real data. I came to this through scientific computing and later backend work, so I care about two things at once, getting the numbers right and keeping the code readable for the next person. My lessons are...
See Felix's profile and tutorsMore like this
Other tutors on the same or nearby topics.