Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Handling Large Data Files in Python

Work with datasets bigger than memory using chunks, smart dtypes, columnar formats and streaming.

Your first 3 replies from the tutor are free. Then you can continue the lesson for $7 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Handling Large Data Files in Python AI tutor following Felix Brandt's plan
Student:

I process a big CSV in chunks and average a column. Is mean of each chunk, then averaged, correct?

Tutor:

Only if every chunk has exactly the same number of valid values, which is rarely true, especially in the last chunk or when some values are missing. The safe way is to keep two running totals: the sum of the column and the count of non missing values in each chunk. Add them up across chunks and divide once at the end. The same idea works for grouped averages with a sum and count per group. How are you handling missing values in that column right now?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Estimate the memory a dataset needs before loading it
  • Cut memory with column selection and appropriate dtypes
  • Process files in chunks and combine partial results correctly
  • Stream JSON Lines with generators and write results incrementally
  • Choose columnar formats, out of core engines or sampling for a task

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Memory as a budget Estimate how much memory a dataset needs and where it goes. Start
  2. 2 Read less, store smaller Load only the data you need with types that fit it. Start
  3. 3 Chunks and partial aggregates Process a file in pieces and combine results without errors. Start
  4. 4 Streaming JSON Lines Handle huge JSON Lines files with constant memory. Start
  5. 5 Columnar formats and out of core tools Store data in columnar files and query data larger than memory. Start
  6. 6 Sampling and smarter questions Know when a sample answers the question as well as the full dataset. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For intermediate Python users whose CSV or JSON Lines files have grown to millions of rows and now crash a laptop or take forever to load. You learn to estimate memory needs, read only the columns you need, shrink data with the right types, process files in chunks, stream JSON Lines with generators, and move to a columnar format that loads faster and keeps types. The tutor introduces out of core tools in general terms, such as lazy dataframe engines and embedded analytical databases you can query with SQL, and explains when sampling is the smarter move.

Reviews

Students can review a tutor after a paid lesson. Nobody has yet.

About the teacher

Felix Brandt

Numerical Python and code quality for data and AI projects that have outgrown a single notebook

9 tutors 4.5(18) 374 lessons taught Sample

I work on the part of AI projects nobody photographs: the arrays, dataframes, tests and packaging that let a prototype survive contact with real data. I came to this through scientific computing and later backend work, so I care about two things at once, getting the numbers right and keeping the code readable for the next person. My lessons are...

See Felix's profile and tutors