Handling Large Data Files in Python
Work with datasets bigger than memory using chunks, smart dtypes, columnar formats and streaming.
A taste of a lesson
I process a big CSV in chunks and average a column. Is mean of each chunk, then averaged, correct?
Only if every chunk has exactly the same number of valid values, which is rarely true, especially in the last chunk or when some values are missing. The safe way is to keep two running totals: the sum of the column and the count of non missing values in each chunk. Add them up across chunks and divide once at the end. The same idea works for grouped averages with a sum and count per group. How are you handling missing values in that column right now?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Estimate the memory a dataset needs before loading it
- Cut memory with column selection and appropriate dtypes
- Process files in chunks and combine partial results correctly
- Stream JSON Lines with generators and write results incrementally
- Choose columnar formats, out of core engines or sampling for a task
Lesson plan
- 1 Memory as a budget Estimate how much memory a dataset needs and where it goes. Start
- 2 Read less, store smaller Load only the data you need with types that fit it. Start
- 3 Chunks and partial aggregates Process a file in pieces and combine results without errors. Start
- 4 Streaming JSON Lines Handle huge JSON Lines files with constant memory. Start
- 5 Columnar formats and out of core tools Store data in columnar files and query data larger than memory. Start
- 6 Sampling and smarter questions Know when a sample answers the question as well as the full dataset. Start
Try asking
About this tutor
For intermediate Python users whose CSV or JSON Lines files have grown to millions of rows and now crash a laptop or take forever to load. You learn to estimate memory needs, read only the columns you need, shrink data with the right types, process files in chunks, stream JSON Lines with generators, and move to a columnar format that loads faster and keeps types. The tutor introduces out of core tools in general terms, such as lazy dataframe engines and embedded analytical databases you can query with SQL, and explains when sampling is the smarter move.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
Numerical Python and code quality for data and AI projects that have outgrown a single notebook
9 tutors 374 lessons taught Sample
I work on the part of AI projects nobody photographs: the arrays, dataframes, tests and packaging that let a prototype survive contact with real data. I came to this through scientific computing and later backend work, so I care about two things at once, getting the numbers right and keeping the code readable for the next person. My lessons are...
See Felix's profile and tutorsMore like this
Other tutors on the same or nearby topics.