Profiling and Speeding Up Python
Find where your Python data and AI code really spends time and memory, then fix the parts that matter.
A taste of a lesson
My script uses df.apply with a lookup function on 2 million rows and takes 20 minutes.
That pattern is a classic. apply calls your Python function once per row, so you pay interpreter overhead two million times, plus whatever the lookup costs. If the lookup maps a key to a value, turn it into a single operation: df['col'].map(lookup_dict), or build a small DataFrame of keys and values and merge. Both run the matching in compiled code. Profile first to confirm apply dominates, then compare outputs before and after. What does your lookup function actually do inside?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Measure runtime and memory correctly with timing and profiling tools
- Read profiler output to find the most expensive call paths
- Apply algorithmic, vectorisation, caching and batching fixes in order of payoff
- Choose concurrency or multiprocessing based on whether work waits or computes
- Verify each optimisation keeps results identical and actually helps
Lesson plan
- 1 Measure before you change Set a target and measure current runtime and memory reliably. Start
- 2 Reading a profile Profile a full program and identify the call paths that cost the most. Start
- 3 Algorithms and data structures Remove unnecessary work with better lookups and joins. Start
- 4 Vectorising NumPy and pandas code Replace loops, iterrows and apply with whole column operations. Start
- 5 Caching, batching and I/O Avoid repeated expensive work and reduce the number of slow operations. Start
- 6 Concurrency, processes and when to stop Use parallelism appropriately and verify the final result. Start
Try asking
About this tutor
For experienced Python users whose data preparation, feature code or evaluation scripts are too slow, or run out of memory. You learn to measure before changing anything: timing snippets correctly, profiling whole programs, reading profiler output, and tracking memory. Then you apply the fixes that usually matter, in order of payoff: better algorithms and data structures, vectorised NumPy and pandas, removing row by row apply calls, caching repeated work, batching I/O, concurrency for waiting and processes for computing. The tutor is honest that in LLM apps the network is often the bottleneck and that some code is fast enough already.
Reviews
4.7
3 ratingsSample
- Sebastian K.Sample
Replaced apply with map and a merge, from many minutes to seconds. The insistence on comparing outputs before and after caught one subtle change.
- Thandiwe M.Sample
Rigorous and clear. It told me my LLM pipeline was network bound and Python tuning would barely matter, which saved me a week.
- Ioana D.Sample
Profiling showed half my runtime was logging inside a loop, not the maths I was trying to optimise. Humbling and very useful.
About the teacher
Numerical Python and code quality for data and AI projects that have outgrown a single notebook
9 tutors 374 lessons taught Sample
I work on the part of AI projects nobody photographs: the arrays, dataframes, tests and packaging that let a prototype survive contact with real data. I came to this through scientific computing and later backend work, so I care about two things at once, getting the numbers right and keeping the code readable for the next person. My lessons are...
See Felix's profile and tutorsMore like this
Other tutors on the same or nearby topics.