Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Rate Limits, Retries and Backoff

Handle 429s and overloads gracefully with backoff, jitter, client side pacing and retries that never cause storms.

By Farid Haddad Intermediate Building with LLM APIs 4.7(3) 53 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $7 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Rate Limits, Retries and Backoff AI tutor following Farid Haddad's plan
Student:

During our launch, 429s got worse after we added retries. How is that possible?

Tutor:

Your retries probably fired in sync and without enough waiting, so every failed request came back almost immediately, adding load exactly when you were over the limit. Check three things: that waits grow exponentially, that you add random jitter so clients spread out, and that you are not stacking your retries on top of the SDK's own, which multiplies attempts. Then add pacing so you send below the limit in the first place. How many attempts does your SDK make by default?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Describe request, token and concurrency limits and read limit headers
  • Retry only transient errors with capped exponential backoff and jitter
  • Prevent retry storms and multiplied retries across layers
  • Pace requests proactively with a token bucket or worker queue
  • Keep retries from repeating side effects such as emails or writes

Lesson plan

6 lessons. Pick one to start there.

  1. 1 How limits work Understand the kinds of limits providers apply and how they show up. Start
  2. 2 Retry the right failures Separate transient errors from errors that retrying cannot fix. Start
  3. 3 Backoff and jitter Implement capped exponential backoff with jitter and attempt limits. Start
  4. 4 Avoiding retry storms Stop retries from making a struggling service worse. Start
  5. 5 Pacing instead of bouncing Stay under limits proactively with pacing and queues. Start
  6. 6 Retries and side effects Make sure retried work never duplicates actions. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For developers whose LLM features fail under load, launch day traffic or large jobs. You learn how providers limit requests and tokens per time window, how to read limit information from responses, and how to retry correctly: which errors to retry, exponential backoff with jitter, respecting Retry-After, capping attempts and total time, and keeping retries from multiplying across layers. Then you go further with client side pacing, queues and concurrency limits so you stay under limits instead of bouncing off them, and with care for side effects so retries never send an email twice.

Reviews

4.7

3 ratingsSample

  • Oskar B.Sample

    We were retrying inside an SDK that already retried three times. Nine attempts per failure. Fixed that and added jitter, and our next big job ran clean.

  • Precious N.Sample

    The token bucket lesson was the most useful part. Our side effect issue with duplicate emails was covered too. A bit fast in the middle.

  • Akira H.Sample

    Clear timelines for what happens when many clients retry. I finally understand why jitter matters.

About the teacher

Farid Haddad

Keeps LLM features fast, affordable and available as traffic grows: cost, caching, retries and observability

9 tutors 4.5(19) 347 lessons taught Sample

Most of my working life has been on platform and reliability teams, and these days I spend it on LLM features: the bills that surprise people, the 429 errors on launch day, the logs nobody can read. I teach the operational side of building with model APIs. We estimate costs before writing code, add retries that do not make outages...

See Farid's profile and tutors