Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Build a Chat Backend That Holds Up

Design and build a chat backend with auth, storage, streaming, quotas and cost tracking that survives real use.

By Gabriela Sousa Intermediate Building with LLM APIs 4.5(2) 18 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $8 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Build a Chat Backend That Holds Up AI tutor following Gabriela Sousa's plan
Student:

Sometimes users get two identical replies after a network hiccup. How do I prevent that?

Tutor:

The client is retrying a request your server already processed. Fix it with an idempotency key: the frontend generates a unique id per message and sends it with the post. Your server stores it with the message, and if the same key arrives again, it returns the existing message and reply instead of creating new ones. Add a unique constraint on conversation id plus key so even simultaneous retries cannot both insert. What should the server return if the first attempt is still generating?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Design endpoints and tables for conversations and messages
  • Wrap model access with timeouts, retries, logging and cost tracking
  • Prevent duplicate and interleaved messages with idempotency and locking
  • Protect budget and limits with per user quotas and input checks
  • Track latency and cost per user and feature from usage data

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Endpoints and data model Sketch the API surface and tables for a chat feature. Start
  2. 2 One wrapper for model access Centralise prompts, retries, timeouts and logging in a single module. Start
  3. 3 The message flow with streaming Implement posting a message end to end, including partial failures. Start
  4. 4 Duplicates and concurrency Handle client retries and simultaneous requests safely. Start
  5. 5 Quotas, abuse and context budgets Protect your budget and keep requests within limits. Start
  6. 6 Observability and data care Measure cost and latency and handle stored conversations responsibly. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For developers who can call a model API and now need a proper backend for a chat feature used by real people. You design endpoints for conversations and messages, store history safely per user, wrap the model client in one place, relay streamed replies, trim context, and add the protections a demo lacks: authentication, per user quotas, idempotent message posting, handling two requests on one conversation, timeouts and clear errors. You also track cost and latency per user and feature. Examples stay framework neutral, so the design fits whatever web stack you use.

Reviews

4.5

2 ratingsSample

  • Sana Q.Sample

    Strong on quotas and abuse, which our demo had none of. Framework neutral examples took a little translating but the design was solid.

  • Mikael J.Sample

    Idempotency keys and per conversation locking fixed the duplicate replies we saw in beta. Designing tables before code saved a migration later.

About the teacher

Gabriela Sousa

Teaches developers and product teams to make their first LLM API calls and design simple apps around them

9 tutors 4.4(18) 337 lessons taught Sample

I help people go from having used a chatbot to having an app that calls a model. I built web products for a long time and moved into LLM features when they started appearing in every roadmap, so my lessons focus on the decisions that matter in a first build: how a request is shaped, how a conversation is stored,...

See Gabriela's profile and tutors