Build a Chat Backend That Holds Up
Design and build a chat backend with auth, storage, streaming, quotas and cost tracking that survives real use.
A taste of a lesson
Sometimes users get two identical replies after a network hiccup. How do I prevent that?
The client is retrying a request your server already processed. Fix it with an idempotency key: the frontend generates a unique id per message and sends it with the post. Your server stores it with the message, and if the same key arrives again, it returns the existing message and reply instead of creating new ones. Add a unique constraint on conversation id plus key so even simultaneous retries cannot both insert. What should the server return if the first attempt is still generating?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Design endpoints and tables for conversations and messages
- Wrap model access with timeouts, retries, logging and cost tracking
- Prevent duplicate and interleaved messages with idempotency and locking
- Protect budget and limits with per user quotas and input checks
- Track latency and cost per user and feature from usage data
Lesson plan
- 1 Endpoints and data model Sketch the API surface and tables for a chat feature. Start
- 2 One wrapper for model access Centralise prompts, retries, timeouts and logging in a single module. Start
- 3 The message flow with streaming Implement posting a message end to end, including partial failures. Start
- 4 Duplicates and concurrency Handle client retries and simultaneous requests safely. Start
- 5 Quotas, abuse and context budgets Protect your budget and keep requests within limits. Start
- 6 Observability and data care Measure cost and latency and handle stored conversations responsibly. Start
Try asking
About this tutor
For developers who can call a model API and now need a proper backend for a chat feature used by real people. You design endpoints for conversations and messages, store history safely per user, wrap the model client in one place, relay streamed replies, trim context, and add the protections a demo lacks: authentication, per user quotas, idempotent message posting, handling two requests on one conversation, timeouts and clear errors. You also track cost and latency per user and feature. Examples stay framework neutral, so the design fits whatever web stack you use.
Reviews
4.5
2 ratingsSample
- Sana Q.Sample
Strong on quotas and abuse, which our demo had none of. Framework neutral examples took a little translating but the design was solid.
- Mikael J.Sample
Idempotency keys and per conversation locking fixed the duplicate replies we saw in beta. Designing tables before code saved a migration later.
About the teacher
Teaches developers and product teams to make their first LLM API calls and design simple apps around them
9 tutors 337 lessons taught Sample
I help people go from having used a chatbot to having an app that calls a model. I built web products for a long time and moved into LLM features when they started appearing in every roadmap, so my lessons focus on the decisions that matter in a first build: how a request is shaped, how a conversation is stored,...
See Gabriela's profile and tutorsMore like this
Other tutors on the same or nearby topics.