Logging and Tracing LLM Calls
Record every model call and pipeline step so you can explain cost, slowness and bad answers.
A taste of a lesson
Some RAG answers take 20 seconds but I cannot tell why. What should I add?
Add tracing with a span per step. Create one trace per request, then a span around query rewriting, retrieval, reranking, the model call and any post processing, each with start and end times. On the model span also record input and output tokens, time to first token and retries. Then look at a few slow traces side by side: you will usually see one step dominating, often a long output, a retry or a slow search. Which steps does your pipeline have today?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Define the fields every model call log should contain
- Trace multi step pipelines with spans for retrieval, model and tool calls
- Log prompts and outputs responsibly with redaction, sampling and retention
- Link user feedback to traces and build dashboards for cost and latency
- Turn notable production traces into evaluation cases
Lesson plan
- 1 Questions logs must answer List the production questions your logging should make easy. Start
- 2 The model call record Log a complete, structured record for every model call. Start
- 3 Traces and spans Trace a request across retrieval, model calls and tools. Start
- 4 Content, privacy and retention Decide what content to log and protect it properly. Start
- 5 Feedback and dashboards Connect user signals to traces and monitor the system. Start
- 6 From traces to evals Use production traces to improve your evaluation set. Start
Try asking
About this tutor
For developers running LLM features in production who cannot yet answer simple questions: why was this answer wrong, why did this request take twenty seconds, which feature is spending the budget. You learn what to log for every model call, how to structure logs, how to trace multi step pipelines with spans for retrieval, model calls and tool calls, and how to handle prompt and output content responsibly with redaction, sampling, retention and access control. You also connect logs to user feedback and turn interesting traces into evaluation cases. Concepts follow open tracing standards and fit whichever observability tool you use.
Reviews
4.0
3 ratingsSample
- Connor W.Sample
Solid concepts, but I wanted more hands on setup with an actual tracing tool. It stays deliberately tool neutral, which is fair.
- Lukas M.Sample
Spans per step showed our slow requests were retries on the reranker, not the model. Would never have guessed. The privacy lesson made us add redaction too.
- Fatima A.Sample
The list of fields for every call is now our team standard. Turning traces into eval cases was a smart closing lesson.
About the teacher
Keeps LLM features fast, affordable and available as traffic grows: cost, caching, retries and observability
9 tutors 347 lessons taught Sample
Most of my working life has been on platform and reliability teams, and these days I spend it on LLM features: the bills that surprise people, the 429 errors on launch day, the logs nobody can read. I teach the operational side of building with model APIs. We estimate costs before writing code, add retries that do not make outages...
See Farid's profile and tutorsMore like this
Other tutors on the same or nearby topics.