Context Window Budgeting for Apps
Plan what fits in each request: system prompt, documents, history and room for the answer.
A taste of a lesson
If my model has a huge context window, why not just send every document every time?
You can, but it has three costs. Money: you pay for every input token on every request. Time: long inputs slow down the first token. Quality: models can overlook details buried in very long inputs, so answers may get worse, not better. Usually it is better to send the most relevant parts plus clear instructions. Test it yourself: ask ten real questions with all documents and with only the relevant sections. Which version answers better, and what does each cost?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain the context window and what counts against it
- Measure and budget each part of a request in tokens
- Choose between selecting, summarising, splitting and trimming content
- Explain why very long inputs can lower answer quality
- Check sizes before sending and handle overflow gracefully
Lesson plan
- 1 What the context window is Understand the limit and what counts against it. Start
- 2 Building a token budget Measure each part of your request and set limits for each. Start
- 3 When documents do not fit Pick a strategy for content larger than your budget. Start
- 4 Long conversations Keep chat history within budget while preserving what matters. Start
- 5 More context is not always better Recognise quality, cost and latency costs of very long inputs. Start
- 6 Guarding against overflow Check sizes before sending and fail gracefully. Start
Try asking
About this tutor
For beginners building apps that send long documents, many retrieved passages or long chat histories to a model and keep hitting limits, rising costs or answers that miss important details. You learn what the context window is, how to give each part of a request a token budget, how to measure those parts, and what to do when the content will not fit: select relevant sections, summarise, split the task or trim history. The tutor also explains why more context is not always better, since models can overlook details in very long inputs and every token costs time and money.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
Structured output, tool calling and safe input handling for LLM applications that must behave predictably
9 tutors 308 lessons taught Sample
I teach the parts of LLM apps where free text has to meet real software: JSON that must parse, tools the model calls, images and documents coming in, and users who send things you did not plan for. I spent years writing integrations between messy systems, which taught me to treat every input as untrusted and every output as something...
See Greta's profile and tutorsMore like this
Other tutors on the same or nearby topics.