Long Context Versus Retrieval
Choose between sending whole documents to a long context model and retrieving passages, using evidence.
A taste of a lesson
Our whole handbook is 150 pages. Should we skip RAG and send it with every question?
Possibly, and it is worth testing. A handbook that size may fit in a large context window, which removes retrieval misses and index maintenance. The costs are price per query and latency, and prompt caching can reduce both if many questions arrive within the cache lifetime. Run a test: thirty real questions, answered once with the full handbook and once with retrieval, graded the same way, with cost and p95 latency recorded. How many questions per day do you expect?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Compare long context and retrieval on quality, cost, latency and freshness
- Explain how prompt caching changes the cost of repeated long prompts
- Recognise detail loss in very long inputs and how to test for it
- Design hybrid approaches that combine both
- Run a head to head comparison and decide from measurements
Lesson plan
- 1 Two ways to give a model documents Describe long context and retrieval approaches and their basic trade offs. Start
- 2 Cost, latency and caching Estimate per query cost and latency for each approach. Start
- 3 Quality in very long inputs Test whether a model actually uses details buried in long prompts. Start
- 4 Freshness, scale and permissions Factor in operational needs beyond answer quality. Start
- 5 Hybrid designs Combine retrieval and long context for the best of both. Start
- 6 Deciding with a head to head test Run both approaches on your questions and choose on evidence. Start
Try asking
About this tutor
For developers and technical leads deciding how to give a model access to documents now that context windows are large. You compare the two approaches on what matters: answer quality, cost per query, latency, freshness, permissions, corpus size and cross document reasoning. You learn how prompt caching changes the economics of repeated long prompts, why models can still miss details in very long inputs, and how hybrid designs (retrieve whole documents, then read them in full) often work best. You finish by running both approaches on the same question set and deciding from measured quality, cost and latency.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
Takes retrieval systems from demo to dependable: parsing, citations, freshness, retrieval evaluation and debugging
9 tutors 267 lessons taught Sample
Most RAG demos work on the ten documents someone picked. I teach what happens after that: scanned PDFs, tables, documents that change every week, answers that cite the wrong page and users who ask things the documents never covered. My background is in document processing and internal knowledge tools, so I am practical about formats and sceptical of any setup...
See Fumiko's profile and tutorsMore like this
Other tutors on the same or nearby topics.