LectionesPart XI
Retrieval
Giving a model access to things it was not trained on.
The acronym has drifted. In the paper that coined it, the retriever is trained jointly with the generator; in most systems now called RAG, an off-the-shelf embedding model is bolted to a prompt. Read the first two papers to see what was given up in that trade, and then decide for yourself whether you want it back.
The survey is the one survey in this series. Use it as an index rather than reading it through.
The last two papers are the interesting ones, because both start from an admission. GraphRAG says plainly that some questions are not retrieval questions at all. CRAG says that retrieval can return confident rubbish and nothing downstream will notice.
The reading
- RAG
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Lewis et al. · NeurIPS · 2020
- Claim
- Coupling a parametric model to a non-parametric index, trained end to end, beats either alone on knowledge-intensive tasks.
- Why
- The paper the acronym comes from, and stricter than current usage. Worth reading precisely to see the difference between what was proposed and what shipped.
- Read
- Sections 2 and 4.
- REALM
REALM: Retrieval-Augmented Language Model Pre-Training
Guu, Lee, Tung, Pasupat & Chang · ICML · 2020
- Claim
- A retriever can be pre-trained together with the language model, using masked-token prediction itself as the training signal.
- Why
- Published months before RAG and arguably the deeper idea. The asynchronous index refresh is the practical problem nobody warns you about until you hit it.
- Read
- Section 3.3.
- RAG survey
Retrieval-Augmented Generation for Large Language Models: A Survey
Gao et al. · arXiv · 2023
- Claim
- A map of the field: naive, advanced and modular retrieval-augmented generation, with retrieval, generation and augmentation separated as axes.
- Why
- The only survey on this list. Read the taxonomy, then follow the one branch your problem sits on and ignore the rest.
- Read
- Section 3, then use the reference list.
- GraphRAG
From Local to Global: A Graph RAG Approach to Query-Focused Summarization
Edge et al. · Microsoft Research · 2024
- Claim
- Building an entity graph over a corpus and summarising its communities answers global questions that chunk retrieval cannot answer at all.
- Why
- It names the failure honestly: what are the themes in this corpus is not a retrieval question. The fix is expensive, so read Section 2 and decide whether your problem is that shape before you build it.
- Read
- Section 2.
- CRAG
Corrective Retrieval Augmented Generation
Yan, Xu, Cai, Han & Yu · arXiv · 2024
- Claim
- A lightweight evaluator grades the retrieved documents and triggers a correction — decompose, discard, or search the web — when they are wrong.
- Why
- The first move toward retrieval that knows when it has failed. Everything before it assumes the index returned something useful, which it frequently did not.