Learn RAG in 7 lessons
Intermediate Free
RAG in 7 lessons
How retrieval-augmented generation actually works end to end - chunking, embeddings, hybrid retrieval, grounded prompting, evaluation, and the failure modes that show up in production but never in a demo.
- 7 lessons
- 53 min total
- Quiz in every lesson
- Runnable Python
What you'll learn
- RAG = retrieve relevant text at request time, then put it in the prompt so the model answers from it instead of from memory.
- Chunk size is a trade-off: too small loses context, too big dilutes relevance and wastes prompt tokens on a partial match.
- An embedding is a fixed-length vector where geometric closeness approximates semantic closeness, learned by a model trained for that purpose.
- Pure vector search can miss exact matches - an error code, a product SKU, an acronym - because embeddings favor meaning over exact tokens.
- Put retrieved context in its own clearly-delimited section, separate from instructions and from the user's question.
- A golden set is a fixed list of real questions with known-correct answers (and ideally, known-correct source chunks) that you re-run every time retrieval or prompting changes.
- A RAG pipeline can fail silently - it returns a confident, well-formatted answer with no visible error, which is why these bugs survive so long in production.
Curriculum
- 1. What RAG is, and when not to use it Retrieval-augmented generation in one picture, why it exists, and the cases where it's the wrong tool. 7 min (completed)
- 2. Chunking Why documents need to be split before they're indexed, the trade-offs between chunk sizes, and a runnable toy chunker. 8 min (completed)
- 3. Embeddings and vector search What an embedding actually represents, how similarity search finds relevant chunks, and a runnable TF-IDF + cosine retriever that mirrors the same idea without a model. 8 min (completed)
- 4. Retrieval quality and hybrid search Why pure semantic search misses obvious matches, what hybrid search fixes, and how re-ranking tightens the final top-k. 7 min (completed)
- 5. Prompting with context and citations How to lay out retrieved chunks in a prompt so the model answers from them, refuses when nothing relevant was found, and cites its sources. 7 min (completed)
- 6. Evaluating RAG Golden sets, faithfulness vs relevance, and why 'it looks right in the demo' isn't evaluation. 8 min (completed)
- 7. Production failure modes The RAG failures that don't show up in a demo: silent misses, stale embeddings, citation-answer mismatch, and the retrieval trace that catches them. 8 min (completed)