Retrieval-Augmented Generation with Claude
Ground Claude in your own data with retrieval-augmented generation. Build hybrid vector plus full-text pipelines, chunk documents for optimal retrieval, add reranking for precision, and evaluate whether answers stay faithful to sources. Lightweight in-process vector search keeps the focus on RAG mechanics, not infrastructure.
About This Course
Learn how to ground Claude in your own data using retrieval-augmented generation (RAG). Build vector-search + full-text hybrid pipelines, chunk documents for optimal retrieval, add reranking for precision, and evaluate whether Claude answers are faithful to sources. Uses the Anthropic Python SDK plus lightweight in-process vector search (no external DB required) so students can focus on RAG mechanics rather than infrastructure.
Course Curriculum
10 Lessons
Introduction to Retrieval-Augmented Generation with Claude
By the end of this lesson you will know what Retrieval-Augmented Generation (RAG) actually is (retrieve then generate, not one thing), why it beats naive prompt-stuffing at any scale beyond a demo, how the citation pattern turns a Claude answer into an auditable artifact, and where RAG fits alongside the agent patterns from CLD-AI-103. Sets the stage for the L2 hands-on where you build a working retrieve-and-generate pipeline against a small Orion Analytics knowledge base.
Build a basic RAG pipeline - Lab Exercises
By the end of this hands-on lab you will have implemented a working retrieval-augmented generation pipeline over an Orion Analytics knowledge base of 12 chunks (SLA runbook, refund policy, API docs, compliance memos), composed a cited-answer Claude prompt, and verified 8 sample questions retrieve the correct chunks + get correctly cited answers. Uses stdlib TF-IDF (no external embedding dependency) so the exercise focuses on RAG mechanics, not infrastructure.
Vector embeddings and semantic retrieval
Learn what embeddings are (fixed-length vectors representing text meaning), how nearest-neighbor search finds semantically similar chunks even when word overlap is zero, why embedding-based retrieval beats keyword TF-IDF on synonym-heavy queries, and how to choose an embedding model. Sets up the L4 hands-on where you build a real vector-search retriever.
Build a vector-search retriever - Lab Exercises
By the end of this hands-on lab you will have built a vector-search retriever using SentenceTransformers embeddings (with a deterministic hash-based fallback if the container lacks the model), compared its recall against L2's TF-IDF baseline on 8 synonym-heavy queries, and seen the ~20-40 point recall lift semantic embeddings provide over pure keyword matching.
Chunking strategies for RAG
Learn how to split long documents into retrieval-ready chunks. Compare fixed-size, recursive-character, and semantic-boundary chunking. Understand the chunk-size sweet spot (500-800 chars for most enterprise knowledge bases), overlap between chunks, and preserving structural context (headings, code fences). Sets up the L6 hands-on where you measure retrieval quality vs chunk size on the Orion corpus.
Chunking experiment - Lab Exercises
Take one long Orion runbook doc, chunk it 5 ways (no-split, fixed-200, fixed-500, fixed-1000, recursive-500 with overlap), build a TF-IDF index per config, run 6 test queries against each, and report precision@1 + recall@3 per config. Confirms L5's theoretical "sweet spot 500-800" claim with your own numbers.
Hybrid search and reranking for higher-quality RAG
Combine BM25 keyword search + dense vector search into a hybrid retriever that catches both exact-term queries (product IDs, error codes) and synonym-heavy natural queries. Layer a cross-encoder reranker to prune wide initial retrieval into sharp top-K. Sets up L8 hands-on where you measure precision@5 with hybrid + reranker vs baseline TF-IDF and dense-only.
Hybrid search + reranker - Lab Exercises
Build a hybrid BM25+vector retriever with Reciprocal Rank Fusion, layer a stub reranker, and compare P@5 across 4 configurations (BM25 alone / vector alone / hybrid / hybrid+rerank) on 10 queries covering natural-language synonyms, exact error codes, and exact term matches.
Evaluating RAG systems
Learn the 4 dimensions of RAG evaluation (retrieval quality, answer faithfulness, answer relevance, citation accuracy), how to measure each with rule-based + LLM-as-judge scoring, and why "faithfulness" is the single most important metric to guard against hallucination in production RAG systems. Sets up L10 capstone where you wire the full eval harness into a production RAG pipeline.
RAG production pipeline capstone - Lab Exercises
Capstone for CLD-AI-104. Combine everything: recursive chunker (L5-L6), TF-IDF retriever with hybrid pattern (L7-L8), cited Claude answers (L1-L2), and faithfulness LLM-judge (L9). Run against 10 test cases and produce a 3-metric scorecard (must_contain relevance, citation validity, avg faithfulness).