RAG Systems Deep Dive
Deep dive into Retrieval-Augmented Generation: embeddings, vector databases, chunking strategies, and hybrid search. What is RAG? Retrieval-Augmented Generation combines a retrieval system (finding relevant documents) with a generator (LLM). The LLM receives both the query and retrieved context, enabling it to answer questions about specific data without training. RAG reduces hallucinations by grounding responses in provided context. Embeddings & Vector Search Embeddings convert text to numerical vectors. Similar texts have similar vectors. Use cosine similarity or dot product to find nearest neighbors. Common embedding models: text-embedding-3-small (OpenAI), embedding-multilingual (Google), BGE (open-source). Store embeddings in vector databases: Pinecone, Weaviate, Milvus, Chroma, pgvector (PostgreSQL). Chunking Strategies Smaller chunks (200-500 tokens): more precise retrieval, less context. Larger chunks (1000-2000 tokens): more context for the LLM, may include irrelevant info. Overlapping chunks prevent information loss at boundaries. Recursive character splitting: split by paragraphs, sentences, then words. Semantic chunking: split at topic boundaries using embedding similarity. Hybrid Search Combine vector search (semantic meaning) with keyword search (exact terms) and re-rank results. BM25 (keyword) + vector search + reciprocal rank fusion (RRF). This catches both conceptual matches and specific term matches. Re-rankers (cross-encoders) score retrieved chunks against the query for the best results. Set high top-k retrieval then re-rank to top-n.