Embeddings & Vector Databases: Text Embeddings and Similarity Search
Learn about text embeddings, cosine similarity, ANN indexing, and how vector databases power semantic search and RAG systems. Embeddings are the foundation of modern AI's ability to understand and compare text. By converting words, sentences, and documents into numerical vectors, embeddings enable semantic search, clustering, classification, and the retrieval components of RAG systems. What Are Embeddings? An embedding is a numerical representation of data — text, images, audio, or any other type — as a vector of floating-point numbers. The key property of embeddings is that similar items produce similar vectors. A text embedding model takes a string like "I love machine learning" and outputs a vector of typically 768 to 4096 numbers. The same model transforms "I enjoy deep learning" into a similar vector — close in the embedding space — while "I hate broccoli" produces a very different vector. This vector representation captures semantic meaning, not just keyword overlap. Two sentences about the same topic but using completely different words will still have similar embeddings, while two sentences containing the same words but with different meanings will have different embeddings. Text Embedding Models Several excellent embedding models are available. OpenAI text-embedding-3-small (1536 dimensions) and text-embedding-3-large (3072 dimensions) offer state-of-the-art performance with straightforward API access. They handle 8K token inputs and support dimensions parameter for flexibility. Google Gecko provides multilingual embeddings with strong performance across 100+ languages. BAAI/bge-large-en-v1.