Embeddings
Numerical representations of content positioned so that items with similar meaning sit near one another in a vector space.
Quick help: Embeddings are numerical representations of content designed so that items with similar meaning are positioned near one another in a multidimensional vector space. They allow AI systems to search by semantic similarity rather than relying only on exact words.
Why they matter
Embeddings are a foundational component of many RAG systems. Documents are divided into chunks, each chunk is embedded, and the resulting vectors are stored in a vector database such as pgvector.
When a user asks a question, the question is embedded using the same embedding model. The system searches for stored vectors that are closest to the query vector.
Key engineering considerations
Model choice. Different embedding models produce different vector representations and dimensions. Embeddings from incompatible models should not normally be mixed in the same vector space.
Dimensions. Higher dimensionality does not automatically mean better retrieval. Quality, storage, indexing cost and latency should be evaluated.
Normalization and similarity. Common measures include cosine similarity, dot product and Euclidean distance. The indexing and query strategy must match the representation used.
Re-embedding. Changing embedding models usually requires regenerating stored vectors.
Evaluation
Embedding quality should ultimately be evaluated through retrieval performance rather than model reputation alone.
Useful measures include:
- Hit@K
- Recall@K
- MRR
- downstream answer quality
- latency
- storage and processing cost
KB Sandbox Principle
An embedding model is an engineering choice to evaluate, not an architectural constant.