Retrieval-Augmented Generation (RAG)

Combines information retrieval with an LLM so answers are grounded in retrieved evidence rather than model memory alone.

Quick help: Retrieval-Augmented Generation, or RAG, combines information retrieval with an LLM. Rather than expecting the model to contain all required knowledge internally, the system retrieves relevant external evidence and provides it to the model when answering.

Typical architecture

Documents -> Parse -> Chunk -> Embed -> Vector Store

User Question -> Retrieve Evidence -> LLM + Evidence -> Grounded Answer

Why RAG is useful

RAG can provide: access to private knowledge, fresher information, source attribution, reduced dependence on model training knowledge, and domain-specific answers.

RAG does not eliminate hallucination

Retrieval can fail. The system may retrieve the wrong evidence, fail to retrieve relevant evidence, retrieve incomplete context, or generate claims unsupported by retrieved evidence.

Therefore retrieval and generation should be evaluated separately.

KB Sandbox Principle

RAG quality is an empirical property of the complete system, not simply the choice of LLM.

Related Learning