
Retrieval-Augmented Generation (RAG)
Combines information retrieval with an LLM so answers are grounded in retrieved evidence rather than model memory alone.
Quick help: Retrieval-Augmented Generation, or RAG, combines information retrieval with an LLM. Rather than expecting the model to contain all required knowledge internally, the system retrieves relevant external evidence and provides it to the model when answering.
Typical architecture
Documents -> Parse -> Chunk -> Embed -> Vector Store
User Question -> Retrieve Evidence -> LLM + Evidence -> Grounded Answer
Why RAG is useful
RAG can provide: access to private knowledge, fresher information, source attribution, reduced dependence on model training knowledge, and domain-specific answers.
RAG does not eliminate hallucination
Retrieval can fail. The system may retrieve the wrong evidence, fail to retrieve relevant evidence, retrieve incomplete context, or generate claims unsupported by retrieved evidence.
Therefore retrieval and generation should be evaluated separately.
KB Sandbox Principle
RAG quality is an empirical property of the complete system, not simply the choice of LLM.