Skip to content
NakodaAI

AI EncyclopediaTechniques & methods

Retrieval-Augmented Generation (RAG)

Also known as RAG

A technique that pairs a language model with a search step over an external knowledge source, so the model's answer is generated from retrieved, citable text rather than from memory alone.

RAGretrievalgroundingsearch

In plain English

Instead of asking a model to answer purely from what it memorised during training, RAG first fetches the most relevant documents from a knowledge base - your company's own policy documents, say - and hands those to the model along with the question. The model then writes its answer grounded in that retrieved text, which it can also cite.

Technical explanation

Introduced by Lewis et al. (Facebook AI Research, 2020) in 'Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,' the original formulation combines parametric memory (a pre-trained sequence-to-sequence model, BART in the paper) with non-parametric memory (an external corpus indexed as dense vectors via Dense Passage Retrieval). At inference time the system retrieves the top-k relevant passages and conditions generation on both the query and that retrieved context. Modern RAG systems generalise this with any embedding model plus a vector database for the retrieval step.

Why it matters

RAG is the standard way to make an LLM answer accurately about information it was not trained on - a company's internal documents, a fast-changing regulation, yesterday's news - without the cost of retraining the model itself. It is the most common architecture behind enterprise AI assistants and document Q&A tools deployed in the UAE and globally.

Real-world example

The original Lewis et al. paper indexed Wikipedia as the retrieval corpus and showed RAG models generated more specific, diverse and factual language than generation from model parameters alone - the finding that established the technique.

Common misunderstanding

That RAG eliminates hallucination entirely. It substantially reduces it by grounding answers in retrieved text, but a model can still misread or misquote what was retrieved, or the retrieval step itself can surface the wrong document - RAG narrows the failure mode, it does not close it.

Something wrong here?

Every entry is hand-researched and hand-written by Nakoda. If a fact is stale, a source has changed or a definition needs sharpening, tell us and we will check it.