Retrieval-Augmented Generation (RAG)
Last updated:
What RAG is
A language model only knows what was in its training data. RAG adds a lookup step: before the model answers, the system retrieves relevant text from a knowledge base and places it in the prompt. The model then writes its answer from that material instead of relying on memory alone.
The term comes from a 2020 research paper by Patrick Lewis and colleagues at Facebook AI Research. Today it describes most chatbots and assistants that answer from an organization’s own content.
How RAG works
A RAG system has a preparation phase and a per-question phase:
The quality of a RAG system depends heavily on retrieval. If the right passage is not found, even a strong model cannot answer correctly.
- Indexing, done ahead of time: documents are split into chunks, each chunk is turned into an embedding, and the chunks are stored in a searchable index.
- Retrieval: the question is matched against the index with semantic search, keyword search or both, and the best-matching chunks are selected.
- Augmentation: the selected chunks are added to the prompt together with the instructions and the conversation so far.
- Generation: the language model writes an answer based on the retrieved text and, if it is set up well, says when that text does not contain the answer.
Example
A software company’s chatbot is asked: “Can I export my invoices as CSV?” Retrieval finds a help article about exports and a recent changelog entry. The model reads both and explains that CSV export is available in the billing settings, repeating the steps from the article. When the company edits the article, the next answer reflects the change as soon as the content is re-indexed, without touching the model.
Why RAG matters for business chatbots
RAG is the most common way to make a general-purpose model useful for one specific business:
RAG does not remove every error. Outdated or contradictory documents lead to wrong answers, and the model can still misread a passage. Keeping sources current matters as much as the technology.
- Answers stay specific to your products, prices and policies.
- Updating knowledge means editing content, not retraining a model.
- Answers can be traced back to source passages, which makes mistakes easier to find and fix.
- Rules can tell the model to answer only from retrieved content and to admit when nothing relevant was found.
How intoCHAT uses RAG
intoCHAT is built on RAG. When you add website pages, files, text snippets or Q&A pairs, intoCHAT splits them into chunks and creates embeddings. For every visitor message it runs hybrid retrieval, combining semantic and keyword search, then passes the best excerpts to the model with rules that keep the answer grounded in them. Q&A answers and starred sources rank ahead of other content, and changes to a source take effect after you retrain it.