Skip to content

Retrieval-Augmented Generation (RAG)

Last updated:

What RAG is

A language model only knows what was in its training data. RAG adds a lookup step: before the model answers, the system retrieves relevant text from a knowledge base and places it in the prompt. The model then writes its answer from that material instead of relying on memory alone.

The term comes from a 2020 research paper by Patrick Lewis and colleagues at Facebook AI Research. Today it describes most chatbots and assistants that answer from an organization’s own content.

How RAG works

A RAG system has a preparation phase and a per-question phase:

The quality of a RAG system depends heavily on retrieval. If the right passage is not found, even a strong model cannot answer correctly.

  • Indexing, done ahead of time: documents are split into chunks, each chunk is turned into an embedding, and the chunks are stored in a searchable index.
  • Retrieval: the question is matched against the index with semantic search, keyword search or both, and the best-matching chunks are selected.
  • Augmentation: the selected chunks are added to the prompt together with the instructions and the conversation so far.
  • Generation: the language model writes an answer based on the retrieved text and, if it is set up well, says when that text does not contain the answer.

Example

A software company’s chatbot is asked: “Can I export my invoices as CSV?” Retrieval finds a help article about exports and a recent changelog entry. The model reads both and explains that CSV export is available in the billing settings, repeating the steps from the article. When the company edits the article, the next answer reflects the change as soon as the content is re-indexed, without touching the model.

Why RAG matters for business chatbots

RAG is the most common way to make a general-purpose model useful for one specific business:

RAG does not remove every error. Outdated or contradictory documents lead to wrong answers, and the model can still misread a passage. Keeping sources current matters as much as the technology.

  • Answers stay specific to your products, prices and policies.
  • Updating knowledge means editing content, not retraining a model.
  • Answers can be traced back to source passages, which makes mistakes easier to find and fix.
  • Rules can tell the model to answer only from retrieved content and to admit when nothing relevant was found.

How intoCHAT uses RAG

intoCHAT is built on RAG. When you add website pages, files, text snippets or Q&A pairs, intoCHAT splits them into chunks and creates embeddings. For every visitor message it runs hybrid retrieval, combining semantic and keyword search, then passes the best excerpts to the model with rules that keep the answer grounded in them. Q&A answers and starred sources rank ahead of other content, and changes to a source take effect after you retrain it.

Frequently asked questions

What does RAG stand for?

RAG stands for retrieval-augmented generation. Retrieval is the search step that finds relevant text, and generation is the language model writing the answer. Augmented means the model’s prompt is enriched with the retrieved text.

What is the difference between RAG and fine-tuning?

RAG gives the model relevant documents at the moment it answers, while fine-tuning changes the model itself by training it on extra examples. RAG usually suits factual, frequently changing business knowledge better, because updates only require editing content. Fine-tuning is more useful for teaching a consistent style or output format.

Does RAG stop hallucinations?

It reduces them but does not eliminate them. The model can still misread a passage or fill gaps when retrieval misses the relevant text. Clear grounding rules, well-maintained content and testing lower the risk further.

Do I need a vector database for RAG?

Most RAG systems store embeddings in a vector index, either in a dedicated vector database or in a regular database with vector support. Hosted chatbot platforms such as intoCHAT handle this storage for you. Many systems also add keyword search, because exact terms like product codes are easy to miss with vectors alone.

See it answer from your own website

Paste your website address and chat with an agent built from your pages. It takes about a minute.

Create your agent free

Free plan, no credit card needed.