Skip to content

Chunking

Last updated:

What chunking is

Search systems and language models work better with focused passages than with whole documents. A 40-page manual covers dozens of topics, so a single embedding of the whole manual would be a blurred average of all of them. Chunking cuts the manual into passages that each cover one topic, so each one can be found on its own.

Every chunk is stored with a reference to its source, so the system knows which page or file an answer came from.

How it works

There are several common strategies, and many systems combine them:

  • Fixed-size chunking: split every N characters or tokens. Simple, but it can cut sentences and ideas in half.
  • Structure-aware chunking: split along paragraphs, headings or sentences, so natural units stay together.
  • Overlap: repeat a little text from the end of one chunk at the start of the next, so context at the boundary is not lost.
  • Semantic chunking: start a new chunk where the topic changes, detected with embeddings.

Choosing a chunk size

Chunk size is a trade-off. Small chunks are precise but can lose context, such as a sentence that says “this applies to all orders” without the sentence that explains what “this” is. Large chunks keep context but dilute the embedding, so they match fewer specific questions, and they take up more of the model's context window when retrieved.

Example: a hotel's FAQ page covers check-in times, parking, pets and breakfast. Without chunking, a question about pets would pull in the whole page, and the answer could mix in unrelated details. With paragraph-based chunking, the pets section becomes its own chunk and gives the model a clean, focused source.

Why it matters for business chatbots

Chunking is easy to overlook, yet it affects every answer. Poor chunks cause two common failures: the right information is not retrieved because it is buried in an unrelated passage, or it is retrieved without the sentence that states its conditions. Clear source content helps. Short sections with descriptive headings chunk well, while long walls of text or tables flattened into plain text are harder to split cleanly.

How intoCHAT handles it

intoCHAT chunks your content automatically when you train a source. It splits text along paragraph and sentence boundaries and adds a small overlap between neighboring chunks, so ideas that cross a boundary stay intact. Each chunk then gets an embedding and is searched with hybrid search. You do not choose a chunk size. The most useful thing you can do is keep your source content well structured.

Frequently asked questions

What is a good chunk size for RAG?

There is no single best size. It depends on the content, the embedding model and the kind of questions people ask. Many systems use passages of a few hundred words or less and adjust after testing retrieval with real questions.

Why do chunks overlap?

Overlap repeats a little text from the end of one chunk at the start of the next. It keeps sentences and context that fall on a boundary from being separated, so they can still be retrieved together. Too much overlap creates duplicate results, so it is usually kept small.

Can I control chunking in intoCHAT?

intoCHAT handles chunking automatically, so there is no chunk setting to adjust. You can improve results by structuring source content with clear headings and short paragraphs. Q&A pairs are also a direct way to give the agent a precise answer to an important question.

See it answer from your own website

Paste your website address and chat with an agent built from your pages. It takes about a minute.

Create your agent free

Free plan, no credit card needed.