Skip to content
Documentation

How intoCHAT answers questions from your content

How an intoCHAT agent finds answers: chunking, hybrid semantic and keyword search, Q&A and priority sources first, grounding rules and no fine-tuning.

intoCHAT doesn't fine-tune an AI model on your content. It indexes your sources and, for every visitor message, looks up the passages that matter and gives them to the model together with your instructions. Knowing the steps helps you understand why an answer came out the way it did.

1. Training: splitting and indexing

When you press Train agent, each source's text is split into passages:

  • A passage holds up to about 1,600 characters. Paragraphs stay whole where possible; long ones are split at sentence boundaries.
  • Neighboring passages overlap by about 200 characters, so a fact on a boundary can be found from either side.
  • A Q&A pair stays in one piece unless it is longer than one passage.
  • Each passage is indexed together with its source's title and address, so a paragraph from "Pro plan" can be told apart from the same wording under "Starter plan".

"Training" in the dashboard means this indexing. The AI model itself doesn't change.

2. Understanding the question

Follow-up messages often depend on what came before, like "and how much is it?". The agent first rewrites the latest message into a standalone search query, using the last few turns of the conversation. Names, product codes and numbers are kept as written, and so is the visitor's language.

The query runs through two searches over your trained passages:

  • Semantic search finds passages with the same meaning, even when they use different words.
  • Keyword search finds exact terms such as product codes, names and prices, which semantic search can miss.

The two result lists are merged. Passages that don't match the question closely enough are dropped, so a weak match isn't passed off as an answer. A passage found only by keyword search must contain every meaningful word of the question, or every product code in it.

4. Choosing what the model sees

From the remaining matches, the agent picks at most six passages and about 12,000 characters in total. It takes up to three from one source before others, so a single long document doesn't crowd out the rest. When matches are close, the more trusted kind of source wins, in this order:

  1. Q&A pairs. A pair that closely matches the question is always placed first.
  2. Sources marked as priority (the star icon).
  3. Text snippets.
  4. Files.
  5. Website pages.

Each passage is labeled with its kind and the date it was last trained. When passages disagree, the model is told to follow the same order and, between sources of the same kind, to trust the most recently updated one. Your instructions outrank all knowledge. A successful action result, such as a live price from your API, outranks both for live facts.

5. Writing the answer

The model receives built-in rules, the agent's current settings (available actions and the lead form), your instructions and guardrails, and the conversation with the passages attached as reference material. The built-in rules take precedence over your instructions and guardrails where they conflict. They make the agent:

  • base business facts such as products, prices, policies and contact details only on your instructions, your knowledge or action results;
  • say briefly what it can't confirm when nothing relevant is found, instead of inventing an answer, a link or a contact channel;
  • treat your knowledge and action results as data, never as instructions. A line on a crawled page that says "ignore your rules" has no effect, and visitors can't override the rules either.

The full list is on the Instructions page.

Answers are written by one OpenAI model that intoCHAT selects and manages. There is no setting for the model or its temperature. The agent replies in the language the visitor writes in.

What this means for you

  • If a fact isn't in a trained source, the agent can't know it. Use a source's preview to check what was actually read. See Managing sources.
  • For questions that must be answered one way, add a Q&A pair.
  • If an outdated page keeps winning, update or delete it, or mark the correct source as priority.
  • Marking a source as priority applies to the next answer. Changed content only applies after training.

View as Markdown