Skip to content
Documentation

Crawl your website to train your AI agent

Add your website to an intoCHAT agent: find pages from the sitemap and links, pick what to read, exclude paths, handle JavaScript sites and crawl limits.

The quickest way to give your agent knowledge is to let it read your website. You enter an address, intoCHAT lists the pages it finds, and you choose which ones to add.

Find and add pages

  1. Open the agent's Knowledge tab and choose Website under Add knowledge.
  2. Enter your site in Website URL, for example https://example.com.
  3. Optional: in Exclude URLs containing, list parts of addresses to skip, separated by commas, for example /blog, /careers, ?lang=.
  4. Click Find pages.
  5. Review the list. Use Filter pages…, Select all or None, and tick pages one by one.
  6. Click Add … selected pages and wait for the progress bar to finish.
  7. Click Train agent at the top of the tab.

Added pages are saved right away as Pending. The agent only uses them after training.

How pages are found

Find pages combines two methods:

  • Sitemap: intoCHAT looks for a sitemap at common locations such as /sitemap.xml and /wp-sitemap.xml, and in the Sitemap: line of robots.txt. Entries for login, sign-up and admin pages are skipped.
  • Link map: a crawler collects the links on the same domain.

Each method returns up to 500 addresses. Duplicates such as /about and /about/, images, feeds and other files that aren't pages are removed. Links to PDF files stay in the list and are read like pages.

On a multilingual site with language folders such as /de/ and /fr/, a Languages row appears above the list. Pages in the language of the address you entered start selected, plus the home page; click a language to select or deselect all of its pages. Pages already in your knowledge show already added and start unselected.

What is read from each page

intoCHAT reads the main content of each page. Chat widgets and cookie banners are removed, and menus or banners that repeat on every page are kept only once.

Some sites build their content with JavaScript after the page loads. When several pages come back as the same empty shell, intoCHAT reads them again with a few seconds to load, and does the same for the rest of the run, so such sites take longer. If the crawler is busy, the run pauses with a countdown and then continues by itself.

Pages that time out are tried once more at the end. A page that still fails is marked failed; hover over it to see why, then select it and add it again.

Single page mode

Turn on Single page mode to add one address without scanning the site. The button changes to Add page. Use it for a page that discovery missed. To update a page you already added, use Refresh from website on its source instead.

Crawl limits

PlanWebsite crawls per monthCharacters per agent
Free1100K
Starter9500K
Pro201M
EnterpriseCustomCustom
  • Scanning with Find pages doesn't use a crawl.
  • Adding pages uses one crawl, whether a selection or a single page. All pages of one run share it. Refresh from website on a source counts like adding a single page.
  • Crawls are counted per account, across all agents, and reset every month.
  • Crawled text counts toward the agent's character limit. If a run reaches it, the remaining pages are marked failed.

When none are left, intoCHAT shows "Monthly crawl limit reached". You can still add files, text and Q&A.

Sites that block crawlers

Some sites turn away automated visitors with bot protection or a firewall. You then see "No pages found", or pages fail with little or no content. Try Single page mode for the pages that matter most, save key pages as PDF and upload them as files, or paste their text as a text snippet.

Only public http and https addresses can be crawled. Pages behind a login, addresses that contain a user name and password, and private or internal network addresses are refused.

Keeping pages up to date

There is no automatic re-sync, and Re-index saved text only trains the text read before again. To pick up changes to a page, click Refresh from website (the cloud icon) on its source. intoCHAT reads that page again and, if its text changed, saves and retrains it right away; otherwise you see "No changes found". A page read within the last hour may come from a cached copy.

To update many pages at once, add them again with Find pages, then train. Existing sources are updated rather than duplicated. See Managing sources.

View as Markdown