Skip to content

Web Crawler

Last updated:

What a web crawler is

A web crawler, also called a spider or bot, starts from one or more addresses and works through a site page by page. Search engines run the best-known crawlers, such as Googlebot and Bingbot, to discover and index pages. AI companies run crawlers such as OpenAI’s GPTBot to collect content for their models, and chatbot platforms use crawlers to read one specific company’s website.

Site owners give crawlers instructions through a robots.txt file and can list their pages in an XML sitemap. Well-behaved crawlers follow those signals and identify themselves with a user agent name.

How a web crawler works

Most crawlers go through four stages:

Crawlers usually limit how many pages they fetch and how quickly, so they do not overload the site they are reading.

  • Discover: collect page addresses from a sitemap, from links on pages already read, or both.
  • Fetch: download each page. Pages that build their content with JavaScript have to be rendered in a browser-like environment first, or the crawler sees an almost empty page.
  • Extract: separate the main content from navigation, footers, cookie banners and other repeated elements.
  • Store: save the text for its purpose, such as a search index or a chatbot’s knowledge base.

Example

A dental practice wants a chatbot that answers questions about treatments and prices. Instead of copying text by hand, the owner enters the website address. The crawler reads the sitemap and finds 40 pages, and the owner deselects the blog archive and job ads. The remaining pages about treatments, prices, opening hours and insurance become the chatbot’s knowledge.

Why it matters for business chatbots

A business website is usually the most complete, already approved description of what the company offers, so crawling it is the quickest way to give a chatbot useful knowledge. The quality of the crawl shapes the quality of the answers: if important pages are missed, if JavaScript content is not rendered or if menus and cookie notices end up in the text, retrieval gets worse.

Websites also change. When prices, policies or products are updated, the crawled content has to be refreshed, or the chatbot keeps answering from the old version.

How intoCHAT crawls a website

In intoCHAT you paste a URL, and the crawler discovers pages from the sitemap and a link map, up to 500 URLs. You pick which pages to include and can exclude paths with patterns. JavaScript-heavy pages are rendered before reading, and chat widgets and cookie banners are stripped from the text. The number of crawls per month depends on your plan, and when your site changes you retrain the affected sources manually; there is no automatic re-sync. To try it first, the free preview tool builds a temporary agent from 5 pages of any website.

Frequently asked questions

What is the difference between a web crawler and a web scraper?

A crawler focuses on discovering and visiting pages by following links or sitemaps. A scraper focuses on extracting specific data from pages, such as prices or contact details. Many tools do both, crawling to find pages and then extracting their content.

Can a chatbot crawler read JavaScript websites?

Only if it renders the pages first, the way a browser does. Sites that load their content through JavaScript can look almost empty to a crawler that only reads the raw HTML. intoCHAT renders JavaScript-heavy pages before reading them.

How do I keep crawlers away from parts of my site?

The standard method is a robots.txt file that tells crawlers which paths not to visit, although it relies on crawlers choosing to respect it. For a chatbot you set up yourself, it is simpler to leave pages out in the crawl settings. In intoCHAT you deselect pages or exclude paths with patterns before training.

Does intoCHAT re-crawl my website automatically?

No. intoCHAT has no automatic re-sync, so when your website changes you retrain the affected sources manually. The number of website crawls per month depends on your plan.

See it answer from your own website

Paste your website address and chat with an agent built from your pages. It takes about a minute.

Create your agent free

Free plan, no credit card needed.