# Upload files to train your AI agent

> Upload PDF, Word, PowerPoint, Excel, CSV, EPUB, text and Markdown files to an intoCHAT agent. Size and character limits, what is extracted, and tips.

Upload documents that your website doesn't cover, such as price lists, manuals, product sheets or terms. intoCHAT extracts the text, and each file becomes its own source.

## Upload files

1. Open the agent's **Knowledge** tab and choose **Files** under **Add knowledge**.
2. Drop files onto the upload area, or click it to choose files. You can pick several at once; they are read one after another.
3. When the files are added, click **Train agent** at the top of the tab.

Files larger than 4 MB are uploaded first, with a progress bar, and then read.

## Supported formats

| Type | Extensions |
|---|---|
| PDF | `.pdf` |
| Word | `.doc`, `.docx`, `.docm` |
| PowerPoint | `.ppt`, `.pptx`, `.pptm`, `.pps`, `.ppsx`, `.ppsm`, `.pot` |
| Excel | `.xls`, `.xlsx`, `.xlsm`, `.xlsb` |
| OpenDocument | `.odt`, `.ods`, `.odp` |
| Other documents | `.rtf`, `.epub`, `.csv` |
| Plain text | `.txt`, `.md`, `.markdown` |

## Limits

- **File size:** up to 50 MB per file.
- **Text per file:** up to 4 MB of extracted text. A larger file fails with a message asking you to split it.
- **Characters per agent:** all of an agent's sources together may hold up to 100K characters on Free, 500K on Starter and 1M on Pro. If a file would go over the limit, the upload stops there and the remaining files aren't added.

The character count is the length of the extracted text, not the file size. A large PDF full of photos may contain only a few thousand characters, while a small spreadsheet can contain many more.

## What is extracted

Documents are converted to Markdown. Headings, paragraphs, lists and tables are kept, which helps intoCHAT split them into meaningful passages. Plain text and Markdown files are read as they are.

Not extracted:

- **Scans and images.** There is no OCR. A scanned PDF without a text layer fails with a message saying so. Run OCR on it first, or paste the text as a [text snippet](/en/docs/text-and-qa).
- **Text inside pictures**, such as charts, diagrams or screenshots.
- **Password-protected files.** Remove the password and upload the file again.

intoCHAT keeps the extracted text, not the original file. Large uploads are deleted from temporary storage once they have been read.

## Updating a file

To update a file, upload the new version under the same file name. You don't need to delete the old source first:

- **The content changed:** the new text replaces the old text of that source, and you see "Updated" with the file name. Click **Train agent** to use it. Until then the agent answers from the previous version.
- **The file is identical:** it is skipped with "Skipped 1 unchanged file".

A file with a different name is added as a new source, so delete the outdated one yourself. See [Managing sources](/en/docs/managing-sources).

> Anything in the agent's knowledge can end up in an answer. Leave out internal notes, personal data and anything you wouldn't publish on your website.

## Tips

- **Check the extracted text.** Open the file's preview with the eye icon in the source list. If something is missing there, the agent can't know it. Tables and multi-column layouts are worth a look.
- **Use descriptive file names.** The file name becomes the source title and is indexed with every passage, so `shipping-rates-2026.pdf` helps more than `scan_0042.pdf`.
- **Split long documents by topic.** Separate files are easier to update, mark as priority or remove one at a time.
- **Keep file names stable.** Re-uploading `prices.pdf` updates the existing source, while `prices-v2.pdf` is added next to it. Two versions of the same price list lead to inconsistent answers, so remove outdated ones.
- **Prefer the source document over a PDF export.** A Word or Excel file often keeps tables cleaner than a PDF made from it.
