Tokens
Last updated:
What tokens are
Language models do not see letters or whole words. Before text reaches the model, a tokenizer splits it into tokens from a fixed vocabulary. Common words are often a single token, while rare words, names and long compound words are split into several pieces. Numbers, punctuation and spaces count too.
A common rule of thumb is that one token of English text is about four characters, or about three quarters of a word. This is only an approximation: each model family has its own tokenizer, and text in many other languages, including German, Italian, French and Spanish, often needs more tokens than the same content in English.
How tokens are counted
Each time a chatbot answers, tokens are counted in two directions:
Both count against the model’s context window. Providers that charge per use usually price input and output tokens separately, and output tokens often cost more. Embedding models also count tokens when they turn text into vectors.
- Input tokens: everything sent to the model, such as instructions, conversation history and retrieved documents.
- Output tokens: the answer the model generates, produced one token at a time.
Example
The sentence “Our store opens at 9 am on weekdays.” has eight words and 36 characters, so most tokenizers turn it into roughly ten tokens. A full chatbot request is much larger: instructions, a few retrieved passages and a short conversation can add up to a few thousand input tokens before the model writes a single word of its answer.
Why tokens matter for business chatbots
If you build a chatbot directly on a model API, tokens drive your costs and limits. Long instructions, large retrieved passages and long conversations all add input tokens to every message, which is a good reason to keep instructions focused and retrieval selective.
Tokens also explain practical limits: why a model cannot read an entire website in one request, why very long conversations eventually have to be shortened, and why an answer can be cut off when it reaches a maximum output length.
Tokens and intoCHAT
With intoCHAT you do not count tokens. Plans are measured in messages per month and characters of training content: Free includes 25 messages per month and 100K characters, Starter ($19 per month) 2,500 messages per month and 500K characters per agent, and Pro ($49 per month) 10,000 messages per month and 1M characters per agent. intoCHAT sends the model only the most relevant excerpts of your content for each message, and the model and its token usage are managed for you.