Context Window
Last updated:
What a context window is
A language model has no memory in the human sense. For every reply, the application sends it a block of text, and the model reads that block before writing. The context window is the upper limit on how large that block, plus the reply, can be. It is measured in tokens, the small pieces of text that models process.
Context windows have grown quickly. Current models range from several thousand to hundreds of thousands of tokens, and some accept more. A larger window lets a model read longer documents or conversations in one go, but it usually costs more and takes longer per request.
How the context window fills up
For a single chatbot reply, the context window typically has to hold:
If the total exceeds the limit, something has to give: older messages are dropped or summarized, fewer passages are included, or the request fails. Studies have also found that models can pay less attention to information buried in the middle of a very long context, so a bigger window does not guarantee better use of it.
- The system prompt with the bot’s instructions and rules.
- The conversation so far, or a shortened version of it.
- Passages retrieved from the knowledge base.
- Results from tools or API calls, if any.
- Room for the answer the model is about to write.
Example
A visitor talks with a furniture shop’s chatbot about sofas for a while, then asks about delivery. The chatbot does not resend the whole catalog and every policy page. It retrieves the few passages about delivery, adds the recent conversation and its instructions, and stays well within the window. If it tried to include every page of the website with each message, it would quickly hit the limit and become slow and expensive.
Why it matters for business chatbots
The context window explains why a chatbot cannot simply read your whole website on every question. Good systems decide what goes into the window: the most relevant chunks of content, the instructions that matter and enough conversation history to follow the thread. That selection, more than the raw window size, shapes answer quality, speed and cost. It also explains why a very long conversation can lose track of details mentioned much earlier.
Context windows and intoCHAT
intoCHAT keeps your knowledge base outside the context window. For each message, it retrieves the most relevant excerpts of your content with hybrid search, combining semantic and keyword matching, and sends those excerpts, your instructions and the conversation to the model, not your whole knowledge base. You do not manage context size or tokens yourself: plans are measured in messages per month and characters of training content.