Temperature
Last updated:
What temperature is
When a language model writes, it calculates a probability for every possible next token and then picks one. Temperature adjusts those probabilities before the pick. At a low temperature, the most likely token wins almost every time. At a higher temperature, the probabilities are flattened, so less likely tokens are chosen more often.
Many APIs accept values from 0 up to 1 or 2, depending on the provider. A value of 0 comes close to always taking the top choice, although even then outputs are not always perfectly identical.
How it changes the output
In practice, temperature settings fall into three broad ranges:
Temperature does not add knowledge and does not make a model more truthful. A low setting makes the model repeat its most likely answer more reliably, which can mean repeating the same mistake. Some newer reasoning models do not expose temperature at all and manage sampling themselves. A related setting, top-p, limits the choice to the most probable tokens instead.
- Low: more consistent wording, well suited to factual answers, data extraction and classification.
- Medium: some variation in wording while staying on topic, a common default for chat.
- High: more varied and surprising output, useful for brainstorming or creative writing, but more likely to drift off topic or make errors.
Example
Ask a model five times to write a tagline for a bakery. At a low temperature you get nearly the same line each time. At a high temperature you get five quite different ideas, some useful and some odd. Ask the same model about the bakery’s opening hours, and you want the low-temperature behavior: one consistent answer based on the opening-hours page.
Why it matters for business chatbots
Customer-facing chatbots usually work best with low to moderate temperatures, because consistency matters more than creativity when the questions are about prices, policies or products. Still, temperature is a minor adjustment, not a fix. Missing or outdated content, vague instructions and weak retrieval are more common causes of bad answers than a poorly chosen temperature.
Temperature and intoCHAT
intoCHAT does not let customers set temperature or choose a model. Every agent runs on one OpenAI model chosen and managed by intoCHAT, with its settings handled for you. To change how your agent answers, you adjust what it knows and how it is told to behave: your sources, your instructions, Q&A pairs for answers that must be exact, and the welcome message and suggested questions.