Skip to content

Temperature

Last updated:

What temperature is

When a language model writes, it calculates a probability for every possible next token and then picks one. Temperature adjusts those probabilities before the pick. At a low temperature, the most likely token wins almost every time. At a higher temperature, the probabilities are flattened, so less likely tokens are chosen more often.

Many APIs accept values from 0 up to 1 or 2, depending on the provider. A value of 0 comes close to always taking the top choice, although even then outputs are not always perfectly identical.

How it changes the output

In practice, temperature settings fall into three broad ranges:

Temperature does not add knowledge and does not make a model more truthful. A low setting makes the model repeat its most likely answer more reliably, which can mean repeating the same mistake. Some newer reasoning models do not expose temperature at all and manage sampling themselves. A related setting, top-p, limits the choice to the most probable tokens instead.

  • Low: more consistent wording, well suited to factual answers, data extraction and classification.
  • Medium: some variation in wording while staying on topic, a common default for chat.
  • High: more varied and surprising output, useful for brainstorming or creative writing, but more likely to drift off topic or make errors.

Example

Ask a model five times to write a tagline for a bakery. At a low temperature you get nearly the same line each time. At a high temperature you get five quite different ideas, some useful and some odd. Ask the same model about the bakery’s opening hours, and you want the low-temperature behavior: one consistent answer based on the opening-hours page.

Why it matters for business chatbots

Customer-facing chatbots usually work best with low to moderate temperatures, because consistency matters more than creativity when the questions are about prices, policies or products. Still, temperature is a minor adjustment, not a fix. Missing or outdated content, vague instructions and weak retrieval are more common causes of bad answers than a poorly chosen temperature.

Temperature and intoCHAT

intoCHAT does not let customers set temperature or choose a model. Every agent runs on one OpenAI model chosen and managed by intoCHAT, with its settings handled for you. To change how your agent answers, you adjust what it knows and how it is told to behave: your sources, your instructions, Q&A pairs for answers that must be exact, and the welcome message and suggested questions.

Frequently asked questions

What is a good temperature for a customer service chatbot?

Low to moderate values are a common choice, because consistent, factual answers matter more than variety. The best value depends on the model and the task, so test with real questions. Temperature matters less than good content, clear instructions and solid retrieval.

Does a temperature of 0 stop hallucinations?

No. A temperature of 0 makes the model pick its most likely continuation, but that continuation can still be wrong. Hallucinations are reduced by grounding answers in reliable sources and instructing the model to say when it does not know.

What is the difference between temperature and top-p?

Both control randomness in how the next token is chosen. Temperature reshapes the whole probability distribution, while top-p limits the choice to the smallest set of tokens whose combined probability reaches a threshold. Providers usually recommend adjusting one or the other, not both.

Can I change the temperature in intoCHAT?

No. intoCHAT manages the model and its settings for every agent, so there is no temperature control. You shape answers through your content, instructions and Q&A pairs instead.

See it answer from your own website

Paste your website address and chat with an agent built from your pages. It takes about a minute.

Create your agent free

Free plan, no credit card needed.