# Choose the AI model your agent answers with

> Pick the AI model that writes your intoCHAT agent's replies: models from OpenAI, Anthropic and Google, how many messages each reply uses, which plans include them, Auto, fallbacks, testing in the Playground and where visitors' messages go.

Each agent writes its replies with one AI model. By default that is GPT-6 Luna from OpenAI, which suits most support conversations. On a paid plan you can pick another model for each agent, from OpenAI, Anthropic or Google, or let **Auto** pick one for each message.

Larger models use more of your plan's messages per reply. The rest of your agent stays the same with any model: its knowledge, instructions, guardrails, actions, forms and procedures.

## Pick a model

1. Open your agent and go to the **Setup** tab. The **Model** card is below **Reply language**.
2. Pick a model, or **Auto**. Each model shows its provider, its tier, how many messages it uses per reply and the cheapest plan that includes it.
3. Click **Save changes** in the header.

Models your plan doesn't include are listed but greyed out, with the plan they need, such as **Needs Pro**. A model only appears when intoCHAT has it set up; if a provider isn't set up, its models aren't listed.

The choice is saved with the agent's other settings, so it shows in [Settings history](/en/docs/settings-history) as **Model** and can be restored from there.

## Models and what they cost

| Model | Provider | Tier | Messages per reply | Plans |
| --- | --- | --- | --- | --- |
| GPT-6 Luna (default) | OpenAI | Standard | 1 | All plans |
| Claude Haiku 4.5 | Anthropic | Fast | 1 | Starter and up |
| Gemini Flash | Google | Fast | 1 | Starter and up |
| Claude Sonnet 5.5 | Anthropic | Advanced | 2 | Pro and up |
| Gemini Pro | Google | Advanced | 2 | Pro and up |
| Claude Opus 5.5 | Anthropic | Premium | 5 | Pro and up |
| Auto | Picks per message | | 1 | Starter and up |

"Messages per reply" is how many of your plan's monthly messages one reply uses. With Claude Opus 5.5, for example, 10 replies use 50 messages. A reply that runs actions or a web search still counts once, at its model's rate. Each agent's [monthly message cap](/en/docs/limits-and-access#monthly-message-cap) counts the same way.

## Auto

With **Auto**, intoCHAT picks a model for each visitor message, from the models in the list that your plan includes. The choice is based on the message, how long the conversation is, whether files are attached and whether actions are on offer.

- Every reply with Auto uses one message, whichever model writes it.
- Auto is available from Starter. On Free it is greyed out.
- Auto never picks a model your plan doesn't include or one that isn't set up.

## How Auto chooses

- Auto uses one message per reply, so it only chooses between the models that use one message: Claude Haiku 4.5, Gemini Flash and GPT-6 Luna (the default).
- No model is asked to decide. The choice is made instantly from simple rules about the visitor's message, so it adds no delay and uses no messages.
- A message with an image or a file goes to a model that can read images, GPT-6 Luna first.
- Complex messages go to GPT-6 Luna:
  - a request to do something, such as an order, refund, booking, cancellation, change, tracking or return, or a message with an order number like #1234, when the agent has something it can do in that conversation (actions, forms, procedures, booking, orders, hand-off or web search, for example);
  - a question in several parts, or a list of points;
  - a message over 400 characters;
  - a request for code or a table;
  - a comparison;
  - any message once the conversation has more than 12 earlier messages.
- Everything else, such as greetings, thanks, yes or no answers, small talk and short questions, goes to the fastest model available: Claude Haiku 4.5, then Gemini Flash, then GPT-6 Luna.
- Only models that are set up on intoCHAT and that your plan includes are considered. If none of them fits, GPT-6 Luna answers.
- The language doesn't change the choice: every one of these models is multilingual, and the rules understand English, German, French, Italian and Spanish.

You can see which models wrote your agents' replies on the **Dashboard**, under **Outcomes**, in the **Models used** card, for the period and agent you pick. With Auto, it shows the model picked for each reply, or GPT-6 Luna when it answered instead after a provider problem.

## Fallbacks

Your visitors always get a reply. The default model, GPT-6 Luna, answers instead of the agent's model when:

- **The plan no longer includes the model**, for example after a downgrade from Pro to Starter. The **Model** card then says that replies use GPT-6 Luna until you upgrade or pick another model. Nothing is changed in your settings, and saving other settings still works.
- **The model isn't set up** on intoCHAT any more.
- **Fewer messages are left this month than the model uses**, but at least one. The reply then uses one message. When no message is left at all, your agents stop answering as usual; see [Plans and limits](/en/docs/plans-and-limits#what-counts-as-a-message).
- **The provider has a problem** before the reply starts, for example an outage. intoCHAT tries GPT-6 Luna once instead, and the rest of that reply stays on it. The reply then uses one message, like any GPT-6 Luna reply: the extra messages of a larger model are given back.

To see which model wrote a reply, open **Why this answer** under it on the **Conversations** tab. The **Model** line names the model, says when Auto picked it and why, and says when the default model answered instead and the reason. See [Conversations and stats](/en/docs/conversations-and-dashboard#why-this-answer).

## Try models in the Playground

The **Playground** uses the model on screen, saved or not, like the instructions and guardrails. Pick a model in the **Model** card, then chat in the Playground to try it before you save. Only people who may edit the agent can try an unsaved model; everyone else gets the saved one.

In **Compare**, each side can use its own model:

- **B** has a **Model** picker next to its other draft settings. **Use B** copies it into the settings form with the rest of the draft.
- **A** answers with the saved model unless you pick another under **A's model**.

Each send in Compare uses the messages of both sides' models, for example two with two one-message models. A side whose model your plan doesn't include, or that isn't set up, is answered by GPT-6 Luna and uses one. The line above the chats shows the number.

## Everywhere your agent answers

The agent's model writes every reply visitors get: in the chat widget, the inline iframe and chat link, the Help Page, [Slack](/en/docs/slack-agent), [email](/en/docs/email-agent), the [REST API](/en/docs/rest-api), the [MCP server](/en/docs/mcp-server) and [tests with simulated customers](/en/docs/agent-tests). In a test, only your agent's replies use its model; the simulated customer and the judge don't use your messages.

intoCHAT's own AI work behind the scenes always uses its standard OpenAI model and doesn't depend on your choice: searching your knowledge, topic analysis, spam checks, the copilot, asking questions about your conversations, live translation and the conflicting facts check.

The REST API returns the agent's model as `model` on an agent, and the model that wrote each reply as `model` on a conversation's messages. See [REST API](/en/docs/rest-api).

## Files and photos

Every listed model reads images visitors attach. PDFs are read by OpenAI's models only: with an Anthropic or Google model, the agent is told that a PDF was attached but can't open it, and answers from what the visitor wrote.

## Privacy

The visitor's messages, the conversation so far, your instructions and the knowledge excerpts used for the reply go to the provider of the model that writes it: OpenAI, Anthropic or Google. With Auto, that can be any of the three, depending on the model picked for each message. When GPT-6 Luna answers instead after a provider problem, the reply's request also goes to OpenAI.

intoCHAT's own AI work behind the scenes, such as searching your knowledge, spam checks and topic analysis, always uses OpenAI, so visitors' messages also reach OpenAI whichever model you pick. If your privacy policy names the companies that process chat messages, list OpenAI and the providers of the models you use.

## Usage by model

For the account owner, **Settings**, **Billing** lists how many replies each model wrote in the current billing period, under **Replies by model this period**. It counts replies saved in conversations, so temporary chats and tests aren't in it. Your plan's message count above it is the number that matters for your limit.

## Limits

- One model per agent. Different agents can use different models.
- A reply that is already streaming when a provider fails can't switch models; the visitor may see the reply stop short. Fallbacks happen before any text is shown.
- Answers differ between models in style and length. Try a model in the Playground before you switch a busy agent.
- Larger models usually take longer to answer than GPT-6 Luna and the fast models, especially in replies that run actions. intoCHAT waits up to 25 seconds for Anthropic or Google to respond before GPT-6 Luna answers instead, so a slow provider can make a reply take longer still.
