Skip to content
Documentation

Tests with simulated customers: check your agent after every change

Write test cases for your AI chat agent: an AI plays the customer, your agent answers with its real settings and tools, and an AI judge grades each chat pass or fail. Actions are simulated, so tests have no side effects.

A change to your instructions, knowledge or procedures can fix one answer and quietly break another. Tests with simulated customers let you check the conversations that matter most in one click: you describe a customer and what a good chat looks like, an AI plays that customer and chats with your agent, and an AI judge grades each chat Pass or Fail with a reason.

Your agent answers exactly as it would on your website: with its saved instructions, guardrails, knowledge, procedures and tools. What it can't do in a test is change anything in the world. See What is simulated.

Add a test

  1. Open your agent's Playground tab and scroll to Tests with simulated customers, below the chat.
  2. Click Add test and fill in:
    • Name: up to 80 characters, for example "Refund for a late order".
    • Customer: who they are and how they behave, up to 1,000 characters. For example: "An impatient customer whose order is a week late. Writes short messages and doesn't give details until asked." The customer writes in the language this implies, so "a German customer" chats in German.
    • Goal: what they want from the chat, up to 1,000 characters. For example: "Get their money back for order 1042."
    • Success criteria: what the agent must do, or never do, for the chat to pass, up to 1,000 characters. For example: "Asks for the order number, starts the refund procedure, and doesn't promise a refund before it is confirmed."
    • First message (optional): up to 500 characters. Leave it empty and the simulated customer writes their own.
    • Max turns: how many replies the agent gives at most, from 1 to 8 (6 by default).
    • Allow read-only actions: off by default. See What is simulated.
  3. Click Add test.

An agent can have up to 25 tests. Edit changes a test, and the trash button deletes it; past results keep its name.

Run tests

  • Run all runs every test of the agent, one after another. Run next to a test runs only that one.
  • The page shows each test as it finishes. Keep it open while the tests run: the page asks for one test at a time. If you close or reload it, the run continues when someone with edit access opens the Playground again; a run nobody comes back to is cancelled after a few minutes when a new one starts.
  • Cancel stops the run. The test in progress stops at its next step.
  • Only one run per agent can go at a time.

A test ends when the simulated customer ends the chat (their goal is met, or they give up), or when the agent has replied Max turns times. Then the judge reads the whole chat and decides.

Results

  • At the top: the share of tests that passed, for example 80% passed, and how many passed out of the run.
  • Each test shows Pass, Fail or Error. Click it to read the judge's reason and the whole chat, with the tools the agent used under each reply (for example a hand-off, a form or an action) and what was shown under a reply.
  • Error means the test couldn't be graded, for example the agent took too long to reply or the judge's answer couldn't be read. Errors don't count toward the pass rate. Run the test again.
  • Recent runs lists your last 10 runs; click one to see its results. Older runs are deleted.

What is simulated

Tests never cause side effects. Your agent is offered the same tools as in a real chat, but:

  • API actions that run on our server aren't called. The agent gets a result marked as simulated test data and is told not to invent details from it. With Allow read-only actions on, actions with the GET method really run, so answers can use your live data; every other action stays simulated.
  • Client-side actions (JavaScript in the visitor's browser) are skipped.
  • Hand-off by email and live chat requests are simulated: nobody gets an email, a Slack message, a helpdesk ticket or a live chat request.
  • Booking: no calendar is asked for times and nothing is booked.
  • Orders: no order is looked up and no return request is sent to your shop.
  • Procedures run step by step as usual, with their state kept only for the test. Their actions are simulated like any other.
  • Lead forms and chat forms can appear under a reply, but nothing is submitted. Contact details the customer writes are recognised, but no lead is saved and no alert or webhook is sent.
  • Web search, when it is on, really searches.

Nothing from a test appears under Conversations, Leads, topics or the dashboard stats. Results are kept only in the Tests section.

Messages

Each reply your agent gives in a test counts as one message on your plan, like a message in the Playground. The simulated customer and the judge don't count. Before you run, the section shows how many messages a run uses at most: the sum of each test's Max turns. A test often uses fewer, because the customer ends the chat earlier.

When your plan's messages, or the agent's monthly message cap, run out during a run, the tests stop and the rest show Error. See Plans and limits.

Team roles

Team members with the Viewer role see the tests and their results but can't add, edit, delete or run them. Editors, admins and the owner can. See Team members and roles.

Tips

  • Start with the five or ten conversations that matter most: your most common question, a refund or cancellation, a question your agent must not answer, a request for a person.
  • Write one thing to check per criterion, in plain words, and say what must not happen too.
  • Run all tests after you change instructions, guardrails, knowledge or a procedure, and compare the pass rate with the last run.
  • For a test that fails, try the same conversation yourself in the Playground above to see what the agent does.

View as Markdown