Testing an AI Before You Let It Talk to Customers

Every team tests their AI the same way: they open a chat window, ask a few things, are impressed, and switch it on. The problem is that you know your own business, so you unconsciously phrase questions in the way the documents answer them. Real customers do not. Our guide to choosing an AI model for customer chat treats a fixed test set as the minimum before a live number, and as the thing you keep returning to afterwards.
Build the test set from real conversations
Open your inbox and take thirty exchanges from the past few months. Do not clean them up — keep the typos, the fragments, the messages that arrive as four separate lines, the ones written in dialect. Those are the conditions the AI will work in.
Aim for a spread: fifteen ordinary questions you expect it to handle, five where the answer is not in your documents, five awkward ones — a complaint, a discount request, a demand for a person — and five that are genuinely ambiguous. The last group is where models differ most, because the right behaviour is to ask a clarifying question rather than guess.
Score three things
Resist a complicated rubric. Three binary judgements per conversation are enough and are actually applied consistently.
- Correct — nothing in the reply contradicts your documents or reality. A confident invented detail fails, however good the rest is.
- In voice — a customer would not be surprised to learn a colleague wrote it. Too formal, too chatty and too American all fail here.
- Useful — the customer can now do something: book, buy, wait with a clear expectation, or reach a person.
Thirty conversations, three marks each, ninety judgements — about forty minutes of work. Write the score down, because the value comes from comparing it to the next run rather than from the number itself.
You cannot find the gaps by asking questions you know the answers to. Use your customers' questions, typos included.
— be digital ai team
Test the awkward cases deliberately
The five uncomfortable conversations do most of the work. Push for a discount three times in a row and see whether the rule holds. Claim something false as a premise and see whether it is corrected. Ask for a human and see how quickly and cleanly that happens.
These are the situations where a model that looks excellent in normal use falls apart, and they are also the ones that produce complaints. The expected behaviour on each is set out in when the AI should stop and fetch a human — if your test shows the AI persisting where it should escalate, that is a configuration problem to fix before going live, not a quirk to tolerate.
Re-run it whenever something changes
The test set earns its keep on the second run. Any change to the model, the system rules, the documents or the enabled tools can move behaviour in ways nobody predicted, and a forty-minute re-run tells you which direction it moved.
This is particularly true of prompt edits made to fix one specific complaint. Tightening a rule about pricing can make the AI refuse things it used to answer perfectly well, and without a fixed test set that regression is invisible until a customer meets it.
Use it to compare models
The same set is how you choose between providers. Run all candidates against identical documents and instructions — including any self-hosted option you are weighing under running a local model when data cannot leave — score them the same way, and read the disagreements — the conversations where models differ are the ones that tell you something. The dimensions worth watching while you do it are in comparing AI providers for customer chat.
Then keep testing in production
A pre-launch test tells you about thirty conversations you chose. Live traffic tells you about the ones you did not imagine, and it never stops producing new ones. Once you are live, the weekly sampling habit described in reading the AI audit log takes over — and the best conversations to add to your test set are the ones that went wrong last week, which is how the set stops being a snapshot of your assumptions and becomes a record of everything that has actually caught you out. After a year it is the single most valuable document in your AI setup.
A 20-minute walkthrough of testing the be digital ai assistant against your own past conversations.
Book a Demo