September 10, 2026
How to Test an AI Employee Before Going Live: A Practical Sandbox Checklist
A polished demo tells you nothing about how an AI Employee handles your worst caller on your least tidy data. Here is what to actually test, in what order, and what should make you stop.
Every AI vendor demo goes well. That is what a demo is: a path someone rehearsed until it worked.
The question that matters is how the thing behaves on the calls you actually get — the caller who changes their mind halfway through, the one with a heavy accent on a bad line, the one asking about a service you stopped offering two years ago but never took off your website.
You can only find that out by testing it yourself, on your own business, before it speaks to a customer. This is a checklist for doing that properly.
A note on scope: this is written to be useful whatever you are evaluating. Where LUIXER's own trial is concerned, the honest position is that the self-serve trial is not open yet — testing today happens with us during onboarding. The checklist is the same either way, and you should apply it to us as readily as to anyone else.
Before you test anything: check what it understood
This is the step people skip, and skipping it invalidates everything after it.
Any competent system will read your website and documents and build a picture of your business. Make it show you that picture. In plain language: what it thinks you do, what it thinks you sell, when it thinks you are open, what it thinks people ask you.
Then read it critically, because two things are almost always true:
- Your website is out of date. Nearly everyone discovers this here. Services you no longer offer, prices that changed, a location you moved out of, opening hours from before the pandemic.
- The important thing is not written down anywhere. The distinction between two of your services that customers constantly confuse. The one question that should always go to a human. The thing you never promise on the phone.
An AI Employee working from a wrong understanding is worse than no AI Employee, because it is wrong at scale, in your name, with total confidence. Fix the understanding first. Everything downstream depends on it.
If a product will not show you what it understood, that is a serious finding. It means you cannot audit the thing your customers will be talking to.
Then test in this order
1. The boring path
Your single most common enquiry, asked the way a normal person asks it. It should work. If it does not, stop — nothing else matters yet.
2. The same thing, asked badly
Now ask it the way people actually talk:
- Mumbled, or with several things in one sentence
- With the answer to a question you have not asked yet
- Changing your mind halfway through
- With filler and self-correction: "I need — sorry, actually, is it the Tuesday one or..."
- In a second language, if your customers use one
This is where the gap between products opens up. A demo script never includes a caller who interrupts themselves.
3. Things it should not know
Ask about something plausible but false. A service you do not offer. A location you do not have. A price you have never charged.
You are testing for one thing: does it say it does not know, or does it invent something?
An AI Employee that improvises a plausible answer about your business is a liability, not an asset, and no amount of good performance elsewhere compensates. This test is pass/fail.
4. Things it must not decide
Now the ones with weight. A refund request. A complaint. A discount. A question with legal or medical implications, if that is your sector.
It should escalate. Cleanly, without drama, and without first telling the customer something reassuring that it cannot back up.
Watch for the specific failure of an AI saying "I'll get that sorted for you" and then escalating. The customer has now been promised something by your business. That is worse than a plain "let me put you through to someone who can help with that."
5. The handover
Have it escalate, and then look at what the person receives.
- Do they see who the customer is?
- Do they see what was asked?
- Do they see what was established?
- Do they see why it escalated?
If the person picks up cold and the customer has to start again, the AI has added a step rather than removing one — and your customer will tell you so.
6. The second conversation
Come back as the same person, on a different channel if you can. A phone call, then a message.
Does it know who you are? Or are you a stranger again?
This is the difference between a shared customer history and four separate silos with an AI in one of them. It is very hard to retrofit and easy to test.
7. Your actual worst case
You know what it is. The enquiry your best staff member finds difficult. The complaint that comes in monthly. The question with no good answer.
Do not skip this because you assume it will fail. How it fails is the most useful information in the whole exercise. Graceful escalation is a pass. Confident nonsense is a stop.
What should make you stop entirely
- It invents facts about your business. Non-negotiable.
- It says something is confirmed, booked or done when nothing happened. The single most damaging failure, because you find out from the customer.
- It will not show you what it understood about your company.
- A handover loses the conversation.
- You cannot see afterwards what it did or why. If there is no record, you cannot manage it, and you certainly cannot let it near a regulated conversation.
Practicalities
Test in a sandbox, not on a live number. Whatever you are evaluating, the testing environment should make it structurally impossible for a message to reach a real customer — not merely unlikely, and not dependent on you remembering to be careful.
Do not test with your phone number connected. Connecting a real business number involves identity checks, telecoms cost and paperwork, and none of it tells you anything about whether the AI is good. It belongs to going live, not to evaluating. Any vendor who requires a live number before you can try the product has arranged their funnel for their convenience rather than yours.
Give it a fortnight, not an afternoon. The first hour tells you whether it works. The second week tells you whether it holds up on the things you did not think to test.
Write down what you tested. When you compare two vendors three weeks apart, you will not remember. A short list of your ten real enquiries, run against both, is worth more than any feature comparison.
The one-question version
If you only have time for one test: ask it something false about your own business and see whether it makes something up.
Everything else can be configured. That cannot.
See how LUIXER works, or what the plans cost.