Voice AI
How to evaluate a voice AI platform without getting burned
Demos are optimised. Here is a practical evaluation framework covering interruption handling, escalation, grounding, governance and the questions vendors hope you skip.
SyntheStudio · VoicaX product team · · 2 min read
- evaluation
- voice ai
- buying guide
Every voice AI demo sounds impressive, because every demo is a happy path. The interesting behaviour is what happens when a caller interrupts, changes their mind, mumbles a postcode or asks something the agent cannot answer.
Test the interruption, not the script
Ask to interrupt the agent mid-sentence. Then ask it something unrelated. Then change your mind. Natural conversation is mostly repair — correcting, clarifying, backtracking — and a runtime that handles turn-taking badly will feel robotic within thirty seconds regardless of voice quality.
Ask what happens when it does not know
There are two acceptable behaviours: say it does not know and escalate, or say it does not know and take a message. There is one unacceptable behaviour, which is a confident, plausible, wrong answer. Ask the vendor to demonstrate the failure path deliberately.
Check where the answers come from
- Is the answer retrieved from your documents, or generated from the model's general knowledge?
- Can you scope a source so a customer-facing agent cannot read internal material?
- When you change a document, how quickly does the answer change?
- Can you see which questions had no supporting content?
Interrogate escalation quality
A transfer that drops the caller into a cold queue where they repeat everything is worse than no transfer at all. Ask whether the human receives the transcript and the actions the agent already attempted.
Ask the governance questions
- Who in our organisation can change a script, and is that change recorded?
- Is our data isolated from other customers, and how is that enforced?
- What is the retention policy for recordings and transcripts, and can we set it?
- Can we review one hundred percent of conversations, or only a sample?
Understand the cost model properly
Per-minute pricing is easy to compare and easy to misjudge. Model your real mix: average handle time, concurrency at peak, escalation rate and the model cost behind each conversation. A platform that resolves in ninety seconds at a higher per-minute rate can be cheaper than one that takes four minutes.
Finally, ask about lock-in
If the speech, language and telephony providers are fixed, a price change or a quality regression at any one vendor becomes your problem. Ask whether you can swap a layer without rebuilding the agent.
