The most misleading metric in a voice-AI purchase is demo performance. A ten-minute demo can sound flawless. The differences show up under load, across accents, and on calls where the connection drops halfway through. This checklist exists to keep the decision anchored in operational evidence rather than in how smooth the sales call felt.
Treat it as a working document, not a scorecard. For every line, ask how you would test it. If nobody can describe a test, that line is still a promise.
Ask for latency as a percentile, not an average
End-to-end first-response time is the sum of speech recognition, intent resolution, response generation, speech synthesis and telephony. Ask the vendor to break those numbers out; that is how you find the bottleneck. Put a 95th-percentile ceiling in the contract, not an average. What exhausts a caller is not the mean, it is the occasional long silence.
Budget latency per call type. In step-by-step flows such as identity verification, a short acknowledgement (“one moment, I am opening your record”) lowers the perceived wait. Silence is the most common reason people hang up, so a fast filler line often beats a technically faster but quiet flow.
- Load test: 50 back-to-back calls from the same number, measured at the 95th percentile
- Drop test: reconnect behaviour under packet loss and carrier handover
- Queue test: the priority rule when two calls reach a human at the same moment
- Accent test: the same script read in at least two non-native accents
- Messy-input test: background noise, barge-in and half-finished sentences
Get the human handoff rules in writing
A good handoff happens when defined conditions are met, not when the agent gives up. It should fire when verification fails, when the same topic comes back a second time, when the caller explicitly asks for a person, or when negative sentiment crosses a threshold. Those conditions belong in the contract as separate lines; “escalates when needed” tells your operations team nothing.
- Pre-handoff summary: topic, verified identity and steps already attempted reach the operator
- Wait announcement: position in queue and an honest estimate
- Post-handoff label: which condition fired and the agent's last three utterances
- Return rule: how the record is classified if the operator closes the call
Handoff rate is not a success metric on its own. A very low rate can mean the agent is absorbing calls it cannot actually resolve; a very high rate means it is punting work it should have finished. Read it next to resolution rate, never alone.
Logging, masking and retention
You do not need a full transcript of every call. For most operations, a record of decision points is enough: which step ran, which field was verified, which answer was produced, which tool was called. If you do store audio, define retention, access rights and masking up front. Card numbers, national ID numbers and health details should be masked before anything is written.
A demo shows the agent on its best day. Logs show it on its worst.
Plan the exit as well. In what format will transcripts, call metadata and configuration be handed over? If that line is missing from the contract, the migration two years from now may cost more than the original rollout.
Test language coverage against your own data
A claim of support for more than fifty languages tells you very little on its own. Measure which languages actually arrive in your recordings, in what proportion and with which accents, then test only your top three in depth. For the rest, plan on the assumption that recognition works but verification is limited.
The last step is moving evidence into the contract: latency ceiling, handoff conditions, retention period and language scope written as measurable lines. A commitment nobody can measure becomes an argument in the first quarterly review.
The hard part of adapting this list is deciding which items genuinely matter. If eighty per cent of your volume comes from three scenarios, weight the checklist towards those three and put the rest into a second phase.


