Putting a voice agent live is not a switch. The call queue is a shared resource: every mistake the agent makes affects not only that call but the waiting time for everyone else on the line. The goal for month one is therefore not speed. It is controlled learning.
This plan covers the month that follows a typical six-week build. Its spine is simple: run invisible first, take a small share of real traffic second, tune the handoff rules last.
Week one: shadow mode
In shadow mode the agent does not answer calls. It listens to live traffic and drafts its own response, which is recorded but never shown to the operator. By Friday you have a comparison set built from real language, real background noise and real scenarios.
- The agent's draft answer is stored next to the operator's actual answer
- Recognition errors are marked word by word
- Questions the agent could not handle drop into a separate list
- Silences and interruptions are timestamped
- Every call gets a suitable / not suitable for automation label
The most valuable output of shadow mode is what the agent does not know. A campaign missing from the catalogue, an undefined return rule, a question about new regulation — all of it surfaces in week one. Closing those gaps before launch prevents most of the bad calls that would otherwise happen on day one.
Week two: a small share of real traffic
In week two the agent starts answering real calls, on a deliberately small share of traffic. The point is to observe real behaviour without putting the queue at risk during peak hours. Tying that share to a single number or a single scenario makes the comparison clean.
Watch one metric daily this week: the reason for each handoff. Seeing which condition sends a call to a human tells you both where the agent's limits are and how it affects the queue. Resolution rate is still too noisy to be useful at this sample size.
Weeks three and four: tuning the handoff
As the traffic share steps up, the handoff thresholds get adjusted. An agent that escalates too often wears out the operations team; one that resists loses callers. The right setting is per call type, not global.
Month one is not judged by how much the agent resolves, but by how much it resolves reliably.
The most common mistake in this period is raising the handoff threshold to make the numbers look better. It works for a fortnight; then repeat calls climb. Handoff decisions belong next to the listening notes, not in a dashboard on their own.
- Record the trigger condition for every handoff
- Ask the operator one question: was the summary enough?
- Flag callers who ring back about the same topic
- Review the three most common handoff topics weekly
What month one should produce
Thirty days in, you should hold three things: a scenario list validated against real calls, a measured latency profile, and handoff rules the operations team has agreed to. Raising the traffic share without those three grows risk rather than learning.
Weekly listening is the part everyone skips. Dashboards show the trend; only the recording explains why a call went badly. One hour a week here sets the roadmap for the next three months.
Write the month's outputs down: which scenario hands over under which condition, which step fails in which language, which call type exceeds the latency budget. Learning that is not written down disappears when the team changes, and the next operator repeats the same mistakes.
Finally, treat the month-one plan as a learning plan, not a performance target. The traffic share should rise when findings are closed, not when the calendar says so. When the findings list is empty, going live is safe.


