Our voice agent work has taught us a small, specific lesson that nobody warns you about: instant response is not always an improvement, and the reason is about pacing, not accuracy.
What hold music was actually doing
Nobody likes hold music, but it does real work: it gives a caller a moment to gather their thoughts about why they called, and it signals that something is happening on the other end. A voice agent that picks up in under a second and immediately asks a precise question can catch a caller mid-thought, before they've settled into the conversation, and the resulting call is noticeably more halting than the transcripts of calls answered by an unhurried human.
The fix wasn't slowing the agent down
We tried adding an artificial delay, which felt dishonest and tested badly — callers could tell. What worked was changing the opening line from a direct question to a brief, open acknowledgement, giving the caller a beat to start talking on their own terms rather than being interrogated from the first syllable.
What else we tuned
- Filler and backchannel sounds — small acknowledgements while the caller is speaking — which humans use constantly and early voice agent scripts omitted entirely, making the agent feel like it had gone silent even while it was listening correctly.
- Turn-taking pauses tuned to the specific call type. A booking call needs less thinking room than a complaint call, and we now vary the pacing by the call's detected purpose rather than using one pace for everything.
- An explicit "let me check" pause before answering anything that requires a live lookup, so the wait has a stated reason rather than feeling like a stall.
Removing all the friction from a conversation doesn't make it feel more human. It makes it feel like an interrogation with good manners.
The measurable effect
Call completion rate — callers who reached a resolved outcome rather than hanging up or asking for a human partway through — improved noticeably once pacing was tuned, on a system whose actual comprehension and grounding hadn't changed at all in that period.
What this means for how we build these
The accuracy of the model answering the question turns out to be necessary and not sufficient. The shape of the conversation carries a good part of whether a caller trusts what they're talking to, and it's the part that's easiest to skip because it doesn't show up in any transcript-accuracy metric.
