Take Sales
Back to blogVoice

Latency is what decides whether your agent sounds human

In text, two seconds of waiting go unnoticed: the person is looking at the screen and knows something is being written. In voice, two seconds of silence is a problem. The person on the other end assumes the call dropped, repeats the question, and now the agent has two overlapping questions to answer.

In a voice conversation the budget between the user’s last syllable and the agent’s first sound is roughly 500 to 800 milliseconds. Above one second it stops sounding like a conversation. That budget has to cover detecting the end of speech, transcribing, retrieving context, generating the answer and synthesising audio, which is why chained architectures rarely fit inside it.

Where the time goes

  • Detecting that the person finished speaking. It looks free and it is not: waiting for too much silence sounds slow, waiting for too little makes the agent cut the customer off.
  • Transcribing the audio into text.
  • Finding the right passage in the knowledge base.
  • Generating the answer.
  • Turning the answer back into audio.

Added up in series, each with its own network round trip, those five steps clear two seconds easily. That is why real-time voice uses models that process audio end to end in streaming, rather than a chain of separate calls. It is not an architectural preference, it is the only way to fit the budget.

What people perceive is not total latency, it is the time to the first sound. An answer that starts at 400 ms and runs for six seconds sounds better than one that starts at two seconds and runs for three.

First token, not full answer

This is the metric worth asking a vendor for, and the one that rarely appears in sales material. Median latency to the complete answer hides exactly what matters. Ask for the median and the 95th percentile to the first token, because it is the tail that ruins the call: if one answer in twenty takes four seconds, the customer will find that one.

800 ms

The ceiling at which a voice answer still sounds like conversation. Above it, people start repeating the question.

The thirty-second test

There is one test that exposes any chained architecture in half a minute: talk over the agent while it is answering. A genuinely streaming system stops, listens and picks up again. A chained system finishes the whole sentence, because the audio had already been generated before you opened your mouth.

What to do with the slack

Low latency is not the goal, it is room. Every 100 ms saved in transport is 100 ms that can become a better search over the knowledge base, or one more check before answering. Teams that optimise latency purely to hit a number end up with an agent that is fast and shallow; the real gain is spending the slack on a better answer inside the same perceived time.

The flip side is that voice does not tolerate the same checks text does. A full faithfulness pass costs time the conversation does not have, so what stays in voice is the cheap and mandatory part: masking sensitive data before it leaves. See response latency and PII masking in the glossary, and the comparison with an IVR if the scenario is telephony.

#Voice#Latency#Architecture

Start the conversation today

Launch an AI agent that qualifies leads and books meetings around the clock. Live in under ten minutes.

A guided demo · your agent live in under 10 minutes