Why doesn’t an AI answer as fast as a person?
In short
Mainly because it first has to detect that the question is over. People plan their answer while the question is still running and start answering a yes-or-no question 208 milliseconds after it ends, on average. An AI waits for silence. In our live tests, the first sound came after 1.5–2 s, a multiple of the everyday pause.
Last updated 2 October 2026
01Too early
Answer faster and you cut in more often
People who ask a question freely sometimes stop in the middle of a sentence, search for a word and then carry on. To a voice AI, that pause sounds like the end of the question. It has to infer from silence that someone is done, and it can get that wrong in either direction.
A test participant put a German question to the AI avatar: “… wie sollte sich eine Marke verhalten auf Social Media?” (“… how should a brand behave on social media?”). The person paused for just under a second, then kept talking. The system released the answer anyway while they were still speaking. Just under half a second after their last word, the AI avatar came in, although their sentence was not finished.
It has also gone the other way. For a short question, the system decided 0.8 seconds too late that the speaker was done. Pause detection had been set too sensitive and counted even breathing and rustling as speech. After asking your question, keep the microphone quiet until the AI avatar comes in.
To answer sooner, a voice AI would have to bet on silence sooner, and the sooner it decides, the more often it is wrong. With our characters, the first sound hardly ever comes in under about 1.5 s anyway, not even with a faster language model. So if you put a question to an AI avatar, do your thinking before the question, not halfway through it.
02Between people
People plan their answer before the question is over
A study measured how fast people answer yes-or-no questions in ten languages on five continents. On average, the answer began 208 milliseconds after the question ended. The average was 7 milliseconds in Japanese and 469 in Danish. In all ten languages, the most common gap was between 0 and 200 milliseconds.
That is not enough time to build an answer. A review paper puts planning a single word at about 600 milliseconds and the start of a simple sentence at about 1,500. To answer on time, people plan while the other person is still talking and predict when that person’s sentence will end.
In conversation, a pause means something. In all ten languages, a confirming answer came 100 to 500 milliseconds sooner on average than a disconfirming one, a statistically significant difference in seven of them. In English telephone calls, responses to requests, offers or invitations were more often rejections than acceptances once the gap reached about 700 milliseconds.
- 208 msaverage time from question to answeryes-or-no questions in ten languages (Stivers et al., 2009)
- 700 msfrom here on, more rejections than acceptancesresponses to requests, offers, invitations; telephone calls (Kendrick and Torreira, 2015)
03On stage
On stage, the moderator hands out the floor
The 208 milliseconds come from conversations in which anyone can take the floor at any moment. On a stage, the moderator calls the AI avatar by name. The room then knows who speaks next.
There, the pause follows a call. The telephone study looked at responses to requests, offers and invitations, meaning things a person can turn down. As we read it, a pause after a call therefore carries less meaning. Still, about two seconds is almost ten times the 208 milliseconds.
The moderator is poorly placed to judge how long this pause feels in the room, because they are already reading the answer on the tablet while the room is still waiting. At the dress rehearsal, sit in one of the back rows while the moderator calls on the AI avatar, and listen to how long the silence feels from there.
Sources
- Stivers et al.: Universals and cultural variation in turn-taking in conversation. PNAS 106 (26), 2009
- Levinson and Torreira: Timing in turn-taking and its implications for processing models of language. Frontiers in Psychology 6, 2015
- Kendrick and Torreira: The Timing and Construction of Preference: A Quantitative Study. Discourse Processes 52 (4), 2015