How fast does an AI avatar answer on stage?
1.5–2 s
to the first audio
measured · live
live with people in the room, median about 2 s · 1.7–1.8 s in the lab
In short
In 1.5–2 s. That is the time from the moderator’s last word to the first audible sound, live at a laptop with people in the room; the median is about two seconds. It is rarely faster than about 1.5 s, because the AI avatar detects the sentence end, then generates its voice. We do not promise answers in under a second.
Last updated 2 October 2026
01First test
Almost four seconds in the first live test
For the first time, the AI avatar heard people speaking freely instead of recorded voices. Two people sat in front of a laptop, its built-in microphone the only ear it had. The median time to the first audio was almost four seconds, more than twice as long as in the lab.
The avoidable delay sat mainly at a point nobody in the room could hear. For almost every question, the system first ran a separate check on whether the AI avatar was the one being addressed. That alone took about a second live, and once, with a wait, three seconds.
The way people talk added to it. Someone speaking freely picks up again after a short pause or trails off with an “um.” Speech recognition then registered the end of the sentence later than with the clean lab recordings.
Starting the question with the avatar’s name spared it this check. Where the name did not come first, it took less than a second in the next run. Later, the median came down to about two seconds.
02Measurement
The clock starts at the last word
An example in German: the moderator asks, “Albert, wie siehst du das?” (“Albert, how do you see this?”). The clock starts after “das” and stops at the first sound you hear from the AI avatar. Whatever the moderator says before, such as an introduction, does not count.
With the name in the question, at the start or at the end, the AI avatar could be heard live after about two seconds. When the name came on its own after the question, it took just over four seconds.
- 1.5–2 sto the first audio, liveat a laptop, moderator and guest speaking freely · median about two seconds
- 1.7–1.8 sto the first audio, in the labrecorded discussion · sentence endings detected sooner than live
- 1.5 slower limit, roundeddetect the end of the sentence, then generate the voice · rarely faster
A median smooths things out. Even within one discussion, the fastest answer came after just over one and a half seconds, the slowest after just over two and a half.
03Lower limit
Two steps every answer needs
- 01
Pause detection
The AI avatar has to hear that the moderator is done. A short pause to think is not the end of a sentence. It waits for a short silence, and if the moderator picks up again, it keeps waiting.
- 02
Speech synthesis
The text of the answer is turned into a voice. This step alone takes almost a second before the first sound comes out.
Together that comes to about 1.5 s. Waiting less would mean talking into pauses and answering half-finished sentences.
04Opening
What the room hears in those two seconds
After the question, the room is briefly quiet. Then the AI avatar says “Also.”, “Hm.” or “Nun.” (roughly “So.”, “Hm.”, “Well.”). That word is the first sound our measurement captures, and it carries no content yet.
The content follows a little later. Live, it usually came one to almost two seconds after the opening word began. Counted from the last word of the question, the first sentence with content was usually heard after just under three to just over four seconds. The AI avatar composes that sentence knowing which word has already been spoken, and picks up from it.
For the moderator, the pause is shorter than for the room. The text of the answer appears on the tablet before the room hears its content, usually a good second earlier. Only the opening word comes without text.
05Questions
When does it take longer?
Does a follow-up without the name take longer?
Slightly longer. After an answer, it stays in the conversation for a while, and “Und warum?” (“And why?”) is enough. To such a follow-up, for instance a request to explain the point in more detail, it responds after a median of just over two seconds; to a question with its name, after just under two.
Do all four characters answer equally fast?
Yes, the numbers on this page apply to each of them. Albert, Vera, Tina and Ken differ in stance, face and voice. Underneath is the same technology, and all of them wait for the end of a sentence with the same pause detection.
Based on our laptop numbers, plan your run of show with about two seconds after each question until the opening word, and one to two more until the first sentence with content. With ten questions to the AI avatar in an hour, that adds up to half a minute or a little more.