What has been measured on the technology, and under what conditions?
In short
We measured live at a laptop with people in the room, including the name call, the first sound, stopping it with the STOP button or with “Danke” (thank you) and its name, and addressing guests by the names from the briefing. Interjections we measured only in the lab, not yet live. Every number here comes with its condition.
Last updated 2 October 2026
- 1.5–2 sTime to first audiomeasured · live
- 334 msStopmeasured · lab
- 12 sConversationa setting
- 1.3 sInterruptionmeasured · lab
What can you not rely on yet?
How the AI avatar reacts to an interjection is measured only in the lab, not yet live. The characters do not yet understand or speak English; they work in standard German. How it summarizes long discussions we measured only with prerecorded ones. The other capabilities we list were measured live at a laptop, with people in the room.
The measurement log: results, failures and what we changed
We measured live at a laptop, with people in the room, and in the lab. The main failures we found are listed here by topic, each with the change it led to.
01Evidence levels
Two numbers for the same moment
From the last word of the call to the first sound, the AI avatar takes 1.7–1.8 s in the lab. With people in the room speaking freely, it takes 1.5–2 s, with a median of about two seconds. When we talk about response time, we quote the second number, even though its median is above the lab value. It was recorded under conditions closer to a real event.
That is why every capability here carries one of four labels. The label says under which condition we have evidence for it.
- measured · live
- Occurred live and was measured: people in the room, natural speech, a laptop with its built-in microphone.
- measured · lab
- Measured with the AI avatar’s own voice but with nobody in the room, for instance with a prerecorded discussion.
- verified in testing · not yet live
- Automated tests cover the capability; it has not yet been measured with people in the room. What was last at this level has since been measured live, so no capability is here right now.
- built · not yet verified live
- Built, but neither covered by automated tests nor tried live. No capability is at this level right now either.
02Numbers
Six numbers and where they come from
Not every number shifts with the condition. The STOP button took effect in under half a second, in the lab and at the laptop alike.
One of the six is not a measurement. How long the AI avatar keeps listening without its name after the last sentence is a setting, 12 s, not a limit of the technology.
- 1.5–2 sto the first soundafter the last word of the call · live at the laptop, median about two seconds · usually an opener like “Also.” (well) first, the content follows
- 1.7–1.8 sto the first sound, in the lablab, prerecorded discussion
- 1.5 sabout this fast at bestthe technology’s lower limit, rounded · it rarely gets faster
- 334 msuntil it falls silentafter STOP on the tablet · lab · live under half a second
- 12 sconversation without a name callfrom the last sentence · a set value, not a measurement · after that its name is needed again
- 1.3 sto the “Moment” linein the lab · from the triggered interjection, excluding speech recognition · not yet measured live
The measured values apply to all four characters, since the same technology runs underneath Albert, Vera, Tina and Ken. Everything was measured in standard German, as the characters do not understand or speak English yet.
03Capabilities
What came up live and what is still waiting
- measured · live
Listening
Whatever was said in the last few minutes the AI avatar has word for word. Everything earlier it summarizes; that we measured only in the lab, with a prerecorded discussion, not yet live.
- measured · live
Answering when called by name
The moderator calls its name and asks, and it answers. Talking about it is not a call.
- measured · live
Staying in the conversation
Up to 12 s after the last sentence, an “Und warum?” (and why?) without its name is enough.
- measured · live
With reference and a position
It calls guests by name and will contradict them, too.
- measured · live
Choosing its own length
Asking for a “Schlusswort” (closing word) gets one sentence, “erklär uns” (explain to us) a full explanation of a good half minute, everything else three to four sentences.
- measured · live
Starting naturally
A neutral “Also.”, “Hm.” or “Nun.” up front, and the first sentence picks up from it.
- measured · live
Understanding split calls
A call spread over two sentences counts as one.
- measured · live
Letting itself be stopped
The moderator ends the answer with the STOP button on the tablet.
- measured · live
Being stopped by voice
“Danke” (thank you) with its name works like the STOP button. Live, it stopped the AI avatar mid-sentence; it fell silent after about a second.
- measured · live
Knowing who is in the room
The names of the moderator and the guests are in the briefing on the tablet. Live, it used them to address the guests correctly.
- measured · lab
Reacting to interjections
If the moderator calls its name into a running answer, it says a short line such as “Moment — den Gedanken führe ich noch kurz zu Ende” (one moment, let me finish this thought) and resumes at the start of the interrupted clause; measured in the lab, not yet live.
04Pilot
From the laptop to your event
From the lab to the laptop, the time to the first sound more than doubled in the first live test. After that, its median settled at about two seconds, just above the lab values.
At the event, the moderator’s microphone comes from your desk on its own channel, the other stage microphones on an aux send, and the AI avatar speaks through the house PA.
We will take our next measurements at pilot events, with your sound engineer at the desk.