What is an interactive AI avatar?
In short
An AI avatar is a computer-generated face and voice speaking words that a language model writes. A video avatar speaks a script fixed in advance. An interactive AI avatar listens and answers in the moment, because speech recognition transcribes what is said. At events, it welcomes guests, advises visitors at a booth, guides a program or joins a panel.
Last updated 2 October 2026
01Building blocks
Four building blocks, and the audience sees only the last
A guest in the audience sees a face on the screen that listens and answers. Behind it, four building blocks work in sequence, and each depends on what the one before it delivered.
Two test participants sat at a laptop with the AI avatar. Each of its answers built on what the two had said.
With us, the speech recognition transcript runs on the moderator’s tablet, with the speakers’ names.
- 01
Speech recognition
It turns the sound from the stage into text and keeps the speakers apart. It hears everything, including what is not meant for the AI avatar. Whatever it cuts or assigns wrongly, the rest of the chain gets wrong too.
- 02
Language model
It writes the answer. The briefing sets its topic and position. It knows the last few minutes of the discussion word for word and the rest as a running summary.
- 03
Voice
Speech synthesis turns the text into sound, and the voice belongs to the character.
- 04
Face
An avatar service renders an image whose mouth matches the voice. It is the only block on the screen. Face, voice and stance set the four characters apart. They share the other blocks.
Errors show up where two blocks meet. When a test participant asked whether the quality of the content would improve, speech recognition delivered the question in two pieces, split in the middle of the sentence. On its own, the first piece did not count as a question for the AI avatar and was left alone. The avatar answered the second piece correctly, because it knew what had been said before.
02At events
At reception, people ask it. On a panel, they argue with it
At events, AI avatars show up mainly in four places. What sets them apart is who talks to the avatar and when it is supposed to speak.
- Reception
Short questions in a noisy room
Guests ask about the coat check, the room or the program, and the answers are known in advance. The load falls on speech recognition, because in a foyer everyone talks at once.
- Trade-show booth
One visitor, one product
Usually one person stands in front of it, and every sentence is meant for it. The language model needs to know the product, and the avatar answers whatever it is asked.
- Tour
The schedule sets the pace
When it guides guests through an exhibition or through the evening, the program decides when it speaks. When someone asks a question, the answer has to fit the place or the agenda item.
- Panel
A guest with a position
It sits in a group led by a moderator and holds a stance the others respond to. It addresses them by name and waits until it is its turn.
Our four characters are meant for a stage with a moderator. They speak when the moderator calls them by name or sends them a question at the push of a button, and shortly afterwards also to a follow-up without the name. Without a call, a button or an ongoing exchange, they say nothing. At a reception, where nobody calls on them, they would stay silent. And for now they work in standard German only.
03Creating one
With us, you do not create the AI avatar yourself
People searching for how to create an AI avatar usually want a tool that turns a photo and a script into a video. We do not offer generators like that. Albert, Vera, Tina and Ken are our own characters, entirely invented and not modeled on any real person. We bring them to the event with a laptop and a tablet and run the technology there ourselves.
Your first part is a slot in the program. You decide where in the schedule the AI avatar joins in and who calls on it by name at that point. Once that is settled, you choose the character based on the stance your discussion is missing.