AURA51.aiBetaRequest a pilotRequest

Why we publish what doesn’t work yet

In short

Because a flattering number gets caught at the first event, and then nobody trusts the correct ones either. So each capability of our AI avatars shows how far it’s proven, even when it has only been measured in the lab. If an answer drops the caveat, the website won’t build.

Last updated 2 October 2026

01Tests

A passing test isn’t a measurement

An automated test checks whether the software recognizes a phrase, such as a German “Danke” (“thanks”) with the AI avatar’s name, spoken into an answer while it runs. The test can’t show whether the phrase comes at the right moment live, or how long it then takes until the avatar goes quiet.

After one live test, our log showed the automated tests passing, including the ones for stopping by voice. The run itself never exercised that stop, though. Nobody had said “Danke” to the AI avatar at all. So we kept listing stopping by voice as verified in testing, not yet measured live. In a later run, a “Danke” with its name stopped the avatar live mid-sentence, and it fell silent after about a second. Only that measurement changed the evidence level; one more passing test wouldn’t have.

When we write “verified in testing”, the software can do it, and how it plays live is still open. For stopping, both ways now carry the evidence level “measured · live”: the STOP button works in under half a second, the “Danke” with its name after about a second.

02Trust

A flattering number costs you the correct ones too

Live on a laptop with people in the room, the AI avatar’s first sound came after 1.5–2 s, with a median of about two seconds. “Answers in under a second” would look better in a brochure, and nobody would check it before the first event. It would also be false. With our technology its answer hardly ever comes in under about 1.5 s.

If you bring an AI avatar onto your stage, you plan with that number, for instance the pause after a question to it. If the number is flattering, the whole room hears it. After that, the organizer stops believing the numbers that are right, too.

That’s why every capability sits next to the condition under which we proved it, even when that condition is just the lab or an automated test. For response time, that means runs on a laptop with people in the room. The first live runs were much slower, and we logged that too.

03Coupling

Leave out a caveat and no new version goes online

An earlier version of our frequently asked questions said a plain yes three times. It was about capabilities that the same page, further up, listed as “not yet live”. One answer even relied on a capability that didn’t exist. Whoever changed an evidence level at the top couldn’t see which answer further down depended on it.

Since then, every answer records which capabilities it relies on, and pages like this one follow the same rule. If one of them isn’t measured live and the answer lacks the “not yet”, the website build stops and the new version doesn’t go online.

The check only sees what has been recorded, though. Anyone who writes a new claim without recording what it relies on gets past it. It doesn’t replace reading. It catches forgetting.

If you find a statement on this website that promises more than its evidence level supports, write to us. It slipped past the check, and we’ll change the text, not the evidence level.

More in Journal