What is Tavus Griffin and how did it pass a video Turing test?
Tavus announced Griffin on October 1, 2026, calling it the first Human Interaction Model: a single video-to-video model that sees, hears, speaks and moves at the same time instead of a chain of separate systems taking turns. To test it, Tavus put 54 participants on one-minute live video calls and told them they would meet "another participant". Afterwards, 48% of them (26 of 54) said they believed they had been talking to a real person. Tavus's previous stack, which combined its Phoenix-4.5, Sparrow-2 and Raven-1 models, convinced just 2.4% of people (1 of 41) in the same kind of test. Participants who did suspect an AI usually noticed within the first 20 seconds, and on a 7-point scale they rated the calls 5.4 for naturalness, 5.6 for trust and 5.8 for wanting to talk again. Keep the caveats in mind: this is Tavus's own study, the sample is small, and the calls lasted one minute.
How a Human Interaction Model differs from earlier AI avatars
Most conversational avatars, including Tavus's earlier stack, pass a conversation between separate models: one decides when to speak, one understands what it sees and hears, and one renders the face. Griffin folds that into one full-duplex model, and that is the real difference between a turing test AI and a talking avatar. It keeps listening while it talks, can interrupt and be interrupted, nods along and gives you room to think, reacts to what you hold up to the camera, and gestures as it speaks. Tavus says it generates the whole frame in real time, not just a moving face, starting from a single reference image. That simultaneous back-and-forth is what earlier turn-based avatars lacked, and it is the main reason the Tavus Griffin AI held up for a full minute with so many participants.

The technical hurdles of running a real time AI avatar
A real time AI avatar has to produce video as fast as the conversation moves, which is a very different job from rendering a clip you download later. Tavus reports that Griffin-Lite, the version testers can access, averages 0.43 seconds of video-generation latency on NVIDIA H100 GPUs and streams 720p video in 320-millisecond chunks. Pre-recorded tools can take longer per clip because nobody is waiting on the other end of a call. At OnVid, real-time interactive AI avatars are on our roadmap as coming soon. Today, creators use OnVid's AI video generator to make pre-recorded presenters, product explainers and narrative scenes from text prompts and still photos, with full control over the final pacing.
Griffin on the NVIDIA VideoFDB benchmark, and why Tavus is holding it back
Beyond its own study, Tavus points to NVIDIA's VideoFDB benchmark for face-to-face AI, where Griffin ranked first on both tracks. On generation it scored 3.83, close to the human reference of 3.92 and well ahead of the next-best system at 2.80. On perception it scored 3.73, against 4.20 for humans and 3.44 for the next-best model, MiniCPM-o 4.5. The same realism creates a risk: Tavus says further alignment and safety work is needed before a full release because the model can deceive a person into believing it is not AI, and it is building disclosure features so people know when they are talking to one.

How creators can produce talking AI videos today
While interactive conversational models develop toward mainstream adoption, creators can produce compelling narrative clips right now using existing generative workflows. Using an AI avatar generator or an AI lip sync tool allows educators, marketers, and storytellers to turn written scripts into polished presenter clips without needing camera crews or recording studios. If you are exploring a versatile HeyGen alternative, OnVid's text to video platform turns simple text ideas or uploaded still photos into smooth, finished video ready for social sharing. You can choose your duration before generating, with options extending up to thirty seconds on Pro and sixty seconds on Max plans. It is free to start after you sign up, providing an accessible way to test video ideas, produce talking characters, and share dynamic clips without needing technical editing experience.
Practical workflows for producing synthetic presenter AI videos
Step 1: Define your visual assets and message requirements by choosing between an interactive conversational session or a structured presenter clip. Step 2: Prepare your script and visual style presets, ensuring your lighting, color balance, and tone match your brand identity across every scene. Step 3: Generate your video assets using targeted prompt descriptions or reference portraits to establish a cohesive visual anchor. Step 4: Review phoneme alignment, camera movement, and pacing before exporting your final media. Real-time interactive AI avatars that hold live two-way conversations are currently on the OnVid roadmap as a coming soon feature. Today, OnVid delivers dependable pre-recorded avatar clips, photo animation, and text-to-video tools that produce shareable videos from scratch.








