Tavus, a San Francisco company that builds real-time video AI, published a claim on Thursday that is easy to state and harder to sit with. In a live study, 48% of the people who spent one minute on a video call with its new model came away believing they had been talking to another human being.
The number worth putting next to it is 2%. That was the best pass rate Tavus had managed with its own previous stack, which was already regarded as the leading conversational video product on the market.
The model is called Griffin. The company calls it the first Human Interaction Model, a category label it invented for the occasion, and the pitch underneath the branding is narrower and more interesting than “better chatbot.” Griffin is not trying to be smarter than the assistant you already use. It is trying to be smoother.
The short version
- 48% of participants believed Griffin was a real person after a one-minute video call, according to a study Tavus designed and ran
- The previous ceiling was 2%, set by Tavus’s own cascade of Phoenix 4.5, Sparrow-2 and Raven-1. The company’s public posts framed the old bar as “under 3%”
- It is one model, not a relay. Griffin is a full-duplex video-to-video system: it perceives, decides and generates at the same time instead of waiting its turn
- Every pixel is generated, from a single reference image. The face, the hands, the chair, the shadows and the background are all produced live
- Tavus also claims first place on Nvidia’s independent test of face-to-face AI, and a 37% lead over the next best system at reacting in the moment
- Voice cloning needs about 10 seconds of reference audio
- You cannot use it yet. Griffin-Lite is a research preview limited to a small group of early testers
- The obvious problem is that convincing live video of a person who does not exist is the missing piece in a fraud playbook that already works without it
What the 48% actually measured
The framing matters here, because “passed the Turing test” has been claimed and withdrawn enough times to be worth no reader’s trust on its own.
Alan Turing’s 1950 proposal was about text. An interrogator exchanges written messages with a human and a machine, and the machine passes if the interrogator cannot reliably say which is which. Tavus is testing something different, and in one respect much harder: a live, face-to-face video conversation, where a model has to produce a voice, a face, an expression and correct timing all at once, in real time, under the scrutiny of a person looking directly at it.
The study design is simple. Participants were recruited through an independent research platform and told they would meet another participant for a one-minute chat about what they were looking forward to this year. Afterwards they were asked whether the person they had met was real. Just under half said yes.
Two caveats travel with that figure and neither of them makes it uninteresting.
The first is that Tavus designed, ran and reported its own study. That is normal for a model launch and it is also exactly why the number needs independent replication before anyone treats it as a settled fact. The company has published a research breakdown, which is more than most launches of this kind bother with, but a vendor measuring its own product is a vendor measuring its own product.
The second is that one minute is short, and a chat about your plans for the year is about the lowest-stress conversation anyone could design. Nobody in that study was trying to catch a machine out. They were told they were meeting a person, so they met a person. Whether 48% survives ten minutes, an adversarial interviewer or a deliberately strange question is unknown, and Tavus has not claimed otherwise.
That dashed line at 50% is the part people keep missing. Griffin did not beat the coin flip. It arrived just underneath it. If the number were 75%, something stranger would be going on, because participants would be more likely to call a machine human than to call a human human. At 48%, the honest reading is that Griffin became roughly indistinguishable in a brief, friendly exchange, which is both a real result and a much narrower one than “AI is now human.”
Why the architecture is the actual news
Strip out the marketing and the technical claim is about plumbing, which is usually where the interesting part lives.
The assistants most people have used are cascades. Your speech is transcribed, the text goes to a language model, the model’s answer goes to a voice synthesizer, and in the video products a face gets animated to match. Each stage waits for the one before it to finish. That is why talking to a voice assistant feels like using a radio: you stop, it thinks, it speaks, and if you pause to collect a thought it treats the silence as your turn ending and barges in.
Griffin is built as one system with two halves that run continuously. A conversational engine takes in your audio and video and, at sub-second intervals, decides whether to speak, hold, yield, nod or react. A generation engine turns those decisions into voice and video as they arrive, rather than after a hand-off.
| Behavior | Turn-based cascade | Griffin |
|---|---|---|
| When it listens | Only while you speak | Continuously, including while it is talking |
| What it perceives | Audio, transcribed to text | Audio and video, including gaze and expression |
| Deciding to speak | When your audio stops | Every sub-second interval, based on content not silence |
| If you interrupt | Usually talks over you or resets | Stops, yields, keeps its place |
| Backchannels | None. Dead air while it thinks | Nods and “mm-hm” while you are still speaking |
| What is on screen | An animated face on a fixed background | Whole scene generated live, shadows and chair included |
| Audio granularity | Sentence or phrase level | Packets as small as 10 milliseconds |
The speech side runs on a codec Tavus calls Tavec, a convolutional autoencoder that maps 48 kHz audio into 40 values per frame at 100 frames a second with no discrete token lookup. Its decoder is fully causal, meaning it needs no look-ahead and can push audio out the instant a frame exists. A diffusion transformer generates the latent speech progressively as control signals arrive, so the model starts talking before the sentence it is saying has been decided.
That is the mechanism behind the thing testers noticed: it does not pause awkwardly, because it never had to assemble a complete thought before opening its mouth.
The demos Tavus published lean on that. In one, the model coaches someone through a Rubik’s Cube, watching the cube turn and waiting when the person stops mid-solve. In another it plays Simon Says and refuses to copy a gesture that was not prefaced correctly. In a third it guides someone soldering a board, tracking how long a step has taken and speaking up when the next one is due rather than whenever the room goes quiet.
The part that arrives before the product
Griffin-Lite is a research preview for a handful of testers. The fraud economy does not wait for general availability, and it does not need Griffin specifically.
The template is already proven. In early 2024 a finance employee at the engineering firm Arup joined a video call with what appeared to be the company’s chief financial officer and several colleagues, and transferred roughly 25 million dollars. Every face on that call was synthetic. The deepfakes were pre-rendered and non-interactive, which is to say they were considerably worse than what Tavus just demonstrated, and they were good enough.
What a full-duplex model changes is the last defense anyone had. The advice for years has been to interrupt, ask something unexpected, make the person on the other end react in real time. A system that interrupts back, holds a pause, and reads your face while it answers removes that test from the list.
This is the same curve that ran through audio first. Cloning a voice went from a research demo to a commodity in about two years, and voice agents became the fastest-compounding corner of the industry on the back of it. Video is a harder problem and the demand is identical, so the honest expectation is that the gap closes the same way.
It also makes verification advice age badly. The heuristics people learned for still images, the extra fingers and the melted jewelry, stopped working a while ago, which is why the old tricks for spotting synthetic media have largely been replaced by provenance checks. Live video is heading to the same place: the question stops being “does this look real” and becomes “can this channel prove who is on it.”
The romance and investment fraud world will get there first, because it is the use case where a convincing face on a live call is worth the most money. That industry has already produced its own strange feedback loop, where the man convicted in the best known case is now trying to sell an anti-fraud dating app of his own.
What Tavus is actually selling. The company’s stated market is customer-facing conversation: interviews, onboarding, support, coaching, healthcare intake. Those are jobs where the awkward pause is the product’s biggest weakness, so smoothing it out has obvious commercial value. Nothing about Griffin is built for deception. The problem is that indistinguishability is not a feature you can ship to one set of users and withhold from another, and the company’s own headline number is a measure of exactly that.
What to watch next
- Independent replication. Nvidia’s benchmark result is the closest thing to outside validation so far, but the 48% figure is Tavus’s own. A third party running the same protocol is the test that matters
- Longer conversations. One minute is the friendliest possible window. Five minutes, or a participant actively hunting for the seam, is a different number
- Disclosure defaults. Whether Griffin announces itself as AI by default, and whether that can be switched off by whoever deploys it
- Consent on likeness. Full-scene generation from one reference image means one photograph is the whole input. Whose photograph, and with what permission, becomes the entire question
- Who gets access. A research preview with a short list of testers is a very different risk surface from a public API with a credit card form
- Detection. Whether anything in the pipeline leaves a signal a bank or a platform could check on a live call, or whether verification has to move to the account layer instead
The bottom line
Griffin looks like a real engineering result and a badly oversold headline at the same time, which is the normal condition of AI announcements.
The jump from 2% to 48% is too large to dismiss, and the reason for it is specific and checkable: Tavus stopped chaining four models together and built one that perceives and generates concurrently. That is a genuine architectural change, and the conversational texture it produces, the interruptions, the nods, the willingness to wait through a pause, is the part people in the study appear to have responded to.
What it is not is a machine that passed for human. It is a machine that passed for human in a one-minute chat with people who had no reason to suspect anything, measured by the company that built it, and it still landed on the losing side of a coin flip.
Both halves of that sentence are worth holding onto. The second one is why nobody should be rewriting their understanding of AI today. The first one is why the “just ask them to wave” era of fraud advice is over.
Sources and further reading
- Tavus: Introducing Griffin, the First Human Interaction Model
- UNILAD Tech: New AI becomes first ever model to pass video Turing test
- Business Today: Griffin passes video Turing test with 48% human success rate
- Cybernews: Griffin AI model interrupts like a real person on video
- Crypto Briefing: Tavus unveils Griffin, claiming the first video Turing test pass
- Background: Alan Turing’s 1950 imitation game
About this article: GeekBlog covers U.S. technology news, AI, phones, smartwatches and gaming. Every story is written and checked under our Editorial Policy. Spotted a mistake or have a story tip? Contact our editors.

