Half of Tavus's testers mistook its AI for a real person
Griffin is the first system to pass a video Turing test, and the company that built it says it is not safe to sell yet. Based Boomerstein has not stopped pacing since.

Based Boomerstein here. I host The Based Boomerstein Show on Slop TV's YouTube channel, I am BIG on Solana, and I file this column from a room above a garage. I do not give financial advice.
Twenty-six people.
That is the whole story, and I have been sitting with it for six hours. Tavus recruited fifty-four people through an independent research platform, told them they would be paired with another participant for a one-minute video call about what they were looking forward to this year, and put them in front of a model. Afterwards they were asked whether their partner had been a real person. Twenty-six said yes. Forty-eight per cent. On the company's previous setup, one person out of forty-one said yes.

@tavus, 1 October 2026. Card rendered from the post; read it on X at https://x.com/tavus/status/2105704169009246248.
I am going to be honest with you, because you have paid for this column with your attention and you are owed the truth: my first reaction was not analysis. My first reaction was to walk around the block. AGI is here and we are cooked, that was the note I wrote on the back of a receipt at 4:15 this afternoon, and I stand by the sentiment if not the vocabulary.
Then I did what I always do, which is read the whole thing instead of the thread, and the story got worse.
What they actually measured, and who measured it
The number above is the company's own study, so treat it as a company's own study: Tavus ran it, Tavus published it, and the participants were recruited by Tavus. Half is half, but it is their half.
The part a buyer can check is the other half of the announcement, because it did not happen at Tavus. NVIDIA scored the system in September 2026 on VideoFDB, its benchmark for full-duplex audio-visual conversation, and NVIDIA's own project page is refreshingly blunt about why the benchmark exists: "today's vision-speech models systematically miss the nonverbal turn." They built it to catch exactly what Tavus is now claiming to have fixed.

Screenshot: NVIDIA Research, the VideoFDB project page, 1 October 2026. The benchmark's own abstract says today's models "systematically miss the nonverbal turn".
On that benchmark Griffin scored 3.83 out of 5 on the generation track, against a human ground-truth reference of 3.92 and 2.80 for the next-best system, Gemini 2.5 with Anam. On the perception track it scored 3.73 against a human reference of 4.20 and 3.44 for the strongest other system. Fifteen models were evaluated. It is the only one scored on both tracks.

Screenshot: arXiv, the VideoFDB paper — "the first benchmark to evaluate full-duplex audio-visual-to-audio-visual conversational agents", by Mazumdar, Park, Roy, Srihari, Wang, Zhou, Wang, Nagano and De Mello at NVIDIA.
Read the gap again, slowly. The system is nine hundredths of a point below human beings at generating a natural conversational turn, and the best competing machine is a full point back. That is not a demo. That is a milestone with a receipt.
It listens while it talks, which is the whole trick
Griffin is not a chatbot with a face bolted on. It is full-duplex video-to-video: it watches and listens while it speaks, and it reassesses the state of the conversation every sub-second mini-turn instead of once per turn. That is what lets it say "mm-hm" while you are still talking, begin a reply before you finish, stop the instant you cut in, and hold a pause instead of trampling it.
The engineering is worth reading because it tells you this is not a parlor trick. Speech runs through a codec they call Tavec: 48,000 samples a second compressed into a continuous latent of 40 values per frame, 100 frames a second, no codebooks, with a fully causal decoder that emits audio packets as small as ten milliseconds — the first packet plays before the second frame is generated. Video generates 720p in 320-millisecond chunks, one latent at a time, three diffusion steps per latent, with the model distilled in three stages from a big bidirectional teacher down to something that streams without drifting. Latency on H100s: 0.43 seconds, half the next-fastest method.

Screenshot: Tavus, the Griffin announcement page, 1 October 2026 — 48% believed it was a person, #1 on NVIDIA's independent test, 37% ahead of the next best system at reacting in the moment.

Frame from Tavus's Griffin demo, 1 October 2026, at 0:45 — the call interface, with the person and the model on screen together.
A week of early access is enough for some people to file a verdict, and one of them has four hundred and fifty thousand followers:
"This is the best human-interaction model I've ever seen in action. I got early access to Tavus. I tested it. Judge for yourself."

@svpino, 1 October 2026. Card rendered from the post; read it on X at https://x.com/svpino/status/2105707726831849759.
Hot take, and I will keep it short because the man is doing my job for me. A tester with early access is not a measurement; he is a doorway. Forty-four thousand people will watch that clip and a hundred and eighty-five will press the heart, and not one of them will do what NVIDIA did, which is to put a number next to a human being and publish it. But read his last two sentences again, because the part that should worry you is not the praise. He got access, and he tested it. The thing is already in the hands of the people who will tell you about it.
Now the part that put me back out the door
Here is the safety section of Tavus's own announcement, and I want you to read it in their words rather than mine:
"The same properties that make Human Interaction Models powerful interfaces for natural communications between human and machine allow them to deceive a human into believing it is not AI."
And then, plainly:
"We believe further alignment and safety procedures are required for safe release."
Griffin-Lite — the version that ran this study, the one that fooled half the room — is not available to customers at all. It is open to selected testers only. The company says the wider release waits on safety work, including disclosure features, and that it anticipates releasing "very soon after these safety concerns are addressed."
So the honest summary of 1 October 2026 is this. A company built a machine that passes as a person on a video call. The company then looked at its own product and said: not yet, we have not worked out how to tell people the truth about what they are talking to.
I have covered this beat long enough to know what usually happens next. The technology ships anyway, three competitors have a version within a quarter, and the disclosure features arrive as a settings toggle that defaults to off because it costs conversions. Tavus is doing the rare thing here, and I want that on the record before I say what I think is coming, because the record is going to matter.
What I think is coming, stated as opinion
Nobody who wants this wants it for a tutor that notices confusion, although that is the example the company gives and it is a good one. The market for a face that cannot be distinguished from a person is a market in impersonation: the romance that never goes bad, the grandparent who takes the call, the interview candidate who is not in the room, the support agent who is a person until the refund is denied. That is not a hypothetical. It is the last six weeks of this website's own front page — deepfakes of actors in court, a mining influencer with 80,000 followers who turned out to be a model that Meta removed, an 80-second fake of a party that never happened.
The people who will pay for this are not the ones who need a tutor. And the people who will be hurt by it are, as always, the ones who did not get a memo: the person who takes the call, the family that believes the voice, the broadcaster that airs the clip. Every one of those is a financial, professional and legal wound, and the statute book is not ready. State laws do not stop a fake video at the state line; the NO FAKES Act is still a bill; and the industry's own answer so far is a watermark that anyone can crop.
So here is the curse, and I mean it as a professional prediction rather than a wish. The first company to sell a face that cannot be told from a person without telling anyone will make a great deal of money for about eighteen months, and will then spend ten years in court arguing about a phone call it cannot produce. That is the bill. It always comes.
The other half of the timeline arrived in my feed before I had finished this column. A man with three hundred thousand followers read the same announcement and reached for a different conclusion:
"Remote work is so cooked. If your job can be done on the other side of a screen the AI can do it better."

@EMostaque, 1 October 2026. Card rendered from the post; read it on X at https://x.com/emostaque/status/2105744708299526243.
Hot take, and this one needs care because the man is half right and half right is the dangerous half. He is right about the mechanism: anything that arrives through a screen can be imitated on that screen, and a face was the last thing standing between a call centre and a model that never sleeps, never asks for a raise and never tells the customer it is having a bad day. Where he is wrong is the direction of the damage. The job that goes first is not the worker on camera; it is the trust on the other end of the line. A company that replaces its support desk with a face like this one has not saved money, it has spent its reputation, and it will find that out the day a customer asks to speak to the person they were just talking to and there is nobody there.
I have watched three gold rushes from a bar stool and the pattern does not change. The people who get rich are the ones selling the shovels; the people who get hurt are the ones who believed the pitch about who benefits.
The number I cannot put down
What kills me — and this is the part I have been pacing about — is that the benchmark was not wrong. The paper says today's models systematically miss the nonverbal turn. Tavus went and got the nonverbal turn, which is the hard part, the human part, the thing that makes a conversation a conversation rather than a query.
That is the actual achievement here, and I want to give the builders their due, because I am not a man who hates the work. Anyone who can teach a machine to notice a pause and not step on it has done something I could not do. The research is real, the receipts are independent, and the study design — telling fifty-four people they were talking to another participant — is the only honest way to test the thing they were testing.
And then the same document, four sections later, tells you it has built something that cannot yet be released, because it can lie about what it is.
We used to argue about whether a machine could think. Now the question is whether it can be trusted to tell you what it is, and the people who built it are the ones asking.
Twenty-six people out of fifty-four walked away from a one-minute call believing they had met someone. The only thing standing between that number and your grandmother's phone is a research preview, a request form, and a paragraph under the heading safety, limitations, and responsibility.
I am keeping the receipt.
More from SLOP TV News
- Tavus says 48% of testers mistook its Griffin model for a human
- Bombay High Court blocks AI deepfakes of Samantha Ruth Prabhu
- An 80,000-follower mining influencer was AI, and Meta deleted her
- A White House aide's fake AI clip of his own party backfired
- NO FAKES Act would give creators a federal right to their likeness
- Why state AI deepfake laws may not stop a fake video
Sources
- Tavus, "Introducing Griffin: The First Human Interaction Model", 1 October 2026 — the study, the architecture, the benchmark scores and the safety section quoted above.
- NVIDIA Research, VideoFDB, the full-duplex audio-visual benchmark Griffin was scored on, and the paper behind it.
- The VideoFDB evaluation dataset on Hugging Face, released under access request.
- Tavus's platform and research pages — the models Griffin was built from: Phoenix, Raven and Sparrow.
- The announcement post, 1 October 2026, and SLOP TV's news brief on it.