Ask what the moment needs
Voice AI agents cut people off, stop at every “mm-hmm” and miss obvious cues. Here’s why that happens, and what changes when an agent decides what each moment needs.
Every few hundred milliseconds, a voice AI agent has to make a call. Is the user done, or still thinking? Was that “mm-hmm” an interruption, or a nudge to keep going? Is that second voice the user, or the TV?
In conversation, people leave gaps of only about 200 ms between turns, because we hear the end of a turn coming and plan our reply while still listening. Most voice agents don’t ask these questions at all. They ask one simpler question instead: has the user stopped talking?

The status quo: VAD and end-of-turn detection
A VAD silence timer answers that question, sometimes backed by a transcript-based end-of-turn model or semantic_vad. Every other decision gets its own component: one for voicemail, one for frustration, one for background voices. Each is built for one narrow job. When your callers, domain or audio conditions shift, or you need a decision none of them cover, you end up retuning, fine-tuning or bolting on another component.
Where turn detection breaks
In AssemblyAI’s 2026 survey, 47.5% of builders said their agent interrupts users mid-sentence, and 55% named “having to repeat themselves” as the biggest end-user frustration. On the ECHO benchmark, a Chinese speech test, three of the four systems tested kept talking through fewer than 13% of backchannels. Full-duplex models (which listen while they talk) have the opposite problem: they take the floor when asked, not when needed, and rarely speak up to correct a false claim.
Part of the problem is the transcript. Whether someone is finished or mid-thought often shows up in pitch and pauses, not words. One ablation study found that adding text to acoustic and prosodic features increased premature end-of-turn detections without improving performance.
A better question than “are they done?”
Swap the yes/no for a choice: respond, wait, backchannel, yield, ignore, or escalate. Research is moving this way too, with four-state models like Easy Turn and three-way ones like MM-When2Speak. That’s a pick from a known set, not open-ended text, so it shouldn’t have to wait for an LLM to generate a reply.

Where Tazo fits: one audio-native decision model
Tazo is an audio-native decision model. You define the options in the request, and Tazo returns a choice (or a yes/no) with calibrated probabilities. Because the options are an input, not baked in at training, one general-purpose model handles every decision on the call (turn-taking, backchannels, voicemail, escalation), and a new decision means a new option list, not a new fine-tune.
There’s no transcript in between, it’s faster than an LLM call, and it runs as a standalone API next to your existing STT, LLM and TTS. Calibrated probabilities let you set a threshold per action, for example escalating only above 0.9.

Your LLM decides what to say. Tazo decides what the moment needs.