The missing piece in the voice AI stack: decision models

Voice agents now have excellent models for hearing, thinking and speaking. The part that decides what to do next is still, in most deployments, a silence timer. That’s the gap.

Ask anyone who has shipped a voice agent where it breaks, and you rarely hear “the transcription was wrong.” You hear: it cut me off while I was reading my card number. It sat in silence after a simple question. It stopped mid-sentence because I said “mm-hm.” It kept talking after I said goodbye.

None of those are failures of hearing, reasoning or speaking. They’re failures of deciding. And deciding is the one job in the voice AI stack that still has no model of its own.

The voice AI stack, as it’s usually drawn

A production voice agent is a pipeline: audio in over WebRTC or a phone line, STT, an LLM, TTS, audio out. Every layer has dedicated models, leaderboards and vendors. Daily’s voice agent primer walks through a typical voice-to-voice round trip of about 1.3 seconds and sets 1,500 ms as the target.

Now look at what sits between the caller’s audio and the LLM: a voice activity detector and a silence threshold. The same primer’s reference setup stops listening after VAD_STOP_SECS = 0.8. OpenAI’s Realtime API offers the same pattern as server_vad, which “uses periods of silence to automatically chunk the audio,” tuned with silence_duration_ms.

That threshold is a timer. It decides when the agent speaks, which makes it one of the most consequential parts of the stack. It’s also the least intelligent.

Diagram: today a VAD and 800 ms silence timer sits in front of STT, LLM and TTS and decides from silence alone; with a decision model, caller audio goes to a model that chooses respond, wait, backchannel, yield, call a tool or end the call

Humans don’t wait for silence

Natural conversation averages about 230 ms between turns across ten languages, as Kyutai’s Moshi paper notes. That’s too fast to start planning a reply only after the other person stops. A 2026 EEG study found that listeners “use prosodic context to predict upcoming turn-ends, aiding in planning their responses.” They predict where the turn will end and get ready while the other person is still talking.

That prosody is exactly what a transcript throws away. A 2026 study of end-of-turn detection found that “prosodic features have the strongest class separability, while text representations overlap substantially,” and that adding text “increases premature detections without improving performance.” Its best acoustic and prosodic setup reached an F1 of 0.93 with 7.8% false alarms at 400 ms median latency.

A silence timer uses one cue: the absence of sound. A transcript-based detector uses another: the words. People use everything.

End of turn is one decision. A call has many.

Most engineering effort has gone into one question: is the user done? That work has moved fast. LiveKit’s first turn detector ran a small language model on transcripts. Daily’s open-source Smart Turn v3 runs on raw audio with 8M parameters in 12 ms on a CPU. Krisp ships a 6.1M-parameter audio-only model. Deepgram’s Flux folds turn detection into speech recognition. In June 2026, LiveKit’s audio-native Turn Detector v1.0 reported a 9.9% false-cutoff rate at a 300 ms latency budget on its own benchmark, against 12.9% for Flux and 27.7% for ultraVAD.

Vendor benchmarks are vendor benchmarks, but the direction is right: end-of-turn detection is moving off timers and onto audio. The catch is that end of turn is only one of the decisions a live call needs:

  • The caller pauses mid-thought (“my account number is… uh…”). The right call is to wait. A timer replies after 800 ms and cuts them off.
  • The caller finishes a clear question. Respond now. A timer waits out the full threshold anyway.
  • The caller says “mm-hm” while the agent talks. Keep talking. A timer treats it as an interruption and stops.
  • The caller says “wait, no, that’s wrong.” Yield, fast. A timer reacts exactly as it did to “mm-hm.”
  • The caller has just said their order number. Start the lookup. A timer waits for end of turn, then for the LLM.
  • The caller says “great, that’s everything.” Wrap up and end the call. A timer asks “anything else?”
  • The caller turns to talk to someone in the room. Stay silent. A timer responds to anything with energy.

Today each of these gets its own fix, usually tuned in isolation. Krisp describes the state of the art bluntly: “Most current systems use crude heuristics: VAD on any sound (which fires on every ‘uh-huh’), fixed timing thresholds, minimum word counts, or stop-word lists.” An ICLR 2025 study found that spoken dialogue systems “sometimes do not understand when to speak up, can interrupt too aggressively and rarely backchannel.”

The deeper problem is that these decisions aren’t independent. Whether “mm-hm” is a backchannel or the start of an objection depends on what the agent was just saying. Whether a pause means “I’m done” or “I’m finding my card” depends on the question that was asked. Split the decisions across separate rules, and sooner or later the rules disagree. We wrote more about the full set of options in Ask what the moment needs.

Why the LLM, STT or a full-duplex model can’t absorb it

There are three obvious places to put this logic. Each is a poor fit.

The LLM

The LLM is the smartest part of the stack, so it’s tempting to let it decide everything. But it sees text, after transcription, after an endpoint has already fired. In the Daily primer’s round trip, LLM time to first byte alone is 650 ms, three times the human gap between turns. A model that runs once per turn can’t also decide when a turn has happened. Being invoked is the decision.

STT

Flux shows that merging end-of-turn into speech recognition works: Deepgram reports 200 to 600 ms lower response latency and about 30% fewer false interruptions than pipeline approaches. But a recognizer’s job is to emit words. It doesn’t know what the agent just said, which tools exist or whether the request is resolved, and those are exactly the inputs the other decisions need.

Full-duplex speech-to-speech models

Full-duplex models are the closest thing to the human system. Kyutai’s Moshi listens and speaks at once, with 200 ms latency in practice. But in these models, when to speak is tangled up with what to say, which makes it hard to control, threshold or audit.

The evidence on tool use is sobering. The Daily primer notes that “speech-to-speech models do not follow instructions or call tools as reliably as text-mode LLMs.” On Full-Duplex-Bench-v3, a 2026 benchmark of tool use under realistic disfluency, the best system (GPT-Realtime) passes 60% of tasks on the first try, and self-correction remains one of the most consistent failure modes. A recent preprint finds that full-duplex models take the floor when they’re addressed or when there’s silence, not when the content calls for it. Even Sesame says its speech model handles the text and speech of a conversation, but not its structure: turn-taking, pauses and pacing.

That’s why most enterprises still run cascaded stacks. They want to pick their LLM, prompt it, ground it, call their own tools and read the transcript afterwards. What they’re missing isn’t a better LLM. It’s the control layer.

What a decision model is

A decision model is a small, streaming model that listens to the conversation as audio and continuously outputs a probability for each thing the agent could do next. Five properties define it:

  1. Audio-native. It reads the waveform on both sides of the call, the caller and the agent’s own voice. That keeps intonation, hesitation, breath, overlap and pace, the cues a transcript throws away.
  2. Continuous. It decides every few tens of milliseconds, not once per turn. Smart Turn’s 12 ms CPU inference and Krisp’s 6.1M parameters show this class of model can be cheap enough to run on every frame of every call.
  3. One set of actions. Respond, wait, backchannel, yield, call a tool, end the call. One model with one view of the conversation, instead of six heuristics that disagree.
  4. Calibrated. It outputs probabilities, not yes or no. Speed versus cutoffs is a curve, not a point, which is why LiveKit reports false cutoffs at a fixed latency budget. A bank might want almost no cutoffs. A drive-through wants speed. Each sets its own threshold per action, without retraining.
  5. Separate from the LLM. It sits beside the pipeline, not inside it. Swap your LLM, prompts, voice or tools without touching turn-taking.

If you come from networking, think control plane and data plane. STT, the LLM and TTS are the data plane: they move content. The decision model is the control plane: it decides when content moves.

There’s good evidence that learned decisions beat timers with real users. In a 2025 study with 39 participants, Skantze and Irfan replaced a silence-based baseline with general turn-taking models (Voice Activity Projection plus TurnGPT). Median response time fell from 2.7 s to 1.5 s, interruptions fell from 16.6% to 6.9%, and 30 of 39 participants preferred the learned system. Faster and more polite at the same time is exactly what a fixed threshold can’t give you, because with a timer those two goals trade off directly.

How to evaluate a decision model

A single accuracy number hides the tradeoff that matters. Ask for curves and per-decision numbers:

  • False cutoffs at a fixed latency budget, or latency at a fixed cutoff rate. LiveKit’s 300 ms comparison is a good template.
  • Backchannel robustness: the share of caller backchannels that wrongly stop the agent.
  • Barge-in latency: time from a real interruption to the agent going quiet.
  • Precision and recall for actions, like tool calls and hang-ups, not just turn shifts.
  • Behavior under disfluency: self-corrections, restarts and long hesitations, where Full-Duplex-Bench-v3 shows current systems fail most.

Full-Duplex-Bench covers pause handling, backchanneling, smooth turn-taking and user interruption, and it’s a good place to start. The field needs more shared benchmarks: a 2025 survey found that 72% of the turn-taking papers it reviewed don’t compare their methods with previous work.

What a decision model unlocks

Once a model is deciding continuously from audio, some things that used to be impossible get cheap. The biggest is acting before the turn ends. If the decision model can tell mid-sentence that the caller has just finished saying their order number, the lookup can start while they’re still talking. We cover that in Speculative tool calling: the next big thing in voice AI.

The voice stack has learned to hear, think and speak. What it hasn’t learned is when. That’s the piece Tazo fills: an audio-native decision model that hears the conversation and decides what happens next, no transcript in between.

Your stack already knows what to say. A decision model knows when.

References

  1. Kwindla Hultman Kramer and contributors, Daily. Voice AI & Voice Agents: An Illustrated Primer. 2025, updated 2026.
  2. OpenAI. Realtime API: Voice activity detection. Developer docs.
  3. Alexandre Défossez et al., Kyutai. Moshi: a speech-text foundation model for real-time dialogue. 2024.
  4. Yujie Ji, Qingyi Song, Xiaoming Jiang, Yingying Tan, Wenshuo Chang and Xiaolin Zhou. Temporal expectation in turn-taking: prosodic cues modulate behavioral and neural signatures of turn-end prediction in conversation. Brain and Language, 2026.
  5. Rini Sharon, Kadri Hacioglu, Andreas Stolcke et al. Less can be More: What Aspects of Speech Drive End-of-Turn Detection. 2026.
  6. LiveKit. Using a transformer to improve end of turn detection. 2024.
  7. Daily. Announcing Smart Turn v3, with CPU inference in just 12ms. 2025.
  8. Krisp. Audio-only, 6M weights Turn-Taking model for Voice AI Agents. 2025.
  9. Deepgram. Introducing Flux: Conversational Speech Recognition. 2025.
  10. LiveKit. Solving end-of-turn detection: LiveKit Turn Detector v1.0. 2026.
  11. Krisp. A New Approach to Turn-Taking in Voice AI: Turn Prediction v3 and Interruption Prediction v1. 2026.
  12. Siddhant Arora, Zhiyun Lu, Chung-Cheng Chiu, Ruoming Pang and Shinji Watanabe. Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics. ICLR 2025.
  13. Guan-Ting Lin, Chen Chen, Zhehuai Chen and Hung-yi Lee. Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency. 2026.
  14. Linkai Peng, Baorian Nuchged, Kaiqi Fu and Yuyang Yao. Full-Duplex Speech Models Take the Floor When Asked, Not When Needed. 2026 (preprint).
  15. Sesame. Crossing the uncanny valley of conversational voice. 2025.
  16. Gabriel Skantze and Bahar Irfan. Applying General Turn-taking Models to Conversational Human-Robot Interaction. HRI 2025.
  17. Guan-Ting Lin et al. Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities. ASRU 2025.
  18. Galo Castillo-López, Gaël de Chalendar and Nasredine Semmar. A Survey of Recent Advances on Turn-taking Modeling in Spoken Dialogue Systems. IWSDS 2025.