Tazo vs Jev: why voice AI needs an audio decision model
An audio decision model listens to a conversation and decides what happens next. Jev decides from text, and AudioJev answers questions about short clips. Tazo is the audio decision model built for live voice AI.
Short answer: an audio decision model listens to audio and picks the next action from a set of options, with a calibrated probability attached. Tazo is an audio decision model built for live voice AI: it hears the call itself and decides what happens next, with no transcript in between. Jev, from TypeSafe AI, makes similar decisions from text. AudioJev, from mocomoco, is an early attempt at audio that answers simple questions about short clips.
What is an audio decision model?
An audio decision model takes audio in and returns a decision out: respond or wait, keep talking or yield, escalate or carry on. It doesn’t generate text. It picks from options you define and tells you how confident it is.
That makes it the missing control layer in a voice agent. STT, the LLM and TTS decide what to say. An audio decision model decides when and what happens next. We cover why that layer matters in the missing piece in the voice AI stack.
What is Jev?
Jev is TypeSafe AI’s first “System One” model, released in early access on September 15, 2026. TypeSafe describes it as “a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.” It returns calibrated probabilities instead of generated text, with an end-to-end response time of 70 to 500 ms.
That’s the right shape for a decision, and it’s the same shape Tazo uses. The difference is the input.
Why Jev isn’t an audio decision model
Jev reads text. TypeSafe lists its input as “unstructured data (e.g. text)” and program state. In a voice agent, that means STT has to run first, and two things go wrong.
- It’s late. The transcript lands after the words are spoken, often after the turn has ended, and Jev’s own 70 to 500 ms comes on top. People reply to each other in about 230 ms.
- It can’t hear how things are said. A transcript drops pitch, pace, hesitation and tone, the cues that tell you whether a caller is done or still thinking. A 2026 study of end-of-turn detection found prosody separates the classes best, and that adding text “increases premature detections without improving performance.”
Jev can decide what to do with a transcript. It can’t hear the moment the decision is about.
AudioJev: an early attempt at an audio decision model
On September 21, 2026, mocomoco released AudioJev, a “Jev-style” audio decision model with open weights. It “answers questions about short audio clips with Yes/No, a choice, or a number,” across speech, music, animal sounds and ambient noise. The launch example asks whether a dog is barking in a clip.
It’s a welcome step, and it can run in the browser. But by its own model card, it’s built for simple Q&A about a recorded clip, not for decisions inside a live business call:
- Clips, not calls. Input is a 0.25 to 30 second clip. The card notes that VAD, speaker separation and streaming aren’t in the current model yet.
- Not calibrated. “Option scores are not calibrated probabilities,” so you can’t trust a threshold per action.
- Often still transcript-based. The public Cloud demo “uses ASR for language tasks that need a transcript.”
- Narrow evaluation. Its one published result is 88.75% on 240 fixed items about reservation or order intent. Accuracy on new questions “remains under evaluation.”
- Japanese-first. Its question encoders are Japanese ModernBERT models.
Tazo: the audio decision model built for voice AI
Tazo keeps the best part of the Jev idea, typed decisions with calibrated probabilities, and builds it for the live call:
- Audio-native. Tazo reads raw call audio, not a transcript, so it hears pitch, pauses and tone.
- Made for the live call. It decides what the moment needs while the conversation is happening: respond, wait, backchannel, yield or escalate.
- Calibrated. Every answer comes with a probability you can act on, for example escalating only above 0.9.
- Your options, one model. You define the options in the request, so turn-taking, voicemail, escalation and tool timing run on one API instead of one model per job.
- Next to your stack. It runs beside your existing STT, LLM and TTS, and it’s faster than an LLM call.
For the full set of decisions a call needs, see Ask what the moment needs.
Tazo vs Jev vs AudioJev at a glance
| Tazo | Jev | AudioJev | |
|---|---|---|---|
| Input | Live call audio | Text | 0.25 to 30 s clips |
| Needs a transcript | No | Yes | Cloud: yes |
| Calibrated | Yes | Yes | No |
| Built for | Voice agents | Software workflows | Q&A about a clip |
FAQ
Is Jev an audio decision model?
No. TypeSafe’s Jev takes text and structured state. To use it in a voice agent, you have to transcribe the audio first, which adds delay and loses tone and timing.
What is AudioJev?
An open-weight, Jev-style model from mocomoco that answers Yes/No, choice or number questions about 0.25 to 30 second audio clips. Its scores aren’t calibrated, and streaming isn’t supported yet.
How is an audio decision model different from an LLM?
An LLM generates text, one token at a time. An audio decision model picks from a fixed set of options in one step and returns a probability for each, so it’s faster and its output plugs straight into code.
What’s the best audio decision model for voice AI agents?
For live calls, you want a model that’s audio-native, works during the conversation and returns calibrated probabilities. That’s what Tazo is built for.
Jev decides from text. AudioJev answers questions about clips. Tazo decides what the moment needs, straight from the call.
References
- TypeSafe AI. Introducing System One Models & Jev. 2026.
- mocomoco. Released AudioDecisionModel, a Jev-style Audio Decision AI. 2026.
- mocomoco. AudioJev-AudioDecisionModel model card. Hugging Face, 2026.
- Alexandre Défossez et al., Kyutai. Moshi: a speech-text foundation model for real-time dialogue. 2024.
- Rini Sharon, Kadri Hacioglu, Andreas Stolcke et al. Less can be More: What Aspects of Speech Drive End-of-Turn Detection. 2026.