Speculative tool calling: the next big thing in voice AI

Voice agents wait for the caller to stop, then decide to call a tool, then wait for the tool. Most of that dead air is avoidable. The trick is an old one from computer architecture: guess early, check later, throw away what’s wrong.

“Let me look that up for you.”

Every voice agent builder knows this line. It’s filler, there to cover the silence while the agent checks a CRM, an order system, a calendar or a knowledge base. It exists because tool calls are where voice agent latency goes to die. Speculative tool calling is a way to start those calls while the caller is still talking, without changing what the agent does.

Tool calls are the latency nobody budgets for

The industry’s own latency budgets quietly admit the problem. Twilio’s guide to voice agent latency sets a mouth-to-ear target of 1,115 ms, then explicitly leaves out “the occasional spikes that come from tool calls or complex reasoning,” which it suggests covering “with techniques like interstitial fillers.” LiveKit’s latency guide notes that tool calls run before the reply is generated and “can add large, variable latency.” A May 2026 paper from Berkeley puts it plainly: voice apps need under a second of latency to feel seamless, and agentic tool calling “can add several seconds or more of latency, which is prohibitive.”

Meanwhile, people reply to each other in about 230 ms on average. Voice agents are now good at almost everything except doing things quickly.

Where the dead air comes from

Here’s the standard sequence when a caller asks for something that needs a tool:

  1. Wait for the caller to stop talking, then wait some more to be sure.
  2. Send the transcript to the LLM, which emits a tool call.
  3. Run the tool.
  4. Send the result back to the LLM, which starts writing a reply.
  5. Synthesize the first audio.

Every step waits for the one before. Now look at when the information arrives. A caller says: “Hi, I’m calling about an order, the number is 4-8-2-1-7, and it still hasn’t shown up.” Once the last digit lands, everything needed for get_order_status("48217") is known. The rest of the sentence is context for the reply, not the lookup. Yet the agent spends the rest of the sentence, the endpoint wait and an LLM round trip doing nothing about it.

This isn’t rare. In SHANKS, a 2025 system that lets a spoken language model reason while it listens, the model completed 56.9% of tool calls before the user finished their turn. More than half the time, the agent could have known what to look up before the caller stopped talking.

Timeline: sequentially, the agent waits on a silence timer, the LLM, an 800 ms lookup and the reply, giving 2.2 seconds of dead air; speculatively, the lookup starts when the order number is heard, so first audio arrives 1.2 seconds sooner

In this example, most of the gain comes from the lookup running during speech. A faster end-of-turn decision accounts for the rest.

An old idea from computer architecture

Processors hit this problem decades ago. A pipeline that waits for every branch to resolve spends most of its time stalled. The fix was branch prediction: guess the outcome and keep working before it’s known. Right guess, the work is already done. Wrong guess, throw it away and carry on.

LLMs rediscovered the idea as speculative decoding. A small draft model proposes tokens and the large model checks them in one parallel pass. The large model still has the final say, so the output is identical to normal decoding, just faster. EAGLE, for example, reports a 2.7x to 3.5x latency speedup on LLaMA2-Chat 70B “while maintaining the distribution of the generated text.”

Agents are next:

  • Speculative Actions predicts an agent’s next action with a fast model while the slow one is still thinking. It reaches up to 55% next-action accuracy and up to 20% lower latency, without changing outcomes.
  • Interactive Speculative Planning and Dynamic Speculative Agent Planning apply the same draft-and-check pattern to multi-step plans.
  • LLMCompiler runs independent function calls in parallel for up to 3.7x lower latency, and AsyncLM keeps generating while calls run, cutting task latency 1.6x to 5.4x.
  • The Berkeley paper reports 1.3 to 1.7x speedups on real-time cloud APIs from speculative tool calling.
  • A July 2026 paper trains a speculator alongside the agent and raises next-tool-call accuracy (Hit@1) from 44.1 to 61.2.

So why does this matter most in voice? Because the user is talking. A text agent gets its input all at once, and the only window for speculation is its own thinking time. A voice agent gets its input slowly, one word at a time, over seconds. The turn itself is the speculation window. Acting on input before it’s complete is an old idea in dialogue research. It has never had this much latency to win back.

How speculative tool calling works

Speculative tool calling in four steps: predict the next tool call while the caller talks, gate on safe read-only tools above a confidence bar, fire it in the background and cache the result, then at end of turn commit if the LLM makes the same call or discard it

  1. Predict. While the caller speaks, a model watching the live audio proposes likely tool calls, with arguments and a confidence.
  2. Gate. A proposal fires only if the tool is safe to speculate (more below) and the confidence clears a threshold set for that tool.
  3. Fire. The call runs in the background. The result is cached, keyed by the tool name and normalized arguments.
  4. Commit or discard. At end of turn, the LLM runs exactly as it does today and emits its real tool call. If it matches a speculative one, the cached result comes back instantly. If not, the call runs as normal and the guesses are thrown away.

The key property is that this is lossless, in the same sense as speculative decoding. The LLM still makes every real decision, and a speculative result is only used when it exactly matches the call the LLM would have made anyway. The agent behaves identically to one without speculation. Only the latency changes. A wrong guess costs a little compute, never correctness.

# While the caller is speaking: fire safe, confident guesses.
def on_partial(state):
    for guess in predictor.propose(state):  # tool, args, confidence
        spec = registry[guess.tool]
        if spec.speculative_ok and guess.confidence >= spec.threshold:
            key = (guess.tool, normalize(guess.args))
            if key not in inflight:
                inflight[key] = run_async(guess.tool, guess.args)

# At end of turn: the LLM emits its real call, as it does today.
async def execute(call):
    key = (call.tool, normalize(call.args))
    if key in inflight:
        result = await inflight.pop(key)  # hit: done, or nearly
    else:
        result = await run(call.tool, call.args)  # miss: unchanged
    cancel_all(inflight)  # drop leftover guesses
    return result

How it fits with speculative replies

Parts of the industry already speculate on the reply. Deepgram’s Flux sends an EagerEndOfTurn event when it’s “moderately confident the user has finished speaking,” so you can start the LLM early, and a TurnResumed event telling you to “cancel the in-progress response” if the caller keeps going. LiveKit’s preemptive generation “speculatively starts an LLM response before the user’s end of turn is confirmed,” and holds TTS until the turn is confirmed.

Deepgram is candid about the price: trimming “that last 100-200ms of end-to-end latency at the cost of 50-70% more LLM calls.” That’s a good trade, and it combines naturally with speculative tool calls. But reply speculation starts at the very end of the turn and saves the LLM’s time to first token. Tool speculation can start mid-turn and saves the tool’s whole run time, which is usually the bigger number.

The math of a speculative tool call

For one tool, the expected saving per turn is roughly:

saving ≈ P(hit) × min(T_tool, T_lead)
waste  ≈ (1 − P(hit)) / P(hit)    extra calls per useful one

T_tool is how long the tool takes. T_lead is how much earlier the speculative call fired than the real one would have. Take an 800 ms CRM lookup fired 1.2 s before the LLM would have issued it. On a hit, the full 800 ms disappears. At a 60% hit rate, that’s about 480 ms off the average turn that uses this tool, at the cost of about 0.67 wasted lookups per useful one.

Two things make this better in practice than the average suggests. First, the saving is largest where it hurts most: the slow, long-tail lookups gain the most from an early start, as long as the caller gives you enough lead time. Second, the first wins come from the most frequent, most predictable calls: order status after an order number, account lookup after a phone number, availability after a date. Salesforce’s VoiceAgentRAG shows how far prediction alone can go for retrieval: a background agent that pre-fetches likely follow-up topics reached a 75% cache hit rate, and hits took 0.35 ms instead of 110 ms.

Which tool calls are safe to speculate

Speculation is only free if a wrong guess leaves no trace. HTTP’s vocabulary helps: a safe call is read-only, and an idempotent call has the same effect whether it runs once or many times. That gives a simple policy:

Tool Type Speculate?
get_order_status Safe read Yes
check_availability Safe read Yes
lookup_account Sensitive read Only after identity check
update_address Idempotent write Prepare, don’t commit
book_appointment Unsafe write Never
charge_card Unsafe write Never

“Prepare, don’t commit” is useful. You can’t speculatively book an appointment, but you can fetch open slots, validate the address or price the order early, so that when the caller confirms, the write is the only thing left. Where a write might be retried, use idempotency keys so its side effects happen only once.

The Spectre lesson for speculative tool calls

For years, CPU speculation was treated as invisible: wrong guesses were rolled back, so they didn’t matter. Spectre proved otherwise: speculative work can leave measurable side effects even when its results are thrown away.

The same goes for agents. Ghost Tool Calls (2026) shows that speculative tool calls “leak inferred user intent to external services before the agent commits to the branch,” and that no cleanup at commit time can unsend what an outside observer has already seen. A read-only call is still a message. A speculative lookup against a third-party API tells that third party what you think your caller is about to ask.

In practice, that means three rules:

  • Speculate against your own systems first. Be much more careful with third-party APIs, where the request itself is information.
  • A speculative call never outranks the real one. It must pass the same authorization the committed call would. If the caller hasn’t verified their identity, the account lookup doesn’t fire, however confident the prediction.
  • Log speculative calls separately, with the confidence and the evidence that triggered them, so an audit shows exactly what fired and why.

The hard part: speech is messy

Callers correct themselves. “4-8-2-1… sorry, 4-8-1-2.” They restart, trail off and change their minds. On Full-Duplex-Bench-v3, a benchmark of tool use under real-world disfluency, self-correction was among the most consistent failure modes across every system tested. A predictor acting on raw partial transcripts will fire on the wrong order number.

Because commit requires an exact match, a wrong guess can’t leak into the answer. But it wastes a call, and a stale value can crowd out the right one. The fix is a predictor that knows when an argument is stable, not just present. That’s largely an acoustic question. A digit string read with falling, final intonation is done. One followed by a hesitation or a rising “uh” probably isn’t. NVIDIA’s voice agent examples already do a simple version of this, acting on interim transcripts only once the recognizer marks them stable.

Why speculative tool calling belongs in the decision layer

The predictor has to run continuously on the live stream, track whether arguments are settled, know which tools are safe and give calibrated confidences, all within tens of milliseconds. The LLM can’t do it, because it only runs when invoked. STT can’t do it, because it knows nothing about your tools.

That’s the job description of a decision model: a small model that listens to the audio and decides what happens next. “Start this lookup now” is one more decision next to “respond” and “wait.” With Tazo, it’s one more question in the request, such as “has the caller given a complete order number?”, answered from the audio with a calibrated probability you can set a threshold on.

Research on spoken language models is arriving at the same place from the other direction. Besides SHANKS, Meta’s Can Speech LLMs Think while Listening? starts reasoning before the user finishes and reports a 70% latency reduction without losing accuracy. The direction is clear: stop waiting for the turn to end before doing the work.

Why speculative tool calling matters now

Three things have changed at once. Conversational decisions have moved onto fast, cheap, audio-native models, so something finally runs continuously during the turn that can carry a prediction. Voice agents have moved from answering FAQs to transactional work that’s mostly tool calls, which makes tool latency the biggest remaining source of dead air. And the research has caught up, with a wave of 2025 and 2026 papers on speculative actions, speculative tool calls and thinking while listening.

If you want to try it, start here:

  1. Tag every tool as safe, idempotent or unsafe, and as sensitive or not.
  2. Normalize arguments, so equivalent calls produce the same cache key.
  3. Start with your slowest, safest, most frequent tool.
  4. Measure hit rate, wasted calls, and p50 and p95 time from end of turn to first audio, with and without speculation.
  5. Set a confidence threshold per tool from those numbers, and revisit it as the predictor improves.

“Let me look that up for you” should be a choice, not a necessity.

References

  1. Phil Bredeson, Twilio. Core Latency in AI Voice Agents. 2025.
  2. LiveKit. Understand and Improve Voice Agent Latency. 2026.
  3. Coleman Hooper, Minwoo Kang, Suhong Moon et al. Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling. 2026.
  4. Alexandre Défossez et al., Kyutai. Moshi: a speech-text foundation model for real-time dialogue. 2024.
  5. Cheng-Han Chiang, Xiaofei Wang, Linjie Li et al. SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models. 2025.
  6. Yuhui Li, Fangyun Wei, Chao Zhang and Hongyang Zhang. EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty. ICML 2024.
  7. Naimeng Ye, Arnav Ahuja, Georgios Liargkovas et al. Speculative Actions: A Lossless Framework for Faster Agentic Systems. 2025.
  8. Wenyue Hua, Mengting Wan, Shashank Vadrevu et al. Interactive Speculative Planning: Enhance Agent Efficiency through Co-design of System and User Interface. ICLR 2025.
  9. Yilin Guan, Qingfeng Lan, Sun Fei et al. Dynamic Speculative Agent Planning. ICLR 2026.
  10. Sehoon Kim, Suhong Moon, Ryan Tabrizi et al. An LLM Compiler for Parallel Function Calling. ICML 2024.
  11. In Gim, Seung-seob Lee and Lin Zhong. Asynchronous LLM Function Calling. 2024.
  12. Jiabao Ji, Yujian Liu, Li An et al. Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL. 2026.
  13. Deepgram. Flux: Eager End of Turn. Developer docs.
  14. LiveKit. Agents: audio and preemptive generation. Developer docs.
  15. Jielin Qiu, Jianguo Zhang, Zixiang Chen et al., Salesforce. VoiceAgentRAG: Solving the RAG Latency Bottleneck in Real-Time Voice Agents Using Dual-Agent Architectures. 2026.
  16. Bardia Mohammadi, Lars Klein, Akhil Arora and Laurent Bindschaedler. Ghost Tool Calls: Issue-Time Privacy for Speculative Agent Tools. 2026.
  17. Guan-Ting Lin, Chen Chen, Zhehuai Chen and Hung-yi Lee. Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency. 2026.
  18. NVIDIA. Speculative Speech Processing. voice-agent-examples docs.
  19. Yi-Jen Shih, Desh Raj, Chunyang Wu et al. Can Speech LLMs Think while Listening?. 2025.