arXiv Artificial Intelligence

MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

Quick summary

arXiv:2609.20152v1 Announce Type: new Abstract: Generally, most voice agents are cascaded systems, i.e., an ASR model transcribes the caller's audio, a language model reads the transcript and decides what to say and which backend tools to call, and a TTS model speaks the reply. Nearly all of the decision making happens in the language model, but existing evaluations measure it either too broadly or too narrowly. End-to-end voice benchmarks score the full pipeline, so recognition errors and model errors mix into a single number. LLM benchmarks isolate the model but they do not evaluate what mak

Key takeaways

  • arXiv:2609.20152v1 Announce Type: new Abstract: Generally, most voice agents are cascaded systems, i.e., an ASR model transcribes the caller's audio, a language model reads the transcript and decides what to say and which backend tools to call, and a TTS model speaks the reply.
  • Nearly all of the decision making happens in the language model, but existing evaluations measure it either too broadly or too narrowly.
  • End-to-end voice benchmarks score the full pipeline, so recognition errors and model errors mix into a single number.

Why it matters

“MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗