arXiv Artificial Intelligence

$\tau$-Multilingual: Benchmarking Voice Agents Across Languages

$\tau$-Multilingual: Benchmarking Voice Agents Across Languages

Quick summary

arXiv:2609.35820v1 Announce Type: cross Abstract: English-only benchmarks expose only a narrow slice of voice-agent behavior. We introduce $\tau$-Multilingual, extending $\tau$-Voice to Spanish, Brazilian Portuguese, Hindi, Korean, and Mandarin with native-speaker review and evaluation of generated language and spoken output. Across 4,500 full-duplex calls and five voice configurations, Spanish, Portuguese, and Hindi remain within 3.2 task-completion points of English, but Korean and Mandarin fall by 14.7 and 8.4 points. The failure modes also vary: Korean systems miss more responses, Mandarin

Key takeaways

  • arXiv:2609.35820v1 Announce Type: cross Abstract: English-only benchmarks expose only a narrow slice of voice-agent behavior.
  • We introduce $\tau$-Multilingual, extending $\tau$-Voice to Spanish, Brazilian Portuguese, Hindi, Korean, and Mandarin with native-speaker review and evaluation of generated language and spoken output.
  • Across 4,500 full-duplex calls and five voice configurations, Spanish, Portuguese, and Hindi remain within 3.2 task-completion points of English, but Korean and Mandarin fall by 14.7 and 8.4 points.

Why it matters

“$\tau$-Multilingual: Benchmarking Voice Agents Across Languages” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗