arXiv Artificial Intelligence

RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue

RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue

Quick summary

arXiv:2609.16614v1 Announce Type: cross Abstract: Speech-to-speech dialogue models increasingly support persona control, yet existing spoken role-playing benchmarks remain largely character-centric and short-horizon. This leaves open whether spoken dialogue models can sustain diverse roles over extended interactions, especially beyond predefined fictional characters. We introduce RoleBreak, an open benchmark for long-horizon role-playing robustness in spoken dialogue. RoleBreak contains 310 character-based and user-centered roles, 6,688 human-verified dialogue turns, and 11,743 fine-grained ev

Key takeaways

  • arXiv:2609.16614v1 Announce Type: cross Abstract: Speech-to-speech dialogue models increasingly support persona control, yet existing spoken role-playing benchmarks remain largely character-centric and short-horizon.
  • This leaves open whether spoken dialogue models can sustain diverse roles over extended interactions, especially beyond predefined fictional characters.
  • We introduce RoleBreak, an open benchmark for long-horizon role-playing robustness in spoken dialogue.

Why it matters

“RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗