RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue
Quick summary
arXiv:2609.16614v1 Announce Type: cross Abstract: Speech-to-speech dialogue models increasingly support persona control, yet existing spoken role-playing benchmarks remain largely character-centric and short-horizon. This leaves open whether spoken dialogue models can sustain diverse roles over extended interactions, especially beyond predefined fictional characters. We introduce RoleBreak, an open benchmark for long-horizon role-playing robustness in spoken dialogue. RoleBreak contains 310 character-based and user-centered roles, 6,688 human-verified dialogue turns, and 11,743 fine-grained ev
Key takeaways
- arXiv:2609.16614v1 Announce Type: cross Abstract: Speech-to-speech dialogue models increasingly support persona control, yet existing spoken role-playing benchmarks remain largely character-centric and short-horizon.
- This leaves open whether spoken dialogue models can sustain diverse roles over extended interactions, especially beyond predefined fictional characters.
- We introduce RoleBreak, an open benchmark for long-horizon role-playing robustness in spoken dialogue.
Why it matters
“RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments