AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis
Quick summary
arXiv:2607.15755v2 Announce Type: replace-cross Abstract: Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods struggle to render authentic human emotions due to limited predefined emotion label spaces (e.g., seven emotion categories), while redundant multimodal tokens in multi-turn dialogue history interfere with context understanding. To address these issues, we propose AuEmoChat, a CSS framework for authentic emotion understanding and rendering. First, we develop AuEmo
Key takeaways
- arXiv:2607.15755v2 Announce Type: replace-cross Abstract: Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions.
- Existing CSS methods struggle to render authentic human emotions due to limited predefined emotion label spaces (e.g., seven emotion categories), while redundant multimodal tokens in multi-turn dialogue history interfere with context understanding.
- To address these issues, we propose AuEmoChat, a CSS framework for authentic emotion understanding and rendering.
Why it matters
“AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments