Beyond Text: LLM-Based Dimensional Emotion Evaluation in Multimodal Dialogue
Quick summary
arXiv:2609.39072v1 Announce Type: cross Abstract: Emotion recognition in conversation has been widely studied, but applying Large Language Models (LLMs) to continuous dimensional emotion evaluation in multimodal dialogue remains largely unexplored. We propose an LLM-based framework that performs discrete emotion recognition and Valence-Arousal-Dominance (VAD) dimensional evaluation on IEMOCAP, incorporating acoustic cues as natural language descriptions following the SpeechCueLLM approach. We evaluate six models spanning the LLaMA, GPT, and Qwen families under zero-shot prompting, few-shot pro
Key takeaways
- arXiv:2609.39072v1 Announce Type: cross Abstract: Emotion recognition in conversation has been widely studied, but applying Large Language Models (LLMs) to continuous dimensional emotion evaluation in multimodal dialogue remains largely unexplored.
- We propose an LLM-based framework that performs discrete emotion recognition and Valence-Arousal-Dominance (VAD) dimensional evaluation on IEMOCAP, incorporating acoustic cues as natural language descriptions following the SpeechCueLLM approach.
- We evaluate six models spanning the LLaMA, GPT, and Qwen families under zero-shot prompting, few-shot pro
Why it matters
“Beyond Text: LLM-Based Dimensional Emotion Evaluation in Multimodal Dialogue” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments