SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
Quick summary
arXiv:2603.16859v3 Announce Type: replace Abstract: Evaluating omni-modal large language models (OLMs) in multi-party dialogue requires more than answer correctness on pre-segmented inputs. We introduce SocialOmni, an offline diagnostic benchmark that separates three turn-level decisions: identifying who is speaking, deciding when a designated participant should enter at an annotated query time, and determining how that participant should continue the dialogue. SocialOmni contains 2,000 perception items and a quality-controlled core split of 200 interaction-generation items, including naturall
Key takeaways
- arXiv:2603.16859v3 Announce Type: replace Abstract: Evaluating omni-modal large language models (OLMs) in multi-party dialogue requires more than answer correctness on pre-segmented inputs.
- We introduce SocialOmni, an offline diagnostic benchmark that separates three turn-level decisions: identifying who is speaking, deciding when a designated participant should enter at an annotated query time, and determining how that participant should continue the dialogue.
- SocialOmni contains 2,000 perception items and a quality-controlled core split of 200 interaction-generation items, including naturall
Why it matters
“SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments