OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses
Quick summary
arXiv:2608.24160v1 Announce Type: new Abstract: Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as "OmniJudges" for evaluation and automatic annotation. How reliably they understand what they score remains unclear, since existing benchmarks and training data tend to overemphasize positive examples and to conflate distinct failure modes, so a judge may score well without recognizing failures while its capability gaps stay hidden. Motivated by this, we introduce D3-Omni, a balanced and decou
Key takeaways
- arXiv:2608.24160v1 Announce Type: new Abstract: Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as "OmniJudges" for evaluation and automatic annotation.
- How reliably they understand what they score remains unclear, since existing benchmarks and training data tend to overemphasize positive examples and to conflate distinct failure modes, so a judge may score well without recognizing failures while its capability gaps stay hidden.
- Motivated by this, we introduce D3-Omni, a balanced and decou
Why it matters
“OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments