arXiv Artificial Intelligence

Relevant Evidence Decoding for Audio-Visual Hallucination Mitigation

Relevant Evidence Decoding for Audio-Visual Hallucination Mitigation

Quick summary

arXiv:2610.02976v1 Announce Type: new Abstract: Audio-Visual Large Language Models (AV-LLMs) remain prone to cross-modal hallucinations, where one modality incorrectly affects predictions about another. Although contrastive decoding reduces hallucinations in vision-language models, its direct extension to AV-LLMs overlooks a key challenge: different questions require different perceptual evidence, including audio, video, or their interaction. Notably, we observe that joint audio-visual inference can weaken the prediction even when a model can recover the correct answer from a single informativ

Key takeaways

  • arXiv:2610.02976v1 Announce Type: new Abstract: Audio-Visual Large Language Models (AV-LLMs) remain prone to cross-modal hallucinations, where one modality incorrectly affects predictions about another.
  • Although contrastive decoding reduces hallucinations in vision-language models, its direct extension to AV-LLMs overlooks a key challenge: different questions require different perceptual evidence, including audio, video, or their interaction.
  • Notably, we observe that joint audio-visual inference can weaken the prediction even when a model can recover the correct answer from a single informativ

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗