DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs
Quick summary
arXiv:2609.28570v1 Announce Type: new Abstract: Reinforcement learning (RL) is widely used to sharpen reasoning in multimodal large language models (MLLMs), yet its effect on hallucination is uneven. We trace this to two weak points in the \emph{correction chain} from reward to parameter update. At the rollout level, hard queries---those with high semantic entropy---frequently produce unanimously wrong sample groups, collapsing the group-relative advantage to zero exactly where hallucination risk is highest. At the optimization level, confident-but-wrong tokens are gradient-invisible: a catego
Key takeaways
- arXiv:2609.28570v1 Announce Type: new Abstract: Reinforcement learning (RL) is widely used to sharpen reasoning in multimodal large language models (MLLMs), yet its effect on hallucination is uneven.
- We trace this to two weak points in the \emph{correction chain} from reward to parameter update.
- At the rollout level, hard queries---those with high semantic entropy---frequently produce unanimously wrong sample groups, collapsing the group-relative advantage to zero exactly where hallucination risk is highest.
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs” may reshape data collection, model training, output accountability and market access.

Member comments