Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents
Quick summary
arXiv:2607.11433v2 Announce Type: replace Abstract: Omni-modal agents must seek evidence across video, audio, web pages, and computation to answer questions. Their main bottleneck is planning: noisy multimodal observations accumulate in conversation history and disrupt later decisions, while multimodal models have limited capacity for multi-step planning. Controlled backend replacements support this diagnosis: replacing the planner causes a much larger performance loss than replacing the perception backend. We present Omni-Decision, an omni-modal agent built on evidence-ledger planning: it rep
Key takeaways
- arXiv:2607.11433v2 Announce Type: replace Abstract: Omni-modal agents must seek evidence across video, audio, web pages, and computation to answer questions.
- Their main bottleneck is planning: noisy multimodal observations accumulate in conversation history and disrupt later decisions, while multimodal models have limited capacity for multi-step planning.
- Controlled backend replacements support this diagnosis: replacing the planner causes a much larger performance loss than replacing the perception backend.
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments