arXiv Artificial Intelligence

Dual-Latent Memory Routing for Vision-Language Reasoning

Dual-Latent Memory Routing for Vision-Language Reasoning

Quick summary

arXiv:2609.05539v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have recently made strong progress in vision-language reasoning, yet their performance often degrades as generations grow longer. A key factor is that they frequently lose track of earlier visual evidence and intermediate constraints under a monolithic growing context. Inspired by how humans separately recall what they see and what they infer when solving complex tasks, we propose DLMR, a parameter-efficient mechanism that equips MLLMs with Dual Latent Memories: a visual memory that compresses image evid

Key takeaways

  • arXiv:2609.05539v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have recently made strong progress in vision-language reasoning, yet their performance often degrades as generations grow longer.
  • A key factor is that they frequently lose track of earlier visual evidence and intermediate constraints under a monolithic growing context.
  • Inspired by how humans separately recall what they see and what they infer when solving complex tasks, we propose DLMR, a parameter-efficient mechanism that equips MLLMs with Dual Latent Memories: a visual memory that compresses image evid

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗