arXiv Artificial Intelligence

Decide Before You Look: Learning Which Retrieved Memories Deserve Pixels

Decide Before You Look: Learning Which Retrieved Memories Deserve Pixels

Quick summary

arXiv:2610.07984v1 Announce Type: cross Abstract: Multimodal assistants answer questions from long-term memories that contain images. After retrieval, each retrieved image reaches the answering model either as pixels, at about a thousand visual tokens per image, or as a stored text proxy that often misses the detail the question asks about. We find that the benefit of pixels usually comes from one or two retrieved memories, and that it can be predicted before the answering model runs, without reading any full-resolution image. In PixelTriage, a plug-in placed after retrieval, a small model tha

Key takeaways

  • arXiv:2610.07984v1 Announce Type: cross Abstract: Multimodal assistants answer questions from long-term memories that contain images.
  • After retrieval, each retrieved image reaches the answering model either as pixels, at about a thousand visual tokens per image, or as a stored text proxy that often misses the detail the question asks about.
  • We find that the benefit of pixels usually comes from one or two retrieved memories, and that it can be predicted before the answering model runs, without reading any full-resolution image.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗