arXiv Artificial Intelligence

Parametric Multimodal User Memory: Storing What Captions Cannot Carry

Parametric Multimodal User Memory: Storing What Captions Cannot Carry

Quick summary

arXiv:2608.28609v1 Announce Type: cross Abstract: A personalized agent needs a user memory: a persistent model of who its user is. Today it is almost always text -- transcripts and captions retrieved by similarity. This serves the captionable half of a person ("my cat is named Bibi"), but discards the perceptual half no caption can hold: how a voice sounds, how a face reads across age and lighting, how tired someone sounds. We measure this loss across five modalities: a strong caption-based re-identifier recovers as little as 0.11 of a dedicated encoder's recall, collapsing toward chance on no

Key takeaways

  • arXiv:2608.28609v1 Announce Type: cross Abstract: A personalized agent needs a user memory: a persistent model of who its user is.
  • Today it is almost always text -- transcripts and captions retrieved by similarity.
  • This serves the captionable half of a person ("my cat is named Bibi"), but discards the perceptual half no caption can hold: how a voice sounds, how a face reads across age and lighting, how tired someone sounds.

Why it matters

“Parametric Multimodal User Memory: Storing What Captions Cannot Carry” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗