arXiv Artificial Intelligence

COMiT: Learning Structured Visual Tokens through Sequential Communication

COMiT: Learning Structured Visual Tokens through Sequential Communication

Quick summary

arXiv:2602.20731v2 Announce Type: replace-cross Abstract: Discrete image tokenizers provide a sequential interface for vision and multimodal models, but are typically optimized for reconstruction or compression and therefore tend to encode local appearance rather than object-level structure. We introduce COMiT, a communication-inspired framework for learning structured discrete visual representations. COMiT constructs a fixed-length latent message through sequential visual observations: at each step, a transformer processes a localized image crop and updates, refines, and reorganizes the exist

Key takeaways

  • arXiv:2602.20731v2 Announce Type: replace-cross Abstract: Discrete image tokenizers provide a sequential interface for vision and multimodal models, but are typically optimized for reconstruction or compression and therefore tend to encode local appearance rather than object-level structure.
  • We introduce COMiT, a communication-inspired framework for learning structured discrete visual representations.
  • COMiT constructs a fixed-length latent message through sequential visual observations: at each step, a transformer processes a localized image crop and updates, refines, and reorganizes the exist

Why it matters

“COMiT: Learning Structured Visual Tokens through Sequential Communication” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗