COMiT: Learning Structured Visual Tokens through Sequential Communication
Quick summary
arXiv:2602.20731v2 Announce Type: replace-cross Abstract: Discrete image tokenizers provide a sequential interface for vision and multimodal models, but are typically optimized for reconstruction or compression and therefore tend to encode local appearance rather than object-level structure. We introduce COMiT, a communication-inspired framework for learning structured discrete visual representations. COMiT constructs a fixed-length latent message through sequential visual observations: at each step, a transformer processes a localized image crop and updates, refines, and reorganizes the exist
Key takeaways
- arXiv:2602.20731v2 Announce Type: replace-cross Abstract: Discrete image tokenizers provide a sequential interface for vision and multimodal models, but are typically optimized for reconstruction or compression and therefore tend to encode local appearance rather than object-level structure.
- We introduce COMiT, a communication-inspired framework for learning structured discrete visual representations.
- COMiT constructs a fixed-length latent message through sequential visual observations: at each step, a transformer processes a localized image crop and updates, refines, and reorganizes the exist
Why it matters
“COMiT: Learning Structured Visual Tokens through Sequential Communication” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments