arXiv Artificial Intelligence

Bigger Text Encoders Can Hurt CLIP Zero-Shot Performance

Bigger Text Encoders Can Hurt CLIP Zero-Shot Performance

Quick summary

arXiv:2609.05730v1 Announce Type: cross Abstract: Contrastive Language-Image Pretraining (CLIP) is a building block of many machine learning applications. Scaling laws have guided resource allocation for large-scale training, yet prior work treats total CLIP model size as a single variable, without exploring how the capacity split between encoders impacts downstream performance. Here, we train multiple CLIP models with different vision and text encoder sizes, revealing that for most vision encoders, there is an optimal text encoder size beyond which zero-shot performance degrades---even as tot

Key takeaways

  • arXiv:2609.05730v1 Announce Type: cross Abstract: Contrastive Language-Image Pretraining (CLIP) is a building block of many machine learning applications.
  • Scaling laws have guided resource allocation for large-scale training, yet prior work treats total CLIP model size as a single variable, without exploring how the capacity split between encoders impacts downstream performance.
  • Here, we train multiple CLIP models with different vision and text encoder sizes, revealing that for most vision encoders, there is an optimal text encoder size beyond which zero-shot performance degrades---even as tot

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗