arXiv Artificial Intelligence

Subspace Alignment for Vision-Language Model Test-time Adaptation

Subspace Alignment for Vision-Language Model Test-time Adaptation

Quick summary

arXiv:2601.08139v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs), despite their extraordinary zero-shot capabilities, are vulnerable to distribution shifts. Test-time adaptation (TTA) emerges as a predominant strategy to adapt VLMs to unlabeled test data on the fly. However, existing TTA methods heavily rely on zero-shot predictions as pseudo-labels for self-training, which can be unreliable under distribution shifts and misguide adaptation due to two fundamental limitations.First (Modality Gap), distribution shifts induce gaps between visual and textual modalities, maki

Key takeaways

  • arXiv:2601.08139v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs), despite their extraordinary zero-shot capabilities, are vulnerable to distribution shifts.
  • Test-time adaptation (TTA) emerges as a predominant strategy to adapt VLMs to unlabeled test data on the fly.
  • However, existing TTA methods heavily rely on zero-shot predictions as pseudo-labels for self-training, which can be unreliable under distribution shifts and misguide adaptation due to two fundamental limitations.First (Modality Gap), distribution shifts induce gaps between visual and textual modalities, maki

Why it matters

“Subspace Alignment for Vision-Language Model Test-time Adaptation” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗