Unlocking Pretrained Vision Transformers for Time Series Classification
Quick summary
arXiv:2506.08641v3 Announce Type: replace-cross Abstract: Adapting vision models for time series analysis is compelling, yet all existing approaches are falling short of dedicated time series foundation models (TSFMs) in classification. In this work, we propose Time Vision Transformer (TiViT), the first framework that successfully unlocks the representational power of frozen Vision Transformers (ViTs) pretrained on large-scale image datasets for time series classification. TiViT achieves state-of-the-art performance without any finetuning by utilizing the hidden representations of OpenCLIP mod
Key takeaways
- arXiv:2506.08641v3 Announce Type: replace-cross Abstract: Adapting vision models for time series analysis is compelling, yet all existing approaches are falling short of dedicated time series foundation models (TSFMs) in classification.
- In this work, we propose Time Vision Transformer (TiViT), the first framework that successfully unlocks the representational power of frozen Vision Transformers (ViTs) pretrained on large-scale image datasets for time series classification.
- TiViT achieves state-of-the-art performance without any finetuning by utilizing the hidden representations of OpenCLIP mod
Why it matters
The importance of “Unlocking Pretrained Vision Transformers for Time Series Classification” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Member comments