arXiv Artificial Intelligence

Later Is Better: Token Reduction for ViTs Under Distribution Shift

Later Is Better: Token Reduction for ViTs Under Distribution Shift

Quick summary

arXiv:2610.07758v1 Announce Type: cross Abstract: Training-free token reduction accelerates vision transformers by removing redundant tokens across layers, recovering most of the original accuracy at a fraction of the compute. These methods, however, are designed and evaluated primarily on clean data, and under real-world distribution shift their accuracy gap to the uncompressed model widens with the removal rate. We show that this gap is governed by the reduction schedule, the depth profile of removal, usually left fixed as an implementation detail. Concretely, we introduce a one-parameter la

Key takeaways

  • arXiv:2610.07758v1 Announce Type: cross Abstract: Training-free token reduction accelerates vision transformers by removing redundant tokens across layers, recovering most of the original accuracy at a fraction of the compute.
  • These methods, however, are designed and evaluated primarily on clean data, and under real-world distribution shift their accuracy gap to the uncompressed model widens with the removal rate.
  • We show that this gap is governed by the reduction schedule, the depth profile of removal, usually left fixed as an implementation detail.

Why it matters

“Later Is Better: Token Reduction for ViTs Under Distribution Shift” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗