arXiv Artificial Intelligence

StackTok: Accelerating VLMs Inference with Budget-Adaptive Visual Token Selection

StackTok: Accelerating VLMs Inference with Budget-Adaptive Visual Token Selection

Quick summary

arXiv:2609.16841v1 Announce Type: cross Abstract: Increasing image resolution produces ever-longer visual-token sequences in vision-language models (VLMs), substantially raising their inference cost. To reduce this overhead without retraining, existing methods select compact token subsets that prioritize query relevance, visual coverage, or a fixed trade-off between them. The appropriate balance, however, varies across queries and token budgets: localized questions favor relevance, whereas holistic questions demand broader visual coverage. We introduce StackTok, a training-free selector that t

Key takeaways

  • arXiv:2609.16841v1 Announce Type: cross Abstract: Increasing image resolution produces ever-longer visual-token sequences in vision-language models (VLMs), substantially raising their inference cost.
  • To reduce this overhead without retraining, existing methods select compact token subsets that prioritize query relevance, visual coverage, or a fixed trade-off between them.
  • The appropriate balance, however, varies across queries and token budgets: localized questions favor relevance, whereas holistic questions demand broader visual coverage.

Why it matters

The importance of “StackTok: Accelerating VLMs Inference with Budget-Adaptive Visual Token Selection” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗