arXiv Artificial Intelligence

Audio Token Attention Is Predictable Before the Language Model Runs

Audio Token Attention Is Predictable Before the Language Model Runs

Quick summary

arXiv:2609.38878v1 Announce Type: cross Abstract: A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention. Audio tokens draw much more attention there, and their ranking is still far from final, so audio needs a ranking before the language model runs. Surprisingly, the attention an audio token will receive across the language model is already linearly predictable from its encoder output, before the language model runs. A linear map,

Key takeaways

  • arXiv:2609.38878v1 Announce Type: cross Abstract: A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one.
  • Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention.
  • Audio tokens draw much more attention there, and their ranking is still far from final, so audio needs a ranking before the language model runs.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗