arXiv Artificial Intelligence

Stride-k Subsampling: Train-Free Audio Token Reduction for Whisper

Stride-k Subsampling: Train-Free Audio Token Reduction for Whisper

Quick summary

arXiv:2608.30927v1 Announce Type: cross Abstract: Whisper exposes speech through a fixed 1500-token encoder interface, now a default representation for ASR decoders and Whisper-based speech language models (SpeechLMs), yet its redundancy remains largely unexamined. We propose stride-k subsampling, a deterministic indexing operation that retains every k-th token after the convolutional stem or encoder transformer. Across five Whisper scales, k=2 preserves baseline WER at both positions, with CKA attributing this stability to acoustic overlap at the stem and attention-induced redistribution at t

Key takeaways

  • arXiv:2608.30927v1 Announce Type: cross Abstract: Whisper exposes speech through a fixed 1500-token encoder interface, now a default representation for ASR decoders and Whisper-based speech language models (SpeechLMs), yet its redundancy remains largely unexamined.
  • We propose stride-k subsampling, a deterministic indexing operation that retains every k-th token after the convolutional stem or encoder transformer.
  • Across five Whisper scales, k=2 preserves baseline WER at both positions, with CKA attributing this stability to acoustic overlap at the stem and attention-induced redistribution at t

Why it matters

“Stride-k Subsampling: Train-Free Audio Token Reduction for Whisper” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗