arXiv Artificial Intelligence

Look Less, Hear Better: Jointly Rewarded GRPO for Streaming ASR

Look Less, Hear Better: Jointly Rewarded GRPO for Streaming ASR

Quick summary

arXiv:2609.18333v1 Announce Type: cross Abstract: Streaming automatic speech recognition (ASR) must be judged jointly on what it transcribes and on how quickly it commits each word. Delayed streams modeling (DSM) has become the dominant paradigm for streaming large audio-language models, exposing a structural delay $\tau$ that bounds the decoder's lookahead. We show that $\tau$ is a poor proxy for user-perceived latency, and that the alignment-based supervision of DSM leaves latency on the table: the same forced-aligned transcript is used at every $\tau$, forcing the model to withhold words it

Key takeaways

  • arXiv:2609.18333v1 Announce Type: cross Abstract: Streaming automatic speech recognition (ASR) must be judged jointly on what it transcribes and on how quickly it commits each word.
  • Delayed streams modeling (DSM) has become the dominant paradigm for streaming large audio-language models, exposing a structural delay $\tau$ that bounds the decoder's lookahead.
  • We show that $\tau$ is a poor proxy for user-perceived latency, and that the alignment-based supervision of DSM leaves latency on the table: the same forced-aligned transcript is used at every $\tau$, forcing the model to withhold words it

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗