arXiv Artificial Intelligence

Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations

Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations

Quick summary

arXiv:2609.03940v1 Announce Type: cross Abstract: Most previous work on speech enhancement (SE) based on masked generative modeling relied on discrete token representations of audio signals, obtained using neural audio codecs (NACs). However, a recent study has shown that continuous latent representations of NACs can be advantageous for SE in terms of speech quality and intelligibility. In this work, we propose masked autoregressive SE (MARSE), a method for SE based on iterative decoding of masked clean speech frames using continuous NAC representations of speech. In particular, we investigate

Key takeaways

  • arXiv:2609.03940v1 Announce Type: cross Abstract: Most previous work on speech enhancement (SE) based on masked generative modeling relied on discrete token representations of audio signals, obtained using neural audio codecs (NACs).
  • However, a recent study has shown that continuous latent representations of NACs can be advantageous for SE in terms of speech quality and intelligibility.
  • In this work, we propose masked autoregressive SE (MARSE), a method for SE based on iterative decoding of masked clean speech frames using continuous NAC representations of speech.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗