Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations
Quick summary
arXiv:2609.03940v1 Announce Type: cross Abstract: Most previous work on speech enhancement (SE) based on masked generative modeling relied on discrete token representations of audio signals, obtained using neural audio codecs (NACs). However, a recent study has shown that continuous latent representations of NACs can be advantageous for SE in terms of speech quality and intelligibility. In this work, we propose masked autoregressive SE (MARSE), a method for SE based on iterative decoding of masked clean speech frames using continuous NAC representations of speech. In particular, we investigate
Key takeaways
- arXiv:2609.03940v1 Announce Type: cross Abstract: Most previous work on speech enhancement (SE) based on masked generative modeling relied on discrete token representations of audio signals, obtained using neural audio codecs (NACs).
- However, a recent study has shown that continuous latent representations of NACs can be advantageous for SE in terms of speech quality and intelligibility.
- In this work, we propose masked autoregressive SE (MARSE), a method for SE based on iterative decoding of masked clean speech frames using continuous NAC representations of speech.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments