ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding
Quick summary
arXiv:2609.11642v1 Announce Type: cross Abstract: Neural audio codecs are a fundamental component of modern speech generation systems. While recent codecs achieve increasingly low bitrates, reducing frame rate remains challenging, as each token must preserve more information while maintaining reconstruction quality. We present ZipCodec, a streaming neural speech codec operating at 6.25 Hz and 0.80 kbps with a theoretical latency of 160 ms. Our approach combines large-scale WavLM distillation with a redesigned transformer-based architecture, a scalar spherical quantizer, and a latency-aware str
Key takeaways
- arXiv:2609.11642v1 Announce Type: cross Abstract: Neural audio codecs are a fundamental component of modern speech generation systems.
- While recent codecs achieve increasingly low bitrates, reducing frame rate remains challenging, as each token must preserve more information while maintaining reconstruction quality.
- We present ZipCodec, a streaming neural speech codec operating at 6.25 Hz and 0.80 kbps with a theoretical latency of 160 ms.
Why it matters
“ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Member comments