Continuous-Time Acoustic Modelling with Neural Controlled Differential Equations
Quick summary
arXiv:2609.11725v1 Announce Type: cross Abstract: Text-to-speech (TTS) models commonly address text--speech alignment by expanding phone-level encoder states to frame-level decoder inputs using predicted durations. While this length-regulation step resolves alignment structurally, this use of duration typically changes only where and how often latent states appear, not the values of the states themselves. This paper proposes a continuous-time mechanism for duration-aware acoustic modelling in TTS using neural controlled differential equations (CDEs). We formulate the phone representation as a
Key takeaways
- arXiv:2609.11725v1 Announce Type: cross Abstract: Text-to-speech (TTS) models commonly address text--speech alignment by expanding phone-level encoder states to frame-level decoder inputs using predicted durations.
- While this length-regulation step resolves alignment structurally, this use of duration typically changes only where and how often latent states appear, not the values of the states themselves.
- This paper proposes a continuous-time mechanism for duration-aware acoustic modelling in TTS using neural controlled differential equations (CDEs).
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “Continuous-Time Acoustic Modelling with Neural Controlled Differential Equations” may reshape data collection, model training, output accountability and market access.

Member comments