DiTAR+: Dual Optimization for Robust Autoregressive Diffusion Speech Synthesis
Quick summary
arXiv:2609.13909v1 Announce Type: cross Abstract: Continuous-latent Autoregressive Diffusion Transformer (AR-DiT) models have demonstrated immense potential in zero-shot speech generation. However, they still suffer from limited decoding stability when synthesizing long utterances or complex linguistic structures. This instability primarily stems from a restricted historical receptive field and an acoustic inertia dependency within the diffusion decoder, which causes the model to ignore semantic conditions. To address these challenges, we propose DiTAR+, a dual-optimization framework. First, w
Key takeaways
- arXiv:2609.13909v1 Announce Type: cross Abstract: Continuous-latent Autoregressive Diffusion Transformer (AR-DiT) models have demonstrated immense potential in zero-shot speech generation.
- However, they still suffer from limited decoding stability when synthesizing long utterances or complex linguistic structures.
- This instability primarily stems from a restricted historical receptive field and an acoustic inertia dependency within the diffusion decoder, which causes the model to ignore semantic conditions.
Why it matters
“DiTAR+: Dual Optimization for Robust Autoregressive Diffusion Speech Synthesis” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments