arXiv Artificial Intelligence

dots.tts Technical Report

dots.tts Technical Report

Quick summary

arXiv:2606.07080v2 Announce Type: replace-cross Abstract: We present dots$.$tts, a 2B-parameter continuous autoregressive text-to-speech (TTS) foundation model that models speech in a continuous latent space. Compared with existing continuous autoregressive models, our key innovations are threefold. First, we train an AudioVAE with multiple objectives to build a semantically structured and prediction-friendly continuous speech space. Second, we use full-history conditioning in the flow-matching head to preserve long-range consistency and reduce drift during generation. Third, we apply reward-f

Key takeaways

  • arXiv:2606.07080v2 Announce Type: replace-cross Abstract: We present dots$.$tts, a 2B-parameter continuous autoregressive text-to-speech (TTS) foundation model that models speech in a continuous latent space.
  • Compared with existing continuous autoregressive models, our key innovations are threefold.
  • First, we train an AudioVAE with multiple objectives to build a semantically structured and prediction-friendly continuous speech space.

Why it matters

“dots.tts Technical Report” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗