arXiv Artificial Intelligence

GrainSpeech: Less Context, More Detail for Compact Speech Synthesis

GrainSpeech: Less Context, More Detail for Compact Speech Synthesis

Quick summary

arXiv:2609.18856v1 Announce Type: cross Abstract: Compact acoustic models face a challenging quality-capacity trade-off. We investigate two factors in this regime: encoder context and Mel-spectrogram supervision. A receptive-field-scaling study shows that expanding self-attention beyond 15 phonemes provides no consistent gains in pitch, energy, or duration prediction. Guided by this finding, we introduce a fixed-receptive-field convolutional encoder that reduces the respective prediction errors by 36.0%, 17.3%, and 3.4%. We further show that directly transferring image-domain gradient-variance

Key takeaways

  • arXiv:2609.18856v1 Announce Type: cross Abstract: Compact acoustic models face a challenging quality-capacity trade-off.
  • We investigate two factors in this regime: encoder context and Mel-spectrogram supervision.
  • A receptive-field-scaling study shows that expanding self-attention beyond 15 phonemes provides no consistent gains in pitch, energy, or duration prediction.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗