arXiv Artificial Intelligence

HaikuS2S: A Cascaded System For Responding In Verse

HaikuS2S: A Cascaded System For Responding In Verse

Quick summary

arXiv:2609.23951v1 Announce Type: cross Abstract: Expressive speech synthesis has advanced through prosody modeling, yet generating structured poetic speech, such as haiku, remains challenging. Prior work on prosody transfer improves expressiveness, and fine-tuned poetry TTS (text-to-speech) systems capture verse intonation. However, these models do not model haiku's 5-7-5 syllable structure or line-ending pauses. We present a cascaded system, HaikuS2S, combining ASR (automatic speech recognition), LLM (large language model)-generated haiku, and TTS fine-tuning on both prose and custom haiku d

Key takeaways

  • arXiv:2609.23951v1 Announce Type: cross Abstract: Expressive speech synthesis has advanced through prosody modeling, yet generating structured poetic speech, such as haiku, remains challenging.
  • Prior work on prosody transfer improves expressiveness, and fine-tuned poetry TTS (text-to-speech) systems capture verse intonation.
  • However, these models do not model haiku's 5-7-5 syllable structure or line-ending pauses.

Why it matters

“HaikuS2S: A Cascaded System For Responding In Verse” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗