HaikuS2S: A Cascaded System For Responding In Verse
Quick summary
arXiv:2609.23951v1 Announce Type: cross Abstract: Expressive speech synthesis has advanced through prosody modeling, yet generating structured poetic speech, such as haiku, remains challenging. Prior work on prosody transfer improves expressiveness, and fine-tuned poetry TTS (text-to-speech) systems capture verse intonation. However, these models do not model haiku's 5-7-5 syllable structure or line-ending pauses. We present a cascaded system, HaikuS2S, combining ASR (automatic speech recognition), LLM (large language model)-generated haiku, and TTS fine-tuning on both prose and custom haiku d
Key takeaways
- arXiv:2609.23951v1 Announce Type: cross Abstract: Expressive speech synthesis has advanced through prosody modeling, yet generating structured poetic speech, such as haiku, remains challenging.
- Prior work on prosody transfer improves expressiveness, and fine-tuned poetry TTS (text-to-speech) systems capture verse intonation.
- However, these models do not model haiku's 5-7-5 syllable structure or line-ending pauses.
Why it matters
“HaikuS2S: A Cascaded System For Responding In Verse” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments