arXiv Artificial Intelligence

StateTree: Enhancing Long-Term Dialogue Reasoning via Reinforcement Learning

StateTree: Enhancing Long-Term Dialogue Reasoning via Reinforcement Learning

Quick summary

arXiv:2609.38809v1 Announce Type: cross Abstract: Large language models deployed as personalized assistants must reason over long, evolving interaction histories. However, in long-term dialogue reasoning, relevant evidence is scattered across sessions, preferences may be revised over time, and standard long-context training fails to address these challenges under data scarcity and prohibitive computational costs. We propose StateTree, a data-driven RL method that constructs a challenging auxiliary task from scarce dialogues with verifiable ground truth. StateTree augments multi-session dialogu

Key takeaways

  • arXiv:2609.38809v1 Announce Type: cross Abstract: Large language models deployed as personalized assistants must reason over long, evolving interaction histories.
  • However, in long-term dialogue reasoning, relevant evidence is scattered across sessions, preferences may be revised over time, and standard long-context training fails to address these challenges under data scarcity and prohibitive computational costs.
  • We propose StateTree, a data-driven RL method that constructs a challenging auxiliary task from scarce dialogues with verifiable ground truth.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗