arXiv Artificial Intelligence

Self-Play Search Distillation for Large Language Model Reasoning

Self-Play Search Distillation for Large Language Model Reasoning

Quick summary

arXiv:2609.30936v1 Announce Type: new Abstract: Improving reasoning abilities in Large Language Models (LLMs) requires high-quality data that exposes difficult decisions, competing alternatives, and their consequences. Data scarcity is driven by the low quality of synthetic data and the cost of human labeling. We introduce Self-Play Search Distillation (SPSD), a framework for generating superhuman synthetic data via self-play of MuZero-like networks trained on board games. SPSD uses executable environments to turn search into structured reasoning problems. At each state, the expert identifies

Key takeaways

  • arXiv:2609.30936v1 Announce Type: new Abstract: Improving reasoning abilities in Large Language Models (LLMs) requires high-quality data that exposes difficult decisions, competing alternatives, and their consequences.
  • Data scarcity is driven by the low quality of synthetic data and the cost of human labeling.
  • We introduce Self-Play Search Distillation (SPSD), a framework for generating superhuman synthetic data via self-play of MuZero-like networks trained on board games.

Why it matters

“Self-Play Search Distillation for Large Language Model Reasoning” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗