Fusion Training for Mathematical Generalization in Large Language Models
Quick summary
arXiv:2608.09893v1 Announce Type: cross Abstract: Thinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by unifying a non-thinking mode and a thinking mode within a single model. However, its training dynamics, including the \emph{data ratio} and \emph{training schedule} between the two modes, remain underexplored. In this work, we present a systematic study of TMF by analyzing the effects of the training schedule and data ratio between thinking and non-thinking modes. Focusing on mathematical problem solving, we construct a benchmar
Key takeaways
- arXiv:2608.09893v1 Announce Type: cross Abstract: Thinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by unifying a non-thinking mode and a thinking mode within a single model.
- However, its training dynamics, including the \emph{data ratio} and \emph{training schedule} between the two modes, remain underexplored.
- In this work, we present a systematic study of TMF by analyzing the effects of the training schedule and data ratio between thinking and non-thinking modes.
Why it matters
“Fusion Training for Mathematical Generalization in Large Language Models” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments