arXiv Artificial Intelligence

Mitigating the Length-Scaling Tax with Online Distillation

Mitigating the Length-Scaling Tax with Online Distillation

Quick summary

arXiv:2609.38854v1 Announce Type: cross Abstract: Length scaling during reinforcement-learning (RL) post-training is often viewed as a sign of improved reasoning ability, especially on difficult problems, but may also make responses to already-solved problems unnecessarily verbose. We quantify this side effect as the length-scaling tax (LST): excess response length on already-solved queries without a commensurate accuracy gain. To mitigate LST, we propose Length Self-Distillation (LSD), which routes solved prompts to on-policy distillation and retains the original RL objective for unsolved pro

Key takeaways

  • arXiv:2609.38854v1 Announce Type: cross Abstract: Length scaling during reinforcement-learning (RL) post-training is often viewed as a sign of improved reasoning ability, especially on difficult problems, but may also make responses to already-solved problems unnecessarily verbose.
  • We quantify this side effect as the length-scaling tax (LST): excess response length on already-solved queries without a commensurate accuracy gain.
  • To mitigate LST, we propose Length Self-Distillation (LSD), which routes solved prompts to on-policy distillation and retains the original RL objective for unsolved pro

Why it matters

“Mitigating the Length-Scaling Tax with Online Distillation” may affect what data AI products can use and where accountability sits. Product teams should watch compliance duties, rights holders should watch enforcement, and users should watch transparency and appeal mechanisms.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗