arXiv Artificial Intelligence

Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging

Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging

Quick summary

arXiv:2602.03702v2 Announce Type: replace-cross Abstract: Large language models are increasingly trained in continual or open-ended settings, where the total training horizon is not known in advance. Despite this, most existing pretraining recipes are not anytime: they rely on horizon-dependent learning rate schedules and extensive tuning under a fixed compute budget. In this work, we provide a theoretical analysis demonstrating the existence of anytime learning schedules for overparameterized linear regression, and we highlight the central role of weight averaging - also known as model mergin

Key takeaways

  • arXiv:2602.03702v2 Announce Type: replace-cross Abstract: Large language models are increasingly trained in continual or open-ended settings, where the total training horizon is not known in advance.
  • Despite this, most existing pretraining recipes are not anytime: they rely on horizon-dependent learning rate schedules and extensive tuning under a fixed compute budget.
  • In this work, we provide a theoretical analysis demonstrating the existence of anytime learning schedules for overparameterized linear regression, and we highlight the central role of weight averaging - also known as model mergin

Why it matters

“Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗