arXiv Artificial Intelligence

Rufus-Air: An Open LLM Post-Training Recipe

Rufus-Air: An Open LLM Post-Training Recipe

Quick summary

arXiv:2609.29421v1 Announce Type: cross Abstract: Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe. Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals. Training builds on open-source components and public data, much

Key takeaways

  • arXiv:2609.29421v1 Announce Type: cross Abstract: Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF.
  • We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe.
  • Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals.

Why it matters

“Rufus-Air: An Open LLM Post-Training Recipe” exposes the compute, energy and supply-chain layer behind model competition. Capacity shifts can influence model costs, service availability and the ability of smaller companies to compete.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗