arXiv Artificial Intelligence

Best-of-$N$ Guidance for Test-time Diffusion Alignment

Best-of-$N$ Guidance for Test-time Diffusion Alignment

Quick summary

arXiv:2610.05108v2 Announce Type: replace-cross Abstract: Diffusion models achieve strong generative performance but often struggle to align generated samples with human preferences measured by a reward model. A simple yet effective algorithm for test-time alignment is Best-of-$N$ (BoN) sampling, which draws $N$ i.i.d. samples from a pre-trained diffusion model and outputs the single highest-reward sample. Despite its empirical success, BoN makes limited use of reward information, as it is incorporated only at the final selection stage without influencing the reverse diffusion trajectory durin

Key takeaways

  • arXiv:2610.05108v2 Announce Type: replace-cross Abstract: Diffusion models achieve strong generative performance but often struggle to align generated samples with human preferences measured by a reward model.
  • A simple yet effective algorithm for test-time alignment is Best-of-$N$ (BoN) sampling, which draws $N$ i.i.d.
  • samples from a pre-trained diffusion model and outputs the single highest-reward sample.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗