PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents
Quick summary
arXiv:2609.36199v1 Announce Type: cross Abstract: Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute at test time through Best-of-N sampling, but final-sample selection is fixed. Best-of-N can only choose among completed outputs and cannot repair a promising trajectory before it fails. We introduce PreviewDiff, a training-free te
Key takeaways
- arXiv:2609.36199v1 Announce Type: cross Abstract: Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions.
- A common way to improve prompt satisfaction is to spend more compute at test time through Best-of-N sampling, but final-sample selection is fixed.
- Best-of-N can only choose among completed outputs and cannot repair a promising trajectory before it fails.
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments