Seeking Physics in Diffusion Noise
Quick summary
arXiv:2603.14294v3 Announce Type: replace-cross Abstract: Do video diffusion models encode signals predictive of physical plausibility? We probe intermediate denoising representations of pretrained Diffusion Transformers (DiTs) and find that physically plausible and implausible videos are partially separable in mid-layer feature space, even at high noise levels. Within-source and perceptual-quality controls suggest that this signal is not fully explained by generator identity or generic visual quality. We distill the signal into a lightweight, backbone-specific physics verifier trained on froz
Key takeaways
- arXiv:2603.14294v3 Announce Type: replace-cross Abstract: Do video diffusion models encode signals predictive of physical plausibility?
- We probe intermediate denoising representations of pretrained Diffusion Transformers (DiTs) and find that physically plausible and implausible videos are partially separable in mid-layer feature space, even at high noise levels.
- Within-source and perceptual-quality controls suggest that this signal is not fully explained by generator identity or generic visual quality.
Why it matters
“Seeking Physics in Diffusion Noise” is a product decision that may change how people work with AI. Its value depends on task completion, correction effort and data handling—not simply the presence of a new feature.

Member comments