Freeze, Share, Shrink: Rethinking the Action Backbone in Diffusion Policies
Quick summary
arXiv:2511.12101v3 Announce Type: replace-cross Abstract: Many recent Vision-Language-Action models employ diffusion or flow-matching backbones with hundreds of millions of parameters for action generation. However, unlike image synthesis where the output spans millions of diverse pixels, a manipulation policy generates only short sequences of low-dimensional, physically correlated action values, a far simpler target that may not require such capacity. We confirm this intuition and show that, in modulation-conditioned diffusion policies, task adaptation can be routed entirely through the condi
Key takeaways
- arXiv:2511.12101v3 Announce Type: replace-cross Abstract: Many recent Vision-Language-Action models employ diffusion or flow-matching backbones with hundreds of millions of parameters for action generation.
- However, unlike image synthesis where the output spans millions of diverse pixels, a manipulation policy generates only short sequences of low-dimensional, physically correlated action values, a far simpler target that may not require such capacity.
- We confirm this intuition and show that, in modulation-conditioned diffusion policies, task adaptation can be routed entirely through the condi
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “Freeze, Share, Shrink: Rethinking the Action Backbone in Diffusion Policies” may reshape data collection, model training, output accountability and market access.

Member comments