On Trajectory-Aware Training for Masked Diffusion Language Models
Quick summary
arXiv:2609.37974v1 Announce Type: cross Abstract: Masked diffusion models (MDMs) generate text by unmasking several tokens per step, but they are trained and sampled under different conditions. The model is trained on randomly masked sequences, whereas inference follows a trajectory shaped by the model's own predictions. Additionally, each step has no access to what the previous one computed. Recent methods narrow these limitations from separate angles, leaving open how these choices interact. We introduce PUMBA, a unified framework for trajectory-aware training that trains the denoiser on con
Key takeaways
- arXiv:2609.37974v1 Announce Type: cross Abstract: Masked diffusion models (MDMs) generate text by unmasking several tokens per step, but they are trained and sampled under different conditions.
- The model is trained on randomly masked sequences, whereas inference follows a trajectory shaped by the model's own predictions.
- Additionally, each step has no access to what the previous one computed.
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments