Trust Region Q Adjoint Matching
Quick summary
arXiv:2605.27079v2 Announce Type: replace-cross Abstract: Off-policy reinforcement learning of pretrained flow policies remains challenging due to the instability of optimization arising from the multi-step sampling process. Recently, Q-learning with Adjoint Matching (QAM) addressed this by recasting policy improvement as a stochastic optimal control (SOC) problem guided by a learned critic. However, QAM inherits a fundamental fragility of critic-guided improvement, since small critic errors can be exponentially amplified and often lead to performance collapse. This paper introduces Trust Regi
Key takeaways
- arXiv:2605.27079v2 Announce Type: replace-cross Abstract: Off-policy reinforcement learning of pretrained flow policies remains challenging due to the instability of optimization arising from the multi-step sampling process.
- Recently, Q-learning with Adjoint Matching (QAM) addressed this by recasting policy improvement as a stochastic optimal control (SOC) problem guided by a learned critic.
- However, QAM inherits a fundamental fragility of critic-guided improvement, since small critic errors can be exponentially amplified and often lead to performance collapse.
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “Trust Region Q Adjoint Matching” may reshape data collection, model training, output accountability and market access.

Member comments