arXiv Artificial Intelligence

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

Quick summary

arXiv:2609.40165v1 Announce Type: cross Abstract: We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy. Our key idea is to formulate preference learning as preference-conditioned generative modeling: preferred trajector

Key takeaways

  • arXiv:2609.40165v1 Announce Type: cross Abstract: We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories.
  • Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy.
  • Our key idea is to formulate preference learning as preference-conditioned generative modeling: preferred trajector

Why it matters

The significance is not only the legal text but how it changes product design. Decisions around “PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors” may reshape data collection, model training, output accountability and market access.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗