arXiv Artificial Intelligence

PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots

PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots

Quick summary

arXiv:2610.01260v1 Announce Type: cross Abstract: Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy. PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion p

Key takeaways

  • arXiv:2610.01260v1 Announce Type: cross Abstract: Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time.
  • We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy.
  • PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion p

Why it matters

“PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots” may affect what data AI products can use and where accountability sits. Product teams should watch compliance duties, rights holders should watch enforcement, and users should watch transparency and appeal mechanisms.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗