arXiv Artificial Intelligence

Axiom Satisfiability of Linear Rewards in Alignment

Axiom Satisfiability of Linear Rewards in Alignment

Quick summary

arXiv:2610.06892v1 Announce Type: cross Abstract: Learning from human preference data is the dominant route to aligning language models with human values. In linear social choice, where rewards are linear in a fixed feature representation of prompt-response pairs, Ge et al.[2024] show that fitting such a reward by minimizing any non-decreasing convex loss, including BTL, fails PO and PMC. Moreover, no rule that reads only the majority relation can satisfy PO once the output is required to be linearly induced. We ask what it costs to enforce these axioms anyway. To this end, we relax the linear

Key takeaways

  • arXiv:2610.06892v1 Announce Type: cross Abstract: Learning from human preference data is the dominant route to aligning language models with human values.
  • In linear social choice, where rewards are linear in a fixed feature representation of prompt-response pairs, Ge et al.[2024] show that fitting such a reward by minimizing any non-decreasing convex loss, including BTL, fails PO and PMC.
  • Moreover, no rule that reads only the majority relation can satisfy PO once the output is required to be linearly induced.

Why it matters

This development shows AI moving deeper into everyday software. Productivity potential should be weighed against price, data permissions, exportability and the preservation of human control.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗