arXiv Artificial Intelligence

Self-Referenced Social Preferences: Cooperation without Observing Others Rewards

Self-Referenced Social Preferences: Cooperation without Observing Others Rewards

Quick summary

arXiv:2610.07881v1 Announce Type: new Abstract: Social preferences can promote cooperation in multi-agent reinforcement learning, but existing approaches often require agents to observe the rewards of their peers. In many real-world interactions, however, an agent can, as humans do, observe others' behavior and outcomes without access to their private reward signals. We introduce self-referenced social preferences, in which each agent learns a model of its own reward, applies it to other agents' observed transitions to assess their outcomes from its own perspective, and feeds these self-refere

Key takeaways

  • arXiv:2610.07881v1 Announce Type: new Abstract: Social preferences can promote cooperation in multi-agent reinforcement learning, but existing approaches often require agents to observe the rewards of their peers.
  • In many real-world interactions, however, an agent can, as humans do, observe others' behavior and outcomes without access to their private reward signals.
  • We introduce self-referenced social preferences, in which each agent learns a model of its own reward, applies it to other agents' observed transitions to assess their outcomes from its own perspective, and feeds these self-refere

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗