arXiv Artificial Intelligence

Policy Complexity, Reaction Time, and Bounded Rationality in Reinforcement Learning

Policy Complexity, Reaction Time, and Bounded Rationality in Reinforcement Learning

Quick summary

arXiv:2609.28737v1 Announce Type: cross Abstract: Biological agents do not learn under conditions of unlimited computation. For humans, learning and choice are shaped by constraints on perception, attention, and working memory, which limit how much state information guides behavior and therefore bound policy complexity. Standard reinforcement learning models typically optimize reward without explicitly representing these internal costs, making them less suitable as models of biological intelligence. We derive MI-SARSA, an on-policy temporal-difference algorithm that incorporates mutual-informa

Key takeaways

  • arXiv:2609.28737v1 Announce Type: cross Abstract: Biological agents do not learn under conditions of unlimited computation.
  • For humans, learning and choice are shaped by constraints on perception, attention, and working memory, which limit how much state information guides behavior and therefore bound policy complexity.
  • Standard reinforcement learning models typically optimize reward without explicitly representing these internal costs, making them less suitable as models of biological intelligence.

Why it matters

The significance is not only the legal text but how it changes product design. Decisions around “Policy Complexity, Reaction Time, and Bounded Rationality in Reinforcement Learning” may reshape data collection, model training, output accountability and market access.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗