arXiv Artificial Intelligence

Lifted Bellman Linear Programming for Offline Reinforcement Learning

Lifted Bellman Linear Programming for Offline Reinforcement Learning

Quick summary

arXiv:2609.24489v1 Announce Type: cross Abstract: Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets stabilized by target networks with exponential moving average (EMA) updates. Multi-step targets incorporate behavior-policy actions and therefore require off-policy correction. We instead impose in-sample Bellman optimality on the critic through inequality constraints. We formulate the Lifted Bellman Linear Program (LBLP), which lifts the linear programming characterization of Bellman optimality to the joint $(Q,V)$ s

Key takeaways

  • arXiv:2609.24489v1 Announce Type: cross Abstract: Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets stabilized by target networks with exponential moving average (EMA) updates.
  • Multi-step targets incorporate behavior-policy actions and therefore require off-policy correction.
  • We instead impose in-sample Bellman optimality on the critic through inequality constraints.

Why it matters

The significance is not only the legal text but how it changes product design. Decisions around “Lifted Bellman Linear Programming for Offline Reinforcement Learning” may reshape data collection, model training, output accountability and market access.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗