arXiv Artificial IntelligenceStaQ: a Finite Memory Approach to Discrete Action Policy Mirror Descent
arXiv:2506.13862v2 Announce Type: replace-cross Abstract: In Reinforcement Learning (RL), regularization with…
Curated from international AI laboratories, specialist publications and technology outlets. Last update: 3 hours ago.
arXiv Artificial IntelligencearXiv:2506.13862v2 Announce Type: replace-cross Abstract: In Reinforcement Learning (RL), regularization with…
arXiv Artificial IntelligencearXiv:2503.15225v3 Announce Type: replace-cross Abstract: The deployment of autonomous virtual avatars (in…
arXiv Artificial IntelligencearXiv:2503.03156v4 Announce Type: replace-cross Abstract: We propose DiRe, a force-directed dimensionality…
arXiv Artificial IntelligencearXiv:2501.04426v2 Announce Type: replace-cross Abstract: Offline diversity maximization under imitation…
arXiv Artificial IntelligencearXiv:2411.19537v4 Announce Type: replace-cross Abstract: We survey deepfake generation and detection…
arXiv Artificial IntelligencearXiv:2407.02025v5 Announce Type: replace-cross Abstract: Motivated by applications in chemistry and other…
arXiv Artificial IntelligencearXiv:2309.02332v3 Announce Type: replace-cross Abstract: In the mammalian central nervous system, neurons…
arXiv Artificial IntelligencearXiv:2607.24649v2 Announce Type: replace Abstract: Large language models are increasingly used as social…
arXiv Artificial IntelligencearXiv:2607.23802v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has…
arXiv Artificial IntelligencearXiv:2607.22614v2 Announce Type: replace Abstract: RL-based LLM post-training increasingly disaggregates…
arXiv Artificial IntelligencearXiv:2607.22186v3 Announce Type: replace Abstract: Asynchronous reinforcement learning (RL) accelerates…
arXiv Artificial IntelligencearXiv:2607.19338v2 Announce Type: replace Abstract: Coding agents increasingly operate in executable…