arXiv Artificial Intelligence

Instance-Dependent Regret for CMDPs with Step-Wise Constraints

Instance-Dependent Regret for CMDPs with Step-Wise Constraints

Quick summary

arXiv:2610.02520v1 Announce Type: cross Abstract: We study online learning in episodic tabular constrained Markov decision processes with step-wise safety constraints. In such a setting, the constraints induce a safe subgraph that shapes the variance of cumulative rewards under feasible policies and, consequently, the difficulty of learning. Exploiting this structure, however, requires learning which actions are safe while controlling constraint violations. We propose Safe Variance-Adaptive Exploration (SVAE), an efficient algorithm that learns candidate safe subgraphs and performs variance-ad

Key takeaways

  • arXiv:2610.02520v1 Announce Type: cross Abstract: We study online learning in episodic tabular constrained Markov decision processes with step-wise safety constraints.
  • In such a setting, the constraints induce a safe subgraph that shapes the variance of cumulative rewards under feasible policies and, consequently, the difficulty of learning.
  • Exploiting this structure, however, requires learning which actions are safe while controlling constraint violations.

Why it matters

This development is a reminder to test misuse and data-leak scenarios alongside speed and quality. Trust should come from testable controls and clear failure reporting, not protection claims alone.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗