arXiv Artificial Intelligence

Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning

Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning

Quick summary

arXiv:2608.14963v1 Announce Type: cross Abstract: Pareto Conditioned Networks learn multiple multi-objective reinforcement learning behaviours by conditioning a single policy on a desired return command. However, the local mapping from command and state to action remains opaque. We propose command-space counterfactual explanations for PCNs: given a fixed state, original command, and foil action, we search, in a black-box setting, for a minimally changed desired-return command under which the same trained policy would choose the foil. Our contributions are threefold. First, we formulate PCN exp

Key takeaways

  • arXiv:2608.14963v1 Announce Type: cross Abstract: Pareto Conditioned Networks learn multiple multi-objective reinforcement learning behaviours by conditioning a single policy on a desired return command.
  • However, the local mapping from command and state to action remains opaque.
  • We propose command-space counterfactual explanations for PCNs: given a fixed state, original command, and foil action, we search, in a black-box setting, for a minimally changed desired-return command under which the same trained policy would choose the foil.

Why it matters

The significance is not only the legal text but how it changes product design. Decisions around “Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning” may reshape data collection, model training, output accountability and market access.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗