Exact Distinguishability in Non-Markovian Decision Processes
Quick summary
arXiv:2610.01527v1 Announce Type: cross Abstract: Non-Markovian environments are often modeled as Regular Decision Processes (RDPs), where dynamics depend on the interaction history through a finite automaton. Existing offline guarantees for RDPs rely on a distinguishability assumption on the behaviour policy but provide no means of verifying it. When the assumption is violated, distinct models may explain the data equally well. We study when data collected under a fixed behaviour policy can distinguish two candidate RDPs. We prove that the posterior odds between observationally equivalent can
Key takeaways
- arXiv:2610.01527v1 Announce Type: cross Abstract: Non-Markovian environments are often modeled as Regular Decision Processes (RDPs), where dynamics depend on the interaction history through a finite automaton.
- Existing offline guarantees for RDPs rely on a distinguishability assumption on the behaviour policy but provide no means of verifying it.
- When the assumption is violated, distinct models may explain the data equally well.
Why it matters
“Exact Distinguishability in Non-Markovian Decision Processes” may affect what data AI products can use and where accountability sits. Product teams should watch compliance duties, rights holders should watch enforcement, and users should watch transparency and appeal mechanisms.

Member comments