Beyond Discounted Returns: Robust Markov Decision Processes with Average and Blackwell Optimality

成果类型:
Article
署名作者:
Grand-Clement, Julien; Petrik, Marek; Vieille, Nicolas
署名单位:
University System Of New Hampshire; University of New Hampshire
刊物名称:
OPERATIONS RESEARCH
ISSN/ISSBN:
0030-364X
DOI:
10.1287/opre.2023.0694
发表日期:
2026
关键词:
games complexity MDPS
摘要:
Robust Markov decision processes (RMDPs) are a widely used framework for sequential decision-making under parameter uncertainty. RMDPs have been studied extensively when the objective was to maximize the discounted return, but little is known for average optimality (optimizing the long-run average of the rewards obtained over time) and Blackwell optimality (remaining discount optimal for all discount factors sufficiently close to 1). In this paper, we prove several foundational results for RMDPs beyond the discounted return. We show that average optimal policies can be chosen stationary and deterministic for sa-rectangular RMDPs but, perhaps surprisingly, we show that for s-rectangular RMDPs average optimal policies may not exist, and if they do exist, they may need to be history dependent (Markovian). We also study Blackwell optimality for sa-rectangular RMDPs, where we show that epsilon-Blackwell optimal policies always exist, although Blackwell optimal policies may not exist. We also provide a sufficient condition for their existence, which encompasses virtually any examples from the literature. We then discuss the connection between average and Blackwell optimality, and we describe several algorithms to compute the optimal average return. Interestingly, our approach leverages the connections between RMDPs and stochastic games. Overall, our paper emphasizes the superior practical properties of distancebased sa-rectangular models over s-rectangular models for average and Blackwell optimality.