On the Convergence of Modified Policy Iteration in Risk-Sensitive Exponential Cost Markov Decision Processes
成果类型:
Article
署名作者:
Murthy, Yashaswini; Moharrami, Mehrdad; Srikant, Rayadurgam
署名单位:
California Institute of Technology; University of Iowa; University of Illinois System; University of Illinois Urbana-Champaign
刊物名称:
OPERATIONS RESEARCH
ISSN/ISSBN:
0030-364X
DOI:
10.1287/opre.2024.0818
发表日期:
2026
关键词:
algorithm
FORMULA
chains
摘要:
Modified policy iteration (MPI) is a dynamic programming algorithm that combines elements of policy iteration and value iteration. The convergence of MPI is wellstudied in the context of discounted and average-cost Markov decision processes (MDPs). In this work, we consider the exponential cost risk-sensitive MDP formulation, which is known to provide some robustness to model parameters. Although policy iteration and value iteration are well-studied in the context of risk-sensitive MDPs, MPI is unexplored. To the best of our knowledge, we provide the first proof that MPI also converges for the risk-sensitive problem in the case of finite state and action spaces. Because the exponential cost formulation deals with the multiplicative Bellman equation, our main contribution is a convergence proof that is quite different than existing results for discounted and risk-neutral average-cost as well as risk-sensitive value and iteration