Deep Learning for High-Dimensional Continuous-Time Stochastic Optimal Control Without Explicit Solution

成果类型:
Article
署名作者:
Dupret, Jean-Loup; Hainaut, Donatien
署名单位:
Swiss Federal Institutes of Technology Domain; ETH Zurich; Universite Catholique Louvain
刊物名称:
OPERATIONS RESEARCH
ISSN/ISSBN:
0030-364X
DOI:
10.1287/opre.2024.1102
发表日期:
2026
关键词:
partial-differential-equations NEURAL-NETWORK APPROACH algorithm
摘要:
This paper introduces the generalized policy iteration physics-informed neural network algorithm, a novel numerical scheme for solving continuous-time stochastic optimal control problems in high dimensions when the optimal control does not admit an explicit solution. The proposed iterative deep learning algorithm leverages physicsinformed neural networks on the Hamilton-Jacobi-Bellman residuals in an actor-critic fashion, which is built from the generalized policy iteration technique. It employs two separate neural networks to approximate both the value function and the multidimensional optimal control, achieving a global approximation of the optimal solution across all time and space that can be evaluated online rapidly. Theoretical guarantees on convergence and optimality are provided, whereas its accuracy and efficacy are empirically validated through two important numerical examples from operations research. In particular, we generalize the standard Almgren-Chriss model arising from optimal liquidation in finance by allowing for a price impact model with fully nonlinear temporary and permanent impact functions and by considering a multidimensional setting with numerous cointegrated assets.