Finite Time Convergence Analysis of Policy Gradient Learning in Stochastic Stackelberg Games
成果类型:
Article
署名作者:
Das, Pranoy; Zaman, Muhammad Aneeq uz; Gupta, Vijay
署名单位:
Purdue University System; Purdue University; University of Illinois System; University of Illinois Urbana-Champaign
刊物名称:
IEEE TRANSACTIONS ON AUTOMATIC CONTROL
ISSN/ISSBN:
0018-9286
DOI:
10.1109/TAC.2026.3678838
发表日期:
2026
关键词:
摘要:
Designing and analyzing learning algorithms for general sum stochastic Stackelberg games remain challenging. We propose an inner-outer loop policy gradient-based learning algorithm for this problem and analyze its finite time convergence. Our analysis does not assume a time scale separation between the leader and the follower or a special structure of the game, such as being Markov potential or zero-sum.