-
作者:Zhao, Feiran; Sha, Xingyu; You, Keyou
作者单位:Tsinghua University; Tsinghua University
摘要:Learning policies in an asynchronous parallel way is essential to numerous successes of reinforcement learning for solving complex problems. However, their convergence has not been rigorously evaluated. To improve the theoretical understanding, we adopt the asynchronous parallel zero-order policy gradient (AZOPG) method to solve the continuous-time linear quadratic regulation problem. Specifically, multiple workers independently perform system rollouts to estimate zero-order policy gradients (...
-
作者:de Oliveira, Paulo J.; Oliveira, Ricardo C. L. F.; Peres, Pedro L. D.
摘要:This article addresses the problem of robust performance analysis and synthesis for uncertain linear systems with state-space matrices containing parameters in bounded intervals. The novelty of the approach relies on treating the interval bounds as optimization variables, leading to two significant results. First, an analysis tool that evaluates the sensitivity of closed-loop performance criteria, such as the H-infinity norm, with respect to the bounds, is presented. The second contribution is...
-
作者:Coutinho, Pedro Henrique Silva; Oliveira, Tiago Roux; Krstic, Miroslav
作者单位:Universidade do Estado do Rio de Janeiro; University of California System; University of California San Diego
摘要:This article deals with the gradient extremum seeking control for static scalar maps with actuators governed by distributed diffusion partial differential equations (PDEs). To achieve the real-time optimization objective, we design a compensation controller for the distributed diffusion PDE via backstepping transformation in infinite dimensions. A further contribution of this article is the appropriate motion planning design of the so-called probing (or perturbation) signal, which is more invo...
-
作者:Feng, Qian; Xiao, Feng; Wang, Xiaoyu
作者单位:North China Electric Power University
摘要:Dissipative estimator (observer) design for continuous time-delay systems poses a significant challenge when multiple pointwise and general distributed delays (DDs) are present. We propose an effective solution to this semiopen problem using the Krasovski functional (KF) framework in conjunction with a quadratic supply rate function, where both the plant and the estimator can accommodate a finite but unlimited number of pointwise and general DDs with square-integrable kernels. A key contributi...
-
作者:Wang, Peng; Lu, Yang; Lian, Jianming; Pan, Lulu; Shao, Haibin; Li, Ning
作者单位:Shanghai Jiao Tong University; Lancaster University; United States Department of Energy (DOE); Oak Ridge National Laboratory
摘要:A privacy-preserving average consensus algorithm is proposed that synergizes the Beaver triple in secret sharing theory and noise obfuscation. The algorithm safeguards the initial values of agents against passive adversaries in a multiagent system. It is proved that the proposed algorithm can concurrently ensure average consensus and privacy, while also reducing the online computation and communication overhead compared to encryption-based ones. In addition, it imposes a less stringent conditi...
-
作者:Jin, Kaijing; Ye, Dan
作者单位:Northeastern University - China; Northeastern University - China
摘要:To reveal the security vulnerabilities of distributed estimation systems, we investigate the existence of false data injection attacks in this article, where attacks can corrupt sensor data and filter data. Considering the innovation-based and consensus-based strict stealthiness, a successful attack should ensure the divergence of each local estimation and not affect innovation and consensus signals. By analyzing the null space of system matrices, the necessary and sufficient conditions for th...
-
作者:Wan, Hexiang; Wang, Guangchen; Xiong, Jie
作者单位:Shandong University; Southern University of Science & Technology; Southern University of Science & Technology
摘要:This article investigates a broad category of McKean-Vlasov type discrete-time partially observable stochastic optimal control problems. The first goal is to prove the dynamic programming principle (DPP) by means of the measurable selection argument, which provides a methodology for finding both the value function as well as the optimal control. Here, we employ the Nisio semigroup technology, which is an intrinsic characterization of the DPP. Then, we derive the recursive formula for the filte...
-
作者:Gao, Rui; Yang, Guang-Hong
作者单位:Northeastern University - China; Northeastern University - China
摘要:In this article, we are concerned with the problem of distributed state estimation for discrete-time linear systems using a network of agents, where the measurement of each agent suffers from the lack of detectability with the system dynamics. The existing results on this topic require stringent condition either on the network connectivity or on the detectability. In contrast, a new form of distributed observer, which ensures that all agents asymptotically estimate the system state under the m...
-
作者:Sasahara, Hampei; Dan, Gyorgy; Amin, Saurabh; Sandberg, Henrik
作者单位:Institute of Science Tokyo; Royal Institute of Technology; Massachusetts Institute of Technology (MIT); Royal Institute of Technology
摘要:Eco-friendly freight operations are crucial for decarbonizing the transportation sector. Systematic analysis of policy measures requires a principled modeling approach. While the commonly used model referred to as a routing game considers the congestible nature of transportation facilities, existing models fail to account for environmental factors. This article aims at providing a mathematical framework to study strategic interaction between owners of mixed fleets comprising both internal comb...
-
作者:Zhang, Weihao; Zhao, Di; Liang, Shu
作者单位:Tongji University; Tongji University; Tongji University; Tongji University
摘要:This article introduces the concept of soft safety for nonlinear dynamical systems, examined from an input-output perspective, and establishes criteria for verifying soft safety in interconnected systems. Diverging from traditional safety approaches that enforce rigid/hard constraints on system states, our proposal involves integrating soft barrier functions with respect to system inputs and outputs to characterize soft safety. Particularly, when the soft barrier function takes a quadratic for...