-
作者:Ding, Yuhao; Zhang, Junzi; Lee, Hyunin; Lavaei, Javad
作者单位:University of California System; University of California Berkeley; University of California System; University of California Berkeley
摘要:Entropy regularization is an efficient technique for encouraging exploration and preventing a premature convergence of (vanilla) policy gradient (PG) methods in reinforcement learning (RL). However, the theoretical understanding of entropy-regularized RL algorithms has been limited. In this article, we revisit the classical entropy-regularized PG methods with the soft-max policy parametrization, whose convergence has so far only been established assuming access to exact gradient oracles. To go...
-
作者:Basilio, Joao C.; Silva, Guilherme M. Ottoni
作者单位:Universidade Federal Rural do Rio de Janeiro (UFRRJ); Universidade Federal do Rio de Janeiro
摘要:Studies carried out in real plants have shown thatrepeated and/or intermittent fault events occur frequently duringthe operation of a system. Regarding the occurrence of repeatedfault events, the problem becomes that of determining the numberof occurrences of the fault event, namely, for a given kappa is an element of Z & lowast;+,and based on the observation of events, be sure that the faultevent has occurred at least kappa times (kappa-diagnosability). In thisarticle, the problem of repeated...
-
作者:Yang, Yanhua; Mei, Jie; Shi, Xiongtao; Ma, Guangfu
作者单位:Harbin Institute of Technology; Harbin Institute of Technology
摘要:In this article, the consensus problem of multiple Euler-Lagrange (EL) systems with time-varying asymmetric full-state constraints under a directed graph is investigated in a fully distributed way, where the global information dependence is removed. First, to prevent the violation of the full-state constraints of each EL system, a nonlinear state-dependent transformation is adopted for both the leader and followers, where the constraints will not be violated as long as the transformed states r...
-
作者:Krupa, Pablo; Limon, Daniel; Bemporad, Alberto; Alamo, Teodoro
作者单位:IMT School for Advanced Studies Lucca; University of Sevilla
摘要:Harmonic model predictive control (HMPC) is a recent model predictive control (MPC) formulation for tracking piece-wise constant references that includes a parameterized artificial harmonic reference as a decision variable, resulting in an increased performance and domain of attraction with respect to other MPC formulations. This article presents an extension of the HMPC formulation to track periodic harmonic/sinusoidal references and discusses its use for tracking arbitrary trajectories. The ...
-
作者:Ma, Yong-Sheng; Che, Wei-Wei; Wu, Zheng-Guang
作者单位:Northeastern University - China; Zhejiang University
摘要:This article studies the consensus problem in multiagent systems under the challenge of an unknown system model and limited communication resources. A novel model-free adaptive learning algorithm is developed to learn the controller from system data. A model-based event-triggered fully distributed control (ET-FDC) algorithm is proposed to achieve consensus while saving the limited communication resources. Furthermore, a data-driven systematic learning methodology for the ET-FDC algorithm is in...
-
作者:Sun, Yifan; Lu, Jianquan; Ho, Daniel W. C.; Li, Lulu
作者单位:Southeast University - China; Southeast University - China; City University of Hong Kong
摘要:In this article, we develop a new denial-of-service (DoS) estimator, enabling defenders to identify duration and frequency parameters of any DoS attacker, except for three edge cases, exclusively using real-time data. The key advantage of the estimator lies in its capability to facilitate security control in a wide range of practical scenarios, even when the attacker's information is previously unknown. We demonstrate the advantage and application of our new estimator in the context of two cla...
-
作者:Chen, Xin; Poveda, Jorge I.; Li, Na
作者单位:Texas A&M University System; Texas A&M University College Station; University of California System; University of California San Diego; Harvard University
摘要:This article introduces a class of model-free feedback methods for solving generic constrained optimization problems where the mathematical forms of the cost and constraint functions are not available. The proposed methods, termed projected zeroth-order (P-ZO) dynamics, incorporate projection maps into a class of continuous-time zeroth-order dynamics that use direct measurements of the cost function and periodic dithering for the purpose of gradient learning. In particular, the proposed P-ZO a...
-
作者:Feng, Shilun; Shi, Dawei; Chen, Tongwen; Shi, Ling
作者单位:Beijing Institute of Technology; Beijing Institute of Technology; University of Alberta; Hong Kong University of Science & Technology
摘要:Event-triggered control has attracted considerable attention for its effectiveness in resource-restricted applications. To make event-triggered control as an end-to-end solution, a key issue is how to effectively learn unknown system dynamics from event-triggered measurements and consequently, develop a learning-based event-triggered controller. Existing works learn system dynamics based on periodic time-triggered measurements, and it is yet to know how to learn a controller with performance g...
-
作者:Wang, Yunan; Hu, Chuxiong; Li, Zeyang; Lin, Yujie; Lin, Shize; He, Suqin
作者单位:Tsinghua University; Tsinghua University; Tsinghua University
摘要:Time-optimal control for high-order chain-of-integrator systems with full state constraints remains an open and challenging problem within the discipline of optimal control. The behavior of optimal control in high-order problems lacks precise characterization. Even the existence of the chattering phenomenon, i.e., the control switches infinitely many times over a finite period, remains unknown. This article establishes a theoretical framework for chattering in the problem, providing novel find...
-
作者:Krener, Arthur J.
作者单位:United States Department of Defense; United States Navy; Naval Postgraduate School
摘要:The uncontrolled diffusion equation on the unit disk under homogeneous Neumann boundary conditions is only neutrally stable, as zero is an eigenvalue. If we add linear reaction terms, the resulting equation can be unstable. To stabilize such a reaction-diffusion equation to a uniform state, which we conveniently take to be zero, we pose and solve a linear quadratic regulator (LQR). We consider two types of control actuation-boundary arc control and boundary point control. With boundary arc con...