EDBT 2026 Demo / reviewers in the wild / expert
Lin Zhao 0009
dblp:72/2195-9
· DBLP profile ↗
20ranked-venue papers
3as first author
15since 2021 · last 2025
0000-0002-1078-887XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 12 since 2021Systems, architecture and hardware · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Robust Self-Reconfiguration for Fault-Tolerant Control of Modular Aerial Robot SystemsabstractModular Aerial Robotic Systems (MARS) consist of multiple drone units assembled into a single, integrated rigid flying platform. With inherent redundancy, MARS can self-reconfigure into different configurations to mitigate rotor or unit failures and maintain stable flight. However, existing works on MARS self-reconfiguration often overlook the practical controllability of intermediate structures formed during the reassembly process, which limits their applicability. In this paper, we address this gap by considering the control-constrained dynamic model of MARS and proposing a robust and efficient self-reconstruction algorithm that maximizes the controllability margin at each intermediate stage. Specifically, we develop algorithms to compute optimal, controllable disassembly and assembly sequences, enabling robust self-reconfiguration. Finally, we validate our method in several challenging fault-tolerant self-reconfiguration scenarios, demonstrating significant improvements in both controllability and trajectory tracking while reducing the number of assembly steps. The videos and source code of this work are available at https://github.com/RuiHuangNUS/MARS-Reconfig/ Zhiqian Cai, Lin Zhao 0009 |
ICRA | 4 |
| 2025 | MARS-FTCP: Robust Fault-Tolerant Control and Agile Trajectory Planning for Modular Aerial Robot SystemsabstractModular Aerial Robot Systems (MARS) consist of multiple drone units that can self-reconfigure to adapt to various mission requirements and fault conditions. However, existing fault-tolerant control methods exhibit significant oscillations during docking and separation, impacting system stability. To address this issue, we propose a novel fault-tolerant control reallocation method that adapts to an arbitrary number of modular robots and their assembly formations. The algorithm redistributes the expected collective force and torque required for MARS to individual units according to their moment arm relative to the center of MARS mass. Furthermore, we propose an agile trajectory planning method for MARS of arbitrary configurations, which is collision-avoiding and dynamically feasible. Our work represents the first comprehensive approach to enable fault-tolerant and collision avoidance flight for MARS. We validate our method through extensive simulations, demonstrating improved fault tolerance, enhanced trajectory tracking accuracy, and greater robustness in cluttered environments. The videos and source code of this work are available at https://github.com/RuiHuangNUS/MARS-FTCP/ Zhiqian Cai, Lin Zhao 0009 |
IROS | 5 |
| 2025 | Safe Reinforcement Learning-Based Eco-Driving Control for Mixed Traffic Flows With DisturbancesabstractThis paper presents a safe learning-based eco-driving framework tailored for mixed traffic flows, which aims to optimize energy efficiency while guaranteeing system constraints during real-system operations. Even though reinforcement learning (RL) is capable of optimizing energy efficiency in intricate environments, it is challenged by safety requirements during both the training and deployment stages. The lack of safety guarantees impedes the application of RL to real-world problems. Compared with RL, model predicted control (MPC) can handle constrained dynamics systems, ensuring safe driving. However, the major challenges lie in complicated eco-driving tasks and the presence of disturbances, which pose difficulties for MPC design and constraint satisfaction. To address these limitations, the proposed framework incorporates the tube-based enhanced MPC (RMPC) to ensure the safe execution of the RL policy under disturbances, thereby improving the control robustness. RL not only optimizes the energy efficiency of the connected and automated vehicle in mixed traffic but also handles more uncertain scenarios, in which the energy consumption of the human-driven vehicle and its diverse and stochastic driving behaviors are considered in the optimization framework. Simulation results demonstrate that the proposed algorithm achieves an average improvement of 10.88% in holistic energy efficiency compared to the RMPC technique, while effectively preventing inter-vehicle collisions when compared to the RL algorithm. Ke Lu 0003, Kaidi Yang, Lin Zhao 0009, Ziyou Song |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Trust-Region Neural Moving Horizon Estimation for RobotsabstractAccurate disturbance estimation is essential for safe robot operations. The recently proposed neural moving horizon estimation (NeuroMHE), which uses a portable neural network to model the MHE’s weightings, has shown promise in further pushing the accuracy and efficiency boundary. Currently, NeuroMHE is trained through gradient descent, with its gradient computed recursively using a Kalman filter. This paper proposes a trust-region policy optimization method for training NeuroMHE. We achieve this by providing the second-order derivatives of MHE, referred to as the MHE Hessian. Remarkably, we show that many of the intermediate results used to obtain the gradient, especially the Kalman filter, can be efficiently reused to compute the MHE Hessian. This offers linear computational complexity with respect to the MHE horizon. As a case study, we evaluate the proposed trust region NeuroMHE on real quadrotor flight data for disturbance estimation. Our approach demonstrates highly efficient training in under 5 min using only 100 data points. It outperforms a state-of-the-art neural estimator by up to 68.1% in force estimation accuracy, utilizing only 1.4% of its network parameters. Furthermore, our method showcases enhanced robustness to network initialization compared to the gradient descent counterpart. Bingheng Wang, Lin Zhao 0009 |
ICRA | 3 |
| 2024 | Global Optimality of Single-Timescale Actor-Critic under Continuous State-Action Space: A Study on Linear Quadratic Regulator
Jingliang Duan, Lin Zhao 0009 |
IJCAI | 3 |
| 2024 | Optimization Landscape of Policy Gradient Methods for Discrete-Time Static Output FeedbackabstractIn recent times, significant advancements have been made in delving into the optimization landscape of policy gradient methods for achieving optimal control in linear time-invariant (LTI) systems. Compared with state-feedback control, output-feedback control is more prevalent since the underlying state of the system may not be fully observed in many practical settings. This article analyzes the optimization landscape inherent to policy gradient methods when applied to static output feedback (SOF) control in discrete-time LTI systems subject to quadratic cost. We begin by establishing crucial properties of the SOF cost, encompassing coercivity, L -smoothness, and M -Lipschitz continuous Hessian. Despite the absence of convexity, we leverage these properties to derive novel findings regarding convergence (and nearly dimension-free rate) to stationary points for three policy gradient methods, including the vanilla policy gradient method, the natural policy gradient method, and the Gauss-Newton method. Moreover, we provide proof that the vanilla policy gradient method exhibits linear convergence toward local minima when initialized near such minima. This article concludes by presenting numerical examples that validate our theoretical findings. These results not only characterize the performance of gradient descent for optimizing the SOF problem but also provide insights into the effectiveness of general policy gradient methods within the realm of reinforcement learning. Jingliang Duan, Jie Li 0042, Kai Zhao 0004, Shengbo Eben Li, Lin Zhao 0009 |
IEEE Trans. Cybern. | 6 |
| 2024 | Neural Moving Horizon Estimation for Robust Flight ControlabstractEstimating and reacting to disturbances is crucial for robust flight control of quadrotors. Existing estimators typically require significant tuning for a specific flight scenario or training with extensive ground-truth disturbance data to achieve satisfactory performance. In this article, we propose a neural moving horizon estimator (NeuroMHE) that can automatically tune its key parameters modeled by a neural network and adapt to different flight scenarios. We achieve this by deriving the analytical gradients of the MHE estimates with respect to the MHE weighting matrices, which enables a seamless embedding of the MHE as a learnable layer into the neural network for highly effective learning. Interestingly, we show that the gradients can be computed efficiently using a Kalman filter in a recursive form. Moreover, we develop a model-based policy gradient algorithm to train NeuroMHE directly from the quadrotor trajectory tracking error without needing the ground-truth disturbance data. The effectiveness of NeuroMHE is verified extensively via both numerical and physical experiments on quadrotors in various challenging flights. Notably, NeuroMHE outperforms a state-of-the-art neural network-based estimator, reducing force estimation errors by up to 76.7%, while using a portable neural network that has only 7.7% of the learnable parameters of the latter. The proposed method is general and can be applied to robust adaptive control of other robotic systems. Bingheng Wang, Zhengtian Ma, Shupeng Lai, Lin Zhao 0009 |
IEEE Trans. Robotics | 4 |
| 2023 | Global Convergence of Two-Timescale Actor-Critic for Solving Linear Quadratic RegulatorabstractThe actor-critic (AC) reinforcement learning algorithms have been the powerhouse behind many challenging applications. Nevertheless, its convergence is fragile in general. To study its instability, existing works mostly consider the uncommon double-loop variant or basic models with finite state and action space. We investigate the more practical single-sample two-timescale AC for solving the canonical linear quadratic regulator (LQR) problem, where the actor and the critic update only once with a single sample in each iteration on an unbounded continuous state and action space. Existing analysis cannot conclude the convergence for such a challenging case. We develop a new analysis framework that allows establishing the global convergence to an epsilon-optimal solution with at most an order of epsilon to -2.5 sample complexity. To our knowledge, this is the first finite-time convergence analysis for the single sample two-timescale AC for solving LQR with global optimality. The sample complexity improves those of other variants by orders, which sheds light on the practical wisdom of single sample algorithms. We also further validate our theoretical findings via comprehensive simulation comparisons. Jingliang Duan, Yingbin Liang, Lin Zhao 0009 |
AAAI | 4 |
| 2023 | Learning Agile Flight Maneuvers: Deep SE(3) Motion Planning and Control for QuadrotorsabstractAgile flights of autonomous quadrotors in clut-tered environments require constrained motion planning and control subject to translational and rotational dynamics. Tra-ditional model-based methods typically demand complicated design and heavy computation. In this paper, we develop a novel deep reinforcement learning-based method that tackles the challenging task of flying through a dynamic narrow gate. We design a model predictive controller with its adaptive tracking references parameterized by a deep neural network (DNN). These references include the traversal time and the quadrotor SE(3) traversal pose that encourage the robot to fly through the gate with maximum safety margins from various initial conditions. To cope with the difficulty of training in highly dynamic environments, we develop a reinforce-imitate learning framework to train the DNN efficiently that generalizes well to diverse settings. Furthermore, we propose a binary search algorithm that allows online adaption of the SE(3) references to dynamic gates in real-time. Finally, through extensive high-fidelity simulations, we show that our approach is adaptive to different gate trajectories, velocities, and orientations. Bingheng Wang, Shenning Zhang, Han Wei Sia, Lin Zhao 0009 |
ICRA | 5 |
| 2023 | Finite-Time Analysis of Single-Timescale Actor-CriticabstractActor-critic methods have achieved significant success in many challenging applications. However, its finite-time convergence is still poorly understood in the most practical single-timescale form. Existing works on analyzing single-timescale actor-critic have been limited to i.i.d. sampling or tabular setting for simplicity. We investigate the more practical online single-timescale actor-critic algorithm on continuous state space, where the critic assumes linear function approximation and updates with a single Markovian sample per actor step. Previous analysis has been unable to establish the convergence for such a challenging scenario. We demonstrate that the online single-timescale actor-critic method provably finds an $\epsilon$-approximate stationary point with $\widetilde{\mathcal{O}}(\epsilon^{-2})$ sample complexity under standard assumptions, which can be further improved to $\mathcal{O}(\epsilon^{-2})$ under the i.i.d. sampling. Our novel framework systematically evaluates and controls the error propagation between the actor and critic. It offers a promising approach for analyzing other single-timescale reinforcement learning algorithms as well. Lin Zhao 0009 |
NeurIPS | 2 |
| 2023 | Unified Mapping Function-Based Neuroadaptive Control of Constrained Uncertain Robotic SystemsabstractFor the existing adaptive constrained robotic control algorithms, the demanding "feasibility conditions" on virtual controller is normally inevitable and the extra limits on constraining functions have to be imposed, making the corresponding approaches more demanding and less user friendly in control development. Here, we develop a new neuroadaptive constrained control strategy for uncertain robotic manipulators in the presence of position and velocity constraints. First, a novel unified mapping function (UMF) is constructed so that the restriction on constraining boundaries is removed and more kinds of constraining forms can be handled. Second, by integrating the UMF-based coordinate transformation with the "universal" approximation characteristic of neural networks over some compact set, the developed neuroadaptive control completely obviates the complicated yet undesired "feasibility conditions." Furthermore, it is proven that all closed-loop signals are semiglobally bounded and the constraints are not violated. The effectiveness of the proposed control is validated via a two-link rigid robotic manipulator. Kai Zhao 0004, Long Chen 0001, Wenchao Meng, Lin Zhao 0009 |
IEEE Trans. Cybern. | 4 |
| 2022 | Deterministic policy gradient: Convergence analysisabstractThe deterministic policy gradient (DPG) method proposed in Silver et al. [2014] has been demonstrated to exhibit superior performance particularly for applications with multi-dimensional and continuous action spaces. However, it remains unclear whether DPG converges, and if so, how fast it converges and whether it converges as efficiently as other PG methods. In this paper, we provide a theoretical analysis of DPG to answer those questions. We study the single timescale DPG (often the case in practice) in both on-policy and off-policy settings, and show that both algorithms attain an $\epsilon$-accurate stationary policy with a sample complexity of $\mathcal{O}(\epsilon^{-2})$. Moreover, we establish the convergence rate for DPG under Gaussian noise exploration, which is widely adopted in practice to improve the performance of DPG. To our best knowledge, this is the first non-asymptotic convergence characterization for DPG methods. Huaqing Xiong, Tengyu Xu, Lin Zhao 0009, Yingbin Liang, Wei Zhang 0013 |
UAI | 3 |
| 2021 | Parallel Collaborative Motion Planning with Alternating Direction Method of MultipliersabstractCollaborative motion planning for multi-agent systems is a challenging problem because of the existence of highly nonlinear and nonconvex constraints. Such difficulties also lead to inavoidable computational inefficiency, which significantly prohibits applying the existing collaborative motion planning algorithms to complex scenarios. This paper proposes a parallel computational algorithm to achieve collaborative motion planning efficiently, considering the nonlinear dynamics model and the nonconvex collision-avoidance constraints. Specifically, the alternating direction method of multipliers (ADMM) framework is elegantly incorporated to separate the large-scale cooperative nonconvex planning problem as two tractable and manageable subproblems, where the two subproblems handle the dynamics constraints and collision-free constraints, respectively. In the proposed approach, the differential dynamic programming (DDP) method is utilized to effectively solve the nonlinear subproblem with the dynamics constraints; meanwhile, the interior point (IPOPT) method is employed to address the nonconvex subproblem derived from the collision-avoidance constraints. Finally, two simulation scenarios are successfully implemented to illustrate the effectiveness of the proposed algorithm. Zilong Cheng, Jun Ma 0008, Lin Zhao 0009, Cheng Xiang 0001, Tong Heng Lee |
IECON | 4 |
| 2021 | Faster Non-asymptotic Convergence for Double Q-learningabstractDouble Q-learning (Hasselt, 2010) has gained significant success in practice due to its effectiveness in overcoming the overestimation issue of Q-learning. However, the theoretical understanding of double Q-learning is rather limited. The only existing finite-time analysis was recently established in (Xiong et al. 2020), where the polynomial learning rate adopted in the analysis typically yields a slower convergence rate. This paper tackles the more challenging case of a constant learning rate, and develops new analytical tools that improve the existing convergence rate by orders of magnitude. Specifically, we show that synchronous double Q-learning attains an $\epsilon$-accurate global optimum with a time complexity of $\tilde{\Omega}\left(\frac{\ln D}{(1-\gamma)^7\epsilon^2} \right)$, and the asynchronous algorithm achieves a time complexity of $\tilde{\Omega}\left(\frac{L}{(1-\gamma)^7\epsilon^2} \right)$, where $D$ is the cardinality of the state-action space, $\gamma$ is the discount factor, and $L$ is a parameter related to the sampling strategy for asynchronous double Q-learning. These results improve the existing convergence rate by the order of magnitude in terms of its dependence on all major parameters $(\epsilon,1-\gamma, D, L)$. This paper presents a substantial step toward the full understanding of the fast convergence of double-Q learning. Lin Zhao 0009, Huaqing Xiong, Yingbin Liang |
NeurIPS | 1 |
| 2021 | Finite-time theory for momentum Q-learningabstractExisting studies indicate that momentum ideas in conventional optimization can be used to improve the performance of Q-learning algorithms. However, the finite-time analysis for momentum-based Q-learning algorithms is only available for the tabular case without function approximation. This paper analyzes a class of momentum-based Q-learning algorithms with finite-time convergence guarantee. Specifically, we propose the MomentumQ algorithm, which integrates the Nesterov’s and Polyak’s momentum schemes, and generalizes the existing momentum-based Q-learning algorithms. For the infinite state-action space case, we establish the convergence guarantee for MomentumQ with linear function approximation under Markovian sampling. In particular, we characterize a finite-time convergence rate which is provably faster than the vanilla Q-learning. This is the first finite-time analysis for momentum-based Q-learning algorithms with function approximation. For the tabular case under synchronous sampling, we also obtain a finite-time convergence rate that is slightly better than the SpeedyQ (Azar et al., NIPS 2011). Finally, we demonstrate through various experiments that the proposed MomentumQ outperforms other momentum-based Q-learning algorithms. Bowen Weng, Huaqing Xiong, Lin Zhao 0009, Yingbin Liang, Wei Zhang 0013 |
UAI | 3 |
| 2020 | Finite-Time Analysis for Double Q-learningabstractAlthough Q-learning is one of the most successful algorithms for finding the best action-value function (and thus the optimal policy) in reinforcement learning, its implementation often suffers from large overestimation of Q-function values incurred by random sampling. The double Q-learning algorithm proposed in~\citet{hasselt2010double} overcomes such an overestimation issue by randomly switching the update between two Q-estimators, and has thus gained significant popularity in practice. However, the theoretical understanding of double Q-learning is rather limited. So far only the asymptotic convergence has been established, which does not characterize how fast the algorithm converges. In this paper, we provide the first non-asymptotic (i.e., finite-time) analysis for double Q-learning. We show that both synchronous and asynchronous double Q-learning are guaranteed to converge to an $\epsilon$-accurate neighborhood of the global optimum by taking $\tilde{\Omega}\left(\left( \frac{1}{(1-\gamma)^6\epsilon^2}\right)^{\frac{1}{\omega}} +\left(\frac{1}{1-\gamma}\right)^{\frac{1}{1-\omega}}\right)$ iterations, where $\omega\in(0,1)$ is the decay parameter of the learning rate, and $\gamma$ is the discount factor. Our analysis develops novel techniques to derive finite-time bounds on the difference between two inter-connected stochastic processes, which is new to the literature of stochastic approximation. Huaqing Xiong, Lin Zhao 0009, Yingbin Liang, Wei Zhang 0013 |
NeurIPS | 2 |
| 2015 | A stochastic hybrid system approach to aggregated load modeling for demand responseabstractThis abstract presents an unified framework for aggregated modeling of various responsive loads. The general model consists of coupled partial differential equations derived by drawing the connection with stochastic hybrid systems. Lin Zhao 0009, Wei Zhang 0013 |
HSCC | 1 |
| 2013 | Robust Stability and Stabilization of Uncertain T-S Fuzzy Systems With Time-Varying Delay: An Input-Output ApproachabstractAn input–output approach to the stability and stabilization of uncertain Takagi–Sugeno (T–S) fuzzy systems with time-varying delay is proposed in this paper. The time-varying parameter uncertainties are assumed to be norm-bounded, and the delay is intervally time varying. A novel method is employed to approximate the time-varying delay, based on which the considered system is transformed into a feedback interconnection form. The new formulation of the system is comprised of a forward subsystem with constant time delay and a feedback subsystem embedding the uncertainties. By applying the scaled small-gain theorem to the converted system, less conservative stability and stabilization criteria are obtained. Moreover, the applicability of the proposed approach to the robust case is simpler since both delay and parameter uncertainties are processed in a unified framework. Numerical experiments are performed to illustrate the advantage of the proposed techniques. Lin Zhao 0009, Huijun Gao, Hamid Reza Karimi |
IEEE Trans. Fuzzy Syst. | 1 |
| 2013 | Integrated Network-Based Model Predictive Control for Setpoints Compensation in Industrial ProcessesabstractComplex industrial processes are controlled by the local regulation controllers at the field level, and the setpoints for the regulation are usually made by manual decomposition of the overall economic objective according to the operators' experience. If a precise static process model can be built, real-time optimization (RTO) can be used to generate the setpoints. Nevertheless, since the aforementioned control structure is actually open-loop, the desired economic objective of the whole processes may not be tracked when disturbances exist. Aiming at solving this problem, a novel network based model predictive control method (MPC) for setpoints compensation is proposed in this paper. Firstly, a multivariable proportional integral (PI) controller is designed to perform the local regulation control. Secondly, a stochastic packet dropout model is adopted to characterize the measurement and human-in-the-loop delay effect. Then, a model predictive controller considering the random dropout effect is developed to compensate the setpoints dynamically according to the changing conditions of the processes, such that the prescribed performance objective can be obtained. Finally, a flotation process model is employed to demonstrate the effectiveness of the proposed method. Tianyou Chai, Lin Zhao 0009, Jianbin Qiu, Fangzhou Liu 0001, Jialu Fan |
IEEE Trans. Ind. Informatics | 2 |
| 2009 | Stochastic stability of Markovian jumping Hopfield neural networks with constant and distributed delays
Lin Zhao 0009, Zexu Zhang, Yan Ou |
Neurocomputing | 2 |