VLDB 2026 Research / reviewers in the wild / expert
Biao Luo 0001
dblp:22/5293-1
· DBLP profile ↗
113ranked-venue papers
21as first author
81since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 88 · 15 first-author · 66 since 2021Human-computer interaction and ubiquitous computing · 14 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Monotonicity: Revisiting Factorization Principles in Multi-Agent Q-LearningabstractValue decomposition is a central approach in multi-agent reinforcement learning (MARL), enabling centralized training with decentralized execution by factorizing the global value function into local values. To ensure individual-global-max (IGM) consistency, existing methods either enforce monotonicity constraints, which limit expressive power, or adopt softer surrogates at the cost of algorithmic complexity. In this work, we present a dynamical systems analysis of non-monotonic value decomposition, modeling learning dynamics as continuous-time gradient flow. We prove that, under approximately greedy exploration, all zero-loss equilibria violating IGM consistency are unstable saddle points, while only IGM-consistent solutions are stable attractors of the learning dynamics. Extensive experiments on both synthetic matrix games and challenging MARL benchmarks demonstrate that unconstrained, non-monotonic factorization reliably recovers IGM-optimal solutions and consistently outperforms monotonic baselines. Additionally, we investigate the influence of temporal-difference targets and exploration strategies, providing actionable insights for the design of future value-based MARL algorithms. Tianmeng Hu, Yongzheng Cui, Biao Luo 0001, Ke Li 0001 |
AAAI | 4 |
| 2026 | Autonomous Partner Selection for Cooperative Multi-Agent Reinforcement LearningabstractIn cooperative Multi-Agent Reinforcement Learning (MARL), the subgroup-wise learning is employed to assign sub-tasks to agents towards the enhancement of team collaboration. However, the present work is dependent on manually defined allocation criteria, which hinders its capacity to adapt to environmental changes promptly, and also relaxes communication restrictions, thereby constraining the application of algorithms in a range of fields. In order to address these issues, the Autonomous Partner Selection (APS) framework is proposed, which offers an implicit grouping mechanism in an autonomous way. Each agent is capable of autonomously selecting cooperative partners and integrating their own observations with those of partners to harmonise the cooperative behaviour during the training stage. With a view to strictly restricting communication, the intention encoder is trained through information distillation, which enables agents to selectively take more cooperative actions based solely on local observations. Meanwhile, in order to circumvent potential conflicts engendered by homogenization behaviour, we employ a contrastive learning strategy to the cooperative intention generated by agents, thereby ensuring that the behavioural tendencies exhibited by different individuals remain as diverse as possible. Finally, extensive comparative experiments on the StarCraft Multi-Agent Challenge and Google Research Football are conducted. The results demonstrate that APS exhibits superior performance in comparison to the state-of-the-art algorithms across a range of tasks, and agents can adapt their grouping strategies in accordance with the environment to facilitate enhanced cooperation. Biao Luo 0001, Yongzheng Cui |
AAAI | 2 |
| 2026 | Adaptive population classification based multi-strategy evolutionary algorithm for dynamic constrained multi-objective optimization
Biao Luo 0001, Zhanglu Hou, Jinhua Zheng, Ke Li 0001 |
Expert Syst. Appl. | 2 |
| 2026 | Asymmetric ADP-driven event-triggered optimal control of industrial sintering temperature field with input constraints
Binyan Li, Xiaopeng Cao, Biao Luo 0001, Jiayuan Gao, Ning Chen 0009, Shaosheng Fan, Weihua Gui 0001 |
Neurocomputing | 3 |
| 2026 | Neural network with reduced-order GSINDy-Koopman operators for modeling and robust control of nonlinear processes
Zihang Xu, Yu Xiao 0006, Yuan Yuan 0017, Xiaodong Xu 0002, Biao Luo 0001 |
Neurocomputing | 5 |
| 2026 | Adaptive neural network mobile controller design of linear parabolic PDE systems with unknown transmission disturbances
Xiao-Wei Zhang, Hai-Fei Cui, Xiaoli Li 0011, Huai-Ning Wu, Biao Luo 0001 |
Neurocomputing | 5 |
| 2026 | Time-Varying HJBE-Based Adaptive Safe Critic Control Design for Stochastic Asymmetric Constrained Multiagent SystemsabstractIn this article, we investigate the problem of adaptive safe critic control design for stochastic multiagent systems (MASs) subject to asymmetric state and input constraints. To systematically address asymmetric state constraints, a unified transformation function (UTF) is proposed to convert the constrained consensus control problem into the stability analysis of an unconstrained error system. In addition, a nonquadratic cost function is incorporated to address input limitations effectively. Building upon these developments, a time-varying Hamilton-Jacobi-Bellman equation (HJBE) is formulated by integrating the Bellman optimality principle with Itô's lemma, thereby accommodating stochastic disturbances and enhancing controller robustness. To improve data utilization and eliminate reliance on explicit drift dynamics, an integral reinforcement learning (IRL) algorithm is developed within this framework. Furthermore, a time-varying single-critic network is designed to approximate the solution to the HJBE and generate optimal control policies, thereby considerably reducing computational complexity. To further enhance learning efficiency and relax the persistent excitation (PE) condition, the experience replay (ER) technique is incorporated into the update process of the critic weight. Finally, two simulation examples are provided to verify the feasibility and effectiveness of the proposed approach. Yuhao Zhou 0001, Biao Luo 0001, Xiaodong Xu 0002, Yalin Wang 0003, Weihua Gui 0001 |
IEEE Trans. Cybern. | 2 |
| 2026 | DACESR: Degradation-Aware Conditional Embedding for Real-World Image Super-ResolutionabstractMultimodal large models have shown excellent ability in addressing image super-resolution in real-world scenarios by leveraging language class as condition information, yet their abilities in degraded images remain limited. In this paper, we first revisit the capabilities of the Recognize Anything Model (RAM) for degraded images by calculating text similarity. We find that directly using contrastive learning to fine-tune RAM in the degraded space is difficult to achieve acceptable results. To address this issue, we employ a degradation selection strategy to propose a Real Embedding Extractor (REE), which achieves significant recognition performance gain on degraded image content through contrastive learning. Furthermore, we use a Conditional Feature Modulator (CFM) to incorporate the high-level information of REE for a powerful Mamba-based network, which can leverage effective pixel information to restore image textures and produce visually pleasing results. Extensive experiments demonstrate that the REE can effectively help image super-resolution networks balance fidelity and perceptual quality, highlighting the great potential of Mamba in real-world applications. The source code of this work will be made publicly available at: https://github.com/nathan66666/DACESR.git. Xiaoyan Lei, Biao Luo 0001, Weifeng Cao, Qiuting Lin |
IEEE Trans. Image Process. | 3 |
| 2026 | Spatiotemporal Topology-Informed Multiagent Reinforcement Learning Framework for Structured Multiprocess Collaborative OptimizationabstractIndustrial multiprocess collaborative optimization presents significant challenges due to the intricate spatiotemporal dependencies inherent in modern process industries. Traditional optimization and reinforcement learning often treat subprocesses as independent entities, neglecting the fine-grained interdependencies among operational variables across different subprocesses. To fundamentally address this limitation, we introduce, a novel spatiotemporal topology-informed multiprocess collaborative optimization (STI-MCO) framework, which pioneers action-level interdependency modeling through an innovative spatiotemporal graph architecture. Rather than treating subprocesses as monolithic entities, STI-MCO operates at the operational variable level, enabling precise representation of both interprocess relationships and intraprocess dependencies through a hierarchical two-stage decision framework. This approach enables more precise coordination through fine-grained variable interactions, better temporal consistency via dynamic graph structures, and enhanced scalability compared with conventional agent-level methods. This paradigm shift from subprocess-level to variable-level collaboration, combined with dynamic graph-based coordination, enables extensive simulations and experiments conducted across three benchmark environments with progressively complex topologies to demonstrate that STI-MCO consistently outperforms baseline methods, achieving up to 38.9% improvement over centralized methods and 171.9% improvement over existing multiagent strategies. In addition, STI-MCO exhibits superior convergence efficiency, requiring significantly fewer training steps to achieve high performance. Its practical applicability is further validated through deployment in a real-world Salt Lake chemical process. By fundamentally shifting the optimization paradigm from holistic subprocess control to fine-grained variable-level collaboration, this work establishes a new framework for more effective optimization in complex industrial processes, particularly those with strong interunit coupling. Diju Liu, Yalin Wang 0003, Chenliang Liu, Biao Luo 0001, Biao Huang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2026 | IRL-Based Optimal Consensus Control of MASs With Predefined Time Convergence PerformanceabstractThis study addresses the adaptive optimal consensus control problem for nonlinear multiagent systems (MASs). To enhance the convergence speed of the consensus error, a predefined time performance technique is integrated into the optimal control framework. Unlike the conventional Hamilton–Jacobi–Bellman equation (HJBE), a time-varying HJBE is formulated to solve the optimal control problem for MASs. To further improve efficiency, an integral reinforcement learning (IRL) algorithm is developed, which eliminates the need for precise system dynamics during controller design. In addition, a single critic network is employed to simultaneously evaluate system performance and execute control actions, effectively reducing computational complexity. The experience replay technique is incorporated into the update law for the critic network weights, thus alleviating the requirement for persistent excitation. Finally, a simulation example is presented to validate the feasibility and effectiveness of the proposed method. Yuhao Zhou 0001, Biao Luo 0001, Xiaodong Xu 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2025 | Adaptive neural boundary control for multi-agent manipulators system with uncertainties through cooperative disturbance observers network
Zhibo Zhao, Yuan Yuan 0017, Xiaodong Xu 0002, Biao Luo 0001, Tingwen Huang |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | CSFIN: A lightweight network for camouflaged object detection via cross-stage feature interaction
Minghong Li, Yuqian Zhao 0001, Fan Zhang 0106, Gui Gui, Biao Luo 0001, Chunhua Yang 0001, Weihua Gui 0001, Kan Chang |
Expert Syst. Appl. | 5 |
| 2025 | Event-triggered H∞ control for highly dissipative spatiotemporal process via ADP: Application to industrial sintering process
Binyan Li, Jiayuan Gao, Ning Chen 0009, Biao Luo 0001, Shaosheng Fan, Weihua Gui 0001 |
Neurocomputing | 4 |
| 2025 | Auto-tuning strategy for deep Koopman robust model predictive control design based on advanced Metaheuristics
Zijia Meng, Yu Xiao 0006, Xiaodong Xu 0002, Biao Luo 0001, Yuncheng Du, Weihua Gui 0001 |
Neurocomputing | 4 |
| 2025 | Pinning boundary sampled-data synchronization of coupled reaction-diffusion neural networks
Zipeng Wang 0001, Bo-Ming Chen, Junfei Qiao 0001, Biao Luo 0001, Huai-Ning Wu, Tingwen Huang, Guangwei Chen |
Neurocomputing | 4 |
| 2025 | Mine-SSD: Dual-threshold set abstraction and radius-adaptive grouping for 3D object detection in open-pit mines
Zhongyu Xie, Yuqian Zhao 0001, Fan Zhang 0106, Biao Luo 0001, Wenliu Hu, Tenghai Qiu |
Neurocomputing | 4 |
| 2025 | Neural operator-based composite learning adaptive backstepping control for linear 2×2 hyperbolic PDE systems
Yu Xiao 0006, Xiaodong Xu 0002, Biao Luo 0001, Weihua Gui 0001, Tingwen Huang |
Neurocomputing | 4 |
| 2025 | Offline reward shaping with scaling human preference feedback for deep reinforcement learning
Biao Luo 0001, Xiaodong Xu 0002, Tingwen Huang |
Neural Networks | 2 |
| 2025 | Reinforcement Learning-Based 3D Trajectory Tracking Control of Hypersonic Gliding Vehicles With Time-Varying UncertaintiesabstractIn this paper, a robust three-dimensional trajectory tracking control scheme based on reinforcement learning is proposed for the glide phase of a hypersonic gliding vehicle (HGV) with time-varying uncertainties. First, the non-affine nonlinear full-state kinematics and dynamics model of the HGV glide phase is constructed. Then, without linearizing the system, the desired multiplanar reference trajectories for HGVs are planned based on the pseudo-spectral theory under the input constraints, initial conditions, and terminal conditions. Subsequently, the full-state error system is generated by subtracting the reference system state from the actual state of the HGV system with time-varying uncertainty. For the full-state HGV error system with time-varying uncertainty and input constraints, we design a reinforcement learning-based optimal control scheme for its nominal system and establish the equivalence between this optimal control and the robust control of the original HGV error system. A single-evaluation network structure is used in the concrete implementation to reduce the computational cost. A rigorous theory is given to demonstrate the uniform ultimate boundedness of the closed-loop system and the weight error. Finally, we perform simulation traces for reference trajectories with different optimization performances to verify the effectiveness of the proposed method. Note to Practitioners—There are various constraints and uncertainties in the glide phase of HGVs, which is the hinge connecting the initial descent phase and the terminal management phase. How to design robust trajectory tracking controllers for the glide phase of HGVs with complex environments and large span of flight parameters is of great significance to aerial guidance practitioners. In this paper, an RL-based three-dimensional trajectory robust tracking guidance method is proposed for the HGV glide phase system, which can resist time-varying uncertainties and satisfy flight constraints. The uniform ultimate boundedness of the closed-loop system is proved using the Lyapunov method. The proposed tracking algorithm is effective for reference trajectories with different performance indexes. Biao Luo 0001, Xiaodong Xu 0002 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | Reinforcement Learning-Based Optimal Formation Control for Multiple WMRs With Visual ServoingabstractIn this paper, a reinforcement learning (RL) control method is developed for the formation control of multiple wheeled mobile robots (WMRs) with visual servoing. First, a multi-robot system model is constructed based on the kinematic models of mobile robots, the camera model, and the multiple-view geometry principles. The leader-follower structure is then applied to derive the distributed error system. Next, the error term is separated from the optimal performance index function, and the Bellman residual error is obtained based on the Hamilton-Jacobi-Bellman equation (HJBE). Subsequently, the gradient descent method is employed to design the weight update rate, which is implemented in an actor-critic neural network (NN) architecture. The proposed RL control method achieves formation tracking and performance optimization simultaneously, which previous approaches have not accomplished. Furthermore, under the Lyapunov stability theory, it is proven that the follower robots can track the leader in a predefined formation, and the tracking error converges ultimately. The simulation outcomes verify the effectiveness of the developed approach. Biao Luo 0001, Yuhao Zhou 0001, Jialin Xiao, Bing-Chuan Wang, Chunhua Yang 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | Dynamic Event-Triggered Control for Hierarchical Differential GamesabstractThis paper proposes a novel dynamic event-triggered control method for a class of completely unknown nonaffine hierarchical differential games, incorporating asymmetric boundaries in both system states and control strategies. To tackle this problem, dynamic feedback and mapping functions are first introduced to construct an unconstrained affine augmented system. Then, integral reinforcement learning techniques are used to derive the Hamilton-Jacobi equation without the original system dynamics. Furthermore, dynamic event-triggered control is employed to alleviate the network transmission burden. During the algorithm implementation, critic neural networks are designed for each agent. Analysis results show that the states and weights are ultimately uniformly bounded. Finally, simulation results using the torsional pendulum system and RLC circuit system validate the effectiveness of the present method. Shan Xue 0004, Biao Luo 0001, Weidong Zhang 0004, Derong Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2025 | Composite Learning Based Adaptive Control of Linear 2 × 2 Hyperbolic PDE SystemsabstractThis article considers the adaptive stability control of a class of linear hyperbolic PDE systems. The PDE model is subject to constant but in-domain and boundary unknown parameters. A novel adaptive controller is developed by leveraging the swapping design technique and composite parameter learning law. With swapping design, several linear and static combinations, including carefully designed filters, unknown parameters, and error terms, are constructed to express the system states. From the static combinations, a composite learning based forgetting-factor least squares law is introduced to guarantee exponential parameter convergence without the persistent excitation (PE). Although inaccurate parameter estimation in the adaptive backstepping control results in asymptotic stability of the system, accurate parameter estimation ensures the exponential convergence of closed-loop system and concomitantly improves the transient performance. Finally, a comparative numerical simulation is performed to validate the effectiveness and advantage of the developed adaptive control scheme. Yu Xiao 0006, Yun Feng 0001, Biao Luo 0001, Han-Xiong Li, Xiaodong Xu 0002 |
IEEE Trans. Cybern. | 3 |
| 2025 | Multistep Q-Learning-Based Optimal Consensus Control of Linear Discrete-Time Multiagent SystemsabstractThis article considers the optimal consensus control for the multiagent systems problem. By developing the multiagent multistep Q-learning (MaMsQL), the methodology achieves enhanced efficiency while addressing the issue of the complex interaction dynamics between agents, environmental uncertainty, thus ultimately meeting demand of balancing exploration and exploitation. First, associated with the performance index, the Q-function is established to prove that all optimal Q-functions form a Nash equilibrium outcome, thereby the consensus problem is converted to finding the optimal Q-functions. Then, the MaMsQL method is developed with theoretical proof of its convergence. Finally, the method is implemented through a specially designed Actor-Critic network. By virtue of the comparison with multiagent single step Q-learning, the effectiveness and superiority of this method are verified through simulation examples. Jialin Xiao, Biao Luo 0001, Xiaodong Xu 0002, Chunhua Yang 0001, Weihua Gui 0001 |
IEEE Trans. Cybern. | 2 |
| 2025 | Integral Reinforcement Learning-Based Dynamic Event-Triggered Nonzero-Sum Games of USVsabstractIn this article, an integral reinforcement learning (IRL) method is developed for dynamic event-triggered nonzero-sum (NZS) games to achieve the Nash equilibrium of unmanned surface vehicles (USVs) with state and input constraints. Initially, a mapping function is designed to map the state and control of the USV into a safe environment. Subsequently, IRL-based coupled Hamilton-Jacobi equations, which avoid dependence on system dynamics, are derived to solve the Nash equilibrium. To conserve computational resources and reduce network transmission burdens, a static event-triggered control is initially designed, followed by the development of a more flexible dynamic form. Finally, a critic neural network is designed for each player to approximate its value function and control policy. Rigorous proofs are provided for the uniform ultimate boundedness of the state and the weight estimation errors. The effectiveness of the present method is demonstrated through simulation experiments. Shan Xue 0004, Weidong Zhang 0004, Biao Luo 0001, Derong Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2025 | Consensus of Nonlinear Uncertain Delayed Multiagent Systems Modeled by PDEs via Adaptive Boundary ControlabstractUnder the influence of nonlinearity, time-varying delay, and uncertainty, the consensus problem is concerned in this study for multiagent systems modeled by partial differential equations, which means that both the time and space variables are included in the dynamic behavior of each agent. First, with a directed graph, an adaptive boundary controller is developed under boundary measurements, which can effectively reduce the control cost with dynamic control gains and a few actuators and sensors installed at the boundary of the spatial domain. Then, through the designed adaptive boundary controller, the linear matrix inequality (LMI)-based consensus conditions are obtained to ensure the exponential stability of the consensus error systems derived by utilizing the inequality techniques and Lyapunov direct approach. Lastly, two numerical examples demonstrate the effectiveness of the presented adaptive boundary control protocols. Xu Zhang 0051, Biao Luo 0001, Zipeng Wang 0001, Xiaodong Xu 0002, Chunhua Yang 0001 |
IEEE Trans. Cybern. | 2 |
| 2025 | Adaptive Neural Consensus Observer Networks Design for a Class of Semilinear Parabolic PDE SystemsabstractThis article concerns the investigation on the consensus problem for the joint state-uncertainty estimation of a class of parabolic partial differential equation (PDE) systems with parametric and nonparametric uncertainties. We propose a two-layer network consisting of informed and uninformed boundary observers where novel adaptation laws are developed for the identification of uncertainties. Particularly, all observer agents in the network transmit their information with each other across the entire network. The proposed adaptation laws include a penalty term of the mismatch between the parameter estimates generated by the other observer agents. Moreover, for the nonparametric uncertainties, radial basis function (RBF) neural networks are employed for the universal approximation of unknown nonlinear functions. Given the persistently exciting condition, it is shown that the proposed network of adaptive observers can achieve exponential joint state-uncertainty estimation in the presence of parametric uncertainties and ultimate bounded estimation in the presence of nonparametric uncertainties based on the Lyapunov stability theory. The effects of the proposed consensus method are demonstrated through a typical reaction-diffusion system example, which implies convincing numerical findings. Mingxing Cai, Yuan Yuan 0017, Biao Luo 0001, Fanbiao Li, Xiaodong Xu 0002, Chunhua Yang 0001, Weihua Gui 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Cognition-Oriented Multiagent Reinforcement LearningabstractInspired by psychological insights into individual behavior, we propose a novel cognition-oriented multiagent reinforcement learning (CORL) framework. CORL equips agents with two distinct types of cognition-situational and self-cognition-derived from local observations. To enhance the informativeness and precision of these cognition types, we introduce two information-theoretical regularizers: one to align situational cognition with the global state and the other to align self-cognition with each agent's identity for improved role differentiation and team coordination. In addition, the centralized training and decentralized execution framework is adopted to train the policy network. Our simulations demonstrate that CORL effectively harnesses local observations for enriched cooperation, leading to pronounced performance improvements, particularly in challenging tasks. Tenghai Qiu, Shiguang Wu 0001, Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi, Yuqian Zhao 0001, Biao Luo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | A Nonlinear Noise-Resistant Zeroing Neural Network Model for Solving Time-Varying Quaternion Generalized Lyapunov Equation and Applications to Color Image ProcessingabstractThe time-varying Lyapunov equation (TVLE) plays a crucial role in control design and system stability. However, there has been limited research conducted on the time-varying generalized Lyapunov equation in the quaternion field. To tackle the time-varying quaternion generalized Lyapunov equation, a nonlinear noise-resistant zeroing neural network (NNR-ZNN) model with a novel power activation function (NPAF) is devised. The issue of non-commutativity within quaternion is circumvented by utilizing the real representation. The theoretical analyses provide a sufficient explanation for the global stability, fixed-time convergence, and robustness of the NNR-ZNN model. Under several different kinds of noises, the exceptional robustness of the NNR-ZNN model is highlighted by comparison with other existing models. In the end, the successful applications of the NNR-ZNN model to color image fusion and color image denoising confirm the practical value of the NNR-ZNN model. Lin Xiao 0002, Xiangru Yan, Yongjun He 0001, Biao Luo 0001, Qiya Song |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | A Hybrid Adaptive Dynamic Programming for Optimal Tracking Control of USVsabstractThis article presents an efficient method for solving the optimal tracking control policy of unmanned surface vehicles (USVs) using a hybrid adaptive dynamic programming (ADP) approach. This approach integrates data-driven integral reinforcement learning (IRL) and dynamic event-driven (DED) mechanisms into the solution of the control policy of the established augmented system while obtaining both the feedforward and feedback components of the tracking controller. For the USV model and the reference trajectory, an augmented system is established, and the tracking Hamilton-Jacobi-Bellman (HJB) equation is derived based on IRL, aiming to fully utilize system data information and reduce model dependency. For the solution of the tracking HJB equation, the DED-based controller update rule is used to further reduce the burden of network transmission. In implementing the ADP method, the DED experience replay-based weight update rule is utilized to recycle data resources. Experiments show that compared with the static event-driven (SED) approach, the DED approach reduces the sample size by 78% and increases the average interval by about four times. Shan Xue 0004, Weidong Zhang 0004, Biao Luo 0001, Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Adaptive Boundary Control for Synchronization of Reaction-Diffusion Neural Networks With Random Time-Varying DelayabstractThis article addresses the synchronization problem of reaction-diffusion neural networks (RDNNs) with random time-varying delay (RTVD) via boundary control (BC) (including adaptive BC and BC with constant-valued gain) under distributed measurements or boundary measurements. First, a novel BC strategy with constant-valued gain is designed, which considers three cases of the measurements, that is, distributed measurements, boundary measurements, and both coexist. Subsequently, an adaptive BC scheme under boundary measurements is proposed, where the control gain is regulated effectively. Next, based on the inequality techniques and Lyapunov direct approach, the delay-dependent synchronization conditions are gained and some linear matrix inequalities (LMIs) based theorems are given. Then, the BC design for the delayed RDNNs is transformed into an LMI feasibility problem. Finally, the developed BC approaches are validated by the simulation results. Xu Zhang 0051, Biao Luo 0001, Zipeng Wang 0001, Xiaodong Xu 0002, Chunhua Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Dynamic Self-Triggered Control for Nonzero-Sum Games of Unknown Nonlinear Constrained Systems via Generalized Fuzzy Hyperbolic ModelsabstractIn this article, a novel adaptive dynamic programming algorithm is devised to handle the multiplayer nonzero-sum (NZS) games of completely unknown nonlinear systems, subject to state and input constraints under the dynamic self-triggered mechanism. Initially, in order to eliminate the demand for system dynamics, a generalized fuzzy hyperbolic model based identifier is established, which only relies on input–output data. Then, the equivalent transformation of the reconstructed system is implemented by virtue of barrier functions. With the aid of the nonquadratic utility function, the Hamilton–Jacobi equation of the NZS game is derived. After that, an adaptive critic scheme with experience replay is employed to acquire the Nash equilibrium solution. Furthermore, a novel dynamic self-triggered rule is proposed with the dead-zone operation, which not only significantly reduces source consumption but also overcomes the implementation difficulty of monitoring hardware in the event-triggered mechanism. Moreover, the stability of the system and the uniform ultimate boundedness of the critic weights are guaranteed. Ultimately, two simulation examples are given to validate the feasibility of the developed method. Fan Liu 0013, Hanguang Su, Huaguang Zhang, Biao Luo 0001, Jiawei Wang 0015 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2025 | DOMAIN: Mildly Conservative Model-Based Offline Reinforcement LearningabstractModel-based reinforcement learning (RL), which learns an environment model from the offline dataset and generates more out-of-distribution model data, has become an effective approach to the problem of distribution shift in offline RL. Due to the gap between the learned and actual environment, conservatism should be incorporated into the algorithm to balance accurate offline data and imprecise model data. The conservatism of current algorithms mostly relies on model uncertainty estimation. However, uncertainty estimation is unreliable and leads to poor performance in certain scenarios, and the previous methods ignore differences between the model data, which brings great conservatism. To address the above issues, this article proposes a mildly conservative model-based offline RL algorithm (DOMAIN) without estimating model uncertainty, and designs the adaptive sampling distribution of model samples, which can adaptively adjust the model data penalty. In this article, we theoretically demonstrate that theQvalue learned by the DOMAIN outside the region is a lower bound of the trueQvalue, the DOMAIN is less conservative than previous model-based offline RL algorithms, and has the guarantee of safety policy improvement. The results of extensive experiments show that DOMAIN outperforms prior RL algorithms and the average performance has improved by 1.8% on the D4RL benchmark. Xiao-Yin Liu, Xiao-Hu Zhou, Mei-Jiang Gui, Xiaoliang Xie, Shiqi Liu 0004, Shuangyi Wang, Qi-Chao Zhang, Biao Luo 0001, Zeng-Guang Hou |
IEEE Trans. Syst. Man Cybern. Syst. | 9 |
| 2024 | PA2D-MORL: Pareto Ascent Directional Decomposition Based Multi-Objective Reinforcement LearningabstractMulti-objective reinforcement learning (MORL) provides an effective solution for decision-making problems involving conflicting objectives. However, achieving high-quality approximations to the Pareto policy set remains challenging, especially in complex tasks with continuous or high-dimensional state-action space. In this paper, we propose the Pareto Ascent Directional Decomposition based Multi-Objective Reinforcement Learning (PA2D-MORL) method, which constructs an efficient scheme for multi-objective problem decomposition and policy improvement, leading to a superior approximation of Pareto policy set. The proposed method leverages Pareto ascent direction to select the scalarization weights and computes the multi-objective policy gradient, which determines the policy optimization direction and ensures joint improvement on all objectives. Meanwhile, multiple policies are selectively optimized under an evolutionary framework to approximate the Pareto frontier from different directions. Additionally, a Pareto adaptive fine-tuning approach is applied to enhance the density and spread of the Pareto frontier approximation. Experiments on various multi-objective robot control tasks show that the proposed method clearly outperforms the current state-of-the-art algorithm in terms of both quality and stability of the outcomes. Tianmeng Hu, Biao Luo 0001 |
AAAI | 2 |
| 2024 | Exploring heterophily in calibration of graph neural networks
Xiaotian Xie, Biao Luo 0001, Yingjie Li 0008, Chunhua Yang 0001, Weihua Gui 0001 |
Neurocomputing | 2 |
| 2024 | Object detection on low-resolution images with two-stage enhancement
Minghong Li, Yuqian Zhao 0001, Gui Gui, Fan Zhang 0106, Biao Luo 0001, Chunhua Yang 0001, Weihua Gui 0001, Kan Chang, Hui Wang 0069 |
Knowl. Based Syst. | 5 |
| 2024 | Concurrent learning adaptive boundary observer design for linear coupled hyperbolic partial differential equation systems
Linbin Teng, Yuan Yuan 0017, Xiaodong Xu 0002, Chunhua Yang 0001, Biao Luo 0001, Stevan Dubljevic, Tingwen Huang |
Knowl. Based Syst. | 5 |
| 2024 | Multi-scale feature selection network for lightweight image super-resolution
Minghong Li, Yuqian Zhao 0001, Fan Zhang 0106, Biao Luo 0001, Chunhua Yang 0001, Weihua Gui 0001, Kan Chang |
Neural Networks | 4 |
| 2024 | Modular hierarchical reinforcement learning for multi-destination navigation in hybrid crowds
Wen Ou, Biao Luo 0001, Bingchuan Wang |
Neural Networks | 2 |
| 2024 | Neural operators for robust output regulation of hyperbolic PDEs
Yu Xiao 0006, Yuan Yuan 0017, Biao Luo 0001, Xiaodong Xu 0002 |
Neural Networks | 3 |
| 2024 | Event-Triggered Neuro-Adaptive Fixed-Time Control for Nonlinear Switched and Constrained Systems: An Initial Condition-Independent MethodabstractThis paper investigates a neuro-adaptive fixed-time tracking control issue for switched nonlinear systems subject to asymmetric time-varying constraints and unknown control gains. Unlike the current study on constraint problems, the system’s initial condition is unavailable in this article, which causes specific difficulties in constructing the Barrier Lyapunov Function. A novel shifting function is presented to unify the initial values of all system states. In addition, the system convergence time becomes known and adjustable by utilizing the Nussbaum gain technique and fixed-time stability criterion. An adaptive neural tracking control scheme is proposed based on the learning ability of neural networks and fixed-time theory. To alleviate the computational burden, we present the single learning parameter method such that the number of adaptive laws is reduced significantly. Furthermore, a novel switching threshold mechanism that considers the system errors is developed to balance the communication burden and control performance. Finally, the simulation example illustrates the feasibility of the proposed control strategy. Xin Wang 0028, Yuhao Zhou 0001, Biao Luo 0001, Yushuai Li, Tingwen Huang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | RL-Based Adaptive Optimal Bipartite Consensus Control for Nonlinear Heterogeneous MASs via Event-Triggered State FeedbackabstractThis article investigates a leader-following bipartite consensus issue for uncertain nonlinear heterogeneous multiagent systems (MASs). Initially, within the framework of optimal control theory, we employ the reinforcement learning (RL) algorithm to derive an approximate solution to the Hamilton-Jacobi-Bellman equation (HJBE). Specifically, the neural networks (NNs) are utilized to construct the Actor-Critic structure with the aim of implementing control behavior and evaluating system performance, respectively. An additional network is employed to address nonlinear uncertainties existing in the system. Furthermore, we design a static threshold event-triggered mechanism (ETM) to achieve the event-triggered state feedback-based control strategy. By utilizing this event-triggered state information, we reconstruct the approximate optimal controller and update laws of neural network weights, effectively reducing the communication burden while ensuring that all signals of the MASs remain bounded. Finally, two simulation examples are carried out to demonstrate the feasibility of the proposed method. Yuhao Zhou 0001, Biao Luo 0001, Xin Wang 0028, Xiaodong Xu 0002, Lin Xiao 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | Boundary Optimal Control for Parabolic Distributed Parameter Systems With Value IterationabstractA reinforcement learning-based boundary optimal control algorithm for parabolic distributed parameter systems is developed in this article. First, a spatial Riccati-like equation and an integral optimal controller are derived in infinite-time horizon based on the principle of the variational method, which avoids the complex semigroups and operator theories. Using state data along the system trajectory, a value iteration algorithm via the Bellman optimality principle is proposed to obtain the solution of the spatial Riccati-like equation and the optimal control law. The convergence of the value iteration algorithm is proved. Subsequently, an approximation scheme based on weighted residuals is developed to implement the value iteration algorithm, where radial basis functions are chosen as the basic functions to approximate the solution of the spatial Riccati-like equation. Simulations on the diffusion-reaction process demonstrate the effectiveness of the developed method. Biao Luo 0001, Xiaodong Xu 0002, Chunhua Yang 0001 |
IEEE Trans. Cybern. | 2 |
| 2024 | Adaptive Neural Tracking Control of a Class of Hyperbolic PDE With Uncertain Actuator DynamicsabstractThis article investigates the adaptive neural tracking control problem for a class of hyperbolic PDE with boundary actuator dynamics described by a set of nonlinear ordinary differential equations (ODEs). Particularly, the control input appears in the ODE subsystem with unknown nonlinearities requiring to be estimated and compensated, which makes the control task rather difficult. It is the first time to consider tracking control of such a class of systems, rendering our contributions essentially different from the existing literature that merely focus on the stabilization problem. By formulating a virtual exosystem to generate a reference trajectory, we propose a novel design of the adaptive geometric controller for the considered system where neural networks (NNs) are employed to approximately estimate nonlinearities, and finite and infinite-dimensional backstepping techniques are leveraged. Moreover, rigorously theoretical proofs based on the Lyapunov theory are provided to analyze the stability of the closed-loop system. Finally, we illustrate the results through two numerical simulations. Yu Xiao 0006, Yuan Yuan 0017, Chunhua Yang 0001, Biao Luo 0001, Xiaodong Xu 0002, Stevan Dubljevic |
IEEE Trans. Cybern. | 4 |
| 2024 | A Novel Zeroing Neurodynamic Method Based on Discrete Fuzzy Control System: Design, Analysis, and VerificationabstractConsidering the extensive research on zeroing neurodynamic (ZN), a self-adaptive and enhanced fixed-time convergent zeroing neurodynamic (SEFC-ZN) method for addressing time-variant problems is presented in this paper based on a discrete fuzzy matrix (DFM) design parameter and a novel advanced sign-bi-power activation function (NASbpAf). Due to the distinctive design of the DFM design parameter and NASbpAf, the proposed SEFC-ZN method possesses prominent self-adaptivity and enhanced fixed-time convergence. Specifically, the DFM design parameter is actually a matrix with all elements generated from a discrete fuzzy control system, so it can self-adaptively adjust the convergence rate of every error in the SEFC-ZN method resulting in the self-adaptivity. This feature is greatly different from the conventional scalar design parameters whose values are usually fixed or increase indefinitely and different errors in the ZN method can only be adjusted by the same design parameter. By summarizing the characteristic of the activation functions designed previously according to the SbpAf, it is found that keeping two terms of the SbpAf and adding extra terms can improve the performance of the ZN method. Thereout, built on the SbpAf, the NASbpAf is presented which can make the SEFC-ZN method realize the enhanced fixed-time convergence. Three theoretical analyses and proofs, together with relative corollaries, conclude the properties of the SEFC-ZN method and the advantages of the DFM design parameter and NASbpAf. A numerical experiment about solving time-variant nonlinear equations by the SEFC-ZN method and an application to the linear-quadratic optimal control strongly verify the proposed theory and method. Lei Jia 0001, Lin Xiao 0002, Yaonan Wang 0001, Jianhua Dai 0003, Biao Luo 0001 |
IEEE Trans. Fuzzy Syst. | 5 |
| 2024 | Decentralized Multiagent Reinforcement Learning Based State-of-Charge Balancing Strategy for Distributed Energy Storage SystemabstractState-of-charge (SoC) balancing in distributed energy storage systems (DESS) is crucial but challenging. Traditional deep reinforcement learning approaches struggle with real-world multiagent cooperation for SoC balance in these decentralized systems. To address these significant hurdles, this article pioneers an innovative fully-decentralized multiagent reinforcement learning (FDMARL) strategy, specifically tailored for DESS. First, the SoC balancing problem is formulated into a finite decentralized Markov decision process with action constraints derived from the microgrid context. Then, the average consensus algorithm is introduced for expanding the agent's observation and obtaining global information through a communication network. To improve multiagent system cooperation, a novel demand balance algorithm is proposed to refine agent actions for precise demand distribution. By the above modules, the FDMARL reveals outstanding performance in a fully-decentralized system without any expert experience or modeling. Finally, numerous simulations were carried out on pymgrid, (i.e., a Python-based open-source microgrid platform), which shows that: The FDMARL exceeds traditional reinforcement learning in decentralized cooperation, effectively manages large-scale systems with random states and demand/PV series, and maintains robustness with ESU failure or integration. Zheng Xiong, Biao Luo 0001, Bing-Chuan Wang, Xiaodong Xu 0002, Xiaodong Liu 0011, Tingwen Huang |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | ADP-Based Event-Triggered Constrained Optimal Control on Spatiotemporal Process: Application to Temperature Field in Roller KilnabstractThe precise control of the spatiotemporal process in a roller kiln is crucial in the production of Ni-Co-Mn layered cathode material of lithium-ion batteries. Since the product is extremely sensitive to temperature distribution, temperature field control is of great significance. In this article, an event-triggered optimal control (ETOC) method with input constraints for the temperature field is proposed, which takes up an important position in reducing the communication and computation costs. A nonquadratic cost function is adopted to describe the system performance with input constraints. First, we present the problem description of the temperature field event-triggered control, where this field is described by a partial differential equation (PDE). Then, the event-triggered condition is designed according to the information of system states and control inputs. On this basis, a framework of the event-triggered adaptive dynamic programming (ETADP) method that is based on the model reduction technology is proposed for the PDE system. A critic network is used to approach the optimal performance index by a neural network (NN) together with that an actor network is used to optimize the control strategy. Furthermore, an upper bound of the performance index and a lower bound of interexecution times, as well as the stabilities of the impulsive dynamic system and the closed-loop PDE system, are also proved. Simulation verification demonstrates the effectiveness of the proposed method. Binyan Li, Ning Chen 0009, Biao Luo 0001, Chunhua Yang 0001, Weihua Gui 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Human-in-the-Loop Reinforcement Learning in Continuous-Action SpaceabstractHuman-in-the-loop for reinforcement learning (RL) is usually employed to overcome the challenge of sample inefficiency, in which the human expert provides advice for the agent when necessary. The current human-in-the-loop RL (HRL) results mainly focus on discrete action space. In this article, we propose a Q value-dependent policy (QDP)-based HRL (QDP-HRL) algorithm for continuous action space. Considering the cognitive costs of human monitoring, the human expert only selectively gives advice in the early stage of agent learning, where the agent implements human-advised action instead. The QDP framework is adapted to the twin delayed deep deterministic policy gradient algorithm (TD3) in this article for the convenience of comparison with the state-of-the-art TD3. Specifically, the human expert in the QDP-HRL considers giving advice in the case that the difference between the twin Q -networks' output exceeds the maximum difference in the current queue. Moreover, to guide the update of the critic network, the advantage loss function is developed using expert experience and agent policy, which provides the learning direction for the QDP-HRL algorithm to some extent. To verify the effectiveness of QDP-HRL, the experiments are conducted on several continuous action space tasks in the OpenAI gym environment, and the results demonstrate that QDP-HRL greatly improves learning speed and performance. Biao Luo 0001, Zhengke Wu, Bing-Chuan Wang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Adaptive Neural Network-Based Event-Triggered SOC Observer With Application to a Stochastic Battery ModelabstractAccurate state of charge (SOC) is crucial to achieving safe, reliable, and efficient use of batteries. This article proposes an adaptive neural network (NN)-based event-triggered observer to estimate SOC. First, a stochastic battery equivalent circuit model (ECM) is established, where an adaptive NN is employed to approximate the unknown nonlinear part. The learning process of network weight is conducted online to observe the variations of model parameters and avoid time-consuming processes for parameter extraction. Besides, for the purpose of saving computational cost, an event-triggered mechanism (ETM) is employed in the weight updating law, which means the weights only update when it is necessary. Then, an adaptive radial basis function (RBF) NN-based SOC observer is designed, and its stability is proven by the Lyapunov theory. Moreover, the strictly positive lower bound of interevent time is derived, and undesirable Zeno behavior can be excluded. Finally, the accuracy and robustness of the proposed observer are evaluated by experiments and simulations. Results show that the proposed method can estimate SOC accurately in the presence of initial deviation and sensor noises. Chenyang Pan, Zhaoxia Peng, Guoguang Wen, Biao Luo 0001, Tingwen Huang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Predefined-Time Zeroing Neural Networks With Independent Prior Parameter for Solving Time-Varying Plural Lyapunov Tensor EquationabstractAs an extension of the Lyapunov equation, the time-varying plural Lyapunov tensor equation (TV-PLTE) can carry multidimensional data, which can be solved by zeroing neural network (ZNN) models effectively. However, existing ZNN models only focus on time-varying equations in field of real number. Besides, the upper bound of the settling time depends on the value of ZNN model parameters, which is a conservative estimation for existing ZNN models. Therefore, this article proposes a novel design formula for converting the upper bound of the settling time into an independent and directly modifiable prior parameter. On this basis, we design two new ZNN models called strong predefined-time convergence ZNN (SPTC-ZNN) and fast predefined (FP)-time convergence ZNN (FPTC-ZNN) models. The SPTC-ZNN model has a nonconservative upper bound of the settling time, and the FPTC-ZNN model has excellent convergence performance. The upper bound of the settling time and robustness of the SPTC-ZNN and FPTC-ZNN models are verified by theoretical analyses. Then, the effect of noise on the upper bound of settling time is discussed. The simulation results show that the SPTC-ZNN and FPTC-ZNN models have better comprehensive performance than existing ZNN models. Zhaohui Qi, Yingqiang Ning, Lin Xiao 0002, Yongjun He 0001, Biao Luo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Costate-Supplement ADP for Model-Free Optimal Control of Discrete-Time Nonlinear SystemsabstractIn this article, an adaptive dynamic programming (ADP) scheme utilizing a costate function is proposed for optimal control of unknown discrete-time nonlinear systems. The state-action data are obtained by interacting with the environment under the iterative scheme without any model information. In contrast with the traditional ADP scheme, the collected data in the proposed algorithm are generated with different policies, which improves data utilization in the learning process. In order to approximate the cost function more accurately and to achieve a better policy improvement direction in the case of insufficient data, a separate costate network is introduced to approximate the costate function under the actor-critic framework, and the costate is utilized as supplement information to estimate the cost function more precisely. Furthermore, convergence properties of the proposed algorithm are analyzed to demonstrate that the costate function plays a positive role in the convergence process of the cost function based on the alternate iteration mode of the costate function and cost function under a mild assumption. The uniformly ultimately bounded (UUB) property of all the variables is proven by using the Lyapunov approach. Finally, two numerical examples are presented to demonstrate the effectiveness and computation efficiency of the proposed method. Jun Ye 0007, Yougang Bian, Biao Luo 0001, Manjiang Hu, Rongjun Ding |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Neural Network-Based Robust Guaranteed Cost Control for Image-Based Visual Servoing of QuadrotorabstractIn this article, a neural network (NN)-based robust guaranteed cost control design is proposed for image-based visual servoing (IBVS) control of quadrotors. According to the dynamics of three subsystems (yaw, height, and lateral subsystems) derived from the quadrotor IBVS dynamic model, the main control design is to solve the robust control problem for the time-varying lateral subsystem with angle constraints and uncertain disturbances. Considering the system dynamics, a two-loop structure is conducted. The outer loop uses the linear quadratic regulator to solve the Riccati equation for the lateral image feature system, and the inner loop adopts the optimal robust guaranteed cost control to solve the lateral velocity system. For the lateral velocity system, the optimal robust control problem is transformed to solve the modified Hamilton-Jacobi-Bellman equation of the corresponding optimal control problem utilizing adaptive dynamic programming. The implementation is accomplished with the time-varying NN and the designed estimated weight update law. In addition, the stability and effectiveness are proved by the theoretic proof and simulations. Xinning Yi, Biao Luo 0001, Yuqian Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Concurrent Learning Robust Adaptive Fault Tolerant Boundary Regulation of Hyperbolic Distributed Parameter SystemsabstractThis article develops a robust adaptive boundary output regulation approach for a class of complex anticollocated hyperbolic partial differential equations subjected to multiplicative unknown faults in both the boundary sensor and actuator. The regulator design is based on the internal model principle, which amounts to stabilize a coupled cascade system, which consists of a finite-dimensional internal model driven by a hyperbolic distributed parameter system (DPS). To this end, a systematic sliding mode equipped with a backstepping approach is developed such that the robust state feedback control can be realized. Moreover, since the available information is a faulty boundary measurement at the right side point, state estimation is required. However, due to the presence of boundary unknown faults, we need to solve an issue of joint fault-state estimation. Restrictive persistent excitation conditions are usually required to guarantee the exact estimation of faults but are unrealistic in practice. To this end, a novel concurrent learning (CL) adaptive observer is proposed so that exponential convergence is obtained. It is the first time that the spirit of CL is introduced to the field of DPSs. Consequently, the observer-based adaptive boundary fault tolerant control scheme is developed, and rigorous theoretical analysis is given such that the exponential output regulation can be achieved. Finally, the effectiveness of the proposed methodology is demonstrated via comparative simulations. Yuan Yuan 0017, Xiaodong Xu 0002, Chunhua Yang 0001, Biao Luo 0001, Stevan Dubljevic |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | SMONAC: Supervised Multiobjective Negative Actor-Critic for Sequential RecommendationabstractRecent research shows that the sole accuracy metric may lead to the homogeneous and repetitive recommendations for users and affect the long-term user engagement. Multiobjective reinforcement learning (RL) is a promising method to achieve a good balance in multiple objectives, including accuracy, diversity, and novelty. However, it has two deficiencies: neglecting the updating of negative action values and limited regulation from the RL Q-networks to the (self-)supervised learning recommendation network. To address these disadvantages, we develop the supervised multiobjective negative actor-critic (SMONAC) algorithm, which includes a negative action update mechanism and multiobjective actor-critic mechanism. For the negative action update mechanism, several negative actions are randomly sampled during each time updating, and then, the offline RL approach is utilized to learn their values. For the multiobjective actor-critic mechanism, accuracy, diversity, and novelty values are integrated into the scalarized value, which is used to criticize the supervised learning recommendation network. The comparative experiments are conducted on two real-world datasets, and the results demonstrate that the developed SMONAC achieves tremendous performance promotion, especially for the metrics of diversity and novelty. Biao Luo 0001, Zhengke Wu, Tingwen Huang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | ADP-Based Decentralized Controller Design for Nonlinear Time-Delay Interconnected SystemsabstractThe control of industrial processes which can be generally described by nonlinear time-delay interconnected systems is very important. A decentralized control method based on adaptive dynamic programming (ADP) is proposed in this article, which can solve the control and stability problem for nonlinear time-delay interconnected systems using past state data and interconnection information. First, the form of nonlinear interconnected systems with state delay is given. Based on such systems, the subsystems are reconstructed by adding the upper bound of the interconnection part to the cost function. New cost functions are constructed which include a series of Hamilton–Jacobi–Bellman (HJB) equations with Lyapunov–Krasovskii (L–K) function, where the past state data and interconnection information are employed. Then, the control problem can be dealt with by solving the HJB equations. The HJB equations can be solved by the critic learning method based on ADP, where the critic neural networks (NNs) are applied to estimate the optimal cost function. The weights can be calculated by the gradient descent method, and the control law can be obtained. Subsequently, the stability of the closed-loop system is proved in the sense of uniformly ultimately bounded (UUB). Finally, simulation results verify the effectiveness of the proposed method by a nonlinear time-delay interconnected system and two-stage chemical reactors. Ning Chen 0009, Zeng Luo, Binyan Li, Biao Luo 0001, Chunhua Yang 0001, Weihua Gui 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2023 | Multi-vehicle Platoon Overtaking Using NoisyNet Multi-agent Deep Q-Learning Network
Lv He, Dongbo Zhang 0003, Tianmeng Hu, Biao Luo 0001 |
ICONIP (15) | 4 |
| 2023 | Mitigation of Voltage Violation for Battery Fast Charging Based on Data-Driven Optimization
Zheng Xiong, Biao Luo 0001, Bing-Chuan Wang |
ICONIP (14) | 2 |
| 2023 | Event-Triggered Constrained H∞ Control Using Concurrent Learning and ADP
Shan Xue 0004, Biao Luo 0001, Derong Liu 0001, Dongsheng Guo 0001 |
ICONIP (8) | 2 |
| 2023 | Adaptive dynamic programming based composite control for profile tracking with multiple constraints
Biao Luo 0001, Yuxin Liao |
Neurocomputing | 2 |
| 2023 | TSDTVOS: Target-guided spatiotemporal dual-stream transformers for video object segmentation
Yuqian Zhao 0001, Fan Zhang 0106, Biao Luo 0001, Lingli Yu, Baifan Chen, Chunhua Yang 0001, Weihua Gui 0001 |
Neurocomputing | 4 |
| 2023 | LJIR: Learning Joint-Action Intrinsic Reward in cooperative multi-agent reinforcement learning
Biao Luo 0001, Tianmeng Hu, Xiaodong Xu 0002 |
Neural Networks | 2 |
| 2023 | Predictive hierarchical reinforcement learning for path-efficient mapless navigation with moving target
Biao Luo 0001, Wei Song 0008, Chunhua Yang 0001 |
Neural Networks | 2 |
| 2023 | BASeg: Boundary aware semantic segmentation for autonomous driving
Xiaoyang Xiao, Yuqian Zhao 0001, Fan Zhang 0106, Biao Luo 0001, Ling-Li Yu, Baifan Chen, Chunhua Yang 0001 |
Neural Networks | 4 |
| 2023 | MO-MIX: Multi-Objective Multi-Agent Cooperative Decision-Making With Deep Reinforcement LearningabstractDeep reinforcement learning (RL) has been applied extensively to solve complex decision-making problems. In many real-world scenarios, tasks often have several conflicting objectives and may require multiple agents to cooperate, which are the multi-objective multi-agent decision-making problems. However, only few works have been conducted on this intersection. Existing approaches are limited to separate fields and can only handle multi-agent decision-making with a single objective, or multi-objective decision-making with a single agent. In this paper, we propose MO-MIX to solve the multi-objective multi-agent reinforcement learning (MOMARL) problem. Our approach is based on the centralized training with decentralized execution (CTDE) framework. A weight vector representing preference over the objectives is fed into the decentralized agent network as a condition for local action-value function estimation, while a mixing network with parallel architecture is used to estimate the joint action-value function. In addition, an exploration guide approach is applied to improve the uniformity of the final non-dominated solutions. Experiments demonstrate that the proposed method can effectively solve the multi-objective multi-agent cooperative decision-making problem and generate an approximation of the Pareto set. Our approach not only significantly outperforms the baseline method in all four kinds of evaluation metrics, but also requires less computational cost. Tianmeng Hu, Biao Luo 0001, Chunhua Yang 0001, Tingwen Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Event-Triggered Optimal Control for Temperature Field of Roller Kiln Based on Adaptive Dynamic ProgrammingabstractTemperature field control is crucial for the comprehensive performance of Ni-Co-Mn layered cathode material that is the most important part of lithium-ion batteries. Starting from the aspect of a class of distributed parameter systems described by highly dissipative partial differential equations (PDEs), an event-triggered optimal control (ETOC) method based on adaptive dynamic programming (ADP) for the roller kiln temperature field is proposed. First, we formulate the event-triggered control problem of the temperature field under the general framework of PDE systems. Then, an event-triggered condition is designed based on the stability of the closed-loop PDE system, which also guarantees the upper bound of the performance index. Subsequently, ADP technology is adopted to realize the ETOC, where the critic network is employed to approximate the optimal value function. Since the studied system can be regarded as an impulsive dynamic system with flow dynamics and jump dynamics simultaneously, the stability of the impulsive dynamic system combined with the ADP-based closed-loop PDE system is proved. Finally, simulation results on the temperature field verify the effectiveness of the proposed method. Ning Chen 0009, Binyan Li, Biao Luo 0001, Weihua Gui 0001, Chunhua Yang 0001 |
IEEE Trans. Cybern. | 3 |
| 2023 | Adaptive Fuzzy Boundary Observer Design for Uncertain Linear Coupled Hyperbolic Partial Differential Equation SystemsabstractJoint uncertainties and state estimation of a class of linear coupled hyperbolic partial differential equation systems in the presence of unstructured and structured uncertainties are studied in this paper. For unstructured uncertainties which are completely unknown, by employing Takagi-Sugeno fuzzy logic system to approximate the unstructured uncertainties, a novel adaptive fuzzy boundary observer is developed to estimate both unknown system states as well as unknown weights in the fuzzy logic system, and the estimation errors are ultimately bounded. Therein, in the design of the proposed observer, a set of swapping filters and infinite dimensional backstepping technique are combined. On the other hand, for structured uncertainties that can be described in a concrete parameterized form, the proposed method can easily achieve the exact estimation of weights and states to their true values. The rigorous proof is provided to show that the ultimately bounded estimation errors for the case of unstructured uncertainties and the exponential convergent estimation errors for the case of structured uncertainties can be realized. Finally, three illustrative simulations are carried out to show the feasibility and effectiveness of the developed methods in this paper. Linbin Teng, Yuan Yuan 0017, Biao Luo 0001, Chunhua Yang 0001, Stevan Dubljevic, Tingwen Huang, Xiaodong Xu 0002 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2023 | Pinning Spatiotemporal Sampled-Data Synchronization of Coupled Reaction-Diffusion Neural Networks Under Deception AttacksabstractIn this article, we investigate the pinning spatiotemporal sampled-data (SD) synchronization of coupled reaction-diffusion neural networks (CRDNNs), which are directed networks with SD in time and space communications under random deception attacks. In order to handle with the random deception attacks, we establish a directed CRDNN model, which respects the impacts of variable sampling and random deception attacks within a unified framework. Through the designed pinning spatiotemporal SD controller, sufficient conditions are obtained by linear matrix inequalities (LMIs) that guarantee the mean square exponential stability of the synchronization error system (SES) derived by utilizing inequality techniques, the stochastic analysis technique, and Lyapunov-Krasovskii functional (LKF). Finally, a numerical example is utilized to support the presented pinning spatiotemporal SD synchronization method. Zipeng Wang 0001, Huai-Ning Wu, Biao Luo 0001, Tingwen Huang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Video object segmentation based on multi-level target models and feature integration
Bocong Gao, Yuqian Zhao 0001, Fan Zhang 0106, Biao Luo 0001, Chunhua Yang 0001 |
Neurocomputing | 4 |
| 2022 | Neural network-based event-triggered integral reinforcement learning for constrained H∞ tracking control with experience replay
Shan Xue 0004, Biao Luo 0001, Derong Liu 0001, Ying Gao 0004 |
Neurocomputing | 2 |
| 2022 | Adaptive dynamic programming-based visual servoing control for quadrotor
Xinning Yi, Biao Luo 0001 |
Neurocomputing | 2 |
| 2022 | Event-triggered integral reinforcement learning for nonzero-sum games with asymmetric input saturation
Shan Xue 0004, Biao Luo 0001, Derong Liu 0001, Ying Gao 0004 |
Neural Networks | 2 |
| 2022 | Adaptive Appointed-Time Consensus Control of Networked Euler-Lagrange Systems With Connectivity PreservationabstractWith consideration of motion control performance and efficient information communication, the synchronization problem on communication connectivity preservation and guaranteed consensus performance for networked mechanical systems has attracted considerable attention in recent years. Different from the existing works, this article investigates a brand-new appointed-time consensus control approach for uncertain networked Euler-Lagrange systems on a directed graph via exploring the prescribed performance control structure. First, a two-layer prescribed performance envelope is formulated via using an appointed-time convergent function for position-related and velocity-related consensus errors, respectively. Then, a simple state-feedback virtual controller with online adaptive performance adjustment is developed to preserve the communication connectivity. Moreover, to guarantee the velocity consensus of the networked systems and improve the position consensus accuracy, an appointed-time adaptive controller is designed by applying the norm inequality to the system uncertainties and external disturbances. Compared to the existing consensus control approaches, the prime advantage of the proposed one is that the constraints generated from the communication ranges are approximated by a time-varying contractive performance envelope, wherein, the appointed-time convergence and steady-state tracking accuracy are preassigned a priori. Meanwhile, no repeated logarithmic error transformations are required in the relevant controller design, which implies that the complexity of the devised control laws has decreased dramatically. Finally, two groups of illustrative examples are organized to validate the effectiveness of the proposed consensus control approach. Caisheng Wei, Mingzhen Gui, Chengxi Zhang, Yuxin Liao, Ming-Zhe Dai, Biao Luo 0001 |
IEEE Trans. Cybern. | 6 |
| 2022 | Event-Triggered ADP for Tracking Control of Partially Unknown Constrained Uncertain SystemsabstractAn event-triggered adaptive dynamic programming (ADP) algorithm is developed in this article to solve the tracking control problem for partially unknown constrained uncertain systems. First, an augmented system is constructed, and the solution of the optimal tracking control problem of the uncertain system is transformed into an optimal regulation of the nominal augmented system with a discounted value function. The integral reinforcement learning is employed to avoid the requirement of augmented drift dynamics. Second, the event-triggered ADP is adopted for its implementation, where the learning of neural network weights not only relaxes the initial admissible control but also executes only when the predefined execution rule is violated. Third, the tracking error and the weight estimation error prove to be uniformly ultimately bounded, and the existence of a lower bound for the interexecution times is analyzed. Finally, simulation results demonstrate the effectiveness of the present event-triggered ADP method. Shan Xue 0004, Biao Luo 0001, Derong Liu 0001, Ying Gao 0004 |
IEEE Trans. Cybern. | 2 |
| 2022 | Optimal Tracking Control for Uncertain Nonlinear Systems With Prescribed Performance via Critic-Only ADPabstractThis article addresses the tracking control problem for a class of nonlinear systems described by Euler–Lagrange equations with uncertain system parameters. The proposed control scheme is capable of guaranteeing prescribed performance from two aspects: 1) a special parameter estimator with prescribed-performance properties is embedded in the control scheme. The estimator not only ensures the exponential convergence of the estimation errors under relaxed excitation conditions but also can restrict all estimates to predetermined bounds during the whole estimation process and 2) the proposed controller can strictly guarantee the user-defined performance specifications on tracking errors, including convergence rate, maximum overshoot, and residual set. More importantly, it has the optimizing ability for the tradeoff between performance and control cost. A state transformation method is employed to transform the constrained optimal tracking control problem to an unconstrained stationary optimal problem. Then, a critic-only adaptive dynamic programming algorithm is designed to approximate the solution of the Hamilton–Jacobi–Bellman equation and the corresponding optimal control policy. Uniformly ultimately bounded stability is guaranteed via a Lyapunov-based stability analysis. Finally, numerical simulation results demonstrate the effectiveness of the proposed control scheme. Hongyang Dong, Xiaowei Zhao 0001, Biao Luo 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | Constrained Event-Triggered H∞ Control Based on Adaptive Dynamic Programming With Concurrent LearningabstractIn this article, an event-triggered$H_{\infty }$control method is proposed based on adaptive dynamic programming (ADP) with concurrent learning for unknown continuous-time nonlinear systems with control constraints. First, a system identification technique based on neural networks (NNs) is adopted to identify completely unknown systems. Second, a critic NN is employed to approximate the value function. A novel weight updating rule is developed based on the event-triggered control law and time-triggered disturbance law, which reduces controller execution times and guarantees the stability of the system. Subsequently, concurrent learning is applied to the weight updating rule to relax the demand for the traditional persistence of excitation condition that is difficult to implement online. Finally, the comparison between the time-triggered method and event-triggered method in simulation demonstrates the effectiveness of the developed constrained event-triggered ADP method. Shan Xue 0004, Biao Luo 0001, Derong Liu 0001, Yin Yang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2021 | A Combinatorial Recommendation System Framework Based on Deep Reinforcement LearningabstractIn this paper, we propose a combinatorial recommendation system based on deep reinforcement learning (DRL). Specifically, we focus on the combinatorial product recommendation problem, design a consumer behavior simulator, and utilize deep reinforcement learning to find appropriate product combinations that can improve the sales of the platform. In order to replace real users and conduct massive real-time interactive training with recommendation system, user short- term characteristics are extracted from user-click-history by hierarchical recurrent neural network to train a user simulator. Through ingenious modeling, we transform the NP-hard combinatorial optimization problem into a multi-step sequential decision problem, and construct a framework of combinatorial recommendation system based on DRL. Relevant experiments show that the binary accuracy of the user simulator in predicting users’ consumption behavior reaches more than 80%, and the combinatorial DRL based recommendation system improves the platform sales and provides attractive combinations of products for customers. Biao Luo 0001, Tianmeng Hu, Yilin Wen 0003 |
IEEE BigData | 2 |
| 2021 | Integral reinforcement learning-based optimal output feedback control for linear continuous-time systems with input delay
Biao Luo 0001, Shan Xue 0004 |
Neurocomputing | 2 |
| 2021 | Periodic Event-Triggered Suboptimal Control With Sampling Period and Performance AnalysisabstractIn this paper, the periodic event-triggered suboptimal control (PETSOC) method is developed for continuous-time linear systems. Different from event-triggered control, where the triggering condition is monitored continuously, the developed PETSOC method only verifies the triggering condition periodically at sampling instants, which further reduces computational resources. First, the control gain of the PETSOC is designed based on the algebraic Riccati equation. Subsequently, the periodic event-triggering condition is proposed for the suboptimal control method, which is only verified at sampling instants periodically. The sampling period is determined and analyzed based on the continuous form of the triggering condition. Moreover, the stability and the performance upper bound of the closed-loop system with the PETSOC are proved. Finally, the effectiveness of the developed PETSOC is validated through simulation on an unstable batch reactor. Biao Luo 0001, Tingwen Huang, Derong Liu 0001 |
IEEE Trans. Cybern. | 1 |
| 2021 | Policy Iteration Q-Learning for Data-Based Two-Player Zero-Sum Game of Linear Discrete-Time SystemsabstractIn this article, the data-based two-player zero-sum game problem is considered for linear discrete-time systems. This problem theoretically depends on solving the discrete-time game algebraic Riccati equation (DTGARE), while it requires complete system dynamics. To avoid solving the DTGARE, the Q -function is introduced and a data-based policy iteration Q -learning (PIQL) algorithm is developed to learn the optimal Q -function by using data collected from the real system. Writing the Q -function in a quadratic form, it is proved that the PIQL algorithm is equivalent to the Newton iteration method in the Banach space by using the Fréchet derivative. Then, the convergence of the PIQL algorithm can be guaranteed by Kantorovich's theorem. For the realization of the PIQL algorithm, the off-policy learning scheme is proposed using real data rather than the system model. Finally, the efficiency of the developed data-based PIQL method is validated through simulation studies. Biao Luo 0001, Yin Yang 0001, Derong Liu 0001 |
IEEE Trans. Cybern. | 1 |
| 2021 | Event-Triggered Adaptive Dynamic Programming for Unmatched Uncertain Nonlinear Continuous-Time SystemsabstractIn this article, an event-triggered adaptive dynamic programming (ADP) method is proposed to solve the robust control problem of unmatched uncertain systems. First, the robust control problem with unmatched uncertainties is transformed into the optimal control design for an auxiliary system. Subsequently, to reduce controller executions and save computational and communication resources, an event-triggering mechanism is introduced. By using a critic neural network (NN) to approximate the value function, novel concurrent learning is developed to learn NN weights, which avoids the requirement of an initial admissible control and the persistence of excitation condition. Moreover, it is proven that the developed event-triggered ADP controller guarantees the robustness of the uncertain system and the uniform ultimate boundedness of the NN weight estimation error. Finally, by using the F-16 aircraft and the inverted pendulum with unmatched uncertainties as examples, the simulation results show the effectiveness of the developed event-triggered ADP method. Shan Xue 0004, Biao Luo 0001, Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Robust Exponential Synchronization for Memristor Neural Networks With Nonidentical Characteristics by Pinning ControlabstractIn this paper, robust exponential synchronization of memristor-based neural networks (MNNs) with nonidentical characteristics is investigated. Coefficient mismatch, time-varying delay mismatch, and activation function mismatch are considered between the drive and the response MNNs. Pinning control strategy is developed to realize robust exponential synchronization and the stability criteria is established by using the Lyapunov function method and differential inclusion theory. Furthermore, the stable region of controller parameters is computed to guarantee that the synchronization errors enter a predetermined error bound within given settling time. Finally, the effectiveness of the proposed methods is verified by the numerical simulations. The methods presented in this paper offer novel schemes for robust exponential synchronization of nonidentical MNNs. Yueheng Li, Biao Luo 0001, Derong Liu 0001, Yin Yang 0001, Zhanyu Yang |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2021 | Adaptive Dynamic Programming for Control: A Survey and Recent AdvancesabstractThis article reviews the recent development of adaptive dynamic programming (ADP) with applications in control. First, its applications in optimal regulation are introduced, and some skilled and efficient algorithms are presented. Next, the use of ADP to solve game problems, mainly nonzero-sum game problems, is elaborated. It is followed by applications in large-scale systems. Note that although the functions presented in this article are based on continuous-time systems, various applications of ADP in discrete-time systems are also analyzed. Moreover, in each section, not only some existing techniques are discussed, but also possible directions for future work are pointed out. Finally, some overall prospects for the future are given, followed by conclusions of this article. Through a comprehensive and complete investigation of its applications in many existing fields, this article fully demonstrates that the ADP intelligent control method is promising in today's artificial intelligence era. Furthermore, it also plays a significant role in promoting economic and social development. Derong Liu 0001, Shan Xue 0004, Bo Zhao 0015, Biao Luo 0001, Qinglai Wei |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2020 | Multi-scale local LSSVM based spatiotemporal modeling and optimal control for the goethite process
Jiayang Dai, Ning Chen 0009, Biao Luo 0001, Weihua Gui 0001, Chunhua Yang 0001 |
Neurocomputing | 3 |
| 2020 | Adaptive synchronization of memristor-based neural networks with discontinuous activations
Yueheng Li, Biao Luo 0001, Derong Liu 0001, Zhanyu Yang, Yunli Zhu |
Neurocomputing | 2 |
| 2020 | Adaptive dynamic programming based event-triggered control for unknown continuous-time nonlinear systems with input constraints
Shan Xue 0004, Biao Luo 0001, Derong Liu 0001, Yueheng Li |
Neurocomputing | 2 |
| 2020 | Integral reinforcement learning based event-triggered control with input saturation
Shan Xue 0004, Biao Luo 0001, Derong Liu 0001 |
Neural Networks | 2 |
| 2020 | Event-Triggered Optimal Control With Performance Guarantees Using Adaptive Dynamic ProgrammingabstractThis paper studies the problem of event-triggered optimal control (ETOC) for continuous-time nonlinear systems and proposes a novel event-triggering condition that enables designing ETOC methods directly based on the solution of the Hamilton-Jacobi-Bellman (HJB) equation. We provide formal performance guarantees by proving a predetermined upper bound. Moreover, we also prove the existence of a lower bound for interexecution time. For implementation purposes, an adaptive dynamic programming (ADP) method is developed to realize the ETOC using a critic neural network (NN) to approximate the value function of the HJB equation. Subsequently, we prove that semiglobal uniform ultimate boundedness can be guaranteed for states and NN weight errors with the ADP-based ETOC. Simulation results demonstrate the effectiveness of the developed ADP-based ETOC method. Biao Luo 0001, Yin Yang 0001, Derong Liu 0001, Huai-Ning Wu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Balancing Value Iteration and Policy Iteration for Discrete-Time ControlabstractThe optimal control problem of discrete-time nonlinear systems depends on the solution of the Bellman equation. In this paper, an adaptive reinforcement learning (RL) method is developed to solve the complex Bellman equation, which balances value iteration (VI) and policy iteration (PI). By adding a balance parameter, an adaptive RL integrates VI and PI together, which accelerates VI and avoids the need of an initial admissible control. The convergence of the adaptive RL is proved by showing that it converges to the Bellman equation. Subsequently, the adaptive RL is realized by using the neural network (NN) approximation for value function and a least-squares scheme is developed for updating NN weights. Then, the convergence of NN-based adaptive RL is proved with considering NN approximation error. To further improve its performance, an adaptive rule is developed for tuning balance parameter in adaptive RL iteration by iteration. Finally, the effectiveness of the adaptive RL is validated with simulation studies. Biao Luo 0001, Yin Yang 0001, Huai-Ning Wu, Tingwen Huang |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2020 | Event-Triggered Adaptive Dynamic Programming for Zero-Sum Game of Partially Unknown Continuous-Time Nonlinear SystemsabstractIn this paper, the zero-sum game problem is considered for partially unknown continuous-time nonlinear systems, and an event-triggered adaptive dynamic programming (ADP) method is developed to solve the problem. First, an identifier neural network (NN) and a critic NN are applied to approximate the drift system dynamics and the optimal value function, respectively. Subsequently, an event-triggered approach is developed based on ADP, which samples the states and updates the weights of NNs at the same time when the event-triggering condition is violated, such that the computational complexity is reduced. It is proved that the states and the error of NN weights are uniformly ultimately bounded. Finally, the effectiveness of the developed ADP-based event-triggered method is verified through simulation studies. Shan Xue 0004, Biao Luo 0001, Derong Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2020 | Adaptive Synchronization of Delayed Memristive Neural Networks With Unknown ParametersabstractIn this paper, the drive-response synchronization of the delayed memristive neural networks (MNNs) with unknown parameters is studied. With the realization of practical memristors, more and more researchers start to investigate MNNs, and their synchronization problem has became a hot topic. However, the majority of the existing works are based on the strict condition that the weights of MNNs are known and determined. When the parameters are unknown, the obtained results may be inapplicable. Thus, it is worthwhile to investigate the synchronization problem of the delayed MNNs with unknown parameters. Due to the parameter uncertainties of MNNs, a novel response system and an adaptive control method are proposed under different assumptions. The update laws for weights in the response system and the gains of adaptive controllers are developed to synchronize the proposed response system with the delayed MNNs. Furthermore, the proposed methods can be applied to various cases and the corresponding stability theories are established. Finally, the numerical simulations are conducted to verify the effectiveness of the developed methods. Zhanyu Yang, Biao Luo 0001, Derong Liu 0001, Yueheng Li |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2019 | Pinning Control for Synchronization of Drive-Response Memristive Neural Networks with Nonidentical ParametersabstractIn this paper, the asymptotic synchronization for drive-response memristive neural networks(MNNs) with nonidentical parameters is investigated. Parameter inconformity is ubiquitous between drive and response systems due to environmental or internal influence. However, the majority of previous results were based on the well-matched MNNs. Thus, it is meaningful to study the synchronization problem of MNNs with nonidentical parameters. First, coefficient mismatches are dealt within the framework of set-valued maps and differential inclusions. Furthermore, in order to reduce the control cost, a pinning control strategy is adopted to drive two nonidentical MNNs to achieve asymptotic synchronization. And the sufficient stability conditions are given based on Lyapunov functional method. Finally, the effectiveness of proposed pinning controller is verified by a numerical example. Yueheng Li, Biao Luo 0001, Derong Liu 0001, Zhanyu Yang |
IJCNN | 2 |
| 2019 | Output Tracking Control Based on Adaptive Dynamic Programming With Multistep Policy EvaluationabstractIn this paper, the optimal output tracking control problem of discrete-time nonlinear systems is considered. First, the augmented system is derived and the tracking control problem is converted to the regulation problem with a discounted performance index, which relies on the solution of the Bellman equation. It is known that policy iteration and value iteration are two classical algorithms for solving the Bellman equation. Through analysis of the two algorithms, it is found that policy iteration converges fast while requires an initial admissible control policy, and value iteration avoids the requirement of an initial admissible control policy but converges slowly. To achieve the tradeoff between policy iteration and value iteration, the multistep heuristic dynamic programming (MsHDP) is proposed by using multistep policy evaluation scheme. The convergence of MsHDP algorithm is proved by demonstrating that it converges to the solution of the Bellman equation. Subsequently, neural network-based actor-critic structure is developed to implement the MsHDP algorithm. The effectiveness and advantages of the developed MsHDP method are validated through comparative simulation studies. Biao Luo 0001, Derong Liu 0001, Tingwen Huang, Jiangjiang Liu 0004 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2018 | Event-Triggered Adaptive Dynamic Programming for Continuous-Time Nonlinear Two-Player Zero-Sum Game
Shan Xue 0004, Biao Luo 0001, Derong Liu 0001, Yueheng Li |
ICONIP (7) | 2 |
| 2018 | Robust synchronization of memristive neural networks with strong mismatch characteristics via pinning control
Yueheng Li, Biao Luo 0001, Derong Liu 0001, Zhanyu Yang |
Neurocomputing | 2 |
| 2018 | Reinforcement learning for robust adaptive control of partially unknown nonlinear systems subject to unmatched uncertainties
Xiong Yang 0001, Haibo He, Qinglai Wei, Biao Luo 0001 |
Inf. Sci. | 4 |
| 2018 | Adaptive Q-Learning for Data-Based Optimal Output Regulation With Experience ReplayabstractIn this paper, the data-based optimal output regulation problem of discrete-time systems is investigated. An off-policy adaptive -learning (QL) method is developed by using real system data without requiring the knowledge of system dynamics and the mathematical model of utility function. By introducing the -function, an off-policy adaptive QL algorithm is developed to learn the optimal -function. An adaptive parameter in the policy evaluation is used to achieve tradeoff between the current and future -functions. The convergence of adaptive QL algorithm is proved and the influence of the adaptive parameter is analyzed. To realize the adaptive QL algorithm with real system data, the actor-critic neural network (NN) structure is developed. The least-squares scheme and the batch gradient descent method are developed to update the critic and actor NN weights, respectively. The experience replay technique is employed in the learning process, which leads to simple and convenient implementation of the adaptive QL method. Finally, the effectiveness of the developed adaptive QL method is verified through numerical simulations. Biao Luo 0001, Yin Yang 0001, Derong Liu 0001 |
IEEE Trans. Cybern. | 1 |
| 2018 | Adaptive Constrained Optimal Control Design for Data-Based Nonlinear Discrete-Time Systems With Critic-Only StructureabstractReinforcement learning has proved to be a powerful tool to solve optimal control problems over the past few years. However, the data-based constrained optimal control problem of nonaffine nonlinear discrete-time systems has rarely been studied yet. To solve this problem, an adaptive optimal control approach is developed by using the value iteration-based Q-learning (VIQL) with the critic-only structure. Most of the existing constrained control methods require the use of a certain performance index and only suit for linear or affine nonlinear systems, which is unreasonable in practice. To overcome this problem, the system transformation is first introduced with the general performance index. Then, the constrained optimal control problem is converted to an unconstrained optimal control problem. By introducing the action-state value function, i.e., Q-function, the VIQL algorithm is proposed to learn the optimal Q-function of the data-based unconstrained optimal control problem. The convergence results of the VIQL algorithm are established with an easy-to-realize initial condition . To implement the VIQL algorithm, the critic-only structure is developed, where only one neural network is required to approximate the Q-function. The converged Q-function obtained from the critic-only VIQL method is employed to design the adaptive constrained optimal controller based on the gradient descent scheme. Finally, the effectiveness of the developed adaptive control method is tested on three examples with computer simulation. Biao Luo 0001, Derong Liu 0001, Huai-Ning Wu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Synchronization of Memristor-Based Time-Delayed Neural Networks via Pinning Control
Zhanyu Yang, Biao Luo 0001, Derong Liu 0001 |
ICONIP (3) | 2 |
| 2017 | Multi-step heuristic dynamic programming for optimal control of nonlinear discrete-time systems
Biao Luo 0001, Derong Liu 0001, Tingwen Huang, Xiong Yang 0001, Hongwen Ma |
Inf. Sci. | 1 |
| 2017 | Pinning synchronization of memristor-based neural networks with time-varying delays
Zhanyu Yang, Biao Luo 0001, Derong Liu 0001, Yueheng Li |
Neural Networks | 2 |
| 2017 | Policy Gradient Adaptive Dynamic Programming for Data-Based Optimal ControlabstractThe model-free optimal control problem of general discrete-time nonlinear systems is considered in this paper, and a data-based policy gradient adaptive dynamic programming (PGADP) algorithm is developed to design an adaptive optimal controller method. By using offline and online data rather than the mathematical system model, the PGADP algorithm improves control policy with a gradient descent scheme. The convergence of the PGADP algorithm is proved by demonstrating that the constructed Q -function sequence converges to the optimal Q -function. Based on the PGADP algorithm, the adaptive control method is developed with an actor-critic structure and the method of weighted residuals. Its convergence properties are analyzed, where the approximate Q -function converges to its optimum. Computer simulation results demonstrate the effectiveness of the PGADP-based adaptive control method. Biao Luo 0001, Derong Liu 0001, Huai-Ning Wu, Ding Wang 0001, Frank L. Lewis |
IEEE Trans. Cybern. | 1 |
| 2016 | Data-Based Optimal Tracking Control of Nonaffine Nonlinear Discrete-Time Systems
Biao Luo 0001, Derong Liu 0001, Tingwen Huang, Chao Li 0024 |
ICONIP (4) | 1 |
| 2016 | Data-based robust adaptive control for a class of unknown nonlinear constrained-input systems via integral reinforcement learning
Xiong Yang 0001, Derong Liu 0001, Biao Luo 0001, Chao Li 0024 |
Inf. Sci. | 3 |
| 2016 | Model-Free Optimal Tracking Control via Critic-Only Q-LearningabstractModel-free control is an important and promising topic in control fields, which has attracted extensive attention in the past few years. In this paper, we aim to solve the model-free optimal tracking control problem of nonaffine nonlinear discrete-time systems. A critic-only Q-learning (CoQL) method is developed, which learns the optimal tracking control from real system data, and thus avoids solving the tracking Hamilton-Jacobi-Bellman equation. First, the Q-learning algorithm is proposed based on the augmented system, and its convergence is established. Using only one neural network for approximating the Q-function, the CoQL method is developed to implement the Q-learning algorithm. Furthermore, the convergence of the CoQL method is proved with the consideration of neural network approximation error. With the convergent Q-function obtained from the CoQL method, the adaptive optimal tracking control is designed based on the gradient descent scheme. Finally, the effectiveness of the developed CoQL method is demonstrated through simulation studies. The developed CoQL method learns with off-policy data and implements with a critic-only structure, thus it is easy to realize and overcome the inadequate exploration problem. Biao Luo 0001, Derong Liu 0001, Tingwen Huang, Ding Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | An Approximate Optimal Control Approach for Robust Stabilization of a Class of Discrete-Time Nonlinear Systems With UncertaintiesabstractIn this correspondence paper, the robust stabilization of a class of discrete-time nonlinear systems with uncertainties is investigated by using an approximate optimal control approach. The robust control problem is transformed into an optimal control problem under some proper restrictions on the bound of the uncertainties. For the purpose of dealing with the transformed optimal control, the discrete-time generalized Hamilton-Jacobi-Bellman equation is introduced and then solved using the successive approximation method with neural network implementation. In addition, a numerical simulation is included to illustrate the effectiveness of the robust control strategy. Ding Wang 0001, Derong Liu 0001, Hongliang Li 0002, Biao Luo 0001, Hongwen Ma |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2015 | H ∞ Control Synthesis for Linear Parabolic PDE Systems with Model-Free Policy IterationabstractThe H ∞ control problem is considered for linear parabolic partial differential equation (PDE) systems with completely unknown system dynamics. We propose a model-free policy iteration (PI) method for learning the H ∞ control policy by using measured system data without system model information. First, a finite-dimensional system of ordinary differential equation (ODE) is derived, which accurately describes the dominant dynamics of the parabolic PDE system. Based on the finite-dimensional ODE model, the H ∞ control problem is reformulated, which is theoretically equivalent to solving an algebraic Riccati equation (ARE). To solve the ARE without system model information, we propose a least-square based model-free PI approach by using real system data. Finally, the simulation results demonstrate the effectiveness of the developed model-free PI method. Biao Luo 0001, Derong Liu 0001, Xiong Yang 0001, Hongwen Ma |
ISNN | 1 |
| 2015 | Reinforcement learning solution for HJB equation arising in constrained optimal control problem
Biao Luo 0001, Huai-Ning Wu, Tingwen Huang, Derong Liu 0001 |
Neural Networks | 1 |
| 2015 | Off-Policy Reinforcement Learning for H∞ Control DesignabstractThe H∞ control design problem is considered for nonlinear systems with unknown internal system model. It is known that the nonlinear H∞ control problem can be transformed into solving the so-called Hamilton-Jacobi-Isaacs (HJI) equation, which is a nonlinear partial differential equation that is generally impossible to be solved analytically. Even worse, model-based approaches cannot be used for approximately solving HJI equation, when the accurate system model is unavailable or costly to obtain in practice. To overcome these difficulties, an off-policy reinforcement leaning (RL) method is introduced to learn the solution of HJI equation from real system data instead of mathematical system model, and its convergence is proved. In the off-policy RL method, the system data can be generated with arbitrary policies rather than the evaluating policy, which is extremely important and promising for practical systems. For implementation purpose, a neural network (NN)-based actor-critic structure is employed and a least-square NN weight update algorithm is derived based on the method of weighted residuals. Finally, the developed NN-based off-policy RL method is tested on a linear F16 aircraft plant, and further applied to a rotational/translational actuator system. Biao Luo 0001, Huai-Ning Wu, Tingwen Huang |
IEEE Trans. Cybern. | 1 |
| 2015 | Data-Driven H∞ Control for Nonlinear Distributed Parameter SystemsabstractThe data-driven H∞ control problem of nonlinear distributed parameter systems is considered in this paper. An off-policy learning method is developed to learn the H∞ control policy from real system data rather than the mathematical model. First, Karhunen-Loève decomposition is used to compute the empirical eigenfunctions, which are then employed to derive a reduced-order model (ROM) of slow subsystem based on the singular perturbation theory. The H∞ control problem is reformulated based on the ROM, which can be transformed to solve the Hamilton-Jacobi-Isaacs (HJI) equation, theoretically. To learn the solution of the HJI equation from real system data, a data-driven off-policy learning approach is proposed based on the simultaneous policy update algorithm and its convergence is proved. For implementation purpose, a neural network (NN)- based action-critic structure is developed, where a critic NN and two action NNs are employed to approximate the value function, control, and disturbance policies, respectively. Subsequently, a least-square NN weight-tuning rule is derived with the method of weighted residuals. Finally, the developed data-driven off-policy learning approach is applied to a nonlinear diffusion-reaction process, and the obtained results demonstrate its effectiveness. Biao Luo 0001, Tingwen Huang, Huai-Ning Wu, Xiong Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | Adaptive Optimal Control of Highly Dissipative Nonlinear Spatially Distributed Processes With Neuro-Dynamic ProgrammingabstractHighly dissipative nonlinear partial differential equations (PDEs) are widely employed to describe the system dynamics of industrial spatially distributed processes (SDPs). In this paper, we consider the optimal control problem of the general highly dissipative SDPs, and propose an adaptive optimal control approach based on neuro-dynamic programming (NDP). Initially, Karhunen-Loève decomposition is employed to compute empirical eigenfunctions (EEFs) of the SDP based on the method of snapshots. These EEFs together with singular perturbation technique are then used to obtain a finite-dimensional slow subsystem of ordinary differential equations that accurately describes the dominant dynamics of the PDE system. Subsequently, the optimal control problem is reformulated on the basis of the slow subsystem, which is further converted to solve a Hamilton-Jacobi-Bellman (HJB) equation. HJB equation is a nonlinear PDE that has proven to be impossible to solve analytically. Thus, an adaptive optimal control method is developed via NDP that solves the HJB equation online using neural network (NN) for approximating the value function; and an online NN weight tuning law is proposed without requiring an initial stabilizing control policy. Moreover, by involving the NN estimation error, we prove that the original closed-loop PDE system with the adaptive optimal control policy is semiglobally uniformly ultimately bounded. Finally, the developed method is tested on a nonlinear diffusion-convection-reaction process and applied to a temperature cooling fin of high-speed aerospace vehicle, and the achieved results show its effectiveness. Biao Luo 0001, Huai-Ning Wu, Han-Xiong Li |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2013 | Simultaneous policy update algorithms for learning the solution of linear continuous-time H∞ state feedback control
Huai-Ning Wu, Biao Luo 0001 |
Inf. Sci. | 2 |
| 2012 | Neural Network Based Online Simultaneous Policy Update Algorithm for Solving the HJI Equation in Nonlinear H∞ ControlabstractIt is well known that the nonlinear H∞ state feedback control problem relies on the solution of the Hamilton-Jacobi-Isaacs (HJI) equation, which is a nonlinear partial differential equation that has proven to be impossible to solve analytically. In this paper, a neural network (NN)-based online simultaneous policy update algorithm (SPUA) is developed to solve the HJI equation, in which knowledge of internal system dynamics is not required. First, we propose an online SPUA which can be viewed as a reinforcement learning technique for two players to learn their optimal actions in an unknown environment. The proposed online SPUA updates control and disturbance policies simultaneously; thus, only one iterative loop is needed. Second, the convergence of the online SPUA is established by proving that it is mathematically equivalent to Newton's method for finding a fixed point in a Banach space. Third, we develop an actor-critic structure for the implementation of the online SPUA, in which only one critic NN is needed for approximating the cost function, and a least-square method is given for estimating the NN weight parameters. Finally, simulation studies are provided to demonstrate the effectiveness of the proposed algorithm. Huai-Ning Wu, Biao Luo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2012 | Approximate Optimal Control Design for Nonlinear One-Dimensional Parabolic PDE Systems Using Empirical Eigenfunctions and Neural NetworkabstractThis paper addresses the approximate optimal control problem for a class of parabolic partial differential equation (PDE) systems with nonlinear spatial differential operators. An approximate optimal control design method is proposed on the basis of the empirical eigenfunctions (EEFs) and neural network (NN). First, based on the data collected from the PDE system, the Karhunen-Loève decomposition is used to compute the EEFs. With those EEFs, the PDE system is formulated as a high-order ordinary differential equation (ODE) system. To further reduce its dimension, the singular perturbation (SP) technique is employed to derive a reduced-order model (ROM), which can accurately describe the dominant dynamics of the PDE system. Second, the Hamilton-Jacobi-Bellman (HJB) method is applied to synthesize an optimal controller based on the ROM, where the closed-loop asymptotic stability of the high-order ODE system can be guaranteed by the SP theory. By dividing the optimal control law into two parts, the linear part is obtained by solving an algebraic Riccati equation, and a new type of HJB-like equation is derived for designing the nonlinear part. Third, a control update strategy based on successive approximation is proposed to solve the HJB-like equation, and its convergence is proved. Furthermore, an NN approach is used to approximate the cost function. Finally, we apply the developed approximate optimal control method to a diffusion-reaction process with a nonlinear spatial operator, and the simulation results illustrate its effectiveness. Biao Luo 0001, Huai-Ning Wu |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2008 | A new methodology for searching robust Pareto optimal solutions with MOEAsabstractIt is of great importance for a solution with high robustness in the real application, not only with good quality. Searching for robust Pareto optimal solutions for multi-objective optimization problems (MOPs) is a challenge, no exception for multi-objective evolutionary algorithms (MOEAs). Recently, as one of the popular approach to search robust Pareto optimal solutions, “effective objective function” based MOEA (Eff-MOEA) can only find solutions which have average robustness and quality, but cannot find solutions which have the highest robustness and best quality. In this paper, we proposed a new methodology for robust Pareto optimal solutions and presented a novel MOEA named MOEA/R, which convert a multi-objective robust optimization problem (MROP) into a bi-objective optimization problem. Each of two objectives represents a sub-MOP, one of which optimizes solutions’ quality and another optimizes solutions’ robustness. Through the comparison and analysis between MOEA/R, Eff-MOEA and NSGA-II, the experimental results demonstrate that MOEA/R can acquire good purposes. The most important contribution of this paper is that MOEA/R explores a novel methodology for searching robust Pareto optimal solutions. Biao Luo 0001, Jinhua Zheng |
IEEE Congress on Evolutionary Computation | 1 |