Renzhi Lu

dblp:118/2993 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0001-6845-1711ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 6 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Mirror Descent Safe Policy Optimization for Reinforcement Learning Agents
abstract
Embodied intelligence and related disciplines have identified several mechanisms that help embodied agents learn how to solve complex problems. Reinforcement learning (RL) is one of the most promising computational approaches toward enhancement of the learning-based problem-solving abilities of such agents. Given the recent rapid evolution of artificial intelligence, RL has become a keystone technology, accelerating scientific discoveries and also finding applications in many other domains. In RL, an agent collects data when interacting with the environment, which optimizes a policy ensuring a higher return. Further improvement requires more exploration of the action space. However, not all actions in that space are safe and acceptable. The exploration of an agent must be constrained. In this work, a novel mirror descent safe policy optimization (MDSPO) algorithm is proposed to ensure the safety of an RL agent. The algorithm leverages mirror descent optimization to maximize the return while satisfying the safety constraint. A novel optimization objective is formulated, and an innovative three-stage optimization strategy is employed-comprising gradient descent without the cost constraint, projection onto the nonparametric policy space with the cost constraint, and projection onto the parametric policy space. Compared to previous methods, MDSPO is a simple and easy to implement first-order approach, which does not impose a hard constraint on the trust region. Theoretical analysis of the MDSPO reveals a lower bound on return improvement and an upper bound on constraint violation at the time of each policy update. The numerical results obtained from two sets of different constrained locomotive experiments demonstrate that MDSPO improves the average return by about 12% and better satisfies the cost constraints than other state-of-the-art methods do. In a real-world obstacle avoidance experiment using an unmanned surface vessel, MDSPO both finds the optimal path and guarantees agent safety.
Renzhi Lu, Qingqing Xiong, Yifang Shi 0001, Dongrui Wu, Tao Yang 0003, Yaochu Jin, Lihua Xie 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 An EfficientNetV2 Deep Learning Framework for Adolescent Bone-age Prediction Using Hand Radiographs
abstract
Given the increasing demand for assessment of adolescent growth and development, prediction of adolescent hand-bone age has become important in the field of medical image analysis. Recent convolutional neural networks(CNNs), especially the EfficientNetV2 model, have significantly improved the accuracy of bone-age prediction. In this study, we trained a CNN model based on EfficientNetV2 to predict bone age based on 10,000 X-ray images of adolescents hand bones. The model extracted deep features from X-ray images, and after training, predicted bone age with remarkable accuracy. Moreover, when tested on real dataset, the model reduced the mean absolute error (MAE) of predicted bone age to 0.752 years. Our study confirms that deep learning methods aid in medical image analysis. Our bone-age prediction model is both objective and quantitative, and will find applications in clinical practice and when inferring the developmental cycle of minors.
TakMan Lo, Lixin Deng, Yizhu Tang, Kaip Tse, Jiatao Wu, Yuzhi Huang, Renzhi Lu
INDIN8
2025 Bone Age Assessment Using EfficientNet V2 and Multi-Model Fusion
abstract
Bone age assessment is a vital indicator for evaluating the growth and development of children and adolescents. Traditional manual evaluation methods are inefficient and highly subjective, failing to meet modern clinical demands. Recent advancements in deep learning, especially in image recognition and medical image analysis, have provided new opportunities for automated and precise bone age assessment. This paper presents a comprehensive study on deep learning-based bone age prediction, covering data preprocessing, model architecture, training strategies, and evaluation methods. The proposed method, leveraging data augmentation, U-Net segmentation, and the EfficientNet V2 architecture, achieves high accuracy and robustness in bone age prediction, demonstrating significant potential for clinical applications and providing valuable insights for future research.
Lixin Deng, Yizhu Tang, Kaip Tse, Jiatao Wu, Yuzhi Huang, TakMan Lo, Renzhi Lu
INDIN8
2025 RL-Based Joint Latency Optimization for Software-Defined Wireless Cloud Fog Automation
abstract
Cloud-Fog Automation (CFA) represents a fundamental paradigm to facilitate flexible deployment modalities of industrial applications. Adopting B5G/6G technologies, software-defined wireless CFA separates wireless networking and computing functions from proprietary hardware, thereby significantly improving the flexibility and scalability of industrial automation. Despite its advantages, optimizing the end-to-end performance in software-defined wireless CFA is technically challenging, due to computing resource fluctuations and the uncertainties imposed on communication. To tackle these challenges, this study presents a Reinforcement Learning (RL)-based latency optimization scheme for software-defined wireless CFA. Jointly considering the wireless channel conditions and computing resource constraints, it adaptively allocates B5G/6G Radio Access Network (RAN) resources for improved performance on end-to-end (E2E) latency. For validation, a full-stack software-defined CFA testbed was implemented utilizing open-source projects, e.g., srsRAN and Open5GS. The numerical data indicate that the proposed scheme surpasses the benchmark, simultaneously achieving lower end-to-end latency and improved performance stability.
Zhibo Pang, Renzhi Lu, Yuemin Ding
INDIN3
2025 Bone Age Prediction using a Convolutional Neural Network-based Regression Algorithm employing Attention-Directing and Cluster
abstract
Bone age assessment (BAA) is a critical research topic in pediatric radiology, with growing interest in developing automated BAA methods. This study proposes a bone age prediction model integrating cluster analysis and convolutional neural network (CNN) regression, further enhanced by a multi-scale attention mechanism to construct a "divide-and-focus" dual-driven deep learning framework. Targeting age-sensitive regional features in hand radiographs, we innovatively design an adaptive spatial attention module that achieves hierarchical anatomical feature enhancement through saliency detection of attention-guided regions of interest (ROI). The algorithm first uses multiconstrained clustering of K methods to generate age-specific subsets, followed by parallel execution on each subset: 1) attention-guided ROI segmentation and feature enhancement; 2) validation of the base CNN regression networks (including ResNet, DenseNet and EfficientNetV2); 3) set of cross-subset models with Bayesian-optimized weighting strategies for final prediction. By synergistically integrating the data distribution priors with attention-driven anatomical priors, the method delivers interpretable solutions when performing medical image regression tasks. The modular design ensures compatibility with mainstream CNN architectures. The method will aid pediatric growth monitoring and the diagnosis of endocrine disorders.
Tinghong Ye, Lixin Deng, Yizhu Tang, TakMan Lo, Kaip Tse, Jiatao Wu, Yuzhi Huang, Renzhi Lu
INDIN10
2025 A Novel YJQR-LSTM Model for Nonparametric Probabilistic Sustainable Agriculture Wind Power Forecasting Based on Intelligent IoT
abstract
Energy costs associated with the consumption of nonrenewable energy sources have become an important issue in improving the international competitiveness of agriculture. Wind power, as a renewable energy source, can replace nonrenewable energy sources to reduce energy costs and improve the sustainability of agricultural. However, the inherent intermittency, randomness, and volatility within weather conditions and wind speed present a substantial challenge in accurately predicting wind power generation. This work proposes a novel YJQR-LSTM algorithm that leverages Yeo-Johnson quantile regression (YJQR) with a long short-term memory (LSTM) network for nonparametric probabilistic forecasting of wind power generation via the Intelligent Internet of Things. First, an improved YJQR model based on the YJ transformation is designed to obtain a more precise characterization of wind uncertainty, providing a more flexible probability density function for wind power generation. Then, utilizing the unique structure of the LSTM network to learn the parameters of the YJQR model, temporal features can be extracted from time-series data. To mitigate the impact of outliers in the raw data on accuracy and improve computational efficiency, a novel logarithmic-likelihood function is developed as the loss function utilized in the training phase. The effectiveness of the proposed algorithm is validated using a real-world dataset from five wind farms from the Global Energy Forecasting Competition. Numerical results demonstrate that the algorithm provides more accurate wind power prediction results in complex wind power data environments, which is important for making full use of wind energy and thus reducing the consumption of nonrenewable energy in agriculture.
Jie Wang 0163, Junhui Jiang 0001, Xinlong Chen, Defu Cai, Yue Wu 0026, Renzhi Lu
IEEE Internet Things J.6
2025 Fuzzy Actor-Critic Reinforcement Learning for Unmanned Surface Vessels Flexible Tracking Formation Control
Renzhi Lu, Bohan Cen, Zhonghui Hu, Housheng Su, Lijun Zhu 0001, Hai-Tao Zhang
IEEE Trans. Fuzzy Syst.1
2025 SMA-PDPPO: Safe Multiagent Primal-Dual Deep Reinforcement Learning for Industrial Parks Energy Trading
abstract
Energy trading in industrial parks has great potential for reducing carbon emissions and lowering energy bills. This article proposes a safe multiagent deep reinforcement learning algorithm for optimizing the energy trading strategy in industrial parks to achieve less reliance on the main grid and save energy costs. Specifically, an industrial park that contains multiple industrial users with both thermal and electrical load requirements is considered, in which the different users can trade energy with each other and with the main grid based on their own strategies. Unlike the existing studies, the time-phased energy trading problem is transformed into a constrained partially observable Markov game, which models the industrial users and objectives of the buyers and sellers. Finally, a novel multiagent primal-dual proximal policy optimization algorithm that guarantees safety is developed to achieve the optimal trading strategies between the main grid and multiple users. Numerical simulations with real-world data demonstrate that the proposed algorithm allows higher total revenue for sellers and lower total costs for buyers in the park, limits each user's bid or offer to a relatively safe range, and increases the amount of electricity traded locally, while reducing trading with the grid.
Renzhi Lu, Tao Yang 0003, Ying Chen 0017, Dong Wang 0003, Xin Peng 0003
IEEE Trans. Ind. Informatics1
2025 A Novel Sequence-to-Sequence-Based Deep Learning Model for Multistep Load Forecasting
abstract
Load forecasting is critical to the task of energy management in power systems, for example, balancing supply and demand and minimizing energy transaction costs. There are many approaches used for load forecasting such as the support vector regression (SVR), the autoregressive integrated moving average (ARIMA), and neural networks, but most of these methods focus on single-step load forecasting, whereas multistep load forecasting can provide better insights for optimizing the energy resource allocation and assisting the decision-making process. In this work, a novel sequence-to-sequence (Seq2Seq)-based deep learning model based on a time series decomposition strategy for multistep load forecasting is proposed. The model consists of a series of basic blocks, each of which includes one encoder and two decoders; and all basic blocks are connected by residuals. In the inner of each basic block, the encoder is realized by temporal convolution network (TCN) for its benefit of parallel computing, and the decoder is implemented by long short-term memory (LSTM) neural network to predict and estimate time series. During the forecasting process, each basic block is forecasted individually. The final forecasted result is the aggregation of the predicted results in all basic blocks. Several cases within multiple real-world datasets are conducted to evaluate the performance of the proposed model. The results demonstrate that the proposed model achieves the best accuracy compared with several benchmark models.
Renzhi Lu, Ruichang Bai, Ruidong Li 0001, Lijun Zhu 0001, Feng Xiao 0002, Dong Wang 0003, Huaming Wu, Yuemin Ding
IEEE Trans. Neural Networks Learn. Syst.1
2025 Adaptive Optimal Surrounding Control of Multiple Unmanned Surface Vessels via Actor-Critic Reinforcement Learning
abstract
In this article, an optimal surrounding control algorithm is proposed for multiple unmanned surface vessels (USVs), in which actor-critic reinforcement learning (RL) is utilized to optimize the merging process. Specifically, the multiple-USV optimal surrounding control problem is first transformed into the Hamilton-Jacobi-Bellman (HJB) equation, which is difficult to solve due to its nonlinearity. An adaptive actor-critic RL control paradigm is then proposed to obtain the optimal surround strategy, wherein the Bellman residual error is utilized to construct the network update laws. Particularly, a virtual controller representing intermediate transitions and an actual controller operating on a dynamics model are employed as surrounding control solutions for second-order USVs; thus, optimal surrounding control of the USVs is guaranteed. In addition, the stability of the proposed controller is analyzed by means of Lyapunov theory functions. Finally, numerical simulation results demonstrate that the proposed actor-critic RL-based surrounding controller can achieve the surrounding objective while optimizing the evolution process and obtains 9.76% and 20.85% reduction in trajectory length and energy consumption compared with the existing controller.
Renzhi Lu, Xiaotao Wang, Yiyu Ding, Hai-Tao Zhang, Lijun Zhu 0001, Yong He 0003
IEEE Trans. Neural Networks Learn. Syst.1
2024 A Novel Hybrid-Action-Based Deep Reinforcement Learning for Industrial Energy Management
abstract
As environmental pollution becomes increasingly serious and industrial energy consumption continuously rises, an intelligent and efficient industrial energy management policy is urgently needed to reduce costs and maximize the benefits of industrial energy systems. However, modern industrial energy systems are characterized by hybrid industrial equipment actions, diverse objectives, and highly intermittent and stochastically distributed renewable energy sources. Therefore, efficient operation and control are difficult. This article presents a novel, model-free energy management policy using a hybrid action deep reinforcement learning algorithm for energy scheduling of industrial equipments operating in various modes. Specifically, the interaction process between the industrial energy management center and each equipment is modeled as a Markov decision process that minimizes the daily operating cost of the energy system and maximizes the revenue of the production equipment. Then, a double parameterized deep Q-networks that does not require an explicit environmental model is developed to learn the hybrid action signals using actor and critic networks, in which the double Q value mechanism avoids value overestimation and improves the algorithm efficiency. In addition, the policy gradient of the proposed algorithm is derived and its convergence proof is discussed. Finally, numerical studies are conducted using real-world data to evaluate algorithm performance and verify its effectiveness.
Renzhi Lu, Tao Yang 0003, Ying Chen 0017, Dong Wang 0003, Xin Peng 0003
IEEE Trans. Ind. Informatics1
2024 A Novel Data-Driven LSTM-SAF Model for Power Systems Transient Stability Assessment
abstract
Transient stability is an important metric for assessing the operational state of a power system. However, due to the inherent complexity of the power systems, it is difficult to achieve stable and precise transient stability assessment (TSA). This article proposes a novel data-driven long short-term memory with self-attention mechanism and focal loss function (LSTM-SAF) model to achieve a rapid and reliable TSA scheme. First, an improved wrapper approach involving a genetic algorithm is established to obtain concise and effective input features, which can enhance model performance and efficiency. Then, an LSTM network combined with a self-attention mechanism is developed to learn reliable TSA paradigms, in which the self-attention mechanism can further explore the information relationships of temporal features extracted from the LSTM, thereby significantly improving TSA accuracy. In addition, to resolve the lack of insufficient training related to sample imbalance, a new focal loss function is designed to guide model training. This article provides a complete TSA scheme (including offline training and online execution) that considers both assessment performance and response speed. The effectiveness of the proposed model is verified by the numerical testing results on IEEE 39 bus system, NPCC 140 bus system, IEEE 145 bus system and IEEE 300 bus system.
Zonghe Shao, Yuzhe Cao, Defu Cai, Renzhi Lu
IEEE Trans. Ind. Informatics6
2023 Cooperative Target-Surrounding Control of Unmanned Surface Vessels Based on MADDPG
abstract
This article proposes a multi-agent deep reinforce-ment learning algorithm to control a fleet of unmanned surface vessels (USVs) that encircle and capture sea targets. First, a simulation environment for USVs is established based on a dynamic model; two-dimensional control variables are used to control movements in three directions. Second, the multi-agent deep deterministic policy gradient (MADDPG) algorithm is employed to achieve intelligent control of the USVs, using a reward function based on certain prior knowledge. Finally, centralized training and decentralized execution are used to complete the offline learning of multiple agents, and continuous action decisions are made based on the observations of the USV sensors. Simulations demonstrate that the method can capture stationary moving sea targets with any number of multi-agents under disturbances, and exhibits strong robustness and practicality.
Taoman Li, Zihan Gan, Zexing Zhou, Xiaotao Wang, Renzhi Lu
IECON6
2023 Reward Shaping-Based Actor-Critic Deep Reinforcement Learning for Residential Energy Management
abstract
Residential energy consumption continues to climb steadily, requiring intelligent energy management strategies to reduce power system pressures and residential electricity bills. However, it is challenging to design such strategies due to the random nature of electricity pricing, appliance demand, and user behavior. This article presents a novel reward shaping (RS)-based actor–critic deep reinforcement learning (ACDRL) algorithm to manage the residential energy consumption profile with limited information about the uncertain factors. Specifically, the interaction between the energy management center and various residential loads is modeled as a Markov decision process that provides a fundamental mathematical framework to represent the decision-making in situations where outcomes are partially random and partially influenced by the decision-maker control signals, in which the key elements containing the agent, environment, state, action, and reward are carefully designed, and the electricity price is considered as a stochastic variable. An RS-ACDRL algorithm is then developed, incorporating both the actor and critic network and an RS mechanism, to learn the optimal energy consumption schedules. Several case studies involving real-world data are conducted to evaluate the performance of the proposed algorithm. Numerical results demonstrate that the proposed algorithm outperforms state-of-the-art RL methods in terms of learning speed, solution optimality, and cost reduction.
Renzhi Lu, Huaming Wu, Yuemin Ding, Dong Wang 0003, Hai-Tao Zhang
IEEE Trans. Ind. Informatics1
2022 Reward Shaping-based Double Deep Q-networks for Unmanned Surface Vessel Navigation and Obstacle Avoidance
abstract
In this paper, a method for navigation and obstacle avoidance of unmanned surface vessel (USV) based on reinforcement learning and reward shaping is proposed. This approach uses double deep Q networks (DDQN) to make decisions based on the continuous states observed from sensors in USV. In addition, a new reward function is designed based on prior knowledge to accelerate the convergence of the algorithm and improve the performance. For training the neural networks, a simulation platform is developed, in which a 3 degree of freedom mathematical model describes USV dynamic system and two-dimension actions are required to control USV. Simulation results on the platform demonstrate the DDQN hoists USV’s capabilities of navigation and obstacle avoidance, and reward shaping technique improves the speed of convergence.
Zihan Gan, Jinghong Zheng 0002, Renzhi Lu
IECON4