Luolin Xiong

dblp:317/6054 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
9since 2021 · last 2026
0009-0001-0142-7933ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HCPO: Hierarchical Conductor-Based Policy Optimization in Multi-Agent Reinforcement Learning
abstract
In cooperative Multi-Agent Reinforcement Learning (MARL), efficient exploration is crucial for optimizing the performance of joint policy. However, existing methods often update joint policies via independent agent exploration, without coordination among agents, which inherently constrains the expressive capacity and exploration of joint policies. To address this issue, we propose a conductor-based joint policy framework that directly enhances the expressive capacity of joint policies and coordinates exploration. In addition, we develop a Hierarchical Conductor-based Policy Optimization (HCPO) algorithm that instructs policy updates for the conductor and agents in a direction aligned with performance improvement. A rigorous theoretical guarantee further establishes the monotonicity of the joint policy optimization process. By deploying local conductors, HCPO retains centralized training benefits while eliminating inter-agent communication during execution. Finally, we evaluate HCPO on three challenging benchmarks: StarCraft II Multi-agent Challenge, Multi-agent MuJoCo, and Multi-agent Particle Environment. The results indicate that HCPO outperforms competitive MARL baselines regarding cooperative efficiency and stability.
Zejiao Liu, Junqi Tu, Yitian Hong, Luolin Xiong, Yaochu Jin, Yang Tang 0001, Fangfei Li
AAAI4
2026 Offline Deep Reinforcement Learning-Based Home Energy Management Systems With Heterogeneous EV Charging Load Models
abstract
With increasing penetration of Electric Vehicles (EVs) into the transportation system and smart electricity grid, there is a growing need for integrating them into Home Energy Management Systems (HEMS). This integration within HEMS introduces dynamic user behaviors and time-varying charging demand, thus posing challenges for the HEMS. To mitigate these challenges, this paper proposes a charging model for heterogeneous EVs that covers the range of Plug-in Hybrid EVs (PHEVs), Range-Extender EVs (REEVs) and Battery EVs (BEVs) with/without heat pumps. The proposed heterogeneous EV charging model considers weather conditions, estimated mileage and driver’s experience to describe the dynamic charging demand and the anxiety level influencing their behavior. To optimize the HEMS operation, minimizing the energy cost and ensuring comfort, this paper introduces an offline Deep Reinforcement Learning (DRL) algorithm which learns directly from pre-collected datasets, avoiding the cost and safety issues associated with continuous real-world interactions. The algorithm incorporates the Huber loss and a Q-quantile estimator to mitigate performance degradation from dataset anomalies such as data noise, sensor failure and human error, resulting in more robust HEMS optimization strategies. Experimental results demonstrate the method’s effectiveness in reducing total costs and analyze the performance of household devices with two different electricity rates.
Luolin Xiong, Yang Tang 0001, Kankar Bhattacharya, Mo-Yuen Chow, Feng Qian 0004
IEEE Trans. Circuits Syst. I Regul. Pap.1
2025 DRL-Based Distributed Coordination of ISO and DSOs in Bi-Level Electricity Markets
abstract
The increasing penetration of distributed energy resources has prompted distribution system operators (DSOs) at the retail electricity market level to coordinate with the independent system operator (ISO) at the wholesale market level, for greater benefits. However, interaction mechanisms between the ISO and DSOs, and impacts of prices and power injections, have not been adequately investigated in literature. This article proposes a distributed coordination framework for the ISO and DSOs across wholesale-retail (bi-level) electricity markets, considering their interactions more fairly. Moreover, to mitigate the challenges arising from the interdependence between the ISO and heterogeneous DSOs, a coupled training mechanism based on the response model is devised. This mechanism iteratively trains the ISO and DSOs by solely exchanging prices and power injections, ensuring the demand–supply balance at both retail and wholesale levels. In addition, a deep reinforcement learning algorithm is introduced for the three-stage iterative training process of heterogeneous agents. Results demonstrate the effectiveness of the proposed method and its advantages in terms of lowering energy prices, clearing of cheaper clean resources and thus, improving overall market efficiency.
Luolin Xiong, Anshul Goyal, Kankar Bhattacharya, Yang Tang 0001, Zhao Yang Dong, Feng Qian 0004, Venkata Balaji Thummalacherla
IEEE Trans. Ind. Informatics1
2024 Localization of False Data Injection Attacks in Smart Grids With Renewable Energy Integration via Spatiotemporal Network
abstract
The precise localization of false data injection attacks (FDIAs) is vital to ensure the stable operation of smart grids. However, the intermittency and uncertainty of renewable energy (RE) can lead to confusion with unknown FDIA. As a result, previous works encountered difficulties in extracting distinguishable spatiotemporal features to construct accurate behavior models, thereby affecting the effectiveness of the localization task. To address this challenge, we establish a more practical data set for FDIA localization that takes RE into account. Subsequently, we propose a spatiotemporal sequence analysis framework for the task. Specifically, we propose a factorized module to mitigate the impact of temporal fluctuations, which processes data sequence with down sampling and feature aggregation. Additionally, we introduce a fine-tuning matrix to take regional correlations of RE into consideration, where the weights of spatial information aggregation are adjusted. We evaluate the effectiveness of our approach through comprehensive case studies on IEEE 14-bus, IEEE 57-bus, and IEEE 118-bus standard test systems. The experimental results indicate that our method outperforms the compared methods by an average of 2.52% and 3% in terms of recall and F1-score, respectively.
Chensheng Liu, Luolin Xiong, Yang Tang 0001, Feng Qian 0004
IEEE Internet Things J.3
2024 Interpretable Deep Reinforcement Learning for Optimizing Heterogeneous Energy Storage Systems
abstract
Energy storage systems (ESS) are pivotal component in the energy market, serving as both energy suppliers and consumers. ESS operators can reap benefits from energy arbitrage by optimizing operations of storage equipment. To further enhance ESS flexibility within the energy market and improve renewable energy utilization, a heterogeneous photovoltaic-ESS (PV-ESS) is proposed, which leverages the unique characteristics of battery energy storage (BES) and hydrogen energy storage (HES). For scheduling tasks of the heterogeneous PV-ESS, a practical cost function plays a crucial role in guiding operator’s strategies to maximize benefits. We develop a comprehensive cost function that takes into account degradation, capital, and operation/maintenance costs to reflect real-world scenarios. Moreover, while numerous methods excel in optimizing ESS energy arbitrage, they often rely on black-box models with opaque decision-making processes, limiting practical applicability. To overcome this limitation and enable explainable scheduling strategies, a prototype-based policy network with inherent interpretability is introduced. This network employs human-designed prototypes to guide decision-making by comparing similarities between prototypical situations and encountered situations, which allows for naturally explained scheduling strategies. Comparative results across four distinct cases demonstrate the effectiveness and practicality of our proposed pre-hoc interpretable optimization method when contrasted with black-box models.
Luolin Xiong, Yang Tang 0001, Chensheng Liu, Ke Meng 0001, Zhao Yang Dong, Feng Qian 0004
IEEE Trans. Circuits Syst. I Regul. Pap.1
2023 A home energy management approach using decoupling value and policy in reinforcement learning
abstract
Considering the popularity of electric vehicles and the flexibility of household appliances, it is feasible to dispatch energy in home energy systems under dynamic electricity prices to optimize electricity cost and comfort residents. In this paper, a novel home energy management (HEM) approach is proposed based on a data-driven deep reinforcement learning method. First, to reveal the multiple uncertain factors affecting the charging behavior of electric vehicles (EVs), an improved mathematical model integrating driver’s experience, unexpected events, and traffic conditions is introduced to describe the dynamic energy demand of EVs in home energy systems. Second, a decoupled advantage actor-critic (DA2C) algorithm is presented to enhance the energy optimization performance by alleviating the overfitting problem caused by the shared policy and value networks. Furthermore, separate networks for the policy and value functions ensure the generalization of the proposed method in unseen scenarios. Finally, comprehensive experiments are carried out to compare the proposed approach with existing methods, and the results show that the proposed method can optimize electricity cost and consider the residential comfort level in different scenarios.
Luolin Xiong, Yang Tang 0001, Chensheng Liu, Ke Meng 0001, Zhao Yang Dong, Feng Qian 0004
Frontiers Inf. Technol. Electron. Eng.1
2023 Meta-Reinforcement Learning-Based Transferable Scheduling Strategy for Energy Management
abstract
In Home Energy Management System (HEMS), the scheduling of energy storage equipment and shiftable loads has been widely studied to reduce home energy costs. However, existing data-driven methods can hardly ensure the transferability amongst different tasks, such as customers with diverse preferences, appliances, and fluctuations of renewable energy in different seasons. This paper designs a transferable scheduling strategy for HEMS with different tasks utilizing a Meta-Reinforcement Learning (Meta-RL) framework, which can alleviate data dependence and massive training time for other data-driven methods. Specifically, a more practical and complete demand response scenario of HEMS is considered in the proposed Meta-RL framework, where customers with distinct electricity preferences, as well as fluctuating renewable energy in different seasons are taken into consideration. An inner level and an outer level are integrated in the proposed Meta-RL-based transferable scheduling strategy, where the inner and the outer level ensure the learning speed and appropriate initial model parameters, respectively. Moreover, Long Short-Term Memory (LSTM) is presented to extract the features from historical actions and rewards, which can overcome the challenges brought by the uncertainties of renewable energy and the customers’ loads, and enhance the robustness of scheduling strategies. A set of experiments conducted on practical data of Australia’s electricity network verify the performance of the transferable scheduling strategy.
Luolin Xiong, Yang Tang 0001, Chensheng Liu, Ke Meng 0001, Zhao Yang Dong, Feng Qian 0004
IEEE Trans. Circuits Syst. I Regul. Pap.1
2023 Molecular Joint Representation Learning via Multi-Modal Information of SMILES and Graphs
abstract
In recent years, artificial intelligence has played an important role on accelerating the whole process of drug discovery. Various of molecular representation schemes of different modals (e.g., textual sequence or graph) are developed. By digitally encoding them, different chemical information can be learned through corresponding network structures. Molecular graphs and Simplified Molecular Input Line Entry System (SMILES) are popular means for molecular representation learning in current. Previous works have done attempts by combining both of them to solve the problem of specific information loss in single-modal representation on various tasks. To further fusing such multi-modal imformation, the correspondence between learned chemical feature from different representation should be considered. To realize this, we propose a novel framework of molecular joint representation learning via Multi-Modal information of SMILES and molecular Graphs, called MMSG. We improve the self-attention mechanism by introducing bond-level graph representation as attention bias in Transformer to reinforce feature correspondence between multi-modal information. We further propose a Bidirectional Message Communication Graph Neural Network (BMC GNN) to strengthen the information flow aggregated from graphs for further combination. Numerous experiments on public property prediction datasets have demonstrated the effectiveness of our model.
Yang Tang 0001, Qiyu Sun, Luolin Xiong
IEEE ACM Trans. Comput. Biol. Bioinform.4
2022 A Two-Level Energy Management Strategy for Multi-Microgrid Systems With Interval Prediction and Reinforcement Learning
abstract
Setting retail electricity prices is one of the significant strategies for energy management of multi-microgrid (MMG) systems integrated with renewable energy. Nevertheless, the need of privacy preservation, the uncertainties of renewable energy and loads, as well as the time-varying scenarios, bring challenges for pricing problems. In this paper, a two-level pricing framework is proposed based on interval predictions and model-free reinforcement learning to address these challenges. In particular, at the higher level, the distribution system operator (DSO) is viewed as an agent, which sets retail electricity prices without detailed user information for privacy protection to maximize the total revenue from selling energy with reinforcement learning. For time-varying scenarios with intermittent photovoltaic power generation and diverse loads, a differentiable trust region layer is considered in reinforcement learning to improve the robustness of the policy updating process. While at the lower level, operators in microgrids solve three-phase unbalanced optimal power flow (OPF) problems to minimize generation cost and network power loss. Additionally, to deal with the challenges from the uncertainties of renewable power generation and user loads, interval predictions are chosen to quantify prediction errors and improve the flexibility of pricing policies. Finally, a set of experiments are conducted to validate the effectiveness of the proposed method for pricing problems in MMG systems.
Luolin Xiong, Yang Tang 0001, Hangyue Liu, Ke Meng 0001, Zhao Yang Dong, Feng Qian 0004
IEEE Trans. Circuits Syst. I Regul. Pap.1