Yunpeng Qing

dblp:333/0812 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0001-7376-9847ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 80% Multi-agent systems · 20%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Energy systems and smart grids · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.622025
CADP: Towards Better Centralized Learning for Decentralized Execution in MARL · IJCAI 2025
Temporal Prototype-Aware Learning for Active Voltage Control on Power Distribution Networks · KDD 2024
Knowledge, reasoning and agents › Multi-agent systems
agent communication
0.912025
CADP: Towards Better Centralized Learning for Decentralized Execution in MARL · IJCAI 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning › decentralized multi-agent reinforcement learning
centralized training with decentralized execution
0.912025
CADP: Towards Better Centralized Learning for Decentralized Execution in MARL · IJCAI 2025
Energy systems and smart grids
power system operation
0.912025
Powerformer: A Section-adaptive Transformer for Power Flow Adjustment · KDD (1) 2025
Machine learning › Reinforcement learning
offline reinforcement learning
0.812024
A2PO: Towards Effective Offline Reinforcement Learning from an Advantage-aware Perspective · NeurIPS 2024
Energy systems and smart grids
power distribution network
0.812024
Temporal Prototype-Aware Learning for Active Voltage Control on Power Distribution Networks · KDD 2024
Energy systems and smart grids › power system control
voltage control
0.812024
Temporal Prototype-Aware Learning for Active Voltage Control on Power Distribution Networks · KDD 2024
Machine learning › Reinforcement learning
policy adaptation
0.212024
Temporal Prototype-Aware Learning for Active Voltage Control on Power Distribution Networks · KDD 2024

Methods — techniques the papers use, named apart from their topics

transformer · 2.4prototype learning · 1.5meta-learning · 1.5model pruning · 0.9graph neural network · 0.9decentralized pruning · 0.9centralized advising · 0.9attention mechanism · 0.9conditional variational autoencoder · 0.8advantage weighting · 0.8
YearPublicationVenuePosition
2025 CADP: Towards Better Centralized Learning for Decentralized Execution in MARL
Yihe Zhou, Shunyu Liu 0001, Yunpeng Qing, Tongya Zheng, Kai-Xuan Chen 0001, Jie Song 0011, Mingli Song
AAMAS3
2025 CADP: Towards Better Centralized Learning for Decentralized Execution in MARL
abstract
Centralized Training with Decentralized Execution (CTDE) has recently emerged as a popular framework for cooperative Multi-Agent Reinforcement Learning (MARL), where agents can use additional global state information to guide training in a centralized way and make their own decisions only based on decentralized local policies. Despite the encouraging results achieved, CTDE makes an independence assumption on agent policies, which limits agents from adopting global cooperative information from each other during centralized training. Therefore, we argue that the existing CTDE framework cannot fully utilize global information for training, leading to an inefficient joint exploration and perception, which can degrade the final performance. In this paper, we introduce a novel Centralized Advising and Decentralized Pruning (CADP) framework for MARL, that not only enables an efficacious message exchange among agents during training but also guarantees the independent policies for decentralized execution. Firstly, CADP endows agents the explicit communication channel to seek and take advice from different agents for more centralized training. To further ensure the decentralized execution, we propose a smooth model pruning mechanism to progressively constrain the agent communication into a closed one without degradation in agent cooperation capability. Empirical evaluations on different benchmarks and across various MARL backbones demonstrate that the proposed framework achieves superior performance compared with the state-of-the-art counterparts. Our code is available at https://github.com/zyh1999/CADP
Yihe Zhou, Shunyu Liu 0001, Yunpeng Qing, Tongya Zheng, Kai-Xuan Chen 0001, Jie Song 0011, Mingli Song
IJCAI3
2025 Powerformer: A Section-adaptive Transformer for Power Flow Adjustment
abstract
In this paper, we present a novel transformer architecture tailored for learning robust power system state representations, which strives to optimize power dispatch for the power flow adjustment across different transmission sections. Specifically, our proposed approach, named Powerformer, develops a dedicated section-adaptive attention mechanism, separating itself from the self-attention employed in conventional transformers. This mechanism effectively integrates power system states with transmission section information, which facilitates the development of robust state representations. Furthermore, by considering the graph topology of power system and the electrical attributes of bus nodes, we introduce two customized strategies to further enhance the expressiveness: graph neural network propagation and multi-factor attention mechanism. Extensive evaluations are conducted on three power system scenarios, including the IEEE 118-bus system, a realistic China 300-bus system, and a large-scale European system with 9241 buses, where Powerformer demonstrates its superior performance over several popular baseline methods. The code is available at: https://github.com/Cra2yDavid/Powerformer
Kai-Xuan Chen 0001, Shunyu Liu 0001, Yaoquan Wei, Yihe Zhou, Yunpeng Qing, Jie Song 0011, Mingli Song
KDD (1)6
2025 Curricular Subgoals for Inverse Reinforcement Learning
abstract
Inverse Reinforcement Learning (IRL) aims to reconstruct the reward function from expert demonstrations to facilitate policy learning, and has demonstrated its remarkable success in imitation learning. To promote expert-like behavior, existing IRL methods mainly focus on learning global reward functions to minimize the trajectory difference between the imitator and the expert. However, these global designs are still limited by the redundant noise and error propagation problems, leading to the unsuitable reward assignment and thus downgrading the agent capability in complex multi-stage tasks. In this paper, we propose a novel Curricular Subgoal-based Inverse Reinforcement Learning (CSIRL) framework, that explicitly disentangles one task with several local subgoals to guide agent imitation. Specifically, CSIRL firstly introduces decision uncertainty of the trained agent over expert trajectories to dynamically select specific states as subgoals, which directly determines the exploration boundary of different task stages. To further acquire local reward functions for each stage, we customize a meta-imitation objective based on these curricular subgoals to train an intrinsic reward generator. Experiments on the D4RL and autonomous driving benchmarks demonstrate that the proposed methods yields results superior to the state-of-the-art counterparts, as well as better interpretability. Our code is publicly available athttps://github.com/Plankson/CSIRL.
Shunyu Liu 0001, Yunpeng Qing, Shuqi Xu, Jingyuan Cong, Tianhao Chen, Yun-Fu Liu, Mingli Song
IEEE Trans. Intell. Transp. Syst.2
2024 Temporal Prototype-Aware Learning for Active Voltage Control on Power Distribution Networks
abstract
Active Voltage Control (AVC) on the Power Distribution Networks (PDNs) aims to stabilize the voltage levels to ensure efficient and reliable operation of power systems. With the increasing integration of distributed energy resources, recent efforts have explored employing multi-agent reinforcement learning (MARL) techniques to realize effective AVC. Existing methods mainly focus on the acquisition of short-term AVC strategies, i.e., only learning AVC within the short-term training trajectories of a singular diurnal cycle. However, due to the dynamic nature of load demands and renewable energy, the operation states of real-world PDNs may exhibit significant distribution shifts across varying timescales (e.g., daily and seasonal changes). This can render those short-term strategies suboptimal or even obsolete when performing continuous AVC over extended periods. In this paper, we propose a novel temporal prototype-aware learning method, abbreviated as TPA, to learn time-adaptive AVC under short-term training trajectories. At the heart of TPA are two complementary components, namely multi-scale dynamic encoder and temporal prototype-aware policy, that can be readily incorporated into various MARL methods. The former component integrates a stacked transformer network to learn underlying temporal dependencies at different timescales of the PDNs, while the latter implements a learnable prototype matching mechanism to construct a dedicated AVC policy that can dynamically adapt to the evolving operation states. Experimental results on the AVC benchmark with different PDN sizes demonstrate that the proposed TPA surpasses the state-of-the-art counterparts not only in terms of control performance but also by offering model transferability. Our code is available at https://github.com/Canyizl/TPA-for-AVC.
Feiyang Xu, Shunyu Liu 0001, Yunpeng Qing, Yihe Zhou, Mingli Song
KDD3
2024 A2PO: Towards Effective Offline Reinforcement Learning from an Advantage-aware Perspective
abstract
Offline reinforcement learning endeavors to leverage offline datasets to craft effective agent policy without online interaction, which imposes proper conservative constraints with the support of behavior policies to tackle the out-of-distribution problem. However, existing works often suffer from the constraint conflict issue when offline datasets are collected from multiple behavior policies, i.e., different behavior policies may exhibit inconsistent actions with distinct returns across the state space. To remedy this issue, recent advantage-weighted methods prioritize samples with high advantage values for agent training while inevitably ignoring the diversity of behavior policy. In this paper, we introduce a novel Advantage-Aware Policy Optimization (A2PO) method to explicitly construct advantage-aware policy constraints for offline learning under mixed-quality datasets. Specifically, A2PO employs a conditional variational auto-encoder to disentangle the action distributions of intertwined behavior policies by modeling the advantage values of all training data as conditional variables. Then the agent can follow such disentangled action distribution constraints to optimize the advantage-aware policy towards high advantage values. Extensive experiments conducted on both the single-quality and mixed-quality datasets of the D4RL benchmark demonstrate that A2PO yields results superior to the counterparts. Our code is available at https://github.com/Plankson/A2PO.
Yunpeng Qing, Shunyu Liu 0001, Jingyuan Cong, Kai-Xuan Chen 0001, Yihe Zhou, Mingli Song
NeurIPS1