EDBT 2026 Demo / reviewers in the wild / expert
Pengyi Li 0001
dblp:195/6948-1
· DBLP profile ↗
17ranked-venue papers
9as first author
17since 2021 · last 2026
0009-0009-8546-2346ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 9 first-author · 17 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tailoring knowledge for empowered cooperative actions in multi-agent reinforcement learning
Yihua Tan, Pengyi Li 0001 |
Neural Networks | 4 |
| 2026 | Signaling-Driven Incentive Communication for Enhanced Multiagent Reinforcement Learning in Dynamic EnvironmentsabstractCentralized training and decentralized execution (CTDE) frameworks in cooperative multiagent reinforcement learning (MARL) address nonstationarity and scalability in dynamic environments. However, coordination among agents remains challenging due to limited observability, often leading to inefficient exploration of policy spaces and increased communication overhead. Existing communication mechanisms partially alleviate these issues but typically add complexity without adapting to changing conditions. We propose the signaling-driven incentive communication (SDIC) framework, a novel approach that integrates Markov signaling games (MSGs) into CTDE to enable more efficient and targeted interagent communication. By integrating value-based methods with sparse communication, SDIC reduces unnecessary exchanges while generating tailored signals that enhance policy alignment and improve coordination. Furthermore, SDIC incorporates partner modeling, allowing agents to anticipate the behavior of others and thus strike an effective balance between communication efficiency and computational complexity. Our experimental results, including extensive evaluations in StarCraft II and SUMO traffic simulations, demonstrate SDIC's superior coordination, task success, and communication efficiency with manageable computational complexity. Ablation studies validate the critical roles of SDIC's components in reducing overhead and ensuring effective policy alignment. Kexing Peng, Pengyi Li 0001, Jianye Hao |
IEEE Trans. Cybern. | 2 |
| 2025 | R*: Efficient Reward Design via Reward Structure Evolution and Parameter Alignment Optimization with Large Language ModelsabstractReward functions are crucial for policy learning. Large Language Models (LLMs), with strong coding capabilities and valuable domain knowledge, provide an automated solution for high-quality reward design. However, code-based reward functions require precise guiding logic and parameter configurations within a vast design space, leading to low optimization efficiency. To address the challenges, we propose an efficient automated reward design framework, called R, which decomposes reward design into two parts: reward structure evolution and parameter alignment optimization. To design high-quality reward structures, R maintains a reward function population and modularizes the functional components. LLMs are employed as the mutation operator, and module-level crossover is proposed to facilitate efficient exploration and exploitation. To design more efficient reward parameters, R first leverages LLMs to generate multiple critic functions for trajectory comparison and annotation. Based on these critics, a voting mechanism is employed to collect the trajectory segments with high-confidence labels. These labeled segments are then used to refine the reward function parameters through preference learning. Experiments on diverse robotic control tasks demonstrate that R outperforms strong baselines in both reward design efficiency and quality, surpassing human-designed reward functions. Pengyi Li 0001, Jianye Hao, Hongyao Tang, Yifu Yuan, Jinbin Qiao, Zibin Dong, Yan Zheng 0002 |
ICML | 1 |
| 2025 | Enhancing Graph-based Coordination with Evolutionary Algorithms for Episodic Multi-agent Reinforcement Learning
Kexing Peng, Pengyi Li 0001, Jianye Hao |
AAMAS | 2 |
| 2025 | CORE: Collaborative Optimization with Reinforcement Learning and Evolutionary Algorithm for FloorplanningabstractFloorplanning is the initial step in the physical design process of Electronic Design Automation (EDA), directly influencing subsequent placement, routing, and final power of the chip.
However, the solution space in floorplanning is vast, and current algorithms often struggle to explore it sufficiently, making them prone to getting trapped in local optima. To achieve efficient floorplanning, we propose **CORE**, a general and effective solution optimization framework that synergizes Evolutionary Algorithms (EAs) and Reinforcement Learning (RL) for high-quality layout search and optimization.
Specifically, we propose the Clustering-based Diversified Evolutionary Search that directly perturbs layouts and evolves them based on novelty and performance. Additionally, we model the floorplanning problem as a sequential decision problem with B*-Tree representation and employ RL for efficient learning.
To efficiently coordinate EAs and RL, we propose the reinforcement-driven mechanism and evolution-guided mechanism.
The former accelerates population evolution through RL, while the latter guides RL learning through EAs. The experimental results on the MCNC and GSRC benchmarks demonstrate that CORE outperforms other strong baselines in terms of wirelength and area utilization metrics, achieving a 12.9\% improvement in wirelength. CORE represents the first evolutionary reinforcement learning (ERL) algorithm for floorplanning, surpassing existing RL-based methods. The code is available at https://github.com/yeshenpy/CORE. Pengyi Li 0001, Shixiong Kai, Jianye Hao, Ruizhe Zhong, Hongyao Tang, Zhentao Tang, Mingxuan Yuan, Junchi Yan |
NeurIPS | 1 |
| 2025 | LaRes: Evolutionary Reinforcement Learning with LLM-based Adaptive Reward SearchabstractThe integration of evolutionary algorithms (EAs) with reinforcement learning (RL) has shown superior performance compared to standalone methods. However, previous research focuses on exploration in policy parameter space, while overlooking the reward function search.
To bridge this gap, we propose **LaRes**, a novel hybrid framework that achieves efficient policy learning through reward function search. LaRes leverages large language models (LLMs) to generate the reward function population, guiding RL in policy learning. The reward functions are evaluated by the policy performance and improved through LLMs.
To improve sample efficiency, LaRes employs a shared experience buffer that collects experiences from all policies, with each experience containing rewards from all reward functions.
Upon reward function updates, the rewards of experiences are relabeled, enabling efficient use of historical data.
Furthermore, we introduce a Thompson sampling-based selection mechanism that enables more efficient elite interaction.
To prevent policy collapse when improving reward functions, we propose the reward scaling and parameter constraint mechanisms to efficiently coordinate reward search with policy learning.
Across both initialized and non-initialized settings, LaRes consistently achieves state-of-the-art performance, outperforming strong baselines in both sample efficiency and final performance.
The code is available at https://github.com/yeshenpy/LaRes. Pengyi Li 0001, Hongyao Tang, Jinbin Qiao, Yan Zheng 0002, Jianye Hao |
NeurIPS | 1 |
| 2025 | COLA: Towards Efficient Multi-Objective Reinforcement Learning with Conflict Objective Regularization in Latent SpaceabstractMany real-world control problems require continual policy adjustments to balance multiple objectives, which requires the acquisition of high-quality policies to cover diverse preferences. Multi-Objective Reinforcement Learning (MORL) provides a general framework to solve such problems. However, current MORL methods suffer from high sample complexity, primarily due to the neglect of efficient knowledge sharing and conflicts in optimization with different preferences.
To this end, this paper introduces a novel framework, Conflict
Objective Regularization in Latent Space (**COLA**).
To enable efficient knowledge sharing, COLA establishes a shared latent representation space for common knowledge, which can avoid redundant learning under different preferences. Besides, COLA introduces a regularization term for the value function to mitigate the negative effects of conflicting preferences on the value function approximation, thereby improving the accuracy of value estimation. The experimental results across various multi-objective continuous control tasks demonstrate the significant superiority of COLA over the state-of-the-art MORL baselines. Code is available at https://github.com/yeshenpy/COLA. Pengyi Li 0001, Hongyao Tang, Yifu Yuan, Jianye Hao, Zibin Dong, Yan Zheng 0002 |
NeurIPS | 1 |
| 2025 | Bridging Evolutionary Algorithms and Reinforcement Learning: A Comprehensive Survey on Hybrid AlgorithmsabstractEvolutionary reinforcement learning (ERL), which integrates the evolutionary algorithms (EAs) and reinforcement learning (RL) for optimization, has demonstrated remarkable performance advancements. By fusing both the approaches, ERL has emerged as a promising research direction. This survey offers a comprehensive overview of the diverse research branches in ERL. Specifically, we systematically summarize the recent advancements in related algorithms and identify three primary research directions: 1) EA-assisted optimization of RL; 2) RL-assisted optimization of EA; and 3) synergistic optimization of EA and RL. Following that, we conduct an in-depth analysis of each research direction, organizing multiple research branches. We elucidate the problems that each branch aims to tackle and how the integration of EAs and RL addresses these challenges. In conclusion, we discuss potential challenges and prospective future research directions across various research directions. To facilitate researchers in delving into ERL, we organize the algorithms and codes involved onhttps://github.com/yeshenpy/Awesome-Evolutionary-Reinforcement-Learning. Pengyi Li 0001, Jianye Hao, Hongyao Tang, Xian Fu, Yan Zheng 0002, Ke Tang 0001 |
IEEE Trans. Evol. Comput. | 1 |
| 2024 | Sample-Efficient Quality-Diversity by Cooperative CoevolutionabstractQuality-Diversity (QD) algorithms, as a subset of evolutionary algorithms, have emerged as a powerful optimization paradigm with the aim of generating a set of high-quality and diverse solutions. Although QD has demonstrated competitive performance in reinforcement learning, its low sample efficiency remains a significant impediment for real-world applications. Recent research has primarily focused on augmenting sample efficiency by refining selection and variation operators of QD. However, one of the less considered yet crucial factors is the inherently large-scale issue of the QD optimization problem. In this paper, we propose a novel Cooperative Coevolution QD (CCQD) framework, which decomposes a policy network naturally into two types of layers, corresponding to representation and decision respectively, and thus simplifies the problem significantly. The resulting two (representation and decision) subpopulations are coevolved cooperatively. CCQD can be implemented with different selection and variation operators. Experiments on several popular tasks within the QDAX suite demonstrate that an instantiation of CCQD achieves approximately a 200% improvement in sample efficiency. Ke Xue 0001, Ren-Jian Wang, Pengyi Li 0001, Dong Li 0016, Jianye Hao, Chao Qian 0001 |
ICLR | 3 |
| 2024 | EvoRainbow: Combining Improvements in Evolutionary Reinforcement Learning for Policy SearchabstractBoth Evolutionary Algorithms (EAs) and Reinforcement Learning (RL) have demonstrated powerful capabilities in policy search with different principles. A promising direction is to combine the respective strengths of both for efficient policy optimization. To this end, many works have proposed various mechanisms to integrate EAs and RL. However, it is still unclear which of these mechanisms are complementary and can be fully combined. In this paper, we revisit different mechanisms from five perspectives: 1) Interaction Mode, 2) Individual Architecture, 3) EAs and operators, 4) Impact of EA on RL, and 5) Fitness Surrogate and Usage. We evaluate the effectiveness of each mechanism and experimentally analyze the reasons for the more effective mechanisms. Using the most effective mechanisms, we develop EvoRainbow and EvoRainbow-Exp, which outperform strong baselines and provide state-of-the-art performance across various tasks with distinct characteristics. To promote community development, we release the code on https://github.com/yeshenpy/EvoRainbow. Pengyi Li 0001, Yan Zheng 0002, Hongyao Tang, Xian Fu, Jianye Hao |
ICML | 1 |
| 2024 | Value-Evolutionary-Based Reinforcement LearningabstractCombining Evolutionary Algorithms (EAs) and Reinforcement Learning (RL) for policy search has been proven to improve RL performance. However, previous works largely overlook value-based RL in favor of merging EAs with policy-based RL. This paper introduces Value-Evolutionary-Based Reinforcement Learning (VEB-RL) that focuses on the integration of EAs with value-based RL. The framework maintains a population of value functions instead of policies and leverages negative Temporal Difference error as the fitness metric for evolution. The metric is more sample-efficient for population evaluation than cumulative rewards and is closely associated with the accuracy of the value function approximation. Additionally, VEB-RL enables elites of the population to interact with the environment to offer high-quality samples for RL optimization, whereas the RL value function participates in the population's evolution in each generation. Experiments on MinAtar and Atari demonstrate the superiority of VEB-RL in significantly improving DQN, Rainbow, and SPR. Our code is available on https://github.com/yeshenpy/VEB-RL. Pengyi Li 0001, Jianye Hao, Hongyao Tang, Yan Zheng 0002, Fazl Barez |
ICML | 1 |
| 2024 | DiffuserLite: Towards Real-time Diffusion PlanningabstractDiffusion planning has been recognized as an effective decision-making paradigm in various domains. The capability of generating high-quality long-horizon trajectories makes it a promising research direction. However, existing diffusion planning methods suffer from low decision-making frequencies due to the expensive iterative sampling cost. To alleviate this, we introduce DiffuserLite, a super fast and lightweight diffusion planning framework, which employs a planning refinement process (PRP) to generate coarse-to-fine-grained trajectories, significantly reducing the modeling of redundant information and leading to notable increases in decision-making frequency. Our experimental results demonstrate that DiffuserLite achieves a decision-making frequency of $122.2$Hz ($112.7$x faster than predominant frameworks) and reaches state-of-the-art performance on D4RL, Robomimic, and FinRL benchmarks. In addition, DiffuserLite can also serve as a flexible plugin to increase the decision-making frequency of other diffusion planning algorithms, providing a structural design reference for future works. More details and visualizations are available at https://diffuserlite.github.io/. Zibin Dong, Jianye Hao, Yifu Yuan, Fei Ni 0001, Yitian Wang, Pengyi Li 0001, Yan Zheng 0002 |
NeurIPS | 6 |
| 2024 | CleanDiffuser: An Easy-to-use Modularized Library for Diffusion Models in Decision MakingabstractLeveraging the powerful generative capability of diffusion models (DMs) to build decision-making agents has achieved extensive success. However, there is still a demand for an easy-to-use and modularized open-source library that offers customized and efficient development for DM-based decision-making algorithms. In this work, we introduce CleanDiffuser, the first DM library specifically designed for decision-making algorithms. By revisiting the roles of DMs in the decision-making domain, we identify a set of essential sub-modules that constitute the core of CleanDiffuser, allowing for the implementation of various DM algorithms with simple and flexible building blocks. To demonstrate the reliability and flexibility of CleanDiffuser, we conduct comprehensive evaluations of various DM algorithms implemented with CleanDiffuser across an extensive range of tasks. The analytical experiments provide a wealth of valuable design choices and insights, reveal opportunities and challenges, and lay a solid groundwork for future research. CleanDiffuser will provide long-term support to the decision-making community, enhancing reproducibility and fostering the development of more robust solutions. Zibin Dong, Yifu Yuan, Jianye Hao, Fei Ni 0001, Yi Ma 0005, Pengyi Li 0001, Yan Zheng 0002 |
NeurIPS | 6 |
| 2023 | ERL-Re$^2$: Efficient Evolutionary Reinforcement Learning with Shared State Representation and Individual Policy Representation
Jianye Hao, Pengyi Li 0001, Hongyao Tang, Yan Zheng 0002, Xian Fu, Zhaopeng Meng |
ICLR | 2 |
| 2023 | RACE: Improve Multi-Agent Reinforcement Learning with Representation Asymmetry and Collaborative EvolutionabstractMulti-Agent Reinforcement Learning (MARL) has demonstrated its effectiveness in learning collaboration, but it often struggles with low-quality reward signals and high non-stationarity. In contrast, Evolutionary Algorithm (EA) has shown better convergence, robustness, and signal quality insensitivity. This paper introduces a hybrid framework, Representation Asymmetry and Collaboration Evolution (RACE), which combines EA and MARL for efficient collaboration. RACE maintains a MARL team and a population of EA teams. To enable efficient knowledge sharing and policy exploration, RACE decomposes the policies of different teams controlling the same agent into a shared nonlinear observation representation encoder and individual linear policy representations. To address the partial observation issue, we introduce Value-Aware Mutual Information Maximization to enhance the shared representation with useful information about superior global states. EA evolves the population using novel agent-level crossover and mutation operators, offering diverse experiences for MARL. Concurrently, MARL optimizes its policies and injects them into the population for evolution. The experiments on challenging continuous and discrete tasks demonstrate that RACE significantly improves the basic algorithms, consistently outperforming other algorithms. Our code is available at https://github.com/yeshenpy/RACE. Pengyi Li 0001, Jianye Hao, Hongyao Tang, Yan Zheng 0002, Xian Fu |
ICML | 1 |
| 2022 | HyAR: Addressing Discrete-Continuous Action Reinforcement Learning via Hybrid Action Representation
Hongyao Tang, Yan Zheng 0002, Jianye Hao, Pengyi Li 0001, Zhen Wang 0004, Zhaopeng Meng |
ICLR | 5 |
| 2022 | PMIC: Improving Multi-Agent Reinforcement Learning with Progressive Mutual Information CollaborationabstractLearning to collaborate is critical in Multi-Agent Reinforcement Learning (MARL). Previous works promote collaboration by maximizing the correlation of agents’ behaviors, which is typically characterized by Mutual Information (MI) in different forms. However, we reveal sub-optimal collaborative behaviors also emerge with strong correlations, and simply maximizing the MI can, surprisingly, hinder the learning towards better collaboration. To address this issue, we propose a novel MARL framework, called Progressive Mutual Information Collaboration (PMIC), for more effective MI-driven collaboration. PMIC uses a new collaboration criterion measured by the MI between global states and joint actions. Based on this criterion, the key idea of PMIC is maximizing the MI associated with superior collaborative behaviors and minimizing the MI associated with inferior ones. The two MI objectives play complementary roles by facilitating better collaborations while avoiding falling into sub-optimal ones. Experiments on a wide range of MARL benchmarks show the superior performance of PMIC compared with other algorithms. Pengyi Li 0001, Hongyao Tang, Tianpei Yang, Xiaotian Hao, Tong Sang, Yan Zheng 0002, Jianye Hao, Matthew E. Taylor, Wenyuan Tao, Zhen Wang 0004 |
ICML | 1 |