Lipeng Wan 0003

dblp:377/4923 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 65% Transfer learning and domain adaptation · 8% Motion planning and robot control · 8%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
2.742024
Grounded Answers for Multi-agent Decision-making Problem through Generative World Model · NeurIPS 2024
Imagine, Initialize, and Explore: An Effective Exploration Method in Multi-Agent Reinforcement Learning · AAAI 2024
Deep Hierarchical Communication Graph in Multi-Agent Reinforcement Learning · IJCAI 2023
Machine learning › Reinforcement learning › offline reinforcement learning
offline-to-online reinforcement learning
0.912025
State Revisit and Re-explore: Bridging Sim-to-Real Gaps in Offline-and-Online Reinforcement Learning with An Imperfect Simulator · IJCAI 2025
Robotics › Motion planning and robot control › robot learning
robot skill learning
0.912025
State Revisit and Re-explore: Bridging Sim-to-Real Gaps in Offline-and-Online Reinforcement Learning with An Imperfect Simulator · IJCAI 2025
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
0.912025
State Revisit and Re-explore: Bridging Sim-to-Real Gaps in Offline-and-Online Reinforcement Learning with An Imperfect Simulator · IJCAI 2025
Machine learning › Generative modeling
autoregressive model
0.812024
Grounded Answers for Multi-agent Decision-making Problem through Generative World Model · NeurIPS 2024
Machine learning › Reinforcement learning › multi-agent reinforcement learning
multi-agent exploration
0.812024
Imagine, Initialize, and Explore: An Effective Exploration Method in Multi-Agent Reinforcement Learning · AAAI 2024
Machine learning › Reinforcement learning › model-based reinforcement learning
world model
0.812024
Grounded Answers for Multi-agent Decision-making Problem through Generative World Model · NeurIPS 2024
Machine learning › Reinforcement learning › multi-agent reinforcement learning › value-based multi-agent reinforcement learning
value decomposition
0.612022
Greedy based Value Representation for Optimal Coordination in Multi-agent Reinforcement Learning · ICML 2022
Machine learning › Reinforcement learning
sample efficiency
0.312025
State Revisit and Re-explore: Bridging Sim-to-Real Gaps in Offline-and-Online Reinforcement Learning with An Imperfect Simulator · IJCAI 2025
Machine learning › Reinforcement learning
policy learning
0.212024
Grounded Answers for Multi-agent Decision-making Problem through Generative World Model · NeurIPS 2024
Knowledge, reasoning and agents › Multi-agent systems
multi-agent collaboration
0.212023
Deep Hierarchical Communication Graph in Multi-Agent Reinforcement Learning · IJCAI 2023
Computer vision › 3D vision › 3d scene understanding › multi-view understanding
multi-view fusion
0.212023
MMRDN: Consistent Representation for Multi-View Manipulation Relationship Detection in Object-Stacked Scenes · ICRA 2023
Computer vision › 3D vision
point cloud processing
0.212023
MMRDN: Consistent Representation for Multi-View Manipulation Relationship Detection in Object-Stacked Scenes · ICRA 2023
Machine learning › Reinforcement learning › value function approximation
value function representation
0.212022
Greedy based Value Representation for Optimal Coordination in Multi-agent Reinforcement Learning · ICML 2022

Methods — techniques the papers use, named apart from their topics

sub-policy · 0.9online reinforcement learning · 0.9offline reinforcement learning · 0.9meta-policy · 0.9transformer sequence modeling · 0.8intrinsic reward · 0.8image tokenizer · 0.8curriculum learning · 0.8causal transformer · 0.8feature concatenation · 0.7
YearPublicationVenuePosition
2026 MInCo: Mitigating conflicting objectives in distracted visual model-based reinforcement learning
Shiguang Sun, Hanbo Zhang, Zeyang Liu 0001, Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan
Knowl. Based Syst.5
2026 Enhancing Value Decomposition With Target Transformation in Cooperative Multi-Agent Reinforcement Learning
abstract
The increasing need for cooperation among intelligent machines has heightened the importance of cooperative multi-agent reinforcement learning (MARL). However, a dominant class of cooperative MARL approaches relies on monotonic value decomposition, which enables scalable decentralized execution but restricts the representable class of joint action-values. However, existing remedies bias learning targets toward high-value samples, which can be fragile under stochastic returns because optimistic emphasis may amplify lucky but suboptimal trajectories. To solve this challenge, we propose Target Transformation, which maps non-monotonic and stochastic learning targets into a monotonic-representable surrogate while preserving the optimal joint action. Building on this idea, we develop Uncertainty-aware Target Transformation (UT2) with value-based and policy-based instantiations that combine an uncertainty estimator with a best-individual coordination envelope. Experiments on diverse cooperative MARL benchmarks show that UT2 improves both performance and stability over strong baselines, with larger gains as non-monotonicity and stochasticity increase.
Zeyang Liu 0001, Lipeng Wan 0003, Shiguang Sun, Xue Sui, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Offline Multi-Agent Preference-based Reinforcement Learning with Agent-aware Direct Preference Optimization
Qian Kou, Zeyang Liu 0001, Zhuoran Chen, Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan
AAMAS6
2025 State Revisit and Re-explore: Bridging Sim-to-Real Gaps in Offline-and-Online Reinforcement Learning with An Imperfect Simulator
abstract
In reinforcement learning (RL) based robot skill acquisition, a high-fidelity simulator is usually indispensable but unattainable since the real environment dynamics are difficult to model, which leads to severe sim-to-real gaps. Existing methods solve this problem by combining offline and online RL to jointly learn transferable policies from limited offline data and imperfect simulators. However, due to the unrestricted exploration in the imperfect simulator, the hybrid offline-and-online RL methods inevitably suffer from low sample efficiency and insufficient state-action space coverage during training. To solve this problem, we propose a State Revisit and Re-exploration (SR2) hybrid offline-and-online RL framework. In particular, the proposed algorithm employs a meta-policy and a sub-policy, where the meta-policy aims to find high-quality states in the offline trajectories for online exploration, and the sub-policy learns the robot skill using mixed offline and online data. By introducing the state revisit and explore mechanism, our approach efficiently improves performance on a set of sim-to-real robotic tasks. Through extensive simulation and real-world tasks, we demonstrate the superior performance of our approach against other state-of-the-art methods.
Xingyu Chen 0001, Jiayi Xie, Ruixun Liu, Zeyang Liu 0001, Lipeng Wan 0003, Xuguang Lan
IJCAI7
2025 Improving Sample Efficiency Through Stability Enhancement in Deep-Reinforcement Learning
abstract
Prioritizing or reweighting important samples has been recognized as an effective means of improving the efficiency of deep-reinforcement learning (DRL) algorithms. However, many existing techniques encounter stability challenges, limiting efficiency and increasing computational costs and training time. In this study, we aim to improve training efficiency by exploring the intrinsic relationship between sample efficiency and stability. To achieve this, we propose the Stability Contribution Index (SI), which assigns sample priorities based on their impact on stability and employs them to weight the value loss, thereby promoting stable learning and improving efficiency. The effectiveness of our method is validated through comprehensive experiments on two distinct benchmarks: 1) the continuous control domain DMControl and 2) the discrete control environment ProcGen. Compatible with both off-policy and on-policy DRL algorithms, our approach significantly improves sample efficiency and overall performance by fostering greater stability during training. Additionally, experimental results show that our method outperforms well-established sample-efficient reinforcement learning techniques across multiple settings.
Ziru Wang, Wanli Jiang, Ru Peng, Qian Kou, Lipeng Wan 0003, Xuguang Lan
IEEE Trans. Syst. Man Cybern. Syst.5
2024 Imagine, Initialize, and Explore: An Effective Exploration Method in Multi-Agent Reinforcement Learning
abstract
Effective exploration is crucial to discovering optimal strategies for multi-agent reinforcement learning (MARL) in complex coordination tasks. Existing methods mainly utilize intrinsic rewards to enable committed exploration or use role-based learning for decomposing joint action spaces instead of directly conducting a collective search in the entire action-observation space. However, they often face challenges obtaining specific joint action sequences to reach successful states in long-horizon tasks. To address this limitation, we propose Imagine, Initialize, and Explore (IIE), a novel method that offers a promising solution for efficient multi-agent exploration in complex scenarios. IIE employs a transformer model to imagine how the agents reach a critical state that can influence each other's transition functions. Then, we initialize the environment at this state using a simulator before the exploration phase. We formulate the imagination as a sequence modeling problem, where the states, observations, prompts, actions, and rewards are predicted autoregressively. The prompt consists of timestep-to-go, return-to-go, influence value, and one-shot demonstration, specifying the desired state and trajectory as well as guiding the action generation. By initializing agents at the critical states, IIE significantly increases the likelihood of discovering potentially important under-explored regions. Despite its simplicity, empirical results demonstrate that our method outperforms multi-agent exploration baselines on the StarCraft Multi-Agent Challenge (SMAC) and SMACv2 environments. Particularly, IIE shows improved performance in the sparse-reward SMAC tasks and produces more effective curricula over the initialized states than other generative methods, such as CVAE-GAN and diffusion models.
Zeyang Liu 0001, Lipeng Wan 0003, Zhuoran Chen, Xingyu Chen 0001, Xuguang Lan
AAAI2
2024 Grounded Answers for Multi-agent Decision-making Problem through Generative World Model
abstract
Recent progress in generative models has stimulated significant innovations in many fields, such as image generation and chatbots. Despite their success, these models often produce sketchy and misleading solutions for complex multi-agent decision-making problems because they miss the trial-and-error experience and reasoning as humans. To address this limitation, we explore a paradigm that integrates a language-guided simulator into the multi-agent reinforcement learning pipeline to enhance the generated answer. The simulator is a world model that separately learns dynamics and reward, where the dynamics model comprises an image tokenizer as well as a causal transformer to generate interaction transitions autoregressively, and the reward model is a bidirectional transformer learned by maximizing the likelihood of trajectories in the expert demonstrations under language guidance. Given an image of the current state and the task description, we use the world model to train the joint policy and produce the image sequence as the answer by running the converged policy on the dynamics model. The empirical results demonstrate that this framework can improve the answers for multi-agent decision-making problems by showing superior performance on the training and unseen tasks of the StarCraft Multi-Agent Challenge benchmark. In particular, it can generate consistent interaction sequences and explainable reward functions at interaction states, opening the path for training generative models of the future.
Zeyang Liu 0001, Shiguang Sun, Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan
NeurIPS5
2024 Knowledge Graph Enhancement for Fine-Grained Zero-Shot Learning on ImageNet21K
abstract
Fine-grained Zero-shot Learning on the large-scale dataset ImageNet21K is an important task that has promising perspectives in many real-world scenarios. One typical solution is to explicitly model the knowledge passing using a Knowledge Graph (KG) to transfer knowledge from seen to unseen instances. By analyzing the hierarchical structure and the word descriptions on ImageNet21K, we find that the noisy semantic information, the sparseness of seen classes, and the lack of supervision of unseen classes make the knowledge passing insufficient, which limits the KG-based fine-grained ZSL. To resolve this problem, in this paper, we enhance the knowledge passing from three aspects. First, we use more powerful models such as the Large Language Model and Vision-Language Model to get more reliable semantic embeddings. Then we propose a strategy that globally enhances the knowledge graph based on the convex combination relationship of the semantic embeddings. It effectively connects the edges between the non-kinship seen and unseen classes that have strong correlations while assigning an importance score to each edge. Based on the enhanced knowledge graph, we further present a novel regularizer that locally enhances the knowledge passing during training. We extensively conducted comparative evaluations to demonstrate the advantages of our method over state-of-the-art approaches.
Xingyu Chen 0001, Zeyang Liu 0001, Lipeng Wan 0003, Xuguang Lan, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 MMRDN: Consistent Representation for Multi-View Manipulation Relationship Detection in Object-Stacked Scenes
abstract
Manipulation relationship detection (MRD) aims to guide the robot to grasp objects in the right order, which is important to ensure the safety and reliability of grasping in object stacked scenes. Previous works infer manipulation relationship by deep neural network trained with data collected from a predefined view, which has limitation in visual dislocation in unstructured environments. Multi-view data provide more comprehensive information in space, while a challenge of multi-view MRD is domain shift. In this paper, we propose a novel multi-view fusion framework, namely multi-view MRD network (MMRDN), which is trained by 2D and 3D multi-view data. We project the 2D data from different views into a common hidden space and fit the embeddings with a set of Von-Mises-Fisher distributions to learn the consistent representations. Besides, taking advantage of position information within the 3D data, we select a set of$K$Maximum Vertical Neighbors (KMVN) points from the point cloud of each object pair, which encodes the relative position of these two objects. Finally, the features of multi-view 2D and 3D data are concatenated to predict the pairwise relationship of objects. Experimental results on the challenging REGRAD dataset show that MMRDN outperforms the state-of-the-art methods in multi-view MRD tasks. The results also demonstrate that our model trained by synthetic data is capable to transfer to real-world scenarios.
Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
ICRA3
2023 Deep Hierarchical Communication Graph in Multi-Agent Reinforcement Learning
abstract
Sharing intentions is crucial for efficient cooperation in communication-enabled multi-agent reinforcement learning. Recent work applies static or undirected graphs to determine the order of interaction. However, the static graph is not general for complex cooperative tasks, and the parallel message-passing update in the undirected graph with cycles cannot guarantee convergence. To solve this problem, we propose Deep Hierarchical Communication Graph (DHCG) to learn the dependency relationships between agents based on their messages. The relationships are formulated as directed acyclic graphs (DAGs), where the selection of the proper topology is viewed as an action and trained in an end-to-end fashion. To eliminate the cycles in the graph, we apply an acyclicity constraint as intrinsic rewards and then project the graph in the admissible solution set of DAGs. As a result, DHCG removes redundant communication edges for cost improvement and guarantees convergence. To show the effectiveness of the learned graphs, we propose policy-based and value-based DHCG. Policy-based DHCG factorizes the joint policy in an auto-regressive manner, and value-based DHCG factorizes the joint value function to individual value functions and pairwise payoff functions. Empirical results show that our method improves performance across various cooperative multi-agent tasks, including Predator-Prey, Multi-Agent Coordination Challenge, and StarCraft Multi-Agent Challenge.
Zeyang Liu 0001, Lipeng Wan 0003, Xue Sui, Zhuoran Chen, Kewu Sun, Xuguang Lan
IJCAI2
2022 Greedy based Value Representation for Optimal Coordination in Multi-agent Reinforcement Learning
abstract
Due to the representation limitation of the joint Q value function, multi-agent reinforcement learning methods with linear value decomposition (LVD) or monotonic value decomposition (MVD) suffer from relative overgeneralization. As a result, they can not ensure optimal consistency (i.e., the correspondence between individual greedy actions and the best team performance). In this paper, we derive the expression of the joint Q value function of LVD and MVD. According to the expression, we draw a transition diagram, where each self-transition node (STN) is a possible convergence. To ensure the optimal consistency, the optimal node is required to be the unique STN. Therefore, we propose the greedy-based value representation (GVR), which turns the optimal node into an STN via inferior target shaping and eliminates the non-optimal STNs via superior experience replay. Theoretical proofs and empirical results demonstrate that given the true Q values, GVR ensures the optimal consistency under sufficient exploration. Besides, in tasks where the true Q values are unavailable, GVR achieves an adaptive trade-off between optimality and stability. Our method outperforms state-of-the-art baselines in experiments on various benchmarks.
Lipeng Wan 0003, Zeyang Liu 0001, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
ICML1
2019 A Multi-task Convolutional Neural Network for Autonomous Robotic Grasping in Object Stacking Scenes
abstract
Autonomous robotic grasping plays an important role in intelligent robotics. However, how to help the robot grasp specific objects in object stacking scenes is still an open problem, because there are two main challenges for autonomous robots: (1) it is a comprehensive task to know what and how to grasp; (2) it is hard to deal with the situations in which the target is hidden or covered by other objects. In this paper, we propose a multi-task convolutional neural network for autonomous robotic grasping, which can help the robot find the target, make the plan for grasping and finally grasp the target step by step in object stacking scenes. We integrate vision-based robotic grasping detection and visual manipulation relationship reasoning in one single deep network and build the autonomous robotic grasping system. Experimental results demonstrate that with our model, Baxter robot can autonomously grasp the target with a success rate of 90.6%, 71.9% and 59.4% in object cluttered scenes, familiar stacking scenes and complex stacking scenes respectively.
Hanbo Zhang, Xuguang Lan, Cedar Site Bai, Lipeng Wan 0003, Chenjie Yang, Nanning Zheng 0001
IROS4