Zeyang Liu 0001

dblp:57/10146-1 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-3110-8618ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 High-performance multi-agent path finding in high-obstacle-density and large-size maps
Shiguang Sun, Chang Tang, Shi-tao Chen, Zeyang Liu 0001, Xingyu Chen 0001, Xuguang Lan
Neurocomputing4
2026 MInCo: Mitigating conflicting objectives in distracted visual model-based reinforcement learning
Shiguang Sun, Hanbo Zhang, Zeyang Liu 0001, Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan
Knowl. Based Syst.3
2026 Enhancing Value Decomposition With Target Transformation in Cooperative Multi-Agent Reinforcement Learning
abstract
The increasing need for cooperation among intelligent machines has heightened the importance of cooperative multi-agent reinforcement learning (MARL). However, a dominant class of cooperative MARL approaches relies on monotonic value decomposition, which enables scalable decentralized execution but restricts the representable class of joint action-values. However, existing remedies bias learning targets toward high-value samples, which can be fragile under stochastic returns because optimistic emphasis may amplify lucky but suboptimal trajectories. To solve this challenge, we propose Target Transformation, which maps non-monotonic and stochastic learning targets into a monotonic-representable surrogate while preserving the optimal joint action. Building on this idea, we develop Uncertainty-aware Target Transformation (UT2) with value-based and policy-based instantiations that combine an uncertainty estimator with a best-individual coordination envelope. Experiments on diverse cooperative MARL benchmarks show that UT2 improves both performance and stability over strong baselines, with larger gains as non-monotonicity and stochasticity increase.
Zeyang Liu 0001, Lipeng Wan 0003, Shiguang Sun, Xue Sui, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 DualSkill: Unifying Discrete Stability and Continuous Flexibility for Embodied Control
Ziru Wang, Haowen Sun 0003, Zeyang Liu 0001, Xuguang Lan
IEEE Trans Autom. Sci. Eng.4
2025 Offline Multi-Agent Preference-based Reinforcement Learning with Agent-aware Direct Preference Optimization
Qian Kou, Zeyang Liu 0001, Zhuoran Chen, Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan
AAMAS3
2025 State Revisit and Re-explore: Bridging Sim-to-Real Gaps in Offline-and-Online Reinforcement Learning with An Imperfect Simulator
abstract
In reinforcement learning (RL) based robot skill acquisition, a high-fidelity simulator is usually indispensable but unattainable since the real environment dynamics are difficult to model, which leads to severe sim-to-real gaps. Existing methods solve this problem by combining offline and online RL to jointly learn transferable policies from limited offline data and imperfect simulators. However, due to the unrestricted exploration in the imperfect simulator, the hybrid offline-and-online RL methods inevitably suffer from low sample efficiency and insufficient state-action space coverage during training. To solve this problem, we propose a State Revisit and Re-exploration (SR2) hybrid offline-and-online RL framework. In particular, the proposed algorithm employs a meta-policy and a sub-policy, where the meta-policy aims to find high-quality states in the offline trajectories for online exploration, and the sub-policy learns the robot skill using mixed offline and online data. By introducing the state revisit and explore mechanism, our approach efficiently improves performance on a set of sim-to-real robotic tasks. Through extensive simulation and real-world tasks, we demonstrate the superior performance of our approach against other state-of-the-art methods.
Xingyu Chen 0001, Jiayi Xie, Ruixun Liu, Zeyang Liu 0001, Lipeng Wan 0003, Xuguang Lan
IJCAI6
2025 Towards Extrinsic Dexterity Grasping in Unrestricted Environments
abstract
Grasping large and flat objects (e.g., a book or a pan) is often regarded as an ungraspable task, which poses significant challenges due to the unreachable grasping poses. Prior research has exploited environmental interactions through Extrinsic Dexterity, utilizing external structures such as walls or table edges to facilitate object grasping. However, they are confined to task-specific policies while neglecting semantic perception and planning to identify optimal pre-grasp configurations. This limits their operational versatility, impeding effective adaptation to varied extrinsic dexterity constraints. In this work, we present ExDiff, a robot manipulation approach for extrinsic dexterity grasping in unrestricted environments. It utilizes Vision-Language Models (VLMs) to perceive the environmental state and generate instructions, followed by a Goal-Conditioned Action Diffusion (GCAD) model to predict the sequence of low-level actions. This diffusion model learns the low-level policy, conditioned on high-level instructions and cumulative rewards, which improves the generation of robot actions. Simulation experiments and real-world deployment results demonstrate that ExDiff effectively performs ungraspable tasks and generalizes to previously unseen target objects and scenes. Videos at - https://exdiff.github.io/index.html
Chengzhong Ma, Houxue Yang, Hanbo Zhang, Zeyang Liu 0001, Xuguang Lan, Nanning Zheng 0001
IROS4
2025 Flight Mastery in Turbulent Skies: Shared Control and Curriculum Reinforcement Learning for Crosswind Landing
abstract
Landing in crosswind conditions poses significant challenges for aircraft, as traditional control methods often fail to ensure stability in rapidly changing wind environments. While reinforcement learning offers a promising alternative, it typically suffers from low sample efficiency and limited generalization under stochastic wind fields. To address these challenges, we propose a Shared Control and Curriculum Reinforcement Learning framework. We model the crosswind landing task as a Markov Decision Process (MDP), explicitly defining the state space, action space, and wind field representation. To initialize learning, we decompose the multi-objective landing task into four sub-tasks—altitude, attitude, heading, and speed control—and train expert policies for each. These are then distilled into a shared control model via behavior cloning, providing a pre-trained policy with basic flight control capabilities. We further fine-tune this model using curriculum reinforcement learning, progressively increasing the complexity of wind conditions to enhance robustness and generalization. Experimental results across multiple aircraft and wind scenarios show that our method improves landing success rates and trajectory smoothness, while generalizing more effectively to unseen wind conditions, outperforming PID controllers, imitation learning, and mainstream RL baselines.
Zechen Shi, Xingyu Chen 0001, Zeyang Liu 0001, Chi Zhang 0020, Yimeng Yu, Junbin You, Xuguang Lan
IEEE Trans. Intell. Transp. Syst.3
2025 Improving Offline Reinforcement Learning With in-Sample Advantage Regularization for Robot Manipulation
abstract
Offline reinforcement learning (RL) aims to learn the possible policy from a fixed dataset without real-time interactions with the environment. By avoiding the risky exploration of the robot, this approach is expected to significantly improve the robot's learning efficiency and safety. However, due to errors in value estimation from out-of-distribution actions, most offline RL algorithms constrain or regularize the policy to the actions contained within the dataset. The cost of such methods is the introduction of new hyperparameters and additional complexity. In this article, we aim to adapt offline RL to robotic manipulation with minimal changes and to avoid evaluating out-of-distribution actions as much as possible. Therefore, we improve offline RL with in-sample advantage regularization (ISAR). To mitigate the impact of unseen actions, the ISAR learns the state-value function only with the dataset sample to regress the optimal action-value function. Our method calculates the advantage function of action-state pairs based on in-sample value estimation and adds a behavior cloning (BC) regularization term in the policy update. This improves sample efficiency with minimal changes, resulting in a simple and easy-to-implement method. The experiments of the D4RL robot benchmark and multigoal sparse rewards robotic tasks show that the ISAR achieves excellent performance comparable to current state-of-the-art algorithms without the need for complex parameter tuning and too much training time. In addition, we demonstrate the effectiveness of our method on a real-world robot platform.
Chengzhong Ma, Deyu Yang, Zeyang Liu 0001, Houxue Yang, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Imagine, Initialize, and Explore: An Effective Exploration Method in Multi-Agent Reinforcement Learning
abstract
Effective exploration is crucial to discovering optimal strategies for multi-agent reinforcement learning (MARL) in complex coordination tasks. Existing methods mainly utilize intrinsic rewards to enable committed exploration or use role-based learning for decomposing joint action spaces instead of directly conducting a collective search in the entire action-observation space. However, they often face challenges obtaining specific joint action sequences to reach successful states in long-horizon tasks. To address this limitation, we propose Imagine, Initialize, and Explore (IIE), a novel method that offers a promising solution for efficient multi-agent exploration in complex scenarios. IIE employs a transformer model to imagine how the agents reach a critical state that can influence each other's transition functions. Then, we initialize the environment at this state using a simulator before the exploration phase. We formulate the imagination as a sequence modeling problem, where the states, observations, prompts, actions, and rewards are predicted autoregressively. The prompt consists of timestep-to-go, return-to-go, influence value, and one-shot demonstration, specifying the desired state and trajectory as well as guiding the action generation. By initializing agents at the critical states, IIE significantly increases the likelihood of discovering potentially important under-explored regions. Despite its simplicity, empirical results demonstrate that our method outperforms multi-agent exploration baselines on the StarCraft Multi-Agent Challenge (SMAC) and SMACv2 environments. Particularly, IIE shows improved performance in the sparse-reward SMAC tasks and produces more effective curricula over the initialized states than other generative methods, such as CVAE-GAN and diffusion models.
Zeyang Liu 0001, Lipeng Wan 0003, Zhuoran Chen, Xingyu Chen 0001, Xuguang Lan
AAAI1
2024 Grounded Answers for Multi-agent Decision-making Problem through Generative World Model
abstract
Recent progress in generative models has stimulated significant innovations in many fields, such as image generation and chatbots. Despite their success, these models often produce sketchy and misleading solutions for complex multi-agent decision-making problems because they miss the trial-and-error experience and reasoning as humans. To address this limitation, we explore a paradigm that integrates a language-guided simulator into the multi-agent reinforcement learning pipeline to enhance the generated answer. The simulator is a world model that separately learns dynamics and reward, where the dynamics model comprises an image tokenizer as well as a causal transformer to generate interaction transitions autoregressively, and the reward model is a bidirectional transformer learned by maximizing the likelihood of trajectories in the expert demonstrations under language guidance. Given an image of the current state and the task description, we use the world model to train the joint policy and produce the image sequence as the answer by running the converged policy on the dynamics model. The empirical results demonstrate that this framework can improve the answers for multi-agent decision-making problems by showing superior performance on the training and unseen tasks of the StarCraft Multi-Agent Challenge benchmark. In particular, it can generate consistent interaction sequences and explainable reward functions at interaction states, opening the path for training generative models of the future.
Zeyang Liu 0001, Shiguang Sun, Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan
NeurIPS1
2024 Optimal bipartite graph matching-based goal selection for policy-based hindsight learning
Shiguang Sun, Hanbo Zhang, Zeyang Liu 0001, Xingyu Chen 0001, Xuguang Lan
Neurocomputing3
2024 Knowledge Graph Enhancement for Fine-Grained Zero-Shot Learning on ImageNet21K
abstract
Fine-grained Zero-shot Learning on the large-scale dataset ImageNet21K is an important task that has promising perspectives in many real-world scenarios. One typical solution is to explicitly model the knowledge passing using a Knowledge Graph (KG) to transfer knowledge from seen to unseen instances. By analyzing the hierarchical structure and the word descriptions on ImageNet21K, we find that the noisy semantic information, the sparseness of seen classes, and the lack of supervision of unseen classes make the knowledge passing insufficient, which limits the KG-based fine-grained ZSL. To resolve this problem, in this paper, we enhance the knowledge passing from three aspects. First, we use more powerful models such as the Large Language Model and Vision-Language Model to get more reliable semantic embeddings. Then we propose a strategy that globally enhances the knowledge graph based on the convex combination relationship of the semantic embeddings. It effectively connects the edges between the non-kinship seen and unseen classes that have strong correlations while assigning an importance score to each edge. Based on the enhanced knowledge graph, we further present a novel regularizer that locally enhances the knowledge passing during training. We extensively conducted comparative evaluations to demonstrate the advantages of our method over state-of-the-art approaches.
Xingyu Chen 0001, Zeyang Liu 0001, Lipeng Wan 0003, Xuguang Lan, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Deep Hierarchical Communication Graph in Multi-Agent Reinforcement Learning
abstract
Sharing intentions is crucial for efficient cooperation in communication-enabled multi-agent reinforcement learning. Recent work applies static or undirected graphs to determine the order of interaction. However, the static graph is not general for complex cooperative tasks, and the parallel message-passing update in the undirected graph with cycles cannot guarantee convergence. To solve this problem, we propose Deep Hierarchical Communication Graph (DHCG) to learn the dependency relationships between agents based on their messages. The relationships are formulated as directed acyclic graphs (DAGs), where the selection of the proper topology is viewed as an action and trained in an end-to-end fashion. To eliminate the cycles in the graph, we apply an acyclicity constraint as intrinsic rewards and then project the graph in the admissible solution set of DAGs. As a result, DHCG removes redundant communication edges for cost improvement and guarantees convergence. To show the effectiveness of the learned graphs, we propose policy-based and value-based DHCG. Policy-based DHCG factorizes the joint policy in an auto-regressive manner, and value-based DHCG factorizes the joint value function to individual value functions and pairwise payoff functions. Empirical results show that our method improves performance across various cooperative multi-agent tasks, including Predator-Prey, Multi-Agent Coordination Challenge, and StarCraft Multi-Agent Challenge.
Zeyang Liu 0001, Lipeng Wan 0003, Xue Sui, Zhuoran Chen, Kewu Sun, Xuguang Lan
IJCAI1
2022 Greedy based Value Representation for Optimal Coordination in Multi-agent Reinforcement Learning
abstract
Due to the representation limitation of the joint Q value function, multi-agent reinforcement learning methods with linear value decomposition (LVD) or monotonic value decomposition (MVD) suffer from relative overgeneralization. As a result, they can not ensure optimal consistency (i.e., the correspondence between individual greedy actions and the best team performance). In this paper, we derive the expression of the joint Q value function of LVD and MVD. According to the expression, we draw a transition diagram, where each self-transition node (STN) is a possible convergence. To ensure the optimal consistency, the optimal node is required to be the unique STN. Therefore, we propose the greedy-based value representation (GVR), which turns the optimal node into an STN via inferior target shaping and eliminates the non-optimal STNs via superior experience replay. Theoretical proofs and empirical results demonstrate that given the true Q values, GVR ensures the optimal consistency under sufficient exploration. Besides, in tasks where the true Q values are unavailable, GVR achieves an adaptive trade-off between optimality and stability. Our method outperforms state-of-the-art baselines in experiments on various benchmarks.
Lipeng Wan 0003, Zeyang Liu 0001, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
ICML2