Siyuan Li 0003

dblp:63/9705-3 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0001-7965-598XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 7 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios
abstract
Model-based reinforcement learning (MBRL) is a crucial approach to enhance the generalization capabilities and improve the sample efficiency of RL algorithms. However, current MBRL methods focus primarily on building world models for single tasks and rarely address generalization across different scenarios. Building on the insight that dynamics within the same simulation engine share inherent properties, we attempt to construct a unified world model capable of generalizing across different scenarios, named Meta-Regularized Contextual World-Model (MrCoM). This method first decomposes the latent state space into various components based on the dynamic characteristics, thereby enhancing the accuracy of world-model prediction. Further, MrCoM adopts meta-state regularization to extract unified representation of scenario-relevant information, and meta-value regularization to align world-model optimization with policy learning across diverse scenario objectives. We theoretically analyze the generalization error upper bound of MrCoM in multi-scenario settings. We systematically evaluate our algorithm's generalization ability across diverse scenarios, demonstrating significantly better performance than previous state-of-the-art methods.
Xuantang Xiong, Ni Mu, Runpeng Xie, Senhao Yang, Lexiang Wang, Yao Luan 0001, Siyuan Li 0003, Yiqin Yang, Bo Xu 0002
AAAI8
2026 Curriculum reinforcement learning with measurable task representation learning
Yongyan Wen, Siyuan Li 0003, Mingjian Fu 0001, Yiqin Yang, Peng Liu 0008
Neural Networks2
2026 An Imitative Reinforcement Learning Framework for Pursuit-Lock-Launch Missions
abstract
Unmanned combat aerial vehicle (UCAV) within-visual-range (WVR) engagement, referring to a fight between two or more UCAVs at close quarters, plays a decisive role on the aerial battlefields. With the development of artificial intelligence, WVR engagement progressively advances toward intelligent and autonomous modes. However, autonomous WVR engagement policy learning is hindered by challenges such as weak exploration capabilities, low learning efficiency, and unrealistic simulated environments. To overcome these challenges, we propose a novel imitative reinforcement learning framework, which efficiently leverages expert data while enabling autonomous exploration. The proposed framework not only enhances learning efficiency through expert imitation but also ensures adaptability to dynamic environments via autonomous exploration with reinforcement learning. Therefore, the proposed framework can learn a successful policy of “pursuit-lock-launch” for UCAVs. To support data-driven learning, we establish an environment based on the Harfang3D sandbox. The extensive experimental results indicate that the proposed framework excels in this multistage task and significantly outperforms state-of-the-art reinforcement learning and imitation learning methods. Thanks to the ability of imitating experts and autonomous exploration, our framework can quickly learn the critical knowledge in complex aerial combat tasks, achieving up to a 100% success rate and demonstrating excellent robustness.
Siyuan Li 0003, Rongchang Zuo, Bofei Liu, Yaoyu He, Peng Liu 0008, Yingnan Zhao 0002
ACM Trans. Auton. Adapt. Syst.1
2025 Safe Planner: Empowering Safety Awareness in Large Pre-Trained Models for Robot Task Planning
abstract
Robot task planning is an important problem for autonomous robots in long-horizon challenging tasks. As large pre-trained models have demonstrated superior planning ability, recent research investigates utilizing large models to achieve autonomous planning for robots in diverse tasks. However, since the large models are pre-trained with Internet data and lack the knowledge of real task scenes, large models as planners may make unsafe decisions that hurt the robots and the surrounding environments. To solve this challenge, we propose a novel Safe Planner framework, which empowers safety awareness in large pre-trained models to accomplish safe and executable planning. In this framework, we develop a safety prediction module to guide the high-level large model planner, and this safety module trained in a simulator can be effectively transferred to real-world tasks. The proposed Safe Planner framework is evaluated on both simulated environments and real robots. The experiment results demonstrate that Safe Planner not only achieves state-of-the-art task success rates, but also substantially improves safety during task execution.
Siyuan Li 0003, Lingfei Cui, Jiani Lu, Qinqin Xiao, Xirui Yang, Peng Liu 0008, Kewu Sun
AAAI1
2025 SkillTree: Explainable Skill-Based Deep Reinforcement Learning for Long-Horizon Control Tasks
abstract
Deep reinforcement learning (DRL) has achieved remarkable success in various domains, yet its reliance on neural networks results in a lack of transparency, which limits its practical applications in safety-critical and human-agent interaction domains. Decision trees, known for their notable explainability, have emerged as a promising alternative to neural networks. However, decision trees often struggle in long-horizon continuous control tasks with high-dimensional observation space due to their limited expressiveness. To address this challenge, we propose SkillTree, a novel hierarchical framework that reduces the complex continuous action space of challenging control tasks into discrete skill space. By integrating the differentiable decision tree within the high-level policy, SkillTree generates discrete skill embeddings that guide low-level policy execution. Furthermore, through distillation, we obtain a simplified decision tree model that improves performance while further reducing complexity. Experiment results validate SkillTree’s effectiveness across various robotic manipulation tasks, providing clear skill-level insights into the decision-making process. The proposed approach not only achieves performance comparable to neural network based methods in complex long-horizon control tasks but also significantly enhances the transparency and explainability of the decision-making process.
Yongyan Wen, Siyuan Li 0003, Rongchang Zuo, Hangyu Mao, Peng Liu 0008
AAAI2
2025 Auxiliary Reward Generation With Transition Distance Representation Learning
abstract
Reinforcement learning (RL) has shown strengths in challenging sequential decision-making problems. The reward function in RL is crucial to the learning performance, as it quantifies the degree of task completion. In real-world problems, the rewards are predominantly human-designed, which requires laborious tuning, and is susceptible to human cognitive biases. To achieve automatic auxiliary reward generation, we propose a novel representation learning approach that can measure the “transition distance” between states. Building upon these representations, we introduce an auxiliary reward generation technique for both single-task and skill-chaining scenarios without the need for human knowledge. Furthermore, we theoretically show that the proposed auxiliary rewards maintain the policy invariance property, i.e., the generated rewards will not hurt the policy optimality under the original rewards. In the experiment section, we evaluate the proposed approach in both online and offline learning settings in a wide range of tasks, including robot manipulation and locomotion. The experiment results demonstrate the effectiveness of measuring the transition distance and the induced improvement by auxiliary rewards, which promotes better learning efficiency and increases convergent stability. Beyond that, we demonstrate that the learned manipulation policy with the auxiliary rewards in a simulator can be transferred to the real robot, as shown inhttps://sites.google.com/view/transition-distance-rp/tdrp. Note to Practitioners—The motivation for this paper arises from the need for a technique that enhances robot skill-learning efficiency and performance in both single-task and skill-chaining scenarios. Our research primarily focuses on robot arm manipulation tasks. To accelerate the policy learning process and improve policy performance for executing these tasks, we introduce an auxiliary reward generation technique for both single-task and skill-chaining scenarios without requiring human expertise. This technique leverages the proposed novel representation learning approach, which can measure the “transition distance” between states. During each policy training round, the robot receives a dense reshaped reward created by our approach. Using the policy trained by our method, we successfully control a real Franka Panda robot arm to complete various manipulation tasks.
Siyuan Li 0003, Shijie Han, Yingnan Zhao 0002, Yiqin Yang, Qianchuan Zhao, Peng Liu 0008
IEEE Trans Autom. Sci. Eng.1
2024 Robust Visual Imitation Learning with Inverse Dynamics Representations
abstract
Imitation learning (IL) has achieved considerable success in solving complex sequential decision-making problems. However, current IL methods mainly assume that the environment for learning policies is the same as the environment for collecting expert datasets. Therefore, these methods may fail to work when there are slight differences between the learning and expert environments, especially for challenging problems with high-dimensional image observations. However, in real-world scenarios, it is rare to have the chance to collect expert trajectories precisely in the target learning environment. To address this challenge, we propose a novel robust imitation learning approach, where we develop an inverse dynamics state representation learning objective to align the expert environment and the learning environment. With the abstract state representation, we design an effective reward function, which thoroughly measures the similarity between behavior data and expert data not only element-wise, but also from the trajectory level. We conduct extensive experiments to evaluate the proposed approach under various visual perturbations and in diverse visual control tasks. Our approach can achieve a near-expert performance in most environments, and significantly outperforms the state-of-the-art visual IL methods and robust IL methods.
Siyuan Li 0003, Rongchang Zuo, Kewu Sun, Lingfei Cui, Jishiyu Ding, Peng Liu 0008
AAAI1
2024 MacMic: Executing Iceberg Orders via Hierarchical Reinforcement Learning
Hui Niu, Siyuan Li 0003, Jian Li 0015
IJCAI2
2024 IMM: An Imitative Reinforcement Learning Approach with Predictive Representation Learning for Automatic Market Making
Hui Niu, Siyuan Li 0003, Jiahao Zheng 0002, Zhouchi Lin, Bo An 0001, Jian Li 0015
IJCAI2
2024 IOB: integrating optimization transfer and behavior transfer for multi-policy reuse
Siyuan Li 0003, Hao Li 0069, Jin Zhang 0016, Zhen Wang 0004, Peng Liu 0008, Chongjie Zhang
Auton. Agents Multi Agent Syst.1
2023 Flow to Control: Offline Reinforcement Learning with Lossless Primitive Discovery
abstract
Offline reinforcement learning (RL) enables the agent to effectively learn from logged data, which significantly extends the applicability of RL algorithms in real-world scenarios where exploration can be expensive or unsafe. Previous works have shown that extracting primitive skills from the recurring and temporally extended structures in the logged data yields better learning. However, these methods suffer greatly when the primitives have limited representation ability to recover the original policy space, especially in offline settings. In this paper, we give a quantitative characterization of the performance of offline hierarchical learning and highlight the importance of learning lossless primitives. To this end, we propose to use a flow-based structure as the representation for low-level policies. This allows us to represent the behaviors in the dataset faithfully while keeping the expression ability to recover the whole policy space. We show that such lossless primitives can drastically improve the performance of hierarchical policies. The experimental results and extensive ablation studies on the standard D4RL benchmark show that our method has a good representation ability for policies and achieves superior performance in most tasks.
Yiqin Yang, Hao Hu 0006, Siyuan Li 0003, Jun Yang 0028, Qianchuan Zhao, Chongjie Zhang
AAAI4
2023 Behavior Contrastive Learning for Unsupervised Skill Discovery
abstract
In reinforcement learning, unsupervised skill discovery aims to learn diverse skills without extrinsic rewards. Previous methods discover skills by maximizing the mutual information (MI) between states and skills. However, such an MI objective tends to learn simple and static skills and may hinder exploration. In this paper, we propose a novel unsupervised skill discovery method through contrastive learning among behaviors, which makes the agent produce similar behaviors for the same skill and diverse behaviors for different skills. Under mild assumptions, our objective maximizes the MI between different behaviors based on the same skill, which serves as an upper bound of the previous MI objective. Meanwhile, our method implicitly increases the state entropy to obtain better state coverage. We evaluate our method on challenging mazes and continuous control tasks. The results show that our method generates diverse and far-reaching skills, and also obtains competitive performance in downstream tasks compared to the state-of-the-art methods.
Rushuai Yang, Chenjia Bai, Hongyi Guo, Siyuan Li 0003, Bin Zhao 0001, Zhen Wang 0004, Peng Liu 0008, Xuelong Li 0001
ICML4
2023 Learning to Solve Tasks with Exploring Prior Behaviours
abstract
Demonstrations are widely used in Deep Reinforcement Learning (DRL) for facilitating solving tasks with sparse rewards. However, the tasks in real-world scenarios can often have varied initial conditions from the demonstration, which would require additional prior behaviours. For example, consider we are given the demonstration for the task of picking up an object from an open drawer, but the drawer is closed in the training. Without acquiring the prior behaviours of opening the drawer, the robot is unlikely to solve the task. To address this, in this paper we propose an Intrinsic Reward Driven Example-based Control (IRDEC). Our method can endow agents with the ability to explore and acquire the required prior behaviours and then connect to the task-specific behaviours in the demonstration to solve sparse-reward tasks without requiring additional demonstration of the prior behaviours. The performance of our method outperforms other baselines on three navigation tasks and one robotic manipulation task with sparse rewards. Codes are available at https://github.com/Ricky-Zhu/IRDEC.
Siyuan Li 0003, Tianhong Dai, Chongjie Zhang, Oya Çeliktutan
IROS2
2023 Classifying ambiguous identities in hidden-role Stochastic games with multi-agent reinforcement learning
Shijie Han, Siyuan Li 0003, Bo An 0001, Wei Zhao 0008, Peng Liu 0008
Auton. Agents Multi Agent Syst.2
2022 MetaTrader: An Reinforcement Learning Approach Integrating Diverse Policies for Portfolio Optimization
abstract
Portfolio management is a fundamental problem in finance. It involves periodic reallocations of assets to maximize the expected returns within an appropriate level of risk exposure. Deep reinforcement learning (RL) has been considered a promising approach to solving this problem owing to its strong capability in sequential decision making. However, due to the non-stationary nature of financial markets, applying RL techniques to portfolio optimization remains a challenging problem. Extracting trading knowledge from various expert strategies could be helpful for agents to accommodate the changing markets. In this paper, we propose MetaTrader, a novel two-stage RL-based approach for portfolio management, which learns to integrate diverse trading policies to adapt to various market conditions. In the first stage, MetaTrader incorporates an imitation learning objective into the reinforcement learning framework. Through imitating different expert demonstrations, MetaTrader acquires a set of trading policies with great diversity. In the second stage, MetaTrader learns a meta-policy to recognize the market conditions and decide on the most proper learned policy to follow. We evaluate the proposed approach on three real-world index datasets and compare it to state-of-the-art baselines. The empirical results demonstrate that MetaTrader significantly outperforms those baselines in balancing profits and risks. Furthermore, thorough ablation studies validate the effectiveness of the components in the proposed approach.
Hui Niu, Siyuan Li 0003, Jian Li 0015
CIKM2
2022 Active Hierarchical Exploration with Stable Subgoal Representation Learning
Siyuan Li 0003, Jin Zhang 0016, Yang Yu 0001, Chongjie Zhang
ICLR1
2022 CUP: Critic-Guided Policy Reuse
abstract
The ability to reuse previous policies is an important aspect of human intelligence. To achieve efficient policy reuse, a Deep Reinforcement Learning (DRL) agent needs to decide when to reuse and which source policies to reuse. Previous methods solve this problem by introducing extra components to the underlying algorithm, such as hierarchical high-level policies over source policies, or estimations of source policies' value functions on the target task. However, training these components induces either optimization non-stationarity or heavy sampling cost, significantly impairing the effectiveness of transfer. To tackle this problem, we propose a novel policy reuse algorithm called Critic-gUided Policy reuse (CUP), which avoids training any extra components and efficiently reuses source policies. CUP utilizes the critic, a common component in actor-critic methods, to evaluate and choose source policies. At each state, CUP chooses the source policy that has the largest one-step improvement over the current target policy, and forms a guidance policy. The guidance policy is theoretically guaranteed to be a monotonic improvement over the current target policy. Then the target policy is regularized to imitate the guidance policy to perform efficient policy search. Empirical results demonstrate that CUP achieves efficient transfer and significantly outperforms baseline algorithms.
Jin Zhang 0016, Siyuan Li 0003, Chongjie Zhang
NeurIPS2
2021 Learning Subgoal Representations with Slow Dynamics
Siyuan Li 0003, Lulu Zheng, Chongjie Zhang
ICLR1
2021 Offline Reinforcement Learning with Reverse Model-based Imagination
abstract
In offline reinforcement learning (offline RL), one of the main challenges is to deal with the distributional shift between the learning policy and the given dataset. To address this problem, recent offline RL methods attempt to introduce conservatism bias to encourage learning in high-confidence areas. Model-free approaches directly encode such bias into policy or value function learning using conservative regularizations or special network structures, but their constrained policy search limits the generalization beyond the offline dataset. Model-based approaches learn forward dynamics models with conservatism quantifications and then generate imaginary trajectories to extend the offline datasets. However, due to limited samples in offline datasets, conservatism quantifications often suffer from overgeneralization in out-of-support regions. The unreliable conservative measures will mislead forward model-based imaginations to undesired areas, leading to overaggressive behaviors. To encourage more conservatism, we propose a novel model-based offline RL framework, called Reverse Offline Model-based Imagination (ROMI). We learn a reverse dynamics model in conjunction with a novel reverse policy, which can generate rollouts leading to the target goal states within the offline dataset. These reverse imaginations provide informed data augmentation for model-free policy learning and enable conservative generalization beyond the offline dataset. ROMI can effectively combine with off-the-shelf model-free algorithms to enable model-based generalization with proper conservatism. Empirical results show that our method can generate more conservative behaviors and achieve state-of-the-art performance on offline RL benchmark tasks.
Haozhe Jiang, Guangxiang Zhu, Siyuan Li 0003, Chongjie Zhang
NeurIPS5
2019 Hierarchical Reinforcement Learning with Advantage-Based Auxiliary Rewards
abstract
Hierarchical Reinforcement Learning (HRL) is a promising approach to solving long-horizon problems with sparse and delayed rewards. Many existing HRL algorithms either use pre-trained low-level skills that are unadaptable, or require domain-specific information to define low-level rewards. In this paper, we aim to adapt low-level skills to downstream tasks while maintaining the generality of reward design. We propose an HRL framework which sets auxiliary rewards for low-level skill training based on the advantage function of the high-level policy. This auxiliary reward enables efficient, simultaneous learning of the high-level policy and low-level skills without using task-specific knowledge. In addition, we also theoretically prove that optimizing low-level skills with this auxiliary reward will increase the task return for the joint policy. Experimental results show that our algorithm dramatically outperforms other state-of-the-art HRL methods in Mujoco domains. We also find both low-level and high-level policies trained by our algorithm transferable.
Siyuan Li 0003, Minxue Tang, Chongjie Zhang
NeurIPS1
2018 An Optimal Online Method of Selecting Source Policies for Reinforcement Learning
abstract
Transfer learning significantly accelerates the reinforcement learning process by exploiting relevant knowledge from previous experiences. The problem of optimally selecting source policies during the learning process is of great importance yet challenging. There has been little theoretical analysis of this problem. In this paper, we develop an optimal online method to select source policies for reinforcement learning. This method formulates online source policy selection as a multi-armed bandit problem and augments Q-learning with policy reuse. We provide theoretical guarantees of the optimal selection process and convergence to the optimal policy. In addition, we conduct experiments on a grid-based robot navigation domain to demonstrate its efficiency and robustness by comparing to the state-of-the-art transfer learning method.
Siyuan Li 0003, Chongjie Zhang
AAAI1