EDBT 2026 Demo / reviewers in the wild / expert
Yiqin Yang
dblp:180/7725
· DBLP profile ↗
22ranked-venue papers
4as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 4 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-ScenariosabstractModel-based reinforcement learning (MBRL) is a crucial approach to enhance the generalization capabilities and improve the sample efficiency of RL algorithms. However, current MBRL methods focus primarily on building world models for single tasks and rarely address generalization across different scenarios. Building on the insight that dynamics within the same simulation engine share inherent properties, we attempt to construct a unified world model capable of generalizing across different scenarios, named Meta-Regularized Contextual World-Model (MrCoM). This method first decomposes the latent state space into various components based on the dynamic characteristics, thereby enhancing the accuracy of world-model prediction. Further, MrCoM adopts meta-state regularization to extract unified representation of scenario-relevant information, and meta-value regularization to align world-model optimization with policy learning across diverse scenario objectives. We theoretically analyze the generalization error upper bound of MrCoM in multi-scenario settings. We systematically evaluate our algorithm's generalization ability across diverse scenarios, demonstrating significantly better performance than previous state-of-the-art methods. Xuantang Xiong, Ni Mu, Runpeng Xie, Senhao Yang, Lexiang Wang, Yao Luan 0001, Siyuan Li 0003, Yiqin Yang, Bo Xu 0002 |
AAAI | 10 |
| 2026 | Curriculum reinforcement learning with measurable task representation learning
Yongyan Wen, Siyuan Li 0003, Mingjian Fu 0001, Yiqin Yang, Peng Liu 0008 |
Neural Networks | 4 |
| 2025 | DPMT: Dual Process Multi-scale Theory of Mind Framework for Real-time Human-AI Collaboration
Xiyun Li, Yining Ding, Yuhua Jiang, Runpeng Xie, Yuanhua Ni, Yiqin Yang, Bo Xu 0002 |
CogSci | 8 |
| 2025 | Episodic Novelty Through Temporal DistanceabstractExploration in sparse reward environments remains a significant challenge in reinforcement learning, particularly in Contextual Markov Decision Processes (CMDPs), where environments differ across episodes. Existing episodic intrinsic motivation methods for CMDPs primarily rely on count-based approaches, which are ineffective in large state spaces, or on similarity-based methods that lack appropriate metrics for state comparison. To address these shortcomings, we propose Episodic Novelty Through Temporal Distance (ETD), a novel approach that introduces temporal distance as a robust metric for state similarity and intrinsic reward computation. By employing contrastive learning, ETD accurately estimates temporal distances and derives intrinsic rewards based on the novelty of states within the current episode. Extensive experiments on various benchmark tasks demonstrate that ETD significantly outperforms state-of-the-art methods, highlighting its effectiveness in enhancing exploration in sparse reward CMDPs. Yuhua Jiang, Qihan Liu, Yiqin Yang, Xiaoteng Ma, Dianyu Zhong, Hao Hu 0006, Jun Yang 0028, Bin Liang 0001, Bo Xu 0002, Chongjie Zhang, Qianchuan Zhao |
ICLR | 3 |
| 2025 | Fewer May Be Better: Enhancing Offline Reinforcement Learning with Reduced DatasetabstractResearch in offline reinforcement learning (RL) marks a paradigm shift in RL. However, a critical yet under-investigated aspect of offline RL is determining the subset of the offline dataset, which is used to improve algorithm performance while accelerating algorithm training. Moreover, the size of reduced datasets can uncover the requisite offline data volume essential for addressing analogous challenges. Based on the above considerations, we propose identifying Reduced Datasets for Offline RL (ReDOR) by formulating it as a gradient approximation optimization problem. We prove that the common actor-critic framework in reinforcement learning can be transformed into a submodular objective. This insight enables us to construct a subset by adopting the orthogonal matching pursuit (OMP). Specifically, we have made several critical modifications to OMP to enable successful adaptation with Offline RL algorithms. The experimental results indicate that the data subsets constructed by the ReDOR can significantly improve algorithm performance with low computational complexity. Yiqin Yang, Quanwei Wang, Chenghao Li 0002, Hao Hu 0006, Chengjie Wu, Yuhua Jiang, Dianyu Zhong, Ziyou Zhang, Qianchuan Zhao, Chongjie Zhang, Bo Xu 0002 |
ICLR | 1 |
| 2025 | CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous QueriesabstractPreference-based reinforcement learning (PbRL) bypasses explicit reward engineering by inferring reward functions from human preference comparisons, enabling better alignment with human intentions. However, humans often struggle to label a clear preference between similar segments, reducing label efficiency and limiting PbRL’s real-world applicability. To address this, we propose an offline PbRL method: Contrastive LeArning for ResolvIng Ambiguous Feedback (CLARIFY), which learns a trajectory embedding space that incorporates preference information, ensuring clearly distinguished segments are spaced apart, thus facilitating the selection of more unambiguous queries. Extensive experiments demonstrate that CLARIFY outperforms baselines in both non-ideal teachers and real human feedback settings. Our approach not only selects more distinguished queries but also learns meaningful trajectory embeddings. Ni Mu, Hao Hu 0006, Yiqin Yang, Bo Xu 0002, Qing-Shan Jia |
ICML | 4 |
| 2025 | S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement LearningabstractPreference-based reinforcement learning (PbRL) stands out by utilizing human preferences as a direct reward signal, eliminating the need for intricate reward engineering. However, despite its potential, traditional PbRL methods are often constrained by the indistinguishability of segments, which impedes the learning process. In this paper, we introduce Skill-Enhanced Preference Optimization Algorithm (S-EPOA), which addresses the segment indistinguishability issue by integrating skill mechanisms into the preference learning framework. Specifically, we first conduct the unsupervised pretraining to learn useful skills. Then, we propose a novel query selection mechanism to balance the information gain and distinguishability over the learned skill space. Experimental results on a range of tasks, including robotic manipulation and locomotion, demonstrate that S-EPOA significantly outperforms conventional PbRL methods in terms of both robustness and learning efficiency. The results highlight the effectiveness of skill-driven learning in overcoming the challenges posed by segment indistinguishability. Ni Mu, Yao Luan 0001, Yiqin Yang, Bo Xu 0002, Qing-Shan Jia |
IJCAI | 3 |
| 2025 | STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement LearningabstractPreference-based reinforcement learning (PbRL) bypasses complex reward engineering by learning rewards directly from human preferences, enabling better alignment with human intentions. However, its effectiveness in multi-stage tasks, where agents sequentially perform sub-tasks (e.g., navigation, grasping), is limited by **stage misalignment**: Comparing segments from mismatched stages, such as movement versus manipulation, results in uninformative feedback, thus hindering policy learning. In this paper, we validate the stage misalignment issue through theoretical analysis and empirical experiments. To address this issue, we propose **ST**age-**A**l**I**gned **R**eward learning (STAIR), which first learns a stage approximation based on temporal distance, then prioritizes comparisons within the same stage. Temporal distance is learned via contrastive learning, which groups temporally close states into coherent stages, without predefined task knowledge, and adapts dynamically to policy changes. Extensive experiments demonstrate STAIR's superiority in multi-stage tasks and competitive performance in single-stage tasks. Furthermore, human studies show that stages approximated by STAIR are consistent with human cognition, confirming its effectiveness in mitigating stage misalignment. Yao Luan 0001, Ni Mu, Yiqin Yang, Bo Xu 0002, Qing-Shan Jia |
NeurIPS | 3 |
| 2025 | DAIL: Beyond Task Ambiguity for Language-Conditioned Reinforcement LearningabstractComprehending natural language and following human instructions are critical capabilities for intelligent agents.
However, the flexibility of linguistic instructions induces substantial ambiguity across language-conditioned tasks, severely degrading algorithmic performance.
To address these limitations, we present a novel method named DAIL (Distributional Aligned Learning), featuring two key components: distributional policy and semantic alignment.
Specifically, we provide theoretical results that the value distribution estimation mechanism enhances task differentiability.
Meanwhile, the semantic alignment module captures the correspondence between trajectories and linguistic instructions.
Extensive experimental results on both structured and visual observation benchmarks demonstrate that DAIL effectively resolves instruction ambiguities, achieving superior performance to baseline methods. Our implementation is available at https://github.com/RunpengXie/Distributional-Aligned-Learning. Runpeng Xie, Quanwei Wang, Hao Hu 0006, Zherui Zhou, Ni Mu, Xiyun Li, Yiqin Yang, Qianchuan Zhao, Bo Xu 0002 |
NeurIPS | 7 |
| 2025 | Auxiliary Reward Generation With Transition Distance Representation LearningabstractReinforcement learning (RL) has shown strengths in challenging sequential decision-making problems. The reward function in RL is crucial to the learning performance, as it quantifies the degree of task completion. In real-world problems, the rewards are predominantly human-designed, which requires laborious tuning, and is susceptible to human cognitive biases. To achieve automatic auxiliary reward generation, we propose a novel representation learning approach that can measure the “transition distance” between states. Building upon these representations, we introduce an auxiliary reward generation technique for both single-task and skill-chaining scenarios without the need for human knowledge. Furthermore, we theoretically show that the proposed auxiliary rewards maintain the policy invariance property, i.e., the generated rewards will not hurt the policy optimality under the original rewards. In the experiment section, we evaluate the proposed approach in both online and offline learning settings in a wide range of tasks, including robot manipulation and locomotion. The experiment results demonstrate the effectiveness of measuring the transition distance and the induced improvement by auxiliary rewards, which promotes better learning efficiency and increases convergent stability. Beyond that, we demonstrate that the learned manipulation policy with the auxiliary rewards in a simulator can be transferred to the real robot, as shown inhttps://sites.google.com/view/transition-distance-rp/tdrp. Note to Practitioners—The motivation for this paper arises from the need for a technique that enhances robot skill-learning efficiency and performance in both single-task and skill-chaining scenarios. Our research primarily focuses on robot arm manipulation tasks. To accelerate the policy learning process and improve policy performance for executing these tasks, we introduce an auxiliary reward generation technique for both single-task and skill-chaining scenarios without requiring human expertise. This technique leverages the proposed novel representation learning approach, which can measure the “transition distance” between states. During each policy training round, the robot receives a dense reshaped reward created by our approach. Using the policy trained by our method, we successfully control a real Franka Panda robot arm to complete various manipulation tasks. Siyuan Li 0003, Shijie Han, Yingnan Zhao 0002, Yiqin Yang, Qianchuan Zhao, Peng Liu 0008 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2024 | Learning Diverse Risk Preferences in Population-Based Self-PlayabstractAmong the remarkable successes of Reinforcement Learning (RL), self-play algorithms have played a crucial role in solving competitive games. However, current self-play RL methods commonly optimize the agent to maximize the expected win-rates against its current or historical copies, resulting in a limited strategy style and a tendency to get stuck in local optima. To address this limitation, it is important to improve the diversity of policies, allowing the agent to break stalemates and enhance its robustness when facing with different opponents. In this paper, we present a novel perspective to promote diversity by considering that agents could have diverse risk preferences in the face of uncertainty. To achieve this, we introduce a novel reinforcement learning algorithm called Risk-sensitive Proximal Policy Optimization (RPPO), which smoothly interpolates between worst-case and best-case policy learning, enabling policy learning with desired risk preferences. Furthermore, by seamlessly integrating RPPO with population-based self-play, agents in the population optimize dynamic risk-sensitive objectives using experiences gained from playing against diverse opponents. Our empirical results demonstrate that our method achieves comparable or superior performance in competitive games and, importantly, leads to the emergence of diverse behavioral modes. Code is available at https://github.com/Jackory/RPBT. Yuhua Jiang, Qihan Liu, Xiaoteng Ma, Chenghao Li 0002, Yiqin Yang, Jun Yang 0028, Bin Liang 0001, Qianchuan Zhao |
AAAI | 5 |
| 2024 | No Prior Mask: Eliminate Redundant Action for Deep Reinforcement LearningabstractThe large action space is one fundamental obstacle to deploying Reinforcement Learning methods in the real world. The numerous redundant actions will cause the agents to make repeated or invalid attempts, even leading to task failure. Although current algorithms conduct some initial explorations for this issue, they either suffer from rule-based systems or depend on expert demonstrations, which significantly limits their applicability in many real-world settings. In this work, we examine the theoretical analysis of what action can be eliminated in policy optimization and propose a novel redundant action filtering mechanism. Unlike other works, our method constructs the similarity factor by estimating the distance between the state distributions, which requires no prior knowledge. In addition, we combine the modified inverse model to avoid extensive computation in high-dimensional state space. We reveal the underlying structure of action spaces and propose a simple yet efficient redundant action filtering mechanism named No Prior Mask (NPM) based on the above techniques. We show the superior performance of our method by conducting extensive experiments on high-dimensional, pixel-input, and stochastic problems with various action redundancy tasks. Our code is public online at https://github.com/zhongdy15/npm. Dianyu Zhong, Yiqin Yang, Qianchuan Zhao |
AAAI | 2 |
| 2024 | Bayesian Design Principles for Offline-to-Online Reinforcement LearningabstractOffline reinforcement learning (RL) is crucial for real-world applications where exploration can be costly or unsafe. However, offline learned policies are often suboptimal, and further online fine-tuning is required. In this paper, we tackle the fundamental dilemma of offline-to-online fine-tuning: if the agent remains pessimistic, it may fail to learn a better policy, while if it becomes optimistic directly, performance may suffer from a sudden drop. We show that Bayesian design principles are crucial in solving such a dilemma. Instead of adopting optimistic or pessimistic policies, the agent should act in a way that matches its belief in optimal policies. Such a probability-matching agent can avoid a sudden performance drop while still being guaranteed to find the optimal policy. Based on our theoretical findings, we introduce a novel algorithm that outperforms existing methods on various benchmarks, demonstrating the efficacy of our approach. Overall, the proposed approach provides a new perspective on offline-to-online RL that has the potential to enable more effective learning from offline data. Hao Hu 0006, Yiqin Yang, Jianing Ye, Chengjie Wu, Ziqing Mai, Yujing Hu, Tangjie Lv, Changjie Fan, Qianchuan Zhao, Chongjie Zhang |
ICML | 2 |
| 2024 | Planning, Fast and Slow: Online Reinforcement Learning with Action-Free Offline Data via Multiscale PlannersabstractThe surge in volumes of video data offers unprecedented opportunities for advancing reinforcement learning (RL). This growth has motivated the development of passive RL, seeking to convert passive observations into actionable insights. This paper explores the prerequisites and mechanisms through which passive data can be utilized to improve online RL. We show that, in identifiable dynamics, where action impact can be distinguished from stochasticity, learning on passive data is statistically beneficial. Building upon the theoretical insights, we propose a novel algorithm named Multiscale State-Centric Planners (MSCP) that leverages two planners at distinct scales to offer guidance across varying levels of abstraction. The algorithm’s fast planner targets immediate objectives, while the slow planner focuses on achieving longer-term goals. Notably, the fast planner incorporates pessimistic regularization to address the distributional shift between offline and online data. MSCP effectively handles the practical challenges involving imperfect pretraining and limited dataset coverage. Our empirical evaluations across multiple benchmarks demonstrate that MSCP significantly outperforms existing approaches, underscoring its proficiency in addressing complex, long-horizon tasks through the strategic use of passive data. Chengjie Wu, Hao Hu 0006, Yiqin Yang, Ning Zhang 0017, Chongjie Zhang |
ICML | 3 |
| 2023 | Flow to Control: Offline Reinforcement Learning with Lossless Primitive DiscoveryabstractOffline reinforcement learning (RL) enables the agent to effectively learn from logged data, which significantly extends the applicability of RL algorithms in real-world scenarios where exploration can be expensive or unsafe. Previous works have shown that extracting primitive skills from the recurring and temporally extended structures in the logged data yields better learning. However, these methods suffer greatly when the primitives have limited representation ability to recover the original policy space, especially in offline settings. In this paper, we give a quantitative characterization of the performance of offline hierarchical learning and highlight the importance of learning lossless primitives. To this end, we propose to use a flow-based structure as the representation for low-level policies. This allows us to represent the behaviors in the dataset faithfully while keeping the expression ability to recover the whole policy space. We show that such lossless primitives can drastically improve the performance of hierarchical policies. The experimental results and extensive ablation studies on the standard D4RL benchmark show that our method has a good representation ability for policies and achieves superior performance in most tasks. Yiqin Yang, Hao Hu 0006, Siyuan Li 0003, Jun Yang 0028, Qianchuan Zhao, Chongjie Zhang |
AAAI | 1 |
| 2023 | The Provable Benefit of Unsupervised Data Sharing for Offline Reinforcement Learning
Hao Hu 0006, Yiqin Yang, Qianchuan Zhao, Chongjie Zhang |
ICLR | 2 |
| 2023 | Unsupervised Behavior Extraction via Random Intent PriorsabstractReward-free data is abundant and contains rich prior knowledge of human behaviors, but it is not well exploited by offline reinforcement learning (RL) algorithms. In this paper, we propose UBER, an unsupervised approach to extract useful behaviors from offline reward-free datasets via diversified rewards. UBER assigns different pseudo-rewards sampled from a given prior distribution to different agents to extract a diverse set of behaviors, and reuse them as candidate policies to facilitate the learning of new tasks. Perhaps surprisingly, we show that rewards generated from random neural networks are sufficient to extract diverse and useful behaviors, some even close to expert ones. We provide both empirical and theoretical evidences to justify the use of random priors for the reward function. Experiments on multiple benchmarks showcase UBER's ability to learn effective and diverse behavior sets that enhance sample efficiency for online RL, outperforming existing baselines. By reducing reliance on human supervision, UBER broadens the applicability of RL to real-world scenarios with abundant reward-free data. Hao Hu 0006, Yiqin Yang, Jianing Ye, Ziqing Mai, Chongjie Zhang |
NeurIPS | 2 |
| 2022 | Offline Reinforcement Learning with Value-based Episodic Memory
Xiaoteng Ma, Yiqin Yang, Hao Hu 0006, Jun Yang 0028, Chongjie Zhang, Qianchuan Zhao, Bin Liang 0001, Qihan Liu |
ICLR | 2 |
| 2022 | On the Role of Discount Factor in Offline Reinforcement LearningabstractOffline reinforcement learning (RL) enables effective learning from previously collected data without exploration, which shows great promise in real-world applications when exploration is expensive or even infeasible. The discount factor, $\gamma$, plays a vital role in improving online RL sample efficiency and estimation accuracy, but the role of the discount factor in offline RL is not well explored. This paper examines two distinct effects of $\gamma$ in offline RL with theoretical analysis, namely the regularization effect and the pessimism effect. On the one hand, $\gamma$ is a regulator to trade-off optimality with sample efficiency upon existing offline techniques. On the other hand, lower guidance $\gamma$ can also be seen as a way of pessimism where we optimize the policy’s performance in the worst possible models. We empirically verify the above theoretical observation with tabular MDPs and standard D4RL tasks. The results show that the discount factor plays an essential role in the performance of offline RL algorithms, both under small data regimes upon existing offline methods and in large data regimes without other conservative methods. Hao Hu 0006, Yiqin Yang, Qianchuan Zhao, Chongjie Zhang |
ICML | 2 |
| 2021 | Believe What You See: Implicit Constraint Approach for Offline Multi-Agent Reinforcement LearningabstractLearning from datasets without interaction with environments (Offline Learning) is an essential step to apply Reinforcement Learning (RL) algorithms in real-world scenarios.However, compared with the single-agent counterpart, offline multi-agent RL introduces more agents with the larger state and action space, which is more challenging but attracts little attention. We demonstrate current offline RL algorithms are ineffective in multi-agent systems due to the accumulated extrapolation error. In this paper, we propose a novel offline RL algorithm, named Implicit Constraint Q-learning (ICQ), which effectively alleviates the extrapolation error by only trusting the state-action pairs given in the dataset for value estimation. Moreover, we extend ICQ to multi-agent tasks by decomposing the joint-policy under the implicit constraint. Experimental results demonstrate that the extrapolation error is successfully controlled within a reasonable range and insensitive to the number of agents. We further show that ICQ achieves the state-of-the-art performance in the challenging multi-agent offline tasks (StarCraft II). Our code is public online at https://github.com/YiqinYang/ICQ. Yiqin Yang, Xiaoteng Ma, Chenghao Li 0002, Zewu Zheng, Gao Huang 0001, Jun Yang 0028, Qianchuan Zhao |
NeurIPS | 1 |
| 2021 | MHRWR: Prediction of lncRNA-Disease Associations Based on Multiple Heterogeneous NetworksabstractIn the last few years, accumulating evidences had demonstrated that long non-coding RNAs (lncRNAs) participated in the regulation of target gene expression and played an important role in biological processes and human disease development. Thus, prediction of the associations between lncRNAs and disease had become a hot research in the fields of human sophisticated diseases. Most of these methods considered the information of two networks (lncRNA, disease) while neglected other networks. In this study, we designed a multi-layer network by integrating the similarity networks of lncRNAs, diseases and genes, and the known association networks of lncRNA-disease, lncRNAs-gene, and disease-gene, and then we developed a model called MHRWR for predicting the lncRNA-disease potential associations based on random walk with restart. The performance of MHRWR was evaluated by experimentally verified lncRNA-disease associations based on leave-one-out cross validation. MHRWR obtained a reliable AUC value of 0.91344, which significantly outperformed some previous methods. To further validate the reproducibility of performance, we used the model of MHRWR to verify related lncRNAs of colon cancer, colorectal cancer and lung adenocarcinoma in the case studies. The codes of MHRWR is available on: https://github.com/yangyq505/MHRWR. Xiaowei Zhao 0004, Yiqin Yang, Minghao Yin |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2018 | Deep Learning Technique-Based Steering of Autonomous CarabstractDeep neural network (DNN) has many advantages. Autonomous driving has become a popular topic now. In this paper, an improved stack autoencoder based on the deep learning techniques is proposed to learn the driving characteristics of an autonomous car. These techniques realize the input data adjustment and solving diffusion gradient problem. A Raspberry Pi and a camera module are mounted on the top of the car. The camera module provides the images needed for training the DNN. There are two stages in the training. In the pre-training process, an improved autoencoder is trained by the unsupervised learning mechanism, and the characterization of the track is extracted. In the fine-tuning stage, the whole network is trained according to the labeled data, and then this model learns the driving characteristics better according to the samples. In the experimental stage, the car will predict the action of the car by the trained model in the autonomous mode. The experiment exhibits the effectiveness of the proposed model. Compared with the traditional neural network, the improved stack autoencoder has a better generalization ability and faster convergence speed. Yiqin Yang, Qingyang Xu, Fabao Yan |
Int. J. Comput. Intell. Appl. | 1 |