VLDB 2026 Research / reviewers in the wild / expert
Jiafei Lyu
dblp:278/1503
· DBLP profile ↗
29ranked-venue papers
11as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 9 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative ReasoningabstractRecent advancements in Large Language Models (LLMs) have shown that it is promising to utilize Process Reward Models (PRMs) as verifiers to enhance the performance of LLMs. However, current PRMs face three key challenges: (1) limited process supervision and generalization capabilities, (2) dependence on scalar value prediction without leveraging the generative abilities of LLMs, and (3) inability to scale the test-time compute of PRMs. In this work, we introduce GenPRM, a generative process reward model that performs explicit Chain-of-Thought (CoT) reasoning with code verification before providing judgment for each reasoning step. To obtain high-quality process supervision labels and rationale data, we propose Relative Progress Estimation (RPE) and a rationale synthesis framework that incorporates code verification. Experimental results on ProcessBench and several mathematical reasoning tasks show that GenPRM significantly outperforms prior PRMs with only 23K training data from MATH dataset. Through test-time scaling, a 1.5B GenPRM outperforms GPT-4o, and a 7B GenPRM surpasses Qwen2.5-Math-PRM-72B on ProcessBench. Additionally, GenPRM demonstrates strong abilities to serve as a critic model for policy model refinement. This work establishes a new paradigm for process supervision that bridges the gap between PRMs and critic models in LLMs. Jian Zhao 0006, Runze Liu 0002, Zhimu Zhou, Junqi Gao, Dong Li 0016, Jiafei Lyu, Zhouyi Qian, Biqing Qi, Xiu Li 0001, Bowen Zhou 0002 |
AAAI | 7 |
| 2026 | Temporal difference learning with constrained initial representations
Jiafei Lyu, Zhongjian Qiao, Runze Liu 0002, Zeyuan Liu, Deheng Ye, Zongqing Lu 0002, Xiu Li 0001 |
Inf. Sci. | 1 |
| 2025 | Novelty-Guided Data Reuse for Efficient and Diversified Multi-Agent Reinforcement LearningabstractRecently, deep Multi-Agent Reinforcement Learning (MARL) has demonstrated its potential to tackle complex cooperative tasks, pushing the boundaries of AI in collaborative environments. However, the efficiency of these systems is often compromised by inadequate sample utilization and a lack of diversity in learning strategies. To enhance MARL performance, we introduce a novel sample reuse approach that dynamically adjusts policy updates based on observation novelty. Specifically, we employ a Random Network Distillation (RND) network to gauge the novelty of each agent's current state, assigning additional sample update opportunities based on the uniqueness of the data. We name our method Multi-Agent Novelty-GuidEd sample Reuse (MANGER). This method increases sample efficiency while promoting exploration and diverse agent behaviors. Our evaluations confirm substantial improvements in MARL effectiveness in complex cooperative scenarios such as Google Research Football and super-hard StarCraft II micromanagement tasks. Yangkun Chen, Jiafei Lyu |
AAAI | 4 |
| 2025 | SUMO: Search-Based Uncertainty Estimation for Model-Based Offline Reinforcement LearningabstractThe performance of offline reinforcement learning (RL) suffers from the limited size and quality of static datasets. Model-based offline RL addresses this issue by generating synthetic samples through a dynamics model to enhance overall performance. To evaluate the reliability of the generated samples, uncertainty estimation methods are often employed. However, model ensemble, the most commonly used uncertainty estimation method, is not always the best choice. In this paper, we propose a Search-based Uncertainty estimation method for Model-based Offline RL (SUMO) as an alternative. SUMO characterizes the uncertainty of synthetic samples by measuring their cross entropy against the in-distribution dataset samples, and uses an efficient search-based method for implementation. In this way, SUMO can achieve trustworthy uncertainty estimation. We integrate SUMO into several model-based offline RL algorithms including MOPO and Adapted MOReL (AMOReL), and provide theoretical analysis for them. Extensive experimental results on D4RL datasets demonstrate that SUMO can provide accurate uncertainty estimation and boost the performance of base algorithms. These indicate that SUMO could be a better uncertainty estimator for model-based offline RL when used in either reward penalty or trajectory truncation. Zhongjian Qiao, Jiafei Lyu, Kechen Jiao, Xiu Li 0001 |
AAAI | 2 |
| 2025 | VLP: Vision-Language Preference Learning for Embodied ManipulationabstractReward engineering is one of the key challenges in Reinforcement Learning (RL).Preference-based RL effectively addresses this issue by learning from human feedback.However, it is both time-consuming and expensive to collect human preference labels.In this paper, we propose a novel Vision-Language Preference learning framework, named VLP, which learns a vision-language preference model to provide feedback for embodied manipulation tasks.To achieve this, we define three types of language-conditioned preferences and construct a vision-language preference dataset, which contains versatile implicit preference orders.The model learns to extract languagerelated features, and then serves as a predictor in various downstream tasks.The policy can be learned according to the annotated labels via reward learning or direct policy optimization.Extensive empirical results on simulated embodied manipulation tasks demonstrate that our method provides accurate preferences and generalizes to unseen tasks and unseen language instructions, outperforming the baselines by a large margin and shifting the burden from continuous, per-task human annotation to one-time, per-domain data collection. Runze Liu 0002, Chenjia Bai, Jiafei Lyu, Shengjie Sun 0002, Yali Du 0001, Xiu Li 0001 |
EMNLP | 3 |
| 2025 | Cross-Domain Offline Policy Adaptation with Optimal Transport and Dataset ConstraintabstractWe explore cross-domain offline reinforcement learning (RL) where offline datasets from another domain can be accessed to facilitate policy learning. However, the underlying environments of the two datasets may have dynamics mismatches, incurring inferior performance when simply merging the data of two domains. Existing methods mitigate this issue by training domain classifiers, using contrastive learning methods, etc. Nevertheless, they still rely on a large amount of target domain data to function well. Instead, we address this problem by establishing a concrete performance bound of a policy given datasets from two domains. Motivated by the theoretical insights, we propose to align transitions in the two datasets using optimal transport and selectively share source domain samples, without training any neural networks. This enables reliable data filtering even given a few target domain data. Additionally, we introduce a dataset regularization term that ensures the learned policy remains within the scope of the target domain dataset, preventing it from being biased towards the source domain data. Consequently, we propose the Optimal Transport Data Filtering (dubbed OTDF) method and examine its effectiveness by conducting extensive experiments across various dynamics shift conditions (e.g., gravity shift), given limited target domain data. It turns out that OTDF exhibits superior performance on many tasks and dataset qualities, often surpassing prior strong baselines by a large margin. Jiafei Lyu, Mengbei Yan, Zhongjian Qiao, Runze Liu 0002, Xiaoteng Ma, Deheng Ye, Zongqing Lu 0002, Xiu Li 0001 |
ICLR | 1 |
| 2025 | Leveraging Score-based Models for Generating Penalization in Model-based Offline Reinforcement Learning
Zeyuan Liu, Zhirui Fang, Jiafei Lyu, Xiu Li 0001 |
AAMAS | 3 |
| 2025 | CDSA: Conservative Denoising Score-based Algorithm for Offline Reinforcement Learning
Zeyuan Liu, Kai Yang 0050, Jiafei Lyu, Xiu Li 0001 |
AAMAS | 3 |
| 2025 | World Models with Hints of Large Language Models for Goal AchievingabstractZeyuan Liu, Ziyu Huan, Xiyao Wang, Jiafei Lyu, Jian Tao, Xiu Li, Furong Huang, Huazhe Xu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zeyuan Liu, Ziyu Huan, Jiafei Lyu, Xiu Li 0001, Furong Huang, Huazhe Xu |
NAACL (Long Papers) | 4 |
| 2025 | ADG: Ambient Diffusion-Guided Dataset Recovery for Corruption-Robust Offline Reinforcement LearningabstractReal-world datasets collected from sensors or human inputs are prone to noise and errors, posing significant challenges for applying offline reinforcement learning (RL). While existing methods have made progress in addressing corrupted actions and rewards, they remain insufficient for handling corruption in high-dimensional state spaces and for cases where multiple elements in the dataset are corrupted simultaneously. Diffusion models, known for their strong denoising capabilities, offer a promising direction for this problem—but their tendency to overfit noisy samples limits their direct applicability.
To overcome this, we propose **A**mbient **D**iffusion-**G**uided Dataset Recovery (**ADG**), a novel approach that pioneers the use of diffusion models to tackle data corruption in offline RL. First, we introduce Ambient Denoising Diffusion Probabilistic Models (DDPM) from approximated distributions, which enable learning on partially corrupted datasets with theoretical guarantees. Second, we use the noise-prediction property of Ambient DDPM to distinguish between clean and corrupted data, and then use the clean subset to train a standard DDPM. Third, we employ the trained standard DDPM to refine the previously identified corrupted data, enhancing data quality for subsequent offline RL training. A notable strength of ADG is its versatility—it can be seamlessly integrated with any offline RL algorithm. Experiments on a range of benchmarks, including MuJoCo, Kitchen, and Adroit, demonstrate that ADG effectively mitigates the impact of corrupted data and improves the robustness of offline RL under various noise settings, achieving state-of-the-art results. Zeyuan Liu, Zhihe Yang, Rui Yang 0010, Jiafei Lyu, Baoxiang Wang 0001, Yunjian Xu, Xiu Li 0001 |
NeurIPS | 5 |
| 2025 | A large language model-driven reward design framework via dynamic feedback for reinforcement learning
Shengjie Sun 0002, Runze Liu 0002, Jiafei Lyu, Liangpeng Zhang, Xiu Li 0001 |
Knowl. Based Syst. | 3 |
| 2024 | Using Human Feedback to Fine-tune Diffusion Models without Any Reward ModelabstractUsing reinforcement learning with human feedback (RLHF) has shown significant promise in fine-tuning diffusion models. Previous methods start by training a reward model that aligns with human preferences, then leverage RL techniques to fine-tune the underlying models. However, crafting an efficient reward model demands extensive datasets, optimal architecture, and manual hyperparameter tuning, making the process both time and cost-intensive. The direct preference optimization (DPO) method, effective in fine-tuning large language models, eliminates the necessity for a reward model. However, the extensive GPU memory requirement of the diffusion model's denoising process hinders the direct application of the DPO method. To address this issue, we introduce the Direct Preference for De-noising Diffusion Policy Optimization (D3PO) method to directly fine-tune diffusion models. The theoretical analysis demonstrates that although D3PO omits training a reward model, it effectively functions as the optimal re-ward model trained using human feedback data to guide the learning process. This approach requires no training of a reward model, proving to be more direct, cost-effective, and minimizing computational overhead. In experiments, our method uses the relative scale of objectives as a proxy for human preference, delivering comparable results to methods using ground-truth rewards. Moreover, D3PO demonstrates the ability to reduce image distortion rates and generate safer images, overcoming challenges lacking robust reward models. Our code is publicly available at https://github.com/yk7333/D3PO. Kai Yang 0050, Jiafei Lyu, Chunjiang Ge, Weihan Shen, Xiu Li 0001 |
CVPR | 3 |
| 2024 | Mind the Model, Not the Agent: The Primacy Bias in Model-Based RLabstractThe primacy bias in model-free reinforcement learning (MFRL), which refers to the agent’s tendency to overfit early data and lose the ability to learn from new data, can significantly decrease the performance of MFRL algorithms. Previous studies have shown that employing simple techniques, such as resetting the agent’s parameters, can substantially alleviate the primacy bias in MFRL. However, the primacy bias in model-based reinforcement learning (MBRL) remains unexplored. In this work, we focus on investigating the primacy bias in MBRL. We begin by observing that resetting the agent’s parameters harms its performance in the context of MBRL. We further find that the primacy bias in MBRL is more closely related to the primacy bias of the world model instead of the primacy bias of the agent. Based on this finding, we propose world model resetting, a simple yet effective technique to alleviate the primacy bias in MBRL. We apply our method to two different MBRL algorithms, MBPO and DreamerV2. We validate the effectiveness of our method on multiple continuous control tasks on MuJoCo and DeepMind Control Suite, as well as discrete control tasks on Atari 100k benchmark. The experimental results show that world model resetting can significantly alleviate the primacy bias in the model-based setting and improve the algorithm’s performance. We also give a guide on how to perform world model resetting effectively. Zhongjian Qiao, Jiafei Lyu, Xiu Li 0001 |
ECAI | 2 |
| 2024 | Enhancing Visual Generalization in Reinforcement Learning with Cycling Augmentation
Shengjie Sun 0002, Jiafei Lyu, Jiazhe Guo, Mengbei Yan, Runze Liu 0002, Xiu Li 0001 |
ICANN (4) | 2 |
| 2024 | SEABO: A Simple Search-Based Method for Offline Imitation LearningabstractOffline reinforcement learning (RL) has attracted much attention due to its ability in learning from static offline datasets and eliminating the need of interacting with the environment. Nevertheless, the success of offline RL relies heavily on the offline transitions annotated with reward labels. In practice, we often need to hand-craft the reward function, which is sometimes difficult, labor-intensive, or inefficient. To tackle this challenge, we set our focus on the offline imitation learning (IL) setting, and aim at getting a reward function based on the expert data and unlabeled data. To that end, we propose a simple yet effective search-based offline IL method, tagged SEABO. SEABO allocates a larger reward to the transition that is close to its closest neighbor in the expert demonstration, and a smaller reward otherwise, all in an unsupervised learning manner. Experimental results on a variety of D4RL datasets indicate that SEABO can achieve competitive performance to offline RL algorithms with ground-truth rewards, given only a single expert trajectory, and can outperform prior reward learning and offline IL methods across many tasks. Moreover, we demonstrate that SEABO also works well if the expert demonstrations contain only observations. Our code is publicly available at https://github.com/dmksjfl/SEABO. Jiafei Lyu, Xiaoteng Ma, Le Wan, Runze Liu 0002, Xiu Li 0001, Zongqing Lu 0002 |
ICLR | 1 |
| 2024 | PEARL: Zero-shot Cross-task Preference Alignment and Robust Reward Learning for Robotic ManipulationabstractIn preference-based Reinforcement Learning (RL), obtaining a large number of preference labels are both time-consuming and costly. Furthermore, the queried human preferences cannot be utilized for the new tasks. In this paper, we propose Zero-shot Cross-task Preference Alignment and Robust Reward Learning (PEARL), which learns policies from cross-task preference transfer without any human labels of the target task. Our contributions include two novel components that facilitate the transfer and learning process. The first is Cross-task Preference Alignment (CPA), which transfers the preferences between tasks via optimal transport. The key idea of CPA is to use Gromov-Wasserstein distance to align the trajectories between tasks, and the solved optimal transport matrix serves as the correspondence between trajectories. The target task preferences are computed as the weighted sum of source task preference labels with the correspondence as weights. Moreover, to ensure robust learning from these transferred labels, we introduce Robust Reward Learning (RRL), which considers both reward mean and uncertainty by modeling rewards as Gaussian distributions. Empirical results on robotic manipulation tasks from Meta-World and Robomimic demonstrate that our method is capable of transferring preference labels across tasks accurately and then learns well-behaved policies. Notably, our approach significantly exceeds existing methods when there are few human preferences. The code and videos of our method are available at: https://sites.google.com/view/pearl-preference. Runze Liu 0002, Yali Du 0001, Fengshuo Bai, Jiafei Lyu, Xiu Li 0001 |
ICML | 4 |
| 2024 | Cross-Domain Policy Adaptation by Capturing Representation MismatchabstractIt is vital to learn effective policies that can be transferred to different domains with dynamics discrepancies in reinforcement learning (RL). In this paper, we consider dynamics adaptation settings where there exists dynamics mismatch between the source domain and the target domain, and one can get access to sufficient source domain data, while can only have limited interactions with the target domain. Existing methods address this problem by learning domain classifiers, performing data filtering from a value discrepancy perspective, etc. Instead, we tackle this challenge from a decoupled representation learning perspective. We perform representation learning only in the target domain and measure the representation deviations on the transitions from the source domain, which we show can be a signal of dynamics mismatch. We also show that representation deviation upper bounds performance difference of a given policy in the source domain and target domain, which motivates us to adopt representation deviation as a reward penalty. The produced representations are not involved in either policy or value function, but only serve as a reward penalizer. We conduct extensive experiments on environments with kinematic and morphology mismatch, and the results show that our method exhibits strong performance on many tasks. Our code is publicly available at https://github.com/dmksjfl/PAR. Jiafei Lyu, Chenjia Bai, Zongqing Lu 0002, Xiu Li 0001 |
ICML | 1 |
| 2024 | Exploration and Anti-Exploration with Distributional Random Network DistillationabstractExploration remains a critical issue in deep reinforcement learning for an agent to attain high returns in unknown environments. Although the prevailing exploration Random Network Distillation (RND) algorithm has been demonstrated to be effective in numerous environments, it often needs more discriminative power in bonus allocation. This paper highlights the ``bonus inconsistency'' issue within RND, pinpointing its primary limitation. To address this issue, we introduce the Distributional RND (DRND), a derivative of the RND. DRND enhances the exploration process by distilling a distribution of random networks and implicitly incorporating pseudo counts to improve the precision of bonus allocation. This refinement encourages agents to engage in more extensive exploration. Our method effectively mitigates the inconsistency issue without introducing significant computational overhead. Both theoretical analysis and experimental results demonstrate the superiority of our approach over the original RND algorithm. Our method excels in challenging online exploration scenarios and effectively serves as an anti-exploration mechanism in D4RL offline tasks. Our code is publicly available at https://github.com/yk7333/DRND. Kai Yang 0050, Jiafei Lyu, Xiu Li 0001 |
ICML | 3 |
| 2024 | ODRL: A Benchmark for Off-Dynamics Reinforcement LearningabstractWe consider off-dynamics reinforcement learning (RL) where one needs to transfer policies across different domains with dynamics mismatch. Despite the focus on developing dynamics-aware algorithms, this field is hindered due to the lack of a standard benchmark. To bridge this gap, we introduce ODRL, the first benchmark tailored for evaluating off-dynamics RL methods. ODRL contains four experimental settings where the source and target domains can be either online or offline, and provides diverse tasks and a broad spectrum of dynamics shifts, making it a reliable platform to comprehensively evaluate the agent's adaptation ability to the target domain. Furthermore, ODRL includes recent off-dynamics RL algorithms in a unified framework and introduces some extra baselines for different settings, all implemented in a single-file manner. To unpack the true adaptation capability of existing methods, we conduct extensive benchmarking experiments, which show that no method has universal advantages across varied dynamics shifts. We hope this benchmark can serve as a cornerstone for future research endeavors. Our code is publicly available at https://github.com/OffDynamicsRL/off-dynamics-rl. Jiafei Lyu, Jiacheng Xu 0003, Mengbei Yan, Zongzhang Zhang, Chenjia Bai, Zongqing Lu 0002, Xiu Li 0001 |
NeurIPS | 1 |
| 2024 | A two-stage reinforcement learning-based approach for multi-entity task allocation
Aicheng Gong, Kai Yang 0050, Jiafei Lyu, Xiu Li 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Off-policy RL algorithms can be sample-efficient for continuous control via sample multiple reuse
Jiafei Lyu, Le Wan, Xiu Li 0001, Zongqing Lu 0002 |
Inf. Sci. | 1 |
| 2024 | Understanding What Affects the Generalization Gap in Visual Reinforcement Learning: Theory and Empirical EvidenceabstractRecently, there are many efforts attempting to learn useful policies for continuous control in visual reinforcement learning (RL). In this scenario, it is important to learn a generalizable policy, as the testing environment may differ from the training environment, e.g., there exist distractors during deployment. Many practical algorithms are proposed to handle this problem. However, to the best of our knowledge, none of them provide a theoretical understanding of what affects the generalization gap and why their proposed methods work. In this paper, we bridge this issue by theoretically answering the key factors that contribute to the generalization gap when the testing environment has distractors. Our theories indicate that minimizing the representation distance between training and testing environments, which aligns with human intuition, is the most critical for the benefit of reducing the generalization gap. Our theoretical results are supported by the empirical evidence in the DMControl Generalization Benchmark (DMC-GB). Jiafei Lyu, Le Wan, Xiu Li 0001, Zongqing Lu 0002 |
J. Artif. Intell. Res. | 1 |
| 2024 | Enhancing visual reinforcement learning with State-Action Representation
Mengbei Yan, Jiafei Lyu, Xiu Li 0001 |
Knowl. Based Syst. | 2 |
| 2023 | Uncertainty-Driven Trajectory Truncation for Data Augmentation in Offline Reinforcement LearningabstractEquipped with the trained environmental dynamics, model-based offline reinforcement learning (RL) algorithms can often successfully learn good policies from fixed-sized datasets, even some datasets with poor quality. Unfortunately, however, it can not be guaranteed that the generated samples from the trained dynamics model are reliable (e.g., some synthetic samples may lie outside of the support region of the static dataset). To address this issue, we propose Trajectory Truncation with Uncertainty (TATU), which adaptively truncates the synthetic trajectory if the accumulated uncertainty along the trajectory is too large. We theoretically show the performance bound of TATU to justify its benefits. To empirically show the advantages of TATU, we first combine it with two classical model-based offline RL algorithms, MOPO and COMBO. Furthermore, we integrate TATU with several off-the-shelf model-free offline RL algorithms, e.g., BCQ. Experimental results on the D4RL benchmark show that TATU significantly improves their performance, often by a large margin. Code is available here. Jiafei Lyu, Xiaoteng Ma, Jiangpeng Yan, Jun Yang 0028, Le Wan, Xiu Li 0001 |
ECAI | 2 |
| 2023 | Value activation for bias alleviation: Generalized-activated deep double deterministic policy gradients
Jiafei Lyu, Yu Yang 0016, Jiangpeng Yan, Xiu Li 0001 |
Neurocomputing | 1 |
| 2022 | Efficient Continuous Control with Double Actors and Regularized CriticsabstractHow to obtain good value estimation is a critical problem in Reinforcement Learning (RL). Current value estimation methods in continuous control, such as DDPG and TD3, suffer from unnecessary over- or under- estimation. In this paper, we explore the potential of double actors, which has been neglected for a long time, for better value estimation in the continuous setting. First, we interestingly find that double actors improve the exploration ability of the agent. Next, we uncover the bias alleviation property of double actors in handling overestimation with single critic, and underestimation with double critics respectively. Finally, to mitigate the potentially pessimistic value estimate in double critics, we propose to regularize the critics under double actors architecture. Together, we present Double Actors Regularized Critics (DARC) algorithm. Extensive experiments on challenging continuous control benchmarks, MuJoCo and PyBullet, show that DARC significantly outperforms current baselines with higher average return and better sample efficiency. Jiafei Lyu, Xiaoteng Ma, Jiangpeng Yan, Xiu Li 0001 |
AAAI | 1 |
| 2022 | Double Check Your State Before Trusting It: Confidence-Aware Bidirectional Offline Model-Based ImaginationabstractThe learned policy of model-free offline reinforcement learning (RL) methods is often constrained to stay within the support of datasets to avoid possible dangerous out-of-distribution actions or states, making it challenging to handle out-of-support region. Model-based RL methods offer a richer dataset and benefit generalization by generating imaginary trajectories with either trained forward or reverse dynamics model. However, the imagined transitions may be inaccurate, thus downgrading the performance of the underlying offline RL method. In this paper, we propose to augment the offline dataset by using trained bidirectional dynamics models and rollout policies with double check. We introduce conservatism by trusting samples that the forward model and backward model agree on. Our method, confidence-aware bidirectional offline model-based imagination, generates reliable samples and can be combined with any model-free offline RL method. Experimental results on the D4RL benchmarks demonstrate that our method significantly boosts the performance of existing model-free offline RL algorithms and achieves competitive or better scores against baseline methods. Jiafei Lyu, Xiu Li 0001, Zongqing Lu 0002 |
NeurIPS | 1 |
| 2022 | Mildly Conservative Q-Learning for Offline Reinforcement LearningabstractOffline reinforcement learning (RL) defines the task of learning from a static logged dataset without continually interacting with the environment. The distribution shift between the learned policy and the behavior policy makes it necessary for the value function to stay conservative such that out-of-distribution (OOD) actions will not be severely overestimated. However, existing approaches, penalizing the unseen actions or regularizing with the behavior policy, are too pessimistic, which suppresses the generalization of the value function and hinders the performance improvement. This paper explores mild but enough conservatism for offline learning while not harming generalization. We propose Mildly Conservative Q-learning (MCQ), where OOD actions are actively trained by assigning them proper pseudo Q values. We theoretically show that MCQ induces a policy that behaves at least as well as the behavior policy and no erroneous overestimation will occur for OOD actions. Experimental results on the D4RL benchmarks demonstrate that MCQ achieves remarkable performance compared with prior work. Furthermore, MCQ shows superior generalization ability when transferring from offline to online, and significantly outperforms baselines. Our code is publicly available at https://github.com/dmksjfl/MCQ. Jiafei Lyu, Xiaoteng Ma, Xiu Li 0001, Zongqing Lu 0002 |
NeurIPS | 1 |
| 2022 | PRAG: Periodic Regularized Action Gradient for Efficient Continuous Control
Xihui Li, Zhongjian Qiao, Aicheng Gong, Jiafei Lyu, ChengHui Yu, Jiangpeng Yan, Xiu Li 0001 |
PRICAI (3) | 4 |