VLDB 2026 Research / reviewers in the wild / expert
Runze Liu 0002
dblp:235/0682-2
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
0009-0007-4784-5333ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative ReasoningabstractRecent advancements in Large Language Models (LLMs) have shown that it is promising to utilize Process Reward Models (PRMs) as verifiers to enhance the performance of LLMs. However, current PRMs face three key challenges: (1) limited process supervision and generalization capabilities, (2) dependence on scalar value prediction without leveraging the generative abilities of LLMs, and (3) inability to scale the test-time compute of PRMs. In this work, we introduce GenPRM, a generative process reward model that performs explicit Chain-of-Thought (CoT) reasoning with code verification before providing judgment for each reasoning step. To obtain high-quality process supervision labels and rationale data, we propose Relative Progress Estimation (RPE) and a rationale synthesis framework that incorporates code verification. Experimental results on ProcessBench and several mathematical reasoning tasks show that GenPRM significantly outperforms prior PRMs with only 23K training data from MATH dataset. Through test-time scaling, a 1.5B GenPRM outperforms GPT-4o, and a 7B GenPRM surpasses Qwen2.5-Math-PRM-72B on ProcessBench. Additionally, GenPRM demonstrates strong abilities to serve as a critic model for policy model refinement. This work establishes a new paradigm for process supervision that bridges the gap between PRMs and critic models in LLMs. Jian Zhao 0006, Runze Liu 0002, Zhimu Zhou, Junqi Gao, Dong Li 0016, Jiafei Lyu, Zhouyi Qian, Biqing Qi, Xiu Li 0001, Bowen Zhou 0002 |
AAAI | 2 |
| 2026 | Temporal difference learning with constrained initial representations
Jiafei Lyu, Zhongjian Qiao, Runze Liu 0002, Zeyuan Liu, Deheng Ye, Zongqing Lu 0002, Xiu Li 0001 |
Inf. Sci. | 4 |
| 2025 | RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted BehaviorsabstractEvaluating deep reinforcement learning (DRL) agents against targeted behavior attacks is critical for assessing their robustness. These attacks aim to manipulate the victim into specific behaviors that align with the attacker’s objectives, often bypassing traditional reward-based defenses. Prior methods have primarily focused on reducing cumulative rewards; however, rewards are typically too generic to capture complex safety requirements effectively. As a result, focusing solely on reward reduction can lead to suboptimal attack strategies, particularly in safety-critical scenarios where more precise behavior manipulation is needed. To address these challenges, we propose RAT, a method designed for universal, targeted behavior attacks. RAT trains an intention policy that is explicitly aligned with human preferences, serving as a precise behavioral target for the adversary. Concurrently, an adversary manipulates the victim's policy to follow this target behavior. To enhance the effectiveness of these attacks, RAT dynamically adjusts the state occupancy measure within the replay buffer, allowing for more controlled and effective behavior manipulation. Our empirical results on robotic simulation tasks demonstrate that RAT outperforms existing adversarial attack algorithms in inducing specific behaviors. Additionally, RAT shows promise in improving agent robustness, leading to more resilient policies. We further validate RAT by guiding Decision Transformer agents to adopt behaviors aligned with human preferences in various MuJoCo tasks, demonstrating its effectiveness across diverse tasks. Fengshuo Bai, Runze Liu 0002, Yali Du 0001, Ying Wen 0001, Yaodong Yang 0001 |
AAAI | 2 |
| 2025 | VLP: Vision-Language Preference Learning for Embodied ManipulationabstractReward engineering is one of the key challenges in Reinforcement Learning (RL).Preference-based RL effectively addresses this issue by learning from human feedback.However, it is both time-consuming and expensive to collect human preference labels.In this paper, we propose a novel Vision-Language Preference learning framework, named VLP, which learns a vision-language preference model to provide feedback for embodied manipulation tasks.To achieve this, we define three types of language-conditioned preferences and construct a vision-language preference dataset, which contains versatile implicit preference orders.The model learns to extract languagerelated features, and then serves as a predictor in various downstream tasks.The policy can be learned according to the annotated labels via reward learning or direct policy optimization.Extensive empirical results on simulated embodied manipulation tasks demonstrate that our method provides accurate preferences and generalizes to unseen tasks and unseen language instructions, outperforming the baselines by a large margin and shifting the burden from continuous, per-task human annotation to one-time, per-domain data collection. Runze Liu 0002, Chenjia Bai, Jiafei Lyu, Shengjie Sun 0002, Yali Du 0001, Xiu Li 0001 |
EMNLP | 1 |
| 2025 | ReviewRL: Towards Automated Scientific Review with RLabstractSihang Zeng, Kai Tian, Kaiyan Zhang, Yuru Wang, Junqi Gao, Runze Liu, Sa Yang, Jingxuan Li, Xinwei Long, Jiaheng Ma, Biqing Qi, Bowen Zhou. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Sihang Zeng, Yuru Wang, Junqi Gao, Runze Liu 0002, Sa Yang, Xinwei Long, Jiaheng Ma, Biqing Qi, Bowen Zhou 0002 |
EMNLP | 6 |
| 2025 | Cross-Domain Offline Policy Adaptation with Optimal Transport and Dataset ConstraintabstractWe explore cross-domain offline reinforcement learning (RL) where offline datasets from another domain can be accessed to facilitate policy learning. However, the underlying environments of the two datasets may have dynamics mismatches, incurring inferior performance when simply merging the data of two domains. Existing methods mitigate this issue by training domain classifiers, using contrastive learning methods, etc. Nevertheless, they still rely on a large amount of target domain data to function well. Instead, we address this problem by establishing a concrete performance bound of a policy given datasets from two domains. Motivated by the theoretical insights, we propose to align transitions in the two datasets using optimal transport and selectively share source domain samples, without training any neural networks. This enables reliable data filtering even given a few target domain data. Additionally, we introduce a dataset regularization term that ensures the learned policy remains within the scope of the target domain dataset, preventing it from being biased towards the source domain data. Consequently, we propose the Optimal Transport Data Filtering (dubbed OTDF) method and examine its effectiveness by conducting extensive experiments across various dynamics shift conditions (e.g., gravity shift), given limited target domain data. It turns out that OTDF exhibits superior performance on many tasks and dataset qualities, often surpassing prior strong baselines by a large margin. Jiafei Lyu, Mengbei Yan, Zhongjian Qiao, Runze Liu 0002, Xiaoteng Ma, Deheng Ye, Zongqing Lu 0002, Xiu Li 0001 |
ICLR | 4 |
| 2025 | Bohdi: Heterogeneous LLM Fusion with Automatic Data ExplorationabstractHeterogeneous Large Language Model (LLM) fusion integrates the strengths of multiple source LLMs with different architectures into a target LLM with low computational overhead. While promising, existing methods suffer from two major limitations: 1) **reliance on real data from limited domain** for knowledge fusion, preventing the target LLM from fully acquiring knowledge across diverse domains, and 2) **fixed data allocation proportions** across domains, failing to dynamically adjust according to the target LLM's varying capabilities across domains, leading to a capability imbalance. To overcome these limitations, we propose Bohdi, a synthetic-data-only heterogeneous LLM fusion framework. Through the organization of knowledge domains into a hierarchical tree structure, Bohdi enables automatic domain exploration and multi-domain data generation through multi-model collaboration, thereby comprehensively extracting knowledge from source LLMs. By formalizing domain expansion and data sampling proportion allocation on the knowledge tree as a Hierarchical Multi-Armed Bandit problem, Bohdi leverages the designed DynaBranches mechanism to adaptively adjust sampling proportions based on the target LLM's performance feedback across domains. Integrated with our proposed Introspection-Rebirth (IR) mechanism, DynaBranches dynamically tracks capability shifts during target LLM's updates via Sliding Window Binomial Likelihood Ratio Testing (SWBLRT), further enhancing its online adaptation capability. Comparative experimental results on a comprehensive suite of benchmarks demonstrate that Bohdi significantly outperforms existing baselines on multiple target LLMs, exhibits higher data efficiency, and virtually eliminates the imbalance in the target LLM's capabilities. Junqi Gao, Zhichang Guo, Dazhi Zhang, Dong Li 0016, Runze Liu 0002, Pengfei Li 0011, Biqing Qi |
NeurIPS | 5 |
| 2025 | A large language model-driven reward design framework via dynamic feedback for reinforcement learning
Shengjie Sun 0002, Runze Liu 0002, Jiafei Lyu, Liangpeng Zhang, Xiu Li 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Enhancing Visual Generalization in Reinforcement Learning with Cycling Augmentation
Shengjie Sun 0002, Jiafei Lyu, Jiazhe Guo, Mengbei Yan, Runze Liu 0002, Xiu Li 0001 |
ICANN (4) | 6 |
| 2024 | SEABO: A Simple Search-Based Method for Offline Imitation LearningabstractOffline reinforcement learning (RL) has attracted much attention due to its ability in learning from static offline datasets and eliminating the need of interacting with the environment. Nevertheless, the success of offline RL relies heavily on the offline transitions annotated with reward labels. In practice, we often need to hand-craft the reward function, which is sometimes difficult, labor-intensive, or inefficient. To tackle this challenge, we set our focus on the offline imitation learning (IL) setting, and aim at getting a reward function based on the expert data and unlabeled data. To that end, we propose a simple yet effective search-based offline IL method, tagged SEABO. SEABO allocates a larger reward to the transition that is close to its closest neighbor in the expert demonstration, and a smaller reward otherwise, all in an unsupervised learning manner. Experimental results on a variety of D4RL datasets indicate that SEABO can achieve competitive performance to offline RL algorithms with ground-truth rewards, given only a single expert trajectory, and can outperform prior reward learning and offline IL methods across many tasks. Moreover, we demonstrate that SEABO also works well if the expert demonstrations contain only observations. Our code is publicly available at https://github.com/dmksjfl/SEABO. Jiafei Lyu, Xiaoteng Ma, Le Wan, Runze Liu 0002, Xiu Li 0001, Zongqing Lu 0002 |
ICLR | 4 |
| 2024 | PEARL: Zero-shot Cross-task Preference Alignment and Robust Reward Learning for Robotic ManipulationabstractIn preference-based Reinforcement Learning (RL), obtaining a large number of preference labels are both time-consuming and costly. Furthermore, the queried human preferences cannot be utilized for the new tasks. In this paper, we propose Zero-shot Cross-task Preference Alignment and Robust Reward Learning (PEARL), which learns policies from cross-task preference transfer without any human labels of the target task. Our contributions include two novel components that facilitate the transfer and learning process. The first is Cross-task Preference Alignment (CPA), which transfers the preferences between tasks via optimal transport. The key idea of CPA is to use Gromov-Wasserstein distance to align the trajectories between tasks, and the solved optimal transport matrix serves as the correspondence between trajectories. The target task preferences are computed as the weighted sum of source task preference labels with the correspondence as weights. Moreover, to ensure robust learning from these transferred labels, we introduce Robust Reward Learning (RRL), which considers both reward mean and uncertainty by modeling rewards as Gaussian distributions. Empirical results on robotic manipulation tasks from Meta-World and Robomimic demonstrate that our method is capable of transferring preference labels across tasks accurately and then learns well-behaved policies. Notably, our approach significantly exceeds existing methods when there are few human preferences. The code and videos of our method are available at: https://sites.google.com/view/pearl-preference. Runze Liu 0002, Yali Du 0001, Fengshuo Bai, Jiafei Lyu, Xiu Li 0001 |
ICML | 1 |
| 2022 | Meta-Reward-Net: Implicitly Differentiable Reward Learning for Preference-based Reinforcement LearningabstractSetting up a well-designed reward function has been challenging for many reinforcement learning applications. Preference-based reinforcement learning (PbRL) provides a new framework that avoids reward engineering by leveraging human preferences (i.e., preferring apples over oranges) as the reward signal. Therefore, improving the efficacy of data usage for preference data becomes critical. In this work, we propose Meta-Reward-Net (MRN), a data-efficient PbRL framework that incorporates bi-level optimization for both reward and policy learning. The key idea of MRN is to adopt the performance of the Q-function as the learning target. Based on this, MRN learns the Q-function and the policy in the inner level while updating the reward function adaptively according to the performance of the Q-function on the preference data in the outer level. Our experiments on robotic simulated manipulation tasks and locomotion tasks demonstrate that MRN outperforms prior methods in the case of few preference labels and significantly improves data efficiency, achieving state-of-the-art in preference-based RL. Ablation studies further demonstrate that MRN learns a more accurate Q-function compared to prior work and shows obvious advantages when only a small amount of human feedback is available. The source code and videos of this project are released at https://sites.google.com/view/meta-reward-net. Runze Liu 0002, Fengshuo Bai, Yali Du 0001, Yaodong Yang 0001 |
NeurIPS | 1 |