VLDB 2026 Research / reviewers in the wild / expert
Fengshuo Bai
dblp:346/1114
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Reinforcement learning · 60% Language models and text generation · 28% Trustworthy machine learning · 6% |
Topics — the 13 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning |
2.2 | 3 | 2025 | STAR: Efficient Preference-based Reinforcement Learning via Dual Regularization · NeurIPS 2025 PEARL: Zero-shot Cross-task Preference Alignment and Robust Reward Learning for Robotic Manipulation · ICML 2024 Meta-Reward-Net: Implicitly Differentiable Reward Learning for Preference-based Reinforcement Learning · NeurIPS 2022 |
Machine learning › Reinforcement learning
reward learning |
2.2 | 3 | 2025 | STAR: Efficient Preference-based Reinforcement Learning via Dual Regularization · NeurIPS 2025 PEARL: Zero-shot Cross-task Preference Alignment and Robust Reward Learning for Robotic Manipulation · ICML 2024 Meta-Reward-Net: Implicitly Differentiable Reward Learning for Preference-based Reinforcement Learning · NeurIPS 2022 |
Natural language and speech › Language models and text generation › agentic language model › tool-augmented language models
function calling |
1.0 | 1 | 2026 | ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling · ACL (1) 2026 |
Natural language and speech › Language models and text generation › text generation
structured generation |
1.0 | 1 | 2026 | ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling · ACL (1) 2026 |
Machine learning › Reinforcement learning › robust reinforcement learning
adversarial reinforcement learning |
0.9 | 1 | 2025 | RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted Behaviors · AAAI 2025 |
Natural language and speech › Language models and text generation › alignment
inference-time alignment |
0.9 | 1 | 2025 | Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs · ICLR 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted Behaviors · AAAI 2025 |
Machine learning › Reinforcement learning › reward learning
robust reward learning |
0.8 | 1 | 2024 | PEARL: Zero-shot Cross-task Preference Alignment and Robust Reward Learning for Robotic Manipulation · ICML 2024 |
Machine learning › Reinforcement learning
multi-task reinforcement learning |
0.7 | 1 | 2023 | PiCor: Multi-Task Deep Reinforcement Learning with Policy Correction · AAAI 2023 |
Machine learning › Reinforcement learning
sample efficiency |
0.7 | 1 | 2023 | PiCor: Multi-Task Deep Reinforcement Learning with Policy Correction · AAAI 2023 |
Machine learning › Optimization for machine learning
bilevel optimization |
0.6 | 1 | 2022 | Meta-Reward-Net: Implicitly Differentiable Reward Learning for Preference-based Reinforcement Learning · NeurIPS 2022 |
Machine learning › Reinforcement learning
value function estimation |
0.3 | 1 | 2025 | STAR: Efficient Preference-based Reinforcement Learning via Dual Regularization · NeurIPS 2025 |
Machine learning › Reinforcement learning
continuous control |
0.2 | 1 | 2023 | PiCor: Multi-Task Deep Reinforcement Learning with Policy Correction · AAAI 2023 |
Methods — techniques the papers use, named apart from their topics
process reward model · 1.0inference scaling · 1.0replay buffer adjustment · 0.9preference margin regularization · 0.9policy regularization · 0.9online learning · 0.9intention policy · 0.9decoding-time optimization · 0.9decision transformer · 0.9closed-form solution · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function CallingabstractJianghao Lin, Yuanyuan Shi, Xin Peng, Renjie Ding, Hairui Wang, Yuxuan Peng, Bizhe Bai, Weixi Song, Fengshuo Bai, Huacan Chai, Weinan Zhang, Fei Huang, Ying Wen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jianghao Lin, Renjie Ding, Yuxuan Peng, Bizhe Bai, Weixi Song, Fengshuo Bai, Huacan Chai, Weinan Zhang 0001, Ying Wen 0001 |
ACL (1) | 9 |
| 2025 | RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted BehaviorsabstractEvaluating deep reinforcement learning (DRL) agents against targeted behavior attacks is critical for assessing their robustness. These attacks aim to manipulate the victim into specific behaviors that align with the attacker’s objectives, often bypassing traditional reward-based defenses. Prior methods have primarily focused on reducing cumulative rewards; however, rewards are typically too generic to capture complex safety requirements effectively. As a result, focusing solely on reward reduction can lead to suboptimal attack strategies, particularly in safety-critical scenarios where more precise behavior manipulation is needed. To address these challenges, we propose RAT, a method designed for universal, targeted behavior attacks. RAT trains an intention policy that is explicitly aligned with human preferences, serving as a precise behavioral target for the adversary. Concurrently, an adversary manipulates the victim's policy to follow this target behavior. To enhance the effectiveness of these attacks, RAT dynamically adjusts the state occupancy measure within the replay buffer, allowing for more controlled and effective behavior manipulation. Our empirical results on robotic simulation tasks demonstrate that RAT outperforms existing adversarial attack algorithms in inducing specific behaviors. Additionally, RAT shows promise in improving agent robustness, leading to more resilient policies. We further validate RAT by guiding Decision Transformer agents to adopt behaviors aligned with human preferences in various MuJoCo tasks, demonstrating its effectiveness across diverse tasks. Fengshuo Bai, Runze Liu 0002, Yali Du 0001, Ying Wen 0001, Yaodong Yang 0001 |
AAAI | 1 |
| 2025 | Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMsabstractHow to align large language models (LLMs) with user preferences from a static general dataset has been frequently studied. However, user preferences are usually personalized, changing, and diverse. This leads to the problem that the actual user preferences often do not coincide with those trained by the model developers in the practical use of LLMs. Since we cannot collect enough data and retrain for every demand, researching efficient real-time preference adaptation methods based on the backbone LLMs during test time is important. To this end, we introduce **Amulet**, a novel, training-free framework that formulates the decoding process of every token as a separate online learning problem with the guidance of simple user-provided prompts, thus enabling real-time optimization to satisfy users' personalized preferences. To reduce the computational cost brought by this optimization process for each token, we additionally provide a closed-form solution for each iteration step of the optimization process, thereby reducing the computational time cost to a negligible level. The detailed experimental results demonstrate that Amulet can achieve significant performance improvements in rich settings with combinations of different LLMs, datasets, and user preferences, while maintaining acceptable computational efficiency. Zhaowei Zhang 0001, Fengshuo Bai, Chengdong Ma, Zilong Zheng, Yaodong Yang 0001 |
ICLR | 2 |
| 2025 | β-DQN: Improving Deep Q-Learning By Evolving the Behavior
Hongming Zhang 0003, Fengshuo Bai, Chenjun Xiao, Chao Gao 0012, Bo Xu 0002, Martin Müller 0003 |
AAMAS | 2 |
| 2025 | STAR: Efficient Preference-based Reinforcement Learning via Dual RegularizationabstractPreference-based reinforcement learning (PbRL) bypasses complex reward engineering by learning from human feedback. However, due to the high cost of obtaining feedback, PbRL typically relies on a limited set of preference-labeled samples. This data scarcity introduces two key inefficiencies: (1) the reward model overfits to the limited feedback, leading to poor generalization to unseen samples, and (2) the agent exploits the learned reward model, exacerbating overestimation of action values in temporal difference (TD) learning. To address these issues, we propose STAR, an efficient PbRL method that integrates preference margin regularization and policy regularization. Preference margin regularization mitigates overfitting by introducing a bounded margin in reward optimization, preventing excessive bias toward specific feedback. Policy regularization bootstraps a conservative estimate $\widehat{Q}$ from well-supported state-action pairs in the replay memory, reducing overestimation during policy learning. Experimental results show that STAR improves feedback efficiency, achieving 34.8\% higher performance in online settings and 29.7\% in offline settings compared to state-of-the-art methods. Ablation studies confirm that STAR facilitates more robust reward and value function learning. The videos of this project are released at https://sites.google.com/view/pbrl-star. Fengshuo Bai, Rui Zhao 0001, Hongming Zhang 0003, Sijia Cui, Shao Zhang, Bo Xu 0002, Ying Wen 0001, Yaodong Yang 0001 |
NeurIPS | 1 |
| 2025 | DexFlyWheel: A Scalable and Self-improving Data Generation Framework for Dexterous ManipulationabstractDexterous manipulation is critical for advancing robot capabilities in real-world applications, yet diverse and high-quality datasets remain scarce. Existing data collection methods either rely on human teleoperation or require significant human engineering, or generate data with limited diversity, which restricts their scalability and generalization. In this paper, we introduce DexFlyWheel, a scalable data generation framework that employs a self-improving cycle to continuously enrich data diversity. Starting from efficient seed demonstrations warmup, DexFlyWheel expands the dataset through iterative cycles. Each cycle follows a closed-loop pipeline that integrates Imitation Learning (IL), residual Reinforcement Learning (RL), rollout trajectory collection, and data augmentation. Specifically, IL extracts human-like behaviors from demonstrations, and residual RL enhances policy generalization. The learned policy is then used to generate trajectories in simulation, which are further augmented across diverse environments and spatial configurations before being fed back into the next cycle. Over successive iterations, a self-improving data flywheel effect emerges, producing datasets that cover diverse scenarios and thereby scaling policy performance. Experimental results demonstrate that DexFlyWheel generates over 2,000 diverse demonstrations across four challenging tasks. Policies trained on our dataset achieve an average success rate of 81.9\% on the challenge test sets and successfully transfer to the real world through digital twin, achieving a 78.3\% success rate on dual-arm lift tasks. Kefei Zhu, Fengshuo Bai, YuanHao Xiang, Yishuai Cai, Xinglin Chen, Ruochong Li, Hao Dong 0003, Yaodong Yang 0001, Xiaopeng Fan 0001, Yuanpei Chen |
NeurIPS | 2 |
| 2024 | PEARL: Zero-shot Cross-task Preference Alignment and Robust Reward Learning for Robotic ManipulationabstractIn preference-based Reinforcement Learning (RL), obtaining a large number of preference labels are both time-consuming and costly. Furthermore, the queried human preferences cannot be utilized for the new tasks. In this paper, we propose Zero-shot Cross-task Preference Alignment and Robust Reward Learning (PEARL), which learns policies from cross-task preference transfer without any human labels of the target task. Our contributions include two novel components that facilitate the transfer and learning process. The first is Cross-task Preference Alignment (CPA), which transfers the preferences between tasks via optimal transport. The key idea of CPA is to use Gromov-Wasserstein distance to align the trajectories between tasks, and the solved optimal transport matrix serves as the correspondence between trajectories. The target task preferences are computed as the weighted sum of source task preference labels with the correspondence as weights. Moreover, to ensure robust learning from these transferred labels, we introduce Robust Reward Learning (RRL), which considers both reward mean and uncertainty by modeling rewards as Gaussian distributions. Empirical results on robotic manipulation tasks from Meta-World and Robomimic demonstrate that our method is capable of transferring preference labels across tasks accurately and then learns well-behaved policies. Notably, our approach significantly exceeds existing methods when there are few human preferences. The code and videos of our method are available at: https://sites.google.com/view/pearl-preference. Runze Liu 0002, Yali Du 0001, Fengshuo Bai, Jiafei Lyu, Xiu Li 0001 |
ICML | 3 |
| 2023 | PiCor: Multi-Task Deep Reinforcement Learning with Policy CorrectionabstractMulti-task deep reinforcement learning (DRL) ambitiously aims to train a general agent that masters multiple tasks simultaneously. However, varying learning speeds of different tasks compounding with negative gradients interference makes policy learning inefficient. In this work, we propose PiCor, an efficient multi-task DRL framework that splits learning into policy optimization and policy correction phases. The policy optimization phase improves the policy by any DRL algothrim on the sampled single task without considering other tasks. The policy correction phase first constructs an adaptive adjusted performance constraint set. Then the intermediate policy learned by the first phase is constrained to the set, which controls the negative interference and balances the learning speeds across tasks. Empirically, we demonstrate that PiCor outperforms previous methods and significantly improves sample efficiency on simulated robotic manipulation and continuous control tasks. We additionally show that adaptive weight adjusting can further improve data efficiency and performance. Fengshuo Bai, Hongming Zhang 0003, Tianyang Tao, Yanna Wang, Bo Xu 0002 |
AAAI | 1 |
| 2022 | Meta-Reward-Net: Implicitly Differentiable Reward Learning for Preference-based Reinforcement LearningabstractSetting up a well-designed reward function has been challenging for many reinforcement learning applications. Preference-based reinforcement learning (PbRL) provides a new framework that avoids reward engineering by leveraging human preferences (i.e., preferring apples over oranges) as the reward signal. Therefore, improving the efficacy of data usage for preference data becomes critical. In this work, we propose Meta-Reward-Net (MRN), a data-efficient PbRL framework that incorporates bi-level optimization for both reward and policy learning. The key idea of MRN is to adopt the performance of the Q-function as the learning target. Based on this, MRN learns the Q-function and the policy in the inner level while updating the reward function adaptively according to the performance of the Q-function on the preference data in the outer level. Our experiments on robotic simulated manipulation tasks and locomotion tasks demonstrate that MRN outperforms prior methods in the case of few preference labels and significantly improves data efficiency, achieving state-of-the-art in preference-based RL. Ablation studies further demonstrate that MRN learns a more accurate Q-function compared to prior work and shows obvious advantages when only a small amount of human feedback is available. The source code and videos of this project are released at https://sites.google.com/view/meta-reward-net. Runze Liu 0002, Fengshuo Bai, Yali Du 0001, Yaodong Yang 0001 |
NeurIPS | 2 |