VLDB 2026 Research / reviewers in the wild / expert
Xuehai Pan
dblp:333/0877
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Reinforcement learning · 38% Language models and text generation · 27% Transfer learning and domain adaptation · 7% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
1.5 | 2 | 2024 | Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024 Safe RLHF: Safe Reinforcement Learning from Human Feedback · ICLR 2024 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
1.4 | 2 | 2024 | Safe RLHF: Safe Reinforcement Learning from Human Feedback · ICLR 2024 BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset · NeurIPS 2023 |
Machine learning › Reinforcement learning
safe reinforcement learning |
1.4 | 2 | 2024 | OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research · J. Mach. Learn. Res. 2024 Safety Gymnasium: A Unified Safe Reinforcement Learning Benchmark · NeurIPS 2023 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.8 | 2 | 2023 | MATE: Benchmarking Multi-Agent Reinforcement Learning in Distributed Target Coverage Control · NeurIPS 2022 Safety Gymnasium: A Unified Safe Reinforcement Learning Benchmark · NeurIPS 2023 |
Natural language and speech › Language models and text generation
hallucination mitigation |
0.8 | 1 | 2024 | Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024 |
Natural language and speech › Language models and text generation
large language model |
0.8 | 1 | 2024 | Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024 |
Computer vision › 3D vision
3d human pose estimation |
0.7 | 1 | 2023 | Proactive Multi-Camera Collaboration for 3D Human Pose Estimation · ICLR 2023 |
Machine learning › Reinforcement learning › reinforcement learning environment
benchmark environments |
0.7 | 1 | 2023 | Safety Gymnasium: A Unified Safe Reinforcement Learning Benchmark · NeurIPS 2023 |
Machine learning › Optimization for machine learning
differentiable optimization |
0.7 | 1 | 2023 | TorchOpt: An Efficient Library for Differentiable Optimization · J. Mach. Learn. Res. 2023 |
Machine learning › Efficient and distributed learning
distributed training |
0.7 | 1 | 2023 | TorchOpt: An Efficient Library for Differentiable Optimization · J. Mach. Learn. Res. 2023 |
Machine learning › Transfer learning and domain adaptation › meta-learning
meta-gradient |
0.7 | 1 | 2023 | TorchOpt: An Efficient Library for Differentiable Optimization · J. Mach. Learn. Res. 2023 |
Machine learning › Trustworthy machine learning › AI safety
safety alignment |
0.7 | 1 | 2023 | BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset · NeurIPS 2023 |
Machine learning › Transfer learning and domain adaptation › foundation model adaptation
model-agnostic adaptation |
0.2 | 1 | 2024 | Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024 |
Natural language and speech › Language models and text generation
large language model safety |
0.2 | 1 | 2023 | BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset · NeurIPS 2023 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
multi-agent communication |
0.2 | 1 | 2022 | MATE: Benchmarking Multi-Agent Reinforcement Learning in Distributed Target Coverage Control · NeurIPS 2022 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
self-play |
0.2 | 1 | 2022 | MATE: Benchmarking Multi-Agent Reinforcement Learning in Distributed Target Coverage Control · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.4synthetic preference data · 0.8reward modeling · 0.8lagrangian method · 0.8cost modeling · 0.8correctional residual learning · 0.8constrained optimization · 0.8bootstrapping · 0.8human preference annotation · 0.7content moderation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Safe RLHF: Safe Reinforcement Learning from Human FeedbackabstractWith the development of large language models (LLMs), striking a balance between the performance and safety of AI systems has never been more critical. However, the inherent tension between the objectives of helpfulness and harmlessness presents a significant challenge during LLM training. To address this issue, we propose Safe Reinforcement Learning from Human Feedback (Safe RLHF), a novel algorithm for human value alignment. Safe RLHF explicitly decouples human preferences regarding helpfulness and harmlessness, effectively avoiding the crowd workers' confusion about the tension and allowing us to train separate reward and cost models. We formalize the safety concern of LLMs as an optimization task of maximizing the reward function while satisfying specified cost constraints. Leveraging the Lagrangian method to solve this constrained problem, Safe RLHF dynamically adjusts the balance between the two objectives during fine-tuning. Through a three-round fine-tuning using Safe RLHF, we demonstrate a superior ability to mitigate harmful responses while enhancing model performance compared to existing value-aligned algorithms. Experimentally, we fine-tuned the Alpaca-7B using Safe RLHF and aligned it with collected human preferences, significantly improving its helpfulness and harmlessness according to human evaluations.
Code is available at https://github.com/PKU-Alignment/safe-rlhf.
Warning: This paper contains example data that may be offensive or harmful. Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang 0001, Yaodong Yang 0001 |
ICLR | 2 |
| 2024 | Aligner: Efficient Alignment by Learning to CorrectabstractWith the rapid development of large language models (LLMs) and ever-evolving practical requirements, finding an efficient and effective alignment method has never been more critical. However, the tension between the complexity of current alignment methods and the need for rapid iteration in deployment scenarios necessitates the development of a model-agnostic alignment approach that can operate under these constraints. In this paper, we introduce Aligner, a novel and simple alignment paradigm that learns the correctional residuals between preferred and dispreferred answers using a small model. Designed as a model-agnostic, plug-and-play module, Aligner can be directly applied to various open-source and API-based models with only one-off training, making it suitable for rapid iteration. Notably, Aligner can be applied to any powerful, large-scale upstream models. Moreover, it can even iteratively bootstrap the upstream models using corrected responses as synthetic human preference data, breaking through the model's performance ceiling. Our experiments demonstrate performance improvements by deploying the same Aligner model across 11 different LLMs, evaluated on the 3H dimensions (helpfulness, harmlessness, and honesty). Specifically, Aligner-7B has achieved an average improvement of 68.9% in helpfulness and 22.8% in harmlessness across the tested LLMs while also effectively reducing hallucination. In the Alpaca-Eval leaderboard, stacking Aligner-2B on GPT-4 Turbo improved its LC Win Rate from 55.0% to 58.3%, surpassing GPT-4 Omni's 57.5% Win Rate (community report). Jiaming Ji, Boyuan Chen 0008, Hantao Lou, Donghai Hong, Borong Zhang, Xuehai Pan, Tianyi Qiu, Juntao Dai, Yaodong Yang 0001 |
NeurIPS | 6 |
| 2024 | OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning ResearchabstractAI systems empowered by reinforcement learning (RL) algorithms harbor the immense potential to catalyze societal advancement, yet their deployment is often impeded by significant safety concerns. Particularly in safety-critical applications, researchers have raised concerns about unintended harms or unsafe behaviors of unaligned RL agents. The philosophy of safe reinforcement learning (SafeRL) is to align RL agents with harmless intentions and safe behavioral patterns. In SafeRL, agents learn to develop optimal policies by receiving feedback from the environment, while also fulfilling the requirement of minimizing the risk of unintended harm or unsafe behavior. However, due to the intricate nature of SafeRL algorithm implementation, combining methodologies across various domains presents a formidable challenge. This had led to an absence of a cohesive and efficacious learning framework within the contemporary SafeRL research milieu. In this work, we introduce a foundational framework designed to expedite SafeRL research endeavors. Our comprehensive framework encompasses an array of algorithms spanning different RL domains and places heavy emphasis on safety elements. Our efforts are to make the SafeRL-related research process more streamlined and efficient, therefore facilitating further research in AI safety. Jiaming Ji, Borong Zhang, Juntao Dai, Xuehai Pan, Ruiyang Sun, Weidong Huang 0008, Yiran Geng, Mickel Liu, Yaodong Yang 0001 |
J. Mach. Learn. Res. | 5 |
| 2023 | Proactive Multi-Camera Collaboration for 3D Human Pose Estimation
Hai Ci, Mickel Liu, Xuehai Pan, Fangwei Zhong, Yizhou Wang 0001 |
ICLR | 3 |
| 2023 | BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference DatasetabstractIn this paper, we introduce the BeaverTails dataset, aimed at fostering research on safety alignment in large language models (LLMs). This dataset uniquely separates annotations of helpfulness and harmlessness for question-answering pairs, thus offering distinct perspectives on these crucial attributes. In total, we have gathered safety meta-labels for 333,963 question-answer (QA) pairs and 361,903 pairs of expert comparison data for both the helpfulness and harmlessness metrics. We further showcase applications of BeaverTails in content moderation and reinforcement learning with human feedback (RLHF), emphasizing its potential for practical safety measures in LLMs. We believe this dataset provides vital resources for the community, contributing towards the safe development and deployment of LLMs. Our project page is available at the following URL: https://sites.google.com/view/pku-beavertails. Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang 0017, Ce Bian, Boyuan Chen 0008, Ruiyang Sun, Yizhou Wang 0001, Yaodong Yang 0001 |
NeurIPS | 4 |
| 2023 | Safety Gymnasium: A Unified Safe Reinforcement Learning BenchmarkabstractArtificial intelligence (AI) systems possess significant potential to drive societal progress. However, their deployment often faces obstacles due to substantial safety concerns. Safe reinforcement learning (SafeRL) emerges as a solution to optimize policies while simultaneously adhering to multiple constraints, thereby addressing the challenge of integrating reinforcement learning in safety-critical scenarios. In this paper, we present an environment suite called Safety-Gymnasium, which encompasses safety-critical tasks in both single and multi-agent scenarios, accepting vector and vision-only input. Additionally, we offer a library of algorithms named Safe Policy Optimization (SafePO), comprising 16 state-of-the-art SafeRL algorithms. This comprehensive library can serve as a validation tool for the research community. By introducing this benchmark, we aim to facilitate the evaluation and comparison of safety performance, thus fostering the development of reinforcement learning for safer, more reliable, and responsible real-world applications. The website of this project can be accessed at https://sites.google.com/view/safety-gymnasium. Jiaming Ji, Borong Zhang, Xuehai Pan, Weidong Huang 0008, Ruiyang Sun, Yiran Geng, Yifan Zhong, Josef Dai, Yaodong Yang 0001 |
NeurIPS | 4 |
| 2023 | TorchOpt: An Efficient Library for Differentiable OptimizationabstractDifferentiable optimization algorithms often involve expensive computations of various meta-gradients. To address this, we design and implement TorchOpt, a new PyTorch-based differentiable optimization library. TorchOpt provides an expressive and unified programming interface that simplifies the implementation of explicit, implicit, and zero-order gradients. Moreover, TorchOpt has a distributed execution runtime capable of parallelizing diverse operations linked to differentiable optimization tasks across CPU and GPU devices. Experimental results demonstrate that TorchOpt achieves a 5.2× training time speedup in a cluster. TorchOpt is open-sourced at https://github.com/metaopt/torchopt and has become a PyTorch Ecosystem project. Xidong Feng, Bo Liu 0039, Xuehai Pan, Yao Fu 0013, Luo Mai, Yaodong Yang 0001 |
J. Mach. Learn. Res. | 4 |
| 2022 | MATE: Benchmarking Multi-Agent Reinforcement Learning in Distributed Target Coverage ControlabstractWe introduce the Multi-Agent Tracking Environment (MATE), a novel multi-agent environment simulates the target coverage control problems in the real world. MATE hosts an asymmetric cooperative-competitive game consisting of two groups of learning agents--"cameras" and "targets"--with opposing interests. Specifically, "cameras", a group of directional sensors, are mandated to actively control the directional perception area to maximize the coverage rate of targets. On the other side, "targets" are mobile agents that aim to transport cargo between multiple randomly assigned warehouses while minimizing the exposure to the camera sensor networks. To showcase the practicality of MATE, we benchmark the multi-agent reinforcement learning (MARL) algorithms from different aspects, including cooperation, communication, scalability, robustness, and asymmetric self-play. We start by reporting results for cooperative tasks using MARL algorithms (MAPPO, IPPO, QMIX, MADDPG) and the results after augmenting with multi-agent communication protocols (TarMAC, I2C). We then evaluate the effectiveness of the popular self-play techniques (PSRO, fictitious self-play) in an asymmetric zero-sum competitive game. This process of co-evolution between cameras and targets helps to realize a less exploitable camera network. We also observe the emergence of different roles of the target agents while incorporating I2C into target-target communication. MATE is written purely in Python and integrated with OpenAI Gym API to enhance user-friendliness. Our project is released at https://github.com/UnrealTracking/mate. Xuehai Pan, Mickel Liu, Fangwei Zhong, Yaodong Yang 0001, Song-Chun Zhu, Yizhou Wang 0001 |
NeurIPS | 1 |