Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xuehai Pan

dblp:333/0877 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 38% Language models and text generation · 27% Transfer learning and domain adaptation · 7%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
1.522024
Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024
Safe RLHF: Safe Reinforcement Learning from Human Feedback · ICLR 2024
Machine learning › Reinforcement learning
reinforcement learning from human feedback
1.422024
Safe RLHF: Safe Reinforcement Learning from Human Feedback · ICLR 2024
BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset · NeurIPS 2023
Machine learning › Reinforcement learning
safe reinforcement learning
1.422024
OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research · J. Mach. Learn. Res. 2024
Safety Gymnasium: A Unified Safe Reinforcement Learning Benchmark · NeurIPS 2023
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.822023
MATE: Benchmarking Multi-Agent Reinforcement Learning in Distributed Target Coverage Control · NeurIPS 2022
Safety Gymnasium: A Unified Safe Reinforcement Learning Benchmark · NeurIPS 2023
Natural language and speech › Language models and text generation
hallucination mitigation
0.812024
Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model
0.812024
Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024
Computer vision › 3D vision
3d human pose estimation
0.712023
Proactive Multi-Camera Collaboration for 3D Human Pose Estimation · ICLR 2023
Machine learning › Reinforcement learning › reinforcement learning environment
benchmark environments
0.712023
Safety Gymnasium: A Unified Safe Reinforcement Learning Benchmark · NeurIPS 2023
Machine learning › Optimization for machine learning
differentiable optimization
0.712023
TorchOpt: An Efficient Library for Differentiable Optimization · J. Mach. Learn. Res. 2023
Machine learning › Efficient and distributed learning
distributed training
0.712023
TorchOpt: An Efficient Library for Differentiable Optimization · J. Mach. Learn. Res. 2023
Machine learning › Transfer learning and domain adaptation › meta-learning
meta-gradient
0.712023
TorchOpt: An Efficient Library for Differentiable Optimization · J. Mach. Learn. Res. 2023
Machine learning › Trustworthy machine learning › AI safety
safety alignment
0.712023
BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation › foundation model adaptation
model-agnostic adaptation
0.212024
Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model safety
0.212023
BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset · NeurIPS 2023
Machine learning › Reinforcement learning › multi-agent reinforcement learning
multi-agent communication
0.212022
MATE: Benchmarking Multi-Agent Reinforcement Learning in Distributed Target Coverage Control · NeurIPS 2022
Machine learning › Reinforcement learning › multi-agent reinforcement learning
self-play
0.212022
MATE: Benchmarking Multi-Agent Reinforcement Learning in Distributed Target Coverage Control · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.4synthetic preference data · 0.8reward modeling · 0.8lagrangian method · 0.8cost modeling · 0.8correctional residual learning · 0.8constrained optimization · 0.8bootstrapping · 0.8human preference annotation · 0.7content moderation · 0.7
YearPublicationVenuePosition
2024 Safe RLHF: Safe Reinforcement Learning from Human Feedback
abstract
With the development of large language models (LLMs), striking a balance between the performance and safety of AI systems has never been more critical. However, the inherent tension between the objectives of helpfulness and harmlessness presents a significant challenge during LLM training. To address this issue, we propose Safe Reinforcement Learning from Human Feedback (Safe RLHF), a novel algorithm for human value alignment. Safe RLHF explicitly decouples human preferences regarding helpfulness and harmlessness, effectively avoiding the crowd workers' confusion about the tension and allowing us to train separate reward and cost models. We formalize the safety concern of LLMs as an optimization task of maximizing the reward function while satisfying specified cost constraints. Leveraging the Lagrangian method to solve this constrained problem, Safe RLHF dynamically adjusts the balance between the two objectives during fine-tuning. Through a three-round fine-tuning using Safe RLHF, we demonstrate a superior ability to mitigate harmful responses while enhancing model performance compared to existing value-aligned algorithms. Experimentally, we fine-tuned the Alpaca-7B using Safe RLHF and aligned it with collected human preferences, significantly improving its helpfulness and harmlessness according to human evaluations. Code is available at https://github.com/PKU-Alignment/safe-rlhf. Warning: This paper contains example data that may be offensive or harmful.
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang 0001, Yaodong Yang 0001
ICLR2
2024 Aligner: Efficient Alignment by Learning to Correct
abstract
With the rapid development of large language models (LLMs) and ever-evolving practical requirements, finding an efficient and effective alignment method has never been more critical. However, the tension between the complexity of current alignment methods and the need for rapid iteration in deployment scenarios necessitates the development of a model-agnostic alignment approach that can operate under these constraints. In this paper, we introduce Aligner, a novel and simple alignment paradigm that learns the correctional residuals between preferred and dispreferred answers using a small model. Designed as a model-agnostic, plug-and-play module, Aligner can be directly applied to various open-source and API-based models with only one-off training, making it suitable for rapid iteration. Notably, Aligner can be applied to any powerful, large-scale upstream models. Moreover, it can even iteratively bootstrap the upstream models using corrected responses as synthetic human preference data, breaking through the model's performance ceiling. Our experiments demonstrate performance improvements by deploying the same Aligner model across 11 different LLMs, evaluated on the 3H dimensions (helpfulness, harmlessness, and honesty). Specifically, Aligner-7B has achieved an average improvement of 68.9% in helpfulness and 22.8% in harmlessness across the tested LLMs while also effectively reducing hallucination. In the Alpaca-Eval leaderboard, stacking Aligner-2B on GPT-4 Turbo improved its LC Win Rate from 55.0% to 58.3%, surpassing GPT-4 Omni's 57.5% Win Rate (community report).
Jiaming Ji, Boyuan Chen 0008, Hantao Lou, Donghai Hong, Borong Zhang, Xuehai Pan, Tianyi Qiu, Juntao Dai, Yaodong Yang 0001
NeurIPS6
2024 OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research
abstract
AI systems empowered by reinforcement learning (RL) algorithms harbor the immense potential to catalyze societal advancement, yet their deployment is often impeded by significant safety concerns. Particularly in safety-critical applications, researchers have raised concerns about unintended harms or unsafe behaviors of unaligned RL agents. The philosophy of safe reinforcement learning (SafeRL) is to align RL agents with harmless intentions and safe behavioral patterns. In SafeRL, agents learn to develop optimal policies by receiving feedback from the environment, while also fulfilling the requirement of minimizing the risk of unintended harm or unsafe behavior. However, due to the intricate nature of SafeRL algorithm implementation, combining methodologies across various domains presents a formidable challenge. This had led to an absence of a cohesive and efficacious learning framework within the contemporary SafeRL research milieu. In this work, we introduce a foundational framework designed to expedite SafeRL research endeavors. Our comprehensive framework encompasses an array of algorithms spanning different RL domains and places heavy emphasis on safety elements. Our efforts are to make the SafeRL-related research process more streamlined and efficient, therefore facilitating further research in AI safety.
Jiaming Ji, Borong Zhang, Juntao Dai, Xuehai Pan, Ruiyang Sun, Weidong Huang 0008, Yiran Geng, Mickel Liu, Yaodong Yang 0001
J. Mach. Learn. Res.5
2023 Proactive Multi-Camera Collaboration for 3D Human Pose Estimation
Hai Ci, Mickel Liu, Xuehai Pan, Fangwei Zhong, Yizhou Wang 0001
ICLR3
2023 BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset
abstract
In this paper, we introduce the BeaverTails dataset, aimed at fostering research on safety alignment in large language models (LLMs). This dataset uniquely separates annotations of helpfulness and harmlessness for question-answering pairs, thus offering distinct perspectives on these crucial attributes. In total, we have gathered safety meta-labels for 333,963 question-answer (QA) pairs and 361,903 pairs of expert comparison data for both the helpfulness and harmlessness metrics. We further showcase applications of BeaverTails in content moderation and reinforcement learning with human feedback (RLHF), emphasizing its potential for practical safety measures in LLMs. We believe this dataset provides vital resources for the community, contributing towards the safe development and deployment of LLMs. Our project page is available at the following URL: https://sites.google.com/view/pku-beavertails.
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang 0017, Ce Bian, Boyuan Chen 0008, Ruiyang Sun, Yizhou Wang 0001, Yaodong Yang 0001
NeurIPS4
2023 Safety Gymnasium: A Unified Safe Reinforcement Learning Benchmark
abstract
Artificial intelligence (AI) systems possess significant potential to drive societal progress. However, their deployment often faces obstacles due to substantial safety concerns. Safe reinforcement learning (SafeRL) emerges as a solution to optimize policies while simultaneously adhering to multiple constraints, thereby addressing the challenge of integrating reinforcement learning in safety-critical scenarios. In this paper, we present an environment suite called Safety-Gymnasium, which encompasses safety-critical tasks in both single and multi-agent scenarios, accepting vector and vision-only input. Additionally, we offer a library of algorithms named Safe Policy Optimization (SafePO), comprising 16 state-of-the-art SafeRL algorithms. This comprehensive library can serve as a validation tool for the research community. By introducing this benchmark, we aim to facilitate the evaluation and comparison of safety performance, thus fostering the development of reinforcement learning for safer, more reliable, and responsible real-world applications. The website of this project can be accessed at https://sites.google.com/view/safety-gymnasium.
Jiaming Ji, Borong Zhang, Xuehai Pan, Weidong Huang 0008, Ruiyang Sun, Yiran Geng, Yifan Zhong, Josef Dai, Yaodong Yang 0001
NeurIPS4
2023 TorchOpt: An Efficient Library for Differentiable Optimization
abstract
Differentiable optimization algorithms often involve expensive computations of various meta-gradients. To address this, we design and implement TorchOpt, a new PyTorch-based differentiable optimization library. TorchOpt provides an expressive and unified programming interface that simplifies the implementation of explicit, implicit, and zero-order gradients. Moreover, TorchOpt has a distributed execution runtime capable of parallelizing diverse operations linked to differentiable optimization tasks across CPU and GPU devices. Experimental results demonstrate that TorchOpt achieves a 5.2× training time speedup in a cluster. TorchOpt is open-sourced at https://github.com/metaopt/torchopt and has become a PyTorch Ecosystem project.
Xidong Feng, Bo Liu 0039, Xuehai Pan, Yao Fu 0013, Luo Mai, Yaodong Yang 0001
J. Mach. Learn. Res.4
2022 MATE: Benchmarking Multi-Agent Reinforcement Learning in Distributed Target Coverage Control
abstract
We introduce the Multi-Agent Tracking Environment (MATE), a novel multi-agent environment simulates the target coverage control problems in the real world. MATE hosts an asymmetric cooperative-competitive game consisting of two groups of learning agents--"cameras" and "targets"--with opposing interests. Specifically, "cameras", a group of directional sensors, are mandated to actively control the directional perception area to maximize the coverage rate of targets. On the other side, "targets" are mobile agents that aim to transport cargo between multiple randomly assigned warehouses while minimizing the exposure to the camera sensor networks. To showcase the practicality of MATE, we benchmark the multi-agent reinforcement learning (MARL) algorithms from different aspects, including cooperation, communication, scalability, robustness, and asymmetric self-play. We start by reporting results for cooperative tasks using MARL algorithms (MAPPO, IPPO, QMIX, MADDPG) and the results after augmenting with multi-agent communication protocols (TarMAC, I2C). We then evaluate the effectiveness of the popular self-play techniques (PSRO, fictitious self-play) in an asymmetric zero-sum competitive game. This process of co-evolution between cameras and targets helps to realize a less exploitable camera network. We also observe the emergence of different roles of the target agents while incorporating I2C into target-target communication. MATE is written purely in Python and integrated with OpenAI Gym API to enhance user-friendliness. Our project is released at https://github.com/UnrealTracking/mate.
Xuehai Pan, Mickel Liu, Fangwei Zhong, Yaodong Yang 0001, Song-Chun Zhu, Yizhou Wang 0001
NeurIPS1