Josef Dai

dblp:359/3349 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 46% Language models and text generation · 39% Trustworthy machine learning · 15%
Theoretical computer science
1 paper
Coding theory · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
3.442025
Language Models Resist Alignment: Evidence From Data Compression · ACL (1) 2025
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference · ACL (1) 2025
Sequence to Sequence Reward Modeling: Improving RLHF by Language Feedback · AAAI 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
2.332025
Sequence to Sequence Reward Modeling: Improving RLHF by Language Feedback · AAAI 2025
Safe RLHF: Safe Reinforcement Learning from Human Feedback · ICLR 2024
BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset · NeurIPS 2023
Machine learning › Trustworthy machine learning › AI safety
safety alignment
1.522025
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference · ACL (1) 2025
BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset · NeurIPS 2023
Natural language and speech › Language models and text generation › alignment
preference alignment
0.912025
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference · ACL (1) 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
Sequence to Sequence Reward Modeling: Improving RLHF by Language Feedback · AAAI 2025
Machine learning › Reinforcement learning › reinforcement learning from human feedback › learning from human feedback
RLHF
0.912025
Sequence to Sequence Reward Modeling: Improving RLHF by Language Feedback · AAAI 2025
Machine learning › Reinforcement learning › reinforcement learning environment
benchmark environments
0.712023
Safety Gymnasium: A Unified Safe Reinforcement Learning Benchmark · NeurIPS 2023
Machine learning › Reinforcement learning
safe reinforcement learning
0.712023
Safety Gymnasium: A Unified Safe Reinforcement Learning Benchmark · NeurIPS 2023
Machine learning › Trustworthy machine learning
robustness
0.312025
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference · ACL (1) 2025
Natural language and speech › Language models and text generation
text summarization
0.312025
Sequence to Sequence Reward Modeling: Improving RLHF by Language Feedback · AAAI 2025
Coding theory
source coding
0.312025
Language Models Resist Alignment: Evidence From Data Compression · ACL (1) 2025
Natural language and speech › Language models and text generation
large language model safety
0.212023
BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset · NeurIPS 2023
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.212023
Safety Gymnasium: A Unified Safe Reinforcement Learning Benchmark · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

data compression analysis · 1.7sequence-to-sequence reward model · 0.9reinforcement learning from human feedback · 0.9maximum likelihood estimation · 0.9reward modeling · 0.8lagrangian method · 0.8cost modeling · 0.8constrained optimization · 0.8human preference annotation · 0.7content moderation · 0.7
YearPublicationVenuePosition
2025 Sequence to Sequence Reward Modeling: Improving RLHF by Language Feedback
abstract
Aligning the behavior of Large language models (LLMs) with human intentions and values remains a critical challenge. Reinforcement learning from human feedback (RLHF) aligns LLMs by training a reward model (RM) on human preferences and fine-tuning the LLMs to maximize RM feedback. Despite its effectiveness and popularity, RLHF is prone to biased local optimization. It means RM fails to provide feedback that accurately aligns with human preference, causing LLMs to explore unexpected generalizations, and failing to achieve alignment objectives. To mitigate this issue, we propose a novel sequence-to-sequence (seq2seq) reward modeling method. Its key insight is that learning from language feedback rather than scalar feedback improves RLHF without additional annotations. We replaced the reward modeling target from binary maximum likelihood estimation (MLE) with sequence MLE. This method enables richer and fine-grained language feedback without additional annotations, models, or training stages. Our experiments demonstrated its effectiveness, specifically, reducing the refusal-to-response paradigm in single-turn safety dialogues and the long-response bias in text summarization tasks. We provide further analysis that seq2seq RM improves RLHF performance across 2B and 7B LLMs on 3 NLP tasks, achieving an average win rate of 76.9%. We further show that seq2seq RM can still improve the performance of RLHF under out-of-distribution prompts.
Jiaming Ji, Josef Dai, Yaodong Yang 0001
AAAI3
2025 PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
abstract
Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen, Josef Dai, Boren Zheng, Tianyi Alex Qiu, Jiayi Zhou, Kaile Wang, Boxun Li, Sirui Han, Yike Guo, Yaodong Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen 0008, Josef Dai, Boren Zheng, Tianyi Qiu, Kaile Wang, Boxun Li, Sirui Han, Yike Guo, Yaodong Yang 0001
ACL (1)5
2025 Language Models Resist Alignment: Evidence From Data Compression
abstract
Jiaming Ji, Kaile Wang, Tianyi Alex Qiu, Boyuan Chen, Jiayi Zhou, Changye Li, Hantao Lou, Josef Dai, Yunhuai Liu, Yaodong Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jiaming Ji, Kaile Wang, Tianyi Qiu, Boyuan Chen 0008, Changye Li 0003, Hantao Lou, Josef Dai, Yunhuai Liu, Yaodong Yang 0001
ACL (1)8
2024 Safe RLHF: Safe Reinforcement Learning from Human Feedback
abstract
With the development of large language models (LLMs), striking a balance between the performance and safety of AI systems has never been more critical. However, the inherent tension between the objectives of helpfulness and harmlessness presents a significant challenge during LLM training. To address this issue, we propose Safe Reinforcement Learning from Human Feedback (Safe RLHF), a novel algorithm for human value alignment. Safe RLHF explicitly decouples human preferences regarding helpfulness and harmlessness, effectively avoiding the crowd workers' confusion about the tension and allowing us to train separate reward and cost models. We formalize the safety concern of LLMs as an optimization task of maximizing the reward function while satisfying specified cost constraints. Leveraging the Lagrangian method to solve this constrained problem, Safe RLHF dynamically adjusts the balance between the two objectives during fine-tuning. Through a three-round fine-tuning using Safe RLHF, we demonstrate a superior ability to mitigate harmful responses while enhancing model performance compared to existing value-aligned algorithms. Experimentally, we fine-tuned the Alpaca-7B using Safe RLHF and aligned it with collected human preferences, significantly improving its helpfulness and harmlessness according to human evaluations. Code is available at https://github.com/PKU-Alignment/safe-rlhf. Warning: This paper contains example data that may be offensive or harmful.
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang 0001, Yaodong Yang 0001
ICLR1
2023 BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset
abstract
In this paper, we introduce the BeaverTails dataset, aimed at fostering research on safety alignment in large language models (LLMs). This dataset uniquely separates annotations of helpfulness and harmlessness for question-answering pairs, thus offering distinct perspectives on these crucial attributes. In total, we have gathered safety meta-labels for 333,963 question-answer (QA) pairs and 361,903 pairs of expert comparison data for both the helpfulness and harmlessness metrics. We further showcase applications of BeaverTails in content moderation and reinforcement learning with human feedback (RLHF), emphasizing its potential for practical safety measures in LLMs. We believe this dataset provides vital resources for the community, contributing towards the safe development and deployment of LLMs. Our project page is available at the following URL: https://sites.google.com/view/pku-beavertails.
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang 0017, Ce Bian, Boyuan Chen 0008, Ruiyang Sun, Yizhou Wang 0001, Yaodong Yang 0001
NeurIPS3
2023 Safety Gymnasium: A Unified Safe Reinforcement Learning Benchmark
abstract
Artificial intelligence (AI) systems possess significant potential to drive societal progress. However, their deployment often faces obstacles due to substantial safety concerns. Safe reinforcement learning (SafeRL) emerges as a solution to optimize policies while simultaneously adhering to multiple constraints, thereby addressing the challenge of integrating reinforcement learning in safety-critical scenarios. In this paper, we present an environment suite called Safety-Gymnasium, which encompasses safety-critical tasks in both single and multi-agent scenarios, accepting vector and vision-only input. Additionally, we offer a library of algorithms named Safe Policy Optimization (SafePO), comprising 16 state-of-the-art SafeRL algorithms. This comprehensive library can serve as a validation tool for the research community. By introducing this benchmark, we aim to facilitate the evaluation and comparison of safety performance, thus fostering the development of reinforcement learning for safer, more reliable, and responsible real-world applications. The website of this project can be accessed at https://sites.google.com/view/safety-gymnasium.
Jiaming Ji, Borong Zhang, Xuehai Pan, Weidong Huang 0008, Ruiyang Sun, Yiran Geng, Yifan Zhong, Josef Dai, Yaodong Yang 0001
NeurIPS9