Boyuan Chen 0008

dblp:193/7174-8 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 48% Reinforcement learning · 23% Trustworthy machine learning · 20%
Theoretical computer science
1 paper
Coding theory · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
3.442025
Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025
Language Models Resist Alignment: Evidence From Data Compression · ACL (1) 2025
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference · ACL (1) 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
2.432025
Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025
Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback · NeurIPS 2025
BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset · NeurIPS 2023
Machine learning › Trustworthy machine learning › AI safety
safety alignment
2.432025
Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback · NeurIPS 2025
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference · ACL (1) 2025
BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset · NeurIPS 2023
Natural language and speech › Language models and text generation › alignment
preference alignment
1.722025
InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback · NeurIPS 2025
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference · ACL (1) 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
1.432025
InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback · NeurIPS 2025
Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025
Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback · NeurIPS 2025
Machine learning › Reinforcement learning › reward learning › reward modeling
generative reward model
0.912025
Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025
Natural language and speech › Language models and text generation › alignment › preference alignment
human feedback alignment
0.912025
InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback · NeurIPS 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025
Natural language and speech › Language models and text generation
hallucination mitigation
0.812024
Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model
0.812024
Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024
Natural language and speech › Language models and text generation
multimodal language model
0.312026
SafeMT: Multi-turn Safety for Multimodal Language Models · ACL (1) 2026
Machine learning › Trustworthy machine learning
robustness
0.312025
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference · ACL (1) 2025
Coding theory
source coding
0.312025
Language Models Resist Alignment: Evidence From Data Compression · ACL (1) 2025
Machine learning › Transfer learning and domain adaptation › foundation model adaptation
model-agnostic adaptation
0.212024
Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model safety
0.212023
BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

data compression analysis · 1.7safety alignment · 1.0adversarial training · 1.0tool-augmented MLLM · 0.9reinforcement learning from human feedback · 0.9preference annotation · 0.9guardrail filtering · 0.9constrained optimization · 0.9agentic workflow · 0.9RLHF · 0.9
YearPublicationVenuePosition
2026 SafeMT: Multi-turn Safety for Multimodal Language Models
abstract
Han Zhu, Juntao Dai, Jiaming Ji, Haoran Li, Chengkun Cai, Pengcheng Wen, Chi-Min Chan, Boyuan Chen, Yaodong Yang, Sirui Han, Yike Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Juntao Dai, Jiaming Ji, Chengkun Cai, Pengcheng Wen, Chi-Min Chan, Boyuan Chen 0008, Yaodong Yang 0001, Sirui Han, Yike Guo
ACL (1)8
2025 PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
abstract
Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen, Josef Dai, Boren Zheng, Tianyi Alex Qiu, Jiayi Zhou, Kaile Wang, Boxun Li, Sirui Han, Yike Guo, Yaodong Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen 0008, Josef Dai, Boren Zheng, Tianyi Qiu, Kaile Wang, Boxun Li, Sirui Han, Yike Guo, Yaodong Yang 0001
ACL (1)4
2025 Language Models Resist Alignment: Evidence From Data Compression
abstract
Jiaming Ji, Kaile Wang, Tianyi Alex Qiu, Boyuan Chen, Jiayi Zhou, Changye Li, Hantao Lou, Josef Dai, Yunhuai Liu, Yaodong Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jiaming Ji, Kaile Wang, Tianyi Qiu, Boyuan Chen 0008, Changye Li 0003, Hantao Lou, Josef Dai, Yunhuai Liu, Yaodong Yang 0001
ACL (1)4
2025 InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
abstract
As multimodal large models (MLLMs) continue to advance across challenging tasks, a key question emerges: \textbf{\textit{What essential capabilities are still missing? }}A critical aspect of human learning is continuous interaction with the environment -- not limited to language, but also involving multimodal understanding and generation.To move closer to human-level intelligence, models must similarly support \textbf{multi-turn}, \textbf{multimodal interaction}. In particular, they should comprehend interleaved multimodal contexts and respond coherently in ongoing exchanges.In this work, we present \textbf{an initial exploration} through the \textsc{InterMT} -- \textbf{the first preference dataset for \textit{multi-turn} multimodal interaction}, grounded in real human feedback. In this exploration, we particularly emphasize the importance of human oversight, introducing expert annotations to guide the process, motivated by the fact that current MLLMs lack such complex interactive capabilities. \textsc{InterMT} captures human preferences at both global and local levels into nine sub-dimensions, consists of 15.6k prompts, 52.6k multi-turn dialogue instances, and 32.4k human-labeled preference pairs. To compensate for the lack of capability for multi-modal understanding and generation, we introduce an agentic workflow that leverages tool-augmented MLLMs to construct multi-turn QA instances.To further this goal, we introduce \textsc{InterMT-Bench} to assess the ability ofMLLMs in assisting judges with multi-turn, multimodal tasks.We demonstrate the utility of \textsc{InterMT} through applications such as judge moderation and further reveal the \textit{multi-turn scaling law} of judge model.We hope the open-source of our data can help facilitate further research on aligning current MLLMs to the next step.
Boyuan Chen 0008, Donghai Hong, Jiaming Ji, Jiacheng Zheng, Kaile Wang, Juntao Dai, Xuyao Wang, Sirui Han, Yike Guo, Yaodong Yang 0001
NeurIPS1
2025 Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback
abstract
Multimodal large language models (MLLMs) are essential for building general-purpose AI assistants; however, they pose increasing safety risks. How can we ensure safety alignment of MLLMs to prevent undesired behaviors? Going further, it is critical to explore how to fine-tune MLLMs to preserve capabilities while meeting safety constraints. Fundamentally, this challenge can be formulated as a min-max optimization problem. However, existing datasets have not yet disentangled single preference signals into explicit safety constraints, hindering systematic investigation in this direction. Moreover, it remains an open question whether such constraints can be effectively incorporated into the optimization process for multi-modal models. In this work, we present the first exploration of the Safe RLHF-V -- the first multimodal safety alignment framework. The framework consists of: (I) BeaverTails-V, the first open-source dataset featuring dual preference annotations for helpfulness and safety, supplemented with multi-level safety labels (minor, moderate, severe); (II) Beaver-Guard-V, a multi-level guardrail system to proactively defend against unsafe queries and adversarial attacks. Applying the guard model over five rounds of filtering and regeneration significantly enhances the precursor model’s overall safety by an average of 40.9%. (II) Based on dual preference, we initiate the first exploration of multi-modal safety alignment within a constrained optimization. Experimental results demonstrate that Safe RLHF effectively improves both model helpfulness and safety. Specifically, Safe RLHF-V enhances model safety by 34.2% and helpfulness by 34.3%.
Jiaming Ji, Donghai Hong, Boyuan Chen 0008, Kaile Wang, Juntao Dai, Chi-Min Chan, Sirui Han, Yike Guo, Yaodong Yang 0001
NeurIPS7
2025 Generative RLHF-V: Learning Principles from Multi-modal Human Preference
abstract
Training multi-modal large language models (MLLMs) that align with human intentions is a long-term challenge. Traditional score-only reward models for alignment suffer from low accuracy, weak generalization, and poor interpretability, blocking the progress of alignment methods, \textit{e.g.,} reinforcement learning from human feedback (RLHF). Generative reward models (GRMs) leverage MLLMs' intrinsic reasoning capabilities to discriminate pair-wise responses, but their pair-wise paradigm makes it hard to generalize to learnable rewards. We introduce Generative RLHF-V, a novel alignment framework that integrates GRMs with multi-modal RLHF. We propose a two-stage pipeline: \textbf{multi-modal generative reward modeling from RL}, where RL guides GRMs to actively capture human intention, then predict the correct pair-wise scores; and \textbf{RL optimization from grouped comparison}, which enhances multi-modal RL scoring precision by grouped responses comparison. Experimental results demonstrate that, besides out-of-distribution generalization of RM discrimination, our framework improves 4 MLLMs' performance across 7 benchmarks by 18.1\%, while the baseline RLHF is only 5.3\%. We further validate that Generative RLHF-V achieves a near-linear improvement with an increasing number of candidate responses.
Jiaming Ji, Boyuan Chen 0008, Jiapeng Sun, Donghai Hong, Sirui Han, Yike Guo, Yaodong Yang 0001
NeurIPS3
2024 Aligner: Efficient Alignment by Learning to Correct
abstract
With the rapid development of large language models (LLMs) and ever-evolving practical requirements, finding an efficient and effective alignment method has never been more critical. However, the tension between the complexity of current alignment methods and the need for rapid iteration in deployment scenarios necessitates the development of a model-agnostic alignment approach that can operate under these constraints. In this paper, we introduce Aligner, a novel and simple alignment paradigm that learns the correctional residuals between preferred and dispreferred answers using a small model. Designed as a model-agnostic, plug-and-play module, Aligner can be directly applied to various open-source and API-based models with only one-off training, making it suitable for rapid iteration. Notably, Aligner can be applied to any powerful, large-scale upstream models. Moreover, it can even iteratively bootstrap the upstream models using corrected responses as synthetic human preference data, breaking through the model's performance ceiling. Our experiments demonstrate performance improvements by deploying the same Aligner model across 11 different LLMs, evaluated on the 3H dimensions (helpfulness, harmlessness, and honesty). Specifically, Aligner-7B has achieved an average improvement of 68.9% in helpfulness and 22.8% in harmlessness across the tested LLMs while also effectively reducing hallucination. In the Alpaca-Eval leaderboard, stacking Aligner-2B on GPT-4 Turbo improved its LC Win Rate from 55.0% to 58.3%, surpassing GPT-4 Omni's 57.5% Win Rate (community report).
Jiaming Ji, Boyuan Chen 0008, Hantao Lou, Donghai Hong, Borong Zhang, Xuehai Pan, Tianyi Qiu, Juntao Dai, Yaodong Yang 0001
NeurIPS2
2023 BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset
abstract
In this paper, we introduce the BeaverTails dataset, aimed at fostering research on safety alignment in large language models (LLMs). This dataset uniquely separates annotations of helpfulness and harmlessness for question-answering pairs, thus offering distinct perspectives on these crucial attributes. In total, we have gathered safety meta-labels for 333,963 question-answer (QA) pairs and 361,903 pairs of expert comparison data for both the helpfulness and harmlessness metrics. We further showcase applications of BeaverTails in content moderation and reinforcement learning with human feedback (RLHF), emphasizing its potential for practical safety measures in LLMs. We believe this dataset provides vital resources for the community, contributing towards the safe development and deployment of LLMs. Our project page is available at the following URL: https://sites.google.com/view/pku-beavertails.
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang 0017, Ce Bian, Boyuan Chen 0008, Ruiyang Sun, Yizhou Wang 0001, Yaodong Yang 0001
NeurIPS7