EDBT 2026 Demo / reviewers in the wild / expert
Donghai Hong
dblp:367/7553
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 51% Reinforcement learning · 24% Trustworthy machine learning · 14% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
2.5 | 3 | 2025 | Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025 PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference · ACL (1) 2025 Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024 |
Natural language and speech › Language models and text generation › alignment
preference alignment |
1.7 | 2 | 2025 | InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback · NeurIPS 2025 PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference · ACL (1) 2025 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
1.7 | 2 | 2025 | Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025 Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › AI safety
safety alignment |
1.7 | 2 | 2025 | Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback · NeurIPS 2025 PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference · ACL (1) 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.4 | 3 | 2025 | InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback · NeurIPS 2025 Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025 Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback · NeurIPS 2025 |
Machine learning › Reinforcement learning › reward learning › reward modeling
generative reward model |
0.9 | 1 | 2025 | Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025 |
Natural language and speech › Language models and text generation › alignment › preference alignment
human feedback alignment |
0.9 | 1 | 2025 | InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback · NeurIPS 2025 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.9 | 1 | 2025 | Generative RLHF-V: Learning Principles from Multi-modal Human Preference · NeurIPS 2025 |
Natural language and speech › Language models and text generation
hallucination mitigation |
0.8 | 1 | 2024 | Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024 |
Natural language and speech › Language models and text generation
large language model |
0.8 | 1 | 2024 | Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
robustness |
0.3 | 1 | 2025 | PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference · ACL (1) 2025 |
Machine learning › Transfer learning and domain adaptation › foundation model adaptation
model-agnostic adaptation |
0.2 | 1 | 2024 | Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
tool-augmented MLLM · 0.9reinforcement learning from human feedback · 0.9reinforcement learning · 0.9preference annotation · 0.9guardrail filtering · 0.9grouped comparison · 0.9generative reward modeling · 0.9constrained optimization · 0.9agentic workflow · 0.9RLHF · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human PreferenceabstractJiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen, Josef Dai, Boren Zheng, Tianyi Alex Qiu, Jiayi Zhou, Kaile Wang, Boxun Li, Sirui Han, Yike Guo, Yaodong Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen 0008, Josef Dai, Boren Zheng, Tianyi Qiu, Kaile Wang, Boxun Li, Sirui Han, Yike Guo, Yaodong Yang 0001 |
ACL (1) | 2 |
| 2025 | InterMT: Multi-Turn Interleaved Preference Alignment with Human FeedbackabstractAs multimodal large models (MLLMs) continue to advance across challenging tasks, a key question emerges: \textbf{\textit{What essential capabilities are still missing? }}A critical aspect of human learning is continuous interaction with the environment -- not limited to language, but also involving multimodal understanding and generation.To move closer to human-level intelligence, models must similarly support \textbf{multi-turn}, \textbf{multimodal interaction}. In particular, they should comprehend interleaved multimodal contexts and respond coherently in ongoing exchanges.In this work, we present \textbf{an initial exploration} through the \textsc{InterMT} -- \textbf{the first preference dataset for \textit{multi-turn} multimodal interaction}, grounded in real human feedback. In this exploration, we particularly emphasize the importance of human oversight, introducing expert annotations to guide the process, motivated by the fact that current MLLMs lack such complex interactive capabilities. \textsc{InterMT} captures human preferences at both global and local levels into nine sub-dimensions, consists of 15.6k prompts, 52.6k multi-turn dialogue instances, and 32.4k human-labeled preference pairs. To compensate for the lack of capability for multi-modal understanding and generation, we introduce an agentic workflow that leverages tool-augmented MLLMs to construct multi-turn QA instances.To further this goal, we introduce \textsc{InterMT-Bench} to assess the ability ofMLLMs in assisting judges with multi-turn, multimodal tasks.We demonstrate the utility of \textsc{InterMT} through applications such as judge moderation and further reveal the \textit{multi-turn scaling law} of judge model.We hope the open-source of our data can help facilitate further research on aligning current MLLMs to the next step. Boyuan Chen 0008, Donghai Hong, Jiaming Ji, Jiacheng Zheng, Kaile Wang, Juntao Dai, Xuyao Wang, Sirui Han, Yike Guo, Yaodong Yang 0001 |
NeurIPS | 2 |
| 2025 | Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human FeedbackabstractMultimodal large language models (MLLMs) are essential for building general-purpose AI assistants; however, they pose increasing safety risks. How can we ensure safety alignment of MLLMs to prevent undesired behaviors? Going further, it is critical to explore how to fine-tune MLLMs to preserve capabilities while meeting safety constraints. Fundamentally, this challenge can be formulated as a min-max optimization problem. However, existing datasets have not yet disentangled single preference signals into explicit safety constraints, hindering systematic investigation in this direction. Moreover, it remains an open question whether such constraints can be effectively incorporated into the optimization process for multi-modal models. In this work, we present the first exploration of the Safe RLHF-V -- the first multimodal safety alignment framework. The framework consists of: (I) BeaverTails-V, the first open-source dataset featuring dual preference annotations for helpfulness and safety, supplemented with multi-level safety labels (minor, moderate, severe); (II) Beaver-Guard-V, a multi-level guardrail system to proactively defend against unsafe queries and adversarial attacks. Applying the guard model over five rounds of filtering and regeneration significantly enhances the precursor model’s overall safety by an average of 40.9%. (II) Based on dual preference, we initiate the first exploration of multi-modal safety alignment within a constrained optimization. Experimental results demonstrate that Safe RLHF effectively improves both model helpfulness and safety. Specifically, Safe RLHF-V enhances model safety by 34.2% and helpfulness by 34.3%. Jiaming Ji, Donghai Hong, Boyuan Chen 0008, Kaile Wang, Juntao Dai, Chi-Min Chan, Sirui Han, Yike Guo, Yaodong Yang 0001 |
NeurIPS | 6 |
| 2025 | Generative RLHF-V: Learning Principles from Multi-modal Human PreferenceabstractTraining multi-modal large language models (MLLMs) that align with human intentions is a long-term challenge. Traditional score-only reward models for alignment suffer from low accuracy, weak generalization, and poor interpretability, blocking the progress of alignment methods, \textit{e.g.,} reinforcement learning from human feedback (RLHF). Generative reward models (GRMs) leverage MLLMs' intrinsic reasoning capabilities to discriminate pair-wise responses, but their pair-wise paradigm makes it hard to generalize to learnable rewards. We introduce Generative RLHF-V, a novel alignment framework that integrates GRMs with multi-modal RLHF. We propose a two-stage pipeline: \textbf{multi-modal generative reward modeling from RL}, where RL guides GRMs to actively capture human intention, then predict the correct pair-wise scores; and \textbf{RL optimization from grouped comparison}, which enhances multi-modal RL scoring precision by grouped responses comparison. Experimental results demonstrate that, besides out-of-distribution generalization of RM discrimination, our framework improves 4 MLLMs' performance across 7 benchmarks by 18.1\%, while the baseline RLHF is only 5.3\%. We further validate that Generative RLHF-V achieves a near-linear improvement with an increasing number of candidate responses. Jiaming Ji, Boyuan Chen 0008, Jiapeng Sun, Donghai Hong, Sirui Han, Yike Guo, Yaodong Yang 0001 |
NeurIPS | 6 |
| 2024 | Aligner: Efficient Alignment by Learning to CorrectabstractWith the rapid development of large language models (LLMs) and ever-evolving practical requirements, finding an efficient and effective alignment method has never been more critical. However, the tension between the complexity of current alignment methods and the need for rapid iteration in deployment scenarios necessitates the development of a model-agnostic alignment approach that can operate under these constraints. In this paper, we introduce Aligner, a novel and simple alignment paradigm that learns the correctional residuals between preferred and dispreferred answers using a small model. Designed as a model-agnostic, plug-and-play module, Aligner can be directly applied to various open-source and API-based models with only one-off training, making it suitable for rapid iteration. Notably, Aligner can be applied to any powerful, large-scale upstream models. Moreover, it can even iteratively bootstrap the upstream models using corrected responses as synthetic human preference data, breaking through the model's performance ceiling. Our experiments demonstrate performance improvements by deploying the same Aligner model across 11 different LLMs, evaluated on the 3H dimensions (helpfulness, harmlessness, and honesty). Specifically, Aligner-7B has achieved an average improvement of 68.9% in helpfulness and 22.8% in harmlessness across the tested LLMs while also effectively reducing hallucination. In the Alpaca-Eval leaderboard, stacking Aligner-2B on GPT-4 Turbo improved its LC Win Rate from 55.0% to 58.3%, surpassing GPT-4 Omni's 57.5% Win Rate (community report). Jiaming Ji, Boyuan Chen 0008, Hantao Lou, Donghai Hong, Borong Zhang, Xuehai Pan, Tianyi Qiu, Juntao Dai, Yaodong Yang 0001 |
NeurIPS | 4 |