EDBT 2026 Demo / reviewers in the wild / expert
Xuyao Wang
dblp:380/5991
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 56% Vision and language · 15% Generative modeling · 13% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › alignment
preference alignment |
1.6 | 2 | 2025 | InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback · NeurIPS 2025 SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset · NeurIPS 2024 |
Natural language and speech › Language models and text generation › alignment › preference alignment
human feedback alignment |
0.9 | 1 | 2025 | InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback · NeurIPS 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › AI safety
safety alignment |
0.8 | 1 | 2024 | SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset · NeurIPS 2024 |
Machine learning › Generative modeling › video generation
text-to-video generation |
0.8 | 1 | 2024 | SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › foundation model
large vision model |
0.2 | 1 | 2024 | SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
tool-augmented MLLM · 0.9preference annotation · 0.9agentic workflow · 0.9prompt augmentation · 0.8human preference annotation · 0.8fine-tuning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | InterMT: Multi-Turn Interleaved Preference Alignment with Human FeedbackabstractAs multimodal large models (MLLMs) continue to advance across challenging tasks, a key question emerges: \textbf{\textit{What essential capabilities are still missing? }}A critical aspect of human learning is continuous interaction with the environment -- not limited to language, but also involving multimodal understanding and generation.To move closer to human-level intelligence, models must similarly support \textbf{multi-turn}, \textbf{multimodal interaction}. In particular, they should comprehend interleaved multimodal contexts and respond coherently in ongoing exchanges.In this work, we present \textbf{an initial exploration} through the \textsc{InterMT} -- \textbf{the first preference dataset for \textit{multi-turn} multimodal interaction}, grounded in real human feedback. In this exploration, we particularly emphasize the importance of human oversight, introducing expert annotations to guide the process, motivated by the fact that current MLLMs lack such complex interactive capabilities. \textsc{InterMT} captures human preferences at both global and local levels into nine sub-dimensions, consists of 15.6k prompts, 52.6k multi-turn dialogue instances, and 32.4k human-labeled preference pairs. To compensate for the lack of capability for multi-modal understanding and generation, we introduce an agentic workflow that leverages tool-augmented MLLMs to construct multi-turn QA instances.To further this goal, we introduce \textsc{InterMT-Bench} to assess the ability ofMLLMs in assisting judges with multi-turn, multimodal tasks.We demonstrate the utility of \textsc{InterMT} through applications such as judge moderation and further reveal the \textit{multi-turn scaling law} of judge model.We hope the open-source of our data can help facilitate further research on aligning current MLLMs to the next step. Boyuan Chen 0008, Donghai Hong, Jiaming Ji, Jiacheng Zheng, Kaile Wang, Juntao Dai, Xuyao Wang, Sirui Han, Yike Guo, Yaodong Yang 0001 |
NeurIPS | 9 |
| 2024 | SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference DatasetabstractTo mitigate the risk of harmful outputs from large vision models (LVMs), we introduce the SafeSora dataset to promote research on aligning text-to-video generation with human values. This dataset encompasses human preferences in text-to-video generation tasks along two primary dimensions: helpfulness and harmlessness. To capture in-depth human preferences and facilitate structured reasoning by crowdworkers, we subdivide helpfulness into 4 sub-dimensions and harmlessness into 12 sub-categories, serving as the basis for pilot annotations. The SafeSora dataset includes 14,711 unique prompts, 57,333 unique videos generated by 4 distinct LVMs, and 51,691 pairs of preference annotations labeled by humans. We further demonstrate the utility of the SafeSora dataset through several applications, including training the text-video moderation model and aligning LVMs with human preference by fine-tuning a prompt augmentation module or the diffusion model. These applications highlight its potential as the foundation for text-to-video alignment research, such as human preference modeling and the development and validation of alignment algorithms. Our project is available at https://sites.google.com/view/safe-sora.Warning: this paper contains example data that may be offensive or harmful. Juntao Dai, Tianle Chen 0001, Xuyao Wang, Ziran Yang, Taiye Chen, Jiaming Ji, Yaodong Yang 0001 |
NeurIPS | 3 |